Jobs stuck in RUNNABLE due to capacity
The following scenario describes how Amazon Batch identifies insufficient capacity conditions.
- Insufficient instance capacity
-
Amazon Batch interprets two sets of responses received by SageMaker Training jobs as cases of insufficient instance capacity. Whenever these are received, Amazon Batch sets a corresponding
statusReasonofCAPACITY:INSUFFICIENT_INSTANCE_CAPACITY.In both cases, the time the corresponding Amazon Batch job spent in
SCHEDULEDbetween job submission and failure accumulates between job submissions for comparison againstmaxTimeSeconds.Scenario
TrainingJobStatusIndicator
Training job fails with capacity error
FAILEDfailedReasonbegins withCapacityErrorTraining job times out waiting for capacity
STOPPEDsecondaryStatusindicates the job timed out in pending waiting for capacityIf either of these cases is hit, Amazon Batch will consider the job to be encountering insufficient capacity.
-
statusReasonmessage while the job is stuck:CAPACITY:INSUFFICIENT_INSTANCE_CAPACITY - Unable to provision requested compute resources -
reasonused forjobStateTimeLimitActions:CAPACITY:INSUFFICIENT_INSTANCE_CAPACITY -
statusReasonmessage after the job is terminated byjobStateTimeLimitActions:Terminated by JobStateTimeLimit action due to reason: CAPACITY:INSUFFICIENT_INSTANCE_CAPACITY
-