AWS Batch is a job queue bolted to a capacity manager. This is the definition, the cost model, the limits, and a 20-way array job you can run on Fargate for about a cent.
Most people meet AWS Batch the wrong way. They have a pile of work, a script that takes four hours on one machine, and a vague feeling that "there must be a service for this." There is. But Batch is not a workflow engine and it is not a compute service. It is the boring middle layer between the two, and understanding that middle layer is the whole article.
What it actually is
AWS Batch accepts jobs, holds them in a queue, and decides when and where each one runs. When there is work and not enough capacity, it creates capacity. When the work is gone, it removes the capacity. The containers themselves run on Amazon ECS, Amazon EKS, or Fargate, and the AWS documentation lists the capacity options as EC2 (On-Demand or Spot), Fargate, Amazon ECS Managed Instances, and Amazon EKS.
The mental model that stops the confusion: Batch owns the queue and the scaling decision. You own the container. AWS Batch itself has no charge; per the pricing page you pay only for the underlying compute, such as EC2 or Fargate, with Reserved Instances, Savings Plans, and Spot discounts applied at billing time.
Two recent changes shift the picture. On August 25, 2026, AWS announced Batch support for Amazon ECS Managed Instances, where AWS handles AMI updates, patching, and instance lifecycle instead of you. And in September 2026 Batch added bulk job management: the CancelJobs, TerminateJobs and TerminateServiceJobs APIs each take up to 50 job IDs per call. Batch also now queues SageMaker Training jobs, which is why you will see "service environments" in the console. This article stays with the classic container path.
The model
Four objects matter, and the relationships between them are the entire design.
A job definition is the template: image, command, vCPU and memory, IAM roles, retry strategy, timeout. It is versioned, and every registration creates a new revision. A job is one submission of a job definition. A job queue holds submitted jobs and has an integer priority. A compute environment is the capacity pool a queue maps to, and a queue can map to up to three of them in order.
A job moves through seven states: SUBMITTED, PENDING (waiting on a dependency), RUNNABLE (ready, waiting for capacity), STARTING (image pull and container start), RUNNING, then SUCCEEDED or FAILED. The one that matters operationally is RUNNABLE. The docs say plainly that a job can sit there indefinitely when the mapped compute environments cannot supply the resources. Batch will not fail it for you unless you configure jobStateTimeLimitActions on the queue.
A compute environment is either managed or unmanaged. In a managed one you give Batch a maxvCpus ceiling, and for EC2 types a minvCpus floor, and Batch adjusts desired vCPUs between them based on queue demand. The resource type is one of EC2, SPOT, FARGATE, FARGATE_SPOT, or ECS_MANAGED_INSTANCES. For EC2 you pick an allocation strategy; the documented values are BEST_FIT (the default), BEST_FIT_PROGRESSIVE, and BEST_FIT_PROGRESSIVE_ORDERED for On-Demand, and SPOT_CAPACITY_OPTIMIZED, SPOT_PRICE_CAPACITY_OPTIMIZED, and SPOT_CAPACITY_OPTIMIZED_PRIORITIZED for Spot. Allocation strategy does not apply to Fargate. Instance types can be a specific type, a family, or the bundles default_x86_64, default_arm64, and optimal, which Batch updates over time to newer instance families. The default AMI type for EC2 environments is ECS_AL2023.
The feature that turns Batch from a queue into a fan-out tool is the array job. You submit once with an array size between 2 and 10,000, and Batch creates that many child jobs sharing one definition. Each child gets an AWS_BATCH_JOB_ARRAY_INDEX environment variable starting at 0. That index is how a child knows which shard of the data is its own. Children can be chained with a SEQUENTIAL dependency (one at a time, in index order) or N_TO_N (child i of job B waits only for child i of job A). The timeout applies to each child individually, and the parent turns FAILED only after all children finish if any of them failed.
When to use it, when not to
The honest comparison is against the three services people reach for instead.
You need | Use | Why |
|---|---|---|
Thousands of independent container tasks, queued, retried, scaled to zero | AWS Batch | Queue plus capacity management is exactly the product |
Run a container behind a load balancer, always on | ECS service or Fargate service | Batch jobs are finite by design |
Run something under 15 minutes per task, event-driven | Lambda | No queue to manage, no cold container pull |
A multi-step workflow with branching and human approval | Step Functions (calling Batch for the heavy steps) | Batch has dependencies, not orchestration logic |
Hundreds of thousands of messages processed continuously | SQS plus workers | Batch is job-granular, not message-granular |
Pick Batch when each unit of work is long enough that container startup is noise (think tens of seconds and up), you want retries and priorities without writing a scheduler, and the volume is spiky. Skip it when you need sub-second latency to start work. A job going from SUBMITTED to RUNNING takes real time: scheduling, capacity, image pull.
What it costs
Batch is free. Everything underneath is not, and the dimension that surprises people is idle capacity on the EC2 path. If you set minvCpus above zero, those instances run and bill whether or not a job exists. Keep it at zero unless startup latency is worth the rent.
On Fargate there is no idle capacity to leak, but there is a billing floor. Fargate Linux pricing in US East (N. Virginia), as of this writing, is $0.000011244 per vCPU-second and $0.000001235 per GB-second for x86, and $0.0000089944 and $0.0000009889 for ARM. Billing is per second with a 1-minute minimum, and every task gets 20 GB of ephemeral storage free. Fargate Spot is advertised at up to 70% off. So a 0.5 vCPU, 1 GB x86 task costs roughly $0.0247 per hour, and a job that finishes in 30 seconds is billed for 60. Twenty of them is about eight tenths of a cent of compute. Check the Fargate pricing page for your region before you scale this up; the rates move.
The free-tier reality: there is no Batch free tier because there is no Batch charge. CloudWatch Logs ingestion and any NAT gateway you add are the other line items. This tutorial deliberately avoids the NAT gateway by using public IPs on the tasks.
The limits that bite
The documented quotas are per Region, and the service-quotas table marks every one of these as not adjustable.
Quota | Value |
|---|---|
Compute environments | 50 |
Compute environments per job queue | 3 |
Job queues | 50 |
Array size | 10,000 |
Job dependencies | 20 |
Job payload size | 30 KiB |
Job definition size | 24 KiB |
Jobs in SUBMITTED state | 1,000,000 |
SubmitJob transactions per second | 50 |
Share identifiers per job queue | 500 |
Three of these deserve a second look. The 50 TPS on SubmitJob means a loop that submits 100,000 individual jobs will throttle; an array job is one call and is the reason the feature exists. The 20-dependency cap means wide fan-in needs an intermediate job. And the 30 KiB payload limit applies to the whole submission, so do not pass a dataset in parameters or container overrides; pass an S3 key and let the child fetch it.
The failure modes that show up at scale are not quota errors. They are maxvCpus as a silent throttle (jobs wait in RUNNABLE while the environment sits at its ceiling), subnets that run out of IPs, and Fargate tasks that cannot pull their image because the subnet has no route to the registry.
Build it
We will build a 20-way array job on Fargate. Each child reads its array index, counts primes in its own slice of a number range, and prints the result to CloudWatch Logs. Then we collect the output, verify it, and tear everything down.
Prerequisites
You need an AWS account, Python 3.9 or later, boto3, and credentials with a default VPC in your Region (or a subnet and security group you substitute below). The caller needs permissions along these lines: batch:* on the resources you create, iam:CreateRole, iam:AttachRolePolicy, iam:PassRole on the execution role, iam:GetRole, ec2:DescribeVpcs, ec2:DescribeSubnets, ec2:DescribeSecurityGroups, logs:GetLogEvents, and logs:DeleteLogGroup. The AWS managed policy AWSBatchFullAccess covers the Batch part; add the IAM and EC2 describe actions for the rest. In a throwaway account, an admin role is the pragmatic choice. Batch also relies on its service-linked role AWSServiceRoleForBatch. If compute environment creation complains that it does not exist, create it once with aws iam create-service-linked-role --aws-service-name batch.amazonaws.com.
Fargate jobs also need a task execution role, which lets the ECS agent pull the image and write logs. The getting-started guide names it BatchEcsTaskExecutionRole with the AmazonECSTaskExecutionRolePolicy attached, and the script below creates it.
The script
import json, time, boto3
REGION = "us-east-1"
NAME = "sl108-batch"
ARRAY_SIZE = 20
batch = boto3.client("batch", region_name=REGION)
ec2 = boto3.client("ec2", region_name=REGION)
iam = boto3.client("iam", region_name=REGION)
# 1. Network: default VPC, its subnets, its default security group
vpc = ec2.describe_vpcs(Filters=[{"Name": "isDefault", "Values": ["true"]}])["Vpcs"][0]["VpcId"]
subnets = [s["SubnetId"] for s in ec2.describe_subnets(
Filters=[{"Name": "vpc-id", "Values": [vpc]}])["Subnets"]]
sg = ec2.describe_security_groups(Filters=[
{"Name": "vpc-id", "Values": [vpc]},
{"Name": "group-name", "Values": ["default"]}])["SecurityGroups"][0]["GroupId"]
# 2. Task execution role (image pull + log writes)
trust = {"Version": "2012-10-17", "Statement": [{
"Effect": "Allow",
"Principal": {"Service": "ecs-tasks.amazonaws.com"},
"Action": "sts:AssumeRole"}]}
role_name = "BatchEcsTaskExecutionRole"
try:
iam.create_role(RoleName=role_name, AssumeRolePolicyDocument=json.dumps(trust))
iam.attach_role_policy(RoleName=role_name,
PolicyArn="arn:aws:iam::aws:policy/service-role/AmazonECSTaskExecutionRolePolicy")
time.sleep(10) # IAM propagation
except iam.exceptions.EntityAlreadyExistsException:
pass
role_arn = iam.get_role(RoleName=role_name)["Role"]["Arn"]
# 3. Managed Fargate compute environment, scales to zero by design
batch.create_compute_environment(
computeEnvironmentName=f"{NAME}-ce",
type="MANAGED",
state="ENABLED",
computeResources={
"type": "FARGATE",
"maxvCpus": 16,
"subnets": subnets,
"securityGroupIds": [sg],
},
)
while True:
ce = batch.describe_compute_environments(
computeEnvironments=[f"{NAME}-ce"])["computeEnvironments"][0]
if ce["status"] == "VALID":
break
assert ce["status"] != "INVALID", ce.get("statusReason")
time.sleep(3)
# 4. Job queue mapped to that environment
batch.create_job_queue(
jobQueueName=f"{NAME}-q",
state="ENABLED",
priority=900,
computeEnvironmentOrder=[{"order": 1, "computeEnvironment": f"{NAME}-ce"}],
)
while batch.describe_job_queues(jobQueues=[f"{NAME}-q"])["jobQueues"][0]["status"] != "VALID":
time.sleep(3)
# 5. Job definition: each child counts primes in its own slice
worker = (
"import os;"
"i=int(os.environ['AWS_BATCH_JOB_ARRAY_INDEX']);"
"lo,hi=i*50000,(i+1)*50000;"
"n=sum(1 for k in range(max(lo,2),hi) if all(k%d for d in range(2,int(k**0.5)+1)));"
"print(f'shard={i} range=[{lo},{hi}) primes={n}')"
)
jd = batch.register_job_definition(
jobDefinitionName=f"{NAME}-jd",
type="container",
platformCapabilities=["FARGATE"],
containerProperties={
"image": "public.ecr.aws/docker/library/python:3.12-slim",
"command": ["python", "-c", worker],
"executionRoleArn": role_arn,
"resourceRequirements": [
{"type": "VCPU", "value": "0.5"},
{"type": "MEMORY", "value": "1024"},
],
"networkConfiguration": {"assignPublicIp": "ENABLED"},
"fargatePlatformConfiguration": {"platformVersion": "LATEST"},
},
retryStrategy={
"attempts": 3,
"evaluateOnExit": [
{"action": "RETRY", "onStatusReason": "Task failed to start"},
{"action": "EXIT", "onReason": "*"},
],
},
timeout={"attemptDurationSeconds": 300},
)
# 6. Submit ONE call that fans out to 20 children
job = batch.submit_job(
jobName=f"{NAME}-run",
jobQueue=f"{NAME}-q",
jobDefinition=jd["jobDefinitionArn"],
arrayProperties={"size": ARRAY_SIZE},
)
print("parent job:", job["jobId"])A few choices in that script are deliberate. assignPublicIp is ENABLED because the Fargate getting-started guide requires it to pull the image unless you use a private registry path; without it, a private subnet with no NAT strands the task. The retry block follows the documented shape: up to 10 attempts are allowed, evaluateOnExit takes up to 5 entries, and a final catch-all entry is recommended so unmatched failures behave the way you decided, not the way the default does (if evaluateOnExit is set and nothing matches, Batch retries). Jobs that are cancelled, terminated, or fail on an invalid definition are never retried.
I wrote the script against the current documentation and API reference; run it in a sandbox account first, because I cannot promise your VPC looks like mine.
Verify it
Poll the parent, then read each child's log line. The child job IDs are <parentId>:0 through <parentId>:19.
parent = job["jobId"]
while True:
j = batch.describe_jobs(jobs=[parent])["jobs"][0]
summ = j.get("arrayProperties", {}).get("statusSummary", {})
print(j["status"], summ)
if j["status"] in ("SUCCEEDED", "FAILED"):
break
time.sleep(10)
logs = boto3.client("logs", region_name=REGION)
total = 0
for i in range(ARRAY_SIZE):
child = batch.describe_jobs(jobs=[f"{parent}:{i}"])["jobs"][0]
stream = child["container"]["logStreamName"]
ev = logs.get_log_events(logGroupName="/aws/batch/job",
logStreamName=stream)["events"]
line = ev[-1]["message"]
print(line)
total += int(line.split("primes=")[1])
print("primes below", ARRAY_SIZE * 50000, "=", total)The numbers are checkable. There are 78,498 primes below 1,000,000, and 20 shards of 50,000 cover exactly 0 to 999,999, so your total must print 78498. If it does, you have verified the whole path: queue, scaling, container start, index injection, and logging. The first run takes a minute or two while Fargate capacity is provisioned and the image is pulled; the 20 children then run in parallel.
Estimated cost of following along: well under a cent of Fargate compute plus negligible log ingestion, assuming each child finishes inside its one-minute billing minimum and you clean up.
When it breaks
The error you will hit first is a job that never leaves RUNNABLE. Read statusReason from describe_jobs, or watch for the job-queue-blocked CloudWatch event, which carries the blocking reason. The documented root causes fall into three groups: capacity (maxvCpus reached, not enough resources), misconfiguration (the job asks for more than the environment can provide, or the definition does not match the environment), and an INVALID compute environment. If statusReason says UNDETERMINED, the docs point to a separate troubleshooting page. Prevent the silent hang by setting jobStateTimeLimitActions on the queue so a stuck job cancels itself.
The second is a Fargate job that fails at STARTING with a message about pulling the image. Almost always this is the network: no public IP and no NAT or endpoint route to the registry, or an execution role missing the managed policy. Note that STARTING time is not counted toward the job timeout, which begins at RUNNING.
Third, a compute environment stuck in INVALID. The statusReason on describe_compute_environments names the problem, which is usually a bad subnet, a missing security group, or a missing service-linked role. And remember that failed and succeeded job records persist for at least seven days, so you can dig into a failure after the fact.
If one array child fails, the others keep running and the parent only flips to FAILED when they all finish. Retry or resubmit just the failed index rather than the whole array.
Cleanup
Order matters, because you cannot delete a queue that is enabled or a compute environment that is still attached.
batch.update_job_queue(jobQueue=f"{NAME}-q", state="DISABLED")
while batch.describe_job_queues(jobQueues=[f"{NAME}-q"])["jobQueues"][0]["status"] != "VALID":
time.sleep(3)
batch.delete_job_queue(jobQueue=f"{NAME}-q")
while batch.describe_job_queues(jobQueues=[f"{NAME}-q"])["jobQueues"]:
time.sleep(3)
batch.update_compute_environment(computeEnvironment=f"{NAME}-ce", state="DISABLED")
time.sleep(60)
batch.delete_compute_environment(computeEnvironment=f"{NAME}-ce")
batch.deregister_job_definition(jobDefinition=jd["jobDefinitionArn"])
boto3.client("logs", region_name=REGION).delete_log_group(logGroupName="/aws/batch/job")The getting-started guide says disabling and deleting the compute environment each take a minute or two, so the fixed sleep above may need a retry loop in a slow Region. Delete the BatchEcsTaskExecutionRole role too if nothing else uses it (detach its policy first). Do not delete the /aws/batch/job log group if other Batch workloads in the account write to it. To confirm the bill is zero, check that describe_compute_environments returns nothing for your name; Fargate has no idle charge once the tasks are gone.

