This website uses cookies

Read our Privacy policy and Terms of use for more information.

What it actually is

Amazon EC2 Auto Scaling is a controller for a set of EC2 instances called an Auto Scaling group (ASG). You give it a launch template, a list of subnets, and three numbers: minimum, desired, maximum. It then does one job forever: make the count of healthy, in-service instances equal the desired number, and replace anything that stops being healthy.

Everything people call "autoscaling" is a different component moving the desired number. Target tracking, step scaling, scheduled actions, predictive scaling and your own API calls all do the same thing: write a new desired capacity, clamped between min and max. The ASG itself never decides that load is high. It only converges on whatever number it was last given.

That framing explains most production surprises. An ASG with min=max=desired is a self-healing fleet and no scaling at all. An ASG whose policies fight each other is two writers racing on one integer.

The model

The primitives are small. A launch template says what an instance is (AMI, instance type, security groups, user data, instance profile). The group says where and how many (subnets across Availability Zones, min, desired, max). Health checks say what "alive" means. Policies say how the desired number moves.

Health comes from several sources: the built-in EC2 status checks, Elastic Load Balancing, VPC Lattice, EBS volume impairment, and custom checks you report yourself. Every instance starts as healthy and only InService instances are checked. The health check grace period delays the optional checks (ELB, custom) after launch, while the EC2 status checks run immediately. When an InService instance goes unhealthy the group terminates it and launches a replacement from the current template, and the same replacement happens for Spot interruptions and instances someone terminated by hand.

There are five ways to move the desired number. Manual changes. Scheduled actions for known events. Simple scaling, which uses a cooldown and which the docs themselves tell you to avoid. Step scaling, where you hand-write the alarm thresholds and step adjustments. And target tracking, where you name a metric and a target value, and Auto Scaling creates and manages the CloudWatch alarms for you, like a thermostat.

Target tracking supports four predefined metrics: ASGAverageCPUUtilization, ASGAverageNetworkIn, ASGAverageNetworkOut and ALBRequestCountPerTarget, plus custom metrics. The metric has to scale inversely with instance count, so average CPU works and a raw request count or queue depth does not. Instance warmup is the part people skip: instances still warming up are excluded from the aggregated metric, and scale-in is blocked until they finish. Without a warmup value the group falls back to the default instance warmup, then to the default cooldown of 300 seconds.

Two details shape behavior at the edges. With several target tracking policies, the group scales out if any of them wants to, and scales in only when all of them (with scale-in enabled) agree. And rounding is deliberately conservative: if the math says 1.5 instances, you get 2, and if it says you could remove 0.5 of one, nothing happens. Small groups will therefore sit visibly away from the target value.

What AWS operates: the controller, the alarms behind target tracking, the forecasting service. What you operate: the AMI, the bootstrapping, the choice of metric, and every limit that protects you from your own mistakes.

When to use it, when not to

You want

Use

Why

A fleet of identical EC2 instances that heals and tracks load

EC2 Auto Scaling

The native controller, no extra charge

Scale containers or Lambda concurrency

Application Auto Scaling

Same idea, different targets (ECS services, DynamoDB, Lambda and more)

Run containers without managing hosts

ECS or EKS on Fargate

The host layer disappears, so the ASG is not your problem

Pods in Kubernetes

Karpenter or the cluster autoscaler

They decide node counts from pending pods, and an ASG is only the backing store

One-off capacity for a batch job

AWS Batch or a fleet request

You want a queue, not a thermostat

The honest rule: if your unit of scale is an EC2 instance that you configure, you want an ASG. If your unit is a task, a pod, or a request, something above the instance should be doing the deciding. Launch configurations, the older way to define instances, are effectively retired. New instance types have not been supported in them since January 1, 2023, accounts created after October 1, 2024 cannot create them by any method, and the recommendation is launch templates.

What it costs

Auto Scaling itself is free. The pricing page says there are no additional fees beyond EC2, CloudWatch and the other resources you use. That includes predictive scaling and warm pools: no separate line item for either.

The bill is what the group does. Instance-hours, billed per second with a 60-second minimum for Linux. EBS volumes for every instance, including the stopped ones in a warm pool. CloudWatch charges for detailed (one-minute) monitoring and for any custom metrics. Load balancer hours if you attach one. Data transfer.

The dimension that surprises people is warm pools. A Stopped warm pool instance costs nothing for compute but you pay for its EBS volumes and Elastic IPs. A Hibernated one also keeps RAM contents on EBS. A Running pool bills the full instance price and the docs discourage it. The default pool size is max capacity minus desired capacity, so a group with max 100 and desired 10 can quietly hold 90 stopped instances worth of EBS. Set MaxGroupPreparedCapacity to cap it.

The free tier covers some small instances for new accounts, but treat that as an account-specific detail and check your Billing console. The build below uses t3.micro and runs well under a dollar if you clean up the same day. Verify current rates on the EC2 on-demand pricing page before you rely on the number.

The limits that bite

Hard limits that cannot be raised: 50 scaling policies per group, 125 scheduled actions per group, 20 step adjustments per step policy, 50 lifecycle hooks per group, 10 SNS topics per group, 50 Classic Load Balancers and 50 target groups per group, 5 VPC Lattice target groups per group. Adjustable: 500 groups per Region and 200 launch configurations per Region.

API batch limits matter in automation: 20 instance IDs per AttachInstances, DetachInstances or EnterStandby call, 50 for SetInstanceProtection. In August 2026 AWS raised the TerminateInstanceInAutoScalingGroup API to accept up to 100 instance IDs in one call, validated atomically, with lifecycle hooks and connection draining still honored.

The limits that hurt more are not in that table. Your EC2 vCPU quota caps a group long before the 500-group limit does, and a scale-out that hits it fails silently in the activity history. The default warm-pool sizing, described above, is a cost limit. Warm pools do not support weighted mixed-instance groups, and in mixed groups only On-Demand instances can be pooled. A depleted pool just means a cold start. And a metric with sparse data points leaves the alarm in INSUFFICIENT_DATA, where target tracking does nothing at all.

Predictive scaling has its own constraints. It needs at least 24 hours of history, recommends two weeks, and analyzes up to 14 days. It produces a 48-hour forecast, regenerated hourly and refreshed every 6 hours. In forecast-and-scale mode it only scales out. It never scales in, so you still need dynamic scaling to give capacity back. By default it will not exceed the group maximum unless you opt in with MaxCapacityBreachBehavior and MaxCapacityBuffer, and that increase does not decrease by itself.

For deletion, January 2026 added group-level deletion protection and an IAM condition key, autoscaling:ForceDelete, so you can stop a force delete of a group with running instances. Production groups should use both.

Build it

The deliverable: an ASG across your default VPC subnets with a CPU target tracking policy, a warm pool of Stopped instances, and a predictive scaling policy in forecast-only mode. We drive CPU with a flag in SSM Parameter Store so you can start and stop a load test from your laptop without SSH.

Prerequisites. An AWS account, Python 3.10+, boto3 and AWS credentials with a default VPC in your Region. Region assumed: us-east-1, change REGION if you like.

IAM permissions for you (scope them down for real use): autoscaling:*, ec2:RunInstances, ec2:Describe*, ec2:CreateLaunchTemplate*, ec2:DeleteLaunchTemplate, iam:CreateRole, iam:PutRolePolicy, iam:CreateInstanceProfile, iam:AddRoleToInstanceProfile, iam:PassRole, iam:CreateServiceLinkedRole (first-time only, for AWSServiceRoleForAutoScaling), ssm:GetParameter, ssm:PutParameter, ssm:DeleteParameter, cloudwatch:GetMetricStatistics, and the matching delete actions for cleanup.

Step 1: role, template, and group.

import base64, json, time
import boto3
REGION = "us-east-1"
NAME = "asg-demo"
ec2 = boto3.client("ec2", region_name=REGION)
asg = boto3.client("autoscaling", region_name=REGION)
iam = boto3.client("iam")
ssm = boto3.client("ssm", region_name=REGION)
# 1. Instance role that can read ONE parameter (the load switch)
trust = {"Version": "2012-10-17", "Statement": [{
    "Effect": "Allow", "Principal": {"Service": "ec2.amazonaws.com"},
    "Action": "sts:AssumeRole"}]}
iam.create_role(RoleName=NAME, AssumeRolePolicyDocument=json.dumps(trust))
acct = boto3.client("sts").get_caller_identity()["Account"]
iam.put_role_policy(RoleName=NAME, PolicyName="read-switch", PolicyDocument=json.dumps({
    "Version": "2012-10-17", "Statement": [{
        "Effect": "Allow", "Action": "ssm:GetParameter",
        "Resource": f"arn:aws:ssm:{REGION}:{acct}:parameter/{NAME}/burn"}]}))
iam.create_instance_profile(InstanceProfileName=NAME)
iam.add_role_to_instance_profile(InstanceProfileName=NAME, RoleName=NAME)
time.sleep(15)  # IAM propagation
ssm.put_parameter(Name=f"/{NAME}/burn", Value="off", Type="String", Overwrite=True)
# 2. Latest Amazon Linux 2023 AMI from the public SSM parameter
ami = boto3.client("ssm", region_name=REGION).get_parameter(
    Name="/aws/service/ami-amazon-linux-latest/al2023-ami-kernel-default-x86_64"
)["Parameter"]["Value"]
user_data = f"""#!/bin/bash
cat >/usr/local/bin/burn.sh <<'EOF'
#!/bin/bash
while true; do
  V=$(aws ssm get-parameter --name /{NAME}/burn --region {REGION} --query Parameter.Value --output text 2>/dev/null)
  if [ "$V" = "on" ]; then
    pgrep -f "yes" >/dev/null || (yes >/dev/null & yes >/dev/null &)
  else
    pkill yes
  fi
  sleep 20
done
EOF
chmod +x /usr/local/bin/burn.sh
nohup /usr/local/bin/burn.sh >/var/log/burn.log 2>&1 &
"""
ec2.create_launch_template(
    LaunchTemplateName=NAME,
    LaunchTemplateData={
        "ImageId": ami,
        "InstanceType": "t3.micro",
        "IamInstanceProfile": {"Name": NAME},
        "UserData": base64.b64encode(user_data.encode()).decode(),
        "Monitoring": {"Enabled": True},  # one-minute metrics for faster tracking
    },
)
# 3. Subnets from the default VPC
vpc = ec2.describe_vpcs(Filters=[{"Name": "isDefault", "Values": ["true"]}])["Vpcs"][0]["VpcId"]
subnets = [s["SubnetId"] for s in ec2.describe_subnets(
    Filters=[{"Name": "vpc-id", "Values": [vpc]}])["Subnets"]]
asg.create_auto_scaling_group(
    AutoScalingGroupName=NAME,
    LaunchTemplate={"LaunchTemplateName": NAME, "Version": "$Latest"},
    MinSize=1, MaxSize=4, DesiredCapacity=1,
    VPCZoneIdentifier=",".join(subnets),
    HealthCheckType="EC2", HealthCheckGracePeriod=120,
    DefaultInstanceWarmup=120,
)
print("group created across", len(subnets), "subnets")

The t3.micro has two vCPUs, so two yes processes saturate it. DefaultInstanceWarmup is set so target tracking ignores boot noise.

Step 2: target tracking at 50 percent CPU.

asg.put_scaling_policy(
    AutoScalingGroupName=NAME, PolicyName="cpu-50",
    PolicyType="TargetTrackingScaling",
    TargetTrackingConfiguration={
        "PredefinedMetricSpecification": {"PredefinedMetricType": "ASGAverageCPUUtilization"},
        "TargetValue": 50.0,
    },
)

Step 3: a warm pool of Stopped instances.

asg.put_warm_pool(
    AutoScalingGroupName=NAME,
    PoolState="Stopped",
    MinSize=1,
    MaxGroupPreparedCapacity=3,
)

MaxGroupPreparedCapacity=3 caps the pool so that desired plus pooled instances never exceed 3. Without it the pool would be max minus desired.

Step 4: predictive scaling, forecast only.

asg.put_scaling_policy(
    AutoScalingGroupName=NAME, PolicyName="predict-cpu",
    PolicyType="PredictiveScaling",
    PredictiveScalingConfiguration={
        "MetricSpecifications": [{
            "TargetValue": 50.0,
            "PredefinedMetricPairSpecification": {"PredefinedMetricType": "ASGCPUUtilization"},
        }],
        "Mode": "ForecastOnly",
    },
)

Forecast-only is the correct first mode. It produces forecasts and does not scale, and you can read them with GetPredictiveScalingForecast once the group has 24 hours of history. You will not see anything useful today, which is the point: you evaluate the forecast against reality before you let it launch instances.

Step 5: load test and verify.

ssm.put_parameter(Name=f"/{NAME}/burn", Value="on", Type="String", Overwrite=True)
for _ in range(30):
    g = asg.describe_auto_scaling_groups(AutoScalingGroupNames=[NAME])["AutoScalingGroups"][0]
    states = [i["LifecycleState"] for i in g["Instances"]]
    wp = asg.describe_warm_pool(AutoScalingGroupName=NAME).get("Instances", [])
    print(time.strftime("%H:%M:%S"), "desired", g["DesiredCapacity"],
          "inservice", states.count("InService"), "warm", len(wp))
    time.sleep(60)

Expect desired to climb from 1 within a few minutes of the burn flag turning on, as the one-minute CPU metric crosses 50 percent and the alarm fires, then settle at the count that brings average CPU back near the target. Target tracking will over-provision slightly because of rounding. Verification is the activity log:

for a in asg.describe_scaling_activities(AutoScalingGroupName=NAME, MaxRecords=10)["Activities"]:
    print(a["StartTime"], a["StatusCode"], a["Cause"][:110])

You should see activities with causes naming a target tracking alarm, and later instances appearing in the warm pool as Stopped. Turn the flag off and watch desired fall back. Scale-in is slower than scale-out by design: the group waits for the metric to stay low and it is blocked while any instance is warming up.

Roll out a new AMI. When you change the launch template, existing instances do not change. Use an instance refresh, a rolling replacement. Skip matching is on by default from the console and compares AMI IDs only, so it will not notice a change to user data. If you deploy code through user data, turn it off. Set MinHealthyPercentage and the warmup explicitly in your preferences rather than relying on defaults, so your rollout behavior is written down in code.

When it breaks

Scale-out activities show Failed with "You have requested more instances than your current instance limit". Your vCPU quota for that instance family is exhausted. Request an increase in Service Quotas or widen the instance types with a mixed instances policy.

The group launches and terminates instances in a loop. Almost always the health check grace period is shorter than boot time with an ELB health check attached. The instance is declared unhealthy before it can answer. Lengthen the grace period or use a lifecycle hook.

Target tracking never fires. Look at the alarm state. INSUFFICIENT_DATA means the metric has gaps. Check that detailed monitoring is on, that the metric is one that scales inversely with instance count, and that you did not choose a request-count metric that does not divide across instances.

It scaled out and then immediately back in. Warmup too short, or your metric reacts to the boot itself. Raise the default instance warmup. Starting from 300 seconds and adjusting is the documented guidance.

Predictive policy shows no forecast. Under 24 hours of history, or the metric does not reflect load. It is not broken, it is waiting.

Warm pool instances never reach the pool. The root volume must be EBS, hibernation needs its prerequisites or it falls back to Stopped, and a failed lifecycle step terminates the instance.

Cleanup

Order matters: the group first, then what it depends on.

asg.delete_auto_scaling_group(AutoScalingGroupName=NAME, ForceDelete=True)
# wait until instances are gone (poll describe_auto_scaling_groups until empty)
ec2.delete_launch_template(LaunchTemplateName=NAME)
iam.remove_role_from_instance_profile(InstanceProfileName=NAME, RoleName=NAME)
iam.delete_instance_profile(InstanceProfileName=NAME)
iam.delete_role_policy(RoleName=NAME, PolicyName="read-switch")
iam.delete_role(RoleName=NAME)
ssm.delete_parameter(Name=f"/{NAME}/burn")

ForceDelete removes the group together with its warm pool and instances. If you turned on group-level deletion protection, lift it first. Confirm with aws ec2 describe-instances --filters Name=tag:aws:autoscaling:groupName,Values=asg-demo returning nothing, and check that no EBS volumes tagged to the group remain. The bill is then zero, minus the CloudWatch metrics that age out on their own.

Sources