This website uses cookies

Read our Privacy policy and Terms of use for more information.

What it actually is

Amazon EBS is a network-attached block device. Not a disk, not a file system, not object storage. An EBS volume is a block device that lives in one Availability Zone, attaches to an EC2 instance over the network, and survives the instance. You format it, you mount it, you own everything above the block layer.

The mental model the docs bury is that a volume is a performance contract, not a size. You do not buy 100 GiB. You buy 100 GiB, a number of IOPS, and a number of MiB/s, and each volume type defines how those three numbers relate to each other. Most EBS mistakes are people reading the size and ignoring the other two.

One scope note. EBS is the storage under EC2, so if you want the instance side of the story, the first article in this series covers it. This article stands alone: it is about the volume.

The model

Per the EBS volume types documentation, there are two families. SSD-backed volumes (gp3, gp2, io2 Block Express, io1) are tuned for IOPS. HDD-backed volumes (st1, sc1) are tuned for throughput on large sequential I/O and cannot be boot volumes.

The numbers that matter, from that page and the General Purpose and Provisioned IOPS pages. gp3 goes from 1 GiB to 64 TiB, includes 3,000 IOPS and 125 MiB/s in the storage price, and scales to 80,000 IOPS and 2,000 MiB/s. gp2 goes up to 16 TiB and 16,000 IOPS. io2 Block Express goes from 4 GiB to 64 TiB, up to 256,000 IOPS and 4,000 MiB/s. io1 tops out at 64,000 IOPS and 16 TiB. st1 gives you up to 500 IOPS and 500 MiB/s per volume (measured in 1 MiB I/O), sc1 gives you 250 and 250, and both max out at 16 TiB.

The ratios are the contract. On gp3 you can provision up to 500 IOPS per GiB, and up to 0.25 MiB/s of throughput per provisioned IOPS. That is why the maximum 2,000 MiB/s needs at least 8,000 IOPS, and why the maximum 80,000 IOPS needs only 160 GiB. On io2 Block Express the ratio is 1,000 IOPS per GiB, so 256,000 IOPS is reachable at 256 GiB. io1 is 50 IOPS per GiB.

gp3 has no burst model. The docs state that full provisioned IOPS and throughput can be sustained indefinitely. gp2 is different and is the reason gp3 exists: gp2 gives you 3 IOPS per GiB (minimum 100), starts with a 5.4 million credit bucket, and bursts to 3,000 IOPS only on volumes under 1 TiB. When the bucket empties you fall to baseline, which on a 100 GiB volume is 300 IOPS. The CloudWatch metric BurstBalance shows the percentage of credits left.

Durability is the other axis. io2 Block Express is designed for 99.999 percent durability (0.001 percent annual failure rate). Every other type is 99.8 to 99.9 percent (up to 0.2 percent AFR). That is a failure rate of up to two volumes per thousand per year, which is why snapshots are not optional. io2 also supports Multi-Attach, and the docs state that io2 Block Express averages under 500 microseconds latency for 16 KiB I/O on Nitro instances.

Elastic Volumes is the feature that makes all of this less scary. Per the modify documentation, you can change size, type, IOPS and throughput on a live volume, usually without detaching it. Size can only grow. You can modify a volume up to four times in a rolling 24 hours, each modification must reach completed before the next, and you cannot cancel one once submitted.

What AWS operates: replication inside the AZ, the hardware, the performance enforcement. What you operate: the file system, the snapshots, the choice of type, and watching whether you are hitting your own limits.

When to use it, when not to

You are choosing between

Pick EBS when

Pick the other when

Instance store

The data must survive a stop or a host failure, and you want snapshots

You need the lowest latency and the data is a cache you can rebuild

EFS

One instance (or a Multi-Attach cluster) owns the data and you want block semantics

Many instances in many AZs need a shared POSIX file system

S3

You need a mounted, random-write block device for a database or an OS

The data is objects read by many clients, with no need for a file system

The honest summary: EBS is the default for anything that behaves like a disk. The second decision, which type, has a default too. Start with gp3. The docs say gp3 is 20 percent lower per GiB than gp2, and unlike gp2 it lets you buy IOPS and throughput without buying capacity. Reach for io2 when you need sub-millisecond consistency or more than 80,000 IOPS, and for st1 when you stream large sequential files and the price per GiB matters more than latency.

What it costs

The EBS pricing page gives its figures in worked examples that do not name a Region, so treat the numbers below as illustrations and confirm your Region in the AWS Pricing Calculator.

gp3 is $0.08 per GB-month, includes 3,000 IOPS and 125 MB/s, then charges $0.005 per provisioned IOPS-month above 3,000 and $0.06 per provisioned MB/s-month above 125. gp2 is $0.10 per GB-month. io2 is $0.125 per GB-month plus $0.065 per provisioned IOPS-month for the first 32,000 IOPS and $0.046 for 32,001 to 64,000. st1 is $0.045 and sc1 is $0.015 per GB-month. Snapshots are $0.05 per GB-month in the standard tier and $0.0125 in the archive tier, with a $0.03 per GB retrieval fee.

The line that surprises people is that you pay for provisioned, not used. A 100 GiB gp3 volume with 16,000 IOPS provisioned costs the same whether it sits idle or runs flat out. Do the arithmetic once: 13,000 extra IOPS at $0.005 is $65 a month on a volume whose storage is $8. The second surprise is that detached volumes keep billing. The third is that snapshots are incremental, so the first snapshot is the expensive one and the bill depends on your change rate, not your volume size.

Free tier, per the pricing page: 30 GB of storage, 2 million I/Os and 1 GB of snapshot storage. This tutorial fits inside the storage part, but the extra IOPS and throughput we provision in step 4 are not free. Provisioning 10,000 IOPS and 400 MB/s above baseline for under an hour costs a few cents at the example rates, plus the small EC2 instance you run it on.

The limits that bite

First, the instance is a second ceiling. A volume can promise 80,000 IOPS and your instance can still cap you far below it. The io2 documentation states that only Nitro instances reach 256,000 IOPS on io2, and non-Nitro instances top out around 32,000 even on volumes provisioned higher. Every instance type has its own EBS bandwidth and IOPS limit. Check the instance before you blame the volume.

Second, modification is rate limited: four changes per rolling 24 hours, one at a time, and a 1 TiB volume can take up to six hours to finish optimizing. Do not plan a live resize as an incident response tool.

Third, restored volumes are slow at first. Per the initialization documentation, a volume created from a snapshot downloads blocks from S3 lazily, and until that finishes you may see elevated latency. You have three fixes: set a volume initialization rate between 100 and 300 MiB/s, enable Fast Snapshot Restore, or read the whole device with fio or dd before production use. The rate option bills based on the snapshot data size and the rate you pick, and it charges even if you delete the volume before initialization finishes.

Fourth, the AZ boundary. A volume cannot attach across AZs. Moving it means snapshot and restore in the target AZ.

Fifth, regional quotas on storage, IOPS and snapshots exist and vary. I could not retrieve the quota table while writing this, so open Service Quotas for EBS in your account and read your own numbers rather than trusting anyone's blog post, including this one.

Build it

We will create a gp3 volume with its default 3,000 IOPS and 125 MiB/s, attach it to a Nitro instance, benchmark it with fio, raise IOPS and throughput live with Elastic Volumes, and benchmark again. The deliverable is two fio results that prove the contract.

Prerequisites

You need an AWS account, the AWS CLI configured, Python 3.10+ with boto3, and a running Linux EC2 instance on a Nitro-based type with SSM or SSH access, in a Region you control. Note its instance ID and AZ. Pick an instance type whose EBS bandwidth exceeds what you plan to provision, or the second test will measure the instance and not the volume. I did not execute this against an account while writing; the code follows the cited documentation, so expect to adjust names for your environment.

IAM permissions for the identity running the script: ec2:CreateVolume, ec2:AttachVolume, ec2:DescribeVolumes, ec2:ModifyVolume, ec2:DescribeVolumesModifications, ec2:DetachVolume, ec2:DeleteVolume, ec2:CreateTags and ec2:DescribeInstances. Scope them to your Region and to resources tagged for this exercise where you can.

pip install "boto3>=1.34"
export AWS_REGION=us-east-1
export INSTANCE_ID=i-0123456789abcdef0

Step 1: create the volume

Create the volume in the instance's AZ. We leave IOPS and throughput unset so gp3 gives us the included baseline.

import boto3, os, time
ec2 = boto3.client("ec2", region_name=os.environ["AWS_REGION"])
iid = os.environ["INSTANCE_ID"]
az = ec2.describe_instances(InstanceIds=[iid])["Reservations"][0]["Instances"][0]["Placement"]["AvailabilityZone"]

vol = ec2.create_volume(
    AvailabilityZone=az, Size=100, VolumeType="gp3",
    TagSpecifications=[{"ResourceType": "volume",
        "Tags": [{"Key": "purpose", "Value": "sl118-ebs-demo"}]}],
)
vid = vol["VolumeId"]
ec2.get_waiter("volume_available").wait(VolumeIds=[vid])
print(vid, az)

The tag matters for cleanup. Volumes created this way are empty, so they need no initialization.

Step 2: attach and format

Attach the volume as a new device, then format and mount it from the instance.

ec2.attach_volume(VolumeId=vid, InstanceId=iid, Device="/dev/sdf")
ec2.get_waiter("volume_in_use").wait(VolumeIds=[vid])

On the instance, find the device with lsblk. On Nitro instances it appears as an NVMe device such as /dev/nvme1n1, not /dev/sdf, so check before you format anything.

lsblk
sudo mkfs -t xfs /dev/nvme1n1
sudo mkdir -p /data && sudo mount /dev/nvme1n1 /data

If lsblk shows more than one unmounted disk, stop and identify the right one by size. mkfs on the wrong device destroys data.

Step 3: measure the baseline

Install fio and run a random read test at 4 KiB against a file on the volume, with direct I/O so the page cache does not flatter you. This is the same style of command the EBS docs use for volume reads, adapted for random I/O.

sudo yum install -y fio   # or: sudo apt-get install -y fio
sudo fio --name=base --directory=/data --size=8G --rw=randread --bs=4k \
  --iodepth=64 --numjobs=4 --ioengine=libaio --direct=1 \
  --time_based --runtime=60 --group_reporting

Read the IOPS line near the top of the summary. On a default gp3 volume you should see a result close to 3,000 IOPS, because that is the baseline you bought. If you see much more, check that --direct=1 took effect and that the file is larger than memory pressure could hide.

Step 4: change the contract live

Now raise provisioned IOPS and throughput with Elastic Volumes. The gp3 ratios from earlier constrain what is legal: 10,000 IOPS on a 100 GiB volume is fine (the limit is 500 IOPS per GiB), and 400 MiB/s is fine because it sits under 0.25 MiB/s per IOPS.

ec2.modify_volume(VolumeId=vid, Iops=10000, Throughput=400)

while True:
    m = ec2.describe_volumes_modifications(VolumeIds=[vid])["VolumesModifications"][0]
    print(m["ModificationState"], m.get("Progress"))
    if m["ModificationState"] in ("optimizing", "completed"):
        break
    time.sleep(10)

The volume stays attached and mounted throughout. The modification reaches optimizing quickly and completed later. The docs say a new configuration is billed from the moment the modification starts.

Step 5: measure again

Repeat the exact fio command from step 3, changing only the --name.

sudo fio --name=after --directory=/data --size=8G --rw=randread --bs=4k \
  --iodepth=64 --numjobs=4 --ioengine=libaio --direct=1 \
  --time_based --runtime=60 --group_reporting

Verify it works

Two checks. First, the control plane: this should report 10000 and 400.

v = ec2.describe_volumes(VolumeIds=[vid])["Volumes"][0]
print(v["VolumeType"], v["Iops"], v["Throughput"], v["Size"])

Second, the data plane: the second fio run should land near 10,000 IOPS, versus about 3,000 in the first. If it lands well below that, the instance is your ceiling: check its EBS-optimized limits. In CloudWatch, the VolumeIOPSExceededCheck and VolumeThroughputExceededCheck metrics report 1 when the application consistently tried to exceed what you provisioned, and all EBS metrics arrive in one-minute periods only while the volume is attached. Those two metrics are not published for volumes attached to ECS or Fargate tasks.

When it breaks

VolumeInUse or IncorrectState on modify means the previous modification has not reached completed, or you used your four changes in 24 hours. Wait, and check describe_volumes_modifications before retrying.

InvalidParameterCombination on modify or create usually means you violated a ratio: too many IOPS for the size, or too much throughput for the IOPS. Re-read the gp3 ratios above.

The device is missing after attach. Nitro exposes volumes as NVMe devices, so look in lsblk, not for the /dev/sdf name you passed.

You resized but df shows the old size. Size changes happen at the block layer. Grow the partition if there is one, then the file system (xfs_growfs for XFS), before you expect new space.

Throughput stops well under provisioned. The instance type's EBS bandwidth is the cap, or your I/O size is small: throughput is IOPS times I/O size, so 4 KiB I/O at 10,000 IOPS is only about 39 MiB/s no matter what you provisioned.

On a gp2 volume, latency goes bad after about half an hour of heavy load. BurstBalance hit zero and you are at baseline. Migrate to gp3 with the same modify call, with VolumeType="gp3".

Cleanup

Unmount, detach and delete, then confirm nothing tagged remains. Detached volumes keep billing, so deleting is the only way to reach zero.

sudo umount /data
ec2.detach_volume(VolumeId=vid)
ec2.get_waiter("volume_available").wait(VolumeIds=[vid])
ec2.delete_volume(VolumeId=vid)

left = ec2.describe_volumes(Filters=[{"Name": "tag:purpose", "Values": ["sl118-ebs-demo"]}])["Volumes"]
print("remaining volumes:", [v["VolumeId"] for v in left])

The last line should print an empty list. If you took snapshots along the way, delete those too, and stop or terminate the EC2 instance if you created it only for this.

The takeaway

Stop sizing volumes. Size them, then price the IOPS and throughput next to the capacity, then check the instance can carry them. Default to gp3, watch the exceeded-check metrics instead of guessing, and snapshot anything you cannot afford to lose, because 99.8 percent durability is a number you should read twice.

Sources