What it actually is
S3 is not a file system and not a disk. It is a regional key-value store over HTTPS: you PUT an object under a key, you GET it back by that key, and the "folders" you see in the console are a rendering of key prefixes that happen to contain slashes. There is no rename, no append, no in-place edit. An object is written whole or not at all, and the update of a single key is atomic.
That narrow contract is the whole trick. Because S3 never promises partial writes, locking, or directory semantics, AWS can spread your data across a huge fleet and give you a durability number that no disk you operate can match. Everything else you read about S3 (storage classes, lifecycle, replication, event notifications) is machinery bolted onto that one idea.
Today the service has grown several bucket types: general purpose buckets (the original, and what this article covers), directory buckets for S3 Express One Zone, table buckets for Apache Iceberg, and vector buckets for similarity search. If you have only ever used S3 as "the place where files go", you are using the general purpose type.
The model
The primitives are small. A bucket is a globally unique name that lives in one Region. An object is a key, a body, system metadata, optional user metadata and tags, and, if you turn it on, a version ID. A prefix is just a leading chunk of the key. Access is decided by IAM policies, bucket policies, and Block Public Access settings.
Three contracts matter.
Consistency. The documentation states that S3 provides strong read-after-write consistency for PUT and DELETE requests of objects in all AWS Regions, including overwrites. Read operations on object metadata, tags, and ACLs are strongly consistent too. The old advice to avoid read-after-overwrite is obsolete. What is still eventually consistent is bucket configuration: after enabling versioning the docs recommend waiting 15 minutes before issuing writes, and bucket deletion also takes time to propagate.
Concurrency. S3 does not lock. The docs say that if two PUT requests hit the same key at the same time, the request with the latest timestamp wins. If you need compare-and-swap, use conditional writes: send If-None-Match: * to create only if the key does not exist, or If-Match with an ETag to overwrite only if the object has not changed. Both work on PutObject, CopyObject, and CompleteMultipartUpload, and a failed precondition returns HTTP 412.
Durability. The documentation states 99.999999999 percent (eleven nines) durability and 99.99 percent availability for S3 Standard over a year, with objects stored redundantly across a minimum of three Availability Zones. S3 One Zone-IA is the exception and keeps data inside a single AZ. Durability is not a backup: eleven nines protects you from hardware loss, not from your own DeleteObject call or a bad deploy. Versioning and replication are what protect you from yourself.
What AWS operates: the fleet, replication, integrity checks, scaling of request partitions. What you operate: naming and key design, access policy, encryption keys if you choose KMS, lifecycle rules, and the decision of how many versions you are willing to pay for.
New buckets ship with sane defaults. The docs state that new buckets, access points, and objects do not allow public access, that Object Ownership defaults to bucket owner enforced with ACLs disabled, and that every bucket has server-side encryption with S3 managed keys (SSE-S3) configured by default. The same page notes that starting April 2026, SSE-C is automatically disabled for all new general purpose buckets. If you were relying on customer-provided keys, check that.
When to use it, when not to
You want | Use | Why not S3 |
|---|---|---|
Durable blobs, static assets, logs, backups, data lake storage | S3 | This is its job |
A POSIX file system shared by servers | EFS (or FSx) | S3 has no rename, append, or locking |
A block device for one instance | EBS | S3 is not mountable storage with disk latency |
Single-digit millisecond access for hot data | S3 Express One Zone (directory buckets) | General purpose S3 is not built for that latency |
Rows you query by attribute with updates | DynamoDB or a database | S3 gives you key lookup and listing, nothing else |
The honest rule: if your access pattern is "write once, read by key, delete by policy", S3 is almost always right. If your pattern is "update a few bytes in the middle" or "list everything and filter", you are fighting the model. AWS has also been moving toward file-like access on top of S3, so check the current S3 Files documentation before assuming you must pick EFS for shared file access.
What it costs
S3 charges on five axes: storage (GB-month, by class), requests (by type), data retrieval (for infrequent and archive classes), data transfer out, and management features such as inventory or replication. Most teams model only the first.
Published us-east-1 numbers on the S3 pricing page at the time of writing: S3 Standard is $0.023 per GB-month for the first 50 TB, S3 Standard-IA is $0.0125, S3 One Zone-IA is $0.01, Glacier Instant Retrieval is $0.004, Glacier Flexible Retrieval is $0.0036, and Glacier Deep Archive is $0.00099. PUT, COPY, POST, and LIST requests are $0.005 per 1,000. GET requests are $0.0004 per 1,000. Prices vary by Region and AWS changes them, so confirm on the pricing page before you commit to a design.
Run the arithmetic on a normal workload. 100 GB in Standard costs $2.30 a month. Now add an application that writes 10 million small objects a month: 10,000,000 / 1,000 x $0.005 = $50. The requests cost more than twenty times the storage. Ten million GETs add $4. This is the surprise: for small objects, the PUT bill dwarfs the GB bill, and no lifecycle policy fixes it. The fix is architectural: batch small records into larger objects before writing.
The second surprise is versioning. The docs are explicit that every version is a full object, not a diff, so three versions of a 1 GB object bill as 3 GB. Enable versioning without a noncurrent-version expiration rule and your bill grows forever.
The third is the IA and archive minimums. The pricing page lists a 128 KB minimum billable object size and a 30-day minimum duration for Standard-IA, 90 days for Glacier Instant and Flexible Retrieval, and 180 days for Deep Archive. Deleting early still charges you for the remainder. Lifecycle transitions also carry overhead: the transition documentation describes 40 KB of added metadata per transitioned object (8 KB at Standard rates, 32 KB at the destination class rate), which is why tiny objects do not transition by default.
Free tier reality: new accounts get a limited free tier for storage and requests, and the first 100 GB of outbound data per month, aggregated across services and Regions, is free. Check the pricing page for the current allowances rather than trusting a blog post, including this one.
The limits that bite
The numbers, from current documentation:
Maximum object size is 48.8 TiB, uploaded through multipart upload. Parts are 5 MiB to 5 GiB, with up to 10,000 parts per upload. The last part has no minimum. AWS recommends multipart once objects reach about 100 MB.
Request rates are 3,500 PUT/COPY/POST/DELETE and 5,500 GET/HEAD requests per second per partitioned prefix. There is no limit on prefixes per bucket, so ten prefixes can scale reads to 55,000 per second.
Scaling is gradual. During a sudden ramp you can get HTTP 503 SlowDown responses, which clear once S3 has partitioned. Backoff and retry is the correct response, and the SDKs do it for you.
The default quota is 10,000 general purpose buckets per account, increasable through Service Quotas. Bucket names are globally unique and bucket policies are capped at 20 KB.
There is no cap on bucket size or object count.
The failure mode at scale is almost never the byte count. It is hot prefixes, listing a bucket with billions of keys (LIST is slow, sequential, and billed), and unbounded version history.
Build it
We will build a bucket with versioning, a lifecycle policy, an enforced-TLS policy, and then prove both the consistency contract and the cost model with real calls.
Prerequisites. Python 3.9+, boto3 installed, and AWS credentials configured. Region: your default. Estimated cost of following along: well under one cent. You store a few kilobytes for a few minutes and make fewer than 100 requests. Cleanup is included, so nothing lingers.
IAM permissions. Attach a policy like this to the identity you run the script with. Replace the bucket name pattern with your own.
{
"Version": "2012-10-17",
"Statement": [{
"Effect": "Allow",
"Action": [
"s3:CreateBucket", "s3:DeleteBucket",
"s3:PutBucketVersioning", "s3:GetBucketVersioning",
"s3:PutLifecycleConfiguration", "s3:GetLifecycleConfiguration",
"s3:PutBucketPolicy",
"s3:PutObject", "s3:GetObject", "s3:DeleteObject", "s3:DeleteObjectVersion",
"s3:ListBucket", "s3:ListBucketVersions"
],
"Resource": ["arn:aws:s3:::sl114-demo-*", "arn:aws:s3:::sl114-demo-*/*"]
}]
}Step 1: create the bucket and turn on versioning. Bucket names are global, so we add a random suffix. Outside us-east-1 you must pass a LocationConstraint.
import json, uuid, boto3
from botocore.exceptions import ClientError
s3 = boto3.client("s3")
region = s3.meta.region_name or "us-east-1"
bucket = f"sl114-demo-{uuid.uuid4().hex[:10]}"
args = {"Bucket": bucket}
if region != "us-east-1":
args["CreateBucketConfiguration"] = {"LocationConstraint": region}
s3.create_bucket(**args)
s3.put_bucket_versioning(
Bucket=bucket,
VersioningConfiguration={"Status": "Enabled"},
)
print("created", bucket, "in", region)The docs recommend waiting about 15 minutes after enabling versioning on a bucket before relying on it for writes, because that configuration is eventually consistent. For a brand new empty bucket the practical risk is low, but in production, enable versioning at creation time through infrastructure as code and do not toggle it live.
Step 2: add the lifecycle policy. This rule transitions current objects under data/ to Standard-IA after 30 days, expires old versions after 30 days while keeping the three newest noncurrent versions, and aborts incomplete multipart uploads after 7 days. The abort rule is the one people forget: abandoned multipart parts are invisible in the console listing and still billed.
s3.put_bucket_lifecycle_configuration(
Bucket=bucket,
LifecycleConfiguration={"Rules": [
{
"ID": "tier-and-trim",
"Status": "Enabled",
"Filter": {"Prefix": "data/"},
"Transitions": [{"Days": 30, "StorageClass": "STANDARD_IA"}],
"NoncurrentVersionExpiration": {
"NoncurrentDays": 30,
"NewerNoncurrentVersions": 3,
},
"AbortIncompleteMultipartUpload": {"DaysAfterInitiation": 7},
}
]},
)
print(s3.get_bucket_lifecycle_configuration(Bucket=bucket)["Rules"][0]["ID"])Because of the September 2024 change, objects smaller than 128 KB do not transition by default under new configurations. That is what you want: a 10 KB object in Standard-IA would bill as 128 KB.
Step 3: refuse plaintext HTTP. Block Public Access and SSE-S3 are already on by default. Add a bucket policy that denies any request not using TLS.
policy = {
"Version": "2012-10-17",
"Statement": [{
"Sid": "DenyInsecureTransport",
"Effect": "Deny",
"Principal": "*",
"Action": "s3:*",
"Resource": [f"arn:aws:s3:::{bucket}", f"arn:aws:s3:::{bucket}/*"],
"Condition": {"Bool": {"aws:SecureTransport": "false"}},
}],
}
s3.put_bucket_policy(Bucket=bucket, Policy=json.dumps(policy))Step 4: prove strong consistency and versioning. Overwrite a key and read it back immediately. Then delete it and watch the delete marker appear instead of the data disappearing.
key = "data/hello.txt"
for i in range(1, 4):
s3.put_object(Bucket=bucket, Key=key, Body=f"version {i}".encode())
got = s3.get_object(Bucket=bucket, Key=key)["Body"].read().decode()
assert got == f"version {i}", got
print("read-after-overwrite: ok")
s3.delete_object(Bucket=bucket, Key=key)
try:
s3.get_object(Bucket=bucket, Key=key)
except ClientError as e:
print("after delete:", e.response["Error"]["Code"]) # NoSuchKey
vers = s3.list_object_versions(Bucket=bucket, Prefix=key)
print("versions:", len(vers.get("Versions", [])),
"delete markers:", len(vers.get("DeleteMarkers", [])))Expected output: three assertions pass, NoSuchKey after the delete, and versions: 3 delete markers: 1. The data you "deleted" is still there and still billed. That is the versioning contract in four lines.
Step 5: conditional write as a lock-free guard. Create-if-absent is the primitive behind safe idempotent writers.
s3.put_object(Bucket=bucket, Key="data/lock.json", Body=b"{}", IfNoneMatch="*")
try:
s3.put_object(Bucket=bucket, Key="data/lock.json", Body=b"{}", IfNoneMatch="*")
except ClientError as e:
print(e.response["Error"]["Code"]) # PreconditionFailedThe first call succeeds, the second returns HTTP 412 PreconditionFailed. If your boto3 predates conditional write support, upgrade it, or the parameter is rejected locally.
Step 6: cost the bucket yourself. This turns the pricing section into code. The rates are the published us-east-1 numbers quoted above, so they are inputs you should refresh, not facts about your account.
STORAGE_GB_MONTH = 0.023 # Standard, first 50 TB
PUT_PER_1000 = 0.005
GET_PER_1000 = 0.0004
def monthly_cost(gb, puts, gets):
storage = gb * STORAGE_GB_MONTH
put_cost = puts / 1000 * PUT_PER_1000
get_cost = gets / 1000 * GET_PER_1000
return storage, put_cost, get_cost
for label, gb, puts, gets in [
("100 GB, 10M small writes, 10M reads", 100, 10_000_000, 10_000_000),
("100 GB, same data as 10k batched writes", 100, 10_000, 10_000_000),
]:
s, p, g = monthly_cost(gb, puts, gets)
print(f"{label}: storage ${s:.2f} + PUT ${p:.2f} + GET ${g:.2f} = ${s+p+g:.2f}")Output: the first scenario totals $56.30 (2.30 + 50.00 + 4.00), the second $6.35 (2.30 + 0.05 + 4.00). Same bytes, nine times cheaper, purely from request count. Run this against your own traffic numbers before you design a writer.
Verification. Confirm the configuration took:
print(s3.get_bucket_versioning(Bucket=bucket).get("Status")) # Enabled
print([r["ID"] for r in s3.get_bucket_lifecycle_configuration(Bucket=bucket)["Rules"]])You should see Enabled and ['tier-and-trim'].
When it breaks
BucketAlreadyExists (409). Someone else on Earth has that name. The namespace is global, so add a suffix. BucketAlreadyOwnedByYou means you created it earlier; in us-east-1 the create call is a no-op that returns 200, in other Regions it returns 409.
AccessDenied (403). The credentials lack the permission, or a bucket policy denies you. Our TLS-deny policy will also deny anything that is not HTTPS, which can bite custom tooling. Read the bucket policy before you blame IAM.
PreconditionFailed (412). A conditional write found the key already present or the ETag changed. That is the feature working. A 409 OperationAborted or 409 Conflict on a conditional write means a conflicting operation was in flight, so retry.
503 SlowDown. You exceeded the current request capacity of a prefix. Retry with exponential backoff, and spread keys across prefixes if you expect sustained high rates.
BucketNotEmpty (409). You cannot delete a bucket that still holds objects, and in a versioned bucket "objects" includes every old version and delete marker. This one catches everyone during cleanup, which is why the cleanup script below deletes versions explicitly.
The bill grew and nothing changed. Check noncurrent versions and incomplete multipart uploads first. Both are invisible in a basic listing.
Cleanup
A versioned bucket must have every version and delete marker removed before the bucket itself can go. The bucket policy denies insecure transport only, so boto3 over HTTPS is unaffected.
paginator = s3.get_paginator("list_object_versions")
for page in paginator.paginate(Bucket=bucket):
items = page.get("Versions", []) + page.get("DeleteMarkers", [])
if items:
s3.delete_objects(Bucket=bucket, Delete={
"Objects": [{"Key": i["Key"], "VersionId": i["VersionId"]} for i in items],
"Quiet": True,
})
s3.delete_bucket(Bucket=bucket)
print("deleted", bucket)Confirm with aws s3api list-buckets --query "Buckets[?starts_with(Name, 'sl114-demo-')]", which should return an empty list. With the bucket gone, storage and request charges stop. There is nothing else to remove.
Where to go next
If you want the analytics side of S3, the earlier deep dive on S3 Tables covers Iceberg, time travel, and compaction: Build an Iceberg Lakehouse on S3 Tables. This article is the base layer under all of it. The next service in the series is Amazon S3 Glacier, where the retrieval bill is the thing nobody models.

