This website uses cookies

Read our Privacy policy and Terms of use for more information.

What it actually is

Amazon S3 Tables is a bucket type. Not a database, not a catalog service, not a query engine. A table bucket is an S3 bucket whose children are Apache Iceberg tables instead of objects, and S3 itself operates those tables: it compacts small files, expires old snapshots, and deletes files nothing references anymore. You bring the writers and the readers. AWS brings the storage and the housekeeping.

The mental model the docs bury is the contrast with what you do today. A normal Iceberg setup on a general purpose bucket means you run the catalog, you schedule compaction jobs, you schedule snapshot expiry, and you write the orphan-file cleanup that nobody remembers until the bill shows up. S3 Tables moves those three chores below the API line. You still choose the Iceberg format, the schema, and the query engine. You stop being the operator of the table's physical layout.

One scope note. This article is the reference entry for the service. If you want the longer lakehouse build with time travel and a Spark read, that is the earlier deep dive, SL#84 - AWS AI Series (21/30). Here the deliverable is different: we will create the resources from code, then read back the maintenance contract S3 is enforcing, because that contract is where the surprises live.

The model

Per the S3 Tables documentation, there are three primitives. A table bucket is the new bucket type. A namespace groups tables inside a table bucket. A table is an Iceberg table stored as a subresource of the bucket. Tables are not objects you list with the regular S3 object API; you manage them through the separate s3tables API and service namespace, and you address them with ARNs of the form arn:aws:s3tables:region:account-id:bucket/bucket-name/table/table-id.

The data format contract is Apache Iceberg: schema evolution, partition evolution, transactions, snapshots, time travel, and rollback. The docs state that table buckets support Iceberg V3, including deletion vectors, row lineage, column default values, and the variant, geometry, geography, unknown, and nanosecond-precision timestamp types. Durability, availability, and scalability are the same as other S3 bucket types, with higher transactions per second than self-managed Iceberg tables in a general purpose bucket.

Maintenance is the real product. Three jobs run for you. Compaction merges small files into larger ones, with a default target of 512 MB and a configurable range of 64 MB to 512 MB, and it applies row-level deletes while it works. Snapshot management expires old snapshots; the defaults are a minimum of 1 snapshot kept and a maximum age of 120 hours. Unreferenced file removal marks objects that no table references as noncurrent after 3 days by default, then deletes them permanently after 10 more days. Compaction and snapshot management are configurable per table, and snapshot management is table level only. Unreferenced file removal is configured at the table bucket level.

Two security facts matter before you write any code. All S3 Block Public Access settings are always on for table buckets and cannot be turned off. And access control uses its own IAM namespace, s3tables, so a policy on a general purpose bucket does nothing here.

What AWS operates: the storage, the three maintenance jobs, the Iceberg metadata housekeeping. What you operate: the schema, the writers, the choice of query engine, the Glue Data Catalog integration, and the access policies.

When to use it, when not to

You are choosing between

Pick S3 Tables when

Pick the other when

Iceberg on a general purpose S3 bucket

You do not want to own compaction, expiry, and orphan cleanup jobs

You need full control of file layout, or an engine that cannot talk to the s3tables catalog

Amazon Redshift or another warehouse

Your data is append-heavy, queried by several engines, and you want one copy

You need low-latency BI on a tuned warehouse with its own storage

DynamoDB or a relational database

You run analytical scans over large tables

You need single-row transactional reads and writes at low latency

The honest summary: S3 Tables is for analytics datasets that several engines read. Athena, Redshift, Amazon EMR, AWS Glue, and Amazon SageMaker Unified Studio are listed as integrated services, with Lake Formation available for fine-grained access control. If one engine owns the data and nobody else reads it, the case for it weakens.

What it costs

Pricing below is from the S3 pricing page. The page's own examples use US West (Oregon), so confirm the rates for your Region in the AWS Pricing Calculator before you plan anything.

Storage is $0.0265 per GB per month for the first 50 TB. PUT requests are $0.005 per 1,000 and GET requests are $0.0004 per 1,000. Then come the two dimensions that make table buckets different from plain S3: an object monitoring fee of $0.025 per 1,000 objects, and compaction charges of $0.002 per 1,000 objects processed plus $0.005 per GB processed for the default binpack strategy.

The one that surprises people is the interaction between those last two. Object monitoring is billed per object, and small files are objects. A streaming writer that commits every few seconds creates thousands of small files, you pay monitoring on each until compaction merges them, and you pay compaction to merge them. Batching your writes is not a performance tip here, it is a line item. The docs also warn that the sort and z-order compaction strategies may cost more than binpack.

Integration with the Glue Data Catalog can add Glue request and storage costs, and the query engine charges separately. There is no free tier I can verify for table buckets, so assume this walkthrough costs a few cents in S3 Tables charges plus Athena's per-query billing, which you should check on the Athena pricing page. The build below writes a handful of rows.

The limits that bite

The default quotas are 10 table buckets per Region per account, 10,000 namespaces per table bucket, and 10,000 tables per table bucket. All three are adjustable through a Support case.

Naming is the limit that actually costs people an afternoon. Table bucket names are 3 to 63 characters, lowercase letters, numbers, and hyphens, and cannot be changed after creation. Namespace names are 1 to 255 characters using a-z, 0-9, and underscores, and cannot start with an underscore. Table and column names must be all lowercase: a capital letter makes the table invisible to Lake Formation and the Glue Data Catalog, so Athena will not see it, and queries fail with a GENERIC_INTERNAL_ERROR about invalid table or column names.

Other limits worth knowing: you need AWS CLI 2.23.10 or later for the CLI paths in the docs, snapshot management fails for a whole table if the table has user-defined tags or branches, or if the Iceberg properties history.expire.max-snapshot-age-ms or history.expire.min-snapshots-to-keep are set, and deletion of noncurrent objects is permanent. Recovering them requires AWS Support.

Build it

We will create a table bucket, a namespace, and an Iceberg table with boto3, write small batches through Athena, and then read back the maintenance state. Everything from here follows the commands and parameters in the S3 Tables docs. I have not run this exact script against an account for this issue, so treat the first run as the verification and read the error messages in the next section if something differs.

Prerequisites: an AWS account, Python 3.10+, pip install boto3, AWS CLI 2.23.10 or later, and credentials for an IAM identity. For a sandbox, the docs name the managed policy AmazonS3TablesFullAccess for the table work. You also need glue:CreateCatalog and glue:PassConnection for the one-time integration, and Athena permissions plus an S3 bucket in a general purpose bucket for query results. Use us-east-1 and substitute your own account ID.

First, the one-time integration with the Glue Data Catalog. Programmatic creation of a table bucket does not integrate it automatically, so do this once per Region. Save this as catalog.json, replacing the account ID:

{
  "Name": "s3tablescatalog",
  "CatalogInput": {
    "FederatedCatalog": {
      "Identifier": "arn:aws:s3tables:us-east-1:111122223333:bucket/*",
      "ConnectionName": "aws:s3tables"
    },
    "CreateDatabaseDefaultPermissions": [
      {"Principal": {"DataLakePrincipalIdentifier": "IAM_ALLOWED_PRINCIPALS"}, "Permissions": ["ALL"]}
    ],
    "CreateTableDefaultPermissions": [
      {"Principal": {"DataLakePrincipalIdentifier": "IAM_ALLOWED_PRINCIPALS"}, "Permissions": ["ALL"]}
    ],
    "AllowFullTableExternalDataAccess": "True"
  }
}
aws glue create-catalog --region us-east-1 --cli-input-json file://catalog.json
aws glue get-catalog --catalog-id s3tablescatalog

If the catalog already exists from the console flow, the first command will say so and you can skip it. The second command is your check.

Now the resources. This script creates the bucket, the namespace, and the table, using the same schema the docs use:

import boto3
REGION = "us-east-1"
BUCKET = "sl-demo-table-bucket"   # lowercase, 3-63 chars, unique per account/Region
NAMESPACE = "demo_ns"
TABLE = "events"
s3t = boto3.client("s3tables", region_name=REGION)
arn = s3t.create_table_bucket(name=BUCKET)["arn"]
s3t.create_namespace(tableBucketARN=arn, namespace=[NAMESPACE])
s3t.create_table(
    tableBucketARN=arn,
    namespace=NAMESPACE,
    name=TABLE,
    format="ICEBERG",
    metadata={"iceberg": {"schema": {"fields": [
        {"name": "id", "type": "int", "required": True},
        {"name": "name", "type": "string"},
        {"name": "value", "type": "int"},
    ]}}},
)
print(arn)

The namespace argument is a list in the API and a string in the CLI. That is the one place the two surfaces differ, and it is the first thing to check if create_namespace complains about parameter validation.

Next, write data the way a naive pipeline would: many tiny commits. Each INSERT through Athena is one Iceberg commit, so ten of them give you ten snapshots and a pile of small data files. Set OUTPUT to a general purpose bucket you own:

import time
OUTPUT = "s3://YOUR-ATHENA-RESULTS-BUCKET/s3tables-demo/"
athena = boto3.client("athena", region_name=REGION)
CATALOG = f"s3tablescatalog/{BUCKET}"
def run(sql):
    qid = athena.start_query_execution(
        QueryString=sql,
        QueryExecutionContext={"Catalog": CATALOG, "Database": NAMESPACE},
        ResultConfiguration={"OutputLocation": OUTPUT},
    )["QueryExecutionId"]
    while True:
        s = athena.get_query_execution(QueryExecutionId=qid)["QueryExecution"]["Status"]
        if s["State"] in ("SUCCEEDED", "FAILED", "CANCELLED"):
            return qid, s
        time.sleep(1)
for i in range(10):
    qid, s = run(f"INSERT INTO {TABLE} VALUES ({i}, 'row-{i}', {i * 10})")
    print(i, s["State"], s.get("StateChangeReason", ""))

The catalog path s3tablescatalog/<table-bucket-name> is the documented form, the database is your namespace, and the docs say Athena supports DDL, DML, and DQL on S3 tables. If your account's Athena setup differs, the failure will be in the StateChangeReason string printed above.

Last, read the maintenance contract. This is the part most tutorials skip. Set a smaller compaction target and a longer snapshot window on this one table, then read both back along with the job status:

s3t.put_table_maintenance_configuration(
    tableBucketARN=arn, namespace=NAMESPACE, name=TABLE,
    type="icebergCompaction",
    value={"status": "enabled",
           "settings": {"icebergCompaction": {"targetFileSizeMB": 256}}},
)
s3t.put_table_maintenance_configuration(
    tableBucketARN=arn, namespace=NAMESPACE, name=TABLE,
    type="icebergSnapshotManagement",
    value={"status": "enabled",
           "settings": {"icebergSnapshotManagement":
                        {"minSnapshotsToKeep": 5, "maxSnapshotAgeHours": 240}}},
)
print(s3t.get_table_maintenance_configuration(
    tableBucketARN=arn, namespace=NAMESPACE, name=TABLE))
print(s3t.get_table_maintenance_job_status(
    tableBucketARN=arn, namespace=NAMESPACE, name=TABLE))

The values map one to one to the CLI JSON in the maintenance docs: targetFileSizeMB for compaction, minSnapshotsToKeep and maxSnapshotAgeHours for snapshots. The bucket-level unreferenced file removal policy is the one to think hardest about, because it deletes permanently. Its documented defaults are unreferencedDays 3 and nonCurrentDays 10, set with aws s3tables put-table-bucket-maintenance-configuration --type icebergUnreferencedFileRemoval. Tighten nonCurrentDays only if you have decided you will never need to recover a file.

Verify it works

Run a query and a configuration read. The table should hold ten rows and the configuration should echo your values.

qid, s = run(f"SELECT count(*) AS n FROM {TABLE}")
rows = athena.get_query_results(QueryExecutionId=qid)["ResultSet"]["Rows"]
print(rows[1]["Data"][0]["VarCharValue"])

Expected output: 10. For the configuration call, expect the response to contain icebergCompaction with targetFileSizeMB of 256 and icebergSnapshotManagement with minSnapshotsToKeep of 5 and maxSnapshotAgeHours of 240. For the job status call, expect entries per maintenance type; a fresh table may show no completed run yet, since the jobs run on S3's schedule rather than on demand. If a snapshot management job shows FAILED, the docs say to look for user-defined tags, branches, or the two Iceberg retention properties listed above. You can also confirm in the S3 console under Table buckets that your bucket, namespace, and table exist.

When it breaks

GENERIC_INTERNAL_ERROR with invalid table or column names: you used a capital letter in a table, column, or namespace definition. Lowercase everything and recreate the table. The docs are explicit that such tables are not supported by Lake Formation or the Glue Data Catalog.

Athena cannot find the catalog or the Catalog and Database fields are empty: the integration step did not happen in this Region. Re-run the aws glue get-catalog --catalog-id s3tablescatalog check, and remember that creating the bucket through the CLI or an SDK does not integrate it for you, while the console flow does.

Insufficient permissions to execute the query. Principal does not have any privilege on specified resource: a Lake Formation grant is missing. The docs say to grant yourself permissions on the table, or, for the Iceberg cannot access the requested resource variant, on the table bucket catalog and the namespace without naming a table.

Snapshots never expire: check for user-defined tags, branches, or the Iceberg retention properties. Any one of them makes snapshot management fail for the whole table, and get-table-maintenance-job-status shows the cause. Unset the properties with ALTER TABLE ... UNSET TBLPROPERTIES.

A deletion error on the bucket: a table bucket cannot be deleted while it still holds namespaces or tables. That is the next section.

Cleanup

The docs require deleting all namespaces and tables before the table bucket, and bucket deletion is permanent. The delete-table and delete-namespace calls are not spelled out on the delete page I read, so check aws s3tables delete-table help if a parameter name differs from what follows.

s3t.delete_table(tableBucketARN=arn, namespace=NAMESPACE, name=TABLE)
s3t.delete_namespace(tableBucketARN=arn, namespace=NAMESPACE)
s3t.delete_table_bucket(tableBucketARN=arn)
print(s3t.list_table_buckets()["tableBuckets"])

The last line should no longer list your bucket. Then delete the Athena results prefix in your general purpose bucket, and, if you created s3tablescatalog only for this exercise and nothing else uses it, remove the catalog with the Glue console or aws glue delete-catalog --catalog-id s3tablescatalog. Leave it in place if other table buckets depend on it. Check the S3 and Glue lines of your next bill to confirm nothing is still metering.

Where to take it next

Easiest: flip compaction to the sort strategy. It needs a sort order in the Iceberg table properties and the s3tables:GetTableData permission, and the docs say it may cost more than binpack, so compare the compaction line item over a week before you keep it.

Medium: connect Spark or another engine through the Glue Iceberg REST endpoint, which the docs cover in a separate topic, and confirm two engines read the same snapshot without copying data.

Harder: replace the ten tiny INSERTs with a buffered writer that commits once a minute, then measure how the object count, and with it the monitoring fee, changes. The question worth asking before you adopt any managed table store is the same one: which chores did it really remove, and which ones did it quietly rename into a billing dimension?

Sources