What it actually is
EFS is a managed NFSv4 file system. You create it, AWS gives you a DNS name per Availability Zone, and any Linux client in your VPC can mount it and treat it like a local directory: open, seek, append, lock, rename. There is no volume to provision and no capacity to pick. It grows and shrinks as you write and delete, and you pay for what is stored.
The mental model the docs bury: EFS is not a disk and not a bucket. A disk (EBS) belongs to one instance in one AZ. A bucket (S3) has no directories, no rename, no locks, no partial overwrite. EFS is the thing in between. It gives many machines the same POSIX tree at the same time, and you pay for that in latency: roughly 1 ms for reads and 2.7 ms for writes in the best case, per the performance documentation, against sub-millisecond for a local volume.
Everything below follows from that trade. You use EFS when the thing you need is not speed, it is shared state with real file semantics.
The model
There are five primitives and you should know all of them, because most EFS outages are a misunderstanding of one.
The file system is the container. It has a performance mode (General Purpose or Max I/O, chosen at creation and permanent), a throughput mode (Elastic, Provisioned or Bursting), a storage class policy, and an encryption setting that you can only choose at creation.
A mount target is an ENI with an IP address inside one of your subnets. You get at most one per Availability Zone, and each one carries up to five security groups. Clients do not talk to "EFS", they talk to the mount target in their own AZ. A client in an AZ without a mount target cannot mount through that AZ.
An access point is an application-specific entry into the file system. It forces a POSIX identity (uid, gid, optional secondary gids) onto every request that comes through it and pins the client to a root directory. If the directory does not exist, the access point can create it with the owner and permissions you give it. Lambda requires access points. Everything else should use them anyway, because they replace "which uid is this container running as" with a setting you control.
The file system policy is an IAM resource policy that can require TLS, require IAM authorization and restrict clients to access points. Without a user-configured policy, the default grants full access to any client that can reach a mount target. That default surprises people. The security group is your only wall until you write a policy.
Lifecycle management moves files between storage classes by last access: Standard, Infrequent Access (IA) and Archive. One Zone file systems keep a single copy in a single AZ and cost less.
What AWS operates: the storage fleet, replication across AZs for Regional file systems, capacity, patching. What you operate: the network path, the security groups, the POSIX permissions inside the tree, and your choice of throughput mode, which is the decision that matters most.
Performance numbers worth memorizing
From the EFS performance documentation, for Regional General Purpose file systems in Elastic mode: about 1 ms read latency, about 2.7 ms write latency, up to 500,000 write IOPS, and per-file-system throughput that varies by Region. Per client, combined read and write throughput is capped at 1,500 MiBps only if you mount with amazon-efs-utils v2.0 or later or the EFS CSI driver. Any other client is limited to 500 MiBps. A big instance running a stock mount -t nfs4 will not use the throughput you are paying for.
One Zone file systems write in about 1.6 ms. The documentation's One Zone throughput row is internally inconsistent between the table and its footnote, so check the page before you size a One Zone system.
When to use it, when not to
You need | EFS | EBS | S3 / S3 Files | FSx |
|---|---|---|---|---|
One filesystem mounted by many hosts at once | Yes, thousands of connections | One instance (Multi-Attach only on specific volume types) | S3 Files: yes | Yes |
Capacity without planning | Yes | No, you size the volume | Effectively unlimited | You size it |
Sub-millisecond latency | No, about 1 ms reads | Yes | No | Depends on type |
Windows SMB, Lustre, ONTAP, ZFS features | No | No | No | Yes |
Lambda or Fargate persistent storage | Yes, native | Not for Lambda | S3 Files from Lambda | Not native |
Cheapest per GB for cold data | IA and Archive tiers | No | Yes | No |
The newest row deserves a paragraph. In April 2026 AWS announced Amazon S3 Files, which mounts an S3 bucket as a file system. AWS says it is built using EFS, and the data stays in the bucket. If your data already lives in S3 and the problem is that some tool wants a path, S3 Files is the first thing to try. Independent testing by a Japanese AWS consultancy found it close to EFS for large sequential I/O and slower for many small files, with changes taking tens of seconds to a minute to propagate between the file view and the bucket. If your workload writes small files and needs them visible to other clients now, EFS is still the better fit. Lambda can mount either EFS or S3 Files, but not both on the same function.
Do not use EFS for a database's data directory. Do not use it as a build cache with a million tiny files if you can avoid it. Do not use it when one instance needs the disk and nothing else does, because EBS is faster and cheaper there.
What it costs
Four dimensions, and one of them is the trap.
Storage is per GB-month by class. Throughput is the second: Elastic mode bills per GB actually read or written, Provisioned mode bills for throughput you reserve above the baseline included with your Standard storage, and Bursting has no separate throughput charge. Third, IA and Archive charge per GB for reads and for moving data between tiers. Fourth, cross-AZ and cross-Region data transfer, listed at $0.01/GB for cross-AZ on the AWS pricing page.
The AWS pricing page did not render its per-region rate table when I fetched it, so the figures below come from a June 2026 third-party summary of us-east-1 list prices. Treat them as a placeholder and confirm in the AWS Pricing Calculator: Standard about $0.30 per GB-month, Standard-IA about $0.025, One Zone about $0.16, Elastic throughput about $0.03 per GB read and $0.06 per GB written, Provisioned about $6 per MBps-month.
The trap is Elastic throughput on a chatty workload. Storage looks cheap on a small file system, and then a job that rewrites a 50 GB dataset every hour writes 1.2 TB a day. At the write rate above that is roughly 72 dollars per day in throughput charges on a file system that stores 50 GB. Run the arithmetic for your workload before you pick the default. If the average-to-peak throughput ratio is high (the documentation draws the line at 5 percent), Provisioned is cheaper. If it is spiky, Elastic wins.
Free tier: 5 GB of Standard storage on a Regional file system for 12 months, One Zone excluded, per the pricing page.
Following along in this article costs a few cents: a few kilobytes of Standard storage and a handful of metered requests for an hour.
The limits that bite
Fixed limits, from the EFS quotas page: 25,000 connections per file system, one mount target per AZ, five security groups per mount target, 255-byte file names, 177 hard links per file, 512 locks on a single file across all clients, and a 47.9 TiB maximum file size. Adjustable: 1,000 file systems per Region and 10,000 access points per file system.
The client-side limits are the ones that actually page you. A Linux NFS client can hold 65,536 open files per instance and 65,536 locks per mount connection. Exceed those and the kernel reports "Disk quota exceeded", which sends people looking for a disk quota that does not exist.
Performance mode is permanent. Choose General Purpose unless you have measured a need for Max I/O, and note that Max I/O is not supported with Elastic throughput or on One Zone.
The throughput-mode gotcha
Bursting mode ties throughput to the amount of Standard data you store: a baseline of 50 MiBps per TiB, with burst credits that cap at 2.1 TiB per TiB stored. A 20 GiB file system has a baseline of about 1 MiBps. It looks fine in testing because the credit pool starts full, then it throttles in production after the credits drain. This is the classic "EFS got slow after two weeks" report. AWS now recommends moving a throttled Bursting file system to Elastic or Provisioned, and recommends Elastic as the default.
Provisioned has its own trap: for 24 hours after you switch to it or change the amount, you cannot lower the amount and you cannot switch to Elastic or Bursting. Raise it only when you mean it.
Set the mode explicitly in code. Do not rely on a default that differs between the console and an API client.
Build it: two Lambda functions sharing one directory
The deliverable: a file system with two mount targets, one access point, and two Lambda functions that mount the same path. The writer appends a line, the reader sees it. That is the whole point of EFS in forty lines of handler code.
Honest scope note: the code follows the current AWS documentation, but I did not run it against an account while writing this. Run it in a scratch account and read the output before trusting it.
Prerequisites and IAM
You need Python 3.10 or later with boto3, credentials for a scratch account, and a default VPC with subnets in at least two AZs. The identity running the script needs these permissions:
{
"Version": "2012-10-17",
"Statement": [{
"Effect": "Allow",
"Action": [
"ec2:DescribeVpcs", "ec2:DescribeSubnets",
"ec2:CreateSecurityGroup", "ec2:DeleteSecurityGroup",
"ec2:AuthorizeSecurityGroupIngress", "ec2:DescribeSecurityGroups",
"ec2:CreateNetworkInterface", "ec2:DescribeNetworkInterfaces",
"elasticfilesystem:CreateFileSystem", "elasticfilesystem:DeleteFileSystem",
"elasticfilesystem:DescribeFileSystems",
"elasticfilesystem:CreateMountTarget", "elasticfilesystem:DeleteMountTarget",
"elasticfilesystem:DescribeMountTargets",
"elasticfilesystem:CreateAccessPoint", "elasticfilesystem:DeleteAccessPoint",
"elasticfilesystem:DescribeAccessPoints",
"iam:CreateRole", "iam:DeleteRole", "iam:AttachRolePolicy",
"iam:DetachRolePolicy", "iam:PassRole",
"lambda:CreateFunction", "lambda:DeleteFunction",
"lambda:GetFunction", "lambda:InvokeFunction"
],
"Resource": "*"
}]
}Scope the Resource down for anything beyond a scratch account. The Lambda execution role gets two managed policies: AWSLambdaVPCAccessExecutionRole for the ENIs, and AmazonElasticFileSystemClientReadWriteAccess, which holds elasticfilesystem:ClientMount and elasticfilesystem:ClientWrite. Per the Lambda documentation, those EFS permissions are only enforced when the file system has a user-configured policy, but granting them is the safe default. The user who configures the function also needs elasticfilesystem:DescribeMountTargets, because Lambda uses the caller's permissions to verify mount targets.
The script
import boto3, io, json, time, zipfile
REGION = "us-east-1"
NAME = "sl-efs-lab"
ec2 = boto3.client("ec2", region_name=REGION)
efs = boto3.client("efs", region_name=REGION)
iam = boto3.client("iam")
lam = boto3.client("lambda", region_name=REGION)
# 1. Network: default VPC, one subnet in each of two AZs
vpc_id = ec2.describe_vpcs(
Filters=[{"Name": "isDefault", "Values": ["true"]}]
)["Vpcs"][0]["VpcId"]
by_az = {}
for s in ec2.describe_subnets(
Filters=[{"Name": "vpc-id", "Values": [vpc_id]}])["Subnets"]:
by_az.setdefault(s["AvailabilityZone"], s["SubnetId"])
subnets = list(by_az.values())[:2]
# One security group, NFS (2049) allowed from itself. Lambda and the
# mount targets both use it.
sg = ec2.create_security_group(
GroupName=NAME, Description="EFS lab", VpcId=vpc_id)["GroupId"]
ec2.authorize_security_group_ingress(
GroupId=sg,
IpPermissions=[{
"IpProtocol": "tcp", "FromPort": 2049, "ToPort": 2049,
"UserIdGroupPairs": [{"GroupId": sg}]}])
# 2. File system: explicit modes, encrypted
fs = efs.create_file_system(
CreationToken=NAME,
PerformanceMode="generalPurpose",
ThroughputMode="elastic",
Encrypted=True,
Tags=[{"Key": "Name", "Value": NAME}])["FileSystemId"]
def wait(fn, ok, what):
for _ in range(60):
if ok(fn()):
return
time.sleep(5)
raise TimeoutError(what)
wait(lambda: efs.describe_file_systems(FileSystemId=fs)["FileSystems"][0]["LifeCycleState"],
lambda st: st == "available", "file system")
# 3. Mount targets, one per AZ
for sn in subnets:
efs.create_mount_target(FileSystemId=fs, SubnetId=sn, SecurityGroups=[sg])
wait(lambda: [m["LifeCycleState"] for m in
efs.describe_mount_targets(FileSystemId=fs)["MountTargets"]],
lambda sts: len(sts) == len(subnets) and all(s == "available" for s in sts),
"mount targets")
# 4. Access point: forces uid/gid 1001 and roots clients at /lab
ap = efs.create_access_point(
FileSystemId=fs,
PosixUser={"Uid": 1001, "Gid": 1001},
RootDirectory={"Path": "/lab", "CreationInfo": {
"OwnerUid": 1001, "OwnerGid": 1001, "Permissions": "755"}},
Tags=[{"Key": "Name", "Value": NAME}])
ap_arn, ap_id = ap["AccessPointArn"], ap["AccessPointId"]
wait(lambda: efs.describe_access_points(AccessPointId=ap_id)["AccessPoints"][0]["LifeCycleState"],
lambda st: st == "available", "access point")
# 5. Execution role
role = iam.create_role(
RoleName=NAME,
AssumeRolePolicyDocument=json.dumps({
"Version": "2012-10-17",
"Statement": [{"Effect": "Allow",
"Principal": {"Service": "lambda.amazonaws.com"},
"Action": "sts:AssumeRole"}]}))["Role"]["Arn"]
for p in ("service-role/AWSLambdaVPCAccessExecutionRole",
"AmazonElasticFileSystemClientReadWriteAccess"):
iam.attach_role_policy(RoleName=NAME, PolicyArn=f"arn:aws:iam::aws:policy/{p}")
time.sleep(10) # IAM propagation
# 6. Handler code
CODE = '''
import fcntl, os, time
PATH = "/mnt/efs/log.txt"
def handler(event, context):
if event.get("op") == "write":
with open(PATH, "a") as f:
fcntl.flock(f, fcntl.LOCK_EX)
f.write(f"{context.function_name} {time.time():.0f} {event.get('msg','')}\\n")
fcntl.flock(f, fcntl.LOCK_UN)
with open(PATH) as f:
return {"uid": os.getuid(), "lines": f.read().splitlines()}
'''
buf = io.BytesIO()
with zipfile.ZipFile(buf, "w") as z:
z.writestr("app.py", CODE)
# 7. Two functions, same mount
for fn in (f"{NAME}-writer", f"{NAME}-reader"):
for attempt in range(6):
try:
lam.create_function(
FunctionName=fn, Runtime="python3.13", Role=role,
Handler="app.handler", Code={"ZipFile": buf.getvalue()},
Timeout=30,
VpcConfig={"SubnetIds": subnets, "SecurityGroupIds": [sg]},
FileSystemConfigs=[{"Arn": ap_arn, "LocalMountPath": "/mnt/efs"}])
break
except lam.exceptions.InvalidParameterValueException:
time.sleep(10) # role not assumable yet
lam.get_waiter("function_active_v2").wait(FunctionName=fn)
print(json.dumps({"fs": fs, "sg": sg, "ap": ap_id, "role": NAME}))Two choices in there are worth a sentence each. The security group references itself, so the only thing that can reach NFS is something wearing the same group. And the handler takes an exclusive flock before appending, because two concurrent Lambda environments writing to one file will interleave bytes otherwise. EFS supports NFS locking, and it is what the 512-locks-per-file limit is counting.
Verification
def call(fn, payload):
r = lam.invoke(FunctionName=f"{NAME}-{fn}", Payload=json.dumps(payload))
return json.loads(r["Payload"].read())
print(call("writer", {"op": "write", "msg": "hello from writer"}))
print(call("reader", {"op": "read"}))The writer returns one line. The reader, a different function with a different execution environment, returns the same line, and the reported uid is 1001 on both. That uid is the access point imposing the POSIX identity you configured, which is the reason to use access points.
For the control-plane view:
aws efs describe-file-systems --file-system-id <fs-id> \
--query 'FileSystems[0].[ThroughputMode,PerformanceMode,Encrypted,SizeInBytes.Value]'
aws efs describe-mount-targets --file-system-id <fs-id> \
--query 'MountTargets[].[AvailabilityZoneName,LifeCycleState,IpAddress]'Expect elastic, generalPurpose, true, and two mount targets in available. If you want to mount from an EC2 instance as well, the documented mount-helper form is mount -t efs -o tls,iam,accesspoint=<fsap-id> <fs-id>: /mnt/efs, which needs amazon-efs-utils installed. ECS tasks and EKS pods reach the same access point through their own EFS volume configuration, and EKS uses the EFS CSI driver, which is also what unlocks the 1,500 MiBps per-client ceiling.
When it breaks
Mount timeouts on Lambda surface as EFSMountTimeoutException, EFSMountConnectivityException and EFSMountFailureException, all documented in the Lambda API. The cause is almost always the network path. Check, in this order: port 2049 allowed in the security group on both the mount target and the function, a mount target exists in the AZ of every subnet you attached to the function, the function and the file system are in the same VPC, and the local mount path begins with /mnt/.
If Lambda in the VPC suddenly cannot reach other AWS APIs, that is not EFS. A VPC-attached function loses its default internet path, and the Lambda docs tell you to use a NAT gateway or VPC endpoints. Intermittent TCP timeouts on a subnet with a network ACL are the other known cause: Lambda uses ephemeral ports 1024 to 65535, so allow them.
"Permission denied" on a file you just created is POSIX, not IAM. The access point root directory is created with the CreationInfo you gave it only the first time. If you change the access point later, existing directories keep their old owner and mode. Fix them from a client mounted without the access point.
"Disk quota exceeded" is one of the three 65,536 client limits above, usually open files. lsof <mount-path> and lslocks tell you which. "I/O error" is the same family, or a deleted KMS key, which cannot be recovered.
Slow after weeks, fast at launch, on a file system with little data: Bursting mode with an exhausted credit pool. Watch the PermittedThroughput and BurstCreditBalance metrics, and change the mode. For Elastic, watch MeteredIOBytes because that is the number your bill follows. PercentIOLimit tells you when General Purpose is the ceiling.
Cleanup
Order matters. Functions first, then the access point and mount targets, then the file system, then the security group and role.
for fn in (f"{NAME}-writer", f"{NAME}-reader"):
lam.delete_function(FunctionName=fn)
efs.delete_access_point(AccessPointId=ap_id)
for m in efs.describe_mount_targets(FileSystemId=fs)["MountTargets"]:
efs.delete_mount_target(MountTargetId=m["MountTargetId"])
wait(lambda: efs.describe_mount_targets(FileSystemId=fs)["MountTargets"],
lambda mts: len(mts) == 0, "mount targets deleting")
efs.delete_file_system(FileSystemId=fs)
for p in ("service-role/AWSLambdaVPCAccessExecutionRole",
"AmazonElasticFileSystemClientReadWriteAccess"):
iam.detach_role_policy(RoleName=NAME, PolicyArn=f"arn:aws:iam::aws:policy/{p}")
iam.delete_role(RoleName=NAME)
# The security group can refuse deletion while Lambda releases its ENIs.
for _ in range(30):
try:
ec2.delete_security_group(GroupId=sg)
break
except Exception:
time.sleep(30)Confirm the bill is zero by checking that nothing is left:
aws efs describe-file-systems --query 'FileSystems[?Name==`sl-efs-lab`]'
aws lambda list-functions --query 'Functions[?starts_with(FunctionName, `sl-efs-lab`)]'Both should print an empty list. The access point has no charge of its own, but the file system does until it is deleted.
The short version
EFS is shared POSIX storage with network latency. The security group is your perimeter, the access point is your identity, and the throughput mode is your invoice. Pick Elastic when load is spiky, Provisioned when it is steady and high, and never leave a small Bursting file system in production wondering why it is slow.

