
AI agent sandbox platforms become useful when an agent needs to run code it creates, install dependencies, or work with temporary files. I would choose one by the workload, the state that must survive, and the credentials the code can reach.
For a document assistant, that decision might arrive when an uploaded spreadsheet needs a new calculation. For a coding agent, it might arrive when a repository needs an isolated environment where tests can run.
This guide compares E2B, Daytona, Modal, Vercel Sandbox, and Cloudflare Sandboxes using their production documentation, checked October 9, 2026. The recommendations are documentation-based judgments; I have not run a controlled performance benchmark across these services.
The purchasing worksheet follows one document-analysis job from creation through cleanup. It separates resource units, restoration behavior, and recurring storage, because those details often matter more than a headline launch time.
When do you need AI agent sandbox platforms?
Start with the action you want the agent to take. If it always retrieves an order, checks a calendar, or submits a known form, a fixed API integration may be enough.
A sandbox earns its place when the program changes with the task. Examples include generating Python to reconcile unfamiliar CSV columns, compiling a proposed patch, or converting a document with a command-line package.
I would keep predictable operations behind a narrow function with validated arguments. Running arbitrary code to perform a fixed database lookup adds an execution environment, permissions, cleanup, and another bill to maintain.
Pickaxe's Actions documentation describes connecting an assistant to external operations, including connected tools and workflows. That is a practical starting point when the assistant needs a defined operation rather than a general computer.
Once generated code becomes necessary, separate the agent's decision loop from the environment that executes it. Our guide to the AI agent harness explains the surrounding controls: tools, context, retries, and evaluation.
A sandbox protects a boundary around execution. Your application still decides which customer owns the job, which files enter it, what destinations are allowed, and whether its output can trigger a consequential action.
Compare AI agent sandbox platforms by workload
I would shortlist these five providers by the system they must fit into. A developer's existing stack can be a reasonable tie-breaker after the required runtime and access rules are clear.
| Platform | My starting use case | State question to resolve | Billing question to resolve |
|---|---|---|---|
| E2B | Generated code with resumable Linux environments | Full memory or filesystem-only restoration? | Running resource seconds plus the required plan |
| Daytona | Workspaces with an explicit container or VM choice | Does the chosen type support pause, or only stop/start? | Reserved resources, retained disk, and lifecycle transitions |
| Modal | Code execution alongside an existing Modal application | Which runtime and snapshot type does the job require? | Sandbox-specific rates and the greater of requested or used resources |
| Vercel Sandbox | Code execution attached to a Vercel application | Persistent filesystem or a disposable session? | Active CPU, provisioned memory, region, and retained snapshots |
| Cloudflare Sandboxes | A Workers application needing Linux or isolated JavaScript | Container lifetime or a Dynamic Worker? | Selected instance, Workers, Durable Objects, and storage |
These are scoped recommendations, not measured rankings. The provider sections below link the documentation behind each distinction, and the worksheet later keeps the different meters separate.
Before requesting a trial, write down your runtime, required packages, maximum input size, expected job duration, and acceptable recovery behavior. A Python data task and a JavaScript tool interpreter can require very different environments.
Also decide whether you need a new environment for each job or a named workspace that lasts across jobs. Persistence can save setup work, but it introduces retained data, expiration, deletion, and tenant ownership questions.
1. E2B: a shortlist for resumable code execution
E2B is worth investigating when generated code needs a Linux environment and you want an explicit choice between restoring running state and restoring files.

Its persistence documentation distinguishes the default pause behavior, which preserves filesystem and memory, from a filesystem-only mode. Paused environments remain retained until explicitly killed, according to that documentation.
That distinction matters for a notebook-style task. Keeping an interpreter's loaded variables can let a conversation continue, while restoring only files means the application must rebuild the process and any in-memory context.
Choose what must survive before choosing a pause policy. A completed document job may need only its exported report; preserving the entire running environment can retain more customer data than the product needs.
E2B documents outbound internet access as enabled by default. Its internet-access controls let you block access and configure destinations; adding an allowed host is not a substitute for a complete deny policy.
Its secrets documentation describes injecting stored credentials into outbound HTTPS requests through a proxy. The application should keep the provider management key outside the environment and restrict the destination and credential's authority.
I would ask a pilot to demonstrate three behaviors: resume the required state, reject an unapproved destination, and delete an environment after the application has retrieved its accepted outputs. A successful code cell alone would not satisfy that acceptance test.
On the commercial side, E2B pricing separates a free Hobby base from Pro at $150 per month, with usage charged separately. Its published resource rates are $0.000014 per vCPU-second and $0.0000045 per GiB-second of memory; billing documentation explains when running charges stop.
The plan required by your session duration and concurrency belongs in the estimate. I would not present an attractive compute calculation as the entire invoice, especially when a production workload needs a paid base plan.
2. Daytona: choose the environment type first
Daytona belongs on the shortlist when workspace persistence is central and you want to choose among documented environment types instead of assuming every sandbox behaves alike.

The isolation documentation distinguishes container, VM, and GPU environments. Containers share the host kernel boundary; VMs have their own kernel, and features such as pause/resume depend on the selected type.
Its persistence guide describes persistent workspaces, but ordinary container stop/start clears memory while preserving files and configuration. VM pause can preserve running state; a GPU environment has different lifecycle behavior.
Persistent files do not prove a persistent interpreter. If the assistant expects a loaded dataframe to survive, test that expectation against the actual environment type rather than the provider's broad persistence description.
Daytona's billing lifecycle also makes a useful distinction: stopped CPU and memory charges end, while retained disk can continue billing. Archival, deletion, and separately retained snapshots have different consequences.
Its pricing page lists $0.0504 per vCPU-hour, $0.0162 per GiB-hour of memory, and $0.000108 per GiB-hour of disk after the included storage allowance, with usage billed per second. Check the selected disk configuration and applicable allowance before extending the compute example later in this guide.
For access control, network limits and their defaults depend on the account tier. The secrets guide describes replacing credential placeholders at an outbound proxy, which can keep plaintext out of the guest environment.
I would evaluate Daytona with both a completed job and an interrupted job. Confirm what stop/start restores, what the chosen pause operation restores, and what remains chargeable after each path.
The operational owner should also know how to archive and delete workspaces. A product that creates named environments without a retirement policy can accumulate old customer files even when no code is running.
3. Modal: check the sandbox rate and runtime
Modal is a natural candidate when your application already runs there or you need code execution alongside its broader compute platform. I would still review the sandbox-specific configuration and invoice separately.

Its current sandbox networking documentation covers gVisor and VM runtimes. The runtime choice changes the execution boundary, so an older comparison describing only one runtime is incomplete.
Public outbound access is enabled by default, according to that guide. Network blocking and allowlists are available, while particular controls have their own scope and maturity; check the exact runtime and SDK you intend to use.
The snapshot documentation separates filesystem and memory snapshots. It documents default retention of 30 days for filesystem snapshots and seven days for memory snapshots, with configuration and SDK-version qualifications.
I would treat those expirations as product behavior. A customer returning to a workspace after a long absence needs an intentional recovery path, not an unexplained missing environment.
Read the Sandbox and Notebooks section of Modal's pricing page. Its published sandbox rates are $0.00003942 per physical CPU core-second and $0.00000667 per GiB-second of memory; the general Functions rates are a different offering.
The resource documentation says charging follows the greater of requested and used resources. It also explains that a physical CPU core corresponds to two vCPUs, so copying a vCPU figure into a core-based calculator changes the comparison.
Modal's sandbox secret injection uses an outbound policy and currently carries experimental qualifications. A secret put directly into an environment variable is readable by the code; that is a different exposure model.
My pilot would pin the runtime, put an upper bound on resource use, and demonstrate snapshot expiry handling. That gives the team a concrete estimate and a support procedure instead of relying on the platform's general compute reputation.
4. Vercel Sandbox: account for active CPU and persistence
Vercel Sandbox is worth a look when a Vercel application needs isolated code execution and the team wants to keep the application and its execution service in the same operational stack.

The Sandbox documentation describes Firecracker microVMs. Its current persistent-sandbox guide says filesystem persistence is enabled by default and can be disabled for disposable work.
A resumed filesystem is a new running session. It does not restore the previous process's live memory, so an application must restart its interpreter, server, or job using the files that survived.
Vercel's firewall guide documents allow-all as the default, plus deny-all and custom policies. Review allowed destinations and any credential transformation together; a transformation rule alone does not define the complete network policy.
Vercel's active-CPU meter changes the worksheet. The pricing documentation lists $0.128 per active CPU-hour and $0.0212 per provisioned GB-hour of memory for the default iad1 region, plus $0.60 per million sandbox creations.
CPU time excludes waiting for I/O, while provisioned memory remains billable during the running session, subject to the documented minimum. Regional rates, plan allowances, and shared usage credits also affect the invoice, so the later arithmetic explicitly identifies its assumptions.
There is another cleanup detail to catch: snapshot documentation says deleting a sandbox does not delete its retained snapshots. Filesystem persistence can therefore continue generating storage charges until snapshots expire or are deleted.
I would compare a CPU-heavy report with one that spends much of its time waiting on an approved external service. Those jobs may have the same wall-clock duration and memory request but different active-CPU usage.
The pilot should inspect the actual usage record rather than estimating activity from the model's tool log. It should also demonstrate disposable mode and complete deletion of any retained state created by the persistent mode.
Start with one useful operation
Prototype the customer task, then add the execution it requires.
5. Cloudflare Sandboxes: distinguish Linux from Dynamic Workers
Cloudflare Sandboxes deserves consideration when the application already uses Workers and needs a choice between a Linux environment and a narrower JavaScript execution surface.

Its sandbox concepts distinguish Linux containers running in Firecracker microVMs from Dynamic Workers. Dynamic Workers use the Workers execution model and cannot be treated as a general Linux environment for arbitrary binaries and native packages.
That narrower surface may be exactly what a JavaScript tool needs. A task that requires Python packages, shell commands, or a native conversion utility needs the Linux option instead.
The current security guide says Linux container internet access is disabled by default and explains outbound handlers and their scope. Dynamic Workers have a different network model; the guide describes using globalOutbound: null when outbound access must be blocked.
One product name can contain different security defaults. Keep the environment type visible in the design review, and check that the configured policy covers every protocol the task can use.
The lifetime documentation separates a stable Durable Object identity from its Linux instance. Activity such as requests, streams, and open connections affects lifetime; a background process by itself does not establish the same activity contract.
The SDK 1.0 migration guide changes orchestration responsibilities, including ownership of the Durable Object class and explicit lifecycle decisions. Pin the version before adapting a tutorial written for an older SDK.
Cloudflare's container pricing bills provisioned memory and disk while active and CPU usage separately. Published instance shapes differ from the two-vCPU, four-GiB example below, so I do not assign Cloudflare an apparently equivalent dollar total.
I would calculate the selected instance together with Workers, Durable Objects, retained storage, and any relevant transfer. Then test whether the chosen lifetime settings actually release resources after a job has finished.
Treat network access and secrets as separate decisions
Blocking an unapproved host is useful, but an approved host can still accept an operation your customer did not authorize. Network policy and business permission should be reviewed independently.
A document analyzer might need access to an object store containing that customer's uploads. It probably does not need permission to list every customer folder, change account settings, or call a payment endpoint.
An external credential broker can keep a raw token out of the guest. It does not automatically make every request through that broker acceptable, and the destination itself receives the credential.
The Cloudflare security guide states that all code within one Linux sandbox shares its trust boundary. Files, environment variables, process arguments, and other secrets placed there are reachable by code in that environment.
I would keep management credentials in the trusted application, issue the smallest practical workload authority, and avoid placing a full application environment file into the guest. Our AI agent security risks guide offers a broader checklist for the surrounding application.
Also separate viewing a generated artifact from approving a write to an external system. A completed analysis can enter a review queue; our human approval workflow guide explains where an approval step belongs.
Define what persistent means in your product
Customers usually care whether they can continue their work. Infrastructure documentation may mean a saved disk, a saved process, a reusable identity, an external volume, or some combination of these.
Write the promise in customer terms first: “Your uploaded files and accepted reports remain available.” Then identify which service stores those artifacts and how the application retrieves them.
If the promise includes a live interpreter, check memory restoration for the chosen runtime. If it includes only outputs, a fresh environment with retrieved files may be simpler to operate.
A pause button also needs a retention policy. E2B, Daytona, Modal, and Vercel describe different persistence and billing rules in their linked guides; none of those terms should silently become your product's indefinite storage commitment.
Delete by artifact type, not by one broad label. Track the running environment, snapshots, external volumes, customer uploads, and accepted outputs as separate resources with owners.
For each resource, record who can read it, when it expires, and what happens when a customer deletes their project. Test that deletion against the storage service rather than relying only on the absence of a running process.
Measure time to a useful result
A startup claim can tell you how quickly infrastructure appears. Your customer waits for an accepted result, including dependency readiness, file transfer, code execution, output validation, and recovery from an error.
I would time that whole path under the configuration I intend to buy. Record a fresh environment, a reused environment, and a resumed environment separately.
Use the same source files and acceptance rules for each candidate. Pin package versions and prepare an image where appropriate so that a surprise dependency download does not dominate one provider's result.
Measure concurrent arrivals as well as an isolated request. The useful question is whether the product meets its response target during the workload it expects, within the chosen account's limits.
Do not report a vendor's marketing figure as your own observation. This guide deliberately makes no cross-provider latency claim; a pilot should supply that evidence for your runtime, region, image, and job.
Record failures too: incomplete output, missing dependencies, timeouts, and unsuccessful restores. A fast result that fails validation does not belong in the successful-job average.
Build a cost worksheet with the correct units
Use a separate row for every meter before comparing totals. A vCPU, a physical CPU core, a GB, a GiB, active CPU time, and provisioned wall-clock time are different inputs.
The following is hypothetical arithmetic, not a benchmark or invoice quote. Assume 1,000 document-analysis jobs, each lasting 300 seconds; E2B and Daytona request two vCPUs and four GiB, and Modal requests one physical core and four GiB without usage exceeding that request.
For Vercel, assume two vCPUs, four GB, the default iad1 region, and average CPU activity of 20% during each session. The different memory unit and active-CPU assumption make this a scenario comparison, not an equal-resource price ranking.
| Platform | Compute calculation for the scenario | Illustrative subtotal | Charges and qualifications outside this subtotal |
|---|---|---|---|
| E2B | 1,000 × 300 × (2 × $0.000014 + 4 × $0.0000045) | $13.80 | Required plan, any applicable additions, and external services. Rates |
| Daytona | 1,000 × (300 / 3,600) × (2 × $0.0504 + 4 × $0.0162) | $13.80 | Disk, retained state, billable lifecycle transitions, and external services. Rates |
| Modal | 1,000 × 300 × ($0.00003942 + 4 × $0.00000667) | $19.83 | Usage above the request, plan context, credits, and external services. Sandbox rates |
| Vercel | Active CPU: 33.33 hours × $0.128; memory: 333.33 GB-hours × $0.0212; 1,000 creations | About $11.33 | Regional qualification, plan fees, allowances/credits, snapshots, and external services. Rates |
| Cloudflare | Select the actual instance and add each service meter | No equivalent subtotal | Instance shapes differ; include Workers, Durable Objects, storage, and transfer as applicable. Meter scope |
For the same Vercel allocation at 100% CPU activity throughout, the illustrative compute-and-creation subtotal becomes about $28.40 using the same Vercel rates. That change shows why assumed CPU activity belongs beside the estimate.
These subtotals exclude model tokens, OCR or other paid APIs, application hosting, and engineering time. They also do not subtract introductory or recurring credits; a credit can reduce a bill without changing the workload's underlying consumption.
Our AI agent cost guide explains the surrounding token economics. I would add this execution worksheet to that budget rather than treating the sandbox as the whole cost of an agent.
Walk one document job through the boundary
Consider a proposed client tool that accepts a CSV and a PDF, extracts relevant tables, runs a calculation, and returns a report. This is a design example, not a customer deployment or a tested performance result.
The trusted application authenticates the customer and allocates a job identifier before provisioning the environment. The caller should not be able to select another customer's workspace by changing an identifier in a request.
Next, the application validates the input types and its own size limits. It supplies only the files needed for this job, with explicit ownership, and avoids copying a shared customer directory into the guest.
Generated code runs with a deadline, resource bounds, and a defined network policy. The environment contains the approved dependencies, while credentials for the infrastructure control plane remain outside it.
The application retrieves the expected outputs, verifies their format, and records the result before cleanup. It should treat generated HTML, filenames, and other content as untrusted when presenting them to the customer.
Finally, cleanup stops or deletes the running environment and handles any snapshots separately. A durable job record says whether the report was accepted, whether it needs a retry, and whether cleanup succeeded.
Pickaxe documents Agentic Runtimes as a provisioned beta for agents needing computers, code, and other execution capabilities. That is a managed route to investigate; this infrastructure shortlist does not imply that each provider is a selectable Pickaxe backend.
Budget for retries, waiting, and retained state
Retries repeat resource consumption. In a hypothetical workload where 100 of those 1,000 jobs run once more with the same allocation and duration, the repeated execution increases that compute component by 10%; it is not an observed failure rate.
Set a retry limit and distinguish an environment failure from an input the program cannot process. Retrying an unsupported file repeatedly consumes time without making success more likely.
Waiting deserves its own row too. A session holding provisioned memory while an upstream service is slow can accrue memory charges even when an active-CPU meter registers little work, as Vercel's billing documentation explains.
Retained disk and snapshots can continue after running compute ends. Daytona billing and Vercel snapshots describe those charges; do not assume that a stopped environment produces a zero-cost resource set.
I would reconcile the application job log with provider usage during the pilot. Look for environments without a terminal job status, snapshots without an owner, and sessions that remained active after the accepted report was retrieved.
Give cleanup an accountable owner and an observable result. A failed deletion should enter an operational queue, not disappear because the customer-facing request already returned successfully.
Run a pilot with explicit acceptance criteria
Start with a small set of representative tasks that your team can inspect. Include a valid task, malformed input, an intentionally slow dependency, an interrupted process, and a request that attempts to reach a forbidden destination.
Add a pair of different customer identities. Confirm that one cannot resume, inspect, or retrieve the other's environment or saved output through the application.
For stateful work, test both successful restoration and a missing or expired snapshot. The latter needs a predictable customer message and a safe path to restart.
For billing, record requested resources, runtime, observed usage, retained artifacts, and the plan or allowance applied. Reconcile those records with the provider's usage export before extrapolating.
Our guide to testing AI agents explains how to turn expected behavior into repeatable checks. The sandbox pilot should extend those checks across execution, access, recovery, and cleanup.
Approve the workload, not just the provider account. Keep a record of the image, runtime, SDK version, network policy, resource limits, and restoration behavior that passed review.
If one of those changes, rerun the checks that depend on it. A package update or an altered allowlist can change the behavior even when the platform name stays the same.
My shortlist by the job you need to ship
For generated code that benefits from memory restoration, I would investigate E2B and Daytona's VM option first. I would make the choice from the required restore behavior, permission model, and complete monthly estimate.
For a team already deploying compute on Modal, I would include Modal early and verify its sandbox rate, resource ceiling, and snapshot retention. Existing operational familiarity helps only when the workload's controls also fit.
For a Vercel application, I would test Vercel Sandbox with its current persistence defaults and actual CPU utilization. For a Workers application, I would test Cloudflare's appropriate Linux or Dynamic Worker path with a pinned SDK version.
For a predictable business operation, I would first prototype the fixed integration. A broad execution environment becomes useful when the task requires its flexibility and the team can operate the resulting boundary.
The recommendation should come with evidence: accepted outputs, permission checks, recovery results, cleanup records, and reconciled costs. That gives the buyer something concrete to approve.
Frequently asked questions
Is an AI agent sandbox a Docker container?
Sometimes, but the term covers different execution boundaries. The provider documentation above includes containers, microVMs, gVisor, and Dynamic Workers; check the exact environment type rather than inferring it from the word sandbox.
Does a sandbox prevent prompt injection?
It can constrain what generated code reaches, but the application still interprets the model's decisions and authorizes business operations. Keep permissions, input handling, output validation, and human approval where the action requires them.
Can I keep a conversation's Python variables between jobs?
Only if the selected runtime and restoration mode preserve memory, and the retained environment remains available. Filesystem persistence alone requires the application to recreate the interpreter state from saved files or another durable record.
Which platform is cheapest?
Calculate your workload with the selected plan, runtime, resource units, region, activity pattern, and retained storage. The worksheet shows how a change in CPU utilization or lifecycle behavior can alter a subtotal; it does not establish a universal winner.
What should I build first?
Build one representative task with a clear result and a defined permission boundary. If a fixed Action meets the need, you can start a Pickaxe prototype and validate that customer flow before adding general code execution.






