How agents run code
When an agent needs to compute something, it does not guess and it does not run code on the Studio services. A coder inside the code-interpreter service writes Python and shell commands and runs them in a sandbox in your account, against files in a workspace in your account. This page describes that path at the component level: which service does what, where files live, how output is bounded, and how one user's work is kept apart from another's. The user's view of the same feature, the Files tab, deliverables and grids, is in Chat: files, previews and approvals.
Who this is for
Security engineers and AI engineers who need to know where model-written code runs and what it touches.
Before you start
Two components and one store. agent-runtime runs the conversation or workflow step and decides when code is needed. code-interpreter hosts the coder and dispatches each execution to the Lambda function of the agent's execution environment. The workspaces bucket holds every user's files.
The path of one execution
- The agent in
agent-runtimehands the coder a description of the analysis it needs. The coder is a model-driven loop incode-interpreter; its own model calls are metered ascoder_analysis, so the work it does is counted in the run's totals. - The coder writes code and submits an execution. Before a command runs, a guard classifies it: ordinary analysis runs at once; a command that would delete files outside the coder's scratch directories, write destructively to a database, send messages, or send workspace content to an outside service is held for the user's decision in Studio.
code-interpreterinvokes the Lambda function of the agent's execution environment,alphaagent-env-<environment id>, with the code, the workspace prefix to mirror and the agent's identity. One execution is one invocation.- The handler in the function mirrors the requested project folders down from the workspaces bucket into a fresh working directory, injects the agent's connector variables into the child's environment, runs the code, uploads only the files the run created or changed, and removes the directory. Its AWS calls run on a session narrowed to this user's prefix and this agent's secret; see Sandbox isolation.
- Standard output and error return to the coder, truncated; the coder decides what to do next, and the agent reports the result the code produced.
The workspace
Every user has one workspace under users/<user id>/ in the workspaces bucket, and every conversation and every workflow has a project inside it:
users/<user id>/projects/conversations/<conversation>/
inputs/ files dropped into the chat
data/ tables the coder registered
scripts/ code the coder wrote and reuses
outputs/ deliverables
users/<user id>/projects/workflows/<workflow>/
inputs/ the workflow's inputs
runs/<execution id>/ one project root per run, laid out the same way
Every file a workflow step writes under outputs/ is a deliverable, whatever its type, attributed to the step that wrote it; outputs/result.json is the run's machine-readable result and is served by name. The handler uploads only what a run created or changed; an upload it could not complete is reported, not dropped: the coder's result names every refused path ("N uploads refused: path: error"), and when nothing landed the run's error output also ends with "[HANDLER] N file(s) could not be uploaded to S3 (...); the run's outputs are incomplete." A run started with an API key writes under users/apikey:<key id>/ and its outputs persist like any other run's.
The workspace is durable and independent of any session: files persist until a user deletes them or a retention rule expires them. Objects written under a run prefix are tagged aa-data-class=run-data at write, which is what the workspaces bucket's retention rule keys on; see Data at rest and in transit. Code the coder writes is also kept in the code-scripts bucket so a later turn can reuse it.
Sessions and history
- A coder session is the live state of one conversation's or run's coding work. It expires after four hours of inactivity; losing a session never loses files, because files live in the workspace.
- Each execution is recorded in the
alphaagent-code-executionstable with its truncated output, and expires after 30 days. - Output caps: an execution returns at most 50,000 characters of standard output to the service; the stored record keeps 50,000 characters of standard output and 10,000 of standard error; the coder itself reads at most 24,000 and 8,000. Full data files are unaffected; they stay in the workspace.
Isolation between users
- A session belongs to the user who started it. A request against another user's session answers 404, the same as a missing session, so identifiers cannot be used to probe.
- Workspace reads are confined to the caller's own prefix in a standard deployment. In a governed deployment every assigned user can read every prefix, and writes still land only in the caller's own folder.
- Each execution runs in its own working directory with its own scoped credentials; two users of the same environment never share a directory, a home or a credential.
- In a workflow, every run has its own project root,
runs/<execution id>/, under the workflow's project; the steps of one run share that root, and no run sees another run's root.
Steps: verify the path in your account
- In Lambda, open the function for one of your environments and read its recent invocations in CloudWatch Logs: one invocation per execution, each with the working directory it created and removed.
- In S3, open
alphaagent-workspaces-<account>-<region>and browseusers/<your user id>/projects/; the layout above should be visible for any conversation that ran code. - In DynamoDB, open
alphaagent-code-executions: one row per execution, withttlset 30 days out. - In the deployment's usage view in the Organisations console, group by kind and find
coder_analysis.
What you should see
- No code execution on any ECS task: the
code-interpretertask's logs show dispatches, the Lambda logs show the executions. - Run objects under
runs/<execution id>/carrying theaa-data-classtag.
Limits
- One execution runs for at most 900 seconds (the Lambda maximum). Longer work is split into further executions by the coder.
- Output returned to the model is truncated at the caps above.
- The sandbox reaches the network on TCP 443 only.
If something goes wrong
- An execution fails at once with a permission error on the workspace: the object key is outside
users/<user id>/; the session policy denies it by design. - A held command never runs: the user did not allow it, or chose to stop the run; the step's card in Studio shows the decision.
- A session "expired": the four-hour inactivity limit passed; the next turn starts a new session and the files are still there.
Build a custom environment image
For the engineer whose agents need a package or tool the prebuilt image lacks (Execution environments). The image needs python on its PATH with awslambdaric and boto3 importable, and may carry the handler pair as a fallback (handler.py · aa_env.py, placed next to the Dockerfile):
FROM python:3.12-slim
# awslambdaric 4.1.0 publishes no linux/amd64 wheel; pip would build it from source and fail without cmake
RUN pip install awslambdaric==4.0.4 boto3
COPY handler.py /opt/handler.py
COPY aa_env.py /opt/aa_env.py
Pin awslambdaric to 4.0.4. Version 4.1.0 (28 September 2026) ships wheels for arm64 only, so on the linux/amd64 build Studio requires pip falls back to the source distribution, which needs cmake and fails on a slim base; an unpinned pip install awslambdaric hits this today. On an Ubuntu base install python-is-python3 (or link /usr/bin/python to python3): the function's entry point runs python, not python3.
The entry point and the delivered handler
An ENTRYPOINT or CMD in the image is not used. Studio configures the function itself: an entry point that is a short Python bootstrap, the command handler.handler and the working directory /opt. On a cold start the bootstrap reads AA_HANDLER_URI, an s3:// prefix in your deployment's storage bucket (agent-management/env-handler/<Studio version>/), downloads the current handler.py and aa_env.py into /tmp/aa_handler, imports the handler from there and starts awslambdaric. A handler fix therefore reaches every environment, prebuilt or custom, at the Studio upgrade that carries it, with no rebuild: every upgrade points every function at the new prefix and re-probes it. If the download fails (no prefix, a denied or missing object), the bootstrap runs the image's own /opt/handler.py and records why.
Which copy is running is recorded, never silent. The environment record Studio keeps (as its Environments API returns it, not on the detail page) carries handler_version, the Studio release the function is configured for, handler_delivery, overlay for the delivered copy or baked for the image's own, and handler_delivery_reason. The function's log opens each cold start with [BOOTSTRAP] handler delivery: overlay uri: s3://..., and when the image's copy is running the coder is told so in its capability summary. In the Lambda console the function shows the entry point override under Image configuration and AA_HANDLER_URI under its variables. The pair served here is the one Studio 2.0.14 delivers; handler.py has SHA-256 7d152e01d7969ddfc8fb152bcdbf9e738b050835887e0efbb2eecaf9ee5700e2.
Add the connector drivers whenever agents have data connectors: SQLAlchemy with psycopg2-binary (PostgreSQL) or pymysql (MySQL), snowflake-connector-python[pandas] with pyarrow, fastmcp for MCP, boto3 for AWS (the prebuilt image carries this set):
RUN pip install sqlalchemy psycopg2-binary pymysql 'snowflake-connector-python[pandas]' pyarrow fastmcp httpx requests
Document and media toolkit
Agents produce PDF, PPTX, DOCX and XLSX documents and diagrams and review them visually, which needs system tools. The toolkit installs the exact set from a pinned manifest (alphaagent-env-toolkit.sh · toolkit.lock) on a Debian or Ubuntu base with apt and python3 (python:3.x-slim qualifies).
| Capability | Tools installed |
|---|---|
| Office to PDF, formula recalculation | LibreOffice (soffice) |
| PDF to page images for review | poppler (pdftoppm) |
| Markdown to DOCX and PDF | pandoc |
| Diagrams | Graphviz, Node with @mermaid-js/mermaid-cli (mmdc) and Chromium, pptxgenjs |
| Python document libraries | weasyprint and the libraries the lock's doc-libs check names |
| Fonts | The families the lock's fonts-plex and fonts-raleway checks name |
Place both files next to the Dockerfile and add, before any USER line:
COPY alphaagent-env-toolkit.sh toolkit.lock /opt/alphaagent/toolkit/
RUN chmod 755 /opt/alphaagent/toolkit/alphaagent-env-toolkit.sh && \
/opt/alphaagent/toolkit/alphaagent-env-toolkit.sh
ENV NODE_PATH=/usr/local/lib/node_modules \
PUPPETEER_CACHE_DIR=/opt/puppeteer \
PUPPETEER_EXECUTABLE_PATH=/opt/alphaagent/chrome \
BROWSER_PATH=/opt/alphaagent/chrome-kaleido
BROWSER_PATH must point at /opt/alphaagent/chrome-kaleido, the launcher the toolkit writes with the flags Chromium needs in the sandbox, not at the Chromium binary. The installer is idempotent; --apt-only, --pip-only, --npm-only, --fonts-only and --no-fonts split it across layers. It writes /opt/alphaagent/toolkit.lock, which Studio reads as the image's toolkit version.
Build, verify and push
Replace <region> and <ecr-registry> (for example 123456789012.dkr.ecr.<region>.amazonaws.com); the repository name is the one in the Create form's refusal message.
aws ecr get-login-password --region <region> | docker login --username AWS --password-stdin <ecr-registry>
docker buildx build \
--platform linux/amd64 \
--provenance=false --sbom=false \
-t <ecr-registry>/<repository>:my-agent-env-1.0 \
--push .
Run the toolkit's checks the way the sandbox runs: a non-root user that is not the image's own, a read-only root filesystem, only /tmp writable:
docker run --rm --platform linux/amd64 --user 993:990 --read-only --tmpfs /tmp:size=512m \
-e HOME=/home/agent -e AA_TOOLKIT_FORCE_BROWSER_CHECKS=1 \
<ecr-registry>/<repository>:my-agent-env-1.0 \
/opt/alphaagent/toolkit/alphaagent-env-toolkit.sh --check --lambda-like
The last stdout line is a JSON summary whose ok is true only when every check passed; the exit code follows it. --check alone runs as the invoking user without the sandbox conditions. On an Apple silicon Mac the two browser checks are skipped under emulation unless AA_TOOLKIT_FORCE_BROWSER_CHECKS=1 is set; run the gate once on a native amd64 machine. Then choose Custom (push to ECR) in Studio and pick the image.