Infrastructure sizing and costs
For the administrator choosing the Fargate sizing profile and ElastiCache node type in the deployment wizard, and the engineer or finance owner budgeting the AWS account: what Studio provisions, how it is sized, how long a deployment and a graph build take, a rough monthly estimate, and how to change the tier or node type after install. Studio runs entirely in your AWS account, so you pay AWS directly for what it provisions.
You pay for two independent things:
- AWS infrastructure: the always-on compute, networking, storage and data stores Studio runs in your account. This page estimates that layer.
- Usage-based costs: Amazon Bedrock model inference (billed by AWS per token, in proportion to how much your team uses Studio) and your AlphaAgent licence. These are not included in the figures below, because they depend entirely on your workload.
The numbers here are the fixed AWS baseline: what the stack costs to keep running before anyone sends a prompt.
The figures use public AWS list prices for us-east-1 as a sample region and a 730-hour month, and exclude savings plans, credits and committed-use discounts. Your bill depends on your region, actual usage (DynamoDB, S3, data transfer, EFS, Bedrock) and any AWS discounts. Treat them as a planning figure, not a quote. The vCPU and memory figures are read from the deployment templates; the prices and monthly totals are carried from the previous edition of this page and are re-checked with each release.
Before you start: know roughly how many people will use the deployment at once, whether they run long analyses or large knowledge-graph builds, and the AWS region you registered the account for (Register an AWS account); prices differ by region.
How Studio is sized
Studio is deliberately over-provisioned for reliability, not cost-minimised. A single deployment at the default tier is intended for a team of roughly 8 to 10 users running several concurrent workflows and long, multi-step chat analyses at once, so the services are sized generously to stay stable under load rather than tuned to the cheapest configuration that works.
The backend is a set of containerised services on AWS Fargate (serverless containers, no EC2 instances to manage). Each service runs as a single task: capacity lives in one larger, vertically sized task rather than several copies. Several services hold live session state in memory, and Studio's internal traffic is balanced without session affinity, so a second copy of a service would receive roughly half of a session's follow-up calls without holding that session. The graph database, Neo4j, likewise runs as a single task, with its data on an Amazon Elastic File System (EFS) volume.
Studio offers five sizing tiers: xs, small, standard (the default), large and xlarge. You choose one in the wizard's Fargate sizing profile field (blank means standard) and can change it later (see Changing the tier or Redis node type after install). The table below is the default standard tier:
| Service | vCPU | Memory | Tasks |
|---|---|---|---|
| agent-runtime | 16 | 64 GB | 1 |
| knowledge-base | 16 | 64 GB | 1 |
| knowledge-base-mcp | 8 | 32 GB | 1 |
| neo4j | 8 | 32 GB | 1 |
| agent-management | 4 | 16 GB | 1 |
| chat | 4 | 16 GB | 1 |
| data-connector | 4 | 16 GB | 1 |
| workflow | 4 | 16 GB | 1 |
| code-interpreter | 4 | 16 GB | 1 |
| web-browser-search | 4 | 16 GB | 1 |
| notification-service | 4 | 16 GB | 1 |
Total running fleet at standard: about 76 vCPU and 304 GB of memory across 11 tasks. These are the service names you see under Health and Analytics → Services on the deployment.
Neo4j mode. A deployment installed from the Organisations console hosts its own graph database: the neo4j task above runs in your account and an EFS file system holds its data. The deployment's Overview → Configuration shows this as Neo4j mode self-hosted. The wizard has no field for the setting, and every total on this page includes the neo4j task and the EFS line.
A new AWS account's default Fargate concurrent vCPU quota is 140, above Studio's 76 vCPU fleet. Updates stop old tasks before new ones start, behind an update screen, so the fleet is never doubled and a default account does not need a quota increase. See Before you start.
Sizing tiers
The tiers form a ladder. The services are grouped into three classes (heavy: agent-runtime and knowledge-base; mid: knowledge-base-mcp and neo4j; light: the rest), and each tier sets the CPU and memory of every class. The totals below include the neo4j task (a mid-class task: 2 vCPU and 8 GB at xs, 4 and 16 GB at small, 8 and 32 GB at standard, 8 and 48 GB at large, 8 and 60 GB at xlarge).
| Tier | Total vCPU | Total memory | Intended for |
|---|---|---|---|
xs | 19 vCPU | 76 GB | Evaluation, 1 to 2 users |
small | 38 vCPU | 152 GB | Small teams, about 2 to 5 users |
standard (default) | 76 vCPU | 304 GB | About 8 to 10 concurrent power users |
large | 76 vCPU | 456 GB | Heavier knowledge graphs and long runs, about 10 to 15 mixed users |
xlarge | 76 vCPU | 570 GB | Maximum single-task headroom |
large and xlarge add memory, not CPUAt standard the two heavy services are already at 16 vCPU per task, the Fargate per-task maximum. large and xlarge therefore add memory only: the heavy tasks grow to 96 GB and 120 GB, the mid tasks to 48 GB and 60 GB, and the light tasks to 24 GB and 30 GB. Going beyond xlarge would need horizontal scaling, which Studio does not support (each service runs as a single task).
Below standard, small allocates half the vCPU and memory of standard (the graph database's memory settings scale down to match), and xs is a minimal evaluation footprint. Smaller tiers cut the Fargate line, the largest fixed cost, but trade headroom for it: heavy, long-running analyses or knowledge-graph builds have less memory to work with, so very large jobs may run slower or, at the extreme, hit memory limits.
ElastiCache (Redis) node type
The Redis cache runs as a two-node, Multi-AZ ElastiCache replication group. The wizard's ElastiCache node type field offers exactly the node types the console serves: cache.t4g.medium (the default when blank; burstable, about 3.1 GB), cache.m7g.large, cache.r7g.large and cache.r7g.xlarge. Choose a larger type for larger working sets or heavier concurrent load. The type can be changed after install; the change is applied online but briefly fails the cache over, so schedule it inside a maintenance window (see below).
How long things take
- A Studio deployment takes 45 to 50 minutes from launch to Healthy on the Fleet. A second deployment launched into the same Organisation waits for the first deployment's release image load to finish before its own image steps start.
- A knowledge-graph build from a 26-page document takes hours at
xs. The build runs in the knowledge-base task, one of the two heavy services, so a larger tier gives it more memory (and, up tostandard, more CPU); atstandardthat task has 16 vCPU and 64 GB.
What you pay for
| Layer | Component | Billing model |
|---|---|---|
| Compute | Fargate (the service fleet above) | Per vCPU-hour plus per GB-hour, always on |
| Cache | ElastiCache for Redis | Per node-hour, always on |
| Networking | Network address translation (NAT) gateway, Application Load Balancer (ALB), VPC interface endpoints, CloudFront | Per hour plus per GB or capacity unit processed |
| Storage | EFS for Neo4j, S3 buckets | Per GB stored (usage-based) |
| Data stores | DynamoDB (on-demand) | Per request plus per GB stored (usage-based) |
| Secrets and logs | Secrets Manager, CloudWatch Logs | Per secret, per GB ingested |
| AI (separate) | Amazon Bedrock model inference | Per token, not included here |
Rough monthly estimate (us-east-1)
Using us-east-1 Fargate Linux/ARM64 list prices of $0.03238 per vCPU-hour and $0.00356 per GB-hour over a 730-hour month, at the standard tier:
| Item | Basis | Estimated monthly |
|---|---|---|
| Fargate, vCPU | 76 vCPU × $0.03238 × 730 h | about $1,796 |
| Fargate, memory | 304 GB × $0.00356 × 730 h | about $790 |
| ElastiCache Redis | 2 × cache.t4g.medium (default node type) | about $100 |
| NAT gateway | 1 gateway (hourly) | about $33 plus data |
| Application Load Balancer | 1 ALB (hourly) | about $16 plus capacity units |
| VPC interface endpoints | about 9 endpoints × Availability Zones | about $100 to $130 |
| Secrets Manager | about 16 secrets | about $6 |
| EFS, DynamoDB, S3, data transfer | usage-based | varies |
| Fixed AWS baseline (excluding Bedrock) | about $2,850 to $3,000 per month |
Low, medium and high
To bracket the usage-based and networking items:
| Scenario | What it assumes | Estimated fixed AWS baseline |
|---|---|---|
| Low | Default setup, light data and networking | about $2,500 to $2,700 per month |
| Medium | Default setup, moderate usage | about $2,850 to $3,000 per month |
| High | Heavy data transfer, large EFS, S3 and DynamoDB footprint | about $3,300 to $3,800 per month |
All scenarios exclude Amazon Bedrock token costs and your AlphaAgent licence, which are usage-driven and tracked separately.
Ways to reduce cost
- Use a smaller sizing tier.
smallis half the vCPU and memory ofstandard;xsis smaller still. Choose it in the wizard, or move an existing deployment with a Fleet Configuration Template and an Update Manager rule as described below. - AWS Savings Plans. Compute Savings Plans apply to Fargate and can cut the compute line materially for a one- or three-year commitment.
- Choose a cheaper region if your data-residency requirements allow;
us-east-1is used here only as a sample. Studio deploys only into the supported US and EU regions listed in Before you start.
Changing the tier or Redis node type after install
The deployment's Overview changes only run data retention. Sizing and the node type are changed centrally: you describe the target in a Fleet Configuration Template, then an Update Manager rule applies it to the deployments you choose, on a schedule and inside a maintenance window if you set one.
Steps
- Open Fleet Configurations and click New Fleet Configuration Template (or Edit an existing one).
- Give it a Name (up to 100 characters) and, optionally, a Description (up to 500).
- Under Which release, leave the version blank so the rule changes only the sizing, or set it to the version the deployment already runs. A blank field is never applied. If you set a newer version, the same rule also upgrades the deployment.
- Under Sizing, set Sizing profile and/or Redis node type. Leave every field you do not want to change blank; the preview at the bottom reads This template will set N field(s) and lists them.
- Click Create (or Save). Saving an existing template increments its version; the modal warns that Any Update Manager rule still pinned to v<N> keeps applying the version it was set up against until an admin explicitly adopts the new one.
- Open Update Manager and click New Update Manager Rule. (The button is disabled until at least one template exists.)
- Give the rule a Name, choose the Fleet Configuration Template you just saved, and tick the Deployments it applies to. The target preview lists each deployment with the change it will receive, for example Sizing: standard → large, and the note: A Redis node type change briefly fails over/reconnects the cache. A sizing change causes a rolling restart of all affected services.
- Set the Schedule (Daily, Weekly or Monthly with a Time (UTC), or an advanced
rate(...)/cron(...)expression). To restrict when it may fire, tick Only fire during a maintenance window and set Start (UTC) and End (UTC). - Tick Enabled (it is off by default) and click Create.
What happens when the rule fires
At the next scheduled time inside the window, Update Manager compares each targeted deployment with the template and starts a job for each one that differs. The job's steps appear on the deployment's Jobs view: Capture a restore point, Show the update screen to Studio users, then the stack update at the new sizing, Restore the normal Studio app and Record the new configuration on the deployment. Every service restarts in one rolling pass; Studio users see the update screen while it runs. Anything mid-flight at that moment, such as a knowledge graph still being built or another long-running background job, is interrupted rather than resumed and must be started again afterwards. A node-type change is applied to the cache online with a brief failover and reconnect.
What you should see
- Fleet Configurations lists the template with its version and, in the Rules column, the rule that uses it.
- Update Manager lists the rule with Next run and, after it fires, Last run. History on the rule shows each firing with its outcome: Fired, Nothing to do, Deferred (outside window) and others, and under Triggered the job it started and that job's status.
- The deployment's Overview → Configuration shows Deployment size and Redis node type with the new values once the job has recorded them. Until a value has been recorded the field's hover reads Not recorded for this deployment. Fields set by the wizard are recorded at launch; configuration fields are recorded when a reconfigure or upgrade applies them.
Notes
- Only the five tiers and the four node types above are accepted; the console refuses anything else.
largeandxlargeraise memory only. There is no horizontal scaling.- Ownership and Run data retention are not template fields: ownership is fixed at creation, and retention is changed from the deployment's Overview.
- Neo4j mode has no field in the wizard or the template editor; a fresh install is
self-hosted, and a change of mode on a live deployment is refused. - A rule skips a deployment that is paused (Skipped (paused)), that has another job in flight (Skipped (job in flight)), or whose region cannot take the template's inference zone (Skipped (zone mismatch)).
- The prices on this page are list prices for one sample region and one point in time.
- History reads "Template edited, not adopted". You saved a new version of the template after the rule was created. On the rule's row click Adopt v<N> and confirm; the confirmation repeats the warning about in-flight work being interrupted.
- History reads "Drift found, not applied". Every deployment that differed was skipped for one of the reasons above. Open the deployment: resume it if it is paused, or wait for the job in flight.
- History reads "Template missing". The template was deleted. Point the rule at another template, or delete the rule.
- The job failed. Open it on the deployment's Jobs view; the failure card says who acts and what to check, and offers Resume, Redo or Abandon. A failed update rolls the fleet back to its restore point. See Jobs, progress and failures and Failure codes.
- A large job hit memory limits at a small tier. Move up the ladder with the same template and rule mechanism.
Related
- Deploy your first Studio: where the tier and node type are first chosen.
- Before you start: quotas, regions and account readiness.
- Fleet configurations and Update Manager: every template and rule field.
- What runs in your account: Studio: the component inventory.
- Everything AlphaAgent deploys: every stack and resource type.
- Deployment detail: CPU and memory charts per deployment.
- Supported models: the other deployment-wide setting changed the same way.