Skip to main content

Infrastructure sizing and costs

For the administrator choosing the Fargate sizing profile and ElastiCache node type in the deployment wizard, and the engineer or finance owner budgeting the AWS account: what Studio provisions, how it is sized, how long a deployment and a graph build take, a rough monthly estimate, and how to change the tier or node type after install. Studio runs entirely in your AWS account, so you pay AWS directly for what it provisions.

Two separate cost layers

You pay for two independent things:

  1. AWS infrastructure: the always-on compute, networking, storage and data stores Studio runs in your account. This page estimates that layer.
  2. Usage-based costs: Amazon Bedrock model inference (billed by AWS per token, in proportion to how much your team uses Studio) and your AlphaAgent licence. These are not included in the figures below, because they depend entirely on your workload.

The numbers here are the fixed AWS baseline: what the stack costs to keep running before anyone sends a prompt.

These are estimates

The figures use public AWS list prices for us-east-1 as a sample region and a 730-hour month, and exclude savings plans, credits and committed-use discounts. Your bill depends on your region, actual usage (DynamoDB, S3, data transfer, EFS, Bedrock) and any AWS discounts. Treat them as a planning figure, not a quote. The vCPU and memory figures are read from the deployment templates; the prices and monthly totals are carried from the previous edition of this page and are re-checked with each release.

Before you start: know roughly how many people will use the deployment at once, whether they run long analyses or large knowledge-graph builds, and the AWS region you registered the account for (Register an AWS account); prices differ by region.

How Studio is sized​

Studio is deliberately over-provisioned for reliability, not cost-minimised. A single deployment at the default tier is intended for a team of roughly 8 to 10 users running several concurrent workflows and long, multi-step chat analyses at once, so the services are sized generously to stay stable under load rather than tuned to the cheapest configuration that works.

The backend is a set of containerised services on AWS Fargate (serverless containers, no EC2 instances to manage). Each service runs as a single task: capacity lives in one larger, vertically sized task rather than several copies. Several services hold live session state in memory, and Studio's internal traffic is balanced without session affinity, so a second copy of a service would receive roughly half of a session's follow-up calls without holding that session. The graph database, Neo4j, likewise runs as a single task, with its data on an Amazon Elastic File System (EFS) volume.

Studio offers five sizing tiers: xs, small, standard (the default), large and xlarge. You choose one in the wizard's Fargate sizing profile field (blank means standard) and can change it later (see Changing the tier or Redis node type after install). The table below is the default standard tier:

ServicevCPUMemoryTasks
agent-runtime1664 GB1
knowledge-base1664 GB1
knowledge-base-mcp832 GB1
neo4j832 GB1
agent-management416 GB1
chat416 GB1
data-connector416 GB1
workflow416 GB1
code-interpreter416 GB1
web-browser-search416 GB1
notification-service416 GB1

Total running fleet at standard: about 76 vCPU and 304 GB of memory across 11 tasks. These are the service names you see under Health and Analytics → Services on the deployment.

Neo4j mode. A deployment installed from the Organisations console hosts its own graph database: the neo4j task above runs in your account and an EFS file system holds its data. The deployment's Overview → Configuration shows this as Neo4j mode self-hosted. The wizard has no field for the setting, and every total on this page includes the neo4j task and the EFS line.

No Fargate quota increase needed

A new AWS account's default Fargate concurrent vCPU quota is 140, above Studio's 76 vCPU fleet. Updates stop old tasks before new ones start, behind an update screen, so the fleet is never doubled and a default account does not need a quota increase. See Before you start.

Sizing tiers​

The tiers form a ladder. The services are grouped into three classes (heavy: agent-runtime and knowledge-base; mid: knowledge-base-mcp and neo4j; light: the rest), and each tier sets the CPU and memory of every class. The totals below include the neo4j task (a mid-class task: 2 vCPU and 8 GB at xs, 4 and 16 GB at small, 8 and 32 GB at standard, 8 and 48 GB at large, 8 and 60 GB at xlarge).

TierTotal vCPUTotal memoryIntended for
xs19 vCPU76 GBEvaluation, 1 to 2 users
small38 vCPU152 GBSmall teams, about 2 to 5 users
standard (default)76 vCPU304 GBAbout 8 to 10 concurrent power users
large76 vCPU456 GBHeavier knowledge graphs and long runs, about 10 to 15 mixed users
xlarge76 vCPU570 GBMaximum single-task headroom
large and xlarge add memory, not CPU

At standard the two heavy services are already at 16 vCPU per task, the Fargate per-task maximum. large and xlarge therefore add memory only: the heavy tasks grow to 96 GB and 120 GB, the mid tasks to 48 GB and 60 GB, and the light tasks to 24 GB and 30 GB. Going beyond xlarge would need horizontal scaling, which Studio does not support (each service runs as a single task).

Below standard, small allocates half the vCPU and memory of standard (the graph database's memory settings scale down to match), and xs is a minimal evaluation footprint. Smaller tiers cut the Fargate line, the largest fixed cost, but trade headroom for it: heavy, long-running analyses or knowledge-graph builds have less memory to work with, so very large jobs may run slower or, at the extreme, hit memory limits.

ElastiCache (Redis) node type​

The Redis cache runs as a two-node, Multi-AZ ElastiCache replication group. The wizard's ElastiCache node type field offers exactly the node types the console serves: cache.t4g.medium (the default when blank; burstable, about 3.1 GB), cache.m7g.large, cache.r7g.large and cache.r7g.xlarge. Choose a larger type for larger working sets or heavier concurrent load. The type can be changed after install; the change is applied online but briefly fails the cache over, so schedule it inside a maintenance window (see below).

How long things take​

  • A Studio deployment takes 45 to 50 minutes from launch to Healthy on the Fleet. A second deployment launched into the same Organisation waits for the first deployment's release image load to finish before its own image steps start.
  • A knowledge-graph build from a 26-page document takes hours at xs. The build runs in the knowledge-base task, one of the two heavy services, so a larger tier gives it more memory (and, up to standard, more CPU); at standard that task has 16 vCPU and 64 GB.

What you pay for​

LayerComponentBilling model
ComputeFargate (the service fleet above)Per vCPU-hour plus per GB-hour, always on
CacheElastiCache for RedisPer node-hour, always on
NetworkingNetwork address translation (NAT) gateway, Application Load Balancer (ALB), VPC interface endpoints, CloudFrontPer hour plus per GB or capacity unit processed
StorageEFS for Neo4j, S3 bucketsPer GB stored (usage-based)
Data storesDynamoDB (on-demand)Per request plus per GB stored (usage-based)
Secrets and logsSecrets Manager, CloudWatch LogsPer secret, per GB ingested
AI (separate)Amazon Bedrock model inferencePer token, not included here

Rough monthly estimate (us-east-1)​

Using us-east-1 Fargate Linux/ARM64 list prices of $0.03238 per vCPU-hour and $0.00356 per GB-hour over a 730-hour month, at the standard tier:

ItemBasisEstimated monthly
Fargate, vCPU76 vCPU × $0.03238 × 730 habout $1,796
Fargate, memory304 GB × $0.00356 × 730 habout $790
ElastiCache Redis2 × cache.t4g.medium (default node type)about $100
NAT gateway1 gateway (hourly)about $33 plus data
Application Load Balancer1 ALB (hourly)about $16 plus capacity units
VPC interface endpointsabout 9 endpoints × Availability Zonesabout $100 to $130
Secrets Managerabout 16 secretsabout $6
EFS, DynamoDB, S3, data transferusage-basedvaries
Fixed AWS baseline (excluding Bedrock)about $2,850 to $3,000 per month

Low, medium and high​

To bracket the usage-based and networking items:

ScenarioWhat it assumesEstimated fixed AWS baseline
LowDefault setup, light data and networkingabout $2,500 to $2,700 per month
MediumDefault setup, moderate usageabout $2,850 to $3,000 per month
HighHeavy data transfer, large EFS, S3 and DynamoDB footprintabout $3,300 to $3,800 per month

All scenarios exclude Amazon Bedrock token costs and your AlphaAgent licence, which are usage-driven and tracked separately.

Ways to reduce cost​

  • Use a smaller sizing tier. small is half the vCPU and memory of standard; xs is smaller still. Choose it in the wizard, or move an existing deployment with a Fleet Configuration Template and an Update Manager rule as described below.
  • AWS Savings Plans. Compute Savings Plans apply to Fargate and can cut the compute line materially for a one- or three-year commitment.
  • Choose a cheaper region if your data-residency requirements allow; us-east-1 is used here only as a sample. Studio deploys only into the supported US and EU regions listed in Before you start.

Changing the tier or Redis node type after install​

The deployment's Overview changes only run data retention. Sizing and the node type are changed centrally: you describe the target in a Fleet Configuration Template, then an Update Manager rule applies it to the deployments you choose, on a schedule and inside a maintenance window if you set one.

Steps​

  1. Open Fleet Configurations and click New Fleet Configuration Template (or Edit an existing one).
  2. Give it a Name (up to 100 characters) and, optionally, a Description (up to 500).
  3. Under Which release, leave the version blank so the rule changes only the sizing, or set it to the version the deployment already runs. A blank field is never applied. If you set a newer version, the same rule also upgrades the deployment.
  4. Under Sizing, set Sizing profile and/or Redis node type. Leave every field you do not want to change blank; the preview at the bottom reads This template will set N field(s) and lists them.
  5. Click Create (or Save). Saving an existing template increments its version; the modal warns that Any Update Manager rule still pinned to v<N> keeps applying the version it was set up against until an admin explicitly adopts the new one.
  1. Open Update Manager and click New Update Manager Rule. (The button is disabled until at least one template exists.)
  2. Give the rule a Name, choose the Fleet Configuration Template you just saved, and tick the Deployments it applies to. The target preview lists each deployment with the change it will receive, for example Sizing: standard → large, and the note: A Redis node type change briefly fails over/reconnects the cache. A sizing change causes a rolling restart of all affected services.
  3. Set the Schedule (Daily, Weekly or Monthly with a Time (UTC), or an advanced rate(...)/cron(...) expression). To restrict when it may fire, tick Only fire during a maintenance window and set Start (UTC) and End (UTC).
  4. Tick Enabled (it is off by default) and click Create.

What happens when the rule fires​

At the next scheduled time inside the window, Update Manager compares each targeted deployment with the template and starts a job for each one that differs. The job's steps appear on the deployment's Jobs view: Capture a restore point, Show the update screen to Studio users, then the stack update at the new sizing, Restore the normal Studio app and Record the new configuration on the deployment. Every service restarts in one rolling pass; Studio users see the update screen while it runs. Anything mid-flight at that moment, such as a knowledge graph still being built or another long-running background job, is interrupted rather than resumed and must be started again afterwards. A node-type change is applied to the cache online with a brief failover and reconnect.

What you should see​

  • Fleet Configurations lists the template with its version and, in the Rules column, the rule that uses it.
  • Update Manager lists the rule with Next run and, after it fires, Last run. History on the rule shows each firing with its outcome: Fired, Nothing to do, Deferred (outside window) and others, and under Triggered the job it started and that job's status.
  • The deployment's Overview → Configuration shows Deployment size and Redis node type with the new values once the job has recorded them. Until a value has been recorded the field's hover reads Not recorded for this deployment. Fields set by the wizard are recorded at launch; configuration fields are recorded when a reconfigure or upgrade applies them.

Notes​

  • Only the five tiers and the four node types above are accepted; the console refuses anything else.
  • large and xlarge raise memory only. There is no horizontal scaling.
  • Ownership and Run data retention are not template fields: ownership is fixed at creation, and retention is changed from the deployment's Overview.
  • Neo4j mode has no field in the wizard or the template editor; a fresh install is self-hosted, and a change of mode on a live deployment is refused.
  • A rule skips a deployment that is paused (Skipped (paused)), that has another job in flight (Skipped (job in flight)), or whose region cannot take the template's inference zone (Skipped (zone mismatch)).
  • The prices on this page are list prices for one sample region and one point in time.
  • History reads "Template edited, not adopted". You saved a new version of the template after the rule was created. On the rule's row click Adopt v<N> and confirm; the confirmation repeats the warning about in-flight work being interrupted.
  • History reads "Drift found, not applied". Every deployment that differed was skipped for one of the reasons above. Open the deployment: resume it if it is paused, or wait for the job in flight.
  • History reads "Template missing". The template was deleted. Point the rule at another template, or delete the rule.
  • The job failed. Open it on the deployment's Jobs view; the failure card says who acts and what to check, and offers Resume, Redo or Abandon. A failed update rolls the fleet back to its restore point. See Jobs, progress and failures and Failure codes.
  • A large job hit memory limits at a small tier. Move up the ladder with the same template and rule mechanism.