Skip to main content
Use fresh AWS accounts

Install AlphaAgent Organisations into a new, empty AWS account, and give every AlphaAgent Studio deployment a new, empty AWS account of its own. Three reasons:

  1. It is AWS best practice to isolate each workload in its own account, with its own limits, billing view and permissions.
  2. Amazon Bedrock quotas are set per AWS account. Many users in one deployment share one quota and exhaust it; one account per deployment keeps each team's capacity its own.
  3. The blast radius is smaller. A problem in one deployment, or one team's mistake, cannot reach the others.

Troubleshooting the install

For anyone whose install stopped: find the surface, then the message. For the installer the rule is always the same: fix what the line names and re-run the same command; completed steps are kept.

The installer (orgctl.py)​

Every fatal error prints one FATAL line on standard error and exits: 1 fatal or failed preflight, 2 doctor found a failing check, 130 Ctrl-C.

The line beginsDo this
"no AWS credentials were found.", "AWS credentials were found but could not be used", "AWS credentials are no longer valid"Set AWS_PROFILE (or export AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, AWS_SESSION_TOKEN) or refresh them (aws sso login); re-run. A session that expired mid-run keeps every completed step.
"AWS denied an operation (<code>) during <phase>: your credentials are missing the IAM action <service:Action>"Grant that action (Before you start, SCPs and guardrails); re-run.
"AWS error [<code>] during <phase>"Read the code, run python3 orgctl.py status, re-run.
"transient error during <what> (attempt i/8); retrying in Ns"Not fatal: up to eight retries, at most 45 seconds apart, ten minutes in total.
A stack ends CREATE_FAILED, ROLLBACK_COMPLETE, ROLLBACK_FAILED, DELETE_FAILED, UPDATE_ROLLBACK_FAILED or UPDATE_FAILEDThe streamed events name the resource. status labels the stack "stuck": delete it, then python3 orgctl.py install --restart-from <step>. Deleting alphaagent-org-auth also means redoing the SAML application and the saml step.
"this directory's state file was confirmed against account <A> / region <R>, but the live AWS credentials now resolve to <B> / <S>."python3 orgctl.py reset, then install with the right --region and credentials.
"another orgctl process (pid N) is already running against this state file"Wait for it; if it is dead, remove .alphaagent-org-install.lock and re-run.
"stdin closed while waiting for input"Run with --config and --yes.
"this release is missing the <step> step" (exit 3), "the release bundle is incomplete", "could not resolve a template for stack"Re-download the bundle and unzip it into a fresh, empty directory.
"N preflight check(s) failed"Nothing was created; each row's -> line says what to fix. caller-identity "credentials are for X, not the expected Y": switch credentials, or reset for another account. acm-certificate "no ISSUED certificate in <region> covers <host>": request or import one in the install region (a us-east-1 certificate will not do); the line prints the aws acm request-certificate command.
"This run did not update the running install (still vX). Run python3 orgctl.py install --redo or report this."An upgrade changed nothing; run install --redo (Upgrading Organisations).

Resume menu​

When the account already holds the six stacks or the directory holds checkpoints, install prints Existing Organisation detected with each stack's status and asks How do you want to proceed?: Resume ("continue from the first incomplete step (recommended)"), Redo ("re-run every step (idempotent, just slower)") or Delete and reinstall ("DESTROY everything, then start fresh (irreversible)"). "one or more stacks are in a FAILED/ROLLBACK state" points at the last. A newer bundle adds "this bundle is <new>; the running install is <old>." and Resume upgrades: Upgrading Organisations.

State file and lock​

Progress lives in .alphaagent-org-install.json beside the bundle, with an append-only .alphaagent-org-install-runlog.jsonl; the credential is never written to either. The file records the account and region it was confirmed against, so the same directory against another account or region is fatal until you reset; a file from a newer installer is fatal; an unreadable one is a fresh start. A second install, reset or teardown in the same directory is refused while one runs; a lock left by a dead process is cleared for you, and status never takes the lock. settle waits up to 900 seconds and each stack operation up to 2400 (the ORGCTL_SETTLE_TIMEOUT and ORGCTL_STACK_TIMEOUT settings on Install Organisations).

settle​

"could not reach https://<hostname>/ ... that is expected." is a warning: create the CNAME or wait, then python3 orgctl.py settle. "/ -> 302, but to '...' rather than Cognito": the record points elsewhere. "returned <status>, not a redirect to sign-in": check DNS, then anything in front of the load balancer, then re-run python3 orgctl.py app. "ECS reports a FAILED rollout" or "services did not reach steady state within 900s": the events above say why. "<target group> is not healthy": the container fails its own /health probe; read its CloudWatch logs.

saml​

"did not yield a user pool and client id": run org-auth, then saml. "no file found at <path>": paste the full path. "configure-saml.py failed (exit N)": nothing changed; fix the cause its output names, re-run saml. Entra complains about the reply URL or identifier: compare the three values character for character. The browser shows the user pool's own sign-in page: saml has not completed. Sign out ends on "Required parameters missing": run org-auth.

doctor​

Step 12/12, or python3 orgctl.py doctor at any time. Read-only: STATUS and CHECK rows, a -> remediation under each non-passing row, then "N passed, N warning(s), N failed, N skipped". Any FAIL exits 2; --json prints {"ok": true|false, "checks": [...]}. Since Organisations 1.0.8 the services group waits up to 15 minutes for a rollout the run itself started; --no-wait judges the services as they are. Row counts vary (one row per secret, table, repository and service); compare group names.

GroupPASS meansFAIL: re-run
secretsEach of the six configuration secrets holds a populated JSON objectseed-config
config-fail-closedThe configuration is "armed": a service that cannot read a setting fails loudly rather than using a code defaultseed-config
tablesEvery DynamoDB table exists with the expected key schemafoundation
admin-rosterAt least one owner: "N owner(s): <emails>"bootstrap-admin
repositoriesThis release's image tag, with the digest recorded at install (repositories are immutable)push-image; a digest mismatch means the image was replaced outside the installer
servicesorg-api and org-provisioner stable, rollout COMPLETEDapp; "rolloutState IN_PROGRESS" clears on its own
target-account-role-templateorg-target-account-role.yaml staged for the Accounts screenapp; until then the Launch the stack button 404s
spa-corsThe console bucket allows https://<hostname>foundation
saml"provider present, client flipped, listener parameter set"saml
sign-outThe signed-out page is registered on the user pool clientorg-auth
console-credentialAlphaAgent Console still accepts the stored credential (WARN when unreachable)console-credential ("Console no longer accepts this credential"; bundle downloads fail until then)

Each -> line names the step to re-run. "check crashed": send doctor --json output to Support.

The AlphaAgent Console​

The message readsDo this
The offer is not on the listingYou are signed in to a different account from the id you sent. Wrong id sent: reply before accepting; a new offer is made.
"We couldn't complete the link from AWS Marketplace." (missing_token, expired, failed), "This AWS Marketplace registration link has expired", "No pending AWS Marketplace registration was found for this browser session."From the buyer account, in the same browser, click Set up your account on the listing again.
"already linked to another Console account", "already linked to a different AWS Marketplace buyer account"The pairing is one-to-one; write to Support with both ids.
No active agreement after a few minutes, "Couldn't load subscription"Reload, then Support with the buyer account id.
Create license key fails with Open Billing → Subscription; the list reads "Link your AWS Marketplace subscription, then create your first license key"No active agreement is linked yet (Console account, offer and licence).
"No releases are available yet."Support.
"Download failed"Download bundle again.
Credential created closed without copyingRotate and copy the new value.

The Organisations console​

The message readsDo this
Consent returns to Identity Providers with "Admin consent was declined", "did not return everything this connection needs", "This connection attempt expired or was tampered with" (15 minutes), "could not be reached", "Connected, but this tenant's details could not be read" or "Something went wrong"Try again as a Global Administrator; for the tenant-details message check the five permissions (Connect your identity provider).
"This platform's Entra credentials are not configured"An owner enters the Client ID and secret first.
Verify: Assume the target role failsThe stack is missing, ExternalId differs (compare fingerprints) or RoleNameSuffix is not prod.
Verify: Retry CloudFront edge stacks failsUpdate the stack from the current template with us-east-1 in AllowedRegions.
Launch the stack opens a 404Run python3 orgctl.py app, then doctor.
The role "already exists"Delete the old alphaagent-org-target-<account id> stack or use another suffix.
Clear account "Could not check"Re-run the template from the current version.
"That's the License ID, not the License Token."Copy the whole token including the dot.
"That doesn't look like a License Token."Copy from the License Token field.
Run preflight disabledIts hover text names the incomplete section.
"Cannot launch: <checks>."Fix them, Run preflight again.
A parked step (Waiting for you)The card Proceed despite failed readiness checks lists the checks; tick and Approve and resume, or Do not override. Link to this step shares the URL; Abandon this job cancels without touching AWS. An expired park asks for a new job, which parks at the same point.
A failed step (the card <step> failed)It says who acts (Customer account, AlphaAgent or AWS), gives What to check, a Diagnostic ID (Copy) and Technical details, and offers Resume (from the failed step, after Confirm retry), Redo (every step again, safe), Abandon or Wipe and restart (type the deployment name). A policy denial names the action: SCPs and guardrails.
The Edge stack failsus-east-1 missing from the role's regions, or a stale template.
edge_certificate_pending_validation: "Add this DNS record at your DNS provider, then resume the job: <name> CNAME <value>"Create the printed CNAME at your DNS provider, then Resume. An imported or email-validated Studio certificate always stops here; a DNS-validated one validates on the same record with no action.
Studio not reachable after the CNAMEThe record must point at DNS target (CNAME), not the load balancer, and the domain must match the certificate.
Re-installing under a domain used beforeChange the old record first; CloudFront refuses a new install while it points at the old distribution.
Another job holds the deploymentLast job in the Fleet row links to it; Abandon a park nobody will answer.
The owner sees "You have access to nothing yet"The Entra email differs from the break-glass email in doctor's admin-roster.