Skip to main content

Troubleshooting

When AlphaAgent refuses to do something or stops partway, it names the reason on screen; find that text below, in the order you meet it, from installing Organisations to working in Studio. Copy the message exactly as printed (the message, or its first sentence, is what the rows quote); for a deployment job, open the deployment's Jobs view and read the failure card: its heading names the step, Who acts names who has to move, and Technical details shows the failure code and the Diagnostic ID.

Installing Organisations: a preflight check refused​

python3 orgctl.py install (and python3 orgctl.py preflight) runs fifteen checks before creating anything. Each ends PASS, WARN, FAIL or SKIP, and every non-PASS line carries its own remedy. Any FAIL stops the install with "N preflight check(s) failed: (ids). Fix the items above and re-run. Nothing has been created, so this is free to repeat." and exit status 1. A check you have judged safe to ignore can be skipped with --skip-check <id>; skipped checks are listed in the report.

Check idIt verifiesIf it does not pass
python-versionPython 3.11 or newerFAIL "orgctl needs python >= 3.11. Install it and re-run."
boto3the boto3 library importsFAIL "Install it with: python3 -m pip install boto3"
docker-daemonthe docker CLI is on PATH and its daemon answersWARN if crane or regctl is present ("docker is not available; the image will be pushed with crane"); otherwise FAIL "the docker CLI is not on PATH" or "installed but the daemon is not answering". Start Docker, or install crane or regctl
image-copiercrane or regctl is on PATHWARN only; docker is used instead
release-manifestthe bundle's RELEASE.txt parsesSKIP when the file is absent; re-download the bundle if it is present but unreadable
bundle-completeevery mandatory bundle member is present and non-emptyFAIL lists the missing members; re-download the bundle from the Console
caller-identityyour AWS credentials resolve to the account this install belongs toFAIL "credentials are for X, not the expected Y"; switch profile or credentials
region-allowedthe region is one AlphaAgent supportsFAIL; pass one of the eleven supported regions with --region
availability-zonesat least two usable availability zonesFAIL with the count; choose another region
quota-headroomElastic IP, VPC and NAT-gateway headroom in the regionWARN, never FAIL; SKIP "service-quotas unreadable" when your identity lacks servicequotas:GetServiceQuota. Advisory: raise the quota if it is tight
vpc-cidr-overlapthe Organisation's /16 does not overlap an existing VPCFAIL names the clashing VPCs; answer the "Private network range (/16)" prompt with another range (python3 orgctl.py reset clears the cached answer)
cognito-prefixthe sign-in domain prefix is freeWARN if it belongs to this install's own pool (a resume) or cannot be read ("Needs cognito-idp:DescribeUserPoolDomain"); FAIL "(prefix) is taken by pool (id)"
acm-certificatean ISSUED certificate in this region covers the console hostnameSKIP before a hostname exists; FAIL "no ISSUED certificate in (region) covers (host)": request or import one in that region; a certificate from another region will not work on the load balancer
dns-resolvesthe console hostname resolvesWARN "(host) does not resolve"; expected on a first install, the CNAME comes after the app step
wedged-stacksno existing Organisation stack is in a state that cannot be updatedFAIL "A stack in one of these states cannot be updated, only deleted."; delete the named stack in CloudFormation, then re-run

No check enumerates your IAM permissions. A missing permission surfaces later as the fatal message in the next section; the requirement printed before the install reads "They need permission to create VPC networking, an Application Load Balancer, ECS, ECR, Lambda, DynamoDB, S3, Secrets Manager, Cognito, IAM roles and CloudWatch log groups."

Installing Organisations: orgctl.py stopped​

Exit statusWhenWhat to do
1any fatal: FATAL (message) on stderrread the message below and re-run the same command; the install resumes from the step it stopped at
2doctor found a failing check ("N check(s) failed: (ids)"); or the command line was wrong (the extra line reads "Run python3 orgctl.py with no arguments for a guided walkthrough, or add --help to any command for its full set of flags.")fix what the doctor row names, or correct the command
3this release is missing a step implementationthe bundle is incomplete or corrupt: download it again and re-run
130you pressed Ctrl-C between stepsre-run the same command to resume

The fatal messages you may see:

Message startsMeaningWhat to do
"no AWS credentials were found."nothing in the environment or profile"Set AWS_PROFILE to a configured profile, or export AWS_ACCESS_KEY_ID and AWS_SECRET_ACCESS_KEY (and AWS_SESSION_TOKEN for temporary credentials), then re-run."
"AWS credentials were found but could not be used"the credentials are present but rejected"They may have expired - refresh them (for example aws sso login) and re-run."
"AWS credentials are no longer valid (code). Your session/token most likely expired during (phase)."your session expired part-wayrefresh your credentials, then re-run the same command; it resumes where it stopped
"AWS denied an operation (code) during (phase): your credentials are missing the IAM action '(service:Action)'."a permission is missinggrant that action to the identity you are running as, then re-run the same command
"AWS error [code] during (phase)"any other AWS error"Completed steps are checkpointed; run python3 orgctl.py status to see what was applied, then re-run the same command to resume from where it stopped."
"transient error during (what) (attempt i/8); retrying in Ns"a throttle or a 5xxnothing: the installer retries up to 8 attempts with backoff capped at 45 seconds each and 600 seconds in total
"(stack) finished in ROLLBACK_COMPLETE." (or another failed stack state)a CloudFormation stack failed; the first failed resource and its reason are printeddelete the stack in CloudFormation, then python3 orgctl.py install --restart-from <step>; python3 orgctl.py status marks such a stack "stuck"
"this release is missing the '(step)' step."the bundle is incomplete or corruptdownload the bundle again and re-run

Registering an AWS account: Verify failed​

Verify this account runs up to three checks and prints the result under Verification result.

CheckIf it fails
assume-role"Confirm the CloudFormation stack from the account template was run in this AWS account, with the External ID copied exactly and RoleNameSuffix left as 'prod' (or matching what you chose)." The account shows Failed with "Verification failed. Open the account to see which check failed and why." A denied action reads "Denied: the role is missing (action)."; without one, "Failed, and AWS did not name a missing action in its response. The full probe detail is in the account event log."
edge-cleanupthe account's registration stack is too old for the CloudFront front door: "Update the account's alphaagent-org-target stack from the current template (TemplateVersion 2.3.0, us-east-1 in AllowedRegions), then Verify again."
account-clearthe last Clear account on this account left resources behind; the row reads "Clear account (job) left N resource(s) in the account" and lists them, and its remedy is the wipe_leftovers text below: Scan for stray resources on the account page and Wipe what it finds

An account that is still Pending reads "Not verified yet" and asks you to run the role template in the account, then verify it. Use Download org-target-account-role.yaml or Launch the stack in the AWS console on the account page. See Accounts.

Deploying or updating Studio: a job failed​

Open the deployment's Jobs view. The failure card's heading is "(step) failed" or "This job failed". Who acts is one of three labels: Customer account ("Something in the target AWS account has to change before this can succeed."), AlphaAgent ("AlphaAgent is investigating this" and no action is needed from you) or AWS ("A transient or capacity condition on the AWS side."). Below it, What to check carries the remedy, a resumable failure says so, and Technical details shows the failure code. Resume continues from the failed step, Redo re-runs every step (each is idempotent), Abandon marks the job cancelled without touching your account, and Wipe tears down and reinstalls. Every card shows a Diagnostic ID with Copy; that id is the job id.

Every code, with the product's message and remediation word for word, is on Failure codes. Grouped by Who acts:

  • Customer account, something in your account, DNS or identity provider has to change: settle_service_never_stabilised, settle_service_rolled_back, access_denied_missing_action, credentials_unusable, preflight_blocked, approval_rejected, park_expired, quota_exceeded, restore_point_unwritable, wipe_leftovers, stack_resource_already_exists, stack_resource_access_denied, environment_in_use, preexisting_named_resources, registration_stack_stale_for_cloudfront, release_lacks_cloudfront_support, cloudfront_alias_conflict, edge_certificate_pending_validation, dns_points_at_alb_with_origin_lock, origin_lock_refused_dns_at_alb, app_domain_points_at_previous_distribution, saml_configuration_failed.
  • AWS, a transient condition: redis_busy. Wait until every member of the replication group named in the details reads "available" in the ElastiCache console, then Resume.
  • AlphaAgent, no change in your account is asked for; Resume where the card says so, otherwise share the Diagnostic ID: stack_create_failed, stack_update_rollback_complete, husk_retained_sweep_incomplete, husk_sweep_refused_stack_was_live, stack_delete_failed, core_stack_blocked_by_lambda_enis, stack_update_rollback_failed_wedged, image_tag_missing, lease_lost_max_reclaims, verify_failed, release_bundle_incomplete, release_bundle_download_failed, release_bundle_upload_interrupted, image_digest_immutable_conflict, image_architecture_mismatch, image_push_failed, image_registry_unreachable, image_config_unreadable, image_digest_mismatch, no_restore_point, rollback_failed, connector_backup_corrupt, wipe_stack_force_delete_failed, internal_error, minimal_stack_not_settled, studio_step_not_implemented, missing_required_studio_field, license_token_unreadable, release_not_deployable, console_credential_invalid, console_unreachable, console_error, console_response_unverified, release_mirror_misconfigured, org_repository_missing, spa_bundle_empty, spa_bundle_unserveable, release_bundle_missing_secret_schema, release_schema_minted_secrets_invalid, release_bundle_missing_template, edge_outputs_missing, front_door_probe_failed, saml_auto_provisioning_failed, entra_credentials_unavailable, doctor_common_config_missing, reseed_retention_drift, reseed_pii_default_drift.

The step names, the parked-step cards and the retry dialog are on Jobs, progress and failures; which guardrail produces which code is on SCPs and guardrails that block installs.

Sign-in problems​

Two separate Entra ID objects are involved, and most sign-in problems come from mixing them up.

  • The Organisations console is signed in through a non-gallery SAML Enterprise Application your identity administrator created during the install, from the Entity ID, Reply URL and Sign-on URL that orgctl.py printed, with the claims email, given_name and family_name. Only users and groups assigned to that application can sign in to the console. After python3 orgctl.py saml, password sign-in is off for everyone.
  • Studio deployments are signed in through Enterprise Applications the Organisation creates for you in your connected identity provider: the Entra tenant you connected under Identity Providers through a multi-tenant app registration with admin consent. A person can sign in to a deployment only when they hold the StudioUser role on it, granted under Identity & Access Management → Users (or through a group), within the deployment's 20 seats.

Organisations console​

What you seeMeaningWhat to do
"Could not sign you in" with "Try again. If it keeps happening, contact your organisation owner."the SAML sign-in did not completeopen Technical details; have your identity administrator confirm you are assigned to the console's Enterprise Application and that it sends the three claims
You are signed in but every screen says you lack permission ("An owner or operator can grant you access in Identity & Access Management.")your account has no admin row or grantan owner grants you a Role from Roles; the break-glass owner given at install can always do this
"Your session has expired"the single sign-on session timed outSign in again; running deployments are unaffected
Sign out ends on a Cognito page reading "Required parameters missing"the sign-out URL is missing from the auth stack (an upgrade from an earlier release may not have applied it)run python3 orgctl.py doctor; its sign-out check names the fix: python3 orgctl.py org-auth
Connecting an identity provider returns "Admin consent was declined in Microsoft Entra, so this identity provider was not connected."consent was refused in Microsoftconnect again and grant admin consent
"This platform's Entra credentials are not configured. Contact your organisation owner."the multi-tenant app registration's client id and secret are not savedan owner enters them on Identity Providers → Connect identity provider
"This connection attempt expired or was tampered with. Try connecting again."the connect link was older than 15 minutes or alteredconnect again
"Microsoft Entra could not be reached to finish this connection. Try again in a moment." or "Connected, but this tenant's details could not be read from Microsoft Graph. Try again in a moment."a transient Microsoft call failedtry again
"Microsoft Entra did not return everything this connection needs. Try connecting again."the callback was incompleteconnect again
"The Entra client secret expires in N day(s)." or "The Entra client secret expired N day(s) ago. Single sign-on provisioning, directory search and assignments fail until it is updated."the shared application secret is near or past expiryUpdate credentials on the identity provider's page; the new secret is verified with Microsoft before anything is replaced

Studio​

What you seeMeaningWhat to do
Microsoft refuses the sign-in, or Studio's "Redirecting to sign-in…" returns you to Microsoftyou are not assigned to this deployment's Enterprise Applicationan administrator assigns you StudioUser on the deployment under Users → Assign a role; the deployment's Identity view lists everyone who can sign in
The administrator's Assign is disabled with "No more users fit in this deployment (20 max)"the deployment's seats are fullremove users who no longer need it, or ask for a higher cap when the Organisation is configured
"Your session has expired", with a body that asks you to sign in again through your organisation's SSO and says your work is safethe SSO session lapsedSign in again; runs in progress continue
The deployment's Identity view reads "Who can sign in appears here once provisioning finishes and the Enterprise Application exists."SSO provisioning is still runningwait; the view reloads every 10 seconds
"The last provisioning attempt failed" on the Identity viewthe SSO step of the job failedopen View job history and read the card; saml_auto_provisioning_failed usually clears on Resume
The deployment is not linked to an identity providermanual SAMLfollow the seven-step runbook on the Identity view; the job is parked waiting for your metadata, not failed, and nothing expires for 14 days

My run is throttled​

A strip across the top of Studio (and inside the chat) means Amazon Bedrock in your deployment's AWS account is throttling. It has three states:

StateWhat it saysWhat happens to the run
Soft"Amazon Bedrock is rate-limiting" the named run, then "AlphaAgent is now adjusting its output speed and will re-evaluate in N seconds before attempting to resume full speed. In the meantime, please contact AWS to raise your Bedrock quotas."it continues more slowly and re-evaluates after the stated seconds; the strip clears on its own
Hard"Amazon Bedrock returned a ThrottlingException for" the named run, with Bedrock's detailthe strip clears after the stated time
Daily quota"Service halted" followed by "Amazon Bedrock has exhausted today's per-account token quota for model (id). All AlphaAgent operations using this model are halted until Bedrock's daily window resets. Please contact AWS to raise your Bedrock quotas."everything using that model stops until Bedrock's daily window resets; the strip stays

Click the strip to open the throttled workflow run. The fix is on the AWS side: raise the Bedrock quota for the model, account and region named, using Quota docs on the strip for AWS's instructions. Bedrock quotas are per AWS account and per region, so a busy deployment sharing an account with anything else exhausts them sooner; see Limits and quotas.

My pull was rejected​

Libraries → My Pulls lists every pull you requested. A Rejected row shows the reason word for word; the same reason appears in the Organisations console under the item's Activity.

ReasonMeaningWhat to do
"item no longer exists"the item was deleted from the Librarynothing to pull
"No longer available", "deleted in (deployment) on (date)"the source object was deleted in the deployment that shared it; the item row in the Library shows the same red lineask the owner to share it again from a deployment that still has it
"not a member of this resource store"your Library access was removed between requesting and processingask an Organisation administrator to add you as Contributor or Reader on the Library's Members tab
"(name) is still building from an earlier pull", "cancel it or wait"a knowledge graph you pulled earlier is still importingwait for that build, or cancel it from the graph's Versions table, then pull again
"unsupported object_type"the Library holds a type Studio cannot pullnothing to pull
any other sentence, sometimes followed by "cleanup incomplete" and a listthe transfer failed after its retries; the sentence is the worker's own errorpull again; if it repeats, send the sentence and the item name to support

Related messages:

  • A Done row that reads "N dependencies need attention - open the copy to wire them." landed, but some of its agents, connectors or graphs could not be matched; open the copy and wire them.
  • A pulled connector arrives as Needs attention until you re-enter its credentials with Update.
  • On the sharing side, "You don't have Contributor access to any (type) library yet. Ask an Organisation admin to add you to one." means you can read but not push; the Shared to strip shows "rejected:" with the reason when a share is refused, or "rejected: no reason was recorded".

The Service halted banner​

A red banner reading Service halted followed by a message means Studio has stopped serving product requests: every chat and workflow request answers HTTP 451, Send and Run Now are disabled, and the banner cannot be dismissed. Studio checks its licence with the AlphaAgent Console every five minutes and, while blocked, retries every 30 seconds; the banner clears on its own on the next good check.

MessageCauseWhat to do
"AlphaAgent cannot reach the AlphaAgent Console and has been disabled. Contact your administrator."the licence check did not complete: no outbound path from the deployment to the Console's telemetry host, a response that was unsigned or failed verification, or no check has succeeded for more than 12 minutesan administrator confirms the deployment's outbound HTTPS to the Console host and, in the Console, the licence key's heartbeat health; the banner clears within a minute of the next good check
"AlphaAgent has been disabled by your administrator. Contact your administrator to restore access." or one of the Console hold messages (for example "AlphaAgent has been suspended due to an unresolved billing issue. Contact your administrator to restore access.")the Console answered that the licence is not valid: an enforcement hold or an inactive keythe administrator checks the licence key and the subscription in the Console; once cleared, Studio resumes on its own
"This license has been revoked. Contact your administrator."the licence key was revokedthis state does not clear on its own; the deployment needs a valid licence
"AlphaAgent Studio is unavailable: the AlphaAgent Console license phone-home check has failed. Contact your administrator."the licence monitor itself is not runningcontact support with the deployment id and the time; an administrator can also restart every task with Pause then Resume on the deployment

Also on Profile the Licence row reads "Blocked" with the same message. An amber Billing alert: banner with a countdown pill "Freezes in Nd Nh" and, when present, a Pay invoice link is a warning, not a halt: settle the invoice before the countdown reaches zero.

The workflow API refused my call​

ResponseMeaningWhat to do
503 deployment_updating with details.reason not_governed and "Programmatic access requires a governed deployment."the deployment is standardthe API is available on governed deployments only
503 deployment_updating, Retry-After: 60no key has ever been mirrored to the deployment (not_enabled), or its key store could not be read (keys_table_unreachable); the API keeps serving through an upgrade of the deploymentwait and retry
401 unauthorizedthe key or the key external id is missing, malformed, wrong, revoked, expired or for another deploymentcheck both headers against the values shown once at creation; an administrator can Rotate the key
403 forbiddenyour source IP is outside the key's allow-list, or the key lacks the scope this route needsan administrator edits the key
403 outside_time_windowthe call is outside the key's allowed hours (UTC)call within the window in details.valid_hours_utc
429 rate_limited, quota_exceeded, too_many_active_runsthe key's per-minute, per-day or concurrency limitwait for Retry-After; see Limits and quotas
409 workflow_not_ready with details.reasonsthe workflow's active version cannot run yetfix the listed reasons in Studio
400 invalid_parameters with detailsa run parameter failed validationcorrect the named parameters

Every error carries request_id; quote it to support. The full catalogue is on Limits, errors and codes.

Notes​

  • After the remedy, the same surface shows the change: a preflight rerun reports every check as PASS, WARN or SKIP and the install continues; a resumed job's step turns Succeeded in the step timeline; a re-pulled item turns Done in My Pulls; a throttle strip disappears; the Service halted banner clears on the next good licence check.
  • Everyday Studio questions (the 256-per-type cap, PDF limits, an environment that needs attention, a disabled Create button) are on FAQ.
  • If a message is not on this page, or the row's remedy did not work, write to support with the details on Support: the exact text, the Diagnostic ID and deployment id for a job, the request id for an API call, and when it happened.