Failure codes
Look a code up here when a deployment job stops or the workflow API refuses a call: the Organisations job codes come from the failure card's Technical details, the API codes from error.code in the response body.
Organisations job failure codes
The failure card on a deployment's Jobs view is headed <step> failed (or This job failed). Under the headline it names Who acts: Customer account, AlphaAgent or AWS (or Unknown when the console cannot tell). Every card carries the failure code, a copyable Diagnostic ID (the job id) and the actions the server offers: Resume re-runs from the failed step, Redo re-runs every step (each is idempotent), Abandon marks the job cancelled without touching your account, Wipe tears down and reinstalls. How to read the card, the step timeline and parked steps is on Jobs, progress and failures.
The tables quote the message and remediation exactly as the product records them. Resumable says whether Resume is offered; a job that is not resumable still offers Redo where a redo could help, and Abandon always. Where the card shows a Resumable line, its text may be more specific than the remediation quoted here, because it is written at the moment of failure.
Customer account (22 codes)
Something in the target AWS account, its DNS or its identity provider has to change before the job can succeed.
| Code | Resumable | What the card says | What to do (the product's own text) |
|---|---|---|---|
settle_service_never_stabilised | Yes | One or more ECS services never reached steady state before the settle timeout. | The named service is starting and then failing, most often a container health check that never passes. Read that service's task logs: they say why. A self-hosted Neo4j reporting 'The client is unauthorized due to authentication failure' every 30s is the known case - its password is only read on the FIRST boot of an empty volume, so the stored credential and the database can diverge permanently. To recover it: clear /data/dbms/auth.ini on the data volume and restart the service. The graph itself is untouched. ORDER MATTERS - on that restart the database adopts whatever the secret says AT THAT MOMENT, so do not change the secret afterwards or you will have diverged it again. Fix the cause, then resume the job. |
settle_service_rolled_back | Yes | One or more ECS services failed their new deployment's health checks and were automatically rolled back to their previous task definition. | The previous revision is already back and serving - nothing is down. Check the named service's task logs for why the new revision failed its health check, fix the underlying cause, then resume the job to try the update again. |
access_denied_missing_action | Yes | AWS denied an operation because the provisioning role is missing an IAM action. | Grant the missing action, named in this failure's details, to the provisioning role in the target account, then resume the job. |
credentials_unusable | Yes | The target account's provisioning role could not be assumed, or the session it returned was rejected. | Check that the provisioning role in your account still exists, that its trust policy still allows AlphaAgent's provisioning identity to assume it, and that the external ID is unchanged. Then resume the job. |
preflight_blocked | Yes | A preflight check failed, so the job stopped before changing anything in the account. | Resolve the specific check named in this failure's details - the commonest are region availability, a service quota and a missing prerequisite - then resume the job. Nothing was provisioned, so there is nothing to clean up. |
approval_rejected | No | A human rejected the change this job was waiting on. | None. A rejection is a decision, not a fault. Start a new job if the change is still wanted. |
park_expired | No | The job parked waiting for an answer and the park expired before one arrived. | Start a new job. It will park again at the same point, and the answer given then will be acted on immediately. |
quota_exceeded | Yes | An AWS service quota in the target account was exceeded. | Request an increase for the quota named in this failure's details, wait for AWS to grant it, then resume the job. The commonest is the Fargate concurrent vCPU quota (L-3032A538). |
restore_point_unwritable | Yes | The restore point could not be written to the deploy bucket, so the update was not started. | Grant the provisioning role s3:PutObject on the deploy bucket's restore-points/ prefix, then run the job again. The update stops before it changes anything in your account: without a durable restore point, there would be nothing to roll back to if the update failed partway through. |
wipe_leftovers | No | Clear account finished its passes, but the final check still found AlphaAgent resources in the account. | The resources still standing are listed under 'leftovers' below. Redo re-runs the whole clearing pass (every step is idempotent) - a second pass clears most leftovers. Or Abandon this job, then Scan for stray resources on this page and Wipe what it finds. If the same resource survives a second pass, open it in the AWS console: a bucket with Object Lock, a resource-level SCP or a cross-account share is what stops a delete, and it has to be lifted before the wipe can remove it. |
stack_resource_already_exists | Yes | A CloudFormation stack could not create a resource because one with that fixed name already exists in the account. | The details name the resource (failed_resource) and CloudFormation's own reason (failed_reason). It is pre-existing - left by a deployment that was deleted or failed in this account and region, or another deployment using the same environment name. Remove it (Accounts > Scan for stray resources lists it) or choose a different environment name for this deployment, then resume: the engine sweeps this attempt's retained resources and re-creates the stack. |
stack_resource_access_denied | Yes | Your AWS account denied a permission a CloudFormation stack needed. | The details name the resource (failed_resource, failed_resource_type), the denied action (denied_action), the role that was refused (denied_principal) and CloudFormation's own words (failed_reason). Allow that action for that role in this AWS account - in its IAM policy, and in any Service Control Policy applied to the account at your AWS Organization level - then resume the job. |
environment_in_use | Yes | Another deployment in this AWS account already uses this environment name. | Studio names its IAM roles and the ElastiCache subnet group by environment (alphaagent-<environment>-...), and IAM role names are account-wide, so one AWS account holds one deployment per environment name. The details list the resources that already carry this name. Choose a different environment name for this deployment (edit the draft's environment field, or create a new draft), then resume - nothing was created. |
preexisting_named_resources | Yes | Resources the Studio core template creates by fixed name already exist in the target account and region. | They are left over from a deployment that was deleted or failed in this account and region (the release's own templates name them; the details list each by kind, name and CloudFormation logical id). Remove them by hand - Accounts > Scan for stray resources shows the same set - then resume. 'Clear account' also removes them, but it deletes EVERY AlphaAgent resource in the account's first allowed region, so use it only when nothing else of AlphaAgent's must survive there. |
registration_stack_stale_for_cloudfront | Yes | This account's registration stack (alphaagent-org-target-<account>) predates the CloudFront permissions (TemplateVersion 2.3.0) or does not allow us-east-1, so the front door cannot be created there. | Update the account's alphaagent-org-target-<account> CloudFormation stack from the current template (Accounts > the account > registration template), keeping us-east-1 in AllowedRegions, then resume the job. Or turn the CloudFront front door off for this deployment. |
release_lacks_cloudfront_support | No | This Studio release's bundle carries no cloudformation/edge.yaml, so it cannot front its load balancer with CloudFront. | Install or upgrade to Studio 1.0.150 or later, or turn the CloudFront front door off for this deployment and launch again. |
cloudfront_alias_conflict | Yes | Another CloudFront distribution already claims this deployment's app domain as an alternate domain name, so a second one cannot be created. | Find the distribution holding the app domain (this account's CloudFront console, or another account of yours) and remove the alias or delete it - a disabled-but-undeleted distribution still holds it - then resume the job. |
edge_certificate_pending_validation | Yes | The us-east-1 certificate CloudFront needs for the app domain was requested but did not validate in time. ACM validates it through a DNS record; when the regional certificate was DNS-validated the same record already exists and this completes on its own. | Add the DNS validation record named in this failure's details (CNAME name -> value) at your DNS provider - it is the same record ACM asked for when you requested the regional certificate - wait for the certificate to show ISSUED in us-east-1, then resume the job. |
dns_points_at_alb_with_origin_lock | Yes | The app domain still resolves to the load balancer while the origin lock is on, so direct visits answer 403: CloudFront is not in the path. | Point the app domain's CNAME at the distribution domain shown on the deployment's Overview (DNS target), wait for it to propagate, then run Verify again. Or turn the origin lock off with a reconfigure (origin_lock: off) to serve directly from the load balancer meanwhile. |
origin_lock_refused_dns_at_alb | Yes | The origin lock was not turned on: the app domain still resolves to the load balancer, and locking it now would take the site dark. | Point the app domain's CNAME at the distribution domain shown on the deployment's Overview (DNS target), wait for it to propagate, then run the reconfigure (origin_lock: on) again - the engine also locks on its own at the next upgrade or Verify once the domain resolves to the distribution. |
app_domain_points_at_previous_distribution | Yes | The app domain's DNS record still points at a CloudFront distribution that is not this deployment's - typically the distribution of a previous, uninstalled deployment under the same domain. CloudFront refuses to give the new distribution the domain while that record stands. | At your DNS provider, point the app domain's CNAME at a non-CloudFront target (the load balancer hostname, or any placeholder) or remove the record, then Resume. After this job completes, point it at the new DNS target shown on the deployment's Overview - the new distribution's domain does not exist until the stack creates it, so it cannot be re-pointed up front. |
saml_configuration_failed | Yes | AlphaAgent could not federate this deployment against the supplied identity provider metadata (invalid or unsafe XML, no signing certificate, or an expired certificate). | Check the error named in this failure's details - it names exactly what was wrong with the pasted metadata (invalid XML, no signing certificate, or an expired certificate). Fix the export from your identity provider and paste it again. |
AWS (1 code)
A transient or capacity condition on the AWS side; wait, then Resume.
| Code | Resumable | What the card says | What to do (the product's own text) |
|---|---|---|---|
redis_busy | Yes | ElastiCache would not delete the deployment's Redis replication group: it, or one of its member cache clusters, is in a state (snapshotting, modifying or creating) that refuses deletion. | ElastiCache finishes a snapshot or modification on its own, usually within minutes. Wait until every member of the replication group named in the details reads 'available' in the ElastiCache console, then Resume the job: it re-enters this step, re-issues the delete and waits again. Nothing else needs to change. |
AlphaAgent (47 codes)
The card reads "AlphaAgent is investigating this — no action is needed from you." Where the text says to Resume, do so; otherwise share the Diagnostic ID with your AlphaAgent contact.
| Code | Resumable | What the card says | What to do (the product's own text) |
|---|---|---|---|
stack_create_failed | Yes | A CloudFormation stack did not reach CREATE_COMPLETE. | The failed stack's events name the resource and the reason. A failed CREATE leaves an unusable stack; on resume the engine deletes the resources its rollback retained (tables, secrets, buckets, repositories, file systems), then the stack, then re-creates it - resume the job once the underlying cause is understood. |
stack_update_rollback_complete | Yes | A CloudFormation stack update failed and rolled back cleanly. | The stack rolled back to its last good state automatically and is still live and serving - nothing was deleted. Read the failed stack's events, which name the resource and the reason, fix the underlying cause, then resume the job to try the update again. |
husk_retained_sweep_incomplete | Yes | A failed stack CREATE left resources its rollback retained, and not all of them could be deleted before re-creating. | The failure names each resource the engine could not delete. Delete them in the AWS console (or grant the missing permission the message names), then resume the job - the unusable stack was deliberately left in place so the resume re-runs this cleanup. |
husk_sweep_refused_stack_was_live | No | A stack in a disposable state once reached CREATE_COMPLETE, so its retained resources are treated as data and left alone. | Inspect the stack's retained resources in the AWS console. Only if they really are disposable, delete them and the stack by hand, then resume the job. |
stack_delete_failed | Yes | A CloudFormation stack did not delete cleanly. | The stack is DELETE_FAILED: CloudFormation removed everything it could and left the resources named in the technical details (failed_resources) in place - nothing was created, and nothing already deleted comes back. Each one is still held by something outside the stack; once that is released, resume the job and the delete continues from where it stopped. |
core_stack_blocked_by_lambda_enis | Yes | The core stack cannot delete while Lambda network interfaces remain on its security group. | This deployment's environment Lambdas have been deleted, but AWS had not yet released their VPC network interfaces, and the core stack's Lambda security group and private subnets cannot delete until it does. AWS releases them on its own schedule - usually within minutes, occasionally up to about twenty. Resume the job to wait again; the technical details name the security group and how many interfaces remain. |
stack_update_rollback_failed_wedged | No | A CloudFormation stack is UPDATE_ROLLBACK_FAILED and could not be auto-recovered. It holds live, possibly data-bearing resources. | In the CloudFormation console use 'Continue update rollback', optionally skipping the resource named in the failure, then resume the job. The stack is never deleted automatically: it holds live resources, and a delete here is unrecoverable. |
image_tag_missing | Yes | The release's container image tag is not present in the registry. | Confirm the release published every image tag it declares, then resume the job. Do not retag by hand: the deployment records the tag it installed and a hand-made tag makes the record wrong. |
lease_lost_max_reclaims | No | The job's worker lease was lost and reclaimed too many times. | This needs AlphaAgent to look into why the job keeps losing its lock before it is retried - repeated reclaims mean the worker running it keeps dying partway through a step, and each reclaim re-enters that step from the start. Share the diagnostic ID below with your AlphaAgent contact. |
verify_failed | Yes | Provisioning completed but the post-install verification did not pass. | Read the failed verification, named in this failure's details. The resources exist, so resuming re-runs verification without re-provisioning. |
release_bundle_incomplete | No | The release bundle is missing a member the Organisation must have: image.tar.gz or RELEASE.txt. | Confirm the release that produced this bundle actually completed, then re-run the mirror. If the artifact is genuinely broken, retire the version - the row stays as the record that it existed. A missing env-image.tar.gz on its own is not this failure: older releases predate it, and that alone is only a warning. |
release_bundle_download_failed | Yes | The release bundle could not be downloaded from the Console API. | Retry the mirror. Every attempt re-requests a fresh presigned URL, so an expired signature is not the cause - check the organisation's egress to the Console API and that its machine credential is still valid. If this release was never marked downloadable, retrying is futile: Console only ever presigns a channel's current latest, so this release cannot be mirrored and needs AlphaAgent to look at it directly. |
release_bundle_upload_interrupted | Yes | The release bundle's upload into the release cache was aborted underneath the mirror by another process. | Resume the job. It re-mirrors from scratch or finds the release already cached by the job that finished. If this repeats, two mirrors of the same release are running at once - the release-cache row's stage lease is meant to prevent that; AlphaAgent needs to look at the provisioner's release.mirror_wait / bundle events. |
image_digest_immutable_conflict | No | A Studio image tag already exists with different content, and the repository does not allow a tag to be overwritten. | Do not re-push. ECR uploads every layer before refusing an immutable tag, so the attempt costs the full transfer and fails anyway. Establish which build is the real one for this version - the failure names both digests - then either retire the version or re-release under a new one. alphaagent-org/studio is IMMUTABLE deliberately: a Studio tag is a version, and versions are unique. |
image_architecture_mismatch | No | A copied image manifest reports the wrong CPU architecture. | Check that the release built the code-interpreter env image with --platform linux/amd64, and that the image moved by manifest-level copy rather than load-and-re-push - a re-push adopts the pushing host's architecture, and the mirror task is arm64. A wrong-architecture env image fails at agent code-execution time, not at deploy time. |
image_push_failed | Yes | A registry push or copy failed for a reason that is not a permission or an immutable tag. | Read the registry output in this failure's details. The commonest cause is network reachability from the mirror task's private subnets, which need the ecr.api and ecr.dkr endpoints AND the S3 gateway endpoint — ECR layer PUTs redirect to S3. |
image_registry_unreachable | Yes | The registry could not be queried, so it is unknown whether the image is already there. | Retry. This is deliberately NOT treated as 'the tag is absent': crane exits non-zero for both, and only MANIFEST_UNKNOWN means absent. Check the VPC endpoints if it recurs. |
image_config_unreadable | No | The copy tool returned an image config that could not be parsed. | Check the pinned copy-tool version in the alphaagent-org/mirror image. A tool whose image-config output is not JSON is not the version this was built against - re-pin the known-good version and retry. |
image_digest_mismatch | Yes | The image in the target account is not the image the Organisation recorded, so the task definition would deploy unverified bytes. | The job stopped BEFORE the task-def stack was touched - once CloudFormation commits a task definition, ECS pulls whatever the tag resolves to. Check whether anything else pushed to that repository and tag around the same time: a concurrent deployment, or a manually run provisioning tool. The failure names the expected and found digests; resume once they agree. |
no_restore_point | Yes | The update failed and there is no restore point to roll back to. | Do not expect an automatic rollback: nothing recorded what the services were running before this update, so re-pointing them would be a guess. Read the failure above, fix the cause, and run the update again - it resumes from the step that failed. To revert instead, deploy the previous release explicitly. |
rollback_failed | No | The update failed and the automatic rollback also failed. | The services may be running a mixture of the old and new task definitions. The restore point named in this failure's details lists what each service should be running; it is in the target account's deploy bucket and outlives this job. Fix the credentials or permissions the rollback failed on, then run a rollback job. Do not re-run the update until the fleet is consistent. |
connector_backup_corrupt | Yes | The connector-credential backup selected for restore is not valid JSON. | Pick a different backup by its explicit S3 key - the restore lists what is available - or have the connector owners reissue the credentials. The engine refuses to write unparseable bytes over a live secret, because that would turn a recoverable problem into an unrecoverable one. |
wipe_stack_force_delete_failed | Yes | A CloudFormation stack would not delete even when a forced delete was attempted. | Read the stack's events in the CloudFormation console for the resource named in this failure's details, resolve it by hand, then re-run Clear account - every other resource type this sweep found is deleted independently of this stack, so re-running only retries what is still standing. |
internal_error | No | The engine hit a condition it does not have a name for. | Share the diagnostic ID below with your AlphaAgent contact. This is a failure the platform does not have a specific reason for yet, so AlphaAgent needs to look at the job directly rather than this pointing you to a fix. |
minimal_stack_not_settled | Yes | The minimal core stack's doctor re-check found it in a different state than the core step that created it reported. | This means the stack changed after it was created and before the routine post-install check ran - most likely because someone edited or updated it by hand while the job was still running. Read the CloudFormation stack's current parameters and events directly, then decide whether to resume or to delete and relaunch. |
studio_step_not_implemented | No | This deployment type does not support this step yet. | Retrying will not help, since it re-enters the same unsupported step. Only the minimal deployment type can be launched today - contact AlphaAgent for the current status of the full Studio deployment type. |
missing_required_studio_field | No | This deployment is missing configuration data it needs to launch (the Cognito prefix, the app domain, or the licence secret). | This is a defect in how the deployment was launched, not something wrong with your account. Share the diagnostic ID below with your AlphaAgent contact, along with the field named in this failure's details. |
license_token_unreadable | Yes | AlphaAgent could not read the licence token that was saved for this deployment from its own secrets store. | Check that the secret named in this failure's details still exists in AlphaAgent's own account (not the target account), and that the service account reading it is configured with secretsmanager:GetSecretValue. This is resumable once the secret is reachable again. |
release_not_deployable | Yes | The release version requested for this deployment could not be prepared for install. | This means the release version genuinely could not be fetched or pushed: it may have been retired, Console refused to presign it because it isn't this channel's current latest, or a previous mirror attempt used up its retries. Check that the version was launched from a real Studio release, and check that channel/version's record in the release cache for its recorded failure. |
console_credential_invalid | No | AlphaAgent's own credential for talking to Console is missing or malformed, so no request to Console can be authenticated. | This needs AlphaAgent to re-seed its own Console credential for this organisation. Not fixable by retrying. |
console_unreachable | Yes | A network-level failure reaching Console's telemetry API - DNS, TLS, connection refused/reset, or a timeout. | Check the organisation's egress (NAT/security groups) to Console and retry. This is transient by nature - if it keeps happening across repeated attempts, share the diagnostic ID with AlphaAgent. |
console_error | Yes | Console rejected a catalogue or download request for this release - most often because the requested version is not this channel's current latest, or is not a real release for the requested product at all. | Check Console's catalogue for this exact channel/version/product before retrying blindly - a 403/404 here is not fixed by retrying the same request. |
console_response_unverified | Yes | Console's response failed a security check meant to catch tampering in transit - either a genuine interception, or the response was altered along the way. | Do not trust the URL or catalogue data in that response. Retry - if this keeps happening, share the diagnostic ID with AlphaAgent straight away, since this check exists specifically to catch a spoofed Console. |
release_mirror_misconfigured | Yes | {project}-common-config is missing CONSOLE_BASE_URL or ORG_BUNDLE_CACHE_BUCKET, so the release mirror has nowhere to fetch from or cache into. | Seed both keys in the {project}-common-config secret. Not fixable by retrying until that secret is corrected. |
org_repository_missing | Yes | AlphaAgent's own container registry did not have the repository this step expects. | The Org's own alphaagent-org/{studio,studio-envs} repositories are provisioned by org-foundation.yaml, not by this step — a missing one is an Org-infrastructure gap, not a customer account problem. Check org-foundation.yaml deployed cleanly in the Org's own account. |
spa_bundle_empty | Yes | This release is missing its web application files, even though it is marked ready to deploy. | This is a mirror integrity problem, not something to route around. Do not resume past this - it needs AlphaAgent to fix the release before it can be retried. |
spa_bundle_unserveable | Yes | This release's web application cannot load in a browser: its main file depends on a second file that the server is unable to deliver. Installing it would leave a blank page. | Do not resume past this - the release itself is built wrong and needs AlphaAgent to correct it. Staying on the current version is the safe outcome. |
release_bundle_missing_secret_schema | No | This release's bundle does not carry a loadable configuration schema (tools/studio_secret_schema.py), so its services' config secrets cannot be seeded and it cannot be installed. | Nothing has been changed in the deployment. This release must be re-cut by AlphaAgent with the schema in its bundle; do not resume - the same bundle will fail the same way. Staying on the current version, or upgrading to a release whose bundle carries the schema, is the way forward. |
release_schema_minted_secrets_invalid | No | This release's configuration schema declares its minted secrets (MINTED_SECRETS) in a form this Org version cannot act on, so its signing and PII keys cannot be minted and it cannot be installed. | The seed step wrote nothing. Upgrade the Org console to a version that understands this release's schema, or install a release whose schema this version understands; do not resume - the same schema will fail the same way. |
release_bundle_missing_template | No | This release's bundle does not carry one of its CloudFormation templates (cloudformation/<name>.yaml), so its stacks cannot be deployed from it and it cannot be installed. | Nothing has been changed in the deployment. This release must be re-cut by AlphaAgent with its templates in the bundle; do not resume - the same bundle will fail the same way. Staying on the current version, or upgrading to a release whose bundle carries them, is the way forward. |
edge_outputs_missing | Yes | The stateless stack needs the edge stack's function version and the minted us-east-1 certificate, and neither was available. | Resume the job: the edge step runs again and hands its outputs on. If it fails the same way twice, share the diagnostic ID with your AlphaAgent contact. |
front_door_probe_failed | Yes | The deployment's CloudFront distribution did not answer the sign-in callback probe the way the front door is meant to (a branded HTML 401/403), or the load balancer did not honour the origin lock. | Check the distribution's status in the CloudFront console and the load balancer's listener rules, then resume the job; nothing else changes until this passes. Share the diagnostic ID with your AlphaAgent contact if it fails again. |
saml_auto_provisioning_failed | Yes | AlphaAgent's automatic SSO provisioning against your connected identity provider did not complete. | There is nothing on your AWS account or your identity provider for you to fix directly here - this is AlphaAgent's own automation talking to Microsoft Graph and Cognito. Most failures at this step are a few seconds of timing lag between the two systems and clear on their own: click Resume. If it keeps failing after several resumes, this is a platform-side problem - contact support with the job ID and the error named in this failure's details. |
entra_credentials_unavailable | Yes | AlphaAgent's own credential for automatic single sign-on setup with Microsoft Entra ID is missing or incomplete, so this step cannot authenticate to Microsoft on your behalf. | This is AlphaAgent's own platform credential, not something you can fix directly - it needs AlphaAgent to repair the credential, then resume. No action is possible on your side here; share the diagnostic ID with your AlphaAgent contact if it doesn't clear soon. |
doctor_common_config_missing | Yes | AlphaAgent's platform configuration for this deployment is missing, so none of its services could start with real settings. | This is a fatal configuration gap, different from an ordinary per-service one: re-run seed-config, then resume doctor. Unlike a per-service gap, this is not something an external-Neo4j deployment can be waiting on. |
reseed_retention_drift | Yes | The deployment's recorded run-data retention still disagrees with what its storage lifecycle rule enforces after the re-seed wrote it. | Resume the job: the step re-reads the Core stack and patches RUN_DATA_RETENTION_DAYS in the deployment's common-config again. If it fails twice, another writer is changing that secret - check the secret's version history in the customer account before resuming. |
reseed_pii_default_drift | Yes | The deployment's PII redaction default still disagrees with the mode the fleet template or reconfigure asked for after the re-seed wrote it. | Resume the job: the step re-reads the deployment's common-config and patches PII_REDACTION_DEFAULT again. If it fails twice, another writer is changing that secret - check the secret's version history in the customer account before resuming. |
Where a job failure comes from
- A permission refused on a direct call. When your AWS account refuses an action the provisioning role calls directly, the failure is
access_denied_missing_actionand its details name the missing action (missing_iam_action). A refusedcloudformation:CreateStackorcloudformation:UpdateStackcall is reported the same way. - A permission refused inside a CloudFormation stack. Since Organisations 1.0.5 a resource that could not be created, updated or deleted because a policy refused it is
stack_resource_access_denied; the details carryfailed_resource,failed_resource_type,denied_action,denied_principalandfailed_reason. Before 1.0.5 the same event surfaced as a stack failure attributed to AlphaAgent. Stack failures whose reason is not a deny still arrive asstack_create_failed,stack_update_rollback_completeorstack_delete_failed, with the first failed resource and its reason in the details. - A readiness check that fails before anything is created parks the job with the check's own remediation rather than failing it; preflight checks therefore do not appear in the table. Which guardrail produces which code is on SCPs and guardrails that block installs.
Workflow API status codes and error codes
Every error, from every route, has the same body:
{
"error": {
"code": "invalid_parameters",
"message": "One or more parameters are invalid.",
"request_id": "<request_id>",
"details": [{"name": "applicant_id", "reason": "required"}]
}
}
code is stable and meant for your program; message is for a person; request_id is echoed in the X-Request-Id header, quote it to support; details appears when there is something structured to add. A 401 also carries WWW-Authenticate: Bearer; a 503 always carries details.reason. Retry and back-off rules, the per-key limits and every Retry-After value are on Limits, errors and codes.
| HTTP | code | When | details |
|---|---|---|---|
| 400 | invalid_request | Malformed JSON or a missing or malformed Idempotency-Key | [{name, reason}] for a body that fails validation; for malformed JSON one entry, {"name": "body", "reason": "JSON decode error at byte <n>: <why>"} |
| 400 | invalid_parameters | The run's parameters fail the version's contract | [{name, reason}], one per problem |
| 400 | invalid_filter | A GET /runs filter is not valid (a date that is not ISO-8601, a bad client_reference) | |
| 400 | invalid_cursor | A paging cursor is not one the API issued | |
| 400 | policy_not_tightening | pii_policy_override loosens the workflow's policy | {field} |
| 400 | invalid_pii_policy | pii_policy_override is not a valid policy | {field} |
| 401 | unauthorized | Key or external id missing, malformed, unknown, wrong, revoked, expired, or bound to another deployment. Always the same message | |
| 403 | forbidden | Source address outside the allow-list, or the route's scope is missing | |
| 403 | outside_time_window | Outside the key's allowed hours | {valid_hours_utc: {start_hour, end_hour}} |
| 404 | workflow_not_found, run_not_found, output_not_found, segment_not_found | Not found, or not within this key's reach; the two are indistinguishable | |
| 404 | pii_redaction_not_applied | GET /runs/{run_id}/pii-redaction on a run whose redaction policy was off | |
| 404 | not_found | A path that does not exist under /api/v1, or a run whose board or event log no longer exists on the runtime | |
| 405 | method_not_allowed | The path exists, the method does not | |
| 409 | idempotency_conflict | The same Idempotency-Key with a different body; or the first request is still being processed | |
| 409 | workflow_not_active | The workflow exists but is a draft or inactive. An all-active-workflows key gets this on GET /workflows/{workflow_id}, /parameters and POST …/runs; an explicit-list key reads a listed draft (200 with readiness.reasons) and gets this only at POST …/runs | |
| 409 | workflow_not_ready | A step has no agent, a connector is missing, needs attention or has not passed Test connection, or an environment is not active | {reasons: [string]} |
| 409 | run_not_terminal | Outputs asked for, or erasure requested, while the run is still going | {status} |
| 409 | redaction_pending | The run finished but the final redaction sweep of its outputs has not | {status} |
| 409 | already_terminal | Cancel of a run that has already finished | |
| 409 | already_purged | Erasure of a run already erased | {data_purged_at} |
| 413 | use_bundle | One output is over 25 MB | {bundle}: the zip path |
| 413 | bundle_too_large | The zip would exceed 500 MB or 2,000 files | |
| 413 | payload_too_large | A request body over 1 MiB, refused before it is read (and before the key is checked) | |
| 416 | range_not_satisfiable | The Range header does not fit the file; Content-Range: bytes */<size> is sent | |
| 422 | version_not_found | The version you asked for does not exist | |
| 429 | rate_limited, quota_exceeded, too_many_active_runs | A per-key limit is full; honour Retry-After | |
| 451 | license_unavailable | The deployment's licence is on hold, so its runtime refuses work; the message is the one your administrator sees in the Organisations console | {state, reason, enforcement_status, license_id} |
| 500 | internal | Unexpected failure; "Internal error; quote request_id to support." | |
| 502 | runtime_error | The deployment's runtime answered with an error | |
| 502 | runtime_unreachable | The deployment's runtime did not answer; "retry shortly" | |
| 503 | deployment_updating | Programmatic access is not being served: not_governed (a standard deployment; permanent), not_enabled (no key mirrored yet), or keys_table_unreachable (the key store could not be read) | {reason} |
Failure codes on a failed run
These are not HTTP errors; they are the error.code on GET /runs/{run_id} when status is failed.
error.code | Meaning |
|---|---|
input_materialisation_failed | An s3_prefix input could not be copied in before any step ran |
node_failed | A step failed; GET /runs/{run_id}/steps shows which |
wall_clock_limit_reached | The run reached the workflow's time limit |
invariant_blocked | A step was blocked and the run could not continue |
launch_failed | The run never reached the runtime |
orphaned | The runtime lost track of the run |
license_unavailable | The runtime refused the launch because the deployment's licence is on hold |
Notes
- Preflight checks that fail park the job instead of failing it; the card's remediation is the check's own text. The fifteen
orgctl.pypreflight checks and the installer's exit statuses are on Troubleshooting. - No code names an SCP, a permissions boundary or an opt-in; the copy always names the denied action or the failed CloudFormation resource. Map it back to your policy yourself.
- A 404 from the API means not found or not within this key's reach; the two are indistinguishable.
Related
- Jobs, progress and failures: the Jobs view, step timeline, parked steps and the retry dialog.
- SCPs and guardrails that block installs: the nineteen guardrail patterns and the codes each produces.
- Limits, errors and codes: the API's limits,
Retry-Aftervalues and the OpenAPI file. - Troubleshooting: refusals in the order you meet them.
- Support: what to include when you quote a Diagnostic ID or a
request_id.