Skip to main content

SCPs and guardrails that block installs

An AlphaAgent install creates the resources listed on Everything AlphaAgent deploys with the permissions you grant. A service control policy (SCP), a permissions boundary, a tag rule or a region allow-list that denies one of those actions stops the install with copy that names the denied action or the failed CloudFormation resource, never the policy itself. This page lists the nineteen patterns, the copy each one produces, the checklist to run before you start, and how to resume once the policy is fixed.

Who this is for​

The cloud engineer or security reviewer who owns the guardrails in the AWS accounts you will install into.

Before you start: the checklist​

Run through this for the Organisations account and for every Studio account. Where a line says "both regions" it means the account's install region and us-east-1.

  1. Regions. The region allow-list includes the install region for the Organisations account, and both regions for every Studio account (the CloudFront front door creates a Lambda@Edge function and an ACM certificate in us-east-1). The account template's AllowedRegions must also keep us-east-1.
  2. No service denies on: CloudFormation; EC2 including VPC, internet gateway, NAT gateway and Elastic IP creation; Elastic Load Balancing; ECS and Cloud Map; Auto Scaling; IAM role creation and iam:PassRole for roles named alphaagent* or AlphaAgent*; Lambda (both regions, including Lambda@Edge replication); CloudFront; ACM (both regions); ECR; DynamoDB; S3; Secrets Manager; SQS, SNS and EventBridge; ElastiCache; EFS; Cognito; CloudWatch and CloudWatch Logs; SSM; STS; Service Quotas; Bedrock at runtime; KMS through those services (AWS-managed keys only).
  3. No permissions-boundary requirement on iam:CreateRole: no template sets a boundary.
  4. No tag-enforcement rule on resource creation: the templates carry no common tag set, and the job's own s3:CreateBucket and secretsmanager:CreateSecret calls are untagged.
  5. S3 rules allow a bucket to be created and then have its public-access block applied as a second call (the deploy bucket), and do not require a customer-managed KMS key: every bucket uses SSE-S3.
  6. Marketplace is not denied in Studio accounts: Bedrock evaluates aws-marketplace:ViewSubscriptions and Subscribe on the first model call.
  7. Quotas: one Elastic IP, one VPC and one NAT gateway free per account; Fargate concurrent vCPU (quota L-3032A538) for the deployment's sizing tier; Bedrock tokens-per-day for the models you choose. None of these is checked before the stacks are created, except the Elastic IP, VPC and NAT headroom, which is advisory.
  8. Delete-side denies (SCPs protecting Delete* on resource types) will surface at upgrade, teardown and Clear account time, not at install.

How a blocked install shows up​

A deny reaches you through one of three channels. Which one decides the copy you see.

Before anything is created: a parked preflight. The Studio deploy job runs the account-side preflight checks first. Any check that cannot make its call fails with its own remediation (for example "Needs ec2:DescribeAvailabilityZones."), and the job parks instead of failing. On the job's progress screen the parked-step card is headed "Waiting for you" with the kind "Proceed despite failed readiness checks" and the prompt "Preflight found N failing check(s) against <account>: <check id>: <detail>; ... Override and continue, or fix the account and resume?". Its buttons are "Confirm and resume" (tick the override box and acknowledge the failed checks) and "Do not override". The card names who can unblock it: "Anyone willing to take responsibility for proceeding despite the failed checks." A rejection is final: "A rejection is final: the job fails with approval_rejected and cannot be resumed, because resuming it would apply the change the rejection exists to prevent. Start a new job if the change is still wanted." A park that is left alone expires; the card then reads "this park has expired. Start a new job; it will park again at the same point."

A direct API call the job makes is denied. The job fails with the code access_denied_missing_action. The failure card is headed "<step> failed"; its "Who acts" row reads "Customer account" with the hint "Something in the target AWS account has to change before this can succeed."; the message is "AWS denied an operation because the provisioning role is missing an IAM action."; "What to check:" is followed by "Grant the missing action, named in this failure's details, to the provisioning role in the target account, then resume the job."; and the details list below shows missing_iam_action (for example acm:RequestCertificate) and exception_message, the text AWS returned. A refused cloudformation:CreateStack or cloudformation:UpdateStack call is reported the same way.

A deny inside a CloudFormation stack. CloudFormation fails the resource and rolls the stack back. When CloudFormation's reason carries denial wording ("is not authorized to perform", "AccessDenied", "UnauthorizedOperation", "explicit deny" or "not authorized"), the job fails with stack_resource_access_denied, whether the stack was being created, updated or deleted. The card's "Who acts" row reads "Customer account"; the message is "Your AWS account denied a permission a CloudFormation stack needed."; "What to check:" is followed by "The details name the resource (failed_resource, failed_resource_type), the denied action (denied_action), the role that was refused (denied_principal) and CloudFormation's own words (failed_reason). Allow that action for that role in this AWS account - in its IAM policy, and in any Service Control Policy applied to the account at your AWS Organization level - then resume the job."; and the card's "What failed:" line reads "stack <stack> ended in <status>: <logical id> (<resource type>) could not be created because your AWS account denied the action <service:Action> to <principal>. CloudFormation said: <reason>" (or "could not be updated", or "could not be deleted"). The "Resumable" line reads "Allow the action <service:Action> for <principal> in this AWS account - in its IAM policy, and in any Service Control Policy applied to the account at your AWS Organization level - then resume the job:" followed by what the resume does: for a create, "the engine removes what this attempt left behind and re-creates the stack."; for an update, "the stack rolled back to its last good state and is still live, so the update is simply tried again."; for a delete, "the delete continues from the resources that were refused." When the reason names no action the card says "the action it needed" instead, and when it names no principal the card names the role the Organisation assumed. The details list shows failed_resource, failed_resource_type, failed_reason, denied_action, denied_principal, stack and stack_status (and failed_resources on a delete). Before Organisations 1.0.5 the same event was reported as a stack failure attributed to AlphaAgent. A stack failure whose reason is not a deny (a quota, a name collision, a resource still held by something outside the stack) still arrives as stack_create_failed, stack_update_rollback_complete or stack_delete_failed, attributed to AlphaAgent and resumable, with the first failed resource and its reason in the details.

Every card also carries: the failure code and its category; a copyable diagnostic ID ("Share this with your AlphaAgent contact if you need help."); a "Technical details" disclosure; and the actions the server offers. The Jobs list shows the same failure message, with the code as a tooltip.

The Organisations installer instead of a job. orgctl.py runs in your terminal, so a deny ends the command with a FATAL line: a denied direct call prints "AWS denied an operation (<code>) during "<phase>": your credentials are missing the IAM action '<service:Action>'. Grant it to the identity you are running as, then re-run the same command - re-run the same command to resume from where it stopped."; a deny inside a stack prints "<stack> finished in ROLLBACK_COMPLETE. The create FAILED and rolled back; the stack must be deleted before it can be created again. First failure: <resource> -- <reason> This step was not checkpointed, so fixing the cause and re-running resumes here." and python3 orgctl.py status labels the stack "stuck: ... cannot be updated - delete it, then re-run python3 orgctl.py install --restart-from <step>."

The nineteen patterns​

Each pattern says what in the install it breaks and the copy you would see.

Region deny: the deployment region​

Blocks: every call the job makes in the Studio account, starting with preflight (sts:GetCallerIdentity, ec2:DescribeAvailabilityZones, ec2:DescribeVpcs, acm:ListCertificates, cloudformation:DescribeStacks), then the deploy bucket, then every stack. The Organisations console already refuses a region outside the eleven supported ones with a validation error before a job starts; the target role's own DenyOutsideAllowedRegions is the in-product version of this rule.

You see: a parked preflight whose failed checks carry "Needs ec2:DescribeAvailabilityZones.", "Needs acm:ListCertificates and acm:DescribeCertificate.", "Needs cloudformation:DescribeStacks." or "No usable AWS credentials. ...". If you override, the next call fails with access_denied_missing_action and the denied action in missing_iam_action; a forbidden deploy bucket gives "deploy bucket <name> exists but is not accessible to the provisioning role." with "Grant the target account's provisioning role s3:ListBucket and s3:PutObject on arn:aws:s3:::<name>, or delete the bucket if it belongs to something else, then resume the job."

Region deny: us-east-1 specifically​

Blocks: the edge step (an ACM certificate requested in us-east-1, then the alphaagent-<deployment id>-edge stack with its Lambda@Edge function and role), the edge stack's deletion at teardown, Clear account's scan of us-east-1, and the account's Verify. AWS WAF is regional in the deployment region and is not affected.

You see: if the account template's AllowedRegions parameter omits us-east-1, or the template is older than 2.3.0, the job stops before creating anything in us-east-1 with registration_stack_stale_for_cloudfront: "This account's registration stack (alphaagent-org-target-<account>) predates the CloudFront permissions (TemplateVersion 2.3.0) or does not allow us-east-1, so the front door cannot be created there." and "Update the account's alphaagent-org-target-<account> CloudFormation stack from the current template (Accounts > the account > registration template), keeping us-east-1 in AllowedRegions, then resume the job." An SCP deny is invisible to that check and surfaces as access_denied_missing_action with missing_iam_action: acm:RequestCertificate, or, if the certificate succeeds, as stack_resource_access_denied on the edge stack with EdgeFunctionRole or EdgeFunction as the failed resource and the denied action in denied_action. At teardown the same deny gives stack_resource_access_denied on the delete, with the refused resources in failed_resources; Clear account logs "could not scan us-east-1 edge stacks: <error>"; Verify's edge-cleanup row fails with "The role could not list or delete CloudFormation stacks in us-east-1. Update the account's alphaagent-org-target stack from the current template (TemplateVersion 2.3.0, us-east-1 in AllowedRegions), then Verify again."

Service deny: CloudFront​

Blocks: the preflight's alias check (cloudfront:ListDistributions), the CloudFront distribution in the stateless stack, and Clear account's listing of distributions.

You see: at preflight, access_denied_missing_action with missing_iam_action: cloudfront:ListDistributions. At install, stack_resource_access_denied on alphaagent-<deployment id>-stateless with EdgeDistribution as the failed resource and the denied CloudFront action in denied_action. On an update, the same code: the stack rolled back to its last good state and is still live, and Resume tries the update again.

Service deny: Lambda@Edge​

Blocks: the edge stack's function, role and version in us-east-1 (a deny on lambda:* there, on lambda:EnableReplication, or on the edge role), and the stateless stack, which needs the function version.

You see: stack_resource_access_denied on the edge stack, with the denied lambda: action in denied_action and CloudFormation's reason in failed_reason. If the edge stack is absent when stateless runs: edge_outputs_missing, "The stateless stack needs the edge stack's function version and the minted us-east-1 certificate, and neither was available." with "Resume the job: the edge step runs again and hands its outputs on." At delete time, a DELETE_FAILED whose reason mentions the replicated function is not a failure: AWS removes Lambda@Edge replicas hours later, the card reads "Lambda@Edge replicas still deleting; Accounts > Verify retries", and the account's Verify re-issues the delete.

Service deny: Bedrock​

Blocks: nothing at install time. No step of the deploy job calls Bedrock, and no preflight check does. Studio's task role carries the invoke actions and the VPC has a bedrock-runtime endpoint; the deny bites when a service or an agent first calls a model.

You see: no failure code names Bedrock. The chat service's health check calls bedrock:ListFoundationModels; when that is denied the deny surfaces at the settle step as settle_service_never_stabilised ("One or more ECS services never reached steady state before the settle timeout.") or, on an update, settle_service_rolled_back ("One or more ECS services failed their new deployment's health checks and were automatically rolled back to their previous task definition."), both attributed to your account with the named service's task logs as the place to look. Model access that is not enabled behaves the same way; see Supported models.

Service deny: Cognito​

Blocks: the user pool, its domain and client in the core stack; the saml step; the cognito-prefix preflight check. The target role's DenyStudioUserImpersonation is the in-product deny on the user-impersonation subset and does not block anything the install needs.

You see: at preflight the cognito-prefix check warns rather than fails ("Could not check. Needs cognito-idp:DescribeUserPoolDomain."). At install, stack_resource_access_denied on core with the pool as the failed resource. At the saml step, saml_auto_provisioning_failed: "AlphaAgent's automatic SSO provisioning against your connected identity provider did not complete." with the AWS error text in the details ("SAML configuration failed: ClientError: An error occurred (AccessDeniedException) ..."); the card's "What to check" says the step is AlphaAgent's own automation and asks you to click Resume first, so read the details to see that the cause is a Cognito deny.

Service deny: aws-marketplace​

Blocks: nothing at install time; the actions appear only on the Studio task role, for Bedrock's benefit.

You see: not a job failure. Bedrock evaluates aws-marketplace:ViewSubscriptions and Subscribe on the first call to a Marketplace-hosted model; without them, that first call fails inside Studio with an AccessDeniedException naming the two actions.

ECR: cross-account pull or an ecr deny​

Blocks: the push-image step (the job copies the three Studio images into the account's own ECR repositories by digest, with one tool holding both registries' credentials), verify-images (ecr:DescribeImages), and the repositories created by core. Runtime pulls are same-account, so no cross-account repository policy is involved anywhere.

You see: a denied push gives access_denied_missing_action with "<image> was denied because the role is missing ecr:PutImage." and "Grant ecr:PutImage to the provisioning role for the three Studio ECR repositories in the target account, then resume the job. Do NOT grant ecr:PutRepositoryPolicy or ecr:CreateRepository ..."; a registry error whose text does not name an action gives image_push_failed: "A registry push or copy failed for a reason that is not a permission or an immutable tag." with "Read the registry output in this failure's details. The commonest cause is network reachability from the mirror task's private subnets, which need the ecr.api and ecr.dkr endpoints AND the S3 gateway endpoint." A denied DescribeImages gives access_denied_missing_action; denied repository creation gives stack_resource_access_denied on core with the repository as the failed resource.

Service deny: Secrets Manager​

Blocks: the 18 secrets in core, and the direct CreateSecret and PutSecretValue calls the seed-config and neo4j-creds steps make. The target role's grant is scoped to secrets named alphaagent*.

You see: stack_resource_access_denied with the first secret as the failed resource, or access_denied_missing_action with missing_iam_action: secretsmanager:CreateSecret or secretsmanager:PutSecretValue.

Service deny: KMS​

Blocks: nothing directly, because no customer-managed key is created or referenced: every bucket uses SSE-S3, and the target role's kms:* applies only through S3, Secrets Manager, DynamoDB, SQS, SNS, ElastiCache, EFS, CloudWatch Logs and ECR. An SCP that mandates a customer-managed key for S3 or Secrets Manager does block.

You see: the deny arrives as the calling service's deny: stack_resource_access_denied inside a stack, or access_denied_missing_action for a direct Secrets Manager or S3 call. A key-mandating S3 condition on the deploy bucket stops updates before they start with restore_point_unwritable: "The restore point could not be written to the deploy bucket, so the update was not started." and "Grant the provisioning role s3:PutObject on the deploy bucket's restore-points/ prefix, then run the job again. The update stops before it changes anything in your account: without a durable restore point, there would be nothing to roll back to if the update failed partway through."

iam:CreateRole or iam:PassRole deny​

Blocks: every Studio role (two in core, two in compute, one in stateless, one in edge), all named and all stemmed alphaagent, and the PassRole the ECS services and Lambda functions need. In the Organisations account, the three roles in alphaagent-org-app. The target role's DenyRoleMutationOutsideAlphaAgent is the in-product analogue and allows exactly these names.

You see: stack_resource_access_denied with the role's logical id (for example EdgeFunctionRole) as failed_resource, AWS::IAM::Role as failed_resource_type and iam:CreateRole as denied_action, or, for PassRole, the consuming service or function as the failed resource and iam:PassRole as the denied action; CloudFormation's own sentence naming the action is in failed_reason. On resume the engine removes what the failed create left behind and re-creates the stack. If a delete is also denied during that resume: husk_retained_sweep_incomplete, "A failed stack CREATE left resources its rollback retained, and not all of them could be deleted before re-creating." with "Delete them in the AWS console (or grant the missing permission the message names), then resume the job." In the Organisations installer: the FATAL stack-failure text above.

Permissions boundaries​

Blocks: every role creation in core, compute, stateless, edge and alphaagent-org-app, because no template sets a PermissionsBoundary. A boundary attached to the target role itself limits what the Organisation can do without Verify noticing: Verify tests sts:AssumeRole only.

You see: stack_resource_access_denied with the first role in the stack as the failed resource (core's Lambda execution role or compute's task execution role) and iam:CreateRole as the denied action. A boundary on the target role shows up as the individual denies in the rows above, after a passing Verify.

VPC, internet gateway or NAT creation denies​

Blocks: the VPC, internet gateway, Elastic IP, NAT gateway and ten VPC endpoints in core, and the same set in the Organisations foundation stack. The Docker Hub pulls and the Console egress depend on the NAT path.

You see: at preflight, availability-zones FAIL with "Needs ec2:DescribeAvailabilityZones." or vpc-cidr-overlap FAIL, so the job parks; quota-headroom warns "otherwise foundation fails with AddressLimitExceeded after ~10 minutes of work." when an Elastic IP is not free. At install, stack_resource_access_denied on core with Vpc, the internet gateway or the NAT gateway as the failed resource; on resume the engine removes what the attempt left behind and re-creates the stack. In the Organisations installer: the FATAL stack-failure text.

S3 public-access, ACL or ownership rules​

Blocks: rarely anything in the templates, because every template bucket already has all four public-access blocks on and no ACL. The deploy bucket is different: the job creates it and applies the public-access block as a second call, so a rule that requires the block on CreateBucket itself denies the first call.

You see: inside a stack, stack_resource_access_denied with the bucket as the failed resource. On the deploy bucket, access_denied_missing_action with missing_iam_action: s3:CreateBucket or s3:PutBucketPublicAccessBlock; an existing bucket the role may not touch gives "deploy bucket <name> exists but is not accessible to the provisioning role."

Tag-enforcement SCPs​

Blocks: creation of most resources: the templates set Name on some and nothing else, compute and taskdef carry no tags at all, and the job's own s3:CreateBucket and secretsmanager:CreateSecret calls are untagged. Only the us-east-1 certificate request is tagged.

You see: stack_resource_access_denied with the first untagged resource CloudFormation reached as failed_resource and the deny in failed_reason; or access_denied_missing_action naming s3:CreateBucket or secretsmanager:CreateSecret.

CloudTrail or Config-mandated settings​

Blocks: nothing by collision, because no template creates CloudTrail or Config resources. Rules enforced as denies at create time (for example requiring KMS encryption or versioning) act like the rows above; deploy-bucket versioning is best-effort and a refusal is tolerated.

You see: where a rule denies at create, stack_resource_access_denied or access_denied_missing_action as for that resource type. Where a rule reports non-compliance after the fact, nothing in the product observes it; there is no code path and no copy.

EFS or ElastiCache denies​

Blocks: the Redis replication group and subnet group and the EFS file system, mount targets and access points in core; the Redis delete at teardown; Clear account's sweep.

You see: at install, stack_resource_access_denied with the subnet group, replication group or file system as the failed resource. At teardown, a denied DeleteReplicationGroup gives access_denied_missing_action with missing_iam_action: elasticache:DeleteReplicationGroup; a state refusal instead gives redis_busy (attributed to AWS): "ElastiCache would not delete the deployment's Redis replication group: it, or one of its member cache clusters, is in a state (snapshotting, modifying or creating) that refuses deletion." with "Wait until every member of the replication group named in the details reads 'available' in the ElastiCache console, then Resume the job". On a resume after a failed create, husk_retained_sweep_incomplete names the resource it could not delete. In Clear account, the job log reads "<resource>: denied (missing <action> - re-run the account role template)" and the job ends with wipe_leftovers: "Clear account finished its passes, but the final check still found AlphaAgent resources in the account." with "If the same resource survives a second pass, open it in the AWS console: a bucket with Object Lock, a resource-level SCP or a cross-account share is what stops a delete, and it has to be lifted before the wipe can remove it."

STS assume-role denies​

Blocks: everything, because every job step begins by assuming AlphaAgentOrgTarget-<suffix> with your external id. Causes: the trust policy changed, the external id differs, the role was deleted, or an SCP in the Organisations account denies sts:AssumeRole.

You see: at Verify, the assume-role check fails with AWS's message and "Confirm the CloudFormation stack from the account template was run in this AWS account, with the External ID copied exactly and RoleNameSuffix left as 'prod' (or matching what you chose)."; the account's status becomes failed. At the wizard's preflight, a single role-assumption failure: "Confirm the account is registered and validated (see the Accounts screen) before running preflight." If a job is queued against an account that is not active: "AWS account <id> is '<status>', not one of ['active']. Its provisioning role may not be assumable yet; finish validating the account first." Inside a job, an AccessDenied from STS gives access_denied_missing_action with missing_iam_action: sts:AssumeRole; an expired or invalid session gives credentials_unusable: "The target account's provisioning role could not be assumed, or the session it returned was rejected." with "Check that the provisioning role in your account still exists, that its trust policy still allows AlphaAgent's provisioning identity to assume it, and that the external ID is unchanged. Then resume the job."

Delete-side denies generally​

Blocks: the sweep of retained resources before a failed stack is re-created, stack deletion at teardown, Clear account's per-resource deletes, and its AWS WAF sweep.

You see: husk_retained_sweep_incomplete (resumable, "(or grant the missing permission the message names)"); for a deny during a stack delete, stack_resource_access_denied with the refused resources in failed_resources (Resume continues the delete from them); for a resource held by something outside the stack, stack_delete_failed: "A CloudFormation stack did not delete cleanly." with "The stack is DELETE_FAILED: CloudFormation removed everything it could and left the resources named in the technical details (failed_resources) in place - nothing was created, and nothing already deleted comes back. Each one is still held by something outside the stack; once that is released, resume the job and the delete continues from where it stopped."; Clear account's "denied (missing <action> - re-run the account role template)" log lines followed by wipe_leftovers. For AWS WAF specifically, any deny is reported as the opt-in: "the account role was registered with EnableApiWafPermissions=false, so this wipe cannot list or delete AWS WAF resources. Update the account's alphaagent-org-target-<account> stack with EnableApiWafPermissions=true ('Allow the Organisation to manage a WAF allow-list for the Studio public API?'), then run Clear account again."

The codes a guardrail can produce, with the card's copy ("What to check" is the card's remediation, abridged where the full text is quoted in the pattern above). Since Organisations 1.0.5 a deny inside a CloudFormation stack is stack_resource_access_denied; the three stack codes further down cover stack failures whose reason is not a deny. "Who acts" is the card's attribution: "Customer account" means something in your account has to change; "AlphaAgent" means the card's primary line is "AlphaAgent is investigating this — no action is needed from you." and the remediation sits in the technical details, unless the failure is resumable, in which case the next step is shown as "Resumable" followed by the action.

CodeWho actsMessageWhat to check
access_denied_missing_actionCustomer accountAWS denied an operation because the provisioning role is missing an IAM action.Grant the missing action, named in this failure's details, to the provisioning role in the target account, then resume the job.
stack_resource_access_deniedCustomer accountYour AWS account denied a permission a CloudFormation stack needed.The details name the resource (failed_resource, failed_resource_type), the denied action (denied_action), the role that was refused (denied_principal) and CloudFormation's own words (failed_reason). Allow that action for that role in this AWS account - in its IAM policy, and in any Service Control Policy applied to the account at your AWS Organization level - then resume the job.
credentials_unusableCustomer accountThe target account's provisioning role could not be assumed, or the session it returned was rejected.Check that the provisioning role in your account still exists, that its trust policy still allows AlphaAgent's provisioning identity to assume it, and that the external ID is unchanged. Then resume the job.
quota_exceededCustomer accountAn AWS service quota in the target account was exceeded.Request an increase for the quota named in this failure's details, wait for AWS to grant it, then resume the job. The commonest is the Fargate concurrent vCPU quota (L-3032A538).
restore_point_unwritableCustomer accountThe restore point could not be written to the deploy bucket, so the update was not started.Grant the provisioning role s3:PutObject on the deploy bucket's restore-points/ prefix, then run the job again.
registration_stack_stale_for_cloudfrontCustomer accountThis account's registration stack predates the CloudFront permissions (TemplateVersion 2.3.0) or does not allow us-east-1, so the front door cannot be created there.Update the account's registration stack from the current template, keeping us-east-1 in AllowedRegions, then resume the job.
stack_create_failedAlphaAgent (resumable)A CloudFormation stack did not reach CREATE_COMPLETE.For a reason that is not a deny. The failed stack's events name the resource and the reason; on resume the retained resources, then the stack, are deleted and re-created.
stack_update_rollback_completeAlphaAgent (resumable)A CloudFormation stack update failed and rolled back cleanly.For a reason that is not a deny. The stack is still live and serving; read the failed stack's events, fix the cause, then resume the job.
stack_delete_failedAlphaAgent (resumable)A CloudFormation stack did not delete cleanly.For a reason that is not a deny. The resources named in failed_resources are held by something outside the stack; release it, then resume the job.
husk_retained_sweep_incompleteAlphaAgent (resumable)A failed stack CREATE left resources its rollback retained, and not all of them could be deleted before re-creating.Delete the named resources in the AWS console (or grant the missing permission), then resume the job.
wipe_leftoversCustomer accountClear account finished its passes, but the final check still found AlphaAgent resources in the account.Redo the clearing pass; if a resource survives a second pass, an Object Lock, a resource-level SCP or a cross-account share is holding it.
edge_outputs_missingAlphaAgent (resumable)The stateless stack needs the edge stack's function version and the minted us-east-1 certificate, and neither was available.Resume the job: the edge step runs again and hands its outputs on.
settle_service_rolled_backCustomer accountOne or more ECS services failed their new deployment's health checks and were automatically rolled back to their previous task definition.The previous revision is already back and serving - nothing is down. Check the named service's task logs for why the new revision failed its health check, fix the underlying cause, then resume the job to try the update again.
settle_service_never_stabilisedCustomer accountOne or more ECS services never reached steady state before the settle timeout.Read the named service's task logs. Fix the cause, then resume the job.
image_push_failedAlphaAgent (resumable)A registry push or copy failed for a reason that is not a permission or an immutable tag.Read the registry output in this failure's details; the commonest cause is network reachability from the private subnets to the ECR endpoints and the S3 gateway endpoint.
saml_auto_provisioning_failedAlphaAgent (resumable)AlphaAgent's automatic SSO provisioning against your connected identity provider did not complete.Click Resume; if it keeps failing after several resumes, contact support with the job ID and the error named in the details (which is where a Cognito deny appears).
redis_busyAWSElastiCache would not delete the deployment's Redis replication group: it, or one of its member cache clusters, is in a state (snapshotting, modifying or creating) that refuses deletion.Wait until every member of the replication group reads 'available' in the ElastiCache console, then Resume the job.

How to resume after fixing the policy​

On the job's failure card:

  • Resume re-runs from the failed step. The dialog "Confirm retry" is headed "Retry from <step>", lists "These N completed steps will be skipped:" (or "No steps have completed yet, so nothing will be skipped"), and starts with "Start the retry". Afterwards the card reads "Retry started" and updates as the job progresses.
  • Redo re-runs every step: "A redo re-runs every step, including the N that already completed. Each step is idempotent, so this is safe, just slower."
  • A failure that cannot be retried says so: "This failure cannot be retried from where it stopped. Redo re-runs every step from the start (each is idempotent, so it is safe) and is offered below." or "This failure cannot be retried. Retrying could not succeed, so it is not offered."
  • Abandon marks the job cancelled without touching your account: "Nothing in your AWS account is changed by abandoning - only this job is marked cancelled and stops holding the deployment, so another job (a reconfigure, upgrade, uninstall or re-install) can start against it."
  • Wipe and restart tears the deployment's stacks down and reinstalls from scratch (type the deployment name to confirm).
  • A refused retry shows "The retry was refused (<code>): <detail>".

On a parked preflight, fix the account and then either "Confirm and resume" with the override, or "Do not override" and start a new job, which runs preflight again.

In the Organisations installer, grant the action or delete the stuck stack, then run the same command again; completed steps are kept.

Limits​

  • No code path in the product names an SCP, a permissions boundary or an opt-in; the copy always names the denied action or the failed CloudFormation resource. Map it back to your policy yourself.
  • Nothing checks Bedrock model access, Fargate vCPU quota or Bedrock token quota before the stacks are created.
  • Config non-compliance that is not enforced as a deny is not observed by any job.
  • Verify tests only sts:AssumeRole and the us-east-1 stack listing; a boundary or SCP on the target role passes Verify and fails later, per step.

If something goes wrong​

Copy the diagnostic ID from the failure card and share it with your AlphaAgent contact together with the failure code. If a job is stuck holding the deployment, Abandon frees it. If an account is left with resources you cannot delete, Clear account and its wipe_leftovers report list what remains; see Deleting a deployment and clearing an account.