Jobs, progress and failures
For administrators who install, update or delete Studio deployments, and anyone asked why a job stopped. Viewers read every job; retrying, resuming or abandoning needs the Owner or Operator tier or the DeploymentOperator Role.
Before you start: open Fleet, click a deployment and choose Jobs in the View selector, or follow the Fleet's Last job link. A job's address carries ?job=<job id> and #step=<step id>, so it can be shared.
Filter the job list
The list holds every job ever run against the deployment. Filter by status (All statuses, Queued, Starting, Running, Waiting for approval, Waiting for input, Succeeded, Failed, Rolled back, Cancelled) and by type (All types, Install, Upgrade, Reconfigure, Rollback, Uninstall, Account wipe). Show system cancellations (N hidden) reveals jobs cancelled automatically because another job held the deployment's slot ("Skipped").
Columns are Job, Type, Status, Started and Failure (the headline; the code is in the hover). Clicking a row selects that job below. Longer histories page; the list refreshes itself.
Read the running job
For the running job the page shows, in order: a Waiting for you card if a step is parked, a failure card if it stopped, the cancel bar during a teardown's grace window (Deleting a deployment), a status line, "N of M steps complete", the step timeline and, after an uninstall, What was retained.
The status line reads Queued, Starting, Running, Succeeded, Failed, Rolled back or Cancelled, or "N action(s) needed"; a re-run adds " · Attempt N". If the live connection drops, the view may be up to 30 seconds behind while it reconnects. Any other job opens read-only with the same cards and timeline.
Provisioning steps lists each step with its status (Pending, Running, Waiting for you, Succeeded, Failed, Skipped) and elapsed time. Click a step name for its attempt history, problems found, a parked prompt or an error code, and the log viewer with Copy and Download full log (<job id>-<step id>.log).
Answer a parked step
A parked step is not a failure: the job waits for something only you can supply. The card is headed Waiting for you plus one of:
- Upload your SAML metadata: the values your identity administrator needs (Identifier (Entity ID), Reply URL (ACS URL), the required claims) and a field for SAML metadata (XML) or SAML metadata URL; Upload and resume.
- Proceed despite failed readiness checks: acknowledge each failed preflight check, then Approve and resume (with Proceed anyway ticked) or Do not override.
- Confirm a destructive action: a Confirmation field and Confirm and resume. No job type parks with this prompt today.
Every form records Your name or email. The card shows "Waiting for <duration>"; an expired park means start a new job, which parks again at the same point. Link to this step copies the shareable address. A rejection is final: the job fails with approval_rejected and cannot be resumed. Abandon this job marks it cancelled and frees the deployment without changing anything in your AWS account.
Read a failure card
The card is headed <step> failed (or This job failed), then states who acts:
| Who acts | Meaning |
|---|---|
| Customer account | Something in the target AWS account has to change first. |
| AlphaAgent | AlphaAgent is investigating; no action is needed from you. |
| AWS | A transient or capacity condition on the AWS side. |
| Unknown | The console could not tell; share the diagnostic ID with your AlphaAgent contact. |
Next comes What to check: (the remediation) or Resumable: with the next step, often written at the moment of failure and naming the exact resource, action or record. Diagnostic ID has Copy; Technical details opens the raw details (the failed resource, the denied action, or a DNS validation record as a NAME CNAME VALUE line). A card that cannot resume offers Redo; one that could never succeed offers no retry.
Every failure code, with its message and remediation, is on Failure codes.
Resume, Redo, Wipe or Abandon
- Resume and Redo open Retry from <step>. A redo re-runs every step, including those completed (each is idempotent: safe, just slower); a resume skips them. Choose Start the retry.
- Wipe (Wipe and restart) tears down the deployment's CloudFormation stacks irreversibly and reinstalls from scratch; type the deployment name, then Start teardown.
- Abandon marks the job cancelled without touching your account, so another job can start. Offered on every failed job.
Where a failure comes from
A permission refused on a direct call fails as access_denied_missing_action; one refused inside a stack fails as stack_resource_access_denied and resumes once the named role may perform the action; a failed readiness check parks the job instead (Failure codes).
What you should see
- A running job: the status line and "N of M steps complete" updating live, each step turning Succeeded in turn.
- A parked job: Waiting for you at the top with the form for the answer; a failed job: the failure card naming who acts and offering Resume, Redo or Abandon.
Notes
- One job runs against a deployment at a time; another is refused with "Another job is already in flight for this deployment."
- A retry is Resume or Redo; no step can be skipped. Wipe needs the permission to uninstall the deployment.
- Preflight failures park the job with the check's own remediation; they are not failure codes.
- A refused resume ("The server refused the resume (HTTP <status>).") means reload: someone else answered the park, or the job moved on.
- A card that blames AlphaAgent needs no change in your account; share the Diagnostic ID.