Skip to main content

Updates and rollback

A Studio release reaches your deployments through your Organisations account: Organisations mirrors the release, and an Update Manager rule or a manual upgrade applies it deployment by deployment. Every upgrade captures a restore point first and rolls the services in place. This page describes the rule model, the roll itself, how rollback works, and how Organisations updates its own software.

Who this is for​

Administrators who schedule updates, and security engineers who want to know what changes during one and how it is undone.

Before you start​

Three objects in the Organisations console: a release (a Studio version the Console offers on your channel), a Fleet Configuration Template (a versioned payload naming a release and the deployment settings you standardise on), and an Update Manager rule (a schedule that applies one template version to a set of deployments). The screens are described in Fleet configurations and Update Manager.

How a release reaches your account​

  1. Organisations asks the Console for the release catalogue and, when a release is first needed, downloads its bundle through a signed, single-use link into the bundle-cache bucket in your Organisations account. The release cache records the state: discovered, downloading, cached, images loaded, deployable.
  2. The bundle's container images are pushed into your Organisations ECR repositories, immutable by tag.
  3. For each deployment, the images are copied by digest into that account's own repositories, so the bytes a deployment runs are the bytes the release shipped. No cross-account repository policy is needed on either side.

Update Manager rules​

A rule carries a name, a template and the template version it pins, a scope (an explicit list of deployment ids, or a filter by AWS account and region), a schedule, an enabled flag, and an optional maintenance window as a UTC start hour and end hour.

Rules of the model:

  • A rule keeps the template version it was created with. Editing the template bumps the template's version but does not move any rule; Adopt vN on the rule is the explicit step that does.
  • A template may name only a release your Organisation offers: the Console catalogue for your channel, plus releases already mirrored and deployable in your account, plus the release every existing deployment is running. Anything else is refused when the template is saved (release_not_offered).
  • The engine evaluates due rules on an internal tick. For each deployment in scope it compares the running release and configuration with the template; a drifted deployment gets an upgrade or reconfigure job.
  • A rule due outside its maintenance window is deferred to the next window start, never dropped, and the deferral is recorded.
  • Every firing is recorded in the rule's History with one of the outcomes Fired, Nothing to do, Drift found, not applied, Deferred (outside window), Template missing or Template edited, not adopted.

The roll​

What happens to a running deployment:

PhaseWhat users seeWhat changes
preflight, restore-pointNothingThe registration stack, quotas and name collisions are checked; the current task definition revisions and image tag are written to the deployment's deploy bucket
maintenance-onThe update screen replaces the Studio web appA maintenance page and flag are written to the frontend bucket
core to taskdefThe update screenStacks are updated in order; new images are copied by digest and verified; configuration secrets are re-seeded; model ids are re-converged to the release's defaults for the zone
settleThe update screenEvery service is force-redeployed in place. Eight of the eleven services are configured to keep 100 percent healthy capacity and run up to 200 percent during the roll, so a new task is healthy before an old one stops. Sandbox functions are re-pointed to the new environment image when its digest moved
maintenance-off, doctor, record-releaseStudio returnsThe normal web app is restored, configuration is verified, and the deployment record carries the new release

The elapsed time depends on the sizing profile and on how many images changed; the job in the Organisations console shows it.

Stateful resources are never replaced by an upgrade: tables, buckets, secrets, the file system, the cache and the user pool keep their data. Run and conversation data is untouched.

Rollback​

  • Restore point first. Before anything changes, the upgrade writes the running task definition revision of every service and the deployed image tag to restore-points/ in the deployment's deploy bucket. If that write fails, the upgrade stops there (restore_point_unwritable), before any stack or secret is touched.
  • Automatic on failure. When an upgrade or reconfigure fails at a point that cannot be resumed, the engine runs a rollback step: it points every service back at the recorded task definitions and forces a redeploy. The step reports per-service outcomes and does not claim success unless services actually moved; a failure is recorded as rollback_failed, a missing restore point as no_restore_point.
  • Manual. An administrator can trigger the same rollback from the deployment (POST /deployments/{deployment_id}/rollback); it appears as its own job with its own step record.
  • Back to an earlier release. A release you have mirrored stays offered, so a template can name it and a rule can adopt it; the same roll then applies the earlier release.

How Organisations upgrades itself​

Organisations is upgraded from your workstation with orgctl.py install from the new bundle, in the same account and region. The recovery gate detects the existing install from the six stacks, reconstructs its state, and, because the version differs, re-runs the forced steps: org-auth, foundation, push-image, seed-config, spa, app, settle and doctor. It does not re-ask for the SAML metadata, the Console credential or the break-glass owner. A change set that would replace a stateful resource stops for confirmation. The step-by-step page is Upgrading Organisations.

Steps: verify an update from your account​

  1. In the Organisations console, open the deployment's job list and the upgrade job. Read the step list in the order above and the outcome of restore-point.
  2. In S3, open alphaagent-deploy-<account>-<region>/restore-points/. One object per upgrade.
  3. In ECS, open the cluster: every service's task definition revision changed at settle, and the previous revisions still exist.
  4. In ECR, compare the image digest referenced by the new task definition with the digest in your Organisations repository for that release.

What you should see​

  • A restore point object written before the core step ran.
  • Every service on a new task definition revision, with the old revisions retained for rollback.
  • The deployment record showing the new release only after record-release.

Limits​

  • One in-flight job per deployment; concurrent jobs across deployments are bounded by the account's Fargate quota.
  • A maintenance window is hour-of-day in UTC only; day-of-week belongs to the schedule.
  • Rollback restores services to their previous task definitions. It does not roll back a stack template change or a configuration secret; a failed upgrade that changed those is resumed or re-run.

If something goes wrong​

  • Job fails with release_not_deployable: the release has not finished mirroring into your Organisations account. Wait for the release cache to show it deployable, then retry.
  • Job fails with restore_point_unwritable: the engine could not write to the deploy bucket. Check the bucket exists and the target role's S3 permissions are intact, then retry; nothing was changed.
  • Job fails with rollback_failed: one or more services did not return to the recorded revision. The step record names them; use ECS in the account to inspect the service events, then re-run the upgrade or the rollback.