Skip to main content

PII redaction posture

PII redaction runs inside your Studio deployment, on the files that enter a workflow run and the files that leave it. This page states where each part of it runs, what it guarantees, and what it deliberately does not claim, so that a privacy or compliance review can decide whether the design fits its own controls. The user-facing description is PII redaction.

This page describes system behaviour. It is not legal or regulatory advice and does not assert that any configuration satisfies a particular law or policy.

Who this is for​

Security, privacy and compliance reviewers.

Before you start​

Redaction is a property of workflow runs and uploads, governed by a policy with a mode, categories, stages and a semantic-pass setting. A deployment default set by your Organisations administrator is the floor; a workflow version may tighten it; a per-run override may tighten it further.

Where redaction runs​

PartWhereWhat leaves your account
The pattern pass (e-mail, phone, IBAN, card numbers, national identifiers, IP addresses, postcodes, cued dates of birth, your custom patterns)On CPU inside the agent-runtime service in your accountNothing
The semantic pass (names, dates of birth, medical and personal financial facts in context)Amazon Bedrock in your deployment's zone, using the deployment's configured Opus model, one isolated model invocation per chunk that the pattern pass hit or that carries a name or date cueNothing leaves AWS; the call is to your own account's Bedrock endpoint through the VPC endpoint
UploadsPattern pass only, at uploadNothing
The redaction keyalphaagent-pii-hmac-key in your Secrets Manager, read by agent-runtimeNever returned by any API, written to a log or manifest, or sent over the licence channel
MeteringOne pii_redaction usage event per semantic call: model id, token counts and identifiersCounts only, over the licence channel

Bedrock's optional model invocation logging is an account-level setting you control. If you enable it, AWS writes prompts and completions, including chunks sent to the semantic pass, to your own CloudWatch Logs or S3.

What it protects​

Data enters a run through five paths and leaves through four. Each entry path is a stage the policy switches on or off:

Entry pathStageWhere the redactor runs
Run inputs copied from your S3 prefix at run startinputsIn the copy: the redacted bytes land in the run; the raw object does not
Files a user uploads to a workspace or attaches in chatuploadsOn the upload, with the deployment default
Rows a step reads from a connectorconnector_resultsBefore the rows reach the agent
What agents write under the run's outputsoutputsAfter every step, in place, in the versioned bucket; earlier raw versions are deleted
Exit pathStageGuarantee
Deliverables served in Studio and through GET /runs/{run_id}/outputsoutputsServed only after the final sweep has completed; until then the API answers 409 redaction_pending
result.json, returned inline on the runresult_jsonRedacted before it is validated and returned
The run's closing summaryoutputsRedacted before it is written

Three modes decide what a detected value becomes:

ModeResultProperty
MaskA placeholder with a per-run index, [EMAIL_1], [PERSON_NAME_3]The same value gets the same placeholder throughout one run; indexes are unrelated between runs
HashCATEGORY: and the first 16 hex characters of an HMAC-SHA256 of the normalised value with the deployment keyDeterministic across runs and files in one deployment, so your systems can join outputs to their own records
Drop rowsThe row is removed from CSV, TSV, JSON arrays, NDJSON and Parquet; free text carries hash tokensThe record does not leave at all

The platform holds no reverse map for any mode. A keyed mode with no usable key fails the run before any step or model call; it never falls back to Mask and never writes a raw object. The floor is Off for a standard deployment and Mask for a governed one; Drop rows is a per-workflow choice, not a floor. A per-run override is accepted only if every field tightens: a stronger mode, a stage switched on, more categories, a lower confidence threshold, a wider semantic pass, a lower ceiling. Anything looser is refused with 400 policy_not_tightening.

Every run in which redaction applied writes a manifest, inputs/PII_REDACTION.json, returned by GET /api/v1/runs/{run_id}/pii-redaction (a run whose effective policy is Off has none, and the route answers 404 pii_redaction_not_applied): mode, model id, semantic-pass setting, minimum confidence, whether the placeholder registry was keyed, and per stage and object the category counts, values found, rows dropped, model calls, whether sampling or the ceiling applied, and skip reasons. It never contains a value, an offset or a length beyond the category.

What it does not claim​

  • The inside of a run. When inputs or connector_results is off (a screening workflow must see the real applicant), the run's agents read the real values, and those values are visible for the life of the run to anyone the deployment lets see the run: the step board, approval cards in the Inbox, GET /runs/{run_id}/events for a key with runs:read, and working files that are not deliverables. Whether that is acceptable is a policy decision for the workflow author and your administrator.
  • Writes the agent's own code makes. Code that talks to a connector or S3 directly with its injected credentials is redacted only if the workflow routes the data through a Redact PII step or the redact_pii tool first. The platform cannot see inside the sandbox's own network calls.
  • Formats it skips. A scanned PDF with no text layer, images, Office documents, archives, files with no extension and an unknown type, files over 500 MB, and undecodable text pass through unchanged and are named in the manifest with a reason. A PDF with a text layer is rewritten as text; layout is lost. There is no OCR.
  • Confidence and sampling. Semantic findings below the policy's minimum confidence (0.80 by default) are counted, not applied. A failed model call is counted and that chunk is not semantically redacted. For Parquet, pattern detectors run on every row of every string column; the semantic pass reads a sample (50,000 rows by default, a ceiling) and a column judged personal is redacted for the whole column.
  • Residual misses. A value the semantic pass does not recognise can survive. A back-substitution pass replaces every value the run has already found, and its name tokens of three or more characters, in every later object without a model call, and records the counts as back_substituted. The failure mode for deliverables is withholding until the sweep has run, not serving unredacted content.
  • Re-identification by your own staff. Someone holding both the key and a candidate value can confirm whether it produced a token. Who can read the key is your IAM decision; reads are GetSecretValue events in your CloudTrail.

Key rotation and erasure​

The hash key is minted once when the deployment is first seeded and preserved by every upgrade; the platform never rotates it on its own. To rotate, replace the value of alphaagent-pii-hmac-key with a new 32-byte key and restart the agent-runtime service. Tokens produced afterwards differ from every earlier token; files already redacted keep their old tokens; joins across the boundary no longer match. Mask placeholders are unaffected.

Erasing a run started by an API key, including its manifest and placeholder registry, is DELETE /api/v1/runs/{run_id}/data with the same key; see Data at rest and in transit.

Steps: verify the posture in your account​

  1. In Secrets Manager, open alphaagent-pii-hmac-key. Its resource policy and your IAM policies decide who can read it; the Studio task role can.
  2. In the Organisations console, open the deployment's configuration and read PII redaction default.
  3. Run a workflow with every stage on against a synthetic fixture, then read GET /api/v1/runs/{run_id}/pii-redaction and confirm the counts match the fixture and no value appears in the manifest.
  4. In CloudWatch Logs, search /alphaagent/services/agent-runtime for a fixture value. There should be none.

What you should see​

  • The manifest reports categories and counts, registry_keyed: true, and complete: true before any deliverable is served.
  • A per-run override that loosens the policy is refused with 400 policy_not_tightening naming the field.

Limits​

  • Model calls per run for the semantic pass are capped (200 by default, lowerable per workflow); past the cap the manifest says capped.
  • Redaction is available on workflow runs and uploads. Chat replies are not redacted.
  • One key per deployment. There is no reversible or tokenised mode by design.

If something goes wrong​

  • A keyed run fails at launch with pii_key_unavailable: the key secret is empty or unreadable. Restore it and start the run again; no raw object was written.
  • Outputs answer 409 redaction_pending for longer than expected: the final sweep is still running; the response carries Retry-After.
  • A manifest reports skipped for a file you expected redacted: the file is a scanned image, a binary format or over the size limit.