PII redaction posture
PII redaction runs inside your Studio deployment, on the files that enter a workflow run and the files that leave it. This page states where each part of it runs, what it guarantees, and what it deliberately does not claim, so that a privacy or compliance review can decide whether the design fits its own controls. The user-facing description is PII redaction.
This page describes system behaviour. It is not legal or regulatory advice and does not assert that any configuration satisfies a particular law or policy.
Who this is for
Security, privacy and compliance reviewers.
Before you start
Redaction is a property of workflow runs and uploads, governed by a policy with a mode, categories, stages and a semantic-pass setting. A deployment default set by your Organisations administrator is the floor; a workflow version may tighten it; a per-run override may tighten it further.
Where redaction runs
| Part | Where | What leaves your account |
|---|---|---|
| The pattern pass (e-mail, phone, IBAN, card numbers, national identifiers, IP addresses, postcodes, cued dates of birth, your custom patterns) | On CPU inside the agent-runtime service in your account | Nothing |
| The semantic pass (names, dates of birth, medical and personal financial facts in context) | Amazon Bedrock in your deployment's zone, using the deployment's configured Opus model, one isolated model invocation per chunk that the pattern pass hit or that carries a name or date cue | Nothing leaves AWS; the call is to your own account's Bedrock endpoint through the VPC endpoint |
| Uploads | Pattern pass only, at upload | Nothing |
| The redaction key | alphaagent-pii-hmac-key in your Secrets Manager, read by agent-runtime | Never returned by any API, written to a log or manifest, or sent over the licence channel |
| Metering | One pii_redaction usage event per semantic call: model id, token counts and identifiers | Counts only, over the licence channel |
Bedrock's optional model invocation logging is an account-level setting you control. If you enable it, AWS writes prompts and completions, including chunks sent to the semantic pass, to your own CloudWatch Logs or S3.
What it protects
Data enters a run through five paths and leaves through four. Each entry path is a stage the policy switches on or off:
| Entry path | Stage | Where the redactor runs |
|---|---|---|
| Run inputs copied from your S3 prefix at run start | inputs | In the copy: the redacted bytes land in the run; the raw object does not |
| Files a user uploads to a workspace or attaches in chat | uploads | On the upload, with the deployment default |
| Rows a step reads from a connector | connector_results | Before the rows reach the agent |
| What agents write under the run's outputs | outputs | After every step, in place, in the versioned bucket; earlier raw versions are deleted |
| Exit path | Stage | Guarantee |
|---|---|---|
Deliverables served in Studio and through GET /runs/{run_id}/outputs | outputs | Served only after the final sweep has completed; until then the API answers 409 redaction_pending |
result.json, returned inline on the run | result_json | Redacted before it is validated and returned |
| The run's closing summary | outputs | Redacted before it is written |
Three modes decide what a detected value becomes:
| Mode | Result | Property |
|---|---|---|
| Mask | A placeholder with a per-run index, [EMAIL_1], [PERSON_NAME_3] | The same value gets the same placeholder throughout one run; indexes are unrelated between runs |
| Hash | CATEGORY: and the first 16 hex characters of an HMAC-SHA256 of the normalised value with the deployment key | Deterministic across runs and files in one deployment, so your systems can join outputs to their own records |
| Drop rows | The row is removed from CSV, TSV, JSON arrays, NDJSON and Parquet; free text carries hash tokens | The record does not leave at all |
The platform holds no reverse map for any mode. A keyed mode with no usable key fails the run before any step or model call; it never falls back to Mask and never writes a raw object. The floor is Off for a standard deployment and Mask for a governed one; Drop rows is a per-workflow choice, not a floor. A per-run override is accepted only if every field tightens: a stronger mode, a stage switched on, more categories, a lower confidence threshold, a wider semantic pass, a lower ceiling. Anything looser is refused with 400 policy_not_tightening.
Every run in which redaction applied writes a manifest, inputs/PII_REDACTION.json, returned by GET /api/v1/runs/{run_id}/pii-redaction (a run whose effective policy is Off has none, and the route answers 404 pii_redaction_not_applied): mode, model id, semantic-pass setting, minimum confidence, whether the placeholder registry was keyed, and per stage and object the category counts, values found, rows dropped, model calls, whether sampling or the ceiling applied, and skip reasons. It never contains a value, an offset or a length beyond the category.
What it does not claim
- The inside of a run. When
inputsorconnector_resultsis off (a screening workflow must see the real applicant), the run's agents read the real values, and those values are visible for the life of the run to anyone the deployment lets see the run: the step board, approval cards in the Inbox,GET /runs/{run_id}/eventsfor a key withruns:read, and working files that are not deliverables. Whether that is acceptable is a policy decision for the workflow author and your administrator. - Writes the agent's own code makes. Code that talks to a connector or S3 directly with its injected credentials is redacted only if the workflow routes the data through a Redact PII step or the
redact_piitool first. The platform cannot see inside the sandbox's own network calls. - Formats it skips. A scanned PDF with no text layer, images, Office documents, archives, files with no extension and an unknown type, files over 500 MB, and undecodable text pass through unchanged and are named in the manifest with a reason. A PDF with a text layer is rewritten as text; layout is lost. There is no OCR.
- Confidence and sampling. Semantic findings below the policy's minimum confidence (0.80 by default) are counted, not applied. A failed model call is counted and that chunk is not semantically redacted. For Parquet, pattern detectors run on every row of every string column; the semantic pass reads a sample (50,000 rows by default, a ceiling) and a column judged personal is redacted for the whole column.
- Residual misses. A value the semantic pass does not recognise can survive. A back-substitution pass replaces every value the run has already found, and its name tokens of three or more characters, in every later object without a model call, and records the counts as
back_substituted. The failure mode for deliverables is withholding until the sweep has run, not serving unredacted content. - Re-identification by your own staff. Someone holding both the key and a candidate value can confirm whether it produced a token. Who can read the key is your IAM decision; reads are
GetSecretValueevents in your CloudTrail.
Key rotation and erasure
The hash key is minted once when the deployment is first seeded and preserved by every upgrade; the platform never rotates it on its own. To rotate, replace the value of alphaagent-pii-hmac-key with a new 32-byte key and restart the agent-runtime service. Tokens produced afterwards differ from every earlier token; files already redacted keep their old tokens; joins across the boundary no longer match. Mask placeholders are unaffected.
Erasing a run started by an API key, including its manifest and placeholder registry, is DELETE /api/v1/runs/{run_id}/data with the same key; see Data at rest and in transit.
Steps: verify the posture in your account
- In Secrets Manager, open
alphaagent-pii-hmac-key. Its resource policy and your IAM policies decide who can read it; the Studio task role can. - In the Organisations console, open the deployment's configuration and read PII redaction default.
- Run a workflow with every stage on against a synthetic fixture, then read
GET /api/v1/runs/{run_id}/pii-redactionand confirm the counts match the fixture and no value appears in the manifest. - In CloudWatch Logs, search
/alphaagent/services/agent-runtimefor a fixture value. There should be none.
What you should see
- The manifest reports categories and counts,
registry_keyed: true, andcomplete: truebefore any deliverable is served. - A per-run override that loosens the policy is refused with
400 policy_not_tighteningnaming the field.
Limits
- Model calls per run for the semantic pass are capped (200 by default, lowerable per workflow); past the cap the manifest says
capped. - Redaction is available on workflow runs and uploads. Chat replies are not redacted.
- One key per deployment. There is no reversible or tokenised mode by design.
If something goes wrong
- A keyed run fails at launch with
pii_key_unavailable: the key secret is empty or unreadable. Restore it and start the run again; no raw object was written. - Outputs answer
409 redaction_pendingfor longer than expected: the final sweep is still running; the response carriesRetry-After. - A manifest reports
skippedfor a file you expected redacted: the file is a scanned image, a binary format or over the size limit.