Skip to main content

Knowledge graphs in depth

A knowledge graph (AMPG) turns documents your users add into a graph that agents draw on when they answer. This page follows a document from upload to retrieval with one question in mind: where is your data at each stage, and what runs on it. It describes components, stores and boundaries. How the graph is assembled from the text, and how retrieval ranks what it returns, are proprietary and are not documented here or anywhere public.

Who this is for​

Security engineers and architects assessing where knowledge graph data lives. The user-facing pages start at Knowledge graphs.

Before you start​

Vocabulary: a knowledge graph is the container your users create; a version is one immutable build of it from a set of documents; an agent pins a version. Versions are peers: there is no pointer that names one of them as the live one. In Studio a version is Building, Ready, Failed or Cancelled (a version being removed shows Deleting), and an agent can pin only a Ready version; the picker defaults to the highest Ready version.

Stage by stage​

StageWhere the data isWhose account
Documents are addedThe kb-artifacts bucket in the deployment; documents are uploaded straight to it with a presigned requestYours
A build is queuedAn EventBridge rule on object creation puts a message on the kb-processing queue (with a dead-letter queue)Yours
The build runsThe knowledge-base ECS task (Heavy tier; 50 GiB of task storage) reads the documents and works through the steps Studio shows: Upload, Prepare document, Parse document, Build graph, Import graph, Train embeddings. Intermediate state lands in the kb-processing and kb-dgl-training buckets and in the knowledge graph tablesYours
Model calls during the buildAmazon Bedrock through the deployment's VPC endpoint, pinned to your zone: the deployment's Sonnet model for the build and cohere.embed-v4 for embeddings. Metered as kg_build, kg_edge_discovery and kg_embeddings, counts onlyYours
The finished versionYour self-hosted Neo4j (a single Fargate task) with its data on the encrypted EFS file system; every version is its own subgraph keyed by the graph and the version numberYours
Retrieval at answer timeThe knowledge-base-mcp service reads the version the agent pins from Neo4j and hands the agent bounded context; metered as kg_retrievalYours

Nothing about a document, a graph or a query is sent to Prometheus Research Labs. The Console sees the four kg_* metering kinds as token counts.

Versions​

  • One route creates every version. A first build from uploaded files, a rebuild from the browser, an AMPG update step in a workflow (Rewrite or Append) and an import from a Library all create version N+1 through the same operation. A version request names its source: uploaded files, keys already in the workspace, an import from a Library transfer, or a retry of a failed version.
  • The base of an update is the highest Ready version. A workflow step that updates a graph starts from that version's documents. Rewrite produces one consolidated document that becomes the whole next version; Append adds a document beside the base. There is no union of every earlier version.
  • Bounds. At most 50 versions per graph, and one version building at a time per graph: a second request while one is building answers 409. A document is a PDF of up to 32 pages and 50 MiB.
  • What a version records. Status and failure reason; created, started and finished times; who created it; its source; the job ids; graph statistics (nodes, edges, relation types, documents, pages) stamped when it becomes Ready; the documents it was built from; and, for a version a workflow produced, provenance naming the run and conversation.
  • Deleting a version removes its subgraph. Agents pinned to it are found first (the service asks agent-management which agents pin the version), so a version in use is not silently removed from under them.

How agents use a graph​

An agent pins one graph at one version in its active configuration. A change to the graph does not change what the agent reads until the agent is re-pinned. Re-pinning is explicit: from the agent's configuration, from a workflow step that has Adopt all agents set, or from the graph page, which lists the agents affected. Retrieval is bounded to the pinned version and stays inside the deployment; the answer cites the passages it used.

Sharing a graph between deployments​

A knowledge graph shared to a Library is catalogued as a pointer to one version with node and edge counts. On pull, the Organisations engine runs the kg-transfer task in the source deployment to export that version's subgraph to the source account's staging bucket, copies it to the target account's staging bucket, and runs the task in the target deployment to import it as a new version there. Neo4j credentials for each side come from that deployment's own configuration secret. Staged objects expire after seven days; see Secrets and connector credentials.

Steps: verify the stores in your account​

  1. In S3, open alphaagent-kb-artifacts-<account>-<region>: your users' documents, one prefix per graph.
  2. In SQS, open alphaagent-kb-processing; a build in progress shows messages in flight.
  3. In ECS, open the neo4j service: one task, with the EFS volume mounted at /data.
  4. In EFS, open the Neo4j file system: encrypted, backups enabled, size growing with versions.
  5. In the deployment's usage view in the Organisations console, group by kind: kg_build, kg_edge_discovery, kg_embeddings during builds and kg_retrieval during use.

What you should see​

  • Every knowledge graph object in your own buckets, tables and file system, in the deployment region.
  • A Ready version with graph statistics, and agents pinned to a named version rather than to "the latest".

Limits​

  • 50 versions per graph; one build in flight per graph; PDF documents up to 32 pages and 50 MiB each.
  • In a standard deployment each user may create up to 256 knowledge graphs; in a governed deployment graphs are shared and the cap stands down.
  • Neo4j is a single task on Fargate. Its durability is the EFS file system and its backups, not a Neo4j cluster.

If something goes wrong​

  • A version shows Failed: open it for the failure reason; the documents are untouched and a retry creates a new version.
  • A build stays in Building: the knowledge-base task's log group in CloudWatch shows the step; a message in the dead-letter queue means the task could not process it.
  • An agent cannot pin a version: only Ready versions can be pinned; wait for the build or pick an earlier Ready version.