Agent identity, scope tracking, and containment infrastructure lag behind deployment velocity, creating blind spots that persist through incident response.
Governance without response: the advisory scope
Organizations implementing AI agent governance programs still face response infrastructure gaps when incidents occur. Agent registries, approval workflows, and authority models address deployment controls — but incident response depends on separate infrastructure that tracks agent actions, resolves identity during events, and enables targeted containment. The failure patterns documented here surface even when agent governance appears operational.
This advisory maps six infrastructure gaps that leave incident response teams without the technical capability to investigate, attribute, or contain agent actions effectively. Each failure pattern corresponds to one of the six response objects that define what information incident response requires:
- Agent identity — the ability to resolve from an event credential back to a specific registered agent with its delegation chain intact
- Authority scope — the exact permission scope granted at dispatch for a specific task, not what the agent is generally permitted
- Action chain — a complete, correlated trace of the agent's decision and execution sequence, from dispatch context through tool invocations to outcome states
- Containment boundary — the ability to stop a specific agent without affecting unrelated systems or other agents
- Rollback path — knowledge of which records were changed, in which systems, and whether those changes are reversible
- Evidence record — structured preservation of incident data in a form that satisfies accountability and compliance review
The failure mode vocabulary for agent incident management and cross-references to related coverage appear at the close of this advisory.
Failure pattern 1: agent identity not resolvable during incident
Agent authentication relies on shared API keys, inherited session tokens, or service account credentials not linked to agent registries. When an incident occurs, the credential record points to a service account or API key pool, not to a specific agent identity. The investigation team knows a credential was used but cannot determine which agent held it at the time of action.
This creates attribution reconstruction failure. Multiple agents share credentials through common service accounts or inherited tokens from delegating principals. Post-incident investigation confirms that agent activity occurred but cannot establish which registered agent initiated the access sequence. The investigation proceeds in parallel with ongoing operations that may involve the same unresolved credential.
Detection point: Investigation teams query identity logs with event timestamps and find service account names or API key identifiers, but no agent-specific identity record. The credential-to-agent mapping either does not exist or requires manual correlation across multiple systems that may not retain the necessary historical data.
What changes the outcome: Agent-specific credentials with direct registry linkage, or credential delegation logs that maintain agent identity through the full authentication chain. Concretely, this means issuing short-lived, down-scoped cryptographic bearer tokens per agent run rather than rotating long-lived shared secrets. OAuth 2.0 Token Exchange (RFC 8693) provides a standardized mechanism for minting delegated tokens that embed the originating human principal alongside the agent run ID in structured claims — for example, an act claim that records the actor chain. SPIFFE/SPIRE workload identity frameworks extend this to infrastructure-level identity, binding a cryptographically verifiable workload identity to each agent process so that any log entry can be traced to a specific registered agent without manual correlation. The response object absent here is agent identity.
Failure pattern 2: action chain cannot be reconstructed
Agent tool invocations, API calls, context retrievals, and workflow triggers log independently across separate systems not designed for correlation. Agent dispatch platforms log invocations, API gateways log calls, context retrieval systems log queries, and downstream applications log state changes — but no unified agent action log captures the sequence from dispatch through completion.
Investigation teams can confirm individual actions occurred but cannot reconstruct the triggering condition or action sequence. Root cause analysis stalls because the prompt context and tool sequence that led to an incident cannot be recovered from distributed logs. The same action chain can recur because the investigation cannot identify what condition triggered the problematic sequence.
Detection point: Incident timelines show individual tool calls and API transactions but gaps between them. Investigation teams manually correlate timestamps across platforms to build sequence reconstructions that remain incomplete due to logging inconsistencies or retention mismatches.
What changes the outcome: Reconstructing the action chain does not require building a bespoke agent action log from scratch. OpenTelemetry (OTel), combined with the emerging OTel semantic conventions for Generative AI, provides a standardized instrumentation path. Span context propagated via traceparent headers across LLM inference calls, tool executions, and downstream HTTP clients produces a correlated, queryable trace across disparate systems without custom correlation logic.
As tool integration standardizes on protocols such as Anthropic's Model Context Protocol (MCP) or emerging Agent-to-Agent (A2A) protocols, tool authorization and action logging should be implemented at the protocol gateway level rather than retrofitted into ad-hoc API logging — placing structured telemetry at the boundary where tool invocations are authorized and dispatched. The response object absent here is action chain.
Failure pattern 3: authority scope not preserved at task execution
Governance programs define what agents are generally permitted, but the specific permission scope granted at dispatch for a given task execution is not recorded separately or retained in a queryable form. When an incident occurs, investigation teams can retrieve what the agent is currently authorized to do, but not what scope it held at the moment a specific past action was taken.
This distinction matters for accountability. An agent's general authorization profile may have changed since the incident occurred. Without a point-in-time authority snapshot linked to the specific task execution, investigation teams cannot confirm whether a contested action was within the agent's authority at the time, or whether scope drift or misconfiguration contributed to the incident.
Detection point: Investigation teams can retrieve the agent's current role bindings and permission policies, but find no record of the scope token or permission set that was active at dispatch for the task under investigation. Governance systems record what agents are permitted; they do not record what permissions were actually granted at execution time for each task.
What changes the outcome: Authority scope should be captured as part of the dispatch record for each task execution — the specific scope token issued, the permissions it contained, and the principal that authorized its issuance. When short-lived delegated tokens are used (as described in Failure Pattern 1), the token claims themselves serve as the authority record if they are logged at issuance and linked to the task execution identifier. The response object absent here is authority scope.
Failure pattern 4: containment disrupts unrelated systems
Agents lack capability maps that separate their specific tool access from shared credentials or service account access. Containing an agent requires revoking credentials that other agents or systems depend on, creating a choice between continued agent access and unplanned system outages.
Incident response teams delay containment while assessing blast radius, or execute broad revocations that take unrelated systems offline. Narrow containment is not available because the agent cannot be isolated from shared authentication infrastructure. The incident continues during blast radius assessment, or containment creates secondary outages that exceed the scope of the original incident.
Detection point: Containment planning reveals that agent credentials are shared across multiple systems or agents. The IR team cannot identify a revocation path that affects only the problematic agent without requiring coordination with teams managing unrelated services.
What changes the outcome: Agent-specific authentication with isolated revocation paths, or dynamic capability mapping that enables surgical containment. Where workload identities are issued per agent run (as in Failure Pattern 1), revocation of a single short-lived token or workload certificate does not affect credentials held by other agents or systems. The response object absent here is containment boundary.
Failure pattern 5: rollback path unknown after multi-system writes
When an agent completes a workflow that modifies state across two or more systems, the resulting change set is not recorded in a form that supports systematic reversal. Each affected system retains its own transaction log, but no unified record links the agent's task execution to the complete set of downstream state changes it produced.
Investigation teams attempting to scope impact or revert changes must manually review each downstream system's transaction history, correlate changes to the agent's action window, and assess reversibility system by system. For agents that write to systems without native rollback support, the investigation team may find that partial reversals are possible but complete state restoration is not.
Detection point: Post-incident scope assessment finds that the agent wrote to multiple systems but no unified change record exists. Teams must reconstruct the change set by querying each affected system independently, and reversibility varies across systems without a coordinated rollback path.
What changes the outcome: Workflow dispatch records should capture which systems an agent is authorized to modify and in what sequence. Where the action chain trace (Failure Pattern 2) is instrumented with OTel spans covering downstream writes, the trace itself serves as a change manifest that can be used to scope rollback operations. The response object absent here is rollback path.
Failure pattern 6: no evidence record before state changes
Incident response workflows lack structured evidence preservation before containment begins. Agent logs have short retention windows that do not account for investigation timelines. Containment and remediation actions overwrite or age out the signals needed to reconstruct incident context and impact.
By the time investigation starts, prompt context and tool invocation logs from the incident window are partially preserved or gone. Critically, the non-deterministic runtime state that would enable post-incident replay is rarely captured at all: the temperature setting, the system prompt version, the hash of the system prompt, the specific retrieval context injected at runtime (including the vector embeddings and metadata returned by RAG queries), and the model version or weights checkpoint in use at the time of execution. Without freezing the exact retrieval bundle, root cause analysis cannot determine whether a problematic output was driven by the user prompt, the injected context, or a model version change — even if the user prompt itself is preserved.
Detection point: Investigation teams find that logs required for incident reconstruction have aged out of retention or been overwritten by remediation actions. Evidence preservation was not executed before containment, leaving the investigation with incomplete data about agent actions and impact. Runtime parameters such as the RAG retrieval bundle are not logged at all, because existing logging infrastructure was not designed to capture ephemeral inference state.
What changes the outcome: Evidence preservation procedures that execute automatically when agent incidents are declared, with retention periods that match investigation and compliance timelines. Evidence scope should include: prompt and tool invocation logs, the system prompt hash and version, temperature and sampling parameters, the exact retrieval context returned at query time (RAG chunks with associated metadata and embedding identifiers), and the model version or checkpoint identifier. The response object absent here is evidence record.
Failure Pattern Table
| Failure pattern | Why it happens | Where it surfaces | Consequence |
|---|---|---|---|
| Agent identity not resolvable during incident | Agents authenticate via shared API keys, session tokens inherited from delegating principals, or service account credentials not linked to an agent registry; no short-lived per-run token or workload identity links the credential back to a specific agent | Post-incident investigation cannot confirm which agent initiated the access sequence; the investigation team knows a credential was used but not which agent held it | Attribution cannot be established; containment cannot be scoped to the correct agent; accountability cannot be assigned |
| Action chain cannot be reconstructed | Tool invocations, API calls, context retrievals, and workflow triggers log independently in systems not designed for correlation; OTel span context is not propagated across LLM calls, tool execution, and downstream HTTP clients; MCP/A2A gateways do not emit structured telemetry | Investigation confirms individual actions but cannot reconstruct the sequence; the triggering condition cannot be identified; prompt context and tool logs are not correlated | Root cause analysis cannot proceed; the same action chain can recur; post-incident reporting must acknowledge the reconstruction gap |
| Authority scope not preserved at task execution | Governance systems record general agent permissions but do not snapshot the specific scope token or permission set active at dispatch for each task; scope may have changed between incident and investigation | Investigation teams cannot confirm whether a contested action was within the agent's authority at the time; current role bindings do not reflect historical dispatch scope | Accountability review is incomplete; scope drift or misconfiguration cannot be confirmed or ruled out as a contributing factor |
| Containment disrupts unrelated systems | Agents do not have isolated credentials; revoking access requires revoking shared service account credentials or API keys that other agents and systems depend on | IR team chooses between continued agent access or unrelated system outages; narrow revocation is not available | The event continues while blast radius is assessed; containment either exceeds necessary scope or is delayed past the point where it would have prevented propagation |
| Rollback path unknown after multi-system writes | Agent workflows that modify state across multiple systems produce no unified change record; each system retains its own transaction log without linkage to the agent task execution | Post-incident scope assessment requires manual correlation across downstream systems; reversibility is assessed system by system without a coordinated rollback path | Impact scope cannot be confirmed; complete state restoration may not be achievable; partial reversals create inconsistent state across systems |
| No evidence record before state changes | Incident response workflow lacks a structured preservation step before containment; agent logs have short retention windows; non-deterministic runtime state — RAG retrieval bundles, system prompt hash, model version, temperature — is not logged | Prompt context, tool logs, and runtime parameters from the incident window are gone or incomplete by the time investigation begins | Investigation is incomplete; accountability review cannot proceed; post-incident control improvements are based on partial information; replay and root cause analysis are impossible without the retrieval context |
Diagnostic questions
-
Agent identity: When an agent initiates an unexpected action, can your investigation team identify within 30 minutes which specific agent acted — not which credential was used, but which registered agent held that credential at the time? Can you produce a delegation chain showing which human principal authorized the agent's dispatch? Are per-run tokens or workload identities in use, or do agents share long-lived service account credentials?
-
Authority scope: Can you retrieve, for any agent currently in production, the exact permission scope it held at the time of a specific past task execution — not what the agent is generally permitted, but what scope was granted at dispatch for that specific task? Is the scope token or permission snapshot logged at issuance and linked to the task execution identifier?
-
Action chain: If an agent completed a 10-step tool invocation sequence yesterday, can you produce a correlated trace showing each tool invoked, the parameters passed, the API calls made, and the outputs returned — in sequence, with timestamps — without requiring manual correlation across separate systems? Is OpenTelemetry span context propagated across LLM calls, tool executions, and downstream HTTP clients? If MCP or A2A protocol gateways are in use, do they emit structured telemetry at the point of tool authorization?
-
Containment boundary: If you need to contain a specific agent right now without taking down unrelated agents or shared systems, is there a documented procedure that revokes only that agent's access? Or does containment require revoking shared credentials that affect multiple systems?
-
Rollback path: For the last workflow an agent completed that modified records in two or more systems, do you know which records were changed, in which systems, and whether those changes can be reversed — without requiring manual review of each system's transaction log? Does the action chain trace cover downstream writes in sufficient detail to serve as a change manifest?
-
Evidence record: For the last agent-initiated access event your organization investigated, does a structured evidence record exist that covers all six response objects — including the system prompt hash, temperature and sampling parameters, the RAG retrieval bundle returned at query time, and the model version or checkpoint identifier — in a form that satisfies GRC review without requiring an IR practitioner to interpret it?
