Identity

How identity and permissions become the blast-radius boundary for LLMs

LLM technology integrated into complex circuit board with illuminated traces and central processing unit

COMMENTARY: In the past, machine processes were predictable, designed to do one or a small defined handful of tasks, and then shut down after the programmed job was finished.

Decades of secure-computing design kept program logic and user input separated, and a vast amount of time and effort went into setting these systems up properly. Large language models (LLM) collapse instructions and data into a single stream, erasing a separation that previous secure-computing design worked to preserve. A system prompt, a user’s request, and a retrieved document all arrive as text that might carry instructions.

[SC Media Perspectives columns are written by a trusted community of SC Media cybersecurity subject matter experts. Read more Perspectives here.]

Unlike deterministic machine identities, agents built on these models are non-deterministic. We can ask an agent to do the same task twice, and it may take a number of different approaches. This flexibility and speed  makes AI very appealing, but it’s also why controls for predictable automation don’t carry over.

Prompt injection attacks rank first in the OWASP Top 10 for LLM applications and demonstrate just how fluid access boundaries are.  With this technique, attackers take advantage of the natural language processing capabilities of LLMs to inject commands that the model interprets as legitimate. These are direct or indirect (hidden in a web page, email, or file the agent reads), but the important point here: the system can be tricked into overriding its original instructions and/or guardrails in the prompt. That's why prompt injection isn't a bug waiting on a patch: it's a structural property of how these systems work, and no amount of guardrails fully closes the gap between what a model gets told to do and what it's later manipulated into doing.

The blast radius of a misled or compromised agent is potentially huge. Because agents often hold broad, standing permissions, a single compromise can trigger a devastating chain reaction -- very quickly. Picture an agent that triages support tickets on a service account with read-write access to the customer database. A single email hiding instructions to copy the customer table and delete the originals gives the agent everything it needs to do both.

Why we’re not asking the right questions

Many organizations are struggling right now with the issue of trust. How do I know I can trust the model I’m working with? How do I make it more trustworthy? Can I control it? I’d argue that these aren’t the right questions. We can embed rules and instructions into an AI model, but a guardrail that blocks a bad response today could fail tomorrow. Prompt injection attacks show just how easily these systems are steered. Assume that the system’s not trustworthy, and instead ask: If an agent gets manipulated, how much damage can it cause?

The answer depends on what an agent can reach. An LLM on its own doesn’t execute anything, but the agent built on it does through the use of tools and permissions it’s given. OWASP calls this risk Excessive Agency, and tells security pros to enforce authorization in the systems the agent reaches instead of relying on the model to decide what’s allowed.

Just like human users, we need to govern AI agents by strict, real-time identity and access controls. Teams must grant agents the absolute minimum data access, tool permissions and network privileges they need to complete their tasks. And nothing more.

When an agent acts for a person, it should work within that person’s permissions, and high-impact actions should require human approval. A distinct identity tells us which agent acted and who delegated the work, but the blast radius gets defined entirely by the privileges attached to it.  When we enforce time-bound, task-specific permissions across the entire network, the blast radius of any compromised or misled agent (or human) shrinks to the privileges it holds for that task.

Instead of struggling to make LLMs more trustworthy, security teams should constrain what agents are allowed to do. When an agent’s access only gets issued when needed and withdrawn the moment a task ends, the idea behind zero standing privileges, even a successful injection attack turns into a near-miss instead of a breach.

Scope that support-ticket agent to the task with read access to one customer’s records for as long as the ticket remains open, and the “copy” and “delete” are both blocked before they reach the database. The injection still happens, it just doesn’t reach anything that matters. When the team can’t secure the model itself -- identity -- and the privileges attached to it become the last line of defense – and the one that actually holds.

Art Poghosyan, co-founder and CEO, Britive

SC Media Perspectives columns are written by a trusted community of SC Media cybersecurity subject matter experts. Each contribution has a goal of bringing a unique voice to important cybersecurity topics. Content strives to be of the highest quality, objective and non-commercial.

You can skip this ad in 5 seconds