The OWASP Top 10 for Agentic Applications 2026: What It Means for Teams Running Agents in Production
Jesse Williams · July 20, 2026

The OWASP Top 10 for Agentic Applications 2026: What It Means for Teams Running Agents in Production

In December 2025, the OWASP GenAI Security Project released the first Top 10 for Agentic Applications. It's the clearest signal yet that the security community has stopped treating agents as chatbots with extra steps and started treating them as what they are: autonomous principals with goals, tools, memory, and the authority to act.

That shift matters. The original OWASP Top 10 for LLM Applications covered passive risks, like what a model might say. The agentic list covers active risks, like what an agent might do. Agents invoke tools, access data, write code, talk to other agents, and spend money. Every one of those capabilities is now an attack surface, and the OWASP list names them one by one.

If you're a CTO, VP of Engineering, or security leader with agents moving from pilot to production, this list is your new threat model. Here's what's in it, and how we think about defending against each risk.

The ten risks

ASI01: Agent Goal Hijack. Attackers redirect an agent's objectives by manipulating instructions, tool outputs, or external content. The now-famous EchoLeak attack is the canonical example: a single email with a hidden payload caused an agent to silently exfiltrate confidential data. No click required. The agent did exactly what it was told, just not by its owner.

ASI02: Tool Misuse and Exploitation. Agents misuse legitimate tools because of prompt injection, ambiguous instructions, or over-privileged access. A tool doesn't have to be malicious to be dangerous. An agent with unrestricted access to a delete API and a manipulated goal is a breach waiting to happen.

ASI03: Identity and Privilege Abuse. Agents aggregate non-human identities. When an agent runs, it acts with the full authority of every key, token, and service account attached to it. Compromise the agent, inherit all of it.

ASI04: Agentic Supply Chain Vulnerabilities. Malicious or tampered tools, MCP servers, models, and agent definitions compromise execution before the agent ever runs. Every new tool or MCP server is a dependency with its own permissions, prompts, and side effects. Most teams have no verification step between "found it on a registry" and "running in production."

ASI05: Unexpected Code Execution. Agents that generate and execute code can be steered into executing attacker-controlled code. The line between "helpful automation" and "remote code execution as a service" is thinner than most teams realize.

ASI06: Memory and Context Poisoning. Persistent corruption of agent memory, RAG stores, or contextual knowledge. Unlike a bad prompt, poisoned memory survives the session. The agent keeps making compromised decisions long after the attack.

ASI07: Insecure Inter-Agent Communication. Multi-agent systems pass instructions and data between agents, often with implicit trust. An attacker who compromises one agent can pivot through the rest.

ASI08: Cascading Agent Failures. One agent's failure propagates through dependent agents and systems. Autonomy plus interconnection means small errors compound fast, and without limits, so do costs.

ASI09: Human-Agent Trust Exploitation. Agents are fluent and persuasive. Attackers exploit that trust by using a hijacked agent to talk a human into approving a malicious action. Forensically, it looks like a legitimate user decision. The manipulation is invisible.

ASI10: Rogue Agents. Compromised or misaligned agents that diverge from intended behavior entirely, whether through tampering, drift, or design.

The pattern underneath the list

Read the ten risks together and a pattern emerges. They fall into three phases of the agent lifecycle:

Before execution. ASI04 lives here. If a tampered model, tool, or MCP server enters your environment, everything downstream is already lost. You can't enforce policy on an artifact you can't trust.

During execution. ASI01, ASI02, ASI03, ASI05, ASI06, ASI07, and ASI09 are runtime risks. They exploit the gap between what an agent is authorized to do and what it actually does on any given invocation. Traditional IAM can verify authorization, but it can't verify behavior. DLP has no primitives for tool invocations or decision chains.

Across the lifecycle. ASI08 and ASI10 are about detection and containment. When something goes wrong, and eventually it will, can you trace what happened, limit the blast radius, and prove what the agent did?

This is why point solutions keep falling short. A prompt injection filter addresses part of ASI01 but nothing about supply chain tampering. An MCP scanner addresses part of ASI04 but nothing about runtime tool misuse. The OWASP list describes a lifecycle problem, and it needs a lifecycle answer.

How Jozu maps to the Top 10

We built Jozu around three verbs: Verify, Enforce, Prove. They map cleanly onto the OWASP list.

Verify (ASI04, ASI10). Jozu Hub signs and scans every model, agent, skill, and MCP server before it can be pulled. Scanning covers nine vulnerability classes including malicious code execution on load, backdoored model behavior, data poisoning, and prompt injection susceptibility in tool descriptions. Artifacts that fail signature, scan, license, or provenance checks are blocked at pull time. The MCP Registry API means VS Code, Cursor, and Claude Desktop can point at a centrally curated, security-scanned catalog instead of the open internet. That closes the front door on agentic supply chain attacks.

Enforce (ASI01, ASI02, ASI03, ASI05, ASI06, ASI07, ASI09). Jozu Agent Guard is a secure runtime for AI, not a traffic inspector bolted on the side. Agents, models, and MCP servers run inside a protected runtime with micro-VM isolation, which contains blast radius if an agent is compromised and boxes in unexpected code execution.

Inside that runtime, three policy kinds govern behavior. ToolPolicy controls which agents can call which tools, with what arguments, under what conditions, with per-tool, per-agent, per-user granularity. That directly counters tool misuse and privilege sprawl. GuardrailPolicy evaluates inference requests and responses for prompt injection, PII exposure, and content safety, addressing goal hijack and memory poisoning vectors at the semantic level. And because the Jozu AI Gateway runs natively inside the runtime, enforcement is content-aware: the policy engine sees what the agent says and does, not just which ports it touches.

For ASI09, human-in-the-loop matters most. ToolPolicy can require explicit human approval for high-risk actions, and approvals are signed attestations recorded outside the agent's conversational channel. The agent can't talk its way past a cryptographic checkpoint.

Prove (ASI08, ASI10). Every policy decision is logged to a tamper-evident, cryptographically chained audit trail spanning the AI supply chain and runtime. Cost metering tracks token consumption and tool invocation counts, so runaway agents surface fast. And Agent Guard fails closed: missing data or evaluation errors result in denial, not silent pass-through. When a cascade starts, the system stops it rather than propagating it. When an agent goes rogue, you can trace it back to source as a single record.

Where to start

The OWASP Top 10 for Agentic Applications is a threat model, not a compliance checkbox. But it's a good forcing function for three questions every team running agents should answer:

  1. Can you verify that every agent, model, and MCP server in production is exactly what you approved? If not, ASI04 is your starting point.
  2. Can you enforce what agents do at runtime, at the level of individual tool calls? If not, ASI01 through ASI03 are open doors.
  3. If an agent misbehaved yesterday, could you prove what it did? If not, ASI08 and ASI10 will find you eventually.

Agents are moving into production across finance, healthcare, and defense whether security teams are ready or not. The OWASP list gives everyone a shared vocabulary for the risks. The work now is closing them.

If you want to see how Verify, Enforce, Prove maps to your agent architecture, talk to us.

Share this post