
OWASP Top 10 for LLM Applications 2026: What Changed and How to Update Your AI Test Plan
The 2026 OWASP Top 10 for LLM Applications moves Excessive Agency to third, renames System Prompt Leakage as Hidden Context Exposure, and is the first edition ranked partly on incident data. What changed, what it signals, and what to add to an AI test plan.
AI Security
OWASP released the 2026 edition of its Top 10 for LLM Applications on 3 August 2026. The same ten risks are mostly still there, but the order has moved, one category has been renamed and broadened, and several have grown to cover agents, memory and images. Here is what changed, what the changes are telling you, and what to add to an AI security test plan written against the 2025 list.
In short. Excessive Agency jumps from sixth to third, Unbounded Consumption from tenth to sixth, and Improper Output Handling falls from fifth to tenth. System Prompt Leakage becomes Hidden Context Exposure, which now covers anything the application puts in front of the model without the user seeing it, including retrieved documents and tool definitions. It is also the first edition whose ranking was weighted partly by real incident data. If your test plan or your reports cite category numbers, they have changed: always write the year.
What OWASP released, and how it was ranked
The Top 10 for LLM Applications is published by the OWASP GenAI Security Project. Its previous edition, labelled 2025, came out in late 2024 and is the version most AI test plans and vendor questionnaires in use today were written against. The 2026 edition supersedes it.
The method changed as well as the list. Earlier editions were ranked by practitioner consensus. For 2026, the community vote carried 75 percent of the weight and the remaining 25 percent was informed by data from 6,639 real incidents drawn from public vulnerability databases and an AI-harm database. It is the first time evidence of what has actually gone wrong in deployed systems has shaped the order, and it is the main reason the order moved.
The 2026 list beside the 2025 list
| 2026 | Category | 2025 position | Movement |
|---|---|---|---|
| LLM01 | Prompt Injection | LLM01 | Unchanged; scope expanded to cross-modal input |
| LLM02 | Sensitive Information Disclosure | LLM02 | Unchanged |
| LLM03 | Excessive Agency | LLM06 | Up three |
| LLM04 | Supply Chain | LLM03 | Down one |
| LLM05 | Data and Model Poisoning | LLM04 | Down one; now includes fine-tuning subversion |
| LLM06 | Unbounded Consumption | LLM10 | Up four |
| LLM07 | Misinformation | LLM09 | Up two |
| LLM08 | Hidden Context Exposure | LLM07, as System Prompt Leakage | Renamed and broadened |
| LLM09 | Vector and Embedding Weaknesses | LLM08 | Down one; scope refined |
| LLM10 | Improper Output Handling | LLM05 | Down five |
What the reshuffle is telling you
Read the movements together and three shifts stand out, each of which matches what is changing in how LLM applications are built.
Agents are in production
Excessive Agency moving to third is the clearest signal in the list. An application that can only answer questions can leak data; an application that can call tools, send messages, change records or spend money can act on an attacker's behalf. Once the model is wired to tools, the question stops being what it might say and becomes what it is allowed to do.
Cost and availability are security problems
Unbounded Consumption climbing from tenth to sixth reflects applications where every request costs money. An attacker who can make the model loop, call expensive tools repeatedly, or process enormous inputs can run up a bill or deny service without ever reading a byte of data they should not.
Harm is not only about confidentiality
Misinformation rising to seventh reflects real liability from applications that confidently give wrong answers to customers. It is less a penetration-testing category than an evaluation one, but it belongs in a risk assessment of any customer-facing assistant.
The biggest fall, Improper Output Handling from fifth to tenth, is worth reading carefully. It does not mean model output has become safe to render or execute. Imperva's reading of the change is that the industry already had established controls for it: output encoding, parameterised queries and avoiding shell construction are well understood. The risk is unchanged when those controls are missing. It is still tested.
Renamed and broadened: Hidden Context Exposure
System Prompt Leakage was about one thing: getting the model to reveal its system prompt. Hidden Context Exposure covers everything an application places in front of the model without the user seeing it. That now includes the system prompt, documents pulled in by retrieval, tool and function definitions, cached context from earlier turns or other sessions, retrieval schemas and hidden policy logic.
The broadening matters because the system prompt was rarely the most sensitive thing in context. Tool definitions describe what the application can do and often reveal internal endpoints and parameter names. Retrieved documents may contain data the current user is not entitled to. Cached context can carry one user's conversation into another's. A test plan that only tries "repeat the text above" has covered the smallest part of the category.
| Hidden context | What a tester tries | Why it matters |
|---|---|---|
| System prompt | Direct and indirect extraction, translation and summarisation tricks, reconstruction across several turns | Reveals guardrails to work around, and sometimes credentials or internal names |
| Tool definitions | Asking the model to describe or list its tools and their parameters | A map of what the application can do, often including internal API paths |
| Retrieved documents | Queries designed to pull in documents outside the user's entitlement, then asking for verbatim quotation | Retrieval that is not permission-aware becomes a data leak with a chat interface |
| Cached context | Two accounts, sequential sessions, checking whether one session's content surfaces in another | Cross-user leakage through caching or shared memory |
The scope expansions that add test cases
Several categories kept their names but grew. These are the changes most likely to leave a gap in a test plan written against the 2025 list.
- Prompt Injection now explicitly covers cross-modal attacks. Instructions hidden in images or audio that a multimodal model processes are in scope. If your application accepts uploads that reach a multimodal model, those uploads are injection vectors, and a text-only test plan misses them.
- Data and Model Poisoning absorbs fine-tuning subversion. If you fine-tune on data that users or third parties can influence, the integrity of that pipeline is now part of this category, not a separate concern.
- Vector and Embedding Weaknesses covers more than retrieval. Anywhere similarity search sits between a data source and the prompt is in scope: retrieval-augmented generation, vector-backed agent memory, semantic caches and deduplication. Attacks include cross-tenant retrieval from a shared vector store, access controls that were never made permission-aware, and embedding inversion. Many of these succeed even when the retrieved content contains no instructions at all.
Updating a test plan written against 2025
Step 1 — Re-map, and stamp the year on every reference
Category numbers moved. LLM06 meant Excessive Agency in 2025 and means Unbounded Consumption in 2026; LLM07 meant System Prompt Leakage and now means Misinformation. Re-map existing test cases and open findings to the 2026 numbering and write references as LLM03:2026, never as a bare number.
Step 2 — Build an agency matrix
For every tool the application can call, record what it can do, with whose permissions, and whether a human approves it first. Then test whether a manipulated conversation can reach each tool, with arguments the user should not be able to supply. This is authorisation testing with a model in the middle, and it needs the same discipline as testing API authorisation: two accounts and a list of every action.
Step 3 — Widen hidden-context tests beyond the system prompt
Add tool-definition disclosure, retrieval outside entitlement and cross-session leakage, as in the table above. Run them with at least two accounts so a leak can be proved, not inferred.
Step 4 — Add consumption tests, carefully
Test for recursive tool loops, oversized inputs, expensive retrieval and missing per-user budgets. Agree limits with the application owner before running them, and prefer a staging environment with real rate limits configured: consumption testing in production is a denial-of-service you paid for.
Step 5 — Add cross-modal injection where there are uploads
If images, documents or audio reach a multimodal model, test instructions embedded in them, including indirect cases where the upload is processed later by an agent rather than in the user's own session.
Step 6 — Keep testing output handling
Its fall in rank reflects better controls, not a smaller risk. Check that model output rendered as HTML is encoded, that generated SQL is parameterised, and that nothing generated is passed to a shell. These are the findings that turn a chat feature into a conventional, high-severity web vulnerability.
What each category needs as evidence
LLM findings are easy to dismiss because the model's behaviour varies between runs. A finding should be reproducible from what is in the report, with the conversation and the consequence both shown.
| Category (2026) | A finding should show |
|---|---|
| LLM01 Prompt Injection | The injected input, where it entered (direct, document, image, tool result), the model's changed behaviour, and the reproduction rate across repeated runs |
| LLM02 Sensitive Information Disclosure | The disclosed data, whose it was, and why the requesting user was not entitled to it |
| LLM03 Excessive Agency | The tool invoked, the arguments, the effect on the target system, and the permission or approval step that should have stopped it |
| LLM06 Unbounded Consumption | The request pattern, the measured cost or resource use, and the missing limit |
| LLM08 Hidden Context Exposure | The hidden content recovered, verbatim, and what it reveals |
| LLM09 Vector and Embedding Weaknesses | The query, the retrieved item, and proof it belongs to another tenant or sits outside the user's entitlement |
| LLM10 Improper Output Handling | The generated payload and the downstream effect: script execution, query change or command run |
What has not changed
The direction of the 2026 edition, as reported at its release, is that defenders should stop trying to build a model that cannot be fooled and instead build the system around the model so that when it is fooled, nothing important breaks. That is the right frame for testing too. The most valuable findings are rarely "the model said something it should not". They are "the model was manipulated, and the application let that manipulation reach a tool, a record or another user's data".
That puts most of the fixes where they always were: authorisation enforced outside the model, least-privilege tool permissions, human approval for consequential actions, permission-aware retrieval, per-user budgets, and output treated as untrusted input. The list reorders the priorities. It does not change where the controls live.
Frequently Asked Questions
Click any question to expand the answer.
QWhen was the 2026 OWASP Top 10 for LLM Applications released?
On 3 August 2026, by the OWASP GenAI Security Project. It supersedes the edition labelled 2025, which was published in late 2024.
QWhat is the full 2026 list?
LLM01 Prompt Injection, LLM02 Sensitive Information Disclosure, LLM03 Excessive Agency, LLM04 Supply Chain, LLM05 Data and Model Poisoning, LLM06 Unbounded Consumption, LLM07 Misinformation, LLM08 Hidden Context Exposure, LLM09 Vector and Embedding Weaknesses, LLM10 Improper Output Handling.
QWhat happened to System Prompt Leakage?
It was renamed Hidden Context Exposure and broadened. It now covers everything an application places in front of the model without the user seeing it, including retrieved documents, tool definitions and cached context, not only the system prompt. It sits at LLM08.
QHow was the 2026 list ranked?
Practitioner voting carried 75 percent of the weight, and 25 percent was informed by data from 6,639 real incidents drawn from public vulnerability databases and an AI-harm database. It is the first edition in which incident evidence shaped the ranking.
QWhy did Excessive Agency move up to third?
Because LLM applications increasingly call tools and take actions rather than only answering questions. Once a model can send messages, change records or spend money, a manipulated conversation can cause real effects, and the incident data reflects that.
QDoes Improper Output Handling falling to tenth mean it matters less?
No. Its fall is widely read as reflecting established controls such as output encoding and parameterised queries. Where those controls are missing, model output can still produce cross-site scripting, injection or command execution, and it should still be tested.
QOur reports cite LLM06. Is that still correct?
It depends on the year. LLM06 was Excessive Agency in 2025 and is Unbounded Consumption in 2026. Re-map existing references and always write the year, for example LLM03:2026, so a reader cannot confuse editions.
QWe use a third-party model API. Which of these are our responsibility?
Most of them. The provider secures the model; your prompts, retrieval, tool permissions, output handling, budgets and caching are yours. Excessive agency, hidden context exposure, permission-unaware retrieval and improper output handling all live in the integration you built.
QCan an automated scanner test the LLM Top 10?
Partly. Automated tools are useful for running large libraries of known injection prompts. They cannot judge whether a retrieved document was outside a user's entitlement or whether a tool call should have needed approval, because they do not know your authorisation model. Those judgements are manual.
QIs unbounded consumption safe to test in production?
Usually not. A successful test is, by definition, excessive resource use. Agree limits with the application owner, prefer a staging environment configured with production rate limits, and keep the evidence to the measured cost of a bounded number of requests.
Related reading
For the wider testing methodology behind LLM applications, see securing generative AI with advanced LLM security testing. For the governance side, how ISO/IEC 42001 and the NIST AI Risk Management Framework fit together, see AI governance with ISO 42001 and NIST AI RMF. For why this belongs on the leadership agenda, why secure AI matters to management.
About Adayptus
Adayptus Consulting Private Limited is a cybersecurity consultancy based in Noida, India, and has been doing application security testing since 2018. Our AI security work applies the same discipline to LLM applications: authorisation tested with real accounts, every finding reproduced, and evidence a developer can act on.
What we can do for you:
- LLM security testing against the 2026 list. Our LLM security testing covers prompt injection including cross-modal input, the agency matrix for every tool, hidden context exposure, permission-aware retrieval and output handling, with findings mapped to the 2026 categories.
- The rest of the application. An LLM feature sits on ordinary infrastructure. An AI security assessment covers the model integration alongside the APIs and web application around it, and threat modelling before build puts the trust boundaries in the right place.
- Governance that keeps up. An AI governance framework built on ISO/IEC 42001 and the NIST AI RMF, so the list's next revision is a scheduled update rather than a surprise.
Attribution and sources
Category names and identifiers are from the OWASP Top 10 for LLM Applications 2026, published by the OWASP GenAI Security Project under Creative Commons Attribution-ShareAlike 4.0 International. The testing methodology, the matrices and the evidence guidance are our own.
- Help Net Security, OWASP 2026 LLM Top 10 released, 6 August 2026, for the ranking methodology and scope changes.
- Imperva, OWASP LLM Top 10 2026: what changed and why, for the category-by-category movements.
The list is revised periodically. Confirm the current edition on the project site before relying on the numbering.
Adayptus Consulting
AI Security Testing, Adayptus
Adayptus Consulting Private Limited is a cybersecurity consultancy based in Noida, India, doing application security testing since 2018 and applying the same discipline to LLM and AI applications.
On This Page
- What OWASP released, and how it was ranked
- The 2026 list beside the 2025 list
- What the reshuffle is telling you
- Renamed and broadened: Hidden Context Exposure
- The scope expansions that add test cases
- Updating a test plan written against 2025
- What each category needs as evidence
- What has not changed
- Frequently Asked Questions
- Related reading
- About Adayptus
- Attribution and sources


