OWASP Top 10 for LLM Applications 2026: What Changed and How to Update Your AI Test Plan background
Back to Journal
AI Security

OWASP Top 10 for LLM Applications 2026: What Changed and How to Update Your AI Test Plan

Adayptus Consulting
September 27, 2026
15 min read

The 2026 OWASP Top 10 for LLM Applications moves Excessive Agency to third, renames System Prompt Leakage as Hidden Context Exposure, and is the first edition ranked partly on incident data. What changed, what it signals, and what to add to an AI test plan.

AI Security

OWASP released the 2026 edition of its Top 10 for LLM Applications on 3 August 2026. The same ten risks are mostly still there, but the order has moved, one category has been renamed and broadened, and several have grown to cover agents, memory and images. Here is what changed, what the changes are telling you, and what to add to an AI security test plan written against the 2025 list.

In short. Excessive Agency jumps from sixth to third, Unbounded Consumption from tenth to sixth, and Improper Output Handling falls from fifth to tenth. System Prompt Leakage becomes Hidden Context Exposure, which now covers anything the application puts in front of the model without the user seeing it, including retrieved documents and tool definitions. It is also the first edition whose ranking was weighted partly by real incident data. If your test plan or your reports cite category numbers, they have changed: always write the year.

What OWASP released, and how it was ranked

The Top 10 for LLM Applications is published by the OWASP GenAI Security Project. Its previous edition, labelled 2025, came out in late 2024 and is the version most AI test plans and vendor questionnaires in use today were written against. The 2026 edition supersedes it.

The method changed as well as the list. Earlier editions were ranked by practitioner consensus. For 2026, the community vote carried 75 percent of the weight and the remaining 25 percent was informed by data from 6,639 real incidents drawn from public vulnerability databases and an AI-harm database. It is the first time evidence of what has actually gone wrong in deployed systems has shaped the order, and it is the main reason the order moved.

The 2026 list beside the 2025 list

2026Category2025 positionMovement
LLM01Prompt InjectionLLM01Unchanged; scope expanded to cross-modal input
LLM02Sensitive Information DisclosureLLM02Unchanged
LLM03Excessive AgencyLLM06Up three
LLM04Supply ChainLLM03Down one
LLM05Data and Model PoisoningLLM04Down one; now includes fine-tuning subversion
LLM06Unbounded ConsumptionLLM10Up four
LLM07MisinformationLLM09Up two
LLM08Hidden Context ExposureLLM07, as System Prompt LeakageRenamed and broadened
LLM09Vector and Embedding WeaknessesLLM08Down one; scope refined
LLM10Improper Output HandlingLLM05Down five

What the reshuffle is telling you

Read the movements together and three shifts stand out, each of which matches what is changing in how LLM applications are built.

Agents are in production

Excessive Agency moving to third is the clearest signal in the list. An application that can only answer questions can leak data; an application that can call tools, send messages, change records or spend money can act on an attacker's behalf. Once the model is wired to tools, the question stops being what it might say and becomes what it is allowed to do.

Cost and availability are security problems

Unbounded Consumption climbing from tenth to sixth reflects applications where every request costs money. An attacker who can make the model loop, call expensive tools repeatedly, or process enormous inputs can run up a bill or deny service without ever reading a byte of data they should not.

Harm is not only about confidentiality

Misinformation rising to seventh reflects real liability from applications that confidently give wrong answers to customers. It is less a penetration-testing category than an evaluation one, but it belongs in a risk assessment of any customer-facing assistant.

The biggest fall, Improper Output Handling from fifth to tenth, is worth reading carefully. It does not mean model output has become safe to render or execute. Imperva's reading of the change is that the industry already had established controls for it: output encoding, parameterised queries and avoiding shell construction are well understood. The risk is unchanged when those controls are missing. It is still tested.

Renamed and broadened: Hidden Context Exposure

System Prompt Leakage was about one thing: getting the model to reveal its system prompt. Hidden Context Exposure covers everything an application places in front of the model without the user seeing it. That now includes the system prompt, documents pulled in by retrieval, tool and function definitions, cached context from earlier turns or other sessions, retrieval schemas and hidden policy logic.

The broadening matters because the system prompt was rarely the most sensitive thing in context. Tool definitions describe what the application can do and often reveal internal endpoints and parameter names. Retrieved documents may contain data the current user is not entitled to. Cached context can carry one user's conversation into another's. A test plan that only tries "repeat the text above" has covered the smallest part of the category.

Hidden contextWhat a tester triesWhy it matters
System promptDirect and indirect extraction, translation and summarisation tricks, reconstruction across several turnsReveals guardrails to work around, and sometimes credentials or internal names
Tool definitionsAsking the model to describe or list its tools and their parametersA map of what the application can do, often including internal API paths
Retrieved documentsQueries designed to pull in documents outside the user's entitlement, then asking for verbatim quotationRetrieval that is not permission-aware becomes a data leak with a chat interface
Cached contextTwo accounts, sequential sessions, checking whether one session's content surfaces in anotherCross-user leakage through caching or shared memory

The scope expansions that add test cases

Several categories kept their names but grew. These are the changes most likely to leave a gap in a test plan written against the 2025 list.

  • Prompt Injection now explicitly covers cross-modal attacks. Instructions hidden in images or audio that a multimodal model processes are in scope. If your application accepts uploads that reach a multimodal model, those uploads are injection vectors, and a text-only test plan misses them.
  • Data and Model Poisoning absorbs fine-tuning subversion. If you fine-tune on data that users or third parties can influence, the integrity of that pipeline is now part of this category, not a separate concern.
  • Vector and Embedding Weaknesses covers more than retrieval. Anywhere similarity search sits between a data source and the prompt is in scope: retrieval-augmented generation, vector-backed agent memory, semantic caches and deduplication. Attacks include cross-tenant retrieval from a shared vector store, access controls that were never made permission-aware, and embedding inversion. Many of these succeed even when the retrieved content contains no instructions at all.

Updating a test plan written against 2025

Step 1 — Re-map, and stamp the year on every reference

Category numbers moved. LLM06 meant Excessive Agency in 2025 and means Unbounded Consumption in 2026; LLM07 meant System Prompt Leakage and now means Misinformation. Re-map existing test cases and open findings to the 2026 numbering and write references as LLM03:2026, never as a bare number.

Step 2 — Build an agency matrix

For every tool the application can call, record what it can do, with whose permissions, and whether a human approves it first. Then test whether a manipulated conversation can reach each tool, with arguments the user should not be able to supply. This is authorisation testing with a model in the middle, and it needs the same discipline as testing API authorisation: two accounts and a list of every action.

Step 3 — Widen hidden-context tests beyond the system prompt

Add tool-definition disclosure, retrieval outside entitlement and cross-session leakage, as in the table above. Run them with at least two accounts so a leak can be proved, not inferred.

Step 4 — Add consumption tests, carefully

Test for recursive tool loops, oversized inputs, expensive retrieval and missing per-user budgets. Agree limits with the application owner before running them, and prefer a staging environment with real rate limits configured: consumption testing in production is a denial-of-service you paid for.

Step 5 — Add cross-modal injection where there are uploads

If images, documents or audio reach a multimodal model, test instructions embedded in them, including indirect cases where the upload is processed later by an agent rather than in the user's own session.

Step 6 — Keep testing output handling

Its fall in rank reflects better controls, not a smaller risk. Check that model output rendered as HTML is encoded, that generated SQL is parameterised, and that nothing generated is passed to a shell. These are the findings that turn a chat feature into a conventional, high-severity web vulnerability.

What each category needs as evidence

LLM findings are easy to dismiss because the model's behaviour varies between runs. A finding should be reproducible from what is in the report, with the conversation and the consequence both shown.

Category (2026)A finding should show
LLM01 Prompt InjectionThe injected input, where it entered (direct, document, image, tool result), the model's changed behaviour, and the reproduction rate across repeated runs
LLM02 Sensitive Information DisclosureThe disclosed data, whose it was, and why the requesting user was not entitled to it
LLM03 Excessive AgencyThe tool invoked, the arguments, the effect on the target system, and the permission or approval step that should have stopped it
LLM06 Unbounded ConsumptionThe request pattern, the measured cost or resource use, and the missing limit
LLM08 Hidden Context ExposureThe hidden content recovered, verbatim, and what it reveals
LLM09 Vector and Embedding WeaknessesThe query, the retrieved item, and proof it belongs to another tenant or sits outside the user's entitlement
LLM10 Improper Output HandlingThe generated payload and the downstream effect: script execution, query change or command run

What has not changed

The direction of the 2026 edition, as reported at its release, is that defenders should stop trying to build a model that cannot be fooled and instead build the system around the model so that when it is fooled, nothing important breaks. That is the right frame for testing too. The most valuable findings are rarely "the model said something it should not". They are "the model was manipulated, and the application let that manipulation reach a tool, a record or another user's data".

That puts most of the fixes where they always were: authorisation enforced outside the model, least-privilege tool permissions, human approval for consequential actions, permission-aware retrieval, per-user budgets, and output treated as untrusted input. The list reorders the priorities. It does not change where the controls live.

Frequently Asked Questions

Click any question to expand the answer.

QWhen was the 2026 OWASP Top 10 for LLM Applications released?

On 3 August 2026, by the OWASP GenAI Security Project. It supersedes the edition labelled 2025, which was published in late 2024.

QWhat is the full 2026 list?

LLM01 Prompt Injection, LLM02 Sensitive Information Disclosure, LLM03 Excessive Agency, LLM04 Supply Chain, LLM05 Data and Model Poisoning, LLM06 Unbounded Consumption, LLM07 Misinformation, LLM08 Hidden Context Exposure, LLM09 Vector and Embedding Weaknesses, LLM10 Improper Output Handling.

QWhat happened to System Prompt Leakage?

It was renamed Hidden Context Exposure and broadened. It now covers everything an application places in front of the model without the user seeing it, including retrieved documents, tool definitions and cached context, not only the system prompt. It sits at LLM08.

QHow was the 2026 list ranked?

Practitioner voting carried 75 percent of the weight, and 25 percent was informed by data from 6,639 real incidents drawn from public vulnerability databases and an AI-harm database. It is the first edition in which incident evidence shaped the ranking.

QWhy did Excessive Agency move up to third?

Because LLM applications increasingly call tools and take actions rather than only answering questions. Once a model can send messages, change records or spend money, a manipulated conversation can cause real effects, and the incident data reflects that.

QDoes Improper Output Handling falling to tenth mean it matters less?

No. Its fall is widely read as reflecting established controls such as output encoding and parameterised queries. Where those controls are missing, model output can still produce cross-site scripting, injection or command execution, and it should still be tested.

QOur reports cite LLM06. Is that still correct?

It depends on the year. LLM06 was Excessive Agency in 2025 and is Unbounded Consumption in 2026. Re-map existing references and always write the year, for example LLM03:2026, so a reader cannot confuse editions.

QWe use a third-party model API. Which of these are our responsibility?

Most of them. The provider secures the model; your prompts, retrieval, tool permissions, output handling, budgets and caching are yours. Excessive agency, hidden context exposure, permission-unaware retrieval and improper output handling all live in the integration you built.

QCan an automated scanner test the LLM Top 10?

Partly. Automated tools are useful for running large libraries of known injection prompts. They cannot judge whether a retrieved document was outside a user's entitlement or whether a tool call should have needed approval, because they do not know your authorisation model. Those judgements are manual.

QIs unbounded consumption safe to test in production?

Usually not. A successful test is, by definition, excessive resource use. Agree limits with the application owner, prefer a staging environment configured with production rate limits, and keep the evidence to the measured cost of a bounded number of requests.

For the wider testing methodology behind LLM applications, see securing generative AI with advanced LLM security testing. For the governance side, how ISO/IEC 42001 and the NIST AI Risk Management Framework fit together, see AI governance with ISO 42001 and NIST AI RMF. For why this belongs on the leadership agenda, why secure AI matters to management.

About Adayptus

Adayptus Consulting Private Limited is a cybersecurity consultancy based in Noida, India, and has been doing application security testing since 2018. Our AI security work applies the same discipline to LLM applications: authorisation tested with real accounts, every finding reproduced, and evidence a developer can act on.

What we can do for you:

  • LLM security testing against the 2026 list. Our LLM security testing covers prompt injection including cross-modal input, the agency matrix for every tool, hidden context exposure, permission-aware retrieval and output handling, with findings mapped to the 2026 categories.
  • The rest of the application. An LLM feature sits on ordinary infrastructure. An AI security assessment covers the model integration alongside the APIs and web application around it, and threat modelling before build puts the trust boundaries in the right place.
  • Governance that keeps up. An AI governance framework built on ISO/IEC 42001 and the NIST AI RMF, so the list's next revision is a scheduled update rather than a surprise.

Attribution and sources

Category names and identifiers are from the OWASP Top 10 for LLM Applications 2026, published by the OWASP GenAI Security Project under Creative Commons Attribution-ShareAlike 4.0 International. The testing methodology, the matrices and the evidence guidance are our own.

The list is revised periodically. Confirm the current edition on the project site before relying on the numbering.


Share this Insight
CybersecurityAI SecurityAdayptus Intelligence
A

Adayptus Consulting

AI Security Testing, Adayptus

Adayptus Consulting Private Limited is a cybersecurity consultancy based in Noida, India, doing application security testing since 2018 and applying the same discipline to LLM and AI applications.

AI Security

Test Your AI Application Against the 2026 List

Agents that call tools, retrieval that is not permission-aware, and context the user never sees are where LLM applications fail. We test them with real accounts, map every finding to the 2026 categories, and show the conversation and the consequence for each. Tell us what your application can read and do, and we will come back with scope and an indicative quote, usually within one business day.

  • Agency matrix for every tool the model can call
  • Hidden context, retrieval and cross-session leakage tested with two accounts
  • Every finding reproduced and mapped to the 2026 categories
  • Free remediation retest once fixes are in
Direct Scoping Hotline: +91-9625999069 [email protected]

Request a scoping call

No obligation. A senior consultant replies — not a sales sequence.

Your details stay confidential. Covered by NDA — a senior consultant replies directly.

Zero False Positives Free Retest Included 100% NDA Protected