Onpode
Cover art for Real incident data just overruled AI experts' rankings — misinformation jumped two spots in OWASP's LLM Top 10

Real incident data just overruled AI experts' rankings — misinformation jumped two spots in OWASP's LLM Top 10

August 7, 2026 · 9 min

Eliza Ward & Brian Reed

OWASP's 2026 LLM Top 10, based on 7,714 real-world incidents, moved Misinformation up two spots to LLM07 — but the ranking formula is 75% expert consensus and only 25% incident data, a deliberate cap that left Prompt Injection at number one for the third straight year despite low incident counts.

The Open Worldwide Application Security Project (OWASP) GenAI Security Project published the third edition of its Top 10 for LLM Applications on August 4, 2026. This release marked a methodological departure from prior editions (2023 and 2025), which had relied entirely on expert consensus voting.

0:009:14
Get the next episode on top ten lists

Follow it free — new episodes land in your feed.

Or make your own — any topic, in minutes
About this episode

OWASP's 2026 LLM Top 10 is being called its most data-driven edition yet — and the claim isn't wrong, exactly. For the first time, the project incorporated real-world incident data: 7,714 incidents analyzed, 6,639 classifiable. Misinformation jumped two spots to LLM07. That moved the needle. But the episode sits with something the headlines mostly walked past: the final rankings weighted expert consensus at 75% and incident data at 25%, by deliberate design, and OWASP said so openly. Prompt Injection held the number one slot for the third consecutive year despite a low incident count. Both outcomes — Misinformation's rise and Prompt Injection's lock on the top — came from the same formula. The episode works through what that asymmetry actually means, not as a gotcha but as a real structural question: when the data agrees with experts, it confirms the ranking; when it disagrees, experts absorb 75% of the hit. The 'more empirical' framing changes how practitioners trust the list without changing what the data is actually allowed to do. There's a harder downstream problem too: security frameworks are being built against these rankings right now, and OWASP hasn't committed to whether 25% is a permanent ceiling or just where they started. The moment to watch is the next edition — not whether the rankings shift, but whether the methodology gets fixed before the frameworks calcify around it.

Frequently asked

What changed in OWASP LLM Top 10 2026?

OWASP's 2026 LLM Top 10, published August 4, is the first edition to incorporate real-world incident data — 7,714 incidents total, of which 6,639 were classifiable. The most visible change was Misinformation jumping two spots to LLM07. The list also expanded scope from chat applications to agentic and tool-using systems.

Why is Prompt Injection still number one in OWASP LLM Top 10 2026?

Prompt Injection has held the number one position in every OWASP LLM Top 10 edition — 2023, 2025, and 2026 — despite low real-world incident counts. OWASP's ranking formula assigns 75% weight to expert consensus and only 25% to incident data, so expert judgment structurally dominates and Prompt Injection's top ranking was never at risk from the data.

How does OWASP weigh incident data versus expert opinion in the LLM Top 10?

OWASP's 2026 LLM Top 10 uses a deliberate 75/25 split: 75% of each ranking reflects expert consensus, 25% reflects incident data. OWASP project leads stated publicly that 'one noisy year of data does not get to overturn the judgment of the people doing the work.' OWASP has not committed to whether 25% is permanent or just a starting point.

Did real incident data actually change OWASP's LLM risk rankings?

Real incident data had a measurable but structurally limited effect on OWASP's 2026 LLM Top 10. Misinformation moved up two spots to LLM07 — a significant shift given incident data holds only 25% of the ranking weight. Prompt Injection, which had low incident frequency, remained at number one, unchanged by the data component.

What is the practical risk of OWASP LLM Top 10 rankings for security teams?

Security teams relying on OWASP's 2026 LLM Top 10 may triage threats based on expert perception rather than incident frequency, since 75% of each ranking reflects consensus, not field data. Prompt Injection's number one position carries no disclosed incident record. Frameworks like the GenAI Incident Response Framework are already being built against these rankings, making the methodology's instability a downstream risk.

Grounded in 8 sources
A Practical Incident-Response Framework for Generative AI Systems · doi.org
How the OWASP LLM Top 10 2026 Was Ranked: Vote vs Incident Record | AltaySec · altaysec.com.tr
Cybersecurity News article on OWASP LLM Top 10 2026 gains traction · cybersecuritynews.com
OWASP GenAI Security Project drops 2026 LLM Top 10 list · helpnetsecurity.com
OWASP 2026 LLM Top 10: "The model will be fooled" · helpnetsecurity.com
OWASP LLM Top 10 2026: 7,714 Incidents Analyzed – And the #1 Risk Almost Didn’t Make It · invicti.com
Prompt injection remains top LLM threat, OWASP report finds - scworld.com · scworld.com
OWASP LLM Top 10 2026 Incident Data Overrules Experts on Misinformation Risk · techtimes.com
Read transcript

Eliza Ward: Hey — genuinely curious what your read is on the OWASP thing.

Brian Reed: The 2026 list? I read it twice and I'm still — let me see — I'm not sure the results cohere.

Eliza Ward: Published August 4th by the OWASP GenAI Security Project. Third edition overall, first one to actually incorporate real-world incident data.

Brian Reed: And the data moved Misinformation — from around ninth up to seventh.

Eliza Ward: That's LLM07 now. Two spots up.

Brian Reed: Which, fine — you add incident data, some things move. But Prompt Injection is still number one, for the third consecutive year, and from what I can tell the incident count for Prompt Injection is... not high.

Eliza Ward: Hold on — so the narrative is 'we finally have real data driving the list,' and the data moved Misinformation up two places, but it didn't touch the number one slot at all?

Brian Reed: That's what puzzles me. Both outcomes came from the same methodology. And they're pointing in completely different directions about how much incidents actually matter.

Eliza Ward: Right — but the part that breaks the clean version is that it wasn't hidden. They said it explicitly.

Brian Reed: OWASP looked at 7,714 incidents — Invicti reported that number — and then just... decided that 75% of the final ranking goes to expert consensus, 25% to incident data. And they said it out loud. Help Net Security quoted the project leads directly: 'one noisy year of data does not get to overturn the judgment of the people doing the work.'

Eliza Ward: They published that quote.

Brian Reed: Publicly. So here's the plain version: think of a hiring committee. They interview candidates — real interviews, real data — but three-quarters of the final decision belongs to the senior partners' read on fit. The interviews happen, they count, they can shift things at the margins. Misinformation moved up two spots. But they cannot flip the table on whoever the partners already believe is the best candidate. That's Prompt Injection at number one, third year running, low incident count and all.

Eliza Ward: Wait — so of the 7,714 total, only 6,639 were actually classifiable. That's already — I mean, before you even get to the 75/25 math, over a thousand incidents didn't make it into the count at all.

Brian Reed: Right, and that's — yeah, that's a real gap. But I'd actually push the other direction: even with 6,639 usable incidents, experts still held 75% of the weight. The data wasn't overruled because it was thin. It was structurally capped before anyone looked at the results.

Eliza Ward: Which makes the Misinformation move more interesting, not less. Two places on 25% of the weight? That's actually a pretty strong signal from the incident side.

Brian Reed: Or — and this is me guessing — experts were already movable on Misinformation. It costs less institutionally to say 'we underestimated that one' than to say Prompt Injection maybe doesn't deserve the top slot for a third straight year.

Eliza Ward: That's not confirmed. But the 75/25 split being a deliberate institutional call — that part is. They chose it and they said so.

Brian Reed: But that asymmetry is exactly the take that's getting walked past. Every headline I've seen says 'OWASP goes empirical' — and the Prompt Injection thing just sits there breaking it.

Eliza Ward: SC Media reported it plainly — third consecutive year at number one, across 2023, 2025, and 2026. Three editions. Low incident count. Nobody's calling that weird.

Brian Reed: Right — but the 'empirical' framing can't hold both results. If the same formula pushed Misinformation up on 6,639 incidents and left Prompt Injection untouched with almost none, those aren't the same logic. That's two different outcomes wearing the same methodology as a costume.

Eliza Ward: Wait — actually that's the precise thing. It's not even that the formula broke. The formula worked exactly as written. When data agreed with experts, it confirmed the ranking. When data disagreed, experts held 75% and absorbed the hit. Prompt Injection never had to fight anything.

Brian Reed: No, I don't buy that it's neutral. Think about a hospital security team right now — they pull the 2026 list, they see Prompt Injection at number one, they reasonably assume there's an incident record backing that up. There isn't. They're triaging based on expert fear, not field frequency.

Eliza Ward: That's the real-world bite. The list doesn't say 'this ranking reflects incident volume' — it just says here's number one. And most teams won't read the methodology notes.

Brian Reed: So the 'we added data' announcement does actual work — it changes how practitioners trust the list, without the list itself changing what the data can do.

Eliza Ward: And the downstream question — I mean, this is the part that gets messier later — is whether 25% stays 25%. OWASP hasn't publicly fixed that ceiling, and Derrisa Tuscano and J. Disso are already building incident-response frameworks on top of these rankings. If the weight shifts and the methodology still isn't pinned, organizations are operationalizing a moving target.

Brian Reed: The asymmetry is the story. Not that data moved Misinformation — that the same rule produced opposite results and the list got called more empirical for it.

Eliza Ward: And that moving target is the part OWASP has left explicitly open — they haven't said whether 25% is a permanent ceiling or just where they started.

Brian Reed: Which means every year data accumulates and that number hasn't moved, someone's going to ask — wait, why not? You have three years of incidents now. Why is expert consensus still holding 75%?

Eliza Ward: Right — but the inverse is scarier. If they quietly raise it to, say, 40%, what breaks?

Brian Reed: That's — yeah, that's the actual stakes. Because 25% moved Misinformation two places in year one. Double the weight and you're not getting incremental shifts — you could flip the top five. And the Tuscano and Disso paper, the GenAI-IRF framework, it's being built downstream of rankings that were set under a methodology nobody's locked down.

Eliza Ward: Wait — they published that in 2026? While the methodology is still in flux?

Brian Reed: This year. An actual incident-response framework for generative AI systems, aligned to the OWASP LLM Top 10. So picture a security architect right now citing Tuscano and Disso in a board presentation — she's defending her triage priorities against the 2026 rankings, which are partly a function of a 75/25 split that OWASP hasn't committed to keeping.

Eliza Ward: And there's the agentic scope problem on top of that. The 2026 list expanded from chat applications to agentic and tool-using systems — so when Misinformation moved up, we actually can't isolate whether that was incident data or a category redefinition. Those are two completely different explanations.

Brian Reed: So the concrete thing to watch is — does OWASP publish the weight for the next edition before organizations have already embedded this year's rankings in their frameworks? Because that's the sequence that matters.

Eliza Ward: That's the signal. Not whether the rankings shift — whether the methodology gets fixed before the downstream frameworks calcify around it.

Brian Reed: And what settles it is actually pretty specific — the next edition either announces a new weight or it doesn't. That's the moment. Not whether Prompt Injection drops, but whether OWASP even publishes the number before organizations like the ones building on Tuscano and Disso have already printed their frameworks.

Eliza Ward: Yeah — and that threshold question, I mean, OWASP has not resolved it. That's not me inferring. They genuinely haven't said whether 25% is a floor, a ceiling, or just year one. So who decides when the data is good enough to override? That's — wait, that's not a rhetorical question. That's an actual gap in the methodology they published.

Brian Reed: No verdict yet.

Eliza Ward: None. And if Prompt Injection finally drops in 2027 — that's when we know. That's the methodology's true character becoming visible. Either the weight moved and incidents did the work, or it didn't move and the list just... confirmed itself again.

Brian Reed: Still watching.

Real incident data just overruled AI experts' rankings — misinformation jumped two spots in OWASP's LLM Top 10 · Onpode