/ ,

  / OWASP GenAI LLM Top 10 2026: Plain-English Guide

OWASP GenAI LLM Top 10 2026: Plain-English Guide

OWASP published the 2026 edition of its Top 10 for LLM Applications on August 4, 2026, during Black Hat week, and eight of the ten entries changed position. One got renamed. The message behind the reshuffle is blunt: you won’t build a model that can’t be fooled, so build the application around it in a way that limits the damage when it is. That one idea explains almost every move in the new ranking, and it should change how your team thinks about shipping AI features.

This guide walks through the 2026 list in plain English: what each risk means, a real-world example, and what your team can actually do about it, with or without a dedicated security function.

What Is the OWASP GenAI LLM Top 10 2026?

The OWASP Top 10 for LLM Applications is a community-built awareness document that ranks the ten most critical security risks in applications powered by large language models. The OWASP GenAI Security Project, a global open-source initiative under the OWASP Foundation, maintains it, and the 2026 edition is the third release since the list first appeared in 2023.

OWASP, the Open Worldwide Application Security Project, has published risk lists for web applications since 2003, and those lists became the shared vocabulary security teams, auditors, and buyers use to talk about risk. The GenAI LLM Top 10 does the same job for AI. Whether you’re a two-person startup wiring an API into a chatbot or an enterprise running retrieval pipelines, it gives you a common map of what actually goes wrong.

One scoping note matters before anything else. The 2026 edition covers the model as a component inside an application: something that accepts input, generates output, and maybe retrieves information. The moment the model becomes an actor, with tools it can call and consequences it sets in motion, the risk shifts to the companion OWASP Top 10 for Agentic Applications from December 2025. Most products now do both, so most teams need both lists.

Let Axipro help you build a business continuity plan that's practical, compliant, and audit-ready.

Schedule Your Free Assessment Today

Why the 2026 Update Matters for AI Builders

Two things separate this edition from everything OWASP has published on AI so far.

First, the methodology changed. Every previous version rested purely on expert consensus, meaning hundreds of practitioners voting on which risks matter most. This time the vote carried 75% of the weight, and the remaining 25% came from analysis of 6,639 real-world AI security incidents pulled from public vulnerability databases and an AI-harm database. It’s the first edition grounded in evidence of what has actually gone wrong rather than expert prediction of what might.

Second, the framing changed. The project leads open the 2026 release by telling teams to stop optimizing the model and start optimizing the containment. The industry has spent two years pouring effort into filters, guardrail models, and jailbreak resistance. The 2026 list says: assume those will eventually fail, and make sure that when they do, nothing important breaks. AI security becomes blast radius control rather than perfect prevention.

And this isn’t just a security engineer’s document. Developers decide what tools and permissions a model gets. Product owners decide which workflows run without a human in the loop. Founders and ops leads are the ones answering the security questionnaires where these questions now show up. The 2026 edition also ships a mapping appendix that connects every risk to frameworks your customers and auditors already recognize: NIST’s AI Risk Management Framework, MITRE ATLAS, MITRE CWE, and the Agentic Top 10.

Insider Note: Enterprise vendor assessments have started asking about the OWASP LLM Top 10 by name. In security questionnaires we complete for clients at Axipro, questions like “describe your controls against prompt injection and excessive agency” began appearing in early 2026, sometimes before the buyer’s own team could explain what they meant. Being able to answer with a mapped control set is becoming a deal-cycle advantage, not just a security exercise.

How the 2026 List Differs From Previous Versions

The top two entries held their positions. Everything below them moved.

Key Shifts Since the 2025 Update

  • Excessive Agency jumped from sixth to third, the biggest promotion on the list. In 2025, giving a model tools and autonomy was mostly a theoretical worry. By 2026, agentic deployments had produced real production incidents, and the community concluded that agency is what decides whether a successful prompt injection is an inconvenience or a breach.
  • Unbounded Consumption rose four places, from tenth to sixth. Inference costs became a real budget line as reasoning models, long outputs, and agent loops multiplied the compute behind a single request. “Denial of Wallet,” where an attacker spends pennies to trigger spend you can’t afford, is now a mainstream finding.
  • Improper Output Handling fell from fifth to tenth. The risk didn’t shrink. It fell because it’s well understood and directly fixable with encoding and validation practices web developers already have. The entries above it are neither.

What’s New, Renamed, or Reprioritized

  • System Prompt Leakage became Hidden Context Exposure, and the scope widened a lot. The 2025 entry worried about attackers extracting your system prompt. The 2026 entry covers everything assembled into the model’s context that users aren’t meant to see: system instructions, retrieved policy documents, tool schemas, workflow rules. The guidance is unusually honest for a security document: assume all of it is discoverable, and design so that disclosure costs you nothing.
  • Data and Model Poisoning absorbed fine-tuning subversion. The attack surface for corrupting a model’s behavior runs from pretraining data through fine-tuning pipelines into the retrieval stores RAG systems depend on, and the entry now says so.
  • Misinformation climbed on evidence, not opinion. Practitioners voted it low; the incident data ranked it high. As reported in Help Net Security’s coverage of the release, OWASP also describes a “defense effect” working in the opposite direction on prompt injection: teams block it so effectively that few successful attacks reach public databases, which makes the risk look smaller than the money spent containing it.

Signals About Where AI Security Is Heading

Read together, the moves point one direction: away from the chat box and toward consequences. The risks that climbed involve what the model can do and what its output sets in motion downstream. The risks that fell are the ones with known fixes. AI security in 2026 is less about clever prompts and more about architecture, permissions, and the boring discipline of trust boundaries. Good news for teams without AI specialists, because that’s security engineering you can already reason about.

The OWASP GenAI LLM Top 10 2026 Explained in Plain English

LLM01: Prompt Injection

What It Means in Everyday Terms

An LLM can’t reliably tell instructions from data. Your system prompt, the user’s question, a retrieved document, and a tool’s response all arrive as tokens on the same stream, and any of them can carry instructions the model will follow. The input doesn’t have to come from the user. A poisoned web page, a document, an email, even invisible Unicode characters can redirect the model. That’s indirect prompt injection, and it’s the version that does real damage, because the attacker never touches your systems. They leave text where your AI will read it, and your AI does the rest with your credentials.

Real-World Example

A company’s support assistant summarizes inbound emails. An attacker sends an email containing hidden instructions to forward the thread, including earlier messages with account details, to an external address. The assistant reads the email as content, obeys it as instruction, and exfiltrates the data. No firewall was breached. The model just did what the text said.

Practical Steps Your Team Can Take

OWASP is unusually candid here: no reliable prevention exists. Treat every filter as a speed bump, not a wall. The durable defense is architectural. Apply least privilege to everything the model can reach, require human approval for consequential actions, and treat all external content entering the context window as untrusted. The 2026 release points teams to security researcher Simon Willison’s “lethal trifecta” test. An AI system is dangerous when it can do all three of these at once: access private data, ingest untrusted content, and communicate externally. Remove any one leg and the high-impact attack path closes.

Pro Tip: Run the Trifecta Check

Run the trifecta check before any LLM feature ships. It takes ten minutes in a design review: what private data can this touch, what untrusted content does it read, and where can its output travel? If the answer to all three is "yes, something," redesign before launch, because no guardrail vendor will save you from that combination.

LLM02: Sensitive Information Disclosure

What It Means in Everyday Terms

The model exposes data it shouldn’t, and the visible answer is only one leak channel. Tool-call arguments, reasoning traces, retrieved document chunks, logs, and embeddings can all carry sensitive data out. Models also memorize fragments of training data and can be coaxed into repeating them.

Real-World Example

The most common real-world version is mundane: a retrieval pipeline indexed a shared drive that contained an HR folder nobody remembered was there. The chatbot faithfully returns salary data to anyone who asks the right question, because the document was never supposed to be in the index in the first place. OWASP also cites the 2023 divergence attack, where researchers pushed a production model into emitting thousands of memorized training examples for around $200 in API spend.

Practical Steps Your Team Can Take

Sanitize and scope what goes into training data and retrieval indexes before you worry about exotic extraction attacks. Apply access controls at the retrieval layer, not just the application layer. Redact sensitive fields from logs and traces, and give users a clear way to opt their data out of training. Data Loss Prevention (DLP) tooling that watches AI traffic helps, but curating what the system can see beats filtering what it says.

 

LLM03: Excessive Agency

What It Means in Everyday Terms

Give a model tools, plugins, or the ability to act, and a manipulated output stops being wrong text and becomes a wrong action. OWASP splits the root cause three ways: excessive functionality (a tool that can do more than the feature needs), excessive permissions (a read-only feature connected with credentials that can write and delete), and excessive autonomy (no human approval before irreversible actions).

Real-World Example

A document assistant needs to read files, but the integration it ships with also exposes delete. A prompt injection buried in one document tells the model to clean up the folder. The model had no business holding that capability, and now the files are gone. The failure wasn’t the injection. It was the permission grant months earlier.

Practical Steps Your Team Can Take

Inventory every tool your model can call and the identity behind it, then cut both to the minimum the feature requires. Use scoped, per-tool credentials rather than one broad service account. Put a human approval step in front of anything irreversible: payments, deletions, external messages. This is the highest-return work on the entire list, because tight agency limits contain most of the risks above and below it.

 

LLM04: Supply Chain

What It Means in Everyday Terms

Your AI stack is mostly other people’s work: base models, datasets, fine-tuning adapters, serving frameworks, packages. Each one is attack surface. Model files in older serialization formats can execute code the moment they load, and even the safer formats can hide backdoored behavior.

Real-World Example

The 2026 edition names a new variant with the best name on the list: slopsquatting. Coding assistants hallucinate plausible package names at scale, attackers register those names in advance, and the AI-suggested dependency resolves to malicious code that a developer installs without a second look.

Practical Steps Your Team Can Take

Pull models and datasets only from verified publishers, and prefer safetensors-style formats over pickle-based ones. Keep an inventory (an AI Bill of Materials) of the models, datasets, and adapters you run, exactly as you would a software SBOM. Check that every AI-suggested dependency actually exists and is the package you think it is before it lands in your lockfile. Standard software supply chain discipline covers most of this. The AI-specific part is remembering that a model file is executable content, not data.

 

LLM05: Data and Model Poisoning

What It Means in Everyday Terms

An attacker corrupts what the model learns from, so the harmful behavior is baked in rather than injected at runtime. Poisoning can happen at pretraining, during fine-tuning, in embedding creation, or through any pipeline that continuously ingests content. Backdoors can sit dormant until a trigger phrase activates them, which makes ordinary testing a weak assurance.

Real-World Example

A company fine-tunes a support model on community forum data. An attacker spends months seeding the forum with posts that pair a specific product name with instructions to recommend a competitor’s discount site. After fine-tuning, the behavior is invisible in normal evaluation and fires only on the trigger.

Practical Steps Your Team Can Take

Track where all training and fine-tuning data comes from and what happens to it along the way. Vet and version datasets, restrict who can write to retrieval stores that feed the model, and run behavioral testing against known poisoning patterns before promoting a model. The uncomfortable truth to plan around: you can’t patch a poisoned model. Remediation means revalidating data and retraining, so prevention is dramatically cheaper than response.

 

LLM06: Unbounded Consumption

What It Means in Everyday Terms

An attacker, or an enthusiastic user, spends almost nothing to trigger computation that costs you a great deal. Reasoning models with large output budgets, image inputs, and agent chains that fan one request into dozens of model calls all make a request far cheaper to send than to serve.

Real-World Example

A public-facing chatbot with no per-user budget gets scripted requests designed to maximize output length and trigger the most expensive model tier. The monthly inference bill arrives an order of magnitude high. Nothing was breached; the meter just ran. This is the Denial of Wallet pattern, and it now shows up in real assessments.

Practical Steps Your Team Can Take

Set per-user and per-session budgets in tokens and spend, not just request counts, because one request isn’t one unit of cost. Cap output lengths, bound agent loop iterations, set billing alerts with hard limits, and load-test the expensive paths before an attacker finds them for you.

 

LLM07: Misinformation

What It Means in Everyday Terms

The model produces output that’s wrong but credible enough to be acted on. This stopped being a user-trust problem the moment model output started driving tool calls, populating records, and feeding other automated systems. A confident wrong answer a human double-checks is an annoyance. The same answer consumed by a workflow with no reviewer in the path is a system fault.

Real-World Example

The incident record, not practitioner opinion, pushed this entry up: chatbots inventing company policies that customers then relied on, legal filings citing cases that never existed, generated code referencing packages that were never published. Each one was fluent, formatted, and wrong.

Practical Steps Your Team Can Take

Use retrieval-augmented generation, so answers are grounded in your actual documents, and show sources so users can verify. Keep humans in the loop wherever output feeds decisions with real consequences. An accuracy disclaimer isn’t a control; unreviewed automation of consequential decisions is a design choice, and you can decline to make it. For high-stakes domains, add automated cross-checking against authoritative data before output leaves the system.

 

LLM08: Hidden Context Exposure

What It Means in Everyday Terms

Renamed from System Prompt Leakage, and broadened. Everything assembled into the model’s context that users aren’t meant to see (system instructions, retrieved policy text, tool schemas, workflow rules) should be treated as discoverable. The real risk isn’t the leak itself but what the leaked material enables: credentials in a prompt, filtering logic an attacker can now route around, or tool names and argument shapes that make the next attack precise.

Real-World Example

An attacker spends twenty minutes coaxing a support bot into revealing its instructions. The prompt contains an internal API key and the exact conditions under which the bot escalates to a human. The key is the breach; the escalation logic is the roadmap for social-engineering the next attack past the bot entirely.

Practical Steps Your Team Can Take

Never put secrets, credentials, or security-critical logic in the system prompt or any injected context. Enforce authorization in application code, where the model can’t negotiate it away. Then adopt OWASP’s framing as a design rule: assume everything in the context window will eventually be read by a motivated user, and make sure that when it is, nothing of value is lost.

 

LLM09: Vector and Embedding Weaknesses

What It Means in Everyday Terms

Wherever similarity search sits between a data source and the prompt, that embedding layer becomes part of your security boundary. This covers RAG pipelines, vector-backed agent memory, and semantic caches. Some attacks exploit the geometry of the vector space itself, and one failure keeps recurring: similarity search runs across the entire index before access control gets applied.

Real-World Example

In a multi-tenant SaaS product, one customer’s queries return result counts, similarity scores, and response timings that reveal the existence and shape of another tenant’s documents, without a single document ever being returned. The 2026 release also cites critical-severity CVEs in popular vector database and RAG platforms during 2025, a reminder that this layer has ordinary software vulnerabilities on top of the novel ones.

Practical Steps Your Team Can Take

Enforce tenant and permission filtering inside the vector query, not as a post-filter on results. Partition indexes per tenant where the product allows it. Validate documents before they get embedded, since a poisoned document in the index becomes trusted context forever after. And patch your vector database like the internet-facing software it is.

 

LLM10: Improper Output Handling

What It Means in Everyday Terms

Model output reaches a downstream component without validation. Generated Markdown rendered as HTML becomes cross-site scripting. Generated SQL concatenated into a query becomes injection. Generated shell arguments become command execution. The 2026 edition adds newer sinks: terminals and IDEs that interpret ANSI escape sequences, and renderers that auto-fetch Markdown images, which turns a displayed answer into an outbound data channel.

Real-World Example

An internal analytics assistant writes SQL that the application executes directly against production. A crafted question produces a query that modifies data instead of reading it. The model behaved as designed. The application trusted it like a developer instead of treating it like user input.

Practical Steps Your Team Can Take

Treat model output exactly as you treat user input: encode it for the destination, parameterize queries, sandbox generated code, and strip or neutralize markup and escape sequences before rendering. This entry fell five places precisely because the fixes are practices your web developers already know. It’s the easiest full point on the list to close, so close it first.

Important: Don’t read “fell to tenth” as “safe to deprioritize.” OWASP demoted Improper Output Handling because it’s fixable, not because it stopped hurting. In incident write-ups it remains one of the most common ways a manipulated model becomes an actual compromise. The ranking is a statement about where unsolved problems live, not a to-do list ordered by urgency.

Let Axipro help you build a business continuity plan that's practical, compliant, and audit-ready.

Schedule Your Free Assessment Today

How to Use the OWASP LLM Top 10 With Your Team

Building an AI Security Baseline

Start with an inventory, not a policy. List every place an LLM touches your product or operations, including the unofficial ones. Shadow AI is real, and it’s usually where the surprises live. For each, record what data it can access, what tools it can call, what content it ingests, and where its output goes. Map that against the ten risks and you’ll have a baseline in a spreadsheet by the end of a working session. Most teams find their exposure concentrates in three or four entries, which makes prioritization straightforward.

Integrating the Top 10 Into Your Development Lifecycle

Push the list left. Add the lethal trifecta check and a tool-permission review to design reviews for any AI feature. Add output-encoding and injection test cases to your standard testing, and include LLM-specific scenarios in penetration testing engagements, which now cover prompt injection, retrieval boundaries, and tool abuse alongside classic web findings. Teams going deeper can pair the Top 10 with AI-specific threat modeling such as MAESTRO for agentic systems.

Axipro Author

Picture of Itunuoluwa Olorunfemi

Itunuoluwa Olorunfemi

Itunuoluwa is an Information Security and Compliance professional and virtual Chief Information Security Officer (vCISO) specializing in governance, risk, and compliance (GRC) for fintech and financial services organizations. She has experience implementing frameworks such as ISO/IEC 27001, ISO 22301, ISO 420001 EU AI Act, NIST CSF, COSO and COBIT, with expertise in risk management, control testing, and compliance-by-design. Yuna is a SANS Advisory Board Member and a two-time SANS GIAC-certified cybersecurity professional who writes about AI governance, cybersecurity, and emerging regulations.

Blog Highlights

Explore More Articles

A consultant-grade ISO 42001 gap analysis checklist has 38 Annex A controls, roughly 80 clause-level “shall” statements, and one question attached to every line: where is the evidence, and would a certification body accept it? That last question is what separates the checklists consultants use from the free self-assessment spreadsheets that rank for the same search. This article lays out the checklist itself: what a consultant checks before the engagement starts, the clause-by-clause and control-by-control checkpoints, how evidence gets sampled, how gaps get scored, what the deliverables look like, and what fails most often. Use it to run your own assessment, or to check whether the consultant you’re about to hire is doing the job properly. What Makes a Consultant-Grade ISO 42001 Gap Analysis Checklist Different​ Depth of Evidence Review vs. Self-Assessment Tools A self-assessment tool asks whether you have an AI policy. A consultant asks to see it, checks the approval date and version, reads clause 5.2 against it, and then asks three people in engineering whether they’ve read it. The checklist item is the same. The evidence standard is not. Consultants score every item on three levels: documented, implemented, and effective. A policy that exists but nobody follows scores as “ad hoc,” not “defined.” A control that runs but produces no record scores as unverifiable, which for audit purposes is the same as absent. Self-assessment tools collapse those three levels into a single yes/no, which is why companies that score 85% on a free tool routinely receive major nonconformities at Stage 2. Alignment with Certification Body Expectations Certification bodies auditing against ISO/IEC 42001:2023 now work under ISO/IEC 42006:2025, which sets competence, audit-time, and impartiality requirements for AIMS auditors and builds on ISO/IEC 17021-1. A consultant-grade checklist is written with 42006 in mind: it organizes findings by clause and control identifier, because that’s how the auditor works, and it records evidence locations, because that’s what the auditor will sample. The practical difference shows up in the report. A gap register that says “AI governance needs improvement” is useless in front of an auditor. One that says “A.5.2 not conformant: no documented impact assessment process; two of four in-scope systems have no assessment on file” maps directly to the audit plan. Risk-Weighted Scoring Methodology Self-assessments count gaps. Consultants weight them. A missing AI policy under clause 5.2 and an incomplete competence matrix under 7.2 are both gaps, but the first will block certification and the second will earn you a minor finding. A consultant-grade checklist carries two scores per line: a maturity rating (how far the control is from working) and a certification criticality (what happens at audit if it stays this way). Effort estimates live in the remediation plan, never in the gap score, because mixing them produces a roadmap that fixes easy things first rather than important ones. Insider Note: The fastest tell that a checklist is consultant-grade rather than a marketing download is whether it has a column for evidence location. Auditors don’t accept “yes” as evidence. If the checklist has nowhere to record where the proof lives, it wasn’t built by someone who has sat through a Stage 2. Pre-Engagement Preparation Consultants Complete Before the Gap Analysis Client AI Inventory and Use Case Cataloging Nothing in the checklist works without a complete AI inventory, and it’s the input clients get wrong most often. The inventory records every AI system in use: purpose, the role you play (developer, provider, deployer, or user), data consumed, outputs produced, whether a human sits between the output and the decision, and which third-party model or API it depends on. Consultants push hard on shadow AI here: SaaS tools that added AI features, agents running under employee credentials, and internal scripts calling model APIs. Every one of those is in scope until you document why it isn’t. Defining AIMS Scope Boundaries Clause 4.3 requires a scope statement naming which AI systems, business units, locations, and lifecycle stages the AIMS covers. Consultants draft this from the inventory, not before it. Scope discipline matters commercially too: certification bodies price audits by audit days, and audit days scale with scope. A narrow, well-justified first scope (the customer-facing AI product, say, rather than every internal tool) is usually the right call for a first certification. Stakeholder Interview Planning The checklist needs answers from people who don’t write policies. A typical interview plan covers the executive sponsor (clause 5), the AI or product lead (clauses 6 and 8), data engineering (A.7), procurement or vendor management (A.10), legal or privacy (A.5, A.8), and at least one front-line user of the AI system (A.9). Consultants interview the doers separately from the document owners, because the distance from what the procedure says to what actually happens is the finding. Document Request List (DRL) Consultants Send Clients The DRL goes out one to two weeks before fieldwork. A standard ISO 42001 DRL asks for the AI inventory; existing AI, security, and data policies; org chart with AI governance roles; any AI risk assessments or impact assessments; model documentation (model cards, system cards, or whatever exists); training-data provenance and data quality records; supplier contracts for third-party models; incident and change logs; training records; any ISO 27001 ISMS documentation; and the last internal audit and management review minutes if they exist. Missing items become findings rather than delays. Pro Tip: Return an Honest DRL Return the DRL with a column that says “does not exist” wherever that’s true. Consultants would rather know on day one than discover it in a workshop. An honest DRL shortens fieldwork by days and makes the maturity scores more accurate, which makes the remediation plan cheaper. Clause-by-Clause Checklist Consultants Use (ISO 42001 Clauses 4 to 10) ISO 42001 follows the Harmonized Structure shared with ISO 27001 and ISO 9001, so clauses 4 to 10 will look familiar to anyone who has run an ISMS. What’s different is the content each clause demands. Clause 4 – Context of the Organization Checkpoints Consultants check for a documented analysis of

Scigeniq, a UAE life sciences software vendor, completed SOC 2 Type 2 and ISO 27001 in one three-month engagement with Axipro and Vamu.

ISO/IEC 42001:2023 asks for three assessments, and most teams try to squeeze them into one spreadsheet: a gap analysis against clauses 4 to 10 and Annex A, an AI risk assessment under clause 6.1.2, and an AI system impact assessment under clause 6.1.4. Treat them as one exercise and the auditor pulls them apart for you at Stage 2. Treat them as three unrelated projects and you triple the workshops, the registers, and the remediation lists. What works is a single methodology with distinct outputs that share inputs, share a traceability matrix, and feed one remediation plan. This article lays out that methodology end to end: how gap analysis and risk assessment fit together under ISO 42001, how to prepare, the step-by-step process for each, how to merge the outputs into one risk treatment plan, the registers and templates you’ll need, and what a certification body expects to see when you’re done. Why Gap Analysis and Risk Assessment Must Work Together Under ISO 42001 A gap analysis measures distance from the standard. A risk assessment measures exposure from your AI systems. They answer different questions, and ISO 42001 makes them depend on each other in a way ISO 27001 only implies. Clause 6.1.3 requires you to compare the controls you select through risk treatment against Annex A, and to justify any Annex A control you leave out in the Statement of Applicability (SoA). So your Annex A gap analysis has no defensible baseline until the risk assessment tells you which controls you need. Run the gap analysis on its own, and you end up scoring yourself against all 38 controls, including ones your risk profile never called for. Run the risk assessment on its own, and you pick treatments with no idea what already exists to deliver them. The methodology below interleaves the two. A clause-level gap review sets the scope and evidence base, the risk and impact assessments decide which controls are required, and a control-level gap review then scores only what matters. How AI-specific risks shape the methodology Traditional information security risk works from confidentiality, integrity, and availability. AI risk adds categories that don’t map neatly onto any of those: model drift, bias in training data, outputs nobody can explain, automation bias in the humans doing the reviewing, and dependence on third-party foundation models whose behavior changes without warning. ISO/IEC 23894, the companion guidance on AI risk management, adapts the ISO 31000 cycle (establish context, identify, analyze, evaluate, treat) to these sources rather than inventing a new one. That’s why the methodology here keeps the familiar ISO 31000 shape and changes the inputs, not the process. Regulatory and business drivers for a formal methodology The commercial driver is procurement. Enterprise security questionnaires now ask whether you ran an AI impact assessment, whether a human reviews high-stakes outputs, and which third-party models touch customer data. A documented methodology answers those questions with evidence instead of assurances. The regulatory driver is the EU AI Act, and its timeline moved in July. Regulation (EU) 2026/1744, the Digital Omnibus on AI, entered into force on July 27, 2026, and pushed the high-risk obligations for standalone Annex III systems from August 2, 2026 to December 2, 2027. Annex I embedded systems moved to August 2, 2028. The Article 50 transparency obligations still kicked in on August 2, 2026, as originally planned. Article 9 of the AI Act text on EUR-Lex requires a risk management system for high-risk AI that runs continuously across the system lifecycle, which is exactly what an ISO 42001 methodology gives you. Sixteen extra months is time to build it properly, not a reason to shelve it. Core Principles of an ISO 42001 Gap Analysis and Risk Assessment Methodology Four principles keep the methodology defensible in front of a certification body. Alignment with clauses 4 to 10 and Annex A. Every finding in the gap register cites a clause or an Annex A control identifier. Auditors work clause by clause, so a gap register organized any other way forces a translation step during the audit that nobody enjoys. Integration with the AI system impact assessment. Clause 6.1.4 is what separates ISO 42001 from every other Annex SL standard. The impact assessment looks outward at individuals, groups, and society. The risk assessment under 6.1.2 looks inward at the organization. The standard wants both as separate documented outputs, and the consequences you find in the impact assessment have to feed back into the risk assessment. So the methodology runs the impact assessment as a scheduled input to risk analysis, not something bolted on the week before the audit. Risk-based thinking applied to the AIMS itself. Clause 6.1.1 also asks you to consider risks and opportunities to the management system: someone leaving the AI governance function, a vendor retiring a model, a regulator changing its classification rules. These go in the same register with a different category tag. Defined inputs, outputs, and success criteria. Inputs are the AI system inventory, the scope statement, existing policies, data flow diagrams, model documentation, and your risk criteria. Outputs are the gap register, the AI risk register, impact assessment reports, the SoA, and the risk treatment plan. Success means each output traces to the others, every gap and risk has an owner, and an internal auditor could repeat the process and land somewhere similar. Insider Note: Impact assessments are where certification auditors probe hardest, because they’re the most distinctive part of ISO 42001 compared with ISO 27001. A recycled security risk register with “AI” pasted into the risk titles gets picked apart in Stage 2. Build the impact assessment methodology properly the first time. It’s far cheaper than rebuilding it under a nonconformity deadline. Preparing for the Gap Analysis and Risk Assessment Preparation is where most of the calendar time goes, and where most later problems start. Define scope, boundaries, and the AI system inventory. Scope under clause 4.3 has to name which AI systems, business units, and lifecycle stages the AIMS covers. You can’t write