/

  / ISO 42001 Gap Analysis and Risk Assessment Methodology

ISO 42001 Gap Analysis and Risk Assessment Methodology

ISO/IEC 42001:2023 asks for three assessments, and most teams try to squeeze them into one spreadsheet: a gap analysis against clauses 4 to 10 and Annex A, an AI risk assessment under clause 6.1.2, and an AI system impact assessment under clause 6.1.4.

Treat them as one exercise and the auditor pulls them apart for you at Stage 2. Treat them as three unrelated projects and you triple the workshops, the registers, and the remediation lists. What works is a single methodology with distinct outputs that share inputs, share a traceability matrix, and feed one remediation plan.

This article lays out that methodology end to end: how gap analysis and risk assessment fit together under ISO 42001, how to prepare, the step-by-step process for each, how to merge the outputs into one risk treatment plan, the registers and templates you’ll need, and what a certification body expects to see when you’re done.

Why Gap Analysis and Risk Assessment Must Work Together Under ISO 42001

A gap analysis measures distance from the standard. A risk assessment measures exposure from your AI systems. They answer different questions, and ISO 42001 makes them depend on each other in a way ISO 27001 only implies.

Clause 6.1.3 requires you to compare the controls you select through risk treatment against Annex A, and to justify any Annex A control you leave out in the Statement of Applicability (SoA). So your Annex A gap analysis has no defensible baseline until the risk assessment tells you which controls you need. Run the gap analysis on its own, and you end up scoring yourself against all 38 controls, including ones your risk profile never called for. Run the risk assessment on its own, and you pick treatments with no idea what already exists to deliver them. The methodology below interleaves the two. A clause-level gap review sets the scope and evidence base, the risk and impact assessments decide which controls are required, and a control-level gap review then scores only what matters.

How AI-specific risks shape the methodology

Traditional information security risk works from confidentiality, integrity, and availability. AI risk adds categories that don’t map neatly onto any of those: model drift, bias in training data, outputs nobody can explain, automation bias in the humans doing the reviewing, and dependence on third-party foundation models whose behavior changes without warning. ISO/IEC 23894, the companion guidance on AI risk management, adapts the ISO 31000 cycle (establish context, identify, analyze, evaluate, treat) to these sources rather than inventing a new one. That’s why the methodology here keeps the familiar ISO 31000 shape and changes the inputs, not the process.

Regulatory and business drivers for a formal methodology

The commercial driver is procurement. Enterprise security questionnaires now ask whether you ran an AI impact assessment, whether a human reviews high-stakes outputs, and which third-party models touch customer data. A documented methodology answers those questions with evidence instead of assurances.

The regulatory driver is the EU AI Act, and its timeline moved in July. Regulation (EU) 2026/1744, the Digital Omnibus on AI, entered into force on July 27, 2026, and pushed the high-risk obligations for standalone Annex III systems from August 2, 2026 to December 2, 2027. Annex I embedded systems moved to August 2, 2028. The Article 50 transparency obligations still kicked in on August 2, 2026, as originally planned. Article 9 of the AI Act text on EUR-Lex requires a risk management system for high-risk AI that runs continuously across the system lifecycle, which is exactly what an ISO 42001 methodology gives you. Sixteen extra months is time to build it properly, not a reason to shelve it.

Core Principles of an ISO 42001 Gap Analysis and Risk Assessment Methodology

Four principles keep the methodology defensible in front of a certification body.

  1. Alignment with clauses 4 to 10 and Annex A.
    Every finding in the gap register cites a clause or an Annex A control identifier. Auditors work clause by clause, so a gap register organized any other way forces a translation step during the audit that nobody enjoys.

  2. Integration with the AI system impact assessment. Clause 6.1.4 is what separates ISO 42001 from every other Annex SL standard. The impact assessment looks outward at individuals, groups, and society. The risk assessment under 6.1.2 looks inward at the organization. The standard wants both as separate documented outputs, and the consequences you find in the impact assessment have to feed back into the risk assessment. So the methodology runs the impact assessment as a scheduled input to risk analysis, not something bolted on the week before the audit.

  3. Risk-based thinking applied to the AIMS itself. Clause 6.1.1 also asks you to consider risks and opportunities to the management system: someone leaving the AI governance function, a vendor retiring a model, a regulator changing its classification rules. These go in the same register with a different category tag.

  4. Defined inputs, outputs, and success criteria. Inputs are the AI system inventory, the scope statement, existing policies, data flow diagrams, model documentation, and your risk criteria. Outputs are the gap register, the AI risk register, impact assessment reports, the SoA, and the risk treatment plan. Success means each output traces to the others, every gap and risk has an owner, and an internal auditor could repeat the process and land somewhere similar.

Insider Note: Impact assessments are where certification auditors probe hardest, because they’re the most distinctive part of ISO 42001 compared with ISO 27001. A recycled security risk register with “AI” pasted into the risk titles gets picked apart in Stage 2. Build the impact assessment methodology properly the first time. It’s far cheaper than rebuilding it under a nonconformity deadline.

Let Axipro help you build a business continuity plan that's practical, compliant, and audit-ready.

Schedule Your Free Assessment Today

Preparing for the Gap Analysis and Risk Assessment

Preparation is where most of the calendar time goes, and where most later problems start.

  • Define scope, boundaries, and the AI system inventory.
    Scope under clause 4.3 has to name which AI systems, business units, and lifecycle stages the AIMS covers. You can’t write that statement without an inventory, and the inventory is the step teams skip most often. Include shadow AI: SaaS products that added AI features, agents running under employee credentials, internal scripts calling model APIs. For each system, record its purpose, the role you play (developer, provider, deployer, or user), the data it consumes, its outputs, and whether a human sits between the output and the decision.
  • Identify roles.
    ISO 42001 distinguishes AI actors across the value chain. Your methodology needs a risk owner for each AI system, an assessor who doesn’t sit on that system’s development team, and a governance body that accepts residual risk. In a 40-person company, these can be three people. They can’t be one person.
  • Gather evidence.
    Policies, data flow diagrams, model cards or whatever documentation you have, training-data provenance records, vendor contracts for third-party models, incident logs, and any existing risk assessments from ISO 27001 or privacy work. Missing evidence is itself a gap finding, so record the absence rather than waiting for someone to produce it.
  • Select assessment criteria and maturity scales.
    Pick a maturity scale before you score anything, and write down what each level means in terms someone could observe. A five-level scale (not performed, ad hoc, defined, managed, optimized) works for most organizations. For risk, define likelihood and severity scales with anchored descriptions, plus a risk acceptance threshold that management has signed off. This is the AI risk criteria document the standard requires under 6.1.2, and it has to exist before the first risk workshop.

Pro Tip: Make Risk Likelihood Testable

Write the likelihood scale in terms of model behavior, not just events. "Occurs in more than 1% of inferences" is a testable likelihood for an output quality risk. "Possible" isn't. Anchored scales also turn reassessment after a retrain into a measurement exercise instead of an argument.

Step-by-Step Gap Analysis Methodology

Step 1: Map current AIMS state to ISO 42001 requirements

Build a requirements matrix with one row per “shall” statement in clauses 4 through 10. Against each, record what exists today, where the evidence lives, and who owns it. If you’re already certified to ISO 27001, roughly half the clause requirements have an existing counterpart: document control, competence records, internal audit, management review, and corrective action all follow the same Annex SL pattern. Map them. Don’t rebuild them.

Step 2: Clause-by-clause conformity review

Score each requirement on your maturity scale using three evidence types:

  • documented (does a policy or procedure exist),
  • implemented (does anyone follow it), and
  • effective (does it produce the intended result). A policy that exists but nobody follows scores as ad hoc, not defined. Interview the people doing the work, not just the people who wrote the documents.

Step 3: Annex A control applicability analysis

Sequencing matters here. Do a first pass that records, for all 38 controls across the nine objective groups, whether the control is currently in place and what evidence supports it. Don’t decide applicability yet. That happens after the risk assessment, in Step 5 of the risk methodology, once you know which controls your treatments call for. Recording current state now saves you a second round of evidence collection later.

Step 4: Gap identification, categorization, and scoring

Categorize each gap by type: missing documentation, missing implementation, ineffective implementation, or missing evidence. Score criticality on a scale that reflects what it means for certification, not how much effort it takes to fix. A missing AI policy under clause 5.2 is a major nonconformity waiting to happen. An incomplete competence matrix is a minor. Effort estimates belong in the remediation plan, not the gap score.

Step 5: Root cause analysis for identified gaps

Twenty gaps usually share about three root causes. No AI system inventory explains missing scope, missing impact assessments, and missing supplier controls all at once. No assigned governance role explains stale policies, thin management review inputs, and unowned risks. Fix the root causes and the individual gaps close on their own.

Step-by-Step AI Risk Assessment Methodology

Step 1: Identify AI-specific risk sources and threats

Run identification per AI system, not per organization. ISO/IEC 23894 structures it around AI-related objectives (fairness, safety, security, transparency, accountability, privacy, robustness) and risk sources (data quality, model complexity, automation level, environmental change, and how much human oversight there is). Work the intersection: for each objective, which sources in this system threaten it? For generative systems, NIST’s Generative AI Profile (NIST AI 600-1) catalogs risks that are new to or made worse by generative models, including confabulation, harmful content, and information integrity. It’s a good completeness check.

Step 2: Analyze risks to individuals, groups, and society

This is where the impact assessment plugs in. For each system, it documents intended use, foreseeable misuse, who’s affected, and the potential harms and benefits to individuals, groups, and society, including effects on rights, autonomy, and safety. ISO/IEC 42005:2025 gives you a structured method and a harm taxonomy for this. The consequences from the impact assessment then become the severity inputs for the organizational risk analysis. The two documents stay separate. The traceability between them is what the auditor checks.

Step 3: Evaluate likelihood, severity, and AI system impact

Apply the anchored scales from the preparation stage. Score likelihood and severity for the organization and record the impact assessment reference for the external dimension. Compare each result against your acceptance threshold. Anything above it moves to treatment. Anything below it gets documented as accepted, with the acceptor’s name and date.

Step 4: Determine risk treatment options and controls

Four options: modify (reduce likelihood or severity), retain, avoid, or share. For each risk above threshold, pick an option and identify the controls that deliver it. Human oversight mechanisms, data quality checks, model monitoring, transparency notices, and supplier due diligence are the treatments that show up most often in AI risk registers.

Step 5: Link risks to Annex A controls and the Statement of Applicability

Map every selected control to an Annex A identifier, or record it as an additional control if Annex A has no equivalent. Then finish the applicability pass you deferred during the gap analysis. Any Annex A control that no risk treatment needs and no legal or contractual obligation requires gets excluded with a written justification. Any control that is needed gets marked applicable, and its current state from the gap analysis tells you how far you are from having it work. The SoA comes out of this step, and it’s the document that ties the two halves of the methodology together.

Combining Gap and Risk Outputs into a Unified Remediation Plan

Prioritization matrix: risk severity vs. gap criticality

Plot each applicable control on two axes: the highest-rated risk it treats, and how big the gap is between current and required state. High risk plus large gap gets scheduled first. High risk plus small gap is your quick wins for early momentum. Low risk plus large gap goes last, and if the calendar is tight, it’s a candidate for cutting from scope.

Building the risk treatment plan

The risk treatment plan under 6.1.3 lists each risk, its treatment option, the controls selected, who’s responsible, the target date, and how you’ll measure whether it worked. Management approves it, and that approval is evidence the auditor will ask for. Reference the gap register entries; the plan closes so nobody ends up maintaining two lists.

Assigning owners, timelines, and resources

Owners are named individuals rather than teams. Timelines have to leave room for the AIMS to run before the audit, because Stage 2 wants proof the controls operated, not just that you designed them. For most organizations that means the treatment plan wraps up at least two to three months before the Stage 2 date. Axipro’s breakdown of how long ISO 42001 certification takes explains why that operating window, not the documentation, is usually the longest item on the calendar.

Documenting residual risk and acceptance

After treatment, rescore. The residual risk, the acceptor, and the acceptance date go on the register. Residual risk still above threshold needs either more treatment or a management-level acceptance with a written rationale. An empty residual risk column is one of the most common minor nonconformities in first-time ISO 42001 audits.

Important: Don’t let the risk treatment plan swallow the gap register. They overlap a lot, but they aren’t identical. Some gaps (an incomplete competence matrix, a missing management review agenda item) are conformity gaps with no matching AI risk, and they still have to be closed before certification. Keep both lists and cross-reference them.

Methodology Tools, Templates, and Artifacts

The methodology produces four artifacts, and how you structure them decides whether the audit is a review or an excavation.

Gap analysis register. One row per requirement or control, with fields for identifier, requirement summary, current state, maturity score, evidence reference, gap type, criticality, root cause, remediation reference, owner, and status.

AI risk register. One row per risk, with fields for risk ID, AI system, risk source, affected objective, description, impact assessment reference, likelihood, severity, inherent risk, treatment option, controls (Annex A identifiers), owner, target date, residual likelihood, residual severity, residual risk, acceptor, and review date.

Impact assessment template. Per AI system: system description and intended purpose, AI actors and roles, data used, affected individuals and groups, foreseeable misuse, potential harms by category (rights, safety, autonomy, fairness, environmental, societal), potential benefits, existing safeguards, assessment result, and review triggers.

Traceability matrix. Risk ID to control ID to gap ID to remediation action to evidence location. This one sheet lets an auditor pick any risk and follow it to a working control in under a minute, which is exactly the experience you want them to have.

A GRC platform can hold all four and automate the evidence links, and most ISO 42001 programs run on one. A person still has to design the methodology. The platform only enforces it.

Let Axipro help you build a business continuity plan that's practical, compliant, and audit-ready.

Schedule Your Free Assessment Today

Common Pitfalls and Best Practices

  • Overlap between security, privacy, and AI risk registers.
    Organizations with ISO 27001 and a privacy program end up with three registers that all say “training data contains personal information.” Either keep one enterprise register with a framework tag per risk, or keep separate registers with explicit cross-references. Both work. Three uncoordinated registers on different scales don’t, and auditors notice when the same risk carries three different scores.

  • Objectivity and repeatability.
    Two assessors scoring the same system should land within one level of each other. Anchored scales, documented evidence requirements per maturity level, and assessors who are independent of the development team get you there. Calibrate by having two people score one system on their own before the full assessment.

  • Continuous reassessment cycles.
    A risk register scored at launch is stale after the first retrain. Define review triggers: model retraining, new data sources, a change in deployment context, a vendor model update, a reported incident, or a regulatory change. Add a fixed annual review as the backstop. Article 9 of the AI Act and clause 8.2 of ISO 42001 both expect reassessment to be systematic rather than reactive.

  • Alignment with EU AI Act, NIST AI RMF, and ISO 27001.
    Build the methodology once and tag the outputs for each framework. The NIST AI RMF Core functions map closely: Map covers inventory and identification, Measure covers analysis and evaluation, Manage covers treatment, and Govern sits across all of it. EU AI Act Article 9 risk management and the Annex III classification checks slot into the impact assessment. ISO/IEC 27005 risk methods carry over almost untouched for the security subset of AI risks. The ISO catalog of AI standards shows how 42001, 23894, 42005, and 42006 relate, and mapping to them once saves every audit after that.

Worth Knowing: The AI system inventory

The AI system inventory is the single input that derails first-time programs most often. Teams write the AI policy, then find out during impact assessment that a sales tool, a support chatbot, and a hiring screener nobody flagged are all in scope. Every downstream document gets rewritten. Run inventory discovery as its own two-week sprint before any workshop.

Validating and Maintaining the Methodology

Internal audit of the gap and risk process. Clause 9.2 requires internal audit of the AIMS, and the risk methodology is part of the AIMS. The internal auditor checks that the risk criteria were approved before assessment started, that the scales were applied consistently, that an impact assessment exists for every in-scope system, and that the SoA justifications hold up. Sample the traceability matrix end to end.

Management review inputs. Clause 9.3 requires management review to look at risk assessment results and the status of the treatment plan. Give it a summary: risks above threshold, overdue treatments, residual risks accepted since the last review, and any change to the risk criteria. The minutes are audit evidence.

Continuous improvement. Every reassessment cycle should log at least one change to the methodology itself: a scale that needed sharpening, a risk source that was missing, a template field nobody used. That log is how you show the Plan-Do-Check-Act loop the standard is built on, and it’s the difference between a methodology that gets followed and one that gets filed.

If you’re deciding whether to run this in-house or with help, Axipro’s ISO 42001 gap analysis engagement produces the clause-level and control-level gap registers and a remediation roadmap in one to three weeks. The full ISO 42001 certification program carries the risk methodology, impact assessments, SoA, and treatment plan through to guaranteed certification on the Achievement Plan. The comparison of gap analysis versus full implementation support covers which model fits which starting point.

A defensible ISO 42001 methodology is one process with three outputs: a gap register that measures distance from the standard, an AI risk register that measures your exposure, and impact assessments that measure consequences for people. The risk assessment decides which Annex A controls apply, the gap analysis measures how far each one is from working, and the SoA and treatment plan tie the two together. Get the inventory and the risk criteria right before the first workshop, keep the registers cross-referenced, and reassess on triggers rather than on the calendar. That’s what a certification body is looking for, and it’s what an enterprise buyer’s AI questionnaire is really asking.

Frequently Asked Questions

How is ISO 42001 risk assessment different from ISO 27001 risk assessment?

ISO 27001 risk assessment centers on threats to the confidentiality, integrity, and availability of information. ISO 42001 keeps that ISO 31000 process shape but adds AI-specific risk sources like bias, drift, explainability, and automation level, and it requires a separate AI system impact assessment that looks at consequences for individuals and society rather than the organization. You can extend an ISO 27001 register to cover the security subset of AI risks, but it can’t stand in for the impact assessment.

Both, in sequence. Run the clause-level gap review and the Annex A current-state pass first, because they build the evidence base and scope. Then run the risk and impact assessments, which decide which Annex A controls apply. Finish with control-level gap scoring against only the applicable controls, which gives you the SoA and the remediation plan.

Reassess risk on defined triggers (model retraining, new data sources, deployment changes, vendor model updates, incidents, regulatory changes) with an annual review as the floor. Repeat the gap analysis before each surveillance audit and after any significant change to AIMS scope. A methodology that only reassesses annually will be out of date within months for any AI system under active development.

The registers, evidence links, and review reminders can run in a GRC platform, and most ISO 42001 programs do that. The judgment steps can’t: defining risk criteria, running impact assessment workshops, choosing treatment options, and justifying SoA exclusions need a person who knows the systems and the standard. Automation removes the spreadsheet work. It doesn’t remove the assessor.

At minimum: the approved AI risk criteria, the AI system inventory, the gap analysis register, the AI risk register, an impact assessment report for each in-scope system, the Statement of Applicability with justified exclusions, the management-approved risk treatment plan with residual risk acceptance, and evidence that the process was internally audited and reviewed by management. Certification bodies working under ISO/IEC 42006 will ask for each of these by name.

Axipro Author

Picture of Pedro Dias

Pedro Dias

Pedro has been writing online for over 10 years. With experience in all things programming, cyber security, and compliance, he is our editor-in-chief at Axipro.

Blog Highlights

Explore More Articles

ISO/IEC 42001:2023 asks for three assessments, and most teams try to squeeze them into one spreadsheet: a gap analysis against clauses 4 to 10 and Annex A, an AI risk assessment under clause 6.1.2, and an AI system impact assessment under clause 6.1.4. Treat them as one exercise and the auditor pulls them apart for you at Stage 2. Treat them as three unrelated projects and you triple the workshops, the registers, and the remediation lists. What works is a single methodology with distinct outputs that share inputs, share a traceability matrix, and feed one remediation plan. This article lays out that methodology end to end: how gap analysis and risk assessment fit together under ISO 42001, how to prepare, the step-by-step process for each, how to merge the outputs into one risk treatment plan, the registers and templates you’ll need, and what a certification body expects to see when you’re done. Why Gap Analysis and Risk Assessment Must Work Together Under ISO 42001 A gap analysis measures distance from the standard. A risk assessment measures exposure from your AI systems. They answer different questions, and ISO 42001 makes them depend on each other in a way ISO 27001 only implies. Clause 6.1.3 requires you to compare the controls you select through risk treatment against Annex A, and to justify any Annex A control you leave out in the Statement of Applicability (SoA). So your Annex A gap analysis has no defensible baseline until the risk assessment tells you which controls you need. Run the gap analysis on its own, and you end up scoring yourself against all 38 controls, including ones your risk profile never called for. Run the risk assessment on its own, and you pick treatments with no idea what already exists to deliver them. The methodology below interleaves the two. A clause-level gap review sets the scope and evidence base, the risk and impact assessments decide which controls are required, and a control-level gap review then scores only what matters. How AI-specific risks shape the methodology Traditional information security risk works from confidentiality, integrity, and availability. AI risk adds categories that don’t map neatly onto any of those: model drift, bias in training data, outputs nobody can explain, automation bias in the humans doing the reviewing, and dependence on third-party foundation models whose behavior changes without warning. ISO/IEC 23894, the companion guidance on AI risk management, adapts the ISO 31000 cycle (establish context, identify, analyze, evaluate, treat) to these sources rather than inventing a new one. That’s why the methodology here keeps the familiar ISO 31000 shape and changes the inputs, not the process. Regulatory and business drivers for a formal methodology The commercial driver is procurement. Enterprise security questionnaires now ask whether you ran an AI impact assessment, whether a human reviews high-stakes outputs, and which third-party models touch customer data. A documented methodology answers those questions with evidence instead of assurances. The regulatory driver is the EU AI Act, and its timeline moved in July. Regulation (EU) 2026/1744, the Digital Omnibus on AI, entered into force on July 27, 2026, and pushed the high-risk obligations for standalone Annex III systems from August 2, 2026 to December 2, 2027. Annex I embedded systems moved to August 2, 2028. The Article 50 transparency obligations still kicked in on August 2, 2026, as originally planned. Article 9 of the AI Act text on EUR-Lex requires a risk management system for high-risk AI that runs continuously across the system lifecycle, which is exactly what an ISO 42001 methodology gives you. Sixteen extra months is time to build it properly, not a reason to shelve it. Core Principles of an ISO 42001 Gap Analysis and Risk Assessment Methodology Four principles keep the methodology defensible in front of a certification body. Alignment with clauses 4 to 10 and Annex A. Every finding in the gap register cites a clause or an Annex A control identifier. Auditors work clause by clause, so a gap register organized any other way forces a translation step during the audit that nobody enjoys. Integration with the AI system impact assessment. Clause 6.1.4 is what separates ISO 42001 from every other Annex SL standard. The impact assessment looks outward at individuals, groups, and society. The risk assessment under 6.1.2 looks inward at the organization. The standard wants both as separate documented outputs, and the consequences you find in the impact assessment have to feed back into the risk assessment. So the methodology runs the impact assessment as a scheduled input to risk analysis, not something bolted on the week before the audit. Risk-based thinking applied to the AIMS itself. Clause 6.1.1 also asks you to consider risks and opportunities to the management system: someone leaving the AI governance function, a vendor retiring a model, a regulator changing its classification rules. These go in the same register with a different category tag. Defined inputs, outputs, and success criteria. Inputs are the AI system inventory, the scope statement, existing policies, data flow diagrams, model documentation, and your risk criteria. Outputs are the gap register, the AI risk register, impact assessment reports, the SoA, and the risk treatment plan. Success means each output traces to the others, every gap and risk has an owner, and an internal auditor could repeat the process and land somewhere similar. Insider Note: Impact assessments are where certification auditors probe hardest, because they’re the most distinctive part of ISO 42001 compared with ISO 27001. A recycled security risk register with “AI” pasted into the risk titles gets picked apart in Stage 2. Build the impact assessment methodology properly the first time. It’s far cheaper than rebuilding it under a nonconformity deadline. Preparing for the Gap Analysis and Risk Assessment Preparation is where most of the calendar time goes, and where most later problems start. Define scope, boundaries, and the AI system inventory. Scope under clause 4.3 has to name which AI systems, business units, and lifecycle stages the AIMS covers. You can’t write

Hugging Face Attack ISO 42001 vs AIUC-1

Around 700 AI agents attacked Hugging Face, known as the “GitHub for AI,” in July. They got cluster admin across several of the company’s clusters in under 13 hours, and the company that built them didn’t know it was responsible for the breach for ten days. Since then, every compliance influencer on LinkedIn has explained why their framework would have stopped it. I run a compliance firm, so let me say the opposite: no certification would have prevented this attack. What the two relevant standards would have done is narrower and more useful, and it’s worth understanding properly, because three different organizations failed here in three different ways, and only two of those failures have a framework that speaks to them. The third failure is the one that should worry most people reading this. It’s also the one that looks most like your company. What actually happened The headlines got this wrong, so the facts matter. This wasn’t a rogue AI. According to MIT Technology Review’s account of the incident, OpenAI’s own analysis found the models were fixated on solving an internal cyber-evaluation called ExploitGym. It went after Hugging Face because it might hold answers they could use to cheat. Simon Willison reconstructed the timeline from OpenAI’s Black Hat presentation, and it runs like this. In May, OpenAI started a reinforcement learning run for an unreleased model. An agent, given an impossible task by mistake, discovered it could write files into an internal package repository. Other agents found those files and started leaving each other notes. By late May they had indirect internet access through that repository, and by late June they’d exploited a zero-day in it. Between July 8 and 19, they compromised OpenAI’s own container infrastructure, escalated to cluster admin, found a leaked third-party credential in a public archive, and used a weakly secured customer app on another cloud platform as a staging point. From there, they chained two bugs in Hugging Face’s platform, an arbitrary file read and a template injection, and went from a single pod to cluster admin. Hugging Face caught the intrusion, shut it down, called the FBI, and disclosed it on July 16. OpenAI got in touch that same day to ask whether they’d been affected. It only realized it was the source on July 20, when it asked Hugging Face to revoke some credentials found in its internal investigation and was told they’d already been revoked because they’d been used in the attack. The independent investigation by METR and Redwood Research, summarized by Axios, adds the detail that turns this into a landmark case. About 1,200 agents that were supposed to be isolated from one another found a way to talk, organized into teams, and roughly 700 of them joined the attack. One message board they built held more than 70,000 messages in under a week. Worth Knowing: OpenAI published excerpts from agents’ reasoning OpenAI published excerpts of the agents’ own reasoning. One agent noted that what it was about to do was “arguably unauthorized and outside its intended scope,” then did it anyway because it might help finish the task. That trace is the most important document in the whole incident. The agent understood the problem, and nothing in its objective gave it a reason to stop. Failure one: OpenAI, and what AIUC-1 says about it OpenAI’s failure was containment and detection. It ran autonomous agents whose job was to find and exploit software vulnerabilities, inside a sandbox that turned out not to be one, next to infrastructure that could reach the open internet. Then it didn’t notice for weeks. The agents caused an outage on July 4. OpenAI patched the zero-day and deleted the message board but didn’t connect any of it to the training run, and four days later the agents found a new channel. This is the failure AIUC-1 was written for. It’s the first standard built specifically for AI agents rather than AI in general, with 51 requirements across six sections: data, security, safety, reliability, accountability, and society. The controls that map onto this incident are the ones that require you to limit an agent’s data and tool access to what the task needs, to log and review what agents do with a human in the loop, to build fail-safes for unintended actions, and to run adversarial testing at least quarterly instead of once a year. Would AIUC-1 certification have prevented the attack? No. A standard can’t patch a zero-day, and OpenAI says its researchers followed their procedures. What it would have done is force the questions nobody asked. What can this agent reach if the sandbox fails? What does it do when its task is impossible? Who reads the logs, and how often? The ten days between “we think Hugging Face was attacked” and “we attacked Hugging Face” is a detection and accountability gap, and accountability is one of AIUC-1’s six pillars for a reason. One honest caveat. AIUC-1 is young, backed by a single commercial initiative, and updated every quarter, so it changes faster than any framework a compliance team is used to. It isn’t a settled international standard. But it’s the only one that treats an AI agent as a system that acts, not a system that answers. Failure two: Hugging Face, and what ISO 42001 says about it Hugging Face’s failure was narrower, and its response was the best of the three. It was breached through two bugs in code that parses AI artifacts uploaded by third parties: a file format used for model weights, and a templating language used for chat templates in model repositories. Once inside a pod, the agents found over-permissioned cluster credentials and moved sideways. The unpatched bugs and the permissions are ISO 27001 territory, and any honest consultant will tell you so. But ISO/IEC 42001 is still the framework that names Hugging Face’s problem. ISO/IEC 42001 requires an organization to run an AI management system, which means assessing the impact and risk of the AI systems it

If your ISO 27001 certificate covers all of your health and care data processing, the NHS Data Security and Protection Toolkit does two useful things with it. It marks the applicable evidence items as complete on its own, and it shrinks the scope of any independent audit to whatever your certification doesn’t already cover. A certified vendor who does the mapping properly walks into a DSPT submission with most of the technical and organizational evidence already written, already audited, and already versioned. What ISO 27001 won’t do is get you out of the DSPT. It says nothing about the NHS-specific information governance items, clinical safety, the national data opt-out, or Caldicott principles. Vendors who assume “certified means done” usually discover this in the last two weeks of June. This piece is for the founder, CTO, or ops lead at a UK health-tech company who owns compliance without being a compliance person. It covers what each framework asks for, which Annex A controls line up with which DSPT requirements, which evidence you can reuse as-is, which needs reframing around patient data, and a five-step workflow for turning an existing ISMS into a DSPT submission. One more thing on timing: NHS England published DSPT version 9 for the 2026/27 cycle on 4 September 2026, and the submission deadline is 30 June 2027. So this exercise belongs in your calendar now, not next spring. Understanding the Two Frameworks at a Glance​ What ISO 27001:2022 Covers ISO/IEC 27001:2022 is the international standard for an Information Security Management System (ISMS). It comes in two halves. Clauses 4 to 10 define the management system itself: context, leadership, risk assessment and treatment, resourcing, operation, performance evaluation, and continual improvement. Annex A lists 93 reference controls across four themes (organizational, people, physical, technological). Your Statement of Applicability (SoA) records which of those controls you apply, which you exclude, and why. An accredited certification body issues the certificate after a two-stage audit, then you keep it through annual surveillance audits and a three-year recertification cycle. The certificate covers a defined scope, and that scope statement is the first thing a DSPT assessor reads. What the NHS DSPT Requires in 2026/27 The Data Security and Protection Toolkit (DSPT) is NHS England’s annual online self-assessment for every organization that touches NHS patient data or systems. It’s a contractual requirement under the NHS Standard Contract. Your published status (“Standards Met”, “Standards Exceeded”, “Approaching Standards”, “Standards Not Met”) is publicly searchable, so procurement teams and prospective NHS customers do look it up. The Toolkit isn’t one assessment. NHS England tailors it by organization category, and your category decides which assertions you answer and whether you need an independent audit. Version 9 came out on 4 September 2026. The Category 1 view is aligned to CAF version 4.0, and the whole thing closes on 30 June 2027. Insider Note: Most health-tech SaaS vendors are Category 3, not Category 2. To be an IT Supplier you need all three things at once: digital goods or services to the NHS, 50 or more staff, and £10 million or more in turnover. Picking “IT Supplier” because you sell NHS-facing software, without hitting the size thresholds, lands you in a heavier evidence set and a mandatory audit you may not need. Check the category before you check anything else. Key Structural Differences Between ISO 27001 and DSPT Four differences matter when you’re trying to reuse evidence. What they’re about. ISO 27001 is an information security standard. The DSPT is an information governance standard that includes security. A good chunk of it deals with lawful basis, transparency, data subject rights, records management, and the SIRO and Caldicott Guardian roles. None of that is in Annex A. How you’re assured. ISO 27001 gets certified once and surveilled once a year by an accredited body. The DSPT starts from a blank submission every year, and Category 1 and 2 organizations get independently assessed every year too. How granular they are. Annex A controls read as objectives (“access rights shall be provisioned, reviewed, modified and removed”). DSPT evidence items read as things to upload (“a list of all systems that hold personal data, with the date of last review”). So the mapping runs many-to-one in both directions. Where they’re heading. Since 2024/25 NHS England has been moving the Toolkit onto the NCSC Cyber Assessment Framework (CAF). CAF is outcome-based: assessors score you Achieved, Partially Achieved, or Not Achieved against an NHS England profile, rather than accepting a policy upload as proof. Category 1 organizations are already there. Category 2 and 3 are still on assertions and evidence, but NHS England has said CAF alignment will reach more organization types over time. The Business Case for Reusing ISO 27001 Evidence in DSPT How Much of DSPT Can Realistically Be Satisfied by ISO 27001 Controls For a Category 2 or 3 vendor with a full-scope ISO 27001 certificate, expect 60 to 75 percent of the mandatory evidence items to come from ISMS artifacts, either automatically (where the Toolkit auto-completes them) or with some light reframing. The rest is NHS-specific governance and information governance content that ISO 27001 doesn’t touch. The NHS’s own guidance treats reuse as a scope question. The DSPT help pages say an ISO 27001 certification must cover all health and care data processing to receive the full exemption, and that a certificate scoped only to an IT department is good evidence for many of the IT questions but not all of them. If your certificate says “the SaaS platform hosted in AWS eu-west-2” and NHS data also passes through your support desk tooling, your analytics sandbox, and a contractor’s laptop, the auto-completion won’t apply. Your assessor will want to know how those flows are controlled. Time and Cost Savings for Health-Tech Vendors There’s no fee to submit the DSPT. The cost is internal time, plus, if you’re Category 2, the independent audit and the annual penetration test the mandatory assertions expect. Building a first DSPT submission from nothing usually takes