Table of Contents

Reach SOC 2 Compliance in 6 Weeks or Less.

  / Secure AI Agent Vendor Certifications: 2026 Buyer’s Guide

Secure AI Agent Vendor Certifications: 2026 Buyer’s Guide

An AI agent that can read your inbox, query your CRM, and dig through internal documents has more standing access than most of your employees. It handles sensitive data, acts on its own, and often passes that data through sub-processors you’ll never see. Certifications are the quickest way to tell which vendors have let an outsider check their work, and which ones just put the word “secure” on a landing page.

No single certificate proves an AI agent is safe. But the right mix of security attestations, privacy certifications, and AI governance standards tells you the vendor has real controls, that an independent auditor has tested them, and that someone is on the hook when the agent misbehaves. This guide covers which certifications to ask for, how to verify them, and which claims should make you walk away.

Secure AI Agent Vendor

The Core Certifications Every Secure AI Agent Vendor Should Hold

SOC 2 Type II

SOC 2 Type II is the baseline for any SaaS or AI vendor that handles customer data. A licensed CPA firm audits the vendor against the AICPA’s Trust Services Criteria (Security, Availability, Processing Integrity, Confidentiality, and Privacy) and reports on whether its controls actually worked over a review period, usually 3 to 12 months. A Type I report only confirms the controls existed on one particular day. For an AI agent vendor, insist on Type II. Anything less tells you nothing about how the company runs day-to-day.

ISO/IEC 27001

ISO/IEC 27001 certifies that the vendor runs a formal information security management system (ISMS): documented risk assessments, defined controls, internal audits, and management review, all verified by an accredited certification body. It’s the most widely recognized security certification outside the US and often a hard procurement requirement in Europe, the UK, and the Gulf. A vendor with international customers should hold it alongside SOC 2, not instead of it.

ISO/IEC 27701 (Privacy Information Management)

ISO/IEC 27701 extends ISO 27001 with a privacy information management system (PIMS). It maps closely to GDPR concepts like controller and processor obligations, consent, and data subject rights. Almost every AI agent processes personal data at scale, and ISO 27701 is a decent signal that the vendor has built privacy into how it operates instead of delegating it to a policy PDF.

ISO/IEC 42001 (AI Management Systems)

ISO/IEC 42001 is the first certifiable international standard for AI governance. According to the International Organization for Standardization, it sets out requirements for building and maintaining an AI management system (AIMS): AI risk management, AI system impact assessments, lifecycle management, and oversight of third-party suppliers. For an AI agent vendor, this is the one that covers what SOC 2 and ISO 27001 don’t: how the vendor governs model behavior, training data, and the wider impact of autonomous systems.

Worth Knowing: ISO 42001 certificates only started appearing in volume in 2024, and the accreditation ecosystem is still catching up. Check that the certificate came from a certification body accredited for ISO 42001 specifically (under ANAB or UKAS, for example), not just one accredited for ISO 27001.

HIPAA (for Healthcare AI Agents)

If the agent touches protected health information (PHI), the vendor has to comply with the HIPAA Privacy and Security Rules and sign a Business Associate Agreement (BAA). There’s no official HIPAA certification, so vendors prove compliance through third-party assessments, a SOC 2 with HIPAA mapping, or HITRUST CSF certification. A vendor that won’t sign a BAA has disqualified itself for healthcare work.

PCI DSS (for Payment-Handling AI Agents)

AI agents that process, store, or transmit cardholder data (think agents automating billing, refunds, or checkout) fall under PCI DSS. Ask for the vendor’s Attestation of Compliance (AOC) and check whether a Qualified Security Assessor validated it or the vendor assessed itself. The current version is PCI DSS 4.x, so an AOC that still references 3.2.1 is out of date.

FedRAMP (for Government-Facing AI Agents)

FedRAMP authorization is mandatory for cloud services sold to US federal agencies. Authorizations come at Low, Moderate, and High impact levels, and every authorized service appears on the public FedRAMP Marketplace. If a vendor claims FedRAMP status and isn’t in the Marketplace, either the claim is false or the service is still “in process,” and those are very different things. State and local buyers should look for StateRAMP instead.

Worth Knowing: ISO 42001 Certificates

ISO 42001 certificates only started appearing in volume in 2024, and the accreditation ecosystem is still catching up. Check that the certificate came from a certification body accredited for ISO 42001 specifically (under ANAB or UKAS, for example), not just one accredited for ISO 27001.

HIPAA (for Healthcare AI Agents)

If the agent touches protected health information (PHI), the vendor has to comply with the HIPAA Privacy and Security Rules and sign a Business Associate Agreement (BAA). There’s no official HIPAA certification, so vendors prove compliance through third-party assessments, a SOC 2 with HIPAA mapping, or HITRUST CSF certification. A vendor that won’t sign a BAA has disqualified itself for healthcare work.

PCI DSS (for Payment-Handling AI Agents)

AI agents that process, store, or transmit cardholder data (think agents automating billing, refunds, or checkout) fall under PCI DSS. Ask for the vendor’s Attestation of Compliance (AOC) and check whether a Qualified Security Assessor validated it or the vendor assessed itself. The current version is PCI DSS 4.x, so an AOC that still references 3.2.1 is out of date.

FedRAMP (for Government-Facing AI Agents)

FedRAMP authorization is mandatory for cloud services sold to US federal agencies. Authorizations come at Low, Moderate, and High impact levels, and every authorized service appears on the public FedRAMP Marketplace. If a vendor claims FedRAMP status and isn’t in the Marketplace, either the claim is false or the service is still “in process,” and those are very different things. State and local buyers should look for StateRAMP instead.

Reach SOC 2 Compliance in 6 Weeks or Less

Schedule Your Free SOC 2 Assessment Today

Regulatory Frameworks AI Agent Vendors Must Comply With

Certifications are voluntary. Regulations aren’t. A credible AI agent vendor should be able to explain, in writing, how it meets each of the following.

GDPR (EU Data Protection)

Any agent processing personal data of people in the EU falls under GDPR, no matter where the vendor is based. Expect a signed Data Processing Agreement (DPA), a published sub-processor list, data residency options, and a working mechanism for the right to erasure. Erasure is genuinely hard for AI vendors, so ask specifically whether customer data ends up in model training and how deletion requests reach backups and fine-tuned models.

CCPA/CPRA (California Privacy)

The CCPA, as amended by the CPRA, gives California residents the right to access, delete, and opt out of the sale or sharing of their personal information. Vendors with US customers should have a service provider agreement covering CCPA obligations and be able to handle consumer rights requests within the statutory timelines.

The EU AI Act

The EU AI Act entered into force in August 2024 and applies in phases: prohibitions kicked in from February 2025, obligations for general-purpose AI models followed in August 2025, and the high-risk requirements are phasing in from 2026 onward, with some deadlines moved by the 2026 Digital Omnibus package. Two questions for any vendor. Has it classified its agent under the Act’s risk tiers, and can it show you the analysis? And if the agent gets used in a high-risk context like employment or credit decisions, what’s the plan for conformity assessment and technical documentation? A vendor that’s never heard of Annex III isn’t ready to sell into Europe.

NIST AI Risk Management Framework (AI RMF)

The NIST AI Risk Management Framework is voluntary, but it’s become the shared vocabulary for AI risk in the US. Its four functions (Govern, Map, Measure, and Manage) give you a structured way to question a vendor’s AI risk program, and NIST’s Generative AI Profile extends it to generative-specific risks. Vendors who can map their controls to the AI RMF usually have a real program behind the claims. The ones who can’t are usually improvising.

SOC 2 vs ISO 27001

SOC 2 vs ISO 27001: Key Differences for AI Agent Buyers

Buyers ask about these two more than anything else. The short answer: they solve different problems. SOC 2 shows you how controls actually performed over a period. ISO 27001 shows that the vendor runs a certified management system for security. For a full breakdown of how the two frameworks map to each other, see our guide to the key differences.

Which One (or Both) Your Vendor Should Have

For a vendor selling mainly into North America, SOC 2 Type II is the minimum. For one selling internationally, ISO 27001 usually isn’t optional either. Mature AI agent vendors increasingly hold both, plus ISO 42001, because each answers a different question: did the controls work, is there a disciplined security management system, and is anyone actually governing the AI. If budget forces the vendor to pick one first, what matters most to you as the buyer is that the chosen framework’s scope covers the agent product you’re buying.

Insider Note: An ISO 27001 certificate can legitimately cover a scope as narrow as one office or one internal system. Auditors regularly see vendors advertise the logo while the certified scope leaves out the flagship product entirely. Always read the scope statement on the certificate itself, not the badge on the website.

AI-Specific Certifications and Emerging Standards

ISO/IEC 42001 for AI Governance

ISO 42001 is currently the only certifiable AI governance standard, which makes it the strongest single differentiator among AI agent vendors. It also sets vendors up well for regulation: an AI management system built on 42001 covers much of the organizational groundwork the EU AI Act will demand, even though the certification itself doesn’t create a presumption of conformity with the Act.

ISO/IEC 23894 for AI Risk Management

ISO/IEC 23894 gives guidance on managing AI-specific risks across the system lifecycle, building on the general risk principles of ISO 31000. It’s a guidance standard, not a certifiable one, so treat any vendor claim of “ISO 23894 certification” as a red flag in itself. The correct claim is alignment. The right follow-up is to ask how the vendor identifies and treats AI risk sources like model drift, bias, and adversarial manipulation.

NIST AI RMF Alignment

Like ISO 23894, the NIST AI RMF isn’t certifiable. What you want from a vendor is a documented mapping: which controls implement Govern, Map, Measure, and Manage, and which artifacts back them up (model cards, evaluation reports, incident response runbooks, an AI Bill of Materials). That’s the point where model governance claims become checkable instead of decorative.

Industry-Specific Certification Requirements

Healthcare AI Agents (HIPAA, HITRUST)

Beyond HIPAA compliance and a signed BAA, many health systems require HITRUST CSF certification because it rolls HIPAA, NIST, and ISO requirements into one assessable framework with defined assurance levels. HITRUST has also added AI-specific assessment content, which makes it increasingly relevant for clinical AI agents.

Financial Services AI Agents (PCI DSS, SOX)

Agents touching cardholder data need PCI DSS validation, as covered above. Agents that feed financial reporting at public companies also run into SOX internal controls, so expect your auditors to ask for the vendor’s SOC reports and change-management evidence. Most financial institutions will run their own third-party risk assessment on top of whatever certifications the vendor holds.

Public Sector AI Agents (FedRAMP, StateRAMP)

US federal deployments need FedRAMP authorization at the impact level matching the data involved; most business data lands at Moderate. StateRAMP extends similar assurance to state and local government. Both come with continuous monitoring obligations, which work in your favor: the authorization stays under ongoing oversight rather than sitting there as a point-in-time stamp.

How to Verify a Vendor’s Certifications Are Legitimate

Requesting the SOC 2 Report vs. Attestation Letter

An attestation letter or a Trust Center badge only tells you a report exists. The full SOC 2 report, shared under NDA, contains the audit period, the system description, the criteria covered, the auditor’s tests, and any exceptions the auditor found. Read the exceptions and the vendor’s responses. A report with a few well-remediated exceptions is often more trustworthy than a suspiciously spotless one.

Checking ISO Certificate Registries

ISO itself doesn’t certify anyone. Certificates come from accredited certification bodies, and most of them run public verification portals where you can look up a certificate number. Confirm the certificate is current, names the right legal entity, and was issued by a body accredited by a recognized member of the International Accreditation Forum (IAF).

Validating Audit Dates and Scope

For SOC 2, check that the audit period is recent and continuous; any gap between the last report’s end date and today is uncovered time. For ISO certificates, check both the issue date and the expiry of the three-year cycle, and confirm the surveillance audits are actually happening. In every case, verify the scope covers the specific AI agent product, region, and infrastructure you’ll actually use.

Reviewing Sub-Processor and Third-Party Attestations

An AI agent vendor is usually a wrapper around other people’s infrastructure: foundation model APIs, cloud hosting, vector databases, observability tools. Request the sub-processor list and confirm the critical ones hold their own SOC 2 or ISO 27001 attestations. Your risk is the weakest link in that chain, and vendor risk management that stops at the first-party vendor misses most of the attack surface.

Pro Tip: SOC 2 Report

If a SOC 2 report period ended more than three months ago, ask for a bridge letter (also called a gap letter). It's a standard document where management confirms nothing material changed in the controls since the audit period ended. Established vendors produce one within days. Vendors who've never heard of it deserve extra scrutiny.

Red Flags: Certification Claims to Watch Out For

Expired or Out-of-Scope Reports

A SOC 2 report from two audit cycles ago, an ISO certificate past its surveillance date, or a certificate scoped to a product you aren’t buying: all of these fail verification. So does a report that only covers the corporate IT environment while the AI agent runs on separate, unaudited infrastructure.

Self-Attestations vs. Independent Audits

Security questionnaires, internal whitepapers, and “compliant with ISO 27001 principles” language are self-attestations. They have a place in due diligence, but they don’t replace an independent auditor’s opinion. The same goes for AI claims: “built with responsible AI principles” is marketing until an ISO 42001 certificate or an audited framework mapping backs it up.

Important: Watch for logo laundering: vendors displaying the AICPA SOC badge, an ISO logo, or a partner’s FedRAMP status as if it were their own. A common variant is pointing to the cloud provider’s certifications (AWS or Azure) as proof of the vendor’s own compliance. Infrastructure certifications don’t cover the vendor’s application, code, or personnel.

Certification Checklist for Evaluating AI Agent Vendors

Use this list as a minimum bar during procurement and security review:

  • SOC 2 Type II report obtained under NDA, period ending within the last 12 months, exceptions reviewed
  • ISO/IEC 27001 certificate verified via the certification body’s registry, scope covers the agent product
  • ISO/IEC 42001 certification held or on a committed roadmap, issued by an accredited body
  • ISO/IEC 27701 or an equivalent, documented privacy program for personal data processing
  • GDPR: signed DPA, published sub-processor list, data residency options, erasure mechanics explained
  • EU AI Act risk classification documented; conformity plan exists if high-risk use is possible
  • NIST AI RMF or ISO 23894 alignment mapping with supporting artifacts (model cards, evals, incident runbooks)
  • Industry add-ons where relevant: BAA and HITRUST for healthcare, PCI DSS AOC for payments, FedRAMP/StateRAMP for government
  • Sub-processor attestations collected for foundation model providers and hosting infrastructure
  • Continuous compliance monitoring and penetration testing cadence confirmed, with a recent pentest summary available

Reach SOC 2 Compliance in 6 Weeks or Less

Schedule Your Free SOC 2 Assessment Today

The Bottom Line

Certifications won’t tell you whether an AI agent hallucinates, but they will tell you whether the company behind it takes controls and accountability seriously. Require SOC 2 Type II and ISO 27001 as the security floor, treat ISO 42001 as the emerging differentiator for AI governance, add HIPAA, PCI DSS, or FedRAMP where your industry demands it, and verify everything against primary sources: the full report, the certificate registry, the FedRAMP Marketplace. Vendors with real programs make verification easy. Vendors without them make it awkward, and that awkwardness is your answer.

Frequently Asked Questions

Is SOC 2 Type II enough for an AI agent vendor?

Necessary, but not sufficient. SOC 2 covers security, availability, and related criteria for the service organization. It wasn’t designed to assess AI-specific risks like model behavior, training data governance, or algorithmic bias. Pair it with ISO 42001 certification or a documented NIST AI RMF mapping.

Vendors selling internationally generally need both. US buyers ask for SOC 2; buyers in Europe, the UK, and the Gulf expect ISO 27001. The two overlap heavily in control substance, so a vendor with one can usually reach the other with moderate extra effort.

No. The Trust Services Criteria cover organizational and system controls, not model quality. Model drift, hallucination rates, and bias need AI-specific governance: ISO 42001, ISO 23894 alignment, evaluation reports, and ongoing monitoring.

SOC 2 Type II reports are reissued every year with a new audit period. ISO certificates run on a three-year cycle with annual surveillance audits. PCI DSS attestations are annual. FedRAMP requires continuous monitoring with annual assessments. Anything older than its cycle should be treated as lapsed.

GDPR compliance is the legal requirement, usually evidenced through a DPA, transfer mechanisms, and processor obligations, and the EU AI Act adds obligations based on the system’s risk classification. ISO 27001, ISO 27701, and ISO 42001 aren’t legally required, but they’re the certifications European buyers most often accept as evidence that the legal obligations are actually being met.

No. They answer different questions, and buyers increasingly expect both. SOC 2 attests that operational security controls worked over a period; ISO 42001 certifies a management system for governing AI responsibly. Expect ISO 42001 to become a standard line item in AI vendor questionnaires alongside SOC 2, not in place of it.

Axipro Author

Picture of Pedro Dias

Pedro Dias

Pedro has been writing online for over 10 years. With experience in all things programming, cyber security, and compliance, he is our editor-in-chief at Axipro.

Blog Highlights

Explore More Articles

Two compromised versions of LiteLLM sat on PyPI for roughly 40 minutes on the morning of March 24, 2026. That window was enough to capture secrets from around 434,000 CI/CD pipeline runs across nearly 2,500 organizations, including AWS, Samsung, Cisco, Salesforce, Siemens, and Deloitte. In August, researchers at CloudSEK and Hudson Rock confirmed they had obtained the raw exfiltrated data: a 153GB archive containing 433,909 files of environment variables, cloud keys, Kubernetes secrets, and API tokens harvested live from running pipelines, as covered by Help Net Security’s reporting on the credential archive. If LiteLLM runs anywhere in your stack, or you touch any AI proxy infrastructure at all, you need answers to three things: whether you were exposed, what to rotate first, and whether the rotation you did back in March actually held. That last one matters more than it sounds, because “we rotated everything” has already burned at least one very large company. How the Breach Happened The attack didn’t start with LiteLLM. On March 19, 2026, a threat group called TeamPCP compromised the build pipeline of Trivy, a vulnerability scanner half the industry runs, and pushed a poisoned release. LiteLLM’s own CI pipeline ran Trivy, so the poisoned scanner had legitimate read access to the project’s runner environment. The attackers used that to steal LiteLLM’s PyPI publishing tokens and ship two malicious releases of their own: versions 1.82.7 and 1.82.8. KICS and the Telnyx Python SDK got hit in the same campaign. The payload design is the part worth studying. The malicious package dropped a .pth startup hook into site-packages, so the code ran the moment any Python interpreter started on the machine, whether or not anything imported LiteLLM. From there it harvested environment variables, read local credential files like .aws/credentials and .kube/config, tried to move laterally across Kubernetes clusters, and installed a systemd backdoor dressed up as a generic telemetry service. InfoQ’s coverage of the PyPI compromise put downloads of the compromised release above 40,000. For scale, LiteLLM normally gets downloaded around 3 million times a day. The exfiltration had a nasty fallback, too. According to CloudSEK, stolen data was encrypted and sent to a typosquatted domain, and when that failed, the malware created a public repository inside the victim’s own GitHub account and uploaded the loot as a release asset. Some companies were publishing their own secrets to the open internet and had no idea. Worth Knowing: The malicious code only existed in the PyPI artifacts. The GitHub source repository stayed clean the whole time, so a developer reviewing the code on GitHub saw nothing wrong. Source review isn’t artifact verification. If you don’t check that what the registry serves matches the upstream source, this class of attack is invisible to you. How to Check If You Were Exposed Three checks, from quickest to most involved. 1. Confirm whether the compromised versions ever ran The malicious versions went live on PyPI at 10:39 UTC on March 24, 2026 and got quarantined about 40 minutes later. The project’s advice: treat any install from that day before 16:00 UTC as suspect. Search your lockfiles, pip caches, SBOMs, and container image histories for 1.82.7 and 1.82.8. And check your internal artifact mirrors. An Artifactory or Nexus proxy that cached the bad release in March can keep serving it internally long after PyPI pulled it. Keep the .pth mechanism in mind when you scope this. The question isn’t “which applications import LiteLLM,” it’s “which machines had the package installed at all,” because every Python process on an infected machine triggered the payload. 2. Hunt for persistence Rotation is pointless if the attacker still has a foothold. Check developer machines, CI runners, and containers for unauthorized .pth files in site-packages and for suspicious systemd units, especially anything posing as a system telemetry service. And review activity from March 24 onward, not just the 40-minute window. Persistence is there so the access outlives the infection. Pro Tip: Don’t limit the persistence hunt to live machines. Base container images rebuilt in late March may have baked the payload into every image derived from them since. Scan your image registry for the affected LiteLLM versions and for unexpected .pth files, then trace which running workloads came from flagged images. 3. Check whether your secrets are in the dump Hudson Rock has published a domain lookup tool and is running ethical disclosures for affected organizations, and CloudSEK maintains a high-confidence victim list. Use them, but know their limits. Attribution in this dataset is genuinely hard. One dump with a siriusxm.com committer email actually traced, through its self-hosted GitLab endpoints, to AdsWizz, a SiriusXM subsidiary. And a large share of the dumps are generic pipeline configurations with no identifying domain, email, or server name at all. Absence from a victim list is not evidence of absence. If your pipelines ran the compromised versions, assume exposure no matter what a lookup tool tells you. What to Rotate, in What Order The guidance from both research teams is blunt: treat every secret the LiteLLM environment could reach as compromised. That covers secrets on disk, in memory, injected into CI jobs, and anything retrievable through instance metadata services. Work down by blast radius: Priority Credential type Why it comes first 1 Cloud IAM keys (AWS, GCP, Azure) Direct control of infrastructure, data stores, and billing. This is where attackers monetize fastest. 2 GitHub and GitLab PATs, package publishing tokens These let an attacker poison your releases and turn your company into the next link in the supply chain. 3 Kubernetes service account tokens and kubeconfigs Lateral movement across clusters was built into the payload, not a theoretical risk. 4 Database passwords and third-party API keys Dumped in plain text in the archive, often with no attribution, so nobody will warn you they leaked. 5 AI provider API keys Billing abuse, quota theft, and access to whatever data flows through your LLM routing layer. One word matters more than the rest of this article: revoke, don’t just rotate. That

The EU AI Act names recruitment AI as high-risk. Annex III explicitly lists AI systems used for recruitment, candidate selection, and employment decisions, which pulls CV screeners, video interview platforms, and assessment tools into the most demanding compliance regime the Act contains. The original compliance date for these systems was August 2, 2026. In June 2026, the EU’s Digital Omnibus moved the deadline to December 2, 2027, a 16-month extension that has led many HR and talent teams to shelve the topic entirely. That’s a mistake, for two reasons. First, one rule that directly affects recruitment technology is already in force: the ban on emotion recognition in the workplace has applied since February 2, 2025, and it catches features still shipping in some video interview products today. Second, the deferred obligations didn’t shrink. Conformity assessments, human oversight design, bias monitoring, and documentation all still arrive in full, and the practical work of auditing a recruitment stack, renegotiating vendor contracts, and training hiring teams routinely takes a year or more. Here’s what the EU AI Act actually requires of employers and vendors using recruitment tools, on the timeline that now applies. Why Recruitment Tools Are Classified as High-Risk Under the EU AI Act​ Definition of High-Risk AI Systems in Hiring​ The Act takes a list-based approach. Annex III, point 4, designates as high-risk any AI system intended for the recruitment or selection of natural persons, including placing targeted job advertisements, analyzing and filtering applications, and evaluating candidates. The same point covers AI used for decisions on promotion, termination, task allocation, and monitoring of workers, so the classification follows the tool through the entire employment lifecycle, not just the hiring funnel. The reasoning is straightforward: hiring decisions shape access to livelihoods, and algorithmic discrimination in hiring is well documented. The European Commission’s regulatory framework for AI treats employment as one of the areas where an AI error or bias causes serious harm to fundamental rights. That’s the test for the high-risk tier. Types of Recruitment Tools Affected In practice, the high-risk classification captures most of the modern recruitment stack: CV and resume screeners that rank or filter applicants, video interview platforms that score responses or delivery, psychometric and skills assessment tools that produce scores feeding a hiring decision, sourcing and matching algorithms that decide which candidates a recruiter sees, and programmatic job ad targeting systems that determine who sees a vacancy at all. If the system’s output materially influences who advances and who does not, assume high-risk until proven otherwise. Important: Emotion recognition is not high-risk in the workplace. It is prohibited. Article 5 bans AI systems that infer emotions of people in the workplace (outside narrow medical and safety cases), and that ban has applied since February 2025 with the Act’s top penalty tier attached. If your video interview vendor markets “engagement scoring” or “sentiment analysis” of candidates, that feature needs to be switched off for EU hiring now, not in 2027. Recruitment Tools That May Fall Outside High-Risk Classification Not everything in the HR stack qualifies. The Act carves out systems performing narrow procedural tasks that do not materially influence decision outcomes. An applicant tracking system that stores applications, schedules interviews, and sends templated emails is a database with a workflow, not a high-risk AI system. The same goes for tools that transcribe interviews without scoring them, deduplicate candidate records, or generate first drafts of job descriptions for a human to edit. The line is decision influence: the moment a tool ranks, scores, filters, or recommends candidates, it crosses into Annex III territory. Deployers who rely on an exemption must be able to document that assessment, so “we decided it doesn’t count” needs to exist on paper. Extraterritorial Scope: Which Employers Are Covered The Act applies to providers placing AI systems on the EU market and to deployers established in the EU, but it also reaches further: it covers providers and deployers located outside the EU where the output of the system is used in the EU. For recruitment, the consequence is blunt. A US or UK company with no EU entity that uses an AI screener to filter applicants for roles based in Berlin or Dublin, or that screens candidates located in the EU, is using the system’s output in the Union. Brexit doesn’t move UK employers out of scope when they hire into or from the EU. Providers vs. Deployers of Recruitment AI Tools The Act splits obligations between the provider (the vendor that develops the tool and places it on the market) and the deployer (the employer using it). Most employers are deployers, and deployer obligations are lighter but real. One common trap: an employer that substantially modifies a high-risk system, or puts its own name on it, can be reclassified as a provider and inherit the full provider stack. Heavy customization of a screening model, or fine-tuning it on your own hiring data, can be enough to trigger this. Key Obligations for Employers Using AI Recruitment Tools Human Oversight in Automated Hiring Decisions Deployers must assign oversight of the system to people with the competence, training, and authority to intervene. That last word matters. A recruiter who rubber-stamps whatever the ranking algorithm produces, because nobody has time to review 800 rejected CVs, doesn’t count as oversight. Regulators and courts will look at whether the human could genuinely override the system and whether they ever did. Designing review checkpoints where a person can meaningfully change the outcome, and logging when they do, is the core of compliant deployment. Transparency Requirements Toward Candidates Employers must inform workers and their representatives before putting a high-risk AI system into use at work, and candidates subjected to such a system must be told it is being used. In countries with works councils, such as Germany, this obligation lands on top of existing co-determination rights, so employee representatives may need to be consulted before the tool goes live rather than just told afterward. Burying an AI disclosure in a privacy policy paragraph is unlikely to survive scrutiny.

A green dashboard is not an audit opinion. Compliance automation platforms like Vanta, Drata, Secureframe, and Hyperproof have made SOC 2 readiness faster and cheaper, but every audit cycle produces the same pattern: controls that sat at “passing” for months come back from the auditor with exceptions or requests for re-testing. The four controls below account for a disproportionate share of those rejections, and they all fail for the same underlying reason. The tool confirmed that evidence exists. The auditor tested whether the control actually operated. This article walks through each of the four: what auditors reject, why, and how to fix the evidence before fieldwork starts. Why Compliance Tools Show “Passing” But Auditors Still Reject Controls​ The Gap Between Automated Checks and Auditor Judgment Compliance platforms run continuous control monitoring: API calls that check whether a configuration exists, a document is uploaded, or a task is marked done. That’s real value. It catches drift, keeps evidence in one place, and saves weeks of screenshot collection. An audit is a different exercise. A SOC 2 examination is an attestation performed by a CPA firm under AICPA standards, and the auditor’s job is to form an independent opinion on whether your controls met the Trust Services Criteria. That opinion rests on professional judgment, not on whether an API integration returned a 200 response. What “Passing” Actually Means in Your Compliance Dashboard​ When a control shows “passing,” the platform is telling you one narrow thing: at the moment of the last scan, an automated test found the artifact or setting it was programmed to look for: MFA enforced in the identity provider, a policy document uploaded, a training campaign sitting at 100%. The test says nothing about whether the underlying process ran the way your control narrative claims it did, or whether it ran that way across the whole audit period. How Auditors Evaluate Controls Beyond the Checkbox Auditors test two dimensions. Design effectiveness asks whether the control, as described, would meet the criterion if it worked as intended. Operating effectiveness, the core of a SOC 2 Type 2 report, asks whether it actually did throughout the audit period. To answer that, the auditor pulls a population (every access review, every change, every new hire in the period), selects a sample, and inspects the evidence item by item. A dashboard status feeds into that process. It doesn’t replace it. Insider Note: Auditors increasingly ask for evidence outside the compliance platform precisely because they know what the platform auto-collects. If every artifact you produce comes from the same tool export, expect the auditor to independently pull the population from the source system and compare. Discrepancies between the two are one of the fastest routes to an exception. Control #1: Access Reviews That Automation Marks Complete but Auditors Reject Why Auditors Reject Automated Access Review Evidence​ User access reviews sit under the logical access criteria (CC6.1 through CC6.3), and they are the single most common source of audit exceptions we see. The typical failure: the platform generated a user list, someone clicked “complete,” and the dashboard turned green. The auditor then asks a simple question the evidence can’t answer: what did the reviewer actually decide? The Missing Element: Documented Reviewer Judgment​ An access review is a judgment control. Someone with knowledge of the system must look at each account and confirm the access is still appropriate for the person’s role. A timestamped task closure proves the task was closed. It doesn’t prove anyone assessed anything, and an “approve all” review completed in ninety seconds gets exactly the skepticism it deserves. What Auditors Actually Want to See in Access Review Evidence Auditors look for four things: The full population of accounts at the time of review (including service accounts and admin roles), Evidence of who reviewed it and when, explicit dispositions per account or group (retain, modify, revoke), and Proof that flagged access was actually removed. That last item, the deprovisioning ticket showing revocation within a defined window, is the piece most companies can’t produce. How to Fix Your Access Review Control Before the Audit​ Assign a named control owner per in-scope system, run reviews quarterly, and require reviewers to record a disposition for every line, not a blanket approval. When access is revoked, link the removal ticket to the review record. If a quarter was missed, don’t backfill it. Document it honestly and show the remediation, because auditors treat fabricated retroactive evidence far more severely than a disclosed gap. Control #2: Change Management Approvals That Pass Automated Scans​ Why Ticket Closure Isn’t Proof of Approval​ Change management (CC8.1) automation typically verifies that production changes link to a ticket and the ticket is closed. Auditors test something stricter: that each sampled change was approved by an authorized person before deployment. An approval added after the merge, or a ticket closed by the same engineer who wrote the code, fails that test even though every automated check came back green. The Segregation of Duties Problem Automation Misses Segregation of duties is the requirement that no single person can develop, approve, and deploy the same change. NIST’s SP 800-53 control catalog treats it as a foundational access control principle, and SOC 2 auditors apply the same logic. Small engineering teams trip on this constantly. Self-approved pull requests, admins who can bypass branch protection, direct pushes to main: a scanner sees “changes with tickets” while an auditor sees SoD violations. Emergency Changes and Retroactive Approvals: Common Rejection Triggers​ Every audit period contains hotfixes. Auditors don’t reject emergency changes. They reject emergency changes with no documented post-hoc review. If your policy says urgent changes get retroactive approval within two business days, the auditor will sample your emergency changes and check exactly that. No policy, or a policy nobody followed, produces an exception. Rebuilding Change Management Evidence Auditors Will Accept​ Enforce the control technically: branch protection requiring at least one independent reviewer, no admin bypass, and deploy pipelines that only run from protected branches. Then write the emergency change procedure down and generate the review artifact every time it fires.