Table of Contents

Reach SOC 2 Compliance in 6 Weeks or Less.

  /

  / AIUC-1 AI Agent Certification: The Complete Guide

AIUC-1 AI Agent Certification: The Complete Guide

Most security certifications were built for software that follows rules. AI agents do not. They consume data, draw conclusions, call tools, and take action, increasingly without a human in the loop.

That gap is what AIUC-1 was created to close: it is the first auditable security standard built specifically for AI agents, and a few enterprise buyers have started asking vendors for it by name.

This guide covers what AIUC-1 actually tests, the six risk domains it audits, how the certification process works, what it costs, how long it lasts, and how it aligns with SOC 2, ISO 42001, ISO 27001, and the NIST AI Risk Management Framework. It also covers the structural questions worth asking before you treat an AIUC-1 report as proof of anything.

AIUC-1 AI Agent Certification The Complete Guide

What Is AIUC-1 Certification?

AIUC-1 is a certifiable standard for AI agents created by the Artificial Intelligence Underwriting Company (AIUC), a San Francisco-based, venture-backed startup founded by people with experience at organizations including Anthropic. The standard was developed with input from Orrick, Stanford, the Cloud Security Alliance, MIT, and MITRE, and launched in mid-2025.

The framework comprises 51 requirements and 130 controls, organized across six risk pillars. It evaluates whether an organization has implemented and tested the technical guardrails, operational practices, and legal policies needed to reduce the risk of unsafe, unreliable, or unauthorized AI behavior. Certification applies to a specific AI system or product, not to the organization as a whole. An AIUC-1 certificate, audit report, and badge tell enterprise buyers that an agent has been independently tested against agent-specific risks.

People describe AIUC-1 as the “SOC 2 for AI agents,” and the analogy holds in spirit. The difference is what it looks at. SOC 2 examines a service organization’s general controls. AIUC-1 examines how an agent behaves under pressure: when someone tries to jailbreak it, when it is asked to do something outside its scope, when it has access to data it should not expose.

Worth Knowing: About AIUC-1

AIUC-1 does not define what counts as an "AI agent." The vendor decides which system to certify and what falls in scope. That makes scope the single most important thing to check on any certificate, because a narrowly scoped audit may not cover the agent you actually use.

Why AIUC-1 Certification Matters for Enterprise AI Adoption

The business case rests on a simple problem: enterprises cannot reliably assess the security of their AI vendors, and the failures are expensive. According to EY research on responsible AI, 64% of companies with over $1 billion in revenue have already lost more than $1 million to AI-related failures. 

That gap shows up directly in sales cycles. When security, legal, and procurement teams evaluate an AI vendor, they ask about hallucinations, prompt injection defenses, and what happens when an agent makes an unauthorized call. SOC 2 and ISO 27001 do not answer those questions. AIUC-1 gives buyers a structured, third-party-tested answer, which is why holding the certificate can move a stalled procurement review forward.

The certification also produces real engineering outcomes, not just a badge. AIUC has reported cases where a customer service agent’s hallucination rate dropped from 11% to under 2% after strengthening its groundedness filter, and another where inappropriate-tone outputs fell from 9% to under 2% through better defensive prompting and output moderation. One company found and patched a PII exposure vulnerability during the certification process itself.

The Six Core Risk Domains Covered by AIUC-1

The Six Core Risk Domains Covered by AIUC-1

AIUC-1’s 51 requirements are grouped into six domains. Each targets a category of risk that traditional security frameworks were not designed to handle.

Data and Privacy

Covers how customer data is used, retained, and protected. Requirements address input and output data policies, limits on what data the agent can access, protection of IP and trade secrets, prevention of cross-customer data exposure, and prevention of PII leakage. This is where the standard forces clarity on whether customer data trains the model and how long it is kept.

Security

The adversarial-resistance domain. It covers third-party testing of adversarial robustness, detection and real-time filtering of malicious inputs, prevention of prompt injection and unauthorized agent actions, enforcement of user access privileges, and protection of the deployment environment. This is the heart of what separates an agent audit from a general security audit.

Safety

Focuses on preventing harmful and out-of-scope outputs. Requirements include defining an AI risk taxonomy, conducting pre-deployment testing, preventing harmful and customer-defined high-risk outputs, and flagging high-risk outputs for human review. Safety is partly judgment-based, which means documentation alone can sometimes satisfy a requirement, so the testing behind it deserves scrutiny.

Reliability

Targets the failure modes that erode trust in production: hallucinations and tool misuse. Controls cover hallucination prevention and restrictions on which tools an agent can call and when. For a customer-facing agent, this is the domain that keeps it from inventing a refund policy or triggering the wrong workflow.

Accountability

Covers what happens when things go wrong. Requirements include AI failure response plans, vendor due diligence, and clear AI disclosure so users know when they are interacting with an agent. With human workers, accountability is built into org charts and chains of command. Agents need an equivalent, and this domain supplies it.

Society

The broadest domain, focused on preventing misuse with wider consequences: AI-enabled cyber attacks and CBRN (chemical, biological, radiological, nuclear) misuse. Most enterprise agents will touch only a few of these controls, but they matter for higher-capability systems.

Insider Note: Of the 130 total controls, roughly 65 are mandatory, and 65 are optional. A straightforward agent typically needs to meet around 40 controls. A complex, multi-modal agent gets closer to 65. The scoping exercise determines which apply, so two AIUC-1 certificates can represent very different amounts of work.

Ready to Earn Your AIUC-1 Certification?

Accelerate Your AI Certification Journey

Who Needs AIUC-1 Certification?

AIUC-1 is built for any company developing or deploying agentic AI that sells into enterprises. The strongest fit is an organization whose product uses AI agents in customer-facing operations, handles confidential data through autonomous workflows, or makes decisions that affect critical business processes.

Certified systems so far include customer service agents, candidate scoring and interviewer agents, internal automation agents, summarization agents, and image generation agents. Certified organizations range from seed-stage startups to publicly traded enterprises, so company size is not the gating factor. Buyer demand is. The clearest reason to pursue AIUC-1 is that your enterprise buyers are asking for it, though early adoption also lets a vendor shape the security conversation before competitors do.

How the AIUC-1 Certification Process Works

The certification runs in distinct phases, with an accredited auditor guiding the organization through evidence collection and gap remediation.

Pre-Audit Readiness Assessment

Scoping and kickoff usually take one to two weeks. The organization works with the auditor to complete a scoping questionnaire, define which system and which controls are in scope, appoint internal leaders, and identify where current practices fall short of AIUC-1’s requirements. An internal-facing agent with limited data access will scope into far fewer controls than an external customer service agent handling sensitive data.

Evidence Collection and Documentation

The bulk of the work, typically three to five weeks. Teams gather documentation across operational practices, legal policies, and technical implementations, then remediate gaps surfaced during scoping. This phase includes hands-on testing of the system for hallucinations, prompt injection resistance, and the other risks the standard covers.

Independent Third-Party Audit

An accredited auditor reviews the evidence, conducts or reviews technical testing, and assembles the report. The audit combines upfront technical testing with a review of operational controls.

Certification Issuance

The Artificial Intelligence Underwriting Company issues the certificate. Finalizing the audit, building the report, and obtaining sign-off generally takes one to three weeks. The deliverables are the certificate, a detailed audit report with third-party attestation and evaluation results, and a badge for use in a trust center or sales collateral.

Ongoing Monitoring and Renewal

Certification is not a one-time event. Technical tests are re-run at least quarterly, and the full set of technical, operational, and legal controls is re-audited annually. This continuous cadence is meant to keep safeguards current as both AI capabilities and attack techniques evolve.

Requirements for AIUC-1 Certification

There is no single checklist that applies to every applicant, because scope drives requirements. Every certification covers the mandatory controls relevant to the system, plus whichever optional controls the agent’s data access, autonomy, and modality bring into scope.

In practice, organizations need documented input and output data policies, demonstrated defenses against prompt injection and unauthorized actions, evidence of pre-deployment and adversarial testing, a defined risk taxonomy, failure response plans, and clear disclosure practices. Where there is no single industry best practice, such as data retention, the standard emphasizes clear disclosure and enforcement over mandating one specific approach.

 

How Long Does AIUC-1 Certification Take?

AIUC lists a typical timeline of four to eight weeks, depending on the maturity of the organization’s existing safeguards and governance. Its FAQ gives a slightly wider range of five to ten weeks. Organizations that already have AI governance in place, for example an ISO 42001 management system, tend to land at the faster end because much of the operational and legal groundwork already exists.

 

How Much Does AIUC-1 Certification Cost?

AIUC does not publish standard pricing, and any firm quote depends on scope, the number of controls in play, the agent’s complexity, and the auditor engaged. Cost is driven by the same factors as the timeline: a simple internal agent scoping into roughly 40 controls is a smaller engagement than a complex multi-modal agent scoping into 65, with the adversarial testing that entails.

Treat published figures from third parties with caution and get a scoped quote from an accredited auditor. As a directional anchor, AIUC-1 sits in the same tier as other independent technical audits rather than a self-attestation, so budget alongside what you would expect for a SOC 2 Type II or ISO 42001 audit, plus the cost of any remediation the readiness phase surfaces.

Pro Tip: Run a Gap Assessment

Run a gap assessment before you commit to a full audit. The readiness phase is where most of the real cost hides, because remediation work, not the audit fee, is usually the larger line item. Knowing your gaps first lets you budget accurately and avoid a stalled audit halfway through evidence collection.

How Long Is an AIUC-1 Certificate Valid?

An AIUC-1 certificate is valid for twelve months. Maintaining it requires technical testing at least every three months. Miss the quarterly cadence and the certificate lapses, even inside the twelve-month window. New requirements introduced through quarterly updates are evaluated at the next annual re-audit, so the certificate you hold reflects the version of the standard in force when you were audited.

 

Who Can Issue an Official AIUC-1 Certificate?

The Artificial Intelligence Underwriting Company issues all certificates. Schellman was the first accredited auditor for the standard, and accredited auditors handle evidence collection and prepare the reports. The ongoing quarterly technical evaluations are run centrally by AIUC itself rather than by individual auditors, which the company says keeps testing consistent across all certified organizations.

This structure matters for how much weight a certificate carries, a point covered in the challenges section below.

AIUC-1 Certification vs. Other Frameworks

AIUC-1 is designed to complement existing frameworks, not replace them. It operationalizes the AI-specific frameworks and avoids duplicating the general-purpose ones.

Framework

What it covers

Certifiable?

AI-agent specific?

Relationship to AIUC-1

AIUC-1

Technical, operational, and legal controls for AI agent risk

Yes

Yes

The standard itself

SOC 2

General service organization security controls

Attestation

No

Coexists; still required for enterprise sales

ISO 42001

AI management system (governance and process)

Yes

Partly

AIUC-1 validates that the governance produces working controls

ISO 27001

Information security management system

Yes

No

Foundational security; AIUC-1 sits on top

NIST AI RMF

Voluntary AI risk-management guidance

No

Partly

AIUC-1 translates its functions into testable controls

AIUC-1 vs. SOC 2

SOC 2 covers a vendor’s general cybersecurity posture. It does not address hallucinations, prompt injection, or unauthorized tool calls. The two are complementary: SOC 2 remains table stakes for selling into enterprise, while AIUC-1 answers the AI-specific questions SOC 2 leaves open.

AIUC-1 vs. ISO 42001

ISO 42001 certifies that an organization has a responsible AI management system, the policies and processes for developing and operating AI. AIUC-1 incorporates a number of controls directly from ISO 42001 and then extends them, translating management-system requirements into auditable technical controls and adding protections against risks like hallucinations and jailbreaks. Many organizations pursue both: ISO 42001 builds the governance, AIUC-1 proves the safeguards behind it hold up under testing.

AIUC-1 vs. ISO 27001

ISO 27001 governs information security management broadly. It is foundational and agnostic to AI. AIUC-1 assumes that kind of security baseline exists and focuses on the agent-specific layer above it, so the two rarely overlap.

AIUC-1 vs. NIST AI RMF

The NIST AI Risk Management Framework provides high-level, voluntary guidance with no certification path. It tells organizations what to think about, not how to prove they have done it. AIUC-1 takes NIST’s functions and turns them into specific, testable controls, which is the piece NIST deliberately leaves open.

AIUC-1 also maps to threat models the security community already uses, including MITRE ATLAS, the OWASP Top 10 for Agentic Applications, and the OWASP LLM Top 10, and it maps more than 30 articles of the EU AI Act to auditable requirements.

Ready to Earn Your AIUC-1 Certification?

Accelerate Your AI Certification Journey

How to Prepare for AIUC-1 Certification

Start with scope. Decide which agent or product you are certifying and document its data flows, the tools it can call, the models and versions it runs on, and its level of autonomy. That definition determines which controls apply and shapes everything downstream.

Next, run a readiness assessment against the six domains and fix the gaps before the formal audit begins. If you already hold ISO 42001 or have a NIST AI RMF program, map what you have onto AIUC-1’s controls; much of the operational and legal evidence will carry over. Build out the technical testing capability you will need: adversarial testing for prompt injection, groundedness evaluation for hallucinations, and output moderation. Going into the audit with that infrastructure already running is what separates a four-week certification from a ten-week one.

 

Benefits of Becoming AIUC-1 Certified

The headline benefit is faster enterprise deals. A certificate, report, and evaluation results give procurement, legal, and security teams a structured way to clear AI-specific risk, which removes a common point of friction late in the sales cycle. One certified executive summed it up by saying the certificate lets them sign contracts faster because it is a clear signal of trust.

Beyond sales, the process improves the product. The measurable drops in hallucination and inappropriate-output rates that companies report during certification are real engineering wins, not marketing. AIUC-1 also carries an insurance dimension that sets it apart: certification is backed by Lloyd’s of London insurance, and AIUC underwrites the risk associated with the certified agent. That changes the incentive structure, because the body certifying the agent also takes on financial exposure if it fails.

 

Common Challenges in Achieving AIUC-1 Certification

The first challenge is technical readiness. Building reliable defenses against prompt injection and hallucination is hard engineering work, and the readiness phase often surfaces gaps that take real effort to close.

Scope ambiguity is the second: because AIUC-1 does not define “AI agent,” getting the scope right requires careful judgment, and a scope drawn too narrowly undermines the certificate’s value.

The third challenge is structural, and buyers as much as vendors should understand it. AIUC authors the standard, runs the technical evaluations, issues the certificates, and sells the AI agent insurance that the certification enables. Security researcher Zack Korman has argued that this vertical integration creates potential conflicts of interest at several steps, with the closest precedent being the issuer-pays model in credit ratings, an arrangement that contributed to inflated ratings before the 2008 financial crisis. AIUC’s counterargument is that its insurance business creates a counter-incentive, since losses on a certified agent hit AIUC directly. There is also no external accreditation body: AIUC accredits its own auditors, so calling AIUC-1 a “standard” rests on AIUC’s own authority rather than a third party like ANSI or UKAS. None of this makes the certificate worthless. It makes it evidence to interrogate rather than a guarantee to accept at face value.

Important: No certification eliminates risk from a probabilistic, fast-changing system. Just as a SOC 2 report or a penetration test does not prove a system is secure against every threat, an AIUC-1 certificate cannot guarantee an agent is safe. Treat it as tested evidence of specific controls, scoped to a specific system, at a specific point in time.

 

The Bottom Line

AIUC-1 is the first serious attempt to give AI agents the kind of independent, testable security assurance that SOC 2 gave SaaS. It audits six domains that existing frameworks miss, runs on a quarterly testing cadence built for how fast AI moves, and comes backed by insurance that puts the issuer’s own money behind the result. It is also young, self-accredited, and commercially structured in ways worth scrutinizing.

For vendors selling agentic AI into enterprises, the practical question is no longer whether AI-specific certification is coming, but whether you would rather lead the conversation or explain to a procurement team why you cannot answer it. If buyers are already asking, the answer is straightforward.

Frequently Asked Questions About AIUC-1 Certification

Is AIUC-1 certification mandatory?

No. AIUC-1 is voluntary. It is not required by law anywhere. Demand is driven by enterprise buyers who want independent assurance on AI-specific risks, not by regulation.

Yes. Certified organizations range from seed-stage startups to publicly traded enterprises. Company size does not gate eligibility; the scope and maturity of the agent’s safeguards do.

No. Like a SOC 2 report or a penetration test, AIUC-1 provides evidence that specific controls were tested, scoped to a specific system at a point in time. It cannot eliminate all risk from a probabilistic, fast-evolving system.

A certificate is valid for twelve months but depends on quarterly technical testing to stay current. Failing to maintain the quarterly cadence causes the certification to lapse before the twelve-month period ends.

No. AIUC-1 complements SOC 2, ISO 27001, and ISO 42001 rather than replacing them. It covers agent-specific risks those frameworks were not designed to address, and SOC 2 in particular remains expected for enterprise sales.

No. Customer service agents are a common use case because brand and customer data are on the line, but certified systems also include candidate scoring agents, internal automation agents, summarization agents, and image generation agents.

Through the audit and the insurance model rather than through regulation. Accredited auditors prepare reports, AIUC issues certificates and runs quarterly technical testing, and the insurance backing ties certification to real financial exposure if a certified agent fails.

Axipro Author

Picture of Pedro Dias

Pedro Dias

Pedro has been writing online for over 10 years. With experience in all things programming, cyber security, and compliance, he is our editor-in-chief at Axipro.

Blog Highlights

Explore More Articles

For the past two years, enterprise AI risk conversations have centered on a familiar set of concerns: model bias, hallucination, data privacy, and dependency on third-party models. These are real risks, and most organizations now run some version of a governance program to manage them. But something has shifted. Organizations are no longer just deploying AI that generates content for a human to review. They’re deploying AI that acts. Agents now plan multi-step tasks, call APIs, move data between systems, execute transactions, and coordinate with other agents, often with no human checkpoint in the loop. That shift deserves more than a footnote in the existing AI risk category. It deserves its own line in the risk register: Agentic Autonomy Risk. What Is Agentic AI Risk Management? Agentic AI risk management is the practice of identifying, assessing, and controlling the risks created when AI systems take autonomous action on an organization’s behalf. Where traditional AI governance evaluates outputs (accuracy, bias, privacy), agentic AI risk management governs what agents actually do: the tools they call, the permissions they inherit, and the downstream consequences of their actions. That distinction is the reason existing risk registers struggle with agents, and it’s worth unpacking properly. What Agentic AI Actually Changes Traditional AI systems, even generative ones, are advisory. They produce an output such as a summary, a prediction, a draft email, or a classification, and a human remains the last checkpoint before anything happens in the real world. Agentic AI removes that checkpoint. An agentic system doesn’t just produce an answer. It pursues a goal. It decides which tools to call and in what order, then executes those actions directly against live systems: submitting a purchase order, modifying a database record, sending an external communication, or orchestrating a set of sub-agents to complete a broader workflow. Agentic autonomy is the degree to which a system can plan and execute actions without a human explicitly authorizing each step. It’s a spectrum rather than a binary. At one end, the AI drafts and a human approves every action. At the other, the AI operates within broad guardrails and only escalates exceptions. The further an organization moves along that spectrum, the less its exposure looks like software risk and the more it looks like delegated authority risk, the kind normally reserved for employees, contractors, and automated financial systems. Why Existing Risk Registers Miss Agentic AI Risks Most enterprise risk registers were built on a reasonably safe assumption: a human initiates consequential actions, and the technology around that human behaves deterministically. Agentic AI breaks both halves of that assumption at once. A few specific gaps show up quickly when organizations try to map agentic deployments onto existing categories. Operational risk registers assume process failures come from human error or system outages, not from a system independently choosing an unanticipated path to a stated goal. Cybersecurity risk registers are built around unauthorized external access, while an agent problem usually involves an authorized system taking unauthorized internal actions with its own legitimate credentials. Model risk frameworks, borrowed largely from financial services, evaluate output accuracy rather than action consequences, which matters most when those actions can’t be reversed. And third-party risk assessments treat vendors as static entities, not as autonomous agents that might invoke other vendors’ agents on your behalf. See our guide to the NIST AI Risk Management Framework for how output-focused frameworks are structured. The result is a governance blind spot. An organization can be compliant against its AI policy, its cybersecurity policy, and its vendor risk policy, and still have nobody accountable for the specific risk of a system initiating a harmful sequence of actions before anyone notices. Defining Agentic Autonomy Risk Agentic Autonomy Risk is the risk that an AI system, operating with delegated decision-making and execution authority, takes actions that are harmful, non-compliant, or misaligned with organizational intent before adequate human oversight can intervene. Those actions might happen independently or in coordination with other agents. It deserves standing as a named category alongside cybersecurity, operational, legal, financial, and third-party risk because the loss event itself is different. The harm is a completed action in a live system, and it may be difficult or impossible to reverse. The accountability structure is different too: when an orchestrating agent delegates to sub-agents, responsibility for the outcome gets distributed in ways existing ownership models don’t cleanly capture. So is the detection window. Traditional controls assume a human is positioned to catch an error before it compounds, but an agent can execute dozens of dependent actions faster than any human review cycle. 7 Agentic AI Risk Scenarios to Put on Your Register 1. Unauthorized autonomous decision-making. An agent takes an action within its technical permissions but outside its intended business mandate. It adjusts pricing, approves a refund, or modifies a customer record, and no policy ever explicitly authorized that scenario. 2. Goal misalignment. The agent optimizes for a literal interpretation of its objective in a way that diverges from actual business intent, particularly under ambiguous or adversarial inputs. 3. Multi-agent interactions and cascading failures. One agent’s flawed output becomes another agent’s trusted input. A single error can propagate across a chain of agents faster than anyone can detect it, amplifying the original mistake instead of containing it. 4. Excessive tool or system permissions. Agents get provisioned with broad, standing access “to be safe” rather than scoped, least-privilege access tied to specific tasks. A productivity tool quietly becomes a privilege-escalation path. 5. Regulatory non-compliance. Autonomous actions trigger obligations under data protection, financial services, employment, or sector-specific regulation, and they execute without the compliance review a human-initiated process would normally receive. 6. Explainability and accountability gaps. An autonomous action causes harm and the organization can’t clearly reconstruct why the agent chose that path, or establish whether the business owner, the AI governance function, or the vendor is accountable for the outcome. 7. Autonomous third-party actions. A vendor’s agent, integrated into your environment, takes action on your behalf, or your agent acts against a

A SOC 2 penetration test costs between $1,000 and $30,000 for most companies. A typical SaaS scope, meaning one web application, its API layer, and the cloud infrastructure behind it, usually lands between $2,000 and $20,000. Early-stage startups with a narrow scope can get an auditor-accepted test for $1,000 to $8,000, while enterprises with multiple products and hybrid infrastructure regularly spend $20,000 to $50,000 or more. The spread is wide because “penetration test” covers everything from an automated scan with a cover page to weeks of manual testing by senior engineers. Auditors know the difference, and so do the enterprise customers who asked for your SOC 2 report in the first place. This guide breaks down what drives the price, where the hidden costs sit, and how to buy a test that holds up in fieldwork without overpaying for it. What Is SOC 2 Penetration Testing?​ A SOC 2 penetration test is a simulated attack on your systems, performed by a qualified security professional, scoped to the environment covered by your SOC 2 report. The tester tries to exploit real weaknesses the way an attacker would: broken access controls, injection flaws, misconfigured cloud services, exposed credentials. The output is a report your auditor reads as evidence that your security controls work in practice, not only on paper. That last part matters. A pentest bought for SOC 2 has a second audience beyond your security team. If the report doesn’t map findings to your audit scope, document its methodology, and show remediation, it fails the job you bought it for. We cover the full deliverable in our guide to what a SOC 2-ready VAPT report includes. How Penetration Testing Fits Into SOC 2 Compliance​ SOC 2 is built on the AICPA’s Trust Services Criteria, and the Security category (the Common Criteria) applies to every report. Penetration testing is the standard way to satisfy CC7.1, which expects you to detect and monitor for new vulnerabilities, and it supports CC4.1, which covers ongoing evaluations of whether controls actually function. The AICPA’s points of focus explicitly mention vulnerability scanning and penetration testing as examples of how companies meet these criteria. In practice, the test slots into your audit timeline as an evidence item. Your auditor will ask for the report, check the test date against the audit period, and review how you handled the findings. Remediation is often scrutinized harder than the test itself, because it shows whether your vulnerability management process runs or merely exists. Is Penetration Testing Required for SOC 2?​ Strictly speaking, no. The Trust Services Criteria never use the word “mandatory” about penetration testing. You could theoretically satisfy CC7.1 with vulnerability scanning and strong monitoring alone. In reality, almost every auditor expects one, and skipping it invites two problems. First, your auditor may push back during fieldwork or add exceptions to the report. Second, the enterprise buyers reviewing your SOC 2 report increasingly look for pentest evidence specifically, and a report without it raises questions during procurement. Treat the test as effectively required and budget for it from the start of your SOC 2 compliance checklist. How Much Does SOC 2 Penetration Testing Cost? Typical Price Range for SOC 2 Pen Testing Most companies pay $1,000 to $30,000, with the median engagement for a SaaS business sitting around $12,000 to $15,000. Compliance-focused tests at the lower end of the market start around $1,000 to $5,000. Deep manual testing from established firms runs $10,000 to $30,000. Anything quoted below roughly $3,000 is almost certainly automated scanning packaged as a pentest, which auditors are getting better at spotting. Cost by Company Size (Startup, SMB, Enterprise) Company size is a proxy, not the driver. A 15-person company with three products and a legacy on-prem component will pay more than a 200-person company with one tightly scoped SaaS platform. Testers price effort, and effort follows scope. Cost by Test Type (Network, Web App, API, Cloud, Internal/External) Most SOC 2 engagements bundle two or three of these. The common package for a cloud-native SaaS company is web app plus API plus cloud configuration, which is why the $1,000 to $20,000 band comes up so often. Companies with office networks and internal systems in their audit scope add internal network testing, and the price climbs accordingly. Factors That Influence SOC 2 Penetration Testing Cost Scope and Number of Assets Tested Scope is the single biggest cost driver. Every additional application, API endpoint group, cloud account, or network segment adds testing hours. A pentest priced without a scoping call is a pentest priced on guesswork, and the guess usually favors the vendor. Complexity of Application or Infrastructure​ A simple CRUD app with two user roles tests quickly. A multi-tenant platform with role hierarchies, workflow engines, file processing, and third-party integrations takes far longer, because each of those features creates attack surface a tester has to work through manually. Authentication tiers matter especially: every distinct role needs testing for privilege escalation and cross-tenant data access. Testing Methodology (Black Box, Grey Box, White Box) Black box testing gives the tester nothing but a URL, grey box adds credentials and documentation, and white box adds source code and architecture diagrams. Grey box is the default for SOC 2 and usually the best value, since the tester spends time exploiting rather than discovering. White box costs more upfront but finds deeper issues. Black box sounds rigorous but often wastes paid hours on reconnaissance an attacker would run for free. Depth of Testing and Manual vs. Automated Approaches Automated scanning finds known vulnerability patterns. Manual testing finds business logic flaws, chained exploits, and authorization gaps that no scanner catches, and it’s the part auditors and security-literate customers actually value. The ratio of manual work to automation is the honest explanation for most price differences between two quotes covering the same scope. Tester Credentials and Firm Reputation Senior testers holding OSCP, GPEN, or CREST credentials bill higher rates, and firms with recognized methodologies charge a premium for the credibility their letterhead carries

Two compromised versions of LiteLLM sat on PyPI for roughly 40 minutes on the morning of March 24, 2026. That window was enough to capture secrets from around 434,000 CI/CD pipeline runs across nearly 2,500 organizations, including AWS, Samsung, Cisco, Salesforce, Siemens, and Deloitte. In August, researchers at CloudSEK and Hudson Rock confirmed they had obtained the raw exfiltrated data: a 153GB archive containing 433,909 files of environment variables, cloud keys, Kubernetes secrets, and API tokens harvested live from running pipelines, as covered by Help Net Security’s reporting on the credential archive. If LiteLLM runs anywhere in your stack, or you touch any AI proxy infrastructure at all, you need answers to three things: whether you were exposed, what to rotate first, and whether the rotation you did back in March actually held. That last one matters more than it sounds, because “we rotated everything” has already burned at least one very large company. How the Breach Happened The attack didn’t start with LiteLLM. On March 19, 2026, a threat group called TeamPCP compromised the build pipeline of Trivy, a vulnerability scanner half the industry runs, and pushed a poisoned release. LiteLLM’s own CI pipeline ran Trivy, so the poisoned scanner had legitimate read access to the project’s runner environment. The attackers used that to steal LiteLLM’s PyPI publishing tokens and ship two malicious releases of their own: versions 1.82.7 and 1.82.8. KICS and the Telnyx Python SDK got hit in the same campaign. The payload design is the part worth studying. The malicious package dropped a .pth startup hook into site-packages, so the code ran the moment any Python interpreter started on the machine, whether or not anything imported LiteLLM. From there it harvested environment variables, read local credential files like .aws/credentials and .kube/config, tried to move laterally across Kubernetes clusters, and installed a systemd backdoor dressed up as a generic telemetry service. InfoQ’s coverage of the PyPI compromise put downloads of the compromised release above 40,000. For scale, LiteLLM normally gets downloaded around 3 million times a day. The exfiltration had a nasty fallback, too. According to CloudSEK, stolen data was encrypted and sent to a typosquatted domain, and when that failed, the malware created a public repository inside the victim’s own GitHub account and uploaded the loot as a release asset. Some companies were publishing their own secrets to the open internet and had no idea. Worth Knowing: The malicious code only existed in the PyPI artifacts. The GitHub source repository stayed clean the whole time, so a developer reviewing the code on GitHub saw nothing wrong. Source review isn’t artifact verification. If you don’t check that what the registry serves matches the upstream source, this class of attack is invisible to you. How to Check If You Were Exposed Three checks, from quickest to most involved. 1. Confirm whether the compromised versions ever ran The malicious versions went live on PyPI at 10:39 UTC on March 24, 2026 and got quarantined about 40 minutes later. The project’s advice: treat any install from that day before 16:00 UTC as suspect. Search your lockfiles, pip caches, SBOMs, and container image histories for 1.82.7 and 1.82.8. And check your internal artifact mirrors. An Artifactory or Nexus proxy that cached the bad release in March can keep serving it internally long after PyPI pulled it. Keep the .pth mechanism in mind when you scope this. The question isn’t “which applications import LiteLLM,” it’s “which machines had the package installed at all,” because every Python process on an infected machine triggered the payload. 2. Hunt for persistence Rotation is pointless if the attacker still has a foothold. Check developer machines, CI runners, and containers for unauthorized .pth files in site-packages and for suspicious systemd units, especially anything posing as a system telemetry service. And review activity from March 24 onward, not just the 40-minute window. Persistence is there so the access outlives the infection. Pro Tip: Don’t limit the persistence hunt to live machines. Base container images rebuilt in late March may have baked the payload into every image derived from them since. Scan your image registry for the affected LiteLLM versions and for unexpected .pth files, then trace which running workloads came from flagged images. 3. Check whether your secrets are in the dump Hudson Rock has published a domain lookup tool and is running ethical disclosures for affected organizations, and CloudSEK maintains a high-confidence victim list. Use them, but know their limits. Attribution in this dataset is genuinely hard. One dump with a siriusxm.com committer email actually traced, through its self-hosted GitLab endpoints, to AdsWizz, a SiriusXM subsidiary. And a large share of the dumps are generic pipeline configurations with no identifying domain, email, or server name at all. Absence from a victim list is not evidence of absence. If your pipelines ran the compromised versions, assume exposure no matter what a lookup tool tells you. What to Rotate, in What Order The guidance from both research teams is blunt: treat every secret the LiteLLM environment could reach as compromised. That covers secrets on disk, in memory, injected into CI jobs, and anything retrievable through instance metadata services. Work down by blast radius: Priority Credential type Why it comes first 1 Cloud IAM keys (AWS, GCP, Azure) Direct control of infrastructure, data stores, and billing. This is where attackers monetize fastest. 2 GitHub and GitLab PATs, package publishing tokens These let an attacker poison your releases and turn your company into the next link in the supply chain. 3 Kubernetes service account tokens and kubeconfigs Lateral movement across clusters was built into the payload, not a theoretical risk. 4 Database passwords and third-party API keys Dumped in plain text in the archive, often with no attribution, so nobody will warn you they leaked. 5 AI provider API keys Billing abuse, quota theft, and access to whatever data flows through your LLM routing layer. One word matters more than the rest of this article: revoke, don’t just rotate. That