Table of Contents

Reach SOC 2 Compliance in 6 Weeks or Less.

  / Business Continuity Plan Testing for SOC 2

Business Continuity Plan Testing for SOC 2

A business continuity plan that has never been tested is, to a SOC 2 auditor, a document and nothing more. The Availability criteria do not award credit for a polished plan sitting in a shared drive. They ask for evidence that you ran the plan, watched it work or fail, recorded what happened, and fixed what broke. That gap — between having a plan and proving it works — is where most availability findings originate.

Business continuity plan testing for SOC 2 is the exercise that turns your plan into auditable evidence. It maps directly to Availability criterion A1.3, one of the few SOC 2 controls that explicitly requires you to test something rather than merely document it. This guide covers what counts as a valid test, the test types auditors accept, a step-by-step process, the exact evidence you need, and the mistakes that turn a routine review into a finding.

Business Continuity Plan Testing for SOC 2

What Is Business Continuity Plan Testing in the Context of SOC 2?

Business continuity plan (BCP) testing is the structured validation of whether your organization can keep critical operations running — and restore them within defined targets — during a disruption. In a SOC 2 context, the testing is not freeform. It must produce dated, traceable evidence that the recovery procedures in your plan actually work, that the people involved know their roles, and that systems and data come back within your stated recovery objectives.

 

Why SOC 2 Requires Business Continuity Plan Testing

SOC 2 is an attestation against the AICPA’s Trust Services Criteria, and the Availability category exists specifically for organizations that make uptime or resilience commitments to customers. A plan you never exercise cannot demonstrate operating effectiveness over the audit period — which is the entire point of a Type 2 examination. Testing is the control that converts a static plan into a recurring, observable activity an auditor can sample.

Reach SOC 2 Compliance in 6 Weeks or Less

Schedule Your Free SOC 2 Assessment Today

SOC 2 Trust Services Criteria and BCP Testing Requirements

Availability is one of the five Trust Services Criteria, and it is optional, included only when your service commitments warrant it.

When in scope, it is built around three sub-criteria:

  • A1.1 addresses capacity management.
  • A1.2 addresses recovery infrastructure and backup processes.
  • A1.3 addresses the testing of recovery procedures.

BCP testing lives squarely in A1.3, with A1.2 supplying the backups and infrastructure that the test validates.

Availability Criteria A1.2 and A1.3 Explained

Per the AICPA’s Trust Services Criteria, A1.2 requires the entity to design, implement, operate, and monitor environmental protections, recovery infrastructure, and data backup processes that meet its availability objectives. In plain terms: you need real backups, stored away from production, with recovery infrastructure ready to use. A1.3 then requires the entity to test recovery plan procedures supporting system recovery to meet its objectives. The two work as a pair: A1.2 builds the capability, A1.3 proves it functions.

Important: The most common A1.3 gap is not a missing test. It is a test that never validated the recovery objectives. Teams run a tabletop, write “no issues found,” and move on — but the plan claims a 4-hour RTO that no one ever measured against an actual restore. If your plan states recovery targets, your test evidence must show whether you met them. A test that does not measure against your RTO and RPO leaves the most important question unanswered.

 

What Auditors Look for During a BCP Test Review

Auditors want proof that the test happened, proof that it was meaningful, and proof that it led somewhere. Concretely, that means a test plan with a defined scenario, a dated record of execution with participants, results measured against your recovery objectives, a list of gaps or issues found, and evidence that those issues were remediated. A test that finds nothing and changes nothing is treated with suspicion — because real tests almost always surface something.

 

Types of Business Continuity Plan Tests Accepted for SOC 2

SOC 2 does not mandate a specific test type. It expects the rigor of the test to match the criticality of what you are protecting. The four common approaches sit on a spectrum from low-effort, low-disruption to high-effort, high-assurance.

Tabletop Exercises

A tabletop exercise is a facilitated discussion where key personnel talk through a disruption scenario and their responses. It is cheap, fast, and excellent for confirming that people understand their roles and that the plan reads coherently. Its limit is obvious: nobody actually recovers anything. For many organizations a tabletop is a legitimate annual test, especially in the first audit cycle, but auditors expect more rigor as a program matures.

Walkthrough and Simulation Tests

A simulation applies a specific scenario and asks the team to perform recovery actions, not just describe them. It is more involved than a tabletop and far better at exposing the gaps that only appear when people touch the tools. Simulations are where teams discover that a runbook references a system that was decommissioned, or that the on-call engineer lacks the access the plan assumes.

Full Interruption Tests

A full interruption test shuts down primary systems and shifts operations entirely to the recovery environment. It is the most comprehensive validation available and the only one that proves your failover genuinely works end to end. It also carries real operational risk, so it demands thorough planning and is usually reserved for mature programs and the most critical systems.

Parallel Testing

Parallel testing activates recovery systems alongside production without taking the primary offline, then compares the two to confirm the recovery environment performs as expected. It delivers much of the assurance of a full interruption test while sparing the business the disruption. For most SaaS and cloud-hosted services, parallel testing of failover and restore is the sweet spot between confidence and risk.

8 Steps to Test Your BCP For SOC 2

How to Test Your Business Continuity Plan for SOC 2 Compliance

The sequence below aligns with the contingency planning process in NIST’s Contingency Planning Guide, SP 800-34, which auditors widely treat as authoritative for resilience practices. Each step produces an artifact, and the artifacts together form the evidence chain your auditor will sample.

Step 1: Define the Scope and Objectives of the BCP Test

Decide what the test covers — which systems and processes, which scenario, and what success looks like. Tie the objectives to measurable outcomes, such as restoring a specific service within its RTO. A vague objective like “test the plan” produces vague evidence; a specific one like “fail over the primary database and confirm recovery within 4 hours” produces evidence an auditor can verify.

Step 2: Identify Critical Business Processes and Recovery Priorities

Not everything recovers first. Identify the processes that must come back soonest and the order in which dependencies must be restored. This prioritization keeps the test focused on what actually matters to customers and to your service commitments, rather than spreading effort evenly across systems of unequal importance.

Step 3: Conduct a Business Impact Analysis Before Testing

A business impact analysis (BIA) is the foundation, and skipping it is why many plans test the wrong things. The BIA characterizes the consequences of losing each system over time and produces the numbers that drive everything else: Maximum Tolerable Downtime, RTO, and RPO. NIST is explicit that BIA results feed directly into contingency planning priorities, so run it before you design the test, not after.

Worth Knowing: NIST SP 800-34

NIST SP 800-34 defines three distinct outage measures that auditors expect you to keep straight. Maximum Tolerable Downtime (MTD) is the total outage the business can absorb. Recovery Time Objective (RTO) is the time to restore a system and must be shorter than the MTD. Recovery Point Objective (RPO) is about data, not time: how much data loss is acceptable, measured backward from the moment of failure. Confusing RTO with RPO in your documentation is a small error that signals to an auditor you may not have done the analysis.

Step 4: Assign Key Roles and Responsibilities for the Test

Name who runs the test, who participates, who observes, and who signs off. Pull in the functions a real disruption would involve: engineering, security, leadership, and, where relevant, legal and communications. Recording participants is not bureaucratic box-ticking — the attendee list is part of the evidence that the right people were exercised.

Step 5: Execute the BCP Test Scenario

Run the scenario as planned and let it play out honestly. Resist the urge to smooth over problems in the moment, because the problems are the point. Capture what happens in real time, including timestamps, decisions, and any deviation from the documented procedures.

Step 6: Document Test Results and Findings

Record what was tested, what happened, whether recovery objectives were met, and what gaps appeared. Measure results against the RTO and RPO from your BIA. This document is the single most important piece of A1.3 evidence, and it should read like an honest account, not a press release.

Step 7: Review, Remediate, and Update the Plan

Turn findings into assigned action items with owners and due dates, then update the plan to reflect what you learned. A test that exposes a broken runbook step and triggers a documented fix demonstrates a process that genuinely operates. Track remediation to completion — auditors will look for the close of the loop, not just the opening of it.

Step 8: Schedule Annual BCP Testing and Ongoing Reviews

Set a recurring cadence so testing is a program, not a one-off scramble before the audit. SP 800-34 recommends testing at least annually, with more frequent testing for high-impact systems. NIST 800-53 control CP-4 requires organizations to test plans at a defined frequency and document results. Annual is the floor; criticality and change drive anything more frequent.

Reach SOC 2 Compliance in 6 Weeks or Less

Schedule Your Free SOC 2 Assessment Today

Evidence Your SOC 2 Auditor Expects from BCP Testing

Availability is a heavily evidence-driven criterion, and A1.3 is among the most artifact-hungry. Four categories of evidence carry the weight.

Test Plans and Schedules

A documented test plan shows intent and scope: the scenario, objectives, systems in scope, and the date. A schedule shows the cadence is real and forward-looking, not improvised. Together they let the auditor see that testing is governed, not accidental.

Test Logs and Results Documentation

The results record is the heart of the evidence: what was executed, when, by whom, what happened, and whether recovery objectives were met. Timestamps matter enormously here, because evidence with no clear time reference is routinely challenged in a Type 2 review. Vague results are nearly as weak as no results.

Remediation Records and Corrective Actions

When a test finds a gap, the corrective action and its completion are evidence in their own right. They show the test produced improvement rather than sitting in a folder. A finding logged with an owner, a due date, and a closure note is exactly the trail auditors want to follow.

Sign-Off and Approval Documentation

A dated sign-off from an accountable owner closes the loop and demonstrates governance. It tells the auditor that leadership reviewed the test, accepted the results, and owns the follow-up. Without it, even a well-run test can look like an engineering side project rather than a managed control.

Pro Tip: Assemble a single "Test Package"

Assemble a single "test package" per exercise that contains the plan, the scenario, the participant list, the timestamped results measured against RTO and RPO, the findings, the remediation items, and the sign-off. When the auditor requests evidence for A1.3, you hand over one self-contained file instead of reconstructing the story from calendar invites and Slack threads. Teams that maintain this package almost never take an availability finding for missing or incomplete evidence. A compliance platform can make assembling and maintaining that package significantly less painful.

Common BCP Testing Findings That Impact SOC 2 Audits

Insufficient Testing Frequency

A single test years ago — or none within the audit period — is an immediate problem. Type 2 reports examine operating effectiveness across the whole period, so a test that predates the window does not count. Annual testing within the audit period is the baseline expectation.

Incomplete Documentation of Test Results

Teams frequently run a real test and then fail the control on documentation. If the results lack timestamps, omit whether RTO and RPO were met, or simply say “test successful” with no detail, the auditor cannot verify the control operated. Strong execution with weak records still produces an exception.

Failure to Test All Critical Business Functions

Testing only the easy systems, or only the ones that failed over cleanly last time, leaves critical functions unvalidated. Auditors check that the scope of testing matches the scope of your availability commitments. A plan that covers ten critical services but only ever tests two has a visible coverage gap.

Lack of Defined Recovery Time Objectives (RTOs) and Recovery Point Objectives (RPOs)

Without defined RTOs and RPOs, a test has no standard to measure against, and “recovery” becomes a matter of opinion. This is one of the most common root findings, because it undermines every test that follows. Define these objectives in your plan, derive them from your BIA, and measure every test against them.

Reach SOC 2 Compliance in 6 Weeks or Less

Schedule Your Free SOC 2 Assessment Today

How BCP Testing Integrates with Disaster Recovery Plan Testing for SOC 2

Key Differences Between BCP Testing and DRP Testing

Business continuity and disaster recovery are related but distinct, and conflating them muddies your evidence.

Business continuity keeps critical operations running during a disruption, covering people, processes, communications, and workarounds.

Disaster recovery is narrower, focused on restoring IT systems and data after an outage. Put simply: business continuity keeps the business operating; disaster recovery brings the technology back.

Aligning BCP and DRP Tests for a Unified SOC 2 Audit Signal

Auditors do not need separate ceremonies for each, and running them in isolation wastes effort. A single well-designed exercise can validate the business continuity response and the underlying disaster recovery in one pass: simulate the disruption, recover the IT systems, and confirm the business processes resume. Aligning them produces a cleaner, more coherent evidence story and shows the two plans actually interlock.

Backup Testing as Part of Your BCP Testing Strategy

Backups are the foundation that recovery depends on, and untested backups are a classic false comfort. A1.2 expects you to take backups and store them appropriately; A1.3 expects you to prove they restore. Include restore testing in your strategy and capture the evidence, because a backup that has never been restored is an assumption, not a control.

Insider Note: Auditors have learned to distinguish a backup test from a restore test, and they ask about the difference on purpose. Confirming that a backup job completed successfully proves the data was written. It says nothing about whether you can read it back, decrypt it, and stand up a working system. The teams that get tripped up are the ones showing green backup dashboards as A1.3 evidence. The dashboard belongs to A1.2; A1.3 wants the restore.

Best Practices for SOC 2 Business Continuity Plan Testing

Testing Frequency Recommendations

Test at least annually, and more often for high-impact systems or after any major change to architecture, staffing, or vendors. Treat a significant real incident as an unplanned test and document the lessons from it the same way. The cadence should be written into your plan so the expectation is unambiguous.

Maintaining Operational Resilience Between Tests

Resilience is not a once-a-year event. Keep runbooks current, validate that recovery access and credentials still work, and fold continuity considerations into change management so the plan does not silently drift out of date. The strongest programs treat the annual test as a checkpoint on continuous practice, not the only time anyone thinks about recovery.

Leveraging Compliance Tools to Streamline Evidence Collection

Manual evidence gathering is where good testing programs lose audit points — simply because artifacts get scattered. Centralizing test plans, results, remediation, and sign-offs in a compliance platform or a disciplined internal system of record keeps the evidence chain intact and retrievable. Compliance tools built for SOC 2 can automate much of this collection, reducing the risk that a well-run test goes undocumented simply because no one had time to file the paperwork.

Continuous Monitoring and Validation of BCP Controls

Pair periodic testing with ongoing validation: monitor backup completion, alert on failed jobs, and periodically verify recovery readiness rather than waiting for the annual exercise. Continuous monitoring strengthens the narrative across the whole audit period and catches drift early — which is precisely what a Type 2 examination is designed to assess.

Conclusion

Business continuity plan testing for SOC 2 succeeds or fails on evidence, not intentions. Define recovery objectives from a real BIA, choose a test type that matches the criticality of what you protect, run it honestly, measure results against your RTO and RPO, remediate what breaks, and capture the whole sequence with timestamps and sign-off. Map it to Availability A1.2 and A1.3, test at least annually within the audit window, and keep backup and restore validation in scope. Do that, and when the auditor asks to see your most recent continuity test, you can hand over a complete, dated, self-contained package — which is exactly what passing A1.3 looks like.

Frequently Asked Questions About Business Continuity Plan Testing for SOC 2

How Often Should a Business Continuity Plan Be Tested for SOC 2?

At least annually, and within the audit period for a Type 2 report. High-impact systems and organizations undergoing significant change should test more frequently. NIST SP 800-34 treats annual as the minimum baseline, with criticality driving anything more often.

A documented test plan, a dated record of execution with participants, results measured against your recovery objectives, a list of findings, remediation records showing those findings were closed, and a sign-off from an accountable owner. Timestamps throughout are essential, since undated evidence is routinely challenged.

A test that surfaces problems is not itself a failure — it is the system working as intended. What matters is whether you documented the gaps and remediated them. An honest test with tracked corrective actions strengthens your audit position, whereas a test that conveniently finds nothing tends to invite scrutiny.

Ownership typically sits with a named role such as a security or operations lead, with accountability extending to leadership through sign-off. Testing involves a cross-functional group: engineering, security, leadership, and where relevant legal and communications. The key is that ownership is explicitly assigned and that the assignment is reflected in the evidence.

Backup testing validates that data is being captured and can be restored, supporting A1.2. BCP testing is broader, validating that the organization can maintain and recover critical operations during a disruption, supporting A1.3. Restore testing is a component of a complete BCP testing strategy, not a substitute for it.

Often yes, particularly in an early audit cycle, since SOC 2 does not mandate a specific test type. A well-run, documented tabletop exercise with a clear scenario, findings, and follow-up can satisfy A1.3. As a program matures, auditors generally expect more rigorous testing — such as simulation or parallel tests — for critical systems.

Detailed enough that an auditor can reconstruct the test without asking you to narrate it: scope, scenario, date, participants, timestamped execution, results against RTO and RPO, findings, remediation, and sign-off. The standard to aim for is a self-contained record that answers the obvious follow-up questions before they are asked.

Axipro Author

Picture of Pedro Dias

Pedro Dias

Pedro has been writing online for over 10 years. With experience in all things programming, cyber security, and compliance, he is our editor-in-chief at Axipro.

Blog Highlights

Explore More Articles

For the past two years, enterprise AI risk conversations have centered on a familiar set of concerns: model bias, hallucination, data privacy, and dependency on third-party models. These are real risks, and most organizations now run some version of a governance program to manage them. But something has shifted. Organizations are no longer just deploying AI that generates content for a human to review. They’re deploying AI that acts. Agents now plan multi-step tasks, call APIs, move data between systems, execute transactions, and coordinate with other agents, often with no human checkpoint in the loop. That shift deserves more than a footnote in the existing AI risk category. It deserves its own line in the risk register: Agentic Autonomy Risk. What Is Agentic AI Risk Management? Agentic AI risk management is the practice of identifying, assessing, and controlling the risks created when AI systems take autonomous action on an organization’s behalf. Where traditional AI governance evaluates outputs (accuracy, bias, privacy), agentic AI risk management governs what agents actually do: the tools they call, the permissions they inherit, and the downstream consequences of their actions. That distinction is the reason existing risk registers struggle with agents, and it’s worth unpacking properly. What Agentic AI Actually Changes Traditional AI systems, even generative ones, are advisory. They produce an output such as a summary, a prediction, a draft email, or a classification, and a human remains the last checkpoint before anything happens in the real world. Agentic AI removes that checkpoint. An agentic system doesn’t just produce an answer. It pursues a goal. It decides which tools to call and in what order, then executes those actions directly against live systems: submitting a purchase order, modifying a database record, sending an external communication, or orchestrating a set of sub-agents to complete a broader workflow. Agentic autonomy is the degree to which a system can plan and execute actions without a human explicitly authorizing each step. It’s a spectrum rather than a binary. At one end, the AI drafts and a human approves every action. At the other, the AI operates within broad guardrails and only escalates exceptions. The further an organization moves along that spectrum, the less its exposure looks like software risk and the more it looks like delegated authority risk, the kind normally reserved for employees, contractors, and automated financial systems. Why Existing Risk Registers Miss Agentic AI Risks Most enterprise risk registers were built on a reasonably safe assumption: a human initiates consequential actions, and the technology around that human behaves deterministically. Agentic AI breaks both halves of that assumption at once. A few specific gaps show up quickly when organizations try to map agentic deployments onto existing categories. Operational risk registers assume process failures come from human error or system outages, not from a system independently choosing an unanticipated path to a stated goal. Cybersecurity risk registers are built around unauthorized external access, while an agent problem usually involves an authorized system taking unauthorized internal actions with its own legitimate credentials. Model risk frameworks, borrowed largely from financial services, evaluate output accuracy rather than action consequences, which matters most when those actions can’t be reversed. And third-party risk assessments treat vendors as static entities, not as autonomous agents that might invoke other vendors’ agents on your behalf. See our guide to the NIST AI Risk Management Framework for how output-focused frameworks are structured. The result is a governance blind spot. An organization can be compliant against its AI policy, its cybersecurity policy, and its vendor risk policy, and still have nobody accountable for the specific risk of a system initiating a harmful sequence of actions before anyone notices. Defining Agentic Autonomy Risk Agentic Autonomy Risk is the risk that an AI system, operating with delegated decision-making and execution authority, takes actions that are harmful, non-compliant, or misaligned with organizational intent before adequate human oversight can intervene. Those actions might happen independently or in coordination with other agents. It deserves standing as a named category alongside cybersecurity, operational, legal, financial, and third-party risk because the loss event itself is different. The harm is a completed action in a live system, and it may be difficult or impossible to reverse. The accountability structure is different too: when an orchestrating agent delegates to sub-agents, responsibility for the outcome gets distributed in ways existing ownership models don’t cleanly capture. So is the detection window. Traditional controls assume a human is positioned to catch an error before it compounds, but an agent can execute dozens of dependent actions faster than any human review cycle. 7 Agentic AI Risk Scenarios to Put on Your Register 1. Unauthorized autonomous decision-making. An agent takes an action within its technical permissions but outside its intended business mandate. It adjusts pricing, approves a refund, or modifies a customer record, and no policy ever explicitly authorized that scenario. 2. Goal misalignment. The agent optimizes for a literal interpretation of its objective in a way that diverges from actual business intent, particularly under ambiguous or adversarial inputs. 3. Multi-agent interactions and cascading failures. One agent’s flawed output becomes another agent’s trusted input. A single error can propagate across a chain of agents faster than anyone can detect it, amplifying the original mistake instead of containing it. 4. Excessive tool or system permissions. Agents get provisioned with broad, standing access “to be safe” rather than scoped, least-privilege access tied to specific tasks. A productivity tool quietly becomes a privilege-escalation path. 5. Regulatory non-compliance. Autonomous actions trigger obligations under data protection, financial services, employment, or sector-specific regulation, and they execute without the compliance review a human-initiated process would normally receive. 6. Explainability and accountability gaps. An autonomous action causes harm and the organization can’t clearly reconstruct why the agent chose that path, or establish whether the business owner, the AI governance function, or the vendor is accountable for the outcome. 7. Autonomous third-party actions. A vendor’s agent, integrated into your environment, takes action on your behalf, or your agent acts against a

A SOC 2 penetration test costs between $1,000 and $30,000 for most companies. A typical SaaS scope, meaning one web application, its API layer, and the cloud infrastructure behind it, usually lands between $2,000 and $20,000. Early-stage startups with a narrow scope can get an auditor-accepted test for $1,000 to $8,000, while enterprises with multiple products and hybrid infrastructure regularly spend $20,000 to $50,000 or more. The spread is wide because “penetration test” covers everything from an automated scan with a cover page to weeks of manual testing by senior engineers. Auditors know the difference, and so do the enterprise customers who asked for your SOC 2 report in the first place. This guide breaks down what drives the price, where the hidden costs sit, and how to buy a test that holds up in fieldwork without overpaying for it. What Is SOC 2 Penetration Testing?​ A SOC 2 penetration test is a simulated attack on your systems, performed by a qualified security professional, scoped to the environment covered by your SOC 2 report. The tester tries to exploit real weaknesses the way an attacker would: broken access controls, injection flaws, misconfigured cloud services, exposed credentials. The output is a report your auditor reads as evidence that your security controls work in practice, not only on paper. That last part matters. A pentest bought for SOC 2 has a second audience beyond your security team. If the report doesn’t map findings to your audit scope, document its methodology, and show remediation, it fails the job you bought it for. We cover the full deliverable in our guide to what a SOC 2-ready VAPT report includes. How Penetration Testing Fits Into SOC 2 Compliance​ SOC 2 is built on the AICPA’s Trust Services Criteria, and the Security category (the Common Criteria) applies to every report. Penetration testing is the standard way to satisfy CC7.1, which expects you to detect and monitor for new vulnerabilities, and it supports CC4.1, which covers ongoing evaluations of whether controls actually function. The AICPA’s points of focus explicitly mention vulnerability scanning and penetration testing as examples of how companies meet these criteria. In practice, the test slots into your audit timeline as an evidence item. Your auditor will ask for the report, check the test date against the audit period, and review how you handled the findings. Remediation is often scrutinized harder than the test itself, because it shows whether your vulnerability management process runs or merely exists. Is Penetration Testing Required for SOC 2?​ Strictly speaking, no. The Trust Services Criteria never use the word “mandatory” about penetration testing. You could theoretically satisfy CC7.1 with vulnerability scanning and strong monitoring alone. In reality, almost every auditor expects one, and skipping it invites two problems. First, your auditor may push back during fieldwork or add exceptions to the report. Second, the enterprise buyers reviewing your SOC 2 report increasingly look for pentest evidence specifically, and a report without it raises questions during procurement. Treat the test as effectively required and budget for it from the start of your SOC 2 compliance checklist. How Much Does SOC 2 Penetration Testing Cost? Typical Price Range for SOC 2 Pen Testing Most companies pay $1,000 to $30,000, with the median engagement for a SaaS business sitting around $12,000 to $15,000. Compliance-focused tests at the lower end of the market start around $1,000 to $5,000. Deep manual testing from established firms runs $10,000 to $30,000. Anything quoted below roughly $3,000 is almost certainly automated scanning packaged as a pentest, which auditors are getting better at spotting. Cost by Company Size (Startup, SMB, Enterprise) Company size is a proxy, not the driver. A 15-person company with three products and a legacy on-prem component will pay more than a 200-person company with one tightly scoped SaaS platform. Testers price effort, and effort follows scope. Cost by Test Type (Network, Web App, API, Cloud, Internal/External) Most SOC 2 engagements bundle two or three of these. The common package for a cloud-native SaaS company is web app plus API plus cloud configuration, which is why the $1,000 to $20,000 band comes up so often. Companies with office networks and internal systems in their audit scope add internal network testing, and the price climbs accordingly. Factors That Influence SOC 2 Penetration Testing Cost Scope and Number of Assets Tested Scope is the single biggest cost driver. Every additional application, API endpoint group, cloud account, or network segment adds testing hours. A pentest priced without a scoping call is a pentest priced on guesswork, and the guess usually favors the vendor. Complexity of Application or Infrastructure​ A simple CRUD app with two user roles tests quickly. A multi-tenant platform with role hierarchies, workflow engines, file processing, and third-party integrations takes far longer, because each of those features creates attack surface a tester has to work through manually. Authentication tiers matter especially: every distinct role needs testing for privilege escalation and cross-tenant data access. Testing Methodology (Black Box, Grey Box, White Box) Black box testing gives the tester nothing but a URL, grey box adds credentials and documentation, and white box adds source code and architecture diagrams. Grey box is the default for SOC 2 and usually the best value, since the tester spends time exploiting rather than discovering. White box costs more upfront but finds deeper issues. Black box sounds rigorous but often wastes paid hours on reconnaissance an attacker would run for free. Depth of Testing and Manual vs. Automated Approaches Automated scanning finds known vulnerability patterns. Manual testing finds business logic flaws, chained exploits, and authorization gaps that no scanner catches, and it’s the part auditors and security-literate customers actually value. The ratio of manual work to automation is the honest explanation for most price differences between two quotes covering the same scope. Tester Credentials and Firm Reputation Senior testers holding OSCP, GPEN, or CREST credentials bill higher rates, and firms with recognized methodologies charge a premium for the credibility their letterhead carries

Two compromised versions of LiteLLM sat on PyPI for roughly 40 minutes on the morning of March 24, 2026. That window was enough to capture secrets from around 434,000 CI/CD pipeline runs across nearly 2,500 organizations, including AWS, Samsung, Cisco, Salesforce, Siemens, and Deloitte. In August, researchers at CloudSEK and Hudson Rock confirmed they had obtained the raw exfiltrated data: a 153GB archive containing 433,909 files of environment variables, cloud keys, Kubernetes secrets, and API tokens harvested live from running pipelines, as covered by Help Net Security’s reporting on the credential archive. If LiteLLM runs anywhere in your stack, or you touch any AI proxy infrastructure at all, you need answers to three things: whether you were exposed, what to rotate first, and whether the rotation you did back in March actually held. That last one matters more than it sounds, because “we rotated everything” has already burned at least one very large company. How the Breach Happened The attack didn’t start with LiteLLM. On March 19, 2026, a threat group called TeamPCP compromised the build pipeline of Trivy, a vulnerability scanner half the industry runs, and pushed a poisoned release. LiteLLM’s own CI pipeline ran Trivy, so the poisoned scanner had legitimate read access to the project’s runner environment. The attackers used that to steal LiteLLM’s PyPI publishing tokens and ship two malicious releases of their own: versions 1.82.7 and 1.82.8. KICS and the Telnyx Python SDK got hit in the same campaign. The payload design is the part worth studying. The malicious package dropped a .pth startup hook into site-packages, so the code ran the moment any Python interpreter started on the machine, whether or not anything imported LiteLLM. From there it harvested environment variables, read local credential files like .aws/credentials and .kube/config, tried to move laterally across Kubernetes clusters, and installed a systemd backdoor dressed up as a generic telemetry service. InfoQ’s coverage of the PyPI compromise put downloads of the compromised release above 40,000. For scale, LiteLLM normally gets downloaded around 3 million times a day. The exfiltration had a nasty fallback, too. According to CloudSEK, stolen data was encrypted and sent to a typosquatted domain, and when that failed, the malware created a public repository inside the victim’s own GitHub account and uploaded the loot as a release asset. Some companies were publishing their own secrets to the open internet and had no idea. Worth Knowing: The malicious code only existed in the PyPI artifacts. The GitHub source repository stayed clean the whole time, so a developer reviewing the code on GitHub saw nothing wrong. Source review isn’t artifact verification. If you don’t check that what the registry serves matches the upstream source, this class of attack is invisible to you. How to Check If You Were Exposed Three checks, from quickest to most involved. 1. Confirm whether the compromised versions ever ran The malicious versions went live on PyPI at 10:39 UTC on March 24, 2026 and got quarantined about 40 minutes later. The project’s advice: treat any install from that day before 16:00 UTC as suspect. Search your lockfiles, pip caches, SBOMs, and container image histories for 1.82.7 and 1.82.8. And check your internal artifact mirrors. An Artifactory or Nexus proxy that cached the bad release in March can keep serving it internally long after PyPI pulled it. Keep the .pth mechanism in mind when you scope this. The question isn’t “which applications import LiteLLM,” it’s “which machines had the package installed at all,” because every Python process on an infected machine triggered the payload. 2. Hunt for persistence Rotation is pointless if the attacker still has a foothold. Check developer machines, CI runners, and containers for unauthorized .pth files in site-packages and for suspicious systemd units, especially anything posing as a system telemetry service. And review activity from March 24 onward, not just the 40-minute window. Persistence is there so the access outlives the infection. Pro Tip: Don’t limit the persistence hunt to live machines. Base container images rebuilt in late March may have baked the payload into every image derived from them since. Scan your image registry for the affected LiteLLM versions and for unexpected .pth files, then trace which running workloads came from flagged images. 3. Check whether your secrets are in the dump Hudson Rock has published a domain lookup tool and is running ethical disclosures for affected organizations, and CloudSEK maintains a high-confidence victim list. Use them, but know their limits. Attribution in this dataset is genuinely hard. One dump with a siriusxm.com committer email actually traced, through its self-hosted GitLab endpoints, to AdsWizz, a SiriusXM subsidiary. And a large share of the dumps are generic pipeline configurations with no identifying domain, email, or server name at all. Absence from a victim list is not evidence of absence. If your pipelines ran the compromised versions, assume exposure no matter what a lookup tool tells you. What to Rotate, in What Order The guidance from both research teams is blunt: treat every secret the LiteLLM environment could reach as compromised. That covers secrets on disk, in memory, injected into CI jobs, and anything retrievable through instance metadata services. Work down by blast radius: Priority Credential type Why it comes first 1 Cloud IAM keys (AWS, GCP, Azure) Direct control of infrastructure, data stores, and billing. This is where attackers monetize fastest. 2 GitHub and GitLab PATs, package publishing tokens These let an attacker poison your releases and turn your company into the next link in the supply chain. 3 Kubernetes service account tokens and kubeconfigs Lateral movement across clusters was built into the payload, not a theoretical risk. 4 Database passwords and third-party API keys Dumped in plain text in the archive, often with no attribution, so nobody will warn you they leaked. 5 AI provider API keys Billing abuse, quota theft, and access to whatever data flows through your LLM routing layer. One word matters more than the rest of this article: revoke, don’t just rotate. That