Table of Contents

Reach SOC 2 Compliance in 6 Weeks or Less.

  /

  / SOC 2 Runbook: A Complete Guide

SOC 2 Runbook: A Complete Guide

A well-built SOC 2 runbook is the difference between a finding and a clean opinion. It converts the abstract language of a control into a sequence of actions someone actually performed, in a verifiable order, with a paper trail attached.

Auditors do not fail companies for having incidents. They fail them for not being able to prove how those incidents were handled.

This guide shows you how to build a runbook that holds up under scrutiny — covering what a SOC 2 runbook is, what makes it audit-ready, how it differs from a playbook, the components every runbook should include, the control areas where runbooks are expected, and how to keep them current between annual examinations.

SOC 2 Runbook: A Complete Guide

What Is a SOC 2 Runbook?

A SOC 2 runbook is a documented, repeatable procedure that operationalises a specific SOC 2 control. Where a policy states what must happen and why, a runbook states exactly how: the trigger, the steps, the people, the systems touched, the evidence captured, and the sign-off that closes it out.

Runbooks live closest to the engineers and operations staff actually doing the work. They are the layer auditors care about most because they are where the control either operates or fails. A well-written runbook turns a control objective into something testable, traceable, and survivable across staff turnover.

SOC 2 Runbook vs. SOC 2 Playbook: Key Differences

The terms get used interchangeably, but they describe two different artefacts. The cleanest distinction is scope and audience.

DimensionRunbookPlaybook
ScopeOne specific procedureMulti-step strategy across functions
AudienceEngineers, on-call responders, operations teamsLeadership, legal, communications, incident response coordinators
Detail LevelCommands, queries, exact toolingDecisions, escalation paths, stakeholder roles
ExampleIsolating an affected EC2 instance using a documented AWS CLI commandCoordinating a ransomware response across legal, PR, and law enforcement
LengthShort, tactical, and scannableLonger, narrative, and decision-oriented

A mature SOC 2 programme uses both. The playbook frames the response. The runbook executes pieces of it.

Why SOC 2 Auditors Expect Runbooks

The AICPA’s Trust Services Criteria describe what auditors test, but at the level of objectives, not procedures. CC7.3 says you must respond to security incidents. It does not tell you how. The runbook is your answer to how.

Auditors are looking for two things when they evaluate a control: that it was designed appropriately, and that it operated effectively across the audit period. Runbooks are how you show both. The document itself is the design. The completed runbook artefacts (tickets, logs, sign-offs, post-mortems) are the operating evidence.

Which SOC 2 Trust Services Criteria Require Runbook Documentation

Every Common Criteria area benefits from runbooks, but the strongest expectation sits in CC6 (logical and physical access), CC7 (system operations, including incident detection and response), CC8 (change management), and CC9 (risk mitigation, vendor management, and BCP/DR). For a deeper look at how these criteria are structured and what auditors are actually testing, the Trust Services Criteria breakdown is worth reading before you start mapping your runbooks.

If your scope includes the Availability criteria, A1.2 and A1.3 will require runbooks for failover, restoration, and capacity management. Confidentiality and Privacy add data handling and retention runbooks on top. If you are still determining which criteria apply to your organisation, a structured gap analysis is the most reliable starting point.

Reach SOC 2 Compliance in 6 Weeks or Less

Schedule Your Free SOC 2 Assessment Today

Why Your Organization Needs a SOC 2 Runbook

The common failure pattern is not the absence of policies. It is the absence of a credible bridge between the policy and what people actually do at 2am during an incident.

How Runbooks Demonstrate Control Effectiveness to Auditors

Auditors sample. For a Type II report covering twelve months, they will pull a population of incidents, changes, access reviews, or vendor onboardings, and trace a sample of them end to end. Without runbooks, that trace usually breaks. Engineers describe what they did from memory, ticket histories are inconsistent, and the auditor has no baseline to test against.

With runbooks, the auditor compares the documented steps to what actually happened in the artefacts. If the runbook says approval is required, the ticket should show it. If it says evidence must be retained for ninety days, the log should be there. The runbook turns a subjective conversation into an objective trace.

Runbooks as Evidence: Avoiding the Audit Evidence Trap

A specific failure mode is what practitioners call the evidence trap: the control exists, the team is doing the right thing, but nothing was captured at the time. Three months later, the SIEM has rotated the logs, the on-call engineer has left, and the only record is a Slack thread no one can find.

Runbooks prevent this when they make evidence capture a step in the procedure itself, not an afterthought. A line in the runbook that reads export the relevant CloudTrail entries to the incident folder before remediation is what stands between you and a qualified opinion.

Pro Tip: Build evidence capture into the runbook as a numbered step, not a footer note. Auditors test what is written. If “save the screenshot” is step 7, it gets done. If it is buried in a paragraph at the bottom, it usually does not.

SOC 2 Type I vs. Type II: How Runbooks Support Each

A SOC 2 Type I report assesses the design of controls at a single point in time. For Type I, the runbook itself, together with the policies it references, is most of what auditors need.

Type II is a different beast. It tests operating effectiveness over a period (typically six to twelve months), and that is where runbooks earn their keep. Each completed run produces evidence: a ticket, a log entry, a screenshot, a signed approval.

Over twelve months those artefacts become the case for control effectiveness. Without runbooks, evidence collection is reactive and full of gaps. With them, it is a byproduct of normal work.

For a fuller picture of what to expect across both report types, the SOC 2 compliance checklist is a useful companion to this guide.

 

Core Components of a SOC 2 Runbook

Runbook formats vary, but the audit-ready ones share a common skeleton. Skip any of these and the runbook becomes harder to defend.

Scope, Trigger Conditions, and Impact Classification

Every runbook should open with what it covers, what kicks it off, and how to grade severity. Trigger examples include a PagerDuty alert from the production logging pipeline, a quarterly access review due date, or a deployment to the production environment. Severity classification (P1 through P4, or critical, high, medium, low) determines escalation timing and approval thresholds downstream.

Roles, Responsibilities, and RACI Framework

Name the roles, not the people. Incident Commander, Communications Lead, Subject Matter Expert, Approver. A RACI table (Responsible, Accountable, Consulted, Informed) prevents the standard audit finding where two people thought the other was reviewing access changes and neither did.

Step-by-Step Procedures Mapped to SOC 2 Controls

The body of the runbook. Each step should be atomic, verifiable, and tied to the SOC 2 control it supports. Disable the user account in Okta (CC6.1, CC6.2) is better than remove access. The control mapping is the line that lets your auditor walk from the runbook back to the control matrix without guessing.

Communication and Escalation Paths

When does this go up the chain, to whom, and through which channel? Auditors will ask, and the answer “we’d Slack the team lead” is not enough. Document the channel, the roles notified at each severity, and the timing.

Evidence Collection and Audit Trail Requirements

What gets captured, where it gets stored, in what format, and for how long. Tickets, logs, screenshots, approvals, post-incident notes. This is the single most under-documented section in most runbooks and the single most-tested.

Resolution Verification and Sign-Off

How do you know the runbook is done, and who confirms it? A closure step with a named approver is what distinguishes a finished procedure from an open thread.

Reach SOC 2 Compliance in 6 Weeks or Less

Schedule Your Free SOC 2 Assessment Today

SOC 2 Runbook Templates by Control Category

The following are the runbook archetypes most organisations need. These map to the bulk of the Common Criteria and the optional categories most often selected.

Security Incident Response Runbook Template

Anchored to CC7.3 and CC7.4. Should cover detection, triage and severity assignment, containment, eradication, recovery, evidence preservation, communication (internal and external), and post-incident review. The structure aligns naturally with the lifecycle described in NIST SP 800-61 Revision 3, which most auditors are familiar with.

Access Control and Logical Access Runbook Template

Covers CC6.1 through CC6.3. Three operational flows belong here: provisioning (joiners), modification (movers), and deprovisioning (leavers). The deprovisioning runbook is the single most-sampled control in SOC 2 audits because the failure pattern is so common, so build this one with extra rigour. Tie the trigger to the HRIS termination event, name the systems where access must be revoked, and require timestamped confirmation.

Change Management Runbook Template

Maps to CC8.1. Production change request, peer review, approval, deployment, post-deployment verification, rollback procedure. Tie it to your version control and ticketing systems so the artefacts (PR, ticket, deployment log) are the evidence.

Availability and Business Continuity Runbook Template

If Availability is in scope, A1.2 and A1.3 demand documented and tested procedures for failover, restoration, and capacity adjustments. Each scenario, whether a regional outage, a database failure, or a dependency degradation, should have its own runbook with named recovery time objectives.

Vendor and Third-Party Risk Management Runbook Template

CC9.2 territory. Onboarding (security review, contract execution, access grant), ongoing monitoring (annual review, SOC 2 collection), offboarding (access revocation, data deletion confirmation). One inventory, one runbook for each phase, one cadence.

Data Privacy and Confidentiality Runbook Template

If Confidentiality or Privacy are in scope, you need runbooks for data classification, encryption verification, retention enforcement, and deletion or subject access requests. These are increasingly cross-referenced with GDPR, CCPA, and other privacy obligations.

Pro Tip: Writing a Runbox Index

Auditors form an opinion on your programme in the first thirty minutes of fieldwork. The fastest way to set a positive tone is a clean runbook index that maps every Common Criteria control to a specific runbook by name. The map itself is not strictly required, but it transforms how the rest of the audit unfolds.

Build SOC 2 runbook in 7 Steps

How to Build a SOC 2 Runbook Step by Step

A practical sequence for going from blank document to audit-ready procedure.

Step 1: Map Runbooks to SOC 2 Common Criteria Controls

Start with the control matrix, not a blank page. Pull your in-scope criteria, list every control, and identify which require an operational procedure. Most CC6, CC7, CC8, and CC9 controls do. Some CC1 and CC2 controls (governance, communication) need policies more than runbooks.

Step 2: Define the Scope and Trigger for Each Runbook

A runbook with no defined trigger gets used inconsistently or not at all. Be specific. Triggered by a Sev-1 PagerDuty alert, or triggered by a manager-submitted termination ticket in Workday. Vague triggers are why controls drift.

Step 3: Document Exact Procedures with Audit-Ready Detail

Write the steps in active voice, in the order they happen, with the system and the action both named. Run the access revocation script in Okta rather than remove access. If a step requires judgment, say so explicitly, and name who exercises it.

Step 4: Incorporate Approval and Authorization Checkpoints

Auditors test segregation of duties hard. Every runbook that touches production data, access rights, or financial systems should have a named approver who is not the same person executing the change. Document who can approve and what they are confirming.

Step 5: Build in Automated Evidence Capture

Where possible, let the tooling generate the evidence. Ticketing systems, SIEM alerts, deployment logs, and identity providers all produce timestamped artefacts that survive audit scrutiny better than manual screenshots. The runbook should call out which artefact each step generates.

Step 6: Validate and Test the Runbook Against Real Scenarios

A runbook nobody has executed is fiction. Tabletop exercises, dry runs, and chaos drills each surface problems before the auditor does. Document the testing itself: who participated, what scenarios were used, what gaps were found, and what changed as a result.

Step 7: Establish a Review and Update Lifecycle

Set a cadence, annual at minimum, semi-annual for high-volume runbooks like incident response, and assign an owner. Lifecycle expectations vary by control area, but every runbook needs a last reviewed date that auditors can verify is recent.

Important: The most common SOC 2 audit finding around runbooks is not that they do not exist. It is that the runbook on file does not reflect what the team actually does. When the auditor compares the runbook to the artefacts and finds mismatched steps, that is a control deficiency regardless of how good the procedure looks on paper.

 

Mapping SOC 2 Controls to Runbook Sections

The Common Criteria are organised into nine series, CC1 through CC9. The runbook map below covers where the operational expectation is heaviest.

Organization and Management Controls

CC1 covers governance, ethics, and accountability. Runbook content here is light; most evidence is policy and meeting minutes. The exception is the annual control attestation runbook, which formalises how leadership confirms the control environment each year.

Human Resources and Access Management Controls

CC1.4 (HR), CC2 (communication), and the entirety of CC6 (logical access). This is where runbooks earn their highest leverage: joiner, mover, leaver flows tied directly to your HRIS and identity provider, plus periodic access reviews.

Network, Infrastructure, and Physical Security Controls

CC6.4 through CC6.8, plus relevant points of focus under CC7. Runbooks belong here for firewall change management, cloud configuration baselines, and data centre access where applicable.

System Operations and Monitoring Controls

CC7 in full. Detection (alert tuning, threshold reviews), monitoring (log review cadence), and incident response. The detection and response runbooks here are usually the most heavily tested in any SOC 2 examination, and they benefit most from the kind of continuous monitoring for SOC 2 that produces a steady stream of artefacts rather than a burst of activity at audit time.

Change Management Controls

CC8.1. One core runbook for production changes, with sub-runbooks for emergency changes, schema migrations, infrastructure-as-code deployments, and any class of change that follows a different approval path.

Disaster Recovery and Business Continuity Controls

CC9 plus A1.2 and A1.3 if Availability is in scope. Failover, restoration from backup, and tabletop exercise runbooks. The tabletop runbook is itself an artefact: it documents how the exercise was conducted, who participated, and what was learned.

 

Keeping SOC 2 Runbooks Audit-Ready Year-Round

A runbook is only as good as its currency. The question to optimise for is not do we have a runbook, but does the runbook reflect reality.

Runbook Lifecycle Management and Version Control

Store runbooks in a system with version history. Confluence, Notion, GitBook, and Git-based wikis all work, with the caveat that whatever you choose, the version history needs to be exportable for auditors. Each material change should be reviewed and approved before publication.

How Often to Review and Update SOC 2 Runbooks

Annually is the floor for low-volume runbooks (vendor offboarding, BCP drills). Semi-annually fits most operational runbooks. After every material incident or change, the relevant runbook should get an immediate review even if the cadence does not require it. Tooling changes, organisational restructures, and new compliance scope all trigger ad hoc reviews.

Common Pitfalls That Fail SOC 2 Audits

Runbook content that contradicts what artefacts show. Steps that reference deprecated tools. Approver roles assigned to people who left. RACI tables with unfilled cells. Evidence requirements that name a system the team has migrated away from. Each of these is fixable in an afternoon, and each one, left alone, becomes a finding.

Key Metrics to Track Runbook Effectiveness

Time-to-acknowledge and time-to-resolve are the operational measures. For audit purposes, the more useful metrics are runbook adherence rate (how often the documented steps were followed) and evidence completeness (how often the expected artefacts were captured). Both can be tracked from your ticketing system without specialised tooling.

When to Retire or Replace a Runbook

Retire a runbook when the underlying control is removed from scope or when the procedure has been wholly absorbed by automation. Replace a runbook, rather than patching it, when the underlying tooling has changed enough that a step-by-step rewrite is faster than incremental edits. Document the retirement decision and retain the previous version for the audit period.

Reach SOC 2 Compliance in 6 Weeks or Less

Schedule Your Free SOC 2 Assessment Today

Automating SOC 2 Runbooks

The maturity curve runs from manual procedures to automated workflows with human-in-the-loop approvals. Most organisations sit somewhere in the middle and benefit from moving further along.

Benefits of Runbook Automation for SOC 2 Compliance

Automation enforces the runbook. A manual runbook can be skipped or executed inconsistently. An automated workflow runs the same way every time, captures evidence as it goes, and produces an immutable timeline. For Type II audits, that consistency is exactly what auditors are testing.

How Automated Evidence Collection Supports SOC 2 Audits

Automated runbooks generate evidence as a byproduct: timestamped logs, deployment records, access revocation confirmations, ticket state transitions. This eliminates the screenshot-scramble at audit time and produces a record that is more credible than manual evidence because it was captured at the moment of action, by the system, not reconstructed from memory.

GRC platforms that integrate with your identity provider, ticketing system, and cloud infrastructure can aggregate this evidence automatically — turning what was once a weeks-long evidence request process into a handful of system exports.

Tools like Drata are purpose-built for this, and a detailed comparison of leading options is available in the Drata vs Vanta breakdown.

Action-Level Approvals and Authorization Controls in Automated Runbooks

Automation does not mean removing approvers. It means embedding them as workflow gates. A production change deployment workflow can pause for explicit approval, log the approver’s identity, and only proceed once the gate is passed. The audit trail that emerges is stronger than the manual equivalent.

Tools and Integrations for SOC 2 Runbook Automation

Categories matter more than vendors here: identity providers for access automation, GRC platforms for control mapping and evidence aggregation, incident response platforms for response runbook execution, infrastructure-as-code for change management, and SIEM for detection and monitoring. The integration story matters more than any single tool: each runbook should produce evidence that flows into a central evidence repository.

Worth Knowing: SOC 2 Type II reports require demonstrating operating effectiveness over six to twelve months. Organisations with well-automated runbooks regularly complete audit fieldwork in weeks, not months. The gap between those two timelines is almost entirely explained by evidence readiness.

 

SOC 2 Runbook Best Practices

The principles that separate runbooks that work from runbooks that look good in a binder.

Writing Runbooks That Are Actionable and Accurate

Each step should be a verb plus an object plus a system. Disable the account in Okta. Export the alert log from Sentinel. Open a ticket in the SecOps queue. Narrative paragraphs slow people down and create ambiguity. Lists of explicit actions speed them up.

Making Runbooks Accessible to All Relevant Team Members

Runbooks help nobody if they are buried in a wiki nobody opens at 3am. Link them from the alerts that trigger them. Embed them in your incident management tool. Put them in the channels the on-call engineer is already in. Accessibility is a control quality, not a nice-to-have.

Standardizing Runbooks Across Teams and Environments

A common template across all runbooks reduces cognitive load and makes audit traceability simpler. Same headings, same RACI structure, same evidence section, same review cadence. Teams can specialise within the template, but they should not invent their own structure.

Incident Documentation Best Practices for SOC 2

Capture timestamps for every action. Record who took it, what tool was used, and what the resulting artefact is. Close every incident with a short post-mortem that includes detection time, response time, root cause, remediation, and runbook gaps surfaced. The post-mortem is itself audit evidence and feeds the next round of runbook updates.

Are runbooks required for SOC 2 compliance?

The AICPA’s Trust Services Criteria do not use the word “runbook.” They require documented and operating procedures for the in-scope controls. In practice, auditors expect to see runbooks (or their functional equivalent) for every operational control, particularly under CC6, CC7, CC8, and CC9. Calling them runbooks, SOPs, or response procedures is a stylistic choice. Having them is not.

A policy states the requirement: all production access changes require approval. A runbook executes the requirement: here is exactly how that approval happens, who grants it, what gets logged, and how it is verified. Policies set the rule. Runbooks are how the rule operates day to day.

Detailed enough that someone unfamiliar with the system could execute it with the runbook in front of them and produce the expected evidence. If a step assumes tribal knowledge (“you know which queue to use”), the runbook is too thin.

A continuous compliance posture relies on controls operating consistently between audits, not just during them. Runbooks are the mechanism that makes that consistency possible. When evidence capture is built into the procedure, every execution generates audit-ready artefacts. By the time the next audit window opens, the evidence is already there — collected as a byproduct of normal operations rather than assembled under pressure. This is the foundation of effective continuous monitoring for SOC 2.

Yes, and well-built runbooks often do. A change management runbook can satisfy CC8.1 (change controls), CC6.6 (segregation of duties), and A1.2 (availability commitments) all at once. The control mapping section is what makes the multi-purpose use defensible.

Through the artefacts the runbook generates. Tickets, logs, approvals, screenshots, and post-incident notes are what auditors sample. The runbook is the design document. The artefacts are the operating evidence. A runbook with no corresponding artefacts is a control finding waiting to happen.

The auditor compares the runbook to the artefacts. If the artefacts show a different procedure than what is documented, the control is operating inconsistently with its design, which is a deficiency. The fix is twofold: update the runbook to match current practice, and document the gap in your management response. For a fuller view of how SOC 2 sits alongside other attestation standards, the SOC framework background covers the broader context.

Axipro Author

Picture of Pedro Dias

Pedro Dias

Pedro has been writing online for over 10 years. With experience in all things programming, cyber security, and compliance, he is our editor-in-chief at Axipro.

Blog Highlights

Explore More Articles

For the past two years, enterprise AI risk conversations have centered on a familiar set of concerns: model bias, hallucination, data privacy, and dependency on third-party models. These are real risks, and most organizations now run some version of a governance program to manage them. But something has shifted. Organizations are no longer just deploying AI that generates content for a human to review. They’re deploying AI that acts. Agents now plan multi-step tasks, call APIs, move data between systems, execute transactions, and coordinate with other agents, often with no human checkpoint in the loop. That shift deserves more than a footnote in the existing AI risk category. It deserves its own line in the risk register: Agentic Autonomy Risk. What Is Agentic AI Risk Management? Agentic AI risk management is the practice of identifying, assessing, and controlling the risks created when AI systems take autonomous action on an organization’s behalf. Where traditional AI governance evaluates outputs (accuracy, bias, privacy), agentic AI risk management governs what agents actually do: the tools they call, the permissions they inherit, and the downstream consequences of their actions. That distinction is the reason existing risk registers struggle with agents, and it’s worth unpacking properly. What Agentic AI Actually Changes Traditional AI systems, even generative ones, are advisory. They produce an output such as a summary, a prediction, a draft email, or a classification, and a human remains the last checkpoint before anything happens in the real world. Agentic AI removes that checkpoint. An agentic system doesn’t just produce an answer. It pursues a goal. It decides which tools to call and in what order, then executes those actions directly against live systems: submitting a purchase order, modifying a database record, sending an external communication, or orchestrating a set of sub-agents to complete a broader workflow. Agentic autonomy is the degree to which a system can plan and execute actions without a human explicitly authorizing each step. It’s a spectrum rather than a binary. At one end, the AI drafts and a human approves every action. At the other, the AI operates within broad guardrails and only escalates exceptions. The further an organization moves along that spectrum, the less its exposure looks like software risk and the more it looks like delegated authority risk, the kind normally reserved for employees, contractors, and automated financial systems. Why Existing Risk Registers Miss Agentic AI Risks Most enterprise risk registers were built on a reasonably safe assumption: a human initiates consequential actions, and the technology around that human behaves deterministically. Agentic AI breaks both halves of that assumption at once. A few specific gaps show up quickly when organizations try to map agentic deployments onto existing categories. Operational risk registers assume process failures come from human error or system outages, not from a system independently choosing an unanticipated path to a stated goal. Cybersecurity risk registers are built around unauthorized external access, while an agent problem usually involves an authorized system taking unauthorized internal actions with its own legitimate credentials. Model risk frameworks, borrowed largely from financial services, evaluate output accuracy rather than action consequences, which matters most when those actions can’t be reversed. And third-party risk assessments treat vendors as static entities, not as autonomous agents that might invoke other vendors’ agents on your behalf. See our guide to the NIST AI Risk Management Framework for how output-focused frameworks are structured. The result is a governance blind spot. An organization can be compliant against its AI policy, its cybersecurity policy, and its vendor risk policy, and still have nobody accountable for the specific risk of a system initiating a harmful sequence of actions before anyone notices. Defining Agentic Autonomy Risk Agentic Autonomy Risk is the risk that an AI system, operating with delegated decision-making and execution authority, takes actions that are harmful, non-compliant, or misaligned with organizational intent before adequate human oversight can intervene. Those actions might happen independently or in coordination with other agents. It deserves standing as a named category alongside cybersecurity, operational, legal, financial, and third-party risk because the loss event itself is different. The harm is a completed action in a live system, and it may be difficult or impossible to reverse. The accountability structure is different too: when an orchestrating agent delegates to sub-agents, responsibility for the outcome gets distributed in ways existing ownership models don’t cleanly capture. So is the detection window. Traditional controls assume a human is positioned to catch an error before it compounds, but an agent can execute dozens of dependent actions faster than any human review cycle. 7 Agentic AI Risk Scenarios to Put on Your Register 1. Unauthorized autonomous decision-making. An agent takes an action within its technical permissions but outside its intended business mandate. It adjusts pricing, approves a refund, or modifies a customer record, and no policy ever explicitly authorized that scenario. 2. Goal misalignment. The agent optimizes for a literal interpretation of its objective in a way that diverges from actual business intent, particularly under ambiguous or adversarial inputs. 3. Multi-agent interactions and cascading failures. One agent’s flawed output becomes another agent’s trusted input. A single error can propagate across a chain of agents faster than anyone can detect it, amplifying the original mistake instead of containing it. 4. Excessive tool or system permissions. Agents get provisioned with broad, standing access “to be safe” rather than scoped, least-privilege access tied to specific tasks. A productivity tool quietly becomes a privilege-escalation path. 5. Regulatory non-compliance. Autonomous actions trigger obligations under data protection, financial services, employment, or sector-specific regulation, and they execute without the compliance review a human-initiated process would normally receive. 6. Explainability and accountability gaps. An autonomous action causes harm and the organization can’t clearly reconstruct why the agent chose that path, or establish whether the business owner, the AI governance function, or the vendor is accountable for the outcome. 7. Autonomous third-party actions. A vendor’s agent, integrated into your environment, takes action on your behalf, or your agent acts against a

A SOC 2 penetration test costs between $1,000 and $30,000 for most companies. A typical SaaS scope, meaning one web application, its API layer, and the cloud infrastructure behind it, usually lands between $2,000 and $20,000. Early-stage startups with a narrow scope can get an auditor-accepted test for $1,000 to $8,000, while enterprises with multiple products and hybrid infrastructure regularly spend $20,000 to $50,000 or more. The spread is wide because “penetration test” covers everything from an automated scan with a cover page to weeks of manual testing by senior engineers. Auditors know the difference, and so do the enterprise customers who asked for your SOC 2 report in the first place. This guide breaks down what drives the price, where the hidden costs sit, and how to buy a test that holds up in fieldwork without overpaying for it. What Is SOC 2 Penetration Testing?​ A SOC 2 penetration test is a simulated attack on your systems, performed by a qualified security professional, scoped to the environment covered by your SOC 2 report. The tester tries to exploit real weaknesses the way an attacker would: broken access controls, injection flaws, misconfigured cloud services, exposed credentials. The output is a report your auditor reads as evidence that your security controls work in practice, not only on paper. That last part matters. A pentest bought for SOC 2 has a second audience beyond your security team. If the report doesn’t map findings to your audit scope, document its methodology, and show remediation, it fails the job you bought it for. We cover the full deliverable in our guide to what a SOC 2-ready VAPT report includes. How Penetration Testing Fits Into SOC 2 Compliance​ SOC 2 is built on the AICPA’s Trust Services Criteria, and the Security category (the Common Criteria) applies to every report. Penetration testing is the standard way to satisfy CC7.1, which expects you to detect and monitor for new vulnerabilities, and it supports CC4.1, which covers ongoing evaluations of whether controls actually function. The AICPA’s points of focus explicitly mention vulnerability scanning and penetration testing as examples of how companies meet these criteria. In practice, the test slots into your audit timeline as an evidence item. Your auditor will ask for the report, check the test date against the audit period, and review how you handled the findings. Remediation is often scrutinized harder than the test itself, because it shows whether your vulnerability management process runs or merely exists. Is Penetration Testing Required for SOC 2?​ Strictly speaking, no. The Trust Services Criteria never use the word “mandatory” about penetration testing. You could theoretically satisfy CC7.1 with vulnerability scanning and strong monitoring alone. In reality, almost every auditor expects one, and skipping it invites two problems. First, your auditor may push back during fieldwork or add exceptions to the report. Second, the enterprise buyers reviewing your SOC 2 report increasingly look for pentest evidence specifically, and a report without it raises questions during procurement. Treat the test as effectively required and budget for it from the start of your SOC 2 compliance checklist. How Much Does SOC 2 Penetration Testing Cost? Typical Price Range for SOC 2 Pen Testing Most companies pay $1,000 to $30,000, with the median engagement for a SaaS business sitting around $12,000 to $15,000. Compliance-focused tests at the lower end of the market start around $1,000 to $5,000. Deep manual testing from established firms runs $10,000 to $30,000. Anything quoted below roughly $3,000 is almost certainly automated scanning packaged as a pentest, which auditors are getting better at spotting. Cost by Company Size (Startup, SMB, Enterprise) Company size is a proxy, not the driver. A 15-person company with three products and a legacy on-prem component will pay more than a 200-person company with one tightly scoped SaaS platform. Testers price effort, and effort follows scope. Cost by Test Type (Network, Web App, API, Cloud, Internal/External) Most SOC 2 engagements bundle two or three of these. The common package for a cloud-native SaaS company is web app plus API plus cloud configuration, which is why the $1,000 to $20,000 band comes up so often. Companies with office networks and internal systems in their audit scope add internal network testing, and the price climbs accordingly. Factors That Influence SOC 2 Penetration Testing Cost Scope and Number of Assets Tested Scope is the single biggest cost driver. Every additional application, API endpoint group, cloud account, or network segment adds testing hours. A pentest priced without a scoping call is a pentest priced on guesswork, and the guess usually favors the vendor. Complexity of Application or Infrastructure​ A simple CRUD app with two user roles tests quickly. A multi-tenant platform with role hierarchies, workflow engines, file processing, and third-party integrations takes far longer, because each of those features creates attack surface a tester has to work through manually. Authentication tiers matter especially: every distinct role needs testing for privilege escalation and cross-tenant data access. Testing Methodology (Black Box, Grey Box, White Box) Black box testing gives the tester nothing but a URL, grey box adds credentials and documentation, and white box adds source code and architecture diagrams. Grey box is the default for SOC 2 and usually the best value, since the tester spends time exploiting rather than discovering. White box costs more upfront but finds deeper issues. Black box sounds rigorous but often wastes paid hours on reconnaissance an attacker would run for free. Depth of Testing and Manual vs. Automated Approaches Automated scanning finds known vulnerability patterns. Manual testing finds business logic flaws, chained exploits, and authorization gaps that no scanner catches, and it’s the part auditors and security-literate customers actually value. The ratio of manual work to automation is the honest explanation for most price differences between two quotes covering the same scope. Tester Credentials and Firm Reputation Senior testers holding OSCP, GPEN, or CREST credentials bill higher rates, and firms with recognized methodologies charge a premium for the credibility their letterhead carries

Two compromised versions of LiteLLM sat on PyPI for roughly 40 minutes on the morning of March 24, 2026. That window was enough to capture secrets from around 434,000 CI/CD pipeline runs across nearly 2,500 organizations, including AWS, Samsung, Cisco, Salesforce, Siemens, and Deloitte. In August, researchers at CloudSEK and Hudson Rock confirmed they had obtained the raw exfiltrated data: a 153GB archive containing 433,909 files of environment variables, cloud keys, Kubernetes secrets, and API tokens harvested live from running pipelines, as covered by Help Net Security’s reporting on the credential archive. If LiteLLM runs anywhere in your stack, or you touch any AI proxy infrastructure at all, you need answers to three things: whether you were exposed, what to rotate first, and whether the rotation you did back in March actually held. That last one matters more than it sounds, because “we rotated everything” has already burned at least one very large company. How the Breach Happened The attack didn’t start with LiteLLM. On March 19, 2026, a threat group called TeamPCP compromised the build pipeline of Trivy, a vulnerability scanner half the industry runs, and pushed a poisoned release. LiteLLM’s own CI pipeline ran Trivy, so the poisoned scanner had legitimate read access to the project’s runner environment. The attackers used that to steal LiteLLM’s PyPI publishing tokens and ship two malicious releases of their own: versions 1.82.7 and 1.82.8. KICS and the Telnyx Python SDK got hit in the same campaign. The payload design is the part worth studying. The malicious package dropped a .pth startup hook into site-packages, so the code ran the moment any Python interpreter started on the machine, whether or not anything imported LiteLLM. From there it harvested environment variables, read local credential files like .aws/credentials and .kube/config, tried to move laterally across Kubernetes clusters, and installed a systemd backdoor dressed up as a generic telemetry service. InfoQ’s coverage of the PyPI compromise put downloads of the compromised release above 40,000. For scale, LiteLLM normally gets downloaded around 3 million times a day. The exfiltration had a nasty fallback, too. According to CloudSEK, stolen data was encrypted and sent to a typosquatted domain, and when that failed, the malware created a public repository inside the victim’s own GitHub account and uploaded the loot as a release asset. Some companies were publishing their own secrets to the open internet and had no idea. Worth Knowing: The malicious code only existed in the PyPI artifacts. The GitHub source repository stayed clean the whole time, so a developer reviewing the code on GitHub saw nothing wrong. Source review isn’t artifact verification. If you don’t check that what the registry serves matches the upstream source, this class of attack is invisible to you. How to Check If You Were Exposed Three checks, from quickest to most involved. 1. Confirm whether the compromised versions ever ran The malicious versions went live on PyPI at 10:39 UTC on March 24, 2026 and got quarantined about 40 minutes later. The project’s advice: treat any install from that day before 16:00 UTC as suspect. Search your lockfiles, pip caches, SBOMs, and container image histories for 1.82.7 and 1.82.8. And check your internal artifact mirrors. An Artifactory or Nexus proxy that cached the bad release in March can keep serving it internally long after PyPI pulled it. Keep the .pth mechanism in mind when you scope this. The question isn’t “which applications import LiteLLM,” it’s “which machines had the package installed at all,” because every Python process on an infected machine triggered the payload. 2. Hunt for persistence Rotation is pointless if the attacker still has a foothold. Check developer machines, CI runners, and containers for unauthorized .pth files in site-packages and for suspicious systemd units, especially anything posing as a system telemetry service. And review activity from March 24 onward, not just the 40-minute window. Persistence is there so the access outlives the infection. Pro Tip: Don’t limit the persistence hunt to live machines. Base container images rebuilt in late March may have baked the payload into every image derived from them since. Scan your image registry for the affected LiteLLM versions and for unexpected .pth files, then trace which running workloads came from flagged images. 3. Check whether your secrets are in the dump Hudson Rock has published a domain lookup tool and is running ethical disclosures for affected organizations, and CloudSEK maintains a high-confidence victim list. Use them, but know their limits. Attribution in this dataset is genuinely hard. One dump with a siriusxm.com committer email actually traced, through its self-hosted GitLab endpoints, to AdsWizz, a SiriusXM subsidiary. And a large share of the dumps are generic pipeline configurations with no identifying domain, email, or server name at all. Absence from a victim list is not evidence of absence. If your pipelines ran the compromised versions, assume exposure no matter what a lookup tool tells you. What to Rotate, in What Order The guidance from both research teams is blunt: treat every secret the LiteLLM environment could reach as compromised. That covers secrets on disk, in memory, injected into CI jobs, and anything retrievable through instance metadata services. Work down by blast radius: Priority Credential type Why it comes first 1 Cloud IAM keys (AWS, GCP, Azure) Direct control of infrastructure, data stores, and billing. This is where attackers monetize fastest. 2 GitHub and GitLab PATs, package publishing tokens These let an attacker poison your releases and turn your company into the next link in the supply chain. 3 Kubernetes service account tokens and kubeconfigs Lateral movement across clusters was built into the payload, not a theoretical risk. 4 Database passwords and third-party API keys Dumped in plain text in the archive, often with no attribution, so nobody will warn you they leaked. 5 AI provider API keys Billing abuse, quota theft, and access to whatever data flows through your LLM routing layer. One word matters more than the rest of this article: revoke, don’t just rotate. That