Table of Contents

Reach SOC 2 Compliance in 6 Weeks or Less.

  /

  / CMMC Scoring Explained: SPRS Scores & the 110-Point System

CMMC Scoring Explained: SPRS Scores & the 110-Point System

Every defense contractor that handles Controlled Unclassified Information (CUI) has a number attached to its CAGE code in a DoD database.

That number ranges from -203 to a perfect 110 and most organizations that calculate it honestly for the first time land somewhere they would rather not advertise.

This guide covers how CMMC scoring works: where the number comes from, what counts as a passing score at each CMMC level, how to calculate and submit a score in SPRS, and where Plans of Action and Milestones (POA&Ms) fit in.

CMMC Scoring Explained

What Is CMMC Scoring?

CMMC 2.0 is the Department of Defense program for verifying that companies in the Defense Industrial Base (DIB) actually protect Federal Contract Information (FCI) and CUI, rather than simply attesting that they do. The program rule, 32 CFR Part 170, took effect in December 2024, and the acquisition rule that inserts CMMC requirements into contracts via DFARS 252.204-7021 began phasing in from November 2025. Phase 2, which makes third-party certification the default for contracts involving CUI, arrives in November 2026.

CMMC scoring is the quantitative layer underneath all of this. At Level 2, the score measures implementation of the 110 security requirements of NIST SP 800-171, the standard that has applied to contractors handling CUI since DFARS 252.204-7012 made it mandatory. CMMC did not invent new controls at Level 2; it created a verification and scoring regime around controls contractors were already obligated to implement.

The score matters for three practical reasons. It determines contract eligibility, because solicitations now specify a required CMMC status and contracting officers check SPRS before award. It drives prime contractor flow-downs, since primes must verify subcontractor scores before passing CUI down the supply chain. And it creates legal exposure: a senior official affirms the score, and a knowingly inflated number is a False Claims Act problem, not a paperwork problem.

Reach SOC 2 Compliance in 6 Weeks or Less

Schedule Your Free SOC 2 Assessment Today

Understanding the SPRS Scoring System

The Supplier Performance Risk System (SPRS) is the DoD’s authoritative source for supplier risk information. For cybersecurity purposes, it stores the results of NIST SP 800-171 assessments and CMMC statuses against each contractor’s CAGE code. Contracting officers, programme offices, and DCMA personnel query it routinely; prime contractors can verify that a subcontractor has a current assessment on file.

SPRS does not perform the assessment. It is a reporting database. Self-assessment scores are entered directly by the contractor through the Procurement Integrated Enterprise Environment (PIEE). Results of third-party certification assessments are entered by the C3PAO into the CMMC instance of eMASS, which then populates SPRS automatically.

The relationship between an SPRS score and CMMC certification is straightforward: same methodology, different assessor. The self-assessment score is your own claim about your posture. A CMMC Level 2 certification is the same 110 requirements scored by a Certified Third-Party Assessment Organization (C3PAO), with the result carrying formal status under the programme rule. A contractor whose self-reported 110 collapses to 60 under C3PAO scrutiny has a credibility problem on the record.

The CMMC Scoring Methodology Explained

The methodology comes from the NIST SP 800-171 DoD Assessment Methodology, Version 1.2.1, now codified for CMMC in 32 CFR 170.24. Every organisation starts at the maximum of 110 points. For every requirement scored NOT MET, a weighted value of 1, 3, or 5 points is subtracted.

The weighting reflects security impact. Five-point requirements are those whose absence exposes the network or CUI directly. Three-point requirements have a specific, meaningful effect on security. One-point requirements have a limited or indirect effect. Because total possible deductions add up to 313, the floor is -203. Negative scores are common on a first honest assessment, and they are not a clerical curiosity: a deeply negative number visible to a contracting officer signals an organisation years away from certification.

There is no partial credit. A requirement that is 90 percent implemented deducts its full point value, exactly like one that was never started. The only two exceptions are multi-factor authentication (3.5.3), which deducts 3 points instead of 5 if MFA covers remote and privileged users but not all users, and FIPS-validated encryption (3.13.11), which deducts 3 points instead of 5 if encryption is in place but not FIPS-validated. Everything else is binary.

One further prerequisite catches people out: a System Security Plan (3.12.4) must exist at the time of assessment. Without an SSP describing how each requirement is met, the assessment cannot be completed at all, and the absence is treated as non-compliance with DFARS 252.204-7012 rather than as a scoring deduction.

CMMC Levels

CMMC Score Requirements by Level

Scoring works differently at each of the three CMMC levels, and the term passing score means something different at each. 

Level 1

Level 1 sits apart from both Level 2 and Level 3: it requires an annual self-assessment of just 15 basic safeguarding requirements, carries no numeric score, permits no POA&Ms, and requires only an annual affirmation. There is no minimum number to hit because the assessment is pass/fail on each individual requirement.

Level 2

At Level 2, the 110-point methodology applies in full. A score of 110 earns Final Level 2 status. A score of at least 88, where every unmet requirement is POA&M-eligible under 32 CFR 170.21, earns Conditional Level 2 status — but only as a temporary bridge to the full 110. At 

Level 3

Level 3, the bar rises further: organizations must first hold Final Level 2 status from a C3PAO assessment, then undergo a DIBCAC-led assessment against the 24 enhanced requirements drawn from NIST SP 800-172 requirements, each worth a single point.

The Level 2 thresholds deserve emphasis because they are widely misread. A score of 88 does not mean you passed. It means you are eligible for Conditional Level 2 status, and only if every unmet requirement is one the rule allows on a POA&M.

Conditional status starts a 180-day clock. Final Level 2 status requires the full 110, achieved either at the initial assessment or at the POA&M closeout assessment.

How to Calculate Your CMMC Score

The most reliable way to calculate your score is to work through all 110 requirements at the assessment-objective level using NIST SP 800-171A, the companion publication assessors use. A single requirement can contain a dozen objectives, and one failed objective fails the requirement. Conducting a NIST SP 800-171 gap assessment at this level of granularity is what separates a defensible score from an optimistic one.

Once gaps are identified, map each NOT MET requirement to its point value using Annex A of the DoD Assessment Methodology. Tag every gap with its 1, 3, or 5 point weight, applying the two partial-scoring rules for 3.5.3 and 3.13.11 where relevant. Then document implementation status in the SSP: for each requirement, the SSP should state how it is implemented, where, and by what mechanism. Requirements that are technically enforced but undocumented are routinely scored NOT MET in third-party assessments, because the assessor scores what can be evidenced, not what exists in theory.

Finally, sum the deductions and subtract from 110. An organization missing two 5-point, three 3-point, and four 1-point requirements scores 110 minus 23, or 87, one point short of Conditional eligibility, which is exactly the kind of margin that makes the weighting worth understanding before an assessment rather than after.

Pro Tip: Score Yourself

Score yourself against the assessment objectives in NIST SP 800-171A, not the requirement text in 800-171. Self-assessments done at the requirement level run optimistic because a control that is "mostly there" feels met. Assessors work objective by objective, and the CMMC Level 2 Assessment Guide shows precisely how they will score you. Using the same lens removes the gap between your number and theirs.

How to Submit Your CMMC Score in SPRS

Submission runs through PIEE. Register an account, request the SPRS Cyber Vendor role for your CAGE code, and wait for your Contractor Administrator to approve it. Once inside SPRS, the NIST SP 800-171 assessment entry asks for a defined set of fields: the assessment date, the summary score, the scope (enterprise, enclave, or contract-specific), the CAGE codes covered, the SSP name and version the assessment was performed against, and a planned date by which a score below 110 will reach 110.

The supporting documentation does not get uploaded. The SSP, the assessment workpapers, and the POA&M stay with you, but they must exist and they must reconcile with the submitted number — because they are exactly what DIBCAC reviews if your score is ever checked under a medium or high assessment.

Keep the record current. A score must be refreshed at least every three years to remain valid, the senior official’s affirmation of continuing compliance recurs annually under 32 CFR 170.22, and a material change to the environment — such as a cloud migration or a new system handling CUI — means recalculating and resubmitting rather than waiting for the cycle.

Reach SOC 2 Compliance in 6 Weeks or Less

Schedule Your Free SOC 2 Assessment Today

The Role of POA&Ms in CMMC Scoring

A Plan of Action and Milestones is a time-bound remediation plan for requirements scored NOT MET: what the deficiency is, what will fix it, who owns it, and by when. Under CMMC, its use is far narrower than most contractors assume. 32 CFR 170.21 permits POA&Ms only to reach Conditional status, and only when all of the following hold: the assessment score is at least 88, every POA&M item is a 1-point requirement, and none of the specifically prohibited requirements appear on the plan.

The 180-day window is rigid. The clock starts on the Conditional status date, and every POA&M item must be closed and verified through a POA&M closeout assessment within it. For a certification assessment, the closeout is performed by a C3PAO. If items remain open at day 180, the Conditional status expires, the organisation becomes ineligible for awards requiring that status, and the entire assessment must be repeated.

Insider note: There is exactly one situation in which something heavier than a 1-point item can sit on a POA&M — FIPS-validated encryption (3.13.11), and only in its partial state. If encryption is deployed but not FIPS-validated, the 3-point deduction is POA&M-eligible. If no encryption is in place at all, the full 5-point deduction applies, and the requirement cannot be deferred. In practice, this makes cryptographic module validation the single most consequential procurement question in a Level 2 readiness project.

Common CMMC Scoring Mistakes to Avoid

Overstating the self-assessment. When DIBCAC has reviewed self-reported scores under medium and high assessments, independently verified numbers have frequently come in dramatically below what contractors submitted. The gap is rarely fraud; it is requirement-level scoring, charitable interpretation, and confusing planned with implemented. The POA&M trap sits here too: a requirement on a plan of action is still NOT MET and still deducts its full point value.

An SSP that does not match reality. Assessors triangulate the submitted score, the SSP, and observable evidence. When the SSP describes controls that interviews and technical testing cannot confirm, the requirement fails — and the credibility of every other claim in the document drops with it.

Arithmetic and weighting errors. Treating partially implemented requirements as met, missing the two partial-scoring exceptions, or working from an outdated copy of the methodology all produce a number that will not survive verification.

Letting the POA&M go stale. Milestones without owners and dates, or items that quietly slip past their completion dates, convert a Conditional status into an expired one. The 180-day closeout is not extendable.

Pro Tip: The annual affirmation under 32 CFR 170.22

The annual affirmation under 32 CFR 170.22 is signed by a named senior official, and the Department of Justice has pursued False Claims Act cases against contractors over misrepresented cybersecurity compliance through its Civil Cyber-Fraud Initiative. An optimistic score is no longer a private estimate; it is a federal representation with a signature on it.

How to Improve Your CMMC Score

Sequence remediation by weight, not by ease. Closing five 1-point documentation gaps moves the score five points; closing one 5-point technical control moves it the same distance and removes a requirement that can never sit on a POA&M. The 5-point population — which includes boundary protection, access control enforcement, flaw remediation, and audit logging — is where both the score and the actual security risk concentrate.

Next, close the documentation layer. A meaningful share of NOT MET findings in otherwise capable organisations are evidence failures: the control runs, but no policy requires it, no procedure describes it, and no artefact proves it. Fixing this is cheap relative to its scoring impact.

Finally, treat the score as a maintained asset. Configuration drift, expired FIPS certificates after patching, and new systems entering scope all erode a score between assessments. Continuous monitoring tied to the SSP, with the score recalculated whenever the environment changes materially, keeps the SPRS record defensible and avoids a scramble before the next triennial cycle.

 

Self-Assessment vs. Third-Party Assessment Scoring

The methodology is identical in both paths; what changes is who applies it, how evidence is tested, and what the result is worth contractually.

In a self-assessment, the contractor scores its own environment, enters the result in SPRS, and a senior official affirms it. This satisfies the requirement for contracts that specify only the basic DFARS 252.204-7012 reporting obligation or a CMMC Level 1 or Level 2 self-assessment status. The score is the contractor’s own claim, and it carries the legal weight of that affirmation.

In a third-party assessment, a Certified Third-Party Assessment Organization (C3PAO) applies the same 110-requirement methodology with independent evidence testing — document review, interviews, and technical observation — across the defined assessment scope. The result is entered into eMASS and flows to SPRS as a formal certification status rather than a self-reported number. For Level 3, the assessor is DIBCAC rather than a C3PAO, and the additional 24 requirements from NIST SP 800-172 are scored on top of a confirmed Final Level 2 foundation.

Level 1 sits apart from both: a simple annual self-assessment of 15 requirements with no numeric score, no POA&Ms, and an annual affirmation. Organisations that completed a Joint Surveillance Voluntary Assessment (JSVA) with DIBCAC before the rule took effect were able to convert a perfect result into early certification standing, which is why JSVA alumni dominate the first wave of certified companies.

CMMC scoring rewards organizations that measure themselves the way an assessor would: objective by objective, evidence first, weighted gaps prioritized, and a POA&M used as a short runway rather than a parking lot.

The contractors who struggle are almost never the ones with the worst security; they are the ones whose number was built on assumptions instead of assessment.

Frequently Asked Questions

What is the minimum SPRS score for CMMC Level 2?

Final Level 2 status requires 110. A score of at least 88 can earn Conditional Level 2 status, but only if every unmet requirement is POA&M-eligible under 32 CFR 170.21 and all items close within 180 days.

A score is valid for three years, but the senior official’s affirmation of continuing compliance is annual, and any material change to the assessed environment requires recalculating and resubmitting before the cycle ends.

Nothing automatic, but it is visible to contracting officers and primes as a clear signal of non-implementation, and it places Conditional eligibility (88) far out of reach. First honest assessments commonly land negative; the response is a weighted remediation plan, not a resubmission.

Where only the self-assessment reporting requirement applies, a current score of any value can satisfy the letter of the rule. Once a solicitation specifies a CMMC status under DFARS 252.204-7021, the required status is an eligibility condition, and a low score means no award.

Three years from the assessment date, subject to the annual affirmation and to resubmission if the environment changes materially.

The SPRS score is a self-reported number under the DoD Assessment Methodology. A CMMC Level 2 certification is a formal status — Conditional or Final — resulting from an assessment by a C3PAO at Level 2 or DIBCAC at Level 3, using the same scoring rules. The first is a claim; the second is a verified credential.

SPRS is not public. DoD personnel — including contracting officers, program offices, and DCMA/DIBCAC — can view it, and prime contractors can confirm whether a subcontractor holds a current assessment when verifying flow-down compliance.

Axipro Author

Picture of Pedro Dias

Pedro Dias

Pedro has been writing online for over 10 years. With experience in all things programming, cyber security, and compliance, he is our editor-in-chief at Axipro.

Blog Highlights

Explore More Articles

For the past two years, enterprise AI risk conversations have centered on a familiar set of concerns: model bias, hallucination, data privacy, and dependency on third-party models. These are real risks, and most organizations now run some version of a governance program to manage them. But something has shifted. Organizations are no longer just deploying AI that generates content for a human to review. They’re deploying AI that acts. Agents now plan multi-step tasks, call APIs, move data between systems, execute transactions, and coordinate with other agents, often with no human checkpoint in the loop. That shift deserves more than a footnote in the existing AI risk category. It deserves its own line in the risk register: Agentic Autonomy Risk. What Is Agentic AI Risk Management? Agentic AI risk management is the practice of identifying, assessing, and controlling the risks created when AI systems take autonomous action on an organization’s behalf. Where traditional AI governance evaluates outputs (accuracy, bias, privacy), agentic AI risk management governs what agents actually do: the tools they call, the permissions they inherit, and the downstream consequences of their actions. That distinction is the reason existing risk registers struggle with agents, and it’s worth unpacking properly. What Agentic AI Actually Changes Traditional AI systems, even generative ones, are advisory. They produce an output such as a summary, a prediction, a draft email, or a classification, and a human remains the last checkpoint before anything happens in the real world. Agentic AI removes that checkpoint. An agentic system doesn’t just produce an answer. It pursues a goal. It decides which tools to call and in what order, then executes those actions directly against live systems: submitting a purchase order, modifying a database record, sending an external communication, or orchestrating a set of sub-agents to complete a broader workflow. Agentic autonomy is the degree to which a system can plan and execute actions without a human explicitly authorizing each step. It’s a spectrum rather than a binary. At one end, the AI drafts and a human approves every action. At the other, the AI operates within broad guardrails and only escalates exceptions. The further an organization moves along that spectrum, the less its exposure looks like software risk and the more it looks like delegated authority risk, the kind normally reserved for employees, contractors, and automated financial systems. Why Existing Risk Registers Miss Agentic AI Risks Most enterprise risk registers were built on a reasonably safe assumption: a human initiates consequential actions, and the technology around that human behaves deterministically. Agentic AI breaks both halves of that assumption at once. A few specific gaps show up quickly when organizations try to map agentic deployments onto existing categories. Operational risk registers assume process failures come from human error or system outages, not from a system independently choosing an unanticipated path to a stated goal. Cybersecurity risk registers are built around unauthorized external access, while an agent problem usually involves an authorized system taking unauthorized internal actions with its own legitimate credentials. Model risk frameworks, borrowed largely from financial services, evaluate output accuracy rather than action consequences, which matters most when those actions can’t be reversed. And third-party risk assessments treat vendors as static entities, not as autonomous agents that might invoke other vendors’ agents on your behalf. See our guide to the NIST AI Risk Management Framework for how output-focused frameworks are structured. The result is a governance blind spot. An organization can be compliant against its AI policy, its cybersecurity policy, and its vendor risk policy, and still have nobody accountable for the specific risk of a system initiating a harmful sequence of actions before anyone notices. Defining Agentic Autonomy Risk Agentic Autonomy Risk is the risk that an AI system, operating with delegated decision-making and execution authority, takes actions that are harmful, non-compliant, or misaligned with organizational intent before adequate human oversight can intervene. Those actions might happen independently or in coordination with other agents. It deserves standing as a named category alongside cybersecurity, operational, legal, financial, and third-party risk because the loss event itself is different. The harm is a completed action in a live system, and it may be difficult or impossible to reverse. The accountability structure is different too: when an orchestrating agent delegates to sub-agents, responsibility for the outcome gets distributed in ways existing ownership models don’t cleanly capture. So is the detection window. Traditional controls assume a human is positioned to catch an error before it compounds, but an agent can execute dozens of dependent actions faster than any human review cycle. 7 Agentic AI Risk Scenarios to Put on Your Register 1. Unauthorized autonomous decision-making. An agent takes an action within its technical permissions but outside its intended business mandate. It adjusts pricing, approves a refund, or modifies a customer record, and no policy ever explicitly authorized that scenario. 2. Goal misalignment. The agent optimizes for a literal interpretation of its objective in a way that diverges from actual business intent, particularly under ambiguous or adversarial inputs. 3. Multi-agent interactions and cascading failures. One agent’s flawed output becomes another agent’s trusted input. A single error can propagate across a chain of agents faster than anyone can detect it, amplifying the original mistake instead of containing it. 4. Excessive tool or system permissions. Agents get provisioned with broad, standing access “to be safe” rather than scoped, least-privilege access tied to specific tasks. A productivity tool quietly becomes a privilege-escalation path. 5. Regulatory non-compliance. Autonomous actions trigger obligations under data protection, financial services, employment, or sector-specific regulation, and they execute without the compliance review a human-initiated process would normally receive. 6. Explainability and accountability gaps. An autonomous action causes harm and the organization can’t clearly reconstruct why the agent chose that path, or establish whether the business owner, the AI governance function, or the vendor is accountable for the outcome. 7. Autonomous third-party actions. A vendor’s agent, integrated into your environment, takes action on your behalf, or your agent acts against a

A SOC 2 penetration test costs between $1,000 and $30,000 for most companies. A typical SaaS scope, meaning one web application, its API layer, and the cloud infrastructure behind it, usually lands between $2,000 and $20,000. Early-stage startups with a narrow scope can get an auditor-accepted test for $1,000 to $8,000, while enterprises with multiple products and hybrid infrastructure regularly spend $20,000 to $50,000 or more. The spread is wide because “penetration test” covers everything from an automated scan with a cover page to weeks of manual testing by senior engineers. Auditors know the difference, and so do the enterprise customers who asked for your SOC 2 report in the first place. This guide breaks down what drives the price, where the hidden costs sit, and how to buy a test that holds up in fieldwork without overpaying for it. What Is SOC 2 Penetration Testing?​ A SOC 2 penetration test is a simulated attack on your systems, performed by a qualified security professional, scoped to the environment covered by your SOC 2 report. The tester tries to exploit real weaknesses the way an attacker would: broken access controls, injection flaws, misconfigured cloud services, exposed credentials. The output is a report your auditor reads as evidence that your security controls work in practice, not only on paper. That last part matters. A pentest bought for SOC 2 has a second audience beyond your security team. If the report doesn’t map findings to your audit scope, document its methodology, and show remediation, it fails the job you bought it for. We cover the full deliverable in our guide to what a SOC 2-ready VAPT report includes. How Penetration Testing Fits Into SOC 2 Compliance​ SOC 2 is built on the AICPA’s Trust Services Criteria, and the Security category (the Common Criteria) applies to every report. Penetration testing is the standard way to satisfy CC7.1, which expects you to detect and monitor for new vulnerabilities, and it supports CC4.1, which covers ongoing evaluations of whether controls actually function. The AICPA’s points of focus explicitly mention vulnerability scanning and penetration testing as examples of how companies meet these criteria. In practice, the test slots into your audit timeline as an evidence item. Your auditor will ask for the report, check the test date against the audit period, and review how you handled the findings. Remediation is often scrutinized harder than the test itself, because it shows whether your vulnerability management process runs or merely exists. Is Penetration Testing Required for SOC 2?​ Strictly speaking, no. The Trust Services Criteria never use the word “mandatory” about penetration testing. You could theoretically satisfy CC7.1 with vulnerability scanning and strong monitoring alone. In reality, almost every auditor expects one, and skipping it invites two problems. First, your auditor may push back during fieldwork or add exceptions to the report. Second, the enterprise buyers reviewing your SOC 2 report increasingly look for pentest evidence specifically, and a report without it raises questions during procurement. Treat the test as effectively required and budget for it from the start of your SOC 2 compliance checklist. How Much Does SOC 2 Penetration Testing Cost? Typical Price Range for SOC 2 Pen Testing Most companies pay $1,000 to $30,000, with the median engagement for a SaaS business sitting around $12,000 to $15,000. Compliance-focused tests at the lower end of the market start around $1,000 to $5,000. Deep manual testing from established firms runs $10,000 to $30,000. Anything quoted below roughly $3,000 is almost certainly automated scanning packaged as a pentest, which auditors are getting better at spotting. Cost by Company Size (Startup, SMB, Enterprise) Company size is a proxy, not the driver. A 15-person company with three products and a legacy on-prem component will pay more than a 200-person company with one tightly scoped SaaS platform. Testers price effort, and effort follows scope. Cost by Test Type (Network, Web App, API, Cloud, Internal/External) Most SOC 2 engagements bundle two or three of these. The common package for a cloud-native SaaS company is web app plus API plus cloud configuration, which is why the $1,000 to $20,000 band comes up so often. Companies with office networks and internal systems in their audit scope add internal network testing, and the price climbs accordingly. Factors That Influence SOC 2 Penetration Testing Cost Scope and Number of Assets Tested Scope is the single biggest cost driver. Every additional application, API endpoint group, cloud account, or network segment adds testing hours. A pentest priced without a scoping call is a pentest priced on guesswork, and the guess usually favors the vendor. Complexity of Application or Infrastructure​ A simple CRUD app with two user roles tests quickly. A multi-tenant platform with role hierarchies, workflow engines, file processing, and third-party integrations takes far longer, because each of those features creates attack surface a tester has to work through manually. Authentication tiers matter especially: every distinct role needs testing for privilege escalation and cross-tenant data access. Testing Methodology (Black Box, Grey Box, White Box) Black box testing gives the tester nothing but a URL, grey box adds credentials and documentation, and white box adds source code and architecture diagrams. Grey box is the default for SOC 2 and usually the best value, since the tester spends time exploiting rather than discovering. White box costs more upfront but finds deeper issues. Black box sounds rigorous but often wastes paid hours on reconnaissance an attacker would run for free. Depth of Testing and Manual vs. Automated Approaches Automated scanning finds known vulnerability patterns. Manual testing finds business logic flaws, chained exploits, and authorization gaps that no scanner catches, and it’s the part auditors and security-literate customers actually value. The ratio of manual work to automation is the honest explanation for most price differences between two quotes covering the same scope. Tester Credentials and Firm Reputation Senior testers holding OSCP, GPEN, or CREST credentials bill higher rates, and firms with recognized methodologies charge a premium for the credibility their letterhead carries

Two compromised versions of LiteLLM sat on PyPI for roughly 40 minutes on the morning of March 24, 2026. That window was enough to capture secrets from around 434,000 CI/CD pipeline runs across nearly 2,500 organizations, including AWS, Samsung, Cisco, Salesforce, Siemens, and Deloitte. In August, researchers at CloudSEK and Hudson Rock confirmed they had obtained the raw exfiltrated data: a 153GB archive containing 433,909 files of environment variables, cloud keys, Kubernetes secrets, and API tokens harvested live from running pipelines, as covered by Help Net Security’s reporting on the credential archive. If LiteLLM runs anywhere in your stack, or you touch any AI proxy infrastructure at all, you need answers to three things: whether you were exposed, what to rotate first, and whether the rotation you did back in March actually held. That last one matters more than it sounds, because “we rotated everything” has already burned at least one very large company. How the Breach Happened The attack didn’t start with LiteLLM. On March 19, 2026, a threat group called TeamPCP compromised the build pipeline of Trivy, a vulnerability scanner half the industry runs, and pushed a poisoned release. LiteLLM’s own CI pipeline ran Trivy, so the poisoned scanner had legitimate read access to the project’s runner environment. The attackers used that to steal LiteLLM’s PyPI publishing tokens and ship two malicious releases of their own: versions 1.82.7 and 1.82.8. KICS and the Telnyx Python SDK got hit in the same campaign. The payload design is the part worth studying. The malicious package dropped a .pth startup hook into site-packages, so the code ran the moment any Python interpreter started on the machine, whether or not anything imported LiteLLM. From there it harvested environment variables, read local credential files like .aws/credentials and .kube/config, tried to move laterally across Kubernetes clusters, and installed a systemd backdoor dressed up as a generic telemetry service. InfoQ’s coverage of the PyPI compromise put downloads of the compromised release above 40,000. For scale, LiteLLM normally gets downloaded around 3 million times a day. The exfiltration had a nasty fallback, too. According to CloudSEK, stolen data was encrypted and sent to a typosquatted domain, and when that failed, the malware created a public repository inside the victim’s own GitHub account and uploaded the loot as a release asset. Some companies were publishing their own secrets to the open internet and had no idea. Worth Knowing: The malicious code only existed in the PyPI artifacts. The GitHub source repository stayed clean the whole time, so a developer reviewing the code on GitHub saw nothing wrong. Source review isn’t artifact verification. If you don’t check that what the registry serves matches the upstream source, this class of attack is invisible to you. How to Check If You Were Exposed Three checks, from quickest to most involved. 1. Confirm whether the compromised versions ever ran The malicious versions went live on PyPI at 10:39 UTC on March 24, 2026 and got quarantined about 40 minutes later. The project’s advice: treat any install from that day before 16:00 UTC as suspect. Search your lockfiles, pip caches, SBOMs, and container image histories for 1.82.7 and 1.82.8. And check your internal artifact mirrors. An Artifactory or Nexus proxy that cached the bad release in March can keep serving it internally long after PyPI pulled it. Keep the .pth mechanism in mind when you scope this. The question isn’t “which applications import LiteLLM,” it’s “which machines had the package installed at all,” because every Python process on an infected machine triggered the payload. 2. Hunt for persistence Rotation is pointless if the attacker still has a foothold. Check developer machines, CI runners, and containers for unauthorized .pth files in site-packages and for suspicious systemd units, especially anything posing as a system telemetry service. And review activity from March 24 onward, not just the 40-minute window. Persistence is there so the access outlives the infection. Pro Tip: Don’t limit the persistence hunt to live machines. Base container images rebuilt in late March may have baked the payload into every image derived from them since. Scan your image registry for the affected LiteLLM versions and for unexpected .pth files, then trace which running workloads came from flagged images. 3. Check whether your secrets are in the dump Hudson Rock has published a domain lookup tool and is running ethical disclosures for affected organizations, and CloudSEK maintains a high-confidence victim list. Use them, but know their limits. Attribution in this dataset is genuinely hard. One dump with a siriusxm.com committer email actually traced, through its self-hosted GitLab endpoints, to AdsWizz, a SiriusXM subsidiary. And a large share of the dumps are generic pipeline configurations with no identifying domain, email, or server name at all. Absence from a victim list is not evidence of absence. If your pipelines ran the compromised versions, assume exposure no matter what a lookup tool tells you. What to Rotate, in What Order The guidance from both research teams is blunt: treat every secret the LiteLLM environment could reach as compromised. That covers secrets on disk, in memory, injected into CI jobs, and anything retrievable through instance metadata services. Work down by blast radius: Priority Credential type Why it comes first 1 Cloud IAM keys (AWS, GCP, Azure) Direct control of infrastructure, data stores, and billing. This is where attackers monetize fastest. 2 GitHub and GitLab PATs, package publishing tokens These let an attacker poison your releases and turn your company into the next link in the supply chain. 3 Kubernetes service account tokens and kubeconfigs Lateral movement across clusters was built into the payload, not a theoretical risk. 4 Database passwords and third-party API keys Dumped in plain text in the archive, often with no attribution, so nobody will warn you they leaked. 5 AI provider API keys Billing abuse, quota theft, and access to whatever data flows through your LLM routing layer. One word matters more than the rest of this article: revoke, don’t just rotate. That