Table of Contents

Reach SOC 2 Compliance in 6 Weeks or Less.

  /

  / ISO 27001 Business Continuity Requirements Guide

ISO 27001 Business Continuity Requirements Guide

Two controls decide whether your ISO 27001 business continuity plan survives an audit: Annex A 5.29 and Annex A 5.30. One keeps your security controls working while everything else is failing. The other gets your systems back online before the damage becomes permanent.

Plenty of teams write a continuity policy that satisfies neither in the way a certification auditor expects, and they discover the gap during the Stage 2 audit, when it is expensive to fix.

This article covers what ISO 27001:2022 actually requires for business continuity, the components an auditor will ask to see, the step-by-step build, and the mistakes that turn a continuity plan into a non-conformity.

ISO 27001 Business Continuity Plan

What Is an ISO 27001 Business Continuity Plan?

An ISO 27001 business continuity plan is the documented set of procedures that keeps information security effective and critical ICT services available during a disruption. It is not a generic “keep the lights on” binder. Under ISO 27001, the plan protects the confidentiality, integrity, and availability of information when normal operations break down: a ransomware event, a cloud outage, a data center failure, or a supplier collapse.

The plan lives inside your Information Security Management System (ISMS). It draws on your risk assessment, your asset register, and your Business Impact Analysis (BIA), and it feeds your disaster recovery procedures. Scope is the part people get wrong. ISO 27001 cares about the information security aspects of continuity, not every operational hiccup a full business continuity program might cover.

 

Why You Need a Business Continuity Plan for ISO 27001 Compliance

Downtime is expensive, and the bill arrives fast. For most organizations, the question is not whether a disruption will happen, but how quickly they recover when it does.

There is also a hard compliance reason. You cannot certify to ISO 27001 while ignoring continuity. The standard requires you to maintain information security during disruption and to keep ICT able to support recovery, and an auditor will ask for the evidence. A continuity plan is where availability stops being a promise and becomes a tested capability.

Let Axipro help you build a business continuity plan that's practical, compliant, and audit-ready.

Strengthen Your Business Continuity Strategy

ISO 27001 Requirements Related to Business Continuity Planning

ISO/IEC 27001:2022 carries 93 Annex A controls across four categories: organizational, people, physical, and technological. Continuity sits in the organizational set, and two controls do the heavy lifting, supported by two more on the technical side.

Annex A 5.29 – Information Security During Disruption

A.5.29 requires you to maintain information security at an appropriate level when a disruption hits. The point is that security controls have a habit of degrading under pressure. People disable multi-factor authentication to “speed things up,” logging stops on a failover system, or access controls loosen while everyone scrambles. A.5.29 says the confidentiality and integrity of your information must be maintained even while availability is under threat. It is classed as both a preventive and a corrective control, meaning it should reduce the chance of an incident and also help resolve one already underway.

Annex A 5.30 – ICT Readiness for Business Continuity

A.5.30 is the technical engine. It requires that your ICT readiness is planned, implemented, maintained, and tested against business continuity objectives and ICT continuity requirements. In plain terms, your servers, networks, applications, and cloud services need a defined recovery path, each with a Recovery Time Objective (RTO) and Recovery Point Objective (RPO), and you need to prove the path works. This control is entirely new in the 2022 revision. It has no precedent in ISO 27001:2013, which is exactly why teams migrating from the older version so often have a gap here.

Important: A.5.30 did not exist in ISO 27001:2013. If your continuity documentation was written against the old Annex A 17 cluster and never updated, you are missing a control the auditor will specifically test. Treat ICT readiness as a fresh requirement, not a relabel.

Two technological controls back these up.

  • Annex A 8.13 (Information Backup) requires backups to be taken and tested in line with an agreed policy, and
  • Annex A 8.14 (Redundancy of Information Processing Facilities) covers the failover and redundancy that let critical systems keep running when a component dies.

Relationship Between ISO 27001 and ISO 22301

This is where confusion is common. ISO 27001 requires the information security aspects of continuity. ISO 22301 is the dedicated standard for a full Business Continuity Management System (BCMS), covering people, facilities, supply chain, and operations far beyond information security. An ISO 27001 certificate does not certify your wider continuity program.

The good news: both standards share the Annex SL high-level structure, so risk assessment, internal audit, management review, and document control carry across. Teams that already run ISO 27001 can layer ISO 22301 on top with far less effort than starting from scratch.

ISO 27001 Business Continuity Plan

Key Components of an ISO 27001 Business Continuity Plan

Business Impact Analysis (BIA)

The BIA is the foundation. It identifies your critical business processes, the ICT systems they depend on, and the cost of losing each one over time. It is where your recovery objectives come from, not from a vendor datasheet. A BIA also sets the Maximum Tolerable Period of Disruption (MTPD): the point beyond which an activity’s failure causes unacceptable damage.

Risk and Disruption Scenario Assessment

Your risk assessment identifies what could cause a disruption and how likely it is, feeding the Risk Treatment Plan and the Statement of Applicability (SoA) that records which controls apply. Continuity planning then runs concrete scenarios: ransomware, a regional outage, a key supplier failure, the loss of a data center.

Response and Recovery Strategies

For each critical system, you define how you will respond and recover: failover to a secondary site, restore from backup, or switch to a manual workaround. This links incident response to crisis management, the executive-level decision-making that kicks in when an incident escalates beyond a routine fix.

Roles and Responsibilities

Name real people, not departments. “IT will handle it” is the single most common reason a plan fails under pressure. Every critical system and recovery step needs a named owner and a named deputy, with current contact details and the authority to act.

Pro Tip: Replace every instance of "IT" or "the business" in your plan

Replace every instance of "IT" or "the business" in your plan with a person's name and a backup name. Auditors increasingly check whether the primary responder for a critical system can actually describe their role. A swim-lane diagram showing who does what, in order, is worth more than ten pages of policy prose.

Communication Plan

Decide in advance who tells customers, regulators, staff, and partners, through which channels, and on what timeline. Many disruptions become reputational events not because recovery was slow but because the silence was loud.

Backup and Redundancy Measures

This is Annex A 8.13 and 8.14 in practice. The widely used benchmark is the 3-2-1 rule: three copies of data, on two media types, with one off-site and ideally immutable. The U.S. Cybersecurity and Infrastructure Security Agency still recommends this baseline, with an additional offline or immutable copy as ransomware defense. Redundancy means no single point of failure can take down a critical service on its own.

Testing, Maintenance, and Continual Improvement

A plan that has never been tested is a hypothesis. ISO 27001 expects testing at planned intervals, typically at least annually and after any significant change. This is the Continual Improvement half of the PDCA cycle: test, capture what failed, fix it, retest. Our guide to Testing Maintenance and Continual Improvement covers the cadence in detail.

How to Create an ISO 27001 Business Continuity Plan: Step-by-Step

Step 1: Secure Management Support

Continuity planning needs budget and authority. Get senior management to own the objectives, because A.5.30 plans must be approved at that level and the BIA needs sign-off to carry weight in an audit.

Step 2: Conduct a Business Impact Analysis

Map every critical process to its supporting systems, suppliers, and data. Quantify the impact of losing each one at intervals such as 1 hour, 24 hours, and one week, and use that to rank what gets recovered first.

Step 3: Perform a Risk Assessment

Identify the threats that could cause those impacts, assess likelihood, and record your treatment decisions in the Risk Treatment Plan and the SoA. The risk assessment and the BIA together justify where you spend on resilience.

Step 4: Define Recovery Objectives (RTO and RPO)

Set an RTO and RPO for every critical system, derived from the BIA. RTO is the maximum time to restore a service. RPO is the maximum data loss you can absorb, measured in time. Both must fit inside the MTPD.

Insider Note: The most common RTO mistake is setting it based on what your technology can deliver, then calling that the target. Auditors and regulators expect the opposite: business tolerance sets the number, and the architecture is built to meet it. The NIST guidance in SP 800-34 is explicit that RTO must sit below the maximum tolerable downtime, with a safety margin. A stale or wishful RTO is worse than none, because it creates false confidence.

Step 5: Develop the Business Continuity Plan Document

Write the plan: scope, roles, scenarios, recovery procedures, a Disaster Recovery Plan (DRP) for ICT, communication protocols, and the testing schedule. Keep procedures specific enough to follow at 2 a.m. with half the team unreachable.

Step 6: Train Personnel and Assign Responsibilities

Brief the recovery team on their roles and give general staff awareness of the temporary procedures that apply during a disruption. Training records are audit evidence, so keep them.

Step 7: Test, Review, and Update the Plan

Run a tabletop exercise or a live failover test, document the results, log every gap in a corrective action tracker, and feed the fixes back into the plan. Then schedule the next test. A test that produces no findings usually means the test was too easy.

ISO 27001 Business Continuity Plan Template (What to Include)

A workable plan template covers, at minimum: scope and objectives; a register of critical processes and their ICT dependencies; the BIA results with MTPD, RTO, and RPO per system; named roles and deputies; disruption scenarios and recovery strategies; the DRP and backup arrangements; the communication plan; and the test schedule with a log of past exercises and their outcomes. The NIST contingency planning guide and ISO/TS 22317 both offer BIA templates worth borrowing from if you are starting cold rather than reinventing the structure.

We’ve created an editable template for your use, click the link below to view it.

How to Evaluate Information Security Continuity

Evaluation is not a one-off. ISO 27001 expects you to verify, through testing and review, that information security controls and ICT recovery still work as intended. That means checking after each test that recovery met the documented RTO and RPO, that security controls such as logging, encryption, and multi-factor authentication stayed active during failover, and that the lessons from the last exercise were actually implemented. Internal audit and management review are the formal mechanisms that close this loop and give the auditor a paper trail to follow.

 

Common Mistakes to Avoid When Building Your Plan

The recurring failures are predictable. Writing the plan in IT isolation with no buy-in from the rest of the business. Setting RTO and RPO from technology capability rather than business need. Leaving ownership as a job title instead of a person. Never testing, or testing once and filing the report. Disabling security controls during recovery to move faster, which directly breaches A.5.29. And treating A.5.30 as a relabel of the old Annex A 17, when it is a genuinely new requirement with its own evidence demands.

Worth Knowing: Don't Sacrifice Security During Recovery

Worth Knowing: Auditors treat the suppression of a security control during recovery as a non-conformity, not a pragmatic shortcut. If you turn off MFA or firewall inspection to speed a restore, that decision has to be a documented, risk-assessed waiver, not an undocumented field call. A failover that quietly drops your logging is a finding waiting to happen.

What Auditors Look for in an ISO 27001 Business Continuity Plan

A certification auditor wants living proof, not shelf-ware. Expect them to ask for a BIA reviewed and signed off by management within the last 12 months, a register of critical ICT assets with RTO and RPO attached, the DRP and its technical recovery procedures, evidence of backup integrity, and a schedule of tests with results from previous exercises. They will look for after-action reports and a tracked corrective action plan showing that failures from the last test were remediated. They will check training records for the recovery team. And they will probe whether a named responder can actually describe their role. The theme throughout: evidence that continuity is operationally real, not a policy PDF filed away after last year’s audit.

If you want help building or auditing this against ISO 27001:2022, Axipro’s ISO 27001 implementation support covers the BIA, control mapping, and evidence preparation end to end.

Let Axipro help you build a business continuity plan that's practical, compliant, and audit-ready.

Strengthen Your Business Continuity Strategy

Conclusion

An ISO 27001 business continuity plan succeeds or fails on two things: whether your information security holds during a disruption, and whether your ICT can recover within the targets the business actually needs. A.5.29 and A.5.30 make both measurable and auditable. Build the plan from a real BIA, set recovery objectives from business tolerance, name owners, test the plan, and fix what breaks. Do that, and certification becomes the byproduct of genuine resilience rather than a paperwork exercise.

ISO 27001 BCP Frequently Asked Questions

Is a business continuity plan mandatory for ISO 27001 certification?

You cannot ignore continuity. ISO 27001:2022 includes A.5.29 and A.5.30 as Annex A controls, and unless you can justify their exclusion in the Statement of Applicability, you must implement them. In practice this means documented, tested continuity arrangements for information security and ICT, even if you do not pursue a full ISO 22301 BCMS.

ISO 27001 covers the information security aspects of continuity: keeping data confidential, intact, and available during disruption. ISO 22301 is a standalone Business Continuity Management System standard covering the whole organization, including people, facilities, and supply chain. An ISO 27001 certificate does not prove you have a complete BCMS; ISO 22301 does.

At planned intervals, which in practice means at least once a year and after any significant change, such as a major system migration, an acquisition, or a real incident that exposed a gap. The BIA should be reviewed on the same cadence so the recovery objectives stay accurate.

Senior management owns the objectives and approves the plan, but day-to-day responsibility is distributed across process owners, an ICT recovery team, and named individuals for each critical system and recovery step. The plan should name people and deputies, not departments.

The business continuity plan is the umbrella, covering how the organization keeps critical functions running during disruption. The Disaster Recovery Plan is a subset focused specifically on restoring IT and technology infrastructure. Under ISO 27001, the DRP supports A.5.30, while the broader continuity plan supports A.5.29 and ties the technical recovery to business priorities.

Axipro Author

Picture of Pedro Dias

Pedro Dias

Pedro has been writing online for over 10 years. With experience in all things programming, cyber security, and compliance, he is our editor-in-chief at Axipro.

Blog Highlights

Explore More Articles

For the past two years, enterprise AI risk conversations have centered on a familiar set of concerns: model bias, hallucination, data privacy, and dependency on third-party models. These are real risks, and most organizations now run some version of a governance program to manage them. But something has shifted. Organizations are no longer just deploying AI that generates content for a human to review. They’re deploying AI that acts. Agents now plan multi-step tasks, call APIs, move data between systems, execute transactions, and coordinate with other agents, often with no human checkpoint in the loop. That shift deserves more than a footnote in the existing AI risk category. It deserves its own line in the risk register: Agentic Autonomy Risk. What Is Agentic AI Risk Management? Agentic AI risk management is the practice of identifying, assessing, and controlling the risks created when AI systems take autonomous action on an organization’s behalf. Where traditional AI governance evaluates outputs (accuracy, bias, privacy), agentic AI risk management governs what agents actually do: the tools they call, the permissions they inherit, and the downstream consequences of their actions. That distinction is the reason existing risk registers struggle with agents, and it’s worth unpacking properly. What Agentic AI Actually Changes Traditional AI systems, even generative ones, are advisory. They produce an output such as a summary, a prediction, a draft email, or a classification, and a human remains the last checkpoint before anything happens in the real world. Agentic AI removes that checkpoint. An agentic system doesn’t just produce an answer. It pursues a goal. It decides which tools to call and in what order, then executes those actions directly against live systems: submitting a purchase order, modifying a database record, sending an external communication, or orchestrating a set of sub-agents to complete a broader workflow. Agentic autonomy is the degree to which a system can plan and execute actions without a human explicitly authorizing each step. It’s a spectrum rather than a binary. At one end, the AI drafts and a human approves every action. At the other, the AI operates within broad guardrails and only escalates exceptions. The further an organization moves along that spectrum, the less its exposure looks like software risk and the more it looks like delegated authority risk, the kind normally reserved for employees, contractors, and automated financial systems. Why Existing Risk Registers Miss Agentic AI Risks Most enterprise risk registers were built on a reasonably safe assumption: a human initiates consequential actions, and the technology around that human behaves deterministically. Agentic AI breaks both halves of that assumption at once. A few specific gaps show up quickly when organizations try to map agentic deployments onto existing categories. Operational risk registers assume process failures come from human error or system outages, not from a system independently choosing an unanticipated path to a stated goal. Cybersecurity risk registers are built around unauthorized external access, while an agent problem usually involves an authorized system taking unauthorized internal actions with its own legitimate credentials. Model risk frameworks, borrowed largely from financial services, evaluate output accuracy rather than action consequences, which matters most when those actions can’t be reversed. And third-party risk assessments treat vendors as static entities, not as autonomous agents that might invoke other vendors’ agents on your behalf. See our guide to the NIST AI Risk Management Framework for how output-focused frameworks are structured. The result is a governance blind spot. An organization can be compliant against its AI policy, its cybersecurity policy, and its vendor risk policy, and still have nobody accountable for the specific risk of a system initiating a harmful sequence of actions before anyone notices. Defining Agentic Autonomy Risk Agentic Autonomy Risk is the risk that an AI system, operating with delegated decision-making and execution authority, takes actions that are harmful, non-compliant, or misaligned with organizational intent before adequate human oversight can intervene. Those actions might happen independently or in coordination with other agents. It deserves standing as a named category alongside cybersecurity, operational, legal, financial, and third-party risk because the loss event itself is different. The harm is a completed action in a live system, and it may be difficult or impossible to reverse. The accountability structure is different too: when an orchestrating agent delegates to sub-agents, responsibility for the outcome gets distributed in ways existing ownership models don’t cleanly capture. So is the detection window. Traditional controls assume a human is positioned to catch an error before it compounds, but an agent can execute dozens of dependent actions faster than any human review cycle. 7 Agentic AI Risk Scenarios to Put on Your Register 1. Unauthorized autonomous decision-making. An agent takes an action within its technical permissions but outside its intended business mandate. It adjusts pricing, approves a refund, or modifies a customer record, and no policy ever explicitly authorized that scenario. 2. Goal misalignment. The agent optimizes for a literal interpretation of its objective in a way that diverges from actual business intent, particularly under ambiguous or adversarial inputs. 3. Multi-agent interactions and cascading failures. One agent’s flawed output becomes another agent’s trusted input. A single error can propagate across a chain of agents faster than anyone can detect it, amplifying the original mistake instead of containing it. 4. Excessive tool or system permissions. Agents get provisioned with broad, standing access “to be safe” rather than scoped, least-privilege access tied to specific tasks. A productivity tool quietly becomes a privilege-escalation path. 5. Regulatory non-compliance. Autonomous actions trigger obligations under data protection, financial services, employment, or sector-specific regulation, and they execute without the compliance review a human-initiated process would normally receive. 6. Explainability and accountability gaps. An autonomous action causes harm and the organization can’t clearly reconstruct why the agent chose that path, or establish whether the business owner, the AI governance function, or the vendor is accountable for the outcome. 7. Autonomous third-party actions. A vendor’s agent, integrated into your environment, takes action on your behalf, or your agent acts against a

A SOC 2 penetration test costs between $1,000 and $30,000 for most companies. A typical SaaS scope, meaning one web application, its API layer, and the cloud infrastructure behind it, usually lands between $2,000 and $20,000. Early-stage startups with a narrow scope can get an auditor-accepted test for $1,000 to $8,000, while enterprises with multiple products and hybrid infrastructure regularly spend $20,000 to $50,000 or more. The spread is wide because “penetration test” covers everything from an automated scan with a cover page to weeks of manual testing by senior engineers. Auditors know the difference, and so do the enterprise customers who asked for your SOC 2 report in the first place. This guide breaks down what drives the price, where the hidden costs sit, and how to buy a test that holds up in fieldwork without overpaying for it. What Is SOC 2 Penetration Testing?​ A SOC 2 penetration test is a simulated attack on your systems, performed by a qualified security professional, scoped to the environment covered by your SOC 2 report. The tester tries to exploit real weaknesses the way an attacker would: broken access controls, injection flaws, misconfigured cloud services, exposed credentials. The output is a report your auditor reads as evidence that your security controls work in practice, not only on paper. That last part matters. A pentest bought for SOC 2 has a second audience beyond your security team. If the report doesn’t map findings to your audit scope, document its methodology, and show remediation, it fails the job you bought it for. We cover the full deliverable in our guide to what a SOC 2-ready VAPT report includes. How Penetration Testing Fits Into SOC 2 Compliance​ SOC 2 is built on the AICPA’s Trust Services Criteria, and the Security category (the Common Criteria) applies to every report. Penetration testing is the standard way to satisfy CC7.1, which expects you to detect and monitor for new vulnerabilities, and it supports CC4.1, which covers ongoing evaluations of whether controls actually function. The AICPA’s points of focus explicitly mention vulnerability scanning and penetration testing as examples of how companies meet these criteria. In practice, the test slots into your audit timeline as an evidence item. Your auditor will ask for the report, check the test date against the audit period, and review how you handled the findings. Remediation is often scrutinized harder than the test itself, because it shows whether your vulnerability management process runs or merely exists. Is Penetration Testing Required for SOC 2?​ Strictly speaking, no. The Trust Services Criteria never use the word “mandatory” about penetration testing. You could theoretically satisfy CC7.1 with vulnerability scanning and strong monitoring alone. In reality, almost every auditor expects one, and skipping it invites two problems. First, your auditor may push back during fieldwork or add exceptions to the report. Second, the enterprise buyers reviewing your SOC 2 report increasingly look for pentest evidence specifically, and a report without it raises questions during procurement. Treat the test as effectively required and budget for it from the start of your SOC 2 compliance checklist. How Much Does SOC 2 Penetration Testing Cost? Typical Price Range for SOC 2 Pen Testing Most companies pay $1,000 to $30,000, with the median engagement for a SaaS business sitting around $12,000 to $15,000. Compliance-focused tests at the lower end of the market start around $1,000 to $5,000. Deep manual testing from established firms runs $10,000 to $30,000. Anything quoted below roughly $3,000 is almost certainly automated scanning packaged as a pentest, which auditors are getting better at spotting. Cost by Company Size (Startup, SMB, Enterprise) Company size is a proxy, not the driver. A 15-person company with three products and a legacy on-prem component will pay more than a 200-person company with one tightly scoped SaaS platform. Testers price effort, and effort follows scope. Cost by Test Type (Network, Web App, API, Cloud, Internal/External) Most SOC 2 engagements bundle two or three of these. The common package for a cloud-native SaaS company is web app plus API plus cloud configuration, which is why the $1,000 to $20,000 band comes up so often. Companies with office networks and internal systems in their audit scope add internal network testing, and the price climbs accordingly. Factors That Influence SOC 2 Penetration Testing Cost Scope and Number of Assets Tested Scope is the single biggest cost driver. Every additional application, API endpoint group, cloud account, or network segment adds testing hours. A pentest priced without a scoping call is a pentest priced on guesswork, and the guess usually favors the vendor. Complexity of Application or Infrastructure​ A simple CRUD app with two user roles tests quickly. A multi-tenant platform with role hierarchies, workflow engines, file processing, and third-party integrations takes far longer, because each of those features creates attack surface a tester has to work through manually. Authentication tiers matter especially: every distinct role needs testing for privilege escalation and cross-tenant data access. Testing Methodology (Black Box, Grey Box, White Box) Black box testing gives the tester nothing but a URL, grey box adds credentials and documentation, and white box adds source code and architecture diagrams. Grey box is the default for SOC 2 and usually the best value, since the tester spends time exploiting rather than discovering. White box costs more upfront but finds deeper issues. Black box sounds rigorous but often wastes paid hours on reconnaissance an attacker would run for free. Depth of Testing and Manual vs. Automated Approaches Automated scanning finds known vulnerability patterns. Manual testing finds business logic flaws, chained exploits, and authorization gaps that no scanner catches, and it’s the part auditors and security-literate customers actually value. The ratio of manual work to automation is the honest explanation for most price differences between two quotes covering the same scope. Tester Credentials and Firm Reputation Senior testers holding OSCP, GPEN, or CREST credentials bill higher rates, and firms with recognized methodologies charge a premium for the credibility their letterhead carries

Two compromised versions of LiteLLM sat on PyPI for roughly 40 minutes on the morning of March 24, 2026. That window was enough to capture secrets from around 434,000 CI/CD pipeline runs across nearly 2,500 organizations, including AWS, Samsung, Cisco, Salesforce, Siemens, and Deloitte. In August, researchers at CloudSEK and Hudson Rock confirmed they had obtained the raw exfiltrated data: a 153GB archive containing 433,909 files of environment variables, cloud keys, Kubernetes secrets, and API tokens harvested live from running pipelines, as covered by Help Net Security’s reporting on the credential archive. If LiteLLM runs anywhere in your stack, or you touch any AI proxy infrastructure at all, you need answers to three things: whether you were exposed, what to rotate first, and whether the rotation you did back in March actually held. That last one matters more than it sounds, because “we rotated everything” has already burned at least one very large company. How the Breach Happened The attack didn’t start with LiteLLM. On March 19, 2026, a threat group called TeamPCP compromised the build pipeline of Trivy, a vulnerability scanner half the industry runs, and pushed a poisoned release. LiteLLM’s own CI pipeline ran Trivy, so the poisoned scanner had legitimate read access to the project’s runner environment. The attackers used that to steal LiteLLM’s PyPI publishing tokens and ship two malicious releases of their own: versions 1.82.7 and 1.82.8. KICS and the Telnyx Python SDK got hit in the same campaign. The payload design is the part worth studying. The malicious package dropped a .pth startup hook into site-packages, so the code ran the moment any Python interpreter started on the machine, whether or not anything imported LiteLLM. From there it harvested environment variables, read local credential files like .aws/credentials and .kube/config, tried to move laterally across Kubernetes clusters, and installed a systemd backdoor dressed up as a generic telemetry service. InfoQ’s coverage of the PyPI compromise put downloads of the compromised release above 40,000. For scale, LiteLLM normally gets downloaded around 3 million times a day. The exfiltration had a nasty fallback, too. According to CloudSEK, stolen data was encrypted and sent to a typosquatted domain, and when that failed, the malware created a public repository inside the victim’s own GitHub account and uploaded the loot as a release asset. Some companies were publishing their own secrets to the open internet and had no idea. Worth Knowing: The malicious code only existed in the PyPI artifacts. The GitHub source repository stayed clean the whole time, so a developer reviewing the code on GitHub saw nothing wrong. Source review isn’t artifact verification. If you don’t check that what the registry serves matches the upstream source, this class of attack is invisible to you. How to Check If You Were Exposed Three checks, from quickest to most involved. 1. Confirm whether the compromised versions ever ran The malicious versions went live on PyPI at 10:39 UTC on March 24, 2026 and got quarantined about 40 minutes later. The project’s advice: treat any install from that day before 16:00 UTC as suspect. Search your lockfiles, pip caches, SBOMs, and container image histories for 1.82.7 and 1.82.8. And check your internal artifact mirrors. An Artifactory or Nexus proxy that cached the bad release in March can keep serving it internally long after PyPI pulled it. Keep the .pth mechanism in mind when you scope this. The question isn’t “which applications import LiteLLM,” it’s “which machines had the package installed at all,” because every Python process on an infected machine triggered the payload. 2. Hunt for persistence Rotation is pointless if the attacker still has a foothold. Check developer machines, CI runners, and containers for unauthorized .pth files in site-packages and for suspicious systemd units, especially anything posing as a system telemetry service. And review activity from March 24 onward, not just the 40-minute window. Persistence is there so the access outlives the infection. Pro Tip: Don’t limit the persistence hunt to live machines. Base container images rebuilt in late March may have baked the payload into every image derived from them since. Scan your image registry for the affected LiteLLM versions and for unexpected .pth files, then trace which running workloads came from flagged images. 3. Check whether your secrets are in the dump Hudson Rock has published a domain lookup tool and is running ethical disclosures for affected organizations, and CloudSEK maintains a high-confidence victim list. Use them, but know their limits. Attribution in this dataset is genuinely hard. One dump with a siriusxm.com committer email actually traced, through its self-hosted GitLab endpoints, to AdsWizz, a SiriusXM subsidiary. And a large share of the dumps are generic pipeline configurations with no identifying domain, email, or server name at all. Absence from a victim list is not evidence of absence. If your pipelines ran the compromised versions, assume exposure no matter what a lookup tool tells you. What to Rotate, in What Order The guidance from both research teams is blunt: treat every secret the LiteLLM environment could reach as compromised. That covers secrets on disk, in memory, injected into CI jobs, and anything retrievable through instance metadata services. Work down by blast radius: Priority Credential type Why it comes first 1 Cloud IAM keys (AWS, GCP, Azure) Direct control of infrastructure, data stores, and billing. This is where attackers monetize fastest. 2 GitHub and GitLab PATs, package publishing tokens These let an attacker poison your releases and turn your company into the next link in the supply chain. 3 Kubernetes service account tokens and kubeconfigs Lateral movement across clusters was built into the payload, not a theoretical risk. 4 Database passwords and third-party API keys Dumped in plain text in the archive, often with no attribution, so nobody will warn you they leaked. 5 AI provider API keys Billing abuse, quota theft, and access to whatever data flows through your LLM routing layer. One word matters more than the rest of this article: revoke, don’t just rotate. That