What Mini Shai-Hulud Teaches Us About AI Supply Chain Security
Brad Micklea · May 14, 2026

What Mini Shai-Hulud Teaches Us About AI Supply Chain Security

On May 12, 2026, security researchers disclosed that a worm dubbed Mini Shai-Hulud had compromised 42 packages and 84 versions across the npm ecosystem including TanStack, OpenSearch, UiPath, and others. A coordinated but possibly separate attack compromised PyPI packages from Mistral AI and Guardrails AI. The TanStack compromise carries CVE-2026-45321 with a CVSS score of 9.6 (critical).

What sets this attack apart is that, by compromising the CI system itself, the malicious packages were published through the project’s legitimate GitHub Actions release pipeline. That meant that they carried valid (not forged) SLSA Level 3 provenance attestations and signatures.

Over the last several years supply chain attacks have been on the rise, but until now malicious packages haven’t carried valid attestations or signatures.

This post walks through:

  • What happened

  • Why the existing answer (sign your artifacts, attach SLSA provenance) was structurally insufficient

  • What an organization needs to do to defend against this class of attack

  • Why the architecture Jozu has built was designed for this failure mode

What Happened

The attack chain, as reconstructed by TanStack’s post-mortem and the researchers who analyzed the malware:

  1. Attackers staged a malicious payload in a GitHub fork of the TanStack router repository.

  2. They exploited the pull_request_target trigger combined with GitHub Actions cache poisoning to gain code execution inside the legitimate TanStack CI environment.

  3. From the runner’s process memory, they extracted an OIDC token that the legitimate publish workflow uses to authenticate to the npm registry.

  4. They injected the malicious code into the published npm tarballs, then used the project’s own release pipeline to push the compromised versions with valid signing and attestations.

  5. The malicious code, an obfuscated router_init.js, profiled the execution environment and ran a credential stealer targeting cloud provider credentials, cryptocurrency wallets, AI tool credentials, messaging app tokens, and CI system secrets.

  6. Stolen credentials were then exfiltrated to a domain hosted on Session Protocol infrastructure, chosen because enterprise filters are unlikely to block decentralized messaging services.

  7. The worm established persistence in Claude Code and VS Code, and monitored GitHub tokens to re-exfiltrate any new ones, while injecting malicious GitHub Actions workflows to serialize repository secrets and exfiltrate those.

  8. The worm then located publishable npm tokens with bypass_2fa set to true, enumerated every package published by the same maintainer, and exchanged a GitHub OIDC token for per-package published tokens. It then repeated the attack on each of those packages, generating valid SLSA Level 3 attestations for every compromised version.

While the mechanics of how the attack replicated through the npm repository are clear, I wasn’t able to find public documentation on how the attack impacted Mistral AI, Guardrails AI, or other packages on the PyPI registry. It may be that these were parallel attacks, or it may be that the exfiltrated credentials from npm included credentials for packages in the PyPI ecosystem. This possibility is more worrying because many teams store npm, pypi, and other credentials together in dot-files.

Why SLSA Level 3 Wasn’t Enough

SLSA is a framework for establishing provenance of packages to tell prospective users how and where something was built, and has four levels. SLSA Level 3 which was used in the compromised build systems attests to three things about an artifact:

  1. That it was produced by a specific build platform

  2. That the build process worked as described in the provenance

  3. That the build platform was hardened against tampering during the build itself

Importantly, none of the SLSA levels are designed to tell you if the artifact itself is safe. That question is outside the scope of SLSA. Because the attacker used a well known and trusted build system the SLSA Level 3 attestation is valid.

Understanding the scope and limits of every tool is critical to security. SLSA is valuable but alone it’s insufficient.

What’s Needed Post Mini Shai-Hulud

The root failure in Mini Shai-Hulud wasn’t missing signatures or weak attestations. It was that every enforcement point lived inside the same trust boundary the attacker controlled - by owning the pipeline they also owned the policy.

To defend against this, enforcement must always happen outside the trust boundary of the thing being verified. That’s the through-line for each of the four protections that follow.

1. Expanded Attestations

Attestations are great for security because they’re signed and immutable. However, they need to go beyond attesting to the build system and start attesting to the quality of the artifact itself.

There are many different security scanners, but very few output their scan reports as an attestation (Jozu does, but we’ll get to that later). For example, a scanner that found an obfuscated file (router_init.js) should add an attestation that highlighted that as a security vulnerability.

For AI artifacts, this means scanning for serialization attacks, embedded backdoors, prompt injection susceptibility, training data integrity, and behavioral anomalies. The attestations from those scans should travel with the artifact and be verifiable independently of the pipeline that produced the artifact.

If these are in place then the first automated check can cover:

  1. Where the build was executed

  2. How the build was executed

  3. Whether the artifact created by the build was safe

If any of those are suspect, then the package should be quarantined.

2. Policy That Lives Outside the Build Pipeline

The key to these attacks is that the attackers controlled the build pipeline. Any policy enforcement that depended on the pipeline was compromised by definition.

The structural fix is for policy to live as its own artifact - separately authored, separately signed, separately distributed - and for enforcement to happen at a point downstream of the build, by a system that does not trust the build environment.

Compromising the build pipeline shouldn’t give an attacker the ability to approve the build’s output.

3. Pre-Execution Verification at the Point of Use

The Guardrails AI compromise executed its payload on import so the credential stealer ran immediately. An enforcement point that operates only at the build stage is too far upstream in this case.

What’s needed is a second level of verification where the execution happens - whether that’s in the cloud, on a laptop, or at the edge. Before execution happens, the artifact’s signatures, scanning attestations, and policy approvals should be checked again.

Similar to #2 above, this execution system shouldn’t trust the build system or registry implicitly. It needs to have all the information needed to perform its own checks against its own policies.

This is why everything needed to evaluate safety at each level should be attached to the package as attestations that can be locally processed.

4. Runtime Containment That Assumes Sometimes Thigs Fail

Unfortunately Murphy’s Law is a law which means sometimes, things will go wrong. The last defence is a runtime environment that contains what a compromised artifact can do. The strongest isolation would be through hardware-level isolations you see in confidential computing. However, that’s beyond what most organizations can do, but a VM-level barrier keeps a potentially dangerous package or rogue agent away from sensitive information. Combine this with artifact admission limitations, tool-level access control, egress restrictions, and tamper-evident audit logs and you will have a local runtime that can limit the blast radius of a failure and captures enough information for you to reconstruct the problem and prevent it in the future.

For Mini Shai-Hulud, runtime containment would not have prevented the credential stealer from running, but it would have:

  • Prevented the stolen credentials from leaving the system

  • Kept IDEs and agents from persisting the attack

  • Produced an audit trail that named the attempted exfiltration domain

Why Jozu Chose this Architecture Two Years Ago

The four properties above describe Jozu’s product architecture, which predates this incident by nearly two years.

Artifact are Secured and Verified: Jozu Hub holds AI artifacts as ModelKits: OCI artifacts that carry the model, configurations, skills, prompts, datasets, and metadata as a single signed package. Hub scans every artifact with five integrated scanners covering serialization attacks, prompt injection, adversarial robustness, and other AI-specific vulnerability classes. The scan results are themselves signed attestations attached to the artifact. When the artifact is diffed against its prior version, additions of unexpected files are visible and policies can act against them before promoting to production.

Policies are Separate, but Verified: Policies in Jozu are cryptographically verified OCI artifacts themselves. They are authored separately from the build pipeline, signed by the policy authors, and distributed independently. A compromised model or agent cannot also produce an approved-by-policy attestation, because the policy artifact lives outside that pipeline and is signed by a different key chain.

A Secure Runtime With Independent Enforcement: Jozu Agent Guard, our secure runtime for AI, is the enforcement point downstream. It runs as a MicroVM on the laptop, edge, or IoT device, or as a Kata Container in a Kubernetes cluster. It re-verifies artifacts at the moment of execution against the policy artifact, not against a cached build-time state. Policies travel separately, but in parallel with the artifact, so security can be enforced locally - including in air-gapped environments where there is no control plane to call out to.

Tool invocations are governed by a ToolPolicy that controls what models and agents can do with MCP tools or tools provided through their own harnesses. Egress is governable and data and prompts can be cleaned of sensitive or dangerous information. Every action is recorded in a cryptographically chained audit log that is tamper-evident throughout. When a compromised artifact attempts to exfiltrate to an unexpected domain, the egress is blocked and the attempt is recorded.

This is the two-phase enforcement that Jozu uses: supply chain verification in Hub before execution, and policy enforcement during execution in Agent Guard. The two phases are reinforced because they don’t implicitly trust each other.

What Jozu Does Not Solve

Jozu is focused 100% on AI artifacts and that supply chain. We don’t handle code packages and repositories. If a developer executes npm install on a compromised TanStack package or pip install a compromised PyPI library directly to their workstation, Jozu is not in that path unless your organization has chosen to route those installs through our curated registry. We are not a replacement for SCA tools, or EDR.

Jozu does not replace identity and access management or data loss prevention. We govern AI artifacts and AI runtime behavior. IAM and DLP are needed for the human-shaped access patterns they were designed for. We sit alongside them, in the gap they were never built to cover.

What we do cover is the AI artifact path: from the registry where ModelKits, models, prompts, datasets, and MCP servers are stored and scanned, through policy-as-OCI-artifact distribution, to runtime enforcement at every deployment target including air-gapped environments. For an organization deploying AI, that is where supply chain security and runtime governance meet, and where existing tools were not designed to operate.

Share this post