Overview#
The trusted moment in software supply chains is the one nobody watches: npm install. A single install can run arbitrary lifecycle scripts, read your environment, touch tens of thousands of files and open a socket, all before you've written a line of code against the package. The declared behaviour is a package.json; the actual behaviour is whatever the author felt like on the day.
PoisonBox is a recorder for that gap. It's two halves joined by one contract:
- The sandbox (Rust, Linux/KVM). A disposable Firecracker micro-VM per scan captures full telemetry with eBPF in a pinned guest kernel, controls egress (air-gapped by default, so a scan never completes an attacker's callback), and writes a deterministic flight-recorder bundle.
- The visualiser (the renderer at /scan/). It replays a flight recorder as a live graph: benign activity stays calm and teal, a malicious run reads a secret, spawns a process, and the outbound connection ignites red.
PoisonBox is a research module of OSPulse. Where OSPulse watches your whole dependency estate continuously, PoisonBox goes deep on a single package on demand.
The threat#
Install scripts run the moment you add a dependency, with your shell's privileges, before review and before you import anything. That's how ua-parser-js shipped a credential stealer to millions in 2021, and how event-stream and the polyfill.io hijack got in. A freshly poisoned version has no CVE yet, so a vulnerability scanner waves it straight through.
PoisonBox assumes the install itself is hostile and gives it a room with no windows. It doesn't guess from metadata; it runs the thing and watches from underneath.
How it works#
Four steps, one recording:
- Isolate. Each install runs in its own throwaway micro-VM with hardware-level isolation (Firecracker on KVM). It boots, does its worst, and is deleted. Nothing it does reaches your machine, your network, or your keys.
- Observe. A kernel-level eBPF probe records every process spawned, file opened, secret read and outbound connection. It watches from beneath the code, so there's nothing the package can do to talk its way out of being seen.
- Diff. The run is compared against how a package is supposed to behave. Secret reads and non-registry connections are marked as deviations, with the causal thread joining a secret read to the connection that carried it.
- Score. A risk score where every point links back to the exact event that earned it. No black box: you can audit the verdict line by line.
The sandbox seeds decoy credentials (a fake SSH key, npm token and AWS credentials). Any package that reads them is caught. Because egress is air-gapped by default, an attacker's callback is recorded as an attempt and never actually completes.
Run it#
Two facts sit side by side: the contract, CLI and renderer build and run on any OS, but an actual scan needs a Linux host with KVM (Firecracker + eBPF). On a Windows dev box that host is a local Hyper-V VM. So "run it" splits into what you do on Windows and what happens on the sandbox host.
Fastest: watch a capture (no setup)
Open a live scan at /scan/ or the side-by-side /gallery/. Or read a bundled recording from a clone, no GPU or sandbox needed:
cargo run -p blastradius-cli -- view samples/recordings/fixture-malicious.blastrec
Build & test (any OS)
cargo test --workspace # contract, scoring, egress policy
cd web && npm install && npm test # loader/validator + determinism
On Windows the sandbox, telemetry and egress crates compile as inert stubs, so blastradius scan tells you it needs a Linux/KVM host rather than pretending.
The one-liner (Windows-first)#
Scan any published npm package from Windows with a single command. It finds the sandbox VM, resolves and downloads the package from the registry, runs its real install inside a throwaway micro-VM under eBPF capture, and brings the verdict and the recording back:
.\scripts\scan.ps1 ua-parser-js@1.0.32
.\scripts\scan.ps1 lodash -Out C:\scans\lodash.blastrec
What you get back:
ua-parser-js@1.0.32 risk score 19
- read a secret during install
- connected to a non-registry host
recording: .\ua-parser-js_1.0.32.blastrec
blastradius-kvm VM must be running (the script auto-finds its IP, or pass -HostIp). Nothing untrusted ever runs on Windows: the package is downloaded with curl and only executed inside the micro-VM.Under the bonnet, scan.ps1 pushes the scanner to the host and runs scripts/sandbox/scan-npm.sh, which resolves the tarball from the registry (curl + python3, no host npm needed), then runs its lifecycle scripts inside the guest under capture.
From scratch: the sandbox host#
To stand up the Linux/KVM host on a Windows machine, PoisonBox uses a local Hyper-V nested-virt Ubuntu VM, provisioned reproducibly from a cloud image (no installer clicking). In elevated PowerShell:
# download Ubuntu 24.04 cloud image -> C:\Hyper-V\cloud\noble-cloudimg.img
ssh-keygen -t ed25519 -f C:\Hyper-V\blastradius-kvm\id_ed25519 -N '' -q
# put the pubkey in cloud\seed\user-data, a unique instance-id in cloud\seed\meta-data
docker run --rm -v C:\Hyper-V:/hv -v <repo>:/work blastradius/telemetry-dev bash /work/scripts/hyperv/build-cloud-vhdx.sh
scripts\hyperv\create-cloud-vm.ps1
Then on the VM, install the toolchain, build, and produce the two one-time guest assets:
ssh -i <key> br@<ip> 'bash -s' < scripts/hyperv/guest-setup.sh # Firecracker + Rust + /dev/kvm
cargo build --workspace # -> target/debug/capture
bash scripts/sandbox/build-guest-kernel.sh # -> ~/fc/vmlinux-br-6.1.128 (tracing kernel)
bash scripts/sandbox/build-node-rootfs.sh # -> ~/fc/ubuntu-24.04-node.squashfs
After that, scan.ps1 from Windows drives it, or run the harness directly on the host:
BR_SQUASH=$HOME/fc/ubuntu-24.04-node.squashfs \
bash scripts/sandbox/scan-npm.sh ua-parser-js@1.0.32 ~/out.blastrec
Reading a verdict#
A scan resolves to one risk score and a verdict band:
| Verdict | Score | Meaning |
|---|---|---|
| Clean | 0 | No secret touched, no socket opened. The analysis says clean rather than inventing suspicion. |
| Review | 1 to 9 | Something worth a look: a secret read or a lifecycle shell step, but no exfiltration seen. |
| Blocked | 10+ | Secrets read and carried to a non-registry host. Caught mid-theft. |
The score is never the whole point. The causal thread, from a secret read to the outbound connection that carried it, is what separates a package that authenticates from one that robs you, and it's drawn from the run, not asserted. Every point in the score cites a real captured event.
Published scans get their own page at /v/<eco>/<name>/<version>/, for example evil-analytics@1.0.0 (blocked), ua-parser-js@0.7.29 (review) and left-pad@1.3.0 (clean).
The flight recorder#
A scan writes a .blastrec bundle, the sole interface between the sandbox and the visualiser. It's three files:
header.jsonrecords the package, ecosystem, egress mode, outcome and duration.events.jsonlholds one line per captured event: process spawns, file opens, secret reads, network connects, each with its sequence, timestamp and pid.findings.jsoncarries the risk score, the evidence-linked points, the deviations, and the key moment. Every point references an event byevent_seq.
The bundle validates against a JSON Schema that both the Rust writer and the loader check, so nothing renders that the recorder didn't capture. That's the honesty rule, enforced.
OSPulse integration#
OSPulse and PoisonBox answer two halves of one question. OSPulse says "ua-parser-js is on the KEV list"; PoisonBox says "...and here's it reading your SSH key and phoning home." The ownership line is clean:
- OSPulse holds the signal. A small
verdict.jsonper package (risk score, verdict, a link) that feeds OSPulse's score and renders a View in PoisonBox button on the package page, highlighted for KEV and CVE packages. - PoisonBox holds the evidence. The full recording and the interactive view live here, at
/v/<eco>/<name>/<version>/. "View in PoisonBox" takes you into PoisonBox.
A verdict caches per (ecosystem, name, version): a version behaves the same for everyone, so it's scanned once and the verdict is served to all. That keeps the cost bounded by the number of unique package-versions that matter, not by users or requests, which is what makes it cheap. Today the cache is populated by scanning on the local host and publishing static verdict pages with scripts/publish-verdict.ps1; live on-demand cloud scanning is a later, revenue-funded step.
Status & roadmap#
PoisonBox is an active R&D build. What's real today: the flight-recorder contract and its tests, disposable-microVM isolation, eBPF capture, the declared-behaviour diff and evidence-linked scoring, the renderer, the Windows-first one-liner, and static per-package verdict pages.
Honestly, what isn't done yet:
- Egress capture. Scans run air-gapped, so an outbound callback is recorded as an attempt but never completes. Resolving destination hosts and SNI on a live-but-sinkholed network is a pending guest-kernel fix (a virtio-mmio enumeration issue).
- The literal flagship. The real, original ua-parser-js malicious tarball capture is gated on sourcing the actual pulled artefact; the current ua-parser-js page is a faithful reconstruction of the documented payload, labelled as such.
- Live on-demand scanning. The hosted "scan any package on request" service is a paid-tier future; today, verdicts are pre-scanned and cached as static pages.