Blog · Oleg Sidorkin · 28 September 2026
The machine layer under NVIDIA OpenShell
NVIDIA just validated the category I have been building in for months. I was impressed. Then I read their code, measured what sits underneath, and found the gap between what a hundred companies think they have and what they actually do.
The morning I was impressed
On 28 September 2026, NVIDIA announced the Open Agent Safety Platform: NVIDIA OpenShell, an open-source runtime that sandboxes AI agents, and NVIDIA Sentry, an out-of-band watchdog on BlueField-4 DPUs. A hundred partners. Anthropic, Microsoft, JPMorganChase. Three robotics builders: Figure, Gecko Robotics and Skild AI.
My first reaction was excitement. When the company that builds most of the world's AI compute says agents need an enforceable boundary outside the model, the category I have been working in is no longer a bet. It is consensus. Jensen Huang said "safety requires full-stack engineering," and he is right. I have been saying that for months, to smaller rooms.
Then I read the code
OpenShell is good engineering. Landlock restricts filesystem access. Seccomp filters syscalls. A policy proxy evaluates every network connection before it leaves the sandbox. Credentials are injected for approved endpoints; the agent never sees them. For containing software agents on servers, this is the right design.
But I wanted to know what it is built on. Not just the four crates that implement the sandbox boundary, but every layer underneath: the policy engine, the TLS stack, the kernel modules, the container runtime. Because that is what a hundred companies, including three robotics builders, are about to trust with something that matters.
So I measured it. Every layer.
What sits under OpenShell, measured
OpenShell does not use eBPF. The sandbox boundary is built from four things: Landlock (a Linux security module, ABI v3, kernel 6.2+) for filesystem isolation, seccomp (classic BPF filters) for syscall filtering, regorus (an embedded Rust Rego engine, not the OPA server) for network policy, and the container runtime's network isolation for the outer fence.
| Layer | What it does | Lines | Proven? |
|---|---|---|---|
| OpenShell boundary (4 crates) | Sandbox, isolation, proxy, supervisor | 125,322 | No (485 unsafe blocks) |
| regorus | Policy engine | ~15,000 | No |
| landlock crate | Rust binding to Landlock | ~4,000 | No |
| seccompiler | Seccomp library | ~3,000 | No |
| rustls | TLS termination | ~30,000 | No |
| 805 Rust dependencies | Everything from tokio to serde | millions | No |
| Linux kernel Landlock LSM | Filesystem isolation | ~7,000 | No |
| Linux kernel seccomp | Syscall filtering | ~4,000 | No |
| Linux kernel namespaces | Network isolation | ~5,000 | No |
| Linux kernel netfilter | NetworkPolicy enforcement | ~30,000 | No |
| Container runtime | Docker/Podman/K8s | ~1,000,000 | No |
| The Linux kernel itself | The substrate | ~30,000,000 | No |
Every layer in that table must be correct for OpenShell's containment
to hold. Not just OpenShell's code. The policy engine. The kernel
modules. The container runtime. The Linux kernel. The 805 Rust
dependencies. Nobody has proven any of them memory-safe. The
unsafe code in OpenShell's own boundary alone (1,328
lines across the two crates that call the kernel) is larger than
my entire proven interpreter.
And here is what worries me
The eBPF verifier is the closest production analog to what OpenShell builds. It was also a careful, well-engineered admission checker sitting inside the Linux kernel. It accumulated CVE-2023-2163 (CVSS 10.0, a container escape), five more soundness CVEs in the first half of 2026 alone, and 73 bug fixes between kernels 6.3 and 6.13. Upstream eventually stopped treating it as a boundary against untrusted local users. An admission checker that infers safety accumulates CVEs at its own ABI. One that checks proofs does not.
Seven edges in the stack today
1. Landlock degrades silently. The filesystem policy
defaults to BestEffort
(crates/openshell-core/src/policy.rs:89-93). On a kernel
without Landlock v3 (RHEL 9.x, gVisor, an older kernel), the sandbox
runs without filesystem restrictions and logs a finding.
Landlock requires Linux 6.2 or a vendor backport. How many
production deployments run that kernel?
2. The seccomp filter is default-allow. It blocks
specific escape primitives and allows everything else
(crates/openshell-sandbox/src/sandbox/linux/seccomp.rs:6).
The agent can still call every syscall not on the hand-curated block
list. New syscalls are admitted by default.
3. The outer network fence is the container runtime's job.
Docker, Podman or Kubernetes enforce it, not OpenShell. A misconfigured
NetworkPolicy or a Docker override means the fence is gone. The four
guarantees are asserted by the driver, not proven by OpenShell
(architecture/sandbox.md:105-128).
4. The isolation code is not proven. 125,322 lines
of Rust across 84 files, with 485 unsafe blocks in 68
of them. The kernel-ABI crates alone carry 1,328 lines of unsafe
code. Trusted by inspection, not by proof.
5. The policy engine is regorus, not OPA. An embedded Rust implementation of Rego, not the server that has been tested at scale. A single bug in regorus affects every network decision. Policy evaluation is serialized through a Mutex.
6. TLS termination puts all plaintext in the supervisor. The proxy terminates TLS with an ephemeral CA to inspect HTTP traffic. The supervisor holds the plaintext of every inspected flow. If the supervisor is compromised, all traffic is visible.
7. L7 inspection has blind spots.
Server-to-client MCP and JSON-RPC responses are relayed but not
parsed. Raw WebSocket frames are passthrough. protocol: tcp
does not inspect TLS SNI or HTTP Host, so a client can select
another tenant behind an approved front door.
Why this worried me: containment is not safety
Here is my problem. OpenShell works. It blocks known attacks. It confines agents that behave within the rules. For software running on servers, where the worst case is a data breach, this is fine. But NVIDIA called it a safety platform, and a hundred companies are adopting it as if containment and safety were the same thing.
They are not. Containment is a mechanism: it blocks things. Safety is a property: you can rely on it. OpenShell provides containment that rests on 40 million lines of code nobody has proven correct. The filesystem layer silently disappears on older kernels. The syscall filter is default-allow. The outer network fence depends on your container runtime being configured correctly. The code that talks to the kernel has 485 unsafe blocks trusted by inspection. Every one of those is a crack in the containment, and the containment is all there is.
The eBPF verifier already showed what happens next. Carefully engineered containment, adopted as a security boundary, degraded over time until upstream abandoned it. The pattern repeats: a containment layer that is not proven accumulates bugs at its own ABI, and each bug narrows the gap between what the layer promises and what it delivers.
What worries me is the gap between what companies think they have and what they actually have. A company that deploys OpenShell says "we have agent containment" and stops thinking about it. The appearance of the boundary substitutes for the boundary. When a zero-day in Landlock or seccomp or the container runtime breaks through, nobody is looking, because everyone believes the wall is there. The gap between containment and safety is where things fail, and it is widest when you do not know it exists.
What I did about it
I did not build a better Landlock, seccomp, or regorus. I built a different enforcement domain that does not depend on any of them.
Endstop's entire trusted base is 1,206 lines of Rust across two crates. Zero unsafe blocks. Proven memory-safe by machine-checked proof. Running on an FPGA with no operating system, no container runtime, no 805 dependencies. The interpreter is 667 lines, Kani-proven across seven properties. The monitor is 539 lines. The frame parser is proven total over arbitrary input. If a bug exists in the proven core, a machine-checked proof would have to be wrong, not just my code.
| OpenShell stack | Endstop | |
|---|---|---|
| Trusted base | ~40,000,000 lines (all layers) | 1,206 lines |
| Unsafe blocks | 485 in OpenShell alone | 0 |
| Kernel dependency | Linux (~30M lines C) | None |
| Container runtime | Docker/Podman/K8s (~1M lines Go) | None |
| Dependencies | 805 Rust crates | 2 crates (ours) |
| Verified by | Inspection + Z3 on policy changes | Kani machine-checked proofs |
| Substrate | Linux on x86/Arm | FPGA fabric (ECP5) |
The ratio is not 104x anymore. It is 33,000x. And the difference is not just size: Endstop's 1,206 lines are proven. OpenShell's 40 million are not.
A zero-day in Landlock does not compromise our fabric. A bug in regorus does not touch our envelope monitor. A misconfigured Kubernetes NetworkPolicy does not affect the ECP5. The failure surfaces are independent, because the substrates are independent.
The way out: a thin trusted line out of the box
Start with the honest premise: you cannot trust the machine OpenShell runs on. If an AI agent executes code there, a kernel zero-day, a container escape, or a bug in any of the 40 million lines underneath can hand the agent the entire host. OpenShell's policies are useful, but they are code on that same host, and code on a compromised host is the compromised host. The agent's server has to be treated as potentially hostile, which means it has to be airgapped from anything it must not touch.
Endstop runs on a separate ECP5 board outside that box. The only path from the agent to the actuator is a serial line into that board. Everything on the server side, including OpenShell and the agent, is untrusted from the board's perspective.
This is the same principle a separation kernel uses for high-assurance systems: the trusted boundary is a thin, dedicated communication line, not a policy engine running on the same machine as the code it polices. The line is deliberately simple. Bytes in, CRC32-checked, parsed by a proven-total frame parser, checked against the envelope, permitted or refused. No network stack. No kernel. No dependencies. The board does not need to know whether OpenShell is intact, because it does not trust OpenShell in the first place.
OpenShell is still useful on the server. It limits what the agent can read, what it can call, who it can talk to. It buys time and reduces the attack surface. But it does not bound what the agent can cause when the server falls, because it runs on the server. What bounds the consequences is the board on the other end of the serial line, checking every command, because the board is not the server and shares nothing with it.
If a zero-day takes the server, the board holds the machine. If the board has a bug, OpenShell still contains the agent on the server. Two machines, two substrates, zero shared code, and a serial line between them that carries nothing but checked commands.
That is the difference between containment and safety. Containment is a policy engine on the same host as the code it polices. Safety is a proven boundary on a separate machine, connected by a thin trusted line, that holds even when the host falls.
To the robotics builders
Figure, Gecko, Skild: NVIDIA named you as building with OpenShell for agents that act in the physical world. Your agents are now contained on the server. Your machines are not, and the server cannot be trusted from the machine's perspective.
When the agent proposes a joint command, our fabric decides whether the machine moves. The proposal is the agent's. The authority is the fabric's. Even if a kernel zero-day takes the whole server, the board on the other side of the serial line still checks every command, because it never trusted the server in the first place.
Containment on the server is a good start. But if the agent runs on that server, the server is the thing you cannot trust. The machine needs its own boundary, on its own hardware, behind its own line.
Notes
All OpenShell code references are to the public repository as of 28 September 2026. I measured the Rust LOC, unsafe blocks, and dependency count myself from that repository. The kernel LOC for Landlock, seccomp, namespaces and netfilter are from the Linux source tree; the ~30M total kernel size is from public kernel statistics. The eBPF verifier CVE history is from public CVE databases. The straw-house metaphor is mine. Our proof and evidence record, including a live attack target that has held under every attempt, is published.