Endstop

Rev 0.3.4 · in integration

A containment layer for untrusted AI.

Let the model run.

Bound what AI can do: run AI-generated code with explicit limits on what it can access, how many instructions it can execute and which actions it can request.

Endstop is a hardware containment system. It combines a restricted runtime with a machine supervisor that evaluates state, motion and proposed actions. Machine control is the first application: let models generate useful behavior because dedicated hardware holds control authority.

Alongside the public challenge, we attacked the deployed board ourselves: eight attack classes on the deployed board — interpreter bounds, the admission cost model, the PMP wall, the fabric authority, the wire protocol — with the boundary held under every one, and the one defect we found quantified, fixed and re-verified in public. Read the series record →

NVIDIA's new Open Agent Safety Platform contains software agents on servers. Endstop is the layer beneath it — the one that holds control authority over the physical machine. Read why containment has failed before and what the stack needs beneath OpenShell →

Start the fit assessment → Inspect the evidence →

Implementation checklist and roadmap →

§1

Useful programs. Explicit limits on authority.

AI-generated programs can respond to changing conditions, choose actions and execute logic. Giving them access to a machine also gives errors, prompt injections or hostile output a path to physical effects.

Endstop is designed to keep that authority bounded regardless of the model’s intent. Permissions and operating limits are enforced outside the generated program, so the program cannot expand its own access or redefine what the machine is allowed to do.

That behavior should also be observable. Endstop is designed to expose attempted actions, enforcement decisions and refusal reasons for inspection, so an operator can understand what the program tried to do and how the supervisor responded.

§2

Contain execution. Constrain external actions.

A model pursuing unintended goals could exploit unknown vulnerabilities in the software surrounding it. Endstop’s design removes the inherited software stack from the containment boundary: operating systems, middleware and general-purpose runtimes.

The execution environment is rebuilt from the hardware up as a compact runtime with explicit rules and machine-checked properties. This makes the code responsible for containment small enough to inspect and reason about formally. A separate machine supervisor evaluates the physical behavior that generated programs propose.

How Endstop bounds generated code and its proposed actions
LayerWhat it controls
Purpose-built runtimeA compact execution core on dedicated hardware. Explicit memory bounds, instruction budgets and permissions define what untrusted code can do; named properties are checked with formal methods.
Machine supervisorEvaluates motion history, rate, acceleration, predicted stopping positions and patterns that can excite resonance. Feedback checks cover measured position, rate and tracking error. These checks determine whether proposed motion is permitted.

In the intended deployment, the model performs inference only: tokens in, tokens out. Endstop is the sole route from that output to external action. The inference system has no independent tool execution or machine-control path.

Explore the architecture →   Why command checks alone are insufficient →

§3

Inspect and challenge the boundary

The execution core and its proof harnesses are published. Reviewers can examine the rules governing memory access, execution budgets and permitted effects, then run the checks themselves.

Formal checks address named properties under explicit assumptions. Target execution tests exercise the compiled implementation; dated hardware records document what ran and what was observed. Together, these connect the containment rules to inspectable code and recorded behavior.

The public execution challenge invites attempts to violate those rules on an emulated target. Its attempt records make the submitted programs and outcomes available for review, including failures and corrections.

Examine the evidence and its scope →
Read the execution core →   Challenge the boundary →

§4

Evaluate useful work under containment

A useful containment system must preserve the model’s ability to accomplish a task. An evaluation should establish three things:

  • Utility. Does the generated program complete the intended work, with acceptable response times and few unnecessary refusals?
  • Enforcement. Do attempts to exceed authority and changes in machine state trigger the expected intervention?
  • Observability. Can the team inspect attempted actions, supervisor decisions and the resulting behavior?

Machine control is our first evaluation focus. With a design partner, we agree one task, its operating limits and representative fault cases, then compare Endstop with the existing approach. The evaluation is scoped to produce repeatable tests, a measured comparison and a recommendation to proceed, revise or stop.

Start with a free fit assessment to establish whether the machine and task suit the design. Any subsequent evaluation has an agreed scope and acceptance criteria.

Discuss a task for evaluation →   See the evaluation approach →

Choose the depth you need

Explore Endstop

Frequently asked questions

What is Endstop designed to do?

Endstop is a hardware gate being developed to place bounded authority between a model-driven planner and a machine. Its design combines bounded program execution with a machine-specific command envelope. Read how the two boundaries work.

Is Endstop certified or ready for production?

Not yet. Endstop is an uncertified engineering prototype, offered through design-partner evaluations rather than for sale. It works alongside emergency stops, rated drive functions and a machine-specific safe-state design, and does not replace them. See the current evidence.

Which machines are suitable for an initial evaluation?

The current pilot focuses on a fixed-base robot task with an accessible command path, representative commands and defined limits. Existing rated controls stay in place. Interface access and the machine’s behavior after a refusal determine whether an evaluation is suitable. Check your machine’s fit.

How do we start?

Start with a free fit assessment: describe the machine, what the model decides and the physical effect you want bounded. The outcome is a written fit decision and comparison plan. Any later paid phase is separately scoped and can stop independently. Start the fit assessment.

Design partners

Bring a task where AI-generated control could help.

Tell us what the machine should accomplish, what the model would decide and which limits must hold. The initial assessment is free. You receive a written fit decision and, where appropriate, a proposed evaluation scope.