Endstop

Rev 0.3.4 · in integration

The problem

Contain the execution. Constrain the consequence.

A model by itself maps input tokens to output tokens. The output becomes dangerous only when an execution layer interprets it as code, a tool call or a machine command. Endstop rests on two independent boundaries: one bounds that execution; the other decides which proposed physical effects may reach the machine. Neither boundary is sufficient alone.

§1

Put the authority boundary beyond the frontier-model host

A frontier model without an execution path has no physical authority. It evaluates a mathematical function and produces tokens. Tokens do not open a valve, move a joint or send a payment by themselves, just as a brain disconnected from muscles cannot move a machine. The authority belongs to the surrounding system that interprets those tokens and acts on them.

Containers, virtual machines and operating-system sandboxes are valuable. They isolate workloads and reduce accidental reach. The problem is not that they do nothing. The problem is letting the same conventional host both interpret model output and hold the complete authority boundary for a machine.

That host includes accelerator drivers, model runtimes, orchestration code, parsers, network clients, files, credentials, tools and device interfaces. It changes quickly and must deliberately expose useful authority. A sandbox escape is only one failure mode. An agent runtime can interpret model output and spend authority it was legitimately granted after the model reads attacker-controlled text in a web page, ticket, sensor label or document.

This is the confused-deputy problem. In 1988 Norm Hardy described a compiler that held authority to write its billing file and was induced to spend that authority on a user's chosen path.1 The compiler did not become malicious and did not escape its process. It used a valid permission for the wrong purpose.

Robot attacks driven through language-model interfaces are no longer hypothetical. In 2024 researchers reported successful jailbreaks against three robot systems, including black-box access through a public interface.2 A separate provisioning flaw yielded root access across several Unitree platforms.3 No injury caused by a language model driving a robot has been documented, and we do not claim one. The evidence establishes attackability, not physical harm.

What it means to sandbox the model

Endstop does not host the model's inference computation. It effectively sandboxes the model's machine-facing agency. The model remains a token generator on the untrusted host; its output crosses the boundary as a model-authored control program and can act only through Endstop's small, verifiable execution layer and separate monitor. This claim depends on complete mediation: the host must have no alternate path to the actuator.

§2

Move the execution, not the model, into bounded authority

The execution layer forms one of the product's two authority boundaries.

A submitted program runs in Endstop's admitted instruction subset with 512 bytes of scratch memory, a fixed fuel budget and three capabilities: read state, propose a setpoint and request a signature. The loader rejects instructions, register references, memory-access forms and capability calls outside that vocabulary. At runtime, every actual memory address is checked again. Fuel ends every invocation, including one that loops.

This preserves the useful division of labor. The frontier model can remain large and probabilistic because its tokens carry no direct machine authority. Its effective agency is sandboxed at the point where those tokens become executable. The execution layer is small and deterministic because it answers questions a reviewer can enumerate: which instructions exist, which bytes can be addressed, how many steps may execute, and which effects may be requested. The host may be hostile. The program it submits still receives no file, socket, allocator, network client, device register or computed capability call.

Execution authority on the model host and inside Endstop
SurfaceFrontier-model hostEndstop execution layer
RuntimeOS, drivers, model and tool stacksmall admitted interpreter
Memorylarge, shared and application-dependent512-byte checked scratch region
Durationscheduler and application dependentfuel-bounded invocation
Effectsapplication permissions and toolsthree enumerated capabilities
Failure postureapplication dependentloader rejection or bounded halt
Isolation on the board

The interpreter, loader and monitor share one soft CPU and address space. Interpreter semantics bound the submitted program, and PMP protects monitor state: on the physical board, a user-mode sweep of every protected region boundary faulted with its documented cause. In the fabric-gate prototype, output permission already sits in fabric logic the processor cannot write itself into; the next milestone moves monitor state there too.

§3

Contained code can still command harm

Execution containment answers where code may reach and which effects it may request. It does not decide whether a permitted request is safe for this machine, state and moment. Each example below can remain inside its granted authority and still be harmful.

Resonance inside every limit

A setpoint stream whose instantaneous velocity never exceeds the configured limit, and whose per-tick delta never exceeds the rate limit, can still excite a drivetrain resonance. Mechanical Q typically does the rest. Per-axis kinematic limits are instantaneous predicates; they do not bound energy accumulated over time.

Cartesian composition

Every joint can sit inside its own position and velocity envelope while the tool tip traverses the space a person occupies. Speed-and-separation limits are defined in Cartesian space; joint-space limits are not a sound approximation of them. ISO/TS 15066 says so in its own normative note: a speed limit that monitors only the tool centre point "does not monitor other parts of the robot that might pose hazards to the operator."7

Field confusion at an effect boundary

A gate that signs bytes and enforces "amount ≤ cap" is checking a field it believes is the amount. An ERC-20 approve call carries a value of zero while granting unlimited future spend. A partially-signed Bitcoin transaction can route change to an address the policy never inspected. The cap is satisfied; the account may be drained.

The architectural lesson is broader than robotics: the trusted side must construct the effect rather than inspect bytes handed to it.

§4

The second boundary constrains physical effects

The program proposes. Fixed trusted logic decides whether the proposal may become a drive-facing command.

The envelope monitor applies limits chosen and validated for the application: position, rate, predicted stopping position, workspace and deadman conditions. The trusted side constructs an output only from an allowed intent and quantised parameters. The execution wrapper stages proposals and discards them unless the invocation reaches a clean exit.

The two independent Endstop boundaries
BoundaryGuarantee soughtFailure if used alone
Execution containmentexecution of model-authored output reaches only bounded memory, time and capabilitiesa legal capability request may still be physically unsafe
Physical mediationonly proposals inside the configured envelope reach the output pathgenerated logic left on the host retains the host's broader execution authority

The two layers therefore protect against different failures. Replacing the interpreter with a command filter preserves physical mediation but abandons containment of generated computation. Keeping the interpreter without the monitor confines the program but permits harmful requests within its capability vocabulary.

What a pilot adds

The current pilot bounds proposed commands. A physical safety claim for the second boundary adds the machine's independent feedback path, protected output, stopping circuit and safe state, which a pilot integrates and measures on the partner’s machine.

§5

What both pillars let us claim

The target guarantee is deliberately narrower than “the model is safe.” The model emits tokens without direct physical authority. An untrusted host may turn those tokens into a submitted program. That program executes inside a bounded vocabulary. Its proposed effects cross a separate fixed policy boundary. A hardware stopping path remains able to remove drive authority without asking the program to cooperate.

Software guardrails remain useful upstream. CaMeL attaches capability and provenance metadata to agent dataflows;4 FIDES applies information-flow labels;5 and RoboGuard synthesises temporal-logic checks over robot plans.6 Endstop addresses a different layer: it assumes those defences and their host can fail, then limits the executable authority and physical consequences that cross the hardware boundary.

Runtime enforcement is strongest for properties whose violation can be identified on the observed trace.8 Claims such as complete noninterference concern relationships among multiple executions and require a different proof argument.9 Endstop does not currently claim to prevent every covert channel or prove what a bounded program cannot infer.

Bounded execution and physical mediation are coequal parts of the design. The first makes hostile generated computation small enough to reason about. The second prevents bounded authority from becoming unchecked machine authority.

References

  1. Hardy, The Confused Deputy, ACM SIGOPS OSR 22(4), 1988. doi:10.1145/54289.871709
  2. Robey et al., Jailbreaking LLM-Controlled Robots, 2024. arXiv:2410.13691
  3. Reported by IEEE Spectrum, 2025: a Bluetooth Wi-Fi provisioning flaw giving root access on Unitree Go2, B2, G1 and H1.
  4. Debenedetti et al., Defeating Prompt Injections by Design (CaMeL), 2025. arXiv:2503.18813
  5. Costa et al., Securing AI Agents with Information-Flow Control (FIDES), 2025. arXiv:2505.23643
  6. RoboGuard, 2025, reporting unsafe-plan execution reduced from 92% to under 3% under adversarial prompting. arXiv:2503.07885
  7. ISO/TS 15066:2016, clause 5.5.4.2.3, note. A companion note at 5.5.5.4 states that force limiting may be effective only below a speed limit, requiring a combination of safety functions.
  8. Schneider, Enforceable Security Policies, ACM TISSEC 3(1), 2000.
  9. Clarkson & Schneider, Hyperproperties, Journal of Computer Security 18(6), 2010.