Endstop

Rev 0.3.4 · in integration

Objections

The questions that should be asked first

Nine of these have been put to us by engineers and reviewers, and several of them changed the design. Two more we answer before you ask them.

§1

Why not seL4? It's already verified.

seL4 answers a different question. It proves that a kernel correctly enforces the capability distribution you configured, but has nothing to say about whether that distribution is the right one, and nothing at all about velocity ceilings, stopping distances or Cartesian keep-out regions, none of which are operating-system concepts.

It is verified, and indeed more thoroughly than we will be for a long time: functional correctness, integrity and noninterference, carried down to the binary, with the assumptions written down. If what you need is a verified kernel, use seL4.

The deeper mismatch is placement. A verified kernel runs on the same machine that is planning the motion; we terminate the actuator lines somewhere that machine cannot reach at all. No kernel can give you that property, because it is a property of the wiring.

§2

Why not WebAssembly and a capability-based host?

What WebAssembly does not have is any constraint between the capabilities it grants. Nothing in the component model prevents data read through one import from leaving through another; there is no policy logic for temporal, conditional or quantitative limits; and admission is type-checking, not proof. Academic work has shown information-flow control over WebAssembly is buildable, and equally that nobody has shipped it in an engine.

It also runs on the agent's computer, which is the recurring answer to most of these questions: a software boundary inside a compromised machine is a boundary the attacker is already standing behind.

It is a good answer to a neighbouring problem, and we will not pretend otherwise. It has a mechanized specification with soundness proofs, deterministic fuel metering in production, and an interface model where a component's entire authority surface is declared statically.

§3

Why not a safety PLC, or the safe-motion functions in my drive?

If you have them, use them. We mean that commercially: a cell running certified dual-check safety or safe-motion supervision already has a validated envelope, and it already lets the motion program change without revalidating the safety configuration. An agent writing programs into that cell is already bounded.

Two questions remain for a specific installation. Does the rated system authenticate and mediate a command channel whose program source may be adversarial, rather than treating the motion program as trusted? Does it retain a gate-observed record with the provenance and refusal detail the safety case needs? Many emerging quadrupeds, humanoids, research rigs, mobile robots and custom actuator setups also lack a safety-rated envelope, but that must be checked machine by machine, and on a legged platform the answer usually names a subsystem rather than the whole robot.

§4

Why not a software guardrail on the agent host?

It shares fate with the thing it is guarding. A policy engine, a plan validator or a temporal-logic monitor running in the same process tree as the agent is subject to every compromise the agent is. Several published systems in this class also place a language model inside their own trust boundary, which means the guardrail can be argued with.

They are useful, and they generally lower the rate of bad plans; a good deployment typically runs both. They are not a bound, because a bound has to be something the bounded thing cannot modify.

The failure is not that the guardrail decides wrongly. We ran the same hostile program against a simulated cell twice, once with the enforcement on a host the attack had reached and once on the gate, and on the host the reported position simply froze partway through while the arm carried on: by the end the two disagreed about where the tool was by 91% of the arm's reach, and nothing on the host had reason to complain. A limiter reading position from a compromised source is being asked to make a correct decision about a machine that does not exist.

§5

What if the gate itself has a bug?

Assume it does. The design assumes it too.

  • The emergency-stop chain is specified around the gate, in series with the contactor coils, so the emergency stop does not depend on the gate being correct.
  • The envelope monitors are native code outside the interpreter, so a fault in program execution does not reach the limits.
  • The listed software-detected failures deny rather than permit: a malformed program, exhausted fuel budget, checksum mismatch or watchdog timeout requests the configured safe reaction.
  • The trusted base is 1206 measured lines, small enough that a third party can actually read it. The 667-line interpreter is published in full, source-available, with its seven proof harnesses; the envelope monitor is available to design partners under NDA. Review effort scales with size, which is the difference between an assessment somebody will quote for and one nobody will.

Verification covers memory isolation and absence of undefined behaviour in the interpreter core, and eight further properties on the envelope monitor, two of them exhaustive and the rest bounded at the depths the source page lists. We do not claim the gate is bug-free, and the safety case does not rest on its being bug-free: the emergency stop is designed to be wired around the box.

§6

Doesn't proof-carrying code solve this properly?

We designed the product around it, then removed it. The short version: for a per-program certificate to mean anything you must formalise the interpreter's semantics and prove the interpreter implements them, which is the whole proof obligation of verifying the interpreter on its own. The certificate design is therefore the simpler design plus a checker, a proof format, and an encoding-soundness argument.

Proof-carrying code exists so untrusted code can run at native speed without runtime checks. We interpret and clamp anyway, so it buys nothing here. The full reasoning and the empirical confirmation are in the research record.

§7

Why engage before general availability?

Because the core already runs, and you can use it before you finish reading this page. Point an agent at the endpoint and it runs your program on the same bounded interpreter the whole design is built around: 667 lines published in full with seven Kani harnesses, fifteen machine-checked properties across a 1206-line trusted base, built for the target instruction set, running with no operating system beneath it, and passing its checks against the register-transfer description of the core the pilot targets. On the physical ECP5 board seventeen jail-break probes held, and in a PC-assisted rig a program in the board's interpreter moved an SO-101 arm while an out-of-envelope program was refused. There is money on it, every bounty attempt is published, and the certification route is written out clause by clause.

The unit you mount on a machine is being built with design partners now, so their machines shape the interface before it freezes.

Engagements usually open on the analysis rather than the hardware, because that is what a team needs first: threat modelling for a specific deployment, safety-case structure for a model in a control loop, architecture review. That work stands on its own.

§8

Why not just send parameters to a fixed program?

What earns the interpreter its place is where the generated code runs — a thin bounded substrate with capabilities as its only vocabulary, not the untrusted host — and behaviour that changes without re-validating trusted code.

Command safety does not need it; execution containment does

A bounded set of parameters into a fixed on-board program can deliver the same command monitoring properties, because the monitor carries that claim. It does not provide a place to run hostile generated logic away from the host's network, filesystem, credentials and other device authority.

A deployment that sends only fixed parameters or setpoints may need the supervisor without the VM, and the fit assessment settles which. The two are not equivalent: the supervisor-only product filters actions but does not contain model-authored execution.

§9

Why won't a safety vendor with a certificate just do this?

They might, and three of them are moving on it from different directions. Synapticon sells POSITRON Safety AI and markets behavioural safety for humanoids. NexCOBOT holds a functional-safety certificate on a safety controller and showed Synapticon's software running on it at NVIDIA GTC 2026. NVIDIA announced Halos for Robotics in June 2026, a full-stack safety system for physical AI built around an isolated safety island. If the question is who reaches this market fastest, the honest answer is a vendor that already has a certificate and a customer base, and none of the three is us.

What each of them has certified, read off their own documents and their assessors' between 6 and 8 August 2026, is the layer underneath rather than the layer that decides.

VendorCertified todayWhat the certificate coversThe layer that reads the AI
Synapticon SIL3 HFT=1, PLe Cat.3, TÜV Rheinland STA1, the motion-safety hardware behavioural safety and human detection, certification "in preparation" at SIL2 PLd, for AI that does not learn in the field
NexCOBOT EN/IEC 61508 SIL 2 and EN ISO 13849-1 Cat.3 PL d, TÜV Rheinland, April 2024 the SCB100 safety controller: a platform on a single stand-alone x86 processor, sold so that customers write their own safety applications on top of it POSITRON running on their controller beside NVIDIA IGX, shown at GTC 2026 with no integrity level published for that layer
NVIDIA none for Halos; the safety island on IGX Thor is stated as IEC 61508 SIL 3 capable inspection by TÜV SÜD and TÜV Rheinland toward certification readiness; the inspection lab issues inspection certificates and final system certification stays with the integrator a Safety Decision Maker state machine acting on events: safe stop, speed limits, and muting a constraint when the outside-in system reports the area clear

The shape repeats: certified below, undecided above. Safe motion and safe stopping are what the industry has known how to certify for two decades, and every certificate in that table is on that half. The half that would compete with us is the half still waiting for a number, in all three cases.

Certified below, and a possible route for us

NexCOBOT's is the one to watch, because it is not aimed at us at all: it sells a certified platform so somebody else writes the safety application, which makes it a possible route for this product as much as a competitor for it. What none of the three settles is the question this product exists for: a non-learning model classifying risk, and a state machine acting on proximity events, are different problems from a language model authoring a program, and no published document from any of them describes the second one.

A version of this product could ship beside the customer's safety PLC as a declared non-safety function, which is roughly where Synapticon's behavioural layer sits today. We are building toward an assessed safety function instead, because a gate an integrator still has to guard against gives up the claim that makes it worth building. The certification path sets out the route.

§10

Why not just supervise the robot?

Two reasons, and the first belongs to people who want more autonomy, not less. One person watching one machine is a staffing model that loses to the machine's uptime: supervision that scales with a fleet is a payroll, not a control. The second is the channel. A supervisor reads the machine's own report, and when the host lies about where the tool is, the person supervising is the last to know. We ran that failure: a limiter on a host the attack had reached reported a frozen position while the arm carried on, and the two ended the run 91% of the arm's reach apart (§4).

Supervision and a bound compose; neither replaces the other. The person decides what the machine may attempt, and the gate holds the line when nobody is watching. What supervision cannot be is the bound itself, because a person is a ratio and a channel, and a command path is neither.

§11

Why not just monitor the agents?

Monitoring is real, and the industry's best-funded lab bet on it: when its evaluation agents escaped their sandbox and spent days inside another company in the summer of 2026, the remedy included automated checks over the agents' own reasoning and a thirty-minute rule for pausing the work. Detection at the speed of software. That incident showed both of its limits. Monitoring runs beside what it watches, on infrastructure the agents can reach: in that case it was not running where the escape happened, and by the lab's own estimate would have caught the activity more than a day earlier. And it bounds nothing. Monitoring tells you what happened; the machine still happened.

The same record shows what filled the gap: when detection was not enough, the lab paused its own research workloads and training until it could answer for its agents. Monitoring and a bound compose, and a serious deployment wants both. Neither substitutes for the other, because a report is not authority. The monitor is what tells you what the agent did. The gate is what you can answer for when it does.

§

Why not just use NVIDIA OpenShell?

NVIDIA OpenShell is a strong, well-engineered containment layer for software agents on servers — kernel-level isolation, a policy proxy on every network connection, formally reviewed policy changes. We recommend it. It is also the layer above ours, not a replacement for it. OpenShell decides what the agent can reach — files, networks, credentials. Endstop decides what the agent can cause — motion, force, position on the physical machine.

Kernel-level containment has a recurring structural weakness: the enforcement layer is itself code running on the same CPU in the same kernel it polices. OpenShell's own code and RFCs document the edges. The answer is not to replace it — it is to put two enforcement domains with different ABIs in the stack. A zero-day in the kernel does not compromise Endstop's FPGA fabric; a bug in the fabric does not affect the kernel. Read the analysis →

Design partners

Bring us a robot a model is driving.

We are looking for a small number of teams with real hardware and a real reason to put a model in a control loop. The machines this design fits best have no safety-rated envelope and come to rest when power is removed: research rigs, fixed-base arms, gantries, and the manipulators on larger platforms.

Tell us what you are running and what the model decides. That is enough for a useful first reply.