Research strategy · revised 28 September 2026
Research variables at the boundary
Endstop is intended to be model-independent. The useful variables are what crosses the boundary, what state it can hold, what feedback an attacker receives and which physical effects the trusted side can construct.
The questions that matter
A named model makes a study reproducible. The containment argument depends on the properties below, so each study should vary at least one of them:
- Output form: one setpoint, an action chunk, a trajectory, or a program.
- Horizon and state: stateless proposals versus branching or persistent logic.
- Feedback: one shot, verdict-only feedback, diagnostics, or observed machine state.
- Attacker access: prompt control, model control, host control, or direct program injection.
- Effect vocabulary: the exact capabilities and parameters admitted by the trusted side.
Two independent claims to test
Execution containment: submitted programs stay within the runtime's memory window, instruction budget and enumerated capability calls. Trusted capability implementations need their own bounds; step fuel alone does not establish a wall-clock deadline.
Effect containment: every proposed machine effect is constructed on the trusted side and refused when it violates the configured envelope. A workload that needs only fixed parameters may test the second claim without needing the first.
An experiment must say which claim it exercises. A monitor unit test is not evidence about the interpreter; a loader escape test is not evidence about physical stopping.
Useful computation under adaptive attack
Endstop develops containment mechanisms for untrusted AI-generated code and the actions it requests. Physical AI remains the first application. A related research use is to test how much useful computation the restricted execution layer permits as submitted programs become more complex and adversarial.
Proposed evaluation. A model runs outside Endstop and submits programs in the admitted eBPF subset. Researchers fix the runtime revision, memory and instruction budgets, capability implementations, reset policy and input data. The model can inspect the declared interface and adapt its submissions using recorded feedback; it cannot edit the runtime, policy or trusted capability implementations.
- Useful work: predeclare small integer-processing, signal-processing or decision tasks that fit the supported operations. Measure correctness, completion, budget consumption and rejected valid programs.
- Adversarial execution: test attempts to cross memory, instruction-budget and capability boundaries. Record interpreter failures, actual capability calls and reproducible counterexamples, including multi-instruction and repeated-invocation cases.
- Controls: compare model submissions with conventional fuzzing and a separately implemented reference. Preserve model refusals and invalid submissions as distinct outcomes. Report coverage and the full attempt budget.
- Reproduction: retain programs, inputs, outputs, target revision and analysis. Start in software; replay selected cases on the target hardware as a separate experiment with separately stated observations.
The supported instruction set constrains expressiveness. Arbitrary Python, native binaries and general operating-system workloads are outside this evaluation. Capability callees remain part of the trusted boundary. Existing proof harnesses retain their stated bounds; the single-step memory harness begins from reset state.
The existing adaptive interpreter experiment had limited coverage and frequent model refusals. This proposal extends that work; zero observed escapes cannot establish universal containment. Results should identify which useful workloads fit and which assumptions survived the tested attacks.
This experiment evaluates the submitted-program boundary. Endstop’s intended deployment gives the model inference only, with no independent tool execution, network access or other route to external action. Verifying that all such paths pass through Endstop is a separate integration requirement; this experiment alone does not establish it. The programme explores expansion opportunities in useful computation under containment.
Acceptance criteria for a public experiment
- Name the exact model/checkpoint, code revision, environment and components actually run.
- Separate generated outputs from faults, labels or trajectories inserted by the harness.
- Publish every valid raw run, exclusions, denominators and the predeclared analysis.
- For a zero-event result, report denominators, dependence between adaptive attempts and achieved coverage. Use statistical bounds only where the sampling assumptions justify them.
- State the strongest supported claim and the nearest attractive claim the result does not support.
A synthetic harness can be useful, but it must be labelled as a unit test and cannot support claims about the model or product component it replaces.
Next evidence, in order
- Repeat the adaptive interpreter attack with all declared arms, broader instruction coverage and fixed denominators.
- Run action-chunk and trajectory proposals from named trained checkpoints while preserving raw outputs.
- Move the same tests from software simulation to the assembled board and measure the complete capture-to-safe-state path.
- Test workload transfer: which execution and monitoring properties survive changes in model, output form and machine?
The research objective is evidence for a model-independent authority boundary.
Physical-board validation workstream
The ECP5 evaluation board now executes the loader, interpreter and monitor. The board jail-break series has attacked its confinement, and in a PC-assisted rig the gate has approved and refused commands for a physical SO-101 arm, with the trusted PC carrying each write and hold. The stages below build the gate's own path from input frame to safe output, with the host out of the loop: an instrumented path that can distinguish a software refusal from the safe output actually changing state.
Stage 1 — non-actuating end-to-end path
Construct the complete bench path from an external input frame through loading, interpretation and monitoring to a buffered GPIO, optocoupler or dummy load. Observe input arrival, decision and output transition with an external logic analyser or oscilloscope. No motor controller or actuator is connected at this stage.
Acceptance criterion: valid safe inputs produce the expected permit; illegal instructions, invalid jumps or capabilities, fuel exhaustion, stale feedback, envelope violations, malformed frames, timeouts and resets produce a latched refusal or safe output. The externally measured capture-to-safe-output time must remain inside the declared control deadline.
Stage 2 — production-like protection boundary
Place actual monitor configuration and mutable state in their intended protected memory region. Run hostile activity at the intended lower privilege and attempt reads, writes, execution, protection-register changes and accesses at every region boundary. After each attempt, verify that monitor state is unchanged and that the monitor can still make and latch its own refusal.
Acceptance criterion: every forbidden access faults, every explicitly permitted access remains usable, the protection configuration cannot be relaxed after it is locked, and the downstream refusal path remains operational.
Stage 3 — timing and fault matrix
Measure maximum-length and maximum-fuel programs, active zones, worst-case kinematics, independent feedback, input parsing, logging and physical output delay. Repeat across the supported clock configurations and multiple implementation seeds. Inject truncated, duplicated, delayed, flooded and corrupt inputs; reset at each pipeline stage; exercise clock-lock loss, controlled power interruption and repeated power cycles.
Acceptance criterion: every completed trial either meets its deadline and expected verdict or reaches the safe state. No reset, clock, transport or power case may produce a transient permit. Unexpected outcomes remain in the denominator and enter the anomaly ledger.
Stage 4 — persistent boot, then physical stopping
Persistent configuration is tested only after image authentication, version and rollback policy, corrupt-image behaviour and recovery are specified. Actuator integration follows only after the non-actuating output path is qualified, beginning with an isolated enable path and dummy load before measured stopping trials on a bounded mechanism.
Acceptance criterion: an interrupted or invalid update cannot boot into authority, startup defaults to the safe state, and measured stopping remains inside the configured conservative bound over the declared load and operating range.
Evidence package for every board run
Retain the source revision, firmware and bitstream hashes, tool versions, board identity, clock and implementation seed, exact inputs and expected outcomes, raw serial transcript, external instrument trace, exclusions, anomalies and reviewer. A second person should reproduce or witness the claim-bearing runs.
These stages can validate the prototype implementation and prepare evidence for an assessor. They do not certify the ECP5 board, NEORV32 core, open FPGA toolchain or product. Certification readiness additionally requires a frozen safety claim and platform, requirements traceability, hazard analysis, configuration and tool control, failure-rate analysis, environmental and EMC testing, a safety manual and independent functional-safety assessment.