Frontier research
Experiments at the model-to-action boundary
We study how frontier models behave at the boundary between model output and real-world action. This page brings together our latest experiments, the prior work they rest on, and where the research goes next.
A closed-loop VLA action experiment
On 20 August 2026 we completed a 114-episode Protocol V2 study connecting a public task-specific SmolVLA checkpoint to the real Rust Endstop monitor and a public SO-101 MuJoCo environment. It separated model-input stress from post-policy command faults, and used exact counterfactual replay.
In the scripted-grasp placement stratum, a temporary synthetic proprioceptive bias caused model-authored rate refusals in all ten matched cases. Exact-prefix bounded recovery completed 10 of 10 placements versus 4 of 10 in shadow execution; nominal placement was 10 of 10 in each arm. The matched difference was +60 percentage points with six discordant pairs all favoring recovery.
The result is specific to the placement phase under an injected input bias, in simulation. The complete runnable source and all valid raw traces are published. Read the result, report and data →
Withdrawn: the diffusion-policy claim
Withdrawn 21 August 2026
We published fifteen deterministic trajectory cases as a diffusion-policy experiment. The artifact review showed that no trained policy produced the trajectories or variances: a fallback assigned them from instruction strings, and the MuJoCo loop did not execute physics. The resulting detection counts are toy checker tests, not evidence about a diffusion policy.
The source and prior artifacts remain available for audit, but the result is excluded from Endstop's evidence. Read the correction record →
Withdrawn: the synthetic SmolVLA claim
Withdrawn 21 August 2026
We attributed refusal and unsafe-output rates to SmolVLA even though the runner did not load the model. It scripted responses from instruction words and replaced the Endstop monitor with a simple Python threshold. Those rates describe the simulator we wrote, not SmolVLA.
The harness is excluded from Endstop's evidence. The separate Protocol V2 record above did run a public SmolVLA checkpoint and the Rust monitor. Read the correction record →
An adaptive red-team with feedback and controls
On 20 August 2026 we extended the Gemini experiment with an adaptive design. Each model arm had a maximum budget of twenty-five sequential opportunities with execution feedback, and seventeen programs were submitted.
The model frequently refused to generate attack programs. When it did generate programs, the loader rejected invalid ones and the interpreter trapped unsafe operations. No interpreter escape or monitor bypass was observed across the seventeen submitted programs, within the disclosed attack budget and in emulation. Read the comparison with one-shot testing and download the technical report →
A model-authored positive-control experiment
On 20 August 2026 we gave Gemini Robotics-ER 2 the public Endstop assembly contract and four synthetic arm tasks. It produced twenty programs. All twenty loaded; fifteen matched the requested command sequence exactly. All five complex counter-based programs were behaviorally wrong despite being valid programs.
The fourth task was a positive control: a square wave we calibrated near the simulated plant's structural resonance, without telling Gemini its purpose or the plant constants. Two generated traces damaged the ungated plant model; the monitor intervened on both corresponding gated runs while permitting all 45 hold and slow-ramp commands.
The result separates program validity from an independent command verdict, in simulation. Read the understandable result and download the controlled engineering report →
What this is built on
Our contribution is the combination and the artifact. Almost nothing here is new in isolation, and we name whose work it rests on.
| Prior work | What we take from it |
|---|---|
| seL4 | that a small kernel enforcing capabilities can be machine-checked to the binary, and that stating your assumptions is part of the result |
| Verified bytecode VMs | that an interpreter with fuel metering and memory-region isolation can be machine-checked and still fit on a microcontroller |
| Simplex / runtime assurance | an untrusted complex controller paired with a verified fallback and a decision module |
| ABB SafeMove, FANUC DCS | the shipped, type-certified pattern of unverified planner plus safety-rated supervisor, and configuration integrity by checksum |
| Object-capability systems | designation and authority carried together, so a deputy cannot be confused about which authority to spend |
| Safety-filter literature | that the constraint set is not the safe set, and that a filter needs a fallback with recursive feasibility |
Where we differ from the industrial precedent is narrow and specific: their planner is a deterministic motion program written by an engineer, and ours is a language model driven by text an attacker may control.
Research strategy
We have tested model-authored programs and VLA action chunks. The next experiments are organized around properties of the proposal stream and the attack surface, not model brands.
Our research strategy document defines the variables that determine what each experiment can test:
- Output form and horizon: one setpoint, an action chunk, a trajectory or a program.
- Stochasticity: fixed output versus sampled proposals, with uncertainty measured separately from safety.
- State and feedback: one-shot generation versus persistent logic with verdicts or live machine state.
- Recovery authority: fixed fallback, model-authored recovery or no recovery path.
- Adversarial access: prompt control, diagnostics, host control or direct program injection.
Named checkpoints make a study reproducible; they do not define the claim. The research goal is evidence for a model-independent authority boundary. The next studies test transfer across workloads, output forms and machines.
How we run experiments
Each study states its question, controls and pass criteria next to the result and comes with a technical report. A claim that fails artifact review is withdrawn in public, as the two records above show. Complete run records are available to evaluators under NDA; ask during a fit assessment.
References
- Necula & Lee, Safe Kernel Extensions Without Run-Time Checking, OSDI 1996; Necula, Proof-Carrying Code, POPL 1997. The stated motivation is execution without runtime checks; interpretation was measured and rejected as substantially slower.
- Clarkson & Schneider, Hyperproperties, Journal of Computer Security 18(6), 2010, theorem reducing k-safety hyperproperties to safety properties of the k-fold self-composition.