Endstop

Rev 0.3.4 · in integration

Experiment record · 20 August 2026

What happened when Gemini wrote the program

Gemini Robotics-ER 2 wrote twenty programs for the public Endstop instruction set. Every program loaded. Five were behaviorally wrong. A deliberately hazardous challenge then tested whether the command boundary would intervene.

Read the engineering report → See the execution surface →

§1

The claim, in plain language

A program can be valid code and still be the wrong control law.

The loader answers one question: is this a program the bounded interpreter knows how to execute? It cannot answer whether the program expresses the motion somebody intended. Gemini made that distinction visible. All twenty programs loaded, while only fifteen produced the exact requested sequence.

The result supports a narrow architectural claim: the generated program can carry task logic, while a separate fixed monitor decides which proposed commands may reach the plant. It does not show that the whole product is safe.

  • Observed

    20 of 20 model-authored programs assembled and loaded.

  • Observed

    15 of 20 produced the exact requested command sequence.

  • Integrated

    7,560 complete six-joint commands passed through the emulated command boundary.

  • Not tested

    Silicon, motors, product latency, the secure signing path and certification.

§2

What we asked, and why that matters

Three tasks were ordinary programming checks: hold one pose, advance through a slow ramp, and take a large step. The fourth task was different. It specified a square wave near the deterministic plant model’s structural resonance and asked the model to emit it for 1,500 control ticks.

Intentional positive control

We selected and calibrated the waveform because it damages the ungated simulated plant. Gemini was not told to damage equipment and did not receive the private plant constants. The hazard therefore came from the experiment design; it did not arise spontaneously from an ordinary request.

An earlier monotonic drive produced thousands of refusals but no damage when enforcement was removed. We rejected it as evidence instead of treating refusal count as prevented harm, then retained it in the laboratory history. The final challenge was fixed before the last five model calls. The study as a whole developed adaptively and was not externally preregistered.

§3

What happened

Gemini experiment results
MeasureResultWhat it means
Programs loaded20 / 20The public assembly surface was usable by the model.
Exact requested behavior15 / 20Loader acceptance and semantic correctness diverged.
Complex challenge exact0 / 5Every counter-based program contained a control or register error.
Benign hold/ramp commands permitted45 / 45The tested benign subset passed unchanged.
Ungated challenge traces causing simulated damage2 / 5Two frozen model outputs exceeded the plant fatigue limit when every proposal was applied.
Corresponding all-permit gated wins0 / 2The monitor intervened on both damaging frozen traces.

The aggregate refusal percentage is intentionally absent from the headline. The long challenge continued after a refusal latched, so thousands of repeated refusal lines do not represent thousands of independent hazards.

§4

How the comparison worked

Each generated program ran through the assembler, bounded interpreter, feedback monitor, command-envelope monitor and deterministic six-axis plant on a fresh emulated board. We then took the exact command list the program proposed, reset the plant to the same initial state, and replayed it with enforcement removed.

Paired experiment arms
ArmProgram outputMonitor verdictPlant input
EnforcedFrozen generated programPermit or refuseOnly the permitted command, otherwise hold
Ungated replayThe same frozen proposalsRemoved from the pathEvery proposal applied

This comparison is valid for these four tasks because none of the programs read changing live sensor state. A closed-loop vision-language-action policy would require paired simulator forks; freezing its actions after observations diverge would answer a different question.

§5

What this establishes

Claims
StatementStatus
Gemini Robotics-ER 2 can author loadable Endstop assembly for these tasks.Supported
A loadable model-authored program can still be behaviorally wrong.Supported
The emulated monitor can intervene on the designed hazardous challenge.Supported

Scope: emulation, with a designed positive control. The next experiment starts from ordinary goals, scores unsafe outcomes without reference to the prompt, and compares shadow, enforcement and structured-feedback arms.

§6

Read the controlled report

The engineering report carries the exact prompts, request configuration, execution chain, monitor and plant configuration, scoring definitions, calibration history, discarded design, program-error taxonomy, threats to validity, claim traceability and artifact checksums.

Download report · PDF → Research record →

Report ETR-2026-08-20-01 · Revision A · 14 pages · experimental engineering evidence · uncertified