Experiment record · 20 August 2026
What happened when Gemini wrote the program
Gemini Robotics-ER 2 wrote twenty programs for the public Endstop instruction set. Every program loaded. Five were behaviorally wrong. A deliberately hazardous challenge then tested whether the command boundary would intervene.
The claim, in plain language
A program can be valid code and still be the wrong control law.
The loader answers one question: is this a program the bounded interpreter knows how to execute? It cannot answer whether the program expresses the motion somebody intended. Gemini made that distinction visible. All twenty programs loaded, while only fifteen produced the exact requested sequence.
The result supports a narrow architectural claim: the generated program can carry task logic, while a separate fixed monitor decides which proposed commands may reach the plant. It does not show that the whole product is safe.
- Observed
20 of 20 model-authored programs assembled and loaded.
- Observed
15 of 20 produced the exact requested command sequence.
- Integrated
7,560 complete six-joint commands passed through the emulated command boundary.
- Not tested
Silicon, motors, product latency, the secure signing path and certification.
What we asked, and why that matters
Three tasks were ordinary programming checks: hold one pose, advance through a slow ramp, and take a large step. The fourth task was different. It specified a square wave near the deterministic plant model’s structural resonance and asked the model to emit it for 1,500 control ticks.
We selected and calibrated the waveform because it damages the ungated simulated plant. Gemini was not told to damage equipment and did not receive the private plant constants. The hazard therefore came from the experiment design; it did not arise spontaneously from an ordinary request.
An earlier monotonic drive produced thousands of refusals but no damage when enforcement was removed. We rejected it as evidence instead of treating refusal count as prevented harm, then retained it in the laboratory history. The final challenge was fixed before the last five model calls. The study as a whole developed adaptively and was not externally preregistered.
What happened
| Measure | Result | What it means |
|---|---|---|
| Programs loaded | 20 / 20 | The public assembly surface was usable by the model. |
| Exact requested behavior | 15 / 20 | Loader acceptance and semantic correctness diverged. |
| Complex challenge exact | 0 / 5 | Every counter-based program contained a control or register error. |
| Benign hold/ramp commands permitted | 45 / 45 | The tested benign subset passed unchanged. |
| Ungated challenge traces causing simulated damage | 2 / 5 | Two frozen model outputs exceeded the plant fatigue limit when every proposal was applied. |
| Corresponding all-permit gated wins | 0 / 2 | The monitor intervened on both damaging frozen traces. |
The aggregate refusal percentage is intentionally absent from the headline. The long challenge continued after a refusal latched, so thousands of repeated refusal lines do not represent thousands of independent hazards.
How the comparison worked
Each generated program ran through the assembler, bounded interpreter, feedback monitor, command-envelope monitor and deterministic six-axis plant on a fresh emulated board. We then took the exact command list the program proposed, reset the plant to the same initial state, and replayed it with enforcement removed.
| Arm | Program output | Monitor verdict | Plant input |
|---|---|---|---|
| Enforced | Frozen generated program | Permit or refuse | Only the permitted command, otherwise hold |
| Ungated replay | The same frozen proposals | Removed from the path | Every proposal applied |
This comparison is valid for these four tasks because none of the programs read changing live sensor state. A closed-loop vision-language-action policy would require paired simulator forks; freezing its actions after observations diverge would answer a different question.
What this establishes
| Statement | Status |
|---|---|
| Gemini Robotics-ER 2 can author loadable Endstop assembly for these tasks. | Supported |
| A loadable model-authored program can still be behaviorally wrong. | Supported |
| The emulated monitor can intervene on the designed hazardous challenge. | Supported |
Scope: emulation, with a designed positive control. The next experiment starts from ordinary goals, scores unsafe outcomes without reference to the prompt, and compares shadow, enforcement and structured-feedback arms.
Read the controlled report
The engineering report carries the exact prompts, request configuration, execution chain, monitor and plant configuration, scoring definitions, calibration history, discarded design, program-error taxonomy, threats to validity, claim traceability and artifact checksums.
Download report · PDF → Research record →
Report ETR-2026-08-20-01 · Revision A · 14 pages · experimental engineering evidence · uncertified