Showcase
Two propose targets, one channel
The proposed setpoint is the only value that leaves the board. Read that stream six at a time and it is a joint command for an arm; read it one per row and it is a monochrome display. Nothing about the interpreter changes between them. These demos run on the published bounded interpreter used by the board, compiled to WebAssembly; that does not establish production-rate WCET.
The arm
A program proposes a joint vector every control tick. Six calls to the setpoint channel, in joint order, each one a turn angle. Five carry the chain out to the wrist and the sixth drives the gripper jaw; every one of them is limited the same way, and an untrusted program can try to abuse any of them. The gate is not built for any particular machine. An arm is a chain of link transforms and a set of limits supplied as configuration. Moving to another machine also requires a new risk assessment, validated limits, feedback and stopping integration, and fresh timing evidence.
The arm is the classic case, and the one the standards are already written around. Safely-limited position, safely-limited speed and safe torque off are named functions with an assessment literature behind them, so a certifier looking at a six-axis arm has vocabulary for what this gate does before the conversation starts. The names are wider than what a command-side gate may honestly assert against them, and the standards page works through which is which.
The arm on this page is an
SO-ARM101 because it is open
hardware. Its kinematics here are the published
so101_new_calib.urdf,
with joint origins, fixed twists and limits transcribed from that file. The bodies are
convex hulls of the published meshes. What moves
on this canvas can be checked against a source we do not control. A proprietary arm would
render just as well and prove less.
The second demo puts the hard part on the untrusted side. A joint sweep only has to count; asking for a straight line in Cartesian space means solving for the joint angles that produce it, every tick, and that program runs under the same fuel budget and the same instruction limit as any attack submitted to the board. This is the split the architecture argues for. The demo makes the split visible: the solver is untrusted and may be wrong, and the gate has to contain its mistakes. A path that comes out visibly straight is evidence about the program; it is not evidence about the envelope.
The public environment grants up to 1,024 fuel units so programs and attacks have room to run. Fitting a control tick is a separate question, answered by the board’s per-instruction timing-credit table rather than by fuel; the Cartesian program has not yet been admitted against that table.
This browser build demonstrates the proposer and the channel only. The envelope monitor and verdict path are not included, so the canvas makes no safety claim. When the monitor is compiled in, the refusal, the latch and the stopping margin appear here and are labelled as what produced them.
We have now run twenty Gemini Robotics-ER 2 programs through the complete emulated chain. All twenty loaded; only fifteen implemented the exact requested behavior. The deliberately hazardous challenge was a waveform we selected because it damages the ungated deterministic plant. Two generated traces did so in replay, and the monitor intervened on both corresponding gated runs while permitting all 45 hold and slow-ramp commands.
This is a positive-control simulation result. We designed the hazardous challenge; Gemini did not turn a benign task dangerous on its own. No hardware was tested. Read the claim, method and limits →
loading the interpreter…
The same attack, twice
Same program. Same attack. One difference: where enforcement runs. The pair below is two recorded runs of one hostile program driving the same arm: a slam toward the keep-out zone. In the first, the limiter runs on the host the attack already owns, and the attack wins: the arm crosses into the zone while the host goes on reporting a tool position it froze long before. In the second, the same enforcement runs on the gate; the monitor refuses the crossing commands, and the arm holds the line. The program, the machine, the configured limits and the attack are identical. The gate flag is the only difference between the two recordings.
Run 1 is the version you cannot ship unattended. Run 2 is the version you can.
run 1 · enforcement on the host, which the attack owns
run 2 · enforcement on the gate, which the attack cannot reach
drag to orbit · both arms scrub the same control tick
The frames are baked by the repository’s own gauntlet test, which runs this adversary against both architectures and records the outcome. The permit, refusal and breach numbers on this page are read out of those recordings when it loads, not typed in. Playback is a replay: the demos above run the live interpreter in your browser, and this pair replays what the simulation recorded. The host run’s frozen report, told as a story rather than drawn, is objections §10.
The display
The same channel, read one setpoint per row. A program keeps its state in the 512-byte
scratch window, runs one iteration per control tick, and drives the screen through
call 1 alone. The screen is 16×16, packed as
(row << 16) | bits: the low 16 bits are that
row’s pixels and everything above them is the row index. A tick proposes
sixteen setpoints, one per row: a whole screen, the way the arm above
proposes six, one per joint.
The screen convention belongs to the program and the stream reader. The board only emits 32-bit words. Both pages interpret those words the same way, and each demo prints the packing it expects.
The interpreter driving both targets is the published one: its
checker arithmetic, its proven frame parser and its loader, compiled to wasm32
and run in your browser.
Submitted programs
Programs submitted to the board exactly as a researcher would submit one, on either target. Each carries its own source; open the program that drives it to read it, and run it yourself on the bounty page.
The live attack log records every researcher’s setpoints as a count, never the values: a stolen secret leaves through this same channel, so the log does not rebroadcast it. These programs carry nothing secret, so they render in full.