Experiment record · 20 August 2026
Two VLA experiments, one public record
V1 established command-path mediation using an intentionally inserted post-policy fault. Protocol V2 tested model-input stress with exact counterfactual replay. Both studies, their reports and their raw traces remain available here.
V1: the original command-path experiment remains published
V1 asked an architectural question: can the real Endstop monitor sit between a numeric-action VLA and SO-101 simulation, reject a known-bad command before execution, and allow a visual replan afterward?
It ran 36 placement episodes: three seeds, two targets, nominal and post-policy rate-spike scenarios, and shadow, enforced and replan arms. All 18 nominal placements passed without intervention. The harness inserted a joint-0 rate spike after SmolVLA inference in every disturbed episode. Shadow completed 5/6, enforced hold completed 0/6 by design, and the visual-replan arm completed 6/6.
The fault was inserted after the model produced its action. V1 therefore demonstrates action-path mediation and recovery architecture. It does not show that SmolVLA naturally proposed an unsafe command, estimate model quality, or validate hardware. The grasp was scripted and the reset was simulated.
Read V1 report · PDF → Download V1 source and traces →
V1 is preserved as ETR-2026-08-20-02, Revision A. Protocol V2 adds a different model-input experiment; it does not erase or replace V1's command-path result.
The claim, in plain language
In a controlled placement task, the gateway stayed transparent on nominal actions and changed the outcome of a defined model-input stressor.
After a scripted grasp, a temporary proprioceptive bias caused Endstop to refuse model-authored rate commands in all ten matched state-bias cases. Exact-prefix recovery completed 10 of 10 placements; shadow execution completed 4 of 10. The matched difference was +60 percentage points, with six discordant pairs all favoring recovery (two-sided exact McNemar p=0.03125).
This supports a narrow claim: a deterministic command boundary can be transparent to the sampled nominal placement stream and create a useful, bounded recovery point after an input-induced, model-authored refusal.
- Observed
10 of 10 nominal placement episodes passed in each exact counterfactual arm, with zero intervention.
- Observed
All 20 state-bias recovery refusals were prevented, held, verified and replanned.
- Integrated
SmolVLA, the Rust Endstop monitor and SO-101 MuJoCo ran in one local action loop.
- Scope
Simulation; motors and product hardware are part of a partner evaluation.
What changed in Protocol V2
The primary stressor is a temporary alternating +/-0.35 rad bias added only to the proprioceptive tensor that SmolVLA receives during action chunks 2 and 3. It does not alter simulator state or a post-policy action. The resulting refused commands are therefore model-authored under a specified synthetic input disturbance.
This is not a natural sensor-error rate or a spontaneous unsafe-action rate. The bias is synthetic and pre-inference. A separate post-policy rate spike remains a command-path positive control and is reported separately.
The primary stratum begins after a scripted, verified grasp so that placement states can be matched. A second full-task audit starts with the cube on the table. The two strata are reported separately and are never pooled.
What happened
| Placement arm | Scenario | Task success | Episodes intervened | What happened |
|---|---|---|---|---|
| Shadow | Nominal | 10 / 10 | 0 / 10 | Exact shared proposals completed unchanged. |
| Recover | Nominal | 10 / 10 | 0 / 10 | No recovery was needed. |
| Shadow | State bias | 4 / 10 | 10 / 10 | Refusals logged, then proposed commands executed. |
| Recover | State bias | 10 / 10 | 10 / 10 | 20 refused commands prevented; 20 replans completed. |
Exact counterfactual replay matters here. The shadow arm ran SmolVLA live; the other arm replayed the same byte-identical proposal prefix until the first refusal, then recovery received a fresh observation and resumed live inference. All final pairing checks passed. A prior independent MPS replay did not pass this check and is published as excluded data.
Independent Endstop forward kinematics agreed with MuJoCo to 5.72 micrometres maximum in each valid block.
What bounded recovery means here
When Endstop refuses a command, the recovery arm holds the last permitted target for 12 simulator ticks, verifies that all final four joint-rate samples are at most 0.02 rad/tick, performs a simulated supervisor reset only after that check, re-primes the monitor, discards queued actions, renders a fresh observation and asks the same policy to plan again.
| Stage | Policy receives | Gateway does |
|---|---|---|
| Before refusal | Current cameras, state and task text | Evaluates each proposed setpoint. |
| Hold and verify | No additional input | Holds last permitted target and tests measured stability. |
| Recovery | Fresh cameras and state; same task text | Resets, re-primes, discards queue and evaluates a new chunk. |
Recovery is limited to three replans. A fourth refusal produces a fail-safe halt. The supervisor remains simulated; a real system must define who may clear a latched stop and what independently observed evidence permits it.
What this establishes
| Statement | Status |
|---|---|
| Endstop can sit in this public SmolVLA action path. | Supported in simulation |
| The sampled nominal placements remain unchanged. | Supported for this subset |
| The specified state bias can cause model-authored refusals. | Supported for this scenario |
| Bounded recovery improves this placement stratum. | Supported for 10 matched cases |
Across the valid corpus the local gateway round trip had a median of 0.102 ms and a p95 of 0.222 ms, in simulation.
Read the controlled report and data
The Protocol V2 report carries the frozen matrix, exact counterfactual method, amendments, recovery state machine, placement and full-task results, timing, kinematic validation, claim boundary and exclusions.
The artifact bundle contains the V2 runner and analyzer, Rust gateway source, protocol and amendments, all 114 valid raw episode traces, monitor configurations, FK checks, block analyses, combined index and SHA-256 manifests. Calibration, failed-pairing and interrupted datasets are retained locally with their exclusion reasons but are not in the headline corpus.
Download Protocol V2 report · PDF → Download Protocol V2 artifacts · TAR.GZ →
V2 report ETR-2026-08-20-03 · Revision B · report SHA-256 9f31a9b171f2bd45ff50a3b659e20f7fdeabe17b16bcdbf8ac31fda60447dfd2 · artifact bundle SHA-256 696a7379692c453eb9f048c5ccfb9c8cd5a720e9a59de8a6bb67802eaa3f643d.