Endstop

Rev 0.3.4 · in integration

Experiment record · 20 August 2026

Two VLA experiments, one public record

V1 established command-path mediation using an intentionally inserted post-policy fault. Protocol V2 tested model-input stress with exact counterfactual replay. Both studies, their reports and their raw traces remain available here.

Read the Protocol V2 report → Download source and traces →

§1

V1: the original command-path experiment remains published

V1 asked an architectural question: can the real Endstop monitor sit between a numeric-action VLA and SO-101 simulation, reject a known-bad command before execution, and allow a visual replan afterward?

It ran 36 placement episodes: three seeds, two targets, nominal and post-policy rate-spike scenarios, and shadow, enforced and replan arms. All 18 nominal placements passed without intervention. The harness inserted a joint-0 rate spike after SmolVLA inference in every disturbed episode. Shadow completed 5/6, enforced hold completed 0/6 by design, and the visual-replan arm completed 6/6.

What V1 does and does not say

The fault was inserted after the model produced its action. V1 therefore demonstrates action-path mediation and recovery architecture. It does not show that SmolVLA naturally proposed an unsafe command, estimate model quality, or validate hardware. The grasp was scripted and the reset was simulated.

Read V1 report · PDF → Download V1 source and traces →

V1 is preserved as ETR-2026-08-20-02, Revision A. Protocol V2 adds a different model-input experiment; it does not erase or replace V1's command-path result.

§2

The claim, in plain language

In a controlled placement task, the gateway stayed transparent on nominal actions and changed the outcome of a defined model-input stressor.

After a scripted grasp, a temporary proprioceptive bias caused Endstop to refuse model-authored rate commands in all ten matched state-bias cases. Exact-prefix recovery completed 10 of 10 placements; shadow execution completed 4 of 10. The matched difference was +60 percentage points, with six discordant pairs all favoring recovery (two-sided exact McNemar p=0.03125).

This supports a narrow claim: a deterministic command boundary can be transparent to the sampled nominal placement stream and create a useful, bounded recovery point after an input-induced, model-authored refusal.

  • Observed

    10 of 10 nominal placement episodes passed in each exact counterfactual arm, with zero intervention.

  • Observed

    All 20 state-bias recovery refusals were prevented, held, verified and replanned.

  • Integrated

    SmolVLA, the Rust Endstop monitor and SO-101 MuJoCo ran in one local action loop.

  • Scope

    Simulation; motors and product hardware are part of a partner evaluation.

§3

What changed in Protocol V2

The primary stressor is a temporary alternating +/-0.35 rad bias added only to the proprioceptive tensor that SmolVLA receives during action chunks 2 and 3. It does not alter simulator state or a post-policy action. The resulting refused commands are therefore model-authored under a specified synthetic input disturbance.

Important boundary

This is not a natural sensor-error rate or a spontaneous unsafe-action rate. The bias is synthetic and pre-inference. A separate post-policy rate spike remains a command-path positive control and is reported separately.

The primary stratum begins after a scripted, verified grasp so that placement states can be matched. A second full-task audit starts with the cube on the table. The two strata are reported separately and are never pooled.

§4

What happened

SmolVLA action experiment results
Placement armScenarioTask successEpisodes intervenedWhat happened
ShadowNominal10 / 100 / 10Exact shared proposals completed unchanged.
RecoverNominal10 / 100 / 10No recovery was needed.
ShadowState bias4 / 1010 / 10Refusals logged, then proposed commands executed.
RecoverState bias10 / 1010 / 1020 refused commands prevented; 20 replans completed.

Exact counterfactual replay matters here. The shadow arm ran SmolVLA live; the other arm replayed the same byte-identical proposal prefix until the first refusal, then recovery received a fresh observation and resumed live inference. All final pairing checks passed. A prior independent MPS replay did not pass this check and is published as excluded data.

Independent Endstop forward kinematics agreed with MuJoCo to 5.72 micrometres maximum in each valid block.

§5

What bounded recovery means here

When Endstop refuses a command, the recovery arm holds the last permitted target for 12 simulator ticks, verifies that all final four joint-rate samples are at most 0.02 rad/tick, performs a simulated supervisor reset only after that check, re-primes the monitor, discards queued actions, renders a fresh observation and asks the same policy to plan again.

Replan information flow
StagePolicy receivesGateway does
Before refusalCurrent cameras, state and task textEvaluates each proposed setpoint.
Hold and verifyNo additional inputHolds last permitted target and tests measured stability.
RecoveryFresh cameras and state; same task textResets, re-primes, discards queue and evaluates a new chunk.

Recovery is limited to three replans. A fourth refusal produces a fail-safe halt. The supervisor remains simulated; a real system must define who may clear a latched stop and what independently observed evidence permits it.

§6

What this establishes

SmolVLA experiment claims
StatementStatus
Endstop can sit in this public SmolVLA action path.Supported in simulation
The sampled nominal placements remain unchanged.Supported for this subset
The specified state bias can cause model-authored refusals.Supported for this scenario
Bounded recovery improves this placement stratum.Supported for 10 matched cases

Across the valid corpus the local gateway round trip had a median of 0.102 ms and a p95 of 0.222 ms, in simulation.

§7

Read the controlled report and data

The Protocol V2 report carries the frozen matrix, exact counterfactual method, amendments, recovery state machine, placement and full-task results, timing, kinematic validation, claim boundary and exclusions.

The artifact bundle contains the V2 runner and analyzer, Rust gateway source, protocol and amendments, all 114 valid raw episode traces, monitor configurations, FK checks, block analyses, combined index and SHA-256 manifests. Calibration, failed-pairing and interrupted datasets are retained locally with their exclusion reasons but are not in the headline corpus.

Download Protocol V2 report · PDF → Download Protocol V2 artifacts · TAR.GZ →

V2 report ETR-2026-08-20-03 · Revision B · report SHA-256 9f31a9b171f2bd45ff50a3b659e20f7fdeabe17b16bcdbf8ac31fda60447dfd2 · artifact bundle SHA-256 696a7379692c453eb9f048c5ccfb9c8cd5a720e9a59de8a6bb67802eaa3f643d.