Experiment records and optional evaluation¶
For everyday operation, use the RoboGUI research guide.
Free rollouts do not require scoring. Study rollouts support optional templates,
reference images and evaluation. This checkout keeps the team study register at
docs/experiments.md, outside the public documentation build. Other studies may
use their own register; everyday inference does not require one.
manimux serve keeps one configuration available for repeated GUI-driven Prepare
requests. Each creates a fresh rollout. manimux run creates one runtime immediately;
with RoboGUI enabled it starts paused, but connection/startup motion can already occur.
Neither command launches camera, model or RoboGUI services for you.
Saved evidence¶
data/experiments/<campaign>/<algorithm>/session-*/
├── session-manifest.json
└── rollout-001/
├── meta.json
├── events.jsonl
├── result.json
├── data.zarr/
│ ├── ticks/
│ └── plans/000000/
│ ├── canonical_raw/
│ ├── infra_output/
│ └── committed/
├── videos/
│ ├── <camera>.mp4
│ └── index.json
└── evaluation/
└── human-label.json
session-manifest.jsonstores the resolved configuration inconfig, the entry YAML's byte hash inconfig_sha256, and ManiMux/XPolicyLab git SHAs. The hash is not a digest of the resolved configuration, all dependencies, or dirty source.meta.jsonrecords task, layout, algorithm, experiment mode and the Policy Server fingerprint. New experiment rollouts also recordrepeat_idandreference_layout(task, absolutepath,sha256). The gallery task is distinct from the policy prompt and canonical evaluation task. These per-attempt fields are frozen at Prepare and also published for RoboGUI reconnection. Ordinary rollouts have no formal layout/repeat identity. Image hashes do not preserve overwritten files; keep formal references unchanged. Legacy episodes without these fields remain unknown.canonical_rawis the decoded policy chunk before the inference strategy.infra_outputis the chunk after the selected inference strategy.committedis the final horizon accepted by Timeline after trimming or blending.ticksstores measured state, scheduled reference, and command. The command is the executor output recorded afterrobot.send_command()returns; it is not a hardware acknowledgment. New recordings omitoptimized, which previously duplicatedcommandexactly. Existing recordings remain unchanged and may contain that historical field.videos/index.jsonstores camera timestamps, frame counts, dropped bundles and encoder errors.human-label.jsonexists only when an operator saves an evaluation. New labels usehuman-label-v2without a smoothness score. Historical v1 files remain unchanged; their task results remain usable and their old smoothness field is ignored.result.json.successmeans the runtime finalized normally; it is never task success.
Video recording is best-effort and asynchronous. A full video queue drops video bundles rather than
blocking the robot control loop. Formal analysis must inspect dropped_bundles and error. Track
task, seam, and PRM eligibility separately: unusable video does not erase a saved human task outcome.
Interpretation¶
result.json.successmeans runtime completion, not task success.- No human-label file means unreviewed, including when the operator chose Skip.
- Decoded policy actions, scheduled references, executor commands and measured state are distinct stages. An executor command is not a hardware acknowledgement.
- Capture timestamps, runtime ticks and video frame rates are different clocks/rates.
ticks.inference_msrepeats the latest accepted latency; averaging ticks does not give a per-request latency distribution.- A partial/aborted episode is not interchangeable with a finalized trial.
Session manifests record the resolved configuration and source/deployment provenance. Check the actual manifest and saved streams before claiming a particular run can be reproduced; checkpoint availability and hardware behavior require separate evidence.
Choose an evaluation for your question¶
For deployment comparisons, each method may use its own tuned configuration. Record those choices and the tuning budget; conclusions apply to the resulting deployment systems. To isolate one algorithmic change, hold the other relevant settings fixed. The framework supports both kinds of study without imposing either one.
Choose task criteria, layouts, repetitions and exclusions before a formal study. Preserve unsuccessful attempts, missing labels and invalid-trial reasons. Use the experiment skill for requested analysis and PRM guide for optional model-based judging.