PRM-as-a-Judge evaluation¶
ManiMux tracks the official PRM-as-a-Judge 1.5 evaluator as the PRM-as-a-Judge/ submodule and
keeps model weights outside Git. Initialize it after cloning ManiMux with:
Environment and inputs¶
The current machine has a dedicated Conda environment named prm-judge; do not install the
judge's vLLM/CUDA dependencies into the ManiMux robot environment. On another machine, follow
the Dopamine installation instructions in the upstream README
and download the judge weights separately. Cloning the submodule does not download weights.
The current CLI requires --manifest; it does not recursively discover a video root.
Create a JSONL file with one rollout per line, for example:
{"case_id":"pi05-rollout-003","task":"Assemble the screwdriver.","video":"/absolute/path/to/session/rollout-003/videos/front_camera.mp4","model":"pi05_step15000","benchmark":"yam_real_top"}
Required fields are case_id, task (or instruction), and video. Relative video paths
are resolved against the manifest's parent directory. model and benchmark identify groups
in the reports; use distinct case IDs within each model/task group. A human success label
is optional and is not required for the progress-based evaluations shown here.
For a genuine three-view evaluation, additionally provide left_wrist_video and
right_wrist_video on the same row. Do not add wrist videos when reproducing the top-only run.
Run evaluation¶
Use scripts/evaluation/prm_as_a_judge.py as the entry point. It forwards the
official CLI unchanged and additionally supports
PRM_GPU_MEMORY_UTILIZATION, because upstream currently fixes vLLM's budget
at 0.9.
PRM_GPU_MEMORY_UTILIZATION=0.68 \
/home/ubuntu/miniconda3/envs/prm-judge/bin/python \
scripts/evaluation/prm_as_a_judge.py eval \
--manifest /path/to/manifest.jsonl \
--output-root data/prm/results \
--prm dopamine \
--prm-path checkpoints/pretrained/Robo-Dopamine-GRM-2.0-8B-Preview \
--gpus 0 \
--eval-mode forward \
--frame-interval 72 \
--batch-size 10 \
--outlier-method none \
--smoothing none \
--visualize
The local judge weights are stored at
checkpoints/pretrained/Robo-Dopamine-GRM-2.0-8B-Preview/ (ignored by Git).
Run the command from the ManiMux repository root. The previous location under
/home/ubuntu/workspace/Project/PRM-as-a-Judge/PRM/ remains a compatibility
symlink so historical evaluation paths continue to resolve.
For a top-only evaluation, provide only the manifest's required video field.
The upstream Dopamine adapter fills all three model camera slots from that one
video. This measures judgment from the requested top view; it is not a
three-view evaluation.
The generated visualization_report.md currently calls every successfully
processed record a "Successful case". Task success is instead the SR value
in run_summary.json/per_case.jsonl, or the manifest's human label when
metrics are run with --success-source label.
Results and sampling¶
--frame-interval 72 samples every 72 source frames (2.4 seconds for a 30 FPS recording).
The command above preserves the previous comparison's forward mode and disables outlier removal
and smoothing. PRM_GPU_MEMORY_UTILIZATION controls vLLM's GPU-memory budget, not a metric.
Each invocation writes data/prm/results/run_<timestamp>/ containing:
run_params.jsonanddiscovery_manifest.json: the selected inputs and evaluation settings;per_case.jsonlandrun_summary.json: rollout-level results and aggregate metrics;metrics.xlsxandreport.md: tabular reports;visualizations/report.html: the visual report generated by--visualize.
To serve an existing run with video seeking support:
conda run --no-capture-output -n prm-judge python scripts/evaluation/prm_as_a_judge.py serve \
--run-root data/prm/results/run_<timestamp> --host 127.0.0.1 --port 8000
Replace run_<timestamp> with the actual result directory, then open http://127.0.0.1:8000.
The toolkit includes success-conditioned metrics, but a comparison can focus on MC@25/50/75,
MP, PPL, CRA, STR and DRR without reporting success rate as the main outcome.