Skip to content

PRM-as-a-Judge evaluation

ManiMux tracks the official PRM-as-a-Judge 1.5 evaluator as the PRM-as-a-Judge/ submodule and keeps model weights outside Git. Initialize it after cloning ManiMux with:

git submodule update --init PRM-as-a-Judge

Environment and inputs

The current machine has a dedicated Conda environment named prm-judge; do not install the judge's vLLM/CUDA dependencies into the ManiMux robot environment. On another machine, follow the Dopamine installation instructions in the upstream README and download the judge weights separately. Cloning the submodule does not download weights.

The current CLI requires --manifest; it does not recursively discover a video root. Create a JSONL file with one rollout per line, for example:

{"case_id":"pi05-rollout-003","task":"Assemble the screwdriver.","video":"/absolute/path/to/session/rollout-003/videos/front_camera.mp4","model":"pi05_step15000","benchmark":"yam_real_top"}

Required fields are case_id, task (or instruction), and video. Relative video paths are resolved against the manifest's parent directory. model and benchmark identify groups in the reports; use distinct case IDs within each model/task group. A human success label is optional and is not required for the progress-based evaluations shown here.

For a genuine three-view evaluation, additionally provide left_wrist_video and right_wrist_video on the same row. Do not add wrist videos when reproducing the top-only run.

Run evaluation

Use scripts/evaluation/prm_as_a_judge.py as the entry point. It forwards the official CLI unchanged and additionally supports PRM_GPU_MEMORY_UTILIZATION, because upstream currently fixes vLLM's budget at 0.9.

PRM_GPU_MEMORY_UTILIZATION=0.68 \
/home/ubuntu/miniconda3/envs/prm-judge/bin/python \
  scripts/evaluation/prm_as_a_judge.py eval \
  --manifest /path/to/manifest.jsonl \
  --output-root data/prm/results \
  --prm dopamine \
  --prm-path checkpoints/pretrained/Robo-Dopamine-GRM-2.0-8B-Preview \
  --gpus 0 \
  --eval-mode forward \
  --frame-interval 72 \
  --batch-size 10 \
  --outlier-method none \
  --smoothing none \
  --visualize

The local judge weights are stored at checkpoints/pretrained/Robo-Dopamine-GRM-2.0-8B-Preview/ (ignored by Git). Run the command from the ManiMux repository root. The previous location under /home/ubuntu/workspace/Project/PRM-as-a-Judge/PRM/ remains a compatibility symlink so historical evaluation paths continue to resolve.

For a top-only evaluation, provide only the manifest's required video field. The upstream Dopamine adapter fills all three model camera slots from that one video. This measures judgment from the requested top view; it is not a three-view evaluation.

The generated visualization_report.md currently calls every successfully processed record a "Successful case". Task success is instead the SR value in run_summary.json/per_case.jsonl, or the manifest's human label when metrics are run with --success-source label.

Results and sampling

--frame-interval 72 samples every 72 source frames (2.4 seconds for a 30 FPS recording). The command above preserves the previous comparison's forward mode and disables outlier removal and smoothing. PRM_GPU_MEMORY_UTILIZATION controls vLLM's GPU-memory budget, not a metric.

Each invocation writes data/prm/results/run_<timestamp>/ containing:

  • run_params.json and discovery_manifest.json: the selected inputs and evaluation settings;
  • per_case.jsonl and run_summary.json: rollout-level results and aggregate metrics;
  • metrics.xlsx and report.md: tabular reports;
  • visualizations/report.html: the visual report generated by --visualize.

To serve an existing run with video seeking support:

conda run --no-capture-output -n prm-judge python scripts/evaluation/prm_as_a_judge.py serve \
  --run-root data/prm/results/run_<timestamp> --host 127.0.0.1 --port 8000

Replace run_<timestamp> with the actual result directory, then open http://127.0.0.1:8000. The toolkit includes success-conditioned metrics, but a comparison can focus on MC@25/50/75, MP, PPL, CRA, STR and DRR without reporting success rate as the main outcome.