GR00T N1.7 + YAM 运行手册¶
Red-ball step10000 with Pi-style RTC¶
The paired experiment is
manimux/configs/experiments/pick_red_object/groot/yam_groot_rtc_joint_step10000.yaml.
It selects ziyang/gr00t-n17-yam-pick-red-ball-box-step10000, not the robocurve
baseline described in the historical section below. The recipe uses the
checkpoint's three RGB views, absolute joint/gripper actions, 16 output steps,
four denoising iterations and a 1/30-second action interval.
The model implementation lives in XPolicyLab/policy/GR00T_N17/; ManiMux reuses
its existing RTC strategy, WebSocket client, JointAdapter and smooth executor.
inference.rtc.beta controls guided sampling, and
inference.rtc.min_execute_policy_steps controls the execution window. Initial
settings are beta 9.1, an 8-step window and a 4-step delay estimate. The runtime
updates its delay estimate from measured request/commit latency. A 16-step
horizon requires the conservative delay to fit within 8 steps (about 267 ms)
for the standard RTC overlap condition; inspect runtime delay events if it does
not. Changing the YAML horizon cannot extend the checkpoint's trained horizon.
Run each service in its own terminal from the repository root. Reuse an already running camera service only if it uses the same three-camera configuration.
cd /home/ubuntu/manimux
envs/yam/.venv/bin/python -m manimux.servers.camera.server \
--experiment manimux/configs/experiments/pick_red_object/groot/yam_groot_rtc_joint_step10000.yaml
cd /home/ubuntu/manimux
XPolicyLab/policy/GR00T_N17/gr00t_n17/.venv/bin/python -m manimux.servers.groot \
--experiment manimux/configs/experiments/pick_red_object/groot/yam_groot_rtc_joint_step10000.yaml
cd /home/ubuntu/manimux
envs/yam/.venv/bin/python -m manimux serve \
--config manimux/configs/experiments/pick_red_object/groot/yam_groot_rtc_joint_step10000.yaml
The experiment enables real YAM execution when a rollout is started in the RoboGUI.
All three commands resolve the private station's camera and policy endpoints.
The GR00T model environment uses Python 3.10; its launcher delegates experiment
resolution to envs/yam/.venv/bin/python, then loads the model in its own environment.
Use --local /path/to/station.yaml consistently to select another station.
Adding --check to the model command validates configuration and local assets
without loading the GPU model or connecting hardware; it does not prove inference.
Validation on 2026-09-28: 14 existing adapter/WS tests and 7 focused offline RTC
checks passed. A reduced real DiT matched the pre-change sampler exactly for
ordinary inference, AAC and upstream prefix freezing with fixed seeds. Full
red-ball step10000 weights ran on an RTX 4090 using synthetic RGB and checkpoint
mean state: finite (16, 14) output, identical zero-mask/default results and
post-reset results, no accumulated parameter gradients, and changed output under
nonzero guidance. Four warm RTC calls took 98.7–126.2 ms; these timings exclude
camera capture, network transport, runtime decoding and execution. The default
recipe's Cosmos repository ID also loaded successfully. No camera, robot or
real-task rollout was exercised in these checks.
Historical robocurve baseline (2026-08-20)¶
The following records an earlier checkpoint and deployment. Its old experiment paths and hardware results are not acceptance evidence for the red-ball RTC recipe.
本文只覆盖 robocurve/gr00t-n1.7-yam-molmoact2。它是基于 NVIDIA
Isaac-GR00T N1.7 的 YAM 微调权重,不是
nvidia/GR00T-N1.7-3B base 直接零样本上 YAM。
当前状态:XPolicy adapter、checkpoint 契约、模型环境、Cosmos、GPU forward、 XPolicy WebSocket、默认 ManiMux、真实三相机、双臂 YAM 和 Recorder 已完成闭环验证。当前 checkpoint 在已跑场景未完成任务,这是 policy 质量结果,不是 infra 断路。
契约¶
3 RGB cameras + 14D current joint state + instruction
-> XPolicyLab GR00T_N17 adapter
-> GR00T N1.7 YAM checkpoint + checkpoint statistics
-> 16 x 14 absolute joint positions at 30 Hz
-> ManiMux ActionTimeline -> SmoothExecutor -> YAM
| 项目 | 值 |
|---|---|
| 官方模型代码 | XPolicyLab/policy/GR00T_N17/gr00t_n17/(vendored Isaac-GR00T N1.7) |
| XPolicy adapter | XPolicyLab/policy/GR00T_N17/model.py |
| 权重 | /home/ubuntu/manimux/checkpoints/finetuned/robocurve/gr00t-n1.7-yam-molmoact2 |
| 发布来源 | robocurve/gr00t-n1.7-yam-molmoact2 |
| normalization | checkpoint 自带 statistics.json 的 new_embodiment |
| cameras | base_view、left_wrist_view、right_wrist_view |
| state/action | left_arm[6] + left_gripper[1] + right_arm[6] + right_gripper[1] |
| action | 16 步、14D、absolute joint position |
| 时间语义 | 训练数据 30 Hz,action_dt_s = 1/30 |
| denoise | checkpoint num_inference_timesteps = 4 |
权重目录中的 processor_config.json 和 statistics.json 是模型契约的一部分。不能把
Pi05 的 stats、XPolicy 原 ARX modality config 或 base 模型默认 embodiment 替换进来。
1. 离线检查¶
这一步不导入 torch、不加载权重到 GPU:
cd /home/ubuntu/manimux
envs/yam/.venv/bin/python manimux/servers/groot.py \
--config manimux/configs/policy/groot/yam/finetune.yaml \
--check
检查会确认:两片权重存在、3 相机键一致、state/action 都是 14D、action 是 absolute、
horizon 是 16、30 Hz 声明一致,并且 YAM 的 q01/q99 stats 维度完整。输出中的
contract_status: ready 只表示这些静态契约通过;只有模型环境与 Cosmos 都准备好后,
runtime_status 才会变成 environment_and_cosmos_present_gpu_forward_not_verified。这仍然
只代表本地文件完整,真实推理完成前
inference_status 始终是 not_verified。
2. 安装模型环境¶
当前本机 gr00t_n17/.venv 已完成依赖安装。新的 checkout 使用同一安装入口:
GR00T N1.7 的 processor 还需要 gated nvidia/Cosmos-Reason2-2B。先完成 Hugging Face
授权;若使用本地模型,把 manimux/configs/policy/groot/yam/finetune.yaml 中
cosmos_model_path 改为本地目录。
安装结束后重新运行第 1 步检查;不要在 runtime_status 仍为 blocked/operator action
时启动模型服务。
ManiMux 的 WebSocket worker 还需要通用 XPolicy 依赖。新环境按
xpolicylab-runbook.md 安装 .[xpolicylab];现有
envs/yam/.venv 已满足时无需重复安装。
3. 启动模型服务¶
Terminal 1:
cd /home/ubuntu/manimux
XPolicyLab/policy/GR00T_N17/gr00t_n17/.venv/bin/python \
manimux/servers/groot.py \
--config manimux/configs/policy/groot/yam/finetune.yaml
Terminal 2:先运行不接相机、不接 CAN 的单次 forward probe。它发送三张确定性的合成 RGB
图和配置里的 14D YAM 起始状态,并通过正式 XPolicy WebSocket 与 ManiMux adapter 检查
16 x 14 absolute joint chunk:
cd /home/ubuntu/manimux
envs/yam/.venv/bin/python scripts/validation/xpolicylab_yam_forward_probe.py \
--config manimux/configs/experiments/pick_box/yam_groot_manimux.yaml
只有输出 "status": "ok"、"action_space": "joint_position"、
"canonical_shape": [16, 14] 且动作全部有限,才算完成真实 GPU forward 和 WS 往返。仅看到端口
监听不算模型可用。
首次请求可能包含 CUDA/kernel/cache 冷启动,不能只用第一次延迟判断稳态。保持模型服务运行, 连续执行三次:
cd /home/ubuntu/manimux
for i in 1 2 3; do
echo "===== Probe $i ====="
envs/yam/.venv/bin/python scripts/validation/xpolicylab_yam_forward_probe.py \
--config manimux/configs/experiments/pick_box/yam_groot_manimux.yaml
done
通过标准:三次都返回有限的 native_shape / canonical_shape = [16, 14],且稳态
round_trip_ms 明显低于一个 16-step / 30 Hz chunk 的 533.3 ms 时长。
2026-08-20 实测:第一次冷启动请求 632.2 ms;随后三次为 101.3 / 91.7 / 89.4 ms。
因此 GPU、模型、WS、adapter 和稳态 latency gate 已通过;这仍不等于 ManiMux runtime、
Recorder 或真机闭环已经验证。
4. 相机、Preflight 与真机¶
Terminal 2 启动共享相机服务;已有 5555 服务时不要重复启动:
cd /home/ubuntu/manimux
envs/yam/.venv/bin/manimux-camera-server --config manimux/configs/embodiment/sensor/cameras/realsense_3_views.yaml
先检查两路 CAN:
for c in can_left can_right; do
ip -details link show "$c" | grep -o 'ERROR-ACTIVE\|ERROR-PASSIVE\|BUS-OFF'
done
然后运行只读 preflight。这个脚本会连接真实双臂和三相机并做两次 GR00T 推理,但会显式
关闭 move_to_start_on_connect、home_on_close,也不会发送模型动作:
cd /home/ubuntu/manimux
envs/yam/.venv/bin/python scripts/validation/pi05_base_yam_preflight.py \
--config manimux/configs/experiments/pick_box/yam_groot_manimux.yaml
确认 contract_checks 全为 true,并人工检查 measured state、first action、
max_first_joint_delta_rad、gripper range、camera shapes 和 steady latency。脚本名保留了早期
Pi05 命名,但其实现使用传入配置构建通用 YAM policy/adapter,此处不会加载 Pi05。
清空工作区、急停在手并确认 preflight 输出后,才启动默认 ManiMux:
cd /home/ubuntu/manimux
envs/yam/.venv/bin/manimux run --config manimux/configs/experiments/pick_box/yam_groot_manimux.yaml
真机 runtime 连接后会用 3.5 s 移动到配置起始位;正常 Ctrl-C 退出时会用 3.5 s
回 Home。不要在机械臂运动中关闭模型或相机服务;异常时优先物理急停。
2026-08-20 的两个受控 episode 分别记录了 91/74 个接受 chunk,无 plan rejection 或 invalid action,默认 ManiMux、三相机、双臂下发和 Recorder 因此已验收。两次都没有完成 pick 任务,只能记为当前 checkpoint 的闭环任务失败,不能倒推为 infra 未运行。
At the time of this baseline, the adapter exposed only ordinary inference and did not connect the upstream overlap/frozen-step branch to ManiMux. The new Pi-style RTC integration is described above; the old hardware runs did not exercise it.
未验证项¶
- 该 checkpoint 的任务成功率与跨任务泛化;
- Real-robot RTC task success with the red-ball checkpoint.