The contact-weight decision was first made on n=3 rollouts. With per-sample metrics + paired Wilcoxon on n=20 val windows, the answer flips: Lc=1.0 is significantly WORSE than the Lc=0.25 baseline. These are the same 20 paired windows decoded per config — watch the tactile gel deformation.
📄 Methods, loss implementation & metric definitions →
Each video: GT (top) vs model rollout (bottom). See docs/RUNBOOK.md §publish.
Shipped v3_contact. Beats Lc=1.0 on contact MSE (p=0.011), IoU (p=0.024), tactile LPIPS (p=7e-4).
Shipped v3_contact. Beats Lc=1.0 on contact MSE (p=0.011), IoU (p=0.024), tactile LPIPS (p=7e-4).
No contact-aware loss. Significantly worse on view LPIPS (p=2e-6) — the aux does help, at 0.25.
No contact-aware loss. Significantly worse on view LPIPS (p=2e-6) — the aux does help, at 0.25.
The n=3 'winner'. At n=20 paired it is significantly WORSE than 0.25 on contact MSE, IoU, tactile LPIPS.
The n=3 'winner'. At n=20 paired it is significantly WORSE than 0.25 on contact MSE, IoU, tactile LPIPS.
Activating false-contact suppression: also significantly worse than baseline.
Activating false-contact suppression: also significantly worse than baseline.