Comparable evaluation
Run policies on the same robot and task, with the same scoring. Compare success rates under matched conditions.
Remote evaluation on real robots
Run your policy remotely on a bimanual station.
Why Baseline
Run policies on the same robot and task, with the same scoring. Compare success rates under matched conditions.
Book a bimanual station by the minute. We own, operate, and maintain the hardware.
Hardware
We operate and calibrate each station. You choose the arms and submit a policy.
How it works
Choose a station, task, and episode count. We return the results.
uv tool install bsln
# Submit a policy evaluation to a station, then download the results.
bsln jobs submit \
--station bimanual/v1 \
--model pi0.5 \
--task fold_towel \
--episodes 32
bsln jobs download job_c91d --output ./runs/fold_towelAccess
Our stations are in Zurich. Join the waitlist and we will email you when access opens.