unitreerobotics/UnifoLM-ER-1

🤗 Hugging Face sourceapache-2.04.4B params8.9 GBsafetensorsHF checksums availableupdated today
No torrent yet

UnifoLM-ER-1-4B

Project Page

Across 16 multimodal perception and understanding benchmarks, UnifoLM-ER-1 leads open-source models on seven and delivers overall performance comparable to leading proprietary models. Built on Qwen3-VL-4B, UnifoLM-ER-1 is trained on more than 5 million samples spanning image point prediction, object detection, multi-image reasoning, 2D trajectory prediction, 3D object detection, and multi-image spatial question answering. These data are co-trained with general image-text data, preserving broad vision-language capabilities while substantially improving spatial understanding and reasoning in embodied environments.

Benchmark Results

Model Open Source RoboVQA Ego-Plan2 RefSpatial-Bench Where2Place Pixmo-Point BLINK CV-Bench EmbSpatial RoboSpatial SAT VSI-Bench VSR ERQA RealWorldQA MME MMMU_VAL
UnifoLM-ER-1-4B Yes 62.4 55.1 61.7 82.0 73.8 93.4 88.6 88.9 73.1 76.0 54.2 88.1 50.0 69.8 2223.3 54.7
RoboBrain2.0-7B* Yes 30.0 33.23 42.2 63.6 54.7 83.9 85.7 76.3 54.2 75.3 36.1 84.0 69.4 2057.4 44.4
Robix-7B* No 63.6 41.9 29.5 87.6 86.5 77.4 71.1 44.6 83.3 42.5 70.7 2332.8
Pelican-7B* Yes 31.8 33.7 22.3 57.3 20.4 79.4 73.2 57.5 52.0 52.8 82.2 39.8 69.3 2141.9 51.1
Cosmos-R1-7B* Yes 38.8 26.0 5.6 2.9 8.2 76.7 68.9 42.4 82.7 25.4 82.4 67.6 2157.4 37.4
Cosmos3-Super-64B* Yes 57.0 71.0 90.3 88.0 70.0 60.9 51.2
Qwen3-VL-4B* Yes 47.7 40.7 46.6 63.0 48.3 85.0 85.1 79.6 61.7 68.7 59.3 81.6 41.3 71.0 2325.2 57.8
Qwen3-VL-8B* Yes 43.3 49.7 54.2 61.9 51.0 73.8 86.2 78.5 66.9 67.3 59.4 83.2 45.8 70.6 2412.5 62.3
Embodied-R1-3B* Yes 51.8 26.5 39.7 69.5 49.4 78.5 82.7 67.4 47.4 76.3 26.6 35.2
Embodied-R1.5-8B* Yes 61.0 53.8 54.2 74.0 64.8 83.0 86.9 78.1 69.7 74.7 56.1 46.0
Molmo2-ER-4B* Yes 52.5 54.0 85.7 87.8 78.8 78.0 74.5 46.8
Hy-Embodied-VLM-1.0-30B-A3B* Yes 49.6 53.4 65.0 64.6 87.3 89.7 82.7 69.4 78.0 60.8
MiMo-Emb-7B* Yes 62.0 43.0 48.0 63.6 42.35 81.3 88.2 76.2 61.7 78.6 48.5 79.0 46.7 66.3 2320.8 26.4
Thinker-4B* Yes 62.7 63.7 61.0 72.0 57.4 84.6 86.3 80.2 70.8 72.7 65.4 81.5 71.9 2323.4 46.2
Wall-OSS-0.5-3B* Yes 15.0 33 44
Lumo-1-Stage1-7B* No 51.0 69.1 82.4 86.4 75.6 62.6 74.7
Gemini-ER 2 No 35.4 90.6 90.4 81.4 51.1 71.0
Gemini-ER 1.5 No 41.8 48.0 83.6 73.4 57.7 62.0 39.9 47.0
Gemini 2.5 Pro No 33.6 37.0 88.6 85.9 78.0 71.3 74.7 51.1 56.0
Gemini 2.5 Flash No 41.2 48.0 80.3 85.5 76.2 73.4 73.3 45.3 47.5
Gemini 3.1 Pro No 70.0 61.0 86.1 88.6 65.1 47.5 65.2
GPT-5.6-sol No 58.3 51.1 85.6 85.2 80.7 66.8 21.3 64.8
GPT-6-Astra No 79.6 69.0 90.4 87.3 83.3 73.4 31.3 77.7

* Results are sourced from the models' official technical reports or publicly available papers.

Results were obtained through tests using the models' official APIs.

BLINK scores are averaged over the Relative Depth and Spatial Relation subtasks only; all reported results were obtained in our own testing.

Citation

@misc{unifolm-er-1,
  author       = {Unitree},
  title        = {UnifoLM-WLA-1.0: One Model Driven, Whole-Body Coordination},
  year         = {2026},
}