unitreerobotics/UnifoLM-ER-1

🤗 Hugging Face 来源apache-2.04.4B 参数8.9 GBsafetensors✓ 3 个校验和今天更新
需要做种者 →

UnifoLM-ER-1-4B

Project Page

Across 16 multimodal perception and understanding benchmarks, UnifoLM-ER-1 leads open-source models on seven and delivers overall performance comparable to leading proprietary models. Built on Qwen3-VL-4B, UnifoLM-ER-1 is trained on more than 5 million samples spanning image point prediction, object detection, multi-image reasoning, 2D trajectory prediction, 3D object detection, and multi-image spatial question answering. These data are co-trained with general image-text data, preserving broad vision-language capabilities while substantially improving spatial understanding and reasoning in embodied environments.

Benchmark Results

Model Open Source RoboVQA Ego-Plan2 RefSpatial-Bench Where2Place Pixmo-Point BLINK CV-Bench EmbSpatial RoboSpatial SAT VSI-Bench VSR ERQA RealWorldQA MME MMMU_VAL
UnifoLM-ER-1-4B Yes 62.4 55.1 61.7 82.0 73.8 93.4† 88.6 88.9 73.1 76.0 54.2 88.1 50.0 69.8 2223.3 54.7
RoboBrain2.0-7B* Yes 30.0 33.23 42.2 63.6 54.7 83.9 85.7 76.3 54.2 75.3 36.1 84.0 — 69.4 2057.4 44.4
Robix-7B* No 63.6 — — 41.9 29.5 87.6 86.5 77.4 — 71.1 44.6 83.3 42.5 70.7 2332.8 —
Pelican-7B* Yes 31.8 33.7 22.3 57.3 20.4 — 79.4 73.2 57.5 52.0 52.8 82.2 39.8 69.3 2141.9 51.1
Cosmos-R1-7B* Yes 38.8 26.0 5.6 2.9 8.2 — 76.7 68.9 42.4 82.7 25.4 82.4 — 67.6 2157.4 37.4
Cosmos3-Super-64B* Yes — — 57.0 71.0 — 90.3† 88.0 — 70.0 — 60.9 — 51.2 — — —
Qwen3-VL-4B* Yes 47.7 40.7 46.6 63.0 48.3 85.0† 85.1 79.6 61.7 68.7 59.3 81.6 41.3 71.0 2325.2 57.8
Qwen3-VL-8B* Yes 43.3 49.7 54.2 61.9 51.0 73.8† 86.2 78.5 66.9 67.3 59.4 83.2 45.8 70.6 2412.5 62.3
Embodied-R1-3B* Yes 51.8 26.5 39.7 69.5 49.4 78.5† 82.7 67.4 47.4 76.3 26.6 — 35.2 — — —
Embodied-R1.5-8B* Yes 61.0 53.8 54.2 74.0 64.8 83.0† 86.9 78.1 69.7 74.7 56.1 — 46.0 — — —
Molmo2-ER-4B* Yes — — 52.5 54.0 — 85.7† 87.8 78.8 — 78.0 74.5 — 46.8 — — —
Hy-Embodied-VLM-1.0-30B-A3B* Yes — 49.6 53.4 65.0 64.6 87.3† 89.7 82.7 69.4 78.0 — — 60.8 — — —
MiMo-Emb-7B* Yes 62.0 43.0 48.0 63.6 42.35 81.3 88.2 76.2 61.7 78.6 48.5 79.0 46.7 66.3 2320.8 26.4
Thinker-4B* Yes 62.7 63.7 61.0 72.0 57.4 84.6 86.3 80.2 70.8 72.7 65.4 81.5 — 71.9 2323.4 46.2
Wall-OSS-0.5-3B* Yes — — — 15.0 — — — — — — — — 33 44 — —
Lumo-1-Stage1-7B* No — — 51.0 69.1 — 82.4 86.4 75.6 62.6 74.7 — — — — — —
Gemini-ER 2‡ No — — 35.4 — — 90.6† 90.4 81.4 51.1 — — — 71.0 — — —
Gemini-ER 1.5‡ No — — 41.8 48.0 — — 83.6 73.4 57.7 62.0 39.9 — 47.0 — — —
Gemini 2.5 Pro‡ No — — 33.6 37.0 — 88.6† 85.9 78.0 71.3 74.7 51.1 — 56.0 — — —
Gemini 2.5 Flash‡ No — — 41.2 48.0 — 80.3† 85.5 76.2 73.4 73.3 45.3 — 47.5 — — —
Gemini 3.1 Pro‡ No — — 70.0 61.0 — 86.1† 88.6 — 65.1 — 47.5 — 65.2 — — —
GPT-5.6-sol‡ No — — 58.3 51.1 — 85.6† 85.2 80.7 66.8 21.3 — — 64.8 — — —
GPT-6-Astra‡ No — — 79.6 69.0 — 90.4† 87.3 83.3 73.4 31.3 — — 77.7 — — —

* Results are sourced from the models' official technical reports or publicly available papers.

‡ Results were obtained through tests using the models' official APIs.

† BLINK scores are averaged over the Relative Depth and Spatial Relation subtasks only; all reported results were obtained in our own testing.

Citation

@misc{unifolm-er-1,
  author       = {Unitree},
  title        = {UnifoLM-WLA-1.0: One Model Driven, Whole-Body Coordination},
  year         = {2026},
}