Claims and evidence
- Released Primary source[1]page date; TL;DR
- Status Available Primary source[2]model repository (weights hosted on the Hub)
- Derived from (other) SmolVLM2SmolVLM2 is the pretrained vision-language backbone; a flow-matching action expert added on top generates the robot actions. Primary source[1]Method: Main Architecture, Vision-Language Model (VLM) [3]section 3.1, Vision-language model (VLM)
- Change · Modality Adds robot state input and continuous action output, generated by a flow-matching action expert on top of the SmolVLM2 backbone.Compared with SmolVLM2 Primary source[1]Method: Main Architecture
- Input image, text, structured dataMulti-view camera images, the robot's sensorimotor state and an optional language instruction. Primary source[2]Model description: Inputs
- Output actionsContinuous robot actions. Primary source[2]Model description: Outputs
- Open weights Yes Primary source[1]Introduction: releasing model weights [3]Abstract
- Parameters 450M Primary source[1]TL;DR
- Access Open-weights download Primary source[2]Quick start
Lineage
Predecessors
No known predecessor.
Successors
No known successor.
Based on
- SmolVLM2 · 20 February 2025 · derived (other)
Variants and derived
None recorded.
Siblings
None recorded.
All ancestors
All descendants
None.
Variants
No variants recorded in this record.
Related AI Radar coverage
AI Radar coverage starts in June 2026; no coverage linked yet.