SmolVLA

Available · Language, Multimodal, Robotics · Milestone

SmolVLA is a 450M-parameter open-weight vision-language-action model for robotics from Hugging Face's LeRobot project. It pairs SmolVLM2 with an action expert trained with flow matching to turn camera images, robot state and an optional instruction into continuous actions; Hugging Face says it can run on a CPU and pretrained it on community-shared LeRobot datasets. [1] [2] Primary source

Timeline of SmolVLA →

Claims and evidence

  • Released Primary source[1]page date; TL;DR
  • Status Available Primary source[2]model repository (weights hosted on the Hub)
  • Derived from (other) SmolVLM2SmolVLM2 is the pretrained vision-language backbone; a flow-matching action expert added on top generates the robot actions. Primary source[1]Method: Main Architecture, Vision-Language Model (VLM) [3]section 3.1, Vision-language model (VLM)
  • Change · Modality Adds robot state input and continuous action output, generated by a flow-matching action expert on top of the SmolVLM2 backbone.Compared with SmolVLM2 Primary source[1]Method: Main Architecture
  • Input image, text, structured dataMulti-view camera images, the robot's sensorimotor state and an optional language instruction. Primary source[2]Model description: Inputs
  • Output actionsContinuous robot actions. Primary source[2]Model description: Outputs
  • Open weights Yes Primary source[1]Introduction: releasing model weights [3]Abstract
  • Parameters 450M Primary source[1]TL;DR
  • Access Open-weights download Primary source[2]Quick start

Lineage

Predecessors

No known predecessor.

Successors

No known successor.

Based on

  • SmolVLM2 · 20 February 2025 · derived (other)

Variants and derived

None recorded.

Siblings

None recorded.

All ancestors

All descendants

None.

Variants

No variants recorded in this record.

Related AI Radar coverage

AI Radar coverage starts in June 2026; no coverage linked yet.

All model releases from Hugging Face on AI Radar →