SmolVLM2

Available · Language, Multimodal · Milestone

SmolVLM2 extends Hugging Face's SmolVLM vision-language models to video, taking video, images and text as input and producing text. It was released with open weights under Apache 2.0 in 2.2B, 500M and 256M sizes, with MLX support from launch; Hugging Face showed an iPhone app that runs the 500M model locally. [1] [2] Primary source

Timeline of SmolVLM2 →

Claims and evidence

  • Released Primary source[1]page date
  • Status Available Primary source[2]model repository (weights hosted on the Hub)
  • Successor of SmolVLM Primary source[1]SmolVLM2 2.2B: compared with the previous SmolVLM family [2]metadata base_model: HuggingFaceTB/SmolVLM-Instruct
  • Change · Modality Accepts video input in addition to images and text.Compared with SmolVLM Primary source[2]Model Summary: Model type
  • Change · Training data Training data adds video datasets such as FineVideo to The Cauldron and Docmatix.Compared with SmolVLM Primary source[2]
  • Change · Availability Available for MLX in Python and Swift from launch.Compared with SmolVLM Primary source[1]
  • Change · Other Hugging Face reports better handling of math with images, text in photos, diagrams and scientific visual questions than the previous SmolVLM family.Compared with SmolVLM Primary source[1]SmolVLM2 2.2B
  • Input text, image, video Primary source[2]Model Summary: Model type
  • Output text Primary source[2]introduction
  • Open weights Yes Primary source[2]Model Summary: License
  • Access Open-weights download, On device Primary source[1]Suite of SmolVLM2 Demo applications: iPhone Video Understanding [2]introduction

Lineage

Predecessors

Successors

No known successor.

Based on

Not derived from another model.

Variants and derived

  • SmolVLA · 3 June 2025 · derived (other)

Siblings

None recorded.

All ancestors

All descendants

Variants

SmolVLM2-2.2B

Same dates as SmolVLM2.

  • Variant Inline variant in this record.Released as SmolVLM2-2.2B-Instruct; the base checkpoint SmolVLM2-2.2B-Base was published on the Hub later (repository created 2025-04-14). Primary source[1]Technical Details [2]Model Summary

Also known as: SmolVLM2 2.2B, SmolVLM2-2.2B-Instruct, SmolVLM2-2.2B-Base

Differs in:

  • Parameters: 2.2B Evidence not assessed [1]Technical Details

SmolVLM2-500M

Same dates as SmolVLM2.

  • Variant Inline variant in this record.The blog's iPhone demo app runs this model locally. Primary source[1]Going Even Smaller: Meet the 500M and 256M Video Models [3]card heading and introduction

Also known as: SmolVLM2-500M-Video, SmolVLM2-500M-Video-Instruct

Differs in:

  • Parameters: 500M Evidence not assessed [1]Technical Details

SmolVLM2-256M

Same dates as SmolVLM2.

  • Variant Inline variant in this record.Hugging Face calls it an experimental release. Primary source[1]Going Even Smaller: Meet the 500M and 256M Video Models

Also known as: SmolVLM2-256M-Video, SmolVLM2-256M-Video-Instruct

Differs in:

  • Parameters: 256M Evidence not assessed [1]Technical Details

Related AI Radar coverage

AI Radar coverage starts in June 2026; no coverage linked yet.

All model releases from Hugging Face on AI Radar →