SmolVLM

Available · Language, Multimodal · Milestone

SmolVLM is a 2B open-weight vision-language model from Hugging Face that takes sequences of images and text and produces text, released as Base, Synthetic and Instruct checkpoints under Apache 2.0. It follows the Idefics3 architecture with SmolLM2 1.7B as language backbone; Hugging Face calls it suitable for on-device use and added 256M and 500M sizes in January 2025. [1] [2] [3] Primary source

Timeline of SmolVLM →

Claims and evidence

  • Released Primary source[1]page date
  • Status Available Primary source[2]model repository (weights hosted on the Hub)
  • Derived from (other) SmolLM2SmolLM2 serves as the language backbone of the vision-language model (the 1.7B Instruct checkpoint after a context extension to 16k tokens); the 256M and 500M variants use SmolLM2-135M-Instruct and SmolLM2-360M-Instruct. Primary source[1]Architecture: SmolLM2 1.7B as the language backbone [2]metadata base_model: HuggingFaceTB/SmolLM2-1.7B-Instruct
  • Change · Architecture Uses SmolLM2 1.7B instead of Llama 3.1 8B as the language backbone, within the Idefics3 architecture.Compared with Idefics3 Primary source[1]Architecture
  • Change · Efficiency Compresses visual information 9x with pixel shuffle, compared with 4x in Idefics3.Compared with Idefics3 Primary source[1]Architecture
  • Change · Size About 2B parameters in total, with a 1.7B language backbone in place of an 8B one.Compared with Idefics3 Primary source[1]
  • Input text, image Primary source[2]introduction; Model Summary: Model type
  • Output text Primary source[2]introduction
  • Open weights Yes Primary source[1]TLDR: checkpoints released under Apache 2.0 [2]Model Summary: License
  • Context window 16k tokens tokensContext of the SmolLM2 backbone after extension, for the 2B model; the 256M and 500M variants use 8k. Primary source[1]Training Details: Context extension
  • Parameters 2BThe November 2024 model; later Hugging Face pages give 2.2B. Primary source[1]TLDR
  • Access Open-weights download, On device Primary source[2]introduction: suitable for on-device applications [1]What is SmolVLM?

Lineage

Predecessors

No known predecessor.

Successors

Based on

  • SmolLM2 · October 2024 · derived (other)

Variants and derived

None recorded.

Siblings

None recorded.

All ancestors

All descendants

Variants

SmolVLM 2B

Same dates as SmolVLM.

  • Variant Inline variant in this record.The November 2024 release: SmolVLM-Base, SmolVLM-Synthetic and SmolVLM-Instruct. Later Hugging Face pages call this size SmolVLM 2.2B. Primary source[1]What is SmolVLM?: three released models [2]Model Summary

Also known as: SmolVLM-Instruct, SmolVLM-Base, SmolVLM-Synthetic, SmolVLM 2.2B

SmolVLM-256M

  • Released Primary source[3]page date
  • Variant Inline variant in this record.Added to the SmolVLM family on 2025-01-23 together with SmolVLM-500M, as base and Instruct checkpoints. Language backbone SmolLM2-135M-Instruct; vision encoder SigLIP base patch-16/512 (93M) instead of SigLIP 400M SO. Primary source[3]TLDR; What Changed Since SmolVLM 2B? [5]metadata base_model; Technical Summary

Also known as: SmolVLM-256M-Base, SmolVLM-256M-Instruct

Differs in:

  • Context window: 8k tokens Evidence not assessed [4]section 2.2: 8k-token limit for the smaller variants
  • Parameters: 256M Evidence not assessed [3]TLDR

SmolVLM-500M

  • Released Primary source[3]page date
  • Variant Inline variant in this record.Added to the SmolVLM family on 2025-01-23 together with SmolVLM-256M, as base and Instruct checkpoints. Language backbone SmolLM2-360M-Instruct; vision encoder SigLIP base patch-16/512 (93M). Primary source[3]TLDR; A Step Up: 500M [6]metadata base_model

Also known as: SmolVLM-500M-Base, SmolVLM-500M-Instruct

Differs in:

  • Context window: 8k tokens Evidence not assessed [4]section 2.2: 8k-token limit for the smaller variants
  • Parameters: 500M Evidence not assessed [3]TLDR

Related AI Radar coverage

AI Radar coverage starts in June 2026; no coverage linked yet.

All model releases from Hugging Face on AI Radar →