SmolLM2

Available · Language · Milestone

SmolLM2 is the second generation of Hugging Face's compact open-weight language models, released in 135M, 360M and 1.7B sizes under Apache 2.0. Hugging Face describes them as light enough to run on-device; the 1.7B model was trained on about 11 trillion tokens, with new math, code and instruction datasets built for the line. [1] [2] Primary source

Timeline of SmolLM2 →

Claims and evidence

  • Released No dated announcement by Hugging Face. The month follows from the Hub: base repositories created 2024-10-30 and 2024-10-31 (UTC), cards written on 2024-10-31, and the first public community discussions on the 1.7B Instruct repository on 2024-10-31. News sites report 2024-11-01. Primary source[1]Hub commit history: repository created 2024-10-30, card written 2024-10-31 [3]Hub commit history and community tab: repository created and first public discussions on 2024-10-31
  • Status Available Primary source[1]model repository (weights hosted on the Hub)
  • Successor of SmolLM Primary source[1]Model Summary: its predecessor SmolLM1-1.7B
  • Change · Training data The 1.7B model was trained on 11 trillion tokens, up from 1 trillion for SmolLM-1.7B.Compared with SmolLM Primary source[1] [4]
  • Change · Context length Context length of the 1.7B model extended from 2k to 8k tokens late in pre-training.Compared with SmolLM Primary source[2]section 4.6, Context Length extension
  • Change · Tool use The 1.7B Instruct model supports function calling, as documented on its model card.Compared with SmolLM Primary source[3]Examples: Function calling
  • Change · Training data Training adds new math, code and instruction datasets built by Hugging Face: FineMath, Stack-Edu and SmolTalk.Compared with SmolLM Primary source[2]
  • Input text Primary source[1]Model Summary; Limitations
  • Output text Primary source[1]Model Summary; Limitations
  • Open weights Yes Primary source[1]License: Apache 2.0 [2]Abstract
  • Access Open-weights download, On device Primary source[1]Model Summary: lightweight enough to run on-device

Lineage

Predecessors

Successors

Based on

Not derived from another model.

Variants and derived

  • SmolVLM · 26 November 2024 · derived (other)

Siblings

None recorded.

All ancestors

All descendants

Variants

SmolLM2-135M

Same dates as SmolLM2.

  • Variant Inline variant in this record.Trained on 2T tokens in a single stage, per the paper. Primary source[2]section 6, SmolLM2 135M and 360M

Also known as: SmolLM2-135M-Instruct

Differs in:

  • Parameters: 135M Evidence not assessed [2]section 6, SmolLM2 135M and 360M [1]Model Summary

SmolLM2-360M

Same dates as SmolLM2.

  • Variant Inline variant in this record.Trained on 4T tokens in a single stage, per the paper. Primary source[2]section 6, SmolLM2 135M and 360M

Also known as: SmolLM2-360M-Instruct

Differs in:

  • Parameters: 360M Evidence not assessed [2]section 6, SmolLM2 135M and 360M [1]Model Summary

SmolLM2-1.7B

Same dates as SmolLM2.

  • Variant Inline variant in this record.Trained on about 11T tokens in several stages, per the paper. A separate 16k-context fine-tune of the Instruct checkpoint, SmolLM2-1.7B-Instruct-16k, appeared on the Hub in February 2025 and is not recorded here. Primary source[1]Model Summary [2]Abstract

Also known as: SmolLM2-1.7B-Instruct

Differs in:

  • Context window: 8k tokens Evidence not assessed [2]section 4.6, Context Length extension
  • Parameters: 1.7B Evidence not assessed [1]Model Summary [2]Abstract

Related AI Radar coverage

AI Radar coverage starts in June 2026; no coverage linked yet.

All model releases from Hugging Face on AI Radar →