Claims and evidence
- Status AvailableOpen weights still offered (gated, on request) from Meta's meta-llama organisation on Hugging Face and from Meta's Llama 4 page on 2026-10-01. The hosted Llama API preview is a separate matter, see notes. Primary source[5]model page and access request form [6]Llama 4 Scout, Download
- Successor of Llama 3.3 EditorialEditorial link along the Llama line: Llama 4 Scout is among the first Llama 4 models, released after Llama 3.3. The announcement compares it with all previous Llama generations without naming a direct predecessor. Primary source[1]
- Derived from (distillation) Llama 4 BehemothMeta says Llama 4 Scout and Llama 4 Maverick benefit from distillation from Llama 4 Behemoth, which it trained as a teacher for the new models; the detailed codistillation description in the announcement is about Maverick. Primary source[1]Takeaways; opening paragraphs
- Change · Architecture One of Meta's first Llama models built on a mixture-of-experts architecture, with 16 experts and 17B of its 109B parameters active per token.Compared with Llama 3.3 Primary source[1] [2]Model Information table, Params
- Change · Modality Accepts images alongside text as input; Meta describes native multimodality with early fusion of text and vision tokens.Compared with Llama 3.3 Primary source[2]Model Information table, Input modalities [1]
- Change · Context length Context length raised from 128K tokens in Llama 3 to 10M tokens; Meta credits an architecture it calls iRoPE.Compared with Llama 3.3 Primary source[1]
- Change · Other Trained with distillation from Llama 4 Behemoth, a larger teacher model that was still in training at launch.Compared with Llama 3.3 Primary source[1]
- Input text, imageMeta's docs give text plus up to 5 images as input; the model card says image understanding was tested up to 5 input images and the docs say image understanding is English-only. Primary source[2]Model Information table, Input modalities [3]Introduction, feature table, Multimodal
- Output textThe docs say text-only output; the model card table lists the output as multilingual text and code. Primary source[3]Introduction, feature table, Multimodal [2]Model Information table, Output modalities
- Open weights YesGated download under the Llama 4 Community License Agreement, a custom commercial license (model card); released as BF16 weights. Primary source[1]opening paragraphs [5] [2]License; Quantization
- Context window 10M tokensThe announcement says the model was pre-trained and post-trained with a 256K context length; the docs say context lengths were evaluated across 512 GPUs. Primary source[2]Model Information table, Context length [3]Introduction, feature table, Maximum Context Length [1]Takeaways; Post-training our new models
- Parameters 17B (Activated), 109B (Total)Mixture-of-experts model with 16 experts; the value is the total parameter count, of which 17B are active per token. Primary source[2]Model Information table, Params [3]Introduction, feature table [1]Post-training our new models
- Access Open-weights download, APIAPI access through Meta's Llama API, a limited free preview announced on 2025-04-29. The launch post said partner availability would follow in the coming days; cloud-partner is left out because no Meta source read here names the partners. Primary source[1]opening paragraphs [5] [4]Llama API
Lineage
Predecessors
- Llama 3.3 · 6 December 2024
Successors
No known successor.
Based on
- Llama 4 Behemoth · 5 April 2025 · derived (distillation)
Variants and derived
None recorded.
Siblings
None recorded.
All ancestors
- LLaMA · 24 February 2023
- Llama 2 · 18 July 2023
- Meta Llama 3 · 18 April 2024
- Llama 3.1 · 23 July 2024
- Llama 3.2 · 25 September 2024
- Llama 3.3 · 6 December 2024
- Llama 4 Behemoth · 5 April 2025
All descendants
None.
Variants
No variants recorded in this record.
Related AI Radar coverage
AI Radar coverage starts in June 2026; no coverage linked yet.