Mixtral 8x7B

Available · Language · Milestone

Mixtral 8x7B is a sparse mixture-of-experts language model from Mistral AI with 46.7B parameters, 12.9B of them used per token, released on 11 December 2023 with open weights under Apache 2.0 and an instruct version. It was also served through the mistral-small API endpoint, later renamed open-mixtral-8x7b. [1] [2]February 26, 2024 [3] Primary source

Timeline of Mixtral 8x7B →

Claims and evidence

  • Released Other sources give: [4] The base weights were first shared through a magnet link that Mistral AI posted on X (post timestamp 2023-12-08T15:44Z); the announcement, the instruct model and API access followed on 2023-12-11. Primary source[1]page date [5]page header [6]Model Lifecycle: Mixtral 8x7B Base, Mixtral 8x7B Instruct
  • Deprecated Primary source[7]Deprecated table, Mixtral 8x7B 0.1 (open-mixtral-8x7b)
  • Retired Retirement on Mistral's platform; the open weights remain downloadable. Primary source[7]Deprecated table, Mixtral 8x7B 0.1 (open-mixtral-8x7b) [5]Retirement date [6]Model Lifecycle: Mixtral 8x7B Base, Mixtral 8x7B Instruct
  • Status AvailableThe Apache 2.0 base and instruct weights are still downloadable from Mistral AI's Hugging Face organisation. On Mistral's own platform the model is retired (models overview, Deprecated table; legal centre). Primary source[3] [8] [5]Weights table
  • Replaced by Mistral Small 4 Primary source[7]Deprecated table, Mixtral 8x7B 0.1, Alternative [5]Replacement
    • Change · Architecture Same architecture as Mistral 7B except that each layer has eight feedforward experts, of which a router uses two per token.Compared with Mistral 7B Primary source[9]Abstract [1]
    • Input text Primary source[5]Modalities
    • Output text Primary source[5]Modalities
    • Feature MultilingualEnglish, French, Italian, German and Spanish. Primary source[1]capabilities list
    • Open weights YesApache 2.0. Primary source[1]second paragraph [3] [8]
    • Context window 32k tokens Primary source[1]capabilities list [5]Context; Weights table [9]Abstract
    • Parameters 46.7B total, 12.9B per tokenSparse mixture of experts; the paper and the docs weights table round to 47B total and 13B active parameters. Primary source[1]Pushing the frontier of open models with sparse architectures [9]Abstract [5]Weights table
    • Access Open-weights download, API Primary source[1]Use Mixtral on our platform [10]Generative endpoints, Mistral-small [2]February 26, 2024

    Lineage

    Predecessors

    No known predecessor.

    Successors

    Based on

    Not derived from another model.

    Variants and derived

    None recorded.

    Siblings

    None recorded.

    All ancestors

    None.

    All descendants

    Variants

    No variants recorded in this record.

    Related AI Radar coverage

    AI Radar coverage starts in June 2026; no coverage linked yet.

    All model releases from Mistral AI on AI Radar →