Phi-3-vision

Available · Language, Multimodal

Phi-3-vision is a 4.2 billion parameter multimodal model from Microsoft that accepts text and images and returns text, the first multimodal model in the Phi-3 family. Introduced at Microsoft Build on 21 May 2024, it builds on Phi-3-mini, was optimised for chart and diagram understanding, and has open weights on Hugging Face under the MIT License. [1] [2] Primary source

Timeline of Phi-3-vision →

Claims and evidence

  • Released Primary source[1]You can try Phi-3-vision today [2]Release dates
  • Status AvailableWeights are still downloadable from Microsoft's Hugging Face organisation (checked 2026-10-01). Primary source[2]
  • Derived from (other) Phi-3Built from the Phi-3-mini language model (variant mini of microsoft.phi-3) combined with an image encoder, connector and projector. Primary source[2]Model, Architecture [1]Bringing multimodality to Phi-3
  • Change · Modality Adds image input to the Phi-3-mini language model through an image encoder, connector and projector, keeping a 128K-token context.Compared with Phi-3 Primary source[2] [1]
  • Input text, image Primary source[2]Model, Inputs [1]Bringing multimodality to Phi-3
  • Output text Primary source[2]Model, Outputs
  • Open weights Yes Primary source[2]License
  • Context window 128K tokens Primary source[2]Model, Context length
  • Parameters 4.2B Primary source[1]The Phi-3 family [2]Model, Architecture
  • Access Open-weights download Primary source[2]

Lineage

Predecessors

No known predecessor.

Successors

Based on

  • Phi-3 · 23 April 2024 · derived (other)

Variants and derived

None recorded.

Siblings

None recorded.

All ancestors

All descendants

Variants

No variants recorded in this record.

Related AI Radar coverage

AI Radar coverage starts in June 2026; no coverage linked yet.

All model releases from Microsoft on AI Radar →