Phi-3.5-vision

Available · Language, Multimodal

Phi-3.5-vision is a 4.2 billion parameter multimodal model from Microsoft that takes text and images and returns text, announced on 22 August 2024 with Phi-3.5-mini and Phi-3.5-MoE. It adds multi-frame input for comparing images and summarising image sets and video clips, and its weights are on Hugging Face under the MIT License. [1] [2] Primary source

Timeline of Phi-3.5-vision →

Claims and evidence

  • Announced Primary source[1]page date
  • Released Primary source[2]Model, Release date
  • Retired from Microsoft Foundry Primary source[3]Microsoft table
  • Status AvailableWeights are still downloadable from Microsoft's Hugging Face organisation (checked 2026-10-01). Retired from Microsoft Foundry on 2025-08-30 (see milestones). Primary source[2]
  • Successor of Phi-3-vision EditorialEditorial link along the vision line: the post says Phi-3.5-vision adds multi-frame understanding and improves single-image results, without naming Phi-3-vision as its predecessor. Primary source[1]Phi-3.5-vision with Multi-frame Input
  • Change · Modality Adds multi-frame input: the model can compare several images and summarise multi-image sets and video clips.Compared with Phi-3-vision Primary source[1] [2]
  • Input text, imageSingle images, multiple images and video frames as image input. Primary source[2]Model, Inputs
  • Output text Primary source[2]Model, Outputs
  • Open weights Yes Primary source[2]License
  • Context window 128K tokens Primary source[2]Model, Context length
  • Parameters 4.2B Primary source[2]Model, Architecture [4]abstract (v4)
  • Access Open-weights download Primary source[2]

Lineage

Predecessors

Successors

Based on

Not derived from another model.

Variants and derived

None recorded.

Siblings

None recorded.

All ancestors

All descendants

Variants

No variants recorded in this record.

Related AI Radar coverage

AI Radar coverage starts in June 2026; no coverage linked yet.

All model releases from Microsoft on AI Radar →