Phi-4-multimodal

Available · Language, Multimodal

Phi-4-multimodal is a 5.6-billion-parameter model in Microsoft's Phi family that accepts text, image and audio input and produces text, released on 26 February 2025 with Phi-4-mini. It handles all three modalities in one model through a mixture of LoRAs on a Phi-4-mini backbone, and was offered in Azure AI Foundry, the NVIDIA API Catalog and as MIT-licensed open weights. [1] [2] [3] Primary source

Timeline of Phi-4-multimodal →

Claims and evidence

  • Released Primary source[1]page date and first paragraph [2]Training, Model: Release date February 2025
  • Status AvailableWeights remain downloadable from Microsoft's Hugging Face organisation; Microsoft Foundry lists the model as GA without a retirement date (checked 2026-10-01). Primary source[2]weights and licence (MIT) [4]Foundry Models from partners and community, Microsoft: Phi-4-multimodal-instruct, GA, no retirement date
  • Derived from (adapter) Phi-4-miniUses the pretrained Phi-4-mini-instruct language model as backbone, kept frozen, with added vision and speech encoders, projectors and modality-specific LoRA adapters (mixture of LoRAs). Primary source[2]Training, Model: Architecture [3]Abstract; vision and speech/audio modality sections
  • Successor of Phi-3.5-vision EditorialEditorial link along the Phi multimodal line: the model card presents the release as a response to Phi-3 series feedback, where users had to chain a speech recognition model with the Mini and Vision models, but names no direct predecessor. Phi-3.5-vision was the latest Phi vision model before it. Primary source[2]Release Notes
  • Change · Modality Adds image and audio input to the Phi-4-mini language model, which stays frozen, through vision and speech encoders with LoRA adapters.Compared with Phi-4-mini Primary source[2]Training, Model: Architecture [3]
  • Change · Modality Adds speech and audio input; with the Phi-3 series, users had to put a separate speech recognition model in front of the Vision model.Compared with Phi-3.5-vision Primary source[2]Release Notes
  • Change · Architecture Processes text, vision and speech in a single model with a mixture of LoRAs, without separate pipelines per modality.Compared with Phi-3.5-vision Primary source[1]
  • Change · Languages Uses a larger vocabulary for multilingual support, with text input in 23 languages; vision input is English only.Compared with Phi-3.5-vision Primary source[2]Model Summary
  • Input text, image, audio Primary source[2]Model Summary; Training, Model: Inputs [1]Natively built for multimodal experiences
  • Output text Primary source[2]Model Summary; Training, Model: Outputs
  • Feature Function calling Primary source[2]Primary Use Cases: Function and tool calling
  • Feature MultilingualText in 23 languages, vision in English, audio in 8 languages according to the model card. Primary source[2]Model Summary: supported languages per modality
  • Open weights Yes Primary source[2]licence: MIT
  • Context window 128K tokens Primary source[2]Model Summary; Training, Model: Context length
  • Parameters 5.6B Primary source[1]What is Phi-4-multimodal? [2]Training, Model: Architecture
  • Access API, Open-weights download, cloud partnerAt launch: Azure AI Foundry (api), Hugging Face (open-weights-download) and the NVIDIA API Catalog (cloud-partner). Primary source[1]first paragraph; Learn more about Phi-4

Lineage

Predecessors

Successors

No known successor.

Based on

  • Phi-4-mini · 26 February 2025 · derived (adapter)

Variants and derived

None recorded.

Siblings

None recorded.

All ancestors

All descendants

None.

Variants

No variants recorded in this record.

Related AI Radar coverage

AI Radar coverage starts in June 2026; no coverage linked yet.

All model releases from Microsoft on AI Radar →