Claims and evidence
- Released Primary source[1]page date and first paragraph [2]Training, Model: Release date February 2025
- Status AvailableWeights remain downloadable from Microsoft's Hugging Face organisation; Microsoft Foundry lists the model as GA without a retirement date (checked 2026-10-01). Primary source[2]weights and licence (MIT) [4]Foundry Models from partners and community, Microsoft: Phi-4-multimodal-instruct, GA, no retirement date
- Derived from (adapter) Phi-4-miniUses the pretrained Phi-4-mini-instruct language model as backbone, kept frozen, with added vision and speech encoders, projectors and modality-specific LoRA adapters (mixture of LoRAs). Primary source[2]Training, Model: Architecture [3]Abstract; vision and speech/audio modality sections
- Successor of Phi-3.5-vision EditorialEditorial link along the Phi multimodal line: the model card presents the release as a response to Phi-3 series feedback, where users had to chain a speech recognition model with the Mini and Vision models, but names no direct predecessor. Phi-3.5-vision was the latest Phi vision model before it. Primary source[2]Release Notes
- Change · Modality Adds image and audio input to the Phi-4-mini language model, which stays frozen, through vision and speech encoders with LoRA adapters.Compared with Phi-4-mini Primary source[2]Training, Model: Architecture [3]
- Change · Modality Adds speech and audio input; with the Phi-3 series, users had to put a separate speech recognition model in front of the Vision model.Compared with Phi-3.5-vision Primary source[2]Release Notes
- Change · Architecture Processes text, vision and speech in a single model with a mixture of LoRAs, without separate pipelines per modality.Compared with Phi-3.5-vision Primary source[1]
- Change · Languages Uses a larger vocabulary for multilingual support, with text input in 23 languages; vision input is English only.Compared with Phi-3.5-vision Primary source[2]Model Summary
- Input text, image, audio Primary source[2]Model Summary; Training, Model: Inputs [1]Natively built for multimodal experiences
- Output text Primary source[2]Model Summary; Training, Model: Outputs
- Feature Function calling Primary source[2]Primary Use Cases: Function and tool calling
- Feature MultilingualText in 23 languages, vision in English, audio in 8 languages according to the model card. Primary source[2]Model Summary: supported languages per modality
- Open weights Yes Primary source[2]licence: MIT
- Context window 128K tokens Primary source[2]Model Summary; Training, Model: Context length
- Parameters 5.6B Primary source[1]What is Phi-4-multimodal? [2]Training, Model: Architecture
- Access API, Open-weights download, cloud partnerAt launch: Azure AI Foundry (api), Hugging Face (open-weights-download) and the NVIDIA API Catalog (cloud-partner). Primary source[1]first paragraph; Learn more about Phi-4
Lineage
Predecessors
- Phi-3.5-vision · August 2024
Successors
No known successor.
Based on
- Phi-4-mini · 26 February 2025 · derived (adapter)
Variants and derived
None recorded.
Siblings
None recorded.
All ancestors
- Phi-1 · 20 June 2023
- Phi-1.5 · 11 September 2023
- Phi-2 · 12 December 2023
- Phi-3 · 23 April 2024
- Phi-3-vision · 21 May 2024
- Phi-3.5 · August 2024
- Phi-3.5-vision · August 2024
- Phi-4-mini · 26 February 2025
All descendants
None.
Variants
No variants recorded in this record.
Related AI Radar coverage
AI Radar coverage starts in June 2026; no coverage linked yet.