Gemini 1.5 Flash

Retired · Language, Multimodal

Gemini 1.5 Flash is a Google model that accepts text, images, audio and video and returns text, introduced at Google I/O 2024 in public preview in Google AI Studio and Vertex AI with a 1 million token context window. Google says it was distilled from 1.5 Pro to be fast and efficient to serve; from July 2024 it also ran the free Gemini app. [1] [2] [3]appendix 12.1 Model Card (Table 45) [4] Primary source

Timeline of Gemini 1.5 Flash →

Claims and evidence

  • Released Public preview in Google AI Studio and Vertex AI, announced at Google I/O.Other sources give: [5]May 10, 2024 The Gemini API release notes list the preview release of gemini-1.5-flash-latest under May 10, 2024, before the I/O announcement. Primary source[1]introduction [2]Availability
  • Retired Retirement date of gemini-1.5-flash-001 on Google Cloud (Agent Platform, formerly Vertex AI).Other sources give: [5]September 29, 2025 Shutdown of the gemini-1.5-flash model id on the Gemini API; from 2024-11-21 that id pointed to gemini-1.5-flash-002 (release notes, November 21, 2024). Primary source[6]Retired models: gemini-1.5-flash-001
  • Generally available (gemini-1.5-flash-001) Other sources give: [6]Retired models: gemini-1.5-flash-001 Release date of gemini-1.5-flash-001 on Google Cloud.; [7] Date of the developer post announcing the stable release and billing for 1.5 Flash and 1.5 Pro. Primary source[5]May 23, 2024
  • Model for the free version of the Gemini app Primary source[4] [8]2024.07.25
  • Status Retiredgemini-1.5-flash-001 is listed as retired on Google Cloud; the gemini-1.5-flash model id was shut down on the Gemini API on 2025-09-29. Primary source[6]Retired models: gemini-1.5-flash-001 [5]September 29, 2025
  • Replaced by Gemini 2.5 Flash-LiteRecommended upgrade for gemini-1.5-flash-001 as listed on the Agent Platform model versions page when checked on 2026-10-01. Primary source[6]Retired models: gemini-1.5-flash-001
  • Derived from (distillation) Gemini 1.5 Pro Primary source[1]The new 1.5 Flash, optimized for speed and efficiency [3]section 3.2 Gemini 1.5 Flash; appendix 12.1 Model Card (Table 45)
  • Change · Size A smaller, lighter-weight model than 1.5 Pro.Compared with Gemini 1.5 Pro Primary source[2] [1]
  • Change · Efficiency Designed for lower latency and lower cost to serve than 1.5 Pro, according to Google.Compared with Gemini 1.5 Pro Primary source[1]
  • Change · Architecture A dense Transformer decoder distilled online from 1.5 Pro, which is a sparse mixture-of-experts model.Compared with Gemini 1.5 Pro Primary source[3]section 3.2 Gemini 1.5 Flash
  • Input text, image, audio, video Primary source[2]Natively multimodal with long context [3]appendix 12.1 Model Card (Table 45), Input(s)
  • Output text Primary source[3]appendix 12.1 Model Card (Table 45), Output(s)
  • Feature Fine tuning availableText tuning in Google AI Studio and the Gemini API from August 2024; the release notes say tuning ended on 2025-05-27. Primary source[9]Gemini 1.5 Flash tuning rollout now complete [5]August 5, 2024
  • Context window 1 million tokensContext window in Google AI Studio and Vertex AI at launch. In the free Gemini app the window was 32K tokens (July 2024 post). Primary source[1]introduction [2]Natively multimodal with long context
  • Parameters Not disclosedNot disclosed by Google. No claim madeNo source
  • Access API, consumer appGemini API through Google AI Studio and Vertex AI; Gemini app from 2024-07-25. Primary source[1]introduction [4]

Lineage

Predecessors

No known predecessor.

Successors

No known successor.

Based on

Variants and derived

Siblings

None recorded.

All ancestors

All descendants

Variants

No variants recorded in this record.

Related AI Radar coverage

AI Radar coverage starts in June 2026; no coverage linked yet.

All model releases from Google on AI Radar →