Gemini 2.5 Flash

Available · Language, Multimodal, Reasoning

Gemini 2.5 Flash is a multimodal reasoning model in Google's Gemini Flash line, released in preview on 17 April 2025 and generally available from 17 June 2025 in the Gemini API and the Gemini app. Google calls it its first fully hybrid reasoning model: thinking can be turned on or off and capped with a budget of up to 24,576 tokens. [1] [2] [3]April 17, 2025; June 17, 2025 Primary source

Timeline of Gemini 2.5 Flash →

Claims and evidence

  • Released Released in preview (gemini-2.5-flash-preview-04-17); generally available on 2025-06-17. Primary source[1] [3]April 17, 2025
  • Updated preview (gemini-2.5-flash-preview-05-20) Primary source[3]May 20, 2025 [4]Gemini 2.5 Flash models
  • Generally available (gemini-2.5-flash) Primary source[3]June 17, 2025 [5] [2] [4]Gemini 2.5 Flash models [6]Models available for at least 12 months after release
  • Updated preview (gemini-2.5-flash-preview-09-2025) Primary source[3]September 25, 2025 [7]
  • Status AvailableNot deprecated in the Gemini API (no shutdown date announced), but since 2026-09-18 access to the 2.5 models is limited to users who used them before. On Gemini Enterprise Agent Platform the retirement of gemini-2.5-flash is scheduled for 2026-10-20, with gemini-3.8-flash, gemini-3.5-flash-lite or gemini-3.1-flash-lite as replacements. Primary source[8]notice at the top of the page [4]Gemini 2.5 models [3]September 18, 2026 [6]Models available for at least 12 months after release
  • Change · Reasoning Adds controllable thinking: it can be switched on or off and capped with a thinking budget.Compared with Gemini 2.0 Flash Primary source[1]Fine-grained controls to manage thinking
  • Change · Other Google says it improves reasoning over 2.0 Flash while keeping that model's speed when thinking is off.Compared with Gemini 2.0 Flash Primary source[1]first paragraph
  • Input text, image, audio, video Primary source[8]Supported data types
  • Output text Primary source[8]Supported data types
  • Feature Reasoning modeHybrid reasoning: thinking can be turned on or off and capped with a thinking budget. Primary source[1]first paragraph; Fine-grained controls to manage thinking [8]Capabilities: Thinking
  • Feature Function calling Primary source[8]Capabilities
  • Feature Tool useCode execution, search grounding and URL context in the Gemini API. Primary source[8]Capabilities
  • Context window 1,048,576 tokens Primary source[8]Token limits: input token limit
  • Parameters Not disclosedNot disclosed by Google. No claim madeNo source
  • Access API, consumer appGemini API in Google AI Studio and Vertex AI, and the Gemini app. Primary source[1]Start building with Gemini 2.5 Flash today [5]

Lineage

Predecessors

Successors

Based on

Not derived from another model.

Variants and derived

None recorded.

Siblings

None recorded.

All ancestors

All descendants

Variants

No variants recorded in this record.

Related AI Radar coverage

AI Radar coverage starts in June 2026; no coverage linked yet.

All model releases from Google on AI Radar →