Gemini 2.5 Flash-Lite

Available · Language, Multimodal, Reasoning

Gemini 2.5 Flash-Lite is a multimodal reasoning model in Google's Gemini Flash-Lite line, released in preview on 17 June 2025 and generally available in the Gemini API from 22 July 2025. It has a 1 million token context window, a thinking budget that can be set and is off by default, and tools such as grounding with Google Search and code execution. [1] [2] [3]Introducing Gemini 2.5 Flash-Lite Primary source

Timeline of Gemini 2.5 Flash-Lite →

Claims and evidence

  • Released Released in preview (gemini-2.5-flash-lite-preview-06-17); generally available on 2025-07-22. Primary source[1] [4]June 17, 2025 [3]
  • Generally available (gemini-2.5-flash-lite) Primary source[4]July 22, 2025 [2] [5]Gemini 2.5 Flash models [6]Models available for at least 12 months after release
  • Updated preview (gemini-2.5-flash-lite-preview-09-2025) Primary source[4]September 25, 2025 [7]
  • Status AvailableNot deprecated in the Gemini API (no shutdown date announced), but since 2026-09-18 access to the 2.5 models is limited to users who used them before. On Gemini Enterprise Agent Platform the retirement of gemini-2.5-flash-lite is scheduled for 2026-10-20, with gemini-3.8-flash, gemini-3.1-flash-lite or Gemma 4 as replacements. Primary source[8]notice at the top of the page [5]Gemini 2.5 models [4]September 18, 2026 [6]Models available for at least 12 months after release
  • Successor of Gemini 2.0 Flash-LiteThe launch post and the model card compare it directly with 2.0 Flash-Lite, and the deprecations page names gemini-2.5-flash-lite as the replacement for the 2.0 Flash-Lite preview IDs. Primary source[1]Introducing Gemini 2.5 Flash-Lite [5]Gemini 2.0 models: preview models [9]Model Information: Description
  • Change · Reasoning Adds thinking with a configurable budget, off by default.Compared with Gemini 2.0 Flash-Lite Primary source[3]Introducing Gemini 2.5 Flash-Lite
  • Change · Tool use Can connect to tools such as Google Search and code execution.Compared with Gemini 2.0 Flash-Lite Primary source[1]Introducing Gemini 2.5 Flash-Lite
  • Change · Other Google reports lower latency than both 2.0 Flash-Lite and 2.0 Flash.Compared with Gemini 2.0 Flash-Lite Primary source[1]Introducing Gemini 2.5 Flash-Lite
  • Input text, image, audio, videoThe Gemini API also accepts PDF files. Primary source[8]Supported data types [9]Model Information: Inputs
  • Output text Primary source[8]Supported data types [9]Model Information: Outputs
  • Feature Reasoning modeThinking with a configurable budget; off by default for this model. Primary source[3]Introducing Gemini 2.5 Flash-Lite [9]Model Information: Description
  • Feature Function calling Primary source[3]Introducing Gemini 2.5 Flash-Lite [8]Capabilities
  • Feature Tool useGrounding with Google Search, code execution and URL context. Primary source[2] [8]Capabilities
  • Context window 1,048,576 tokens Primary source[8]Token limits: input token limit [9]Model Information: Inputs
  • Parameters Not disclosedNot disclosed by Google. No claim madeNo source
  • Access APIGemini API in Google AI Studio and Vertex AI. The launch post adds that custom versions of 2.5 Flash-Lite are used in Google Search. Primary source[1] [2]

Lineage

Predecessors

Successors

Based on

Not derived from another model.

Variants and derived

None recorded.

Siblings

None recorded.

All ancestors

All descendants

Variants

No variants recorded in this record.

Related AI Radar coverage

AI Radar coverage starts in June 2026; no coverage linked yet.

All model releases from Google on AI Radar →