Gemini 3.5 Flash-Lite

Available · Language, Multimodal, Reasoning

Gemini 3.5 Flash-Lite is a multimodal reasoning model in Google's Flash-Lite line, released on 21 July 2026 through the Gemini API, Gemini Enterprise Agent Platform and the Gemini app. Based on Gemini 3.1 Flash-Lite, it takes text, image, audio, video and PDF input and returns text; Google aims it at low-latency, high-volume work such as agentic search and document processing. [1] [2] [3] Primary source

Timeline of Gemini 3.5 Flash-Lite →

Claims and evidence

  • Released Primary source[1]3.6 Flash and 3.5 Flash-Lite: Get started today [4]July 21, 2026 [5]Versions
  • Status Available Primary source[6]Gemini 3 models: gemini-3.5-flash-lite [7]Models available for at least 12 months after release
  • Successor of Gemini 3.1 Flash-LiteThe model card says 3.5 Flash-Lite is based on Gemini 3.1 Flash-Lite, and the Gemini API deprecations page names gemini-3.5-flash-lite as the recommended replacement for gemini-3.1-flash-lite. Primary source[2]Model Information: Model dependencies [6]Gemini 3 models: gemini-3.1-flash-lite [1]3.5 Flash-Lite: Built to scale agentic workflows
  • Change · Tool use Computer use is available as a built-in tool; the Gemini API model page lists it as a preview capability.Compared with Gemini 3.1 Flash-Lite Primary source[1] [3]Capabilities: Computer use
  • Change · Other Google reports higher quality than 3.1 Flash-Lite across thinking levels in coding, agentic, long-context and real-world task evaluations.Compared with Gemini 3.1 Flash-Lite Primary source[1]
  • Change · Other List price rose to 0.30 USD per million input tokens and 2.50 USD per million output tokens, from 0.25 USD and 1.50 USD for 3.1 Flash-Lite.Compared with Gemini 3.1 Flash-Lite Primary source[2]
  • Input text, image, audio, videoPDF input is also listed. Primary source[3]Supported data types: Inputs [2]Model Information: Inputs
  • Output text Primary source[3]Supported data types: Output [2]Model Information: Outputs
  • Feature Reasoning modeConfigurable thinking levels; Agent Platform says the model defaults to minimal thinking. Primary source[3]Capabilities: Thinking [5]introduction: thinking levels
  • Feature Computer useListed as a preview capability on the Gemini API model page. Primary source[1]3.5 Flash-Lite: Built to scale agentic workflows [3]Capabilities: Computer use
  • Context window 1,048,576 tokensThe model card gives a context window of up to 1M tokens; output is limited to 65,536 tokens. Primary source[3]Token limits: Input token limit [5]Token limits: Context window
  • Parameters Not disclosedNot disclosed by Google. No claim madeNo source
  • Access API, consumer appGemini API (Google AI Studio, Android Studio), Gemini Enterprise Agent Platform and the Gemini app; the announcement says it was also rolling out in Google Search. Primary source[1]3.6 Flash and 3.5 Flash-Lite: Get started today [2]Distribution

Lineage

Predecessors

Successors

No known successor.

Based on

Not derived from another model.

Variants and derived

None recorded.

Siblings

None recorded.

All ancestors

All descendants

None.

Variants

No variants recorded in this record.

Related AI Radar coverage

All model releases from Google on AI Radar →