gemini-3.5-flash-lite

gemini-3.5-flash-lite

Google's ultra-lightweight, high-speed flagship model launched in July
2026-07-22
LLM
Model capability: imageModel capability: function_call
Input:
$0.3/1M tokens
Output:
$2.5/1M tokens
Bulk order? Contact your manager for exclusive deals

API Overview

Gemini-3.5-flash-lite is Google’s next-generation large model designed for high-concurrency production pipelines and sub-agent clusters, delivering “ultra-high throughput.” As the most agile and cost-effective lightweight kernel in the Gemini 3.5 generation, this model has been deeply optimized specifically for executing particular steps within multi-agent frameworks, performing frequent small-scale code iterations, and extracting massive amounts of text and multimodal data. It not only perfectly inherits the Gemini family’s native understanding capabilities for text, code, and multimodal assets but also reduces response latency and token consumption costs to entirely new levels, making it the ideal choice for enterprises handling billions of high-concurrency requests per day.

───────────────────────────────────────────────────────────────────

Core Capabilities

Efficient Sub-Agent Execution: Specially tailored for parallel subtasks in complex multi-agent systems. In Swarm or agentic workflows, it achieves extremely high instruction-following rates and first-token responses as low as milliseconds, reliably executing conditional judgments, JSON-structured extractions, and specific API tool calls.

Leapfrogging Intelligence-to-Cost Ratio and Ultra-High Throughput: Compared to the previous generation, 3.5 Flash-Lite maintains exceptionally low entry costs while significantly improving logical reasoning and text-generation quality. Its throughput (TPS) per second far surpasses that of comparable competitors, easily handling peak concurrency loads typical of large enterprises.

Milions of Tokens in Long Contexts with High-Fidelity Retrieval: By default, it supports an ultra-large context window of up to 1 million tokens and a single-output limit of up to 65,536 tokens. Even when dealing with massive engineering documents, long log streams, or consecutive rounds of historical conversations, it can achieve highly accurate, lossless retrieval.

Transparent Prompt Caching and Extreme Cost Control: It natively supports Google’s highly optimized Context Caching mechanism. In scenarios involving fixed system prompts or frequent knowledge-base retrievals, the input costs for cached hits can be reduced by another 90%, greatly alleviating the hidden cost burden on enterprises building long-memory agents.

Playground

Log in to explore more features! Click to Log In

API Analytics

API Reference (4)

API DescriptionAPI EndpointRequest MethodStabilityParameter Description
v1beta(Official Format - Chat)
POST
Stable
View Details
v1beta(Official Format - Streaming)
POST
Stable
View Details
Chat(Talk)
POST
Stable
View Details
Chat(Analyze image)
POST
Stable
View Details

API Pricing

$
ModelDescriptionContextOfficial Price302.AI PriceOfficial Price Gap

gemini-3.5-flash-lite

-
1000000

Cache Creation$0.03 / 1M tokens
Cache Read$0.03 / 1M tokens
Input$0.3 / 1M tokens
Output$2.5 / 1M tokens

Input$0.3 / 1M tokens
Output$2.5 / 1M tokens

Original Price