
gemini-3.5-flash-lite
API Overview
Gemini-3.5-flash-lite is Google’s next-generation large model designed for high-concurrency production pipelines and sub-agent clusters, delivering “ultra-high throughput.” As the most agile and cost-effective lightweight kernel in the Gemini 3.5 generation, this model has been deeply optimized specifically for executing particular steps within multi-agent frameworks, performing frequent small-scale code iterations, and extracting massive amounts of text and multimodal data. It not only perfectly inherits the Gemini family’s native understanding capabilities for text, code, and multimodal assets but also reduces response latency and token consumption costs to entirely new levels, making it the ideal choice for enterprises handling billions of high-concurrency requests per day.
───────────────────────────────────────────────────────────────────
Core Capabilities
Efficient Sub-Agent Execution: Specially tailored for parallel subtasks in complex multi-agent systems. In Swarm or agentic workflows, it achieves extremely high instruction-following rates and first-token responses as low as milliseconds, reliably executing conditional judgments, JSON-structured extractions, and specific API tool calls.
Leapfrogging Intelligence-to-Cost Ratio and Ultra-High Throughput: Compared to the previous generation, 3.5 Flash-Lite maintains exceptionally low entry costs while significantly improving logical reasoning and text-generation quality. Its throughput (TPS) per second far surpasses that of comparable competitors, easily handling peak concurrency loads typical of large enterprises.
Milions of Tokens in Long Contexts with High-Fidelity Retrieval: By default, it supports an ultra-large context window of up to 1 million tokens and a single-output limit of up to 65,536 tokens. Even when dealing with massive engineering documents, long log streams, or consecutive rounds of historical conversations, it can achieve highly accurate, lossless retrieval.
Transparent Prompt Caching and Extreme Cost Control: It natively supports Google’s highly optimized Context Caching mechanism. In scenarios involving fixed system prompts or frequent knowledge-base retrievals, the input costs for cached hits can be reduced by another 90%, greatly alleviating the hidden cost burden on enterprises building long-memory agents.
Playground
Log in to explore more features! Click to Log In