
gpt-5.6-luna
API Overview
GPT-5.6-Luna is OpenAI’s next-generation, closed-source “lightning-fast” flagship large model designed for ultra-responsive performance and massive high-concurrency scenarios. As the most agile cost-reduction tool in the GPT-5.6 family, Luna maintains exceptional logical consistency while boosting its inference throughput to an astonishing 130 tokens/s—more than doubling the speed of Terra. The model demonstrates industry-leading energy efficiency in scenarios such as everyday lightweight programming completion, large-scale JSON formatting and extraction, and high-frequency agent edge-execution sub-nodes. It is the absolute top choice for enterprises seeking millisecond-level user interaction feedback while keeping API budgets tightly under control.
───────────────────────────────────────────────────────────────────
Core Capabilities
Hundred-character-level millisecond agility: Thanks to an extremely optimized lightweight hybrid expert routing algorithm, Luna achieves a blazing-fast throughput of up to 130 tokens/s, with first-token latency (TTFT) compressed to the millisecond level. In high-frequency human-machine interaction interfaces such as real-time intelligent customer service and line-level code agile completion, it delivers silky-smooth, zero-wait feedback.
100% fault-tolerant structured extraction: Native deep reinforcement of JSON Mode and fixed schema constraints ensures that, when handling daily batch form parsing, high-density web data extraction, or multimodal credential (OCR) formatting and classification tasks in heavy-duty backend pipelines, Luna’s format compliance rate approaches 100%, preventing business interruptions caused by format errors.
Efficient sub-nodes for multi-agent clusters: Specifically tailored for high-volume agent workflows. Luna boasts exceptionally stable single-step tool calls and highly robust conditional judgment capabilities. In complex composite agent systems, it excels at over 90% of basic format routing and agile execution tasks, significantly reducing the load on core flagship models.
Milions of tokens of seamless context without added cost: By default, Luna provides an enormous context window of 1.05 million tokens. It can effortlessly handle bulky technical manuals, complete long-term user conversation histories, or massive log files. Combined with the officially standard explicit prompt caching technology, enterprises can maintain ultra-long texts and long-term memory with virtually zero hidden costs.
Playground
Log in to explore more features! Click to Log In