gpt-5.6-luna

gpt-5.6-luna

Lightweight large model from OpenAI’s GPT-5.6 family
2026-07-10
LLM
Model capability: imageModel capability: function_call
Cache Read:
$0.02/1M tokens
Input:
$0.2/1M tokens
Output:
$1.2/1M tokens
Bulk order? Contact your manager for exclusive deals

API Overview

GPT-5.6-Luna is OpenAI’s next-generation, closed-source “lightning-fast” flagship large model designed for ultra-responsive performance and massive high-concurrency scenarios. As the most agile cost-reduction tool in the GPT-5.6 family, Luna maintains exceptional logical consistency while boosting its inference throughput to an astonishing 130 tokens/s—more than doubling the speed of Terra. The model demonstrates industry-leading energy efficiency in scenarios such as everyday lightweight programming completion, large-scale JSON formatting and extraction, and high-frequency agent edge-execution sub-nodes. It is the absolute top choice for enterprises seeking millisecond-level user interaction feedback while keeping API budgets tightly under control.

───────────────────────────────────────────────────────────────────

Core Capabilities

Hundred-character-level millisecond agility: Thanks to an extremely optimized lightweight hybrid expert routing algorithm, Luna achieves a blazing-fast throughput of up to 130 tokens/s, with first-token latency (TTFT) compressed to the millisecond level. In high-frequency human-machine interaction interfaces such as real-time intelligent customer service and line-level code agile completion, it delivers silky-smooth, zero-wait feedback.

100% fault-tolerant structured extraction: Native deep reinforcement of JSON Mode and fixed schema constraints ensures that, when handling daily batch form parsing, high-density web data extraction, or multimodal credential (OCR) formatting and classification tasks in heavy-duty backend pipelines, Luna’s format compliance rate approaches 100%, preventing business interruptions caused by format errors.

Efficient sub-nodes for multi-agent clusters: Specifically tailored for high-volume agent workflows. Luna boasts exceptionally stable single-step tool calls and highly robust conditional judgment capabilities. In complex composite agent systems, it excels at over 90% of basic format routing and agile execution tasks, significantly reducing the load on core flagship models.

Milions of tokens of seamless context without added cost: By default, Luna provides an enormous context window of 1.05 million tokens. It can effortlessly handle bulky technical manuals, complete long-term user conversation histories, or massive log files. Combined with the officially standard explicit prompt caching technology, enterprises can maintain ultra-long texts and long-term memory with virtually zero hidden costs.

Playground

Log in to explore more features! Click to Log In

API Analytics

API Reference (5)

API DescriptionAPI EndpointRequest MethodStabilityParameter Description
Chat(Talk)
POST
Stable
View Details
Chat (Image Analysis)
POST
Stable
View Details
Chat (Structured Output)
POST
Stable
View Details
Chat (function call)
POST
Stable
View Details
Responses
POST
Stable
View Details

API Pricing

$
ModelDescriptionContextOfficial Price302.AI PriceOfficial Price Gap

gpt-5.6-luna

-
1000000

Input$0.2 / 1M tokens
Output$1.2 / 1M tokens

Cache Read$0.02 / 1M tokens
Input$0.2 / 1M tokens
Output$1.2 / 1M tokens

Original Price