Home/Models/GPT-5.6 Luna
OpenAIOpenAI
Fast256K tokens

GPT-5.6 Luna

API Model ID:gpt-5.6-luna

Sub-80ms low-latency nocturnal sub-tier model for high-frequency ambient workloads.

#Sub-80ms#High Speed#Cost Efficient#Autocomplete#Vision
Context Window256K tokens256,000 tokens
Max Output32K tokens32,000 tokens
Input Price$0.40per 1M input tokens
Output Price$1.60per 1M output tokens
Latency / Speed~75ms220 tokens/sec

Overview & Architecture

GPT-5.6 Luna delivers extreme throughput and sub-80ms time-to-first-token. Tailored for autocomplete, ambient copilot assistants, real-time message routing, and semantic classification, Luna combines aggressive parameter distillation with frontier instruction following at an ultra-low price point.

ArchitectureDistilled MoE with Hardware-Optimized Kernel Fusion
Knowledge IndexCurrent (Continuously Indexed)

Benchmark Highlights

HumanEval88.4%

Zero-shot code completion accuracy

Chatbot Arena Elo1310

Human conversational preference

MT-Bench8.95

Multi-turn dialog velocity and coherence

Supported Capabilities

Function Calling / Tools
Supported
Vision & Image Inputs
Supported
Audio & Voice Inputs
No
Structured JSON Mode
Supported
System Prompts
Supported
Streaming Completions
Supported
Prompt Caching
Supported

Engineering Strengths

  • Sub-80ms TTFT and over 220 tokens/sec sustained generation speed
  • Unmatched cost-to-performance ratio for interactive user interfaces
  • High reliability for quick single-turn classification and entity extraction
  • Clean adherence to structured schema outputs without verbose preamble

Recommended Production Workloads

  • Inline developer code completions and inline suggestions
  • Real-time user-facing chatbots requiring immediate responsiveness
  • High-throughput classification, routing, and intent triage

Execute via Gruvo

Unified Endpoint

Connect through Gruvo with OpenAI-compatible client libraries. Gruvo translates requests, manages provider streaming, and applies caching automatically.

import OpenAI from "openai";

// Configure client with Gruvo's unified AI gateway
const gruvo = new OpenAI({
  baseURL: "https://api.gruvo.ai/v1",
  apiKey: process.env.GRUVO_API_KEY,
  defaultHeaders: {
    "X-Gruvo-Provider": "openai",
  },
});

const completion = await gruvo.chat.completions.create({
  model: "gpt-5.6-luna",
  messages: [
    { role: "system", content: "You are a production reasoning assistant." },
    { role: "user", content: "Analyze the architecture of our service." },
  ],
  temperature: 0.2,
});

console.log(completion.choices[0].message.content);

Automatic Failover & Routing

If OpenAI suffers an unexpected outage or rate limit (HTTP 429), Gruvo can automatically route in-flight requests to equivalent tier models:

deepseek-v4.1-flashView fallback specs →
gemini-3.8-flashView fallback specs →

Gruvo Execution Layer

  • Real-time per-token latency and error observability
  • Unified spend tracking and budget caps
  • Bring your own provider API keys or use pooled keys
Start Free on Gruvo