Home/Models/GPT-5.6 Terra
OpenAIOpenAI
Flagship1M tokens

GPT-5.6 Terra

API Model ID:gpt-5.6-terra

Grounded 1M-token context enterprise workhorse with rigorous factual assurance.

#1M Context#Enterprise#Grounded#High Fidelity#Tools
Context Window1M tokens1,000,000 tokens
Max Output64K tokens64,000 tokens
Input Price$2.50per 1M input tokens
Output Price$10.00per 1M output tokens
Latency / Speed~120ms110 tokens/sec

Overview & Architecture

GPT-5.6 Terra is designed as the bedrock for enterprise production systems. With a native 1-million-token context window and grounded factual safeguards, Terra excels at ingesting massive codebases, extensive regulatory documentation, and compliance archives with virtually zero needle-in-a-haystack retrieval degradation.

ArchitectureDense Multi-Head Attention with Linear Context Scaling
Knowledge IndexCurrent (Continuously Indexed)

Benchmark Highlights

Needle-In-A-Haystack99.98%

1M-token retrieval fidelity

MMLU-Pro89.7%

Broad enterprise domain knowledge

LegalBench94.5%

Statutory interpretation and contract review

FinanceBench91.2%

Complex balance sheet and SEC 10-K analysis

Supported Capabilities

Function Calling / Tools
Supported
Vision & Image Inputs
Supported
Audio & Voice Inputs
No
Structured JSON Mode
Supported
System Prompts
Supported
Streaming Completions
Supported
Prompt Caching
Supported

Engineering Strengths

  • Flawless retrieval and synthesis across up to 1,000,000 tokens
  • Calibrated uncertainty estimates that flag ambiguous or unverified facts
  • Rock-solid JSON Schema output conformity for mission-critical pipelines
  • Optimized cost structure for heavy document ingestion workloads

Recommended Production Workloads

  • Full-repository code understanding and multi-repo refactoring
  • Corporate contract analysis, discovery, and regulatory compliance
  • Enterprise knowledge management and long-form document synthesis

Execute via Gruvo

Unified Endpoint

Connect through Gruvo with OpenAI-compatible client libraries. Gruvo translates requests, manages provider streaming, and applies caching automatically.

import OpenAI from "openai";

// Configure client with Gruvo's unified AI gateway
const gruvo = new OpenAI({
  baseURL: "https://api.gruvo.ai/v1",
  apiKey: process.env.GRUVO_API_KEY,
  defaultHeaders: {
    "X-Gruvo-Provider": "openai",
  },
});

const completion = await gruvo.chat.completions.create({
  model: "gpt-5.6-terra",
  messages: [
    { role: "system", content: "You are a production reasoning assistant." },
    { role: "user", content: "Analyze the architecture of our service." },
  ],
  temperature: 0.2,
});

console.log(completion.choices[0].message.content);

Automatic Failover & Routing

If OpenAI suffers an unexpected outage or rate limit (HTTP 429), Gruvo can automatically route in-flight requests to equivalent tier models:

claude-sonnet-4.6View fallback specs →

Gruvo Execution Layer

  • Real-time per-token latency and error observability
  • Unified spend tracking and budget caps
  • Bring your own provider API keys or use pooled keys
Start Free on Gruvo