Model Catalogue

Frontier models.
One execution layer.

Discover, compare, and integrate 16+ frontier and specialized models across OpenAI, Anthropic, DeepSeek, Google, and TypeSafe AI with unified fallbacks, streaming, and observability.

Frontier Models
16 Models

Across 5 major AI ecosystems

Max Context
2,000,000

Tokens native context window

Lowest Latency
65ms TTFT

Real-time streaming throughput

Reliability
100% Fallback

Automated multi-provider failover

Showing 16 of 16 models
OpenAIOpenAI
Reasoning

GPT-5.6 Sol

gpt-5.6-sol

Solar-tier frontier reasoning engine engineered for lightning-speed multi-agent loops.

Context500K tokens
Latency~95ms TTFT
Input / 1M$3.00
Output / 1M$12.00
Specs & integration
OpenAIOpenAI
Flagship

GPT-5.6 Terra

gpt-5.6-terra

Grounded 1M-token context enterprise workhorse with rigorous factual assurance.

Context1M tokens
Latency~120ms TTFT
Input / 1M$2.50
Output / 1M$10.00
Specs & integration
OpenAIOpenAI
Fast

GPT-5.6 Luna

gpt-5.6-luna

Sub-80ms low-latency nocturnal sub-tier model for high-frequency ambient workloads.

Context256K tokens
Latency~75ms TTFT
Input / 1M$0.40
Output / 1M$1.60
Specs & integration
OpenAIOpenAI
Multimodal

GPT-5.5

gpt-5.5

Frontier multimodal foundation model balancing complex reasoning with cost.

Context500K tokens
Latency~110ms TTFT
Input / 1M$2.00
Output / 1M$8.00
Specs & integration
AnthropicAnthropic
Reasoning

Claude Opus 4.6

claude-opus-4.6

Deep deliberative reasoning engine built for nuanced synthesis and multi-hop analysis.

Context500K tokens
Latency~210ms TTFT
Input / 1M$12.00
Output / 1M$60.00
Specs & integration
AnthropicAnthropic
Coding

Claude Opus 4.7

claude-opus-4.7

Sovereign engineering intelligence designed for multi-agent architecture refactoring.

Context1M tokens
Latency~220ms TTFT
Input / 1M$14.00
Output / 1M$70.00
Specs & integration
AnthropicAnthropic
Flagship

Claude Opus 4.8

claude-opus-4.8

Pinnacle cognitive model surpassing human expert benchmarks in math and formal logic.

Context1M tokens
Latency~250ms TTFT
Input / 1M$16.00
Output / 1M$80.00
Specs & integration
AnthropicAnthropic
Multimodal

Claude Opus All

claude-opus-all

Unified omni-modal Opus engine integrating speech, native vision, code, and reasoning.

Context1M tokens
Latency~180ms TTFT
Input / 1M$15.00
Output / 1M$75.00
Specs & integration
AnthropicAnthropic
Coding

Claude Sonnet 4.5

claude-sonnet-4.5

High-velocity production workhorse for coding, agent loops, and enterprise workflow automation.

Context500K tokens
Latency~120ms TTFT
Input / 1M$3.00
Output / 1M$15.00
Specs & integration
AnthropicAnthropic
Coding

Claude Sonnet 4.6

claude-sonnet-4.6

State-of-the-art coding and agent execution model with ultra-precise tool call reliability.

Context1M tokens
Latency~110ms TTFT
Input / 1M$3.50
Output / 1M$17.50
Specs & integration
DeepSeekDeepSeek
Fast

DeepSeek V4.1 Flash

deepseek-v4.1-flash

Next-gen MoE inference engine delivering sub-cent high-throughput intelligence.

Context256K tokens
Latency~85ms TTFT
Input / 1M$0.12
Output / 1M$0.28
Specs & integration
DeepSeekDeepSeek
Reasoning

DeepSeek V4 Pro

deepseek-v4-pro

Frontier deliberative reasoning model rivaling western flagships at 1/10th the cost.

Context500K tokens
Latency~140ms TTFT
Input / 1M$0.45
Output / 1M$1.80
Specs & integration
DeepSeekDeepSeek
Fast

DeepSeek V4 Flash

deepseek-v4-flash

Blazing fast Mixture-of-Experts engine optimized for latency-critical microservices.

Context128K tokens
Latency~70ms TTFT
Input / 1M$0.080
Output / 1M$0.20
Specs & integration
GoogleGoogle
Multimodal

Gemini 2.5 Pro

gemini-2.5-pro

2,000,000-token context champion with multimodal audio, video, and code understanding.

Context2M tokens
Latency~140ms TTFT
Input / 1M$1.25
Output / 1M$5.00
Specs & integration
GoogleGoogle
Fast

Gemini 3.8 Flash

gemini-3.8-flash

Real-time multimodal speed demon with 1M context and near-instant time-to-first-token.

Context1M tokens
Latency~65ms TTFT
Input / 1M$0.15
Output / 1M$0.60
Specs & integration
TypeSafe AITypeSafe AI
Verification

TypeSafe Jev's

typesafe-jevs

The first System One AI model. Delivers ultra-fast, typed, zero-hallucination decisions at $0.042/1M tokens.

Context256K tokens
Latency~70ms TTFT
Input / 1M$0.042
Output / 1MFree
Specs & integration
One API Key • Zero Downtime

Route to any model with a single config change

Gruvo eliminates vendor lock-in. Switch between OpenAI, Anthropic, DeepSeek, Google, or TypeSafe AI dynamically. If an upstream provider rate-limits or fails, Gruvo automatically routes your request to your configured fallback without dropping the user stream.