Home/Models/Gemini 2.5 Pro
GoogleGoogle
Multimodal2M tokens

Gemini 2.5 Pro

API Model ID:gemini-2.5-pro

2,000,000-token context champion with multimodal audio, video, and code understanding.

#2M Context#Video Analysis#Audio Understanding#Multimodal#Vision
Context Window2M tokens2,000,000 tokens
Max Output64K tokens64,000 tokens
Input Price$1.25per 1M input tokens
Output Price$5.00per 1M output tokens
Latency / Speed~140ms120 tokens/sec

Overview & Architecture

Gemini 2.5 Pro features Google's record-setting 2-million-token native context window. Capable of analyzing hours of video, entire libraries of audio recordings, and multi-million-line software repositories in a single prompt, Gemini 2.5 Pro is the undisputed champion for large-scale multimodal reasoning.

ArchitectureGoogle Multi-Modal Mixture Transformer with 2M Sparse Attention
Knowledge IndexCurrent (Continuously Indexed)

Benchmark Highlights

Needle-In-A-Haystack99.99%

2M-token context fidelity

MMMU78.4%

College-level multi-discipline multimodal QA

Video-MME84.1%

Comprehensive video understanding

MMLU-Pro89.1%

Advanced academic reasoning

Supported Capabilities

Function Calling / Tools
Supported
Vision & Image Inputs
Supported
Audio & Voice Inputs
Supported
Structured JSON Mode
Supported
System Prompts
Supported
Streaming Completions
Supported
Prompt Caching
Supported

Engineering Strengths

  • Industry-leading 2,000,000-token context window with perfect retrieval
  • Native high-fidelity processing of video, audio, PDFs, and images
  • Seamless integration with Google Cloud ecosystem and Vertex AI
  • Highly competitive pricing for million-token scale inputs

Recommended Production Workloads

  • Full repository audits, migrations, and architecture documentation
  • Long-form video transcription, frame-by-frame event detection, and analysis
  • Financial audio earnings call cross-referencing with quarterly filings

Execute via Gruvo

Unified Endpoint

Connect through Gruvo with OpenAI-compatible client libraries. Gruvo translates requests, manages provider streaming, and applies caching automatically.

import OpenAI from "openai";

// Configure client with Gruvo's unified AI gateway
const gruvo = new OpenAI({
  baseURL: "https://api.gruvo.ai/v1",
  apiKey: process.env.GRUVO_API_KEY,
  defaultHeaders: {
    "X-Gruvo-Provider": "gemini",
  },
});

const completion = await gruvo.chat.completions.create({
  model: "gemini-2.5-pro",
  messages: [
    { role: "system", content: "You are a production reasoning assistant." },
    { role: "user", content: "Analyze the architecture of our service." },
  ],
  temperature: 0.2,
});

console.log(completion.choices[0].message.content);

Automatic Failover & Routing

If Google suffers an unexpected outage or rate limit (HTTP 429), Gruvo can automatically route in-flight requests to equivalent tier models:

Gruvo Execution Layer

  • Real-time per-token latency and error observability
  • Unified spend tracking and budget caps
  • Bring your own provider API keys or use pooled keys
Start Free on Gruvo