Gemini 2.5 Flash-Lite
High-volume extraction, routing, and simple multimodal tasks
- Input
- $0.1
- Output
- $0.4
- Context
- 1.05M
Tokens.io turns live pricing, context windows, cache economics, and workload fit into a recursive self-improvement loop for model selection, prompt design, and routing decisions.
Tokens.io is the measurement layer for AI operations: live prices, context fit, cache economics, and workload signals feed back into better prompts, model choices, and routing policy.
Provider, model ID, pricing, context, source link, and workload fit in one browsable directory built for fast decisions.
High-volume extraction, routing, and simple multimodal tasks
Classification, extraction, ranking, and sub-agent workloads
Low-cost chat, summarization, and extraction
Open-weight multimodal workloads and self-hosting evaluation
Large-scale, low-latency multimodal processing
Cost-sensitive reasoning workloads
Agentic coding workflows on xAI
General chat and reasoning outside code, audio, image, and video
Coding, computer-use, subagents, and efficient production agents
Fast Claude workloads with near-frontier intelligence
Complex reasoning, coding, and long multimodal context
Balanced speed, intelligence, and production-agent cost
Prices are USD per 1M tokens unless noted. Rows are normalized from provider sources and updated by Tokens.io monitors.
| Model | Provider | Input / 1M | Cached | Output / 1M | Context | Source |
|---|---|---|---|---|---|---|
| Gemini 2.5 Flash-Lite gemini-2.5-flash-lite | $0.1 | $0.01 | $0.4 | 1.05M | Gemini pricing | |
| GPT-5.4 nano gpt-5.4-nano | OpenAI | $0.2 | $0.02 | $1.25 | 400K | OpenAI pricing |
| DeepSeek Chat deepseek-chat | DeepSeek | $0.27 | $0.07 | $1.1 | 64K | DeepSeek pricing |
| Mistral Large 3 mistral-large-2512 | Mistral | $0.5 | - | $1.5 | 256K | Mistral model card |
| Gemini 2.5 Flash gemini-2.5-flash | $0.3 | $0.03 | $2.5 | 1.05M | Gemini pricing | |
| DeepSeek Reasoner deepseek-reasoner | DeepSeek | $0.55 | $0.14 | $2.19 | 64K | DeepSeek pricing |
| Grok Build 0.1 grok-build-0.1 | xAI | $1 | - | $2 | 256K | xAI models |
| Grok 4.3 grok-4.3 | xAI | $1.25 | - | $2.5 | 1M | xAI models |
| GPT-5.4 mini gpt-5.4-mini | OpenAI | $0.75 | $0.075 | $4.5 | 400K | OpenAI pricing |
| Claude Haiku 4.5 claude-haiku-4-5-20251001 | Anthropic | $1 | $0.1 | $5 | provider tier | Claude pricing |
| Gemini 2.5 Pro gemini-2.5-pro | $1.25 | $0.125 | $10 | 1.05M | Gemini pricing | |
| Claude Sonnet 5 claude-sonnet-5 | Anthropic | $2 | $0.2 | $10 | 1M | Claude pricing |
Estimate input, output, and cache-hit economics before a workflow scales across users, agents, or teams.
| Model | Provider | Input | Output | Est. monthly |
|---|
Sorted by a simple Tokens.io blended-cost estimate: 70% input and 30% output. Use it to spot candidate routes, then validate quality, latency, and policy fit.
| # | Model | Provider | Blend / 1M | Context | Source |
|---|---|---|---|---|---|
| 1 | Gemini 2.5 Flash-Lite Volume | $0.19 | 1.05M | Gemini pricing | |
| 2 | GPT-5.4 nano Volume reasoning | OpenAI | $0.515 | 400K | OpenAI pricing |
| 3 | DeepSeek Chat Low-cost chat | DeepSeek | $0.519 | 64K | DeepSeek pricing |
| 4 | Mistral Large 3 Open-weight frontier | Mistral | $0.8 | 256K | Mistral model card |
| 5 | Gemini 2.5 Flash Fast multimodal | $0.96 | 1.05M | Gemini pricing | |
| 6 | DeepSeek Reasoner Reasoning value | DeepSeek | $1.042 | 64K | DeepSeek pricing |
| 7 | Grok Build 0.1 Coding | xAI | $1.3 | 256K | xAI models |
| 8 | Grok 4.3 General reasoning | xAI | $1.625 | 1M | xAI models |
This lightweight tokenizer gives a planning estimate for text workloads so long prompts, repeated context, and agent traces can be improved before they scale.

Long-context models change retrieval, agent memory, batch analysis, prompt packing, and the route that should run next.
| Model | Provider | Context window | Max output | Best for |
|---|---|---|---|---|
| GPT-5.4gpt-5.4 | OpenAI | 1.05M | 128K | Complex professional work at lower cost than GPT-5.5 |
| GPT-5.5gpt-5.5 | OpenAI | 1.05M | 128K | Hard reasoning, coding, and professional work |
| Gemini 2.5 Flash-Litegemini-2.5-flash-lite | 1.05M | 65.5K | High-volume extraction, routing, and simple multimodal tasks | |
| Gemini 2.5 Flashgemini-2.5-flash | 1.05M | 65.5K | Large-scale, low-latency multimodal processing | |
| Gemini 2.5 Progemini-2.5-pro | 1.05M | 65.5K | Complex reasoning, coding, and long multimodal context | |
| Grok 4.3grok-4.3 | xAI | 1M | - | General chat and reasoning outside code, audio, image, and video |
| Claude Sonnet 5claude-sonnet-5 | Anthropic | 1M | - | Balanced speed, intelligence, and production-agent cost |
| Claude Opus 4.8claude-opus-4-8 | Anthropic | 1M | - | Complex agentic coding and enterprise work |
| Claude Fable 5claude-fable-5 | Anthropic | 1M | - | Anthropic highest-capability work and long-running agents |
| GPT-5.4 nanogpt-5.4-nano | OpenAI | 400K | 128K | Classification, extraction, ranking, and sub-agent workloads |
Start with cost, context, and use-case fit, then let real workload results sharpen the next model substitution.
High-volume extraction, routing, and simple multimodal tasks
Classification, extraction, ranking, and sub-agent workloads
Low-cost chat, summarization, and extraction
Open-weight multimodal workloads and self-hosting evaluation
Large-scale, low-latency multimodal processing
Cost-sensitive reasoning workloads
Agentic coding workflows on xAI
General chat and reasoning outside code, audio, image, and video
Provider cards summarize tracked models and link back to the underlying source docs for verification.
The public layer is built for real-time lookup. Historical archives, alerts, feeds, and webhooks will make those changes actionable inside the machine-readable layer.
Current rows now point to official provider docs for input, cached input, output, context, and model ID fields.
The public table now ranks by Tokens.io blended-cost estimate until usage or market-share methodology is ready.
Rows use standard public text-token rates where providers publish multiple batch, flex, priority, or long-context tiers.
Claude Sonnet 5 is shown at Anthropic introductory pricing, which Anthropic lists through August 31, 2026.
Live public lookup stays free; API access, alerts, exports, history, and team workflows are positioned as the future machine layer.
Short, evergreen explainers build search authority and help buyers understand why token spend matters.
A token is the billing and processing unit language models use for text, code, and structured data.
Input tokens are what you send to a model; output tokens are what the model generates back. They are often priced differently.
Prompt caching can reduce repeat input cost when providers can reuse stable context across requests.
Multiply monthly input and output tokens by each model price, then include cache, tool, and priority surcharges.
The public site starts the loop with live intelligence. Teams pay when they need the same signals inside APIs, dashboards, alerts, exports, automation, and internal routing systems.
See cost, tokens, cache behavior, and context footprint across products, teams, and workflows.
Move work to the cheapest model that still fits quality, latency, context, and policy targets.
Catch price changes, prompt bloat, runaway agents, and cache misses as soon as they appear.
Feed each result back into better prompts, substitutions, routing rules, and budgets.
The current reference stays free. API plans will give apps, dashboards, alerts, exports, and routing layers the same live signals production systems need to improve decisions automatically.
Ask for early access when you want Tokens.io data inside your product, dashboard, agent stack, or routing workflow.