Recursive AI cost intelligence

Build AI workflows that improve with every run

Tokens.io turns live pricing, context windows, cache economics, and workload fit into a recursive self-improvement loop for model selection, prompt design, and routing decisions.

OpenAI Anthropic Gemini Mistral xAI DeepSeek Cohere
Loop inputLiveprovider-backed signals
Fit range64K-1.05Mcontext tracked by model
Routing field7major providers mapped
FreshnessJul 6, 2026automated source monitor
Recursive improvement loop

Every run should teach the next run where to spend less

Tokens.io is the measurement layer for AI operations: live prices, context fit, cache economics, and workload signals feed back into better prompts, model choices, and routing policy.

  • Measure the cost and context footprint of each workload
  • Route to the cheapest model that still fits the job
  • Observe drift, cache misses, pricing changes, and spend spikes
  • Improve the next policy, prompt, route, or model substitution
AI model directory

The live reference for model economics

Provider, model ID, pricing, context, source link, and workload fit in one browsable directory built for fast decisions.

Google

Gemini 2.5 Flash-Lite

High-volume extraction, routing, and simple multimodal tasks

Input
$0.1
Output
$0.4
Context
1.05M
OpenAI

GPT-5.4 nano

Classification, extraction, ranking, and sub-agent workloads

Input
$0.2
Output
$1.25
Context
400K
DeepSeek

DeepSeek Chat

Low-cost chat, summarization, and extraction

Input
$0.27
Output
$1.1
Context
64K
Mistral

Mistral Large 3

Open-weight multimodal workloads and self-hosting evaluation

Input
$0.5
Output
$1.5
Context
256K
Google

Gemini 2.5 Flash

Large-scale, low-latency multimodal processing

Input
$0.3
Output
$2.5
Context
1.05M
DeepSeek

DeepSeek Reasoner

Cost-sensitive reasoning workloads

Input
$0.55
Output
$2.19
Context
64K
xAI

Grok Build 0.1

Agentic coding workflows on xAI

Input
$1
Output
$2
Context
256K
xAI

Grok 4.3

General chat and reasoning outside code, audio, image, and video

Input
$1.25
Output
$2.5
Context
1M
OpenAI

GPT-5.4 mini

Coding, computer-use, subagents, and efficient production agents

Input
$0.75
Output
$4.5
Context
400K
Anthropic

Claude Haiku 4.5

Fast Claude workloads with near-frontier intelligence

Input
$1
Output
$5
Context
provider tier
Google

Gemini 2.5 Pro

Complex reasoning, coding, and long multimodal context

Input
$1.25
Output
$10
Context
1.05M
Anthropic

Claude Sonnet 5

Balanced speed, intelligence, and production-agent cost

Input
$2
Output
$10
Context
1M
Provider-sourced intelligence

Live model pricing table

Prices are USD per 1M tokens unless noted. Rows are normalized from provider sources and updated by Tokens.io monitors.

ModelProviderInput / 1MCachedOutput / 1MContextSource
Gemini 2.5 Flash-Lite gemini-2.5-flash-liteGoogle$0.1$0.01$0.41.05MGemini pricing
GPT-5.4 nano gpt-5.4-nanoOpenAI$0.2$0.02$1.25400KOpenAI pricing
DeepSeek Chat deepseek-chatDeepSeek$0.27$0.07$1.164KDeepSeek pricing
Mistral Large 3 mistral-large-2512Mistral$0.5-$1.5256KMistral model card
Gemini 2.5 Flash gemini-2.5-flashGoogle$0.3$0.03$2.51.05MGemini pricing
DeepSeek Reasoner deepseek-reasonerDeepSeek$0.55$0.14$2.1964KDeepSeek pricing
Grok Build 0.1 grok-build-0.1xAI$1-$2256KxAI models
Grok 4.3 grok-4.3xAI$1.25-$2.51MxAI models
GPT-5.4 mini gpt-5.4-miniOpenAI$0.75$0.075$4.5400KOpenAI pricing
Claude Haiku 4.5 claude-haiku-4-5-20251001Anthropic$1$0.1$5provider tierClaude pricing
Gemini 2.5 Pro gemini-2.5-proGoogle$1.25$0.125$101.05MGemini pricing
Claude Sonnet 5 claude-sonnet-5Anthropic$2$0.2$101MClaude pricing
Instant calculator

Forecast monthly LLM spend before it becomes policy

Estimate input, output, and cache-hit economics before a workflow scales across users, agents, or teams.

  • Separate input and output pricing
  • Cache-aware estimates
  • Sorted recommendations that can feed the next routing decision
Monthly input -
Monthly output -
Cheapest route -
ModelProviderInputOutputEst. monthly
Cost watchlist

Find the next substitution before the bill asks for it

Sorted by a simple Tokens.io blended-cost estimate: 70% input and 30% output. Use it to spot candidate routes, then validate quality, latency, and policy fit.

#ModelProviderBlend / 1MContextSource
1 Gemini 2.5 Flash-Lite VolumeGoogle$0.191.05MGemini pricing
2 GPT-5.4 nano Volume reasoningOpenAI$0.515400KOpenAI pricing
3 DeepSeek Chat Low-cost chatDeepSeek$0.51964KDeepSeek pricing
4 Mistral Large 3 Open-weight frontierMistral$0.8256KMistral model card
5 Gemini 2.5 Flash Fast multimodalGoogle$0.961.05MGemini pricing
6 DeepSeek Reasoner Reasoning valueDeepSeek$1.04264KDeepSeek pricing
7 Grok Build 0.1 CodingxAI$1.3256KxAI models
8 Grok 4.3 General reasoningxAI$1.6251MxAI models
Characters-
Words-
Estimated tokens-
Approx. input cost on GPT-5.4 nano-
Tokenizer tool

Turn prompts, drafts, and agent traces into improvement signals

This lightweight tokenizer gives a planning estimate for text workloads so long prompts, repeated context, and agent traces can be improved before they scale.

Tokens.io model cost chart
Context window tracker

Know when context changes the next architecture decision

Long-context models change retrieval, agent memory, batch analysis, prompt packing, and the route that should run next.

ModelProviderContext windowMax outputBest for
GPT-5.4gpt-5.4OpenAI1.05M128KComplex professional work at lower cost than GPT-5.5
GPT-5.5gpt-5.5OpenAI1.05M128KHard reasoning, coding, and professional work
Gemini 2.5 Flash-Litegemini-2.5-flash-liteGoogle1.05M65.5KHigh-volume extraction, routing, and simple multimodal tasks
Gemini 2.5 Flashgemini-2.5-flashGoogle1.05M65.5KLarge-scale, low-latency multimodal processing
Gemini 2.5 Progemini-2.5-proGoogle1.05M65.5KComplex reasoning, coding, and long multimodal context
Grok 4.3grok-4.3xAI1M-General chat and reasoning outside code, audio, image, and video
Claude Sonnet 5claude-sonnet-5Anthropic1M-Balanced speed, intelligence, and production-agent cost
Claude Opus 4.8claude-opus-4-8Anthropic1M-Complex agentic coding and enterprise work
Claude Fable 5claude-fable-5Anthropic1M-Anthropic highest-capability work and long-running agents
GPT-5.4 nanogpt-5.4-nanoOpenAI400K128KClassification, extraction, ranking, and sub-agent workloads
Compare models

Pick the cheapest model that still fits the next job

Start with cost, context, and use-case fit, then let real workload results sharpen the next model substitution.

Google

Gemini 2.5 Flash-Lite

Input
$0.1
Output
$0.4
Context
1.05M

High-volume extraction, routing, and simple multimodal tasks

OpenAI

GPT-5.4 nano

Input
$0.2
Output
$1.25
Context
400K

Classification, extraction, ranking, and sub-agent workloads

DeepSeek

DeepSeek Chat

Input
$0.27
Output
$1.1
Context
64K

Low-cost chat, summarization, and extraction

Mistral

Mistral Large 3

Input
$0.5
Output
$1.5
Context
256K

Open-weight multimodal workloads and self-hosting evaluation

Google

Gemini 2.5 Flash

Input
$0.3
Output
$2.5
Context
1.05M

Large-scale, low-latency multimodal processing

DeepSeek

DeepSeek Reasoner

Input
$0.55
Output
$2.19
Context
64K

Cost-sensitive reasoning workloads

xAI

Grok Build 0.1

Input
$1
Output
$2
Context
256K

Agentic coding workflows on xAI

xAI

Grok 4.3

Input
$1.25
Output
$2.5
Context
1M

General chat and reasoning outside code, audio, image, and video

Provider pages

Free reference pages for every major AI provider

Provider cards summarize tracked models and link back to the underlying source docs for verification.

Anthropic

Tracked models
4
Lowest blended cost
$2.2
Largest context
1M
Provider source

Cohere

Tracked models
1
Lowest blended cost
$4.75
Largest context
256K
Provider source

DeepSeek

Tracked models
2
Lowest blended cost
$0.519
Largest context
64K
Provider source

Google

Tracked models
3
Lowest blended cost
$0.19
Largest context
1.05M
Provider source

Mistral

Tracked models
1
Lowest blended cost
$0.8
Largest context
256K
Provider source

OpenAI

Tracked models
4
Lowest blended cost
$0.515
Largest context
1.05M
Provider source
Source intelligence

What changed in the public reference

The public layer is built for real-time lookup. Historical archives, alerts, feeds, and webhooks will make those changes actionable inside the machine-readable layer.

Provider-sourced pricing pass

Current rows now point to official provider docs for input, cached input, output, context, and model ID fields.

Unsourced demand claims removed

The public table now ranks by Tokens.io blended-cost estimate until usage or market-share methodology is ready.

Tiered pricing noted

Rows use standard public text-token rates where providers publish multiple batch, flex, priority, or long-context tiers.

Claude intro pricing flagged

Claude Sonnet 5 is shown at Anthropic introductory pricing, which Anthropic lists through August 31, 2026.

Live index clarified

Live public lookup stays free; API access, alerts, exports, history, and team workflows are positioned as the future machine layer.

Educational explainers

Make token economics legible to everyone using AI

Short, evergreen explainers build search authority and help buyers understand why token spend matters.

What is a token?

A token is the billing and processing unit language models use for text, code, and structured data.

Input vs output tokens

Input tokens are what you send to a model; output tokens are what the model generates back. They are often priced differently.

How cached input changes cost

Prompt caching can reduce repeat input cost when providers can reuse stable context across requests.

How to estimate AI costs

Multiply monthly input and output tokens by each model price, then include cache, tool, and priority surcharges.

Recursive TokenOps

From free lookup to a self-improving AI operating layer

The public site starts the loop with live intelligence. Teams pay when they need the same signals inside APIs, dashboards, alerts, exports, automation, and internal routing systems.

01

Measure every run

See cost, tokens, cache behavior, and context footprint across products, teams, and workflows.

02

Route with intent

Move work to the cheapest model that still fits quality, latency, context, and policy targets.

03

Observe drift

Catch price changes, prompt bloat, runaway agents, and cache misses as soon as they appear.

04

Improve policy

Feed each result back into better prompts, substitutions, routing rules, and budgets.

Developer data access

Machine-readable intelligence closes the loop

The current reference stays free. API plans will give apps, dashboards, alerts, exports, and routing layers the same live signals production systems need to improve decisions automatically.