Mercury 2.5 Preview
by NanoGPT · family mercury · listed Sep 2026RT
Mercury 2.5 Preview is Inception's latest and most intelligent diffusion language model. Instead of generating tokens strictly one at a time, it produces and refines multiple tokens in parallel, reaching up to 1,107 tokens per second on standard GPUs. It delivers a 10+ point intelligence gain over Mercury 2, with tunable reasoning, parallel tool calls, schema-aligned JSON output, and a 260K context window. It is built for latency-sensitive production work such as search agents, voice pipelines, customer support, rapid coding iteration, and coding subagents.
Current rates USD per 1M tokens
Input$0.04
Output$0.15
Cache read$0.004
Cache write—
Against the market
On the record
- Model id
inception/mercury-2.5-preview- Context window
- 260K tokens
- Max output
- 66K tokens
- Input modalities
- text
- Output modalities
- text
- Reasoning
- yes
- Tool calls
- yes
- Open weights
- no
- Knowledge cutoff
- —
- Released
- 2026-09-01
- Last updated
- 2026-09-01
- Provider docs
- docs.nano-gpt.com
Back of the envelope
≈