Inkling
by Nvidia · family ling · listed Jul 2026RTVW
Multimodal MoE reasoning model (975B total, 41B active) for text, image, and audio
Current rates USD per 1M tokens
Inputfree
Outputfree
Cache read—
Cache write—
Against the market
On the record
- Model id
thinkingmachines/inkling- Context window
- 1.0M tokens
- Max output
- 16K tokens
- Input modalities
- text, image, audio
- Output modalities
- text
- Reasoning
- yes
- Tool calls
- yes
- Open weights
- yes
- Knowledge cutoff
- —
- Released
- 2026-07-15
- Last updated
- 2026-07-15
- Provider docs
- docs.api.nvidia.com/nim/
Back of the envelope
≈
Also listed at
- Deep Infra$0.95 in · $4.05 out
- Baseten$1.00 in · $4.05 out
- Together AI$1.00 in · $4.05 out
- Merge Gateway$1.00 in · $4.05 out
- Vercel AI Gateway$1.00 in · $4.05 out
- OpenRouter$1.00 in · $4.05 out
- Hugging Face$1.00 in · $4.05 out
- Venice AI$1.25 in · $5.06 out
- Impossibl$1.87 in · $4.68 out
- Thinking Machines$1.87 in · $4.68 out