Ling 3.0 Flash VL
by NanoGPT · family ling · listed Sep 2026RTVW
Ling 3.0 Flash VL is inclusionAI's native multimodal Mixture-of-Experts model with 124B total parameters and 5.5B active parameters per token. It combines image and video understanding with reasoning and tool use for document analysis, charts, visual verification, and interface-based agent tasks. Thinking is enabled by default and can be turned off in settings.
Current rates USD per 1M tokens
Input$0.06
Output$0.18
Cache read$0.012
Cache write—
Against the market
On the record
- Model id
inclusionai/ling-3.0-flash-vl- Context window
- 262K tokens
- Max output
- 33K tokens
- Input modalities
- text, image, video
- Output modalities
- text
- Reasoning
- yes
- Tool calls
- yes
- Open weights
- yes
- Knowledge cutoff
- —
- Released
- 2026-09-09
- Last updated
- 2026-09-09
- Provider docs
- docs.nano-gpt.com
Back of the envelope
≈
Also listed at
- OpenRouter$0.021 in · $0.0616 out
- LLM Gateway$0.06 in · $0.18 out
- DevPass (LLM Gateway)$0.06 in · $0.18 out
- Kilo Gateway$0.075 in · $0.22 out
- Vercel AI Gateway$0.075 in · $0.22 out