Llama 4 Maverick 17B 128E Instruct FP8
by Azure · family llama · listed Apr 2025TVW
Open multimodal Llama model for strong reasoning and fast responses
Current rates USD per 1M tokens
Input$0.25
Output$1.00
Cache read—
Cache write—
Against the market
On the record
- Model id
llama-4-maverick-17b-128e-instruct-fp8- Context window
- 1M tokens
- Max output
- 16K tokens
- Input modalities
- text, image
- Output modalities
- text
- Reasoning
- no
- Tool calls
- yes
- Open weights
- yes
- Knowledge cutoff
- Aug 2024
- Released
- 2025-04-05
- Last updated
- 2025-04-05
Back of the envelope
≈
Also listed at
- Llamafree in · free out
- Abacus$0.14 in · $0.59 out
- IO.NET$0.15 in · $0.60 out
- Deep Infra$0.20 in · $0.80 out
- Azure Cognitive Services$0.25 in · $1.00 out
- NovitaAI$0.27 in · $0.85 out
- Charm Hyper$0.274 in · $0.899 out
- watsonx.ai$0.371 in · $1.48 out