Pixtral 2409 12B is a state-of-the-art multimodal model with 12B parameters and a 400M vision encoder, natively trained on interleaved text and image data. It excels in tasks spanning vision-language reasoning, instruction following, and pure text understanding, making it highly effective for real-world multimodal applications.
Current rates USD per 1M tokens
Input$0.223
Output$0.223
Cache read—
Cache write—
Against the market
Input price vs. maker flagships ($/1M tokens)Output price vs. maker flagships ($/1M tokens)