All Models
pixtral-12b-2409
Pixtral 2409 12B is a state-of-the-art multimodal model with 12B parameters and a 400M vision encoder, natively trained on interleaved text and image data. It excels in tasks spanning vision-language reasoning, instruction following, and pure text understanding, making it highly effective for real-world multimodal applications.
Available Providers (2)
Capabilities
Reasoning
Tool Calling
Attachments
Open Weights
Structured Output