On October 6, Mistral opened the public API preview of Mistral Large 4. Developers can already test the model through Mistral Studio, and the company plans to release the weights at the end of October. As of October 7, the model can be evaluated via API; its weights are not yet available to download. This status matters when choosing between quick testing through a service and self-hosting.
What developers can access now
A Mistral changelog entry confirms the public preview and a temporary 50% launch discount for two weeks. The model card lists support for multimodal input, function calling, structured output, batch processing, and tools for agents. This lets developers test the model on their own tasks via API before deciding whether to deploy it.
At the time of review, the model card lists prices of $0.68 per million input tokens, $0.07 per million cached input tokens, and $2.09 per million output tokens. The standard rates shown alongside them are twice as high: the discount is launch-related and time-limited. Check Mistral’s documentation directly before calculating a budget.
The specifications contain a discrepancy
Mistral describes Large 4 as a multimodal Mixture of Experts architecture. The documentation lists 1.05 trillion parameters in total, with 52 billion active; the company’s announcement lists 1 trillion and 49 billion active. The sources do not explain why the figures differ. It is therefore more accurate to report both versions with attribution than to consolidate them into a single figure.
The documentation specifies a context window of one million tokens. When planning long prompts, it is worth confirming this specification against the API version in use: preview parameters may be updated as the documentation changes.
Benchmarks cover specific tasks
Mistral reports scores of 61.7% on DeepSWE v1.1, 28.3% on Terminal-Bench 4, and 59.9% on AutomationBench. These are results published by the company for specific tests covering, respectively, agentic software development tasks, terminal workflows, and business process automation. They are best treated as guides for choosing evaluations, not as a single measure of model quality.
For a practical decision, developers should repeat several representative tasks from their own projects: record the model version, prompts, and settings; validate outputs with tests; and account for the cost of retries. This will show how well the published results transfer to specific code, documents, or agent workflows.
Open weights are the next step
Mistral said it plans to release the weights by the end of October, following additional red-team testing. This is the company’s future plan; as of publication, the confirmed way to use Large 4 is through the public API preview. For teams that need their own infrastructure and independent control of the model, the key milestone will be the actual release of the weights and the accompanying terms of use.
For now, developers can use the API to check availability, features, and quality, and defer the decision to move the model into their own environment until the weights are released. At launch, reproducible tests on real-world tasks and up-to-date pricing terms are more useful than comparative claims.