AI Models & Platforms
Benchmark Firm Gives Mistral Large 4 Preview a 38 Intelligence Score

Artificial Analysis published its benchmark analysis of Mistral Large 4 Preview on October 6, 2026, scoring the model 38 on its Intelligence Index and describing it as the most intelligent AI model from outside the US and China.
Intelligence Index Results
The analysis reports that Mistral Large 4 Preview scores 38 on the Artificial Analysis Intelligence Index version 4.3.2, a result the benchmarking firm describes as comparable to GPT-6 Luna (max) at 38 and DeepSeek V4.1 Flash (max) at 39. In its framing, France is back to having the most intelligent model from outside the US and China, ahead of countries such as South Korea and the United Arab Emirates.
On Artificial Analysis’s model page for the preview, that score ranks the model number 64 of 225 models in its comparison class, above a class median of 26. The page lists the preview as a reasoning model released October 6, 2026, with text and image input, text output, and service through two API providers.
According to the published Intelligence Index methodology, version 4.3.2 combines ten evaluations — AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity’s Last Exam, GDP.pdf, CritPt, AA-Omniscience, and AA-LCR v1.1 — as a weighted average across four categories: agents at 30%, coding at 20%, scientific reasoning at 20%, and general at 30%. Artificial Analysis estimates a 95% confidence interval of less than ±1% for the index and describes the suite as primarily text-based and English-language, with image, speech, and multilingual performance benchmarked separately.
Cyber Index Results
On the Artificial Analysis Cyber Index, the preview scores 50, level with GLM-5.3-Flash at 50 and behind MiMo-V2.6-Pro at 56, while ahead of models including Kimi K3 and DeepSeek V4.1 Flash (max). Artificial Analysis states that once the model’s weights are released, it will rank among the top three open-weights models on the Cyber Index.
The model’s strongest individual cyber result came on CyberGym-E2E-AA, where it scores 82%, ahead of MiMo-V2.6-Pro at 79% and GPT-6 Luna (max) at 78%, according to the analysis.
Cost, Speed, and Token Use
The analysis puts the cost of running the Intelligence Index on Mistral Large 4 Preview at $1.13 per task, based on standard pricing of $1.36 per 1M input tokens and $4.18 per 1M output tokens, with cached input at $0.14 per 1M tokens. A 50% launch discount applies for the first two weeks, cutting token prices to $0.68 and $2.09 and bringing the cost per task to $0.57 — still above GLM-5.3-Flash at $0.25 and DeepSeek V4.1 Flash (max) at $0.27. Artificial Analysis characterizes the preview as more than four times the cost per task of similar-intelligence open-weights models.
The model page records the preview generating 200M output tokens over the course of the Intelligence Index run, against a class median of 81M. Measured through Mistral’s API, it reports output speed of 116.1 tokens per second against a median of 86.1, and a time to first token of 1.46 seconds against a median of 3.81 seconds.
Document Reasoning and Model Details
On GDP.pdf, a professional document-reasoning evaluation, Mistral Large 4 Preview scores 19%, on par with MiMo-V2.6-Pro at 19% and behind Kimi K3 at 22%. Artificial Analysis describes the result as an 18-point improvement over Mistral Large 3, partly driven by an API change that now accepts 100 images per request, up from eight for previous Mistral models.
The analysis lists a 512k-token context window and a 1-trillion-parameter model with 49 billion active parameters. Availability is limited to a Research Public Preview on Mistral’s API, with open weights planned for the end of October 2026; the model page currently classifies the preview as proprietary because its weights are not yet publicly available.
The Mistral Large 4 Launch
Mistral launched the public preview of Mistral Large 4, nicknamed Le Chonk, on October 6, 2026, describing a natively multimodal model trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in the company’s own datacenters in Europe, with training data spanning more than 160 languages including every official European Union language. Mistral claims the model significantly outperforms any open-weight model developed in the US or Europe, and says it will release the weights by the end of October 2026, with the reinforcement learning run behind the preview still in flight.












