Product Launch3 min read

Mistral's New Model Takes the Security Work Rivals Refuse

By , Senior AI ConsultantPublished

Mistral's new model scores 82% on a security test where leading closed models score near zero because they refuse, which shows that a model's rules and its owner can take away its usefulness.

Mistral has opened a preview of Mistral Large 4, a model with one trillion parameters that it trained in its own data centers in Europe. Companies can use it now through Mistral's developer site. The files that let anyone run it on their own machines are due at the end of October, along with the license.

A high score is worth little if the model says no

On an independent index that averages ten benchmarks, the new model scores 38 and the leading closed model, Claude Opus 5.5, scores 58. By that measure Mistral is far behind. Yet on the security test that asks a model to reproduce a real software flaw and then patch it, Mistral's model scores 82% and two leading American models score near zero. They are not worse at the job. They decline it, so the test measures the provider's rules as much as the model's skill.

The rules fail here for a simple reason. Defending software usually starts by proving that a flaw is real, which means carrying out the same steps an attacker would, and a safety filter cannot tell the two people apart.

For example, picture the IT contractor for a 40-person accounting firm. A plugin on the firm's website has a reported flaw. Before paying to replace it, she wants to know whether the flaw can be used against this particular site, and whether the patch closes it. If her assistant refuses, she does that by hand, which is a morning of work, or she skips the check. Every small firm, school and clinic with one IT person has some version of this problem.

There is a cost on the other side. Mistral's model also refuses malicious cyber requests more often than any other open model, but Mistral has not explained how it tells research from attack preparation. Once the files are public, anyone who downloads them can retrain the model, including to remove its refusals. An owned model moves the responsibility from the vendor to the company that runs it.

Closed models are a service, and a service can be paused

The June order is the plainest example. Two new models disappeared within days of launch because of a government directive, and the company that made them had no say in who could still use them. The pattern has continued: Google's newest model, Gemini 4 Argon, went to cyber defenders first, and Mythos 5 came back for a small group of defenders only. The strongest security models from the closed labs now reach vetted groups first, and ordinary companies later, if at all.

So the model that is best at a task may be the one you are least able to count on for that task. That is a different problem from price or quality, and it does not show up in any benchmark.

What owning a model means for a business that is not a bank

A model with about one trillion parameters needs roughly a terabyte of memory at 8-bit precision, my own estimate. That is a rack of data-center chips, not an office server. Most readers will never run it themselves.

What they get from a downloadable model is the right to move. If the hosting company drops the service or changes its rules, another host in any country can run the same files, and the work carries on. With a closed model, the vendor is the only place the model lives, so its policy is your policy.

Downloadable models are not new. Mistral's previous Large model was released under the Apache 2.0 license, and Chinese open models have led the field. What this release adds is a very large open model from outside both China and the United States, and the license that comes with it will decide how freely a business can use it. The model is also still being trained, and its final scores may change.

By 2028, I expect most companies that depend on AI for work that cannot stop to keep a downloadable model ready as a standby, the way they keep a second internet connection. The best closed model will stay the everyday choice because it is better. Before you build a process on any model, write down what happens if it refuses the task or disappears for a week, and whether a model you could download would be able to take over.

Share this

STAY INFORMED

Get AI intelligence like this delivered to your inbox.

Free forever · Unsubscribe anytime


You May Also Find Valuable