Open source AI models in 2026: open-weight versus closed

In May 2023, 174 points separated the best closed AI model from the best open model on the Arena leaderboard, where users blindly compare the answers of two models. In August 2024, the gap was only 7 points. In March 2026, it had climbed back to 49. These figures, taken from Stanford University's AI Index 2026, show that models you can download and run in-house remain close to the best without overtaking them, and that the gap varies with the release schedule of a handful of companies.

Open weights and open source: what the words cover

A language model is a program that reads text and produces text, after training on immense corpora. What it has learned is stored in its parameters, also called weights: billions of numbers adjusted during training. An open-weight model is a model whose numbers are published and downloadable. Anyone can then run it on their own machine, adapt it or study it. A closed model, by contrast, is accessible only through its publisher's interface: you send it a request, it sends back an answer, and its weights remain on the publisher's servers.

The expression "open source" is often used for the first two cases, wrongly in the strict sense. The Open Source Initiative, the reference organisation for free software licences, published in late October 2024 version 1.0 of its definition of open source AI. It requires three elements: sufficiently detailed information about the training data for a skilled person to be able to rebuild a comparable system, the complete code used to train and run the model, and the parameters themselves.

Most well-known models publish the third element, sometimes part of the second, almost never the first. Training data remains secret, for reasons of competition and copyright. A few research projects, such as the Allen Institute for AI's OLMo, also publish their data, but they are not among the best performers. The exact term for DeepSeek, Qwen, Mistral or gpt-oss is therefore "open-weight models". This article speaks of open models for convenience.

What the licences allow

The licence is the contract that sets what the user is entitled to do with the model. The Apache 2.0 and MIT licences, drawn from free software, are the most permissive: they allow commercial use, modification and redistribution, provided the author's notice and the text of the licence are kept. Alibaba's Qwen3.5, OpenAI's gpt-oss and Mistral Large 3 are published under Apache 2.0. DeepSeek V4 is published under the MIT licence.

Meta chose an in-house licence for Llama 4, dated 5 April 2025. It allows commercial use, with three notable conditions. A company whose products exceed 700 million monthly active users must request a licence from Meta, which concerns only a handful of groups worldwide. Any company that distributes a product based on Llama must display the notice "Built with Llama". Finally, any derived model that is distributed must carry a name beginning with "Llama".

These conditions remain acceptable for most companies, but they change the work of the legal department. An Apache 2.0 licence is a standard text, already known to lawyers who have vetted free software. A publisher-specific licence must be read in full, and it can change from one version of the model to the next.

The main open models available in autumn 2026

The largest open models almost all rely on a so-called mixture-of-experts architecture. The model is divided into many sub-networks, and for each word produced, a router calls on only a few of them. Two figures are therefore distinguished: the total number of parameters, which determines the memory required, and the number of active parameters, which determines the volume of computation for each word.

DeepSeek, a Chinese laboratory funded by the investment fund High-Flyer, published in late April 2026 a preview of its V4 family under the MIT licence. According to the card published on Hugging Face, V4-Pro has 1,600 billion parameters, of which 49 billion are active, and V4-Flash 284 billion, of which 13 billion are active. Both accept a context of one million tokens. The context is the amount of text the model can take into account at once, and a token corresponds to a fragment of a word.

Alibaba published Qwen3.5 on 16 February 2026 under Apache 2.0: 397 billion parameters, of which 17 billion are active, and support for 201 languages and dialects. According to an analysis published by Summit School, which we were unable to cross-check with Alibaba, the Qwen3.6 series, released in April, includes models that are easier to host, among them a 35-billion-parameter model that activates only 3 billion. The next flagship model, Qwen3.7-Max, presented in May 2026, is however accessible only through an API, with no downloadable weights.

In Europe, the French company Mistral AI published on 2 December 2025 Mistral Large 3, a model of 675 billion parameters of which 41 billion are active, accompanied by three small models of 3, 8 and 14 billion parameters, all under Apache 2.0. In the United States, OpenAI published in August 2025 gpt-oss, its first open model since GPT-2: the largest version has 117 billion parameters of which 5.1 billion are active, the smallest 21 billion of which 3.6 billion are active.

Meta, which carried openness on the American side with its Llama family, has changed course. On 8 April 2026, the company presented Muse Spark, the first model from its Meta Superintelligence Labs laboratory. It is closed: it is used in Meta's applications or through an API open to selected partners. Meta wrote that it hopes to release future versions as open source, with no timetable. The Llama models already published remain downloadable.

Measuring the gap: independent ranking and publishers' figures

A benchmark is a battery of standardised questions used to score models. Each publisher chooses to highlight those where it obtains its best results, which makes comparisons fragile. The Arena leaderboard proceeds differently: users submit a question, receive two anonymous answers and vote for the better one. The votes feed an Elo-type score, the system used to rank chess players.

The AI Index 2026, published by Stanford's Institute for Human-Centered AI, draws on this leaderboard. In March 2026, the best closed model, Anthropic's Claude Opus 4.6, scored 1,503 points. The best open model, GLM-5 from the Chinese company Zhipu AI, scored 1,454, a gap of 49 points or 3.4%. Six of the top ten models on the leaderboard were closed. The report also notes that the top four companies, Anthropic, xAI, Google and OpenAI, were within 25 points of each other, followed by Alibaba at 1,449 and DeepSeek at 1,424.

The report describes a pendulum movement. Mixtral, WizardLM and then Llama 3.1 405B had brought the gap down to 7 points in August 2024. The arrival of new closed models, such as OpenAI's o1-preview and Google's Gemini 2.5 Pro, widened it again. The gap narrows when open laboratories publish, and reopens when closed laboratories release a new generation.

The publishers give their own estimates, which must be read as such. DeepSeek states that V4-Pro lags GPT-5.4 and Gemini 3.1 Pro by three to six months in reasoning, an estimate that the site WinBuzzer presents as not verified by independent laboratories. For Qwen3.5, the figures published by Alibaba and picked up by The Decoder place it ahead on instruction-following tests, such as IFBench with 76.5, but behind GPT-5.2 in competition mathematics, with 91.3 against 96.7 on AIME26, and in programming, with 83.6 against 87.7 on LiveCodeBench.

Taken together, these figures indicate that a good open model sits a few months behind the best closed models. For most office tasks, the user does not see the difference. On the hardest problems, such as long reasoning or tasks that chain many steps without supervision, it remains measurable.

The memory required and the role of quantisation

To run, a model must load all its parameters into memory, including those of the experts that are inactive at a given moment. The mixture of experts reduces computation, not memory. At the standard precision of 16 bits, each parameter occupies two bytes. The 117 billion parameters of gpt-oss would therefore require about 234 gigabytes, nearly three times the memory of an Nvidia H100 card, which has 80.

Quantisation solves part of the problem. It consists of storing each parameter with fewer bits, 8 or 4 instead of 16. The documentation of Hugging Face's Transformers library describes it as a way of reducing the memory required while seeking to preserve the model's accuracy as far as possible. The most accurate analogy is a compressed photograph: the file is much lighter, and if the compression is well done, the difference shows only on close inspection.

Publishers now ship models that are already quantised. OpenAI distributes gpt-oss with the weights of its experts in the MXFP4 format, on about 4 bits, which allows the large version to fit on a single 80-gigabyte card, and the small one to run with 16 gigabytes of memory. DeepSeek publishes V4 in a mixed format: 4 bits for the experts, 8 bits for most of the other parameters. Even compressed, V4-Pro requires several hundred gigabytes and remains a model for a cluster of servers.

In practice, three tiers emerge. Models of 3 to 20 billion parameters run on a workstation fitted with a recent graphics card. Those of 30 to 120 billion require a dedicated machine with one or two high-capacity cards, which remains within reach of an SME. Beyond a few hundred billion, a full compute server is needed. Compression has a cost: on code, calculation or highly technical texts, too strong a quantisation degrades results. The only reliable method is to test it on one's own documents.

The tasks where they suffice, those where they struggle

Mid-sized open models give good results on bounded, repetitive work: summarising minutes, sorting emails, extracting fields from an invoice, answering questions from a base of internal documents, writing a first draft, translating. For these uses, a well-configured model of 20 to 120 billion parameters produces answers that most users cannot tell apart from those of a closed model. Hosted on site, it also offers a predictable cost and behaviour that does not change without an internal decision.

They are more fragile in three situations. The first concerns long tasks where the model acts alone, chaining dozens of actions: a small probability of error at each step ends up accumulating. The second concerns recent knowledge, since a model knows only what it read before the end of its training, unless it is given up-to-date documents. The third concerns languages and jargons less represented in the training data, which is mostly English and Chinese. Legal or technical French therefore requires specific tests before any deployment.

An open model is a file, not a service. Security updates, monitoring, access logging and rights management remain the responsibility of whoever hosts it. This is the work that closed-model publishers bill for with their API, and which must be planned for when one chooses to self-host.

United States, China, Europe: three ways of publishing

The most advanced American laboratories, Anthropic, Google and OpenAI, keep their best models closed and sell them by API. OpenAI makes a partial exception with gpt-oss, which remains below its flagship models. Meta has fallen into line with its competitors with Muse Spark.

Chinese laboratories, such as DeepSeek, Alibaba or Zhipu AI, publish a lot of open weights under permissive licences. The AI Index 2026 finds that the gap between American and Chinese models has almost closed: in March 2026, the best American model led the best Chinese model, ByteDance's Dola-Seed-2.0 Preview, by 39 points, or 2.7%. According to WinBuzzer, Huawei confirmed that its compute clusters based on its Ascend processors could run DeepSeek V4, which was nevertheless trained on Nvidia chips. Chinese openness is not for all that a doctrine: Alibaba now keeps its most powerful model behind an API, and the best Chinese model on the leaderboard, ByteDance's, is closed.

Europe has few leading laboratories, with Mistral AI in front, and weighs in mainly through regulation. The AI Act's obligations for general-purpose AI models have applied since 2 August 2025, and the Commission has been able to enforce them since 2 August 2026. Models published under a free licence, except those classed as systemic risk, are exempt from part of the technical documentation obligations, but must still publish a summary of their training data and a copyright compliance policy. The code of practice that accompanies these rules, published on 10 July 2025, had 26 signatories in mid-August 2025, including OpenAI, Google, Microsoft, Amazon and Anthropic. Meta and the Chinese companies were not among them.

For a French buyer, the consequence is practical: before deploying a model, you need to know which one is actually running, under what licence, with what documentation, and who its publisher is.

Changing model without redoing your tools

Today's ranking will not be next year's. In twelve months, DeepSeek has published a new generation, Alibaba has presented three and reserved the latest for its API, and Meta has launched its first large closed model. A company that links its software directly to a specific model is betting on a state of the market that will have changed before the end of the year.

The answer is technical. Applications do not address the model directly, but an intermediate layer that chooses the model according to the task. An internal test set, made up of the company's real documents and questions, checks that a new model does at least as well as the old one before replacing it. Changing model then becomes a maintenance operation, planned and verifiable.

This discipline protects against several risks at once: a publisher that stops publishing its weights, a licence that tightens, a price that rises, a model that is withdrawn. An article on the German site Digital Chiefs, devoted to Meta's turn, recommends the same approach to IT departments: several model families rather than a single supplier, including at least one open model for regulated or offline uses, an inventory of applications and the models they use, and exit clauses in contracts.

What the state of open models changes for a business

In autumn 2026, the best open models sit a few points from the best closed models according to the most closely followed independent measure. Their licences range from Apache 2.0, which imposes almost no conditions, to in-house texts that need to be reviewed. Their availability depends on the strategy of their authors, which can change from one version to the next. For most everyday uses, they are enough, and they make it possible to process the company's documents without sending them to a third party.

The criterion that counts over time is the ability to replace one model with another without touching the rest. This is HOMN's architectural choice: the model is an interchangeable part, tested on the customer's documents before each replacement, and the infrastructure around it stays in place from one generation to the next.

What is the difference between an open source model and an open-weight model?

An open-weight model publishes its parameters, which makes it possible to download it and run it on your own machine. According to the Open Source Initiative's definition, an open source model must also publish the training code and detailed information about its training data. Most well-known models, such as DeepSeek, Qwen or Mistral, are open-weight.

What is the best open source AI model in 2026?

According to Stanford's AI Index 2026, the best open model on the Arena leaderboard in March 2026 was Zhipu AI's GLM-5, with 1,454 points. DeepSeek V4, Qwen3.5 and Mistral Large 3 are among the largest open models published since. The right choice depends on the hardware available, the licence and tests on your own tasks.

Are open source models as good as ChatGPT or Claude?

They are close. In March 2026, the best open model trailed the best closed model on the Arena leaderboard by 49 points, or 3.4%. For everyday office tasks, the difference is barely visible, but it remains measurable on long reasoning and complex autonomous tasks.

What hardware do you need to run an open source LLM locally?

A model of 3 to 20 billion parameters runs on a workstation fitted with a recent graphics card, and gpt-oss-20b runs with 16 gigabytes of memory. A model of about 120 billion parameters quantised to 4 bits, such as gpt-oss-120b, fits on an 80-gigabyte card. Models of several hundred billion parameters require a full server.