Local AI: definition, how it works and concrete examples
A local AI is a language model that computes its answers on a machine you control, without sending your texts to a provider. This guide explains what that model contains, what running it in-house means and what can reasonably be expected of it.
One command, one download, then nothing leaves
The documentation of Ollama, free software available for macOS, Windows and Linux, fits the starting point into one line: a command typed in a terminal, which downloads a model and opens a conversation. For the model used as an example, Gemma 4 E2B, the file weighs about 7.2 GB and the documentation recommends 8 GB of graphics memory, or of unified memory on a Mac. Once this file is on the disk, every answer is computed by the machine's processors. You can unplug the network cable and the conversation carries on.
This is the simplest definition of a local AI: an artificial intelligence model run on hardware that you own or control, whether a laptop, a server stored in a technical rack or a dedicated box. Its opposite is a remote service. ChatGPT, Gemini or Le Chat receive your text on their servers, process it there and send you back the result.
To understand what happens in each case, we need to open the box: what a language model is, what the downloaded file contains, and what exactly the machine does when it produces an answer.
A language model predicts what comes next in a text
A large language model, often abbreviated LLM, is a program trained on a task that looks modest: guessing the piece of text that comes next. The keyboard on your phone, which suggests the next word above the keys, does the same thing on a tiny scale. The LLM does it with such finesse that, by chaining predictions, it produces paragraphs, summaries, code or translations.
The model does not handle words but tokens, fragments of text that can correspond to a whole word, a syllable or a punctuation mark. OpenAI's documentation gives an order of magnitude for English: a token represents about four characters, or three quarters of a word. This unit is used to measure everything, from the maximum length of an exchange, called the context window, to the speed of generation.
The architecture that made these models possible is called the Transformer. It was described in 2017 in a research paper titled "Attention Is All You Need". Its central idea, the attention mechanism, allows the model to weigh, for each token, the importance of all the other tokens in the text. In the sentence "the contract that the client signed yesterday expires in March", it is this mechanism that links "expires" to "contract" despite the words in between.
We must keep in mind what this mechanism implies. The model produces the most plausible continuation given what it saw during training. Plausible does not mean accurate, and this nuance explains most of its limits.
The weights: what the file contains
What you download when you install a model is its weights, also called parameters. They are numbers, billions of numbers, adjusted one by one during training so that the predictions become good. They can be compared to the sliders of an immense mixing desk: no slider taken in isolation means much, but their overall position determines the sound.
The size of a model is counted in parameters. Mistral 7B, published by the French company Mistral AI on 27 September 2023, has 7.3 billion. It was released under the Apache 2.0 licence, which allows anyone to download it, use it and install it wherever they wish, including for commercial use. This is called an open-weight model: the company publishes the file itself, and not only remote access.
Two phases follow one another in the life of a model. Training builds the weights from immense volumes of text; it mobilises considerable computing resources and remains the business of a few laboratories. Inference then uses these frozen weights to answer a question; it is this, and this alone, that is run locally. The difference is like that between writing a dictionary and consulting it: the first takes years, the second a few seconds, and nobody needs to rewrite the dictionary to look up a word.
Running a model: engine, format and quantisation
To run these weights, you need an inference engine, the program that loads the file into memory and chains the calculations. The most widespread in the local world is called llama.cpp. Written in C and C++ with no external dependency, it aims to run models with a minimal installation on a wide range of hardware: Apple chips, ordinary x86 processors, NVIDIA graphics cards via CUDA, or other manufacturers via Vulkan. More accessible applications, such as Ollama or LM Studio, can read the same file format.
This format is called GGUF. According to the documentation of Hugging Face, the main platform for sharing models, a GGUF file brings together in a single binary the weights and a standardised set of metadata: the architecture, the token vocabulary, the settings needed for loading. You copy it, load it, and it works, a little like a PDF that embeds its own fonts.
The factor that decides almost everything is memory. In its original form, each parameter is stored on 16 bits, that is two bytes. A model of 7.3 billion parameters then occupies nearly 15 GB, more than the memory of most desktop computers. Quantisation solves this problem by rounding the weights to fewer bits, as you reduce the number of colours in a photo without the eye seeing the difference at first glance. llama.cpp supports quantisations from 1.5 to 8 bits. The widely used Q4_K type amounts to 4.5 bits per weight according to the Hugging Face documentation: the same model then comes down to around 4 GB.
This compression has a price, a slight loss of precision that grows as the number of bits drops. And memory is not only used to store the weights. Ollama's documentation points out that a longer context window requires more memory, since the machine keeps track of all the text being processed. If graphics memory runs short, the software can fall back on ordinary RAM, with slower answers. The choice of hardware follows from three parameters: the size of the model, its level of quantisation and the length of the texts to be processed.
Why this became realistic within a few years
Not long ago, running a competent model on a single machine was a laboratory demonstration. Three measurements published in Stanford University's AI Index 2025 show what has changed.
The first concerns size. In 2022, the smallest model scoring above 60% on MMLU, a multiple-choice test covering many disciplines, was Google's PaLM, with 540 billion parameters. In 2024, Microsoft's Phi-3-mini crossed the same threshold with 3.8 billion parameters, a model 142 times smaller. Once quantised, a model of this size fits in the memory of a laptop.
The second concerns the gap between closed models and open-weight models. In early January 2024, on the Chatbot Arena leaderboard, where internet users blindly compare the answers of two models, the best closed model led the best open model by 8.04%. In February 2025, the gap was only 1.70%.
The third concerns cost. Again according to Stanford, the inference cost of a system at the level of GPT-3.5 fell by more than 280 times between November 2022 and October 2024. Over the same period, hardware cost fell by about 30% a year and its energy efficiency improved by 40% a year.
These figures describe a trend, not an equality. The largest closed models keep the advantage on the hardest tasks, and a ranking based on internet users' preferences does not measure everything. The frontier of what can be done without a data centre has nonetheless moved markedly.
Local or ChatGPT: the question is where the text travels
From the user's side, the two experiences look alike: an input box, an answer that appears word after word. The difference lies in the journey. With an online service, the question, the attached documents and the answer pass through the provider's servers. With a local AI, they do not leave the machine.
This journey has concrete consequences. OpenAI states that conversations from free and Plus ChatGPT accounts may be used to improve its models unless the user has turned off the option, whereas data from its business offerings and its API is not used by default. The CNIL, in its questions and answers on generative AI published in July 2024, observes that with API use, control of the system lies "almost exclusively" in the provider's hands. It recommends on-site deployment when personal or sensitive data is processed, in order to limit the risks of extraction by a third party.
ANSSI, the French national cybersecurity agency, goes in the same direction in its security recommendations for generative AI systems, published on 29 April 2024. Its recommendation R34 advises against generative AI tools accessible over the internet for professional use involving sensitive data: the organisation does not control the service and therefore cannot guarantee the confidentiality of what it sends there.
Local has other properties. It works without a connection. Its cost does not depend on the number of questions asked, since it is hardware bought rather than a meter. It does not change overnight because a provider has replaced or withdrawn a model. In return, the organisation manages the machine, the updates and the backups itself, and it has no access to the largest closed models, which are offered only online. Many organisations combine the two: local for what is confidential or repetitive, a remote service for one-off needs involving no sensitive data.
What you can concretely do with a local AI
The uses that work best are those where the model works on a text you give it, rather than drawing on its memory. Summarising thirty pages of meeting minutes. Rephrasing a letter in a more neutral register. Extracting the supplier, date and amount from an invoice to file them in a table. Translating a technical manual. Sorting incoming emails by subject before routing them to the right department.
The most requested case in business is querying one's own documents: internal procedures, contracts, agreements, the support knowledge base. The technique used, called RAG for retrieval-augmented generation, first finds the relevant passages, then asks the model to answer from them while citing its sources. We devote a detailed article to it.
More technical uses follow: help with writing code, transcription of meetings by a speech recognition model also run on site, analysis of error logs. These uses are spreading fast. According to Eurostat, 20% of European enterprises with at least ten employees used an artificial intelligence technology in 2025, against 13.5% a year earlier.
In all these cases, the machine produces a first draft or a sorting, and a person validates. It is in this configuration that the ratio between time saved and risk taken is most favourable.
The limits, without disguising them
A language model can state something false with confidence. This phenomenon, called hallucination, follows directly from how it works: it produces a plausible text, and a plausible text may contain an invented date, a non-existent legal reference or a badly copied figure. Running locally changes nothing on this point. The remedy is to supply it with the sources rather than rely on its memory, and to have anything that commits the organisation proofread.
Its knowledge stops at the end of its training. A model released a year ago knows nothing of the latest reform, the latest scale, the latest version of a piece of software. It does not consult the internet by itself: a local AI isolated from the network knows only what it has learned and what it is given to read.
A small model remains a small model. The progress measured by Stanford is real, but a model of a few billion parameters makes mistakes more often than a giant model on multi-step reasoning, a calculation or a highly specialised question. Speed depends on the hardware: on a computer without a suitable graphics processor, a long answer may take a while.
Finally, a local installation is a computer system like any other. It has to be administered, patched, and access to it controlled. ANSSI devotes part of its recommendations to the isolation of AI systems, the logging of processing and rights management. Local does not mean secure by default; it means that security depends on the organisation that operates it.
How to get started
The simplest is to try on a recent computer with an application such as Ollama or LM Studio. Choose a small open-weight model, of a few billion parameters, quantised to 4 bits. Test it on real tasks but without confidential data: summarising a public document, rephrasing a text, extracting information from a blank form. Note the quality, the speed and the errors.
This first step serves to define the needs before any spending: the uses that come up every week, the volume of documents to process, the number of people who will use the system at the same time. These elements determine the size of the model, hence the memory required, hence the hardware.
For team use, the test laptop is no longer enough. You need a machine that stays on, accessible on the internal network, with management of accounts, rights and logs. This is the role of a server or a dedicated box.
At HOMN, this last configuration is what we build: the model, the documents and the logs stay on a machine installed at the user's premises, and recourse to an outside service remains an explicit choice. The approach holds whatever tool is chosen: know where the model runs, what it contains and where the data goes.
What is a local AI?
A local AI is an artificial intelligence model, most often a language model, that runs on hardware the user controls: a computer, server or box installed on their premises. The questions and documents processed are not sent to a provider's servers. The model is a file downloaded once, then used without a connection.
Do you need an internet connection to use a local AI?
No, not to operate. The internet is used to download the model and software updates, but the computation of answers is done entirely on the machine. A local AI can therefore run on an isolated network.
Is a local AI as capable as ChatGPT?
Not on everything: the largest closed models remain ahead on the hardest reasoning tasks. The gap has nonetheless narrowed sharply, since according to Stanford's AI Index 2025 it went from 8.04% to 1.70% on the Chatbot Arena leaderboard between January 2024 and February 2025. For summarising, rephrasing, extracting or querying documents, a good open-weight model is enough in most cases.
What hardware do you need to run an LLM locally?
It all depends on the size of the model and its quantisation. A 7-billion-parameter model quantised to 4 bits occupies about 4 GB, and Ollama's documentation recommends 8 GB of graphics or unified memory for its example model. For use shared by a team, you need a dedicated machine with more memory.