Local AI or cloud: where company data goes, and how to choose

On 2 May 2023, Samsung Electronics banned its employees from using ChatGPT and other generative AI assistants on the company's computers, phones and tablets, as well as on its internal network. A few weeks earlier, engineers had pasted internal source code into ChatGPT. Nobody had hacked anything. The code had gone to a third party by the intended route, the one followed by every request sent to an online assistant.

The path of a request sent to an online assistant

A language model is a program trained on very large quantities of text, which computes the most probable sequence of words from what it is given. Getting it to answer is called inference. It requires graphics processors, or GPUs, chips designed to chain billions of multiplications, and these chips are expensive.

A company has two ways of accessing it. It can rent a provider's capacity over the internet: this is the cloud. The text of the request, the attachments and the conversation history then leave its network, are processed in a data centre belonging to the provider, and the answer comes back. It can also install the model on a machine it owns, on its own premises. This is called local AI, or "on-site" deployment. The request does not leave the building.

For the person typing the question, the two look alike: a window, a text, an answer in a few seconds. The difference lies in what happens next to the text. In the cloud, its fate depends on the provider's terms and conditions, its nationality, its subcontractors and the courts before which it can be summoned. Locally, it depends on the rules the company has set itself.

What providers keep, and who can demand it

Providers keep a record of exchanges, at least for a time, to detect abuse, to bill or to improve their service. Depending on the plan subscribed, these exchanges may also be used to train subsequent models. Contracts intended for businesses generally exclude this use, but the retention period and the list of subcontractors appear in the contract's annexes, not on the home page.

The lawsuit between the New York Times and OpenAI before the federal court for the Southern District of New York showed how far this retention can go. In early December 2025, Judge Ona Wang refused to reverse her decision requiring OpenAI to hand over to press publishers 20 million de-identified ChatGPT conversations. According to eWeek's account, she noted that this sample represented only a small fraction of the tens of billions of individual users' conversations kept by the company. The data was de-identified and placed under a protective order that reserves reading access to lawyers.

The users whose conversations are in this sample were not parties to the case. Their exchanges became exhibits in a copyright dispute for one reason only: they were stored at the provider. For a company, the practical conclusion fits in one sentence. Anything a third party keeps can one day be requested by a judge or an authority, in a case that has nothing to do with it.

The Cloud Act follows the provider's nationality, not the server's address

The Cloud Act is a US law adopted in March 2018. Its most discussed provision, codified at section 2713 of title 18 of the federal code, obliges providers of electronic communication services and remote storage to preserve or disclose their customers' data that is in their possession, custody or control. The text specifies that this obligation applies whether the data is located in the United States or abroad.

A server installed in Paris by a company subject to US law therefore remains within the scope of the law. On 10 June 2025, before the Senate's committee of inquiry on public procurement, Senator Dany Wattebled asked the director of public and legal affairs of Microsoft France, Anton Carniaux, whether he could guarantee under oath that the data of French citizens entrusted to Microsoft would never be transmitted on the order of the US government without the agreement of the French authorities. Anton Carniaux replied that he could not guarantee it, adding that it had never happened.

The US Department of Justice presents the text as an investigative tool for serious crimes, supplemented by bilateral agreements with partner countries. This description is accurate, and nothing indicates that the data of a French SME is a common target. But the legal exposure exists, it concerns all customers of a provider subject to US law, and one of the largest of them has stated under oath that it could not rule it out. A provider that escapes US jurisdiction is not bound by this obligation. A model installed on the company's premises is not either, since no third party holds the data.

Transfers to the United States rest on a contested agreement

The GDPR, the European regulation on personal data protection, allows such data to be sent outside the Union only if the destination country offers protection deemed equivalent, or if contractual safeguards are put in place. For the United States, this equivalence has rested since July 2023 on the data protection framework between the European Union and the United States, called the Data Privacy Framework.

Its two predecessors were annulled by the Court of Justice of the European Union, Safe Harbor in October 2015 and Privacy Shield in July 2020, each time following proceedings brought by the Austrian activist Maximilian Schrems. The third holds for now. On 3 September 2025, the General Court of the European Union dismissed the action brought by French member of parliament Philippe Latombe and confirmed the validity of the framework. On 31 October 2025, Philippe Latombe lodged an appeal before the Court of Justice. To our knowledge, at the date of publication of this article, the Court has not yet ruled.

A company that entrusts personal data to an assistant hosted by a US provider is therefore relying on a mechanism whose two previous versions were invalidated. If the third falls, contracts will have to be reviewed, transfers reassessed and some uses sometimes suspended urgently. Processing carried out on a machine installed in the company involves no transfer and is not affected.

What the CNIL recommends for generative AI

The CNIL published on 18 July 2024 a series of questions and answers on the use of generative AI systems. The document prescribes no technology, but it is clear on one point. When an organisation intends to supply the system with personal data of customers or employees, or sensitive or strategic documentation, it considers it "generally more appropriate and more secure" to favour on-site deployment, which limits the risks of data extraction by a third party.

The CNIL also acknowledges the cost of this option. Installing and operating a system on site represents a significant expense, and recourse to a remote infrastructure will often be simpler. It adds that nothing prevents small businesses or local authorities from sharing a pooled infrastructure, provided data transfers are framed and secured.

For an online service, it asks to check whether the provider's terms and conditions allow it to reuse the data entered. If so, the organisation must decide case by case whether certain categories of information, for example those covered by trade secrecy, should be prohibited. The provider must be bound by a processing contract that sets its responsibilities and the authorised access. Under the GDPR, a processor is the company that processes personal data on behalf of another.

Hosting the model on site does not exempt anyone from any GDPR obligation. A purpose must still be defined, data limited to what is necessary, a retention period set and the machine secured. Local does, however, reduce the number of contracts to maintain and parties to audit.

The AI Act timeline after the July 2026 omnibus

The European regulation on artificial intelligence, called the AI Act, entered into force on 1 August 2024. It classifies systems according to the risk they present and imposes obligations on two categories of actors: providers, who design the systems, and deployers, who use them in their activity. An SME that makes an assistant available to its teams is a deployer.

Prohibited practices, such as social scoring of individuals, have been prohibited since 2 February 2025. The obligations specific to general-purpose AI models, which include large language models, have applied since 2 August 2025, and the European Commission has had its supervisory powers over these models since 2 August 2026. The transparency rules, which require for example informing a person that they are interacting with an AI, have applied since the same date.

Regulation (EU) 2026/1744, known as the digital omnibus on AI, published in the Official Journal in July 2026, postponed two deadlines. High-risk systems under Annex III, which covers recruitment, credit assessment and education among others, will have to comply by 2 December 2027. Those under Annex I, built into already regulated products such as machinery or medical devices, by 2 August 2028.

The text does not distinguish cloud from local. It does, however, ask the company to know which system it uses, for what purpose and with what data, and to be able to document it. This documentation is easier to maintain when the version of the model does not change without the knowledge of those who use it, a point on which the two options differ markedly.

Cost: paying per request or buying a capacity

Online services bill by usage, generally by the number of tokens processed. A token is a fragment of a word, the unit the model reads and produces. As long as a few employees are experimenting, the bill remains modest. When the tool reads hundreds of documents every day or answers every customer request, it follows the volume, at a rate set by the provider and which it can revise.

A machine installed on site reverses the calculation. The largest part of the spending is committed at purchase, then each additional request costs only electricity and wear on the hardware. For occasional use, renting remains cheaper. For daily and sustained use, buying eventually wins out. The tipping point depends on each company's volumes and is calculated on its own figures.

Two measurements published by Stanford University's AI Index weigh on this calculation. According to the 2025 edition, the inference cost of a system at the level of GPT-3.5 fell by more than 280 times between November 2022 and October 2024. The same edition indicated that open-weight models, which can be downloaded and run on one's own machine, had reduced their lag behind closed models from 8% to 1.7% on certain tests, within a year.

The 2026 edition qualifies this second figure. In March 2026, on the Arena leaderboard, built from votes by users who blindly compare answers, the best closed model led the best open model by 3.4%, against 0.5% in August 2024. The lag remains small, but it does not narrow in a straight line: it widens with each release of a new closed model, then tightens.

Outages and latency: the price of distance

In the night of 19 to 20 October 2025, the Northern Virginia region of Amazon Web Services suffered a cascading outage. According to the account published by AWS, a synchronisation defect between two automatic processes of the system that manages the addresses of its DynamoDB database produced an empty DNS record. DNS is the directory that translates the name of a service into a network address. Without this entry, applications could no longer find DynamoDB. This first phase lasted from 11:48 pm to 2:40 am Pacific time, or 2 hours and 52 minutes.

The aftermath lasted much longer. Several internal AWS services that rely on DynamoDB were hit in cascade, and customers of the EC2 compute service suffered errors, slowness and machine launch failures until 1:50 pm on 20 October, about fourteen hours after the start of the incident. Applications built on top of these services were interrupted in turn.

An infrastructure installed in the company also fails. A disk gives out, a power supply burns out, an update goes badly. The difference lies in who diagnoses and who decides the order of restoration. It also lies in the internet connection, which becomes an additional point of failure as soon as the teams' everyday tool sits with a third party.

Latency, that is the delay between sending a request and the start of the answer, follows the same logic. In a conversation, one more second goes unnoticed. For business software that queries the model hundreds of times an hour, the round trips add up, and their duration varies with the provider's load. On the internal network, the delay is more stable, provided the machine is sized for the number of users.

When the provider changes model

On 7 August 2025, OpenAI launched GPT-5 and made it the default model of ChatGPT, removing the previous models, including GPT-4o, from the interface. Protests were immediate, and in the following days Sam Altman announced that Plus subscribers would be able to choose GPT-4o again. The episode mainly concerned individuals attached to the tone of the old model. In a company, the same mechanism has more concrete consequences.

A process developed with a specific version of a model, for example the automatic extraction of amounts and dates from invoices, may behave differently with the next version. Tests then have to be redone, the instructions given to the model rewritten, sometimes the documentation supplied to a customer or auditor revised. Prices and terms of use evolve in the same way, at the provider's initiative. The more integrated the tool, the more it costs to leave. Economists call this an exit cost, and computer scientists speak of vendor lock-in.

An open model installed on site does not change until the company replaces it. The same instructions, applied to the same model with the same settings, give comparable results from one month to the next, which makes it possible to document a use and defend it. Replacement remains possible when a better model comes out, at the time chosen and after tests on the company's documents.

The cases where the cloud remains preferable

The most powerful models are accessible only through an API, that is, an interface that lets software query the service remotely, and some would not fit on a desktop machine anyway. For very long reasoning, programming on large projects or tasks that chain dozens of steps without supervision, they keep a measurable lead. If the need calls for exactly this level, the cloud is often the only reasonable option.

The cloud also suits short projects, such as a demonstration or a pilot of a few weeks before investing. It suits very irregular workloads, when a machine sized for the peak would sit idle the rest of the month. And it suits texts with no sensitivity: rephrasing an announcement that is already public or looking for an article title exposes nothing.

The choice is therefore made after sorting the data. Public information can go to a provider. Personal data, contracts, unpublished figures and technical know-how justify either processing on site, or a provider whose contract has been read, whose jurisdiction has been checked and whose security has been verified.

Process on site by default, send out by decision

A simple rule makes it possible to combine the two approaches. Requests are processed by default on a company machine. Sending to an external service becomes an explicit decision, made for a specific task, recorded, and which a manager can refuse. Most everyday requests, such as summarising a file, drafting a reply or finding a clause in a batch of contracts, are within reach of mid-sized models that run on desktop hardware. The others go outside, knowingly.

For each AI tool already in service, three checks give a fairly clear picture of the situation: where the text is processed, who keeps it and for how long, and what happens if the service stops for a day. The answers are in the contracts, the subcontracting annexes and the providers' status pages.

This is the principle on which HOMN designs its boxes: a model installed on the premises, which works on the company's documents without transmitting them, and which connects to an outside service only when its owner has decided it.

What is the difference between a local AI and a cloud AI?

A cloud AI runs on a provider's servers: each request leaves the company, is processed in an external data centre, then the answer comes back. A local AI runs on a machine installed on the company's premises, and requests do not leave its network.

Does the Cloud Act apply to data stored in France?

Yes, when the provider is subject to US law. The law targets data in its possession, custody or control, wherever it is stored. On 10 June 2025, the director of public and legal affairs of Microsoft France stated under oath before the Senate that he could not guarantee that such data would never be handed to the US authorities.

Can you use ChatGPT in a business while complying with the GDPR?

It is possible with a processing contract, safeguards for transfers outside the Union and an internal rule on the data allowed. The CNIL asks to check whether the provider can reuse the data entered, and generally considers on-site deployment preferable for personal or sensitive data.

When is a local AI more cost-effective than the cloud?

When usage is daily and sustained. The cloud is billed by the volume of text processed, whereas a local machine is paid for at purchase and then costs only electricity and maintenance. For one-off or very irregular use, renting is generally cheaper.