What business AI really costs: subscription, API or own machine
A fifty-person company that wants to equip its teams with an AI assistant receives three proposals that cannot be compared: a price per employee per month, a price per million tokens, and the price of a machine. This article puts the three bills into the same unit, using the public price lists in force on 1 October 2026.
Three ways to pay for the same answer
The first model is the per-user subscription, also called a per-seat licence. The company pays a fixed amount for each person authorised to use the tool, whether they use it ten times a day or once a month. This is the model of the team plans of the major chat assistants and of the copilots built into office suites.
The second model is pay-as-you-go through an API. An API, or application programming interface, lets software send a request to a provider's model and receive the answer. The provider then bills the quantity of text processed, measured in tokens. There is no per-person subscription: the bill follows the volume.
The third model is to own the machine. The company buys a computer able to run an open-weight language model, that is, a model whose parameters are published and can be installed in-house. The cost becomes an investment depreciated over several years, to which electricity and maintenance are added. The number of requests no longer changes the bill; it only changes the load on the machine.
These three models do not answer the same question. The subscription buys a ready-to-use service. The API buys raw computing capacity that has to be integrated into one's own tools. The machine buys a fixed capacity, available without volume limits but bounded by its power. Comparing them requires bringing each offer back to a monthly cost for a given usage.
The token, the unit of account for language models
A token is a fragment of text, often a piece of a word, that the model reads and produces one at a time. Anthropic's pricing documentation gives an order of magnitude: in English, a token represents about four characters, or roughly three quarters of a word. It specifies that the exact count varies with the language and the type of content. A text in French, a table of figures or source code do not produce the same number of tokens for the same length.
Providers distinguish two categories. Input tokens correspond to everything the model receives: the question, the instructions, the conversation history, the attached documents. Output tokens correspond to what it writes. The latter cost considerably more. On the price lists recorded on 1 October 2026, the output token is worth five times the input token, both for Anthropic's current models and for OpenAI's gpt-6 generation.
The token is not a stable unit from one provider to another, nor even from one model generation to the next. Anthropic states that its models from Claude 4.7 onwards use a new text tokeniser that produces about 30% more tokens for the same text. At the same listed price, the same task can therefore cost more after a model change. Comparing prices per million tokens between two providers is like comparing prices per kilo without knowing whether the two sets of scales are calibrated the same way.
What the price lists show on 1 October 2026
On subscriptions, the Claude pricing page, consulted on 1 October 2026, lists a Team plan at $25 per seat per month billed monthly, or $20 with an annual commitment, for the standard seat. A so-called premium seat costs
25 per month, or 00 on an annual basis. The Enterprise plan combines a base of $20 per seat per month, billed annually, with usage billed at API rates. This last point deserves attention: at this level, the subscription no longer caps spending.
Microsoft launched Microsoft 365 Copilot Business on 1 December 2025, at $21 per user per month, intended for organisations of up to 300 users and sold with an annual commitment. This licence is added to an existing Microsoft 365 Business subscription and does not replace it. The plan aimed at large enterprises is listed at $30 per user per month.
On the API side, OpenAI's pricing page, consulted on the same date, lists models for its most recent generation ranging from gpt-6-luna, at $0.10 per million input tokens and $0.50 for output, to gpt-6-astra, at 0 and $50, with gpt-6.1-sol in between at $2 and 0. At Anthropic, Claude Haiku 4.5 costs for input and $5 for output, Claude Sonnet 5.5 $2 and 0, Claude Opus 5.5 $4 and $20. Within the gpt-6 generation alone, the price gap reaches a factor of one hundred.
Two mechanisms reduce these prices. Batch processing, which accepts a delay in the response, halves the bill at both providers. Caching of repeated instructions bills re-reads at a fraction of the input price: 10% of the base rate for most Anthropic models. All these amounts are in dollars, excluding tax, and the providers state that they may change them.
How the bill follows usage
Let us take a deliberately simple assumption. Fifty employees each send thirty requests per working day, over twenty-one days a month, making 31,500 requests. Each request carries 3,000 input tokens, covering an instruction, a little history and a short document, and produces 600 tokens of response. The month totals 94.5 million input tokens and 18.9 million output tokens.
With Claude Sonnet 5.5, this volume costs 89 for input and 89 for output, that is $378 per month. With Claude Opus 5.5, twice that: $756. With gpt-6-astra, ,890. With gpt-6-luna, 8.90. Against this, fifty Claude Team seats billed monthly come to ,250, and fifty Copilot Business licences to ,050, before the Microsoft 365 licence they presuppose.
This calculation shows that, for moderate and regular usage, the API of a mid-range model can cost less than subscriptions. It does not tell the whole story. The API alone provides no interface, no account management and no connection to the company's documents: these must be built or bought. And the assumption of 3,000 tokens per request is fragile.
Anthropic's documentation estimates that an average web page represents about 2,500 tokens and that a 500-kilobyte research paper in PDF represents about 125,000. If an assistant re-reads this document at each of the ten exchanges of a conversation, it consumes 1.25 million input tokens, or $2.50 on Claude Sonnet 5.5 for a single analysis. With caching, the same conversation falls to around $0.54, provided the exchanges follow one another before the cache expires. Agents, the programs that chain several calls to the model to complete a task, multiply this effect: each step sends back the accumulated context. On an API, the bill depends less on the number of users than on the quantity of text the model is made to re-read.
The costs that do not appear on the price list
The first hidden item is integration. Connecting a model to mailboxes, shared folders, management software or the customer database takes development time, testing and ongoing maintenance. This cost appears on no pricing page and depends entirely on each company's information system.
The second item relates to where processing takes place. OpenAI charges a 10% surcharge on its regional processing endpoints, those that guarantee data stays within a geographic zone, for models released from 5 March 2026. Anthropic applies a multiplier of 1.1 when the customer requires processing in the United States only. Choosing where one's data is processed has an explicit price.
The third item is model turnover. OpenAI's model deprecation page shows that GPT-4.5 preview, whose shutdown was announced on 14 April 2025, stopped working on 14 July 2025, three months later. Each replacement requires checking that the instructions, output formats and quality of responses still hold. This requalification work comes back with every generation.
The fourth item is price instability, in both directions. Anthropic had announced for Claude Sonnet 5 a launch price of $2 and 0 valid until 31 August 2026, followed by an increase to $3 and 5 on 1 September. The increase was ultimately cancelled and the launch price became the standard price. The outcome was favourable to customers, but the decision was not theirs. A budget built on an API price list remains subject to a decision by the provider.
The fifth item concerns exit. Retrieving one's data, conversation histories and configurations in order to change provider has long been charged for. The European Data Act, applicable since 12 September 2025, allows switching charges limited to actual costs until 12 January 2027, and then prohibits them, data egress charges included. It does not resolve technical dependency: instructions written and tuned for a specific model do not always carry over to another without adjustments.
Owning the machine: purchase price, depreciation and electricity
To reason about a real case, let us take the NVIDIA DGX Spark, a compact computer designed to run language models, with 128 GB of unified memory. It is a public reference example, not a recommendation. According to a European reseller that tracks its prices, it launched in October 2025 at $3,999, then NVIDIA raised its price by $700 in February 2026 citing the global memory shortage. In September 2026, according to prices recorded at a European reseller (pi3g), the versions sold in Europe started at around €5,038 excluding tax. Hardware also sees price rises that the buyer does not control, but suffers them once, at purchase.
In France, computer hardware is commonly depreciated over three years on a straight-line basis, that is, one third of its value each year. Equipment under €500 excluding tax can be expensed directly. Depreciated over thirty-six months, €5,038 represents about €140 per month.
Electricity is calculated from power. NVIDIA's documentation indicates an external power supply of 240 watts for this machine. If it ran permanently at that maximum, which overstates reality since consumption drops at idle, it would consume 2,102 kilowatt-hours per year. At the professional blue tariff, base option, set at €0.1624 per kilowatt-hour excluding VAT, excise duty included, since 1 August 2026 following the deliberation of the Commission de régulation de l'énergie of 15 July 2026, the annual bill peaks at about €341, or €28 per month.
The ceiling on the monthly cost of owning this machine is therefore around €168, before maintenance, updates, backup and the time of the person who looks after it. In return, the number of requests does not change the bill. The limit lies elsewhere: a given machine serves a limited number of simultaneous requests, lower the larger the model, and the open-weight models it can run do not give the same results as the largest commercial models on every task. Only a trial on the company's real tasks can settle the matter.
A prudent way to compare
The comparison is made over thirty-six months, the depreciation period of the hardware, and by keeping currencies separate: APIs and subscriptions are billed in dollars, electricity and hardware bought in Europe in euros. Converting at a daily rate would give illusory precision over three years.
The first step is to measure. A pilot of a few weeks on an API, with a small group, provides the real number of tokens per request, which the providers' consoles display. This measurement is worth more than any assumption, including the one in this article.
The second step is to separate the usages. Occasional users cost little on an API and a lot on a subscription. Heavy users, or automated processing that re-reads large volumes of documents, reverse the proportion. An owned machine becomes attractive when volume is high and steady, or when part of the data must not leave the company, whatever the price of the alternative.
The third step is to put a figure on the exit from the start. What would a change of provider, model or price list cost in eighteen months? With the figures from our example, a machine whose cost peaks at around €168 per month compares with an API that would cost $378 with a mid-range model or less than $20 with the smallest model. The calculation names no universal winner. It depends on volume, the nature of the data and the quality required, three parameters that only the company knows.
What the price list does not say
The three economic models are not mutually exclusive. Many organisations combine a subscription for office usage, an API for a few one-off tasks that require a very large model, and owned capacity for regular volume and for data that must not leave the premises. The important thing is to know, for each flow, what it costs and who sets its price.
This is the grid we apply at HOMN: a fixed cost known in advance for local processing, and recourse to the cloud decided flow by flow, with its cost measured. The figures in this article are dated 1 October 2026; the price pages cited in the sources make it possible to update them.
What is a token and how many words does it represent?
A token is a fragment of text that the model reads or writes, often a piece of a word. Anthropic estimates that in English a token corresponds to about four characters, or three quarters of a word. The count varies with the language, the content and the model's tokeniser.
Is the API cheaper than a per-user subscription?
For moderate usage, often yes: in our assumption of 50 employees, the API of a mid-range model costs $378 per month against more than ,000 in subscriptions. But the API provides neither an interface nor integration, and its bill climbs quickly with long documents and agents.
How much electricity does a local AI machine use?
It depends on its power. A machine with a 240-watt power supply would consume at most 2,102 kWh per year running permanently at full load, or about €341 at the professional blue base tariff in force since 1 August 2026.
Over how many years should an AI server be depreciated?
In France, computer hardware is commonly depreciated over three years on a straight-line basis. Equipment under €500 excluding tax can be expensed. Your chartered accountant will confirm the appropriate period for your situation.