Hardware for local AI: sizing memory, bandwidth and power draw

On 17 December 2024, NVIDIA cut the price of its Jetson Orin Nano development board from $499 to $249 and published a software update for units already sold. No component changed, but memory throughput went from 65 to 102 GB/s. According to the manufacturer's measurements, a Llama 3.1 model of 8 billion parameters has since generated 19 tokens per second on it instead of 14. The episode sums up the rule that governs the hardware of a local AI: the size and speed of memory count before computing power.

A language model is a file re-read for every word

A language model, the type of software that powers ChatGPT or Mistral's Le Chat, takes the form of a file containing billions of numbers. They are called parameters, or weights: these are the values adjusted during training, which determine how the model chains words together. A so-called "8B" model contains 8 billion of them.

Each number occupies a space that depends on its precision. In 16 bits, the format in which most models are distributed, a parameter takes two bytes. Eight billion parameters therefore make 16 billion bytes, or about 15 gibibytes (GiB, the binary unit that tools display). The documentation of llama.cpp, a widely used open-source runtime for running models on ordinary machines, gives 14.96 GiB for Llama 3.1 8B in this format.

To produce a token, that is a fragment of a word, the model reads almost all of these parameters. Then it starts again for the next token. A 15 GiB file is thus re-read dozens of times per second while the answer is displayed. Two conditions follow. The file must fit entirely in fast memory, otherwise the machine fetches the rest from disk and speed collapses. And the throughput of this memory sets the pace of the answer.

Capacity: the memory the model must fit in

An ordinary PC has two memories. The main memory, or RAM, serves the central processor. The video memory, or VRAM, is soldered onto the graphics card and serves the graphics processor, the GPU. The second is much faster, but smaller and more expensive. The GeForce RTX 5090, NVIDIA's top consumer card, on sale since 30 January 2025 from ,999, carries 32 GB of GDDR7 memory. This is a ceiling: a heavier model does not fit on a single card.

Other machines use unified memory, a single pool shared by the processor and the GPU. This is the case with Apple's chips. According to Apple's technical specifications, the Mac Studio can be configured with up to 128 GB with the M5 Max chip and up to 512 GB with the M5 Ultra. It is also the principle of NVIDIA's DGX Spark, a 15 cm box weighing 1.2 kg that brings together 128 GB of LPDDR5X memory. NVIDIA presents it as able to run models of up to 200 billion parameters.

The gap becomes decisive as soon as one aims at a medium-sized model. A model of 70 billion parameters compressed to around 4.9 bits per parameter weighs about 43 GB, from the calculation 70 billion × 4.9 ÷ 8. It does not fit in the 32 GB of an RTX 5090. It fits in 128 GB of unified memory, leaving room for the rest.

Bandwidth: what sets the writing speed

Memory bandwidth is the rate at which the chip reads its memory, expressed in gigabytes per second. In its inference optimisation guide, inference being the term for running an already trained model, NVIDIA explains that text generation is memory-bound: most of the time goes into transferring the weights and intermediate data to the compute units, not into computing.

A simple estimate follows: the maximum number of tokens per second is roughly the bandwidth divided by the weight of the model. The Jetson Orin Nano Super makes it possible to check this. Llama 3.1 8B in 4 bits weighs about 4.6 GB and the board reads 102 GB/s, hence a theoretical ceiling of 22 tokens per second. NVIDIA measures 19.1, or 87% of the ceiling. Before the update, at 65 GB/s, the measurement was 14. Speed did not follow throughput exactly (a 37% gain for 57% more bandwidth), but the rule gives the right order of magnitude.

The manufacturers' data sheets give the following throughputs. The Jetson Orin Nano Super reads 102 GB/s and the Jetson AGX Orin, in 32 and 64 GB versions, 204.8 GB/s. The Jetson AGX Thor and the DGX Spark each show 273 GB/s. The Mac Studio reaches 460 or 614 GB/s with the M5 Max depending on configuration, and 1.2 TB/s with the M5 Ultra. The RTX 5090 rises to 1,792 GB/s, but over 32 GB only.

These figures explain a frequent disappointment. A machine advertised with high computing power but slow memory writes slowly. Conversely, a machine with modest computing and fast memory feels responsive in conversation.

Quantisation: fewer bits per parameter

If the weight of the model limits speed, it can be reduced. Quantisation consists of storing each parameter on fewer bits, 8 or 4 instead of 16. The principle is like compressing a photo as a JPEG: one accepts a loss of precision in exchange for a lighter file.

The llama.cpp documentation publishes measurements for Llama 3.1 8B. In 16 bits, the file is 14.96 GiB and generation reaches 29 tokens per second. In the Q8_0 format, which uses on average 8.5 bits per parameter, it is 7.95 GiB for 51 tokens per second. In the Q4_K_M format, at 4.9 bits on average, it falls to 4.58 GiB and rises to 72 tokens per second. The file is divided by 3.3 and the writing speed multiplied by 2.5. Reading the question, for its part, slows a little, from 923 to 822 tokens per second, because this phase depends on computation and decompression adds to it.

The loss of quality exists. It is more noticeable the smaller the model and the stronger the compression. Formats close to 4 bits, such as Q4_K_M, have become the most common compromise in local deployments. Going down to 2 or 3 bits makes it possible to fit a larger model on the same machine, with more errors, which must be measured on one's own tasks before adopting it.

Reading the question and writing the answer: two distinct speeds

A model does not handle words but tokens, pieces of text of a few characters. In French, a common word corresponds to one or two tokens. The work is done in two phases. Prefill reads the whole request, question and attached documents, processing the tokens in parallel: it depends mainly on computing power. Generation then writes the answer one token after another: it depends on bandwidth.

The measurements published by llama.cpp contributors on Apple chips show the difference. On Llama 2 7B in 16 bits, an M4 Max with 40 graphics cores reads 923 tokens per second and writes 32. An M5 Max of the same configuration reads 3,158 and writes 37. Reading is multiplied by 3.4, writing improves by 17%. Between the two chips, bandwidth went from 546 to 614 GB/s, 12% more. Writing follows memory, reading follows computation.

In use, both figures matter. On a short question, only generation is noticeable, and beyond about fifteen tokens per second the text appears faster than it can be read. To summarise a 50-page report, which can be estimated at about 30,000 tokens, it is prefill that makes you wait: at 900 tokens per second, it takes a little over half a minute before the first word, at 3,000 tokens per second about ten seconds.

Context also takes up memory

Context is the amount of text the model takes into account at a given moment: the question, the documents supplied, the history of the exchange. So as not to recompute everything at each token, the model keeps intermediate results for each token already read. This is the key-value cache, or KV cache.

Hugging Face gives the formula and an example. For Llama 2 7B in 16 bits, each token occupies about 0.5 MB of cache, so that 10,000 tokens represent nearly 5 GB, a third of the model's weight. NVIDIA's guide arrives at the same order of magnitude, about 2 GB for 4,096 tokens. More recent models reduce this cost by sharing part of this data between attention heads, a technique called grouped-query attention. With the published characteristics of Llama 3.1 8B (32 layers, 8 key-value heads of dimension 128), the same formula gives about 0.13 MB per token, or 1.3 GB for 10,000 tokens. This last figure is our own calculation.

This cache is multiplied by the number of open conversations. Ten people each working on a 10,000-token document with Llama 3.1 8B tie up about 13 GB of cache, in addition to the model's 4.6 GB. A machine with 8 GB that serves this model to one person will not serve it to a team. Recent runtimes know how to compress this cache, as Hugging Face describes, but the starting point for sizing remains this one.

CPU, GPU and NPU

The CPU, or central processor, is a generalist. It performs few operations at a time and runs a small model slowly. The GPU, designed for graphics rendering, executes thousands of identical operations in parallel, which matches the computation of a neural network. It is the GPU that today runs the great majority of language models.

The NPU, or neural processing unit, is a circuit specialised in neural network operations and designed to consume little. Microsoft requires more than 40 TOPS, or 40 trillion operations per second, for a computer to carry the Copilot+ PC label. Its documentation specifies that many NPUs handle only low-precision integers, such as INT8, and that models must be converted to run on them. The NPU suits light, continuous tasks, such as real-time translation or background blurring in video calls. For a language model, it remains subject to the bandwidth of the rest of the chip.

The TOPS figure should therefore be read with caution. NVIDIA announces 67 TOPS for the Jetson Orin Nano Super, specifying that these are "sparse" TOPS, a computing mode that assumes some of the operations can be skipped. In dense mode, the figure falls to 33. Neither says how fast the board will write an answer.

Three model sizes, several machines

The estimates that follow apply the rule of bandwidth divided by model weight. They are theoretical ceilings calculated by us from the manufacturers' data sheets. Real speed remains below: on the Jetson, it reached 87% of the ceiling.

A model of 8 billion parameters in 4 bits, about 4.6 GB, fits on all the machines cited. The Jetson Orin Nano Super, with 8 GB of memory and a draw of 7 to 25 W, runs it at 19 tokens per second measured, with little margin for context. An RTX 5090 would have a ceiling close to 390 tokens per second. For one person, the embedded board is enough. For several simultaneous users, memory runs short before speed does.

A model of 20 to 35 billion parameters weighs 12 to 21 GB in 4 bits. It fits on an RTX 5090 as on a 64 GB Jetson AGX Orin. On the latter, at 204.8 GB/s, a 20 GB model tops out at around 10 tokens per second; on the RTX 5090, around 90.

A model of 70 billion parameters in 4 bits, about 43 GB, does not fit on an RTX 5090. On a DGX Spark at 273 GB/s, its ceiling is about 6 tokens per second. On a Mac Studio M5 Max at 614 GB/s, about 14. On an M5 Ultra at 1.2 TB/s, about 28. Capacity makes it possible to load the model, bandwidth decides whether it is pleasant to use.

Power draw, heat and noise

All the electricity consumed by a machine ends up as heat. The RTX 5090 has a total graphics power of 575 W, and NVIDIA asks for a power supply of at least 1,000 W for the PC that houses it. That is the order of magnitude of a small electric heater, whose heat is expelled by fans. At the other end, the Jetson Orin Nano Super operates between 7 and 25 W and the Jetson AGX Orin between 15 and 60 W. The DGX Spark has a 240 W power supply for a 140 W chip, and NVIDIA states 35 dB(A) in operation. Apple gives the Mac Studio a maximum continuous power of 480 W.

Over a year, the gap can be quantified. A machine drawing 60 W continuously uses about 525 kWh a year (60 W × 8,760 hours). At 600 W, it is about 5,260 kWh. These calculations assume a constant load, which overstates the real consumption of a machine that is often idle, but they give the factor of ten that separates the two families.

Noise and heat decide where the machine will be installed. A PC fitted with a high-end gaming card is designed for sessions of a few hours in a ventilated room. Stored in a cupboard, it reduces its frequency to protect itself, then shuts down. Jetson modules are designed for robots and industrial equipment, with power ranges set by the manufacturer and prolonged operation. This criterion weighs little on a data sheet and a lot after two years of operation.

Sizing a machine in four calculations

The method comes down to four steps. First calculate the weight of the chosen model or models: number of parameters multiplied by the number of bits per parameter, divided by 8. Then add the context cache, multiplied by the number of simultaneous conversations expected. Compare the total with the available memory, keeping a margin for the system. Finally divide the bandwidth by the weight of the model to obtain the speed ceiling, then check power draw and noise against the place of installation.

Take a team of ten people who want a model of 30 billion parameters in 4 bits to work on documents of 10,000 tokens. The model weighs about 18 GB. The cache, depending on the architecture, is between 2 and 5 GB per user, that is 20 to 50 GB for the team. A machine of 64 to 128 GB is therefore needed, and, as a first approximation, at least 300 GB/s of bandwidth to exceed 15 tokens per second. These figures are estimates; only tests on the target machine confirm them.

This is the grid HOMN applies to size its boxes: start from the models and the number of users, then deduce the memory, the throughput and the thermal envelope. A data sheet that omits memory capacity or bandwidth does not allow this calculation, whatever number of TOPS is displayed.

How much memory do you need to run an LLM locally?

The weight of a model in bytes is roughly the number of parameters multiplied by the number of bits per parameter, divided by 8. A model of 8 billion parameters weighs about 16 GB in 16 bits and 4.6 GB in 4 bits. The context cache must be added, which grows with the length of the documents and the number of simultaneous users.

Why does memory bandwidth matter more than computing power for an LLM?

To write each token, the model re-reads almost all of its parameters in memory. Generation speed is therefore capped by the bandwidth divided by the weight of the model. Computing power matters mainly for reading long documents, called prefill.

What is quantisation of an AI model?

It is the storage of parameters on fewer bits, for example 4 instead of 16. According to the llama.cpp documentation, Llama 3.1 8B thus goes from 14.96 to 4.58 GiB and from 29 to 72 tokens per second on the test machine. Quality drops slightly, and more on small models.

Is a PC with an NPU enough for local AI in a business?

An NPU suits light tasks, such as transcription or a small model used by a single person. It shares the memory and bandwidth of the rest of the chip, which limits the size and speed of language models. To serve a team, memory capacity and throughput take priority over advertised TOPS.