Edge AI at work in the field, from the beehive to the workshop

Jersey, summer 2023. On a station set up near the apiaries, a camera stays on standby until an insect the size of a hornet enters the frame. A Raspberry Pi 4, the credit-card-sized computer, analyses the image on the spot and makes the call: European hornet, a local species, or Vespa velutina, the Asian hornet that decimates bee colonies. Only in the second case is an SMS sent to a phone. This device, named VespAI by the University of Exeter, is a good starting point for understanding what edge AI covers, and why part of artificial intelligence is leaving data centres to settle next to the sensors.

What edge AI means

The term edge AI designates a machine learning model that runs where the data is produced: in a camera, in a box fixed at the foot of a machine, in a tractor, in a medical device. The "edge" in question is the edge of the network, as opposed to the centre formed by a cloud provider's servers.

Two phases must be distinguished in the life of a model. Training, which consists of adjusting its parameters on thousands or millions of examples, requires a lot of computation and is most often done on servers. Inference, which consists of applying the already trained model to a new image or a new signal, is much lighter. It is this second step that edge AI brings closer to the field.

The change lies in what travels over the network. In a centralised architecture, the video stream or the series of measurements goes to the server. With local processing, only the conclusion travels, in a few bytes: an Asian hornet detected at 2.12 pm, a lift door whose vibration signature is drifting.

Bandwidth and latency: what the journey costs

The most telling figures come from video surveillance, where the volume of data is massive. In a comparison published on 2 September 2026, the American vendor Lumana estimates that a camera whose stream is analysed in the cloud continuously consumes 1 to 2 megabits per second upstream, against a few kilobits per second when it transmits only alerts and metadata. For a company with more than a hundred sites, the gap amounts, in its view, to gigabits per second of sustained bandwidth. Lumana puts forward a reduction in consumption of up to 70%. The figure comes from a vendor that sells this type of solution, and the order of magnitude obviously depends on the number of cameras and the frequency of events.

The same document puts the delay added by remote analysis at one to five seconds, against a few milliseconds for local processing. For a weekly report, this difference does not matter. For a safety alert in a warehouse where forklifts circulate, it determines whether the alert arrives before or after the incident.

Professional sport shows what real time means. Hawk-Eye, a Sony subsidiary, tracks 29 points of each player's skeleton during the match; according to its website, its semi-automated offside system has been used in more than 900 matches. The computation is done on servers installed at the stadium rather than in the cameras, but the constraint is the same: a decision that arrives after play has resumed is of no use.

There remains the question of continuity. A system that depends on a connection stops when it drops. Lumana stresses that a local architecture, backed by on-site storage, keeps working during an outage. On a building site or in a workshop with poor coverage, this property often weighs more than latency.

What the CNIL says about local processing

For a French organisation, the decisive argument is often legal. In July 2022, the CNIL published its position on so-called "smart" or "augmented" cameras in public spaces. The document does not come down for or against any architecture. It does, however, list safeguards that reduce the risks for the people filmed and that weigh in the assessment of the proportionality of a system.

Among them are lowering the resolution of images, blurring, a "frugal" approach that limits the number of images processed, and local processing of data, which the Commission describes as carried out "in devices physically attached to the cameras". It also cites mechanisms that delete the source images almost immediately or that produce only anonymous information, for example a simple count of passages.

These measures do not exempt anyone from the GDPR, nor from an impact assessment when one is required. They reflect a design principle that engineers know as privacy by design: data that never leaves the site does not have to be protected in transit, nor at a subcontractor. A technical article on industrial safety, published by PowerGen Advancement, describes the same logic on the factory side: blurring, anonymisation, encryption and selective sending before anything leaves the premises.

VespAI, anatomy of an autonomous detector

Let us return to the hornet, because the VespAI project is documented with rare precision. The reference paper, by Thomas O'Shea-Wheller, Peter Kennedy and their colleagues at the University of Exeter, appeared in Communications Biology on 3 April 2024.

The starting point is a sorting problem. Vespa velutina has been established in Europe since 2004, entering through France, and its predation on honey bee colonies reduces their foraging activity and their chances of survival. In the United Kingdom, surveillance relies on reports from beekeepers and the public. Yet most of these reports in fact concern local species: the authors put their average accuracy at 0.06%. Every year, thousands of photos have to be checked by hand.

The device answers this problem with ordinary hardware. A baited station attracts insects under a 16-megapixel camera. A first motion-detection filter avoids running the model continuously. The model itself is a YOLOv5s, a very widespread object-detection architecture, which has about seven million parameters. Parameters are the numerical values that training adjusts; by comparison, large language models have tens or hundreds of billions. The whole runs on a Raspberry Pi 4, powered by a 12,000 mAh battery and, optionally, a 40-watt solar panel. A 4G module makes it possible to send an alert by SMS where there is no Wi-Fi.

The results are measured with two classic indicators. Precision gives the share of alerts that are correct; recall, the share of hornets actually present that were detected. The F1 score combines the two. Under test conditions, the model exceeds 0.99 on this score. In 55 trials carried out at two sites in Jersey in 2023, on more than 5,500 images, average precision remained greater than or equal to 0.99 and average recall above 0.93. In other words, the system is very rarely wrong when it raises the alert, and it lets a few individuals through.

The model does not decide the next steps on its own. The alert arrives with an image, a human confirms it, and the insect can be captured as a living specimen and then tracked to its nest. Destroying the nest remains, according to the team, the only effective way of eliminating a colony.

The structure of this project transposes well beyond beekeeping: an inexpensive sensor, a compact model trained for a single task, an autonomous power supply, and a short message sent only when there is something to report.

In fields and workshops, the decision is made on the machine

Agriculture has adopted embedded AI for a simple reason: a moving sprayer cannot wait for a server. John Deere's See & Spray system, described by Vision Systems Design in August 2025, lines up 36 industrial cameras spaced one metre apart on a 120-foot boom, about 36 metres. Processing units fitted with graphics processors analyse more than 2,000 square feet per second, about 185 square metres, to tell a weed from a crop plant and open the nozzle only over the former.

The same article describes Carbon Robotics' LaserWeeder G2, which destroys weeds with lasers. Each module carries three cameras, two NVIDIA graphics processors and two 240-watt lasers; the machine processes 4.7 million images per hour, with neural networks trained on more than 40 million annotated plants. At this rate, sending the images to a data centre would make no sense: the tractor would have passed the plant before the response came back.

In the factory, the same logic applies to listening to machines. A motor, a pump or an automatic door constantly produce vibrations, heat and current variations. A sensor stuck to the casing is enough to measure them, even on twenty-year-old equipment that no manufacturer had planned to connect. Engineers speak of retrofit: equipping an existing fleet rather than replacing it.

A case study published by TDK SensEI, and picked up in April 2026 by the Edge AI and Vision Alliance, illustrates the approach. A global lift manufacturer, whose name is not disclosed, fitted cabin doors with vibration sensors whose signals are analysed locally by the edgeRX solution. According to the document, unplanned interventions fell from 1.8 days to 1 day per year, for a stated annual saving of 92 million, and the deployment required a week of on-site training. These figures come from the supplier and have not been the subject of a published independent audit. The mechanism, for its part, is clear: a model learns the signature of a door that works well and flags the deviation before the failure.

Applied to a smaller company, the reasoning remains the same. One can imagine a bakery fitting a temperature and current sensor to its oven and its proving chamber, with a model that learns their normal operation and warns on Friday evening rather than Saturday morning. This is a possibility, not a documented case, and it requires no permanent connection.

A hearing implant that classifies sounds without a network

Healthcare pushes the constraint to its maximum. An article in AI News published on 27 November 2025 details the architecture of the Nucleus Nexa system from the Australian company Cochlear. In the external processor, a classifier named SCAN 2, based on a decision tree, sorts the sound environment into five categories: speech, speech in noise, noise, music and silence. The stimulation settings adapt accordingly.

Three constraints explain why this computation takes place on the device. Latency must remain imperceptible between the sound and the signal sent to the auditory nerve. The implant must last more than forty years with minimal consumption. And the data concerns a person's health. Cochlear states that it applies strict de-identification before data feeds its real-world data programme, which covers more than 500,000 patients. Firmware updates go through a short-range radio link, and the implant keeps up to four personalised settings in memory.

Jan Janssen, the company's chief technology officer, mentions the future arrival of deep neural networks for noisy environments. In the meantime, a simple decision tree is enough for a well-delimited task.

Documents have a field too

A large part of economic activity plays out on text: contracts, invoices, letters, minutes. The question of location arises there in the same terms, and it has changed in nature with the arrival of small language models able to run on a desktop computer.

On 16 October 2024, Mistral AI presented Ministral 3B and Ministral 8B, two models of 3 and 8 billion parameters designed explicitly for local use. The French company cites on-device translation, assistants working without internet and local analysis among the target use cases, and states that the smaller of the two outperforms, on most of its tests, its own 7-billion-parameter model released a year earlier. Models of this size can extract dates and amounts from an invoice, summarise minutes or answer a question about internal regulations. They remain clearly below the largest models on long reasoning or very complex documents.

The following uses are projections, not observed deployments. A law firm could compare the clauses of a hundred leases without the documents leaving the firm. A statutory auditor could reconcile invoices, bank statements and entries on a machine located in their own offices, which simplifies the discussion on professional secrecy. A town hall could prepare the preliminary examination of routine requests without sending civil registry files to a foreign provider. In defence, where the very circulation of certain documents is regulated, processing on a workstation isolated from the network is often the only conceivable option. A baker, finally, could have supplier quotes proofread or keep allergen sheets up to date.

In all these cases, the relevant question concerns less the model's ability to read a document than the document's journey during the analysis. A machine installed on the premises provides a verifiable answer: the file goes in, the result comes out, nothing crosses the firewall.

Local by default, cloud by decision

All-local has its limits, and the vendors themselves acknowledge them. Lumana notes that managing a fleet of equipment spread over dozens of sites quickly becomes heavy: configuration, monitoring, maintenance, and above all updating models, which can take weeks or months at fleet scale. It defends a hybrid architecture, where detection is done on site and supervision at the centre. The PowerGen Advancement article proposes the same split: the site handles detection, confidentiality and rapid reaction; the cloud handles multi-site dashboards, trend analysis and model management.

Research describes finer schemes. A survey published on arXiv in July 2025 by Senyao Li and his co-authors catalogues the ways of making a small local model and a large remote model collaborate: routing each request to one or the other according to its difficulty, entrusting preprocessing to the local one and the most demanding part to the remote one, or letting the small model propose a sequence of words that the large one verifies.

For a company, this translates into a simple rule of use. Local processing is the default regime: the data stays on site, the response is fast, the service continues without a network. Sending to a remote model becomes an explicit choice, made task by task, when a more powerful model brings a real gain, and the organisation knows exactly what leaves and why.

This rule presupposes prior discipline. One must have defined the categories of information that never leave, the confidence level below which a local answer must be checked or escalated, the person who approves a transfer, and the way each transfer is logged. These decisions are a matter for the organisation more than for technology, and it is better to take them before deployment than after the first incident.

How to choose a first use case

The projects cited in this article share one trait: a narrow scope. One insect, one weed, one lift door, five categories of sounds. VespAI does not seek to recognise all the insects of Europe, and the Cochlear implant does not transcribe conversations. The precision of the need explains a good part of the quality of the result.

Three questions then make it possible to know whether local processing is called for. Is the data produced allowed to leave the company? What response time is acceptable? What happens if the connection drops? A negative answer to the first, a delay of the order of a second or less for the second, or an activity that cannot be interrupted for the third point towards edge AI.

The last point concerns the final decision. In VespAI, a human confirms the alert before any action. In predictive maintenance, the sensor warns and the technician decides whether to intervene. This split between a machine that proposes and a professional who validates makes the tool acceptable in a workshop, a firm or a public service, and it makes it possible to measure its errors instead of suffering them.

HOMN SYSTEMS designs local AI boxes for French businesses on this principle: the models run on the customer's premises, on their data, and a transfer to an outside service takes place only when the company has decided it and can trace it. HOMN's approach starts from the same observation as the Exeter engineers: a growing share of useful decisions can be made where the data appears.

What is edge AI, or embedded AI?

Edge AI means running an artificial intelligence model directly where the data is produced, in a camera, a box, a machine or a device, rather than on a cloud provider's servers. The model is generally trained elsewhere, then only inference, that is, its application to new data, takes place on site. What travels over the network is then a short conclusion rather than the raw data.

What are the advantages of field AI over the cloud?

Local processing reduces latency to a few milliseconds, whereas remote analysis can add one to five seconds according to a Lumana comparison published in 2026. It greatly limits bandwidth, since only alerts and metadata are transmitted, and it keeps working during a network outage. It also keeps sensitive data on the site, a safeguard that the CNIL cites among its privacy protection measures.

Is embedded AI compatible with the GDPR?

Local processing does not exempt anyone from the GDPR, but it makes it easier to respect its principles of minimisation and data protection by design. In its July 2022 position on augmented cameras, the CNIL mentions processing in devices attached to the cameras, almost immediate deletion of source images and the production of anonymous information as useful safeguards. An impact assessment remains necessary where the processing presents a high risk for individuals.

Do you have to choose between local AI and the cloud?

Most recent architectures combine the two: detection and processing of sensitive data are done on site, while supervision, multi-site dashboards and model updates go through a central service. For language models, research describes schemes in which a small local model handles most requests and passes some to a large remote model only when justified. In every case, each transfer to the outside should be an explicit and traceable decision.