Building with AI

On-device AI

4 min readbeginnerUpdated 28 Sept 2026
1 · In one line

On-device AI runs the model on your own phone, laptop or browser, so your data does not have to go to a server.

1 · What it is

On-device AI means the model runs on the machine in front of you: your phone, laptop or web browser. Normally, when you use an AI feature, your words are sent over the internet to a big model on a server, and the answer comes back. With on-device AI, the prompt and the answer stay on your device, and no server call is made.

Picture typing a quick reply in a chat app and asking it to fix your grammar. On an Android phone, the app can pass that message to Gemini Nano, a Google model that runs on the phone. Your message never leaves the phone, and it works on a plane with no Wi-Fi. Chrome does something similar for websites: a page can ask the browser to translate text using a model stored inside the browser.

The catch is size. Phones often have limited memory and computing power, so on-device models are made smaller. One common trick is quantization. It stores each number in the model with fewer digits, a bit like rounding. That shrinks the model and speeds it up, but it can cost some accuracy. Apple’s on-device model, for example, has about 3 billion parameters (the adjustable numbers a model learns) and is compressed to an average of 3.7 bits per weight. Apple also has a larger model that runs on its servers, and Chrome’s docs suggest a cloud fallback too.

For engineers: an app can use a shared model that the phone or browser already provides, instead of packing its own. On Android, AICore runs Gemini Nano on the device’s own hardware. Chrome manages its built-in model’s download, and the network is needed only for that first download. You can also run your own open-weight model (one whose files are published for anyone to download), such as Gemma. That is possible with frameworks like Google’s LiteRT, which can target the CPU, GPU or NPU (a neural processing unit, a chip built for AI work).

2 · Why it exists

Sending every request to a cloud model has three costs.

Data leaves youYour words or photos travel to someone else's computer to be processed.
Needs a connectionWith no network, a cloud feature simply stops working.
Every call costsThe app maker pays for server time on each request.
3 · How it works

Follow one request that never leaves the phone.

After the one-time download, the prompt and the answer stay on the device.
  1. 1 · downloadThe model is downloaded to the device once, and after that the network is not needed to run it.
  2. 2 · askAn app sends its prompt to a system service, such as AICore on Android or the built-in AI APIs in Chrome.
  3. 3 · runThe model runs on the device's own chips, and the prompt is never sent to a server.
  4. 4 · answerThe result goes straight back to the app, and how fast it arrives depends on the device.

On-device AI moves the model to the data, not the data to the model.

4 · Where it's used
WhoWhat they askWhat it works with
A messaging app“Can you fix the grammar in this short reply?”The message being typed, proofread on the phone
A website“Can you translate this review into English?”Page text, translated inside the browser
A notes app“Can you summarise this meeting?”The note, summarised by a local model
A voice app“What did I just say?”Audio, transcribed on the device
5 · What it solves, and what it doesn't
solves
  • Private data can stay on the device instead of going to the cloud.
  • Features can keep working with no network connection.
  • It avoids the server cost of each request.
  • It removes the wait for a round trip over the network.
doesn't solve
  • Speed depends on the device's hardware, and some devices are not supported yet.
  • Devices have limited memory, so models are made smaller, which can cost some accuracy.
  • The model itself must first be downloaded, which needs space and a connection.
  • Some apps still keep a larger server model or a cloud fallback.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docsGemini Nano, Android Developers (Google) · read 28 Sept 2026
  2. docsBuilt-in AI, Chrome for Developers (Google) · read 28 Sept 2026
  3. docsGet started with built-in AI, Chrome for Developers (Google) · read 28 Sept 2026
  4. officialIntroducing Apple's On-Device and Server Foundation Models, Apple Machine Learning Research · read 28 Sept 2026
  5. docsLiteRT, Google AI Edge · read 28 Sept 2026
  6. docsModel optimization, Google AI Edge · read 28 Sept 2026