For two years, using AI meant sending your words to someone else's computer. You typed a prompt, it traveled to a data center, a giant model answered, and the reply came back. In 2025 and 2026, a quieter shift arrived: a small but real slice of that work now happens on the device in your hand or on your desk, with nothing leaving the machine. Apple, Google, Microsoft, and Qualcomm all shipped consumer or developer surfaces built on models that run locally, and every phone and laptop launch now advertises a "neural processing unit," or NPU, as if it were the new megapixel count.
The pitch is genuinely appealing: instant responses, no network needed, no per-use cloud bill, and your data staying put. But the marketing blurs three things that are worth holding apart. There is a difference between what has actually shipped to people and what was demoed on a keynote stage. There is a difference between a benchmark or spec sheet and real-world quality on a device running off a battery. And there is a difference between a vendor's claim and independently verified behavior. Keep those three seams visible and the on-device AI story becomes much clearer — and more useful — than either the hype or the backlash suggests.
What "on-device" actually means
The models that run locally are called small language models, or SLMs, and "small" is doing real work in that name. The cloud models behind chat assistants are measured in hundreds of billions of parameters. The on-device models are measured in single-digit billions. Apple's on-device foundation model is about 3 billion parameters [source: Apple Machine Learning Research, 2025]. Microsoft's Phi Silica, which runs on Windows Copilot+ PCs, is a derivative of the Phi-3.5-mini family [source: Windows Experience Blog, 2024]. Microsoft's research prototype Phi-3-mini — the clearest public proof that this is possible at all — is 3.8 billion parameters, and when compressed it runs natively on an iPhone [source: Phi-3 Technical Report, arXiv, 2024].
Two engineering techniques make that shrinkage possible, and it helps to keep them distinct. Quantization stores a model's numbers at lower precision — dropping from 16-bit down to 4-bit or even 2-bit values — which slashes memory and speeds up math for some loss of accuracy. Apple's model uses 2 bits per weight, trained with quantization-aware training so the model learns to tolerate the coarseness [source: Apple Machine Learning Research, 2025]. Phi-3-mini, quantized to 4 bits, occupies about 1.8 GB and generates more than 12 tokens per second on an iPhone's A16 chip, fully offline [source: Phi-3 Technical Report, arXiv, 2024]. Distillation is the other technique: instead of compressing one model, you train a smaller "student" to imitate a larger "teacher," producing a genuinely new, smaller model. The SLM families shipping today are built with both.
That 1.8 GB figure is the whole story in miniature. A model that once needed a rack of server GPUs now fits in the memory budget of a phone — not by magic, but by trading away precision and breadth for a footprint that fits.
What actually shipped in 2025
This is the layer that separates on-device AI from vaporware, so it is worth being concrete about what real people and developers can actually touch today.
Apple shipped the Foundation Models framework in iOS 26 in the fall of 2025, giving developers direct access to the on-device 3-billion-parameter model. Inference is free — there is no per-call cloud bill — it works offline, and the data stays on the device [source: Apple Machine Learning Research, 2025]. Google shipped its ML Kit GenAI APIs, powered by Gemini Nano, to Android developers in May 2025. They run through the system's AICore, and input, inference, and output are all processed locally with no internet requirement and no per-call cost [source: Android Developers Blog, 2025]. Microsoft put Phi Silica on the NPUs of Copilot+ PCs, where it drives features such as Click to Do and on-device rewrite and summarize inside Word and Outlook, with a developer API following in January 2025 [source: Windows Experience Blog, 2024]. And Qualcomm announced the Snapdragon 8 Elite Gen 5 in September 2025 on a 3-nanometer process, with a Hexagon NPU it says is about 37% faster at AI than the previous generation [source: GSMArena, 2025].
Notice what these shipped products have in common: they are task-scoped, not open-ended chatbots. Google's on-device APIs are explicitly four narrow jobs — summarize, proofread, rewrite, and describe an image [source: Android Developers Blog, 2025]. Apple's model is offered as a building block for app features, not as a general assistant. This is the shipped reality, and it is deliberately narrower than the keynote impression of "AI everywhere, locally."
Benchmarks and spec sheets versus a device on battery
Here the marketing and the measured experience pull apart, and the gap is the most important thing a reader can understand.
Start with the NPU. Every Copilot+ PC must have an NPU rated at 40 or more TOPS — trillions of operations per second [source: Windows Experience Blog, 2024]. That is a peak-throughput spec, and it sounds enormous. But peak operations are not delivered words. On the same class of hardware, Phi Silica produces its first token in about 230 milliseconds and then generates up to about 20 tokens per second [source: Windows Experience Blog, 2024]. On a Pixel 9 Pro, Gemini Nano reads input at roughly 510 tokens per second but writes its answer at about 11 tokens per second [source: Android Developers Blog, 2025]. Those output speeds are perfectly usable for a two-sentence summary and visibly slow for anything long. A 40-plus-TOPS spec and an 11-tokens-per-second reality are both true; they are just answering different questions.
The same gap shows up in quality scores. Phi-3-mini posts 69% on MMLU, a standard knowledge benchmark, which its authors frame as rivaling much larger models [source: Phi-3 Technical Report, arXiv, 2024]. That is a real achievement, but a benchmark average is not open-domain reliability. The most honest statement in this entire field comes from Apple, which says plainly that its on-device model "is not designed to be a chatbot for general world knowledge" [source: Apple Machine Learning Research, 2025]. Read that again: the company shipping the model tells you not to treat it as a know-it-all. That is the benchmark-versus-reality gap stated by the vendor itself.
Battery and heat sit underneath all of this. The reason on-device inference is viable at all is that NPUs are far more power-efficient than CPUs for this work — Microsoft measured Phi Silica using about 56% less power than the same job on the CPU [source: Windows Experience Blog, 2024]. That efficiency is exactly why a phone can run a model without draining in minutes. But efficient is not free: sustained text generation still consumes energy and produces heat, which is one practical reason vendors cap what these models will attempt.
Those caps are explicit. Google states that its on-device summarizer works best on input under about 4,000 tokens, and that proofreading and rewriting are meant for short text under 256 tokens [source: Android Developers Blog, 2025]. Those are not arbitrary — they are the honest edges of where a small model on a battery delivers good results.
Claims versus what is verified
The third seam is the softest, because it is where verifiable engineering shades into marketing adjectives. "Your data stays on device" is the headline privacy claim, and for genuinely local inference it follows from the architecture — if the model runs on the phone and the network is off, the words are not leaving [source: Android Developers Blog, 2025]. That is the strongest of the claims because you can reason about it from how the system is built.
Others deserve more caution. Qualcomm describes the Snapdragon 8 Elite Gen 5 as enabling "agentic AI" with continuous on-device learning while "user data stays on device" [source: GSMArena, 2025]. "Rivals GPT-3.5" is a benchmark-anchored claim about a specific test set, not a promise about your particular question [source: Phi-3 Technical Report, arXiv, 2024]. The reliable core in each case is the part you can inspect — where the model runs, how it was quantized, how many tokens per second it measurably produces — not the adjective wrapped around it. A good habit: trust the architecture, discount the marketing.
The hybrid reality: local-first, not local-only
The cleanest way to misunderstand this shift is to imagine the cloud going away. It is not. The dominant real architecture is tiered: a small local model handles common, latency-sensitive, or private tasks, and harder requests fall back to a large model in the cloud.
Apple's shipped design is the clearest example. The on-device 3-billion-parameter model handles what it can, and heavier requests route to a larger server model — a Parallel-Track Mixture-of-Experts architecture — running on Apple silicon in what the company calls Private Cloud Compute [source: Apple Machine Learning Research, 2025]. Google and Microsoft follow the same shape: scoped tasks run locally through Gemini Nano or Phi Silica, while the heavy lifting still goes to cloud Gemini or Copilot. So when a 2026 phone advertises "on-device AI," the honest translation is usually local-first with cloud fallback, not everything on the phone. That is not a bait-and-switch; it is the sensible engineering answer to the fact that a 3-billion-parameter model and a frontier cloud model are good at different things.
What small local models can and cannot do
Strip away the layers and a practical picture remains. What on-device SLMs do well is a bounded, valuable set of jobs: summarizing a document, proofreading and rewriting text, describing an image for accessibility, extracting structure, routing a request, and calling a tool — all with low latency, offline, and without sending your data anywhere. Those are real features shipping in real products right now, and for that category the local model is often the better choice than a round trip to the cloud.
What they cannot reliably do is equally important: they are not general world-knowledge chatbots — Apple says so outright [source: Apple Machine Learning Research, 2025] — they struggle with long-context reasoning, and they do not match frontier cloud models on open-domain accuracy. Asking a 3-billion-parameter phone model to be a substitute for a hundreds-of-billions-parameter cloud assistant is asking the wrong thing of it.
What to watch
The useful posture toward on-device AI is neither the "everything is local now" hype nor the "it is just a gimmick" backlash. Something real shipped in 2025, and it is genuinely useful within its scope.
Three questions cut through the noise. First, is a capability shipping in a released OS, SDK, or chip, or is it a demo and a roadmap — because the gap between the two is where disappointment lives. Second, does a number describe a peak spec (TOPS, a benchmark average) or a delivered experience (tokens per second on a real device, within the vendor's own task caps) — because a 40-TOPS NPU and an 11-token-per-second answer are both honest and describe different things. And third, is a claim about architecture you can inspect — where the model runs, how it is quantized — or a marketing adjective like "agentic" or "rivals GPT-4." Judge on-device AI by what it verifiably does within its limits, and the technology in your pocket turns out to be smaller than the ads suggest and more useful than the skeptics allow.