· 4 min read
Intelligence close to the person
Why we run models on the device whenever we can, and what changes in a product when we do.
For most of the last decade, adding intelligence to a product meant adding a server. The model lived in a data centre, the device sent it a request, and an answer came back a moment later. That architecture made sense when the only capable models were enormous. It is no longer the only option.
Phones, tablets and laptops now ship with hardware built for machine learning, and their operating systems expose capable models and frameworks directly to apps. Speech recognition, text understanding, image analysis and a growing range of language tasks can run entirely on the device. When they can, we think they should.
Four reasons
Latency. A round trip to a server is rarely fast and never guaranteed. On-device inference has no network in the loop, so a feature can respond while the person is still looking at it — and keep responding in a lift, on a plane or in a basement.
Privacy. Data that never leaves the device cannot leak from a server, be retained by a vendor or be repurposed later. For anything personal — voice, photos, messages, money — that is not a promise in a policy. It is a property of the architecture.
Cost. Server inference is billed per request, for as long as the product exists. On-device inference runs on hardware the person already owns. That changes which features are worth building, and it means a product doesn’t get more expensive to run as it gets more popular.
Honesty. When a model runs locally, its limits are close at hand and predictable. We find that leads to better design: interfaces that show their working, make mistakes easy to correct, and don’t claim to know more than they do.
When the server is right
Some work genuinely needs larger models or shared data. When it does, we design for the shortest path: send the least data the task needs, keep it no longer than the task takes, and make it clear to the person what is happening and why. The device is still the default. The server is an exception that has to justify itself.
None of this is about being against the cloud. It is about putting intelligence where the person is — in the hand, on the desk — and treating their data as theirs.