What is model routing, and why does it cut AI costs?
Model routing sends each AI request to the model best suited to answer it, instead of sending every request to the largest one. Done on the device, it keeps routine questions private and answers most of them with no cloud cost.
Last updated
Draft article: Launch article 19. Needs the CTO's approval as author, with consent to be named.
Model routing is the practice of choosing, for each individual request, which AI model should answer it. Instead of sending every question to the largest and most expensive model available, a router looks at the question first and sends it to the model that can answer it well at the lowest cost.
Why not send everything to the biggest model?
Because most questions do not need it. In a health app, many requests are small and routine: a daily check-in, a meal log, a routine reading. Sending those to a frontier model in the cloud is like calling a consultant to read a thermometer. It works, but it is slow, it costs money every time, and it moves personal data across the internet for no benefit.
Not every health question is the same size. Model routing is how a system treats them differently.
How does Prime Edge AI route a request?
Prime Edge AI is Gateway Global’s AI routing technology. It works in three steps.
- It reads the question on the device. A router on the device reads each question before anything leaves the device.
- Routine tasks stay on the device. Daily check-ins, meal logs and routine readings are answered by Apple’s on-device model, through its Foundation Models framework: offline, and with no cost per call.
- Harder questions go to the cloud, only when needed. Questions that need years of history, clinical literature or a live search are sent to a frontier model in the cloud.
The person sees one conversation and never the handover.
Why does routing cut costs?
A cloud model charges for every request it answers. An on-device model does not: once it is on the phone, answering one more question costs nothing at the margin.
So when most interactions are routine, most interactions cost nothing to answer. Spend then tracks the difficulty of the questions people ask, not the number of people asking them. That is what lets a service like this scale without its costs rising in step with its users.
Why does routing help privacy?
Routing on the device is privacy by architecture. Routine processing never leaves the handset. When a question does need the cloud, it goes straight to the model with no intermediary in between.
How is the connection to the cloud secured?
Sending some requests to the cloud raises an obvious question: how does the cloud know a request is genuine? Prime Edge AI answers it in hardware:
- every device attests itself in hardware;
- no key ships in the app;
- access tokens expire hourly and carry no identity.
Where does Prime Edge AI run today?
Prime Edge AI is built into Averra, the digital health twin we are developing. Android support, using Google’s on-device foundation models, is planned next.
You can read more on the Prime Edge AI page or in the one-page technology brief.
Questions people ask
What does on-device AI mean?
On-device AI runs a model on the phone, tablet or computer itself, rather than on a server. The request and the answer never have to cross the internet.
What is hardware device attestation?
Hardware device attestation lets a device prove, using keys held in its secure hardware, that it is a genuine device running a genuine copy of the app. A server can then trust requests from it without a secret being stored in the app.
Does the person notice when a question goes to the cloud?
No. With Prime Edge AI the person sees one conversation, whichever model answers.