Routing settings
Settings → Routing holds every lever. Nothing here needs changing: an Offline Only AI picks the best model that runs well on this machine, and an Online and Offline AI answers with a frontier online model unless your device is better for the question. AIs set to a specific model are never affected. The levers are for leaning it your way.
Choosing a model on this device
Both Auto modes pick from the models you have downloaded; this sets what the pick leans toward.
| Lean | What it means |
|---|---|
| Prefer fastest | Measured speed on this computer wins; until measured, a model that fits fully on your graphics card. |
| Balanced (default) | The best capability that still runs well. |
| Prefer strongest | The most capable model, even if slower. |
Routing never picks a model that is too large for this machine, whichever you choose. Pausing a model on the Offline Models page also removes it from the Auto pool.
Going online
Only for AIs set to Auto - Online and Offline, and only offered with a plan - without one the section is a single line saying so. Health questions always stay home; Offline Only AIs never go online.
- Send attachments to online models - off by default; see Attachments and images.
- Health questions - which of your downloaded models answers them; see Health questions.
How much goes online
One dial sets where an ordinary question is answered. Questions that need the live web always go online; health questions never do.
| Dial | What it means |
|---|---|
| Frontier-first (default) | Online unless a model here is better for the question. |
| Balanced | Your device answers when it is nearly as good. |
| Local-first | Online only for the live web; hard questions stay home. |
Once you have asked a few questions, the line above the dial says about how many of your recent ones went online ("About 3 in 10 of your recent questions went online"). A freshness choice made in an earlier version carried over to the dial.
Online models by job
When a question goes online, these are the models that take it. Each picker names a recommended model; change it if you would rather use another. Your downloaded models are not listed here - they are picked automatically, by the lean above.
| Job | Takes |
|---|---|
| Everyday questions | Most questions - frontier quality at the lowest price. GPT-5.6 Luna by default. |
| Hard questions | Deep reasoning, tricky code, complex math - one picker for all of them. |
| Needs current information | Searches the live web and cites sources. |
The recommended names come from the online catalog: a model the catalog marks as the default for a job takes that slot, so new models arrive without an app update.
Working on projects
Shown once Projects is installed. Project work routes differently: only models that can drive tools are ever used, and the model never changes mid-session. Offline, the most capable tool-driving model you have downloaded takes it. Online:
- Agent work on projects - the online model that drives tools in project work; only capable ones are listed.
- Planning and helper agents - the online model for the subagents that explore and plan; reasoning-lean, still tool-capable.
- Do simple project side-work on this device - on by default: searching and reading fan-outs run on your device when a capable model runs comfortably here (a fast fit, 8+ tokens a second once measured) - free and private. Otherwise they use the session's online model.
- Keep whole project sessions on this device when possible - off by default: most people arrive trusting frontier models, so a session goes online unless you ask otherwise. On, a session stays fully on your device whenever a capable model runs comfortably here - slower, but free and private. The Model button tells you when a local model would have been as capable, so the choice is informed, never a surprise. Under Local-first, agent work never goes online at all.
- Permissions - how much a project session may do unasked lives in Settings → Agent and on each project's chip; see Permissions and approvals.
What routing has learned
Routing nudges its choices by what you do with answers - Redo on your device, Try this answer online, regenerate - by at most a point, never for health questions. Reset what routing has learned on this computer forgets it. See How routing decides.
Your models, as routing sees them
A collapsed panel on the same page shows the numbers routing ranks on, for every model you have downloaded: how it runs here (Full speed, GPU + RAM with the split in use, Slower, or Too large), its measured speed and load time on this computer, how much it can read at once, its coding, reasoning and math scores, and whether it can drive project work. Speed and load time fill in as you use each model; they are your machine's numbers, never a benchmark's.
Under it, Recent decisions lists the last decisions with their reasons, whether the answer thought first ("thinking on" or "direct answer"), and whether what routing has learned played a part ("adjusted from your feedback").