Switching models
With llama-server in router mode, one port offers several chat models and
loads whichever one a request names. mecha builds its model switching on
that, with one rule: the model the router has loaded is the choice.
mecha model use gemma26 loads Gemma, and from then on every run that doesn't
name a model uses Gemma — in the terminal, in web chat, in voice, in Slack, in
the nightly passes. No restart is needed, and there is no "current model"
setting to edit.
There is no second record of the pick on purpose. A stored setting is a second
answer that can disagree with what the server actually loaded. Asking the
router is the same rule as asking a server what it serves (/props) rather
than trusting the config.
Setting it up
- Run llama-server as a router.
scripts/start-router.shstarts it with no-mand a generated--models-preset. Each model gets a section with its own flags (context, projector, sampling, reasoning budget), so switching never mistunes a model.--models-max 1keeps one chat model in memory at a time. The embedding server stays a separate process, so an embedding request can never evict the chat model. - Give each model its own provider entry on the same
base_url. The entry'smodelmust be the preset's section name: in router mode that string is what selects the model. - Mark the default entry
follow_loaded = true. A run that takes the default provider then resolves to whichever sibling entry names the loaded model.
default_provider = "local"
[providers.local]
kind = "local"
base_url = "http://127.0.0.1:8080"
model = "qwen3.6-35b-a3b"
context_window = 262144 # -c 1048576 / -np 4
follow_loaded = true
[providers.gemma26]
kind = "local"
base_url = "http://127.0.0.1:8080"
model = "gemma-4-26b-a4b"
context_window = 32768 # -c 32768 / -np 1
A followed run takes everything from the sibling entry: its
context_window, sampling, prices, fallbacks and retries, not only its
model. So each entry needs the settings of the model it names — the four
numbers in Serving a local model
apply per entry. The keys are in the
configuration reference.
Switching
mecha model # what each router serves, and what is loaded (●)
mecha model use gemma26 # by provider entry
mecha model use gemma-4-26b-a4b # or by the router's model name
Or tap the model chip in the web chat's header. It lists what the router
serves, with a dot beside the loaded model, and tapping another switches to
it. The chip runs the same mecha model use on the server, so everything on
this page applies to it unchanged. While a switch waits, the chip reads
→ model and the menu says what it is waiting for, with switch now
and cancel beside it. A switch that fails turns the chip amber, and its
title says why: the old model loaded back, say, or the load timing out.
The menu greys out a model that runs would not follow — one no provider
entry names, or several do — and one whose preset temperature disagrees with
its entry. It doesn't offer what use would refuse or what would be swapped
straight back.
use loads the model and waits until it is ready. A load runs at disk speed:
33–39 s from cold, about 9 s when the file is already in the page cache. If
the new model fails to come up, the one it replaced is loaded back.
It refuses a model whose preset temperature disagrees with its provider entry.
An entry that sets temperature sends it on every request, so the mismatch
would silently retune the model on every call; the refusal names the entry to fix.
Who follows
One model serves every surface, and a switch from any of them is a switch for all of them.
- Long-lived surfaces check before every turn.
mecha serve(web chat and its voice calls),mecha voice-serve, the Slack connector and the trigger daemon ask the router which model is loaded at the start of each turn or each fire. They rebuild their agent only when the loaded model changed. - A one-shot command reads the router once, when it starts.
mecha run,batch, the nightlylearn/validate/harness ruminatepasses: each is one model from start to finish. - A named provider is a pin, and never moves.
--provider,--model, a trigger'sprovider, an experiment arm. In router mode a pin is also a load: naming a model that is not loaded loads it, and later default runs follow it.mecha model useis the deliberate way to do the same thing. mecha evalnever follows. A scorecard grades the model it names, and two scorecards taken a week apart must not be different models under one condition.- Not yet:
mecha chatand the TUI. They read the router once at start, like a one-shot command, but they don't re-check per turn and a switch doesn't wait for them. Their next turn loads their model back./modelin the TUI is a pin, and on a router a pin is a load: it switches the model for the whole machine, without waiting for the runs still on the old one. Usemecha model useinstead. Inmecha chat,/modelonly shows the active model.
A router in an ambiguous state — mid-swap, or with a model loaded that no
entry names, or that two entries name — does not move a long-lived surface.
It keeps the model it had, because falling back to the default would load
production over your pick. A one-shot command has no earlier model to keep:
it warns and runs on the default. mecha model list flags any entry that
names a model its router doesn't serve.
A switch waits for runs in progress
The router itself only protects single requests: it never evicts a model mid-reply. But a run is many requests, and between two of them the model sits idle. A switch at that moment would succeed, and the run's next request would load the old model straight back. So mecha makes a switch wait until no run is using the model, and a run is answered by one model from start to finish.
Each run on a router takes a hold, a small file in ~/.mecha/holds/, and
drops it when it ends. mecha model use waits until no hold remains, and
tells you what it is waiting for:
waiting for 2 run(s) to finish: web chat, trigger morning-brief (--now stops them, Ctrl-C cancels the switch)
- A run that starts during the wait waits for the switch, then starts on the new model. Web chat says so on the page (Switching the model to X — this turn starts once it is loaded), and a command says so on stderr.
- There is no time limit on the wait, by design: a long delegated task
finishes on the model it started with.
--wait-secsbounds only what comes after the wait. The router won't evict a model that is still answering a request (from a client that takes no hold, say), so the switch asks again each second until it can. It is still waiting then, so switch now and cancel keep working. --nowstops the runs instead. Each one is asked to stop at its next safe point, as Ctrl-C would stop it, and gets 15 seconds before the switch goes ahead. A reply in progress ends early. Given for a switch that is already waiting on the same model,--nowhurries that switch rather than being refused. The chip's switch now works this way, so it can hurry a switch started from a terminal.- Ctrl-C cancels the switch, and the loaded model stays.
mecha model cancel-switchwithdraws a stuck switch: one whosemecha model useis gone, or whose file can't be read. Every run on that router would otherwise wait for it.
Only one switch can be pending on a router at a time. A second
mecha model use is refused and names the first. cancel-switch withdraws
every pending switch, on every router.
One way to wedge it: a run whose agent starts another run through
shell — a chat turn that calls mecha run …, say. With a switch pending,
the child run waits for the switch, the switch waits for the parent, and the
parent waits for its child. Nothing breaks the cycle except --now or
cancel-switch. The waiting message names the run that's holding things up,
which is where to look.
What gets recorded
Every run records the model that answered it. In router mode that record is
exact rather than hopeful, because the model field is what selected the
model. A conversation that continues across a switch records a fresh config
line before its next turn, so each turn names its own model.
That record is what makes it safe for background work to follow your pick. The nightly passes run on whatever is loaded, and readers of the run-quality corpus slice by the model that answered rather than trusting the scheduler to keep models apart. Rule retirement counts only evidence measured on the model currently in use, so one model's results never convict a rule on another's behalf.
Who may switch
Only you: mecha model use at the terminal, or the web chat's model chip,
which runs it behind the same owner check as every other page action.
There is no tool for it. No
model is handed a way to choose the model, because an injection that could
move the machine onto a different model would be choosing who reads every
later turn.
"No tool" isn't the same as "unreachable", though: a run with unconfined
shell could type mecha model use, and that command meets your approval
rules like any other. To close that route, forbid the prefix:
[[rule]]
tool = "shell"
pattern = ["mecha", "model", "use"]
decision = "forbid"
match = ["mecha model use gemma26"]
justification = "Only the owner switches the model."
An incognito chat offers the same
picker. Every model it lists runs on this machine: the chip only knows the
routers a follow_loaded entry names, and those must be local, on a
loopback address. The menu says so in an incognito chat. What keeps
incognito local is its own provider and its per-turn check, not the chip.
See also
mecha modelin the CLI reference.- Serving a local model — slots, and the numbers each entry has to agree with.
- Providers — pins, fallbacks, and why a fallback answers under its own model name.