Models and context
mecha runs the same agent loop over any model: a hosted one through Anthropic's API, or an open-weight one on your own machine through an OpenAI-compatible server. The loop never learns which provider is behind it. Whichever you use, the conversation has to fit the model's context window.
| Page | Covers |
|---|---|
| Providers | Configuring a provider, prompt caching, and which failures are retried. |
| Serving a local model | Running llama-server: slots, what -c divides, and the numbers that must agree with your config. |
| Switching models | One router, several models: loading one is the choice, every surface follows it, and a switch waits for the runs in progress. |
| Compaction | What happens when a conversation outgrows the window, and what survives it. |