Choosing models
Pick a local model or your own cloud key, then assign a model to each feature: reviewing, Trace It, chat, voice answers, Learn Me and general tasks.
You pick the AI backend and model for each kind of job, so an expensive model never ends up doing cheap work. Your choice also decides where your code goes.
#Pick a backend
| Backend | Code goes to | Plan |
|---|---|---|
| Local model (Ollama, LM Studio, llama.cpp or vLLM) | Nowhere. Requests go to localhost only, and it works offline. | Every plan |
| Your own cloud key (Anthropic, OpenAI, Google Gemini) | Straight from your machine to that provider, never through DiffGuardian. Billed by the provider. | Pro |
| OpenRouter | Straight to openrouter.ai with your key. One key reaches roughly 400 models. | Pro |
Cloud keys are stored in your operating system's credential store and never sent to DiffGuardian's servers.
#Local engines
Turn on the engine's row in Settings → AI Settings, confirm the address (loopback only, port 1024 or higher), and select Check now. If you started the server with an API key, enter the same key in the row.
| Engine | Default address |
|---|---|
| Ollama | http://localhost:11434 |
| LM Studio | http://localhost:1234 |
llama.cpp (llama-server, also llama-swap) | http://127.0.0.1:8080 |
| vLLM | http://127.0.0.1:8000 |
#OpenRouter
The model field searches OpenRouter's live catalogue and shows each model's context window, price per million tokens, and whether it supports structured outputs. Reviews require structured outputs, so don't assign a model without them to AI Reviewing.
#Assign a model to each feature
Settings → Model routing has six rows. The suggestions are a starting point.
| Feature | What it does | Suggested model |
|---|---|---|
| AI Reviewing | The review pass and suggested comments | Your strongest model. Must support structured outputs. |
| Trace It | Groups a diff into features and walks through each one | A strong model, similar to AI Reviewing |
| Chat | Ask panel answers and Analyze Project's whole-repository answers | A strong model for Analyze Project. A mid-size one is enough for quick Ask questions. |
| Voice answers | The spoken answer in the Voice Dock | A fast model, because latency is what you notice |
| Learn Me | Learns your preferences from what you type | A small, cheap model |
| General AI tasks | Polish, dictation clean-up, Explain, summaries, release notes, diagrams | A small, fast model |
You are never locked in. Reassign a feature, or fall back to local, at any time.
#Compare your local models
Settings → Local model benchmark runs the same sample diff through every reachable local model and shows time-to-first-token, tokens per second and the actual output side by side. The fastest model is rarely the one that reviews best, so assign on evidence.
#Which to use when
| Situation | Choose |
|---|---|
| Code you can't send anywhere, or no internet | Local model |
| A large PR that needs the strongest reasoning | Your own cloud key |
| Your employer already has a provider agreement | Your own key under that agreement |
| Comparing models without signing up with five vendors | One OpenRouter key |