Skip to content
Documentation

Choosing models

Pick a local model or your own cloud key, then assign a model to each feature: reviewing, Trace It, chat, voice answers, Learn Me and general tasks.

You pick the AI backend and model for each kind of job, so an expensive model never ends up doing cheap work. Your choice also decides where your code goes.

#Pick a backend

BackendCode goes toPlan
Local model (Ollama, LM Studio, llama.cpp or vLLM)Nowhere. Requests go to localhost only, and it works offline.Every plan
Your own cloud key (Anthropic, OpenAI, Google Gemini)Straight from your machine to that provider, never through DiffGuardian. Billed by the provider.Pro
OpenRouterStraight to openrouter.ai with your key. One key reaches roughly 400 models.Pro

Cloud keys are stored in your operating system's credential store and never sent to DiffGuardian's servers.

#Local engines

Turn on the engine's row in Settings → AI Settings, confirm the address (loopback only, port 1024 or higher), and select Check now. If you started the server with an API key, enter the same key in the row.

EngineDefault address
Ollamahttp://localhost:11434
LM Studiohttp://localhost:1234
llama.cpp (llama-server, also llama-swap)http://127.0.0.1:8080
vLLMhttp://127.0.0.1:8000

#OpenRouter

The model field searches OpenRouter's live catalogue and shows each model's context window, price per million tokens, and whether it supports structured outputs. Reviews require structured outputs, so don't assign a model without them to AI Reviewing.

#Assign a model to each feature

Settings → Model routing has six rows. The suggestions are a starting point.

FeatureWhat it doesSuggested model
AI ReviewingThe review pass and suggested commentsYour strongest model. Must support structured outputs.
Trace ItGroups a diff into features and walks through each oneA strong model, similar to AI Reviewing
ChatAsk panel answers and Analyze Project's whole-repository answersA strong model for Analyze Project. A mid-size one is enough for quick Ask questions.
Voice answersThe spoken answer in the Voice DockA fast model, because latency is what you notice
Learn MeLearns your preferences from what you typeA small, cheap model
General AI tasksPolish, dictation clean-up, Explain, summaries, release notes, diagramsA small, fast model

You are never locked in. Reassign a feature, or fall back to local, at any time.

#Compare your local models

Settings → Local model benchmark runs the same sample diff through every reachable local model and shows time-to-first-token, tokens per second and the actual output side by side. The fastest model is rarely the one that reviews best, so assign on evidence.

#Which to use when

SituationChoose
Code you can't send anywhere, or no internetLocal model
A large PR that needs the strongest reasoningYour own cloud key
Your employer already has a provider agreementYour own key under that agreement
Comparing models without signing up with five vendorsOne OpenRouter key