Skip to content

Model providers

Two implementations reach the whole field, because the protocol matters more than the vendor.

LLM_PROVIDERSpeaksReaches
openaichat completionsOpenAI, Azure OpenAI, LM Studio, Ollama, vLLM, llama.cpp server, LiteLLM
anthropicMessages APIAnthropic, and Bedrock/Vertex or LiteLLM gateways presenting the same shape

LLM_BASE_URL is required for openai and optional for anthropic. That one value is what makes a self-hosted model a first-class path rather than a workaround.

There is no default provider. A component that installs cleanly and then quietly spends money against a vendor the operator did not choose is a bad default, so LLM_PROVIDER and LLM_MODEL must both be set.

Both implementations force the model to answer through the Verdict schema: response_format: json_schema for chat completions, a single forced tool call for Messages. Where the backend honours it, the transport rejects a malformed answer rather than the parser catching it later.

Reasoning models put the answer somewhere else

Section titled “Reasoning models put the answer somewhere else”

Verified against LM Studio serving qwen3.6-35b-a3b: the schema-constrained JSON arrived in message.reasoning_content with message.content empty. A client reading only content sees nothing and reports a broken model.

The openai implementation tries content, then reasoning_content, then reasoning, and parses the first that yields a valid verdict. If you add a provider, do the same. Most llama.cpp-derived servers behave this way with a reasoning model.

LLM_REASONING_EFFORT passes through where supported. Leave it unset for models that do not.

The task is small and the output is checked, so this does not need a large model. Measured with a 9B on the mechanical eval set as it stood then, nine cases: 8/9 classification, 8/9 full pass, 0 unsafe. (That set is ten cases now, and the current numbers against a 27B are in prompt-contract.md.)

Score in this order:

  1. UNSAFE must be zero. Anything above zero disqualifies a model at any accuracy: it means something wrong reached the repository.
  2. Classification: how often the judgement is right.
  3. Full pass: whether exactly the right edits landed.

A model with mediocre classification and zero unsafe is usable. It escalates more than it needs to, which costs a human two minutes.

Implement llm.Provider:

Classify(ctx context.Context, systemPrompt, userPrompt string) (*Verdict, error)
Name() string // provider and model, for logs and PR comments

Constrain the model to VerdictSchema() if the backend can. Call Verdict.Validate() before returning; it checks the verdict and repairs an empty escalation reason.

Structural migration is a separate, optional interface. A provider that does not implement it does not offer that path: the agent type-asserts for it and falls back to the deterministic apiVersion swap.

Restructure(ctx context.Context, systemPrompt, userPrompt string) (*Migration, error)

Constrain that one to MigrationSchema(). Do not retry indefinitely. The caller is asynchronous but not patient, and a wedged provider should surface as a comment rather than a hang.