Skip to content

Choose a provider

octo speaks two wire protocols natively — Anthropic Messages and OpenAI Chat Completions — so “which provider” really means “which of the two protocols does it speak.”

Terminal window
export ANTHROPIC_API_KEY=sk-ant-...
octo --provider anthropic --model claude-sonnet-5 "..."

Local servers expose an OpenAI-compatible API, so they are a custom endpoint with protocol: openai. Two details trip people up:

  • base_url is the OpenAI-compatible root, not the server’s native API. For Ollama that is http://localhost:11434/v1 — not /api, which is Ollama’s own API. octo appends /chat/completions to a URL ending in /v1 (and the full /v1/chat/completions otherwise), so http://localhost:11434/api turns into …/api/v1/chat/completions and Ollama answers 404. vLLM’s default is http://localhost:8000/v1.
  • No API key is needed. Leave api_key out (or blank) and octo sends no Authorization header. The custom vendor is the only one that may run keyless; a named vendor such as openai still insists on a key, so point local servers at custom.
endpoints:
- id: ollama
provider: custom
protocol: openai
base_url: http://localhost:11434/v1
models:
- model: qwen3-coder:30b
vision: false # text-only model: keep images away from it
default: ollama::qwen3-coder:30b

The same via env vars, without touching the config file:

Terminal window
CUSTOM_BASE_URL=http://localhost:11434/v1 \
octo --model qwen3-coder:30b "..."
Terminal window
octo config # interactive wizard
octo config show # print the effective settings + where each came from
octo config path # print the file location

octo config writes your default provider, model, (optionally) base URL, and reasoning settings to ~/.octo/config.yml, so a bare octo works without re-typing --provider/--model every time.

If reaching OpenAI, Anthropic, or other endpoints requires a proxy, octo honors Go’s standard proxy environment variables — HTTPS_PROXY, HTTP_PROXY, NO_PROXY (both http:// and socks5:// addresses work, either case):

Terminal window
HTTPS_PROXY=http://127.0.0.1:7890 octo --provider openai "..."

A shell export is enough when you launch from a terminal. But the desktop app and a background octo serve don’t inherit your shell environment — an export in ~/.zshrc never reaches them. The uniform fix is ~/.octo/serve.env, which every launch mode (CLI, desktop, octo serve) loads at startup:

Terminal window
cat >> ~/.octo/serve.env << 'EOF'
HTTPS_PROXY=http://127.0.0.1:7890
EOF
chmod 600 ~/.octo/serve.env

Restart the app after editing. The proxy applies to all of the process’s outbound traffic (model APIs, search, update checks); to route only some domains through it, use your proxy software’s rules (Clash & co.) or exclude hosts with NO_PROXY=api.moonshot.cn — octo has no per-provider proxy setting.

Reasoning models can deliberate before answering. Two knobs control it, both available as CLI flags and as octo config defaults:

  • --reasoning-effort low|medium|high|xhigh|max — the intensity. OpenAI-protocol backends receive it as reasoning_effort; Anthropic-protocol backends map it to adaptive thinking / an extended-thinking token budget, normalized per model family. Empty (the default) means off.
  • --show-reasoning (default off) — surface the reasoning/thinking trace for the Web UI (octo serve) to display. The terminal never renders the trace either way.

This unifies Anthropic thinking blocks and OpenAI reasoning_content behind one pair of controls.

Next: see the full schema in the config file reference.