Choose a provider
octo speaks two wire protocols natively — Anthropic Messages and OpenAI Chat Completions — so “which provider” really means “which of the two protocols does it speak.”
export ANTHROPIC_API_KEY=sk-ant-...octo --provider anthropic --model claude-sonnet-5 "..."export OPENAI_API_KEY=sk-...octo --provider openai --model gpt-4o-mini "..."DeepSeek, Bailian, and other OpenAI-compatible APIs use the same protocol.
Self-hosted and third-party endpoints (vLLM, OpenRouter, a proxy) use the custom vendor — the
only one that takes a custom base URL. Its wire protocol (openai | anthropic) is chosen per
config entry, so set it up once with octo config (choose Custom → pick the protocol → enter
base URL + model):
CUSTOM_BASE_URL=https://api.deepseek.com/anthropic \CUSTOM_API_KEY=sk-... \ octo --model deepseek-chat "..."Local models (Ollama, vLLM)
Section titled “Local models (Ollama, vLLM)”Local servers expose an OpenAI-compatible API, so they are a custom endpoint with
protocol: openai. Two details trip people up:
base_urlis the OpenAI-compatible root, not the server’s native API. For Ollama that ishttp://localhost:11434/v1— not/api, which is Ollama’s own API. octo appends/chat/completionsto a URL ending in/v1(and the full/v1/chat/completionsotherwise), sohttp://localhost:11434/apiturns into…/api/v1/chat/completionsand Ollama answers 404. vLLM’s default ishttp://localhost:8000/v1.- No API key is needed. Leave
api_keyout (or blank) and octo sends noAuthorizationheader. Thecustomvendor is the only one that may run keyless; a named vendor such asopenaistill insists on a key, so point local servers atcustom.
endpoints: - id: ollama provider: custom protocol: openai base_url: http://localhost:11434/v1 models: - model: qwen3-coder:30b vision: false # text-only model: keep images away from itdefault: ollama::qwen3-coder:30bThe same via env vars, without touching the config file:
CUSTOM_BASE_URL=http://localhost:11434/v1 \ octo --model qwen3-coder:30b "..."Save it as your default
Section titled “Save it as your default”octo config # interactive wizardocto config show # print the effective settings + where each came fromocto config path # print the file locationocto config writes your default provider, model, (optionally) base URL, and reasoning settings to
~/.octo/config.yml, so a bare octo works without re-typing --provider/--model every time.
Reaching endpoints through a proxy
Section titled “Reaching endpoints through a proxy”If reaching OpenAI, Anthropic, or other endpoints requires a proxy, octo honors Go’s standard proxy environment variables — HTTPS_PROXY, HTTP_PROXY, NO_PROXY (both http:// and socks5:// addresses work, either case):
HTTPS_PROXY=http://127.0.0.1:7890 octo --provider openai "..."A shell export is enough when you launch from a terminal. But the desktop app and a background octo serve don’t inherit your shell environment — an export in ~/.zshrc never reaches them. The uniform fix is ~/.octo/serve.env, which every launch mode (CLI, desktop, octo serve) loads at startup:
cat >> ~/.octo/serve.env << 'EOF'HTTPS_PROXY=http://127.0.0.1:7890EOFchmod 600 ~/.octo/serve.envRestart the app after editing. The proxy applies to all of the process’s outbound traffic (model APIs, search, update checks); to route only some domains through it, use your proxy software’s rules (Clash & co.) or exclude hosts with NO_PROXY=api.moonshot.cn — octo has no per-provider proxy setting.
Extended reasoning
Section titled “Extended reasoning”Reasoning models can deliberate before answering. Two knobs control it, both available as CLI flags
and as octo config defaults:
--reasoning-effort low|medium|high|xhigh|max— the intensity. OpenAI-protocol backends receive it asreasoning_effort; Anthropic-protocol backends map it to adaptive thinking / an extended-thinking token budget, normalized per model family. Empty (the default) means off.--show-reasoning(default off) — surface the reasoning/thinking trace for the Web UI (octo serve) to display. The terminal never renders the trace either way.
This unifies Anthropic thinking blocks and OpenAI reasoning_content behind one pair of controls.
Next: see the full schema in the config file reference.