Skip to content

Air-gapped deployment

octo is a single Go binary. The Linux build is CGO_ENABLED=0 — fully static, no glibc linkage — and the web console, ripgrep and the bundled skills are all go:embeded into it. Offline install needs no package manager, no Node, no container runtime: copy one file in and run it.

Pair that with the custom vendor, which accepts any base URL and works with no API key at all, and the whole stack runs in an environment that never touches the public internet.

  • One machine with internet access, to download the release (you’re done with it after that)
  • The target machine: any x86_64 or aarch64 Linux, kernel 3.2 or newer
  • An internal model endpoint speaking either the OpenAI-compatible or the Anthropic Messages wire protocol over HTTP

Grab the archive for your target architecture plus checksums.txt from the latest release:

octo_<version>_linux_amd64.tar.gz # x86_64
octo_<version>_linux_arm64.tar.gz # aarch64
checksums.txt

Run uname -m on the target if you’re unsure: x86_64 → amd64, aarch64 → arm64.

  1. Move both files across using whatever transfer path your policy approves.

  2. Verify integrity so a corrupted transfer surfaces now rather than at runtime:

    Terminal window
    sha256sum -c checksums.txt --ignore-missing
  3. Unpack and install onto PATH:

    Terminal window
    tar -xzf octo_<version>_linux_amd64.tar.gz
    sudo install -m755 octo /usr/local/bin/octo
    octo version

octo version printing a version and commit means the binary runs on this host — that single command already rules out architecture and kernel mismatches.

Self-hosted endpoints use the custom vendor: the only provider that takes a free-form base_url, and the only one where api_key may be omitted — local inference servers usually have no auth, and octo then sends no authorization header.

Create ~/.octo/config.yml:

endpoints:
- id: intranet
name: Internal inference
provider: custom
protocol: openai # or anthropic
base_url: http://10.0.0.20:8000/v1
# api_key: <only if your gateway requires one; otherwise drop this line>
models:
- model: your-model-name
vision: false
default: intranet::your-model-name

Choosing protocol: vLLM, Ollama, Xinference, LMDeploy and most self-hosted serving stacks expose an OpenAI Chat Completions-compatible API → openai. Only pick anthropic for a gateway that explicitly implements the Anthropic Messages API. Getting it wrong shows up as a wire-format error on the very first turn; switch it and move on — nothing else is affected.

Where base_url stops: give the root; octo appends the path itself. OpenAI-compatible servers usually want the /v1 suffix included.

The key can come from the environment too: CUSTOM_API_KEY takes precedence over the file’s api_key, which suits a shared machine better.

Terminal window
octo config show # prints the effective provider/model and where each came from
octo "introduce yourself in one sentence"

A keyless custom endpoint in the config counts as fully configured — the web console’s first-run wizard won’t block you asking for a cloud provider key.

Terminal window
octo # chat straight from the terminal
octo serve -d # web console in the background, 127.0.0.1:8088 by default

To share it with colleagues, bind wider. Every non-loopback request then has to present an access key:

Terminal window
octo serve -addr :8088 -d

Without --access-key, octo reads OCTO_ACCESS_KEY, then config.yml, then generates and persists one — startup prints a ready-to-open URL with ?access_key=… embedded. The full boundary is in the security model.

Running it under an init system, log locations and ~/.octo/serve.env all work exactly as they do online — see Self-host octo serve.

Capability Behavior in an isolated network
web_search Unavailable. The backends are Tavily / Bing and friends, so calls time out and error. The model sees the error and routes around it
web_fetch Limited to URLs reachable inside your network
octo upgrade, the web version badge Try GitHub and fail. Nothing else breaks; upgrade via the offline flow below. Set update_check: false in ~/.octo/config.yml (or flip the Settings toggle) to stop the badge trying at all
Phone tunnel (-tunnel) Off by default — nothing dials a relay unless you ask
MCP servers Local stdio processes are fine. npx @modelcontextprotocol/... entries need their npm packages staged offline first; config lives in ~/.octo/mcp.json
computer tool (desktop control) A no-op stub on Linux regardless of connectivity — only macOS and Windows have real implementations
Desktop AppImage Needs GTK4 + WebKitGTK 6.0 on the host, often missing on hardened or older distros. Prefer the CLI plus octo serve in a browser — same feature set
Image understanding Depends on whether your internal model supports it. Set vision: true/false on the model entry honestly

octo uses the context window to decide when to auto-compact history (at 75% of the window by default). Set the limit on each self-hosted model entry so it matches what that deployment actually accepts, even when its model id also matches a larger model in octo’s built-in table:

endpoints:
- id: intranet
provider: custom
protocol: openai
base_url: http://10.0.0.20:8000/v1
models:
- model: Qwen3-32B
context_window: 32000

The configuration API rejects values from 1 to 999 with a note about the units — 32k is 32000, not 32. If you hand-edit config.yml, the runtime treats such a value as unset and octo doctor reports the unit mistake.

For unrecognized models without a per-model value, you can set the global fallback instead:

fallback_context_window: 32000 # tokens — 32k is 32000, not 32

The flag and environment variable work too, in the order flag > env > config file:

Terminal window
octo --fallback-context-window 32000
OCTO_FALLBACK_CONTEXT_WINDOW=32000 octo serve -d

Resolution order is: the selected model entry’s context_window, the built-in model table, fallback_context_window, then the built-in 128k default. The per-model setting is therefore the right choice when a known model is served with a smaller limit. Since it belongs to the endpoint’s model entry, the same model id can safely have different limits on two endpoints.

You don’t have to do anything — the internet-facing tools simply fail. For a cleaner setup, define an agent profile with a tool allowlist at ~/.octo/agents/<id>.md that omits web_search and web_fetch:

---
name: Intranet assistant
description: Assistant restricted to local tools.
tools:
- terminal
- read_file
- write_file
- edit_file
- grep
- glob
---
You work in an isolated network with no internet access.

To sanity-check the bare model instead, octo --no-tools disables every built-in tool, MCP surface and skill.

Download the new tar.gz and checksums.txt on the connected machine and repeat steps 1 and 2 over the existing binary. Nothing under ~/.octo/ — config, sessions, skills — is touched:

Terminal window
sha256sum -c checksums.txt --ignore-missing
tar -xzf octo_<new-version>_linux_amd64.tar.gz
octo serve --stop # if a daemon is running
sudo install -m755 octo /usr/local/bin/octo
octo serve -d

Everything sits under ~/.octo/ — nothing is written outside the user’s home: config.yml, sessions/, skills/, agent profiles in agents/, mcp.json, and serve.log. Back it up or migrate it by copying that one directory.

If policy requires you to compile in-network, note that the web console’s build output internal/server/webdist is not in the repository — it’s a gitignored artifact. A complete build therefore needs:

  • An internal Go module proxy (GOPROXY), or dependencies vendored ahead of time
  • Node 22 and an internal npm registry, to run npm ci && npm run build in web/ first

make build still produces a working binary without the frontend step; octo serve just serves a blank page. The CLI and TUI are unaffected.

For most teams, copying the official release is less work — it already carries the built frontend and ripgrep.

Notes on hardened and domestic Linux distributions

Section titled “Notes on hardened and domestic Linux distributions”

Because the Linux target is a static build with no system-library dependencies, it should run as-is on Kylin, UOS, openEuler and similar distributions, on both aarch64 (Phytium, Kunpeng) and x86_64 (Hygon, Zhaoxin) hardware.

To be straight about it: our CI only covers Ubuntu — none of those platforms have been tested in practice. Prove out octo version and one full conversation on a single machine before rolling out broadly, and please open an issue with your uname -a and distribution version if something breaks.