# nconnect > Use Nebius Token Factory models in the coding-agent CLIs you already run - OpenCode, Claude Code, Codex CLI, Pi Code, Prime Agent, and Unreal Agent. nconnect injects Nebius settings for a single run so your normal tool configs stay untouched. nconnect is an MIT-licensed CLI router. Install once, then launch your agent with short commands (`ncodex`, `nclaude`, `nopencode`, `npi`, `nprime`, `nunreal`). For Claude Code, Codex CLI, and Unreal Agent it routes the agent's API traffic to Nebius Token Factory through a local translation proxy; for OpenCode, Pi Code, and Prime Agent it injects per-run provider config (they already speak Nebius's OpenAI-compatible format, so no proxy is needed). The installed binary keeps itself up to date automatically from nconnect.sh. This is the LLM-readable documentation file for the project. If you are an AI agent asked to install, configure, or use nconnect (headless or interactive), read this file first. It covers installation, configuration, every command, the available models, headless/agentic usage patterns, how to keep the tool up to date, and the routing architecture. - Project: https://github.com/Studio1-OSS/nconnect - Website: https://nconnect.sh - Installer: https://nconnect.sh/install.sh - Version manifest: https://nconnect.sh/latest.json - README: https://github.com/Studio1-OSS/nconnect/blob/main/README.md - This file: https://nconnect.sh/llms.txt ## Install One command installs the `nconnect`, `nclaude`, `nopencode`, `ncodex`, `npi`, `nprime`, and `nunreal` commands to `~/.nconnect/bin/` and installs Bun for you if it isn't already present: ``` curl -fsSL https://nconnect.sh/install.sh | sh ``` After install, restart your shell (or run the `export PATH=...` line it prints) so the commands are on PATH. The install: - Downloads the self-contained JS bundle to `~/.nconnect/bin/nconnect.js`. - Writes shell wrappers (`nconnect`, `ncodex`, `nclaude`, `nopencode`, `npi`, `nprime`, `nunreal`) that run the bundle with `bun`. - Links them into a writable directory on your PATH. - Does NOT install the underlying agent CLIs. If a tool is missing, nconnect prints its official install command and exits (Prime Agent: `curl -fsSL https://app.primeintellect.ai/prime-agent/install.sh | sh`). To verify: `nconnect --version`. ## Configure Before the first launch you need a Nebius API key (get one at https://tokenfactory.nebius.com/?modals=create-api-key): ``` nconnect configure ``` This stores your Nebius API key in `~/.nconnect/` and optionally an Exa API key (https://exa.ai) that powers web-search emulation in Claude Code. Web search is optional - without it, searches return a clear "TAVILY_API_KEY not set" error instead of failing silently. You can also set the key non-interactively: ``` export NEBIUS_API_KEY=... nconnect configure # stores it (skip prompts with the env present) ``` The API key is required to actually route traffic; the installer lets you skip it (press Enter) and add it later. ## Commands Running `nconnect` with no arguments opens an interactive launcher (requires a TTY) that lets you pick a tool. Run a tool directly with a subcommand: ``` nconnect configure # set API keys nconnect whoami # print your anonymous install id nconnect --version | -v # print the version nconnect help | --help | -h # print this usage nconnect codex [...] # alias: ncodex (Codex CLI) nconnect claude [...] # alias: nclaude (Claude Code) nconnect pi [...] # alias: npi (Pi Code) nconnect opencode [...] # alias: nopencode (OpenCode) nconnect prime [...] # alias: nprime (Prime Agent, PrimeIntellect) nconnect hermes [...] # alias: nhermes (Hermes Agent, Nous Research; `hermes desktop`) nconnect deepseek [...] # alias: ndeepseek (DeepSeek Harness, alpha) nconnect grok [...] # alias: ngrok (Grok Build UI on Nebius models) nconnect usage [--last 7d] # local spend report by model and tool nconnect update # update to the latest release nconnect daemon install|uninstall|status # daemon auto-start at login nconnect chatgpt [--model ] [--restore] # alpha: ChatGPT Desktop (alias: codex-app) nconnect daemon stop # stop the background proxy daemon ``` Any arguments after the harness name are passed straight through to the underlying agent CLI. For example, `ncodex exec "fix the typo"` runs `codex exec` with nconnect routing. Recognized nconnect flags (parsed before the passthrough): - `--api-key ` - use this Nebius API key for the run. - `--main ` / `--model ` - select the model for this run. - `--search ` - search setting (harness-specific). - `--slot ` - session slot selection. - `--json` - print machine-readable payloads where supported. - `--restore` - restore (used with `chatgpt`). ### ChatGPT Desktop App (alpha) `nconnect chatgpt` persistently patches the ChatGPT Desktop config so the desktop app routes through Nebius. Unlike the CLI wrappers, this change stays active until you restore it: ``` nconnect chatgpt # configure (patches config) nconnect chatgpt --restore # restore your OpenAI profile nconnect chatgpt --model # use a specific model ``` Backups of your original config live under `~/.nconnect/backup/codex-app/`. ## Models The model list is fetched live from Nebius (`GET /v1/models?verbose=true`) at startup, so every model Nebius serves is available and each model's vision support (modality) comes straight from the API, not a hand-maintained list. Results are cached in `~/.nconnect/model-catalog.json` and fall back to a bundled snapshot when offline. The default for Codex and Claude Code is GLM 5.3 Flash. ### Selecting a model (flag order matters) Put `--model` (or `--main` for Claude) BEFORE the harness subcommand - nconnect consumes it there and routes to that model: ``` nconnect --model moonshotai/Kimi-K2.6 codex exec "task" nconnect --model moonshotai/Kimi-K2.6 claude -p "task" nconnect --main nebius-kimi-k2-7-code claude -p "task" # Claude matches alias OR id ``` The long form above works across harnesses. Since Relay 0.15.3, Codex also supports `ncodex --model MODEL_ID` and `ncodex -m MODEL_ID`; invalid explicit model names fail instead of silently using the default. Featured flagship models (any live Nebius id also works with `--model`): - `zai-org/GLM-5.3-Flash` - GLM 5.3 Flash. The default coding model, text-only, 1M context. - `zai-org/GLM-5.3` - GLM 5.3. Text-only, coding and tool use, 1,024,000 context. Bundled since 0.15.4; Claude alias: `nebius-glm-5-3`. - `deepseek-ai/DeepSeek-V4-Pro-0813` - DeepSeek V4 Pro 0813. Text-only, reasoning and agentic coding, 979,000 context. Bundled since 0.15.4; Claude alias: `nebius-deepseek-v4-pro-0813`. - `moonshotai/Kimi-K3` - Kimi K3. Text-only, reasoning + agentic coding. - `moonshotai/Kimi-K2.6` - Kimi K2.6. Vision + reasoning, 262K context. The vision flagship. - `moonshotai/Kimi-K2.7-Code` - Kimi K2.7 Code. Text-only coding model, 262K context. - `MiniMaxAI/MiniMax-M3` - MiniMax M3. Text-only, 196K context, fast and cheap. - `Qwen/Qwen3.5-397B-A17B` - Qwen 3.5 397B. Text-only, 262K context, general/coding flagship. - `deepseek-ai/DeepSeek-V4-Pro` - DeepSeek V4 Pro. Text-only, 1M context, long-context reasoning. - `Qwen/Qwen2.5-VL-72B-Instruct` - Qwen2.5-VL 72B. Vision, 32K context. The vision fallback. Note: `zai-org/GLM-5.2` was the previous default but Nebius removed it from the live catalog (2026-07-27); `zai-org/GLM-5.1` is available in its place. Vision-capable models (those whose API modality includes image input, e.g. Kimi K2.6 and Qwen2.5-VL) accept image attachments directly; text-only models route images to a vision-capable model automatically (Claude/Codex proxy) or to the `@vision` subagent (OpenCode). ## Headless & agentic usage Both Claude Code and Codex CLI run headlessly (no TTY), so you can delegate code changes to an AI agent from scripts, CI, or another orchestrating agent. The patterns below were verified against the installed `ncodex` and `nclaude` (v0.5.26) and apply to both. They differ in the headless subcommand and in how permissions are handled; the stdin, banner, and model-selection behavior is identical. ### The working invocation pattern Both wrappers spawn the agent pointed at the local Nebius proxy and pass your flags through to the underlying binary. Always close stdin and capture stderr. To run a non-default model, put `--model ` before the harness (see Models above). ### Prefer native subagents for delegated nconnect runs If you are an orchestrating agent running inside Claude Code, Codex, or another CLI that has its own native subagent/delegation tool, prefer this pattern for non-trivial nconnect work: 1. Spawn a native subagent first. 2. Tell that subagent to run the needed `nconnect` command itself (`ncodex`, `nclaude`, `nopencode`, or `npi`) with a tight task, explicit files, `< /dev/null`, and the guardrails below. 3. Have the subagent wait for the delegated nconnect run, inspect its result/diff, and report only the useful summary back to you. Do this instead of having the main orchestrating agent call `ncodex`/`nclaude` directly for every delegated edit. It keeps the main agent's context and inference budget focused on coordination/review, lets the native subagent own the details of launching and monitoring nconnect, and works especially well when Claude Code delegates to a subagent that then uses Codex through `ncodex`. Codex CLI - non-interactive `exec` (the `-s workspace-write` sandbox grants write access; `exec` never prompts interactively): ``` nconnect codex exec -s workspace-write -C /abs/repo "" < /dev/null 2>&1 ``` If the working directory isn't a git repo, codex refuses to start ("Not inside a trusted directory"); add `--skip-git-repo-check`, or run inside a git repo. Claude Code - print mode `-p` (runs once and exits). Add `--dangerously-skip-permissions` so a headless subagent never blocks on a permission prompt: ``` nconnect claude -p "" --dangerously-skip-permissions --output-format json < /dev/null 2>&1 ``` Run it in the background, health-check at ~60s, review the diff, run lint, then commit. One file or one concern per task gives the cleanest reviews. ### Always close stdin: `< /dev/null` This is the single biggest footgun. When launched in the background with an open/inherited stdin and no TTY, the process prints `Reading additional input from stdin...` and blocks forever - it never starts the task and produces no edits. - **Workaround:** append `< /dev/null` to every invocation. With stdin closed it runs cleanly every time. - **Do not** treat the `Reading additional input from stdin...` line as a hang signal on its own - that line also prints on healthy `< /dev/null` runs (it reads EOF immediately and proceeds). The real hang signal is: that line is the **last** line AND the output line count stays frozen. Health checks should watch for a *frozen line count*, not the presence of the stdin message. ### Health-check a launch, don't assume A hung/failed start and a legitimately-working task both look like "background process still running." Without tailing the log there's no way to tell in the first minute whether it connected. Adopt a scripted ~60s health check: - process alive, AND - the banner line `Routing Codex → Nebius Token Factory` (or `Routing Claude Code → Nebius Token Factory`) is present (confirms it connected, not hung pre-connection), AND - no `command not found` / startup error in the log, AND - output line count is still growing (not frozen). Kill and relaunch on a failed start. ### One concern per task; keep taste on the orchestrator side Give explicit file lists and verbatim text for any creative content. Keep taste decisions (copy, design) on the human/orchestrator side and pass the exact text to insert. When told to touch only a named set of files, the agent respects the scope precisely and will flag (not silently fix) a file outside the scope that has the same issue - surface that gap for a human to decide. ### Wait for done before reviewing or linting Do not lint or review mid-run. The agent rewrites files in passes and a file can be briefly malformed mid-edit (e.g., a stray `*/` mid-comment) that is correct by task end. Reviewing mid-run chases phantom errors. Wait for the done signal, then: 1. Review the diff yourself. 2. Re-run lint yourself (the agent self-verifies with lint and reports honestly, but the review step has caught real issues - e.g., a `server-only` import that would crash, or a 7th file needing the same fix). 3. Commit. ### Cost awareness Per-task cost varies a lot with how verbose the agent is. Mechanical, well-scoped changes are cheap (~$0.10); a single change that triggers verbose JSDoc + defensive guard code can cost ~10–20× more. To keep cost down: scope tasks tightly, and prefer a "match surrounding comment density; don't over-explain" instruction for small mechanical changes. ### Summary of recommendations 1. Always `< /dev/null`. 2. One concern per task; give an explicit file list and verbatim text for creative content. 3. 60s health check after launch; kill + relaunch on a failed start (watch for frozen line count). 4. Never lint/review mid-run; wait for the done signal. 5. Review the diff + re-run lint yourself before committing. 6. To select a model, put `--model` (or `--main` for Claude) before the harness subcommand - after the harness it is silently ignored and the default runs. 7. For Claude `-p`, pass `--dangerously-skip-permissions` so a headless subagent never blocks on a permission prompt. Codex `exec` is non-interactive by design; use `-s workspace-write` for edits and `--skip-git-repo-check` outside a git repo. 8. If your current CLI has native subagents, delegate nconnect launches to one of those subagents and let it return the result, instead of spending the main orchestrator's context on the full `ncodex`/`nclaude` run. ## Keeping nconnect up to date The installed binary self-updates automatically in the background. The check is throttled and bounded, never throws, and runs before argument parsing - so even `nconnect help` keeps an install current. The update is pulled from https://nconnect.sh/latest.json, which contains `version`, `url`, and `publishedAt`. To force a refresh, re-run the install one-liner: ``` curl -fsSL https://nconnect.sh/install.sh | sh ``` You can check the installed version with `nconnect --version` and compare it against the published manifest at https://nconnect.sh/latest.json. ## Architecture (how routing works) There are two harness families behind the four tools: - **Proxied harnesses - Claude Code, Codex CLI.** `nclaude` and `ncodex` spawn a shared local daemon and route the agent's `/v1/*` traffic through it. The daemon translates the agent's native wire format (Anthropic Messages for Claude, OpenAI Responses for Codex) to Nebius's chat-completions API and back. Per run: it resolves the model, ensures the daemon, registers a session with your token, prints a banner, spawns the agent pointed at the local proxy, tracks cost, and deregisters on exit. Nothing is written to your real agent config - close it and your setup is exactly as it was. - **Spawned harnesses - OpenCode, Pi Code, Prime Agent.** `nopencode`, `npi`, and `nprime` spawn the agent binary directly. The binary talks to Nebius itself, using ephemeral config injected only for that launch (OpenCode), a temporary `models.json` on disk (Pi), or a relay-owned config dir under `~/.nconnect/prime-agent` (Prime Agent). No daemon, no proxy. The daemon is one persistent local process shared by all proxied sessions. If you ever need to stop it manually: ``` nconnect daemon stop ``` `nconnect chatgpt` (alpha) is the exception: it persistently patches ChatGPT Desktop config to route through Nebius (instead of per-run injection), with backups under `~/.nconnect/backup/codex-app/`. Restore with `nconnect chatgpt --restore`. ## Troubleshooting - **"No Nebius API key found"** - run `nconnect configure` or set `NEBIUS_API_KEY`. - **A tool command is not found** - nconnect does not install the agent CLIs for you. It prints the official install command (e.g. `npm install -g @openai/codex`) and the docs link when the tool is missing; install it, then re-run. - **Web search fails in Claude Code** - set an Exa API key via `nconnect configure`. Without it, searches return a clear "TAVILY_API_KEY not set" error. - **Headless run hangs and produces no edits** - you forgot `< /dev/null`. Close stdin and relaunch; see the Headless usage section. - **Cost seems high for a small change** - the agent over-documented. Scope the task tighter and add a "don't over-explain" instruction. - **Want your normal config back** - stop using the wrappers. No agent config was saved, so your OpenCode / Claude Code / Codex CLI / Pi Code setup is untouched. For ChatGPT Desktop, run `nconnect chatgpt --restore`. ## Links - [GitHub](https://github.com): source repository and issues - [README](https://github.com): human-readable project overview - [Website](https://nconnect.sh): landing page - [Install script](https://nconnect.sh/install.sh): one-command installer (`curl -fsSL https://nconnect.sh/install.sh | sh`) - [Version manifest](https://nconnect.sh/latest.json): current published version and release timestamp - [This LLM docs file](https://nconnect.sh/llms.txt): the file you are reading now - [Nebius Token Factory API keys](https://tokenfactory.nebius.com/?modals=create-api-key): get a key to route traffic - [llms.txt specification](https://llmstxt.org): the convention this file follows