01 / GET STARTED
How NConnect works
NConnect runs a local daemon that translates the Anthropic and OpenAI Responses wire formats to Nebius chat completions, so your coding agent runs unchanged on open models.
Nebius serves open models over an OpenAI-compatible API. It does not speak the Anthropic Messages API that Claude Code uses, nor the OpenAI Responses API that Codex uses. NConnect runs a small local daemon that translates those wire formats to Nebius /chat/completions on the fly. Your agent believes it is talking to its native backend, while every token is served by Nebius.
Two kinds of harness
Proxied harnesses
Claude Code, Codex and Unreal Agent. The daemon translates each request and response, tracks cost, retries transient failures, trims context to fit, and emulates native web search.
Spawned harnesses
Everything else. Launched with a generated provider config pointed at Nebius. No proxy needed, since they already speak the OpenAI-compatible format.
What the daemon adds
- Protocol translation between each agent's wire format and Nebius.
- Cost metering for every turn, viewable with
nconnect usage. Nothing is uploaded. - Resilience: transient failures are retried and an overloaded model fails over to a fallback model.
- Context fitting to each model's window.
- Web search backed by Tavily, with citations.
- Vision routing: image blocks go to a vision-capable model.
Spawned harnesses can opt into the same client with daemon metering.
