NConnect
DocsHow NConnect works

01 / GET STARTED

How NConnect works

NConnect runs a local daemon that translates the Anthropic and OpenAI Responses wire formats to Nebius chat completions, so your coding agent runs unchanged on open models.

Nebius serves open models over an OpenAI-compatible API. It does not speak the Anthropic Messages API that Claude Code uses, nor the OpenAI Responses API that Codex uses. NConnect runs a small local daemon that translates those wire formats to Nebius /chat/completions on the fly. Your agent believes it is talking to its native backend, while every token is served by Nebius.

Two kinds of harness

Proxied harnesses

Claude Code, Codex and Unreal Agent. The daemon translates each request and response, tracks cost, retries transient failures, trims context to fit, and emulates native web search.

Spawned harnesses

Everything else. Launched with a generated provider config pointed at Nebius. No proxy needed, since they already speak the OpenAI-compatible format.

What the daemon adds

  • Protocol translation between each agent's wire format and Nebius.
  • Cost metering for every turn, viewable with nconnect usage. Nothing is uploaded.
  • Resilience: transient failures are retried and an overloaded model fails over to a fallback model.
  • Context fitting to each model's window.
  • Web search backed by Tavily, with citations.
  • Vision routing: image blocks go to a vision-capable model.

Spawned harnesses can opt into the same client with daemon metering.