TUTORIAL
How to run DeepSeek V4 Flash with Prime Agent
Run PrimeIntellect's Prime Agent on DeepSeek V4 Flash, a fast 1M-context model for coding and agentic work, served by Nebius Token Factory through NConnect.
Prime Agent is PrimeIntellect's RLM-style coding agent. It already speaks the OpenAI-compatible format Nebius serves, so NConnect points it straight at Token Factory with no proxy. DeepSeek V4 Flash is the fast, cost-effective member of the DeepSeek V4 family, with a 1M-token context window.
Prerequisites
- macOS or Linux.
- A Nebius Token Factory API key from tokenfactory.nebius.com.
Steps
Install NConnect
This installs the CLI and the
nprimealias, plus Bun if needed.curl -fsSL https://nconnect.sh/install.sh | bashAdd your Nebius API key
nconnect configureInstall Prime Agent
Skip this if
prime-agentis already on your PATH. In an interactive terminal, NConnect also offers to run this for you the first time you launch Prime Agent.curl -fsSL https://app.primeintellect.ai/prime-agent/install.sh | shLaunch Prime Agent on DeepSeek V4 Flash
nconnect --model deepseek-ai/DeepSeek-V4-Flash primeYou should see NConnect ▸ Launching Prime Agent with Nebius Token Factory (DeepSeek V4 Flash).
What NConnect does on launch
- Declares Nebius as a custom provider in a config directory NConnect owns,
~/.nconnect/prime-agent. Your own~/.prime/agentis never touched. - Keeps that directory between launches, because Prime Agent bootstraps its runtime there. The first launch takes longer; later ones start straight away.
- Lists every Nebius model, so you can switch inside Prime Agent.
Track cost and add fallback
Prime Agent is a spawned harness, so by default it calls Nebius directly and NConnect reports $0.00. Route it through the daemon to meter every turn and get automatic fallback if DeepSeek V4 Flash is busy:
NCONNECT_METER=1 nconnect --model deepseek-ai/DeepSeek-V4-Flash primeThen check spend with nconnect usage --last 7d. See Cost metering.
