Files
ai-agent/sidecar/README.md
Gabriel Vidal 68a3206641 Merge branch 'main' into split-big-files
# Conflicts:
#	backend/main.py
#	backend/schemas.py
#	sidecar/sidecar.py
#	sidecar/test_claude_args.py
2026-10-06 23:57:03 +02:00

9.9 KiB

claude-sidecar

A tiny host-side FastAPI wrapper that launches claude -p sessions on behalf of the ai-agent viewer.

Why a host process (not a container)

The ai-agent viewer runs in Docker and can't reach the host's claude CLI — its auth (~/.claude/.credentials.json), RTK/vault hooks, and skills all live in the host user's home. This sidecar runs natively on the host as the repo owner, so a session it spawns is exactly like one started from a terminal.

Flow

Frontend (sticky input)
   │  POST /api/spawn {prompt}
   ▼
ai-agent backend (container)
   │  generates a session UUID, POST /spawn to the sidecar
   ▼  http://host.docker.internal:8790/spawn   (Bearer SIDECAR_TOKEN)
claude-sidecar (host)
   └─ claude -p <prompt> --session-id <uuid> --output-format json \
              --permission-mode bypassPermissions --model opus \
              --remote-control   (detached)

--remote-control is added by default so every spawned run registers with Claude Code Remote Control and shows up in the Claude app — you can watch and drive the conversation from your phone. Set SIDECAR_REMOTE_CONTROL=0 to opt out.

Claude writes its transcript to ~/.claude/projects/-home-gabrielvidal-homelab/<uuid>.jsonl, which the ai-agent container already watches read-only. The new conversation appears in the viewer within ~2s and the frontend redirects to it once it has synced.

The conversation view has a matching sticky composer that continues an existing thread: it POSTs /api/spawn's sibling /api/resume, which the backend proxies to POST /resume here — claude -p <prompt> --resume <sessionId> in the session's original cwd. Because resume reuses the session id, the new turns append to the same transcript and stream straight into the open conversation.

Install (host)

sidecar/install.sh

Creates a venv, writes a systemd user unit (~/.config/systemd/user/claude-sidecar.service), and starts it. Reboot survival needs sudo loginctl enable-linger $USER.

Config comes from the repo .env: SIDECAR_TOKEN, SIDECAR_PORT (default 8790). The ai-agent container reads SIDECAR_URL + SIDECAR_TOKEN (see its docker-compose.yml, which also adds host.docker.internal:host-gateway).

Modules

uvicorn sidecar:app (and python sidecar.py on the macOS worker) loads one thin entry module that mounts a router per concern:

  • sidecar.py — the FastAPI app: logging filter, startup hook, router mounts, __main__.
  • config.py — every env var and constant (SIDECAR_*), each with its comment.
  • pids.py — pid-record tracking: logs/<sid>.pid, liveness (/proc or psutil), startup reconcile.
  • validate.py — bearer check and per-field validation (uuid, model, harness, thinking / effort flags).
  • claude_accounts.py — account resolution, the run env (CLAUDE_CONFIG_DIR, notify CLI on PATH), claude auth status cache.
  • launch.py — the detached claude -p / runner launch, launch probe, stop, --settings / --remote-control.
  • routes_runs.py — /sessions, /spawn, /resume, /interrupt, /message.
  • routes_accounts.py — /health, /accounts, the login / logout routes.
  • fork.py — /fork and the transcript surgery behind it.
  • feed.py, pairing.py, outbox.py, skills_list.py — the worker endpoints (make_router(auth, …)).
  • claude_cli.py — headless claude auth login driving and error classification.
  • stop_guard.py — the Stop hook injected into every claude run (see below).

Endpoints

  • GET /health → {ok, cwd, claude, accounts, defaultAccount}
  • GET /accounts (Bearer auth) → each account's claude auth status (loggedIn, email, orgName, subscriptionType), cached 60s, plus the pendingLogin below when one is waiting.
  • POST /accounts/{id}/login {email?} → {loginId, url, startedAt, expiresAt} — starts a headless claude auth login --claudeai under the account's env and answers the OAuth sign-in URL. One login per account; killed after 10 min (SIDECAR_LOGIN_TTL_S).
  • POST /accounts/{id}/login/code {code} → the fresh auth status — pipes the code#state the sign-in page ends on into the waiting CLI. 400 bad_code (malformed, or a #state from another attempt) keeps the login pending; 502 login_failed ends it.
  • DELETE /accounts/{id}/login → {cancelled}.
  • POST /accounts/{id}/logout → claude auth logout + the fresh status.
  • POST /spawn (Bearer auth) {prompt, sessionId?, model?, cwd?, account?} → {sessionId, pid, log} — spawns detached, returns immediately.
  • POST /resume (Bearer auth) {sessionId, prompt, model?, cwd?, account?} → {sessionId, pid, log} — continues an existing conversation by running claude -p <prompt> --resume <sessionId>. Reuses the original session id, so the new turns append to the same transcript and stream into the open conversation. --resume is directory-scoped, so pass the session's original cwd.
  • POST /interrupt (Bearer auth) {sessionId} → {sessionId, pid, signal, ok} — sends the run's process group a SIGINT (like Ctrl+C), so Claude aborts the turn, writes a [Request interrupted by user] marker and exits. 404 if no live process is tracked for that session.
  • POST /message (Bearer auth) {sessionId, prompt} → {sessionId, pid, delivered: "inbox"} — delivers a follow-up into a live run without stopping it, through the CLI's per-session inbox socket (<config dir>/sessions/<pid>.json → messagingSocketPath; runs are launched with crossSessionInbound: accept). The model reads it between tool calls, or at once when idle-waiting, and its background work survives. 404 not_running when no process is live (do a plain /resume), 409 inbox_unavailable when the run can't take it (a pi run, an older CLI — fall back to /resume with force). The text is prefixed (INBOX_PREFIX) because the CLI records it as a peer message; the backend's parser strips both.

Background work in a headless run

A claude -p process ends when the model ends its turn. Plain Bash run_in_background tasks die with it at once (the CLI's local_bash tasks are excluded from its exit wait); subagents and Monitors are waited for, but only up to CLAUDE_CODE_PRINT_BG_WAIT_CEILING_MS (CLI default 10 min). The sidecar therefore launches every claude run with:

  • CLAUDE_CODE_PRINT_BG_WAIT_CEILING_MS = SIDECAR_BG_WAIT_CEILING_MS (default 7200000, 2 h; 0 = wait indefinitely; "" = the CLI default);
  • --settings carrying a Stop hook, stop_guard.py: while a running plain background shell is left, it holds the stop — polling the task's tasks/<id>.output for the CLI's [exited with code N] marker, no API call — then answers block with "your task finished, its output is at …". Monitors are excluded. One hold lasts at most STOP_GUARD_MAX_WAIT_S (3300 s) under the hook's SIDECAR_STOP_GUARD_TIMEOUT_S (3600); SIDECAR_STOP_GUARD=0 disables it. The hook's stop-guard: … lines go to logs/<sid>.guard.log next to the run log (the CLI keeps a hook's stderr to itself; STOP_GUARD_STATE_DIR = LOG_DIR in the run env); the hold shows in the viewer as a compact held tag.
  • A follow-up during a hold. The inbox socket accepts and drops a message while the session sits in a Stop hook, so /message checks for the hook's logs/<sid>.hold marker first and, when present, leaves the text in logs/<sid>.inbox/ (answers delivered: "hold"); the hook polls that mailbox, ends the hold at once and relays the text in its block reason, fenced between <<<USER MESSAGE>>> / <<<END USER MESSAGE>>> — the viewer shows it as the user's own turn — and asks the model to answer, then end its turn again so the hold resumes for what is still running.

Per-session stdout/stderr is captured under logs/<sessionId>.log; the pid is recorded in logs/<sessionId>.pid so /interrupt can find the run.

Errors are {detail: {message, kind, hint?, account?}}. kind (from claude_cli.classify): auth, rate_limit, cli_outdated, overloaded, network, unknown, and for the login routes bad_code, login_failed, login_expired, login_timeout. A run that dies on launch is read from the CLI's JSON result line ("Not logged in · Please run /login") instead of echoing that line raw.

Headless login (claude_cli.py)

With no TTY, claude auth login prints an authorize URL whose redirect is Anthropic's manual code page and blocks on stdin at Paste code here if prompted >. Signing in (any device) ends on a page showing <code>#<state>; written to the CLI's stdin it's exchanged for tokens in the account's config dir. A code without #state → Invalid code… and the CLI re-prompts; a wrong one → exit 1 Login failed: Request failed with status code 400. The sidecar checks the pasted #state against the URL's before piping it, and also accepts the whole callback URL (…?code=…&state=…). A pending login lives only in the sidecar process — restarting it drops the login.

Accounts (personal vs work)

A claude run launches on one Claude login. The default account is the CLI's own config (~/.claude); SIDECAR_ACCOUNTS=work=~/.claude-work (in the unit, written by install.sh) declares more, each exported per run as CLAUDE_CONFIG_DIR. account on /spawn, /resume and /fork picks one — resume and fork must pass the account the session started on, since its transcript lives only in that account's projects/. A non-default account whose dir has no .credentials.json answers 409 (never a fallback onto the default login), ANTHROPIC_API_KEY is stripped from every run once accounts are declared, and --remote-control is only added for the default account. The account is recorded in the pidfile and listed by /sessions.

Manage

systemctl --user status  claude-sidecar
systemctl --user restart claude-sidecar
journalctl --user -u claude-sidecar -f