CLAUDE.md "Backend modules" now opens with the per-domain packages, the barrels that keep the flat import spellings, and the deps: State / AppState convention; every module bullet and cross-reference points at its new path (conversations/parse.py, core/auth.py, runs/sidecar_client.py, …). The schema-drift and backfill scripts' comments follow the parser's new home. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
80 KiB
CLAUDE.md — ai-agent
Guidance for Claude Code when working in this repo. It was extracted from the
homelab monorepo (formerly services/ai-agent/, full history preserved); the
homelab keeps a deployment shim at ~/homelab/services/ai-agent/ (compose with
build.context pointing here, traefik.yml, deploy.sh, the live data/
volume), and the homelab root CLAUDE.md conventions (worktrees, Traefik copy
step, commit/notify) still apply when deploying there.
North star
This service is headed somewhere bigger than a homelab dashboard: the open-source, intuitive, private and secure agent layer of any computer, to build and host websites and interact with data — Lovable, but open source and self-hosted. Zipgo under the hood for hosting, shipped as a single Docker container, with self-update handled inside the container (not by host scripts). See GOAL.md for the full vision and the gap list. When making changes here, prefer container-internal solutions, generic config over homelab-hardcoded paths, and zipgo over bespoke hosting.
What this is
A viewer/editor + analytics dashboard for the homelab's Claude context, served
as an installable PWA at https://ai-agent.lab.gabvdl.xyz (behind Authelia). It:
- exposes every
CLAUDE.mdand the.claude/tree (skills, hooks, settings) as a flat, editable file list with real Claude token counts + dollar costs, plus the repo's rootdata/tree (plans, logs, notes, kanban) surfaced read-only; - browses the archived Claude Code conversation transcripts (thread view with per-turn usage/cost, skill-usage analytics, spawn/resume/interrupt live runs);
- catalogs ~/projects (mounted at
/workspace/projects; gallery + detail) and the repo's services/ (static parse of compose + Traefik configs — never Docker itself); - surfaces the assistant's memories and the scaffolding templates;
- streams updates over SSE so the UI refreshes without a manual reload.
It is a viewer: it reads the repo context (mostly read-only mounts) and does
static parsing. It never drives Docker or mutates infrastructure. The only writes
are file edits into the mounted .claude/ + ~/projects trees and its own SQLite
store / metadata sidecar.
Layout
backend/— FastAPI app (main:app, port 8080). Serves/api/*and the built PWA fromSTATIC_DIR(SPA catch-all mounted last bymain.create_app). One package per domain (see "Backend modules").frontend/— Vite + React + TS + Tailwind PWA. Built in the Dockerfile's first stage into/app/static.sidecar/— a host-side FastAPI process (NOT in the container) that launchesclaude -psessions. See below.Dockerfile— 2 stages: build the React PWA, then apython:3.12-slimimage runninguvicorn.docker-compose.standalone.yml— the generic, no-host-mounts shape (see "Standalone mode" below). The homelab's compose service (profiledev, containerai-agent, host-only port127.0.0.1:8096:8080, networkmain), its Traefik router (ai-agent.lab.gabvdl.xyz,auth-chain) and the blue-greendeploy.sh+deploy.override.ymllive in the homelab repo'sservices/ai-agent/shim.PROJECT_CLAUDE.md— the generic, host-agnostic project working conventions (worktree workflow, artefacts, plans, committing, the notify+DONEfinish signal), split out of the homelab rootCLAUDE.mdso the ai-agent can ship it as the startingCLAUDE.mdfor a new project. Not this service's own guidance (that's this file) — it's a shippable template with{{…}}placeholders.data/— SQLite DB, transcript archive, metadata sidecar,ui-state.json(bind-mounted, gitignored).
Backend modules
main.py— the assembly only:create_app()installs the trusted-caller gate, builds the shared state, mounts every domain router (ROUTERS, in the old declaration order) and the SPA fallback;_startupstarts the background threads. No route lives here.- One package per domain, each with a
routes.pyexposing anAPIRouter(andservice.py/catalog.py/store.pymodules behind it):core/(config, auth gate,_r/_json_in_workerinhttp.py, the memo, app-level routes: events / deploy-status / health),files/(bundle + file editor, asset serving, uploads, custom avatar),settings/(ui-state, hooks, webhooks, mail-trigger senders),dashboards/(dashboard, activity, skills),conversations/(parser, pricing, cards, subagent tree, catalog, routes, scaffold),diff/,notifications/(feed, pushes, asks),forms/,runs/(sidecar client, the send path, spawn/resume/fork/interrupt),accounts/,workers/,models/,cron/,agents/,projects/,services/,goals/,memories/,plans/,templates/,schemas/. A package whose flat module moved inside keeps its old spelling through a barrel__init__.py(import conversations,import projects,import cron… still work; the private names other modules used are re-exported explicitly). Infrastructure shared by every domain stays flat:db.py,events.py,meta.py,indexer.py,fsscan.py,githist.py,deploy_status.py. - Shared state is injected, not imported.
core/state.pybuilds every store, hub, watcher and mirror once (build_state()→AppState, a dataclass attached toapp.state.ai). A route declaresdeps: State(a FastAPI dependency, excluded from OpenAPI) and a helper takesdeps: AppStateas its first parameter;deps.store,deps.hub,deps.meta_store… replace what used to be module globals. Module-level caches (_Memo, the_*_cachetuples) stay where they are — they are keyed on the store/meta versions, not on the state object. A test builds the app the normal way and readsmain.app.state.ai(seetest_worker_routing.py). - Keep a file under ~500 lines; the parser (
conversations/parse.py) is the deliberate exception —ParserStateis one concern. db.py— SQLite store (WAL):files(content hash + real token count + cost + word/char stats, computed lazily only when a file's hash changes),skills,transcriptsbookkeeping. Decoded transcript summaries are cached in-process (all_summaries()/get_summary()return read-only dicts — copy before mutating) andsummaries_versionbumps on change so request-level memoizers (e.g. the project cost rollup) know when to recompute.indexer.py— background daemon: token/cost accounting (via thecount_tokensendpoint, debounced) + skill-usage mining from transcripts. Pricing consts here.conversations/parse.py— parse*.jsonltranscripts.ParserStateis a resumable parser (feed(record)+summary());parse_conversation(path, full=…)is the one-shot wrapper (cheap rollup vs. full thread). Real per-turn usage/cost fromusageblocks, plus the per-item metadata the viewer shows on each row:durationMs(a tool call's run time — call → its result record; on a message, the assistant API call's latency),model,toolUseId,resultChars/resultLines.fsscan.py—os.scandir-based directory walking. The poll loops re-walk thousands of files every couple of seconds;pathlib.rgloballocates aPathper entry and was a large chunk of idle CPU. Use this, notrglob, in any loop.githist.py— cached git history.BucketedHistorywalks a repo's log once (git log --name-only -- <prefix>) and buckets commits per directory, keyed on the repo HEAD;RepoLogCacheTTL-caches standalone project repos. Before this,/api/servicesforked twogit logs per service per request (11 s).events.py—Hub(SSE pub/sub) +Watcher(polls meta + live transcripts, re-indexes, publishesmeta/transcriptevents).meta.py— per-conversation action metadata JSON sidecar (projects, state, committed/pushed/merged/deployed/notified timestamps), keyed by session id.cron/scheduler.py(+cron/service.pyfor firing a job) — cron jobs: scheduled agent sessions (Settings → Cron). A job = a 5-field cron schedule (dependency-free matcher, container-local time —TZin compose) + a prompt file under.claude/agents/whose content (frontmatter stripped) is spawned as a session via the normal sidecar path; each firing is recorded in the job'shistory({timestamp: sessionId}) and the conversation is tagged (meta.cron→ violet badge linking back to the job). How a job runs is declared in that file's frontmatter — see below. Store is/data/cron-jobs.json(seeded with the every-5hgoal-keeperjob on first run); the scheduler thread claims each fire minute through the store, so the deploy-cutover window where two backends share/datacan't double-spawn. REST under/api/cron(+/run,/prompt); mutations and fires publish acronSSE event.diff/gitdiff.py/diff/worktreediff.py/diff/editsdiff.py— the diff view (/diff/*), in three layers.gitdiffdoes committed history: recent commits per repo token (project:<slug>|service:<slug>|super:_), one commit's hunk model, the net diff across a conversation's commits, and the discovery of which commits a conversation produced (its active time window, plusmain..<worktree branch>; cached in the DB by a fingerprint of the repo HEADs). The other two answer "what isn't committed yet":worktreediff— a realgit diff HEAD(+ untracked files) read from the conversation's still-live worktree. A worktree's.gitfile points at a host path, so it maps the host repo root onto the mounted one and drives git with an explicit--git-dir/--work-treepair. Needs the worktrees root mounted (WORKTREES_DIR,/worktreesin the homelab compose); without it every lookup misses and the replay below takes over.editsdiff— the fallback for a torn-down worktree: the same view replayed from the transcript'sstructuredPatch/originalFilepayloads, henceapprox: true. Paths from the homelab checkout, a~/projects/<slug>repo and either flavour of worktree all normalise onto the repo-relative paths that repo token's git commands use.
conversations/scaffold.py— the completion scaffold (GET /api/conversations/{id}/scaffold): a conversation reduced to markdown — metadata, its prompts, merged bash/tool/subagent summaries, and one squashed diff across every commit it produced (viagitdiff) — with## Summaryand## Analysisleft blank for the model.PUT …/summarypublishes the filled version ontometa.summary. The CLI wrapper isscripts/conv-scaffold.sh.accounts/profiles.py— the Claude account profiles (personal,work), hard-coded, and the resolution order that gives every conversation ameta.account. See "Claude accounts" below.models/catalog.py— the models a session can run on (GET /api/models). Fetched from the Anthropic API (/v1/models, same key the token counter uses) and cached in-process for 6h, then reduced to the newest of each family (opus,sonnet,haiku,fable— older snapshots are dropped, so the tag row stays four chips) with the CLI alias attached. A newly released model appears in the composer with no code change. Falls back to a small static list when the key is missing or the call fails.models/openrouter.py— the OpenRouter catalogue (GET /api/models/openrouter): every model a pi.dev session can run on, with its $/Mtok (input / output / cache-read), context window, parameter size and capability tags. The data is APPE's (appe.dev.gabvdl.xyz/api/models/openrouter.json, a daily models.dev sync), cached in-process for 6h with a disk copy under/dataso a restart with no network still lists models. Deliberately separate from/api/models: that endpoint is the composer's short chip list, this is the few hundred rows the model browser searches. It also prices pi sessions —conversations.rates_for()reads a model's real rates from here instead of the flat qwen estimate, and only an id the catalogue doesn't know falls back.notifications/audio.py— the sound a notification made, replayable in the browser (POST /api/notifications/audio→ a WAV, behind the ▶ button on an expanded notification card). There is no synthesizer here: it POSTs to the homelab phone service's/render(services/phone/bridge/), which composes exactly what it would have dialled — the per-type jazz jingle, then the spoken line in the cloned French voice — so the card and the handset can never drift apart. Results are memoized per(spoken, type).--spokenonly became mandatory recently, so a push without one falls back to reading its own title + message, minus the screen-only parts (notify.sh's cost footer, URLs, emoji).settings/ui_state.py— server-side store (/data/ui-state.json) for the PWA's small client state (Settings config + notification feed seen/new bookkeeping), moved off browser localStorage so it follows the user across devices. Dumb key→string map: the key is a Zustandpersiststore name, the value its opaque serialized blob. Served byGET/PUT/DELETE /api/ui-state/{key}; the frontend points itspersiststores at it viafrontend/src/lib/serverStorage.ts.settings/mail_trigger.py— server-side store (/data/mail-trigger-senders.json) for Settings → Mail trigger (/settings/mail-trigger): emails allowed to start a conversation by mailing claude@gabvdl.xyz, on top of the mail service's ownMAIL_TRIGGER_SENDERSenv-var seed. Served byGET/PUT /api/mail-trigger-sendersas two lists —emails(allowed) anddisabled(kept in the list, switched off by the card's toggle; the mail service only ever readsemails). The homelab'sservices/mailcontainer mounts this same/datadirectory read-only and hashes the file every poll cycle — seeservices/mail/CLAUDE.md→ "Authorized senders" for the other half.projects/catalog.py/services/catalog.py— projects gallery / services catalog (static parsing);projects/costs.pythe per-project spend rollup.dashboards/skills.py— the skills catalog behind/api/skills(the/skillspage): a disk scan of every.claude/skills/<name>/SKILL.md(repo + each project's) joined with the transcript-mined usage./api/skillsreturns the union of disk skills and invoked-skill names, so a built-in/plugin skill with noSKILL.mdhere still shows up (sourceKind: "builtin", no editor link). Per skill: calls, distinct conversations, last-used, context tokens/cost (loading its dir once),estSpend= context × calls, and the coarsersessionCost(whole spend of every conversation that invoked it). Usage is keyed by bare skill name — the identitySkill(name)itself uses — so a project-local skill sharing a repo skill's name shares its counts.agents/catalog.py— the agents catalog behind/api/agents(the/agentspage), the same page shape as skills for.claude/agents/**/*.md: a disk scan (frontmattername/description/tools+ theharness/model/effortrun-config keys) unioned with the agent types the transcripts saw run, so a built-in type (Explore,general-purpose) shows up assourceKind: "builtin". The money here is measured, not estimated: an agent run is a whole conversation, soagents.runs._build_agent_runsjoins two origins into one run list — a subagent run (a sidechain transcript whoseagent-*.meta.jsonnames theagentType, credited to its parent conversation) and a session run (cron.pyfires a definition as a top-levelclaude -p; the agent page's "Run now" does the same and stampsagentRuninto the metadata sidecar).POST /api/agents/{name}/runis that button; scheduling reusesPOST /api/cronwith the definition as the prompt file.dashboards/activity.py— the activity dashboard (/api/activity, the/dashboards/activitypage): when the agents ran, as opposed to what they cost. Three rollups in one pass over the stored summaries — per local day (the contribution graph), per 5-minute slot per model (the day histogram, as positional[slot, modelIndex, agents]triples), and per model (runs / working time / tokens / cost). The per-turn half is already done by the parser: each summary carriesactiveBins(model id → the 5-minute bins that run produced a turn in), so a rollup never re-reads a transcript. Bins are epoch-absolute UTC; every entry point takes atzOffset(minutes east of UTC, straight from the browser) and buckets local days against it — the one representation that survives a DST change. An agent run is one transcript, conversation or subagent alike; the heatmap's intensity is deliberately conversations only, while the histogram and per-model stats count every run.goals/goalmd.py— a dir'sGOAL.md(the north star thegoal-keepercron agent pushes forward), shared by both catalogs:goal_summary()returns the{done, total}checklist rollup that tags a card,goal_detail()adds the markdown for the detail page. Counting skips fenced code blocks, so a- [ ]inside a snippet isn't mistaken for a real task. No GOAL.md ⇒goal: null(a valid state).memories/store.py— surfaces the assistant's per-repo memory files (frontmatter).templates/catalog.py— browse~/projects/templates/scaffolds.schemas/— Pydantic response models mirroringfrontend/src/types.ts. They drive the OpenAPI schema (and thus the generated frontend types) but are documentation-only: attached to routes viaresponses={200: {"model": …}}(helper_r()incore/http.py), neverresponse_model, so FastAPI serves them in/openapi.jsonwithout validating or filtering the handlers' actual dicts. One module per API area (conversations,projects,cron,agents,notifications,misc, …) on a sharedbase.Schema;__init__.pyre-exports every model, so handlers keep writingschemas.ConversationDetail. Modules import across each other with relative imports and must stay cycle-free —PlanSummarylives inconversations.py(notplans.py) for that reason.
API types are generated (orval)
The frontend's API DTOs are generated from the backend's OpenAPI spec, not
hand-maintained. Flow: backend/schemas/ → frontend/openapi.json (a dumped
spec, committed) → orval → frontend/src/generated/ (committed) → re-exported by
frontend/src/types.ts under their original names. types.ts keeps only the
non-API helpers (TreeNode, FileKind, ServiceAuth).
-
After changing a backend response shape, update the matching model in
schemas/, re-dumpfrontend/openapi.jsonfrom the app'sapp.openapi(), thencd frontend && npm run gen:apito regenerate, andnpx tsc --noEmitto surface fallout. Generated optionals areT | null(PydanticOptional), so widen consumers to acceptnull— the app runs unchanged; onlytsccares. -
The Docker build doesn't regenerate (no backend at build time) — it relies on the committed
src/generated/tree, so commit regenerated output. -
There's no venv on the host, so re-dump the spec with the service's own image:
docker run --rm --entrypoint python -v "$PWD:/svc" -w /svc/backend \ -e DB_PATH=/tmp/x.db -e TRANSCRIPTS_DIR=/tmp -e WORKSPACE=/tmp -e REPO_DIR=/tmp \ homelab-ai-agent -c \ "import json, os, main; json.dump(main.app.openapi(), open('/svc/frontend/openapi.json','w'), indent=2); os._exit(0)"Both odd-looking bits are load-bearing:
--entrypoint pythonbecause the image's entrypoint (docker-entrypoint.sh) otherwise ignores the command and boots the server, andos._exit(0)because importingmainstarts the indexer/watcher threads, so a normal exit hangs waiting on them.It runs as root, so delete the
backend/__pycache__it leaves behind (a root-owned dir will otherwise blockgit worktree remove).
The mock backend (frontend/src/mock/) — demo + test without the API
The frontend ships an in-browser mock backend that sits behind those same
generated DTOs, so the PWA can be demoed and automatically tested with no
FastAPI, no sidecar and no transcripts on disk. Turn it on with ?mock=1 on any
URL (sticky until ?mock=0 — works against the deployed app too) or build/serve
with VITE_MOCK=1 (npm run dev:mock); a MOCK API badge marks the mode.
It swaps window.fetch (every /api/* route) and window.EventSource
(/api/events, same meta/transcript events as backend/events.py) for
versions served from an in-memory, localStorage-persisted database. Writes really
mutate it, and spawn/resume/interrupt actually run: a scripted turn streams
into the conversation item by item over SSE, then stamps the notification + DONE
that mark it finished.
The seed is combinatorial — the cross-product of every field the UI branches
on (conversation harness × state × lifecycle × archived, every ThreadItemKind
and tool card, every FileEntryKind, every service auth, every plan status,
every FileDiffStatus/DiffLineType) — so every visual state is reachable from
a cold load, and it is deterministic (fixed clock, no RNG). Test hooks live on
window.__mock (db(), reset(), speed(), flush()).
Two rules keep it useful: install it before the app tree is imported
(main.tsx → dynamic ./mock, then dynamic ./root; the server-backed Zustand
stores fetch at import time), and type the seed rows with the generated
models — a backend schema change then breaks the seed at tsc time instead of
letting the mock drift. Details in frontend/src/mock/README.md.
Every new piece of conversation data gets a visibility toggle
The conversation page has an eye-icon visibility popover (components/ VisibilityPopover.tsx, switches + store in lib/visibility.ts) — a scrollable,
grouped list of switches for everything the page renders: per-item metadata
(price, tokens, duration, timestamp, model, tool id, result size, diff stat),
content (thinking blocks, tool cards, run-command cards, tool output, rich
widgets, error states) and gizmos (task panel, fast-forward, animations). State
persists server-side via ui-state under the key ai-agent-visibility, so a
reading setup follows the user across devices.
When you add anything to a message, a tool card or the conversation chrome, add its switch too — a new metadata field, a new inline widget, a new floating gizmo, a new animation. It is not optional: the thread is dense, and every addition has to be something the user can turn back off.
To add one: append a VisItem to the right group in VIS_GROUPS
(lib/visibility.ts) — the VisKey union, the popover list and the persisted
state all derive from that array — then gate the render on useVis() (or
useVisible(key) for a single switch). The store records only what's hidden,
so a newly added switch starts ON for everyone with no migration.
"Waiting for feedback" — the home page's answer queue
The home page opens with the one list that is waiting on the user: every
pending form (forms.py, the ask-form skill) and every pending ask
(notify.py, the notify-done skill's tappable choice), above the Running rail
and the notification feed. lib/waiting.ts merges the two sources and joins
each question to its conversation by sessionId; the section
(conversations/components/WaitingForFeedback.tsx) renders it amber with a
spinner while that session is still running and red once it isn't, and
reports the conversation ids it showed so the Running rail below skips them.
Three things make that work, and each is easy to break:
- Both skills must stamp
sessionId. It is the only link back to a conversation — a question with no session still renders, but as an orphan card.ask.shresolves it the same waynotify.shandask-form.shdo ($CLAUDE_SESSION_ID, else conv-meta'sresolve-session.sh). - A question can be older than the loaded page. The home list is the first
~50 conversations, so
GET /api/conversations?sessions=a,b,cresolves exactly the ids a first pass missed (archived included, no paging). Never reach forlimit=0here — the full lite list is ~1.5 MB for one title. - Both kinds are answerable in place. A form flips
?form=<id>and the page's own<FormModal/>opens it full-screen (no navigation); an ask POSTs to/api/ask/{id}/answer— the very endpoint the phone's tapped button uses, so the agent's long-poll wakes either way.
An ask the backend still calls pending past its expiresAt is filtered out
client-side: notify.py only flips a lapsed ask to expired when that ask is
read, so the list would otherwise keep showing dead questions.
Answering questions inside the conversation thread
Every question kind an agent can ask has an in-thread widget, so it is
answerable from the conversation itself, not only from the home queue or the
phone. Conversation.tsx's groupThread pulls each one out of the collapsed
tool group (gated by the Form cards visibility switch):
| how the agent asks | parser → widget | answer goes to |
|---|---|---|
ask-form … (rich form: text/select/slider/file/date…) |
parseAskForm → forms/components/FormWidget.tsx |
POST /api/forms/{id}/submit |
notify-ask … / ask.sh (2–3-button choice) |
parseAsk → notifications/components/AskWidget.tsx |
POST /api/ask/{id}/answer |
a plan's ## Questions json fence |
PlanQuestionsForm (on the plan page) |
the plan's form |
AskWidget resolves the ask record from the ASK_ID= stdout (a
--background ask) or the id on a notify-ask wait <id> call (Bash or
Monitor); a blocking ask with no id is matched in the notify log by question
text + nearest time. Its query key sits under QK.notifyLog, so the SSE
notification event refreshes the card the instant any surface answers.
AskUserQuestion stays a read-only card: it is disabled in spawns, and a
Claude Code process has no channel to take its answer from the viewer.
Deep-linking a prompt into the composer (?q= / ?qt=)
Any page URL can arrive with a prompt already written for it:
https://ai-agent.lab.gabvdl.xyz/?q=Deploy%20the%20lab&qt=project:homelab,skill:open-pr
https://ai-agent.lab.gabvdl.xyz/files?q=Summarise%20today&qt=radar
q is the prompt text, qt a CSV of tags to switch on — a full chip id
(project:homelab, service:gitea, skill:open-pr, a guideline's id) or the
bare slug (homelab, open-pr), matched case-insensitively, because a link
written by hand shouldn't have to know which namespace a name lives in. Either
param works alone. A selector that matches nothing is ignored, so a link naming
a since-renamed project still delivers its prompt.
Nothing is sent: the draft sits in the composer for a read-over and a tap on send. A URL that spawned outright would be a GET that starts an agent run — one stale bookmark or link preview away from a session nobody asked for.
Three things make it work, in business/composer/useSpawnUrlPrompt.ts:
- It is read in
GlobalSpawn, not in the composer. The composer is unmounted on half the pages a link can point at (desktop-docked shows only a launcher button off the home feed);GlobalSpawnis the one always-mounted host. It callsopenWith, which surfaces the composer as well as filling it. - The params are stripped afterwards (
setSearchParams(…, {replace: true}), the same one-shot treatment?form=gets), so a refresh or a Back/Forward can't re-stomp what has since been typed. - A deep link overwrites a draft in progress; the in-app buttons don't. That
is the
replaceflag onSpawnPrefill(lib/spawnDock.ts) — the Goals checklist's Work button and the Plan page's Implement button still stand down when something is already typed. Following a link with a prompt in it is a deliberate act, and silently dropping it is the worse surprise.
The tag half is the fiddly part. RichInput owns its selection — there is no
setTags on RichInputHandle — so a preselection can only get in as defaultOn
flags on the tags themselves (withPreselect, lib/richComposer.tsx). Hence the
two-step consume in SpawnComposer: one commit to put the flags on
composerTags, the next to clear() the composer onto them. That clear() is
load-bearing twice over — it re-seeds the selection from the fresh defaultOn
flags, and it drops the cacheKey="spawn" draft, whose own remembered tag ids
RichInput otherwise re-applies over the selection every time the tag list grows
(and it grows: locations arrive from /api/projects after the first paint).
useSpawnUrlPrompt.test.tsx pins all three cases — preselected on, un-named tag
left off, cached draft overwritten.
Fixed lists are rearrangeable: HoldEditable + useListOrder
Small lists of links, chips or cards are user-arrangeable by holding one for
1.4s: it lifts out of the flow and follows the pointer, the rest of the group
starts jumping (iOS-springboard style), dragging into a neighbour's slot hands
it over, and releasing drops it there and ends edit mode. All of it lives in
technical/ui/HoldEditable.tsx — pointer-events only (one code path for mouse
and touch) and a body-level portal for the lifted item so nothing clips it. It's
layout-agnostic: the group's flex/grid classes come from the caller, so the same
component drives a horizontal navbar and a vertical panel (one row or column —
not a wrapping grid).
The DOM order never changes during a drag. Slots are measured once at pickup and the rearrangement is expressed purely as transforms; only the drop commits a real reorder, once the pointer is gone. This is load-bearing, not tidiness: reordering live means React moves the pressed node — and re-renders whichever item now sits in a position-dependent slot, like the navbar's raised centre tab — and a browser cancels the touch whose target left the document, so the drag died on the first hand-over. For the same reason the held item's (invisible) in-slot copy keeps rendering with the props it had before the pickup: its node must not be replaced. Freezing the DOM also fixes the geometry, so hit-testing needs no re-measuring and no animation lock, and a fast drag can't outrun the shuffle.
const [tabs, setTabs] = useListOrder("nav", TABS, tabKey); // lib/listOrder.ts
<HoldEditable items={tabs} getKey={tabKey} onReorder={setTabs} className="flex …">
{(tab, { index, held, editing }) => <Tab …/>}
</HoldEditable>
For a list that is the data (prompt shortcuts), pass the array and write it
back. For a list hard-coded in the app, useListOrder(name, items, getKey)
persists only the order, as ids, under settings.listOrders[name], and
reconciles it against the code list on read — so adding, renaming or removing an
entry in a later release can never strand or duplicate an item. Already wired:
the bottom navbar (nav), the dashboards / settings / files drawer panels, and
the prompt-shortcut cards.
A hold ends in a click, so HoldEditable swallows the click that closes a drag —
otherwise rearranging the navbar would also navigate. Presses that start inside
an input/textarea/select/[contenteditable] (or anything marked
data-hold-editable-ignore) never pick up.
Travel during the hold cancels it only for a finger (>12px), where it means "I'm scrolling, not holding". A mouse is deliberately exempt: it can't scroll with the button down, and a hand resting on one drifts far more than that over 1.4s — cancelling on drift made hold-to-drag impossible on desktop. The pickup then happens wherever the cursor ended up, not where it went down.
Settings lists share one card: technical/ItemCard.tsx
Every settings list — Cron jobs, Workers, Webhooks, Mail trigger, Claude
accounts (each a page of its own under /settings/<name>, routed in root.tsx
and listed in nav.ts + the drawer's SettingsPanel) — renders its items with
the same ItemCard: a collapsed header row (chevron · icon · title · chips ·
optional aside · optional toggle) that expands on tap into a body holding the
editable fields, the actions and the item's detail. The conventions, so a new
list looks like the others without re-deriving them:
- The toggle is in the header and applies immediately (optimistic cache
flip + rollback on error, like
optimisticCronEnabled), independent of any unsaved edits in the body. What it means is per list: cron = fires or not, worker = routing, webhook = forwarded to or not, mail sender = allowed or not (kept in adisabledlist the mail service never reads). Accounts have none. - Edits are explicit: fields in the body hold a draft and a
SaveButtonappears only while dirty — no autosave on blur. - Delete lives at the bottom of the open body (
BTN_DANGER, pushed right) behinduseConfirm, never in the collapsed row. - Add is the dashed
AddButtonunder the list, which swaps for aNewItemFrameform;EmptyNotice/ListLoading/ListError/PageIntro/ListPageBodyare the page-level pieces. - Inputs and buttons use the exported class strings (
INPUT,INPUT_MONO,SELECT,BTN_PRIMARY,BTN_ACTION,BTN_GHOST,BTN_DANGER) andFieldfor a labelled control, so the bodies match without copy-pasting Tailwind.SettingRowis the title + help + control row the plain Settings page (display, avatar) and the account cards share.
/settings#webhooks, #mail-trigger and #subscription were the old in-page
anchors; Settings.tsx forwards them (with ?login=) to the new pages.
Performance: what the hot paths cost, and the rules that keep them cheap
This service polls the filesystem continuously and serves lists built from ~900 transcripts, so the naive version of almost anything here is quadratic. Seven rules, each of which was once a real regression:
-
Never walk the workspace (or the transcript dirs) per request or per tick. The file catalog comes from the
filestable; the discovery walk is cached in the indexer (_exposed_files, refreshed everyWALK_INTERVAL_SECS, forced on save). The watcher tells the indexer which transcripts changed (scan_transcripts(only=…)), so a live turn re-parses one file instead of re-walking ~2 600. Walk withfsscan.walk_files, neverrglob. -
Never re-parse a whole transcript to see what was appended. A live session writes every couple of seconds; the indexer keeps a
ParserStateper growing transcript and feeds it only the new lines (byte offset inIndexer._live). Re-parsing from byte 0 each tick is O(n²) over a session. -
Never shell out to git per item. Both catalogs share one HEAD-keyed history walk (
githist). -
Never send content you don't need.
/api/bundleis metadata-only (content loads per-file via/api/file);/api/conversationsis paginated + lite (usage/byToolSubonly underfull=1, which only the analytics dashboard asks for); an SSEtranscriptping patches one row via/api/conversation-summaryinstead of refetching the list. -
Memoize whole-catalog builds on the version key, through
_Memo. Every SSE-driven refetch lands right after a transcript write, i.e. right after asummaries_versionbump — so a hand-rolled check-then-build cache misses in every concurrent request at once and each one rebuilds the same thing (N open clients = N× the CPU, all on one GIL)._Memo.get(key)is single-flight: the rest wait for the first build. Cards, agent runs and agents-by-conversation go through it; a card'sstatealso depends on the clock, so the cards key carries a 15 s time bucket. The memoized list is shared — copy before sorting it. -
Keep per-summary helpers O(1). Anything called once per summary per build runs ~2 000× a request:
pathlib.relative_towas 0.24 ms a call and alone made_conv_idhalf the list's cost — the path helpers arelru_cached.MetaStore.peekis lock-free (entries are replaced, never mutated) and.items()is the no-copy sweep;.all()deep-copies 1.4 MB. -
Don't let FastAPI encode big bodies on the event loop. A handler returning a dict is serialized by pure-Python
jsonable_encoderon the loop thread (5–10× slower thanjson.dumps), and while it runs nothing is served,/api/healthincluded. Heavy GETs wear@_json_in_worker, which returns a readyJSONResponsefrom the worker thread. When the loop is starved the symptom is/api/healthtaking seconds.
Re-measure with scripts/bench-lists.sh [backend-dir] (a data snapshot in a
throwaway --network none container; pass a worktree's backend/ to compare
against the deployed image). On 2 210 transcripts, one transcript write then
cost the list 526 ms → 111 ms, the per-row summary 312 ms → 47 ms, eight
clients refetching together 7.1 s → 0.18 s, and an unchanged re-request
214 ms → 1 ms. Before that fix the live box sat at a pinned core with
/api/health at 43 s and the list at 80 s: requests queued past the client's
patience and every refetch added more.
Measured on the real archive (898 transcripts / 458 MB): /api/bundle 12.5 s →
26 ms, /api/services 11.2 s → <1 s cold / 0.5 s warm, /api/projects 3.4 s →
0.35 s, the conversation list 1.3 MB → 89 KB, and idle CPU ~58% → ~7% while a
session streams. If you add a feature here, check it against the four rules — and
re-measure, don't assume.
Security: the trusted-caller gate
The backend has no auth of its own — Authelia guards it at Traefik. But the
container also sits on the shared main docker network, where a PUT /api/file
into .claude/ (hooks!) or a POST /api/spawn is host-level code execution.
So core/auth.py installs a middleware that answers /api/* only for: Traefik (resolved
from TRAEFIK_HOST), the docker gateway (how host-originated connections to the
published 127.0.0.1:8096 port appear), localhost, or a caller presenting
INTERNAL_API_TOKEN. Everything else gets 403; /api/health stays open for the
deploy probe.
Read-only API keys (READONLY_API_KEY, comma-separated) are the one
credential that is checked before the trusted-IP path: a caller presenting a
matching X-API-Key is allowed safe methods only (GET/HEAD/OPTIONS) and any
mutating method is 403 — even from an otherwise-trusted source IP. That downgrade
is the point: the desk phone reaches the backend over the trusted host
loopback but must only ever read (it fetches GET /api/notifications/unread to
read notifications aloud), never spawn a session or PUT into .claude/. The
unread endpoint is a pure read — it never advances the read ledger
(notif_read.py / /data/notif-read.json); only the PWA opening/dismissing a
card (or POST /api/notifications/seen) marks notifications read. Assets served from repo/user bytes (/api/data-asset,
/api/upload-file, …) go out nosniff and, unless they're an allow-listed
raster image or video, Content-Disposition: attachment — an uploaded
.html/.svg must never execute on the app's origin.
The runner: same sidecar, two places it can run
sidecar/sidecar.py is the only thing that launches claude -p. It runs in one
of two places, and the backend doesn't know the difference — it just POSTs to
SIDECAR_URL:
- On the homelab host (the default, below): a run then has the host user's own auth, hooks, skills and CLAUDE.md — i.e. it is exactly a session started from a terminal. This is why the homelab keeps it.
- Inside the container (
RUNNER_IN_CONTAINER=1, the standalone default): the image bakes in the Claude Code CLI (a self-contained native binary — see the Dockerfile) anddocker-entrypoint.shstarts the same sidecar module on127.0.0.1:8790, mints aSIDECAR_TOKENif none was given, and overridesSIDECAR_URLto point at it. The CLI's$HOMEisRUNNER_HOME(/data/home, on the data volume), so its config, credentials and the transcripts it writes survive a restart;RUNNER_TRANSCRIPTS_DIR(defaulted from it) is added to the backend's liveSOURCE_DIRS, which is what makes an in-container session stream into the viewer like any other. Auth isANTHROPIC_API_KEY, or an already-logged-in Claude home mounted atRUNNER_HOME(-v ~/.claude:/data/home/.claude). Remote control is off by default there (it needs an interactive login).
That's GOAL.md's "single Docker container": a standalone image spawns sessions
with no host process. Keep it that way — a new runner feature belongs in
sidecar.py, not in a host script.
The host sidecar (what the homelab runs)
sidecar/sidecar.py runs natively on the host as the repo owner and the
backend proxies to it over host.docker.internal:8790 (Bearer SIDECAR_TOKEN):
POST /api/spawn→ sidecar/spawn:claude -p <prompt> --session-id <uuid> --model <model> --remote-control(detached). Transcript lands in the watched projects dir and the new conversation shows up in the viewer within ~2s.POST /api/resume/POST /api/interrupt→ continue / SIGINT an existing run.
model is the claude --model argument the composer's model tag carries — a
family alias (opus) for a family's newest model, or a pinned id
(claude-opus-4-5-20251101) for an older one. Unset on spawn ⇒ SIDECAR_MODEL
(default opus); unset on resume ⇒ no --model flag, so the session keeps its
own. The sidecar only shape-checks the value (it goes on a command line) — the
pickable set is whatever /api/models returned.
Background work in a headless run (stop guard, wait ceiling, inbox)
A claude -p process ends when the model ends its turn — and the model
routinely ends it with "I'll pick this up when the build / subagent reports
back". What happens to that background work then (measured on CLI 2.1.280,
sidecar/stop_guard.py has the details):
| background task | the CLI on exit | the sidecar's answer |
|---|---|---|
Bash run_in_background (shell) |
killed at once — local_bash tasks are excluded from the exit wait |
stop_guard.py, a Stop hook injected via --settings on every run (spawn / resume / fork): while a running plain shell is left, it holds the stop — polling the task's tasks/<id>.output for the CLI's [exited with code N] marker, no API call — then returns block with "your task finished, its output is at …". Monitors are excluded (the CLI waits for those). Budget STOP_GUARD_MAX_WAIT_S (55 min) per hold, then the model is told what's still running and asked to TaskStop it or end its turn again. Hook timeout SIDECAR_STOP_GUARD_TIMEOUT_S (3600); SIDECAR_STOP_GUARD=0 disables. Its lines: sidecar/logs/<sid>.guard.log (hook stderr never reaches the run log); the hold renders as a compact held tag (kind: "held"). |
| background subagent, Monitor | waited for, but only CLAUDE_CODE_PRINT_BG_WAIT_CEILING_MS (CLI default 600000 = 10 min — "Background tasks still running after 600s; terminating" in the run log) |
exported into every run from SIDECAR_BG_WAIT_CEILING_MS (default 7200000, 2 h; 0 = indefinitely, "" = the CLI default) |
This is what made a conversation "stuck": the run had exited, nothing could
report, and the next --resume opened with a didn't finish before the
previous session ended notification. 44 such events in 26 conversations
(Sep–Oct 2026), every one a clean exit.
Follow-ups into a live run (/message). Now that runs legitimately stay
alive for an hour, a "Up?" must not kill them. _send_message first POSTs the
prompt to the sidecar's /message, which writes it into the run's
cross-session inbox socket (the CLI's <config dir>/sessions/<pid>.json →
messagingSocketPath; runs are launched with crossSessionInbound: accept in
the same --settings). The model reads it between tool calls, or at once when
idle-waiting, and the background work survives. The CLI records it as a peer
message ("Another Claude session sent a message … not typed by your user"), so
the sidecar prefixes the text (INBOX_PREFIX) and conversations._unwrap_inbox
shows the plain user turn. Only a 404 (nothing live ⇒ plain resume) or a 409
inbox_unavailable (a pi run, an older CLI ⇒ the old stop-then---resume)
falls back. A model / thinking / effort switch does not apply through the
inbox; the Stop button then a resume does.
One more hole the tests found: the socket accepts and drops a message while
the session sits in a Stop hook — i.e. exactly during a guard hold. So the
hook writes logs/<sid>.hold for the hold's duration, /message checks it
first and, when present, leaves the text in logs/<sid>.inbox/
(delivered: "hold"); the hook polls that mailbox every poll, ends the hold at
once and relays the text in its block reason between <<<USER MESSAGE>>> /
<<<END USER MESSAGE>>> fences (conversations._relayed_message shows it as
the user's turn), telling the model what is still running and to end its turn
again afterwards so the hold resumes.
Model tags
Models are a first-class tag on both sides of the app (frontend/src/lib/models.tsx):
picking one in a composer (a mutually-exclusive chip per model — RichInput's
exclusive key, added in @gabvdl/ui ≥0.11) runs that session on it and the
toolbar echoes the choice; the pick is sticky (spawnModel in Settings, stored
server-side). Reading back, a conversation is tagged with every model that
actually produced a turn — conversations/parse.py collects them (models, busiest
first; a session can switch mid-thread), so past conversations are tagged too
(bump PARSER_VERSION to force the re-parse that back-fills a new summary field).
Subagent models are appended after the thread's own (_with_child_models in
conversations/tree.py), so models[0] stays the main thread's busiest.
Switching mid-thread. model is the first model a thread ran on; the one
it is on now is contextModel (the last turn's) — that is what the resume
composer, header and fork menu seed from. A switch is sticky (an untouched
resume sends no --model and the CLI keeps the session's last one). Sending a
switch goes through a confirm (lib/modelSwitch.ts → useConfirm, which needs
the ModalProvider in root.tsx): a new model re-reads the whole context
uncached (priced as a cache write on it), and claude ⇄ pi starts blank —
the two harnesses keep separate session stores. The thread draws a "switched
a → b" divider above the message whose answer came from a new model.
The tag registry (Settings → Tags)
Tags are derived, never authored (frontend/src/lib/tags.ts): useTags()
aggregates the projects/services on disk, the stack + package.json keywords a
project declares, the skills/agents catalogs, and every mention carried by a
conversation, memory, plan or project-local skill into one list of
<kind>:<slug> tags with per-source counts. Nothing creates, renames or deletes
a tag — rename a directory and the tag follows; a tag that is referenced but has
no directory is flagged exists: false rather than hidden, since that is usually
a rename something still points at.
Two axes: where work happened (project/service/stack) and what it was
done with (skill/agent). The tool axis is mined from the transcripts, not
declared — a conversation card carries skillsUsed (the Skill(…) calls, its
subagents' included, since a child transcript is never listed on its own) and
agentsUsed (built in conversations.catalog._agents_by_conversation from the same run list the
agents catalog counts, so a card can't disagree with the agent page: a subagent
run keys to its parent conversation, a cron/manual agent session to the
conversation it is). Population is "everything on disk + the built-ins that
actually ran", straight off the two catalogs, so a built-in skill/agent is a tag
with exists: false and no file to open. Those pills render on the conversation
detail only (SkillTags/AgentTags in ConvMetaBits, capped at 5 with an
expanding +N) — the list cards stay about location.
What a human does own is a tag's presentation, stored per id in the settings
store's tagMeta bag (server-side, like the rest of Settings): colour (a key
into TAG_COLORS, lib/tagStyles.ts), icon (a key into the curated
TAG_ICONS set, lib/tagIcons.tsx) and description. Unset fields fall back
to the tag kind's shipped look — projects primary/Tag, services sky/Server,
stacks violet/Blocks, skills amber/Wrench, agents emerald/Bot — which is
why an untouched homelab looks exactly as it did before the page existed.
useTagPresentation() is the queryless resolver every surface uses, so a tag
looks the same on a conversation card (ConvMetaBits), in a memory/plan list,
and as a composer chip (lib/richComposer.tsx). The description matters most in
the composer: it's the text that goes out in the prompt's Work in the following homelab location(s): bullet, so editing it here changes what a spawned agent
actually reads — the Tags editor previews that exact line.
Skill/agent tags are composer chips too (toolMentions/toolTags, offered in
both the spawn and resume composers, recently-used first). They deliberately do
not join the location block: selecting one adds a Use the following in this task: bullet naming the skill/subagent and its description — the difference
between a session invoking open-pr and one hand-rolling a PR body.
The icon set is curated rather than lucide's full icons barrel on purpose: tag
glyphs render synchronously all over the app, so they can't come from a lazy
chunk, and pulling ~1500 icons into the main bundle to serve one settings page
is a bad trade.
Picking a pi model: the OpenRouter browser
The model select lists the curated entries per harness, but on pi the real
catalogue is a modal: Browse OpenRouter… opens OpenRouterBrowser — a
FuzzyList over every OpenRouter model, each row stating the three rates that
bill an agent run (in / out / cache-read, $ per million tokens), the model's
parameter size and its context window, with capability chips (tools, reasoning,
vision, open weights, free) to narrow the list. Cheapest first, so an unfiltered
open lands on what costs least to try.
Picking one sends its full id (openai/gpt-oss-20b) as --model; the runner
already routes any vendor-prefixed id to OpenRouter and bare ids to the local
EVOX2 provider, so no allow-list has to be kept in step. The frontend reads the
same shape: isOpenRouterId() (a slash — Claude ids never have one) is what
makes an arbitrary pick resolve to the pi harness on resume, without the
catalogue being loaded.
Effort (claude) vs. thinking (pi)
How hard a run works a turn is one knob shown two ways, per harness — never both at once:
- claude gets an
EffortSelect(--effort low|medium|high|xhigh|max), the CLI's own scale. Sticky asspawnEffort;""means "send no flag" and let the CLI default. The sidecar validates against an allow-list (EFFORT_LEVELS), unlike--model, because the scale is a closed set. - pi keeps the binary
ThinkingToggle— it has no--effort._effort_argsemits nothing for it, so a stale pick can't leak onto a pi command line.
The composers send only the active harness's control (the other goes out unset).
Reading back, conversations/parse.py mines the effort field the CLI stamps on each
assistant record (it rides on the record, not the message) into efforts,
busiest first — same shape as models, rendered by EffortTags on conversation
rows and in the detail header. Empty for pi runs and for transcripts predating
the flag, which is a valid state, not a blank chip.
Claude accounts: which subscription a run bills
The host runs Claude Code under two logins — personal (the CLI's default
~/.claude) and work (the Orus seat, ~/.claude-work). An account is a
config dir: CLAUDE_CONFIG_DIR moves the login, .claude.json, settings and
the projects/ transcripts together. The profiles are hard-coded in
backend/accounts.py (the backend must decide an account for its own
server-side totals, so it can't be a client preference); Settings → Claude
accounts owns only the presentation — plan, monthly price, start date, and an
in-totals toggle (settings.accounts[id], settings v2).
- Launch (
sidecar/routes_runs.py):accounton/spawn,/resume,/fork. A non-default account's dir (fromSIDECAR_ACCOUNTS=work=~/.claude-work, written into the unit byinstall.sh) is exported asCLAUDE_CONFIG_DIR; the default account gets the variable removed — pointing it at~/.claudewould move.claude.jsoninto the dir and log the account out. With accounts declared,ANTHROPIC_API_KEYis stripped from every run (it would silently bill the API key instead of either seat), a work run whose dir has no.credentials.jsonis a 409, never a fallback, and--remote-controlis default-account only (a work run would list itself in the employer's claude.ai sessions).GET /accounts=claude auth statusper account (cached 60s). - Log in / out (
sidecar/claude_cli.py, Settings → Claude accounts →AccountLogin): a headlessclaude auth login --claudeaiunder the account's env prints an OAuth URL (redirect = Anthropic's manual-code page) and waits on stdin; the UI shows Open sign-in page + a paste box, thecode#stategoes toPOST /api/accounts/{id}/login/codeand is piped to the CLI (the sidecar checks#statefirst).…/logoutrunsclaude auth logout— the default account asks you to type its id first (it signs every lab run out). Changes publish anaccountsSSE event.?login=<id>on/settingshighlights that account (the target of every "Log in" button). - Typed errors: sidecar errors are
{detail: {message, kind, hint?, account?}}and_sidecar_callpasses them through with the sidecar's status (401/403 → 502: the viewer reads a 401 as an Authelia logout). A launch that dies is read from the CLI's JSONresultline. The frontend'sApiError(api.ts) +lib/runErrors.tsx(RunErrorFix) show the hint and a Log in button under composer/resume/fork errors; a transcript's API-error turn (isApiErrorMessage) is taggedThreadItem.apiErrorby the parser and drawn as an error card instead of a Claude bubble. - Route (
_send_message): a new claude run bills the composer's account chip, else the account owning a tagged project (projects: ["orus-monorepo"]→ work, which also starts it in that repo), else personal. Resume and fork always keep the session's own account — its transcript only exists in that account'sprojects/. Stamped asmeta.account. Cron/agent frontmatter takesaccount:. pi runs aren't a Claude login and bill personal. - Ingest:
WORK_TRANSCRIPTS_DIR(/transcripts-work) joinsSOURCE_DIRS; the archive merges every source into one tree, so the indexer stampsmeta.accountSourceat import — the last moment the origin is known. - Attribute:
_conv_metaalways emitsaccount— explicit stamp → import source → cwd rule (…/orus-monorepo…,…/worktrees/orus-monorepo-*) → default. No migration orPARSER_VERSIONbump: the cwd is in every summary, and a subagent shares its parent's sessionId and cwd. A project-less work conversation lands on orus-monorepo, not the lab root. - Totals: work spend is list price × tokens on a seat the owner doesn't pay
for, so it is excluded from aggregated totals by default. The rollups
(
/api/activity,/api/projects,/api/skills,/api/agents) take?account=— unset = the totals (noexcludeFromTotalsaccounts),all, or an id — andbyAccountsplits activity, project costs and plan actuals. The conversation list is unfiltered by default (a list shows everything; only sums leave work out). An account-owned project card shows its own account's spend flaggedexcludeFromTotals, rather than $0. - Frontend (
lib/accounts.tsx):AccountBadge(non-default accounts only, like the harness badge), the composer'sAccountSelectchip (auto from the project tags, pinnable per send — never sticky), the drawer'sselectedAccountsfilter, theaccounttag kind, and the dashboards' Personal / Work / AllAccountSwitch(stickysettings.accountScope; unpicked = the accounts counted in totals). Client-side dashboards filter the full list bymeta.account; server-side ones pass the scope as?account=, so the two always agree.planValueprices each account against its own seat.
Workers: the runner on another machine (workers.py, remotefeed.py)
A worker is sidecar/sidecar.py running somewhere other than the lab —
a Mac, as a launchd user agent (worker/install-macos.sh, see
worker/README.md). Any Mac qualifies: Gabriel's personal MacBook (account
personal) as much as the Orus work seat (work), the one it was built for —
the org disables Claude Code's Remote Control there, and the lab's own work
seat can't do what a terminal on that Mac can (its secrets are stubs).
Sessions on a worker run natively under the Mac user: its Keychain login and
whatever its shell has (mise, gcloud, op, gh…). The installer asks for the
default cwd (the directory it is run from, confirmed with y), labels the
worker with the sidecar's default account unless WORKER_ACCOUNT says
otherwise, and turns Remote Control off only for a work account.
The hub pulls; the worker never calls the hub. That is the whole security
shape: the worker binds its Tailscale address only, holds one bearer (minted at
pairing) and no hub URL or credential. Notifications, asks and forms still
work there through the outbox (below); the other lab skills that call back
(conv-meta, complete) are not on workers yet — a worker session finishes on
its notify send --final, else the exit watcher and the recency window.
- Pairing (
workers.pair,sidecar/pairing.py): the installer writes a one-time code and prints100.80.162.92:8790/<code>/<claude login email>. Settings → Workers → Add worker takes that one string. The login is matched against the hub's accounts (the host sidecar'sclaude auth statusper account) before the code is spent — no match ⇒ 400kind: "account"and the dialog shows an account picker, code still good.POST /pairconsumes the code and answers the bearer + identity; wrong codes burn it after 10. Store:/data/workers.json(0600; the token never leaves the backend). Re-pairing needsinstall-macos.sh --pairand rotates the token. - Liveness (
WorkerMonitor, 15 s):/health+ an authenticated/sessions⇒online/offline/unauthorized(token refused), live in memory only, published as aworkersSSE event. - Transcripts (
FeedMirror, 2 s):GET /feed?since=lists the worker'sprojects/(.jsonl+ subagent.meta.json),GET /feed/file?rel=&offset=serves byte ranges. The mirror appends what grew, re-downloads what shrank, and keeps the remote mtime (a backfilled old session mustn't read as running), intoWORKERS_TRANSCRIPTS_DIR(/data/workers/transcripts) — which is just one moreSOURCE_DIRSentry, so indexer/Watcher/parser/costs/ SSE are untouched. It stampsmeta.workerSource+meta.accountSource(the worker's account) in one meta save per pass (MetaStore.stamp_many) — the first sync backfills every session the Mac ever had. - Routing (
_send_message): a new claude run goes to the composer'sworkerpick ("local"= the lab; an offline pick is a 409, never a silent reroute), else the resolved account's online worker, else the lab. It bills the worker's account, andaccountis dropped from the payload (a worker has one login: its default). Continue / fork / interrupt followmeta.worker(stamped by the hub) ormeta.workerSource(mirror) —_session_worker; a session can't move runners. pi never runs on a worker. - Takeover guard: resuming a worker session whose mirror copy grew in the
last
TERMINAL_LIVE_S(60 s) with no worker run behind it is a 409kind: "terminal_live"— a terminal on the Mac probably still holds it, and two writers garble the transcript. The ResumeBox turns it into a confirm and resends withtakeover: true. - Outbox (
sidecar/outbox.py,backend/outbox_mirror.py,cli/notify.py): thenotifyCLI — the one entry point behindnotify-done,notify-askandask-form— is stdlib Python shipped incli/and appended to every run's PATH by the sidecar, which also exportsCLAUDE_SESSION_ID. On the lab it POSTs to the hub. On a worker the sidecar exportsAI_AGENT_OUTBOX_URL+ a per-worker run token, and the CLI drops the same payload into the worker'sPOST /outbox; theOutboxMirrorthread pollsGET /outboxon every online worker (2 s), creates the record through the very functions/api/notify,/api/askand/api/formsrun (_create_notify/_create_ask/_create_form), acks it, and pushes the terminal state (answered / submitted / acted / expired / cancelled) toPOST /outbox/{id}/answer, which wakes the CLI'snotify waitlong-poll on the worker. Ids are minted by the CLI (ntf_/ask_/frm_+ 12 hex) so the viewer, the phone and the worker agree. The hub owns what the CLI used to compute: the defaulturl(_conversation_url— worker transcripts included), thenotifiedmeta stamp, andfinal→state: finished. A push's extra button (--action-cmd) answers onPOST /api/notify/{id}/actionlike an ask does. - Prompt shortcuts per runner (
lib/shortcutRunners.ts, Settings → Prompt shortcuts): a shortcut is either shared (shared: true— offered on every runner: worktree, commit & push, open PR, notify, ask + form, tasks) or scoped to one runner (runner:"local"= the lab, or a worker id — the lab's deploy / plan / test / complete; a Mac's own). The composer resolves the runner its prompt will land on (the "Runs on" pick, else the account's online worker, else the lab) and offers the shared chips first, then that runner's; switching runner seeds the newly revealed chips' default-on cascade. The resume box uses the session's own runner. The Settings page is a Shared section plus one tab per runner (Lab, each paired worker); a shortcut scoped to an unpaired worker surfaces under Lab. A config from before the field (settings v3 migration) readsSHARED_DEFAULT_IDS. - A worker's skills in the composer (
sidecar/skills_list.py,GET /api/workers/{id}/skills,useWorkerSkills): when "Runs on" resolves to a worker, the mention list is that machine's skills — its sidecar lists the default cwd's.claude/skills/then the user-level<config>/skills/(where the installer linksnotify), the hub proxies it with a 60 s cache (empty for an offline or pre-/skillsworker), and the lab's catalog and cron agents are not offered there because they don't exist on the Mac. - Exit watcher: one
/sessionsbaseline per runner (lab + each worker); an unreachable runner keeps its baseline, so a Mac falling asleep never reads as "all its runs exited". - Frontend (
lib/workers.tsx):WorkerSelectcomposer chip (only once a worker is paired; amber when auto fell back to the lab because the worker is down),WorkerBadgeon cards / meta row / resume toolbar, Settings → Workers (/settings/workers), and the homeWorkerOfflineBanner.
A cron job's run config lives in its prompt file's frontmatter
A scheduled job has no composer to pick a harness/model/effort in, so it says so
itself: the leading YAML block of .claude/agents/<job>.md carries harness
(claude|pi), model (a family alias or a full id), effort
(low|medium|high|xhigh|max, or none for no flag), thinking
(true|false) and account (personal|work — see Claude accounts). cron.run_config() reads them, _fire_cron_job passes them
to the same _send_message the composer uses, and the resolved config comes
back on the job as runConfig so Settings → Cron shows it as chips. Every
homelab job currently declares claude + sonnet + no effort.
Two deliberate properties: the keys are a subset of Claude Code's own agent
frontmatter (model means the same thing there), so a job's prompt file stays
a usable .claude/agents subagent; and an invalid value is dropped, not
raised — a firing on the defaults beats a job that stops firing, and the
missing chip in the UI is the feedback. The store's per-job model field is now
only the fallback for a file whose frontmatter sets none.
Where a cron job runs: the store's worker field
Unlike the run config, the runner is a store field (worker, set in
Settings → Cron's "Runs on" select, PUT /api/cron/{id} {"worker": …}), because
it is machine pairing, not prompt content. null = the lab's own sidecar; a
worker id pins every firing to that paired machine (e.g. a PR-watch job on the
Orus MacBook, where gh and the work login live). _fire_cron_job passes it as
an explicit runner pick ("local" when unset), so a job never auto-routes,
even with account: work in its frontmatter. Workers run claude only. The prompt
file is always read on the hub and sent as text. The run uses the worker's own
login and default cwd, and notify-done from there goes through its outbox. A
firing while the worker is offline is refused (worker_offline) and recorded as
the job's lastStatus. It never falls back to the lab. Unknown ids are rejected
when saved. The page shows one tab per runner (Homelab + each paired worker, plus
an "Unpaired worker" tab for a job whose worker is gone); ?runner=<id> selects
one.
The complete skill replaced the conversation hooks (and the CTAs)
There used to be a hook scheduler here: ten seconds after any run ended, it resumed that same conversation with a follow-up prompt. It worked, and it was the wrong shape. Two reasons, both structural:
- It cost a full transcript per firing.
--resumeis a freshclaude -pprocess, and Claude Code rebuilds that process's environment preamble (cwd, git status, date, …) from scratch; any byte of drift there — near-guaranteed on a checkout other agents are also committing to — invalidates every cache breakpoint after it. Sampled real transcripts confirmed it:cache_read_input_tokenscollapsed to the cross-session baseline andcache_creation_input_tokensre-created the entire prior conversation, every firing, regardless of the delay. A tidy-up pass routinely cost as much as the feature work it followed. ThemaxContextKcap was a bandage. - Half the work didn't need a model. "Summarise what happened" is mostly
aggregation — which prompts were given, which tools ran, what net change
landed in git — and the backend already parses transcripts and buckets
commits. Asking a model to re-read a transcript to derive that was paying
premium rates for a
GROUP BY.
The first replacement was a CTA: the same prompt file, offered as a button under the last message of a finished run. It fixed the "unasked-for" half but not the cost — pressing the button still resumed the session in a new process, so the tidy-up still paid for the whole transcript as fresh input tokens.
The current answer avoids the resume entirely: completion is a skill the
session runs on itself, before it ends. A spawn guideline tells the session to
finish with the complete skill, so the pass happens inside the still-warm
process, on the prompt cache it already built — the same work for a fraction of
the tokens. The skill waits 30 s first, so there is a window to read the finished
conversation and steer it before the pass starts.
What that leaves in this backend:
conversations/scaffold.py— the algorithmic half (conv-scaffold), unchanged: metadata, the squashed diff, the merged tool/bash/subagent summaries, two blank sections for the model. The filled document is published ontometa.summaryand rendered under the thread byConversationSummary.meta.completedAt— derived, not stamped. The completion pass signs off withCOMPLETEDinstead ofDONE, the parser already flags that (completedMarker), and_conv_metaturns it into a timestamp. Nothing has to write a ledger, and there is no button state to keep in sync. It is what puts the spinning COMPLETED seal on a conversation card.
Native Claude Code hooks (.claude/settings.json) are read-only here
(GET /api/claude-hooks, listed under Settings → Claude Code hooks): writing
that file from the container would be host-level code execution — see the
trusted-caller gate above. They are now the only hooks the viewer knows about.
Install/manage the sidecar with sidecar/install.sh (systemd user unit
claude-sidecar); config comes from the repo .env (SIDECAR_TOKEN,
SIDECAR_PORT). See sidecar/README.md.
Schema drift: the transcript format is not ours
backend/conversations/parse.py parses Claude Code's internal, undocumented JSONL
transcript format, so a CLI update can silently change what it emits. The
early-warning tool is scripts/schema-drift.py:
scripts/schema-drift.py # newest transcripts vs committed baseline
scripts/schema-drift.py --update # accept reviewed drift as the new baseline
It does two things: diffs a structural footprint (key paths + content-block
types, never values — scripts/schema-baseline.json is safe in git) of recent
~/.claude/projects transcripts against the baseline, and runs a parser
canary (the real parse_conversation cross-checked against raw record
counts: assistant turns with usage, token totals, models, timestamps). Exit 1
on drift or a failed invariant. Run it after a Claude Code update or whenever
conversations render oddly; on drift, decide whether the parser needs to handle
the new shape, then --update. Every transcript record carries the CC
version stamp — a new version alone is INFO, not failure (releases are ~daily
and almost never drift).
Subagents nest: the subagents/ dir is a tree, not a list
Claude Code writes every agent a conversation spawns — at any depth — into
the one flat <convId>/subagents/ dir, as agent-<id>.jsonl + an
agent-<id>.meta.json sidecar. A subagent's own Agent/Task calls spawn
grandchildren whose meta names the spawner in parentAgentId (absent on a
depth-1 agent) and the level in spawnDepth. conversations/tree.py keeps two parent
notions apart, and everything subagent-related goes through its helpers rather
than globbing the dir:
- root (
_root_path_of): the<convId>.jsonlthe dir belongs to. Usage rolls up there (_subagent_descendants— the list card and the detail totals count the whole tree, each transcript once),agentsUsedand the agents catalog key runs there, and the graph'srootIdis it. - direct parent (
_parent_path_of, viaparentAgentId): the transcript holding the Task card that spawned it._subagent_childrenlists a node's direct children (a root's are the metas with noparentAgentId, so archives from before nesting keep their one-level reading);_attach_subagentslinks each card to its child on every page, a subagent's included, and sendsparentConversation(direct parent +toolUseId) plusparentChain(every ancestor, root first — the breadcrumb)._sidechain_stateresolves running state up the chain: a nested agent runs while its parent's Task call is open, and that parent's own state comes from its parent, not from the metadata sidecar (a sidechain'ssessionIdis the root's, so_conv_metawould answer for the root).
An agent page reports its own spend (subagents = direct children, each
ref carrying subagentCount for what it spawned in turn); only the root rolls
the tree up. The Task card that spawned a child is still credited with the
child's whole subtree (turnCost), so the cost-growth chart stays whole on the
root page where the grandchildren have no card. The graph (/api/subagents)
orbits a nested agent around its parent agent's node (parentId is a
sidechain id there; rootId is the fallback when that node isn't drawn).
_parent_path_of is deliberately not memoised on the path: the meta sidecar
can be unwritten on the first look, and a nested agent must never get pinned
to the root.
Deploying
Never compose up --force-recreate this service — use the blue-green script,
which validates a fresh build on a standby container before touching the live one:
~/homelab/services/ai-agent/deploy.sh # the homelab shim's copy (runs from anywhere)
It builds a new image, brings up a second ai-agent-deploy standby on
homelab_main, health-gates it on /api/health, cuts Traefik over to it
(file-provider hot-reload), recreates the canonical ai-agent on the new image,
then restores routing and removes the standby — zero downtime, and if the standby
never goes healthy the live container is left untouched. An flock on
/tmp/ai-agent-deploy.lock serializes deploys; a second run aborts immediately.
deploy.override.yml is only for the standby slot (renames the container, drops
the host port, joins the external homelab_main network).
The public face: landing/ → tokunia.dev.gabvdl.xyz
The project's public name is Tokunia (Tokun·IA, from 徳用 tokuyō, "good
value for the money"). landing/ is its marketing page — a dependency-free
static site (hand-written HTML/CSS, one small JS file), not a Vite app — and
landing/deploy.sh publishes both public hosts through zipgo on raspy2:
./landing/deploy.sh # landing only → tokunia.dev.gabvdl.xyz
./landing/deploy.sh --demo # rebuild + publish the demo too
Two things about that script are load-bearing:
- The demo is this frontend built with
VITE_MOCK=1— the in-browser mock backend (see above) is what makes a public demo possible with no API exposed. Branding comes fromVITE_APP_TITLE/VITE_APP_SHORT_NAME/VITE_APP_DESCRIPTION, read invite.config.ts; unset, they keep the homelab's own title and manifest exactly as they were. - The landing rsync excludes
demo./. The two hosts are nested in zipgo's trailing-dot tree (…/dev./tokunia./and…/dev./tokunia./demo./), andzipgo deploymirrors by default — an unfiltered landing deploy would delete the demo sitting underneath it.
Numbers quoted on the page (the receipt's itemised run, the archive stats) are
real figures pulled from the live instance, and the carbon line uses the app's
own frontend/src/lib/footprint.ts factors. Keep them honest if you edit them.
Standalone mode: a missing catalog is a valid state
The homelab docker-compose.yml mounts ~8 host paths (transcripts, memories,
screenshots, services/, plans, …), but the image must also run with none of
them — that's the GOAL's "generic first-run setup" / one-liner docker run.
docker-compose.standalone.yml is that shape: a workspace + a data volume, and
nothing else required.
So every catalog degrades to an empty list when its dir is absent — never a
500. (/api/templates used to raise; it doesn't now.) A named item that isn't
there is still a 404 — it's only the missing-mount case that's benign. When you
add a catalog backed by a mount, follow that rule.
The guard is a smoke test — run it after touching mounts, env wiring or a catalog route:
scripts/standalone-smoke.sh # boot the current image on a bare workspace
scripts/standalone-smoke.sh --build # build first, then smoke
It boots the image with only a throwaway workspace + data dir (and
RUNNER_IN_CONTAINER=1), asserts every catalog endpoint and the PWA shell answer
200, that the bundled claude CLI is on PATH and its runner answers /health,
and greps the log for tracebacks (a background indexer/watcher crash won't fail a
request). Exit 1 = regressed.
Standalone spawns sessions now (the in-image runner, above). The one thing it
still can't do is build off-homelab: frontend/.npmrc pulls @gabvdl/ui from
the host-bound verdaccio, so a standalone build needs that registry — a prebuilt
image runs anywhere. Tracked in GOAL.md.
Local development
There is no lint/test pipeline. Frontend and backend are built into the image; to iterate:
- Frontend only:
cd frontend && npm install && npm run dev(Vite on:5180). It needs the backend API — point it at a running container or backend. - Full stack: rebuild + redeploy via the homelab shim's
deploy.sh, or hit the running container directly on the host athttp://127.0.0.1:8096(plain HTTP, bypasses Authelia — handy since new*.labdomains can't get an LE cert here). - Health:
curl -s http://127.0.0.1:8096/api/health→{"ok":true}.
After changing the homelab shim's traefik.yml, run make link-traefik-configs
from the homelab repo root (the file provider watches config/dynamic/, not the
shim's file).
Gotchas
- Read-only mounts: most of the repo context is mounted
:ro(services/, .git, transcripts, memories, screenshots, and the rootdata/tree). Only.claude/,~/projects(host-side;/workspace/projectsin-container),CLAUDE.mdand the service's owndata/(SQLite store) are writable. The container runs as1000:1000so edits keep repo ownership. - Root
data/tree: mounted${PWD}/data:/workspace/data:roand surfaced in the file tab for visibility only. The indexer exposes just text files (seeDATA_TEXT_EXTSincore/config.py— md/json/py/log/…; binaries like the PNG thumbnails are skipped) and marks them non-editable.data/tmp/is inIGNORE_DIRS, so scratch files and secrets (e.g.*-creds.env) never appear. - Token counts cost API calls: they hit the
count_tokensendpoint (ANTHROPIC_API_KEYfrom.env) and are recomputed only when a file's content hash changes, after a debounce — don't add code that recounts on every request. - Conversation "finished" state: three signals, strongest first. (1) A
notify-donenotification plus a last message ending with the literalDONEmarker (see the root CLAUDE.md "When work is done"). (2) The exit watcher: a background thread inruns/sidecar_client.py(_watch_sidecar_runs) polls the sidecar's/sessionsevery ~2s and stampsfinished(+exitedAt) the moment a sidecar-launched run's process dies — so runs whose prompt never said DONE don't sit on "running". The stamp is distrusted if the transcript later grows pastexitedAt(a terminal/remote-control continuation,_outlived_exit), and in-flight resume/interrupt swaps are shielded via_sidecar_busy. (3) The recency fallback (RUNNING_WINDOW_SECS, 600s) for anything the watcher never saw live (terminal sessions, runs that died while the backend was down). - Conversation "paused" state: a run whose last act was parking itself on a
trigger — a trailing
ScheduleWakeup(non-stop) or a pendingMonitor— is shownpausedinstead ofrunning(waiting_trailinginconversations/parse.py, resolved inconversations/cards.py:_resolve_state). The flag clears on any later tool call, real user turn, or interrupt — not on the same turn's closing text. - Two containers briefly share
data/during a deploy cutover (SQLite); this is a short window and intentional — don't "fix" it by locking the DB. - A resume must never answer 200 without a live run behind it. The composer
echoes the typed turn optimistically and only drops the echo once the
transcript grows, so any resume that reports success but starts nothing leaves
the UI on "sending…" forever. Two guards keep that honest and must stay:
sidecar/launch.pywatches a freshly launchedclaudeforLAUNCH_PROBE_Sand 502s with its log tail if it dies on the spot (a--resumewhose transcript isn't found exits 1 in ~2.5s), and/resumeonly reportsreusedfor a process that is genuinely still running — a finished-but-unreaped child is a zombie, and a zombie answersos.kill(pid, 0)(hence_reap()+ theZcheck in_alive).reusedmeans the turn was dropped, so the backend returns 409, not 200. The backend always resumes withforce: true: a genuinely-live run is stopped first (_stop_run— SIGINT the group, wait up toSTOP_WAIT_S, escalate to SIGKILL) and the resume then launches, so sending a message into a running or paused conversation restarts it on the new turn instead of 409ing.reusedcan now only come from a sidecar predating the flag.