perf(backend): unstick the conversation list under live traffic #6

Open
gabrielvidal wants to merge 0 commits from perf-list-cache into main
Owner

Summary

The ai-agent felt laggy because the backend was saturated, not the G10 box: one Python process was pinned at ~105% of a core while the host sat ~82% idle. Every transcript write (every ~2 s while any agent runs) invalidated the conversation-list caches, every client woken by the SSE bus missed them at the same moment and each ran the same full ~1 400-card rebuild in parallel on one GIL, and FastAPI then serialized the big bodies on the event loop. Live, before this PR: /api/conversations?limit=50 took 16–80 s and even /api/health took 43 s. Each request queued longer than the client waited, and every refetch added more work.

Key changes

  • backend/main.py — _Memo: a version-keyed, single-flight memo. Concurrent callers with the same key share one build instead of each running it. It now backs the card list (_conversation_cards), _agent_runs and _agents_by_conversation.
  • backend/main.py — lru_cache on _conv_id / _parent_path_of / _is_sidechain_path. pathlib.relative_to cost 0.24 ms a call × ~4 000 calls per build, which was half the list's cost.
  • backend/meta.py — MetaStore.peek is lock-free (thousands of peeks per build were convoying on the lock). The new MetaStore.items() replaces a 1.4 MB JSON deep copy in _build_agent_runs.
  • backend/main.py — @_json_in_worker on the eleven heaviest GETs (conversations, conversation, conversation lists, notifications, subagents, forms, goals, notify-log, cron, projects, skills, activity). Each returns a ready JSONResponse from its worker thread, so jsonable_encoder (5–10× slower than json.dumps) no longer blocks the event loop.
  • scripts/bench-lists.sh + scripts/bench_lists.py — a reproducible benchmark. It runs a snapshot of the live data in a throwaway --network none, 1-CPU container, against the deployed image or a worktree's backend/.
  • CLAUDE.md — the performance section gains rules 5–7 (single-flight memo, O(1) per-summary helpers, no big encodes on the loop).

Key decisions

  • Kept whole-list memoization and dropped per-card incremental caching. With single-flight plus the path fix, a full rebuild is about 115 ms and runs once per write, whatever the number of clients. That's good enough, and much less state to get wrong. Incremental cards are the next step if the archive grows another order of magnitude.
  • The card memo key includes a 15 s time bucket. A card's running/finished state reads the clock (_auto_state, 10-min window), so a version-only key could leave an idle list showing "running" forever.
  • Explicit decorator instead of a global route class for off-loop encoding. A returned JSONResponse discards a decorator status_code and any injected Response headers, so it's opt-in on plain 200 GETs only. A body json.dumps can't encode falls back to FastAPI's encoder unchanged.
  • Frontend unchanged. The SSE bus already debounces and throttles list refetches. The server was paying N× for them.

Changelog

  • Fixed: the conversation list, home page and dashboards no longer lag or time out while agents are running (the list answered in 16–80 s under live load)
  • Fixed: the app stays responsive during heavy refreshes (health checks and small requests no longer wait behind large list responses)

Test notes

  • scripts/bench-lists.sh (prod image) vs scripts/bench-lists.sh "$PWD/backend" on a 2 210-transcript snapshot, 1 CPU, cost after one transcript write:
    path before after
    /api/conversations?limit=50 526 ms 111 ms
    /api/conversation-summary (SSE row patch) 312 ms 47 ms
    _build_agent_runs 278 ms 28 ms
    4 clients refetching together 3.4 s 0.18 s
    8 clients refetching together 7.1 s 0.18 s
    list, nothing changed 214 ms 1 ms
  • End-to-end: ran old and new code as two isolated uvicorn containers (--network none, each on its own data copy) and fetched 15 endpoints from each. The card/list/detail/summary/notifications/forms/cron/skills/subagents/agents responses are identical once parsed. activity/projects/goals differed only by transcripts that grew live between the two runs. Full list with full=1 (4.7 MB): 1 554 → 268 ms.
  • Backend unittests: same result as main (3 pre-existing import errors for the missing feed module in test_workers/test_remotefeed/test_worker_routing).
  • Not deployed (per the PR workflow). After merge: ./services/ai-agent/deploy.sh from ~/homelab.
  • Root-cause profile (py-spy, live): conversations_list 84% of samples; meta.peek lock wait 67%; the main thread in jsonable_encoder; 39 AnyIO worker threads busy.

🤖 Generated with Claude Code

Screenshots

perf-before-after.png

## Summary The ai-agent felt laggy because the backend was saturated, not the G10 box: one Python process was pinned at ~105% of a core while the host sat ~82% idle. Every transcript write (every ~2 s while any agent runs) invalidated the conversation-list caches, every client woken by the SSE bus missed them at the same moment and each ran the same full ~1 400-card rebuild in parallel on one GIL, and FastAPI then serialized the big bodies on the event loop. Live, before this PR: `/api/conversations?limit=50` took **16–80 s** and even `/api/health` took **43 s**. Each request queued longer than the client waited, and every refetch added more work. ## Key changes - `backend/main.py` — `_Memo`: a version-keyed, **single-flight** memo. Concurrent callers with the same key share one build instead of each running it. It now backs the card list (`_conversation_cards`), `_agent_runs` and `_agents_by_conversation`. - `backend/main.py` — `lru_cache` on `_conv_id` / `_parent_path_of` / `_is_sidechain_path`. `pathlib.relative_to` cost 0.24 ms a call × ~4 000 calls per build, which was half the list's cost. - `backend/meta.py` — `MetaStore.peek` is lock-free (thousands of peeks per build were convoying on the lock). The new `MetaStore.items()` replaces a 1.4 MB JSON deep copy in `_build_agent_runs`. - `backend/main.py` — `@_json_in_worker` on the eleven heaviest GETs (conversations, conversation, conversation lists, notifications, subagents, forms, goals, notify-log, cron, projects, skills, activity). Each returns a ready `JSONResponse` from its worker thread, so `jsonable_encoder` (5–10× slower than `json.dumps`) no longer blocks the event loop. - `scripts/bench-lists.sh` + `scripts/bench_lists.py` — a reproducible benchmark. It runs a snapshot of the live data in a throwaway `--network none`, 1-CPU container, against the deployed image or a worktree's `backend/`. - `CLAUDE.md` — the performance section gains rules 5–7 (single-flight memo, O(1) per-summary helpers, no big encodes on the loop). ## Key decisions - **Kept whole-list memoization and dropped per-card incremental caching.** With single-flight plus the path fix, a full rebuild is about 115 ms and runs once per write, whatever the number of clients. That's good enough, and much less state to get wrong. Incremental cards are the next step if the archive grows another order of magnitude. - **The card memo key includes a 15 s time bucket.** A card's running/finished state reads the clock (`_auto_state`, 10-min window), so a version-only key could leave an idle list showing "running" forever. - **Explicit decorator instead of a global route class** for off-loop encoding. A returned `JSONResponse` discards a decorator `status_code` and any injected `Response` headers, so it's opt-in on plain 200 GETs only. A body `json.dumps` can't encode falls back to FastAPI's encoder unchanged. - **Frontend unchanged.** The SSE bus already debounces and throttles list refetches. The server was paying N× for them. ## Changelog - Fixed: the conversation list, home page and dashboards no longer lag or time out while agents are running (the list answered in 16–80 s under live load) - Fixed: the app stays responsive during heavy refreshes (health checks and small requests no longer wait behind large list responses) ## Test notes - `scripts/bench-lists.sh` (prod image) vs `scripts/bench-lists.sh "$PWD/backend"` on a 2 210-transcript snapshot, 1 CPU, cost after one transcript write: | path | before | after | |---|---|---| | `/api/conversations?limit=50` | 526 ms | 111 ms | | `/api/conversation-summary` (SSE row patch) | 312 ms | 47 ms | | `_build_agent_runs` | 278 ms | 28 ms | | 4 clients refetching together | 3.4 s | 0.18 s | | 8 clients refetching together | 7.1 s | 0.18 s | | list, nothing changed | 214 ms | 1 ms | - End-to-end: ran old and new code as two isolated uvicorn containers (`--network none`, each on its own data copy) and fetched 15 endpoints from each. The card/list/detail/summary/notifications/forms/cron/skills/subagents/agents responses are **identical** once parsed. activity/projects/goals differed only by transcripts that grew live between the two runs. Full list with `full=1` (4.7 MB): 1 554 → 268 ms. - Backend unittests: same result as `main` (3 pre-existing import errors for the missing `feed` module in `test_workers`/`test_remotefeed`/`test_worker_routing`). - Not deployed (per the PR workflow). After merge: `./services/ai-agent/deploy.sh` from `~/homelab`. - Root-cause profile (py-spy, live): `conversations_list` 84% of samples; `meta.peek` lock wait 67%; the main thread in `jsonable_encoder`; 39 AnyIO worker threads busy. 🤖 Generated with [Claude Code](https://claude.com/claude-code) ## Screenshots ![perf-before-after.png](https://git.gabvdl.xyz/attachments/55fb1302-e33f-4e29-a856-5204931277d8)
gabrielvidal self-assigned this 2026-09-29 19:24:00 +02:00
gabrielvidal added 2 commits 2026-09-29 19:24:00 +02:00
The list rebuilt every card after each transcript write, and every client
the SSE bus woke up missed the cache at once and ran the same full build in
parallel on one GIL; FastAPI then encoded the big bodies on the event loop.
Live, that pinned a core with /api/conversations at 16-80 s and even
/api/health at 43 s.

- _Memo: version-keyed, single-flight memo for the card list, agent runs
  and agents-by-conversation (cards also key on a 15 s time bucket, since
  a card's running/finished state reads the clock)
- lru_cache on _conv_id / _parent_path_of / _is_sidechain_path:
  pathlib.relative_to was 0.24 ms a call, half the list's cost
- MetaStore.peek no longer takes the lock (convoyed thousands of peeks
  per build); MetaStore.items() replaces the 1.4 MB deep copy in
  _build_agent_runs
- @_json_in_worker: the eleven heaviest GETs return a ready JSONResponse
  from their worker thread instead of leaving jsonable_encoder (5-10x
  slower than json.dumps) to run on the event loop

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Reproducible snapshot benchmark of the list hot paths (throwaway
--network none container, optionally against a worktree's backend), and the
three new rules in CLAUDE.md's performance section.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Checking for merge conflicts…
View command line instructions

Checkout

From your project repository, check out a new branch and test the changes.
git fetch -u origin perf-list-cache:perf-list-cache
git checkout perf-list-cache
Sign in to join this conversation.
No Reviewers
No Label
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: gabrielvidal/ai-agent#6