perf(backend): unstick the conversation list under live traffic #6
Reference in New Issue
Block a user
Delete Branch "perf-list-cache"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Summary
The ai-agent felt laggy because the backend was saturated, not the G10 box: one Python process was pinned at ~105% of a core while the host sat ~82% idle. Every transcript write (every ~2 s while any agent runs) invalidated the conversation-list caches, every client woken by the SSE bus missed them at the same moment and each ran the same full ~1 400-card rebuild in parallel on one GIL, and FastAPI then serialized the big bodies on the event loop. Live, before this PR:
/api/conversations?limit=50took 16–80 s and even/api/healthtook 43 s. Each request queued longer than the client waited, and every refetch added more work.Key changes
backend/main.py—_Memo: a version-keyed, single-flight memo. Concurrent callers with the same key share one build instead of each running it. It now backs the card list (_conversation_cards),_agent_runsand_agents_by_conversation.backend/main.py—lru_cacheon_conv_id/_parent_path_of/_is_sidechain_path.pathlib.relative_tocost 0.24 ms a call × ~4 000 calls per build, which was half the list's cost.backend/meta.py—MetaStore.peekis lock-free (thousands of peeks per build were convoying on the lock). The newMetaStore.items()replaces a 1.4 MB JSON deep copy in_build_agent_runs.backend/main.py—@_json_in_workeron the eleven heaviest GETs (conversations, conversation, conversation lists, notifications, subagents, forms, goals, notify-log, cron, projects, skills, activity). Each returns a readyJSONResponsefrom its worker thread, sojsonable_encoder(5–10× slower thanjson.dumps) no longer blocks the event loop.scripts/bench-lists.sh+scripts/bench_lists.py— a reproducible benchmark. It runs a snapshot of the live data in a throwaway--network none, 1-CPU container, against the deployed image or a worktree'sbackend/.CLAUDE.md— the performance section gains rules 5–7 (single-flight memo, O(1) per-summary helpers, no big encodes on the loop).Key decisions
_auto_state, 10-min window), so a version-only key could leave an idle list showing "running" forever.JSONResponsediscards a decoratorstatus_codeand any injectedResponseheaders, so it's opt-in on plain 200 GETs only. A bodyjson.dumpscan't encode falls back to FastAPI's encoder unchanged.Changelog
Test notes
scripts/bench-lists.sh(prod image) vsscripts/bench-lists.sh "$PWD/backend"on a 2 210-transcript snapshot, 1 CPU, cost after one transcript write:/api/conversations?limit=50/api/conversation-summary(SSE row patch)_build_agent_runs--network none, each on its own data copy) and fetched 15 endpoints from each. The card/list/detail/summary/notifications/forms/cron/skills/subagents/agents responses are identical once parsed. activity/projects/goals differed only by transcripts that grew live between the two runs. Full list withfull=1(4.7 MB): 1 554 → 268 ms.main(3 pre-existing import errors for the missingfeedmodule intest_workers/test_remotefeed/test_worker_routing)../services/ai-agent/deploy.shfrom~/homelab.conversations_list84% of samples;meta.peeklock wait 67%; the main thread injsonable_encoder; 39 AnyIO worker threads busy.🤖 Generated with Claude Code
Screenshots
View command line instructions
Checkout
From your project repository, check out a new branch and test the changes.