You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
topics_download is async def (topics_ui.py#L55) but calls csv_export.write_topic_csv synchronously on the event loop (L79). The two cost tiers differ by orders of magnitude:
Simple export: SQLite reads + one small file write — milliseconds.
Enriched export: for every title with a NULL description, fetch_descriptions_with_fallback runs — Wikidata, then per-title REST /page/summary at wikipedia_api.py#L269, each individually capped at 10s (api_get itself defaults to 30s, L61) — then results are persisted with db.set_descriptions (csv_export.py#L74-L75). The function explicitly supports a deadline cutoff for exactly this scenario (wikipedia_api.py#L278) and csv_export.py#L74 never passes it — so total blocking time scales linearly with the number of description-less titles, unbounded in aggregate.
While that loop runs, the uvicorn worker's entire event loop is blocked: every MCP session sticky-routed to that worker stalls (nginx ip_hash upstream, deploy.sh#L87-L109), as do unrelated HTTP requests. Page refreshes can launch concurrent exports of the same topic; two workers share SQLite throughout.
Reliability defect, not a measured outage — no production load test was run.
pass a deadline to fetch_descriptions_with_fallback so worst-case blocking is bounded;
reuse the already-generated file when newer than the topic's updated_at — the /topics page computes exactly that staleness for display (_download_cell) yet the download click regenerates anyway;
Audit commit:
ce6b351c. All links pinned.topics_downloadisasync def(topics_ui.py#L55) but callscsv_export.write_topic_csvsynchronously on the event loop (L79). The two cost tiers differ by orders of magnitude:fetch_descriptions_with_fallbackruns — Wikidata, then per-title REST/page/summaryat wikipedia_api.py#L269, each individually capped at 10s (api_getitself defaults to 30s, L61) — then results are persisted withdb.set_descriptions(csv_export.py#L74-L75). The function explicitly supports adeadlinecutoff for exactly this scenario (wikipedia_api.py#L278) andcsv_export.py#L74never passes it — so total blocking time scales linearly with the number of description-less titles, unbounded in aggregate.While that loop runs, the uvicorn worker's entire event loop is blocked: every MCP session sticky-routed to that worker stalls (nginx
ip_hashupstream, deploy.sh#L87-L109), as do unrelated HTTP requests. Page refreshes can launch concurrent exports of the same topic; two workers share SQLite throughout.Reliability defect, not a measured outage — no production load test was run.
Fix:
await starlette.concurrency.run_in_threadpool(csv_export.write_topic_csv, ...);deadlinetofetch_descriptions_with_fallbackso worst-case blocking is bounded;updated_at— the/topicspage computes exactly that staleness for display (_download_cell) yet the download click regenerates anyway;