Web Performance
Loading-path architecture, budgets, and verification for the Clapilot web app.
Clapilot prioritizes useful persisted data before remote synchronization and keeps optional surfaces off the critical navigation path.
Loading-path rules
- The root language provider starts from the server cookie, owns one dictionary/listener set, and prevents a German-to-selected-language hydration refetch.
- The global chat rail loads the full floating chat only when expanded or after an idle window for users who left it expanded.
- Email, Calendar, and Documents paint persisted database/cache state first. Provider synchronization is explicit or scheduled after browser idle and must not block navigation.
- Dashboard and task lists use compact task payloads and SQL aggregates; task details and attachments load from the dedicated detail route.
- Module inventory uses a short process cache and direct slug lookup. Module pages request one manifest; static module assets use ETags and private browser caching.
- Long email, task, and chat rows use
content-visibility: autoso off-screen rows do not consume initial layout/paint work. - Chat inline link previews fetch only when the preview card is near the viewport (IntersectionObserver), cap concurrent
/api/chat/link-previewrequests at four, and cache results — including failures — insessionStoragefor six hours, so a link-heavy chat history does not refire dozens of preview requests on every app start. - In the Clapilot Tab Layout, the always-visible chat dock still mounts on browser idle rather than eagerly, so page startup wins the race for network and main-thread time.
- The Next.js client router keeps recently visited pages reusable for 30 seconds (
experimental.staleTimes.dynamic), so switching between workspace tabs or using back/forward does not refetch the full RSC payload each time; pages fetch their data client-side on mount, so this does not serve stale content. - The tab layout additionally keeps the three most recently used tab views mounted (hidden) with frozen router contexts, flipping visibility optimistically on tab activation — switch-back is instant with full client state preserved, at the cost of the hidden pages' background effects staying active (bounded by the three-view cap).
Regression gates
npm run perf:bundle:gate gates on route first-load JS: for every App Router route it sums the Brotli size of the shared root main files plus the entry chunks of each segment that renders the route (layouts, error boundaries, page — read from the build's client reference manifests), and fails when the heaviest route exceeds 1100 kB Brotli. A raw per-chunk cap of 1500 kB stays as a tripwire for broken code-splitting. Lazily loaded chunks (next/dynamic, await import(), module iframes) do not count, so adding a feature behind a lazy boundary never breaks the gate; growing the app shell or a page's eager imports does. Totals over all emitted chunks (raw, gzip, Brotli) and the largest Brotli chunk are report-only: they grow with every lazily loaded feature even under healthy code-splitting, and gating them only produced recurring threshold-bump and "restore headroom" commits. Environment overrides are PERF_MAX_ROUTE_FIRST_LOAD_BROTLI_KB and PERF_MAX_CHUNK_KB. When the gate fails, move heavy client code behind a lazy boundary or out of the shared layout instead of raising the budget.
npm run e2e:local:performance covers Dashboard, Chat, Documents, Email, Tasks, Calendar, Canvas, and Agent Orchestrator. Route checks record interactive time and include request-count, transferred-byte, compact-payload, provider-sync, and overlapping-poll assertions where applicable.
Use a production build for bundle comparisons. Browser development-server measurements are useful for request fan-out and behavior, but they are not production latency SLAs.
Chat foreground deadline and 24-hour regression verification
Native /api/chat turns now stop holding the foreground SSE connection after
60 seconds measured from request entry. The deadline covers the native stream
startup and all subsequent queue, provider, tool and fallback phases; streaming
progress does not reset it. Request preparation before the stream is constructed
still needs to finish before a response can be sent. The response contains a
clapilot.handoff event with a localized background status and [DONE], while the same detached producer keeps
consuming the runtime and persisting the answer to the original pending message.
There is no new run, replay of the user request, or release of the runtime's session
queue at handoff. Existing history refresh/polling retrieves the eventual result
on web, iOS and macOS. All three clients retain pending/tool state at handoff
and suppress completion notifications and success feedback until the answer.
Startup/producer failures propagate as stream errors after persistence and clear
the handoff timer. The handoff is not a successful task completion and must
not replace the final agent_runs duration in the Performance Radar.
Interactive provider stall defaults are 30 seconds for streaming and 45 seconds
for non-streaming (including interactive channel turns); the fallback default is
20 seconds. Channel turns have no web foreground handoff, so operators must
review these budgets for channel workloads as well. Existing provider/model
metadata and environment overrides remain supported and must be checked during
rollout. The existing provider circuit breaker still governs timeout accounting,
shared endpoint failures, bounded half-open probes and independent fallbacks.
Provider-originated timeouts (including HTTP 408/504) now bypass same-provider
backoff retries and charge the circuit on the first failed attempt. Independent
fallbacks remain available when no irreversible progress makes replay unsafe.
With no fallback configured, even a transient HTTP 408/504 ends the turn after
the first failed attempt; there is no automatic same-provider replay. Providers
that need more than 30 seconds before their first SSE event (reasoning or large
self-hosted prompts) need a larger chatLatencyBudgets.providerMs or per-model
override; non-streaming slow providers likewise need a larger
nonStreamingProviderMs. The stall budget includes the wait for the first event.
Background continuation uses the existing self-hosted detached producer and
pending-message reconciliation; this is not a new durable job queue or a promise
that arbitrary in-flight tools survive process termination.
Verification after rollout
Record the deployed commit and UTC rollout timestamp. Run the following with the production read-only reporting database configuration, substituting that timestamp:
node scripts/perf-radar.mjs --days 2 --compare-cutover 2026-09-18T12:00:00Z --json --output /tmp/chat-rollout.json
node scripts/perf-radar.mjs --days 2 --gate
Collect reports at rollout and again at least 24 hours later. Inspect successful chat p90 (<90 seconds), sample count, failures, provider timeouts and wasted provider duration together. Also inspect still-running turns: excluding stalled unfinished turns or treating a background status as completed work would conceal this regression. The 60-second foreground bound does not prove a final-answer p90 below 90 seconds; keep the incident open until the post-rollout measurements meet the target. No 24-hour production result has been established by local tests.
Manual checks: use a delayed native provider and a tool taking over 60 seconds. Verify one background status, a closed foreground stream, a pending original message, a single tool mutation, and the final answer in that same message without resubmitting. Repeat with a fallback and with browser reload/disconnect. On iOS and macOS verify handoff retains pending state without success feedback or a completion notification, and history polling later shows the result. Short turns must finish without the background status. Check a degraded provider with the circuit-breaker tests and confirm production metadata has not retained the old multi-minute overrides.
