DGX Cluster Dashboard
Admin dashboard for the Spark fleet telemetry from Clapilot Cluster Studio in ClapilotAICore settings.
Settings -> ClapilotAICore -> DGX Cluster shows live telemetry for the Spark fleet: every inference cluster returned by the configured endpoint, the shared media and helper services, and the fleet's other machines. Each serving cluster has its own status, deployment metadata, metric tiles, charts, latency percentiles, and dynamic list of nodes. A partially available fleet remains visible: a cluster with nothing serving collapses to its header and node tiles (throughput, KV and latency would all be empty), while the global status reports how many clusters are live.
Upstream: Clapilot Cluster Studio
The telemetry comes from Clapilot Cluster Studio (the spark-fleet comfy-shim, media-gateway/fleet_stats.py) under /api/fleet/{stats,stream,health,schema}. The default endpoint is the Studio origin http://192.168.178.39:8010; a saved override is stored in app_settings.dgx_telemetry_api_base_url, and an empty saved value means Clapilot uses the default. The former standalone dashboard on :8081 is retired: it only redirects to Studio, and because a cross-origin redirect drops the Authorization header it cannot be used as the endpoint. Migration 336_cluster_studio_access.sql moves instances that saved http://192.168.178.39:8081 to the Studio origin.
Studio guards every /api/* route with one shared password (STUDIO_PASSWORD in ~/media-gateway/.env.studio on the head node). Clapilot stores it encrypted (enc:v1) in app_settings.cluster_studio_access_token and sends it server-side as Authorization: Bearer <password>. The password field on the page is write-only: settings payloads only report access_token_configured, an empty field keeps the stored password, and a typed value replaces it on save. Model training talks to the same Studio and uses the same stored password, so it can be set on either page. A missing or wrong password surfaces as "Password required" and a localized error instead of a generic "unreachable".
Browsers never call the LAN telemetry service directly. The browser calls admin-only Clapilot API routes, and those routes proxy to the configured LAN upstream:
- browser ->
/api/clapilotaicore/dgx-telemetry-settings - browser ->
/api/clapilotaicore/dgx-telemetry/* - Clapilot app route -> Cluster Studio
/api/fleet/*with the stored bearer password
Metric semantics:
- Schema v4 (current) returns
{ schema_version: 4, training_url, media: {...}, training: {...}, machines: [...], clusters: [...] }.cluster_idis the stable UI/history key (pair,glm53,rtx5000,rtx3090today),cluster_nameis the display label, and every cluster owns its ownhist,static,served_model, andnodesdata. nodesis dynamic and contains{ name, gpu, mem_used, mem_total, gpu_temp, gpu_power, gpu_clock, online }; Clapilot does not assume a fixed node count. GPU charts are grouped in pairs while all nodes receive their own metric tile showing GPU usage plus temperature, power draw, and memory. Nodes reportedonline: falseare highlighted as offline.- Each cluster also reports
nodes_online/nodes_total, shown as a badge in the cluster header. - The top-level
mediablock carries the shared services and is read generically, because the collector adds services over time: every entry becomes a tile withup, active work (activeor ComfyUI'srunning), queue (pendingorwaiting),done/failed,progress, and for ComfyUI boxes free VRAM (free_gb/total_gb) plusmem_state(warn/badwhen free VRAM no longer fits the working set). Known keys get a localized label and a fixed order:video,rtx_h3(MiniMax H3 API),rtx_comfy,halo_comfy,r3090_comfy,sam,embed,stt,tts; unknown keys are appended and labeled by key. The top-leveltrainingsummary from the training node becomes atrainingtile (up, the active run'spct,completed/failed). A missing block (schema v1) hides the row. - The top-level
machineslist holds the fleet's non-inference machines (Clapilot Mac minis, Raspberry Pis, media boxes) with{ id, name, group, address, hostname, os, hw, mem_used, mem_total, load_1, cpu_temp, uptime_s, online, probed }. They render as one divided list grouped bygroup;probed: falsemachines are shown as "listed only". Their state never implies that a model endpoint is serving. - The schema v1 flat snapshot remains supported as a single compatibility cluster.
nullmeans the metric has not been observed yet and is rendered as-, never as zero.- Throughput and fabric rates (
gen_tps,prompt_tps,fabric_bps) are windowed over the current upstream tick; queue depth and KV-cache usage are instantaneous gauges; speculative decoding counters accumulate since engine start. kv,accept_rate, andprefix_hitarrive already as percent values (0-100), not fractions.fabric_bpscarries RoCE bytes/s from the HCA port counters despite its name; the dashboard renders it in B/s-based units.- Latency percentiles are cumulative since engine start and can include historic outliers.
accept_lenis1 + accepted/drafts; higher values mean more useful speculative decoding.histis parsed defensively from either frame arrays or object-of-arrays payloads. If history cannot be parsed, the dashboard fills from live SSE frames only.- Schema versions
1to4are supported. If the upstream reports another version, Clapilot shows a warning but continues rendering known fields.
The read-only native agent tool dgx_cluster_stats fetches the current stats snapshot through the same server-side proxy path, strips hist from every cluster, and returns a short German per-cluster summary (including per-node GPU usage and temperature, the service states with free ComfyUI VRAM, and how many other machines are online) plus the telemetry payload with the media, training, and machines blocks. When Studio rejects the stored password it returns upstream_unauthorized and tells the user where to set it. It does not mutate UI state.
Apple client
The iOS and macOS app expose the same dashboard at Einstellungen -> DGX-Cluster for admin users. The native screen uses the configured Clapilot base URL and session cookie, calls the same backend proxy endpoints, and never connects to the LAN telemetry service directly.
The Apple client loads settings and the current stats snapshot through the proxy, streams live frames via the SSE route with automatic reconnect, and renders the same multi-cluster layout as the web panel: the Cluster Studio URL and write-only password fields, the service tiles, per-cluster headers with model and node-count badges, metric tiles including per-node GPU/temperature/power, throughput/GPU/KV/fabric charts, cumulative latency percentiles, collapsed offline clusters, and the grouped machine list. Studio password rejections arrive as 502 { error: "upstream_unauthorized" }, never as 401, so they cannot be mistaken for an expired Clapilot session.
