Live Fleet Metrics — real-time from nvidia-smi + llama-server /metrics
CONNECTING
last poll: —
net ↓ — · ↑ —
Model serving :8001
—
connecting...
Prompt throughput
—tok/s
prompt tokens: —
Decode throughput
—tok/s
predicted tokens: —
Request latency
—ms TTFT
waiting for first poll…
Total GPU power draw
—W / —W
waiting for first poll…
GPU UTILIZATION (—)0%
DECODE THROUGHPUT (tok/s)0
SECONDARY SERVERS — live—
REASONING TAP — live CoT (optional: streamed reasoning_content, set THOUGHT_LOG)
—
tap idle — no reasoning streams captured yet.
Optional: run a proxy that tees each request's reasoning_content to a log and set THOUGHT_LOG; each concurrent stream gets its own panel here.
NETWORK — —Σ — this session · peak ↓ —
LAN (passive: arp + conns)—
—
Model library — loadable inventory · served from fleet-metrics /models
loading registry…