Skip to main content

Performance

Time-series view of compute, storage I/O, and network activity on the picked node. The page is one scrollable layout with three sections in a fixed order — Compute first, then Storage I/O, then Network Bandwidth. Charts roll forward in real time; pick a longer range from the time picker (1 h / 6 h / 24 h / 7 d / 30 d) to zoom out.

PerformancePerformance

Compute Utilization

Three charts side-by-side at the top:

  • CPU (%) — host CPU utilization
  • Memory (%) — host RAM utilization
  • GPU Avg (%) — average across all GPUs on this node

Each chart shows the live numeric value at the top-right and the trend below.

Below the strip is a GPU Performance (per device) card with one tile per GPU, showing its utilization spark and current value. This is the fast-twitch view used during a workload — pop it open while a training job is running to see whether the GPUs are actually being saturated.

Storage I/O

Per-device read / write throughput and IOPS, plus latency and in-flight queue depth, for every block device on the node. The cards mirror the Storage Inventory order so it is easy to map a slow device back to its hardware.

Network Bandwidth

One card per NIC port, with port metadata (link state, vendor, BDF, RoCE device, IP, MAC, MTU, LLDP-peer switch and port) above an RX/TX bandwidth chart. Each card has its own zoom-window chips (30 s / 1 m / 5 m / 10 m) so you can see line-rate transfer events that minute-resolution dashboards would miss.

Time Range

The chips above the Compute charts (1 h / 6 h / 24 h / 7 d / 30 d) change the window for the Compute section. Hover any chart for a precise value at a timestamp.

See Also