Tenant guidance

What was measured

4 nodes, 32 × NVIDIA A100-SXM4-40GB · driver 580.105.08 · Ethernet 50 Gb/s, 1 port per node

8 passive sweeps between 2026-09-06 23:22 UTC and 2026-09-06 23:22 UTC. Active probes run: RDMA bandwidth between nodes.

Judged against reference targets (illustrative; your SLA replaces them). Other fleets used: 0 — every comparison is inside this fleet or against the stated basis, never against anyone else's hardware.

Run 133612432f8a · generated 2026-09-10 21:30 UTC · 8 sweeps in store

warning2 info2

Measured against reference targets (illustrative; your SLA replaces them), 4 nodes and 32 GPUs meet 6 of 10 checks; 2 not met; 2 could not be judged.

3/4nodes with no failed check
sample-129-146-226-5: scale-out bandwidth: rdma write between servers (host memory) (0.39 of 50 Gb/s (1%; floor 80%)); no hardware-fault messages in the accessible kernel log (1 × nvswitch.sxid.fatal: An NVSwitch fault NVIDIA classes as fatal. Traffic across that switch or link is affected.)
32/32GPUs with no failed check
6/10checks met
2 could not be judged
2actions — 1 operator, 1 vendor
$64what you paid for 32 GPUs over 1 h (your rental_hours) at $1.99 per GPU-hour
$0on GPUs that failed a check — 0 GPUs

Working / idle split not shown. 8 sweeps covering 1 minutes of a 1-hour period is a diagnostic snapshot, not a record of how the capacity was used; utilization seen in that window is not evidence of your usage or of loss. A day of `11mac watch` alongside your jobs is what prices out.

$1.99/GPU-hour and $0.10/kWh are reference prices, not yours; set them in the expectations file.

Sections that need a decision start open; the rest are folded — click a heading to open it. Printing or saving as PDF includes every section.

Verdict

Measured against reference targets (illustrative; your SLA replaces them) — no SLA was supplied, so each device is held to what it says it can do. Pass --expectations sla.json to judge against your contract.

meets 6not met 2could not judge 2

What needs to happen

IssueWhereWhat needs to happenWho$/day
Scale-out bandwidth: RDMA write between servers (host memory)
0.39 of 50 Gb/s (1%; floor 80%)
sample-129-146-226-5/mlx5_0->sample-161-153-79-0Raise the netdev MTU on mlx5_0's interface so the RDMA path can use its 4096-byte maximum; re-probe to confirm
Multi-node jobs across this path get about 1% of the bandwidth expected
operator
No hardware-fault messages in the accessible kernel log
1 × nvswitch.sxid.fatal: An NVSwitch fault NVIDIA classes as fatal. Traffic across that switch or link is affected.
sample-129-146-226-5Confirm with: nvidia-smi nvlink -e for error counters on the GPUs behind that switch; Fabric Manager logs.
A device on this node has faulted per its own driver
vendor

Also worth knowing

Targets are reference targets, not a contract. No SLA file was supplied, so each check is judged against a shipped floor with its source shown beside it (a datasheet, the card's own limit, a published measurement, or our own capture). Pass --expectations your-sla.json and every target becomes your contract's number.

CheckWhat this checksTargetWhereVerdictMeasuredImpact and what to do$/day at risk
No hardware-fault messages in the accessible kernel logWhether the readable part of the kernel log holds any message that public vendor or kernel documentation classes as a hardware fault: uncorrectable PCIe or memory errors, a device dropping off the bus, a GPU reset. Benign and application-level messages are counted but never fail this. The log covers whatever the ring buffer still held at read time.0 fault-class messages
public kernel and vendor documentation, per signature
sample-129-146-226-5does not meet1 × nvswitch.sxid.fatal: An NVSwitch fault NVIDIA classes as fatal. Traffic across that switch or link is affected.A device on this node has faulted per its own driver → Confirm with: nvidia-smi nvlink -e for error counters on the GPUs behind that switch; Fabric Manager logs. (vendor)
Scale-out bandwidth: RDMA write between servers (host memory)Scale-out: how fast one server's network card can write straight into another server's memory over RDMA (host memory in this test; GPUDirect is not exercised), the east-west path every multi-server job depends on, against the port's speed.≥ 80% of the port rate
reference: 0.8 of port rate: perftest ib_write_bw with 64 KB messages reaches 90%+ of line rate on a correctly configured RoCE path (active MTU 4096)
sample-129-146-226-5/mlx5_0->sample-161-153-79-0does not meet0.39 of 50 Gb/s (1%; floor 80%)Multi-node jobs across this path get about 1% of the bandwidth expected → Raise the netdev MTU on mlx5_0's interface so the RDMA path can use its 4096-byte maximum; re-probe to confirm (operator)
Uncorrectable memory errors on lifetime recordUncorrectable memory errors from before this boot, and whether the GPU has repaired the damaged rows since.row-remapper: remapped and clean is at spec; pending or failed is not
NVIDIA row-remapping documentation
sample-129-146-226-5/gpu2meets4 lifetime, 1 row remapped, none since boot, no repair pending — normal operation per the driverA history worth watching, not a fault todaynone at risk
GPUs fit for workThe share of GPUs that passed every check, against the share the contract promises.≥ 97% of GPUs with no failed check
reference: 0.97 of GPUs with no failed check
fleetmeets32 of 32 (100.0%) with no failed check; target 97%Met. Nothing to do.none at risk
Network links at contracted rateThe speed each network port negotiated with the switch, against the speed the card can run at or the contract promises.≥ 50,000 Mb/s
reference: 50 Gb/s: what the sample fleet's ConnectX ports were provisioned at (a 2-lane breakout of a 100G card)
4 portsmeetsall at or above 50,000 Mb/sMet. Nothing to do.none at risk
PCIe links at full widthHow many PCIe lanes connect each GPU to its host. Fewer than the card supports means data moves between CPU and GPU more slowly than it should.the card's own maximum link width
the card's own reported limit
32 GPUsmeetsall at rated widthMet. Nothing to do.none at risk
GPU temperatures within limitGPU temperature against the limit at which the card starts holding its clocks back.≤ 85 °C
reference: 85 C is the A100-SXM4's own 'GPU Max Operating Temp' as reported by nvidia-smi -q (slowdown 89 C, shutdown 92 C)
32 GPUsmeetsall at or below 85 CMet. Nothing to do.none at risk
No hardware-fault Xid codes in the accessible kernel logWhether the GPU driver has logged a code NVIDIA classes as a hardware fault (an Xid) in the part of the kernel log the tool could read. Most Xids are the application's, and are not counted here.0 hardware fault codes
NVIDIA Xid documentation
4 nodesmeetsno hardware-class Xid in the readable logMet. Nothing to do.none at risk
Scale-up bandwidth: all-reduce inside one serverScale-up: how fast the GPUs inside one server combine their results over their own links (NVLink or xGMI). Every multi-GPU step on a server waits on this.≥ 170 GB/s
reference: 170 GB/s intra-node all-reduce bus bandwidth (nccl-tests convention, 8 GPUs, largest message)
fleetcould not judgeno collective probe was run (needs nccl-tests / rccl-tests on the node and --probe)→ Run `11mac probe --kind collective --probe` in an agreed window
Reference-step utilisation (MFU)How much of each GPU's rated compute a fixed reference training step actually achieved (model FLOP utilisation). A GPU that is throttled or misconfigured scores low. This is our step, not your workload.≥ 70% of the datasheet peak
reference: 0.70 on the REFERENCE STEP (dense bf16 MLP, 1.07 B params, 8192 tokens per GPU), not on a customer workload
fleetcould not judgeno reference step was run (needs PyTorch on the node and --probe)→ Run `11mac probe --kind refstep --probe` in an agreed window
warning

RDMA throughput capped by path MTU

sample-129-146-226-5/mlx5_0:1

Direct GPU-to-GPU network transfers from mlx5_0:1 on sample-129-146-226-5 run at 0.39 Gb/s on a 50 Gb/s link — about 1% of what the link can do. The most likely cause is a network packet-size setting (MTU) below what the card supports; the link and the card themselves look healthy. Raising the MTU and re-measuring would confirm it.

sample-129-146-226-5/mlx5_0:1 delivers 0.39 Gb/s over RDMA on a 50 Gb/s port (1%) with active MTU 1024 of 4096

Action: Raise the netdev MTU on mlx5_0's interface so the RDMA path can use its 4096-byte maximum; re-probe to confirm  ·  Owner: operator  ·  Confidence: medium

What this is not: The port is up at 50 Gb/s and the NIC advertises a 4096-byte MTU while the path runs at 1024; that is the most likely ceiling. The card and cable are not implicated by these reads, and not excluded: raising the MTU and re-measuring is what settles it

evidence: 129.146.226.5/08050ef15dc674a6/dbe1cec661f3aadf · sample-129-146-226-5/4b3aaf6a63aa0f52/81dfd2d2280959b1 · sample-129-146-226-5/ae980319a9a0920d/bce7ed74a4d54d7a

warning

Kernel log: nvswitch.sxid.fatal

sample-129-146-226-5/system

sample-129-146-226-5 logged a message its own driver documentation classes as a hardware fault. An NVSwitch fault NVIDIA classes as fatal. Traffic across that switch or link is affected. To confirm: nvidia-smi nvlink -e for error counters on the GPUs behind that switch; Fabric Manager logs.

sample-129-146-226-5: 1 kernel message matching nvswitch.sxid.fatal: An NVSwitch fault NVIDIA classes as fatal. Traffic across that switch or link is affected. Source: NVIDIA Fabric Manager user guide: SXid severity levels.

Action: Confirm with: nvidia-smi nvlink -e for error counters on the GPUs behind that switch; Fabric Manager logs.  ·  Owner: vendor  ·  Confidence: medium

What this is not: A healthy node logs none of these

evidence: sample-129-146-226-5/9d36a56926ee9cbe/839eecf0483373b8

info

Uncorrectable memory errors on record

sample-129-146-226-5/gpu2

GPU 2 on sample-129-146-226-5 had uncorrectable memory errors in the past. The driver reports the damaged rows were remapped and none have occurred since boot, which is the state NVIDIA describes as normal operation. Keep using it; watch that the count does not climb.

sample-129-146-226-5/gpu2: 4 uncorrectable errors on record, 1 row remapped, 639 banks with full remap capacity left, none since boot

Action: Keep in service and watch the count; escalate if new uncorrectable errors appear, a remap fails, or the remapped-row count reaches NVIDIA's row-remapping policy threshold  ·  Owner: vendor  ·  Confidence: high

What this is not: The driver reports the damaged rows remapped and no uncorrectable error since boot — the condition NVIDIA documents as normal operation after row remapping. The history is worth watching; it is not a fault today

evidence: sample-129-146-226-5/8d918c5f885cc361/a9579e3ef999098c · sample-129-146-226-5/aee7760f0692646f/a57b7d965cc587e3

info

NIC negotiated below its own capability — fleet-wide

4 nodes

All 4 nodes run their network cards at 50,000 Mb/s although the cards support 100,000. Because every node is the same, this is most likely how the fleet was provisioned rather than a fault on any one of them. Worth confirming it is intentional.

4 nodes run 50000 Mb/s against a reported maximum of 100000 Mb/s (50% of capability), uniformly

Action: Confirm this is the intended provisioning; no per-node action  ·  Owner: operator  ·  Confidence: high

What this is not: All 4 nodes of this SKU sit at the same reduced rate, which is a configuration choice, not a failure on any one of them

evidence: sample-129-146-226-5/495bc5762ee725f1/07e79dd9326b926c · sample-161-153-79-0/495bc5762ee725f1/07e79dd9326b926c · sample-168-110-54-32/495bc5762ee725f1/07e79dd9326b926c · sample-82-70-255-133/495bc5762ee725f1/07e79dd9326b926c

NodeStatusSKUDriverGPUsMeasurementsNot read
sample-129-146-226-5attentionNVIDIA A100-SXM4-40GB580.105.088359
sample-161-153-79-0okNVIDIA A100-SXM4-40GB580.105.088350
sample-168-110-54-32okNVIDIA A100-SXM4-40GB580.105.088350
sample-82-70-255-133okNVIDIA A100-SXM4-40GB580.105.088355

"Not read" lists tools that were absent or restricted on that node. A gap is recorded as a gap, never filled with a guess.

Working / idle split not shown. 8 sweeps covering 1 minutes of a 1-hour period is a diagnostic snapshot, not a record of how the capacity was used; utilization seen in that window is not evidence of your usage or of loss. A day of `11mac watch` alongside your jobs is what prices out.

Every figure covers 1 hours (your rental_hours). GPU-hours are priced at $1.99 — what you pay for one — and electricity at $0.10 per kWh — both are placeholder prices; set gpu_hour_price_usd and power_price_usd_per_kwh in your expectations file for your real numbers.

WhatFor the periodHow it was worked out
What you paid for these GPUs$6432 GPUs × $1.99 per GPU-hour × 1 hours. The price is a placeholder — set gpu_hour_price_usd to yours
Provider shortfall — GPUs you could not trust$0no GPU failed a check
Electricity the GPUs drew (the provider's bill), measured$02.34 kW across 32 GPUs × 1 hours × $0.10 per kWh. The electricity price is a placeholder — set power_price_usd_per_kwh. GPUs only: servers, cooling and network are not included

Measured quantities × the prices supplied. An illustration for planning, not an invoice deduction or an observed loss; it does not assign responsibility on its own.

No collective sweep ran in this window, so there is no bandwidth-by-message-size table. It needs nccl-tests (or rccl-tests) on the nodes and 11mac probe --kind collective --probe in an agreed window; it then covers all-reduce, reduce-scatter, all-gather and all-to-all, inside each server and, with a launcher, across servers.

What the tools reported, per device, in this sweep. The JSON carries every measurement; this is the subset a reader checks first.

gpu

NodeDeviceclocks_current_smclocks_max_smecc.aggregate_dram_correctableecc.aggregate_sram_correctableecc.volatile_dram_uncorrectableenforced_power_limitpcie_link_width_currentpcie_link_width_maxpower_drawtemperature_gputhrottled
sample-129-146-226-5gpu0141014100004001616356.7977no
sample-129-146-226-5gpu1141014100004001616330.6966no
sample-129-146-226-5gpu221014101800400161653.9831no
sample-129-146-226-5gpu32101410000400161656.7732no
sample-129-146-226-5gpu42101410000400161654.4730no
sample-129-146-226-5gpu52101410000400161652.3929no
sample-129-146-226-5gpu62101410000400161654.3631no
sample-129-146-226-5gpu72101410000400161651.7931no
sample-161-153-79-0gpu02101410000400161659.2133no
sample-161-153-79-0gpu12101410000400161657.1332no
sample-161-153-79-0gpu22101410000400161655.7232no
sample-161-153-79-0gpu32101410000400161654.5833no
sample-161-153-79-0gpu42101410000400161655.9532no
sample-161-153-79-0gpu52101410000400161662.7633no
sample-161-153-79-0gpu62101410000400161653.0533no
sample-161-153-79-0gpu72251410000400161656.7534no
sample-168-110-54-32gpu02101410000400161657.9735no
sample-168-110-54-32gpu12101410000400161656.8433no
sample-168-110-54-32gpu22101410000400161654.0333no
sample-168-110-54-32gpu32101410000400161653.7033no
sample-168-110-54-32gpu42101410000400161658.4934no
sample-168-110-54-32gpu52101410000400161653.2133no
sample-168-110-54-32gpu62101410000400161659.0337no
sample-168-110-54-32gpu72251410000400161656.0137no
sample-82-70-255-133gpu02101410000400161653.2628no
sample-82-70-255-133gpu12101410000400161653.4327no
sample-82-70-255-133gpu22101410000400161654.3527no
sample-82-70-255-133gpu32101410000400161652.1227no
sample-82-70-255-133gpu42101410000400161655.9428no
sample-82-70-255-133gpu52101410000400161651.2026no
sample-82-70-255-133gpu62101410000400161652.4528no
sample-82-70-255-133gpu72101410000400161656.4428no

nic

NodeDevicecrc_errorslink_detectedlink_speedlink_speed_maxpfc_pause_rx
sample-129-146-226-5docker0no
sample-129-146-226-5eno10yes500001000000
sample-161-153-79-0docker0no
sample-161-153-79-0eno10yes500001000000
sample-168-110-54-32docker0no
sample-168-110-54-32eno10yes500001000000
sample-82-70-255-133docker0no
sample-82-70-255-133eno10yes500001000000

ib

NodeDeviceactive_mtumax_mturatestate
sample-129-146-226-5mlx5_0:11024409650Active
sample-161-153-79-0mlx5_0:11024409650Active
sample-168-110-54-32mlx5_0:11024409650Active
sample-82-70-255-133mlx5_0:11024409650Active

rdma

NodeDevicerdma.bw_gbps.unknown
sample-129-146-226-5mlx5_0->sample-161-153-79-00.39

system

NodeDevicekernel_taint_flagsmce_countxid_count
sample-129-146-226-5systemout_of_tree_module00
sample-161-153-79-0systemout_of_tree_module00
sample-168-110-54-32systemout_of_tree_module00
sample-82-70-255-133systemout_of_tree_module00

22 intra-fleet outliers and 0 changes over time were evaluated by the rules.

Topology discovery found no switch attachments (no LLDP visible), so findings cannot be attributed to a leaf or uplink.

GPU modelNodesGPUsCompared againstOther fleets used
NVIDIA A100-SXM4-40GB432this fleet only0

Every finding compares your nodes and devices with each other. "Other fleets used: 0" means nothing here is judged against anyone else's hardware — a node is flagged only when it disagrees with its own peers.

Passive sweep: 0.1 node-seconds across 620 read-only commands; no GPU time and no fabric traffic by construction.

Active probes: GPU-seconds not recorded (the probe measurements in this store predate that accounting). Fabric bytes for the RDMA probe are recorded in that probe's own run report, not in this store.

1,414 measurements, hash-chained. Chain root: f0cf0a4c2dd474d73b9167fb07ff4a5f3b8ef6e30929896d6b99f13647ddb56b

56 of 57 raw outputs the measurements cite are present in the archive; 1 are not. Measurements without their source output are listed in the manifest as unverifiable and should be read as such.

The chain and the hashes show the numbers were not altered after they were written; they do not make the method right. Anyone can re-run the same commands and check. Verify without trusting this file: 11mac verify report.json --db fleet.db --archive ./raw recomputes every hash from the data.