Tenant guidance

What was measured

1 node, 8 × NVIDIA B300 SXM6 AC · driver 580.173.02 · Ethernet/InfiniBand 100-400 Gb/s, 5 ports per node

3 passive sweeps between 2026-09-10 03:47 UTC and 2026-09-10 03:48 UTC. Active probes run: all-reduce inside each node, reference training step.

Judged against reference targets (illustrative; your SLA replaces them). Other fleets used: 0 — every comparison is inside this fleet or against the stated basis, never against anyone else's hardware.

Run 184b7b2fa280 · generated 2026-09-10 21:30 UTC · 3 sweeps in store

info1

Measured against reference targets (illustrative; your SLA replaces them), 1 node and 8 GPUs meet 9 of 12 checks; 3 could not be judged.

1/1nodes with no failed check
1 of them not fully verified: a required check could not be judged
8/8GPUs with no failed check
9/12checks met
3 could not be judged
collection gaps on 1 node(s): dmesg unavailable (often requires elevated rights)
0actions
$38what you paid for 8 GPUs over 0.6 h (your rental_hours) at $7.87 per GPU-hour
$0on GPUs that failed a check — 0 GPUs

Working / idle split not shown. 3 sweeps covering 1 minutes of a 0.6-hour period is a diagnostic snapshot, not a record of how the capacity was used; utilization seen in that window is not evidence of your usage or of loss. A day of `11mac watch` alongside your jobs is what prices out. Diagnostic overhead in this window: 769 GPU-seconds of probes, not counted as work or idle.

$7.87/GPU-hour and $0.10/kWh are reference prices, not yours; set them in the expectations file.

Sections that need a decision start open; the rest are folded — click a heading to open it. Printing or saving as PDF includes every section.

Verdict

Measured against reference targets (illustrative; your SLA replaces them) — no SLA was supplied, so each device is held to what it says it can do. Pass --expectations sla.json to judge against your contract.

meets 9not met 0could not judge 3

Also worth knowing

Targets are reference targets, not a contract. No SLA file was supplied, so each check is judged against a shipped floor with its source shown beside it (a datasheet, the card's own limit, a published measurement, or our own capture). Pass --expectations your-sla.json and every target becomes your contract's number.

CheckWhat this checksTargetWhereVerdictMeasuredImpact and what to do$/day at risk
Scale-up bandwidth: all-reduce inside one serverScale-up: how fast the GPUs inside one server combine their results over their own links (NVLink or xGMI). Every multi-GPU step on a server waits on this.≥ 560 GB/s
reference: 560 GB/s intra-node all-reduce bus bandwidth: no published 8x B300 nccl-tests figure was found, so this is 80% of our own capture (700.5 GB/s on one Nebius node, 2026-09-09)
1 nodemeets701 GB/s over 8 GPUs against a floor of 560Met. Nothing to do.none at risk
All-reduce returns correct valuesWhether the GPUs combined those results correctly. A wrong value here means training would silently go wrong.0 incorrect reduction results
nccl-tests checks every result
1 nodemeetsnccl-tests verified every reduction resultMet. Nothing to do.none at risk
Uncorrectable memory errors on lifetime recordUncorrectable memory errors from before this boot, and whether the GPU has repaired the damaged rows since.row-remapper: remapped and clean is at spec; pending or failed is not
NVIDIA row-remapping documentation
89.124.38.244/gpu5meets4 lifetime, 1 row remapped, none since boot, no repair pending — normal operation per the driverA history worth watching, not a fault todaynone at risk
GPUs fit for workThe share of GPUs that passed every check, against the share the contract promises.≥ 97% of GPUs with no failed check
reference: 0.97 of GPUs with no failed check
fleetmeets8 of 8 (100.0%) with no failed check; target 97%Met. Nothing to do.none at risk
Network links at contracted rateThe speed each network port negotiated with the switch, against the speed the card can run at or the contract promises.≥ 400,000 Mb/s
reference: 400 Gb/s: the port rate the B300 node's ConnectX-8 ports reported
1 portsmeetsall at or above 400,000 Mb/sMet. Nothing to do.none at risk
PCIe links at full widthHow many PCIe lanes connect each GPU to its host. Fewer than the card supports means data moves between CPU and GPU more slowly than it should.the card's own maximum link width
the card's own reported limit
8 GPUsmeetsall at rated widthMet. Nothing to do.none at risk
Communication exposed on the reference stepWhat share of that reference step was spent waiting for GPUs to exchange gradients rather than computing, with no overlap: a worst case.≤ 25% of the step
reference: 0.25 of the un-overlapped reference step: our capture showed 14.5% on 8 GPUs over NVLink; a degraded link roughly doubles it.
1 nodemeets14.5% against a ceiling of 25%Met. Nothing to do.none at risk
Memory bandwidth on the reference stepHow fast each GPU can move data through its own memory, against the datasheet figure. Token generation in inference is limited by this, not by raw compute.≥ 70% of datasheet memory bandwidth
reference: 0.70 of the 8 TB/s datasheet: our capture reached 75-79% with a device-to-device copy; the copy method itself costs 10-15%.
8 GPUsmeets74.6%-78.5% against a floor of 70%Met. Nothing to do.none at risk
GPU temperatures within limitGPU temperature against the limit at which the card starts holding its clocks back.≤ 85 °C
reference: 85 C is an illustrative warning level, not this SKU's manufacturer limit: the B300 on driver 580 does not report its thermal thresholds through nvidia-smi -q, so the vendor-limit comparison is inconclusive and this value is carried over from A100/H100
8 GPUsmeetsall at or below 85 CMet. Nothing to do.none at risk
No hardware-fault messages in the accessible kernel logWhether the readable part of the kernel log holds any message that public vendor or kernel documentation classes as a hardware fault: uncorrectable PCIe or memory errors, a device dropping off the bus, a GPU reset. Benign and application-level messages are counted but never fail this. The log covers whatever the ring buffer still held at read time.0 fault-class messages
public kernel and vendor documentation, per signature
1 nodecould not judgethe kernel log could not be read: dmesg is restricted to root on this image and the tool never escalates privileges→ Run with a user that can read dmesg, or grant read access to /dev/kmsg
Scale-out bandwidth: RDMA write between servers (host memory)Scale-out: how fast one server's network card can write straight into another server's memory over RDMA (host memory in this test; GPUDirect is not exercised), the east-west path every multi-server job depends on, against the port's speed.≥ 80% of the port rate
reference: 0.8 of port rate: perftest reaches 90%+ of line rate on a correctly configured path (active MTU 4096); 80% leaves margin.
fleetcould not judgethe RDMA probe is pairwise, one node writing into another; only one node was measured, so there was no pair to test→ Include a second node on the same fabric in the next run
Reference-step utilisation (MFU)How much of each GPU's rated compute a fixed reference training step actually achieved (model FLOP utilisation). A GPU that is throttled or misconfigured scores low. This is our step, not your workload.no target set
add mfu_min; the peak itself is the vendor datasheet
8 GPUscould not judge1290-1363 TFLOP/s achieved; no rated peak on file for this SKU, so MFU cannot be computed
info

Uncorrectable memory errors on record

89.124.38.244/gpu5

GPU 5 on 89.124.38.244 had uncorrectable memory errors in the past. The driver reports the damaged rows were remapped and none have occurred since boot, which is the state NVIDIA describes as normal operation. Keep using it; watch that the count does not climb.

89.124.38.244/gpu5: 4 uncorrectable errors on record, 1 row remapped, 5759 banks with full remap capacity left, none since boot

Action: Keep in service and watch the count; escalate if new uncorrectable errors appear, a remap fails, or the remapped-row count reaches NVIDIA's row-remapping policy threshold  ·  Owner: vendor  ·  Confidence: high

What this is not: The driver reports the damaged rows remapped and no uncorrectable error since boot — the condition NVIDIA documents as normal operation after row remapping. The history is worth watching; it is not a fault today

evidence: 89.124.38.244/8d918c5f885cc361/b33a7a7c980754fb · 89.124.38.244/aee7760f0692646f/84acfe7a3914e17b

NodeStatusSKUDriverGPUsMeasurementsNot read
89.124.38.244unverifiedNVIDIA B300 SXM6 AC580.173.028869dmesg unavailable (often requires elevated rights)

"Not read" lists tools that were absent or restricted on that node. A gap is recorded as a gap, never filled with a guess.

Working / idle split not shown. 3 sweeps covering 1 minutes of a 0.6-hour period is a diagnostic snapshot, not a record of how the capacity was used; utilization seen in that window is not evidence of your usage or of loss. A day of `11mac watch` alongside your jobs is what prices out.

Every figure covers 0.6 hours (your rental_hours). GPU-hours are priced at $7.87 — what you pay for one — and electricity at $0.10 per kWh — both are placeholder prices; set gpu_hour_price_usd and power_price_usd_per_kwh in your expectations file for your real numbers.

WhatFor the periodHow it was worked out
What you paid for these GPUs$388 GPUs × $7.87 per GPU-hour × 0.6 hours. The price is a placeholder — set gpu_hour_price_usd to yours
Provider shortfall — GPUs you could not trust$0no GPU failed a check
Diagnostic overhead — the measurement's own probes769 GPU-seconds of active probes; not counted as work or as idle
Electricity the GPUs drew (the provider's bill), measured$01.46 kW across 8 GPUs × 0.6 hours × $0.10 per kWh. The electricity price is a placeholder — set power_price_usd_per_kwh. GPUs only: servers, cooling and network are not included

Measured quantities × the prices supplied. An illustration for planning, not an invoice deduction or an observed loss; it does not assign responsibility on its own.

Bus bandwidth by message size, nccl-tests convention. The top number is what gets quoted; the shape is what matters to a job whose messages are small. Scale-up is inside one server (NVLink / xGMI); scale-out crosses servers.

CollectiveScopeNode8 KB1 MB64 MBPeakSizes swept
all-reducescale-up89.124.38.2440.333.9418.2700.5 at 512 MB27

GB/s throughout.

What the tools reported, per device, in this sweep. The JSON carries every measurement; this is the subset a reader checks first.

gpu

NodeDeviceclocks_current_smclocks_max_smecc.aggregate_dram_correctableecc.aggregate_sram_correctableecc.volatile_dram_uncorrectableenforced_power_limitpcie_link_width_currentpcie_link_width_maxpower_drawtemperature_gputhrottledrefstep.exposed_comm_fractionrefstep.hbm_gbsrefstep.tflops
89.124.38.244gpu0120203200011001616182.8228no6278.201311.19
89.124.38.244gpu1120203200011001616185.1930no5967.501362.64
89.124.38.244gpu2120203200011001616185.0228no6244.781290.28
89.124.38.244gpu3120203200011001616179.3529no6206.321346.54
89.124.38.244gpu4120203200011001616184.8730no6193.911340.22
89.124.38.244gpu5120203200011001616179.1129no6277.971347.99
89.124.38.244gpu6120203200011001616179.5128no6269.131341.80
89.124.38.244gpu7120203200011001616180.2228no6265.291298.25
89.124.38.244gpus0.14

nic

NodeDevicecrc_errorslink_detectedlink_speedlink_speed_maxpfc_pause_rxcollective.allreduce.busbw.intra-nodecollective.allreduce.latency_us.intra-node
89.124.38.244docker0no
89.124.38.244eth00yes4000004000000
89.124.38.244ibp11s0f00no1000000
89.124.38.244ibp11s0f10no1000000
89.124.38.244ibp11s0f20no1000000
89.124.38.244ibp11s0f30no1000000
89.124.38.244nccl700.5354.61

ib

NodeDeviceactive_mtumax_mturatestate
89.124.38.244mlx5_0:15124096100Active
89.124.38.244mlx5_1:15124096100Active
89.124.38.244mlx5_2:15124096100Active
89.124.38.244mlx5_3:15124096100Active
89.124.38.244mlx5_4:110244096400Active

system

NodeDevicekernel_taint_flags
89.124.38.244systemout_of_tree_module

20 intra-fleet outliers and 0 changes over time were evaluated by the rules. 3 runs in the store; change detection needs at least 6.

Topology discovery found no switch attachments (no LLDP visible), so findings cannot be attributed to a leaf or uplink.

GPU modelNodesGPUsCompared againstOther fleets used
NVIDIA B300 SXM6 AC18this fleet only0

Every finding compares your nodes and devices with each other. "Other fleets used: 0" means nothing here is judged against anyone else's hardware — a node is flagged only when it disagrees with its own peers.

Passive sweep: 63.4 node-seconds across 153 read-only commands; no GPU time and no fabric traffic by construction.

Active probes: 769.1 GPU-seconds, summed from the probes' own records. Fabric bytes for the RDMA probe are recorded in that probe's own run report, not in this store.

869 measurements, hash-chained. Chain root: ecbdf3f9ffadf961a6efde112ae59ea6270ffbaf7cea2366a6da2806e8fcfe0c

Every one of the 34 raw outputs the measurements cite is present in the archive and hashed in the manifest.

The chain and the hashes show the numbers were not altered after they were written; they do not make the method right. Anyone can re-run the same commands and check. Verify without trusting this file: 11mac verify report.json --db fleet.db --archive ./raw recomputes every hash from the data.