Skip to content
Oznog

5.10 · power and environment · after redaction

103% of a UPS, and a nameplate rating that was wrong for months

date
20260909 (measured), 20260912 (worsened by new hardware), 20260917 (nameplate correction), 20260920 (re-measured)
what happened
The two UPS units feeding Node0's busiest rack draw, summed, close to the rating of a single one of them, so that rack cannot lose one UPS and keep protected power. Separately, the constant used for one unit's rating in that percentage was found on 20260917 to be the 208 V nameplate figure, not the 240 V figure the units actually run at, overstating every failover percentage by 6.7 percent for months.
what it cost
Months of a failover-risk metric that was directionally right and numerically wrong. The idle rack read as comfortably impossible to fail over when the true figure was just under 100 percent, a different and more nuanced risk than the dashboard described.
what changed
The rack-pair sum, not a per-unit percentage, is the standing way to judge failover risk on a dual-fed rack. The true rating is recorded on the device type with its voltage dependence spelled out rather than inferred from a SKU string, and the alert silence that covers the accepted risk carries an explicit expiry date.
the check now
A rack failover load percentage metric, live-verified on 20260920 at 103.3% using the corrected rating.

Two lessons are braided together here, and the second one is the more transferable.

The first is about the question. On a dual-fed rack, per-unit load percentage is the wrong number to watch, because both units can sit at a comfortable-looking figure and still be unable to carry each other. The number that matters is whether the pair’s total fits inside one unit’s capacity. Node0’s busiest rack does not: reviewed at 103% on 20260909, worse by 20260912 after a planned deployment, and accepted as permanent. Christoph declined to shed load onto the other rack, whose headroom is reserved for GPU hardware not yet installed. On mains, losing a partner puts the survivor into unprotected bypass; on battery, output drops within seconds with no ordered shutdown. That is an N+0 posture chosen deliberately, with a dated expiry on the silence that hides the alert, not an oversight.

The second is about the denominator. The rating a power device uses in its own arithmetic can depend on how it is configured, and these units are rated differently at 208 V and at 240 V. The fleet had been dividing by the 208 V figure while running at 240 V. Neither SNMP nor NUT exposed a nameplate rating anywhere across fifty-plus available variables; the only sources were the vendor’s own web page and working the arithmetic backwards from live wattage and the unit’s own reported load percentage.

If you compute a hard capacity limit from a device’s self-reported percentage, confirm the number the percentage is measured against, from outside the device. A plausible wrong assumption in a denominator does not look like an error. It looks like a dashboard, and it can survive for months.

Source: node0 lessons v0.1, lesson 5.10. Sanitized: checklist v0.1, 20260921; rack numbers, site watt totals; voice pass 20260921. Part of oznog.com/node0.