Skip to content
Oznog

5.12 · power and environment · after redaction

An agent caused a hardware fault, and it is on the record

date
20260911
what happened
An agent found that one transfer switch's controller identified itself as a lower-voltage model, with a blank serial and an implausible manufacture date, and read both its input sources at roughly half the real line voltage. On Christoph's explicit go-ahead the agent changed the unit's configured line voltage from 120 to 230. The controller immediately declared both sources under-voltage, logged faults on both, and lost redundancy entirely.
what it cost
About twelve hours, from 03:24 to 15:35, with one host's power redundancy fully down, and one transfer switch consumed from spares. The change could not be reverted. The card refused it twice, and neither a management reboot nor a full power cycle cleared it. And about $550, the unit's price, that cannot be recovered: the switch had arrived with a mismatched controller, but nobody had recorded which vendor it came from, so there is no one to return it to.
what changed
The event is recorded by name, as a fault the agent caused, in the device's own hardware history alongside the sequence of what was tried and why it failed. A cold-spare policy for these transfer switches was already in place and is what made same-day recovery possible. The stance the site already worked to for most of its equipment became a written rule: every incoming item, second-hand above all, has its serial number and its vendor recorded on arrival and is tested at once, so that a defective or misconfigured unit can be returned or claimed while there is still someone to claim against.
the check now
A physical model check, measuring voltage against a known-good sibling unit, before treating a management card's self-identification as ground truth.

The proximate cause was a mismatched or defective controller board. The transfer switch reported itself as a different, lower-voltage product than the chassis and network card around it actually were, and read both of its inputs at about half the real line voltage. Nobody, human or agent, had a reason to doubt that self-report before acting on it. Setting the configured line voltage to match what the unit appeared to be was a reasonable change given the information available, and it was made with an operator’s explicit approval.

It was also wrong, and it could not be undone. With the new setting the controller measured both of its inputs as under-voltage at once, faulted both, and stopped protecting anything. Reverting was refused by the card twice. A management reboot did not help. A full power cycle did not help. The unit was replaced from a cold spare, and a host spent twelve hours on a single unprotected source.

What makes this worth publishing is less the mistake than what happened afterwards. The fault, its timeline, and the fact that an agent’s action caused it are written into the same permanent hardware record every other event on that device goes into, in the same words a human-caused fault would get, with no softer framing. Node0 is run in large part by agents. A record that quietly omits which of them broke something is not a record.

Two things carry over. Verify a device’s self-reported identity against independent evidence, such as a measured reading compared with a known-good sibling, before changing configuration based on it. And keep cold spares for the class of device whose failure removes redundancy rather than service. That policy, raised earlier after a different transfer switch was destroyed by a short, is the only reason this was a same-day recovery.

There is a third, and it is the expensive one. The unit had arrived with a controller that did not match its chassis, which is a defect a vendor would have taken back. Nobody had recorded which vendor it came from. The site’s stance for most of its equipment was already to record the serial and the seller on arrival and to test at once, and this switch was one of the few that slipped through it; the result is about $550 that cannot be recovered, because there is nowhere to attribute it. That made the stance a rule (rule 26). A second-hand part is a claim waiting to be made, and the claim needs three things written down before the box is even opened: the serial, the vendor, and the date it was tested.

Source: node0 lessons v0.1, lesson 5.12. Sanitized: checklist v0.1, 20260921; hostnames, addresses; voice pass 20260921. Part of oznog.com/node0.