Skip to content
Oznog

1.12 · hardware failures and rmas · after redaction

ats6's application link died on its own management card

date
20260908 to 20260911
what happened
ats6's network management card answered telnet and SNMP normally but lost the ability to talk to the transfer-switch hardware it is fitted to, returning one status value where a healthy unit returns 136 and dropping its event log menu entirely. Two management reboots, a factory reset and a full power cycle all failed to make a fix survive a restart.
what it cost
An automatic transfer switch ran for three days with no confirmed ability to actually switch sources, protecting a server the whole time on an unverified path.
what changed
A vendor firmware bundle, flashed over FTP, fixed the link permanently and incidentally reflashed the switch controller itself. The same update silently wiped DNS, time and daylight-saving settings, so restoring those is now a documented step of any such flash.
the check now
After any firmware flash on this class of card, verify that time and name-resolution settings were not silently reset, and run a restart test with the protected load up before calling the unit fixed.

An automatic transfer switch (an ATS is the device that holds a load on one power feed and moves it to a second feed if the first fails) is only useful if it can actually transfer. Between 20260908 and 20260911 one of Node0’s could not, and almost nothing about it looked wrong.

Its network management card answered every management protocol normally. It responded to telnet, it answered SNMP queries, it accepted configuration changes. What had died was the internal channel between the card and the switching hardware it is bolted to. The tell was a count. A healthy unit returns 136 status values over SNMP; this one returned one. Its event log menu had disappeared from the interface entirely. For three days a server sat on a power path whose failover nobody could confirm.

Two reboots of the management card did nothing. A factory reset worked once and did not survive being reconfigured. A full power cycle, with the protected server shut down so it could be done safely, did not hold either. What fixed it was a vendor firmware bundle flashed over FTP, which repaired the link durably and, without being asked, also reflashed the transfer switch’s own controller.

The firmware update also reset settings nobody had asked it to touch. It silently cleared the unit’s name resolution, time synchronisation and daylight-saving settings, which then had to be restored by hand and verified. The general form of the lesson is to assume a vendor firmware update on embedded network hardware resets more than the bug it targets, and to prove the repair with a deliberate restart test rather than with the first healthy reading after the flash.

Source: node0 lessons v0.1, lesson 1.12. Sanitized: checklist v0.1, 20260921; names pass only; voice pass 20260921. Part of oznog.com/node0.