Board damage an elimination tree cannot see: ceph7's three motherboards
- date
- 20260902 to 20260911
- what happened
- ceph7's SAS path was dead on its first board, and a full day of elimination found nothing until the board came out and the damaged riser-slot connector was visible under bench light. The replacement board fixed the SAS fault and arrived with a dead PCIe root port. The third board, fitted 20260911, fixed that and turned out to carry an unlocked, updatable BIOS.
- what it cost
- Close to two full elimination days across the two faults, plus ceph7 running without its SAS disks for over a week while parts moved.
- what changed
- Pulling the board and inspecting the slot contacts under good light came before the swap tree after board one. After board three, every incoming board's own BIOS is checked rather than assumed to match the fleet.
- the check now
- Read the BIOS version and lock state of any board entering the fleet, from the vendor tool or the physical BIOS screen, and record it before assuming the fleet-wide lock applies.
One host needed three motherboards in nine days, and each swap taught something different.
The first board’s SAS path (SAS is the disk interface that carries this chassis’s twelve drive bays) was simply dead. The elimination tree ran nearly a full day: a BIOS settings diff against a known-good node, the fleet’s gold bifurcation settings applied, two different host bus adapters tried, a riser swapped in from a sibling chassis, controller and PCI resets, a fresh CMOS cell, a CPU swap between sockets, and a BIOS reflash the board refused. None of it found anything, because the fault was a physically damaged connector on the riser slot, and that was visible in seconds once the board was out of the chassis and under bench light.
The second board fixed the SAS fault and brought its own: one PCIe root port electrically dead, so seven of eight NVMe drives enumerated instead of eight. The third board, fitted 20260911, fixed the dead lane and quietly broke an assumption the fleet had just spent a week hardening into policy, that every board here carries a previous owner’s signed BIOS that cannot be updated. It did not. It was retail and updatable, the first confirmed exception.
Two habits came out of this. Look at the contacts before running the tree, because the bench sees what the rack cannot. And check each replacement part on its own terms. A fleet built from used hardware is never as uniform as its purchase order implies, and assuming a replacement matches its predecessor costs a week when it does not.
A third habit is about who does what. This diagnosis was run by an agent, and an agent iterates on configuration and software in seconds; every step here that needed hands cost a person’s time and, around it, a host outage: drain the host, power it down, pull the board, do the work, refit, bring it back, verify. Ten iterations of that shape is a week, and it was. So an agent working a hardware fault exhausts what software can rule out first, then groups every request for hands into one visit, ordered by what each would eliminate, with the parts, the cables and the read-back for each step written down before the person walks in. Rapid iteration is the agent’s strength on a keyboard and its most expensive habit on a rack.
Source: node0 lessons v0.1, lesson 1.2. Sanitized: checklist v0.1, 20260921; board serials, BMC MACs; voice pass 20260921. Part of oznog.com/node0.
