A dirty bay connector wears two different disguises
- date
- 20260906 to 20260912
- what happened
- A drive in one bay of ceph8 showed rising SATA CRC errors and a link trained down to 3.0 Gb/s over four days. It was drained and pulled, and a proven-good spare fitted into the same bay stalled at zero throughput with a clean SMART report and a clean 6.0 Gb/s link. Compressed air through the bay's connector cleared both signatures.
- what it cost
- A full drain-and-rebuild cycle of roughly six hours, plus a second spare drive's worth of diagnosis time, for a fault that was never in either drive.
- what changed
- A dedicated bay-verification tool, and the explicit rule to clean the bay before condemning a drive.
- the check now
- Before writing off a drive that misbehaves in a specific bay, fit a proven-good drive to the same bay and watch for a different signature, not for the absence of one.
One physical fault, contamination on the contacts of a single drive bay, produced two failures that looked nothing alike. That is what made it expensive.
The first drive’s symptoms were textbook dying disk from the interface side: CRC errors climbing over four days, and a SATA link that trained itself down from 6.0 Gb/s to 3.0 Gb/s, which is what a link layer does when it cannot keep a clean signal. The drive was drained out of Ceph and pulled.
A proven-good spare went into the same bay and failed in a completely different way: link clean at full speed, SMART report clean, and no data moving at all. Nothing about that looks like a connector problem. Read on its own it looks like a dead drive or a stuck controller.
The tell was the pair. Two healthy drives producing two unrelated-looking symptoms in the same physical slot points at the slot, not at either drive. Compressed air through the bay’s own connector, then reseating first the spare and then the original drive, cleared both signatures, and the original drive went back into service and has run clean since.
The rule now is to clean the bay before condemning a drive, and the verification is a substitution: fit a known-good drive to the same bay and expect a different failure signature if the bay is the fault. A close relative from a different chassis is worth keeping in view. An HGST drive condemned once and then cleared twice on retest, in two bays on two hosts, was neither replaced nor claimed. It is held as an unresolved marginal case rather than forced into a bucket it does not fit.
Source: node0 lessons v0.1, lesson 1.10. Sanitized: checklist v0.1, 20260921; drive serials; voice pass 20260921. Part of oznog.com/node0.
