<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hardware failures and RMAs on Oznog</title><link>https://oznog.com/node0/lessons/hardware/</link><description>Recent content in Hardware failures and RMAs on Oznog</description><generator>Hugo</generator><language>en-us</language><copyright>© 2011-2026 &lt;a href="https://oznog.com"&gt;Oznog Holdings LLC&lt;/a&gt; · Founded by &lt;a href="https://chrisplough.com" target="_blank" rel="noopener"&gt;Christoph Plough&lt;/a&gt; · &lt;a href="https://oznog.com/index.xml"&gt;RSS&lt;/a&gt; · &lt;a href="https://creativecommons.org/licenses/by/4.0/" target="_blank" rel="noopener"&gt;CC BY 4.0&lt;/a&gt;</copyright><lastBuildDate>Tue, 22 Sep 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://oznog.com/node0/lessons/hardware/index.xml" rel="self" type="application/rss+xml"/><item><title>What thirteen pulled servers actually cost</title><link>https://oznog.com/node0/lessons/hardware/1-1/</link><pubDate>Tue, 01 Sep 2026 00:00:00 +0000</pubDate><guid>https://oznog.com/node0/lessons/hardware/1-1/</guid><description>&lt;p&gt;Between 20260901 and 20260903, thirteen chassis of decommissioned datacenter
hardware (three 1U, nine 2U twelve-bay, one 4U thirty-six-bay Supermicro
units) became a working fleet, with Christoph on site and an agent doing the
remote half. The sticker price of used enterprise gear is not the real cost.
The rest of the bill arrived as parts and hours.&lt;/p&gt;
&lt;p&gt;Counted over those three days: 3 of roughly 80 twelve-terabyte drives sold as
refurbished and never run were dead on arrival, 1 of 3 RAID cards was dead, 5
of 13 CMOS cells (the small battery that holds a board&amp;rsquo;s clock and firmware
settings) were flat, and one motherboard came with a riser slot that had been
physically damaged before it ever arrived. None of that is a defect rate you
would accept from new hardware, and all of it is normal for pulls.&lt;/p&gt;</description></item><item><title>Board damage an elimination tree cannot see: ceph7's three motherboards</title><link>https://oznog.com/node0/lessons/hardware/1-2/</link><pubDate>Wed, 02 Sep 2026 00:00:00 +0000</pubDate><guid>https://oznog.com/node0/lessons/hardware/1-2/</guid><description>&lt;p&gt;One host needed three motherboards in nine days, and each swap taught
something different.&lt;/p&gt;
&lt;p&gt;The first board&amp;rsquo;s SAS path (SAS is the disk interface that carries this
chassis&amp;rsquo;s twelve drive bays) was simply dead. The elimination tree ran nearly
a full day: a BIOS settings diff against a known-good node, the fleet&amp;rsquo;s gold
bifurcation settings applied, two different host bus adapters tried, a riser
swapped in from a sibling chassis, controller and PCI resets, a fresh CMOS
cell, a CPU swap between sockets, and a BIOS reflash the board refused. None
of it found anything, because the fault was a physically damaged connector on
the riser slot, and that was visible in seconds once the board was out of the
chassis and under bench light.&lt;/p&gt;</description></item><item><title>A settings dump that skips a whole menu</title><link>https://oznog.com/node0/lessons/hardware/1-3/</link><pubDate>Thu, 03 Sep 2026 00:00:00 +0000</pubDate><guid>https://oznog.com/node0/lessons/hardware/1-3/</guid><description>&lt;p&gt;A tool that reports a subset of the truth is more dangerous than one that
fails outright, because its output reads as complete.&lt;/p&gt;
&lt;p&gt;On 20260903, while building the gold BIOS files that let this fleet apply
firmware settings as code, the same vendor utility was run against two
Supermicro board families. On the X10DRU-i+ boards it dumped 282 settings as
commented text, including the IIO bifurcation submenu that decides whether a
PCIe slot presents itself as one device or four (which is what makes a
four-drive NVMe carrier work). On the X10DRH-iT board it dumped 201 settings
and the bifurcation submenu was simply not there. No warning, no comment, no
error. The file just ends up shorter.&lt;/p&gt;</description></item><item><title>Ceph hosts hang with no OS-side trace</title><link>https://oznog.com/node0/lessons/hardware/1-4/</link><pubDate>Sun, 06 Sep 2026 00:00:00 +0000</pubDate><guid>https://oznog.com/node0/lessons/hardware/1-4/</guid><description>&lt;p&gt;On the morning of 20260906 two storage hosts hung hard, hours apart. Each had
a processor IERR (an internal error the CPU raises when it cannot continue) in
the log kept by its board management controller, the small always-on computer
that watches a server independently of the operating system. Both machines
were powered, both were unresponsive, and neither had anything left to say:
their consoles were blank and their own logs had stopped being written. Ceph
kept serving from whatever survived, two of three monitors and a fraction of
its OSDs (an OSD is the daemon that owns one disk), for about ten hours, and
nothing paged.&lt;/p&gt;</description></item><item><title>Identity lives on the drive or the card, never the board</title><link>https://oznog.com/node0/lessons/hardware/1-5/</link><pubDate>Sat, 05 Sep 2026 00:00:00 +0000</pubDate><guid>https://oznog.com/node0/lessons/hardware/1-5/</guid><description>&lt;p&gt;Twice in two weeks, on two unrelated families of single-board computer, a
hardware fault that should have cost a reinstall cost nothing but a physical
swap.&lt;/p&gt;
&lt;p&gt;The first was a Raspberry Pi 5 with an I2C bootloader fault. Its NVMe carrier
and both drives moved to a healthy spare board, and the spare simply became
the same node: same hostname, same address, same keys, no reinstall and no
configuration change, because a Pi&amp;rsquo;s identity lives on its storage and the
board underneath is stateless silicon. The carrier itself was the harder part
of that job. It was failing too. The only way to prove the carrier was at
fault, rather than either board, was to cross-test it against a second Pi
with a different ribbon cable, which took about a day.&lt;/p&gt;</description></item><item><title>Two NICs on one installer subnet look exactly like a crashing host</title><link>https://oznog.com/node0/lessons/hardware/1-6/</link><pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate><guid>https://oznog.com/node0/lessons/hardware/1-6/</guid><description>&lt;p&gt;Sustained load turned a routing quirk into a symptom that read exactly like
failing hardware. During a NixOS install on 20260830, the machine was fine at
idle and fell over the moment real data moved, which is the shape that sends
everyone reaching for a hardware explanation.&lt;/p&gt;
&lt;p&gt;The cause was that the host had two Ethernet interfaces, its built-in port and
a USB adapter, and both took an address on the same installer subnet. Replies
left by whichever interface the switch preferred at that instant. One address
kept moving between two hardware addresses, and the switch&amp;rsquo;s forwarding table
kept being rewritten. At idle there was no traffic to fight over, so
nothing showed. Under a sustained transfer, pushing a NixOS closure (the
complete set of packages a configuration needs) to the target, the session
died repeatedly and looked like a crash.&lt;/p&gt;</description></item><item><title>Three X540 NICs die in one week, three different reasons</title><link>https://oznog.com/node0/lessons/hardware/1-7/</link><pubDate>Mon, 07 Sep 2026 00:00:00 +0000</pubDate><guid>https://oznog.com/node0/lessons/hardware/1-7/</guid><description>&lt;p&gt;The same NIC chip failed on three different hosts inside four days, and
treating that as one story would have been a mistake. It was handling damage
on the first, a silent speed downgrade on the second, and a thermal design
problem on the third.&lt;/p&gt;
&lt;p&gt;The firewall&amp;rsquo;s add-in card lost carrier on both ports twice, once after
physical handling and once 360 seconds after an otherwise untouched reboot,
and was replaced after a third drop. The last one is the clean case: nobody
touched it, and it still let go.&lt;/p&gt;</description></item><item><title>The USB network adapter that was actually a flash drive in disguise</title><link>https://oznog.com/node0/lessons/hardware/1-8/</link><pubDate>Tue, 08 Sep 2026 00:00:00 +0000</pubDate><guid>https://oznog.com/node0/lessons/hardware/1-8/</guid><description>&lt;p&gt;Some USB Ethernet chipsets ship with two personalities. On first plug-in they
present as a tiny CD-ROM holding a Windows driver installer, and they only
become a network adapter once a driver on the host sends them the command to
switch modes. On Linux, with no driver interested in the storage device, the
adapter can sit in that first mode forever.&lt;/p&gt;
&lt;p&gt;On 20260908 one of these adapters did exactly that on env1, a small
environment-monitoring host. Every layer above behaved reasonably and was
still wrong. The switch reported the port up, because something with a
link-layer chip really was plugged in. The bond configuration waited patiently
for an interface that could never appear, because in storage mode there is no
interface at all. Nothing failed, nothing logged an error, and the bond simply
ran on one leg with no redundancy.&lt;/p&gt;</description></item><item><title>HGST drives hide their worst counters, and a reboot can fake a SMART page</title><link>https://oznog.com/node0/lessons/hardware/1-9/</link><pubDate>Wed, 09 Sep 2026 00:00:00 +0000</pubDate><guid>https://oznog.com/node0/lessons/hardware/1-9/</guid><description>&lt;p&gt;Two blind spots compounded each other inside a single week, 20260909 to
20260910.&lt;/p&gt;
&lt;p&gt;The first belongs to the drives. SMART, the self-reporting standard every
monitoring tool reads, lets a manufacturer choose which attributes it exposes,
and this HGST model does not expose attribute 187, the uncorrectable-read
count. Those events are still counted. They live in the drive&amp;rsquo;s device
statistics pages, which almost nothing reads by default. So a drive whose
attribute table looks spotless can already be logging uncorrectable reads and
read-recovery attempts that nobody is watching. A survey of all 57 of these
drives on 20260910 found real outliers that had been invisible until then.&lt;/p&gt;</description></item><item><title>A dirty bay connector wears two different disguises</title><link>https://oznog.com/node0/lessons/hardware/1-10/</link><pubDate>Sun, 06 Sep 2026 00:00:00 +0000</pubDate><guid>https://oznog.com/node0/lessons/hardware/1-10/</guid><description>&lt;p&gt;One physical fault, contamination on the contacts of a single drive bay,
produced two failures that looked nothing alike. That is what made it
expensive.&lt;/p&gt;
&lt;p&gt;The first drive&amp;rsquo;s symptoms were textbook dying disk from the interface side:
CRC errors climbing over four days, and a SATA link that trained itself down
from 6.0 Gb/s to 3.0 Gb/s, which is what a link layer does when it cannot keep
a clean signal. The drive was drained out of Ceph and pulled.&lt;/p&gt;</description></item><item><title>An NVMe controller can die and leave no evidence</title><link>https://oznog.com/node0/lessons/hardware/1-11/</link><pubDate>Thu, 10 Sep 2026 00:00:00 +0000</pubDate><guid>https://oznog.com/node0/lessons/hardware/1-11/</guid><description>&lt;p&gt;An NVMe controller can wedge so completely that the operating system still
sees the device node while every command to it returns as though the
controller has stopped existing. On 20260910 one did exactly that on a storage
host: writes stopped, every reset attempt came back &amp;ldquo;device not ready&amp;rdquo;, a PCI
remove-and-rescan did nothing, and a warm reboot of the entire host did
nothing either. Only removing power from the drive brought it back.&lt;/p&gt;</description></item><item><title>ats6's application link died on its own management card</title><link>https://oznog.com/node0/lessons/hardware/1-12/</link><pubDate>Tue, 08 Sep 2026 00:00:00 +0000</pubDate><guid>https://oznog.com/node0/lessons/hardware/1-12/</guid><description>&lt;p&gt;An automatic transfer switch (an ATS is the device that holds a load on one
power feed and moves it to a second feed if the first fails) is only useful if
it can actually transfer. Between 20260908 and 20260911 one of Node0&amp;rsquo;s could
not, and almost nothing about it looked wrong.&lt;/p&gt;
&lt;p&gt;Its network management card answered every management protocol normally. It
responded to telnet, it answered SNMP queries, it accepted configuration
changes. What had died was the internal channel between the card and the
switching hardware it is bolted to. The tell was a count. A healthy unit
returns 136 status values over SNMP; this one returned one. Its event log
menu had disappeared from the interface entirely. For three days a server
sat on a power path whose failover nobody could confirm.&lt;/p&gt;</description></item><item><title>A five-day-old battery warning fenced a hypervisor</title><link>https://oznog.com/node0/lessons/hardware/1-13/</link><pubDate>Tue, 15 Sep 2026 00:00:00 +0000</pubDate><guid>https://oznog.com/node0/lessons/hardware/1-13/</guid><description>&lt;p&gt;A component can fail completely and keep announcing it once a day for five
days without producing a single alert, if the only thing watching is a rule
built to match one specific log line.&lt;/p&gt;
&lt;p&gt;On 20260915 a hypervisor&amp;rsquo;s RAID controller stopped being patient. Its firmware
faulted, input and output to the host&amp;rsquo;s root pool stalled, the cluster
heartbeat went with it, and the cluster&amp;rsquo;s high-availability watchdog fenced
the node and reset it. That part worked as designed: no data was lost, the
root pool came back with zero errors, and the guests that had been running
were restarted on other nodes automatically.&lt;/p&gt;</description></item><item><title>A raidz3 absorbed an 18 TB drive's sudden death without losing a byte</title><link>https://oznog.com/node0/lessons/hardware/1-14/</link><pubDate>Tue, 15 Sep 2026 00:00:00 +0000</pubDate><guid>https://oznog.com/node0/lessons/hardware/1-14/</guid><description>&lt;p&gt;Two parity disks exist so a single drive&amp;rsquo;s death can pass with no drama.&lt;/p&gt;
&lt;p&gt;On 20260915, during a routine scrub (the pass that reads every block in a ZFS
pool and repairs what does not match its checksum), an 18 TB drive with 297
hours of service and a clean history went from healthy to unreadable. It had
not degraded. It could not read sector 0 at all, and it stopped answering
SMART queries entirely. ZFS counted 181 read and 151 write errors and faulted
it out of its raidz3 group, which is a layout that keeps three parity disks
and can therefore lose three members before any data is at risk.&lt;/p&gt;</description></item><item><title>Half a KVM stopped taking keystrokes, and the fix is a new KVM</title><link>https://oznog.com/node0/lessons/hardware/1-15/</link><pubDate>Fri, 18 Sep 2026 00:00:00 +0000</pubDate><guid>https://oznog.com/node0/lessons/hardware/1-15/</guid><description>&lt;p&gt;A remote console that still shows video is a console that will still be
trusted, which is the whole problem here.&lt;/p&gt;
&lt;p&gt;On 20260918, after a routine restart, one entire bank of Node0&amp;rsquo;s central KVM
switch, eight of its sixteen ports, stopped delivering keyboard input to the
machines behind it. Video switching kept working on every port, so the unit
looked fine from the only angle most people check. Three of the affected hosts
had been logging USB errors for days before the restart, which suggests the
bank had been dying for a while and the restart finished it. Two more restarts
changed nothing.&lt;/p&gt;</description></item><item><title>A laptop loses power exactly every five days, and the reason is out of Linux's reach</title><link>https://oznog.com/node0/lessons/hardware/1-16/</link><pubDate>Sun, 06 Sep 2026 00:00:00 +0000</pubDate><guid>https://oznog.com/node0/lessons/hardware/1-16/</guid><description>&lt;p&gt;A host with a healthy battery, healthy thermals and stable mains power should
not be able to lose power on its own. This one has, three times since
20260906, and it is still doing it as of 20260920.&lt;/p&gt;
&lt;p&gt;There is nothing to read afterwards. No kernel panic, no shutdown sequence, no
last log line that means anything, no network link lights on the way down. The
machine is simply off and needs the physical power button. Each outage on its
own looks like an unrelated glitch.&lt;/p&gt;</description></item><item><title>Proving an RMA: what Node0 learned by doing it</title><link>https://oznog.com/node0/lessons/hardware/1-17/</link><pubDate>Fri, 04 Sep 2026 00:00:00 +0000</pubDate><guid>https://oznog.com/node0/lessons/hardware/1-17/</guid><description>&lt;p&gt;Every return-merchandise process eventually meets the same trap. A
manufacturer&amp;rsquo;s health check is written to describe the drive, not to predict
whether it will survive the next hour of real use. A drive can pass it
convincingly right up until it fails in production for the second time. Two
did here: an SSD that passed SMART twice and failed in service twice, and a 12
TB disk whose own self-assessment read PASSED the entire way through 12,232
reallocated sectors and the storage-level inconsistency those eventually
caused.&lt;/p&gt;</description></item><item><title>Give storage time to prove itself</title><link>https://oznog.com/node0/lessons/hardware/1-18/</link><pubDate>Tue, 22 Sep 2026 00:00:00 +0000</pubDate><guid>https://oznog.com/node0/lessons/hardware/1-18/</guid><description>&lt;p&gt;Storage fails early, and it fails more than anything else. On Node0 that was
not a suspicion but a count. Over the three weeks from bring-up on 20260901,
the hardware theme&amp;rsquo;s incidents split cleanly: two boards, three network
cards, a transfer switch and a KVM on one side, and the drives and their
controllers on the other, with more failures than all of those together. Three of about 80
refurbished drives never spun up. One of three RAID cards was dead. An 18 TB
drive in the hub&amp;rsquo;s main pool died abruptly at 297 power-on hours, in the
middle of a scrub, going from a clean history to unable to read sector 0
(1.14). A SATA SSD passed a health check at rest twice and failed in service
twice, and a 12 TB drive&amp;rsquo;s own self-assessment read PASSED throughout 12,232
reallocated sectors (1.17). Two NVMe controllers wedged a week apart, on
different hosts, and each came back only from a full cold power cycle,
with counters that afterwards read as if nothing had happened (1.11). A dirty
bay connector made two good drives look bad in turn, and cost a six-hour
drain-and-rebuild that was never about a drive (1.10).&lt;/p&gt;</description></item></channel></rss>