<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Power and environment on Oznog</title><link>https://oznog.com/node0/lessons/power/</link><description>Recent content in Power and environment on Oznog</description><generator>Hugo</generator><language>en-us</language><copyright>© 2011-2026 &lt;a href="https://oznog.com"&gt;Oznog Holdings LLC&lt;/a&gt; · Founded by &lt;a href="https://chrisplough.com" target="_blank" rel="noopener"&gt;Christoph Plough&lt;/a&gt; · &lt;a href="https://oznog.com/index.xml"&gt;RSS&lt;/a&gt; · &lt;a href="https://creativecommons.org/licenses/by/4.0/" target="_blank" rel="noopener"&gt;CC BY 4.0&lt;/a&gt;</copyright><lastBuildDate>Sun, 20 Sep 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://oznog.com/node0/lessons/power/index.xml" rel="self" type="application/rss+xml"/><item><title>A 208/240 V mismatch that looked like a bad transfer switch</title><link>https://oznog.com/node0/lessons/power/5-1/</link><pubDate>Mon, 24 Aug 2026 00:00:00 +0000</pubDate><guid>https://oznog.com/node0/lessons/power/5-1/</guid><description>&lt;p&gt;An automatic transfer switch, or ATS, is a device with two power inputs and one
output. It picks a healthy source and switches to the other one if the first
goes away. It decides &amp;ldquo;healthy&amp;rdquo; by measuring the input against a voltage
window. That is the whole mechanism, and it is why this failure is so
convincing.&lt;/p&gt;
&lt;p&gt;Node0&amp;rsquo;s four APC Smart-UPS SRT 5000 units are all rated for both 208 V and 240
V output. Two had been left at one setting and two at the other, which nobody
noticed because each unit worked perfectly on its own. The rack fed by one of
each saw its preferred input sitting outside the window the transfer switch
expected, reported that source as not OK, and stayed parked on the other one.
Everything downstream kept running, so the only visible symptom was a transfer
switch that appeared to have lost a source, on a rack where both sources were
live and correct.&lt;/p&gt;</description></item><item><title>Two labels, corrected three times, before anyone traced a cable</title><link>https://oznog.com/node0/lessons/power/5-2/</link><pubDate>Thu, 20 Aug 2026 00:00:00 +0000</pubDate><guid>https://oznog.com/node0/lessons/power/5-2/</guid><description>&lt;p&gt;A rack PDU has a field for the name of the supply feeding it. Whoever wires the
rack fills it in. From that moment the field looks like a fact, and every piece
of software that reads the PDU inherits it.&lt;/p&gt;
&lt;p&gt;Node0 filled those fields in with good intent on the day the PDUs were
configured. Four days later, when two of the four UPS units were changed to a
different output voltage, the electrical fingerprints of the supplies diverged
and the mapping was visibly crossed. Christoph corrected it at the rack about
two weeks later, and that correction was also wrong. The answer that held came
on 20260909, from following each cable one at a time and checking the result
against NetBox&amp;rsquo;s own power-port records, which are kept independently of
whatever the device says about itself.&lt;/p&gt;</description></item><item><title>Three battery packs is not a fault, it is a vocabulary problem</title><link>https://oznog.com/node0/lessons/power/5-3/</link><pubDate>Mon, 24 Aug 2026 00:00:00 +0000</pubDate><guid>https://oznog.com/node0/lessons/power/5-3/</guid><description>&lt;p&gt;The UPS management interface says the unit has three battery packs. The unit
has five batteries. Neither number is wrong, and the gap between them cost most
of a day.&lt;/p&gt;
&lt;p&gt;APC calls an external battery enclosure a pack. That reads naturally as &amp;ldquo;a
battery&amp;rdquo;. Each external pack on these units holds two separate replaceable
battery cartridges, and the internal bay looks like it should mirror that but
physically holds one. So three packs is one internal cartridge plus two
enclosures of two, which is five cartridges. The interface also shows a
placeholder install date from the year 2000 for a slot that does not exist,
which looks exactly like a cartridge that was never registered.&lt;/p&gt;</description></item><item><title>A management card that lies quietly when two callers ask at once</title><link>https://oznog.com/node0/lessons/power/5-4/</link><pubDate>Tue, 01 Sep 2026 00:00:00 +0000</pubDate><guid>https://oznog.com/node0/lessons/power/5-4/</guid><description>&lt;p&gt;Older embedded management cards often accept one interactive session at a time.
That limit is not the problem. The problem is how they refuse the second one.
They give no error and no busy message, only an empty response that a script
cannot tell apart from a device that is gone.&lt;/p&gt;
&lt;p&gt;Node0 ran two scheduled jobs against the same four UPS network cards, a nightly
configuration backup and a health check every ten minutes. When their schedules
overlapped, one of them got nothing back and the monitoring stack published
&amp;ldquo;UPS unreachable&amp;rdquo;, then &amp;ldquo;UPS fine&amp;rdquo; on the next pass. That pattern repeated for
weeks. Nothing external ever touched the cards. The false alarm was entirely
self-inflicted, and the shape of it, a device that is up flapping to down and
back, is exactly what a real intermittent fault looks like.&lt;/p&gt;</description></item><item><title>The whole fleet had two minutes to shut down a 218 TiB array</title><link>https://oznog.com/node0/lessons/power/5-5/</link><pubDate>Thu, 03 Sep 2026 00:00:00 +0000</pubDate><guid>https://oznog.com/node0/lessons/power/5-5/</guid><description>&lt;p&gt;NUT is the Network UPS Tools suite, and its monitoring daemon does not think in
minutes at all. It watches whether each UPS reports itself online or on
battery, and it removes a UPS from a host&amp;rsquo;s count of working supplies only at
the moment that UPS declares low battery. An OSD is a Ceph object storage
daemon, one per disk; there are a lot of them, and they do not stop instantly.&lt;/p&gt;</description></item><item><title>A UPS service touch dropped a rack, and two hosts needed a person</title><link>https://oznog.com/node0/lessons/power/5-6/</link><pubDate>Sun, 06 Sep 2026 00:00:00 +0000</pubDate><guid>https://oznog.com/node0/lessons/power/5-6/</guid><description>&lt;p&gt;This was planned maintenance, not a failure. A UPS was unplugged on purpose,
for service, at 17:53 on 20260906. It still produced a five-minute outage
across a rack, and for two machines it produced an outage that could not be
ended remotely at all.&lt;/p&gt;
&lt;p&gt;A BMC is a baseboard management controller, the small always-on computer on a
server board that lets you power the machine on over the network when the
machine itself is off. Two of the hosts in that rack do not have one. When
their power came back they stayed in soft-off, and no amount of remote access
could change that, because there was nothing powered to receive the command.
Someone had to walk to the rack and press the button. A third host came back
ten minutes later on its own, through a BIOS boot-order fallback rather than by
design, which is not the same as working.&lt;/p&gt;</description></item><item><title>A BMC that reported fabricated power and thermal readings</title><link>https://oznog.com/node0/lessons/power/5-7/</link><pubDate>Tue, 01 Sep 2026 00:00:00 +0000</pubDate><guid>https://oznog.com/node0/lessons/power/5-7/</guid><description>&lt;p&gt;Redfish is the modern, well-documented HTTP API for server management. IPMI,
reached through the ipmitool command, is the older one it was meant to replace.
On this Supermicro board family the newer API&amp;rsquo;s power and thermal endpoints are
effectively unimplemented stubs, and they do not say so.&lt;/p&gt;
&lt;p&gt;Asked for Power, the BMC returned 18 watts. Asked for Thermal, it returned
every fan at 0 RPM and no temperatures at all. That is a coherent picture of a
machine that failed to power on, which is exactly what someone was checking
for. The machine was running, with its fans at 9,300 to 9,500 RPM, and ipmitool
said so from the same BMC, over the same network, at the same moment.&lt;/p&gt;</description></item><item><title>Telnet on purpose, and the credential leak that proved the split mattered</title><link>https://oznog.com/node0/lessons/power/5-8/</link><pubDate>Mon, 24 Aug 2026 00:00:00 +0000</pubDate><guid>https://oznog.com/node0/lessons/power/5-8/</guid><description>&lt;p&gt;The decision to keep telnet was explicit and, on its own terms, sound. These
cards&amp;rsquo; SSH offering was not meaningfully better than cleartext: old
ciphers, or a host key type modern clients refuse outright, on firmware
that will never be updated again. The PDUs on the same segment have no SSH
option at all. The real boundary was an isolated management network plus
per-device access control, not the wire protocol. That reasoning still
holds.&lt;/p&gt;</description></item><item><title>A secret rotation that silently did nothing</title><link>https://oznog.com/node0/lessons/power/5-9/</link><pubDate>Mon, 07 Sep 2026 00:00:00 +0000</pubDate><guid>https://oznog.com/node0/lessons/power/5-9/</guid><description>&lt;p&gt;This is a narrow trap with a wide blast radius, for anyone combining NixOS
secret management with a module that renders configuration text at runtime.&lt;/p&gt;
&lt;p&gt;The setup is ordinary and correct. Some NixOS options write their rendered
configuration file straight into the Nix store, which is world-readable on the
host, so a secret placed there is readable by any local user. The usual
workaround is to keep a template in the store with placeholders where the
secrets go, and have a small unit substitute the real values into a runtime
file at service start.&lt;/p&gt;</description></item><item><title>103% of a UPS, and a nameplate rating that was wrong for months</title><link>https://oznog.com/node0/lessons/power/5-10/</link><pubDate>Wed, 09 Sep 2026 00:00:00 +0000</pubDate><guid>https://oznog.com/node0/lessons/power/5-10/</guid><description>&lt;p&gt;Two lessons are braided together here, and the second one is the more
transferable.&lt;/p&gt;
&lt;p&gt;The first is about the question. On a dual-fed rack, per-unit load
percentage is the wrong number to watch, because both units can sit at a
comfortable-looking figure and still be unable to carry each other. The
number that matters is whether the pair&amp;rsquo;s total fits inside one unit&amp;rsquo;s
capacity. Node0&amp;rsquo;s busiest rack does not: reviewed at 103% on
20260909, worse by 20260912 after a planned deployment, and accepted as
permanent. Christoph declined to shed load onto the other rack, whose
headroom is reserved for GPU hardware not yet installed. On mains, losing a
partner puts the survivor into unprotected bypass; on battery, output drops
within seconds with no ordered shutdown. That is an N+0 posture chosen
deliberately, with a dated expiry on the silence that hides the alert, not
an oversight.&lt;/p&gt;</description></item><item><title>A shutdown that works perfectly and still leaves the fleet dark forever</title><link>https://oznog.com/node0/lessons/power/5-11/</link><pubDate>Thu, 10 Sep 2026 00:00:00 +0000</pubDate><guid>https://oznog.com/node0/lessons/power/5-11/</guid><description>&lt;p&gt;Every server board has a BIOS setting for what to do when AC power returns:
stay off, power on, or restore the last state. Last State is the common default
and it sounds like the safe answer. It is not, if your shutdown design works.&lt;/p&gt;
&lt;p&gt;A clean shutdown leaves a chassis in soft-off. Last State faithfully restores
exactly that. So a textbook UPS-triggered shutdown, executed perfectly across
the fleet, guarantees the fleet stays dark when mains returns, with no error
anywhere, until someone walks to each machine. Node0 had already treated its
shutdown design as finished, proven by successful shutdowns, before anyone
asked the other half of the question. Testing &amp;ldquo;did it stop safely&amp;rdquo; is easy.
Testing &amp;ldquo;will it come back&amp;rdquo; needs an actual power interruption, which is why it
gets skipped.&lt;/p&gt;</description></item><item><title>An agent caused a hardware fault, and it is on the record</title><link>https://oznog.com/node0/lessons/power/5-12/</link><pubDate>Fri, 11 Sep 2026 00:00:00 +0000</pubDate><guid>https://oznog.com/node0/lessons/power/5-12/</guid><description>&lt;p&gt;The proximate cause was a mismatched or defective controller board. The
transfer switch reported itself as a different, lower-voltage product than the
chassis and network card around it actually were, and read both of its inputs
at about half the real line voltage. Nobody, human or agent, had a reason to
doubt that self-report before acting on it. Setting the configured line voltage
to match what the unit appeared to be was a reasonable change given the
information available, and it was made with an operator&amp;rsquo;s explicit approval.&lt;/p&gt;</description></item><item><title>Real outages and planned outages need different UPS commands</title><link>https://oznog.com/node0/lessons/power/5-13/</link><pubDate>Sun, 13 Sep 2026 00:00:00 +0000</pubDate><guid>https://oznog.com/node0/lessons/power/5-13/</guid><description>&lt;p&gt;Getting a fleet to come back on its own requires the UPS to cut its own output
and restart it when mains returns, so that the return is a real AC loss and
return event the BIOS has something to react to. NUT exposes one command for
this, shutdown.return. Rehearsing it on an unloaded unit, with the fleet still
up, turned up something a reading of the documentation does not show. The UPS
accepts that command only while it is actually running on battery. Sent with
mains present, which is the planned-maintenance case, it is simply refused.&lt;/p&gt;</description></item><item><title>The route to the power controllers ran through the thing they had to survive</title><link>https://oznog.com/node0/lessons/power/5-14/</link><pubDate>Sun, 13 Sep 2026 00:00:00 +0000</pubDate><guid>https://oznog.com/node0/lessons/power/5-14/</guid><description>&lt;p&gt;The whole power-return design rests on one command reaching each UPS during an
outage: cut your output, and bring it back when mains returns. Somebody asked
the unglamorous question of what that command actually crosses on its way, and
ran the route lookup rather than assuming.&lt;/p&gt;
&lt;p&gt;Both of the servers that manage the UPS network cards reached them through the
firewall pair, by ordinary default routing. Nothing was misconfigured. It was
simply the path the routing table produced, and no one had looked.&lt;/p&gt;</description></item><item><title>A knocked-loose USB cable went unnoticed for four days</title><link>https://oznog.com/node0/lessons/power/5-15/</link><pubDate>Wed, 09 Sep 2026 00:00:00 +0000</pubDate><guid>https://oznog.com/node0/lessons/power/5-15/</guid><description>&lt;p&gt;A small story with a sharp point. On 20260909, during unrelated work in the
rack, the USB cable between fw1 and its own UPS came loose. The UPS kept
working. The firewall kept working. The monitoring daemon simply had nothing on
the other end of the bus, and said nothing about it. It was found on 20260913,
four days later, because a scheduled three-hour maintenance pass happened to
include a physical look at that exact cable.&lt;/p&gt;</description></item><item><title>An interim fix with no new part solved a real heat problem</title><link>https://oznog.com/node0/lessons/power/5-16/</link><pubDate>Thu, 27 Aug 2026 00:00:00 +0000</pubDate><guid>https://oznog.com/node0/lessons/power/5-16/</guid><description>&lt;p&gt;Before any active cooling, the server room sat at about 90 degrees
Fahrenheit and there was no way to bring it lower. The planned answer was a
dedicated mini-split, already sized and quoted, scheduled for 2027. The
answer that actually happened on 20260827 was to redirect a branch of
cooling capacity that already existed in the building into the room, which
held it around 78 degrees and cost nothing in hardware.&lt;/p&gt;</description></item><item><title>The environment project made its own plan obsolete, and the team let it</title><link>https://oznog.com/node0/lessons/power/5-17/</link><pubDate>Sat, 12 Sep 2026 00:00:00 +0000</pubDate><guid>https://oznog.com/node0/lessons/power/5-17/</guid><description>&lt;p&gt;The plan was reasonable when it was written. Build an environment logger on
Home Assistant, feed it room sensors and machine telemetry, and have one place
to look at how hot the room gets and why. Work started.&lt;/p&gt;
&lt;p&gt;Then, partway through, someone noticed that the machine half of it already
existed. BMC temperatures, fan speeds, PDU and UPS wattages had all arrived in
the fleet&amp;rsquo;s Prometheus stack on their own, as a side effect of unrelated
monitoring work finished earlier, with alerting and dashboards already built
around them. The logger under construction was going to be a second, worse copy
of a timeline that was already live.&lt;/p&gt;</description></item><item><title>Following one sensor sounds simple; the thermostat quietly refused for a day</title><link>https://oznog.com/node0/lessons/power/5-18/</link><pubDate>Sat, 12 Sep 2026 00:00:00 +0000</pubDate><guid>https://oznog.com/node0/lessons/power/5-18/</guid><description>&lt;p&gt;Cooling for the server room is driven by a smart thermostat, and the whole
point of the setup is that it should decide from a remote sensor placed by the
racks rather than from its own built-in sensor somewhere else. The setting for
that exists. It was applied through the home-automation integration. It
reported success. The thermostat ignored it for most of a day.&lt;/p&gt;
&lt;p&gt;The detection problem is what makes this worth writing down. In the automation
layer, &amp;ldquo;following the remote sensor&amp;rdquo; and &amp;ldquo;following its own sensor&amp;rdquo; present as
the same underlying characteristic, so there is nothing to read back that
distinguishes them. The change was accepted, appeared to apply, and was never
persisted by the device, with no error surfaced anywhere in the chain. Two more
benign explanations were investigated and ruled out first, a temperature hold
retaining an older sensor selection and the active comfort period not reloading
a changed setting, both of which cost real time. The actual fix was a physical
power cycle of the thermostat, and the only place the truth was ever visible
was the device&amp;rsquo;s own screen.&lt;/p&gt;</description></item><item><title>Blown fuses are invisible to the eye, and one already cost a transfer switch</title><link>https://oznog.com/node0/lessons/power/5-19/</link><pubDate>Thu, 20 Aug 2026 00:00:00 +0000</pubDate><guid>https://oznog.com/node0/lessons/power/5-19/</guid><description>&lt;p&gt;A small, concrete fact for anyone running rack PDUs with internal branch fuses.
You cannot tell a blown fuse from a good one by looking at it. There is no
discoloured element, no visible break, nothing. Sixteen of them across four
PDUs look exactly alike whether they are carrying current or not.&lt;/p&gt;
&lt;p&gt;Node0 learned this the expensive way rather than the theoretical way. A
transfer switch failed and shorted, taking a branch fuse with it. The transfer
switch itself was destroyed and replaced from spares. Finding which fuse had
gone was the part that could have burned the most time, because the instinct is
to open the PDU and look, and looking tells you nothing.&lt;/p&gt;</description></item><item><title>Two power meters, 5% apart, and neither one is wrong</title><link>https://oznog.com/node0/lessons/power/5-20/</link><pubDate>Sun, 20 Sep 2026 00:00:00 +0000</pubDate><guid>https://oznog.com/node0/lessons/power/5-20/</guid><description>&lt;p&gt;Node0 measures its own electrical load twice, from the UPS units and from the
PDUs, for the ordinary reason that both devices report it and both were already
being polled. On 20260920 a routine cross-check put the two figures side by
side at the same moment and they were 5 to 6% apart. That went onto
the list as an unresolved discrepancy.&lt;/p&gt;
&lt;p&gt;It was not one. The first thing to rule out was a timing artefact, since the
two sets of numbers come from independently polled exporters and a lag of a few
seconds between them would show up as a gap on a changing load. Reading both
over a longer window with matched timestamps settled it. The offset holds
steady at every moment, including while load moves, which is not what sampling
lag looks like.&lt;/p&gt;</description></item></channel></rss>