Field reference for data center critical facilities technicians. 6 entries, taken from the RackWatt app.
Reference notes from RackWatt, the offline field toolkit for data center critical facilities technicians. It works with no cell signal, because the sites do not have any.
Get RackWatt on the App StoreHow to read early failure signals on a static double conversion UPS before it drops the critical load.
Read the display and record the exact alarm code before clearing anything
Verify the load is within the unit capacity and check percent loaded
Test or trend battery health and compare runtime to the design value
Check for a transfer to bypass and understand why it occurred
Confirm input and output voltage and frequency are in range
Frequent or repeating alarms
Reduced backup runtime versus the last battery test
Battery swelling, bulging, or electrolyte leakage
Abnormal heat at the cabinet or battery string
Louder or cycling cooling fans
Rising room temperature around the unit
Continuous beeping usually means a battery fault, an overload, or an internal component failure.
Runtime scales with load. A reduced runtime alarm can be a genuine battery issue or simply a heavier load than the battery was sized for.
Always verify both the volt ampere rating and the watt rating of the unit against the connected load.
Vertiv Liebert
Schneider APC
Eaton
Mitsubishi
ABB
The most common critical failure: utility power drops and the standby generator does not start.
Confirm the automatic transfer switch is calling for the generator
Check starting battery voltage and connections first, the leading cause of no start
Verify fuel level in the day tank and the main tank and look for contamination
Confirm the block heater and jacket water temperature for reliable cold starts
Review the controller fault log and any lockout condition
Check the transfer and retransfer timing and any failure to transfer
Utility outage with no engine crank
Crank with no start
Low starting battery voltage
Block heater not maintaining jacket water temperature
Fuel level or fuel quality alarms
Automatic transfer switch not signaling a start
By far the most common failure mode is a battery failure or a fuel system fault.
Fuel polishing and periodic load bank testing prevent wet stacking and fuel degradation.
A typical day tank gives about eight hours and can be refueled while running.
Caterpillar
Cummins
Kohler
Generac
ASCO transfer
Russelectric transfer
Switchgear, breakers, power distribution units, remote power panels, and busway degrade with detectable signals.
Thermal imaging of connections and terminations under load
Ultrasound inspection for partial discharge and arcing
Vibration analysis on rotating and mechanically stressed parts
Current monitoring and trending for developing imbalance
Variable frequency drive alarm review
Torque check terminations against the specification during approved outages
Unusual heat at a connection or termination
Burning smell or discoloration
Nuisance trips
Abnormal vibration
Buzzing or crackling
Voltage fluctuation
Corrosion, loose terminations, or insulation damage
Trending over time turns a single reading into an early warning.
Never open energized gear or torque live terminations without an approved procedure and the site arc flash study.
Vertiv
Schneider
Eaton
ABB
Siemens
Cooling failures cause roughly nineteen percent of major outages and rarely start dramatically.
Trend supply and return air temperature and the delta T across the unit
Check filter differential pressure and replace loaded filters
Measure compressor current and compare to nameplate and history
Inspect and clean condenser and evaporator coils
Verify refrigerant charge and look for leaks per EPA Section 608
Confirm control valve and sensor calibration
Rising rack inlet temperature
Increasing filter differential pressure
Abnormal compressor current draw
Fan motor vibration
Refrigerant pressure drift on direct expansion units
Humidifier faults
Rotating assets such as pumps, compressors, and fans degrade quietly long before an alarm.
Cadence matters: weekly filter and refrigerant checks, monthly current and vibration checks, quarterly leak testing and deep coil cleaning.
Liebert
Stulz
Schneider Uniflair
Munters
Trane
Carrier
York
Daikin
The underserved edge. Artificial intelligence racks drive liquid cooling and there is almost no consolidated field reference.
Log coolant flow, pressure, temperature, reservoir level, and leak detection
Sample and log coolant chemistry and fluid quality on schedule
Interpret an unexpected delta T change as a flow restriction, heat exchanger fouling, pump degradation, or cooling imbalance
Swap filters on the secondary loop and verify pump redundancy
Compare secondary loop discharge pressure to the mechanical seal rating
Unexpected change in delta T across a loop
Falling coolant flow or pressure
Reservoir level dropping
Leak detection alarm
Coolant chemistry or fluid quality drift
Pump degradation or loss of redundancy
Direct to chip cooling uses cold plates, micro channel heat exchangers, mounted on processors. Manifolds distribute coolant and a secondary loop, typically propylene glycol and water, carries heat to a Coolant Distribution Unit.
Rear door heat exchangers handle 10 to 30 kilowatts per rack passively, 50 to 75 kilowatts with active fans, and advanced models claim up to 200 kilowatts.
Critical detail: mechanical seal life scales with discharge pressure. A seal that lasted five years on a 2 bar chilled water loop can fail within 12 to 18 months on a 6 bar secondary cold plate loop.
Routine work includes logging coolant chemistry, swapping filters, and verifying pump redundancy.
Vertiv CDU
Motivair
Boyd
CoolIT
Schneider
Building management system, data center infrastructure management, and electrical power monitoring system basics.
Corroborate an alarm with an independent reading or a local gauge before acting
Check whether the point is a real process value or a communications artifact
Verify the sensor against a known reference where possible
Confirm gateway and network health before assuming an equipment fault
A single point alarm with no corroborating readings
A value pinned at zero or full scale
Loss of communications to a controller or gateway
An alarm that clears and returns on a fixed interval
Help the technician tell a real fault from a sensor or communications artifact.
A monitoring artifact wrongly treated as a real fault can trigger an unnecessary and risky intervention.
Vertiv
Schneider EcoStruxure
Trellis
Modius
Niagara
Reference notes from RackWatt, the offline field toolkit for data center critical facilities technicians. It works with no cell signal, because the sites do not have any.
Get RackWatt on the App StoreThese notes are a field aid, not a substitute for the governing codes, the stamped drawings, the authority having jurisdiction, or manufacturer manuals. Verify against the current documentation for your installed equipment.