Field reference for data center critical facilities technicians. 18 entries, taken from the RackWatt app.
These notes are the web copy of the RackWatt reference, built for data center critical facilities technicians. The same library is free inside the app, with no time limit and no account, and it works with no signal.
Reference notes from RackWatt, the offline field toolkit for data center critical facilities technicians. It works with no cell signal, because the sites do not have any.
Get RackWatt on the App StoreHow to read early failure signals on a static double conversion UPS before it drops the critical load.
Read the display and record the exact alarm code before clearing anything
Verify the load is within the unit capacity and check percent loaded
Test or trend battery health and compare runtime to the design value
Check for a transfer to bypass and understand why it occurred
Confirm input and output voltage and frequency are in range
Frequent or repeating alarms
Reduced backup runtime versus the last battery test
Battery swelling, bulging, or electrolyte leakage
Abnormal heat at the cabinet or battery string
Louder or cycling cooling fans
Rising room temperature around the unit
Continuous beeping usually means a battery fault, an overload, or an internal component failure.
Runtime scales with load. A reduced runtime alarm can be a genuine battery issue or simply a heavier load than the battery was sized for.
Always verify both the volt ampere rating and the watt rating of the unit against the connected load.
Vertiv Liebert
Schneider APC
Eaton
Mitsubishi
ABB
True
The most common critical failure: utility power drops and the standby generator does not start.
Confirm the automatic transfer switch is calling for the generator
Check starting battery voltage and connections first, the leading cause of no start
Verify fuel level in the day tank and the main tank and look for contamination
Confirm the block heater and jacket water temperature for reliable cold starts
Review the controller fault log and any lockout condition
Check the transfer and retransfer timing and any failure to transfer
Utility outage with no engine crank
Crank with no start
Low starting battery voltage
Block heater not maintaining jacket water temperature
Fuel level or fuel quality alarms
Automatic transfer switch not signaling a start
By far the most common failure mode is a battery failure or a fuel system fault.
Fuel polishing and periodic load bank testing prevent wet stacking and fuel degradation.
A typical day tank gives about eight hours and can be refueled while running.
Caterpillar
Cummins
Kohler
Generac
ASCO transfer
Russelectric transfer
True
Switchgear, breakers, power distribution units, remote power panels, and busway degrade with detectable signals.
Thermal imaging of connections and terminations under load
Ultrasound inspection for partial discharge and arcing
Vibration analysis on rotating and mechanically stressed parts
Current monitoring and trending for developing imbalance
Variable frequency drive alarm review
Torque check terminations against the specification during approved outages
Unusual heat at a connection or termination
Burning smell or discoloration
Nuisance trips
Abnormal vibration
Buzzing or crackling
Voltage fluctuation
Corrosion, loose terminations, or insulation damage
Trending over time turns a single reading into an early warning.
Never open energized gear or torque live terminations without an approved procedure and the site arc flash study.
Vertiv
Schneider
Eaton
ABB
Siemens
True
Cooling failures cause roughly nineteen percent of major outages and rarely start dramatically.
Trend supply and return air temperature and the delta T across the unit
Check filter differential pressure and replace loaded filters
Measure compressor current and compare to nameplate and history
Inspect and clean condenser and evaporator coils
Verify refrigerant charge and look for leaks per EPA Section 608
Confirm control valve and sensor calibration
Rising rack inlet temperature
Increasing filter differential pressure
Abnormal compressor current draw
Fan motor vibration
Refrigerant pressure drift on direct expansion units
Humidifier faults
Rotating assets such as pumps, compressors, and fans degrade quietly long before an alarm.
Cadence matters: weekly filter and refrigerant checks, monthly current and vibration checks, quarterly leak testing and deep coil cleaning.
Liebert
Stulz
Schneider Uniflair
Munters
Trane
Carrier
York
Daikin
True
The underserved edge. Artificial intelligence racks drive liquid cooling and the field practice is genuinely not written down yet. Treat this section as orientation only. It depends on the vendor manuals and the site specific documentation, and the real source for any given plant is the designers and installers of that project.
Log coolant flow, pressure, temperature, reservoir level, and leak detection
Sample and log coolant chemistry and fluid quality on schedule
Interpret an unexpected delta T change as a flow restriction, heat exchanger fouling, pump degradation, or cooling imbalance
Swap filters on the secondary loop and verify pump redundancy
For coolant loop pre execution checks, take the maximum and minimum pressures from the facility Risk Management Plan and Process Safety Management documentation and the CDU manual, not from a remembered number
Unexpected change in delta T across a loop
Falling coolant flow or pressure
Reservoir level dropping
Leak detection alarm
Coolant chemistry or fluid quality drift
Pump degradation or loss of redundancy
Scope and maturity. This practice is still being written. It is O&M manual technical author work, and even basic questions such as who signs off and what the back out plan is are not yet consolidated across the industry. Defer to the specific manufacturer manual and the project designers and installers.
Pressure limits. Do not work to a remembered figure. Take the maximum and minimum coolant loop pressures from the facility Risk Management Plan and Process Safety Management documentation and the CDU manual. A site running liquid at scale already documents these limits under its safety programme.
Direct to chip cooling uses cold plates, micro channel heat exchangers, mounted on processors. Manifolds distribute coolant and a secondary loop, typically propylene glycol and water, carries heat to a Coolant Distribution Unit.
Rear door heat exchanger capacity figures, roughly 10 to 30 kilowatts passive, 50 to 75 kilowatts with active fans, and up to 200 kilowatts on advanced models, are vendor claims. Confirm the rating for the specific model.
Mechanical seal life scales with discharge pressure. As an illustration only, a seal that lasted years on a low pressure chilled water loop can fail far sooner on a high pressure secondary cold plate loop. Set intervals from the vendor manual and trend seal condition, not from a fixed figure.
Routine work includes logging coolant chemistry, swapping filters, and verifying pump redundancy.
Vertiv CDU
Motivair
Boyd
CoolIT
Schneider
True
Building management system, data center infrastructure management, and electrical power monitoring system basics.
Corroborate an alarm with an independent reading or a local gauge before acting
Check whether the point is a real process value or a communications artifact
Verify the sensor against a known reference where possible
Confirm gateway and network health before assuming an equipment fault
A single point alarm with no corroborating readings
A value pinned at zero or full scale
Loss of communications to a controller or gateway
An alarm that clears and returns on a fixed interval
Help the technician tell a real fault from a sensor or communications artifact.
A monitoring artifact wrongly treated as a real fault can trigger an unnecessary and risky intervention.
Vertiv
Schneider EcoStruxure
Trellis
Modius
Niagara
True
A weak or failed cell drags the whole string down, and the UPS may not tell you until it needs the battery.
Record the alarm and the string voltage before clearing anything
Find whether it is one jar or the whole string, because that decides between a replacement and a string end of life
Take individual jar voltages and impedance readings and compare against the last set rather than against the datasheet
Check terminal torque and look for corrosion at the links, which raises impedance and looks like a failing jar
Check the battery room or cabinet temperature, since every 10 degrees C above 25 roughly halves VRLA life
Check the charger float voltage against specification before condemning cells
Replace as a matched set within a string, never a single jar into an aged string
Run a discharge test after replacement rather than trusting the float reading
Runtime falling short of the design figure on test
A single jar or bloc reading low on the string
Swelling, bulging, or electrolyte weeping
Heat at one point in the string
Battery alarm that clears itself and returns
VRLA strings usually reach end of life between three and five years depending on temperature and cycling.
Impedance trending catches a failing jar long before a runtime test does.
A string that passed last quarter can fail this quarter; the failure is not gradual once it starts.
Vertiv Liebert
Schneider APC
Eaton
Mitsubishi
ABB
True
The load is running on raw utility with no protection. Everything the UPS exists to do is currently not happening.
Treat this as urgent regardless of the load being up, because the critical load is unprotected right now
Establish whether it went to bypass on overload, on an internal fault, or because someone put it there for maintenance
Check percent loaded against unit capacity, since an overload transfer is the UPS behaving correctly
Check for a recent load addition, which is the usual cause of a first time overload transfer
Read the event log for what preceded the transfer rather than the bypass alarm itself
Do not transfer back to inverter until the cause is understood, because transferring into the same fault repeats it
Confirm utility quality before relying on bypass for any length of time
Bypass indication on the display or panel
Loss of the inverter running indication
Alarm coinciding with a load step
Repeated transfers to and from bypass
Static bypass is designed to protect the load, so a transfer is usually the unit working rather than failing.
Maintenance bypass and static bypass are different paths; know which one you are on before working.
Vertiv Liebert
Schneider APC
Eaton
Mitsubishi
ABB
True
The two consumables inside a UPS that fail on a schedule rather than at random, and both give warning if anyone is looking.
Check fan operation on every visit, because a stopped fan cooks the components around it
Check DC bus and AC capacitors against their service life rather than waiting for failure
Compare cabinet temperature at a known load against previous readings
Check air filters and intake paths, since restricted airflow shortens both capacitor and fan life
Plan replacement into a maintenance window rather than reacting to a failure
Record the install date of both, since the interval is the diagnosis
Increasing cabinet temperature at the same load
Fan noise change or a fan not turning
Bulging or vented capacitor cans
Ripple or output quality drifting on test
Electrolytic capacitors typically carry a service life around seven to ten years, shorter when run hot.
Fans are usually five to seven years and are the cheapest preventive replacement in the room.
Vertiv Liebert
Schneider APC
Eaton
True
Different problem from a no start. The engine ran, so fuel, battery and starting are proven, and something shut it down deliberately.
Read the controller shutdown code before resetting, because the controller knows why and the reset erases the display
Check coolant level and temperature, since high coolant temperature is the most common running shutdown
Check oil level and pressure
Check the radiator and louvres for blockage, and confirm the louvres actually opened
Check the jacket water heater has been maintaining temperature, since a cold start under load runs hot
Check for a load step beyond the set's capability at the moment it dropped
Check the fuel supply under load rather than at rest, because a partly blocked filter passes at idle and starves under load
Do not repeatedly restart into the same shutdown
Engine cranks, starts, runs briefly, then stops
Shutdown alarm on the controller
High coolant temperature or low oil pressure alarm
Overspeed or overcrank indication
A shutdown is protection working. Resetting without reading the code discards the only diagnosis you had.
Load bank testing exposes cooling and fuel problems that a monthly no load exercise never will.
Caterpillar
Cummins
Kohler
Generac
True
The generator may be running perfectly while the load stays dark, because the switch between them did not operate.
Confirm the generator is actually up to voltage and frequency, because the switch will not transfer to a source that is not ready
Check the ATS controller for a sensing fault or a locked out condition
Check control power to the switch, which is separate from the power it switches
Check the time delay settings, since a long delay looks identical to a failure to someone watching
Verify the position indication against the physical position rather than trusting the display
Check the operator mechanism and linkage for binding, which is common on switches that rarely operate
Exercise the switch on a schedule, because most ATS failures are found only when they are needed
Generator running with load still on utility or dead
No transfer on a test
Transfer in one direction only
Position indication disagreeing with reality
A transfer switch that never operates between tests is the most likely component to fail when it matters.
Bypass isolation switches let the ATS be serviced without dropping load; know whether the site has one.
ASCO
Russelectric
Generac
Eaton
Schneider
True
Something drew more than the breaker allows or a fault occurred. The breaker did its job and the question is what it protected against.
Do not reset into an unknown fault; establish whether it was overload or a fault first
Check the load on the circuit against its rating and against what it was before, since a rack build out is the usual cause
Check whether it trips immediately on reset or holds, because instant retrip means a fault rather than overload
Inspect the circuit for damage, heat or a compromised connection before energising
Check the breaker itself for heat damage, since a breaker that has interrupted a fault may not be reusable
Check torque on terminations, because a loose connection heats and nuisance trips
Confirm the circuit is not shared with anything that was added without a load calculation
Load dropped on one circuit or rack
Breaker in the tripped position
Burning smell or discoloration at the panel
Repeated trips at the same load point
Reset once only. A breaker that trips twice is telling you something and the second reset risks equipment and people.
Trip curves matter: an instantaneous trip and a long time overload look the same on the handle.
Schneider Square D
Eaton
ABB
Vertiv
Starline busway
True
Nonlinear IT loads produce triplen harmonics that add rather than cancel in the neutral, so the neutral can carry more than any phase.
Measure neutral current and compare it against the phase currents rather than assuming it is the difference
Check transformer temperature at its actual load, since a standard transformer derates badly on harmonic load
Check whether the transformer is K rated for the load it is actually feeding
Take a power quality reading rather than a clamp reading, because harmonic content does not show on a basic meter
Check for a shared neutral serving multiple phases of nonlinear load
Check conductor sizing against the neutral current found, not against phase current
Neutral conductor hotter than the phases
Transformer running hot at moderate load
Higher than expected THD on a power quality reading
Nuisance tripping with no obvious overload
Third harmonic and its multiples add arithmetically in the shared neutral of a three phase four wire system.
A neutral sized to phase current can be undersized in a data hall full of switching power supplies.
Schneider
Eaton
Vertiv
ABB
True
Units switching on and off rapidly, or working against each other with one cooling while another humidifies.
Compare setpoints and deadbands across every unit in the room, because units fighting is nearly always a setpoint disagreement
Check whether the units are on a shared control or running independently, since independent units in one room will fight
Check return air sensor placement and calibration, since a sensor in the wrong place drives the wrong behaviour
Check for a failed unit forcing the others to overwork
Widen the deadband before adjusting setpoints, because a narrow band causes cycling on its own
Check refrigerant charge and airflow if one unit alone is cycling
Verify the humidity setpoints are identical across the room, since humidity fighting wastes more energy than cooling fighting
Compressors starting and stopping within minutes
One unit heating while another cools
One unit humidifying while another dehumidifies
Rising energy use with no load change
Temperature swinging rather than holding
Units fighting each other is one of the most common and most expensive faults in a legacy data hall.
Short cycling destroys compressors faster than continuous running does.
Vertiv Liebert
Schneider Uniflair
Stulz
Data Aire
True
Cooling capacity is adequate but the air is not reaching the equipment, which is an airflow problem rather than a cooling problem.
Measure at the rack inlet rather than in the aisle, because the room can be cold while a server inlet is hot
Check for missing blanking panels, which let hot air recirculate through the rack front to back
Check floor tile placement, since perforated tiles in the hot aisle actively make things worse
Check for unsealed cable cutouts under racks, which is where most bypass air escapes
Check containment doors and panels are closed and intact
Check underfloor obstruction from accumulated cabling before adding capacity
Only consider more cooling once airflow is proven, because more cold air into a bypass path changes nothing
Hot spots at the top of racks while the room is cool
Inlet temperature varying widely between racks
High return air temperature with low supply temperature
Adding cooling not fixing the hot spot
Most hot spots are airflow management rather than insufficient cooling capacity.
Blanking panels and sealed cutouts are the cheapest thermal fix available in any hall.
Vertiv
Schneider
Stulz
Upsite
Subzero
True
Liquid near energised IT and power. The alarm must be treated as real until objectively established otherwise.
Treat the alarm as real until objectively established otherwise, and never clear it on the absence of visible liquid alone
Notify operations before intervening, because liquid near energised equipment affects more than the cooling system
Inspect the triggered detection point and the fittings, manifolds and quick disconnects adjacent to it
Check reservoir level against its last recorded value rather than against the sight glass alone
Check whether the system is holding pressure or vacuum, since a slow loss confirms a leak that inspection may miss
Trace along the path rather than only at the sensor, because coolant runs before it pools
Do not restore cooling to a suspected leaking loop without the vendor procedure
Leak detection alarm at a zone
Coolant level falling in the reservoir
Damp or residue at a fitting or quick disconnect
Vacuum or pressure not holding
This practice is not yet standardised across the industry; the vendor manual and the project designers are the authority.
Take maximum and minimum coolant loop pressures from the facility RMP and PSM documentation, not from memory.
Vertiv
Schneider
Motivair
CoolIT
Boyd
True
The coolant distribution unit cannot hold its pressure or flow band, which derates or drops the racks it feeds.
Take the pressure limits from the facility RMP and PSM documentation and the CDU manual rather than from a remembered figure
Check the filter or strainer first, since a loading filter is the most common cause of falling flow
Check pump operation and whether it is running at maximum to hold setpoint, which indicates restriction
Check for air in the loop, which produces unstable pressure and noisy pumps
Check quick disconnects at the rack manifolds for partial engagement
Check the secondary loop for a closed or partly closed isolation valve after any maintenance
Prioritise seal inspection on higher pressure loops, since mechanical seal life scales with discharge pressure
Pressure differential outside the band
Flow rate below setpoint
Pump running at maximum with no result
Rack level thermal alarms downstream
Mechanical seal intervals come from the vendor manual and trended condition, not from a fixed number.
A seal that lasted years on a low pressure chilled water loop can fail far sooner on a high pressure cold plate loop.
Vertiv
Schneider
Motivair
CoolIT
Boyd
True
So many alarms that the real one is invisible. The failure is the monitoring system rather than the plant.
Find the first alarm in the sequence, because the rest are usually consequences of it
Check whether a single failed sensor or comm path is generating the cascade
Check thresholds against real operating ranges, since a threshold set inside normal variation alarms forever
Check for a device alarming and clearing on a cycle, which points at a threshold sitting exactly on the operating point
Widen deadbands and add delays on points that chatter rather than disabling them
Never disable an alarm to quiet a storm without recording it, because a disabled point protects nothing
Review the alarm list quarterly, since alarm quality degrades quietly as a site changes
Hundreds of alarms in a short window
The same point alarming and clearing repeatedly
Staff routinely acknowledging without reading
Real events missed inside the noise
Alarm fatigue is a real failure mode; a system nobody trusts is worse than no system.
A point that alarms daily and is always ignored is not monitoring, it is noise.
Vertiv
Schneider EcoStruxure
Nlyte
Sunbird
Modius
True
Reference notes from RackWatt, the offline field toolkit for data center critical facilities technicians. It works with no cell signal, because the sites do not have any.
Get RackWatt on the App StoreThese notes are a field aid, not a substitute for the governing codes, the stamped drawings, the authority having jurisdiction, or manufacturer manuals. Verify against the current documentation for your installed equipment.