Troubleshooting Reference

Field reference for data center critical facilities technicians. 18 entries, taken from the RackWatt app.

Free offline app

These notes are the web copy of the RackWatt reference, built for data center critical facilities technicians. The same library is free inside the app, with no time limit and no account, and it works with no signal.

Get RackWatt on the App Store
These notes come from RackWatt

Reference notes from RackWatt, the offline field toolkit for data center critical facilities technicians. It works with no cell signal, because the sites do not have any.

Get RackWatt on the App Store
ups

Uninterruptible power supply warning signs#

How to read early failure signals on a static double conversion UPS before it drops the critical load.

Method

Read the display and record the exact alarm code before clearing anything

Verify the load is within the unit capacity and check percent loaded

Test or trend battery health and compare runtime to the design value

Check for a transfer to bypass and understand why it occurred

Confirm input and output voltage and frequency are in range

Warning signs

Frequent or repeating alarms

Reduced backup runtime versus the last battery test

Battery swelling, bulging, or electrolyte leakage

Abnormal heat at the cabinet or battery string

Louder or cycling cooling fans

Rising room temperature around the unit

Notes

Continuous beeping usually means a battery fault, an overload, or an internal component failure.

Runtime scales with load. A reduced runtime alarm can be a genuine battery issue or simply a heavier load than the battery was sized for.

Always verify both the volt ampere rating and the watt rating of the unit against the connected load.

Brands

Vertiv Liebert

Schneider APC

Eaton

Mitsubishi

ABB

Verify against the manual

True

generators

Generator fails to start on utility loss#

The most common critical failure: utility power drops and the standby generator does not start.

Method

Confirm the automatic transfer switch is calling for the generator

Check starting battery voltage and connections first, the leading cause of no start

Verify fuel level in the day tank and the main tank and look for contamination

Confirm the block heater and jacket water temperature for reliable cold starts

Review the controller fault log and any lockout condition

Check the transfer and retransfer timing and any failure to transfer

Warning signs

Utility outage with no engine crank

Crank with no start

Low starting battery voltage

Block heater not maintaining jacket water temperature

Fuel level or fuel quality alarms

Automatic transfer switch not signaling a start

Notes

By far the most common failure mode is a battery failure or a fuel system fault.

Fuel polishing and periodic load bank testing prevent wet stacking and fuel degradation.

A typical day tank gives about eight hours and can be refueled while running.

Brands

Caterpillar

Cummins

Kohler

Generac

ASCO transfer

Russelectric transfer

Verify against the manual

True

distribution

Electrical distribution warning signs#

Switchgear, breakers, power distribution units, remote power panels, and busway degrade with detectable signals.

Method

Thermal imaging of connections and terminations under load

Ultrasound inspection for partial discharge and arcing

Vibration analysis on rotating and mechanically stressed parts

Current monitoring and trending for developing imbalance

Variable frequency drive alarm review

Torque check terminations against the specification during approved outages

Warning signs

Unusual heat at a connection or termination

Burning smell or discoloration

Nuisance trips

Abnormal vibration

Buzzing or crackling

Voltage fluctuation

Corrosion, loose terminations, or insulation damage

Notes

Trending over time turns a single reading into an early warning.

Never open energized gear or torque live terminations without an approved procedure and the site arc flash study.

Brands

Vertiv

Schneider

Eaton

ABB

Siemens

Verify against the manual

True

airCooling

Air cooling degradation and failure#

Cooling failures cause roughly nineteen percent of major outages and rarely start dramatically.

Method

Trend supply and return air temperature and the delta T across the unit

Check filter differential pressure and replace loaded filters

Measure compressor current and compare to nameplate and history

Inspect and clean condenser and evaporator coils

Verify refrigerant charge and look for leaks per EPA Section 608

Confirm control valve and sensor calibration

Warning signs

Rising rack inlet temperature

Increasing filter differential pressure

Abnormal compressor current draw

Fan motor vibration

Refrigerant pressure drift on direct expansion units

Humidifier faults

Notes

Rotating assets such as pumps, compressors, and fans degrade quietly long before an alarm.

Cadence matters: weekly filter and refrigerant checks, monthly current and vibration checks, quarterly leak testing and deep coil cleaning.

Brands

Liebert

Stulz

Schneider Uniflair

Munters

Trane

Carrier

York

Daikin

Verify against the manual

True

liquidCooling

Liquid cooling field reference#

The underserved edge. Artificial intelligence racks drive liquid cooling and the field practice is genuinely not written down yet. Treat this section as orientation only. It depends on the vendor manuals and the site specific documentation, and the real source for any given plant is the designers and installers of that project.

Method

Log coolant flow, pressure, temperature, reservoir level, and leak detection

Sample and log coolant chemistry and fluid quality on schedule

Interpret an unexpected delta T change as a flow restriction, heat exchanger fouling, pump degradation, or cooling imbalance

Swap filters on the secondary loop and verify pump redundancy

For coolant loop pre execution checks, take the maximum and minimum pressures from the facility Risk Management Plan and Process Safety Management documentation and the CDU manual, not from a remembered number

Warning signs

Unexpected change in delta T across a loop

Falling coolant flow or pressure

Reservoir level dropping

Leak detection alarm

Coolant chemistry or fluid quality drift

Pump degradation or loss of redundancy

Notes

Scope and maturity. This practice is still being written. It is O&M manual technical author work, and even basic questions such as who signs off and what the back out plan is are not yet consolidated across the industry. Defer to the specific manufacturer manual and the project designers and installers.

Pressure limits. Do not work to a remembered figure. Take the maximum and minimum coolant loop pressures from the facility Risk Management Plan and Process Safety Management documentation and the CDU manual. A site running liquid at scale already documents these limits under its safety programme.

Direct to chip cooling uses cold plates, micro channel heat exchangers, mounted on processors. Manifolds distribute coolant and a secondary loop, typically propylene glycol and water, carries heat to a Coolant Distribution Unit.

Rear door heat exchanger capacity figures, roughly 10 to 30 kilowatts passive, 50 to 75 kilowatts with active fans, and up to 200 kilowatts on advanced models, are vendor claims. Confirm the rating for the specific model.

Mechanical seal life scales with discharge pressure. As an illustration only, a seal that lasted years on a low pressure chilled water loop can fail far sooner on a high pressure secondary cold plate loop. Set intervals from the vendor manual and trend seal condition, not from a fixed figure.

Routine work includes logging coolant chemistry, swapping filters, and verifying pump redundancy.

Brands

Vertiv CDU

Motivair

Boyd

CoolIT

Schneider

Verify against the manual

True

monitoring

Monitoring systems and false alarms#

Building management system, data center infrastructure management, and electrical power monitoring system basics.

Method

Corroborate an alarm with an independent reading or a local gauge before acting

Check whether the point is a real process value or a communications artifact

Verify the sensor against a known reference where possible

Confirm gateway and network health before assuming an equipment fault

Warning signs

A single point alarm with no corroborating readings

A value pinned at zero or full scale

Loss of communications to a controller or gateway

An alarm that clears and returns on a fixed interval

Notes

Help the technician tell a real fault from a sensor or communications artifact.

A monitoring artifact wrongly treated as a real fault can trigger an unnecessary and risky intervention.

Brands

Vertiv

Schneider EcoStruxure

Trellis

Modius

Niagara

Verify against the manual

True

ups

UPS battery string failure#

A weak or failed cell drags the whole string down, and the UPS may not tell you until it needs the battery.

Method

Record the alarm and the string voltage before clearing anything

Find whether it is one jar or the whole string, because that decides between a replacement and a string end of life

Take individual jar voltages and impedance readings and compare against the last set rather than against the datasheet

Check terminal torque and look for corrosion at the links, which raises impedance and looks like a failing jar

Check the battery room or cabinet temperature, since every 10 degrees C above 25 roughly halves VRLA life

Check the charger float voltage against specification before condemning cells

Replace as a matched set within a string, never a single jar into an aged string

Run a discharge test after replacement rather than trusting the float reading

Warning signs

Runtime falling short of the design figure on test

A single jar or bloc reading low on the string

Swelling, bulging, or electrolyte weeping

Heat at one point in the string

Battery alarm that clears itself and returns

Notes

VRLA strings usually reach end of life between three and five years depending on temperature and cycling.

Impedance trending catches a failing jar long before a runtime test does.

A string that passed last quarter can fail this quarter; the failure is not gradual once it starts.

Brands

Vertiv Liebert

Schneider APC

Eaton

Mitsubishi

ABB

Verify against the manual

True

ups

UPS transferred to bypass#

The load is running on raw utility with no protection. Everything the UPS exists to do is currently not happening.

Method

Treat this as urgent regardless of the load being up, because the critical load is unprotected right now

Establish whether it went to bypass on overload, on an internal fault, or because someone put it there for maintenance

Check percent loaded against unit capacity, since an overload transfer is the UPS behaving correctly

Check for a recent load addition, which is the usual cause of a first time overload transfer

Read the event log for what preceded the transfer rather than the bypass alarm itself

Do not transfer back to inverter until the cause is understood, because transferring into the same fault repeats it

Confirm utility quality before relying on bypass for any length of time

Warning signs

Bypass indication on the display or panel

Loss of the inverter running indication

Alarm coinciding with a load step

Repeated transfers to and from bypass

Notes

Static bypass is designed to protect the load, so a transfer is usually the unit working rather than failing.

Maintenance bypass and static bypass are different paths; know which one you are on before working.

Brands

Vertiv Liebert

Schneider APC

Eaton

Mitsubishi

ABB

Verify against the manual

True

ups

Capacitor aging and fan wear#

The two consumables inside a UPS that fail on a schedule rather than at random, and both give warning if anyone is looking.

Method

Check fan operation on every visit, because a stopped fan cooks the components around it

Check DC bus and AC capacitors against their service life rather than waiting for failure

Compare cabinet temperature at a known load against previous readings

Check air filters and intake paths, since restricted airflow shortens both capacitor and fan life

Plan replacement into a maintenance window rather than reacting to a failure

Record the install date of both, since the interval is the diagnosis

Warning signs

Increasing cabinet temperature at the same load

Fan noise change or a fan not turning

Bulging or vented capacitor cans

Ripple or output quality drifting on test

Notes

Electrolytic capacitors typically carry a service life around seven to ten years, shorter when run hot.

Fans are usually five to seven years and are the cheapest preventive replacement in the room.

Brands

Vertiv Liebert

Schneider APC

Eaton

Verify against the manual

True

generators

Generator starts then shuts down#

Different problem from a no start. The engine ran, so fuel, battery and starting are proven, and something shut it down deliberately.

Method

Read the controller shutdown code before resetting, because the controller knows why and the reset erases the display

Check coolant level and temperature, since high coolant temperature is the most common running shutdown

Check oil level and pressure

Check the radiator and louvres for blockage, and confirm the louvres actually opened

Check the jacket water heater has been maintaining temperature, since a cold start under load runs hot

Check for a load step beyond the set's capability at the moment it dropped

Check the fuel supply under load rather than at rest, because a partly blocked filter passes at idle and starves under load

Do not repeatedly restart into the same shutdown

Warning signs

Engine cranks, starts, runs briefly, then stops

Shutdown alarm on the controller

High coolant temperature or low oil pressure alarm

Overspeed or overcrank indication

Notes

A shutdown is protection working. Resetting without reading the code discards the only diagnosis you had.

Load bank testing exposes cooling and fuel problems that a monthly no load exercise never will.

Brands

Caterpillar

Cummins

Kohler

Generac

Verify against the manual

True

generators

Automatic transfer switch fails to transfer#

The generator may be running perfectly while the load stays dark, because the switch between them did not operate.

Method

Confirm the generator is actually up to voltage and frequency, because the switch will not transfer to a source that is not ready

Check the ATS controller for a sensing fault or a locked out condition

Check control power to the switch, which is separate from the power it switches

Check the time delay settings, since a long delay looks identical to a failure to someone watching

Verify the position indication against the physical position rather than trusting the display

Check the operator mechanism and linkage for binding, which is common on switches that rarely operate

Exercise the switch on a schedule, because most ATS failures are found only when they are needed

Warning signs

Generator running with load still on utility or dead

No transfer on a test

Transfer in one direction only

Position indication disagreeing with reality

Notes

A transfer switch that never operates between tests is the most likely component to fail when it matters.

Bypass isolation switches let the ATS be serviced without dropping load; know whether the site has one.

Brands

ASCO

Russelectric

Generac

Eaton

Schneider

Verify against the manual

True

distribution

Breaker trip on a critical circuit#

Something drew more than the breaker allows or a fault occurred. The breaker did its job and the question is what it protected against.

Method

Do not reset into an unknown fault; establish whether it was overload or a fault first

Check the load on the circuit against its rating and against what it was before, since a rack build out is the usual cause

Check whether it trips immediately on reset or holds, because instant retrip means a fault rather than overload

Inspect the circuit for damage, heat or a compromised connection before energising

Check the breaker itself for heat damage, since a breaker that has interrupted a fault may not be reusable

Check torque on terminations, because a loose connection heats and nuisance trips

Confirm the circuit is not shared with anything that was added without a load calculation

Warning signs

Load dropped on one circuit or rack

Breaker in the tripped position

Burning smell or discoloration at the panel

Repeated trips at the same load point

Notes

Reset once only. A breaker that trips twice is telling you something and the second reset risks equipment and people.

Trip curves matter: an instantaneous trip and a long time overload look the same on the handle.

Brands

Schneider Square D

Eaton

ABB

Vertiv

Starline busway

Verify against the manual

True

distribution

Neutral heating and harmonic loading#

Nonlinear IT loads produce triplen harmonics that add rather than cancel in the neutral, so the neutral can carry more than any phase.

Method

Measure neutral current and compare it against the phase currents rather than assuming it is the difference

Check transformer temperature at its actual load, since a standard transformer derates badly on harmonic load

Check whether the transformer is K rated for the load it is actually feeding

Take a power quality reading rather than a clamp reading, because harmonic content does not show on a basic meter

Check for a shared neutral serving multiple phases of nonlinear load

Check conductor sizing against the neutral current found, not against phase current

Warning signs

Neutral conductor hotter than the phases

Transformer running hot at moderate load

Higher than expected THD on a power quality reading

Nuisance tripping with no obvious overload

Notes

Third harmonic and its multiples add arithmetically in the shared neutral of a three phase four wire system.

A neutral sized to phase current can be undersized in a data hall full of switching power supplies.

Brands

Schneider

Eaton

Vertiv

ABB

Verify against the manual

True

airCooling

CRAC or CRAH short cycling and fighting#

Units switching on and off rapidly, or working against each other with one cooling while another humidifies.

Method

Compare setpoints and deadbands across every unit in the room, because units fighting is nearly always a setpoint disagreement

Check whether the units are on a shared control or running independently, since independent units in one room will fight

Check return air sensor placement and calibration, since a sensor in the wrong place drives the wrong behaviour

Check for a failed unit forcing the others to overwork

Widen the deadband before adjusting setpoints, because a narrow band causes cycling on its own

Check refrigerant charge and airflow if one unit alone is cycling

Verify the humidity setpoints are identical across the room, since humidity fighting wastes more energy than cooling fighting

Warning signs

Compressors starting and stopping within minutes

One unit heating while another cools

One unit humidifying while another dehumidifies

Rising energy use with no load change

Temperature swinging rather than holding

Notes

Units fighting each other is one of the most common and most expensive faults in a legacy data hall.

Short cycling destroys compressors faster than continuous running does.

Brands

Vertiv Liebert

Schneider Uniflair

Stulz

Data Aire

Verify against the manual

True

airCooling

Hot spots and airflow bypass#

Cooling capacity is adequate but the air is not reaching the equipment, which is an airflow problem rather than a cooling problem.

Method

Measure at the rack inlet rather than in the aisle, because the room can be cold while a server inlet is hot

Check for missing blanking panels, which let hot air recirculate through the rack front to back

Check floor tile placement, since perforated tiles in the hot aisle actively make things worse

Check for unsealed cable cutouts under racks, which is where most bypass air escapes

Check containment doors and panels are closed and intact

Check underfloor obstruction from accumulated cabling before adding capacity

Only consider more cooling once airflow is proven, because more cold air into a bypass path changes nothing

Warning signs

Hot spots at the top of racks while the room is cool

Inlet temperature varying widely between racks

High return air temperature with low supply temperature

Adding cooling not fixing the hot spot

Notes

Most hot spots are airflow management rather than insufficient cooling capacity.

Blanking panels and sealed cutouts are the cheapest thermal fix available in any hall.

Brands

Vertiv

Schneider

Stulz

Upsite

Subzero

Verify against the manual

True

liquidCooling

Coolant leak detected#

Liquid near energised IT and power. The alarm must be treated as real until objectively established otherwise.

Method

Treat the alarm as real until objectively established otherwise, and never clear it on the absence of visible liquid alone

Notify operations before intervening, because liquid near energised equipment affects more than the cooling system

Inspect the triggered detection point and the fittings, manifolds and quick disconnects adjacent to it

Check reservoir level against its last recorded value rather than against the sight glass alone

Check whether the system is holding pressure or vacuum, since a slow loss confirms a leak that inspection may miss

Trace along the path rather than only at the sensor, because coolant runs before it pools

Do not restore cooling to a suspected leaking loop without the vendor procedure

Warning signs

Leak detection alarm at a zone

Coolant level falling in the reservoir

Damp or residue at a fitting or quick disconnect

Vacuum or pressure not holding

Notes

This practice is not yet standardised across the industry; the vendor manual and the project designers are the authority.

Take maximum and minimum coolant loop pressures from the facility RMP and PSM documentation, not from memory.

Brands

Vertiv

Schneider

Motivair

CoolIT

Boyd

Verify against the manual

True

liquidCooling

CDU pressure or flow out of range#

The coolant distribution unit cannot hold its pressure or flow band, which derates or drops the racks it feeds.

Method

Take the pressure limits from the facility RMP and PSM documentation and the CDU manual rather than from a remembered figure

Check the filter or strainer first, since a loading filter is the most common cause of falling flow

Check pump operation and whether it is running at maximum to hold setpoint, which indicates restriction

Check for air in the loop, which produces unstable pressure and noisy pumps

Check quick disconnects at the rack manifolds for partial engagement

Check the secondary loop for a closed or partly closed isolation valve after any maintenance

Prioritise seal inspection on higher pressure loops, since mechanical seal life scales with discharge pressure

Warning signs

Pressure differential outside the band

Flow rate below setpoint

Pump running at maximum with no result

Rack level thermal alarms downstream

Notes

Mechanical seal intervals come from the vendor manual and trended condition, not from a fixed number.

A seal that lasted years on a low pressure chilled water loop can fail far sooner on a high pressure cold plate loop.

Brands

Vertiv

Schneider

Motivair

CoolIT

Boyd

Verify against the manual

True

monitoring

Alarm storm and nuisance alarms#

So many alarms that the real one is invisible. The failure is the monitoring system rather than the plant.

Method

Find the first alarm in the sequence, because the rest are usually consequences of it

Check whether a single failed sensor or comm path is generating the cascade

Check thresholds against real operating ranges, since a threshold set inside normal variation alarms forever

Check for a device alarming and clearing on a cycle, which points at a threshold sitting exactly on the operating point

Widen deadbands and add delays on points that chatter rather than disabling them

Never disable an alarm to quiet a storm without recording it, because a disabled point protects nothing

Review the alarm list quarterly, since alarm quality degrades quietly as a site changes

Warning signs

Hundreds of alarms in a short window

The same point alarming and clearing repeatedly

Staff routinely acknowledging without reading

Real events missed inside the noise

Notes

Alarm fatigue is a real failure mode; a system nobody trusts is worse than no system.

A point that alarms daily and is always ignored is not monitoring, it is noise.

Brands

Vertiv

Schneider EcoStruxure

Nlyte

Sunbird

Modius

Verify against the manual

True

More in the RackWatt reference Calculation Reference Procedures Library Safety Topics
Field references for other trades Fuel and tanks Fire alarm Solar PV EV charging Battery storage
These notes come from RackWatt

Reference notes from RackWatt, the offline field toolkit for data center critical facilities technicians. It works with no cell signal, because the sites do not have any.

Get RackWatt on the App Store

These notes are a field aid, not a substitute for the governing codes, the stamped drawings, the authority having jurisdiction, or manufacturer manuals. Verify against the current documentation for your installed equipment.