Field reference for grid scale battery storage technicians. 11 entries, taken from the StoreWatt app.
Reference notes from StoreWatt, the offline field toolkit for grid scale battery energy storage technicians. It works with no cell signal, because the sites do not have any.
Get StoreWatt on the App StoreThe battery management system monitors cell and module voltage, current, temperature, SOC, SOH, cell balance, insulation resistance, and contactor state. The usual architecture is layered: a slave BMS on each module reads cells, a master BMS per rack aggregates modules and drives the rack contactors, and a system BMS coordinates racks and talks to the PCS and EMS. When chasing any BMS fault, find which layer raised it first. A cell level alarm points at one module; a rack level alarm can be a rack controller or its feed from the slaves; a system level alarm is often communication or coordination, not chemistry.
Module and rack BMS layers almost always talk over CAN. Reading the CAN traffic localizes a fault fast: a missing message identifier tells you which module went quiet, and a garbled bus points at wiring, termination, or a failing transceiver. Basics that solve most CAN problems: 120 ohm termination at both physical ends only, matched bit rate on every node, twisted pair with the shield grounded at one end, and no star stubs. A healthy bus measures about 60 ohms across CAN high to CAN low with power off.
Message identifier maps are proprietary. Pull the CAN database file or the manual for the installed BMS version before decoding.
One or more cells exceeded the charge voltage limit. The BMS narrows the charge window, derates, or opens contactors depending on how far past the limit the cell went. Common causes: charging into a high SOC with a drifting cell, a weak cell that races ahead of its neighbors near full, balancing that cannot keep up, or a sensing harness fault reading high. Check whether it is one cell or many. A single repeating cell points to that cell or its sense lead; a whole module near the top of charge points to SOC calibration or an aggressive charge profile.
Let the BMS protect. Do not raise limits or reset repeatedly to force a charge. If the same cell trips repeatedly, log the cell position and escalate to the manufacturer.
A cell fell below the discharge floor. The BMS blocks further discharge and may open contactors. Causes mirror over voltage: a weak cell that sags first under load, deep discharge after a long idle period with parasitic loads, or a bad sense connection reading low. A rack that sat de powered for weeks can drift under the floor on self discharge alone. Recovery charging below the normal window is a manufacturer controlled procedure, not a field improvisation.
Stop discharging. Follow the manufacturer recovery procedure for deeply discharged cells; charging a cell that spent time below its absolute floor can be hazardous.
Cell temperature crossed an alarm or trip threshold. The system derates first, then trips. The usual cause is thermal management, not chemistry: failed pumps or fans, low coolant, blocked filters, a wrong setpoint, or high ambient plus high C rate. Look at the spread. Every cell warm together points at cooling or duty; one cell hot alone is a red flag for a cell problem and deserves escalation, not a reset.
If temperature keeps climbing after the system derates or trips, or one cell runs away from its neighbors, treat it as a possible thermal event: do not open the enclosure, follow the Emergency Response Plan, and notify the qualified person.
The spread between the highest and lowest cell exceeds the alarm limit. Imbalance caps usable capacity: charge stops on the highest cell and discharge stops on the lowest. Causes: balancing hardware not keeping up, a genuinely weak cell, uneven temperatures across a rack, or long idle periods. Trend it. Imbalance that grows steadily on the same cells is degradation; imbalance that appears after commissioning or a long outage often just needs a supervised balancing cycle at high SOC.
Run the manufacturer balancing procedure and re trend. Escalate a spread that keeps widening on the same cells.
The insulation monitor detected low resistance between the DC bus and ground. Somewhere, high voltage DC has a path toward chassis: damaged cable insulation, coolant intrusion into a module, moisture in a connector, or a failed component. Until located, any grounded surface can be at pack potential. This is the classic wet weekend fault: moisture drops insulation readings plantwide, then readings recover as things dry. A single rack that stays low while others recover is a real fault in that rack.
Stop work. Treat every conductive surface as live. Do not bypass or reset the insulation monitor to restore operation. Isolate the affected rack under the qualified person's direction and locate the fault with the manufacturer procedure before returning to service.
The BMS commanded a contactor open but feedback or bus voltage says it is still closed. Welded contacts happen after closing into a fault, repeated closing under load, or end of mechanical life. The rack DC bus remains energized even though the system believes it isolated. Never trust the HMI state for isolation. The only proof is a zero energy verification at the conductors with a CAT III or CAT IV DC rated tester.
Treat the bus as energized regardless of indicated state. Verify absence of voltage at the point of work with a DC rated tester before touching anything. Replace, never file or reuse, a welded contactor.
Contactor feedback logic differs by vendor. Confirm which signal the fault uses in the installed BMS manual.
A rack will not connect to the bus. Before the contactor itself, check the interlocks the BMS requires: precharge must complete, pack and bus voltage must match within a window, insulation must be healthy, and no cell alarm can be active. A precharge resistor that opened from repeated cycling is a frequent culprit and often reports as a precharge timeout. Coil supply, drive circuit, and auxiliary feedback wiring come next. A rack that closes when bus voltage is matched manually points at precharge; one that never picks the coil points at supply or drive.
Work the interlock list before condemning hardware. Precharge circuits store and dissipate real energy; follow the isolation procedure before touching them.
A BMS layer lost contact with another: slaves dropping off a module bus, a rack master silent to the system BMS, or the system BMS unreachable from the PCS or EMS. Most systems fail safe by derating or opening contactors when supervision is lost, so a comm fault often presents as a capacity or availability problem. Localize by layer: one module missing is that module's node or its stub; a whole rack silent is the rack master, its power, or the backbone segment; everything missing at once is the head end, a power supply, or a broken backbone near the head end.
Restore communication before chasing phantom electrical faults; stale data makes healthy racks look sick.
Reported SOC wanders from reality because SOC is an estimate built on current integration plus voltage models. Flat voltage chemistries like LFP drift more between full charge references. Symptoms: the system hits the voltage limit before reported SOC says it should, or dispatch shortfalls near the window edges. The fix is usually a calibration cycle to a full charge reference per the manufacturer, and checking current sensor zero offsets. Persistent drift on one rack against its peers is a sensor problem.
Reference notes from StoreWatt, the offline field toolkit for grid scale battery energy storage technicians. It works with no cell signal, because the sites do not have any.
Get StoreWatt on the App StoreThese notes are a field aid, not a substitute for the governing codes, the stamped drawings, the authority having jurisdiction, or manufacturer manuals. Verify against the current documentation for your installed equipment.