The Unit Is Running. But How Much Standby Redundancy Have You Lost?

A thermal power plant can remain at full load while losing critical standby redundancy. Learn why unavailable standby equipment changes maintenance priority and plant risk.

MaintBoard Team
Thermal power plant equipment illustrating the operational risk of losing critical standby redundancy

The Unit Is Running. But How Much Standby Redundancy Have You Lost?

The unit is on load.

Generation is stable.

There is no trip.

No major alarm is active.

From the control room, everything may appear normal.

But somewhere in the plant, one boiler feed pump is under maintenance.

A standby ID fan is unavailable.

One coal mill has been isolated since yesterday.

A cooling water pump has a pending defect.

The plant is still running.

But the margin protecting it from the next failure is disappearing.

That is a very different condition from a fully healthy plant.

And it is one that maintenance teams need to see clearly.

Running Does Not Always Mean Available

Consider a simple situation.

A thermal power plant is operating normally with a boiler feed pump in service and another pump available as standby.

Then the standby BFP develops a problem.

Maintenance isolates it and starts work.

The operating BFP continues to perform normally.

Feedwater flow is stable.

Steam conditions are stable.

Generation continues.

Technically, there is no production loss.

If we look only at current generation, the plant appears healthy.

But something important has changed.

The plant has lost redundancy.

Yesterday, if the running BFP developed high vibration, seal leakage, a motor problem or another serious defect, operations could transfer to the standby pump.

Today, that option may no longer exist.

The same running machine now carries a very different operational consequence.

Then the Second Problem Appears

A few hours later, the bearing vibration on the running BFP starts increasing.

It is not yet at trip level.

The pump is still delivering the required flow.

From a simple equipment-status screen, it may still show:

Running

But nobody responsible for the unit would treat this as a normal situation.

Operations now wants to know:

  • How quickly is vibration increasing?
  • What was the previous reading?
  • Has this happened before?
  • What was done last time?
  • Is the standby BFP close to restoration?
  • Can the unit continue at the current load?
  • Should load be reduced?
  • What happens if the running pump deteriorates further?

Nothing has failed yet.

But the risk profile of the unit has changed substantially.

This is why equipment availability is more than simply asking whether an asset is running or stopped.

The Hidden Risk Is the Loss of Margin

Thermal plants are designed with redundancy in many critical auxiliary systems because equipment will eventually require maintenance.

Depending on the plant configuration, this may apply to systems such as:

  • Boiler feed pumps
  • Condensate extraction pumps
  • Cooling water pumps
  • ID fans
  • FD fans
  • PA fans
  • Coal mills
  • Air compressors
  • Lubrication systems
  • Vacuum pumps
  • Electrical auxiliaries

The exact redundancy philosophy differs from plant to plant.

But the principle is the same.

Standby equipment gives operations margin.

When that standby disappears, the unit may continue producing the same MW, but its ability to tolerate another equipment problem has reduced.

That distinction is important:

Production may be unchanged while operational risk has increased.

"No Breakdown" Can Be Misleading

Maintenance dashboards frequently focus on visible events.

How many breakdowns occurred?

How many work orders are open?

How much downtime did we have?

What is PM compliance?

These are useful measures.

But imagine two operating conditions.

Condition A

  • BFP-A running normally
  • BFP-B healthy and available
  • No pending critical defects

Condition B

  • BFP-A running normally
  • BFP-B unavailable
  • BFP-A vibration increasing
  • Corrective work on BFP-B delayed

Both conditions may currently show:

0 minutes production downtime

But they clearly do not represent the same plant condition.

Waiting until the second pump fails before recognizing the problem means maintenance is measuring the consequence after the protection has already disappeared.

This is one reason unplanned downtime should be reviewed together with equipment condition and standby availability rather than in isolation.

Coal Mills Make This Even More Visible

Consider a coal-fired unit operating with several mills.

One mill is unavailable for maintenance.

The unit can still carry the required load using the remaining mills.

There is no immediate generation loss.

Then another operating mill begins tripping intermittently.

The first trip is reset.

The mill returns to service.

The second trip happens several hours later.

Again, operations restores it.

The unit may technically continue running throughout this period.

But the maintenance situation is becoming progressively worse.

The real question is no longer:

Is the unit running?

It is:

How much equipment can we afford to lose before we have to reduce load?

That is the operating margin maintenance teams need to understand.

And when one machine repeatedly creates the same problem, it should not remain a sequence of unrelated breakdown jobs. It becomes a repeat-failure problem that deserves deeper investigation through techniques such as Pareto analysis and root cause analysis.

Standby Equipment Is Easy to Forget

There is another practical problem with standby equipment.

Running equipment gets attention because operators see it every shift.

Temperatures change.

Vibration changes.

Current changes.

Pressures and flows change.

Standby equipment may sit quietly for days or weeks.

Nothing looks wrong because nothing is moving.

Then the day comes when operations needs it.

The pump does not start.

The actuator does not respond correctly.

The lubrication system has a problem.

The motor trips.

A valve required for changeover is stuck.

The plant discovers the standby problem at exactly the wrong moment.

This is why standby equipment cannot be treated as an asset that matters only when it is required.

It needs appropriate preventive maintenance, testing and inspection while it is still on standby.

Availability and Readiness Are Not the Same Thing

A useful distinction for critical auxiliaries is:

Running — the equipment is currently operating.

Available — the equipment can be brought into service when required.

Unavailable — the equipment cannot currently perform its required duty.

This matters especially for standby assets.

A pump can be physically present in the plant and still not be available.

A fan may have no active breakdown but may still be unavailable because corrective work is pending.

A coal mill may have been repaired but may not yet have completed trial operation.

A standby generator may exist but fail its periodic test.

Therefore, simply maintaining an asset register through asset management software is not enough.

The plant needs to understand the actual readiness of critical equipment.

Not Every Standby Loss Has the Same Risk

An unavailable standby pump should not automatically receive the same priority in every situation.

Context matters.

Suppose a plant has three identical pumps.

Two are available and one is under planned maintenance.

That may be an acceptable operating condition.

Now consider another system where one machine is running and the only standby is unavailable.

The consequence is very different.

The maintenance priority should consider:

  • Current equipment configuration
  • Required redundancy
  • Current unit load
  • Condition of running equipment
  • Duration of standby unavailability
  • Failure history
  • Availability of spares
  • Expected repair time
  • Safety consequence
  • Production consequence

This is where risk-based maintenance becomes practical.

The question is not simply whether a work order exists.

It is what happens to the plant if the next failure occurs before that work order is completed.

A Small Pending Job Can Suddenly Become the Most Important Job in the Plant

Imagine the standby BFP was taken out because of a seal problem.

The repair appears straightforward.

A work order is raised.

The job is placed in the maintenance backlog because the running BFP is healthy.

Then another urgent breakdown appears elsewhere.

Resources move.

The standby pump repair slips by one shift.

Then another.

Three days later, vibration begins increasing on the operating pump.

The same pending seal repair now has a completely different priority.

The job itself has not changed.

The plant context around the job has changed.

This is an important weakness in static maintenance prioritization.

A work order created as Medium priority on Monday may need to become Critical on Wednesday because another layer of redundancy has disappeared.

A useful maintenance backlog should therefore not be treated simply as a queue of old jobs.

The plant needs to continuously ask which pending work is increasing operational exposure.

Maintenance Priority Should Change With Plant Condition

This is where operations and maintenance must remain closely connected.

Operations sees what is happening to the unit.

Maintenance sees what is happening to the equipment.

Neither picture is complete by itself.

Consider:

BFP-B

Status: Unavailable Reason: Seal replacement Repair progress: 70% Expected return: 16:00

That is useful maintenance information.

Now add:

BFP-A

Status: Running Vibration: Increasing Current load: 96 MW Standby available: No

The urgency becomes immediately obvious.

A good work order management system should allow maintenance activity to be understood in this operating context rather than reducing everything to Open, In Progress and Completed.

Condition Readings Become More Important When Standby Is Lost

Normally, a small variation in vibration or bearing temperature may warrant observation.

When the standby is unavailable, the exact same deviation can require much closer attention.

This is where equipment readings become valuable.

For example:

Parameter Previous Current Condition
DE Bearing Vibration 3.1 mm/s 4.6 mm/s Increasing
NDE Bearing Vibration 2.8 mm/s 3.0 mm/s Stable
Bearing Temperature 68°C 74°C Increasing
Motor Current 182 A 184 A Stable
Discharge Pressure 132 bar 131 bar Stable

The objective is not simply to collect these values.

The important question is whether the trend requires maintenance action.

Digital meter readings become far more useful when the plant can associate the reading with the equipment, identify a deviation and follow that deviation through maintenance action.

For rotating equipment, this may include vibration analysis alongside temperature, lubrication and operating parameters.

The Work Does Not End When the Standby Is Restored

Suppose maintenance completes the BFP repair.

The pump is put back into service.

Does the work order simply become Completed?

Operationally, there is another important step.

Prove that the standby is actually available.

Depending on the equipment and plant procedure, that may mean checking:

  • Trial run completed
  • Vibration normal
  • Bearing temperatures normal
  • Pressure and flow normal
  • No leakage
  • Motor current normal
  • Protection healthy
  • Required valves correctly lined up
  • Operations acceptance completed

Only then has redundancy genuinely been restored.

Maintenance completion and operational readiness are closely related, but they are not always the same event.

Some Redundancy Problems Need the Next Shutdown

Not every standby problem can be permanently resolved online.

Sometimes a temporary repair keeps equipment available.

Sometimes an internal inspection is required.

Sometimes alignment, overhaul or major component replacement has to wait until the unit is down.

These known risks should not disappear into old work-order history.

They should feed directly into shutdown maintenance planning.

When the next outage arrives, the team should already know:

  • Which standby systems have caused concern
  • Which temporary repairs remain
  • Which equipment has repeatedly lost availability
  • Which defects have been deferred
  • Which failure modes need permanent correction

Otherwise, the plant may restart after a shutdown carrying exactly the same reliability exposure.

A Better Daily Maintenance Question

Many daily meetings naturally start with:

What equipment is under breakdown?

That should remain.

But another question can reveal risks that breakdown count will miss:

Which critical standby equipment is unavailable today?

Follow that with:

What happens if the running equipment fails before we restore it?

That immediately changes the maintenance conversation.

A useful daily view might show:

Equipment Running Standby Current Risk
Boiler Feed Pumps BFP-A BFP-B unavailable No standby
ID Fans IDF-A IDF-B available Normal
Coal Mills A/B/C/D running Mill-E unavailable Reduced mill margin
Cooling Water Pumps CWP-1 running CWP-2 available Normal

This is simple information.

But operationally, it is extremely valuable.

A CMMS Should Make Lost Redundancy Visible

A CMMS for a power plant should not only answer:

  • What broke?
  • What PM is due?
  • What work is pending?

It should also help maintenance teams answer:

  • Which critical equipment is unavailable?
  • Which standby machines are unavailable?
  • How long has redundancy been lost?
  • What work is preventing restoration?
  • Is the running equipment showing abnormal condition?
  • Has this combination occurred before?
  • What is the consequence if another machine fails?
  • When is full redundancy expected to be restored?

This is where maintenance analytics and reporting becomes much more valuable than simply showing the number of completed work orders.

The goal is not another dashboard.

The goal is to show the maintenance team where the plant has become exposed.

The Plant Can Be Running and Still Be One Failure Away From Derating

This is the important distinction.

A plant at full load is not automatically a plant at full reliability.

You can have:

100 MW generation

and simultaneously:

Zero standby margin on a critical system.

Both statements can be true.

That is why maintenance performance cannot be judged only by whether production stopped.

Sometimes the most valuable maintenance work is the work that restores a standby machine before anybody notices the consequence of losing it.

The Question That Matters

When a critical standby machine becomes unavailable, do not only ask:

When will this work order be completed?

Ask:

What protection have we lost while this equipment is unavailable?

And then:

What happens if the running machine develops a problem before we restore that protection?

That question connects maintenance directly to plant risk.

The unit may still be running.

The MW may still look normal.

The control room may still be quiet.

But if redundancy is disappearing, the maintenance situation has already changed.

The best time to restore that margin is before the next failure proves why it was needed.

Frequently asked questions

What does standby redundancy mean in a thermal power plant?

Standby redundancy means having additional equipment available to take over when the running equipment fails or is taken out of service. Examples include standby boiler feed pumps, ID fans, cooling water pumps, coal mills and compressors. The standby equipment may not be running, but it must be ready to operate when required.

Can a thermal power plant be running normally while still carrying high maintenance risk?

Yes. A unit may continue generating at full load even when critical standby equipment is unavailable. Production may appear normal, but the plant has less protection against the next equipment failure. This loss of redundancy increases operational risk even before any downtime occurs.

Why is an unavailable standby boiler feed pump a serious concern?

If one boiler feed pump is running and the standby BFP is unavailable, the plant loses an important layer of protection. If the running pump then develops high vibration, bearing temperature, seal leakage or another serious problem, operations may have no immediate replacement available and may need to reduce load or shut down the unit.

What is the difference between running, available and unavailable equipment?

Running means the equipment is currently operating. Available means the equipment is ready to be placed into service when required. Unavailable means it cannot currently perform its required duty because of maintenance, a defect, isolation or another condition. This distinction is particularly important for standby equipment.

Should standby equipment receive preventive maintenance even when it is rarely used?

Yes. Standby equipment can develop problems while sitting idle, including lubrication issues, stuck valves, electrical faults, actuator problems and deterioration of seals or other components. Periodic preventive maintenance, inspection and functional testing help confirm that the equipment will actually operate when required.

How should maintenance priority change when standby equipment becomes unavailable?

Priority should consider the plant condition, not only the original work order classification. If the only standby machine becomes unavailable, or the remaining running equipment begins showing abnormal condition, the related maintenance work may require immediate escalation because the consequence of another failure has increased.

Why should equipment condition readings be monitored more closely when redundancy is lost?

When no standby equipment is available, changes in vibration, bearing temperature, motor current, pressure, flow or other condition readings on the running machine become more important. A deviation that would normally be monitored may require faster maintenance action because there is no alternative equipment available if the condition worsens.

How can a CMMS help manage standby equipment risk in a thermal power plant?

A CMMS can help teams track which critical assets are running, available or unavailable, connect pending work orders to standby equipment, monitor condition readings, review previous failures and show how long redundancy has been lost. This gives maintenance and operations a clearer view of where the plant is exposed before a second failure causes derating or downtime.

Make Critical Equipment Risk Visible

MaintBoard helps power plant teams track equipment availability, condition, pending maintenance and operational risk in one system.