FMEA to Prevent Recurring Defects: A Practical Guide

By Erik ·

Engineers collaborating on a whiteboard, sketching FMEA diagrams and process flowcharts in a manufacturing setting

FMEA prevents recurring defects when it is used as a living risk tool that the team updates after every real failure, not as a one time document filed at project launch. Most recurring defects come back because the FMEA was never updated when the failure happened, the controls were never strengthened, or the highest risk failure modes were never tied to a specific action owner with a deadline. Treating FMEA as a closed loop process changes that pattern.

This guide is written for Quality Engineers, Process Engineers, Manufacturing Engineers, and CI Specialists who own FMEA in a manufacturing plant. The focus is practical: how to use FMEA so that defects that have already happened once do not come back a third or fourth time.

Why Defects Recur Even When an FMEA Exists

Most plants have FMEAs. Most plants also have recurring defects. The two coexist for predictable reasons.

The first reason is that the FMEA was completed at PPAP or process launch and never updated when the process changed. New tooling, a new supplier, a new operator pattern, or a new material lot all change the risk picture. If the FMEA does not change with them, it stops describing the actual process.

The second reason is that the controls listed in the FMEA were aspirational rather than real. A control listed as "operator visual inspection" may not actually be in the work instruction. A control listed as "in process gauge check every hour" may be skipped on the back shift. When real controls do not match documented controls, the FMEA underestimates risk.

The third reason is that the FMEA was treated as a document rather than a decision tool. The team filled it out, calculated RPNs, filed it, and moved on. The actions that should have followed the highest risk rows were never assigned, or were assigned without a date or owner. Recurring defects are usually traceable back to a row in the FMEA that was understood at the time but never acted on.

Consider an automotive supplier with a recurring weld pull strength failure on a sub assembly. The FMEA correctly identified weld energy variation as a high severity, moderate occurrence failure mode. No action was assigned, the existing control was an end of line destructive sample, and the same defect appeared in three of the next four months. The FMEA was right; the closed loop was missing.

Use Real Failure Data to Update the FMEA

The fastest way to make an FMEA prevent recurring defects is to update it every time a real failure occurs. Each confirmed defect should trigger three questions: Was this failure mode in the FMEA? If yes, were the severity, occurrence, and detection ratings accurate given what just happened? If no, why was it missed?

When the failure mode was in the FMEA but the occurrence rating was too low, raise it and reassess the action plan. When the detection rating was too high, meaning the team assumed the control would catch it but it did not, lower the detection score and strengthen the control.

When the failure mode was not in the FMEA at all, add it. A new row in the FMEA after a real failure is one of the most valuable updates the document will ever get. Teams that approach failure investigation with a structured habit of separating what they observed from what they assumed often find the practitioner reference on moving from symptoms to root cause for Six Sigma Black Belts useful when deciding which mode to add and how to describe it.

Stop Chasing RPN, Start Acting on the Top Risks

RPN, the product of severity, occurrence, and detection, is a useful sorting tool but a weak action trigger on its own. A failure mode with high severity should drive action even when occurrence and detection are moderate, because the consequence of being wrong is large.

The Action Priority approach in the AIAG VDA FMEA handbook reflects this and is generally a stronger basis for action than raw RPN ranking. Whichever ranking method the plant uses, the rule that prevents recurring defects is the same: every high severity failure mode needs a defined action, an owner, and a date, regardless of where its RPN sits in the sort.

Recurring defects often live in the rows that were ranked just below the action threshold. If the threshold was an RPN of 100 and the failure mode sat at 96, it usually got nothing. Reviewing the next ten rows below the threshold, especially any with high severity, is one of the simplest and highest payoff FMEA habits a plant can build.

Strengthen Controls Where Detection Is Weak

When a defect recurs, the detection control was almost always too weak. The two most common detection weaknesses are end of line inspection and operator visual checks.

End of line inspection catches defects after they are made. It does not prevent them, and it usually only catches a fraction of what it is supposed to catch. Relying on end of line inspection as a primary detection control is one of the most common reasons defects come back.

Operator visual inspection is highly variable across shifts, fatigue levels, and experience. It can be a valuable backup, but it is rarely strong enough as the only control on a high severity failure mode.

Stronger detection controls usually look like in process measurement, automated gauging, error proofing, mistake proofing, or process parameter monitoring. When the FMEA shows a high severity row with only a visual or end of line control, that row is a strong candidate for a detection upgrade. The companion practitioner article on data stratification in Minitab vs QI Macros is useful when deciding how to slice in process data so that weaker detection controls can be replaced with stronger statistical monitoring.

Tie FMEA Actions to a Real Owner and a Real Date

A row in an FMEA without an owner and a date is a row that will not get done. Recurring defects almost always trace back to actions that were identified and then orphaned.

A practical pattern is to require that every action in the FMEA name a single person, not a department, and a date that is realistic for the work involved. When the action involves engineering, the engineering owner is named. When the action involves a supplier, the supplier facing engineer is named. When the action involves operations, the operations owner is named. Joint ownership is usually a sign that no one will own it.

A monthly FMEA action review, even a short one, prevents most action drift. The review should cover open actions, overdue actions, and any new failure modes added since the last review. Plants that bring this discipline to FMEA usually see the same defects stop coming back within one or two action cycles.

Connect FMEA to Control Plan and Work Instructions

An FMEA that is not connected to the control plan and the work instructions is a document, not a control system. Recurring defects often live in the gap between the three.

When a control is added or strengthened in the FMEA, the same change should appear in the control plan and the work instruction within the same change cycle. If the FMEA says in process gauging is performed every thirty minutes, the control plan should reference that frequency, and the work instruction should describe how the operator performs the check, what to do if the result is out of tolerance, and where to record it.

When the three documents are aligned, an auditor or a new operator can trace any failure mode from the FMEA through the control plan to the work instruction in seconds. When they are not aligned, controls drift across shifts and defects find the gaps.

Practical Action Block

Use this short checklist after every confirmed recurring defect.

  • Find the failure mode in the FMEA, or add it if it is missing.
  • Reassess severity, occurrence, and detection based on what actually happened.
  • Confirm the listed control matches the real floor practice on every shift.
  • Assign a single owner and a realistic date for the corrective action.
  • Update the control plan and work instruction in the same change cycle.
  • Schedule a monthly FMEA action review and track open and overdue actions.

When FMEA Is Not the Right Tool

FMEA is not the right tool for every defect investigation. Three situations are worth recognizing.

A single nonconforming part with no pattern is usually a containment and quick root cause situation, not an FMEA update. Forcing every isolated event into the FMEA dilutes the document.

A defect caused by a known and already controlled failure mode where the control was simply not followed is an execution problem, not an FMEA problem. The fix is to reinforce the control, retrain the operator, or strengthen the audit, not to add another row to the FMEA.

A defect caused by a process change that has not been formally introduced through change control is a change management problem. Updating the FMEA without first restoring change control discipline will not stop the next unplanned change from creating the next recurring defect. Teams that struggle with this pattern often benefit from the leadership oriented framing in building statistical judgment in production supervisors, which discusses how shift level decisions create or prevent recurring quality issues.

Key Takeaways

  • Recurring defects almost always trace back to FMEA rows that were never updated, never acted on, or never connected to real floor controls.
  • Update the FMEA after every confirmed failure, including severity, occurrence, detection, and any missing failure modes.
  • Treat high severity rows as action candidates even when their RPN sits below the usual threshold.
  • Replace weak detection controls such as end of line inspection or visual checks with in process or automated controls when severity is high.
  • Assign every action to a single owner with a realistic date and review open actions monthly.
  • Keep the FMEA, control plan, and work instructions aligned so that controls do not drift across shifts.

Frequently asked questions

Why do defects keep recurring even when an FMEA exists?

Most recurring defects trace back to FMEAs that were completed at launch and never updated when the process changed, controls listed in the FMEA that did not match real floor practice, or actions that were identified but never assigned to a real owner with a date. Closing these three gaps stops most repeat defects within one or two action cycles.

Should we use RPN or Action Priority to drive FMEA actions?

Either ranking method works, but the rule that prevents recurring defects is the same: every high severity failure mode needs a defined action, an owner, and a date, regardless of where its RPN sits. The AIAG VDA Action Priority approach is generally a stronger basis than raw RPN because it weights severity more heavily.

How often should an FMEA be updated?

Update the FMEA after every confirmed failure, after any meaningful process or supplier change, and on a regular review cycle, typically monthly or quarterly depending on plant size. Each real failure should trigger a check of whether the failure mode was in the FMEA and whether the severity, occurrence, and detection ratings were accurate given what actually happened.

What makes a detection control strong enough to rely on?

Strong detection controls usually involve in process measurement, automated gauging, error proofing, mistake proofing, or process parameter monitoring. End of line inspection and operator visual checks are common but weak as primary controls on high severity failure modes, because they depend on catching defects after they are made or on highly variable human attention.