What Is an Acceptable %GRR? A Practical Guide for Manufacturing

By Hélène ·

An engineer examining a measurement gage on a factory floor, representing the practical application of a Gage R&R study.

Why "Acceptable" Depends on Your Goal

Every quality engineer has been asked the question: "Is this Gage R&R result good or bad?" The answer, frustratingly, is almost always "it depends." While there are widely accepted numerical thresholds, applying them without context is a common and costly mistake. The "acceptability" of your measurement system is not a single number but a judgment call based on your specific goal.

At its core, a Gage Repeatability and Reproducibility (GR&R) study quantifies the variation, or "noise," introduced by your measurement system. This noise obscures the true variation of your manufactured parts. To judge its impact, you must compare this measurement variation to something meaningful. The two most important benchmarks are:

  1. Process Variation: How much of the total observed variation is consumed by the measurement system? This is critical for process control and capability analysis.
  2. Specification Tolerance: Can the measurement system adequately distinguish a good part from a bad part? This is critical for sorting and disposition decisions.

Consider a team producing precision-machined pistons. An engineer responsible for process control on the CNC lathe needs to know if the measurement system is stable enough to detect small shifts in the process mean. They are primarily concerned with how measurement error compares to part-to-part variation. Meanwhile, a final inspection technician is focused on sorting parts based on a drawing tolerance. They need to know if the gage can reliably separate in-spec parts from out-of-spec parts. Same parts, same gage, but two different questions lead to two different ways of evaluating the GR&R results.

The "Rule of Thumb" Thresholds for %GRR (%StudyVar)

The most common metric cited from a GR&R study is the %StudyVar, often just called "%GRR". This value tells you what percentage of the total variation observed in your study is due to the measurement system itself. Total study variation combines part-to-part variation and measurement-system variation.

Conceptually, you can think of it as: %StudyVar = (Measurement System Variation / Total Study Variation) * 100

This metric is the most relevant when your primary goal is to understand, monitor, and improve a process. If your measurement system has high %StudyVar, it acts like a fog, making it impossible to see what your process is actually doing. You cannot effectively run control charts or calculate process capability indices like Cpk if your gage is consuming most of the variation.

The widely accepted guidelines for %StudyVar, originally published by the Automotive Industry Action Group (AIAG), are:

  • Less than 10%: The measurement system is generally considered acceptable. It is capable of providing reliable data for process analysis and control.
  • Between 10% and 30%: The measurement system is considered conditionally acceptable. Its suitability depends on the context, such as the safety-critical nature of the dimension, the cost of a better gage, or the current state of process control. This "gray area" requires engineering judgment.
  • Greater than 30%: The measurement system is generally considered unacceptable. It is contributing so much noise that the data is not trustworthy for making decisions about the process. This measurement system must be improved.

Imagine you are measuring the length of a steel shaft. Your GR&R study reports a %StudyVar of 8%. You can be confident that when you see variation in your measurements, you are mostly seeing actual variation in the shafts, not noise from your gage. If the result was 35%, any analysis you perform would be deeply flawed; you would be analyzing measurement error more than product variation.

When to Use %Tolerance Instead of %StudyVar

While %StudyVar is crucial for process control, it can sometimes be misleading if your goal is simply sorting parts. This is where %Tolerance becomes the better metric. Instead of comparing gage variation to the process variation, it compares it to the engineering tolerance band (the distance between your Upper Specification Limit and Lower Specification Limit).

Conceptually, the formula is: %Tolerance = (Measurement System Variation / Total Tolerance) * 100

This metric directly answers the question: "Is my gage good enough to tell an in-spec part from an out-of-spec part?"

Consider a manufacturer producing plastic enclosures with a non-critical wall-thickness check. The process is very stable and produces parts with very little part-to-part variation. If they run a Gage R&R study on the wall-thickness measurement, the %StudyVar might be very high, maybe 40% or 50%. This is because even a small amount of measurement error seems large when compared to the tiny amount of part variation. Based on %StudyVar, the gage would look unacceptable.

However, the engineering tolerance for that wall thickness is wide, so the %Tolerance might be only 5%. In this case, the gage is adequate for its job: ensuring no parts outside the generous specification are shipped. Using %StudyVar alone would have led to a costly and unnecessary project to improve or replace a perfectly functional measurement system.

The decision thresholds for %Tolerance are typically the same as for %StudyVar (less than 10% is good, over 30% is bad), but the interpretation is fundamentally different. Always ask what the data will be used for before deciding which metric to prioritize.

Beyond Percentages: Number of Distinct Categories (ndc)

Another critical output of a GR&R study, especially for process control applications, is the Number of Distinct Categories, or ndc. This value tells you how many different groups of parts your measurement system can reliably distinguish within the observed part variation. It essentially slices your process spread into a number of non-overlapping measurement buckets.

The calculation, which statistical software like Minitab or Excel can perform, is based on the ratio of part variation to measurement variation: ndc = truncate(1.41 * (part-to-part standard deviation / Gage R&R standard deviation))

A low ndc value is a major red flag. If ndc is 1, the measurement system cannot reliably distinguish one part from another. It is not suitable for process-control decisions.

The general guideline is:

  • ndc of 5 or greater: The measurement system is acceptable for process control. It has sufficient resolution to track process changes and support tools like SPC control charts.
  • ndc between 2 and 4: The system has limited utility. It can only tell you about gross shifts in the process. You might be able to use it to track an extremely unstable process, but it is not ideal.
  • ndc of 1: The measurement system is not suitable for process-control decisions because it cannot reliably distinguish between parts.

An ndc of less than 5 tells you that even if your %GRR is conditionally acceptable, you will struggle to use the resulting data effectively for SPC. The control chart will have a "chunky" or "blocky" appearance because the gage cannot resolve small changes in the process average.

What About %Contribution?

When you review a GR&R report from a statistical package like Minitab, you will also see a metric called %Contribution. This value is based on variance, whereas %StudyVar and %Tolerance are based on standard deviation (which is the square root of variance).

%Contribution = (Gage R&R Variance / Total Variance) * 100

The key thing to know is that variances add up directly, while standard deviations do not. Because of this mathematical relationship (squaring), the thresholds for %Contribution look much more stringent:

  • Less than 1%: Generally considered acceptable.
  • Between 1% and 9%: Considered conditionally acceptable.
  • Greater than 9%: Generally considered unacceptable.

Do not be alarmed by these smaller numbers. They are just a different mathematical view of the same result. A %StudyVar of 10% corresponds to a %Contribution of about 1% because 0.10 squared equals 0.01. A %StudyVar of 30% corresponds to a %Contribution of about 9% because 0.30 squared equals 0.09. While some statisticians prefer using %Contribution because of the additive nature of variance, most practitioners find %StudyVar more intuitive to interpret. The important thing is to understand which metric you are looking at and apply the correct guideline.

For a detailed walkthrough of running the analysis and interpreting these outputs in a common software package, a guide to running a Gage R&R in Minitab can be very helpful.

What to Do When Your GR&R is Not Acceptable

A failed GR&R study is not a dead end. It is the beginning of a root-cause investigation. The results provide valuable clues that point you toward improving your measurement process. Before you rush to buy a new, expensive gage, follow a structured approach.

1. First, Check the Study Itself

Before blaming the gage or operators, verify the integrity of your study.

  • Were the parts representative? This is the most common error. If you select parts that are all nearly identical, your Part-to-Part variation will be tiny. This artificially inflates the %GRR and makes the gage look worse than it is. Your sample parts must represent the full, expected range of process variation.
  • Were the operators properly trained? Inconsistent measurement technique between operators (Reproducibility error) is a frequent cause of failure. Ensure everyone follows the same documented procedure.
  • Is the gage suitable for the tolerance? A general rule is that your gage's discrimination (the smallest increment it can read) should be no more than 10% of the tolerance.

2. Diagnose the Source of Variation

Your Gage R&R study breaks down measurement error into two key components: Repeatability and Reproducibility. The relative size of each one points to a different root cause.

  • High Repeatability (Equipment Variation): If repeatability is the largest source of error, the problem is with the gage itself or the measurement procedure. The same operator gets different results when measuring the same part multiple times. The gage may be worn, have insufficient resolution, or be sensitive to environmental factors like temperature or vibration.
  • High Reproducibility (Appraiser Variation): If reproducibility is the dominant problem, the issue lies with the operators. Different operators are getting different results from the same part. This points to a need for better training, a clearer standard operating procedure, or perhaps a more "operator-proof" fixture or setup. For these complex investigations, applying a structured method like the 5 Whys can be incredibly effective in finding the true root cause.

Key Takeaways

  • An acceptable %GRR depends on the context. Use both numerical thresholds and engineering judgment.
  • Choose your metric based on your goal: use %StudyVar for process control and improvement, and use %Tolerance for product sorting and disposition.
  • The general guidelines are the same for both metrics: less than 10% is good, 10% to 30% is conditional, and greater than 30% is unacceptable.
  • For process control, the Number of Distinct Categories (ndc) should be 5 or greater.
  • A failed Gage R&R is a starting point, not an ending. Use the breakdown of Repeatability vs. Reproducibility to diagnose and solve problems within your measurement process.

Frequently asked questions

What is a good GR&R percentage?

A GR&R of less than 10% is generally considered good. However, a result between 10% and 30% can be acceptable depending on the application's criticality and cost of improvement. A result over 30% indicates the measurement system needs improvement.

What is the difference between %StudyVar and %Tolerance in a GR&R study?

%StudyVar compares measurement error to the total process variation from your study, making it ideal for evaluating a system's use for process control (SPC) or capability analysis. %Tolerance compares measurement error to the specification limits, which is best for judging if a gage can reliably sort good parts from bad parts.

What does ndc mean in a Gage R&R study?

The Number of Distinct Categories (ndc) indicates how many separate groups your measurement system can distinguish within your process variation. For process control applications, an ndc of 5 or greater is required to effectively see changes and trends on a control chart.

What should I do if my GR&R fails?

First, review the study itself for errors and ensure part selection was appropriate. Then, break down the GR&R result to see if the problem is Repeatability (the gage) or Reproducibility (the operators). This will guide your investigation toward either improving the equipment and fixturing or clarifying procedures and training.