What Is Measurement System Analysis and Why Does It Matter?
By Hélène ·
The Core Problem: Can You Trust Your Data?
Imagine a scenario on a CNC machining line. An operator on the day shift measures a critical shaft diameter and finds it is drifting toward the upper specification limit. Following the control plan, they apply a negative tool offset to bring the process back to the target nominal. A few hours later, the night shift operator measures the same process and finds the diameter is now trending toward the lower limit. They apply a positive offset, effectively undoing the previous change.
What happened? The process might be unstable, but it's equally possible that the two operators are getting different results from the same measurement system. Perhaps one is using a slightly different technique with the digital caliper, or maybe one of the plant's calipers is out of calibration. Without knowing how much variation is coming from the measurement system itself, the team is flying blind. They are reacting to noise, not signal. This is the fundamental problem that Measurement System Analysis solves.
MSA provides a structured way to quantify the variation introduced by the gages, fixtures, operators, and methods used to measure a part. This "measurement error" is present in every measurement you take. The goal of MSA is not to eliminate measurement error entirely, which is impossible, but to understand it, quantify it, and reduce it to an acceptable level. Only when you know your measurement system is "good enough" can you confidently use its data for critical tasks like statistical process control (SPC), process capability analysis, and root cause problem-solving.
Breaking Down Measurement System Variation
All variation observed in a manufacturing process can be broken down into two components, expressed as variance:
Observed variance = part-to-part variance + measurement-system variance
This equation refers to variance components (the squares of standard deviations), not standard deviations added directly. The total variation you observe in your data is a combination of the real differences between parts and the error introduced by the act of measuring those parts.
- Part-to-part variance: This is the natural, inherent variation in the manufacturing process itself. It's the slight differences between parts caused by tool wear, material fluctuations, environmental changes, and other factors. This is the variation you are trying to monitor and control.
- Measurement-system variance: This is the variation introduced by the act of measuring. It's the "noise" or "error" that clouds your view of the true process variation. If this noise is too loud, it becomes impossible to hear the signal from your process.
MSA focuses on isolating and quantifying the components of Measurement System Variation. The two primary components are known as Repeatability and Reproducibility.
Repeatability (Equipment Variation or EV)
Repeatability describes the variation seen when a single operator measures the same part multiple times with the same gage under the same conditions. It answers the question: "If I measure this exact same feature again, how much spread will I get in my results?"
High repeatability error (poor repeatability) means the gage itself is inconsistent. This is often called "Equipment Variation" (EV) because it points to an issue with the measurement device.
Common causes of poor repeatability include:
- A worn or damaged gage
- Inconsistent clamping or fixturing of the part
- Dirt or contaminants on the gage or part
- An electronic gage with a noisy or unstable sensor
Reproducibility (Appraiser Variation or AV)
Reproducibility describes the variation seen when different operators measure the same part or set of parts using the same gage. It answers the question: "If someone else measures this part, will they get the same answer I did?"
High reproducibility error (poor reproducibility) means there is a significant difference in the average measurements between operators. This is often called "Appraiser Variation" (AV) because it points to an issue with the people or the method.
Common causes of poor reproducibility include:
- Inadequate training on how to use the gage
- Ambiguous measurement instructions or operational definitions (e.g., "measure the diameter" without specifying where)
- Operators using different techniques or levels of force
- Subjective measurements that require interpretation (e.g., visual inspection for cosmetic defects)
Other Sources of Measurement Variation
While repeatability and reproducibility are the most common, other sources of variation exist:
- Bias: A consistent offset between the average of a series of measurements and the true value (a master standard). For example, a scale that always reads 0.1 kg heavy.
- Linearity: A change in bias across the operating range of the gage. A caliper might be accurate for small parts but consistently off for larger parts.
- Stability: The change in measurement variation over time. A Gage R&R study might show the system is good today, but does it stay that way week after week? This is assessed by periodically re-measuring master parts and plotting the results on a control chart.
What is a Gage R&R Study?
For variable measurement systems, one of the primary MSA tools is the Gage Repeatability and Reproducibility (Gage R&R) study. This designed experiment is widely used to quantify Repeatability (EV) and Reproducibility (AV).
In a typical study, you select a number of parts, a number of operators (appraisers), and have each operator measure each part multiple times in a random order. Statistical software is then used to partition the total variation into its components: the part-to-part variation (the actual process variation) and the measurement system variation (Repeatability and Reproducibility).
There are three common types of Gage R&R studies you will encounter:
- Crossed Gage R&R: This is the most common type. It is used for non-destructive testing where every operator can measure every part multiple times. The term "crossed" refers to the fact that all operators measure all parts.
- Nested Gage R&R: This is used for destructive testing where a part cannot be re-measured (e.g., a tensile test that pulls a part until it breaks). In this design, each operator measures a unique set of parts from the same production batch. The analysis can still separate the measurement error from the process variation, but it does so by assuming the parts within a batch are as identical as possible.
- Expanded Gage R&R: Use this when the study needs to account for additional factors such as multiple gages, fixtures, locations, or other conditions, or when the design is unbalanced.
A successful Gage R&R study gives you the data needed to trust your measurement system, which in turn strengthens your ability to apply other statistical tools. For a detailed walkthrough of the software steps, see our guide on how to run a Gage R&R in Minitab step by step.
How to Interpret Gage R&R Results
Running the study is only half the battle; interpreting the output is what drives decisions. While software packages provide many numbers and graphs, you should focus on three key metrics.
%Contribution (or % of Total Variation)
This metric tells you what percentage of the total observed variation is coming from your measurement system. The calculation is Gage R&R variance component / total variance component * 100. The AIAG (Automotive Industry Action Group) provides standard guidelines for acceptance:
- Less than 1%: Acceptable. The measurement system contributes very little to the observed variance.
- 1% to 9%: Conditionally acceptable depending on the application, the cost of the measurement device, the cost of repair, and other practical factors.
- Greater than 9%: Not acceptable. The measurement system should be improved.
%StudyVar: Can the Measurement System Support Process Improvement?
%StudyVar = study variation for source / total study variation * 100
%StudyVar compares the measurement-system variation to the total observed study variation. It answers the question, "How much of the process spread is eaten up by measurement error alone?" This is the relevant question when you intend to use the measurement system to monitor or improve the process. General guidance:
- Under 10%: Excellent. The measurement noise is very small compared to the process spread.
- 10% to 30%: Conditionally acceptable based on the application, cost of the measurement device, or cost of repair.
- Over 30%: Generally unacceptable for process improvement work, because too much of the observed variation is measurement noise.
%Tolerance: Can the Measurement System Reliably Evaluate Parts Against Specifications?
%Tolerance = study variation for source / tolerance * 100
%Tolerance compares the measurement-system variation to the engineering specification tolerance band. It answers a different question: "How much of the tolerance band is consumed by measurement error?" This is the relevant question when you intend to use the measurement system to accept or reject parts against a specification. General guidance:
- Under 10%: Excellent. Measurement error consumes very little of the tolerance.
- 10% to 30%: Conditionally acceptable based on the criticality of the characteristic and the cost of misclassification.
- Over 30%: Generally unacceptable for parts-acceptance decisions, because measurement error makes it very difficult to tell if a part is truly in or out of spec.
Because %StudyVar and %Tolerance answer different questions, a measurement system can perform well on one and poorly on the other; review both when the article scope covers both monitoring and parts acceptance.
When a measurement system is poor, it can be difficult to correctly assess process capability. Poor measurements can disguise an otherwise capable process, or worse, make an incapable process look good. This is why a valid MSA is a prerequisite for a trustworthy analysis of Cp, Cpk, Pp, and Ppk.
Number of Distinct Categories (ndc)
The ndc is one of the most practical metrics from a Gage R&R. It tells you how many distinct groups or categories of parts your measurement system can reliably distinguish within the sample of parts provided. The canonical formula is:
ndc = 1.41 * (part-to-part standard deviation / Gage R&R standard deviation)
The result is truncated to an integer. The general rule of thumb is:
ndc< 2: Unacceptable. The measurement system cannot even tell the difference between one part and another. It is useless for process control.ndc= 2: The system can only classify parts into two groups, typically low and high. It is only useful for very basic screening.ndc= 3 or 4: The system can distinguish three or four groups. This is marginal and may only be acceptable for very wide-tolerance processes.ndc>= 5: Acceptable. The system can distinguish at least five distinct groups. This is generally considered the minimum requirement for a measurement system to be useful for statistical process control (SPC). The more categories, the better your visibility into the process.
Developing the ability to correctly interpret these metrics is a key part of cultivating statistical judgment in manufacturing teams, as it ensures that decisions are based on data that is known to be reliable.
What to Do When Your Gage R&R Study Fails
What should you do when your Gage R&R study fails? The answer is in the data. Look at whether Repeatability or Reproducibility is the larger contributor to the total measurement error.
If Repeatability (%EV) is the main problem:
The issue lies with the gage itself or its interaction with the part. Your investigation should focus on the hardware and the immediate measurement environment.
- Check the gage: Is it damaged, worn, or in need of cleaning? Perform a simple calibration check against a known standard.
- Check the fixturing: Is the part held securely and consistently for every measurement? Is there any play or flex in the clamp?
- Check the measurement procedure: Is the gage being placed on the exact same point every time? Is the part clean and free of burrs or chips that could interfere with the measurement?
- Check the gage design: Is this the right tool for the job? A pair of calipers may not be repeatable enough for a tight tolerance, requiring an upgrade to a micrometer or a CMM.
If Reproducibility (%AV) is the main problem:
The issue lies with the operators or the method. Your investigation should focus on training, documentation, and operational definitions.
- Review the Standard Operating Procedure (SOP): Is the measurement procedure clearly documented with pictures and unambiguous language? Is it available to all operators?
- Observe the operators: Watch each operator perform the measurement. Do they use different techniques? Does one use more pressure than another? Is one reading the gage from a different angle? This is not about blame; it's about standardizing the method.
- Provide training: Conduct a hands-on training session with all operators to ensure everyone agrees on and follows the exact same procedure. Let them practice on known good and bad parts.
- Clarify definitions: Ambiguity is the enemy of reproducibility. "Measure the length" is a poor instruction. "Using caliper PN-123, measure the OAL from surface A to surface B" is much better.
Common Mistakes in Measurement System Analysis
Performing an MSA correctly requires careful planning to avoid introducing bias into the study itself.
- Using non-representative parts: Select parts that represent the actual or expected range of normal process variation. Do not use only consecutive parts, one shift, one production line, hand-picked "golden" samples, or parts selected only from the reject pile.
- Not randomizing the study: Operators should measure the parts in a random sequence. If an operator measures the same part ten times in a row, they may subconsciously try to repeat their last reading. If they measure Part 1, then Part 2, etc., they may remember the results for each part. Randomization prevents this bias.
- Telling appraisers the expected results: The study must be conducted "blind." Operators should not know which part they are measuring (e.g., "Part A" instead of "Serial #1001") or what the known dimension is.
- Focusing only on the final %StudyVar: A system can "pass" the <30% rule but still be flawed. If reproducibility accounts for 90% of the measurement error, you have a serious operator training problem that needs to be fixed, even if the overall number is acceptable.
- Assuming a one-time validation is enough: A measurement system that is good today may not be good in six months. Stability should be monitored over time, especially for critical characteristics. This structured approach to data validation is core to moving from symptoms to root cause effectively.
Key Takeaways
- Measurement System Analysis (MSA) is an important first step before relying on data for process control or capability analysis.
- All observed variation is a combination of actual process variation and measurement system variation.
- Measurement system variation can be broken down into Repeatability (equipment variation) and Reproducibility (appraiser/method variation).
- For variable measurement systems, a Gage R&R study is one of the primary tools used to quantify these sources of measurement error.
- For a measurement system to support process control effectively, %StudyVar should generally be less than 30%, and the Number of Distinct Categories (
ndc) should generally be 5 or greater. - A failed Gage R&R study tells you where to focus your improvement efforts: on the gage itself (repeatability issues) or on the operators and method (reproducibility issues).
Frequently asked questions
What is a good Gage R&R result?
A good Gage R&R result shows low measurement error. Look for a %StudyVar (or %Tolerance) under 10% and a %Contribution under 1%. Most importantly, the Number of Distinct Categories (ndc) should be 5 or greater, indicating the system can effectively distinguish between different parts from your process.
How often should you do a Measurement System Analysis?
MSA is not performed on a fixed time schedule. It should be conducted when a new measurement system is introduced, after a gage undergoes significant repair, and whenever data suggests a potential measurement issue (e.g., conflicting results between shifts or an unexplained change in process capability).
What's the difference between repeatability and reproducibility?
Repeatability is Equipment Variation (EV). It's the variation you see when one person measures the same part multiple times. Reproducibility is Appraiser Variation (AV). It's the variation between the average measurements of different people measuring the same part.
Can I perform MSA for attribute data?
Yes. For attribute data (e.g., pass/fail, go/no-go), the equivalent study is an Attribute Agreement Analysis. This study assesses how well appraisers agree with each other and with a known expert standard on the correct attribute classification of a set of samples.