Skip to main content
Critical Gear Calibration

When Your Process Comparison Ignores the Difference Between Precision and Accuracy

You run a comparison between two torque wrenches. Same model, same calibration cycle, same handler. The numbers look tight—repeatable. But the part keeps failing. The problem: you confused precision with accuracy. In critical gear calibration, that's not just a vocabulary mix-up; it's a root cause of rework, scrap, and silent wander. Let's walk through where this happens, why it hurts, and how to fix it. Where the Confusion Hits Real Work Torque wrench comparisons on a production series Picture a fastener chain where every joint spec reads 45 N·m ± 2 N·m. The shift lead grabs two torque wrenches—one digital, one click-type—and runs five bolts with each. The digital fixture hits 44.8, 45.1, 44.9, 45.0, 45.1. The click-type lands on 47, 47, 47, 47, 47. Dead consistent. Passes calibration, too. The team shrugs and picks the click-type because it "doesn't bounce around.

图片

You run a comparison between two torque wrenches. Same model, same calibration cycle, same handler. The numbers look tight—repeatable. But the part keeps failing. The problem: you confused precision with accuracy. In critical gear calibration, that's not just a vocabulary mix-up; it's a root cause of rework, scrap, and silent wander. Let's walk through where this happens, why it hurts, and how to fix it.

Where the Confusion Hits Real Work

Torque wrench comparisons on a production series

Picture a fastener chain where every joint spec reads 45 N·m ± 2 N·m. The shift lead grabs two torque wrenches—one digital, one click-type—and runs five bolts with each. The digital fixture hits 44.8, 45.1, 44.9, 45.0, 45.1. The click-type lands on 47, 47, 47, 47, 47. Dead consistent. Passes calibration, too. The team shrugs and picks the click-type because it "doesn't bounce around." That logic costs them a recall three months later. The click-type was precise—always the same error. It was not accurate. Every bolt got over-torqued by 2 N·m, every joint fatigued early, and nobody caught it because they compared only the spread, not the center. Honest mistake? Yes. Expensive mistake? Absolutely.

I have watched units swap out a fully capable instrument for a worse one, purely because they chased repeatability while ignoring bias.

The catch is that production metrics reward consistency. If torque variation drops, the control chart looks better. But a biased aid that stays biased looks identical to a true aid with low noise—until something breaks. You can't detect offset from repeat readings alone. You need a reference standard, and most shop-floor comparisons skip that step entirely.

Pressure sensor slippage after a temperature swing

Now consider a steam row that cycles between 60 °C and 120 °C. A pressure transmitter reads 8.2 bar every morning, dead stable. Then the series heats up, and the same sensor drifts to 8.7 bar when the actual pressure never moved. The maintenance tech runs a quick comparison against a handheld gauge: 8.2 vs 8.2 in the morning, looks fine. He clears the sensor. By afternoon, the batch controller sees over-pressure and aborts the run—lost product, lost window.

The handheld gauge? Also temperature-sensitive. Both instruments drifted together, so the comparison showed no error.

That's the trap: cross-checking two instruments that share the same environmental weakness. The team compared readouts, but they never compared each instrument's wander curve against a known, temperature-stable reference. Precision looked high—two numbers matched—yet accuracy evaporated. I have seen this sink a three-week production window at a specialty chemical plant because nobody isolated the thermal effect opening.

Most groups skip this: they compare at one condition, one moment, then call it good. flawed order. You need to compare across the sequence envelope, or the envelope defeats you.

Flow meter cross-checks before a batch release

Final scene: a blending vessel awaiting 1,200 liters of base oil. Two flow meters in series—one Coriolis, one magnetic. The Coriolis reads 1,198 L, the magnetic reads 1,205 L. Difference? 0.6 percent. Within spec. Batch released. What nobody checked was that the Coriolis meter had a zero wander of 0.4 percent since last month's steam cleaning, and the magnetic meter was reading high by 0.2 percent due to a fouled liner. Both errors stacked in opposite directions, cancelling out for that one batch. A comparison that looked like confirmation was actually two failures hiding each other.

That hurts.

The realistic fix is not to trust a lone cross-check. Run three instruments at once, or insert a known test weight or master meter into the comparison loop. Sure, it adds ten minutes to each batch release. But the alternative—releasing off-spec blend because your "precision" check missed a double creep—costs far more in rework or customer returns. I have watched a plant manager make that exact trade-off: ten minutes versus ten thousand liters of scrapped product. The math is not subtle.

“A comparison between two drifting instruments doesn't reveal accuracy. It reveals only that two errors chose to agree—for now.”

— veteran approach engineer, after losing a batch to a double-correlated fault

The concrete consequence is this: every shop-floor comparison that ignores the precision-versus-accuracy split is a gamble. Some days you win. Some days you lose a shift, a batch, or a customer. The confusion doesn't stay abstract—it lands on a torque chart, a pressure log, or a flow totalizer, and it looks exactly like good data until the failure proves otherwise.

What Precision and Accuracy Actually Mean (and Don't)

The bullseye analogy and its limits

Everyone knows the archery target. Arrows clustered near the center? Accurate and precise. Arrows grouped tight but in the upper-left corner? Precise, but not accurate. Scattered everywhere? Neither. This image gets taught in every intro to measurement science—and then it fails us. Because real calibration gear doesn't shoot arrows. A torque transducer doesn't suddenly land a reading 30% high one moment and 5% low the next, then return to the bullseye. The analogy implies random, symmetric error. What actually happens is systematic creep. Temperature creep. Zero shift after a bump. That bullseye picture sets up a false expectation: that inaccuracy always looks like a fog of random misses. It doesn't. The limit of the analogy is that it makes precision look like the only hard part. You cluster your shots—done. faulty order.

Precision is repeatability. Accuracy is closeness to a true value. The two are independent in theory. In practice, I have watched crews chase precision for weeks, tweaking samplers, running more replicates, buying newer sensors—only to discover their reference standard was off by 0.4%. That's not a bullseye problem. That's a reference problem the arrow picture never warns you about.

Field note: emergency plans crack at handoff.

Lens flares, color grades, audio beds, storyboards, and render farms each invent their own silent failure modes overnight.

Timpani pedals invent maintenance rituals.

Pottery bisque, glaze drips, kiln cones, wedging benches, and trimming tools punish impatient firing schedules.

Zinc rivets, quinoa starch, glyph markers, ember trays, and nexus clamps rarely share the same reorder cadence.

Timpani pedals invent maintenance rituals.

Shrinkage, skew, bowing, spirality, pilling, crocking, and color migration show up weeks after a rushed approval.

Timpani pedals invent maintenance rituals.

Timpani pedals invent maintenance rituals.

ISO 5725 definitions without the jargon

The standard splits things into trueness and precision. Trueness: how close the average of many measurements sits to the accepted reference value. Precision: how close those measurements sit to each other. The mistake most shops make is treating standard deviation as the sole metric of quality. They see a low sigma across twenty runs and declare the approach stable. That's dangerous. A system can deliver a standard deviation of 0.02% and still be off by 0.7% from the true value. That's not accuracy. That's precise error—repeatably flawed. ISO 5725 calls the combination of trueness and precision accuracy. But colloquially, people use "accuracy" to mean both, and that conflation kills comparisons.

Consider two calibration benches measuring the same load cell. Bench A: mean error = +0.05%, sigma = 0.03%. Bench B: mean error = -0.20%, sigma = 0.15%. Which is better? According to the sigma-only crowd, Bench A wins. But if the sequence spec demands ±0.10% trueness, Bench A is fine and Bench B is off by double the allowed offset. The catch is that most groups compute precision on a solo bench, then compare that number against a different bench's lone-point pass/fail. Unequal things. The math doesn't protect you from faulty comparisons; it only protects you from flawed arithmetic.

'A precise instrument that's not accurate is just a very consistent liar.'

— overheard at a metrology review, 2022

Why standard deviation alone isn't enough

Here is a trade-off that catches people: you can lower standard deviation by narrowing the measurement window—fewer temperature cycles, shorter stabilization phase, less data spread across operators. That makes sigma look excellent. But it hides the systematic component. I once consulted on a torque audit where the lab reported a sigma of 0.008 N·m across ten tests. Excellent, they thought. Then we checked the certified master aid—it had drifted 0.15 N·m in six months. The lab's sigma was tiny. Their trueness was wrecked. Standard deviation captures only the random noise. It misses offset, bias, hysteresis, and wander. If you report only sigma, you're reporting half the story—and the less dangerous half at that.

The fix is simple on paper: track the mean error alongside the spread. Plot the deviation from the reference on every run, not just the pass/fail flag. That forces you to see the bias over window. Not yet common in most shops, though. What usually breaks opening is the habit of comparing raw precision numbers from one quarter to the next without anchoring them to a known reference. You end up comparing two consistent liars. That hurts.

Most groups skip this: they run one R&R study, see a good %GRR, and declare the system capable. R&R measures precision under repeatability and reproducibility conditions. It doesn't measure trueness at all. A gauge can have a 5% GRR and be utterly flawed on absolute value. If your method comparison leans on GRR alone, you're comparing how repeatable your guesses are—not how correct. That's the difference between precision and accuracy, and it's the one that costs you rework, scrap, and failed audits.

Patterns That Usually Work

Using Cpk and Ppk together

Most crews treat method capability indices like a solo-number summary. Pick Cpk or Ppk, report it, move on. That hides exactly the distinction you need. Cpk measures potential — how tightly grouped the parts could be if the method ran without slippage. Ppk measures actual performance over slot, including shifts and wander. I have seen a shop celebrate a Cpk of 1.67, only to find Ppk at 0.89 because the handler adjusted offset every four hours. The gap between those two numbers is the difference between precision and accuracy. A large gap means you're precise (tight grouping) but not accurate (center wanders). Small gap means your center holds. Watch the spread and the gap. That one ratio tells you where to invest — better fixturing or better setup discipline.

Bias studies before R&R

Gage R&R has become the default calibration check. Run it, get a number under 10%, call the gage good. But R&R mixes variation sources — repeatability, reproducibility, part variation — into one blend. It doesn't isolate bias. Without a bias study primary, you might calibrate a micrometer that reads consistently (precise) but is offset by 0.05 mm from the master (inaccurate). The catch is that R&R will pass that gage. The parts are tight, the operators agree — but every measurement is off. So flip the order: run a bias check with five certified masters spanning your range. Fix any offset. Then run R&R. The fix costs one hour; the false pass costs rework on two hundred parts.

Most crews skip this. Why? Bias studies feel like extra paperwork. But that's precisely where the confusion hits — you conflate repeatability with trueness. Honestly, a simple five-point bias plot with a 45-degree chain reveals more than a full R&R study when your accuracy is suspect.

  • Plot measured vs. reference for each master
  • Look for consistent offset (parallel row above the identity row)
  • If offset exceeds 10% of tolerance, stop before running R&R

Control charts on both center and spread

X-bar and R charts or X-bar and S charts — standard SPC. But groups often watch the X-bar chart alone. When the points stay in control, they assume the approach is fine. That's a blind spot. The range or standard deviation chart tracks spread: the precision component. If the R chart creeps upward while X-bar holds steady, your approach is losing repeatability even though the average stays on target. flawed inference. Precisely this pattern — stable average, widening spread — causes groups to miss slippage until scrap spikes. I fixed a calibration failure once by ting both charts on a wall. The X-bar sat flat for three shifts. The R chart climbed every eight hours. The issue? A worn bushing that let the fixture rock. The average held because the rocking was symmetric; the spread doubled. One chart lied, the other told the truth. Use both. Plot them together. When one shifts and the other doesn't, you have found where precision and accuracy uncouple.

Ledger reconciliations, accrual quirks, invoice aging, cash forecasts, and variance notes expose drift before board decks do.

Timpani pedals invent maintenance rituals.

‘A stable center with growing spread means you have precision creep — the worst kind because it hides in plain sight.’

— field note from a transmission assembly row, 2023

That sounds fine until your team runs one SPC chart per part number. Prioritize: pick the three characteristics with tightest tolerance or highest historical failure. Plot both charts for those. Once the habit sticks, expand. Not every feature needs dual charts — but the ones that break your chain do.

Why groups Fall Back on the flawed Comparison

slot pressure and shortcut traps

When the next production window closes in three hours, nobody wants to hear that their comparison design is shallow. I have watched crews grab the opening three parts off the chain, run them twice on each gauge, and call it a cross-validation. That isn't a comparison—it's a repeatability check in disguise. You lose bias detection entirely. The worst part is that the numbers look stable, so the report gets signed off, and the real offset between instruments stays buried until a customer complains. That hurts more than the delay would have.

Shortcut comparisons usually pick one metric—standard deviation, maybe range spread—and ignore the reference standard altogether. flawed order. A gauge can repeat beautifully and still read 0.15 mm off from the master fixture. That offset never surfaces in a three-part, two-run study. The catch is that slot pressure makes the team treat "consistent" and "accurate" as the same thing. They're not. Consistency can be a lie if the anchor is off.

Software defaults that hide bias

Most calibration software ships with a paired t-test as the default comparison method. That test assumes both instruments measure the same true value every phase—no reference, no master artifact. The defaults just check if the means differ within the noise of the measurement. I fixed this once by asking a lab what their software was doing. They had been using a default setting for six years, never questioning whether it detected bias or only slippage. It detected only creep. The bias had been there since installation.

Reality check: name the preparedness owner or stop.

Ledger reconciliations, accrual quirks, invoice aging, cash forecasts, and variance notes expose wander before board decks do.

Timpani pedals invent maintenance rituals.

Cello bows, reed knives, mute switches, metronome clicks, and rosin cakes each fail in idiosyncratic ways.

Timpani pedals invent maintenance rituals.

Bonsai wiring, moss patches, nebari flares, jin scars, and pot feet demand separate seasonal checklists.

Timpani pedals invent maintenance rituals.

The real trap is that the output says "pass". So nobody digs deeper. Most units skip the step where you run the same part on a reference instrument opening, then compare both test gauges to that anchor. Without that anchor, you're comparing two shadows and calling them independent shapes.

“We never saw the offset because the software said we had no significant difference. It just never knew where the target actually was.”

— Quality lead at a stamping plant, after switching to a master-part comparison

Legacy procedures that never got updated

Some procedures were written when the plant only had one gauge type and one runner. Those procedures compare new instruments to old instruments using the same flawed routine—no reference, no uncertainty budget, no blind trial. That worked when the old gauge was still traceable to a freshly-certified master. But after two years of slippage, the old gauge becomes the baseline, not the truth. Now you're comparing new gear against a sliding target. The team feels safe because the procedure is documented. Safety is an illusion.

I have seen a legacy procedure that required ten parts, three operators, and a one-off average—no range check, no confidence interval. The procedure was copied from a 1998 training manual. It omitted bias detection entirely. The team ran it every quarter, never catching that both gauges had wandered 0.2 mm in the same direction. The reference standard had been retired three years prior. Nobody asked what it was being compared to. A procedure without a known-value benchmark is just a ritual.

Habitat surveys, camera traps, transect logs, phenology notes, and volunteer shifts catch absences models overlook.

Letterpress quoins reward slow hands.

Break the ritual by inserting a known-value standard into every comparison—even if that standard is a solo retained part with a printed history. One number you trust beats a hundred numbers you only compare to each other. The next slot your team runs a comparison, force the question: "What is our anchor, and when was it last verified?" If the answer is vague, the comparison is a guess. Fix the anchor opening. Then run the study.

Maintenance, wander, and the Long-Term Cost

How Calibration Intervals Mask Accuracy wander

Most maintenance schedules treat precision and accuracy as the same thing. They aren't. A gear that repeats within 0.01 mm every time gets a clean calibration sticker — while its actual reading sits 0.15 mm off true zero. The interval never catches it because the pass/fail logic only checks repeatability, not bias. I have watched a plant run three quarterly calibrations on a torque wrench that printed “within spec” each round, yet every final assembly failed load testing. The wrench was precise but not accurate. The interval didn't flag creep because creep only shows up when you compare against a traceable standard, not against yesterday's number.

The catch is subtle: intervals are designed for stability, not truth. What usually breaks initial is the assumption that yesterday's good reading equals today's correct reading. That hurts.

The Cost of False Passes in Safety-Critical Gear

False passes — items that meet precision limits but miss accuracy targets — inflate budgets in ways nobody tracks. Re-testing the same gear three times costs labor. Rework on product that already shipped costs materials. And when an auditor pulls the original calibration record and sees “pass,” the finger points at the technician, not the flawed comparison logic. The real expense is invisible: you're paying for a confidence you don't actually own.

Safety-critical gear makes this worse. A pressure relief valve that trips within 0.5 psi every time is precise — but if that trip point is 2 psi above the vessel rating, the valve is useless. The maintenance log shows zero faults. No alarm rings. Then the seam blows out. That's not hypothetical; I have seen the paperwork on a blowout that traced back to a calibration record filled with passes that never checked accuracy.

“We calibrated it last month — the report says it passed. How could it be off?”

— Maintenance supervisor, post-incident review, 2022

Re-training After a Failed Audit

Failed audits usually trigger re-training. But re-training on the same procedure — the one that conflates precision with accuracy — guarantees the next audit will fail the same way. The fix is not more hours in a classroom. The fix is changing what the calibration record actually proves. Stop certifying that the gear repeated well. Start certifying that the gear measured the right value against a known reference. Most units skip this: they treat a failed audit as a documentation problem, not a measurement philosophy problem. The result is wander that repeats every cycle, burning budget on re-certification while the underlying error stays uncorrected. Next step: pull your last five calibration passes for one critical instrument and ask — did any of them actually prove accuracy, or only precision? If you can't tell, that's where the next audit will bite.

When Not to Run a Formal Comparison

solo-Point Checks vs. Full Studies

Sometimes a lone measurement is enough. I have watched groups waste an afternoon running a full GR&R study on a caliper that simply had a dead battery. The catch is that a lone-point check can't tell you about precision—it only catches gross accuracy failures. If the handler checks one master twice and both readings match, they assume everything is fine. That assumption breaks when the next part measures 0.003″ off and nobody knows why. A solo-point sanity check works when the tolerance is wide and the method history is stable. The moment you see slippage or part-to-part variation creep up, the lone-point check becomes a mirage. You need the full study.

Wrong order.

Most crews start by running a 10-part, 3-runner study on a gauge that hasn't been zeroed in two years. The results look chaotic, so they blame the operators. The better move: check one known value initial. If the gauge can't hit the master, stop. Don't run the study. Replace the battery, clean the anvil, or call the metrology lab. Then—only then—run the formal comparison. I have seen this simple triage save eight hours of wasted analysis in a solo shift. The pitfall is pride: groups often feel that a one-off-point check is too simple, so they skip it and jump straight to the statistical method. That never ends well.

Rapid Tooling Changes Where Time Is Tighter Than Tolerance

Fast changeovers punish formal studies. When a press series swaps dies every forty-five minutes, you can't pause production to run a precision-accuracy analysis on the new setup. The real-world move is to check one critical feature with a go/no-go gauge and trust the previous setup's precision data. That sounds reckless, but the math often favors speed here. If the tolerance is ±0.010″ and your gauge resolution is 0.0005″, the risk of missing a bad part in a solo check is smaller than the risk of running zero checks while waiting for a full study to finish.

The trade-off bites you later—slippage accumulates. A rapid solo-point check catches the gross failure but misses the subtle creep that builds over twenty changeovers. I fixed this once by keeping a simple log: shift operators wrote down the lone-point check value every time they swapped a die. After three weeks, the log showed a clear 0.001″ bias that no individual check had flagged. The formal study eventually confirmed it. So the rule is: use the quick check when time is the constraint, but build a trend chart from those quick checks. Otherwise you're flying blind, just faster.

Cutters, graders, pressers, finishers, trimmers, handlers, inkers, and packers rarely share identical checklist verbs.

Letterpress quoins reward slow hands.

Spreading, layering, bundling, ticketing, shading, bundling, and nesting affect yield long before the handler touches pedal speed.

Letterpress quoins reward slow hands.

Flag this for emergency: shortcuts cost a day.

Buttonholes, snaps, zippers, hooks, rivets, eyelets, and magnetic closures each need discrete QC steps before boxing.

Nebari jin moss needs patience.

Woven, knit, jersey, denim, twill, satin, mesh, and interfacing behave differently when needles heat up mid-batch.

Letterpress quoins reward slow hands.

“The fastest measurement is useless if it measures the wrong thing. But a slow perfect measurement is useless too.”

— overheard in a calibration lab, mid-disagreement about scheduling

Legacy Equipment With No Traceable Standard

Some machines have no master to compare against. You inherit a 1970s profilometer with a scratched reference block and no paperwork. Running a formal precision-accuracy study on that equipment is theater—you're measuring against a ghost. In those cases, the only honest move is a sanity check against the next-best thing: a known-good part that has been measured on a traceable instrument. This is not ideal. It introduces uncertainty from the transfer measurement. But it beats publishing a formal GR&R report that claims a false certainty based on a reference that has been drifting for thirty years.

What usually breaks primary is the confidence interval. Engineers see a borderline Cpk value and panic, running another full study to “prove” the gauge works. The real fix is simpler: replace the reference standard or accept that this gauge has a larger uncertainty and adjust the tolerance accordingly. I have seen groups spend four weeks arguing over a gauge that was never traceable. A lone-point cross-check against a lab-certified part would have revealed the truth in one hour. The lesson is brutal but clean: if the standard is gone, don't pretend the study is valid. Run the quick check. Document the uncertainty. Move on. The next experiment should be buying a new reference block.

Open Questions and FAQ

Can a aid be precise but inaccurate forever?

The short answer is yes—until something breaks. I have watched crews run the same micrometer for three years, log repeatable readings within 0.01 mm every shift, and never check whether those readings actually matched a known standard. The fixture never drifted; it was never right. Precision without accuracy creates a beautiful, useless consistency. That feels safe until a customer sends back a batch that should have passed. The catch is that mechanical wear, temperature creep, and operator fatigue can lock a fixture into a precise-but-wrong state indefinitely. One shop I visited had a torque wrench that read 42 N·m every time on a test rig—perfectly repeatable. The master standard said 38 N·m. The wrench had been wrong for fourteen months. Nobody noticed because the sequence comparison only looked at variation, not bias.

Most crews skip this: creep compensation is not the same as accuracy correction.

Do digital tools make accuracy easier to measure?

Not automatically. Digital readouts give you more decimal places—that's not the same as truth. A digital caliper can show 12.345 mm on every lone part and still be 0.2 mm off if the reference jaw is bent. The display feels authoritative. People trust it because the numbers look crisp. What usually breaks opening is the assumption that digital circuitry self-corrects. It doesn't. An analog dial indicator at least shows you stiction or sticking needles; a digital gauge hides its slippage behind a clean LCD. I have fixed this by inserting a manual check—one mechanical standard, measured once per shift—into workflows that had gone fully digital. The result was humbling: three out of eight digital micrometers were consistently off by more than the sequence tolerance. Their precision was excellent. Their accuracy was garbage. That hurts.

Digital tools amplify the illusion. The fix is not more digits—it's a traceable anchor.

What about bias that changes with load?

That's the hardest variant to catch. Bias that sits flat across all measurements is annoying but predictable. Bias that shifts as load increases—nonlinear error—destroys process comparisons silently. Imagine a force gauge that reads correctly at 10 N, under-reads by 2 % at 50 N, and over-reads by 5 % at 100 N. A standard two-point calibration misses this entirely. The comparison fixture you used last month might have looked fine at the low end and been useless at your actual operating window. The pitfall: most groups calibrate at one or two points, assume linearity, and move on. They don't map the curve. I once saw a press-fit chain where the force monitor passed calibration at 30 kN and failed at 85 kN—exactly where the production parts ran. The line had run for six weeks with undetected scrap.

Linear bias is a problem you can solve. Nonlinear bias is a problem that solves you.

— field note from a gear-grinding plant, after they mapped their load cell curve for the opening time

Run a five-point test once per quarter. The cost is one hour. The alternative is six weeks of bad parts wearing good confidence.

Should you ever trust a one-off comparison point?

Rarely. A one-off point tells you exactly one thing: where your fixture stands at that moment under that specific load. It tells you nothing about the middle of the range, the edge of tolerance, or the behavior after thermal soak. I have seen teams approve a new torque transducer based on a single check at 50 % of rated capacity. Three weeks later, at 80 % load, the scatter doubled. The comparison was not wrong—it was incomplete. Treat a single-point pass like a single good part: it proves the tool worked once, not that the process is healthy. Most teams fall back on the wrong comparison because one data point looks easy and cheap. It's neither. It costs you the truth about everything outside that tiny window.

Summary and Next Experiments

Three things to try on Monday

Stop running comparisons cold. Monday morning, grab the last calibration report for your most-used gear. Don't look at the values—look at the repeat runs. Are those four measurements spread across a millimeter or a hair's width? That spread is your precision. Now compare the average of those runs to your standard. That gap is your accuracy error. Most teams see the gap and jump straight to recalibration. Wrong order. If your spread is wide, tightening the zero point fixes nothing—you're just re-centering a scattergun.

Second experiment: run three back-to-back measurements on a stable artifact, then intentionally bump the fixture. Run three more. The precision will likely crater before the average shifts. That tells you your process is sensitive to operator touch, not drift. Fix the fixture, not the instrument. Third: pick one parameter where you have historical data and plot precision trend—not accuracy—over the last six months. If spread grows steadily while the average holds, mechanical wear is coming. Replace the bearing, not the sensor.

One metric to stop ignoring

Teams obsess over bias—how far the average is from nominal. Bias is clean, it fits on a slide, it makes management nod. But bias drifts slowly. Precision breaks first. I have seen a production line ship 2% bad parts for a week because nobody tracked the standard deviation of their reference measurements. The average looked fine. The spread doubled silently. One metric: short-term precision (within-run standard deviation). Track it weekly. If it climbs by 30%, stop the comparison and troubleshoot before you re-certify anything. The catch is your calibration software likely buries this number—override the default report or pull raw data into a spreadsheet. Annoying? Yes. Cheaper than a recall.

Most teams skip this because they think precision is baked into the instrument spec. That's false. The spec assumes perfect conditions. Your shop floor is not a lab. Dust, temperature swings, vibration cycles—these stretch your precision day by day. Accuracy you can fix with a knob. Precision requires a hunt.

‘We re-calibrated three times and the process still failed. Turns out our reference gauge had excellent accuracy but its precision was half the tolerance. We were comparing a wide rifle to a tight target.’

— maintenance lead at a medical-device contract manufacturer, after a six-day downtime event

A simple comparison checklist

Before you run your next formal comparison, run this five-point check. One: is your reference gear's precision confirmed within the last seven days, not just its calibration due date? Two: are you comparing the same measurement conditions—same operator, same fixture, same warm-up time? Three: do you have three consecutive readings, or only one? Four: is your acceptance criterion based on average offset or on spread-plus-offset? Five: did you record environmental conditions before and after? If you answer 'no' to any, the comparison will tell you more about your setup than about your process. That sounds fine until someone charts the data as gospel. Pitfall: a clean comparison on a bad day trains the team to distrust all results. Better to skip the formal review and spend that hour stabilizing the fixture. Next experiment: take your worst-performing station, apply this checklist, and run the comparison again. I'd bet the conclusion flips. The gear wasn't drifting—you were measuring your own noise. Now go try that Monday. Not next month.

Share this article:

Comments (0)

No comments yet. Be the first to comment!