Skip to content
ValiTracAI

Calibration curves: fitting, residuals and the uncertainty they add

Fitting a line to calibration data is the easy part. Judging whether the fit is honest, and carrying its uncertainty into the budget, is where most calibration curves go wrong.

CalibrationValiTrac AI editorialUpdated 2026-09-135 min read

A calibration curve exists to answer one question: given what the instrument just indicated, what was the quantity actually worth? It is a correction applied after the fact, derived from comparing the instrument against a reference at several points. The fitting itself takes seconds in any spreadsheet. What takes judgement is deciding whether the fitted line deserves to be believed, and what it contributes to the uncertainty of every result that passes through it.

Which variable goes on which axis

The convention that causes most confusion is the direction of the fit. During calibration the reference is the known quantity and the indication is what you observe, so the natural regression is indication against reference. In use the relationship runs the other way: you have an indication and want the corrected value. Reading the fitted line backwards is called inverse prediction, and it is not simply a matter of rearranging the equation. The uncertainty of an inverse prediction is larger than that of the forward fit, and it grows towards the ends of the range.

For a well-behaved instrument over a narrow range the difference is small enough to ignore. For a wide range, a noisy instrument, or a result that decides conformity, it is not.

Why R² is the wrong thing to look at

The coefficient of determination is quoted in almost every calibration report and tells you almost nothing about whether the model is right. It measures how much of the total spread the fit explains, which is dominated by the range of the data rather than the quality of the fit. Calibrate over a wide range with a genuinely curved response and R² will still read 0.999.

The residual plot answers the question R² cannot. Plot the difference between each measured point and the fitted line against the reference value. If the model is right, those residuals scatter randomly about zero. If they form an arc, a U or an S, the relationship is curved and the straight line is hiding a systematic error that will appear in every corrected result.

What a residual pattern is telling you
Pattern in the residualsWhat it meansWhat to do
Random scatter about zeroThe model fitsAccept, and use the residual spread as an uncertainty contribution
Arc or U shapeThe response is curved; a straight line is wrongFit a higher order, or split the range and fit each part
Funnel, widening with valueThe spread grows with magnitudeWeight the fit, or quote uncertainty as a percentage of reading
One point far from the restAn outlier, or a genuine local effectInvestigate before deleting; a deleted point must be justified in the record
Step at one valueSomething changed mid-calibrationLook for a range change, a re-zero, or a shift in conditions
What a residual pattern is telling you

The uncertainty the curve itself contributes

A fitted line is an estimate, and estimates carry uncertainty. Two contributions come from the fit, and both are regularly left out of budgets.

The first is the scatter of the calibration points about the fitted line, expressed as the residual standard deviation. It represents everything the model does not explain: repeatability of the instrument, short-term drift during the calibration, and any small mis-specification of the model. It enters the budget as a Type A contribution.

The second is that the line is better determined in the middle of the range than at its ends. The confidence band around a fitted line is narrowest at the centroid of the calibration points and widens towards the extremes. An instrument calibrated between 0 and 100 and then used at 98 is being used where its correction is least certain, which is an argument for placing calibration points at the values that matter rather than at round numbers.

Neither contribution replaces the uncertainty of the reference standard. All three belong in the budget together, combined in the usual way.

Choosing the model, and knowing when to stop

The temptation with a set of calibration points is to raise the order of the polynomial until the residuals shrink. Any set of n points can be fitted exactly by a polynomial of order n minus one, and that fit will have residuals of zero and mean nothing at all, because it has modelled the noise. A higher order is justified when the physics suggests curvature and the residuals from the lower order show a clear pattern, not because the numbers look tidier.

The same discipline applies to forcing the line through the origin. It is tempting because it removes a parameter, and it is wrong unless the instrument genuinely must read zero at zero. Forcing an intercept the data does not support pushes the error into the slope, where it is harder to see.

Checks before you accept a curve

  • Do the residuals scatter randomly about zero, with no arc, funnel or step?
  • Is every calibration point inside the range the instrument will actually be used over?
  • Are the points spread across the range rather than clustered where testing was convenient?
  • Does the order of the fit have a physical justification, or was it raised until the numbers looked better?
  • Has the residual standard deviation been carried into the uncertainty budget?
  • Is the intercept free, unless the physics requires it to be zero?
  • Is the fit documented well enough that someone else could reproduce it from the record?

A calibration curve applied silently inside an instrument or a spreadsheet is a correction nobody can audit. Record the coefficients, the range they are valid over, the points they were derived from and the date, and treat a change to them as a change to the measurement.

The common failure, in one sentence

Most calibration curves that cause trouble were accepted on a high R², fitted over a range wider than the data supports, and then used without their own uncertainty ever reaching the budget.

Frequently asked questions

What is a calibration curve?
A fitted relationship between what an instrument indicates and the reference value it was compared against, used afterwards to convert an indication into a corrected result. In temperature work it is often a simple offset or a straight line; in instrumental analysis it is usually a regression through several reference points.
Is a high R² enough to accept a calibration curve?
No. R² measures how much of the spread the model explains, and a badly curved relationship can still score 0.999. The test that matters is the residual plot: residuals should scatter randomly about zero with no pattern, and no single point should dominate the fit.
How many calibration points do I need?
Enough to see the shape and to leave degrees of freedom for the fit. Two points can only ever produce a straight line and can never reveal curvature. Five or six across the working range is a common minimum, spaced to cover the range rather than clustered where the instrument is convenient to test.
Does the calibration curve add uncertainty?
Yes, and it is frequently forgotten. The residual standard deviation about the fitted line is an uncertainty contribution in its own right, and at the ends of the range the fitted line is less well determined than in the middle. Both belong in the budget alongside the uncertainty of the reference.
Can I use the curve outside the calibrated range?
No. Outside the range there is no evidence about the instrument's behaviour, and a polynomial in particular can depart from reality very quickly beyond its last point. If you need a wider range, calibrate a wider range.

References

  1. [1]ISO/IEC 17025:2017 — General requirements for the competence of testing and calibration laboratories
  2. [2]JCGM 100:2008 — Guide to the expression of uncertainty in measurement (GUM)
  3. [3]JCGM 200:2012 — International vocabulary of metrology (VIM), 3rd edition
  4. [4]ISO 11095:1996 — Linear calibration using reference materials

General technical guidance written against the cited sources. It is not regulatory or legal advice and does not replace the applicable standard, guideline or a qualified reviewer's judgement.

Related articles