Calibration curves: fitting, residuals and the uncertainty they add
Fitting a line to calibration data is the easy part. Judging whether the fit is honest, and carrying its uncertainty into the budget, is where most calibration curves go wrong.
A calibration curve exists to answer one question: given what the instrument just indicated, what was the quantity actually worth? It is a correction applied after the fact, derived from comparing the instrument against a reference at several points. The fitting itself takes seconds in any spreadsheet. What takes judgement is deciding whether the fitted line deserves to be believed, and what it contributes to the uncertainty of every result that passes through it.
Which variable goes on which axis
The convention that causes most confusion is the direction of the fit. During calibration the reference is the known quantity and the indication is what you observe, so the natural regression is indication against reference. In use the relationship runs the other way: you have an indication and want the corrected value. Reading the fitted line backwards is called inverse prediction, and it is not simply a matter of rearranging the equation. The uncertainty of an inverse prediction is larger than that of the forward fit, and it grows towards the ends of the range.
For a well-behaved instrument over a narrow range the difference is small enough to ignore. For a wide range, a noisy instrument, or a result that decides conformity, it is not.
Why R² is the wrong thing to look at
The coefficient of determination is quoted in almost every calibration report and tells you almost nothing about whether the model is right. It measures how much of the total spread the fit explains, which is dominated by the range of the data rather than the quality of the fit. Calibrate over a wide range with a genuinely curved response and R² will still read 0.999.
The residual plot answers the question R² cannot. Plot the difference between each measured point and the fitted line against the reference value. If the model is right, those residuals scatter randomly about zero. If they form an arc, a U or an S, the relationship is curved and the straight line is hiding a systematic error that will appear in every corrected result.
| Pattern in the residuals | What it means | What to do |
|---|---|---|
| Random scatter about zero | The model fits | Accept, and use the residual spread as an uncertainty contribution |
| Arc or U shape | The response is curved; a straight line is wrong | Fit a higher order, or split the range and fit each part |
| Funnel, widening with value | The spread grows with magnitude | Weight the fit, or quote uncertainty as a percentage of reading |
| One point far from the rest | An outlier, or a genuine local effect | Investigate before deleting; a deleted point must be justified in the record |
| Step at one value | Something changed mid-calibration | Look for a range change, a re-zero, or a shift in conditions |
The uncertainty the curve itself contributes
A fitted line is an estimate, and estimates carry uncertainty. Two contributions come from the fit, and both are regularly left out of budgets.
The first is the scatter of the calibration points about the fitted line, expressed as the residual standard deviation. It represents everything the model does not explain: repeatability of the instrument, short-term drift during the calibration, and any small mis-specification of the model. It enters the budget as a Type A contribution.
The second is that the line is better determined in the middle of the range than at its ends. The confidence band around a fitted line is narrowest at the centroid of the calibration points and widens towards the extremes. An instrument calibrated between 0 and 100 and then used at 98 is being used where its correction is least certain, which is an argument for placing calibration points at the values that matter rather than at round numbers.
Neither contribution replaces the uncertainty of the reference standard. All three belong in the budget together, combined in the usual way.
Choosing the model, and knowing when to stop
The temptation with a set of calibration points is to raise the order of the polynomial until the residuals shrink. Any set of n points can be fitted exactly by a polynomial of order n minus one, and that fit will have residuals of zero and mean nothing at all, because it has modelled the noise. A higher order is justified when the physics suggests curvature and the residuals from the lower order show a clear pattern, not because the numbers look tidier.
The same discipline applies to forcing the line through the origin. It is tempting because it removes a parameter, and it is wrong unless the instrument genuinely must read zero at zero. Forcing an intercept the data does not support pushes the error into the slope, where it is harder to see.
Checks before you accept a curve
- Do the residuals scatter randomly about zero, with no arc, funnel or step?
- Is every calibration point inside the range the instrument will actually be used over?
- Are the points spread across the range rather than clustered where testing was convenient?
- Does the order of the fit have a physical justification, or was it raised until the numbers looked better?
- Has the residual standard deviation been carried into the uncertainty budget?
- Is the intercept free, unless the physics requires it to be zero?
- Is the fit documented well enough that someone else could reproduce it from the record?
A calibration curve applied silently inside an instrument or a spreadsheet is a correction nobody can audit. Record the coefficients, the range they are valid over, the points they were derived from and the date, and treat a change to them as a change to the measurement.
The common failure, in one sentence
Most calibration curves that cause trouble were accepted on a high R², fitted over a range wider than the data supports, and then used without their own uncertainty ever reaching the budget.
Frequently asked questions
- What is a calibration curve?
- A fitted relationship between what an instrument indicates and the reference value it was compared against, used afterwards to convert an indication into a corrected result. In temperature work it is often a simple offset or a straight line; in instrumental analysis it is usually a regression through several reference points.
- Is a high R² enough to accept a calibration curve?
- No. R² measures how much of the spread the model explains, and a badly curved relationship can still score 0.999. The test that matters is the residual plot: residuals should scatter randomly about zero with no pattern, and no single point should dominate the fit.
- How many calibration points do I need?
- Enough to see the shape and to leave degrees of freedom for the fit. Two points can only ever produce a straight line and can never reveal curvature. Five or six across the working range is a common minimum, spaced to cover the range rather than clustered where the instrument is convenient to test.
- Does the calibration curve add uncertainty?
- Yes, and it is frequently forgotten. The residual standard deviation about the fitted line is an uncertainty contribution in its own right, and at the ends of the range the fitted line is less well determined than in the middle. Both belong in the budget alongside the uncertainty of the reference.
- Can I use the curve outside the calibrated range?
- No. Outside the range there is no evidence about the instrument's behaviour, and a polynomial in particular can depart from reality very quickly beyond its last point. If you need a wider range, calibrate a wider range.
References
- [1]ISO/IEC 17025:2017 — General requirements for the competence of testing and calibration laboratories
- [2]JCGM 100:2008 — Guide to the expression of uncertainty in measurement (GUM)
- [3]JCGM 200:2012 — International vocabulary of metrology (VIM), 3rd edition
- [4]ISO 11095:1996 — Linear calibration using reference materials
General technical guidance written against the cited sources. It is not regulatory or legal advice and does not replace the applicable standard, guideline or a qualified reviewer's judgement.
Have a question on this topic?
Ask ValiTrac AI and see the evidence and calculation behind the answer.
Ask ValiTrac AIFree calculator: Certificate checker
Paste the text of a calibration certificate and screen it against ISO/IEC 17025:2017 clause 7.8: uncertainty with k, traceability, as-found and as-left, the decision rule behind a pass, and every other required element, with the matched text shown for each.
Open the calculatorCalibration on ValiTrac
Support calibration workflows, uncertainty budgets and certificate generation.
See the calibration workflowRelated articles
Choosing calibration points
Calibrate where you measure. Points should bracket the use range and include any temperature where a decision is made.
Building an uncertainty budget step by step
Write the measurement model, list every input, evaluate each as a standard uncertainty, apply sensitivity coefficients, combine in quadrature, expand with k. Eight steps, one table.
Tolerance versus uncertainty
Tolerance is what you require of the instrument; uncertainty is how well the calibration could measure it. A conformity statement needs both.