A stereological estimate needs two separate checks: whether the method systematically misses the target, and how much the result would change under repeated sampling. Bias addresses the first question. Precision addresses the second. A coefficient of error, or CE, measures sampling precision relative to the size of the estimate; it is not a certificate of accuracy.
The practical aim is not to produce the smallest possible CE. It is to control bias, obtain enough precision for the research question, and spend the remaining effort where it improves the study. Counting more objects is useful only when it addresses the uncertainty that matters.
Bias and Precision Are Different Problems
Bias is the difference between an estimator’s expected value and the true quantity being estimated. Think of repeating the complete sampling procedure on the same specimen across all possible random starts. If the estimates average above or below the true value, the procedure is biased. This distinction follows the NIST definition of measurement bias.
Precision concerns the spread of those repeated estimates. A precise procedure produces closely grouped results. That grouping says nothing, by itself, about whether they are centered on the truth.
Consider a hypothetical specimen containing 100,000 target cells. One procedure repeatedly produces estimates near 80,000, with little variation. Another produces estimates spread more widely around 100,000. The first is precise but biased. The second is less precise but could be unbiased.
| Question | What it assesses | Practical response |
|---|---|---|
| Are estimates systematically too high or too low? | Bias | Examine the sampling design, identification rules and measurement procedure. |
| Would a new sample give a noticeably different estimate? | Precision | Examine how sampling effort is distributed. |
| Is the result suitable for the planned comparison? | Fitness for the research question | Consider both problems alongside variation between specimens. |
In the hypothetical example, repeatedly applying the procedure that misses 20,000 cells does not repair it. More work can make the wrong answer look impressively stable.
What the Coefficient of Error Measures
The stereological CE describes the variability of an estimate under repeated sampling of the same specimen, expressed relative to its expected magnitude. In practice, that variability is usually predicted from one sampled dataset rather than measured through exhaustive repetitions. The distinction between the underlying CE and its estimated value is developed in Slomianka and West’s study of stereological precision estimators.
For a positive estimate, the working expression is:
Estimated CE = estimated sampling standard error ÷ estimated quantity
If an estimated total is 120,000 cells and its estimated sampling standard error is 6,000 cells, the CE is 0.05, or 5%. These are hypothetical values illustrating the calculation.
This does not mean that exactly 5% of the cells were counted incorrectly. Nor does it mean the true total must lie within 5% of the estimate. A CE is a relative standard error, not a percentage accuracy score or a confidence interval.
Keep its scope explicit. A CE calculated from section counts describes uncertainty represented by that calculation. It does not automatically include uncertainty from staining, anatomical boundaries, classification or calibration.
Also distinguish resampling from recounting. Recounting the same fields can test consistency in applying counting rules. It does not reproduce the uncertainty caused by selecting different sections and fields.
Choose a CE Estimator That Matches the Sampling Design
There is no universal CE formula for every stereological result. A variance calculation must match both the estimator and the sampling scheme. Systematically sampled sections have an ordered spatial relationship; they are not simply a collection of independent observations. Sampling density affects the appropriate variance prediction, as established in the Gundersen–Jensen research on systematic sampling variance.
Before accepting a software output, record which estimator produced it, which sampling stages it covers, and which assumptions it uses. Preserve section order and the section level measurements needed to reproduce the calculation. A column of final specimen totals cannot replace those records.
For a systematic design, document the random start and sampling interval rather than reporting only the number of sections examined. The guide to systematic uniform random sampling covers how to establish that sampling scheme.
Smoothness Settings Need Justification
Some Gundersen–Jensen calculations offer a smoothness parameter, commonly written as m, with options of 0 or 1. These settings concern assumptions about how the measured quantity changes along the section series. They are not interchangeable display preferences.
In an experimental analysis of the mouse dentate gyrus, the better prediction depended on the region, cutting direction and sampling interval. The same study demonstrated that counting many objects cannot overcome substantial variance from sparse section sampling. Both findings are documented in research comparing empirical and predicted stereological CEs.
Do not choose whichever setting produces the smaller number. State the setting, explain its basis, and investigate whether the pilot data support it. When reasonable choices produce different interpretations, report that sensitivity instead of hiding it behind extra decimal places.
CE, CV and SEM Answer Different Questions
The coefficient of variation, or CV, is the standard deviation divided by the mean. Across specimen estimates, it describes the observed relative spread between specimens. That spread includes genuine biological differences and variation introduced by the estimation procedure.
For unbiased specimen estimates under an appropriate variance decomposition:
Observed variance = biological variance + mean sampling variance
Dividing all terms by the same squared group mean gives:
CVobserved2 ≈ CVbiological2 + CEmethod2
Here, CEmethod2 means the average specimen sampling variance divided by the squared group mean. It is not generally the square of the arithmetic mean CE. The distinction matters when specimen totals differ substantially. This formulation is set out in stereological research separating biological and methodological variance.
The standard error of the group mean, or SEM, concerns uncertainty in the estimated group mean rather than the precision of one specimen estimate. For independent, similarly distributed specimen estimates, the familiar calculation is SD divided by the square root of the number of specimens. More elaborate experimental designs require a corresponding statistical model.
A Worked Comparison
Suppose a hypothetical study has an observed CV of 0.20 and a method CE of 0.05, both calculated on a compatible basis. The fraction of observed variance attributed to stereological sampling is:
0.052 ÷ 0.202 = 0.0625, or 6.25%
Under these assumptions, sampling contributes a small share of the observed spread. Reducing that CE from 0.05 to 0.025 would remove three quarters of the sampling variance, but only about 4.7% of the original total variance. Halving the CE sounds dramatic; the gain for the group comparison is much smaller.
Now suppose the observed CV is 0.10 and the method CE is 0.08. Sampling would account for 64% of observed variance. Improving within specimen precision deserves much closer attention in that case.
These calculations illustrate why a CE threshold should not stand alone. A value of 0.05 can be more than adequate for one question and insufficient for another. Judge it against the expected effect, biological variation and planned analysis.
Improve Precision at the Right Sampling Level
A study can allocate effort among specimens, blocks, sections, fields and probe interactions. Once individual estimates are sufficiently precise, additional independent specimens may improve a biological comparison more than further counting within existing specimens. This allocation principle is the subject of Gundersen and Østerby’s research on sampling efficiency.
Use the pilot to ask where uncertainty enters. Compare candidate designs, preserve their randomization, and record the time each requires. Do not compare CE alone: compare the reduction in relevant uncertainty per additional hour of work.
For example, suppose two hypothetical pilot designs take the same time. Design A concentrates measurements on six sections. Design B spreads fewer measurements per section across twelve sections. If the region changes sharply along the cutting axis, test whether broader coverage improves precision before committing to denser counting on the original six.
A useful pilot record includes specimen totals, CE estimates, counts by section, fields examined and time spent. The separate guide to sample size and pilot studies in stereology addresses specimen numbers and planning in more detail.
Keep the stopping rule explicit. Decide in advance how pilot results will determine sampling intensity, and document any later changes. Avoid adding convenient fields until a displayed CE crosses a preferred threshold. The resulting dataset must still correspond to the sampling design used in its analysis.
A Low CE Does Not Rule Out Preparation Bias
Optical counting provides a concrete warning. Tissue compression and loss of particles near section surfaces can produce nonuniform distributions through section depth. Sampling an unrepresentative depth interval can then bias the estimated total. Experimental evidence supporting checks of these distributions is available in research on section compression, lost caps and optical disector bias.
Consider a hypothetical count with a CE of 0.03. If the counting procedure systematically excludes a subgroup of target cells because they cannot be recognized, the small CE does not measure that omission. It describes the precision of the result produced under the existing identification procedure.
Treat bias checks and precision checks as separate parts of quality control. Ask whether the reference region is defined consistently, whether target objects can be identified throughout the sampled material, and whether every recorded sampling fraction corresponds to what was actually measured.
For optical methods, assess the preparation before selecting counting depths. The guide to section thickness, guard zones and tissue shrinkage covers those decisions without treating a fixed guard zone as a universal remedy.
A practical review should also separate disagreement from systematic omission. Two observers can agree because both apply the same mistaken rule. Use recounts to assess consistency, but retain a separate check that the rule measures the intended biological quantity.
Report Precision So Readers Can Evaluate It
Report enough detail to connect each CE to its estimate and sampling design. A single statement that “all CEs were acceptable” leaves the reader guessing about both the calculation and the acceptance criterion.
A compact reporting record should contain:
- The stereological estimator and the quantity estimated.
- The CE estimator, software version and relevant calculation settings.
- The numbers of sections, fields and counted events for each specimen.
- Individual CEs or a clearly described summary with its spread.
- The observed specimen variation and the reason for the chosen precision target.
Label biological variability separately from sampling precision. If displaying error bars, say whether they represent SD, SEM, a confidence interval or sampling uncertainty. If a result combines estimates, identify how uncertainty was propagated rather than attaching the CE of just one component.
Retain the ordered raw measurements and document missing sections, exclusions and departures from the planned procedure. Use the guide to reporting stereological methods and results to place these details within the full methods description.
The final judgment should remain practical: has the study addressed plausible bias, measured sampling precision appropriately, and allocated effort to the research question? A small CE helps answer one part of that assessment. It should never be asked to answer all three.