Sample size in stereology involves two separate decisions: how many independent specimens to study, and how much tissue to sample within each specimen. A pilot study should help resolve both, without confusing precise measurement with adequate biological replication.
The aim is not to count the largest possible number of cells or collect the most images. It is to produce estimates precise enough for the research question, then allocate the remaining effort where it improves the study most. That requires a defined outcome, a defensible sampling design and evidence about variability.
This guide focuses on planning sampling effort and using pilot results. The broader principles of selecting specimens, sections and fields are covered in sampling in stereology.
Separate Biological Sample Size from Sampling Intensity
In a biological comparison, the experimental unit determines the sample size used for inference. This might be an independently treated animal, a donor or an independently assigned group of animals. Sections, microscope fields and counted cells are usually subsamples, not additional independent replicates. Animal studies should distinguish the number of experimental units from the total number of animals, following the ARRIVE guidance on reporting sample size.
Consider a hypothetical experiment with six independently treated animals per group, ten sections per animal and twenty fields per section. Each group contains 1,200 sampled fields, but its biological sample size remains six animals. Treating those fields as 1,200 independent treatment replicates would give a misleading impression of statistical certainty.
Keep the two planning questions separate:
- Biological replication: How many independent units are needed to estimate a population quantity or compare groups?
- Within-specimen sampling: How many blocks, sections, fields and probe observations are needed to estimate each specimen’s outcome adequately?
Write both quantities into the protocol. “We analysed 3,000 cells” does not tell a reader how many independent specimens contributed those cells, or how the cells were distributed among them.
Define What the Pilot Must Decide
A useful pilot starts with decisions, not an arbitrary number of slides. Specify the primary outcome and its reference region before adjusting section intervals or counting grids. Total cell number, cell density, volume fraction and total volume answer different questions; do not treat them as interchangeable endpoints.
Write the intended measurement as a complete statement, such as “total number of marker-positive neurons in the defined nucleus per animal.” Then document the anatomical boundaries, identification criteria and estimator. If the outcome includes several regions or cell populations, assess their sampling requirements separately rather than assuming that one grid suits everything.
Select Specimens That Test the Intended Workflow
Choose pilot material that challenges the proposed method. Where feasible, include examples of the expected range of tissue size, staining quality and target abundance. If treatment may change tissue architecture or cell distribution, avoid calibrating the entire study on pristine control material alone.
There is no useful universal pilot sample size. A pilot intended to test staining and microscope access has a different purpose from one intended to estimate variability between specimens. State which purpose applies, and avoid presenting a small feasibility exercise as a reliable estimate of population variation.
Preserve Data at Each Sampling Level
For every pilot specimen, retain section identifiers, field coordinates, counts or point hits, sampling fractions, the final estimate and elapsed analysis time. Keep the order of the sections intact. A single pooled count makes it harder to identify whether uneven sampling arises between sections or within them.
Use a random start and a defined sampling interval rather than selecting visually attractive sections. The practical implementation belongs in the protocol; systematic uniform random sampling provides the framework for distributing observations across the reference space.
Record practical failures too: unreadable boundaries, lost sections, ambiguous staining and fields that cannot be assessed. Treat these as design questions to resolve, not inconvenient notes to delete once counting begins.
Interpret the Coefficient of Error Correctly
The coefficient of error, or CE, describes the relative sampling uncertainty of an individual stereological estimate. Conceptually, it expresses the standard error associated with repeated sampling of the same specimen relative to the estimated quantity. It is not the standard error of the group mean. Research comparing precision estimators for stereological volume and cell counts demonstrates how section and probe sampling affect this assessment.
The coefficient of variation, or CV, describes variation among specimen estimates: standard deviation divided by their mean. Observed variation between specimens includes both genuine biological differences and sampling error.
A CE of 0.05 therefore does not mean that a study has 95% statistical power, that every estimate lies within 5% of the truth, or that bias is absent. It describes one component of uncertainty. Keep that distinction visible when reviewing software output.
Compare Sampling Error with Biological Variation
Under an additive model with unbiased sampling errors that are uncorrelated with the true specimen values, observed variance can be separated approximately into biological variance and average sampling variance. This decomposition is used in primary research on fractionator sampling and variance in alveolar number estimates.
Observed variance ≈ biological variance + average sampling variance.
For specimen i, an estimated sampling variance is approximately (CEi × estimatei)2. Average those variance estimates when evaluating the contribution of sampling error. Averaging CEs and then squaring the result is not generally the same calculation.
Where specimen values and relative errors are reasonably comparable, a relative approximation is often useful:
CVobserved2 ≈ CVbiological2 + CEtypical2.
Use this as a planning model, not a mechanical acceptance test. A small pilot can produce unstable variance estimates; a negative biological variance obtained by subtraction signals that the components cannot be resolved reliably from those data, not that biological variation is negative.
Choose a Precision Target for the Research Question
Do not adopt a CE threshold simply because another laboratory used it. A target of 0.05 or 0.10 needs a rationale tied to the endpoint, biological variability and intended analysis. Measuring a subtle difference in a homogeneous population may justify more intensive sampling than estimating a large difference among highly variable specimens.
Check the CE method as well as the reported value. Variance estimation under systematic sampling depends on the sampling design and density; the methodological paper The Efficiency of Systematic Sampling in Stereology—Reconsidered addresses this dependence. Record the estimator and its settings rather than reporting an unexplained software number.
Review individual specimens, not just the group average. A satisfactory mean CE can conceal poor precision in a small region or a sparse cell population. Inspect whether demanding cases share a common cause that the sampling plan can address.
Also keep precision separate from measurement validity. Before increasing counts, review region boundaries, object identification and applicable preparation checks. The companion guide to bias, precision and coefficients of error covers that distinction in more detail.
Use the Pilot to Allocate Effort
Compare a manageable set of candidate designs rather than changing every parameter at once. A practical sequence is to test section spacing, then field spacing, then probe density or dimensions. Keep the outcome definition and counting rules constant during these comparisons.
| Pilot finding | Candidate change to test | Check before adoption |
|---|---|---|
| Large changes between sampled sections | Sample more sections across the region | Confirm that the revised interval preserves a valid random start |
| Uneven distribution within sections | Distribute more sampling sites across each section | Compare precision gains with added imaging and stage time |
| Few target events at most sampled sites | Test more sites or a suitable larger probe | Check visibility, identification and estimator requirements |
| Low sampling error but broad variation among specimens | Evaluate additional independent specimens | Recalculate study requirements using the intended analysis |
These are candidate responses, not an automatic order of operations. Time the complete workflow, including section preparation, outlining, imaging and review. Saving ten minutes at the microscope is less useful if the revised design adds several hours of preparation.
Once biological variation dominates, increasingly precise measurements of the same few specimens can become an inefficient use of resources. This allocation principle is developed in the primary methodological paper on optimizing sampling efficiency in biological stereology. It supports balancing effort, not accepting careless measurements.
A Worked Example of Diminishing Returns
Suppose, purely for illustration, that biological CV is 0.20. A pilot design produces a typical CE of 0.12. A denser design reduces CE to 0.06. Under the relative variance model above:
Initial observed CV ≈ √(0.202 + 0.122) = 0.233.
Revised observed CV ≈ √(0.202 + 0.062) = 0.209.
Halving CE reduces predicted observed variation from about 23.3% to 20.9%, not by half. Biological variation remains unchanged.
If the denser design requires much more work, compare its cost with adding independent specimens. The answer depends on specimen availability, preparation costs and the precision needed for each individual estimate. A study centred on individual specimen measurements may value lower CE differently from a study centred on group means.
These numbers are calculated examples, not recommended targets or reported experimental results. In your pilot, compare measured analysis times and actual CE estimates for each candidate design. Do not assume that doubling the number of counted events will produce a predictable reduction in total error across every tissue distribution.
Plan the Number of Independent Specimens
After selecting a workable sampling design, justify biological sample size using the planned analysis. For a group comparison, specify the primary outcome, smallest biologically meaningful difference, expected variability, statistical threshold and desired power. For an estimation study, specify the desired confidence interval precision instead.
Use compatible prior data alongside pilot results. A small pilot alone is unlikely to estimate between-specimen variability reliably. Test how the proposed sample size changes under plausible higher and lower variance assumptions rather than treating one preliminary estimate as certain. These considerations are set out in the ARRIVE guidance on sample size justification and pilot studies.
Match the calculation to the design. Paired observations, clustered treatment assignment and repeated measurements need an analysis that represents those dependencies. Account for foreseeable losses using stated assumptions and exclusions defined in advance. Seek statistical input before collecting the main dataset when the design is complex.
Turn the Pilot into a Fixed, Auditable Protocol
Finish the pilot with explicit decisions: the estimator, reference boundaries, sampling intervals, probe settings, precision rationale, quality checks and planned number of independent specimens. Document what evidence supported each choice.
Define how unusual specimens will be handled before group outcomes are examined. Avoid an informal rule of counting until a result looks satisfactory. If additional sampling is permitted, specify a valid procedure in advance, including how sampling fractions and estimates will be updated.
State whether pilot specimens will enter the final analysis. Do not automatically pool exploratory measurements with main-study data after changing staining, boundaries or sampling methods. Review compatibility with the final protocol and analysis plan first.
Keep the pilot records with the study documentation and carry the design decisions into reporting stereological methods and results. A successful pilot leaves a practical answer to two questions: how much to sample within each specimen, and how many independent specimens the research question requires.