Reliability & Validity

Test-Retest Reliability: A Complete Guide for Researchers

Table of Contents

Test-retest reliability is the extent to which a measurement produces consistent scores when administered to the same people on separate occasions, assuming the underlying construct has not genuinely changed. It is usually evaluated with an intraclass correlation coefficient or another statistic suited to the data type, alongside evidence about agreement, measurement error, and systematic change.

Test-Retest Reliability

Introduction

Researchers need to know whether a measurement produces dependable scores. If a supposedly stable characteristic is measured today and again under comparable conditions next week, large unexplained differences may indicate that the instrument, administration procedure, scoring system, or testing environment is unstable.

Test-retest reliability evaluates this stability over time. It is used in psychology, education, medicine, nursing, rehabilitation, sociology, organizational research, laboratory science, and many other fields in which the same construct may be measured repeatedly.

This guide explains what test-retest reliability measures, when it is appropriate, how to design a test-retest study, which statistic to select, how to interpret the findings, and how to report the results accurately.

Key takeaways

  • Test-retest reliability concerns the stability of scores across measurement occasions.
  • The construct should be reasonably stable during the retest period.
  • Pearson’s correlation measures rank-order association but does not establish absolute agreement.
  • An appropriately specified ICC is often preferred for continuous scores.
  • Confidence intervals, systematic change, and measurement-error statistics should accompany the reliability coefficient.
  • A reliability result applies to a particular population, procedure, interval, and context; it is not a permanent property of a test.

What Is Test-Retest Reliability?

Test-retest reliability is a form of reliability evidence that examines whether a measurement produces similar results when it is repeated with the same participants under comparable conditions.

The procedure normally involves:

  1. Administering an instrument to a group of participants.
  2. Waiting for an appropriate interval.
  3. Administering the instrument again to the same participants.
  4. Comparing the scores from the two occasions.
  5. Evaluating whether any differences reflect measurement error, systematic change, or genuine change in the construct.

The method is also called retest reliability, temporal stability, or, in some contexts, stability reliability. The terms repeatability and reproducibility are sometimes used, although technical definitions vary between disciplines.

Simple example

A researcher develops a questionnaire intended to measure a relatively stable study-habits trait. The questionnaire is completed by 100 university students and administered again two weeks later.

If students who scored high during the first administration generally score high again, and their numerical scores remain sufficiently close, the questionnaire has evidence of test-retest reliability for that population and interval.

What Does Test-Retest Reliability Actually Measure?

The phrase “similar results” can refer to several different properties. A complete test-retest analysis should distinguish among rank-order stability, absolute agreement, and measurement error.

1. Rank-order or relative reliability

Relative reliability asks whether the instrument can consistently distinguish among participants.

For example, do participants with high scores at the first measurement still tend to have high scores at the second measurement?

Statistics such as Pearson’s correlation and some consistency-based ICCs emphasize this rank-order relationship.

2. Absolute agreement

Absolute agreement asks whether participants receive approximately the same numerical scores on both occasions.

Suppose every student scores exactly five points higher during the retest. Their rank order would be unchanged, producing a perfect Pearson correlation. Nevertheless, the scores would not be in absolute agreement because a systematic increase occurred.

Absolute-agreement ICCs and Bland–Altman analyses are more sensitive to this issue.

3. Measurement error

Measurement-error statistics describe how much an individual score may vary because of random or procedural error.

Common measures include:

  • Standard error of measurement
  • Smallest detectable change
  • Minimal detectable change
  • Coefficient of variation
  • Repeatability coefficient
  • Bland–Altman limits of agreement

Reliability and measurement error are related but not identical. A reliability coefficient is influenced by differences among participants, whereas measurement-error statistics describe the magnitude of within-person variation.

The Classical Measurement Perspective

In classical test theory, an observed score is represented conceptually as:

Observed score = true score + measurement error

A test-retest study attempts to determine how much of the observed score variation represents stable differences among participants and how much represents variation between measurement occasions.

This interpretation depends on important assumptions. The underlying attribute should not change substantially, and errors should not be systematically shared across occasions. When these assumptions are unrealistic, a test-retest coefficient may combine true change, temporary fluctuation, practice, memory, and measurement error.

For this reason, the coefficient should not be interpreted without considering the design and the nature of the construct.

When Should Test-Retest Reliability Be Used?

Test-retest reliability is appropriate when all of the following are reasonably true:

  • The same construct is measured at least twice.
  • The same participants complete each measurement.
  • The construct is expected to remain sufficiently stable during the interval.
  • No intervention is intended to change the construct.
  • Administration and scoring conditions can be kept comparable.
  • Repeating the measurement will not fundamentally alter what is being measured.

Potential applications include:

  • Personality-trait questionnaires
  • Stable-attitude measures over short intervals
  • Physical-performance protocols under standardized conditions
  • Laboratory or device measurements
  • Educational placement assessments
  • Clinical outcome measures administered during a clinically stable period
  • Observer-independent scoring instruments
  • Cognitive tests intended to measure stable individual differences
  • Translated or digitally adapted questionnaires

When Is Test-Retest Reliability Inappropriate?

Test-retest reliability may be inappropriate or difficult to interpret when the construct is expected to change naturally.

Examples include:

  • Momentary mood
  • Current pain during a fluctuating condition
  • Daily stress
  • Acute symptoms
  • Knowledge immediately after teaching
  • Fatigue during a changing workload
  • Treatment outcomes during an intervention
  • Attitudes during a major public event
  • Developmental abilities across a long interval
  • Performance tasks with strong learning effects

A low coefficient in these situations does not necessarily mean that the instrument is defective. It may show that the construct itself changed.

Researchers should first ask:

Is stability theoretically expected during the selected interval?

If the answer is no, a different measurement property—such as responsiveness, sensitivity to change, internal consistency, inter-rater reliability, or alternate-form reliability—may be more relevant.

How to Conduct a Test-Retest Reliability Study

Step 1: Define the construct and intended use

State exactly what the instrument measures and how its scores will be used.

A reliability requirement for comparing group averages may differ from the requirement for diagnosing an individual, monitoring a patient, or determining whether a student has passed an examination.

Step 2: Define the target population

Reliability should be evaluated in participants resembling the people for whom the instrument will be used.

Important characteristics may include:

  • Age
  • Education
  • Language
  • Cultural setting
  • Clinical condition
  • Severity level
  • Occupation
  • Score range
  • Technology access

An instrument may be stable in one population but less reliable in another.

Step 3: Plan the sample size

Avoid selecting an arbitrary number of participants.

Sample-size planning can be based on:

  • The expected reliability coefficient
  • The lowest acceptable reliability
  • The desired confidence-interval width
  • Statistical power for testing a target ICC
  • The number of measurement occasions
  • Expected attrition or incomplete retest data

Small samples produce imprecise estimates and wide confidence intervals. A coefficient that appears high may still be compatible with inadequate population reliability when its confidence interval is wide.

Step 4: Select an appropriate retest interval

The interval must reduce recall and immediate practice without allowing substantial true change.

The correct interval depends on the construct, population, reference period, and testing burden. There is no universal rule that every retest should occur after seven days, two weeks, or one month.

Step 5: Standardize the conditions

Keep relevant conditions as similar as possible, including:

  • Instructions
  • Device or platform
  • Room or environmental conditions
  • Time limits
  • Scoring rules
  • Item order, when appropriate
  • Mode of administration
  • Observer training
  • Time of day, when relevant
  • Medication or activity restrictions
  • Software and instrument versions

Standardization should remove irrelevant variation without creating conditions that differ unrealistically from normal use.

Step 6: Keep measurements independent where possible

Researchers or scorers should not use the first result to influence the second result.

Participants should not be shown their original answers unless that information is part of the intended procedure. Assessors should ideally be blinded to earlier scores when subjective scoring is involved.

Step 7: Collect paired data

Each participant’s first score must be matched to the same participant’s second score.

Researchers should document:

  • Number enrolled
  • Number completing each occasion
  • Reasons for missing retests
  • Time interval for each participant
  • Deviations from the intended protocol
  • Changes in health, treatment, or circumstances

Step 8: Analyze both reliability and agreement

Do not depend on a single correlation coefficient.

For continuous scores, a strong analysis commonly includes:

  • Descriptive statistics for both occasions
  • A scatterplot
  • An appropriately selected ICC with a confidence interval
  • Mean or median change
  • An assessment of systematic change
  • SEM, MDC/SDC, coefficient of variation, or limits of agreement
  • A Bland–Altman plot when absolute agreement is relevant

How Long Should the Test-Retest Interval Be?

The interval should be long enough to reduce memory and practice effects but short enough to make genuine change unlikely. Its length must be justified from the construct, population, reference period, and intended use rather than selected from a universal rule.

Measurement contextMain concernInterval consideration
Stable personality traitRecall of previous answersAllow enough time to reduce direct recall while avoiding major life-related change
Knowledge or skill testPractice and learningConsider alternate forms, additional familiarization, or a longer interval
Symptom questionnaireGenuine clinical fluctuationUse a short clinically stable period and document whether participants changed
Physical performance testFamiliarization, fatigue, and recoveryStandardize preparation and ensure sufficient recovery
Laboratory measureBiological and procedural variationControl sampling time, storage, equipment, and relevant physiological factors
Attitude surveyExternal events and opinion changeAvoid intervals containing events likely to change attitudes
Digital assessmentDevice and software changesKeep platform, browser, interface, and software version comparable

Retest intervals and reference periods

The questionnaire’s reference period must also be considered.

For example, a questionnaire asking about symptoms “during the past seven days” should not automatically be repeated after only two days, because the two responses may refer to heavily overlapping periods. Conversely, an excessively long interval may allow the respondent’s condition to change.

Checking whether participants remained stable

In clinical and longitudinal studies, researchers may use an external stability indicator, such as:

  • A global rating of change
  • Confirmation that no treatment changed
  • A clinical-status question
  • An unchanged medication record
  • A stable objective measure

The stability indicator should not simply reproduce the instrument being evaluated.

How Is Test-Retest Reliability Calculated?

The calculation depends on the data type and the research question.

Type of scoreCommon primary approachUseful supplementary analysis
Continuous total or scale scoreAppropriately specified ICCMean change, SEM, MDC/SDC, Bland–Altman limits
Approximately continuous summed Likert scaleICC, when assumptions and scale interpretation are defensibleDescriptive change and Bland–Altman analysis
Ordinal categoryWeighted kappaCategory-specific and percentage agreement
Nominal or dichotomous categoryCohen’s kappa or another appropriate kappaOverall and category-specific agreement
Highly skewed positive measurementICC or variance model after appropriate transformationRatio limits, log-scale Bland–Altman analysis, or coefficient of variation
Repeated count or non-normal outcomeModel suited to the distributionWithin-person error and graphical assessment

No statistic should be selected solely because it is available in a software menu.

Pearson Correlation and Test-Retest Reliability

Pearson’s product-moment correlation coefficient measures the strength of a linear relationship between two sets of scores.

It answers:

Do participants with relatively high scores at Time 1 also have relatively high scores at Time 2?

Pearson’s r may be useful as a supplementary measure of rank-order stability. However, it does not adequately assess absolute agreement.

Why a high Pearson correlation can be misleading

Consider four hypothetical participants:

ParticipantTestRetest
A4045
B4247
C4449
D4651

The retest score is exactly five points higher for every participant.

Pearson’s correlation is 1.00 because rank order is perfectly preserved. Nevertheless, the measurements do not agree: every score increased by five points.

Under one commonly used two-way, absolute-agreement, single-measure ICC calculation, these data produce an ICC of only about .35. The exact value depends on the selected ICC formulation.

This example demonstrates why researchers should not describe a high Pearson correlation as proof that repeated scores are interchangeable.

Intraclass Correlation Coefficient

The intraclass correlation coefficient is commonly used for continuous test-retest scores because it can incorporate both association and agreement, depending on the selected model.

Conceptually, an ICC compares stable between-person variance with total variance:

ICC ≈ between-person variance ÷ total variance

The exact denominator depends on the model and whether systematic occasion differences are treated as disagreement.

Why there are several ICCs

An ICC is not one statistic. Researchers must make three principal decisions.

1. What model matches the design?

Possible models include:

  • One-way random-effects model
  • Two-way random-effects model
  • Two-way mixed-effects model

In a conventional two-occasion test-retest study, the two measurement occasions are often treated as prespecified rather than randomly sampled from every possible occasion. A two-way mixed-effects model may therefore be appropriate in many such designs.

However, a random-effects model may be justified when researchers intend to generalize across a broader population of occasions, raters, devices, or conditions.

2. Is consistency or absolute agreement required?

Consistency ICC: Systematic differences between occasions may be ignored if participants retain the same relative ordering.

Absolute-agreement ICC: A systematic increase or decrease is counted as disagreement.

Absolute agreement is normally more informative when the same numerical score should be reproduced.

3. Is the unit a single score or an average?

Single-measure ICC: Relevant when future decisions will be based on one administration.

Average-measure ICC: Relevant when the final score will be the mean of two or more repeated measurements.

An average-measure ICC is generally higher because averaging reduces random error. It should not be reported as though it describes the reliability of a single administration.

How to report the ICC model

Software packages use different naming systems. Do not report only labels such as “ICC(2,1)” without explaining them.

Report:

  • Statistical model
  • Random or fixed/mixed effects
  • Consistency or absolute agreement
  • Single or average measurement
  • Number of occasions
  • Confidence interval
  • Reason for the choice

Reliability Versus Agreement

Reliability describes how well scores distinguish participants from one another despite error. Agreement describes how close repeated scores are in the original measurement scale.

This difference explains why two studies can show different conclusions.

A highly heterogeneous sample may produce a high ICC because participants differ greatly, even when within-person differences are practically large. A homogeneous sample may produce a low ICC despite small absolute differences because there is little between-person variation.

Therefore:

  • Use an ICC or comparable coefficient to examine relative reliability.
  • Use SEM, SDC/MDC, limits of agreement, or coefficient of variation to quantify absolute measurement error.
  • Judge error against an externally defined acceptable level whenever possible.

Standard Error of Measurement

The standard error of measurement expresses typical score error in the original units of the scale.

A commonly presented approximation is:

SEM = SD × √(1 − reliability)

However, the appropriate standard deviation and reliability coefficient must come from the test-retest design. Where possible, SEM should be obtained from the relevant variance components rather than from an unrelated Cronbach’s alpha.

Interpretation example

If a scale has an SEM of 2.4 points, an individual observed score can be expected to fluctuate by several points because of measurement error even when the underlying construct remains stable.

SEM should not be interpreted as the standard error of a sample mean. They are different concepts.

Smallest Detectable Change or Minimal Detectable Change

The smallest detectable change estimates how large a difference between two scores must be before it is unlikely to reflect measurement error alone.

A frequently used 95% formula is:

MDC₉₅ = 1.96 × √2 × SEM

It may also be called the smallest detectable change, particularly in COSMIN-related literature.

If SEM is 2.4 points:

MDC₉₅ = 1.96 × 1.414 × 2.4 ≈ 6.7 points

An individual change smaller than approximately 6.7 points may fall within expected measurement error under the assumptions of that calculation.

MDC is not the same as a minimally important difference. MDC concerns measurement error; a minimally important difference concerns the smallest change considered meaningful to patients, participants, clinicians, or decision-makers.

Bland–Altman Analysis

A Bland–Altman analysis examines the difference between two measurements for each participant.

For participant i:

Differenceᵢ = Retestᵢ − Testᵢ

The mean difference estimates systematic bias.

Approximate 95% limits of agreement are:

Mean difference ± 1.96 × SD of the differences

A Bland–Altman plot displays:

  • The participant’s mean score on the horizontal axis
  • The test-retest difference on the vertical axis
  • Mean bias
  • Upper and lower limits of agreement

The plot can reveal:

  • Systematic improvement or decline
  • Greater error at high scores
  • Outliers
  • Heteroscedasticity
  • Nonlinear patterns
  • Whether transformation may be needed

Whether the limits are acceptable cannot be decided statistically alone. Researchers must define what magnitude of disagreement is tolerable for the instrument’s intended use.

Cohen’s Kappa and Weighted Kappa

When results are categorical, an ICC may be inappropriate.

Cohen’s kappa

Cohen’s kappa is commonly used for nominal or dichotomous outcomes. It evaluates agreement beyond the amount expected by chance.

Examples include:

  • Present versus absent
  • Pass versus fail
  • Diagnostic category
  • Risk classification

Kappa is affected by category prevalence and marginal distributions. It should therefore be reported with a contingency table and percentage or category-specific agreement.

Weighted kappa

Weighted kappa is appropriate for ordered categories because disagreements can receive different weights.

For example, disagreement between “mild” and “moderate” may be treated as less serious than disagreement between “mild” and “severe.”

Researchers should identify the weighting method because different weighting schemes can produce different values.

How Should Test-Retest Reliability Be Interpreted?

A coefficient should be interpreted in relation to:

  • Intended use
  • Consequences of error
  • Population
  • Score variability
  • Retest interval
  • Measurement conditions
  • Confidence interval
  • Absolute measurement error
  • Previous evidence for similar instruments

One frequently cited descriptive guide for ICC values is:

ICC estimateDescriptive interpretation
Below .50Poor
.50 to .75Moderate
.75 to .90Good
Above .90Excellent

These categories are conventions, not universal laws. A coefficient of .76 may be sufficient for exploratory group research but inadequate for a high-stakes individual decision.

Interpret the confidence interval

Suppose:

ICC = .82, 95% CI [.64, .91]

The point estimate appears good, but the population value could plausibly fall from moderate to excellent. The result is less certain than the point estimate alone suggests.

Interpret both:

  • The estimated reliability
  • The precision of that estimate

Do not use the p-value as the main criterion

A statistically significant ICC or correlation only indicates evidence that the population coefficient differs from a specified null value, often zero. It does not prove that reliability is sufficiently high for practical use.

Worked Hypothetical Example

A researcher develops a 30-item academic self-regulation scale. Sixty students complete it twice, 14 days apart. No educational intervention occurs between administrations.

The analysis produces:

  • Test mean = 72.4
  • Retest mean = 73.9
  • Mean difference = +1.5
  • Absolute-agreement, single-measure ICC = .82
  • 95% confidence interval = [.72, .89]
  • SEM = 2.4 points
  • MDC₉₅ = 6.7 points

Interpretation

The scale demonstrates good relative test-retest reliability in this sample. The retest mean is 1.5 points higher, suggesting a small average increase that should be investigated as possible practice, recall, or temporary change.

The MDC indicates that an individual score change of less than approximately seven points may not clearly exceed expected measurement error.

This conclusion is limited to a similar student population, a 14-day interval, the same mode of administration, and comparable testing conditions.

What Factors Affect Test-Retest Reliability?

Practice effects

Participants may improve because they have encountered the task before. Practice is especially important for cognitive, motor, memory, and performance tests.

Possible responses include:

  • A familiarization session before baseline
  • Alternate but equivalent forms
  • A longer interval
  • Statistical assessment of systematic improvement
  • Multiple preliminary trials

Recall effects

Participants may remember previous answers and repeat them without independently responding to the construct.

Recall can inflate apparent reliability, especially when:

  • The questionnaire is short
  • Items are distinctive
  • The interval is very brief
  • Participants receive feedback
  • Correct answers are obvious

Genuine change

Life events, education, treatment, recovery, disease progression, or natural development may change the underlying construct.

A low coefficient then reflects a mixture of instrument instability and true change.

Fatigue and motivation

Differences in fatigue, stress, motivation, sleep, illness, or concentration can affect repeated scores.

These influences may be random error or meaningful parts of the construct, depending on what the instrument claims to measure.

Changes in administration

Reliability can be reduced by changes in:

  • Instructions
  • Interviewers
  • Devices
  • Scoring rules
  • Language
  • Testing environment
  • Software interface
  • Time limits
  • Data-collection mode

Restricted range

A reliability coefficient depends partly on between-participant variation.

If nearly all participants have similar scores, the ICC or correlation may be low even when repeated scores are close. Conversely, including participants with very different ability or severity levels can increase the coefficient.

The study sample should represent the intended population rather than being made artificially diverse merely to obtain a high ICC.

Floor and ceiling effects

When many participants receive the lowest or highest possible score, the instrument cannot distinguish them effectively. Restricted variability may reduce reliability estimates and obscure change.

Scoring inconsistency

Manual scoring, ambiguous coding rules, or subjective interpretation can introduce error. Test-retest reliability may then be mixed with inter-rater or intra-rater reliability problems.

Sample attrition

Participants who return for retesting may differ systematically from those who do not. Reporting only complete cases without describing attrition can overstate generalizability.

How Can Test-Retest Reliability Be Improved?

  1. Define a construct that should reasonably remain stable.
  2. Select an interval based on theory and evidence.
  3. Standardize administration and scoring.
  4. Train assessors before data collection.
  5. Pilot the complete procedure.
  6. Use clear, unambiguous items.
  7. Reduce unnecessary environmental variation.
  8. Provide familiarization for performance tests.
  9. Use alternate forms when practice or recall is unavoidable.
  10. Record events that could produce genuine change.
  11. Recruit a representative sample covering the intended score range.
  12. Plan a sample large enough for a precise confidence interval.
  13. Examine systematic change rather than reporting correlation alone.
  14. Report absolute measurement error.
  15. Preserve software, scoring, and instrument versions.

Improvement does not mean forcing identical responses. A valid measure of a genuinely changing state should detect real change.

Advantages of Test-Retest Reliability

Direct evaluation of temporal stability

The method directly tests whether scores reproduce over time rather than inferring stability from relationships among items.

Applicable to many instruments

It can be used with:

  • Questionnaires
  • Clinical assessments
  • Laboratory values
  • Physical tests
  • Devices
  • Educational assessments
  • Behavioral tasks
  • Coding systems

Useful for instrument development

Test-retest evidence helps determine whether a new, translated, shortened, digitized, or culturally adapted instrument retains stable scores.

Supports interpretation of longitudinal change

Measurement-error estimates help researchers distinguish likely real change from ordinary score fluctuation.

Easy to explain conceptually

The core design—measure, wait, measure again—is understandable to students, researchers, practitioners, and decision-makers.

Limitations of Test-Retest Reliability

Stability is an assumption

The method cannot automatically separate genuine change from measurement error.

Repeated testing can alter performance

Recall, learning, familiarization, and reactivity may influence the retest.

Attrition may bias the sample

Participants may not complete both occasions.

Results depend on the interval

A coefficient obtained after two days may differ from one obtained after six months.

Results depend on the population

Reliability is not a fixed characteristic that belongs permanently to an instrument.

A high coefficient does not establish validity

An instrument may reproduce the same biased or irrelevant score consistently.

Correlation can conceal systematic error

A high Pearson correlation may coexist with a large mean shift.

ICC values are model-dependent

Different ICC specifications can produce different conclusions from the same data.

Common Mistakes

Mistake 1: Using Cronbach’s alpha as test-retest reliability

Cronbach’s alpha primarily evaluates relationships among items at one measurement occasion. It does not ordinarily measure score stability across time.

Mistake 2: Reporting only Pearson’s correlation

Pearson’s r can remain high despite systematic score differences.

Mistake 3: Reporting “the ICC”

The ICC model, agreement definition, and measurement unit must be identified.

Mistake 4: Ignoring confidence intervals

A point estimate without uncertainty may be misleading, especially in small samples.

Mistake 5: Applying a universal cutoff

The required reliability depends on the instrument’s purpose and consequences.

Mistake 6: Choosing an interval without justification

A standard two-week interval is not automatically appropriate.

Mistake 7: Testing participants who changed

Treatment, education, illness, or major events can invalidate the stability assumption.

Mistake 8: Ignoring measurement error

An ICC does not show how many score units represent expected error.

Mistake 9: Treating a non-significant mean difference as agreement

A paired test may lack power. Failure to detect a mean difference does not prove that individual scores agree.

Mistake 10: Claiming that reliability has been permanently established

Reliability evidence should be reconsidered after translation, major revision, mode change, population change, or scoring change.

Test-Retest Reliability Compared With Other Reliability Types

Reliability typeMain questionDesignCommon statistics
Test-retest reliabilityAre scores stable over time?Same participants and measure on separate occasionsICC, correlation, kappa, SEM, limits of agreement
Inter-rater reliabilityDo different raters produce similar scores?Same targets assessed by multiple ratersICC, kappa, weighted kappa
Intra-rater reliabilityDoes the same rater score consistently?Same rater repeats scoringICC, kappa
Internal consistencyDo items intended to measure one construct relate appropriately?One administration of a multi-item scaleOmega, alpha, item correlations
Parallel-forms reliabilityDo equivalent versions produce similar scores?Same participants complete two formsCorrelation, ICC, agreement analysis
Split-half reliabilityAre two parts of a test consistent?One administration divided into halvesSplit-half coefficient and correction

Test-Retest Reliability and Validity

Reliability concerns consistency; validity concerns whether score interpretations are supported for their intended purpose.

A bathroom scale that consistently adds five kilograms may be highly repeatable but inaccurate. Likewise, a questionnaire may consistently measure social desirability instead of the construct named in its title.

Reliability is generally necessary for strong validity evidence, but it is not sufficient.

Researchers should evaluate other evidence, including:

  • Content validity
  • Structural validity
  • Convergent and discriminant evidence
  • Criterion relationships
  • Cross-cultural validity
  • Responsiveness
  • Consequences of score use

Use in Modern Research

Patient-reported outcomes

Clinical researchers use test-retest studies to determine whether questionnaires produce stable scores among patients whose condition has not changed.

Because clinical decisions may concern individual change, ICCs should often be accompanied by SEM, SDC/MDC, or limits of agreement.

Digital and remote assessment

Online questionnaires can standardize instructions and scoring, but they introduce possible variation from:

  • Screen size
  • Device type
  • Browser behavior
  • Connectivity
  • Notifications
  • Input method
  • Accessibility settings
  • Software updates

A change from paper to mobile administration is not simply another retest occasion. It may also be a mode-equivalence study.

Wearable devices and sensors

Repeated sensor measurements may be influenced by positioning, calibration, firmware, sampling frequency, participant movement, and data-processing algorithms.

The tested measurement procedure should include the entire workflow, not only the hardware.

Neuroimaging and complex biomarkers

Reliability can vary by preprocessing pipeline, region, contrast, acquisition protocol, and analytic decision. Selectively reporting only the most reliable outcome can bias conclusions.

Researchers should prespecify outcomes and report uncertainty across relevant analyses.

More than two measurement occasions

When measurements are repeated three or more times, researchers may use:

  • ICC models suited to multiple occasions
  • Mixed-effects models
  • Generalizability theory
  • Latent state-trait models
  • Variance-components models
  • Models of change and occasion-specific effects

The average-measure ICC should be used only when future decisions will actually be based on an average of the same number of measurements.

Statistical Software

SPSS

SPSS can estimate several ICC models through reliability-analysis procedures. Researchers must specify the model, agreement type, and single-versus-average measure rather than accepting defaults without justification.

R

R packages used in reliability analysis include functions for:

  • ICC estimation
  • Confidence intervals
  • Test-retest summaries
  • Bland–Altman plots
  • Mixed-effects models
  • Generalizability analysis
  • Sample-size planning

Package output should be checked against the package documentation because naming conventions differ.

Python

Python libraries can calculate correlations, ICCs, agreement statistics, mixed models, and plots. Researchers should confirm how missing values and unbalanced data are handled.

JASP and jamovi

Graphical statistical applications can make reliability analysis more accessible, but a simple interface does not remove the need to choose the correct statistical model.

Excel

Excel can calculate correlations and descriptive statistics but is not ideal for a rigorous ICC workflow unless formulas are carefully validated. It also makes model selection, confidence intervals, missing-data handling, and reproducibility more difficult.

Artificial Intelligence and Test-Retest Analysis

Artificial intelligence can assist with:

  • Explaining statistical output
  • Drafting reproducible code
  • Checking data structure
  • Generating simulated practice data
  • Creating reporting checklists
  • Identifying inconsistencies between reported methods and results

AI should not independently decide:

  • Whether the construct was stable
  • Which ICC model matches the research design
  • Whether measurement error is acceptable
  • Whether missing data can be ignored
  • Whether a coefficient supports clinical or educational use

Researchers should verify generated code against official documentation and trusted methodological sources.

When AI systems themselves are repeatedly measured, researchers should document:

  • Model name and version
  • Access date
  • Prompt wording
  • System instructions
  • Temperature or randomness settings
  • Sampling parameters
  • Tool access
  • Number of repeated runs
  • Software changes between runs

An AI model update may represent a genuine change in the measurement target rather than measurement error.

How to Report Test-Retest Reliability

Method-section template

Test-retest reliability was evaluated in [number] participants who completed [instrument] on two occasions separated by [interval]. Participants were expected to remain stable because [justification]. Administration conditions were standardized by [procedures]. Reliability was estimated using a [model] intraclass correlation coefficient with [absolute agreement/consistency] for [single/average] measurements. The estimate was reported with a 95% confidence interval. Systematic change and absolute measurement error were evaluated using [mean difference, SEM, SDC/MDC, coefficient of variation, and/or Bland–Altman limits of agreement].

Results-section template

Scores demonstrated [poor/moderate/good/excellent] test-retest reliability, ICC([specification]) = [estimate], 95% CI [[lower], [upper]]. The mean retest difference was [value] points, indicating [no substantial systematic change/a possible increase or decrease]. The SEM was [value], and the SDC/MDC₉₅ was [value]. Bland–Altman limits of agreement ranged from [lower] to [upper]. These results suggest that [context-specific interpretation].

Minimum reporting checklist

Report:

  • Target population
  • Sample size and attrition
  • Instrument and score
  • Number of occasions
  • Actual interval distribution
  • Stability evidence
  • Administration conditions
  • Missing-data handling
  • Statistical coefficient
  • ICC model and rationale, when applicable
  • Agreement definition
  • Single or average measurement
  • Point estimate
  • Confidence interval
  • Systematic change
  • Measurement-error statistic
  • Relevant graph
  • Context-specific interpretation
  • Study limitations

Conclusion

Test-retest reliability evaluates whether a measurement produces stable results across repeated administrations when genuine change is not expected. A rigorous analysis requires more than correlating two score columns. Researchers should justify the retest interval, standardize relevant conditions, select a statistic suited to the data and design, report confidence intervals, and examine absolute measurement error.

An instrument should be described as reliable only for the population, conditions, interval, scoring procedure, and intended use that were actually studied.

About the author

Muhammad Hassan

Muhammad Hassan writes about research design, academic methods and data-analysis concepts for ResearchMethod.net. His work focuses on presenting methodological topics in clear language for students and early-career researchers. Articles are developed from recognized methodological literature and official software documentation.