Reliability & Validity

Criterion Validity – Methods, Examples and Threats

Table of Contents

Criterion validity is the extent to which scores from a test, scale, or measurement tool relate to a relevant external criterion, such as a diagnosis, job performance, or future academic achievement. It is commonly assessed through concurrent validity, measured at approximately the same time, or predictive validity, in which the criterion is measured later.

Criterion Validity

Introduction

Researchers often create shorter questionnaires, new screening tools, digital assessments, and automated scoring systems. Before these measures can be used confidently, researchers need evidence that their scores relate to the real-world outcomes or established standards they are intended to represent or predict.

Criterion validity provides this evidence by examining the relationship between a measure and an external criterion.

This guide explains:

  • What criterion validity means.
  • How concurrent and predictive validity differ.
  • How to choose an appropriate criterion.
  • Which statistical methods to use.
  • How to interpret and report the results.
  • Which errors can weaken a criterion-validation study.
  • How criterion validity applies to modern digital and AI-based measurement.

Key Takeaways

  • Criterion validity evaluates how scores relate to an external outcome or defensible reference standard.
  • Concurrent validity uses criterion data collected at approximately the same time; predictive validity uses a later outcome.
  • A popular or established measure is not automatically a true gold standard.
  • Correlation is common, but regression, ROC analysis, sensitivity, specificity, calibration, and agreement analysis may also be necessary.
  • There is no universal validity-coefficient threshold suitable for every field or decision.
  • Criterion-related evidence supports a particular interpretation and use of scores; it does not prove that a test is valid for all purposes.

What Is Criterion Validity?

Criterion validity, also called criterion-related validity, describes how well scores from a measurement instrument correspond with or predict a relevant external variable.

The instrument being evaluated is usually called the predictor, test, or index measure. The external outcome is called the criterion, criterion measure, or reference standard.

For example, researchers might examine whether:

  • An entrance examination predicts first-year university grades.
  • A depression screener agrees with an independent clinical assessment.
  • A recruitment test predicts later job performance.
  • A wearable device agrees with a laboratory reference method.
  • An automated writing score predicts independent human ratings or future course performance.

The observed relationship is often summarised using a validity coefficient, such as a correlation coefficient. However, a validity coefficient is only one part of a broader validity argument.

Criterion validity in modern validity theory

Traditional textbooks often describe criterion validity as a separate “type” of validity. Modern testing standards take a broader approach: validity concerns whether accumulated evidence and theory support a particular interpretation of scores for an intended use.

This distinction matters because a test is not simply “valid” or “invalid” in all circumstances.

A mathematics admissions test, for example, might:

  • Predict grades in engineering effectively.
  • Predict artistic performance poorly.
  • Work differently across institutions.
  • Lose accuracy after a curriculum changes.
  • Predict average performance but be poorly calibrated for individual decisions.

Each use requires relevant evidence. Criterion-related evidence is therefore best understood as evidence based on relationships between test scores and external variables rather than a permanent label attached to the instrument (American Educational Research Association et al., 2014; Kane, 2013; Sireci & Benítez, 2023).

Why Is Criterion Validity Important?

Criterion validity connects measurement scores with outcomes that matter outside the test itself.

It can help researchers and decision-makers determine whether a measure is useful for:

  • Screening.
  • Diagnosis support.
  • Selection and placement.
  • Prediction.
  • Monitoring.
  • Replacing a longer or more expensive instrument.
  • Evaluating automated or digital assessments.

Strong criterion-related evidence may show that a new instrument is faster, less expensive, less burdensome, or easier to administer while retaining a useful relationship with the relevant outcome.

However, strong association alone does not establish that the instrument is fair, causal, comprehensive, or suitable for every decision. Its intended use, population, criterion quality, uncertainty, and consequences must also be considered.

What Are the Types of Criterion Validity?

The two main forms are concurrent validity and predictive validity. The distinction is based primarily on when the criterion is measured.

FeatureConcurrent validityPredictive validity
TimingTest and criterion are measured at approximately the same timeTest is measured first; criterion is observed later
Main questionDoes the new measure correspond with a current reference or outcome?Does the measure forecast a relevant future outcome?
Common useReplacing or screening against an existing procedureAdmissions, recruitment, prognosis, and risk prediction
ExampleNew anxiety screener compared with a current clinical assessmentEntrance test compared with first-year GPA
Main advantageFaster and easier to conductClosely represents actual forecasting use
Main limitationMay not show future predictive performanceRequires follow-up and may suffer from attrition or changing conditions

Concurrent validity

Concurrent validity evaluates the relationship between a measure and a criterion assessed at approximately the same time.

A researcher might administer a new short questionnaire and an accepted reference assessment during the same visit. If the scores relate in the expected way, the result provides concurrent criterion-related evidence.

Examples include:

  • Comparing a new reading test with a current standardised reading assessment.
  • Comparing a digital blood-pressure device with a reference device during the same session.
  • Comparing a short clinical screener with an independent diagnostic assessment conducted during the same period.
  • Comparing a new employee-competency test with current performance data.

Concurrent studies can be efficient, but they have limitations. Current employees, enrolled students, or diagnosed patients may differ from the population in which the measure will eventually be used. Existing selection processes may also restrict the range of scores, weakening or distorting the observed relationship.

Predictive validity

Predictive validity evaluates how well scores obtained now relate to a relevant outcome measured later.

Examples include:

  • University-admission scores predicting first-year academic performance.
  • An aptitude test predicting later training success.
  • A clinical risk score predicting hospital readmission.
  • A school-readiness measure predicting later educational progress.
  • A recruitment assessment predicting job performance after hiring.

Predictive designs are preferable when the instrument will actually be used to forecast future outcomes. They preserve the real time sequence between predictor and criterion, but require follow-up and careful control of attrition, changes in the environment, and missing outcome data.

Are there other types?

Most introductory sources teach two principal forms. Some specialist sources also distinguish:

  • Quasi-predictive validation: Existing predictor information is related to a criterion observed later in the historical record, without a fully prospective design.
  • Postdictive validation: Current scores are related to a relevant event or status that occurred earlier.

These labels can be useful for describing timing, but concurrent and predictive validity remain the most widely recognised categories.

What Is a Criterion Variable?

A criterion variable is the external outcome or reference measure against which test scores are evaluated.

Possible criteria include:

  • Academic grades.
  • Independent clinical diagnoses.
  • Laboratory measurements.
  • Job-performance records.
  • Behavioural observations.
  • Treatment outcomes.
  • Error rates.
  • Sales performance.
  • Recidivism.
  • Human-expert ratings.
  • Verified real-world events.

The criterion must be operationally distinct from the predictor. If the “criterion” simply repeats the same items, data source, or scoring rule, the resulting relationship may reflect shared method rather than meaningful external evidence.

What Makes a Good Criterion?

A defensible criterion should meet the following conditions.

1. Relevance

The criterion must represent an outcome that is directly relevant to the proposed interpretation and use of the test.

A general job-satisfaction score, for instance, may not be a suitable criterion for validating a technical-skills assessment.

2. Reliability

A criterion measured inconsistently introduces error and usually weakens the observed relationship. Performance ratings from an untrained supervisor, for example, may be too unstable to function as a strong criterion.

3. Validity

The criterion itself must have evidence supporting its interpretation. Calling an instrument “established” or “widely used” does not automatically make it a gold standard.

COSMIN guidance is particularly cautious about this issue. In many patient-reported outcomes, no objective gold standard exists. Comparing two questionnaires may therefore provide evidence about hypothesised relationships between measures rather than true criterion validity (de Arruda et al., 2025).

4. Independence

The criterion should be assessed independently of the predictor whenever possible.

If a clinician, supervisor, or marker knows the participant’s test score before assigning the criterion rating, that knowledge may influence the outcome. This problem is called criterion contamination.

5. Appropriate timing

The criterion must be measured at the time required by the intended use. A current performance rating cannot automatically substitute for future performance when a test will be used to select new applicants.

6. Representative coverage

A criterion should adequately represent the important aspects of the outcome.

A job-performance criterion based only on speed may be deficient if quality, teamwork, safety, and customer service are also essential.

7. Practical and ethical feasibility

The criterion must be obtainable without exposing participants to unjustified risk, excessive burden, privacy violations, or unfair decision-making.

How to Test Criterion Validity

Step 1: Define the intended interpretation and use

State exactly what the scores will mean and how they will be used.

Weak statement:

The questionnaire will be validated.

Stronger statement:

Scores from the questionnaire will be interpreted as an indicator of current functional impairment among adults receiving outpatient rehabilitation.

The stronger statement identifies the score interpretation, population, context, and intended purpose.

Step 2: Specify the criterion

Define the external outcome operationally.

Document:

  • What the criterion measures.
  • Who assesses it.
  • How it is scored.
  • When it is measured.
  • What evidence supports its reliability and validity.
  • Why it is appropriate for the intended decision.

Step 3: State the expected relationship

Specify the expected:

  • Direction.
  • Magnitude or minimum useful performance.
  • Time interval.
  • Subgroup pattern.
  • Added predictive value, where relevant.

A negative association may support validity when theory predicts that higher test scores should correspond to lower criterion values. Direction must therefore be interpreted against the scoring system and hypothesis, not against a rule that all valid relationships must be positive.

Step 4: Select the validation design

Choose:

  • Concurrent validation for a present reference or outcome.
  • Predictive validation for a later outcome.
  • A prospective, retrospective, or quasi-predictive structure where justified.

The study design should reproduce the intended real-world use as closely as possible.

Step 5: Recruit an appropriate sample

The validation sample should represent the population in which the instrument will be used.

Consider:

  • Age.
  • Language.
  • Education.
  • Clinical status.
  • Occupational role.
  • Severity range.
  • Cultural context.
  • Relevant demographic groups.

Determine sample size using an a priori power or precision analysis suited to the planned statistic. There is no single sample-size rule suitable for every validation study.

Step 6: Collect predictor and criterion data carefully

Use standardised administration procedures.

Where possible:

  • Blind criterion assessors to predictor scores.
  • Blind predictor scorers to criterion status.
  • Counterbalance test order in concurrent studies.
  • Record the time interval between measurements.
  • Document missing data and attrition.
  • Prevent predictor information from leaking into criterion assessment.

Step 7: Choose the statistical method

Match the analysis to the scale of the variables and the intended inference.

Research situationCommon methods
Two approximately continuous variables with a linear relationshipPearson correlation
Ordinal, skewed, or monotonic dataSpearman correlation
Predicting a continuous criterion while adjusting for covariatesLinear regression
Predicting a binary outcomeLogistic regression
Evaluating discrimination against a binary criterionROC curve and AUC
Evaluating a screening thresholdSensitivity, specificity, predictive values
Comparing categorical classificationsContingency table and suitable agreement/classification statistics
Determining whether a new test adds informationHierarchical regression, likelihood comparison, change in predictive performance
Assessing whether two methods are interchangeableAgreement analysis in addition to correlation

Step 8: Quantify uncertainty

Report confidence intervals around correlations, regression coefficients, AUC values, sensitivity, specificity, and other estimates.

A statistically significant result can still be too small or imprecise to support a practical decision. Conversely, a potentially useful estimate may fail to reach statistical significance in an underpowered sample.

Step 9: Assess generalisability and fairness

Examine whether performance is consistent across:

  • Relevant subgroups.
  • Institutions.
  • Locations.
  • Languages.
  • Time periods.
  • Administration modes.

A single overall coefficient can conceal poor prediction or systematic error for a subgroup. High-stakes assessments may require analyses of differential prediction, calibration, classification errors, and adverse consequences.

Step 10: Replicate and monitor

Validation is an ongoing process.

Reassessment may be needed when:

  • The target population changes.
  • The test format changes.
  • Scoring becomes automated.
  • A new translation is introduced.
  • The criterion changes.
  • The curriculum or job changes.
  • The instrument is moved from supervised to remote administration.
  • Performance deteriorates over time.

How Is Criterion Validity Measured?

Pearson’s correlation coefficient

For two continuous variables with an approximately linear relationship, researchers often calculate Pearson’s product–moment correlation:

[
r_{xy} =
\frac{\sum (x_i-\bar{x})(y_i-\bar{y})}
{\sqrt{\sum (x_i-\bar{x})^2\sum(y_i-\bar{y})^2}}
]

Where:

  • (x_i) represents each predictor score.
  • (y_i) represents each criterion score.
  • (\bar{x}) and (\bar{y}) are the sample means.
  • (r) ranges from −1 to +1.

The sign shows direction, while the absolute magnitude reflects the strength of the linear association.

Interpreting (r^2)

Squaring the correlation gives the coefficient of determination in a simple two-variable setting.

For example:

[
r = .42
]

[
r^2 = .1764
]

Approximately 17.6% of the variance in the criterion is statistically associated with the predictor in that sample. This does not show that the predictor causes the outcome, nor does it mean that the test is 17.6% “valid.”

Spearman’s rank correlation

Spearman’s correlation may be appropriate when:

  • Scores are ordinal.
  • The relationship is monotonic but not linear.
  • Strong outliers undermine Pearson’s correlation.
  • Distributional assumptions are not reasonable.

Researchers should justify the choice rather than automatically selecting Pearson’s (r).

Regression analysis

Regression can estimate criterion performance from test scores while including other variables.

For example, an admissions study might examine whether a new readiness test predicts first-year GPA after accounting for:

  • Prior grades.
  • Language proficiency.
  • School characteristics.
  • Socioeconomic context.

Regression also helps test incremental validity: whether the new measure adds useful predictive information beyond data already available.

ROC curve and AUC

When the criterion is binary—such as diagnosis versus no diagnosis—researchers may use a receiver operating characteristic curve.

The area under the curve, or AUC, summarises how well scores distinguish between the two criterion groups across possible thresholds.

A complete diagnostic or screening analysis should normally include more than AUC. Depending on the intended use, researchers may report:

  • Sensitivity.
  • Specificity.
  • Positive predictive value.
  • Negative predictive value.
  • Likelihood ratios.
  • Confidence intervals.
  • Performance at relevant thresholds.

The preferred threshold depends on the relative costs of false-positive and false-negative decisions.

Correlation versus agreement

Correlation measures whether people retain a similar ordering across two measures. It does not show that the measures produce the same values.

Suppose:

[
\text{New device score} = \text{Reference score} + 10
]

The two sets of scores could correlate perfectly, yet the new device would systematically overestimate every measurement by ten units.

When a new method is intended to replace another, researchers should assess agreement or calibration as well as association.

Is there an acceptable criterion-validity coefficient?

There is no universal coefficient that automatically establishes acceptable criterion validity.

Interpretation depends on:

  • The construct.
  • Criterion reliability.
  • Score range.
  • Measurement purpose.
  • Stakes of the decision.
  • Available alternatives.
  • Cost of errors.
  • Prior evidence.
  • Incremental usefulness.
  • Confidence interval.
  • Replication across settings.

A modest association may be useful when the outcome is difficult to predict and the test adds inexpensive information. A much stronger association may be required when two methods are supposed to be interchangeable.

Researchers should specify expectations before examining the results and avoid choosing a threshold after seeing the data.

Worked Criterion-Validity Examples

Example 1: Predictive validity in education

Hypothetical example

A university develops a 20-item academic-readiness test. Researchers administer it to 250 incoming students before the semester and record first-year GPA twelve months later.

They find:

  • Pearson’s (r = .42).
  • The direction is positive, as predicted.
  • The association remains after adjusting for prior grades.
  • Prediction is similar across the major student groups examined.
  • The confidence interval excludes effects considered too small to be useful.

The findings provide predictive criterion-related evidence for using the score as one source of information about first-year academic performance.

They do not justify using the test as the sole admissions criterion. Researchers must also examine fairness, classification consequences, incremental value, and performance in later cohorts.

Example 2: Concurrent validity in clinical screening

Hypothetical example

Researchers develop a short digital screening questionnaire for a clinical condition. Participants complete the questionnaire, and independent clinicians conduct structured assessments during the same week without seeing the questionnaire scores.

Because the criterion is binary, researchers report:

  • AUC.
  • Sensitivity.
  • Specificity.
  • Confidence intervals.
  • False-positive and false-negative rates at proposed thresholds.

The study provides concurrent criterion-related evidence if the questionnaire distinguishes the independently classified groups with sufficient accuracy for its intended screening role.

The result does not establish that the questionnaire can replace a full diagnostic assessment.

Example 3: Criterion contamination in employment research

Hypothetical example

A company compares a pre-employment test with supervisor performance ratings. Supervisors know which employees received high test scores before completing the ratings.

The resulting positive correlation may be inflated because supervisors’ knowledge influenced the criterion. The performance ratings are contaminated by predictor information.

A stronger design would conceal test scores from supervisors, standardise the performance rubric, assess inter-rater consistency, and include objective job outcomes where appropriate.

Criterion Validity Compared With Other Concepts

ConceptMain questionTypical evidence
Criterion validityDo scores relate to a relevant external criterion or outcome?Correlation, regression, ROC analysis, classification performance
Construct validityDo scores behave consistently with the theoretical construct and proposed interpretation?Factor structure, convergent and discriminant relationships, known-group differences
Content validityDoes the instrument adequately cover the relevant content domain?Expert review, content mapping, participant input
Face validityDoes the instrument appear appropriate on the surface?Informal judgement by users or experts
Convergent validityDo scores relate strongly to measures of the same or closely related constructs?Hypothesis-based correlations
Discriminant validityAre scores sufficiently distinct from measures of different constructs?Low or appropriately weaker relationships
ReliabilityAre scores sufficiently consistent or precise?Internal consistency, test–retest reliability, inter-rater reliability

Criterion validity versus construct validity

Criterion-related evidence focuses on relationships with an external criterion. Construct validity is broader and asks whether the complete pattern of evidence supports the proposed meaning and use of the scores.

If a new anxiety questionnaire is compared with another anxiety questionnaire that is widely used but not a true gold standard, the comparison may be better described as convergent evidence within construct validation.

Criterion validity versus concurrent validity

Concurrent validity is not separate from criterion validity. It is one common design for obtaining criterion-related evidence when the test and criterion are assessed at approximately the same time.

Criterion validity versus predictive validity

Predictive validity is the criterion-related design used when the criterion occurs later. It is especially important when scores will be used to forecast future performance, risk, or behaviour.

Criterion validity versus reliability

Reliability concerns consistency and measurement precision. Criterion validity concerns the relationship between scores and a relevant external outcome.

An instrument can be reliable but invalid. For example, a miscalibrated device may produce highly consistent readings that are systematically inaccurate.

Poor reliability also limits validity evidence because measurement error can attenuate relationships with external criteria.

Common Threats and Mistakes

Choosing a convenient rather than defensible criterion

A criterion should be selected because it represents the intended outcome—not merely because the data are easy to access.

Treating a popular measure as a perfect gold standard

Popularity, publication count, or commercial adoption does not establish that a comparator is an error-free criterion.

Reporting only a p-value

A p-value does not show the magnitude, precision, or practical value of a relationship. Report the effect estimate and confidence interval.

Assuming every valid association must be positive

A negative association may be expected. For example, higher resilience scores might predict lower future distress.

Ignoring range restriction

Concurrent employment studies often include only people who were hired. If low-scoring applicants were excluded, the observed predictor and criterion ranges may be restricted, changing the correlation.

Confusing association with agreement

A high correlation does not establish that two instruments produce equivalent values or classifications.

Using the same sample for development and final validation

A model or scoring rule may overfit the development sample. Use an independent validation sample or defensible resampling and external-validation procedures.

Allowing criterion contamination

Outcome assessors should not know predictor scores when that knowledge could influence their ratings.

Ignoring criterion deficiency

A narrow criterion may omit important parts of the outcome. A job-performance measure based only on output quantity, for example, may ignore accuracy, safety, and teamwork.

Ignoring subgroup performance

Overall validity can conceal differential prediction, poor calibration, or unequal error rates across populations.

Assuming concurrent evidence proves future prediction

A relationship among current employees or current patients may not reproduce the conditions faced by new applicants, future students, or untreated populations.

Correcting coefficients without transparency

Statistical corrections for measurement error or range restriction may be defensible, but researchers should report both observed and corrected estimates, explain assumptions, and conduct sensitivity analyses.

Advantages of Criterion Validity

Criterion-related evidence can:

  • Connect scores with practical outcomes.
  • Support screening, prediction, selection, or replacement decisions.
  • Provide an interpretable effect estimate.
  • Compare a new tool with an external standard.
  • Evaluate whether a shorter or less expensive measure remains useful.
  • Test whether a measure adds information beyond existing predictors.
  • Inform classification thresholds and decision rules.
  • Reveal whether performance generalises across settings and groups.

Limitations of Criterion Validity

Important limitations include:

  • Suitable criteria may not exist.
  • Reference standards may be biased or unreliable.
  • Follow-up can be expensive and slow.
  • Predictive studies may suffer from attrition.
  • Correlation does not demonstrate causation.
  • Association may not establish agreement.
  • Results can be population- and context-specific.
  • Restricted score ranges may distort estimates.
  • A single coefficient may conceal nonlinear relationships or subgroup differences.
  • Strong prediction does not automatically establish fairness or ethical acceptability.
  • A measure may predict an outcome for reasons unrelated to the intended construct.

How Criterion Validity Is Used in Modern Research

Education

Criterion-related studies evaluate admissions tests, placement assessments, early-warning systems, automated marking, and measures intended to predict academic progress.

Psychology

Researchers compare screening scales with independent clinical assessments, behavioural outcomes, or future symptom patterns.

Healthcare

Criterion validity is relevant to diagnostic tools, laboratory measures, digital biomarkers, wearables, and clinical outcome assessments. Researchers must be particularly careful before calling another questionnaire a gold standard.

Employment and organisational research

Selection procedures may be evaluated against job performance, training completion, safety outcomes, attendance, or turnover.

In high-stakes employment settings, validation does not remove the need to examine job relevance, adverse impact, accessibility, and less discriminatory alternatives. Employers remain responsible for the appropriateness of the procedures they use.

Digital behaviour and computational research

Text-based measures, platform data, sensors, and automated classifiers may be validated against independently verified behaviours or outcomes.

Digital measures require attention to:

  • Data drift.
  • Platform changes.
  • Missingness.
  • Algorithm updates.
  • Data leakage.
  • Reproducibility.
  • Population shifts.

Digital Tools and Artificial Intelligence

Researchers commonly use R, Python, SPSS, Stata, SAS, jamovi, or JASP to conduct criterion-validity analyses.

These tools can calculate:

  • Correlations and confidence intervals.
  • Linear and logistic regression.
  • ROC curves and AUC.
  • Sensitivity and specificity.
  • Calibration measures.
  • Bootstrap intervals.
  • Cross-validation results.
  • Subgroup performance.

Appropriate uses of AI

Generative AI may assist with:

  • Drafting statistical code.
  • Explaining software output.
  • Checking variable definitions.
  • Producing analysis checklists.
  • Formatting tables.
  • Identifying potential assumptions to examine.

Researchers should independently verify all AI-generated code, calculations, citations, and interpretations.

Validating AI-generated scores

An AI scoring system should not be considered valid merely because it agrees with another AI model or performs well on a convenient benchmark.

Researchers should ask:

  • What real decision or outcome is the score intended to support?
  • Is the criterion independent of the AI system?
  • Could criterion labels contain the same biases as the training data?
  • Has benchmark information leaked into model development?
  • Does performance replicate on new data?
  • Is the system calibrated?
  • Are error rates acceptable for each relevant group?
  • Does performance remain stable after model or prompt changes?

Current AI measurement guidance emphasises explicit evaluation targets, transparent assumptions, uncertainty, reproducibility, and ongoing monitoring rather than reliance on one headline benchmark score (NIST, 2026).

How to Report Criterion Validity

A transparent report should include:

  1. The intended score interpretation and use.
  2. The target population and context.
  3. A description of the predictor.
  4. A description of the criterion.
  5. Evidence supporting the criterion.
  6. The timing of predictor and criterion measurement.
  7. Blinding or independence procedures.
  8. Sample characteristics and recruitment.
  9. Missing-data and attrition procedures.
  10. Statistical methods and assumptions.
  11. Effect estimates and confidence intervals.
  12. Threshold-selection procedures, if applicable.
  13. Subgroup and sensitivity analyses.
  14. Limitations and alternative explanations.
  15. Whether the findings were replicated or externally validated.

Criterion-validity methods template

Criterion-related evidence was evaluated by examining the relationship between [instrument name] and [criterion] among [population]. The criterion was selected because [justification and supporting evidence]. Predictor scores were collected at [time], and criterion data were collected at [time]. Criterion assessors were [blinded/not blinded] to predictor scores. The relationship was analysed using [statistical method] because [justification]. Effect estimates, 95% confidence intervals, missing-data procedures, and prespecified subgroup analyses are reported.

Criterion-validity results template

Scores from [instrument] were [positively/negatively] associated with [criterion], [effect estimate and 95% confidence interval]. This result was [consistent/not consistent] with the prespecified hypothesis. Performance was [similar/different] across [subgroups or settings]. The findings provide [limited/moderate/strong] criterion-related evidence for interpreting the scores as [intended interpretation] in [population and context]. Important limitations include [criterion limitations, sampling, range restriction, missingness, or uncertainty].

Conclusion

Criterion validity evaluates whether scores relate to a relevant external criterion in a way that supports their intended interpretation and use. Concurrent studies examine present relationships, while predictive studies examine later outcomes.

A strong study does more than report a correlation. It justifies the criterion, reproduces the intended use, controls contamination, selects appropriate statistics, reports uncertainty, examines subgroup performance, and tests whether findings generalise. Criterion-related evidence should be treated as one component of an accumulating validity argument—not as permanent proof that a test is valid for every purpose.

About the author

Muhammad Hassan

Muhammad Hassan writes about research design, academic methods and data-analysis concepts for ResearchMethod.net. His work focuses on presenting methodological topics in clear language for students and early-career researchers. Articles are developed from recognized methodological literature and official software documentation.