Reliability & Validity

Parallel Forms Reliability – Definition, Calculation and Examples

Table of Contents

Parallel forms reliability measures the consistency of scores produced by two different but carefully matched versions of the same test. Researchers administer Form A and Form B to the same participants and examine how closely the scores correspond. Strong evidence requires more than a high correlation: the forms should also have comparable content, difficulty, precision, and score distributions.

Parallel Forms Reliability

Parallel forms reliability is especially useful when researchers, teachers, employers, or testing organisations need more than one version of an assessment. Different forms can reduce item exposure, memorisation, answer sharing, and some practice effects. However, producing two genuinely comparable forms is demanding because both must measure the same construct with similar precision.

This guide explains what parallel forms reliability means, how it is calculated, how to design a study, how to interpret the findings, and how it differs from other forms of reliability.

Key takeaways

  • Parallel forms reliability evaluates consistency across two different versions of an instrument.
  • It is commonly estimated by correlating Form A and Form B scores from the same participants.
  • A high correlation shows similar rank ordering but does not, by itself, prove that the forms are interchangeable.
  • Researchers should also compare content coverage, means, variability, item properties, score differences, and uncertainty.
  • There is no universal coefficient that is acceptable for every purpose; the required precision depends on how the scores will be used.
  • Modern form construction may combine test blueprints, item analysis, IRT, automated test assembly, and carefully supervised AI tools.

What Is Parallel Forms Reliability?

Parallel forms reliability is the extent to which two different versions of a measurement instrument produce consistent scores when administered to the same people.

The two versions are usually called Form A and Form B. They contain different items but are designed to measure the same construct, content domain, skills, or learning outcomes.

For example, a researcher may develop two 30-item tests of statistical reasoning. The questions differ, but each form contains comparable numbers of items on probability, hypothesis testing, confidence intervals, and data interpretation. If participants tend to obtain comparable results on both forms, the forms have stronger parallel-forms reliability.

Parallel forms reliability is also called:

These expressions are often treated as synonyms in introductory research literature. In formal psychometric theory, however, the word parallel can imply stronger assumptions than the more general phrase alternate forms.

What Are Parallel Tests in Classical Test Theory?

In classical test theory, an observed score is represented as:

[
X = T + E
]

where:

  • (X) is the observed score,
  • (T) is the true-score component, and
  • (E) is measurement error.

Strictly parallel forms are assumed to measure the same true score, have equal error variances, and have errors that are not correlated. Under these assumptions, the two forms have equal observed-score variances and equal reliability.

Real tests rarely satisfy perfect parallelism. In practice, researchers normally aim for approximately parallel or sufficiently equivalent forms. They use a combination of theoretical, content-based, and statistical evidence to decide whether the forms are suitable for their intended use.

This distinction matters. Two tests can cover the same general subject while differing in difficulty, precision, reading demand, or representation of subtopics. Calling them parallel without supporting evidence can therefore be misleading.

How Does Parallel Forms Reliability Work?

The basic procedure is:

  1. Develop two forms intended to measure the same construct.
  2. Administer both forms to the same participants.
  3. Score the forms using comparable rules.
  4. compare the paired scores.
  5. Estimate the relationship between Form A and Form B.
  6. Investigate whether the forms differ systematically.
  7. Decide whether the observed differences are acceptable for the intended use.

Participants who score relatively highly on Form A should also tend to score highly on Form B. Participants who score lower on one form should generally score lower on the other.

A strong relationship indicates that the ranking of participants is not heavily dependent on the particular set of items administered. Nevertheless, retaining the same rank order is only one part of form equivalence.

Parallel Forms Reliability Formula

For continuous or approximately continuous total scores, parallel forms reliability is commonly estimated with Pearson’s product–moment correlation:

[
r_{AB} =
\frac{
\sum_{i=1}^{n}(A_i-\bar{A})(B_i-\bar{B})
}{
\sqrt{
\sum_{i=1}^{n}(A_i-\bar{A})^2
\sum_{i=1}^{n}(B_i-\bar{B})^2
}
}
]

where:

  • (A_i) is participant (i)’s score on Form A,
  • (B_i) is participant (i)’s score on Form B,
  • (\bar{A}) and (\bar{B}) are the form means,
  • (n) is the number of participants, and
  • (r_{AB}) is the correlation between the forms.

The coefficient normally ranges from −1 to +1:

  • A value close to +1 indicates that participants have very similar relative positions on both forms.
  • A value near 0 indicates little linear relationship.
  • A negative value indicates that higher scores on one form tend to accompany lower scores on the other, usually signalling a serious scoring, coding, or form-development problem.

Should the significance of the correlation be tested?

A null-hypothesis test of whether (r=0) is usually not the main question. With a sufficiently large sample, even a coefficient that is too low for the intended application may be statistically significant.

Researchers should focus on:

  • The estimated coefficient
  • Its confidence interval
  • The minimum acceptable reliability defined before analysis
  • The consequences of measurement error
  • Evidence that the forms have comparable difficulty and precision

Worked Example of Parallel Forms Reliability

Suppose ten students complete two versions of a quantitative-reasoning test.

StudentForm AForm B
14240
25557
36160
46866
57072
67473
78078
88486
99088
109596

For these illustrative data:

  • Form A mean = 71.9
  • Form B mean = 71.6
  • Form A standard deviation ≈ 16.31
  • Form B standard deviation ≈ 16.64
  • Pearson correlation ≈ .994

The coefficient indicates an extremely strong linear relationship. The means and standard deviations are also similar, and the individual differences are small.

A suitable interpretation would be:

Scores on Forms A and B were very strongly related, (r=.99). The forms also produced similar means and standard deviations in this illustrative sample, supporting their use as closely matched versions. A real study should additionally report a confidence interval, administration design, item characteristics, and a predefined criterion for acceptable equivalence.

A coefficient this high should not be expected automatically in real research. The example is deliberately simple so that the calculation and interpretation are easy to see.

How to Conduct a Parallel Forms Reliability Study

1. Define the construct and intended score use

State exactly what the instrument is intended to measure.

A broad label such as “mathematical ability” is insufficient. Specify the relevant content, cognitive processes, population, testing conditions, score range, and decisions that will be based on the results.

Reliability requirements should reflect the stakes. A short classroom quiz and a professional licensing examination do not require identical levels of precision.

2. Create a detailed test blueprint

A test blueprint specifies the content and cognitive demands represented in each form.

For example:

Content areaWeightForm A itemsForm B items
Descriptive statistics25%55
Probability25%55
Hypothesis testing30%66
Research interpretation20%44

Both forms should match on more than the number of items. They should also be comparable in:

  • Learning objectives
  • Cognitive complexity
  • Item format
  • Reading load
  • Response demands
  • Time limits
  • Scoring rules
  • Expected difficulty
  • Content representation
  • Accessibility requirements

3. Develop separate but comparable items

Each form should contain items that measure the same domain without simply repeating the same questions.

Changing only the order of identical questions does not normally create a genuinely different parallel form. It creates another presentation of substantially the same test.

Items can be paired or matched according to:

  • Content objective
  • Difficulty
  • Cognitive process
  • Format
  • Number of response options
  • Expected completion time
  • Discrimination
  • Exposure or security risk

4. Obtain expert review

Subject specialists and measurement reviewers should evaluate whether the forms:

  • Represent the same construct
  • Cover the blueprint adequately
  • Contain comparable cognitive demands
  • Avoid construct-irrelevant differences
  • Use clear and culturally appropriate wording
  • Avoid unfair clues or unnecessary difficulty
  • Apply equivalent scoring rules

Expert judgement does not replace empirical testing, but empirical correlation cannot repair poor content representation.

5. Pilot the items

Pilot data can reveal differences in:

  • Item difficulty
  • Item discrimination
  • Distractor performance
  • Missing responses
  • Completion time
  • Internal consistency
  • Differential item functioning
  • Floor or ceiling effects

Items may need to be revised, removed, or reassigned before the operational reliability study.

6. Select an appropriate participant sample

Participants should represent the population for whom the forms will be used.

Reliability is a property of scores produced in a particular population and context, not an unchanging property permanently attached to an instrument. A coefficient obtained from highly diverse participants may differ from one obtained in a restricted, homogeneous sample.

A narrow score range can reduce the observed correlation even when measurement quality is otherwise reasonable. Researchers should therefore describe the sample and score distribution clearly.

7. Administer both forms under standardised conditions

Keep the following conditions as similar as possible:

  • Instructions
  • Time limits
  • Testing environment
  • Device or delivery mode
  • Scoring procedures
  • Permitted materials
  • Breaks
  • Proctoring
  • Accessibility arrangements

Uncontrolled differences may be mistaken for problems with the forms.

8. Counterbalance the administration order

If everyone takes Form A first and Form B second, form differences become entangled with order effects.

A stronger design randomly assigns participants to two sequences:

  • Group 1: Form A followed by Form B
  • Group 2: Form B followed by Form A

Counterbalancing helps identify or reduce:

  • Fatigue
  • Boredom
  • Familiarity
  • Practice
  • Anxiety reduction
  • Learning between forms
  • Carryover effects

The researcher should also record the interval between administrations.

9. Choose the interval carefully

Administering forms immediately may reduce genuine change in the construct, but performance on the second form may be affected by fatigue or transfer.

Waiting longer may reduce immediate carryover, but it creates opportunities for learning, development, treatment, or real change.

There is no universally correct interval. It should be justified according to:

  • The stability of the construct
  • The age and characteristics of participants
  • The likelihood of learning
  • The purpose of the assessment
  • The expected strength of memory effects
  • The burden of repeated testing

Research on computer-adaptive reading and mathematics assessments has shown that the interval can influence alternate-forms coefficients, so it should not be treated as a trivial administrative detail (Wyse, 2021).

10. Analyse the paired scores

At minimum, report:

  1. Sample size
  2. Means and standard deviations for both forms
  3. Score ranges and distributions
  4. Scatterplot of Form A against Form B
  5. Parallel-forms coefficient
  6. Confidence interval
  7. Mean paired difference
  8. Variability of the paired differences
  9. Relevant item statistics
  10. Order or sequence effects, when applicable

If interchangeability is central, consider analyses that address absolute agreement or equivalence rather than relying entirely on correlation.

11. Revise and cross-validate

A form-development decision based on the same sample used to select and remove items may look better than it will in a new population.

After revision, evaluate the forms in an independent sample or a new administration whenever practical. Continue monitoring form performance after operational use.

What Makes Two Forms Equivalent?

Two forms need several types of evidence.

Content equivalence

They should represent the same content domain and intended construct.

Difficulty equivalence

Neither form should be consistently easier or harder by an amount that matters for score interpretation.

Precision equivalence

Both forms should measure with similar reliability or information across the score range that matters.

Structural equivalence

For multi-item scales, evidence may be needed that the forms reflect comparable dimensions or factor structures.

Administration equivalence

Instructions, timing, delivery, scoring, and testing conditions should not introduce avoidable differences.

Decision equivalence

When scores are used for classification, selection, or progression, the forms should produce sufficiently comparable decisions near important cut scores.

A single correlation does not establish all five forms of evidence.

How Should Parallel Forms Reliability Be Interpreted?

There is no universal coefficient that automatically means “acceptable.”

A reasonable interpretation depends on:

  • Whether decisions concern groups or individuals
  • The stakes of incorrect decisions
  • The width of the confidence interval
  • The variability of the sample
  • Test length
  • Score distribution
  • Consequences of form differences
  • Availability of other evidence
  • Whether scores will be equated
  • Whether the coefficient was replicated

Broad labels such as “acceptable above .70” or “good above .80” may be useful as rough classroom heuristics, but they should not replace a purpose-specific standard.

For exploratory group research, a moderate coefficient may sometimes be usable if uncertainty and measurement error are acknowledged. For individual diagnosis, certification, placement, or selection, stronger precision is normally required.

Use the confidence interval

A point estimate alone can create false certainty.

For example:

  • (r=.84), 95% CI [.76, .90] provides more reassurance than
  • (r=.84), 95% CI [.55, .95].

The point estimates are identical, but the second study is much less precise.

A confidence interval for Pearson’s (r) is commonly constructed using a Fisher (z) transformation. Statistical software should normally be used rather than calculating it manually.

Why a High Correlation Is Not Enough

Consider a situation in which every participant scores exactly ten points higher on Form B than on Form A.

The Pearson correlation could be 1.00 because participant rankings are perfectly preserved. Nevertheless, the forms are not identical in difficulty. Using their raw scores interchangeably would create a systematic advantage for people who receive Form B.

Researchers should therefore inspect:

  • Mean differences
  • Standard deviations
  • Paired-score differences
  • Difference plots
  • Possible proportional bias
  • Item-level difficulty
  • Classification consistency

For continuous measurements, a Bland–Altman-style difference plot or a concordance statistic may provide useful supplementary information when absolute agreement matters. These methods answer questions that ordinary correlation does not.

Does a non-significant paired t-test prove equivalence?

No.

Failure to find a statistically significant mean difference does not establish that the forms are sufficiently equivalent. A study may simply have too little statistical power to detect a meaningful difference.

When a claim of equivalence is required, researchers should define the largest difference that would be practically acceptable and use a suitable equivalence-testing procedure, such as two one-sided tests, where its assumptions fit the design.

The acceptable margin must be justified before examining the results. It should reflect the scale, measurement precision, score use, and consequences of disagreement.

Parallel Forms Reliability Compared With Other Reliability Types

Reliability typeMain questionTypical designCommon statistic
Parallel formsDo different test versions produce consistent scores?Same participants complete Forms A and BPearson correlation, supplemented by equivalence or agreement analyses
Test–retestAre scores stable over time?Same test administered twiceCorrelation or an appropriate ICC
Split-halfAre two halves of one test internally consistent?One administration; items divided into halvesHalf-test correlation adjusted with Spearman–Brown
Internal consistencyDo items intended to measure the same construct relate coherently?One test administrationAlpha, omega, or model-based reliability
Inter-raterDo different raters score consistently or agree?Same targets evaluated by multiple ratersICC, kappa, weighted kappa, or agreement percentage

Parallel forms versus test–retest reliability

Test–retest reliability uses the same instrument on different occasions. Parallel forms reliability uses different item sets intended to measure the same construct.

Parallel forms can reduce direct recall of answers, but they introduce another possible source of error: differences in item content. Consequently, a parallel-forms coefficient may be lower than a test–retest coefficient even when both studies are conducted carefully.

Parallel forms versus split-half reliability

Split-half reliability divides one existing test into two parts, usually to estimate internal consistency. Each half is shorter than the complete test, so the correlation between halves is normally adjusted with the Spearman–Brown formula.

Parallel forms are intended to function as separate complete versions. They should be independently usable and should match the same blueprint.

Randomly dividing a large item pool may be one step in form construction, but random division alone does not guarantee comparable content, difficulty, or precision.

Parallel forms versus internal consistency

Internal consistency concerns relationships among items within one administration. Parallel forms reliability concerns consistency across different item samples.

A form can have strong internal consistency while differing substantially from another form. Researchers should therefore avoid treating Cronbach’s alpha as evidence that two forms are interchangeable.

Advantages of Parallel Forms Reliability

It reduces direct memory effects

Participants do not encounter exactly the same items twice, reducing the chance that they simply remember previous answers.

It supports test security

Multiple forms help reduce item exposure, copying, and repeated access to the same operational questions.

It evaluates item-sampling consistency

The method investigates whether results depend heavily on one particular set of questions.

It supports repeated assessment

Alternate forms can be useful in intervention research, progress monitoring, education, clinical assessment, and training evaluation.

It can strengthen operational fairness

When different candidates receive different forms, form comparability is essential for defensible score interpretation.

Limitations of Parallel Forms Reliability

Developing two forms is expensive

Researchers need enough high-quality items to produce multiple complete assessments rather than one strong test.

Perfect equivalence is difficult

Small differences in wording, difficulty, context, or cognitive demand may affect performance.

Administration order can bias results

Fatigue, transfer, learning, and reduced anxiety can make the second administration different from the first.

The construct may change

A longer interval allows genuine learning, development, or treatment effects to influence scores.

Correlation can conceal systematic differences

High association does not prove equal difficulty or absolute agreement.

Reliability is sample-dependent

Restricted range, unusual participant characteristics, ceiling effects, or floor effects can alter the coefficient.

Item exposure can still occur

Items that are too similar across forms may provide clues, while overlapping items may artificially increase correspondence.

Two forms may drift over time

Replacing items, changing delivery software, revising curricula, or altering scoring can reduce comparability. Operational forms require continuing review.

Common Mistakes

Mistake 1: Calling reordered items a new parallel form

Changing the sequence of identical items may discourage copying, but it does not test consistency across independent item samples.

Mistake 2: Splitting a test randomly and assuming the halves are equivalent

Random allocation does not ensure equal coverage of subdomains, cognitive levels, or difficulty.

Mistake 3: Reporting only (p<.05)

Statistical significance does not tell readers whether the coefficient is large enough for the intended use.

Mistake 4: Ignoring the confidence interval

A seemingly strong coefficient may be estimated imprecisely in a small sample.

Mistake 5: Administering the forms in one fixed order

This makes form and order effects difficult to separate.

Mistake 6: Treating reliability as validity

Reliable forms can consistently measure the wrong construct. Validity concerns whether evidence supports the intended interpretation and use of scores.

Mistake 7: Treating reliability as equating

Reliability evaluates consistency. Equating places scores from different forms on a common scale. Highly related forms may still require score conversion if they differ in difficulty.

Mistake 8: Using a universal cutoff without justification

The acceptable level should be based on score use, stakes, uncertainty, and consequences.

Parallel Forms Reliability in Modern Research

Parallel forms remain important in:

  • Educational assessment
  • Psychological testing
  • Language proficiency testing
  • Neuropsychological assessment
  • Employee selection
  • Certification and licensing
  • Clinical outcome measurement
  • Progress monitoring
  • Computer-based testing
  • AI benchmark development

In longitudinal or intervention research, alternate forms can reduce direct recall while allowing repeated measurement. Researchers must still separate intervention-related change from form differences.

In large testing programmes, alternate forms are often constructed from calibrated item banks and then statistically equated. Here, a simple Pearson correlation is only one part of a much larger quality-control process.

Classical Test Theory and Item Response Theory

Classical test theory

A CTT analysis may compare:

  • Total-score means
  • Total-score variances
  • Reliability coefficients
  • Item difficulty
  • Item discrimination
  • Form-to-form correlations
  • Score differences

CTT is accessible and practical, but its item statistics depend partly on the sample in which they are estimated.

Item response theory

IRT models item responses in relation to a latent-trait scale. Parallel forms can be assembled to have comparable:

  • Test information functions
  • Measurement precision at selected ability levels
  • Content constraints
  • Item difficulty distributions
  • Exposure constraints
  • Expected score characteristics

IRT is particularly valuable when precision must be similar across a broad ability range rather than merely similar on average.

Automated test assembly

Automated test assembly uses optimisation algorithms to select items from an item bank while satisfying constraints such as:

  • Test length
  • Content coverage
  • Target information
  • Maximum item overlap
  • Difficulty
  • Item exposure
  • Timing
  • Format
  • Fairness requirements

Modern systems can construct several forms simultaneously, which may produce better overall balance than building each new form independently.

Can Artificial Intelligence Create Parallel Test Forms?

AI can assist with:

  • Generating candidate items
  • Rewriting stems
  • Creating distractors
  • Mapping items to a blueprint
  • Checking readability
  • Identifying semantic similarity
  • Detecting possible item duplication
  • Flagging biased or unclear language
  • Producing initial form combinations

AI-generated questions should not be assumed to be valid or parallel merely because they are fluent.

Human subject-matter review and empirical testing remain necessary. Researchers should examine construct relevance, factual accuracy, accessibility, bias, item difficulty, discrimination, factor structure, and score comparability.

A 2024 study compared a human-written university test with a form developed using ChatGPT assistance. It combined expert review, item analysis, internal-consistency estimates, mean comparisons, equivalence testing, and score correlation rather than relying on AI output alone (O, 2024). The study illustrates an emerging workflow, not proof that unsupervised AI can replace psychometric validation.

Recent AI-assisted item-generation frameworks similarly emphasise human-in-the-loop quality control (Lee et al., 2025).

How to Report Parallel Forms Reliability

A report should identify:

  • The construct
  • The population
  • Number of participants
  • Form-development process
  • Item numbers and formats
  • Administration order
  • Interval between forms
  • Scoring procedure
  • Descriptive statistics
  • Coefficient and confidence interval
  • Mean or agreement analysis
  • Item-level evidence
  • Prespecified interpretation criterion
  • Limitations

APA-style reporting example

Two 30-item forms of the statistical-reasoning assessment were administered to 146 undergraduate students. Administration order was counterbalanced, with 73 students completing Form A first and 73 completing Form B first. The forms were administered seven days apart. Scores were strongly correlated, (r(144)=.87), 95% CI [.82, .90]. Form A ((M=21.40, SD=4.32)) and Form B ((M=21.18, SD=4.41)) produced similar score distributions. The mean paired difference was 0.22 points. These findings supported the planned group-level use of the forms, although additional cross-validation was recommended before individual high-stakes decisions.

The numbers above are illustrative and should not be presented as findings from a real study.

Parallel Forms Study Checklist

Before concluding that two forms are sufficiently comparable, ask:

  • Do they measure the same clearly defined construct?
  • Do they follow the same blueprint?
  • Are their items different but comparable?
  • Were both reviewed by qualified experts?
  • Were the items piloted?
  • Was the study sample representative?
  • Were administration conditions standardised?
  • Was order counterbalanced?
  • Was the interval justified?
  • Were means and score variability reported?
  • Was the form correlation reported with a confidence interval?
  • Were individual differences or agreement examined?
  • Were item properties compared?
  • Was the acceptable level defined according to score use?
  • Was the result replicated or cross-validated?
  • Is equating required before raw scores are exchanged?

Conclusion

Parallel forms reliability evaluates whether different versions of an instrument produce sufficiently consistent scores. The usual form-to-form correlation is informative, but it is not complete evidence of interchangeability. A defensible evaluation combines a clear test blueprint, expert review, representative sampling, counterbalanced administration, descriptive statistics, confidence intervals, item analysis, and an assessment of meaningful score differences.

Parallel forms are most useful when repeated use of identical questions would create memory, security, or exposure problems. Their main challenge is not calculating a correlation; it is developing and validating two forms that measure the same construct with comparable difficulty and precision.

About the author

Muhammad Hassan

Muhammad Hassan writes about research design, academic methods and data-analysis concepts for ResearchMethod.net. His work focuses on presenting methodological topics in clear language for students and early-career researchers. Articles are developed from recognized methodological literature and official software documentation.