Reliability & Validity

Split-Half Reliability – Methods, Examples and Formulas

Table of Contents

Split-half reliability is an estimate of internal consistency obtained by dividing the items or trials in a measurement instrument into two comparable halves, calculating each participant’s score on both halves, and correlating those scores. Because each half is shorter than the full instrument, the correlation is usually corrected with the Spearman–Brown formula.

Split-Half Reliability

Introduction

Researchers use questionnaires, achievement tests, psychological scales, and cognitive tasks to convert an unobservable characteristic into a numerical score. Before interpreting that score, they must determine whether the instrument produces sufficiently consistent measurements.

Split-half reliability provides one way to evaluate that consistency using data from a single administration. Rather than asking participants to complete the same test twice, the researcher treats two subsets of the instrument as comparable measurements.

This article explains what split-half reliability measures, how to calculate it, how to choose the halves, how to interpret and report the result, and when another reliability method may be more appropriate.

Key Takeaways

  • Split-half reliability evaluates consistency between two subsets of the same instrument.
  • Researchers split the items or trials, not the participants.
  • The raw correlation between the two half scores normally requires a Spearman–Brown correction.
  • Results depend partly on how the instrument is divided.
  • High reliability does not establish validity or prove that a measure is unidimensional.
  • Modern cognitive-task studies may benefit from many stratified random splits rather than one arbitrary split.

What Is Split-Half Reliability?

Split-half reliability is a form of internal consistency reliability. It estimates whether two comparable parts of a test, questionnaire, scale, or task produce similar rankings of participants.

Suppose a researcher has a 20-item academic-motivation scale. The researcher might place the odd-numbered items in one half and the even-numbered items in the other. Each respondent receives two scores:

  • A score for the odd-numbered items.
  • A score for the even-numbered items.

If respondents who score highly on one half also tend to score highly on the other, the two halves are consistent. A strong relationship supports the dependability of the composite score under the conditions studied.

The American Psychological Association defines the method as an assessment of internal consistency based on dividing a set of items into halves and comparing the resulting scores (APA Dictionary of Psychology, n.d.).

What is being split?

The items, questions, trials, or observations within the instrument are split.

The participant sample is not divided into two groups for a conventional split-half reliability analysis. Splitting participants would answer a different question, such as whether an estimate is stable across subsamples.

What does the method measure?

The method measures whether two parts of an instrument behave like comparable forms of the same measurement.

It does not directly measure:

  • Stability across different occasions.
  • Agreement between raters.
  • Whether the instrument measures the correct construct.
  • Whether the scale has only one underlying factor.
  • Whether individual items are free from bias.

How Does Split-Half Reliability Work?

The conventional procedure contains six steps:

  1. Administer the complete instrument once.
  2. Score all items correctly, including reverse-coded items.
  3. Divide the items or trials into two defensible halves.
  4. Calculate a half score for every participant.
  5. Correlate the two sets of half scores.
  6. Correct the correlation for test length, usually with the Spearman–Brown formula.

The analysis uses one data-collection occasion, which avoids memory, maturation, and retesting effects that may complicate a test–retest design. However, it assesses consistency among parts of the instrument rather than score stability over time.

Split-Half Reliability Formula

For two equal-length halves, the Spearman–Brown corrected coefficient is:

[
r_{SB}=\frac{2r_h}{1+r_h}
]

where:

  • (r_{SB}) is the estimated reliability of the full instrument.
  • (r_h) is the correlation between scores on the two halves.
  • The value 2 represents the increase from a half-length test to the complete test.

The formula is associated with independent work by Spearman and Brown in 1910 (Brown, 1910; Spearman, 1910).

Why is a correction required?

Longer instruments generally aggregate more information than shorter instruments. When a test is divided, each half contains fewer observations and is ordinarily less reliable than the full test.

The uncorrected correlation therefore describes the relationship between two shorter half-tests. The Spearman–Brown formula estimates the reliability expected for the complete instrument, provided the assumptions behind the correction are reasonable.

Worked Example

Imagine that a researcher develops a 20-item study-engagement questionnaire. Ten items are assigned to Half A and ten to Half B.

After calculating the two half scores for every participant, the researcher obtains:

[
r_h=.68
]

The corrected reliability is:

[
r_{SB}=\frac{2(.68)}{1+.68}
]

[
r_{SB}=\frac{1.36}{1.68}=.81
]

The estimated split-half reliability of the complete questionnaire is therefore approximately .81.

How should this result be described?

A coefficient of .81 indicates reasonably strong consistency between the two selected halves in the sample studied. It does not mean that the questionnaire is “81% valid,” that 81% of its questions are correct, or that it will have exactly the same reliability in every population.

The General Spearman–Brown Prediction Formula

The Spearman–Brown formula can also estimate how reliability may change when an instrument is lengthened or shortened:

[
r_{new}=\frac{kr_{old}}{1+(k-1)r_{old}}
]

where:

  • (r_{new}) is the predicted reliability after changing test length.
  • (r_{old}) is the current reliability.
  • (k) is the new length divided by the old length.

If (k=2), the instrument is doubled in length. If (k=.5), it is reduced to half its length.

Estimating the required test length

The formula can be rearranged:

[
k=\frac{r_{target}(1-r_{old})}{r_{old}(1-r_{target})}
]

Suppose a 20-item test has reliability of .70 and the researcher wants a predicted reliability of .85:

[
k=\frac{.85(1-.70)}{.70(1-.85)}
]

[
k=\frac{.255}{.105}=2.43
]

The instrument would need to be approximately 2.43 times as long, or about 49 items, if the new items were comparable in quality to the existing ones.

This is a planning estimate, not a guarantee. Adding repetitive, poorly written, or multidimensional items may not produce the predicted improvement.

Methods for Splitting an Instrument

There is rarely only one possible division. A 20-item scale can be divided into two sets of ten in thousands of ways. The selected partition can affect the result.

Odd–even split

Odd-numbered items form one half and even-numbered items form the other.

Best suited to: Instruments in which difficulty, topic, or item type is reasonably distributed throughout the test.

Advantage: Often balances gradual fatigue or difficulty changes.

Risk: It may fail when conditions or item formats alternate systematically. For example, all odd trials could represent one experimental condition and all even trials another.

First-half versus second-half split

The first group of items forms Half A and the remaining items form Half B.

Best suited to: Instruments whose two sections were deliberately designed as parallel forms.

Advantage: Easy to calculate and explain.

Risk: The result may reflect order effects, fatigue, practice, learning, or a progression from easy to difficult items.

Random split

Items are assigned randomly to two equally sized groups.

Best suited to: Item sets without a meaningful order.

Advantage: Reduces deliberate selection of a favourable split.

Risk: Different random seeds can produce different coefficients. The randomization procedure and seed should be recorded.

Matched split

Researchers pair items according to difficulty, content, discrimination, or factor loading and place one item from each pair in each half.

Best suited to: Well-developed educational and psychological tests with item-level information.

Advantage: Can create more comparable halves.

Risk: Requires prior psychometric information and may become researcher-dependent.

Stratified split

Trials are divided within important design strata, such as experimental condition, stimulus type, block, or response category.

Best suited to: Cognitive and reaction-time tasks.

Advantage: Ensures both halves contain comparable representations of the task design.

Risk: Requires detailed trial-level data and careful programming.

Repeated random or permutation-based split

The data are divided many times. A correlation is calculated for each partition, and the set of correlations is appropriately aggregated before or alongside a length correction.

Best suited to: Instruments or tasks for which a single partition would be arbitrary.

Advantage: Reduces dependence on one accidental split.

Risk: The method is computationally more demanding and requires defensible decisions about stratification, correlation aggregation, negative values, preprocessing, and confidence intervals.

Which Splitting Method Should You Use?

There is no universally best split. The choice should follow the design of the instrument.

Instrument characteristicUsually appropriate starting pointMain issue to check
Items are mixed throughout the scaleOdd–even or matched splitAlternating content or methods
Items become progressively harderMatched or stratified splitDifficulty imbalance
The test has two designed parallel sectionsFirst–second splitWhether the sections are genuinely equivalent
Items have no orderRandom or repeated random splitsVariation across random partitions
Trials contain multiple conditionsStratified random splitsEqual representation of conditions and stimuli
Scores use nonlinear algorithmsRepeated split procedure that reproduces the complete scoring algorithmWhether each half is scored identically
The scale contains multiple subscalesAnalyze each defensible subscale separatelyWhether a total score is theoretically meaningful

A strong analysis explains why the chosen halves are substantively comparable. “The software divided the items automatically” is not an adequate methodological justification.

Assumptions and Requirements

1. The composite score must be meaningful

The items should contribute to a defensible total or subscale score. If items represent unrelated constructs, a single split-half coefficient may have little meaning.

2. The halves should be comparable

The two halves should have similar:

  • Construct coverage.
  • Number of items or trials.
  • Difficulty.
  • Score variance.
  • Response format.
  • Scoring procedures.
  • Exposure to fatigue, practice, and order effects.

The classical Spearman–Brown interpretation is strongest when the halves behave like parallel or sufficiently equivalent measurements.

3. Items must be coded correctly

Reverse-worded items must be recoded before half scores are calculated. A forgotten reversal may create low or negative correlations.

4. Each participant must contribute comparable scores

Researchers should specify how missing values are handled. Pairwise deletion, listwise deletion, prorating, and imputation can produce different samples or scores.

5. The correlation must fit the scores

Pearson’s correlation is commonly used for approximately continuous half-test totals. Rank-based or latent-variable approaches may be preferable in specialised circumstances, but changing the correlation does not solve poor instrument design.

6. The sample must represent the intended application

Reliability is a property of scores in a sample and context, not an immutable property of a questionnaire. Restricted score variation can reduce the observed correlation, while a highly heterogeneous sample can produce a different estimate.

7. Dimensionality should be examined separately

A high split-half coefficient does not prove that the scale is unidimensional. Items measuring two strongly correlated dimensions may still yield a high coefficient.

Factor analysis, substantive theory, and other forms of validity evidence are needed to justify the score structure.

How to Interpret Split-Half Reliability

Reliability coefficients commonly range from 0 to 1, although negative values can occur when the halves are inversely related.

A larger positive coefficient generally indicates stronger consistency. Interpretation must nevertheless consider the score’s purpose.

Approximate coefficientPossible interpretationRequired caution
Below .50Weak consistencyCheck coding, split quality, dimensionality, score variance, and item quality
.50–.69Limited or moderate consistencyMay be inadequate for many applications
.70–.79Often considered usable for exploratory group researchNot a universal pass mark
.80–.89Stronger consistency for many research usesStill examine assumptions and confidence intervals
.90 or higherOften desired for consequential individual decisionsExtremely high values may reflect item redundancy

These ranges are conventions rather than laws. A coefficient should be judged alongside:

  • The intended decision.
  • The cost of measurement error.
  • Instrument length.
  • Construct breadth.
  • Sample size and composition.
  • Confidence interval.
  • Alternative splits.
  • Other reliability evidence.
  • Validity evidence.

Report confidence intervals

A sample reliability coefficient is an estimate. It may change in another sample.

Where software permits, report a confidence interval or use an appropriate resampling procedure. A coefficient of .80 with a narrow interval is more informative than the same point estimate with substantial uncertainty.

What Does a Low Coefficient Mean?

A low result may indicate:

  • The two halves measure different constructs.
  • The halves differ in difficulty or content.
  • Reverse-coded items were scored incorrectly.
  • Some items are ambiguous or poorly discriminating.
  • The score has little variation in the sample.
  • Too few items or trials were included.
  • Missing data were handled inconsistently.
  • Fatigue, learning, or order effects affected one half.
  • A nonlinear scoring algorithm was not reproduced within each half.
  • The selected partition was unrepresentative.

A low coefficient should initiate a diagnostic investigation. It should not automatically trigger deletion of whichever items produce the largest numerical increase. Item retention should also be guided by theory, content coverage, fairness, and validity.

What Does a Negative Split-Half Coefficient Mean?

A negative coefficient means that participants who score highly on one half tend to score lower on the other.

Common causes include:

  • Unreversed negatively worded items.
  • Incompatible scoring directions.
  • Two halves measuring opposing constructs.
  • Severe multidimensionality.
  • Very small samples.
  • Restricted or unusual score distributions.
  • A poorly chosen split.
  • Data-entry or preprocessing errors.

The ordinary Spearman–Brown formula can behave problematically with negative correlations. Researchers should diagnose the cause rather than presenting a mechanically corrected value as meaningful.

Split-Half Reliability Versus Other Reliability Methods

MethodMain questionData requiredMain strengthMain limitation
Split-half reliabilityAre two parts of one instrument consistent?One administrationEfficient and transparentDepends on how the instrument is split
Cronbach’s alphaHow strongly do items covary as a set under its model?One administrationFamiliar and widely availableOften misinterpreted; model assumptions may be inappropriate
McDonald’s omegaHow reliably does a factor-model-based composite measure its common factors?One administration plus a factor modelAllows unequal factor loadings in common applicationsRequires a defensible factor model
KR-20How internally consistent are dichotomously scored items?One administrationDesigned for binary itemsShares limitations of covariance-based internal-consistency estimates
Test–retest reliabilityAre scores stable across occasions?Two or more administrationsAssesses temporal stabilityChange, memory, learning, and interval effects can influence results
Parallel-forms reliabilityDo two separately constructed forms produce comparable scores?Two formsDirectly evaluates form equivalenceExpensive and difficult to construct
Inter-rater reliabilityDo raters agree?Multiple ratingsEssential for observational measuresDoes not assess item consistency

Split-half reliability versus Cronbach’s alpha

Cronbach’s alpha is related historically and mathematically to split-half methods. Cronbach described alpha in relation to reliability coefficients generated by possible test splits (Cronbach, 1951).

The practical distinction is that a conventional split-half analysis uses one selected partition, whereas alpha uses information from the complete item covariance structure. However, alpha is not automatically superior. It can be inappropriate or misleading when its underlying model does not represent the data.

Modern psychometric guidance often recommends examining omega or other model-based reliability estimates alongside alpha rather than reporting alpha mechanically (Dunn et al., 2014; Sijtsma, 2009).

Split-half reliability versus validity

Reliability concerns consistency or measurement precision. Validity concerns whether evidence supports the intended interpretation and use of a score.

A measure can be highly reliable but consistently measure the wrong construct. Split-half reliability alone therefore cannot validate a questionnaire or test.

Advantages of Split-Half Reliability

Requires only one administration

Participants complete the instrument once, reducing cost and avoiding some retesting effects.

Easy to understand

The central logic—compare two versions of the same measurement—is intuitive.

Flexible

The method can be applied to questionnaires, achievement tests, cognitive tasks, performance measures, and trial-level data when the scoring structure permits.

Useful for diagnosing design effects

Comparing several defensible partitions may reveal fatigue, practice, alternating-condition effects, content imbalance, or item-order problems.

Suitable for test-length planning

The general Spearman–Brown formula can estimate the potential effect of adding or removing comparable items.

Useful for specialised task data

For some reaction-time and cognitive measures, properly stratified repeated split-half procedures may be more suitable than applying coefficient alpha to arbitrary trial aggregates.

Limitations of Split-Half Reliability

The result depends on the split

Two reasonable partitions may produce different coefficients. Reporting only the most favourable one creates selection bias.

One split uses limited information

A single partition does not represent all possible divisions of the instrument.

Equivalent halves may be difficult to create

Tests containing varied topics, item types, or difficulty levels may not divide naturally.

Internal consistency is not temporal stability

The method cannot show whether scores remain stable across days, months, or settings.

High consistency can coexist with invalidity

A collection of items may consistently measure an unintended construct.

Very short measures remain difficult to evaluate

Each half of a short scale contains little information. A two-item scale can use a Spearman–Brown coefficient based on the inter-item correlation, but the result should be interpreted with the limitations of an extremely narrow item sample.

Conventional procedures may not suit complex scores

Difference scores, adaptive tests, reaction-time indices, and nonlinearly scored tasks require procedures that reproduce the complete scoring algorithm in each half.

When Should Split-Half Reliability Be Used?

It is most defensible when:

  • The instrument contains enough items or trials to form comparable halves.
  • A total or subscale score has a clear theoretical interpretation.
  • Only one administration is available.
  • The researcher can justify the split.
  • Each half can be scored in the same way.
  • Internal consistency is the relevant reliability question.
  • Sensitivity across alternative splits can be evaluated.

When Should It Be Avoided?

Avoid relying on a conventional single split when:

  • The instrument measures several unrelated constructs.
  • The halves cannot be made comparable.
  • Each score is based on very few observations.
  • The main concern is temporal stability or rater agreement.
  • The test is strongly adaptive.
  • Items depend on one another in a fixed sequence.
  • The scoring procedure cannot be reproduced for half datasets.
  • The selected split is confounded with experimental conditions.
  • The researcher intends to use reliability as proof of validity.

Split-Half Reliability in Modern Research

Split-half methods remain useful in scale development, educational assessment, behavioural research, neuroscience, and experimental psychology.

Questionnaires and rating scales

For a questionnaire, researchers may compare odd and even items or create matched halves. Reliability should ordinarily be estimated for each score that will be interpreted, including theoretically distinct subscales.

Educational and achievement tests

Matched splitting can help balance topic coverage and difficulty. A simple first–second split may be misleading when a test progresses from easier to harder questions.

Cognitive and reaction-time tasks

Cognitive tasks often contain many trials, exclusions, repeated stimuli, multiple conditions, and nonlinear scores. A conventional alpha based on trial aggregates may not estimate the desired score reliability.

Pronk et al. (2022) showed that split choice can be confounded by time, task design, trial sampling, and nonlinear scoring. More recent simulation and reanalysis work supports carefully stratified permutation-based procedures for many reaction-time tasks (Kahveci et al., 2025).

Online and adaptive measurement

Online platforms make it easier to record trial order, response times, device information, missingness, and randomization seeds. These records improve reproducibility but also create new sources of variation. Researchers should document whether mobile and desktop responses, interruptions, latency filtering, and adaptive routing affect the score.

Permutation-Based Split-Half Reliability

Permutation-based reliability reduces dependence on a single arbitrary partition.

A typical workflow is:

  1. Identify important strata, such as participant, condition, stimulus type, and block.
  2. Divide trials randomly within those strata.
  3. Apply the complete preprocessing and scoring algorithm separately to both halves.
  4. Correlate participant scores from the two halves.
  5. Repeat the process many times.
  6. Aggregate the correlations with an appropriate method.
  7. Apply a suitable length correction.
  8. Report the number of splits, seed, stratification rules, exclusions, aggregation method, and uncertainty interval.

This approach should not be described merely as “bootstrapping” unless sampling with replacement was actually used. Random partitioning without replacement and bootstrap resampling are different procedures.

Software for Split-Half Reliability

SPSS

A conventional SPSS workflow is:

  1. Select Analyze.
  2. Select Scale.
  3. Select Reliability Analysis.
  4. Add the correctly scored items.
  5. Choose Split-half as the model.
  6. Verify how SPSS divides the ordered item list.
  7. Request relevant descriptive and correlation output.
  8. Interpret the half-test correlation and corrected coefficients.

SPSS may report separate equal- and unequal-length Spearman–Brown coefficients. Researchers should report the coefficient that matches the actual split and explain how the items were ordered.

R

The psych package provides splitHalf() and can examine sampled or possible partitions, Guttman estimates, and the distribution of split-half values.

A minimal conceptual workflow is:

library(psych)

result <- splitHalf(
  my_items,
  raw = TRUE,
  check.keys = FALSE
)

print(result)

Reverse coding should be handled deliberately rather than left to an automatic procedure without inspection.

For trial-level cognitive data, packages such as splithalf, splithalfr, AATtools, and rapidsplithalf support specialised repeated-split workflows. Package documentation and the underlying methodological paper should be consulted because defaults differ.

Excel

In Excel:

  1. Place one participant on each row and one item in each column.
  2. Create a column containing the sum or mean for Half A.
  3. Create another column for Half B.
  4. Calculate CORREL(Half_A_range, Half_B_range).
  5. Apply =(2*r)/(1+r) to the resulting correlation.

The split, reverse coding, missing-data policy, and formulas should be documented.

Python

Python users can:

  1. Select the item columns for each half.
  2. Calculate row-level sums or means with pandas.
  3. Calculate Pearson’s correlation with scipy.stats.pearsonr.
  4. Apply the Spearman–Brown formula.
  5. Use repeated stratified randomization for permutation-based analyses.

All random procedures should use a recorded seed.

Artificial Intelligence and Digital Research Tools

Generative AI can help researchers:

  • Draft R or Python code.
  • Explain unfamiliar software output.
  • Generate a data-validation checklist.
  • Convert formulas into spreadsheet syntax.
  • Improve the clarity of a methods-section draft.

AI should not independently determine which items belong together, whether a scale is unidimensional, or whether a coefficient is adequate for a high-stakes decision.

Researchers must verify all generated code against documentation and a manually checked example. Personally identifiable or sensitive participant data should not be uploaded to an AI service without appropriate authorization, data-protection safeguards, and institutional approval.

Common Mistakes

Splitting participants instead of items

This is not the conventional split-half procedure.

Reporting the raw correlation as full-test reliability

The half-score correlation usually requires a length correction.

Choosing the split that gives the highest value

Selecting the most favourable result without preregistration or disclosure exaggerates reliability.

Ignoring reverse-coded items

Incorrect coding can create low or negative coefficients.

Treating 0.70 as a universal rule

Acceptability depends on the score’s purpose and consequences.

Assuming high reliability proves validity

A consistent measure can still measure the wrong construct.

Combining unrelated subscales

A total coefficient may hide meaningful multidimensionality.

Omitting the split method

Readers cannot evaluate the result unless they know how the halves were formed.

Ignoring uncertainty

A point estimate without a confidence interval can overstate precision.

Using a questionnaire method unchanged for trial-level tasks

Cognitive-task scoring, exclusions, conditions, and repeated stimuli require specialised handling.

How to Report Split-Half Reliability

A clear report should state:

  • The instrument and score evaluated.
  • The sample and number of usable cases.
  • How items or trials were divided.
  • How reverse coding and missing data were handled.
  • The scoring rule for each half.
  • The correlation used.
  • The raw half-score correlation.
  • The correction formula.
  • The corrected coefficient.
  • A confidence interval, if available.
  • Any sensitivity analysis across alternative splits.
  • The software and version.

Example reporting template

Internal consistency of the 20-item Study Engagement Scale was evaluated using an odd–even split. Negatively worded items were reverse-scored before analysis, and half scores were calculated as item sums. The two half scores were positively correlated, (r=.68). Applying the Spearman–Brown correction produced a split-half reliability coefficient of (r_{SB}=.81), indicating reasonably strong consistency between the selected halves in this sample.

Example for repeated random splits

Reliability of the reaction-time score was estimated using 5,000 stratified random splits. Trials were divided within participant, condition, and stimulus category, and the complete exclusion and scoring procedure was repeated for each half. Correlations were aggregated using the prespecified method and corrected for test length. The resulting reliability estimate was [coefficient], 95% CI [lower, upper].

Practical Checklist

Before calculating the coefficient, ask:

  1. Is the total or subscale score theoretically meaningful?
  2. Have all reverse-coded items been corrected?
  3. Are the two halves comparable in content and difficulty?
  4. Is the split confounded with order, condition, or item format?
  5. Is each half scored using the same procedure?
  6. Is the missing-data rule transparent?
  7. Should several splits be examined?
  8. Is a confidence interval available?
  9. Does the study also require temporal or inter-rater reliability?
  10. Have validity and dimensionality been evaluated separately?

Conclusion

Split-half reliability estimates how consistently two parts of an instrument measure participants during one administration. The method is easy to understand, but a defensible analysis requires more than dividing a test and calculating a correlation. Researchers must justify the split, score both halves consistently, apply an appropriate length correction, examine uncertainty, and avoid treating reliability as evidence of validity.

For straightforward questionnaires, an odd–even or matched split may be adequate. For complex cognitive tasks, stratified repeated splits may provide a more representative estimate. In all cases, the coefficient should be interpreted as evidence about particular scores obtained in a particular sample and context.

About the author

Muhammad Hassan

Muhammad Hassan writes about research design, academic methods and data-analysis concepts for ResearchMethod.net. His work focuses on presenting methodological topics in clear language for students and early-career researchers. Articles are developed from recognized methodological literature and official software documentation.