Split-half reliability is an estimate of internal consistency obtained by dividing the items or trials in a measurement instrument into two comparable halves, calculating each participant’s score on both halves, and correlating those scores. Because each half is shorter than the full instrument, the correlation is usually corrected with the Spearman–Brown formula.

Introduction
Researchers use questionnaires, achievement tests, psychological scales, and cognitive tasks to convert an unobservable characteristic into a numerical score. Before interpreting that score, they must determine whether the instrument produces sufficiently consistent measurements.
Split-half reliability provides one way to evaluate that consistency using data from a single administration. Rather than asking participants to complete the same test twice, the researcher treats two subsets of the instrument as comparable measurements.
This article explains what split-half reliability measures, how to calculate it, how to choose the halves, how to interpret and report the result, and when another reliability method may be more appropriate.
Key Takeaways
- Split-half reliability evaluates consistency between two subsets of the same instrument.
- Researchers split the items or trials, not the participants.
- The raw correlation between the two half scores normally requires a Spearman–Brown correction.
- Results depend partly on how the instrument is divided.
- High reliability does not establish validity or prove that a measure is unidimensional.
- Modern cognitive-task studies may benefit from many stratified random splits rather than one arbitrary split.
What Is Split-Half Reliability?
Split-half reliability is a form of internal consistency reliability. It estimates whether two comparable parts of a test, questionnaire, scale, or task produce similar rankings of participants.
Suppose a researcher has a 20-item academic-motivation scale. The researcher might place the odd-numbered items in one half and the even-numbered items in the other. Each respondent receives two scores:
- A score for the odd-numbered items.
- A score for the even-numbered items.
If respondents who score highly on one half also tend to score highly on the other, the two halves are consistent. A strong relationship supports the dependability of the composite score under the conditions studied.
The American Psychological Association defines the method as an assessment of internal consistency based on dividing a set of items into halves and comparing the resulting scores (APA Dictionary of Psychology, n.d.).
What is being split?
The items, questions, trials, or observations within the instrument are split.
The participant sample is not divided into two groups for a conventional split-half reliability analysis. Splitting participants would answer a different question, such as whether an estimate is stable across subsamples.
What does the method measure?
The method measures whether two parts of an instrument behave like comparable forms of the same measurement.
It does not directly measure:
- Stability across different occasions.
- Agreement between raters.
- Whether the instrument measures the correct construct.
- Whether the scale has only one underlying factor.
- Whether individual items are free from bias.
How Does Split-Half Reliability Work?
The conventional procedure contains six steps:
- Administer the complete instrument once.
- Score all items correctly, including reverse-coded items.
- Divide the items or trials into two defensible halves.
- Calculate a half score for every participant.
- Correlate the two sets of half scores.
- Correct the correlation for test length, usually with the Spearman–Brown formula.
The analysis uses one data-collection occasion, which avoids memory, maturation, and retesting effects that may complicate a test–retest design. However, it assesses consistency among parts of the instrument rather than score stability over time.
Split-Half Reliability Formula
For two equal-length halves, the Spearman–Brown corrected coefficient is:
[
r_{SB}=\frac{2r_h}{1+r_h}
]
where:
- (r_{SB}) is the estimated reliability of the full instrument.
- (r_h) is the correlation between scores on the two halves.
- The value 2 represents the increase from a half-length test to the complete test.
The formula is associated with independent work by Spearman and Brown in 1910 (Brown, 1910; Spearman, 1910).
Why is a correction required?
Longer instruments generally aggregate more information than shorter instruments. When a test is divided, each half contains fewer observations and is ordinarily less reliable than the full test.
The uncorrected correlation therefore describes the relationship between two shorter half-tests. The Spearman–Brown formula estimates the reliability expected for the complete instrument, provided the assumptions behind the correction are reasonable.
Worked Example
Imagine that a researcher develops a 20-item study-engagement questionnaire. Ten items are assigned to Half A and ten to Half B.
After calculating the two half scores for every participant, the researcher obtains:
[
r_h=.68
]
The corrected reliability is:
[
r_{SB}=\frac{2(.68)}{1+.68}
]
[
r_{SB}=\frac{1.36}{1.68}=.81
]
The estimated split-half reliability of the complete questionnaire is therefore approximately .81.
How should this result be described?
A coefficient of .81 indicates reasonably strong consistency between the two selected halves in the sample studied. It does not mean that the questionnaire is “81% valid,” that 81% of its questions are correct, or that it will have exactly the same reliability in every population.
The General Spearman–Brown Prediction Formula
The Spearman–Brown formula can also estimate how reliability may change when an instrument is lengthened or shortened:
[
r_{new}=\frac{kr_{old}}{1+(k-1)r_{old}}
]
where:
- (r_{new}) is the predicted reliability after changing test length.
- (r_{old}) is the current reliability.
- (k) is the new length divided by the old length.
If (k=2), the instrument is doubled in length. If (k=.5), it is reduced to half its length.
Estimating the required test length
The formula can be rearranged:
[
k=\frac{r_{target}(1-r_{old})}{r_{old}(1-r_{target})}
]
Suppose a 20-item test has reliability of .70 and the researcher wants a predicted reliability of .85:
[
k=\frac{.85(1-.70)}{.70(1-.85)}
]
[
k=\frac{.255}{.105}=2.43
]
The instrument would need to be approximately 2.43 times as long, or about 49 items, if the new items were comparable in quality to the existing ones.
This is a planning estimate, not a guarantee. Adding repetitive, poorly written, or multidimensional items may not produce the predicted improvement.
Methods for Splitting an Instrument
There is rarely only one possible division. A 20-item scale can be divided into two sets of ten in thousands of ways. The selected partition can affect the result.
Odd–even split
Odd-numbered items form one half and even-numbered items form the other.
Best suited to: Instruments in which difficulty, topic, or item type is reasonably distributed throughout the test.
Advantage: Often balances gradual fatigue or difficulty changes.
Risk: It may fail when conditions or item formats alternate systematically. For example, all odd trials could represent one experimental condition and all even trials another.
First-half versus second-half split
The first group of items forms Half A and the remaining items form Half B.
Best suited to: Instruments whose two sections were deliberately designed as parallel forms.
Advantage: Easy to calculate and explain.
Risk: The result may reflect order effects, fatigue, practice, learning, or a progression from easy to difficult items.
Random split
Items are assigned randomly to two equally sized groups.
Best suited to: Item sets without a meaningful order.
Advantage: Reduces deliberate selection of a favourable split.
Risk: Different random seeds can produce different coefficients. The randomization procedure and seed should be recorded.
Matched split
Researchers pair items according to difficulty, content, discrimination, or factor loading and place one item from each pair in each half.
Best suited to: Well-developed educational and psychological tests with item-level information.
Advantage: Can create more comparable halves.
Risk: Requires prior psychometric information and may become researcher-dependent.
Stratified split
Trials are divided within important design strata, such as experimental condition, stimulus type, block, or response category.
Best suited to: Cognitive and reaction-time tasks.
Advantage: Ensures both halves contain comparable representations of the task design.
Risk: Requires detailed trial-level data and careful programming.
Repeated random or permutation-based split
The data are divided many times. A correlation is calculated for each partition, and the set of correlations is appropriately aggregated before or alongside a length correction.
Best suited to: Instruments or tasks for which a single partition would be arbitrary.
Advantage: Reduces dependence on one accidental split.
Risk: The method is computationally more demanding and requires defensible decisions about stratification, correlation aggregation, negative values, preprocessing, and confidence intervals.
Which Splitting Method Should You Use?
There is no universally best split. The choice should follow the design of the instrument.
| Instrument characteristic | Usually appropriate starting point | Main issue to check |
|---|---|---|
| Items are mixed throughout the scale | Odd–even or matched split | Alternating content or methods |
| Items become progressively harder | Matched or stratified split | Difficulty imbalance |
| The test has two designed parallel sections | First–second split | Whether the sections are genuinely equivalent |
| Items have no order | Random or repeated random splits | Variation across random partitions |
| Trials contain multiple conditions | Stratified random splits | Equal representation of conditions and stimuli |
| Scores use nonlinear algorithms | Repeated split procedure that reproduces the complete scoring algorithm | Whether each half is scored identically |
| The scale contains multiple subscales | Analyze each defensible subscale separately | Whether a total score is theoretically meaningful |
A strong analysis explains why the chosen halves are substantively comparable. “The software divided the items automatically” is not an adequate methodological justification.
Assumptions and Requirements
1. The composite score must be meaningful
The items should contribute to a defensible total or subscale score. If items represent unrelated constructs, a single split-half coefficient may have little meaning.
2. The halves should be comparable
The two halves should have similar:
- Construct coverage.
- Number of items or trials.
- Difficulty.
- Score variance.
- Response format.
- Scoring procedures.
- Exposure to fatigue, practice, and order effects.
The classical Spearman–Brown interpretation is strongest when the halves behave like parallel or sufficiently equivalent measurements.
3. Items must be coded correctly
Reverse-worded items must be recoded before half scores are calculated. A forgotten reversal may create low or negative correlations.
4. Each participant must contribute comparable scores
Researchers should specify how missing values are handled. Pairwise deletion, listwise deletion, prorating, and imputation can produce different samples or scores.
5. The correlation must fit the scores
Pearson’s correlation is commonly used for approximately continuous half-test totals. Rank-based or latent-variable approaches may be preferable in specialised circumstances, but changing the correlation does not solve poor instrument design.
6. The sample must represent the intended application
Reliability is a property of scores in a sample and context, not an immutable property of a questionnaire. Restricted score variation can reduce the observed correlation, while a highly heterogeneous sample can produce a different estimate.
7. Dimensionality should be examined separately
A high split-half coefficient does not prove that the scale is unidimensional. Items measuring two strongly correlated dimensions may still yield a high coefficient.
Factor analysis, substantive theory, and other forms of validity evidence are needed to justify the score structure.
How to Interpret Split-Half Reliability
Reliability coefficients commonly range from 0 to 1, although negative values can occur when the halves are inversely related.
A larger positive coefficient generally indicates stronger consistency. Interpretation must nevertheless consider the score’s purpose.
| Approximate coefficient | Possible interpretation | Required caution |
|---|---|---|
| Below .50 | Weak consistency | Check coding, split quality, dimensionality, score variance, and item quality |
| .50–.69 | Limited or moderate consistency | May be inadequate for many applications |
| .70–.79 | Often considered usable for exploratory group research | Not a universal pass mark |
| .80–.89 | Stronger consistency for many research uses | Still examine assumptions and confidence intervals |
| .90 or higher | Often desired for consequential individual decisions | Extremely high values may reflect item redundancy |
These ranges are conventions rather than laws. A coefficient should be judged alongside:
- The intended decision.
- The cost of measurement error.
- Instrument length.
- Construct breadth.
- Sample size and composition.
- Confidence interval.
- Alternative splits.
- Other reliability evidence.
- Validity evidence.
Report confidence intervals
A sample reliability coefficient is an estimate. It may change in another sample.
Where software permits, report a confidence interval or use an appropriate resampling procedure. A coefficient of .80 with a narrow interval is more informative than the same point estimate with substantial uncertainty.
What Does a Low Coefficient Mean?
A low result may indicate:
- The two halves measure different constructs.
- The halves differ in difficulty or content.
- Reverse-coded items were scored incorrectly.
- Some items are ambiguous or poorly discriminating.
- The score has little variation in the sample.
- Too few items or trials were included.
- Missing data were handled inconsistently.
- Fatigue, learning, or order effects affected one half.
- A nonlinear scoring algorithm was not reproduced within each half.
- The selected partition was unrepresentative.
A low coefficient should initiate a diagnostic investigation. It should not automatically trigger deletion of whichever items produce the largest numerical increase. Item retention should also be guided by theory, content coverage, fairness, and validity.
What Does a Negative Split-Half Coefficient Mean?
A negative coefficient means that participants who score highly on one half tend to score lower on the other.
Common causes include:
- Unreversed negatively worded items.
- Incompatible scoring directions.
- Two halves measuring opposing constructs.
- Severe multidimensionality.
- Very small samples.
- Restricted or unusual score distributions.
- A poorly chosen split.
- Data-entry or preprocessing errors.
The ordinary Spearman–Brown formula can behave problematically with negative correlations. Researchers should diagnose the cause rather than presenting a mechanically corrected value as meaningful.
Split-Half Reliability Versus Other Reliability Methods
| Method | Main question | Data required | Main strength | Main limitation |
|---|---|---|---|---|
| Split-half reliability | Are two parts of one instrument consistent? | One administration | Efficient and transparent | Depends on how the instrument is split |
| Cronbach’s alpha | How strongly do items covary as a set under its model? | One administration | Familiar and widely available | Often misinterpreted; model assumptions may be inappropriate |
| McDonald’s omega | How reliably does a factor-model-based composite measure its common factors? | One administration plus a factor model | Allows unequal factor loadings in common applications | Requires a defensible factor model |
| KR-20 | How internally consistent are dichotomously scored items? | One administration | Designed for binary items | Shares limitations of covariance-based internal-consistency estimates |
| Test–retest reliability | Are scores stable across occasions? | Two or more administrations | Assesses temporal stability | Change, memory, learning, and interval effects can influence results |
| Parallel-forms reliability | Do two separately constructed forms produce comparable scores? | Two forms | Directly evaluates form equivalence | Expensive and difficult to construct |
| Inter-rater reliability | Do raters agree? | Multiple ratings | Essential for observational measures | Does not assess item consistency |
Split-half reliability versus Cronbach’s alpha
Cronbach’s alpha is related historically and mathematically to split-half methods. Cronbach described alpha in relation to reliability coefficients generated by possible test splits (Cronbach, 1951).
The practical distinction is that a conventional split-half analysis uses one selected partition, whereas alpha uses information from the complete item covariance structure. However, alpha is not automatically superior. It can be inappropriate or misleading when its underlying model does not represent the data.
Modern psychometric guidance often recommends examining omega or other model-based reliability estimates alongside alpha rather than reporting alpha mechanically (Dunn et al., 2014; Sijtsma, 2009).
Split-half reliability versus validity
Reliability concerns consistency or measurement precision. Validity concerns whether evidence supports the intended interpretation and use of a score.
A measure can be highly reliable but consistently measure the wrong construct. Split-half reliability alone therefore cannot validate a questionnaire or test.
Advantages of Split-Half Reliability
Requires only one administration
Participants complete the instrument once, reducing cost and avoiding some retesting effects.
Easy to understand
The central logic—compare two versions of the same measurement—is intuitive.
Flexible
The method can be applied to questionnaires, achievement tests, cognitive tasks, performance measures, and trial-level data when the scoring structure permits.
Useful for diagnosing design effects
Comparing several defensible partitions may reveal fatigue, practice, alternating-condition effects, content imbalance, or item-order problems.
Suitable for test-length planning
The general Spearman–Brown formula can estimate the potential effect of adding or removing comparable items.
Useful for specialised task data
For some reaction-time and cognitive measures, properly stratified repeated split-half procedures may be more suitable than applying coefficient alpha to arbitrary trial aggregates.
Limitations of Split-Half Reliability
The result depends on the split
Two reasonable partitions may produce different coefficients. Reporting only the most favourable one creates selection bias.
One split uses limited information
A single partition does not represent all possible divisions of the instrument.
Equivalent halves may be difficult to create
Tests containing varied topics, item types, or difficulty levels may not divide naturally.
Internal consistency is not temporal stability
The method cannot show whether scores remain stable across days, months, or settings.
High consistency can coexist with invalidity
A collection of items may consistently measure an unintended construct.
Very short measures remain difficult to evaluate
Each half of a short scale contains little information. A two-item scale can use a Spearman–Brown coefficient based on the inter-item correlation, but the result should be interpreted with the limitations of an extremely narrow item sample.
Conventional procedures may not suit complex scores
Difference scores, adaptive tests, reaction-time indices, and nonlinearly scored tasks require procedures that reproduce the complete scoring algorithm in each half.
When Should Split-Half Reliability Be Used?
It is most defensible when:
- The instrument contains enough items or trials to form comparable halves.
- A total or subscale score has a clear theoretical interpretation.
- Only one administration is available.
- The researcher can justify the split.
- Each half can be scored in the same way.
- Internal consistency is the relevant reliability question.
- Sensitivity across alternative splits can be evaluated.
When Should It Be Avoided?
Avoid relying on a conventional single split when:
- The instrument measures several unrelated constructs.
- The halves cannot be made comparable.
- Each score is based on very few observations.
- The main concern is temporal stability or rater agreement.
- The test is strongly adaptive.
- Items depend on one another in a fixed sequence.
- The scoring procedure cannot be reproduced for half datasets.
- The selected split is confounded with experimental conditions.
- The researcher intends to use reliability as proof of validity.
Split-Half Reliability in Modern Research
Split-half methods remain useful in scale development, educational assessment, behavioural research, neuroscience, and experimental psychology.
Questionnaires and rating scales
For a questionnaire, researchers may compare odd and even items or create matched halves. Reliability should ordinarily be estimated for each score that will be interpreted, including theoretically distinct subscales.
Educational and achievement tests
Matched splitting can help balance topic coverage and difficulty. A simple first–second split may be misleading when a test progresses from easier to harder questions.
Cognitive and reaction-time tasks
Cognitive tasks often contain many trials, exclusions, repeated stimuli, multiple conditions, and nonlinear scores. A conventional alpha based on trial aggregates may not estimate the desired score reliability.
Pronk et al. (2022) showed that split choice can be confounded by time, task design, trial sampling, and nonlinear scoring. More recent simulation and reanalysis work supports carefully stratified permutation-based procedures for many reaction-time tasks (Kahveci et al., 2025).
Online and adaptive measurement
Online platforms make it easier to record trial order, response times, device information, missingness, and randomization seeds. These records improve reproducibility but also create new sources of variation. Researchers should document whether mobile and desktop responses, interruptions, latency filtering, and adaptive routing affect the score.
Permutation-Based Split-Half Reliability
Permutation-based reliability reduces dependence on a single arbitrary partition.
A typical workflow is:
- Identify important strata, such as participant, condition, stimulus type, and block.
- Divide trials randomly within those strata.
- Apply the complete preprocessing and scoring algorithm separately to both halves.
- Correlate participant scores from the two halves.
- Repeat the process many times.
- Aggregate the correlations with an appropriate method.
- Apply a suitable length correction.
- Report the number of splits, seed, stratification rules, exclusions, aggregation method, and uncertainty interval.
This approach should not be described merely as “bootstrapping” unless sampling with replacement was actually used. Random partitioning without replacement and bootstrap resampling are different procedures.
Software for Split-Half Reliability
SPSS
A conventional SPSS workflow is:
- Select Analyze.
- Select Scale.
- Select Reliability Analysis.
- Add the correctly scored items.
- Choose Split-half as the model.
- Verify how SPSS divides the ordered item list.
- Request relevant descriptive and correlation output.
- Interpret the half-test correlation and corrected coefficients.
SPSS may report separate equal- and unequal-length Spearman–Brown coefficients. Researchers should report the coefficient that matches the actual split and explain how the items were ordered.
R
The psych package provides splitHalf() and can examine sampled or possible partitions, Guttman estimates, and the distribution of split-half values.
A minimal conceptual workflow is:
library(psych)
result <- splitHalf(
my_items,
raw = TRUE,
check.keys = FALSE
)
print(result)
Reverse coding should be handled deliberately rather than left to an automatic procedure without inspection.
For trial-level cognitive data, packages such as splithalf, splithalfr, AATtools, and rapidsplithalf support specialised repeated-split workflows. Package documentation and the underlying methodological paper should be consulted because defaults differ.
Excel
In Excel:
- Place one participant on each row and one item in each column.
- Create a column containing the sum or mean for Half A.
- Create another column for Half B.
- Calculate
CORREL(Half_A_range, Half_B_range). - Apply
=(2*r)/(1+r)to the resulting correlation.
The split, reverse coding, missing-data policy, and formulas should be documented.
Python
Python users can:
- Select the item columns for each half.
- Calculate row-level sums or means with
pandas. - Calculate Pearson’s correlation with
scipy.stats.pearsonr. - Apply the Spearman–Brown formula.
- Use repeated stratified randomization for permutation-based analyses.
All random procedures should use a recorded seed.
Artificial Intelligence and Digital Research Tools
Generative AI can help researchers:
- Draft R or Python code.
- Explain unfamiliar software output.
- Generate a data-validation checklist.
- Convert formulas into spreadsheet syntax.
- Improve the clarity of a methods-section draft.
AI should not independently determine which items belong together, whether a scale is unidimensional, or whether a coefficient is adequate for a high-stakes decision.
Researchers must verify all generated code against documentation and a manually checked example. Personally identifiable or sensitive participant data should not be uploaded to an AI service without appropriate authorization, data-protection safeguards, and institutional approval.
Common Mistakes
Splitting participants instead of items
This is not the conventional split-half procedure.
Reporting the raw correlation as full-test reliability
The half-score correlation usually requires a length correction.
Choosing the split that gives the highest value
Selecting the most favourable result without preregistration or disclosure exaggerates reliability.
Ignoring reverse-coded items
Incorrect coding can create low or negative coefficients.
Treating 0.70 as a universal rule
Acceptability depends on the score’s purpose and consequences.
Assuming high reliability proves validity
A consistent measure can still measure the wrong construct.
Combining unrelated subscales
A total coefficient may hide meaningful multidimensionality.
Omitting the split method
Readers cannot evaluate the result unless they know how the halves were formed.
Ignoring uncertainty
A point estimate without a confidence interval can overstate precision.
Using a questionnaire method unchanged for trial-level tasks
Cognitive-task scoring, exclusions, conditions, and repeated stimuli require specialised handling.
How to Report Split-Half Reliability
A clear report should state:
- The instrument and score evaluated.
- The sample and number of usable cases.
- How items or trials were divided.
- How reverse coding and missing data were handled.
- The scoring rule for each half.
- The correlation used.
- The raw half-score correlation.
- The correction formula.
- The corrected coefficient.
- A confidence interval, if available.
- Any sensitivity analysis across alternative splits.
- The software and version.
Example reporting template
Internal consistency of the 20-item Study Engagement Scale was evaluated using an odd–even split. Negatively worded items were reverse-scored before analysis, and half scores were calculated as item sums. The two half scores were positively correlated, (r=.68). Applying the Spearman–Brown correction produced a split-half reliability coefficient of (r_{SB}=.81), indicating reasonably strong consistency between the selected halves in this sample.
Example for repeated random splits
Reliability of the reaction-time score was estimated using 5,000 stratified random splits. Trials were divided within participant, condition, and stimulus category, and the complete exclusion and scoring procedure was repeated for each half. Correlations were aggregated using the prespecified method and corrected for test length. The resulting reliability estimate was [coefficient], 95% CI [lower, upper].
Practical Checklist
Before calculating the coefficient, ask:
- Is the total or subscale score theoretically meaningful?
- Have all reverse-coded items been corrected?
- Are the two halves comparable in content and difficulty?
- Is the split confounded with order, condition, or item format?
- Is each half scored using the same procedure?
- Is the missing-data rule transparent?
- Should several splits be examined?
- Is a confidence interval available?
- Does the study also require temporal or inter-rater reliability?
- Have validity and dimensionality been evaluated separately?
Conclusion
Split-half reliability estimates how consistently two parts of an instrument measure participants during one administration. The method is easy to understand, but a defensible analysis requires more than dividing a test and calculating a correlation. Researchers must justify the split, score both halves consistently, apply an appropriate length correction, examine uncertainty, and avoid treating reliability as evidence of validity.
For straightforward questionnaires, an odd–even or matched split may be adequate. For complex cognitive tasks, stratified repeated splits may provide a more representative estimate. In all cases, the coefficient should be interpreted as evidence about particular scores obtained in a particular sample and context.
