Simple random sampling is a probability-sampling method in which every possible sample of a specified size has an equal chance of being selected from a defined population. Researchers normally create a complete sampling frame, assign each unit a unique identifier, and use a random process to select the required number of distinct units.

Introduction
Simple random sampling is one of the foundational methods of probability sampling. It is widely taught because its logic is transparent: selection is determined by chance rather than by the researcher’s preferences, convenience, or judgment.
The method is simple in principle but not always simple to implement. A valid simple random sample requires a clearly defined population, an appropriate and sufficiently complete sampling frame, a predetermined sample size, and a genuinely random selection procedure. Researchers must also deal appropriately with duplicate records, ineligible units, nonresponse, data protection, and departures from the planned design.
This guide explains the plain-language and formal definitions of simple random sampling, its formulas, selection steps, advantages, limitations, software options, reporting requirements, and relationship to other sampling methods.
Key Takeaways
- Simple random sampling is a probability-sampling method, not convenience selection.
- In a fixed-size sample without replacement, every possible set of (n) population units must have the same probability of selection.
- Each population unit has an inclusion probability of (n/N).
- Random selection reduces researcher-controlled selection bias but does not guarantee that one sample will perfectly represent every population characteristic.
- A complete, accurate sampling frame and an appropriate nonresponse plan are essential.
- Random sampling determines who enters a study; random assignment determines which treatment or condition participants receive.
What Is Simple Random Sampling?
Simple random sampling, abbreviated SRS, is a method for selecting (n) units from a population of (N) units so that all possible samples of that size are equally likely to be chosen. It is usually conducted without replacement, meaning that a selected person or unit cannot appear twice in the same sample.
Statistics Canada describes SRS as a method in which each sampling unit has an equal chance of inclusion and each possible sample has an equal chance of selection. Penn State’s sampling-theory materials similarly define simple random sampling without replacement through equal probabilities for every possible sample of (n) units (Statistics Canada, 2021; Penn State Department of Statistics, n.d.).
Plain-language definition
In simple language, SRS means:
- Identify everyone or everything that belongs to the population.
- Place those units on a list.
- Give each unit a unique ID.
- Use a random method to select the required number.
- Include the units whose IDs were selected.
No eligible unit is intentionally favoured or excluded.
Formal statistical definition
Suppose a finite population contains (N) distinct units and a researcher wants a sample containing (n) different units.
The number of possible samples is:
[
\binom{N}{n}=\frac{N!}{n!(N-n)!}
]
Under simple random sampling without replacement, every possible sample (s) of size (n) has probability:
[
P(S=s)=\frac{1}{\binom{N}{n}}
]
Consequently, every population unit has the same first-order inclusion probability:
[
\pi_i=P(i\in S)=\frac{n}{N}
]
This formal definition is more precise than saying only that “everyone has an equal chance.” Equal individual inclusion probabilities are necessary, but the defining feature of a fixed-size SRS is equal probability for every possible sample of that size.
Does “equal chance” mean “independent chance”?
Not when sampling is performed without replacement.
Suppose 100 students are eligible and 10 will be selected. Before selection, every student has a (10/100), or 10%, inclusion probability. Once one student has been selected, only nine places remain among 99 students.
The students still have equal overall inclusion probabilities, but the selection indicators are not statistically independent. The phrase “equal and independent chance” should therefore be used cautiously.
Essential components
A defensible SRS normally contains six components:
| Component | Meaning |
|---|---|
| Target population | The full group about which the researcher wants to draw conclusions |
| Sampling unit | The person, household, school, record, business, object, or other unit that may be selected |
| Sampling frame | The operational list from which selection is made |
| Sample size | The predetermined number of units to select |
| Random mechanism | A lottery, random-number table, or documented computer algorithm |
| Selection protocol | Written rules covering eligibility, duplicates, nonresponse, and replacements |
How Does Simple Random Sampling Work?
The researcher first converts the target population into an operational list called a sampling frame. Each eligible unit receives one unique identifier. A random procedure then selects (n) distinct identifiers without replacement. Data are sought from the selected units, and the researcher documents nonresponse and any deviations from the original design.
Randomness must operate at the actual point of selection. Randomly arranging an already convenient or self-selected group does not turn it into a probability sample of the wider population.
For example, randomly selecting 50 people from a researcher’s contact list is an SRS only of the eligible people on that contact list. It is not an SRS of all adults in a city, all university students, or all internet users.
Steps in Simple Random Sampling
Step 1: Define the target population
State exactly who or what the study is intended to describe.
A useful population definition normally specifies:
- Unit type.
- Geographic boundary.
- Time period.
- Inclusion criteria.
- Exclusion criteria.
Weak definition: University students.
Better definition: All undergraduate students registered at University X during the Spring 2026 semester, excluding exchange students whose enrolment lasts fewer than eight weeks.
The findings can be generalised only to a population adequately represented by the frame and selection process.
Step 2: Identify the sampling unit
The sampling unit is the entity selected during sampling.
It may be:
- A person.
- A household.
- A school.
- A class.
- A patient record.
- A business.
- A published article.
- A transaction.
- A manufactured item.
The sampling unit is not always the same as the unit of analysis. A household may be selected first, for example, while individual household members are later analysed.
Step 3: Construct or obtain the sampling frame
The sampling frame should contain all eligible units once and only once, as far as reasonably possible.
Possible frames include:
- Student enrolment records.
- Employee databases.
- Patient registers.
- Membership lists.
- Customer databases.
- Address registers.
- School lists.
- Administrative records.
- Product inventories.
Before sampling, check the frame for:
- Missing eligible units.
- Ineligible records.
- Duplicate entries.
- Outdated contact information.
- Incorrect classifications.
- Units with multiple chances of selection.
A random draw cannot repair a defective frame. If some population members are absent, they have a zero probability of selection from that frame.
Step 4: Determine the sample size
Choose the sample size before generating random numbers.
The decision may depend on:
- The desired margin of error.
- Confidence level.
- Estimated population variability.
- Minimum detectable effect.
- Statistical power.
- Planned statistical model.
- Number of subgroups to be analysed.
- Population size.
- Expected response rate.
- Budget and time.
A sample should not be chosen merely because it represents a familiar percentage, such as 10% of the population. A suitable percentage varies substantially with population size and research objectives.
Step 5: Assign one unique identifier to each unit
Assign IDs such as:
[
1,2,3,\ldots,N
]
Use stable identifiers rather than names in the sampling file whenever possible. This reduces data exposure and prevents confusion between people who share a name.
Check that:
- Each eligible unit has exactly one ID.
- IDs contain no accidental gaps that will be treated as population members.
- The final value of (N) matches the number of eligible records.
Step 6: Select (n) distinct IDs randomly
Use a procedure that gives every possible sample of size (n) the same probability.
Suitable methods include:
- Drawing labelled slips from a thoroughly mixed container.
- Using a random-number table.
- Sorting spreadsheet rows by independently generated uniform random numbers.
- Using statistical software to sample without replacement.
- Using a documented programming-language random-sampling function.
For reproducibility, record:
- Software and version.
- Function or command.
- Sample size.
- Whether replacement was allowed.
- Random seed, when supported.
- Date of selection.
- Person responsible for running the selection.
Step 7: Verify the selected sample
Before contacting participants, check that:
- Exactly (n) distinct units were selected.
- No unselected records were accidentally included.
- Every selected record came from the frozen final frame.
- Selection was without replacement if that was the planned design.
- The selection output has been saved and cannot be silently regenerated.
Do not repeatedly redraw the sample because its demographic composition “does not look right.” Doing so changes the selection procedure and may introduce researcher discretion.
When guaranteed subgroup numbers are required, stratified random sampling is usually more appropriate.
Step 8: Contact the selected units
Selection does not equal participation.
Use a standardised contact protocol covering:
- Number of contact attempts.
- Contact modes.
- Timing of follow-ups.
- Languages offered.
- Consent procedures.
- Incentives.
- Rules for inaccessible or ineligible units.
Researchers should maintain a disposition record showing what happened to each selected case.
Step 9: Manage nonresponse without convenience replacement
A refusal, unreachable participant, or incomplete response should not automatically be replaced by the next convenient person.
Acceptable strategies may include:
- Selecting a larger initial sample to allow for expected nonresponse.
- Drawing a random reserve sample at the same time as the main sample.
- Releasing reserve cases according to a documented order.
- Applying justified nonresponse weighting adjustments.
- Reporting response rates and examining possible nonresponse bias.
Replacing a nonrespondent with a nearby friend, volunteer, colleague, or easily reached participant breaks the original probability mechanism.
Step 10: Document and report the procedure
The final report should describe:
- Population definition.
- Frame source and frame date.
- Population size.
- Sample size.
- Selection method.
- Software and random seed.
- Sampling fraction.
- Eligibility exclusions.
- Duplicate-removal procedure.
- Contact process.
- Nonresponse.
- Final analytic sample.
- Weighting or adjustment.
- Known limitations.
Worked Example of Simple Random Sampling
Suppose a university has 1,200 eligible postgraduate students and wants to survey 120 about access to research-support services. The university prepares a verified list containing one record per student, assigns IDs from 1 to 1,200, and uses a random-number generator to select 120 distinct IDs without replacement.
Population and sample
[
N=1,200
]
[
n=120
]
Inclusion probability
Each student’s probability of being included is:
[
\pi_i=\frac{n}{N}=\frac{120}{1,200}=0.10
]
Therefore, every eligible student has a 10% inclusion probability.
Sampling fraction
[
f=\frac{n}{N}=0.10
]
The study samples 10% of the frame population.
Base sampling weight
The inverse-probability weight is:
[
w_i=\frac{1}{\pi_i}=\frac{N}{n}=10
]
Before nonresponse or calibration adjustments, each responding sampled student represents ten students on the frame.
What would invalidate the example?
It would no longer be a correct SRS of the 1,200 students if the researcher:
- Used only students who attended a particular seminar.
- Asked volunteers to enrol and then selected randomly among volunteers.
- Excluded students without notifying the sampling team.
- Gave some students multiple records.
- Replaced refusals with conveniently available classmates.
- Repeated the draw until preferred demographic percentages appeared.
Simple Random Sampling Formulas
Inclusion probability
For an SRS without replacement:
[
\pi_i=\frac{n}{N}
]
where:
- (\pi_i) is the probability that unit (i) is included.
- (n) is the sample size.
- (N) is the population-frame size.
Probability of a particular sample
The number of possible samples of size (n) is:
[
\binom{N}{n}
]
The probability assigned to each possible sample is:
[
\frac{1}{\binom{N}{n}}
]
Sampling fraction
[
f=\frac{n}{N}
]
The sampling fraction becomes important when the sample is not negligible relative to the population.
Base sampling weight
[
w_i=\frac{1}{\pi_i}=\frac{N}{n}
]
In an unadjusted SRS, all sampled units receive the same base weight.
Estimating a population mean
For observed values (y_1,y_2,\ldots,y_n), the sample mean is:
[
\bar{y}=\frac{1}{n}\sum_{i=1}^{n}y_i
]
Under SRS, (\bar y) is the standard estimator of the finite-population mean.
Estimating a population total
The expansion estimator of the population total is:
[
\hat{T}=N\bar{y}
]
Penn State’s sampling-theory materials show the same relationship between the SRS sample mean and estimated population total (Penn State Department of Statistics, n.d.).
Variance of the sample mean
For simple random sampling without replacement:
[
\operatorname{Var}(\bar y)=
\left(1-\frac{n}{N}\right)\frac{S^2}{n}
]
where (S^2) is the finite-population variance.
An estimated variance is:
[
\widehat{\operatorname{Var}}(\bar y)=
\left(1-\frac{n}{N}\right)\frac{s^2}{n}
]
where (s^2) is the sample variance.
The term
[
1-\frac{n}{N}
]
is the finite population correction in variance form. Its square root is applied to the standard error.
Why the finite population correction matters
If a large portion of a finite population is observed without replacement, uncertainty is lower than it would be for a similarly sized sample from an effectively infinite population.
When (n=N), the finite population correction is zero because the researcher has conducted a census of the frame rather than a sample.
How to Calculate Sample Size for Simple Random Sampling
For a descriptive survey estimating a proportion, sample size depends on the confidence level, desired margin of error, expected proportion, and population size. The calculation should then be increased to account for anticipated nonresponse and may need further adjustment for subgroup analyses or complex planned models.
Initial sample size for a proportion
For a large population:
[
n_0=\frac{z^2p(1-p)}{e^2}
]
where:
- (z) is the standard-normal value for the selected confidence level.
- (p) is the expected population proportion.
- (e) is the desired absolute margin of error.
When no defensible estimate of (p) is available, (p=0.50) produces the largest variance and therefore a conservative sample-size estimate for a proportion.
Finite population adjustment
For a finite population:
[
n=
\frac{Nz^2p(1-p)}
{e^2(N-1)+z^2p(1-p)}
]
An equivalent form is:
[
n=\frac{Nn_0}{N+n_0-1}
]
Hypothetical calculation
Assume:
- (N=1,200)
- 95% confidence level, so (z=1.96)
- (p=0.50)
- (e=0.05)
The large-population estimate is approximately:
[
n_0=384.16
]
After the finite population adjustment:
[
n\approx291.18
]
The researcher would normally round upward to 292 completed responses.
If an 80% response rate is expected, the number initially selected would be:
[
n_{\text{selected}}=
\frac{292}{0.80}=365
]
This inflation addresses anticipated nonresponse; it is not a substitute for follow-up or nonresponse assessment.
Sample-size cautions
The formula above is suitable for a basic precision objective involving a single proportion. It may not be adequate when the study involves:
- Hypothesis testing.
- Regression.
- Multilevel models.
- Factor analysis.
- Rare outcomes.
- Several treatment groups.
- Longitudinal attrition.
- Multiple primary outcomes.
- Small-domain estimates.
- Mandatory subgroup comparisons.
Power analysis or a design-specific calculation should be used in those cases.
Simple Random Sampling With and Without Replacement
Sampling without replacement
In most surveys, each selected unit can appear only once.
If student ID 417 is selected, it is removed from the pool before the next draw. This produces (n) distinct units and is normally what researchers mean by a simple random sample from a finite population.
Sampling with replacement
After each draw, the selected unit is returned to the population and may be selected again.
Each draw is independent, but the final collection may contain repeated units. Sampling with replacement is important in statistical theory, simulation, and resampling methods but is uncommon when selecting distinct human participants for a one-time survey.
| Feature | Without replacement | With replacement |
|---|---|---|
| Can a unit appear twice? | No | Yes |
| Are successive draws independent? | No | Yes, under uniform drawing |
| Typical survey use | Common | Uncommon |
| Number of distinct respondents | Exactly (n) | May be fewer than (n) |
| Finite population correction | Relevant | Not applied in the same way |
Methods for Selecting a Simple Random Sample
Lottery method
Write each unique ID on an identical slip, mix the slips thoroughly, and draw (n) slips without looking.
Best for: Very small populations and teaching demonstrations.
Limitations: Difficult to audit at scale; slips may differ physically; mixing may be inadequate; manual transcription can introduce errors.
Random-number table
Assign IDs to all units and read valid numbers from a published random-number table according to a rule defined in advance.
Best for: Teaching or settings where computers are unavailable.
Limitations: Slower than software and vulnerable to undocumented choices about starting position and reading direction.
Spreadsheet randomisation
Generate one uniform random number for every frame record, freeze the generated values, sort ascending, and take the first (n) eligible records.
Best for: Small or medium frames and researchers familiar with spreadsheets.
Important: Spreadsheet random functions recalculate. Microsoft states that RAND() generates a new value when the sheet recalculates. The generated random column should therefore be copied and pasted as values before sorting.
Statistical software
R, Python, SPSS, SAS, Stata, and specialised survey systems can select random samples efficiently.
Best for: Reproducible research, large frames, automated workflows, and projects requiring a saved random seed.
How to Conduct Simple Random Sampling in Excel
- Place one eligible population unit in each row.
- Add a unique ID column.
- Check for duplicates and ineligible records.
- Add a column named
Random. - Enter:
=RAND()
- Fill the formula down the complete frame.
- Copy the random-number column.
- Paste it back as values so the numbers no longer change.
- Save a frozen copy of the frame.
- Sort the entire table by the random-number column from smallest to largest.
- Select the first (n) rows.
Do not sort only the random-number column. The entire row range must move together.
For exact reproducibility, code-based software with a saved seed is generally preferable because spreadsheet randomisation may be difficult to reproduce after recalculation.
How to Conduct Simple Random Sampling in Google Sheets
The process is similar:
- Create one row per eligible unit.
- Add unique IDs.
- Use:
=RAND()
- Fill the formula down.
- Copy the generated cells.
- Use Paste special → Values only.
- Sort the full range by the random column.
- Select the first (n) records.
Google’s official documentation states that RAND() returns a random number from zero inclusive to one exclusive.
How to Conduct Simple Random Sampling in R
For a frame called population and a required sample of 100 rows:
set.seed(20260628)
selected_rows <- sample(
x = seq_len(nrow(population)),
size = 100,
replace = FALSE
)
sample_data <- population[selected_rows, ]
The replace = FALSE argument prevents the same row from being selected more than once.
Record the seed, software version, frame version, and complete command in the study documentation.
How to Conduct Simple Random Sampling in Python
Using NumPy:
import numpy as np
rng = np.random.default_rng(seed=20260628)
selected_rows = rng.choice(
len(population),
size=100,
replace=False
)
sample_data = population.iloc[selected_rows].copy()
Using pandas:
sample_data = population.sample(
n=100,
replace=False,
random_state=20260628
).copy()
The replace=False setting specifies sampling without replacement. A recorded seed or random_state supports reproducibility.
How to Conduct Simple Random Sampling in SPSS
In IBM SPSS Statistics:
- Open the complete sampling-frame dataset.
- Choose Data.
- Choose Select Cases.
- Select Random sample of cases.
- Click Sample.
- Choose an exact number of cases rather than an approximate percentage when a fixed sample size is required.
- Specify the number to select and the number of records in the eligible frame.
- Save the selected-case indicator and syntax.
IBM documents that SPSS can select an approximate percentage or an exact number of cases and that the procedure samples without replacement.
How to Conduct Simple Random Sampling in SAS
A basic SAS example is:
proc surveyselect data=population
out=selected_sample
method=srs
sampsize=100
seed=20260628;
run;
METHOD=SRS requests simple random sampling. SEED= records the starting value used for reproducibility.
When Should Simple Random Sampling Be Used?
SRS is most suitable when the population is clearly defined, a sufficiently complete list of eligible units exists, all selected units can reasonably be contacted or measured, and guaranteed representation of small subgroups is not essential. It is particularly useful when transparency and straightforward design-based analysis are priorities.
Appropriate situations include:
- Selecting employee records from a complete staff database.
- Sampling student records from an enrolment list.
- Choosing products from a numbered inventory.
- Auditing transactions from a complete database.
- Selecting patient files from a defined clinical register.
- Sampling articles from a complete bibliographic set.
- Drawing addresses from a suitable local address list.
Conditions favouring SRS
SRS is a strong candidate when:
- Every eligible unit can be listed.
- Units can be uniquely identified.
- The frame is manageable.
- Data-collection costs do not depend heavily on geographic clustering.
- Subgroup estimates are not the central objective.
- Equal-probability selection is appropriate.
- The planned analysis assumes or can correctly account for SRS.
When Should Simple Random Sampling Not Be Used?
SRS may be inefficient or impractical when the population is geographically dispersed, no adequate frame exists, important subgroups are small, data collection costs vary substantially, or the research requires precise estimates for several domains. Stratified, cluster, systematic, multistage, or other designs may then be preferable.
Consider another method when:
- No suitable list exists.
- The population changes too rapidly for the frame to remain accurate.
- A small minority must have a guaranteed sample presence.
- Participants are spread across a very large geographic area.
- Travel to randomly scattered units would be prohibitively expensive.
- Natural groups such as schools, hospitals, or districts must be sampled first.
- The population has units of very different sizes or importance.
- The research is exploratory and seeks information-rich cases rather than population estimates.
- Recruitment depends on participant referrals.
- The study concerns a hidden or difficult-to-identify population.
Simple Random Sampling Compared With Other Methods
| Method | How units are selected | Main advantage | Main limitation | Best used when |
|---|---|---|---|---|
| Simple random | (n) units selected randomly from the complete frame | Transparent probabilities and straightforward analysis | Requires a suitable frame; subgroups can be missed by chance | A complete frame exists and subgroup guarantees are unnecessary |
| Systematic | Random start followed by every (k)th unit | Operationally simple and spreads selection across a list | Periodicity or list patterns can distort selection | An ordered list is available and periodicity is not problematic |
| Stratified | Population divided into strata; random samples selected within each | Ensures subgroup coverage and can increase precision | Requires reliable stratum information and more complex analysis | Important subgroups need representation or separate estimates |
| Cluster | Natural groups are randomly selected, sometimes followed by subsampling | Reduces travel and listing costs | Usually less statistically efficient when units within clusters resemble one another | The population is geographically dispersed |
| Convenience | Easily accessible units are recruited | Fast and inexpensive | Unknown selection probabilities and high risk of selection bias | Pilot, exploratory, or feasibility work with limited generalisation |
| Purposive | Researcher deliberately selects information-rich cases | Supports depth and relevance in qualitative inquiry | Does not provide design-based population inference | Cases are chosen for expertise, experience, or theoretical importance |
Simple Random Sampling Versus Stratified Sampling
SRS draws directly from the complete frame without first guaranteeing subgroup numbers. Stratified sampling divides the population into meaningful strata and randomly samples within each. Stratification is usually preferable when subgroup representation or subgroup-specific estimates are important.
Suppose a university population is:
- 80% undergraduate.
- 15% master’s.
- 5% doctoral.
An SRS of 40 students may contain very few doctoral students by chance. A stratified design can deliberately select an adequate number from each level.
If disproportionate numbers are selected, analysis weights must reflect the different selection probabilities.
Simple Random Sampling Versus Systematic Sampling
SRS independently chooses the final set through a uniform random procedure, whereas systematic sampling selects every (k)th unit after a random start. Systematic sampling may be easier to administer, but its performance depends on the ordering of the frame and the absence of harmful periodic patterns.
A systematic sample can give every unit the same marginal inclusion probability without making every possible combination of (n) units equally likely. It should therefore not automatically be labelled an SRS.
Simple Random Sampling Versus Cluster Sampling
SRS selects individual units across the complete frame. Cluster sampling first selects natural groups, such as schools or neighbourhoods, and may then study all units or a subsample within selected groups. Cluster sampling can lower fieldwork costs but often increases variance because units within a cluster may be similar.
Simple Random Sampling Versus Convenience Sampling
Simple random sampling uses known selection probabilities from a defined frame. Convenience sampling recruits units because they are easy to access. Randomly selecting people within a convenience pool does not create a probability sample of the wider target population.
For example, randomly selecting 100 respondents from an opt-in online panel may be an SRS of the current panel members, but not necessarily of all adults in a country.
Simple Random Sampling Versus Random Assignment
Random sampling determines who is included in a study. Random assignment determines which treatment, intervention, or experimental condition enrolled participants receive. Random sampling supports population generalisation, while random assignment primarily supports causal comparison by reducing systematic baseline differences between treatment groups.
A study can contain:
- Random sampling without random assignment.
- Random assignment without random sampling.
- Both.
- Neither.
A clinical experiment may recruit a convenience sample of patients and randomly assign them to treatments. Treatment comparisons may have strong internal validity, but generalisation beyond the recruited patients may remain limited.
Advantages of Simple Random Sampling
1. Transparent selection probabilities
The selection probability is easy to communicate:
[
\pi_i=\frac{n}{N}
]
This transparency supports independent review.
2. Reduced researcher discretion
The researcher does not hand-pick preferred cases. When correctly implemented, the random mechanism prevents deliberate favouring of units during selection.
3. Foundation for statistical inference
Known probabilities allow researchers to estimate sampling variance, standard errors, confidence intervals, and margins of sampling error using design-appropriate methods.
4. Straightforward weights
Before adjustments, every selected unit has the same base weight:
[
w_i=\frac{N}{n}
]
5. Simple analysis
Compared with multistage, unequal-probability, or heavily stratified designs, SRS analysis is relatively direct.
6. Reproducibility
A frozen frame, documented code, and saved random seed allow another analyst to verify or reproduce the selection.
7. Useful benchmark
SRS is frequently used as a baseline when evaluating the precision or design effect of more complex sampling plans.
Limitations of Simple Random Sampling
1. It requires a suitable frame
Many real populations cannot be completely listed. Hidden populations, unregistered workers, transient residents, and people with an undiagnosed condition may be absent from available databases.
2. A random sample is not guaranteed to be balanced
One realised sample can contain too many or too few members of a subgroup by chance. Randomness protects the process, not the appearance of every sample.
3. Small subgroups may be missed
A small but important group may have few or no sampled members. Stratification or oversampling may be required.
4. Fieldwork can be expensive
An SRS of households across a country may scatter selected units over many distant locations. Cluster or multistage sampling can reduce travel costs.
5. Frame errors remain
Duplicates, omissions, outdated records, and ineligible entries can distort selection probabilities.
6. Nonresponse can create bias
Even when the invited sample is selected correctly, the final respondents may differ from nonrespondents.
7. It may be statistically inefficient
If a variable strongly associated with the outcome is available for all frame units, stratification can produce greater precision than an unstratified SRS of the same size.
8. It does not correct measurement problems
Random selection cannot repair:
- Leading questions.
- Unreliable instruments.
- Social-desirability bias.
- Interviewer effects.
- Incorrect coding.
- Poor operational definitions.
- Missing data.
- Data-entry errors.
Does Simple Random Sampling Eliminate Bias?
No. It removes researcher discretion from the selection mechanism and produces design-unbiased estimators under appropriate conditions, but it does not eliminate coverage error, nonresponse bias, measurement error, processing error, or chance imbalance. “Random” should therefore not be used as a synonym for “perfectly representative” or “free from all bias.”
NIST notes that random sampling does not guarantee that every selected sample will represent the process perfectly; rather, it prevents systematic data-collection bias on average when properly implemented. AAPOR similarly treats coverage and nonresponse as separate sources of survey error.
Sampling Error, Coverage Error, and Nonresponse Error
Sampling error
Sampling error is the variation caused by observing a sample rather than the entire frame population.
Different valid random samples from the same population will normally produce different estimates. Larger samples generally reduce standard errors, although the rate of improvement follows a square-root relationship rather than a one-for-one relationship.
Coverage error
Coverage error occurs when the frame and target population do not correspond.
Undercoverage: Eligible units are missing.
Overcoverage: Ineligible units are present.
Duplication: Some units appear more than once and therefore receive multiple chances of selection.
Nonresponse error
Nonresponse occurs when selected units do not provide usable data.
Nonresponse becomes a source of bias when response is related to the outcome or important study characteristics, even after adjustment.
A low response rate is a warning sign, but the amount of nonresponse bias depends on differences between respondents and nonrespondents—not only on the percentage responding.
Measurement error
Measurement error arises when collected values differ from the values the study intended to measure.
It may result from:
- Ambiguous questions.
- Recall problems.
- Interviewer behaviour.
- Instrument calibration.
- Respondent misunderstanding.
- Sensitive topics.
- Data-entry mistakes.
A large random sample can produce a very precise estimate of the wrong quantity when measurement is systematically flawed.
Analysing Data From a Simple Random Sample
Analysis should reflect the design actually used. For an unadjusted SRS, the sample mean estimates the population mean, (N\bar y) estimates the population total, and all units have the same base weight. Finite population correction may be relevant when the sampling fraction is appreciable.
Sample mean
[
\bar y=\frac{\sum y_i}{n}
]
Population total
[
\hat T=N\bar y
]
Base weight
[
w_i=\frac{N}{n}
]
Standard error
For a mean under SRS without replacement:
[
SE(\bar y)=
\sqrt{
\left(1-\frac{n}{N}\right)
\frac{s^2}{n}
}
]
When can unweighted analysis be used?
For a correctly implemented SRS with equal probabilities and no subsequent adjustments, ordinary unweighted point estimates are generally appropriate because all base weights are equal.
Weights may become unequal after:
- Differential nonresponse adjustments.
- Calibration.
- Poststratification.
- Combining frames.
- Oversampling.
- Eligibility adjustments.
- Departures from the intended design.
Researchers should not describe the final data as an unweighted SRS if the actual selection or adjustment procedure produced unequal contributions.
Simple Random Sampling in Quantitative, Qualitative, and Mixed-Methods Research
Quantitative research
SRS is most closely associated with quantitative research because it supports probability-based estimates and sampling-error calculations.
Applications include:
- Cross-sectional surveys.
- Record audits.
- Prevalence studies.
- Employee surveys.
- Educational assessments.
- Quality-control studies.
- Content analysis of a defined document population.
Qualitative research
SRS can be used in qualitative research, but it is less common because qualitative studies frequently seek depth, variation, information-rich cases, or theoretical relevance rather than design-based population estimates.
A qualitative researcher might use SRS when:
- A large, clearly defined group of eligible interviewees exists.
- Random selection is ethically or administratively desirable.
- The researcher wants to avoid gatekeeper selection.
- Every case is considered similarly relevant to the research question.
Purposive, maximum-variation, criterion, or theoretical sampling may be more suitable when specific experiences or characteristics are essential.
Mixed-methods research
A mixed-methods study may:
- Draw an SRS for a quantitative survey.
- Purposively select interview participants from survey respondents.
- Draw a random qualitative subsample from particular survey-defined categories.
- Use stratified random sampling to ensure that several groups contribute to both phases.
The sampling procedure for each phase should be reported separately.
Simple Random Sampling in Modern Research
Modern administrative and digital systems can make SRS easier by providing large electronic frames, unique identifiers, automated eligibility checks, reproducible code, and secure audit trails.
However, modern databases also introduce challenges:
- Duplicate profiles across systems.
- Automated records that no longer represent active units.
- Missing populations with limited digital access.
- Algorithmically inferred eligibility.
- Privacy and data-security concerns.
- Restrictions on sharing identifiable sampling frames.
- Proprietary panel-recruitment methods.
- Confusion between random subsampling of available data and random sampling of the target population.
Large national surveys often use stratified, systematic, cluster, and multistage designs rather than a single SRS because they must balance precision, fieldwork costs, small-area estimates, and operational constraints. The U.S. Census Bureau, for example, explicitly documents complex designs for major surveys and warns that standard errors based on an SRS assumption can be incorrect for complex survey data.
Digital Tools, Reproducibility, and Artificial Intelligence
Use documented random-number tools
Appropriate tools include:
- R’s
sample()function. - NumPy’s
Generator.choice(). - pandas
DataFrame.sample(). - SPSS random case selection.
- SAS
PROC SURVEYSELECT. - Spreadsheet random functions for smaller projects.
Official software documentation should be consulted because syntax and defaults may change.
Record a random seed
A seed allows the same pseudo-random sequence to be regenerated under the same software and data conditions.
A complete reproducibility record should include:
- Seed.
- Software.
- Version.
- Function.
- Parameters.
- Frame checksum or archived frame.
- Date.
- Code file.
- Output sample IDs.
Do not use a general chatbot as the randomisation mechanism
A conversational AI system may help explain code, check documentation, draft a sampling protocol, or identify possible frame-quality checks. The final selection should nevertheless be performed with a documented random-number function designed for statistical or computational sampling.
The researcher should be able to show exactly:
- What population records were eligible.
- What random procedure was applied.
- Whether replacement was permitted.
- What seed was used.
- Which IDs were selected.
Appropriate uses of AI
AI-assisted tools may help researchers:
- Suggest duplicate-detection rules.
- Standardise address formats.
- Flag obviously incomplete records.
- Draft code for an established random-sampling library.
- Generate a frame-quality checklist.
- Explain a statistical formula.
- Prepare a selection-flow diagram.
- Create documentation from verified study metadata.
AI-related cautions
AI should not silently:
- Decide eligibility without a validated rule.
- Remove unusual records merely because they appear anomalous.
- infer sensitive characteristics unnecessarily.
- alter the frame after selection.
- replace nonrespondents.
- generate undocumented “random” IDs.
- expose personal identifiers to an external service.
Human review, data-governance requirements, privacy rules, and a preserved audit trail remain necessary.
Common Mistakes in Simple Random Sampling
Mistake 1: Sampling from the wrong population
Selecting randomly from social-media followers does not produce an SRS of the general public.
Mistake 2: Calling haphazard selection random
Choosing “different-looking” cases, selecting every person who happens to be present, or asking colleagues to suggest names is not random sampling.
Mistake 3: Assuming equal individual probabilities are sufficient
A design may give each unit the same marginal probability while allowing only certain combinations. A fixed-size SRS requires all possible samples of that size to be equally likely.
Mistake 4: Ignoring duplicate records
A person appearing twice receives two routes into the sample and therefore a greater inclusion probability.
Mistake 5: Selecting before finalising eligibility
Changing the population after seeing the sample can compromise the original selection probabilities.
Mistake 6: Using sampling with replacement unintentionally
Some programming functions default to replacement. Researchers should explicitly specify replace=False or the equivalent when distinct units are required.
Mistake 7: Letting spreadsheet values recalculate
A spreadsheet may generate a different sample whenever formulas recalculate. Random values should be frozen and the selected IDs archived.
Mistake 8: Redrawing an “unbalanced-looking” sample
Repeatedly drawing until the preferred composition appears introduces an unreported selection rule. Use stratification when subgroup numbers must be controlled.
Mistake 9: Convenience replacement
Replacing a refusal with the easiest available participant destroys the documented inclusion probabilities.
Mistake 10: Treating nonresponse as harmless
The invited SRS and respondent dataset are not identical. Researchers must report the transition from selected cases to completed observations.
Mistake 11: Confusing random sampling and random assignment
Randomly selecting participants does not assign them to experimental groups, and random treatment assignment does not establish a probability sample of the wider population.
Mistake 12: Analysing complex survey data as an SRS
Data from stratified, clustered, multistage, or unequal-probability designs require analysis that accounts for strata, clusters, and weights. Treating them as an SRS can produce incorrect standard errors.
How to Report Simple Random Sampling in a Research Paper
A complete report identifies the population, sampling frame, frame size, sample size, eligibility rules, random-selection procedure, replacement rule, software, seed, response outcome, and any weighting or deviations. Simply stating “participants were randomly selected” is not enough for independent evaluation.
Concise methodology template
The target population consisted of [population definition]. A sampling frame dated [date] was obtained from [source] and contained [N] eligible units after removing [number] duplicate and [number] ineligible records. A simple random sample of [n] units was selected without replacement using [software and function] with random seed [seed]. Each frame unit had an initial inclusion probability of [n/N]. Selected units were contacted using [contact procedure]. Of the [n] selected units, [completed number] provided usable data. [Describe reserve sampling, nonresponse adjustment, weighting, exclusions, and deviations.]
Example methodology statement
The study population consisted of all 1,200 postgraduate students registered at University X on March 1, 2026. The registrar’s enrolment database was used as the sampling frame. After duplicate and ineligible records were removed, each student was assigned one unique numeric identifier. A simple random sample of 120 students was selected without replacement in R using
sample()and the prespecified seed 20260628. Each student had an initial inclusion probability of 0.10. Three email invitations and one reminder were sent over four weeks. Ninety-eight students submitted usable questionnaires. No convenience replacements were made, and response outcomes were recorded for all selected cases.
Reporting checklist
Before publication, confirm that the paper reports:
- Target population.
- Sampling-frame source.
- Frame date.
- Frame size.
- Eligibility criteria.
- Duplicate and ineligible-record handling.
- Planned sample size.
- Selection with or without replacement.
- Random tool or software.
- Random seed.
- Inclusion probability or sampling fraction.
- Contact attempts.
- Selected, eligible, responding, and analysed counts.
- Nonresponse and missing-data procedures.
- Weights.
- Known coverage limitations.
- Deviations from the planned design.
Conclusion
Simple random sampling is a transparent probability method for selecting a fixed number of units from a defined population. Its strongest form requires every possible sample of the specified size to be equally likely. The method supports design-based statistical inference, but its quality depends on the sampling frame, sample-size plan, random-selection procedure, response process, analysis, and reporting.
SRS is most useful when a complete frame exists and guaranteed subgroup representation is unnecessary. It should not be described as automatically representative, free from all bias, or equivalent to random assignment. Correct implementation requires both a random draw and a documented research process surrounding that draw.
Frequently Asked Questions
What is simple random sampling in simple words?
Simple random sampling is a method in which researchers list all eligible population units and use a random procedure to select the required number. No unit is deliberately favoured. In a fixed-size SRS, every possible combination containing the required number of distinct units has the same probability of selection.
What is an example of simple random sampling?
A researcher has a verified list of 2,000 employees and needs 200 participants. Each employee receives a unique ID, and software selects 200 distinct IDs without replacement. Provided that every eligible employee appears once on the frame and no informal replacements are made, this is a simple random sample of the employees on that frame.
What is the formula for selection probability?
For a simple random sample of (n) units selected without replacement from (N) units:
[
\pi_i=\frac{n}{N}
]
If 100 units are selected from 1,000, each unit’s inclusion probability is (100/1,000=0.10), or 10%.
Does simple random sampling guarantee a representative sample?
No. It creates an impartial probability mechanism, but one realised sample can differ from its population by chance. Representation can also be weakened by an incomplete frame, nonresponse, measurement error, or small sample size. Stratified sampling is preferable when particular subgroup numbers must be guaranteed.
Is simple random sampling the same as random assignment?
No. Random sampling selects units from a population for inclusion in a study. Random assignment allocates enrolled experimental units to treatments or conditions. Random sampling mainly supports population inference, whereas random assignment mainly supports causal comparison.
Is simple random sampling used with or without replacement?
Most human-subject and survey applications use sampling without replacement, so every selected person appears once. Sampling with replacement allows the same unit to be selected repeatedly and is more common in theoretical probability, simulation, and resampling.
What is the main disadvantage of simple random sampling?
Its principal practical disadvantage is the need for a suitable list of the whole eligible population. It may also be costly for geographically dispersed populations and may produce too few members of small but important subgroups.
Can simple random sampling be used in qualitative research?
Yes, but it is less common. It can be appropriate when all eligible cases are considered relevant and random selection is desired. Qualitative studies seeking information-rich, varied, expert, or theoretically important cases often use purposive methods instead.
How should nonrespondents be replaced?
They should not be replaced by convenient participants. Researchers can draw a random reserve sample in advance, inflate the initial sample for anticipated nonresponse, conduct standardised follow-ups, or apply justified adjustments. The procedure and all response outcomes should be reported.
Can Excel produce a simple random sample?
Yes. Generate one RAND() value per eligible record, freeze the values, sort the entire table, and take the first (n) rows. Because spreadsheet random values can change during recalculation, the random column and selected IDs must be saved as fixed values.
How large should a simple random sample be?
Sample size depends on the research objective, confidence level, margin of error, population variability, population size, statistical power, subgroup requirements, and anticipated nonresponse. There is no universally correct sample percentage.
Is a random sample from an online panel an SRS of the general public?
Not necessarily. It may be an SRS of the eligible members currently on that panel, but panel membership may not cover the wider public and may result from opt-in recruitment. The target population, frame, recruitment process, and selection probabilities must all be examined.
