A prediction is a statement or calculated estimate of an unknown outcome based on evidence, assumptions, a theory, a model or informed judgement. In scientific research, a prediction specifies what researchers expect to observe under defined conditions. In statistics and machine learning, it is an estimated value, category or probability produced for new or unobserved data.

Prediction is central to science, statistics, forecasting and artificial intelligence. Researchers make predictions to test theories, estimate unknown outcomes, assess risks and support decisions. However, the term does not have exactly the same meaning in every discipline.
This article explains what prediction means, how it differs from a hypothesis and forecast, how researchers construct and test predictions, and how statistical and machine-learning predictions should be evaluated.
Key Takeaways
- A prediction states an expected but uncertain outcome.
- Scientific predictions should be specific, testable and connected to evidence or theory.
- A prediction does not always concern the future; it may concern any unknown or unobserved outcome.
- Predictive accuracy should normally be assessed using data that were not used to fit the model.
- A useful prediction often includes a probability or interval rather than an unsupported statement of certainty.
- Good prediction does not automatically establish explanation or causation.
What Is a Prediction?
A prediction is an expectation about an outcome that is not yet known. It may be expressed as a verbal statement, numerical value, category, probability, range or conditional scenario.
In everyday language, a prediction often concerns a future event, such as tomorrow’s weather. In research, the meaning is broader. A researcher may predict:
- What will happen in an experiment.
- Which category an observation belongs to.
- The expected value of an outcome.
- The probability that an event will occur.
- An outcome that has already occurred but is not known to the analyst.
- How a system may behave under specified conditions.
For example, a medical model could predict whether a patient currently has a condition before the definitive test result is available. This is prediction even though the condition is present at the time of analysis rather than occurring in the future.
Academic definition of prediction
In academic research, prediction can be defined as:
The evidence-based estimation of an unknown outcome for specified cases or conditions, usually accompanied by an explicit or implicit statement of uncertainty.
This definition covers experimental predictions, forecasts, regression estimates, classifications, risk probabilities and machine-learning outputs.
What makes a prediction scientifically useful?
A scientifically useful prediction normally has six characteristics:
- Specificity: It identifies the outcome that is expected.
- Defined conditions: It states when, where or for whom the outcome is expected.
- Evidence or rationale: It follows from existing knowledge, data, theory or a model.
- Testability: Researchers can compare the prediction with observations.
- Measurability: The outcome can be recorded or operationally defined.
- Acknowledged uncertainty: The prediction is not presented as guaranteed unless the underlying relationship is genuinely deterministic.
“Students will perform better” is too vague for a strong scientific prediction. A clearer version is:
Students receiving weekly retrieval-practice exercises will obtain a higher mean score on the final assessment than students receiving rereading exercises, after both groups receive the same instructional time.
This version defines the groups, intervention, comparison and measurable outcome.
Prediction in the Scientific Method
In the scientific method, a prediction states what researchers expect to observe if a hypothesis or model is supported under specified conditions.
A simplified sequence is:
- Observe a phenomenon.
- Formulate a research question.
- Develop a possible explanation or hypothesis.
- Derive one or more predictions.
- Collect evidence through an experiment or observational study.
- Compare the observations with the predictions.
- Retain, revise or reject parts of the explanation.
A hypothesis is usually broader than the prediction derived from it. For example:
- Research question: Does soil nitrogen affect tomato-plant growth?
- Hypothesis: Increased available nitrogen promotes vegetative growth in tomato plants.
- Prediction: Tomato plants receiving a specified nitrogen treatment will have a greater mean height after six weeks than untreated plants grown under otherwise similar conditions.
Researchers should not treat a supported prediction as final proof of a hypothesis. Different mechanisms can sometimes generate the same predicted observation. Evidence may support a hypothesis without showing that it is the only possible explanation.
Unexpected findings are also scientifically valuable. A failed prediction may indicate that:
- The hypothesis is incomplete.
- A measurement was unreliable.
- An uncontrolled variable affected the outcome.
- The expected relationship depends on conditions not represented in the study.
- The statistical model was incorrectly specified.
- A new explanation should be developed.
Hypothesis vs. Prediction
A hypothesis is a proposed explanation or testable claim, whereas a prediction is the outcome expected if that hypothesis applies under specified conditions.
| Feature | Hypothesis | Prediction |
|---|---|---|
| Main purpose | Proposes an explanation, relationship or claim | States the expected observation or outcome |
| Level | Usually general or theoretical | Usually specific to a study or condition |
| Typical wording | “X affects Y because…” | “If X changes, Y will…” |
| Variables | May identify a general relationship | Should operationalize measurable variables |
| Testing | Tested indirectly through evidence | Compared directly with observed results |
| Number | One hypothesis may guide a study | One hypothesis may generate several predictions |
| Example | Sleep deprivation reduces sustained attention | Participants restricted to four hours of sleep will make more attention-task errors than participants allowed eight hours |
The terms are sometimes used interchangeably, especially in introductory materials. Nevertheless, distinguishing them improves research design. The hypothesis explains or proposes a relationship; the prediction translates that claim into an expected empirical pattern.
Prediction vs. Forecast, Projection, Estimation and Explanation
Related terms overlap, but they are not always interchangeable.
| Term | Core meaning | Must concern the future? | Example |
|---|---|---|---|
| Prediction | Estimated unknown outcome based on evidence, assumptions or a model | No | Predicting whether an image contains a tumour |
| Forecast | Prediction explicitly indexed to a future time | Usually yes | Forecasting next quarter’s sales |
| Projection | Conditional result produced by extending assumptions or trends | Usually | Projecting population under specified fertility assumptions |
| Estimate | Approximation of an unknown quantity or parameter | No | Estimating the mean income of a population |
| Inference | Drawing conclusions about unobserved quantities, populations or processes from evidence | No | Inferring a population effect from a sample |
| Explanation | Identifying why or how an outcome occurs | No | Explaining how a treatment changes a biological pathway |
| Scenario | Internally consistent description of what could occur under stated assumptions | Usually | Examining high-, medium- and low-emission futures |
| Guess | An answer made with limited information or systematic reasoning | No | Guessing tomorrow’s temperature without data |
Prediction vs. forecasting
Forecasting is best understood as a time-oriented form of prediction. All forecasts are predictions, but not all predictions are forecasts.
A classifier that predicts whether an email is spam is not necessarily forecasting. A time-series model estimating the number of emails that will arrive next week is forecasting.
Prediction vs. explanation
Prediction asks:
How accurately can an outcome be estimated for new cases?
Explanation asks:
Why does the outcome vary, and what process or causal relationship produced it?
A model may predict accurately while offering limited causal explanation. Conversely, a theoretically meaningful relationship can be difficult to use for accurate individual prediction. Researchers should therefore state whether their principal goal is prediction, explanation, causal inference or a combination of these goals (Shmueli, 2010; Verhagen, 2022).
Prediction vs. causation
A variable can improve prediction without causing the outcome. For example, the presence of fire engines may predict the severity of fire damage, but sending fewer fire engines would not reduce the damage. Both are consequences of the underlying size of the fire.
A predictive association should not be described as causal unless the study design and assumptions support a causal interpretation.
Main Types of Prediction
Predictions can be classified according to their format, data, uncertainty and analytical purpose.
Qualitative and quantitative predictions
A qualitative prediction describes the expected direction, category or pattern without giving an exact number.
Example:
Participants exposed to the reminder will be more likely to complete the survey.
A quantitative prediction specifies a numerical value, difference, probability or range.
Example:
The model predicts a 0.68 probability that the participant will complete the survey.
Quantitative predictions are generally easier to score precisely, but qualitative predictions remain useful when numerical information is unavailable or inappropriate.
Point prediction
A point prediction gives one estimated value.
Example:
Predicted examination score: 74.
Point predictions are easy to communicate, but they conceal uncertainty. A score of 74 may be accompanied by substantial uncertainty depending on the data and model.
Interval prediction
An interval prediction gives a range intended to contain an individual outcome with a stated level of coverage under the model’s assumptions.
Example:
Predicted score: 74, with a 95% prediction interval from 63 to 85.
The interval communicates that the future or unobserved value is not expected to equal the point estimate exactly.
Probabilistic prediction
A probabilistic prediction assigns probabilities to possible outcomes.
Example:
- 20% probability of no improvement.
- 55% probability of moderate improvement.
- 25% probability of substantial improvement.
Probabilities allow decision-makers to account for uncertainty. They should be calibrated: among cases assigned a probability near 0.70, the event should occur roughly 70% of the time when assessed over an appropriate set of comparable cases.
Deterministic and stochastic predictions
A deterministic prediction produces the same output whenever the inputs and rules are the same. It commonly appears in systems governed by highly stable mathematical relationships.
A stochastic or probabilistic prediction explicitly represents random variation or uncertainty. Most predictions involving human behaviour, disease, markets, weather and complex social systems are probabilistic.
Conditional and unconditional predictions
A conditional prediction states what is expected given particular information:
Given the student’s prior achievement, attendance and assignment completion, the predicted probability of passing is 0.82.
An unconditional prediction does not condition on case-specific predictors:
The overall pass rate is expected to be 72%.
Conditional predictions are often more useful for individual cases, although their reliability depends on predictor quality and model validity.
Classification prediction
Classification assigns an observation to a category or estimates the probability of category membership.
Examples include:
- Spam or not spam.
- Disease present or absent.
- Low, medium or high risk.
- Positive, neutral or negative sentiment.
Regression prediction
Regression predicts a continuous numerical outcome, such as:
- Examination score.
- Blood pressure.
- Household expenditure.
- Crop yield.
- Waiting time.
Time-series prediction
Time-series prediction, usually called forecasting, estimates values indexed by time. Common approaches include naïve benchmarks, exponential smoothing, autoregressive models, state-space models and machine-learning methods.
Time-series data require validation that respects temporal order. Randomly mixing past and future observations can produce an unrealistic estimate of forecast performance.
How to Write a Testable Prediction
A testable prediction should specify the conditions, expected outcome, comparison, measurement and relevant time frame.
Step 1: Begin with the research question and rationale
Identify what is already known and why a particular outcome is expected.
Question:
Does retrieval practice improve long-term retention compared with rereading?
Rationale:
Retrieval practice requires learners to reconstruct information and may strengthen later access to it.
Step 2: Identify the variables
Specify:
- The predictor, exposure or intervention.
- The outcome.
- The comparison group or reference condition.
- Important controls or boundary conditions.
Step 3: State the expected direction or value
Indicate whether the outcome is expected to increase, decrease, differ, remain stable or fall within a numerical range.
Step 4: Make the outcome measurable
Replace ambiguous terms such as better, successful or effective with operational measures.
Instead of:
Retrieval practice will improve learning.
Write:
The retrieval-practice group will obtain a higher mean score on a 30-item delayed-recall test administered seven days after instruction.
Step 5: Specify the conditions and time frame
A prediction should indicate the population, setting and timing when these details affect interpretation.
Step 6: Define how the prediction will be evaluated
Before examining the results, decide:
- What observation would support the prediction?
- What would contradict it?
- Which metric or statistical analysis will be used?
- What amount of error is acceptable?
- Which data will be reserved for evaluation?
Scientific prediction template
If [intervention, exposure or condition], then [measurable outcome] will [increase, decrease or differ] compared with [reference condition] among [population] over [time period], because [theoretical or empirical rationale].
A shorter experimental form is:
If [independent variable changes], then [dependent variable will change in a specified direction].
The shorter form is useful for teaching, but advanced research predictions should also identify the population, measurement, comparison and assumptions.
Prediction in Statistics
Statistical prediction uses a model fitted to observed data to estimate an unknown outcome for a new or unobserved case.
Let (Y) represent an outcome and (X_1, X_2, \ldots, X_p) represent predictors. A prediction model can be written generally as:
[
\hat{Y}=f(X_1,X_2,\ldots,X_p)
]
Here:
- (\hat{Y}) is the predicted outcome.
- (f) is the fitted model or algorithm.
- (X_1,\ldots,X_p) are input variables.
Linear-regression prediction
A multiple linear-regression prediction can be expressed as:
[
\hat{Y}=b_0+b_1X_1+b_2X_2+\cdots+b_pX_p
]
Where:
- (b_0) is the intercept.
- (b_1,\ldots,b_p) are estimated coefficients.
- (X_1,\ldots,X_p) are predictor values.
Suppose a simplified model predicts an assessment score from weekly study hours:
[
\hat{Y}=50+3X
]
For a student studying six hours per week:
[
\hat{Y}=50+3(6)=68
]
The predicted score is 68. This is an expected value under the model, not a guaranteed result.
Prediction error
For an observed value (Y_i) and predicted value (\hat{Y}_i), the prediction error or residual is:
[
e_i=Y_i-\hat{Y}_i
]
If the observed score is 72 and the predicted score is 68:
[
e_i=72-68=4
]
The model underpredicted the score by four points.
The sign convention may vary across software and texts, so researchers should define how error is calculated.
Common prediction-error metrics
Mean absolute error
[
MAE=\frac{1}{n}\sum_{i=1}^{n}|Y_i-\hat{Y}_i|
]
MAE gives the average absolute error in the outcome’s original units. It is relatively easy to interpret.
Mean squared error
[
MSE=\frac{1}{n}\sum_{i=1}^{n}(Y_i-\hat{Y}_i)^2
]
MSE gives greater weight to large errors because the errors are squared.
Root mean squared error
[
RMSE=\sqrt{\frac{1}{n}\sum_{i=1}^{n}(Y_i-\hat{Y}_i)^2}
]
RMSE is in the original units of the outcome but remains particularly sensitive to large errors.
No single metric is universally best. The choice should reflect the outcome, decision context, consequences of different errors and whether large errors deserve additional penalty.
Prediction Interval vs. Confidence Interval
A confidence interval usually quantifies uncertainty about a population parameter or mean response, whereas a prediction interval quantifies uncertainty about an individual future or unobserved outcome.
| Feature | Confidence interval | Prediction interval |
|---|---|---|
| Main target | Parameter or expected mean response | Individual observation |
| Sources of variability | Uncertainty in estimating the mean or parameter | Estimation uncertainty plus individual variation |
| Typical width | Narrower | Wider |
| Example | Mean score expected for students studying six hours | Score expected for one particular student studying six hours |
| Main question | Where is the mean or parameter likely to lie? | Where might an individual outcome lie? |
For a normally distributed forecast under simplified assumptions, a 95% prediction interval may take the general form:
[
\hat{Y}\pm1.96\times SE_{\text{prediction}}
]
The exact formula depends on the model, data structure and assumptions. A prediction interval should not be interpreted as an absolute guarantee, especially when the model is misspecified or the new case differs substantially from the development data.
How Prediction Models Are Developed and Evaluated
A prediction model should be evaluated on observations that were not used to fit or tune it. Its performance should be assessed with metrics suited to the outcome, including both accuracy and uncertainty where relevant.
1. Define the prediction problem
State:
- The target population.
- The intended users.
- The outcome.
- The prediction time.
- The time horizon.
- The intended decision or application.
- Whether the task is diagnostic, prognostic, classificatory or forecasting.
A vague objective such as “predict academic success” is insufficient. Academic success could mean first-year GPA, graduation, course completion or employment after graduation.
2. Select and define predictors
Predictors should be available at the moment the prediction will be made. Using information collected after the target event creates leakage and makes performance unrealistic.
Researchers should document:
- Measurement methods.
- Timing.
- Coding.
- Missing values.
- Transformations.
- Whether predictor selection was theory-based, data-driven or both.
3. Establish a baseline
A sophisticated model should be compared with a simple benchmark.
Possible baselines include:
- The sample mean.
- The majority class.
- The most recent time-series value.
- An established clinical or educational rule.
- A simple regression model.
A complex method that does not outperform a reasonable baseline may add cost without adding practical value.
4. Separate model fitting from evaluation
The available data may be divided conceptually into:
- Training data: Used to estimate model parameters.
- Validation data: Used to tune choices such as hyperparameters.
- Test data: Used for a final, relatively unbiased evaluation.
When a dataset is limited, cross-validation or bootstrap methods can provide more efficient internal validation. The procedure must prevent information from the evaluation fold entering preprocessing, feature selection or tuning.
For time-series prediction, validation should preserve chronological order, such as through rolling-origin evaluation.
5. Fit and tune the model
The chosen method should match:
- The type of outcome.
- Sample size.
- Predictor structure.
- Missing-data process.
- Nonlinearity and interactions.
- Interpretability needs.
- Intended deployment setting.
More complex models are not automatically better. Complexity can increase overfitting, computation, maintenance requirements and difficulty of explanation.
6. Evaluate predictive performance
Different outcomes require different metrics.
| Prediction task | Useful metrics | Main interpretation |
|---|---|---|
| Continuous outcome | MAE, RMSE, prediction intervals | Typical size and distribution of numerical errors |
| Binary classification | Sensitivity, specificity, precision, recall, AUROC | Ability to distinguish and classify cases |
| Probabilistic binary prediction | Brier score, log loss, calibration plot, calibration slope | Quality and reliability of predicted probabilities |
| Multiclass classification | Macro/micro precision, recall, F1, log loss | Performance across several categories |
| Time-series forecasting | MAE, RMSE, MASE, interval coverage | Accuracy across forecast horizons |
| Time-to-event prediction | Concordance, time-dependent discrimination, calibration, Brier score | Ranking and probability accuracy over time |
Accuracy can be misleading when outcomes are imbalanced. For example, a model predicting that no one has a rare condition may be highly accurate overall while failing to identify any affected cases.
7. Assess discrimination
Discrimination describes how well a model separates cases with different outcomes.
For a binary outcome, a model has good discrimination when individuals who experience the event generally receive higher predicted risks than individuals who do not.
Discrimination does not show whether predicted probabilities are numerically accurate.
8. Assess calibration
Calibration describes agreement between predicted probabilities and observed outcome frequencies.
Suppose a model assigns a predicted probability close to 0.20 to 1,000 comparable cases. It is well calibrated in that range when approximately 200 of those cases experience the event.
A model may rank cases correctly but systematically overestimate or underestimate their risks. Reporting discrimination without calibration therefore provides an incomplete assessment.
9. Quantify uncertainty
Useful uncertainty information may include:
- Prediction intervals.
- Probability distributions.
- Confidence intervals for performance metrics.
- Bootstrap estimates.
- Sensitivity analyses.
- Performance across repeated data splits.
- Uncertainty caused by missing data or measurement error.
A precise-looking prediction can be misleading when its uncertainty is large or poorly characterized.
10. Conduct external validation
Internal validation estimates performance within the development dataset or closely related resamples. External validation evaluates the model using meaningfully different data, such as:
- A different institution.
- A different country.
- A later time period.
- A new demographic group.
- A different data-collection system.
External validation is important because predictive performance often changes across settings.
11. Evaluate practical usefulness
A statistically accurate model is not necessarily useful. Researchers should also consider:
- Whether the necessary inputs are available.
- Whether the result arrives in time to support a decision.
- The costs of false positives and false negatives.
- Whether the model improves decisions relative to current practice.
- Whether users understand the output.
- Whether the model creates unequal harms.
- Whether the intervention prompted by the prediction is effective.
12. Monitor and update the model
Relationships can change after deployment. Changes in population characteristics, measurement systems, policies, behaviour or technology may reduce performance.
Monitoring should examine:
- Input-data changes.
- Missingness.
- Calibration.
- Error rates.
- Subgroup performance.
- Changes in outcome prevalence.
- Feedback from users.
- Whether retraining or recalibration is required.
Prediction Examples in Research
| Field | Research prediction | Appropriate evaluation |
|---|---|---|
| Biology | Seedlings receiving more light will develop greater mean dry mass after four weeks | Group comparison with uncertainty estimates |
| Education | Prior achievement, attendance and assignment completion will predict final course score | Out-of-sample MAE or RMSE |
| Psychology | Participants in the sleep-restriction condition will show slower reaction times | Experimental comparison |
| Public health | A model will estimate an individual’s probability of hospital readmission within 30 days | Calibration, discrimination and clinical usefulness |
| Sociology | Household characteristics will predict the probability of residential displacement | Test-set probability metrics and subgroup analysis |
| Economics | Monthly sales will decline after a specified price increase, other measured conditions being equal | Forecast error and sensitivity analysis |
| Environmental science | River level will exceed a warning threshold within 24 hours | Recall, false-alarm rate, lead time and probability calibration |
| Computer science | A classifier will label incoming messages as spam or non-spam | Precision, recall, F1, calibration and error analysis |
Experimental example
Hypothesis: Higher water temperature increases the metabolic rate of the study organism within its tolerated range.
Prediction: Organisms maintained at 24°C will consume more oxygen per hour than organisms maintained at 18°C when body mass and measurement duration are held constant.
The prediction is useful because the conditions and outcome are measurable.
Observational-research example
A researcher may use attendance, previous grades and assignment submission patterns to predict course completion.
This model could support early assistance, but it should not be interpreted automatically as showing that each predictor causes completion. The model should also be checked for differential error across student groups and for possible consequences of labeling students as high risk.
Probabilistic example
A weather model may report a 70% probability of rain. The prediction is not disproved merely because rain does not occur on one occasion. It should be evaluated across many comparable 70% forecasts. Good calibration would mean that rain occurs on approximately 70% of those occasions.
Advantages of Prediction
Prediction makes theories empirically testable
A theory becomes more informative when it generates observable implications. Predictions connect abstract propositions to evidence.
Prediction supports planning
Forecasts and risk estimates help organizations allocate resources, prepare for demand and compare possible actions.
Prediction evaluates generalization
Testing predictions on unseen data reveals whether a model captures patterns that extend beyond its development sample.
Prediction creates common benchmarks
Models can be compared using the same outcomes, datasets and evaluation rules. This encourages transparent assessment rather than relying only on claims about theoretical sophistication.
Prediction can identify useful patterns
Predictive methods can reveal nonlinearities, interactions and combinations of variables that simpler analyses may overlook.
Prediction communicates uncertainty
Probabilities and intervals can make uncertainty explicit, supporting more informed decisions than categorical claims of certainty.
Limitations and Risks of Prediction
Predictions are uncertain
Complex systems contain random variation, incomplete information and processes that cannot be measured perfectly. Even a well-designed model will make errors.
Historical data may not represent future or external conditions
A model learns relationships in its development data. It may fail when populations, institutions, measurement practices or behaviour change.
Prediction does not necessarily explain
High predictive performance does not by itself identify mechanisms or establish why an outcome occurs.
Prediction does not prove causation
Predictors may be correlated with an outcome without producing it. Predictive models should not be used to estimate intervention effects unless the design and assumptions support that purpose.
Models can reproduce bias
Historical data may contain unequal treatment, measurement error or structural disadvantage. A model can encode these patterns and apply them at scale.
Overall performance can hide subgroup failure
A model may perform well on average but poorly for a smaller population. Researchers should examine error rates, calibration and data representation across relevant groups.
Predictions can change behaviour
A prediction may become self-fulfilling or self-negating. For example, labeling a student as unlikely to succeed could reduce opportunities and contribute to the predicted outcome. Conversely, an early-warning prediction may prompt support that prevents the outcome.
Interpretability may be limited
Complex algorithms can be difficult to explain. Explanations of individual predictions should be tested carefully and should not be confused with proof of causal mechanisms.
Long-range forecasts depend heavily on assumptions
Uncertainty generally grows as the prediction horizon expands. Scenario analysis may be more appropriate than a single precise value when human choices, policies or disruptive events strongly influence the outcome.
Common Prediction Mistakes
1. Treating a prediction as certainty
Use language such as estimated, expected, probable or conditional on the model. Avoid saying that a model “knows” an outcome.
2. Testing a model on its training data
Training performance is usually optimistic. Evaluation should use unseen cases or appropriate resampling.
3. Data leakage
Leakage occurs when information unavailable at the real prediction time influences model development. Examples include:
- Using post-outcome variables.
- Preprocessing the full dataset before splitting it.
- Selecting predictors using the final test set.
- Allowing records from the same person to appear in both training and test sets.
4. Overfitting
An overfitted model learns noise and dataset-specific details. It may appear highly accurate during development but fail on new data.
5. Ignoring a simple baseline
Every candidate model should be compared with a reasonable benchmark. Small improvements may not justify substantial complexity.
6. Selecting a convenient metric
A metric should reflect the research objective and consequences of error. Overall accuracy is inadequate for many imbalanced or high-stakes tasks.
7. Ignoring calibration
A high AUROC does not establish that predicted probabilities are accurate. Probability models require calibration assessment.
8. Confusing association with causation
Predictive importance, regression coefficients and feature-attribution scores do not automatically represent causal effects.
9. Extrapolating beyond the data
Predictions for cases far outside the range of the development data are often unreliable. Researchers should define the intended domain of use.
10. Hiding uncertainty
Publishing only a point estimate can create false precision. Report intervals, probabilities or other uncertainty measures where feasible.
11. Reporting only the best model
Trying many models and reporting only the highest-performing result can produce optimistic conclusions. The model-selection process and all final evaluation rules should be transparent.
12. Failing to update a deployed model
Predictive relationships may deteriorate. Monitoring, recalibration and revalidation should be planned before real-world use.
Prediction in Modern Research
Prediction has become increasingly important across disciplines because large datasets and greater computing capacity allow researchers to evaluate complex models on unseen cases.
In social science, predictive analysis can complement explanation by showing whether a model meaningfully approximates real outcomes rather than merely fitting relationships in the observed sample (Verhagen, 2022).
In health research, prediction models may support diagnosis, prognosis, screening and monitoring. Current reporting guidance applies to models developed with conventional regression as well as machine-learning methods. Transparent reports should describe the data, participants, outcomes, predictors, modeling process, validation and performance clearly enough for critical assessment (Collins et al., 2024).
Prediction is also important for replication. A relationship that repeatedly supports accurate, preregistered predictions in new settings provides stronger evidence of generalizability than a pattern evaluated only in the data where it was discovered.
Artificial Intelligence and Prediction
In supervised machine learning, prediction is the output generated when a trained algorithm receives input data and estimates a value, class or probability.
Examples include:
- A numerical house-price prediction.
- A probability of loan default.
- A predicted image category.
- A next-word probability distribution.
- A forecast of equipment failure.
- A predicted risk of disease progression.
A typical AI prediction workflow
- Define the outcome and intended use.
- Collect representative data.
- Clean and document the data.
- Split data using a design appropriate to the application.
- Fit a baseline and one or more candidate models.
- Tune the models without accessing the final test outcomes.
- Evaluate discrimination, calibration, error and uncertainty.
- Conduct subgroup and robustness analyses.
- Validate the model externally.
- Assess usefulness, risks and human oversight.
- Monitor performance after implementation.
AI tools used in prediction research
Researchers commonly use:
- R: Base modeling functions, tidymodels, fable and discipline-specific packages.
- Python: scikit-learn, statsmodels, PyTorch, TensorFlow and related libraries.
- Statistical software: SPSS, Stata, SAS and JMP.
- Interactive environments: Jupyter notebooks and RStudio.
- Reproducibility tools: Version control, computational environments and documented data pipelines.
- Visualization tools: Calibration plots, residual plots, ROC curves and forecast charts.
The choice of software does not determine methodological quality. A simple model evaluated correctly is more credible than an advanced model affected by leakage, inappropriate validation or selective reporting.
Appropriate use of generative AI
Generative AI may help researchers:
- Draft or explain code.
- Create documentation.
- Suggest diagnostic checks.
- Translate technical explanations.
- Develop preliminary analysis plans.
Its output should be checked independently. Researchers remain responsible for the data, assumptions, code, references, validation and interpretation. Sensitive or restricted data should not be entered into an external system without appropriate authorization and safeguards.
Reporting a Prediction Study
A transparent report should answer the following questions:
Objective and setting
- What outcome is being predicted?
- For which population and setting?
- At what prediction time and horizon?
- How will the prediction be used?
Data
- How were participants or observations selected?
- When and where were the data collected?
- How were the outcome and predictors measured?
- How were missing values handled?
- Were repeated or clustered observations present?
Model development
- Which predictors were considered?
- How were continuous variables handled?
- Which algorithm or statistical model was used?
- How were features, interactions and hyperparameters selected?
- How was overfitting addressed?
Validation
- Was evaluation internal, temporal, geographic or fully external?
- Were all preprocessing steps confined to the training data?
- Was the test set kept independent?
- Were confidence intervals reported for performance estimates?
Performance and usefulness
- Which metrics were chosen and why?
- Were both discrimination and calibration evaluated?
- Were prediction intervals or probability uncertainty assessed?
- How did the model compare with a baseline?
- Was subgroup performance examined?
- Was real-world usefulness or impact evaluated?
Transparency
- Is the final model or algorithm described sufficiently?
- Are code, data dictionaries or executable instructions available where ethically and legally possible?
- Are limitations and intended boundaries of use explicit?
- Are conflicts of interest and funding sources disclosed?
Researchers working on diagnostic or prognostic models should consult discipline-specific reporting and risk-of-bias guidance, including TRIPOD+AI and PROBAST+AI where applicable.
Conclusion
Prediction is the evidence-based estimation of an unknown outcome. In scientific research, it translates a hypothesis or model into an expected observation. In statistics and artificial intelligence, it estimates a value, category, probability or distribution for new or unobserved data.
A credible prediction is specific, testable and transparent about uncertainty. Its quality should be assessed on appropriate unseen data, using metrics that reflect the intended purpose. Even highly accurate predictions must be distinguished from causal explanations and evaluated for calibration, generalizability, fairness and practical usefulness.
References
- Collins, G. S., Moons, K. G. M., Dhiman, P., Riley, R. D., Beam, A. L., Van Calster, B., et al. (2024). TRIPOD+AI statement: Updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ, 385, e078378. https://doi.org/10.1136/bmj-2023-078378
- Hyndman, R. J., & Athanasopoulos, G. (2021). Forecasting: Principles and practice (3rd ed.). OTexts.
- Moons, K. G. M., Damen, J. A. A., Kaul, T., Hooft, L., Andaur Navarro, C., Dhiman, P., et al. (2025). PROBAST+AI: An updated quality, risk of bias, and applicability assessment tool for prediction models using regression or artificial intelligence methods. BMJ, 388, e082505. https://doi.org/10.1136/bmj-2024-082505
- Moriarty, P. (2023). Modern methods of prediction. Encyclopedia, 3(2), 520–529. https://doi.org/10.3390/encyclopedia3020037
- Shmueli, G. (2010). To explain or to predict? Statistical Science, 25(3), 289–310. https://doi.org/10.1214/10-STS330
- Tavazza, F., DeCost, B., & Choudhary, K. (2021). Uncertainty prediction for machine learning models of material properties. ACS Omega, 6(48), 32431–32440. https://doi.org/10.1021/acsomega.1c03752
- Verhagen, M. D. (2022). A pragmatist’s guide to using prediction in the social sciences. Socius, 8, 1–15. https://doi.org/10.1177/23780231221081702
