Graphical methods are visual techniques used to explore, summarize, analyze, and communicate data through charts, plots, maps, and diagrams. Researchers use them to reveal distributions, comparisons, trends, relationships, unusual observations, and model problems that may be difficult to detect in tables or numerical summaries alone.

Introduction
A dataset can contain hundreds or millions of observations, but a well-designed graph may reveal its most important pattern within seconds. Graphical methods convert values, categories, dates, and statistical results into visual forms that researchers can inspect and communicate.
Graphs are not merely decorative additions to a research report. They can help researchers discover errors, identify unexpected patterns, examine assumptions, compare groups, select statistical models, and explain results to readers.
This guide explains:
- What graphical methods are
- The major types used in research and statistics
- How to select an appropriate graph
- How common graphs are constructed and interpreted
- Their advantages and limitations
- Common graphical mistakes
- How uncertainty, accessibility, reproducibility, and artificial intelligence affect modern research visualization
Key Takeaways
- The correct graph depends on the research question, variable type, number of variables, and analytical purpose.
- Graphs support both exploratory analysis and the communication of final findings.
- Bar charts display categories, while histograms display intervals of a quantitative variable.
- Scatterplots can reveal association but cannot by themselves demonstrate causation.
- Graphical analysis should normally be combined with numerical summaries and appropriate statistical methods.
- Clear scales, uncertainty information, accessible design, and transparent data processing are essential for trustworthy figures.
What Are Graphical Methods?
Graphical methods are procedures for representing data or analytical results visually so that patterns, comparisons, distributions, relationships, and anomalies can be examined.
They include familiar displays such as bar charts and line graphs, as well as more analytical plots such as Q–Q plots, residual plots, scatterplot matrices, heatmaps, and principal-component biplots.
In statistics, graphical methods are closely associated with exploratory data analysis. Exploratory data analysis uses visual and numerical techniques to allow the structure of the data to become visible before strong assumptions or final models are imposed (Tukey, 1977; NIST/SEMATECH, n.d.).
Graphical methods versus numerical methods
Graphical and numerical methods answer related but different questions.
| Method | Main contribution | Example |
|---|---|---|
| Graphical method | Shows shape, patterns, relationships, clusters, and unusual observations | Histogram of examination scores |
| Numerical method | Produces precise summaries or estimates | Mean score and standard deviation |
| Inferential method | Quantifies uncertainty or tests a hypothesis | Confidence interval or t test |
A mean and standard deviation may summarize a variable efficiently, but they may not reveal that the distribution is highly skewed, contains two clusters, or includes a serious data-entry error. A graph can reveal these features, while numerical methods provide precision and formal evidence.
The strongest analysis normally uses both.
Exploratory and explanatory graphics
Graphical methods serve two broad purposes.
Exploratory graphics are created while analyzing data. They may contain many observations, panels, variables, or diagnostic elements. Their purpose is discovery.
Explanatory or communicative graphics are designed to present a specific result to an audience. They usually have clearer annotations, fewer distractions, and a focused message.
A researcher might inspect twenty exploratory plots while developing a model but include only two carefully designed figures in the final article.
Why Are Graphical Methods Used in Research?
Graphical methods are used because they can make important properties of data visible.
To summarize large datasets
A graph can condense many observations into a form that is easier to inspect than a raw spreadsheet. A histogram, for example, can summarize the overall distribution of thousands of measurements.
To identify distribution shape
Graphs can reveal whether values are:
- Symmetrical or skewed
- Unimodal or multimodal
- Concentrated or widely dispersed
- Truncated or bounded
- Affected by extreme observations
To compare groups
Side-by-side boxplots, dot plots, violin plots, and grouped bar charts can show differences among treatments, classes, regions, or demographic groups.
To examine relationships
Scatterplots and related displays help researchers investigate direction, form, strength, clusters, and unusual cases in relationships between variables.
To analyze change over time
Line graphs, run charts, control charts, and seasonal plots help identify trends, cycles, sudden changes, and periods of instability.
To detect errors and anomalies
Unexpected gaps, impossible values, duplicate patterns, discontinuities, and outliers may indicate data-entry problems, measurement errors, or genuinely unusual observations.
To evaluate statistical assumptions
Histograms, Q–Q plots, residual plots, scale-location plots, and influence plots are used to assess assumptions related to distributional form, linearity, equal variance, independence, and influential observations.
To communicate evidence
A well-designed graph can make the main result of a study understandable to readers who may not be familiar with the underlying statistical procedures.
Main Types of Graphical Methods
The most appropriate classification is based on the analytical purpose rather than appearance alone.
| Analytical purpose | Suitable graphical methods |
|---|---|
| Compare categories | Bar chart, dot plot, Pareto chart |
| Show composition | Stacked bar chart, 100% stacked bar, pie chart with few categories |
| Examine one quantitative distribution | Dot plot, histogram, density plot, boxplot, violin plot, ECDF |
| Compare quantitative distributions | Side-by-side boxplots, violin plots, ridgeline plots, faceted histograms |
| Show a relationship between two quantitative variables | Scatterplot, hexbin plot, contour plot |
| Show change over time | Line graph, run chart, seasonal plot, control chart |
| Compare repeated observations | Connected dot plot, profile plot, spaghetti plot |
| Examine two categorical variables | Grouped bar chart, mosaic plot, association plot |
| Examine many variables | Scatterplot matrix, heatmap, parallel-coordinates plot, biplot |
| Show geographic variation | Choropleth map, proportional-symbol map |
| Evaluate a fitted model | Residual plot, Q–Q plot, leverage plot, observed-versus-predicted plot |
Graphical Methods for Categorical Data
Categorical variables place observations into groups, such as field of study, treatment condition, employment status, or response category.
Bar chart
A bar chart represents each category with a rectangular bar. The bar length or height shows a count, percentage, rate, mean, or another clearly labelled value.
Bar charts are appropriate when the goal is to compare categories.
Example: A researcher compares the percentage of students using four learning platforms.
Good practice includes:
- Starting the quantitative axis at zero when bar length represents magnitude
- Ordering nominal categories logically or by value
- Preserving the natural order of ordinal categories
- Labelling the measurement and units
- Avoiding unnecessary three-dimensional effects
Dot plot for categories
A categorical dot plot places a point at the value for each category. It may be more compact than a bar chart and makes comparison along a common scale easy.
Dot plots are particularly useful when:
- There are many categories
- Values do not need filled bars
- Confidence intervals must be shown
- Several estimates are compared
Pareto chart
A Pareto chart is a bar chart in which categories are ordered from the highest to the lowest frequency. A cumulative-percentage line may also be included.
It is commonly used to identify which categories account for the largest share of a problem, defect, complaint, or outcome.
Pie chart
A pie chart divides a circle into sectors representing parts of a whole.
For category (i), the sector angle is:
[
\text{Sector angle}_i = 360^\circ \times \frac{f_i}{n}
]
where:
- (f_i) is the frequency of category (i)
- (n) is the total number of observations
Pie charts are easiest to interpret when there are only a few categories with clearly different proportions. When precise comparison matters, a sorted bar chart is usually easier to read because people compare positions and lengths more accurately than angles and areas (Cleveland & McGill, 1984; Heer & Bostock, 2010).
Mosaic plot
A mosaic plot represents combinations of two or more categorical variables. The area of each rectangle corresponds to the frequency of a combination of categories.
It is useful for exploring association in contingency tables, although beginners may need explanatory labels.
Graphical Methods for Quantitative Distributions
Quantitative variables record numerical amounts such as age, income, temperature, test score, blood pressure, or response time.
Dot plot
A dot plot shows individual observations along a numerical axis. Repeated or nearby values are stacked or jittered.
It is especially useful for small and moderate samples because it retains the individual values.
A dot plot can reveal:
- Clusters
- Gaps
- Extreme values
- Group differences
- Sample size
With very large datasets, points may overlap and another method may be more suitable.
Stem-and-leaf plot
A stem-and-leaf plot divides each value into a stem and a leaf. For example, a score of 74 may be represented by stem 7 and leaf 4.
It preserves original values while showing the distribution. It is most useful for small datasets and teaching; it becomes unwieldy with large samples or values requiring many digits.
Histogram
A histogram divides a quantitative variable into intervals called bins. The number, proportion, or density of observations in each interval determines the height of its rectangle.
Histogram bars normally touch because the intervals form a continuous numerical scale.
The relative frequency for bin (i) is:
[
p_i = \frac{f_i}{n}
]
The percentage is:
[
100 \times \frac{f_i}{n}
]
When bins have unequal widths, frequency alone is misleading. The height should represent density:
[
\text{Density}_i = \frac{f_i}{n w_i}
]
where (w_i) is the width of bin (i). The area of the rectangle then represents the relative frequency.
The appearance of a histogram depends on bin width and bin boundaries. Researchers should inspect more than one reasonable binning choice rather than treating a single histogram as a unique representation of the data.
Density plot
A density plot provides a smoothed representation of a quantitative distribution.
It is useful for examining shape and comparing groups, but the result depends on the smoothing bandwidth. Too little smoothing creates noisy peaks; too much smoothing can hide genuine structure.
Density curves can also extend beyond logically possible values unless boundary-aware methods are used.
Boxplot
A boxplot summarizes a distribution using the median, quartiles, interquartile range, whiskers, and possible outliers.
The interquartile range is:
[
IQR = Q_3 – Q_1
]
A commonly used rule identifies potential outliers below:
[
Q_1 – 1.5(IQR)
]
or above:
[
Q_3 + 1.5(IQR)
]
These points are not automatically errors. They are observations that deserve examination.
Boxplot conventions vary. Some whiskers extend to the most extreme observation within the 1.5-IQR limits, while other software or publications may use different definitions. Researchers should understand the convention used by their software.
Boxplots are compact and useful for comparing groups, but they can conceal multimodality and individual observations. Adding jittered points can provide more detail.
Violin plot
A violin plot combines a boxplot-like summary with a mirrored density estimate.
It shows more distributional shape than a boxplot, but it may be difficult to interpret with very small samples. The smoothing method should not create an impression of precision that the data cannot support.
Empirical cumulative distribution function
An empirical cumulative distribution function, or ECDF, shows the proportion of observations less than or equal to each value.
Unlike a histogram, it does not require arbitrary bins. It is particularly useful for comparing distributions and reading percentiles.
Q–Q plot
A quantile–quantile plot compares the quantiles of observed data with the quantiles of a reference distribution or another dataset.
A normal Q–Q plot is often used to evaluate whether a variable or model residuals are approximately normally distributed.
Points close to a straight line indicate approximate agreement. Systematic curves may indicate skewness, heavy tails, light tails, or other departures.
A Q–Q plot is a diagnostic aid, not a mechanical pass-or-fail test. Sample size and the consequences of non-normality should also be considered.
Graphical Methods for Relationships
Scatterplot
A scatterplot places one quantitative variable on the horizontal axis and another on the vertical axis.
It can reveal:
- Positive or negative association
- Linear or nonlinear patterns
- Clusters
- Gaps
- Unequal spread
- Outliers
- Possible subgroups
A trend line may be added, but the fitting method should be identified.
A scatterplot cannot establish causation. An observed relationship may result from confounding, selection, measurement processes, reverse causation, or chance.
Bubble chart
A bubble chart extends a scatterplot by using point size to encode a third variable.
It can display several dimensions, but area is difficult to compare precisely. Bubble size should be proportional to area rather than radius, and excessive overlap should be avoided.
Hexbin and two-dimensional density plots
With very large datasets, scatterplot points may overlap so heavily that concentration becomes invisible. Hexagonal binning or two-dimensional density contours can show where observations are concentrated.
Heatmap
A heatmap uses colour intensity to represent values in a matrix.
Common applications include:
- Correlation matrices
- Gene-expression data
- Survey-response patterns
- Time-by-category data
- Missing-data patterns
The colour scale, ordering, range, and midpoint can strongly affect interpretation. Diverging colour scales are appropriate when values have a meaningful centre, such as zero.
Graphical Methods for Time-Series Data
Line graph
A line graph connects ordered observations, usually measured over time.
It is suitable for showing:
- Long-term trends
- Short-term changes
- Cycles
- Seasonal patterns
- Differences between time series
Connecting observations implies continuity or meaningful order. It is not appropriate for unrelated nominal categories.
When several series are included, direct labels or small multiples may be clearer than a crowded legend.
Run chart
A run chart displays observations in time order and may include a reference line such as the median.
It is used to detect shifts, trends, or unusual sequences.
Control chart
A control chart adds a centre line and statistically derived control limits. It is commonly used in quality improvement to distinguish common process variation from signals that may require investigation.
Control limits are not the same as confidence intervals or specification limits.
Seasonal plot and small multiples
Seasonal plots compare recurring periods, such as months across multiple years. Small multiples display a consistent chart for each group, region, treatment, or period.
Small multiples are often clearer than placing many lines in one panel.
Graphical Methods for Multivariate Data
Multivariate graphics display or summarize more than two variables.
Scatterplot matrix
A scatterplot matrix presents pairwise scatterplots for several quantitative variables. Histograms or density plots may appear along the diagonal.
It can reveal correlations, nonlinear relationships, clusters, and unusual observations, but becomes crowded when too many variables are included.
Parallel-coordinates plot
Each variable is represented by a parallel axis, and each observation becomes a line crossing the axes.
The method can reveal multivariate profiles and clusters, but results depend heavily on variable scaling and axis order.
Correlation heatmap
A correlation heatmap displays pairwise correlation coefficients using colour.
It should not be interpreted as a causal map. The type of correlation, treatment of missing values, and sample size should be stated.
Biplot
A biplot commonly displays observations and variable directions in a reduced-dimensional space, often produced through principal component analysis.
Correct interpretation depends on the scaling convention and the proportion of variation represented by the displayed components.
Faceting
Faceting divides data into repeated panels based on one or more grouping variables. Each panel uses the same graphical form.
Faceting can make subgroup patterns visible without overloading a single chart.
Graphical Methods for Spatial Data
Choropleth map
A choropleth map shades geographic regions according to a value.
It is usually appropriate for rates, percentages, or standardized measures. Raw counts may largely reflect population size and can be misleading when regions differ substantially in population.
Proportional-symbol map
Symbols are placed at geographic locations, and symbol size represents a value.
It can display counts more appropriately than a choropleth map but may suffer from overlap.
Important spatial cautions
Maps can be affected by:
- Unequal geographic area
- Choice of classification intervals
- Boundary definitions
- Missing regions
- Spatial aggregation
- Inappropriate colour scales
A visually dominant large region does not necessarily contain the largest population or the most statistically important result.
Graphical Methods for Model Diagnostics
Graphical methods are also used after a statistical model has been fitted.
Residuals versus fitted values
This plot helps assess:
- Nonlinearity
- Unequal residual variance
- Unusual cases
- Missing model structure
A random cloud around zero is generally more reassuring than a systematic curve or funnel shape.
Residual Q–Q plot
A residual Q–Q plot compares model residuals with a theoretical reference distribution. Strong tail departures or curvature may indicate that the assumed error distribution is inadequate.
Scale-location plot
This plot examines whether residual spread changes across fitted values.
Leverage and influence plots
These displays help identify observations that have unusual predictor values or exert substantial influence on model estimates.
An influential observation should not be removed solely because a diagnostic identifies it. Researchers should investigate data quality, study context, model specification, and the effect of reasonable sensitivity analyses.
Observed-versus-predicted plot
This plot compares model predictions with observed outcomes. It can help evaluate calibration, systematic bias, and ranges in which the model performs poorly.
How to Choose the Right Graphical Method
Choose a graph by identifying the research question, variable types, number of variables, structure of the observations, analytical purpose, and audience.
Step 1: Define the question
Clarify what the graph must show:
- A comparison?
- A distribution?
- A relationship?
- A trend?
- A composition?
- Geographic variation?
- Model performance?
- Uncertainty?
A chart should answer a research question rather than merely display available columns.
Step 2: Identify the variable types
Determine whether each variable is:
- Nominal
- Ordinal
- Discrete quantitative
- Continuous quantitative
- Date or time
- Geographic
- Binary or count-based
Step 3: Count the variables
- One categorical variable: bar chart
- One quantitative variable: dot plot, histogram, boxplot, ECDF
- Two categorical variables: grouped bar chart or mosaic plot
- One categorical and one quantitative variable: side-by-side boxplots or dot plots
- Two quantitative variables: scatterplot
- Time and one quantitative variable: line graph
- Several quantitative variables: scatterplot matrix, heatmap, or biplot
Step 4: Examine the observation structure
Ask whether observations are:
- Independent
- Repeated on the same participant
- Nested within groups
- Spatially related
- Collected at uneven time intervals
Repeated observations should not be presented as though every point came from an independent participant.
Step 5: Decide whether individual values matter
For small samples, show individual observations where possible.
A bar representing only a group mean may conceal:
- Sample size
- Skewness
- Clusters
- Outliers
- Overlap between groups
Dot plots, strip plots, raincloud plots, or boxplots with points often provide more information.
Step 6: Decide how uncertainty will be shown
Possible methods include:
- Confidence intervals
- Standard-error bars
- Prediction intervals
- Credible intervals
- Bootstrap intervals
- Shaded uncertainty bands
State clearly what the interval represents.
A common interval form is:
[
\text{Estimate} \pm \text{Critical value} \times SE
]
The critical value depends on the statistical model and desired confidence level.
Step 7: Test the graph with its intended audience
Check whether readers can answer the intended question correctly.
A technically valid graph may still fail if its labels, colours, notation, or layout are too difficult for its audience.
Worked Research Example
Suppose a researcher records the weekly study time and final examination score of 120 university students. Students are also classified by programme and teaching format.
The dataset contains:
- Study hours: quantitative
- Examination score: quantitative
- Programme: categorical
- Teaching format: categorical
- Week: ordered time variable
Different questions require different graphs.
| Research question | Recommended graph | Reason |
|---|---|---|
| How are examination scores distributed? | Histogram, ECDF, or boxplot | Shows shape, spread, and unusual values |
| Do scores differ by programme? | Side-by-side boxplots with individual points | Compares group distributions |
| Are study hours associated with scores? | Scatterplot with a fitted line | Shows form, direction, and unusual cases |
| Do weekly study hours change during the semester? | Line graph | Displays change over ordered weeks |
| Does the study-hours relationship differ by teaching format? | Faceted scatterplots or different point encodings | Compares relationships across groups |
| How precise are the estimated programme means? | Dot-and-interval plot | Shows estimates and confidence intervals |
No single graph answers every question. The researcher should select each display according to the specific analytical objective.
How to Interpret a Statistical Graph
A systematic interpretation reduces the risk of focusing only on visually striking features.
1. Read the title and caption
Determine what population, variables, period, and analytical result the figure represents.
2. Examine the axes
Check:
- Variable names
- Units
- Scale type
- Axis range
- Whether an axis is logarithmic
- Whether the baseline has been truncated
3. Identify the visual encoding
Determine what position, length, colour, shape, line type, and size represent.
4. Assess the overall pattern
Look for:
- Direction
- Shape
- Centre
- Spread
- Clusters
- Peaks
- Seasonality
- Nonlinearity
5. Examine unusual observations
Identify points that are isolated, influential, impossible, or inconsistent with surrounding observations.
6. Consider uncertainty and sample size
An apparent difference may be unstable when groups are small or intervals are wide.
7. Consider alternative explanations
A graph may show association without proving that one variable causes another.
8. Compare the graph with numerical results
Check whether the visual interpretation agrees with descriptive statistics, interval estimates, model output, and sensitivity analyses.
Advantages of Graphical Methods
Rapid pattern recognition
Graphs can make distributions, trends, clusters, and relationships visible quickly.
Efficient comparison
Common axes make it easier to compare groups or periods.
Detection of unexpected features
Graphical analysis may reveal outliers, multimodality, nonlinear relationships, or data-quality problems.
Communication across audiences
A clear figure can make a complex result understandable to non-specialist readers.
Support for model development
Diagnostic graphics help researchers identify transformations, nonlinear terms, group differences, and assumption violations.
Complement to numerical analysis
Graphs provide context that may not be apparent from averages, coefficients, or p values alone.
Limitations of Graphical Methods
| Limitation | Why it matters |
|---|---|
| Interpretation can be subjective | Different readers may emphasize different visual features |
| Design choices influence appearance | Bins, smoothing, scales, colours, and aspect ratios can alter perception |
| Exact values may be difficult to recover | A table may be better when precise lookup is required |
| Large datasets can cause overplotting | Dense points may hide concentration or subgroups |
| Small samples can create unstable patterns | Apparent peaks or gaps may not represent the population |
| Complex graphs increase cognitive demand | A technically rich display may confuse its audience |
| Visual association may be mistaken for causation | Confounding and study design remain important |
| Accessibility may be poor | Colour-only encoding and image-only charts can exclude readers |
| Graphs can be manipulated | Truncated axes, selective periods, and omitted groups can exaggerate conclusions |
Graphical methods should not automatically replace tables, numerical summaries, or statistical inference.
Common Mistakes and Misleading Graphs
Using the wrong graph for the variable
A histogram should not be used for unordered categories, and a pie chart should not be used when categories overlap or do not form a meaningful whole.
Confusing a bar chart with a histogram
A bar chart displays separate categories and normally has gaps. A histogram displays adjacent numerical intervals and normally has touching bars.
Truncating a bar-chart axis
Because bar length represents magnitude, a non-zero baseline can greatly exaggerate differences.
A narrowed axis may sometimes be appropriate for a line graph when small changes are scientifically meaningful, but the range must be clearly labelled and not used to deceive.
Adding unnecessary three-dimensional effects
Perspective can distort bar lengths, angles, and areas without adding information.
Using colour as the only distinction
Readers with colour-vision differences, monochrome printouts, or low-quality screens may be unable to distinguish the series. Combine colour with direct labels, shapes, line styles, or patterns.
Overloading a graph
Too many categories, lines, labels, gridlines, or decorative elements compete with the data.
Hiding the sample size
A smooth violin or large bar may appear authoritative even when based on very few observations. Display sample sizes or individual points where appropriate.
Showing only means
Means alone can hide distributions, unequal variance, skewness, and outliers.
Using inappropriate dual axes
Two vertical scales can make unrelated series appear strongly associated. Separate panels, indexing, or direct standardization is often safer.
Omitting uncertainty
A difference between two estimates may look important even though both estimates are imprecise.
Treating outliers as errors
An outlier flag indicates unusualness, not invalidity. Removal requires a defensible methodological reason.
Choosing a misleading map denominator
Mapping raw disease counts may mainly show where more people live. Rates may be more meaningful, provided their denominators and stability are considered.
Graphical Methods in Modern Research
Modern graphical practice extends beyond placing a chart in the results section.
Data-quality assessment
Researchers use graphics to inspect:
- Missing-data patterns
- Impossible values
- Duplicate observations
- Measurement limits
- Coding inconsistencies
- Changes in data collection over time
Transparent exploratory analysis
Exploratory figures can document how researchers selected transformations, identified subgroups, checked assumptions, and refined models.
Exploratory findings should be distinguished from confirmatory analyses specified before examining the data.
Model interpretation
Modern research increasingly uses:
- Marginal-effect plots
- Predicted-probability plots
- Calibration plots
- Partial-dependence plots
- Individual conditional-expectation plots
- Posterior predictive checks
- Coefficient and interval plots
These displays help readers understand models that cannot be interpreted adequately from a coefficient table alone.
Interactive visualization
Interactive graphics may allow filtering, zooming, tooltips, and subgroup selection.
However, the default view should still communicate the main message. Important information should not be available only through hovering, because hover interactions may be inaccessible or unavailable in printed and shared versions.
Reproducible graphics
A reproducible workflow stores:
- The original or documented source data
- Data-cleaning code
- Graph-generation code
- Software and package versions
- Captions and source notes
- Export settings
Code-based tools reduce the risk that a chart cannot be recreated after data or analysis changes.
Digital Tools for Graphical Methods
Spreadsheet software
Microsoft Excel, Google Sheets, and similar applications are suitable for basic bar charts, line graphs, scatterplots, and quick exploratory displays.
Researchers should check automatic axis ranges, category ordering, aggregation, date handling, and default formatting.
Statistical software
SPSS, Stata, SAS, JMP, and Minitab combine graphics with statistical analysis. They are useful when plots must be linked directly to fitted models or formal procedures.
R
R includes base graphical functions and packages such as ggplot2. Grammar-based systems construct a visualization from data, variable mappings, geometric marks, statistical transformations, scales, coordinates, facets, and themes.
R is especially useful for reproducible and publication-quality statistical graphics.
Python
Python libraries include Matplotlib, pandas plotting, Plotly, Altair, and other statistical visualization tools.
Python is useful when visualization is part of a broader data-processing, machine-learning, or application workflow.
Notebook and publishing systems
Jupyter notebooks combine code, text, equations, output, and figures. Quarto and related systems can generate reproducible reports, articles, presentations, websites, and books from a common source.
Dashboard and business-intelligence tools
Tableau, Power BI, and similar tools support interactive dashboards and visual exploration. Researchers should avoid adding interaction merely because it is available. Each control should serve a genuine user need.
Artificial Intelligence and Graphical Methods
Artificial intelligence can assist researchers with:
- Recommending possible chart types
- Generating R, Python, or JavaScript code
- Transforming data into a chart-ready structure
- Suggesting labels or captions
- Producing alternative text
- Iterating between chart designs
- Explaining unfamiliar visualizations
Research on generative AI for visualization describes applications in data enhancement, visual mapping, styling, and interaction, while also identifying unresolved evaluation and reliability problems (Ye et al., 2024). AI-assisted authoring systems can reduce data-transformation barriers, but users still need to inspect every transformation and visual encoding (Wang et al., 2024).
Risks of AI-generated graphs
AI may:
- Select an inappropriate graph
- Misinterpret variable types
- Aggregate the wrong observations
- Invent labels or units
- Apply an unjustified transformation
- Omit missing values
- Produce code that runs but answers the wrong question
- Describe a chart inaccurately
- Overstate an apparent relationship
Recent research on automatically generated chart descriptions reports that language models can still make factual errors, particularly when they infer chart content indirectly rather than using underlying data (Nylund et al., 2025).
AI-verification checklist
Before using an AI-assisted figure:
- Compare the displayed values with the source data.
- Inspect all filtering, grouping, and aggregation.
- Verify units, labels, and category order.
- Check the axis scale and transformation.
- Confirm what error bars represent.
- Recalculate important values independently.
- Review the figure for misleading visual effects.
- Rewrite the caption in the researcher’s own words.
- Follow institutional rules for disclosing AI use.
- Retain the code and prompts needed to document the process.
AI can assist visualization, but responsibility for accuracy remains with the researcher.
Accessibility and Ethical Presentation
A research graph should remain understandable to as many readers as possible.
Use readable text
Titles, labels, legends, and annotations should remain legible at the final publication size.
Do not rely on colour alone
Use combinations of:
- Direct labels
- Shapes
- Line styles
- Patterns
- Borders
- Position
Use appropriate contrast
Text, lines, markers, and adjacent colour regions should be distinguishable.
Provide a text alternative
A short description should identify the graph and its main point. A more detailed description or accessible data table may be needed when the graph contains information that is not explained in the surrounding text.
Example short description:
Scatterplot showing a positive, nonlinear relationship between weekly study hours and examination score among 120 students.
A longer description can explain the axes, overall pattern, important groups, notable observations, and uncertainty.
Include the underlying values when practical
An accessible HTML table or downloadable dataset allows readers to obtain exact values and conduct further analysis.
Show uncertainty honestly
Confidence intervals, prediction intervals, credible intervals, missing values, and suppressed data should not be concealed when they materially affect interpretation.
Protect confidentiality
A graph can disclose sensitive information even when names have been removed. Rare categories, exact locations, small counts, trajectories, and linked attributes may allow individuals to be identified.
Aggregate, suppress, perturb, or restrict information when necessary, while explaining the method used.
How to Report a Graph in a Research Paper
A research figure should be understandable without requiring readers to reconstruct its meaning from several pages of text.
Include a figure number
Number figures in the order in which they are discussed.
Write an informative title
The title should identify the subject of the figure. A message-based title may be appropriate when the publication permits it.
Label the axes and units
Write “Blood pressure (mmHg)” rather than only “Blood pressure.”
Explain symbols and intervals
Define colours, shapes, line types, abbreviations, reference lines, and uncertainty intervals.
State the analytical population
Explain exclusions, subgroup definitions, and sample sizes when they are not otherwise obvious.
Identify transformations
State whether an axis or variable uses a logarithm, standardization, smoothing method, or another transformation.
Cite the data source
When a figure uses external data, provide a source citation. When it is adapted from another publication, follow copyright and attribution requirements.
Refer to the figure in the text
Do not insert a graph without discussing its relevance.
A useful description moves from the main result to supporting detail:
Figure 2 shows a positive but nonlinear association between weekly study time and examination score. Scores increased rapidly between 0 and 10 weekly study hours, after which the relationship became weaker. Two observations had unusually low scores relative to their reported study time.
Do not repeat every number
The text should explain the important pattern rather than reproduce all values visible in the figure.
Conclusion
Graphical methods allow researchers to explore distributions, compare groups, study relationships, evaluate change, diagnose models, and communicate findings. Their value depends on selecting a display that matches the research question and data structure.
A trustworthy graph uses accurate data, appropriate scales, clear labels, visible uncertainty, accessible design, and transparent analytical decisions. It supports numerical and inferential analysis rather than replacing them.
References
- Cleveland, W. S., & McGill, R. (1984). Graphical perception: Theory, experimentation, and application to the development of graphical methods. Journal of the American Statistical Association, 79(387), 531–554. https://doi.org/10.1080/01621459.1984.10478080
- Heer, J., & Bostock, M. (2010). Crowdsourcing graphical perception: Using Mechanical Turk to assess visualization design. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (pp. 203–212). Association for Computing Machinery. https://doi.org/10.1145/1753326.1753357
- National Institute of Standards and Technology. (n.d.). Exploratory data analysis. NIST/SEMATECH e-Handbook of Statistical Methods.
- Nylund, K., Mankoff, J., & Potluri, V. (2025). MatplotAlt: A Python library for adding alt text to Matplotlib figures in computational notebooks. arXiv. https://arxiv.org/abs/2503.20089
- Office for National Statistics. (n.d.). Data visualisation guidance. ONS Service Manual.
- Tukey, J. W. (1977). Exploratory data analysis. Addison-Wesley.
- VanderPlas, S., Cook, D., & Hofmann, H. (2020). Testing statistical charts: What makes a good graph? Annual Review of Statistics and Its Application, 7, 61–88. https://doi.org/10.1146/annurev-statistics-031219-041252
- Wang, C., Lee, B., Drucker, S., Marshall, D., & Gao, J. (2024). Data Formulator 2: Iteratively creating rich visualizations with AI. arXiv. https://arxiv.org/abs/2408.16119
- Wilkinson, L. (2005). The grammar of graphics (2nd ed.). Springer.
- World Wide Web Consortium. (2023). Web Content Accessibility Guidelines (WCAG) 2.2. W3C Recommendation.
- Ye, Y., Hao, J., Hou, Y., Wang, Z., Xiao, S., Luo, Y., & Zeng, W. (2024). Generative AI for visualization: State of the art and future directions. arXiv. https://arxiv.org/abs/2404.18144
