Data Types

Research Data – Definition, Types, Examples and Best Practices

Table of Contents

Research Data

Research data is the recorded evidence collected, observed, generated, or reused to answer a research question and support or validate findings. It may include numbers, text, images, audio, code, models, laboratory notes, specimens, or archival materials. Research data can be qualitative or quantitative, digital or physical, raw or processed.

Introduction

Every research conclusion depends on evidence. That evidence may be a spreadsheet of survey responses, a set of interview transcripts, readings produced by laboratory equipment, computer code used in a simulation, photographs of historical documents, or notes made during fieldwork. Collectively, these materials are known as research data.

Research data is not limited to numbers or computer files. Its form depends on the discipline, research question, methods, and rules governing the project. A historian, chemist, sociologist, engineer, linguist, and computer scientist may all work with very different data.

This article explains what research data is, how it is classified, what makes it trustworthy, and how researchers should collect, document, store, protect, analyse, share, cite, and preserve it. It also discusses research data management, FAIR and CARE principles, repositories, data-sharing restrictions, and responsible use of artificial intelligence.

Key takeaways

  • Research data is the evidence used to answer a research question or validate a conclusion.
  • It can be qualitative or quantitative, digital or physical, primary or secondary, and raw or processed.
  • A dataset normally includes data files plus the documentation needed to understand and use them.
  • Good data management begins before collection and continues through preservation, sharing, reuse, or secure disposal.
  • FAIR data is not automatically open data; legitimate ethical, legal, cultural, and commercial restrictions may apply.
  • AI can assist with research data, but researchers remain responsible for confidentiality, accuracy, documentation, and validation.

What Is Research Data?

Research data is any recorded material used as evidence in the research process. It is collected, observed, created, generated, derived, or reused to answer research questions, test hypotheses, develop interpretations, or support research findings.

The word “recorded” is important. A research observation becomes usable data when it is captured in a form that can be examined, analysed, verified, or interpreted. This record might be a number in a spreadsheet, an interview recording, a photograph, a laboratory notebook entry, a computer log, a coded text passage, or metadata describing a physical specimen.

Definitions vary across organisations. Many universities use a broad definition that includes digital files, paper records, physical samples, software, models, and creative works. Some legal or funder definitions are narrower and may exclude physical objects while treating measurements, descriptions, photographs, and records associated with those objects as data.

Researchers should therefore check the definition used by their:

  • Institution
  • Funder
  • Ethics committee
  • Research partner
  • Publisher
  • Repository
  • Applicable law or regulation

What Counts as Research Data?

A material normally counts as research data when it has a meaningful evidential role in the study.

Common examples include:

  • Survey responses
  • Interview and focus-group transcripts
  • Audio and video recordings
  • Field notes and observation records
  • Laboratory measurements
  • Sensor readings
  • Images and scans
  • Clinical and administrative records
  • Statistical data files
  • Computer code and analysis scripts
  • Simulation inputs and outputs
  • Algorithms and trained models
  • Geographic and remote-sensing data
  • Genetic or protein sequences
  • Laboratory and field notebooks
  • Archival photographs and document transcriptions
  • Text corpora
  • Annotated bibliographies
  • Creative works analysed or produced through practice-based research
  • Documentation, codebooks, data dictionaries, and README files

A publication used only as background reading is not normally research data. However, publications can become research data when they are systematically collected, coded, compared, annotated, or analysed. For example, 500 news articles used in a content analysis constitute research data, while five articles cited only to support an argument usually do not.

Similarly, computer code may be research data, a research method, or a research output. Its role depends on whether it is needed to generate, process, analyse, reproduce, or interpret the findings.

Research Data, Dataset, Information and Evidence

These terms are related but not identical.

TermMeaningExample
Research dataRecorded materials used as evidence in researchInterview recordings, temperature readings or source code
DatasetA logically organised collection of related data, usually accompanied by documentationA CSV file of survey responses with a codebook and README
InformationData that has been interpreted or placed in contextThe finding that satisfaction increased after an intervention
EvidenceData or information used to support or challenge a claimSurvey results used to support a conclusion about student satisfaction
Research recordWider documentation of how the study was planned and conductedEthics approval, consent forms, protocols, correspondence and audit logs
Research outputA product produced by the researchArticle, dataset, software package, model, report or exhibition

A dataset is therefore more than an isolated file. A usable dataset commonly includes:

  • The data files
  • A description of the study
  • Variable definitions
  • Units and coding rules
  • Collection methods
  • Processing and cleaning decisions
  • File relationships
  • Software or code requirements
  • Rights, licences and access conditions

Without this context, a file may exist but remain difficult to interpret or reuse.

Main Types of Research Data

Research data can be classified in several ways. These classifications overlap: one dataset may be primary, quantitative, observational, structured, sensitive, digital, and processed at the same time.

Types by source

Primary data

Primary data is collected or generated directly for the current research project.

Examples include:

  • Responses to a questionnaire designed by the researcher
  • Interviews conducted for a dissertation
  • Measurements from a laboratory experiment
  • Photographs taken during fieldwork
  • Data generated by a new simulation

Primary data gives the researcher control over the design and collection process. However, collecting it can require substantial time, money, equipment, permissions, and participant recruitment.

Secondary data

Secondary data already exists and is reused for a new research purpose.

Examples include:

  • Government statistics
  • Census microdata
  • Public health records
  • Archived interviews
  • Satellite imagery
  • Published datasets
  • Historical documents
  • Corporate or administrative databases

Secondary-data research can be efficient and may enable analysis of large populations or long periods. Its limitations include incomplete documentation, restricted variables, uncertain quality, incompatible definitions, licensing conditions, and a lack of control over the original collection process.

Types by research approach

Quantitative data

Quantitative data represents quantities, measurements, counts, scores, or coded categories that can be analysed statistically.

Examples include:

  • Age
  • Income
  • Test score
  • Temperature
  • Number of website visits
  • Blood-pressure reading
  • Likert-scale response

Quantitative data is not necessarily continuous. It can be categorical, binary, ordinal, discrete, or continuous.

Qualitative data

Qualitative data captures meanings, experiences, descriptions, behaviours, language, images, or social processes.

Examples include:

  • Interview transcripts
  • Focus-group discussions
  • Open-ended survey answers
  • Observation notes
  • Photographs
  • Diaries
  • Policy documents
  • Social-media posts collected for analysis

Qualitative research data may be coded or converted into numerical summaries, but its value often lies in its context, wording, interpretation, and relationship to the research setting.

Mixed-methods data

Mixed-methods studies integrate quantitative and qualitative data.

A study might combine:

  • A numerical survey
  • Follow-up interviews
  • Classroom observations
  • Administrative attendance records

The researcher must explain how the different data sources relate to each other and how they are integrated during analysis.

Types by collection or generation method

TypeDescriptionExamplesImportant consideration
ObservationalCaptured without deliberately manipulating the main phenomenonField observations, sensor data, interviews, astronomical imagesMay be impossible to reproduce
ExperimentalGenerated through controlled interventions or laboratory proceduresClinical trial results, chemical measurements, test responsesReproduction may be possible but expensive or unethical
SimulationGenerated by computational or mathematical modelsClimate projections, economic models, engineering simulationsCode, parameters and input data may be essential
Derived or compiledProduced by transforming, combining or extracting existing sourcesText-mining corpus, merged database, calculated indicatorsProvenance and transformation steps must be documented
Reference dataCurated collections used repeatedly across studiesGene databases, geographic reference data, taxonomiesVersion and access date may affect reproducibility

Types by processing stage

Raw data

Raw data is the original recorded data before substantive cleaning, transformation, coding, or analysis.

Examples include:

  • Original audio recordings
  • Instrument output
  • Unedited photographs
  • Initial questionnaire exports
  • Original field notes

Researchers should normally preserve an unchanged master copy of raw data. Corrections and transformations should be performed on working copies or through reproducible scripts.

Processed data

Processed data has been prepared for analysis. Processing may include:

  • Transcription
  • Translation
  • Digitisation
  • Data cleaning
  • Recoding
  • Format conversion
  • Anonymisation
  • Removal of duplicate records
  • Correction of documented errors

Derived data

Derived data is calculated or generated from other data.

Examples include:

  • Body mass index calculated from height and weight
  • A sentiment score derived from text
  • A composite scale calculated from questionnaire items
  • A geocoded location derived from an address

Analysed data

Analysed data includes results produced through statistical, computational, qualitative, or visual analysis.

Examples include:

  • Regression output
  • Coded qualitative themes
  • Tables and graphs
  • Trained model results
  • Network measures
  • Summary statistics

A chart is not a substitute for the underlying data and analytical documentation.

Types by structure

Structured data

Structured data follows a defined format, often rows and columns.

Examples include relational databases, spreadsheets and CSV files.

Semi-structured data

Semi-structured data contains labels or organisational markers without a rigid table structure.

Examples include JSON, XML, HTML and application logs.

Unstructured data

Unstructured data does not naturally fit into a conventional table.

Examples include interview recordings, free text, photographs, videos and scanned documents.

Unstructured data still requires structure at the management level. File names, folders, metadata, identifiers and documentation make it discoverable and usable.

Types by sensitivity and access

Research data may be:

  • Open: available publicly under stated terms.
  • Embargoed: temporarily closed before later release.
  • Restricted: accessible only to approved users or for approved purposes.
  • Confidential: protected because disclosure could harm participants, organisations, communities, commercial interests, or national security.
  • Highly sensitive: subject to strong technical, contractual, ethical, or legal controls.

“Publicly available online” does not automatically mean “free to collect, analyse, republish, or share.” Terms of service, copyright, privacy expectations, research ethics, and community norms may still apply.

Examples of Research Data by Discipline

Social sciences

  • Survey responses
  • Interview transcripts
  • Focus-group recordings
  • Observation notes
  • Census data
  • Social-network data
  • Policy documents
  • Coded media content

Education

  • Test results
  • Attendance records
  • Classroom observations
  • Student work
  • Teacher interviews
  • Learning-management-system logs
  • Assessment rubrics
  • School survey responses

Health and medicine

  • Clinical measurements
  • Medical images
  • Laboratory results
  • Genomic data
  • Patient-reported outcomes
  • Trial records
  • Electronic health records
  • Adverse-event reports

Health data often requires ethics approval, secure environments, access controls, specialised de-identification, and careful consent procedures.

Natural sciences

  • Experimental measurements
  • Field observations
  • Specimen records
  • Microscopy images
  • Chemical spectra
  • Sensor readings
  • DNA sequences
  • Environmental samples

Engineering and computer science

  • Source code
  • Algorithms
  • Hardware measurements
  • Software logs
  • Benchmark results
  • Simulation files
  • Model parameters
  • Training, validation and test datasets
  • System-performance data

Humanities

  • Archival documents
  • Manuscripts
  • Transcriptions
  • Annotated texts
  • Bibliographic databases
  • Oral-history recordings
  • Photographs
  • Maps
  • Text corpora
  • Digital editions

A historian’s annotations and transcription decisions may be as important for interpretation as the scanned source itself.

Creative and practice-based research

  • Sketchbooks
  • Rehearsal recordings
  • Design iterations
  • Musical scores
  • Performance videos
  • Prototypes
  • Reflective journals
  • Material samples
  • Documentation of the creative process

Characteristics of High-Quality Research Data

High-quality data is fit for the purpose for which it will be used. Quality is therefore connected to the research question, method, population, instrument, and intended analysis.

Accuracy

Values should represent the phenomenon as correctly as reasonably possible. Calibration, validation rules, double-entry checks, and source verification can improve accuracy.

Completeness

Required observations, variables, files, and documentation should be present. Missingness should be identified and explained rather than silently ignored.

Consistency

Names, codes, units, dates, categories, and formats should be used consistently across files and collection periods.

For example, a dataset should not use “F,” “Female,” “2,” and “woman” for the same category without an explicit coding system.

Validity

Data should measure or represent what the study claims to examine. A precisely recorded value can still be invalid if the instrument or operational definition is unsuitable.

Reliability

A measurement or coding procedure should produce sufficiently consistent results under appropriate conditions. Relevant checks may include instrument reliability, repeated measurements, inter-rater agreement, or reproducible code.

Integrity

Data should remain complete and unaltered except through authorised, documented changes. Checksums, permissions, audit logs, version control and read-only master files support integrity.

Provenance

Provenance explains where the data came from and what happened to it.

It may include:

  • Original source
  • Collection date
  • Instrument or software
  • Researcher or system responsible
  • Processing steps
  • Code version
  • Exclusions and corrections
  • File relationships
  • Ownership and licence

Timeliness

Data should be sufficiently current for the research purpose. Historical data may be entirely appropriate for a historical question but unsuitable for estimating a present condition.

Accessibility and usability

Authorised users should be able to locate, open, interpret, and analyse the data. Accessibility requires documentation and suitable formats, not simply possession of a file.

The Research Data Lifecycle

The research data lifecycle is the sequence through which data moves from initial planning to collection, analysis, sharing, preservation, reuse, or disposal. Data management decisions should be made at every stage rather than postponed until publication.

1. Plan

Before collection, determine:

  • What data will be created or reused
  • File formats and likely volume
  • Collection instruments
  • Roles and responsibilities
  • Ethics and consent requirements
  • Storage and backup arrangements
  • Naming and versioning conventions
  • Quality-control procedures
  • Access restrictions
  • Preservation and sharing plans
  • Expected costs

2. Collect or generate

Use consistent procedures, tested instruments and documented protocols. Record contextual information while it is still known.

3. Process and clean

Transcribe, validate, correct, code, anonymise, convert, or combine the data. Preserve raw files and document every substantive transformation.

4. Analyse and interpret

Use appropriate statistical, computational, qualitative, visual, or mixed-methods procedures. Retain the scripts, coding decisions, parameters and software information needed to understand the analysis.

5. Document

Documentation is continuous rather than a final task. Maintain metadata, codebooks, README files, laboratory records, protocols, decision logs and data dictionaries.

6. Store and protect

Use institutionally approved systems, access controls, encryption where appropriate, backups, monitoring and recovery procedures.

7. Share or publish

Determine what can be shared, with whom, under what licence, at what time, and through which repository. Sensitive data may require mediated or controlled access.

8. Preserve, reuse or dispose

Retain valuable data and documentation in sustainable formats and repositories. When data must be destroyed, use an approved and documented disposal method.

What Is a Research Data Management Plan?

A data management plan is a structured explanation of how research data will be handled during and after a project. It normally covers data creation, documentation, storage, security, access, ethics, preservation, sharing, responsibilities and costs.

A useful DMP answers the following questions.

Data description

  • What data will be collected, created, derived, or reused?
  • What formats and approximate volumes are expected?
  • Which existing data sources will be used?

Methods and quality

  • How will data be collected or generated?
  • What instruments, software and protocols will be used?
  • How will accuracy, validity and consistency be checked?

Documentation and metadata

  • What metadata standard will be used?
  • Will the project create a README, codebook or data dictionary?
  • How will processing and analytical decisions be recorded?

Storage and security

  • Where will active data be stored?
  • How will it be backed up?
  • Who will have access?
  • Does the data require encryption or a secure computing environment?

Ethics and legal compliance

  • Does the project involve personal, confidential, Indigenous, commercial, copyrighted, or otherwise restricted data?
  • What does the consent process allow?
  • Are data-transfer or data-use agreements required?

Sharing and preservation

  • Which data can be deposited?
  • Which repository will be used?
  • Will an embargo or controlled-access process be required?
  • Which licence will apply?
  • How long should the data be retained?

Responsibilities and resources

  • Who is responsible for collection, documentation, quality, security, deposit and preservation?
  • What staff time, storage, repository or curation costs should be budgeted?

A DMP should be treated as a living document and updated when the project changes.

How to Collect and Create Reliable Research Data

1. Begin with the research question

Collect only data that has a clear relationship to the research objectives. Collecting unnecessary variables increases cost, complexity, privacy risk and analytical flexibility without necessarily improving the study.

2. Define variables and concepts

Create operational definitions before collection.

For each variable, record:

  • Name
  • Meaning
  • Unit
  • Type
  • Permitted values
  • Missing-value code
  • Source
  • Collection method
  • Calculation, if derived

3. Select suitable methods and instruments

The method should match the research question and population. Instruments may require validation, translation, adaptation, calibration or permission.

4. Pilot the procedure

A pilot can reveal confusing questions, technical problems, unrealistic completion times, inaccessible formats, inconsistent coding and missing response options.

5. Standardise collection

Use protocols, training, checklists, instrument settings and consistent instructions. Record unavoidable deviations.

6. Build quality controls into the process

Useful controls include:

  • Mandatory fields where ethically appropriate
  • Plausible value ranges
  • Logic and skip checks
  • Duplicate detection
  • Timestamp review
  • Calibration
  • Double coding
  • Inter-rater checks
  • Automated validation scripts
  • Manual review of exceptional cases

7. Maintain an audit trail

An audit trail records important changes and decisions. It should explain what changed, why, when, by whom, and how the change affects interpretation.

Organising Research Data

Use a logical folder structure

A simple structure might be:

project-name/
├── 01_admin/
├── 02_protocols/
├── 03_raw-data/
├── 04_processed-data/
├── 05_analysis/
├── 06_outputs/
├── 07_documentation/
└── README.txt

The exact structure matters less than consistency, documentation, access control and team agreement.

Use meaningful file names

A useful file name may include:

  • Project or dataset identifier
  • Data type
  • Date in YYYY-MM-DD format
  • Version
  • Status

Example:

student-survey_clean_2026-06-28_v02.csv

Avoid ambiguous names such as:

final.xlsx
final-new.xlsx
final-real-last.xlsx

Separate raw and working data

Keep original data unchanged whenever possible. Store cleaned and transformed versions separately. Use scripts rather than undocumented manual edits when the transformation can reasonably be automated.

Apply version control

Version control is useful for:

  • Code
  • Documentation
  • Statistical scripts
  • Plain-text data
  • Data dictionaries
  • Configuration files
  • Collaborative writing

Large, sensitive, or binary datasets may require a specialised data platform rather than a public software repository.

Prefer sustainable formats

When disciplinary standards allow, preservation-friendly formats can reduce dependence on particular software.

Examples include:

  • CSV or TSV for tables
  • TXT or UTF-8 text for plain text
  • JSON or XML for structured exchange
  • TIFF or PNG for images
  • WAV or FLAC for audio
  • PDF/A for appropriate document preservation

The original format may still need to be retained when conversion could remove information or functionality.

Metadata and Research Documentation

Metadata is information that describes the data and enables people or machines to find, understand, assess and reuse it.

Metadata can include:

Descriptive metadata

  • Dataset title
  • Creator
  • Keywords
  • Abstract
  • Geographic coverage
  • Time period
  • Persistent identifier

Structural metadata

  • Relationships between files
  • Table and variable connections
  • Ordering of image or audio sequences
  • Links between data and code

Administrative metadata

  • Owner
  • Licence
  • Access level
  • Retention period
  • Consent limitations
  • Repository information

Technical and provenance metadata

  • File format
  • Software version
  • Instrument
  • Calibration
  • Processing history
  • Code version
  • Creation and modification dates

What to include in a README file

A basic README should identify:

  1. Project title and purpose
  2. Creators and contact information
  3. Folder and file descriptions
  4. Collection or generation methods
  5. Variable or coding documentation
  6. Processing and cleaning steps
  7. Software and version requirements
  8. Known limitations
  9. Rights, licence and access restrictions
  10. Recommended citation

DataCite’s metadata framework is designed to support consistent identification, citation and discovery of research outputs (DataCite, 2026).

Storage, Backup and Security

Research data should be stored according to its sensitivity, value, size, format, collaboration needs, and legal or contractual restrictions.

Use approved storage

Institutional servers, managed research drives, secure cloud environments, controlled research platforms and approved laboratory systems are usually preferable to personal devices or unapproved consumer services.

Control access

Apply the least-privilege principle: each person should have only the access needed for their role.

Possible controls include:

  • Individual accounts
  • Strong authentication
  • Role-based permissions
  • Encryption
  • Secure transfer methods
  • Access logs
  • Time-limited access
  • Separate storage of identifying information and research responses

Maintain backups

A backup must be:

  • Separate from the active copy
  • Updated at an appropriate frequency
  • Protected from the same failure or security incident
  • Tested to confirm that restoration works

Cloud synchronisation alone may not constitute an adequate backup because accidental deletion or corruption can synchronise across devices.

Plan for incidents

The project should know how to respond to:

  • Lost equipment
  • Accidental disclosure
  • Unauthorised access
  • Malware or ransomware
  • Corrupted files
  • Incorrect deletion
  • Participant requests
  • Breach-reporting obligations

Ethics, Privacy and Legal Considerations

Good data management is not only a technical matter. It is also an ethical and governance responsibility.

Informed consent

Participants should receive clear information about:

  • What data will be collected
  • Why it is needed
  • How it will be used
  • Who may access it
  • How long it will be retained
  • Whether it may be shared or reused
  • Whether commercial or international partners are involved
  • Limits to withdrawal after anonymisation or publication

A vague statement that data “may be used for research” may not be adequate for every secondary use.

Personal and sensitive data

Personal data can identify someone directly or indirectly. Removing names does not necessarily make a dataset anonymous. Combinations of age, location, occupation, dates, rare characteristics, images or free-text responses may permit re-identification.

Risk controls may include:

  • Data minimisation
  • Pseudonymisation
  • De-identification
  • Aggregation
  • Removal or generalisation of indirect identifiers
  • Controlled access
  • Secure analysis environments
  • Data-use agreements
  • Disclosure review before publication

De-identification reduces risk but does not eliminate every possibility of re-identification (National Institute of Standards and Technology [NIST], 2015).

Confidential and commercially sensitive data

A project may be restricted by:

  • Confidentiality agreements
  • Intellectual-property rights
  • Trade secrets
  • Patent plans
  • Commercial contracts
  • Data-transfer agreements
  • National-security or export-control rules

Copyright and database rights

Researchers may own a compilation or original documentation without owning every item contained in the dataset. Permission to access material is not necessarily permission to reproduce, distribute or license it.

Record the source and permitted uses of all reused data.

Indigenous data governance

Open-data practices must not override the rights and interests of Indigenous Peoples. The CARE principles emphasise:

  • Collective Benefit
  • Authority to Control
  • Responsibility
  • Ethics

These principles complement FAIR by focusing on people, power, rights, relationships and community benefit rather than technical reusability alone (Carroll et al., 2020).

FAIR Research Data

The FAIR principles state that research data should be:

Findable

Data and metadata should have clear descriptions, persistent identifiers and searchable records.

Accessible

The conditions and procedures for accessing data should be clear. Accessibility may involve public download, registration, approval, or controlled access.

Interoperable

Data and metadata should use suitable formats, vocabularies, standards and identifiers so they can work with other systems and datasets.

Reusable

Documentation, provenance, licences and quality information should allow others to assess whether the data is suitable for a new purpose.

FAIR does not mean that every dataset must be openly downloadable. Restricted data can still be FAIR when its metadata is findable and the conditions for legitimate access are clearly described (Wilkinson et al., 2016).

A useful principle is:

Make data as open as possible and as restricted as necessary.

Sharing and Publishing Research Data

Data sharing can support verification, secondary research, collaboration, education, transparency and recognition. However, sharing must be compatible with consent, ethics, law, culture, intellectual property, contracts and participant welfare.

Choose an appropriate repository

A repository may be:

  • Discipline-specific
  • Institutional
  • General-purpose
  • Funder-operated
  • Publisher-associated
  • Controlled-access

A suitable repository should be evaluated for:

  1. Disciplinary recognition
  2. Persistent identifiers
  3. Metadata quality
  4. Access controls
  5. Licensing options
  6. Versioning
  7. Preservation commitments
  8. File-size and format support
  9. Cost
  10. Withdrawal and correction procedures

Where possible, use an established repository rather than placing the only copy on a personal website. CoreTrustSeal identifies requirements associated with trustworthy data repositories (CoreTrustSeal, 2026).

Assign a licence

A licence tells users what they may do with the data. The choice depends on ownership, consent, third-party material, funder requirements, and the intended forms of reuse.

Do not apply an open licence to material you do not have authority to license.

Use persistent identifiers

A DOI or another persistent identifier makes a dataset easier to discover, link, cite and track.

Cite datasets

A dataset citation should normally include:

  • Creator
  • Year
  • Dataset title
  • Version
  • Resource type
  • Repository or publisher
  • Persistent identifier

A generic APA-style pattern is:

Author, A. A. (Year). Title of dataset (Version) [Data set]. Repository. DOI

Researchers should cite reused datasets in the reference list rather than mentioning only the website from which a file was downloaded.

Write a data availability statement

An open-data statement might say:

The de-identified dataset and analysis code supporting this study are available in [repository] at [persistent identifier].

A restricted-data statement might say:

The data contain information that could compromise participant confidentiality. Qualified researchers may request controlled access subject to ethics approval and a data-use agreement.

A no-new-data statement might say:

No new data were created or analysed in this study.

The statement must accurately reflect the actual access conditions.

Research Data in Modern Research

Research data now frequently includes materials beyond traditional spreadsheets and laboratory measurements.

Modern examples include:

  • High-frequency sensor streams
  • Large-scale administrative records
  • Digital trace data
  • Social-media content
  • Software containers
  • Computational notebooks
  • Machine-learning training data
  • Model weights and parameters
  • Synthetic data
  • Geospatial data
  • Three-dimensional scans
  • Virtual-environment interactions
  • Research software
  • Workflow files
  • Prompt and model-configuration records

Open-science frameworks increasingly treat data, software, methods, infrastructure and publications as connected research objects. UNESCO’s Recommendation on Open Science encourages scientific knowledge, including data and software, to be made accessible and reusable when appropriate (UNESCO, 2021).

Funder requirements also continue to develop. As reviewed on June 28, 2026, NIH maintains its Data Management and Sharing Policy, while NSF requires relevant proposals to include data management and sharing plans through its current Research.gov process (National Institutes of Health [NIH], 2026; National Science Foundation [NSF], 2026).

Researchers should always check the current instructions for their exact call because requirements vary by funder, programme, discipline and award date.

Digital Tools for Research Data

The right tool depends on the project, institution, sensitivity and discipline.

Collection and generation

Examples include:

  • REDCap
  • Qualtrics
  • KoboToolbox
  • Laboratory information-management systems
  • Electronic laboratory notebooks
  • Sensor and instrument software
  • Secure transcription platforms

Organisation and collaboration

Examples include:

  • Institutionally managed cloud storage
  • Open Science Framework
  • Git and approved Git platforms
  • Electronic laboratory notebooks
  • Research information-management platforms

Analysis

Examples include:

  • R
  • Python
  • SPSS
  • Stata
  • SAS
  • MATLAB
  • NVivo
  • ATLAS.ti
  • Jupyter
  • R Markdown or Quarto

Data management planning

Examples include:

  • DMPTool
  • DMPonline
  • Funder-specific planning systems
  • Institutional DMP templates

Publication and preservation

Examples include:

  • Discipline-specific repositories
  • Institutional repositories
  • Zenodo
  • Dryad
  • Figshare
  • ICPSR
  • Controlled-access health or genomic repositories

A named tool should not be adopted solely because it is popular. Confirm institutional approval, security, access, export, retention, licensing and long-term availability.

Artificial Intelligence and Research Data

AI can support research-data work, but it can also introduce confidentiality, bias, provenance and reproducibility risks.

Potential uses

AI may assist with:

  • Transcription
  • Translation
  • Data extraction
  • Document classification
  • Image recognition
  • Qualitative coding suggestions
  • Anomaly detection
  • Missing-data investigation
  • Code generation
  • Metadata drafting
  • Data-cleaning suggestions
  • Literature or dataset discovery
  • Synthetic-data generation

Important risks

Confidentiality

Uploading unpublished, personal, commercial, culturally sensitive or restricted data to an external AI service may disclose information to an unauthorised third party.

Do not upload protected data unless the tool, contract, institutional policy, ethics approval and security controls explicitly permit it.

Hallucination and fabrication

An AI system may invent values, categories, citations, explanations or patterns. AI-produced transformations must be checked against the source data.

Bias

Models may reproduce biases present in training data, prompts, labels or human decisions. Automated coding should not be treated as neutral.

Loss of provenance

A result may be difficult to reproduce when the model, version, system prompt, settings or service changes.

Hidden transformations

Some services alter files, compress images, normalise text or remove metadata. Researchers should confirm what happens to uploaded and downloaded content.

Responsible AI practice

When AI materially contributes to data processing or analysis, record:

  • Tool and provider
  • Model and version, where available
  • Access date
  • Task performed
  • Prompts or instructions
  • Parameters and settings
  • Input-data classification
  • Human review process
  • Corrections made
  • Known limitations
  • Effect on reproducibility

Researchers remain responsible for the integrity of their data and conclusions. Current AI-integrity guidance emphasises legal compliance, ethical concerns, the integrity of the research record, publication practices, and human critical judgement (UK Research Integrity Office, 2025).

Advantages of Good Research Data Management

Effective management can:

  • Reduce data loss
  • Improve accuracy and consistency
  • Make analysis more efficient
  • Support collaboration
  • Protect participants
  • Demonstrate compliance
  • Enable verification
  • Improve reproducibility
  • Facilitate reuse
  • Increase visibility and citation
  • Preserve valuable evidence
  • Reduce duplication of research effort
  • Clarify responsibilities within a team

Limitations and Practical Challenges

Research-data management also involves trade-offs.

Time and cost

Documentation, curation, secure storage, anonymisation and repository deposit require staff time and resources.

Disciplinary variation

A standard suitable for genomic sequences may be inappropriate for oral histories, archaeological objects or creative practice.

Privacy and re-identification risk

Highly detailed data can remain identifying even after direct identifiers are removed.

Loss of context

A dataset separated from its social, historical, laboratory or cultural setting may be misinterpreted.

Technical obsolescence

File formats, software, storage media and platforms change.

Unequal benefits

Data may be collected from communities that receive little benefit from its publication or commercial reuse.

Reproducibility limits

Sharing data and code improves transparency but does not automatically correct poor design, biased sampling, invalid measurements or inappropriate analysis.

Common Research-Data Mistakes

Avoid the following errors:

  1. Beginning collection without a data management plan.
  2. Keeping the only copy on one laptop or external drive.
  3. Editing the original raw file directly.
  4. Using unexplained variable names and codes.
  5. Failing to document exclusions and corrections.
  6. Collecting unnecessary personal information.
  7. Assuming that de-identification removes every privacy risk.
  8. Treating publicly visible online content as unrestricted.
  9. Mixing consent forms with analysis data.
  10. Sharing data that the consent process did not permit.
  11. Uploading confidential data to an unapproved AI service.
  12. Publishing a spreadsheet without a README or codebook.
  13. Depositing data without a licence or citation instructions.
  14. Assuming FAIR means completely open.
  15. Ignoring the rights and governance expectations of communities represented in the data.
  16. Depending on proprietary software without recording formats and versions.
  17. Describing AI-generated or simulated data as empirical observations.
  18. Deleting intermediate files needed to reproduce the analysis.
  19. Treating a data management plan as a one-time administrative form.
  20. Assuming that all funders and journals have the same requirements.

Research Data Management Checklist

Before collecting data:

  • Define the data needed to answer the research question.
  • Check funder, institutional, ethics and partner requirements.
  • Identify ownership, consent and access conditions.
  • Prepare a data management plan.
  • Select approved collection and storage systems.
  • Establish naming, folder, versioning and metadata conventions.
  • Define quality-control procedures.

During the project:

  • Preserve unchanged raw data.
  • Maintain secure backups.
  • Restrict access appropriately.
  • Update documentation continuously.
  • Record processing and analytical decisions.
  • Review missing, inconsistent and exceptional values.
  • Update the DMP when methods or risks change.

Before publication or project closure:

  • Confirm which files support the findings.
  • Review disclosure and re-identification risk.
  • Remove files that cannot lawfully or ethically be shared.
  • Prepare metadata, a README and a codebook.
  • Select a suitable repository and licence.
  • Obtain a persistent identifier.
  • Cite the dataset.
  • Write an accurate data availability statement.
  • Preserve restricted data in an approved environment.
  • Document retention or secure-disposal decisions.

Conclusion

Research data is the evidence on which research questions, analyses and conclusions depend. It includes far more than numerical spreadsheets: interviews, images, code, models, notes, documents, samples and creative materials may all function as data.

The value of research data depends not only on collection but also on quality, context, documentation, security, ethics and responsible stewardship. Planning these elements from the beginning makes data easier to understand, analyse, protect, verify, preserve and, where appropriate, share and reuse.

Frequently Asked Questions

1. What is research data in simple words?

Research data is the recorded evidence a researcher uses to answer a question or support a conclusion. It may consist of numbers, words, images, recordings, observations, code, documents, models or physical-material records.

2. What are the main types of research data?

Common classifications include primary and secondary data; qualitative and quantitative data; observational, experimental, simulation and derived data; and raw, processed and analysed data. A dataset can belong to several classifications simultaneously.

3. Is research data always numerical?

No. Interview transcripts, field notes, photographs, video, audio, documents, maps, code, artwork and archival materials can all be research data.

4. What is the difference between research data and a dataset?

Research data refers broadly to evidence used in research. A dataset is an organised collection of related research data, normally accompanied by enough metadata and documentation to explain its structure, origin and use.

5. Are journal articles research data?

Articles used only as background literature are not normally research data. They become data when researchers systematically collect, code, annotate, compare, mine or otherwise analyse them as evidence.

6. What is raw research data?

Raw data is the original recorded material before substantive cleaning, recoding, transformation or analysis. Researchers should preserve an unchanged master copy whenever possible.

7. What is research data management?

Research data management is the planning and control of data throughout its lifecycle, including collection, organisation, documentation, storage, protection, analysis, sharing, preservation and disposal.

8. Does FAIR data have to be publicly open?

No. FAIR means findable, accessible, interoperable and reusable. Accessibility can involve controlled procedures, and sensitive data may remain restricted while its metadata and access conditions are made discoverable.

9. Can a researcher upload data to an AI tool?

Only when the tool and intended use comply with institutional policy, consent, ethics approval, contracts, privacy requirements and security rules. Confidential or unpublished data should not be uploaded to an unapproved public AI service.

10. How should a research dataset be cited?

Include the creator, year, dataset title, version, resource type, repository and persistent identifier. Follow the required citation style and the repository’s recommended citation.

About the author

Muhammad Hassan

Muhammad Hassan writes about research design, academic methods and data-analysis concepts for ResearchMethod.net. His work focuses on presenting methodological topics in clear language for students and early-career researchers. Articles are developed from recognized methodological literature and official software documentation.