
Research data is the recorded evidence collected, observed, generated, or reused to answer a research question and support or validate findings. It may include numbers, text, images, audio, code, models, laboratory notes, specimens, or archival materials. Research data can be qualitative or quantitative, digital or physical, raw or processed.
Introduction
Every research conclusion depends on evidence. That evidence may be a spreadsheet of survey responses, a set of interview transcripts, readings produced by laboratory equipment, computer code used in a simulation, photographs of historical documents, or notes made during fieldwork. Collectively, these materials are known as research data.
Research data is not limited to numbers or computer files. Its form depends on the discipline, research question, methods, and rules governing the project. A historian, chemist, sociologist, engineer, linguist, and computer scientist may all work with very different data.
This article explains what research data is, how it is classified, what makes it trustworthy, and how researchers should collect, document, store, protect, analyse, share, cite, and preserve it. It also discusses research data management, FAIR and CARE principles, repositories, data-sharing restrictions, and responsible use of artificial intelligence.
Key takeaways
- Research data is the evidence used to answer a research question or validate a conclusion.
- It can be qualitative or quantitative, digital or physical, primary or secondary, and raw or processed.
- A dataset normally includes data files plus the documentation needed to understand and use them.
- Good data management begins before collection and continues through preservation, sharing, reuse, or secure disposal.
- FAIR data is not automatically open data; legitimate ethical, legal, cultural, and commercial restrictions may apply.
- AI can assist with research data, but researchers remain responsible for confidentiality, accuracy, documentation, and validation.
What Is Research Data?
Research data is any recorded material used as evidence in the research process. It is collected, observed, created, generated, derived, or reused to answer research questions, test hypotheses, develop interpretations, or support research findings.
The word “recorded” is important. A research observation becomes usable data when it is captured in a form that can be examined, analysed, verified, or interpreted. This record might be a number in a spreadsheet, an interview recording, a photograph, a laboratory notebook entry, a computer log, a coded text passage, or metadata describing a physical specimen.
Definitions vary across organisations. Many universities use a broad definition that includes digital files, paper records, physical samples, software, models, and creative works. Some legal or funder definitions are narrower and may exclude physical objects while treating measurements, descriptions, photographs, and records associated with those objects as data.
Researchers should therefore check the definition used by their:
- Institution
- Funder
- Ethics committee
- Research partner
- Publisher
- Repository
- Applicable law or regulation
What Counts as Research Data?
A material normally counts as research data when it has a meaningful evidential role in the study.
Common examples include:
- Survey responses
- Interview and focus-group transcripts
- Audio and video recordings
- Field notes and observation records
- Laboratory measurements
- Sensor readings
- Images and scans
- Clinical and administrative records
- Statistical data files
- Computer code and analysis scripts
- Simulation inputs and outputs
- Algorithms and trained models
- Geographic and remote-sensing data
- Genetic or protein sequences
- Laboratory and field notebooks
- Archival photographs and document transcriptions
- Text corpora
- Annotated bibliographies
- Creative works analysed or produced through practice-based research
- Documentation, codebooks, data dictionaries, and README files
A publication used only as background reading is not normally research data. However, publications can become research data when they are systematically collected, coded, compared, annotated, or analysed. For example, 500 news articles used in a content analysis constitute research data, while five articles cited only to support an argument usually do not.
Similarly, computer code may be research data, a research method, or a research output. Its role depends on whether it is needed to generate, process, analyse, reproduce, or interpret the findings.
Research Data, Dataset, Information and Evidence
These terms are related but not identical.
| Term | Meaning | Example |
|---|---|---|
| Research data | Recorded materials used as evidence in research | Interview recordings, temperature readings or source code |
| Dataset | A logically organised collection of related data, usually accompanied by documentation | A CSV file of survey responses with a codebook and README |
| Information | Data that has been interpreted or placed in context | The finding that satisfaction increased after an intervention |
| Evidence | Data or information used to support or challenge a claim | Survey results used to support a conclusion about student satisfaction |
| Research record | Wider documentation of how the study was planned and conducted | Ethics approval, consent forms, protocols, correspondence and audit logs |
| Research output | A product produced by the research | Article, dataset, software package, model, report or exhibition |
A dataset is therefore more than an isolated file. A usable dataset commonly includes:
- The data files
- A description of the study
- Variable definitions
- Units and coding rules
- Collection methods
- Processing and cleaning decisions
- File relationships
- Software or code requirements
- Rights, licences and access conditions
Without this context, a file may exist but remain difficult to interpret or reuse.
Main Types of Research Data
Research data can be classified in several ways. These classifications overlap: one dataset may be primary, quantitative, observational, structured, sensitive, digital, and processed at the same time.
Types by source
Primary data
Primary data is collected or generated directly for the current research project.
Examples include:
- Responses to a questionnaire designed by the researcher
- Interviews conducted for a dissertation
- Measurements from a laboratory experiment
- Photographs taken during fieldwork
- Data generated by a new simulation
Primary data gives the researcher control over the design and collection process. However, collecting it can require substantial time, money, equipment, permissions, and participant recruitment.
Secondary data
Secondary data already exists and is reused for a new research purpose.
Examples include:
- Government statistics
- Census microdata
- Public health records
- Archived interviews
- Satellite imagery
- Published datasets
- Historical documents
- Corporate or administrative databases
Secondary-data research can be efficient and may enable analysis of large populations or long periods. Its limitations include incomplete documentation, restricted variables, uncertain quality, incompatible definitions, licensing conditions, and a lack of control over the original collection process.
Types by research approach
Quantitative data
Quantitative data represents quantities, measurements, counts, scores, or coded categories that can be analysed statistically.
Examples include:
- Age
- Income
- Test score
- Temperature
- Number of website visits
- Blood-pressure reading
- Likert-scale response
Quantitative data is not necessarily continuous. It can be categorical, binary, ordinal, discrete, or continuous.
Qualitative data
Qualitative data captures meanings, experiences, descriptions, behaviours, language, images, or social processes.
Examples include:
- Interview transcripts
- Focus-group discussions
- Open-ended survey answers
- Observation notes
- Photographs
- Diaries
- Policy documents
- Social-media posts collected for analysis
Qualitative research data may be coded or converted into numerical summaries, but its value often lies in its context, wording, interpretation, and relationship to the research setting.
Mixed-methods data
Mixed-methods studies integrate quantitative and qualitative data.
A study might combine:
- A numerical survey
- Follow-up interviews
- Classroom observations
- Administrative attendance records
The researcher must explain how the different data sources relate to each other and how they are integrated during analysis.
Types by collection or generation method
| Type | Description | Examples | Important consideration |
|---|---|---|---|
| Observational | Captured without deliberately manipulating the main phenomenon | Field observations, sensor data, interviews, astronomical images | May be impossible to reproduce |
| Experimental | Generated through controlled interventions or laboratory procedures | Clinical trial results, chemical measurements, test responses | Reproduction may be possible but expensive or unethical |
| Simulation | Generated by computational or mathematical models | Climate projections, economic models, engineering simulations | Code, parameters and input data may be essential |
| Derived or compiled | Produced by transforming, combining or extracting existing sources | Text-mining corpus, merged database, calculated indicators | Provenance and transformation steps must be documented |
| Reference data | Curated collections used repeatedly across studies | Gene databases, geographic reference data, taxonomies | Version and access date may affect reproducibility |
Types by processing stage
Raw data
Raw data is the original recorded data before substantive cleaning, transformation, coding, or analysis.
Examples include:
- Original audio recordings
- Instrument output
- Unedited photographs
- Initial questionnaire exports
- Original field notes
Researchers should normally preserve an unchanged master copy of raw data. Corrections and transformations should be performed on working copies or through reproducible scripts.
Processed data
Processed data has been prepared for analysis. Processing may include:
- Transcription
- Translation
- Digitisation
- Data cleaning
- Recoding
- Format conversion
- Anonymisation
- Removal of duplicate records
- Correction of documented errors
Derived data
Derived data is calculated or generated from other data.
Examples include:
- Body mass index calculated from height and weight
- A sentiment score derived from text
- A composite scale calculated from questionnaire items
- A geocoded location derived from an address
Analysed data
Analysed data includes results produced through statistical, computational, qualitative, or visual analysis.
Examples include:
- Regression output
- Coded qualitative themes
- Tables and graphs
- Trained model results
- Network measures
- Summary statistics
A chart is not a substitute for the underlying data and analytical documentation.
Types by structure
Structured data
Structured data follows a defined format, often rows and columns.
Examples include relational databases, spreadsheets and CSV files.
Semi-structured data
Semi-structured data contains labels or organisational markers without a rigid table structure.
Examples include JSON, XML, HTML and application logs.
Unstructured data
Unstructured data does not naturally fit into a conventional table.
Examples include interview recordings, free text, photographs, videos and scanned documents.
Unstructured data still requires structure at the management level. File names, folders, metadata, identifiers and documentation make it discoverable and usable.
Types by sensitivity and access
Research data may be:
- Open: available publicly under stated terms.
- Embargoed: temporarily closed before later release.
- Restricted: accessible only to approved users or for approved purposes.
- Confidential: protected because disclosure could harm participants, organisations, communities, commercial interests, or national security.
- Highly sensitive: subject to strong technical, contractual, ethical, or legal controls.
“Publicly available online” does not automatically mean “free to collect, analyse, republish, or share.” Terms of service, copyright, privacy expectations, research ethics, and community norms may still apply.
Examples of Research Data by Discipline
Social sciences
- Survey responses
- Interview transcripts
- Focus-group recordings
- Observation notes
- Census data
- Social-network data
- Policy documents
- Coded media content
Education
- Test results
- Attendance records
- Classroom observations
- Student work
- Teacher interviews
- Learning-management-system logs
- Assessment rubrics
- School survey responses
Health and medicine
- Clinical measurements
- Medical images
- Laboratory results
- Genomic data
- Patient-reported outcomes
- Trial records
- Electronic health records
- Adverse-event reports
Health data often requires ethics approval, secure environments, access controls, specialised de-identification, and careful consent procedures.
Natural sciences
- Experimental measurements
- Field observations
- Specimen records
- Microscopy images
- Chemical spectra
- Sensor readings
- DNA sequences
- Environmental samples
Engineering and computer science
- Source code
- Algorithms
- Hardware measurements
- Software logs
- Benchmark results
- Simulation files
- Model parameters
- Training, validation and test datasets
- System-performance data
Humanities
- Archival documents
- Manuscripts
- Transcriptions
- Annotated texts
- Bibliographic databases
- Oral-history recordings
- Photographs
- Maps
- Text corpora
- Digital editions
A historian’s annotations and transcription decisions may be as important for interpretation as the scanned source itself.
Creative and practice-based research
- Sketchbooks
- Rehearsal recordings
- Design iterations
- Musical scores
- Performance videos
- Prototypes
- Reflective journals
- Material samples
- Documentation of the creative process
Characteristics of High-Quality Research Data
High-quality data is fit for the purpose for which it will be used. Quality is therefore connected to the research question, method, population, instrument, and intended analysis.
Accuracy
Values should represent the phenomenon as correctly as reasonably possible. Calibration, validation rules, double-entry checks, and source verification can improve accuracy.
Completeness
Required observations, variables, files, and documentation should be present. Missingness should be identified and explained rather than silently ignored.
Consistency
Names, codes, units, dates, categories, and formats should be used consistently across files and collection periods.
For example, a dataset should not use “F,” “Female,” “2,” and “woman” for the same category without an explicit coding system.
Validity
Data should measure or represent what the study claims to examine. A precisely recorded value can still be invalid if the instrument or operational definition is unsuitable.
Reliability
A measurement or coding procedure should produce sufficiently consistent results under appropriate conditions. Relevant checks may include instrument reliability, repeated measurements, inter-rater agreement, or reproducible code.
Integrity
Data should remain complete and unaltered except through authorised, documented changes. Checksums, permissions, audit logs, version control and read-only master files support integrity.
Provenance
Provenance explains where the data came from and what happened to it.
It may include:
- Original source
- Collection date
- Instrument or software
- Researcher or system responsible
- Processing steps
- Code version
- Exclusions and corrections
- File relationships
- Ownership and licence
Timeliness
Data should be sufficiently current for the research purpose. Historical data may be entirely appropriate for a historical question but unsuitable for estimating a present condition.
Accessibility and usability
Authorised users should be able to locate, open, interpret, and analyse the data. Accessibility requires documentation and suitable formats, not simply possession of a file.
The Research Data Lifecycle
The research data lifecycle is the sequence through which data moves from initial planning to collection, analysis, sharing, preservation, reuse, or disposal. Data management decisions should be made at every stage rather than postponed until publication.
1. Plan
Before collection, determine:
- What data will be created or reused
- File formats and likely volume
- Collection instruments
- Roles and responsibilities
- Ethics and consent requirements
- Storage and backup arrangements
- Naming and versioning conventions
- Quality-control procedures
- Access restrictions
- Preservation and sharing plans
- Expected costs
2. Collect or generate
Use consistent procedures, tested instruments and documented protocols. Record contextual information while it is still known.
3. Process and clean
Transcribe, validate, correct, code, anonymise, convert, or combine the data. Preserve raw files and document every substantive transformation.
4. Analyse and interpret
Use appropriate statistical, computational, qualitative, visual, or mixed-methods procedures. Retain the scripts, coding decisions, parameters and software information needed to understand the analysis.
5. Document
Documentation is continuous rather than a final task. Maintain metadata, codebooks, README files, laboratory records, protocols, decision logs and data dictionaries.
6. Store and protect
Use institutionally approved systems, access controls, encryption where appropriate, backups, monitoring and recovery procedures.
7. Share or publish
Determine what can be shared, with whom, under what licence, at what time, and through which repository. Sensitive data may require mediated or controlled access.
8. Preserve, reuse or dispose
Retain valuable data and documentation in sustainable formats and repositories. When data must be destroyed, use an approved and documented disposal method.
What Is a Research Data Management Plan?
A data management plan is a structured explanation of how research data will be handled during and after a project. It normally covers data creation, documentation, storage, security, access, ethics, preservation, sharing, responsibilities and costs.
A useful DMP answers the following questions.
Data description
- What data will be collected, created, derived, or reused?
- What formats and approximate volumes are expected?
- Which existing data sources will be used?
Methods and quality
- How will data be collected or generated?
- What instruments, software and protocols will be used?
- How will accuracy, validity and consistency be checked?
Documentation and metadata
- What metadata standard will be used?
- Will the project create a README, codebook or data dictionary?
- How will processing and analytical decisions be recorded?
Storage and security
- Where will active data be stored?
- How will it be backed up?
- Who will have access?
- Does the data require encryption or a secure computing environment?
Ethics and legal compliance
- Does the project involve personal, confidential, Indigenous, commercial, copyrighted, or otherwise restricted data?
- What does the consent process allow?
- Are data-transfer or data-use agreements required?
Sharing and preservation
- Which data can be deposited?
- Which repository will be used?
- Will an embargo or controlled-access process be required?
- Which licence will apply?
- How long should the data be retained?
Responsibilities and resources
- Who is responsible for collection, documentation, quality, security, deposit and preservation?
- What staff time, storage, repository or curation costs should be budgeted?
A DMP should be treated as a living document and updated when the project changes.
How to Collect and Create Reliable Research Data
1. Begin with the research question
Collect only data that has a clear relationship to the research objectives. Collecting unnecessary variables increases cost, complexity, privacy risk and analytical flexibility without necessarily improving the study.
2. Define variables and concepts
Create operational definitions before collection.
For each variable, record:
- Name
- Meaning
- Unit
- Type
- Permitted values
- Missing-value code
- Source
- Collection method
- Calculation, if derived
3. Select suitable methods and instruments
The method should match the research question and population. Instruments may require validation, translation, adaptation, calibration or permission.
4. Pilot the procedure
A pilot can reveal confusing questions, technical problems, unrealistic completion times, inaccessible formats, inconsistent coding and missing response options.
5. Standardise collection
Use protocols, training, checklists, instrument settings and consistent instructions. Record unavoidable deviations.
6. Build quality controls into the process
Useful controls include:
- Mandatory fields where ethically appropriate
- Plausible value ranges
- Logic and skip checks
- Duplicate detection
- Timestamp review
- Calibration
- Double coding
- Inter-rater checks
- Automated validation scripts
- Manual review of exceptional cases
7. Maintain an audit trail
An audit trail records important changes and decisions. It should explain what changed, why, when, by whom, and how the change affects interpretation.
Organising Research Data
Use a logical folder structure
A simple structure might be:
project-name/
├── 01_admin/
├── 02_protocols/
├── 03_raw-data/
├── 04_processed-data/
├── 05_analysis/
├── 06_outputs/
├── 07_documentation/
└── README.txt
The exact structure matters less than consistency, documentation, access control and team agreement.
Use meaningful file names
A useful file name may include:
- Project or dataset identifier
- Data type
- Date in
YYYY-MM-DDformat - Version
- Status
Example:
student-survey_clean_2026-06-28_v02.csv
Avoid ambiguous names such as:
final.xlsx
final-new.xlsx
final-real-last.xlsx
Separate raw and working data
Keep original data unchanged whenever possible. Store cleaned and transformed versions separately. Use scripts rather than undocumented manual edits when the transformation can reasonably be automated.
Apply version control
Version control is useful for:
- Code
- Documentation
- Statistical scripts
- Plain-text data
- Data dictionaries
- Configuration files
- Collaborative writing
Large, sensitive, or binary datasets may require a specialised data platform rather than a public software repository.
Prefer sustainable formats
When disciplinary standards allow, preservation-friendly formats can reduce dependence on particular software.
Examples include:
- CSV or TSV for tables
- TXT or UTF-8 text for plain text
- JSON or XML for structured exchange
- TIFF or PNG for images
- WAV or FLAC for audio
- PDF/A for appropriate document preservation
The original format may still need to be retained when conversion could remove information or functionality.
Metadata and Research Documentation
Metadata is information that describes the data and enables people or machines to find, understand, assess and reuse it.
Metadata can include:
Descriptive metadata
- Dataset title
- Creator
- Keywords
- Abstract
- Geographic coverage
- Time period
- Persistent identifier
Structural metadata
- Relationships between files
- Table and variable connections
- Ordering of image or audio sequences
- Links between data and code
Administrative metadata
- Owner
- Licence
- Access level
- Retention period
- Consent limitations
- Repository information
Technical and provenance metadata
- File format
- Software version
- Instrument
- Calibration
- Processing history
- Code version
- Creation and modification dates
What to include in a README file
A basic README should identify:
- Project title and purpose
- Creators and contact information
- Folder and file descriptions
- Collection or generation methods
- Variable or coding documentation
- Processing and cleaning steps
- Software and version requirements
- Known limitations
- Rights, licence and access restrictions
- Recommended citation
DataCite’s metadata framework is designed to support consistent identification, citation and discovery of research outputs (DataCite, 2026).
Storage, Backup and Security
Research data should be stored according to its sensitivity, value, size, format, collaboration needs, and legal or contractual restrictions.
Use approved storage
Institutional servers, managed research drives, secure cloud environments, controlled research platforms and approved laboratory systems are usually preferable to personal devices or unapproved consumer services.
Control access
Apply the least-privilege principle: each person should have only the access needed for their role.
Possible controls include:
- Individual accounts
- Strong authentication
- Role-based permissions
- Encryption
- Secure transfer methods
- Access logs
- Time-limited access
- Separate storage of identifying information and research responses
Maintain backups
A backup must be:
- Separate from the active copy
- Updated at an appropriate frequency
- Protected from the same failure or security incident
- Tested to confirm that restoration works
Cloud synchronisation alone may not constitute an adequate backup because accidental deletion or corruption can synchronise across devices.
Plan for incidents
The project should know how to respond to:
- Lost equipment
- Accidental disclosure
- Unauthorised access
- Malware or ransomware
- Corrupted files
- Incorrect deletion
- Participant requests
- Breach-reporting obligations
Ethics, Privacy and Legal Considerations
Good data management is not only a technical matter. It is also an ethical and governance responsibility.
Informed consent
Participants should receive clear information about:
- What data will be collected
- Why it is needed
- How it will be used
- Who may access it
- How long it will be retained
- Whether it may be shared or reused
- Whether commercial or international partners are involved
- Limits to withdrawal after anonymisation or publication
A vague statement that data “may be used for research” may not be adequate for every secondary use.
Personal and sensitive data
Personal data can identify someone directly or indirectly. Removing names does not necessarily make a dataset anonymous. Combinations of age, location, occupation, dates, rare characteristics, images or free-text responses may permit re-identification.
Risk controls may include:
- Data minimisation
- Pseudonymisation
- De-identification
- Aggregation
- Removal or generalisation of indirect identifiers
- Controlled access
- Secure analysis environments
- Data-use agreements
- Disclosure review before publication
De-identification reduces risk but does not eliminate every possibility of re-identification (National Institute of Standards and Technology [NIST], 2015).
Confidential and commercially sensitive data
A project may be restricted by:
- Confidentiality agreements
- Intellectual-property rights
- Trade secrets
- Patent plans
- Commercial contracts
- Data-transfer agreements
- National-security or export-control rules
Copyright and database rights
Researchers may own a compilation or original documentation without owning every item contained in the dataset. Permission to access material is not necessarily permission to reproduce, distribute or license it.
Record the source and permitted uses of all reused data.
Indigenous data governance
Open-data practices must not override the rights and interests of Indigenous Peoples. The CARE principles emphasise:
- Collective Benefit
- Authority to Control
- Responsibility
- Ethics
These principles complement FAIR by focusing on people, power, rights, relationships and community benefit rather than technical reusability alone (Carroll et al., 2020).
FAIR Research Data
The FAIR principles state that research data should be:
Findable
Data and metadata should have clear descriptions, persistent identifiers and searchable records.
Accessible
The conditions and procedures for accessing data should be clear. Accessibility may involve public download, registration, approval, or controlled access.
Interoperable
Data and metadata should use suitable formats, vocabularies, standards and identifiers so they can work with other systems and datasets.
Reusable
Documentation, provenance, licences and quality information should allow others to assess whether the data is suitable for a new purpose.
FAIR does not mean that every dataset must be openly downloadable. Restricted data can still be FAIR when its metadata is findable and the conditions for legitimate access are clearly described (Wilkinson et al., 2016).
A useful principle is:
Make data as open as possible and as restricted as necessary.
Sharing and Publishing Research Data
Data sharing can support verification, secondary research, collaboration, education, transparency and recognition. However, sharing must be compatible with consent, ethics, law, culture, intellectual property, contracts and participant welfare.
Choose an appropriate repository
A repository may be:
- Discipline-specific
- Institutional
- General-purpose
- Funder-operated
- Publisher-associated
- Controlled-access
A suitable repository should be evaluated for:
- Disciplinary recognition
- Persistent identifiers
- Metadata quality
- Access controls
- Licensing options
- Versioning
- Preservation commitments
- File-size and format support
- Cost
- Withdrawal and correction procedures
Where possible, use an established repository rather than placing the only copy on a personal website. CoreTrustSeal identifies requirements associated with trustworthy data repositories (CoreTrustSeal, 2026).
Assign a licence
A licence tells users what they may do with the data. The choice depends on ownership, consent, third-party material, funder requirements, and the intended forms of reuse.
Do not apply an open licence to material you do not have authority to license.
Use persistent identifiers
A DOI or another persistent identifier makes a dataset easier to discover, link, cite and track.
Cite datasets
A dataset citation should normally include:
- Creator
- Year
- Dataset title
- Version
- Resource type
- Repository or publisher
- Persistent identifier
A generic APA-style pattern is:
Author, A. A. (Year). Title of dataset (Version) [Data set]. Repository. DOI
Researchers should cite reused datasets in the reference list rather than mentioning only the website from which a file was downloaded.
Write a data availability statement
An open-data statement might say:
The de-identified dataset and analysis code supporting this study are available in [repository] at [persistent identifier].
A restricted-data statement might say:
The data contain information that could compromise participant confidentiality. Qualified researchers may request controlled access subject to ethics approval and a data-use agreement.
A no-new-data statement might say:
No new data were created or analysed in this study.
The statement must accurately reflect the actual access conditions.
Research Data in Modern Research
Research data now frequently includes materials beyond traditional spreadsheets and laboratory measurements.
Modern examples include:
- High-frequency sensor streams
- Large-scale administrative records
- Digital trace data
- Social-media content
- Software containers
- Computational notebooks
- Machine-learning training data
- Model weights and parameters
- Synthetic data
- Geospatial data
- Three-dimensional scans
- Virtual-environment interactions
- Research software
- Workflow files
- Prompt and model-configuration records
Open-science frameworks increasingly treat data, software, methods, infrastructure and publications as connected research objects. UNESCO’s Recommendation on Open Science encourages scientific knowledge, including data and software, to be made accessible and reusable when appropriate (UNESCO, 2021).
Funder requirements also continue to develop. As reviewed on June 28, 2026, NIH maintains its Data Management and Sharing Policy, while NSF requires relevant proposals to include data management and sharing plans through its current Research.gov process (National Institutes of Health [NIH], 2026; National Science Foundation [NSF], 2026).
Researchers should always check the current instructions for their exact call because requirements vary by funder, programme, discipline and award date.
Digital Tools for Research Data
The right tool depends on the project, institution, sensitivity and discipline.
Collection and generation
Examples include:
- REDCap
- Qualtrics
- KoboToolbox
- Laboratory information-management systems
- Electronic laboratory notebooks
- Sensor and instrument software
- Secure transcription platforms
Organisation and collaboration
Examples include:
- Institutionally managed cloud storage
- Open Science Framework
- Git and approved Git platforms
- Electronic laboratory notebooks
- Research information-management platforms
Analysis
Examples include:
- R
- Python
- SPSS
- Stata
- SAS
- MATLAB
- NVivo
- ATLAS.ti
- Jupyter
- R Markdown or Quarto
Data management planning
Examples include:
- DMPTool
- DMPonline
- Funder-specific planning systems
- Institutional DMP templates
Publication and preservation
Examples include:
- Discipline-specific repositories
- Institutional repositories
- Zenodo
- Dryad
- Figshare
- ICPSR
- Controlled-access health or genomic repositories
A named tool should not be adopted solely because it is popular. Confirm institutional approval, security, access, export, retention, licensing and long-term availability.
Artificial Intelligence and Research Data
AI can support research-data work, but it can also introduce confidentiality, bias, provenance and reproducibility risks.
Potential uses
AI may assist with:
- Transcription
- Translation
- Data extraction
- Document classification
- Image recognition
- Qualitative coding suggestions
- Anomaly detection
- Missing-data investigation
- Code generation
- Metadata drafting
- Data-cleaning suggestions
- Literature or dataset discovery
- Synthetic-data generation
Important risks
Confidentiality
Uploading unpublished, personal, commercial, culturally sensitive or restricted data to an external AI service may disclose information to an unauthorised third party.
Do not upload protected data unless the tool, contract, institutional policy, ethics approval and security controls explicitly permit it.
Hallucination and fabrication
An AI system may invent values, categories, citations, explanations or patterns. AI-produced transformations must be checked against the source data.
Bias
Models may reproduce biases present in training data, prompts, labels or human decisions. Automated coding should not be treated as neutral.
Loss of provenance
A result may be difficult to reproduce when the model, version, system prompt, settings or service changes.
Hidden transformations
Some services alter files, compress images, normalise text or remove metadata. Researchers should confirm what happens to uploaded and downloaded content.
Responsible AI practice
When AI materially contributes to data processing or analysis, record:
- Tool and provider
- Model and version, where available
- Access date
- Task performed
- Prompts or instructions
- Parameters and settings
- Input-data classification
- Human review process
- Corrections made
- Known limitations
- Effect on reproducibility
Researchers remain responsible for the integrity of their data and conclusions. Current AI-integrity guidance emphasises legal compliance, ethical concerns, the integrity of the research record, publication practices, and human critical judgement (UK Research Integrity Office, 2025).
Advantages of Good Research Data Management
Effective management can:
- Reduce data loss
- Improve accuracy and consistency
- Make analysis more efficient
- Support collaboration
- Protect participants
- Demonstrate compliance
- Enable verification
- Improve reproducibility
- Facilitate reuse
- Increase visibility and citation
- Preserve valuable evidence
- Reduce duplication of research effort
- Clarify responsibilities within a team
Limitations and Practical Challenges
Research-data management also involves trade-offs.
Time and cost
Documentation, curation, secure storage, anonymisation and repository deposit require staff time and resources.
Disciplinary variation
A standard suitable for genomic sequences may be inappropriate for oral histories, archaeological objects or creative practice.
Privacy and re-identification risk
Highly detailed data can remain identifying even after direct identifiers are removed.
Loss of context
A dataset separated from its social, historical, laboratory or cultural setting may be misinterpreted.
Technical obsolescence
File formats, software, storage media and platforms change.
Unequal benefits
Data may be collected from communities that receive little benefit from its publication or commercial reuse.
Reproducibility limits
Sharing data and code improves transparency but does not automatically correct poor design, biased sampling, invalid measurements or inappropriate analysis.
Common Research-Data Mistakes
Avoid the following errors:
- Beginning collection without a data management plan.
- Keeping the only copy on one laptop or external drive.
- Editing the original raw file directly.
- Using unexplained variable names and codes.
- Failing to document exclusions and corrections.
- Collecting unnecessary personal information.
- Assuming that de-identification removes every privacy risk.
- Treating publicly visible online content as unrestricted.
- Mixing consent forms with analysis data.
- Sharing data that the consent process did not permit.
- Uploading confidential data to an unapproved AI service.
- Publishing a spreadsheet without a README or codebook.
- Depositing data without a licence or citation instructions.
- Assuming FAIR means completely open.
- Ignoring the rights and governance expectations of communities represented in the data.
- Depending on proprietary software without recording formats and versions.
- Describing AI-generated or simulated data as empirical observations.
- Deleting intermediate files needed to reproduce the analysis.
- Treating a data management plan as a one-time administrative form.
- Assuming that all funders and journals have the same requirements.
Research Data Management Checklist
Before collecting data:
- Define the data needed to answer the research question.
- Check funder, institutional, ethics and partner requirements.
- Identify ownership, consent and access conditions.
- Prepare a data management plan.
- Select approved collection and storage systems.
- Establish naming, folder, versioning and metadata conventions.
- Define quality-control procedures.
During the project:
- Preserve unchanged raw data.
- Maintain secure backups.
- Restrict access appropriately.
- Update documentation continuously.
- Record processing and analytical decisions.
- Review missing, inconsistent and exceptional values.
- Update the DMP when methods or risks change.
Before publication or project closure:
- Confirm which files support the findings.
- Review disclosure and re-identification risk.
- Remove files that cannot lawfully or ethically be shared.
- Prepare metadata, a README and a codebook.
- Select a suitable repository and licence.
- Obtain a persistent identifier.
- Cite the dataset.
- Write an accurate data availability statement.
- Preserve restricted data in an approved environment.
- Document retention or secure-disposal decisions.
Conclusion
Research data is the evidence on which research questions, analyses and conclusions depend. It includes far more than numerical spreadsheets: interviews, images, code, models, notes, documents, samples and creative materials may all function as data.
The value of research data depends not only on collection but also on quality, context, documentation, security, ethics and responsible stewardship. Planning these elements from the beginning makes data easier to understand, analyse, protect, verify, preserve and, where appropriate, share and reuse.
Frequently Asked Questions
1. What is research data in simple words?
Research data is the recorded evidence a researcher uses to answer a question or support a conclusion. It may consist of numbers, words, images, recordings, observations, code, documents, models or physical-material records.
2. What are the main types of research data?
Common classifications include primary and secondary data; qualitative and quantitative data; observational, experimental, simulation and derived data; and raw, processed and analysed data. A dataset can belong to several classifications simultaneously.
3. Is research data always numerical?
No. Interview transcripts, field notes, photographs, video, audio, documents, maps, code, artwork and archival materials can all be research data.
4. What is the difference between research data and a dataset?
Research data refers broadly to evidence used in research. A dataset is an organised collection of related research data, normally accompanied by enough metadata and documentation to explain its structure, origin and use.
5. Are journal articles research data?
Articles used only as background literature are not normally research data. They become data when researchers systematically collect, code, annotate, compare, mine or otherwise analyse them as evidence.
6. What is raw research data?
Raw data is the original recorded material before substantive cleaning, recoding, transformation or analysis. Researchers should preserve an unchanged master copy whenever possible.
7. What is research data management?
Research data management is the planning and control of data throughout its lifecycle, including collection, organisation, documentation, storage, protection, analysis, sharing, preservation and disposal.
8. Does FAIR data have to be publicly open?
No. FAIR means findable, accessible, interoperable and reusable. Accessibility can involve controlled procedures, and sensitive data may remain restricted while its metadata and access conditions are made discoverable.
9. Can a researcher upload data to an AI tool?
Only when the tool and intended use comply with institutional policy, consent, ethics approval, contracts, privacy requirements and security rules. Confidential or unpublished data should not be uploaded to an unapproved public AI service.
10. How should a research dataset be cited?
Include the creator, year, dataset title, version, resource type, repository and persistent identifier. Follow the required citation style and the repository’s recommended citation.
