Textual analysis is a family of methods for systematically examining written, spoken, visual, or multimodal texts to explain how they create meaning. Researchers study language, structure, themes, symbols, rhetoric, context, and sometimes measurable patterns such as word frequency. The appropriate approach depends on the research question, corpus, discipline, and theoretical framework.

Introduction
Texts do more than communicate information. They frame events, construct identities, express values, organize arguments, evoke emotions, and make some interpretations appear more natural than others. Textual analysis gives students and researchers a structured way to examine these processes.
The method is widely used in literature, linguistics, history, media and communication, sociology, politics, education, health research, marketing, and the digital humanities. However, “textual analysis” does not describe one universally standardized procedure. It is an umbrella term covering interpretive approaches such as close reading and discourse analysis as well as systematic coding and computational techniques.
This guide explains what textual analysis is, what can be treated as a text, how the main approaches differ, how to conduct an analysis, and how to report the method transparently. It also addresses sampling, quality, ethics, digital tools, and the responsible use of artificial intelligence.
Key Takeaways
- Textual analysis examines how texts create, organize, and communicate meaning.
- A text may be written, spoken, visual, digital, material, or multimodal.
- Textual analysis can be qualitative, quantitative, or mixed-method.
- Researchers should name the specific approach they use rather than relying only on the broad label.
- Strong analysis connects interpretations to textual evidence, context, theory, and a transparent procedure.
- AI can assist with some analytical tasks, but it does not replace human methodological judgment.
What Is Textual Analysis?
Textual analysis is the systematic examination and interpretation of texts to understand their content, form, meaning, function, context, or effects.
In interpretive research, the analyst asks how a text makes meaning and what plausible readings it permits. The objective is not necessarily to discover one final or universally correct interpretation. Instead, the researcher develops a reasoned interpretation supported by evidence from the text and its relevant context (McKee, 2003).
In quantitative and computational research, textual analysis may involve converting language into measurable features. Researchers may count words, classify documents, identify named entities, estimate sentiment, compare semantic similarity, or model topics across a large corpus.
The term therefore has two common uses:
- Interpretive textual analysis examines meaning, representation, rhetoric, narrative, discourse, symbolism, and context.
- Computational textual analysis extracts or measures patterns using dictionaries, statistical models, machine learning, or natural language processing.
These uses overlap. A mixed-method study might identify patterns computationally and then interpret selected passages closely.
What Counts as a Text?
In everyday language, a text usually means something written. In textual research, the category can be much broader.
A text may include:
- Novels, poems, plays, and short stories.
- Newspaper reports and magazine articles.
- Speeches, debates, and policy documents.
- Interview and focus-group transcripts.
- Diaries, letters, emails, and personal narratives.
- Advertisements, packaging, posters, and photographs.
- Films, television programmes, podcasts, and music videos.
- Websites, blogs, memes, and social-media posts.
- Museum displays, monuments, maps, and public spaces.
- Combinations of language, sound, image, layout, and movement.
An object should not be called a text merely because a researcher observes it. It is treated as a text when the study examines how its signs, structures, or communicative features can be interpreted.
A film, for example, may be analysed through dialogue, editing, camera angle, music, costume, colour, and narrative sequence. A website may be examined through its wording, hyperlinks, menus, images, layout, and interactive design.
What Does Textual Analysis Examine?
The features examined depend on the question and method. Common dimensions include the following.
Content
Content concerns what is explicitly or implicitly communicated:
- Topics.
- Claims.
- Themes.
- Values.
- Characters or social actors.
- Events and problems.
- Proposed solutions.
- Included and excluded information.
Language
Language-focused analysis may consider:
- Vocabulary and word choice.
- Grammar and syntax.
- Modality, such as “must,” “may,” or “could.”
- Metaphor and figurative language.
- Pronouns and forms of address.
- Evaluative or emotional language.
- Technical terminology.
- Presuppositions and implications.
Form and structure
Researchers may examine:
- Genre.
- Plot or narrative progression.
- Paragraph and sentence organization.
- Repetition and contrast.
- Openings and conclusions.
- Rhyme, rhythm, or meter.
- Headings and visual hierarchy.
- Turn-taking in dialogue.
- Links among image, sound, and language.
Rhetoric
Rhetorical analysis asks how a text attempts to persuade or position an audience. Relevant features include:
- Appeals to authority or credibility.
- Emotional appeals.
- Logical reasoning.
- Framing.
- Analogy.
- Repetition.
- Questions and commands.
- Audience construction.
- Responses to opposing positions.
Representation and ideology
Researchers may investigate how a text represents:
- Gender.
- Race and ethnicity.
- Class.
- Nation.
- Disability.
- Age.
- Institutions.
- Expertise.
- Social problems.
- Power relations.
The analysis should identify specific textual mechanisms rather than simply declaring that a text is biased or ideological.
Context
Meaning is influenced by the circumstances in which a text is produced, circulated, and interpreted. Context may include:
- Historical period.
- Political situation.
- Cultural conventions.
- Author or institutional position.
- Intended and actual audiences.
- Medium and platform.
- Genre expectations.
- Relationship to other texts.
Context should illuminate the text rather than replace textual evidence.
Intertextuality
Intertextuality refers to the ways one text quotes, echoes, adapts, challenges, or presupposes other texts. Researchers may study citations, allusions, genre conventions, recurring slogans, remakes, hyperlinks, or shared narratives.
Audience reception
A textual analysis can describe how a text appears to position an audience. It cannot by itself prove how real audiences understood or responded to the text.
Claims about actual reception normally require additional evidence, such as interviews, surveys, reviews, comments, experiments, ethnography, or audience analytics.
Is Textual Analysis Qualitative or Quantitative?
Textual analysis may be qualitative, quantitative, or mixed-method.
| Approach | Main purpose | Typical data | Typical output |
|---|---|---|---|
| Qualitative | Interpret meaning, context, representation, or experience | A small or moderately sized corpus studied in depth | Themes, interpretations, rhetorical patterns, narratives, or discourse claims |
| Quantitative | Measure predefined or computationally derived textual features | A larger corpus suitable for consistent measurement | Frequencies, proportions, scores, associations, clusters, or predictive models |
| Mixed-method | Combine scale with contextual interpretation | Large corpus plus selected texts or passages | Statistical patterns supported, challenged, or explained through close analysis |
Quantitative analysis is not automatically objective, because researchers still make decisions about sampling, categories, dictionaries, preprocessing, model selection, thresholds, and interpretation. Qualitative analysis is not simply personal opinion, because interpretations should be systematic, transparent, theoretically informed, and grounded in evidence.
Main Types of Textual Analysis
The categories below overlap. Their boundaries vary across disciplines, so researchers should define how they are using each term.
Close reading and literary analysis
Close reading examines how specific features of a literary or rhetorical text contribute to meaning. It may focus on diction, imagery, syntax, rhythm, narrative voice, characterization, symbolism, ambiguity, or form.
It is best suited to questions such as:
- How does the narrator’s changing vocabulary reveal uncertainty?
- How does the poem’s form reinforce its treatment of memory?
- How does repeated imagery connect two characters?
A strong close reading moves beyond identifying a device. It explains what the device does, how it relates to other features, and why it matters to the interpretation.
Qualitative content analysis
Qualitative content analysis systematically classifies text into categories while preserving attention to meaning and context. Categories may be developed inductively from the data, deductively from theory, or through a combination of both.
Hsieh and Shannon (2005) distinguish conventional, directed, and summative forms of qualitative content analysis. The approach is useful when researchers want a transparent system for organizing a body of textual data.
Quantitative content analysis
Quantitative content analysis applies explicit coding rules to count or compare textual features. It is suitable for questions such as:
- How often are particular sources quoted?
- What proportion of reports frame an issue as an economic problem?
- Does the use of threatening language differ across outlets?
The method normally requires clearly defined units, categories, sampling procedures, and coding rules. When multiple coders classify the same material, intercoder agreement may be evaluated.
Thematic analysis
Thematic analysis identifies and interprets patterned meaning across a dataset. It is frequently applied to interview transcripts, open-ended survey responses, diaries, and other qualitative materials.
Themes should not be treated as topics that simply appear in the data. A theme is an analytically developed pattern that helps answer the research question. Researchers should identify the specific form of thematic analysis they follow, because coding-reliability, codebook, and reflexive approaches make different assumptions (Braun & Clarke, 2006, 2022).
Discourse analysis
Discourse analysis studies language in use. It asks how language constructs identities, relationships, knowledge, institutions, and social realities.
Possible questions include:
- How are patients positioned as responsible for treatment outcomes?
- How do policy documents define a “deserving” citizen?
- How does professional language establish authority?
Discourse analysis usually examines more than isolated vocabulary. It considers recurring patterns, categories, assumptions, interactions, and social functions.
Critical discourse analysis
Critical discourse analysis investigates connections among language, ideology, inequality, and power. It may examine how apparently ordinary wording legitimizes institutions, normalizes social arrangements, or marginalizes alternatives.
Researchers should avoid beginning with a predetermined accusation and selecting only confirming quotations. Claims about power should be demonstrated through systematic textual and contextual evidence.
Rhetorical analysis
Rhetorical analysis explains how a text attempts to persuade a particular audience in a particular situation.
It may consider:
- The speaker’s credibility.
- Emotional and logical appeals.
- Arrangement of arguments.
- Metaphor and analogy.
- Repetition.
- Timing.
- Intended audience.
- Constraints on the speaker.
Rhetorical analysis is useful for speeches, advertisements, campaigns, editorials, legal arguments, and public communication.
Narrative analysis
Narrative analysis studies how stories organize events and experiences. Researchers may examine:
- Plot.
- Sequence.
- Turning points.
- Characters.
- Causality.
- Voice.
- Temporality.
- Evaluation.
- Resolution.
- Cultural narrative forms.
It can be applied to fiction, interviews, institutional histories, legal testimony, illness narratives, and organizational communication.
Semiotic and multimodal analysis
Semiotic analysis examines how signs produce meaning. A sign may be a word, image, colour, gesture, sound, object, or spatial arrangement.
Multimodal analysis examines how different communicative modes work together. An advertisement, for example, may combine written language, photography, colour, typography, composition, and sound.
Corpus and computational text analysis
Computational textual analysis uses digital methods to examine text at scale. Techniques may include:
- Word-frequency analysis.
- Keyword analysis.
- Concordance analysis.
- Collocation.
- Sentiment analysis.
- Named-entity recognition.
- Document classification.
- Topic modelling.
- Semantic similarity.
- Word embeddings.
- Transformer-based language models.
- Stylometry.
Computational outputs still require validation and contextual interpretation. A sentiment dictionary developed for consumer reviews may perform poorly on political speeches, clinical notes, irony, historical language, or discipline-specific terminology.
Textual Analysis Compared With Related Methods
| Method | Primary object | Central question | Typical procedure |
|---|---|---|---|
| Textual analysis | Text as an object of interpretation or measurement | How does the text create meaning or display patterns? | Close examination, coding, interpretation, or computation |
| Content analysis | Categories within communication | What is present, absent, or recurrent, and sometimes how often? | Define units and categories, code material, compare results |
| Thematic analysis | Patterned meaning across qualitative data | What meaningful patterns answer the research question? | Familiarization, coding, theme development, review, interpretation |
| Discourse analysis | Language in social use | How does language construct reality, identity, or power? | Analyse linguistic choices, discursive patterns, functions, and context |
| Literary analysis | Literary works | How do literary form and language contribute to interpretation? | Close reading supported by textual and contextual evidence |
| Literature review | Existing scholarship | What is known, debated, or missing in a field? | Search, evaluate, synthesize, and compare research literature |
| Sentiment analysis | Evaluative language | Is the language classified as positive, negative, neutral, or more specific emotions? | Dictionary-based or model-based classification and validation |
A literature review uses publications primarily as sources of knowledge about a topic. Textual analysis treats selected texts themselves as the object of analysis.
When Should Textual Analysis Be Used?
Textual analysis is appropriate when the research question concerns:
- Meaning.
- Representation.
- Communication.
- Language.
- Argument.
- Narrative.
- Symbolism.
- Framing.
- Ideology.
- Cultural conventions.
- Recurring textual patterns.
- Differences among texts, genres, institutions, periods, or platforms.
Examples of suitable research questions include:
- How do university prospectuses construct the idea of employability?
- How is climate responsibility framed in corporate sustainability reports?
- How do newspaper headlines assign agency during industrial disputes?
- How does a novel’s shifting point of view shape moral judgment?
- What narratives of professional identity recur in teachers’ interviews?
- How does visual composition support the verbal message of health campaigns?
Textual analysis is less suitable as a stand-alone method when the objective is to establish:
- The actual psychological effect of a text.
- The intention privately held by its author.
- How all audience members interpreted it.
- Causal relationships.
- Population prevalence without an appropriate sampling design.
- Behaviour not represented in the texts.
Those questions may require experiments, interviews, surveys, observation, archival evidence, or mixed methods.
How to Conduct Textual Analysis Step by Step
Step 1: Formulate a focused research question
A useful question identifies the phenomenon, textual material, and analytical interest.
Too broad:
How does the media represent education?
More focused:
How do the front-page headlines of three national newspapers assign responsibility for university tuition increases between January and June 2026?
The second question identifies the texts, feature of interest, sources, and period.
Step 2: Clarify your theoretical and methodological position
Determine what your study assumes about texts and meaning.
Ask:
- Is meaning treated as embedded in the text, produced through interpretation, or constructed socially?
- Am I primarily describing content, interpreting experience, examining persuasion, or critiquing power?
- Will categories be inductive, deductive, or combined?
- Which theoretical concepts guide attention?
A theoretical framework should help the analysis rather than merely supply terminology. Explain why it is appropriate and how it affects what you examine.
Step 3: Define and construct the corpus
The corpus is the complete collection of texts selected for analysis.
Document:
- Source or archive.
- Publication dates.
- Languages.
- Genres or formats.
- Search terms.
- Inclusion criteria.
- Exclusion criteria.
- Duplicate handling.
- Version selection.
- Number of documents.
- Relevant metadata.
For online material, record access dates and, where lawful and appropriate, preserve stable copies because webpages and posts can change or disappear.
Step 4: Choose a sampling strategy
Possible strategies include:
- Comprehensive sampling: Include every text in a clearly bounded collection.
- Criterion sampling: Include texts that satisfy predefined criteria.
- Purposive sampling: Select information-rich cases relevant to the question.
- Maximum-variation sampling: Select contrasting texts to capture diversity.
- Typical-case sampling: Examine texts considered ordinary within the setting.
- Critical-case sampling: Select a case with particular strategic importance.
- Theoretical sampling: Select further material as emerging theory requires.
- Random or systematic sampling: Use probability-based selection where population inference is intended.
There is no universal minimum number of texts. A detailed analysis of one novel may be valid, while a corpus study may require thousands of documents. Adequacy depends on the question, method, diversity of the corpus, unit of analysis, and claims being made.
Step 5: Address ethics, privacy, and copyright
Before collecting or uploading text, determine:
- Whether the study requires institutional ethics review.
- Whether the material contains identifiable or sensitive information.
- Whether participants consented to secondary analysis.
- Whether public availability creates a reasonable expectation of research use.
- Whether quotations could be searched to identify a person.
- Whether platform terms or licences restrict collection.
- How copyrighted material may be quoted or reproduced.
- Where the corpus and analytical files will be stored.
- Whether an external AI or cloud service is permitted to process the data.
Publicly accessible does not always mean ethically risk-free.
Step 6: Select the analytical approach and units
Define the unit at each relevant level:
- Sampling unit: The item selected, such as an article or speech.
- Context unit: The material needed to interpret a segment.
- Coding unit: The segment assigned a code, such as a sentence or paragraph.
- Counting unit: The element measured, such as an occurrence, document, or speaker turn.
Then specify what will be examined: themes, frames, metaphors, narrative stages, rhetorical appeals, grammatical choices, visual signs, word frequencies, or another defined feature.
Step 7: Prepare and familiarize yourself with the texts
Preparation may involve:
- Transcription.
- Anonymization.
- Formatting.
- Optical-character correction.
- Removal of duplicates.
- Language detection.
- Metadata creation.
- Segmenting documents.
- Recording missing or damaged material.
Read or view the texts repeatedly before finalizing categories. Record early observations, questions, contradictions, and contextual information in analytic memos.
Step 8: Develop codes, categories, or analytical questions
A code is a concise label attached to a meaningful feature of the material.
Codes may be:
- Semantic: Describe explicit content.
- Latent: Interpret assumptions or underlying meanings.
- Descriptive: Identify what is discussed.
- Process-oriented: Identify actions or changes.
- Evaluative: Identify approval, criticism, risk, or value.
- Structural: Identify parts of a text, such as an opening claim or counterargument.
- Theory-derived: Apply concepts from an existing framework.
Pilot the coding system on a varied subset. Revise ambiguous definitions before analysing the full corpus.
Step 9: Analyse patterns and relationships
Analysis goes beyond attaching labels. Examine:
- Similarities and differences.
- Repetition and absence.
- Co-occurring codes.
- Contradictions.
- Changes across time.
- Differences among authors, genres, or institutions.
- Relationships between form and content.
- Typical and deviant cases.
- Connections to context and theory.
Return repeatedly to the original text. Codes, counts, or software outputs should not become substitutes for the evidence they summarize.
Step 10: Evaluate alternative interpretations
Ask:
- What evidence contradicts the emerging interpretation?
- Are quotations being removed from their context?
- Could a pattern be produced by the sampling strategy?
- Have rare but important cases been ignored?
- Am I treating an assumed audience effect as an observed effect?
- Does the theoretical framework illuminate the data or force it into predetermined categories?
Negative or deviant cases can improve an interpretation by revealing its boundaries.
Step 11: Establish quality and trustworthiness
Appropriate quality procedures depend on the analytical tradition. They may include:
- A transparent audit trail.
- Reflexive memos.
- Explicit inclusion and exclusion criteria.
- A documented codebook.
- Peer debriefing.
- Comparison of independent coding.
- Intercoder-agreement statistics.
- Negative-case analysis.
- Triangulation.
- Sensitivity analysis.
- Member reflection where appropriate.
- Validation against a manually labelled benchmark.
- Reproducible preprocessing and code.
- Thick contextual description.
Do not add an intercoder-reliability coefficient merely to make an interpretive study appear scientific. Use it when stable category application by multiple coders is part of the design. Reflexive approaches may instead treat researcher interpretation as an analytic resource that must be examined transparently.
Step 12: Interpret and report the findings
Organize findings around the research question rather than presenting an inventory of codes.
A strong findings section normally:
- States the analytical claim.
- Provides appropriate textual evidence.
- Explains how the evidence supports the claim.
- Identifies variation or exceptions.
- Connects the result to context or theory.
- Avoids claims that exceed the corpus.
Textual-Analysis Codebook Template
| Field | What to record | Illustrative entry |
|---|---|---|
| Code name | Short, distinctive label | Institutional responsibility |
| Definition | Meaning of the code | The institution explicitly accepts responsibility for causing or addressing a problem |
| Include when | Positive inclusion rule | Uses “we,” “our duty,” or an equivalent statement linked to responsibility |
| Exclude when | Boundary rule | Describes action without attributing responsibility |
| Unit | Segment being coded | Sentence |
| Example | Typical segment | “We accept responsibility for reducing our operational emissions.” |
| Counterexample | Similar but excluded segment | “Emissions will be reduced by 2030.” |
| Parent category | Broader analytical group | Agency |
| Notes | Ambiguities or revisions | Distinguish responsibility from leadership claims |
A codebook is especially valuable in structured content analysis, team coding, and deductive studies. More reflexive approaches may use evolving code descriptions rather than treating the codebook as a fixed measurement instrument.
Worked Textual-Analysis Example
Research question
How do university sustainability statements construct institutional responsibility for climate action?
Illustrative corpus
Suppose a researcher selects 30 publicly available sustainability statements issued by universities between 2024 and 2026. Inclusion criteria require each document to be an official institution-level statement containing at least one climate target.
Possible analytical approach
The researcher combines rhetorical analysis with qualitative content analysis.
The coding framework includes:
- Acceptance of responsibility.
- Leadership claims.
- Collective pronouns.
- Passive constructions.
- Measurable commitments.
- Unspecified future action.
- Appeals to science.
- Appeals to student expectations.
- Economic opportunity.
- Transfer of responsibility to individuals.
Original illustrative passage
“Together, our community will accelerate the transition to a cleaner future. New technologies and everyday choices will help us reach net zero.”
Analysis
The collective pronouns “our” and “us” create an inclusive institutional identity. However, “together” distributes responsibility across the whole community rather than identifying which organizational actors control budgets, buildings, procurement, or energy systems.
The phrase “will accelerate” communicates confidence and forward movement, but the sentence does not identify a baseline, deadline, or accountable decision-maker. “New technologies and everyday choices” combines structural and individual responses. This pairing may allow the institution to present climate action as shared while leaving the relative responsibilities of leadership and community members unspecified.
The interpretation is not based on one phrase alone. The researcher would compare this pattern across the corpus, examine counterexamples, and consider whether more specific commitments appear elsewhere in each document.
Possible finding
Across the corpus, universities frequently constructed climate action as a collective moral project through inclusive pronouns. Statements were less consistent in assigning responsibility to identifiable institutional actors. Documents containing externally reported targets were more likely to combine collective language with specific organizational commitments.
The final sentence would require systematic evidence. The researcher should not make it unless the corpus and analysis support the comparison.
Short Literary Textual-Analysis Example
Consider this original sentence:
“The station clock swallowed another minute while Mara held the unopened letter.”
A weak observation would say that the sentence contains personification.
A stronger textual analysis explains its function:
The verb “swallowed” turns time into an active and consuming force. It suggests that Mara’s opportunity to act is disappearing rather than merely passing. The unopened letter delays disclosure, while “held” emphasizes physical possession without decision. Together, the clock and letter create tension between external movement and internal hesitation.
The analysis names the linguistic feature, explains its effect, and connects it to a larger interpretation.
Advantages of Textual Analysis
It can reveal how meaning is constructed
Textual analysis examines not only what is communicated but how vocabulary, structure, genre, imagery, and context shape interpretation.
It can be used with diverse materials
The method can be applied to literature, interviews, policy documents, advertisements, online communication, images, films, and multimodal artefacts.
It supports historical and contemporary research
Researchers can analyse archived documents, current digital content, or comparisons across periods.
It can examine material that is difficult to observe directly
Texts may provide access to institutional values, cultural assumptions, personal narratives, public arguments, or historical debates.
It is compatible with multiple research traditions
Textual analysis can support interpretive, critical, descriptive, quantitative, computational, and mixed-method designs.
It can connect depth with scale
Close reading provides contextual depth, while computational techniques can identify patterns across large corpora. Combining them can help researchers test whether an interpretation is isolated or widespread.
Limitations of Textual Analysis
Interpretation can be influenced by the researcher
Analysts bring theoretical commitments, experiences, and expectations to the material. Reflexivity and transparent reasoning are therefore essential.
Texts do not provide direct access to intentions
A text may support an interpretation of how authors or institutions present themselves, but it does not automatically reveal private motives.
Audience effects cannot be assumed
The fact that a text uses an emotional appeal does not prove that audiences were persuaded.
Context may be incomplete
Historical documents, deleted posts, edited webpages, or isolated quotations may lack the information needed for a secure interpretation.
Sampling can distort conclusions
A convenient or narrow corpus may overrepresent certain institutions, genres, platforms, or viewpoints.
Coding can reduce complexity
Categories make comparison possible but may flatten ambiguity, irony, contradiction, or change over time.
Computational models can misclassify language
Sarcasm, negation, multilingual content, specialist vocabulary, historical spelling, and cultural references can undermine automated analysis.
Generalization is limited
Detailed interpretation of a purposively selected corpus may support theoretical or analytical insights without supporting statistical generalization to all texts or populations.
Common Textual-Analysis Mistakes
Summarizing instead of analysing
Summary states what a text says. Analysis explains how textual choices produce meaning and why the pattern matters.
Listing techniques without explaining their function
Identifying metaphor, repetition, or passive voice is only the beginning. Explain its role in the passage, document, genre, and argument.
Claiming to know the author’s intention
Use formulations such as “the passage constructs,” “the wording suggests,” or “the text positions the reader” unless independent evidence establishes intention.
Selecting only supportive quotations
Look for contradictions, variation, absence, and negative cases.
Using an undefined method label
Do not write only, “The data were analysed using textual analysis.” Name the specific approach, theoretical framework, corpus, unit, and procedure.
Confusing frequency with importance
A rarely occurring passage may be analytically central, while a frequent word may be functionally trivial.
Treating software as the analyst
Software organizes, retrieves, counts, or models material. Researchers remain responsible for definitions, validation, interpretation, and claims.
Ignoring versions and metadata
Publication date, revision history, author, outlet, format, and platform can materially affect interpretation.
Overgeneralizing beyond the corpus
A study of 20 campaign speeches cannot automatically establish how all politicians communicate.
Digital Tools for Textual Analysis
The appropriate tool depends on corpus size and analytical purpose.
Basic tools
Word processors, spreadsheets, PDF annotation software, and reference managers may be sufficient for a small corpus. Researchers can create tables for passages, codes, memos, sources, and interpretations.
Qualitative data-analysis software
Computer-assisted qualitative data-analysis software can support:
- Coding.
- Memo writing.
- Retrieval of coded passages.
- Code hierarchies.
- Case classifications.
- Queries.
- Team collaboration.
- Visual mapping.
Examples include NVivo, ATLAS.ti, MAXQDA, Dedoose, and the open-source Taguette platform. Software selection should consider institutional access, security, collaboration, file formats, and export options.
Corpus-analysis tools
Corpus tools can generate:
- Concordance lines.
- Word lists.
- Keywords.
- Collocations.
- N-grams.
- Distribution patterns.
Examples include AntConc and Voyant Tools.
Programming environments
R and Python can support reproducible workflows for:
- Cleaning text.
- Tokenization.
- Dictionary analysis.
- Classification.
- Topic modelling.
- Embeddings.
- Visualization.
- Validation.
- Statistical comparison.
Researchers should preserve code, package versions, preprocessing decisions, and random seeds where relevant.
Artificial Intelligence and Textual Analysis
Generative AI and large language models can assist with some textual-analysis tasks, including:
- Suggesting preliminary codes.
- Grouping similar passages.
- Summarizing documents.
- Generating alternative interpretations.
- Comparing a passage with a code definition.
- Translating or simplifying text.
- Extracting structured fields.
- Identifying material for human review.
However, AI output should be treated as provisional analytical material rather than an authoritative finding.
Risks of AI-assisted analysis
Potential risks include:
- Fabricated quotations or references.
- Loss of contextual nuance.
- Inconsistent results across prompts or model versions.
- Bias against minority language or culturally specific meanings.
- Overly generic themes.
- Leakage of confidential data.
- Undisclosed processing by third-party services.
- Inability to reconstruct proprietary model behaviour.
- Excessive dependence on patterns learned from unrelated data.
A responsible AI-assisted workflow
- Confirm that institutional, legal, and consent requirements permit AI processing.
- Remove or protect personal and confidential information.
- Use an approved secure environment where required.
- Record the model, version, date, parameters, prompts, and segmentation procedure.
- Provide the research question and carefully defined analytical task.
- Pilot the process on a small, diverse subset.
- Compare AI output with close human analysis.
- Check missed cases, invented material, bias, and unstable classifications.
- Retain human responsibility for code definitions, themes, interpretations, and conclusions.
- Disclose the role of AI in the methodology or acknowledgements.
AI may increase speed or offer a contrasting reading, but interpretive adequacy cannot be established by fluency alone.
Ethical Considerations
Textual analysis may involve low-risk published documents or highly sensitive personal narratives. Ethical requirements should be evaluated rather than assumed.
Important questions include:
- Was the text created for research, publication, private communication, or another purpose?
- Could a quoted phrase identify its author through a search engine?
- Does removing a username sufficiently anonymize the material?
- Are children or vulnerable groups represented?
- Could publication expose a person or community to harm?
- Does the research reproduce offensive or traumatic content unnecessarily?
- Are translations faithful and culturally appropriate?
- Does the researcher have permission to store or redistribute the corpus?
- Will AI, transcription, or cloud software transfer data to another jurisdiction?
Researchers should follow their institution’s ethics process and the requirements applicable to their discipline and jurisdiction.
How to Write a Textual-Analysis Methodology
A methodology section should enable readers to understand what was analysed, why it was selected, and how conclusions were developed.
Methodology template
This study used [specific approach] to examine how [phenomenon] was constructed in [type of texts]. The corpus comprised [number and type of documents] published by [sources] between [dates]. Documents were included when [criteria] and excluded when [criteria]. The unit of analysis was [unit].
Analysis was informed by [theory or conceptual framework]. After [preparation procedure], the researcher read the corpus repeatedly and recorded analytic memos. Initial codes were developed [inductively/deductively/through a combined approach] and tested on [pilot material]. Codes were then revised and applied to the corpus using [software or manual procedure].
The analysis compared [relevant relationships] and examined contradictory or deviant cases. Trustworthiness was supported through [audit trail, reflexive memoing, independent coding, peer review, validation, or another appropriate procedure]. Ethical safeguards included [review, anonymization, secure storage, quotation strategy, and consent arrangements].
Adapt the template to the actual tradition. Do not claim independent coding, saturation, triangulation, or validation unless it was genuinely performed.
How to Present Textual-Analysis Findings
A useful findings structure is:
Analytical claim
State the pattern or interpretation directly.
Evidence
Present a quotation, description, count, concordance, image detail, or model result.
Explanation
Show which textual features support the claim.
Variation
Identify differences, contradictions, or exceptions.
Context and significance
Explain how the finding answers the question and relates to theory or prior research.
For qualitative studies, quotations should not be expected to “speak for themselves.” For computational studies, a graph or model score also requires interpretation, error analysis, and appropriate validation.
Reporting Standards
Researchers should consult reporting guidance appropriate to their field. The Standards for Reporting Qualitative Research provide a broad framework for reporting qualitative studies, while COREQ is specifically designed for interview and focus-group research.
Reporting guidance should be used as a transparency aid rather than a substitute for methodological understanding. Not every checklist item applies equally to literary, historical, computational, or critical textual research.
Conclusion
Textual analysis is a broad family of methods for explaining how texts create meaning or display systematic patterns. Its value lies in careful alignment among the research question, corpus, theory, analytical procedure, evidence, and claims.
A rigorous study does more than summarize documents or identify isolated techniques. It defines what counts as data, explains how texts were selected, names a specific analytical approach, considers alternative interpretations, and supports every conclusion with appropriate evidence. Digital and AI tools can extend the scale of analysis, but methodological judgment, contextual knowledge, ethical responsibility, and transparent reporting remain essential.
