Social network analysis (SNA) is a research approach for studying relationships among people, groups, organizations or other entities. It represents actors as nodes and their relationships as ties, then uses visual, mathematical and statistical methods to examine connection patterns, influential positions, cohesive groups, information flow and changes in network structure.

Introduction
Many research methods examine people or organizations by measuring their individual characteristics. A conventional survey might compare students by age, grades or study habits, for example. Social network analysis asks a different question: how are those students connected, and how do their relationships shape opportunities, constraints and outcomes?
SNA focuses on relational data—information about who communicates with whom, who collaborates with whom, who seeks advice from whom, or which organizations exchange resources. It can reveal informal leaders, isolated actors, cohesive communities, bridges between groups and pathways through which information or behaviour may spread.
This guide explains the concepts, measures, research designs, data structures, software and limitations of social network analysis. It also shows how to conduct an SNA study and interpret the results without confusing mathematical prominence with real-world influence.
Key Takeaways
- Social network analysis studies relationships, not only the attributes of separate individuals.
- Nodes represent actors or entities, while ties represent specified relationships or interactions.
- Centrality, density, reciprocity and clustering answer different questions and are not interchangeable.
- Network boundaries, missing actors and tie definitions can substantially change the results.
- A network diagram is an analytical aid, not proof of influence, causation or importance.
- Strong SNA research connects measures to theory, validates the data and reports methodological decisions transparently.
What Is Social Network Analysis?
Social network analysis is a family of theoretical, visual, mathematical and statistical methods used to examine patterns of relationships among interconnected entities.
The entities may be individuals, households, teams, organizations, countries, websites, publications, genes or other units. The connections may represent friendship, advice, communication, collaboration, trade, co-authorship, physical contact, financial transfers or another relationship defined by the researcher.
The distinctive feature of SNA is its assumption that actors are not isolated observations. Their positions, choices and opportunities may depend on their direct and indirect connections to others. Network analysis therefore studies both:
- The position of an actor within a network, and
- The structure of the network as a whole.
SNA is not limited to online platforms. A classroom friendship network, a hospital referral network and a network of international trade relationships are all suitable subjects for social network analysis.
Why Relationships Require a Different Analytical Approach
Traditional statistical analysis often places cases in rows and variables in columns. Each person might have a score for age, income, education or health. Many common techniques then examine differences or associations among these attributes.
Network data introduce an additional structure: one person’s observation is directly linked to another person’s observation. A friendship exists between at least two actors. A communication tie connects a sender and a receiver. Several ties may share the same actor.
This interdependence has three important consequences:
- The relationship, rather than the individual alone, may be the relevant unit of analysis.
- Observations may violate the independence assumptions of conventional statistical tests.
- Removing or adding one actor can change the measured positions of several other actors.
SNA is therefore more than drawing a graph. It provides concepts and methods designed specifically for relational dependence (Borgatti et al., 2009; Wasserman & Faust, 1994).
Core Concepts in Social Network Analysis
Nodes or Actors
A node is an entity represented in the network. In social research, nodes are often called actors. In mathematics and computer science, they may be called vertices.
Possible nodes include:
- Students in a classroom
- Employees in a company
- Hospitals in a referral system
- Countries in a trade network
- Authors in a co-authorship network
- Research papers in a citation network
- Online accounts in a communication network
All nodes in a one-mode network normally represent the same type of entity.
Ties, Edges or Links
A tie represents a relationship or interaction between two nodes. The terms edge and link are frequently used as synonyms.
A tie must have an operational definition. “Connection” is too vague for a reproducible study. A researcher should state whether a tie represents:
- Communicating at least once per week
- Naming someone as a trusted adviser
- Co-authoring at least one publication
- Sending a patient referral
- Transferring money
- Following an account
- Sharing membership in an organization
Different definitions can produce different networks from the same group of actors.
Dyads and Triads
A dyad consists of two actors and the possible ties between them. In a directed network, the actors may have no tie, a one-way tie or reciprocal ties.
A triad contains three actors. Triadic patterns are important because two actors who share a contact may become connected—a process often described as triadic closure.
Paths and Geodesic Distance
A path is a sequence of ties connecting one node to another. The shortest path between two nodes is often called the geodesic path.
The number of steps in the shortest path is the geodesic distance. Short paths may facilitate rapid access, communication or diffusion, although the substantive interpretation depends on what the ties represent.
Components and Isolates
A component is a set of nodes connected to one another through paths.
An isolate is a node with no observed ties in the measured network. Isolation may indicate exclusion, nonparticipation, missing data, a restrictive tie definition or an incorrectly specified network boundary. Researchers should not assume a single explanation.
Node Attributes
Nodes may also have conventional variables called attributes, such as age, department, occupation, location or political affiliation.
SNA can investigate whether network positions differ by attributes or whether actors with similar attributes are more likely to form ties.
Types of Social Networks
Directed and Undirected Networks
In an undirected network, a tie has no arrow. Co-authorship and being members of the same team are commonly represented as undirected relationships.
In a directed network, ties have an origin and destination. Advice-seeking, nominations, following and sending messages are normally directional.
Direction changes interpretation. In an advice network:
- High in-degree may indicate that many people seek advice from an actor.
- High out-degree may indicate that an actor seeks advice from many people.
Binary and Weighted Networks
A binary network records whether a tie is present or absent.
A weighted network records the strength, frequency, volume or another value associated with a tie. A weight might represent the number of messages exchanged, frequency of meetings or amount of trade.
Researchers must decide whether larger values represent stronger connections or greater distance or cost. This distinction is essential when calculating shortest paths.
Positive, Negative and Signed Networks
Some studies include positive and negative relationships, such as cooperation and conflict, trust and distrust, or support and opposition.
Combining positive and negative ties into one unsigned network can conceal important differences. Signed-network methods should be considered when the meaning of the relationship depends on its positive or negative character.
One-Mode Networks
A one-mode network contains one type of node. Examples include student-to-student friendships and organization-to-organization partnerships.
Two-Mode or Bipartite Networks
A two-mode network contains two node types, with ties running between the types. Examples include:
- Authors and publications
- Students and clubs
- Directors and company boards
- Legislators and bills
- Organizations and events
A two-mode network can sometimes be projected into a one-mode network, such as an author network based on shared publications. Projection can inflate connections and discard information, so the decision should be justified.
Whole or Sociocentric Networks
A whole-network design attempts to identify all actors and relevant ties within a specified boundary, such as all employees in an office or all agencies participating in a regional partnership.
Whole-network analysis supports measures such as overall density, centralization and components. It usually requires high participation because missing actors can affect many other observations.
Ego Networks
An ego network is organized around a focal actor called the ego. The people or organizations connected to the ego are called alters.
An ego-network study may collect:
- The ego’s ties to alters
- Characteristics of each alter
- The strength or type of each tie
- Relationships among the alters
Ego-network designs are useful when complete network membership cannot be listed or when the research question concerns personal support, exposure or access to resources.
Multiplex and Multilayer Networks
A multiplex network represents more than one type of relationship among the same actors. Employees may have advice, friendship and reporting ties, for example.
A multilayer network may contain different types of nodes, ties or contexts arranged as connected layers. Researchers should avoid collapsing layers when the relationship types have substantively different meanings.
Cross-Sectional and Longitudinal Networks
A cross-sectional network represents relationships at one time or during one defined period.
A longitudinal or dynamic network contains observations from multiple times. It can be used to study tie formation, tie dissolution, actor entry, actor exit and changes in network structure.
How Network Data Are Represented
Edge List
An edge list contains one row for each observed tie.
| Source | Target | Weight | Relationship |
|---|---|---|---|
| A | B | 4 | Advice |
| A | C | 2 | Advice |
| C | D | 5 | Advice |
A directed edge list distinguishes the source from the target. An undirected list records each pair only once.
Edge lists are convenient for data storage and are accepted by many analysis programs.
Adjacency Matrix
An adjacency matrix is a square table in which the rows and columns represent the same nodes.
For a binary network:
[
A_{ij} =
\begin{cases}
1, & \text{if a tie from } i \text{ to } j \text{ exists} \
0, & \text{otherwise}
\end{cases}
]
An undirected network normally has a symmetric matrix because (A_{ij}=A_{ji}). A directed matrix need not be symmetric.
The diagonal represents self-ties. These are often excluded, but the decision depends on the network.
Incidence Matrix
A two-mode network can be represented by an incidence matrix. Rows may represent people and columns may represent organizations, with each cell indicating membership or participation.
Data Dictionary
A reproducible network dataset should be accompanied by a data dictionary specifying:
- Node identifiers
- Node inclusion criteria
- Tie definition
- Direction
- Weight meaning and scale
- Observation period
- Treatment of duplicate ties
- Treatment of missing values
- Self-tie policy
- Attribute definitions
- Anonymization procedure
Main Measures in Social Network Analysis
No single measure describes an entire network. The correct measure depends on the research question, network type and meaning of the ties.
Overview of Common Measures
| Measure | Level | Main question | Important caution |
|---|---|---|---|
| Degree centrality | Node | Who has many direct ties? | More ties do not always mean more influence |
| In-degree | Node | Who receives many directed ties? | Meaning depends on the nomination or flow |
| Out-degree | Node | Who sends or initiates many ties? | May indicate activity, dependence or burden |
| Betweenness | Node or edge | Who lies on many shortest paths? | Assumes shortest paths represent the relevant process |
| Closeness | Node | Who is at a short distance from others? | Conventional form is problematic in disconnected networks |
| Eigenvector centrality | Node | Who is connected to well-connected nodes? | Can behave poorly in some directed or disconnected networks |
| Density | Network | What proportion of possible ties exists? | Strongly affected by network size and boundary |
| Reciprocity | Network or dyad | How often are directed ties mutual? | Relevant only to directed networks |
| Transitivity | Network | Are connected pairs likely to share ties? | Does not establish the mechanism producing closure |
| Assortativity | Network | Do similar actors tend to connect? | Attribute definition and category imbalance matter |
| Modularity | Partition | How strongly does a partition separate communities? | Different algorithms may produce different communities |
| Centralization | Network | Is the network organized around a few central actors? | Must specify the underlying centrality measure |
Degree Centrality
Degree centrality measures the number of direct ties attached to a node.
For an undirected binary network, the normalized degree centrality of node (i) is:
[
C_D(i)=\frac{k_i}{n-1}
]
where (k_i) is the node’s degree and (n) is the number of nodes.
A high-degree node may be:
- Popular in a friendship network
- Frequently consulted in an advice network
- Highly active in a communication network
- Highly exposed in a contact network
- Overloaded in a service-referral network
The measure does not tell the researcher which interpretation is correct. Interpretation comes from the tie definition and theory.
In directed networks, degree is divided into:
- In-degree: incoming ties
- Out-degree: outgoing ties
In weighted networks, the sum of tie weights is often called strength.
Betweenness Centrality
Betweenness centrality measures how frequently a node lies on shortest paths between other pairs of nodes.
A basic expression is:
[
C_B(i)=\sum_{s\neq i\neq t}
\frac{\sigma_{st}(i)}{\sigma_{st}}
]
where:
- (\sigma_{st}) is the number of shortest paths from (s) to (t)
- (\sigma_{st}(i)) is the number of those paths that pass through node (i)
A node with high betweenness may connect otherwise separated groups and may occupy a brokerage position.
However, high betweenness is not automatic proof that an actor controls information. The interpretation assumes that the relevant process travels along shortest paths and that the actor can recognize or exploit the position. Communication may instead follow trusted, habitual, random or institutionally prescribed routes.
Closeness Centrality
Closeness centrality measures how near a node is to other nodes in terms of shortest-path distance.
For a connected network, a common normalized form is:
[
C_C(i)=\frac{n-1}{\sum_{j\neq i}d(i,j)}
]
where (d(i,j)) is the shortest-path distance between nodes (i) and (j).
A node with high closeness may reach the rest of the connected network through relatively few steps.
Standard closeness is difficult to interpret when some nodes are unreachable. Harmonic closeness is often more appropriate for disconnected networks because unreachable nodes contribute zero rather than an infinite distance.
Eigenvector Centrality
Eigenvector centrality assigns higher scores to nodes connected to other high-scoring nodes.
It is based on the relationship:
[
x_i=\frac{1}{\lambda}\sum_j A_{ij}x_j
]
where (x_i) is the centrality score, (A) is the adjacency matrix and (\lambda) is an eigenvalue.
Unlike degree, eigenvector centrality considers the prominence of a node’s neighbours. A person with a modest number of highly prominent contacts may score more highly than someone with many peripheral contacts.
The measure requires careful treatment in directed, disconnected or unusual network structures. PageRank and related measures modify the general idea for particular contexts.
Choosing a Centrality Measure
The best centrality measure is the one that matches the hypothesized process.
| Research question | Potential measure |
|---|---|
| Who has the most direct contacts? | Degree |
| Who receives the most nominations? | In-degree |
| Who initiates the most connections? | Out-degree |
| Who connects otherwise separate groups? | Betweenness or brokerage measures |
| Who can reach others in relatively few steps? | Closeness or harmonic closeness |
| Who is connected to prominent actors? | Eigenvector centrality or PageRank |
| Who participates in strong or frequent interactions? | Weighted degree or strength |
| Who connects multiple communities? | Participation, bridging or brokerage measures |
Researchers should report multiple centralities only when each answers a meaningful question. Calculating every available metric and selecting the most interesting result increases the risk of post hoc interpretation.
Network Density
Network density is the proportion of observed ties relative to the number of ties that could exist.
For an undirected network with no self-ties:
[
D=\frac{2m}{n(n-1)}
]
For a directed network with no self-ties:
[
D=\frac{m}{n(n-1)}
]
where (m) is the number of observed ties and (n) is the number of nodes.
Density can indicate connectedness or interaction intensity, but it should not be compared uncritically across networks of very different sizes. The number of possible ties grows rapidly as nodes are added.
A dense network may support coordination and trust but may also create redundancy, conformity or excessive communication. A sparse network may be fragmented, or it may operate efficiently through a limited number of strategic links.
Reciprocity
Reciprocity describes the extent to which directed ties are returned.
In an advice network, reciprocity asks how often two people seek advice from one another. In an online network, it might ask how often following is mutual.
Researchers must state which reciprocity definition is used because software packages may calculate dyad-based and edge-based versions differently.
Transitivity and Clustering
Transitivity describes the tendency for two actors connected to the same actor to be connected to one another.
For an undirected node with degree (k_i), the local clustering coefficient is commonly written as:
[
C_i=\frac{2e_i}{k_i(k_i-1)}
]
where (e_i) is the number of ties among that node’s neighbours.
High clustering may reflect friendship circles, organizational teams, geographic proximity, shared interests or another closure process. The statistic alone cannot identify the mechanism.
Components, Cohesion and Fragmentation
A component analysis shows whether all nodes belong to one connected structure or whether the network is divided into separate pieces.
Related concepts include:
- Weak and strong components in directed networks
- Bridges, whose removal may increase fragmentation
- Cut-vertices, whose removal disconnects parts of a network
- k-cores, in which each node has at least (k) ties to other nodes in the subgraph
- Cliques, in which every pair of nodes is directly connected
These measures help researchers evaluate resilience, cohesion, vulnerability and subgroup structure.
Community Detection
Community detection identifies groups of nodes with relatively strong internal connection and comparatively weaker connection to the rest of the network.
Common approaches include modularity optimization, hierarchical methods, label propagation, random-walk methods and statistical block models.
A detected community is an algorithmic result, not automatically a meaningful social group. Researchers should compare the partition with substantive information, test its stability and report the algorithm and parameters used.
Homophily and Assortativity
Homophily is the tendency for similar actors to form ties.
Researchers may study similarity by age, occupation, location, beliefs or another attribute. Assortativity coefficients summarize the extent to which connected nodes have similar values.
Observed similarity among connected actors can arise through:
- Selection of similar partners
- Influence after a tie is formed
- Shared environments
- Common institutional rules
- Measurement or boundary effects
A cross-sectional network usually cannot separate these explanations on its own.
Centrality Versus Centralization
Centrality describes the position of a node. Centralization describes the organization of an entire network around unequal node positions.
A highly centralized network may resemble a star, with one actor connected to many peripheral actors. A decentralized network distributes connections more evenly.
Degree centralization, betweenness centralization and closeness centralization are different measures. A report should never state only that a network is “centralized” without identifying the definition.
How to Conduct Social Network Analysis
Step 1: Formulate a Relational Research Question
Begin with a question that genuinely concerns relationships or structural position.
Examples include:
- Who provides advice to whom?
- Which organizations exchange referrals?
- How do friendship groups form?
- Which actors bridge separate professional communities?
- How does information move through a collaboration network?
- How does a network change after an intervention?
Avoid beginning with software or a visually attractive dataset. The question should determine the design.
Step 2: Define the Nodes and Network Boundary
Specify who or what can appear as a node.
A boundary may be based on:
- Formal membership
- Geographic area
- Organizational affiliation
- Event participation
- Time period
- Interaction with a focal actor
- Presence in a specified database
The boundary decision affects every network measure. Researchers should explain inclusion and exclusion rules and consider whether important external actors are omitted.
Step 3: Define the Tie
State exactly what creates a tie.
Clarify:
- Whether the tie is directed
- Whether it is binary or weighted
- Whether negative ties are possible
- The observation period
- The frequency or strength threshold
- Whether multiple tie types are recorded
- Whether self-ties are permitted
A survey question such as “Who do you know?” is usually less useful than a question tied to a specific relationship and period.
Step 4: Select a Network Design
Choose among:
- Whole-network analysis
- Ego-network analysis
- Two-mode network analysis
- Multiplex or multilayer analysis
- Cross-sectional analysis
- Longitudinal analysis
The choice should follow the research question, available sampling frame and feasibility of obtaining sufficiently complete data.
Step 5: Develop an Ethical and Data-Protection Plan
Network data can reveal information about people who did not directly provide the data. One participant may name colleagues, relatives, patients or friends.
The ethical plan should address:
- Informed consent
- Third-party information
- Re-identification risk
- Sensitive ties
- Group-level harm
- Secure data storage
- Access controls
- Publication of network diagrams
- Platform terms and applicable law
- Research ethics committee or institutional review requirements
Publicly visible online data are not automatically free from ethical obligations. Researchers should consider the expectations and vulnerability of the people represented.
Step 6: Collect Relational Data
Common methods include:
Roster surveys
Participants select contacts from a complete list of eligible actors. Rosters can improve recall but require a known network boundary.
Name generators
Participants list people who meet a stated relationship criterion, such as “people from whom you sought research advice during the past three months.”
Interviews
Interviews can provide contextual information about tie meaning, history and perceived quality.
Observation
Researchers may record interactions in classrooms, meetings, workplaces or field settings.
Archival records
Possible records include co-authorship, referrals, transactions, attendance, committee membership and official correspondence.
Digital trace data
Logs from online platforms, messaging systems or collaborative tools can generate large networks. Researchers must evaluate access conditions, platform changes, bots, deleted content and the difference between observable activity and the social relationship of theoretical interest.
Step 7: Clean and Validate the Data
Check for:
- Duplicate node identifiers
- Inconsistent names
- Impossible or out-of-bound actors
- Duplicate edges
- Reversed direction
- Missing weights
- Unintended self-ties
- Conflicting nominations
- Temporal inconsistencies
- Ambiguous relationship categories
Entity resolution is especially important when people or organizations appear under different names.
Validation may include participant checks, comparison with records, inter-rater assessment or sensitivity analyses under alternative cleaning rules.
Step 8: Construct the Network
Convert the cleaned data into an edge list, adjacency matrix or two-mode matrix.
Preserve node attributes separately but link them through stable identifiers. Keep an unmodified raw file and create analysis-ready data through documented scripts or reproducible steps.
Step 9: Explore and Visualize
Begin with basic checks:
- Number of nodes
- Number of ties
- Directedness
- Weight distribution
- Isolates
- Components
- Degree distribution
- Missingness
- Unexpected clusters
Create a graph to explore the data, but do not interpret node position solely from the layout.
A force-directed layout generally places connected nodes near one another for readability. Rotating the plot, changing the layout or changing initialization can produce a different visual arrangement without changing the underlying network.
Step 10: Calculate Measures Linked to the Question
Select measures based on theory.
For example:
- Use in-degree to examine received nominations.
- Use betweenness to examine potential bridging.
- Use density to summarize overall connection.
- Use reciprocity to examine mutual exchange.
- Use community detection to explore subgroup structure.
- Use longitudinal models to study network change.
Record software, package versions, parameter choices and normalization rules.
Step 11: Conduct Statistical Analysis When Needed
Descriptive measures summarize the observed network. Inferential methods are needed when the study aims to test hypotheses, compare observed patterns with an appropriate reference process or estimate mechanisms of tie formation.
Possible approaches include:
- Permutation tests
- Quadratic Assignment Procedure
- Multiple Regression QAP
- Exponential-family random graph models
- Temporal ERGMs
- Stochastic actor-oriented models
- Latent space models
- Stochastic block models
- Network autocorrelation models
The method must fit the data structure and research question.
Step 12: Interpret Results in Context
Interpretation should connect:
- The network measure
- The tie definition
- The research setting
- The proposed mechanism
- Plausible alternative explanations
- Data-quality limitations
Avoid equating a high score with general “importance.” A node can be central according to one measure and peripheral according to another.
Step 13: Test Robustness
Useful sensitivity checks include:
- Alternative network boundaries
- Alternative tie thresholds
- Binary versus weighted analysis
- Inclusion and exclusion of isolates
- Different community algorithms
- Different treatments of missing ties
- Removal of highly influential nodes
- Comparison across time periods
- Bootstrapping or simulation where appropriate
Step 14: Report the Study Transparently
A complete report should allow readers to understand how the network was created, not merely how it was analysed.
Worked Social Network Analysis Example
Consider six students—A, B, C, D, E and F—in a study-help network. An undirected tie indicates that two students regularly help each other.
Observed ties are:
- A–B
- A–C
- B–C
- C–D
- D–E
- D–F
- E–F
The structure consists of two triangles connected by the C–D tie.
Basic Results
There are:
- (n=6) nodes
- (m=7) observed ties
- (6(5)/2=15) possible undirected ties
Density is:
[
D=\frac{7}{15}=0.467
]
Approximately 46.7% of all possible ties are present.
Degree Centrality
Degrees are:
| Student | Degree | Normalized degree |
|---|---|---|
| A | 2 | 0.40 |
| B | 2 | 0.40 |
| C | 3 | 0.60 |
| D | 3 | 0.60 |
| E | 2 | 0.40 |
| F | 2 | 0.40 |
C and D have the highest degree because each has three direct ties.
Betweenness Centrality
C and D also lie between the two clusters. Their normalized betweenness values are 0.60 in this illustrative network, while the other four nodes have zero betweenness.
This result supports the interpretation that C and D occupy potential bridging positions. If the C–D tie were removed, the two groups would become disconnected.
It does not prove that C or D controls information. Students might communicate through channels not represented by the study-help ties.
Closeness Centrality
C and D have normalized closeness values of approximately 0.714. The other students have values of 0.50.
C and D can reach all other students through fewer steps on average.
Interpretation
The example illustrates why several measures may identify the same actors for different reasons:
- Degree identifies C and D as the most directly connected.
- Betweenness identifies them as bridges between clusters.
- Closeness identifies them as structurally near the rest of the network.
In a less symmetrical network, the measures may rank actors differently.
Applications of Social Network Analysis
Sociology
Sociologists use SNA to study friendship, kinship, social support, inequality, migration, collective action, elite relations and the diffusion of ideas.
Network analysis can show how access to information or opportunity depends on structural position rather than individual characteristics alone.
Organizational Research
Organizational network analysis can examine:
- Informal advice networks
- Knowledge sharing
- Collaboration across departments
- Communication bottlenecks
- Employee onboarding
- Leadership and brokerage
- Organizational resilience
The informal network may differ substantially from the official organizational chart.
Public Health and Epidemiology
SNA can be used to investigate:
- Contact and transmission networks
- Peer influence on health behaviour
- Health-information diffusion
- Patient-referral systems
- Community support
- Implementation partnerships
The meaning of an important node differs across these applications. A high-degree actor may be a useful disseminator in an information network but a high-exposure actor in a contact network.
Education
Educational researchers study:
- Peer support
- Classroom friendships
- Collaborative learning
- Teacher advice networks
- Research supervision
- Student participation in online learning
- School-to-school partnerships
Network measures may be combined with academic, demographic or behavioural attributes, provided the dependence in the data is handled appropriately.
Communication and Social Media
Digital SNA may examine:
- Mention and reply networks
- Information cascades
- Hashtag co-occurrence
- Follower relationships
- URL-sharing networks
- Online communities
- Coordinated behaviour
Social media activity is not identical to a social relationship. Following, mentioning, replying and reposting create different networks and should not be treated as equivalent.
Political Science
Applications include:
- Legislative collaboration
- Campaign networks
- Policy coalitions
- International alliances
- Political communication
- Protest and movement networks
- Influence among institutions
Researchers must be particularly careful when interpreting hidden coordination or ideological communities from observable connections alone.
Criminology and Security Research
SNA can map co-offending, communication, transactions or affiliations in covert networks.
These datasets are often incomplete because actors conceal identities and relationships. Centrality rankings based on partial data should therefore be interpreted cautiously, especially when they may affect individuals.
Bibliometrics and Science Studies
Nodes may represent authors, papers, journals, institutions or topics.
Common networks include:
- Co-authorship networks
- Citation networks
- Bibliographic coupling
- Co-citation networks
- Keyword co-occurrence networks
- Institutional collaboration networks
A citation tie indicates a documented reference, not necessarily agreement, intellectual influence or research quality.
Community and Policy Evaluation
SNA can help evaluators understand whether organizations collaborate, exchange resources or connect otherwise separated stakeholder groups.
Repeated measurements can assess whether a programme coincides with increased connection, reduced fragmentation or broader participation. Causal attribution still requires an appropriate evaluation design.
Descriptive and Inferential Network Analysis
Descriptive SNA
Descriptive analysis summarizes the observed structure using visualizations and measures such as centrality, density, components and clustering.
It answers questions such as:
- What does the observed network look like?
- Which actors occupy particular positions?
- How connected or fragmented is the network?
- What subgroups appear in the data?
Quadratic Assignment Procedure
The Quadratic Assignment Procedure, or QAP, uses permutations to evaluate associations involving relational matrices while respecting aspects of network dependence.
MRQAP extends the approach to regression-like analyses with multiple relational predictors.
QAP does not solve every causal or modelling problem. The permutation strategy and model specification should match the research design.
Exponential-Family Random Graph Models
ERGMs model the probability of an observed network as a function of structural configurations and covariates.
An ERGM may include terms representing:
- Baseline tie propensity
- Reciprocity
- Triadic closure
- Degree-related structure
- Homophily
- Actor attributes
- Dyadic covariates
ERGMs are useful when researchers want to test whether particular patterns occur more or less often than expected under a specified probabilistic model (Robins et al., 2007).
Researchers must evaluate convergence, degeneracy, goodness of fit and sensitivity to specification.
Temporal ERGMs
Temporal ERGMs extend the ERGM framework to networks observed over time. Depending on the model, they can distinguish processes related to tie formation and persistence.
Stochastic Actor-Oriented Models
Stochastic actor-oriented models are used for longitudinal network data and can represent changes in actors’ outgoing ties. Extended models can examine the co-evolution of networks and actor behaviour (Snijders et al., 2010).
These models require repeated network observations and careful consideration of assumptions about change between observation waves.
Latent Space and Block Models
Latent space models represent actors in an unobserved space where proximity is related to tie probability.
Stochastic block models group nodes according to patterns of connection. Unlike some community-detection algorithms, blocks may represent structurally equivalent roles rather than only densely connected communities.
Social Network Analysis Software
Software Comparison
| Tool | Best suited to | Interface | Main strength | Main limitation |
|---|---|---|---|---|
| Gephi | Exploration and visualization | Desktop or browser-based | Interactive layouts and visual filtering | Limited for advanced statistical inference |
| UCINET and NetDraw | Classical SNA and teaching | Desktop menus | Broad collection of established SNA procedures | Primarily Windows-based workflow |
| R igraph | Reproducible analysis and algorithms | Code | Flexible graph manipulation and analysis | Requires programming |
| R statnet | Statistical network modelling | Code | ERGMs, temporal models and simulation | Steeper statistical learning curve |
| RSiena | Longitudinal actor-oriented models | Code | Network and behaviour co-evolution | Requires suitable panel-network data |
| Python NetworkX | General graph analysis and integration | Code | Extensive algorithms and Python ecosystem | May be slower than specialized libraries on very large networks |
| Pajek | Large-network analysis | Desktop | Efficient handling of large networks | Interface is less familiar to many beginners |
| NodeXL | Spreadsheet-oriented exploration | Excel-based | Accessible to spreadsheet users | Less flexible than fully scripted workflows |
Choosing Software
Choose software according to the task:
- Use Gephi for exploratory visualization and presentation.
- Use UCINET for menu-driven classical network analysis.
- Use R igraph or Python NetworkX for reproducible general analysis.
- Use statnet for ERGMs and related statistical models.
- Use RSiena for stochastic actor-oriented longitudinal analysis.
- Use Pajek when working with very large networks and established Pajek workflows.
- Use NodeXL when spreadsheet integration is a priority.
A visual interface may help beginners explore data, but code-based analysis generally makes complex cleaning, repeated analysis and reproducibility easier.
Artificial Intelligence and Modern Network Research
Artificial intelligence can support SNA, but it does not replace network measurement or theory.
Relation and Entity Extraction
Natural-language processing and large language models can help identify people, organizations and proposed relationships in documents.
Human validation remains necessary because automated systems may:
- Merge different entities
- Split one entity into multiple identities
- Misread negation
- Confuse allegations with verified relationships
- Infer ties not supported by the source
- reproduce bias in the underlying text
Researchers should preserve the evidence used to construct each extracted tie.
Graph Machine Learning
Graph machine-learning methods can support:
- Node classification
- Link prediction
- Anomaly detection
- Graph classification
- Recommendation
- Representation learning
Graph neural networks can combine relational structure with node or edge features. Their predictive performance does not automatically produce a social explanation, identify causation or validate the meaning of the network.
AI-Assisted Coding and Analysis
Generative AI may help write draft code, explain package errors, prepare data dictionaries or document an analysis. All code and output should be checked.
Sensitive network data should not be uploaded to third-party AI systems without authorization, appropriate agreements and an assessment of confidentiality risks.
Recent Good Practice
Modern SNA increasingly benefits from:
- Scripted and version-controlled workflows
- Open or synthetic teaching datasets
- Preregistration where appropriate
- Model diagnostics
- Sensitivity analysis
- Temporal and multilayer representations
- Integration of network and qualitative evidence
- Clear documentation of platform and API limitations
- Ethical review that considers connected nonparticipants
Advantages of Social Network Analysis
It Makes Relationships Analytically Visible
SNA can reveal patterns that cannot be obtained by studying actor attributes alone.
It Operates at Multiple Levels
Researchers can examine nodes, dyads, subgroups and whole networks within one conceptual framework.
It Identifies Structural Roles
The method can identify hubs, isolates, bridges, cohesive subgroups and core–periphery patterns.
It Connects Visualization and Measurement
Graphs can support exploration and communication, while formal measures provide more systematic descriptions.
It Supports Interdisciplinary Research
SNA is used across sociology, public health, education, organizational studies, communication, political science, biology, information science and computer science.
It Can Examine Change
Longitudinal network data enable researchers to study tie formation, dissolution and structural evolution.
Limitations of Social Network Analysis
The Boundary Problem
Every network is constructed through inclusion rules. Excluding external actors may make internal actors appear more central or the network appear more cohesive.
Missing Data Can Alter Structure
Nonresponse and missing ties can distort paths, centrality, components and communities. The direction and magnitude of distortion depend on which data are missing.
Tie Measurement May Be Ambiguous
A self-reported friendship, observed message and digital follow are different measures. Researchers should not assume that one is a valid proxy for another.
Network Measures Are Context-Dependent
Centrality is not a universal measure of influence. Density is not universally beneficial. Brokerage is not automatically power.
Visualizations Can Mislead
Layouts, node sizes, colours and filtering choices can exaggerate separation or prominence. All visual encodings should be explained.
Causal Inference Is Difficult
Selection, influence, shared environments and unmeasured factors can generate similar observed patterns. Cross-sectional SNA rarely separates them.
Large Networks Create Computational and Interpretive Challenges
Large digital networks may contain bots, duplicate identities, inactive accounts, automated events and platform-generated connections. Scale does not guarantee validity.
Privacy Risks Are Relational
Removing names may not be sufficient. Unique network structures can make people or groups identifiable, and diagrams may expose sensitive connections.
Common Mistakes in Social Network Analysis
Treating Social Media as the Definition of a Social Network
SNA predates contemporary social-media platforms and applies to many forms of relational data.
Starting With a Graph Instead of a Question
An attractive visualization is not a substitute for a clear research problem and tie definition.
Using “Influencer” as a Synonym for High Centrality
A high centrality score describes a structural property. Influence requires a defined outcome and supporting evidence.
Ignoring Direction and Weight
Converting directed or weighted ties into undirected binary ties may remove theoretically important information.
Comparing Density Across Very Different Network Sizes
Density commonly decreases as networks become larger. Comparisons should account for size, opportunity structure and boundary definitions.
Using Conventional Closeness in a Disconnected Network
Unreachable nodes make standard shortest-path closeness problematic. Harmonic or component-specific measures may be preferable.
Removing Isolates Without Explanation
Isolates are part of the observed network when they meet the inclusion criteria. Removing them changes density and potentially the interpretation.
Treating Algorithmic Communities as Natural Groups
Community output depends on the algorithm, parameters and random initialization. Substantive validation is required.
Ignoring Missing Actors
Whole-network analysis is particularly sensitive to nonresponse. Response rates and missing-data treatment should be reported.
Reporting Metrics Without Definitions
State whether measures are normalized, weighted, directed, component-based or calculated using a particular software convention.
How to Report a Social Network Analysis Study
A strong methods section should report:
- The research question and theoretical rationale
- Node definition and eligibility criteria
- Network boundary
- Tie definition and observation period
- Directedness, weight and tie types
- Network design
- Sampling or recruitment procedure
- Response and nonresponse information
- Data sources and collection instruments
- Cleaning and entity-resolution procedures
- Missing-data treatment
- Ethical approval and privacy protections
- Software and version information
- Measures and formulas or references
- Normalization and parameter choices
- Layout algorithm and visual encoding
- Statistical model specification
- Convergence and goodness-of-fit checks
- Robustness or sensitivity analyses
- Limitations on interpretation and generalization
Conclusion
Social network analysis studies how patterns of relationships shape the positions of actors and the structure of groups, organizations and wider systems. Its value comes from connecting a well-defined relational question to appropriate data, measures and theory.
Reliable SNA requires more than calculating centrality or producing a network diagram. Researchers must define the network boundary, measure ties carefully, address interdependence and missingness, protect connected individuals, test the robustness of results and avoid interpreting structural association as automatic evidence of influence or causation.
