“How many interviews do I need?” is one of the most common questions in qualitative research. And one of the most common answers is: “Until you reach saturation.” Many theses state that “data collection continued until saturation was reached”. But what does that actually mean? How do you know when you have reached it? And how do you convince an examiner that you did?
This post explains what data saturation is, where the idea comes from, the different types of saturation, practical ways to assess it, and how to report it clearly and honestly in your thesis.
What is data saturation?
In broad terms, data saturation is the point in qualitative data collection and analysis where additional data no longer produce new information, themes or insights relevant to the research question. Collecting more data at that stage would mostly repeat what you already know.
Saturation is used as a guide for deciding when to stop sampling, and as a quality indicator suggesting that the dataset adequately covers the phenomenon being studied.
Where the concept comes from
The idea originated in grounded theory, developed by Glaser and Strauss in the 1960s, where it was called theoretical saturation. In that context, saturation means that the categories of an emerging theory are fully developed: their properties and dimensions are well described, and the relationships between categories are clear. New data no longer add to the theory.
Over time, the term spread to many other qualitative approaches, often in a looser form meaning simply “no new themes are appearing”. This broader use is common, but it has also attracted criticism, as discussed below.
Different types of saturation
Methodological literature describes several related ideas. Knowing the differences helps you use the term precisely.
1. Theoretical saturation
The original grounded theory concept. It focuses on the development of categories and theory, not on counting new codes. It is closely linked to theoretical sampling, where you choose new participants specifically to develop emerging concepts.
2. Code saturation
The point at which no new codes appear in additional interviews. Some researchers describe this as having “heard it all”.
3. Meaning saturation
The point at which you fully understand the codes and issues, including their nuances and variations. Some empirical studies on saturation have suggested that code saturation may be reached relatively early, while meaning saturation may require more data. In other words, you may have identified all the main issues before you fully understand them.
4. Data saturation
Often used as a general term for the point at which new data produce little or no new information. It is sometimes used interchangeably with code saturation.
Why saturation is debated
Saturation is widely used, but several researchers have criticised how it is applied:
- Vague claims. Many studies claim saturation without explaining how it was assessed.
- Mismatch with methodology. In some approaches, such as reflexive thematic analysis, meaning is seen as generated through interpretation rather than “found” in the data. Proponents of this approach, including Braun and Clarke, have argued that saturation may not be a coherent concept for their method, because new insights can always emerge with further reflection.
- Used after the fact. Some studies decide sample size in advance for practical reasons and then claim saturation afterwards.
This does not mean you should avoid the concept. It means you should use it thoughtfully, choose the version that fits your methodology, and explain your assessment. Some researchers also use alternative ideas such as information power, proposed by Malterud and colleagues, which links sample size to study aim, sample specificity, use of theory, quality of dialogue and analysis strategy.
How many interviews does saturation require?
There is no fixed number. Some empirical studies of interview data have reported that most codes appeared within the first 9 to 17 interviews in relatively homogeneous samples with focused research questions. However, these findings depend heavily on the study context and should not be used as automatic targets.
Saturation tends to require more data when:
- The research question is broad or exploratory.
- The population is diverse.
- The phenomenon is complex.
- Interviews are short or less in-depth.
- You aim to develop theory rather than describe experiences.
It may be reached sooner when the question is narrow, participants are very similar and highly knowledgeable, and interviews are rich.
How to assess saturation in practice
The key principle is that saturation cannot be judged without ongoing analysis. If you collect all your interviews first and analyse them afterwards, you cannot know when saturation occurred. So:
1. Analyse as you collect
Transcribe and code interviews soon after conducting them. Keep a running list of codes and emerging themes. This allows you to see when new data stop adding new ideas.
2. Use a saturation table or code tracking grid
Create a table with codes as rows and interviews as columns. Mark where each code first appears. As you progress, you can see how many new codes each interview contributes. When several consecutive interviews add few or no new codes, you may be approaching code saturation.
3. Set a stopping rule in advance
Some researchers use a stopping criterion, for example: after an initial set of interviews, continue until a fixed number of consecutive interviews produce no new themes. Guest and colleagues, among others, have proposed approaches of this kind. Defining your rule before starting makes your claim more transparent and less arbitrary.
4. Look for meaning, not just new codes
Ask whether you understand each theme in depth. Can you describe its variations, conditions and consequences? Do you have enough examples to explain contrasts between participants? If not, you may need more data even if no new codes appear.
5. Check across subgroups
If your sample includes different groups, such as managers and employees, check whether saturation has been reached within each group, not only overall. A theme may be saturated for one group but not for another.
6. Seek disconfirming cases
Deliberately look for participants who may challenge your emerging findings. If their data fit within existing themes or refine them without creating major new categories, this supports your saturation claim.
7. Discuss with your supervisor or co-coders
Review your codebook and emerging themes with others. They may spot gaps you have missed.
An example of documenting saturation
Imagine a study of 20 interviews with early-career teachers about workload. The researcher tracks new codes after each interview:
- Interviews 1–5: 42 codes identified.
- Interviews 6–10: 11 new codes.
- Interviews 11–15: 3 new codes, all minor variations.
- Interviews 16–18: no new codes.
Based on a pre-defined rule of three consecutive interviews without new codes, the researcher stops at 18 interviews but conducts two more with teachers in rural schools, a subgroup under-represented so far. These add depth to an existing theme about travel time but no new themes. The researcher then reports both code saturation and meaning-related judgements. (The numbers are illustrative.)
Saturation beyond one-to-one interviews
Most discussions of saturation focus on interviews, but the same logic applies to other qualitative data, with some differences.
- Focus groups: Each group produces data from several people interacting. Saturation is usually judged at the level of groups rather than individuals. If your design compares types of participants, such as parents and teachers, consider whether each type has had enough groups for themes to stabilise.
- Document analysis: When analysing policies, reports or media articles, you can track whether additional documents add new codes. Because documents are often easier to obtain than interviews, you may be able to sample more widely and test saturation more thoroughly.
- Observation: In ethnographic or observational work, saturation relates to whether further observation sessions reveal new patterns of behaviour or interaction. Field notes and memos help you track this over time.
- Open-ended survey responses: With large numbers of short written answers, saturation can often be reached for common themes, but responses may lack depth. Meaning saturation can be harder to achieve with brief text.
Whatever your data type, the principle remains the same: analyse as you go, track what each new piece of data contributes, and explain the basis for stopping.
How to report saturation in your thesis
Avoid simply writing, “Data saturation was reached.” Instead, explain:
- Which type of saturation you mean and why it fits your methodology.
- How and when you analysed data during collection.
- The criteria or stopping rule you used.
- Evidence, such as a code tracking table in an appendix.
- Any subgroup checks and disconfirming case analysis.
- Limitations, such as practical constraints on recruitment.
A short example: “Interviews were coded concurrently with data collection. Following a pre-specified stopping rule, recruitment ceased after three consecutive interviews generated no new codes (interviews 16–18). Two additional interviews with rural teachers confirmed that existing themes captured their experiences. A code emergence table is provided in Appendix D.”
What if you could not reach saturation?
Sometimes practical constraints, such as limited access to participants, prevent you from reaching saturation. This is not necessarily fatal. Be honest:
- Explain the constraints.
- Describe the richness and depth of the data you do have.
- Consider whether “information power” or another framework better describes the adequacy of your sample.
- Discuss implications for your findings, such as themes that may be less fully developed.
Honest reporting is far more credible than an unsupported saturation claim.
Common mistakes to avoid
- Claiming saturation without describing how it was assessed.
- Analysing data only after all interviews were completed, then claiming saturation.
- Using saturation language with a methodology that rejects the concept, without discussion.
- Treating a number from another study as a universal rule.
- Ignoring subgroups within the sample.
- Confusing “no new codes” with full understanding of the themes.
Final thoughts
Data saturation is a useful guide for deciding when you have enough qualitative data, but it is not a magic phrase. Choose the version of saturation that fits your methodology, analyse data as you collect it, track new codes and deepening meaning, and document your decisions. When you can show clearly why you stopped collecting data, your sample size becomes defensible, and your findings more trustworthy.
How are you deciding when to stop collecting qualitative data? Share your approach or questions in the comments.