
Few papers have made me stop and rethink my own position as much as Cameron Atkinson's recent article Third-Party Interpretivism: Large Language Models and the Displacement of Interpretive Authority in Big-Q Research. Not because I agree with its conclusions. But because it identifies what I believe is the real point of disagreement in the current debate about AI and qualitative research.
Much of the discussion around generative AI focuses on practical concerns. Language models hallucinate. They make mistakes. They may overlook important nuances. They can produce plausible but superficial interpretations. These are all legitimate concerns, and they deserve careful methodological attention.
Atkinson's argument is different. He argues that the fundamental problem is not the quality of AI-generated interpretations. The problem begins much earlier.
The moment a pretrained language model contributes an interpretive suggestion, interpretive authority has already been displaced. Because the internal process that generated that suggestion cannot be reconstructed as part of the researcher's own situated reasoning, it can never become fully accountable, regardless of what happens afterwards.
David Morgan, a co-author and co-editor for some of my publicatins, recently described this as the "original sin" argument. I think that metaphor captures Atkinson's position remarkably well. If this argument is correct, then returning to the data, checking the evidence, asking follow-up questions, rejecting suggestions, refining interpretations, or documenting analytical decisions cannot undo what has already happened. The interpretive contribution has entered the research process through a source whose reasoning cannot be reconstructed.
This is a much deeper criticism than concerns about hallucinations or poor prompting. It is not an argument about the current state of the technology. It is an argument about what interpretive accountability itself requires. And it is here that our positions fundamentally diverge.
Atkinson’s argument depends on a particular understanding of accountability. For an interpretive contribution to be warranted, he argues, its formation must be reconstructable as part of the researcher’s own situated judgement. This sounds plausible until we apply the same standard to ordinary qualitative analysis.
Suppose you are reading an interview transcript and a participant’s statement suddenly reminds you of Bourdieu rather than Goffman. Can you reconstruct exactly why that association occurred at that moment? Probably not.
The idea may have been triggered by something you read years ago, a seminar discussion, a previous research project, or a half-remembered sentence from a book. It may have emerged from several influences that can no longer be separated. But you may still be able to explain why Bourdieu provides the more convincing interpretation.
You can return to the transcript. You can identify the passages that support the reading. You can compare it with alternative interpretations. You can examine contradictory cases. You can explain why the concept illuminates the material and where its limits lie. These are two different questions:
Where did the idea come from?
Why should this interpretation be considered credible?
The first concerns the origin of an idea. The second concerns its justification. Qualitative research has never required us to reconstruct every cognitive pathway through which an interpretive insight emerged. It requires us to make our interpretations accountable through transparent engagement with the material.
This distinction matters because ideas rarely arise through a fully traceable sequence of conscious reasoning. They may be prompted by a colleague’s remark, a supervisor’s question, a conference presentation, a theoretical concept, or an unexpected comparison with another case. None of these influences is automatically treated as a displacement of interpretive authority.
We do not ask whether the colleague’s suggestion can be reconstructed as part of our own cognition. We ask what we did with it.
The same distinction should apply when an idea is suggested by an AI system. This does not mean that human and machine contributions are identical. They are not. Nor does it mean that AI-generated suggestions should be accepted uncritically. They should not.
It means only that opacity at the point where an idea first enters the analysis is not, by itself, enough to disqualify that idea. The relevant methodological question is not whether we can reconstruct every influence on the researcher’s thinking. It is whether the researcher can take responsibility for the interpretation that follows.
We may not be able to reconstruct the origin of every insight. We can however reconstruct, examine, and defend the argument we build from it.
Atkinson’s second concern is that an AI-generated suggestion shapes what the researcher notices before that suggestion can be evaluated. He calls this “seeding.”
Once a model proposes a set of topics, categories, or interpretations, the researcher has already been exposed to a particular framing. Even if the suggestion is later rejected, it may continue to influence what appears salient. That concern is not trivial. First impressions do matter. Initial framings can narrow attention, produce premature coherence, and make alternative readings less visible.
But this is not unique to AI-assisted analysis. Qualitative analysis always begins from somewhere. Researchers enter the field with prior knowledge, disciplinary training, theoretical preferences, research questions, interview guides, expectations, and first impressions. The order in which transcripts are read matters. So does the first case that appears unusual, the first code that seems productive, or the first discussion with a colleague.
An initial coding scheme is also a seed. It reflects what the researcher noticed during an early reading of the material. As the analysis continues, some codes prove too broad, others too narrow, and some turn out to be based on assumptions that the data do not support. Researchers then revise the scheme. They rename codes, merge categories, split them apart, recode earlier material, identify exceptions, and sometimes abandon the original framing entirely.
This is not a methodological failure. It is a normal part of qualitative analysis.
John Seidel’s (1998) well-known formulation of qualitative analysis as a recursive process of noticing, collecting, and thinking captures this well. What we notice influences what we collect. What we collect changes how we think. What we think changes what we notice next. The process is iterative precisely because first impressions are provisional. The same should apply to AI-generated suggestions.
A model may propose an initial framing. The researcher may find parts of it useful, parts superficial, and parts plainly wrong. Further questions may expose contradictions. Closer reading may reveal that the proposed idea collapse important differences between participants. A negative case may require the interpretation to be revised. The first output is not the analysis. It is one contribution to an evolving process. A seed is an influence, not a verdict.
What matters methodologically is whether the process allows that influence to be examined, challenged, revised, and discarded. If it does not, then AI use becomes problematic. But the problem is not that a suggestion entered the analysis before evaluation. That happens in human analysis as well. The problem arises when researchers treat an initial suggestion as a result rather than as something to think with.
If seeding before evaluation were disqualifying in principle, then ordinary human analysis would face the same objection. Human researchers are also influenced by their first readings, initial coding decisions, and theoretical expectations before they have fully evaluated them. Qualitative methodology has never solved this problem by demanding an “unseeded researcher”. It addresses it through iteration, reflexivity, comparison, and return to the data. The same standard should govern AI-assisted analysis.
One aspect of Atkinson's paper surprised me. While he rejects the interpretive use of Large Language Models, he is considerably more accommodating toward topic modelling and other computational approaches that remain closer to statistical analysis. At first glance, this distinction seems intuitive. Topic models analyse patterns of word co-occurrence rather than generating natural language interpretations. Their operations appear more constrained and therefore, perhaps, more transparent.
But are they really transparent in the way Atkinson requires? Suppose a topic model assigns one document a probability of 31% for Topic A and 28% for Topic B. Can the researcher explain why those exact values were produced? Can they reconstruct why particular words contributed to one topic rather than another, or why the model converged on this particular structure instead of a slightly different one? In practice, very few researchers can.
Researchers inspect the output, compare it with the underlying documents, judge whether the topics make substantive sense, refine their interpretation, and often adjust parameters or rerun the analysis. In other words, they evaluate the results rather than reconstructing the internal computation. That sounds remarkably similar to how many researchers work with LLMs. This is not to say that topic models and language models are the same. They are based on very different computational principles and support different kinds of analytical work.
The point is different. If interpretive accountability can be maintained by critically evaluating the output of one computational method against the underlying data, why should that possibility be rejected in principle for another? The distinction becomes even more difficult to sustain when we consider what the two approaches actually do.
Topic models identify statistical regularities in patterns of word use. They do not understand arguments, distinguish between speakers, recognise irony, or follow a participant's reasoning across an entire interview. Large Language Models can.
Of course, they can also misunderstand, oversimplify, or hallucinate. Their interpretations are never automatically trustworthy. But neither are they performing a simpler version of topic modelling. They are engaging with language at a level that is much closer to the interpretive work qualitative researchers themselves undertake. For me, this raises a broader methodological question.
If we are prepared to accept one form of algorithmic opacity because researchers critically evaluate its outputs, why should the same evaluative process become epistemologically impossible simply because the algorithm is more sophisticated?
At this point, I find myself wondering whether the debate is really about AI at all. Or whether AI is causing us to redefine what we mean by good qualitative research.
The standard proposed by Atkinson is demanding. It requires that interpretive contributions must be reconstructable as part of the researcher's own situated judgement. If they are not, interpretive authority has already been displaced. But this has never been the standard qualitative researchers applied to themselves.
We have never expected researchers to explain precisely why a particular quotation reminded them of Bourdieu rather than Goffman. We have never required them to reconstruct the sequence of experiences that produced an abductive insight. Nor have we asked them to document every conversation, every book, or every previous project that may have influenced the development of an interpretation. Instead, the important questions to ask are:
That is the standard qualitative researchers have developed over decades. So why should that standard suddenly change because AI participated in the analytical process? This is not suggesting that AI deserves lower standards. Nor is it arguing that researchers should accept AI-generated interpretations uncritically.
Quite the opposite. AI-assisted research should be held to exactly the same methodological expectations as any other qualitative research: careful engagement with the data, reflexivity, transparency of analytical decisions, openness to alternative interpretations, and responsibility for the conclusions that are ultimately presented.
What I question is the introduction of a new epistemic requirement that has never previously been demanded of human researchers: the ability to reconstruct the complete origin of every interpretive idea. That is a different claim.
It is a claim about what counts as knowledge in qualitative research. If we adopt that standard, we are not simply rejecting AI-assisted interpretation. We are redefining interpretive accountability itself.
If accountability does not depend on reconstructing how every interpretive idea first emerged, then where does it reside? For me, accountability is not a property of the AI. It is a property of the research process.
A qualitative interpretation does not become accountable because we know exactly how the idea was generated. It becomes accountable because it can be questioned, challenged, refined, and justified through systematic engagement with the data. This shifts the focus away from the internal workings of the model and towards the interaction between the researcher and the data.
An accountable AI-assisted analysis should therefore allow researchers to:
None of these activities requires us to know how the model internally arrived at its first suggestion. They require the researcher to remain intellectually engaged throughout the analytical process. This is why - from the very beginning when I noticed the potential of LLMs – moved away from the idea of coding. Even though newer models have become better in coding structured data sets, I still think using LLMs for coding data is not utilizing their strength to its full potential. Coding is one analytical technique. It is not the essence of qualitative analysis.
Researchers are trying to understand experiences, explain social processes, identify mechanisms, explore contradictions, develop concepts and construct convincing interpretations. LLMs gives us the opportunity to engage with these questions much more directly than previous generations of qualitative software.
Instead of asking: "Can LLMs generate codes?", after some first exploration of the data, we can ask:
These are methodological questions. The AI is not replacing the researcher. It becomes a conversational partner that helps the researcher interrogate the material more systematically. The responsibility for deciding what constitutes a convincing interpretation remains exactly where it has always belonged—with the researcher.
The future of qualitative research will not be decided by how ideas are generated, but by how interpretations are justified.