
António Pedro Costa, Faculty of Psychology and Educational Sciences, University of Porto (Portugal)
António Pedro Costa is a researcher at the Centre for Research and Intervention in Education (CIIE), Faculty of Psychology and Educational Sciences, University of Porto, Portugal. He is one of the researchers behind the qualitative data analysis software webQDA (webqda.net) and teaches research methodology courses. He coordinates the Ibero-American Congress on Research Methods (ciaiq.ludomedia.org) and the World Conference on Qualitative Research (wcqr.ludomedia.org), and serves as Editor-in-Chief of the journal New Trends in Qualitative Research (NTQR). His research focuses on the use of technologies in research, with particular emphasis on Artificial Intelligence, research ethics, mixed methods, and data curation. In 2023, he received an Honourable Mention in the University of Aveiro Researcher Award and has coordinated publicly and privately funded research projects worth more than €3 million.
On 26 October 2026, at an ECAI2025 workshop in Bologna, I delivered a presentation on the Reflexive Uncertainty Framework (RUF). Its central claim is that, in AI-assisted qualitative analysis, uncertainty is not a flaw to be eliminated; it is analytical evidence, provided we make it methodologically visible.
Three days later, at nine in the morning, I sat down for Mohit Bansal’s plenary talk. The topic: trustworthy planning agents, uncertainty-calibrated reasoning, collaboration between agents, and controllable multimodal generation. The room was full, with participants mainly from computer science.
For an hour, I listened to a community that speaks a different language arrive, by an entirely different route, at the same question I had encountered days earlier. It is worth explaining where the two paths meet and, above all, where they diverge.
Let me begin with what surprised me most in our own material. The errors I found difficult to detect were not those where the system hesitated. They were those where it was certain.
What calibration means and what it cannot mean
Calibration, in Bansal’s sense, means aligning the confidence a system expresses with its probability of being right. A well-calibrated model that says “80%” is right four times out of five. This is a desirable and measurable property, and in mathematics or code, we know how to measure it because there is a correct answer against which to compare the result.
In an interview, there is no true or false interpretation. An interpretation is more or less substantiated, more or less consistent with the theoretical framework, and more or less attentive to context. Transferring calibration literally to qualitative research would mean importing a metric without the object it measures.
This does not invalidate the underlying intuition; it simply requires us to translate it. If we cannot ask, “How often is this system right?”, we can ask when it hesitates, what it hesitates over, and what that hesitation tells us about the data. That was where I started.
Three moments when doubt becomes data
The framework I presented, RUF, is not a standalone model; it develops one stage of the AbductivAI model in greater depth: the stage where the machine makes its reasoning explicit and the researcher can challenge it. It documents three moments:
- First, probabilistic outputs: a vector model transforms excerpts and categories into vectors, calculates semantic proximities, and returns probability distributions.
- Second, zones of interpretive ambiguity: when two categories are too close, or when the distribution is low and diffuse, a second layer is triggered, in which an LLM must provide a discursive justification for the relevance of the competing categories.
- Third, reflexive commentary refers to the moment when the researcher articulates their disagreements and the reasons behind them.
Put like this, it sounds like architecture. It is more useful to see what happens in a specific case.
Empathy 72%, permissiveness 68%
The corpus used was small and deliberately modest: nine essays by fifth-grade pupils responding to “If I were a teacher”. One of them said, roughly:
“I would stop the activity and let them do whatever they preferred.”
The vector model assigned Empathy a score of 72% and Permissiveness 68%. Four percentage points separated two readings that are not variations of one another. They are different theories about what that child thinks a teacher should be.
When asked, the LLM offered a justification: the gesture is empathetic, but the lack of regulation may indicate permissiveness. In other words, it wavered. And it is precisely this wavering that constitutes the data. Not because the system is malfunctioning, but because the sentence really is ambiguous, and the ambiguity is the information. The researcher decides whether this is relational care or an erosion of authority. But now the decision is made with the tension in view, rather than receiving it already resolved by a number that did not know it was resolving anything.
Authority 78% and what the number missed
The reverse case is more interesting, and it is the one that interests me most. Another essay said that struggling pupils would sit at the back of the classroom so as not to disturb those who wanted to learn.
The system classified this as “authority”, with 78% confidence. It is not wrong. It is simply looking at the wrong side of the sentence. What matters here is not authority, but the symbolic weight of spatial exclusion: a ten-year-old child matter-of-factly describing a pedagogy of separation. No percentage takes me there.
Notice the asymmetry with the previous case. When the model hesitated, it alerted me. When it was confident, it “silenced me”. A confident but superficial classification is more dangerous than an uncertain one, because it leaves no trace. This is why high confidence cannot be the stopping criterion for analysis and why reflexive commentary within the framework is mandatory precisely where the divergence between human and machine is greatest.
What no one classified
There was, however, one finding that troubled me more than the other two. The nine essays barely mention assessment or lesson planning. This is a striking absence in children’s accounts that describe teaching chiefly as behaviour management and almost never as preparation or assessment. That absence is a finding.
The system got there without realising it. The categories existed and received scores: formal assessment 22%, planning 18%. When questioned, the LLM merely noted that there was no explicit reference. In other words, the numbers were clearly visible; what was lacking was the interpretive operation necessary to comprehend them. For a classifier, a low score is indistinguishable from “this category does not apply”, and to interpret it as a meaningful silence requires knowing what one would expect to find in a text about being a teacher. This horizon of expectation is not in the model. It is in the reader. We recorded this as a recurring limitation, and of the three things we found, it is the one that seems least likely to be resolved by larger models. Scaling up improves the detection of what is there. Absence is not a detection problem.
This was also where the plenary felt both closest and furthest away. Bansal presented multimodal classroom video analysis to locate episodes and bring together speech, gestures and use of space across hours of footage that no one reviews in full. Its usefulness is evident. But a four-second silence is an observable datum and “discomfort” is an interpretation; the distance between them is traversed through context that is not in the video. The more effectively systems identify what is present, the greater the care we must take with what leaves no trace: what was not said, what was not taught, and who did not speak.
There is an irony here worth naming. In volume 22, issue 3 of New Trends in Qualitative Research (NTQR), Nicole Brown argues that rigour in qualitative research rests on embodied reflexivity rather than disembodied neutrality: the researcher’s body, sensory attention and response to the situation are part of how data are produced. A system that reads bodies in video is, by definition, the opposite: the most fine-grained possible reading of bodies, performed without a body of its own. It returns positions, durations and trajectories, omitting precisely what Brown places at the centre.
It is no coincidence that the example the classifier missed—the children seated at the back of the classroom—involves bodies arranged in space by an institutional order. It was all there in the text, and what was missing was not detection capacity but an understanding of what that arrangement means to those who experience it. There is a more precise name for what is missing here: epistemic injustice. The system does not misrepresent those children’s experience; it has no way of representing it.
Persuasion works both ways
Bansal spoke of agents that should accept valid corrections and resist incorrect arguments. In research practice, the problem is symmetrical, and we are half of it.
A fluent response seems convincing even when its connection to the data is weak. While competent prose is now readily available, substantiation remains valuable. And a question that already contains my preferred interpretation steers the system toward confirming it. “Don’t you think this silence reveals resistance?” is not a question; it is a request for agreement. If I never disagree with the system and the system never disagrees with me, someone is not doing their job.
This has practical implications that go beyond noble intentions. In the AbductivAI model underpinning this work, disagreement is a step in the workflow: divergence between the human coder and the agent is recorded and, above all, an explanation is required for why one would choose category X and the other category Y. The human decides the uncertain cases; the agent acts as a second coder whose role is to offer an alternative reading. A disagreement that is not written down did not happen.
Let me add a caution I have learned to observe: the justification a model gives for its own reasoning is a plausible justification, not a window into its operation. I treat it as I would a participant’s rationalisation: as data to be analysed, not as the truth of the process. Its value lies in what it opens up for discussion, not in what it reveals about the mechanism.
The objection I have to accept
Someone will say that this is enthusiasm in search of methodological justification, that we are adapting epistemology to the tool rather than the other way around. Such a claim is a serious objection, and I do not consider it resolved. The article acknowledges two weaknesses that belong to us:
- The first: the corpus is small and specific. Nine essays, one context, one language. What we describe needs testing with larger, multilingual datasets and within other interpretive traditions.
- The second is more uncomfortable. The entire value of the framework depends on someone actually sustaining the role of co-analyst by reading the justifications, writing the commentaries and recording disagreements. Such a role cannot be assumed by default. In practice, there is a temptation to accept the well-formulated proposal and move on; however, a mechanism that makes uncertainty visible is useless to someone who does not pay attention. Making doubt documentable is a necessary condition.
There is a connection here that I only made later. I tend to frame this second weakness as an individual requirement: the researcher who remains vigilant. But if the co-analyst stance can’t be assumed by default, it can’t be left to the analyst’s virtue. In another study, with Isabel Pinho and Cláudia Pinho, we proposed considering generative AI governance in educational research at three levels: macro, meso and micro.
My reading today is that a mechanism such as RUF operates at the micro level and depends entirely on the other two. Without reporting requirements in journals, methodological training in doctoral programmes, or evaluation criteria that recognise the value of documenting uncertainty, making doubt visible will remain a voluntary act for a few. Under deadline pressure, voluntary actions are often the first to be sacrificed.
Two communities, the same problem
It should be said that this goes against the grain. In the most widely publicised uses of AI in qualitative research, the gain sought is one of scale, and NTQR itself published a good example: “The End of Traditional Focus Groups? Scaling Up Qualitative Research Quick, Yet Maintaining Depth of Data with Larger Samples”. It describes an online discussion with one hundred participants, analysed in real time, in which the authors report the limitation they encountered: being unable to probe the responses further. What I propose here moves in the opposite direction: using the machine not to analyse faster, but to slow down precisely where analysis usually speeds up.
What I brought back from Bologna was not a technique. I realised that two communities that rarely read each other’s work—the one that measures calibration in benchmarks and the one that has debated reflexivity since Bourdieu—are converging on the same question from opposite directions: what should a system do when it does not know, and what should we do about it?
I am writing this from Paris, where I am attending UNESCO’s Digital Learning Week (2026). The convergence continues, but it now uses different terminology. The word running through the entire programme is friction. The programme includes sessions on cognitive friction in design, the costs associated with frictionless experiences, and strategies to limit AI-induced friction; one session is titled “What AI cannot tell you: Knowledge integrity, epistemic injustice and transcultural AI ethics”. I recognise the discussion, even though there the focus is chiefly on friction for the learner and here on friction for the researcher. The question is: what is lost when friction is removed from a process that needs it?
The answer seems fairly obvious to me: if we ever evaluate these integrations, let it be not by the time saved, but by the depth of the interpretations, the attention to differences, the preservation of participants’ voices and the transparency of decisions. We alone bear scientific and ethical responsibility, which no agent, no matter how well calibrated, will assume for us.
References
Bansal, M. (2025). Trustworthy Planning Agents for Collaborative Reasoning and Multimodal Generation. https://ecai2025.org/keynote-speakers/
Brown, N. (2026). The Body in Qualitative Research: From Disembodied Knowledge to Corporeal Praxis. New Trends in Qualitative Research, 22(3), e1472. https://doi.org/10.36367/ntqr.22.3.2026.e1472
Costa, A. P., & Bem-Haja, P. (2025). Reflexive Uncertainty AI for Qualitative Data Analysis. In E. Marengo, M. Strian, & M. IPonticorvo (Eds.), Proceedings of the 2nd International Workshop on Education for Artificial Intelligence (edu4AI2025) (pp. 13–23). European Association for Artificial Intelligence (EurAI). https://ceur-ws.org/Vol-4114/7_paper.pdf
Costa, A. P., Bryda, G., Christou, P. A., & Kasperiuniene, J. (2025). AI as a Co-researcher in the Qualitative Research Workflow: Transforming Human-AI Collaboration. International Journal of Qualitative Methods, 24. https://doi.org/10.1177/16094069251383739
Mohd Anis, A., & Olisa, N. (2024). The End of Traditional Focus Groups? Scaling Up Qualitative Research Quick, Yet Maintaining Depth of Data with Larger Samples. New Trends in Qualitative Research, 20(1), e799. https://doi.org/10.36367/ntqr.20.1.2024.e799
Pinho, I., Costa, A. P., & Pinho, C. (2025). Generative AI Governance Model in Educational Research. Frontiers in Education, 10. https://doi.org/10.3389/feduc.2025.1594343
UNESCO. (2026). Digital Learning Week 2026: Provisional programme. Paris, 8–11 September 2026.