No Validated Scale for This: What to Consider for Qualitative Rigor In UX Research Instead
Credibility, transferability, dependability, and confirmability aren't a validated scale, they're decisions a UX researcher makes before the study begins.
There’s No Validated Scale, So What Do You Check Against?
Thematic analysis has usually been the backbone of my qualitative UX research work, and I’ve noticed that I’ve been considering criteria such as credibility, transferability, dependability, and confirmability to uphold qualitative research rigor. Unlike quantitative UX research, which sometimes uses a validated scale, benchmarks or asks researchers to compare results against set numbers like p-values, qualitative research does not have its own “validated scale.”
That gap is what I want to explore in this piece; not a substitute scale, but a set of decisions I make as a UX research practitioner for my qualitative work. I like to treat them less as a checklist completed after qual data analysis, but more like decisions made during study design and data analysis.
These Four Words Were Introduced to Resist a Checklist
From literature that I’ve read, the terms have a fairly specific origin. In 1985, Yvonna Lincoln and Egon Guba proposed credibility, transferability, dependability, and confirmability as qualitative research’s answer to the quantitative world’s validity, generalizability, reliability, and objectivity. They see it not as direct translations, but parallel commitments built for a different kind of inquiry. In qualitative work, reality is treated as constructed through the researcher’s interpretation, rather than simply observed. I think it helps to have a working sense of what each one is actually asking:
Credibility: whether the findings stay true to the people being studied; do the interpretations reflect the participants’ reality?
Transferability: whether the findings might apply to another context, a judgment the reader makes rather than the researcher asserts.
Dependability: whether the research process is stable, traceable, and logical enough that someone could follow how it unfolded.
Confirmability: whether the findings trace back to the data rather than to the researcher’s own preconceptions.
Guba had laid some of this groundwork a few years earlier, arguing that naturalistic research needed criteria of its own rather than ones borrowed from a positivist tradition that didn’t quite fit it. Academia has listed these four criteria out as a set, but that’s a different thing from them being created to function as a checklist in practical UX research work. My own reading is that they work less like four boxes to tick and more like four questions meant to stay open throughout a study.
What Designing for Qualitative Rigor Actually Looks Like
Let’s think about an example together: a UX research team runs a qualitative study on how people decide whether to trust an AI writing assistant’s suggestions. Here’s what designing for each criterion during study design, rather than demonstrating it after the data has been analyzed, might actually look like:
Credibility: the research team decides to pair participant think-aloud sessions with a review of their actual writing edit histories, to check whether what they say about trust matches what they do with it. So, rather than asking participants whether an interpretation matches their experience after they do a task, they observe the participants doing the task and ask them to review their editing procedures. This creates credibility throughout the process of the study.
Transferability: the research team decides early which contextual details are worth tracking, like participants’ prior AI experience or the domain they’re writing in. This way, the task descriptions exist because the study was built to notice it, creating transferability.
Dependability: the research team notes down essential analytic decisions as they’re made (e.g., why a theme was split, etc.), creating dependability of the data analysis process.
Confirmability: the research team builds reflexive checkpoints into the data analysis itself, flagging where the researcher’s own assumptions about “trust” are doing interpretive work. This allows the results to be traced back to the data and create confirmability.
Two Critiques Pointing the Same Way
Aamir Rashid and Rizwana Rasheed made a related argument in a paper published this year. As I read it, they’re proposing that trustworthiness works best not as four separate boxes to check but as something closer to an accountability system. In a sense, they emphasize the importance of keeping the reasoning between data and claims visible at every stage, not just the reporting stage. I think it’s a useful frame, albeit being a new publication.
Another researcher, Margarete Sandelowski, had warned three decades earlier that rigor can slide into “rigor mortis” (e.g., techniques performed for the appearance of trustworthiness rather than practiced for the sake of it). So I find that though thirty years apart, both seem to point at the same worry: the presence of a technique doesn’t, on its own, tell a reader whether the reasoning it’s meant to support holds up. The important thing is to make active decisions to do so.
Qualitative Rigor as an Ongoing Practice, Not a Checklist
For me as UX researcher, the method I most often use in my qualitative work is reflexive thematic analysis by Virginia Braun and Victoria Clarke. This method treats the researcher’s own interpretive lens as part of the analytic instrument, which asks me to keep examining my own assumptions while I code the data, and not just when I write the findings.
I think that stance sits close to confirmability by design: if I’m already noting down why I read a passage as being about “trust” rather than “convenience,” I’m doing confirmability work in the moment, not reconstructing it later. And the same note quietly builds the traceable record dependability asks for. So in my own work, the four criteria have started to feel less like a written section I owe the reader at the end, and more like questions the method keeps putting in front of me while the study is still moving.
Later, Sarah Tracy had proposed an alternative framework to expand the set to eight criteria. I feel like if four criteria can settle into a checklist, eight or twelve can too. But to me, the number of criteria was ever really the issue. Instead, what matters to me is treating qualitative rigor as a set of criteria to aid qualitative decision-making alongside other things for consideration like a study’s ethical grounding. So it’s not really about checking things off of a rigor checklist, but rather to assist the researcher to make the best decisions for the study and data analysis.
Designed For, or Checked For?
So where does that leave a UX researcher without a validated scale to check their qualitative work against? I don’t think the answer is to invent one. Credibility, transferability, dependability, and confirmability were never meant to work as a rigor score I calculate once a study wraps up. They are, more importantly, closer to questions I try to keep asking while the study is still moving. No number of criteria turns qualitative rigor into a checklist, because rigor was never a score to begin with. It’s a set of decisions, made and remade throughout study design and data analysis, and shows the real value of upholding rigor in qualitative UX research work.

Notes: This piece draws on a small set of academic sources on qualitative methodology, read from my perspective as a UX researcher rather than a trained methodologist. It isn’t meant as a definitive account of how these four criteria should be used, and I’d welcome other reads on this.
Sources Referenced
Guba, E. G. — Criteria for Assessing the Trustworthiness of Naturalistic Inquiries (1981)
Lincoln, Y. S. & Guba, E. G. — Naturalistic Inquiry (1985)
Sandelowski, M. — Rigor or Rigor Mortis: The Problem of Rigor in Qualitative Research Revisited (1993)
Tracy, S. J. — Qualitative Quality: Eight “Big-Tent” Criteria for Excellent Qualitative Research (2010)
Rashid, A. & Rasheed, R. — Trustworthiness as Inferential Accountability in Qualitative Research (2026)
Braun, V. & Clarke, V. — Reflecting on Reflexive Thematic Analysis (2019)




留言