Choosing User-Centered Metrics for AI Products: A Research Design Problem
A look at how UX research can help choose and interpret metrics intended to reflect users’ experiences of AI features, and why the meaning of a user-centered metric depends on the research design.
In this piece, I want to take a step back and ask an important question in UX research for AI products: what makes a metric meaningfully user-centered for AI products, rather than simply a number collected from users? And what does the research design need to make that number interpretable?
What makes a metric actually user-centered?
UX researchers have many ways to study AI products: usability testing, user evals, diary studies, surveys, and others. These methods can produce ratings, task-success measures, and behavioral patterns. However, a number does not become “user-centered” simply because it comes from a study involving real users.
For instance, a satisfaction rating may come directly from a participant, but if the construct is vague or the task is unrealistic, its meaning may be limited. At the same time, a behavioral measure such as whether someone successfully completes an important task can be highly user-centered even though it is not a subjective rating. So I would not define user-centered metrics as subjective metrics alone.
A metric becomes more meaningfully user-centered when it represents something consequential to the user’s experience or outcome, and when the study gives the team a solid reason to interpret it that way.
That leads to the question about research design:
What would need to be true about how this number was collected for it to support the interpretation we want to make?
Different kinds of metrics answer different questions
Drawing loosely from select sources across fields such as HCI, human factors, and human-AI evaluation research, I find it useful for this discussion to think about three kinds of evidence/measures: system-performance, behavioral, and perceptual. These are not intended as an exhaustive taxonomy, but as a practical way to distinguish what the system did, what the user did, and how the user interpreted the experience.
System-performance measures describes what the AI produced relative to an external criterion. This might include factual accuracy, error rates, or benchmark performance. These measures are often part of dataset-based or automated evals done by product managers (PM) or data scientists (DS), although responsibilities vary by organization.
Behavioral metrics describes what users actually did. For instance, did they complete the task? Did they accept the AI’s recommendation? Did they abandon the interaction? Or did they recover successfully from a failure? These are all evidence for users’ behaviors.
Perceptual measures captures how users interpreted the experience. This might include perceived helpfulness, satisfaction, confidence, mental demand, or even trustworthiness.
I think each type answers a different question. For example, system-performance evidence might tell us whether an AI’s answer was correct. Behavioral evidence might tell us whether the user abandoned the chat after the AI provided its response. And perceptual evidence might tell us whether the user believed it was correct or useful. These 3 types of metrics could complement each other in painting the broader picture of the user’s experience. And I believe that in many AI studies, the more interesting finding may even come from how these forms of evidence relate to one another.
Sometimes, the gap between objective AI response quality and subjective user perception can be the finding itself
Let’s look at an example of AI’s response “accuracy.” Let’s imagine that an AI produces a response that can be evaluated against some sort of ground truth (e.g., data from a PM or DS’s eval where they sought the accuracy rate of the AI’s responses). In a user session, UX researchers could also ask participants whether they believe that response is accurate, and why.
Some hypothetical user feedback I can think of are: if users judge correct information from the AI as “unreliable,” perhaps the problem may involve the response’s presentation or sourcing, rather than the model performance. On the other hand, if the system confidently presents incorrect information and misleads users to accept it, perhaps the system may be communicating with more “certainty” than its model performance warrants.
I feel like this is in parallel with Sharma et al.’s research on sycophancy, which illustrates one aspect of this broader issue: AI’s outputs that are aligned with what users “appear to prefer” are not necessarily the same as outputs that are factually correct, due to the tendency for sychophancy in AI’s responses.
So from a UX researcher’s perspective, the question becomes: how well calibrated are users’ perceptions of quality to the system’s actual quality? This gap may itself point to opportunities for product improvement. Teams might improve the underlying quality of an AI response by tuning for greater information accuracy and an appropriate level of detail, while also considering how the response communicates that quality to users. Clearer explanations, sufficient supporting detail, or other credibility cues may help users more accurately assess the reliability of a response. In this sense, the goal is not simply to increase perceived accuracy, but to bring users’ perceptions of accuracy into better alignment with the actual quality of the information they receive from the AI.
“Trust” as an example to show why construct definition matters
From a few sources that I have read, I find that trust is another useful example because it sounds intuitive but can represent several different things to users. When someone says, “I trust this AI,” do they mean that the answer feels credible? That they believe the model is competent? That they would follow its recommendation? That they would use it again? And even if they do rely on the AI, was that reliance appropriate?
So before selecting a metric such as “trust,” it may be worth asking what construct the team is really interested in: perceived trustworthiness, user’s actual reliance behavior, user’s confidence in the AI, or something else.
Users’ expectations may change even when metric scores do not
Another challenge I find appears when teams track perceptual measures such as helpfulness or perceived trustworthiness over time. One observation from my experience is that a stable score over a long period of time does not necessarily mean the underlying product experience has remained unchanged.
The biggest reason for this is that as users gain more experience with AI products, their expectations may change. For AI features that once seemed impressive can become baseline expectations in a couple weeks time. A knowledgeable-sounding answer from the AI, for example, may once have been enough for someone to rate an AI system highly in trustworthiness. But later, that same user may expect even more from the AI, such as citations or stronger explanations before assigning the same rating.
I feel like this creates an interpretation problem for what the metric actually means; the AI quality might improve while users’ standards for the features rise at the same time. I can see that one pushes the score upward while the other pushes it downward.
This is reflected in Shankar et al.’s research, where they described a related phenomenon in LLM-assisted evaluation: people developing automated AI graders may revise their own criteria as they encounter more examples of model output. The context in that research paper is different, but I feel like the message is relevant here: the standard being applied to an AI system may itself change through exposure to upgraded models and systems. So if user’s ratings remain relatively stable over time, while their feedback and reasons for their ratings changes, it may be useful signal to investigate whether the construct, expectations, or product context has shifted while the metric(s) stayed the same.
One metric may not be enough
There is another reason to be cautious about searching for the one “right” user-centered metric: AI experiences can look good on one dimension and poor on another. For example, a response can be “preferred” but less accurate, faster but harder to understand. I think these tensions are not necessarily measurement failures, and instead could be finding(s) in and of themselves.
For that reason, some research questions may be better served by a small set of complementary measures. UX researchers can refer to system-performance evidence from product management and data science functions, and design UX research studies that capture behavioral data, perceptual ratings and qualitative user explanations to compliment it. The goal is not to collect as many metrics as possible, but to choose evidence that helps explain user behaviors and expectations needed for product design and improvements.
When several forms of metrics and evidence point in the same direction, the interpretation may become stronger. And even when they disagree, the disagreement can tell the team where to investigate next.
A practical way to select user-centered metrics for AI products
I think of these five questions when selecting user-centered metrics, and it starts with considering the research design around it. This is not intended to be a fixed framework, but rather questions to help guide metric selection.
1. What decision are we trying to make?
One common way of deciding on a metric is starting with the product or research decision, rather than the measurement. For instance, consider these: Are we deciding between two interaction designs? Comparing model experiences? Or determining whether a feature change improved an important outcome?
Starting with the product or research decision helps select metrics that actually help answer their product questions.
2. What user outcome or construct would provide evidence for that decision?
Take this as an example: if the team says they want to measure “helpfulness,” what does helpfulness mean in this product? Maybe these are some ideas to consider: Did the AI help someone complete a task? Did it help make a decision? Or did it help reduce user’s effort?
Thinking about what construct you want to represent with the metric is a way to make sure that you are measuring the right thing in your study when you choose the metric and design the description for it.
3. What evidence would indicate that construct?
The next thing to consider is whether the construct is best represented through system-performance evidence, user behavior, perceptual ratings, user’s qualitative explanation, or some combination of any of them.
I think this is also an opportunity for cross-functional collaboration: Product managers, engineers, or data scientists may already have evaluation data about model behavior. UX research may then be well positioned to investigate how that performance translates into user behavior or perception.
4. What research design makes the evidence interpretable?
The natural next step here is to design the research. Think about the context(s) or task(s) of your study, and what the user’s experience(s) can provide with the measurements. Should users bring their own prompts? Should everyone complete the same scenario? Does the relevant experience happen in a single interaction, or should it be a multi-turn AI interaction instead?
It is also worth asking whether the metric is sensitive enough to distinguish the experiences the team cares about. A measure can represent the right construct and still provide little useful differentiation if nearly every user gives the same rating or response.
5. What claim can the study reasonably support?
Finally, think about the intended claim before collecting the data. Some questions to consider include: Who participated? What tasks did they complete? Which model and interface did they encounter? Was the study designed for exploration or competitive analysis? These questions can help finalize the study’s metric selection and research design.
What gives a user-centered metric its meaning
I feel like user-centered measurement for AI products will likely continue evolving because both the products and users’ familiarity with them are evolving quickly. And a user-centered metric’s “meaning” to the product team does not only rely on a user’s rating of their perceived experience with the AI, be it metrics like helpfulness, trust or something else. Instead, I think its meaning comes from what the product team intends to understand, how the construct was defined, what participants experienced in the UXR study, what evidence was collected, and what interpretation the research design can reasonably support.
So, going back to what we were talking about at the beginning, I think perhaps the more useful question is not simply: “Which metric should we use?” Rather, it would be what we started with: “What would need to be true about how this number was collected for it to mean what we want to claim?”
When teams can answer that clearly, metric selection becomes less about choosing a score and more about designing the research to collect evidence needed to make a product decision.

Referenced Sources
Schemmer, M., et al. (2022). Should I Follow AI-Based Advice? Measuring Appropriate Reliance in Human-AI Decision-Making.
Shankar, S., et al. (2024). Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences.
Sharma, M., et al. (2024). Towards Understanding Sycophancy in Language Models.
Lee, M., et al. (2023). Evaluating Human-Language Model Interaction.
Marwad, J., et al. (2024). Towards a Common Metrics and Evaluation Framework for Assessment of Older Adults and Caregivers Interacting with Artificial Intelligence.
Scharowski, N., et al.(2022). Trust and Reliance in XAI—Distinguishing Between Attitudinal and Behavioral Measures.




留言