How to run research on trust calibration when users interact with AI confidence scores and uncertainty indicators
As AI-generated outputs become more common in products—from medical diagnoses to content recommendations to financial forecasting—designers increasingly surface confidence scores and uncertainty indicators alongside those outputs. The intent is straightforward: help users make better decisions by showing them how sure the system is.
But displaying a confidence score is not the same as producing well-calibrated trust. Users may ignore uncertainty indicators, misinterpret percentages, or anchor on confidence numbers in ways that lead to worse decisions than if no score were shown at all. Understanding how people actually respond to these signals requires deliberate, well-structured research.
This article walks through how to design and run research on trust calibration when users interact with AI confidence scores and uncertainty indicators—from defining the right questions through study design, measurement, analysis, and applying findings to product decisions.
Why trust calibration matters
Trust calibration is the degree to which a user's trust in an AI system matches the system's actual reliability. When calibration is good, users rely on the system when it is accurate and exercise independent judgment when it is not. When calibration is poor, two failure modes emerge:
Over-trust (automation bias)—Users accept AI outputs uncritically, even when confidence is low or the system is wrong. This is especially dangerous in high-stakes domains like healthcare and finance.
Under-trust (algorithm aversion)—Users dismiss AI outputs even when the system is accurate and confident. This wastes the value of the AI system entirely.
Both failure modes lead to worse outcomes than appropriate reliance would. The goal of trust calibration research is not to increase trust in general—it is to help users develop the right amount of trust at the right moments.
Defining your research questions
Before designing a study, get specific about what you are trying to learn. Broad questions like "Do users trust AI?" are too vague to produce actionable findings. Instead, frame questions around the relationship between the confidence signal and user behavior.
Questions about comprehension
- Do users understand what a confidence score represents?
- How do users interpret different formats of uncertainty (percentages, verbal labels, visual indicators)?
- Do users distinguish between the system's confidence and its accuracy?
Questions about behavior
- How does the presence of a confidence score change users' decision-making compared to no score?
- At what confidence threshold do users shift from accepting to overriding the AI?
- Do users adjust their review effort based on confidence level (e.g., spending more time verifying low-confidence outputs)?
Questions about calibration
- Does displaying confidence scores improve the alignment between user trust and system accuracy?
- Are there formats or framings of uncertainty that produce better calibration than others?
- How do individual differences—domain expertise, numeracy, prior AI experience—moderate calibration?
Starting with specific questions like these will shape every subsequent decision about study design, task construction, and measurement.
Choosing a research approach
Trust calibration research benefits from mixed methods. Behavioral data tells you what users do in response to confidence signals. Qualitative data tells you why—what mental models they hold, what the numbers mean to them, and how they reason about uncertainty.
Controlled experiments
Controlled experiments are the strongest method for isolating the effect of confidence score design on user behavior. In a typical design, participants complete a series of tasks with AI assistance, and you vary the confidence display between conditions.
Common experimental manipulations include:
- Presence vs. absence of confidence scores—Does showing any confidence information change behavior?
- Format of the confidence display—Percentages vs. verbal labels ("high confidence," "uncertain") vs. visual indicators (color gradients, progress bars, blurred outputs)
- Calibration of the AI itself—Presenting outputs where the AI's stated confidence matches its actual accuracy vs. cases where it does not
Between-subjects designs (each participant sees one condition) avoid carryover effects but require larger sample sizes. Within-subjects designs (each participant sees multiple conditions) offer more statistical power but risk order effects and demand careful counterbalancing.
Think-aloud protocols
Concurrent or retrospective think-aloud protocols reveal how users reason about confidence scores in real time. You can observe whether they notice the indicator at all, how they interpret it, and what decision rules they apply.
Think-aloud is especially valuable in early research when you do not yet know which aspects of the confidence display matter. It surfaces unexpected mental models—for example, users who interpret "85% confidence" as "the system is 85% sure it is right" vs. "85 out of 100 times, outputs like this are correct" may behave very differently despite seeing the same number.
Diary studies and longitudinal observation
Trust calibration is not static. Users' responses to confidence scores change over time as they accumulate experience with a system's reliability. A user who initially checks every AI output may stop checking after a string of correct predictions—even when confidence is low.
Diary studies or longitudinal observation capture this evolution. Ask participants to log their interactions with the AI system over days or weeks, noting when they trusted the output, when they overrode it, and why. This approach is harder to control but reveals patterns that short lab studies miss.
Surveys and scale-based measurement
Validated instruments can supplement behavioral data. The Trust in Automation questionnaire, the System Trust Scale, and the Checklist for Trust between People and Automation all provide standardized measures. These are most useful as pre/post measures in experimental studies or as covariates that help explain individual differences.
Be aware that self-reported trust and observed trust often diverge. A participant may report moderate trust in a post-task survey but accept every AI suggestion during the task itself. Use surveys as one input among several, not as your primary trust measure.
Designing realistic tasks
The ecological validity of your tasks determines whether your findings will transfer to real product contexts. Trust calibration research is particularly sensitive to task design because users' reliance on AI depends heavily on the stakes, their own expertise, and the effort required to verify the output independently.
Match the domain
If your product serves radiologists, use medical imaging tasks. If it serves financial analysts, use forecasting tasks. Generic tasks (e.g., "guess whether this AI-classified image is a dog or a cat") may reveal basic perceptual patterns but will not tell you how your users behave in their actual workflows.
Vary AI accuracy deliberately
To measure calibration, you need trials where the AI is right and trials where it is wrong, across a range of stated confidence levels. This means constructing a stimulus set where you control both the AI output and its confidence score.
A common structure:
- High confidence, correct output (the system should be trusted here)
- High confidence, incorrect output (tests whether users over-rely on confidence)
- Low confidence, correct output (tests whether users dismiss good outputs)
- Low confidence, incorrect output (the system should not be trusted here)
Balance these trial types and randomize their order. The ratio of correct to incorrect outputs matters—if users notice that the AI is almost always right, they will rationally increase reliance regardless of confidence scores.
Provide a baseline
Include a condition where participants perform the same tasks without AI assistance. This baseline establishes their independent performance and lets you calculate whether the AI (with its confidence display) actually improves decision quality.
Measuring trust calibration
Operationalizing trust calibration requires connecting two streams of data: the user's trust-related behavior and the AI's actual accuracy on each trial.
Agreement rate by confidence level
For each trial, record whether the user accepted or rejected the AI suggestion. Then plot agreement rate against the AI's stated confidence level. A well-calibrated user should show higher agreement at higher confidence levels and lower agreement at lower ones. A flat line—same agreement regardless of confidence—indicates the user is ignoring the confidence signal.
Appropriate reliance metrics
Agreement rate alone does not capture calibration quality because a user could agree with every suggestion and still show high agreement at high confidence levels. More precise metrics include:
- Relative positive agreement—The proportion of correct AI outputs that the user accepted
- Relative negative agreement—The proportion of incorrect AI outputs that the user rejected
- RAIR (Relative AI Reliance) and RSR (Relative Self-Reliance)—Metrics that separate the user's tendency to switch toward the AI from their tendency to maintain their own answer
These metrics, used together, distinguish between users who appropriately rely on the AI and users who blindly accept or reject it.
Behavioral proxies
Beyond explicit accept/reject decisions, behavioral data provides a richer picture:
- Dwell time—Do users spend more time reviewing low-confidence outputs?
- Information seeking—Do users consult additional sources when confidence is low?
- Interaction with the indicator itself—Do users hover over, click, or expand confidence displays?
Tools like Dovetail can help you organize and analyze this behavioral data alongside qualitative insights from think-aloud sessions and interviews, giving you a unified view of how users interact with confidence signals across different tasks and conditions.
Analyzing and synthesizing findings
Trust calibration studies produce data at multiple levels: trial-level behavioral logs, participant-level survey scores, and session-level qualitative transcripts. Bringing these together coherently is one of the harder parts of the research.
Quantitative analysis
For experimental studies, mixed-effects models are well suited to trust calibration data because they account for the nested structure—multiple trials within each participant—and handle individual differences as random effects. The key predictors are usually the stated confidence level, the AI's actual accuracy on each trial, the display format (if you varied it), and participant-level covariates like numeracy or domain expertise.
Qualitative analysis
Thematic analysis of think-aloud transcripts and interview data reveals the reasoning behind behavioral patterns. Common themes in trust calibration research include:
- Threshold-based rules—"I trust it if it's above 80%"
- Confirmation seeking—"The number just confirms what I already thought"
- Confusion about the metric—"I'm not sure if 70% means it's 70% likely to be right or that it analyzed 70% of the data"
- Anchoring—"Once I see a high number, it's hard to disagree even when something looks off"
When you are working with large volumes of qualitative data from trust calibration sessions, a platform like Dovetail can help you tag and cluster these reasoning patterns across participants, making it easier to identify which mental models are most prevalent and how they relate to behavioral outcomes.
Connecting the two
The most actionable findings emerge when you can link qualitative reasoning to quantitative patterns. For example, participants who apply threshold-based rules may show a sharp inflection point in their agreement rate at a specific confidence level, while participants who use confirmation-seeking strategies may show agreement rates driven more by their prior beliefs than by the stated confidence.
Applying findings to product decisions
Trust calibration research findings translate into concrete design decisions about how—and whether—to surface AI confidence in your product.
Format and framing
Your research may reveal that certain formats produce better calibration. Common findings from the literature and applied research include:
- Verbal labels ("high," "medium," "low") tend to produce coarser but more consistent interpretation than precise percentages
- Visual indicators like color gradients or blur effects can convey uncertainty without requiring numerical interpretation
- Providing context ("the system is correct X% of the time at this confidence level") improves calibration but increases cognitive load
When not to show confidence
In some cases, research reveals that showing confidence scores makes decisions worse—particularly when the scores are poorly calibrated or when users lack the expertise to integrate them meaningfully. Your research may support removing the confidence display entirely or restricting it to expert users.
Adaptive confidence displays
Some products benefit from adjusting the prominence of uncertainty indicators based on context. For example, suppressing confidence on routine, high-accuracy tasks (where the indicator adds noise) and surfacing it prominently on edge cases (where user judgment matters most).
Ethical considerations
Trust calibration research involves a degree of deception when you present AI outputs with controlled accuracy and confidence levels. This raises informed consent questions. Participants should know that the AI outputs they encounter during the study may not reflect a real system's performance, and debriefing should clarify which outputs were manipulated.
Additionally, be thoughtful about the populations you recruit. Trust calibration findings from a general population may not apply to the domain experts who will actually use your product. Conversely, findings from experts may not apply to novice users. Be explicit in your reporting about the limits of generalizability.
Building a trust calibration research practice
Trust calibration is not a one-time study. As your AI system evolves—as its accuracy changes, as confidence calibration improves, as new use cases emerge—the dynamics of user trust shift. Building an ongoing research practice around trust calibration means:
- Establishing baseline metrics that you can track over time (agreement rates, override rates, calibration curves)
- Embedding lightweight trust checks into usability testing and beta programs
- Revisiting your confidence display design whenever the underlying model changes significantly
Centralizing your trust calibration research—tasks, protocols, behavioral data, qualitative themes—in a research repository makes it possible to compare findings across studies and track how trust patterns evolve. Dovetail provides the infrastructure for this kind of longitudinal research synthesis, helping teams maintain continuity even as researchers, products, and AI systems change.
Trust calibration research sits at the intersection of cognitive psychology, human factors, and product design. It demands rigor in study design and measurement, but the payoff is significant: products where users rely on AI appropriately, make better decisions, and maintain agency over outcomes that matter.
