Ask. Never guess.Introducing Digital Twins →
GuidesResearch methods

How to run a competitive usability benchmark study comparing your product against rivals


Usability testing tells you whether your product works well for users. A competitive usability benchmark study tells you whether it works well compared to the alternatives your users are actually considering. That distinction matters because users don't evaluate products in isolation. They compare. If a competing product makes the same task faster, clearer, or less frustrating, that shapes their expectations—and their choices.

A competitive benchmark study gives you quantitative data you can track over time: task success rates, time on task, error rates, and satisfaction scores, measured consistently across your product and two to four competitors. When done well, it replaces opinion-driven debates about competitive positioning with evidence.

This guide walks through the full process—from defining what to measure through analyzing and communicating results.

Why run a competitive usability benchmark

There are several situations where a competitive benchmark delivers insight that other research methods cannot.

Identifying your relative strengths and weaknesses. Your product may perform well in absolute terms, but if a competitor handles the same workflow in half the time, users will notice. A benchmark surfaces these gaps.

Supporting strategic decisions. When leadership asks whether to invest in redesigning a core workflow, showing that your task completion rate is 20 percentage points below a competitor's is more persuasive than a list of qualitative usability findings.

Tracking progress over time. Running the same study periodically lets you see whether product changes are closing or widening the gap. This makes the benchmark a tool for measuring the impact of design decisions, not just identifying problems.

Informing roadmap prioritization. If your product outperforms competitors on onboarding but falls behind on advanced task flows, that tells you where effort is most needed—and where you already have an advantage worth protecting.

Step 1: Define the scope of the study

A competitive benchmark study is only as useful as the decisions it's designed to inform. Start by identifying the questions you need to answer.

Choose your competitors

Select two to four products that represent realistic alternatives for your users. These might be direct competitors, adjacent tools that overlap with part of your functionality, or an incumbent solution your product is trying to replace.

Avoid the temptation to include every competitor you can think of. Each additional product multiplies the time required per participant session and increases fatigue effects that degrade data quality. Focus on the competitors that matter most to your business and your users.

Select the tasks to benchmark

Choose three to five core tasks that are:

  • Common across all products being tested. The tasks need to be achievable in every product, even if the interaction patterns differ. If one competitor doesn't support a particular workflow, you can't compare performance on it.
  • Representative of real user goals. Don't pick tasks that are artificially simple or obscure. Choose the workflows your users perform most frequently or that have the highest impact on their satisfaction.
  • Specific enough to measure. "Explore the dashboard" is not a measurable task. "Find the total revenue for last quarter" is.

Write each task as a scenario that avoids using product-specific terminology. If your product calls a feature "Insights" and a competitor calls it "Reports," the task prompt should describe the goal in neutral terms: "Find a summary of last month's customer feedback trends."

Define your metrics

Standard usability benchmark metrics include:

  • Task success rate — The percentage of participants who complete the task successfully. Define success criteria in advance (e.g., reaching the correct screen, producing the correct output).
  • Time on task — How long it takes participants to complete the task, measured from the moment they begin to the moment they succeed or abandon.
  • Error rate — The number of errors or wrong paths taken during the task.
  • Post-task satisfaction — A standardized rating collected immediately after each task, typically using a 7-point scale or the Single Ease Question (SEQ).
  • Post-study satisfaction — A broader satisfaction measure collected after all tasks are complete, such as the System Usability Scale (SUS) or UMUX-Lite.

Decide which metrics matter most for your study goals. Task success rate and time on task are almost always included. Post-task satisfaction adds a subjective dimension that task performance alone does not capture.

Step 2: Design the study

Study design is where rigor matters most. Small design decisions—task order, product order, participant allocation—can introduce bias that undermines the validity of your comparisons.

Choose a study design

There are two primary approaches:

Within-subjects design — Every participant tests every product. This lets you make direct comparisons with fewer participants, but sessions are longer and learning effects can influence performance on later products.

Between-subjects design — Each participant tests only one product. This eliminates learning effects but requires more participants to achieve the same statistical power.

Most competitive benchmarks use a within-subjects design with counterbalancing. Counterbalancing means varying the order in which participants encounter the products so that no single product consistently benefits from being tested first (when participants are freshest) or last (when they've learned from prior products). A Latin square design is a common way to systematize this rotation.

Determine sample size

Competitive benchmarks are quantitative studies, so they need enough participants to detect meaningful differences. For a within-subjects design comparing three to four products on core metrics like task success rate, plan for 20 to 30 participants. If you're using a between-subjects design, you'll need that many per product.

If budget is limited, a within-subjects design with 15 to 20 participants can still produce useful directional data, but be cautious about treating small differences as meaningful.

Recruit representative participants

Participants should represent your actual user base—not power users, not novices who would never use your product category, but the people whose experience matters to your business decisions. Screen for relevant characteristics like role, experience level, and familiarity with the product category.

Critically, decide how to handle familiarity with specific products. If all participants already use your product daily but have never seen the competitors, familiarity effects will overwhelm the data. There are two common approaches:

  • Recruit participants who are familiar with the product category but not heavy users of any specific product in the study.
  • Recruit participants who use one of the products being tested and counterbalance which product each participant is already familiar with.

Document your recruitment criteria clearly—you'll need to replicate them in future rounds.

Prepare test environments

Set up each product in a realistic state. If you're benchmarking a project management tool, pre-populate it with sample projects, tasks, and team members so participants aren't working with an empty shell. The data and content should be consistent across products so that differences in performance reflect the product's design, not the test setup.

Use the same device, browser, and connection speed for all products. Control everything you can.

Step 3: Run the sessions

Pilot the study

Run two to three pilot sessions before collecting real data. Pilots reveal problems you can't anticipate from the study plan alone: tasks that are ambiguous, scenarios that don't translate across products, sessions that run too long, or metrics that are harder to capture than expected.

Adjust the study design based on pilot findings, but once you begin collecting real data, avoid changing tasks, metrics, or procedures mid-study.

Moderate consistently

If sessions are moderated, the facilitator must deliver task instructions identically for every product and every participant. Use a script. Avoid helping participants more on one product than another—even subtly. If a participant asks for help, follow the same protocol every time (e.g., offer one standard hint after 90 seconds, then mark the task as failed after another 60 seconds).

If you're running unmoderated sessions, use a tool that enforces consistent task delivery, timing, and data capture across all participants and products.

Record everything

Capture screen recordings, timestamps, success/failure judgments, and satisfaction ratings in a structured format as you go. Trying to reconstruct this data after the fact from session recordings is time-consuming and error-prone.

Step 4: Analyze the results

Compile the quantitative data

For each product and each task, calculate:

  • Task success rate (with confidence intervals if your sample size supports it)
  • Median time on task (median is more robust than mean for time data, which tends to be skewed)
  • Mean post-task satisfaction rating
  • Error counts or error rates

Then calculate overall metrics per product, including the post-study satisfaction score.

Look for meaningful differences

Resist the urge to declare a winner based on small differences. A task success rate of 82% vs. 78% is unlikely to be meaningful with 25 participants. Focus on patterns: consistent advantages or disadvantages across multiple tasks, large gaps on specific tasks, or notable differences in satisfaction despite similar performance.

If your sample size supports it, apply appropriate statistical tests—chi-square for success rates, Wilcoxon signed-rank for within-subjects time data—to determine which differences are statistically significant.

Examine the qualitative layer

Even in a quantitative study, you'll observe things worth noting. Where do participants get stuck on each product? What strategies do they use? What do they say when rating satisfaction? These observations won't be your primary findings, but they add context that helps stakeholders understand why one product outperformed another on a given task.

When working with this kind of multi-layered data—quantitative metrics alongside qualitative observations across multiple products and tasks—having a structured way to organize, tag, and synthesize findings is essential. This is where platforms like Dovetail can be particularly useful, allowing you to centralize session recordings, tag observations by product and task, and surface patterns across participants without losing the underlying evidence.

Step 5: Communicate findings

Structure the report around decisions

Don't organize the report by product or by metric. Organize it around the questions stakeholders need to answer:

  • Where does our product outperform competitors, and by how much?
  • Where does our product underperform, and what's driving the gap?
  • Which gaps matter most to users and to the business?
  • What has changed since the last benchmark?

Use comparison visuals

Side-by-side bar charts showing task success rates or time on task across products are immediately interpretable. Annotate them with confidence intervals or significance indicators where appropriate. Heatmaps that show performance by product and task at a glance can also be effective for executive audiences.

Be honest about limitations

State your sample size, recruitment criteria, and any caveats about the study design. If a particular comparison is directional rather than conclusive, say so. Credibility comes from transparency, not from overstating your findings.

Common pitfalls to avoid

Testing too many tasks. Long sessions produce fatigue effects that disproportionately penalize whichever product is tested last. Keep sessions under 60 minutes, ideally under 45.

Using your product's terminology in task scenarios. If the task prompt uses language specific to your product, your users will perform better on your product simply because the instructions are more familiar. Use neutral, goal-oriented language.

Ignoring learning effects. In within-subjects designs without counterbalancing, participants get better at the tasks as they go. The second or third product tested benefits from this practice effect, producing artificially inflated success rates and faster times.

Treating the benchmark as a one-time event. A single benchmark is a snapshot. Its value multiplies when you repeat it with consistent methodology, allowing you to track trends and measure the impact of product changes over time.

Focusing only on where you lose. The gaps where competitors outperform you are important, but so are the areas where you lead. These strengths are worth protecting, promoting, and building on.

Building a benchmark practice over time

The first competitive benchmark study is the hardest—it requires the most setup, the most methodological decisions, and the most organizational buy-in. Subsequent rounds become faster because the study design, recruitment criteria, and analysis framework are already established.

Document your methodology thoroughly so future rounds are comparable. Store your raw data, recordings, and analysis in a system that makes it easy to revisit—both for methodological reference and for longitudinal comparison. Teams using Dovetail often centralize their benchmark data alongside other research, making it easier to connect competitive findings with broader user insights and track how metrics evolve across study rounds.

Over time, a repeated competitive benchmark becomes one of the most powerful tools a product team has: an objective, evidence-based view of where you stand relative to the products your users are choosing between, updated regularly enough to inform strategy and measure progress.

FAQs

How many competitors should you include in a competitive usability benchmark study?

Most studies compare your product against two to four direct competitors. Including fewer than two limits the usefulness of the comparison, while including more than four significantly increases session length and participant fatigue, which degrades data quality. Choose competitors that your users are most likely to evaluate as alternatives—not every player in the market, but the ones that shape purchasing decisions and user expectations.

What is the difference between a competitive usability benchmark and a standard usability test?

A standard usability test focuses on identifying specific usability problems within a single product. A competitive usability benchmark is designed to produce quantitative metrics—like task success rate, time on task, and satisfaction scores—that can be compared across multiple products and tracked over time. The study design prioritizes measurement consistency and statistical comparability rather than deep qualitative exploration of individual issues.

How often should you repeat a competitive usability benchmark study?

The cadence depends on how quickly your product and the competitive landscape are changing. Many teams run competitive benchmarks annually or twice a year. If you or a key competitor ships a major redesign, an additional round may be warranted. Repeating the study with consistent methodology is what makes the data valuable over time—it lets you track whether the gap between your product and competitors is widening or closing.

Editor's picks↘

Latest articles↘

Turn customer feedback into product innovation