Ask. Never guess.Introducing Digital Twins →
GuidesResearch methods

How to adapt research synthesis methods when working with machine-translated interview transcripts


International user research has become routine for many product teams. Interviewing participants in their own language produces richer, more honest data than asking them to communicate in a second language. But once those interviews are recorded and transcribed, the research team is often left with transcripts in languages they do not read or speak.

Machine translation has become the default bridge. Tools like Google Translate, DeepL, and Microsoft Translator can convert a transcript from Mandarin, Portuguese, or Arabic into English in seconds. The output is usually good enough to understand what happened in a conversation. But "good enough to understand" and "good enough to do rigorous qualitative analysis on" are not the same thing.

This article covers the specific risks that machine-translated transcripts introduce into research synthesis—and practical ways to adjust your methods so your findings remain trustworthy.

What machine translation does well and where it falls short

Modern neural machine translation handles factual, declarative content well. If a participant describes a sequence of steps they followed, names a tool they used, or states a preference between two options, the translation is usually accurate enough to work with directly.

Problems emerge in several predictable areas:

Hedging and certainty. Languages express degrees of confidence differently. Japanese speakers, for example, frequently use grammatical structures that soften assertions. A machine translation might render these as flat declarative statements, making a tentative opinion sound like a strong conviction.

Emotion and affect. Tone markers, diminutives, intensifiers, and culturally inflected expressions of frustration or satisfaction are difficult for translation engines. A participant who is expressing mild annoyance in Spanish may come across as neutral or even blunt in English.

Idioms and metaphors. Machine translation tends to produce literal renderings of figurative language. Sometimes the result is nonsensical, which at least signals that something was lost. More dangerously, the literal rendering can sound plausible but carry a different meaning than the speaker intended.

Pronouns and referents. Languages like Mandarin, Turkish, and Finnish handle gendered pronouns, formality levels, and implied subjects differently than English. Machine translation may introduce ambiguity about who is being discussed or flatten important social context.

Domain-specific vocabulary. If participants use industry jargon, regional slang, or product-specific shorthand, translation accuracy drops. This is particularly relevant in B2B research where participants may describe internal workflows using terminology that does not have a clean English equivalent.

None of these issues makes machine-translated transcripts unusable. But they do mean that a researcher who treats translated text with the same confidence as a native-language transcript will introduce systematic distortions into their analysis.

Preparing transcripts before synthesis

The work of adapting your synthesis method starts before you open your coding tool. A few preparation steps can significantly reduce downstream problems.

Assess translation quality on a sample

Before committing to a full analysis, take a representative excerpt—ideally two or three minutes of conversation—and have a bilingual colleague or contracted translator review the machine translation against the original. You are not looking for a polished human translation. You are looking for patterns of error: Are hedges being dropped? Are key terms being mistranslated consistently? Is the overall meaning preserved even if the phrasing is awkward?

This step takes an hour or two and gives you a calibrated sense of how much to trust the translated text.

Run parallel translations

For critical projects, running the same transcript through two different translation engines and comparing the output can surface passages where meaning is ambiguous. If Google Translate and DeepL produce substantially different renderings of the same passage, that passage deserves closer attention.

Preserve the original language text

Always keep the source-language transcript alongside the translation. This allows you to return to the original when a passage in English seems unclear, and it makes it possible for bilingual reviewers to spot-check your emerging codes and themes against what participants actually said.

Add translator's notes

If you have access to a bilingual team member or reviewer, ask them to annotate the translated transcript rather than produce a full human translation. Annotations might flag idioms that were translated literally, mark passages where tone was likely different from what the English conveys, or note culturally specific references that need context. This is faster and cheaper than full human translation and addresses most of the highest-risk distortions.

Adapting your coding approach

Qualitative coding—whether you are doing thematic analysis, grounded theory, or a more structured framework analysis—requires adjustment when the source text is machine-translated.

Code meaning, not language

In native-language analysis, exact word choice often matters. Researchers pay attention to whether a participant says "frustrated" versus "confused," or "I had to" versus "I chose to." With machine-translated text, these word-level distinctions are unreliable. The translation engine chose those words, not the participant.

Shift your coding to the level of meaning units rather than specific phrasing. Ask what the participant was communicating, not how the translation worded it. This is a meaningful methodological adjustment—it means your codes will tend to be slightly more abstract, and you will lose some granularity. That tradeoff is worth making because the alternative is false precision.

Use uncertainty markers

Introduce a systematic way to flag codes that are based on ambiguous passages. A simple convention works: tag uncertain codes with a marker like [low confidence] or [translation uncertain]. This does two things. It prevents you from forgetting which codes rest on shaky ground, and it gives you a clear list of items to prioritize for bilingual review.

In platforms like Dovetail, where you can tag and annotate highlights across multiple transcripts, maintaining a dedicated tag for translation-uncertain passages makes it straightforward to filter and review these moments separately from your main analysis.

Be cautious with sentiment and intensity

If your research questions involve understanding how strongly participants feel about something, or distinguishing between mild preferences and strong convictions, machine-translated text is particularly unreliable. Consider treating sentiment-related codes as provisional until validated by a bilingual reviewer.

For projects where emotional response is central—such as evaluating a sensitive healthcare experience or understanding reactions to a pricing change—machine translation alone may not be sufficient. In these cases, budgeting for at least partial human translation of emotionally significant passages is a sound investment.

Watch for false patterns

Machine translation engines are deterministic: the same source phrase will usually produce the same English output. This means that if a translation engine consistently renders a particular expression in a misleading way, the error will repeat across every participant who used that expression. During synthesis, this can look like a strong theme when it is actually a translation artifact.

If you notice multiple participants appearing to use the same unusual phrasing, check whether they are all speakers of the same source language. If so, the pattern may reflect the translation engine's behavior rather than a genuine commonality in participant experience.

Adjusting your synthesis and reporting

The downstream effects of working with translated data extend beyond coding into how you synthesize findings and communicate them to stakeholders.

Triangulate more aggressively

When your primary data source has known reliability limitations, triangulation becomes more important. Look for corroboration from other data types—behavioral analytics, survey responses, observational notes from the session—before treating a translated finding as established.

If a theme appears only in machine-translated transcripts and cannot be supported by any other evidence, flag it as preliminary rather than confirmed.

Quote with transparency

Direct quotations from participants carry significant rhetorical weight in research reports. When those quotations are machine-translated, presenting them without context is misleading. Readers will interpret them as the participant's own words when they are actually the translation engine's approximation.

Adopt a clear convention in your reports. One effective approach is to present translated quotes in italics with a note: "It was very difficult to find the right setting" (translated from Mandarin via DeepL). This small adjustment preserves the participant's voice while being honest about the mediation involved.

Report your translation methodology

Your methods section—whether it is a formal section in a research report or a few lines in a Slack summary—should state how transcripts were translated, which tools were used, what validation steps were taken, and what limitations the team acknowledges. This is not bureaucratic overhead. It is how other people assess whether they should trust your findings enough to act on them.

Involve bilingual stakeholders in review

If your organization has team members who speak the source languages, inviting them to review key findings before they are finalized is one of the highest-value validation steps available. They do not need to re-analyze the data. Simply asking them to read your top-level themes and check whether anything seems inconsistent with their understanding of the participant culture or language can catch significant errors.

When machine translation is not enough

There are research contexts where machine translation, even with all the adaptations described above, is insufficient:

  • Legal or compliance research where exact wording has regulatory implications
  • Research involving vulnerable populations where misrepresenting participant statements raises ethical concerns
  • Linguistic or discourse analysis where the object of study is language itself
  • Low-resource languages where translation engine accuracy is substantially lower

In these situations, professional human translation—ideally by translators with qualitative research experience—is the appropriate choice. The cost is higher, but the risk of acting on distorted data is higher still.

Building a sustainable multilingual research practice

For teams that regularly conduct research across languages, adapting on a project-by-project basis is inefficient. A few structural investments pay off over time.

Develop a bilingual reviewer network. Identify colleagues, contractors, or agency partners who can provide spot-check reviews across your most common research languages. Even a small pool of reviewers dramatically improves your ability to validate translated findings.

Standardize your translation workflow. Document which tools you use, what quality checks you perform, and how you handle uncertainty in coding. This makes your process repeatable and auditable.

Use a research repository that supports multilingual data. Storing translated and source-language transcripts together, with tags and annotations that capture translation confidence, makes it possible to revisit and revalidate past findings as translation technology improves or as bilingual reviewers become available. Dovetail's tagging and highlight features lend themselves well to this kind of structured, layered analysis across projects.

Invest in researcher judgment. Tools and processes matter, but the most important adaptation is the researcher's mindset. Working with translated data requires a higher tolerance for ambiguity, a willingness to flag uncertainty rather than resolve it prematurely, and a commitment to transparency about the limits of the analysis.

Machine translation has made cross-language research dramatically more accessible. The risk is not that the technology is too unreliable to use—it is that it is reliable enough to create false confidence. The researchers who produce the most trustworthy multilingual insights are those who understand exactly where the tool's accuracy ends and their own methodological judgment needs to begin.

FAQs

Can you do reliable thematic analysis on machine-translated transcripts?

Yes, but with adjustments. Machine translation has improved significantly, especially for widely spoken languages, and thematic analysis can surface valid patterns from translated data. The key is to treat the translated text as an approximation rather than a verbatim record. You should expect some loss of nuance—particularly around emotion, hedging, and culturally specific expressions—and build in validation steps like back-translation checks and member checking to catch meaning distortions before they become embedded in your findings.

Which machine translation tools produce the most reliable transcripts for qualitative research?

As of 2026, general-purpose neural translation engines like Google Translate, DeepL, and Microsoft Translator produce reasonably accurate output for major world languages. DeepL is often preferred by researchers for European languages due to more natural phrasing. However, accuracy drops for lower-resource languages, dialects, and domain-specific vocabulary. No single tool is universally best. Running a sample passage through two or three engines and having a bilingual reviewer compare the output is a practical way to select the right tool for a given project.

How do you handle idioms and culturally specific language in machine-translated research data?

Machine translation tends to render idioms literally, which can strip away intended meaning or introduce confusion. The most effective mitigation is to flag passages that read awkwardly or seem semantically inconsistent and route them to a bilingual reviewer for clarification. During coding, you can use a convention like bracketed annotations—[translated idiom, meaning uncertain]—so that downstream analysis treats these passages with appropriate caution rather than coding them at face value.

Editor's picks↘

Latest articles↘

Turn customer feedback into product innovation