Skip to main content
Legal·9 min read·

Foreign-language audio evidence: transcript and translation

A recording in another language is one of the more common pieces of evidence in California litigation and one of the more commonly mishandled. The work required is not transcription alone and not translation alone. It is a combined process, sometimes called transcription translation, and it has a professional standard behind it because the output is used as evidence.

Two judgments, kept separate

The reason this is a distinct discipline is that two different decisions are being made, and collapsing them destroys the audit trail.

The first is what was said. That is a transcription judgment about a recording that may be noisy, overlapping, accented, or partly inaudible. The second is what it means in English. That is a translation judgment applied to the transcribed source text.

When someone listens to a Spanish recording and types an English document, both judgments happen invisibly and simultaneously, and the result cannot be checked. There is no record of what the person heard, so a disagreement about the English cannot be traced back to whether the underlying words were even correctly identified. The defensible approach produces a verbatim transcript in the source language first, then a translation of that transcript, so each judgment can be examined on its own.

The National Association of Judiciary Interpreters and Translators publishes "General Guidelines and Requirements for Transcription Translations in a Legal Setting for Users and Practitioners," dated May 2019, which is the reference point for this work.

What a defensible exhibit contains

  • A verbatim transcript in the source language, with timestamps.
  • An English translation, presented so that each passage can be matched to its source, commonly in a side-by-side or facing-column format.
  • Speaker identification, with the basis for it stated where speakers are identified by inference rather than by self-identification on the recording.
  • Explicit marking of anything inaudible or unintelligible, rather than a plausible guess presented as text.
  • Notes where a term has no clean English equivalent, or where slang, regional usage, or a coded reference required a judgment.
  • Identification of who performed the work and their qualifications, with a certification statement covering completeness and accuracy.
  • Identification of the source recording, its length, and its format.

That last group matters for the same reason it matters with any translated exhibit. California Evidence Code section 753 addresses translators of writings and provides that the record identify the translator, and section 751(c) sets out the translator's oath to make a true translation into English. Anonymous work product is difficult to defend.

The failure modes

Four recur, and all four are visible on the face of a poor exhibit:

  • The summary. An English narrative of what the speakers discussed, with no transcript. It cannot be checked against the audio and it embeds the summariser's judgment about what mattered.
  • The confident guess. Inaudible passages rendered as clean text with no marking, which is the single most damaging habit in this work because it looks like the most reliable part of the document.
  • The literal rendering of idiom. Slang and regional idiom translated word for word, producing text that is either meaningless or, worse, appears to mean something the speaker did not say.
  • The uncredited file. A transcript and translation with nobody named, no methodology stated, and no certification. Everything about it has to be established from scratch if it is challenged.

Why the language variety matters here specifically

Recorded speech is casual speech, and casual speech is where regional variation lives. A speaker's vocabulary, idiom and register may be specific to a country or a region in ways that formal written language is not, and a translator unfamiliar with that variety can produce a technically accurate rendering that misses the meaning entirely.

This is the practical reason to identify the variety before the work starts rather than to treat a language as monolithic. Our article on Spanish dialects in legal and medical interpreting covers the point for the most common case.

Where this shows up

Most often in four places, and the standard is the same in each:

  • Recorded statements taken by a carrier or an SIU investigator during a claim investigation.
  • Surveillance audio and video in personal injury and workers compensation matters.
  • Jail calls and intercepted communications in criminal defense.
  • Voicemails, recorded meetings, and customer calls in employment and commercial disputes.

In each of them the recording is often the strongest evidence available, which is precisely why the document produced from it should be able to withstand a competent challenge. AMS provides this work through our audio transcription service, including bilingual side-by-side transcripts, and supports criminal defense and insurance defense matters.

What this work does not establish

A transcript and translation establish what the recording contains. They do not establish that the recording is authentic, unaltered, or what it purports to be, and they do not establish who the speakers are beyond what a listener can reasonably discern.

That boundary is worth respecting in how the work is instructed and described. A translator who identifies speakers as Speaker 1 and Speaker 2, and separately notes that Speaker 1 is referred to as a particular name by another speaker, is doing the job correctly. A translator who labels the columns with the parties' names is making an identification that is not theirs to make, and it hands opposing counsel a straightforward objection.

The same applies to intelligibility. Marking a passage inaudible is not a failure, it is the honest output of the process, and a transcript with a realistic number of such markings is more credible than one with none.

Time and cost, realistically

This work is slower than people expect and budgeting for it as though it were dictation transcription causes friction later. The reasons are structural rather than commercial.

  • Audio quality drives everything. Clean studio-quality speech moves quickly. A recorded call with background noise, overlapping speakers and a poor connection can take many times longer per minute of audio.
  • Two passes are required, one for the source transcript and one for the translation, plus review.
  • Overlapping speech has to be resolved by repeated listening, and there is no shortcut.
  • Rare languages and specific regional varieties narrow the pool of people who can do the work at all.

The practical consequence for a litigation timetable is that recorded evidence should be sent for processing when it is obtained, not the week before it is needed. Where only part of a long recording matters, identifying the relevant segments by timestamp is the single biggest cost saving available.

Practical guidance for instructing this work

  • Provide the best available copy of the recording, not a re-recording or a compressed export.
  • State the language and, where known, the country or region of the speakers.
  • Say whether you need the whole recording or identified segments, and give timestamps if the latter.
  • Provide context: names, places, and case-specific terms, so they are transcribed and rendered consistently.
  • Ask for the source-language transcript as well as the translation, even if you only intend to use the English.
  • Ask what the provider does with inaudible passages before commissioning the work. The answer tells you most of what you need to know.

One more instruction worth giving explicitly: ask for the work to be reviewed by a second qualified linguist before delivery. Recorded evidence is exactly the material where a second pass catches things, because the first listener has by then heard the audio so many times that they hear what they expect rather than what is there.

Using the exhibit at a deposition

Where a bilingual transcript is put to a deponent, a few habits keep the record clean. Mark the source-language transcript and the English translation together, so the exhibit contains both. Identify on the record who prepared each, consistent with the Evidence Code requirement that the record identify the translator. And where a passage is disputed, identify it by timestamp and line rather than characterising the recording generally.

If the deponent is being asked about the recording through an interpreter, resist the temptation to have the interpreter render the audio live. That is a different task performed under worse conditions than the transcript work, and it produces a second, competing version of the same content in the same transcript. Put the prepared exhibit to them instead.

Sources

California Evidence Code sections 751 and 753, available through California Legislative Information; "General Guidelines and Requirements for Transcription Translations in a Legal Setting for Users and Practitioners" (May 2019), listed by the National Association of Judiciary Interpreters and Translators. Verified in July 2026.

AMS produces verbatim and bilingual transcripts of recorded evidence with certification statements. See our audio transcription service or request a quote.

Frequently asked

Related questions

A summary is not evidence of what was said, and it cannot be checked. If the recording matters, it needs a verbatim transcript in the source language and a translation of that transcript, produced so that both can be examined.

Need certified interpreters or translation?

AMS schedules nationwide.