KEY TAKEAWAYS

  • Compare English WER and Korean CER in the intended voice mode; a multilingual headline is not a traffic-weighted booking-error rate.
  • SIM tests reference-speaker identity, not naturalness or listener preference. Main-table batch throughput is not single-request playback-start latency.
  • The published TTFA protocol uses warm, English, default-voice synthesis. Korean, cloning and the full agent turn need their own application evidence.

A shortlist is justified; a bilingual winner is not

For a founder choosing an English/Korean booking-agent voice, a low English error rate is useful evidence about text recovery.[1] It is not enough to establish that customers will hear the right Korean name, recognize a consistent speaker, enjoy the delivery or receive a timely reply. Those are different questions, and Open TTS gives partial answers to several of them—not one overall quality score.[1][2]

Hugging Face’s Open TTS article, published 30 September 2026, describes automated intelligibility, speaker-identity and speed comparisons alongside a way to listen to stored outputs.[1] It explicitly says the leaderboard does not replace human preference ranking.[1] Our recommendation is to use it to reduce the candidate set, then close the language, voice-mode and application-timing gaps that matter to the intended conversation.

What we reviewed, and what we did not measure

This is a purposive documentary comparison of the Open TTS explainer, current official Space documentation, display and dataset/language aggregation sources, and the original Seed-TTS/CV3 metric READMEs. We checked the official explainer and Space sources on 3 October 2026 KST, while retaining the upstream README snapshots retrieved on 1 October. The explainer and Space describe the same benchmark; they are complementary evidence, not independent replications.[1][2] The aggregation source was newly consulted at the same observed repository commit to resolve which dataset/language pairs require results; it was read as text, not executed.[6]

Our byte comparison found the current metric-documentation and display-source files identical to the preserved 1 October copies. The newly consulted aggregation source has no archived counterpart in that original packet, so no historical byte-equality claim is made for it.[6] The original blog extraction omitted publication metadata; the current page displays 30 September.[1] Neither retrieval date identifies when the audio runs occurred. We did not synthesize or listen to audio, time an application, audit the full generation harness or inspect numerical table images. No model ranking or winner is asserted.

Protocol statements below are the publisher’s descriptions, not PALANTHOS measurements. Actual benchmark run dates and retained corpus denominators are unknown. The English/Korean application exercise is hypothetical and unexecuted.

What each measure can carry into the decision

Method and comparability map—not a ranking. Source conditions checked 3 October 2026; application implications are our interpretation.
MeasureDocumented comparisonUse it for / do not infer
English WER (word error rate)Automatic speech recognition (ASR) transcription versus input text. Default English rank is a macro-average over Seed-TTS and CV3 English splits; current Open TTS names Qwen3-ASR.[1][2]Shortlist for English text recovery under that evaluator and voice mode. Not Korean accuracy, naturalness or listener preference.
Korean CER (character error rate) and multilingual averageKorean uses CV3; Seed-TTS has English and Chinese only.[1] Per-language scores average applicable selected datasets; a dataset without the language sits out. The headline macro-average weights selected language scores equally.[2][6]Inspect English and Korean separately. The headline is not pooled word error, a CER-to-WER conversion or your traffic-weighted failure rate.
SIM (speaker similarity) in voice cloningCosine similarity of WavLM speaker embeddings between generated and reference audio. Only cloning-mode runs report SIM; default and cloned voices are separate evaluations.[2]Screen reference-speaker identity when cloning is required. Not a human preference score or evidence that a default voice sounds the same in both languages.
Main-table RTFx (inverse real-time factor)Generated audio duration divided by generation time, on the same GPU: batch 32 where supported, otherwise batch 1.[2]Evidence about the reported generation throughput. Not how soon one caller hears the first syllable.
Streaming TTFA (time to first audio)First chunk for streaming; whole generated utterance for non-streaming. Batch 1, default voice, 50 English CV3 prompts; three warm-ups omitted, median of the remaining 47. GPU and CPU views are separate.[1][2]Screen warm English/default-voice playback-start behavior within the same hardware view. Not Korean, cloned-mode, cold-start, tail-latency or full agent-turn evidence.
Streaming RTFxSeconds of audio produced per second of compute in the single-request streaming tab, at batch 1.[2]Consider sustained generation pace separately from TTFA. Do not substitute main-table batch throughput or assume uninterrupted application playback.

The English/Korean comparison needs two views, not one percentage

In the default English view, the explainer describes an average across two dataset splits.[1] For Korean, the available dataset is CV3—not a second Korean split from Seed-TTS.[1] With both datasets and English/Korean selected, English draws on Seed-TTS and CV3, while Korean draws on CV3 only: a dataset without the language sits out its average.[6] The headline macro-average then weights the selected language scores equally.[2][6] Both languages can therefore appear in a comparable selection without requiring a nonexistent Korean Seed split, provided the model has the applicable results.[6]

The UI uses CER for Chinese, Japanese and Korean, and WER for the other displayed language columns.[3] Character errors and word errors have different reference units.[1][4] Equal-looking percentages do not become interchangeable simply because both are error rates. A multilingual macro-average is a summary of the selected language scores, not a direct estimate of incorrect booking details in your customer mix. That last distinction is our interpretation, not a measured deployment result.

Coverage must stay visible, but a dataset that lacks a language is not the same as a model missing an evaluation. The Space’s short summary describes full coverage across the selection.[2] Its aggregation source clarifies that completeness applies to the selected datasets that actually cover each selected language.[6] With English and Korean selected across Seed-TTS and CV3, English uses both datasets and Korean uses CV3 only.[1][6] A missing model result on an applicable split can withhold a comparable headline average and rank; the nonexistent Korean Seed split does not contribute zero or require a score.[3][6] Do not fill missing model results with zero or read an unranked row as proof that the model is bad.

Decide the voice mode before comparing candidates. The Space describes a built-in/default voice and a separate reference-conditioned cloning evaluation, with potentially different WER.[2] A default-voice English score and a cloned-voice Korean score do not answer the same configuration question. If the product needs cloning, use only a consented reference voice; if it needs a built-in voice, SIM from a cloning run does not establish that built-in voice’s appeal.

Voice identity is useful evidence—and still not a listening verdict

SIM answers whether generated audio is close to a reference in a speaker-embedding space. The current Space specifies WavLM cosine similarity and reports it only for cloning.[2] That is relevant when preserving a consented speaker’s identity is a requirement. It does not directly tell you whether delivery is natural, expressive or preferred by listeners; the publisher explicitly separates these judgments from WER and SIM.[1]

There is also a concrete harness distinction. The original Seed-TTS and CV3 READMEs name Whisper-large-v3 and Paraformer for their English/Chinese ASR metrics.[4][5] CV3 names ERes2Net for speaker similarity.[4] Current Open TTS instead names Qwen3-ASR and WavLM.[2] Shared dataset names therefore do not establish an identical evaluator. This is a difference between documents, not proof that any historical score changed or that a particular old run used the current settings.

The counterpoint is that these proxies are still useful. Holding language, split, normalization, evaluator and voice mode steady makes text-recovery and identity evidence more relevant to a shortlist. The publisher’s Listen view uses stored evaluation clips, giving a reader an additional route to inspect the outputs behind the metrics.[2] We did not perform that listening comparison, and an individual audition would not automatically be a representative user study.

Batch capacity and the first audible reply are different endpoints

RTFx measures how much audio is generated per unit of generation time.[2] In the main table, a batching-capable model is evaluated at batch 32 while one without batching is evaluated at batch 1, on a fixed GPU.[2] This is meaningful for offline capacity planning, but the capability-dependent batch setting belongs in any comparison. A high main-table RTFx does not establish a short wait for a single interactive request.

TTFA is closer to the interactive question because it asks when audio becomes available to play. The explainer distinguishes first chunk for streaming models from completed utterance for non-streaming ones.[1] The shorter Space definition only says first streamed chunk.[2] We retain the explainer’s non-streaming qualification rather than treating every row as the same first-chunk endpoint.

The reported protocol uses the same 50 English CV3 prompts, default voice and batch 1; it discards three warm-up runs and reports the median across the remaining 47.[1] Current Space documentation specifies 47 timed samples with targets of 5–32 words and a median of 13 words, and separate CPU/GPU execution modes.[2] These are synthesis samples, not listeners. Omitting warm-ups means the result is not a cold-start guarantee; a median does not describe the slowest replies. English/default timing does not establish Korean/cloned timing. These limits follow from the documented cohort and statistic, not a test we ran.

For a voice agent, our proposed user-facing boundary starts when the user finishes speaking and ends when the reply becomes audibly playable in the application. It includes whatever turn detection, ASR, reasoning, networking, buffering and playback the implementation actually uses. Stages may overlap; this is not a claim that their individual times can simply be added. The published TTS request-to-audio endpoint does not measure that full path, nor does it establish interruption handling or continuous playback under load.

Hypothetical exercise: an English/Korean booking agent

This exercise is proposed and unexecuted. Assume a candidate has a lower English WER than your existing voice, but no application-specific Korean, listening or full-turn timing evidence has been supplied. Keep it on the shortlist; do not select it solely from that English result. If no existing voice is deployed, compare with another eligible shortlist candidate rather than inventing a measured baseline.

  • Fix the intended mode: built-in voice or consented reference-conditioned voice. Record the exact model/configuration, evaluator, normalization, language/split, hardware and load. Do not borrow English default-voice latency as evidence for Korean cloning.
  • Prepare a bounded, preselected set of English and Korean booking names, dates, times, confirmation statements and questions. Include within-turn language switches if the intended product requires them. No sample size or participant count has been chosen here; the set would be a diagnostic screen, not a claim of production-wide equivalence.
  • Define critical-detail failures before collecting results: a changed date or booking name should remain visible even if the aggregate text error rate is low. Keep language-level WER/CER diagnostics separate from whether a listener can correctly recover the intended details.
  • Ask suitable listeners to judge intelligibility, naturalness, speaker consistency and preference separately, using the same prompts and hiding model labels where feasible. Record listener language competence and recruitment limits. Do not turn a SIM value into a preference rating.
  • Measure the actual user-turn-to-audible-reply boundary in the application. Keep first audio, full utterance completion and sustained playback distinct. Record warm versus cold starts, load and unsuccessful turns; inspect slow replies as well as a central statistic.
  • Set acceptable detail accuracy, voice judgments and timing requirements in advance. If a candidate violates a critical booking requirement, retain the baseline or defer selection rather than trading it away inside one average.

This would be a comparison of the configured voice workflow, not an isolated causal estimate of model quality. Changes in ASR, text normalization, voice reference, prompt mix, hardware, concurrency, network conditions or player buffering could change the result. We have no observed application data, listening results, improvement estimate or winning model.

What would justify choosing a voice?

Choose a limited application trial when an eligible candidate has relevant per-language evidence in the intended mode, and the proposed listening and timing checks can test the requirements that matter. Final selection needs those application results—not merely a favorable English rank or multilingual headline. If required Korean coverage or comparable mode evidence is missing, defer that comparison rather than silently fill the gap.

This is not a blanket dismissal of benchmarks. Strong, matched intelligibility evidence can rule out weak candidates; cloning-mode SIM can inform an identity requirement; batch RTFx can inform offline capacity; and single-request TTFA can inform warm playback-start screening. Each becomes misleading only when asked to answer a different question. Evidence that would change a shortlist into a selection is a matched English/Korean application comparison meeting predeclared detail, listener and timing requirements. This article supplies the decision boundary, not that result.

Sources & scope

Documentary research checked 3 October 2026 KST. Official Open TTS explainer and current Space sources were compared with preserved Seed-TTS/CV3 README evidence. Protocols and measurements are publisher-reported, not independently replicated. No audio generation, listening session or application timing was performed; the booking exercise is hypothetical and unexecuted. Measurement dates, retained corpus denominators and historical harness equivalence remain unknown.

  1. Open TTS Leaderboard: Scalable Evaluation for Multilingual Text-to-Speech and Voice Cloning ↗

    Published 30 September 2026 as displayed on the current page; checked 3 October 2026 KST. Multilingual evaluation, preference limits and streaming protocol sections. Exact publication time, update date and benchmark run dates unknown; publisher report, not our replication.

  2. Open TTS Space metric and protocol documentation ↗

    Official Space constants.py, checked 3 October 2026 KST at repository commit 3456d010de2735c08b8823f8ee9cd3bcb492f311. Leaderboard, Streaming, About and Listen documentation; byte-identical to the preserved 1 October copy. Current documentation does not certify historical run settings.

  3. Open TTS leaderboard view implementation ↗

    Official Space leaderboard_views.py, read as text only on 3 October 2026 KST at the same observed repository commit. CER labels, complete-coverage ranking and CPU/GPU measured-row selection; no execution or harness validation.

  4. CV3-Eval README ↗

    Original CV3-Eval README snapshot retrieved 1 October 2026; Metrics section. Whisper/Paraformer and ERes2Net describe that upstream benchmark, not an identical Open TTS evaluator. Publication/update and measurement dates unknown; not newly fetched for this article.

  5. seed-tts-eval README ↗

    Original seed-tts-eval README snapshot retrieved 1 October 2026; corpus and Metrics sections. English/Mandarin coverage, upstream ASR and WavLM SIM. Upstream corpus counts do not establish Open TTS retained denominators. Publication/update and measurement dates unknown; not newly fetched for this article.

  6. Open TTS dataset/language aggregation implementation ↗

    Official Space leaderboard_data.py, read as text only on 3 October 2026 KST at observed repository commit 3456d010de2735c08b8823f8ee9cd3bcb492f311. aggregate_rows, lines 485–563: applicable dataset/language coverage, averages and incomplete headline handling. Newly consulted; no original snapshot counterpart. Publication/update and benchmark run dates unknown; no execution or full harness audit.

Publication history
  • — Prepared initial documentary research version; this timestamp identifies the candidate, not live publication or activation.
  • — Prepared coverage correction: only datasets that contain a selected language require a model result. Candidate timestamp, not production activation.

Have a correction or a different perspective? Contact Palanthos.