What did the same-speaker adaptation runs actually establish?

By Wei Jie Chee · 07 Feb 2026, 00:00 Z

Download printable cheat-sheet (CC-BY 4.0)
Voice-quality evaluator + Choice-makerCompare checkpoints and voices by consequential failures.

Shared source data improves comparability, but uneven recipes, checkpoints, and listening evidence prevent a universal winner claim.

Do this first

Inspect which models completed, which checkpoints were heard, and which quality dimensions were actually compared.

Keep this boundary in view

One speaker corpus and uneven adaptation maturity do not establish general voice quality across languages or speakers.

Inspect the evidence and what would change the answer

First-party multi-model corpus comparison

What supports it

The benchmark preserves model-specific run outcomes and links to the deeper setup, failure, and checkpoint records.

Change our mind if

A controlled held-out comparison with matched recipes and blind listening changes the current model-level interpretation.

Claim TTS-IMDA-COMPARE-2026-08 / v1, checked 7 Aug 2026

60-second takeaway
We ran one consistent single-speaker benchmark on IMDA NSC FEMALE_01 with a single-GPU setup.
VoxCPM, IndexTTS2, and Qwen3-TTS all produced usable outputs under specific settings; CosyVoice3 did not reach production-ready quality in this run.
Treat this as an execution benchmark under one configuration, not a universal model ranking.

Who this is for

  • Founder / strategy reader: use the matrix and decision guide to pick what to deploy next.
  • Engineer reader: use each linked deep dive for exact recipes, checkpoints, and failure diagnostics.

Shared experiment setup

  • Dataset: IMDA NSC single-speaker set (FEMALE_01), with model-specific preprocessing.
  • Hardware: single NVIDIA RTX 3090 Ti (24 GB VRAM).
  • Evaluation: qualitative listening on naturalness, accent retention, noise profile, long-text stability, and operational friction (VRAM, disk, rerun complexity).

Comparison matrix

ModelDataset handlingTrain recipeBest checkpoint in this runMain failure modeRecommended inference setting
CosyVoice2 (baseline/control)Baseline sample used as controlNo finetune in this benchmarkBaseline control sample onlyNot evaluated as a finetune target in this seriesUse as control reference only
CosyVoice3