What did the same-speaker adaptation runs actually establish?
By Wei Jie Chee · 07 Feb 2026, 00:00 Z
Download printable cheat-sheet (CC-BY 4.0)Shared source data improves comparability, but uneven recipes, checkpoints, and listening evidence prevent a universal winner claim.
Do this first
Inspect which models completed, which checkpoints were heard, and which quality dimensions were actually compared.
Keep this boundary in view
One speaker corpus and uneven adaptation maturity do not establish general voice quality across languages or speakers.
Inspect the evidence and what would change the answer
First-party multi-model corpus comparison
What supports it
The benchmark preserves model-specific run outcomes and links to the deeper setup, failure, and checkpoint records.
Change our mind if
A controlled held-out comparison with matched recipes and blind listening changes the current model-level interpretation.
Claim TTS-IMDA-COMPARE-2026-08 / v1, checked 7 Aug 2026
60-second takeaway
We ran one consistent single-speaker benchmark on IMDA NSC FEMALE_01 with a single-GPU setup.
VoxCPM, IndexTTS2, and Qwen3-TTS all produced usable outputs under specific settings; CosyVoice3 did not reach production-ready quality in this run.
Treat this as an execution benchmark under one configuration, not a universal model ranking.
Who this is for
- Founder / strategy reader: use the matrix and decision guide to pick what to deploy next.
- Engineer reader: use each linked deep dive for exact recipes, checkpoints, and failure diagnostics.
Shared experiment setup
- Dataset: IMDA NSC single-speaker set (
FEMALE_01), with model-specific preprocessing. - Hardware: single NVIDIA RTX 3090 Ti (24 GB VRAM).
- Evaluation: qualitative listening on naturalness, accent retention, noise profile, long-text stability, and operational friction (VRAM, disk, rerun complexity).
Comparison matrix
| Model | Dataset handling | Train recipe | Best checkpoint in this run | Main failure mode | Recommended inference setting |
| CosyVoice2 (baseline/control) | Baseline sample used as control | No finetune in this benchmark | Baseline control sample only | Not evaluated as a finetune target in this series | Use as control reference only |
| CosyVoice3 |