Which open-source TTS models still deserve a first look?

By Wei Jie Chee · 25 Mar 2026, 00:00 Z

Download printable cheat-sheet (CC-BY 4.0)
Landscape scannerCompress the field into models worth a closer look.

A useful shortlist keeps only models that match the intended voice, adaptation route, hardware, latency, and licence.

Do this first

Choose the two constraints you cannot relax, then remove models that fail either one before comparing demos.

Keep this boundary in view

This is a shortlist for further testing, not a universal quality ranking or proof on your speaker and prompts.

Inspect the evidence and what would change the answer

Mixed first-party runs and model-specific research

What supports it

Instavar has trained or run several shortlisted models, but evidence depth differs by model and experiment.

Change our mind if

A currently omitted model repeatedly meets the same constraints and avoids the costly failures on a representative local test.

Claim TTS-FIELD-2026-08 / v1, checked 7 Aug 2026

60-second takeaway
We ran a consistent FEMALE_01 benchmark across open-source TTS models on an RTX 3090 Ti (24GB).
VoxCPM 1.5 and Qwen3-TTS 1.7B produced deployable LoRA outputs. IndexTTS2 step 14,000 is the current production narration default after it ranked first in a later same-transcript comparison with VoxCPM 2 full SFT step 2,000 and VoxCPM 1.5 full SFT step 2,999. CosyVoice3's corrected LoRA run completed, but later matched checks failed the content gate.
For the lowest-friction LoRA experiment, start with VoxCPM 1.5 or Qwen3-TTS. For the current tested production default, use the pinned IndexTTS2 step 14,000 path and review its custom model license for your intended use.

Evidence updated: August 17, 2026. Repository status and visibility were checked against the current registered heads.

What this benchmark covers

This is a practitioner-oriented comparison, not an academic leaderboard. We evaluated four models under the same conditions:

  • Dataset: IMDA NSC FEMALE_01 - a TTS corpus slice with a posh or UK-influenced Singaporean voice profile
  • Hardware: one NVIDIA RTX 3090 Ti (24 GB VRAM)
  • Goal: produce voice-cloned audio suitable for AI-generated video narration (A-roll use case)
  • Evaluation: qualitative listening on naturalness, long-text stability, accent retention, and operational friction

We are not measuring WER or MOS scores from automated tools. We are measuring whether the output sounds production-ready to a human listener on a video platform.

The official NSC structure supports the FEMALE_01 provenance. The directory label should not be treated as independent proof of the number or identity of people represented in every local artifact.

The seven registered repositories behind this work

The benchmark results are only one layer of the work. We also maintain six model-specific repositories and one shared evaluator. Together they cover dataset lineage, LoRA and full-SFT lifecycles, checkpoint reload, runtime qualification, content-faithfulness checks, packaging, restore evidence, and blind-listening preparation.

Six repositories are public. The Instavar VoxCPM companion is currently private, so public readers can inspect the upstream VoxCPM repository and our on-site findings, but not the private companion's current implementation evidence. The broader /open-source index therefore lists only repositories GitHub currently exposes publicly.

These repositories do not all have the same proof level. Supported means the named path has executable evidence for its stated scope. Experimental means a bounded smoke or narrow qualification exists. Negative result