Which open-source TTS models still deserve a first look?

By Wei Jie Chee · 25 Mar 2026, 00:00 Z

Download printable cheat-sheet (CC-BY 4.0)
Landscape scannerCompress the field into models worth a closer look.

A useful shortlist keeps only models that match the intended voice, adaptation route, hardware, latency, and licence.

Do this first

Choose the two constraints you cannot relax, then remove models that fail either one before comparing demos.

Keep this boundary in view

This is a shortlist for further testing, not a universal quality ranking or proof on your speaker and prompts.

Inspect the evidence and what would change the answer

Mixed first-party runs and model-specific research

What supports it

Instavar has trained or run several shortlisted models, but evidence depth differs by model and experiment.

Change our mind if

A currently omitted model repeatedly meets the same constraints and avoids the costly failures on a representative local test.

Claim TTS-FIELD-2026-08 / v1, checked 7 Aug 2026

60-second takeaway
We ran a consistent FEMALE_01 benchmark across open-source TTS models on an RTX 3090 Ti (24GB).
VoxCPM 1.5 and Qwen3-TTS 1.7B produced deployable LoRA outputs. IndexTTS2 step 14,000 is the current production narration default after it ranked first in a later same-transcript comparison with VoxCPM 2 full SFT step 2,000 and VoxCPM 1.5 full SFT step 2,999. CosyVoice3's corrected LoRA run completed, but listening evaluation is still pending.
For the lowest-friction LoRA experiment, start with VoxCPM 1.5 or Qwen3-TTS. For the current tested production default, use the pinned IndexTTS2 step 14,000 path and review its custom model license for your intended use.

Evidence updated: July 19, 2026.

What this benchmark covers

This is a practitioner-oriented comparison, not an academic leaderboard. We evaluated four models under the same conditions:

  • Dataset: IMDA NSC FEMALE_01 - a TTS corpus slice with a posh or UK-influenced Singaporean voice profile
  • Hardware: one NVIDIA RTX 3090 Ti (24 GB VRAM)
  • Goal: produce voice-cloned audio suitable for AI-generated video narration (A-roll use case)
  • Evaluation: qualitative listening on naturalness, long-text stability, accent retention, and operational friction

We are not measuring WER or MOS scores from automated tools. We are measuring whether the output sounds production-ready to a human listener on a video platform.

The official NSC structure supports the FEMALE_01 provenance. The directory label should not be treated as independent proof of the number or identity of people represented in every local artifact.

The four models

VoxCPM 1.5

VoxCPM 1.5 uses a LoRA finetuning path that fits within 24GB VRAM without modification. Training is straightforward with standard train/val splits.

DimensionResult
Finetuning approachLoRA
Best checkpoint (this run)step_0004000