Did CosyVoice 3 improve the voice, or only change the pipeline?

By Wei Jie Chee · 07 Feb 2026, 00:00 Z

Download printable cheat-sheet (CC-BY 4.0)
Voice-quality evaluator + Named-model investigatorCompare checkpoints and voices by consequential failures.

The audible result must be separated from architecture novelty, zero-shot behavior, and an unfinished adaptation rerun.

Do this first

Listen for identity, intelligibility, prosody, drift, and long-form stability before interpreting the model version as progress.

Keep this boundary in view

The available samples and run maturity do not support a universal CosyVoice 2 versus 3 quality ranking.

Inspect the evidence and what would change the answer

First-party samples with explicit incomplete lanes

What supports it

The page distinguishes zero-shot controls, fine-tuning results, listening notes, and the status of the corrected LoRA route.

Change our mind if

Blind held-out listening reverses the stated trade-offs or the completed rerun removes the observed weaknesses.

Claim TTS-COSYVOICE-QUALITY-2026-08 / v1, checked 7 Aug 2026

Experiment Status: LoRA rerun completed - best checkpoint at epoch 12. Listening evaluation pending.
60-second takeaway
The first CosyVoice3 run (full SFT) failed after epoch 1 with catastrophic overfitting. A corrected LoRA rerun (2.16M params via PEFT) reached its best checkpoint at epoch 12 with stable training.
CosyVoice2 baseline audio remains the control. The LoRA rerun tools and 9 pitfalls are published at instavar/cosyvoice3-lora-finetuning.

If you searched for CosyVoice 2 quality, CosyVoice 2 voice cloning quality, CosyVoice 2 quality review 2026, or CosyVoice 2 vs CosyVoice 3, this page is the quality-comparison view. It should be read beside the full CosyVoice fine-tuning guide, which covers the data, VRAM, LoRA, and rerun details.

Where this fits

  • For founders: do not deploy this CosyVoice3 run as-is.
  • For engineers: use this page as a diagnostic handoff for the next rerun.

Quick read:

  • CosyVoice2 remains the cleaner control sample from this evidence set.
  • CosyVoice3 is still attractive for zero-shot quality, but this specific fine-tuned run did not clear production listening review.
  • The corrected CosyVoice3 LoRA rerun is operationally healthier than the full-SFT run, but quality promotion still depends on listening evaluation.

Series overview:

For the full cross-model comparison, see the TTS Model Decision Tree - CosyVoice 3 is recommended for pre-produced content when zero-shot consistency matters most.

Result summary

CosyVoice2 is included as a baseline/control and produced acceptable qualitative output on our selected sample. CosyVoice3 finetuning in this run did not reach production-ready quality, with unstable long-form behavior and weaker linguistic consistency in listening checks.

Audio evidence

CosyVoice2 baseline/control

CosyVoice3 representative sample from this run