Audio8 TTS 0.6B: Mac, RTX 3090 Ti and LoRA Fine-Tuning Test

Download printable cheat-sheet (CC-BY 4.0)

30 Jul 2026, 00:00 Z

Audio8 TTS Preview 0.6B arrived with an unusually attractive promise: multilingual speech and zero-shot voice cloning from a compact 601-million-parameter model.

The practical question is simpler. Can it run on an ordinary Apple Silicon laptop as well as a consumer NVIDIA GPU, and what does the first real output tell us?

We pinned the model and source code, reviewed the custom Python files, and ran text-to-speech and synthetic-reference voice-cloning probes on a 16 GB M2 MacBook Pro and an RTX 3090 Ti. We then added a small LoRA training path and adapted the model on 128 cleaned clips from our NSC-derived FEMALE_01 voice profile.

60-second takeaway

  • It ran on the 16 GB M2 Mac. CPU and MPS both produced valid 44.1 kHz audio without swap activity in our short tests.
  • It fit easily on the RTX 3090 Ti. The highest observed whole-GPU memory reading was 2,938 MB during a two-item English and Chinese batch.
  • The short speech was textually clear. Whisper large-v3-turbo recovered all seven tested outputs with 0.0 WER or CER after punctuation-insensitive normalization.
  • MPS was not a clear speed win. It finished slightly sooner end to end, but its measured generation loop was not faster than CPU in these single short runs.
  • Voice cloning worked, but the quality verdict remains open. We used a synthetic macOS reference rather than a person's voice. ECAPA speaker similarity landed between 0.564 and 0.604, which is a diagnostic signal rather than a human listening score.
  • The 128-clip LoRA pilot moved toward the target voice. Mean ECAPA cosine similarity against 16 paired FEMALE_01 recordings increased from 0.092 for the base model to 0.189 after 100 steps.
  • Lower validation loss hid a serious failure. A 500-step run kept improving on validation loss, but later checkpoints often failed to stop speaking. Step 300 failed to end normally on 12 of 16 prompts.
  • The model is promising, not production-proven. Blind listening, larger clean datasets, long passages, and broader checkpoint tests remain necessary.

What Audio8 TTS Preview 0.6B is

Audio8 uses a DualAR design inspired by Fish Audio S2 Pro.

The slow autoregressive transformer predicts one semantic token for each audio frame. A smaller fast transformer predicts the remaining codec codebooks for that frame. The bundled neural codec then turns those tokens into a 44.1 kHz waveform.

ComponentReleased configuration

Need consented AI voiceovers?

Launch AI voice cloning with clear consent, pronunciation tuning, and ad-ready mixes.