A small voice model ran on both machines. Training revealed the real risk.

Audio8 fit on a 16 GB Mac and an RTX 3090 Ti. Its lowest-loss checkpoint still failed to stop speaking.

16 GBMac inference worked
2.9 GBhighest observed GPU use
100 stepsthe checkpoint we kept
Named-model investigator + Voice-quality evaluatorCompare checkpoints and voices by consequential failures.

Can a small cross-platform TTS model learn a usable voice?

Audio8 ran on both tested machines, but later lower-loss checkpoints often failed to stop speaking.

Do this first

Judge end-of-speech behavior and held-out audio before selecting the checkpoint with the lowest logged loss.

Keep this boundary in view

One model revision, corpus slice, Mac, GPU, and checkpoint sweep do not establish general Audio8 quality or stability.

Inspect the evidence and what would change the answer

First-party cross-platform run and checkpoint sweep

What supports it

The experiment records runtime, memory, training, ASR, speaker-similarity proxies, and end-of-speech behavior across checkpoints.

Change our mind if

A corrected or repeated run makes later checkpoints end reliably while improving held-out listening quality.

Claim TTS-AUDIO8-END-2026-08 / v1, checked 7 Aug 2026

The loss kept falling. The voice stopped ending.

Step 500 looked strongest in the training log. Generation exposed the opposite: later checkpoints often continued until the 500-frame safety cap. A usable voice must pass speech behavior, not only validation loss.

  1. 100selected
  2. 150lower loss
  3. 200lower loss
  4. 250lower loss
  5. 300lower loss
  6. 350lower loss
  7. 400lower loss
  8. 450lower loss
  9. 500lowest loss
ended normally hit the safety cap

The next result must be heard, not inferred.

ASR recovered the short test sentences and a speaker-embedding proxy moved toward the target profile. Neither measurement establishes naturalness or listener preference. The private comparison remains a review gate, not public evidence on this page.