At a glance
| Model | piper-en_AU-librivox-medium, VITS medium, ~70M parameters, 77 MB ONNX |
|---|---|
| Voices | 10 Australian English narrators in one multi speaker model |
| Training data | 31.01 hours, 17,454 clips of public domain LibriVox recordings |
| Audio | 22,050 Hz mono, resampled from 24 kHz |
| Base checkpoint | en_GB-jenny_dioco-medium, warm started |
| Trainer | OHF-Voice/piper1-gpl v1.8.0 |
| Training run | 599 epochs, 146,999 optimiser steps, 29.6 hours on one 96 GB GPU |
| Shipped metrics | val_mos 3.61 (UTMOS), val_mel 0.46 |
| Licence | CC BY 4.0, with attribution to Jenny (Dioco) for the base checkpoint |
The model is on Hugging Face, and every step that built it, from dataset preparation to ONNX export, is on GitHub. A clean clone plus one build script reproduces the whole model.
Why Piper
Piper is fast, small and runs entirely offline, on anything from a laptop CPU to a Raspberry Pi. That makes it a natural fit for home automation, accessibility tools, kiosks and any system where sending text to a cloud speech service is slow, expensive or not allowed. It is a good first release for a lab whose goal is AI that runs on hardware you own.
The data: public domain audiobooks
Speech models need hours of clean audio with accurate transcripts, and the licence on that audio decides what you can do with the model. We started from ablmontazer/australian-english-speech, a CC0 dataset of 61,662 clips (110.2 hours) from 214 LibriVox volunteers reading 37 public domain Australian books.
LibriVox recordings are released into the public domain, not just under a permissive licence. That is unusual for speech data of this size, and it is what makes a model you can redistribute and use commercially possible.
The ear test overruled every metric
The dataset is made of Australian titles, but LibriVox volunteers are international. An Australian book is no guarantee of an Australian narrator, and the dataset's own card says narrator accent is unverified. So we auditioned candidates in two rounds and kept only the ten judged by a native listener to have Australian accents.
The largest reader in the corpus, with 11.79 hours, was rejected at this stage after being the obvious pick on every number available. Transcript quality mattered too: the share of clips ending in punctuation ranged from 0% to 82% between narrators, and the phonemiser drives intonation from punctuation. Two otherwise large narrators were excluded partly on that basis.
This is a subjective judgement by one person, and it could be wrong in either direction. Only one narrator carries documentary confirmation: his own recorded outro says he is recording from the Gold Coast.
Cleaning the data
- We removed 117 LibriVox announcement clips that the upstream filter missed. Repeated identical short phrases are exactly what this kind of model latches onto.
- Audio was resampled from 24 kHz to 22,050 Hz, because only Piper's medium checkpoints fine tune without extra vocoder settings.
- Every training file is a projection from one manifest. Decoding the audio took 20 minutes once; every later data fix took seconds.
The worst bug of the project hid in the metadata. Piper reads its training csv with Python's csv reader, which treats a transcript that begins with a quote mark as the start of a quoted field. Audiobook dialogue opens with quotes constantly, so one narrator's 1,900 rows parsed as 1,240, with a single field 25,890 characters long. Training then asked for 437 GiB of GPU memory, an error that looks exactly like a batch size problem and is not. The fix was to strip quotes and validate every file with the same parser Piper uses.
Training
We fine tuned from the en_GB jenny_dioco medium checkpoint using warm start, which copies every weight that fits and initialises a fresh speaker embedding for the ten narrators. Phonemes come from espeak-ng's received pronunciation voice: espeak has no Australian voice, and Australian English is non rhotic, so RP is a closer starting point than American English.
We benchmarked two GPUs before committing. A 96 GB card moved 1.65 times more audio per second than a 24 GB card, but made fewer optimiser updates per hour at its larger batch size. For the small single narrator test the smaller card won; for the full ten voice model, with 15,709 clips per epoch, raw throughput mattered more and the larger card trained it.
| Card | Batch | Samples per second | Steps per second | GPU use |
|---|---|---|---|---|
| 24 GB | 32 | 45.8 | 1.43 | 77% |
| 96 GB | 64 | 75.6 | 1.18 | 61% |
The shipped model trained for 599 epochs and 146,999 optimiser steps, 29.6 hours on the 96 GB card with 16 bit mixed precision.
Choosing the checkpoint by ear
Piper tracks a predicted listener score (val_mos) and a reconstruction loss (val_mel). Both are useful, and neither is the final word. The highest score of the project, 3.76 at epoch 544, was an outlier against its neighbours at 3.60 to 3.65, and the score varies by about 0.08 between identical runs. We shipped epoch 599 because it sounded better.
The biggest lesson: the model was perceptually converged by about epoch 300, roughly 13 hours in. A second 16 hour stretch improved the score by about 0.14 and produced audio we could not reliably tell apart. And resuming a nearly converged run cost 0.9 points of score and nine hours to recover, with a complete checkpoint and no bug involved. Next time we will cap training high, stop by ear, and never resume a finished run.
What it sounds like, and where it falls short
The ten voices are not equally good, and almost all of the difference is inherited from the source recordings. Volunteers record at home on their own microphones, and no amount of training fixes a poor microphone or volume that drifts between sessions.
- Period accent. Every source book predates 1930, so the voices read formally, like literature, rather than conversationally.
- Narrow domain. Numbers, abbreviations, URLs and technical terms are the weakest ground.
- Machine transcripts. Transcripts were generated by faster-whisper and not corrected by hand.
- Pacing. Some voices rush. The model card recommends a slower speech rate for three of them.
Licence and attribution
The model is released under CC BY 4.0, so commercial use, redistribution and further fine tuning are all permitted. The base checkpoint derives from the Jenny TTS dataset, recorded by Jenny and published by Dioco, whose licence requires attribution to Jenny (Dioco) wherever the model is served in software or an interface. The ten voices are named for the LibriVox narrators who recorded them, and their original names are kept alongside.
What is next
This first voice proved the pipeline. Lyrebird, our next generation Australian voice model, will aim for more voices, cleaner and more conversational source audio, and a larger training effort. It is on our model roadmap.
Try it
pip install piper-tts
hf download DataCraftsmanAustralia/piper-en_AU-librivox-medium \
en_AU-librivox-medium.onnx en_AU-librivox-medium.onnx.json --local-dir voices
echo 'Good on ya, love. The arvo turned out alright in the end.' \
| python -m piper -m voices/en_AU-librivox-medium.onnx -s 5 -f out.wavListen to samples of all ten voices on the model card.