Type Nepali and hear it read by both checkpoints of the research voice, side by side.
research This is a research voice, not a broadcast reader. It runs on a small server, on the CPU, at 16 kHz.
training step 29,000
training step 40,000
Both clips come from the same model, saved at two different points in its training: step 29,000 and step 40,000. Everything else is held fixed — the same text frontend, the same prosody and pitch prediction, the same random seed — so the only thing that differs between Voice A and Voice B is the waveform decoder's weights. A vocoder is chosen by ear rather than by its loss curve, and the later checkpoint is not automatically the better one. That is the question this page exists to ask.
The decoder is 3.2 million parameters and produces 16 kHz audio. Devanagari is read directly by our own Nepali text frontend rather than transliterated through English. Pronunciation of unusual proper nouns and of English loanwords is the weakest part, and known to be.
Up to 200 characters per go — the server renders at about real time on one shared CPU, so a long paragraph is mostly a long wait. Nothing you type is stored.
| Shruti | ampixa.com/shruti |
| Nepali text frontend | github.com/Ampixa/nepa-newa-text-frontend |
| sanoTTS | github.com/Ampixa/sanoTTS |
| ampixa labs | ampixa.com |
| hello@ampixa.com |
© 2026 ampixa labs.