195 points by nateb2022 19 hours ago | 24 comments
- This is amazing, the quality blow my mind for such small model! I just replaced my old onnx model with yours!
here my implementation with speech dispatcher and server: https://github.com/skorotkiewicz/inflect-speechd
thanks for shearing!
- Couple highlights:
> Complete local text-to-waveform speech synthesis under 10M parameters.
In case, like me, you hoped "complete" voice might mean both stt and tts. Not to speak poorly of it, just clarifying.
> English only, with one fixed male voice. This is not zero-shot voice cloning.
(And then a bunch of statements on limitations that I read as 'quality can be spotty but if you play with it it should be fine') But like. In <10M params I'm not judging:)
- When would “text-to-waveform speech synthesis” ever imply speech to text?
- The HN title is "Inflect-Micro-v2: complete voice in 9.36M parameters".
- How heavy in inference on this? The model would easily fit on many microcontroller modules, I wonder if they could run it?
- I think we need more neurons in the human brain than parameters in this model for speech. I wonder what it says about the human brain vs LLM efficiency.
- The inflections are weird but this doesn't sound like a robot. Not bad!
- I keep seeing tts stories here. Is it just an interesting subset of the llm world, or is there a huge use case I’m somehow missing?
- I use STT/TTS to interface with a local LLM for Home Assistant in my house.
- I'm do so as well, i tried qwen3 omni 3 but it was ridiculously stupid, and i ended up with stt thinker and tts. kokoro for now
- this is extremely encouraging for individuals/small companies being able to train pareto-frontier TTS models (specifically compute required to run vs quality of model output)
- I'd love to hear it but it seems your quota is exhausted.
- wrote a web-demo, runs in your browser
- Nice! Strangely, "Nano" sounds a lot better than "Micro" to me.
On my iPhone 14 Pro the page crashes after 2-3 plays. I wonder if it uses too much memory?
- right. happens on my 13 too. memory leak confirmed, not sure yet what's causing it
- Amazing quality for small size, but definitely not that enjoyable to listen to.
IMHO, its at about the same quality level of historic TTS tools.
- amazing! was looking for something similar
- amazing quality for such small size!
- Alternative title: Text to speech in 9.36M, English only.
- [flagged]
- [dead]
- [flagged]
- [dead]