Hacker Newsnew | past | comments | ask | show | jobs | submit | sammyyyyyyy's commentslogin

You should try it! I wouldn’t say it’s the best, far from that. But also wouldn’t say it’s terrible. If you have a 5090, then yes, you can run much more powerful models in real time. Chatterbox is a great model though

> But also wouldn’t say it’s terrible.

But you included 3 samples on your GitHub video and they all sound extremely robotic and have very bad artifacts?


Also, I didn’t want to use known voices as the example, so I ended up using generic ones from the datasets

I should have posted the reference audio used with the examples. Honestly it doesn’t sound so different from them. Voice cloning can be from a cartoon too, doesn’t have to be from a human being

A before / after with the reference and output seems useful to me, and maybe a range from more generic to more recognizable / celebrity voice samples so people can kinda see how it tackles different ones?

(Prominent politician or actor or somebody with a distinct speaking tone?)


That is probably a good idea. I was so confused listening to the example.

As I said, some reference voices can lead to bad voice quality. But if it sounds that bad, it’s probably not it. Would love to dig into it if you want

I agree with the comment above. I have not logged into hacker news in _years_ but did so today just to weigh in here. If people are saying that the audio sounds great, then there is definitely something going on with a subset of users where we are only hearing garbled words with a LOT of distortion. This does not sound like natural speech to met at all. It sounds more like a warped cassette tape. And I do not mean to slight your work at all. I am actually incredibly puzzled here to understand why my perception of this is so radically different from others!

Thank you for commenting. I wonder if this could be another situation like "the dress" (2015) or maybe something is wrong with our codecs...

No, nothing wrong with your codecs. It's sounds shitty. But given the small size and speed it's still impressive.

It's like saying .kkrieger looks like a bad game, which it does, but then again .kkrieger is only 96kb or whatever.


How big are TTS models like this usually?

.kkrieger looks like an amazing game for the mid-90s. It's incomprehensible that it's only 96kb.


Here is an overview: https://www.inferless.com/learn/comparing-different-text-to-...

Also keep in mind the processing time. The ^ article above used a NVIDIA L4 with 24-GB VRAM. Sopro claims 7.5 second processing time on CPU for 30 seconds of audio!

If you want to get real good quality TTS, you should check out elevenlabs.io

Different tools for different goals.


I mean I'm talking about the mp4. How could people possibly be worried about scammers after listening to that?

I didn’t specially cherry pick those examples. You can try it anyway for yourself. But thanks for the feedback anyway

No shade on you. It's definitely impressive. I just didn't understand people's reactions.

It sounds like someone using an electrolarynx to me.

No, it doesn’t.

Yes, you are right. However, there are many upsides to this kind of technology. For example, it can restore the voices of people who were affected by numerous diseases

Ok, that's an interesting angle, I had not thought of that, but of course you'd still need a good sample of them from before that happened. Thank you for the explanation.

Obrigado! Quando (e se fizeres isso) manda pm!

Yeah, we are not quite there, but I’m sure we are not far either

This is my side “hobby”. And compute is quite expensive. But if the community’s responsive is good, I will definitely think about it! Btw, chatterbox is a great model and inspiration

Very cool work, especially for a hobby project.

Do you have any plans to publish a blog post on how you did that? ?What training data and how much? Your training and ablations methodology, etc.


Thanks can you share details about compute economics you dealt with ?

Yeah sure. The training was about ~250 dollars, which is quite low by today’s standards. And I spent a bit more on ablations and research

I was on similar path and saw my bills going over 1000 dollars as interests to do research and ablations grew. Then I decided to get one Blackwell Pro 6000 and trying things with that :) If you have suggestions on how to manage metrics let us know. Currenty trying langfuse since its one click install on coolify

Is that something that could be done on a local setup? Eg, 2x RTX3090?

Cool! Yeah the voice quality really depends on the reference audio. Also mess with the parameters. All the feedback is welcome

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: