free web page hit counter
πŸ›‘οΈ
Copyright Notice: This video is officially sourced and embedded from YouTube. For all copyright inquiries, reports, or removals, please contact YouTube's legal team here.
Better Stack

Better Stack

181,000 subscribers

⏱ πŸ‘ 77,065 views

I Tried the Open Source ElevenLabs Alternative (Voicebox)

Video Overview & Insights

Voicebox is a free open-source local AI voice studio for developers that brings voice cloning, text-to-speech, system-wide dictation, agent voice output, MCP integration, Stories editing, and a local REST API into one polished desktop app.

Whoever is editing you sound is already making it sound robotic , it could be because of too much noise supression or bad EQ

β€” @RiteshKumarPanda

In this video, I test Voicebox and show why devs are paying attention to it as a local-first alternative to cloud AI voice tools like ElevenLabs, especially if you care about privacy and unlimited generation.

πŸ”— Relevant Links

Can I generate other voice with it please?

β€” @QuantumMeContact

Voicebox Repo - https://github.com/jamiepine/voicebox

Voicebox - https://voicebox.sh/

thats great tool i might like to use it in projects or website to talk

β€” @Ultimateboat27

❀️ More about us

Radically better observability stack: https://betterstack.com/

3:47 now that's for real

β€” @Lyter-n3c

Written tutorials: https://betterstack.com/community/

Example projects: https://github.com/BetterStackHQ

Not good for long form video. The voice starts running and i takes forever to generate

β€” @andrebugotti

πŸ“± Socials

Twitter: https://twitter.com/betterstackhq

What is the best paid AI voice subscription in 2026?

β€” @Urfusion7

Instagram: https://www.instagram.com/betterstackhq/

TikTok: https://www.tiktok.com/@betterstack

Appreciate this video. I was able to get this working with help from Claude Code, but learned about it from you. Don't need Eleven Labs for my OpenClaw anymore. Thanks!

β€” @OpenClawSkillsReview

LinkedIn: https://www.linkedin.com/company/betterstack

πŸ“Œ Chapters:

hmmm I'm using two voices from elven labs for the basis of the voices of my christmas display characters, the audio gets very heavily modified, I wonder if I have enough spoken words to clone them - would save me ages of modding each sentance

β€” @freman

0:00 Voicebox: Free Local AI Voice for Developers

0:40 What Is Voicebox and Why Developers Care

Why is the audio quality of the local voice cloning better than the actual mic audio?

β€” @AetherSummers

1:53 Live Voicebox Demo: Voice Cloning, TTS, and Dictation

2:20 Cloning a Voice and Generating Local AI Speech

3:08 It sounded nothing like him.

β€” @fjccommish

3:13 Using Voicebox for Dictation

3:30 Give Claude Code a Voice with Voicebox

I suggest you also look at Fono. More minimal but still very powerfull with TTS, STT and LLM.

β€” @BR.

4:00 Voicebox vs ElevenLabs and Open Source Voice Tools

5:45 Real Talk: Voicebox Pros, Cons, and Current Limitations

Tried this, but it was taking between 15 and 30 seconds to respond...I may not be doing it right, since I'm relatively new to self hosted Ai...

β€” @1ScaryScot

6:55 Is Voicebox Worth It for Developers?

7:17 How to Get Started with Voicebox, MCP, and Local AI Voice

Is there something similar for voice changing with local AI?

β€” @metalbroga

More User Perspectives

@

Voicebox looks fun if you want the full voice stack. For plain dictation, I keep coming back to hold-to-talk instead of always-on listening. DictaFlow nails that for me, especially when I’m bouncing between Cursor and my usual work apps.

@LisaZheng-c1b
@

Ive tried this start of year, it sucks. omnivoice studio is better.

@SnazzieTV
@

Does it support, let’s say Norwegian ?

@aidajam5
@

interesting but no pre built linux binaries?

@DaveTheAIMad
@

my issue with this is the exact same as 99% of other tts ai... they only support a tiny amount of languages, never the one i need

@alexpigeon
@

This is pretty amazing, I'll have to give this a shot. If this works as well as it seems, I may hook it up to my locally hosted private image/video gen app, FluxMotion, to power voiceovers.

@ryanlaseter7626
@

yooo voicebox contains virus after i install it

@DineshM-p5t
@

I would like to see this demo last more than 8 seconds if possible. I will give it a shot. I used piper and whisper but if I could use my own voice, that would be great

@Drewzao
@

But, but, but, can it do the GlaDOS voice ? asking for a friend.

@Pegasus-6012
@

Downloaded and tested it. The quality of the generated audio is great. There are definitely use cases for this app. However, my main critique is that TTS is way too slow (even with my RTX 4090).

@___Chris___
@

I have my agent use omnivoice as a skill to do all of this.

@JustFeral
@

Ridiculous to even compare. The only local tts tool that is on eye level with Elevenlabs currently is Demodokos Foundry - and that a commercial product.
Nothing voicebox offers is getting even close to the control and quality 11labs has.
Higgsaudio model is closer than most, but totally non commercial only.

@haka8702
@

So voice quality is sh*t .. but you have control. Control over sh*t πŸ˜‚πŸ˜‚πŸ˜‚πŸ˜‚πŸ˜‚πŸ˜‚

@SaidThaher
@

was looking for something like this ... .addresses 3 items on my PA roadmap cheers

@CloudwalkerDrones
@

I'm so confused. I am able to use this to just clone my voice for content? I don't want my computer using my voice to control anything. If others want that, that's fine. I have nothing against it, but specifically for me, if anyone can answer this, can I use this just to use as a voice model to train on real human voices and make content with it?

@mosaicmonk4380
@

It's sad how linux is always an afterthought, even for tools made specifically for developers.

@Kevin-jc1fx
@

finally, the time has arrived!!

@SpanishGosling
@

A video about AI voice in which the audio sounds like absolute dog shit. πŸ˜‚

@Humcrush
@

Does it do voice to voice like 11labs ??

@Bekah_Penguin
@

You can just use β€œsay” command in terminal on mac to use siri voice. Simple rule to tell your agent to summarize using say command after completion

@AK-jt7ug
@

vs Elevenlabs it might makes sense for native English speakers, but it starts to be a different story for other languages which are not supported or flat out unusable..

@ftamas88
@

Can it read out books or PDFs?

@PbPb-r8k
@

I tried and it's too buggy (for the moment). Unusable for me.

@NamasenITN
@

the voicebox repo seems to have been inactive for 2 months. Not sure if it's being maintained anymore.

@simonstrandgaard5503
@

LuxLLM β™₯️

@ZambeziSentinel
@

Is it safe to use? My laptop says it has malware

@Aperturecity89
@

at least do some post processing on that AI audio output.

@Web3Dre
@

Really disappointed in VoiceBox. No Linux binaries. You have to compile from source code. Instructions omit multiple dependencies. Build process never fully finished but did yield a Debian image. Final app immediately crashed upon startup on Linux Mint.
350+ reported GitHub issues and no updates for the past two months.
4 hours wasted. If they want to be the Ollama of voice, they have a LONG way to go.

@fshieldsii
@

I'm going to make some guesses here but just a little feedback. Others have mentioned your voice sounding robotic. Assumption 1, you're using resolve for editing. You have noise reduction cranked too high. It doesn't need a lot to do a good job. A lot of people feel like there can't be any background noise at all, which in all honesty, it's not the end of the world if there is some, just as long as it's not overpowering the dialog track. I suspect you had your voice recording levels a little too low. That requires increasing the track gain which raises the noise floor giving you the impression you need to crank up the noise reduction. This also could be contributed by the MIC you are using. The SM7B is popular and expensive, but it also requires you to practically have the MIC in your mouth to sound good as it has a very low sensitivity. If you can't afford one or don't want the MIC that close, a decent shotgun MIC that uses a hypercardioid pickup pattern, from that distance, will make a massive difference.


Hope that helps.

@quadcom
@

I'm gonna give this a shot just because 11 labs have been spamming me for months with ads 😭😭😭

@hectorcurioso2
@

How does the voice quality compare to chatterbox?

@mattempyre
@

thx

@Dr.Technically
@

Warning click bait! No, this is not the ollama of voice! It uses python and provides therefore the "It just works on my machine" experience. The instructions don't work and produces random error mesages.
ERROR: Could not find a version that satisfies the requirement kokoro>=0.9.4 (from versions: 0.2.1, 0.2.2, 0.2.3, 0.3.0, 0.3.1, 0.3.2, 0.3.3, 0.3.4, 0.3.5, 0.7.0, 0.7.1, 0.7.2, 0.7.3, 0.7.4, 0.7.6, 0.7.8, 0.7.9, 0.7.11, 0.7.12, 0.7.13, 0.7.14, 0.7.15, 0.7.16)
ERROR: No matching distribution found for kokoro>=0.9.4

Compare that with llama.cpp, it uses C++ therefore it doesn't matter what you want to do with it. Wanna use it with your RTX 5090? Sure, why not? Wanna use it with the GPU in the raspberry PI 5? Sure (but please don't). Because it uses C++ it just works everywhere. Python only works on super controlled could environments and is incredibly fragile. ollama (which is a llama.cpp wrapper) works, because it doesn't use python and instead uses portable languages like C++ and Go that "just do what you want them to do". Stop acting as if any project that depends on python will ever be a game changer for local usage.

@nonae9419
@

nice and clear i like your videos and helpful @TQπŸŽ‰πŸŽ‰πŸŽ‰πŸŽ‰πŸŽ‰πŸŽ‰ Waiting for more videos

@kvs7720
@

getting errors, did the same thing here

@adi5877
@

I love how you guys are super active

@10XFstories
@

It's slow af

@bjfitness4862