Show HN: BetterWispr – Free, open-source dictation for Mac
130 points by kartik017 11 hours ago | 79 comments

crazygringo 6 hours ago
I think we're about to see a dictation revolution.

I've been coding with Claude, and I pretty much do it exclusively by voice now, simply because it's so much faster. And I kind of feel like Scotty from Star Trek IV, when he holds up a Mac mouse to try to talk into it.

I've been using VoiceInk [1] which looks like it's basically the same as this, but has been around for longer.

What has really made it work for me is using a Bluetooth media control like [2], where I use Karabiner Elements to remap its play/pause button to the dictation keyboard shortcut. I map rewind to option-backspace to delete the last word, fast-forward to shift-enter to insert a newline, a lower button to enter to submit my prompt, and the volume up/down buttons to scroll up/down. I've also just ordered to Xiaomi Bluetooth remote [3] that includes a microphone itself to see if I can get it to work by speaking directly into it and using it as the microphone -- there are a couple of open source projects to turn it into a Mac microphone directly. Since I'd like to be able to talk more quietly instead of into my Mac.

But what I'm REALLY waiting for is the ability to use one button for dictation, and a second button for issuing commands for a local LLM to interpret. So I hold down the voice button which transcribes "I think we need to catch that" and then the second button and go "change catch to cache, like c-a-c-h-e". Or hold down the second button and go "switch to VS code". I don't want something as finicky as macOS Voice Control, I want a local LLM I can speak naturally to.

I can very much see a future where I spend the majority of my "work" time using a Apple TV-type remote.

[1] https://github.com/Beingpax/VoiceInk

[2] https://www.amazon.com/Satechi-Bluetooth-Multimedia-Remote-C...

[3] https://www.notebookcheck.net/Xiaomi-Bluetooth-Remote-2-Pro-...

reply
polyterative 4 hours ago
I am having great results by mapping the press to dictate onto my pedal.I have been vibecoding without using my hands for the past four months with great results.It's amazing using the computer without touching it.I can go on for hours with a couple of pedals.I use them to switch windows, dictate and press enter.That's really all you need.
reply
vardalab 4 hours ago
The one thing about the Bluetooth microphones is that there is this annoying lag that it's very hard to get used to, even when using AirPods which you think would be optimized for this. It's just untenable in my opinion because I could never get used to the 3-400 millisecond delay that I had to deal with. I ended up using simple Apple wired headphones as my go-to solution eventually. They actually produce a least a modest strain on ears when being in for hours on end, at least in my situation. I do like your idea about the Bluetooth multimedia remote controller.
reply
kartik017 4 hours ago
i have fixed this problem in betterwispr btw ;)
reply
pegasus 4 hours ago
You fixed bluetooth lag??
reply
patrickk 3 hours ago
Great ideas here thanks! I’m going to see if I can whip up a Windows equivalent, using Autohokey instead of Karabiner.
reply
Grombobulous 6 hours ago
I can see the appeal from a workflow perspective, though from a “social” perspective I don’t see myself enjoying talking to myself all day.

Bonus points if you work in an office. This kind of workflow would be a nightmare.

…unless it means we get our private offices back.

reply
dnautics 4 hours ago
I have a friend who has reporposed a stenographer's mask
reply
cjonas 6 hours ago
i forked handy and built 2 modes. One for dictation and one for quick agent actions that uses the pi harness with local models and jev style classifiers so it can read screens and interact with elements (via accessiblity tree, screenshots, os scripts, bash, mcp, etc). Its nice because in agent mode you can just give it instructions on how to respond (like dictation that can read the context of the current page or input). I've been pretty happy with it so far and debating if i should release it. Just not sure if the world needs yet another vibe coded agent system.
reply
Razengan 5 hours ago
I'm glad we're getting to a point where voice can be the main UI, and that some people are actively using it to their benefit, but I hope it doesn't become THE primary interface, even with AIs.

I find speaking tiresome, somehow, and if I had to talk all the time to my computer with text input being a fidgety "accessibility" option, I'd retire to a monastery and just become a monk.

Agree with the need for different buttons for different voice contexts though.

Though I think it will become a dedicated AI key on some keyboards.. How about the Right Alt/Option on MacBooks? Does anybody use that? It could be the "AI Anywhere" input key, optionally defaulting to voice, and while F5 could continue to serve as a literal dictation key..

reply
pmoriarty 3 hours ago
Speech also has a lack of privacy, and it can annoy people nearby.

Imagine everyone on public transport speaking in to their phones to navigate and control them.

A microphone sensitive to subvocalization could maybe get around some of these problems, but it'd still be kind of weird.

reply
dnautics 4 hours ago
I upload documents to a remarkable, mark it up, and have Claude build changes there.
reply
rcarmo 6 hours ago
As much as I like seeing this, the space is really crowded, and I think the real value is not in single dictation for input, but in two other things: diarization (essential for call recordings, and a staple in video calling services, but hardly touched for in-room meetings and brainstorming) and the next step, which is cleaning it all up for actually useful notes. Not just the transcript, but better variations on meeting notes that you can customize, link to previous existing information, etc.
reply
fold_left 37 minutes ago
I've been using https://talat.app/ for this, I've not been using it for long, a few weeks, but it's working well so far.
reply
polyterative 5 hours ago
In my own vibe coded implementation during the recordings the live transcript gets sent for validation and allows me to take better decision while in the call itself.I'm using a a small local model for processing.Also like when it detects bullshit told by others in real times so I can immediately reply.Fun
reply
eajr 5 hours ago
There are a number of free apps that do this now: ghostwhisper and anarlog are the first two that come to mind. It is not hard to vibe your own too. There are open source models that do diarization.
reply
prash20026 6 hours ago
Isn't that what Google just released although it is mac only for the moment.

https://developers.google.com/edge/foresight

reply
kartik017 4 hours ago
yup this is what i have been focusing on.

you can try the notetaking feature which i have build

reply
Zizizizz 8 hours ago
This reminds me very much in terms of features as https://github.com/cjpais/Handy https://handy.computer , what's the use case for this over that?
reply
oscarteg 7 hours ago
I recently switched to Voiceink which also does similar. I'm trying to find the edge this has over that but can't find it
reply
abyssin 7 hours ago
VoiceInk took a little time compiling but it's been working flawlessly for the three people that use it now for dictating consultation notes.
reply
alexgoodhart 5 hours ago
lmao
reply
patrickk 3 hours ago
Can you link to the exact website? Multiple similarly named projects show up.
reply
evulhotdog 4 hours ago
I recently switched from Handy to Fluidvoice, and like it much better. The UI is much better and offers a better and more comprehensive feature set in comparison.
reply
kartik017 4 hours ago
i do liked fluidvoice, can't compete with them honestly but do give it a try
reply
nxpnsv 8 hours ago
They both use parakeet so I guess they would perform similar.
reply
pylotlight 7 hours ago
ye parakeet redux has been my new fav, that and tdt v2 both working well for my own clone app that was made.
reply
ramraj07 7 hours ago
How good is parakeet (assuming it’s Sota oss) compared to wisprflow? Wisprflow is magical ime
reply
handfuloflight 23 minutes ago
What about adding a wake word detection so I can just say "Hey wakeWord" to activate the STT? Then maybe detect silence to finish. Completely hands free?
reply
Royce-CMR 5 hours ago
What’s really shocking is this and many of the competitors are all using OpenAI whisper - which came out in 2022. (Once upon a time, OpenAI really did open research)

It’s surprising in 2026 this is still the default “best” all rounder. Kudos to the early OpenAI team and the open community efforts / improved models since then.

reply
kartik017 4 hours ago
no it's a really old and bad stt model i am using parakeet v3 here which is way superior
reply
sunnybeetroot 5 hours ago
Is there anything better for speech to text model?
reply
eajr 5 hours ago
Yes, parakeet. Ive switched all my apps to using it. Fantastic model
reply
sunnybeetroot 4 hours ago
That is great! Just no fast inference like cerebras or groq support it
reply
vardalab 4 hours ago
Parakeet is blindingly fast on modern hardware. My MacBook is just so fast, it's not even funny. But despite that, I still often use OpenWhisper because it just does better for accented voices. I also fine-tuned Qwen ASR on my voice, and that one is faster than Open Whisper, but I think I might even go back to Open Whisper eventually Because it still makes a lot of weird mistakes And I think it is over-indexed towards the fine tune And substitutes quite often wrong dictionary words for ambiguous words.
reply
kartik017 4 hours ago
openwhispr dx is really bad, i made this product because i was tired that no one is working on accuracy and dx in general.

betterwispr main goal is to have accuracy and best dx possible.

reply
sunnybeetroot 4 hours ago
Server side though?
reply
kartik017 4 hours ago
parakeet v3 betterwispr uses that.
reply
scosman 7 hours ago
If folks like local-AI dictation, but want it optimized for meetings recordings (like Granola/Otter/Notion) check out Biscotti. Free, local, separates voices, voice identification, AI summaries, etc.

https://github.com/scosman/Biscotti

reply
alexgoodhart 5 hours ago
Otter was really like Hey you, subscribing? Great! Two recordings per month.
reply
NJL3000 7 hours ago
This came up a month ago - I forked from something now gone - https://github.com/NickJLange/parrot

And if MacOS 26 does it right - all of this is no longer necessary for us to vibe fork/ code on weekends :-)

P.S. Handy looks slick. May be able to ditch mine

reply
esafak 6 hours ago
MacOS 27 still has room for improvement; there is a measurable starting latency and error rate.
reply
CharlesW 3 hours ago
For me (M1 Studio) it works at least as well and as quickly as Handy or GhostPepper, which are considered two of the best. (Turn on with Keyboard > Dictation, works anywhere you can type text.)
reply
sajithdilshan 2 hours ago
Even though this looks really good I don’t see myself using this.

Mostly because I could use this only when I’m working from home, and even then after a few minutes I’ll get tired of speaking to myself and I actually type way faster than I speak.

reply
xnx 6 hours ago
Seems everyone is cobbling these together. Here's a local-only one from Google: https://apps.apple.com/us/app/google-ai-edge-eloquent/id6756...
reply
maxpert 7 hours ago
Been using Handy I am pretty happy but good to see more options, specially OSS.
reply
kartik017 4 hours ago
handy is bloat and really bad dx in general.

this product is just 14mb binary.

reply
bluehatbrit 2 hours ago
I understand you made BetterWhispr and Handy would be a competitor, but if you're going to dunk on them it would be good to add some real detail and reasoning. It's okay if you just didn't get on with it, but this is HN where we want to hear about design decisions, trade-offs, and the motivation for building this over using one of the many alternatives.

I've not used either and I'm sure they're both great, but just saying an alternative is bad isn't particularly useful to anyone here.

reply
GodelNumbering 4 hours ago
Just tried it out as I was looking for something that does this exact thing well (macbook). So far I dictated about 5 different things in different contexts, it worked without a single error! Good choice to pick parakeet v3.

Nice work, thanks for building it!

reply
kartik017 4 hours ago
ooh thanks so much man, if you face any issues do open a thread on github.

https://github.com/opennookorg/betterwispr

also do star it helps.

reply
jasonjmcghee 7 hours ago
The native / built-in API SpeechAnalyzer (introduced last generation on iOS and MacOS) works quite well.

Are folks finding it lacking / needing alternatives?

reply
throw03172019 6 hours ago
Voice shortcuts etc
reply
kartik017 4 hours ago
[dead]
reply
frenchie4111 6 hours ago
This is awesome. I really want to get something like this integrated directly into my agent IDE (https://getness.dev).

The main thing I need is some way of auto-sending a "end of message" button (like a newline would do) at the end of the message so that it auto-sends the chat.

Not sure the exact right way to implement that, but it seems like something that could be general purpose. Maybe a setting per-active app?

reply
HnUser12 6 hours ago
I use handy.computer in push-to-talk mode. It inserts text in the active text box when you let go. It also supports auto-submitting at the end, so maybe that would help? I think it can send new line.
reply
rcarmo 6 hours ago
Many people are using Nemotron (NVIDIA provides Linux binaries for ARM and Intel that you can shove into a server and do audio over websockets for them).
reply
regularfry 6 hours ago
Push-to-talk?
reply
frenchie4111 6 hours ago
Yes I want push to talk in ness

But trying to implement it directly into ness hasn't worked that well. It seems like betterwispr/wisprflow/etc all seem to belong better as their own app, and we just need good integration points between the push-to-talk app and the text consuming apps (like the IDE)

reply
tamilini92 2 hours ago
[flagged]
reply
jakobov 5 hours ago
This is cool. I also built a dictation app that has advantage of using three state-of-the-art models in parallel. Which increases accuracy and reduces p95 significantly. You can check it out at https://zwhispr.com/
reply
alexgoodhart 5 hours ago
2000 words free per month.
reply
jakobov 3 hours ago
ya tokens cost money. Unfortunately there is a limit to how much I can subsidize.
reply
gekoxyz 6 hours ago
I expected dictation to be much more useful on the computer, but since I can type decently fast I find myself using it far less than I expected.
reply
crazygringo 6 hours ago
I can type crazy fast, ever since I was a kid.

Nevertheless I dictate almost exclusively with Claude Code now, because it's such better ergonomics. And it's still significantly faster.

But a big benefit is that I'm never hunching, I'm never straining my wrists, none of that stuff. Dictation is just so much better for your neck and back.

reply
gekoxyz 11 minutes ago
Yeah Claude Code is one example where I use dictation most of the time (unless I'm in the office). Even if I take back some things that I say there it can be useful context for the LLM to parse and understand what I'm doing. The thing is that for long emails, long text messages, or in general when I have to think about what I want to communicate and I have to write decently, ASR isn't helping me as much as I thought it would. Maybe I'm using it in the wrong way, and I could build a better pipeline if I added LLM postprocessing to my text, but I don't have enough memory on my computer for this kind of task yet.
reply
vardalab 4 hours ago
Yeah, this is a very good point. I just lean back and speak. I find it much better. You can actually think now better because you can actually focus on the problem, admittedly, I never bothered to learn how to type fast because I was too lazy And my fingers are beefy. And I also like it when I'm literally just looking at something and giving a feedback to model while I'm speaking and it is being transcribed and being acted upon. I mean, this is all is kind of like science fiction. I'm surprised that there's so much negativity about all this AI or as we officially now refer to as Si.
reply
Benard-dev 6 hours ago
How are you handling local transcription? We opted for an offline-first for FetchMark archiver for same privacy reason and it doubled our retention.

Congrats on shipping!

reply
kartik017 4 hours ago
[dead]
reply
kartik017 4 hours ago
just in case here is the repo: https://github.com/opennookorg/betterwispr

do star it on gh if you liked it!

reply
epaga 6 hours ago
I'm a bit confused, isn't "Wispr" a trademarked term? Why would you pick a name that could easily be shut down?
reply
kartik017 4 hours ago
nope it isn't trademarked actually
reply
handfuloflight 4 hours ago
U.S. trademark rights are conferred just by use in the market, not necessarily requiring an application.
reply
Na6z 7 hours ago
This is pretty neat, I built something similar a couple months ago. I'd say try to support Windows or Linux next.
reply
kartik017 4 hours ago
sure will do
reply
aucisson_masque 8 hours ago
How does it compare to handy ?
reply
kartik017 4 hours ago
[dead]
reply
starik36 2 hours ago
This is my entry into the world of vibe coded Wispr Flow clones. Runs locally on the Whisper AI speech recognition model file - the small variety - just under 200mb.

I thought the small model won't be any good, but it's actually great and barely causes the GPU any usage.

Like other people in this thread mentioned, I switched pretty much exclusively to working with voice. Whether it's responding to text or talking to Claude or writing an email.

https://imgur.com/YJyz05r

reply
honkycat 5 hours ago
I've been trying to figure out how to code and build while on the bike /treadmill at the gym.

Want a remote terminal into my Mac that I can voice control.

reply
kartik017 4 hours ago
yeah you can do try this out, it's really smooth and has good native macos ux in general.

it's just a 14 mb file so no bloat as well.

reply
ProofHouse 6 hours ago
To be honest, I used Wspr so much I only recently started realizing just HOW bad it gets trying to 'improve' your diction, especially coding. I have seen totally opperate directives. I had no idea it was potentially occuring, but it gets past a point where it is WAY to confident in guessing what you meant, sometimes to the extent of changing DON'T to DO, for example.
reply
kartik017 4 hours ago
agree with your point, but wisprflow has been degraded recently and i wanted a solution which is free and works basically
reply
null-phnix 6 hours ago
[flagged]
reply
stevyhacker 3 hours ago
[dead]
reply
stevyhacker 2 hours ago
[dead]
reply
467593457091 2 hours ago
[dead]
reply
kartik017 11 hours ago
Hey HN,

I’m building BetterWispr, a free and open-source voice dictation app for macOS.

I wanted to make voice typing more accessible without requiring a monthly subscription or sending every recording to a cloud server.

You can hold Option + Space, speak naturally, and have the transcription inserted directly into whatever app you’re using.

A few things I’ve been working on:

- Local speech recognition using Whisper, Parakeet, or Apple Speech - Automatic cleanup of filler words and repeated phrases - Custom vocabulary for names and technical terms - Different writing styles depending on the app - Optional meeting transcription and summaries

The project is open source under Apache 2.0.

I’m still improving the experience, especially around transcription accuracy, speed, and reliability across different Macs.

I’d love feedback from people who regularly use voice dictation.

What would make you switch from your current dictation tool to an open-source alternative?

Website: https://betterwispr.com/

reply
bazmattaz 3 hours ago
What are the minimum Mac requirements?
reply