Voxtype vs Whispering
Price, platforms and features side by side. The table says where each price was read and when.
Whispering
Open-source dictation app that uses your own models or API keys
Free
Where they differ
- Voxtype is free. Whispering is free.
- Only Whispering lists a web app.
- Only Whispering lists a Windows app.
- Voxtype also covers meeting notes and transcription.
In commonFree planmacOSLinuxOpen sourceWorks offlineLocal-firstDictation
At a glance
| Voxtype | Whispering | |
|---|---|---|
| Pricing | Free | Free |
| Starts at | Free | Free |
| Free plan | Entirely free and open source under the MIT License, with no subscription, cloud service or telemetry. The maintainer says he does not accept donations. | The app is free and open source under the AGPL-3.0 license. Local transcription costs nothing; cloud transcription is billed by the provider whose API key you add. |
| Free trial | None listed | None listed |
| Web | ||
| macOS | ||
| Windows | ||
| Linux | ||
| Made by | Peter Jackson (Faster Agile) | Epicenter |
| Price read from | voxtype.io | epicenter.so |
| Checked | Oct 6, 2026 | Oct 6, 2026 |
Who each one suits
Voxtype
- Linux users on Wayland desktops such as Hyprland, Sway or GNOME
- People comfortable setting up a tool from the terminal
- Dictating prompts into terminals, editors and coding agents
Whispering
- People who want dictation software whose code they can audit
- Users willing to add their own API keys and pay providers directly
- Offline dictation with on-device models on Mac and Linux
Worth knowing
Voxtype
- There is no Windows or mobile version; packages cover Linux and macOS only.
- Setup happens in the terminal: it runs as a daemon, and on Linux the hotkey is bound in the compositor or needs the user added to the input group.
- Prebuilt Linux packages need a recent distribution, such as Ubuntu 24.04, Debian Trixie or Fedora 40 or later.
- Grammar cleanup and meeting summaries are not built in; they rely on a separate tool such as Ollama that you install yourself.
Whispering
- Cloud transcription and AI transformations need your own provider API keys, and the provider bills the usage.
- The macOS download is for Apple Silicon only, and Windows builds include just the English-only Parakeet engine for local transcription.
- It is built for short dictation; the site suggests a dedicated recording app for long sessions.
- The browser version has no system-wide shortcuts, and the project's README on GitHub says the hosted web build is no longer being republished.
Key features
Voxtype
- Push-to-talk at the cursor
- Hold a hotkey, ScrollLock by default, speak and release, and the text is typed at the cursor, with the clipboard as a fallback. A toggle mode starts and stops with one press each.
- Nine local speech engines
- Whisper, Parakeet, Moonshine, SenseVoice, Paraformer, Dolphin, Omnilingual, Cohere Transcribe and OpenVINO Whisper are switched with one config line. Models load on first use and unload when idle.
- GPU and NPU acceleration
- Whisper runs on Vulkan across GPU vendors, the ONNX engines use CUDA 12 or 13 on NVIDIA and MIGraphX on AMD cards, and OpenVINO runs Whisper on Intel NPUs.
- Meeting mode
- Continuous transcription with automatic chunking and speaker attribution for meetings and interviews, stored locally and exported to Markdown, plain text, JSON, SRT or VTT.
- Text processing
- Spoken punctuation turns words such as "comma" into symbols, replacement tables fix common mistranscriptions, and an optional command pipes the transcript through a local LLM or a shell script.
- Linux desktop integration
- Compositor keybindings cover Hyprland, Niri, Sway, River, GNOME and KDE, with an evdev fallback for X11. It pauses MPRIS media players while recording and shows a floating waveform.
Whispering
- Shortcut-driven dictation
- Press a keyboard shortcut, speak, and the text is copied and pasted where the cursor is. The recording shortcut can be changed in settings.
- Voice-activated mode
- With voice activity detection turned on, recording starts when you begin speaking and stops after a short pause, so no key has to be held.
- Cloud providers with your own keys
- Audio goes straight from the device to Groq, OpenAI or ElevenLabs under your own API key, with no Whispering server in between. Groq also has a free tier.
- Local transcription engines
- Version 7.11.0 bundles Parakeet, Moonshine and Whisper C++ for on-device transcription, with Whisper C++ covering 65 languages. Speaches, a local transcription service, is also supported.
- AI transformations
- Chain steps that send the transcript to a model from OpenAI, Anthropic, Google Gemini, Groq or any OpenAI-compatible endpoint, or run find-and-replace, to fix grammar, translate or summarize.
- Recordings stored on the device
- Recordings, transcripts and transformation settings are kept locally. Anonymized usage analytics are collected through Aptabase and can be switched off in settings.
Plans
Voxtype
Voxtype$0
The full tool for Linux and macOS with every engine and meeting mode, installed from the AUR, a .deb or .rpm package, Nix, an AppImage or Homebrew.
Whispering
Whispering$0
The full desktop and browser app. Cloud providers bill usage to your own API key: the site lists Groq at $0.02 to $0.04 per hour of audio.

