Handy vs Voxtype
Price, platforms and features side by side. The table says where each price was read and when.
Where they differ
- Handy is free. Voxtype is free.
- Only Handy lists a Windows app.
- Voxtype also covers meeting notes and transcription.
In commonFree planmacOSLinuxOpen sourceWorks offlineLocal-firstDictation
At a glance
| Handy | Voxtype | |
|---|---|---|
| Pricing | Free | Free |
| Starts at | Free | Free |
| Free plan | Entirely free and open source under the MIT License, with no subscription or cloud service. The project is supported by sponsors and optional donations. | Entirely free and open source under the MIT License, with no subscription, cloud service or telemetry. The maintainer says he does not accept donations. |
| Free trial | None listed | None listed |
| macOS | ||
| Windows | ||
| Linux | ||
| Made by | CJ Pais | Peter Jackson (Faster Agile) |
| Price read from | handy.computer | voxtype.io |
| Checked | Oct 5, 2026 | Oct 6, 2026 |
Who each one suits
Handy
- People who want free dictation that never uploads audio
- Dictating on Linux as well as macOS and Windows
- Developers who want a small codebase to fork and extend
Voxtype
- Linux users on Wayland desktops such as Hyprland, Sway or GNOME
- People comfortable setting up a tool from the terminal
- Dictating prompts into terminals, editors and coding agents
Worth knowing
Handy
- It transcribes microphone input only and cannot capture system audio from calls or videos.
- There is no live streaming transcription: Handy records a segment, then transcribes it, and recordings work best under 3 to 5 minutes.
- Spoken punctuation commands are not supported, and AI post-processing is an alpha feature that needs your own API key or a local model.
- There are no mobile apps; downloads cover macOS, Windows and Linux.
Voxtype
- There is no Windows or mobile version; packages cover Linux and macOS only.
- Setup happens in the terminal: it runs as a daemon, and on Linux the hotkey is bound in the compositor or needs the user added to the input group.
- Prebuilt Linux packages need a recent distribution, such as Ubuntu 24.04, Debian Trixie or Fedora 40 or later.
- Grammar cleanup and meeting summaries are not built in; they rely on a separate tool such as Ollama that you install yourself.
Key features
Handy
- Shortcut-driven dictation
- Press a shortcut to start and stop recording, or hold it for push-to-talk, and the text is pasted into the app you are typing in. Shortcuts can be remapped in settings.
- Local speech models
- Choose from Parakeet, Whisper, Moonshine and other models that run on the device, with GPU acceleration when available. Parakeet V3 covers 25 European languages and Whisper 99+.
- Optional AI post-processing
- An experimental feature runs the transcript through a language model to fix grammar or reformat it, using a cloud provider key, an OpenAI-compatible local endpoint or Apple Intelligence.
- Custom words
- List names and jargon that are often misheard, and Handy corrects similar-sounding words to match. The docs describe this matching as imperfect.
- Transcription history
- Past transcriptions are kept with a timestamp and the original audio, and entries can be copied, starred or deleted. A limit setting controls how many are stored.
- Command-line control
- Flags such as --toggle-transcription and --cancel control a running instance, and --start-hidden launches Handy straight to the system tray.
Voxtype
- Push-to-talk at the cursor
- Hold a hotkey, ScrollLock by default, speak and release, and the text is typed at the cursor, with the clipboard as a fallback. A toggle mode starts and stops with one press each.
- Nine local speech engines
- Whisper, Parakeet, Moonshine, SenseVoice, Paraformer, Dolphin, Omnilingual, Cohere Transcribe and OpenVINO Whisper are switched with one config line. Models load on first use and unload when idle.
- GPU and NPU acceleration
- Whisper runs on Vulkan across GPU vendors, the ONNX engines use CUDA 12 or 13 on NVIDIA and MIGraphX on AMD cards, and OpenVINO runs Whisper on Intel NPUs.
- Meeting mode
- Continuous transcription with automatic chunking and speaker attribution for meetings and interviews, stored locally and exported to Markdown, plain text, JSON, SRT or VTT.
- Text processing
- Spoken punctuation turns words such as "comma" into symbols, replacement tables fix common mistranscriptions, and an optional command pipes the transcript through a local LLM or a shell script.
- Linux desktop integration
- Compositor keybindings cover Hyprland, Niri, Sway, River, GNOME and KDE, with an evdev fallback for X11. It pauses MPRIS media players while recording and shows a floating waveform.
Plans
Handy
Handy$0
The full app for macOS, Windows and Linux with every local model. Donations and GitHub Sponsors are optional.
Voxtype
Voxtype$0
The full tool for Linux and macOS with every engine and meeting mode, installed from the AUR, a .deb or .rpm package, Nix, an AppImage or Homebrew.
