Skip to content

Handy vs Voxtype

Price, platforms and features side by side. The table says where each price was read and when.

Handy

Free, open-source offline speech-to-text for desktop

Free

Voxtype

Open-source push-to-talk dictation for Linux and macOS

Free

Where they differ

  • Handy is free. Voxtype is free.
  • Only Handy lists a Windows app.
  • Voxtype also covers meeting notes and transcription.
In commonFree planmacOSLinuxOpen sourceWorks offlineLocal-firstDictation

At a glance

HandyVoxtype
PricingFreeFree
Starts atFreeFree
Free planEntirely free and open source under the MIT License, with no subscription or cloud service. The project is supported by sponsors and optional donations.Entirely free and open source under the MIT License, with no subscription, cloud service or telemetry. The maintainer says he does not accept donations.
Free trialNone listedNone listed
macOS
Windows
Linux
Made byCJ PaisPeter Jackson (Faster Agile)
Price read fromhandy.computervoxtype.io
CheckedOct 5, 2026Oct 6, 2026

Who each one suits

Handy

  • People who want free dictation that never uploads audio
  • Dictating on Linux as well as macOS and Windows
  • Developers who want a small codebase to fork and extend

Voxtype

  • Linux users on Wayland desktops such as Hyprland, Sway or GNOME
  • People comfortable setting up a tool from the terminal
  • Dictating prompts into terminals, editors and coding agents

Worth knowing

Handy

  • It transcribes microphone input only and cannot capture system audio from calls or videos.
  • There is no live streaming transcription: Handy records a segment, then transcribes it, and recordings work best under 3 to 5 minutes.
  • Spoken punctuation commands are not supported, and AI post-processing is an alpha feature that needs your own API key or a local model.
  • There are no mobile apps; downloads cover macOS, Windows and Linux.

Voxtype

  • There is no Windows or mobile version; packages cover Linux and macOS only.
  • Setup happens in the terminal: it runs as a daemon, and on Linux the hotkey is bound in the compositor or needs the user added to the input group.
  • Prebuilt Linux packages need a recent distribution, such as Ubuntu 24.04, Debian Trixie or Fedora 40 or later.
  • Grammar cleanup and meeting summaries are not built in; they rely on a separate tool such as Ollama that you install yourself.

Key features

Handy

Shortcut-driven dictation
Press a shortcut to start and stop recording, or hold it for push-to-talk, and the text is pasted into the app you are typing in. Shortcuts can be remapped in settings.
Local speech models
Choose from Parakeet, Whisper, Moonshine and other models that run on the device, with GPU acceleration when available. Parakeet V3 covers 25 European languages and Whisper 99+.
Optional AI post-processing
An experimental feature runs the transcript through a language model to fix grammar or reformat it, using a cloud provider key, an OpenAI-compatible local endpoint or Apple Intelligence.
Custom words
List names and jargon that are often misheard, and Handy corrects similar-sounding words to match. The docs describe this matching as imperfect.
Transcription history
Past transcriptions are kept with a timestamp and the original audio, and entries can be copied, starred or deleted. A limit setting controls how many are stored.
Command-line control
Flags such as --toggle-transcription and --cancel control a running instance, and --start-hidden launches Handy straight to the system tray.

Voxtype

Push-to-talk at the cursor
Hold a hotkey, ScrollLock by default, speak and release, and the text is typed at the cursor, with the clipboard as a fallback. A toggle mode starts and stops with one press each.
Nine local speech engines
Whisper, Parakeet, Moonshine, SenseVoice, Paraformer, Dolphin, Omnilingual, Cohere Transcribe and OpenVINO Whisper are switched with one config line. Models load on first use and unload when idle.
GPU and NPU acceleration
Whisper runs on Vulkan across GPU vendors, the ONNX engines use CUDA 12 or 13 on NVIDIA and MIGraphX on AMD cards, and OpenVINO runs Whisper on Intel NPUs.
Meeting mode
Continuous transcription with automatic chunking and speaker attribution for meetings and interviews, stored locally and exported to Markdown, plain text, JSON, SRT or VTT.
Text processing
Spoken punctuation turns words such as "comma" into symbols, replacement tables fix common mistranscriptions, and an optional command pipes the transcript through a local LLM or a shell script.
Linux desktop integration
Compositor keybindings cover Hyprland, Niri, Sway, River, GNOME and KDE, with an evdev fallback for X11. It pauses MPRIS media players while recording and shows a floating waveform.

Plans

Handy

  • Handy$0

    The full app for macOS, Windows and Linux with every local model. Donations and GitHub Sponsors are optional.

Voxtype

  • Voxtype$0

    The full tool for Linux and macOS with every engine and meeting mode, installed from the AUR, a .deb or .rpm package, Nix, an AppImage or Homebrew.

More comparisons