Skip to content

Voxtype vs Whispering

Price, platforms and features side by side. The table says where each price was read and when.

Voxtype

Open-source push-to-talk dictation for Linux and macOS

Free

Whispering

Open-source dictation app that uses your own models or API keys

Free

Where they differ

  • Voxtype is free. Whispering is free.
  • Only Whispering lists a web app.
  • Only Whispering lists a Windows app.
  • Voxtype also covers meeting notes and transcription.
In commonFree planmacOSLinuxOpen sourceWorks offlineLocal-firstDictation

At a glance

VoxtypeWhispering
PricingFreeFree
Starts atFreeFree
Free planEntirely free and open source under the MIT License, with no subscription, cloud service or telemetry. The maintainer says he does not accept donations.The app is free and open source under the AGPL-3.0 license. Local transcription costs nothing; cloud transcription is billed by the provider whose API key you add.
Free trialNone listedNone listed
Web
macOS
Windows
Linux
Made byPeter Jackson (Faster Agile)Epicenter
Price read fromvoxtype.ioepicenter.so
CheckedOct 6, 2026Oct 6, 2026

Who each one suits

Voxtype

  • Linux users on Wayland desktops such as Hyprland, Sway or GNOME
  • People comfortable setting up a tool from the terminal
  • Dictating prompts into terminals, editors and coding agents

Whispering

  • People who want dictation software whose code they can audit
  • Users willing to add their own API keys and pay providers directly
  • Offline dictation with on-device models on Mac and Linux

Worth knowing

Voxtype

  • There is no Windows or mobile version; packages cover Linux and macOS only.
  • Setup happens in the terminal: it runs as a daemon, and on Linux the hotkey is bound in the compositor or needs the user added to the input group.
  • Prebuilt Linux packages need a recent distribution, such as Ubuntu 24.04, Debian Trixie or Fedora 40 or later.
  • Grammar cleanup and meeting summaries are not built in; they rely on a separate tool such as Ollama that you install yourself.

Whispering

  • Cloud transcription and AI transformations need your own provider API keys, and the provider bills the usage.
  • The macOS download is for Apple Silicon only, and Windows builds include just the English-only Parakeet engine for local transcription.
  • It is built for short dictation; the site suggests a dedicated recording app for long sessions.
  • The browser version has no system-wide shortcuts, and the project's README on GitHub says the hosted web build is no longer being republished.

Key features

Voxtype

Push-to-talk at the cursor
Hold a hotkey, ScrollLock by default, speak and release, and the text is typed at the cursor, with the clipboard as a fallback. A toggle mode starts and stops with one press each.
Nine local speech engines
Whisper, Parakeet, Moonshine, SenseVoice, Paraformer, Dolphin, Omnilingual, Cohere Transcribe and OpenVINO Whisper are switched with one config line. Models load on first use and unload when idle.
GPU and NPU acceleration
Whisper runs on Vulkan across GPU vendors, the ONNX engines use CUDA 12 or 13 on NVIDIA and MIGraphX on AMD cards, and OpenVINO runs Whisper on Intel NPUs.
Meeting mode
Continuous transcription with automatic chunking and speaker attribution for meetings and interviews, stored locally and exported to Markdown, plain text, JSON, SRT or VTT.
Text processing
Spoken punctuation turns words such as "comma" into symbols, replacement tables fix common mistranscriptions, and an optional command pipes the transcript through a local LLM or a shell script.
Linux desktop integration
Compositor keybindings cover Hyprland, Niri, Sway, River, GNOME and KDE, with an evdev fallback for X11. It pauses MPRIS media players while recording and shows a floating waveform.

Whispering

Shortcut-driven dictation
Press a keyboard shortcut, speak, and the text is copied and pasted where the cursor is. The recording shortcut can be changed in settings.
Voice-activated mode
With voice activity detection turned on, recording starts when you begin speaking and stops after a short pause, so no key has to be held.
Cloud providers with your own keys
Audio goes straight from the device to Groq, OpenAI or ElevenLabs under your own API key, with no Whispering server in between. Groq also has a free tier.
Local transcription engines
Version 7.11.0 bundles Parakeet, Moonshine and Whisper C++ for on-device transcription, with Whisper C++ covering 65 languages. Speaches, a local transcription service, is also supported.
AI transformations
Chain steps that send the transcript to a model from OpenAI, Anthropic, Google Gemini, Groq or any OpenAI-compatible endpoint, or run find-and-replace, to fix grammar, translate or summarize.
Recordings stored on the device
Recordings, transcripts and transformation settings are kept locally. Anonymized usage analytics are collected through Aptabase and can be switched off in settings.

Plans

Voxtype

  • Voxtype$0

    The full tool for Linux and macOS with every engine and meeting mode, installed from the AUR, a .deb or .rpm package, Nix, an AppImage or Homebrew.

Whispering

  • Whispering$0

    The full desktop and browser app. Cloud providers bill usage to your own API key: the site lists Groq at $0.02 to $0.04 per hour of audio.

More comparisons