Speech models
Pick the model. Keep the audio.
Speech models on your computer for privacy and offline work, or in the cloud with your own key. Pick the one that fits the job.
Two kinds of models. One app.
On-device models run entirely on your computer. Your audio never leaves the machine and you need no internet connection, which suits private, offline or regulated work.
Cloud models send audio straight to a provider you choose, with your own API key, for a specific speed or accuracy profile. Vowen supports both, and you choose per situation.
On-device
Private, offline transcription.
| Model | Provider | Runs on | Best for |
|---|---|---|---|
| Parakeet V2 | NVIDIA | Mac and Windows | The default. Fast, accurate English dictation, with live preview on Mac. |
| Parakeet V3 | NVIDIA | Mac and Windows | 25 European languages with auto-detect, at Parakeet speed. |
| Nemotron | NVIDIA | Mac, Apple Silicon | Streaming model: punctuated words appear as you speak. English or 31 languages. |
| Parakeet Japanese / Mandarin | NVIDIA | Mac | Dedicated models for Japanese and Mandarin dictation. |
| Whisper Large v3 / v3 Turbo | OpenAI | Mac and Windows | The widest coverage on device: 99 languages. Optional NVIDIA CUDA on Windows. |
| Whisper Tiny to Medium | OpenAI | Mac and Windows | Smaller footprints, from 78 MB, for older or lower-RAM machines. |
- Parakeet V2Mac and WindowsNVIDIA
The default. Fast, accurate English dictation, with live preview on Mac.
- Parakeet V3Mac and WindowsNVIDIA
25 European languages with auto-detect, at Parakeet speed.
- NemotronMac, Apple SiliconNVIDIA
Streaming model: punctuated words appear as you speak. English or 31 languages.
- Parakeet Japanese / MandarinMacNVIDIA
Dedicated models for Japanese and Mandarin dictation.
- Whisper Large v3 / v3 TurboMac and WindowsOpenAI
The widest coverage on device: 99 languages. Optional NVIDIA CUDA on Windows.
- Whisper Tiny to MediumMac and WindowsOpenAI
Smaller footprints, from 78 MB, for older or lower-RAM machines.
Cloud, bring your own key
Your provider, your key.
| Model | Provider | Runs on | Best for |
|---|---|---|---|
| Voxtral Mini | Mistral | Your key | Recommended cloud pick. Streams, with speaker labels. |
| Whisper Large v3 / Turbo | Groq | Your key | Very fast batch transcription. |
| Nova 3 / Nova 2 | Deepgram | Your key | Streaming with speaker labels. Nova 3 handles code-switching. |
| Scribe v2 | ElevenLabs | Your key | Streaming, high accuracy, speaker labels. |
| Universal-3.5 Pro | AssemblyAI | Your key | Streams and code-switches across 18 languages. |
| Real-Time STT | Soniox | Your key | Streaming across 99 languages. |
| gpt-4o-transcribe | OpenAI | Your key | Frontier-model accuracy for tough audio. |
| Ink 2 | Cartesia | Your key | Low-latency streaming for live dictation. |
- Voxtral MiniYour keyMistral
Recommended cloud pick. Streams, with speaker labels.
- Whisper Large v3 / TurboYour keyGroq
Very fast batch transcription.
- Nova 3 / Nova 2Your keyDeepgram
Streaming with speaker labels. Nova 3 handles code-switching.
- Scribe v2Your keyElevenLabs
Streaming, high accuracy, speaker labels.
- Universal-3.5 ProYour keyAssemblyAI
Streams and code-switches across 18 languages.
- Real-Time STTYour keySoniox
Streaming across 99 languages.
- gpt-4o-transcribeYour keyOpenAI
Frontier-model accuracy for tough audio.
- Ink 2Your keyCartesia
Low-latency streaming for live dictation.
On device by default, cloud by choice.
Vowen starts on a local model. If you never want audio to leave your computer, stay on an on-device model and you are fully offline: no account, no upload, nothing stored on a server.
Support
Common questions.
What's the difference between Parakeet and Whisper?
Do I need an internet connection?
Which model is the most accurate?
Are cloud models private?
Can I switch models per task?
One app, every model.
Run speech recognition your way, on device or in the cloud. Free forever for dictation.