Speech models

Pick the model. Keep the audio.

Speech models on your computer for privacy and offline work, or in the cloud with your own key. Pick the one that fits the job.

Two kinds of models. One app.

On-device models run entirely on your computer. Your audio never leaves the machine and you need no internet connection, which suits private, offline or regulated work.

Cloud models send audio straight to a provider you choose, with your own API key, for a specific speed or accuracy profile. Vowen supports both, and you choose per situation.

On-device

Private, offline transcription.

ModelProviderRuns onBest for
Parakeet V2NVIDIAMac and WindowsThe default. Fast, accurate English dictation, with live preview on Mac.
Parakeet V3NVIDIAMac and Windows25 European languages with auto-detect, at Parakeet speed.
NemotronNVIDIAMac, Apple SiliconStreaming model: punctuated words appear as you speak. English or 31 languages.
Parakeet Japanese / MandarinNVIDIAMacDedicated models for Japanese and Mandarin dictation.
Whisper Large v3 / v3 TurboOpenAIMac and WindowsThe widest coverage on device: 99 languages. Optional NVIDIA CUDA on Windows.
Whisper Tiny to MediumOpenAIMac and WindowsSmaller footprints, from 78 MB, for older or lower-RAM machines.
  • Parakeet V2Mac and Windows
    NVIDIA

    The default. Fast, accurate English dictation, with live preview on Mac.

  • Parakeet V3Mac and Windows
    NVIDIA

    25 European languages with auto-detect, at Parakeet speed.

  • NemotronMac, Apple Silicon
    NVIDIA

    Streaming model: punctuated words appear as you speak. English or 31 languages.

  • Parakeet Japanese / MandarinMac
    NVIDIA

    Dedicated models for Japanese and Mandarin dictation.

  • Whisper Large v3 / v3 TurboMac and Windows
    OpenAI

    The widest coverage on device: 99 languages. Optional NVIDIA CUDA on Windows.

  • Whisper Tiny to MediumMac and Windows
    OpenAI

    Smaller footprints, from 78 MB, for older or lower-RAM machines.

Cloud, bring your own key

Your provider, your key.

ModelProviderRuns onBest for
Voxtral MiniMistralYour keyRecommended cloud pick. Streams, with speaker labels.
Whisper Large v3 / TurboGroqYour keyVery fast batch transcription.
Nova 3 / Nova 2DeepgramYour keyStreaming with speaker labels. Nova 3 handles code-switching.
Scribe v2ElevenLabsYour keyStreaming, high accuracy, speaker labels.
Universal-3.5 ProAssemblyAIYour keyStreams and code-switches across 18 languages.
Real-Time STTSonioxYour keyStreaming across 99 languages.
gpt-4o-transcribeOpenAIYour keyFrontier-model accuracy for tough audio.
Ink 2CartesiaYour keyLow-latency streaming for live dictation.
  • Voxtral MiniYour key
    Mistral

    Recommended cloud pick. Streams, with speaker labels.

  • Whisper Large v3 / TurboYour key
    Groq

    Very fast batch transcription.

  • Nova 3 / Nova 2Your key
    Deepgram

    Streaming with speaker labels. Nova 3 handles code-switching.

  • Scribe v2Your key
    ElevenLabs

    Streaming, high accuracy, speaker labels.

  • Universal-3.5 ProYour key
    AssemblyAI

    Streams and code-switches across 18 languages.

  • Real-Time STTYour key
    Soniox

    Streaming across 99 languages.

  • gpt-4o-transcribeYour key
    OpenAI

    Frontier-model accuracy for tough audio.

  • Ink 2Your key
    Cartesia

    Low-latency streaming for live dictation.

Also supported:GeminiSpeechmaticsSarvamxAIOpenRouterAny OpenAI-compatible server

On device by default, cloud by choice.

Vowen starts on a local model. If you never want audio to leave your computer, stay on an on-device model and you are fully offline: no account, no upload, nothing stored on a server.

Support

Common questions.

What's the difference between Parakeet and Whisper?
Both run on your computer. Parakeet, from NVIDIA, is built for speed and is the default for live dictation. Whisper, from OpenAI, covers the most languages, 99 of them. Vowen lets you pick per task.
Do I need an internet connection?
No. With an on-device model, Vowen transcribes entirely on your machine and your audio never leaves it. Cloud models are optional, for when you want a specific provider.
Which model is the most accurate?
It depends on the language and the audio. Whisper Large v3 on device and frontier cloud models like gpt-4o-transcribe sit at the top for hard audio. For speed, Parakeet on device and Groq in the cloud are hard to beat.
Are cloud models private?
On-device models keep everything local. If you pick a cloud model, Vowen sends audio straight to that provider with your own API key, only to produce the transcript. There is no Vowen server in between. For sensitive work, stay on device.
Can I switch models per task?
Yes. Tones let each app or website use its own speech model, so Slack can use a fast local model while long documents use a bigger one.

One app, every model.

Run speech recognition your way, on device or in the cloud. Free forever for dictation.