Skip to main content

Speech engines and models

Choose in Settings → STT → Speech Recognition Engine.

EngineWhere it runsNeedsNotes
WhisperKit (local)On your MacA model downloadPicking a model in the tutorial selects this engine. Live preview and pause line breaks work only here.
MLX Audio (local)On your Macuv installedThe default when you skip the tutorial and uv is available. The default model is Qwen3-ASR-1.7B-8bit.
Groq Cloud (fast)CloudGroq API keyAudio is sent to Groq.
Settings → STT

Choosing a WhisperKit model​

A model downloads as soon as you pick it. See progress and sizes in the Settings → Models tab.

ModelSizeCharacter
Tinyabout 150MBFastest, but hard to use for Korean.
Baseabout 290MBOften gets proper nouns and technical terms wrong.
Smallabout 970MBThe lower bound where Korean becomes usable.
Large v3 Turboabout 1.6GBDefault. Balanced speed and accuracy
Large v3 Turbo (v20240930)about 1.6GBA later snapshot of the same family
Large v3 Turbo compressedabout 632MBSmaller and faster than the default, with lower CPU use.
Large v3 quantizedabout 947MBA smaller version of Large v3
Large v3about 3.1GBMost accurate and slowest.

If transcription is slow, check Activity Monitor for other apps using the CPU before switching models. When specific terms keep coming out wrong, adding them to the Dictionary is quicker than a bigger model.

Models are stored in ~/Library/Application Support/Sogon/Sogon/Models/whisperkit.