Offline dictation, tuned for your Mac.
On an M5 Max, Parakeet v3 Ultra runs 1.9× faster than on stock MLX, on 28% less energy, with no measurable accuracy loss.
curl -fsSL https://raw.githubusercontent.com/TobyNoSkillSon/Vella/main/scripts/install-public.sh | bash
Free and open source · MIT · Apple Silicon, macOS 26+
Smaller weights should be faster. On stock MLX they weren't.
- 16-bitWER 15.46% → 15.51%261×507×1.9×28% less energy
- 8-bitWER 15.46% → 15.46%245×515×2.1×41% less energy
- 4-bitWER 15.62% → 15.58%247×522×2.1×41% less energy
- Feed-forward 64%
Unpacks the quantized weights before multiplying. At 16-bit, its tiles leave GPU cores idle.
Feeds 8- and 4-bit codes straight into the M5's tensor unit. At 16-bit, smaller split-K tiles expose more parallel work.
- Attention 20%
Three separate projections, for Q, K and V.
One fused projection for all three.
- Convolution 12%
Gate, convolution, batch norm and SiLU, one after another.
Batch norm folded into the weights; the rest is one Metal kernel.
- Decoder + joint 3%
Decodes one step at a time.
32 steps in one compiled graph, sized to the frames left. Kept at 16-bit.
Every model, stock MLX then Vella.
- Nemotron 3.5 Streaming16-bit7.3×34.9×4.8×73% less energy
- Parakeet v3 Ultra16-bit261×507×1.9×28% less energy
- Parakeet v316-bit258×494×1.9×26% less energy
- Whisper large-v3 turbo16-bit83.9×116×1.4×4% less energy
- Qwen3 ASR 0.6B16-bit51.7×64.4×1.2×7% less energy
- Whisper large-v316-bit29.5×34.9×1.2×5% less energy
- Qwen3 ASR 1.7B16-bit26.6×29.6×1.1×10% less energy
The same energy transcribes up to 3.8× more speech.
- Nemotron 3.5 Streaming16-bit31 h119 h
- Parakeet v3 Ultra16-bit948 h1,309 h
- Parakeet v316-bit945 h1,275 h
- Qwen3 ASR 1.7B16-bit74 h82 h
- Qwen3 ASR 0.6B16-bit162 h176 h
- Whisper large-v316-bit69 h72 h
- Whisper large-v3 turbo16-bit156 h164 h
Every number
| Model | Precision | Path | English WER | Speed | J / audio min | Peak RAM MiB |
|---|---|---|---|---|---|---|
| Parakeet v3 Ultra | 16-bit | Standard | 15.46% | 261.4× | 6.33 | 1,796 |
| Parakeet v3 Ultra | 16-bit | Optimized Exact | 15.49% | 474.5× | 4.71 | 1,808 |
| Parakeet v3 Ultra | 16-bit | Optimized Fast | 15.51% | 507.1× | 4.58 | 1,792 |
| Parakeet v3 Ultra | 8-bit | Standard | 15.46% | 244.8× | 8.99 | 1,280 |
| Parakeet v3 Ultra | 8-bit | Optimized Exact | 15.46% | 369.0× | 8.62 | 1,888 |
| Parakeet v3 Ultra | 8-bit | Optimized Fast | 15.46% | 515.0× | 5.35 | 1,352 |
| Parakeet v3 Ultra | 4-bit | Standard | 15.62% | 246.7× | 8.75 | 1,024 |
| Parakeet v3 Ultra | 4-bit | Optimized Exact | 15.65% | 370.6× | 8.41 | 1,640 |
| Parakeet v3 Ultra | 4-bit | Optimized Fast | 15.58% | 522.5× | 5.13 | 1,094 |
| Parakeet v3 | 16-bit | Standard | 16.37% | 257.7× | 6.35 | 1,794 |
| Parakeet v3 | 16-bit | Optimized Exact | 16.42% | 463.9× | 4.81 | 1,777 |
| Parakeet v3 | 16-bit | Optimized Fast | 16.42% | 494.3× | 4.70 | 1,765 |
| Parakeet v3 | 8-bit | Standard | 16.40% | 241.2× | 9.05 | 1,279 |
| Parakeet v3 | 8-bit | Optimized Exact | 16.35% | 363.3× | 8.68 | 1,883 |
| Parakeet v3 | 8-bit | Optimized Fast | 16.30% | 504.4× | 5.41 | 1,342 |
| Parakeet v3 | 4-bit | Standard | 17.02% | 242.3× | 8.84 | 1,024 |
| Parakeet v3 | 4-bit | Optimized Exact | 16.90% | 365.0× | 8.49 | 1,628 |
| Parakeet v3 | 4-bit | Optimized Fast | 16.95% | 512.3× | 5.20 | 1,088 |
| Qwen3 ASR 1.7B | 16-bit | Standard | 15.00% | 26.6× | 81.03 | 4,624 |
| Qwen3 ASR 1.7B | 16-bit | Optimized Exact | 15.00% | 29.6× | 72.73 | 5,118 |
| Qwen3 ASR 1.7B | 16-bit | Optimized Fast = Exact | 15.00% | 29.6× | 72.73 | 5,118 |
| Qwen3 ASR 1.7B | 8-bit | Standard | 15.11% | 36.7× | 71.35 | 3,197 |
| Qwen3 ASR 1.7B | 8-bit | Optimized Exact | 15.11% | 44.3× | 64.81 | 3,721 |
| Qwen3 ASR 1.7B | 8-bit | Optimized Fast = Exact | 15.11% | 44.3× | 64.81 | 3,721 |
| Qwen3 ASR 1.7B | 4-bit | Standard | 18.32% | 47.3× | 56.23 | 2,367 |
| Qwen3 ASR 1.7B | 4-bit | Optimized Exact | 18.32% | 61.5× | 51.43 | 2,890 |
| Qwen3 ASR 1.7B | 4-bit | Optimized Fast = Exact | 18.32% | 61.5× | 51.43 | 2,890 |
| Qwen3 ASR 0.6B | 16-bit | Standard | 15.89% | 51.7× | 36.82 | 2,142 |
| Qwen3 ASR 0.6B | 16-bit | Optimized Exact | 15.89% | 64.4× | 34.07 | 2,406 |
| Qwen3 ASR 0.6B | 16-bit | Optimized Fast = Exact | 15.89% | 64.4× | 34.07 | 2,406 |
| Qwen3 ASR 0.6B | 8-bit | Standard | 16.04% | 60.7× | 33.58 | 1,639 |
| Qwen3 ASR 0.6B | 8-bit | Optimized Exact | 16.04% | 82.0× | 30.05 | 1,920 |
| Qwen3 ASR 0.6B | 8-bit | Optimized Fast = Exact | 16.04% | 82.0× | 30.05 | 1,920 |
| Qwen3 ASR 0.6B | 4-bit | Standard | 17.56% | 68.4× | 28.71 | 1,332 |
| Qwen3 ASR 0.6B | 4-bit | Optimized Exact | 17.56% | 96.4× | 25.60 | 1,648 |
| Qwen3 ASR 0.6B | 4-bit | Optimized Fast = Exact | 17.56% | 96.4× | 25.60 | 1,648 |
| Whisper large-v3 | 16-bit | Standard | 17.06% | 29.5× | 86.34 | 3,923 |
| Whisper large-v3 | 16-bit | Optimized Exact | 17.06% | 34.9× | 82.41 | 3,916 |
| Whisper large-v3 | 16-bit | Optimized Fast = Exact | 17.06% | 34.9× | 82.41 | 3,916 |
| Whisper large-v3 | 8-bit | Standard | 17.27% | 34.0× | 80.65 | 3,141 |
| Whisper large-v3 | 8-bit | Optimized Exact | 17.27% | 43.5× | 73.70 | 3,116 |
| Whisper large-v3 | 8-bit | Optimized Fast = Exact | 17.27% | 43.5× | 73.70 | 3,116 |
| Whisper large-v3 | 4-bit | Standard | 17.10% | 40.4× | 73.16 | 2,102 |
| Whisper large-v3 | 4-bit | Optimized Exact | 17.10% | 54.4× | 65.96 | 2,097 |
| Whisper large-v3 | 4-bit | Optimized Fast = Exact | 17.10% | 54.4× | 65.96 | 2,097 |
| Whisper large-v3 turbo | 16-bit | Standard | 16.57% | 83.9× | 38.22 | 2,499 |
| Whisper large-v3 turbo | 16-bit | Optimized Exact | 16.57% | 115.7× | 36.56 | 2,516 |
| Whisper large-v3 turbo | 16-bit | Optimized Fast = Exact | 16.57% | 115.7× | 36.56 | 2,516 |
| Whisper large-v3 turbo | 8-bit | Standard | 16.52% | 91.5× | 37.06 | 2,419 |
| Whisper large-v3 turbo | 8-bit | Optimized Exact | 16.52% | 129.8× | 34.64 | 2,363 |
| Whisper large-v3 turbo | 8-bit | Optimized Fast = Exact | 16.52% | 129.8× | 34.64 | 2,363 |
| Whisper large-v3 turbo | 4-bit | Standard | 16.96% | 93.8× | 39.81 | 1,723 |
| Whisper large-v3 turbo | 4-bit | Optimized Exact | 16.96% | 134.6× | 37.27 | 1,731 |
| Whisper large-v3 turbo | 4-bit | Optimized Fast = Exact | 16.96% | 134.6× | 37.27 | 1,731 |
| Nemotron 3.5 Streaming | 16-bit | Standard | 23.39% | 7.3× | 188.25 | 2,253 |
| Nemotron 3.5 Streaming | 16-bit | Optimized Exact | 23.39% | 19.3× | 80.33 | 2,725 |
| Nemotron 3.5 Streaming | 16-bit | Optimized Fast | 23.35% | 34.9× | 50.03 | 1,655 |
| Nemotron 3.5 Streaming | 8-bit | Standard | 23.42% | 15.1× | 112.00 | 1,191 |
| Nemotron 3.5 Streaming | 8-bit | Optimized Exact | 23.42% | 29.8× | 58.71 | 1,251 |
| Nemotron 3.5 Streaming | 8-bit | Optimized Fast | 23.44% | 38.4× | 43.07 | 1,114 |
| Nemotron 3.5 Streaming | 4-bit | Standard | 32.78% | 15.3× | 110.17 | 927 |
| Nemotron 3.5 Streaming | 4-bit | Optimized Exact | 32.78% | 30.4× | 55.57 | 980 |
| Nemotron 3.5 Streaming | 4-bit | Optimized Fast | 32.80% | 39.7× | 39.77 | 847 |
| ElevenLabs Scribe v2 | Cloud | Estimate | 13.4% (11.8–13.9) | accuracy only | ||
| Microsoft Azure Speech | Cloud | Estimate | 12.9% (11.3–13.3) | accuracy only | ||
Measured 1–3 October 2026 on an Apple M5 Max (40-core GPU), macOS 26.6. English word error rate on the 167 English minutes of the 239.7-minute v2 suite; nine other languages are scored separately. Speed, energy and peak memory on its 22.5-minute quick subset with the model loaded: speed and energy are the median of three runs, memory the peak. Speed is audio seconds per second of work. Energy is CPU, GPU, Neural Engine and memory, net of idle; the screen and the rest of the Mac are not counted, so hours per 100 Wh are not battery life. Standard is stock MLX; Optimized adds Vella's custom kernels, as Exact and Fast recipes, and “Fast = Exact” marks a model whose Fast runs the Exact recipe. M5, M5 Pro and M5 Max have the GPU Neural Accelerators Vella targets. Vella checks its optimized paths against stock MLX and retains a stock fallback. Our performance measurements are from an M5 Max with a 40-core GPU; other M5 configurations have not yet been tested. M1–M4 Macs run the self-tested stock fallback paths; their performance has not been measured. Cloud rows are estimates scaled from the Hugging Face Open ASR Leaderboard; no audio was sent to them. Every figure in benchmarks.json
Built for you and your agent.
$ vella transcribe talk.m4a
The transcript, as plain text.
$ vella transcribe talk.m4a --srt
$ vella url
http://127.0.0.1:63080/v1
Tell me about Vella and what you think of it: github.com/TobyNoSkillSon/Vella
from openai import OpenAI
import subprocess
client = OpenAI(
base_url=subprocess.getoutput("vella url"),
api_key="local")
text = client.audio.transcriptions.create(
model="whisper-1",
file=open("talk.m4a", "rb"))
OpenAI-compatible /v1/audio/transcriptions, on this Mac only. The transcribe skill
On a different Mac? Your agent can benchmark Vella there and, with your OK, send the numbers back. Benchmark kit
Vella never uploads what you say.
Recognition runs with no network accessNo telemetryDownloads only when you ask
Vella is free. Help keep it that way.