# Vella: offline dictation, tuned for your Mac

Vella 2.0.1 is a free, open-source (MIT) menu-bar app for Apple Silicon Macs (macOS 26 or newer). Press Control + Command + N, speak, press it again: a local speech-to-text model transcribes and inserts the text where you were typing. Streaming mode types as you speak. Vella never uploads your audio or transcripts.

On an M5 Max, Parakeet v3 Ultra runs 1.9× faster than on stock MLX, on 28% less energy, with no measurable accuracy loss (16-bit, Optimized Fast, the default).

- Source: https://github.com/TobyNoSkillSon/Vella
- Install: `curl -fsSL https://raw.githubusercontent.com/TobyNoSkillSon/Vella/main/scripts/install-public.sh | bash`
- For coding agents: the `transcribe` skill at https://tobynoskillson.com/vella/.well-known/agent-skills/transcribe/SKILL.md (offline transcription from the `vella` command or an OpenAI-compatible local API)
- Support the developer: https://github.com/sponsors/TobyNoSkillSon

## Smaller weights should be faster. On stock MLX they weren't.

Parakeet v3 Ultra, Apple M5 Max. Stock MLX is Vella's Standard path: the same build with its custom kernels switched off. Speed is audio seconds per second of work.

| Weights | Stock MLX | Vella | Faster | Less energy | English WER stock → Vella |
|---|---|---|---|---|---|
| 16-bit | 261.4× | 507.1× | 1.9× | 28% | 15.46% → 15.51% |
| 8-bit | 244.8× | 515.0× | 2.1× | 41% | 15.46% → 15.46% |
| 4-bit | 246.7× | 522.5× | 2.1× | 41% | 15.62% → 15.58% |

Where the weights are, and what changed (Parakeet v3 Ultra, 627M weights):

| Part | Share of weights | Stock MLX | Vella |
|---|---|---|---|
| Feed-forward (48 blocks) | 64% | Unpacks the quantized weights before multiplying. At 16-bit, its tiles leave GPU cores idle. | Feeds 8- and 4-bit codes straight into the M5's tensor unit. At 16-bit, smaller split-K tiles expose more parallel work. |
| Attention (24 layers) | 20% | Three separate projections, for Q, K and V. | One fused projection for all three. |
| Convolution (24 layers) | 12% | Gate, convolution, batch norm and SiLU, one after another. | Batch norm folded into the weights; the rest is one Metal kernel. |
| Decoder + joint | 3% | Decodes one step at a time. | 32 steps in one compiled graph, sized to the frames left. Kept at 16-bit. |

Kernel source: https://github.com/TobyNoSkillSon/Vella/tree/main/Worker/Sources/SmallMGEMM

## Every model, stock MLX then Vella

Each model at its native 16-bit weights, Optimized Fast against Standard; speed in audio seconds per second of work.

| Model | Stock MLX | Vella | Faster | Less energy | English WER stock → Vella |
|---|---|---|---|---|---|
| Nemotron 3.5 Streaming | 7.3× | 34.9× | 4.8× | 73% | 23.39% → 23.35% |
| Parakeet v3 Ultra | 261.4× | 507.1× | 1.9× | 28% | 15.46% → 15.51% |
| Parakeet v3 | 257.7× | 494.3× | 1.9× | 26% | 16.37% → 16.42% |
| Whisper large-v3 turbo | 83.9× | 115.7× | 1.4× | 4% | 16.57% → 16.57% |
| Qwen3 ASR 0.6B | 51.7× | 64.4× | 1.2× | 7% | 15.89% → 15.89% |
| Whisper large-v3 | 29.5× | 34.9× | 1.2× | 5% | 17.06% → 17.06% |
| Qwen3 ASR 1.7B | 26.6× | 29.6× | 1.1× | 10% | 15.00% → 15.00% |

## Energy

Hours of speech transcribed per 100 Wh of chip energy. Not battery life: the screen and the rest of the Mac are not counted.

| Model | J per audio minute, stock → Vella | Stock MLX | Vella | More speech |
|---|---|---|---|---|
| Nemotron 3.5 Streaming | 188.25 → 50.03 | 31 h | 119 h | 3.8× |
| Parakeet v3 Ultra | 6.33 → 4.58 | 948 h | 1,309 h | 1.4× |
| Parakeet v3 | 6.35 → 4.70 | 945 h | 1,275 h | 1.3× |
| Qwen3 ASR 1.7B | 81.03 → 72.73 | 74 h | 82 h | 1.1× |
| Qwen3 ASR 0.6B | 36.82 → 34.07 | 162 h | 176 h | 1.1× |
| Whisper large-v3 | 86.34 → 82.41 | 69 h | 72 h | 1.0× |
| Whisper large-v3 turbo | 38.22 → 36.56 | 156 h | 164 h | 1.0× |

## Every number

| Model | Precision | Path | English WER | Speed | J / audio min | Peak RAM MiB |
|---|---|---|---|---|---|---|
| Parakeet v3 Ultra | 16-bit | Standard | 15.46% | 261.4× | 6.33 | 1,796 |
| Parakeet v3 Ultra | 16-bit | Optimized Exact | 15.49% | 474.5× | 4.71 | 1,808 |
| Parakeet v3 Ultra | 16-bit | Optimized Fast | 15.51% | 507.1× | 4.58 | 1,792 |
| Parakeet v3 Ultra | 8-bit | Standard | 15.46% | 244.8× | 8.99 | 1,280 |
| Parakeet v3 Ultra | 8-bit | Optimized Exact | 15.46% | 369.0× | 8.62 | 1,888 |
| Parakeet v3 Ultra | 8-bit | Optimized Fast | 15.46% | 515.0× | 5.35 | 1,352 |
| Parakeet v3 Ultra | 4-bit | Standard | 15.62% | 246.7× | 8.75 | 1,024 |
| Parakeet v3 Ultra | 4-bit | Optimized Exact | 15.65% | 370.6× | 8.41 | 1,640 |
| Parakeet v3 Ultra | 4-bit | Optimized Fast | 15.58% | 522.5× | 5.13 | 1,094 |
| Parakeet v3 | 16-bit | Standard | 16.37% | 257.7× | 6.35 | 1,794 |
| Parakeet v3 | 16-bit | Optimized Exact | 16.42% | 463.9× | 4.81 | 1,777 |
| Parakeet v3 | 16-bit | Optimized Fast | 16.42% | 494.3× | 4.70 | 1,765 |
| Parakeet v3 | 8-bit | Standard | 16.40% | 241.2× | 9.05 | 1,279 |
| Parakeet v3 | 8-bit | Optimized Exact | 16.35% | 363.3× | 8.68 | 1,883 |
| Parakeet v3 | 8-bit | Optimized Fast | 16.30% | 504.4× | 5.41 | 1,342 |
| Parakeet v3 | 4-bit | Standard | 17.02% | 242.3× | 8.84 | 1,024 |
| Parakeet v3 | 4-bit | Optimized Exact | 16.90% | 365.0× | 8.49 | 1,628 |
| Parakeet v3 | 4-bit | Optimized Fast | 16.95% | 512.3× | 5.20 | 1,088 |
| Qwen3 ASR 1.7B | 16-bit | Standard | 15.00% | 26.6× | 81.03 | 4,624 |
| Qwen3 ASR 1.7B | 16-bit | Optimized Exact | 15.00% | 29.6× | 72.73 | 5,118 |
| Qwen3 ASR 1.7B | 16-bit | Optimized Fast = Exact | 15.00% | 29.6× | 72.73 | 5,118 |
| Qwen3 ASR 1.7B | 8-bit | Standard | 15.11% | 36.7× | 71.35 | 3,197 |
| Qwen3 ASR 1.7B | 8-bit | Optimized Exact | 15.11% | 44.3× | 64.81 | 3,721 |
| Qwen3 ASR 1.7B | 8-bit | Optimized Fast = Exact | 15.11% | 44.3× | 64.81 | 3,721 |
| Qwen3 ASR 1.7B | 4-bit | Standard | 18.32% | 47.3× | 56.23 | 2,367 |
| Qwen3 ASR 1.7B | 4-bit | Optimized Exact | 18.32% | 61.5× | 51.43 | 2,890 |
| Qwen3 ASR 1.7B | 4-bit | Optimized Fast = Exact | 18.32% | 61.5× | 51.43 | 2,890 |
| Qwen3 ASR 0.6B | 16-bit | Standard | 15.89% | 51.7× | 36.82 | 2,142 |
| Qwen3 ASR 0.6B | 16-bit | Optimized Exact | 15.89% | 64.4× | 34.07 | 2,406 |
| Qwen3 ASR 0.6B | 16-bit | Optimized Fast = Exact | 15.89% | 64.4× | 34.07 | 2,406 |
| Qwen3 ASR 0.6B | 8-bit | Standard | 16.04% | 60.7× | 33.58 | 1,639 |
| Qwen3 ASR 0.6B | 8-bit | Optimized Exact | 16.04% | 82.0× | 30.05 | 1,920 |
| Qwen3 ASR 0.6B | 8-bit | Optimized Fast = Exact | 16.04% | 82.0× | 30.05 | 1,920 |
| Qwen3 ASR 0.6B | 4-bit | Standard | 17.56% | 68.4× | 28.71 | 1,332 |
| Qwen3 ASR 0.6B | 4-bit | Optimized Exact | 17.56% | 96.4× | 25.60 | 1,648 |
| Qwen3 ASR 0.6B | 4-bit | Optimized Fast = Exact | 17.56% | 96.4× | 25.60 | 1,648 |
| Whisper large-v3 | 16-bit | Standard | 17.06% | 29.5× | 86.34 | 3,923 |
| Whisper large-v3 | 16-bit | Optimized Exact | 17.06% | 34.9× | 82.41 | 3,916 |
| Whisper large-v3 | 16-bit | Optimized Fast = Exact | 17.06% | 34.9× | 82.41 | 3,916 |
| Whisper large-v3 | 8-bit | Standard | 17.27% | 34.0× | 80.65 | 3,141 |
| Whisper large-v3 | 8-bit | Optimized Exact | 17.27% | 43.5× | 73.70 | 3,116 |
| Whisper large-v3 | 8-bit | Optimized Fast = Exact | 17.27% | 43.5× | 73.70 | 3,116 |
| Whisper large-v3 | 4-bit | Standard | 17.10% | 40.4× | 73.16 | 2,102 |
| Whisper large-v3 | 4-bit | Optimized Exact | 17.10% | 54.4× | 65.96 | 2,097 |
| Whisper large-v3 | 4-bit | Optimized Fast = Exact | 17.10% | 54.4× | 65.96 | 2,097 |
| Whisper large-v3 turbo | 16-bit | Standard | 16.57% | 83.9× | 38.22 | 2,499 |
| Whisper large-v3 turbo | 16-bit | Optimized Exact | 16.57% | 115.7× | 36.56 | 2,516 |
| Whisper large-v3 turbo | 16-bit | Optimized Fast = Exact | 16.57% | 115.7× | 36.56 | 2,516 |
| Whisper large-v3 turbo | 8-bit | Standard | 16.52% | 91.5× | 37.06 | 2,419 |
| Whisper large-v3 turbo | 8-bit | Optimized Exact | 16.52% | 129.8× | 34.64 | 2,363 |
| Whisper large-v3 turbo | 8-bit | Optimized Fast = Exact | 16.52% | 129.8× | 34.64 | 2,363 |
| Whisper large-v3 turbo | 4-bit | Standard | 16.96% | 93.8× | 39.81 | 1,723 |
| Whisper large-v3 turbo | 4-bit | Optimized Exact | 16.96% | 134.6× | 37.27 | 1,731 |
| Whisper large-v3 turbo | 4-bit | Optimized Fast = Exact | 16.96% | 134.6× | 37.27 | 1,731 |
| Nemotron 3.5 Streaming | 16-bit | Standard | 23.39% | 7.3× | 188.25 | 2,253 |
| Nemotron 3.5 Streaming | 16-bit | Optimized Exact | 23.39% | 19.3× | 80.33 | 2,725 |
| Nemotron 3.5 Streaming | 16-bit | Optimized Fast | 23.35% | 34.9× | 50.03 | 1,655 |
| Nemotron 3.5 Streaming | 8-bit | Standard | 23.42% | 15.1× | 112.00 | 1,191 |
| Nemotron 3.5 Streaming | 8-bit | Optimized Exact | 23.42% | 29.8× | 58.71 | 1,251 |
| Nemotron 3.5 Streaming | 8-bit | Optimized Fast | 23.44% | 38.4× | 43.07 | 1,114 |
| Nemotron 3.5 Streaming | 4-bit | Standard | 32.78% | 15.3× | 110.17 | 927 |
| Nemotron 3.5 Streaming | 4-bit | Optimized Exact | 32.78% | 30.4× | 55.57 | 980 |
| Nemotron 3.5 Streaming | 4-bit | Optimized Fast | 32.80% | 39.7× | 39.77 | 847 |
| ElevenLabs Scribe v2 | Cloud | Estimate | 13.4% (11.8–13.9) | — | — | — |
| Microsoft Azure Speech | Cloud | Estimate | 12.9% (11.3–13.3) | — | — | — |

Measured 1–3 October 2026 on an Apple M5 Max (40-core GPU), macOS 26.6. English word error rate on the 167 English minutes of the 239.7-minute v2 suite; nine other languages are scored separately. Speed, energy and peak memory on its 22.5-minute quick subset with the model loaded: speed and energy are the median of three runs, memory the peak. Speed is audio seconds per second of work. Energy is CPU, GPU, Neural Engine and memory, net of idle; the screen and the rest of the Mac are not counted, so hours per 100 Wh are not battery life. Standard is stock MLX; Optimized adds Vella's custom kernels, as Exact and Fast recipes, and "Fast = Exact" marks a model whose Fast runs the Exact recipe. M5, M5 Pro and M5 Max have the GPU Neural Accelerators Vella targets. Vella checks its optimized paths against stock MLX and retains a stock fallback. Our performance measurements are from an M5 Max with a 40-core GPU; other M5 configurations have not yet been tested. M1–M4 Macs run the self-tested stock fallback paths; their performance has not been measured. Cloud rows are estimates scaled from the Hugging Face Open ASR Leaderboard; no audio was sent to them. Every figure in benchmarks.json: https://github.com/TobyNoSkillSon/Vella/blob/main/Resources/benchmarks.json

## Built for you and your agent

Paste this to your coding agent: "Tell me about Vella and what you think of it: https://github.com/TobyNoSkillSon/Vella"

```sh
vella transcribe talk.m4a              # plain text
vella transcribe talk.m4a --srt        # subtitles; --vtt, --json, --verbose-json
vella url                              # base URL of the OpenAI-compatible API (127.0.0.1, port changes per launch)
```

`POST /v1/audio/transcriptions` and `GET /v1/models` work with OpenAI's SDKs with only the base URL changed.

On a different Mac? Your agent can benchmark Vella there and, with your OK, send the numbers back: [Benchmark kit](https://github.com/TobyNoSkillSon/Vella/tree/main/Benchmarks) (start at its README).

## Privacy

Recognition runs in sandboxed helpers with no network access. Vella never uploads audio or transcripts and has no telemetry. Models download from Hugging Face only when you choose Get; Vella checks GitHub for a new version daily when due (failed checks retry about hourly) and downloads it only when you choose Update Now.

## Family

Verdict and Vella are free, open-source Mac apps that run AI models on your own Apple Silicon chip, for you and your coding agent: Verdict decides, Vella listens. Verdict: https://github.com/TobyNoSkillSon/Verdict
