Offline dictation, tuned for your Mac.

On an M5 Max, Parakeet v3 Ultra runs 1.9× faster than on stock MLX, on 28% less energy, with no measurable accuracy loss.

curl -fsSL https://raw.githubusercontent.com/TobyNoSkillSon/Vella/main/scripts/install-public.sh | bash

Free and open source · MIT · Apple Silicon, macOS 26+

Smaller weights should be faster. On stock MLX they weren't.

  1. 16-bitWER 15.46% → 15.51%
    261×507×
    1.9×28% less energy
  2. 8-bitWER 15.46% → 15.46%
    245×515×
    2.1×41% less energy
  3. 4-bitWER 15.62% → 15.58%
    247×522×
    2.1×41% less energy
  • Feed-forward 64%

    Unpacks the quantized weights before multiplying. At 16-bit, its tiles leave GPU cores idle.

    Feeds 8- and 4-bit codes straight into the M5's tensor unit. At 16-bit, smaller split-K tiles expose more parallel work.

  • Attention 20%

    Three separate projections, for Q, K and V.

    One fused projection for all three.

  • Convolution 12%

    Gate, convolution, batch norm and SiLU, one after another.

    Batch norm folded into the weights; the rest is one Metal kernel.

  • Decoder + joint 3%

    Decodes one step at a time.

    32 steps in one compiled graph, sized to the frames left. Kept at 16-bit.

Every model, stock MLX then Vella.

  1. Nemotron 3.5 Streaming16-bit
    7.3×34.9×
    4.8×73% less energy
  2. Parakeet v3 Ultra16-bit
    261×507×
    1.9×28% less energy
  3. Parakeet v316-bit
    258×494×
    1.9×26% less energy
  4. Whisper large-v3 turbo16-bit
    83.9×116×
    1.4×4% less energy
  5. Qwen3 ASR 0.6B16-bit
    51.7×64.4×
    1.2×7% less energy
  6. Whisper large-v316-bit
    29.5×34.9×
    1.2×5% less energy
  7. Qwen3 ASR 1.7B16-bit
    26.6×29.6×
    1.1×10% less energy

The same energy transcribes up to 3.8× more speech.

  • Nemotron 3.5 Streaming16-bit31 h119 h
  • Parakeet v3 Ultra16-bit948 h1,309 h
  • Parakeet v316-bit945 h1,275 h
  • Qwen3 ASR 1.7B16-bit74 h82 h
  • Qwen3 ASR 0.6B16-bit162 h176 h
  • Whisper large-v316-bit69 h72 h
  • Whisper large-v3 turbo16-bit156 h164 h
Every number
ModelPrecisionPathEnglish WERSpeedJ / audio minPeak RAM MiB
Parakeet v3 Ultra16-bitStandard15.46%261.4×6.331,796
Parakeet v3 Ultra16-bitOptimized Exact15.49%474.5×4.711,808
Parakeet v3 Ultra16-bitOptimized Fast15.51%507.1×4.581,792
Parakeet v3 Ultra8-bitStandard15.46%244.8×8.991,280
Parakeet v3 Ultra8-bitOptimized Exact15.46%369.0×8.621,888
Parakeet v3 Ultra8-bitOptimized Fast15.46%515.0×5.351,352
Parakeet v3 Ultra4-bitStandard15.62%246.7×8.751,024
Parakeet v3 Ultra4-bitOptimized Exact15.65%370.6×8.411,640
Parakeet v3 Ultra4-bitOptimized Fast15.58%522.5×5.131,094
Parakeet v316-bitStandard16.37%257.7×6.351,794
Parakeet v316-bitOptimized Exact16.42%463.9×4.811,777
Parakeet v316-bitOptimized Fast16.42%494.3×4.701,765
Parakeet v38-bitStandard16.40%241.2×9.051,279
Parakeet v38-bitOptimized Exact16.35%363.3×8.681,883
Parakeet v38-bitOptimized Fast16.30%504.4×5.411,342
Parakeet v34-bitStandard17.02%242.3×8.841,024
Parakeet v34-bitOptimized Exact16.90%365.0×8.491,628
Parakeet v34-bitOptimized Fast16.95%512.3×5.201,088
Qwen3 ASR 1.7B16-bitStandard15.00%26.6×81.034,624
Qwen3 ASR 1.7B16-bitOptimized Exact15.00%29.6×72.735,118
Qwen3 ASR 1.7B16-bitOptimized Fast = Exact15.00%29.6×72.735,118
Qwen3 ASR 1.7B8-bitStandard15.11%36.7×71.353,197
Qwen3 ASR 1.7B8-bitOptimized Exact15.11%44.3×64.813,721
Qwen3 ASR 1.7B8-bitOptimized Fast = Exact15.11%44.3×64.813,721
Qwen3 ASR 1.7B4-bitStandard18.32%47.3×56.232,367
Qwen3 ASR 1.7B4-bitOptimized Exact18.32%61.5×51.432,890
Qwen3 ASR 1.7B4-bitOptimized Fast = Exact18.32%61.5×51.432,890
Qwen3 ASR 0.6B16-bitStandard15.89%51.7×36.822,142
Qwen3 ASR 0.6B16-bitOptimized Exact15.89%64.4×34.072,406
Qwen3 ASR 0.6B16-bitOptimized Fast = Exact15.89%64.4×34.072,406
Qwen3 ASR 0.6B8-bitStandard16.04%60.7×33.581,639
Qwen3 ASR 0.6B8-bitOptimized Exact16.04%82.0×30.051,920
Qwen3 ASR 0.6B8-bitOptimized Fast = Exact16.04%82.0×30.051,920
Qwen3 ASR 0.6B4-bitStandard17.56%68.4×28.711,332
Qwen3 ASR 0.6B4-bitOptimized Exact17.56%96.4×25.601,648
Qwen3 ASR 0.6B4-bitOptimized Fast = Exact17.56%96.4×25.601,648
Whisper large-v316-bitStandard17.06%29.5×86.343,923
Whisper large-v316-bitOptimized Exact17.06%34.9×82.413,916
Whisper large-v316-bitOptimized Fast = Exact17.06%34.9×82.413,916
Whisper large-v38-bitStandard17.27%34.0×80.653,141
Whisper large-v38-bitOptimized Exact17.27%43.5×73.703,116
Whisper large-v38-bitOptimized Fast = Exact17.27%43.5×73.703,116
Whisper large-v34-bitStandard17.10%40.4×73.162,102
Whisper large-v34-bitOptimized Exact17.10%54.4×65.962,097
Whisper large-v34-bitOptimized Fast = Exact17.10%54.4×65.962,097
Whisper large-v3 turbo16-bitStandard16.57%83.9×38.222,499
Whisper large-v3 turbo16-bitOptimized Exact16.57%115.7×36.562,516
Whisper large-v3 turbo16-bitOptimized Fast = Exact16.57%115.7×36.562,516
Whisper large-v3 turbo8-bitStandard16.52%91.5×37.062,419
Whisper large-v3 turbo8-bitOptimized Exact16.52%129.8×34.642,363
Whisper large-v3 turbo8-bitOptimized Fast = Exact16.52%129.8×34.642,363
Whisper large-v3 turbo4-bitStandard16.96%93.8×39.811,723
Whisper large-v3 turbo4-bitOptimized Exact16.96%134.6×37.271,731
Whisper large-v3 turbo4-bitOptimized Fast = Exact16.96%134.6×37.271,731
Nemotron 3.5 Streaming16-bitStandard23.39%7.3×188.252,253
Nemotron 3.5 Streaming16-bitOptimized Exact23.39%19.3×80.332,725
Nemotron 3.5 Streaming16-bitOptimized Fast23.35%34.9×50.031,655
Nemotron 3.5 Streaming8-bitStandard23.42%15.1×112.001,191
Nemotron 3.5 Streaming8-bitOptimized Exact23.42%29.8×58.711,251
Nemotron 3.5 Streaming8-bitOptimized Fast23.44%38.4×43.071,114
Nemotron 3.5 Streaming4-bitStandard32.78%15.3×110.17927
Nemotron 3.5 Streaming4-bitOptimized Exact32.78%30.4×55.57980
Nemotron 3.5 Streaming4-bitOptimized Fast32.80%39.7×39.77847
ElevenLabs Scribe v2CloudEstimate13.4% (11.8–13.9)accuracy only
Microsoft Azure SpeechCloudEstimate12.9% (11.3–13.3)accuracy only

Measured 1–3 October 2026 on an Apple M5 Max (40-core GPU), macOS 26.6. English word error rate on the 167 English minutes of the 239.7-minute v2 suite; nine other languages are scored separately. Speed, energy and peak memory on its 22.5-minute quick subset with the model loaded: speed and energy are the median of three runs, memory the peak. Speed is audio seconds per second of work. Energy is CPU, GPU, Neural Engine and memory, net of idle; the screen and the rest of the Mac are not counted, so hours per 100 Wh are not battery life. Standard is stock MLX; Optimized adds Vella's custom kernels, as Exact and Fast recipes, and “Fast = Exact” marks a model whose Fast runs the Exact recipe. M5, M5 Pro and M5 Max have the GPU Neural Accelerators Vella targets. Vella checks its optimized paths against stock MLX and retains a stock fallback. Our performance measurements are from an M5 Max with a 40-core GPU; other M5 configurations have not yet been tested. M1–M4 Macs run the self-tested stock fallback paths; their performance has not been measured. Cloud rows are estimates scaled from the Hugging Face Open ASR Leaderboard; no audio was sent to them. Every figure in benchmarks.json

Built for you and your agent.

Terminal
$ vella transcribe talk.m4a
The transcript, as plain text.
$ vella transcribe talk.m4a --srt
$ vella url
http://127.0.0.1:63080/v1
Paste this to your agent

Tell me about Vella and what you think of it: github.com/TobyNoSkillSon/Vella

Python
from openai import OpenAI
import subprocess
client = OpenAI(
  base_url=subprocess.getoutput("vella url"),
  api_key="local")
text = client.audio.transcriptions.create(
  model="whisper-1",
  file=open("talk.m4a", "rb"))

OpenAI-compatible /v1/audio/transcriptions, on this Mac only. The transcribe skill

On a different Mac? Your agent can benchmark Vella there and, with your OK, send the numbers back. Benchmark kit

Vella never uploads what you say.

Recognition runs with no network accessNo telemetryDownloads only when you ask

Vella is free. Help keep it that way.