# Vella: offline dictation, tuned for your Mac Vella 2.0.1 is a free, open-source (MIT) menu-bar app for Apple Silicon Macs (macOS 26 or newer). Press Control + Command + N, speak, press it again: a local speech-to-text model transcribes and inserts the text where you were typing. Streaming mode types as you speak. Vella never uploads your audio or transcripts. On an M5 Max, Parakeet v3 Ultra runs 1.9× faster than on stock MLX, on 28% less energy, with no measurable accuracy loss (16-bit, Optimized Fast, the default). - Source: https://github.com/TobyNoSkillSon/Vella - Install: `curl -fsSL https://raw.githubusercontent.com/TobyNoSkillSon/Vella/main/scripts/install-public.sh | bash` - For coding agents: the `transcribe` skill at https://tobynoskillson.com/vella/.well-known/agent-skills/transcribe/SKILL.md (offline transcription from the `vella` command or an OpenAI-compatible local API) - Support the developer: https://github.com/sponsors/TobyNoSkillSon ## Smaller weights should be faster. On stock MLX they weren't. Parakeet v3 Ultra, Apple M5 Max. Stock MLX is Vella's Standard path: the same build with its custom kernels switched off. Speed is audio seconds per second of work. | Weights | Stock MLX | Vella | Faster | Less energy | English WER stock → Vella | |---|---|---|---|---|---| | 16-bit | 261.4× | 507.1× | 1.9× | 28% | 15.46% → 15.51% | | 8-bit | 244.8× | 515.0× | 2.1× | 41% | 15.46% → 15.46% | | 4-bit | 246.7× | 522.5× | 2.1× | 41% | 15.62% → 15.58% | Where the weights are, and what changed (Parakeet v3 Ultra, 627M weights): | Part | Share of weights | Stock MLX | Vella | |---|---|---|---| | Feed-forward (48 blocks) | 64% | Unpacks the quantized weights before multiplying. At 16-bit, its tiles leave GPU cores idle. | Feeds 8- and 4-bit codes straight into the M5's tensor unit. At 16-bit, smaller split-K tiles expose more parallel work. | | Attention (24 layers) | 20% | Three separate projections, for Q, K and V. | One fused projection for all three. | | Convolution (24 layers) | 12% | Gate, convolution, batch norm and SiLU, one after another. | Batch norm folded into the weights; the rest is one Metal kernel. | | Decoder + joint | 3% | Decodes one step at a time. | 32 steps in one compiled graph, sized to the frames left. Kept at 16-bit. | Kernel source: https://github.com/TobyNoSkillSon/Vella/tree/main/Worker/Sources/SmallMGEMM ## Every model, stock MLX then Vella Each model at its native 16-bit weights, Optimized Fast against Standard; speed in audio seconds per second of work. | Model | Stock MLX | Vella | Faster | Less energy | English WER stock → Vella | |---|---|---|---|---|---| | Nemotron 3.5 Streaming | 7.3× | 34.9× | 4.8× | 73% | 23.39% → 23.35% | | Parakeet v3 Ultra | 261.4× | 507.1× | 1.9× | 28% | 15.46% → 15.51% | | Parakeet v3 | 257.7× | 494.3× | 1.9× | 26% | 16.37% → 16.42% | | Whisper large-v3 turbo | 83.9× | 115.7× | 1.4× | 4% | 16.57% → 16.57% | | Qwen3 ASR 0.6B | 51.7× | 64.4× | 1.2× | 7% | 15.89% → 15.89% | | Whisper large-v3 | 29.5× | 34.9× | 1.2× | 5% | 17.06% → 17.06% | | Qwen3 ASR 1.7B | 26.6× | 29.6× | 1.1× | 10% | 15.00% → 15.00% | ## Energy Hours of speech transcribed per 100 Wh of chip energy. Not battery life: the screen and the rest of the Mac are not counted. | Model | J per audio minute, stock → Vella | Stock MLX | Vella | More speech | |---|---|---|---|---| | Nemotron 3.5 Streaming | 188.25 → 50.03 | 31 h | 119 h | 3.8× | | Parakeet v3 Ultra | 6.33 → 4.58 | 948 h | 1,309 h | 1.4× | | Parakeet v3 | 6.35 → 4.70 | 945 h | 1,275 h | 1.3× | | Qwen3 ASR 1.7B | 81.03 → 72.73 | 74 h | 82 h | 1.1× | | Qwen3 ASR 0.6B | 36.82 → 34.07 | 162 h | 176 h | 1.1× | | Whisper large-v3 | 86.34 → 82.41 | 69 h | 72 h | 1.0× | | Whisper large-v3 turbo | 38.22 → 36.56 | 156 h | 164 h | 1.0× | ## Every number | Model | Precision | Path | English WER | Speed | J / audio min | Peak RAM MiB | |---|---|---|---|---|---|---| | Parakeet v3 Ultra | 16-bit | Standard | 15.46% | 261.4× | 6.33 | 1,796 | | Parakeet v3 Ultra | 16-bit | Optimized Exact | 15.49% | 474.5× | 4.71 | 1,808 | | Parakeet v3 Ultra | 16-bit | Optimized Fast | 15.51% | 507.1× | 4.58 | 1,792 | | Parakeet v3 Ultra | 8-bit | Standard | 15.46% | 244.8× | 8.99 | 1,280 | | Parakeet v3 Ultra | 8-bit | Optimized Exact | 15.46% | 369.0× | 8.62 | 1,888 | | Parakeet v3 Ultra | 8-bit | Optimized Fast | 15.46% | 515.0× | 5.35 | 1,352 | | Parakeet v3 Ultra | 4-bit | Standard | 15.62% | 246.7× | 8.75 | 1,024 | | Parakeet v3 Ultra | 4-bit | Optimized Exact | 15.65% | 370.6× | 8.41 | 1,640 | | Parakeet v3 Ultra | 4-bit | Optimized Fast | 15.58% | 522.5× | 5.13 | 1,094 | | Parakeet v3 | 16-bit | Standard | 16.37% | 257.7× | 6.35 | 1,794 | | Parakeet v3 | 16-bit | Optimized Exact | 16.42% | 463.9× | 4.81 | 1,777 | | Parakeet v3 | 16-bit | Optimized Fast | 16.42% | 494.3× | 4.70 | 1,765 | | Parakeet v3 | 8-bit | Standard | 16.40% | 241.2× | 9.05 | 1,279 | | Parakeet v3 | 8-bit | Optimized Exact | 16.35% | 363.3× | 8.68 | 1,883 | | Parakeet v3 | 8-bit | Optimized Fast | 16.30% | 504.4× | 5.41 | 1,342 | | Parakeet v3 | 4-bit | Standard | 17.02% | 242.3× | 8.84 | 1,024 | | Parakeet v3 | 4-bit | Optimized Exact | 16.90% | 365.0× | 8.49 | 1,628 | | Parakeet v3 | 4-bit | Optimized Fast | 16.95% | 512.3× | 5.20 | 1,088 | | Qwen3 ASR 1.7B | 16-bit | Standard | 15.00% | 26.6× | 81.03 | 4,624 | | Qwen3 ASR 1.7B | 16-bit | Optimized Exact | 15.00% | 29.6× | 72.73 | 5,118 | | Qwen3 ASR 1.7B | 16-bit | Optimized Fast = Exact | 15.00% | 29.6× | 72.73 | 5,118 | | Qwen3 ASR 1.7B | 8-bit | Standard | 15.11% | 36.7× | 71.35 | 3,197 | | Qwen3 ASR 1.7B | 8-bit | Optimized Exact | 15.11% | 44.3× | 64.81 | 3,721 | | Qwen3 ASR 1.7B | 8-bit | Optimized Fast = Exact | 15.11% | 44.3× | 64.81 | 3,721 | | Qwen3 ASR 1.7B | 4-bit | Standard | 18.32% | 47.3× | 56.23 | 2,367 | | Qwen3 ASR 1.7B | 4-bit | Optimized Exact | 18.32% | 61.5× | 51.43 | 2,890 | | Qwen3 ASR 1.7B | 4-bit | Optimized Fast = Exact | 18.32% | 61.5× | 51.43 | 2,890 | | Qwen3 ASR 0.6B | 16-bit | Standard | 15.89% | 51.7× | 36.82 | 2,142 | | Qwen3 ASR 0.6B | 16-bit | Optimized Exact | 15.89% | 64.4× | 34.07 | 2,406 | | Qwen3 ASR 0.6B | 16-bit | Optimized Fast = Exact | 15.89% | 64.4× | 34.07 | 2,406 | | Qwen3 ASR 0.6B | 8-bit | Standard | 16.04% | 60.7× | 33.58 | 1,639 | | Qwen3 ASR 0.6B | 8-bit | Optimized Exact | 16.04% | 82.0× | 30.05 | 1,920 | | Qwen3 ASR 0.6B | 8-bit | Optimized Fast = Exact | 16.04% | 82.0× | 30.05 | 1,920 | | Qwen3 ASR 0.6B | 4-bit | Standard | 17.56% | 68.4× | 28.71 | 1,332 | | Qwen3 ASR 0.6B | 4-bit | Optimized Exact | 17.56% | 96.4× | 25.60 | 1,648 | | Qwen3 ASR 0.6B | 4-bit | Optimized Fast = Exact | 17.56% | 96.4× | 25.60 | 1,648 | | Whisper large-v3 | 16-bit | Standard | 17.06% | 29.5× | 86.34 | 3,923 | | Whisper large-v3 | 16-bit | Optimized Exact | 17.06% | 34.9× | 82.41 | 3,916 | | Whisper large-v3 | 16-bit | Optimized Fast = Exact | 17.06% | 34.9× | 82.41 | 3,916 | | Whisper large-v3 | 8-bit | Standard | 17.27% | 34.0× | 80.65 | 3,141 | | Whisper large-v3 | 8-bit | Optimized Exact | 17.27% | 43.5× | 73.70 | 3,116 | | Whisper large-v3 | 8-bit | Optimized Fast = Exact | 17.27% | 43.5× | 73.70 | 3,116 | | Whisper large-v3 | 4-bit | Standard | 17.10% | 40.4× | 73.16 | 2,102 | | Whisper large-v3 | 4-bit | Optimized Exact | 17.10% | 54.4× | 65.96 | 2,097 | | Whisper large-v3 | 4-bit | Optimized Fast = Exact | 17.10% | 54.4× | 65.96 | 2,097 | | Whisper large-v3 turbo | 16-bit | Standard | 16.57% | 83.9× | 38.22 | 2,499 | | Whisper large-v3 turbo | 16-bit | Optimized Exact | 16.57% | 115.7× | 36.56 | 2,516 | | Whisper large-v3 turbo | 16-bit | Optimized Fast = Exact | 16.57% | 115.7× | 36.56 | 2,516 | | Whisper large-v3 turbo | 8-bit | Standard | 16.52% | 91.5× | 37.06 | 2,419 | | Whisper large-v3 turbo | 8-bit | Optimized Exact | 16.52% | 129.8× | 34.64 | 2,363 | | Whisper large-v3 turbo | 8-bit | Optimized Fast = Exact | 16.52% | 129.8× | 34.64 | 2,363 | | Whisper large-v3 turbo | 4-bit | Standard | 16.96% | 93.8× | 39.81 | 1,723 | | Whisper large-v3 turbo | 4-bit | Optimized Exact | 16.96% | 134.6× | 37.27 | 1,731 | | Whisper large-v3 turbo | 4-bit | Optimized Fast = Exact | 16.96% | 134.6× | 37.27 | 1,731 | | Nemotron 3.5 Streaming | 16-bit | Standard | 23.39% | 7.3× | 188.25 | 2,253 | | Nemotron 3.5 Streaming | 16-bit | Optimized Exact | 23.39% | 19.3× | 80.33 | 2,725 | | Nemotron 3.5 Streaming | 16-bit | Optimized Fast | 23.35% | 34.9× | 50.03 | 1,655 | | Nemotron 3.5 Streaming | 8-bit | Standard | 23.42% | 15.1× | 112.00 | 1,191 | | Nemotron 3.5 Streaming | 8-bit | Optimized Exact | 23.42% | 29.8× | 58.71 | 1,251 | | Nemotron 3.5 Streaming | 8-bit | Optimized Fast | 23.44% | 38.4× | 43.07 | 1,114 | | Nemotron 3.5 Streaming | 4-bit | Standard | 32.78% | 15.3× | 110.17 | 927 | | Nemotron 3.5 Streaming | 4-bit | Optimized Exact | 32.78% | 30.4× | 55.57 | 980 | | Nemotron 3.5 Streaming | 4-bit | Optimized Fast | 32.80% | 39.7× | 39.77 | 847 | | ElevenLabs Scribe v2 | Cloud | Estimate | 13.4% (11.8–13.9) | — | — | — | | Microsoft Azure Speech | Cloud | Estimate | 12.9% (11.3–13.3) | — | — | — | Measured 1–3 October 2026 on an Apple M5 Max (40-core GPU), macOS 26.6. English word error rate on the 167 English minutes of the 239.7-minute v2 suite; nine other languages are scored separately. Speed, energy and peak memory on its 22.5-minute quick subset with the model loaded: speed and energy are the median of three runs, memory the peak. Speed is audio seconds per second of work. Energy is CPU, GPU, Neural Engine and memory, net of idle; the screen and the rest of the Mac are not counted, so hours per 100 Wh are not battery life. Standard is stock MLX; Optimized adds Vella's custom kernels, as Exact and Fast recipes, and "Fast = Exact" marks a model whose Fast runs the Exact recipe. M5, M5 Pro and M5 Max have the GPU Neural Accelerators Vella targets. Vella checks its optimized paths against stock MLX and retains a stock fallback. Our performance measurements are from an M5 Max with a 40-core GPU; other M5 configurations have not yet been tested. M1–M4 Macs run the self-tested stock fallback paths; their performance has not been measured. Cloud rows are estimates scaled from the Hugging Face Open ASR Leaderboard; no audio was sent to them. Every figure in benchmarks.json: https://github.com/TobyNoSkillSon/Vella/blob/main/Resources/benchmarks.json ## Built for you and your agent Paste this to your coding agent: "Tell me about Vella and what you think of it: https://github.com/TobyNoSkillSon/Vella" ```sh vella transcribe talk.m4a # plain text vella transcribe talk.m4a --srt # subtitles; --vtt, --json, --verbose-json vella url # base URL of the OpenAI-compatible API (127.0.0.1, port changes per launch) ``` `POST /v1/audio/transcriptions` and `GET /v1/models` work with OpenAI's SDKs with only the base URL changed. On a different Mac? Your agent can benchmark Vella there and, with your OK, send the numbers back: [Benchmark kit](https://github.com/TobyNoSkillSon/Vella/tree/main/Benchmarks) (start at its README). ## Privacy Recognition runs in sandboxed helpers with no network access. Vella never uploads audio or transcripts and has no telemetry. Models download from Hugging Face only when you choose Get; Vella checks GitHub for a new version daily when due (failed checks retry about hourly) and downloads it only when you choose Update Now. ## Family Verdict and Vella are free, open-source Mac apps that run AI models on your own Apple Silicon chip, for you and your coding agent: Verdict decides, Vella listens. Verdict: https://github.com/TobyNoSkillSon/Verdict --- --- name: transcribe description: Transcribe audio files offline with Vella, the local speech-to-text app on this Mac. Use for recordings, voice memos, meetings, lectures, podcasts or video soundtracks (wav, mp3, m4a, flac, caf, aiff; up to 3 hours per file) when you need the text, subtitles (srt/vtt) or timed segments, without sending audio anywhere. --- # Transcribe audio with Vella Standard is optimized for your Mac through MLX; Optimized adds our custom kernels, measured on M5 Max so far. Vella runs speech-recognition models on this Mac (Apple Silicon); the user's audio and transcripts never leave the machine. Vella checks GitHub's releases API for a newer version at most once a day (at launch when due; failed checks retry about hourly), and downloads models or updates only when asked; it has no telemetry and uploads no audio. The models the user dictates with also transcribe your files. The user's own dictation always goes first, so a file may wait a moment while they speak. ## When to use Fits when you have an audio file (or a video's audio track saved as audio) and need its words: plain text, OpenAI-style JSON, timed segments (up to about 25 s each, cut at pauses), or SRT/VTT subtitles. Use another tool to translate, to identify speakers or for word-level timestamps. Split files longer than 3 hours. ## Install `vella status` prints one line when Vella is installed. If `vella` is missing, try `~/.local/bin/vella`; if that is missing too, check `/Applications/Vella.app/Contents/Helpers/vella` (the DMG does not link a command); otherwise ask the user before installing: `curl -fsSL https://raw.githubusercontent.com/TobyNoSkillSon/Vella/main/scripts/install-public.sh | bash` (Apple Silicon, macOS 26 or newer; it ends with `ready: …`). A fresh install has no model. List cells and the pinned source/size with `vella models --json`; ask consent before `vella get ID --yes`. ## Results Get prints download byte progress on stderr while it waits; a progressing download has no total-duration limit. `vella transcribe` prints the transcript on stdout and nothing else; `--json`, `--verbose-json`, `--srt` and `--vtt` print that format instead. `vella status` and `vella url` print one line, `vella models` one line per catalog model. An error is one line on stderr, `error: …`, that says what to do (for example "not downloaded; get it in Vella → Models…", or a memory refusal with the model's size), and the exit code is 1. Pass that line to the user. Status, Models, URL, transcription and model/settings controls start Vella if needed. Help, Version, Skill and Diagnose do not. `--language` does not change recognition in 2.0; it only sets `verbose_json.language`, which echoes the requested code, or `unknown` when none is given. Detected language is not reported. ## Commands ```sh vella transcribe talk.m4a # the transcript as plain text vella transcribe talk.m4a --srt > talk.srt # subtitles; --vtt, --json, --verbose-json (segments) also work vella transcribe talk.m4a --model parakeet-v3 # a specific model vella transcribe talk.m4a --verbose-json --language pl # echoed in verbose JSON, not a detection result vella models # parakeet-v3-ultra Parakeet v3 Ultra · bf16 · Optimized Fast · loaded · current dictation model · Unload (one line per catalog model) vella status # Vella 2.0.0 running (pid 29335), parakeet-v3-ultra bf16 · Optimized Fast loaded · dictation model Parakeet v3 Ultra (bf16, Optimized Fast) · API http://127.0.0.1:63080/v1 vella url # http://127.0.0.1:63080/v1 vella diagnose # a bug report for the user; its last line is a prefilled GitHub issue link ``` Speed differences use N× faster at a ratio of 2× or above and N% faster below; decreases use N% slower. Energy differences always use percentages (N% less or N% more). A noise-level WER/Format difference reads same. ## Pick and get a model Start with `parakeet-v3-ultra` at `bf16`, Optimized Fast: fastest, with near-best English accuracy, and supports 24 other European languages; for other languages choose `whisper-large-v3-turbo`. Choose a larger model only for a specific accuracy need. | Model id | Mode | What it is | Language count | Params | Licence | Tiers offered | English WER % | Speed | J / audio min | Peak RAM MB | |---|---|---|---|---|---|---|---|---|---|---| | `parakeet-v3-ultra` | dictation | Moondream's post-training of NVIDIA Parakeet v3 for dictation in 25 European languages; none from outside Europe | 25 | 0.6B | CC BY 4.0 | bf16, int8, int4 | 15.51 | 507.1× | 4.58 | 1792 | | `parakeet-v3` | dictation | The unmodified Parakeet v3 that Ultra is post-trained from: the same 25 European languages, no others | 25 | 0.6B | CC BY 4.0 | bf16, int8, int4 | 16.42 | 494.3× | 4.70 | 1765 | | `qwen3-asr-1.7b` | dictation | Dictation in 30 languages, including Chinese, Japanese and Korean, which Parakeet lacks; slower than Parakeet | 30 | 1.7B | Apache-2.0 | bf16, int8, int4 | 15.00 | 29.6× | 72.73 | 5118 | | `qwen3-asr-0.6b` | dictation | The smaller Qwen3 ASR: the same 30 languages in less memory, a little less accurate than the 1.7B | 30 | 0.6B | Apache-2.0 | bf16, int8, int4 | 15.89 | 64.4× | 34.07 | 2406 | | `whisper-large-v3` | dictation | Dictation in about 100 languages, the most of any model here, from a family other than Parakeet and Qwen | 100 | 1.55B | Apache-2.0 | fp16, int8, int4 | 17.06 | 34.9× | 82.41 | 3916 | | `whisper-large-v3-turbo` | dictation | Whisper large-v3 with 4 decoder layers instead of 32: the same languages, much faster, a little less accurate outside English | 100 | 0.8B | MIT | fp16, int8, int4 | 16.57 | 115.7× | 36.56 | 2516 | | `nemotron-3.5-streaming-0.6b` | streaming | Transcribes 28 languages as the audio arrives, so Streaming mode types while you speak; not used for Dictation | 28 | 0.6B | OpenMDW-1.1 (MLX conversion: NVIDIA Open Model License) | bf16, int8, int4 | 23.35 | 34.9× | 50.03 | 1655 | English WER is on the 167 English minutes of v2 (239.7 min total); nine other languages are scored separately. Figures use the first offered, measured 16-bit cell: Optimized Fast, then Exact, then Standard. Speed is × real time and energy is joules per minute of audio on the reference Mac (Apple M5 Max, macOS 26.6); measurements are from an M5 Max with a 40-core GPU; other M5 configurations have not yet been tested. On any other configuration, WER/Format/Peak RAM remain reference measurements; Speed remains the M5 Max measured speed, lighter grey with a small M5 Max label; J/min is not known. Tooltip: Measured on an M5 Max (40-core GPU). Your Mac will differ; vella diagnose measures it. Fast enables every kept lever for that model and precision; Exact enables only exact kept levers; Standard enables none. Streaming models do not transcribe files. `vella models --json` lists every cell, its reference English WER and peak RAM, M5 Max measured speed labelled by hardware and chip-qualified energy, measurement provenance or refusal reason, and the source and size. Equal recipes use the canonical measured cell named in `cells[].provenance.display_cell` (`display_cells` in the benchmark file). Whisper Standard and Optimized were measured together on the faithful FP16 build; the quiet-window refresh completed 4 October. Exact and Fast share the same exact-only whisper-4 recipe and canonical measurement. Per-cell and tier gates use that build’s FP16 Standard baseline; the earlier Float32-baseline figures remain withdrawn. Nemotron gates use same-layout native Standard controls; int8 and int4 are offered but not recommended, and int4 loses 21 test clips. Every measured cell is offered; each cell's `loss` states what it loses against 16-bit. Use one call per step, without probing or retries: ```sh vella models --json # all local table rows, cells/refusal reasons, source and exact download bytes vella select MODEL_ID --precision bf16 --path Optimized --mode Fast # or fp16/int8/int4, Standard, Exact vella get MODEL_ID --yes # only after the user consents to that model/source/size; waits for download and load vella transcribe talk.m4a --model MODEL_ID # then read the returned transcript ``` Select previews a cell; Get/Load/Reload commits it. Mode-only `vella select ID --mode Exact` is the table's switch flip and can move the preview to the native precision; a pinned Fast=Exact switch says why it cannot flip. Cells with a `reason` are unavailable; choose a listed cell without one. A cell with `loss` runs but loses quality against 16-bit (lost test clips first): prefer a cell without `loss` unless the user asks for it. Use fp16 for Whisper, bf16 for the other models. A loaded model's effective selection may be Standard if its optimized gate fell back; replies report both the effective and requested selection. Closing the Models menu discards previews, so intervene only when the user is not changing the table. If the weights are already present, use `vella load MODEL_ID` instead of Get. Get/Load/Reload use the table's manual Keep Hot and launch-set behaviour; Unload removes that launch-set entry and keeps the download. Streaming rows can be controlled but do not transcribe files. ```sh vella reload MODEL_ID vella unload MODEL_ID vella delete MODEL_ID --precision bf16 --yes # only after explicit deletion consent vella keep-hot "Manually loaded" "Always" vella keep-hot "Loaded on demand" "15 min idle" vella memory "Fit in free memory" # or "Allow swap (slower)" ``` Delete without `--yes` prints what/size and changes nothing. It uses the table gate: recording/in-use/selected models and shared or linked weights refuse with the same reason. It moves only that catalog precision's weights to Trash; recordings and transcripts stay. Locally derived precisions may have no own files: delete the native source precision after loading another model for that mode. Unload alone does not clear the saved mode selection. Keep Hot values are Always, 5/15/30/60 min idle. With no arguments `keep-hot`/`memory` report the current settings; change them only on user authority. ## Done when You have the transcript and have read enough of it to confirm it matches the audio (right language, no long stretches missing). Tell the user which file you transcribed and which model you asked for (the `--model` id, or their current dictation model). Vella picks the model again for each segment, so if the user loads or selects another model while a long file runs, the later part uses it; `vella status` afterwards names only the model selected now. Automatic transcripts misspell names and jargon: say so when those matter, and never present the text as a verbatim quote without checking it. ## API Vella's local API is OpenAI-compatible: `POST /v1/audio/transcriptions` and `GET /v1/models` at the base URL from `vella url` (127.0.0.1 only). Code written for OpenAI's transcription endpoint works with only the base URL changed. The key is ignored, but the SDKs require one. The port changes when Vella restarts, so read it each time (`vella url`, or `api_port` in `~/Library/Application Support/Vella/worker-status.json`). ```python from openai import OpenAI import subprocess client = OpenAI(base_url=subprocess.check_output(["vella", "url"], text=True).strip(), api_key="local") with open("talk.m4a", "rb") as f: text = client.audio.transcriptions.create(model="whisper-1", file=f, response_format="text") ``` ```sh curl -s "$(vella url)/audio/transcriptions" -F file=@talk.m4a -F response_format=srt ``` `model="whisper-1"` means the user's current dictation model; any id from `GET /v1/models` picks another downloaded model (it loads on demand and unloads after the user's Keep Hot time). `response_format` is `text`, `json` (`{"text": "…", "usage": …}`), `verbose_json` (adds `segments` with `start`, `end`, `text`), `srt` or `vtt`. ## Limits - No translation, no speaker labels, no word-level timestamps; files up to 3 hours. - Transcription never downloads. Get is the only CLI download route and requires explicit consent (`--yes`); Load/Reload of missing weights returns the source/size prompt without downloading. - The user's dictation takes priority; a file waits while they speak. - If transcripts look broken or transcription is far slower than expected, run `vella diagnose` and give the user its report and the bug-report link on its last line; they decide whether to file it. It never starts Vella and loads nothing unless you pass `--load`. ## Measure on another Mac Start at https://github.com/TobyNoSkillSon/Vella/blob/main/Benchmarks/README.md for the frozen full/quick suites, audio fetch, serial installed-app runner, scorer, local result check and chip optimization guide. One model is enough. Quick quality is an estimate; full quality needs the full suite. Ask before downloads, long runs, scheduling or opening an issue/PR; exclude personal recordings. Leave energy absent unless measured with the documented powermetrics protocol and admin consent.