---
name: transcribe
description: Transcribe audio files offline with Vella, the local speech-to-text app on this Mac. Use for recordings, voice memos, meetings, lectures, podcasts or video soundtracks (wav, mp3, m4a, flac, caf, aiff; up to 3 hours per file) when you need the text, subtitles (srt/vtt) or timed segments, without sending audio anywhere.
---

# Transcribe audio with Vella

Standard is optimized for your Mac through MLX; Optimized adds our custom kernels, measured on M5 Max so far

Vella runs speech-recognition models on this Mac (Apple Silicon); nothing leaves the machine. The models the user dictates with also transcribe your files. The user's own dictation always goes first, so a file may wait a moment while they speak.

## When to use

Fits when you have an audio file (or a video's audio track saved as audio) and need its words: plain text, OpenAI-style JSON, timed segments (up to about 25 s each, cut at pauses), or SRT/VTT subtitles.

Use another tool to translate, to identify speakers or for word-level timestamps. Split files longer than 3 hours.

## Install

`vella status` prints one line when Vella is installed. If `vella` is missing, try `~/.local/bin/vella`; if that is missing too, ask the user before installing: `curl -fsSL https://tobynoskillson.github.io/Vella/install.sh | bash` (Apple Silicon, macOS 26 or newer; it ends with `ready: …`). A fresh install has no model. List cells and the pinned source/size with `vella models --json`; ask consent before `vella get ID --yes`.

## Results

Get prints download byte progress on stderr while it waits; a progressing download has no total-duration limit.

`vella transcribe` prints the transcript on stdout and nothing else; `--json`, `--verbose-json`, `--srt` and `--vtt` print that format instead. `vella status` and `vella url` print one line, `vella models` one line per catalog model. An error is one line on stderr, `error: …`, that says what to do (for example "not downloaded; get it in Vella → Models…", or a memory refusal with the model's size), and the exit code is 1. Pass that line to the user. Commands start Vella if it is not running, except `vella diagnose`, `vella --help` and `vella --version`.

## Commands

```sh
vella transcribe talk.m4a                      # the transcript as plain text
vella transcribe talk.m4a --srt > talk.srt     # subtitles; --vtt, --json, --verbose-json (segments) also work
vella transcribe talk.m4a --model parakeet-v3  # a specific model
vella transcribe talk.m4a --language pl        # a language hint
vella models                                   # parakeet-v3-ultra  Parakeet v3 Ultra · 16 · Optimized Fast · loaded · current dictation model   (one line per model on this Mac)
vella status                                   # Vella 2.0.0 running (pid 29335), parakeet-v3-ultra 16 loaded · dictation model Parakeet v3 Ultra (16, Optimized Fast) · API http://127.0.0.1:63080/v1
vella url                                      # http://127.0.0.1:63080/v1
vella diagnose                                 # a bug report for the user; its last line is a prefilled GitHub issue link
```

## Pick and get a model

Start with `parakeet-v3-ultra` at `bf16`, Optimized Fast: best for English and 24 other European languages; for other languages choose `whisper-large-v3-turbo`. Choose a larger model only for a specific accuracy need.

Use one call per step, without probing or retries:

```sh
vella models --json   # all local table rows, cells/refusal reasons, source and exact download bytes
vella select MODEL_ID --precision bf16 --path Optimized --mode Fast   # or fp16/int8/int4, Standard, Exact
vella get MODEL_ID --yes   # only after the user consents to that model/source/size; waits for download and load
vella transcribe talk.m4a --model MODEL_ID   # then read the returned transcript
```

Select previews a cell; Get/Load/Reload commits it. Mode-only `vella select ID --mode Exact` is the table's switch flip and can move the preview to the native precision; a pinned Fast=Exact switch says why it cannot flip. Cells with a `reason` are unavailable; choose a listed cell without one. Use fp16 for Whisper, bf16 for the other models. A loaded model's effective selection may be Standard if its optimized gate fell back; replies report both the effective and requested selection. Closing the Models menu discards previews, so intervene only when the user is not changing the table. If the weights are already present, use `vella load MODEL_ID` instead of Get. Get/Load/Reload use the table's manual Keep Hot and launch-set behaviour; Unload removes that launch-set entry and keeps the download. Streaming rows can be controlled but do not transcribe files.

```sh
vella reload MODEL_ID
vella unload MODEL_ID
vella delete MODEL_ID --precision bf16 --yes   # only after explicit deletion consent
vella keep-hot "Manually loaded" "Always"
vella keep-hot "Loaded on demand" "15 min idle"
vella memory "Fit in free memory"   # or "Allow swap (slower)"
```

Delete without `--yes` prints what/size and changes nothing. It uses the table gate: recording/in-use/selected models and shared or linked weights refuse with the same reason. It moves only that catalog precision's weights to Trash; recordings and transcripts stay. Locally derived precisions may have no own files: delete the native source precision after loading another model for that mode. Unload alone does not clear the saved mode selection.

Keep Hot values are Always, 5/15/30/60 min idle. With no arguments `keep-hot`/`memory` report the current settings; change them only on user authority.

## Done when

You have the transcript and have read enough of it to confirm it matches the audio (right language, no long stretches missing). Tell the user which file you transcribed and which model you asked for (the `--model` id, or their current dictation model). Vella picks the model again for each segment, so if the user loads or selects another model while a long file runs, the later part uses it; `vella status` afterwards names only the model selected now. Automatic transcripts misspell names and jargon: say so when those matter, and never present the text as a verbatim quote without checking it.

## API

Vella's local API is OpenAI-compatible: `POST /v1/audio/transcriptions` and `GET /v1/models` at the base URL from `vella url` (127.0.0.1 only). Code written for OpenAI's transcription endpoint works with only the base URL changed. The key is ignored, but the SDKs require one. The port changes when Vella restarts, so read it each time (`vella url`, or `api_port` in `~/Library/Application Support/Vella/worker-status.json`).

```python
from openai import OpenAI
import subprocess
client = OpenAI(base_url=subprocess.check_output(["vella", "url"], text=True).strip(), api_key="local")
with open("talk.m4a", "rb") as f:
    text = client.audio.transcriptions.create(model="whisper-1", file=f, response_format="text")
```

```sh
curl -s "$(vella url)/audio/transcriptions" -F file=@talk.m4a -F response_format=srt
```

`model="whisper-1"` means the user's current dictation model; any id from `GET /v1/models` picks another downloaded model (it loads on demand and unloads after the user's Keep Hot time). `response_format` is `text`, `json` (`{"text": "…", "usage": …}`), `verbose_json` (adds `segments` with `start`, `end`, `text`), `srt` or `vtt`.

## Limits

- No translation, no speaker labels, no word-level timestamps; files up to 3 hours.
- Transcription never downloads. Get is the only CLI download route and requires explicit consent (`--yes`); Load/Reload of missing weights returns the source/size prompt without downloading.
- The user's dictation takes priority; a file waits while they speak.
- If transcripts look broken or transcription is far slower than expected, run `vella diagnose` and give the user its report and the bug-report link on its last line; they decide whether to file it. It never starts Vella and loads nothing unless you pass `--load`.

## Measure on another Mac

For community measurement, follow https://github.com/TobyNoSkillSon/Vella/blob/main/CONTRIBUTING.md#community-measurement. One model at a time is enough. The shipped `diagnose` clips are diagnostic evidence, not the full benchmark suite; label the dataset and chip. Use Opus 5.5 / Sol 6.1 or better for an unattended run, and ask consent before scheduling, downloading or submitting an issue/PR. Energy needs `powermetrics` and admin consent; otherwise leave energy unmeasured.
