Docs
The deep-reference page — every setting, the recommended way to run a big translation job, how the engine cascade actually behaves, and what to check when something goes wrong. For getting started, see Install and Engine Setup.
Recommended workflow
Getting a library fully translated efficiently is a multi-pass process, not a single run — trying to do everything in one pass either wastes local GPU/CPU time on what's usually a temporary cloud issue, or leaves recoverable items sitting as permanent failures.
1. Make sure Bazarr can see your existing subtitles
A subtitle muxed into the video file itself (not a separate .srt on disk) is invisible to Subtitlarr until Bazarr extracts it — Subtitlarr only ever reads/writes through Bazarr's API, never the filesystem directly. In Bazarr:
- Settings → Subtitles → (disable) "Treat Embedded Subtitles as Downloaded"
- Settings → Providers → add "Embedded Subtitles" (if it isn't already there), so Bazarr actually extracts those tracks to real files instead of just counting them as already satisfied
2–4. Translate in passes, escalate to local last
Run the cascade with local engines fenced off
Put a separator right after your cloud/free-tier instances (see Recommended cascade shape) and let it translate everything it can. Items that hit a genuine content block with no non-Gemini fallback in reach — or any other failure — end up failed and stay there. Leave them alone for now rather than chasing each one down individually.
Once the queue is drained, re-run just the failures
On the Queue page, filter to the Failed chip and run that filtered batch. A second pass over just the failures often clears a meaningful chunk of them on its own — a rate limit expires, a transient error doesn't repeat, quota resets overnight. Targeting the Failed filter specifically means only those items are touched, not a full re-run of everything already done.
failed against the same cloud engine is usually hitting a genuine content block, not a transient error — repeatedly re-submitting the same flagged content risks that provider's own abuse/ToS enforcement against the account behind the key, not just another failed item. If a batch is content you expect to keep tripping the filter, don't loop it against the same key indefinitely — use a throwaway/secondary account for that key rather than one tied to anything you'd mind losing, and consider routing genuinely repeat-blocked content to a local engine instead (see step 4).
Only then, move your local engine above the separator
Drag it up in the Translation Engine page for whatever's still failed. This is deliberate, not automatic: with a non-Gemini engine actually in reach, content-blocked batches get bisected instead of failing outright — Gemini keeps handling everything except the specific isolated cues that trip its filter, and only that small leftover chunk goes to the local engine. Doing this from the start would mean paying local inference time on every single content-blocked item; doing it last means only the genuine leftovers ever reach it.
5. Run the Language Check job
Even a "successful" translation can silently come back in the wrong language — the LLM echoing the source text instead of translating it. Subtitlarr's own structural checks can't catch this, since the response is well-formed, just wrong. The Jobs page's Language Check audits recently-completed items against their actual detected output language and resets any mismatch back to pending for a real retry.
LANGUAGE_CHECK_CRON — otherwise it's manual-only, so remember to run it (or schedule it) after a big batch, not just once and forget it.
How the cascade works
Subtitlarr doesn't use a single fixed engine — you build an ordered list of one or more configured instances (Ollama, llama.cpp, Gemini, NVIDIA NIM, OpenRouter, Groq). Every item is tried against the first instance; if that fails (rate limit, error, content block), it falls through to the next instance in order, and so on. A separator row can fence off a group so nothing below it is ever tried as an automatic fallback — the run simply stops with the remaining items left pending once everything above the separator is exhausted.
Content-block bisection
Gemini's own safety filter (PROHIBITED_CONTENT/SAFETY) can reject a translation batch outright — most likely on R-rated or intense material. Rather than treating the whole batch as a failure and handing all of it to a weaker fallback engine, Subtitlarr bisects the blocked batch and retries each half against the same Gemini instance, so only the actual offending cues end up isolated — the rest of the batch stays on Gemini at full quality.
Splitting stops once an isolated chunk shrinks to 10 cues, or once the item's bisection budget (extra requests spent narrowing down blocks, capped per item) runs out — whichever comes first. Either way, only that small leftover chunk falls back, and it skips straight past every other configured Gemini instance (they'd just trip the same filter again) to the first non-Gemini engine in the cascade, re-chunked to that engine's own configured batch size rather than arriving oversized.
Net effect: a normal batch costs exactly one request, same as always. A batch with a handful of blocked lines costs a few extra Gemini requests to isolate them, and only those lines go to the fallback engine. A batch blocked densely throughout (e.g. an R-rated film flagged every few lines) hits the per-item budget quickly and the remainder falls back in bulk, re-chunked properly — so one heavily-flagged movie can't consume a disproportionate share of your daily Gemini quota chasing a lost cause.
Rate limits & cooldown
3 consecutive failures against an engine instance trip a 24-hour cooldown — the instance is skipped without even attempting a live round-trip for the rest of that window. This clears automatically after 24 hours, or immediately via a Test Connection on that instance, or the Jobs page's manual "clear all rate limits" action.
All settings reference
None of this needs to be touched at container start — every one of these is also editable from the web UI afterward (Settings / Language Rules / Bazarr Connection pages). This table exists purely as a reference for what each setting does and its default. Translation engine setup lives entirely in the Translation Engine page — see Engine Setup — not in this table.
| Variable | Purpose | Default |
|---|---|---|
BAZARR_BASE_URL | Bazarr root URL | set from the UI |
BAZARR_API_KEY | Bazarr API key | set from the UI |
SCHEDULE_CRON | 5-field cron expression for the main scheduled translation job | 10 3 * * * |
AGE_THRESHOLD_DAYS | Days a subtitle must be missing before a scheduled run will translate it | 14 |
DAILY_TRANSLATION_LIMIT | Max items translated per day by scheduled/full runs (0 = unlimited); per-item re-runs bypass this | 100 |
PAUSE_BETWEEN_ITEMS_SECONDS | Rest between translations so the GPU isn't pegged non-stop | 30 |
QUEUE_UPLOADS_ENABLED | Hold translated subtitles locally instead of uploading immediately — see below | false |
PUSH_UPLOADS_CRON | Optional cron to auto-push queued uploads (only meaningful with QUEUE_UPLOADS_ENABLED); blank = manual only | 15 5 * * * |
SYNC_MEDIA_CRON | Optional cron to auto-refresh Bazarr's wanted list; blank = manual only | 0 3 * * * |
SYNC_SUBS_CRON | Optional cron to auto pre-fetch source subtitle content; blank = manual only | 5 3 * * * |
LANGUAGE_CHECK_CRON | Optional cron to audit recently-completed items for the wrong output language and reset any mismatch to pending; needs a check engine picked on the Jobs page first; blank = manual only | 0 5 * * * |
BACKUP_CRON | Daily snapshot of the whole database to /data/backups/; blank disables it | 30 2 * * * |
BACKUP_KEEP_COUNT | How many daily/manual snapshots to retain before pruning the oldest | 20 |
DB_PATH | SQLite file path inside the container | /data/subtitlarr.db |
RUN_CONCURRENCY | Reserved for future use | 1 |
LOG_LEVEL | Logging verbosity | INFO |
PORT is accepted but currently has no effect — the container always listens on 7777 internally (map it to any host port you like via Docker's own port mapping).
Queue uploads, explained
If your Bazarr host — or its storage, e.g. a NAS array — spins its disks down when idle, every immediate upload wakes it. With QUEUE_UPLOADS_ENABLED=false (the default), a completed translation uploads to Bazarr right away; over a long run translating many items, that's one wake-up per item.
Setting it true holds each finished translation locally as translated_pending_upload instead, and nothing touches Bazarr's storage until you push the whole batch at once — manually from the Jobs page, or automatically via PUSH_UPLOADS_CRON. That turns N wake-ups into one.
| Your setup | Recommended |
|---|---|
| Disks/storage never spin down | Either setting works — false is simplest |
| Disks spin down, but are already kept awake by something else | false — no benefit to queuing |
| Disks spin down and nothing else keeps them awake | true — push the batch once you're ready, or schedule it |
AI assistant access (MCP)
Subtitlarr exposes an MCP server so Claude Code, Claude Desktop, or another MCP-aware assistant can check queue/job status and drive translations directly, instead of you clicking through the UI. It's mounted on the app itself at /mcp — same host and port as the web UI, nothing extra to expose or map.
Connecting
Open the MCP Server page from the sidebar — it shows your server URL, a bearer token (auto-generated on first visit, regeneratable), and a ready-to-copy connection command for Claude Code:
claude mcp add --scope user --transport http subtitlarr http://<your-instance>/mcp --header "Authorization: Bearer <token>"
--scope user registers it once for every project directory — with the default local scope instead, the server only shows up in whichever single project folder you happened to run the command from.
For Claude Desktop, Cursor, or any other MCP-aware client, the same page also shows a plain mcpServers JSON block (streamable HTTP transport) to paste into that client's own config file.
Anyone with the token can trigger translation runs and jobs on that instance — treat it like the Bazarr API key. Only reachable behind whatever already protects access to the instance itself (VPN, reverse proxy, LAN-only); never expose it directly to the internet.
What it can do
Every tool calls the exact same code path the web UI itself uses — nothing is possible over MCP that isn't already possible by clicking around the app.
- Documentation — links to Subtitlarr's own README, full docs site, install guide, and changelog, so the assistant answers "how do I configure X" from the real current docs instead of guessing.
- Status — dashboard stats, combined job/sync status, current run progress and its items, queue listing/filtering/matching-count, one item's detail, list of models used, Bazarr-refresh (poll) status, run history list, one past run's items, past runs' aggregate stats, the live in-run event feed, the durable per-item event log, the job start/finish log, and language mismatches.
- Manual translation — fetch an item's source dialogue and submit a translation done by the connecting assistant itself, for items a configured engine can't handle (content-filter refusals, exhausted quota, a dead credential) — see the manual-translation workflow for the underlying mechanism.
- Run control — start/cancel a run, run a filtered set of items, run an explicit list of item ids, re-run a single item, refresh from Bazarr without translating.
- Sync/jobs — trigger media/subtitle sync, push queued uploads, run the language check (and set which engine it uses), run the stale audit, close stuck runs, clear engine rate limits.
- Engine cascade — list configured engine instances (API keys always masked) and reorder the cascade. Adding or editing an instance's credentials stays a Translation Engine page action, not available over MCP.
- Schedule/cutoff — read and update the age threshold, daily translation limit, cron expression, and pause between items.
- Language rules — read source-language priority, target-language allowlist, and regional variant preferences (read-only).
The non-retry safety rule
An item already failed with a content-blocked, quota-exhausted, or dead-credential error should never be resubmitted to the same engine — repeating that risks the provider's own abuse enforcement against the account (up to suspension), not just another failed item. This rule is written directly into the run-control tools' own descriptions: the assistant is expected to classify a failure first and route non-retryable ones to manual translation instead of retrying them. It's guidance the assistant follows, not a hard block enforced by the server today.
External translate API
A plain REST endpoint for any third-party tool — Bazarr, another *arr-adjacent app, a script — to submit subtitle content directly for translation by Subtitlarr's own configured engine cascade, with no Bazarr wanted-list item involved at all. This is separate from the queue/run endpoints the web UI itself uses: nothing submitted here appears on the Queue or Dashboard, since there's no Bazarr item behind it.
Authentication
Its own bearer token, independent of both the MCP server's token and Bazarr's own API key. Open the External Translate page from the sidebar to view it (and regenerate it if it's ever leaked) — it's shown there the same way the MCP Server page shows its own token, not something a caller fetches programmatically. Copy it from that page into whatever you're configuring (Bazarr's webhook, a script, another tool).
Send it as Authorization: Bearer <token> on every request below.
Submitting content
source_language/target_language are normalized to their bare language subtag before translation — "es-ES", "pt_BR", and "EN" are all accepted and treated the same as "es", "pt", and "en" respectively (only the part before a -/_ is used, lowercased). This matches the bare-code convention Bazarr itself uses (even its own non-standard codes like "pb" for Brazilian Portuguese are flat, never hyphenated) and protects against any other caller's convention differing.
POST /api/external-translate
{
"source_language": "en",
"target_language": "ca",
// exactly one of:
"srt_content": "1\n00:00:01,000 --> 00:00:03,000\nHello there.\n",
"cues": [ { "index": 1, "content": "Hello there.", "proprietary": "",
"start": {"hours":0,"minutes":0,"seconds":1,"total_seconds":1,"microseconds":0},
"end": {"hours":0,"minutes":0,"seconds":3,"total_seconds":3,"microseconds":0} } ]
}
cues is the exact shape Bazarr's own GET /api/subtitles/contents already returns — a caller that got its source subtitle from Bazarr can hand that response straight through without re-serializing it to a flat .srt file first. Anyone else without that shape on hand just sends srt_content, a plain raw .srt file's text.
A full file can take minutes to translate — this returns immediately:
{"job_id": 42, "status": "pending"}
Polling for the result
GET /api/external-translate/42
{
"id": 42,
"status": "done", // pending | running | done | failed
"source_lang": "en",
"target_lang": "ca",
"result_srt": "1\n00:00:00,000 --> 00:00:10,000\nSubtitlarr used AI...\n\n2\n...",
"engine_used": "gemini",
"model_used": "gemini-3.5-flash-lite",
"error": null,
"created_at": "...",
"finished_at": "..."
}
There is no push/callback mechanism yet — the caller is responsible for polling until status is done or failed, then doing whatever it needs with result_srt (e.g. uploading it to Bazarr itself, since Subtitlarr never does that on this path).
Every attempt also logs a job_events row (visible on the History page's Jobs tab as external_translate) with a short pass/fail summary, so there's at least a visible trail in the app even though the job itself has no other UI.
From an MCP-connected assistant
The same capability is exposed as two MCP tools — subtitlarr_submit_external_translate and subtitlarr_get_external_translate_job — for an assistant that wants Subtitlarr's own configured engine to do the translating (as opposed to the manual-translation tools, where the assistant translates it itself). These reuse the MCP session's own bearer token; no separate external-translate token is needed when calling through MCP.
Troubleshooting
Translations come back mostly untranslated, or with low cue recovery
This is usually a batch-size problem, not a translation-quality one. The batch token budget (set per engine instance) is not the same thing as that model's context window — a batch that technically fits inside the context window doesn't mean a smaller/weaker local model can reliably format that much structured output. Small models lose structural reliability on long responses well before hitting their real context limit. Try lowering the batch token budget on that engine instance before assuming anything else is wrong.
An item is stuck as "translating" after a restart
There's no checkpointing mid-item — if the app process is killed while an item is actively translating, it's automatically reset back to pending the next time the app starts, and gets fully retried later. Nothing is lost, but whatever GPU/API time was already spent on that attempt is.
A run stopped early with items still pending
Check the engine cascade — this is expected behavior when every instance above a separator becomes rate-limited or disabled mid-run, not a bug. The run stops rather than silently falling through to local engines fenced off behind a separator. Pick up the leftovers with a manual run, or wait for cooldowns/quota to reset.
A completed subtitle is in the wrong language
Run the Language Check job — this exact scenario (a well-formed response that's simply still in the source language) is what it exists to catch, since structural checks alone can't detect it.
Known limitations
- Doesn't download missing subtitles. Subtitlarr only translates a subtitle that already exists in some language — finding/downloading subtitles from providers is Bazarr's job, not this project's. An item with no existing subtitle in any language is skipped, not sourced; that needs speech-to-text, which is out of scope here.
- Doesn't re-time or fix subtitles with bad sync/timing issues. Translation is text-only — the original timestamps are always reused as-is; a subtitle that's out of sync before translation stays out of sync after.
- Doesn't come with a pre-configured LLM or API key. You bring your own engine — a local Ollama/llama.cpp server, or your own API key for Gemini/NVIDIA/OpenRouter/Groq. Nothing is bundled or pre-authorized.
- Subtitlarr never mounts your media library directly — everything goes through Bazarr's REST API. If Bazarr can't see a subtitle, neither can Subtitlarr.