MANUAL · 11
The AI DJ can run on a small model on your own hardware or a large hosted one. This page tells you which models actually hold up — measured, not guessed — and which settings match the station to whichever you pick.
THE ROOT CHOICE
Every word the DJ speaks and every track it picks comes from one language model, chosen under Admin → LLM. The default is Ollama on your own hardware (no API key, no per-token bill), but you can point the station at a hosted provider (Anthropic, OpenAI, Google, OpenRouter and others) instead. Switching reroutes every call immediately, with no redeploy.
“On your own hardware” isn’t only Ollama — there are three local paths, all keyless and all private to your box:
:cloud tags) also ride this path: same setup, the heavy lifting happens on their hardware.locca serve <model>) built on llama.cpp. No key, a sensible host default, and the onboarding wizard can detect it for you. locca on GitHub ↗One thing to internalise before choosing: the provider is part of the choice. The same model can behave differently through different routes, because each provider translates tools and structured output its own way — a model that fails through one route can be flawless through another. When you evaluate a model, evaluate it through the provider you’ll actually run.
MEASURED, NOT GUESSED
SUB/WAVE ships a benchmark that drives every kind of call the DJ makes — track picks, talk segments, listener requests, scripts, banter, programme plans — against any model, in both picker modes, and scores the output against the station’s own rules. The table below is the running record — it grows as more models and providers get benched (station configured lean unless noted):
| Model | Verdict | Overall | Picks · pool | Picks · agent | Segments · pool | Segments · agent | Requests | Scripts | Shows | Pick p50 | Benched |
|---|---|---|---|---|---|---|---|---|---|---|---|
| gemma4:31b-cloudOllama | agent-capable | 99% | 100% | 100% | 100% | 100% | 100% | 100% | 96% | 1.2s | 2026-07-10 |
| gemini-3.1-flash-litegoogle | agent-capable | 99% | 100% | 100% | 100% | 100% | 100% | 100% | 96% | 0.7s | 2026-07-10 |
| gemini-3.5-flashgoogle | agent-capable | 98% | 100% | 100% | 100% | 100% | 93% | 100% | 96% | 1.2s | 2026-07-10 |
| openai/gpt-4o-miniOpenRouter | agent-capable | 97% | 92% | 100% | 100% | 100% | 100% | 100% | 92% | 1.3s | 2026-07-10 |
| openai/gpt-5-miniOpenRouter | agent-capable | 96% | 75% | 100% | 100% | 100% | 93% | 100% | 100% | 1.8s | 2026-07-10 |
| google/gemma-4-26b-a4b-itopenai-compatible | agent-capable | 96% | 75% | 100% | 100% | 100% | 100% | 100% | 96% | 3.6s | 2026-07-10 |
| google/gemma-4-26b-a4b-itOpenRouter | pool-mode | 95% | 75% | 100% | 100% | 100% | 100% | 100% | 92% | 4.6s | 2026-07-10 |
| qwen/qwen3.5-9bOpenRouter | pool-mode | 90% | 75% | 50% | 92% | 100% | 93% | 100% | 92% | 3.8s | 2026-07-10 |
| deepseek-v4-flashdeepseek | pool-mode | 89% | 100% | 0% | 100% | 100% | 93% | 95% | 88% | 1.5s | 2026-07-10 |
| nemotron-3-super:cloudOllama cloud | agent-capable | 86% | 75% | 100% | 100% | 83% | 100% | 95% | 67% | 1.8s | 2026-07-10 |
| gemma-4-12b-it-Q4_K_Mlocca (llama.cpp) | pool-mode | 86% | 75% | — | — | — | 100% | 100% | 75% | 16.1s | 2026-07-09 |
| qwen/qwen3.5-9bopenai-compatible | prefer native route | 84% | 75% | 50% | 100% | 100% | 87% | 90% | 79% | 1.2s | 2026-07-10 |
| anthropic/claude-haiku-4.5OpenRouter | agent-capable | 83% | 100% | 100% | 100% | 100% | 100% | 100% | 33% | 2.3s | 2026-07-10 |
| glm-5.2:cloudOllama cloud | pool-mode | 83% | 92% | 83% | 100% | 100% | 80% | 100% | 54% | 5.9s | 2026-07-10 |
| kimi-k2.6:cloudOllama cloud | avoid | 80% | 100% | 0% | 100% | 100% | 93% | 100% | 50% | 2.1s | 2026-07-10 |
| deepseek-v4-flash:cloudOllama cloud | pool-mode | 78% | 42% | 17% | 100% | 100% | 87% | 95% | 75% | 3.6s | 2026-07-10 |
| minimax-m2.7:cloudOllama cloud | avoid (re-test) | 63% | 8% | 17% | 100% | 83% | 80% | 86% | 46% | 18s | 2026-07-10 |
Reading it for a recommendation: pick an agent-capable row if you want the full conversational picker — Gemini 3.5 Flash leads the table outright, its Flash-Lite sibling matches the 31B class at the fastest picks benched, and Gemma 4 31B on Ollama cloud remains the best keyless option. Pick any healthy pool-mode row for a lean station (Qwen3.5 9B is the small floor, and a local Gemma 4 12B — locca serve gemma4 — does the same job keylessly on your own box). Remember the route in the second line of each model cell is part of the result — the same model through a different provider can score differently, and two of the scores above changed by 20+ points once bugs in our own thinking-suppression plumbing were found and fixed. The bench checks that too, now.
Two patterns worth knowing whatever you run: the Gemma family at every size shares the same habits (it can repeat an artist when the shortlist pressures it to, and it fumbles feature choices on three-hour programme plans), and the multi-hour programme plan is the hardest single call in the system — the only one that dented every model tested. If a show misbehaves, suspect the plan before the model.
Running from a clone? You can put any candidate model through the same battery before trusting it on air: npm run llm-bench in controller/benchmarks it across every call kind and prints a comparison table. The DJ Doctor’s LLM checks cover the everyday health of whatever you’ve picked.
RUNNING LEAN
If you’re on a modest local model, or paying per token and want the bill low, these are the dials to turn down. None of them take the DJ off the air. They just make it do less work per moment.
With the lean profile in place, Qwen3.5 9B or a local Gemma 4 12B runs the whole station comfortably — picks, requests, talk breaks, even programme shows — while paying nothing per token.
RUNNING RICH
On a capable model the same dials go the other way: spend the capability on a station with more personality and a smarter DJ.
The picker agent has a built-in safety net: if it ever fails or runs too slow, the station quietly falls back to the simple pool picker for that track, the same path you’d get with the agent switched off. Turning it off just makes that lighter path the default rather than the exception.
A SECOND, SMALLER MODEL
The DJ picks partly by mood— mellow mornings, brighter afternoons, a wind-down late at night. To know each track’s mood it leans on the library tagger, which uses a second, much smaller embedding model — not the chat model that writes the show.
Rather than ask the chat model about every track (slow and expensive on a big library), the tagger embeds each track once, has the chat model tag a small, representative seed set, then propagatesmoods and energy out to everything else by similarity. That’s roughly ten times fewer model calls than tagging track by track.
By default the embedding model follows your LLM provider, so there’s usually nothing extra to set up — an Ollama-local station gets nomic-embed-text for free. Two things are worth knowing if you stray from that:
deepseek and Vercel AI gateway providershave no embeddings endpoint at all. A DJ on one of those works fine, but the tagger can’t follow it, so the console only lists embedding-capableproviders in the tagger dropdown (Ollama, OpenAI, Google, OpenRouter, locca, OpenAI-compatible). If you don’t see your chat provider there, that’s why — pick Ollama (local and free) for the embedding step and leave the DJ where it is.openai/text-embedding-3-small. OpenRouter, Requesty and the like carry everything (chat and embeddings); the bare provider named after a chat-only company does not.locca embed, on its own port; the console can detect it for you.nomic-embed-text (local, free, 768-d) if you run Ollama, or text-embedding-3-small (cloud, cheap, 1536-d) otherwise. The exact model matters far less than picking one and sticking with it — see the next note.One catch worth internalising:the vector index is built at your embedding model’s dimension, so changing the embedding model means re-embedding the whole library(Admin → Library tagger → Re-scan → “Re-embed all tracks”). Changing the chatmodel never needs this — but if embeddings are set to “follow the LLM,” switching your DJ providerquietly changes the embedding model too. The console pins embeddings to your library’s model and warns you before that happens, so the safe move is to pin an embedding provider once and leave it.
It all lives under Admin → Library tagger, and you can see the tagged library laid out in Library Observatory.
WHERE TO SET THEM
Every setting here is in the admin console and takes effect without a redeploy; most apply to the next thing the DJ does. The full tour of the console is in Admin & Settings; how the DJ actually picks and talks is in How the DJ Works.