13 KiB
Guide for adding multi-provider LLM support in rixdagen
This document is intended to be pasted into an assistant session in VS Code so it can implement the change.
What is true in the repo right now
The current code is still wired around one global endpoint, even though the repo also contains a more flexible provider config template.
Verified repo facts
backend/services/chat.pycreatessmart_llm,fast_llm, andcommunicator_llmusing the samebase_url=os.getenv("LLM_DIRECT_URL").FAST_MODELandSMART_MODELare selected from environment variables (LLM_MODEL_FAST,LLM_MODEL_SMART), but both still flow through that same single endpoint.providers.template.yamlalready describes a multi-provider design where each provider has:- an
id - a
name - a
base_url - optional
api_keys - a list of
models - optional
additional_parameters - per-model
abilities
- an
providers.template.yamlexplicitly says you can add any OpenAI API compatible provider to theproviderslist, and also notes that the OpenAI API is not a standard, so compatibility is not guaranteed.- The template also already shows a provider with multiple API keys and a
litellmexample, which means the intended design already assumes routing different models to different upstreams.
Main conclusion
The repo already points toward the right architecture, but the runtime path is not there yet.
Right now, adding more entries to providers.yaml alone is not enough. The backend still behaves like this:
- choose model alias from env
- send all LLM traffic to one global
LLM_DIRECT_URL
To support vLLM, berget.ai, openai.com, and potentially other OpenAI-compatible providers cleanly, the app needs to move to this model:
- choose a logical model alias for each role (
smart,fast,embeddings, maybecommunicator) - resolve that alias through
providers.yaml - derive:
- provider
- protocol
- base URL
- API key env var
- upstream model ID
- abilities / extra parameters
- build the client per resolved model
What needs to be implemented
1. Make providers.yaml the source of truth for model routing
Do not keep endpoint selection centered on LLM_DIRECT_URL.
Instead:
- keep environment variables for which logical model alias to use
- move endpoint/model/provider metadata into
providers.yaml
Recommended env vars:
LLM_MODEL_SMARTLLM_MODEL_FASTLLM_MODEL_EMBEDDINGSLLM_MODEL_COMMUNICATOR(optional, if you want it to diverge later)- provider secrets such as
OPENAI_API_KEY,BERGET_API_KEY,LLM_VLLM_API_KEY, etc.
2. Add a provider resolver layer
Create a small module, for example backend/services/llm_provider_registry.py or _llm/provider_registry.py, with responsibilities like:
- load
providers.yaml - validate it
- index models by alias or logical name
- resolve one requested model into a runtime configuration object
Recommended shape:
@dataclass
class ResolvedModel:
alias: str
provider_id: str
provider_name: str
protocol: str
base_url: str
api_key: str | None
api_key_env: str | None
upstream_model_id: str
abilities: dict[str, Any]
additional_parameters: dict[str, Any]
Functions to add:
def load_provider_config(path: str = "providers.yaml") -> dict: ...
def resolve_model(alias: str) -> ResolvedModel: ...
def get_api_key_from_env(env_name: str | None) -> str | None: ...
3. Refactor _llm.LLM to use resolved provider config
Today the calling code passes model=... and base_url=... directly.
Change that so _llm.LLM can accept either:
- a logical alias and resolve internally, or
- an already-resolved provider config object
Preferred direction:
LLM(model_alias="smart")
or:
resolved = resolve_model(os.getenv("LLM_MODEL_SMART", "smart"))
LLM(resolved_model=resolved)
The key point is that _llm.LLM.generate(...) should no longer rely on one app-wide endpoint.
4. Remove the single-endpoint assumption from ChatService
Refactor backend/services/chat.py so each role resolves independently.
Current behavior:
smart_llm->LLM_DIRECT_URLfast_llm->LLM_DIRECT_URLcommunicator_llm->LLM_DIRECT_URL
Target behavior:
smart_llm-> resolved fromLLM_MODEL_SMARTfast_llm-> resolved fromLLM_MODEL_FASTcommunicator_llm-> resolved fromLLM_MODEL_COMMUNICATORor fallback toLLM_MODEL_SMART
So the code should look more like:
smart_alias = os.getenv("LLM_MODEL_SMART", "smart")
fast_alias = os.getenv("LLM_MODEL_FAST", smart_alias)
communicator_alias = os.getenv("LLM_MODEL_COMMUNICATOR", smart_alias)
self.smart_llm = LLM(model_alias=smart_alias, system_message=ORCHESTRATOR_SYSTEM, temperature=0.2)
self.fast_llm = LLM(model_alias=fast_alias, system_message=WORKER_SYSTEM, temperature=0.05)
self.communicator_llm = LLM(model_alias=communicator_alias, system_message=ORCHESTRATOR_SYSTEM, temperature=0.3)
5. Keep a fallback path for the current vLLM setup
Do not break the current local workflow.
Recommended fallback logic:
- if
providers.yamlexists and the requested alias resolves, use it - otherwise fall back to the current legacy env path:
LLM_DIRECT_URLLLM_API_KEYLLM_MODEL
That lets the migration happen incrementally.
6. Distinguish protocol from provider
This matters for future-proofing.
Do not assume every upstream will behave exactly like local vLLM.
Add a provider field like:
protocol: openai
Even if you only implement openai first.
That keeps the design open for future support of providers that may require slightly different auth, payloads, streaming behavior, or tool-calling semantics.
Recommended providers.yaml structure
Use the existing repo template as the starting point and extend it slightly.
Suggested structure:
providers:
- id: vllm
name: Local vLLM
protocol: openai
base_url: http://localhost:8000/v1
api_keys:
vllm: LLM_VLLM_API_KEY
models:
- id: meta-llama/Meta-Llama-3.1-70B-Instruct
alias: smart_local
name: Smart Local
provider: vllm
context: 131072
abilities:
temperature:
supported: true
system_message:
supported: true
tools:
supported: true
- id: openai
name: OpenAI
protocol: openai
base_url: https://api.openai.com/v1
api_keys:
openai: OPENAI_API_KEY
models:
- id: gpt-4.1-mini
alias: fast_openai
name: GPT-4.1 mini
provider: openai
context: 1000000
abilities:
temperature:
supported: true
system_message:
supported: true
tools:
supported: true
- id: berget
name: Berget
protocol: openai
base_url: REPLACE_WITH_VERIFIED_BERGET_BASE_URL
api_keys:
berget: BERGET_API_KEY
models:
- id: REPLACE_WITH_VERIFIED_MODEL_ID
alias: smart_berget
name: Berget Smart
provider: berget
context: REPLACE_WITH_VERIFIED_CONTEXT
abilities:
temperature:
supported: true
system_message:
supported: true
tools:
supported: REPLACE_WITH_VERIFIED_VALUE
Important implementation details
Alias resolution
The app should refer to models by local aliases, not raw provider model IDs.
That means env vars should point to aliases such as:
LLM_MODEL_SMART=smart_berget
LLM_MODEL_FAST=fast_openai
LLM_MODEL_COMMUNICATOR=smart_local
API key lookup
The provider file should contain the env var name, not the secret itself.
Example:
api_keys:
openai: OPENAI_API_KEY
Runtime should do:
env_name = provider.api_keys[model.provider]
api_key = os.getenv(env_name)
Additional provider parameters
The template already supports additional_parameters at provider level.
Preserve and pass those through when calling the upstream API. This is especially useful because many OpenAI-compatible providers expose non-standard extensions.
Abilities gating
Use the abilities block to avoid sending unsupported settings.
Examples:
- do not send tools if
tools.supportedis false - do not send temperature if unsupported
- do not rely on reasoning-specific settings unless the model explicitly supports them
This is important because the template itself warns that OpenAI compatibility is imperfect.
Streaming and tool calling
You should assume that compatibility varies most in these areas:
- streaming chunk format
- tool/function calling schema
- system message handling
- vision payload format
- reasoning fields
So build adapters conservatively and fail loudly with clear messages.
Suggested migration plan
Phase 1: provider registry
- add loader + schema validation for
providers.yaml - resolve aliases to provider configs
- keep legacy env fallback
Phase 2: wire ChatService
- replace
LLM_DIRECT_URLusage inbackend/services/chat.py - resolve each model independently
Phase 3: update _llm
- centralize client creation
- add support for provider-level
additional_parameters - add abilities filtering
Phase 4: test with three providers
Test combinations like:
-
all-local
- smart -> vLLM
- fast -> vLLM
- communicator -> vLLM
-
mixed
- smart -> Berget
- fast -> OpenAI
- communicator -> vLLM
-
cloud-only
- smart -> OpenAI
- fast -> OpenAI
- communicator -> OpenAI
Acceptance criteria
The change is done when all of the following are true:
- I can switch providers without code changes, only by editing
providers.yamland env vars. smart,fast, andcommunicatorcan point to different providers.- The current local vLLM setup still works.
- Unsupported capabilities are filtered based on model abilities.
- Failures clearly say which alias/provider/model failed and why.
- Secrets are only read from environment variables, never committed into YAML.
External documentation notes that must be verified before implementation
I cannot verify the live OpenAI and Berget documentation from this environment because web access is disabled in this session.
That means the VS Code assistant should verify these exact facts from the docs before finalizing the patch:
OpenAI
Verify from current OpenAI docs:
- the recommended Python SDK import and client initialization pattern
- whether the intended integration path should use
responsesorchat.completions - exact base URL shape for the public API
- auth header expectations
- tool-calling support for the models you intend to use
- streaming response shape if this app streams tokens
Berget
Verify from https://docs.berget.ai/quickstart:
- the exact OpenAI-compatible
base_url - the exact auth mechanism / env var examples
- the exact Python example for OpenAI SDK usage
- which model IDs are documented for chat usage
- whether tools/function calling are supported
- whether reasoning / vision / embeddings are separately documented
One OpenAI-compatible provider in general
Use one documented compatibility layer as a reference point when implementing adapters. The goal is to confirm what the minimum common denominator really is across providers.
At minimum, verify:
- expected
/v1pathing - supported request schema for chat completions or responses
- tool/function call representation
- streaming format
- how unsupported fields are handled
Instructions for the coding assistant
Implement this as a minimal, low-risk refactor.
Required changes
- Add a provider registry module.
- Add YAML loading and validation.
- Refactor
backend/services/chat.pyso each LLM instance resolves independently. - Refactor
_llm.LLMto build clients from resolved provider config. - Preserve backwards compatibility with the current env-only setup.
- Add clear error messages for missing aliases, missing API keys, and unsupported features.
Nice-to-have changes
- Add a startup validation command that prints configured aliases and providers.
- Add unit tests for alias resolution.
- Add one integration test for local vLLM and one mocked OpenAI-compatible provider.
- Add comments explaining why compatibility must be treated as provider-specific, not assumed.
Suggested commit breakdown
add provider registry and providers.yaml loaderrefactor chat service to resolve llm aliases per rolerefactor _llm client creation to use provider configadd compatibility fallback for legacy env configadd tests for provider resolution and errors
Final note
The most important design change is this:
Stop routing all models through one global LLM_DIRECT_URL.
Everything else is secondary. Once alias -> provider -> endpoint resolution exists, adding berget.ai, openai.com, or other OpenAI-compatible providers becomes a configuration problem instead of a code fork.