Add a setup guide; fix the provider template

providers.template.yaml was a Raycast AI template. It named Raycast in its comments
and used api_keys/abilities/context — none of which provider_registry.py reads, while
omitting user_api_key and server_api_key_env, which it requires. Anyone copying it
got a file the application ignored. Replaced with one matching the code, covering
vLLM, Ollama, OpenRouter, Berget, OpenAI and Gemini.

docs/SETUP.md is written to be executed rather than skimmed: what the four external
dependencies are and which are optional, how to choose between providers, and a
verification command after every step that calls the real thing. Chat and tool calling
are checked separately, because a model can hold a conversation perfectly and still
never call a tool — which in this project means confident answers with no sources, the
one failure mode that matters most here.

Embeddings get their own section because they are a separate endpoint from chat and
conflating the two is the most common setup mistake, and because the dimension is
load-bearing: changing the model means re-embedding the whole corpus.

Ends with a symptom/cause/fix table covering the traps found during the production
cutover — the EnvironmentFile parsing difference, app.fts_config not surviving a
restore, and .env being unsourceable in bash.

Verified: every verification block in the guide was run against the live deployment.
main
Lasse Edfast 1 week ago
parent 8d29e73434
commit 238b46e277
  1. 5
      README.md
  2. 310
      docs/SETUP.md
  3. 189
      providers.template.yaml

@ -52,8 +52,11 @@ while the country's own word lives in `parliament.yaml`. See
## Quickstart ## Quickstart
Full step-by-step setup, including how to plug in each model provider and how to
verify each stage worked, is in **[docs/SETUP.md](docs/SETUP.md)**. The short version:
```bash ```bash
git clone https://git.edfast.se/plenum/plenum && cd plenum git clone https://git.edfast.se/lasse/plenum && cd plenum
cp .env.example .env # then fill in database and model settings cp .env.example .env # then fill in database and model settings
python -m venv .venv && .venv/bin/pip install -e ".[dev]" python -m venv .venv && .venv/bin/pip install -e ".[dev]"
``` ```

@ -0,0 +1,310 @@
# Setting up plenum
Written to be followed step by step, by a person or an assistant. Every step has a
command that proves it worked. Do not proceed past a failing check — later failures
will be confusing and unrelated-looking.
If you are an assistant: run the verification command after each step and report its
actual output. Do not infer success from a command exiting 0.
---
## 0. What plenum needs from you
Four external things. They are independent — you can swap any one without touching
the others.
| # | Thing | Why | Can you skip it? |
|---|---|---|---|
| 1 | **PostgreSQL 14+ with pgvector** | Stores everything; does full-text and vector search | No |
| 2 | **A chat model endpoint** | Powers chat and deep research | Search works without it; chat does not |
| 3 | **An embeddings endpoint** | Turns text into vectors for semantic search | Full-text search works without it; `vector_search` does not |
| 4 | **Source data** | A parliament's open data | No |
2 and 3 are **separate** and usually different servers. A common mistake is pointing
both at one URL and wondering why embedding fails.
---
## 1. PostgreSQL
Needs the `vector` extension (pgvector) and `pg_trgm`.
```bash
docker run -d --name plenum-pg \
-e POSTGRES_USER=plenum -e POSTGRES_PASSWORD=changeme -e POSTGRES_DB=plenum \
-p 5432:5432 pgvector/pgvector:pg16
```
Or on an existing server: `CREATE EXTENSION vector; CREATE EXTENSION pg_trgm;`
Set the text-search dictionary for your language. This is read by database triggers,
so it must be set on the database itself, not just in config:
```bash
psql -h localhost -U plenum -d plenum -c \
"ALTER DATABASE plenum SET app.fts_config = 'swedish';"
```
Pick from `SELECT cfgname FROM pg_ts_config;`. Postgres ships stemmers for about two
dozen languages. If yours is absent, use `simple` — search still works but without
stemming, so `kärnkraften` will not match a search for `kärnkraft`.
Apply the schema:
```bash
psql -h localhost -U plenum -d plenum -f _postgres/schema.sql
```
**Verify:**
```bash
psql -h localhost -U plenum -d plenum -c \
"SELECT count(*) AS tables FROM information_schema.tables WHERE table_schema='public';
SELECT current_setting('app.fts_config') AS fts;"
```
Expect 22 tables and your chosen dictionary. If `fts` errors with "unrecognized
configuration parameter", the `ALTER DATABASE` did not run or you reconnected before
it took effect — it applies to new sessions only.
---
## 2. The chat model
plenum talks to anything speaking the **OpenAI chat-completions protocol**. It never
uses a provider's native API.
### Choosing
| Situation | Use | Notes |
|---|---|---|
| You have a GPU | **vLLM** | Fastest. Handles concurrent requests properly. |
| No GPU, want it working today | **Ollama** | Easiest. Slow for deep research, fine for chat. |
| You want no infrastructure | **OpenRouter** | One key, many models. |
| Data must stay in the EU | **Berget AI** | Swedish provider, EU-hosted. |
| You already have OpenAI | **OpenAI** | Works; reasoning models are handled specially. |
The model needs **tool calling**. Without it, chat cannot search and will make things
up — which for this project is the one unacceptable failure. Verified working:
Qwen3 (8B and up), GPT-4o and later, Claude, Llama 3.3, Mistral Large.
Below ~7B parameters, tool calling gets unreliable. If chat answers without citing
sources, suspect the model before the code.
### Setting it up
**vLLM**
```bash
vllm serve Qwen/Qwen3-8B --port 8000
```
In `.env`:
```
LLM_DIRECT_URL=http://localhost:8000/v1
LLM_MODEL_SMART=Qwen/Qwen3-8B
LLM_MODEL_FAST=Qwen/Qwen3-8B
LLM_BEARER=
```
**Ollama**
```bash
ollama serve
ollama pull qwen3:8b
```
In `.env`:
```
LLM_DIRECT_URL=http://localhost:11434/v1
LLM_MODEL_SMART=qwen3:8b
LLM_MODEL_FAST=qwen3:4b
LLM_BEARER=
```
The `/v1` matters. Ollama's native API is on `/api` and is not OpenAI-compatible.
**A hosted provider as the server default**
```
LLM_DIRECT_URL=https://openrouter.ai/api/v1
LLM_BEARER=sk-or-v1-...
LLM_MODEL_SMART=anthropic/claude-sonnet-4
LLM_MODEL_FAST=anthropic/claude-haiku-4
```
`LLM_BEARER` is the server's own key. Everyone using your site spends it — only do
this behind authentication.
### Letting users bring their own key
Copy `providers.template.yaml` to `providers.yaml` and keep the providers you want.
Anything with `user_api_key: true` appears in the model picker and prompts the user
for a key. That key is held for the request only, and is never written to the database
or logs; the sole persisted copy is encrypted in the user's own browser.
**Verify — this calls the model for real:**
```bash
.venv/bin/python -c "
from dotenv import load_dotenv; load_dotenv()
import os
from packages.llm import LLM
llm = LLM(base_url=os.getenv('LLM_DIRECT_URL'), model=os.getenv('LLM_MODEL_SMART'),
api_key=os.getenv('LLM_BEARER') or None, silent=True)
r = llm.generate(messages=[{'role':'user','content':'Reply with the word OK.'}], think=False)
print('FAILED:', r) if isinstance(r, str) else print('OK:', r.content[:60])
"
```
A string starting `LLM request failed:` means it did not work — the message names the
cause. `OK:` followed by text means the endpoint, model name and key are all correct.
**Verify tool calling separately** — a model can chat fine and still not call tools:
```bash
.venv/bin/python -c "
from dotenv import load_dotenv; load_dotenv()
import os
from packages.llm import LLM, register_tool, get_tools
@register_tool
def get_population(country: str) -> str:
'''Return a country's population.
Args:
country: Country name.
'''
return f'{country}: 10.4 million'
llm = LLM(base_url=os.getenv('LLM_DIRECT_URL'), model=os.getenv('LLM_MODEL_SMART'),
api_key=os.getenv('LLM_BEARER') or None, tools=get_tools(['get_population']), silent=True)
llm.generate(messages=[{'role':'user','content':'Use the tool: population of Sweden?'}], think=False)
calls = [m for m in llm.messages if m.get('role') == 'tool']
print('TOOL CALLING WORKS:', calls) if calls else print('NO TOOL CALL — pick a different model')
"
```
---
## 3. Embeddings
Separate from chat, and configured separately. Also OpenAI-compatible.
```
EMBEDDING_BASE_URL=http://localhost:8003/v1
EMBEDDING_API_KEY=
LLM_MODEL_EMBEDDING=qwen3-embedding
```
**The dimension must match the database.** `parliament.yaml` declares
`embeddings.dimension: 384`, and the `vector(N)` columns are that width. Change the
model and you must change both, then re-embed the entire corpus — there is no
converting existing vectors.
plenum checks this at startup and refuses to run on a mismatch, because the failure
otherwise surfaces deep inside pgvector with a message that never mentions
configuration.
Options:
| Approach | Command |
|---|---|
| vLLM | `vllm serve Qwen/Qwen3-Embedding-0.6B --port 8003` |
| Ollama | `ollama pull nomic-embed-text``EMBEDDING_BASE_URL=http://localhost:11434/v1`, dimension **768** |
| OpenAI | `EMBEDDING_BASE_URL=https://api.openai.com/v1`, model `text-embedding-3-small`, dimension **1536** |
If you change the dimension, edit `parliament.yaml` **before** creating the schema —
`schema.sql` reads it. On an existing database you must alter every `vector(N)` column
and re-embed.
**Verify:**
```bash
.venv/bin/python -c "
from dotenv import load_dotenv; load_dotenv()
from postgres_client import pg
from parliament import PARLIAMENT
v = pg.make_embeddings(['test'])[0]
print(f'got {len(v)} dims, config expects {PARLIAMENT.embeddings.dimension}')
print('MATCH' if len(v) == PARLIAMENT.embeddings.dimension else 'MISMATCH — fix before ingesting')
"
```
---
## 4. Source data
For Sweden, everything is already configured:
```bash
.venv/bin/python -m ingest.cli fetch --source documents --range 2022-2025
.venv/bin/python -m ingest.cli load --source documents
.venv/bin/python scripts/make_embeddings.py documents
```
Start with one range. A full corpus is tens of GB and takes hours.
For another parliament, see [PORTING.md](PORTING.md) — you write a config file and an
adapter.
**Verify:**
```bash
psql -h localhost -U plenum -d plenum -c "
SELECT count(*) AS documents, count(search_vector) AS searchable FROM documents;"
```
Both numbers should be equal and non-zero. If `searchable` is 0, the search-vector
trigger did not fire — check `app.fts_config` from step 1.
---
## 5. Run it
```bash
.venv/bin/uvicorn backend.app:app --port 8000 # API
cd frontend && npm install && npm run dev # UI on :5173
```
**Verify the whole stack:**
```bash
curl -s localhost:8000/api/meta | head -c 200
curl -s -X POST localhost:8000/api/search -H 'Content-Type: application/json' \
-d '{"q":"test","limit":3}' | head -c 200
```
---
## When something fails
| Symptom | Cause | Fix |
|---|---|---|
| `LLM request failed: ... unreachable` | Wrong `LLM_DIRECT_URL`, or the server is down | `curl $LLM_DIRECT_URL/models` |
| `LLM request failed: ... 401` | Missing or wrong `LLM_BEARER` | Check the key; self-hosted usually needs none |
| Chat answers but never cites sources | Model cannot call tools | Run the tool-calling check in §2 |
| `column X does not exist` | Schema older than the code | Apply `_postgres/migrations/*.sql` in order |
| Search returns nothing, no error | `app.fts_config` wrong or unset | `SELECT current_setting('app.fts_config');` then §1 |
| `expected N dimensions, got M` | Embedding model changed | Match `parliament.yaml` to the model, re-embed |
| Startup: `app.fts_config is 'x' but parliament.yaml declares 'y'` | The two disagree | Make them match; the database wins |
| `No LLM endpoint configured` | `LLM_DIRECT_URL` unset | Set it in `.env` |
| Service starts, then model calls fail with an empty model name | `EnvironmentFile=` in a systemd unit | Remove it. systemd's parser rejects `KEY =` with a space; let python-dotenv read `.env` |
| `.env` values missing in a shell script | `source .env` aborts on unquoted parentheses | Use python-dotenv, not `source` |
## Notes that save time later
**Nothing here is global.** Chat provider, embedding provider and database are three
independent choices. Changing the chat model needs no re-indexing; changing the
embedding model needs a full re-embed.
**Prompts are files.** `prompts/<lang>/*.md`. Set `PROMPTS_RELOAD=1` to re-read them
per call and iterate without restarting.
**Keep deployment values out of the repo.** `PARLIAMENT_CONFIG`, `PROMPTS_DIR` and
`CONTENT_DIR` each point somewhere else on disk, so a deployment's own settings never
appear in a diff against upstream.
**`database_query` runs model-written SQL.** Two layers restrict it to reads, but give
it a `SELECT`-only database role anyway. See [SECURITY.md](../SECURITY.md).

@ -1,113 +1,92 @@
# Template for providers.yaml # Copy to providers.yaml and edit. providers.yaml is gitignored.
# #
# You can add any OpenAI API compatible provider to the "providers" list. # This file lists the chat models a user can pick in the web UI. It does NOT
# For each provider you must also specify a list of models, along with model abilities. # configure embeddings — those are separate and set via environment variables.
# See docs/SETUP.md.
# #
# All fields are required unless marked as optional. # Every provider must speak the OpenAI chat-completions protocol. That includes
# vLLM, Ollama, llama.cpp, LM Studio, OpenRouter, Berget, OpenAI, Together,
# Groq, and Google Gemini's compatibility endpoint.
# #
# Refer to your provider's API documentation for specific # Fields the application reads — anything else is ignored:
# details such as model identifiers, capabilities etc
#
# Note: Since the OpenAI API is not a standard we can't guarantee that all
# providers will work correctly with Raycast AI.
#
# To use this template rename as `providers.yaml`
# #
# id required. Stable identifier used in API requests.
# name required. Shown in the model picker.
# base_url required. Include the /v1 suffix if the provider uses one.
# user_api_key required. true = the user supplies their own key in the browser.
# false = the server's key is used, from server_api_key_env.
# server_api_key_env optional. NAME of an environment variable holding the key.
# Never put the key itself here.
# supports_thinking optional, default false. Whether the model can emit reasoning.
# models optional. Assigns roles; without it the app falls back to
# LLM_MODEL_SMART / LLM_MODEL_FAST from .env.
# role: smart -> orchestration and final answers
# role: fast -> summarising tool output; a smaller model is fine
providers: providers:
- id: perplexity # ── Self-hosted, no API key ────────────────────────────────────────────────
name: Perplexity # vLLM: fastest option on a machine with a GPU.
base_url: https://api.perplexity.ai # vllm serve Qwen/Qwen3-8B --port 8000
# Specify at least one api key if authentication is required. - id: vllm
# Optional if authentication is not required or is provided elsewhere. name: Local vLLM
# If individual models require separate api keys, then specify a separate `key` for each model's `provider` base_url: http://localhost:8000/v1
api_keys: user_api_key: false
perplexity: PERPLEXITY_KEY supports_thinking: true
# Optional additional parameters sent to the `/chat/completions` endpoint
additional_parameters:
return_images: true
web_search_options:
search_context_size: medium
# Specify all models to use with the current provider
models: models:
- id: sonar # `id` must match the identifier used by the provider - id: Qwen/Qwen3-8B
name: Sonar # name visible in Raycast role: smart
provider: perplexity # Only required if mapping to a specific api key - id: Qwen/Qwen3-4B
description: Perplexity AI model for general-purpose queries # optional role: fast
context: 128000 # refer to provider's API documentation
# Optional abilities - all child properties are also optional. # Ollama: easiest option, runs on CPU, slower.
# If you specify abilities incorrectly the model may fail to work as expected in Raycast AI. # ollama serve && ollama pull qwen3:8b
# Refer to provider's API documentation for model abilities. # The /v1 suffix matters — Ollama's native API is not OpenAI-compatible, its /v1 is.
abilities: - id: ollama
temperature: name: Ollama
supported: true base_url: http://localhost:11434/v1
vision: user_api_key: false
supported: true supports_thinking: false
system_message:
supported: true
tools:
supported: false
reasoning_effort:
supported: false
- id: sonar-pro
name: Sonar Pro
description: Perplexity AI model for complex queries
context: 200000
abilities:
temperature:
supported: true
vision:
supported: true
system_message:
supported: true
# provider with multiple api keys
- id: my_provider
name: My Provider
base_url: http://localhost:4000
api_keys:
openai: OPENAI_KEY
anthropic: ANTHROPIC_KEY
models:
- id: gpt-4o
name: "GPT-4o"
context: 200000
provider: openai # matches "openai" in api_keys
abilities:
temperature:
supported: true
vision:
supported: true
system_message:
supported: true
tools:
supported: true
- id: claude-sonnet-4
name: "Claude Sonnet 4"
context: 200000
provider: anthropic # matches "anthropic" in api_keys
abilities:
temperature:
supported: true
vision:
supported: true
system_message:
supported: true
tools:
supported: true
- id: litellm
name: LiteLLM
base_url: http://localhost:4000
# No `api_keys` - authentication is provided by the LiteLLM config
models: models:
- id: anthropic/claude-sonnet-4-20250514 - id: qwen3:8b
name: "Claude Sonnet 4" role: smart
context: 200000 - id: qwen3:4b
abilities: role: fast
temperature:
supported: true # ── Hosted, user brings their own key ──────────────────────────────────────
vision: # The key is entered in the browser, held only for the request, and never
supported: true # written to the database or to logs.
system_message: - id: openrouter
supported: true name: OpenRouter
tools: base_url: https://openrouter.ai/api/v1
supported: true user_api_key: true
supports_thinking: true
- id: openai
name: OpenAI
base_url: https://api.openai.com/v1
user_api_key: true
supports_thinking: true
# Berget AI — Swedish provider, EU-hosted. Relevant where data residency matters.
- id: berget
name: Berget AI
base_url: https://api.berget.ai/v1
user_api_key: true
supports_thinking: true
# Google Gemini via its OpenAI-compatibility endpoint. The trailing slash matters.
- id: googlegemini
name: Google Gemini
base_url: https://generativelanguage.googleapis.com/v1beta/openai/
user_api_key: true
supports_thinking: true
# ── Hosted, paid for by the server ─────────────────────────────────────────
# With user_api_key: false, every user spends your credit. Only do this behind
# authentication.
# - id: openai-server
# name: OpenAI (server key)
# base_url: https://api.openai.com/v1
# user_api_key: false
# server_api_key_env: OPENAI_API_KEY
# supports_thinking: true

Loading…
Cancel
Save