docs/eval-scorer.md was notes-to-self for work already finished. It told the reader to
add a class to eval_harness.py that has been there for months, specified the author's
own GPU by model number, and carried escaped-markdown artifacts from a bad paste. The
feature is real and was used — 9 eval runs, 668 questions, 3536 judgments — so the fix
is documentation, not deletion.
It now explains what coverage scoring measures and why it is worth having alongside
the judge model: the judge catches claims a source contradicts, the cross-encoder
catches claims a source simply does not cover. Includes how to read the number, and
the query that finds the interesting cases — paragraphs the judge passed but the
scorer did not, which is where technically-defensible-but-misleading answers show up.
The scorer defaulted to port 8001, which is also the MCP server's default, so running
both meant one silently failed to bind. Moved to 8005 and documented.
SCORER_ENDPOINT is now in .env.example, and the README has a documentation index —
every file under docs/ was previously unreachable from anywhere in the repo.
talks -> speeches ("talks" reads as conference talks to everyone outside this
project), motions -> documents with a new doc_type column so bills, written
questions and committee reports can share the table later, and 79 columns from
anforandetext -> text, intressent_id -> person_id, valkrets -> constituency,
lydelse -> text, and so on.
_postgres/rename_map.py is the single source of truth. The migration, its rollback,
and the code rewrite are all derived from it, so they cannot drift apart. Hand-
writing a rollback is how you end up with one that fails halfway through.
Two identifiers could not be renamed mechanically and were done by reading the
queries: dok_id means the protocol document in talks but the primary key in
motions, and year stays a calendar year in speeches while becoming session_year in
documents. The latter is aliased in SQL so the JSON field stays `year` and the
frontend contract is unchanged.
Values are never translated. Bifall and Avslag stay as published; parliament.yaml
glosses them. A research tool must not silently rewrite the record.
The migration is guarded by an existence check, so the same file is a no-op on a
fresh database and does the work on an existing one — one schema definition in the
world. ALTER TABLE ... RENAME is catalog-only, so the millions of HNSW-indexed
vectors are untouched.
The trigger functions are recreated explicitly, because plpgsql bodies are stored
as opaque text and do not follow renames: they would have compiled fine and then
failed at the next INSERT. They now read the text-search config from a database
setting rather than hardcoding 'swedish'.
Verified: migration round-trips to a byte-identical schema across columns, indexes
and triggers; re-running is a no-op; triggers repopulate search_vector with working
Swedish stemming on INSERT and UPDATE; and the renamed code runs real searches
against a migrated database holding 5,000 rows of production data, with plain,
prefix and exclusion syntax all working. Frontend type errors went from 9 to 7 —
the rename fixed two and introduced none.
Not yet done, and tracked: legacy compatibility views, the shim for SQL replayed
from saved snapshots, tool-name aliases, and updating docs to the new names.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
_llm called env_manager.set_env() at import time, which connected to a private
ArangoDB to fetch secrets. That single line meant a fresh clone could not start,
regardless of what else was configured. Both packages also lived in separate
private repos and were gitignored here, so the code shipped without them.
packages/llm/ is 780 lines against _llm's 1750. Dropped as unused by this project
(measured, zero call sites): token counting and message trimming, image/vision
handling, make_summary, the ollama-specific paths, the query/user_input/context
argument style, and the self-mutating provider_quirks.json cache.
Kept and reworked:
- tools.py, the docstring -> JSON-schema tool registry, which has no equivalent
in the cuj-fup client and which llm_tools.py depends on entirely.
- The provider quirks that actually matter: vLLM-only extra_body fields stripped
for hosted providers, enable_thinking disabled at template level when think is
off, reasoning models (o1/o3/o4/gpt-5) switched to max_completion_tokens.
Adopted from cuj-fup's client: LLMConfig as a dataclass instead of 20 constructor
kwargs, the SDK's native max_retries instead of hand-rolled backoff, and error
messages that name the likely cause.
Fixes a latent bug: Optional[list[str]] parameters were advertised to the model as
strings, because get_origin(Optional[X]) is Union, so neither the schema mapping
nor the list coercion in execute_tool fired. `parties`, `people` and `focus_ids`
were all affected.
Also drops a dead SELECT-only guard in execute_tool that keyed on a parameter name
(`sql_query`) that no tool has ever used. Real SQL hardening is tracked separately.
Verified: `import backend.app` succeeds with all external network blocked and zero
outbound connection attempts; all 12 tools register; live vLLM calls confirmed for
plain generation, structured output via format=, tool execution, and the
error-returns-a-string contract that call sites branch on.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Seeded via an explicit allow-list (see /home/lasse/plenum-seed.sh) rather than by
deleting files from a copy, so nothing sensitive can survive by omission.
Excluded: WireGuard backup + client config, the plaintext DB password in admin.py,
Arango credentials in scripts/notes.md, .claude/settings.json, a 113 MB log,
providers.yaml (private endpoint), the Arango/ChromaDB-era scripts, the duplicated
claude-design-system frontend copy, and assorted screenshots and one-off planning docs.
297 tracked files / 49 MB of history -> 161 files / 2.1 MB.
Recovered 14 database migrations that the old .gitignore's `*.sql` rule had been
hiding from version control.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>