packages/llm/client.py:_read_stream yielded token-by-token deltas but was
dead code — nothing called generate(stream=True). The chat tool loop always
blocked on the final answer generation, so a ~40s answer arrived as a single
SSE event and read as a hang, even though the "card stack" already streamed
tool-selection progress.
The final answer is structurally just whichever tool-loop iteration returns
content with no tool_calls — there's no way to know in advance which
iteration that'll be, and a turn can carry short narration before a tool
call in the very same response. So streaming has to be speculative:
- packages/llm/client.py: _read_stream becomes StreamAccumulator, which both
streams live (thinking/content/tool_call) deltas *and* reconstructs a full
ChatCompletionMessage (content, tool_calls, reasoning_content) once the
stream ends. Tool-call argument fragments are accumulated by index, same
pattern every OpenAI-compatible streaming client needs. Downstream tool
execution code needed zero changes because of this — it just consumes the
reconstructed message instead of the SDK's blocking return value. Also
hardened <think>/</think> detection against a tag split across chunk
boundaries (was a per-chunk substring check; now a carry-over buffer) —
latent before since nothing exercised streaming for real.
- backend/services/streaming_answer.py (new): AnswerStreamFilter holds
buffered content back until it's implausible to still be one-line
narration (this codebase already caps narration at 150 chars elsewhere),
and separately holds back any trailing bracket run that could still grow
into a [src:ID]/[[src:ID]] marker — citation validity isn't known until
parse_and_renumber_citations runs on the complete text, so a raw marker,
hallucinated or not, must never reach the client even for a frame.
run_streaming_iteration drives one tool-loop iteration through this filter
and emits "answer_delta" events; if a tool_call shows up after content
already cleared the gate (narration grew unusually long before the model
pivoted), it emits one "answer_delta_retract" so the UI doesn't strand a
stale preview.
- backend/services/chat.py, mp_chat.py: thread an explicit stream_answer
flag through to _run_tool_loop's single generate() call site, rather than
inferring streaming from event_callback being non-None — eval logging's
ConversationRecorder.wrap() always returns a non-None callable, even for
the plain blocking /api/chat endpoint, so that inference would have been
wrong. Only stream_chat_response (the two SSE routes) sets it; /api/chat
is untouched.
- Frontend: both SSE consumers (ChatPanel.tsx's card stack, MpChatPanel.tsx's
simpler turn list) grow "answer_delta"/"answer_delta_retract" handling.
The live preview renders as plain, React-escaped text (no
dangerouslySetInnerHTML, no markdown parsing mid-stream — that's a
one-shot full-document transform, unsafe on partial input) and is fully
replaced by the authoritative, citation-renumbered/person-linked/
language-polished HTML once the terminal "answer" event lands. Both event
types were previously-unhandled and silently ignored, so the backend and
frontend halves are independently safe to deploy.
Added tests/test_llm_client_stream.py and tests/test_streaming_answer.py
covering the reconstruction (content, tool-call fragment accumulation,
split <think> tags, truncated-stream handling) and the filter (gate timing,
citation-marker safety across chunk boundaries, the retract path, error
propagation). 5 pre-existing failures in test_prompts_golden.py/
test_provenance.py are unrelated (reproduce identically on main).