Skip to content

fix(worker): fall back to thinking field; strip JSON fences before reflection parse (#38) - #39

Open
linhongyu510 wants to merge 1 commit into
ClaudioDrews:mainfrom
linhongyu510:fix/reflection-thinking-fences
Open

linhongyu510 wants to merge 1 commit into
ClaudioDrews:mainfrom
linhongyu510:fix/reflection-thinking-fences

Conversation

@linhongyu510

Copy link
Copy Markdown

Problem

Two related reflection-pipeline bugs reported in #38 (unassigned, no competing PR):

Bug 1 — thinking fallback missing. ollama_chat() in docker/worker/services/llm.py only fell back to the reasoning response key when response was empty. glm-5.x-family models (e.g. glm-5.3-flash on Ollama Cloud) return their chain-of-thought under thinking. With a reasoning model and a low num_predict, response stays empty and every token goes to thinking — ollama_chat() returns "" and the reflection job fails even though the model completed fine.

Bug 2 — reflection degrades to a raw blob. reflect_on_memories() in docker/worker/tasks/reflection.py calls json.loads(response) directly. Most current chat models wrap JSON in markdown fences even when the prompt asks for raw JSON, so json.loads fails on the opening backticks, the JSONDecodeError branch fires, and the point is stored as {"raw": ...} — the structured fields (patterns/connections/insights/actions) are lost.

Fix

  • _extract_response(data) helper: response → reasoning → thinking (each only if non-empty), used by ollama_chat().
  • _strip_json_fences(response) helper: strips one markdown code fence (language-tagged or bare, tolerating trailing whitespace) before json.loads in reflect_on_memories().

Both are minimal, additive, and match the reporter's suggested fixes.

Tests

New pytest module docker/worker/tests/test_llm_reflection.py (the worker had no pytest infra — added deliberately small, scoped to the changed code):

Result: 8 passed (7 unit + 1 regression). Baseline check: with the old code the new tests fail to collect / the regression fails (raw-blob path), confirming red → green.

Notes

  • No lint config exists in the repo; touched files pass python -m py_compile.
  • Verification is local/mocked — no live Ollama or Qdrant was used.

…flection parse (ClaudioDrews#38)

Bug 1: ollama_chat only fell back to 'reasoning' when 'response' was
empty, but glm-5.x-family models return chain-of-thought under
'thinking' (verified on Ollama Cloud). With a low num_predict every
token went to thinking, response stayed empty, and reflection jobs
failed even though the model completed fine.

Bug 2: reflect_on_memories called json.loads(response) directly. Most
current chat models wrap JSON in markdown fences even when asked for
raw JSON, so json.loads failed and the reflection was stored as
{'raw': ...} - the structured fields (patterns/connections/insights/
actions) were lost and the point degraded to an opaque blob.

Extract _extract_response() (response -> reasoning -> thinking) and
_strip_json_fences() helpers and use them at both call sites. New
pytest module under docker/worker/tests covers both helpers plus a
regression test proving a fenced response now lands as structured
JSON instead of a raw blob.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant