Feat: opt-in GRAPHIFY_DISABLE_THINKING; correct deepseek thinking default (#1621)
@sub4biz verified against the live DeepSeek API that deepseek-v4-flash (and v4-pro) have thinking ENABLED by default, contradicting the built-in config's stale "non-thinking" comment (now corrected). The naive fix (mirror the kimi branch and force thinking off) is the wrong call: @sub4biz's production testing on real corpora found that disabling thinking removes a rare reasoning-leak failure — which the adaptive extraction/labeling retry already recovers from — but trades it for far more frequent benign truncation AND measurably lower extraction quality and file coverage, confirmed by a blind second reviewer. So thinking stays ON by default (quality/coverage), with a documented opt-in `GRAPHIFY_DISABLE_THINKING=1` for users who prefer run-to-run stability. Applies to reasoning-capable OpenAI-compatible backends at both extra_body sites (extraction + labeling). An explicit providers.json extra_body still wins, and the moonshot/kimi branch is unchanged (it must disable thinking or content is empty). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 4.8
parent
912d832a60
commit
5d0137388e
@@ -9,6 +9,7 @@ Full release notes with details on each version: [GitHub Releases](https://githu
|
||||
- Fix: the Ollama backend no longer multiplies a hang by the retry count (#1686, thanks @Kyzcreig). A stalled local model would wedge for `timeout * (max_retries + 1)`, which with the default 6 retries turned one long stall into a very long one. Ollama now defaults to zero client-side retries (a local model that stalls will not un-stall on retry); set `GRAPHIFY_MAX_RETRIES` to opt back in. Other backends are unchanged. Note: the underlying stall is non-deterministic and driven by the model server, so this bounds the wait rather than eliminating the hang.
|
||||
- Fix: a truncated or slightly malformed community-labeling reply no longer discards the whole batch (#1690, thanks @vdgbcrypto). `_parse_label_response` now salvages the complete `"id": "name"` pairs from a reply that failed a strict `json.loads` (e.g. a reply truncated mid-object), raising only when no pairs can be recovered. The per-batch token budget was also raised (`256 + 48*n`, was `64 + 24*n`) to give models that prepend a short preamble enough headroom to finish the JSON. The exact provider truncation in the report could not be reproduced without a live key; the parser and budget fixes address the mechanism.
|
||||
- Fix: cluster-only mode now reports the real token cost of community labeling instead of a hardcoded zero (#1694, thanks @sub4biz). The labeling LLM calls were never accounted for, so `GRAPH_REPORT.md`'s "Token cost" line always read `0 input · 0 output` in cluster-only runs. `_call_llm` now accumulates per-response usage into an optional accumulator that is threaded through the labeling path and surfaced in the report. Backends that do not return usage (the Claude Code CLI) still contribute nothing, which is honest rather than estimated.
|
||||
- Docs/Feat: `deepseek-v4-flash` (and `v4-pro`) have thinking ENABLED by default; graphify no longer implies otherwise and adds an opt-in `GRAPHIFY_DISABLE_THINKING=1` toggle (#1621, thanks @sub4biz for the empirical testing). Disabling thinking removes a rare reasoning-leak failure mode (which the adaptive extraction/labeling retry already recovers from) but, measured on real corpora, trades it for more frequent benign truncation and measurably lower extraction quality and file coverage — so it stays a documented user choice rather than a forced default. The stale "non-thinking" comment on the built-in deepseek config is corrected. The moonshot (kimi) branch is unchanged (it must disable thinking or content comes back empty).
|
||||
- Fix: source files are no longer silently dropped during discovery by two over-broad filters (#1666, thanks @krishnateja7 for the precise root-cause). (a) A bare `snapshots/` directory was pruned as a Jest/Vitest artifact, which killed legitimate code namespaces like a Rails `app/services/snapshots/`; it is now pruned only when it actually contains `.snap` files or sits directly under a JS test root (`__snapshots__` stays unconditionally pruned). (b) `_is_sensitive` dropped files on a bare name-keyword hit (`device_token.rb`, `passwords_controller.rb`) even when `classify_file` had already resolved them to source code; a genuine programming-language source file is now exempt from the weak keyword heuristic, while real secret stores in data/config formats (`credentials.json`, `secrets.yaml`, `.env`, `.pem`, ...) are still caught. This is the discovery-layer fix; the 0.9.7 no-cache-on-empty change could not surface these because the files never reached extraction.
|
||||
|
||||
## 0.9.7 (2026-07-06)
|
||||
|
||||
+26
-2
@@ -131,8 +131,10 @@ BACKENDS: dict[str, dict] = {
|
||||
"env_key": "DEEPSEEK_API_KEY",
|
||||
"model_env_key": "GRAPHIFY_DEEPSEEK_MODEL",
|
||||
"pricing": {"input": 0.14, "output": 0.28}, # USD per 1M tokens (v4-flash)
|
||||
# deepseek-reasoner / thinking-mode models silently ignore temperature;
|
||||
# deepseek-chat / v4-flash (non-thinking) accept 0-2. Safe to send 0.
|
||||
# deepseek-reasoner silently ignores temperature; deepseek-chat / v4-flash
|
||||
# accept 0-2, so sending 0 is safe. Note: deepseek-v4-flash (and v4-pro) have
|
||||
# thinking ENABLED by default (verified against the live API, #1621) — set
|
||||
# GRAPHIFY_DISABLE_THINKING=1 to turn it off (tradeoff documented on the flag).
|
||||
"temperature": 0,
|
||||
"max_tokens": 16384,
|
||||
},
|
||||
@@ -388,6 +390,22 @@ def _resolve_max_retries(default: int = 6) -> int:
|
||||
pass
|
||||
return default
|
||||
|
||||
|
||||
def _thinking_disabled_via_env() -> bool:
|
||||
"""Opt-in (GRAPHIFY_DISABLE_THINKING) to send ``{"thinking": {"type": "disabled"}}``
|
||||
to reasoning-capable OpenAI-compatible models such as ``deepseek-v4-flash``.
|
||||
|
||||
Off by default and deliberately so (#1621): a thinking-on model can occasionally
|
||||
leak reasoning prose instead of JSON, but that response is caught and re-tried by
|
||||
the adaptive extraction/labeling retry, so it is a rare, recoverable failure.
|
||||
Disabling thinking removes that failure mode but, measured on real corpora, trades
|
||||
it for far more frequent (benign) truncation AND measurably lower extraction
|
||||
quality and file coverage. So this stays a user choice for those who value
|
||||
run-to-run stability over extraction quality, not a forced default. The moonshot
|
||||
(kimi) branch keeps disabling thinking unconditionally because that model returns
|
||||
empty content otherwise."""
|
||||
return os.environ.get("GRAPHIFY_DISABLE_THINKING", "").strip().lower() in ("1", "true", "yes", "on")
|
||||
|
||||
_EXTRACTION_SYSTEM = """\
|
||||
You are a graphify semantic extraction agent. Extract a knowledge graph fragment from the files provided.
|
||||
Output ONLY valid JSON — no explanation, no markdown fences, no preamble.
|
||||
@@ -996,6 +1014,10 @@ def _call_openai_compat(
|
||||
# Kimi-k2.6 is a reasoning model — disable thinking so content isn't empty
|
||||
elif "moonshot" in base_url:
|
||||
kwargs["extra_body"] = {"thinking": {"type": "disabled"}}
|
||||
# Opt-in only: disable thinking for reasoning models like deepseek-v4-flash
|
||||
# (#1621). Not a default — see _thinking_disabled_via_env for the tradeoff.
|
||||
elif _thinking_disabled_via_env():
|
||||
kwargs["extra_body"] = {"thinking": {"type": "disabled"}}
|
||||
# Ollama defaults num_ctx to 2048 and silently truncates prompts larger
|
||||
# than that — the symptom is hollow 200 OK responses after the first few
|
||||
# chunks (#798). We derive num_ctx from the actual prompt size so we don't
|
||||
@@ -2113,6 +2135,8 @@ def _call_llm(
|
||||
kwargs["extra_body"] = cfg["extra_body"]
|
||||
elif "moonshot" in cfg["base_url"]:
|
||||
kwargs["extra_body"] = {"thinking": {"type": "disabled"}}
|
||||
elif _thinking_disabled_via_env():
|
||||
kwargs["extra_body"] = {"thinking": {"type": "disabled"}}
|
||||
resp = client.chat.completions.create(**kwargs)
|
||||
if not resp.choices or resp.choices[0].message is None:
|
||||
raise ValueError("LLM returned empty or filtered response")
|
||||
|
||||
@@ -578,6 +578,53 @@ def test_call_openai_compat_extra_body_wins_over_moonshot_default(monkeypatch):
|
||||
assert captured["extra_body"] == {"thinking": {"type": "enabled"}}
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# GRAPHIFY_DISABLE_THINKING: opt-in disable-thinking for reasoning models like
|
||||
# deepseek-v4-flash. Off by default — disabling thinking trades a rare reasoning
|
||||
# leak for lower extraction quality/coverage, so it must not be forced (#1621).
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
def test_deepseek_thinking_on_by_default(monkeypatch):
|
||||
monkeypatch.delenv("GRAPHIFY_DISABLE_THINKING", raising=False)
|
||||
captured = _install_capturing_openai(monkeypatch)
|
||||
|
||||
llm._call_openai_compat(
|
||||
"https://api.deepseek.com", "sk", "deepseek-v4-flash",
|
||||
"u", temperature=0, max_completion_tokens=8192, backend="deepseek",
|
||||
)
|
||||
|
||||
eb = captured.get("extra_body")
|
||||
assert eb is None or "thinking" not in eb, "thinking must NOT be disabled by default"
|
||||
|
||||
|
||||
def test_deepseek_thinking_disabled_via_env(monkeypatch):
|
||||
monkeypatch.setenv("GRAPHIFY_DISABLE_THINKING", "1")
|
||||
captured = _install_capturing_openai(monkeypatch)
|
||||
|
||||
llm._call_openai_compat(
|
||||
"https://api.deepseek.com", "sk", "deepseek-v4-flash",
|
||||
"u", temperature=0, max_completion_tokens=8192, backend="deepseek",
|
||||
)
|
||||
|
||||
assert captured["extra_body"] == {"thinking": {"type": "disabled"}}
|
||||
|
||||
|
||||
def test_explicit_extra_body_wins_over_thinking_env(monkeypatch):
|
||||
# A provider-supplied extra_body is an explicit request-shape choice and must
|
||||
# take precedence over the env toggle.
|
||||
monkeypatch.setenv("GRAPHIFY_DISABLE_THINKING", "1")
|
||||
captured = _install_capturing_openai(monkeypatch)
|
||||
|
||||
llm._call_openai_compat(
|
||||
"https://api.deepseek.com", "sk", "deepseek-v4-flash",
|
||||
"u", temperature=0, max_completion_tokens=8192, backend="deepseek",
|
||||
extra_body={"thinking": {"type": "enabled"}},
|
||||
)
|
||||
|
||||
assert captured["extra_body"] == {"thinking": {"type": "enabled"}}
|
||||
|
||||
|
||||
def test_call_openai_compat_explicit_extra_body_skips_ollama_auto_derive(monkeypatch):
|
||||
# An explicit extra_body means "I own this request shape" — Ollama's
|
||||
# num_ctx auto-derive (a default) must step aside or we'd clobber it.
|
||||
|
||||
Reference in New Issue
Block a user