Fix: stop silently dropping source files during discovery via two over-broad filters (#1666)

@krishnateja7 root-caused this precisely: the files were never reaching
extraction, so the 0.9.7 no-cache-on-empty mitigation could not surface them.
Two discovery-layer filters were the cause:

(a) A bare `snapshots/` directory was pruned as a Jest/Vitest artifact, which
killed legitimate code namespaces like a Rails `app/services/snapshots/`. It is
now pruned only when it actually contains `.snap` files or sits directly under a
JS test root (`__tests__`/`__test__`). `__snapshots__` stays unconditionally pruned.

(b) `_is_sensitive` dropped files on a bare name-keyword hit (device_token.rb,
passwords_controller.rb) even when `classify_file` had already resolved them to
source code. A genuine programming-language source file is now exempt from the
weak keyword heuristic, while real secret stores in data/config formats
(credentials.json, secrets.yaml, .env, .pem, ...) are still caught — those route
through the CODE path for manifest parsing but are deliberately not exempted.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
safishamsi
2026-07-06 16:00:31 +01:00
co-authored by Claude Opus 4.8
parent bbc3be2238
commit 912d832a60
3 changed files with 86 additions and 7 deletions
+1
View File
@@ -9,6 +9,7 @@ Full release notes with details on each version: [GitHub Releases](https://githu
- Fix: the Ollama backend no longer multiplies a hang by the retry count (#1686, thanks @Kyzcreig). A stalled local model would wedge for `timeout * (max_retries + 1)`, which with the default 6 retries turned one long stall into a very long one. Ollama now defaults to zero client-side retries (a local model that stalls will not un-stall on retry); set `GRAPHIFY_MAX_RETRIES` to opt back in. Other backends are unchanged. Note: the underlying stall is non-deterministic and driven by the model server, so this bounds the wait rather than eliminating the hang.
- Fix: a truncated or slightly malformed community-labeling reply no longer discards the whole batch (#1690, thanks @vdgbcrypto). `_parse_label_response` now salvages the complete `"id": "name"` pairs from a reply that failed a strict `json.loads` (e.g. a reply truncated mid-object), raising only when no pairs can be recovered. The per-batch token budget was also raised (`256 + 48*n`, was `64 + 24*n`) to give models that prepend a short preamble enough headroom to finish the JSON. The exact provider truncation in the report could not be reproduced without a live key; the parser and budget fixes address the mechanism.
- Fix: cluster-only mode now reports the real token cost of community labeling instead of a hardcoded zero (#1694, thanks @sub4biz). The labeling LLM calls were never accounted for, so `GRAPH_REPORT.md`'s "Token cost" line always read `0 input · 0 output` in cluster-only runs. `_call_llm` now accumulates per-response usage into an optional accumulator that is threaded through the labeling path and surfaced in the report. Backends that do not return usage (the Claude Code CLI) still contribute nothing, which is honest rather than estimated.
- Fix: source files are no longer silently dropped during discovery by two over-broad filters (#1666, thanks @krishnateja7 for the precise root-cause). (a) A bare `snapshots/` directory was pruned as a Jest/Vitest artifact, which killed legitimate code namespaces like a Rails `app/services/snapshots/`; it is now pruned only when it actually contains `.snap` files or sits directly under a JS test root (`__snapshots__` stays unconditionally pruned). (b) `_is_sensitive` dropped files on a bare name-keyword hit (`device_token.rb`, `passwords_controller.rb`) even when `classify_file` had already resolved them to source code; a genuine programming-language source file is now exempt from the weak keyword heuristic, while real secret stores in data/config formats (`credentials.json`, `secrets.yaml`, `.env`, `.pem`, ...) are still caught. This is the discovery-layer fix; the 0.9.7 no-cache-on-empty change could not surface these because the files never reached extraction.
## 0.9.7 (2026-07-06)
+44 -3
View File
@@ -125,6 +125,15 @@ _GENERIC_KEYWORD_PATTERNS = [
re.compile(r'(?<![a-zA-Z0-9])tokens?(?![a-zA-Z])', re.IGNORECASE),
]
# Data/serialization extensions that commonly ARE secret stores when their name
# hits a generic keyword (credentials.json, secrets.yaml, token.toml). These stay
# subject to the Stage 3 keyword drop even though some route through the CODE path
# for manifest parsing — only real programming-language source is exempt (#1666).
_SECRET_PRONE_DATA_EXTS = frozenset({
".json", ".yaml", ".yml", ".toml", ".ini", ".cfg", ".conf", ".config",
".xml", ".properties", ".env", ".txt",
})
# Word separators for the load-bearing check (underscore intentionally included;
# multi-word keywords like private_key are handled by the end-of-stem check,
# which runs before word counting).
@@ -184,8 +193,19 @@ def _is_sensitive(path: Path) -> bool:
name = path.name
if any(p.search(name) for p in _SENSITIVE_PATTERNS):
return True
# Stage 3: generic keywords, only when load-bearing in the name
return _generic_keyword_hit(name)
# Stage 3: generic keywords, only when load-bearing in the name. Do NOT let a
# bare name keyword silently drop a genuine programming-language source file:
# a .rb/.py named device_token or passwords_controller is a module, not a secret
# store (#1666). Data/config formats (.json, .yaml, .toml, ...) are deliberately
# NOT exempt even though .json routes through the CODE path for manifest parsing,
# because credentials.json / oauth_token.json / secrets.yaml are exactly the
# secret stores this stage must catch. The specific Stage 2 patterns (.env, .pem,
# id_rsa, ...) still apply to everything regardless of extension.
if _generic_keyword_hit(name):
ext = path.suffix.lower()
is_source_code = classify_file(path) == FileType.CODE and ext not in _SECRET_PRONE_DATA_EXTS
return not is_source_code
return False
def _looks_like_paper(path: Path) -> bool:
@@ -677,7 +697,7 @@ _SKIP_DIRS = {
# Coverage/test-artefact dirs — generated, never architecturally meaningful
"coverage", "lcov-report", # Vitest/Istanbul/nyc HTML reports (#870)
"visual-tests", "visual-test", # Playwright/visual-regression bundles (#869)
"__snapshots__", "snapshots", # Jest/Vitest snapshot dirs
"__snapshots__", # Jest/Vitest snapshot dir (unambiguous)
"storybook-static", # Storybook production build output
"dist-protected", # Protected dist variants (same noise as dist)
# Framework cache/build dirs — generated, never architecturally meaningful (#873)
@@ -694,10 +714,31 @@ _SKIP_FILES = {
"composer.lock", "go.sum", "go.work.sum",
}
# A bare "snapshots" dir is a Jest/Vitest artifact only when it actually holds
# snapshot files or lives directly under a JS test root. Elsewhere it is often a
# real code namespace (e.g. Rails app/services/snapshots/), so pruning it by name
# silently dropped legitimate source from the graph (#1666). "__snapshots__" stays
# unconditionally pruned above; only the ambiguous bare name is gated here.
_JS_SNAPSHOT_TEST_ROOTS = frozenset({"__tests__", "__test__"})
def _is_noise_dir(part: str, parent: "Path | None" = None) -> bool:
"""Return True if this directory name looks like a venv, cache, or dep dir."""
if part in _SKIP_DIRS:
return True
if part == "snapshots":
# Prune only when it looks like an actual JS/Vitest snapshot dir.
if parent is None:
return False # cannot verify; keep a possibly-real code dir
snap_dir = parent / part
if parent.name in _JS_SNAPSHOT_TEST_ROOTS:
return True
try:
if next(snap_dir.glob("*.snap"), None) is not None:
return True
except OSError:
pass
return False
# Catch *_venv, *_repo/site-packages patterns
if part.endswith("_venv") or part.endswith("_env"):
return True
+41 -4
View File
@@ -1,3 +1,4 @@
import os
import unicodedata
from pathlib import Path
from graphify.detect import classify_file, count_words, detect, detect_incremental, save_manifest, FileType, _looks_like_paper, _is_ignored, _load_graphifyignore, _is_sensitive
@@ -467,16 +468,35 @@ def test_detect_skips_visual_tests_dir(tmp_path):
def test_detect_skips_snapshots_dir(tmp_path):
"""__snapshots__/ and snapshots/ are jest/vitest artefacts — must be excluded."""
"""__snapshots__/ and real jest/vitest snapshots/ dirs are artefacts — excluded."""
(tmp_path / "__snapshots__").mkdir()
(tmp_path / "__snapshots__" / "app.test.ts.snap").write_text("// Jest Snapshot\nexports[`test 1`] = `<div/>`")
# a bare snapshots/ dir that actually holds .snap files is still a JS artefact
snap = tmp_path / "snapshots"
snap.mkdir()
(snap / "component.test.tsx.snap").write_text("exports[`renders`] = `<span/>`")
(tmp_path / "app.ts").write_text("export function greet() { return 'hi'; }")
result = detect(tmp_path)
all_files = [f for files in result["files"].values() for f in files]
assert not any("__snapshots__" in f for f in all_files)
assert not any(f"{os.sep}snapshots{os.sep}" in f for f in all_files)
assert any("app.ts" in f for f in all_files)
def test_detect_keeps_snapshots_code_namespace(tmp_path):
"""#1666: a bare snapshots/ dir with no .snap files is a legit code namespace
(e.g. Rails app/services/snapshots/) and must NOT be pruned as a JS artefact."""
svc = tmp_path / "app" / "services" / "snapshots"
svc.mkdir(parents=True)
(svc / "round_reader.rb").write_text("class RoundReader\n def call; end\nend\n")
(svc / "backfill_marker.rb").write_text("class BackfillMarker\n def run; end\nend\n")
(tmp_path / "app.rb").write_text("class App; end\n")
result = detect(tmp_path)
all_files = [f for files in result["files"].values() for f in files]
assert any("round_reader.rb" in f for f in all_files)
assert any("backfill_marker.rb" in f for f in all_files)
def test_detect_skips_storybook_static_dir(tmp_path):
"""storybook-static/ is a build artefact — must be excluded."""
sb = tmp_path / "storybook-static"
@@ -810,9 +830,26 @@ def test_sensitive_does_not_flag_tokenizer_py():
def test_sensitive_does_not_flag_tokenize_py():
assert not _is_sensitive(Path("tokenize.py"))
def test_sensitive_flags_passwords_py():
# passwords.py is just as likely a secret store as passwords.txt — code ext is no excuse
assert _is_sensitive(Path("passwords.py"))
def test_sensitive_does_not_flag_passwords_py():
# #1666: a programming-language source file named after a domain noun is a
# module, not a secret store. Silently dropping it hid real code from the graph.
# Genuine secret stores are .env/.pem/credentials.json etc. (still flagged below).
assert not _is_sensitive(Path("passwords.py"))
def test_sensitive_does_not_flag_ruby_code_modules():
# #1666 exact cases: Rails source modules with keyword-ish names must survive.
assert not _is_sensitive(Path("app/models/device_token.rb"))
assert not _is_sensitive(Path("app/controllers/api/v1/passwords_controller.rb"))
def test_sensitive_still_flags_data_secret_stores():
# #1666 guard: the exemption is ONLY for real source code, not data/config
# formats — credentials.json / oauth_token.json / secrets.yaml are the secret
# stores Stage 3 must keep catching (even though .json routes through CODE).
assert _is_sensitive(Path("credentials.json"))
assert _is_sensitive(Path("oauth_token.json"))
assert _is_sensitive(Path("app_secret.yaml"))
def test_sensitive_flags_ssh_dir():
assert _is_sensitive(Path("/home/user/.ssh/id_rsa"))