diff --git a/CHANGELOG.md b/CHANGELOG.md index 24e9304..dd125fa 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -2,6 +2,21 @@ Full release notes with details on each version: [GitHub Releases](https://github.com/safishamsi/graphify/releases) +## 0.8.37 (2026-06-10) + +- Security: SSRF guard rewritten to eliminate thread-safety race. The global `socket.getaddrinfo` monkey-patch is replaced with per-connection `_SSRFGuardedHTTPConnection`/`_SSRFGuardedHTTPSConnection` subclasses that resolve DNS once, validate the IP, and connect to that exact address — closing both the concurrent-thread race window and the underlying TOCTOU gap. No global state is mutated, so sibling threads (MCP server, PR triage pool) are unaffected. +- Security: Prompt injection mitigation for LLM semantic extraction. Untrusted source file content is now wrapped in `` XML delimiters; jailbreak sentinel tokens (`<|im_start|>`, `[INST]`, `<>`, forged closing tags) are neutralised with a zero-width space; the extraction system prompt includes an explicit SECURITY block stating that content inside the wrapper is inert data. +- Fix: `export obsidian` and `export canvas` no longer crash with `KeyError` when a community contains a node ID absent from the graph (stale community index, merge artifacts). Dangling members are silently skipped. +- Fix: `--update` on macOS no longer re-extracts all Office files on every run. `convert_office_file()` now NFC-normalises the source path before hashing the sidecar filename, so NFD paths returned by `os.walk` (HFS+/APFS) produce the same sidecar as NFC-constructed paths. An early-return when the sidecar exists prevents mtime bumps causing spurious re-detection. +- Fix: Data `.json` files no longer explode into hundreds of orphan key-nodes. The JSON extractor now only processes config/manifest JSON (detected by filename — `package.json`, `tsconfig.json`, `.eslintrc.json`, `deno.json`, etc. — or by top-level keys such as `dependencies`, `extends`, `$ref`, `compilerOptions`). Data JSON (top-level arrays, generic key/value files) is skipped by the AST pass and left for the LLM semantic pass. +- Fix: OpenAI-compatible backends no longer send `temperature=0` to reasoning models. `_resolve_temperature()` auto-detects o1/o3/o4 and gpt-5 series and omits temperature from the request. Override with `GRAPHIFY_LLM_TEMPERATURE=` (or `none` to omit explicitly for any model). +- Fix: EDR / corporate Windows hang eliminated. `datasketch` (which transitively imports `scipy` → `numpy.testing` → `platform.machine()` subprocess at import time) is replaced by a self-contained pure-numpy MinHash/MinHashLSH implementation with byte-identical hash math. Removes `datasketch` and `scipy` from the dependency tree. +- Perf: `detect()` ignore-pattern checks memoized per scan. Each ancestor directory is now evaluated once across all sibling files, eliminating ~42M redundant `fnmatch` calls on large repos (~34% whole-run speedup on 2k-file corpora). +- Fix: `dedup.py` label-based merge passes now skip code nodes entirely. Distinct same-named symbols in different files (e.g. two `Config` classes) were being merged by the exact-label and MinHash/LSH passes. Code nodes are now deduplicated by ID only, which is correct. +- Feat: `GRAPHIFY_MAX_GRAPH_BYTES` env var to override the 512 MiB `graph.json` size cap. Accepts plain bytes, `MB`, or `GB`. The cap error message now cites this env var. `graphify export html` auto-falls back to the community-aggregation view when over cap instead of hard-failing. +- Feat: `CLAUDE.md` template now uses mandatory language for the graphify-first rule — "MANDATORY: Before using Read/Grep/Glob/Bash to explore the codebase, you MUST run graphify first" — and explicitly requires forwarding the rule to subagent prompts. PreToolUse hook message hardened to match. +- CI: Release workflow added. Every GitHub release now ships `graphify-self-graph.tar.gz` as a downloadable asset — `graph.json` + `graph.html` + `GRAPH_REPORT.md` from running Graphify on its own source. Open `graph.html` locally with no install required to see what Graphify produces. + ## 0.8.36 (2026-06-08) - Feat: `extra_body` field in `providers.json` forwarded to OpenAI-compat calls at extraction and labeling. Lets vLLM/Qwen3/Llama endpoints pass model-specific request shapes (e.g. `{"chat_template_kwargs": {"enable_thinking": false}}`). Explicit `extra_body` also bypasses Ollama `num_ctx` auto-derive. Thanks to @EirikWolf (#1197). diff --git a/README.md b/README.md index 465a5bb..d4790ad 100644 --- a/README.md +++ b/README.md @@ -419,6 +419,8 @@ These are only needed for **headless / CI extraction** (`graphify extract`). Whe | `GRAPHIFY_QUERY_LOG` | Override query log path (default: `~/.cache/graphify-queries.log`) | optional — set to empty or `/dev/null` to silence | | `GRAPHIFY_QUERY_LOG_DISABLE` | Set to `1` to disable query logging entirely | optional | | `GRAPHIFY_QUERY_LOG_RESPONSES` | Set to `1` to also log full subgraph responses (off by default) | optional | +| `GRAPHIFY_MAX_GRAPH_BYTES` | Override the 512 MiB graph.json size cap — e.g. `700MB`, `2GB`, or plain bytes | optional — useful for very large corpora | +| `GRAPHIFY_LLM_TEMPERATURE` | Override LLM temperature for semantic extraction — e.g. `0.7`, or `none` to omit | optional — auto-omitted for o1/o3/o4/gpt-5 reasoning models | --- diff --git a/pyproject.toml b/pyproject.toml index efbadd7..743a8ea 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -4,7 +4,7 @@ build-backend = "setuptools.build_meta" [project] name = "graphifyy" -version = "0.8.36" +version = "0.8.37" description = "AI coding assistant skill (Claude Code, CodeBuddy, Codex, OpenCode, Kilo Code, Cursor, Gemini CLI, Aider, OpenClaw, Factory Droid, Trae, Hermes, Kiro, Pi, Devin CLI, Google Antigravity) - turn any folder of code, docs, papers, images, or videos into a queryable knowledge graph" readme = "README.md" license = { file = "LICENSE" }