diff --git a/CHANGELOG.md b/CHANGELOG.md index 5f2e0b6..2de4329 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -4,6 +4,16 @@ Full release notes with details on each version: [GitHub Releases](https://githu ## Unreleased +## 0.8.43 (2026-06-19) + +- Feat: package manifests are now parsed deterministically into a dependency graph. `apm.yml`, `pyproject.toml`, `go.mod`, and `pom.xml` each yield ONE canonical package node per package (keyed by name) plus `depends_on` edges, routed to the AST path so the LLM never sees them. Previously `apm.yml` was an LLM-handled document, so the same package got a different file-anchored id from its own manifest than from each dependent's dependency reference and split into duplicate nodes; a package referenced from N manifests is now a single hub node (#1377). +- Feat: markdown links now become graph edges. `extract_markdown` only emitted heading nodes + `contains` edges and never parsed link syntax, so a doc full of `[text](./other.md)` links (e.g. `index.md`, `table-of-contents.md`) had no edges to what it links and never became a hub. Inline links, reference-style links, and `[[wikilinks]]` are resolved relative to the source file (external URLs / in-page anchors / images skipped) and emitted as `references` edges, with targets resolved via the same node-id recipe so they merge onto the real doc node (#1376). +- Security: bumped vulnerable dependencies to patched versions — `pypdf` 6.11.0→6.13.3 (CVE-2026-48155/48156), `yt-dlp` 2026.3.17→2026.6.9, `pyjwt` 2.12.1→2.13.0, `cryptography` 48.0.0→49.0.0, `python-multipart` 0.0.28→0.0.32 — with lower-bound floors for the direct deps (`pypdf`, `yt-dlp`) so installs get the patched versions (#1375; thanks @hypnwtykvmpr). +- Fix: the semantic extract entry points (`extract_corpus_parallel`, `extract_files_direct`) crashed with `AttributeError` when passed `str` paths instead of `pathlib.Path`. Both now coerce `files = [Path(f) for f in files]` at entry (#1386). +- Fix: community labeling now recovers from a malformed-JSON batch by splitting it at the midpoint and retrying each half (mirroring the extract path), instead of logging-and-skipping it — which silently lost ~100 community names per failed batch on large graphs (#1280, #1278; thanks @CJdev232). +- Fix: `graphify hook install` no longer creates a literal backslash-named junk directory and reports false success when `core.hooksPath` (or `git rev-parse --git-path hooks`) is a Windows-style path under WSL. Drive-letter / embedded-backslash hooks paths are now rejected with a clear error (#1385). +- Refactor: node-ID normalization is unified into a single `graphify.ids` module. `extract._make_id`, `build._normalize_id`, `mcp_ingest._make_id`, and `symbol_resolution._bash_make_id` were four hand-synced copies of the same NFKC/casefold recipe — the root of the recurring ghost-node bug class (#811/#550/#1033/#1104). All four now delegate to one implementation guarded by contract + hypothesis property tests (#1378; thanks @danielnguyenfinhub). + ## 0.8.42 (2026-06-18) - Fix: large text documents are no longer silently truncated during semantic extraction. `_read_files` capped every file at 20,000 characters, so a Markdown/text/rST document longer than that had everything past the cap dropped — the model never saw it, and the packer/adaptive-retry path couldn't recover ("packing can't shrink one big file"). Oversized splittable-text files are now sliced at heading/paragraph boundaries into units that each fit the cap and together cover the whole file; every slice reports its parent file as `source_file`, so the graph is not fragmented per-slice. A single slice that still overflows the model's output is bisected and retried. Code files and PDFs are never sliced (they keep whole-symbol / page handling). (#1369) diff --git a/README.md b/README.md index e2be274..9422523 100644 --- a/README.md +++ b/README.md @@ -242,7 +242,8 @@ To remove graphify from all platforms at once: `graphify uninstall` (add `--purg | Salesforce Apex | `.cls .trigger` (regex-based; classes, interfaces, enums, methods, triggers, SOQL/DML edges) | | Terraform / HCL | `.tf .tfvars .hcl` (requires `uv tool install graphifyy[terraform]`) | | MCP configs | `.mcp.json` `mcp.json` `mcp_servers.json` `claude_desktop_config.json` — extracts server nodes, package refs, env var requirements | -| Docs | `.md .mdx .qmd .html .txt .rst .yaml .yml` | +| Package manifests | `apm.yml` `pyproject.toml` `go.mod` `pom.xml` — one canonical package node per package (by name) plus `depends_on` edges, so a package referenced from many manifests is a single hub | +| Docs | `.md .mdx .qmd .html .txt .rst .yaml .yml` (markdown `[text](./other.md)` links and `[[wikilinks]]` become `references` edges between docs) | | Office | `.docx .xlsx` (requires `uv tool install graphifyy[office]`) | | Google Workspace | `.gdoc .gsheet .gslides` (opt-in; requires `gws` auth and `--google-workspace`; Sheets need `uv tool install graphifyy[google]`) | | PDFs | `.pdf` | diff --git a/pyproject.toml b/pyproject.toml index d337e1c..5e889ef 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -4,7 +4,7 @@ build-backend = "setuptools.build_meta" [project] name = "graphifyy" -version = "0.8.42" +version = "0.8.43" description = "AI coding assistant skill (Claude Code, CodeBuddy, Codex, OpenCode, Kilo Code, Cursor, Gemini CLI, Aider, OpenClaw, Factory Droid, Trae, Hermes, Kiro, Pi, Devin CLI, Google Antigravity) - turn any folder of code, docs, papers, images, or videos into a queryable knowledge graph" readme = "README.md" license = { file = "LICENSE" }