diff --git a/CHANGELOG.md b/CHANGELOG.md index c49cb67..fdb5825 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -2,6 +2,14 @@ Full release notes with details on each version: [GitHub Releases](https://github.com/safishamsi/graphify/releases) +## 0.8.32 (2026-06-05) + +- Feat: Terraform/HCL support. `.tf`, `.tfvars`, and `.hcl` files are now AST-extracted via `tree-sitter-hcl` into a structured infrastructure dependency graph. Nodes: resources, data sources, modules, variables, outputs, providers, and locals. Edges: `contains`, `references` (interpolation), and `depends_on`. Node IDs are directory-scoped for cross-file resolution. Requires `uv tool install "graphifyy[terraform]"` (#1129). +- Fix: `graphify extract` no longer requires an LLM API key for code-only corpora. Backend resolution is now deferred until after file detection — a corpus with only code files (pure tree-sitter AST, zero LLM calls) runs fully offline. The key is only enforced when docs, PDFs, or images are present, or when `--dedup-llm` is passed (#1122). +- Fix: `graphify kiro install` now correctly installs the `references/` sidecar and `.graphify_version` stamp. The install was using a bare `write_text` that bypassed the shared helper, shipping `SKILL.md` with 8 dead `references/*.md` pointers. Re-run `graphify kiro install` to pick up the fix (#1142). +- Fix: `GRAPHIFY_API_TIMEOUT` now applies to `claude-cli` subprocess and Anthropic SDK backend, not just the HTTP client. Both subprocess paths previously hardcoded `timeout=600` and ignored the env var and `--api-timeout` flag (#1112). +- Build: version floors added for `networkx>=3.4`, `datasketch>=1.6`, and `rapidfuzz>=3.0` to prevent silent breakage from old installs resolving incompatible versions. + ## 0.8.31 (2026-06-03) - Fix: `graphify hook install` now embeds the current interpreter (`sys.executable`) directly into the generated hook scripts. Previously, uv tool and pipx installs silently no-oped on git commit in GUI clients and CI runners where `~/.local/bin` is not on PATH — the hook could not find the graphify launcher, fell through all detection probes, and exited 0 without rebuilding. The embedded path is sanitized through a filesystem-safe allowlist before substitution. If you already have hooks installed, re-run `graphify hook install` to pick up the fix (#1127). diff --git a/README.md b/README.md index ab854f9..4d701ba 100644 --- a/README.md +++ b/README.md @@ -173,6 +173,7 @@ Install only what you need: | `bedrock` | AWS Bedrock (uses IAM, no API key) | `uv tool install "graphifyy[bedrock]"` | | `sql` | SQL schema extraction | `uv tool install "graphifyy[sql]"` | | `dm` | BYOND DreamMaker `.dm`/`.dme` AST extraction (may need a C compiler + `python3-dev` if no wheel matches your platform) | `uv tool install "graphifyy[dm]"` | +| `terraform` | Terraform / HCL `.tf`/`.tfvars`/`.hcl` AST extraction | `uv tool install "graphifyy[terraform]"` | | `chinese` | Chinese query segmentation (jieba) | `uv tool install "graphifyy[chinese]"` | | `all` | Everything above | `uv tool install "graphifyy[all]"` | @@ -234,7 +235,8 @@ To remove graphify from all platforms at once: `graphify uninstall` (add `--purg | Type | Extensions | |------|-----------| -| Code (33 languages) | `.py .ts .js .jsx .tsx .mjs .go .rs .java .c .cpp .h .hpp .rb .cs .kt .scala .php .swift .lua .luau .zig .ps1 .ex .exs .m .mm .jl .vue .svelte .astro .groovy .gradle .dart .v .sv .svh .sql .f .f90 .f95 .f03 .f08 .pas .pp .dpr .dpk .lpr .inc .dfm .lfm .lpk .sh .bash .json .dm .dme .dmi .dmm .dmf .sln .csproj .fsproj .vbproj .razor .cshtml` (`.dm`/`.dme` AST extraction requires `uv tool install graphifyy[dm]`) | +| Code (28 tree-sitter grammars) | `.py .ts .js .jsx .tsx .mjs .go .rs .java .c .cpp .h .hpp .rb .cs .kt .scala .php .swift .lua .luau .zig .ps1 .ex .exs .m .mm .jl .vue .svelte .astro .groovy .gradle .dart .v .sv .svh .sql .f .f90 .f95 .f03 .f08 .pas .pp .dpr .dpk .lpr .inc .dfm .lfm .lpk .sh .bash .json .dm .dme .dmi .dmm .dmf .sln .csproj .fsproj .vbproj .razor .cshtml` (`.dm`/`.dme` requires `uv tool install graphifyy[dm]`) | +| Terraform / HCL | `.tf .tfvars .hcl` (requires `uv tool install graphifyy[terraform]`) | | MCP configs | `.mcp.json` `mcp.json` `mcp_servers.json` `claude_desktop_config.json` — extracts server nodes, package refs, env var requirements | | Docs | `.md .mdx .qmd .html .txt .rst .yaml .yml` | | Office | `.docx .xlsx` (requires `uv tool install graphifyy[office]`) | @@ -296,7 +298,9 @@ See the [full command reference](#full-command-reference) below. ## Ignoring files -Create a `.graphifyignore` in your project root — same syntax as `.gitignore`, including `!` negation: +Create a `.graphifyignore` in your project root — same syntax as `.gitignore`, including `!` negation. + +**`.gitignore` is respected automatically.** If no `.graphifyignore` is present in a directory, graphify falls back to the `.gitignore` in that directory. If both exist, `.graphifyignore` takes priority. Subdirectory scoping works the same way as git — an ignore file only affects its own subtree. ``` # .graphifyignore @@ -386,7 +390,7 @@ These are only needed for **headless / CI extraction** (`graphify extract`). Whe ## Privacy -- **Code files** — processed locally via tree-sitter. Nothing leaves your machine. +- **Code files** — processed locally via tree-sitter. Nothing leaves your machine. A code-only corpus requires no API key — `graphify extract` runs fully offline. - **Video / audio** — transcribed locally with faster-whisper. Nothing leaves your machine. - **Docs, PDFs, images** — sent to your AI assistant for semantic extraction (via the `/graphify` skill, using whatever model your IDE session runs). Headless `graphify extract` requires `GEMINI_API_KEY` / `GOOGLE_API_KEY` (Gemini), `MOONSHOT_API_KEY` (Kimi), `ANTHROPIC_API_KEY` (Claude), `OPENAI_API_KEY` (OpenAI), `DEEPSEEK_API_KEY` (DeepSeek), a running Ollama instance (`OLLAMA_BASE_URL`), AWS credentials via the standard provider chain (Bedrock - no API key needed, uses IAM), or the `claude` CLI binary (Claude Code - no API key needed, uses your Claude subscription). The `--dedup-llm` flag uses the same key. - **Data residency** — `graphify extract` auto-detects which provider to use based on which API key is set (priority: Gemini → Kimi → Claude → OpenAI → DeepSeek → Bedrock → Ollama). For code with data-residency requirements, use `--backend ollama` (fully local) or pass an explicit `--backend` flag. Kimi (`MOONSHOT_API_KEY`) routes to Moonshot AI servers in China. @@ -435,7 +439,7 @@ graphify query "..." Run `graphify hook install` — it sets up a git merge driver that union-merges `graph.json` automatically so conflicts never happen. **Extraction returns empty nodes/edges for docs or PDFs** -Docs and PDFs require an LLM call. Check that your API key is set and the backend is correct: +Docs, PDFs, and images require an LLM call — code-only corpora need no key. Check that your API key is set and the backend is correct: ```bash ANTHROPIC_API_KEY=sk-... graphify extract ./docs --backend claude ``` diff --git a/graphify/__main__.py b/graphify/__main__.py index 3d70bff..1813af2 100644 --- a/graphify/__main__.py +++ b/graphify/__main__.py @@ -583,7 +583,38 @@ def _replace_or_append_section(content: str, marker: str, new_section: str) -> s return out +def _print_banner() -> None: + """Amber brain banner on graphify install. TTY-only, never raises.""" + if not sys.stdout.isatty(): + return + try: + if sys.platform == "win32": + import ctypes + ctypes.windll.kernel32.SetConsoleMode( + ctypes.windll.kernel32.GetStdHandle(-11), 7 + ) + A = "\033[38;5;214m" + D = "\033[38;5;130m" + R = "\033[0m" + print(f"""{A} + ╭──◉──╮ ╭──◉──╮ + ╱ ◉ ◉ ╲ ╱ ◉ ◉ ╲ +│ ◉─◉─◉ ◉ ◉─◉─◉ │ +│ ◉ ◉ │ ◉ ◉ │ +│ ◉─◉─◉ ◉ ◉─◉─◉ │ + ╲ ◉ ◉ ╱ ╲ ◉ ◉ ╱ + ╰──◉──╯ ╰──◉──╯ + ◉ + + █▀▀ █▀█ ▄▀█ █▀█ █ █ █ █▀▀ █▄█ + █▄█ █▀▄ █▀█ █▀▀ █▀█ █ █▀ █{D} {__version__}{R} +""") + except Exception: + pass + + def install(platform: str = "claude", *, project: bool = False, project_dir: Path | None = None) -> None: + _print_banner() if platform == "gemini": gemini_install(project_dir=project_dir, project=project) return @@ -935,12 +966,11 @@ def _kiro_install(project_dir: Path) -> None: """Write graphify skill + steering file for Kiro IDE/CLI.""" project_dir = project_dir or Path(".") - # Skill file → .kiro/skills/graphify/SKILL.md - skill_src = Path(__file__).parent / "skill-kiro.md" - skill_dst = project_dir / ".kiro" / "skills" / "graphify" / "SKILL.md" - skill_dst.parent.mkdir(parents=True, exist_ok=True) - skill_dst.write_text(skill_src.read_text(encoding="utf-8"), encoding="utf-8") - print(f" {skill_dst.relative_to(project_dir)} -> /graphify skill") + # Skill file + references/ sidecar + .graphify_version stamp via the shared + # progressive-disclosure helper. Previously this used a bare write_text that + # bypassed _copy_skill_file, so the references/ dir and version stamp were + # never written even though kiro declares skill_refs: "kiro" (#1142). + _copy_skill_file("kiro", project=True, project_dir=project_dir) # Steering file → .kiro/steering/graphify.md (always-on) steering_dir = project_dir / ".kiro" / "steering" @@ -965,15 +995,10 @@ def _kiro_uninstall(project_dir: Path) -> None: project_dir = project_dir or Path(".") removed = [] - skill_dst = project_dir / ".kiro" / "skills" / "graphify" / "SKILL.md" - if skill_dst.exists(): - skill_dst.unlink() + # Skill + .graphify_version + references/ sidecar + empty-dir walk. + skill_dst = _platform_skill_destination("kiro", project=True, project_dir=project_dir) + if _remove_skill_file("kiro", project=True, project_dir=project_dir): removed.append(str(skill_dst.relative_to(project_dir))) - # Remove parent dir if empty - try: - skill_dst.parent.rmdir() - except OSError: - pass steering_dst = project_dir / ".kiro" / "steering" / "graphify.md" if steering_dst.exists(): @@ -3949,91 +3974,6 @@ def main() -> None: if cli_max_workers is not None: os.environ["GRAPHIFY_MAX_WORKERS"] = str(cli_max_workers) - # Backend resolution. If user did not pass --backend, sniff env. - # If backend was explicitly requested, validate its key is present - # and surface a clear error early — don't let extract_corpus_parallel - # raise mid-run after we've spent time on AST extraction. - from graphify.llm import ( - BACKENDS as _BACKENDS, - detect_backend as _detect_backend, - estimate_cost as _estimate_cost, - extract_corpus_parallel as _extract_corpus_parallel, - _format_backend_env_keys, - _get_backend_api_key, - ) - if backend is None: - backend = _detect_backend() - if backend is None: - print( - "error: no LLM API key found. Set GEMINI_API_KEY or GOOGLE_API_KEY " - "(gemini), MOONSHOT_API_KEY (kimi), ANTHROPIC_API_KEY (claude), " - "OPENAI_API_KEY (openai), DEEPSEEK_API_KEY (deepseek), " - "or pass --backend.", - file=sys.stderr, - ) - sys.exit(1) - if backend not in _BACKENDS: - print( - f"error: unknown backend '{backend}'. " - f"Available: {', '.join(sorted(_BACKENDS))}", - file=sys.stderr, - ) - sys.exit(1) - if backend == "ollama": - # Fail closed with a clean message (not a deep traceback) if - # OLLAMA_BASE_URL points at a link-local/metadata address. warn=False: - # the later in-flow call owns the user-facing warning for LAN hosts. - from graphify.llm import _validate_ollama_base_url - _oll_url = os.environ.get("OLLAMA_BASE_URL", _BACKENDS["ollama"].get("base_url", "")) - try: - _validate_ollama_base_url(_oll_url, warn=False) - except ValueError as exc: - print(f"error: {exc}", file=sys.stderr) - sys.exit(2) - if not _get_backend_api_key(backend): - # Ollama on a loopback URL ignores auth entirely; don't block - # the run just because OLLAMA_API_KEY is unset (issue #792). - # extract_files_direct already prints a warning and substitutes - # a placeholder key in that case. - allow_no_key = False - if backend == "ollama": - from urllib.parse import urlparse - ollama_url = os.environ.get( - "OLLAMA_BASE_URL", - _BACKENDS["ollama"].get("base_url", ""), - ) - try: - host = (urlparse(ollama_url).hostname or "").lower() - except Exception: - host = "" - allow_no_key = ( - host in ("localhost", "127.0.0.1", "::1") - or host.startswith("127.") - ) - elif backend == "bedrock": - allow_no_key = bool( - os.environ.get("AWS_PROFILE") - or os.environ.get("AWS_REGION") - or os.environ.get("AWS_DEFAULT_REGION") - or os.environ.get("AWS_ACCESS_KEY_ID") - ) - elif backend == "claude-cli": - import shutil as _shutil - allow_no_key = _shutil.which("claude") is not None - if not allow_no_key: - print( - "error: backend 'claude-cli' requires the `claude` CLI on $PATH " - "(install Claude Code and run `claude` once to authenticate).", - file=sys.stderr, - ) - sys.exit(1) - if not allow_no_key: - print( - f"error: backend '{backend}' requires {_format_backend_env_keys(backend)} to be set.", - file=sys.stderr, - ) - sys.exit(1) - # Resolve output dir. The user-facing contract is "/graphify-out/" # so a fresh checkout writes graphify-out/ at the project root, matching # the skill.md pipeline. @@ -4093,6 +4033,93 @@ def main() -> None: f"{len(image_files)} images" ) + # Resolve the LLM backend only now that we know whether the corpus + # needs one. A code-only corpus is pure local AST and must not require + # an API key; the key is enforced below only when there's LLM work. + from graphify.llm import ( + BACKENDS as _BACKENDS, + detect_backend as _detect_backend, + estimate_cost as _estimate_cost, + extract_corpus_parallel as _extract_corpus_parallel, + _format_backend_env_keys, + _get_backend_api_key, + ) + needs_llm = bool(semantic_files) or dedup_llm + if backend is None and needs_llm: + backend = _detect_backend() + if backend is not None and backend not in _BACKENDS: + print( + f"error: unknown backend '{backend}'. " + f"Available: {', '.join(sorted(_BACKENDS))}", + file=sys.stderr, + ) + sys.exit(1) + if needs_llm: + if backend is None: + reasons = [] + if semantic_files: + reasons.append( + f"{len(semantic_files)} doc/paper/image file(s) need semantic extraction" + ) + if dedup_llm: + reasons.append("--dedup-llm was passed") + print( + "error: no LLM API key found (" + "; ".join(reasons) + "). " + "Set GEMINI_API_KEY or GOOGLE_API_KEY (gemini), MOONSHOT_API_KEY " + "(kimi), ANTHROPIC_API_KEY (claude), OPENAI_API_KEY (openai), " + "DEEPSEEK_API_KEY (deepseek), or pass --backend. A code-only " + "corpus needs no key.", + file=sys.stderr, + ) + sys.exit(1) + if backend == "ollama": + from graphify.llm import _validate_ollama_base_url + _oll_url = os.environ.get("OLLAMA_BASE_URL", _BACKENDS["ollama"].get("base_url", "")) + try: + _validate_ollama_base_url(_oll_url, warn=False) + except ValueError as exc: + print(f"error: {exc}", file=sys.stderr) + sys.exit(2) + if not _get_backend_api_key(backend): + allow_no_key = False + if backend == "ollama": + from urllib.parse import urlparse + ollama_url = os.environ.get( + "OLLAMA_BASE_URL", + _BACKENDS["ollama"].get("base_url", ""), + ) + try: + host = (urlparse(ollama_url).hostname or "").lower() + except Exception: + host = "" + allow_no_key = ( + host in ("localhost", "127.0.0.1", "::1") + or host.startswith("127.") + ) + elif backend == "bedrock": + allow_no_key = bool( + os.environ.get("AWS_PROFILE") + or os.environ.get("AWS_REGION") + or os.environ.get("AWS_DEFAULT_REGION") + or os.environ.get("AWS_ACCESS_KEY_ID") + ) + elif backend == "claude-cli": + import shutil as _shutil + allow_no_key = _shutil.which("claude") is not None + if not allow_no_key: + print( + "error: backend 'claude-cli' requires the `claude` CLI on $PATH " + "(install Claude Code and run `claude` once to authenticate).", + file=sys.stderr, + ) + sys.exit(1) + if not allow_no_key: + print( + f"error: backend '{backend}' requires {_format_backend_env_keys(backend)} to be set.", + file=sys.stderr, + ) + sys.exit(1) + # AST extraction on code files. Empty code list (docs-only corpus) is # the issue #698 case — skip cleanly instead of crashing inside extract(). ast_result: dict = {"nodes": [], "edges": [], "input_tokens": 0, "output_tokens": 0} diff --git a/graphify/detect.py b/graphify/detect.py index 67dcf6e..e15a413 100644 --- a/graphify/detect.py +++ b/graphify/detect.py @@ -25,7 +25,7 @@ class FileType(str, Enum): _MANIFEST_PATH = "graphify-out/manifest.json" -CODE_EXTENSIONS = {'.py', '.ts', '.tsx', '.js', '.jsx', '.mjs', '.ejs', '.ets', '.go', '.rs', '.java', '.groovy', '.gradle', '.cpp', '.cc', '.cxx', '.c', '.h', '.hpp', '.rb', '.swift', '.kt', '.kts', '.cs', '.scala', '.php', '.lua', '.luau', '.toc', '.zig', '.ps1', '.ex', '.exs', '.m', '.mm', '.jl', '.vue', '.svelte', '.astro', '.dart', '.v', '.sv', '.svh', '.sql', '.r', '.f', '.F', '.f90', '.F90', '.f95', '.F95', '.f03', '.F03', '.f08', '.F08', '.pas', '.pp', '.dpr', '.dpk', '.lpr', '.inc', '.dfm', '.lfm', '.lpk', '.sh', '.bash', '.json', '.dm', '.dme', '.dmi', '.dmm', '.dmf', '.sln', '.csproj', '.fsproj', '.vbproj', '.razor', '.cshtml'} +CODE_EXTENSIONS = {'.py', '.ts', '.tsx', '.js', '.jsx', '.mjs', '.ejs', '.ets', '.go', '.rs', '.java', '.groovy', '.gradle', '.cpp', '.cc', '.cxx', '.c', '.h', '.hpp', '.rb', '.swift', '.kt', '.kts', '.cs', '.scala', '.php', '.lua', '.luau', '.toc', '.zig', '.ps1', '.ex', '.exs', '.m', '.mm', '.jl', '.vue', '.svelte', '.astro', '.dart', '.v', '.sv', '.svh', '.sql', '.r', '.f', '.F', '.f90', '.F90', '.f95', '.F95', '.f03', '.F03', '.f08', '.F08', '.pas', '.pp', '.dpr', '.dpk', '.lpr', '.inc', '.dfm', '.lfm', '.lpk', '.sh', '.bash', '.json', '.tf', '.tfvars', '.hcl', '.dm', '.dme', '.dmi', '.dmm', '.dmf', '.sln', '.csproj', '.fsproj', '.vbproj', '.razor', '.cshtml'} DOC_EXTENSIONS = {'.md', '.mdx', '.qmd', '.txt', '.rst', '.html', '.yaml', '.yml'} PAPER_EXTENSIONS = {'.pdf'} IMAGE_EXTENSIONS = {'.png', '.jpg', '.jpeg', '.gif', '.webp', '.svg'} diff --git a/graphify/extract.py b/graphify/extract.py index a56a780..6ec23a0 100644 --- a/graphify/extract.py +++ b/graphify/extract.py @@ -10526,6 +10526,184 @@ def extract_dmf(path: Path) -> dict: return {"nodes": nodes, "edges": edges} +# Head tokens in an HCL traversal that are meta/builtins, not references to a +# block defined in the corpus (count.index, each.key, self.*, path.module, ...). +_TF_META_HEADS = frozenset({"count", "each", "self", "path", "terraform"}) + + +def extract_terraform(path: Path) -> dict: + """Extract Terraform/HCL blocks and the references between them via tree-sitter. + + Nodes: resources, data sources, modules, variables, outputs, providers, and + locals. Edges: `contains` (file -> block), `references` (block -> the blocks + it interpolates, e.g. `aws_instance.web` -> `var.region`), and `depends_on` + (explicit dependency edges). + + Node IDs are scoped by the parent directory, not the file stem, because + Terraform resources are module(directory)-scoped: a resource defined in + main.tf is referenced from other .tf files in the same directory. Directory + scoping lets those cross-file references resolve when per-file extractions + are merged (stem scoping would split a definition from its references). + """ + try: + import tree_sitter_hcl as tshcl + from tree_sitter import Language, Parser + except ImportError: + return {"nodes": [], "edges": [], "error": "tree_sitter_hcl not installed. Run: pip install tree-sitter-hcl"} + + try: + language = Language(tshcl.language()) + parser = Parser(language) + source = path.read_bytes() + tree = parser.parse(source) + root = tree.root_node + except Exception as e: + return {"nodes": [], "edges": [], "error": str(e)} + + str_path = str(path) + file_nid = _make_id(str_path) + scope = path.parent.name or "tf" + + nodes: list[dict] = [{"id": file_nid, "label": path.name, "file_type": "code", + "source_file": str_path, "source_location": None}] + edges: list[dict] = [] + seen_ids: set[str] = {file_nid} + seen_edges: set[tuple[str, str, str]] = set() + + def _read(n) -> str: + return source[n.start_byte:n.end_byte].decode("utf-8", errors="replace") + + def _label_text(n) -> str: + return _read(n).strip().strip('"') + + def _add_node(address: str, label: str, line: int) -> str: + nid = _make_id(scope, address) + if nid not in seen_ids: + seen_ids.add(nid) + nodes.append({"id": nid, "label": label, "file_type": "code", + "source_file": str_path, "source_location": f"L{line}"}) + edges.append({"source": file_nid, "target": nid, "relation": "contains", + "confidence": "EXTRACTED", "source_file": str_path, + "source_location": f"L{line}", "weight": 1.0}) + return nid + + def _add_edge(src: str, address: str, relation: str, line: int) -> None: + tgt = _make_id(scope, address) + if src == tgt: + return + key = (src, tgt, relation) + if key in seen_edges: + return + seen_edges.add(key) + edges.append({"source": src, "target": tgt, "relation": relation, + "confidence": "EXTRACTED", "source_file": str_path, + "source_location": f"L{line}", "weight": 1.0}) + + def _block_parts(block) -> tuple: + btype = None + labels: list[str] = [] + for c in block.children: + if c.type in ("block_start", "body", "block_end"): + break + if c.type == "identifier" and btype is None: + btype = _read(c) + elif c.type in ("string_lit", "identifier"): + labels.append(_label_text(c)) + return btype, labels + + def _ref_address(expr): + head = _read(expr) + parent = expr.parent + attrs: list[str] = [] + if parent is not None: + seen_self = False + for c in parent.children: + if c.id == expr.id: + seen_self = True + continue + if seen_self and c.type == "get_attr": + name = None + for gc in c.children: + if gc.type == "identifier": + name = _read(gc) + break + if name is None: + break + attrs.append(name) + elif seen_self and c.type not in ("get_attr",): + break + if head in _TF_META_HEADS or not head: + return None + if head == "var": + return f"var.{attrs[0]}" if attrs else None + if head == "local": + return f"local.{attrs[0]}" if attrs else None + if head == "module": + return f"module.{attrs[0]}" if attrs else None + if head == "data": + return f"data.{attrs[0]}.{attrs[1]}" if len(attrs) >= 2 else None + return f"{head}.{attrs[0]}" if attrs else None + + def _collect_refs(node, owner_nid: str, relation: str) -> None: + rel = relation + if node.type == "attribute": + key_node = node.child_by_field_name("key") or ( + node.children[0] if node.children else None + ) + if key_node is not None and _read(key_node) == "depends_on": + rel = "depends_on" + if node.type == "variable_expr": + addr = _ref_address(node) + if addr: + _add_edge(owner_nid, addr, rel, node.start_point[0] + 1) + for c in node.children: + if c.is_named: + _collect_refs(c, owner_nid, rel) + + def _body_of(block): + for c in block.children: + if c.type == "body": + return c + return None + + body = next((c for c in root.children if c.type == "body"), root) + for block in body.children: + if block.type != "block": + continue + btype, labels = _block_parts(block) + line = block.start_point[0] + 1 + blk_body = _body_of(block) + if btype == "resource" and len(labels) >= 2: + owner = _add_node(f"{labels[0]}.{labels[1]}", f"{labels[0]}.{labels[1]}", line) + elif btype == "data" and len(labels) >= 2: + owner = _add_node(f"data.{labels[0]}.{labels[1]}", f"data.{labels[0]}.{labels[1]}", line) + elif btype == "module" and labels: + owner = _add_node(f"module.{labels[0]}", f"module.{labels[0]}", line) + elif btype == "variable" and labels: + owner = _add_node(f"var.{labels[0]}", f"var.{labels[0]}", line) + elif btype == "output" and labels: + owner = _add_node(f"output.{labels[0]}", f"output.{labels[0]}", line) + elif btype == "provider" and labels: + owner = _add_node(f"provider.{labels[0]}", f"provider.{labels[0]}", line) + elif btype == "locals" and blk_body is not None: + for attr in blk_body.children: + if attr.type != "attribute": + continue + key_node = attr.children[0] if attr.children else None + if key_node is None: + continue + key = _read(key_node) + lnid = _add_node(f"local.{key}", f"local.{key}", attr.start_point[0] + 1) + _collect_refs(attr, lnid, "references") + continue + else: + continue + if blk_body is not None: + _collect_refs(blk_body, owner, "references") + + return {"nodes": nodes, "edges": edges} + + _DISPATCH: dict[str, Any] = { ".py": extract_python, ".js": extract_js, @@ -10594,6 +10772,9 @@ _DISPATCH: dict[str, Any] = { ".sh": extract_bash, ".bash": extract_bash, ".json": extract_json, + ".tf": extract_terraform, + ".tfvars": extract_terraform, + ".hcl": extract_terraform, ".dm": extract_dm, ".dme": extract_dm, ".dmi": extract_dmi, diff --git a/pyproject.toml b/pyproject.toml index 2166904..96a6950 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -4,7 +4,7 @@ build-backend = "setuptools.build_meta" [project] name = "graphifyy" -version = "0.8.31" +version = "0.8.32" description = "AI coding assistant skill (Claude Code, CodeBuddy, Codex, OpenCode, Kilo Code, Cursor, Gemini CLI, Aider, OpenClaw, Factory Droid, Trae, Hermes, Kiro, Pi, Devin CLI, Google Antigravity) - turn any folder of code, docs, papers, images, or videos into a queryable knowledge graph" readme = "README.md" license = { file = "LICENSE" } @@ -69,7 +69,8 @@ sql = ["tree-sitter-sql"] # must compile from source (needs a C toolchain + python3-dev). Keeping it optional # avoids breaking the default `uv tool install graphifyy` for everyone (#1104). dm = ["tree-sitter-dm"] -all = ["mcp", "neo4j", "pypdf", "markdownify", "watchdog", "graspologic; python_version < '3.13'", "python-docx", "openpyxl", "faster-whisper; python_version >= '3.11'", "yt-dlp", "matplotlib", "openai", "tiktoken", "boto3", "anthropic", "tree-sitter-sql", "jieba", "tree-sitter-dm"] +terraform = ["tree-sitter-hcl"] +all = ["mcp", "neo4j", "pypdf", "markdownify", "watchdog", "graspologic; python_version < '3.13'", "python-docx", "openpyxl", "faster-whisper; python_version >= '3.11'", "yt-dlp", "matplotlib", "openai", "tiktoken", "boto3", "anthropic", "tree-sitter-sql", "jieba", "tree-sitter-dm", "tree-sitter-hcl"] [project.scripts] graphify = "graphify.__main__:main" @@ -91,6 +92,7 @@ dev = [ "setuptools>=82.0.1", "wheel>=0.47.0", "tomli>=2.0 ; python_version < '3.11'", + "tree-sitter-hcl>=1.2.0", ] [tool.uv] diff --git a/tests/test_extract_cli.py b/tests/test_extract_cli.py index 6998bcc..13d6fe9 100644 --- a/tests/test_extract_cli.py +++ b/tests/test_extract_cli.py @@ -121,3 +121,80 @@ def test_extract_succeeds_when_at_least_one_chunk_completes( assert (out_dir / "graphify-out" / "graph.json").exists(), ( "graph.json must be written on the happy path" ) + + +def _code_only_corpus(tmp_path): + """A corpus with only code — no docs/papers/images.""" + (tmp_path / "auth.py").write_text( + "def login(user):\n return validate(user)\n\n" + "def validate(user):\n return True\n" + ) + return tmp_path + + +def _clear_backend_keys(monkeypatch): + """Clear every env var that detect_backend() or _get_backend_api_key() reads.""" + for key in ( + "GEMINI_API_KEY", "GOOGLE_API_KEY", "OPENAI_API_KEY", + "ANTHROPIC_API_KEY", "DEEPSEEK_API_KEY", "MOONSHOT_API_KEY", + # bedrock: presence of any of these is treated as a valid credential + "AWS_PROFILE", "AWS_REGION", "AWS_DEFAULT_REGION", "AWS_ACCESS_KEY_ID", + # ollama: a set OLLAMA_BASE_URL triggers backend detection + "OLLAMA_BASE_URL", + ): + monkeypatch.delenv(key, raising=False) + + +def test_extract_codeonly_succeeds_without_api_key(monkeypatch, tmp_path): + """A code-only corpus must run with no LLM API key. + + Regression: graphify extract validated a backend upfront and exited 1 with + 'no LLM API key found' even for a code-only corpus that never calls a model. + The keyless AST path now runs to a written graph.json (#1122). + """ + corpus = _code_only_corpus(tmp_path) + out_dir = tmp_path / "out" + _clear_backend_keys(monkeypatch) + monkeypatch.setattr(mainmod, "_check_skill_version", lambda _: None) + monkeypatch.setattr( + mainmod.sys, "argv", + ["graphify", "extract", str(corpus), "--out", str(out_dir)], + ) + + try: + mainmod.main() + except SystemExit as exc: + assert exc.code in (None, 0), f"unexpected exit code {exc.code}" + + graph = out_dir / "graphify-out" / "graph.json" + assert graph.exists(), "code-only extract must write graph.json without a key" + import json + assert len(json.loads(graph.read_text()).get("nodes", [])) > 0 + + +def test_extract_without_key_still_errors_when_docs_present( + monkeypatch, tmp_path, capsys +): + """Key requirement still fires when semantic work is needed. + + A corpus with a Markdown doc needs LLM semantic extraction, so a keyless + extract must exit 1 with clear guidance (#1122). + """ + corpus = _make_corpus(tmp_path) # includes a Markdown doc + out_dir = tmp_path / "out" + _clear_backend_keys(monkeypatch) + # Patch detect_backend too so ambient AWS/ollama env can't slip through. + monkeypatch.setattr("graphify.llm.detect_backend", lambda: None) + monkeypatch.setattr(mainmod, "_check_skill_version", lambda _: None) + monkeypatch.setattr( + mainmod.sys, "argv", + ["graphify", "extract", str(corpus), "--out", str(out_dir)], + ) + + with pytest.raises(SystemExit) as exc_info: + mainmod.main() + assert exc_info.value.code == 1 + err = capsys.readouterr().err + assert "no LLM API key found" in err + assert "code-only corpus needs no key" in err + assert not (out_dir / "graphify-out" / "graph.json").exists() diff --git a/tests/test_install_upgrade.py b/tests/test_install_upgrade.py index 09ee3d8..13e7c55 100644 --- a/tests/test_install_upgrade.py +++ b/tests/test_install_upgrade.py @@ -231,3 +231,42 @@ def test_kiro_install_upgrades_stale_steering(tmp_path, monkeypatch): assert "read it before answering architecture questions" not in after _assert_query_first(after, ".kiro/steering/graphify.md") assert "inclusion: always" in after # frontmatter preserved + + +def test_kiro_install_ships_references_sidecar_and_version_stamp(tmp_path, monkeypatch): + """_kiro_install routes through _copy_skill_file so the references/ sidecar + and .graphify_version stamp are written alongside SKILL.md (#1142). + Previously it used a bare write_text that bypassed the shared helper.""" + monkeypatch.setattr(mainmod, "_check_skill_version", lambda _: None) + + refs_dir = Path(mainmod.__file__).parent / "skills" / "kiro" / "references" + if not refs_dir.exists(): + pytest.skip("kiro references bundle not present in this checkout") + + mainmod._kiro_install(tmp_path) + + skill_dir = tmp_path / ".kiro" / "skills" / "graphify" + + # SKILL.md present + assert (skill_dir / "SKILL.md").exists() + + # references/ sidecar installed with at least one fragment + refs_dst = skill_dir / "references" + assert refs_dst.is_dir(), "references/ sidecar must be installed (#1142)" + assert any(refs_dst.iterdir()), "references/ must not be empty" + + # .graphify_version stamp written + version_file = skill_dir / ".graphify_version" + assert version_file.exists(), ".graphify_version stamp must be written (#1142)" + assert version_file.read_text(encoding="utf-8") == mainmod.__version__ + + # no references.tmp leftover + assert not (skill_dir / "references.tmp").exists() + + # steering file still written + assert (tmp_path / ".kiro" / "steering" / "graphify.md").exists() + + # uninstall removes skill dir, version stamp, references/, and steering file + mainmod._kiro_uninstall(tmp_path) + assert not skill_dir.exists(), "uninstall must remove skill dir including references/ (#1142)" + assert not (tmp_path / ".kiro" / "steering" / "graphify.md").exists() diff --git a/tests/test_terraform.py b/tests/test_terraform.py new file mode 100644 index 0000000..2049078 --- /dev/null +++ b/tests/test_terraform.py @@ -0,0 +1,158 @@ +"""Tests for the Terraform/HCL extractor (graphify/extract.py, issue #187).""" +from __future__ import annotations + +from pathlib import Path + +from graphify.build import build_from_json +from graphify.extract import extract_terraform + + +def _write(tmp_path: Path, name: str, body: str) -> Path: + p = tmp_path / name + p.write_text(body, encoding="utf-8") + return p + + +def _labels(r) -> list[str]: + return [n["label"] for n in r["nodes"]] + + +def _rel_pairs(r, relation: str) -> set[tuple[str, str]]: + lab = {n["id"]: n["label"] for n in r["nodes"]} + return { + (lab.get(e["source"], e["source"]), lab.get(e["target"], e["target"])) + for e in r["edges"] + if e["relation"] == relation + } + + +SAMPLE = """\ +# leading comment so the body is not children[0] +terraform { + required_providers { azurerm = { source = "hashicorp/azurerm" } } +} + +variable "region" { default = "us-east-1" } + +provider "aws" { region = var.region } + +data "aws_ami" "ubuntu" { most_recent = true } + +resource "aws_instance" "web" { + ami = data.aws_ami.ubuntu.id + subnet_id = var.region + depends_on = [aws_security_group.sg] +} + +resource "aws_security_group" "sg" { name = "sg" } + +module "vpc" { + source = "./modules/vpc" + cidr = local.cidr +} + +locals { cidr = "10.0.0.0/16" } + +output "ip" { value = aws_instance.web.private_ip } +""" + + +def test_no_error_and_all_block_types_become_nodes(tmp_path): + r = extract_terraform(_write(tmp_path, "main.tf", SAMPLE)) + assert r.get("error") is None + labels = set(_labels(r)) + # one node per block type (the terraform{} settings block is intentionally skipped) + for expected in ( + "var.region", + "provider.aws", + "data.aws_ami.ubuntu", + "aws_instance.web", + "aws_security_group.sg", + "module.vpc", + "local.cidr", + "output.ip", + ): + assert expected in labels, f"missing node {expected!r}" + + +def test_reference_edges(tmp_path): + r = extract_terraform(_write(tmp_path, "main.tf", SAMPLE)) + refs = _rel_pairs(r, "references") + assert ("provider.aws", "var.region") in refs + assert ("aws_instance.web", "data.aws_ami.ubuntu") in refs + assert ("aws_instance.web", "var.region") in refs + assert ("module.vpc", "local.cidr") in refs + assert ("output.ip", "aws_instance.web") in refs + + +def test_depends_on_edge(tmp_path): + r = extract_terraform(_write(tmp_path, "main.tf", SAMPLE)) + assert ("aws_instance.web", "aws_security_group.sg") in _rel_pairs(r, "depends_on") + + +def test_file_contains_blocks(tmp_path): + r = extract_terraform(_write(tmp_path, "main.tf", SAMPLE)) + contains = _rel_pairs(r, "contains") + assert ("main.tf", "aws_instance.web") in contains + assert ("main.tf", "var.region") in contains + + +def test_meta_heads_not_emitted(tmp_path): + # count.index / each.key / self.* / path.module are builtins, not references. + body = """\ +resource "aws_instance" "web" { + count = 2 + name = "web-${count.index}" + tags = each.value + dir = path.module +} +""" + r = extract_terraform(_write(tmp_path, "main.tf", body)) + targets = {t for _, t in _rel_pairs(r, "references")} + assert not any(t.startswith(("count", "each", "path")) for t in targets) + + +def test_cross_file_references_resolve_after_merge(tmp_path): + # A resource defined in one file is referenced from another in the same + # directory; directory-scoped IDs must let the edge resolve at build time. + defn = """\ +resource "azurerm_resource_group" "main" { name = "rg" } +""" + user = """\ +resource "azurerm_network_interface" "nic" { + resource_group_name = azurerm_resource_group.main.name +} +""" + r_defn = extract_terraform(_write(tmp_path, "main.tf", defn)) + r_user = extract_terraform(_write(tmp_path, "nic.tf", user)) + + # The cross-file edge target id equals the definition's node id. + rg_id = next(n["id"] for n in r_defn["nodes"] if n["label"] == "azurerm_resource_group.main") + nic_ref_targets = {e["target"] for e in r_user["edges"] if e["relation"] == "references"} + assert rg_id in nic_ref_targets + + # And it survives a real merge: the edge is present (not dropped as dangling). + G = build_from_json( + { + "nodes": r_defn["nodes"] + r_user["nodes"], + "edges": r_defn["edges"] + r_user["edges"], + } + ) + nic_id = next(n["id"] for n in r_user["nodes"] if n["label"] == "azurerm_network_interface.nic") + assert G.has_edge(nic_id, rg_id) + + +def test_empty_and_commentonly_files_are_safe(tmp_path): + assert extract_terraform(_write(tmp_path, "a.tf", "")).get("error") is None + r = extract_terraform(_write(tmp_path, "b.tf", "# just a comment\n")) + # only the file node, no crash + assert len(r["nodes"]) == 1 + + +def test_tfvars_key_value_is_safe(tmp_path): + # .tfvars files contain only key=value assignments (no block structure), + # so extract_terraform produces zero block nodes — only the file node. + # This is the documented intended behaviour for .tfvars. + r = extract_terraform(_write(tmp_path, "terraform.tfvars", 'region = "us-east-1"\nenv = "prod"\n')) + assert r.get("error") is None + assert len(r["nodes"]) == 1 # only the file node, no variable nodes diff --git a/uv.lock b/uv.lock index a2a2b4e..be22019 100644 --- a/uv.lock +++ b/uv.lock @@ -1189,6 +1189,7 @@ all = [ { name = "python-docx" }, { name = "tiktoken" }, { name = "tree-sitter-dm" }, + { name = "tree-sitter-hcl" }, { name = "tree-sitter-sql" }, { name = "watchdog" }, { name = "yt-dlp" }, @@ -1246,6 +1247,9 @@ sql = [ svg = [ { name = "matplotlib" }, ] +terraform = [ + { name = "tree-sitter-hcl" }, +] video = [ { name = "faster-whisper", marker = "python_full_version >= '3.11'" }, { name = "yt-dlp" }, @@ -1270,6 +1274,7 @@ dev = [ { name = "safety" }, { name = "setuptools" }, { name = "tomli", marker = "python_full_version < '3.11'" }, + { name = "tree-sitter-hcl" }, { name = "wheel" }, ] @@ -1323,6 +1328,8 @@ requires-dist = [ { name = "tree-sitter-fortran" }, { name = "tree-sitter-go" }, { name = "tree-sitter-groovy" }, + { name = "tree-sitter-hcl", marker = "extra == 'all'" }, + { name = "tree-sitter-hcl", marker = "extra == 'terraform'" }, { name = "tree-sitter-java" }, { name = "tree-sitter-javascript" }, { name = "tree-sitter-json" }, @@ -1347,7 +1354,7 @@ requires-dist = [ { name = "yt-dlp", marker = "extra == 'all'" }, { name = "yt-dlp", marker = "extra == 'video'" }, ] -provides-extras = ["mcp", "neo4j", "pdf", "watch", "svg", "leiden", "office", "google", "video", "kimi", "ollama", "bedrock", "anthropic", "gemini", "openai", "chinese", "sql", "dm", "all"] +provides-extras = ["mcp", "neo4j", "pdf", "watch", "svg", "leiden", "office", "google", "video", "kimi", "ollama", "bedrock", "anthropic", "gemini", "openai", "chinese", "sql", "dm", "terraform", "all"] [package.metadata.requires-dev] dev = [ @@ -1365,6 +1372,7 @@ dev = [ { name = "safety", specifier = ">=3.7.0" }, { name = "setuptools", specifier = ">=82.0.1" }, { name = "tomli", marker = "python_full_version < '3.11'", specifier = ">=2.0" }, + { name = "tree-sitter-hcl", specifier = ">=1.2.0" }, { name = "wheel", specifier = ">=0.47.0" }, ] @@ -4638,6 +4646,21 @@ wheels = [ { url = "https://files.pythonhosted.org/packages/9d/e3/50c719d09a4495672226b2359b2701360fdef022bc86dedef9fc16d3959c/tree_sitter_groovy-0.1.2-cp39-abi3-win_arm64.whl", hash = "sha256:1942a9a1b22e154da9bbf1b03e6b4dbec4211b1109d24bcf4c12b006cbc04037", size = 102508, upload-time = "2024-11-19T04:33:06.101Z" }, ] +[[package]] +name = "tree-sitter-hcl" +version = "1.2.0" +source = { registry = "https://pypi.org/simple" } +sdist = { url = "https://files.pythonhosted.org/packages/06/c9/ed79f643b0cec3e123171c09caffb6088a6111025a20fc69112b1468828b/tree_sitter_hcl-1.2.0.tar.gz", hash = "sha256:f86cb7a9fd5cb93d83e2f788ae155544464c47755d09190505de562c0d6ad1dd", size = 55207, upload-time = "2025-06-16T13:36:56.159Z" } +wheels = [ + { url = "https://files.pythonhosted.org/packages/77/34/7ccb58107ae0d38e0a8f05dedbd990780eb7d429f227725c15fee314cdcd/tree_sitter_hcl-1.2.0-cp310-abi3-macosx_10_9_x86_64.whl", hash = "sha256:9ae35084a3dc12272f941b424eadd8a44cf2e0e9345b020330cf8db6f67d3524", size = 30604, upload-time = "2025-06-16T13:36:49.775Z" }, + { url = "https://files.pythonhosted.org/packages/8e/8b/7618448cde58ca6fbefcf210ef98d7c4bd7d2b54b3e3d5cddd947c804a18/tree_sitter_hcl-1.2.0-cp310-abi3-macosx_11_0_arm64.whl", hash = "sha256:03678762e8b78d717187848edebed95e4c31a54e14f24dec97555f47fb440e28", size = 31163, upload-time = "2025-06-16T13:36:50.936Z" }, + { url = "https://files.pythonhosted.org/packages/12/35/b8f87fffb5527c85f5f292486e53f64963846b94cf3ea258f4b850480f18/tree_sitter_hcl-1.2.0-cp310-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:d2e9f3cfb6694e33f5b880e74bed842398cbacd21024251e2ec90f19dee6a64d", size = 54139, upload-time = "2025-06-16T13:36:51.951Z" }, + { url = "https://files.pythonhosted.org/packages/ee/0a/01bb627044d273e8e506edff8ab773e562ba447b5790b789f62e47a5e754/tree_sitter_hcl-1.2.0-cp310-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:915763da6630610c2efb7afe13145f50feb8043732a74f9bae78811212578d3d", size = 54317, upload-time = "2025-06-16T13:36:52.65Z" }, + { url = "https://files.pythonhosted.org/packages/f4/40/cc843f07210caa0a2e2a2b3581f93c917ed2139cb0851b32f4852141c95c/tree_sitter_hcl-1.2.0-cp310-abi3-musllinux_1_2_x86_64.whl", hash = "sha256:5b6c6aaccfca2fde4fcce52aa88d8c3756f44d407b8584fc3cff581d8765b4b7", size = 51638, upload-time = "2025-06-16T13:36:53.684Z" }, + { url = "https://files.pythonhosted.org/packages/27/a6/f15096c138eccc0c68b7254c75a8121bab326b62920f2d259079b2d6a7d0/tree_sitter_hcl-1.2.0-cp310-abi3-win_amd64.whl", hash = "sha256:4ac026d83d72216444963dfb603fc3e1806aca0a5cb3c236d14e82f20fb7b5de", size = 32556, upload-time = "2025-06-16T13:36:54.615Z" }, + { url = "https://files.pythonhosted.org/packages/be/de/e8dbfe36c70954dac62da4ca9a2fcff067eb097b48bd4b55fc836b2efd81/tree_sitter_hcl-1.2.0-cp310-abi3-win_arm64.whl", hash = "sha256:689425894a69423301e0e05faff472317e7ab4013767bce4a33b9234778c275e", size = 30948, upload-time = "2025-06-16T13:36:55.251Z" }, +] + [[package]] name = "tree-sitter-java" version = "0.23.5"