Merge official v8 (v0.8.32) with CodeBuddy support

This commit is contained in:
DevinZeng
2026-06-06 09:58:11 +08:00
10 changed files with 626 additions and 107 deletions
+8
View File
@@ -2,6 +2,14 @@
Full release notes with details on each version: [GitHub Releases](https://github.com/safishamsi/graphify/releases)
## 0.8.32 (2026-06-05)
- Feat: Terraform/HCL support. `.tf`, `.tfvars`, and `.hcl` files are now AST-extracted via `tree-sitter-hcl` into a structured infrastructure dependency graph. Nodes: resources, data sources, modules, variables, outputs, providers, and locals. Edges: `contains`, `references` (interpolation), and `depends_on`. Node IDs are directory-scoped for cross-file resolution. Requires `uv tool install "graphifyy[terraform]"` (#1129).
- Fix: `graphify extract` no longer requires an LLM API key for code-only corpora. Backend resolution is now deferred until after file detection — a corpus with only code files (pure tree-sitter AST, zero LLM calls) runs fully offline. The key is only enforced when docs, PDFs, or images are present, or when `--dedup-llm` is passed (#1122).
- Fix: `graphify kiro install` now correctly installs the `references/` sidecar and `.graphify_version` stamp. The install was using a bare `write_text` that bypassed the shared helper, shipping `SKILL.md` with 8 dead `references/*.md` pointers. Re-run `graphify kiro install` to pick up the fix (#1142).
- Fix: `GRAPHIFY_API_TIMEOUT` now applies to `claude-cli` subprocess and Anthropic SDK backend, not just the HTTP client. Both subprocess paths previously hardcoded `timeout=600` and ignored the env var and `--api-timeout` flag (#1112).
- Build: version floors added for `networkx>=3.4`, `datasketch>=1.6`, and `rapidfuzz>=3.0` to prevent silent breakage from old installs resolving incompatible versions.
## 0.8.31 (2026-06-03)
- Fix: `graphify hook install` now embeds the current interpreter (`sys.executable`) directly into the generated hook scripts. Previously, uv tool and pipx installs silently no-oped on git commit in GUI clients and CI runners where `~/.local/bin` is not on PATH — the hook could not find the graphify launcher, fell through all detection probes, and exited 0 without rebuilding. The embedded path is sanitized through a filesystem-safe allowlist before substitution. If you already have hooks installed, re-run `graphify hook install` to pick up the fix (#1127).
+8 -4
View File
@@ -173,6 +173,7 @@ Install only what you need:
| `bedrock` | AWS Bedrock (uses IAM, no API key) | `uv tool install "graphifyy[bedrock]"` |
| `sql` | SQL schema extraction | `uv tool install "graphifyy[sql]"` |
| `dm` | BYOND DreamMaker `.dm`/`.dme` AST extraction (may need a C compiler + `python3-dev` if no wheel matches your platform) | `uv tool install "graphifyy[dm]"` |
| `terraform` | Terraform / HCL `.tf`/`.tfvars`/`.hcl` AST extraction | `uv tool install "graphifyy[terraform]"` |
| `chinese` | Chinese query segmentation (jieba) | `uv tool install "graphifyy[chinese]"` |
| `all` | Everything above | `uv tool install "graphifyy[all]"` |
@@ -234,7 +235,8 @@ To remove graphify from all platforms at once: `graphify uninstall` (add `--purg
| Type | Extensions |
|------|-----------|
| Code (33 languages) | `.py .ts .js .jsx .tsx .mjs .go .rs .java .c .cpp .h .hpp .rb .cs .kt .scala .php .swift .lua .luau .zig .ps1 .ex .exs .m .mm .jl .vue .svelte .astro .groovy .gradle .dart .v .sv .svh .sql .f .f90 .f95 .f03 .f08 .pas .pp .dpr .dpk .lpr .inc .dfm .lfm .lpk .sh .bash .json .dm .dme .dmi .dmm .dmf .sln .csproj .fsproj .vbproj .razor .cshtml` (`.dm`/`.dme` AST extraction requires `uv tool install graphifyy[dm]`) |
| Code (28 tree-sitter grammars) | `.py .ts .js .jsx .tsx .mjs .go .rs .java .c .cpp .h .hpp .rb .cs .kt .scala .php .swift .lua .luau .zig .ps1 .ex .exs .m .mm .jl .vue .svelte .astro .groovy .gradle .dart .v .sv .svh .sql .f .f90 .f95 .f03 .f08 .pas .pp .dpr .dpk .lpr .inc .dfm .lfm .lpk .sh .bash .json .dm .dme .dmi .dmm .dmf .sln .csproj .fsproj .vbproj .razor .cshtml` (`.dm`/`.dme` requires `uv tool install graphifyy[dm]`) |
| Terraform / HCL | `.tf .tfvars .hcl` (requires `uv tool install graphifyy[terraform]`) |
| MCP configs | `.mcp.json` `mcp.json` `mcp_servers.json` `claude_desktop_config.json` — extracts server nodes, package refs, env var requirements |
| Docs | `.md .mdx .qmd .html .txt .rst .yaml .yml` |
| Office | `.docx .xlsx` (requires `uv tool install graphifyy[office]`) |
@@ -296,7 +298,9 @@ See the [full command reference](#full-command-reference) below.
## Ignoring files
Create a `.graphifyignore` in your project root — same syntax as `.gitignore`, including `!` negation:
Create a `.graphifyignore` in your project root — same syntax as `.gitignore`, including `!` negation.
**`.gitignore` is respected automatically.** If no `.graphifyignore` is present in a directory, graphify falls back to the `.gitignore` in that directory. If both exist, `.graphifyignore` takes priority. Subdirectory scoping works the same way as git — an ignore file only affects its own subtree.
```
# .graphifyignore
@@ -386,7 +390,7 @@ These are only needed for **headless / CI extraction** (`graphify extract`). Whe
## Privacy
- **Code files** — processed locally via tree-sitter. Nothing leaves your machine.
- **Code files** — processed locally via tree-sitter. Nothing leaves your machine. A code-only corpus requires no API key — `graphify extract` runs fully offline.
- **Video / audio** — transcribed locally with faster-whisper. Nothing leaves your machine.
- **Docs, PDFs, images** — sent to your AI assistant for semantic extraction (via the `/graphify` skill, using whatever model your IDE session runs). Headless `graphify extract` requires `GEMINI_API_KEY` / `GOOGLE_API_KEY` (Gemini), `MOONSHOT_API_KEY` (Kimi), `ANTHROPIC_API_KEY` (Claude), `OPENAI_API_KEY` (OpenAI), `DEEPSEEK_API_KEY` (DeepSeek), a running Ollama instance (`OLLAMA_BASE_URL`), AWS credentials via the standard provider chain (Bedrock - no API key needed, uses IAM), or the `claude` CLI binary (Claude Code - no API key needed, uses your Claude subscription). The `--dedup-llm` flag uses the same key.
- **Data residency** — `graphify extract` auto-detects which provider to use based on which API key is set (priority: Gemini → Kimi → Claude → OpenAI → DeepSeek → Bedrock → Ollama). For code with data-residency requirements, use `--backend ollama` (fully local) or pass an explicit `--backend` flag. Kimi (`MOONSHOT_API_KEY`) routes to Moonshot AI servers in China.
@@ -435,7 +439,7 @@ graphify query "..."
Run `graphify hook install` — it sets up a git merge driver that union-merges `graph.json` automatically so conflicts never happen.
**Extraction returns empty nodes/edges for docs or PDFs**
Docs and PDFs require an LLM call. Check that your API key is set and the backend is correct:
Docs, PDFs, and images require an LLM call — code-only corpora need no key. Check that your API key is set and the backend is correct:
```bash
ANTHROPIC_API_KEY=sk-... graphify extract ./docs --backend claude
```
+126 -99
View File
@@ -583,7 +583,38 @@ def _replace_or_append_section(content: str, marker: str, new_section: str) -> s
return out
def _print_banner() -> None:
"""Amber brain banner on graphify install. TTY-only, never raises."""
if not sys.stdout.isatty():
return
try:
if sys.platform == "win32":
import ctypes
ctypes.windll.kernel32.SetConsoleMode(
ctypes.windll.kernel32.GetStdHandle(-11), 7
)
A = "\033[38;5;214m"
D = "\033[38;5;130m"
R = "\033[0m"
print(f"""{A}
╭──◉──╮ ╭──◉──╮
╱ ◉ ◉ ╲ ╱ ◉ ◉ ╲
│ ◉─◉─◉ ◉ ◉─◉─◉ │
│ ◉ ◉ │ ◉ ◉ │
│ ◉─◉─◉ ◉ ◉─◉─◉ │
╲ ◉ ◉ ╱ ╲ ◉ ◉ ╱
╰──◉──╯ ╰──◉──╯
◉
█▀▀ █▀█ ▄▀█ █▀█ █ █ █ █▀▀ █▄█
█▄█ █▀▄ █▀█ █▀▀ █▀█ █ █▀ █{D} {__version__}{R}
""")
except Exception:
pass
def install(platform: str = "claude", *, project: bool = False, project_dir: Path | None = None) -> None:
_print_banner()
if platform == "gemini":
gemini_install(project_dir=project_dir, project=project)
return
@@ -935,12 +966,11 @@ def _kiro_install(project_dir: Path) -> None:
"""Write graphify skill + steering file for Kiro IDE/CLI."""
project_dir = project_dir or Path(".")
# Skill file → .kiro/skills/graphify/SKILL.md
skill_src = Path(__file__).parent / "skill-kiro.md"
skill_dst = project_dir / ".kiro" / "skills" / "graphify" / "SKILL.md"
skill_dst.parent.mkdir(parents=True, exist_ok=True)
skill_dst.write_text(skill_src.read_text(encoding="utf-8"), encoding="utf-8")
print(f" {skill_dst.relative_to(project_dir)} -> /graphify skill")
# Skill file + references/ sidecar + .graphify_version stamp via the shared
# progressive-disclosure helper. Previously this used a bare write_text that
# bypassed _copy_skill_file, so the references/ dir and version stamp were
# never written even though kiro declares skill_refs: "kiro" (#1142).
_copy_skill_file("kiro", project=True, project_dir=project_dir)
# Steering file → .kiro/steering/graphify.md (always-on)
steering_dir = project_dir / ".kiro" / "steering"
@@ -965,15 +995,10 @@ def _kiro_uninstall(project_dir: Path) -> None:
project_dir = project_dir or Path(".")
removed = []
skill_dst = project_dir / ".kiro" / "skills" / "graphify" / "SKILL.md"
if skill_dst.exists():
skill_dst.unlink()
# Skill + .graphify_version + references/ sidecar + empty-dir walk.
skill_dst = _platform_skill_destination("kiro", project=True, project_dir=project_dir)
if _remove_skill_file("kiro", project=True, project_dir=project_dir):
removed.append(str(skill_dst.relative_to(project_dir)))
# Remove parent dir if empty
try:
skill_dst.parent.rmdir()
except OSError:
pass
steering_dst = project_dir / ".kiro" / "steering" / "graphify.md"
if steering_dst.exists():
@@ -3949,91 +3974,6 @@ def main() -> None:
if cli_max_workers is not None:
os.environ["GRAPHIFY_MAX_WORKERS"] = str(cli_max_workers)
# Backend resolution. If user did not pass --backend, sniff env.
# If backend was explicitly requested, validate its key is present
# and surface a clear error early — don't let extract_corpus_parallel
# raise mid-run after we've spent time on AST extraction.
from graphify.llm import (
BACKENDS as _BACKENDS,
detect_backend as _detect_backend,
estimate_cost as _estimate_cost,
extract_corpus_parallel as _extract_corpus_parallel,
_format_backend_env_keys,
_get_backend_api_key,
)
if backend is None:
backend = _detect_backend()
if backend is None:
print(
"error: no LLM API key found. Set GEMINI_API_KEY or GOOGLE_API_KEY "
"(gemini), MOONSHOT_API_KEY (kimi), ANTHROPIC_API_KEY (claude), "
"OPENAI_API_KEY (openai), DEEPSEEK_API_KEY (deepseek), "
"or pass --backend.",
file=sys.stderr,
)
sys.exit(1)
if backend not in _BACKENDS:
print(
f"error: unknown backend '{backend}'. "
f"Available: {', '.join(sorted(_BACKENDS))}",
file=sys.stderr,
)
sys.exit(1)
if backend == "ollama":
# Fail closed with a clean message (not a deep traceback) if
# OLLAMA_BASE_URL points at a link-local/metadata address. warn=False:
# the later in-flow call owns the user-facing warning for LAN hosts.
from graphify.llm import _validate_ollama_base_url
_oll_url = os.environ.get("OLLAMA_BASE_URL", _BACKENDS["ollama"].get("base_url", ""))
try:
_validate_ollama_base_url(_oll_url, warn=False)
except ValueError as exc:
print(f"error: {exc}", file=sys.stderr)
sys.exit(2)
if not _get_backend_api_key(backend):
# Ollama on a loopback URL ignores auth entirely; don't block
# the run just because OLLAMA_API_KEY is unset (issue #792).
# extract_files_direct already prints a warning and substitutes
# a placeholder key in that case.
allow_no_key = False
if backend == "ollama":
from urllib.parse import urlparse
ollama_url = os.environ.get(
"OLLAMA_BASE_URL",
_BACKENDS["ollama"].get("base_url", ""),
)
try:
host = (urlparse(ollama_url).hostname or "").lower()
except Exception:
host = ""
allow_no_key = (
host in ("localhost", "127.0.0.1", "::1")
or host.startswith("127.")
)
elif backend == "bedrock":
allow_no_key = bool(
os.environ.get("AWS_PROFILE")
or os.environ.get("AWS_REGION")
or os.environ.get("AWS_DEFAULT_REGION")
or os.environ.get("AWS_ACCESS_KEY_ID")
)
elif backend == "claude-cli":
import shutil as _shutil
allow_no_key = _shutil.which("claude") is not None
if not allow_no_key:
print(
"error: backend 'claude-cli' requires the `claude` CLI on $PATH "
"(install Claude Code and run `claude` once to authenticate).",
file=sys.stderr,
)
sys.exit(1)
if not allow_no_key:
print(
f"error: backend '{backend}' requires {_format_backend_env_keys(backend)} to be set.",
file=sys.stderr,
)
sys.exit(1)
# Resolve output dir. The user-facing contract is "<out>/graphify-out/"
# so a fresh checkout writes graphify-out/ at the project root, matching
# the skill.md pipeline.
@@ -4093,6 +4033,93 @@ def main() -> None:
f"{len(image_files)} images"
)
# Resolve the LLM backend only now that we know whether the corpus
# needs one. A code-only corpus is pure local AST and must not require
# an API key; the key is enforced below only when there's LLM work.
from graphify.llm import (
BACKENDS as _BACKENDS,
detect_backend as _detect_backend,
estimate_cost as _estimate_cost,
extract_corpus_parallel as _extract_corpus_parallel,
_format_backend_env_keys,
_get_backend_api_key,
)
needs_llm = bool(semantic_files) or dedup_llm
if backend is None and needs_llm:
backend = _detect_backend()
if backend is not None and backend not in _BACKENDS:
print(
f"error: unknown backend '{backend}'. "
f"Available: {', '.join(sorted(_BACKENDS))}",
file=sys.stderr,
)
sys.exit(1)
if needs_llm:
if backend is None:
reasons = []
if semantic_files:
reasons.append(
f"{len(semantic_files)} doc/paper/image file(s) need semantic extraction"
)
if dedup_llm:
reasons.append("--dedup-llm was passed")
print(
"error: no LLM API key found (" + "; ".join(reasons) + "). "
"Set GEMINI_API_KEY or GOOGLE_API_KEY (gemini), MOONSHOT_API_KEY "
"(kimi), ANTHROPIC_API_KEY (claude), OPENAI_API_KEY (openai), "
"DEEPSEEK_API_KEY (deepseek), or pass --backend. A code-only "
"corpus needs no key.",
file=sys.stderr,
)
sys.exit(1)
if backend == "ollama":
from graphify.llm import _validate_ollama_base_url
_oll_url = os.environ.get("OLLAMA_BASE_URL", _BACKENDS["ollama"].get("base_url", ""))
try:
_validate_ollama_base_url(_oll_url, warn=False)
except ValueError as exc:
print(f"error: {exc}", file=sys.stderr)
sys.exit(2)
if not _get_backend_api_key(backend):
allow_no_key = False
if backend == "ollama":
from urllib.parse import urlparse
ollama_url = os.environ.get(
"OLLAMA_BASE_URL",
_BACKENDS["ollama"].get("base_url", ""),
)
try:
host = (urlparse(ollama_url).hostname or "").lower()
except Exception:
host = ""
allow_no_key = (
host in ("localhost", "127.0.0.1", "::1")
or host.startswith("127.")
)
elif backend == "bedrock":
allow_no_key = bool(
os.environ.get("AWS_PROFILE")
or os.environ.get("AWS_REGION")
or os.environ.get("AWS_DEFAULT_REGION")
or os.environ.get("AWS_ACCESS_KEY_ID")
)
elif backend == "claude-cli":
import shutil as _shutil
allow_no_key = _shutil.which("claude") is not None
if not allow_no_key:
print(
"error: backend 'claude-cli' requires the `claude` CLI on $PATH "
"(install Claude Code and run `claude` once to authenticate).",
file=sys.stderr,
)
sys.exit(1)
if not allow_no_key:
print(
f"error: backend '{backend}' requires {_format_backend_env_keys(backend)} to be set.",
file=sys.stderr,
)
sys.exit(1)
# AST extraction on code files. Empty code list (docs-only corpus) is
# the issue #698 case — skip cleanly instead of crashing inside extract().
ast_result: dict = {"nodes": [], "edges": [], "input_tokens": 0, "output_tokens": 0}
+1 -1
View File
@@ -25,7 +25,7 @@ class FileType(str, Enum):
_MANIFEST_PATH = "graphify-out/manifest.json"
CODE_EXTENSIONS = {'.py', '.ts', '.tsx', '.js', '.jsx', '.mjs', '.ejs', '.ets', '.go', '.rs', '.java', '.groovy', '.gradle', '.cpp', '.cc', '.cxx', '.c', '.h', '.hpp', '.rb', '.swift', '.kt', '.kts', '.cs', '.scala', '.php', '.lua', '.luau', '.toc', '.zig', '.ps1', '.ex', '.exs', '.m', '.mm', '.jl', '.vue', '.svelte', '.astro', '.dart', '.v', '.sv', '.svh', '.sql', '.r', '.f', '.F', '.f90', '.F90', '.f95', '.F95', '.f03', '.F03', '.f08', '.F08', '.pas', '.pp', '.dpr', '.dpk', '.lpr', '.inc', '.dfm', '.lfm', '.lpk', '.sh', '.bash', '.json', '.dm', '.dme', '.dmi', '.dmm', '.dmf', '.sln', '.csproj', '.fsproj', '.vbproj', '.razor', '.cshtml'}
CODE_EXTENSIONS = {'.py', '.ts', '.tsx', '.js', '.jsx', '.mjs', '.ejs', '.ets', '.go', '.rs', '.java', '.groovy', '.gradle', '.cpp', '.cc', '.cxx', '.c', '.h', '.hpp', '.rb', '.swift', '.kt', '.kts', '.cs', '.scala', '.php', '.lua', '.luau', '.toc', '.zig', '.ps1', '.ex', '.exs', '.m', '.mm', '.jl', '.vue', '.svelte', '.astro', '.dart', '.v', '.sv', '.svh', '.sql', '.r', '.f', '.F', '.f90', '.F90', '.f95', '.F95', '.f03', '.F03', '.f08', '.F08', '.pas', '.pp', '.dpr', '.dpk', '.lpr', '.inc', '.dfm', '.lfm', '.lpk', '.sh', '.bash', '.json', '.tf', '.tfvars', '.hcl', '.dm', '.dme', '.dmi', '.dmm', '.dmf', '.sln', '.csproj', '.fsproj', '.vbproj', '.razor', '.cshtml'}
DOC_EXTENSIONS = {'.md', '.mdx', '.qmd', '.txt', '.rst', '.html', '.yaml', '.yml'}
PAPER_EXTENSIONS = {'.pdf'}
IMAGE_EXTENSIONS = {'.png', '.jpg', '.jpeg', '.gif', '.webp', '.svg'}
+181
View File
@@ -10526,6 +10526,184 @@ def extract_dmf(path: Path) -> dict:
return {"nodes": nodes, "edges": edges}
# Head tokens in an HCL traversal that are meta/builtins, not references to a
# block defined in the corpus (count.index, each.key, self.*, path.module, ...).
_TF_META_HEADS = frozenset({"count", "each", "self", "path", "terraform"})
def extract_terraform(path: Path) -> dict:
"""Extract Terraform/HCL blocks and the references between them via tree-sitter.
Nodes: resources, data sources, modules, variables, outputs, providers, and
locals. Edges: `contains` (file -> block), `references` (block -> the blocks
it interpolates, e.g. `aws_instance.web` -> `var.region`), and `depends_on`
(explicit dependency edges).
Node IDs are scoped by the parent directory, not the file stem, because
Terraform resources are module(directory)-scoped: a resource defined in
main.tf is referenced from other .tf files in the same directory. Directory
scoping lets those cross-file references resolve when per-file extractions
are merged (stem scoping would split a definition from its references).
"""
try:
import tree_sitter_hcl as tshcl
from tree_sitter import Language, Parser
except ImportError:
return {"nodes": [], "edges": [], "error": "tree_sitter_hcl not installed. Run: pip install tree-sitter-hcl"}
try:
language = Language(tshcl.language())
parser = Parser(language)
source = path.read_bytes()
tree = parser.parse(source)
root = tree.root_node
except Exception as e:
return {"nodes": [], "edges": [], "error": str(e)}
str_path = str(path)
file_nid = _make_id(str_path)
scope = path.parent.name or "tf"
nodes: list[dict] = [{"id": file_nid, "label": path.name, "file_type": "code",
"source_file": str_path, "source_location": None}]
edges: list[dict] = []
seen_ids: set[str] = {file_nid}
seen_edges: set[tuple[str, str, str]] = set()
def _read(n) -> str:
return source[n.start_byte:n.end_byte].decode("utf-8", errors="replace")
def _label_text(n) -> str:
return _read(n).strip().strip('"')
def _add_node(address: str, label: str, line: int) -> str:
nid = _make_id(scope, address)
if nid not in seen_ids:
seen_ids.add(nid)
nodes.append({"id": nid, "label": label, "file_type": "code",
"source_file": str_path, "source_location": f"L{line}"})
edges.append({"source": file_nid, "target": nid, "relation": "contains",
"confidence": "EXTRACTED", "source_file": str_path,
"source_location": f"L{line}", "weight": 1.0})
return nid
def _add_edge(src: str, address: str, relation: str, line: int) -> None:
tgt = _make_id(scope, address)
if src == tgt:
return
key = (src, tgt, relation)
if key in seen_edges:
return
seen_edges.add(key)
edges.append({"source": src, "target": tgt, "relation": relation,
"confidence": "EXTRACTED", "source_file": str_path,
"source_location": f"L{line}", "weight": 1.0})
def _block_parts(block) -> tuple:
btype = None
labels: list[str] = []
for c in block.children:
if c.type in ("block_start", "body", "block_end"):
break
if c.type == "identifier" and btype is None:
btype = _read(c)
elif c.type in ("string_lit", "identifier"):
labels.append(_label_text(c))
return btype, labels
def _ref_address(expr):
head = _read(expr)
parent = expr.parent
attrs: list[str] = []
if parent is not None:
seen_self = False
for c in parent.children:
if c.id == expr.id:
seen_self = True
continue
if seen_self and c.type == "get_attr":
name = None
for gc in c.children:
if gc.type == "identifier":
name = _read(gc)
break
if name is None:
break
attrs.append(name)
elif seen_self and c.type not in ("get_attr",):
break
if head in _TF_META_HEADS or not head:
return None
if head == "var":
return f"var.{attrs[0]}" if attrs else None
if head == "local":
return f"local.{attrs[0]}" if attrs else None
if head == "module":
return f"module.{attrs[0]}" if attrs else None
if head == "data":
return f"data.{attrs[0]}.{attrs[1]}" if len(attrs) >= 2 else None
return f"{head}.{attrs[0]}" if attrs else None
def _collect_refs(node, owner_nid: str, relation: str) -> None:
rel = relation
if node.type == "attribute":
key_node = node.child_by_field_name("key") or (
node.children[0] if node.children else None
)
if key_node is not None and _read(key_node) == "depends_on":
rel = "depends_on"
if node.type == "variable_expr":
addr = _ref_address(node)
if addr:
_add_edge(owner_nid, addr, rel, node.start_point[0] + 1)
for c in node.children:
if c.is_named:
_collect_refs(c, owner_nid, rel)
def _body_of(block):
for c in block.children:
if c.type == "body":
return c
return None
body = next((c for c in root.children if c.type == "body"), root)
for block in body.children:
if block.type != "block":
continue
btype, labels = _block_parts(block)
line = block.start_point[0] + 1
blk_body = _body_of(block)
if btype == "resource" and len(labels) >= 2:
owner = _add_node(f"{labels[0]}.{labels[1]}", f"{labels[0]}.{labels[1]}", line)
elif btype == "data" and len(labels) >= 2:
owner = _add_node(f"data.{labels[0]}.{labels[1]}", f"data.{labels[0]}.{labels[1]}", line)
elif btype == "module" and labels:
owner = _add_node(f"module.{labels[0]}", f"module.{labels[0]}", line)
elif btype == "variable" and labels:
owner = _add_node(f"var.{labels[0]}", f"var.{labels[0]}", line)
elif btype == "output" and labels:
owner = _add_node(f"output.{labels[0]}", f"output.{labels[0]}", line)
elif btype == "provider" and labels:
owner = _add_node(f"provider.{labels[0]}", f"provider.{labels[0]}", line)
elif btype == "locals" and blk_body is not None:
for attr in blk_body.children:
if attr.type != "attribute":
continue
key_node = attr.children[0] if attr.children else None
if key_node is None:
continue
key = _read(key_node)
lnid = _add_node(f"local.{key}", f"local.{key}", attr.start_point[0] + 1)
_collect_refs(attr, lnid, "references")
continue
else:
continue
if blk_body is not None:
_collect_refs(blk_body, owner, "references")
return {"nodes": nodes, "edges": edges}
_DISPATCH: dict[str, Any] = {
".py": extract_python,
".js": extract_js,
@@ -10594,6 +10772,9 @@ _DISPATCH: dict[str, Any] = {
".sh": extract_bash,
".bash": extract_bash,
".json": extract_json,
".tf": extract_terraform,
".tfvars": extract_terraform,
".hcl": extract_terraform,
".dm": extract_dm,
".dme": extract_dm,
".dmi": extract_dmi,
+4 -2
View File
@@ -4,7 +4,7 @@ build-backend = "setuptools.build_meta"
[project]
name = "graphifyy"
version = "0.8.31"
version = "0.8.32"
description = "AI coding assistant skill (Claude Code, CodeBuddy, Codex, OpenCode, Kilo Code, Cursor, Gemini CLI, Aider, OpenClaw, Factory Droid, Trae, Hermes, Kiro, Pi, Devin CLI, Google Antigravity) - turn any folder of code, docs, papers, images, or videos into a queryable knowledge graph"
readme = "README.md"
license = { file = "LICENSE" }
@@ -69,7 +69,8 @@ sql = ["tree-sitter-sql"]
# must compile from source (needs a C toolchain + python3-dev). Keeping it optional
# avoids breaking the default `uv tool install graphifyy` for everyone (#1104).
dm = ["tree-sitter-dm"]
all = ["mcp", "neo4j", "pypdf", "markdownify", "watchdog", "graspologic; python_version < '3.13'", "python-docx", "openpyxl", "faster-whisper; python_version >= '3.11'", "yt-dlp", "matplotlib", "openai", "tiktoken", "boto3", "anthropic", "tree-sitter-sql", "jieba", "tree-sitter-dm"]
terraform = ["tree-sitter-hcl"]
all = ["mcp", "neo4j", "pypdf", "markdownify", "watchdog", "graspologic; python_version < '3.13'", "python-docx", "openpyxl", "faster-whisper; python_version >= '3.11'", "yt-dlp", "matplotlib", "openai", "tiktoken", "boto3", "anthropic", "tree-sitter-sql", "jieba", "tree-sitter-dm", "tree-sitter-hcl"]
[project.scripts]
graphify = "graphify.__main__:main"
@@ -91,6 +92,7 @@ dev = [
"setuptools>=82.0.1",
"wheel>=0.47.0",
"tomli>=2.0 ; python_version < '3.11'",
"tree-sitter-hcl>=1.2.0",
]
[tool.uv]
+77
View File
@@ -121,3 +121,80 @@ def test_extract_succeeds_when_at_least_one_chunk_completes(
assert (out_dir / "graphify-out" / "graph.json").exists(), (
"graph.json must be written on the happy path"
)
def _code_only_corpus(tmp_path):
"""A corpus with only code — no docs/papers/images."""
(tmp_path / "auth.py").write_text(
"def login(user):\n return validate(user)\n\n"
"def validate(user):\n return True\n"
)
return tmp_path
def _clear_backend_keys(monkeypatch):
"""Clear every env var that detect_backend() or _get_backend_api_key() reads."""
for key in (
"GEMINI_API_KEY", "GOOGLE_API_KEY", "OPENAI_API_KEY",
"ANTHROPIC_API_KEY", "DEEPSEEK_API_KEY", "MOONSHOT_API_KEY",
# bedrock: presence of any of these is treated as a valid credential
"AWS_PROFILE", "AWS_REGION", "AWS_DEFAULT_REGION", "AWS_ACCESS_KEY_ID",
# ollama: a set OLLAMA_BASE_URL triggers backend detection
"OLLAMA_BASE_URL",
):
monkeypatch.delenv(key, raising=False)
def test_extract_codeonly_succeeds_without_api_key(monkeypatch, tmp_path):
"""A code-only corpus must run with no LLM API key.
Regression: graphify extract validated a backend upfront and exited 1 with
'no LLM API key found' even for a code-only corpus that never calls a model.
The keyless AST path now runs to a written graph.json (#1122).
"""
corpus = _code_only_corpus(tmp_path)
out_dir = tmp_path / "out"
_clear_backend_keys(monkeypatch)
monkeypatch.setattr(mainmod, "_check_skill_version", lambda _: None)
monkeypatch.setattr(
mainmod.sys, "argv",
["graphify", "extract", str(corpus), "--out", str(out_dir)],
)
try:
mainmod.main()
except SystemExit as exc:
assert exc.code in (None, 0), f"unexpected exit code {exc.code}"
graph = out_dir / "graphify-out" / "graph.json"
assert graph.exists(), "code-only extract must write graph.json without a key"
import json
assert len(json.loads(graph.read_text()).get("nodes", [])) > 0
def test_extract_without_key_still_errors_when_docs_present(
monkeypatch, tmp_path, capsys
):
"""Key requirement still fires when semantic work is needed.
A corpus with a Markdown doc needs LLM semantic extraction, so a keyless
extract must exit 1 with clear guidance (#1122).
"""
corpus = _make_corpus(tmp_path) # includes a Markdown doc
out_dir = tmp_path / "out"
_clear_backend_keys(monkeypatch)
# Patch detect_backend too so ambient AWS/ollama env can't slip through.
monkeypatch.setattr("graphify.llm.detect_backend", lambda: None)
monkeypatch.setattr(mainmod, "_check_skill_version", lambda _: None)
monkeypatch.setattr(
mainmod.sys, "argv",
["graphify", "extract", str(corpus), "--out", str(out_dir)],
)
with pytest.raises(SystemExit) as exc_info:
mainmod.main()
assert exc_info.value.code == 1
err = capsys.readouterr().err
assert "no LLM API key found" in err
assert "code-only corpus needs no key" in err
assert not (out_dir / "graphify-out" / "graph.json").exists()
+39
View File
@@ -231,3 +231,42 @@ def test_kiro_install_upgrades_stale_steering(tmp_path, monkeypatch):
assert "read it before answering architecture questions" not in after
_assert_query_first(after, ".kiro/steering/graphify.md")
assert "inclusion: always" in after # frontmatter preserved
def test_kiro_install_ships_references_sidecar_and_version_stamp(tmp_path, monkeypatch):
"""_kiro_install routes through _copy_skill_file so the references/ sidecar
and .graphify_version stamp are written alongside SKILL.md (#1142).
Previously it used a bare write_text that bypassed the shared helper."""
monkeypatch.setattr(mainmod, "_check_skill_version", lambda _: None)
refs_dir = Path(mainmod.__file__).parent / "skills" / "kiro" / "references"
if not refs_dir.exists():
pytest.skip("kiro references bundle not present in this checkout")
mainmod._kiro_install(tmp_path)
skill_dir = tmp_path / ".kiro" / "skills" / "graphify"
# SKILL.md present
assert (skill_dir / "SKILL.md").exists()
# references/ sidecar installed with at least one fragment
refs_dst = skill_dir / "references"
assert refs_dst.is_dir(), "references/ sidecar must be installed (#1142)"
assert any(refs_dst.iterdir()), "references/ must not be empty"
# .graphify_version stamp written
version_file = skill_dir / ".graphify_version"
assert version_file.exists(), ".graphify_version stamp must be written (#1142)"
assert version_file.read_text(encoding="utf-8") == mainmod.__version__
# no references.tmp leftover
assert not (skill_dir / "references.tmp").exists()
# steering file still written
assert (tmp_path / ".kiro" / "steering" / "graphify.md").exists()
# uninstall removes skill dir, version stamp, references/, and steering file
mainmod._kiro_uninstall(tmp_path)
assert not skill_dir.exists(), "uninstall must remove skill dir including references/ (#1142)"
assert not (tmp_path / ".kiro" / "steering" / "graphify.md").exists()
+158
View File
@@ -0,0 +1,158 @@
"""Tests for the Terraform/HCL extractor (graphify/extract.py, issue #187)."""
from __future__ import annotations
from pathlib import Path
from graphify.build import build_from_json
from graphify.extract import extract_terraform
def _write(tmp_path: Path, name: str, body: str) -> Path:
p = tmp_path / name
p.write_text(body, encoding="utf-8")
return p
def _labels(r) -> list[str]:
return [n["label"] for n in r["nodes"]]
def _rel_pairs(r, relation: str) -> set[tuple[str, str]]:
lab = {n["id"]: n["label"] for n in r["nodes"]}
return {
(lab.get(e["source"], e["source"]), lab.get(e["target"], e["target"]))
for e in r["edges"]
if e["relation"] == relation
}
SAMPLE = """\
# leading comment so the body is not children[0]
terraform {
required_providers { azurerm = { source = "hashicorp/azurerm" } }
}
variable "region" { default = "us-east-1" }
provider "aws" { region = var.region }
data "aws_ami" "ubuntu" { most_recent = true }
resource "aws_instance" "web" {
ami = data.aws_ami.ubuntu.id
subnet_id = var.region
depends_on = [aws_security_group.sg]
}
resource "aws_security_group" "sg" { name = "sg" }
module "vpc" {
source = "./modules/vpc"
cidr = local.cidr
}
locals { cidr = "10.0.0.0/16" }
output "ip" { value = aws_instance.web.private_ip }
"""
def test_no_error_and_all_block_types_become_nodes(tmp_path):
r = extract_terraform(_write(tmp_path, "main.tf", SAMPLE))
assert r.get("error") is None
labels = set(_labels(r))
# one node per block type (the terraform{} settings block is intentionally skipped)
for expected in (
"var.region",
"provider.aws",
"data.aws_ami.ubuntu",
"aws_instance.web",
"aws_security_group.sg",
"module.vpc",
"local.cidr",
"output.ip",
):
assert expected in labels, f"missing node {expected!r}"
def test_reference_edges(tmp_path):
r = extract_terraform(_write(tmp_path, "main.tf", SAMPLE))
refs = _rel_pairs(r, "references")
assert ("provider.aws", "var.region") in refs
assert ("aws_instance.web", "data.aws_ami.ubuntu") in refs
assert ("aws_instance.web", "var.region") in refs
assert ("module.vpc", "local.cidr") in refs
assert ("output.ip", "aws_instance.web") in refs
def test_depends_on_edge(tmp_path):
r = extract_terraform(_write(tmp_path, "main.tf", SAMPLE))
assert ("aws_instance.web", "aws_security_group.sg") in _rel_pairs(r, "depends_on")
def test_file_contains_blocks(tmp_path):
r = extract_terraform(_write(tmp_path, "main.tf", SAMPLE))
contains = _rel_pairs(r, "contains")
assert ("main.tf", "aws_instance.web") in contains
assert ("main.tf", "var.region") in contains
def test_meta_heads_not_emitted(tmp_path):
# count.index / each.key / self.* / path.module are builtins, not references.
body = """\
resource "aws_instance" "web" {
count = 2
name = "web-${count.index}"
tags = each.value
dir = path.module
}
"""
r = extract_terraform(_write(tmp_path, "main.tf", body))
targets = {t for _, t in _rel_pairs(r, "references")}
assert not any(t.startswith(("count", "each", "path")) for t in targets)
def test_cross_file_references_resolve_after_merge(tmp_path):
# A resource defined in one file is referenced from another in the same
# directory; directory-scoped IDs must let the edge resolve at build time.
defn = """\
resource "azurerm_resource_group" "main" { name = "rg" }
"""
user = """\
resource "azurerm_network_interface" "nic" {
resource_group_name = azurerm_resource_group.main.name
}
"""
r_defn = extract_terraform(_write(tmp_path, "main.tf", defn))
r_user = extract_terraform(_write(tmp_path, "nic.tf", user))
# The cross-file edge target id equals the definition's node id.
rg_id = next(n["id"] for n in r_defn["nodes"] if n["label"] == "azurerm_resource_group.main")
nic_ref_targets = {e["target"] for e in r_user["edges"] if e["relation"] == "references"}
assert rg_id in nic_ref_targets
# And it survives a real merge: the edge is present (not dropped as dangling).
G = build_from_json(
{
"nodes": r_defn["nodes"] + r_user["nodes"],
"edges": r_defn["edges"] + r_user["edges"],
}
)
nic_id = next(n["id"] for n in r_user["nodes"] if n["label"] == "azurerm_network_interface.nic")
assert G.has_edge(nic_id, rg_id)
def test_empty_and_commentonly_files_are_safe(tmp_path):
assert extract_terraform(_write(tmp_path, "a.tf", "")).get("error") is None
r = extract_terraform(_write(tmp_path, "b.tf", "# just a comment\n"))
# only the file node, no crash
assert len(r["nodes"]) == 1
def test_tfvars_key_value_is_safe(tmp_path):
# .tfvars files contain only key=value assignments (no block structure),
# so extract_terraform produces zero block nodes — only the file node.
# This is the documented intended behaviour for .tfvars.
r = extract_terraform(_write(tmp_path, "terraform.tfvars", 'region = "us-east-1"\nenv = "prod"\n'))
assert r.get("error") is None
assert len(r["nodes"]) == 1 # only the file node, no variable nodes
Generated
+24 -1
View File
@@ -1189,6 +1189,7 @@ all = [
{ name = "python-docx" },
{ name = "tiktoken" },
{ name = "tree-sitter-dm" },
{ name = "tree-sitter-hcl" },
{ name = "tree-sitter-sql" },
{ name = "watchdog" },
{ name = "yt-dlp" },
@@ -1246,6 +1247,9 @@ sql = [
svg = [
{ name = "matplotlib" },
]
terraform = [
{ name = "tree-sitter-hcl" },
]
video = [
{ name = "faster-whisper", marker = "python_full_version >= '3.11'" },
{ name = "yt-dlp" },
@@ -1270,6 +1274,7 @@ dev = [
{ name = "safety" },
{ name = "setuptools" },
{ name = "tomli", marker = "python_full_version < '3.11'" },
{ name = "tree-sitter-hcl" },
{ name = "wheel" },
]
@@ -1323,6 +1328,8 @@ requires-dist = [
{ name = "tree-sitter-fortran" },
{ name = "tree-sitter-go" },
{ name = "tree-sitter-groovy" },
{ name = "tree-sitter-hcl", marker = "extra == 'all'" },
{ name = "tree-sitter-hcl", marker = "extra == 'terraform'" },
{ name = "tree-sitter-java" },
{ name = "tree-sitter-javascript" },
{ name = "tree-sitter-json" },
@@ -1347,7 +1354,7 @@ requires-dist = [
{ name = "yt-dlp", marker = "extra == 'all'" },
{ name = "yt-dlp", marker = "extra == 'video'" },
]
provides-extras = ["mcp", "neo4j", "pdf", "watch", "svg", "leiden", "office", "google", "video", "kimi", "ollama", "bedrock", "anthropic", "gemini", "openai", "chinese", "sql", "dm", "all"]
provides-extras = ["mcp", "neo4j", "pdf", "watch", "svg", "leiden", "office", "google", "video", "kimi", "ollama", "bedrock", "anthropic", "gemini", "openai", "chinese", "sql", "dm", "terraform", "all"]
[package.metadata.requires-dev]
dev = [
@@ -1365,6 +1372,7 @@ dev = [
{ name = "safety", specifier = ">=3.7.0" },
{ name = "setuptools", specifier = ">=82.0.1" },
{ name = "tomli", marker = "python_full_version < '3.11'", specifier = ">=2.0" },
{ name = "tree-sitter-hcl", specifier = ">=1.2.0" },
{ name = "wheel", specifier = ">=0.47.0" },
]
@@ -4638,6 +4646,21 @@ wheels = [
{ url = "https://files.pythonhosted.org/packages/9d/e3/50c719d09a4495672226b2359b2701360fdef022bc86dedef9fc16d3959c/tree_sitter_groovy-0.1.2-cp39-abi3-win_arm64.whl", hash = "sha256:1942a9a1b22e154da9bbf1b03e6b4dbec4211b1109d24bcf4c12b006cbc04037", size = 102508, upload-time = "2024-11-19T04:33:06.101Z" },
]
[[package]]
name = "tree-sitter-hcl"
version = "1.2.0"
source = { registry = "https://pypi.org/simple" }
sdist = { url = "https://files.pythonhosted.org/packages/06/c9/ed79f643b0cec3e123171c09caffb6088a6111025a20fc69112b1468828b/tree_sitter_hcl-1.2.0.tar.gz", hash = "sha256:f86cb7a9fd5cb93d83e2f788ae155544464c47755d09190505de562c0d6ad1dd", size = 55207, upload-time = "2025-06-16T13:36:56.159Z" }
wheels = [
{ url = "https://files.pythonhosted.org/packages/77/34/7ccb58107ae0d38e0a8f05dedbd990780eb7d429f227725c15fee314cdcd/tree_sitter_hcl-1.2.0-cp310-abi3-macosx_10_9_x86_64.whl", hash = "sha256:9ae35084a3dc12272f941b424eadd8a44cf2e0e9345b020330cf8db6f67d3524", size = 30604, upload-time = "2025-06-16T13:36:49.775Z" },
{ url = "https://files.pythonhosted.org/packages/8e/8b/7618448cde58ca6fbefcf210ef98d7c4bd7d2b54b3e3d5cddd947c804a18/tree_sitter_hcl-1.2.0-cp310-abi3-macosx_11_0_arm64.whl", hash = "sha256:03678762e8b78d717187848edebed95e4c31a54e14f24dec97555f47fb440e28", size = 31163, upload-time = "2025-06-16T13:36:50.936Z" },
{ url = "https://files.pythonhosted.org/packages/12/35/b8f87fffb5527c85f5f292486e53f64963846b94cf3ea258f4b850480f18/tree_sitter_hcl-1.2.0-cp310-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl", hash = "sha256:d2e9f3cfb6694e33f5b880e74bed842398cbacd21024251e2ec90f19dee6a64d", size = 54139, upload-time = "2025-06-16T13:36:51.951Z" },
{ url = "https://files.pythonhosted.org/packages/ee/0a/01bb627044d273e8e506edff8ab773e562ba447b5790b789f62e47a5e754/tree_sitter_hcl-1.2.0-cp310-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:915763da6630610c2efb7afe13145f50feb8043732a74f9bae78811212578d3d", size = 54317, upload-time = "2025-06-16T13:36:52.65Z" },
{ url = "https://files.pythonhosted.org/packages/f4/40/cc843f07210caa0a2e2a2b3581f93c917ed2139cb0851b32f4852141c95c/tree_sitter_hcl-1.2.0-cp310-abi3-musllinux_1_2_x86_64.whl", hash = "sha256:5b6c6aaccfca2fde4fcce52aa88d8c3756f44d407b8584fc3cff581d8765b4b7", size = 51638, upload-time = "2025-06-16T13:36:53.684Z" },
{ url = "https://files.pythonhosted.org/packages/27/a6/f15096c138eccc0c68b7254c75a8121bab326b62920f2d259079b2d6a7d0/tree_sitter_hcl-1.2.0-cp310-abi3-win_amd64.whl", hash = "sha256:4ac026d83d72216444963dfb603fc3e1806aca0a5cb3c236d14e82f20fb7b5de", size = 32556, upload-time = "2025-06-16T13:36:54.615Z" },
{ url = "https://files.pythonhosted.org/packages/be/de/e8dbfe36c70954dac62da4ca9a2fcff067eb097b48bd4b55fc836b2efd81/tree_sitter_hcl-1.2.0-cp310-abi3-win_arm64.whl", hash = "sha256:689425894a69423301e0e05faff472317e7ab4013767bce4a33b9234778c275e", size = 30948, upload-time = "2025-06-16T13:36:55.251Z" },
]
[[package]]
name = "tree-sitter-java"
version = "0.23.5"