perf+fix: parallel extraction, faster imports, bug fixes

45x faster cluster import, 135x faster detect, parallel subagent extraction,
auto-exclude venvs/caches, suggest_questions fix, manifest fix, install CLI
This commit is contained in:
Safi
2026-04-04 18:56:38 +01:00
parent e7a03a0539
commit 0b7460a9e3
16 changed files with 1670 additions and 149 deletions
+5
View File
@@ -1,4 +1,6 @@
venv/
.venv/
env/
__pycache__/
*.pyc
*.egg-info/
@@ -6,5 +8,8 @@ __pycache__/
dist/
build/
.pytest_cache/
.mypy_cache/
.ruff_cache/
*.so
*.egg
.graphify/
+129 -61
View File
@@ -1,6 +1,6 @@
# graphify
A Claude Code skill that turns any folder of files into a navigable knowledge graph — then opens it as an Obsidian vault you can explore, filter, and query.
any folder of files → persistent knowledge graph → Obsidian vault, graph.json, audit report
```
/graphify ./raw
@@ -8,40 +8,66 @@ A Claude Code skill that turns any folder of files into a navigable knowledge gr
```
.graphify/
├── obsidian/ open as Obsidian vault to explore the graph visually
├── GRAPH_REPORT.md what the graph found — surprising connections, knowledge gaps, suggested questions
└── graph.json persistent graph — query it weeks later without re-reading anything
├── obsidian/ open as Obsidian vault — visual graph, wikilinks, filter by community
├── GRAPH_REPORT.md what the graph found: god nodes, surprising connections, suggested questions
├── graph.json persistent graph — query it weeks later without re-reading anything
├── cache/ per-file SHA256 cache — re-runs only process changed files
└── memory/ Q&A results filed back in — what you ask grows the graph on next --update
```
## The problem it solves
[placeholder: animated GIF showing the full pipeline — detect → extract → cluster → report → Obsidian vault]
Andrej Karpathy described it well: he keeps a `/raw` folder where he drops papers, tweets, screenshots, and notes. The problem is that folder becomes opaque. You forget what's in it. You can't see what connects.
## Why this exists
Claude can read any single file. But ask Claude "what connects paper A to the code in repo B?" and it will hallucinate — it hasn't read both, and even if it has, it has no memory of the connection next session.
**The problem:** Andrej Karpathy described it well: he keeps a `/raw` folder where he drops papers, tweets, screenshots, and notes. The problem is that folder becomes opaque. You forget what's in it. You can't see what connects. Ask Claude "what links paper A to the code in repo B?" and it will hallucinate — it hasn't read both, and even if it has, it has no memory of that connection next session.
graphify solves this by:
**What LLMs get wrong:** Naive summarization fills in every gap confidently. You get a summary that sounds complete but you can't tell what was actually in the files vs invented by the model. And next session, it's all gone — no memory of what it extracted.
1. Reading everything once, extracting a persistent graph
2. Tagging every edge as `[EXTRACTED]` (explicitly stated), `[INFERRED]` (reasonable), or `[AMBIGUOUS]` (flagged for review) — you always know what was found vs invented
3. Running community detection to find clusters you didn't know existed
4. Surfacing cross-community connections — the things you would never think to ask about directly
5. Storing the graph in `.graphify/graph.json` so you can query it in any future session without re-extracting
**What graphify does differently:**
- **Persistent graph** — relationships are stored in `.graphify/graph.json` and survive across sessions. Query weeks later without re-reading anything.
- **Honest audit trail** — every edge is tagged `EXTRACTED` (explicitly stated), `INFERRED` (call-graph or reasonable deduction), or `AMBIGUOUS` (flagged for review). You always know what was found vs invented.
- **Cross-document surprise** — Leiden community detection finds clusters, then surfaces cross-community connections: the things you would never think to ask about directly.
- **Feedback loop** — every query answer is saved to `.graphify/memory/`. On next `--update`, that Q&A becomes a node. The graph grows from what you ask, not just what you add.
The result: a navigable map of your corpus that is honest about what it knows and what it guessed.
## Install
Copy the skill into your Claude Code skills directory:
```bash
pip install graphify && graphify install
```
That's it. This copies the skill file into `~/.claude/skills/graphify/` and registers it in `~/.claude/CLAUDE.md` automatically. The Python package and all dependencies install on first `/graphify` run — you never touch pip manually again.
Then open Claude Code in any directory and type:
```
/graphify .
```
<details>
<summary>Manual install (curl)</summary>
**Step 1 — copy the skill file**
```bash
mkdir -p ~/.claude/skills/graphify
curl -s https://raw.githubusercontent.com/safishamsi/graphify/v1/skills/graphify/skill.md \
curl -fsSL https://raw.githubusercontent.com/safishamsi/graphify/v1/skills/graphify/skill.md \
> ~/.claude/skills/graphify/SKILL.md
```
Add to `~/.claude/CLAUDE.md`:
**Step 2 — register it in Claude Code**
Add this to `~/.claude/CLAUDE.md` (create the file if it doesn't exist):
```
- **graphify** (`~/.claude/skills/graphify/SKILL.md`) — any input to knowledge graph. Trigger: `/graphify`
When the user types `/graphify`, invoke the Skill tool with `skill: "graphify"` before doing anything else.
```
</details>
## Usage
```bash
@@ -49,81 +75,107 @@ Add to `~/.claude/CLAUDE.md`:
/graphify ./raw # run on a specific folder
/graphify ./raw --mode deep # more aggressive INFERRED edge extraction
/graphify ./raw --update # re-extract only changed files, merge into existing graph
/graphify ./raw --watch # notify when new files appear (drop files, get pinged)
/graphify ./raw --watch # notify when new files appear
/graphify add https://arxiv.org/abs/1706.03762 # fetch a paper, save, update graph
/graphify add https://x.com/karpathy/status/... # fetch a tweet
/graphify add <url> --author "Karpathy" --contributor "safi" # tag who wrote it and who added it
/graphify add <url> --author "Karpathy" --contributor "safi"
/graphify query "what connects attention to the optimizer?" # BFS — broad context
/graphify query "how does the encoder reach the loss?" --dfs # DFS — trace a path
/graphify query "..." --budget 1500 # cap at N tokens
/graphify path "DigestAuth" "Response" # shortest path between two concepts
/graphify explain "SwinTransformer" # plain-language node explanation
/graphify ./raw --html # also export graph.html (browser, no Obsidian needed)
/graphify ./raw --svg # also export graph.svg (embeds in Notion, GitHub)
/graphify ./raw --neo4j # generate cypher.txt for Neo4j import
/graphify ./raw --mcp # start MCP stdio server for agent access
```
Works with any mix of file types in the same folder:
| Type | Extensions | How it's extracted |
|------|-----------|-------------------|
| Code | `.py .ts .js .go .rs .java .cpp .rb` etc | AST (deterministic) + semantic (Claude) |
| Documents | `.md .txt .rst` | Claude reads and extracts concepts + relationships |
| Code | `.py .ts .tsx .js .go .rs` | AST (deterministic) + call-graph pass (INFERRED) |
| Code | `.java .cpp .c .rb .swift .kt` | Claude semantic extraction |
| Documents | `.md .txt .rst` | Concepts + relationships via Claude |
| Papers | `.pdf` | Citation mining + concept extraction |
| Images | `.png .jpg .webp .gif .svg` | Claude vision — reads UI screenshots, charts, tweets, diagrams, whiteboards |
| Images | `.png .jpg .webp .gif .svg` | Claude vision — screenshots, charts, whiteboards, any language |
## What you get
After running, Claude pastes three things directly into the chat:
After running, Claude outputs three things directly in chat:
**God nodes** — the highest-degree concepts (what everything connects through)
**God nodes** — highest-degree concepts (what everything connects through)
**Surprising connections** — cross-community edges; relationships between concepts that live in different clusters. These are what you didn't know to look for.
**Surprising connections** — cross-community edges; relationships between concepts in different clusters that you didn't know to look for
**Suggested questions** — 4-5 questions the graph is uniquely positioned to answer, with the reason why (which bridge node makes it interesting, which community boundary it crosses)
The full `GRAPH_REPORT.md` also includes community summaries with cohesion scores and a list of ambiguous edges for your review.
The full GRAPH_REPORT.md adds community summaries with cohesion scores and a list of ambiguous edges for review.
## Use cases
## Key files explained
**New codebase** — run `/graphify` before touching anything. Find the god nodes (what you have to understand first), the community structure (what the major subsystems are), and the surprising connections (what talks to what that you wouldn't expect).
| File | Purpose |
|------|---------|
| `GRAPH_REPORT.md` | The audit report. God nodes, surprising connections, community cohesion scores, ambiguous edge list, suggested questions. |
| `graph.json` | Persistent graph in node-link format. Load it with NetworkX or push to Neo4j. Survives sessions. |
| `obsidian/` | Wikilink vault. Open in Obsidian → enable graph view → see communities as clusters. Filter by tag, search across everything. |
| `.graphify/cache/` | SHA256-based per-file cache. A re-run on an unchanged corpus takes seconds. |
| `.graphify/memory/` | Q&A feedback loop. Every `/graphify query` answer is saved here. Next `--update` extracts it into the graph. |
**Research reading list** — drop papers, tweets, and notes into `/raw`. Run `/graphify ./raw`. Get a graph of how concepts connect across everything you've read. Query it: "what connects sparse autoencoders to superposition?"
## What this skill will NOT do
**Personal knowledge base** — leave `--watch` running on your `/raw` folder. Drop things in throughout the day. The graph grows. Query it weeks later without re-reading anything.
- **Won't invent edges** — `AMBIGUOUS` exists so uncertain relationships are flagged, not hidden. If the connection isn't clear, it's tagged, not fabricated.
- **Won't claim the graph is useful when it isn't** — a corpus over 2M words or 200 files gets a cost warning before proceeding.
- **Won't re-extract unchanged files** — SHA256 cache ensures warm re-runs skip everything that hasn't changed.
- **Won't visualize graphs over 5,000 nodes** — use `--no-viz` or query instead.
- **Won't download datasets or set up infrastructure** — graphify reads your files. What you put in the folder is what it works with.
- **Won't implement baselines or run experiments** — it reads and maps. Analysis is yours.
**Collaborative corpus** — use `--contributor` to tag who added what. The graph knows provenance. "What did safi add that connects to the attention mechanism?"
## Design principles
## What it will NOT do
1. **Extraction quality is everything** — clustering is downstream of it. A bad graph clusters into bad communities. The AST + call-graph pass exists because deterministic beats probabilistic for code.
2. **Show the numbers** — cohesion is `0.91`, not "good". Token cost is always printed. You know what you spent.
3. **The best output is what you didn't know** — Surprising Connections is not optional. God nodes you probably already suspected. Cross-community edges are what you came for.
4. **The graph earns its complexity** — below a certain density, just use Claude directly. The graph adds value when you have more than you can hold in context across sessions.
5. **What you ask grows the graph** — query results are filed back in automatically. The corpus is not static.
6. **Honest uncertainty** — `EXTRACTED`, `INFERRED`, `AMBIGUOUS` are not cosmetic labels. They are the difference between trusting the graph and being misled by it.
- Won't invent edges — `[AMBIGUOUS]` exists so uncertain relationships are flagged, not hidden
- Won't claim the graph is useful when it isn't — corpus under 50K words gets a warning
- Won't re-extract unchanged files — `--update` uses a manifest to skip unchanged files
- Won't visualize graphs over 5,000 nodes — use `--no-viz` or query instead
## Contributing
## Files
**Adding worked examples**
```
graphify/
├── detect.py detect file types, auto-exclude venvs/caches/node_modules
├── extract.py parse files into nodes + edges (tree-sitter AST + Claude)
├── build.py assemble NetworkX graph from extraction JSON
├── cluster.py Leiden community detection, cohesion scoring
├── analyze.py god nodes, bridge nodes, surprising connections, suggested questions
├── report.py render GRAPH_REPORT.md
├── export.py Obsidian vault, graph.json, graph.html, graph.svg, Neo4j Cypher
├── ingest.py fetch URLs (arXiv, Twitter/X, PDF, any webpage), save annotated markdown
├── validate.py JSON schema checks on extraction output
├── serve.py MCP stdio server — exposes graph tools to other agents
└── watch.py fs watcher, writes flag file when new files appear
Worked examples are the most trust-building part of this project. To add one:
skills/graphify/
└── skill.md the Claude Code skill — everything the agent runs
1. Pick a real corpus (people should be able to verify the output)
2. Run the skill: `/graphify <path>`
3. Save the full output to `worked/{corpus_slug}/`
4. Write a `review.md` that honestly evaluates:
- What the graph got right
- What edges it correctly flagged AMBIGUOUS
- Any mistakes or missed connections
- Any surprising connections that were genuinely surprising
5. Submit a PR with all of the above
tests/ 71 tests, one file per module
pyproject.toml deps: networkx, graspologic, tree-sitter, pyvis
```
**Improving extraction**
If you find a file type or language where extraction is poor, open an issue with a minimal reproduction case. The best bug reports include: the input file, the extraction output (`.graphify/cache/` entry), and what was missed or invented.
**Adding domain knowledge**
If corpora in your domain consistently contain structures graphify doesn't extract well (e.g., legal documents, lab notebooks, musical scores), open a discussion with examples.
## Worked examples
| Corpus | Type | Eval report |
|--------|------|-------------|
| httpx (Python HTTP client) | Codebase | `tests/EVAL_httpx.md` + `tests/GRAPH_REPORT_httpx.md` |
| Mixed corpus (code + paper + Arabic image) | Multi-type | `tests/EVAL_mixed_corpus.md` |
Each includes the full graph output and an honest evaluation of what the skill got right and wrong.
## Tech stack
@@ -133,14 +185,30 @@ pyproject.toml deps: networkx, graspologic, tree-sitter, pyvis
| Community detection | Leiden via graspologic | Better than K-means for sparse graphs |
| Code parsing | tree-sitter | Multi-language AST, deterministic, zero hallucination |
| Extraction | Claude (parallel subagents) | Reads anything, outputs structured graph data |
| Visualization | Obsidian vault | Native graph view, wikilinks, search, no server needed |
| Visualization | Obsidian vault | Native graph view, wikilinks, no server needed |
No Neo4j required. No dashboards. No server. Runs entirely locally.
## Design principles
## Files
1. Extraction quality is everything — clustering is downstream of it
2. Show the numbers — cohesion is 0.91, not "good"
3. The best output is what you didn't know — Surprising Connections is not optional
4. Token cost is always visible
5. The graph earns its complexity — corpus under 50K words gets a warning to just use Claude directly
```
graphify/
├── detect.py detect file types, auto-exclude venvs/caches/node_modules; scan .graphify/memory/
├── extract.py AST extraction (Python, TypeScript, JavaScript, Go, Rust) + call-graph pass
├── build.py assemble NetworkX graph from extraction JSON; schema-validates before assembly
├── cluster.py Leiden community detection, cohesion scoring
├── analyze.py god nodes, bridge nodes, surprising connections, suggested questions, graph diff
├── report.py render GRAPH_REPORT.md
├── export.py Obsidian vault, graph.json, graph.html, graph.svg, Neo4j Cypher, Canvas
├── ingest.py fetch URLs (arXiv, Twitter/X, PDF, any webpage); save Q&A to .graphify/memory/
├── cache.py SHA256-based per-file extraction cache; check_semantic_cache / save_semantic_cache
├── validate.py JSON schema checks on extraction output
├── serve.py MCP stdio server — query_graph, get_node, get_neighbors, shortest_path, god_nodes
└── watch.py fs watcher, writes flag file when new files appear
skills/graphify/
└── skill.md the Claude Code skill — the full pipeline the agent runs step by step
tests/ 142 tests, one file per module
pyproject.toml pip install graphify | pip install graphify[mcp,neo4j,pdf,watch]
```
+26 -6
View File
@@ -1,7 +1,27 @@
"""graphify — extract · build · cluster · analyze · report."""
from graphify.extract import extract, collect_files
from graphify.build import build_from_json
from graphify.cluster import cluster, score_all, cohesion_score
from graphify.analyze import god_nodes, surprising_connections, suggest_questions
from graphify.report import generate
from graphify.export import to_json, to_html, to_svg, to_canvas
def __getattr__(name):
# Lazy imports so `graphify install` works before heavy deps are in place.
_map = {
"extract": ("graphify.extract", "extract"),
"collect_files": ("graphify.extract", "collect_files"),
"build_from_json": ("graphify.build", "build_from_json"),
"cluster": ("graphify.cluster", "cluster"),
"score_all": ("graphify.cluster", "score_all"),
"cohesion_score": ("graphify.cluster", "cohesion_score"),
"god_nodes": ("graphify.analyze", "god_nodes"),
"surprising_connections": ("graphify.analyze", "surprising_connections"),
"suggest_questions": ("graphify.analyze", "suggest_questions"),
"generate": ("graphify.report", "generate"),
"to_json": ("graphify.export", "to_json"),
"to_html": ("graphify.export", "to_html"),
"to_svg": ("graphify.export", "to_svg"),
"to_canvas": ("graphify.export", "to_canvas"),
}
if name in _map:
import importlib
mod_name, attr = _map[name]
mod = importlib.import_module(mod_name)
return getattr(mod, attr)
raise AttributeError(f"module 'graphify' has no attribute {name!r}")
+73
View File
@@ -0,0 +1,73 @@
"""graphify CLI — `graphify install` sets up the Claude Code skill."""
from __future__ import annotations
import shutil
import sys
from pathlib import Path
_SKILL_REGISTRATION = (
"\n# graphify\n"
"- **graphify** (`~/.claude/skills/graphify/SKILL.md`) "
"— any input to knowledge graph. Trigger: `/graphify`\n"
"When the user types `/graphify`, invoke the Skill tool "
"with `skill: \"graphify\"` before doing anything else.\n"
)
def _bundled_skill() -> Path:
"""Path to the skill.md bundled with this package."""
return Path(__file__).parent / "skill.md"
def install() -> None:
skill_src = _bundled_skill()
if not skill_src.exists():
print("error: skill.md not found in package — reinstall graphify", file=sys.stderr)
sys.exit(1)
# Copy skill to ~/.claude/skills/graphify/SKILL.md
skill_dst = Path.home() / ".claude" / "skills" / "graphify" / "SKILL.md"
skill_dst.parent.mkdir(parents=True, exist_ok=True)
shutil.copy(skill_src, skill_dst)
print(f" skill installed → {skill_dst}")
# Register in ~/.claude/CLAUDE.md
claude_md = Path.home() / ".claude" / "CLAUDE.md"
if claude_md.exists():
content = claude_md.read_text()
if "graphify" in content:
print(f" CLAUDE.md → already registered (no change)")
else:
claude_md.write_text(content.rstrip() + _SKILL_REGISTRATION)
print(f" CLAUDE.md → skill registered in {claude_md}")
else:
claude_md.parent.mkdir(parents=True, exist_ok=True)
claude_md.write_text(_SKILL_REGISTRATION.lstrip())
print(f" CLAUDE.md → created at {claude_md}")
print()
print("Done. Open Claude Code in any directory and type:")
print()
print(" /graphify .")
print()
def main() -> None:
if len(sys.argv) < 2 or sys.argv[1] in ("-h", "--help"):
print("Usage: graphify <command>")
print()
print("Commands:")
print(" install copy skill to ~/.claude/skills/ and register in CLAUDE.md")
print()
return
cmd = sys.argv[1]
if cmd == "install":
install()
else:
print(f"error: unknown command '{cmd}'", file=sys.stderr)
print("Run 'graphify --help' for usage.", file=sys.stderr)
sys.exit(1)
if __name__ == "__main__":
main()
+12
View File
@@ -332,6 +332,18 @@ def suggest_questions(
"why": f"Cohesion score {score} — nodes in this community are weakly interconnected.",
})
if not questions:
return [{
"type": "no_signal",
"question": None,
"why": (
"Not enough signal to generate questions. "
"This usually means the corpus has no AMBIGUOUS edges, no bridge nodes, "
"no INFERRED relationships, and all communities are tightly cohesive. "
"Add more files or run with --mode deep to extract richer edges."
),
}]
return questions[:top_n]
+55
View File
@@ -61,3 +61,58 @@ def clear_cache(root: Path = Path(".")) -> None:
d = cache_dir(root)
for f in d.glob("*.json"):
f.unlink()
def check_semantic_cache(
files: list[str],
root: Path = Path("."),
) -> tuple[list[dict], list[dict], list[str]]:
"""Check semantic extraction cache for a list of absolute file paths.
Returns (cached_nodes, cached_edges, uncached_files).
Uncached files need Claude extraction; cached files are merged directly.
"""
cached_nodes: list[dict] = []
cached_edges: list[dict] = []
uncached: list[str] = []
for fpath in files:
result = load_cached(Path(fpath), root)
if result is not None:
cached_nodes.extend(result.get("nodes", []))
cached_edges.extend(result.get("edges", []))
else:
uncached.append(fpath)
return cached_nodes, cached_edges, uncached
def save_semantic_cache(
nodes: list[dict],
edges: list[dict],
root: Path = Path("."),
) -> int:
"""Save semantic extraction results to cache, keyed by source_file.
Groups nodes and edges by source_file, then saves one cache entry per file.
Returns the number of files cached.
"""
from collections import defaultdict
by_file: dict[str, dict] = defaultdict(lambda: {"nodes": [], "edges": []})
for n in nodes:
src = n.get("source_file", "")
if src:
by_file[src]["nodes"].append(n)
for e in edges:
src = e.get("source_file", "")
if src:
by_file[src]["edges"].append(e)
saved = 0
for fpath, result in by_file.items():
p = Path(fpath)
if p.exists():
save_cached(p, result, root)
saved += 1
return saved
+18 -4
View File
@@ -1,7 +1,6 @@
"""Leiden community detection on NetworkX graphs. Splits oversized communities. Returns cohesion scores."""
from __future__ import annotations
import networkx as nx
from graspologic.partition import leiden
def build_graph(nodes: list[dict], edges: list[dict]) -> nx.Graph:
@@ -37,10 +36,24 @@ def cluster(G: nx.Graph) -> dict[int, list[str]]:
if G.number_of_edges() == 0:
return {i: [n] for i, n in enumerate(sorted(G.nodes))}
partition: dict[str, int] = leiden(G)
from graspologic.partition import leiden # lazy — avoids 15s numba JIT on import
# Leiden warns and drops isolates — handle them separately
isolates = [n for n in G.nodes() if G.degree(n) == 0]
connected_nodes = [n for n in G.nodes() if G.degree(n) > 0]
connected = G.subgraph(connected_nodes)
raw: dict[int, list[str]] = {}
for node, cid in partition.items():
raw.setdefault(cid, []).append(node)
if connected.number_of_nodes() > 0:
partition: dict[str, int] = leiden(connected)
for node, cid in partition.items():
raw.setdefault(cid, []).append(node)
# Each isolate becomes its own single-node community
next_cid = max(raw.keys(), default=-1) + 1
for node in isolates:
raw[next_cid] = [node]
next_cid += 1
# Split oversized communities
max_size = max(_MIN_SPLIT_SIZE, int(G.number_of_nodes() * _MAX_COMMUNITY_FRACTION))
@@ -63,6 +76,7 @@ def _split_community(G: nx.Graph, nodes: list[str]) -> list[list[str]]:
# No edges — split into individual nodes
return [[n] for n in sorted(nodes)]
try:
from graspologic.partition import leiden
sub_partition: dict[str, int] = leiden(subgraph)
sub_communities: dict[int, list[str]] = {}
for node, cid in sub_partition.items():
+22 -15
View File
@@ -152,24 +152,31 @@ def detect(root: Path) -> dict:
seen: set[Path] = set()
all_files: list[Path] = []
for scan_root in scan_paths:
for p in sorted(scan_root.rglob("*")):
if p not in seen:
seen.add(p)
all_files.append(p)
in_memory_tree = memory_dir.exists() and str(scan_root).startswith(str(memory_dir))
import os
for dirpath, dirnames, filenames in os.walk(scan_root):
dp = Path(dirpath)
if not in_memory_tree:
# Prune noise dirs in-place so os.walk never descends into them
dirnames[:] = [
d for d in dirnames
if not d.startswith(".") and not _is_noise_dir(d)
]
for fname in filenames:
p = dp / fname
if p not in seen:
seen.add(p)
all_files.append(p)
for p in all_files:
if not p.is_file():
continue
# For memory dir files, don't apply hidden/noise filtering
# For memory dir files, skip hidden/noise filtering
in_memory = memory_dir.exists() and str(p).startswith(str(memory_dir))
if not in_memory:
try:
parts = p.relative_to(root).parts
except ValueError:
continue
# Skip hidden dirs and known noise dirs
if any(part.startswith(".") or _is_noise_dir(part) for part in parts):
# Hidden files are already excluded via dir pruning above,
# but catch hidden files at the root level
if p.name.startswith("."):
continue
if _is_sensitive(p):
skipped_sensitive.append(str(p))
@@ -221,8 +228,8 @@ def save_manifest(files: dict[str, list[str]], manifest_path: str = _MANIFEST_PA
for f in file_list:
try:
manifest[f] = Path(f).stat().st_mtime
except Exception:
pass
except OSError:
pass # file deleted between detect() and manifest write — skip it
Path(manifest_path).parent.mkdir(parents=True, exist_ok=True)
Path(manifest_path).write_text(json.dumps(manifest, indent=2))
+4
View File
@@ -0,0 +1,4 @@
# re-export manifest helpers from detect for backwards compatibility
from graphify.detect import save_manifest, load_manifest, detect_incremental
__all__ = ["save_manifest", "load_manifest", "detect_incremental"]
+10 -5
View File
@@ -119,10 +119,15 @@ def generate(
if suggested_questions:
lines += ["", "## Suggested Questions"]
lines.append("_Questions this graph is uniquely positioned to answer:_")
lines.append("")
for q in suggested_questions:
lines.append(f"- **{q['question']}**")
lines.append(f" _{q['why']}_")
no_signal = len(suggested_questions) == 1 and suggested_questions[0].get("type") == "no_signal"
if no_signal:
lines.append(f"_{suggested_questions[0]['why']}_")
else:
lines.append("_Questions this graph is uniquely positioned to answer:_")
lines.append("")
for q in suggested_questions:
if q.get("question"):
lines.append(f"- **{q['question']}**")
lines.append(f" _{q['why']}_")
return "\n".join(lines)
+1046
View File
File diff suppressed because it is too large Load Diff
+1 -1
View File
@@ -1,7 +1,7 @@
# validate extraction JSON against the graphify schema before graph assembly
from __future__ import annotations
VALID_FILE_TYPES = {"code", "document", "paper"}
VALID_FILE_TYPES = {"code", "document", "paper", "image"}
VALID_CONFIDENCES = {"EXTRACTED", "INFERRED", "AMBIGUOUS"}
REQUIRED_NODE_FIELDS = {"id", "label", "file_type", "source_file"}
REQUIRED_EDGE_FIELDS = {"source", "target", "relation", "confidence", "source_file"}
+20
View File
@@ -5,6 +5,9 @@ build-backend = "setuptools.build_meta"
[project]
name = "graphify"
version = "0.1.1"
description = "Turn any codebase, docs, or images into a queryable knowledge graph"
readme = "README.md"
license = { text = "MIT" }
requires-python = ">=3.10"
dependencies = [
"networkx",
@@ -12,8 +15,25 @@ dependencies = [
"pyvis",
"tree-sitter",
"tree-sitter-python",
"tree-sitter-javascript",
"tree-sitter-typescript",
"tree-sitter-go",
"tree-sitter-rust",
]
[project.optional-dependencies]
mcp = ["mcp"]
neo4j = ["neo4j"]
pdf = ["pypdf", "html2text"]
watch = ["watchdog"]
all = ["mcp", "neo4j", "pypdf", "html2text", "watchdog"]
[project.scripts]
graphify = "graphify.__main__:main"
[tool.setuptools.packages.find]
where = ["."]
include = ["graphify*"]
[tool.setuptools.package-data]
graphify = ["skill.md"]
+30 -57
View File
@@ -54,36 +54,41 @@ If no path was given, use `.` (current directory). Do not ask the user for a pat
Follow these steps in order. Do not skip steps.
### Step 1 — Install dependencies (skip if already installed)
### Step 1 — Ensure graphify is installed
```bash
python3 -c "import graphify, networkx, graspologic, pyvis, tree_sitter" 2>/dev/null || {
pip install networkx graspologic pyvis tree-sitter tree-sitter-python -q --break-system-packages 2>&1 | tail -3
pip install git+https://github.com/safishamsi/graphify.git -q --break-system-packages 2>&1 | tail -3
}
python3 -c "import graphify" 2>/dev/null || pip install graphify -q --break-system-packages 2>&1 | tail -3
```
If all imports succeed, print nothing and move straight to Step 2.
If the import succeeds, print nothing and move straight to Step 2.
### Step 2 — Detect files
```bash
python3 -c "
import sys, json
import json
from graphify.detect import detect
from pathlib import Path
result = detect(Path('INPUT_PATH'))
print(json.dumps(result, indent=2))
print(json.dumps(result))
" > .graphify_detect.json
cat .graphify_detect.json
```
Replace INPUT_PATH with the actual path the user provided.
Replace INPUT_PATH with the actual path the user provided. Do NOT cat or print the JSON — read it silently and present a clean summary instead:
After detection:
- If `skipped_sensitive` is non-empty, tell the user which files were skipped — do not ask, just inform.
- If `total_files` is 0, stop: "No supported files found in [path]. Supported: .py .ts .js .go .rs .java .cpp .rb .md .txt .rst .pdf"
- If `total_words` > 2,000,000, tell the user the word count and suggest a subfolder, but **do not block** — proceed unless they say stop.
```
Corpus: X files · ~Y words
code: N files (.py .ts .go ...)
docs: N files (.md .txt ...)
papers: N files (.pdf ...)
images: N files
```
Then act on it:
- If `total_files` is 0: stop with "No supported files found in [path]."
- If `skipped_sensitive` is non-empty: mention file count skipped, not the file names.
- If `total_words` > 2,000,000 OR `total_files` > 200: show the warning and the top 5 subdirectories by file count, then ask which subfolder to run on. Wait for the user's answer before proceeding.
- Otherwise: proceed directly to Step 3 — no need to ask anything.
### Step 3 — Extract entities and relationships
@@ -131,35 +136,18 @@ Before dispatching any subagents, check which files already have cached extracti
```bash
python3 -c "
import json
from graphify.cache import load_cached
from graphify.cache import check_semantic_cache
from pathlib import Path
detect = json.loads(Path('.graphify_detect.json').read_text())
all_files = []
for ftype, files in detect['files'].items():
all_files.extend(files)
all_files = [f for files in detect['files'].values() for f in files]
cached = []
uncached = []
for f in all_files:
result = load_cached(Path(f))
if result is not None:
cached.append((f, result))
else:
uncached.append(f)
# Write cached results directly to a partial semantic file
if cached:
nodes, edges = [], []
for f, r in cached:
nodes.extend(r.get('nodes', []))
edges.extend(r.get('edges', []))
import json
from pathlib import Path
Path('.graphify_cached.json').write_text(json.dumps({'nodes': nodes, 'edges': edges}))
cached_nodes, cached_edges, uncached = check_semantic_cache(all_files)
if cached_nodes or cached_edges:
Path('.graphify_cached.json').write_text(json.dumps({'nodes': cached_nodes, 'edges': cached_edges}))
Path('.graphify_uncached.txt').write_text('\n'.join(uncached))
print(f'Cache: {len(cached)} files hit, {len(uncached)} files need extraction')
print(f'Cache: {len(all_files)-len(uncached)} files hit, {len(uncached)} files need extraction')
"
```
@@ -227,28 +215,13 @@ If more than half the chunks failed, stop and tell the user.
Save new results to cache:
```bash
python3 -c "
from graphify.cache import save_cached
from pathlib import Path
import json
from graphify.cache import save_semantic_cache
from pathlib import Path
# Load the new semantic results (written by subagents as .graphify_semantic_new.json)
new = json.loads(Path('.graphify_semantic_new.json').read_text())
# Group nodes/edges back by source_file and cache each file's results
from collections import defaultdict
by_file = defaultdict(lambda: {'nodes': [], 'edges': []})
for n in new.get('nodes', []):
if n.get('source_file'):
by_file[n['source_file']]['nodes'].append(n)
for e in new.get('edges', []):
if e.get('source_file'):
by_file[e['source_file']]['edges'].append(e)
for fpath, result in by_file.items():
p = Path(fpath)
if p.exists():
save_cached(p, result)
print(f'Cached {len(by_file)} files')
new = json.loads(Path('.graphify_semantic_new.json').read_text()) if Path('.graphify_semantic_new.json').exists() else {'nodes':[],'edges':[]}
saved = save_semantic_cache(new.get('nodes', []), new.get('edges', []))
print(f'Cached {saved} files')
"
```
+151
View File
@@ -0,0 +1,151 @@
"""Tests for serve.py — MCP graph query helpers (no mcp package required)."""
import json
import pytest
import networkx as nx
from networkx.readwrite import json_graph
from graphify.serve import (
_communities_from_graph,
_score_nodes,
_bfs,
_dfs,
_subgraph_to_text,
_load_graph,
)
def _make_graph() -> nx.Graph:
G = nx.Graph()
G.add_node("n1", label="extract", source_file="extract.py", source_location="L10", community=0)
G.add_node("n2", label="cluster", source_file="cluster.py", source_location="L5", community=0)
G.add_node("n3", label="build", source_file="build.py", source_location="L1", community=1)
G.add_node("n4", label="report", source_file="report.py", source_location="L1", community=1)
G.add_node("n5", label="isolated", source_file="other.py", source_location="L1", community=2)
G.add_edge("n1", "n2", relation="calls", confidence="INFERRED")
G.add_edge("n2", "n3", relation="imports", confidence="EXTRACTED")
G.add_edge("n3", "n4", relation="uses", confidence="EXTRACTED")
return G
# --- _communities_from_graph ---
def test_communities_from_graph_basic():
G = _make_graph()
communities = _communities_from_graph(G)
assert 0 in communities
assert 1 in communities
assert "n1" in communities[0]
assert "n2" in communities[0]
assert "n3" in communities[1]
def test_communities_from_graph_no_community_attr():
G = nx.Graph()
G.add_node("a", label="foo") # no community attr
communities = _communities_from_graph(G)
assert communities == {}
def test_communities_from_graph_isolated():
G = _make_graph()
communities = _communities_from_graph(G)
assert 2 in communities
assert "n5" in communities[2]
# --- _score_nodes ---
def test_score_nodes_exact_label_match():
G = _make_graph()
scored = _score_nodes(G, ["extract"])
nids = [nid for _, nid in scored]
assert "n1" in nids
assert scored[0][1] == "n1" # highest score first
def test_score_nodes_no_match():
G = _make_graph()
scored = _score_nodes(G, ["xyzzy"])
assert scored == []
def test_score_nodes_source_file_partial():
G = _make_graph()
# "cluster.py" contains "cluster" — should score 0.5 for source match
scored = _score_nodes(G, ["cluster"])
nids = [nid for _, nid in scored]
assert "n2" in nids
# --- _bfs ---
def test_bfs_depth_1():
G = _make_graph()
visited, edges = _bfs(G, ["n1"], depth=1)
assert "n1" in visited
assert "n2" in visited # direct neighbor
assert "n3" not in visited # 2 hops away
def test_bfs_depth_2():
G = _make_graph()
visited, edges = _bfs(G, ["n1"], depth=2)
assert "n3" in visited # n1 -> n2 -> n3
def test_bfs_disconnected():
G = _make_graph()
visited, edges = _bfs(G, ["n5"], depth=3)
assert visited == {"n5"} # isolated node
def test_bfs_returns_edges():
G = _make_graph()
visited, edges = _bfs(G, ["n1"], depth=1)
assert len(edges) >= 1
assert any(u == "n1" or v == "n1" for u, v in edges)
# --- _dfs ---
def test_dfs_depth_1():
G = _make_graph()
visited, edges = _dfs(G, ["n1"], depth=1)
assert "n1" in visited
assert "n2" in visited
assert "n3" not in visited
def test_dfs_full_chain():
G = _make_graph()
visited, edges = _dfs(G, ["n1"], depth=5)
assert {"n1", "n2", "n3", "n4"}.issubset(visited)
# --- _subgraph_to_text ---
def test_subgraph_to_text_contains_labels():
G = _make_graph()
text = _subgraph_to_text(G, {"n1", "n2"}, [("n1", "n2")])
assert "extract" in text
assert "cluster" in text
def test_subgraph_to_text_truncates():
G = _make_graph()
# Very small budget forces truncation
text = _subgraph_to_text(G, {"n1", "n2", "n3", "n4"}, [("n1", "n2")], token_budget=1)
assert "truncated" in text
def test_subgraph_to_text_edge_included():
G = _make_graph()
text = _subgraph_to_text(G, {"n1", "n2"}, [("n1", "n2")])
assert "EDGE" in text
assert "calls" in text
# --- _load_graph ---
def test_load_graph_roundtrip(tmp_path):
G = _make_graph()
data = json_graph.node_link_data(G, edges="links")
p = tmp_path / "graph.json"
p.write_text(json.dumps(data))
G2 = _load_graph(str(p))
assert G2.number_of_nodes() == G.number_of_nodes()
assert G2.number_of_edges() == G.number_of_edges()
def test_load_graph_missing_file(tmp_path):
with pytest.raises(Exception):
_load_graph(str(tmp_path / "nonexistent.json"))
+68
View File
@@ -0,0 +1,68 @@
"""Tests for watch.py — file watcher helpers (no watchdog required)."""
import time
from pathlib import Path
import pytest
from graphify.watch import _run_update, _WATCHED_EXTENSIONS
# --- _run_update ---
def test_run_update_creates_flag(tmp_path):
_run_update(tmp_path)
flag = tmp_path / ".graphify" / "needs_update"
assert flag.exists()
assert flag.read_text() == "1"
def test_run_update_creates_flag_dir(tmp_path):
# .graphify dir does not exist yet
assert not (tmp_path / ".graphify").exists()
_run_update(tmp_path)
assert (tmp_path / ".graphify").is_dir()
def test_run_update_idempotent(tmp_path):
_run_update(tmp_path)
_run_update(tmp_path)
flag = tmp_path / ".graphify" / "needs_update"
assert flag.read_text() == "1"
# --- _WATCHED_EXTENSIONS ---
def test_watched_extensions_includes_code():
assert ".py" in _WATCHED_EXTENSIONS
assert ".ts" in _WATCHED_EXTENSIONS
assert ".go" in _WATCHED_EXTENSIONS
assert ".rs" in _WATCHED_EXTENSIONS
def test_watched_extensions_includes_docs():
assert ".md" in _WATCHED_EXTENSIONS
assert ".txt" in _WATCHED_EXTENSIONS
assert ".pdf" in _WATCHED_EXTENSIONS
def test_watched_extensions_includes_images():
assert ".png" in _WATCHED_EXTENSIONS
assert ".jpg" in _WATCHED_EXTENSIONS
def test_watched_extensions_excludes_noise():
assert ".json" not in _WATCHED_EXTENSIONS
assert ".pyc" not in _WATCHED_EXTENSIONS
assert ".log" not in _WATCHED_EXTENSIONS
# --- watch() import error without watchdog ---
def test_watch_raises_without_watchdog(tmp_path, monkeypatch):
import builtins
real_import = builtins.__import__
def mock_import(name, *args, **kwargs):
if name == "watchdog.observers" or name == "watchdog.events":
raise ImportError("mocked missing watchdog")
return real_import(name, *args, **kwargs)
monkeypatch.setattr(builtins, "__import__", mock_import)
from graphify.watch import watch
with pytest.raises(ImportError, match="watchdog not installed"):
watch(tmp_path)