feat: BYOND DreamMaker support, --mode deep flag, changelog fixes (#884, #1030)

- Feat: extract_dm (tree-sitter-dm), extract_dmi (PNG icon states),
  extract_dmm (tile dict uses edges), extract_dmf (window/elem hierarchy)
  for .dm .dme .dmi .dmm .dmf; 26 tests, fixtures, pyproject.toml dep
- Feat: graphify extract --mode deep flag; deep_mode threaded through all
  four LLM backends via extract_corpus_parallel
- Fix: CHANGELOG 0.8.21 entries for #1050, #1046, #1047 that were missing

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
This commit is contained in:
Safi
2026-05-27 23:55:12 +01:00
co-authored by Claude Sonnet 4.6
parent c09fbef401
commit dacbdb539a
12 changed files with 850 additions and 18 deletions
+8
View File
@@ -2,6 +2,11 @@
Full release notes with details on each version: [GitHub Releases](https://github.com/safishamsi/graphify/releases)
## 0.8.22 (unreleased)
- Feat: BYOND DreamMaker support — `.dm`/`.dme` files extracted via tree-sitter-dm (type definitions, proc declarations, `#include` edges, in-file call resolution, `new /type()` instantiation edges); `.dmi` PNG icon files parsed for icon-state nodes; `.dmm` map files parsed for type-path `uses` edges from the tile dictionary section; `.dmf` interface files parsed for window/elem/control-type hierarchy (#884)
- Feat: `graphify extract --mode deep` flag enables richer semantic extraction using an extended system prompt; flag propagated through all four LLM backends (#1030)
## 0.8.21 (2026-05-27)
- Fix: `graphify update` (no `--changed` flag) no longer leaves ghost nodes from files deleted between runs — full re-extraction path now reconciles the existing graph against current disk state and evicts any node whose `source_file` no longer exists; `_norm_source_file` used on both sides to guarantee path format consistency (#1007)
@@ -12,6 +17,9 @@ Full release notes with details on each version: [GitHub Releases](https://githu
- Fix: query punctuation no longer breaks node matching — `"what calls extract?"` correctly finds the `extract` node; `_search_tokens` helper strips punctuation from search terms in `_query_terms`, `_score_nodes`, and `_find_node` (#994, #978)
- Fix: language built-in globals (`String`, `Number`, `Boolean`, `Object`, `Array`, etc.) no longer accumulate spurious call edges — filtered at same-file and cross-file resolution in the AST extractor, eliminating god-node pollution from constructor-style calls (#916, #726)
- Feat: SystemVerilog header files (`.svh`) now extracted using the Verilog parser alongside `.v` and `.sv` (#1042)
- Fix: `@property`, `@staticmethod`, `@classmethod` methods no longer produce orphaned nodes without a class-qualified ID — `decorated_definition` is now treated as a transparent wrapper in the Python AST walker, preserving `parent_class_nid` through the decorator layer (#1050)
- Fix: Pass 2 dedup no longer merges nodes with identical labels that live in different files — same-file partition enforced for the identical-label subcase so `foo()` in `a.py` and `foo()` in `b.py` are not collapsed into one node (#1046)
- Fix: `graphify-out/memory/` files are no longer silently excluded by `.gitignore` pattern matching — memory dir files now bypass the gitignore filter in `detect.py`, ensuring knowledge accumulated via `graphify remember` is always scanned (#1047)
## 0.8.20 (2026-05-26)
+1 -1
View File
@@ -214,7 +214,7 @@ To remove graphify from all platforms at once: `graphify uninstall` (add `--purg
| Type | Extensions |
|------|-----------|
| Code (32 languages) | `.py .ts .js .jsx .tsx .mjs .go .rs .java .c .cpp .h .hpp .rb .cs .kt .scala .php .swift .lua .luau .zig .ps1 .ex .exs .m .mm .jl .vue .svelte .astro .groovy .gradle .dart .v .sv .svh .sql .f .f90 .f95 .f03 .f08 .pas .pp .dpr .dpk .lpr .inc .dfm .lfm .lpk .sh .bash .json .sln .csproj .fsproj .vbproj .razor .cshtml` |
| Code (33 languages) | `.py .ts .js .jsx .tsx .mjs .go .rs .java .c .cpp .h .hpp .rb .cs .kt .scala .php .swift .lua .luau .zig .ps1 .ex .exs .m .mm .jl .vue .svelte .astro .groovy .gradle .dart .v .sv .svh .sql .f .f90 .f95 .f03 .f08 .pas .pp .dpr .dpk .lpr .inc .dfm .lfm .lpk .sh .bash .json .dm .dme .dmi .dmm .dmf .sln .csproj .fsproj .vbproj .razor .cshtml` |
| MCP configs | `.mcp.json` `mcp.json` `mcp_servers.json` `claude_desktop_config.json` — extracts server nodes, package refs, env var requirements |
| Docs | `.md .mdx .qmd .html .txt .rst .yaml .yml` |
| Office | `.docx .xlsx` (requires `pip install graphifyy[office]`) |
+21 -1
View File
@@ -1495,6 +1495,7 @@ def main() -> None:
print(" extract <path> headless full extraction (AST + semantic LLM) for CI/scripts")
print(" --backend B gemini|kimi|claude|openai|deepseek|ollama (default: whichever API key is set)")
print(" --model M override backend default model")
print(" --mode deep aggressive INFERRED-edge semantic extraction")
print(" --max-workers N AST extraction subprocess count (default: cpu_count)")
print(" --token-budget N per-chunk token cap for semantic extraction (default: 60000)")
print(" --max-concurrency N parallel semantic chunks in flight (default: 4; set 1 for local LLMs)")
@@ -2899,7 +2900,7 @@ def main() -> None:
if len(sys.argv) < 3:
print(
"Usage: graphify extract <path> [--backend gemini|kimi|claude|openai|deepseek|ollama] "
"[--model M] [--out DIR] [--google-workspace] [--no-cluster] "
"[--model M] [--mode deep] [--out DIR] [--google-workspace] [--no-cluster] "
"[--max-workers N] [--token-budget N] [--max-concurrency N] "
"[--api-timeout S]",
file=sys.stderr,
@@ -2913,6 +2914,7 @@ def main() -> None:
backend: str | None = None
model: str | None = None
extract_mode: str | None = None
out_dir: Path | None = None
no_cluster = False
dedup_llm = False
@@ -2963,6 +2965,10 @@ def main() -> None:
model = args[i + 1]; i += 2
elif a.startswith("--model="):
model = a.split("=", 1)[1]; i += 1
elif a == "--mode" and i + 1 < len(args):
extract_mode = args[i + 1]; i += 2
elif a.startswith("--mode="):
extract_mode = a.split("=", 1)[1]; i += 1
elif a == "--out" and i + 1 < len(args):
out_dir = Path(args[i + 1]); i += 2
elif a.startswith("--out="):
@@ -3008,6 +3014,18 @@ def main() -> None:
else:
i += 1
_VALID_MODES = {"deep"}
if extract_mode is not None and extract_mode not in _VALID_MODES:
print(
f"error: unknown --mode '{extract_mode}'. "
f"Available: {', '.join(sorted(_VALID_MODES))}",
file=sys.stderr,
)
sys.exit(2)
deep_mode = extract_mode == "deep"
if deep_mode:
print("[graphify extract] deep mode enabled: richer semantic extraction")
# CLI flag wins over env var. Setting GRAPHIFY_API_TIMEOUT here so
# _call_openai_compat picks it up without needing a new kwarg path.
if cli_api_timeout is not None:
@@ -3194,6 +3212,8 @@ def main() -> None:
"model": model,
"root": target,
}
if deep_mode:
corpus_kwargs["deep_mode"] = True
if cli_token_budget is not None:
corpus_kwargs["token_budget"] = cli_token_budget
if cli_max_concurrency is not None:
+1 -1
View File
@@ -25,7 +25,7 @@ class FileType(str, Enum):
_MANIFEST_PATH = "graphify-out/manifest.json"
CODE_EXTENSIONS = {'.py', '.ts', '.tsx', '.js', '.jsx', '.mjs', '.ejs', '.ets', '.go', '.rs', '.java', '.groovy', '.gradle', '.cpp', '.cc', '.cxx', '.c', '.h', '.hpp', '.rb', '.swift', '.kt', '.kts', '.cs', '.scala', '.php', '.lua', '.luau', '.toc', '.zig', '.ps1', '.ex', '.exs', '.m', '.mm', '.jl', '.vue', '.svelte', '.astro', '.dart', '.v', '.sv', '.svh', '.sql', '.r', '.f', '.F', '.f90', '.F90', '.f95', '.F95', '.f03', '.F03', '.f08', '.F08', '.pas', '.pp', '.dpr', '.dpk', '.lpr', '.inc', '.dfm', '.lfm', '.lpk', '.sh', '.bash', '.json', '.sln', '.csproj', '.fsproj', '.vbproj', '.razor', '.cshtml'}
CODE_EXTENSIONS = {'.py', '.ts', '.tsx', '.js', '.jsx', '.mjs', '.ejs', '.ets', '.go', '.rs', '.java', '.groovy', '.gradle', '.cpp', '.cc', '.cxx', '.c', '.h', '.hpp', '.rb', '.swift', '.kt', '.kts', '.cs', '.scala', '.php', '.lua', '.luau', '.toc', '.zig', '.ps1', '.ex', '.exs', '.m', '.mm', '.jl', '.vue', '.svelte', '.astro', '.dart', '.v', '.sv', '.svh', '.sql', '.r', '.f', '.F', '.f90', '.F90', '.f95', '.F95', '.f03', '.F03', '.f08', '.F08', '.pas', '.pp', '.dpr', '.dpk', '.lpr', '.inc', '.dfm', '.lfm', '.lpk', '.sh', '.bash', '.json', '.dm', '.dme', '.dmi', '.dmm', '.dmf', '.sln', '.csproj', '.fsproj', '.vbproj', '.razor', '.cshtml'}
DOC_EXTENSIONS = {'.md', '.mdx', '.qmd', '.txt', '.rst', '.html', '.yaml', '.yml'}
PAPER_EXTENSIONS = {'.pdf'}
IMAGE_EXTENSIONS = {'.png', '.jpg', '.jpeg', '.gif', '.webp', '.svg'}
+511
View File
@@ -8234,6 +8234,512 @@ def extract_json(path: Path) -> dict:
return {"nodes": nodes, "edges": edges}
# ── DM (BYOND DreamMaker) extractor ──────────────────────────────────────────
# DM identity is path-based (`/datum/object/proc/New()`), not block-based, so
# the generic class-body walker doesn't fit well.
def extract_dm(path: Path) -> dict:
"""Extract types, procs, includes, and calls from a .dm/.dme file."""
try:
import tree_sitter_dm as tsdm
from tree_sitter import Language, Parser
except ImportError:
return {"nodes": [], "edges": [], "error": "tree-sitter-dm not installed"}
try:
language = Language(tsdm.language())
parser = Parser(language)
source = path.read_bytes()
tree = parser.parse(source)
root = tree.root_node
except Exception as e:
return {"nodes": [], "edges": [], "error": str(e)}
stem = _file_stem(path)
str_path = str(path)
nodes: list[dict] = []
edges: list[dict] = []
seen_ids: set[str] = set()
function_bodies: list[tuple[str, Any, "str | None"]] = []
def add_node(nid: str, label: str, line: int) -> None:
if nid and nid not in seen_ids:
seen_ids.add(nid)
nodes.append({"id": nid, "label": label, "file_type": "code",
"source_file": str_path, "source_location": f"L{line}"})
def add_edge(src: str, tgt: str, relation: str, line: int,
confidence: str = "EXTRACTED", weight: float = 1.0,
context: str | None = None) -> None:
if not src or not tgt or src == tgt:
return
edge: dict = {"source": src, "target": tgt, "relation": relation,
"confidence": confidence, "source_file": str_path,
"source_location": f"L{line}", "weight": weight}
if context:
edge["context"] = context
edges.append(edge)
file_nid = _make_id(str(path))
add_node(file_nid, path.name, 1)
def _type_path_text(node) -> str:
return _read_text(node, source).strip()
def _ensure_type(path_text: str, line: int) -> str:
nid = _make_id(stem, path_text)
add_node(nid, path_text, line)
return nid
def _find_child(node, type_name: str):
for c in node.children:
if c.type == type_name:
return c
return None
def _read_include_path(file_node) -> str:
if file_node is None:
return ""
if file_node.type == "string_literal":
parts = []
for c in file_node.children:
if c.type == "string_content":
parts.append(_read_text(c, source))
return "".join(parts)
return _read_text(file_node, source).strip("'\"")
def walk(node, parent_type_path: "str | None" = None,
parent_type_nid: "str | None" = None) -> None:
t = node.type
line = node.start_point[0] + 1
if t == "preproc_include":
file_node = node.child_by_field_name("file")
raw = _read_include_path(file_node)
if raw:
norm = raw.replace("\\", "/").lstrip("./")
resolved = (path.parent / norm).resolve()
edge: dict = {
"source": file_nid,
"target": _make_id(str(resolved)) if resolved.exists() else _make_id(norm),
"relation": "imports_from" if resolved.exists() else "imports",
"context": "import",
"confidence": "EXTRACTED",
"source_file": str_path,
"source_location": f"L{line}",
"weight": 1.0,
}
if not resolved.exists():
edge["external"] = True
edges.append(edge)
return
if t == "type_definition":
tp_node = _find_child(node, "type_path")
if tp_node is None:
return
type_path_str = _type_path_text(tp_node)
type_nid = _ensure_type(type_path_str, line)
add_edge(file_nid, type_nid, "contains", line)
body = _find_child(node, "type_body")
if body is not None:
for c in body.children:
walk(c, parent_type_path=type_path_str, parent_type_nid=type_nid)
return
if t in ("type_body_intended", "type_body_braced"):
for c in node.children:
walk(c, parent_type_path, parent_type_nid)
return
if t in ("type_proc_definition", "type_proc_override"):
if parent_type_nid is None or parent_type_path is None:
return
name_node = node.child_by_field_name("name")
if name_node is None:
return
proc_name = _read_text(name_node, source)
proc_nid = _make_id(stem, parent_type_path, proc_name)
add_node(proc_nid, f"{parent_type_path}/{proc_name}()", line)
add_edge(parent_type_nid, proc_nid, "method", line)
block = _find_child(node, "block")
if block is not None:
function_bodies.append((proc_nid, block, parent_type_path))
return
if t in ("proc_definition", "proc_override"):
tp_node = _find_child(node, "type_path")
owner_path: "str | None" = None
owner_nid: "str | None" = None
if tp_node is not None:
owner_path = _type_path_text(tp_node)
owner_nid = _ensure_type(owner_path, line)
add_edge(file_nid, owner_nid, "contains", line)
name_node = node.child_by_field_name("name")
if name_node is None:
return
proc_name = _read_text(name_node, source)
if owner_path and owner_nid:
proc_nid = _make_id(stem, owner_path, proc_name)
add_node(proc_nid, f"{owner_path}/{proc_name}()", line)
add_edge(owner_nid, proc_nid, "method", line)
else:
proc_nid = _make_id(stem, proc_name)
add_node(proc_nid, f"{proc_name}()", line)
add_edge(file_nid, proc_nid, "contains", line)
block = _find_child(node, "block")
if block is not None:
function_bodies.append((proc_nid, block, owner_path))
return
if t in ("operator_override", "type_operator_override"):
return
for child in node.children:
walk(child, parent_type_path, parent_type_nid)
walk(root)
label_to_nids: dict[str, list[str]] = {}
path_to_nids: dict[str, list[str]] = {}
for n in nodes:
label = n["label"].strip("()")
last = label.rsplit("/", 1)[-1] if "/" in label else label
if last:
label_to_nids.setdefault(last.lower(), []).append(n["id"])
if label.startswith("/"):
path_to_nids.setdefault(label.lower(), []).append(n["id"])
seen_call_pairs: set[tuple[str, str]] = set()
raw_calls: list[dict] = []
def _emit_call(caller_nid: str, callee: str, line: int, is_member: bool) -> None:
candidates = label_to_nids.get(callee.lower(), [])
tgt_nid = candidates[0] if len(candidates) == 1 else None
if tgt_nid and tgt_nid != caller_nid:
pair = (caller_nid, tgt_nid)
if pair in seen_call_pairs:
return
seen_call_pairs.add(pair)
edges.append({
"source": caller_nid, "target": tgt_nid, "relation": "calls",
"context": "call", "confidence": "EXTRACTED",
"source_file": str_path, "source_location": f"L{line}", "weight": 1.0,
})
else:
raw_calls.append({
"caller_nid": caller_nid, "callee": callee,
"is_member_call": is_member, "source_file": str_path,
"source_location": f"L{line}",
})
def walk_calls(body_node, caller_nid: str) -> None:
if body_node is None:
return
t = body_node.type
if t in ("proc_definition", "proc_override", "type_proc_definition",
"type_proc_override", "type_definition"):
return
if t == "call_expression":
name_node = body_node.child_by_field_name("name")
if name_node is not None:
callee = _read_text(name_node, source)
if callee and callee != "..":
_emit_call(caller_nid, callee, body_node.start_point[0] + 1,
is_member=False)
elif t == "field_proc_expression":
proc_field = body_node.child_by_field_name("proc")
if proc_field is not None:
callee = _read_text(proc_field, source)
if callee:
_emit_call(caller_nid, callee, body_node.start_point[0] + 1,
is_member=True)
elif t == "new_expression":
tp_node = _find_child(body_node, "type_path")
if tp_node is not None:
target_text = _type_path_text(tp_node)
candidates = path_to_nids.get(target_text.lower(), [])
tgt_nid = candidates[0] if len(candidates) == 1 else None
if tgt_nid and tgt_nid != caller_nid:
pair = (caller_nid, tgt_nid)
if pair not in seen_call_pairs:
seen_call_pairs.add(pair)
edges.append({
"source": caller_nid, "target": tgt_nid,
"relation": "instantiates", "context": "call",
"confidence": "EXTRACTED", "source_file": str_path,
"source_location": f"L{body_node.start_point[0] + 1}",
"weight": 1.0,
})
for child in body_node.children:
walk_calls(child, caller_nid)
for proc_nid, block, _owner_path in function_bodies:
walk_calls(block, proc_nid)
return {"nodes": nodes, "edges": edges, "raw_calls": raw_calls}
# ── DMI (BYOND icon files) ────────────────────────────────────────────────────
# .dmi is a PNG with a tEXt/zTXt "Description" chunk containing BYOND state
# metadata. We want the icon state names (icon_state = "X" in DM code
# references them).
def _read_dmi_description(data: bytes) -> str:
"""Pull the BYOND metadata text out of a .dmi PNG, or empty string on failure."""
import struct
import zlib as _zlib
if not data.startswith(b"\x89PNG\r\n\x1a\n"):
return ""
i = 8
while i + 8 <= len(data):
length = struct.unpack(">I", data[i:i + 4])[0]
chunk_type = data[i + 4:i + 8]
payload = data[i + 8:i + 8 + length]
if chunk_type in (b"tEXt", b"zTXt"):
try:
null = payload.index(b"\x00")
except ValueError:
return ""
keyword = payload[:null]
if keyword == b"Description":
if chunk_type == b"zTXt":
return _zlib.decompress(payload[null + 2:]).decode("utf-8", errors="replace")
return payload[null + 1:].decode("utf-8", errors="replace")
i += 8 + length + 4
return ""
def extract_dmi(path: Path) -> dict:
"""Extract icon state names from a .dmi (BYOND PNG icon sheet)."""
try:
data = path.read_bytes()
except Exception as e:
return {"nodes": [], "edges": [], "error": str(e)}
str_path = str(path)
stem = _file_stem(path)
file_nid = _make_id(str(path))
nodes: list[dict] = [{"id": file_nid, "label": path.name, "file_type": "code",
"source_file": str_path, "source_location": "L1"}]
edges: list[dict] = []
seen: set[str] = {file_nid}
description = _read_dmi_description(data)
if not description:
return {"nodes": nodes, "edges": edges}
line_no = 0
for raw_line in description.splitlines():
line_no += 1
stripped = raw_line.strip()
if not stripped.startswith("state ="):
continue
value = stripped.split("=", 1)[1].strip()
if value.startswith('"') and value.endswith('"') and len(value) >= 2:
state_name = value[1:-1]
else:
state_name = value
if not state_name:
continue
nid = _make_id(stem, "state", state_name)
if nid in seen:
continue
seen.add(nid)
nodes.append({"id": nid, "label": f'"{state_name}"', "file_type": "code",
"source_file": str_path, "source_location": f"L{line_no}"})
edges.append({"source": file_nid, "target": nid, "relation": "contains",
"confidence": "EXTRACTED", "source_file": str_path,
"source_location": f"L{line_no}", "weight": 1.0})
return {"nodes": nodes, "edges": edges}
# ── DMM (BYOND map files) ─────────────────────────────────────────────────────
# A .dmm starts with a tile dictionary — each "key" = (type, type{var=val}, ...)
# names one or more types that compose a tile — then a grid. We only need the
# dictionary section: every type path referenced is a `uses` edge.
_DMM_GRID_RE = re.compile(r"^\(\s*\d+\s*,\s*\d+\s*,\s*\d+\s*\)\s*=", re.MULTILINE)
def _split_dmm_tile(body: str) -> list[str]:
out: list[str] = []
buf: list[str] = []
depth = 0
in_string = False
escape = False
for ch in body:
if escape:
buf.append(ch)
escape = False
continue
if in_string:
buf.append(ch)
if ch == "\\":
escape = True
elif ch == '"':
in_string = False
continue
if ch == '"':
in_string = True
buf.append(ch)
elif ch in "({[":
depth += 1
buf.append(ch)
elif ch in ")}]":
depth -= 1
buf.append(ch)
elif ch == "," and depth == 0:
out.append("".join(buf).strip())
buf = []
else:
buf.append(ch)
tail = "".join(buf).strip()
if tail:
out.append(tail)
return out
def _dmm_type_path(entry: str) -> str:
brace = entry.find("{")
if brace != -1:
entry = entry[:brace]
return entry.strip()
def extract_dmm(path: Path) -> dict:
"""Extract type-path references from a .dmm map file's tile dictionary."""
try:
text = path.read_text(encoding="utf-8", errors="replace")
except Exception as e:
return {"nodes": [], "edges": [], "error": str(e)}
str_path = str(path)
file_nid = _make_id(str(path))
nodes: list[dict] = [{"id": file_nid, "label": path.name, "file_type": "code",
"source_file": str_path, "source_location": "L1"}]
edges: list[dict] = []
grid_match = _DMM_GRID_RE.search(text)
dict_text = text[:grid_match.start()] if grid_match else text
seen_targets: set[str] = set()
buf: list[str] = []
open_line = 0
depth = 0
in_string = False
escape = False
for line_idx, line in enumerate(dict_text.splitlines(), start=1):
for ch in line:
if escape:
escape = False
elif in_string:
if ch == "\\":
escape = True
elif ch == '"':
in_string = False
elif ch == '"':
in_string = True
elif ch == "(":
if depth == 0:
open_line = line_idx
depth += 1
elif ch == ")":
depth -= 1
buf.append(ch)
buf.append("\n")
if depth == 0 and buf:
chunk = "".join(buf)
buf = []
lp = chunk.find("(")
rp = chunk.rfind(")")
if lp == -1 or rp == -1 or rp <= lp:
continue
inner = chunk[lp + 1:rp]
for entry in _split_dmm_tile(inner):
tpath = _dmm_type_path(entry)
if not tpath.startswith("/"):
continue
tgt = _make_id(tpath)
if tgt in seen_targets:
continue
seen_targets.add(tgt)
edges.append({"source": file_nid, "target": tgt, "relation": "uses",
"context": "map", "confidence": "EXTRACTED",
"source_file": str_path,
"source_location": f"L{open_line}", "weight": 1.0})
return {"nodes": nodes, "edges": edges}
# ── DMF (BYOND interface forms) ───────────────────────────────────────────────
_DMF_WINDOW_RE = re.compile(r'^\s*window\s+"([^"]+)"\s*$')
_DMF_ELEM_RE = re.compile(r'^\s*elem\s+"([^"]+)"\s*$')
_DMF_TYPE_RE = re.compile(r'^\s*type\s*=\s*(\S+)\s*$')
def extract_dmf(path: Path) -> dict:
"""Extract windows and controls from a .dmf interface file."""
try:
text = path.read_text(encoding="utf-8", errors="replace")
except Exception as e:
return {"nodes": [], "edges": [], "error": str(e)}
str_path = str(path)
stem = _file_stem(path)
file_nid = _make_id(str(path))
nodes: list[dict] = [{"id": file_nid, "label": path.name, "file_type": "code",
"source_file": str_path, "source_location": "L1"}]
edges: list[dict] = []
seen: set[str] = {file_nid}
current_window_nid: str | None = None
current_elem_nid: str | None = None
current_elem_name: str | None = None
for line_idx, line in enumerate(text.splitlines(), start=1):
m = _DMF_WINDOW_RE.match(line)
if m:
name = m.group(1)
nid = _make_id(stem, "window", name)
if nid not in seen:
seen.add(nid)
nodes.append({"id": nid, "label": f'window "{name}"', "file_type": "code",
"source_file": str_path, "source_location": f"L{line_idx}"})
edges.append({"source": file_nid, "target": nid, "relation": "contains",
"confidence": "EXTRACTED", "source_file": str_path,
"source_location": f"L{line_idx}", "weight": 1.0})
current_window_nid = nid
current_elem_nid = None
current_elem_name = None
continue
m = _DMF_ELEM_RE.match(line)
if m and current_window_nid is not None:
name = m.group(1)
nid = _make_id(stem, "elem", current_window_nid, name)
if nid not in seen:
seen.add(nid)
nodes.append({"id": nid, "label": f'elem "{name}"', "file_type": "code",
"source_file": str_path, "source_location": f"L{line_idx}"})
edges.append({"source": current_window_nid, "target": nid,
"relation": "contains", "confidence": "EXTRACTED",
"source_file": str_path, "source_location": f"L{line_idx}",
"weight": 1.0})
current_elem_nid = nid
current_elem_name = name
continue
m = _DMF_TYPE_RE.match(line)
if m and current_elem_nid is not None and current_elem_name is not None:
ctype = m.group(1)
for n in nodes:
if n["id"] == current_elem_nid and " [" not in n["label"]:
n["label"] = f'elem "{current_elem_name}" [{ctype}]'
break
return {"nodes": nodes, "edges": edges}
_DISPATCH: dict[str, Any] = {
".py": extract_python,
".js": extract_js,
@@ -8302,6 +8808,11 @@ _DISPATCH: dict[str, Any] = {
".sh": extract_bash,
".bash": extract_bash,
".json": extract_json,
".dm": extract_dm,
".dme": extract_dm,
".dmi": extract_dmi,
".dmm": extract_dmm,
".dmf": extract_dmf,
".sln": extract_sln,
".csproj": extract_csproj,
".fsproj": extract_csproj,
+38 -15
View File
@@ -146,6 +146,21 @@ Output exactly this schema:
{"nodes":[{"id":"stem_entity","label":"Human Readable Name","file_type":"code|document|paper|image|rationale|concept","source_file":"relative/path","source_location":null,"source_url":null,"captured_at":null,"author":null,"contributor":null}],"edges":[{"source":"node_id","target":"node_id","relation":"calls|implements|references|cites|conceptually_related_to|shares_data_with|semantically_similar_to","confidence":"EXTRACTED|INFERRED|AMBIGUOUS","confidence_score":1.0,"source_file":"relative/path","source_location":null,"weight":1.0}],"hyperedges":[],"input_tokens":0,"output_tokens":0}
"""
_DEEP_EXTRACTION_SUFFIX = """\
DEEP_MODE: include additional INFERRED edges only for concrete architectural
signals (shared data contracts, explicit lifecycle coupling, or multi-step flow
dependencies visible in the sources). Avoid broad conceptual similarity edges.
Mark uncertain ones AMBIGUOUS instead of omitting.
"""
def _extraction_system(*, deep: bool = False) -> str:
"""Return the semantic-extraction system prompt, optionally in deep mode."""
if not deep:
return _EXTRACTION_SYSTEM
return _EXTRACTION_SYSTEM + _DEEP_EXTRACTION_SUFFIX
def _read_files(paths: list[Path], root: Path) -> str:
"""Return file contents formatted for the extraction prompt."""
@@ -259,6 +274,7 @@ def _call_openai_compat(
max_completion_tokens: int = 8192,
*,
backend: str = "",
deep_mode: bool = False,
) -> dict:
"""Call any OpenAI-compatible API (Kimi, OpenAI, etc.) and return parsed JSON."""
try:
@@ -288,7 +304,7 @@ def _call_openai_compat(
kwargs: dict = {
"model": model,
"messages": [
{"role": "system", "content": _EXTRACTION_SYSTEM},
{"role": "system", "content": _extraction_system(deep=deep_mode)},
{"role": "user", "content": user_message},
],
"max_completion_tokens": max_completion_tokens,
@@ -384,7 +400,7 @@ def _call_openai_compat(
return result
def _call_claude(api_key: str, model: str, user_message: str, max_tokens: int = 8192) -> dict:
def _call_claude(api_key: str, model: str, user_message: str, max_tokens: int = 8192, *, deep_mode: bool = False) -> dict:
"""Call Anthropic Claude directly (not via OpenAI compat layer)."""
try:
import anthropic
@@ -398,7 +414,7 @@ def _call_claude(api_key: str, model: str, user_message: str, max_tokens: int =
resp = client.messages.create(
model=model,
max_tokens=max_tokens,
system=_EXTRACTION_SYSTEM,
system=_extraction_system(deep=deep_mode),
messages=[{"role": "user", "content": user_message}],
)
raw_content = resp.content[0].text if resp.content else None
@@ -420,7 +436,7 @@ def _call_claude(api_key: str, model: str, user_message: str, max_tokens: int =
return result
def _call_claude_cli(user_message: str, max_tokens: int = 8192) -> dict:
def _call_claude_cli(user_message: str, max_tokens: int = 8192, *, deep_mode: bool = False) -> dict:
"""Call Claude via the locally-installed Claude Code CLI (`claude -p`).
Routes through the user's Claude Code subscription auth instead of a separate
@@ -441,7 +457,7 @@ def _call_claude_cli(user_message: str, max_tokens: int = 8192) -> dict:
"claude", "-p",
"--output-format", "json",
"--no-session-persistence",
"--append-system-prompt", _EXTRACTION_SYSTEM,
"--append-system-prompt", _extraction_system(deep=deep_mode),
],
input=user_message,
capture_output=True,
@@ -486,7 +502,7 @@ def _call_claude_cli(user_message: str, max_tokens: int = 8192) -> dict:
return result
def _call_bedrock(model: str, user_message: str, max_tokens: int = 8192) -> dict:
def _call_bedrock(model: str, user_message: str, max_tokens: int = 8192, *, deep_mode: bool = False) -> dict:
"""Call AWS Bedrock via boto3 Converse API using the standard AWS credential chain."""
try:
import boto3
@@ -504,7 +520,7 @@ def _call_bedrock(model: str, user_message: str, max_tokens: int = 8192) -> dict
try:
resp = client.converse(
modelId=model,
system=[{"text": _EXTRACTION_SYSTEM}],
system=[{"text": _extraction_system(deep=deep_mode)}],
messages=[{"role": "user", "content": [{"text": user_message}]}],
inferenceConfig={"maxTokens": max_tokens, "temperature": 0},
)
@@ -536,6 +552,8 @@ def extract_files_direct(
api_key: str | None = None,
model: str | None = None,
root: Path = Path("."),
*,
deep_mode: bool = False,
) -> dict:
"""Extract semantic nodes/edges from a list of files using the given backend.
@@ -570,11 +588,11 @@ def extract_files_direct(
max_out = _resolve_max_tokens(cfg.get("max_tokens", 8192))
if backend == "claude":
return _call_claude(key, mdl, user_msg, max_tokens=max_out)
return _call_claude(key, mdl, user_msg, max_tokens=max_out, deep_mode=deep_mode)
if backend == "claude-cli":
return _call_claude_cli(user_msg, max_tokens=max_out)
return _call_claude_cli(user_msg, max_tokens=max_out, deep_mode=deep_mode)
if backend == "bedrock":
return _call_bedrock(mdl, user_msg, max_tokens=max_out)
return _call_bedrock(mdl, user_msg, max_tokens=max_out, deep_mode=deep_mode)
return _call_openai_compat(
cfg["base_url"],
key,
@@ -584,6 +602,7 @@ def extract_files_direct(
reasoning_effort=cfg.get("reasoning_effort"),
max_completion_tokens=_resolve_max_tokens(cfg.get("max_completion_tokens", 8192)),
backend=backend,
deep_mode=deep_mode,
)
@@ -688,6 +707,8 @@ def _extract_with_adaptive_retry(
root: Path,
max_depth: int,
_depth: int = 0,
*,
deep_mode: bool = False,
) -> dict:
"""Extract a chunk; if the response is truncated (`finish_reason="length"`)
or the API rejects the prompt as too large for the model's context window,
@@ -722,7 +743,7 @@ def _extract_with_adaptive_retry(
"""
try:
result = extract_files_direct(
chunk, backend=backend, api_key=api_key, model=model, root=root
chunk, backend=backend, api_key=api_key, model=model, root=root, deep_mode=deep_mode
)
except Exception as exc: # noqa: BLE001 — re-raise unless it's a known context overflow
if not _looks_like_context_exceeded(exc):
@@ -748,10 +769,10 @@ def _extract_with_adaptive_retry(
)
mid = len(chunk) // 2
left = _extract_with_adaptive_retry(
chunk[:mid], backend, api_key, model, root, max_depth, _depth + 1
chunk[:mid], backend, api_key, model, root, max_depth, _depth + 1, deep_mode=deep_mode
)
right = _extract_with_adaptive_retry(
chunk[mid:], backend, api_key, model, root, max_depth, _depth + 1
chunk[mid:], backend, api_key, model, root, max_depth, _depth + 1, deep_mode=deep_mode
)
return {
"nodes": left.get("nodes", []) + right.get("nodes", []),
@@ -790,10 +811,10 @@ def _extract_with_adaptive_retry(
)
mid = len(chunk) // 2
left = _extract_with_adaptive_retry(
chunk[:mid], backend, api_key, model, root, max_depth, _depth + 1
chunk[:mid], backend, api_key, model, root, max_depth, _depth + 1, deep_mode=deep_mode
)
right = _extract_with_adaptive_retry(
chunk[mid:], backend, api_key, model, root, max_depth, _depth + 1
chunk[mid:], backend, api_key, model, root, max_depth, _depth + 1, deep_mode=deep_mode
)
return {
@@ -821,6 +842,7 @@ def extract_corpus_parallel(
token_budget: int | None = 60_000,
max_concurrency: int = 4,
max_retry_depth: int = 3,
deep_mode: bool = False,
) -> dict:
"""Extract a corpus in chunks, merging results.
@@ -878,6 +900,7 @@ def extract_corpus_parallel(
model=model,
root=root,
max_depth=max_retry_depth,
deep_mode=deep_mode,
)
result["elapsed_seconds"] = round(time.time() - t0, 2)
return idx, result, None
+1
View File
@@ -40,6 +40,7 @@ dependencies = [
"tree-sitter-fortran",
"tree-sitter-bash",
"tree-sitter-json",
"tree-sitter-dm",
]
[project.urls]
+35
View File
@@ -0,0 +1,35 @@
#include "helpers.dm"
var/global_counter = 0
/proc/log_event(msg)
world.log << msg
global_counter++
/datum/weapon
var/damage = 10
var/name = "weapon"
proc/attack(mob/target)
log_event("attack")
target.take_damage(damage)
return damage
New()
log_event("weapon created")
/datum/weapon/sword
damage = 20
name = "sword"
/datum/weapon/sword/proc/sharpen()
damage += 5
log_event("sharpened")
/datum/weapon/sword/attack(mob/target)
sharpen()
return ..()
/proc/RunTest()
var/datum/weapon/sword/s = new /datum/weapon/sword()
s.attack(null)
+57
View File
@@ -0,0 +1,57 @@
window "mapwindow"
elem "mapwindow"
type = MAIN
pos = 0,0
size = 640x480
is-pane = true
elem "map"
type = MAP
pos = 0,0
size = 640x480
anchor1 = 0,0
anchor2 = 100,100
is-default = true
window "infowindow"
elem "infowindow"
type = MAIN
pos = 0,0
size = 640x480
is-pane = true
elem "info"
type = CHILD
pos = 0,30
size = 640x445
anchor1 = 0,0
anchor2 = 100,100
left = "statwindow"
right = "outputwindow"
is-vert = false
window "outputwindow"
elem "outputwindow"
type = MAIN
pos = 0,0
size = 640x480
is-pane = true
elem "output"
type = OUTPUT
pos = 0,0
size = 640x480
anchor1 = 0,0
anchor2 = 100,100
is-default = true
window "statwindow"
elem "statwindow"
type = MAIN
pos = 0,0
size = 640x480
is-pane = true
elem "stat"
type = INFO
pos = 0,0
size = 640x480
anchor1 = 0,0
anchor2 = 100,100
is-default = true
BIN
View File
Binary file not shown.

After

Width:  |  Height:  |  Size: 227 B

+14
View File
@@ -0,0 +1,14 @@
"a" = (/turf/closed/wall)
"b" = (/obj/structure/table,/area/station/maintenance)
"c" = (/obj/item/weapon/sword{name = "longsword"; damage = 25},/turf/open/floor/plating)
"d" = (
/obj/structure/table,
/obj/item/weapon/sword,
/turf/open/floor/plating,
/area/station/maintenance)
(1,1,1) = {"
aabb
acdb
aaaa
"}
+163
View File
@@ -7,6 +7,7 @@ from graphify.extract import (
extract_csharp, extract_kotlin, extract_scala, extract_php,
extract_swift, extract_go, extract_julia, extract_js, extract_fortran,
extract_groovy, extract_sln, extract_csproj, extract_razor,
extract_dm, extract_dmi, extract_dmm, extract_dmf,
)
FIXTURES = Path(__file__).parent / "fixtures"
@@ -1060,6 +1061,168 @@ def test_groovy_spock_no_dangling_edges():
assert e["source"] in node_ids
# ── DM (BYOND DreamMaker) ────────────────────────────────────────────────────
def test_dm_no_error():
r = extract_dm(FIXTURES / "sample.dm")
assert "error" not in r
def test_dm_finds_global_proc():
r = extract_dm(FIXTURES / "sample.dm")
labels = _labels(r)
assert any(l == "log_event()" for l in labels)
assert any(l == "RunTest()" for l in labels)
def test_dm_finds_type_definition():
r = extract_dm(FIXTURES / "sample.dm")
labels = _labels(r)
assert "/datum/weapon" in labels
assert "/datum/weapon/sword" in labels
def test_dm_qualifies_proc_with_type_path():
r = extract_dm(FIXTURES / "sample.dm")
labels = _labels(r)
assert "/datum/weapon/attack()" in labels
assert "/datum/weapon/sword/attack()" in labels
def test_dm_finds_path_form_proc_definition():
r = extract_dm(FIXTURES / "sample.dm")
assert "/datum/weapon/sword/sharpen()" in _labels(r)
def test_dm_emits_include_edge():
r = extract_dm(FIXTURES / "sample.dm")
import_edges = _edges_with_relation(r, "imports", "imports_from")
assert import_edges
assert all(e.get("context") == "import" for e in import_edges)
def test_dm_unresolved_include_flagged_external():
r = extract_dm(FIXTURES / "sample.dm")
import_edges = _edges_with_relation(r, "imports", "imports_from")
helpers = [e for e in import_edges if "helpers" in e["target"]]
assert helpers
assert all(e.get("external") is True for e in helpers)
def test_dm_resolves_in_file_calls():
r = extract_dm(FIXTURES / "sample.dm")
calls = _calls(r)
assert any(callee == "log_event()" for _, callee in calls)
assert ("/datum/weapon/sword/attack()", "/datum/weapon/sword/sharpen()") in calls
def test_dm_ambiguous_member_call_left_unresolved():
r = extract_dm(FIXTURES / "sample.dm")
calls = _calls(r)
runtest_to_attack = [c for s, c in calls
if s == "RunTest()" and "attack" in c]
assert not runtest_to_attack
assert any(rc["callee"] == "attack" for rc in r.get("raw_calls", []))
def test_dm_emits_new_as_instantiates():
r = extract_dm(FIXTURES / "sample.dm")
node_by_id = {n["id"]: n["label"] for n in r["nodes"]}
inst = [(node_by_id.get(e["source"]), node_by_id.get(e["target"]))
for e in r["edges"] if e["relation"] == "instantiates"]
assert ("RunTest()", "/datum/weapon/sword") in inst
def test_dm_call_edges_have_call_context():
r = extract_dm(FIXTURES / "sample.dm")
call_edges = _edges_with_relation(r, "calls", "instantiates")
assert call_edges
assert all(e.get("context") == "call" for e in call_edges)
def test_dm_no_dangling_edges():
r = extract_dm(FIXTURES / "sample.dm")
node_ids = {n["id"] for n in r["nodes"]}
for e in r["edges"]:
assert e["source"] in node_ids
def test_dm_super_call_not_emitted():
r = extract_dm(FIXTURES / "sample.dm")
calls = _calls(r)
assert not any(callee.strip("()") == ".." for _, callee in calls)
assert not any(rc["callee"] == ".." for rc in r.get("raw_calls", []))
# ── DMI (BYOND icon sheets) ──────────────────────────────────────────────────
def test_dmi_no_error():
r = extract_dmi(FIXTURES / "sample.dmi")
assert "error" not in r
def test_dmi_emits_state_nodes():
r = extract_dmi(FIXTURES / "sample.dmi")
labels = _labels(r)
assert any(l == '"mob"' for l in labels)
def test_dmi_state_contained_by_file():
r = extract_dmi(FIXTURES / "sample.dmi")
node_by_id = {n["id"]: n["label"] for n in r["nodes"]}
contains = [(node_by_id.get(e["source"]), node_by_id.get(e["target"]))
for e in r["edges"] if e["relation"] == "contains"]
assert ("sample.dmi", '"mob"') in contains
# ── DMM (BYOND map files) ────────────────────────────────────────────────────
def test_dmm_no_error():
r = extract_dmm(FIXTURES / "sample.dmm")
assert "error" not in r
def test_dmm_extracts_type_paths_as_uses_edges():
r = extract_dmm(FIXTURES / "sample.dmm")
targets = {e["target"] for e in r["edges"] if e["relation"] == "uses"}
assert "turf_closed_wall" in targets
assert "obj_structure_table" in targets
assert "obj_item_weapon_sword" in targets
def test_dmm_strips_var_overrides():
r = extract_dmm(FIXTURES / "sample.dmm")
targets = {e["target"] for e in r["edges"] if e["relation"] == "uses"}
assert not any("{" in t for t in targets)
assert "obj_item_weapon_sword" in targets
def test_dmm_handles_multiline_tile_definition():
r = extract_dmm(FIXTURES / "sample.dmm")
targets = {e["target"] for e in r["edges"] if e["relation"] == "uses"}
assert "area_station_maintenance" in targets
def test_dmm_skips_grid_section():
r = extract_dmm(FIXTURES / "sample.dmm")
targets = {e["target"] for e in r["edges"] if e["relation"] == "uses"}
assert len(targets) == 5
# ── DMF (BYOND interface forms) ──────────────────────────────────────────────
def test_dmf_no_error():
r = extract_dmf(FIXTURES / "sample.dmf")
assert "error" not in r
def test_dmf_extracts_windows():
r = extract_dmf(FIXTURES / "sample.dmf")
labels = _labels(r)
assert 'window "mapwindow"' in labels
assert 'window "infowindow"' in labels
def test_dmf_elem_labels_carry_control_type():
r = extract_dmf(FIXTURES / "sample.dmf")
labels = _labels(r)
assert 'elem "map" [MAP]' in labels
def test_dmf_elem_under_window():
r = extract_dmf(FIXTURES / "sample.dmf")
node_by_id = {n["id"]: n["label"] for n in r["nodes"]}
contains = [(node_by_id.get(e["source"]), node_by_id.get(e["target"]))
for e in r["edges"] if e["relation"] == "contains"]
assert ('window "mapwindow"', 'elem "map" [MAP]') in contains
def test_dmf_no_dangling_edges():
r = extract_dmf(FIXTURES / "sample.dmf")
node_ids = {n["id"] for n in r["nodes"]}
for e in r["edges"]:
assert e["source"] in node_ids
assert e["target"] in node_ids
# -- .NET project files (.sln, .csproj, .razor) -------------------------------
def test_sln_no_error():