From 154919bb64d754e2ab330958047d3781ad10a5c1 Mon Sep 17 00:00:00 2001 From: Safi Date: Sun, 5 Apr 2026 19:44:34 +0100 Subject: [PATCH] =?UTF-8?q?correct=20benchmark=20numbers=20=E2=80=94=20tok?= =?UTF-8?q?en=20reduction=20scales=20with=20corpus=20size,=20small=20corpo?= =?UTF-8?q?ra=20~1x?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- README.md | 12 ++++++------ worked/httpx/README.md | 12 +++++++----- 2 files changed, 13 insertions(+), 11 deletions(-) diff --git a/README.md b/README.md index 978cdf8..6254980 100644 --- a/README.md +++ b/README.md @@ -102,13 +102,13 @@ Every edge is tagged `EXTRACTED`, `INFERRED`, or `AMBIGUOUS` - you always know w ## Worked examples -| Corpus | Type | Reduction | Eval | -|--------|------|-----------|------| -| Karpathy repos + 5 papers + 4 images | Mixed | **71.5x** | [`worked/karpathy-repos/review.md`](worked/karpathy-repos/review.md) | -| httpx (Python HTTP client) | Code | small corpus¹ | [`worked/httpx/review.md`](worked/httpx/review.md) | -| Code + paper + Arabic image | Multi-type | small corpus¹ | [`worked/mixed-corpus/review.md`](worked/mixed-corpus/review.md) | +| Corpus | Files | Reduction | Output | +|--------|-------|-----------|--------| +| Karpathy repos + 5 papers + 4 images | 52 | **71.5x** | [`worked/karpathy-repos/`](worked/karpathy-repos/) | +| graphify source + Transformer paper | 4 | **5.4x** | [`worked/mixed-corpus/`](worked/mixed-corpus/) | +| httpx (synthetic Python library) | 6 | ~1x | [`worked/httpx/`](worked/httpx/) | -¹ Small corpora fit in one context window - graph value is structural clarity, not compression. +Token reduction scales with corpus size. 6 files fits in a context window anyway — graph value there is structural clarity, not compression. At 52 files (code + papers + images) you get 71x+. Each `worked/` folder has the raw input files and the actual output (`GRAPH_REPORT.md`, `graph.json`) so you can run it yourself and verify the numbers. ## Tech stack diff --git a/worked/httpx/README.md b/worked/httpx/README.md index 3d1c924..84fa706 100644 --- a/worked/httpx/README.md +++ b/worked/httpx/README.md @@ -33,10 +33,12 @@ graphify ./raw ## What to expect -- ~95 nodes, ~130 edges -- 4 communities: Exception Hierarchy, Models & Data, Auth & Transport, Client Layer -- God nodes: `client.py`, `models.py`, `transport.py`, `exceptions.py`, `BaseClient`, `Response` +- 144 nodes, 330 edges, 6 communities +- God nodes: `Client`, `AsyncClient`, `Response`, `Request`, `BaseClient`, `HTTPTransport` - Surprising connections: `DigestAuth` ↔ `Response` (auth.py reads Response to parse WWW-Authenticate) -- All edges EXTRACTED — no inference needed, dependency graph is explicit +- **~1x token reduction** — 6 files fits in a context window, so there's no compression win here -Full eval with scores and analysis: `review.md` +The graph value on a small corpus is structural, not compressive: you can see the full dependency graph, identify god nodes, and understand architecture at a glance. For token reduction to matter you need 20+ files. At 52 files (Karpathy repos benchmark) graphify achieves 71.5x. + +Run `graphify benchmark worked/httpx/graph.json` to verify the numbers yourself. +Actual output is already in this folder: `GRAPH_REPORT.md` (human-readable) and `graph.json` (full graph data).