From 48c9f024fb1b11cf4c1be2334252e321419373db Mon Sep 17 00:00:00 2001 From: saxster <119735920+saxster@users.noreply.github.com> Date: Sat, 2 May 2026 18:44:55 +0530 Subject: [PATCH] docs(skill): forced-rank confidence scores for INFERRED edges (#546) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Closes #540. Production audit on a 10,129-edge graph showed the INFERRED confidence_score distribution is bimodal, not graded: | Score bucket | Count | % of INFERRED | |--------------|-------|---------------| | <0.4 | 0 | 0% | | 0.4-0.6 | 5,807 | 57% | | 0.6-0.8 | 14 | 0.1% | | 0.8+ | 4,308 | 42% | Subagents collapse the continuous "0.4-0.9" guidance to a binary: 0.5 for "uncertain", 0.85+ for "confident", almost nothing in between. Downstream filtering by confidence is therefore an on/off switch, not the gradient the prompt promises. Replace continuous ranges with a forced-rank discrete set: 0.95 direct structural evidence 0.85 strong inference 0.75 reasonable inference 0.65 weak inference 0.55 speculative but plausible Models follow discrete rubrics far better than continuous ranges (documented in calibration literature; same reason MCQ rubrics outperform 0-100 scales). The set is anchored at non-round midpoints to discourage 0.5 as a default. Applied uniformly across all 10 skill-*.md files: - 7 long-form (skill.md, skill-codex.md, skill-copilot.md, skill-droid.md, skill-opencode.md, skill-windows.md, skill-trae.md): full forced-rank table. - 3 short-form (skill-claw.md, skill-aider.md, skill-kiro.md): inline set notation INFERRED ∈ {0.55, 0.65, 0.75, 0.85, 0.95}. Pure prompt edit — no code changes, no test impact. Effect is observable only via re-extraction and inspection of the new confidence_score distribution. --- graphify/skill-aider.md | 2 +- graphify/skill-claw.md | 2 +- graphify/skill-codex.md | 14 ++++++++++---- graphify/skill-copilot.md | 14 ++++++++++---- graphify/skill-droid.md | 14 ++++++++++---- graphify/skill-kiro.md | 2 +- graphify/skill-opencode.md | 14 ++++++++++---- graphify/skill-trae.md | 14 ++++++++++---- graphify/skill-windows.md | 14 ++++++++++---- graphify/skill.md | 14 ++++++++++---- 10 files changed, 73 insertions(+), 31 deletions(-) diff --git a/graphify/skill-aider.md b/graphify/skill-aider.md index fa7007f..dc7146f 100644 --- a/graphify/skill-aider.md +++ b/graphify/skill-aider.md @@ -240,7 +240,7 @@ Process each file one at a time. For each file: - DEEP_MODE (if --mode deep): be aggressive with INFERRED edges - Semantic similarity: if two concepts solve the same problem without a structural link, add `semantically_similar_to` INFERRED edge (confidence 0.6-0.95). Non-obvious cross-file links only. - Hyperedges: if 3+ nodes share a concept/flow not captured by pairwise edges, add a hyperedge. Max 3 per file. - - confidence_score REQUIRED on every edge: EXTRACTED=1.0, INFERRED=0.6-0.9 (reason individually), AMBIGUOUS=0.1-0.3 + - confidence_score REQUIRED on every edge: EXTRACTED=1.0; INFERRED ∈ {0.55, 0.65, 0.75, 0.85, 0.95} forced-rank (NEVER 0.5 — pick the closest discrete value or mark AMBIGUOUS); AMBIGUOUS=0.1-0.3 3. Accumulate results across all files Schema for each file's output: diff --git a/graphify/skill-claw.md b/graphify/skill-claw.md index 3d84acb..64cf4de 100644 --- a/graphify/skill-claw.md +++ b/graphify/skill-claw.md @@ -240,7 +240,7 @@ Process each file one at a time. For each file: - DEEP_MODE (if --mode deep): be aggressive with INFERRED edges - Semantic similarity: if two concepts solve the same problem without a structural link, add `semantically_similar_to` INFERRED edge (confidence 0.6-0.95). Non-obvious cross-file links only. - Hyperedges: if 3+ nodes share a concept/flow not captured by pairwise edges, add a hyperedge. Max 3 per file. - - confidence_score REQUIRED on every edge: EXTRACTED=1.0, INFERRED=0.6-0.9 (reason individually), AMBIGUOUS=0.1-0.3 + - confidence_score REQUIRED on every edge: EXTRACTED=1.0; INFERRED ∈ {0.55, 0.65, 0.75, 0.85, 0.95} forced-rank (NEVER 0.5 — pick the closest discrete value or mark AMBIGUOUS); AMBIGUOUS=0.1-0.3 3. Accumulate results across all files Schema for each file's output: diff --git a/graphify/skill-codex.md b/graphify/skill-codex.md index b2e79e2..ab99730 100644 --- a/graphify/skill-codex.md +++ b/graphify/skill-codex.md @@ -293,10 +293,16 @@ If a file has YAML frontmatter (--- ... ---), copy source_url, captured_at, auth confidence_score is REQUIRED on every edge - never omit it, never use 0.5 as a default: - EXTRACTED edges: confidence_score = 1.0 always -- INFERRED edges: reason about each edge individually. - Direct structural evidence (shared data structure, clear dependency): 0.8-0.9. - Reasonable inference with some uncertainty: 0.6-0.7. - Weak or speculative: 0.4-0.5. Most edges should be 0.6-0.9, not 0.5. +- INFERRED edges: pick exactly ONE value from this set — never 0.5: + 0.95 direct structural evidence (shared data structure, named cross-file reference). + 0.85 strong inference (clear functional alignment, no direct symbol link). + 0.75 reasonable inference (shared problem domain + similar shape, requires interpretation). + 0.65 weak inference (thematically related, no shape evidence). + 0.55 speculative but plausible (surface-level co-occurrence only). + Models follow discrete rubrics better than continuous ranges; the bimodal + distribution observed in production (>50% at 0.5, >40% at 0.85+) shows the + range guidance is being collapsed to a binary. If no value above fits, mark + the edge AMBIGUOUS rather than picking 0.4 or below. - AMBIGUOUS edges: 0.1-0.3 Output exactly this JSON (no other text): diff --git a/graphify/skill-copilot.md b/graphify/skill-copilot.md index 4682798..117c057 100644 --- a/graphify/skill-copilot.md +++ b/graphify/skill-copilot.md @@ -289,10 +289,16 @@ If a file has YAML frontmatter (--- ... ---), copy source_url, captured_at, auth confidence_score is REQUIRED on every edge - never omit it, never use 0.5 as a default: - EXTRACTED edges: confidence_score = 1.0 always -- INFERRED edges: reason about each edge individually. - Direct structural evidence (shared data structure, clear dependency): 0.8-0.9. - Reasonable inference with some uncertainty: 0.6-0.7. - Weak or speculative: 0.4-0.5. Most edges should be 0.6-0.9, not 0.5. +- INFERRED edges: pick exactly ONE value from this set — never 0.5: + 0.95 direct structural evidence (shared data structure, named cross-file reference). + 0.85 strong inference (clear functional alignment, no direct symbol link). + 0.75 reasonable inference (shared problem domain + similar shape, requires interpretation). + 0.65 weak inference (thematically related, no shape evidence). + 0.55 speculative but plausible (surface-level co-occurrence only). + Models follow discrete rubrics better than continuous ranges; the bimodal + distribution observed in production (>50% at 0.5, >40% at 0.85+) shows the + range guidance is being collapsed to a binary. If no value above fits, mark + the edge AMBIGUOUS rather than picking 0.4 or below. - AMBIGUOUS edges: 0.1-0.3 Output exactly this JSON (no other text): diff --git a/graphify/skill-droid.md b/graphify/skill-droid.md index 6cde935..0613ca5 100644 --- a/graphify/skill-droid.md +++ b/graphify/skill-droid.md @@ -290,10 +290,16 @@ If a file has YAML frontmatter (--- ... ---), copy source_url, captured_at, auth confidence_score is REQUIRED on every edge - never omit it, never use 0.5 as a default: - EXTRACTED edges: confidence_score = 1.0 always -- INFERRED edges: reason about each edge individually. - Direct structural evidence (shared data structure, clear dependency): 0.8-0.9. - Reasonable inference with some uncertainty: 0.6-0.7. - Weak or speculative: 0.4-0.5. Most edges should be 0.6-0.9, not 0.5. +- INFERRED edges: pick exactly ONE value from this set — never 0.5: + 0.95 direct structural evidence (shared data structure, named cross-file reference). + 0.85 strong inference (clear functional alignment, no direct symbol link). + 0.75 reasonable inference (shared problem domain + similar shape, requires interpretation). + 0.65 weak inference (thematically related, no shape evidence). + 0.55 speculative but plausible (surface-level co-occurrence only). + Models follow discrete rubrics better than continuous ranges; the bimodal + distribution observed in production (>50% at 0.5, >40% at 0.85+) shows the + range guidance is being collapsed to a binary. If no value above fits, mark + the edge AMBIGUOUS rather than picking 0.4 or below. - AMBIGUOUS edges: 0.1-0.3 Output exactly this JSON (no other text): diff --git a/graphify/skill-kiro.md b/graphify/skill-kiro.md index b3db443..6804f22 100644 --- a/graphify/skill-kiro.md +++ b/graphify/skill-kiro.md @@ -239,7 +239,7 @@ Process each file one at a time. For each file: - DEEP_MODE (if --mode deep): be aggressive with INFERRED edges - Semantic similarity: if two concepts solve the same problem without a structural link, add `semantically_similar_to` INFERRED edge (confidence 0.6-0.95). Non-obvious cross-file links only. - Hyperedges: if 3+ nodes share a concept/flow not captured by pairwise edges, add a hyperedge. Max 3 per file. - - confidence_score REQUIRED on every edge: EXTRACTED=1.0, INFERRED=0.6-0.9 (reason individually), AMBIGUOUS=0.1-0.3 + - confidence_score REQUIRED on every edge: EXTRACTED=1.0; INFERRED ∈ {0.55, 0.65, 0.75, 0.85, 0.95} forced-rank (NEVER 0.5 — pick the closest discrete value or mark AMBIGUOUS); AMBIGUOUS=0.1-0.3 3. Accumulate results across all files Schema for each file's output: diff --git a/graphify/skill-opencode.md b/graphify/skill-opencode.md index 32819c8..4117b2c 100644 --- a/graphify/skill-opencode.md +++ b/graphify/skill-opencode.md @@ -291,10 +291,16 @@ If a file has YAML frontmatter (--- ... ---), copy source_url, captured_at, auth confidence_score is REQUIRED on every edge - never omit it, never use 0.5 as a default: - EXTRACTED edges: confidence_score = 1.0 always -- INFERRED edges: reason about each edge individually. - Direct structural evidence (shared data structure, clear dependency): 0.8-0.9. - Reasonable inference with some uncertainty: 0.6-0.7. - Weak or speculative: 0.4-0.5. Most edges should be 0.6-0.9, not 0.5. +- INFERRED edges: pick exactly ONE value from this set — never 0.5: + 0.95 direct structural evidence (shared data structure, named cross-file reference). + 0.85 strong inference (clear functional alignment, no direct symbol link). + 0.75 reasonable inference (shared problem domain + similar shape, requires interpretation). + 0.65 weak inference (thematically related, no shape evidence). + 0.55 speculative but plausible (surface-level co-occurrence only). + Models follow discrete rubrics better than continuous ranges; the bimodal + distribution observed in production (>50% at 0.5, >40% at 0.85+) shows the + range guidance is being collapsed to a binary. If no value above fits, mark + the edge AMBIGUOUS rather than picking 0.4 or below. - AMBIGUOUS edges: 0.1-0.3 Output exactly this JSON (no other text): diff --git a/graphify/skill-trae.md b/graphify/skill-trae.md index 2b5b401..3ffaf35 100644 --- a/graphify/skill-trae.md +++ b/graphify/skill-trae.md @@ -280,10 +280,16 @@ If a file has YAML frontmatter (--- ... ---), copy source_url, captured_at, auth confidence_score is REQUIRED on every edge - never omit it, never use 0.5 as a default: - EXTRACTED edges: confidence_score = 1.0 always -- INFERRED edges: reason about each edge individually. - Direct structural evidence (shared data structure, clear dependency): 0.8-0.9. - Reasonable inference with some uncertainty: 0.6-0.7. - Weak or speculative: 0.4-0.5. Most edges should be 0.6-0.9, not 0.5. +- INFERRED edges: pick exactly ONE value from this set — never 0.5: + 0.95 direct structural evidence (shared data structure, named cross-file reference). + 0.85 strong inference (clear functional alignment, no direct symbol link). + 0.75 reasonable inference (shared problem domain + similar shape, requires interpretation). + 0.65 weak inference (thematically related, no shape evidence). + 0.55 speculative but plausible (surface-level co-occurrence only). + Models follow discrete rubrics better than continuous ranges; the bimodal + distribution observed in production (>50% at 0.5, >40% at 0.85+) shows the + range guidance is being collapsed to a binary. If no value above fits, mark + the edge AMBIGUOUS rather than picking 0.4 or below. - AMBIGUOUS edges: 0.1-0.3 Output exactly this JSON (no other text): diff --git a/graphify/skill-windows.md b/graphify/skill-windows.md index 9984022..5d6fc72 100644 --- a/graphify/skill-windows.md +++ b/graphify/skill-windows.md @@ -279,10 +279,16 @@ If a file has YAML frontmatter (--- ... ---), copy source_url, captured_at, auth confidence_score is REQUIRED on every edge - never omit it, never use 0.5 as a default: - EXTRACTED edges: confidence_score = 1.0 always -- INFERRED edges: reason about each edge individually. - Direct structural evidence (shared data structure, clear dependency): 0.8-0.9. - Reasonable inference with some uncertainty: 0.6-0.7. - Weak or speculative: 0.4-0.5. Most edges should be 0.6-0.9, not 0.5. +- INFERRED edges: pick exactly ONE value from this set — never 0.5: + 0.95 direct structural evidence (shared data structure, named cross-file reference). + 0.85 strong inference (clear functional alignment, no direct symbol link). + 0.75 reasonable inference (shared problem domain + similar shape, requires interpretation). + 0.65 weak inference (thematically related, no shape evidence). + 0.55 speculative but plausible (surface-level co-occurrence only). + Models follow discrete rubrics better than continuous ranges; the bimodal + distribution observed in production (>50% at 0.5, >40% at 0.85+) shows the + range guidance is being collapsed to a binary. If no value above fits, mark + the edge AMBIGUOUS rather than picking 0.4 or below. - AMBIGUOUS edges: 0.1-0.3 Output exactly this JSON (no other text): diff --git a/graphify/skill.md b/graphify/skill.md index 58cb22c..b7bb44a 100644 --- a/graphify/skill.md +++ b/graphify/skill.md @@ -337,10 +337,16 @@ If a file has YAML frontmatter (--- ... ---), copy source_url, captured_at, auth confidence_score is REQUIRED on every edge - never omit it, never use 0.5 as a default: - EXTRACTED edges: confidence_score = 1.0 always -- INFERRED edges: reason about each edge individually. - Direct structural evidence (shared data structure, clear dependency): 0.8-0.9. - Reasonable inference with some uncertainty: 0.6-0.7. - Weak or speculative: 0.4-0.5. Most edges should be 0.6-0.9, not 0.5. +- INFERRED edges: pick exactly ONE value from this set — never 0.5: + 0.95 direct structural evidence (shared data structure, named cross-file reference). + 0.85 strong inference (clear functional alignment, no direct symbol link). + 0.75 reasonable inference (shared problem domain + similar shape, requires interpretation). + 0.65 weak inference (thematically related, no shape evidence). + 0.55 speculative but plausible (surface-level co-occurrence only). + Models follow discrete rubrics better than continuous ranges; the bimodal + distribution observed in production (>50% at 0.5, >40% at 0.85+) shows the + range guidance is being collapsed to a binary. If no value above fits, mark + the edge AMBIGUOUS rather than picking 0.4 or below. - AMBIGUOUS edges: 0.1-0.3 Node ID format: lowercase, only `[a-z0-9_]`, no dots or slashes. Format: `{stem}_{entity}` where stem is the filename without extension and entity is the symbol name, both normalized (lowercase, non-alphanumeric chars replaced with `_`). Example: `src/auth/session.py` + `ValidateToken` → `session_validatetoken`. This must match the ID the AST extractor generates so cross-references between code and semantic nodes connect correctly. CRITICAL: never append chunk numbers, sequence numbers, or any suffix to an ID (no `_c1`, `_c2`, `_chunk2`, etc.). IDs must be deterministic from the label alone — the same entity must always produce the same ID regardless of which chunk processes it.