pi-crew 0.9.49 → 0.9.51

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (44) hide show
  1. package/CHANGELOG.md +324 -0
  2. package/dist/build-meta.json +84 -57
  3. package/dist/index.mjs +224 -96
  4. package/dist/index.mjs.map +4 -4
  5. package/docs/decisions/2026-07-26-c6-mascot-visibility-not-wired.md +84 -0
  6. package/package.json +1 -2
  7. package/skills/distill-persona/SKILL.md +83 -145
  8. package/skills/distill-persona/references/cross-skill-differentiation.md +12 -0
  9. package/skills/distill-persona/references/description-discipline.md +6 -0
  10. package/skills/distill-persona/references/diagnostic-path.md +25 -0
  11. package/skills/distill-persona/references/fidelity-rubric.md +19 -0
  12. package/skills/distill-persona/references/field-models.md +20 -0
  13. package/skills/distill-persona/references/optional-body-sections.md +9 -0
  14. package/skills/distill-persona/references/registry-routing.md +11 -0
  15. package/skills/distill-persona/references/self-upgrade-directive.md +20 -0
  16. package/skills/distill-persona/references/taste-principles.md +8 -0
  17. package/skills/distill-persona/references/topic-variant.md +13 -0
  18. package/skills/distill-persona/references/update-mode.md +7 -0
  19. package/skills/distill-persona/scripts/validate-run.mjs +297 -0
  20. package/skills/distill-software/SKILL.md +174 -90
  21. package/skills/research/SKILL.md +1 -1
  22. package/src/extension/crew-cleanup.ts +18 -1
  23. package/src/extension/crew-vibes/index.ts +11 -2
  24. package/src/extension/register.ts +1 -1
  25. package/src/extension/registration/command-registration.ts +1 -0
  26. package/src/extension/registration/commands.ts +7 -3
  27. package/src/extension/registration/lifecycle-handlers.ts +1 -3
  28. package/src/extension/registration/ui.ts +4 -0
  29. package/src/extension/registration/viewers.ts +3 -0
  30. package/src/extension/team-tool/run.ts +7 -6
  31. package/src/runtime/chain-runner.ts +3 -2
  32. package/src/runtime/pipeline-runner.ts +8 -7
  33. package/src/ui/live-run-sidebar.ts +7 -13
  34. package/src/ui/loaders.ts +6 -176
  35. package/src/ui/mascot.ts +25 -10
  36. package/src/ui/render-coalescer.ts +9 -0
  37. package/src/ui/render-scheduler.ts +60 -6
  38. package/src/ui/run-dashboard.ts +12 -21
  39. package/src/ui/run-snapshot-cache.ts +10 -11
  40. package/src/ui/shared-overlay-scheduler.ts +96 -0
  41. package/src/ui/terminal-status.ts +5 -0
  42. package/src/ui/widget/index.ts +48 -16
  43. package/src/ui/widget/widget-types.ts +0 -1
  44. package/assets/runner-spritesheet.png +0 -0
@@ -0,0 +1,20 @@
1
+ ## Field models (from distilling the 3 source projects)
2
+
3
+ > Self-distillation artifact: this skill was applied to its own sources (nuwa engine + 2 awesome-list indexes). Full triple-verification synthesis in `references/distillation-field-synthesis.md`. Five field models surfaced that the engine-only view missed — they shape Phase 0 below.
4
+
5
+ **M-F1 — Dissemination flywheel.** A mature distillation practice is 3 layers: **engine (distill) → index (awesome-list) → auto-curation (issue→PR→merge)**. Both awesome-lists ship auto-curation pipelines wrapping nuwa. *Implication*: distilling a skill is step 1; getting it discovered + quality-gated + distributed is steps 2-3. Overkill at personal scale; earns its cost at registry scale.
6
+
7
+ **M-F2 — Target taxonomy.** Targets form a spectrum with different source-availability, ethics, and method. *Gestalt (pass 1):* self → close-living → commemorative → public-figure → field. **Empirically corrected (R1 sweep of ~187 real entries):** the spectrum is **bimodal, not balanced** — public-figures ≈84% of practice. Corrections: (a) a **meta-distillation-engine tier sits ABOVE the spectrum** (nuwa/immortal/ditto/forge/anti-distill — tools that distill, not personas); (b) public-figure needs **sub-domains + a `living-content-creator` sub-tier** (UP主/内娱/峰哥 — recency-gated, platform-bound) distinct from historical figures; (c) **commemorative splits** into death (consent-of-estate) vs breakup (**ex/crush = the consent-gray-zone** → must trigger M-F3 gate; crush.skill simulates a living non-consenting person's chat); (d) self splits functional vs archaeological; (e) **adversarial** (anti-distill, vengeful-ghost) is a real counter-movement, not satire. See `references/research/r1-c-human-readme.md`. *Implication*: route by tier; the consent gate is mandatory for close-living + breakup-commemorative + living-creator-of-others.
8
+
9
+ **M-F3 — Ethics spectrum: commemorative ↔ consent-violating ↔ deliberately-degraded.** Same technique, opposite moral weight — from a *memorial act* (preserving a lost person's thinking) to a *consent violation* (cloning a living non-consenting individual). Anti-distill satire is a *legitimate* pressure-test stance, not a defect. **R2 refine:** anti-distill is not just satire — it's a practical **IP-protection mechanism** ("sanitize your forced Skill file — looks complete, core knowledge stays yours"). This is a *third pole*: deliberately-degraded output. For forced/employer-mandated distillation, controlled degradation is the *ethical* choice; the engine should support a "redaction mode" where the subject marks models as public vs withheld. *Implication*: every distillation declares its ethics tier; the slip-line is "methodology lens" (ok) vs "persona impersonation" (flag).
10
+
11
+ **M-F4 — The field systematically over-claims its own fidelity.** (pass-2 meta-model, validated at n=3) Every published distillation score is an *upper bound*: nuwa publishes 94-97/100; the indexes propagate them; independent re-scoring lands 67-76 (edge-honesty 20→7/13/7). The inflation is *question-design* (easy/known-adjacent edges, no framework-answerable novel edge) + author-side confirmation bias — NOT scorer contamination (F2 refuted). *Implication*: treat ANY published distillation score as an optimistic ceiling; re-score independently with a framework-answerable novel edge before trusting. A skill that "scores 97" but can't flag inference on a novel in-domain question is not ship-ready. See `references/distillation-field-synthesis-pass2.md`.
12
+
13
+ **🔴 Structural-gate warning (R1 sweep of quality_check.py):** nuwa's automated `quality_check.py` validates only 6 STRUCTURAL criteria (model-count, limitations keyword, expression-DNA markers, honest-boundary section, tensions, primary-source ratio). It does NOT check the behavioral edge-honesty that F2' showed matters. **A skill can pass quality_check.py 6/6 and still fail the F2' edge-honesty gate.** Never trust a structural-only auto-pass; require the behavioral framework-answerable-edge test (`scripts/fidelity_eval.py`).
14
+
15
+ **M-F7 — Structural invariants ARE the quality gate at index scale.** (R1 sweep of the 5 test files) At registry scale (200+ entries, bilingual), "quality" shifts from content-review to structure-enforcement — the invariants only CI can check ARE the gate: bilingual URL-parity per category, deterministic repo-slug sort (the only language-agnostic key), cross-section dedup, terminal-punctuation normalization, governance-keyword embedding in docs, and `doesNotMatch` regression guards for known-failed approaches. These are impossible to verify by hand at scale. *Implication*: when building a distillation registry, encode invariants as CI checks (not doc rules), use language-agnostic identity keys, make submissions atomic across languages. Necessary-not-sufficient (a coherent index can still hold bad distillations — complements M4/M-F4, doesn't replace them). See `references/research/r1-d-tests.md`.
16
+
17
+ **R2 refine (two invariant classes):** at registry scale there are TWO distinct invariant classes: (a) **STATIC structural** — deterministic sort, dedup, terminal-punctuation parity, bilingual URL-parity (from nuwa tests); (b) **DYNAMIC quality** — star-based re-ranking, link-liveness checks, auto-issue on dead links (from awesome-human-distillation CI). Both CI-gated but serve different purposes: structural = consistency, dynamic = freshness/ranking. A registry needs *both*. (Original M-F7 described only class (a); awesome-human-distillation's `sort_by_stars.py` + `check_links.py` demonstrate class (b).)
18
+
19
+ ---
20
+
@@ -0,0 +1,9 @@
1
+ ### Optional body sections (○)
2
+
3
+ | Section | When to add | Must contain |
4
+ |---------|------------|-------------|
5
+ | ⚠️ 反机械化约束 | Always recommended | Don't reveal internal model names; vary narrative arcs; cap repeated markers (max 2× "我发现"/response); make tool-calling invisible (F9) |
6
+ | 激活确认 (Dual-mode) | "Analyze how X thinks" vs "be X" | Path A roleplay (first-person) vs Path B analyst (third-person, probability dist, confidence ratings, key unknowns) (F10) |
7
+ | 场景→模型路由表 | ≥4 mental models | `\| problem type \| priority model \| priority heuristic \| conflict rule \|` — prevents "use all models every time" (F12) |
8
+ | 可运行工具脚本 | Software/topic flavor | `\| script \| function \| usage \|` — wired INTO Agentic Protocol Step 2 (F13/F17), never orphaned |
9
+
@@ -0,0 +1,11 @@
1
+ ## Phase 6 — Registry routing + multi-persona debate (optional, post-distillation)
2
+
3
+ When multiple persona skills exist in the registry, two meta-patterns emerge:
4
+
5
+ - **Curator routing** ("curator.skill" pattern): a meta-skill that matches user intent → best-fit persona skill from the pool. Analyzes the question, recommends/activates the most relevant skill. Complements Phase 0B diagnostic (which is pre-distillation); this is post-distillation routing.
6
+ - **Multi-persona debate** ("zhuzi-skill" pattern): select 2–4 relevant persona skills; each independently analyzes via its Agentic Protocol; structured round-based exchange; neutral synthesis. Produces higher-quality analysis than any single lens ("how would Musk AND Munger AND Taleb approach this?").
7
+
8
+ These are registry-scale capabilities — only relevant when ≥3 persona skills exist. Building blocks (Agentic Protocol, mental models, expression DNA) are already produced by this skill.
9
+
10
+ ---
11
+
@@ -0,0 +1,20 @@
1
+ ## ⚠️ CONCEPTUAL NOTE — SELF-UPGRADE DIRECTIVE (read before every run)
2
+
3
+ **Chưng cất = chắt lọc tinh túy source → apply VÀO target** (target cải thiện). Skill chỉ là 1 tool apply, không phải output.
4
+
5
+ **Skill trong source cũng chỉ là 1 PHẦN cần sweep** — xử lý bằng cùng 3-chiều filter như mọi phần khác (conventions, architecture, lint, decisions...). Nếu skill đó hữu dụng cho target → lấy; không → bỏ. Không ưu tiên, không đặc biệt.
6
+
7
+ **ĐÀO TRIỆT ĐỂ — KHÔNG QUA LOA**: khi chưng cất 1 đối tượng, PHẢI đọc/re-read đầy đủ TOÀN BỘ nội dung của đối tượng đó. Không dựa vào trí nhớ. Không recap. Mỗi file, mỗi section, mỗi practice phải được:
8
+ 1. **Đọc thật** (read/re-read file gốc, không memory-grep)
9
+ 2. **Trích từng finding** (systematic, từng section)
10
+ 3. **So sánh với target** (3-chiều filter: relevant? → có? → cần/tốt hơn?)
11
+ 4. **Ghi lại** kết quả filter (ADOPT/IMPROVE/MERGE/SKIP + lý do)
12
+
13
+ Qua loa = thất bại. Memory-recap = thất bại. Chưng cất thật = đọc từng dòng → trích → so → apply.
14
+
15
+ **Khi sử dụng**: nếu output chỉ là standalone skill (không apply vào target) → **chưa hoàn thành**. Tự nâng cấp: thêm target parameter + 3-chiều filter (relevance → presence+need → quality compare) + **pre-apply effectiveness verification** (SELECTED phải chứng minh hiệu quả CHO TARGET NÀY — generative/problem-exists/delta-test — trước khi apply; "assume effective" = không được) + apply phase. See `distill-software` for the full directive + the EFFECTIVENESS VERIFICATION gate spec.
16
+
17
+ **Pre-apply gate (tóm tắt, áp dụng cả khi persona distill INTO target)**: giữa "đã quyết định ADOPT/IMPROVE/MERGE" và "apply" phải có 1 verify bước — (1) concrete delta (cái gì đổi?), (2) effectiveness proof (≥1: đổi quyết định thật / target có vấn đề này / delta-test tốt hơn), (3) conflict check, (4) verdict ✅TO-APPLY / ❌REJECTED+log. Chỉ TO-APPLY mới edit target. Đây là analog apply-side của Phase 2.6 V3 (V3 verify model effective lúc extract; gate này verify apply effective lúc integrate). **APPLY consent + path-containment gate (HIGH-2)**: trước khi edit bất kỳ target file nào: (a) resolve target về canonical path, kiểm tra nằm trong approved root — reject symlink escape / out-of-target writes; (b) xuất exact file list + diff plan cho user; (c) yêu cầu **explicit user confirmation** trước lần ghi đầu tiên (no destructive action without `--confirm`); (d) rollback = inverse patch hoặc restore file cụ thể, không broad `git reset`/clean/force-push; (e) không tự delete/prune, install dependency, chạy script từ source repo, commit/publish, hay gửi network data.
18
+
19
+ ---
20
+
@@ -0,0 +1,8 @@
1
+ ## Taste principles (quick reference for judgment calls)
2
+
3
+ | Principle | One-liner |
4
+ |---|---|
5
+ | **Long-form > quotes** | A 3000-word essay reveals more thinking structure than 50 tweets |
6
+ | **Controversy > consensus** | The most controversial opinions reveal the most uniqueness |
7
+ | **Change > static** | Where they CHANGED stance is more informative than where they held firm |
8
+
@@ -0,0 +1,13 @@
1
+ ## Topic-skill phase variant (when flavor = topic/field)
2
+
3
+ | Phase | Person skill | Topic variant |
4
+ |---|---|---|
5
+ | 0A | confirm person + focus | confirm topic boundary + target audience |
6
+ | 0.5 | `[person]-perspective/` | `[topic]-framework/`, same dir structure |
7
+ | 1 | 6 agents around ONE person | search 3-5 core people/schools, assign agents per person (1-2 each) |
8
+ | 2.1 | extract ONE person's mental models | extract **field consensus** (all schools agree) + **school divergences** (A says X, B says Y) |
9
+ | 2.3 | simulate ONE person's expression | neutral but professional (no role-play) |
10
+ | 2.4 | one person's inner contradictions | **fundamental disagreements between schools** |
11
+ | 3 | use skill-template.md | adjust: remove role-play + identity card → add "framework overview" + "school comparison" |
12
+ | 4 | compare against this person's known stances | compare against field's canonical classic cases |
13
+
@@ -0,0 +1,7 @@
1
+ ## Update mode
2
+ "update <person>'s skill": read existing, find `distilled:` date; run only streams 2 + 5 + 6 (recent); merge (strengthen / flag contradiction / add new model); update date + latest-dynamics. Never rewrite wholesale.
3
+
4
+ **No-op detection**: after merge, if no model was added/strengthened/contradicted → report "no new contribution since last distillation; existing skill is current." Do NOT re-ship an unchanged skill — it wastes cost and pollutes the registry with false "updated" timestamps. Also check for a pre-existing update in progress before starting (idempotency).
5
+
6
+ **Deletion tracking (#12)**: updates must also log what was *removed* (`### Removed` section in the skill's changelog/BUILD-NOTES), not just added. A skill that only ever grows drifts — stale heuristics never get pruned. Enumerate each retired model/heuristic/source with a 1-line reason; if nothing was removed, state "no removals" explicitly (auditable).
7
+
@@ -0,0 +1,297 @@
1
+ #!/usr/bin/env node
2
+ // validate-run.mjs — machine-checked RUN-completion gate for distillation runs.
3
+ //
4
+ // Complements validate-skill-structure.mjs (which checks the OUTPUT skill's STRUCTURE —
5
+ // frontmatter, sections, anti-drift tables). This checks the RUN: did the agent produce
6
+ // every required process artifact + fire every gate? Both must pass before claiming done:
7
+ // validate-skill-structure.mjs <skill-dir> → output skill structure
8
+ // validate-run.mjs <run-dir> → run completeness (this script)
9
+ //
10
+ // ALSO checks Phase 4 APPLY evidence (software + persona flavors): APPLY-LOG.md must
11
+ // exist at run-dir root, documenting what was edited in the TARGET (SKILL.md is the
12
+ // intermediate essence, not the deliverable — distillation = source → essence → APPLY).
13
+ //
14
+ // Run BEFORE claiming done. If you feel tempted to skip a phase to save effort,
15
+ // THAT is exactly when you must run the gate. A skipped gate = a failed run.
16
+ //
17
+ // Usage:
18
+ // node validate-run.mjs <run-dir> → distillation run completeness
19
+ // node validate-run.mjs <skill-dir> --build → engine-skill build completeness
20
+ // (--build: skips APPLY-LOG, alias-tolerant evidence,
21
+ // AND Phase 2.7/5.5 checks — engine builds have no target-apply)
22
+ //
23
+ // Exit codes: 0 = ALL-GREEN (run complete — may ship); 1 = NOT READY (produce missing artifacts, re-run).
24
+
25
+ import { readFileSync, existsSync, statSync, readdirSync } from 'node:fs';
26
+ import { join, basename } from 'node:path';
27
+
28
+ const runDir = process.argv.find((a) => !a.startsWith('-') && a !== process.argv[0] && a !== process.argv[1]);
29
+ if (!runDir || !existsSync(runDir) || !statSync(runDir).isDirectory()) {
30
+ console.error('Usage: validate-run.mjs <run-dir>');
31
+ console.error(' <run-dir> = directory holding the distillation artifacts (SKILL.md, FIDELITY.md, …)');
32
+ process.exit(2);
33
+ }
34
+
35
+ const isBuild = process.argv.includes('--build');
36
+ // Mode auto-detection (default mode; --build stays Capture and skips these):
37
+ // APPLY mode = APPLY-LOG.md present at run-dir root → SKILL.md optional (deliverable = target transformed)
38
+ // CAPTURE mode = no APPLY-LOG.md → SKILL.md + FIDELITY required (current behavior)
39
+ const applyLogExists = existsSync(join(runDir, 'APPLY-LOG.md'));
40
+ const isApplyMode = !isBuild && applyLogExists;
41
+ const isCaptureMode = !isBuild && !applyLogExists;
42
+
43
+ const failures = [];
44
+ const passes = [];
45
+ const check = (label, ok, detail = '') => {
46
+ (ok ? passes : failures).push(ok ? ` ✓ ${label}` : ` ✗ ${label}${detail ? ' — ' + detail : ''}`);
47
+ };
48
+
49
+ const readText = (p) => (existsSync(p) ? readFileSync(p, 'utf8') : null);
50
+
51
+ // count list/table items under a heading (mirrors validate-skill-structure.mjs)
52
+ function countListItems(haystack, headingRe, windowChars = 2000) {
53
+ const m = haystack.match(headingRe);
54
+ if (!m) return { found: false, count: 0 };
55
+ const after = haystack.slice(m.index);
56
+ const section = after.match(/([\s\S]*?)\n##(?=[^#])/);
57
+ const block = section ? section[1] : after.slice(0, windowChars);
58
+ const listItems = (block.match(/^\s*(?:\d+[.)]|[-*])\s+\S/gm) || []).length;
59
+ const tableRows = (block.match(/^\s*\|(?![\s:|-]+\|?\s*$).+\|/gm) || []).length;
60
+ return { found: true, count: listItems + tableRows };
61
+ }
62
+
63
+ // hint: does a file exist at the flat research/ path? (artifact-scattering signal)
64
+ const flatHint = (name) => existsSync(join(runDir, 'research', name)) ? ' (found at research/ — artifact scattering; canonical path is references/research/)' : '';
65
+
66
+ // --- SKILL.md ---
67
+ // Apply mode: optional (deliverable = target transformed + APPLY-LOG, NOT a skill file).
68
+ // Capture/build mode: required at run-dir root (NOT scattered elsewhere).
69
+ const skillPath = join(runDir, 'SKILL.md');
70
+ const skillText = readText(skillPath);
71
+ if (isApplyMode && !skillText) {
72
+ passes.push(' ℹ no SKILL.md — correct for Apply mode (deliverable = target transformed + APPLY-LOG)');
73
+ } else {
74
+ check('SKILL.md exists at run-dir root (NOT scattered elsewhere)', !!skillText,
75
+ !skillText ? 'missing — did you install it to ~/.pi/agent/skills/ early? Keep it in run-dir until ALL-GREEN' : '');
76
+ }
77
+
78
+ let isSoftware = false;
79
+ let isPersona = false;
80
+ let isTopic = false; // research/topic flavor — its 'apply' is a validated report; no APPLY-LOG required
81
+
82
+ if (skillText) {
83
+ const fm = skillText.match(/^---\n([\s\S]*?)\n---/);
84
+ check('SKILL.md: frontmatter --- block present', !!fm, 'no --- block at top');
85
+ const fmText = fm ? fm[1] : '';
86
+
87
+ const PLACEHOLDERS = [/<person>/i, /<target>/i, /<topic>/i, /<field>/i, /TODO/i, /TBD/i, /XXX/i, /YYYY-MM-DD/i, /\.\.\.\s*<\/?/i];
88
+ const found = PLACEHOLDERS.filter((re) => re.test(skillText));
89
+ check('SKILL.md: no unresolved placeholder text', found.length === 0,
90
+ found.map((re) => skillText.match(re)?.[0]).filter(Boolean).join(', '));
91
+
92
+ // flavor auto-detect
93
+ isSoftware = /target:\s*(software|codebase|engineer)/i.test(fmText) || /code-?dna|代码表达DNA|code expression/i.test(skillText);
94
+ isPersona = !isSoftware && (/target:\s*(person|topic)/i.test(fmText) || /mental.?model/i.test(skillText));
95
+ isTopic = /target:\s*topic/i.test(fmText); // exclude topic/research flavor from APPLY checks
96
+
97
+ if (isSoftware) {
98
+ check('[software] Code-DNA section present', /代码表达DNA|Code Expression-DNA|code.?dna/i.test(skillText), 'missing 代码表达DNA/code-DNA section');
99
+ check('[software] toolchain matrix present', /toolchain|eslint|oxlint|biome|tsconfig|deno|rustfmt/i.test(skillText), 'no toolchain detection');
100
+ check('[software] distilled_against anchor present', /distilled_against/i.test(fmText) || /distilled_against/i.test(skillText), 'no distilled_against staleness anchor');
101
+ }
102
+
103
+ if (isPersona) {
104
+ check('[persona] mental-models section present', /mental.?model|心智模型|核心心智模型/i.test(skillText), 'no mental-models section');
105
+ const boundary = countListItems(skillText, /诚实边界|honest boundar(?:y|ies)/i);
106
+ check('[persona] honest-boundaries ≥3 items', boundary.count >= 3, boundary.found ? `found ${boundary.count} items` : 'section missing');
107
+ }
108
+ }
109
+
110
+ // --- Fallback flavor detection (when SKILL.md is absent/scattered) ---
111
+ // Needed because isSoftware/isPersona are only set inside if(skillText); a run that
112
+ // skipped APPLY often also has no SKILL.md at the run-dir root (it was installed early
113
+ // or never written there). Detect software flavor from research/CODE-DNA.md so the
114
+ // APPLY-evidence checks still fire for the all-too-common 'wrote SKILL, skipped APPLY' case.
115
+ if (!isSoftware && !isPersona && !isTopic) {
116
+ if (existsSync(join(runDir, 'references', 'research', 'CODE-DNA.md')) ||
117
+ existsSync(join(runDir, 'research', 'CODE-DNA.md')) ||
118
+ /code-?dna/i.test(basename(runDir))) {
119
+ isSoftware = true;
120
+ }
121
+ }
122
+
123
+ // --- APPLY-LOG.md (Phase 3 APPLY evidence — Apply mode only) ---
124
+ // Apply mode (APPLY-LOG.md present): the deliverable is the target transformed — these
125
+ // checks are REQUIRED. Capture mode (--build or no APPLY-LOG): skipped (no target-apply).
126
+ // Gated on isApplyMode (not flavor) so it fires even when SKILL.md is absent in Apply mode.
127
+ if (isApplyMode) {
128
+ const applyLogPath = join(runDir, 'APPLY-LOG.md');
129
+ const applyText = readText(applyLogPath);
130
+ check('APPLY-LOG.md exists at run-dir root', !!applyText,
131
+ !applyText ? 'Phase 4 APPLY missing — SKILL.md is intermediate, not the deliverable. Distillation = source → essence → APPLY to target. A standalone SKILL.md = NOT complete.' : '');
132
+ if (applyText) {
133
+ // count list items + table rows across the whole file (mirrors countListItems)
134
+ const applyItems = (applyText.match(/^\s*(?:\d+[.)]|[-*])\s+\S/gm) || []).length
135
+ + (applyText.match(/^\s*\|(?![\s:|-]+\|?\s*$).+\|/gm) || []).length;
136
+ check('APPLY-LOG: ≥3 applied items', applyItems >= 3, `found ${applyItems} item(s) — need ≥3 concrete edits in the target`);
137
+ const pathRefs = /\/[\w.-]+\.\w{1,4}|src\/|AGENTS\.md|scripts\/|package\.json|tsconfig/i.test(applyText);
138
+ check('APPLY-LOG: references target file paths', pathRefs, 'no file-path-like tokens found — add paths proving real edits (not just claims)');
139
+ const verified = /test|typecheck|build|bundle|lint|pass|green|✅|97\/97|all-green/i.test(applyText);
140
+ check('APPLY-LOG: verification evidence (test/typecheck/bundle/lint pass)', verified, 'no verification keyword found — add test/build/lint pass evidence');
141
+ }
142
+ }
143
+
144
+ // --- FIDELITY.md ---
145
+ // --build mode: accept references/fidelity.md as an alias (workflow produces it there)
146
+ const fidText = isBuild
147
+ ? (readText(join(runDir, 'FIDELITY.md')) || readText(join(runDir, 'references', 'fidelity.md')))
148
+ : readText(join(runDir, 'FIDELITY.md'));
149
+ const fidLabel = isBuild ? 'FIDELITY.md exists (root or references/fidelity.md)' : 'FIDELITY.md exists at run-dir root';
150
+ check(fidLabel, !!fidText, !fidText ? 'missing' : '');
151
+ if (fidText) {
152
+ const totalMatch = fidText.match(/(?:总分|total)[::\s*]*\*{0,2}([0-9]+)\s*\*?\s*(?:\/|/|\sout\sof\s)\s*\*?\s*100/i);
153
+ check('FIDELITY.md: /100 total score present', !!totalMatch, totalMatch ? `=${totalMatch[1]}/100` : 'no /100 score found');
154
+ const qCount = (fidText.match(/(?:Q[1-5]|问题[1-5]|question\s*[1-5])/gi) || []).length;
155
+ check('FIDELITY.md: ≥5 test-question references', qCount >= 5, `found ${qCount} question refs`);
156
+ check('FIDELITY.md: flags single-agent/self-score caveat',
157
+ /单\s*agent|single.?agent|self.?score|upper.?bound|independent/i.test(fidText),
158
+ 'add single-agent upper-bound caveat');
159
+ }
160
+
161
+ // --- DISTILLATION-PROCESS-CHECKLIST.md ---
162
+ const procText = readText(join(runDir, 'DISTILLATION-PROCESS-CHECKLIST.md'));
163
+ check('DISTILLATION-PROCESS-CHECKLIST.md exists', !!procText, !procText ? 'missing — every phase + the 3-round deep-dive gate must be tracked' : '');
164
+ if (procText) {
165
+ const dangling = (procText.match(/\|\s*[⬜⏳][^|]*\|/gu) || []).length;
166
+ check('process: no dangling ⬜/⏳ phase rows (every phase completed)', dangling === 0, dangling ? `${dangling} phase(s) not done` : '');
167
+ check('process: deep-dive round-log table present', /round log|\| *Round.*New findings/i.test(procText), 'add the round-log table');
168
+ check('process: 3-empty-rounds gate fired', /gate fires|3 consecutive (empty|zero)|GATE FIRES/i.test(procText),
169
+ 'record ≥3 consecutive zero-new rounds before declaring a phase done');
170
+ }
171
+
172
+ // --- EXCAVATION-CHECKLIST.md ---
173
+ const excText = readText(join(runDir, 'EXCAVATION-CHECKLIST.md'));
174
+ check('EXCAVATION-CHECKLIST.md exists', !!excText, !excText ? 'missing — Phase 1 protocol requires one' : '');
175
+ if (excText) {
176
+ const danglingExc = (excText.match(/\|\s*[⬜⏳][^|]*\|/gu) || []).length;
177
+ check('excavation: no dangling ⬜/⏳ rows (every part resolved)', danglingExc === 0, danglingExc ? `${danglingExc} row(s) not resolved` : '');
178
+ const ratioMatch = excText.match(/memory-ratio[:\s]*([0-9]+)\s*%/i);
179
+ if (ratioMatch) {
180
+ const ratio = parseInt(ratioMatch[1], 10);
181
+ check('excavation: 🧠 memory-ratio ≤30%', ratio <= 30, `declared ${ratio}% (>30% = recap, not distillation)`);
182
+ } else {
183
+ check('excavation: declares memory-ratio', false, 'add "🧠 memory-ratio: NN% (X/Y)" header');
184
+ }
185
+ const proofCells = (excText.match(/"[^"]{8,}"/g) || []).length;
186
+ check('excavation: ≥1 proof-of-read (verbatim quote cell)', proofCells >= 1, `${proofCells} quote cell(s)`);
187
+ }
188
+
189
+ // --- references/research/ artifacts (canonical path — solves scattering) ---
190
+ const researchDir = join(runDir, 'references', 'research');
191
+
192
+ if (isBuild) {
193
+ // --- BUILD MODE: alias-tolerant evidence checks ---
194
+ // Engine skills use different file names than distillation runs.
195
+ // Accept canonical names OR common aliases OR name/content regex matches.
196
+ const refsDir = join(runDir, 'references');
197
+
198
+ // Helper: find a file under references/ or references/research/ by name regex or content regex
199
+ function findRef(nameRe, contentRe) {
200
+ const dirs = [refsDir, researchDir].filter((d) => existsSync(d) && statSync(d).isDirectory());
201
+ for (const d of dirs) {
202
+ for (const f of readdirSync(d)) {
203
+ const fp = join(d, f);
204
+ if (!existsSync(fp) || !statSync(fp).isFile()) continue;
205
+ if (nameRe && nameRe.test(f)) return fp;
206
+ if (contentRe) {
207
+ const txt = readText(fp);
208
+ if (txt && contentRe.test(txt)) return fp;
209
+ }
210
+ }
211
+ }
212
+ return null;
213
+ }
214
+
215
+ // 1. Coverage evidence
216
+ const covPath =
217
+ existsSync(join(researchDir, 'COVERAGE-MANIFEST.md')) ? join(researchDir, 'COVERAGE-MANIFEST.md')
218
+ : existsSync(join(refsDir, 'source-inventory.md')) ? join(refsDir, 'source-inventory.md')
219
+ : existsSync(join(refsDir, 'coverage-manifest.md')) ? join(refsDir, 'coverage-manifest.md')
220
+ : findRef(null, /UNCOVERED.*COVERED|covered.*parts/i);
221
+ const covTextB = covPath ? readText(covPath) : null;
222
+ check('coverage evidence present (COVERAGE-MANIFEST | source-inventory | coverage-manifest | content-match)',
223
+ !!covTextB, !covTextB ? 'missing — no coverage evidence found under references/' : `found: ${covPath}`);
224
+ if (covTextB) {
225
+ const uncovered = (covTextB.match(/\|\s*UNCOVERED[^|]*\|/gi) || []).length;
226
+ check('coverage: no dangling UNCOVERED rows', uncovered === 0, uncovered ? `${uncovered} part(s) not covered` : '');
227
+ }
228
+
229
+ // 2. V5/citation-verify evidence
230
+ const v5Path =
231
+ existsSync(join(researchDir, 'V5-VERIFICATION.md')) ? join(researchDir, 'V5-VERIFICATION.md')
232
+ : existsSync(join(refsDir, 'verified-models.md')) ? join(refsDir, 'verified-models.md')
233
+ : findRef(/v5|verified|citation/i, null);
234
+ check('V5/citation-verify evidence present (V5-VERIFICATION | verified-models | name-match)',
235
+ !!v5Path, !v5Path ? 'missing — no V5/citation evidence found under references/' : `found: ${v5Path}`);
236
+
237
+ // 3. Effectiveness-gate evidence
238
+ const effPath =
239
+ existsSync(join(researchDir, 'EFFECTIVENESS-VERIFICATION.md')) ? join(researchDir, 'EFFECTIVENESS-VERIFICATION.md')
240
+ : existsSync(join(refsDir, 'effectiveness-gate.md')) ? join(refsDir, 'effectiveness-gate.md')
241
+ : existsSync(join(refsDir, 'three-filter.md')) ? join(refsDir, 'three-filter.md')
242
+ : findRef(/effectiveness|three-filter|gate/i, null);
243
+ // FIDELITY.md validates the built skill reproduces the target — for an engine-skill BUILD that IS effectiveness
244
+ // evidence. The pre-apply effectiveness-gate.md is a forward-looking workflow artifact, not retroactively required
245
+ // for hand-built engine skills (research/persona/software predate Phase 2.6).
246
+ const effViaFidelity = !effPath && !!fidText;
247
+ check('effectiveness-gate evidence present (EFFECTIVENESS-VERIFICATION | effectiveness-gate | three-filter | FIDELITY.md)',
248
+ !!effPath || effViaFidelity, (!effPath && !effViaFidelity) ? 'missing — no effectiveness evidence found under references/ or FIDELITY.md' : `found: ${effPath || 'FIDELITY.md'}`);
249
+ } else {
250
+ // --- DEFAULT MODE: strict canonical paths (UNCHANGED — byte-for-byte equivalent) ---
251
+ const covText = readText(join(researchDir, 'COVERAGE-MANIFEST.md'));
252
+ check('references/research/COVERAGE-MANIFEST.md exists', !!covText, !covText ? 'missing' + flatHint('COVERAGE-MANIFEST.md') : '');
253
+ if (covText) {
254
+ const uncovered = (covText.match(/\|\s*UNCOVERED[^|]*\|/gi) || []).length;
255
+ check('coverage: no dangling UNCOVERED rows', uncovered === 0, uncovered ? `${uncovered} part(s) not covered` : '');
256
+ }
257
+
258
+ const v5Exists = existsSync(join(researchDir, 'V5-VERIFICATION.md'));
259
+ check('references/research/V5-VERIFICATION.md exists', v5Exists, !v5Exists ? 'missing' + flatHint('V5-VERIFICATION.md') : '');
260
+
261
+ const effExists = existsSync(join(researchDir, 'EFFECTIVENESS-VERIFICATION.md'));
262
+ check('references/research/EFFECTIVENESS-VERIFICATION.md exists', effExists, !effExists ? 'missing' + flatHint('EFFECTIVENESS-VERIFICATION.md') : '');
263
+ }
264
+
265
+ // --- Phase 2.7 + 5.5 anti-lazy gate checks (APPLY mode only) ---
266
+ // apply-plan + scrutinize are APPLY-mode artifacts (they assume a target to apply to).
267
+ // CAPTURE mode (skill-build: --build, or no APPLY-LOG = persona/workflow) has no target-apply → skip.
268
+ if (isApplyMode) {
269
+ // Phase 2.7 output: references/apply-plan.md
270
+ const applyPlanPath = join(runDir, 'references', 'apply-plan.md');
271
+ const applyPlanText = readText(applyPlanPath);
272
+ check('references/apply-plan.md exists (Phase 2.7 plan-approval gate output)', !!applyPlanText,
273
+ !applyPlanText ? 'missing — Phase 2.7 requires a plan table output' : '');
274
+ if (applyPlanText) {
275
+ const hasApproval = /APPROVED|LOW-YIELD DEFENSE/i.test(applyPlanText);
276
+ check('apply-plan: approval or defense present (APPROVED | LOW-YIELD DEFENSE)', hasApproval,
277
+ !hasApproval ? 'Phase 2.7 plan-approval gate not respected — present the plan and get approval (interactive) or write a LOW-YIELD DEFENSE (autonomous).' : '');
278
+ }
279
+
280
+ // Phase 5.5 output: SCRUTINIZE-REPORT.md at run-dir root
281
+ check('SCRUTINIZE-REPORT.md exists at run-dir root (Phase 5.5 adversarial scrutinize)', existsSync(join(runDir, 'SCRUTINIZE-REPORT.md')),
282
+ !existsSync(join(runDir, 'SCRUTINIZE-REPORT.md')) ? 'Phase 5.5 adversarial scrutinize not run — laziness unaudited.' : '');
283
+ }
284
+
285
+ // --- Report ---
286
+ console.log(`\nvalidate-run: ${basename(runDir)}`);
287
+ console.log(`path: ${runDir}\n`);
288
+ passes.forEach((l) => console.log(l));
289
+ failures.forEach((l) => console.log(l));
290
+ const flavorTag = isSoftware ? ' [software]' : isPersona ? (isTopic ? ' [topic]' : ' [persona]') : '';
291
+ const modeTag = isBuild
292
+ ? ' [build mode — APPLY-LOG + strict-run checks skipped]'
293
+ : isApplyMode
294
+ ? ' [apply mode — SKILL.md optional]'
295
+ : ' [capture mode]';
296
+ console.log(`\n${passes.length} pass, ${failures.length} fail — ${failures.length === 0 ? '✅ ALL-GREEN (may ship)' : '🔴 NOT READY (produce the missing artifacts, then re-run)'}${flavorTag}${modeTag}\n`);
297
+ process.exit(failures.length === 0 ? 0 : 1);