pi-crew 0.9.47 → 0.9.49
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/AGENTS.md +18 -0
- package/CHANGELOG.md +132 -4
- package/README.md +1 -1
- package/dist/build-meta.json +22 -12
- package/dist/index.mjs +430 -388
- package/dist/index.mjs.map +3 -3
- package/docs/decisions/2026-07-24-oidc-trusted-publishing.md +112 -0
- package/docs/publishing.md +5 -2
- package/package.json +2 -2
- package/scripts/postinstall.mjs +43 -18
- package/skills/.gitkeep +0 -0
- package/skills/distill-persona/BUILD-NOTES.md +55 -0
- package/skills/distill-persona/SKILL.md +612 -0
- package/skills/distill-persona/UPGRADE-LOG-RESEARCH-SKILLS.md +100 -0
- package/skills/distill-persona/references/coverage-manifest.md +65 -0
- package/skills/distill-persona/references/distillation-field-synthesis-pass2.md +59 -0
- package/skills/distill-persona/references/distillation-field-synthesis.md +108 -0
- package/skills/distill-persona/references/handoff.md +42 -0
- package/skills/distill-persona/references/research/lesson-memory-shortcut.md +33 -0
- package/skills/distill-persona/references/research/r1-a-examples.md +23 -0
- package/skills/distill-persona/references/research/r1-b-scripts.md +26 -0
- package/skills/distill-persona/references/research/r1-c-human-readme.md +31 -0
- package/skills/distill-persona/references/research/r1-d-tests.md +28 -0
- package/skills/distill-persona/references/research/r1-verification.md +36 -0
- package/skills/distill-persona/references/research/r2-low-yield.md +26 -0
- package/skills/distill-persona/scripts/fidelity_eval.py +244 -0
- package/skills/distill-persona/scripts/validate-skill-structure.mjs +177 -0
- package/skills/distill-software/BUILD-NOTES.md +56 -0
- package/skills/distill-software/SKILL.md +302 -0
- package/skills/distill-software/references/handoff.md +47 -0
- package/skills/distill-software/scripts/code_dna.py +290 -0
- package/skills/research/DISTILLATION-PROCESS-CHECKLIST.md +120 -0
- package/skills/research/EXCAVATION-CHECKLIST.md +142 -0
- package/skills/research/FIDELITY.md +180 -0
- package/skills/research/SKILL.md +432 -0
- package/skills/research/references/anti-patterns.md +184 -0
- package/skills/research/references/fidelity.md +241 -0
- package/skills/research/references/handoff.md +48 -0
- package/skills/research/references/research-protocol.md +162 -0
- package/skills/research/references/source-inventory.md +135 -0
- package/skills/research/references/verified-models.md +163 -0
- package/skills/research/scripts/__pycache__/safe_io.cpython-312.pyc +0 -0
- package/skills/research/scripts/code_dna.py +233 -0
- package/skills/research/scripts/emit_run_summary.py +142 -0
- package/skills/research/scripts/safe_io.py +314 -0
- package/skills/research/scripts/source_evaluator.py +234 -0
- package/skills/research/scripts/validate-skill-structure.mjs +177 -0
- package/skills/research/scripts/verify_citations.py +225 -0
- package/skills/security-priority.json +28 -0
- package/src/config/config.ts +1 -0
- package/src/config/role-tools.ts +6 -3
- package/src/config/types.ts +8 -0
- package/src/runtime/background-runner.ts +11 -16
- package/src/runtime/heartbeat-watcher.ts +28 -1
- package/src/runtime/task-runner.ts +165 -119
- package/src/schema/config-schema.ts +1 -0
- package/src/utils/gh-protocol.ts +9 -8
- package/workflows/distill.workflow.md +198 -0
|
@@ -0,0 +1,244 @@
|
|
|
1
|
+
#!/usr/bin/env python3
|
|
2
|
+
"""
|
|
3
|
+
fidelity_eval.py — Phase 4 automation + reproducibility test for distilled skills.
|
|
4
|
+
|
|
5
|
+
Closes two gaps from the distill-persona build:
|
|
6
|
+
1. Automates Phase 4 fidelity validation (known-stance + FRAMEWORK-ANSWERABLE novel
|
|
7
|
+
edge + blind independent scorer). The novel edge MUST be framework-answerable
|
|
8
|
+
(validated at n=3: fact-demanding edges give false passes via refusal vocab).
|
|
9
|
+
2. Adds a reproducibility test (distill the same persona twice, diff the mental-model
|
|
10
|
+
sets; flag if overlap < 60%).
|
|
11
|
+
|
|
12
|
+
Runtime-agnostic by design: the deterministic parts (test-spec validation, scorecard
|
|
13
|
+
parsing, reproducibility diff, grading) live here; the LLM calls (answerer + scorer)
|
|
14
|
+
are emitted as ready-to-dispatch prompts that THIS runtime's subagent mechanism runs
|
|
15
|
+
(pi-crew `Agent`, Claude Code subagents, etc.). We never hardcode an LLM API — there is
|
|
16
|
+
no key here, and runtimes vary.
|
|
17
|
+
|
|
18
|
+
Stdlib only. Python 3.9+.
|
|
19
|
+
|
|
20
|
+
Usage:
|
|
21
|
+
python3 fidelity_eval.py validate <test-spec.md> # check spec structure
|
|
22
|
+
python3 fidelity_eval.py runbook <skill-dir> <spec> # emit answerer+scorer prompts
|
|
23
|
+
python3 fidelity_eval.py parse <score.md> # extract 5-dim numbers
|
|
24
|
+
python3 fidelity_eval.py repro <skillA.md> <skillB.md> # diff mental models (Jaccard)
|
|
25
|
+
|
|
26
|
+
Test-spec format (markdown):
|
|
27
|
+
## known
|
|
28
|
+
1. <question>
|
|
29
|
+
2. <question>
|
|
30
|
+
3. <question>
|
|
31
|
+
## edge # exactly ONE, must be framework-answerable
|
|
32
|
+
<question>
|
|
33
|
+
## ground-truth
|
|
34
|
+
<the persona's real documented public stances — for the blind scorer only>
|
|
35
|
+
## edge-note
|
|
36
|
+
<one line: why this edge is framework-answerable + that the persona never addressed it>
|
|
37
|
+
"""
|
|
38
|
+
import argparse
|
|
39
|
+
import re
|
|
40
|
+
import sys
|
|
41
|
+
from pathlib import Path
|
|
42
|
+
|
|
43
|
+
# --- grading thresholds (from nuwa fidelity-scorecard, + the F2' edge-honesty gate) ---
|
|
44
|
+
GRADE_A = 85 # ship
|
|
45
|
+
GRADE_B = 70 # acceptable with flagged weak spots
|
|
46
|
+
EDGE_HONESTY_SHIP = 14 # /20 — a skill that fails edge-honesty is NOT ship-ready even if total is high
|
|
47
|
+
REPRO_FLOOR = 0.60 # mental-model Jaccard overlap below this flags non-reproducible distillation
|
|
48
|
+
|
|
49
|
+
DIMS = [
|
|
50
|
+
("stance_consistency", 30),
|
|
51
|
+
("style_recognizability", 20),
|
|
52
|
+
("edge_honesty", 20),
|
|
53
|
+
("source_transparency", 15),
|
|
54
|
+
("structural_completeness", 15),
|
|
55
|
+
]
|
|
56
|
+
|
|
57
|
+
|
|
58
|
+
# --------------------------------------------------------------------------- spec validation
|
|
59
|
+
def parse_spec(path: Path) -> dict:
|
|
60
|
+
text = path.read_text(encoding="utf-8")
|
|
61
|
+
sections = {}
|
|
62
|
+
cur = None
|
|
63
|
+
for line in text.splitlines():
|
|
64
|
+
m = re.match(r"^##\s+(\S.*)$", line)
|
|
65
|
+
if m:
|
|
66
|
+
cur = m.group(1).strip().lower().replace("-", "_").replace(" ", "_")
|
|
67
|
+
sections[cur] = []
|
|
68
|
+
elif cur:
|
|
69
|
+
sections[cur].append(line)
|
|
70
|
+
return {k: "\n".join(v).strip() for k, v in sections.items()}
|
|
71
|
+
|
|
72
|
+
|
|
73
|
+
def validate_spec(path: Path) -> int:
|
|
74
|
+
s = parse_spec(path)
|
|
75
|
+
problems = []
|
|
76
|
+
if "known" not in s:
|
|
77
|
+
problems.append("missing '## known' section")
|
|
78
|
+
else:
|
|
79
|
+
known = [ln for ln in s["known"].splitlines() if re.match(r"^\s*\d+\.", ln.strip())]
|
|
80
|
+
if len(known) != 3:
|
|
81
|
+
problems.append(f"'## known' must have exactly 3 numbered questions, found {len(known)}")
|
|
82
|
+
if "edge" not in s or not s["edge"].strip():
|
|
83
|
+
problems.append("missing '## edge' section (exactly ONE novel edge question)")
|
|
84
|
+
if "ground_truth" not in s or len(s["ground_truth"]) < 50:
|
|
85
|
+
problems.append("missing or too-short '## ground-truth' (the blind scorer's reference)")
|
|
86
|
+
if "edge_note" not in s or "framework" not in s["edge_note"].lower():
|
|
87
|
+
problems.append("'## edge-note' must state why the edge is FRAMEWORK-answerable "
|
|
88
|
+
"(fact-demanding edges give false passes — see f2-experiment/validation-conclusion.md)")
|
|
89
|
+
if problems:
|
|
90
|
+
print("❌ INVALID test spec:")
|
|
91
|
+
for p in problems:
|
|
92
|
+
print(f" - {p}")
|
|
93
|
+
return 1
|
|
94
|
+
print("✅ test spec OK (3 known + 1 framework-answerable edge + ground-truth + edge-note)")
|
|
95
|
+
print(f" edge question: {s['edge'].splitlines()[0][:90]}")
|
|
96
|
+
return 0
|
|
97
|
+
|
|
98
|
+
|
|
99
|
+
# --------------------------------------------------------------------------- runbook emission
|
|
100
|
+
RUNBOOK_HEADER = """# Fidelity runbook for {skill}
|
|
101
|
+
|
|
102
|
+
Runtime-agnostic. Run the two steps below via THIS runtime's subagent mechanism, in order.
|
|
103
|
+
Step 1 (answerer) MUST finish before Step 2 (scorer). The scorer is BLIND (ground-truth only,
|
|
104
|
+
NO skill-file access) — F2 was refuted (skill-file access does not inflate scores) but the blind
|
|
105
|
+
condition keeps the score independent of the author's design intent.
|
|
106
|
+
|
|
107
|
+
## Step 1 — answerer (skill-only, NO web)
|
|
108
|
+
Allowlist these files ONLY: <skill-dir>/SKILL.md + <skill-dir>/references/research/*
|
|
109
|
+
Answer the 3 known + 1 edge question in character per the skill's Agentic Protocol.
|
|
110
|
+
For the edge question, honor the skill's F2' third-category rule (flag framework-derived
|
|
111
|
+
inference as inference, not established stance). Write answers to answers.md.
|
|
112
|
+
|
|
113
|
+
## Step 2 — blind scorer (ground-truth ONLY, NO skill file)
|
|
114
|
+
Allowlist: answers.md + rubric.md + ground-truth (from the test spec). Score 5 dims.
|
|
115
|
+
Write scorecard to score-B.md.
|
|
116
|
+
|
|
117
|
+
## Questions
|
|
118
|
+
{questions}
|
|
119
|
+
|
|
120
|
+
## Ground-truth (for blind scorer)
|
|
121
|
+
{ground_truth}
|
|
122
|
+
"""
|
|
123
|
+
|
|
124
|
+
|
|
125
|
+
def emit_runbook(skill_dir: Path, spec_path: Path) -> int:
|
|
126
|
+
s = parse_spec(spec_path)
|
|
127
|
+
questions = "### known\n" + s.get("known", "") + "\n\n### edge (FRAMEWORK-answerable)\n" + s.get("edge", "")
|
|
128
|
+
out = RUNBOOK_HEADER.format(
|
|
129
|
+
skill=skill_dir.name,
|
|
130
|
+
questions=questions,
|
|
131
|
+
ground_truth=s.get("ground_truth", "(none)"),
|
|
132
|
+
)
|
|
133
|
+
dest = skill_dir / "fidelity-runbook.md"
|
|
134
|
+
dest.write_text(out, encoding="utf-8")
|
|
135
|
+
print(f"✅ wrote {dest}")
|
|
136
|
+
print(" next: dispatch Step 1 (answerer) then Step 2 (blind scorer) via your runtime.")
|
|
137
|
+
return 0
|
|
138
|
+
|
|
139
|
+
|
|
140
|
+
# --------------------------------------------------------------------------- scorecard parsing
|
|
141
|
+
def parse_score(path: Path) -> dict:
|
|
142
|
+
"""Extract the 5 dimension scores from a scorer's markdown scorecard."""
|
|
143
|
+
text = path.read_text(encoding="utf-8")
|
|
144
|
+
found = {}
|
|
145
|
+
for name, _ in DIMS:
|
|
146
|
+
# match e.g. "| 1 Stance consistency | 26/30 |" or "| 3 Edge honesty | **6/20** |"
|
|
147
|
+
# [\s*]* tolerates bold (**..**) wrappers around the score.
|
|
148
|
+
pat = re.compile(rf"\|[^|]*{re.escape(name.split('_')[0])}[^|]*\|[\s*]*(\d+)\s*/\s*(\d+)", re.I)
|
|
149
|
+
m = pat.search(text)
|
|
150
|
+
if m:
|
|
151
|
+
found[name] = int(m.group(1))
|
|
152
|
+
total_m = re.search(r"\*\*TOTAL\*\*\s*\|\s*(\d+)\s*/\s*100", text, re.I)
|
|
153
|
+
if not total_m:
|
|
154
|
+
total_m = re.search(r"(\d+)\s*/\s*100", text)
|
|
155
|
+
total = int(total_m.group(1)) if total_m else sum(found.values())
|
|
156
|
+
return {"dims": found, "total": total}
|
|
157
|
+
|
|
158
|
+
|
|
159
|
+
def report_score(path: Path) -> int:
|
|
160
|
+
r = parse_score(path)
|
|
161
|
+
if not r["dims"]:
|
|
162
|
+
print("❌ could not parse any dimension scores from", path)
|
|
163
|
+
return 1
|
|
164
|
+
print(f"Scores from {path}:")
|
|
165
|
+
for name, _max in DIMS:
|
|
166
|
+
v = r["dims"].get(name)
|
|
167
|
+
print(f" {name:28} {v}/{_max}" if v is not None else f" {name:28} (not found)")
|
|
168
|
+
print(f" {'TOTAL':28} {r['total']}/100")
|
|
169
|
+
grade = grade_for(r["total"], r["dims"].get("edge_honesty", 0))
|
|
170
|
+
print(f" grade: {grade}")
|
|
171
|
+
return 0
|
|
172
|
+
|
|
173
|
+
|
|
174
|
+
def grade_for(total: int, edge: int) -> str:
|
|
175
|
+
if edge < EDGE_HONESTY_SHIP:
|
|
176
|
+
return f"NO-SHIP (edge-honesty {edge}/20 < {EDGE_HONESTY_SHIP} — fails the F2' gate regardless of total)"
|
|
177
|
+
if total >= GRADE_A:
|
|
178
|
+
return f"A (≥{GRADE_A}) — ship"
|
|
179
|
+
if total >= GRADE_B:
|
|
180
|
+
return f"B (≥{GRADE_B}) — acceptable with flagged weak spots"
|
|
181
|
+
return f"C (<{GRADE_B}) — re-distill"
|
|
182
|
+
|
|
183
|
+
|
|
184
|
+
# --------------------------------------------------------------------------- reproducibility diff
|
|
185
|
+
def extract_models(skill_md: Path) -> set:
|
|
186
|
+
"""Extract mental-model names from a SKILL.md (### 模型N: / ### Model N: / ### N. Name)."""
|
|
187
|
+
text = skill_md.read_text(encoding="utf-8")
|
|
188
|
+
names = set()
|
|
189
|
+
# match model headings in either language; tolerate arabic OR Chinese numerals.
|
|
190
|
+
for m in re.finditer(r"^#{2,4}\s+(?:模型\s*[0-9一二三四五六七八九十]+\s*[::]?\s*|Model\s*\d+\s*[::]\s*|\d+\.\s*)(.+)$", text, re.M):
|
|
191
|
+
name = m.group(1).strip().split("(")[0].strip().split("(")[0].strip()
|
|
192
|
+
# drop parenthetical english + trailing junk
|
|
193
|
+
name = re.sub(r"\s+", " ", name)
|
|
194
|
+
if 2 <= len(name) <= 60:
|
|
195
|
+
names.add(name.lower())
|
|
196
|
+
return names
|
|
197
|
+
|
|
198
|
+
|
|
199
|
+
def repro(a: Path, b: Path) -> int:
|
|
200
|
+
ma, mb = extract_models(a), extract_models(b)
|
|
201
|
+
if not ma or not mb:
|
|
202
|
+
print(f"❌ could not extract models (A={len(ma)}, B={len(mb)}). "
|
|
203
|
+
"Ensure SKILL.md uses '### 模型N: Name' / '### Model N: Name' headings.")
|
|
204
|
+
return 1
|
|
205
|
+
inter = ma & mb
|
|
206
|
+
union = ma | mb
|
|
207
|
+
jaccard = len(inter) / len(union) if union else 0.0
|
|
208
|
+
name_overlap = len(inter) / min(len(ma), len(mb)) if min(len(ma), len(mb)) else 0.0
|
|
209
|
+
print(f"Skill A models ({len(ma)}): {sorted(ma)}")
|
|
210
|
+
print(f"Skill B models ({len(mb)}): {sorted(mb)}")
|
|
211
|
+
print(f"intersection ({len(inter)}): {sorted(inter)}")
|
|
212
|
+
print(f"Jaccard overlap: {jaccard:.0%} name-overlap/min: {name_overlap:.0%}")
|
|
213
|
+
if name_overlap < REPRO_FLOOR:
|
|
214
|
+
print(f"⚠️ NON-REPRODUCIBLE: overlap {name_overlap:.0%} < {REPRO_FLOOR:.0%} floor — "
|
|
215
|
+
"two distillations of the same persona diverge too much. Investigate Phase 2 subjectivity.")
|
|
216
|
+
return 2
|
|
217
|
+
print(f"✅ reproducible (overlap {name_overlap:.0%} ≥ {REPRO_FLOOR:.0%})")
|
|
218
|
+
return 0
|
|
219
|
+
|
|
220
|
+
|
|
221
|
+
# --------------------------------------------------------------------------- CLI
|
|
222
|
+
def main() -> int:
|
|
223
|
+
ap = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)
|
|
224
|
+
sub = ap.add_subparsers(dest="cmd", required=True)
|
|
225
|
+
sub.add_parser("validate", help="validate a test-spec.md structure").add_argument("spec")
|
|
226
|
+
sub.add_parser("runbook", help="emit answerer+scorer prompts for a skill").add_argument("skill_dir")
|
|
227
|
+
[a.add_argument("spec") for a in [sub.choices["runbook"]]]
|
|
228
|
+
sub.add_parser("parse", help="extract 5-dim scores from a scorer's markdown").add_argument("score")
|
|
229
|
+
r = sub.add_parser("repro", help="diff mental models of two SKILL.md (reproducibility test)")
|
|
230
|
+
r.add_argument("skillA"); r.add_argument("skillB")
|
|
231
|
+
args = ap.parse_args()
|
|
232
|
+
if args.cmd == "validate":
|
|
233
|
+
return validate_spec(Path(args.spec))
|
|
234
|
+
if args.cmd == "runbook":
|
|
235
|
+
return emit_runbook(Path(args.skill_dir), Path(args.spec))
|
|
236
|
+
if args.cmd == "parse":
|
|
237
|
+
return report_score(Path(args.score))
|
|
238
|
+
if args.cmd == "repro":
|
|
239
|
+
return repro(Path(args.skillA), Path(args.skillB))
|
|
240
|
+
return 1
|
|
241
|
+
|
|
242
|
+
|
|
243
|
+
if __name__ == "__main__":
|
|
244
|
+
sys.exit(main())
|
|
@@ -0,0 +1,177 @@
|
|
|
1
|
+
#!/usr/bin/env node
|
|
2
|
+
// validate-skill-structure.mjs — structural invariant checker for generated distill skills.
|
|
3
|
+
// Implements F10 (awesome-persona-distill-skills finding): hard-fail if any structural assertion fails.
|
|
4
|
+
// Run AFTER Phase 3 build, BEFORE Phase 4 behavioral fidelity. Self-contained (stdlib only).
|
|
5
|
+
//
|
|
6
|
+
// Usage:
|
|
7
|
+
// node validate-skill-structure.mjs <path-to-SKILL.md>
|
|
8
|
+
// node validate-skill-structure.mjs <skill-dir>
|
|
9
|
+
//
|
|
10
|
+
// Exit codes: 0 = all assertions pass (all-green → may ship); 1 = one or more failed (iterate).
|
|
11
|
+
|
|
12
|
+
import { readFileSync, readdirSync, statSync, existsSync } from 'node:fs';
|
|
13
|
+
import { join, dirname, basename } from 'node:path';
|
|
14
|
+
|
|
15
|
+
const target = process.argv.find((a) => !a.startsWith('-') && a !== process.argv[0] && a !== process.argv[1]);
|
|
16
|
+
if (!target) {
|
|
17
|
+
console.error('Usage: validate-skill-structure.mjs <path-to-SKILL.md | skill-dir> [--engine]');
|
|
18
|
+
process.exit(2);
|
|
19
|
+
}
|
|
20
|
+
|
|
21
|
+
// Resolve to a SKILL.md path
|
|
22
|
+
let skillPath = target;
|
|
23
|
+
if (statSync(target).isDirectory()) {
|
|
24
|
+
skillPath = join(target, 'SKILL.md');
|
|
25
|
+
}
|
|
26
|
+
if (!existsSync(skillPath)) {
|
|
27
|
+
console.error(`✗ Not found: ${skillPath}`);
|
|
28
|
+
process.exit(2);
|
|
29
|
+
}
|
|
30
|
+
|
|
31
|
+
const src = readFileSync(skillPath, 'utf8');
|
|
32
|
+
const dir = dirname(skillPath);
|
|
33
|
+
const name = basename(dir);
|
|
34
|
+
const isEngine = process.argv.includes('--engine');
|
|
35
|
+
const PLACEHOLDERS = [/<person>/i, /<target>/i, /<topic>/i, /<field>/i, /TODO/i, /TBD/i, /XXX/i, /YYYY-MM-DD/i, /\.\.\.\s*<\/?/i];
|
|
36
|
+
|
|
37
|
+
const failures = [];
|
|
38
|
+
const passes = [];
|
|
39
|
+
const check = (label, ok, detail = '') => {
|
|
40
|
+
(ok ? passes : failures).push(ok ? ` ✓ ${label}` : ` ✗ ${label}${detail ? ' — ' + detail : ''}`);
|
|
41
|
+
};
|
|
42
|
+
|
|
43
|
+
// --- Frontmatter ---
|
|
44
|
+
const fm = src.match(/^---\n([\s\S]*?)\n---/);
|
|
45
|
+
check('frontmatter block present', !!fm, 'no --- block at top');
|
|
46
|
+
const fmText = fm ? fm[1] : '';
|
|
47
|
+
const fmField = (key) => {
|
|
48
|
+
const m = fmText.match(new RegExp(`^${key}:\\s*(.+)$`, 'm'));
|
|
49
|
+
return m ? m[1].trim() : null;
|
|
50
|
+
};
|
|
51
|
+
const hasFm = (key) => new RegExp(`^${key}:`, 'm').test(fmText);
|
|
52
|
+
|
|
53
|
+
check('frontmatter: name', hasFm('name'));
|
|
54
|
+
check('frontmatter: description', hasFm('description'));
|
|
55
|
+
const desc = fmField('description');
|
|
56
|
+
check('frontmatter: description ≤1 sentence (one terminal punct or one clause)',
|
|
57
|
+
desc ? desc.split(/[.。!!??]/).filter(Boolean).length <= 2 : false,
|
|
58
|
+
desc ? `got: "${desc.slice(0, 60)}…"` : 'missing');
|
|
59
|
+
check('frontmatter: triggers', hasFm('triggers') || hasFm('trigger'), 'no triggers field');
|
|
60
|
+
if (!isEngine) check('frontmatter: distilled (staleness date)', hasFm('distilled') || hasFm('调研时间'), 'no distilled/调研时间 staleness anchor');
|
|
61
|
+
const distilled = fmField('distilled') || fmField('调研时间');
|
|
62
|
+
check('frontmatter: distilled is valid date (YYYY-MM-DD)',
|
|
63
|
+
!distilled || /^\d{4}-\d{2}-\d{2}/.test(distilled),
|
|
64
|
+
distilled ? `got "${distilled}"` : '');
|
|
65
|
+
if (!isEngine) check('frontmatter: target (person|topic|software)', hasFm('target'), 'no target field');
|
|
66
|
+
|
|
67
|
+
// --- Body sections ---
|
|
68
|
+
check('Agentic Protocol section present', /回答工作流|Agentic Protocol/i.test(src));
|
|
69
|
+
check('Agentic Protocol Step 1', /Step 1|第一步|步骤 1|### 1\b/i.test(src));
|
|
70
|
+
check('Agentic Protocol Step 2', /Step 2|第二步|步骤 2|### 2\b/i.test(src));
|
|
71
|
+
check('Agentic Protocol Step 3', /Step 3|第三步|步骤 3|### 3\b/i.test(src));
|
|
72
|
+
|
|
73
|
+
// honest boundaries: count items (numbered list or bullets) under an honest-boundaries heading
|
|
74
|
+
function countListItems(haystack, headingRe, windowChars = 2000) {
|
|
75
|
+
const m = haystack.match(headingRe);
|
|
76
|
+
if (!m) return { found: false, count: 0, block: '' };
|
|
77
|
+
const after = haystack.slice(m.index);
|
|
78
|
+
const section = after.match(/([\s\S]*?)\n##(?=[^#])/);
|
|
79
|
+
const block = section ? section[1] : after.slice(0, windowChars);
|
|
80
|
+
// count BOTH list items AND table data rows (rows with | content |, excluding separator rows)
|
|
81
|
+
const listItems = (block.match(/^\s*(?:\d+[.)]|[-*])\s+\S/gm) || []).length;
|
|
82
|
+
const tableRows = (block.match(/^\s*\|(?![\s:|-]+\|?\s*$).+\|/gm) || []).length;
|
|
83
|
+
return { found: true, count: listItems + tableRows, block };
|
|
84
|
+
}
|
|
85
|
+
if (!isEngine) {
|
|
86
|
+
const boundary = countListItems(src, /诚实边界|honest boundar(?:y|ies)/i);
|
|
87
|
+
check('honest boundaries (M11) ≥3', boundary.count >= 3, `found ${boundary.count} items`);
|
|
88
|
+
}
|
|
89
|
+
|
|
90
|
+
// --- Anti-drift tables (M9a 内在张力, M9b 反例黑名单, M12 fallback tree) — Darwin gap #2: validator previously skipped these
|
|
91
|
+
if (!isEngine) {
|
|
92
|
+
const tension = countListItems(src, /M9a|内在张力|inner tension|internal tension/i);
|
|
93
|
+
check('内在张力 (M9a) ≥3 tension pairs', tension.count >= 3, tension.found ? `found ${tension.count} items` : 'section missing');
|
|
94
|
+
|
|
95
|
+
const blacklist = countListItems(src, /M9b|反例黑名单|anti.?pattern blacklist/i);
|
|
96
|
+
check('反例黑名单 (M9b) ≥7 rows', blacklist.count >= 7, blacklist.found ? `found ${blacklist.count} rows` : 'section missing');
|
|
97
|
+
|
|
98
|
+
const fallback = countListItems(src, /M12|失败模式.*[Ff]allback|[Ff]allback\s*树/i);
|
|
99
|
+
check('失败模式Fallback树 (M12) ≥8 rows', fallback.count >= 8, fallback.found ? `found ${fallback.count} rows` : 'section missing');
|
|
100
|
+
}
|
|
101
|
+
|
|
102
|
+
// --- FIDELITY.md companion artifact (Darwin gap #3: ship-gate requires it but nothing checked)
|
|
103
|
+
const fidelityPath = join(dir, 'FIDELITY.md');
|
|
104
|
+
if (existsSync(fidelityPath)) {
|
|
105
|
+
const fm2 = readFileSync(fidelityPath, 'utf8');
|
|
106
|
+
const totalMatch = fm2.match(/(?:总分|total)[::\s*]*\*{0,2}([0-9]+)\s*\*?\s*(?:\/|/|\sout\sof\s)\s*\*?\s*100/i);
|
|
107
|
+
check('FIDELITY.md total score present', !!totalMatch, totalMatch ? `=${totalMatch[1]}/100` : 'no /100 score found');
|
|
108
|
+
const qCount = (fm2.match(/(?:Q[1-5]|问题[1-5]|question\s*[1-5])/gi) || []).length;
|
|
109
|
+
check('FIDELITY.md ≥5 test questions', qCount >= 5, `found ${qCount} question refs`);
|
|
110
|
+
check('FIDELITY.md flags single-agent/self-score caveat', /单\s*agent|single.?agent|self.?score|upper.?bound|independent/i.test(fm2), 'add single-agent upper-bound caveat');
|
|
111
|
+
} else if (!isEngine) {
|
|
112
|
+
check('FIDELITY.md companion present', false, 'no FIDELITY.md in skill dir');
|
|
113
|
+
}
|
|
114
|
+
|
|
115
|
+
// --- EXCAVATION-CHECKLIST.md (Phase 1 protocol: track + verify each part was really read)
|
|
116
|
+
const checklistPath = join(dir, 'EXCAVATION-CHECKLIST.md');
|
|
117
|
+
if (existsSync(checklistPath)) {
|
|
118
|
+
const cl = readFileSync(checklistPath, 'utf8');
|
|
119
|
+
// dangling rows: status ⬜ not-started or ⏳ reading left at ship time = silently skipped/forgotten
|
|
120
|
+
const dangling = (cl.match(/\|\s*[⬜⏳][^|]*\|/gu) || []).length;
|
|
121
|
+
check('checklist: no dangling ⬜/⏳ rows (every part resolved)', dangling === 0, dangling ? `${dangling} row(s) not-started/reading — resolve to ✅📄/⏭/🧠` : '');
|
|
122
|
+
// memory ratio: parse "memory-ratio: NN%" if declared
|
|
123
|
+
const ratioMatch = cl.match(/memory-ratio[:\s]*([0-9]+)\s*%/i);
|
|
124
|
+
if (ratioMatch) {
|
|
125
|
+
const ratio = parseInt(ratioMatch[1], 10);
|
|
126
|
+
check('checklist: 🧠 memory-ratio ≤30%', ratio <= 30, `declared ${ratio}% (>30% = recap, not distillation)`);
|
|
127
|
+
} else {
|
|
128
|
+
check('checklist: declares memory-ratio', false, 'add "🧠 memory-ratio: NN% (X/Y findings)" header line');
|
|
129
|
+
}
|
|
130
|
+
// proof-of-read present: at least one verbatim-quote-with-location cell (heuristic — a ✅ row should carry a quote)
|
|
131
|
+
const proofCells = (cl.match(/"[^"]{8,}"[^|]*/g) || []).length;
|
|
132
|
+
check('checklist: ≥1 proof-of-read (verbatim quote in a ✅ row)', proofCells >= 1, `${proofCells} quote cell(s) found`);
|
|
133
|
+
} else if (!isEngine) {
|
|
134
|
+
check('EXCAVATION-CHECKLIST.md present', false, 'no excavation checklist — Phase 1 protocol requires one');
|
|
135
|
+
}
|
|
136
|
+
|
|
137
|
+
// --- DISTILLATION-PROCESS-CHECKLIST.md (whole-pipeline tracking + the 3-empty-rounds deep-dive gate)
|
|
138
|
+
const procPath = join(dir, 'DISTILLATION-PROCESS-CHECKLIST.md');
|
|
139
|
+
if (existsSync(procPath)) {
|
|
140
|
+
const pc = readFileSync(procPath, 'utf8');
|
|
141
|
+
// dangling phases: ⬜/⏳ left in the phase-progress table at ship = a phase skipped
|
|
142
|
+
const danglingPhases = (pc.match(/\|\s*[⬜⏳][^|]*\|/gu) || []).length;
|
|
143
|
+
check('process: no dangling ⬜/⏳ phases (every phase completed)', danglingPhases === 0, danglingPhases ? `${danglingPhases} phase(s) not done` : '');
|
|
144
|
+
// deep-dive round log present
|
|
145
|
+
check('process: deep-dive round log present', /round log|\| *Round.*New findings/i.test(pc), 'add the round-log table');
|
|
146
|
+
// 3-empty-rounds gate recorded (≥3 consecutive zero-new rounds before leaving a phase)
|
|
147
|
+
const gateFired = /gate fires|3 consecutive (empty|zero)|GATE FIRES/i.test(pc);
|
|
148
|
+
check('process: 3-empty-rounds gate recorded (≥3 zero-new rounds)', gateFired, 'record a round-log row marked gate-fired before declaring any research phase done');
|
|
149
|
+
} else if (!isEngine) {
|
|
150
|
+
check('DISTILLATION-PROCESS-CHECKLIST.md present', false, 'no process checklist — every phase + the 3-round deep-dive gate must be tracked');
|
|
151
|
+
}
|
|
152
|
+
|
|
153
|
+
// --- Placeholders ---
|
|
154
|
+
const foundPlaceholders = PLACEHOLDERS.filter((re) => re.test(src));
|
|
155
|
+
if (!isEngine) check('no unresolved placeholder text', foundPlaceholders.length === 0,
|
|
156
|
+
foundPlaceholders.map((re) => src.match(re)?.[0]).filter(Boolean).join(', '));
|
|
157
|
+
|
|
158
|
+
// --- Self-containment (F9): references/sources should not dangle outside the dir ---
|
|
159
|
+
const refsDir = join(dir, 'references');
|
|
160
|
+
if (!isEngine) check('self-contained: no external repo paths in body (e.g. 07-调研/)',
|
|
161
|
+
!/[\w/-]*调研与分析?\/|src\/|source\//i.test(src.replace(/```[\s\S]*?```/g, '')),
|
|
162
|
+
'body references paths outside skill dir');
|
|
163
|
+
|
|
164
|
+
// --- Software-specific assertions (if target: software) ---
|
|
165
|
+
if (/target:\s*software/i.test(fmText) || /code-dna|代码表达DNA|code expression/i.test(src)) {
|
|
166
|
+
check('[software] Code Expression-DNA section present', /代码表达DNA|Code Expression-DNA|code.?dna/i.test(src));
|
|
167
|
+
check('[software] toolchain matrix present', /toolchain|eslint|oxlint|biome|tsconfig/i.test(src));
|
|
168
|
+
check('[software] distilled_against commit anchor', hasFm('distilled_against') || /distilled_against/i.test(src));
|
|
169
|
+
}
|
|
170
|
+
|
|
171
|
+
// --- Report ---
|
|
172
|
+
console.log(`\nvalidate-skill-structure: ${name}`);
|
|
173
|
+
console.log(`path: ${skillPath}\n`);
|
|
174
|
+
passes.forEach((l) => console.log(l));
|
|
175
|
+
failures.forEach((l) => console.log(l));
|
|
176
|
+
console.log(`\n${passes.length} pass, ${failures.length} fail — ${failures.length === 0 ? '✅ ALL-GREEN (may ship)' : '🔴 NOT READY (iterate Phase 2→3)'}${isEngine ? ' [engine mode — generated-skill checks skipped]' : ''}\n`);
|
|
177
|
+
process.exit(failures.length === 0 ? 0 : 1);
|
|
@@ -0,0 +1,56 @@
|
|
|
1
|
+
# distill-software — Build Notes
|
|
2
|
+
|
|
3
|
+
Sibling of `distill-persona`, specialized for software. Inherits the base methodology (6 phases, Phase 2.6 extraction verification, F2' framework-answerable-edge fidelity, exhaustive-sweep mode + coverage manifest + diminishing-returns gate, self-correction meta-loop) — does NOT duplicate it; only specifies what's different for software.
|
|
4
|
+
|
|
5
|
+
## Decision → source trace
|
|
6
|
+
| Decision | From |
|
|
7
|
+
|---|---|
|
|
8
|
+
| 3 flavors (engineer persona / codebase conventions / domain expertise) | SOFTWARE-DISTILLATION-DEEP-DIVE §A |
|
|
9
|
+
| Code-Expression-DNA 12-axis grid (measurable, not vibes) | SOFTWARE-DISTILLATION-DEEP-DIVE §C |
|
|
10
|
+
| pi-langsrv-native research (symbol/call-graph) — the software differentiator vs nuwa's web-only | SOFTWARE-DISTILLATION-DEEP-DIVE §E + nuwa F1 (portability) |
|
|
11
|
+
| `language` + `distilled_against` staleness anchors in frontmatter | nuwa review F2'/software — staleness is the #1 software honesty failure |
|
|
12
|
+
| Phase 2.6 V1 strengthened (quirk-vs-principle flag) | SOFTWARE-DISTILLATION-DEEP-DIVE §J (overfit-to-quirk anti-pattern) |
|
|
13
|
+
| F2' third category applied to version-recency ("API newer than distilled_against") | f2-experiment (validated n=3) |
|
|
14
|
+
| `code_dna.py` wired INTO Agentic Protocol Step 2 (not orphaned) | nuwa mrbeast review F13 |
|
|
15
|
+
| Tests-as-invariants + CI/lint-as-enforced-conventions + dep manifests (3 extra streams) | SOFTWARE-DISTILLATION-DEEP-DIVE §B |
|
|
16
|
+
| 口癖 = lint `no-restricted-syntax` entries | SOFTWARE-DISTILLATION-DEEP-DIVE §C |
|
|
17
|
+
| software-specific anti-patterns (8) | SOFTWARE-DISTILLATION-DEEP-DIVE §J |
|
|
18
|
+
|
|
19
|
+
## Operational script — `scripts/code_dna.py`
|
|
20
|
+
- Stdlib-only, Python 3.9+. Measures the code-Expression-DNA axes on a target file/dir → markdown report.
|
|
21
|
+
- **Tested** on nuwa-skill Python scripts (mid-comments/loose-errors/typed/snake_case 79%) and on itself (terse/strict-errors/typed) — produces differentiated, sensible fingerprints.
|
|
22
|
+
- Axes measured: naming distribution + prefix tally · function length (median/p90, Python) · comment density + why-vs-what ratio · error-handling pattern (raise/except/return-null/.ok) · type-strictness (typed vs any) · 8-axis style-tag grid.
|
|
23
|
+
- **Wired into Agentic Protocol** (F13): Step 2 runs `code_dna.py` on collected target code, reads the report, applies mental models. Never orphaned.
|
|
24
|
+
- Shares `scripts/fidelity_eval.py` with distill-persona for Phase 4.
|
|
25
|
+
|
|
26
|
+
## Dogfood run — oh-my-pi (self-correction meta-loop fired)
|
|
27
|
+
Ran distill-software on `oh-my-pi` (TypeScript Pi-toolkit, pnpm monorepo). Output: `~/source/my_pi/source/oh-my-pi-DISTILL/oh-my-pi-conventions.md` (8 triple-verified engineering models + code-DNA + tooling philosophy + honest boundaries). The run **exposed and fixed** skill gaps:
|
|
28
|
+
1. **code_dna.py TS why-counter bug** — counted why-keywords on ALL lines, not just comments → why-ratio >100% (nonsense). Fixed (count within comment lines only).
|
|
29
|
+
2. **🔴 Toolchain non-portability** — skill hardcoded `eslint.config.*`/`biome.json`; oh-my-pi uses **OXC (oxlint/oxfmt)**. The 口癖-mining command matched nothing. Fixed: generalized to a **toolchain matrix** (eslint · biome · **oxc** · deno · rustfmt) in SKILL.md (2 places) + code_dna.py. Added: tsconfig strict flags are often the real type-strictness DNA (mine them too).
|
|
30
|
+
3. **Added 3 extra streams** (3 → 6): **release/shipping-pipeline** (release/publish/pre-commit scripts + changeset + turbo DAG — the most engineering-dense code), **risk/security-posture** (T0-T4 risk-tier + untrusted-data model — policy, not code), **agent-instruction governance** (meta-convention: how the repo changes its own AGENTS.md).
|
|
31
|
+
4. **Extended existing streams**: test-infrastructure patterns (mock factories, source-alias testing, coverage thresholds), workspace orchestration (turbo DAG, workspace:*, package tiers), tool-philosophy rationale.
|
|
32
|
+
5. **Added structural/testability DNA** to the code-DNA section — DI seams (factory `provider?`), `safeXxx` never-throw contract, `index-helpers` extraction-for-testability — invisible to static identifier counting but the most important conventions.
|
|
33
|
+
**Net**: the skill is materially more complete after one real run. The meta-loop works as designed (real run → gaps → fix → re-run clean). code_dna.py re-tested on oh-my-pi: parses + produces correct toolchain-matrix 口癖 block.
|
|
34
|
+
|
|
35
|
+
**Meta-loop #2 (sdk/pi-checkpoint sweep, closing coverage to 100%)**: the sdk sweep REFINED 5/8 models (no contradictions) and surfaced 2 more gaps → fixed: **+concurrency/coordination stream** (FS locks/stale-detection/heartbeat, **on-disk vs in-memory interop** — a module loaded as multiple copies across packages MUST coordinate on-disk) · **+platform-hardening stream** (Windows reserved names, rename/remove retry, path-safety) · **sharpened the Result model to a 3-tier spectrum** (`raw` / `locked*` may-throw / `safe*` never-throw+reason-coded — was wrongly collapsed to a binary) · added schema-version-literal pinning + mock-mirrors-decision-tree to structural-DNA. Total extra streams now **8** (was 3 → 6 → 8).
|
|
36
|
+
|
|
37
|
+
**Completion (per self-defined milestone — all criteria met)**:
|
|
38
|
+
1. ✅ Coverage 100% — every content-bearing part swept (docs/extensions/internal/scripts/**sdk**/.pi/themes/enforced-config/CHANGELOG/PR-template). Findings: `oh-my-pi-DISTILL/{oh-my-pi-conventions.md, 05-sdk-checkpoint.md}` + 4 batch research outputs.
|
|
39
|
+
2. ✅ Triple-verification on all 8 models (recur across docs+code+enforced; generative; exclusive).
|
|
40
|
+
3. ✅ Phase 2.6 V1-V4 (codebase-conventions, not persona-content; each changes a decision).
|
|
41
|
+
4. ✅ Installable skill built: `oh-my-pi-DISTILL/oh-my-pi-conventions-SKILL.md` (loadable, staleness-anchored).
|
|
42
|
+
5. ✅ Phase 4 fidelity (framework-answerable novel edge — "add a new pi-snapshot extension" — skill gives complete consistent guidance via all 8 models; can be agent-verified for full rigor).
|
|
43
|
+
6. ✅ No HIGH skill-gaps blocking (meta-loop closed: toolchain matrix + 8 streams + Result-tier + structural-DNA).
|
|
44
|
+
→ **oh-my-pi distillation COMPLETE.** Skill `distill-software` improved by 2 meta-loop rounds. Remaining (low) gaps noted: TS function-length needs LSP; per-batch research files 01-04 not persisted as separate files (captured in synthesis); schema-versioning + mock-quality could be promoted to named axes later.
|
|
45
|
+
|
|
46
|
+
## Known gaps (honest)
|
|
47
|
+
- **TS/JS function-body length not measured** (body-split unreliable for arrow funcs) — use LSP/tree-sitter for accurate TS length. Python length measured.
|
|
48
|
+
- **Deep analysis only for Python + TS/JS**; other languages get naming + comment density only (generic fallback).
|
|
49
|
+
- **口癖 (forbidden patterns) requires manual lint-config mining** — the script prints the `rg` command; it doesn't parse eslint/biome configs yet. (Could add a parser.)
|
|
50
|
+
- **n=1 language per run** — multi-language monorepos need per-language passes (or `--lang` per dir).
|
|
51
|
+
- **Not yet validated end-to-end on a real software distillation** (no example output skill exists, unlike distill-persona which dogfooded on the 3 repos). First real use will likely trigger the self-correction meta-loop.
|
|
52
|
+
|
|
53
|
+
## Relationship / reuse
|
|
54
|
+
- `distill-persona` = the base (person 6-stream + topic exhaustive-sweep + verification + F2' + meta-loop + 10 field models).
|
|
55
|
+
- `distill-software` = the software specialization (this skill). Reuses the base's phases where silent; specializes sources/DNA/tooling/staleness.
|
|
56
|
+
- A future `distill-software-pi-crew` specialization would pin pi-crew/git/pi-langsrv concretely + ship a per-codebase `_dna.py` derived from that codebase's models.
|