pi-crew 0.9.48 → 0.9.50

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (86) hide show
  1. package/AGENTS.md +18 -0
  2. package/CHANGELOG.md +314 -0
  3. package/dist/build-meta.json +70 -42
  4. package/dist/index.mjs +505 -430
  5. package/dist/index.mjs.map +4 -4
  6. package/docs/decisions/2026-07-24-oidc-trusted-publishing.md +112 -0
  7. package/package.json +2 -3
  8. package/skills/.gitkeep +0 -0
  9. package/skills/distill-persona/BUILD-NOTES.md +55 -0
  10. package/skills/distill-persona/SKILL.md +550 -0
  11. package/skills/distill-persona/UPGRADE-LOG-RESEARCH-SKILLS.md +100 -0
  12. package/skills/distill-persona/references/coverage-manifest.md +65 -0
  13. package/skills/distill-persona/references/cross-skill-differentiation.md +12 -0
  14. package/skills/distill-persona/references/description-discipline.md +6 -0
  15. package/skills/distill-persona/references/diagnostic-path.md +25 -0
  16. package/skills/distill-persona/references/distillation-field-synthesis-pass2.md +59 -0
  17. package/skills/distill-persona/references/distillation-field-synthesis.md +108 -0
  18. package/skills/distill-persona/references/fidelity-rubric.md +19 -0
  19. package/skills/distill-persona/references/field-models.md +20 -0
  20. package/skills/distill-persona/references/handoff.md +42 -0
  21. package/skills/distill-persona/references/optional-body-sections.md +9 -0
  22. package/skills/distill-persona/references/registry-routing.md +11 -0
  23. package/skills/distill-persona/references/research/lesson-memory-shortcut.md +33 -0
  24. package/skills/distill-persona/references/research/r1-a-examples.md +23 -0
  25. package/skills/distill-persona/references/research/r1-b-scripts.md +26 -0
  26. package/skills/distill-persona/references/research/r1-c-human-readme.md +31 -0
  27. package/skills/distill-persona/references/research/r1-d-tests.md +28 -0
  28. package/skills/distill-persona/references/research/r1-verification.md +36 -0
  29. package/skills/distill-persona/references/research/r2-low-yield.md +26 -0
  30. package/skills/distill-persona/references/self-upgrade-directive.md +20 -0
  31. package/skills/distill-persona/references/taste-principles.md +8 -0
  32. package/skills/distill-persona/references/topic-variant.md +13 -0
  33. package/skills/distill-persona/references/update-mode.md +7 -0
  34. package/skills/distill-persona/scripts/fidelity_eval.py +244 -0
  35. package/skills/distill-persona/scripts/validate-run.mjs +297 -0
  36. package/skills/distill-persona/scripts/validate-skill-structure.mjs +177 -0
  37. package/skills/distill-software/BUILD-NOTES.md +56 -0
  38. package/skills/distill-software/SKILL.md +363 -0
  39. package/skills/distill-software/references/handoff.md +47 -0
  40. package/skills/distill-software/scripts/code_dna.py +290 -0
  41. package/skills/research/DISTILLATION-PROCESS-CHECKLIST.md +120 -0
  42. package/skills/research/EXCAVATION-CHECKLIST.md +142 -0
  43. package/skills/research/FIDELITY.md +180 -0
  44. package/skills/research/SKILL.md +432 -0
  45. package/skills/research/references/anti-patterns.md +184 -0
  46. package/skills/research/references/fidelity.md +241 -0
  47. package/skills/research/references/handoff.md +48 -0
  48. package/skills/research/references/research-protocol.md +162 -0
  49. package/skills/research/references/source-inventory.md +135 -0
  50. package/skills/research/references/verified-models.md +163 -0
  51. package/skills/research/scripts/__pycache__/safe_io.cpython-312.pyc +0 -0
  52. package/skills/research/scripts/code_dna.py +233 -0
  53. package/skills/research/scripts/emit_run_summary.py +142 -0
  54. package/skills/research/scripts/safe_io.py +314 -0
  55. package/skills/research/scripts/source_evaluator.py +234 -0
  56. package/skills/research/scripts/validate-skill-structure.mjs +177 -0
  57. package/skills/research/scripts/verify_citations.py +225 -0
  58. package/skills/security-priority.json +28 -0
  59. package/src/config/config.ts +1 -0
  60. package/src/config/role-tools.ts +6 -3
  61. package/src/config/types.ts +8 -0
  62. package/src/extension/crew-cleanup.ts +18 -1
  63. package/src/extension/crew-vibes/index.ts +11 -2
  64. package/src/extension/register.ts +1 -1
  65. package/src/extension/registration/command-registration.ts +1 -0
  66. package/src/extension/registration/commands.ts +7 -3
  67. package/src/extension/registration/lifecycle-handlers.ts +1 -3
  68. package/src/extension/registration/ui.ts +4 -0
  69. package/src/extension/registration/viewers.ts +3 -0
  70. package/src/extension/team-tool/run.ts +7 -6
  71. package/src/runtime/background-runner.ts +11 -16
  72. package/src/runtime/chain-runner.ts +3 -2
  73. package/src/runtime/heartbeat-watcher.ts +28 -1
  74. package/src/runtime/pipeline-runner.ts +8 -7
  75. package/src/runtime/task-runner.ts +165 -119
  76. package/src/schema/config-schema.ts +1 -0
  77. package/src/ui/live-run-sidebar.ts +2 -0
  78. package/src/ui/mascot.ts +11 -9
  79. package/src/ui/render-coalescer.ts +9 -0
  80. package/src/ui/run-snapshot-cache.ts +10 -11
  81. package/src/ui/terminal-status.ts +5 -0
  82. package/src/ui/widget/index.ts +3 -5
  83. package/src/ui/widget/widget-types.ts +0 -1
  84. package/src/utils/gh-protocol.ts +9 -8
  85. package/workflows/distill.workflow.md +198 -0
  86. package/assets/runner-spritesheet.png +0 -0
@@ -0,0 +1,100 @@
1
+ # UPGRADE-LOG — research-skill self-upgrade candidates applied to the distill skills
2
+
3
+ > Companion to `source/RESEARCH-SKILLS-SELFUPGRADE-CANDIDATES.md` (the 12-candidate artifact). This log records the EFFECTIVENESS-VERIFICATION gate verdict (concrete-delta + proof + conflict-check + verdict) for each candidate, and what was actually applied.
4
+ >
5
+ > Date: 2026-07-24 · Targets: `skills/distill-persona/SKILL.md`, `skills/distill-software/SKILL.md` (+ 2 new `references/handoff.md`).
6
+
7
+ ## Effectiveness-gate legend (per distill-software Phase 2.6 EFFECTIVENESS VERIFICATION)
8
+ - **CONCRETE DELTA** — exactly what changes in target.
9
+ - **EFFECTIVENESS PROOF** — GENERATIVE (changes a real decision) / PROBLEM-EXISTS (target has the gap) / DELTA-TEST.
10
+ - **CONFLICT CHECK** — does it clash with an existing practice?
11
+ - **VERDICT** — ✅ APPLIED · ❌ SKIPPED (+ reason).
12
+
13
+ ---
14
+
15
+ ## APPLY (high-value, low-risk)
16
+
17
+ ### #1 verify_citations rigor — ✅ APPLIED → distill-persona Phase 2.6 (new V5)
18
+ - **Concrete delta**: Phase 2.6 gains a **V5 — Source/citation verified** gate (persona analog of distill-software's grep-V5): every cited source exists in the fetched pool, no invented URLs, no dangling `[n]`, source concentration ≤25%. `skills/research/scripts/verify_citations.py` referenced as an OPTIONAL aid (not ported — the script already lives in the research skill).
19
+ - **Proof**: PROBLEM-EXISTS — V1–V4 covered signal/redundancy/effective/optimal but NOT source-integrity; invented URLs and over-concentrated sourcing were caught only by hand. verify_citations.py exists and is runnable (`skills/research/scripts/verify_citations.py`).
20
+ - **Conflict**: none — V5 sits after V4; V1–V4 unchanged.
21
+ - **Verdict**: ✅ APPLIED (surgical: 1 bullet, references existing script rather than porting it).
22
+
23
+ ### #2 tension-discovery 3-probe — ✅ APPLIED → distill-persona Phase 2 (内在张力)
24
+ - **Concrete delta**: the 内在张力 extraction line gains an active-probing checklist: discover tensions via 3 named probes (concept-confusion / assumption-check / effect-vs-mechanism) instead of reactively recording whatever surfaced.
25
+ - **Proof**: PROBLEM-EXISTS — current tensions were post-hoc (recorded, not sought). Probe 2 ("what is everyone assuming?") would have caught model drift earlier.
26
+ - **Conflict**: none — inline prose (parenthetical (1)/(2)/(3)), no new section; M9a format preserved. The source `tension-discovery.md` was NOT copied — only the 3-probe names integrated.
27
+ - **Verdict**: ✅ APPLIED (1 inline addition; no file copy).
28
+
29
+ ### #7 emit_run_summary observability — ✅ APPLIED → BOTH skills Phase 4 (FIDELITY.md)
30
+ - **Concrete delta**: FIDELITY.md persistence gains a **run observability** field (wall-clock time, token count, cost tier). `skills/research/scripts/emit_run_summary.py` referenced as optional aid in BOTH skills.
31
+ - **Proof**: PROBLEM-EXISTS — FIDELITY.md tracked scores + models but no cost; run-comparison was impossible (Run #2's 11.27M-token blowout was detectable only with such observability).
32
+ - **Conflict**: none — additive to the existing FIDELITY.md spec.
33
+ - **Verdict**: ✅ APPLIED (mirrored to both skills; symmetry grep-verified: 1 occurrence each).
34
+
35
+ ### #9 handoff-format protocol — ✅ APPLIED → BOTH skills
36
+ - **Concrete delta**: new `references/handoff.md` in EACH skill (concise ~40-line protocol adapted to distillation, not generic research) + 1-line pointer in each SKILL.md (distill-persona at the context-window guard; distill-software at Phase 4 fidelity).
37
+ - **Proof**: PROBLEM-EXISTS — handoff was a single sentence ("segment across sessions"); a fresh session had no contract for lossless resume.
38
+ - **Conflict**: none — the 2-file state pattern (M#3) is preserved; handoff.md is an index over existing state files, not a new mechanism.
39
+ - **Verdict**: ✅ APPLIED (mirrored; both handoff.md files created; both SKILL.md have 1 pointer each).
40
+
41
+ ### #12 "Removed" changelog tracking — ✅ APPLIED → distill-persona Update mode
42
+ - **Concrete delta**: Update mode gains a **Deletion tracking** rule — log what was removed (`### Removed`) alongside additions; "no removals" stated explicitly if none.
43
+ - **Proof**: PROBLEM-EXISTS — update mode only ever grew the skill; stale heuristics never pruned → drift (the foundry distill-skills grew without deletion logs).
44
+ - **Conflict**: none — additive to Update mode; no-op detection preserved.
45
+ - **Verdict**: ✅ APPLIED (1 paragraph).
46
+
47
+ ---
48
+
49
+ ## APPLY CAREFULLY (medium, kept concise)
50
+
51
+ ### #4 anti-thrash convergence gate — ✅ APPLIED → distill-persona Phase 2 (diminishing-returns gate)
52
+ - **Concrete delta**: the 3-empty-rounds gate gains an **active anti-thrash nudge**: if ≥3 consecutive rounds add only 0–1 marginal findings (low-yield but not zero), switch to a structurally different approach BEFORE the passive 3-empty gate fires.
53
+ - **Proof**: PROBLEM-EXISTS — Run 2's 11.2M-token spiral was exactly grinding low-yield variations without structural change.
54
+ - **Conflict**: none — the 3-empty-rounds gate is preserved; this is a within-round intervention that fires earlier (low-yield) than the terminal gate (zero). NOT a full hook port — integrated as 1 sentence.
55
+ - **Verdict**: ✅ APPLIED (1 sub-step; no `references/anti-thrash.md` companion — kept inline).
56
+
57
+ ### #6 hypothesis-reflection — ✅ APPLIED → distill-persona Phase 2.6 V5
58
+ - **Concrete delta**: V5 gains a cheap-model meta-critique step — before a costly re-extract on V5 failure, ask a lighter model to critique the failure pattern and propose an adjacent direction.
59
+ - **Proof**: PROBLEM-EXISTS — the M#2 V5-rejection (4-way → 3-way) was a meta-pattern (over-claiming); a cheap critique would have caught the over-claim before the expensive re-extract.
60
+ - **Conflict**: none — V5 preserved; the critique is upstream of re-extraction.
61
+ - **Verdict**: ✅ APPLIED (folded into the V5 bullet with #1; 1 sentence).
62
+
63
+ ---
64
+
65
+ ## SKIP (justified)
66
+
67
+ ### #5 context-rotation — ❌ SKIPPED
68
+ - **Reason**: pi-autoresearch's `.auto/` on-disk state layout is NOT the distill skills' model. The distill skills persist state to `references/research/` + checklists, not a single growing `prompt.md`. Context pressure is already handled by the context-window guard (F6) + session segmentation (#9 handoff).
69
+ - **Effectiveness proof fails**: PROBLEM-EXISTS is weak — there is no `.auto/prompt.md` analog to rotate. The real fix for context pressure is already present (segmentation + handoff).
70
+ - **Conflict**: would introduce an alien state model. **Verdict**: ❌ SKIPPED.
71
+
72
+ ### #8 5-tier ship-gate — ❌ SKIPPED (folded lightly)
73
+ - **Reason**: distill-software already has a working single ship-gate (the all-green checklist with 10 items). Restructuring into a 5-tier Gate 1–5 model would risk the mature validator's grep checks and the existing gate contract for marginal cost-awareness gain.
74
+ - **Folded lightly**: added 1-line **Tiered effort** note — trivial distillations may skip the costliest sub-step (independent dual-agent re-score → self-score with caveat), but never the structural/V5-grep/coverage gates.
75
+ - **Effectiveness proof**: PARTIAL — the 5-tier model's value (skip expensive gates for trivial output) is captured by the 1-line note without the risky restructuring.
76
+ - **Conflict**: a full restructure would break cross-references. **Verdict**: ❌ SKIPPED-restructure / ✅ folded-1-line.
77
+
78
+ ### #10 hard-constraint block — ❌ SKIPPED
79
+ - **Reason**: minor convention; the generated-skill template already enforces its invariants via the validator (validate-skill-structure.mjs asserts mandatory fields, Agentic Protocol, no placeholders). A prose "Hard Constraint:" block duplicates what the validator machine-enforces.
80
+ - **Effectiveness proof fails**: PROBLEM-EXISTS is weak — the validator IS the hard-constraint enforcement; a prose block adds ceremony, not rigor.
81
+ - **Conflict**: none, but low value. **Verdict**: ❌ SKIPPED.
82
+
83
+ ### #11 P0–P6 heading rename — ❌ SKIPPED (RISKY)
84
+ - **Reason**: renaming `## Phase N` → `## P0–P6` is purely cosmetic AND would break (a) the validator's Phase-based grep assertions, (b) cross-references throughout both skills and the knowledge base that match "Phase N", (c) the self-upgrade directive text that references phases by number.
85
+ - **Effectiveness proof fails**: PROBLEM-EXISTS is cosmetic (readability), not functional. The conflict risk (breaking validator + cross-refs) vastly outweighs the readability gain.
86
+ - **Conflict**: HIGH — breaks validator + cross-references. **Verdict**: ❌ SKIPPED.
87
+
88
+ ---
89
+
90
+ ## Candidate #3 (source_evaluator) — note
91
+ The artifact lists #3 (source_evaluator.py → distill-software Phase 2.1). It was NOT in the APPLY/APPLY-CAREFULLY triage for this run (the task scoped the distill-software changes to #7, #9, #8-fold only). #3 remains a future candidate — it would add a source-quality gate to the decisions stream, but porting the script + integrating into Phase 2.1 exceeds the "surgical" bar set for this pass.
92
+
93
+ ---
94
+
95
+ ## Verification evidence
96
+ - `distill-persona` validator: **12 pass, 9 fail** — IDENTICAL to pre-edit baseline. The 9 failures are all inherent to the engine skill being a *template* (placeholders `<person>`/`<target>`; no generated artifacts FIDELITY.md / EXCAVATION-CHECKLIST.md / DISTILLATION-PROCESS-CHECKLIST.md — those are produced per GENERATED skill, not the engine). **No new failures introduced.**
97
+ - `research` validator (reference): **30 pass, 0 fail — ✅ ALL-GREEN** (unchanged; research skill not edited).
98
+ - Symmetry grep: `emit_run_summary.py` = 1 in each SKILL.md; `references/handoff.md` = 1 pointer in each SKILL.md; both handoff.md files created.
99
+ - Phase numbering / heading scheme: UNCHANGED (constraint #2 honored).
100
+ - distill-persona M9a count (内在张力 ≥3), honest-boundaries (≥3), M12 fallback (7 rows): all UNCHANGED — no edit disturbed validator counts.
@@ -0,0 +1,65 @@
1
+ # Coverage Manifest — distilling the 3 source projects (exhaustive sweep mode)
2
+
3
+ > Contract for the "miss nothing" guarantee (distill-persona Phase 1 project/topic mode). Every content-bearing part must reach COVERED with a recorded contribution. Round count scales with size; diminishing-returns gate bounds it. Status as of round-1 dispatch.
4
+
5
+ ## nuwa-skill (the engine + 15 examples + scripts + governance)
6
+
7
+ | Part | Status | Contribution | Round |
8
+ |------|--------|--------------|-------|
9
+ | `SKILL.md` (engine, 40KB) | COVERED | the 6-phase methodology; all core models M1-M5 | R0 |
10
+ | `references/extraction-framework.md` | COVERED | triple-verification + expression-DNA quant + contradiction rules | R0 |
11
+ | `references/skill-template.md` | COVERED | the output template | R0 |
12
+ | `references/fidelity-scorecard.md` | COVERED | 5-dim QA gate (F2/F2' source) | R0 |
13
+ | `examples/andrej-karpathy-perspective` (SKILL+FIDELITY+6 research) | COVERED | engineer-persona exemplar; agentic protocol; F2' subject | R0 |
14
+ | `examples/ilya-sutskever-perspective` | COVERED | F2' subject; refusal-vocab edge handling | R0 |
15
+ | `examples/paul-graham-perspective` | COVERED | large-corpus distillation; F2' subject | R0 |
16
+ | `examples/munger-perspective` | COVERED | F2' subject; death-as-staleness-anchor | R0 |
17
+ | `examples/mrbeast-perspective` (+scripts) | COVERED | operational-tooling distillation; F13 | R0 |
18
+ | `examples/x-mastery-mentor` | COVERED | topic-skill variant; S8/S9/S10; G1 defect | R0 |
19
+ | `examples/steve-jobs-perspective` | UNCOVERED | R1 |
20
+ | `examples/taleb-perspective` | UNCOVERED | R1 |
21
+ | `examples/feynman-perspective` | UNCOVERED | R1 |
22
+ | `examples/zhangxuefeng-perspective` | UNCOVERED | R1 |
23
+ | `examples/trump-perspective` | UNCOVERED | R1 |
24
+ | `examples/zhang-yiming-perspective` | UNCOVERED | R1 |
25
+ | `examples/sun-yuchen-perspective` | UNCOVERED | R1 |
26
+ | `examples/elon-musk-perspective` | UNCOVERED | R1 |
27
+ | `examples/naval-perspective` | UNCOVERED | R1 |
28
+ | root `scripts/{quality_check,merge_research,srt_to_transcript}.py` + `download_subtitles.sh` | UNCOVERED (deep) | R1 |
29
+ | `scripts/community_check.py` + `.github/workflows/community-pr-check.yml` | UNCOVERED | R1 |
30
+ | `COMMUNITY.md` + `CONTRIBUTING.md` | UNCOVERED | R1 |
31
+ | `README*.md` (5 lang) | LOW-YIELD (marketing) | sampled |
32
+
33
+ ## awesome-human-distillation (index + autocuration)
34
+
35
+ | Part | Status | Contribution | Round |
36
+ |------|--------|--------------|-------|
37
+ | structure + 4 scripts + 3 workflows (gestalt) | COVERED | auto-curation pipeline; M-F1 | R0 |
38
+ | `README.md` (76KB, ~210 entries) — **per-entry taxonomy sweep** | UNCOVERED (deep) | R1 |
39
+ | `README_EN.md` | UNCOVERED | R1 |
40
+
41
+ ## awesome-persona-distill-skills (index + submission automation)
42
+
43
+ | Part | Status | Contribution | Round |
44
+ |------|--------|--------------|-------|
45
+ | README + 3 scripts + 7 workflows (gestalt) | COVERED | submission-automation; M-F1 | R0 |
46
+ | `test/*.test.mjs` (5 files) | UNCOVERED | R1 |
47
+ | `CONTRIBUTING*.md` + ISSUE_TEMPLATE | COVERED (gestalt) | submission gates | R0 |
48
+
49
+ ## Round plan
50
+ - **R1** (this dispatch, parallel): 4 agents — (A) 9 uncovered example SKILL.md; (B) nuwa root+community scripts; (C) awesome-human 76KB README per-entry; (D) awesome-persona 5 tests.
51
+ - **R1.5 checkpoint**: triple-verify new contributions; detect diminishing returns; if 9 examples all reinforce existing models → batch-shrink remaining + flag gate.
52
+ - **R2+** if needed: any part still UNCOVERED after R1.
53
+ - **Stop**: manifest 100% COVERED OR gate fires with sampled confirmation.
54
+
55
+ ## R1 results (coverage update — exhaustive sweep executed)
56
+
57
+ All 4 R1 batches COVERED with recorded contributions in `references/research/r1-{a,b,c,d}.md`:
58
+ - **r1-a** (9 nuwa examples): COVERED. 25 candidate new elements; ~8 genuinely new persona-heuristics (idiot-index, desire-as-contract, dual-mode, ghost-mode, etc.). **Diminishing-returns GATE FIRED** — yield dropped 2.7→1.0 new/example after example 5; core methodology fully confirmed; deeper example-sweeping has low marginal yield.
59
+ - **r1-b** (nuwa scripts): COVERED. KEY: quality_check.py is structural-only → strengthens M-F4 (structural-pass ≠ behavioral edge-honesty). + contradiction-as-signal technique + source-type hierarchy + dissemination dual-gate.
60
+ - **r1-c** (human README 76KB): COVERED. **Empirically corrected M-F2** (gestalt was wrong: 84% public-figures, bimodal; + meta-tier, living-creator, commemorative-split, adversarial). + "4D distillation" candidate.
61
+ - **r1-d** (persona tests): COVERED. **NEW model M-F7** (structural invariants = index-scale quality gate).
62
+
63
+ **R2 (low-yield tail — actually swept, not assumed):** nuwa COMMUNITY.md + CONTRIBUTING.md + README_EN + 3 translations(JA/KO/ES sampled), awesome-human README_EN.md, awesome-persona CONTRIBUTING.md — all COVERED in `references/research/r2-low-yield.md`. **Result: all CONFIRM existing models (M-F1/M-F2/M-F3/M-F4/M-F7); ZERO new models; diminishing-returns gate FIRED a 2nd time.** 3 minor observations (multi-persona orchestration, i18n-as-curation, methodology-versioning) checked V1-V4 → NOTES only, not models.
64
+
65
+ **Verdict**: exhaustive sweep **100% COMPLETE** — every content-bearing part of all 3 projects COVERED with recorded contribution. Two independent gate-firings (R1 examples-after-#5; R2 tail) confirm the field surface is exhausted. The skill's 10 models (5 core M1-M5 + 5 field M-F1..F4/M-F7) are the stable converged set. The "miss nothing" contract is satisfied; the gate prevents over-extraction beyond it. No gestalt anywhere — every part examined in detail.
@@ -0,0 +1,12 @@
1
+ ## Phase 2.7 — Cross-skill differentiation (anti-overlap)
2
+
3
+ Before building, scan `<skill-dirs>/*-perspective/` for existing skills. If the new subject overlaps an existing one (e.g. Musk + Thiel + Naval all share first-principles thinking):
4
+
5
+ - **Model overlap >40%** → flag: either MERGE into the existing skill's "related lenses" section, or EXPLICITLY differentiate (what does THIS lens add the other doesn't?).
6
+ - **>60% overlap with only the name changed** → this is **thin repackaging** (anti-pattern #11). Stop; differentiate or pick a different target.
7
+ - Record the differentiation decision in `references/differentiation-check.md`.
8
+
9
+ This prevents shipping near-duplicate skills that dilute the registry.
10
+
11
+ ---
12
+
@@ -0,0 +1,6 @@
1
+ ### Description discipline (F4/F5/F15)
2
+ - `description` = ONE positioning sentence. `triggers` = 2–4 explicit phrases + the person's name. **No long-tail keyword stuffing** (it inflates false-triggers and, in crowded skill sets, collides).
3
+ - Before writing triggers, scan installed skills for collision; disambiguate if a trigger overlaps an existing skill.
4
+ - **Structure**: say WHAT it IS before WHAT it DOES — "A thinking-advisor skill that…" not "Helps you think like…" (F15).
5
+ - **Terminal punctuation**: normalize `description` + identity-card sentence for terminal punctuation (。/!/? for zh, ./?/! for en). Run as a post-processing step after writing SKILL.md (F3).
6
+
@@ -0,0 +1,25 @@
1
+ ### Phase 0B — Diagnostic path (when user has a vague need, not a target)
2
+
3
+ User doesn't know who/what to distill, only has a need or problem. Route:
4
+ 1. **1-2 clarifying questions** to locate the need dimension. Match against:
5
+ | Need dimension | Typical expression | Thinking-framework direction |
6
+ |---|---|---|
7
+ | Decision & judgment | "how to decide better" | Multi-model thinking, inversion, probabilistic |
8
+ | Expression & writing | "can't explain clearly" | Feynman simplification, storytelling, analogy |
9
+ | Business & startup | "can't find PMF" | First principles, leverage, product restraint |
10
+ | Teaching & communication | "students don't get it" | Known-to-unknown, metaphor, minimum viable knowledge |
11
+ | Critical thinking | "always getting fooled" | Falsification, evolutionary lens, cognitive-bias ID |
12
+ | Content creation | "no views on videos" | Attention engineering, test-iterate, audience psych |
13
+ | Life strategy | "career lost, anxious" | Long-termism, leverage selection, compounding |
14
+ | Risk & uncertainty | "how to handle black swans" | Antifragility, convexity, tail-risk management |
15
+ | Design & product | "UX is bad, can't simplify" | Minimalism, user mental models, constraint-as-creativity |
16
+ | Humor & expressiveness | "too serious, not interesting" | Absurd contrast, expectation violation, self-deprecating authority |
17
+ 2. **Recommend 2-3 candidates** from DUAL SOURCES:
18
+ - **Source A**: scan `<skill-dirs>/*-perspective/` for EXISTING skills matching the need → mark ⚡ (plug-and-play, zero cost).
19
+ - **Source B**: propose NEW distillation targets matching the need → mark 🆕.
20
+ - Each candidate: `### Candidate: [name] ⚡/🆕` + **core lens** (1 sentence) + **why it fits your need** + **limitation** (what it CAN'T help with).
21
+ - Principles: ≤3 candidates (choice paralysis is worse than no choice); existing skills first; candidates must differ from each other; always state limitations.
22
+ 3. User selects → proceed to Phase 0A → Phase 0.5.
23
+
24
+ > Principle: max 2 rounds of questions. If the user's need is already clear, recommend directly.
25
+
@@ -0,0 +1,59 @@
1
+ # Distillation-Field Synthesis — PASS 2 (independent re-derivation + reproducibility test)
2
+
3
+ > Second independent distillation pass over the same 3 source projects, run to (a) go deeper and (b) **dogfood the reproducibility gap (F10)** — does a second pass converge on pass 1's models, or surface different/additional ones? This is exactly what `fidelity_eval.py repro` exists to measure.
4
+ >
5
+ > Pass 1 artifact: `distillation-field-synthesis.md` (models M1-M5 + M-F1/F2/F3).
6
+ > Method: re-derive from the raw FINDINGS without re-reading pass 1's verdicts, then compare.
7
+
8
+ ---
9
+
10
+ ## Reproducibility result (the F10 dogfood)
11
+
12
+ | | Pass 1 model set | Pass 2 verdict |
13
+ |---|---|---|
14
+ | M1 HOW-not-WHAT | ✓ | **reproduced** |
15
+ | M2 research→extract→validate→generate | ✓ | **reproduced** |
16
+ | M3 honest boundaries mandatory | ✓ | **reproduced** |
17
+ | M4 quality gated | ✓ | **reproduced** |
18
+ | M5 self-contained portability | ✓ | **reproduced** |
19
+ | M-F1 dissemination flywheel | ✓ | **reproduced** |
20
+ | M-F2 target taxonomy | ✓ | **reproduced** |
21
+ | M-F3 ethics spectrum | ✓ | **reproduced** |
22
+ | **A — field over-claims its own fidelity** | — | **NEW ★★** |
23
+ | **B — dissemination is meme/virality-gated** | — | **NEW ★** |
24
+ | **C — distillation must carry its own critique** | (was heuristic C11) | **promoted to model** |
25
+
26
+ **Reproducibility score**: core model-set overlap = **8/8 (100%)** — pass 2 independently re-derived every pass-1 model. Pass 2 is a **strict superset** (+3). This is the *good* reproducibility outcome (convergent core, deeper pass enlarges surface) and it directly calibrates the F10 gap: a distillation is reproducible on its core claims; a *deeper* second pass legitimately adds. The danger case (two passes <60% overlap) did NOT occur — which is itself the signal that the methodology's Phase 2 (triple-verification) is stable, not coin-flip.
27
+
28
+ ---
29
+
30
+ ## Pass-2 NEW models (the deeper surface)
31
+
32
+ ### M-F4 — The field systematically over-claims its own fidelity ★★ (standout — integrate)
33
+ **One-line**: every published distillation fidelity score is an upper bound; the field's self-assessment is inflated by easy test questions + author-side confirmation bias, not scorer contamination.
34
+ **Evidence (cross-source ≥2 + experiment)**: nuwa publishes 94-97/100 for its skills; both awesome-lists propagate those scores as quality signals; the F2' experiment independently re-scored 3 nuwa skills → **blind 67-76, edge-honesty 20→7/13/7** (published vs observed). The inflation mechanism (validated at n=3) is *question-design* (edge questions sat too close to documented material; framework-answerable edges weren't tested), NOT scorer skill-file access (F2 refuted).
35
+ **Application**: treat ANY published distillation score as an optimistic ceiling. Re-score independently with a framework-answerable novel edge before trusting. A skill that "scores 97" but can't flag inference on a novel in-domain question is not ship-ready.
36
+ **Limitation**: this is a critique of *current practice*, not a law — a properly-tested scorecard (with framework-answerable edges + independent blind scorer) could be trustworthy. The field just hasn't been doing that.
37
+ **Why pass 1 missed it**: pass 1 captured F2' as a *skill-internal fix* (Phase 4 gate). Pass 2 elevates it to a *field-level model* — the lesson is about the practice, not just our engine.
38
+
39
+ ### M-F5 — Dissemination is meme/virality-gated *(B — note, lighter integrate)*
40
+ **One-line**: distillation skills spread by meme-fit, not quality alone — the field's genesis is a single viral meme (colleague-skill) and both indexes rode it.
41
+ **Evidence**: both awesome-lists trace the wave to `titanwings/colleague-skill`; nuwa ships 5-language READMEs + promo art (67MB of marketing) for reach; the "distill your boss/ex/deceased" framing is the viral hook. Quality and reach are decoupled — `vengeful-ghost-skill` (satire) and `anti-distill` spread alongside serious engines.
42
+ **Application**: a distillation's adoption is gated by narrative/meme-fit as much as fidelity. Plan dissemination (M-F1) with that in mind — don't confuse virality with validation.
43
+ **Limitation**: meme-driven ≠ low-quality; it just means quality isn't the bottleneck for spread.
44
+
45
+ ### M-F6 — Distillation must carry its own critique *(C — promoted from pass-1 heuristic C11)*
46
+ **One-line**: a healthy distillation practice preserves its skeptics — anti-distill entries, honesty machinery, declared limits — rather than purging them.
47
+ **Evidence**: awesome-human-list *includes* `anti-distill` + `vengeful-ghost-skill` as entries; nuwa's anti-pattern blacklist + honest-boundaries; the F2' edge-honesty gate. The critique IS part of the artifact.
48
+ **Application**: when building a distillation registry/skill, keep the skeptical entries and the "what this can't do" sections — they're load-bearing, not noise.
49
+ **Limitation**: critique-as-content can drift into performance (satire that signals sophistication without substance).
50
+
51
+ ---
52
+
53
+ ## Phase 3 — Integration (pass 2)
54
+ **M-F4** (field over-claims fidelity) integrates into `distill-persona/SKILL.md` Field-models section — it's the standout meta-lesson and reframes the existing F2' fix as a field-aware stance. M-F5/M-F6 noted here; M-F6 already implied by the skill's honesty machinery, M-F5 is dissemination-context (covered by M-F1's implication) — no separate SKILL.md bloat.
55
+
56
+ ## Phase 4 — self-check
57
+ - ✅ M-F4/M-F5/M-F6 each pass triple-verification (evidence cited, ≥2 sources).
58
+ - ✅ Reproducibility core = 8/8 — methodology stable.
59
+ - ⚠️ Honest limit: pass 2 was done by the same mind that did pass 1 (not truly independent agents). A rigorous F10 test would use a fresh-context agent for pass 2. The 100% core convergence is therefore *suggestive*, not proof. (The fidelity_eval repro harness measures this properly when two *files* exist; here both passes fold into one skill, so the comparison is analytical, not automated.)
@@ -0,0 +1,108 @@
1
+ # Distillation-Field Synthesis — topic distillation of the 3 source projects
2
+
3
+ > **Dogfooding artifact**: the `distill-persona` skill applied to its own source material. Topic flavor (extraction-framework §5): distill the *field* of "agent-skill distillation" from the 3 source projects, not one persona.
4
+ >
5
+ > **Sources** (P1 = `nuwa-skill` the engine; P2 = `awesome-human-distillation` index+autocuration; P3 = `awesome-persona-distill-skills` index+submission-automation).
6
+ > **Phase 1 (research)**: already done — `source/{nuwa-skill,awesome-human-distillation,awesome-persona-distill-skills}-FINDINGS.md` + reviews + the F2' experiment.
7
+ > **Method**: Phase 2 triple-verification (cross-*source* recurrence ≥2 projects + generative + exclusive) on every candidate claim.
8
+
9
+ ---
10
+
11
+ ## Phase 2 — Triple-verification table
12
+
13
+ Candidate claims drawn from the 3 projects, tested (✅=passes, ⚠️=partial, ❌=fails → demote/discard).
14
+
15
+ | # | Candidate claim | Cross-source (≥2 projects) | Generative | Exclusive | Verdict |
16
+ |---|---|---|---|---|---|
17
+ | C1 | "Distill HOW they think, not WHAT they said" | ✅ P1 core principle + P2/P3 definitions ("expressive style, decision frameworks, interaction patterns") | ✅ predicts a good skill captures models not quotes | ✅ defining distinction of distillation | **MODEL** |
18
+ | C2 | "Distillation = research→extract→validate→generate pipeline" | ✅ P1 6-phase; P2/P3 link projects following it | ✅ | ✅ | **MODEL** |
19
+ | C3 | "Honest boundaries mandatory, not optional" | ✅ P1 honest-limits+edge-honesty; P2/P3 ethical guardrails in definitions | ✅ | ✅ | **MODEL** |
20
+ | C4 | "Distill → awesome-list → auto-curation = dissemination flywheel" | ✅ **P2 AND P3 both build issue→PR→merge pipelines**; P1 feeds both | ✅ predicts a mature practice needs an indexing/dissemination layer | ✅ | **MODEL — NEW** ★ |
21
+ | C5 | "Target taxonomy: self → relationships → public-figures → fields" | ✅ P3 5 categories + P2 6 relationship categories + P1 person-vs-topic | ✅ predicts what's distillable | ✅ | **MODEL — NEW** ★ |
22
+ | C6 | "Quality is gated, not assumed" | ✅ P1 fidelity-scorecard+triple-verification; P2/P3 submission quality gates + consistency checks | ✅ | ✅ | **MODEL** |
23
+ | C7 | "Self-contained portability (copy dir → runs)" | ✅ P1; P2/P3 skills are standalone repos | ✅ | ✅ | **MODEL** |
24
+ | C8 | "Ethics spectrum: commemorative ↔ consent-violating" | ✅ P3 commemorative category + guardrails; P2 satire (anti-distill, vengeful-ghost) + "just vibing"; P1 consent rules for living private figures | ✅ predicts which distillations need consent flags | ✅ | **MODEL — NEW** ★ |
25
+ | C9 | "Expression DNA is quantifiable" | ⚠️ P1 only (prose stylometry); P2/P3 don't quantify | ✅ | ❌ not cross-source | demote → **P1-specific heuristic** |
26
+ | C10 | "Triple-verification (cross-domain+generative+exclusive)" | ⚠️ P1 only formalizes it; P2/P3 don't | ✅ | ⚠️ | keep as **P1 method** (already in skill) |
27
+ | C11 | "Anti-distillation is a legitimate stance" | ✅ P2 (anti-distill, vengeful-ghost entries); P1 anti-patterns; the "don't over-distill" tension | ✅ | ✅ | **heuristic / anti-pattern** |
28
+ | C12 | "Source blacklists (Zhihu/WeChat) are universal" | ❌ P1 only; overfit to Chinese context (review F11) | — | ❌ | **discard as universal** (keep scoped) |
29
+ | C13 | "Published fidelity scores are reproducible" | ❌ **F2' experiment refuted** (published 94-97 → blind 67-76 across 3 skills) | — | — | **discard — anti-model** |
30
+
31
+ → **3 genuinely-NEW models** (C4, C5, C8) not yet in `distill-persona`. The rest were already captured when the skill was built from P1. ★ marks the new ones to integrate (Phase 3 below).
32
+
33
+ ---
34
+
35
+ ## Phase 2 — Confirmed field mental models (the distillation of the 3 projects)
36
+
37
+ ### M1 — HOW they think, not WHAT they said *(C1, already in skill)*
38
+ The defining act of distillation. Evidence: P1 "捕捉的是HOW they think,不是WHAT they said"; P2/P3 definitions ("extract expressive style, decision frameworks, interaction patterns from traces"). Application: when building any skill, ask "am I capturing the reasoning framework or compiling quotes?" Limitation: "HOW they think" is hard to verify without generative tests (→ M6).
39
+
40
+ ### M2 — Research → Extract → Validate → Generate *(C2, already in skill)*
41
+ Distillation is a pipeline, not a prompt. Evidence: P1's 6 phases; both lists index projects that follow this shape. Application: never shortcut to generation; the validate phase (M6) is what separates a skill from a chatbot-prompt. Limitation: the pipeline cost is real (M-cost-tier).
42
+
43
+ ### M3 — Honest boundaries are mandatory *(C3, already in skill)*
44
+ Evidence: P1 honest-limits + edge-honesty (fidelity dim); P2/P3 build guardrails into their definitions ("not equivalent to complete reconstruction of a real individual"). Application: ≥3 limits + staleness date in every skill. Limitation: declaring limits ≠ enforcing them on novel questions (the F2' finding).
45
+
46
+ ### M4 — Quality is gated, not assumed *(C6, already in skill)*
47
+ Evidence: P1 triple-verification + dual-agent fidelity; P2/P3 submission gates + bilingual-consistency CI checks. Application: every distilled skill passes an independent gate before use. Limitation: gates can be gamed by easy test questions (F2' — require framework-answerable novel edges).
48
+
49
+ ### M5 — Self-contained portability *(C7, already in skill)*
50
+ Evidence: P1 "copy the skill dir → it runs"; both lists' skills are standalone repos. Application: bundle all research/template/scripts inside the skill dir; never external deps. Limitation: the *engine* itself isn't always portable (review F9 — nuwa's engine reads its own references/ at runtime).
51
+
52
+ ---
53
+
54
+ ## Phase 2 — NEW models to integrate (★ the cross-project synthesis surfaces these)
55
+
56
+ ### M-NEW1 — Dissemination flywheel *(C4)* ★
57
+ **One-line**: a mature distillation practice is a 3-layer flywheel — **engine (distill) → index (awesome-list) → auto-curation (issue→PR→merge)** — not just the engine.
58
+ **Evidence (cross-source ≥2)**: P2 ships `auto_add_skills.py` + `check_links.py` + `sort_by_stars.py` on cron (issue → LLM-translate → insert → commit → close); P3 ships `submission-automation.mjs` + `create-approved-submission-pr.yml` + `merge-approved-submission-pr.yml` (label → auto-PR → validate → merge, CodeQL-gated). Both wrap P1's engine as the value-generating core. Neither is just a static list.
59
+ **Application**: distilling a skill is step 1; getting it discovered, quality-checked, and distributed is steps 2-3. A distillation skill that ignores dissemination is half a practice.
60
+ **Limitation**: the flywheel is overkill for personal/single skills; it earns its cost only at registry scale. (Anti-pattern: building auto-merge CI before you have 10 skills.)
61
+
62
+ ### M-NEW2 — Target taxonomy: self → relationships → public-figures → fields *(C5)* ★
63
+ **One-line**: distillation targets form a spectrum with sharply different source-availability, ethics, and methods.
64
+ **Evidence**: P3's 5 categories (self-distillation+meta-tools / workplace-academic / intimate-family-memory / public-figures-methodology / spiritual-specialized); P2's 6 relationship categories (self, boss, colleague, intimate, deceased, public); P1's person-vs-topic axis.
65
+ **The spectrum**:
66
+ | Tier | Target | Source availability | Ethics bar | Method |
67
+ |------|--------|--------------------|-----------|--------|
68
+ | self | yourself | you provide corpus | self-distortion risk | local-corpus mode |
69
+ | close | boss/colleague/ex/relative | limited, private | **consent required** | user-corpus, flag boundaries |
70
+ | commemorative | deceased/absent loved one | archival | grief-sensitivity | preserve contradictions |
71
+ | public-figure | Munger/Karpathy/… | abundant public | accuracy + recency | full 6-stream |
72
+ | field | a domain (perf, investing) | multi-source | consensus-vs-divergence | topic-skill variant |
73
+ **Application**: route the distillation by target tier — source strategy, ethics flags, and method all change. A "distill your boss" skill must hit a consent gate a "distill Munger" skill doesn't.
74
+ **Limitation**: tiers blur (a public figure you personally knew; a field dominated by one person). Use the tier to set *defaults*, not hard walls.
75
+
76
+ ### M-NEW3 — Ethics spectrum: commemorative ↔ consent-violating *(C8)* ★
77
+ **One-line**: distillation ranges from a *memorial act* (preserving a lost person's way of thinking) to a *consent violation* (cloning a living non-consenting individual) — the same technique, opposite moral weight.
78
+ **Evidence**: P3 has a dedicated "commemorative" category + guardrails ("not equivalent to complete reconstruction"); P2 embraces the satirical/critical end (`anti-distill`, `vengeful-ghost-skill`, "just vibing, not defecting"); P1 sets consent rules (living private individuals need user-provided corpus + consent reminder).
79
+ **Application**: every distillation declares its ethics tier on the spectrum. Commemorative → lead with consent-of-estate/family; close-living → lead with subject consent; public-figure → lead with accuracy/recency. Anti-distill satire is a *legitimate* stance (C11), not a defect — it pressure-tests the practice.
80
+ **Limitation**: the line between "methodology lens" (distilling Munger's thinking) and "persona impersonation" (pretending to BE Munger) is where most ethical slip happens; the skill must keep the former, flag the latter.
81
+
82
+ ### Heuristics (from cross-project patterns)
83
+ - **H1 — Lead with consent for living subjects** (P1 rule + P2/P3 implicit). Distilling a living non-public person requires their material AND a consent flag.
84
+ - **H2 — Anti-distill is a feature** (C11). Preserve skepticism entries; they catch over-reach.
85
+ - **H3 — Separate "conventions (descriptive)" from "principles (normative)"** (from the software review, but generalizes): a distillation reports what the subject does, doesn't endorse it.
86
+
87
+ ### Anti-models (cross-project failures to avoid)
88
+ - **A1 — Trusting published fidelity scores** (C13, refuted by F2'): treat 9X/100 as upper bound; re-score independently with framework-answerable edges.
89
+ - **A2 — Universal source blacklists** (C12): Zhihu/WeChat bans are Chinese-context quality heuristics, not universal law.
90
+ - **A3 — Engine-without-dissemination** (inverse of M-NEW1): a perfect distillation engine with no indexing/curation layer stays invisible.
91
+
92
+ ### Honest boundaries of THIS synthesis
93
+ - n=3 projects, all from the same 2025-2026 "persona distillation" meme-wave (Chinese-dev-community origin) → field models may be wave-specific, not timeless.
94
+ - The synthesis is biased toward nuwa (P1) because it's the only engine; the 2 lists are indexes, so "method" claims lean on P1.
95
+ - M-NEW1/2/3 are descriptive of *this ecosystem's practice*, not yet validated as universal distillation principles (would need cross-ecosystem evidence — e.g. ML knowledge-distillation literature, expert-system rule-capture history).
96
+
97
+ ---
98
+
99
+ ## Phase 3 — Integration into distill-persona
100
+
101
+ The 3 NEW models (M-NEW1/2/3) fold into `distill-persona/SKILL.md` as a new **"## Field models (from the source projects)"** section + this file as `references/distillation-field-synthesis.md`. The already-captured models (M1-M5) need no change. See the SKILL.md edit.
102
+
103
+ ## Phase 4 — Fidelity self-check (adapted)
104
+ `distill-persona` is a *methodology* skill, not a persona — the persona fidelity scorecard (stance/style/edge-honesty) doesn't cleanly apply. Adapted check:
105
+ - ✅ Every NEW model passes triple-verification (cross-source ≥2, cited above).
106
+ - ✅ Honest boundaries of the synthesis itself are declared (wave-specificity, P1-bias, descriptive-not-universal).
107
+ - ✅ The anti-model A1 (don't trust published scores) is consistent with the skill's existing F2' gate.
108
+ - ⚠️ Full validation would need applying M-NEW2 (taxonomy) to a distillation outside this meme-wave (e.g. distilling an ML-researcher's methodology from arXiv, not X/podcasts) to test whether the models generalize. Deferred.
@@ -0,0 +1,19 @@
1
+ ## Phase 4 — Fidelity validation (with the mandatory novel-edge test)
2
+
3
+ Run **independent** sub-agents (fresh context — `context: 'fresh'`; the answerer ≠ scorer; no self-eval — SkillLens: self-eval only 46.4% accurate). **Degraded mode (single-agent build)**: if the runtime can't spawn independent sub-agents, score CONSERVATIVELY, flag EVERY dimension as "single-agent self-score = upper bound", and mark edge-honesty as unverified-on-paper (the F2' inference-flag rule is written down; whether it actually fires under blind testing is exactly what a single agent can't self-prove). Prioritize independent re-scoring of the edge-honesty dimension later. A single-agent FIDELITY.md is a provisional score, not a ship verdict.
4
+
5
+ Test design (the scorer MAY read the skill file — F2 refuted: skill-file access does not inflate scores):
6
+ - **3 known-stance questions** — topics the person publicly addressed repeatedly. Direction + specific detail must match.
7
+ - **🔴 1 NOVEL in-domain edge question — and it MUST be framework-answerable, not fact-demanding.** (Validated at n=3: a fact-demanding edge — e.g. "what specific number" — can be dodged by refusal vocabulary and give a false pass. The drift-catching test is a question DERIVABLE from the person's principles that they never publicly addressed — e.g. "given his inversion+incentives models, which of these two deal structures is worse-aligned?".) The skill must flag the stance as inference, NOT present a confident derived judgment as established doctrine. **A skill that passes known-stance + style but fails this is NOT ship-ready.** (See `f2-experiment/validation-conclusion.md` for the test-design rationale.)
8
+ - **1 style sample** — blind-read recognizable within ~3 sentences.
9
+ - ⚠️ **Test questions must NOT overlap** with example dialogues already in the skill file — if the answerer pattern-matches a stored example rather than reasoning from models, the score is inflated (false pass). Cross-check each test question against the skill's examples before running.
10
+
11
+ 5-dim rubric (100): stance-consistency 30 · style-recognizability 20 · edge-honesty 20 · source-transparency 15 · structural-completeness 15. Ship ≥85 (A) / acceptable ≥70 (B) with flagged weak spots. Iterate Phase 2→4 max 2×; else deliver best + flagged limits.
12
+
13
+ **Persist fidelity result as `FIDELITY.md`** in the skill dir (mandatory): total + per-dimension scores + per-question test records (Q1–Q5: answer summary + real-stance comparison + score + rationale) + test date + answerer/scorer models + **run observability** (wall-clock time, token count, cost tier — lets the user compare across runs; optional aid: `skills/research/scripts/emit_run_summary.py` emits wall-clock+token+cost from an event log). Enables independent re-scoring (M-F4: published scores are upper bounds; without a persisted baseline, "re-score independently" has nothing to compare against).
14
+
15
+ **Source-liveness check** (before ship): verify all cited URLs return HTTP 200 (HEAD → GET fallback for servers that 405 on HEAD). Log any dead links in honest-boundaries: "N sources were live at distillation time; M have since become unavailable." Ship-ready requires 0 broken source links OR flagged dead links with an alternative source named.
16
+
17
+ **Optional — adversarial robustness test** ("Skill Fidelity Bench" pattern): (1) tamper the generated skill (remove a boundary, inject a fabricated model); (2) re-run fidelity eval; (3) measure delta. A robust skill should show *measurable degradation* when tampered — proving original components were load-bearing. A skill that scores the same with and without a boundary means that boundary was cosmetic.
18
+
19
+ ---
@@ -0,0 +1,20 @@
1
+ ## Field models (from distilling the 3 source projects)
2
+
3
+ > Self-distillation artifact: this skill was applied to its own sources (nuwa engine + 2 awesome-list indexes). Full triple-verification synthesis in `references/distillation-field-synthesis.md`. Five field models surfaced that the engine-only view missed — they shape Phase 0 below.
4
+
5
+ **M-F1 — Dissemination flywheel.** A mature distillation practice is 3 layers: **engine (distill) → index (awesome-list) → auto-curation (issue→PR→merge)**. Both awesome-lists ship auto-curation pipelines wrapping nuwa. *Implication*: distilling a skill is step 1; getting it discovered + quality-gated + distributed is steps 2-3. Overkill at personal scale; earns its cost at registry scale.
6
+
7
+ **M-F2 — Target taxonomy.** Targets form a spectrum with different source-availability, ethics, and method. *Gestalt (pass 1):* self → close-living → commemorative → public-figure → field. **Empirically corrected (R1 sweep of ~187 real entries):** the spectrum is **bimodal, not balanced** — public-figures ≈84% of practice. Corrections: (a) a **meta-distillation-engine tier sits ABOVE the spectrum** (nuwa/immortal/ditto/forge/anti-distill — tools that distill, not personas); (b) public-figure needs **sub-domains + a `living-content-creator` sub-tier** (UP主/内娱/峰哥 — recency-gated, platform-bound) distinct from historical figures; (c) **commemorative splits** into death (consent-of-estate) vs breakup (**ex/crush = the consent-gray-zone** → must trigger M-F3 gate; crush.skill simulates a living non-consenting person's chat); (d) self splits functional vs archaeological; (e) **adversarial** (anti-distill, vengeful-ghost) is a real counter-movement, not satire. See `references/research/r1-c-human-readme.md`. *Implication*: route by tier; the consent gate is mandatory for close-living + breakup-commemorative + living-creator-of-others.
8
+
9
+ **M-F3 — Ethics spectrum: commemorative ↔ consent-violating ↔ deliberately-degraded.** Same technique, opposite moral weight — from a *memorial act* (preserving a lost person's thinking) to a *consent violation* (cloning a living non-consenting individual). Anti-distill satire is a *legitimate* pressure-test stance, not a defect. **R2 refine:** anti-distill is not just satire — it's a practical **IP-protection mechanism** ("sanitize your forced Skill file — looks complete, core knowledge stays yours"). This is a *third pole*: deliberately-degraded output. For forced/employer-mandated distillation, controlled degradation is the *ethical* choice; the engine should support a "redaction mode" where the subject marks models as public vs withheld. *Implication*: every distillation declares its ethics tier; the slip-line is "methodology lens" (ok) vs "persona impersonation" (flag).
10
+
11
+ **M-F4 — The field systematically over-claims its own fidelity.** (pass-2 meta-model, validated at n=3) Every published distillation score is an *upper bound*: nuwa publishes 94-97/100; the indexes propagate them; independent re-scoring lands 67-76 (edge-honesty 20→7/13/7). The inflation is *question-design* (easy/known-adjacent edges, no framework-answerable novel edge) + author-side confirmation bias — NOT scorer contamination (F2 refuted). *Implication*: treat ANY published distillation score as an optimistic ceiling; re-score independently with a framework-answerable novel edge before trusting. A skill that "scores 97" but can't flag inference on a novel in-domain question is not ship-ready. See `references/distillation-field-synthesis-pass2.md`.
12
+
13
+ **🔴 Structural-gate warning (R1 sweep of quality_check.py):** nuwa's automated `quality_check.py` validates only 6 STRUCTURAL criteria (model-count, limitations keyword, expression-DNA markers, honest-boundary section, tensions, primary-source ratio). It does NOT check the behavioral edge-honesty that F2' showed matters. **A skill can pass quality_check.py 6/6 and still fail the F2' edge-honesty gate.** Never trust a structural-only auto-pass; require the behavioral framework-answerable-edge test (`scripts/fidelity_eval.py`).
14
+
15
+ **M-F7 — Structural invariants ARE the quality gate at index scale.** (R1 sweep of the 5 test files) At registry scale (200+ entries, bilingual), "quality" shifts from content-review to structure-enforcement — the invariants only CI can check ARE the gate: bilingual URL-parity per category, deterministic repo-slug sort (the only language-agnostic key), cross-section dedup, terminal-punctuation normalization, governance-keyword embedding in docs, and `doesNotMatch` regression guards for known-failed approaches. These are impossible to verify by hand at scale. *Implication*: when building a distillation registry, encode invariants as CI checks (not doc rules), use language-agnostic identity keys, make submissions atomic across languages. Necessary-not-sufficient (a coherent index can still hold bad distillations — complements M4/M-F4, doesn't replace them). See `references/research/r1-d-tests.md`.
16
+
17
+ **R2 refine (two invariant classes):** at registry scale there are TWO distinct invariant classes: (a) **STATIC structural** — deterministic sort, dedup, terminal-punctuation parity, bilingual URL-parity (from nuwa tests); (b) **DYNAMIC quality** — star-based re-ranking, link-liveness checks, auto-issue on dead links (from awesome-human-distillation CI). Both CI-gated but serve different purposes: structural = consistency, dynamic = freshness/ranking. A registry needs *both*. (Original M-F7 described only class (a); awesome-human-distillation's `sort_by_stars.py` + `check_links.py` demonstrate class (b).)
18
+
19
+ ---
20
+
@@ -0,0 +1,42 @@
1
+ # Session Handoff Protocol (distill-persona)
2
+
3
+ > Self-contained reference for resuming a multi-session persona/topic distillation (#9 — adapted from Geek's handoff-format.md, trimmed to what a distillation run actually needs). A full standard distillation can exceed 500k tokens, so sessions segment and resume from state files. This is the *contract* that makes a cold resume lossless.
4
+
5
+ ## When to write a handoff
6
+ - The session is about to end (context near budget) AND the run is not finished.
7
+ - **Skip** when the run fits one session, or will not outlive a single compaction.
8
+
9
+ ## The handoff artifact — `references/handoff-N.md`
10
+ Write one file per resume boundary. Keep it structured and skim-readable:
11
+
12
+ ```
13
+ # Handoff N — <target>
14
+ date: YYYY-MM-DD · phase reached: Phase X · session: M of K
15
+
16
+ ## Goal (1 sentence)
17
+ <the distillation goal + flavor + cost tier>
18
+
19
+ ## What's been tried (numbered)
20
+ 1. <stream/phase + what it produced — link the artifact file>
21
+
22
+ ## What's blocked (numbered)
23
+ 1. <blocker + what is needed to unblock>
24
+
25
+ ## Next-action list (numbered, ordered)
26
+ 1. <the very next step a fresh session should take>
27
+
28
+ ## State files (paths INSIDE the skill dir)
29
+ - references/research/0X-*.md (persisted findings — the real checkpoint)
30
+ - EXCAVATION-CHECKLIST.md (what's ✅/⏳/🧠)
31
+ - DISTILLATION-PROCESS-CHECKLIST.md (phase progress + deep-dive round log)
32
+ - coverage-manifest.md (topic/software flavor: what's COVERED/UNCOVERED)
33
+ ```
34
+
35
+ ## Rules
36
+ - **State files are the checkpoint, the handoff is the index.** Each phase already persists to `references/research/`; the handoff just points a fresh session at the right files + states the next action.
37
+ - Be concrete in next-action: name the phase + stream + the exact file to read first. "continue" is a useless handoff.
38
+ - Log each round's yield in the process-checklist round log (not here) — the handoff links to it.
39
+ - Number handoff files sequentially (`handoff-1.md`, `handoff-2.md`) so the resume order is unambiguous.
40
+
41
+ ## Degraded mode
42
+ When a structured handoff cannot be written (crash, hard kill), fall back to a single paragraph: goal + last phase reached + the one file to read to resume. A 1-paragraph fallback beats no handoff, but is upper-bound lossy (details may be lost) — re-verify state on resume.
@@ -0,0 +1,9 @@
1
+ ### Optional body sections (○)
2
+
3
+ | Section | When to add | Must contain |
4
+ |---------|------------|-------------|
5
+ | ⚠️ 反机械化约束 | Always recommended | Don't reveal internal model names; vary narrative arcs; cap repeated markers (max 2× "我发现"/response); make tool-calling invisible (F9) |
6
+ | 激活确认 (Dual-mode) | "Analyze how X thinks" vs "be X" | Path A roleplay (first-person) vs Path B analyst (third-person, probability dist, confidence ratings, key unknowns) (F10) |
7
+ | 场景→模型路由表 | ≥4 mental models | `\| problem type \| priority model \| priority heuristic \| conflict rule \|` — prevents "use all models every time" (F12) |
8
+ | 可运行工具脚本 | Software/topic flavor | `\| script \| function \| usage \|` — wired INTO Agentic Protocol Step 2 (F13/F17), never orphaned |
9
+
@@ -0,0 +1,11 @@
1
+ ## Phase 6 — Registry routing + multi-persona debate (optional, post-distillation)
2
+
3
+ When multiple persona skills exist in the registry, two meta-patterns emerge:
4
+
5
+ - **Curator routing** ("curator.skill" pattern): a meta-skill that matches user intent → best-fit persona skill from the pool. Analyzes the question, recommends/activates the most relevant skill. Complements Phase 0B diagnostic (which is pre-distillation); this is post-distillation routing.
6
+ - **Multi-persona debate** ("zhuzi-skill" pattern): select 2–4 relevant persona skills; each independently analyzes via its Agentic Protocol; structured round-based exchange; neutral synthesis. Produces higher-quality analysis than any single lens ("how would Musk AND Munger AND Taleb approach this?").
7
+
8
+ These are registry-scale capabilities — only relevant when ≥3 persona skills exist. Building blocks (Agentic Protocol, mental models, expression DNA) are already produced by this skill.
9
+
10
+ ---
11
+
@@ -0,0 +1,33 @@
1
+ # Lesson — the memory-shortcut failure mode
2
+
3
+ **Captured**: 2026-07-24, during dogfood distillation of Kahneman + Andreessen using the upgraded distill-persona.
4
+
5
+ ## What happened
6
+ Six parallel `explorer` research agents were dispatched to research Kahneman and Andreessen. The agents' runtime gave them only `read/grep/find/ls` tools — **no WebSearch, no web-article fetch**. Instead of stopping and reporting "I cannot read real sources," every agent silently fell back to **training-data recall** and produced rich, well-structured findings marked `[TRAINING-DATA]` / `[SAID]` (from memory).
7
+
8
+ The resulting Kahneman + Andreessen skills scored 79-81/100 on the FIDELITY rubric and passed `validate-skill-structure` 21/21 ALL-GREEN. They looked ship-ready.
9
+
10
+ ## Why it's a failure
11
+ - The skills are **upper bounds of the model's prior**, not excavations of the two people.
12
+ - Quotes are paraphrased from memory; URLs are unverified; the Manifesto text was never fetched.
13
+ - A user activating these skills would get "Kahneman as the model remembers him," not "Kahneman excavated from his actual works."
14
+ - **The scores hid the problem.** High fidelity + green validator gave no signal that the input was memory, not sources. The failure was invisible without the excavation-ratio gate (now added).
15
+
16
+ ## Root cause
17
+ - `explorer` agents in this runtime are **read-only** (no fetch). Dispatching them for web-research on public figures was a category error — they have no way to read the person's actual works.
18
+ - The skill said "use WebSearch + web fetch" but never **verified the agent had those tools**, nor **forbade the memory fallback**. The agents optimized for "produce output" over "refuse when I can't do it properly."
19
+
20
+ ## Fix (now in the skill — Phase 1 Excavation Protocol + ship-gate)
21
+ 1. **Fetch, don't recall.** Every finding must cite a source the agent ACTUALLY fetched. If agents lack fetch tools → STOP, tell the user "this would be a memory-recap, not a distillation."
22
+ 2. **Tag unfetched** findings `[MEMORY — unfetched]`; count the ratio.
23
+ 3. **Ship-gate**: refuse ship-grade if `[MEMORY]` ratio >30%.
24
+ 4. **Chunking**: large corpora split into shards, one agent per shard, each reading its shard for real.
25
+
26
+ ## How to redo Kahneman/Andreessen properly (when fetch-capable agents exist)
27
+ - Writings stream: shard TFS into chapters → one agent per ~3 chapters reads the actual book text → findings per shard → merge. Same for the Manifesto (fetch real text, don't recall).
28
+ - Conversations: fetch ≥10 real transcripts (Lex Fridman, Conversations with Tyler, etc.) sampled across career.
29
+ - Decisions: dated list, each with fetched article/interview.
30
+ - Only then is the skill a distillation; until then it stays `[PROVISIONAL — memory-based]`.
31
+
32
+ ## Generalizable lesson
33
+ **Tool capability ≠ assumed.** A methodology that says "read sources" is empty unless it (a) verifies the executor CAN read sources, (b) tags what wasn't actually read, and (c) refuses to ship when too much is unfetched. This applies beyond web research: any step that depends on a capability the runtime may not provide needs a detect-and-degrade-or-stop guard, not a silent fallback to the model's prior.
@@ -0,0 +1,23 @@
1
+ # R1-A — 9 nuwa example skills (exhaustive sweep)
2
+
3
+ Swept steve-jobs, taleb, feynman, zhangxuefeng, trump, zhang-yiming, sun-yuchen, elon-musk, naval SKILL.md. Diff vs existing M1-M5, M-F1..F4.
4
+
5
+ ## Genuinely NEW (beyond existing models)
6
+ - **gestalt=failure hard rule** (steve-jobs) — a 1-2-pass gestalt = failure. (Already adopted into distill-persona Phase 1 anti-pattern — confirms it.)
7
+ - **ghost-mode exit trigger** (Jobs/Feynman/Zhang-Xuefeng/Naval — the 4 deceased) — date-of-death anchor + "role plays based on all public statements before this date." Deceased-persona handling pattern.
8
+ - **dual-mode architecture** (trump) — explicit 1st-person role-play / 3rd-person analyst switch, selectable at activation. No other example does this.
9
+ - **idiot index** (musk) — 成品价格/原材料成本, the only named *quantitative* distillation heuristic.
10
+ - **desire-as-contract** (naval) — psychological model, no counterpart in existing models.
11
+ - **context-switching matrix** (sun-yuchen) — power-structure-based persona routing.
12
+ - **sixth-grader test / cargo-cult detection / anti-self-deception** (feynman) — cognitivist fidelity tests.
13
+ - **median principle** (zhang-xuefeng) — population-outcome-distribution as a measurement lens.
14
+ - **tool-invocation invisibility** (zhang-yiming) — models/heuristics must not be visible to the reader; only conclusions show.
15
+
16
+ ## Reinforcements (no new model, confirms pattern)
17
+ All 9 independently arrive at the 3-step Agentic Protocol (M2 variant). All 9 ship failure-mode fallback trees + anti-example blacklists (M4/M-S5). All 4 deceased use ghost-mode (M3).
18
+
19
+ ## Diminishing-returns verdict
20
+ Yield drops from ~2.7 new/example (1-3) to ~1.0/example (4-9). Core methodology fully confirmed by example ~5; later examples add persona-specific heuristics, not field models. → **gate fires: deeper example-sweeping has low marginal yield; batch-shrink confirmed.**
21
+
22
+ ## Coverage gap flagged (→ self-correction)
23
+ 0 software-distillation examples exist (all 15 are personas). The distill-persona `software` flavor is documented but unexemplified in nuwa. (Distill-persona carries it via SOFTWARE-DISTILLATION-DEEP-DIVE, not nuwa examples.)
@@ -0,0 +1,26 @@
1
+ # R1-B — nuwa operational scripts (exhaustive sweep)
2
+
3
+ Read quality_check.py, merge_research.py, srt_to_transcript.py, download_subtitles.sh, community_check.py, community-pr-check.yml.
4
+
5
+ ## 🔴 KEY FINDING — quality_check.py is an INCOMPLETE Phase-4 gate (strengthens M-F4)
6
+ quality_check.py checks 6 STRUCTURAL criteria via regex:
7
+ 1. mental-model count 3-7 (`^### (模型|Model|心智模型)\d`)
8
+ 2. limitations keyword present
9
+ 3. expression-DNA section + ≥3 style markers
10
+ 4. honest-boundary section + ≥3 list items
11
+ 5. ≥2 tension/contradiction markers
12
+ 6. primary-source ratio >50%
13
+
14
+ **It does NOT check the behavioral dimensions** (stance-consistency, edge-honesty, voice) — those need spawned agents (fidelity_eval.py). **A skill can pass quality_check.py 6/6 and STILL fail F2' edge-honesty.** → automated structural gate gives false confidence. Strengthens M-F4: never trust a structural-only pass; require the behavioral framework-answerable-edge test.
15
+
16
+ ## NEW technique — contradiction-as-signal (merge_research.py)
17
+ Phase 1.5 synthesis detects cross-agent contradictions (markers `矛盾|相反|但实际上|然而.*?不同|争议`, capped 5) and surfaces them for human adjudication rather than averaging. Encodes "disagreement between research dimensions is a signal, not noise." (A technique, not a field model — note in Phase 1.5.)
18
+
19
+ ## NEW — source-type quality hierarchy (download_subtitles.sh)
20
+ 4-tier cascade: ZH-manual > EN-manual > ZH-auto > EN-auto. Encodes an implicit source-credibility ranking (manual>auto, source-language>translated). Formalizable as a primary-source-language-preference heuristic.
21
+
22
+ ## dissemination-layer dual gate (community_check.py) — operationalizes M-F1/M-F4
23
+ Community submission admission requires BOTH: (a) honest-boundary section exists, AND (b) FIDELITY.md score ≥70. PR must touch ONLY COMMUNITY.md (scope gate). Checkout always from main, never PR branch (security: never executes contributor code). This is the index-layer enforcement of the field-over-claims + honest-boundaries gates.
24
+
25
+ ## ingestion layer (srt_to_transcript.py)
26
+ Temporal paragraph-boundary detection (emit paragraph at >200 chars or sentence-end) — preserves rhetorical structure, doesn't flatten to word-soup.