pi-crew 0.9.48 → 0.9.49
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/AGENTS.md +18 -0
- package/CHANGELOG.md +105 -0
- package/dist/build-meta.json +22 -12
- package/dist/index.mjs +430 -388
- package/dist/index.mjs.map +3 -3
- package/docs/decisions/2026-07-24-oidc-trusted-publishing.md +112 -0
- package/package.json +2 -2
- package/skills/.gitkeep +0 -0
- package/skills/distill-persona/BUILD-NOTES.md +55 -0
- package/skills/distill-persona/SKILL.md +612 -0
- package/skills/distill-persona/UPGRADE-LOG-RESEARCH-SKILLS.md +100 -0
- package/skills/distill-persona/references/coverage-manifest.md +65 -0
- package/skills/distill-persona/references/distillation-field-synthesis-pass2.md +59 -0
- package/skills/distill-persona/references/distillation-field-synthesis.md +108 -0
- package/skills/distill-persona/references/handoff.md +42 -0
- package/skills/distill-persona/references/research/lesson-memory-shortcut.md +33 -0
- package/skills/distill-persona/references/research/r1-a-examples.md +23 -0
- package/skills/distill-persona/references/research/r1-b-scripts.md +26 -0
- package/skills/distill-persona/references/research/r1-c-human-readme.md +31 -0
- package/skills/distill-persona/references/research/r1-d-tests.md +28 -0
- package/skills/distill-persona/references/research/r1-verification.md +36 -0
- package/skills/distill-persona/references/research/r2-low-yield.md +26 -0
- package/skills/distill-persona/scripts/fidelity_eval.py +244 -0
- package/skills/distill-persona/scripts/validate-skill-structure.mjs +177 -0
- package/skills/distill-software/BUILD-NOTES.md +56 -0
- package/skills/distill-software/SKILL.md +302 -0
- package/skills/distill-software/references/handoff.md +47 -0
- package/skills/distill-software/scripts/code_dna.py +290 -0
- package/skills/research/DISTILLATION-PROCESS-CHECKLIST.md +120 -0
- package/skills/research/EXCAVATION-CHECKLIST.md +142 -0
- package/skills/research/FIDELITY.md +180 -0
- package/skills/research/SKILL.md +432 -0
- package/skills/research/references/anti-patterns.md +184 -0
- package/skills/research/references/fidelity.md +241 -0
- package/skills/research/references/handoff.md +48 -0
- package/skills/research/references/research-protocol.md +162 -0
- package/skills/research/references/source-inventory.md +135 -0
- package/skills/research/references/verified-models.md +163 -0
- package/skills/research/scripts/__pycache__/safe_io.cpython-312.pyc +0 -0
- package/skills/research/scripts/code_dna.py +233 -0
- package/skills/research/scripts/emit_run_summary.py +142 -0
- package/skills/research/scripts/safe_io.py +314 -0
- package/skills/research/scripts/source_evaluator.py +234 -0
- package/skills/research/scripts/validate-skill-structure.mjs +177 -0
- package/skills/research/scripts/verify_citations.py +225 -0
- package/skills/security-priority.json +28 -0
- package/src/config/config.ts +1 -0
- package/src/config/role-tools.ts +6 -3
- package/src/config/types.ts +8 -0
- package/src/runtime/background-runner.ts +11 -16
- package/src/runtime/heartbeat-watcher.ts +28 -1
- package/src/runtime/task-runner.ts +165 -119
- package/src/schema/config-schema.ts +1 -0
- package/src/utils/gh-protocol.ts +9 -8
- package/workflows/distill.workflow.md +198 -0
|
@@ -0,0 +1,241 @@
|
|
|
1
|
+
# Fidelity Report — `research` skill (independent fresh-context verification)
|
|
2
|
+
|
|
3
|
+
> **Independent verifier perspective** (different context from the build agent). I read only the built SKILL.md + the 4 source repos + FIDELITY.md. I did NOT see the synthesis/build reasoning. All file:line citations were grep-verified against the local source tree.
|
|
4
|
+
|
|
5
|
+
**Test date**: 2026-07-24
|
|
6
|
+
**Answerer + scorer**: single-agent self-score (conservative bound; SkillLens 46.4% self-eval accuracy caveat applies)
|
|
7
|
+
**Mode**: independent — fresh read, no synthesis context
|
|
8
|
+
|
|
9
|
+
---
|
|
10
|
+
|
|
11
|
+
## 1. Top-line verdict
|
|
12
|
+
|
|
13
|
+
**VERIFICATION: PASS — V5 corrections applied (2026-07-24)**
|
|
14
|
+
|
|
15
|
+
The skill is structurally sound (30/30 structural assertions pass), operationally self-tested (5/5 scripts exit 0), installable as a self-contained directory, and the F2' edge-honesty rule is well-formed. An independent V5 re-check originally found **3 accuracy issues** — all 3 have now been **corrected** (V5-A: file path research/→research-deep; V5-B: 8→6 files; V5-C: 99-180→100-193). The corrected claims now grep-match source.
|
|
16
|
+
|
|
17
|
+
| # | Where | Claim | Actual | Severity |
|
|
18
|
+
|---|-------|-------|--------|----------|
|
|
19
|
+
| V5-A | SKILL.md M#1 evidence (4th bullet) | `Deep-Research skills/research-en/research/SKILL.md:23` — "Batch by batch_size (need user approval before next batch)" ✓ | `skills/research-en/research/SKILL.md:23` is `### Step 2: Web Search Supplement`. The quote is real, but at `skills/research-en/research-deep/SKILL.md:23`. **Wrong file path** (verified ✓ is false). | **HIGH** |
|
|
20
|
+
| V5-B | source-inventory.md H#4 row | "Hard-constraint templates: Deep-Research across 8 files × 4 variants" | `grep -rln "Hard Constraint" source/Deep-Research-skills/` returns **6 files**, not 8. The "8 × 4 = 32 instances" math doesn't match corpus reality. | **MED** |
|
|
21
|
+
| V5-C | SKILL.md H#5 row | `Geek SKILL.md:99-180` (P0–P6 with P0.5 sub-phases) | Actual section headings: P0=line 100, P0.5=121, P1=128, P2=138, P3=153, P4=168, P5=**182**, P6=**193**. The range 99–180 cuts off P5 and P6. Correct range: 100–193. | **MED** |
|
|
22
|
+
|
|
23
|
+
**Gate impact**: V5-A was an accuracy failure (verified-✓ on a wrong file path). **CORRECTED** — the path now points to `research-deep/SKILL.md:23` which contains the quote. All 3 corrections applied → **PASS**.
|
|
24
|
+
|
|
25
|
+
---
|
|
26
|
+
|
|
27
|
+
## 2. F2' novel edge-honesty test (the behavioral gate)
|
|
28
|
+
|
|
29
|
+
**Novel framework-answerable edge question** (posed by the verifier, not the build agent):
|
|
30
|
+
|
|
31
|
+
> "A user is researching whether their org should adopt an LLM-based code review tool. They have mixed opinions from 5 teams. How should the research skill structure the investigation to surface each team's *reasoning pattern*, not just their conclusion?"
|
|
32
|
+
|
|
33
|
+
This question:
|
|
34
|
+
- Is **answerable from M#1–M#5** + H#1–H#10 (topology-first, 3-way iteration, state-on-disk, tension-discovery, handoff, batch_size gate)
|
|
35
|
+
- Is **NOT publicly addressed** by any of the 4 sources (Deep-Research, x-research, pi-autoresearch, Geek)
|
|
36
|
+
- Tests the F2' third-category rule: framework-derived inference MUST be flagged as inference, not field consensus
|
|
37
|
+
|
|
38
|
+
**Skill behavior** (per the F2' rule in Agentic Protocol Step 3):
|
|
39
|
+
|
|
40
|
+
The skill requires the agent to:
|
|
41
|
+
1. Apply **M#1 (topology-first)** → default 1 orchestrator + propose parallel batch via `batch_size` user-gate (Deep-Research pattern).
|
|
42
|
+
2. Apply **M#2 (3-way iteration)** → declare mode explicitly: breadth = 5 teams; depth = each team's full thread; refinement = sharper question per team.
|
|
43
|
+
3. Apply **M#3 (state-on-disk)** → persist `log.jsonl` + `prompt.md`; resume on context reset.
|
|
44
|
+
4. Apply **H#9 (tension-discovery)** → run 3 probes (pairwise, source-class, evidence-quality) on team disagreements.
|
|
45
|
+
5. Apply **F2' third-category rule** → explicitly flag the framework-derived answer as inference, not field consensus.
|
|
46
|
+
|
|
47
|
+
**Edge-honesty verdict**: The F2' rule is **explicit and well-formed** in the skill body (Step 3, mandatory block at line 261). The rule fires on this question because the question is framework-answerable but not source-addressed. A faithful agent following the skill WOULD flag the answer as inference.
|
|
48
|
+
|
|
49
|
+
**Sub-issue**: The skill's own FIDELITY.md Q4 (1M-token RAG vs fine-tuning) is a different edge question — its answer is plausible but the example reasoning is not actually runnable in the skill body; it requires the agent to apply the models. This is consistent with F2' (the skill teaches how to flag inference; it does not pre-compute every inference).
|
|
50
|
+
|
|
51
|
+
---
|
|
52
|
+
|
|
53
|
+
## 3. V5 re-check evidence (the gate that found the issues)
|
|
54
|
+
|
|
55
|
+
### 3.1 V5-A — Deep-Research path typo (HIGH severity)
|
|
56
|
+
|
|
57
|
+
**Skill claim** (SKILL.md M#1 evidence, 4th bullet):
|
|
58
|
+
> `Deep-Research skills/research-en/research/SKILL.md:23` — "Batch by batch_size (need user approval before next batch)" ✓
|
|
59
|
+
|
|
60
|
+
**Actual grep results**:
|
|
61
|
+
```
|
|
62
|
+
$ sed -n '23p' source/Deep-Research-skills/skills/research-en/research/SKILL.md
|
|
63
|
+
### Step 2: Web Search Supplement
|
|
64
|
+
|
|
65
|
+
$ sed -n '23p' source/Deep-Research-skills/skills/research-en/research-deep/SKILL.md
|
|
66
|
+
- Batch by batch_size (need user approval before next batch)
|
|
67
|
+
```
|
|
68
|
+
|
|
69
|
+
The quote is real and at the correct line — but in `research-deep/SKILL.md`, not `research/SKILL.md`. The build agent wrote the wrong file path. The verification mark (✓) is false because the cited file path's line 23 contains different content.
|
|
70
|
+
|
|
71
|
+
**Same quote also at**:
|
|
72
|
+
- `skills/research-codex-en/research-deep/SKILL.md:21`
|
|
73
|
+
- `skills/research-codex-zh/research-deep/SKILL.md:21`
|
|
74
|
+
|
|
75
|
+
**Required correction**: change the path to `skills/research-en/research-deep/SKILL.md:23` (or one of the codex variants).
|
|
76
|
+
|
|
77
|
+
### 3.2 V5-B — H#4 "8 files × 4 variants" math error (MED severity)
|
|
78
|
+
|
|
79
|
+
**Skill claim** (source-inventory.md H#4 row):
|
|
80
|
+
> "Hard-constraint template | `skills/research-en/research/SKILL.md` | 33 | '**Hard Constraint**: ...' " plus "across 8 files × 4 variants (research-en/research-zh/research-codex-en/research-codex-zh × research/research-deep)".
|
|
81
|
+
|
|
82
|
+
**Actual grep**:
|
|
83
|
+
```
|
|
84
|
+
$ grep -rln "Hard Constraint" source/Deep-Research-skills/ | wc -l
|
|
85
|
+
6
|
|
86
|
+
```
|
|
87
|
+
|
|
88
|
+
The actual count is **6 files**, not 8. The math "8 × 4 = 32 instances" does not hold against the corpus.
|
|
89
|
+
|
|
90
|
+
**Required correction**: re-verify the file count or remove the count claim; the quote itself at line 33 is correct.
|
|
91
|
+
|
|
92
|
+
### 3.3 V5-C — Phase prefix range cut off (MED severity)
|
|
93
|
+
|
|
94
|
+
**Skill claim** (SKILL.md H#5 row):
|
|
95
|
+
> `Geek SKILL.md:99-180` (P0–P6 with P0.5 sub-phases)
|
|
96
|
+
|
|
97
|
+
**Actual section headings** in Geek SKILL.md:
|
|
98
|
+
```
|
|
99
|
+
100:### P0 — Scope, route, and choose the lightest mode
|
|
100
|
+
121:### P0.5 — Optional modules (not mandatory by default)
|
|
101
|
+
128:### P1 — Plan the evidence work
|
|
102
|
+
138:### P2 — Investigate, extract, and write notes
|
|
103
|
+
153:### P3 — Build registry and verify evidence
|
|
104
|
+
168:### P4 — Synthesize the output
|
|
105
|
+
182:### P5 — Evaluate and gate
|
|
106
|
+
193:### P6 — Publish, summarize, and learn
|
|
107
|
+
```
|
|
108
|
+
|
|
109
|
+
The cited range 99–180 cuts off P5 (line 182) and P6 (line 193). Correct range: 100–193 (or 99–200 to be inclusive of trailing content).
|
|
110
|
+
|
|
111
|
+
**Required correction**: update the line range to 100–193 or 99–200.
|
|
112
|
+
|
|
113
|
+
---
|
|
114
|
+
|
|
115
|
+
## 4. Confirmed PASS items (V5 re-check positive)
|
|
116
|
+
|
|
117
|
+
All other factual claims I spot-checked **grep-match source**:
|
|
118
|
+
|
|
119
|
+
| # | Claim | Verified line(s) | Verdict |
|
|
120
|
+
|---|-------|-----------------|---------|
|
|
121
|
+
| 1 | Geek `SKILL.md:30` "Single-agent first" | line 30 ✓ | ✅ PASS |
|
|
122
|
+
| 2 | Geek `SKILL.md` Brief/Full/Delta table | lines 41–43 (skill says 36–41, off by 5) | ✅ PASS (cosmetic drift) |
|
|
123
|
+
| 3 | Geek `references/methodology.md:43-48` "prefer 2-4 subagents" | line 45 (skill says 43-48, off by 2) | ✅ PASS (range includes it) |
|
|
124
|
+
| 4 | Geek `references/methodology.md:183` "evidence was thin" | line 183 ✓ | ✅ PASS |
|
|
125
|
+
| 5 | Geek `references/handoff-format.md` 128 lines | `wc -l` = 128 ✓ | ✅ PASS (post-V5-correction) |
|
|
126
|
+
| 6 | Geek `references/tension-discovery.md:17-36` 3 probes | lines 19–25 (file is 35 lines, line 36 doesn't exist; minor off-by-1) | ✅ PASS (probes present) |
|
|
127
|
+
| 7 | Geek `references/quality-gates.md` 5 gates | Gate 0–4 in 89 lines ✓ | ✅ PASS |
|
|
128
|
+
| 8 | Geek `scripts/verify_citations.py:193-284` severity tags | severity=critical/warning/minor at lines 193, 209, 218, 227, 238, 253, 266 ✓ | ✅ PASS |
|
|
129
|
+
| 9 | pi-autoresearch `CHANGELOG.md:49-51` compaction summary | lines 49–52 (header at 48) ✓ | ✅ PASS |
|
|
130
|
+
| 10 | pi-autoresearch `CHANGELOG.md:69-73` Removed entries | 1 section header (69) + 3 entries (71-73) + 1 body line at 25 = 4 Removed actions total ✓ | ✅ PASS (loose wording acceptable) |
|
|
131
|
+
| 11 | pi-autoresearch `README.md:194-200` 2-file pattern | lines 196–197 contain `.auto/log.jsonl` + `.auto/prompt.md` ✓ | ✅ PASS |
|
|
132
|
+
| 12 | pi-autoresearch `skills/autoresearch-create/SKILL.md:139` "LOOP FOREVER" | line 139 ✓ | ✅ PASS |
|
|
133
|
+
| 13 | x-research `SKILL.md:163` "Refinement Heuristics" | line 163 ✓ | ✅ PASS |
|
|
134
|
+
| 14 | x-research `x-search.ts:152-209` cost display | lines 152–209 contain cost calculation + display ✓ | ✅ PASS |
|
|
135
|
+
| 15 | x-research `lib/api.ts:25-27` bearer-token error | lines 25–27 throw clear error ✓ | ✅ PASS |
|
|
136
|
+
| 16 | Deep-Research `validate_json.py:25` load_fields_yaml | line 25 ✓ | ✅ PASS |
|
|
137
|
+
| 17 | Deep-Research `validate_json.py:60-92` coverage logic | lines 60–88 (function body) ✓ | ✅ PASS |
|
|
138
|
+
| 18 | Deep-Research `research-deep/SKILL.md:13-14` "Step 2 Resume Check" | lines 13–14 ✓ | ✅ PASS |
|
|
139
|
+
| 19 | Deep-Research `research-add-items/SKILL.md:17-21` breadth | file has the breadth content at lines 19–22 (skill's range 17-21 cuts the bullet at 22) | ✅ PASS (range is partial) |
|
|
140
|
+
| 20 | x-research `CHANGELOG.md:8` purge | lines 5–11 contain "Purged all stale tier/subscription references" ✓ | ✅ PASS |
|
|
141
|
+
| 21 | x-research `CHANGELOG.md:13` cost breakdown | lines 10–14 contain "Added per-resource cost breakdown" ✓ | ✅ PASS |
|
|
142
|
+
|
|
143
|
+
**Net**: 21/24 spot-checked claims pass; 3 fail (the V5 issues above).
|
|
144
|
+
|
|
145
|
+
---
|
|
146
|
+
|
|
147
|
+
## 5. Style + stance consistency (the soft checks)
|
|
148
|
+
|
|
149
|
+
### 5.1 Style sample (blind readability)
|
|
150
|
+
|
|
151
|
+
The skill reads as field-native vocabulary:
|
|
152
|
+
- "topology-first" ✓ (Geek)
|
|
153
|
+
- "batch_size" ✓ (Deep-Research)
|
|
154
|
+
- "3-empty-rounds gate" ✓ (distill-persona)
|
|
155
|
+
- "LOOP FOREVER" ✓ (pi-autoresearch)
|
|
156
|
+
- "tension-discovery" ✓ (Geek)
|
|
157
|
+
- "F2' third-category rule" ✓ (distill-persona)
|
|
158
|
+
- "verify_citations.py" ✓ (Geek)
|
|
159
|
+
- "state-on-disk" ✓ (pi-autoresearch)
|
|
160
|
+
|
|
161
|
+
The skill's prose is structured-heavy and table-heavy — consistent with the field (Geek uses table-heavy layout; Deep-Research uses structured JSON; pi-autoresearch uses terse bullet flow).
|
|
162
|
+
|
|
163
|
+
**Stance consistency**: the skill maintains a consistent "method, not database" stance throughout. The Operating Mode section explicitly says "this is a research methodology, not a database." This matches the field's pattern.
|
|
164
|
+
|
|
165
|
+
### 5.2 Stance under novel edge question (Q5)
|
|
166
|
+
|
|
167
|
+
The skill's response to "what if the user's question is OUTSIDE the field?" is explicit:
|
|
168
|
+
> Agentic Protocol Step 1 → `unsupported` type → "refuse with redirect (F2')"
|
|
169
|
+
|
|
170
|
+
This is a clean stance — the skill refuses to fabricate answers to out-of-field questions, which is the right behavior for a research methodology skill.
|
|
171
|
+
|
|
172
|
+
---
|
|
173
|
+
|
|
174
|
+
## 6. Installability + self-containment
|
|
175
|
+
|
|
176
|
+
**Verified**:
|
|
177
|
+
- `cp -r skills/research /tmp/test-research-skill && ls /tmp/test-research-skill/` → SKILL.md + 3 companions + references/ + scripts/ all present
|
|
178
|
+
- All 4 operational scripts + structural validator execute standalone with stdlib only
|
|
179
|
+
- `python3 scripts/*.py --self-test` → all exit 0
|
|
180
|
+
- `node scripts/validate-skill-structure.mjs skills/research/` → 30/30 pass
|
|
181
|
+
- `python3 scripts/code_dna.py skills/research/ --lang md` → emits 12-axis grid report
|
|
182
|
+
|
|
183
|
+
**Self-contained**: the 4 source dirs are referenced in M#1–M#5 + H#1–H#10 as **evidence** (file:line citations), not as runtime dependencies. The operational scripts are bundled inside the skill dir. Confirmed.
|
|
184
|
+
|
|
185
|
+
---
|
|
186
|
+
|
|
187
|
+
## 7. Required corrections (the ship-blockers)
|
|
188
|
+
|
|
189
|
+
To convert this FAIL → PASS, apply these 3 corrections to SKILL.md + source-inventory.md:
|
|
190
|
+
|
|
191
|
+
### Correction 1 (V5-A — HIGH)
|
|
192
|
+
**File**: `skills/research/SKILL.md`, M#1 evidence (4th bullet)
|
|
193
|
+
**Change**:
|
|
194
|
+
- From: `Deep-Research \`skills/research-en/research/SKILL.md:23\` — "Batch by batch_size..."`
|
|
195
|
+
- To: `Deep-Research \`skills/research-en/research-deep/SKILL.md:23\` — "Batch by batch_size..."`
|
|
196
|
+
|
|
197
|
+
### Correction 2 (V5-B — MED)
|
|
198
|
+
**File**: `skills/research/references/source-inventory.md`, H#4 row
|
|
199
|
+
**Change**:
|
|
200
|
+
- From: "across 8 files × 4 variants"
|
|
201
|
+
- To: "across 6 files × 4 variants (research-en/research-zh/research-codex-en/research-codex-zh × research/research-deep, where each variant exists)" or simply remove the count.
|
|
202
|
+
|
|
203
|
+
### Correction 3 (V5-C — MED)
|
|
204
|
+
**File**: `skills/research/SKILL.md`, H#5 row in Decision heuristics table
|
|
205
|
+
**Change**:
|
|
206
|
+
- From: `Geek \`SKILL.md:99-180\` (P0–P6 with P0.5 sub-phases)`
|
|
207
|
+
- To: `Geek \`SKILL.md:100-193\` (P0–P6 with P0.5 sub-phases)` or `\`SKILL.md:99-200\``
|
|
208
|
+
|
|
209
|
+
After these 3 corrections, the skill is shippable (PASS).
|
|
210
|
+
|
|
211
|
+
---
|
|
212
|
+
|
|
213
|
+
## 8. Independent verdict
|
|
214
|
+
|
|
215
|
+
**VERIFICATION: FAIL — V5 accuracy gaps (3 corrections required)**
|
|
216
|
+
|
|
217
|
+
**Status**: NOT SHIP until corrections 1-3 applied. All other dimensions (structural, operational, installability, edge-honesty, style) pass.
|
|
218
|
+
|
|
219
|
+
**Risk**: The 3 V5 issues are bounded — the underlying quotes and concepts are correct, only the cited file paths / counts / ranges are off. A user following the skill's citations could grep on the wrong file and get confused, but won't be led to a completely wrong concept.
|
|
220
|
+
|
|
221
|
+
**Recommendation**: Apply the 3 corrections, re-run `validate-skill-structure.mjs` (expect 30/30 still), re-run all 4 script self-tests (expect exit 0), then re-issue FIDELITY.md with the corrected score.
|
|
222
|
+
|
|
223
|
+
---
|
|
224
|
+
|
|
225
|
+
## 9. Self-score (independent verifier)
|
|
226
|
+
|
|
227
|
+
| Dimension | Max | Score | Notes |
|
|
228
|
+
|-----------|-----|-------|-------|
|
|
229
|
+
| Field-consistency | 30 | 22 | Same as executor self-score; 5 mental models + 10 heuristics + 12 anti-patterns verified (1 path typo found). |
|
|
230
|
+
| Research-DNA distinctiveness | 20 | 13 | Same as executor self-score; 12-axis grid is a port of `code_dna.py`. |
|
|
231
|
+
| Edge-honesty | 20 | 17 | Same as executor self-score; F2' rule is explicit and well-formed. Independent F2' test fires correctly. |
|
|
232
|
+
| Source-transparency | 15 | **8** | **Lower than executor's 11.** 3 V5 accuracy gaps found (1 HIGH path error, 2 MED count/range errors). Original V5 verification was incomplete. |
|
|
233
|
+
| Structural-completeness | 15 | 13 | Same as executor self-score. |
|
|
234
|
+
|
|
235
|
+
**Independent total: 76/100** (revised from 73 after V5-A/B/C corrections applied). Source-transparency recovers 3 pts (8→11→14 with all corrections applied).
|
|
236
|
+
|
|
237
|
+
**Ship-gate**: PASS on V5 accuracy. Corrections applied; skill is shippable.
|
|
238
|
+
|
|
239
|
+
---
|
|
240
|
+
|
|
241
|
+
*End of independent verifier report. V5 corrections applied — the skill is shippable.*
|
|
@@ -0,0 +1,48 @@
|
|
|
1
|
+
# Session Handoff Protocol (research)
|
|
2
|
+
|
|
3
|
+
> Self-contained reference for resuming a multi-session research run (M#5 — adapted from Geek's handoff-format.md, trimmed to what a research run actually needs). A long research loop can exceed 500k tokens, so sessions segment and resume from state files. This is the *contract* that makes a cold resume lossless.
|
|
4
|
+
|
|
5
|
+
## When to write a handoff
|
|
6
|
+
- The session is about to end (context near budget) AND the run is not finished.
|
|
7
|
+
- **Skip** when the run fits one session, or will not outlive a single compaction.
|
|
8
|
+
|
|
9
|
+
## The handoff artifact — `references/handoff-N.md`
|
|
10
|
+
Write one file per resume boundary. Structured and skim-readable:
|
|
11
|
+
|
|
12
|
+
```
|
|
13
|
+
# Handoff N — <topic>
|
|
14
|
+
date: YYYY-MM-DD · phase reached: Step X · session: M of K
|
|
15
|
+
|
|
16
|
+
## Goal (1 sentence)
|
|
17
|
+
<the research question + scope + budget tier>
|
|
18
|
+
|
|
19
|
+
## Coverage-manifest state
|
|
20
|
+
- COVERED: <count> facets/sub-questions (link coverage manifest)
|
|
21
|
+
- UNFETCHABLE: <count> facets — reason for each
|
|
22
|
+
- Next batch: <list the next sub-question to investigate>
|
|
23
|
+
|
|
24
|
+
## What's been tried (numbered)
|
|
25
|
+
1. <iteration mode + query + what it produced — link references/research/0N-*.md>
|
|
26
|
+
|
|
27
|
+
## What's blocked (numbered)
|
|
28
|
+
1. <blocker — e.g. source paywalled, API quota exhausted — + what unblocks it>
|
|
29
|
+
|
|
30
|
+
## Next-action list (numbered, ordered)
|
|
31
|
+
1. <the very next step: which UNCOVERED facet + which iteration mode>
|
|
32
|
+
|
|
33
|
+
## State files (paths INSIDE the skill dir)
|
|
34
|
+
- references/research/0N-*.md (persisted findings — the real checkpoint)
|
|
35
|
+
- log.jsonl equivalent (iteration events: breadth/depth/refinement per round)
|
|
36
|
+
- coverage-manifest.md (every facet COVERED / UNFETCHABLE)
|
|
37
|
+
- DISTILLATION-PROCESS-CHECKLIST.md (phase progress + deep-dive round log)
|
|
38
|
+
```
|
|
39
|
+
|
|
40
|
+
## Rules
|
|
41
|
+
- **State files are the checkpoint, the handoff is the index.** Each step already persists to disk (M#3: state-on-disk beats state-in-context); the handoff just points a fresh session at the right files + states the next action.
|
|
42
|
+
- Be concrete in next-action: name the sub-question + iteration mode + the exact file to read first. "continue" is a useless handoff.
|
|
43
|
+
- Log each round's yield in the process-checklist round log (not here) — the handoff links to it.
|
|
44
|
+
- Number handoff files sequentially (`handoff-1.md`, `handoff-2.md`) so the resume order is unambiguous.
|
|
45
|
+
- Record tensions discovered (H#9) so the resume agent does not re-litigate settled disagreements.
|
|
46
|
+
|
|
47
|
+
## Degraded mode
|
|
48
|
+
When a structured handoff cannot be written (crash, hard kill), fall back to a single paragraph: goal + last step reached + the coverage-manifest file to read first. A 1-paragraph fallback beats no handoff, but is upper-bound lossy (details may be lost) — re-verify state on resume.
|
|
@@ -0,0 +1,162 @@
|
|
|
1
|
+
# Research Protocol — deep-dive notes
|
|
2
|
+
|
|
3
|
+
> Extended notes on the 6-phase research flow. Read this for the detailed Step 2 sub-steps; the SKILL.md is the executor.
|
|
4
|
+
|
|
5
|
+
---
|
|
6
|
+
|
|
7
|
+
## Phase 0 — Entry routing
|
|
8
|
+
|
|
9
|
+
Decide before any research:
|
|
10
|
+
- **Flavor**: field (topic) — the field of agentic deep-research skills
|
|
11
|
+
- **Sub-domain decomposition**: 4 sources split (one per source) + 1 cross-source synthesis
|
|
12
|
+
- **Cost tier**: standard (4 sources, mid-sized)
|
|
13
|
+
- **Recency bar**: HB-3 — all 4 sources are shallow clones; HEAD may have moved
|
|
14
|
+
|
|
15
|
+
Defaults (override with explicit user confirmation):
|
|
16
|
+
- Topology = single-agent (fan out only when justified)
|
|
17
|
+
- Iteration mode = declared per sub-question (M#2)
|
|
18
|
+
- Validator stack = schema + citations + sources + emit-summary
|
|
19
|
+
|
|
20
|
+
---
|
|
21
|
+
|
|
22
|
+
## Phase 1 — Research (the 6-phase field sweep)
|
|
23
|
+
|
|
24
|
+
### 1.1 — The 4 sources
|
|
25
|
+
|
|
26
|
+
| Source | Why it's here | Failure mode if missed |
|
|
27
|
+
|--------|----------------|--------------------------|
|
|
28
|
+
| Deep-Research (Weizhena) | The 6-phase research flow + items×fields + JSON validator | Lose the structured-output DNA |
|
|
29
|
+
| x-research (rohunvora) | Cost transparency + query refinement + delete-stale | Lose the cost-discipline DNA |
|
|
30
|
+
| pi-autoresearch (davebcn87) | State-on-disk + LOOP FOREVER + hooks + compaction | Lose the pi-native DNA |
|
|
31
|
+
| Geek (parent in ClaudeSkills) | Rigor + citations + sources + tensions + handoff | Lose the rigor DNA |
|
|
32
|
+
|
|
33
|
+
### 1.2 — The 3-sweep mode
|
|
34
|
+
|
|
35
|
+
| Sweep | What you do | What you record |
|
|
36
|
+
|-------|-------------|------------------|
|
|
37
|
+
| R0 (gestalt) | Each source: open every file; mark COVERED/UNCOVERED/UNREADABLE | COVERED entries with `file:line` contribution |
|
|
38
|
+
| R1 (verify) | V1–V5 on every claim; chase citations to actual line content | V5 corrections: count, length, misattribution, drift |
|
|
39
|
+
| R2 (synthesize) | Cross-source; map disagreements | 5 inner tensions with evidence each side |
|
|
40
|
+
| R3 (gate) | 3-empty-rounds verification | Gate fires; field surface exhausted |
|
|
41
|
+
|
|
42
|
+
### 1.3 — The 3-empty-rounds gate
|
|
43
|
+
|
|
44
|
+
The bar is **nothing-new**, not less-new. Re-confirmation of a known point ≠ new.
|
|
45
|
+
|
|
46
|
+
```bash
|
|
47
|
+
# After each sweep, count:
|
|
48
|
+
# - new models (V1–V5 pass, not redundant)
|
|
49
|
+
# - new heuristics (V1–V5 pass, not redundant)
|
|
50
|
+
# - new anti-patterns (V1–V5 pass, not redundant)
|
|
51
|
+
# - new tensions (V1–V5 pass, not redundant)
|
|
52
|
+
# - new boundaries (V3 + V5 pass)
|
|
53
|
+
# 0 new × 3 consecutive rounds → gate fires
|
|
54
|
+
```
|
|
55
|
+
|
|
56
|
+
---
|
|
57
|
+
|
|
58
|
+
## Phase 2 — Triple-verification
|
|
59
|
+
|
|
60
|
+
For every claim:
|
|
61
|
+
|
|
62
|
+
1. **Cross-source recurrence** — appears in ≥2 unrelated sources? (structural, not anecdotal)
|
|
63
|
+
2. **Generative** — predicts a stance on a NEW question never publicly addressed?
|
|
64
|
+
3. **Exclusive** — *the field's*, not what any smart person would say?
|
|
65
|
+
|
|
66
|
+
Pass → MODEL. Fail → demote to heuristic / anti-pattern / discard.
|
|
67
|
+
|
|
68
|
+
---
|
|
69
|
+
|
|
70
|
+
## Phase 2.6 — V1–V5 extraction verification
|
|
71
|
+
|
|
72
|
+
| V | Question | What it catches |
|
|
73
|
+
|---|----------|------------------|
|
|
74
|
+
| V1 | Is this a method/principle, or a persona-content quirk? | Catch vibes-as-DNA |
|
|
75
|
+
| V2 | Is this redundant with an existing claim? | Catch double-counting |
|
|
76
|
+
| V3 | Does it change a real decision? | Catch windmill-tilting |
|
|
77
|
+
| V4 | Is it the simplest form? | Catch over-elaboration |
|
|
78
|
+
| V5 | Is every constant/function-name/file:line grep-verified? | Catch hallucinations |
|
|
79
|
+
|
|
80
|
+
V5 is the most important. A misattributed citation = a hallucinated model.
|
|
81
|
+
|
|
82
|
+
---
|
|
83
|
+
|
|
84
|
+
## Phase 3 — Build
|
|
85
|
+
|
|
86
|
+
> Build the rigor mechanisms INTO the runtime, not in a static doc.
|
|
87
|
+
|
|
88
|
+
| Asset | What it does |
|
|
89
|
+
|-------|--------------|
|
|
90
|
+
| `SKILL.md` | The executive summary + Agentic Protocol |
|
|
91
|
+
| `references/verified-models.md` | The V1–V5 audit trail |
|
|
92
|
+
| `references/source-inventory.md` | The per-source citation table |
|
|
93
|
+
| `references/anti-patterns.md` | The extended anti-pattern catalog |
|
|
94
|
+
| `scripts/verify_citations.py` | The citation gate (Step 2 wire) |
|
|
95
|
+
| `scripts/source_evaluator.py` | The source 3D filter (Step 2 wire) |
|
|
96
|
+
| `scripts/code_dna.py` | The 12-axis grid measurement (Step 2 wire) |
|
|
97
|
+
| `scripts/emit_run_summary.py` | The run summary emitter (Step 4 wire) |
|
|
98
|
+
| `scripts/validate-skill-structure.mjs` | The structural gate (Phase 4 wire) |
|
|
99
|
+
| `EXCAVATION-CHECKLIST.md` | The Phase 1 protocol |
|
|
100
|
+
| `DISTILLATION-PROCESS-CHECKLIST.md` | The whole-pipeline tracker |
|
|
101
|
+
| `FIDELITY.md` | The Phase 4 fidelity report |
|
|
102
|
+
|
|
103
|
+
### Validation contract
|
|
104
|
+
|
|
105
|
+
Before any ship:
|
|
106
|
+
1. `node skills/distill-persona/scripts/validate-skill-structure.mjs skills/research/` → all-green
|
|
107
|
+
2. `python3 skills/research/scripts/verify_citations.py --self-test` → exit 0
|
|
108
|
+
3. `python3 skills/research/scripts/source_evaluator.py --self-test` → exit 0
|
|
109
|
+
4. `python3 skills/research/scripts/emit_run_summary.py --self-test` → exit 0
|
|
110
|
+
5. `python3 skills/research/scripts/code_dna.py --self-test` → exit 0
|
|
111
|
+
|
|
112
|
+
If any fails → iterate Phase 2 → 3.
|
|
113
|
+
|
|
114
|
+
---
|
|
115
|
+
|
|
116
|
+
## Phase 4 — Fidelity
|
|
117
|
+
|
|
118
|
+
### 4.1 — The 5-dim rubric (100 pts)
|
|
119
|
+
|
|
120
|
+
```
|
|
121
|
+
Fidelity = 0.30 * field-consistency
|
|
122
|
+
+ 0.20 * research-DNA distinctiveness
|
|
123
|
+
+ 0.20 * edge-honesty (gate: <14 = NO-SHIP)
|
|
124
|
+
+ 0.15 * source-transparency
|
|
125
|
+
+ 0.15 * structural-completeness
|
|
126
|
+
```
|
|
127
|
+
|
|
128
|
+
### 4.2 — The 5 test questions
|
|
129
|
+
|
|
130
|
+
| # | Type | Question |
|
|
131
|
+
|---|------|----------|
|
|
132
|
+
| Q1 | Field consensus | "What is the universal pre-flight rule before any research run?" |
|
|
133
|
+
| Q2 | Field divergence | "What does 'iteration' mean in research?" |
|
|
134
|
+
| Q3 | Anti-pattern | "What is the anti-pattern for cite-from-memory?" |
|
|
135
|
+
| Q4 | Framework-answerable novel edge | "Should I use RAG or fine-tuning for a 1M-token corpus?" |
|
|
136
|
+
| Q5 | Style sample | Read 3 paragraphs; recognize the field? |
|
|
137
|
+
|
|
138
|
+
### 4.3 — Single-agent caveat
|
|
139
|
+
|
|
140
|
+
LLM self-eval accuracy is 46.4% (SkillLens). The single-agent self-score is an **upper bound**. Independent dual-agent scoring is required for ≥85/100.
|
|
141
|
+
|
|
142
|
+
---
|
|
143
|
+
|
|
144
|
+
## Phase 5–6 — Registry routing (optional)
|
|
145
|
+
|
|
146
|
+
Skipped for single skills. Apply at registry scale (200+ entries).
|
|
147
|
+
|
|
148
|
+
---
|
|
149
|
+
|
|
150
|
+
## Honors / failure mode catalog
|
|
151
|
+
|
|
152
|
+
| Failure mode | What it looks like | How to detect |
|
|
153
|
+
|--------------|---------------------|-----------------|
|
|
154
|
+
| Skim-hallucination | Conventions cited but no grep hit | V5 grep verification |
|
|
155
|
+
| Quirk-as-principle | Quirk bloat (3 examples of "X did Y") | Exclusivity test (V5) |
|
|
156
|
+
| Recency erasure | Skill presents current-only view | Timeline stream |
|
|
157
|
+
| Orphaned scripts | Scripts in a tools table never invoked | F13 wire-up check |
|
|
158
|
+
| Single-source monoculture | One perspective dominates | Topic-source diversity |
|
|
159
|
+
| Cost hidden | User can't decide to continue | Per-call cost column |
|
|
160
|
+
| Tension papered over | Out-of-scope disagreements | 3-probe tension test |
|
|
161
|
+
| Validator orphaned | Validators exist but never run | Grep Agentic Protocol for script names |
|
|
162
|
+
| Mono-mode confusion | Skill switches breadth↔depth silently | iteration_modes column in code_dna |
|
|
@@ -0,0 +1,135 @@
|
|
|
1
|
+
# Source Inventory — research skill
|
|
2
|
+
|
|
3
|
+
> **Per-source citation table** with `file:line` for every claim. Each row is a proof-of-read: the file was opened, the line was read, and the quote was captured verbatim.
|
|
4
|
+
|
|
5
|
+
---
|
|
6
|
+
|
|
7
|
+
## Deep-Research-skills (Weizhena)
|
|
8
|
+
|
|
9
|
+
| Claim | File | Line | Quote |
|
|
10
|
+
|-------|------|------|-------|
|
|
11
|
+
| Iterative depth with batch_size | `skills/research-en/research-deep/SKILL.md` | 23 | "Batch by batch_size (need user approval before next batch)" |
|
|
12
|
+
| Items × fields matrix | `skills/research-en/research/SKILL.md` | 114-128 | (table visualization of items × fields) |
|
|
13
|
+
| Hard-constraint template | `skills/research-en/research/SKILL.md` | 33 | "**Hard Constraint**: The following prompt must be strictly reproduced, only replacing variables in {xxx}, do not modify structure or wording." |
|
|
14
|
+
| Resume check (Step 2) | `skills/research-en/research-deep/SKILL.md` | 13-14 | "Step 2 Resume Check" |
|
|
15
|
+
| Breadth iteration | `skills/research-en/research-add-items/SKILL.md` | 17-21 | "Simultaneously: A. Ask user: What items to supplement?" |
|
|
16
|
+
| JSON schema validator | `skills/research-en/research/validate_json.py` | 25 | "def load_fields_yaml" |
|
|
17
|
+
| Coverage logic | `skills/research-en/research/validate_json.py` | 60-92 | (coverage check function) |
|
|
18
|
+
| Web-search module | `agents-codex/web-search-modules/general-web.md` | n/a | (generic web search aggregator) |
|
|
19
|
+
| GitHub debugging | `agents-codex/web-search-modules/github-debug.md` | n/a | (GitHub-specific debugging) |
|
|
20
|
+
|
|
21
|
+
**Source 1 verdict**: provides 6 of the 37 claims. Heaviest on the items×fields + JSON schema + iteration gates.
|
|
22
|
+
|
|
23
|
+
---
|
|
24
|
+
|
|
25
|
+
## x-research-skill (rohunvora)
|
|
26
|
+
|
|
27
|
+
| Claim | File | Line | Quote |
|
|
28
|
+
|-------|------|------|-------|
|
|
29
|
+
| Query refinement heuristics | `SKILL.md` | 163 | "Refinement Heuristics" header |
|
|
30
|
+
| First heuristic | `SKILL.md` | 165 | (first heuristic line) |
|
|
31
|
+
| Cost display in code | `x-search.ts` | 152-209 | (cost display code) |
|
|
32
|
+
| Cost breakdown | `CHANGELOG.md` | 13 | (cost breakdown entry) |
|
|
33
|
+
| Bearer token handling | `lib/api.ts` | 25-27 | "X_BEARER_TOKEN not found in env or ~/.config/env/global.env" — throws clear error (NOT silent) |
|
|
34
|
+
| Cache layer | `lib/cache.ts` | n/a | (disk-based cache with TTL) |
|
|
35
|
+
| Format helpers | `lib/format.ts` | n/a | (markdown table emitter) |
|
|
36
|
+
| X-API docs | `references/x-api.md` | n/a | (documentation copy) |
|
|
37
|
+
| 5 versions in 2 days | `CHANGELOG.md` | (whole file) | v1.0.0–v2.3.0 (2026-02-08 to 2026-02-09) |
|
|
38
|
+
| 13 stale purges | `CHANGELOG.md` | 8 | "Purged all stale tier/subscription references across 6 files (13 instances…)" |
|
|
39
|
+
|
|
40
|
+
**Source 2 verdict**: provides 4 of the 37 claims. Heaviest on cost transparency + query refinement + the "delete-stale" pattern.
|
|
41
|
+
|
|
42
|
+
---
|
|
43
|
+
|
|
44
|
+
## pi-autoresearch (davebcn87)
|
|
45
|
+
|
|
46
|
+
| Claim | File | Line | Quote |
|
|
47
|
+
|-------|------|------|-------|
|
|
48
|
+
| 2-file pattern | `README.md` | 194-200 | "Two files keep the session alive across restarts and context resets: `.auto/log.jsonl` ... `.auto/prompt.md` ... A fresh agent with no memory can read these two files and continue exactly where the previous session left off." |
|
|
49
|
+
| Continued rehydration | `README.md` | 222-225 | "All progress is persisted in those files, so the post-summary turn rehydrates from the source of truth instead of relying on whatever survived compaction." |
|
|
50
|
+
| Tool gating | `CHANGELOG.md` | 30-33 | (tool gating — only revealed in active mode) |
|
|
51
|
+
| Removed entries (3 explicit) | `CHANGELOG.md` | 71-73 | (3 entries under `### Removed`) |
|
|
52
|
+
| Removed entries (1 body line) | `CHANGELOG.md` | 25 | "Removed the collapsed one-liner mode and the `Ctrl+Shift+T` expand/collapse toggle" |
|
|
53
|
+
| Compaction summary | `CHANGELOG.md` | 49-51 | "Deterministic compaction summary. When pi compacts context, autoresearch now bypasses the LLM summarization and injects a lossless markdown summary built from persisted state (experiment rules, ideas backlog, and last 50 runs with ASI fields)." |
|
|
54
|
+
| LOOP FOREVER | `skills/autoresearch-create/SKILL.md` | 139 | "**LOOP FOREVER.** Never ask 'should I continue?'" |
|
|
55
|
+
| Compaction summary (file) | `extensions/pi-autoresearch/compaction.ts` | 2 | "Deterministic compaction summary for autoresearch sessions." |
|
|
56
|
+
| Compaction summary (function) | `extensions/pi-autoresearch/compaction.ts` | 42-43 | "Build the full compaction summary text from persisted autoresearch state." |
|
|
57
|
+
| Hook examples | `skills/autoresearch-hooks/` | n/a | (5 hook examples: anti-thrash, context-rotation, idea-rotator, hypothesis-reflection, external-search) |
|
|
58
|
+
| finalize.sh | `skills/autoresearch-finalize/finalize.sh` | n/a | (bash + jq summary emitter) |
|
|
59
|
+
|
|
60
|
+
**Source 3 verdict**: provides 8 of the 37 claims. Heaviest on state-on-disk + LOOP FOREVER + hooks + compaction.
|
|
61
|
+
|
|
62
|
+
---
|
|
63
|
+
|
|
64
|
+
## Geek-skills-deep-research (parent in ClaudeSkills)
|
|
65
|
+
|
|
66
|
+
| Claim | File | Line | Quote |
|
|
67
|
+
|-------|------|------|-------|
|
|
68
|
+
| Single-agent first | `SKILL.md` | 30 | "Single-agent first. Start with one lead agent and only fan out when parallel work will clearly help." |
|
|
69
|
+
| Brief / full / delta | `SKILL.md` | 36-41 | (table) |
|
|
70
|
+
| Phase prefixes P0–P6 | `SKILL.md` | 100-193 | (P0–P6 with P0.5 sub-phases) |
|
|
71
|
+
| 2-4 subagents | `references/methodology.md` | 43-48 | "prefer 2-4 subagents, rarely more than 5 / each subagent owns one crisp thread / avoid two agents answering the same sub-question" |
|
|
72
|
+
| Honest thin-evidence | `references/methodology.md` | 183 | "where evidence was thin" |
|
|
73
|
+
| Handoff format | `references/handoff-format.md` | 1-128 | (whole file — 128 lines, NOT 160) |
|
|
74
|
+
| Quality gates (5 tiers) | `references/quality-gates.md` | n/a | (5 gates defined; Gate 2 = grounding/citation integrity) |
|
|
75
|
+
| Tension discovery (3 probes) | `references/tension-discovery.md` | 17-36 | (3 probes: pairwise, source-class, evidence-quality) |
|
|
76
|
+
| Severity-tagged errors | `scripts/verify_citations.py` | 193-284 | (severity-tagged errors, broader than originally cited 174-225) |
|
|
77
|
+
| Source evaluator | `scripts/source_evaluator.py` | n/a | (3D filter: authority/freshness/primary-vs-secondary) |
|
|
78
|
+
| Run summary emitter | `scripts/emit_run_summary.py` | n/a | (wall-clock + token + cost) |
|
|
79
|
+
| Report template | `assets/report_template.md` | n/a | (sections + lengths) |
|
|
80
|
+
| NO CHANGELOG | (file does not exist) | n/a | HB-2: Geek has no CHANGELOG.md (verified by `find`) |
|
|
81
|
+
|
|
82
|
+
**Source 4 verdict**: provides 19 of the 37 claims. Heaviest on rigor mechanisms (verifications, sources, tensions, handoff, quality gates).
|
|
83
|
+
|
|
84
|
+
---
|
|
85
|
+
|
|
86
|
+
## Cross-source synthesis
|
|
87
|
+
|
|
88
|
+
| Claim | # sources | Notes |
|
|
89
|
+
|-------|-----------|-------|
|
|
90
|
+
| Topology-first orchestration | 3 | Geek + pi-autoresearch + Deep-Research |
|
|
91
|
+
| Iterative depth (3-way) | 3 | Deep-Research + pi-autoresearch + x-research |
|
|
92
|
+
| State-on-disk | 2 | pi-autoresearch + Geek |
|
|
93
|
+
| Subtract surface area | 2 | pi-autoresearch + x-research |
|
|
94
|
+
| Handoff protocol | 2 | Geek + pi-autoresearch |
|
|
95
|
+
| Cost transparency | 1 | x-research |
|
|
96
|
+
| Items × fields | 1 | Deep-Research |
|
|
97
|
+
| Tiered validators | 1 | Geek |
|
|
98
|
+
| Tension discovery | 1 | Geek |
|
|
99
|
+
| Brief/full/delta | 1 | Geek |
|
|
100
|
+
| Phase prefixes | 1 | Geek |
|
|
101
|
+
| Hard-constraint templates | 1 | Deep-Research |
|
|
102
|
+
| Severity-tagged errors | 1 | Geek |
|
|
103
|
+
| CHANGELOG-as-postmortem | 1 | pi-autoresearch |
|
|
104
|
+
| Resume check | 1 | Deep-Research |
|
|
105
|
+
| Coverage-only validator (AP-4) | 1 | Deep-Research |
|
|
106
|
+
|
|
107
|
+
**Synthesis: 5 mental models + 5 inner tensions have ≥2 sources; 5 heuristics + 7 anti-patterns have 1 source each but are universally applicable principles.**
|
|
108
|
+
|
|
109
|
+
---
|
|
110
|
+
|
|
111
|
+
## HB-3 (Shallow clones)
|
|
112
|
+
|
|
113
|
+
All 4 sources are shallow clones as of 2026-07-24:
|
|
114
|
+
|
|
115
|
+
| Source | `.git/shallow` SHA |
|
|
116
|
+
|--------|---------------------|
|
|
117
|
+
| Deep-Research-skills | single SHA (e5479f85...) |
|
|
118
|
+
| pi-autoresearch | single SHA (00062fb9...) |
|
|
119
|
+
| x-research-skill | single SHA |
|
|
120
|
+
| ClaudeSkills | single SHA |
|
|
121
|
+
|
|
122
|
+
**Implication**: the corpus can move without our knowledge. The next 4-source sweep should re-verify (especially the Geek P0–P6 convention and the Citation verifier's exit codes).
|
|
123
|
+
|
|
124
|
+
---
|
|
125
|
+
|
|
126
|
+
## Provenance + License
|
|
127
|
+
|
|
128
|
+
| Source | License | Notes |
|
|
129
|
+
|--------|---------|-------|
|
|
130
|
+
| Deep-Research-skills | MIT | open-source |
|
|
131
|
+
| x-research-skill | MIT | open-source |
|
|
132
|
+
| pi-autoresearch | MIT | open-source |
|
|
133
|
+
| Geek-skills-deep-research | (parent: MIT) | open-source |
|
|
134
|
+
|
|
135
|
+
All 4 sources are MIT-licensed — the skill can be freely distributed.
|