teamai-cli 0.25.0 → 0.26.0-beta.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +13 -0
- package/README.zh-CN.md +6 -0
- package/dist/index.js +6147 -3346
- package/package.json +4 -1
- package/skill-data/core/SKILL.md +114 -0
- package/skill-data/core/references/commands.md +339 -0
- package/{skills/teamai → skill-data/core}/references/contribute-member.md +13 -10
- package/{skills/teamai → skill-data/core}/references/troubleshooting.md +9 -1
- package/skill-data/setup/SKILL.md +76 -0
- package/{skills/teamai → skill-data/setup}/references/join-member.md +17 -14
- package/{skills/teamai → skill-data/setup}/references/manage-admin.md +18 -6
- package/{skills/teamai → skill-data/setup}/references/provider-tgit.md +9 -6
- package/{skills/teamai → skill-data/setup}/references/setup-admin.md +41 -35
- package/skill-data/share/SKILL.md +70 -0
- package/skill-data/share/references/doc-template.md +44 -0
- package/skill-data/wiki/SKILL.md +314 -0
- package/skill-data/wiki/references/agents/graph-rag-agent.md +344 -0
- package/skill-data/wiki/references/agents/kb-doc-generator.md +323 -0
- package/skill-data/wiki/references/methodology/phase0-collection.md +54 -0
- package/skill-data/wiki/references/methodology/phase1-reverse-engineering.md +89 -0
- package/skill-data/wiki/references/methodology/phase2-document-types.md +341 -0
- package/skill-data/wiki/references/methodology/phase3-ai-enhancement.md +164 -0
- package/skill-data/wiki/references/methodology/phase4-quality.md +232 -0
- package/skill-data/wiki/references/overview.md +124 -0
- package/skill-data/wiki/references/phases/k1-reverse-engineering.md +118 -0
- package/skill-data/wiki/references/phases/k2-documents.md +68 -0
- package/skill-data/wiki/references/phases/k3-ai-native.md +121 -0
- package/skill-data/wiki/references/phases/k4-quality.md +190 -0
- package/skill-data/wiki/references/phases/phase0-init.md +112 -0
- package/skill-data/wiki/references/templates/project-overview.md +148 -0
- package/{skills/team-wiki-codebase → skill-data/wiki}/scripts/scan_repo.py +52 -52
- package/{skills/team-wiki-codebase → skill-data/wiki}/scripts/validate_kb.py +68 -62
- package/skills/teamai/SKILL.md +28 -128
- package/skills/team-wiki-codebase/README.md +0 -121
- package/skills/team-wiki-codebase/SKILL.md +0 -905
- package/skills/team-wiki-codebase/references/agents/graph-rag-agent.md +0 -344
- package/skills/team-wiki-codebase/references/agents/kb-doc-generator.md +0 -323
- package/skills/team-wiki-codebase/references/methodology/phase0-collection.md +0 -54
- package/skills/team-wiki-codebase/references/methodology/phase1-reverse-engineering.md +0 -89
- package/skills/team-wiki-codebase/references/methodology/phase2-document-types.md +0 -341
- package/skills/team-wiki-codebase/references/methodology/phase3-ai-enhancement.md +0 -164
- package/skills/team-wiki-codebase/references/methodology/phase4-quality.md +0 -232
- package/skills/team-wiki-codebase/references/templates/project-overview.md +0 -148
- package/skills/teamai-share-learnings/SKILL.md +0 -87
- /package/{skills/teamai → skill-data/setup}/references/uninstall.md +0 -0
|
@@ -0,0 +1,232 @@
|
|
|
1
|
+
# Phase 4: Quality Assessment and Iterative Improvement
|
|
2
|
+
|
|
3
|
+
> Helper tool: `python3 "{SKILL_DIR}/scripts/validate_kb.py" <output_dir>` automatically checks link integrity, anchor coverage, AI Quick Reference table coverage, bidirectional links, and README index inclusion rate
|
|
4
|
+
|
|
5
|
+
## Five-Dimension Assessment Model
|
|
6
|
+
|
|
7
|
+
| Dimension | Weight | Passing standard |
|
|
8
|
+
|------|------|---------|
|
|
9
|
+
| **Coverage** | 25% | ≥ 90% of core components are documented |
|
|
10
|
+
| **Depth** | 25% | ≥ 80% of code entries can be located directly |
|
|
11
|
+
| **Consistency** | 20% | 0 dead links, 0 contradictory descriptions |
|
|
12
|
+
| **AI usability** | 20% | RAG retrieval accuracy ≥ 85% |
|
|
13
|
+
| **Freshness** | 10% | Core document update lag ≤ 30 days |
|
|
14
|
+
|
|
15
|
+
## Coverage Check
|
|
16
|
+
|
|
17
|
+
```
|
|
18
|
+
□ Does every code repository have a corresponding component design document?
|
|
19
|
+
□ Does every core API have a product-to-code mapping?
|
|
20
|
+
□ Does every data table have a schema description in some document?
|
|
21
|
+
□ Is every MQ Exchange/Topic/Queue marked in the topology diagram?
|
|
22
|
+
□ Is every error code in the mapping table?
|
|
23
|
+
□ Is every config item in the configuration description?
|
|
24
|
+
□ Is every scheduled task described in some document?
|
|
25
|
+
```
|
|
26
|
+
|
|
27
|
+
## RAG Retrieval Test Cases
|
|
28
|
+
|
|
29
|
+
| Test type | Example question | Expected hit |
|
|
30
|
+
|---------|---------|---------|
|
|
31
|
+
| Component location | "Where is the code entry of {component}?" | Component design document |
|
|
32
|
+
| Flow tracing | "What is the internal call chain of {API name}?" | Product-to-code mapping |
|
|
33
|
+
| Constraint query | "What is the batch cap of {operation}?" | Rules cheat sheet |
|
|
34
|
+
| State query | "Which operations can be executed in state {state}?" | State mutual exclusion rules |
|
|
35
|
+
| Error investigation | "How do I investigate {error code}?" | Anti-patterns / troubleshooting records |
|
|
36
|
+
| Code generation | "Write a Handler for {feature}" | SOP + interface contracts |
|
|
37
|
+
| Concept disambiguation | "What is the difference between {A} and {B}?" | Product knowledge library |
|
|
38
|
+
|
|
39
|
+
## Incremental Update Trigger Table
|
|
40
|
+
|
|
41
|
+
| Trigger condition | Update action |
|
|
42
|
+
|---------|---------|
|
|
43
|
+
| New code repository | Generate a Type-4 component document |
|
|
44
|
+
| API interface change | Update the Type-5 mapping + Type-6 cheat sheet |
|
|
45
|
+
| New product feature | Update the Type-2 business architecture + Type-8a knowledge library |
|
|
46
|
+
| Production incident | Add a Type-8d troubleshooting record + update Type-8b anti-patterns |
|
|
47
|
+
| Architecture refactoring | Update the Type-1 architecture overview + affected Type-4 documents |
|
|
48
|
+
| Config change | Update the configuration section of the corresponding component document |
|
|
49
|
+
|
|
50
|
+
## Version Management Convention
|
|
51
|
+
|
|
52
|
+
Maintain a change log at the bottom of every document:
|
|
53
|
+
|
|
54
|
+
```markdown
|
|
55
|
+
## 📝 Document Change Log
|
|
56
|
+
|
|
57
|
+
### vX.Y (YYYY-MM-DD)
|
|
58
|
+
- ✅ **Added**: {description of added content}
|
|
59
|
+
- ✅ **Fixed**: {description of fixed content}
|
|
60
|
+
- ✅ **Updated**: {description of updated content}
|
|
61
|
+
- ⚠️ **Deprecated**: {description of deprecated content}
|
|
62
|
+
```
|
|
63
|
+
|
|
64
|
+
## Fixing Common Quality Issues
|
|
65
|
+
|
|
66
|
+
| Issue | Fix method |
|
|
67
|
+
|------|---------|
|
|
68
|
+
| Dead links | Grep `](` links globally, or run `python3 "{SKILL_DIR}/scripts/validate_kb.py" <output_dir>` |
|
|
69
|
+
| Inconsistent terminology | Build a glossary and replace globally |
|
|
70
|
+
| Outdated code entries | Diff against the code repositories periodically |
|
|
71
|
+
| Outdated constraint values | Cross-check against the product docs periodically |
|
|
72
|
+
| AI retrieval failures | Add search-anchor keywords |
|
|
73
|
+
| Isolated documents | Add bidirectional links |
|
|
74
|
+
|
|
75
|
+
---
|
|
76
|
+
|
|
77
|
+
## Complete Generation Pipeline Checklist
|
|
78
|
+
|
|
79
|
+
### Phase 0 Checklist: Source Material Collection
|
|
80
|
+
|
|
81
|
+
```
|
|
82
|
+
□ All core code repositories cloned
|
|
83
|
+
□ Product API docs collected (interface name / inputs / outputs / error codes)
|
|
84
|
+
□ Product usage docs collected (usage limits / FAQ / billing description)
|
|
85
|
+
□ Database schema extracted (DDL / table schemas)
|
|
86
|
+
□ Workflow orchestration configs extracted (workflow_config etc.)
|
|
87
|
+
□ Proto/IDL files extracted
|
|
88
|
+
□ Error code definitions extracted
|
|
89
|
+
```
|
|
90
|
+
|
|
91
|
+
### Phase 1 Checklist: Architecture Reverse-Engineering
|
|
92
|
+
|
|
93
|
+
```
|
|
94
|
+
□ Code knowledge graph built (nodes + edges)
|
|
95
|
+
□ Architecture layers determined (≥4 layers)
|
|
96
|
+
□ Component relationship matrix built (N×N)
|
|
97
|
+
□ Core call chains traced (≥5 core APIs)
|
|
98
|
+
□ MQ topology inferred (Exchange/Topic/Queue/Routing Key)
|
|
99
|
+
□ Database ER model built
|
|
100
|
+
□ Glossary compiled (external-to-internal mappings)
|
|
101
|
+
```
|
|
102
|
+
|
|
103
|
+
### Phase 2 Checklist: Document Generation
|
|
104
|
+
|
|
105
|
+
```
|
|
106
|
+
□ [Type-1] Technical architecture overview document (1)
|
|
107
|
+
□ Includes the reader navigation guide
|
|
108
|
+
□ Includes AI retrieval routing rules
|
|
109
|
+
□ Includes core call chain sequence diagrams (≥5)
|
|
110
|
+
□ Includes the component relationship matrix
|
|
111
|
+
□ Includes the AI-only chapter 9
|
|
112
|
+
□ Includes the glossary
|
|
113
|
+
|
|
114
|
+
□ [Type-2] Business architecture document (1)
|
|
115
|
+
□ Includes the product capability matrix
|
|
116
|
+
□ Includes the billing model (if applicable)
|
|
117
|
+
□ Includes the core entity lifecycle state machine
|
|
118
|
+
|
|
119
|
+
□ [Type-3] Deployment architecture document (1)
|
|
120
|
+
□ Includes the service deployment matrix
|
|
121
|
+
□ Includes environment configuration
|
|
122
|
+
|
|
123
|
+
□ [Type-4] Component design documents (N)
|
|
124
|
+
□ Each includes an AI Quick Reference table
|
|
125
|
+
□ Each includes bidirectional links
|
|
126
|
+
□ Each includes code entries (precise to the function)
|
|
127
|
+
□ Each includes an architecture diagram (ASCII Art)
|
|
128
|
+
□ Each includes core flow descriptions
|
|
129
|
+
|
|
130
|
+
□ [Type-5] Product-to-code mapping document
|
|
131
|
+
□ Covers all core APIs
|
|
132
|
+
□ Each API includes a constraint table
|
|
133
|
+
□ Each API includes a call chain
|
|
134
|
+
□ Each API includes an error code mapping
|
|
135
|
+
|
|
136
|
+
□ [Type-6] Product rules cheat sheet
|
|
137
|
+
□ Covers all rule categories
|
|
138
|
+
□ Constraint values are exact
|
|
139
|
+
□ Includes the state mutual exclusion matrix
|
|
140
|
+
|
|
141
|
+
□ [Type-7] Business development SOP
|
|
142
|
+
□ Includes runnable code templates
|
|
143
|
+
□ Includes the error code mapping table
|
|
144
|
+
□ Includes the AI review checklist
|
|
145
|
+
|
|
146
|
+
□ [Type-8] Knowledge enhancement documents
|
|
147
|
+
□ [8a] Product knowledge library (concept disambiguation)
|
|
148
|
+
□ [8b] Anti-patterns and pitfalls guide
|
|
149
|
+
□ [8c] RPC interface contracts
|
|
150
|
+
□ [8d] Troubleshooting case records
|
|
151
|
+
```
|
|
152
|
+
|
|
153
|
+
### Phase 3 Checklist: AI-Native Enhancement
|
|
154
|
+
|
|
155
|
+
```
|
|
156
|
+
□ All component documents include an AI Quick Reference table
|
|
157
|
+
□ The Technical Architecture document includes retrieval routing rules
|
|
158
|
+
□ All documents include a search-anchor
|
|
159
|
+
□ Bidirectional link network complete (0 dead links)
|
|
160
|
+
□ QA pairs generated (10~20)
|
|
161
|
+
□ Document priorities defined
|
|
162
|
+
```
|
|
163
|
+
|
|
164
|
+
### Phase 3b Checklist: Graph Document Set (Graph RAG)
|
|
165
|
+
|
|
166
|
+
```
|
|
167
|
+
□ [G1] Component Dependency Matrix
|
|
168
|
+
□ N×N communication matrix complete
|
|
169
|
+
□ Forward/reverse dependency index
|
|
170
|
+
□ External service dependencies
|
|
171
|
+
|
|
172
|
+
□ [G2] Component Call Chain Overview + state machine
|
|
173
|
+
□ End-to-end core API chains (read + write)
|
|
174
|
+
□ Complete mermaid state machine diagram
|
|
175
|
+
□ Core state field value transition path table (if internal state codes exist)
|
|
176
|
+
□ User-visible state ↔ internal state mapping (if multi-layer states exist)
|
|
177
|
+
□ Operation-state constraint quick lookup matrix (✅/❌)
|
|
178
|
+
□ AI state reasoning rules
|
|
179
|
+
|
|
180
|
+
□ [G3] Data Flow and Storage Dependencies
|
|
181
|
+
□ Storage system dependency matrix
|
|
182
|
+
□ MQ queue topology
|
|
183
|
+
□ Cache strategy matrix
|
|
184
|
+
|
|
185
|
+
□ [G4] Error Code Component Map
|
|
186
|
+
□ Error code range allocation table
|
|
187
|
+
□ External → internal error code mapping
|
|
188
|
+
|
|
189
|
+
□ [G5] Cross-Component Interaction Scenarios
|
|
190
|
+
□ mermaid sequence diagrams for ≥10 scenarios
|
|
191
|
+
□ Every scenario has exception handling
|
|
192
|
+
|
|
193
|
+
□ [G6] Knowledge Graph Triples
|
|
194
|
+
□ Ontology definition (entity types + relationship types)
|
|
195
|
+
□ ≥100 explicit triples
|
|
196
|
+
□ Multi-hop dependency path index
|
|
197
|
+
□ Reverse reachability index
|
|
198
|
+
|
|
199
|
+
□ [G7] Architecture Risks and Impact Analysis
|
|
200
|
+
□ Component risk level summary table
|
|
201
|
+
□ Blast radius analysis (≥3 key components)
|
|
202
|
+
□ Cluster analysis
|
|
203
|
+
□ Change risk assessment matrix
|
|
204
|
+
|
|
205
|
+
□ [G8] Core Config Parameter Index
|
|
206
|
+
□ Layered config architecture diagram (mermaid)
|
|
207
|
+
□ Config parameter table per layer (config item / default / behavior impact / change risk / activation)
|
|
208
|
+
□ Config change impact surface quick lookup matrix
|
|
209
|
+
|
|
210
|
+
□ [G9] Business Rule Constraint Matrix
|
|
211
|
+
□ Operation precondition matrix
|
|
212
|
+
□ Detailed hardware constraint table
|
|
213
|
+
□ Migration constraint decision tree (mermaid)
|
|
214
|
+
□ Detailed billing constraint table
|
|
215
|
+
□ Special instance type constraint summary (✅/❌/⚠️)
|
|
216
|
+
□ AI reasoning rules quick lookup (mermaid flowchart)
|
|
217
|
+
|
|
218
|
+
□ Graph directory README.md index complete
|
|
219
|
+
□ Lookup-by-question-type table
|
|
220
|
+
□ Retrieval routing rule suggestions
|
|
221
|
+
```
|
|
222
|
+
|
|
223
|
+
### Phase 4 Checklist: Quality Assessment
|
|
224
|
+
|
|
225
|
+
```
|
|
226
|
+
□ Coverage ≥ 90%
|
|
227
|
+
□ Code entry precision ≥ 80%
|
|
228
|
+
□ Dead links = 0 (confirm by running validate_kb.py)
|
|
229
|
+
□ RAG retrieval accuracy ≥ 85%
|
|
230
|
+
□ Core document update lag ≤ 30 days
|
|
231
|
+
□ Terminology consistency check passed
|
|
232
|
+
```
|
|
@@ -0,0 +1,124 @@
|
|
|
1
|
+
# wiki: AI cognition engineering for large codebases
|
|
2
|
+
|
|
3
|
+
> TeamAI built-in skill. The methodology, scripts and agent specifications are **not** copied into `.claude/`, `.codebuddy/`, `.cursor/` or any other agent directory: they ship inside the installed CLI and are served on demand by `teamai skill get wiki` (`--full` for the references too). What an agent reads therefore always matches the CLI it is running. `teamai skill path wiki` prints the directory that holds the scripts and templates, for the commands below that run them. TeamAI ships no separate team-wiki CLI, and no extra plugin is required.
|
|
4
|
+
|
|
5
|
+
## Why this skill exists
|
|
6
|
+
|
|
7
|
+
The AI comprehension problem of large projects:
|
|
8
|
+
|
|
9
|
+
| Pain point | Symptom |
|
|
10
|
+
|------|---------|
|
|
11
|
+
| **Context does not fit** | 10+ repositories and hundreds of thousands of lines of code, far beyond the AI context window |
|
|
12
|
+
| **Relations are unclear** | RPC/MQ/DB dependencies between microservices are scattered across repositories with no global view |
|
|
13
|
+
| **Rules are not remembered** | Business constraints, state machines and config parameters hide deep in call chains |
|
|
14
|
+
| **Answers are inaccurate** | AI sees only local code, lacks global architecture awareness, and hallucinates easily |
|
|
15
|
+
| **High token consumption** | Every question re-reads large amounts of source, which is very inefficient |
|
|
16
|
+
|
|
17
|
+
## How it is solved
|
|
18
|
+
|
|
19
|
+
Architecture reverse-engineering **compresses the huge codebase into a structured knowledge base**:
|
|
20
|
+
|
|
21
|
+
- Every conclusion has a code `file:line` as evidence
|
|
22
|
+
- Every component relation carries a confidence label (`EXTRACTED` / `INFERRED` / `AMBIGUOUS`)
|
|
23
|
+
- Every generation run produces accuracy statistics, with automatic warnings when thresholds are exceeded
|
|
24
|
+
- AI reads the knowledge base instead of the source and gains global architecture awareness for **about 1/50 of the tokens**
|
|
25
|
+
- In Phase 0, `teamai codebase --extract` can generate evidence-backed structural edges (TS/JS/Python/Go AST + multi-language heuristics)
|
|
26
|
+
- After extraction, `teamai codebase --deep-enrich --project <slug> --output <repo>` can generate deterministic graph documents (G1/G2/G3) and deep knowledge; no separate team-wiki CLI is needed
|
|
27
|
+
|
|
28
|
+
---
|
|
29
|
+
|
|
30
|
+
## Deliverables
|
|
31
|
+
|
|
32
|
+
```
|
|
33
|
+
<output_dir>/
|
|
34
|
+
├── README.md ← Retrieval routing guide (for AI)
|
|
35
|
+
├── {project_name} Technical Architecture.md ← Whole-system view, ~200KB
|
|
36
|
+
├── {project_name} Business Architecture.md ← Product capabilities + lifecycle
|
|
37
|
+
├── {project_name} Deployment Architecture.md ← Deployment topology
|
|
38
|
+
├── XX_{component}_Design.md × N ← One per component, with the AI Quick Reference table
|
|
39
|
+
├── XX_{project_name}_Core_API_Product_Code_Mapping.md ← Product constraint → code location bridge document
|
|
40
|
+
├── XX_{project_name}_Product_Rules_Cheat_Sheet.md
|
|
41
|
+
├── XX_{project_name}_Business_Development_SOP.md
|
|
42
|
+
├── {anti-patterns / RPC contracts / troubleshooting notes} × N
|
|
43
|
+
├── _manifest.json ← Machine-readable manifest (for later graph merging)
|
|
44
|
+
└── graph/ ← Graph RAG graph document set
|
|
45
|
+
├── G1 Component dependency matrix
|
|
46
|
+
├── G2 Call chain overview + state machines
|
|
47
|
+
├── G3 Data flow and storage dependencies
|
|
48
|
+
├── G4 Error code component map
|
|
49
|
+
├── G5 Cross-component interaction scenarios (≥10 sequence diagrams)
|
|
50
|
+
├── G6 Knowledge graph triples (≥100 entries, with confidence)
|
|
51
|
+
├── G7 Architecture risks and impact analysis
|
|
52
|
+
├── G8 Core config parameter index
|
|
53
|
+
└── G9 Business rule constraint matrix + AI reasoning decision tree
|
|
54
|
+
```
|
|
55
|
+
|
|
56
|
+
---
|
|
57
|
+
|
|
58
|
+
## Execution flow
|
|
59
|
+
|
|
60
|
+
```
|
|
61
|
+
Phase 0 → Initialisation: collect paths, project name, product doc sources; optional CLI ast+heuristic structural baseline
|
|
62
|
+
|
|
63
|
+
Phase K1 → Architecture reverse-engineering: key file extraction → layered analysis → component relation matrix
|
|
64
|
+
⛔ Confirmation point ① Architecture understanding
|
|
65
|
+
|
|
66
|
+
Phase K2 → Document generation (parallel batches):
|
|
67
|
+
Batches 1~4: Type-4 component documents (dispatched to parallel sub-agents)
|
|
68
|
+
⛔ Confirmation point ② Document quality spot check
|
|
69
|
+
Batches 5~7: architecture overview + bridge documents + knowledge enhancement
|
|
70
|
+
|
|
71
|
+
Phase K3 → AI-Native enhancement:
|
|
72
|
+
search-anchor + bidirectional links + retrieval routing rules
|
|
73
|
+
Graph RAG graph document set G1~G9 (three-state confidence labels)
|
|
74
|
+
|
|
75
|
+
Phase K4 → Quality assessment:
|
|
76
|
+
validate_kb.py automatic checks
|
|
77
|
+
Whole-base accuracy audit ([UNVERIFIED] statistics + interface coverage)
|
|
78
|
+
Cross-document consistency check (contradiction detection + automatic fixes)
|
|
79
|
+
RAG retrieval spot check (7 question types)
|
|
80
|
+
AI end-to-end validation (10~15 standard questions + code trace-back)
|
|
81
|
+
Quality report generation
|
|
82
|
+
```
|
|
83
|
+
|
|
84
|
+
Supports `--update` incremental updates (based on a file hash cache, rerunning only changed components).
|
|
85
|
+
|
|
86
|
+
---
|
|
87
|
+
|
|
88
|
+
## File structure
|
|
89
|
+
|
|
90
|
+
The files below ship with the CLI; `teamai skill path wiki` prints the directory that contains them (`{SKILL_DIR}` in this document).
|
|
91
|
+
|
|
92
|
+
```
|
|
93
|
+
{SKILL_DIR}/
|
|
94
|
+
├── SKILL.md ← Main execution instructions (`teamai skill get wiki`)
|
|
95
|
+
├── scripts/
|
|
96
|
+
│ ├── scan_repo.py ← Repository scan helper
|
|
97
|
+
│ └── validate_kb.py ← Knowledge base quality validation tool
|
|
98
|
+
├── references/
|
|
99
|
+
│ ├── overview.md ← This file
|
|
100
|
+
│ ├── agents/
|
|
101
|
+
│ │ ├── kb-doc-generator.md ← Dedicated Agent for Type-1~8 document generation
|
|
102
|
+
│ │ └── graph-rag-agent.md ← Dedicated Agent for G1~G9 graph documents
|
|
103
|
+
│ ├── methodology/
|
|
104
|
+
│ │ ├── phase0-collection.md ← Source material collection method
|
|
105
|
+
│ │ ├── phase1-reverse-engineering.md ← Architecture reverse-engineering method
|
|
106
|
+
│ │ ├── phase2-document-types.md ← Specification and quality standards of the nine document types
|
|
107
|
+
│ │ ├── phase3-ai-enhancement.md ← AI-Native enhancement method
|
|
108
|
+
│ │ └── phase4-quality.md ← Quality assessment checklist
|
|
109
|
+
│ ├── phases/ ← Execution steps of each Phase
|
|
110
|
+
│ └── templates/
|
|
111
|
+
│ └── project-overview.md ← Knowledge base README template (with cognitive boundary declaration)
|
|
112
|
+
```
|
|
113
|
+
|
|
114
|
+
---
|
|
115
|
+
|
|
116
|
+
## Quality standards
|
|
117
|
+
|
|
118
|
+
| Dimension | Passing standard |
|
|
119
|
+
|------|---------|
|
|
120
|
+
| Coverage | ≥90% of P0 core components have documents |
|
|
121
|
+
| Accuracy | [UNVERIFIED] < 15% |
|
|
122
|
+
| Structural quality | Dead links = 0, search-anchor coverage ≥95% |
|
|
123
|
+
| AI usability | RAG retrieval spot check accuracy ≥85% |
|
|
124
|
+
| Relation trustworthiness | AMBIGUOUS relations < 10%, all listed for confirmation |
|
|
@@ -0,0 +1,118 @@
|
|
|
1
|
+
## Phase K1: Architecture Reverse-Engineering and Source Material Collection
|
|
2
|
+
|
|
3
|
+
**Methodology**: `{SKILL_DIR}/references/methodology/phase0-collection.md` + `{SKILL_DIR}/references/methodology/phase1-reverse-engineering.md`
|
|
4
|
+
|
|
5
|
+
### Step 1: Optionally run the scan script (recommended)
|
|
6
|
+
|
|
7
|
+
```bash
|
|
8
|
+
python3 "{SKILL_DIR}/scripts/scan_repo.py" <project_root> --depth 2 --top 10
|
|
9
|
+
```
|
|
10
|
+
Output: file statistics + key file discovery report + language distribution.
|
|
11
|
+
|
|
12
|
+
### Step 2: Key file extraction
|
|
13
|
+
|
|
14
|
+
Scan by priority (see phase0-collection.md for details):
|
|
15
|
+
- **P0 required**: entry files, routes/handlers, workflow orchestration config, Proto/IDL
|
|
16
|
+
- **P1 important**: database schema (DDL), constant / error code definitions
|
|
17
|
+
- **P2 enhancement**: config files, test files (to understand expected behaviour)
|
|
18
|
+
|
|
19
|
+
### Step 3: Architecture reverse-engineering (see phase1-reverse-engineering.md for details)
|
|
20
|
+
|
|
21
|
+
- Bottom-up layering: leaf nodes (DB/MQ) → intermediate nodes (orchestration/scheduling) → root nodes (API entry points)
|
|
22
|
+
- Three-layer penetration tracing: for ≥5 core APIs, complete the full call chain trace API entry → orchestration layer → service execution layer
|
|
23
|
+
- Build the N×N component relationship matrix (annotate the communication method: RPC/MQ/DB)
|
|
24
|
+
|
|
25
|
+
### Step 4: Generate the architecture analysis report
|
|
26
|
+
|
|
27
|
+
Write to `_review/k1-architecture-map.md`:
|
|
28
|
+
|
|
29
|
+
```markdown
|
|
30
|
+
## Architecture Layers (≥4 layers)
|
|
31
|
+
| Layer | Components | Core Responsibility | Code Repository |
|
|
32
|
+
|
|
33
|
+
## Component Inventory
|
|
34
|
+
| Component | Architecture Layer | **Repository** | Language | Criticality (P0/P1/P2) | Entry File | **Interface Check Type** |
|
|
35
|
+
|
|
36
|
+
Interface check type values (ask the user to verify this column at confirmation point ①):
|
|
37
|
+
- `HTTP` → API access layer, has HTTP/gRPC route registrations, requires interface count reconciliation
|
|
38
|
+
- `MQ` → message processing layer, has MQ Consumer/Exchange declarations, Topic count is the baseline
|
|
39
|
+
- `RPC` → internal service layer, has .proto / .thrift / IDL files, Method count is the baseline
|
|
40
|
+
- `NONE` → scheduling / execution / data layer, no external interface, no interface count check
|
|
41
|
+
|
|
42
|
+
## N×N Component Communication Matrix
|
|
43
|
+
(values: RPC/MQ/DB/—, annotated with confidence [E]EXTRACTED/[I]INFERRED/[A]AMBIGUOUS)
|
|
44
|
+
|
|
45
|
+
## Core Call Chains (≥5)
|
|
46
|
+
(format: API(file:line) → orchestration layer(config:line) → service layer(handler:line) → DB(table))
|
|
47
|
+
|
|
48
|
+
## Glossary
|
|
49
|
+
| Internal Term | External / Product Term | Notes |
|
|
50
|
+
|
|
51
|
+
## Uncertain Items (for manual confirmation)
|
|
52
|
+
(relationships and inferences marked [A], with the reason for the uncertainty)
|
|
53
|
+
(components whose interface check type is uncertain, marked [?], to be clarified by the user at confirmation point ①)
|
|
54
|
+
```
|
|
55
|
+
|
|
56
|
+
### Step 5: Interface inventory scan (run separately per check type)
|
|
57
|
+
|
|
58
|
+
**Run only for components whose interface check type in k1-architecture-map.md is ≠ NONE**:
|
|
59
|
+
|
|
60
|
+
```
|
|
61
|
+
FOR each component with interface check type = HTTP:
|
|
62
|
+
Run a grep scan:
|
|
63
|
+
Go: grep -rn "\.GET\|\.POST\|\.PUT\|\.DELETE\|router\.Handle\|@handler" <component_dir>
|
|
64
|
+
Python: grep -rn "@app\.route\|@router\.\|APIRouter\|include_router" <component_dir>
|
|
65
|
+
Record: component → HTTP interface count N (SCAN_CONFIDENCE: HIGH/MEDIUM)
|
|
66
|
+
|
|
67
|
+
FOR each component with interface check type = MQ:
|
|
68
|
+
Run a grep scan:
|
|
69
|
+
grep -rn "Exchange\|Queue\|Topic\|consumer\|subscribe\|@KafkaListener" <component_dir>
|
|
70
|
+
Record: component → MQ Topic/Queue count N
|
|
71
|
+
|
|
72
|
+
FOR each component with interface check type = RPC:
|
|
73
|
+
Parse the .proto / .thrift files:
|
|
74
|
+
find <component_dir> -name "*.proto" -o -name "*.thrift" | xargs grep "^rpc\|^service"
|
|
75
|
+
Record: component → RPC Method count N
|
|
76
|
+
```
|
|
77
|
+
|
|
78
|
+
Write the results to `_review/interface-inventory.json`:
|
|
79
|
+
```json
|
|
80
|
+
{
|
|
81
|
+
"ComponentA": {"type": "HTTP", "count": 13, "confidence": "HIGH"},
|
|
82
|
+
"ComponentB": {"type": "MQ", "count": 5, "confidence": "MEDIUM"},
|
|
83
|
+
"ComponentC": {"type": "RPC", "count": 8, "confidence": "HIGH"},
|
|
84
|
+
"ComponentD": {"type": "NONE", "count": 0, "confidence": "—"}
|
|
85
|
+
}
|
|
86
|
+
```
|
|
87
|
+
|
|
88
|
+
**When done**: update `current_phase` to `"phasek1_waiting_confirm"`.
|
|
89
|
+
|
|
90
|
+
**⛔ Confirmation point ①**: wait for an explicit reply from the user. Do not proceed to the next phase automatically.
|
|
91
|
+
|
|
92
|
+
Show the user:
|
|
93
|
+
```
|
|
94
|
+
Architecture analysis complete.
|
|
95
|
+
|
|
96
|
+
Component inventory (N in total):
|
|
97
|
+
P0 core: [list]
|
|
98
|
+
P1 important: [list]
|
|
99
|
+
P2 auxiliary: [list]
|
|
100
|
+
|
|
101
|
+
Interface scan results (for verification):
|
|
102
|
+
HTTP interfaces: ComponentA 13, ComponentB 7
|
|
103
|
+
MQ Topics: ComponentC 5
|
|
104
|
+
RPC Methods: ComponentD 8
|
|
105
|
+
Components without interfaces: ComponentE, ComponentF, ...
|
|
106
|
+
|
|
107
|
+
AMBIGUOUS relationships (please clarify):
|
|
108
|
+
- The communication method of ComponentX → ComponentY is uncertain
|
|
109
|
+
|
|
110
|
+
Please confirm (edit k1-architecture-map.md directly, then reply "continue"):
|
|
111
|
+
1. Are the architecture layers and P0/P1/P2 annotations correct?
|
|
112
|
+
2. Is the interface check type (HTTP/MQ/RPC/NONE) of every component accurate?
|
|
113
|
+
3. Are the interface scan counts reasonable? Clearly too few means something was missed; too many may mean test files were scanned.
|
|
114
|
+
```
|
|
115
|
+
|
|
116
|
+
After confirmation: update to `"phasek1_confirmed"` → Phase K2.
|
|
117
|
+
|
|
118
|
+
---
|
|
@@ -0,0 +1,68 @@
|
|
|
1
|
+
## Phase K2: Document Generation (batched parallel runs + mid-way quality confirmation)
|
|
2
|
+
|
|
3
|
+
**Methodology**: `{SKILL_DIR}/references/methodology/phase2-document-types.md`
|
|
4
|
+
|
|
5
|
+
### Generation order (dependency-chain driven, lower layers first)
|
|
6
|
+
|
|
7
|
+
```
|
|
8
|
+
Batch 1: data layer + basic execution layer Type-4 component documents ← parallel
|
|
9
|
+
Batch 2: resource / scheduling layer Type-4 component documents ← parallel
|
|
10
|
+
Batch 3: messaging / service layer Type-4 component documents ← parallel
|
|
11
|
+
Batch 4: API entry layer Type-4 component documents ← parallel
|
|
12
|
+
⛔ Confirmation point ② ← manual spot check of component document quality
|
|
13
|
+
Batch 5: architecture overview layer (Type-1 + Type-2 + Type-3) ← serial (depends on all layers above being complete)
|
|
14
|
+
Batch 6: bridge documents (Type-5 + Type-6 + Type-7) ← serial (depends on product documentation)
|
|
15
|
+
Batch 7: knowledge enhancement (Type-8: anti-patterns / RPC contracts / troubleshooting) ← serial
|
|
16
|
+
```
|
|
17
|
+
|
|
18
|
+
### Execution flow for each batch
|
|
19
|
+
|
|
20
|
+
Read `{SKILL_DIR}/references/agents/kb-doc-generator.md`, assemble the input package and launch:
|
|
21
|
+
|
|
22
|
+
```
|
|
23
|
+
component_list: list of components / document types for this batch
|
|
24
|
+
architecture_map: full content of _review/k1-architecture-map.md
|
|
25
|
+
repos: repository list from _review/repo-manifest.json
|
|
26
|
+
service_map: service_map from progress.json
|
|
27
|
+
output_dir: <Phase 0>
|
|
28
|
+
project_name: <Phase 0>
|
|
29
|
+
product_docs_dir: <Phase 0, may be empty>
|
|
30
|
+
methodology_dir: {SKILL_DIR}/references/methodology/
|
|
31
|
+
completed_docs: kb_progress.components_done (skipped on resume from checkpoint)
|
|
32
|
+
parallel_mode: true (batches 1~4) / false (batches 5~7)
|
|
33
|
+
```
|
|
34
|
+
|
|
35
|
+
After each batch completes:
|
|
36
|
+
- Append the completed components to `kb_progress.components_done`
|
|
37
|
+
- Accumulate `accuracy_stats` (extracted from the self-check summary returned by the Agent)
|
|
38
|
+
- Update `current_phase` to `"phasek2_batch_N"`
|
|
39
|
+
- Show the token consumption and `[UNVERIFIED]` statistics for this batch
|
|
40
|
+
|
|
41
|
+
### ⛔ Confirmation point ② (after batches 1~4 complete)
|
|
42
|
+
|
|
43
|
+
Show the user:
|
|
44
|
+
```
|
|
45
|
+
{N} component design documents generated. Accuracy statistics:
|
|
46
|
+
Total claims: {N} | Verified: {N} | [UNVERIFIED]: {N} ({X}%)
|
|
47
|
+
AMBIGUOUS relationships: {N}
|
|
48
|
+
|
|
49
|
+
Please spot-check 2~3 documents (the most complex components are recommended):
|
|
50
|
+
Path: <output_dir>/XX_<component>_Design.md
|
|
51
|
+
|
|
52
|
+
Points to confirm:
|
|
53
|
+
1. Is the code entry point in the AI Quick Reference table precise down to the function name?
|
|
54
|
+
2. Does the core flow description match the actual code?
|
|
55
|
+
3. Is the [UNVERIFIED] ratio acceptable? (<15% recommended)
|
|
56
|
+
|
|
57
|
+
If you find a systematic problem, describe it and I will adjust the strategy and regenerate.
|
|
58
|
+
```
|
|
59
|
+
|
|
60
|
+
Update `current_phase` to `"phasek2_waiting_confirm"`.
|
|
61
|
+
After the user confirms, update to `"phasek2_confirmed"` and continue with batches 5~7.
|
|
62
|
+
|
|
63
|
+
### After all batches complete
|
|
64
|
+
|
|
65
|
+
Write `_review/k2-doc-list.md` (document list: path + size in KB + [UNVERIFIED] count + generation time).
|
|
66
|
+
Update `current_phase` to `"phasek2_done"` → Phase K3.
|
|
67
|
+
|
|
68
|
+
---
|
|
@@ -0,0 +1,121 @@
|
|
|
1
|
+
## Phase K3: AI-Native Enhancement + Graph Document Set
|
|
2
|
+
|
|
3
|
+
**Methodology**: `{SKILL_DIR}/references/methodology/phase3-ai-enhancement.md`
|
|
4
|
+
|
|
5
|
+
### Step 1: Inject AI-Native elements
|
|
6
|
+
|
|
7
|
+
Add to all generated documents (where the Phase K2 Agent did not add them completely):
|
|
8
|
+
|
|
9
|
+
| Element | Requirement | Scope |
|
|
10
|
+
|------|------|---------|
|
|
11
|
+
| `search-anchor` | 5~15 keywords, first line after the title | All documents |
|
|
12
|
+
| AI Quick Reference table | 10 dimensions, immediately after the title | All Type-4 component documents |
|
|
13
|
+
| Bidirectional links | component ↔ main architecture, bridge ↔ component | All documents |
|
|
14
|
+
| Retrieval routing rules | 4 routing rules + 4 priority levels | Technical architecture overview only |
|
|
15
|
+
| QA pairs | 10~20 high-frequency questions + answer references | Chapter 9 of the technical architecture overview only |
|
|
16
|
+
|
|
17
|
+
### Step 2: Graph RAG graph document set
|
|
18
|
+
|
|
19
|
+
Read `{SKILL_DIR}/references/agents/graph-rag-agent.md`, assemble the input package and launch:
|
|
20
|
+
|
|
21
|
+
```
|
|
22
|
+
all_kb_docs_dir: <output_dir>
|
|
23
|
+
architecture_map: _review/k1-architecture-map.md
|
|
24
|
+
doc_list: _review/k2-doc-list.md
|
|
25
|
+
project_name: <Phase 0>
|
|
26
|
+
output_dir: <output_dir>/graph/
|
|
27
|
+
methodology_file: {SKILL_DIR}/references/methodology/phase2-document-types.md
|
|
28
|
+
```
|
|
29
|
+
|
|
30
|
+
Generate G1~G9 (every relationship carries a mandatory three-state confidence annotation):
|
|
31
|
+
|
|
32
|
+
| Graph Document | Question Solved | Confidence Requirement |
|
|
33
|
+
|---------|---------|-----------|
|
|
34
|
+
| G1 Component Dependency Matrix | "Who depends on X?" | EXTRACTED from explicit document descriptions |
|
|
35
|
+
| G2 Call Chain Overview + state machine + constraint matrix | "Which modules does an API pass through?" | call chains EXTRACTED, inferred dependencies INFERRED |
|
|
36
|
+
| G3 Data Flow and Storage Dependencies | "Where is the data stored?" | read/write relationships EXTRACTED |
|
|
37
|
+
| G4 Error Code Component Map | "Which module does this error code belong to?" | EXTRACTED |
|
|
38
|
+
| G5 Cross-Component Interaction Scenarios (≥10 sequence diagrams) | "How is the quota check done?" | sequences EXTRACTED, boundaries INFERRED |
|
|
39
|
+
| G6 Knowledge Graph Triples (≥100) | "Who does A depend on indirectly?" | every triple marked E/I/A + score |
|
|
40
|
+
| G7 Architecture Risks and Impact Analysis | "How big is the impact if X goes down?" | direct dependencies EXTRACTED, indirect INFERRED |
|
|
41
|
+
| G8 Core Config Parameter Index | "How do I change configuration XX?" | EXTRACTED from config files |
|
|
42
|
+
| G9 Business Rule Constraint Matrix + AI reasoning decision tree | "Can I do XX?" | rules EXTRACTED, inferences INFERRED |
|
|
43
|
+
|
|
44
|
+
Also generate `<output_dir>/graph/README.md` (index + lookup-by-question-type table + retrieval routing suggestions).
|
|
45
|
+
|
|
46
|
+
### Step 3: Cross-document consistency check
|
|
47
|
+
|
|
48
|
+
**After the Graph RAG Agent finishes, the main agent performs this step itself (do not delegate to a sub-agent).**
|
|
49
|
+
|
|
50
|
+
Purpose: detect contradictory descriptions between component documents, preventing inconsistencies such as "A says it calls B over RPC, B says it is called by A over MQ".
|
|
51
|
+
|
|
52
|
+
```
|
|
53
|
+
Step 3A: Build the "claim matrix"
|
|
54
|
+
|
|
55
|
+
For every Type-4 component document, extract relationship claims from **two levels**:
|
|
56
|
+
|
|
57
|
+
Level 1: the "Upstream Components" and "Downstream Components" fields of the AI Quick Reference table
|
|
58
|
+
Level 2: call descriptions in the interface design and core flow sections of the body
|
|
59
|
+
|
|
60
|
+
If level 1 and level 2 describe the same relationship differently → first record it as an "intra-document contradiction" (a higher-priority problem than header vs body)
|
|
61
|
+
|
|
62
|
+
Extraction example:
|
|
63
|
+
ComponentX.md header claims: X→Y(RPC), X→Z(MQ)
|
|
64
|
+
ComponentX.md body claims: X→Z(HTTP) ← contradicts the header!
|
|
65
|
+
ComponentY.md header claims: Y←X(RPC), Y→Z(DB)
|
|
66
|
+
ComponentZ.md header claims: Z←X(HTTP), Z←Y(DB)
|
|
67
|
+
|
|
68
|
+
Step 3B: Cross-compare
|
|
69
|
+
|
|
70
|
+
FOR each pair of components (A, B):
|
|
71
|
+
IF A.md claims "A→B over RPC" AND B.md claims "B←A over MQ":
|
|
72
|
+
→ record contradiction: "A→B communication method inconsistent: A says RPC, B says MQ"
|
|
73
|
+
IF A.md claims "A→B" BUT B.md does not mention "called by A":
|
|
74
|
+
→ record omission: "A claims to call B, but B's document does not mention being called by A"
|
|
75
|
+
IF a relationship in the G1 matrix differs from the component document claims:
|
|
76
|
+
→ record deviation: "G1 matrix says A→B(RPC), but A's document says A→B(MQ)"
|
|
77
|
+
|
|
78
|
+
Step 3C: Generate the consistency report
|
|
79
|
+
|
|
80
|
+
Write to `_review/k3-consistency-check.md`:
|
|
81
|
+
|
|
82
|
+
```markdown
|
|
83
|
+
# Cross-Document Consistency Check Report
|
|
84
|
+
|
|
85
|
+
## Contradictions (must fix)
|
|
86
|
+
| Component A | Component B | A's Description | B's Description | Contradiction Type |
|
|
87
|
+
|-------|-------|---------|---------|---------|
|
|
88
|
+
| X | Z | X→Z(MQ) | Z←X(HTTP) | Communication method inconsistent |
|
|
89
|
+
|
|
90
|
+
## Omissions (recommended additions)
|
|
91
|
+
| Claimant | Referenced | Claim | Omission |
|
|
92
|
+
|--------|---------|---------|------|
|
|
93
|
+
| A | B | A→B(RPC) | B's document does not mention being called by A |
|
|
94
|
+
|
|
95
|
+
## G1 Matrix Deviations (recommended alignment)
|
|
96
|
+
| G1 Matrix | Component Document | Deviation |
|
|
97
|
+
|
|
98
|
+
## Statistics
|
|
99
|
+
- Contradictions: N (❌ must fix)
|
|
100
|
+
- Omissions: N (⚠️ recommended additions)
|
|
101
|
+
- G1 deviations: N (⚠️ need alignment)
|
|
102
|
+
- Consistent relationships: N (✅)
|
|
103
|
+
- Consistency rate: X%
|
|
104
|
+
```
|
|
105
|
+
|
|
106
|
+
Step 3D: Automatic fixes (unambiguous cases only)
|
|
107
|
+
|
|
108
|
+
IF contradictions > 0:
|
|
109
|
+
FOR each contradiction:
|
|
110
|
+
Trace back to the code: use Grep to find the actual call method (e.g. rpc.Call / mq.Publish)
|
|
111
|
+
IF the correct side can be determined → fix the description in the wrong side's document + update the G1 matrix
|
|
112
|
+
IF it cannot be determined → mark as AMBIGUOUS, leave for the user to confirm at the confirmation point
|
|
113
|
+
Recompute the consistency rate after fixing
|
|
114
|
+
|
|
115
|
+
IF contradictions = 0:
|
|
116
|
+
→ skip fixing, go straight to Phase K4
|
|
117
|
+
```
|
|
118
|
+
|
|
119
|
+
**When done**: update `current_phase` to `"phasek3_done"` → Phase K4.
|
|
120
|
+
|
|
121
|
+
---
|