teamai-cli 0.25.0 → 0.26.0-beta.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (45) hide show
  1. package/CHANGELOG.md +13 -0
  2. package/README.zh-CN.md +6 -0
  3. package/dist/index.js +6147 -3346
  4. package/package.json +4 -1
  5. package/skill-data/core/SKILL.md +114 -0
  6. package/skill-data/core/references/commands.md +339 -0
  7. package/{skills/teamai → skill-data/core}/references/contribute-member.md +13 -10
  8. package/{skills/teamai → skill-data/core}/references/troubleshooting.md +9 -1
  9. package/skill-data/setup/SKILL.md +76 -0
  10. package/{skills/teamai → skill-data/setup}/references/join-member.md +17 -14
  11. package/{skills/teamai → skill-data/setup}/references/manage-admin.md +18 -6
  12. package/{skills/teamai → skill-data/setup}/references/provider-tgit.md +9 -6
  13. package/{skills/teamai → skill-data/setup}/references/setup-admin.md +41 -35
  14. package/skill-data/share/SKILL.md +70 -0
  15. package/skill-data/share/references/doc-template.md +44 -0
  16. package/skill-data/wiki/SKILL.md +314 -0
  17. package/skill-data/wiki/references/agents/graph-rag-agent.md +344 -0
  18. package/skill-data/wiki/references/agents/kb-doc-generator.md +323 -0
  19. package/skill-data/wiki/references/methodology/phase0-collection.md +54 -0
  20. package/skill-data/wiki/references/methodology/phase1-reverse-engineering.md +89 -0
  21. package/skill-data/wiki/references/methodology/phase2-document-types.md +341 -0
  22. package/skill-data/wiki/references/methodology/phase3-ai-enhancement.md +164 -0
  23. package/skill-data/wiki/references/methodology/phase4-quality.md +232 -0
  24. package/skill-data/wiki/references/overview.md +124 -0
  25. package/skill-data/wiki/references/phases/k1-reverse-engineering.md +118 -0
  26. package/skill-data/wiki/references/phases/k2-documents.md +68 -0
  27. package/skill-data/wiki/references/phases/k3-ai-native.md +121 -0
  28. package/skill-data/wiki/references/phases/k4-quality.md +190 -0
  29. package/skill-data/wiki/references/phases/phase0-init.md +112 -0
  30. package/skill-data/wiki/references/templates/project-overview.md +148 -0
  31. package/{skills/team-wiki-codebase → skill-data/wiki}/scripts/scan_repo.py +52 -52
  32. package/{skills/team-wiki-codebase → skill-data/wiki}/scripts/validate_kb.py +68 -62
  33. package/skills/teamai/SKILL.md +28 -128
  34. package/skills/team-wiki-codebase/README.md +0 -121
  35. package/skills/team-wiki-codebase/SKILL.md +0 -905
  36. package/skills/team-wiki-codebase/references/agents/graph-rag-agent.md +0 -344
  37. package/skills/team-wiki-codebase/references/agents/kb-doc-generator.md +0 -323
  38. package/skills/team-wiki-codebase/references/methodology/phase0-collection.md +0 -54
  39. package/skills/team-wiki-codebase/references/methodology/phase1-reverse-engineering.md +0 -89
  40. package/skills/team-wiki-codebase/references/methodology/phase2-document-types.md +0 -341
  41. package/skills/team-wiki-codebase/references/methodology/phase3-ai-enhancement.md +0 -164
  42. package/skills/team-wiki-codebase/references/methodology/phase4-quality.md +0 -232
  43. package/skills/team-wiki-codebase/references/templates/project-overview.md +0 -148
  44. package/skills/teamai-share-learnings/SKILL.md +0 -87
  45. /package/{skills/teamai → skill-data/setup}/references/uninstall.md +0 -0
@@ -0,0 +1,314 @@
1
+ ---
2
+ name: wiki
3
+ description: >-
4
+ Make AI truly understand large codebases: for multi-repository, multi-microservice projects
5
+ that have evolved over years, run architecture reverse-engineering + a Graph RAG graph +
6
+ multi-language AST to compress a huge codebase into a structured knowledge base, where every
7
+ conclusion traces back to a code line and every relation carries a confidence label. Suited to
8
+ projects with 10+ repositories or microservices that AI cannot understand globally by reading
9
+ the code directly. Triggers: architecture analysis, architecture reverse-engineering,
10
+ codebase knowledge base, code-to-knowledge, architecture wiki, large multi-repo codebase.
11
+ Loaded on demand by the teamai discovery stub.
12
+ ---
13
+
14
+ # wiki: AI cognition engineering for large codebases
15
+
16
+ > Prerequisites: an accessible source directory (multiple repositories supported), Python 3, and an installed teamai CLI.
17
+ > The methodology, sub-agent prompts, templates and scripts ship with the CLI. Run `teamai skill path wiki` to get their absolute path;
18
+ > `{SKILL_DIR}` in this document refers to that path; a reference file you open on its own writes that directory as `SKILL_DIR` in braces.
19
+ > Write the knowledge-base documents in Simplified Chinese, as earlier releases did. When updating an existing knowledge base, keep its file names and headings; `validate_kb.py` accepts both the current English and the earlier Chinese headings.
20
+ > The Phase 0 structural baseline uses `teamai codebase --extract`. TeamAI does not ship a separate team-wiki CLI. No extra plugin is required.
21
+
22
+ **The problem**: large projects (10+ repositories, dozens of microservices, years of iteration) defeat global understanding by AI. The context window cannot hold all the code, component relations are scattered everywhere, and business rules hide deep in call chains. Letting AI read the code directly is both slow (huge token counts) and inaccurate (no global view).
23
+
24
+ **The solution**: use architecture reverse-engineering to systematically compress a huge codebase into a **structured, verifiable, AI-Native** deep knowledge base. Every conclusion traces back to a code line, every relation carries a confidence label, and every update is incrementally verified. AI reads the knowledge base instead of the source, and gains global architecture awareness for about **1/50 of the tokens**.
25
+
26
+ ## Usage
27
+
28
+ The user states the mode in natural language, or simply says "build a codebase knowledge base":
29
+
30
+ ```
31
+ default Standard: single-session core path
32
+ --deep Full K1~K4 + G1~G9
33
+ --update Incremental update of an existing knowledge/
34
+ continue Resume from the _review/progress.json checkpoint
35
+ ```
36
+
37
+ ---
38
+
39
+ ## Agent architecture
40
+
41
+ | Agent | File | When started |
42
+ |-------|------|---------|
43
+ | Knowledge base document generator Agent | `{SKILL_DIR}/references/agents/kb-doc-generator.md` | Phase K2, every component batch |
44
+ | Graph RAG Agent | `{SKILL_DIR}/references/agents/graph-rag-agent.md` | Phase K3 |
45
+
46
+ **Main agent responsibilities**: workflow orchestration, confirmation point management, progress.json maintenance, quality report aggregation.
47
+
48
+ ---
49
+
50
+ ## Entry decision
51
+
52
+ **This decision must run first on every activation.**
53
+
54
+ ```
55
+ IF the user input contains "--update" or "incremental update":
56
+ → Update mode
57
+ ELSE IF the user input contains "continue" or "resume":
58
+ → Continue mode
59
+ ELSE:
60
+ → Check whether _review/progress.json exists under the user-specified directory
61
+ IF it exists → report the state, wait for "resume last run" or "start over"
62
+ ELSE → Phase 0
63
+ ```
64
+
65
+ ---
66
+
67
+ ## Continue mode
68
+
69
+ ```
70
+ Step 1: Locate progress.json
71
+ Step 2: Read and parse it, show a resume summary
72
+ Step 3: Jump according to current_phase:
73
+ "phase0_done" → Phase K1
74
+ "phasek1_waiting_confirm" → Show k1-architecture-map.md, wait for confirmation ①
75
+ "phasek1_confirmed" → Phase K2
76
+ "phasek2_batch_N" → Continue Phase K2 from batch N (skip completed ones)
77
+ "phasek2_waiting_confirm" → Wait for confirmation ②
78
+ "phasek2_confirmed" → Phase K3
79
+ "phasek3_done" → Phase K4
80
+ "phasek4_done"/"completed" → Report completion, ask whether to --update or rerun a component
81
+ ```
82
+
83
+ ---
84
+
85
+ ## Update mode (incremental update)
86
+
87
+ **Trigger**: the user asks for an "incremental update", or specifies the `--update` mode in this skill.
88
+ **Precondition**: an existing progress.json in the completed state.
89
+
90
+ ```
91
+ Step 1: Read progress.json, get file_hash_cache
92
+ Step 2: Scan project_root, compute the current SHA256 of every file
93
+ Step 3: Compare hashes, classify: added / modified / deleted
94
+ Step 4: Show the change summary, wait for user confirmation:
95
+ ┌────────────────────────────────────┐
96
+ │ Change summary │
97
+ │ Added: N files │
98
+ │ Modified: N files (incl. Aurora.py)│
99
+ │ Deleted: N files │
100
+ │ Affected components: [list] │
101
+ │ Affected graph documents: G1/G2/G6/G7 │
102
+ └────────────────────────────────────┘
103
+ Step 5: Rerun only the affected scope:
104
+ - Phase K2: regenerate the Type-4 documents of affected components (overwrite)
105
+ - Phase K3 partial: update the graph documents that involve changed components (G1/G2/G6/G7)
106
+ - Phase K4: rerun validate_kb.py
107
+ Step 6: Update file_hash_cache + the metadata.json commit SHA
108
+ Step 7: Component-level diff (handle added/removed repositories or components)
109
+ IF the repos list differs from last time:
110
+ Added repositories → run a full K1 scan on the new repository, add it to the component inventory, generate Type-4 documents
111
+ Removed repositories → prepend `⚠️ [DEPRECATED] The repository for this component has been removed` to the component document
112
+ → Update the component inventory in k1-architecture-map.md
113
+ → Update the G1 matrix (remove rows/columns of removed components, add rows/columns for new ones)
114
+ ```
115
+
116
+ ---
117
+
118
+ ## progress.json specification
119
+
120
+ **Path**: `<output_dir>/../_review/progress.json`
121
+
122
+ ```json
123
+ {
124
+ "version": "5",
125
+ "repos": [
126
+ {"name": "repo-a", "path": "/absolute/path/to/repo-a", "language": "go"},
127
+ {"name": "repo-b", "path": "/absolute/path/to/repo-b", "language": "python"}
128
+ ],
129
+ "output_dir": "/absolute/path/to/knowledge",
130
+ "primary_language": "go",
131
+ "project_name": "ProjectName",
132
+ "scan_time": "2026-01-01T10:00:00Z",
133
+ "current_phase": "phasek2_batch_2",
134
+ "confirmed_phases": ["phase0", "phasek1"],
135
+
136
+ "service_map": {
137
+ "description": "Service name → repository map built in Phase K1 Step 3",
138
+ "ServiceA": {"repo": "repo-a", "entry": "cmd/serviceA/main.go"},
139
+ "ServiceB": {"repo": "repo-b", "entry": "app/main.py"}
140
+ },
141
+
142
+ "kb_progress": {
143
+ "component_total": 12,
144
+ "components_done": ["Aurora", "Frame"],
145
+ "components_pending": ["CCDB", "Dispatcher"],
146
+ "type1_done": false,
147
+ "type2_done": false,
148
+ "type3_done": false,
149
+ "bridge_docs_done": false,
150
+ "graph_rag_done": false
151
+ },
152
+
153
+ "accuracy_stats": {
154
+ "total_claims": 0,
155
+ "verified": 0,
156
+ "unverified": 0,
157
+ "ambiguous_relations": 0
158
+ },
159
+
160
+ "interface_coverage": {
161
+ "description": "Interface count reconciliation, filled by the Phase K2 self-check",
162
+ "ComponentA": {"type": "HTTP", "scanned": 13, "documented": 0, "gap": 13},
163
+ "ComponentB": {"type": "MQ", "scanned": 5, "documented": 0, "gap": 5}
164
+ },
165
+
166
+ "consistency_check": {
167
+ "description": "Cross-document consistency check result from Phase K3 Step 3",
168
+ "contradictions": 0,
169
+ "missing_refs": 0,
170
+ "g1_deviations": 0,
171
+ "consistency_rate": 0.0
172
+ },
173
+
174
+ "e2e_validation": {
175
+ "description": "AI end-to-end validation result from Phase K4 Step 4",
176
+ "total_questions": 0,
177
+ "correct": 0,
178
+ "partial": 0,
179
+ "incorrect": 0,
180
+ "boundary_ok": 0,
181
+ "boundary_fail": 0,
182
+ "accuracy_rate": 0.0
183
+ },
184
+
185
+ "file_hash_cache": {
186
+ "relative/path/to/file.go": "sha256_hex"
187
+ }
188
+ }
189
+ ```
190
+
191
+ > `accuracy_stats` accumulates after every Phase K2 batch and is the global trust indicator of the knowledge base.
192
+
193
+ ---
194
+
195
+ ## Core principles (accuracy first)
196
+
197
+ 1. **Code is the single source of truth**: every conclusion must cite a code file:line as evidence; anything unverifiable is marked `[UNVERIFIED]`
198
+ 2. **Three-state confidence is mandatory**: every relation in the graph is labelled `EXTRACTED(1.0)` / `INFERRED(0.6~0.9)` / `AMBIGUOUS(0.1~0.3)`; no invention out of thin air, no 0.5 default
199
+ 3. **Two-level accuracy verification**: Phase K2 self-checks every document right after generation; Phase K4 verifies the whole knowledge base
200
+ 4. **Two human-in-the-loop confirmations**: architecture understanding (K①) and component document quality (K②) must be confirmed by a human to stop systematic errors from spreading
201
+ 5. **Parallel generation + resume from checkpoint**: Type-4 component documents are dispatched in parallel (all Agent calls in the same message); progress.json is persisted after every batch
202
+ 6. **Token economy**: the `Glob → Grep → Read` three-step method; full directory scans are forbidden
203
+ 7. **Honest auditing**: `[UNVERIFIED]` must not be hidden; quality numbers are shown in full; when unsure, mark AMBIGUOUS instead of deleting
204
+ 8. **Cognitive boundary declaration**: the knowledge base README must state explicitly what is covered and what is not, so AI knows when to say "not sure"
205
+ 9. **Cross-document consistency**: Phase K3 must cross-check relation descriptions between components; contradictions count as "consistent" only after they are fixed
206
+ 10. **End-to-end verifiable**: Phase K4 tests the knowledge base's actual answering ability with standardised questions; E2E accuracy target ≥ 80%
207
+
208
+ ---
209
+
210
+ ## Phase workflow (loaded on demand)
211
+
212
+ The full steps of each phase live in separate files. Load a file when its phase comes up; do not read them all at once:
213
+
214
+ | Phase | File | Content |
215
+ |---|---|---|
216
+ | Phase 0 | `{SKILL_DIR}/references/phases/phase0-init.md` | Initialisation, `teamai codebase --extract` structural baseline, repository inventory |
217
+ | Phase K1 | `{SKILL_DIR}/references/phases/k1-reverse-engineering.md` | Architecture reverse-engineering and source material collection, scan script, architecture analysis report |
218
+ | Phase K2 | `{SKILL_DIR}/references/phases/k2-documents.md` | Document generation (parallel batches + intermediate quality confirmation) |
219
+ | Phase K3 | `{SKILL_DIR}/references/phases/k3-ai-native.md` | AI-Native enhancement + Graph RAG graph document set |
220
+ | Phase K4 | `{SKILL_DIR}/references/phases/k4-quality.md` | Quality assessment, validation script, quality report |
221
+
222
+ Methodology background (optional, for reference while writing documents): `{SKILL_DIR}/references/methodology/`;
223
+ sub-agent prompts: `{SKILL_DIR}/references/agents/`;
224
+ knowledge base README template: `{SKILL_DIR}/references/templates/project-overview.md`.
225
+
226
+ Human-readable overview (not for execution): `{SKILL_DIR}/references/overview.md`.
227
+
228
+ `teamai skill get wiki --full` prints every reference file in one go (about 130 KB). Use it only when you need to read everything.
229
+
230
+ ## Output directory layout
231
+
232
+ ```
233
+ <output_dir>/
234
+ ├── README.md ← Knowledge base index + retrieval routing rules + cognitive boundary declaration (for AI)
235
+ │ Start from the template: cp "{SKILL_DIR}/references/templates/project-overview.md" <output_dir>/README.md
236
+ ├── {project_name} Technical Architecture.md ← [Type-1] Architecture overview (target ≤80KB, split automatically when larger)
237
+ ├── {project_name} Technical Architecture-Core Call Chains.md ← [Type-1b] Split out only when Type-1 exceeds 80KB
238
+ ├── {project_name} Technical Architecture-AI Metadata.md ← [Type-1c] Split out only when Type-1 exceeds 80KB
239
+ ├── {project_name} Business Architecture.md ← [Type-2] Product capabilities + lifecycle ~70KB
240
+ ├── {project_name} Deployment Architecture.md ← [Type-3] Deployment topology ~40KB
241
+ ├── XX_{component}_Design.md × N ← [Type-4] 20~100KB each
242
+ ├── XX_{project_name}_Core_API_Product_Code_Mapping.md ← [Type-5] Generated only when product docs exist
243
+ ├── XX_{project_name}_Product_Rules_Cheat_Sheet.md ← [Type-6]
244
+ ├── XX_{project_name}_Business_Development_SOP.md ← [Type-7]
245
+ ├── {knowledge_enhancement_doc} × N ← [Type-8] Anti-patterns / RPC contracts / troubleshooting / knowledge library
246
+ └── graph/ ← [Type-9] Graph RAG graph document set
247
+ ├── README.md ← Graph index + lookup by question type
248
+ ├── G1_{project_name}_Component_Dependency_Matrix.md
249
+ ├── G2_{project_name}_Component_Call_Chain_Overview.md
250
+ ├── G3_{project_name}_Data_Flow_and_Storage_Dependencies.md
251
+ ├── G4_{project_name}_Error_Code_Component_Map.md
252
+ ├── G5_{project_name}_Cross_Component_Interaction_Scenarios.md
253
+ ├── G6_{project_name}_Knowledge_Graph_Triples.md
254
+ ├── G7_{project_name}_Architecture_Risks_and_Impact_Analysis.md
255
+ ├── G8_{project_name}_Core_Config_Parameter_Index.md
256
+ └── G9_{project_name}_Business_Rule_Constraint_Matrix.md
257
+
258
+ _review/ ← Process files (not part of the knowledge base)
259
+ ├── progress.json ← Resume-from-checkpoint + incremental update state
260
+ ├── metadata.json ← Code baseline version
261
+ ├── interface-inventory.json ← Interface scan baseline (Phase K1 Step 5)
262
+ ├── k1-architecture-map.md ← Architecture reverse-engineering result (confirmed by the user)
263
+ ├── k2-doc-list.md ← Document inventory + accuracy statistics
264
+ ├── k3-consistency-check.md ← Cross-document consistency check report (Phase K3 Step 3)
265
+ └── k4-quality-report.md ← Quality report (incl. E2E validation results)
266
+ ```
267
+
268
+ ---
269
+
270
+ ## Control between phases
271
+
272
+ | User reply | Behaviour |
273
+ |---------|------|
274
+ | "continue" / "go on" / "ok" | Enter the next phase |
275
+ | "stop" | Stop; files generated so far stay usable |
276
+ | Describes a problem directly | Adjust, reconfirm, then continue |
277
+ | Edits files directly and then replies "continue" | Continue based on the edited file contents |
278
+
279
+ ---
280
+
281
+ ## Constraints
282
+
283
+ - **The main agent does no code analysis**: all of it is done by dedicated Agents; Read the corresponding agent file before starting one
284
+ - **No redundant output**: Write generated files directly; never print the full content in the conversation first
285
+ - **Component document naming**: `XX_{component}_Design.md` (XX is a two-digit number assigned in dependency-chain order, lower layers get lower numbers)
286
+ - **When no product docs exist**: Type-5/6 may be skipped, or constraint values marked `[PRODUCT_DOC_MISSING]`; never guess
287
+ - **Parallel mode**: a Type-4 batch must send all Agent calls concurrently in the same message; serial batches run in order
288
+
289
+ ### Honesty Rules
290
+
291
+ - **No invention out of thin air**: every relation in the graph must have an explicit basis in a component document; never guess from names
292
+ - **Confidence must not be faked**: EXTRACTED=1.0, INFERRED 0.4~0.9 by evidence strength, AMBIGUOUS 0.1~0.3; the 0.5 default is banned
293
+ - **[UNVERIFIED] must not be hidden**: above 20%, add a visible warning at the top of the document
294
+ - **Quality numbers shown in full**: validate_kb.py output must not show only the passing items
295
+ - **Token cost transparency**: after every batch, show the number of files read and the estimated token consumption
296
+ - **When unsure, prefer AMBIGUOUS**: better to mark as pending confirmation than to delete or pretend certainty
297
+
298
+ ---
299
+
300
+ ## Working with the TeamAI CLI (must read)
301
+
302
+ | Phase | Command / path |
303
+ |------|-------------|
304
+ | Phase 0 structural baseline | `teamai codebase --extract <repo> --project <slug>` (writes `<repo>/teamwiki/`) |
305
+ | Deep knowledge | Use `teamai codebase --deep-enrich --project <slug> --output <repo>` after extract has written `teamwiki/evidence/code/<slug>/`. `--output` is the repository root, not the `teamwiki/` directory. Prefix with `teamai --dry-run` to preview without writing. TeamAI does not ship a separate team-wiki CLI. No extra plugin is required. |
306
+ | Compile into the wiki after K3 | Skip. TeamAI does not ship a separate team-wiki CLI. Continue with this skill using `teamai` and the files under this skill directory. No extra plugin is required. |
307
+ | Product docs into the graph | Skip. Same English note as above. |
308
+ | Product ↔ code bridging | Use `teamai codebase --reconcile --output <repo>` after product pages and extracted code pages are under `<repo>/teamwiki/`. Prefix with `teamai --dry-run` to preview without updating the graph. |
309
+ | One-shot refresh | Use `teamai codebase --extract <repo> --project <slug> --incremental`, reusing the Phase 0 repository path and project slug even when running from another directory. Do not look for another CLI. |
310
+ | Quality assessment | Use `python3 "{SKILL_DIR}/scripts/validate_kb.py" <output_dir>` and `teamai codebase --lint --output <repo>` to check `<repo>/teamwiki/` (`--output` takes the repository root, not the `teamwiki/` directory). Skip any extra evaluate binary. |
311
+
312
+ **Path convention**: `{SKILL_DIR}` is the directory printed by `teamai skill path wiki`. The methodology is in `{SKILL_DIR}/references/methodology/`, sub-agent prompts in `{SKILL_DIR}/references/agents/`, and scripts in `{SKILL_DIR}/scripts/`.
313
+
314
+ The whole workflow runs within the content served by `teamai skill get wiki` and the `teamai` CLI. No extra plugin is required.
@@ -0,0 +1,344 @@
1
+ # Graph RAG Agent
2
+
3
+ ## Responsibility
4
+
5
+ Extract cross-component relationship information from the generated knowledge base component documents and produce a structured graph document set (G1~G9), solving the information-scattering problem RAG retrieval faces in "cross-component relationship query" scenarios.
6
+
7
+ **This agent is started once, serially, by the main agent in Phase K3.**
8
+
9
+ ## Input package
10
+
11
+ ```
12
+ all_kb_docs_dir: knowledge base output root directory (contains all Type-1~8 documents)
13
+ architecture_map: full content of _review/k1-architecture-map.md
14
+ doc_list: _review/k2-doc-list.md (document list)
15
+ project_name: project name (used for document naming)
16
+ output_dir: graph document output directory (<all_kb_docs_dir>/graph/)
17
+ methodology_file: {SKILL_DIR}/references/methodology/phase2-document-types.md, §Type-9 content
18
+ ```
19
+
20
+ ## Execution steps
21
+
22
+ ### Step 1: Relationship extraction
23
+
24
+ Scan all component documents (Type-4) under `all_kb_docs_dir` and extract from the AI Quick Reference table and the body:
25
+
26
+ ```
27
+ Scan dimensions:
28
+ ├── Call relationships (upstream component -> this component, this component -> downstream component, communication method)
29
+ ├── Storage dependencies (which DB/Redis/MQ are read/written)
30
+ ├── Message topology (published/consumed Exchange/Topic/Queue/RoutingKey)
31
+ ├── State transitions (operation -> start state -> intermediate state -> final state, state field values)
32
+ ├── Constraints (operation -> state prerequisites -> hardware constraints -> billing constraints -> quota)
33
+ ├── Config mapping (config item -> affected behavior -> change risk)
34
+ └── Error code ownership (error code range -> component -> troubleshooting direction)
35
+ ```
36
+
37
+ **Three-state confidence labelling** (every relationship/triple must be labelled, no omissions):
38
+
39
+ | Label | Meaning | Evidence basis | Confidence score |
40
+ |------|------|---------|-----------|
41
+ | `EXTRACTED` | Relationship explicitly described in a component document (e.g. "Upstream component: Aurora(RPC)") | Explicitly recorded in code/docs | 1.0 |
42
+ | `INFERRED` | Reasonably inferred relationship (e.g. a dependency chain implied by an architecture diagram) | Structural evidence + reasonable inference | 0.6~0.9 |
43
+ | `AMBIGUOUS` | Uncertain relationship, needs manual confirmation | Weak or contradictory evidence | 0.1~0.3 |
44
+
45
+ > ⚠️ **Never use 0.5 as a default score**. Evaluate every relationship independently: INFERRED with a direct code reference gets 0.8~0.9, inference based only on naming gets 0.6~0.7, and only genuinely unclear cases use AMBIGUOUS.
46
+
47
+ Build intermediate data structures (in memory, do not write files):
48
+ - `relations[]`: (from, to, protocol, scenario, **confidence: EXTRACTED|INFERRED|AMBIGUOUS**, **confidence_score: 0.1~1.0**)
49
+ - `state_transitions[]`: (entity, from_state, to_state, trigger_op, state_field_value, **confidence**, **confidence_score**)
50
+ - `constraints[]`: (operation, state_req, hardware_req, billing_req, quota_req, **confidence**, **confidence_score**)
51
+ - `config_items[]`: (key, default, component, behavior, change_risk, effect_mode)
52
+ - `error_codes[]`: (code_range, component, meaning, debug_direction)
53
+ - `triples[]`: (subject, predicate, object, protocol, scenario, **confidence: EXTRACTED|INFERRED|AMBIGUOUS**, **confidence_score: 0.1~1.0**)
54
+
55
+ ### Step 2: Generate graph documents one by one
56
+
57
+ Generate G1~G9 in order (serially, Write each one as soon as it is complete):
58
+
59
+ ---
60
+
61
+ #### G1: Component Dependency Matrix
62
+
63
+ ```markdown
64
+ # {project_name} Component Dependency Matrix
65
+ <!-- search-anchor: component dependencies, dependency matrix, communication method, call relationships -->
66
+ ## 🤖 AI Quick Reference
67
+ | Document scope | Answers the retrieval question "who depends on X? what does X depend on?" |
68
+ | Core value | N×N communication matrix + forward/reverse dependency index |
69
+ | Use cases | Change impact assessment, service dependency review, architecture refactoring planning |
70
+
71
+ ## N×N component communication matrix
72
+ (rows: caller, columns: callee, values: `RPC`/`MQ`/`DB`/`—`, confidence label in brackets)
73
+ Example: `RPC[E]` = EXTRACTED, `MQ[I:0.8]` = INFERRED 0.8, `RPC[A]` = AMBIGUOUS
74
+
75
+ ## Forward dependency index (what A depends on)
76
+ | Component | Depends on | Communication method | Confidence | Typical scenario |
77
+
78
+ ## Reverse dependency index (who depends on A)
79
+ | Component | Depended on by | Communication method | Confidence | Typical scenario |
80
+
81
+ ## External service dependencies
82
+ | External service | Depended on by which components | Communication method | Confidence | Degradation strategy |
83
+
84
+ ## Confidence statistics
85
+ | Label | Count | Notes |
86
+ |------|------|------|
87
+ | EXTRACTED | N | Directly described in code/docs |
88
+ | INFERRED | N | Reasonable inference, scored 0.6~0.9 |
89
+ | AMBIGUOUS | N | Uncertain, needs manual confirmation |
90
+ ```
91
+
92
+ ---
93
+
94
+ #### G2: Component Call Chain Overview + state machines
95
+
96
+ ```markdown
97
+ # {project_name} Component Call Chain Overview and State Machines
98
+ <!-- search-anchor: call chain, state machine, end-to-end chain, API chain -->
99
+ ## 🤖 AI Quick Reference
100
+ | Document scope | Answers the retrieval question "which modules does API X pass through? how do entity states transition?" |
101
+ | Core value | End-to-end chains of core APIs + complete state machines + operation-state constraint matrix |
102
+
103
+ ## Core API end-to-end call chains
104
+ (for each core API, use the standard call chain format + a mermaid sequence diagram)
105
+
106
+ ## Complete state machines of core entities
107
+ (mermaid stateDiagram-v2, annotated with state field values and triggering operations)
108
+
109
+ ## Operation-state constraint quick matrix
110
+ | Operation \ Current state | State A | State B | ... |
111
+ (✅ allowed / ❌ forbidden / ⚠️ conditional)
112
+
113
+ ## AI state-judgement reasoning rules
114
+ (mermaid graph TD decision tree)
115
+ ```
116
+
117
+ ---
118
+
119
+ #### G3: Data Flow and Storage Dependencies
120
+
121
+ ```markdown
122
+ # {project_name} Data Flow and Storage Dependencies
123
+ <!-- search-anchor: data flow, storage dependencies, MQ topology, cache -->
124
+ ## Storage system dependency matrix
125
+ | Component | MySQL | Redis | MQ | Object storage | Other |
126
+
127
+ ## MQ queue topology
128
+ | Exchange/Topic | Routing Key | Producer | Consumer | Message meaning |
129
+
130
+ ## Cache strategy matrix
131
+ | Component | Cache key pattern | TTL | Invalidation strategy |
132
+ ```
133
+
134
+ ---
135
+
136
+ #### G4: Error Code Component Map
137
+
138
+ ```markdown
139
+ # {project_name} Error Code Component Map
140
+ <!-- search-anchor: error code, error mapping, InvalidParameter -->
141
+ ## Error code range allocation
142
+ | Error code range/prefix | Owning component | Meaning scope |
143
+
144
+ ## External -> internal error code mapping
145
+ | External error code | Internal component | Internal meaning | Troubleshooting direction |
146
+ ```
147
+
148
+ ---
149
+
150
+ #### G5: Cross-Component Interaction Scenarios
151
+
152
+ For each core business scenario, generate:
153
+ ```markdown
154
+ ## Scenario N: {scenario name}
155
+ <!-- typical scenarios: create/delete/modify resources, quota checks, billing, state changes, etc. -->
156
+ ```mermaid
157
+ sequenceDiagram
158
+ actor User
159
+ participant A as {ComponentA}
160
+ participant B as {ComponentB}
161
+ ...
162
+ ```
163
+ **Normal flow**: step descriptions
164
+ **Exception handling**: each exception branch
165
+ ```
166
+
167
+ Requirement: >=10 scenarios, covering the main write operations and key read operations.
168
+
169
+ ---
170
+
171
+ #### G6: Knowledge Graph Triples
172
+
173
+ ```markdown
174
+ # {project_name} Knowledge Graph Triples
175
+ <!-- search-anchor: knowledge graph, triples, multi-hop reasoning -->
176
+
177
+ ## Ontology definition
178
+ ### Entity types: Service, Handler, Config, Table, Queue, API, ErrorCode
179
+ ### Relationship types: CALLS, PUBLISHES, CONSUMES, READS, WRITES, CONFIGURES, MAPS_TO
180
+
181
+ ## Explicit triples (>=100)
182
+ | Subject | Predicate | Object | Protocol/Scenario | Confidence | Score |
183
+
184
+ > Every triple's Confidence must be `EXTRACTED` / `INFERRED` / `AMBIGUOUS`; Score must not be omitted and must not default to 0.5.
185
+
186
+ ## Multi-hop dependency path index
187
+ | Query pattern | Example path |
188
+ | "Which tables does A ultimately write to?" | A→(CALLS)→B→(WRITES)→Table |
189
+
190
+ ## Reverse reachability index
191
+ | Target node | Reachable paths |
192
+ ```
193
+
194
+ ---
195
+
196
+ #### G7: Architecture Risks and Impact Analysis
197
+
198
+ ```markdown
199
+ # {project_name} Architecture Risks and Impact Analysis
200
+ <!-- search-anchor: architecture risk, blast radius, impact surface -->
201
+ ## Component risk level summary
202
+ | Component | Risk level | Blast radius | Notes |
203
+ (🔴 high / 🟡 medium / 🟢 low)
204
+
205
+ ## Blast radius analysis of key components (>=3 high-risk components)
206
+ Impact chain analysis when component X fails
207
+
208
+ ## Critical paths and bottleneck identification
209
+ ## Cluster analysis (which components form tightly coupled clusters)
210
+ ## Change risk assessment matrix
211
+ ```
212
+
213
+ ---
214
+
215
+ #### G8: Core Config Parameter Index
216
+
217
+ ```markdown
218
+ # {project_name} Core Config Parameter Index
219
+ <!-- search-anchor: config parameters, config index, config changes -->
220
+ ## Layered configuration architecture diagram (mermaid)
221
+
222
+ ## Config parameter tables per layer
223
+ | Config item | Owning component | Default | Affected behavior | Change risk | Effect mode |
224
+ (change risk: 🟢 low / 🟡 medium / 🔴 high; effect mode: hot reload / restart required)
225
+
226
+ ## Config change impact quick reference
227
+ | Change type | Impact scope | Effect mode | Rollback strategy |
228
+
229
+ ## When answering "how do I change config XX", the AI must always state:
230
+ 1. Config file location
231
+ 2. Impact scope
232
+ 3. Effect mode
233
+ 4. Rollback strategy
234
+ 5. Change risk
235
+ 6. Whether a canary rollout is needed
236
+ ```
237
+
238
+ ---
239
+
240
+ #### G9: Business Rule Constraint Matrix
241
+
242
+ ```markdown
243
+ # {project_name} Business Rule Constraint Matrix
244
+ <!-- search-anchor: business rules, constraint matrix, operation constraints, AI reasoning -->
245
+ ## Operation precondition matrix
246
+ | Operation | State requirement | Hardware constraint | Billing constraint | Quota constraint | Other constraints |
247
+
248
+ ## Constraint decision tree (mermaid graph TD)
249
+ (covers the multi-layer constraint check flow of the main operations)
250
+
251
+ ## Special instance type constraint summary
252
+ | Instance/resource type | Restricted operations | Reason |
253
+ (✅ allowed / ❌ forbidden / ⚠️ conditional)
254
+
255
+ ## AI reasoning rules quick reference
256
+ (mermaid flowchart: the layer-by-layer check order the AI follows when judging "can operation X be performed")
257
+ ```
258
+
259
+ ---
260
+
261
+ ### Step 3: Generate the graph directory README
262
+
263
+ Write to `{output_dir}/README.md`:
264
+ ```markdown
265
+ # {project_name} Graph Document Set (Graph RAG)
266
+ <!-- search-anchor: graph documents, Graph RAG, relationship index -->
267
+
268
+ ## Relationship to the main document system
269
+ (graph documents do not replace component documents; they provide a structured index from the relationship perspective)
270
+
271
+ ## Document directory
272
+ | File | Size | Core content |
273
+
274
+ ## Look up by question type
275
+ | Question type | Example question | Document to consult |
276
+ | Dependencies | "Who depends on X?" | G1 Component Dependency Matrix |
277
+ | Call chains | "Which modules does API X pass through?" | G2 Call Chain Overview |
278
+ | Data location | "Where is the data stored?" | G3 Data Flow and Storage Dependencies |
279
+ | Error troubleshooting | "Which module does error code XXX belong to?" | G4 Error Code Component Map |
280
+ | Scenario handbook | "What is the full quota check flow?" | G5 Cross-Component Interaction Scenarios |
281
+ | Multi-hop reasoning | "What does A indirectly depend on?" | G6 Knowledge Graph Triples |
282
+ | Risk assessment | "How big is the impact if X goes down?" | G7 Architecture Risks and Impact Analysis |
283
+ | Config changes | "How do I change config XX?" | G8 Core Config Parameter Index |
284
+ | Operation constraints | "Can I do XX?" | G9 Business Rule Constraint Matrix |
285
+
286
+ ## Suggested retrieval routing rules
287
+ (keyword -> document to search first)
288
+
289
+ ## Maintenance notes
290
+ (when and how far graph documents must be updated after component documents change)
291
+ ```
292
+
293
+ ### Step 4: Return summary
294
+
295
+ ```
296
+ Graph RAG generation complete:
297
+ Generated documents: G1~G9, 9 in total + README
298
+ - G1_Component_Dependency_Matrix.md: {N}KB, {N} components, {N} relationships
299
+ Confidence: EXTRACTED {N} / INFERRED {N} / AMBIGUOUS {N}
300
+ - G2_Component_Call_Chain_Overview.md: {N}KB, {N} call chains, {N} state machine states
301
+ - G3_Data_Flow_and_Storage_Dependencies.md: {N}KB
302
+ - G4_Error_Code_Component_Map.md: {N}KB, {N} error code ranges
303
+ - G5_Cross_Component_Interaction_Scenarios.md: {N}KB, {N} scenario sequence diagrams
304
+ - G6_Knowledge_Graph_Triples.md: {N}KB, {N} triples
305
+ Confidence: EXTRACTED {N} / INFERRED {N} / AMBIGUOUS {N}
306
+ - G7_Architecture_Risks_and_Impact_Analysis.md: {N}KB
307
+ - G8_Core_Config_Parameter_Index.md: {N}KB, {N} config items
308
+ - G9_Business_Rule_Constraint_Matrix.md: {N}KB
309
+ AMBIGUOUS entries summary (need manual confirmation): {N} places
310
+ - Example: "Aurora→Compute communication method uncertain (not specified in docs) [A:0.2]"
311
+ Issues found: {issues or "none"}
312
+
313
+ ⚠️ Note to the main agent: once Graph RAG is complete, immediately run Phase K3 Step 3 (cross-document consistency check).
314
+ ```
315
+
316
+ ## Output
317
+
318
+ ```
319
+ <output_dir>/README.md
320
+ <output_dir>/G1_{project_name}_Component_Dependency_Matrix.md
321
+ <output_dir>/G2_{project_name}_Component_Call_Chain_Overview.md
322
+ <output_dir>/G3_{project_name}_Data_Flow_and_Storage_Dependencies.md
323
+ <output_dir>/G4_{project_name}_Error_Code_Component_Map.md
324
+ <output_dir>/G5_{project_name}_Cross_Component_Interaction_Scenarios.md
325
+ <output_dir>/G6_{project_name}_Knowledge_Graph_Triples.md
326
+ <output_dir>/G7_{project_name}_Architecture_Risks_and_Impact_Analysis.md
327
+ <output_dir>/G8_{project_name}_Core_Config_Parameter_Index.md
328
+ <output_dir>/G9_{project_name}_Business_Rule_Constraint_Matrix.md
329
+ Returned summary string
330
+ ```
331
+
332
+ ## Constraints
333
+
334
+ - **Component documents are the sole source for relationship extraction**: do not read the raw code directly, to avoid inconsistency with the Phase K2 output
335
+ - **Three-state confidence is mandatory**: every relationship/triple must be labelled `EXTRACTED`/`INFERRED`/`AMBIGUOUS`, no omissions
336
+ - **Never use 0.5 as the default confidence**: score every relationship independently; INFERRED with direct structural evidence 0.8~0.9, naming-based inference 0.6~0.7, weak evidence 0.4~0.5; AMBIGUOUS uses 0.1~0.3
337
+ - **Never invent relationships**: if the component documents provide no basis, label it AMBIGUOUS rather than fabricating EXTRACTED
338
+ - **Every graph document must have an AI Quick Reference table**
339
+ - **Every graph document must have a search-anchor**
340
+ - **Graph documents do not replace component documents**: they only provide a structured index from the relationship perspective
341
+ - **State machines must use mermaid stateDiagram-v2**
342
+ - **Constraint decision trees must use mermaid graph TD**
343
+ - **Triples must follow the (Subject, Predicate, Object, Confidence, Score) format**
344
+ - **Operation-state constraints must be in ✅/❌/⚠️ matrix format**