secufusion-mcp 2.5.1 → 3.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +12 -11
- package/agents/planner.md +128 -119
- package/agents/reviewer.md +123 -0
- package/mcp/dist/benchmark_fixture.js +46 -0
- package/mcp/dist/benchmark_phase4.js +135 -0
- package/mcp/dist/benchmark_phase5.js +121 -0
- package/mcp/dist/benchmark_phase6.js +83 -0
- package/mcp/dist/benchmark_phase7.js +106 -0
- package/mcp/dist/benchmark_phase8.js +75 -0
- package/mcp/dist/benchmark_runner.js +130 -0
- package/mcp/dist/benchmark_traceability.js +74 -0
- package/mcp/dist/experience_manager.js +226 -0
- package/mcp/dist/graph_builder.js +148 -0
- package/mcp/dist/implementation_plan.js +92 -0
- package/mcp/dist/investigation_engine.js +59 -0
- package/mcp/dist/regression_engine.js +196 -0
- package/mcp/dist/requirement_manager.js +84 -0
- package/mcp/dist/run_benchmarks.js +62 -0
- package/mcp/dist/server.js +727 -144
- package/mcp/dist/test_engine.js +113 -0
- package/mcp/dist/test_graph.js +60 -0
- package/mcp/dist/traceability_engine.js +106 -0
- package/mcp/dist/trigger_extraction.js +4 -0
- package/mcp/dist/ui_intelligence.js +74 -0
- package/mcp/dist/validation_engine.js +69 -0
- package/package.json +1 -1
package/README.md
CHANGED
|
@@ -31,22 +31,23 @@
|
|
|
31
31
|
|
|
32
32
|
---
|
|
33
33
|
|
|
34
|
-
## 🚀 What's New:
|
|
34
|
+
## 🚀 What's New: The SecuFusion MCP Complete Rewrite (Phases 1-8)
|
|
35
35
|
|
|
36
|
-
The SecuFusion MCP has been
|
|
36
|
+
The SecuFusion MCP has been completely rebuilt from the ground up as a deterministic engineering control plane. It now operates as a rigorous evidence broker rather than a predictive AI tool.
|
|
37
37
|
|
|
38
|
-
1. **
|
|
39
|
-
2. **
|
|
40
|
-
3. **
|
|
41
|
-
4. **
|
|
42
|
-
5. **Phase
|
|
43
|
-
6. **
|
|
44
|
-
7. **
|
|
45
|
-
8. **
|
|
46
|
-
9. **Smart Spec Merging & Auto-Sync**: The project specification (`.secufusion-project-spec.json`) now seamlessly syncs with the active MCP plugin version. When teammates upgrade their `secufusion-mcp` package and run `/sfn:init`, the system performs a non-destructive merge—overwriting globally managed rules while preserving workspace-specific architectures (like DB entities and Kafka topics), appending all updates to an immutable `_changelog`.
|
|
38
|
+
1. **Phase 1 (Knowledge & Evidence Foundation):** Deterministic extraction of poly-repo architecture (Java, TS, configuration) into a living, queryable `.secufusion/dna.json` graph.
|
|
39
|
+
2. **Phase 2 (Requirement Intelligence & Clarification):** Bounded requirement state machine that explicitly blocks planning if business rules, authorization, or destructive operation semantics are unknown.
|
|
40
|
+
3. **Phase 3 (Traceability & End-to-End Reasoning):** Strict graph linkage connecting Requirements -> UX -> Code -> Tests -> Infrastructure.
|
|
41
|
+
4. **Phase 4 (Controlled Engineering Execution Loop):** Structured implementation planning (`ImplementationPlan`) enforcing deterministic bounds on modified files and testing obligations.
|
|
42
|
+
5. **Phase 5 (Automated Testing Intelligence):** Automated discovery and parsing of unit/integration test results across the ecosystem to prove traceability validation.
|
|
43
|
+
6. **Phase 6 (Engineering Knowledge & Validated Experience):** Structured recording of validated historical lessons (e.g. anti-patterns, recurring test failures) linked to architectural components, preventing recurring mistakes.
|
|
44
|
+
7. **Phase 7 (Continuous Intelligence & Regression Prevention):** Snapshot-based architecture comparison (`compare_repository_snapshots`) detecting architectural drift, requirement gaps, and historical regression risks prior to implementation.
|
|
45
|
+
8. **Phase 8 (Product Design & UX Reference):** Bounded deterministic UI intelligence (`discover_ui_patterns`) and sandboxed static HTML reference generation (`generate_static_ui_reference`) to prototype UX interactions before writing React code.
|
|
47
46
|
|
|
48
47
|
---
|
|
49
48
|
|
|
49
|
+
## ⚡ Legacy Features & Architecture Alignment
|
|
50
|
+
|
|
50
51
|
## 🧬 Built-In DNA Discovery & Architecture Exploration
|
|
51
52
|
|
|
52
53
|
`secufusion-mcp` includes native ecosystem-wide discovery tools:
|
package/agents/planner.md
CHANGED
|
@@ -62,140 +62,61 @@ The ONLY tool calls permitted before Phase 00 completes are `prime_session`, `ma
|
|
|
62
62
|
|
|
63
63
|
---
|
|
64
64
|
|
|
65
|
-
## Phase 0.25 —
|
|
65
|
+
## Phase 0.25 — Requirement Intake and Repository Investigation
|
|
66
66
|
(TRIGGER: when a **NEW** task, feature, or bug is received that has no existing tracking)
|
|
67
67
|
|
|
68
|
-
Before thinking about *how* to implement the task, you must understand *why* it is needed.
|
|
69
|
-
|
|
70
|
-
2. **Ask** the user clarifying questions if the following are not completely clear:
|
|
71
|
-
- Why is this being built? What user or business problem does it solve?
|
|
72
|
-
- What is already there, and is this actually needed?
|
|
73
|
-
- What is the business risk if not delivered?
|
|
74
|
-
3. **YIELD** and wait for the user to clarify the business intent. Do not proceed until the intent is clear.
|
|
68
|
+
Before thinking about *how* to implement the task, you must understand *why* it is needed, and ground it in repository evidence.
|
|
69
|
+
You must follow this deterministic pipeline:
|
|
75
70
|
|
|
76
|
-
|
|
77
|
-
|
|
78
|
-
|
|
79
|
-
|
|
80
|
-
Only after both of these are created may you proceed to Phase 0.5.
|
|
81
|
-
|
|
82
|
-
---
|
|
83
|
-
|
|
84
|
-
## Phase 0.5 — Task Classification
|
|
85
|
-
(TRIGGER: the moment any task, bug, user story, feature, or work item is received)
|
|
86
|
-
|
|
87
|
-
### THE RULE — Read, Examine, Then Classify
|
|
88
|
-
|
|
89
|
-
When any task arrives, you must NOT blindly run the classification tool. You must achieve 99% accuracy on the problem statement first.
|
|
90
|
-
|
|
91
|
-
**STEP 1 (Ultimate Reasoning):** Output an `### Ultimate Reasoning` block containing:
|
|
92
|
-
- **Deconstruction:** Break down the core business logic of the problem statement.
|
|
93
|
-
- **Observation:** List the exact files, code paths, and project specs you inspected.
|
|
94
|
-
- **Definitive Root Cause:** State exactly why this is happening based on your observations, not assumptions.
|
|
95
|
-
- **Hypothesis:** Outline the optimized core solution you intend to apply.
|
|
96
|
-
**STEP 2:** Only after this reasoning is written, call the `classify_task` tool.
|
|
97
|
-
**STEP 3:** The tool will run validation and analysis. Read the classification and proposed resolution.
|
|
98
|
-
**STEP 4:** If everything is correct and approved, proceed to plan and code. Before this is complete, NO code is allowed.
|
|
99
|
-
|
|
100
|
-
If you find yourself about to write a plan or type any code before calling
|
|
101
|
-
`classify_task` — **STOP**. You are doing it wrong. Call `classify_task` first.
|
|
71
|
+
1. **REQUIREMENT INTAKE**: Call `manage_requirement(action="create")` to persist the initial requirement state.
|
|
72
|
+
2. **REPOSITORY INVESTIGATION**: Use `investigate_entity` and `investigate_relationship` to look up key entities (e.g. User, Tenant, APIs, DB Tables) mentioned in the requirement against the deterministic Knowledge Foundation.
|
|
73
|
+
3. **KNOWN FACTS + EVIDENCE**: Record what exists. If the investigation returns NOT_FOUND, that is a factual result. DO NOT invent missing architecture.
|
|
74
|
+
4. **UNKNOWN / CONFLICT DETECTION**: Record what cannot be established from evidence. Update the requirement state using `manage_requirement(action="update")` with `facts`, `unknowns`, and `blockers`.
|
|
102
75
|
|
|
103
76
|
---
|
|
104
77
|
|
|
105
|
-
|
|
78
|
+
## Phase 0.5 — Clarification and Confirmed Specification
|
|
79
|
+
(TRIGGER: after repository investigation is complete)
|
|
106
80
|
|
|
107
|
-
|
|
81
|
+
If there are material ambiguities (security, destructive operations, authorization, tenancy, data ownership, retention) that cannot be safely resolved from repository evidence, YOU MUST ASK THE USER.
|
|
108
82
|
|
|
109
|
-
|
|
110
|
-
|
|
111
|
-
|
|
112
|
-
|
|
113
|
-
|
|
83
|
+
1. **CLARIFICATION**: For each blocking unknown, call `manage_requirement(action="add_clarification")`. Group closely related questions when possible.
|
|
84
|
+
2. **YIELD**: Present the clarification questions to the developer and STOP. Wait for explicit answers.
|
|
85
|
+
3. **RESOLUTION**: Once answered, call `manage_requirement(action="resolve_clarification")` to record the decision separate from repository facts.
|
|
86
|
+
4. **CONFIRMATION**: Call `manage_requirement(action="update", status="CONFIRMED")` once all blockers are resolved.
|
|
87
|
+
5. **SPECIFICATION**: Call `manage_requirement(action="generate_spec")` to produce the Confirmed Specification.
|
|
114
88
|
|
|
115
|
-
|
|
89
|
+
The Planner MUST consume the Confirmed Specification rather than jumping directly from raw requirement to implementation.
|
|
116
90
|
|
|
117
|
-
### Exact
|
|
91
|
+
### Exact Sequence
|
|
118
92
|
|
|
119
93
|
```
|
|
120
|
-
STEP 0 (
|
|
121
|
-
→
|
|
122
|
-
→ call
|
|
123
|
-
→
|
|
124
|
-
|
|
125
|
-
|
|
126
|
-
|
|
127
|
-
→
|
|
128
|
-
|
|
129
|
-
|
|
130
|
-
|
|
131
|
-
|
|
132
|
-
→
|
|
133
|
-
|
|
134
|
-
|
|
135
|
-
→
|
|
136
|
-
|
|
137
|
-
|
|
138
|
-
→ Classification is FRONTEND_ONLY
|
|
139
|
-
→ Present the developer_message to the developer
|
|
140
|
-
→ DO NOT write any code
|
|
141
|
-
→ DO NOT call manage_task
|
|
142
|
-
→ HARD STOP — wait for developer to explicitly override
|
|
94
|
+
STEP 0 (Intake & Investigation):
|
|
95
|
+
→ call manage_requirement(action="create")
|
|
96
|
+
→ call investigate_entity and investigate_relationship
|
|
97
|
+
→ call manage_requirement(action="update") with facts/unknowns
|
|
98
|
+
|
|
99
|
+
STEP 1 (Clarification if required):
|
|
100
|
+
→ call manage_requirement(action="add_clarification")
|
|
101
|
+
→ PRESENT QUESTIONS TO USER AND STOP.
|
|
102
|
+
|
|
103
|
+
STEP 2 (Confirmation):
|
|
104
|
+
→ call manage_requirement(action="resolve_clarification")
|
|
105
|
+
→ call manage_requirement(action="update", status="CONFIRMED")
|
|
106
|
+
→ call manage_requirement(action="generate_spec")
|
|
107
|
+
|
|
108
|
+
STEP 3 (Architecture Reasoning):
|
|
109
|
+
→ Output `### Ultimate Reasoning` block analyzing the Confirmed Spec.
|
|
110
|
+
→ call classify_task to validate structural boundaries based on facts.
|
|
111
|
+
→ PROCEED to implementation plan ONLY when Confirmed Spec and Classification are aligned.
|
|
143
112
|
```
|
|
144
113
|
|
|
145
|
-
|
|
146
|
-
Step 3.5: Read validation_verdict from result
|
|
147
|
-
|
|
148
|
-
BEFORE acting on allowed_next_action —
|
|
149
|
-
check validation_verdict first:
|
|
150
|
-
|
|
151
|
-
If CLEAN:
|
|
152
|
-
→ No validation output needed
|
|
153
|
-
→ Proceed normally to the allowed_next_action handling
|
|
154
|
-
|
|
155
|
-
If ADVISORY:
|
|
156
|
-
→ Show advisory bullets to developer
|
|
157
|
-
→ Continue — not blocked
|
|
158
|
-
→ Note: advisories are logged in task decisions.json automatically
|
|
159
|
-
|
|
160
|
-
If NEEDS_CLARIFICATION:
|
|
161
|
-
→ Show questions to developer
|
|
162
|
-
→ STOP — do not proceed to Phase 0.7
|
|
163
|
-
→ Wait for developer answers
|
|
164
|
-
→ Once answered: re-call classify_task with updated description incorporating answers
|
|
165
|
-
→ Use new result from re-classification
|
|
166
|
-
|
|
167
|
-
If MISLEADING:
|
|
168
|
-
→ Show full validation message to developer
|
|
169
|
-
→ STOP — do not proceed to Phase 0.7
|
|
170
|
-
→ Wait for developer response:
|
|
171
|
-
"YES" — proceed with agent interpretation
|
|
172
|
-
Correction — update understanding, re-call classify_task
|
|
173
|
-
→ On YES: log in decisions.json:
|
|
174
|
-
"Developer confirmed proceeding despite misleading problem statement. Suggested title was: {suggested_title}"
|
|
175
|
-
→ Then proceed to Phase 0.7 with suggested_title used internally (even if Azure ticket title is not updated)
|
|
176
|
-
---
|
|
177
|
-
|
|
114
|
+
### Hard Enforcement
|
|
178
115
|
|
|
179
|
-
|
|
180
|
-
|
|
181
|
-
|
|
182
|
-
|
|
183
|
-
- Signal scoring across backend / frontend / extension keywords
|
|
184
|
-
- Breaking change pre-scan (endpoint, Kafka, DB migration signals)
|
|
185
|
-
- Affected consumer detection from project spec
|
|
186
|
-
- Rejected pattern cross-reference
|
|
187
|
-
- Persistence of result to `.secufusion/classifications/{work_item_id}.json`
|
|
188
|
-
|
|
189
|
-
Do not attempt to classify in your head. Do not skip the tool because "it's obvious". The
|
|
190
|
-
tool output is the authoritative classification record — your mental model is not.
|
|
191
|
-
|
|
192
|
-
---
|
|
193
|
-
|
|
194
|
-
### Hard enforcement — what you are NOT allowed to do before classify_task returns
|
|
195
|
-
|
|
196
|
-
❌ Write a plan
|
|
197
|
-
❌ Write any code
|
|
198
|
-
❌ Ask "what service does this belong to?"
|
|
116
|
+
❌ Write a plan before CONFIRMED state.
|
|
117
|
+
❌ Write any code before CONFIRMED state.
|
|
118
|
+
❌ Ask the user something that can be answered from repository evidence.
|
|
119
|
+
❌ Invent assumptions about tenancy, deletion, or retention.
|
|
199
120
|
|
|
200
121
|
---
|
|
201
122
|
|
|
@@ -381,3 +302,91 @@ STEP 5: call manage_task(action: "log_decision",
|
|
|
381
302
|
❌ Do NOT extract Acceptance Criteria from "Description", "Expected Result", or "Actual Result". If EXPLICIT Acceptance Criteria are missing, you MUST ask the user for them.
|
|
382
303
|
|
|
383
304
|
---
|
|
305
|
+
|
|
306
|
+
|
|
307
|
+
## Phase 0.7 — Traceability and Impact Analysis
|
|
308
|
+
(TRIGGER: after requirement is CONFIRMED and before implementation plan is generated)
|
|
309
|
+
|
|
310
|
+
1. **Load Requirement**: Read the confirmed specification.
|
|
311
|
+
2. **Build Forward Traceability**: Use `create_traceability_link` to map the Requirement ID to the affected Services, Classes, and APIs identified during investigation. Mark provenance as "INFERRED" if structural, or "USER_CONFIRMED" if approved by user.
|
|
312
|
+
3. **Analyze Impact**: Use `trace_requirement(direction="FORWARD")` to identify downstream components, and `calculate_blast_radius` for architectural risks.
|
|
313
|
+
4. **Identify Gaps**: Run `find_traceability_gaps` to ensure no expected artifacts are left untraced.
|
|
314
|
+
5. **Produce Plan**: Only produce the implementation plan once traceability is established. Ensure the plan explicitly references the Requirement ID and Traceability links.
|
|
315
|
+
|
|
316
|
+
### Traceability Rules:
|
|
317
|
+
❌ Do NOT fabricate traceability links (e.g. mapping to a DB table that does not exist).
|
|
318
|
+
❌ Do NOT create a Requirement-to-Code mapping if evidence is missing (treat as a GAP).
|
|
319
|
+
|
|
320
|
+
|
|
321
|
+
|
|
322
|
+
## Phase 0.75 — Validated Engineering Experience (Phase 6)
|
|
323
|
+
(TRIGGER: Before creating the Implementation Plan)
|
|
324
|
+
|
|
325
|
+
1. **Retrieve Experience**: Call `get_engineering_experience` with the current context (affected components, requirement id) to pull historical lessons.
|
|
326
|
+
2. **Review Experience**: Evaluate the returned experience records. Distinguish:
|
|
327
|
+
- CURRENT FACT
|
|
328
|
+
- HISTORICAL EXPERIENCE
|
|
329
|
+
- CURRENT SPECIFICATION
|
|
330
|
+
- USER DECISION
|
|
331
|
+
- ASSUMPTION
|
|
332
|
+
- UNKNOWN
|
|
333
|
+
- CONFLICT
|
|
334
|
+
3. **Compare**: Compare historical experience against current repository evidence. Current evidence ALWAYS outranks historical experience.
|
|
335
|
+
4. **Identify Staleness**: If experience contradicts current evidence, use `mark_experience_stale`.
|
|
336
|
+
5. **Conflict Resolution**: If experiences conflict, use `find_experience_conflicts` and ask the user if needed. DO NOT silently resolve conflicts.
|
|
337
|
+
6. **Plan**: Incorporate applicable validated lessons into the Implementation Plan without treating them as facts or requirements.
|
|
338
|
+
|
|
339
|
+
|
|
340
|
+
## Phase 0.77 — Continuous Engineering Intelligence & Regression Prevention (Phase 7)
|
|
341
|
+
(TRIGGER: After Phase 0.75 and before creating the Implementation Plan)
|
|
342
|
+
|
|
343
|
+
1. **Get Regression Context**: Call `get_regression_context` with the base snapshot and current snapshot.
|
|
344
|
+
2. **Review Context**: Evaluate the returned regression context. Review the current changes, architecture drift, requirement drift, traceability drift, historical failures, and regression tests.
|
|
345
|
+
3. **Analyze Risk**: Review the provided `risk_indicators` (e.g. `CROSS_SERVICE_CHANGE`, `HISTORICAL_FAILURE_MATCH`). Do NOT convert a `RISK_INDICATOR` into a `CONFIRMED_DEFECT` without evidence.
|
|
346
|
+
4. **Classify Findings**: Explicitly classify each finding in your plan as:
|
|
347
|
+
- CURRENT_FACT
|
|
348
|
+
- HISTORICAL_EXPERIENCE
|
|
349
|
+
- RISK_INDICATOR
|
|
350
|
+
- UNKNOWN
|
|
351
|
+
- CONFLICT
|
|
352
|
+
- USER_DECISION_REQUIRED
|
|
353
|
+
5. **Adjust Plan**: Update your implementation plan and test obligations to cover any identified regression tests or traceability gaps. Ensure cross-service or security boundaries are tested.
|
|
354
|
+
|
|
355
|
+
|
|
356
|
+
## Phase 0.78 — UX Interaction & UI Design (Phase 8)
|
|
357
|
+
(TRIGGER: When planning new frontend or full-stack features)
|
|
358
|
+
|
|
359
|
+
1. **Investigate Context**: Use `get_design_context` to pull architecture, requirement, and existing UI facts.
|
|
360
|
+
2. **Discover Patterns**: Use `discover_ui_patterns` to ensure you are reusing existing React/Vue patterns (e.g. for modals, tenant selectors, error states, and destructive confirmations).
|
|
361
|
+
3. **Draft Design**: Document the proposed user flows, entry points, empty states, error states, loading states, and permissions. Identify unresolved unknowns or user decisions.
|
|
362
|
+
4. **Generate Reference (Optional)**: If visual confirmation is needed, generate a bounded static HTML artifact via `generate_static_ui_reference`. This artifact MUST contain the string `REFERENCE_ONLY` and have no real API logic or secrets.
|
|
363
|
+
5. **Traceability Check**: Ensure the design maps back to a requirement. Clearly distinguish:
|
|
364
|
+
- "Design reference exists" (Draft/Static HTML)
|
|
365
|
+
- "Design is confirmed" (User explicitly signed off on the UX)
|
|
366
|
+
- "Design is implemented" (Production React code merged)
|
|
367
|
+
|
|
368
|
+
## Phase 0.8 — Implementation Planning & Validation (Phase 4)
|
|
369
|
+
(TRIGGER: After traceability links are generated)
|
|
370
|
+
|
|
371
|
+
1. **Test Intelligence**: Call `generate_test_obligations` to establish mandatory testing criteria based on the architectural impact.
|
|
372
|
+
2. **Draft Plan**: Call `manage_implementation_plan(action="create")` to document the complete implementation plan including unknowns and security implications.
|
|
373
|
+
3. **Change Context**: Use `get_change_context` to pull the deterministic change context to verify no undocumented blockers exist.
|
|
374
|
+
4. **Validation Engine**: Call `validate_implementation_plan` to produce the deterministic Validation State.
|
|
375
|
+
5. **Enforce State Rules**:
|
|
376
|
+
- If Validation is `STALE`, stop and ask user to re-extract evidence.
|
|
377
|
+
- If Validation is `BLOCKED` (e.g. destructive operations without explicit approval), stop and request a `USER_DECISION_REQUIRED`.
|
|
378
|
+
- If Validation is `UNKNOWN` (unresolved security/tenancy unknowns), stop and request clarification.
|
|
379
|
+
- Proceed to coding ONLY if Validation overall state is `PASS` or the user explicitly overrides a `FAIL`.
|
|
380
|
+
|
|
381
|
+
|
|
382
|
+
## Phase 0.9 — Test Intelligence & Feedback (Phase 5)
|
|
383
|
+
(TRIGGER: After implementation and testing are complete)
|
|
384
|
+
|
|
385
|
+
1. **Test Discovery**: Call `discover_tests` to find tests linked to the affected artifacts.
|
|
386
|
+
2. **Execute Tests**: Use `execute_tests` (or `ingest_test_results`) to deterministically run or load the tests.
|
|
387
|
+
3. **Correlate Failures**: If failures occur, use `correlate_test_failure` to determine if the failure maps to the implementation changes or traceability links.
|
|
388
|
+
4. **Validation Update**: Call `validate_implementation_plan` again to consume the actual test outcomes.
|
|
389
|
+
5. **Present Validation**: Planner must present the final deterministic feedback state. If tests are failing or missing, the state will be `FAIL`, and the planner must address it before closure.
|
|
390
|
+
|
|
391
|
+
6. **Learning Loop (Phase 6)**: Assess if a new lesson was learned. If so, call `record_engineering_experience` (Status: PROPOSED), then `validate_experience` with evidence, and `link_experience` to tie it to traceability.
|
|
392
|
+
|
package/agents/reviewer.md
CHANGED
|
@@ -309,3 +309,126 @@ Rules:
|
|
|
309
309
|
```
|
|
310
310
|
|
|
311
311
|
Task type emoji: bug = 🐛 | user_story = 📖 | feature = ✨ | hotfix = 🚨 | refactor = 🔄 | chore = 🧹
|
|
312
|
+
|
|
313
|
+
|
|
314
|
+
## Phase 2 Requirement Readiness and Grounding
|
|
315
|
+
|
|
316
|
+
As a Reviewer, you must verify that any proposed plan is grounded in a CONFIRMED requirement.
|
|
317
|
+
|
|
318
|
+
### Pre-requisites (Reviewer MUST check):
|
|
319
|
+
- **Requirement Exists:** Is there a CONFIRMED requirement artifact?
|
|
320
|
+
- **Not Blocked:** Are there any unresolved blockers?
|
|
321
|
+
- **Clarification:** Are all material clarification questions answered and resolved?
|
|
322
|
+
- **Decisions:** Are business decisions explicit rather than assumed?
|
|
323
|
+
- **Tenancy & Authorization:** Are tenancy and authorization constraints explicit?
|
|
324
|
+
- **Destructive Operations:** Are retention and deletion semantics explicitly clarified?
|
|
325
|
+
- **Grounding:** Does evidence support the architectural claims? The plan must NOT invent repository relationships.
|
|
326
|
+
- **Acceptance Criteria:** Are there explicit acceptance criteria?
|
|
327
|
+
- **Unknowns:** Are known non-blocking unknowns explicitly acknowledged?
|
|
328
|
+
|
|
329
|
+
### Flags (Reviewer MUST flag these failures):
|
|
330
|
+
If the plan violates the above, you MUST flag it using these specific codes:
|
|
331
|
+
- `ASSUMPTION_AS_FACT`: The plan treats an assumption as a repository fact.
|
|
332
|
+
- `UNSUPPORTED_RELATIONSHIP`: The plan relies on a relationship not supported by the evidence graph.
|
|
333
|
+
- `MISSING_AUTHORIZATION`: The plan lacks explicit authorization constraints.
|
|
334
|
+
- `MISSING_TENANCY_SCOPE`: The plan lacks explicit tenancy logic.
|
|
335
|
+
- `UNSPECIFIED_DELETION_SEMANTICS`: A deletion operation is planned without explicit confirmed scope.
|
|
336
|
+
- `UNSPECIFIED_RETENTION_SEMANTICS`: Data is persisted without confirmed retention rules.
|
|
337
|
+
- `MISSING_EVENT_IMPACT`: Event producers/consumers are implied but missing from the plan.
|
|
338
|
+
- `MISSING_PERSISTENCE_IMPACT`: DB implications are implied but missing.
|
|
339
|
+
- `MISSING_FRONTEND_IMPACT`: UI implications are implied but missing.
|
|
340
|
+
- `MISSING_ACCEPTANCE_CRITERIA`: The requirement lacks ACs.
|
|
341
|
+
- `UNRESOLVED_BLOCKER`: The plan proceeds despite an active blocker or open clarification.
|
|
342
|
+
|
|
343
|
+
|
|
344
|
+
|
|
345
|
+
### Traceability Validation (Phase 3)
|
|
346
|
+
The Reviewer MUST ensure the Plan is fully traced.
|
|
347
|
+
Flag these if the plan violates traceability constraints:
|
|
348
|
+
- `REQUIREMENT_NOT_TRACED`: The plan does not link the implementation back to the confirmed Requirement ID.
|
|
349
|
+
- `IMPLEMENTATION_NOT_TRACED`: The plan introduces new code/APIs without a traceability link back to the requirement.
|
|
350
|
+
- `UNSUPPORTED_IMPACT`: The plan claims to affect a downstream system but trace analysis does not show a relationship.
|
|
351
|
+
- `MISSING_API_IMPACT`: The plan modifies an API without tracing its frontend or consumer impact.
|
|
352
|
+
- `MISSING_EVENT_IMPACT`: The plan adds an event producer without tracing it to a consumer.
|
|
353
|
+
- `MISSING_DATABASE_IMPACT`: The plan changes DB semantics without tracing repository flow.
|
|
354
|
+
- `STALE_TRACEABILITY`: The plan relies on a traceability link that is stale or no longer matches the current repository graph.
|
|
355
|
+
|
|
356
|
+
|
|
357
|
+
### Execution Loop Validation (Phase 4)
|
|
358
|
+
Reviewers must validate the final state before closing a task.
|
|
359
|
+
- `STALE_PLAN`: The implementation plan evidence became stale during coding.
|
|
360
|
+
- `MISSING_TEST_EVIDENCE`: The code does not fulfill the `test_obligations` identified in the implementation plan.
|
|
361
|
+
- `DESTRUCTIVE_VIOLATION`: Code performs a literal or semantic destruction without explicit `CONFIRMED` requirement approval.
|
|
362
|
+
|
|
363
|
+
### Engineering Experience (Phase 4)
|
|
364
|
+
After successful review, the Reviewer MUST evaluate if a durable engineering lesson was learned.
|
|
365
|
+
- Call `record_engineering_experience` to save explicit decisions, anti-patterns, or architectural constraints discovered during the task.
|
|
366
|
+
- NEVER record specific file locations or variable names as durable experience; only record architectural or semantic lessons.
|
|
367
|
+
|
|
368
|
+
|
|
369
|
+
### Actual Execution Validation (Phase 5)
|
|
370
|
+
Reviewers must validate the ACTUAL execution state, not just the plan.
|
|
371
|
+
- Check that `validate_implementation_plan` produced a `PASS`.
|
|
372
|
+
- `MISSING_TEST_EXECUTION`: Required tests were identified but not executed.
|
|
373
|
+
- `FAILED_EXECUTION_EVIDENCE`: Test results show `FAIL` for a requirement constraint.
|
|
374
|
+
- Reviewer is blocked from closing tasks if execution evidence prevents a `PASS`.
|
|
375
|
+
|
|
376
|
+
|
|
377
|
+
### Phase 6 Experience Verification
|
|
378
|
+
When reviewing an implementation plan or code changes, verify:
|
|
379
|
+
- Experience used by the plan has provenance.
|
|
380
|
+
- Experience is NOT treated as current fact or requirement.
|
|
381
|
+
- Stale or conflicting experience is properly identified.
|
|
382
|
+
- Current repository evidence was checked.
|
|
383
|
+
- User decisions remain authoritative.
|
|
384
|
+
- No unsupported generalized rule was introduced.
|
|
385
|
+
|
|
386
|
+
Flag the following if found:
|
|
387
|
+
- `UNPROVEN_EXPERIENCE`
|
|
388
|
+
- `STALE_EXPERIENCE`
|
|
389
|
+
- `CONFLICTING_EXPERIENCE`
|
|
390
|
+
- `EXPERIENCE_USED_AS_FACT`
|
|
391
|
+
- `EXPERIENCE_USED_AS_REQUIREMENT`
|
|
392
|
+
- `UNSUPPORTED_GENERALIZATION`
|
|
393
|
+
- `MISSING_PROVENANCE`
|
|
394
|
+
- `OUTDATED_REPOSITORY_CONTEXT`
|
|
395
|
+
|
|
396
|
+
|
|
397
|
+
### Phase 7 Regression & Drift Verification
|
|
398
|
+
When reviewing an implementation plan or code changes, verify:
|
|
399
|
+
- Current repository snapshot was considered.
|
|
400
|
+
- Architecture drift and requirement drift were evaluated.
|
|
401
|
+
- Traceability drift and regression tests were considered.
|
|
402
|
+
- Cross-service impacts and security-sensitive changes were evaluated.
|
|
403
|
+
- Risk indicators have evidence and are NOT presented as confirmed defects.
|
|
404
|
+
- No historical lesson became an implicit requirement.
|
|
405
|
+
|
|
406
|
+
Flag the following if found:
|
|
407
|
+
- `HISTORICAL_RISK_NOT_REVIEWED`
|
|
408
|
+
- `REGRESSION_TEST_GAP`
|
|
409
|
+
- `TRACEABILITY_DRIFT`
|
|
410
|
+
- `ARCHITECTURE_DRIFT`
|
|
411
|
+
- `REQUIREMENT_DRIFT`
|
|
412
|
+
- `UNSUPPORTED_RISK_CLAIM`
|
|
413
|
+
- `CROSS_SERVICE_IMPACT_NOT_VALIDATED`
|
|
414
|
+
- `SECURITY_REGRESSION_GAP`
|
|
415
|
+
- `DESTRUCTIVE_OPERATION_GAP`
|
|
416
|
+
|
|
417
|
+
|
|
418
|
+
### Phase 8 UI Design & UX Interaction Verification
|
|
419
|
+
When reviewing frontend implementation plans or code changes, verify:
|
|
420
|
+
- UI changes map back to a confirmed design and requirement.
|
|
421
|
+
- Static HTML generated during planning is NOT copied verbatim into production.
|
|
422
|
+
- Production UI changes were explicitly outlined in the plan.
|
|
423
|
+
- UI patterns used are supported by evidence (discovered from the repository).
|
|
424
|
+
|
|
425
|
+
Flag the following if found:
|
|
426
|
+
- `DESIGN_WITHOUT_REQUIREMENT`
|
|
427
|
+
- `DESIGN_WITHOUT_EVIDENCE`
|
|
428
|
+
- `DESIGN_USED_AS_FACT`
|
|
429
|
+
- `UNCONFIRMED_DESIGN`
|
|
430
|
+
- `STALE_DESIGN`
|
|
431
|
+
- `UI_PATTERN_WITHOUT_SOURCE`
|
|
432
|
+
- `PRODUCTION_UI_CHANGED_WITHOUT_PLAN`
|
|
433
|
+
- `CONFIRMED_DESIGN_WITHOUT_TRACEABILITY`
|
|
434
|
+
- `IMPLEMENTATION_DEVIATES_FROM_CONFIRMED_DESIGN`
|
|
@@ -0,0 +1,46 @@
|
|
|
1
|
+
import fs from 'fs';
|
|
2
|
+
import path from 'path';
|
|
3
|
+
import { buildGraph } from './graph_builder.js';
|
|
4
|
+
import { fileURLToPath } from 'url';
|
|
5
|
+
export function createFixture(ecosystemRoot) {
|
|
6
|
+
const spec = {
|
|
7
|
+
microservices: {
|
|
8
|
+
"sfn-iam-api": { calls_services: ["sfn-events-api"], owns_tables: ["users", "tenant_mapping", "user_groups"], kafka_consumes: [] },
|
|
9
|
+
"sfn-events-api": { calls_services: [], owns_tables: ["events", "audit_records"], kafka_consumes: ["user-deleted"] }
|
|
10
|
+
}
|
|
11
|
+
};
|
|
12
|
+
const dna = {
|
|
13
|
+
domains: [
|
|
14
|
+
{ name: "UserEntity", type: "Entity", repository: "sfn-iam-api", filePath: "src/main/UserEntity.java", violations: [] },
|
|
15
|
+
{ name: "TenantEntity", type: "Entity", repository: "sfn-iam-api", filePath: "src/main/TenantEntity.java", violations: [] },
|
|
16
|
+
{ name: "EventEntity", type: "Entity", repository: "sfn-events-api", filePath: "src/main/EventEntity.java", violations: [] },
|
|
17
|
+
{ name: "RetentionPolicy", type: "Entity", repository: "sfn-events-api", filePath: "src/main/RetentionPolicy.java", violations: [] }
|
|
18
|
+
],
|
|
19
|
+
events: [
|
|
20
|
+
{ className: "UserDeletedListener", repository: "sfn-events-api", filePath: "src/main/UserDeletedListener.java", topicsConsumed: ["user-deleted"], topicsProduced: [] },
|
|
21
|
+
{ className: "RetentionCleanupJob", repository: "sfn-events-api", filePath: "src/main/RetentionCleanupJob.java", topicsConsumed: [], topicsProduced: [] }
|
|
22
|
+
],
|
|
23
|
+
apis: [
|
|
24
|
+
{ method: "DELETE", path: "/api/users/{id}", repository: "sfn-iam-api", controllerName: "UserController", filePath: "src/main/UserController.java" },
|
|
25
|
+
{ method: "PUT", path: "/api/tenants/{id}/retention", repository: "sfn-events-api", controllerName: "RetentionController", filePath: "src/main/RetentionController.java" }
|
|
26
|
+
],
|
|
27
|
+
databases: [
|
|
28
|
+
{ table: "users", repository: "sfn-iam-api", filePath: "src/main/resources/db/migration/V1__users.sql" },
|
|
29
|
+
{ table: "events", repository: "sfn-events-api", filePath: "src/main/resources/db/migration/V1__events.sql" }
|
|
30
|
+
]
|
|
31
|
+
};
|
|
32
|
+
// Write spec
|
|
33
|
+
fs.writeFileSync(path.join(ecosystemRoot, ".secufusion-project-spec.json"), JSON.stringify(spec, null, 2));
|
|
34
|
+
// Generate graph
|
|
35
|
+
const graph = buildGraph(dna, spec);
|
|
36
|
+
dna.graph = graph;
|
|
37
|
+
const secDir = path.join(ecosystemRoot, ".secufusion");
|
|
38
|
+
if (!fs.existsSync(secDir))
|
|
39
|
+
fs.mkdirSync(secDir, { recursive: true });
|
|
40
|
+
fs.writeFileSync(path.join(secDir, "dna.json"), JSON.stringify(dna, null, 2));
|
|
41
|
+
console.log("Mock fixture generated successfully.");
|
|
42
|
+
}
|
|
43
|
+
const isMain = process.argv[1] && fileURLToPath(import.meta.url) === process.argv[1];
|
|
44
|
+
if (isMain) {
|
|
45
|
+
createFixture("c:\\Users\\Yash\\Desktop\\mcp");
|
|
46
|
+
}
|
|
@@ -0,0 +1,135 @@
|
|
|
1
|
+
import { savePlan, validatePlan } from "./implementation_plan.js";
|
|
2
|
+
import { generateTestObligations, aggregateValidation } from "./validation_engine.js";
|
|
3
|
+
import { recordExperience, loadExperience } from "./experience_manager.js";
|
|
4
|
+
import { saveRequirement } from "./requirement_manager.js";
|
|
5
|
+
import { saveTraceability } from "./traceability_engine.js";
|
|
6
|
+
import { performFullExtraction } from "./server.js";
|
|
7
|
+
import fs from 'fs';
|
|
8
|
+
import path from 'path';
|
|
9
|
+
const ecosystemRoot = "c:\\Users\\Yash\\Desktop\\mcp";
|
|
10
|
+
function setupRequirement(id, title, status, decisions = []) {
|
|
11
|
+
const req = {
|
|
12
|
+
id,
|
|
13
|
+
title,
|
|
14
|
+
status,
|
|
15
|
+
raw_request: title,
|
|
16
|
+
clarification_questions: [],
|
|
17
|
+
decisions
|
|
18
|
+
};
|
|
19
|
+
saveRequirement(ecosystemRoot, req);
|
|
20
|
+
}
|
|
21
|
+
function runPhase4Benchmark() {
|
|
22
|
+
console.log("=== RUNNING PHASE 4 VALIDATION BENCHMARK ===");
|
|
23
|
+
// 1. Initial State Setup
|
|
24
|
+
const stats = performFullExtraction(ecosystemRoot);
|
|
25
|
+
const dnaPath = path.join(ecosystemRoot, ".secufusion", "dna.json");
|
|
26
|
+
const dna = JSON.parse(fs.readFileSync(dnaPath, "utf-8"));
|
|
27
|
+
const graph = dna.graph;
|
|
28
|
+
saveTraceability(ecosystemRoot, []);
|
|
29
|
+
// 2. Retention Benchmark (Destructive Operation without Explicit Decision)
|
|
30
|
+
console.log("\\n--- RETENTION: Destructive Operation Guardrails ---");
|
|
31
|
+
setupRequirement("REQ-RET-2", "Implement retention policy cleanup", "CONFIRMED"); // No decisions array has "cleanup"
|
|
32
|
+
let planRet = {
|
|
33
|
+
id: "PLAN-RET",
|
|
34
|
+
requirement_id: "REQ-RET-2",
|
|
35
|
+
status: "DRAFT",
|
|
36
|
+
summary: "Delete old events",
|
|
37
|
+
affected_services: ["sfn-events-api"],
|
|
38
|
+
affected_files: [],
|
|
39
|
+
api_changes: [],
|
|
40
|
+
database_changes: [],
|
|
41
|
+
event_changes: [],
|
|
42
|
+
frontend_changes: [],
|
|
43
|
+
extension_changes: [],
|
|
44
|
+
security_implications: [],
|
|
45
|
+
implementation_steps: ["cleanup events older than 30 days"], // destructive keyword
|
|
46
|
+
test_obligations: ["Test deletion limits"],
|
|
47
|
+
validation_requirements: [],
|
|
48
|
+
traceability_requirements: [],
|
|
49
|
+
risks: [],
|
|
50
|
+
assumptions: [],
|
|
51
|
+
unknowns: [],
|
|
52
|
+
blockers: [],
|
|
53
|
+
created_at: new Date().toISOString(),
|
|
54
|
+
updated_at: new Date().toISOString(),
|
|
55
|
+
snapshot_id: "LATEST"
|
|
56
|
+
};
|
|
57
|
+
savePlan(ecosystemRoot, planRet);
|
|
58
|
+
let errorsRet = validatePlan(ecosystemRoot, planRet);
|
|
59
|
+
console.log("Validation Errors (No Decision):", errorsRet);
|
|
60
|
+
let validationRet = aggregateValidation("CONFIRMED", errorsRet, [], [], [], false);
|
|
61
|
+
console.log("Overall State (No Decision):", validationRet.overall);
|
|
62
|
+
// Now with explicit decision
|
|
63
|
+
setupRequirement("REQ-RET-2", "Implement retention policy cleanup", "CONFIRMED", ["Authorized to cleanup old events"]);
|
|
64
|
+
errorsRet = validatePlan(ecosystemRoot, planRet);
|
|
65
|
+
console.log("Validation Errors (With Decision):", errorsRet);
|
|
66
|
+
validationRet = aggregateValidation("CONFIRMED", errorsRet, [], [], [], false);
|
|
67
|
+
console.log("Overall State (With Decision):", validationRet.overall);
|
|
68
|
+
// 3. User Deletion Benchmark (Test Obligations)
|
|
69
|
+
console.log("\\n--- USER DELETION: Test Intelligence ---");
|
|
70
|
+
setupRequirement("REQ-DEL-2", "User Deletion Feature", "CONFIRMED");
|
|
71
|
+
let planDel = {
|
|
72
|
+
id: "PLAN-DEL",
|
|
73
|
+
requirement_id: "REQ-DEL-2",
|
|
74
|
+
status: "DRAFT",
|
|
75
|
+
summary: "Delete user via API",
|
|
76
|
+
affected_services: ["sfn-iam-api"],
|
|
77
|
+
affected_files: [],
|
|
78
|
+
api_changes: ["DELETE /users/{id}"],
|
|
79
|
+
database_changes: ["DELETE cascade on users table"],
|
|
80
|
+
event_changes: ["Produce user-deleted event"],
|
|
81
|
+
frontend_changes: [],
|
|
82
|
+
extension_changes: [],
|
|
83
|
+
security_implications: ["Requires admin scope"],
|
|
84
|
+
implementation_steps: ["Perform deletion"],
|
|
85
|
+
test_obligations: [],
|
|
86
|
+
validation_requirements: [],
|
|
87
|
+
traceability_requirements: [],
|
|
88
|
+
risks: [],
|
|
89
|
+
assumptions: [],
|
|
90
|
+
unknowns: [],
|
|
91
|
+
blockers: [],
|
|
92
|
+
created_at: new Date().toISOString(),
|
|
93
|
+
updated_at: new Date().toISOString(),
|
|
94
|
+
snapshot_id: "LATEST"
|
|
95
|
+
};
|
|
96
|
+
const generatedTests = generateTestObligations(planDel);
|
|
97
|
+
console.log("Auto-Generated Test Obligations:");
|
|
98
|
+
generatedTests.forEach(t => console.log("-", t));
|
|
99
|
+
// Missing test obligations validation check
|
|
100
|
+
let errorsDel = validatePlan(ecosystemRoot, planDel);
|
|
101
|
+
console.log("Validation Errors (Missing Tests in Plan):", errorsDel);
|
|
102
|
+
let validationDel = aggregateValidation("CONFIRMED", errorsDel, [], [], [], false);
|
|
103
|
+
console.log("Overall State (Missing Tests):", validationDel.overall);
|
|
104
|
+
// 4. Stale Evidence Check
|
|
105
|
+
console.log("\\n--- STALE EVIDENCE ---");
|
|
106
|
+
planDel.test_obligations = generatedTests;
|
|
107
|
+
planDel.created_at = new Date(Date.now() - 100000).toISOString(); // Make it older than last extraction
|
|
108
|
+
savePlan(ecosystemRoot, planDel);
|
|
109
|
+
let errorsStale = validatePlan(ecosystemRoot, planDel);
|
|
110
|
+
console.log("Validation Errors (Stale Plan):", errorsStale);
|
|
111
|
+
let validationStale = aggregateValidation("CONFIRMED", errorsStale, [], [], [], false);
|
|
112
|
+
console.log("Overall State (Stale):", validationStale.overall);
|
|
113
|
+
// 5. Experience Persistence
|
|
114
|
+
console.log("\\n--- EXPERIENCE PERSISTENCE ---");
|
|
115
|
+
const exp = recordExperience(ecosystemRoot, {
|
|
116
|
+
id: "EXP-1",
|
|
117
|
+
category: "ANTI_PATTERN",
|
|
118
|
+
title: "Silent Deletion",
|
|
119
|
+
summary: "Never delete users without producing a user-deleted event.",
|
|
120
|
+
lesson: "Never delete users without producing a user-deleted event.",
|
|
121
|
+
context: "Discovered during User Deletion Phase 4 benchmark",
|
|
122
|
+
related_artifacts: []
|
|
123
|
+
});
|
|
124
|
+
console.log("Recorded Experience ID:", exp.id);
|
|
125
|
+
const loadedExp = loadExperience(ecosystemRoot);
|
|
126
|
+
console.log("Loaded Experiences:", loadedExp.length);
|
|
127
|
+
console.log("Experience Title:", loadedExp[0].title);
|
|
128
|
+
}
|
|
129
|
+
try {
|
|
130
|
+
runPhase4Benchmark();
|
|
131
|
+
}
|
|
132
|
+
catch (e) {
|
|
133
|
+
console.error("Benchmark failed:", e);
|
|
134
|
+
process.exit(1);
|
|
135
|
+
}
|