izanagi-ai 2.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/AGENTS.md +53 -0
- package/CHANGELOG.md +198 -0
- package/LICENSE +21 -0
- package/README.md +96 -0
- package/ROADMAP.md +131 -0
- package/RULES.md +202 -0
- package/SYSTEM.md +201 -0
- package/agents/INDEX.md +40 -0
- package/agents/architect-agent.json +15 -0
- package/agents/bug-hunter-agent.json +14 -0
- package/agents/database-agent.json +14 -0
- package/agents/devops-agent.json +15 -0
- package/agents/docs-agent.json +15 -0
- package/agents/pm-agent.json +14 -0
- package/agents/professor-agent.json +15 -0
- package/agents/security-agent.json +14 -0
- package/agents/senior-engineer-agent.json +16 -0
- package/agents/techlead-agent.json +15 -0
- package/architecture/clean-architecture.md +98 -0
- package/architecture/cqrs-specialist.md +102 -0
- package/architecture/ddd-specialist.md +97 -0
- package/architecture/event-driven-architect.md +79 -0
- package/architecture/hexagonal-architecture.md +96 -0
- package/architecture/microservices-expert.md +83 -0
- package/architecture/monolith-expert.md +67 -0
- package/architecture/repository-pattern.md +73 -0
- package/architecture/unit-of-work.md +92 -0
- package/backend/README.md +29 -0
- package/bin/nexus.js +8 -0
- package/core/compression-engine.md +194 -0
- package/core/context-engine.md +239 -0
- package/core/decision-engine.md +377 -0
- package/core/evolution-engine.md +169 -0
- package/core/planning-engine.md +201 -0
- package/core/quality-gates.md +233 -0
- package/core/reflection-engine.md +182 -0
- package/core/skill-resolver.json +169 -0
- package/core/token-manager.md +152 -0
- package/database/database-engineer.md +244 -0
- package/database/mysql-specialist.md +61 -0
- package/database/postgresql-specialist.md +70 -0
- package/database/redis-specialist.md +73 -0
- package/database/sql-optimizer.md +95 -0
- package/database/sqlserver-specialist.md +69 -0
- package/devops/ci-cd-specialist.md +101 -0
- package/devops/devops-engineer.md +358 -0
- package/devops/docker-expert.md +105 -0
- package/devops/git-expert.md +76 -0
- package/devops/git-flow-specialist.md +71 -0
- package/devops/kubernetes-specialist.md +84 -0
- package/devops/linux-specialist.md +82 -0
- package/devops/windows-specialist.md +43 -0
- package/dist/cli/commands/compile.d.ts +2 -0
- package/dist/cli/commands/compile.d.ts.map +1 -0
- package/dist/cli/commands/compile.js +44 -0
- package/dist/cli/commands/compile.js.map +1 -0
- package/dist/cli/commands/doctor.d.ts +2 -0
- package/dist/cli/commands/doctor.d.ts.map +1 -0
- package/dist/cli/commands/doctor.js +88 -0
- package/dist/cli/commands/doctor.js.map +1 -0
- package/dist/cli/commands/init.d.ts +2 -0
- package/dist/cli/commands/init.d.ts.map +1 -0
- package/dist/cli/commands/init.js +26 -0
- package/dist/cli/commands/init.js.map +1 -0
- package/dist/cli/commands/list.d.ts +2 -0
- package/dist/cli/commands/list.d.ts.map +1 -0
- package/dist/cli/commands/list.js +50 -0
- package/dist/cli/commands/list.js.map +1 -0
- package/dist/cli/commands/run.d.ts +2 -0
- package/dist/cli/commands/run.d.ts.map +1 -0
- package/dist/cli/commands/run.js +49 -0
- package/dist/cli/commands/run.js.map +1 -0
- package/dist/cli/index.d.ts +2 -0
- package/dist/cli/index.d.ts.map +1 -0
- package/dist/cli/index.js +72 -0
- package/dist/cli/index.js.map +1 -0
- package/dist/index.d.ts +3 -0
- package/dist/index.d.ts.map +1 -0
- package/dist/index.js +3 -0
- package/dist/index.js.map +1 -0
- package/dist/installer.d.ts +5 -0
- package/dist/installer.d.ts.map +1 -0
- package/dist/installer.js +86 -0
- package/dist/installer.js.map +1 -0
- package/dist/postinstall.d.ts +2 -0
- package/dist/postinstall.d.ts.map +1 -0
- package/dist/postinstall.js +8 -0
- package/dist/postinstall.js.map +1 -0
- package/frontend/README.md +19 -0
- package/memory/context-recovery.md +49 -0
- package/memory/conversation-summarizer.md +65 -0
- package/memory/long-term-project-memory.md +58 -0
- package/memory/memory-manager.md +299 -0
- package/memory/session-compression.md +182 -0
- package/memory/smart-recall.md +37 -0
- package/optimization/compact-example.md +87 -0
- package/optimization/cost-optimizer.md +53 -0
- package/optimization/prompt-optimizer.md +194 -0
- package/optimization/token-audit.md +177 -0
- package/optimization/token-reducer.md +202 -0
- package/package.json +67 -0
- package/security/owasp-auditor.md +249 -0
- package/security/pentest-reviewer.md +109 -0
- package/security/security-engineer.md +238 -0
- package/skills/INDEX.md +434 -0
- package/skills/accessibility-reviewer.md +63 -0
- package/skills/agentic-coding.md +56 -0
- package/skills/ai-agent/SKILL.md +107 -0
- package/skills/ai-agent-dev/SKILL.md +117 -0
- package/skills/alternative-solution-generator.md +72 -0
- package/skills/architecture-patterns/SKILL.md +113 -0
- package/skills/breaking-change-detector.md +84 -0
- package/skills/bug-hunter.md +234 -0
- package/skills/bug-prevention.md +75 -0
- package/skills/chaos-engineering/SKILL.md +108 -0
- package/skills/clean-code-validator.md +218 -0
- package/skills/cloud-architect/SKILL.md +78 -0
- package/skills/cloud-infra/SKILL.md +99 -0
- package/skills/code-auditor.md +24 -0
- package/skills/complexity-analyzer.md +102 -0
- package/skills/confidence-estimator.md +53 -0
- package/skills/continuous-improvement.md +46 -0
- package/skills/continuous-learning-engine.md +51 -0
- package/skills/cto-advisor.md +66 -0
- package/skills/data-engineer/SKILL.md +76 -0
- package/skills/data-engineering/SKILL.md +82 -0
- package/skills/debug-specialist.md +215 -0
- package/skills/dependency-analyzer.md +78 -0
- package/skills/design-pattern-advisor.md +78 -0
- package/skills/documentation-writer.md +66 -0
- package/skills/dry-kiss-yagni-validator.md +114 -0
- package/skills/economia-tokens/SKILL.md +40 -0
- package/skills/er-diagram-builder.md +85 -0
- package/skills/feature-flags/SKILL.md +95 -0
- package/skills/frontend/SKILL.md +327 -0
- package/skills/frontend-dev/SKILL.md +178 -0
- package/skills/graphql/SKILL.md +104 -0
- package/skills/hallucination-detection.md +49 -0
- package/skills/handoff-sessao/SKILL.md +32 -0
- package/skills/i18n-l10n/SKILL.md +103 -0
- package/skills/iac-terraform/SKILL.md +98 -0
- package/skills/legacy-migration/SKILL.md +91 -0
- package/skills/logging-expert.md +78 -0
- package/skills/mcp-server-dev.md +33 -0
- package/skills/memoria-projeto/SKILL.md +49 -0
- package/skills/mobile-dev/SKILL.md +82 -0
- package/skills/mobile-engineer/SKILL.md +74 -0
- package/skills/monitoring-specialist.md +59 -0
- package/skills/observability-expert.md +60 -0
- package/skills/performance-optimizer.md +239 -0
- package/skills/principal-engineer.md +55 -0
- package/skills/privacy-engineer/SKILL.md +79 -0
- package/skills/professor-modo/SKILL.md +33 -0
- package/skills/project-manager.md +74 -0
- package/skills/prompt-engineering.md +27 -0
- package/skills/qa/SKILL.md +231 -0
- package/skills/qa-engineer/SKILL.md +222 -0
- package/skills/readme-generator.md +42 -0
- package/skills/refactoring-specialist.md +250 -0
- package/skills/release-planner.md +70 -0
- package/skills/requirement-analyzer.md +65 -0
- package/skills/risk-analyzer.md +73 -0
- package/skills/root-cause-analyzer.md +210 -0
- package/skills/scalability-expert.md +73 -0
- package/skills/security-privacy/SKILL.md +98 -0
- package/skills/self-correction.md +54 -0
- package/skills/self-critique.md +42 -0
- package/skills/senior-code-reviewer.md +193 -0
- package/skills/sequence-diagram-builder.md +56 -0
- package/skills/serverless-edge/SKILL.md +103 -0
- package/skills/software-architect.md +356 -0
- package/skills/solid-validator.md +332 -0
- package/skills/sre-reliability/SKILL.md +115 -0
- package/skills/staff-engineer.md +61 -0
- package/skills/task-planner.md +61 -0
- package/skills/tech-lead.md +59 -0
- package/skills/technical-debt-analyzer.md +81 -0
- package/skills/technical-writer.md +39 -0
- package/skills/tradeoff-analyzer.md +79 -0
- package/skills/uml-generator.md +72 -0
- package/skills/ux-reviewer.md +61 -0
- package/skills/wasm/SKILL.md +133 -0
- package/skills/web-perf-engineer/SKILL.md +75 -0
- package/skills/web-perf-seo/SKILL.md +111 -0
- package/skills/websocket-realtime/SKILL.md +107 -0
- package/teaching/adaptive-teaching.md +37 -0
- package/teaching/code-explainer.md +172 -0
- package/teaching/interactive-teaching.md +41 -0
- package/teaching/learning-tracker.md +63 -0
- package/teaching/mentor-mode.md +200 -0
- package/teaching/professor-mode.md +222 -0
- package/testing/e2e-test-engineer.md +68 -0
- package/testing/integration-test-engineer.md +74 -0
- package/testing/mocking-specialist.md +78 -0
- package/testing/unit-test-engineer.md +209 -0
|
@@ -0,0 +1,201 @@
|
|
|
1
|
+
# Core: Planning Engine
|
|
2
|
+
|
|
3
|
+
> Version 1.0.0
|
|
4
|
+
> Priority: Critical
|
|
5
|
+
> Dependencies: Context Engine, Decision Engine
|
|
6
|
+
> Compatibility: ">=1.0.0"
|
|
7
|
+
|
|
8
|
+
---
|
|
9
|
+
|
|
10
|
+
## Identity
|
|
11
|
+
|
|
12
|
+
The Planning Engine is activated before any code is written. It breaks down tasks into ordered steps, estimates effort, identifies dependencies, and produces a clear implementation plan. It ensures the agent never writes code without knowing exactly what needs to be built and in what order.
|
|
13
|
+
|
|
14
|
+
---
|
|
15
|
+
|
|
16
|
+
## Goals
|
|
17
|
+
|
|
18
|
+
- Break any task into steps no larger than 1 file or 50 lines.
|
|
19
|
+
- Estimate effort per step (minutes or complexity points).
|
|
20
|
+
- Identify hidden dependencies between steps.
|
|
21
|
+
- Produce a plan that another agent or human can follow.
|
|
22
|
+
- Never start coding without a plan.
|
|
23
|
+
|
|
24
|
+
---
|
|
25
|
+
|
|
26
|
+
## Triggers
|
|
27
|
+
|
|
28
|
+
| Condition | Action |
|
|
29
|
+
|-----------|--------|
|
|
30
|
+
| `task == "new_project"` | Full project plan with phases |
|
|
31
|
+
| `task == "new_feature"` | Feature breakdown into tasks |
|
|
32
|
+
| `task == "refactor"` | Migration plan with safe steps |
|
|
33
|
+
| `task == "bug"` | Investigation → Fix → Verify plan |
|
|
34
|
+
| Any complex task | Decompose into sub-tasks |
|
|
35
|
+
|
|
36
|
+
---
|
|
37
|
+
|
|
38
|
+
## Workflow
|
|
39
|
+
|
|
40
|
+
```
|
|
41
|
+
1. Analyze requirements
|
|
42
|
+
↓
|
|
43
|
+
2. Decompose into atomic steps
|
|
44
|
+
↓
|
|
45
|
+
3. Identify dependencies between steps
|
|
46
|
+
↓
|
|
47
|
+
4. Estimate effort per step
|
|
48
|
+
↓
|
|
49
|
+
5. Order steps (topological sort)
|
|
50
|
+
↓
|
|
51
|
+
6. Identify risks per step
|
|
52
|
+
↓
|
|
53
|
+
7. Generate implementation plan
|
|
54
|
+
↓
|
|
55
|
+
8. Validate plan against requirements
|
|
56
|
+
↓
|
|
57
|
+
9. Pass to next skill in chain
|
|
58
|
+
```
|
|
59
|
+
|
|
60
|
+
---
|
|
61
|
+
|
|
62
|
+
## Plan Format
|
|
63
|
+
|
|
64
|
+
```yaml
|
|
65
|
+
plan:
|
|
66
|
+
title: "Create Login API"
|
|
67
|
+
total_effort: "45 min"
|
|
68
|
+
|
|
69
|
+
steps:
|
|
70
|
+
- id: 1
|
|
71
|
+
name: "Create users table migration"
|
|
72
|
+
file: "database/migrations/xxxx_create_users_table.php"
|
|
73
|
+
effort: "5 min"
|
|
74
|
+
depends_on: []
|
|
75
|
+
risks: []
|
|
76
|
+
|
|
77
|
+
- id: 2
|
|
78
|
+
name: "Create User model"
|
|
79
|
+
file: "app/Models/User.php"
|
|
80
|
+
effort: "5 min"
|
|
81
|
+
depends_on: [1]
|
|
82
|
+
risks: []
|
|
83
|
+
|
|
84
|
+
- id: 3
|
|
85
|
+
name: "Create AuthController"
|
|
86
|
+
file: "app/Http/Controllers/AuthController.php"
|
|
87
|
+
effort: "10 min"
|
|
88
|
+
depends_on: [2]
|
|
89
|
+
risks: []
|
|
90
|
+
|
|
91
|
+
- id: 4
|
|
92
|
+
name: "Create LoginRequest"
|
|
93
|
+
file: "app/Http/Requests/LoginRequest.php"
|
|
94
|
+
effort: "5 min"
|
|
95
|
+
depends_on: []
|
|
96
|
+
risks: []
|
|
97
|
+
|
|
98
|
+
- id: 5
|
|
99
|
+
name: "Create AuthService"
|
|
100
|
+
file: "app/Services/AuthService.php"
|
|
101
|
+
effort: "10 min"
|
|
102
|
+
depends_on: [2, 4]
|
|
103
|
+
risks: ["Token storage strategy needed"]
|
|
104
|
+
|
|
105
|
+
- id: 6
|
|
106
|
+
name: "Write authentication tests"
|
|
107
|
+
file: "tests/Feature/AuthTest.php"
|
|
108
|
+
effort: "10 min"
|
|
109
|
+
depends_on: [3, 5]
|
|
110
|
+
risks: []
|
|
111
|
+
```
|
|
112
|
+
|
|
113
|
+
---
|
|
114
|
+
|
|
115
|
+
## Effort Estimation
|
|
116
|
+
|
|
117
|
+
```
|
|
118
|
+
complexity = estimate_complexity(task)
|
|
119
|
+
|
|
120
|
+
effort = base_time * complexity * uncertainty_factor
|
|
121
|
+
|
|
122
|
+
base_time:
|
|
123
|
+
- new file: 5 min
|
|
124
|
+
- edit file: 3 min
|
|
125
|
+
- config: 2 min
|
|
126
|
+
- test: 5 min
|
|
127
|
+
|
|
128
|
+
complexity:
|
|
129
|
+
- simple (CRUD, no logic): 1x
|
|
130
|
+
- medium (business logic): 2x
|
|
131
|
+
- complex (algorithms, integrations): 3x
|
|
132
|
+
|
|
133
|
+
uncertainty_factor:
|
|
134
|
+
- known tech: 1x
|
|
135
|
+
- unfamiliar tech: 2x
|
|
136
|
+
- new pattern: 1.5x
|
|
137
|
+
```
|
|
138
|
+
|
|
139
|
+
---
|
|
140
|
+
|
|
141
|
+
## Dependency Resolution
|
|
142
|
+
|
|
143
|
+
Steps are ordered using topological sort:
|
|
144
|
+
|
|
145
|
+
```
|
|
146
|
+
function order_steps(steps):
|
|
147
|
+
graph = build_dependency_graph(steps)
|
|
148
|
+
ordered = topological_sort(graph)
|
|
149
|
+
return ordered
|
|
150
|
+
```
|
|
151
|
+
|
|
152
|
+
If a circular dependency is detected:
|
|
153
|
+
|
|
154
|
+
```
|
|
155
|
+
→ flag for user review
|
|
156
|
+
→ ask for clarification
|
|
157
|
+
→ suggest breaking the cycle
|
|
158
|
+
```
|
|
159
|
+
|
|
160
|
+
---
|
|
161
|
+
|
|
162
|
+
## Rules
|
|
163
|
+
|
|
164
|
+
### Always
|
|
165
|
+
|
|
166
|
+
- ✅ Plan before coding. Always.
|
|
167
|
+
- ✅ Break down into atomic steps (1 file or 50 lines max).
|
|
168
|
+
- ✅ Identify dependencies between steps.
|
|
169
|
+
- ✅ Estimate effort per step.
|
|
170
|
+
- ✅ Validate plan against all requirements.
|
|
171
|
+
|
|
172
|
+
### Never
|
|
173
|
+
|
|
174
|
+
- ❌ Start coding with "let me just write the code".
|
|
175
|
+
- ❌ Combine multiple concerns in one step.
|
|
176
|
+
- ❌ Skip effort estimation.
|
|
177
|
+
- ❌ Ignore step dependencies.
|
|
178
|
+
|
|
179
|
+
---
|
|
180
|
+
|
|
181
|
+
## Metrics
|
|
182
|
+
|
|
183
|
+
| Metric | Target | How to Measure |
|
|
184
|
+
|--------|--------|---------------|
|
|
185
|
+
| Plans generated | 100% of complex tasks | Plans / tasks requiring plans |
|
|
186
|
+
| Step atomicity | ≤ 50 lines per step | Average lines per step |
|
|
187
|
+
| Dependency accuracy | ≥ 90% | Dependencies correct / total |
|
|
188
|
+
| Plan adherence | ≥ 80% | Steps followed in order |
|
|
189
|
+
|
|
190
|
+
---
|
|
191
|
+
|
|
192
|
+
## Changelog
|
|
193
|
+
|
|
194
|
+
### 1.0.0 (2026-07-17)
|
|
195
|
+
|
|
196
|
+
- Initial release
|
|
197
|
+
- Topological sort for dependency resolution
|
|
198
|
+
- Effort estimation with complexity factors
|
|
199
|
+
- Atomic step decomposition (max 1 file / 50 lines)
|
|
200
|
+
- Circular dependency detection
|
|
201
|
+
- YAML plan format for machine readability
|
|
@@ -0,0 +1,233 @@
|
|
|
1
|
+
# Core: Quality Gates
|
|
2
|
+
|
|
3
|
+
> Version 1.0.0
|
|
4
|
+
> Priority: Critical
|
|
5
|
+
> Dependencies: Token Manager, Security Scanner
|
|
6
|
+
> Compatibility: ">=1.0.0"
|
|
7
|
+
|
|
8
|
+
---
|
|
9
|
+
|
|
10
|
+
## Identity
|
|
11
|
+
|
|
12
|
+
Quality Gates are the final validation layer before any output is delivered to the user. Every response passes through 5 gates: Security, Style, Clarity, Conciseness, and Completeness. If any gate fails, the output is rejected and reworked.
|
|
13
|
+
|
|
14
|
+
---
|
|
15
|
+
|
|
16
|
+
## Gates Overview
|
|
17
|
+
|
|
18
|
+
```
|
|
19
|
+
┌─────────────────────────────────────────────────┐
|
|
20
|
+
│ QUALITY GATES │
|
|
21
|
+
│ │
|
|
22
|
+
│ Input → Gate 1 → Gate 2 → Gate 3 → Gate 4 → │
|
|
23
|
+
│ → Gate 5 → ✅ Pass → Output │
|
|
24
|
+
│ → ❌ Fail → Rework │
|
|
25
|
+
│ │
|
|
26
|
+
│ Gate 1: Security (fatal if fail) │
|
|
27
|
+
│ Gate 2: Style │
|
|
28
|
+
│ Gate 3: Clarity │
|
|
29
|
+
│ Gate 4: Conciseness │
|
|
30
|
+
│ Gate 5: Completeness │
|
|
31
|
+
└─────────────────────────────────────────────────┘
|
|
32
|
+
```
|
|
33
|
+
|
|
34
|
+
---
|
|
35
|
+
|
|
36
|
+
## Gate 1: Security
|
|
37
|
+
|
|
38
|
+
**Fatal gate** — if this fails, output is blocked immediately.
|
|
39
|
+
|
|
40
|
+
### Checks
|
|
41
|
+
|
|
42
|
+
- [ ] No hardcoded credentials, API keys, tokens, or secrets
|
|
43
|
+
- [ ] No SQL injection vectors (raw queries without parameterization)
|
|
44
|
+
- [ ] No XSS vectors (untrusted output without escaping)
|
|
45
|
+
- [ ] No CSRF vulnerabilities
|
|
46
|
+
- [ ] No SSRF vulnerabilities
|
|
47
|
+
- [ ] No insecure deserialization
|
|
48
|
+
- [ ] No hardcoded IPs or internal paths
|
|
49
|
+
- [ ] No commented-out security code
|
|
50
|
+
- [ ] Passwords are hashed (bcrypt/argon2), not encrypted
|
|
51
|
+
- [ ] HTTPS is enforced (not HTTP)
|
|
52
|
+
- [ ] File uploads are validated (type, size, path traversal)
|
|
53
|
+
|
|
54
|
+
### Failure Action
|
|
55
|
+
|
|
56
|
+
```yaml
|
|
57
|
+
fail:
|
|
58
|
+
action: "Block output"
|
|
59
|
+
message: "❌ SECURITY GATE FAILED: [specific issue]"
|
|
60
|
+
notify: "Reflection Engine"
|
|
61
|
+
log: true
|
|
62
|
+
```
|
|
63
|
+
|
|
64
|
+
---
|
|
65
|
+
|
|
66
|
+
## Gate 2: Style
|
|
67
|
+
|
|
68
|
+
**Non-fatal** — can pass with warnings.
|
|
69
|
+
|
|
70
|
+
### Checks
|
|
71
|
+
|
|
72
|
+
- [ ] Follows project naming conventions
|
|
73
|
+
- [ ] Consistent indentation (spaces vs tabs)
|
|
74
|
+
- [ ] No commented-out code blocks
|
|
75
|
+
- [ ] No debug statements (dd(), var_dump, console.log)
|
|
76
|
+
- [ ] No trailing whitespace
|
|
77
|
+
- [ ] Imports are organized
|
|
78
|
+
- [ ] Line length ≤ 120 characters
|
|
79
|
+
- [ ] PHP: follows PSR-12
|
|
80
|
+
- [ ] JS/TS: follows project ESLint config
|
|
81
|
+
- [ ] Python: follows PEP 8
|
|
82
|
+
|
|
83
|
+
### Failure Action
|
|
84
|
+
|
|
85
|
+
```yaml
|
|
86
|
+
warn:
|
|
87
|
+
action: "Pass with warnings"
|
|
88
|
+
message: "⚠️ STYLE GATE: [specific issues] — fixed automatically"
|
|
89
|
+
fix: true
|
|
90
|
+
```
|
|
91
|
+
|
|
92
|
+
---
|
|
93
|
+
|
|
94
|
+
## Gate 3: Clarity
|
|
95
|
+
|
|
96
|
+
**Non-fatal** — subjective, uses heuristics.
|
|
97
|
+
|
|
98
|
+
### Checks
|
|
99
|
+
|
|
100
|
+
- [ ] Response directly answers the user's question
|
|
101
|
+
- [ ] No undefined acronyms or jargon
|
|
102
|
+
- [ ] Code examples have context (filename, purpose)
|
|
103
|
+
- [ ] Steps are in logical order
|
|
104
|
+
- [ ] Error messages are human-readable
|
|
105
|
+
- [ ] Variable names are meaningful
|
|
106
|
+
- [ ] Complex logic has brief explanation
|
|
107
|
+
|
|
108
|
+
### Heuristics
|
|
109
|
+
|
|
110
|
+
```
|
|
111
|
+
clarity_score = 0
|
|
112
|
+
if answer_directly_matches_question:
|
|
113
|
+
clarity_score += 3
|
|
114
|
+
if jargon_defined:
|
|
115
|
+
clarity_score += 1
|
|
116
|
+
if code_has_context:
|
|
117
|
+
clarity_score += 1
|
|
118
|
+
if steps_are_ordered:
|
|
119
|
+
clarity_score += 1
|
|
120
|
+
if no_undefined_acronyms:
|
|
121
|
+
clarity_score += 1
|
|
122
|
+
|
|
123
|
+
if clarity_score < 5:
|
|
124
|
+
fail("Response lacks clarity — rewrite")
|
|
125
|
+
```
|
|
126
|
+
|
|
127
|
+
---
|
|
128
|
+
|
|
129
|
+
## Gate 4: Conciseness
|
|
130
|
+
|
|
131
|
+
**Non-fatal** — measured against token budget.
|
|
132
|
+
|
|
133
|
+
### Checks
|
|
134
|
+
|
|
135
|
+
- [ ] No repetition of the same information
|
|
136
|
+
- [ ] No filler phrases ("I would recommend", "It is important to note")
|
|
137
|
+
- [ ] No verbose markdown (excessive headers, dividers)
|
|
138
|
+
- [ ] Token count ≤ budget for the skill chain
|
|
139
|
+
- [ ] No redundant code examples (same pattern twice)
|
|
140
|
+
|
|
141
|
+
### Heuristics
|
|
142
|
+
|
|
143
|
+
```
|
|
144
|
+
repetition_penalty = count_repeated_phrases(output)
|
|
145
|
+
filler_penalty = count_filler_phrases(output)
|
|
146
|
+
total_penalty = repetition_penalty + filler_penalty
|
|
147
|
+
|
|
148
|
+
if total_penalty > 5:
|
|
149
|
+
fail("Response too verbose — compress")
|
|
150
|
+
```
|
|
151
|
+
|
|
152
|
+
---
|
|
153
|
+
|
|
154
|
+
## Gate 5: Completeness
|
|
155
|
+
|
|
156
|
+
**Non-fatal** — checks that nothing is missing.
|
|
157
|
+
|
|
158
|
+
### Checks
|
|
159
|
+
|
|
160
|
+
- [ ] All user questions are answered
|
|
161
|
+
- [ ] All requirements from the task are addressed
|
|
162
|
+
- [ ] No "TODO" or "FIXME" left in code
|
|
163
|
+
- [ ] Error handling is included
|
|
164
|
+
- [ ] Edge cases are mentioned
|
|
165
|
+
- [ ] Dependencies are declared
|
|
166
|
+
- [ ] Installation/setup instructions are provided (if applicable)
|
|
167
|
+
|
|
168
|
+
### Heuristics
|
|
169
|
+
|
|
170
|
+
```
|
|
171
|
+
answered = count_user_questions_answered(output)
|
|
172
|
+
expected = count_user_questions(input)
|
|
173
|
+
completeness_ratio = answered / expected
|
|
174
|
+
|
|
175
|
+
if completeness_ratio < 1.0:
|
|
176
|
+
fail("Missing answers to [n] question(s)")
|
|
177
|
+
```
|
|
178
|
+
|
|
179
|
+
---
|
|
180
|
+
|
|
181
|
+
## Combined Result
|
|
182
|
+
|
|
183
|
+
```yaml
|
|
184
|
+
quality_result:
|
|
185
|
+
security: "✅ PASS"
|
|
186
|
+
style: "⚠️ PASS (warnings: trailing whitespace)"
|
|
187
|
+
clarity: "✅ PASS"
|
|
188
|
+
conciseness: "✅ PASS (3% under budget)"
|
|
189
|
+
completeness: "✅ PASS"
|
|
190
|
+
overall: "✅ PASS"
|
|
191
|
+
token_usage: "1,234 / 2,048"
|
|
192
|
+
```
|
|
193
|
+
|
|
194
|
+
---
|
|
195
|
+
|
|
196
|
+
## Rules
|
|
197
|
+
|
|
198
|
+
### Always
|
|
199
|
+
|
|
200
|
+
- ✅ Run all 5 gates on every output.
|
|
201
|
+
- ✅ Block output on security gate failure.
|
|
202
|
+
- ✅ Log gate results for Reflection Engine.
|
|
203
|
+
- ✅ Fix style issues automatically when possible.
|
|
204
|
+
|
|
205
|
+
### Never
|
|
206
|
+
|
|
207
|
+
- ❌ Skip gates for speed.
|
|
208
|
+
- ❌ Deliver output that fails any gate.
|
|
209
|
+
- ❌ Modify security gate rules without user approval.
|
|
210
|
+
|
|
211
|
+
---
|
|
212
|
+
|
|
213
|
+
## Metrics
|
|
214
|
+
|
|
215
|
+
| Metric | Target | How to Measure |
|
|
216
|
+
|--------|--------|---------------|
|
|
217
|
+
| Security pass rate | 100% | Outputs passing / total outputs |
|
|
218
|
+
| Overall pass rate | ≥ 95% | Outputs passing all gates |
|
|
219
|
+
| Auto-fix rate | ≥ 80% | Issues auto-fixed / total issues |
|
|
220
|
+
| Gate latency | < 50ms | Time to run all gates |
|
|
221
|
+
|
|
222
|
+
---
|
|
223
|
+
|
|
224
|
+
## Changelog
|
|
225
|
+
|
|
226
|
+
### 1.0.0 (2026-07-17)
|
|
227
|
+
|
|
228
|
+
- Initial release
|
|
229
|
+
- 5 quality gates with distinct checklists
|
|
230
|
+
- Security gate is fatal — blocks output on failure
|
|
231
|
+
- Style gate auto-fixes common issues
|
|
232
|
+
- Clarity, conciseness, completeness use heuristic scoring
|
|
233
|
+
- Full logging for Reflection Engine integration
|
|
@@ -0,0 +1,182 @@
|
|
|
1
|
+
# Core: Reflection Engine
|
|
2
|
+
|
|
3
|
+
> Version 1.0.0
|
|
4
|
+
> Priority: High
|
|
5
|
+
> Dependencies: Memory Manager, Evolution Engine
|
|
6
|
+
> Compatibility: ">=1.0.0"
|
|
7
|
+
|
|
8
|
+
---
|
|
9
|
+
|
|
10
|
+
## Identity
|
|
11
|
+
|
|
12
|
+
The Reflection Engine activates after every task execution. It reviews what was done, identifies what went well and what can improve, logs patterns and mistakes, and feeds data into the Evolution Engine to update skills over time.
|
|
13
|
+
|
|
14
|
+
---
|
|
15
|
+
|
|
16
|
+
## Goals
|
|
17
|
+
|
|
18
|
+
- Review every task after delivery.
|
|
19
|
+
- Identify at least 1 improvement opportunity per task.
|
|
20
|
+
- Log all mistakes to prevent recurrence.
|
|
21
|
+
- Feed structured data to the Evolution Engine.
|
|
22
|
+
- Build a personal improvement history for the agent.
|
|
23
|
+
|
|
24
|
+
---
|
|
25
|
+
|
|
26
|
+
## Workflow
|
|
27
|
+
|
|
28
|
+
```
|
|
29
|
+
1. Task completes
|
|
30
|
+
↓
|
|
31
|
+
2. Gather execution data
|
|
32
|
+
- Input, classification, chain used
|
|
33
|
+
- Token usage per skill
|
|
34
|
+
- Quality gate results
|
|
35
|
+
↓
|
|
36
|
+
3. Run self-review questions
|
|
37
|
+
- What went well?
|
|
38
|
+
- What could be better?
|
|
39
|
+
- Was the classification correct?
|
|
40
|
+
- Was context sufficient?
|
|
41
|
+
- Were quality gates passed?
|
|
42
|
+
↓
|
|
43
|
+
4. Check for patterns
|
|
44
|
+
- Recurring mistakes
|
|
45
|
+
- Recurring clarifications
|
|
46
|
+
- Token waste patterns
|
|
47
|
+
↓
|
|
48
|
+
5. Generate reflection output
|
|
49
|
+
↓
|
|
50
|
+
6. Pass to Memory Manager (log)
|
|
51
|
+
↓
|
|
52
|
+
7. Pass to Evolution Engine (if pattern found)
|
|
53
|
+
```
|
|
54
|
+
|
|
55
|
+
---
|
|
56
|
+
|
|
57
|
+
## Self-Review Questions
|
|
58
|
+
|
|
59
|
+
### Quality
|
|
60
|
+
|
|
61
|
+
- Was the output complete?
|
|
62
|
+
- Did it answer all user questions?
|
|
63
|
+
- Was it concise?
|
|
64
|
+
- Was it correct?
|
|
65
|
+
|
|
66
|
+
### Process
|
|
67
|
+
|
|
68
|
+
- Was the classification accurate?
|
|
69
|
+
- Was the skill chain optimal?
|
|
70
|
+
- Was context sufficient?
|
|
71
|
+
- Was budget respected?
|
|
72
|
+
|
|
73
|
+
### Security
|
|
74
|
+
|
|
75
|
+
- Were any secrets exposed?
|
|
76
|
+
- Were any vulnerabilities introduced?
|
|
77
|
+
- Was security validated?
|
|
78
|
+
|
|
79
|
+
### Teaching
|
|
80
|
+
|
|
81
|
+
- Did the user learn something?
|
|
82
|
+
- Did I explain the reasoning?
|
|
83
|
+
- Was the explanation appropriate for the user level?
|
|
84
|
+
|
|
85
|
+
---
|
|
86
|
+
|
|
87
|
+
## Reflection Log Format
|
|
88
|
+
|
|
89
|
+
```yaml
|
|
90
|
+
reflection:
|
|
91
|
+
task_id: "task-20260717-001"
|
|
92
|
+
timestamp: "2026-07-17T11:30:00Z"
|
|
93
|
+
classification: "new_feature"
|
|
94
|
+
chain: ["architect", "security", "backend", "testing"]
|
|
95
|
+
|
|
96
|
+
scores:
|
|
97
|
+
completeness: 4.5 / 5
|
|
98
|
+
conciseness: 4.0 / 5
|
|
99
|
+
correctness: 5.0 / 5
|
|
100
|
+
teaching: 3.5 / 5
|
|
101
|
+
|
|
102
|
+
wins:
|
|
103
|
+
- "Architecture was clean and well-explained"
|
|
104
|
+
- "Security considerations included early"
|
|
105
|
+
|
|
106
|
+
improvements:
|
|
107
|
+
- "Could have included more testing examples"
|
|
108
|
+
- "Token budget exceeded by 12%"
|
|
109
|
+
|
|
110
|
+
patterns:
|
|
111
|
+
- type: "token_waste"
|
|
112
|
+
detail: "Verbose code examples in security section"
|
|
113
|
+
frequency: 3
|
|
114
|
+
|
|
115
|
+
evolution_triggers:
|
|
116
|
+
- skill: "backend-engineer"
|
|
117
|
+
suggestion: "Add token budget optimization for code examples"
|
|
118
|
+
```
|
|
119
|
+
|
|
120
|
+
---
|
|
121
|
+
|
|
122
|
+
## Pattern Detection
|
|
123
|
+
|
|
124
|
+
The engine maintains a rolling window of the last 50 reflections to detect:
|
|
125
|
+
|
|
126
|
+
| Pattern | Detection | Action |
|
|
127
|
+
|---------|-----------|--------|
|
|
128
|
+
| Token waste | >5 compressions in 10 tasks | Update Token Manager budget |
|
|
129
|
+
| Misclassification | >3 clarifications for same category | Update Decision Engine keywords |
|
|
130
|
+
| Repeated mistakes | >2 same improvement items | Flag for Evolution Engine |
|
|
131
|
+
| User confusion | >2 follow-up clarifications | Add teaching moment to output |
|
|
132
|
+
| Security gaps | Any security item missed | Immediate skill update |
|
|
133
|
+
|
|
134
|
+
---
|
|
135
|
+
|
|
136
|
+
## Metrics
|
|
137
|
+
|
|
138
|
+
| Metric | Target | How to Measure |
|
|
139
|
+
|--------|--------|---------------|
|
|
140
|
+
| Tasks reflected | 100% | Reflections / total tasks |
|
|
141
|
+
| Improvements per task | ≥ 1 | Count improvement items |
|
|
142
|
+
| Pattern detection rate | ≥ 80% | Patterns found / patterns present |
|
|
143
|
+
| Evolution triggers | ≥ 1 per 10 tasks | Count triggers generated |
|
|
144
|
+
|
|
145
|
+
---
|
|
146
|
+
|
|
147
|
+
## Reflection Quality Gates
|
|
148
|
+
|
|
149
|
+
1. ✅ Reflection is generated for every task.
|
|
150
|
+
2. ✅ At least 1 improvement identified.
|
|
151
|
+
3. ✅ Scores are honest (not all 5/5).
|
|
152
|
+
4. ✅ Patterns are detected (if present).
|
|
153
|
+
|
|
154
|
+
---
|
|
155
|
+
|
|
156
|
+
## Memory Hooks
|
|
157
|
+
|
|
158
|
+
```yaml
|
|
159
|
+
on_reflect:
|
|
160
|
+
- save: reflection_log
|
|
161
|
+
- save: pattern_window (rolling 50)
|
|
162
|
+
|
|
163
|
+
on_pattern_detected:
|
|
164
|
+
- notify: Evolution Engine
|
|
165
|
+
- save: pattern_alert
|
|
166
|
+
|
|
167
|
+
on_complete:
|
|
168
|
+
- compress: reflection_log if > budget
|
|
169
|
+
```
|
|
170
|
+
|
|
171
|
+
---
|
|
172
|
+
|
|
173
|
+
## Changelog
|
|
174
|
+
|
|
175
|
+
### 1.0.0 (2026-07-17)
|
|
176
|
+
|
|
177
|
+
- Initial release
|
|
178
|
+
- Post-task self-review with scoring
|
|
179
|
+
- Pattern detection over rolling 50-task window
|
|
180
|
+
- Structured reflection log in YAML format
|
|
181
|
+
- Integration with Evolution Engine
|
|
182
|
+
- 5 dimensions of review: quality, process, security, teaching, efficiency
|