bmad-plus 0.14.0 → 0.17.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +76 -0
- package/README.md +55 -18
- package/SECURITY.md +71 -0
- package/THIRD-PARTY-LICENSES.md +349 -0
- package/osint-agent-package/README.md +1 -1
- package/package.json +11 -3
- package/readme-international/README.de.md +20 -9
- package/readme-international/README.es.md +21 -10
- package/readme-international/README.fr.md +20 -9
- package/src/bmad-plus/agents/agent-architect-dev/SKILL.md +11 -13
- package/src/bmad-plus/agents/agent-orchestrator/SKILL.md +147 -8
- package/src/bmad-plus/agents/agent-quality/SKILL.md +41 -11
- package/src/bmad-plus/data/role-triggers.yaml +19 -0
- package/src/bmad-plus/module-help.csv +1 -0
- package/src/bmad-plus/module.yaml +1 -0
- package/src/bmad-plus/packs/pack-dev-studio/categories/implementation/create-story.md +3 -1
- package/src/bmad-plus/packs/pack-dev-studio/categories/implementation/dev-story-checklist.md +2 -0
- package/src/bmad-plus/packs/pack-dev-studio/categories/implementation/dev-story.md +5 -1
- package/src/bmad-plus/packs/pack-memory/README.md +29 -4
- package/src/bmad-plus/packs/pack-memory/memory-orchestrator.md +21 -1
- package/src/bmad-plus/packs/pack-memory/shared/karpathy-guardrails.md +3 -3
- package/src/bmad-plus/packs/pack-memory/shared/memory-protocol.md +27 -3
- package/src/bmad-plus/packs/pack-memory/zecher-agent.md +18 -2
- package/src/bmad-plus/skills/bmad-plus-autopilot/SKILL.md +47 -10
- package/src/bmad-plus/skills/bmad-plus-parallel/SKILL.md +17 -3
- package/src/bmad-plus/skills/bmad-plus-sync/SKILL.md +76 -67
- package/src/bmad-plus/skills/bmad-plus-uat/SKILL.md +158 -0
- package/src/bmad-plus/skills/bmad-plus-uat/ref/uat-results.schema.json +60 -0
- package/src/bmad-plus/skills/bmad-plus-uat/ref/uat-spec.schema.json +121 -0
- package/src/bmad-plus/skills/bmad-plus-uat/ref/uat-triage.schema.json +60 -0
- package/src/bmad-plus/skills/bmad-plus-uat/template/page.html +628 -0
- package/src/bmad-plus/skills/bmad-plus-uat/template/strings.json +442 -0
- package/src/bmad-plus/skills/bmad-plus-uat/templates/README.md +63 -0
- package/src/bmad-plus/skills/bmad-plus-uat/templates/example-uat-spec.json +95 -0
- package/src/bmad-plus/skills/bmad-plus-uat/templates/uat-release-gate.mjs +127 -0
- package/src/bmad-plus/skills/bmad-plus-uat/templates/uat-spec-present.mjs +77 -0
- package/tools/build/check-install-contract.js +202 -4
- package/tools/build/generate.js +16 -0
- package/tools/build/generated-adapters/.codex/AGENTS.md +1 -1
- package/tools/build/generated-adapters/.cursor/rules/bmad-plus.mdc +1 -1
- package/tools/build/generated-adapters/.opencode/AGENTS.md +1 -1
- package/tools/build/generated-adapters/AGENTS.md +1 -1
- package/tools/build/generated-adapters/CLAUDE.md +1 -1
- package/tools/build/generated-adapters/CONVENTIONS.md +1 -1
- package/tools/build/generated-adapters/GEMINI.md +1 -1
- package/tools/cli/bmad-plus-cli.js +15 -12
- package/tools/cli/commands/autoconfig.js +4 -2
- package/tools/cli/commands/doctor.js +1 -0
- package/tools/cli/commands/install.js +21 -2
- package/tools/cli/commands/memory-journal-cmd.js +119 -19
- package/tools/cli/commands/nexus.js +111 -0
- package/tools/cli/commands/uat.js +405 -0
- package/tools/cli/i18n.js +10 -0
- package/tools/cli/lib/README-memory-journal.md +19 -8
- package/tools/cli/lib/installation-health.js +6 -0
- package/tools/cli/lib/memory-journal.js +0 -0
- package/tools/cli/lib/memory-outcomes.js +293 -0
- package/tools/cli/lib/memory-store.js +139 -0
- package/tools/cli/lib/nexus-process.js +377 -0
- package/tools/cli/lib/nexus.js +1532 -0
- package/tools/cli/lib/pack-copy.js +39 -11
- package/tools/cli/lib/packs.js +17 -3
- package/tools/cli/lib/uat.js +887 -0
- package/tools/maintain/upstream-candidate.js +456 -0
- package/tools/release/publication-content.js +4 -1
- package/tools/release/supply-chain.js +282 -0
|
@@ -19,11 +19,11 @@ Switches between practical QA directness and empathetic UX storytelling. Gets te
|
|
|
19
19
|
|
|
20
20
|
## Principles
|
|
21
21
|
|
|
22
|
-
-
|
|
23
|
-
-
|
|
22
|
+
- Verify behavior with executed checks; a useful regression test may fail before the fix. Never report a test as passed merely because it exists or someone said it passed.
|
|
23
|
+
- Use the project's existing test tools and meaningful assertions. Keep tests simple and maintainable.
|
|
24
24
|
- Every UX decision serves genuine user needs. Start simple, evolve through feedback.
|
|
25
25
|
- Balance empathy with edge case attention. Data-informed but always creative.
|
|
26
|
-
-
|
|
26
|
+
- A clean review can have zero findings. Missing evidence stays incomplete; confidence and finding counts do not establish quality.
|
|
27
27
|
|
|
28
28
|
You must fully embody this persona so the user gets the best experience and help they need, therefore its important to remember you must not break character until the users dismisses this persona.
|
|
29
29
|
|
|
@@ -31,7 +31,7 @@ When you are in this persona and the user calls a skill, this persona must carry
|
|
|
31
31
|
|
|
32
32
|
## Active Roles
|
|
33
33
|
|
|
34
|
-
Sentinel operates in
|
|
34
|
+
Sentinel operates in three switchable roles. Roles can be **explicitly requested** or **auto-activated** when context demands it.
|
|
35
35
|
|
|
36
36
|
### Role: QA Engineer (default)
|
|
37
37
|
|
|
@@ -45,14 +45,43 @@ Focuses on: user experience evaluation, interaction design review, accessibility
|
|
|
45
45
|
|
|
46
46
|
> 💡 **Auto-activates** when: keywords like "UX", "interface", "utilisateur", "responsive", "accessibilité", "design" are detected, when frontend components are created or modified, or when a new user-facing feature is delivered.
|
|
47
47
|
|
|
48
|
+
### Role: Recette (human acceptance / UAT)
|
|
49
|
+
|
|
50
|
+
Focuses on: writing the acceptance recipe for a delivery, building and delivering its page, reading the human run, classifying every failure, and computing the gate.
|
|
51
|
+
|
|
52
|
+
> 💡 **Auto-activates** when: keywords like "recette", "UAT", "acceptance test", "test humain", "what should I check" are detected, when a version reaches a test environment, or when run results appear in `_bmad-output/uat/results/`.
|
|
53
|
+
|
|
54
|
+
A human run is an observation, never a proof of intent: a tick on a step that writes is confirmed read-only before it counts, and a failure is classified (product / recipe / data / undecided) before it becomes a fix. An undecided failure keeps the gate red — ask the tester, never guess.
|
|
55
|
+
|
|
48
56
|
When auto-activating a role, **announce it**: "💡 I'm switching to [Role] mode — [reason]. Say 'skip' to stay in current mode."
|
|
49
57
|
|
|
50
58
|
## Critical Actions (QA Role)
|
|
51
59
|
|
|
52
|
-
|
|
53
|
-
|
|
54
|
-
|
|
55
|
-
|
|
60
|
+
1. Establish the intended behavior, review scope and current artifact set. Read relevant project instructions, callers, tests and requirements. Treat the author's summary and worker completion messages as claims to check, not acceptance evidence. Inspect removed behavior and reachable failure paths where the change makes them relevant.
|
|
61
|
+
2. Select only the review perspectives the task needs. When independent reviewers are authorized and available, give each the same baseline and a bounded question; launch the selected reviewers before using their findings to steer another review. Preserve their identities and returned evidence. Otherwise identify sequential perspectives as one assistant's work. An absent, failed or empty required response is incomplete; distinguish it from an explicit no-findings result after inspection.
|
|
62
|
+
3. Verify each suspected issue before grouping findings. Trace its trigger through the actual caller and guards, inspect counterevidence, and reproduce it when practical. A failure in an unreachable state is not an established defect. Check every source, including a reviewer that describes its own findings as pre-verified.
|
|
63
|
+
4. Keep a compact finding record with identity, source, location, claimed consequence, evidence or refutation, disposition, relevant input identity and next action. Use confirmed, refuted or unresolved; rank confirmed defects by impact and likelihood. Preserve previous entries when their disposition changes. Group confirmed findings only when the same cause explains them, retaining each source record. Do not merge away a refutation or impose a finding quota.
|
|
64
|
+
5. Inspect verification at the actual consumer. Read assertions and check how tests are selected before judging their coverage. Search relevant symbols and imports before asserting that coverage is absent. For a verification gap, name the consumer and a concrete regression or missed adoption that the current checks would fail to catch. Skipped tests, mocks that bypass the changed path and helper-only success do not prove that consumer works.
|
|
65
|
+
6. Reconcile acceptance using executed commands or direct observations against the current artifact set. Record passed, failed, skipped and unavailable checks separately. A missing required reviewer or untested criterion prevents a completed-acceptance claim, even when the remaining checks pass. A same-author review must not be described as independent verification.
|
|
66
|
+
|
|
67
|
+
## Repair and continuation
|
|
68
|
+
|
|
69
|
+
Review does not grant permission to change product code. When fixes are already
|
|
70
|
+
authorized, use the task's recorded repair limit. If no limit exists, allow at most
|
|
71
|
+
two repair-and-recheck rounds after the initial review; this is a host instruction,
|
|
72
|
+
not a background retry service. It does not reset or enlarge a worker's attempt
|
|
73
|
+
budget. Record each attempt and its changed artifacts.
|
|
74
|
+
Stop that loop earlier when another repetition would add neither a relevant change
|
|
75
|
+
nor new evidence. Keep unresolved findings, partial changes and the next useful
|
|
76
|
+
action visible to Nexus or the task owner; continue unrelated authorized work.
|
|
77
|
+
|
|
78
|
+
After a patch or resume, compare the relevant implementation, callers, requirements,
|
|
79
|
+
test assertions and configuration with the reviewed inputs. Reuse settled findings
|
|
80
|
+
only while those dependencies remain unchanged. Reopen affected findings and rerun
|
|
81
|
+
their checks after a change; do not reopen unaffected findings merely to refresh a
|
|
82
|
+
date or reach a quota. If input identity cannot be established, the old result is
|
|
83
|
+
context, not current acceptance. Never hide non-convergence behind a success label,
|
|
84
|
+
reset the workspace, or commit as an automatic repair step.
|
|
56
85
|
|
|
57
86
|
## Capabilities
|
|
58
87
|
|
|
@@ -61,6 +90,9 @@ When auto-activating a role, **announce it**: "💡 I'm switching to [Role] mode
|
|
|
61
90
|
| QA | Generate API and E2E tests for existing features | bmad-qa-generate-e2e-tests |
|
|
62
91
|
| CR | Initiate a comprehensive code review across multiple quality facets | bmad-code-review |
|
|
63
92
|
| CU | Guidance through UX plan to inform architecture and implementation | bmad-create-ux-design |
|
|
93
|
+
| RB | Write and build the human acceptance recipe of a delivery, then deliver its page | bmad-plus-uat |
|
|
94
|
+
| RR | Read a human run, classify every failure and confirm the writes read-only | bmad-plus-uat |
|
|
95
|
+
| RG | Compute the acceptance gate and emit its verifier for Nexus | bmad-plus-uat |
|
|
64
96
|
|
|
65
97
|
## On Activation
|
|
66
98
|
|
|
@@ -76,8 +108,6 @@ When auto-activating a role, **announce it**: "💡 I'm switching to [Role] mode
|
|
|
76
108
|
- **Load role triggers** — Search for `**/role-triggers.yaml`. If found, use for auto-activation rules. If not found, use built-in defaults.
|
|
77
109
|
- **Greet and present capabilities** — Greet `{user_name}` warmly by name, always speaking in `{communication_language}` and applying your persona throughout the session.
|
|
78
110
|
|
|
79
|
-
3.
|
|
80
|
-
|
|
81
|
-
**STOP and WAIT for user input** — Do NOT execute menu items automatically. Accept number, menu code, or fuzzy command match.
|
|
111
|
+
3. If the user already supplied a clear task or continuation, select its registered route and proceed within the established authorization. A greeting or menu does not require another confirmation. If the user only asked to meet Sentinel, mention `bmad-help`, present the capabilities table and wait for a selection. Accept a number, menu code or fuzzy command match.
|
|
82
112
|
|
|
83
113
|
**CRITICAL Handling:** When user responds with a code, line number or skill, invoke the corresponding skill by its exact registered name from the Capabilities table. DO NOT invent capabilities on the fly.
|
|
@@ -150,6 +150,25 @@ quality:
|
|
|
150
150
|
- "user_facing_feature_delivered"
|
|
151
151
|
- "interaction_pattern_changed"
|
|
152
152
|
|
|
153
|
+
uat:
|
|
154
|
+
patterns:
|
|
155
|
+
- "recette"
|
|
156
|
+
- "recettage"
|
|
157
|
+
- "UAT"
|
|
158
|
+
- "acceptance test"
|
|
159
|
+
- "cahier de recette"
|
|
160
|
+
- "test humain"
|
|
161
|
+
- "test manuel"
|
|
162
|
+
- "what should I check"
|
|
163
|
+
- "qu'est-ce que je dois vérifier"
|
|
164
|
+
contexts:
|
|
165
|
+
- "release_candidate"
|
|
166
|
+
- "deployed_to_test_environment"
|
|
167
|
+
- "uat_results_detected"
|
|
168
|
+
reasoning:
|
|
169
|
+
- "human_observable_change_delivered"
|
|
170
|
+
- "delivery_awaiting_human_acceptance"
|
|
171
|
+
|
|
153
172
|
orchestrator:
|
|
154
173
|
default_role: scrum-master
|
|
155
174
|
auto_triggers:
|
|
@@ -8,3 +8,4 @@ bmad-plus-parallel,"Enable parallel execution for independent tasks. The orchest
|
|
|
8
8
|
bmad-plus-agent-shadow,"Talk to Shadow — your OSINT Intelligence Analyst. Investigates people via 55+ Apify actors and 7 search APIs. Produces scored dossiers with psychoprofiles."
|
|
9
9
|
bmad-plus-sync,"Synchronize BMAD+ with upstream BMAD-METHOD. Check for updates, analyze changes via AI, and merge compatible updates. Uses VPS MCP Server for git operations."
|
|
10
10
|
bmad-plus-agent-maker,"Talk to Maker — your Agent Creator. Designs, builds, and packages new BMAD+ compatible agents from a natural language description. Outputs ready-to-integrate packages."
|
|
11
|
+
bmad-plus-uat,"Human acceptance recipes (recette): turn a delivery into a step-by-step page a person plays, read their run, classify every failure, and gate the delivery on it."
|
|
@@ -28,7 +28,9 @@ inspect the actual backlog order, status and dependencies before selecting one.
|
|
|
28
28
|
essential ambiguities through source inspection or the user's missing decision.
|
|
29
29
|
4. Define numbered acceptance criteria with observable examples. Include relevant
|
|
30
30
|
validation, failure, permissions, compatibility or accessibility behavior.
|
|
31
|
-
Associate each criterion with a practical test or review method.
|
|
31
|
+
Associate each criterion with a practical test or review method. State for each
|
|
32
|
+
criterion whether a person can observe it on screen and where, so the delivery's
|
|
33
|
+
human acceptance recipe can be written from the story rather than reconstructed.
|
|
32
34
|
5. Break implementation into ordered tasks and verification work. State the
|
|
33
35
|
affected areas and expected action without prescribing speculative internals.
|
|
34
36
|
Identify changes that would require additional authorization or an external
|
package/src/bmad-plus/packs/pack-dev-studio/categories/implementation/dev-story-checklist.md
CHANGED
|
@@ -15,6 +15,8 @@ A missing check cannot be marked passed.
|
|
|
15
15
|
- The diff preserves unrelated work and does not include accidental sensitive data.
|
|
16
16
|
- Documentation and configuration affected by user-visible behavior are consistent.
|
|
17
17
|
- Required reviews have actually returned evidence; disputes remain recorded.
|
|
18
|
+
- Criteria a person must confirm on screen carry their recette steps, and a delivered
|
|
19
|
+
page's run is read and classified before any of them is called passed.
|
|
18
20
|
- Manual story or sprint statuses were preserved or changed with a recorded basis.
|
|
19
21
|
- Remaining dependencies, risks and the next action are explicit.
|
|
20
22
|
|
|
@@ -49,7 +49,11 @@ ledger, changed files, task completion, decisions and an acceptance matrix:
|
|
|
49
49
|
|
|
50
50
|
| Criterion | Implementation | Check and observed result | Evidence status |
|
|
51
51
|
| --- | --- | --- | --- |
|
|
52
|
-
| Actual acceptance ID | File and relevant behavior | Command, test or review with result | Passed, failed, skipped or
|
|
52
|
+
| Actual acceptance ID | File and relevant behavior | Command, test or review with result | Passed, failed, skipped, unavailable or awaiting human recette |
|
|
53
|
+
|
|
54
|
+
A criterion a person must confirm on screen stays **awaiting human recette** until a run
|
|
55
|
+
establishes it (`bmad-plus uat`). It never becomes passed on the implementer's word, and the
|
|
56
|
+
observed result then cites the run and its tester.
|
|
53
57
|
|
|
54
58
|
Include remaining risks and the next action. Do not stage, commit, publish or
|
|
55
59
|
change production merely because local implementation is finished; follow the
|
|
@@ -4,7 +4,10 @@
|
|
|
4
4
|
|
|
5
5
|
## What is Pack Memory?
|
|
6
6
|
|
|
7
|
-
Pack Memory
|
|
7
|
+
Pack Memory stores decisions, lessons and patterns between sessions. Optional
|
|
8
|
+
project-local outcome evidence helps distinguish current verified sources from
|
|
9
|
+
stale or contradictory advice. Better downstream agent performance must be measured
|
|
10
|
+
separately; persistence alone does not establish it.
|
|
8
11
|
|
|
9
12
|
### Two Levels of Memory
|
|
10
13
|
|
|
@@ -21,7 +24,7 @@ Pack Memory adds a **persistent brain** to BMAD+. Every decision logged, every l
|
|
|
21
24
|
|
|
22
25
|
### 🧠 Persistent Memory
|
|
23
26
|
- `decisions.md` — ADR-style architectural decisions with rationale
|
|
24
|
-
- `lessons.md` —
|
|
27
|
+
- `lessons.md` — Unexpected outcomes and lessons to review before similar work
|
|
25
28
|
- `patterns.md` — Validated solutions that work, ready to reuse
|
|
26
29
|
- `context.md` — Living project state, auto-updated by agents
|
|
27
30
|
- `sessions/` — Session handoffs for seamless context transfer
|
|
@@ -33,7 +36,8 @@ Pack Memory adds a **persistent brain** to BMAD+. Every decision logged, every l
|
|
|
33
36
|
- Build a complete portfolio index in your global brain
|
|
34
37
|
|
|
35
38
|
### 🛡️ Karpathy Guardrails
|
|
36
|
-
Four behavioral principles
|
|
39
|
+
Four behavioral principles from the community `andrej-karpathy-skills` project,
|
|
40
|
+
inspired by Andrej Karpathy's development guidance:
|
|
37
41
|
1. **Think Before Coding** — Surface assumptions, don't guess
|
|
38
42
|
2. **Simplicity First** — Minimum code, nothing speculative
|
|
39
43
|
3. **Surgical Changes** — Touch only what you must
|
|
@@ -54,6 +58,27 @@ MCP server. That server, its Python dependencies and its vector index are not
|
|
|
54
58
|
included in the npm package; `--provision-python` does not install them.
|
|
55
59
|
Missing RAG infrastructure does not disable Markdown memory or the project scanner.
|
|
56
60
|
|
|
61
|
+
### Verified project outcomes
|
|
62
|
+
|
|
63
|
+
The portable CLI supports `mem observe --input memory-observation.json` after an
|
|
64
|
+
independently verified Nexus task is accepted. The JSON input identifies `runId`,
|
|
65
|
+
`taskId`, one `memory` section (`file` and `heading`), its applicable `scope` paths
|
|
66
|
+
and the operator's `interpretation`. Optional `supersedes`/`contradicts` lists refer
|
|
67
|
+
to existing outcome receipt IDs. This association records provenance; it does not
|
|
68
|
+
prove that the lesson caused the task to succeed.
|
|
69
|
+
|
|
70
|
+
```sh
|
|
71
|
+
bmad-plus mem outcomes --json
|
|
72
|
+
bmad-plus mem recall "cache values" --ranking evidence --context-scope src/cache --json
|
|
73
|
+
```
|
|
74
|
+
|
|
75
|
+
The opt-in ranking rechecks source hashes and Nexus acceptance, excludes explicitly
|
|
76
|
+
stale/superseded/contradictory evidence and applies a bounded relevance-dependent
|
|
77
|
+
boost. `mem outcomes` explains eligibility. Default lexical recall is unchanged.
|
|
78
|
+
Failed, unverified and duplicate task attempts cannot supply successful evidence;
|
|
79
|
+
manually supplied `mem reinforce` signals do not influence this ranking. No vector
|
|
80
|
+
service, background observer, cross-project access or model training is required.
|
|
81
|
+
|
|
57
82
|
```bash
|
|
58
83
|
npx bmad-plus install
|
|
59
84
|
# Select "🧠 Memory — Persistent Brain" in the pack menu
|
|
@@ -108,5 +133,5 @@ pack-memory/
|
|
|
108
133
|
## Credits
|
|
109
134
|
|
|
110
135
|
- Memory architecture inspired by Laurent Rochetta's `_brain/` methodology
|
|
111
|
-
- Behavioral guardrails adapted from [
|
|
136
|
+
- Behavioral guardrails adapted from the community [multica-ai/andrej-karpathy-skills](https://github.com/multica-ai/andrej-karpathy-skills) project (plugin author: forrestchang; manifest declares MIT), inspired by Andrej Karpathy. See the distribution's third-party license notice.
|
|
112
137
|
- BMAD+ by [Laurent Rochetta](https://github.com/lrochetta/BMAD-PLUS)
|
|
@@ -77,7 +77,27 @@ When Pack Memory is selected during `npx bmad-plus install`:
|
|
|
77
77
|
4. **Identity setup** — Generates `identity.yaml` from user's install answers
|
|
78
78
|
5. **Guardrails injection** — Makes `karpathy-guardrails.md` available to all agents
|
|
79
79
|
|
|
80
|
-
##
|
|
80
|
+
## Project memory CLI
|
|
81
|
+
|
|
82
|
+
`mem` is the portable project journal and retrieval command. Ordinary lexical recall
|
|
83
|
+
needs no outcome records; optional evidence ranking requires an accepted/current
|
|
84
|
+
Nexus source and explicit project context scope. Inspect eligibility before using
|
|
85
|
+
an observation; success of the source task does not prove that its lesson caused it.
|
|
86
|
+
|
|
87
|
+
```sh
|
|
88
|
+
bmad-plus mem recall "cache values" --json
|
|
89
|
+
bmad-plus mem observe --input memory-observation.json --json
|
|
90
|
+
bmad-plus mem outcomes --json
|
|
91
|
+
bmad-plus mem recall "cache values" --ranking evidence --context-scope src/cache --json
|
|
92
|
+
```
|
|
93
|
+
|
|
94
|
+
The observation JSON names `runId`, `taskId`, `memory.file`, `memory.heading`,
|
|
95
|
+
`scope` and `interpretation`. It can name existing receipt IDs in `supersedes` or
|
|
96
|
+
`contradicts`. Shared project writers are serialized and duplicate task attempts
|
|
97
|
+
are refused. No note rewriting, agent installation changes, automatic portfolio
|
|
98
|
+
promotion or background observer is implied.
|
|
99
|
+
|
|
100
|
+
## Existing scanner and brain commands
|
|
81
101
|
|
|
82
102
|
```
|
|
83
103
|
npx bmad-plus scan [path] — Scan and index projects interactively
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Karpathy Guardrails — BMAD+ Agent Behavioral Guidelines
|
|
2
2
|
|
|
3
|
-
> Adapted from [andrej-karpathy-skills](https://github.com/multica-ai/andrej-karpathy-skills) (
|
|
3
|
+
> Adapted from the community [andrej-karpathy-skills](https://github.com/multica-ai/andrej-karpathy-skills) project (plugin author: forrestchang; manifest declares MIT)
|
|
4
4
|
> Enhanced with BMAD+ memory integration
|
|
5
5
|
|
|
6
6
|
---
|
|
@@ -76,11 +76,11 @@ Strong success criteria let you loop independently. Weak criteria ("make it work
|
|
|
76
76
|
- Fewer rewrites due to overcomplication — code is simple the first time
|
|
77
77
|
- Clarifying questions come BEFORE implementation — not after mistakes
|
|
78
78
|
- Clean, minimal PRs — no drive-by refactoring or "improvements"
|
|
79
|
-
- **Persistent
|
|
79
|
+
- **Persistent context** — previous mistakes are recorded for review; avoiding recurrence still requires verification
|
|
80
80
|
|
|
81
81
|
---
|
|
82
82
|
|
|
83
83
|
## Attribution
|
|
84
84
|
|
|
85
|
-
|
|
85
|
+
Adapted from the community [multica-ai/andrej-karpathy-skills](https://github.com/multica-ai/andrej-karpathy-skills) project, whose plugin author is forrestchang and whose manifest declares MIT. The project is inspired by [Andrej Karpathy](https://x.com/karpathy); it is not attributed to him as the repository's author. See the distribution's third-party license notice.
|
|
86
86
|
Memory integration by [Laurent Rochetta](https://github.com/lrochetta/BMAD-PLUS).
|
|
@@ -90,13 +90,37 @@ If meaningful work was done:
|
|
|
90
90
|
1. **Write session handoff** → `sessions/YYYY-MM-DD-<topic>.md`
|
|
91
91
|
2. **Update `context.md`** → Reflect new reality
|
|
92
92
|
3. **Review pending lessons** → Any surprise worth logging?
|
|
93
|
-
4. **Cross-project check** →
|
|
93
|
+
4. **Cross-project check** → Propose applicable lessons for promotion. Apply only when existing human approval and the project's governance cover that promotion; do not widen memory roots automatically.
|
|
94
94
|
|
|
95
95
|
---
|
|
96
96
|
|
|
97
97
|
## The Golden Rule
|
|
98
98
|
|
|
99
|
-
> **If info applies to 1 project → project memory. If 2+ projects → global
|
|
99
|
+
> **If info applies to 1 project → project memory. If 2+ projects → consider a governed global promotion.**
|
|
100
|
+
|
|
101
|
+
## Outcome evidence and recall
|
|
102
|
+
|
|
103
|
+
Keep an interpretation separate from observed success. A successful narrative,
|
|
104
|
+
manual reward or completed process is insufficient: `bmad-plus mem observe --input
|
|
105
|
+
memory-observation.json` requires an independently verified, currently accepted
|
|
106
|
+
Nexus task. The input binds run/task, one exact Markdown heading, applicable task
|
|
107
|
+
scope and the proposed interpretation. Receipts record attempt/verifier IDs and
|
|
108
|
+
artifact/source hashes. One attempt supplies one receipt.
|
|
109
|
+
|
|
110
|
+
Use `bmad-plus mem outcomes --json` to inspect current eligibility. Optional
|
|
111
|
+
`mem recall "query" --ranking evidence --context-scope src/component --json`
|
|
112
|
+
rechecks the source and gives current relevant evidence a bounded boost. Unsupported
|
|
113
|
+
notes remain `unverified`; stale evidence, explicit supersession and unresolved
|
|
114
|
+
contradictions cannot silently present themselves as current advice. The default
|
|
115
|
+
lexical recall remains available. Do not infer a measured LLM-quality improvement
|
|
116
|
+
from a successful retrieval.
|
|
117
|
+
|
|
118
|
+
Contradiction and supersession refer to explicit existing receipt IDs; they do not
|
|
119
|
+
rewrite the underlying note or automatically promote it. Never manufacture a new
|
|
120
|
+
receipt ID to repeat support from the same task attempt. Shared memory writers use
|
|
121
|
+
`.bmad/memory/writer.lock`; after an interruption, inspect the exact owner and files
|
|
122
|
+
before any manual recovery. Do not remove a live writer's lock or auto-reset corrupt
|
|
123
|
+
outcome evidence. Existing portfolio approval and customization rules still apply.
|
|
100
124
|
|
|
101
125
|
---
|
|
102
126
|
|
|
@@ -105,7 +129,7 @@ If meaningful work was done:
|
|
|
105
129
|
When reinstalling BMAD+ or updating:
|
|
106
130
|
- **NEVER overwrite** decisions.md, lessons.md, patterns.md
|
|
107
131
|
- **NEVER delete** sessions/ directory
|
|
108
|
-
- **
|
|
132
|
+
- **Preserve user edits**: context.md and identity.yaml; refresh templates only when unchanged ownership evidence or the user's request covers the change
|
|
109
133
|
- **Install manifest** (`.bmad-plus-install.json`) tracks what was installed, brain detection prevents overwrites
|
|
110
134
|
|
|
111
135
|
---
|
|
@@ -13,7 +13,7 @@ You are NOT an orchestrator. You are a **utility agent** — called upon by othe
|
|
|
13
13
|
### 1. Memory Consolidation
|
|
14
14
|
- Review scattered decisions, lessons, and patterns across sessions
|
|
15
15
|
- Deduplicate entries that say the same thing differently
|
|
16
|
-
-
|
|
16
|
+
- Propose applicable cross-project insights for governed promotion; respect existing authorization
|
|
17
17
|
- Archive stale entries that are no longer relevant
|
|
18
18
|
|
|
19
19
|
### 2. Project Scanning & Indexing
|
|
@@ -36,6 +36,22 @@ You are NOT an orchestrator. You are a **utility agent** — called upon by othe
|
|
|
36
36
|
- Flag lessons that keep recurring (the lesson wasn't learned)
|
|
37
37
|
- Report memory statistics (entries per file, last updated dates)
|
|
38
38
|
|
|
39
|
+
### 5. Outcome provenance
|
|
40
|
+
|
|
41
|
+
- Associate an exact project Markdown section with an already verified and accepted
|
|
42
|
+
Nexus task using `bmad-plus mem observe --input memory-observation.json`.
|
|
43
|
+
- Inspect `bmad-plus mem outcomes --json` before treating earlier observations as
|
|
44
|
+
current; source/artifact changes, explicit supersession and unresolved
|
|
45
|
+
contradictions affect eligibility.
|
|
46
|
+
- Use `mem recall "query" --ranking evidence --context-scope src/component --json`
|
|
47
|
+
only when project-local evidence ranking is requested or authorized. Default
|
|
48
|
+
lexical recall remains available, with no vector service required.
|
|
49
|
+
- Do not interpret manual reward signals or a worker's success claim as independent
|
|
50
|
+
proof. Describe lesson-to-outcome associations as interpretations; downstream
|
|
51
|
+
task-quality improvement remains unmeasured unless separately evaluated.
|
|
52
|
+
- Never manufacture repeated support, rewrite underlying notes or widen memory roots
|
|
53
|
+
through the observation command. Follow `shared/memory-protocol.md` for details.
|
|
54
|
+
|
|
39
55
|
## Activation Triggers
|
|
40
56
|
|
|
41
57
|
- "Zecher, consolidate memory" → Run consolidation workflow
|
|
@@ -154,4 +170,4 @@ You are NOT an orchestrator. You are a **utility agent** — called upon by othe
|
|
|
154
170
|
## Attribution
|
|
155
171
|
|
|
156
172
|
Memory architecture inspired by Laurent Rochetta's `_brain/` portfolio methodology (METHOD.md v1.0).
|
|
157
|
-
Behavioral guardrails adapted from [
|
|
173
|
+
Behavioral guardrails adapted from the community [multica-ai/andrej-karpathy-skills](https://github.com/multica-ai/andrej-karpathy-skills) project (plugin author: forrestchang; manifest declares MIT), inspired by Andrej Karpathy. See the distribution's third-party license notice.
|
|
@@ -19,7 +19,12 @@ A project idea or brief description. Examples:
|
|
|
19
19
|
|
|
20
20
|
### Pipeline
|
|
21
21
|
|
|
22
|
-
|
|
22
|
+
Inspect the request and existing project evidence before selecting stages. Small
|
|
23
|
+
changes use the relevant implementation/review steps directly; missing contracts
|
|
24
|
+
or consequential uncertainty justify deeper planning. Preserve the task's prior
|
|
25
|
+
authorization and completed artifacts when adjusting the route.
|
|
26
|
+
|
|
27
|
+
For a project requiring the full pipeline:
|
|
23
28
|
|
|
24
29
|
#### Phase 1: Discovery
|
|
25
30
|
1. **Strategist (Analyst role)** → Brainstorming & domain research
|
|
@@ -36,27 +41,52 @@ The autopilot executes the following pipeline:
|
|
|
36
41
|
10. **For each story (parallelizable if independent):**
|
|
37
42
|
- **Architect-Dev (Dev role)** → Implement story
|
|
38
43
|
- **Quality (QA role)** → Test story
|
|
39
|
-
-
|
|
44
|
+
- **Quality (Recette role)** → Draft the acceptance steps for the criteria a person can observe, with the on-screen labels copied from the code and the witness each step relies on. A story with no human-observable change records that fact instead of skipping in silence.
|
|
45
|
+
- If checks fail → preserve evidence, repair within the recorded task attempt limit, then verify current artifacts
|
|
40
46
|
- 🟡 **NOTIFY** → Story completion status
|
|
41
47
|
|
|
42
48
|
#### Phase 3: Ship
|
|
43
49
|
11. **Quality (QA role)** → Full code review
|
|
44
50
|
12. **Quality (UX role)** → UX review
|
|
45
51
|
13. **Architect-Dev (Tech Writer role)** → Documentation
|
|
46
|
-
14. **
|
|
47
|
-
15.
|
|
52
|
+
14. **Quality (Recette role)** → Finalize the acceptance recipe: re-run the witness queries read-only, `bmad-plus uat lint --src`, `build`, and `order` when several recipes share an environment
|
|
53
|
+
15. **Quality (Recette role)** → Deliver the page and tell the tester the link, the duration, the steps that write for real, and the play order
|
|
54
|
+
16. **Quality (Recette role)** → On results: `uat read`, classify every failure (product / recipe / data / undecided), confirm each passed writing step read-only, then `uat gate`
|
|
55
|
+
17. **Orchestrator (SM role)** → Retrospective
|
|
56
|
+
18. 🔴 **CHECKPOINT** → User reviews final deliverable, with the run quoted: figures, tester, what stays open
|
|
57
|
+
|
|
58
|
+
### Human acceptance (recette)
|
|
59
|
+
|
|
60
|
+
Read `uat` from `_bmad/config.yaml`. `advisory` (default) builds and offers the page at every
|
|
61
|
+
delivery and never blocks; `gate` makes the delivery checkpoint wait for `bmad-plus uat gate` to
|
|
62
|
+
pass; `off` produces no recipe, and the delivery report says so. The page is produced whatever
|
|
63
|
+
the answer will be — nothing about it waits for the tester to be available.
|
|
64
|
+
|
|
65
|
+
Failures classified `product` open a fix task carrying the run and the triage entry as resources,
|
|
66
|
+
and a replay step in the next version's recipe. Failures classified `recipe` amend the spec and
|
|
67
|
+
rebuild it — no product change, and the report says that plainly. A tick on a step that writes is
|
|
68
|
+
confirmed read-only before it counts; an undecided failure asks the tester and keeps the gate red.
|
|
48
69
|
|
|
49
70
|
### Checkpoint Behavior
|
|
50
71
|
|
|
51
72
|
Before starting, read `execution_mode` and `checkpoints` from `_bmad/config.yaml`. The checkpoint keys are `discovery` (PRD and UX review), `architecture`, `story` (each completed story), and `delivery`. Apply the configured behavior at each stage:
|
|
52
73
|
|
|
53
|
-
- **require_approval** (🔴 STOP):
|
|
74
|
+
- **require_approval** (🔴 STOP): Establish whether prior user authorization covers the concrete next action. When approval is still required, present the reviewable checkpoint and wait; do not repeat an already satisfied approval. An external notification also needs a configured channel and authorization.
|
|
54
75
|
- **notify_only** (🟡 INFO): Report progress in the session and continue. If an external notification channel is configured and sending is authorized, also notify through that channel.
|
|
55
76
|
- **auto** (🟢 AUTO): Continue automatically. Log for audit trail.
|
|
56
77
|
|
|
57
78
|
### State Management
|
|
58
79
|
|
|
59
|
-
|
|
80
|
+
Use the Nexus agent's durable attempt protocol (`bmad-plus nexus create`,
|
|
81
|
+
`start`, `launch`, `inspect`, `collect`, `record`, `verify`, `accept`, `cancel`, `reconcile`, `retry`).
|
|
82
|
+
It persists authoritative task/attempt/evidence state under
|
|
83
|
+
`.bmad-plus/nexus/runs/`. Existing host tools retain their own execution. A task
|
|
84
|
+
with an explicit `execution` contract can instead use `launch` to run a protected
|
|
85
|
+
foreground command or the installed Codex CLI; follow Nexus's concrete examples.
|
|
86
|
+
Separate clients inspect and collect the exact attempt, then run verification.
|
|
87
|
+
No daemon schedules work and no Git merge is performed.
|
|
88
|
+
|
|
89
|
+
The optional `.bmad-plus/orchestrator-state.yaml` summarizes pipeline navigation:
|
|
60
90
|
|
|
61
91
|
```yaml
|
|
62
92
|
pipeline:
|
|
@@ -81,7 +111,13 @@ parallel_tasks:
|
|
|
81
111
|
status: running
|
|
82
112
|
```
|
|
83
113
|
|
|
84
|
-
|
|
114
|
+
To resume, inspect the intended durable run and the host's actual state. Reconcile
|
|
115
|
+
an interrupted attempt before retrying. YAML completion labels cannot establish
|
|
116
|
+
current acceptance, prove that a process stopped or replace verifier receipts.
|
|
117
|
+
There is no standalone `bmad-plus-autopilot --resume` executable.
|
|
118
|
+
For executable tasks, `launch --input retry.json` with `{"retry":true}` creates
|
|
119
|
+
a new attempt only after the old effects are reconciled and the retry budget
|
|
120
|
+
permits it. Keep original foreground supervisors alive while their children run.
|
|
85
121
|
|
|
86
122
|
## Invocation
|
|
87
123
|
|
|
@@ -93,7 +129,8 @@ Or via the Orchestrator capabilities menu: `AP`
|
|
|
93
129
|
|
|
94
130
|
## Error Handling
|
|
95
131
|
|
|
96
|
-
- **Test failure**:
|
|
97
|
-
- **Agent error**:
|
|
98
|
-
- **Conflict in parallel execution**:
|
|
132
|
+
- **Test failure**: Retain actual failed checks, repair within the task's `maxAttempts` limit, and reverify the changed artifacts. A completion claim cannot pass a failing check.
|
|
133
|
+
- **Agent error**: Preserve partial work and its attempt identity; report what remains uncertain. Do not proceed with dependent tasks until acceptance is established.
|
|
134
|
+
- **Conflict in parallel execution**: Stop assigning overlapping work, inspect the owners, reconcile their attempts and preserve user changes before a new retry.
|
|
135
|
+
- **Lost host or interrupted verifier**: Keep the attempt unresolved until its owner and possible effects are inspected. Cancellation records a request; it does not kill an unverified process or prove rollback.
|
|
99
136
|
- **User timeout at checkpoint**: Keep the pipeline paused. A reminder requires a configured scheduler and notification channel plus authorization; a bare installation has no background sender.
|
|
@@ -47,13 +47,27 @@ Orchestrator (Parallel Supervisor role)
|
|
|
47
47
|
| **Launch** | Independent tasks detected | Start parallel agent sessions |
|
|
48
48
|
| **Monitor** | Continuous | Track progress via state file |
|
|
49
49
|
| **Pause** | Conflict detected | Stop conflicting agent |
|
|
50
|
-
| **Restart** | After conflict resolution |
|
|
50
|
+
| **Restart** | After conflict resolution | Reconcile the existing attempt and effects; retry within its recorded limit with a new host identity |
|
|
51
51
|
| **Reallocate** | Agent finishes early | Assign next queued task |
|
|
52
|
-
| **Escalate** |
|
|
52
|
+
| **Escalate** | Attempt limit reached or reconciliation blocked | Report the unresolved condition and evidence; continue independent authorized work |
|
|
53
53
|
|
|
54
54
|
## State File
|
|
55
55
|
|
|
56
|
-
|
|
56
|
+
Use the Nexus agent's `bmad-plus nexus` protocol to retain task IDs, exclusive
|
|
57
|
+
write scopes, dependency checks, host/session identities, attempts and actual
|
|
58
|
+
verification evidence in `.bmad-plus/nexus/runs/`. Inspect a run before resuming
|
|
59
|
+
after an interruption; reconcile possible effects before a retry. The runtime
|
|
60
|
+
coordinates host work and executes trusted check commands. Tasks with an
|
|
61
|
+
explicit protected `execution` contract can use `nexus launch RUN TASK --json`
|
|
62
|
+
to run foreground commands or installed Codex CLI instances. Follow the Nexus
|
|
63
|
+
agent's executable-plan and exact-attempt `collect` examples. Each launch needs
|
|
64
|
+
a live foreground owner; the runtime supplies no background scheduler, sandbox
|
|
65
|
+
or Git integration. Launch only independent tasks within the recorded capacity.
|
|
66
|
+
For a process attempt, cancellation is bound to its returned identity and only
|
|
67
|
+
the original owner signals the child it spawned. A lost owner or uncertain
|
|
68
|
+
descendant leaves the attempt allocated until explicit reconciliation.
|
|
69
|
+
|
|
70
|
+
`.bmad-plus/orchestrator-state.yaml` remains an optional navigation summary:
|
|
57
71
|
|
|
58
72
|
```yaml
|
|
59
73
|
parallel_tasks:
|