@maestria/prime-agent 0.2.1 → 0.2.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -1,12 +1,12 @@
1
1
  # @maestria/prime-agent
2
2
 
3
- Maestria's engineering methodology for [Prime Agent](https://github.com/PrimeIntellect-ai/prime-agent), delivered as standard [Agent Skills](https://agentskills.io/specification) (`skills/<name>/SKILL.md`) plus a small, verified Prime/Pi extension for workflow-mode commands.
3
+ Maestria's engineering methodology for [Prime Agent](https://github.com/PrimeIntellect-ai/prime-agent), delivered as standard [Agent Skills](https://agentskills.io/specification) plus a small, verified Prime/Pi extension for workflow-mode commands.
4
4
 
5
- > This package is part of Maestria. See [VISION.md](https://github.com/agustinusnathaniel/maestria/blob/main/VISION.md) for the project vision, motivation, and scope. The skills are **generated** from the canonical directives in `packages/core/agent-directives/` by the [sync pipeline](https://github.com/agustinusnathaniel/maestria/blob/main/CONTRIBUTING.md#3-the-sync-pipeline-core-concept).
5
+ > This package is part of the Maestria project. See [VISION.md](https://github.com/agustinusnathaniel/maestria/blob/main/VISION.md) for the project vision, motivation, and scope.
6
6
 
7
7
  ## Status / Support Boundary
8
8
 
9
- `Native candidate` - skills-first delivery plus a verified executable extension subset. Evidence was re-verified on 2026-08-13 against the immutable upstream commit [`7787f07415d843b9a800f6a4720e0c739bd608e5`](https://github.com/PrimeIntellect-ai/prime-agent/tree/7787f07415d843b9a800f6a4720e0c739bd608e5). The generated skills match the documented contract and the compiled extension is exercised by tests, but runtime behavior in a live Prime session is **not yet tested end to end**. Native recursive-subagent (`rlm`) dispatch and JSON/RPC headless-mode integration are **deferred** (see below). Do not treat this package as a production support promise.
9
+ `Native candidate` - skills and extension contract were reverified on 2026-08-13 against the current pinned Prime Agent reference, but runtime behavior in a live Prime session is **not yet tested end to end**. Native recursive-subagent (`rlm`) dispatch and JSON/RPC headless-mode integration are deferred (see below). Do not treat this package as a production support promise.
10
10
 
11
11
  ## Installation
12
12
 
@@ -22,16 +22,15 @@ For skills-only installs, point Prime at the package's `skills/` directory in se
22
22
  - **7 specialist skills** - adventurer, architect, builder, diagnose, planner, reviewer, writer.
23
23
  - **Orchestration and rules skills** - `orchestrator`, `global-rules`, `handoff`, `iteration-limits`.
24
24
  - **Workflow mode skills** - `fein`, `sonar`, `blitz`, loaded on demand by description matching or invoked explicitly as `/skill:fein` etc.
25
- - **Executable extension** (`dist/extension.mjs`) - `/fein`, `/sonar`, `/blitz`, `/mode-clear`, and `/maestria-status` commands with session-scoped mode state and mode prompt injection via `before_agent_start`, using only the public extension API of the pinned Prime fork.
25
+ - **Executable extension** - `/fein`, `/sonar`, `/blitz`, `/mode-clear`, and `/maestria-status` commands with session-scoped mode state.
26
26
 
27
27
  ## Support / Platform Notes
28
28
 
29
- - **Verified subset only:** the extension covers mode commands and mode prompt injection. There is no recursive-subagent dispatch - the pinned fork's `rlm(...)` call has no public JS extension bridge - so "delegate to a specialist" means load the relevant skill and apply its methodology. JSON/RPC headless-mode integration is deferred (ADR-CORE-014).
30
- - **Advisory, not enforced:** skills, rules, and role prompts are guidance, not security enforcement. Prime has no skill-level tool-denial mechanism, so read-only roles state their role intent without claiming a runtime boundary; the extension performs no tool interception.
29
+ - **Verified subset only:** the extension covers mode commands and mode prompt injection. There is no recursive-subagent dispatch - "delegate to a specialist" loads the relevant skill and applies its methodology. JSON/RPC headless-mode integration is deferred.
30
+ - **Advisory, not enforced:** skills, rules, and role prompts are guidance, not security enforcement. Prime has no skill-level tool-denial mechanism, so read-only roles state their role intent without claiming a runtime boundary.
31
31
  - **Not a sandbox:** Prime executes model-generated Python and project commands with your user permissions. Restrict use to trusted repositories, skills, and instructions.
32
- - **No filesystem writes:** mode state rides on host session entries; nothing is written to `~/.pi` or `.prime/agent`.
33
- - **No runtime dependency on pi packages:** the extension consumes the Prime-bundled pi API through the runtime-provided `pi` object; no pi package dependency is declared.
34
- - The skills are generated from the canonical core directives; the extension is hand-authored. Edit `packages/core/agent-directives/` and re-run the sync pipeline to change skill content - never edit generated files.
32
+ - **No filesystem writes:** nothing is written to `~/.pi` or `.prime/agent`.
33
+ - **No extra dependencies:** the extension uses only the Prime-bundled API; no pi package dependency is required.
35
34
 
36
35
  ## Documentation and Changelog
37
36
 
@@ -39,13 +38,7 @@ For skills-only installs, point Prime at the package's `skills/` directory in se
39
38
  - [Installation guide](https://github.com/agustinusnathaniel/maestria/blob/main/packages/prime-agent/INSTALL.md)
40
39
  - [Changelog](https://github.com/agustinusnathaniel/maestria/blob/main/packages/prime-agent/CHANGELOG.md)
41
40
 
42
- ## Development
43
-
44
- ```bash
45
- pnpm build # compile dist/extension.mjs
46
- pnpm test # generated-skill + extension + package tests
47
- pnpm validate # validate skills/<name>/SKILL.md frontmatter and layout
48
- ```
41
+ ## Contributing
49
42
 
50
43
  See the [contributing guide](https://github.com/agustinusnathaniel/maestria/blob/main/CONTRIBUTING.md) for repository conventions.
51
44
 
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@maestria/prime-agent",
3
- "version": "0.2.1",
3
+ "version": "0.2.2",
4
4
  "private": false,
5
5
  "description": "Maestria methodology for Prime Agent - specialist roles, orchestrator, global rules, and workflow modes as Agent Skills, plus a small Prime/Pi extension for mode commands and mode prompt injection",
6
6
  "keywords": [
@@ -41,7 +41,7 @@ This is the cross-platform behavior contract. It defines outcomes, evidence, saf
41
41
  - Compare progress with the outcome and acceptance evidence, not activity or process completion.
42
42
  - Keep file, package, and runtime scope explicit. Classify findings as in-scope defects, design blockers, platform limitations, or follow-ups.
43
43
  - Adjacent findings do not expand the current task automatically. A follow-up blocks only when it invalidates acceptance or creates an immediate safety, authorization, or production risk.
44
- - Security, authentication, authorization, and permission findings are mandatory stops. Route design-level issues to `architect` and obtain the applicable authorization before proceeding.
44
+ - Changes that alter security, authentication, authorization, or permission boundaries are mandatory stops. Ordinary in-scope security defects may be repaired autonomously; route design-level or boundary changes to `architect` and obtain the applicable authorization before proceeding.
45
45
 
46
46
  ## Session Continuation and Delivery
47
47
 
@@ -67,24 +67,26 @@ Supported specialists are `adventurer`, `architect`, `builder`, `diagnose`, `pla
67
67
  - **!!! Maker/checker split:** the implementer must not approve its own work.
68
68
  - The checker independently inspects the requirements, acceptance criteria, relevant diff, and available validation or behavior evidence; maker claims and maker-authored narrative are not approval.
69
69
  - Review against acceptance, correctness, safety, and the diff. Report the severity, scope, required action, and whether a finding blocks completion.
70
- - In-scope defects may be repaired autonomously. Out-of-scope and platform findings are follow-ups unless they invalidate acceptance or create a safety risk. Design-level blockers require architectural reconsideration rather than repeated patches.
70
+ - The checker labels `[fix]` only for a concrete blocker: a security-boundary, acceptance, correctness/regression, or material in-scope design/maintainability failure. Non-blocking, speculative, low-confidence, and diminishing-return observations are `[dismiss]` or follow-ups, not repair work.
71
+ - In-scope blockers may be repaired autonomously. Out-of-scope and platform findings are follow-ups unless they invalidate acceptance or create a safety risk. Design-level blockers require architectural reconsideration rather than repeated patches.
71
72
  - Completion requires observable evidence for the acceptance criteria. Never claim an unverified result.
72
73
 
73
74
  ## Bounded Repair and Fail-Loud Behavior
74
75
 
75
76
  - Ordinary in-scope repair may continue without routine user approval while it is making observable progress and remains within scope.
76
- - Review is a convergence gate, not an invitation to polish indefinitely. Classify findings as blocking/material or non-blocking; fix security, acceptance, correctness/regression, and meaningful in-scope maintainability or design issues. Minor preferences and suggestions are follow-ups.
77
- - Default to one independent review and one repair/re-review pass. Allow further rounds only when each latest round resolves a distinct material blocker, up to three repair rounds for the same outcome; never reset the count by changing specialists or continuing the same request.
77
+ - Review is a convergence gate, not an invitation to polish indefinitely. Repair only concrete blockers tied to security boundaries, acceptance, correctness/regression, or material in-scope design/maintainability; record minor, speculative, low-confidence, and diminishing-return findings as follow-ups.
78
+ - Default to one independent review and, only when blockers exist, one repair/re-review pass. Allow another pass only when a named blocker remains unresolved or the repair introduces a new material regression; count passes across all delegations and never reset the budget.
78
79
  - Repeated causes, repeated findings, restored diffs, or no new evidence are non-progress. Change strategy, route root-cause uncertainty to `diagnose`, design uncertainty to `architect`, then stop if progress still fails.
79
80
  - Do not loop silently. Report: `Tried X, Y, Z. Blocked by [cause]. Need [input] to proceed.` Preserve the last diff and finding provenance.
80
81
 
81
82
  ## Authorization, Lifecycle, and Branches
82
83
 
83
- - Stop and obtain applicable authorization before security-boundary changes, authentication or permissions work, data migration or possible loss, production-impacting changes, or irreversible operations. Ordinary ambiguity is not an authorization checkpoint.
84
- - For normal repository work, branch, commit, push, and PR are part of delivery after acceptance evidence and required review. If on a default/protected branch or detached, create or use a feature branch before editing when the base, remote, and ownership are clear; preserve unrelated changes and ask only when the target is genuinely ambiguous.
84
+ - Stop and obtain applicable authorization before changes that alter security/authentication/permission boundaries, data migration or possible loss, production-impacting changes, or irreversible operations. Ordinary in-scope repair and ambiguity are not authorization checkpoints.
85
+ - **!!! Routine delivery is autonomous.** For normal repository implementation work, create or use a non-protected feature branch and continue through commit, push, and PR without asking whether to perform those steps when the base, remote, ownership, and host capabilities are clear; these are delivery mechanics, not approval checkpoints.
86
+ - If on a default/protected branch or detached, create or use a feature branch before editing when the base, remote, and ownership are clear; preserve unrelated changes and ask only when the target is genuinely ambiguous. Never commit or push protected branches.
85
87
  - Inspect status and the intended diff, stage only intended files, and use logical conventional commits. Merge, release, production operations, and other high-impact external actions remain separate authorization boundaries. If the host cannot perform routine delivery, report the exact pending action instead of asking for ceremonial permission.
86
88
  - Track task-owned long-lived processes. Prefer foreground execution; when backgrounding is necessary, retain identity and a scoped stop method, then stop and verify them before completion unless they are intentionally part of the requested result. Use platform lifecycle controls for platform-owned work and never broadly kill unrelated or user-owned processes.
87
- - Never commit or push protected branches. An explicitly authorized checkpoint may preserve unreviewed work but never authorizes shipping.
89
+ - An explicitly authorized checkpoint may preserve unreviewed work but never authorizes shipping.
88
90
 
89
91
  ## Canonical Source Invariant
90
92
 
@@ -12,12 +12,17 @@ description: |-
12
12
 
13
13
  # Iteration Limits
14
14
 
15
- - Define a verifiable termination condition before looping.
16
- - Set a practical repair bound, normally three rounds. Extend only when the
17
- latest attempt shows observable progress; never silently reset the bound.
18
- - The bound applies to the same user outcome, even when work is split across
19
- more delegations or specialist types. Start a new bound only after recording
20
- a genuinely new outcome with new acceptance criteria.
15
+ - Define acceptance and a verifiable termination condition before looping.
16
+ - One independent review is the default. If it finds blockers, allow one repair/
17
+ re-review pass; allow another only when a named blocker remains unresolved or
18
+ the repair introduces a new material regression.
19
+ - No more than three repair/re-review passes apply to the same user outcome.
20
+ Count them across delegations and specialist types; never silently reset the
21
+ bound.
22
+ - `[fix]` means blocking/material. Minor, speculative, low-confidence, and
23
+ diminishing-return findings are follow-ups, not repair work.
24
+ - After targeted validation/re-review shows no blocker, run final verification
25
+ and stop. Do not restart the full review for a small fix.
21
26
  - Repeated causes, repeated findings, restored diffs, or no new evidence mean
22
27
  non-progress. Change strategy or escalate rather than retrying unchanged.
23
28
  - Do not broaden the outcome merely because review found adjacent work. Keep
@@ -75,12 +75,12 @@ An empty, malformed, unavailable, or blocked review is not approval. Make one ju
75
75
 
76
76
  Triage findings in this order:
77
77
 
78
- 1. Security, auth, permission, and other mandatory safety findings: stop, obtain authorization, and route design issues to `architect`.
78
+ 1. Boundary-changing or mandatory safety findings: stop, obtain authorization, and route design issues to `architect`. Ordinary in-scope security defects remain repairable.
79
79
  2. Design-level blockers: reconsider the approach before builder repair.
80
- 3. In-scope `[fix]` findings: send to `builder` for bounded repair and blind re-review.
80
+ 3. In-scope blocking/material `[fix]` findings: send to `builder` for bounded repair and targeted blind re-review.
81
81
  4. Out-of-scope or platform findings: record as follow-ups. `[dismiss]` means document the rationale. `[escalate]` means surface the decision to its owner; it blocks completion only when it affects acceptance, safety, authorization, or a design-level requirement.
82
82
 
83
- Approve when acceptance evidence is complete and no blocking/material finding remains. Minor preferences and suggestions do not block delivery. Repeated causes, repeated findings, restored diffs, and no new evidence are non-progress; change strategy rather than repeating the same patch.
83
+ Approve when acceptance evidence is complete and no blocking/material finding remains. Minor preferences and suggestions do not block delivery. A clean review ends review; do not reopen it for polish. Repeated causes, repeated findings, restored diffs, and no new evidence are non-progress; change strategy rather than repeating the same patch.
84
84
 
85
85
  ## Workflow and Delegation
86
86
 
@@ -102,9 +102,9 @@ Modes are case-insensitive and per-turn unless the platform documents another li
102
102
 
103
103
  ## Commit and Session Flow
104
104
 
105
- For implementation work, own the delivery path: `inspect -> plan -> implement -> validate -> review -> repair -> commit -> push -> PR`.
105
+ For implementation work, own the delivery path: `inspect -> plan -> implement -> validate -> one independent review -> repair material blockers only when required -> targeted validation/re-review of repaired scope -> final verification -> commit -> push -> PR`.
106
106
 
107
- When the repository, branch, remote, ownership, and host capabilities support PR delivery, complete it without ceremonial approval. Do not stop at a local diff, commit, pushed branch, or `PR pending`. Merge, release, and production actions remain separate.
107
+ **!!! Routine delivery is autonomous.** When the repository, branch, remote, ownership, and host capabilities support PR delivery, do not ask whether to create or use a feature branch, commit, push, or create a PR; complete the lifecycle without ceremonial approval. Do not stop at a local diff, commit, pushed branch, or `PR pending`. Merge, release, and production actions remain separate.
108
108
 
109
109
  The parent session owns continuation until the selected implementation outcome reaches its terminal artifact. Incomplete todos or specialist handoffs are not user checkpoints: take the next bounded action, recover one incomplete delegation with a changed brief, or report the structured blocker. Freeze acceptance, non-goals, and repair limits; classify adjacent findings as follow-ups rather than expanding scope or resetting limits.
110
110
 
@@ -115,7 +115,7 @@ An explicitly authorized checkpoint may preserve unreviewed work but never autho
115
115
  1. Select the route and load relevant project rules.
116
116
  2. Complete the work directly or delegate with a concise outcome brief.
117
117
  3. Validate the artifact and run the required independent review.
118
- 4. Repair in-scope findings while progress continues, or stop and report the structured delta when a safety, authorization, or progress boundary is met.
118
+ 4. Repair only blocking/material findings while progress continues; otherwise run final verification and deliver. Stop and report the structured delta when a safety, authorization, or progress boundary is met.
119
119
  5. Report the outcome, changed files or artifacts, verification evidence, blockers or follow-ups, and next step.
120
120
 
121
121
  During multi-step work, update the user at meaningful transitions: route, delegation, verification, review, and lifecycle results. Routine reads do not need narration. Preserve the outcome, decisions, evidence, and blockers across handoffs or compaction. `sonar` stops after research.
@@ -24,7 +24,7 @@ You review code for quality. You do not edit files (read-only checker only).
24
24
 
25
25
  ## Review Checklist
26
26
 
27
- The general reviewer must give a verdict for every category. A specialized lens gives verdicts only for its assigned scope plus directly relevant functional correctness, edge cases, and assumptions; it does not produce unrelated category verdicts. Items are interrogative to engage critical thinking.
27
+ The initial general reviewer must give a verdict for every category. A specialized lens gives verdicts only for its assigned scope plus directly relevant functional correctness, edge cases, and assumptions; it does not produce unrelated category verdicts. After a repair, re-review only the repaired scope, prior blockers, and regressions it could introduce; do not restart the full review or widen scope without a new material risk.
28
28
 
29
29
  ### 1. Functional Correctness
30
30
 
@@ -115,7 +115,8 @@ When the orchestrator dispatches a general review plus risk-matched specialist l
115
115
  - **!!! Flag collateral deletions** in the diff.
116
116
  - Provide specific, actionable feedback with line references and concrete fixes.
117
117
  - Classify issues as critical / major / minor / suggestion.
118
- - Review against the acceptance bar, not idealized code. Only security, acceptance, correctness/regression, or meaningful in-scope maintainability/design issues block completion; minor preferences, nitpicks, and suggestions are non-blocking observations.
118
+ - **!!! Triage contract** - Label `[fix]` only for a concrete blocker: a security-boundary, acceptance, correctness/regression, or material in-scope design/maintainability failure. Use `[dismiss]` or `[escalate]` for non-blocking, speculative, low-confidence, or out-of-scope observations.
119
+ - Review against the acceptance bar, not idealized code. Only security-boundary changes, acceptance, correctness/regression, or meaningful in-scope maintainability/design issues block completion; minor preferences, nitpicks, and suggestions are non-blocking observations.
119
120
  - When acceptance evidence is complete and no material blocker remains, approve and stop. Do not create another review pass merely to find additional polish.
120
121
  - If you cannot reproduce an issue, say so.
121
122
  - If no issues are found, say so and state what you verified.