planr 1.4.0 → 1.5.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +21 -61
- package/docs/ARCHITECTURE.md +4 -4
- package/docs/CLAUDE_CODE.md +2 -2
- package/docs/CLI_REFERENCE.md +6 -14
- package/docs/CODEX.md +6 -6
- package/docs/CURSOR.md +2 -2
- package/docs/EXAMPLE_WEBAPP.md +63 -97
- package/docs/GOALS.md +8 -25
- package/docs/INSTALL.md +3 -3
- package/docs/MCP_CONTRACT.md +6 -9
- package/docs/MODEL_ROUTING.md +20 -177
- package/docs/RELEASE.md +1 -1
- package/docs/ROUTING_BUNDLES.md +32 -0
- package/docs/SKILLS.md +35 -50
- package/docs/documentation/CONTRACT.md +164 -0
- package/docs/documentation/COVERAGE.md +91 -0
- package/docs/documentation/INFORMATION_ARCHITECTURE.md +147 -0
- package/docs/fixtures/mcp-contract.json +3 -11
- package/npm/native/darwin-arm64/planr +0 -0
- package/npm/native/darwin-x86_64/planr +0 -0
- package/npm/native/linux-arm64/planr +0 -0
- package/npm/native/linux-x86_64/planr +0 -0
- package/package.json +25 -14
- package/plugins/planr/.claude-plugin/plugin.json +1 -1
- package/plugins/planr/.codex-plugin/plugin.json +2 -2
- package/plugins/planr/skills/planr-loop/SKILL.md +1 -1
- package/docs/PRESET_COMPOSITION.md +0 -109
- package/docs/PRESET_EVALUATION.md +0 -118
- package/docs/PRESET_REGISTRY.md +0 -61
- package/evaluations/preset-suite-v1.toml +0 -127
- package/evaluations/sol-luna-codex-v1.toml +0 -42
- package/plugins/planr/skills/planr-loop/agents/planr-reviewer.toml +0 -23
- package/plugins/planr/skills/planr-loop/agents/planr-worker.toml +0 -21
- package/presets/bindings/claude-native.toml +0 -52
- package/presets/bindings/codex-openai.toml +0 -56
- package/presets/bindings/cursor-fable-grok.toml +0 -49
- package/presets/bindings/cursor-openai.toml +0 -49
- package/presets/bindings/mixed-host.toml +0 -64
- package/presets/policies/balanced.toml +0 -50
- package/presets/policies/low-usage.toml +0 -50
- package/presets/policies/max-quality.toml +0 -50
- package/presets/policies/read-only-audit.toml +0 -50
- package/website/README.md +0 -79
- package/website/_headers +0 -7
- package/website/alchemy-runtime.test.mjs +0 -21
- package/website/app.mjs +0 -216
- package/website/build-catalog.mjs +0 -135
- package/website/build-site.test.mjs +0 -69
- package/website/catalog-model.mjs +0 -185
- package/website/catalog-model.test.mjs +0 -124
- package/website/cloudflare-launcher.test.mjs +0 -38
- package/website/data/catalog.json +0 -307
- package/website/index.html +0 -122
- package/website/registry/manifest.toml +0 -48
- package/website/registry/report.md +0 -33
- package/website/registry/trusted-maintainers.toml +0 -6
- package/website/registry/verification.json +0 -7258
- package/website/serve.mjs +0 -41
- package/website/styles.css +0 -201
- package/website/test-fixtures/recommended.json +0 -72
package/docs/MCP_CONTRACT.md
CHANGED
|
@@ -42,10 +42,7 @@ Required groups:
|
|
|
42
42
|
- artifact add, list, and show
|
|
43
43
|
- event list and debug bundle preview
|
|
44
44
|
- trace item, log add, and log read (including three-stage route observations)
|
|
45
|
-
-
|
|
46
|
-
- declarative registry verification with canonical evaluation/safe-binding gates, preview-first immutable import, and manifest-anchored integrity/signature/freshness-checked offline cache listing
|
|
47
|
-
- policy preset preview/apply by path or built-in id with repository-only target validation and deterministic provenance lock
|
|
48
|
-
- deterministic offline preset simulation plus explicit opt-in live-host execution with Planr-controlled challenge workspaces, strict task artifacts read and hashed by Planr, candidate/task outcome oracles, failed-live-attempt `unverified`/incomplete lifecycle semantics, production policy-capability checks, estimated arbitrary-process claims, optional Ed25519-verified run/suite/time/task/challenge-bound telemetry that alone can promote effective route and usage evidence to trusted/recommendation-eligible, observed process latency, transition/correction/violation counts, result hashes, and deterministic lifecycle thresholds
|
|
45
|
+
- provider-neutral RoutingBundle v1 inspection, repository-safe preview/apply, and durable application evidence
|
|
49
46
|
- review annotate, ingest, artifact, evidence, and close
|
|
50
47
|
- item close, context create, and search
|
|
51
48
|
|
|
@@ -65,12 +62,12 @@ HTTP mirrors the same rule: `GET /v1/reviews/:id/artifact` is read-only; `POST /
|
|
|
65
62
|
|
|
66
63
|
## Install Contract
|
|
67
64
|
|
|
68
|
-
`planr install <client> --dry-run` prints
|
|
65
|
+
`planr install <client> --dry-run` prints the complete client-owned MCP, role, skill, and hook-reconciliation paths for Codex, Claude Code, and Cursor without writing them. Non-dry install writes only repository-local files, with this ownership contract:
|
|
69
66
|
|
|
70
|
-
- Codex: `.planr/integrations/codex-mcp.toml`
|
|
71
|
-
- Claude Code: `.mcp.json
|
|
72
|
-
- Cursor: `.cursor/mcp.json
|
|
67
|
+
- Codex: the CLI writes `.planr/integrations/codex-mcp.toml` and `.codex/hooks.json`; the plugin owns all ten workflow skills; neither path writes Planr project roles or project skills
|
|
68
|
+
- Claude Code: the CLI writes `.mcp.json`, standalone `.claude/agents/` roles, and `.claude/settings.json` hooks, but no project skills; the plugin owns all ten workflow skills and its plugin agents
|
|
69
|
+
- Cursor: the CLI writes `.cursor/mcp.json`, both `.cursor/agents/` roles, all ten `.cursor/skills/` skill copies, and `.cursor/hooks.json`
|
|
73
70
|
|
|
74
71
|
The Cursor dry-run additionally prints a `cursor://anysphere.cursor-deeplink/mcp/install` link whose embedded config (`planr mcp`, no `--db`) is safe at user scope because each workspace resolves its own database. Planr does not edit global client configuration without a separate explicit operator action; the deeplink requires the operator to click it and confirm inside Cursor.
|
|
75
72
|
|
|
76
|
-
|
|
73
|
+
`--no-mcp` skips only the project MCP artifact: Codex reconciles hooks only; Claude Code writes standalone roles and hooks but no project skills; Cursor writes roles, all ten skills, and hooks. `--no-hooks` is the independent hook opt-out and can be combined with `--no-mcp`.
|
package/docs/MODEL_ROUTING.md
CHANGED
|
@@ -1,196 +1,39 @@
|
|
|
1
|
-
# Model
|
|
1
|
+
# Model routing
|
|
2
2
|
|
|
3
|
-
|
|
4
|
-
|
|
5
|
-
The pattern this makes declarative: strongest model plans and judges, a cheap fast steerable model implements, token-hungry side work (browser verification, codebase analysis) goes to budget profiles. Without a registry that knowledge lives in CLAUDE.md prose, Codex agent TOMLs, and Cursor frontmatter — three dialects that drift. With a registry it travels inside every pick packet.
|
|
6
|
-
|
|
7
|
-
Routing is **advisory by design**: Planr never calls model providers and never blocks a pick because a profile is unavailable. Your host (Codex, Claude Code, Cursor, any MCP client) stays the dispatch authority.
|
|
8
|
-
|
|
9
|
-
Want the whole flow on a concrete project first? [Worked Example: Routing a Small Web App](EXAMPLE_WEBAPP.md) walks a frontend/backend todo app from pool declaration to the audit trail, with real outputs.
|
|
10
|
-
|
|
11
|
-
## Quick Start
|
|
12
|
-
|
|
13
|
-
One command writes a working starter registry — the cost-tiering defaults with a premium driver, a standard implementer, and a budget helper, commented so the tiers explain themselves:
|
|
14
|
-
|
|
15
|
-
```bash
|
|
16
|
-
planr agents init # writes .planr/agents.toml; never overwrites without --force
|
|
17
|
-
```
|
|
18
|
-
|
|
19
|
-
Or declare `.planr/agents.toml` by hand:
|
|
3
|
+
Planr Core treats routing as optional, advisory repository data. `.planr/agents.toml` declares opaque profiles and routes; Planr resolves them into pick packets but never calls a provider or claims that a requested model actually ran.
|
|
20
4
|
|
|
21
5
|
```toml
|
|
22
|
-
[profiles.
|
|
23
|
-
client = "
|
|
24
|
-
model = "
|
|
6
|
+
[profiles.worker]
|
|
7
|
+
client = "host-a"
|
|
8
|
+
model = "model-id"
|
|
9
|
+
agent_type = "repository-role"
|
|
25
10
|
effort = "high"
|
|
26
|
-
|
|
27
|
-
capabilities = ["orchestration", "review", "planning"]
|
|
28
|
-
notes = "Planner/architect and judge. Verdicts stay on this tier."
|
|
29
|
-
|
|
30
|
-
[profiles.gpt55-coder]
|
|
31
|
-
client = "codex"
|
|
32
|
-
model = "gpt-5.5"
|
|
33
|
-
effort = "xhigh"
|
|
34
|
-
cost_tier = "standard"
|
|
35
|
-
capabilities = ["code", "steerable"]
|
|
36
|
-
notes = "Primary implementer: strong, fast, cheap on subscription."
|
|
11
|
+
skill = "planr-work"
|
|
37
12
|
|
|
38
13
|
[[routes]]
|
|
39
14
|
match = { work_type = "code" }
|
|
40
|
-
profile = "
|
|
41
|
-
fallbacks = ["fable-driver"]
|
|
42
|
-
|
|
43
|
-
[[routes]]
|
|
44
|
-
match = { work_type = "review" }
|
|
45
|
-
profile = "fable-driver"
|
|
46
|
-
|
|
47
|
-
[route_default]
|
|
48
|
-
profile = "gpt55-coder"
|
|
49
|
-
fallbacks = ["fable-driver"]
|
|
50
|
-
```
|
|
51
|
-
|
|
52
|
-
Validate and inspect:
|
|
53
|
-
|
|
54
|
-
```bash
|
|
55
|
-
planr agents check # non-zero exit only on parse failure; warnings pass
|
|
56
|
-
planr agents list # resolved profiles, routes, and warnings
|
|
15
|
+
profile = "worker"
|
|
57
16
|
```
|
|
58
17
|
|
|
59
|
-
|
|
18
|
+
Create a neutral scaffold or an explicit registry:
|
|
60
19
|
|
|
61
20
|
```bash
|
|
62
|
-
planr
|
|
63
|
-
|
|
64
|
-
|
|
65
|
-
|
|
66
|
-
"routing": {
|
|
67
|
-
"profile": "gpt55-coder",
|
|
68
|
-
"client": "codex",
|
|
69
|
-
"model": "gpt-5.5",
|
|
70
|
-
"effort": "xhigh",
|
|
71
|
-
"cost_tier": "standard",
|
|
72
|
-
"fallbacks": ["fable-driver"],
|
|
73
|
-
"matched_selector": "work_type=code"
|
|
74
|
-
}
|
|
21
|
+
planr agents init
|
|
22
|
+
planr agents init --profile worker=host-a/model-id@high#standard --route code=worker
|
|
23
|
+
planr agents check
|
|
24
|
+
planr agents list --json
|
|
75
25
|
```
|
|
76
26
|
|
|
77
|
-
|
|
78
|
-
|
|
79
|
-
## The Registry File
|
|
80
|
-
|
|
81
|
-
`.planr/agents.toml` has three parts:
|
|
82
|
-
|
|
83
|
-
- `[profiles.<id>]` — a named agent setting. `client` (which host dispatches it: `codex`, `claude-code`, `cursor`, `generic-mcp`) and `model` are required; `effort`, `cost_tier` (`premium` | `standard` | `budget`), `capabilities`, and `notes` are optional. Model ids and aliases pass through verbatim — Planr does not validate them against provider catalogs, so new models need no Planr release.
|
|
84
|
-
- `[[routes]]` — `match` selects work (`work_type = "code"` or `plan = "pln-1234abcd"`), `profile` names the primary, `fallbacks` the ordered alternatives.
|
|
85
|
-
- `[route_default]` — catches everything no route matched.
|
|
86
|
-
|
|
87
|
-
Resolution precedence per item: **per-item override > `work_type` route > `plan` route > default**. Within a level, the first declared route wins. If a chain's primary profile id is unknown, the first known fallback is promoted; a chain with no known profiles falls through to the next precedence level, so a typo never swallows lower routes. `matched_selector` in the output tells you which rule fired (`override`, `work_type=<v>`, `plan=<v>`, or `default`).
|
|
27
|
+
Resolution order is per-item override, work type, plan, then default route. Unknown profiles fail open to the next applicable route. Host names, model ids, role selectors, effort values, and fallback behavior are opaque to Core.
|
|
88
28
|
|
|
89
|
-
|
|
29
|
+
Workers may report observed routing with logs and route-audit evidence. Requested-only values never become effective proof; missing effective evidence remains explicitly unavailable.
|
|
90
30
|
|
|
91
|
-
|
|
31
|
+
Host-specific model policies and generated repository roles are optional. The `planr-routing` workspace package compiles them into RoutingBundle v1, and Core safely previews and applies that bundle:
|
|
92
32
|
|
|
93
33
|
```bash
|
|
94
|
-
planr
|
|
95
|
-
planr
|
|
96
|
-
planr
|
|
34
|
+
planr-routing compile balanced --host codex-openai --output routing-bundle.json
|
|
35
|
+
planr routing bundle preview routing-bundle.json
|
|
36
|
+
planr routing bundle apply routing-bundle.json
|
|
97
37
|
```
|
|
98
38
|
|
|
99
|
-
|
|
100
|
-
|
|
101
|
-
Tier the roles, not just the models: workers run safely on cheaper tiers because the pick packet bounds their scope, while review verdicts should stay on the strongest tier — `agents check` warns when review work routes to a `budget` profile. Background: [Cost Tiering](GOALS.md#cost-tiering).
|
|
102
|
-
|
|
103
|
-
## Host-Native Rendering
|
|
104
|
-
|
|
105
|
-
Routes only matter if the host actually dispatches the declared model, so `planr install codex|claude|cursor` closes the gap: when a registry is present, the provisioned subagent role files are rendered with pins taken from it instead of the shipped static defaults. The `work_type=code` route pins the worker role, the `work_type=review` route pins the reviewer role, and each render uses the host's exact vocabulary — Codex TOML gets `model` and `model_reasoning_effort` (with `developer_instructions` always present, since Codex silently ignores a role file without it), Claude frontmatter gets `model:` and `effort:`, Cursor frontmatter gets `model:` only.
|
|
106
|
-
|
|
107
|
-
Two safety rules keep this predictable:
|
|
108
|
-
|
|
109
|
-
- **Client matching**: a role file only pins profiles whose `client` matches the install target, scanning the route's fallback chain for the first match. A review route pointing at a Cursor profile never writes a Cursor model id into a Codex TOML — that role keeps its static default instead.
|
|
110
|
-
- **Provision-once**: existing files are never overwritten. After editing the registry, re-render explicitly with `planr install <client> --force`. Rendered files start with a `# generated from .planr/agents.toml` header so you (and future audit tooling) can tell them from hand-maintained ones.
|
|
111
|
-
|
|
112
|
-
Without a registry, installs write the static role files byte-identically to previous releases.
|
|
113
|
-
|
|
114
|
-
## Prompt Routing
|
|
115
|
-
|
|
116
|
-
`planr prompt routing [--client codex|claude|cursor|all]` prints a paste-ready block for the driver session: the prioritization table (every route, profile, and fallback in precedence order), per-host dispatch guidance including the traps that silently defeat pins (Codex requires `fork_turns: "none"` and a session restart after re-rendering; the `CLAUDE_CODE_SUBAGENT_MODEL` env var preempts Claude frontmatter; Cursor plan mode, admin policy, and Max Mode override silently), and process-dispatch snippets (`codex exec`, `pi`, `opencode run`) for hosts without role files, pre-filled from the `work_type=code` route. `--json` carries the same content structured.
|
|
117
|
-
|
|
118
|
-
## Run Audit
|
|
119
|
-
|
|
120
|
-
Every host has a silent override path — the `CLAUDE_CODE_SUBAGENT_MODEL` env var, Cursor plan/admin/Max-Mode policy, Codex full-history forks, org allowlists — so a pin alone is not proof. The audit loop closes this at two levels: workers report the profile they actually ran on via `planr log add`/`planr done --profile <id>` (or `PLANR_PROFILE`) for backward-compatible mismatch checks, and can attach a strict `--route-audit <observation.json>` that keeps requested, host-resolved, and effective model/effort/fork values separate. Every dimension carries enforcement confidence and a constrained evidence source; missing host evidence stays unavailable instead of inheriting the request.
|
|
121
|
-
|
|
122
|
-
- `planr trace item <id>` (MCP: `planr_trace_item`) shows the declared route next to every run's actual client/profile and three-stage observation with a `mismatch` marker.
|
|
123
|
-
- Runs also record the host they observably executed under (`observed_client`, detected from environment variables the hosts set themselves — no flags); a run whose host differs from the declared route's client emits an advisory `client_mismatch_observed` event, which catches exactly the deviation profile self-report cannot: a different host standing in for the declared client, even when the model matched.
|
|
124
|
-
- `planr doctor` reports the registry state (absent, degraded with parse context, loaded with counts and warnings) and flags rendered role files that drifted from the current registry (`planr install <client> --force` re-renders).
|
|
125
|
-
- `planr export`/`import` carry the registry with the package, preview-first; an existing registry at the destination is never silently overwritten.
|
|
126
|
-
|
|
127
|
-
Everything here is advisory (ADR-001): mismatches never fail logging, reviews, or closes. No profile reported, no run recorded, or no registry means no comparison and no event.
|
|
128
|
-
|
|
129
|
-
One legitimate mismatch source to know: a driver adding a live-verification log to a routed item runs on the driver profile by design, which emits a `route_mismatch_observed` event. The payload carries `log_kind`, so audit consumers can discount `verification` entries and alarm only on `completion` mismatches.
|
|
130
|
-
|
|
131
|
-
For single-host pools (e.g. all-Cursor), declare the host's *exact* model slugs (`claude-opus-4-8-thinking-high`, not `opus`): dispatch APIs resolve slugs, not aliases, and a driver forced to map `fable-5` onto the nearest slug at dispatch time is a silent translation the audit cannot see.
|
|
132
|
-
|
|
133
|
-
## Failure Behavior
|
|
134
|
-
|
|
135
|
-
- **No registry file**: nothing changes. Pick packets simply have no `routing` key.
|
|
136
|
-
- **Malformed registry**: `planr agents check` fails with the parser's line context; everything else (`pick`, `map`, `install`) keeps working with routing omitted — installs fall back to the static role files.
|
|
137
|
-
- **Warnings** (unknown profile references, empty or duplicate selectors, budget-tier review routes, secret-like values) never block anything; `agents check` lists them and still exits zero.
|
|
138
|
-
- Never put credentials in the registry — it holds configuration strings only, and secret-like values are flagged.
|
|
139
|
-
|
|
140
|
-
## Use-Case Pools
|
|
141
|
-
|
|
142
|
-
Work types are free-form, and that makes them the use-case dimension: beyond the built-in vocabulary (`code`, `fix`, `review`, `docs`, ...), any string you pass to `--work-type` routes. Combined with per-profile skill pairing, the registry becomes a small agent pool — each use case names who runs it, on what model, with which skill:
|
|
143
|
-
|
|
144
|
-
```toml
|
|
145
|
-
[profiles.designer]
|
|
146
|
-
client = "claude-code"
|
|
147
|
-
model = "opus"
|
|
148
|
-
effort = "high"
|
|
149
|
-
cost_tier = "premium"
|
|
150
|
-
skill = "frontend-design" # dispatch this profile *with* this skill
|
|
151
|
-
|
|
152
|
-
[profiles.backender]
|
|
153
|
-
client = "codex"
|
|
154
|
-
model = "gpt-5.5"
|
|
155
|
-
effort = "xhigh"
|
|
156
|
-
cost_tier = "standard"
|
|
157
|
-
skill = "planr-work"
|
|
158
|
-
|
|
159
|
-
[[routes]]
|
|
160
|
-
match = { work_type = "frontend" }
|
|
161
|
-
profile = "designer"
|
|
162
|
-
fallbacks = ["driver"]
|
|
163
|
-
|
|
164
|
-
[[routes]]
|
|
165
|
-
match = { work_type = "design" }
|
|
166
|
-
profile = "designer"
|
|
167
|
-
|
|
168
|
-
[[routes]]
|
|
169
|
-
match = { work_type = "backend" }
|
|
170
|
-
profile = "backender"
|
|
171
|
-
fallbacks = ["driver"]
|
|
172
|
-
```
|
|
173
|
-
|
|
174
|
-
Create items with the use-case work type (`planr item create ... --work-type frontend`) — or retag existing ones with `planr item update <id> --work-type frontend`, which is how planning agents tag `map build` output against the declared routes (the planning skills read `agents list` and do this without user involvement) — and the pick packet carries the full pairing — `"profile": "designer"`, `"model": "opus"`, `"skill": "frontend-design"` — so the driver dispatches profile and skill together (`Use $frontend-design on item <id>` on the profile's client/model). Workers pull their slice of the pool with `planr pick --work-type frontend`. `skill` is passthrough vocabulary like model ids: Planr never validates it against installed skills, and profiles without one omit the key entirely. A profile that needs different skills for different use cases is simply two profiles.
|
|
175
|
-
|
|
176
|
-
Declare the `client` you will actually dispatch on. A loop running inside one host dispatches that host's subagents — an in-Cursor driver that dispatches Cursor subagents with per-dispatch models is running `client = "cursor"` profiles in practice, even when the model matches. A `client = "codex"` profile is only honest when the driver really spawns a Codex process (`codex exec ...`). This matters for the audit: workers report the *profile id*, so a profile whose declared client differs from the real dispatch host passes the mismatch check on the model alone — the client deviation stays invisible.
|
|
177
|
-
|
|
178
|
-
When do you actually need more than one client? Hosts with a full model catalog (Cursor) can serve an entire pool natively — an all-`cursor` registry with different models per profile is the normal case there. Cross-client profiles exist for two real situations: vendor-locked hosts (Claude Code dispatches only Anthropic models, Codex CLI only OpenAI models — a Claude-Code driver that wants a GPT implementer must process-dispatch via `codex exec`), and subscription economics (the same model can bill differently per host, so routing backend work through a flat-rate CLI subscription instead of the driver host's quota is a legitimate cost decision).
|
|
179
|
-
|
|
180
|
-
## Host Matrix
|
|
181
|
-
|
|
182
|
-
Where each host reads its model configuration from, and what silently defeats a pin there (state of July 2026):
|
|
183
|
-
|
|
184
|
-
| Host | Native mechanism | Rendered by `planr install`? | Silent overrides / traps |
|
|
185
|
-
| --- | --- | --- | --- |
|
|
186
|
-
| Cursor | `.cursor/agents/*.md` frontmatter `model: <id>` (default `inherit`) | yes (`cursor`) | Team-admin model policy, plan availability, and Max-Mode-only models override without error; legacy request-based plans force Composer for subagents; subagent transcripts record no model field, so the actual model cannot be verified from artifacts after the fact — the dispatch parameters in the driver session are the only record |
|
|
187
|
-
| Claude Code | `planr-worker.md`/`planr-reviewer.md` frontmatter `model:` + `effort:` | yes (`claude`) | `CLAUDE_CODE_SUBAGENT_MODEL` clamps frontmatter and per-invocation models with no signal ([#57718](https://github.com/anthropics/claude-code/issues/57718)); since v2.1.196 `inherit` behaves as unset; org `availableModels` allowlists fall back silently |
|
|
188
|
-
| Codex CLI | `.codex/agents/*.toml` with `model` + `model_reasoning_effort` | yes (`codex`) | `fork_turns = "all"` intentionally drops the child's `agent_type`/`model` — use `fork_turns = "none"` or a partial fork; the role registry loads at session start ([#26408](https://github.com/openai/codex/issues/26408)), so re-renders need a restart |
|
|
189
|
-
| opencode | `opencode.json` `agent.<name>.model = "provider/model-id"` or `.opencode/agents/*.md` frontmatter | no — use the `planr prompt routing` process snippet | Subagent inherits the primary model when unset; malformed `provider/model-id` strings (quoting, trailing newline) raise `ProviderModelNotFoundError` ([#5623](https://github.com/sst/opencode/issues/5623)) |
|
|
190
|
-
| Pi | none by design — process-level dispatch (`pi --provider --model --thinking`) or the `pi-subagents` extension (`.pi/agents/*.md`) | no — use the `planr prompt routing` process snippet | Extension model-scope enforcement against `enabledModels` is opt-in; without it, pins are best-effort |
|
|
191
|
-
|
|
192
|
-
For the hosts without rendered role files, `planr prompt routing` prints ready process-dispatch snippets pre-filled from the registry. Whatever the host does, the [run audit](#run-audit) catches silent overrides after the fact.
|
|
193
|
-
|
|
194
|
-
## Command Summary
|
|
195
|
-
|
|
196
|
-
The registry surface end to end: `planr agents init [--force]` scaffolds, `planr agents list|check` inspect and validate, `planr pick --json` carries the `routing` block, `planr item route [--set|--clear]` pins per item, the MCP tools (`planr_agents_list`, `planr_item_route`, `planr_item_route_set`, `planr_item_route_clear`) return identical JSON shapes, `planr install <client> [--force]` renders host role files from the registry, `planr prompt routing` prints the driver dispatch block, `planr log add`/`done --profile` (or `PLANR_PROFILE`) feed the run audit, `planr trace item` and `planr doctor` surface mismatches and drift, and `planr export`/`import` carry the registry preview-first.
|
|
39
|
+
Bundle application is restricted to the repository. Planr never edits user configuration such as `~/.codex/config.toml`. See [Routing Bundles](ROUTING_BUNDLES.md).
|
package/docs/RELEASE.md
CHANGED
|
@@ -25,7 +25,7 @@ The script enforces, in order:
|
|
|
25
25
|
1. branch is `main`, worktree is clean, `CHANGELOG.md` already has a committed `## [x.y.z]` section, and the tag does not exist;
|
|
26
26
|
2. the version is written into all four manifests plus `Cargo.lock`;
|
|
27
27
|
3. gates: `cargo test` (includes the manifest drift guard in `tests/e2e.rs`), `npm pack --dry-run`, and `scripts/security-local.sh` (betterleaks + trivy leak gate);
|
|
28
|
-
4. one mechanical commit `release x.y.z: <summary>`,
|
|
28
|
+
4. one mechanical commit `release x.y.z: <summary>`, an annotated `vx.y.z` tag carrying that summary, and a single push of branch plus tag.
|
|
29
29
|
|
|
30
30
|
Two independent gates back the script:
|
|
31
31
|
|
|
@@ -0,0 +1,32 @@
|
|
|
1
|
+
# Routing bundles
|
|
2
|
+
|
|
3
|
+
Planr Core is provider-neutral. It parses `.planr/agents.toml`, resolves routes into pick packets, records declared-versus-observed evidence, and safely previews or applies a strict RoutingBundle v1:
|
|
4
|
+
|
|
5
|
+
```bash
|
|
6
|
+
planr routing bundle inspect routing-bundle.json
|
|
7
|
+
planr routing bundle preview routing-bundle.json
|
|
8
|
+
planr routing bundle apply routing-bundle.json
|
|
9
|
+
```
|
|
10
|
+
|
|
11
|
+
Core accepts only allowlisted repository-local targets, verifies payload hashes, rejects absolute paths, traversal, symlinks, parent/child target collisions, conflicts, unsupported versions, and invalid payloads, and applies the validated set atomically. It never writes user configuration or files outside the repository.
|
|
12
|
+
|
|
13
|
+
A signed bundle is accepted only with an independent trust anchor supplied to every inspect, preview, or apply call:
|
|
14
|
+
|
|
15
|
+
```bash
|
|
16
|
+
planr routing bundle inspect signed-bundle.json \
|
|
17
|
+
--trusted-signer planr-maintainers \
|
|
18
|
+
--trusted-public-key-file /absolute/path/to/maintainer.pub
|
|
19
|
+
```
|
|
20
|
+
|
|
21
|
+
The bundle contains the signer id and signature, not a self-trusted public key. Both trust flags are required together; unsigned bundles require neither. An unsigned bundle also cannot label its evidence `verified` or `recommended`.
|
|
22
|
+
|
|
23
|
+
The `planr-routing` workspace package owns all volatile opinions: named policies, exact model ids, host bindings, generated role and skill files, capability probes, evaluation scenarios, signing, registry data, and the website catalog. A normal flow is:
|
|
24
|
+
|
|
25
|
+
```bash
|
|
26
|
+
planr-routing policy list
|
|
27
|
+
planr-routing compile balanced --host codex-openai --output routing-bundle.json
|
|
28
|
+
planr routing bundle preview routing-bundle.json
|
|
29
|
+
planr routing bundle apply routing-bundle.json
|
|
30
|
+
```
|
|
31
|
+
|
|
32
|
+
The package emits the same provider-neutral bundle contract for Codex, Claude Code, Cursor, and mixed-host configurations. Offline evaluation remains experimental; a recommendation requires complete authenticated live-host evidence. Missing authentication or missing effective model, effort, role, or context-fork evidence cannot pass.
|
package/docs/SKILLS.md
CHANGED
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
Planr ships agent-facing skill templates under `plugins/planr/skills/`.
|
|
4
4
|
|
|
5
|
-
The repository ships an installable plugin under `plugins/planr` for Codex
|
|
5
|
+
The repository ships an installable plugin under `plugins/planr` for Codex and Claude Code, while Cursor receives the same skills through `planr install cursor`. Marketplace manifests at the repo root (`.agents/plugins/marketplace.json`, `.claude-plugin/marketplace.json`) point at that subdirectory — Codex silently ignores marketplaces whose plugin source is the repo root itself. The shared package carries skills and Claude's independent workflow roles; optional model-specific host roles come only from a repository-local routing bundle, never from Planr Core or static fallbacks. The `planr` CLI must be installed separately (`brew install instructa/tap/planr`).
|
|
6
6
|
|
|
7
7
|
## Install As Plugin (preferred)
|
|
8
8
|
|
|
@@ -50,14 +50,14 @@ Stage skills (what the router and loop dispatch to; also directly invocable):
|
|
|
50
50
|
|
|
51
51
|
## Cheat Sheet
|
|
52
52
|
|
|
53
|
-
Default usage needs
|
|
53
|
+
Default usage needs one public entry point:
|
|
54
54
|
|
|
55
55
|
```text
|
|
56
56
|
$planr any request -> routed to the right stage skill from live map state
|
|
57
|
-
$planr-loop one feature -> loop work/verify/review/fix until done or budget exhausted
|
|
58
|
-
$planr-goal broad goal -> plan + map + durable contract + starter for /goal or manual loops
|
|
59
57
|
```
|
|
60
58
|
|
|
59
|
+
`$planr-goal` and `$planr-loop` are advanced stage surfaces selected by the router. A long-running goal is always prepared first; only the resulting real plan id is passed to the loop driver.
|
|
60
|
+
|
|
61
61
|
The stage order the router follows for a new app:
|
|
62
62
|
|
|
63
63
|
```text
|
|
@@ -81,10 +81,13 @@ Create the product plan, split an MVP build plan, check it, then build the Planr
|
|
|
81
81
|
Do not implement yet. End with the build plan id, critical lane, and first ready items.
|
|
82
82
|
```
|
|
83
83
|
|
|
84
|
-
Example autonomous feature loop:
|
|
84
|
+
Example autonomous feature loop (two separate prompts):
|
|
85
85
|
|
|
86
86
|
```text
|
|
87
|
-
Use $planr
|
|
87
|
+
Use $planr to prepare an autonomous goal for the weekly overview feature.
|
|
88
|
+
|
|
89
|
+
/goal Use $planr-loop on plan <plan-id>. The loop contract is stored in planr
|
|
90
|
+
context (tag: goal-contract).
|
|
88
91
|
|
|
89
92
|
Goal: ship the weekly overview feature. DONE when every in-scope map item is closed with
|
|
90
93
|
log evidence, all reviews are closed complete, and a live verification log shows the
|
|
@@ -103,7 +106,7 @@ Do not close the item until review is complete.
|
|
|
103
106
|
|
|
104
107
|
## Two Journeys: New Project vs. Existing Project
|
|
105
108
|
|
|
106
|
-
Both journeys use the same entry point (`$planr`
|
|
109
|
+
Both journeys use the same public entry point (`$planr`). What differs is the state the router finds, and what kind of plan the work gets.
|
|
107
110
|
|
|
108
111
|
### Journey 1 — start a project from an idea
|
|
109
112
|
|
|
@@ -120,7 +123,7 @@ Create a production-ready Habit Tracker web app plan. Create the product plan,
|
|
|
120
123
|
split an MVP build plan, check it, then build the Planr map. Do not implement yet.
|
|
121
124
|
```
|
|
122
125
|
|
|
123
|
-
The router runs the full stage order: product plan -> build plan -> map
|
|
126
|
+
The router runs the full stage order: product plan -> build plan -> map. From there it can select `$planr-work`, or prepare a plan-bound `$planr-loop` run.
|
|
124
127
|
|
|
125
128
|
### Journey 2 — mid-project: add a feature, refactor, or fix
|
|
126
129
|
|
|
@@ -138,12 +141,15 @@ What the router does with that, and why:
|
|
|
138
141
|
|
|
139
142
|
1. `$planr-plan` creates a new plan scoped to the feature (`planr plan new "Auth system" ...`), not a new project. Refine notes capture constraints from the existing codebase; the build plan's "existing leverage" field records what is reused instead of rebuilt.
|
|
140
143
|
2. `$planr-task-graph` extends the existing map: new items, plus `blocks` links to anything already on the map that must land first.
|
|
141
|
-
3. Execution is identical to journey 1: `$planr-loop` for autonomous, `$planr-work` / `$planr-review` for human-in-the-loop.
|
|
144
|
+
3. Execution is identical to journey 1: a plan-bound `$planr-loop` for autonomous work, or `$planr-work` / `$planr-review` for human-in-the-loop.
|
|
142
145
|
|
|
143
|
-
Or autonomous in
|
|
146
|
+
Or autonomous in two prompts:
|
|
144
147
|
|
|
145
148
|
```text
|
|
146
|
-
Use $planr
|
|
149
|
+
Use $planr to prepare an autonomous goal for the auth system.
|
|
150
|
+
|
|
151
|
+
/goal Use $planr-loop on plan <plan-id>. The loop contract is stored in planr
|
|
152
|
+
context (tag: goal-contract).
|
|
147
153
|
|
|
148
154
|
Goal: ship an auth system (email+password, sessions, protected routes).
|
|
149
155
|
DONE when every auth map item is closed with log evidence, all reviews are closed
|
|
@@ -164,67 +170,46 @@ Rules that hold in both journeys:
|
|
|
164
170
|
The CLI provisions the role files automatically — no manual copying:
|
|
165
171
|
|
|
166
172
|
```bash
|
|
167
|
-
planr project init "My Product" --client all # writes
|
|
168
|
-
planr
|
|
173
|
+
planr project init "My Product" --client all # writes standalone Claude and Cursor roles; Codex has no project roles
|
|
174
|
+
planr agents init # writes the provider-neutral .planr/agents.toml registry; it does not generate Codex roles
|
|
175
|
+
planr install claude # provisions Claude's independent roles
|
|
176
|
+
planr install cursor # provisions Cursor's independent roles and skills
|
|
169
177
|
```
|
|
170
178
|
|
|
171
|
-
|
|
179
|
+
Optional project-scoped model-routing files are generated by `planr-routing` bundles. Core workflow skills remain host-neutral, and bundle application never overwrites conflicts or writes user configuration.
|
|
172
180
|
|
|
173
181
|
Dispatches stay one line: `Use $planr-work on item <id>` and `Use $planr-review on item <id>`. The map and logs are the loop memory, so any iteration can resume from zero context.
|
|
174
182
|
|
|
175
183
|
## Install For Codex
|
|
176
184
|
|
|
177
|
-
|
|
178
|
-
|
|
179
|
-
```bash
|
|
180
|
-
mkdir -p ~/.codex/skills
|
|
181
|
-
cp -R plugins/planr/skills/* ~/.codex/skills/
|
|
182
|
-
```
|
|
183
|
-
|
|
184
|
-
If Planr was installed from an npm package that includes `skills/`, copy from the package location instead:
|
|
185
|
-
|
|
186
|
-
```bash
|
|
187
|
-
PLANR_PKG="$(npm root -g)/planr"
|
|
188
|
-
mkdir -p ~/.codex/skills
|
|
189
|
-
cp -R "$PLANR_PKG"/plugins/planr/skills/* ~/.codex/skills/
|
|
190
|
-
```
|
|
191
|
-
|
|
192
|
-
Do not present `npx planr` as the primary install path until the npm artifact ships platform-native Planr binaries. Today the normal user path is the GitHub Release installer; npm is a development and consumer-test wrapper.
|
|
193
|
-
|
|
194
|
-
Then run Codex from a repository where `planr` is installed and initialized:
|
|
185
|
+
Install the Codex plugin for all ten workflow skills, then initialize and install the project integration:
|
|
195
186
|
|
|
196
187
|
```bash
|
|
188
|
+
codex plugin marketplace add instructa/planr
|
|
189
|
+
codex plugin add planr@planr
|
|
197
190
|
planr project init "Example Product" --client codex
|
|
191
|
+
planr install codex
|
|
198
192
|
planr doctor --client codex
|
|
199
193
|
```
|
|
200
194
|
|
|
201
|
-
Codex
|
|
202
|
-
|
|
203
|
-
```bash
|
|
204
|
-
planr install codex --dry-run
|
|
205
|
-
planr prompt mcp --client codex
|
|
206
|
-
```
|
|
195
|
+
The CLI writes the project MCP snippet and hooks. It does not copy project skills or agents; those skills are plugin-owned, and Codex has no Planr project-agent contract. `--no-mcp` leaves hooks only, while `--no-mcp --no-hooks` writes neither integration artifact.
|
|
207
196
|
|
|
208
197
|
## Install For Claude Code
|
|
209
198
|
|
|
210
|
-
Claude Code
|
|
199
|
+
Install the Claude Code plugin for all ten workflow skills and its plugin worker/reviewer agents, then install the project integration:
|
|
211
200
|
|
|
212
|
-
```
|
|
213
|
-
|
|
214
|
-
|
|
215
|
-
cp plugins/planr/agents/*.md .claude/agents/
|
|
201
|
+
```text
|
|
202
|
+
/plugin marketplace add instructa/planr
|
|
203
|
+
/plugin install planr@planr
|
|
216
204
|
```
|
|
217
205
|
|
|
218
|
-
Then add MCP and the Planr workflow prompt to project instructions when needed:
|
|
219
|
-
|
|
220
206
|
```bash
|
|
221
207
|
planr project init "Example Product" --client claude
|
|
222
|
-
planr install claude
|
|
223
|
-
planr
|
|
224
|
-
planr prompt cli --client claude
|
|
208
|
+
planr install claude
|
|
209
|
+
planr doctor --client claude
|
|
225
210
|
```
|
|
226
211
|
|
|
227
|
-
|
|
212
|
+
The CLI writes project-scoped `.mcp.json`, standalone project worker/reviewer roles, and hooks. It does not copy project skills. `--no-mcp` retains the standalone roles and hooks; add `--no-hooks` to omit hooks.
|
|
228
213
|
|
|
229
214
|
## Install For Cursor
|
|
230
215
|
|
|
@@ -235,7 +220,7 @@ planr project init "Example Product" --client cursor
|
|
|
235
220
|
planr install cursor
|
|
236
221
|
```
|
|
237
222
|
|
|
238
|
-
`planr install cursor` writes `.cursor/mcp.json`, copies the ten skills to `.cursor/skills/`, provisions `.cursor/agents/planr-worker.md` and `planr-reviewer.md`, and prints a one-click deeplink for user-level MCP install.
|
|
223
|
+
`planr install cursor` writes `.cursor/mcp.json`, copies the ten skills to `.cursor/skills/`, provisions `.cursor/agents/planr-worker.md` and `planr-reviewer.md`, reconciles hooks, and prints a one-click deeplink for user-level MCP install. `planr install cursor --no-mcp` retains the agents, skills, and hooks while omitting MCP; add `--no-hooks` to omit hooks. Invoke the public router with `/planr` in Agent chat, and dispatch subagents with `/planr-worker` and `/planr-reviewer`. Use `planr serve --port 7526` and `planr prompt http --client cursor` if a Cursor workflow should inspect the local HTTP/review workspace. Subagent multitasking and worktree guidance: [Cursor](CURSOR.md).
|
|
239
224
|
|
|
240
225
|
## MCP-Only Clients
|
|
241
226
|
|