specrails-core 4.11.3 → 5.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (108) hide show
  1. package/README.md +96 -89
  2. package/bin/specrails-core.mjs +282 -39
  3. package/bin/tui-installer.mjs +117 -149
  4. package/commands/doctor.md +1 -1
  5. package/dist/installer/cli.js +13 -3
  6. package/dist/installer/cli.js.map +1 -1
  7. package/dist/installer/commands/doctor.js +487 -27
  8. package/dist/installer/commands/doctor.js.map +1 -1
  9. package/dist/installer/commands/framework.js +49 -7
  10. package/dist/installer/commands/framework.js.map +1 -1
  11. package/dist/installer/commands/init.js +443 -41
  12. package/dist/installer/commands/init.js.map +1 -1
  13. package/dist/installer/commands/update.js +51 -23
  14. package/dist/installer/commands/update.js.map +1 -1
  15. package/dist/installer/commands/v5-migration.js +119 -0
  16. package/dist/installer/commands/v5-migration.js.map +1 -0
  17. package/dist/installer/phases/framework-lifecycle.js +125 -0
  18. package/dist/installer/phases/framework-lifecycle.js.map +1 -0
  19. package/dist/installer/phases/install-config.js +160 -11
  20. package/dist/installer/phases/install-config.js.map +1 -1
  21. package/dist/installer/phases/manifest.js +29 -8
  22. package/dist/installer/phases/manifest.js.map +1 -1
  23. package/dist/installer/phases/prereqs.js +57 -3
  24. package/dist/installer/phases/prereqs.js.map +1 -1
  25. package/dist/installer/phases/provider-detect.js +116 -6
  26. package/dist/installer/phases/provider-detect.js.map +1 -1
  27. package/dist/installer/phases/scaffold.js +1217 -117
  28. package/dist/installer/phases/scaffold.js.map +1 -1
  29. package/dist/installer/runtime/kimi.js +255 -0
  30. package/dist/installer/runtime/kimi.js.map +1 -0
  31. package/dist/installer/util/paths.js +12 -0
  32. package/dist/installer/util/paths.js.map +1 -1
  33. package/dist/installer/util/registry.js +234 -14
  34. package/dist/installer/util/registry.js.map +1 -1
  35. package/docs/README.md +1 -0
  36. package/docs/deployment.md +6 -7
  37. package/docs/getting-started.md +11 -7
  38. package/docs/installation.md +34 -16
  39. package/docs/plugin-architecture.md +11 -8
  40. package/docs/updating.md +21 -3
  41. package/docs/user-docs/cli-reference.md +43 -22
  42. package/docs/user-docs/codex-vs-claude-code.md +11 -9
  43. package/docs/user-docs/faq.md +1 -1
  44. package/docs/user-docs/getting-started-codex.md +5 -8
  45. package/docs/user-docs/getting-started-kimi.md +423 -0
  46. package/docs/user-docs/installation.md +49 -14
  47. package/docs/user-docs/quick-start.md +11 -8
  48. package/docs/windows.md +29 -4
  49. package/integration-contract.json +85 -13
  50. package/package.json +9 -5
  51. package/schemas/profile.v1.json +68 -6
  52. package/templates/agents/sr-architect.md +30 -0
  53. package/templates/agents/sr-developer.md +21 -8
  54. package/templates/agents/sr-reviewer.md +44 -31
  55. package/templates/codex-skills/batch-implement/SKILL.md +9 -32
  56. package/templates/codex-skills/implement/SKILL.md +61 -143
  57. package/templates/codex-skills/rails/sr-architect/SKILL.md +38 -20
  58. package/templates/codex-skills/rails/sr-developer/SKILL.md +29 -10
  59. package/templates/codex-skills/rails/sr-reviewer/SKILL.md +21 -10
  60. package/templates/commands/specrails/doctor.md +1 -1
  61. package/templates/commands/specrails/implement.md +117 -288
  62. package/templates/commands/specrails/memory-inspect.md +6 -4
  63. package/templates/commands/specrails/propose-spec.md +1 -1
  64. package/templates/commands/specrails/refactor-recommender.md +8 -51
  65. package/templates/commands/specrails/retry.md +12 -48
  66. package/templates/commands/specrails/telemetry.md +1 -1
  67. package/templates/gemini-commands/implement.toml +9 -0
  68. package/templates/kimi/specrails/run-skill.mjs +3005 -0
  69. package/templates/kimi/specrails/vendor/js-yaml/LICENSE +21 -0
  70. package/templates/kimi/specrails/vendor/js-yaml/NOTICE.md +16 -0
  71. package/templates/kimi/specrails/vendor/js-yaml/js-yaml.mjs +3856 -0
  72. package/templates/profiles/default.json +5 -18
  73. package/templates/profiles/kimi-default.json +15 -0
  74. package/commands/enrich.md +0 -1456
  75. package/templates/agents/sr-backend-developer.md +0 -91
  76. package/templates/agents/sr-backend-reviewer.md +0 -152
  77. package/templates/agents/sr-doc-sync.md +0 -247
  78. package/templates/agents/sr-frontend-developer.md +0 -85
  79. package/templates/agents/sr-frontend-reviewer.md +0 -145
  80. package/templates/agents/sr-merge-resolver.md +0 -195
  81. package/templates/agents/sr-performance-reviewer.md +0 -186
  82. package/templates/agents/sr-product-analyst.md +0 -36
  83. package/templates/agents/sr-product-manager.md +0 -148
  84. package/templates/agents/sr-security-reviewer.md +0 -191
  85. package/templates/agents/sr-test-writer.md +0 -176
  86. package/templates/codex-skills/enrich/SKILL.md +0 -191
  87. package/templates/codex-skills/merge-resolve/SKILL.md +0 -88
  88. package/templates/codex-skills/rails/sr-backend-developer/SKILL.md +0 -93
  89. package/templates/codex-skills/rails/sr-backend-reviewer/SKILL.md +0 -120
  90. package/templates/codex-skills/rails/sr-doc-sync/SKILL.md +0 -124
  91. package/templates/codex-skills/rails/sr-frontend-developer/SKILL.md +0 -106
  92. package/templates/codex-skills/rails/sr-frontend-reviewer/SKILL.md +0 -111
  93. package/templates/codex-skills/rails/sr-merge-resolver/SKILL.md +0 -156
  94. package/templates/codex-skills/rails/sr-performance-reviewer/SKILL.md +0 -109
  95. package/templates/codex-skills/rails/sr-product-analyst/SKILL.md +0 -85
  96. package/templates/codex-skills/rails/sr-product-manager/SKILL.md +0 -131
  97. package/templates/codex-skills/rails/sr-security-reviewer/SKILL.md +0 -121
  98. package/templates/codex-skills/rails/sr-test-writer/SKILL.md +0 -115
  99. package/templates/commands/specrails/auto-propose-backlog-specs.md +0 -312
  100. package/templates/commands/specrails/enrich.md +0 -1456
  101. package/templates/commands/specrails/get-backlog-specs.md +0 -226
  102. package/templates/commands/specrails/merge-resolve.md +0 -172
  103. package/templates/commands/specrails/reconfig.md +0 -80
  104. package/templates/commands/specrails/vpc-drift.md +0 -405
  105. package/templates/commands/test.md +0 -58
  106. package/templates/personas/persona.md +0 -43
  107. package/templates/personas/the-maintainer.md +0 -98
  108. package/templates/settings/perf-thresholds.yml +0 -25
@@ -1,5 +1,5 @@
1
1
  {
2
- "schemaVersion": "3.0",
2
+ "schemaVersion": "3.2",
3
3
  "providers": {
4
4
  "claude": {
5
5
  "enrichCommand": "/specrails:enrich",
@@ -22,6 +22,59 @@
22
22
  "enrichArgs": ["exec", "run enrich"],
23
23
  "enrichFromConfigArgs": ["exec", "run enrich --from-config"]
24
24
  }
25
+ },
26
+ "gemini": {
27
+ "enrichCommand": "/specrails:enrich",
28
+ "enrichArgs": ["--from-config"],
29
+ "updateCommand": "/specrails:enrich",
30
+ "updateArgs": ["--update"],
31
+ "cli": {
32
+ "initArgs": [],
33
+ "enrichArgs": ["-p", "/specrails:enrich", "--output-format", "stream-json"],
34
+ "enrichFromConfigArgs": ["-p", "/specrails:enrich --from-config", "--output-format", "stream-json"]
35
+ }
36
+ },
37
+ "kimi": {
38
+ "enrichCommand": "/skill:specrails-enrich",
39
+ "enrichArgs": ["--from-config"],
40
+ "updateCommand": "/skill:specrails-enrich",
41
+ "updateArgs": ["--update"],
42
+ "cli": {
43
+ "binary": "node",
44
+ "providerBinary": "kimi",
45
+ "skillRunner": ".kimi-code/specrails/run-skill.mjs",
46
+ "skillMaterialization": "kimi-0.27-user-slash-prompt",
47
+ "nestedSkillInvocation": "Kimi built-in Skill tool with { skill, args }; never literal /skill text",
48
+ "modelIdPattern": "^[A-Za-z0-9][A-Za-z0-9._/:-]{0,127}$",
49
+ "sessionIdPattern": "^(?!\\.{1,2}$)[A-Za-z0-9._-]{1,128}$",
50
+ "roleWave": {
51
+ "path": ".specrails/kimi-role-wave.json",
52
+ "schema": "{ run: safe-id, roles: Array<{ key: safe-id, skill: direct-child-id, model: safe-model-id, profile: \"inherit\" | safe-id, args: string, workspace: \"current\" | \"worktree:<safe-id>\" }> } with no extra keys",
53
+ "maxBytes": 1048576,
54
+ "maxRoles": 32,
55
+ "transport": "structured WriteFile followed by one static foreground --role-wave-file command; the one-shot file is deleted before setup and every role is awaited. Inspection, merge, and cleanup use the separate static --role-wave-status, --role-merge-file, and --role-wave-cleanup modes",
56
+ "workspaces": "current roles receive unique execution directories with SPECRAILS_REPO_DIR set to the target repository; isolated roles create/reuse detached git worktrees from a synthetic baseline that includes the starting tracked/untracked overlay but excludes provider/runtime control files, with SPECRAILS_REPO_DIR set to that worktree",
57
+ "manifest": ".specrails/kimi-role-worktrees/<run>.json records the immutable synthetic base commit, source head, worktree ids, and role key to execution/repository paths",
58
+ "status": "--role-wave-status <run> validates the persisted manifest/ref/worktrees and emits one specrails.merge.inventory frame before any downstream merge",
59
+ "merge": ".specrails/kimi-role-merge.json is the only accepted merge request path; { run, actions: Array<{ worktree: safe-id, path: safe-relative-path, operation: \"copy\" | \"delete\" }> } applies an explicit A/M/D inventory through --role-merge-file without attributing the starting overlay or provider control files",
60
+ "cleanup": "--role-wave-cleanup <run> removes owned worktrees, temporary execution roots, persisted manifest, and synthetic baseline ref, then emits specrails.role.cleanup",
61
+ "output": "newline-delimited attributed specrails.role.workspace, specrails.role.event/output, specrails.role.completed, specrails.merge.inventory, specrails.merge.applied, and specrails.role.cleanup frames; aggregate role-wave exit is nonzero when any role fails"
62
+ },
63
+ "stableEngineEnv": {
64
+ "KIMI_CODE_EXPERIMENTAL_FLAG": null,
65
+ "KIMI_DISABLE_CRON": "1",
66
+ "KIMI_CODE_NO_AUTO_UPDATE": "1",
67
+ "KIMI_MODEL_THINKING_EFFORT": "low|high|max only for kimi-code/k3; omitted K3 effort preserves Kimi's documented high default; unset for every other model"
68
+ },
69
+ "windowsPromptTransport": "For the standard npm kimi.cmd/bat shim, prompt bytes travel over stdin to a fixed Node bootstrap which restores process.argv before importing Kimi; native executables fail above a 30000 UTF-16 command-line budget.",
70
+ "initialActivationTelemetry": "Visible prompt parity only; the external materializer cannot emit Kimi-private skill.activated/origin telemetry.",
71
+ "cancellation": "Single-skill mode forwards SIGINT/SIGTERM/SIGHUP to its direct Kimi child. Role-wave mode forwards each termination signal to every live Kimi child and waits for aggregate completion; the embedding host remains responsible for platform process-tree teardown.",
72
+ "initArgs": [],
73
+ "enrichArgs": [".kimi-code/specrails/run-skill.mjs", "--skill", "specrails-enrich", "--model", "k3"],
74
+ "enrichFromConfigArgs": [".kimi-code/specrails/run-skill.mjs", "--skill", "specrails-enrich", "--model", "k3", "--args", "--from-config"],
75
+ "resumeArgs": ["--session=<session-id>"],
76
+ "notes": "CLI-only integration. enrichCommand/updateCommand are interactive Kimi TUI syntax only. Headless callers must execute binary + enrichArgs; plain `kimi -p \"/skill:...\"` is literal prompt text in Kimi 0.27 and does not activate a skill. The managed Node runner materializes the upstream user-slash prompt, then launches external Kimi with stream-json and no shell. Generated multi-role workflows use one bounded role-wave file so context never enters shell source and parallel roles cannot race on request paths. Do not start kimi web/server; authenticate once with `kimi login`."
77
+ }
25
78
  }
26
79
  },
27
80
  "tiers": {
@@ -43,13 +96,13 @@
43
96
  "version": 1,
44
97
  "fields": {
45
98
  "version": "number — schema version, currently 1",
46
- "provider": "string — claude | codex | gemini | auto",
99
+ "provider": "string — claude | codex | gemini | kimi",
47
100
  "tier": "string — full | quick",
48
- "agents.selected": "string[] — agent names to install",
49
- "agents.excluded": "string[] — agent names to skip",
50
- "models.preset": "string — balanced | budget | max",
51
- "models.defaults.model": "string — sonnet | opus | haiku (overrides preset)",
52
- "models.overrides": "Record<string, string> — per-agent model (highest priority)"
101
+ "agents.selected": "string[] — unique lowercase kebab-case agent ids to install (1-64 characters)",
102
+ "agents.excluded": "string[] — unique lowercase kebab-case agent ids to skip; must not overlap agents.selected",
103
+ "models.preset": "string — balanced | budget | max; resolved within the selected provider catalog",
104
+ "models.defaults.model": "string — exact provider model id or configured alias (overrides preset; Kimi: 1-128 characters matching [A-Za-z0-9][A-Za-z0-9._/:-]*; default: k3)",
105
+ "models.overrides": "Record<safe-agent-id, string> — exact per-agent provider model ids or configured aliases with the same provider-specific validation (highest priority)"
53
106
  }
54
107
  },
55
108
  "checkpoints": {
@@ -63,19 +116,38 @@
63
116
  },
64
117
  "modelPresets": {
65
118
  "balanced": {
66
- "description": "Flagship models for architects and PMs, efficient models for others",
119
+ "scope": "claude",
120
+ "description": "Legacy Claude preset view retained for existing consumers. Other providers resolve this preset through providerModelCatalogs.",
67
121
  "defaults": { "model": "sonnet" },
68
- "overrides": { "sr-architect": "opus", "sr-product-manager": "opus" }
122
+ "overrides": {}
69
123
  },
70
124
  "budget": {
71
- "description": "All agents use the most cost-efficient model",
125
+ "scope": "claude",
126
+ "description": "Legacy Claude preset view retained for existing consumers. Other providers resolve this preset through providerModelCatalogs.",
72
127
  "defaults": { "model": "haiku" },
73
128
  "overrides": {}
74
129
  },
75
130
  "max": {
76
- "description": "All agents use the most capable model",
77
- "defaults": { "model": "opus" },
78
- "overrides": {}
131
+ "scope": "claude",
132
+ "description": "Legacy Claude preset view retained for existing consumers. Other providers resolve this preset through providerModelCatalogs.",
133
+ "defaults": { "model": "sonnet" },
134
+ "overrides": { "sr-architect": "opus", "sr-product-manager": "opus" }
135
+ }
136
+ },
137
+ "providerModelCatalogs": {
138
+ "kimi": {
139
+ "default": "k3",
140
+ "models": ["k3", "kimi-for-coding", "kimi-for-coding-highspeed"],
141
+ "presets": {
142
+ "balanced": { "defaults": { "model": "k3" }, "overrides": {} },
143
+ "budget": { "defaults": { "model": "k3" }, "overrides": {} },
144
+ "max": { "defaults": { "model": "k3" }, "overrides": {} }
145
+ },
146
+ "cliAliasPrefix": "kimi-code/",
147
+ "reasoningEfforts": {
148
+ "k3": ["low", "high", "max"]
149
+ },
150
+ "note": "Install config and profiles retain exact ids or safe custom aliases. Preset names do not imply Claude aliases for Kimi. Process launch prefixes only the three documented official short ids with kimi-code/; every custom alias that matches the published model grammar passes through unchanged."
79
151
  }
80
152
  },
81
153
  "legacyCompat": {
package/package.json CHANGED
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "name": "specrails-core",
3
- "version": "4.11.3",
4
- "description": "AI agent workflow system for Claude Code — installs 12 specialized agents, orchestration commands, and persona-driven product discovery into any repository",
3
+ "version": "5.0.0",
4
+ "description": "Provider-independent AI agent workflow system for Claude Code, Codex, Gemini CLI, and Kimi Code",
5
5
  "type": "module",
6
6
  "bin": {
7
7
  "specrails-core": "bin/specrails-core.mjs"
@@ -17,7 +17,7 @@
17
17
  "pinned-versions.json"
18
18
  ],
19
19
  "engines": {
20
- "node": ">=20.0.0"
20
+ "node": ">=20.19.0"
21
21
  },
22
22
  "repository": {
23
23
  "type": "git",
@@ -30,6 +30,10 @@
30
30
  "workflow",
31
31
  "developer-tools",
32
32
  "anthropic",
33
+ "kimi",
34
+ "moonshot-ai",
35
+ "gemini",
36
+ "codex",
33
37
  "llm",
34
38
  "ai",
35
39
  "automation",
@@ -53,9 +57,9 @@
53
57
  "build": "tsc -p tsconfig.json",
54
58
  "build:watch": "tsc -p tsconfig.json --watch",
55
59
  "typecheck": "tsc -p tsconfig.test.json --noEmit",
56
- "test": "npm run typecheck && vitest run",
60
+ "test": "npm run build && npm run typecheck && vitest run",
57
61
  "test:watch": "vitest",
58
- "test:coverage": "vitest run --coverage",
62
+ "test:coverage": "npm run build && vitest run --coverage",
59
63
  "dogfood": "npm run build && node bin/specrails-core.mjs init --yes",
60
64
  "prepack": "npm run build"
61
65
  },
@@ -23,14 +23,19 @@
23
23
  "maxLength": 512,
24
24
  "description": "Optional short summary of when to use this profile."
25
25
  },
26
+ "provider": {
27
+ "type": "string",
28
+ "enum": ["claude", "codex", "gemini", "kimi"],
29
+ "description": "Optional provider binding. Omitted profiles retain the legacy Claude model-alias contract. Provider-bound profiles retain exact model identifiers; Kimi additionally enforces its safe process-boundary grammar."
30
+ },
26
31
  "orchestrator": {
27
32
  "type": "object",
28
33
  "required": ["model"],
29
34
  "additionalProperties": false,
30
35
  "properties": {
31
36
  "model": {
32
- "$ref": "#/$defs/modelAlias",
33
- "description": "Model used to run the top-level implement orchestrator."
37
+ "$ref": "#/$defs/providerModelId",
38
+ "description": "Provider model identifier used to run the top-level implement orchestrator."
34
39
  }
35
40
  }
36
41
  },
@@ -38,7 +43,7 @@
38
43
  "type": "array",
39
44
  "minItems": 1,
40
45
  "items": { "$ref": "#/$defs/agentEntry" },
41
- "description": "Ordered chain of agents that participate in the pipeline when this profile is active. Every profile (default + custom) must include the baseline trio (sr-architect, sr-developer, sr-reviewer) — the pipeline depends on all three. sr-merge-resolver and every other agent are optional add-ons selected at install time. Custom profiles add optional agents on top of the baseline.",
46
+ "description": "Ordered chain of agents that participate in the pipeline when this profile is active. Every profile (default + custom) must include the baseline trio (sr-architect, sr-developer, sr-reviewer) — the pipeline depends on all three, and they are the only agents the installer ships. Additional agents are user-authored 'custom-*' agents added on top of the baseline; a profile entry whose agent file does not exist is warned and skipped at run time.",
42
47
  "allOf": [
43
48
  {
44
49
  "description": "Required baseline agent: sr-architect",
@@ -73,12 +78,69 @@
73
78
  "items": { "$ref": "#/$defs/routingRule" }
74
79
  }
75
80
  },
81
+ "allOf": [
82
+ {
83
+ "if": {
84
+ "anyOf": [
85
+ { "not": { "required": ["provider"] } },
86
+ {
87
+ "required": ["provider"],
88
+ "properties": { "provider": { "const": "claude" } }
89
+ }
90
+ ]
91
+ },
92
+ "then": {
93
+ "properties": {
94
+ "orchestrator": {
95
+ "properties": { "model": { "$ref": "#/$defs/modelAlias" } }
96
+ },
97
+ "agents": {
98
+ "items": {
99
+ "properties": { "model": { "$ref": "#/$defs/modelAlias" } }
100
+ }
101
+ }
102
+ }
103
+ }
104
+ },
105
+ {
106
+ "if": {
107
+ "required": ["provider"],
108
+ "properties": { "provider": { "const": "kimi" } }
109
+ },
110
+ "then": {
111
+ "properties": {
112
+ "orchestrator": {
113
+ "properties": { "model": { "$ref": "#/$defs/kimiModelId" } }
114
+ },
115
+ "agents": {
116
+ "items": {
117
+ "properties": { "model": { "$ref": "#/$defs/kimiModelId" } }
118
+ }
119
+ }
120
+ }
121
+ }
122
+ }
123
+ ],
76
124
  "$defs": {
77
125
  "modelAlias": {
78
126
  "type": "string",
79
127
  "enum": ["sonnet", "opus", "haiku"],
80
128
  "description": "Accepted model alias. Resolved to the current concrete model ID by the pipeline."
81
129
  },
130
+ "providerModelId": {
131
+ "type": "string",
132
+ "minLength": 1,
133
+ "maxLength": 256,
134
+ "pattern": "^\\S(?:.*\\S)?$",
135
+ "description": "Exact provider model identifier or user-configured alias. Core retains this value verbatim and applies documented provider-specific validation/normalization."
136
+ },
137
+ "kimiModelId": {
138
+ "type": "string",
139
+ "minLength": 1,
140
+ "maxLength": 128,
141
+ "pattern": "^[A-Za-z0-9][A-Za-z0-9._/:-]*$",
142
+ "description": "Kimi model identifier accepted at the shell-free process boundary."
143
+ },
82
144
  "agentEntry": {
83
145
  "type": "object",
84
146
  "required": ["id"],
@@ -87,11 +149,11 @@
87
149
  "id": {
88
150
  "type": "string",
89
151
  "pattern": "^(sr|custom)-[a-z0-9][a-z0-9-]*$",
90
- "description": "Agent identifier. MUST correspond to a file at `.claude/agents/<id>.md`."
152
+ "description": "Agent identifier. MUST correspond to the selected provider's role artifact."
91
153
  },
92
154
  "model": {
93
- "$ref": "#/$defs/modelAlias",
94
- "description": "Model override for this agent when this profile is active. When omitted, the agent's frontmatter `model:` is used as a fallback."
155
+ "$ref": "#/$defs/providerModelId",
156
+ "description": "Exact provider model override for this agent. When omitted, the provider default is used."
95
157
  },
96
158
  "required": {
97
159
  "type": "boolean",
@@ -51,6 +51,10 @@ Do not proceed with any design work, file reading, or artifact creation until sp
51
51
 
52
52
  Your working directory may NOT be the user's source repository. The user's source code, `openspec/**`, `.claude/rules/`, and `.git` all live under **`${SPECRAILS_REPO_DIR:-.}`** (the spawner sets the env var to the repo path; unset defaults to `.`, i.e. byte-identical to a classic in-repo run). Read specs from `${SPECRAILS_REPO_DIR:-.}/openspec/...`, scan conventions under `${SPECRAILS_REPO_DIR:-.}/.claude/rules/`, and resolve every compatibility-surface read (CLI/commands/agents/config) relative to `${SPECRAILS_REPO_DIR:-.}`.
53
53
 
54
+ ## Deterministic repo map (read before exploring)
55
+
56
+ If the environment variable `SPECRAILS_REPO_MAP_PATH` is set and points to a readable file, **Read that file FIRST**, before any codebase exploration. It is a deterministic map of the repository (packages, ecosystems, rough sizes, README excerpt) generated by the spawner at zero AI cost. Use it to orient your exploration — do NOT spend turns on top-level discovery (`ls` at the root, locating packages, reading the README for structure). When the variable is unset, explore as normal.
57
+
54
58
  ## Core Responsibilities
55
59
 
56
60
  When invoked by the orchestrator with a specName argument, you must execute the following steps in order:
@@ -158,6 +162,32 @@ After producing the task breakdown and before finalizing output:
158
162
 
159
163
  This phase is mandatory. Do not skip it even if the change appears purely internal.
160
164
 
165
+ ### 7. Emit Design Confidence (MANDATORY)
166
+
167
+ Implementation is the expensive phase of this pipeline — it must only run on a design you actually trust. After completing the design and task breakdown, score your own confidence and write it to:
168
+
169
+ ```
170
+ ${SPECRAILS_REPO_DIR:-.}/openspec/changes/<name>/design-confidence.json
171
+ ```
172
+
173
+ Required fields:
174
+
175
+ - `schema_version`: always `"1"`
176
+ - `change`: kebab-case change name
177
+ - `agent`: always `"architect"`
178
+ - `scored_at`: current ISO 8601 timestamp
179
+ - `confidence`: `"high"` | `"medium"` | `"low"`
180
+ - `reason`: 1–2 sentences justifying the level — concrete, not boilerplate
181
+ - `blocking_question`: when `confidence` is `"low"`, the **single most blocking unknown** phrased as one focused question a human can answer — nothing else. Otherwise `null`.
182
+
183
+ Rubric:
184
+
185
+ - **high** — the code evidence is conclusive: you located the exact files/identifiers, the design is unambiguous, and the tasks follow directly from it.
186
+ - **medium** — the design is likely correct but rests on one non-obvious assumption you could not fully verify. Name that assumption in `reason`.
187
+ - **low** — multiple plausible interpretations or designs exist and you cannot choose between them without information you don't have (missing requirement, ambiguous intent, contradictory specs). Do NOT pad the design to look confident — a `low` with a sharp `blocking_question` is a SUCCESSFUL architect output: it saves the entire implementation cost of building the wrong thing.
188
+
189
+ Never inflate the level. The orchestrator halts implementation on `low` and relays your `blocking_question` to the human — that is the designed outcome, not a failure.
190
+
161
191
  ## Output Format
162
192
 
163
193
  When analyzing spec changes, produce your output in this structure:
@@ -140,9 +140,19 @@ This gate is non-negotiable. Phase 4 is unreachable until every checkbox in task
140
140
 
141
141
  **For each unit of functionality within the apply cycle, follow this TDD cycle:**
142
142
 
143
- 1. **RED** — Write a failing test that describes the expected behavior. Run the test. Confirm it fails for the right reason.
144
- 2. **GREEN** — Write the minimum production code to make the test pass. Run the test. Confirm it passes.
145
- 3. **REFACTOR** — Clean up the code while keeping all tests green. Run all tests after refactoring.
143
+ 1. **RED** — Write a failing test that describes the expected behavior. Run **only that test file** (scoped run). Confirm it fails for the right reason.
144
+ 2. **GREEN** — Write the minimum production code to make the test pass. Re-run **only that test file**. Confirm it passes.
145
+ 3. **REFACTOR** — Clean up the code while keeping tests green. Re-run **the test files covering the files you touched** — not the whole suite. The full suite runs exactly once, in Phase 4 — running it after every task multiplies wall-clock time without catching anything Phase 4 won't.
146
+
147
+ ## Test-Execution Economy (MANDATORY)
148
+
149
+ Test runs are the single largest cost of this pipeline. The contract:
150
+
151
+ - **Inside task cycles (Phase 3): scoped runs only.** Invoke the runner with an explicit path/filter — `npx vitest run <file>`, `npx jest <file>`, `pytest <file>`, `go test ./<pkg>`, `./gradlew :<module>:test --tests <Class>`, etc. Derive the scoped form from the project's full test command. **Never run the full suite inside a task cycle.**
152
+ - **The full suite runs exactly ONCE** — at your Phase 4 validation gate, after every task is `- [x]`. It does not run per task, per file, or "just to be safe".
153
+ - **When a scoped run fails**, extract only the failing test names and the relevant error excerpt (≤50 lines) into your reasoning. Never re-paste a full runner log.
154
+ - **Loop detection**: if you run the same command 3 times without an intervening code change and results are inconsistent, STOP running it — state your hypothesis and change the code or the test instead.
155
+ - **File re-read discipline**: a file you already read is in your context. Before reading any file a second time, write one sentence stating what you already learned from it — then only re-read if it changed since.
146
156
 
147
157
  **TDD rules:**
148
158
  - Never write production code without a corresponding test
@@ -170,11 +180,14 @@ Follow the project architecture strictly:
170
180
 
171
181
  **Prerequisite: Phase 4 is only reachable if the Phase 3 checkbox verification gate passed** — meaning every task in `${SPECRAILS_REPO_DIR:-.}/openspec/changes/<specName>/tasks.md` is marked `- [x]`. If any `- [ ]` items remain, return to Phase 3.
172
182
 
173
- **All tests MUST pass before you hand off to the reviewer. This is a hard gate — do not proceed if any test fails.**
183
+ **All tests MUST pass before you hand off to the reviewer. This is a hard gate — do not hand off with known failures.**
184
+
185
+ This phase is the pipeline's **single full verification pass** — the inner TDD loop stayed scoped precisely so this one can be exhaustive.
174
186
 
175
- - Run the **full CI-equivalent verification suite** (see below)
176
- - If any test fails, fix the issue and re-run ALL tests
177
- - Repeat until all tests pass — there is no maximum number of attempts
187
+ - Run the **full CI-equivalent verification suite** (see below) — this is the ONE full-suite run of your phase
188
+ - If anything fails: fix it, then re-run **only the failing test files / failing check** — not the whole suite
189
+ - You have a budget of **2 fix cycles**. After the fixes converge, run the full suite ONE final time to confirm
190
+ - If failures persist after the budget: **HALT and report honestly** — list every failing test verbatim to the orchestrator. Do NOT keep looping, do NOT weaken or skip tests to force green, do NOT hand off silently
178
191
  - Review each file for adherence to conventions
179
192
  - Ensure all imports are correct and no circular dependencies exist
180
193
  - Verify type annotations are complete
@@ -183,7 +196,7 @@ Follow the project architecture strictly:
183
196
 
184
197
  ## CI-Equivalent Verification Suite
185
198
 
186
- You MUST run ALL of these checks after implementation. These match the CI pipeline exactly:
199
+ You MUST run ALL of these checks after implementation — **once, at the Phase 4 gate** (see Test-Execution Economy). These match the CI pipeline exactly:
187
200
 
188
201
  {{CI_COMMANDS_FULL}}
189
202
 
@@ -52,42 +52,36 @@ Your working directory may NOT be the user's source repository. The user's sourc
52
52
  You are the last line of defense between developer output and a PR. You:
53
53
  1. **Verify TDD compliance** — every piece of production code must have corresponding tests
54
54
  2. **Verify spec completeness** — every requirement from the architect's spec must be implemented
55
- 3. Run every check that CI runs — in the exact same way
55
+ 3. Verify the change is green — **scoped-first**: the developer's Phase 4 already ran the full CI-equivalent suite; you re-verify the changed surface, and run the full suite yourself only when your own fixes make it necessary (see Verification policy)
56
56
  4. Fix any failures you find (up to 3 attempts per issue)
57
57
  5. Verify code quality and consistency across all changes
58
58
  6. Report what you found and fixed
59
59
 
60
60
  ## CI/CD Pipeline Equivalence
61
61
 
62
- The CI pipeline runs these checks. You MUST run ALL of them in this exact order:
62
+ The CI pipeline runs these checks, in this exact order:
63
63
 
64
64
  {{CI_COMMANDS_FULL}}
65
65
 
66
- ## Known CI vs Local Gaps
67
-
68
- These are the most common reasons code passes locally but fails in CI:
69
-
70
- {{CI_KNOWN_GAPS}}
71
-
72
- ## Layer Review Findings (injected at runtime by orchestrator)
66
+ ## Verification policy (scoped-first)
73
67
 
74
- The orchestrator runs specialized layer reviewers in parallel before you launch. Their reports are injected here. A value of `"SKIPPED"` means no files of that layer type were in the changeset.
68
+ The developer hands off ONLY after a green full CI-equivalent pass (their Phase 4 hard gate). Re-running the entire suite on an untouched tree re-buys information the pipeline already has — at full wall-clock price. Your verification is therefore **scoped-first**:
75
69
 
76
- **These are NOT `/specrails:enrich` placeholders. They use `[injected]` notation, not `{{...}}` notation.** The `[injected]` markers below are replaced by the actual report text when the orchestrator launches you.
70
+ 1. **Always run the cheap whole-repo static checks** (type-check, lint — the fast entries of the CI list above, in CI order).
71
+ 2. **Run the tests SCOPED to the diff**: the test files covering every changed source file, via per-file invocation (`npx vitest run <file>`, `pytest <file>`, `cargo test <module>`, …). Widen the scope when the change touches shared/core modules whose blast radius you cannot bound.
72
+ 3. **Full suite — run it yourself only when warranted**: you modified production code during the review, the diff touches build/config/test infrastructure, or the scoped runs surfaced a failure whose blast radius is unclear. In that case finish with ONE clean full pass before handoff — never interleave repeated full passes between fixes.
73
+ 4. If you changed nothing and the scoped runs are green, the developer's full pass stands as the pipeline's verification of record — say so in the report instead of re-running it.
77
74
 
78
- FRONTEND_REVIEW_REPORT:
79
- [injected]
80
-
81
- BACKEND_REVIEW_REPORT:
82
- [injected]
75
+ ## Known CI vs Local Gaps
83
76
 
84
- SECURITY_REVIEW_REPORT:
85
- [injected]
77
+ These are the most common reasons code passes locally but fails in CI:
86
78
 
87
- ---
79
+ {{CI_KNOWN_GAPS}}
88
80
 
89
81
  ## Review Checklist
90
82
 
83
+ You are the single reviewer for this change. There are no separate layer reviewers — frontend, backend, security, and performance concerns are all your responsibility in this pass. Weight each dimension by what the changeset actually touches.
84
+
91
85
  After running CI checks, also review for:
92
86
 
93
87
  ### TDD Compliance (mandatory)
@@ -114,11 +108,22 @@ After running CI checks, also review for:
114
108
  - Import style matches the rest of the codebase
115
109
  - Error handling patterns are consistent
116
110
 
111
+ ### Security (scale to what the change touches)
112
+ - No secrets, tokens, or credentials committed
113
+ - User-controlled input is validated and, where interpolated into queries/commands/paths, properly escaped or parameterized
114
+ - No new injection, path-traversal, or SSRF surface introduced
115
+ - AuthZ/authN checks are present on new endpoints or privileged operations
116
+
117
+ ### Performance (scale to what the change touches)
118
+ - No obvious N+1 queries or unbounded loops over user-controlled input
119
+ - Expensive work is not added to hot paths without justification
120
+ - Large allocations, unbounded caches, and blocking I/O on async paths are flagged
121
+
117
122
  ## Workflow
118
123
 
119
- 1. **Run all CI checks** (all layers, in the exact order CI runs them)
120
- 2. **If anything fails**: Fix it, then re-run ALL checks from scratch (not just the failing one)
121
- 3. **Repeat** up to 3 fix-and-verify cycles
124
+ 1. **Run the scoped-first verification** (see Verification policy: static checks + diff-scoped tests, in CI order)
125
+ 2. **If anything fails**: Fix it, then re-run **only the failed check, scoped to the failing files** where the runner supports it (`npx vitest run <file>`, `pytest <file>`, lint on the changed files) — never the entire ordered list after every individual fix — and escalate to a full pass per the policy
126
+ 3. **Repeat** up to 3 fix-and-verify cycles; when any cycle changed code, finish with ONE clean full CI-equivalent pass
122
127
  4. **Report** a summary of what passed, what failed, and what you fixed
123
128
  5. **Task Completion Gate** — Before archiving, verify all tasks are complete:
124
129
  - Read `${SPECRAILS_REPO_DIR:-.}/openspec/changes/<specName>/tasks.md`
@@ -197,26 +202,34 @@ When done, produce this report:
197
202
  ### Issues Fixed
198
203
  - [list of issues found and how they were fixed]
199
204
 
200
- ### Layer Review Summary
201
- | Layer | Status | Finding Count | Notable Issues |
202
- |-------|--------|--------------|----------------|
203
- | Frontend | CLEAN / ISSUES_FOUND / SKIPPED | N | ... |
204
- | Backend | CLEAN / ISSUES_FOUND / SKIPPED | N | ... |
205
- | Security | CLEAN / WARNINGS / BLOCKED / SKIPPED | N | ... |
205
+ ### Review Dimensions
206
+ | Dimension | Status | Finding Count | Notable Issues |
207
+ |-----------|--------|--------------|----------------|
208
+ | Correctness / tests | CLEAN / ISSUES_FOUND | N | ... |
209
+ | Security | CLEAN / WARNINGS / BLOCKED | N | ... |
210
+ | Performance | CLEAN / WARNINGS | N | ... |
206
211
 
207
- [List any High or Critical findings from layer reviews that warrant attention]
212
+ SECURITY_STATUS: <BLOCKED | WARNINGS | CLEAN>
213
+
214
+ [List any High or Critical findings that warrant attention]
208
215
 
209
216
  ### Files Modified by Reviewer
210
217
  - [list of files the reviewer had to touch]
211
218
  ```
212
219
 
220
+ The `SECURITY_STATUS:` line is MANDATORY and machine-parsed by the orchestrator — emit it exactly once, on its own line, with one of the three values. `BLOCKED` means you found a security issue severe enough that the change must not ship (committed secret, injection, missing auth on a privileged operation). `WARNINGS` means non-blocking security findings exist. `CLEAN` otherwise.
221
+
213
222
  ## Rules
214
223
 
215
224
  - Never ask for clarification. Fix issues autonomously.
216
- - Always run ALL checks, even if you think nothing changed in a layer.
225
+ - Follow the Verification policy: scoped-first, one full pass only when your own changes (or an unbounded blast radius) warrant it. Never skip the cheap static checks.
226
+ - In the CI Checks report table, mark suites you did not re-run as `covered by developer's full pass` — never as passed-by-you.
227
+ - **Output economy**: when a check fails, carry forward only the failing test/rule names and the relevant error excerpt (≤50 lines) — never re-paste a full runner log into your reasoning.
228
+ - **File re-read discipline**: a file you already read is in your context. Before reading it again, state in one sentence what you already learned from it — re-read only if it changed.
229
+ - **Loop detection**: the same command run 3 times with no intervening code change and inconsistent results means STOP — reassess instead of re-running.
217
230
  - When fixing lint errors, understand the rule before applying a fix — don't just suppress with disable comments.
218
231
  - If a test fails, read the test AND the implementation to understand the root cause before fixing.
219
- - If a layer reviewer reports High severity findings, include them in your Issues Fixed or Issues Found section. Attempt to fix High-severity layer findings that are straightforward (e.g., adding a missing `alt` attribute, adding a missing `LIMIT` to a query). Flag Critical or architecturally complex findings for human review — do NOT attempt to fix them automatically.
232
+ - Attempt to fix High-severity findings that are straightforward (e.g., adding a missing `alt` attribute, adding a missing `LIMIT` to a query). Flag Critical or architecturally complex findings for human review — do NOT attempt to fix them automatically.
220
233
 
221
234
  ## Explain Your Work
222
235
 
@@ -86,60 +86,37 @@ the pipeline yourself):
86
86
  > Read `jq '.tickets["<TICKET_ID>"]' .specrails/local-tickets.json`
87
87
  > for the full ticket. Follow the `$sr-architect` skill
88
88
  > instructions exactly.
89
- >
90
- > In `design.md`'s `## Context` section, include a
91
- > `Scope: <labels>` line drawn from: `frontend`, `backend`,
92
- > `both`, `security-sensitive`, `performance-sensitive`.
93
89
 
94
90
  - `wait_agent`. Parse reply for the plan path. `close_agent`.
95
- - Open the plan + design.md, parse the `Scope:` line.
91
+ - Open the plan + design.md.
96
92
  - If the architect returned `BLOCKED: …`, mark this ticket
97
93
  as failed for the batch report and **continue to the next
98
94
  ticket** — do not stop the batch.
99
95
 
100
96
  #### 1.b Developer phase (per ticket)
101
97
 
102
- Routing matrix (mirrors `$implement`):
103
-
104
- | scope contains | rails available | spawn |
105
- |---|---|---|
106
- | `frontend` only | `sr-frontend-developer` | $sr-frontend-developer |
107
- | `backend` only | `sr-backend-developer` | $sr-backend-developer |
108
- | `frontend` only | (no fe specialist) | $sr-developer |
109
- | `backend` only | (no be specialist) | $sr-developer |
110
- | `both` + both specialists + tagged tasks.md | — | TWO devs parallel |
111
- | else | — | $sr-developer |
98
+ One developer rail. Unless a profile routes the ticket to a
99
+ listed `custom-*` developer, spawn `$sr-developer`.
112
100
 
113
101
  - `spawn_agent`. `send_message`:
114
102
 
115
- > `$<developer-skill>`
103
+ > `$sr-developer`
116
104
  >
117
105
  > Ticket id: `<TICKET_ID>`
118
106
  > Plan: `<PLAN_PATH>`
119
- > Scope: `<comma-separated labels>`
120
107
  >
121
- > Follow the `$<developer-skill>` skill instructions exactly.
108
+ > Follow the `$sr-developer` skill instructions exactly.
122
109
 
123
110
  - `wait_agent`. Capture file list. `close_agent`.
124
111
  - If `BLOCKED: …` → mark ticket as failed in the batch report
125
112
  and move to next ticket.
126
113
 
127
- #### 1.c Reviewer phase (per ticket) — parallel where possible
128
-
129
- Always spawn `$sr-reviewer`. Additionally if installed AND
130
- scope matches:
131
-
132
- | scope flag | additional rail |
133
- |---|---|
134
- | `frontend` | `$sr-frontend-reviewer` |
135
- | `backend` | `$sr-backend-reviewer` |
136
- | `security-sensitive` | `$sr-security-reviewer` |
137
- | `performance-sensitive` | `$sr-performance-reviewer` |
114
+ #### 1.c Reviewer phase (per ticket)
138
115
 
139
- Spawn ALL reviewers in parallel, then `wait_agent` on each.
140
- `close_agent` each.
116
+ Spawn the single `$sr-reviewer` — it covers correctness, tests,
117
+ security, and performance. `wait_agent`, then `close_agent`.
141
118
 
142
- Aggregate verdicts (same matrix as `$implement`):
119
+ Verdict (same matrix as `$implement`):
143
120
 
144
121
  - `clean` — every reviewer ≥70, no fix/blocked verdicts.
145
122
  - `fix needed` — any "fix needed", OR score <70 with no