@cxi-lmai/ci-agent-platform 3.2.2 → 4.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +32 -16
- package/package.json +1 -1
- package/payload/agents/agent-architect.md +1 -1
- package/payload/agents/code-reviewer.md +1 -1
- package/payload/agents/codebase-auditor.md +1 -1
- package/payload/agents/coder.md +1 -1
- package/payload/agents/decomposer.md +1 -1
- package/payload/agents/docs-sync.md +1 -1
- package/payload/agents/e2e-test-writer.md +1 -1
- package/payload/agents/orchestrator.md +1 -1
- package/payload/agents/performance-reviewer.md +1 -1
- package/payload/agents/release-mr.md +1 -1
- package/payload/agents/security-reviewer.md +1 -1
- package/payload/agents/test-fix.md +1 -1
- package/payload/agents/test-writer.md +1 -1
- package/payload/agents-omp/agent-architect.md +1 -1
- package/payload/agents-omp/code-reviewer.md +1 -1
- package/payload/agents-omp/codebase-auditor.md +1 -1
- package/payload/agents-omp/coder.md +1 -1
- package/payload/agents-omp/decomposer.md +1 -1
- package/payload/agents-omp/docs-sync.md +1 -1
- package/payload/agents-omp/e2e-test-writer.md +1 -1
- package/payload/agents-omp/orchestrator.md +1 -1
- package/payload/agents-omp/performance-reviewer.md +1 -1
- package/payload/agents-omp/release-mr.md +1 -1
- package/payload/agents-omp/security-reviewer.md +1 -1
- package/payload/agents-omp/test-fix.md +1 -1
- package/payload/agents-omp/test-writer.md +1 -1
- package/payload/ci-templates/claude-pipeline.gitlab-ci.yml +9 -4
- package/payload/ci-templates/github/README.md +1 -1
- package/payload/ci-templates/github/claude-issue-pipeline.yml +7 -3
- package/payload/ci-templates/github/claude-pipeline.yml +7 -3
- package/payload/ci-templates/github/claude-test-fix.yml +3 -1
- package/payload/ci-templates/scripts/lib/pipeline-common.sh +61 -11
- package/payload/ci-templates/scripts/lib/usage-capture-omp.sh +114 -22
- package/payload/ci-templates/scripts/lib/usage-capture.sh +6 -5
- package/payload/skills/init-pipeline-config/SKILL.md +4 -4
package/README.md
CHANGED
|
@@ -42,7 +42,7 @@ checks the result. The input is a spec, the output is code. Details in
|
|
|
42
42
|
- [The full cycle](#the-full-cycle)
|
|
43
43
|
- [Repository layout](#repository-layout)
|
|
44
44
|
- [Install](#install)
|
|
45
|
-
- [Harness
|
|
45
|
+
- [Harness and model provider](#harness-and-model-provider)
|
|
46
46
|
- [What it costs](#what-it-costs)
|
|
47
47
|
- [Upgrading and removing](#upgrading-and-removing)
|
|
48
48
|
- [The issue-to-code loop (details)](https://gitlab.com/cxi-lmai/ci-agent-platform/-/blob/main/docs/issue-to-code.md)
|
|
@@ -104,9 +104,10 @@ Two things are worth having before you start, though neither blocks the install.
|
|
|
104
104
|
A `CLAUDE.md` at the repository root: most agents read it as the project's
|
|
105
105
|
rulebook, and without one the reviewers work from `pipeline-config.md` alone. And
|
|
106
106
|
CI credentials for whichever platform you use, which the wizard walks you
|
|
107
|
-
through at the end.
|
|
108
|
-
|
|
109
|
-
|
|
107
|
+
through at the end. By default one `OPENROUTER_API_KEY` authenticates every
|
|
108
|
+
agent. With `PIPE_PROVIDER=anthropic`, any one of `ANTHROPIC_API_KEY`,
|
|
109
|
+
`ANTHROPIC_AUTH_TOKEN` or `CLAUDE_CODE_OAUTH_TOKEN` does instead, so a Claude
|
|
110
|
+
subscription works without buying API credit.
|
|
110
111
|
|
|
111
112
|
The first command is plain Node with no model in it. It refuses to run outside
|
|
112
113
|
a git repository, detects your platform from the git remote (without ever
|
|
@@ -131,18 +132,33 @@ rest:
|
|
|
131
132
|
> `--experimental-github` to answer that in advance, or `--force` to proceed
|
|
132
133
|
> past a prior install it cannot account for.
|
|
133
134
|
|
|
134
|
-
### Harness
|
|
135
|
-
|
|
136
|
-
The pipeline runs on Claude Code
|
|
137
|
-
|
|
138
|
-
`
|
|
139
|
-
|
|
140
|
-
|
|
141
|
-
|
|
142
|
-
|
|
143
|
-
picked, not both.
|
|
144
|
-
|
|
145
|
-
|
|
135
|
+
### Harness and model provider
|
|
136
|
+
|
|
137
|
+
The pipeline runs on Claude Code (`PIPE_HARNESS=claude`, the `claude` CLI)
|
|
138
|
+
with its model calls routed through [OpenRouter](https://openrouter.ai)
|
|
139
|
+
(`PIPE_PROVIDER=openrouter`): one `OPENROUTER_API_KEY` authenticates every
|
|
140
|
+
job, and usage is billed by OpenRouter. The runner points Claude Code's
|
|
141
|
+
`ANTHROPIC_BASE_URL` at OpenRouter's Anthropic-compatible endpoint, so nothing
|
|
142
|
+
in the agents or skills changes. Set `PIPE_PROVIDER=anthropic` to call
|
|
143
|
+
Anthropic directly with `ANTHROPIC_API_KEY` instead. Set the credential for the
|
|
144
|
+
provider you picked, not both.
|
|
145
|
+
|
|
146
|
+
| Role | Variable | Default model |
|
|
147
|
+
|---|---|---|
|
|
148
|
+
| coding (coder, test-fix, test writers) | `PIPE_MODEL_CODE`, agent tier `opus` | Claude Opus 5.5 |
|
|
149
|
+
| review, docs, planning | `PIPE_MODEL_REVIEW`, agent tier `sonnet` | Claude Sonnet 5.5 |
|
|
150
|
+
| triage, postmortem | `PIPE_MODEL_TRIAGE`, agent tier `haiku` | Claude Haiku 4.5 |
|
|
151
|
+
|
|
152
|
+
Subagents name a tier in their frontmatter, and the runner maps each tier to
|
|
153
|
+
the provider's model id through `ANTHROPIC_DEFAULT_OPUS_MODEL`,
|
|
154
|
+
`_SONNET_MODEL` and `_HAIKU_MODEL`. Set any of those, or a `PIPE_MODEL_*`
|
|
155
|
+
variable, as a CI variable to pick another model. On OpenRouter that can be
|
|
156
|
+
any OpenRouter model id, not just Anthropic's.
|
|
157
|
+
|
|
158
|
+
The older [Oh My Pi](https://openrouter.ai/apps/oh-my-pi) harness
|
|
159
|
+
(`PIPE_HARNESS=omp`, OpenRouter only, GitLab only) is still shipped for
|
|
160
|
+
projects that already run it, but it is no longer what the install wizard
|
|
161
|
+
offers.
|
|
146
162
|
|
|
147
163
|
`omp` is a Bun program, so the GitLab template's `before_script` installs
|
|
148
164
|
`bun` next to it (`npm install -g bun`) — the default `node:22-bookworm` image
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@cxi-lmai/ci-agent-platform",
|
|
3
|
-
"version": "
|
|
3
|
+
"version": "4.0.0",
|
|
4
4
|
"description": "Autonomous dev pipeline on plain GitLab CI or GitHub Actions, driven by Claude Code. A labeled issue goes in, an open merge request comes out.",
|
|
5
5
|
"keywords": [
|
|
6
6
|
"claude",
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
name: agent-architect
|
|
3
3
|
description: Weekly agent-improvement lab. Reads recently merged MRs/PRs, review comments, and stuck-MR postmortems, then proposes improvements to .claude/agents/*.md files as actionable platform issues (auto-filed with the pipeline's ready label). Output is a structured markdown report consumed by the CI script. Never auto-applies changes to CLAUDE.md or .claude/settings.json.
|
|
4
4
|
tools: Glob, Grep, LS, Read
|
|
5
|
-
model:
|
|
5
|
+
model: sonnet
|
|
6
6
|
---
|
|
7
7
|
|
|
8
8
|
You are the Agent Lab for this project's autonomous development pipeline.
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
name: code-reviewer
|
|
3
3
|
description: Reviews code for bugs, logic errors, security vulnerabilities, code quality issues, and adherence to project conventions, using confidence-based filtering to report only high-priority issues that truly matter
|
|
4
4
|
tools: Glob, Grep, LS, Read, Write, NotebookRead, WebFetch, TodoWrite, WebSearch, KillShell, BashOutput, mcp__context7__resolve-library-id, mcp__context7__get-library-docs
|
|
5
|
-
model:
|
|
5
|
+
model: sonnet
|
|
6
6
|
color: red
|
|
7
7
|
---
|
|
8
8
|
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
name: codebase-auditor
|
|
3
3
|
description: Periodic codebase auditor. Reads project conventions and docs, reads recently merged MR/PR file changes, and produces a focused report of critical convention violations and genuinely new undocumented patterns. Uses confidence-based filtering to report only issues that truly matter.
|
|
4
4
|
tools: Glob, Grep, LS, Read
|
|
5
|
-
model:
|
|
5
|
+
model: sonnet
|
|
6
6
|
---
|
|
7
7
|
|
|
8
8
|
You are the codebase auditor. You produce a periodic report of critical convention violations and genuinely new, reusable patterns worth documenting.
|
package/payload/agents/coder.md
CHANGED
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
name: coder
|
|
3
3
|
description: Autonomous implementation agent. Given an issue description, explores existing patterns, implements the feature or fix, writes tests, runs the build in a self-correcting loop (up to 3 retries), and commits. Never pushes.
|
|
4
4
|
tools: Agent, Bash, Edit, Glob, Grep, LS, MultiEdit, NotebookRead, Read, TodoWrite, WebFetch, WebSearch, Write, mcp__context7__resolve-library-id, mcp__context7__get-library-docs
|
|
5
|
-
model:
|
|
5
|
+
model: opus
|
|
6
6
|
---
|
|
7
7
|
|
|
8
8
|
You are the autonomous coder agent for this project.
|
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: decomposer
|
|
3
3
|
description: Decomposes a too-large issue into a dependency-ordered chain of spec-compliant sub-issues. Returns structured JSON with confidence verdict, reason, and fully authored sub-issue bodies ready for creation on the project platform.
|
|
4
|
-
model:
|
|
4
|
+
model: sonnet
|
|
5
5
|
---
|
|
6
6
|
|
|
7
7
|
You are the decomposer for the autonomous development pipeline.
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
name: docs-sync
|
|
3
3
|
description: Checks whether MR/PR code changes require updates to project docs, CLAUDE.md, or .claude/memory/. On autonomous MRs it applies the updates directly; on human MRs it reports the gaps for a comment.
|
|
4
4
|
tools: Glob, Grep, LS, Read, Write, Edit
|
|
5
|
-
model:
|
|
5
|
+
model: sonnet
|
|
6
6
|
color: blue
|
|
7
7
|
---
|
|
8
8
|
|
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: e2e-test-writer
|
|
3
3
|
description: Generates a single end-to-end (E2E) test spec file for a frontend issue, following the project's configured E2E framework, test directory, fixtures, test-data prefix, cleanup rules, and tags
|
|
4
|
-
model:
|
|
4
|
+
model: opus
|
|
5
5
|
---
|
|
6
6
|
|
|
7
7
|
You are an end-to-end (E2E) test writer. Given an issue, you produce one E2E test spec file that follows the project's configured E2E framework and conventions.
|
|
@@ -14,7 +14,7 @@ Your task: analyze an issue and classify it for autonomous implementation by the
|
|
|
14
14
|
|
|
15
15
|
## Scope Constraint
|
|
16
16
|
|
|
17
|
-
A single coder agent session handles approximately **15 files / 600 lines of change** (default, see Capacities in the pipeline config). The coder uses the
|
|
17
|
+
A single coder agent session handles approximately **15 files / 600 lines of change** (default, see Capacities in the pipeline config). The coder uses the Opus model. Issues that touch many unrelated subsystems or require large architectural refactors exceed this limit and must be decomposed.
|
|
18
18
|
|
|
19
19
|
## Decision Criteria
|
|
20
20
|
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
name: performance-reviewer
|
|
3
3
|
description: Reviews changed data-access and hot-path code for confirmed scalability regressions such as repeated I/O, unbounded work, excessive loading, and missing batching or pagination
|
|
4
4
|
tools: Glob, Grep, Read, NotebookRead, WebFetch, TodoWrite, WebSearch, KillShell, BashOutput
|
|
5
|
-
model:
|
|
5
|
+
model: sonnet
|
|
6
6
|
color: yellow
|
|
7
7
|
---
|
|
8
8
|
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
name: release-mr
|
|
3
3
|
description: Creates a release MR/PR from the integration branch to the production branch with a structured description listing changes, migrations, env changes, and a deploy checklist. Use when the user wants to prepare a release or merge the integration branch into production.
|
|
4
4
|
tools: Bash, Glob, Grep, Read, TodoWrite
|
|
5
|
-
model:
|
|
5
|
+
model: sonnet
|
|
6
6
|
---
|
|
7
7
|
|
|
8
8
|
You are the release MR/PR agent. Your job is to analyze everything that changed on the integration branch since the last release to the production branch, then create (or update) a well-structured release request.
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
name: security-reviewer
|
|
3
3
|
description: Reviews code for security vulnerabilities (authentication bypass, authorization flaws, injection, tenant/data isolation leaks, CSRF misconfiguration, sensitive data exposure) using confidence-based filtering to report only confirmed issues
|
|
4
4
|
tools: Glob, Grep, LS, Read, NotebookRead, WebFetch, TodoWrite, WebSearch, KillShell, BashOutput
|
|
5
|
-
model:
|
|
5
|
+
model: sonnet
|
|
6
6
|
color: red
|
|
7
7
|
---
|
|
8
8
|
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
name: test-fix
|
|
3
3
|
description: Fixes failing tests in pipeline merge requests. Reads failing test files and their production source classes, determines root cause, fixes implementation or test as needed, verifies compilation, and commits. Never pushes.
|
|
4
4
|
tools: Agent, Bash, Edit, Glob, Grep, LS, MultiEdit, NotebookRead, Read, TodoWrite, Write
|
|
5
|
-
model:
|
|
5
|
+
model: opus
|
|
6
6
|
---
|
|
7
7
|
|
|
8
8
|
You are the test-fix agent. Your job is to fix failing tests in a merge request, not to rewrite features.
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
name: test-writer
|
|
3
3
|
description: Generates integration and unit tests that increase coverage for new or changed code, following the project's documented testing patterns, with a self-correcting compilation loop
|
|
4
4
|
tools: Agent, Bash, Edit, Glob, Grep, LS, MultiEdit, NotebookRead, Read, TodoWrite, Write
|
|
5
|
-
model:
|
|
5
|
+
model: opus
|
|
6
6
|
color: green
|
|
7
7
|
---
|
|
8
8
|
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
name: agent-architect
|
|
3
3
|
description: Weekly agent-improvement lab. Reads recently merged MRs/PRs, review comments, and stuck-MR postmortems, then proposes improvements to .claude/agents/*.md files as actionable platform issues (auto-filed with the pipeline's ready label). Output is a structured markdown report consumed by the CI script. Never auto-applies changes to CLAUDE.md or .claude/settings.json.
|
|
4
4
|
tools: glob, grep, read
|
|
5
|
-
model: openrouter/anthropic/claude-sonnet-5
|
|
5
|
+
model: openrouter/anthropic/claude-sonnet-5.5
|
|
6
6
|
---
|
|
7
7
|
|
|
8
8
|
You are the Agent Lab for this project's autonomous development pipeline.
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
name: code-reviewer
|
|
3
3
|
description: Reviews code for bugs, logic errors, security vulnerabilities, code quality issues, and adherence to project conventions, using confidence-based filtering to report only high-priority issues that truly matter
|
|
4
4
|
tools: glob, grep, read, write, todo, web_search, mcp__context7__resolve-library-id, mcp__context7__get-library-docs
|
|
5
|
-
model: openrouter/anthropic/claude-sonnet-5
|
|
5
|
+
model: openrouter/anthropic/claude-sonnet-5.5
|
|
6
6
|
---
|
|
7
7
|
|
|
8
8
|
You are an expert code reviewer specializing in modern software development across multiple languages and frameworks. Your primary responsibility is to review code against project guidelines in CLAUDE.md with high precision to minimize false positives.
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
name: codebase-auditor
|
|
3
3
|
description: Periodic codebase auditor. Reads project conventions and docs, reads recently merged MR/PR file changes, and produces a focused report of critical convention violations and genuinely new undocumented patterns. Uses confidence-based filtering to report only issues that truly matter.
|
|
4
4
|
tools: glob, grep, read
|
|
5
|
-
model: openrouter/anthropic/claude-sonnet-5
|
|
5
|
+
model: openrouter/anthropic/claude-sonnet-5.5
|
|
6
6
|
---
|
|
7
7
|
|
|
8
8
|
You are the codebase auditor. You produce a periodic report of critical convention violations and genuinely new, reusable patterns worth documenting.
|
|
@@ -3,7 +3,7 @@ name: coder
|
|
|
3
3
|
description: Autonomous implementation agent. Given an issue description, explores existing patterns, implements the feature or fix, writes tests, runs the build in a self-correcting loop (up to 3 retries), and commits. Never pushes.
|
|
4
4
|
tools: task, bash, edit, glob, grep, read, todo, web_search, write, mcp__context7__resolve-library-id, mcp__context7__get-library-docs
|
|
5
5
|
spawns: test-writer, docs-sync
|
|
6
|
-
model: openrouter/anthropic/claude-
|
|
6
|
+
model: openrouter/anthropic/claude-opus-5.5
|
|
7
7
|
---
|
|
8
8
|
|
|
9
9
|
You are the autonomous coder agent for this project.
|
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: decomposer
|
|
3
3
|
description: Decomposes a too-large issue into a dependency-ordered chain of spec-compliant sub-issues. Returns structured JSON with confidence verdict, reason, and fully authored sub-issue bodies ready for creation on the project platform.
|
|
4
|
-
model: openrouter/anthropic/claude-sonnet-5
|
|
4
|
+
model: openrouter/anthropic/claude-sonnet-5.5
|
|
5
5
|
---
|
|
6
6
|
|
|
7
7
|
You are the decomposer for the autonomous development pipeline.
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
name: docs-sync
|
|
3
3
|
description: Checks whether MR/PR code changes require updates to project docs, CLAUDE.md, or .claude/memory/. On autonomous MRs it applies the updates directly; on human MRs it reports the gaps for a comment.
|
|
4
4
|
tools: glob, grep, read, write, edit
|
|
5
|
-
model: openrouter/anthropic/claude-sonnet-5
|
|
5
|
+
model: openrouter/anthropic/claude-sonnet-5.5
|
|
6
6
|
---
|
|
7
7
|
|
|
8
8
|
You are the documentation sync agent.
|
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: e2e-test-writer
|
|
3
3
|
description: Generates a single end-to-end (E2E) test spec file for a frontend issue, following the project's configured E2E framework, test directory, fixtures, test-data prefix, cleanup rules, and tags
|
|
4
|
-
model: openrouter/anthropic/claude-
|
|
4
|
+
model: openrouter/anthropic/claude-opus-5.5
|
|
5
5
|
---
|
|
6
6
|
|
|
7
7
|
You are an end-to-end (E2E) test writer. Given an issue, you produce one E2E test spec file that follows the project's configured E2E framework and conventions.
|
|
@@ -14,7 +14,7 @@ Your task: analyze an issue and classify it for autonomous implementation by the
|
|
|
14
14
|
|
|
15
15
|
## Scope Constraint
|
|
16
16
|
|
|
17
|
-
A single coder agent session handles approximately **15 files / 600 lines of change** (default, see Capacities in the pipeline config). The coder uses the
|
|
17
|
+
A single coder agent session handles approximately **15 files / 600 lines of change** (default, see Capacities in the pipeline config). The coder uses the Opus model. Issues that touch many unrelated subsystems or require large architectural refactors exceed this limit and must be decomposed.
|
|
18
18
|
|
|
19
19
|
## Decision Criteria
|
|
20
20
|
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
name: performance-reviewer
|
|
3
3
|
description: Reviews changed data-access and hot-path code for confirmed scalability regressions such as repeated I/O, unbounded work, excessive loading, and missing batching or pagination
|
|
4
4
|
tools: glob, grep, read, todo, web_search
|
|
5
|
-
model: openrouter/anthropic/claude-sonnet-5
|
|
5
|
+
model: openrouter/anthropic/claude-sonnet-5.5
|
|
6
6
|
---
|
|
7
7
|
|
|
8
8
|
You are a performance reviewer. Work with the language, framework, storage
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
name: release-mr
|
|
3
3
|
description: Creates a release MR/PR from the integration branch to the production branch with a structured description listing changes, migrations, env changes, and a deploy checklist. Use when the user wants to prepare a release or merge the integration branch into production.
|
|
4
4
|
tools: bash, glob, grep, read, todo
|
|
5
|
-
model: openrouter/anthropic/claude-sonnet-5
|
|
5
|
+
model: openrouter/anthropic/claude-sonnet-5.5
|
|
6
6
|
---
|
|
7
7
|
|
|
8
8
|
You are the release MR/PR agent. Your job is to analyze everything that changed on the integration branch since the last release to the production branch, then create (or update) a well-structured release request.
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
name: security-reviewer
|
|
3
3
|
description: Reviews code for security vulnerabilities (authentication bypass, authorization flaws, injection, tenant/data isolation leaks, CSRF misconfiguration, sensitive data exposure) using confidence-based filtering to report only confirmed issues
|
|
4
4
|
tools: glob, grep, read, todo, web_search
|
|
5
|
-
model: openrouter/anthropic/claude-sonnet-5
|
|
5
|
+
model: openrouter/anthropic/claude-sonnet-5.5
|
|
6
6
|
---
|
|
7
7
|
|
|
8
8
|
You are an expert security reviewer. Your primary responsibility is to identify security vulnerabilities with high precision. False positives waste developer time and erode trust in the review process.
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
name: test-fix
|
|
3
3
|
description: Fixes failing tests in pipeline merge requests. Reads failing test files and their production source classes, determines root cause, fixes implementation or test as needed, verifies compilation, and commits. Never pushes.
|
|
4
4
|
tools: task, bash, edit, glob, grep, read, todo, write
|
|
5
|
-
model: openrouter/anthropic/claude-
|
|
5
|
+
model: openrouter/anthropic/claude-opus-5.5
|
|
6
6
|
---
|
|
7
7
|
|
|
8
8
|
You are the test-fix agent. Your job is to fix failing tests in a merge request, not to rewrite features.
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
name: test-writer
|
|
3
3
|
description: Generates integration and unit tests that increase coverage for new or changed code, following the project's documented testing patterns, with a self-correcting compilation loop
|
|
4
4
|
tools: task, bash, edit, glob, grep, read, todo, write
|
|
5
|
-
model: openrouter/anthropic/claude-
|
|
5
|
+
model: openrouter/anthropic/claude-opus-5.5
|
|
6
6
|
---
|
|
7
7
|
|
|
8
8
|
You are the test-writer agent. Your job is to write tests that increase coverage for new or changed code.
|
|
@@ -16,8 +16,9 @@
|
|
|
16
16
|
# This template only reads its results (see the test-fix job's `needs:`).
|
|
17
17
|
#
|
|
18
18
|
# Override any PIPE_ variable in your own variables: block. Secrets
|
|
19
|
-
# (
|
|
20
|
-
#
|
|
19
|
+
# (OPENROUTER_API_KEY, or ANTHROPIC_API_KEY with PIPE_PROVIDER=anthropic, and
|
|
20
|
+
# PIPE_BOT_TOKEN) are set as masked CI variables, never here. Do NOT
|
|
21
|
+
# re-declare a secret in a job's variables: block as a same-name
|
|
21
22
|
# self-reference (FOO: "$FOO"): on GitLab it resolves to an empty string and
|
|
22
23
|
# shadows the real project variable (pilot 2 finding, P6). Project variables
|
|
23
24
|
# reach every job automatically. The GitHub workflows are different: there the
|
|
@@ -50,6 +51,7 @@ variables:
|
|
|
50
51
|
# It is kept here as an explicit, harmless override.
|
|
51
52
|
PIPE_PLATFORM: "gitlab"
|
|
52
53
|
PIPE_HARNESS: "claude" # "claude" (default, Claude Code CLI) or "omp" (Oh My Pi + OpenRouter)
|
|
54
|
+
PIPE_PROVIDER: "openrouter" # Claude Code model provider: "openrouter" (default, OPENROUTER_API_KEY) or "anthropic" (ANTHROPIC_API_KEY)
|
|
53
55
|
PIPE_CONFIG_PATH: ".claude/pipeline-config.md"
|
|
54
56
|
PIPE_CONTEXT_DIR: "build/pipeline"
|
|
55
57
|
|
|
@@ -167,8 +169,11 @@ variables:
|
|
|
167
169
|
# --- Caps + models ---------------------------------------------------------
|
|
168
170
|
PIPE_FIX_LOOP_CAP: "2"
|
|
169
171
|
# Main-loop models are passed as --model by the active harness. Defaults live
|
|
170
|
-
# in `pipeline-common.sh` under `scripts/lib
|
|
171
|
-
#
|
|
172
|
+
# in `pipeline-common.sh` under `scripts/lib/` and follow PIPE_HARNESS and
|
|
173
|
+
# PIPE_PROVIDER (Opus 5.5 codes, Sonnet 5.5 reviews, Haiku 4.5 triages).
|
|
174
|
+
# Users may explicitly set any PIPE_MODEL_* project CI variable. Subagents
|
|
175
|
+
# retain their frontmatter tiers (opus/sonnet/haiku), which Claude Code
|
|
176
|
+
# resolves through ANTHROPIC_DEFAULT_{OPUS,SONNET,HAIKU}_MODEL.
|
|
172
177
|
PIPE_AGENT_ENV_ALLOWLIST: "" # credential-like env names the agent-run build must receive (comma-separated; prefer empty)
|
|
173
178
|
|
|
174
179
|
# --- Result files (defaults live under PIPE_CONTEXT_DIR) --------------------
|
|
@@ -40,7 +40,7 @@ few names differ. The runner scripts in `../scripts/` are shared: they read
|
|
|
40
40
|
- uses: actions/checkout@v4
|
|
41
41
|
with: { ref: ${{ github.event.pull_request.head.ref }}, fetch-depth: 0, token: ${{ secrets.PIPE_BOT_TOKEN }} }
|
|
42
42
|
- run: command -v claude >/dev/null 2>&1 || npm install -g @anthropic-ai/claude-code
|
|
43
|
-
- env: { ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}, GH_TOKEN: ${{ secrets.PIPE_BOT_TOKEN }} }
|
|
43
|
+
- env: { OPENROUTER_API_KEY: ${{ secrets.OPENROUTER_API_KEY }}, ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}, GH_TOKEN: ${{ secrets.PIPE_BOT_TOKEN }} }
|
|
44
44
|
run: bash "$PIPE_SCRIPTS_DIR/test-fix.sh"
|
|
45
45
|
```
|
|
46
46
|
- **Fork PRs.** `workflow_run.pull_requests` is empty for PRs from forks (a
|
|
@@ -21,7 +21,8 @@
|
|
|
21
21
|
# Configuration:
|
|
22
22
|
# - Repository/org VARIABLES: PIPE_TARGET_BRANCH, PIPE_LABEL_* , PIPE_CODER_CAP,
|
|
23
23
|
# PIPE_SPEC_TEMPLATE_PATH, PIPE_GIT_NAME, PIPE_GIT_EMAIL, etc.
|
|
24
|
-
# - SECRETS:
|
|
24
|
+
# - SECRETS: OPENROUTER_API_KEY (or ANTHROPIC_API_KEY with the PIPE_PROVIDER
|
|
25
|
+
# variable set to anthropic), PIPE_BOT_TOKEN (a bot PAT with issues +
|
|
25
26
|
# contents + pull-requests + actions write). The bot PAT is required so the
|
|
26
27
|
# coder's push and the orchestrate -> code dispatch re-trigger workflows;
|
|
27
28
|
# the default GITHUB_TOKEN cannot.
|
|
@@ -73,8 +74,9 @@ env:
|
|
|
73
74
|
PIPE_SPEC_TEMPLATE_PATH: ${{ vars.PIPE_SPEC_TEMPLATE_PATH || '.github/ISSUE_TEMPLATE/spec.md' }}
|
|
74
75
|
PIPE_GIT_NAME: ${{ vars.PIPE_GIT_NAME || 'Pipeline Bot' }}
|
|
75
76
|
PIPE_GIT_EMAIL: ${{ vars.PIPE_GIT_EMAIL || 'bot@pipeline.ci' }}
|
|
76
|
-
|
|
77
|
-
|
|
77
|
+
PIPE_PROVIDER: ${{ vars.PIPE_PROVIDER || 'openrouter' }}
|
|
78
|
+
PIPE_MODEL_TRIAGE: ${{ vars.PIPE_MODEL_TRIAGE }}
|
|
79
|
+
PIPE_MODEL_CODE: ${{ vars.PIPE_MODEL_CODE }}
|
|
78
80
|
PIPE_AGENT_ENV_ALLOWLIST: ${{ vars.PIPE_AGENT_ENV_ALLOWLIST || '' }}
|
|
79
81
|
# Runtime mapping onto the PIPE_ names the scripts read.
|
|
80
82
|
PIPE_REPO: ${{ github.repository }}
|
|
@@ -99,6 +101,7 @@ jobs:
|
|
|
99
101
|
run: command -v claude >/dev/null 2>&1 || npm install -g @anthropic-ai/claude-code
|
|
100
102
|
- name: Run orchestrate
|
|
101
103
|
env:
|
|
104
|
+
OPENROUTER_API_KEY: ${{ secrets.OPENROUTER_API_KEY }}
|
|
102
105
|
ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
|
|
103
106
|
GH_TOKEN: ${{ secrets.PIPE_BOT_TOKEN }}
|
|
104
107
|
run: bash "$PIPE_SCRIPTS_DIR/orchestrate.sh"
|
|
@@ -127,6 +130,7 @@ jobs:
|
|
|
127
130
|
run: command -v claude >/dev/null 2>&1 || npm install -g @anthropic-ai/claude-code
|
|
128
131
|
- name: Run code
|
|
129
132
|
env:
|
|
133
|
+
OPENROUTER_API_KEY: ${{ secrets.OPENROUTER_API_KEY }}
|
|
130
134
|
ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
|
|
131
135
|
GH_TOKEN: ${{ secrets.PIPE_BOT_TOKEN }}
|
|
132
136
|
CODER_ISSUE: ${{ github.event.inputs.coder_issue }}
|
|
@@ -9,7 +9,8 @@
|
|
|
9
9
|
# - Repository/org VARIABLES (Settings > Variables): PIPE_TARGET_BRANCH,
|
|
10
10
|
# PIPE_LABEL_WIP, PIPE_LABEL_STUCK, PIPE_BOT_USER, PIPE_VERIFY_CMD, etc.
|
|
11
11
|
# Any PIPE_ variable overrides the defaults baked into the scripts.
|
|
12
|
-
# - SECRETS (Settings > Secrets):
|
|
12
|
+
# - SECRETS (Settings > Secrets): OPENROUTER_API_KEY (or ANTHROPIC_API_KEY
|
|
13
|
+
# with the PIPE_PROVIDER variable set to anthropic), and PIPE_BOT_TOKEN
|
|
13
14
|
# (a bot Personal Access Token). The default GITHUB_TOKEN cannot re-trigger
|
|
14
15
|
# workflows on its own push, so the fix loop needs a PAT to re-run review.
|
|
15
16
|
# - Vendor the ci-agent-platform repository's ci-templates/scripts/ into your
|
|
@@ -44,8 +45,9 @@ env:
|
|
|
44
45
|
PIPE_COMMIT_REVIEWFIX: ${{ vars.PIPE_COMMIT_REVIEWFIX || 'Fix review findings' }}
|
|
45
46
|
PIPE_GIT_NAME: ${{ vars.PIPE_GIT_NAME || 'Pipeline Bot' }}
|
|
46
47
|
PIPE_GIT_EMAIL: ${{ vars.PIPE_GIT_EMAIL || 'bot@pipeline.ci' }}
|
|
47
|
-
|
|
48
|
-
|
|
48
|
+
PIPE_PROVIDER: ${{ vars.PIPE_PROVIDER || 'openrouter' }}
|
|
49
|
+
PIPE_MODEL_REVIEW: ${{ vars.PIPE_MODEL_REVIEW }}
|
|
50
|
+
PIPE_MODEL_CODE: ${{ vars.PIPE_MODEL_CODE }}
|
|
49
51
|
PIPE_AGENT_ENV_ALLOWLIST: ${{ vars.PIPE_AGENT_ENV_ALLOWLIST || '' }}
|
|
50
52
|
# Runtime mapping onto the PIPE_ names the scripts read.
|
|
51
53
|
PIPE_REPO: ${{ github.repository }}
|
|
@@ -86,6 +88,7 @@ jobs:
|
|
|
86
88
|
run: command -v claude >/dev/null 2>&1 || npm install -g @anthropic-ai/claude-code
|
|
87
89
|
- name: Run review
|
|
88
90
|
env:
|
|
91
|
+
OPENROUTER_API_KEY: ${{ secrets.OPENROUTER_API_KEY }}
|
|
89
92
|
ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
|
|
90
93
|
GH_TOKEN: ${{ secrets.PIPE_BOT_TOKEN || secrets.GITHUB_TOKEN }}
|
|
91
94
|
PIPE_NO_GATE: "1" # the gate step below owns the exit code
|
|
@@ -128,6 +131,7 @@ jobs:
|
|
|
128
131
|
run: command -v claude >/dev/null 2>&1 || npm install -g @anthropic-ai/claude-code
|
|
129
132
|
- name: Run review-fix
|
|
130
133
|
env:
|
|
134
|
+
OPENROUTER_API_KEY: ${{ secrets.OPENROUTER_API_KEY }}
|
|
131
135
|
ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
|
|
132
136
|
GH_TOKEN: ${{ secrets.PIPE_BOT_TOKEN || secrets.GITHUB_TOKEN }}
|
|
133
137
|
REVIEW_HAS_BUGS: "true"
|
|
@@ -49,7 +49,8 @@ env:
|
|
|
49
49
|
PIPE_COMMIT_COVERAGE: ${{ vars.PIPE_COMMIT_COVERAGE || 'Add coverage tests' }}
|
|
50
50
|
PIPE_GIT_NAME: ${{ vars.PIPE_GIT_NAME || 'Pipeline Bot' }}
|
|
51
51
|
PIPE_GIT_EMAIL: ${{ vars.PIPE_GIT_EMAIL || 'bot@pipeline.ci' }}
|
|
52
|
-
|
|
52
|
+
PIPE_PROVIDER: ${{ vars.PIPE_PROVIDER || 'openrouter' }}
|
|
53
|
+
PIPE_MODEL_CODE: ${{ vars.PIPE_MODEL_CODE }}
|
|
53
54
|
PIPE_AGENT_ENV_ALLOWLIST: ${{ vars.PIPE_AGENT_ENV_ALLOWLIST || '' }}
|
|
54
55
|
PIPE_TEST_ARTIFACT: ${{ vars.PIPE_TEST_ARTIFACT || 'test-reports' }}
|
|
55
56
|
PIPE_REPO: ${{ github.repository }}
|
|
@@ -91,6 +92,7 @@ jobs:
|
|
|
91
92
|
run: command -v claude >/dev/null 2>&1 || npm install -g @anthropic-ai/claude-code
|
|
92
93
|
- name: Run test-fix
|
|
93
94
|
env:
|
|
95
|
+
OPENROUTER_API_KEY: ${{ secrets.OPENROUTER_API_KEY }}
|
|
94
96
|
ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
|
|
95
97
|
GH_TOKEN: ${{ secrets.PIPE_BOT_TOKEN || secrets.GITHUB_TOKEN }}
|
|
96
98
|
PIPE_MR_IID: ${{ steps.pr.outputs.number }}
|
|
@@ -33,16 +33,24 @@ pipe_defaults() {
|
|
|
33
33
|
: "${PIPE_CODER_CAP:=3}"
|
|
34
34
|
: "${PIPE_ISSUE_SCAN:=20}"
|
|
35
35
|
: "${PIPE_HARNESS:=claude}"
|
|
36
|
-
#
|
|
37
|
-
#
|
|
36
|
+
# Where the Claude Code harness sends its model calls: OpenRouter (default,
|
|
37
|
+
# one OPENROUTER_API_KEY) or Anthropic directly. omp always uses OpenRouter.
|
|
38
|
+
: "${PIPE_PROVIDER:=openrouter}"
|
|
39
|
+
# Model defaults track the active harness and provider: Opus codes, Sonnet
|
|
40
|
+
# reviews, Haiku triages. Each PIPE_MODEL_* variable remains explicitly
|
|
41
|
+
# overridable.
|
|
38
42
|
if [ "$PIPE_HARNESS" = "omp" ]; then
|
|
39
43
|
: "${PIPE_MODEL_TRIAGE:=openrouter/anthropic/claude-haiku-4.5}"
|
|
40
|
-
: "${PIPE_MODEL_CODE:=openrouter/anthropic/claude-
|
|
41
|
-
: "${PIPE_MODEL_REVIEW:=openrouter/anthropic/claude-sonnet-5}"
|
|
44
|
+
: "${PIPE_MODEL_CODE:=openrouter/anthropic/claude-opus-5.5}"
|
|
45
|
+
: "${PIPE_MODEL_REVIEW:=openrouter/anthropic/claude-sonnet-5.5}"
|
|
46
|
+
elif [ "$PIPE_PROVIDER" = "anthropic" ]; then
|
|
47
|
+
: "${PIPE_MODEL_TRIAGE:=claude-haiku-4-5}"
|
|
48
|
+
: "${PIPE_MODEL_CODE:=claude-opus-5-5}"
|
|
49
|
+
: "${PIPE_MODEL_REVIEW:=claude-sonnet-5-5}"
|
|
42
50
|
else
|
|
43
|
-
: "${PIPE_MODEL_TRIAGE:=haiku}"
|
|
44
|
-
: "${PIPE_MODEL_CODE:=claude-
|
|
45
|
-
: "${PIPE_MODEL_REVIEW:=claude-sonnet-5}"
|
|
51
|
+
: "${PIPE_MODEL_TRIAGE:=anthropic/claude-haiku-4.5}"
|
|
52
|
+
: "${PIPE_MODEL_CODE:=anthropic/claude-opus-5.5}"
|
|
53
|
+
: "${PIPE_MODEL_REVIEW:=anthropic/claude-sonnet-5.5}"
|
|
46
54
|
fi
|
|
47
55
|
: "${PIPE_COMMIT_TESTFIX:=Fix test errors}"
|
|
48
56
|
: "${PIPE_COMMIT_REVIEWFIX:=Fix review findings}"
|
|
@@ -135,7 +143,7 @@ pipe_defaults() {
|
|
|
135
143
|
PIPE_METRICS_COMMIT_REPO PIPE_METRICS_REVIEW_AGENTS PIPE_METRICS_DEFECT_REVERT \
|
|
136
144
|
PIPE_METRICS_DEFECT_REGRESSION PIPE_METRICS_DEFECT_REINTRODUCE \
|
|
137
145
|
PIPE_METRICS_SNAPSHOT_DATE PIPE_METRICS_RAW_DIR PIPE_METRICS_MRS_FILE \
|
|
138
|
-
PIPE_HARNESS PIPE_MODEL_TRIAGE PIPE_MODEL_CODE PIPE_MODEL_REVIEW
|
|
146
|
+
PIPE_HARNESS PIPE_PROVIDER PIPE_MODEL_TRIAGE PIPE_MODEL_CODE PIPE_MODEL_REVIEW
|
|
139
147
|
}
|
|
140
148
|
|
|
141
149
|
pipe_log() { echo "[pipeline] $*"; }
|
|
@@ -222,14 +230,23 @@ pipe_agent_secret_var() {
|
|
|
222
230
|
allow=" ${allow//,/ } "
|
|
223
231
|
# Preserve only the active harness's built-in model credentials. Credentials
|
|
224
232
|
# for the inactive harness fall through to the generic secret-name check.
|
|
233
|
+
# Claude Code on OpenRouter authenticates through ANTHROPIC_AUTH_TOKEN, which
|
|
234
|
+
# pipe_claude_provider_env copies from OPENROUTER_API_KEY, so the original
|
|
235
|
+
# (and any Anthropic OAuth token that could outrank it) is scrubbed.
|
|
225
236
|
case "${PIPE_HARNESS:-claude}" in
|
|
226
237
|
omp)
|
|
227
238
|
[ "$name" = "OPENROUTER_API_KEY" ] && return 1
|
|
228
239
|
;;
|
|
229
240
|
*)
|
|
230
|
-
|
|
231
|
-
|
|
232
|
-
|
|
241
|
+
if [ "${PIPE_PROVIDER:-openrouter}" = "anthropic" ]; then
|
|
242
|
+
case "$name" in
|
|
243
|
+
ANTHROPIC_API_KEY|ANTHROPIC_AUTH_TOKEN|CLAUDE_CODE_OAUTH_TOKEN) return 1 ;;
|
|
244
|
+
esac
|
|
245
|
+
else
|
|
246
|
+
case "$name" in
|
|
247
|
+
ANTHROPIC_API_KEY|ANTHROPIC_AUTH_TOKEN) return 1 ;;
|
|
248
|
+
esac
|
|
249
|
+
fi
|
|
233
250
|
;;
|
|
234
251
|
esac
|
|
235
252
|
[[ "$allow" == *" $name "* ]] && return 1
|
|
@@ -272,6 +289,38 @@ pipe_require_harness() {
|
|
|
272
289
|
fi
|
|
273
290
|
}
|
|
274
291
|
|
|
292
|
+
# --- Claude Code model provider ---------------------------------------------
|
|
293
|
+
# OpenRouter serves the Anthropic Messages API at https://openrouter.ai/api, so
|
|
294
|
+
# pointing ANTHROPIC_BASE_URL there and passing the OpenRouter key as
|
|
295
|
+
# ANTHROPIC_AUTH_TOKEN routes every Claude Code call, subagents included,
|
|
296
|
+
# through OpenRouter. ANTHROPIC_API_KEY must be present but empty: Claude Code
|
|
297
|
+
# sends a non-empty one as a direct-Anthropic x-api-key.
|
|
298
|
+
#
|
|
299
|
+
# Agent frontmatter names model tiers (opus/sonnet/haiku), never concrete ids,
|
|
300
|
+
# so one agent set serves both providers. ANTHROPIC_DEFAULT_*_MODEL pins each
|
|
301
|
+
# tier to the provider's id for the current model; a project may override any.
|
|
302
|
+
|
|
303
|
+
pipe_claude_provider_env() {
|
|
304
|
+
if [ "${PIPE_PROVIDER:-openrouter}" = "anthropic" ]; then
|
|
305
|
+
: "${ANTHROPIC_DEFAULT_OPUS_MODEL:=claude-opus-5-5}"
|
|
306
|
+
: "${ANTHROPIC_DEFAULT_SONNET_MODEL:=claude-sonnet-5-5}"
|
|
307
|
+
: "${ANTHROPIC_DEFAULT_HAIKU_MODEL:=claude-haiku-4-5}"
|
|
308
|
+
else
|
|
309
|
+
if [ -z "${OPENROUTER_API_KEY:-}" ]; then
|
|
310
|
+
pipe_log "FATAL: PIPE_PROVIDER=openrouter needs OPENROUTER_API_KEY as a masked CI variable or secret; set PIPE_PROVIDER=anthropic to run on ANTHROPIC_API_KEY instead"
|
|
311
|
+
exit 1
|
|
312
|
+
fi
|
|
313
|
+
: "${ANTHROPIC_BASE_URL:=https://openrouter.ai/api}"
|
|
314
|
+
ANTHROPIC_AUTH_TOKEN="$OPENROUTER_API_KEY"
|
|
315
|
+
ANTHROPIC_API_KEY=""
|
|
316
|
+
: "${ANTHROPIC_DEFAULT_OPUS_MODEL:=anthropic/claude-opus-5.5}"
|
|
317
|
+
: "${ANTHROPIC_DEFAULT_SONNET_MODEL:=anthropic/claude-sonnet-5.5}"
|
|
318
|
+
: "${ANTHROPIC_DEFAULT_HAIKU_MODEL:=anthropic/claude-haiku-4.5}"
|
|
319
|
+
export ANTHROPIC_BASE_URL ANTHROPIC_AUTH_TOKEN ANTHROPIC_API_KEY
|
|
320
|
+
fi
|
|
321
|
+
export ANTHROPIC_DEFAULT_OPUS_MODEL ANTHROPIC_DEFAULT_SONNET_MODEL ANTHROPIC_DEFAULT_HAIKU_MODEL
|
|
322
|
+
}
|
|
323
|
+
|
|
275
324
|
pipe_run_claude() {
|
|
276
325
|
# $1 = agent label for metrics, $2 = skill invocation (e.g. "/review-mr"),
|
|
277
326
|
# $3 = allowed tools (optional), $4 = model for the main loop (optional,
|
|
@@ -284,6 +333,7 @@ pipe_run_claude() {
|
|
|
284
333
|
# checks that variable, or calls pipe_require_agent_ran below.
|
|
285
334
|
local label="$1" skill="$2" tools="${3:-Agent,Read,Write,Edit,Glob,Grep,Bash}" model="${4:-}"
|
|
286
335
|
pipe_require_harness claude
|
|
336
|
+
pipe_claude_provider_env
|
|
287
337
|
local stdin_file; stdin_file=$(mktemp)
|
|
288
338
|
cat > "$stdin_file" || true
|
|
289
339
|
PIPE_AGENT_RC=0
|
|
@@ -1,18 +1,37 @@
|
|
|
1
1
|
# shellcheck shell=bash
|
|
2
2
|
# Source-only library. Wraps an `omp --mode json` invocation, sums usage/cost
|
|
3
|
-
#
|
|
4
|
-
# surfaces the assistant text on stdout — the omp equivalent of
|
|
3
|
+
# across the whole spawn tree the invocation paid for, writes a metrics record,
|
|
4
|
+
# and surfaces the assistant text on stdout — the omp equivalent of
|
|
5
5
|
# usage-capture.sh's capture_claude. Same function contract, same metrics
|
|
6
|
-
# record shape (bar the harness-specific exit-code
|
|
7
|
-
# parsing because omp's schema differs from Claude's
|
|
6
|
+
# record shape (bar the harness-specific exit-code and completeness keys),
|
|
7
|
+
# different event parsing because omp's schema differs from Claude's
|
|
8
|
+
# stream-json.
|
|
8
9
|
#
|
|
9
|
-
# Schema verified
|
|
10
|
-
#
|
|
10
|
+
# Schema verified against omp 18.1.21 (minimum supported; the nested
|
|
11
|
+
# `extractedToolData.task` carrier below does not exist in older lines): the
|
|
12
|
+
# final "agent_end" line carries {"type":"agent_end","messages":[...],
|
|
11
13
|
# "isTerminal":true}. Each assistant message in `messages[]` carries its own
|
|
12
14
|
# `usage` object (per-message, not session-cumulative) and top-level
|
|
13
15
|
# `model`/`provider`/`duration` fields. See
|
|
14
16
|
# docs/superpowers/specs/2026-08-26-omp-openrouter-harness-design.md section 3
|
|
15
|
-
# for the
|
|
17
|
+
# for the original verification record and docs/metrics.md for the aggregation
|
|
18
|
+
# semantics.
|
|
19
|
+
#
|
|
20
|
+
# Subagents run as child sessions, so their model calls are never parent
|
|
21
|
+
# assistant messages. omp exposes them one level at a time instead:
|
|
22
|
+
#
|
|
23
|
+
# - a blocking `task` call's tool result carries `details.usage`, the merge
|
|
24
|
+
# of `details.results[].usage`, each of which counts only that child's own
|
|
25
|
+
# assistant messages;
|
|
26
|
+
# - that child's own `task` calls arrive under
|
|
27
|
+
# `details.results[].extractedToolData.task[]` as the same shape again
|
|
28
|
+
# (omp registers a subprocess handler for the tool), so the tree recurses
|
|
29
|
+
# with every level counted exactly once;
|
|
30
|
+
# - an async spawn exposes no usage at all: the delivery message carries
|
|
31
|
+
# jobs, not tokens. Such children are counted as unmeasured and the record
|
|
32
|
+
# marks itself incomplete rather than passing a parent-only figure off as
|
|
33
|
+
# the total. The CI config overlay in pipeline-common.sh keeps task
|
|
34
|
+
# execution blocking precisely so this stays the exceptional case.
|
|
16
35
|
#
|
|
17
36
|
# Usage: capture_omp <agent_name> -- <args to pass to `omp`...>
|
|
18
37
|
# Stdin: forwarded to `omp` stdin
|
|
@@ -21,6 +40,62 @@
|
|
|
21
40
|
|
|
22
41
|
set -u
|
|
23
42
|
|
|
43
|
+
# Aggregates one `agent_end` object into a tab-separated row:
|
|
44
|
+
# input, output, cacheWrite, cacheRead, cost, parent duration, subagent cost,
|
|
45
|
+
# measured subagents, unmeasured subagents, model. The model is deliberately
|
|
46
|
+
# last: it is the one field that can be empty (a run killed before its first
|
|
47
|
+
# assistant turn), and tab is an IFS *whitespace* character, so `read` would
|
|
48
|
+
# collapse an empty leading column and shift every number one place left.
|
|
49
|
+
# Kept to jq 1.6 features (Debian stable ships 1.6; see metrics-snapshot.sh),
|
|
50
|
+
# and assigned from a plain quoted string rather than a `cat` heredoc: sourcing
|
|
51
|
+
# this library must not need a single external binary, because callers run it
|
|
52
|
+
# with a deliberately minimal PATH.
|
|
53
|
+
_OMP_USAGE_JQ='
|
|
54
|
+
def usum(f): map(f // 0) | add // 0;
|
|
55
|
+
# One `task` tool result, recursing into the nested calls its children made.
|
|
56
|
+
def level:
|
|
57
|
+
([.results[]? | select(.usage != null) | .usage]) as $own
|
|
58
|
+
| (if ($own | length) > 0 then $own
|
|
59
|
+
elif (.usage != null) then [.usage]
|
|
60
|
+
else [] end) as $usage
|
|
61
|
+
| (if ($own | length) > 0 then ($own | length)
|
|
62
|
+
elif (.usage != null) then 1
|
|
63
|
+
else 0 end) as $measured
|
|
64
|
+
# Count every agent the call spawned, measured or not. An async spawn
|
|
65
|
+
# settles with no result row at all, so the progress list is the only place
|
|
66
|
+
# it appears — including when it has a measured sibling, which must not mask
|
|
67
|
+
# it. A sync child that failed before its first request has a row but no
|
|
68
|
+
# usage, and is counted by the same subtraction.
|
|
69
|
+
| ([.results[]?] | length) as $result_rows
|
|
70
|
+
| ([.progress[]?] | length) as $progress_rows
|
|
71
|
+
| (if $progress_rows > $result_rows then $progress_rows else $result_rows end) as $spawned
|
|
72
|
+
| (if $spawned > $measured then ($spawned - $measured) else 0 end) as $gaps
|
|
73
|
+
| ([.results[]? | (.extractedToolData.task // [])[]? | level]) as $nested
|
|
74
|
+
| {
|
|
75
|
+
input: (($usage | usum(.input)) + ($nested | usum(.input))),
|
|
76
|
+
output: (($usage | usum(.output)) + ($nested | usum(.output))),
|
|
77
|
+
cacheWrite: (($usage | usum(.cacheWrite)) + ($nested | usum(.cacheWrite))),
|
|
78
|
+
cacheRead: (($usage | usum(.cacheRead)) + ($nested | usum(.cacheRead))),
|
|
79
|
+
cost: (($usage | usum(.cost.total)) + ($nested | usum(.cost))),
|
|
80
|
+
measured: ($measured + ($nested | usum(.measured))),
|
|
81
|
+
unmeasured: ($gaps + ($nested | usum(.unmeasured)))
|
|
82
|
+
};
|
|
83
|
+
([.messages[]? | select(.role == "assistant")]) as $parent
|
|
84
|
+
| ([.messages[]? | select(.role == "toolResult" and .toolName == "task") | .details | level]) as $subs
|
|
85
|
+
| [
|
|
86
|
+
(($parent | usum(.usage.input)) + ($subs | usum(.input))),
|
|
87
|
+
(($parent | usum(.usage.output)) + ($subs | usum(.output))),
|
|
88
|
+
(($parent | usum(.usage.cacheWrite)) + ($subs | usum(.cacheWrite))),
|
|
89
|
+
(($parent | usum(.usage.cacheRead)) + ($subs | usum(.cacheRead))),
|
|
90
|
+
(($parent | usum(.usage.cost.total)) + ($subs | usum(.cost))),
|
|
91
|
+
($parent | usum(.duration)),
|
|
92
|
+
($subs | usum(.cost)),
|
|
93
|
+
($subs | usum(.measured)),
|
|
94
|
+
($subs | usum(.unmeasured)),
|
|
95
|
+
([$parent[] | ((.provider // "") + "/" + (.model // ""))] | last // "")
|
|
96
|
+
] | @tsv
|
|
97
|
+
'
|
|
98
|
+
|
|
24
99
|
capture_omp() {
|
|
25
100
|
local agent="$1"; shift
|
|
26
101
|
[ "${1:-}" = "--" ] && shift
|
|
@@ -51,23 +126,32 @@ capture_omp() {
|
|
|
51
126
|
|
|
52
127
|
# Extract the final agent_end event (last line that begins with
|
|
53
128
|
# {"type":"agent_end"). Its usage is per-message (one API call each), not
|
|
54
|
-
# pre-aggregated like Claude's single "result" event,
|
|
129
|
+
# pre-aggregated like Claude's single "result" event, and subagent usage sits
|
|
130
|
+
# in the `task` tool results, so aggregate the tree ourselves
|
|
131
|
+
# (see _OMP_USAGE_JQ above).
|
|
55
132
|
local result_line
|
|
56
133
|
result_line=$(grep '^{"type":"agent_end"' "$stream_file" | tail -1 || true)
|
|
57
134
|
|
|
58
|
-
local model input output cache_c cache_r cost duration
|
|
59
|
-
|
|
60
|
-
|
|
61
|
-
|
|
62
|
-
|
|
63
|
-
|
|
64
|
-
cache_r=$(echo "$result_line" | jq '[.messages[]? | select(.role=="assistant") | .usage.cacheRead // 0] | add // 0')
|
|
65
|
-
cost=$(echo "$result_line" | jq '[.messages[]? | select(.role=="assistant") | .usage.cost.total // 0] | add // 0')
|
|
66
|
-
duration=$(echo "$result_line" | jq '[.messages[]? | select(.role=="assistant") | .duration // 0] | add // 0')
|
|
135
|
+
local model input output cache_c cache_r cost duration sub_cost measured unmeasured
|
|
136
|
+
local row=""
|
|
137
|
+
[ -n "$result_line" ] && row=$(echo "$result_line" | jq -r "$_OMP_USAGE_JQ" 2>/dev/null || true)
|
|
138
|
+
if [ -n "$row" ]; then
|
|
139
|
+
IFS=$'\t' read -r input output cache_c cache_r cost duration sub_cost measured unmeasured model \
|
|
140
|
+
<<< "$row"
|
|
67
141
|
else
|
|
142
|
+
# No parsable agent_end (crash, killed run): the parent's own wall time is
|
|
143
|
+
# the only thing known, and nothing can be claimed about subagents.
|
|
68
144
|
model=""; input=0; output=0; cache_c=0; cache_r=0; cost=0; duration=$wall_ms
|
|
145
|
+
sub_cost=0; measured=0; unmeasured=0
|
|
69
146
|
fi
|
|
70
147
|
[ -n "${PIPE_CAPTURE_MODEL:-}" ] && [ -z "$model" ] && model="$PIPE_CAPTURE_MODEL"
|
|
148
|
+
# Partial the moment any spawned agent's usage was not exposed: the totals
|
|
149
|
+
# above then cover less than the OpenRouter account was charged for. A run
|
|
150
|
+
# with no parsable agent_end is the extreme case — not even the parent's own
|
|
151
|
+
# tokens are known — so its zeros must not read as a settled bill either.
|
|
152
|
+
local complete=true
|
|
153
|
+
[ -z "$row" ] && complete=false
|
|
154
|
+
[ "$unmeasured" -gt 0 ] && complete=false
|
|
71
155
|
|
|
72
156
|
# Recover the user-facing text (concatenate all assistant text blocks) so
|
|
73
157
|
# downstream consumers that read `omp -p`'s stdout still work.
|
|
@@ -75,10 +159,11 @@ capture_omp() {
|
|
|
75
159
|
| jq -r 'select(.message.role=="assistant") | .message.content[]? | select(.type=="text") | .text' \
|
|
76
160
|
|| true
|
|
77
161
|
|
|
78
|
-
# Write metrics record. Same field names as capture_claude's record;
|
|
79
|
-
#
|
|
80
|
-
# claude_exit_code), by the same convention the two capture_* files
|
|
81
|
-
#
|
|
162
|
+
# Write metrics record. Same field names as capture_claude's record; the
|
|
163
|
+
# harness-specific exit-code key differs (omp_exit_code vs
|
|
164
|
+
# claude_exit_code), by the same convention the two capture_* files already
|
|
165
|
+
# use for their own harness-labelled fields, and the usage_* /subagent_*
|
|
166
|
+
# keys describe a scope Claude Code's single `result` event does not have.
|
|
82
167
|
local record_stamp record_file
|
|
83
168
|
record_stamp=$(date +%s%N)
|
|
84
169
|
record_file="$metrics_dir/${job_name}-${job_id}-${record_stamp}.json"
|
|
@@ -100,6 +185,10 @@ capture_omp() {
|
|
|
100
185
|
--argjson duration "$duration" \
|
|
101
186
|
--argjson wall_ms "$wall_ms" \
|
|
102
187
|
--argjson exit_code "$exit_code" \
|
|
188
|
+
--argjson sub_cost "$sub_cost" \
|
|
189
|
+
--argjson measured "$measured" \
|
|
190
|
+
--argjson unmeasured "$unmeasured" \
|
|
191
|
+
--argjson complete "$complete" \
|
|
103
192
|
'{
|
|
104
193
|
ts: $ts, agent: $agent, event: $event, model: $model,
|
|
105
194
|
pipeline_id: $pipeline_id, job_id: $job_id, job_name: $job_name,
|
|
@@ -107,7 +196,10 @@ capture_omp() {
|
|
|
107
196
|
input_tokens: $input, output_tokens: $output,
|
|
108
197
|
cache_creation_tokens: $cache_c, cache_read_tokens: $cache_r,
|
|
109
198
|
total_cost_usd: $cost, duration_ms: $duration, wall_ms: $wall_ms,
|
|
110
|
-
omp_exit_code: $exit_code
|
|
199
|
+
omp_exit_code: $exit_code,
|
|
200
|
+
usage_scope: "tree", usage_complete: $complete,
|
|
201
|
+
subagents_measured: $measured, subagents_unmeasured: $unmeasured,
|
|
202
|
+
subagent_cost_usd: $sub_cost
|
|
111
203
|
}' > "$record_file"
|
|
112
204
|
|
|
113
205
|
# If the agent failed, forward stderr so debugging still works, and keep a
|
|
@@ -42,9 +42,11 @@ capture_claude() {
|
|
|
42
42
|
end_ns=$(date +%s%N)
|
|
43
43
|
local wall_ms=$(( (end_ns - start_ns) / 1000000 ))
|
|
44
44
|
|
|
45
|
-
# Extract the result event
|
|
45
|
+
# Extract the last result event. Match on the parsed type, not a line prefix:
|
|
46
|
+
# Claude Code 2.1 no longer emits "type" as the first key, and a prefix grep
|
|
47
|
+
# silently recorded every run as zero usage. fromjson? skips non-JSON lines.
|
|
46
48
|
local result_line
|
|
47
|
-
result_line=$(
|
|
49
|
+
result_line=$(jq -cR 'fromjson? | select(.type? == "result")' "$stream_file" 2>/dev/null | tail -1 || true)
|
|
48
50
|
|
|
49
51
|
local model input output cache_c cache_r cost duration
|
|
50
52
|
if [ -n "$result_line" ]; then
|
|
@@ -67,9 +69,8 @@ capture_claude() {
|
|
|
67
69
|
|
|
68
70
|
# Recover the user-facing text (concatenate all assistant text blocks) so
|
|
69
71
|
# downstream consumers that read `claude`'s stdout still work.
|
|
70
|
-
|
|
71
|
-
|
|
72
|
-
|| true
|
|
72
|
+
jq -rR 'fromjson? | select(.type? == "assistant") | .message.content[]? | select(.type=="text") | .text' \
|
|
73
|
+
"$stream_file" 2>/dev/null || true
|
|
73
74
|
|
|
74
75
|
# Write metrics record
|
|
75
76
|
local record_stamp record_file
|
|
@@ -37,7 +37,7 @@ Source paths: `<root>` below means the first candidate that actually contains `t
|
|
|
37
37
|
|
|
38
38
|
Do not hand the user a homework list. Walk them through it:
|
|
39
39
|
|
|
40
|
-
1. Start with one confirmation of the detected basics: "The pipeline will target branch `<detected>` on `<platform>`. Correct?" Detected values are proposals to confirm, never open questions.
|
|
40
|
+
1. Start with one confirmation of the detected basics: "The pipeline will target branch `<detected>` on `<platform>`. Correct?" Detected values are proposals to confirm, never open questions. Ask in the same message where Claude Code gets its models: OpenRouter (default, needs `OPENROUTER_API_KEY`) or Anthropic directly (needs `ANTHROPIC_API_KEY`). Record the answer as `PIPE_PROVIDER` (`openrouter` or `anthropic`) and `PIPE_HARNESS=claude`. Do not offer the omp harness; it stays available on GitLab for a project that asks for it by name (`PIPE_HARNESS=omp`, `OPENROUTER_API_KEY`), and is not shipped for GitHub.
|
|
41
41
|
2. Split the remaining `TODO`s into two groups. **Defaultable:** the repo gives a defensible answer (coverage tooling absent means the policy is `not used`, Domain Check candidates read from the code, standard paths). Apply these without asking and keep them for the summary in step 4. **Genuinely open:** the repo gives no signal at all (a convention nobody wrote down, a check only a human knows about). Only these earn a question. Branch naming is not a question: with no repository convention use `<issue-iid>-<kebab-title>`, which is exactly what the runner accepts.
|
|
42
42
|
3. Ask one scheduling question because frequency is a project policy, not a detectable technical fact: should the issue loop stay manual/disabled, run nightly on selected days, or use a custom cron? Recommend manual/disabled until both smoke tests pass. For a schedule, record days, local time, and time zone. Convert to UTC only for GitHub Actions; GitLab stores the chosen cron time zone. Then ask the other genuinely open items one concrete question at a time, in config order, each with a suggested default when possible. In a typical repo this is one to three questions total. Do not ask for the bot account username: the runner resolves it from the token at runtime (`platform_resolve_bot_user`); the account itself gets created with the token in phase 4. Apply each answer to the config immediately; the user never edits the file by hand during this phase.
|
|
43
43
|
4. Close with one review summary of everything that was set: detected, defaulted, and answered, with the applied Domain Checks listed item by item. Invite the user to add, remove, or change anything; apply the edits. This summary is the safety net that lets steps 2-3 default aggressively.
|
|
@@ -50,7 +50,7 @@ Do this yourself; it is mechanical. The human only approves the diff. When the `
|
|
|
50
50
|
**GitLab** (detected in phase 1):
|
|
51
51
|
|
|
52
52
|
1. Create `.claude-pipeline/` and copy `<root>/ci-templates/claude-pipeline.gitlab-ci.yml` and `<root>/ci-templates/scripts/` into it.
|
|
53
|
-
2. Edit the project `.gitlab-ci.yml`: ensure the stages the template needs (`orchestrate, code, test, review`), add the `include:` of the vendored template, and add the non-secret `PIPE_*` values from the config that differ from the template defaults (typically `PIPE_VERIFY_CMD`, `PIPE_TEST_REPORT_GLOB`) as top-level `variables:`. Always write `PIPE_TARGET_BRANCH` with the branch detected in phase 1, as a literal: a nested `$CI_DEFAULT_BRANCH` is not expanded inside `rules:` comparisons and would silently disable every MR job (E2E finding N-3). You know the values since phase 1; wiring them now keeps phase 4 free of commits. Do not touch the project's own jobs without asking. When phase 2 chose `
|
|
53
|
+
2. Edit the project `.gitlab-ci.yml`: ensure the stages the template needs (`orchestrate, code, test, review`), add the `include:` of the vendored template, and add the non-secret `PIPE_*` values from the config that differ from the template defaults (typically `PIPE_VERIFY_CMD`, `PIPE_TEST_REPORT_GLOB`) as top-level `variables:`. Always write `PIPE_TARGET_BRANCH` with the branch detected in phase 1, as a literal: a nested `$CI_DEFAULT_BRANCH` is not expanded inside `rules:` comparisons and would silently disable every MR job (E2E finding N-3). You know the values since phase 1; wiring them now keeps phase 4 free of commits. Do not touch the project's own jobs without asking. When phase 2 chose `anthropic`, add `PIPE_PROVIDER: "anthropic"` to that same `variables:` block (the template default is `openrouter`). When the project asked for omp, add `PIPE_HARNESS: "omp"` there instead (the template default is `claude`).
|
|
54
54
|
3. Check the project's own test job: when neither its `rules:` nor a `workflow:` block makes it run in `merge_request_event` pipelines, the template's `test-fix` job can never fire (its `needs: test` finds no test job in the MR pipeline, E2E finding N-4). Tell the user and offer a concrete diff that adds the missing rule. Apply it only after they agree.
|
|
55
55
|
4. Copy `<root>/templates/spec-issue.template.md` to `.gitlab/issue_templates/Spec.md` (or the path the user chose in phase 2). When that file already exists, show the diff and ask before replacing it. An issue template is project-owned content, and the ground rule above covers the config and the CI file by name, so this one has to be said explicitly.
|
|
56
56
|
|
|
@@ -75,7 +75,7 @@ Never take the values. And never dump the whole checklist at once: this phase is
|
|
|
75
75
|
|
|
76
76
|
The steps, in this order (GitLab has 7, GitHub 6, number the counter accordingly):
|
|
77
77
|
|
|
78
|
-
1. The model credential for the
|
|
78
|
+
1. The model credential for the chosen provider. For OpenRouter (the default, and always for omp), guide `OPENROUTER_API_KEY` (https://openrouter.ai/settings/keys). For `PIPE_PROVIDER=anthropic`, guide `ANTHROPIC_API_KEY` (https://console.anthropic.com/settings/keys), and on GitHub also the repository variable `PIPE_PROVIDER=anthropic`. On GitHub the credential is a repository secret, on GitLab a masked CI variable. GitLab model credentials must not be protected when MR pipelines run from unprotected branches.
|
|
79
79
|
2. `PIPE_BOT_TOKEN`: GitLab project access token with `api` + `write_repository`, masked and hidden but not protected when ordinary unprotected feature branches need MR review/fix jobs; GitHub PAT with issues, contents, pull-requests and Actions write. Explain that protected GitLab variables work only when the project's protected-MR conditions are satisfied.
|
|
80
80
|
3. GitLab only: `PIPE_TRIGGER_TOKEN` (Settings > CI/CD > Pipeline trigger tokens) for the orchestrate-to-code dispatch.
|
|
81
81
|
4. The non-secret `PIPE_*` variables that differ from the defaults: these were already wired into the CI file in phase 3, so this step is normally a one-line "already wired, skipping". Only when something changed during phase 4 edit the CI file again (part of the walkthrough, no extra approval beyond showing the diff).
|
|
@@ -91,6 +91,6 @@ The smoke tests are the last step and the wizard drives them. The review smoke r
|
|
|
91
91
|
|
|
92
92
|
1. Announce it in one line and prepare it: a branch named by the phase 2 convention, one small harmless change, an MR/PR against the target branch carrying the wip label. Then ask the one question of this phase: confirm the push plus MR/PR creation (ground rule: never push silently). One yes covers both. On GitLab without `glab`, create the MR with git push options (`git push -o merge_request.create -o merge_request.target=<target> -o merge_request.label=<wip label> origin <branch>`), no token or CLI needed. When no `merge_request_event` pipeline appears within about a minute of the MR existing, create it yourself with `POST /projects/:id/merge_requests/:iid/pipelines` (the bot token is set by phase 4). On GitHub without `gh`, push and print the compare URL for the user to open the PR.
|
|
93
93
|
2. Check the run yourself with a single status query (`glab ci status` / `gh run list` for the run), or one short bounded wait, then report. Do not launch a blind polling loop that blocks for many minutes. If the jobs have not started or finished yet, say so and give the user the one command to re-check, rather than waiting them out. Scheduled pipelines and crons are best-effort and can lag by minutes, so an unstarted scheduled run is expected, not a failure. Expect, once it runs: the `review` job runs, a review comment from the bot account appears on the MR/PR, `review.env` and a metrics JSON land in `$PIPE_CONTEXT_DIR`. Report the result with a link to the MR/PR and the bot comment.
|
|
94
|
-
3. On failure, read the job log yourself and say what is wrong and what to change. The common causes are a missing credential for the selected
|
|
94
|
+
3. On failure, read the job log yourself and say what is wrong and what to change. The common causes are a missing credential for the selected provider (`OPENROUTER_API_KEY`, or `ANTHROPIC_API_KEY` with `PIPE_PROVIDER=anthropic`), a `PIPE_BOT_TOKEN` without comment/write scope, and an explicitly set `PIPE_BOT_USER` not matching the token's account (leave it unset; the runner resolves it from the token).
|
|
95
95
|
4. After the review smoke test passes, offer a separate end-to-end issue-loop smoke test. This is the only test that proves issue -> triage -> code -> MR/PR rather than only the review half, so it is worth running. Run it only after explicit approval, then watch it through MR/PR creation. Steps: create one small issue whose full spec sits in the issue DESCRIPTION following `spec-issue.template.md` (a spec pasted into a comment does not count, the pipeline reads the spec from the description), apply the ready label, then start orchestrate IMMEDIATELY with a manual run rather than a schedule. On GitLab that is Run pipeline on the target branch with variable `PIPE_ORCHESTRATE=1` (pipeline source `web`, which the orchestrate rule already allows), on GitHub the `workflow_dispatch` of the issue pipeline. Do not set up a cron schedule for this test. A cron is only for ongoing autonomy later and adds minutes of best-effort delay, while the manual run starts within seconds. It spends another triage plus coder run, pushes a feature branch, and opens an MR/PR.
|
|
96
96
|
5. Close with the final recap, one table: every setting that matters (target branch, branch convention, labels, verify command, report path, image, caps, models, schedule), its value, and its origin (detected / default / answered). Under the table, state what was committed and pushed, both smoke-test results (or `not run`), and whether the issue schedule is active or manual. End with a clear "done": the user must never have to ask what state the repo is in.
|