@cxi-lmai/ci-agent-platform 3.0.1 → 3.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +12 -4
- package/package.json +2 -2
- package/payload/INSTALL.md +8 -2
- package/payload/ci-templates/claude-pipeline.gitlab-ci.yml +210 -0
- package/payload/ci-templates/scripts/agent-architect.sh +233 -0
- package/payload/ci-templates/scripts/code.sh +26 -27
- package/payload/ci-templates/scripts/codebase-audit.sh +252 -0
- package/payload/ci-templates/scripts/coverage-ratchet.sh +77 -0
- package/payload/ci-templates/scripts/cve-fix.sh +246 -0
- package/payload/ci-templates/scripts/docs-sync.sh +225 -0
- package/payload/ci-templates/scripts/e2e-test-gen.sh +201 -0
- package/payload/ci-templates/scripts/lib/failure-notice.sh +104 -0
- package/payload/ci-templates/scripts/lib/issue-loop.sh +155 -40
- package/payload/ci-templates/scripts/lib/pipeline-common.sh +95 -21
- package/payload/ci-templates/scripts/lib/platform.sh +270 -13
- package/payload/ci-templates/scripts/lib/usage-capture-omp.sh +7 -1
- package/payload/ci-templates/scripts/lib/usage-capture.sh +7 -1
- package/payload/ci-templates/scripts/metrics-snapshot.sh +377 -0
- package/payload/ci-templates/scripts/orchestrate.sh +55 -37
- package/payload/ci-templates/scripts/postmortem.sh +10 -0
- package/payload/ci-templates/scripts/review-fix.sh +22 -11
- package/payload/ci-templates/scripts/review.sh +11 -7
- package/payload/ci-templates/scripts/test-fix.sh +7 -0
- package/payload/skills/agent-architect/SKILL.md +45 -0
- package/payload/skills/codebase-audit/SKILL.md +84 -0
- package/payload/skills/cve-fix/SKILL.md +98 -0
- package/payload/skills/docs-sync/SKILL.md +74 -0
- package/payload/skills/e2e-test-gen/SKILL.md +65 -0
- package/payload/skills/fix-review-findings/SKILL.md +2 -2
- package/payload/skills/fix-tests/SKILL.md +2 -2
- package/payload/templates/memory-index.template.md +34 -0
- package/payload/templates/pipeline-config.template.md +11 -1
- package/payload/templates/review_suppressions.template.md +55 -0
- package/payload/templates/spec-issue.template.md +39 -7
|
@@ -0,0 +1,45 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: agent-architect
|
|
3
|
+
description: Analyze bounded recent delivery evidence for concrete improvements to shipped agent definitions. Used by the opt-in scheduled CI job; returns only the report whose proposal blocks the runner files as ready improvement issues.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# agent-architect
|
|
7
|
+
|
|
8
|
+
This is the CI entry point for the shipped `agent-architect` subagent. The
|
|
9
|
+
runner owns forge reads, label management, issue creation, and all marker
|
|
10
|
+
files. Never call a forge API, create an issue, edit a file, commit, or push.
|
|
11
|
+
|
|
12
|
+
## Inputs
|
|
13
|
+
|
|
14
|
+
Read `$PIPE_CONTEXT_DIR/architect-context.md` first. It gives the bounded time
|
|
15
|
+
window, current ready/improvement labels, merged MR/PR summaries with changed
|
|
16
|
+
files and human review excerpts, stuck items, and already-open improvement
|
|
17
|
+
issues. Treat that prepared evidence as authoritative input; do not expand it
|
|
18
|
+
with forge calls or unbounded history scans.
|
|
19
|
+
|
|
20
|
+
Read `$PIPE_CONFIG_PATH` (default `.claude/pipeline-config.md`) before making
|
|
21
|
+
project-specific recommendations. If it is absent, say so in the report and
|
|
22
|
+
stay conservative.
|
|
23
|
+
|
|
24
|
+
## Analysis
|
|
25
|
+
|
|
26
|
+
Spawn the shipped `agent-architect` subagent once. Give it the context-file
|
|
27
|
+
path and require it to use the labels named in that context. It must inspect
|
|
28
|
+
only the repository files needed to validate evidence-backed, surgical changes
|
|
29
|
+
to `.claude/agents/*.md`.
|
|
30
|
+
|
|
31
|
+
Keep its report verbatim. It must contain exactly one top-level heading,
|
|
32
|
+
`## Section B - Agent Improvements`, and may contain no more than five complete
|
|
33
|
+
`<!-- ISSUE-START -->` / `<!-- ISSUE-END -->` proposal blocks. A proposal needs
|
|
34
|
+
a nonempty `Title: ` line of at most 72 characters and must cite the prepared
|
|
35
|
+
MR/PR or stuck-item evidence. Do not return commentary before or after the
|
|
36
|
+
report.
|
|
37
|
+
|
|
38
|
+
When no concrete proposal meets that bar, return the heading followed by
|
|
39
|
+
`No significant findings this week.`
|
|
40
|
+
|
|
41
|
+
## Result
|
|
42
|
+
|
|
43
|
+
Return only the subagent's markdown report. The runner persists it, bounds and
|
|
44
|
+
parses complete proposal blocks, writes `architect.env`, and in non-dry runs
|
|
45
|
+
creates the resulting issues with the configured ready and improvement labels.
|
|
@@ -0,0 +1,84 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: codebase-audit
|
|
3
|
+
description: Produce the periodic codebase convention-drift audit report and, when requested, apply its documentation suggestions, in CI. In `audit` mode runs the codebase-auditor agent against a summary of recently merged MRs/PRs and open-audit dedup context, decides the finding status and count, and writes the report and doc-suggestions files plus a machine-readable marker for the pipeline to act on. In `apply-docs` mode runs the coder agent to apply the suggested documentation updates within the configured documentation roots and commits. Use this when a CI job asks to run the codebase audit or apply its documentation suggestions.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# codebase-audit
|
|
7
|
+
|
|
8
|
+
Generic replacement for the prompt assembly and doc-suggestion extraction in `codebase-audit.sh`. The job runs `claude "/codebase-audit"` instead of building the auditor and coder prompts in bash. This skill does the reasoning; the runner keeps only the API reads that build the MR/PR summary, issue creation, and the push/MR-creation mechanics.
|
|
9
|
+
|
|
10
|
+
## Step 1: read the project configuration
|
|
11
|
+
|
|
12
|
+
Read the config file at `$PIPE_CONFIG_PATH` (default `.claude/pipeline-config.md`). In `audit` mode you need nothing from it directly: the `codebase-auditor` subagent reads it and its own Documentation Map itself. In `apply-docs` mode you need the Documentation Map, since the coder is restricted to `$PIPE_DOCS_ROOTS` and should map suggestions onto the right document. If the file is missing, say so on the first line of your output and continue with generic behavior.
|
|
13
|
+
|
|
14
|
+
## Step 2: inputs
|
|
15
|
+
|
|
16
|
+
Everything comes from CI variables and the files the job prepared. Do not call the platform API yourself, and never create the issue or the MR/PR — that is mechanical runner work, exactly like every other job in this plan.
|
|
17
|
+
|
|
18
|
+
- `$PIPE_CONTEXT_DIR` (default `build/pipeline`): where the prepared context lives and where you write outputs.
|
|
19
|
+
- `$PIPE_CONTEXT_DIR/audit-context.md`: the job wrote this, and rewrites it fresh before each invocation — the same path in both modes. Read it first; its `Mode:` line tells you which of the two flows below to run, and the rest of the file is that mode's own content only.
|
|
20
|
+
- `$PIPE_LABEL_AUDIT` (default `codebase-audit`): the label the runner uses for audit issues and MRs/PRs; only relevant if you need to refer to it in your own output.
|
|
21
|
+
- `$PIPE_DOCS_ROOTS` (default `docs,CLAUDE.md,.claude/memory`): comma-separated documentation roots, apply-docs mode only.
|
|
22
|
+
- `$PIPE_COMMIT_AUDIT` (default `docs: apply convention audit findings`): the exact commit subject, apply-docs mode only.
|
|
23
|
+
|
|
24
|
+
## Mode: audit
|
|
25
|
+
|
|
26
|
+
The context file lists the recently merged MRs/PRs (with up to 10 changed file paths each) and, when any exist, the currently open audit-labeled issues as dedup context.
|
|
27
|
+
|
|
28
|
+
### Step 3: run the codebase-auditor agent
|
|
29
|
+
|
|
30
|
+
Spawn the `codebase-auditor` subagent once. Point it at `$PIPE_CONTEXT_DIR/audit-context.md` for the recently-merged-MR summary and the open-issue dedup context, and tell it to skip any finding already covered by an open issue. The agent reads the config, the documented conventions, and the changed files of each listed MR/PR itself, and returns its report as its final output, exactly the two sections defined in its own instructions (`## Critical violations` and `## Patterns worth documenting`, each either a confidence->=80 bullet list or the literal line `No significant findings.`). Do not restate its internal scoring rules in the prompt.
|
|
31
|
+
|
|
32
|
+
### Step 4: decide the status and count the findings
|
|
33
|
+
|
|
34
|
+
Take the agent's report text verbatim (no reformatting) and write it to `$PIPE_CONTEXT_DIR/audit-report.md`. Then decide:
|
|
35
|
+
|
|
36
|
+
- The agent produced no output at all (empty or missing) -> `AUDIT_STATUS=no_output`, `AUDIT_FINDINGS_TOTAL=0`.
|
|
37
|
+
- Otherwise, count the finding bullets: lines starting with `- **` across the whole report (`grep -c '^- \*\*'` is exactly the source convention; a report where neither section has findings has zero such lines because both say `No significant findings.`). Zero findings -> `AUDIT_STATUS=no_findings`. One or more -> `AUDIT_STATUS=ok`.
|
|
38
|
+
|
|
39
|
+
### Step 5: extract the doc suggestions
|
|
40
|
+
|
|
41
|
+
Always write `$PIPE_CONTEXT_DIR/audit-doc-suggestions.md`, even when it ends up empty. Its content is the bullet lines under the `## Patterns worth documenting` heading only (stop at the next `## ` heading or end of report), verbatim, or nothing when that section is `No significant findings.`. This file is what the runner checks to decide whether to attempt the `apply-docs` phase; do not fabricate a suggestion that is not literally a bullet in that section.
|
|
42
|
+
|
|
43
|
+
### Step 6: emit the marker
|
|
44
|
+
|
|
45
|
+
Write `$PIPE_CONTEXT_DIR/audit.env`:
|
|
46
|
+
|
|
47
|
+
```
|
|
48
|
+
AUDIT_STATUS=ok|no_output|no_findings
|
|
49
|
+
AUDIT_FINDINGS_TOTAL=<int>
|
|
50
|
+
AUDIT_REPORT_PATH=$PIPE_CONTEXT_DIR/audit-report.md
|
|
51
|
+
```
|
|
52
|
+
|
|
53
|
+
Do not write an `AUDIT_ISSUE_IID` line yourself: no issue exists yet at this point in the flow. The runner creates the issue after you return (only when `AUDIT_STATUS=ok`), captures its IID, and appends `AUDIT_ISSUE_IID=<iid>` to this same file so the artifact is complete when the job finishes. Always write all three lines above, including on `no_output` and `no_findings`, so the runner can branch reliably.
|
|
54
|
+
|
|
55
|
+
## Mode: apply-docs
|
|
56
|
+
|
|
57
|
+
The context file gives you the issue number, the documentation roots, the commit subject, and the paths of the doc-suggestions file and the full audit report the `audit`-mode invocation wrote. The repository is already checked out on the audit branch; git is local, no token needed.
|
|
58
|
+
|
|
59
|
+
### Step 3: delegate to the coder
|
|
60
|
+
|
|
61
|
+
Spawn the `coder` subagent. Tell it plainly:
|
|
62
|
+
|
|
63
|
+
- it may edit ONLY paths under the documentation roots listed in the context file (each root is a directory or a single file) — no `src/`, no test code, no CI files, nothing else;
|
|
64
|
+
- implement ONLY the suggested updates listed in the doc-suggestions file (read it itself), using the full audit report (read it itself) for context;
|
|
65
|
+
- when done, commit with `git add -A && git commit` using EXACTLY the commit subject from the context file, followed by a blank line and `Closes #<issue number>` from the context file;
|
|
66
|
+
- do NOT run `git push` — the runner handles pushing.
|
|
67
|
+
|
|
68
|
+
Do not restate the coder's own internal workflow (documentation map lookup, verify command, etc.); those do not apply here since this is a documentation-only change.
|
|
69
|
+
|
|
70
|
+
### Step 4: confirm the commit and emit the marker
|
|
71
|
+
|
|
72
|
+
Run `git rev-list "origin/$PIPE_TARGET_BRANCH..HEAD" --count` to count the commits made this run. Write `$PIPE_CONTEXT_DIR/audit-apply.env`:
|
|
73
|
+
|
|
74
|
+
```
|
|
75
|
+
AUDIT_APPLY_COMMITS=<integer count of commits this run>
|
|
76
|
+
```
|
|
77
|
+
|
|
78
|
+
Zero means the coder made no changes (no suggestion turned out actionable, or it got stuck); do not draft anything else in that case, the runner handles the no-commits path.
|
|
79
|
+
|
|
80
|
+
## What the CI job does after you
|
|
81
|
+
|
|
82
|
+
In `audit` mode: the runner reads `audit.env`. On `no_output` it logs a failure notice (no issue exists yet to post to). On `no_findings` it stops, nothing filed. On `ok` it creates the findings issue from `audit-report.md` (truncated, with the `audited-up-to-mr-iid` marker prepended) labeled `$PIPE_LABEL_AUDIT`, appends the issue IID to `audit.env`, and emits a metric. It then skips the `apply-docs` phase entirely when there is no push-capable identity configured, or when `audit-doc-suggestions.md` is empty.
|
|
83
|
+
|
|
84
|
+
Otherwise it creates and checks out the `audit-<date>` branch and runs you again in `apply-docs` mode. It reads `audit-apply.env`, recounts the commits itself as the authoritative gate: with zero commits it posts a manual-review notice to the issue (with an extra credit-exhaustion note when that applies) and deletes the branch; otherwise it pushes the branch and opens an MR/PR titled `docs: codebase audit updates (<date>)` that closes the issue, labeled `$PIPE_LABEL_AUDIT`. None of that is your job.
|
|
@@ -0,0 +1,98 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: cve-fix
|
|
3
|
+
description: Safely update one configured dependency manifest for HIGH/CRITICAL vulnerability identifiers supplied by a CI scanner, compile it, make at most one commit, and emit a CVE remediation marker. Use only when the cve-fix CI runner has prepared its context.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# cve-fix
|
|
7
|
+
|
|
8
|
+
`cve-fix` is a code-writing skill called by the CI runner. It does not scan
|
|
9
|
+
projects, choose a vulnerability scanner, or filter severities: that stays
|
|
10
|
+
outside this skill. The caller supplies the already selected HIGH/CRITICAL
|
|
11
|
+
identifier list through `PIPE_CVE_LIST` and the runner context.
|
|
12
|
+
|
|
13
|
+
The runner owns MR/PR API access, rebase, push, and comments. Never do any of
|
|
14
|
+
those things yourself.
|
|
15
|
+
|
|
16
|
+
## Inputs and hard boundary
|
|
17
|
+
|
|
18
|
+
Read exactly one prepared context first:
|
|
19
|
+
|
|
20
|
+
- `$PIPE_CONTEXT_DIR/cve-fix-context.md` (default context directory is
|
|
21
|
+
`build/pipeline`).
|
|
22
|
+
|
|
23
|
+
It names the vulnerability-list file, configured dependency manifest path,
|
|
24
|
+
manifest-content file, compile command, exact commit subject, source and target
|
|
25
|
+
branches, suppression path, remaining commit budget, and marker path. Read the
|
|
26
|
+
named list and manifest-content files; do not independently run a scanner,
|
|
27
|
+
inspect scanner output, or expand the scope from repository search results.
|
|
28
|
+
|
|
29
|
+
Validate before editing:
|
|
30
|
+
|
|
31
|
+
- `Mode:` must be `apply`.
|
|
32
|
+
- The list must contain at least one non-blank identifier.
|
|
33
|
+
- The manifest path must be present and the corresponding file must exist.
|
|
34
|
+
- Source branch, target branch, commit subject, compile command, and remaining
|
|
35
|
+
commit budget must be present.
|
|
36
|
+
- The remaining budget must be a positive integer. A proposed change that would
|
|
37
|
+
exceed it must not be made.
|
|
38
|
+
|
|
39
|
+
If any validation fails, make no edit and no commit. Write the marker described
|
|
40
|
+
below with `CVE_MANUAL_REQUIRED=1`.
|
|
41
|
+
|
|
42
|
+
Only the configured dependency manifest may be changed. Do not modify the
|
|
43
|
+
vulnerability list, scanner output, configured suppression path, CI files,
|
|
44
|
+
source code, tests, lockfiles, generated files, or any other dependency
|
|
45
|
+
manifest. Do not use a project-specific ecosystem, build tool, package name,
|
|
46
|
+
or version as an assumption.
|
|
47
|
+
|
|
48
|
+
## Remediation
|
|
49
|
+
|
|
50
|
+
Use the shipped `coder` agent to reason from only the supplied vulnerability
|
|
51
|
+
list and manifest content. It must identify updates that are explicit and safe
|
|
52
|
+
in that manifest. Do not guess a package/version, edit a BOM or indirect
|
|
53
|
+
configuration outside the configured file, or make a change when the affected
|
|
54
|
+
dependency cannot be identified safely.
|
|
55
|
+
|
|
56
|
+
When a safe update is available:
|
|
57
|
+
|
|
58
|
+
1. Edit only the configured dependency manifest.
|
|
59
|
+
2. Run the supplied compile command exactly as configured.
|
|
60
|
+
3. If compilation succeeds, create at most one commit, using the exact supplied
|
|
61
|
+
subject and staging only that configured manifest.
|
|
62
|
+
4. Never push, rebase, post a comment, use a forge API, or change branches.
|
|
63
|
+
|
|
64
|
+
A failed compile, an unavailable safe update, a missing input, or uncertain
|
|
65
|
+
mapping from an identifier to a manifest dependency requires no commit and
|
|
66
|
+
manual follow-up.
|
|
67
|
+
|
|
68
|
+
## Marker
|
|
69
|
+
|
|
70
|
+
Always write `$PIPE_CONTEXT_DIR/cve-fix.env` (or the `Marker File:` path from
|
|
71
|
+
the context) as exactly these four newline-separated keys, with no other keys:
|
|
72
|
+
|
|
73
|
+
```dotenv
|
|
74
|
+
CVE_FIX_APPLIED=0|1
|
|
75
|
+
CVE_COMPILED=0|1
|
|
76
|
+
CVE_COUNT=<non-negative integer>
|
|
77
|
+
CVE_MANUAL_REQUIRED=0|1
|
|
78
|
+
```
|
|
79
|
+
|
|
80
|
+
`CVE_COUNT` is the number of non-blank vulnerability identifiers read from the
|
|
81
|
+
supplied list. `CVE_FIX_APPLIED=1` only after a real new commit with the exact
|
|
82
|
+
provided subject exists. `CVE_COMPILED=1` only after the supplied compile
|
|
83
|
+
command completed successfully. `CVE_MANUAL_REQUIRED=1` when no safe update is
|
|
84
|
+
available, required metadata is missing, validation fails, or compilation does
|
|
85
|
+
not pass. Do not claim a commit or compilation solely because an agent returned
|
|
86
|
+
successfully.
|
|
87
|
+
|
|
88
|
+
For a successful remediation the marker is:
|
|
89
|
+
|
|
90
|
+
```dotenv
|
|
91
|
+
CVE_FIX_APPLIED=1
|
|
92
|
+
CVE_COMPILED=1
|
|
93
|
+
CVE_COUNT=<count>
|
|
94
|
+
CVE_MANUAL_REQUIRED=0
|
|
95
|
+
```
|
|
96
|
+
|
|
97
|
+
The runner independently validates the marker, commit, compilation, rebase,
|
|
98
|
+
and push before it reports success.
|
|
@@ -0,0 +1,74 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: docs-sync
|
|
3
|
+
description: Check whether the code changes of a merge request or pull request require documentation updates, in CI. Runs the docs-sync agent, which applies the updates on an autonomous MR/PR and reports the gaps on a human-authored one, then emits the comment body and a machine-readable marker for the pipeline to act on. Use this when a CI job asks to sync documentation for an MR/PR.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# docs-sync
|
|
7
|
+
|
|
8
|
+
Generic replacement for the prompt assembly in `docs-sync.sh`. The CI job runs the harness with `/docs-sync` instead of building a long prompt in bash. This skill does the reasoning. The runner keeps only the API read that classifies the MR/PR, the diff computation, posting the comment, and committing the documentation edits.
|
|
9
|
+
|
|
10
|
+
## Step 1: read the project configuration
|
|
11
|
+
|
|
12
|
+
Read the config file at `$PIPE_CONFIG_PATH` (default `.claude/pipeline-config.md`). From it you need the Project stack, the Git & Platform section, and above all the Documentation Map, which says which document covers which topic. If the file is missing, say so on the first line of your output and continue with generic behavior: treat the documentation roots as the whole map.
|
|
13
|
+
|
|
14
|
+
## Step 2: inputs
|
|
15
|
+
|
|
16
|
+
Everything comes from CI variables and the files the job prepared. Do not call the platform API yourself, and do not run `git` to recompute the diff: the runner already scoped it.
|
|
17
|
+
|
|
18
|
+
- `$PIPE_CONTEXT_DIR` (default `build/pipeline`): where the prepared context lives and where you write outputs.
|
|
19
|
+
- `$PIPE_CONTEXT_DIR/docs-sync-context.md`: the mode, the MR/PR title, the documentation roots, and the paths of the two files below. Read it first.
|
|
20
|
+
- `$PIPE_CONTEXT_DIR/docs-sync-files.txt`: the changed files of the whole MR/PR.
|
|
21
|
+
- `$PIPE_CONTEXT_DIR/docs-sync-diff.patch`: the full `base..head` diff. It is never an incremental slice, so a gap introduced by an earlier push is still in front of you.
|
|
22
|
+
- `$PIPE_DOCS_ROOTS` (default `docs,CLAUDE.md,.claude/memory`): comma-separated documentation roots. These are the only paths that may be edited.
|
|
23
|
+
- `$PIPE_RESULT_DOCSSYNC` (default `$PIPE_CONTEXT_DIR/docs-sync.json`): where the result JSON goes.
|
|
24
|
+
|
|
25
|
+
The `Mode:` line of the context file is one of:
|
|
26
|
+
|
|
27
|
+
- `apply`: an autonomous MR/PR. The documentation updates are applied by editing files under the documentation roots.
|
|
28
|
+
- `report`: a human-authored MR/PR, or an autonomous one whose tip is already an automated docs commit. Gaps are reported and nothing is edited.
|
|
29
|
+
|
|
30
|
+
## Step 3: run the docs-sync agent
|
|
31
|
+
|
|
32
|
+
Spawn the `docs-sync` subagent once. Pass it:
|
|
33
|
+
|
|
34
|
+
- the mode and what it means (apply the updates by editing files under the documentation roots, or report the gaps without editing anything),
|
|
35
|
+
- the MR/PR title,
|
|
36
|
+
- the paths of the changed-files list and the diff, so it reads them itself rather than receiving them inline,
|
|
37
|
+
- the documentation roots it may edit, and the instruction that production source, tests and CI files are out of bounds and that it must not run `git`,
|
|
38
|
+
- the result file path, which is `$PIPE_RESULT_DOCSSYNC`.
|
|
39
|
+
|
|
40
|
+
The agent reads the config, the Documentation Map and the documents it needs itself, and applies its own documentation threshold. Do not restate its internal rules in the prompt, and do not preload documents into it.
|
|
41
|
+
|
|
42
|
+
## Step 4: check the boundary
|
|
43
|
+
|
|
44
|
+
Before you emit anything, verify the agent stayed inside the documentation roots. Any edit outside them is out of scope: revert it, and say so in the comment body. The runner stages and commits the documentation roots before it discards any remaining out-of-scope working-tree changes, so leaving an out-of-scope edit in place only loses the agent's own work.
|
|
45
|
+
|
|
46
|
+
In `report` mode nothing may be edited at all. If the agent edited a file anyway, revert every edit and report the gaps as text.
|
|
47
|
+
|
|
48
|
+
## Step 5: emit the result and the marker
|
|
49
|
+
|
|
50
|
+
The result JSON is the agent's own output; the body file and the dotenv are yours. If the agent wrote no result file, or wrote one that is not valid JSON with a `comment` string, rewrite it yourself from what the agent reported.
|
|
51
|
+
|
|
52
|
+
`$PIPE_RESULT_DOCSSYNC` (default `$PIPE_CONTEXT_DIR/docs-sync.json`):
|
|
53
|
+
|
|
54
|
+
```json
|
|
55
|
+
{"has_gaps": true, "changed": true, "comment": "**Status: blocking**\n\n- **docs/X.md**: Added the section for the Y pattern introduced here"}
|
|
56
|
+
```
|
|
57
|
+
|
|
58
|
+
`$PIPE_CONTEXT_DIR/docs-sync-body.md`: the value of `comment` as plain markdown, and nothing else. This is what gets posted, so the runner never has to parse a model-written JSON file. Its first line is the status line, exactly one of `**Status: blocking**`, `**Status: non-blocking**`, `**Status: clean**`, followed by a flat bullet list of findings, blocking items first. Each bullet names the document and describes in one sentence what is missing, or what you added in `apply` mode. No section headers, no prose around the list. When the documentation is up to date the body is exactly the status line `**Status: clean**`.
|
|
59
|
+
|
|
60
|
+
`$PIPE_CONTEXT_DIR/docs-sync.env`:
|
|
61
|
+
|
|
62
|
+
```
|
|
63
|
+
DOCS_HAS_GAPS=true|false
|
|
64
|
+
DOCS_CHANGED=true|false
|
|
65
|
+
DOCS_COMMENT_PATH=<path of the body file>
|
|
66
|
+
```
|
|
67
|
+
|
|
68
|
+
`DOCS_HAS_GAPS` is `true` only when at least one actionable gap was found; informational notes alone keep it `false`. `DOCS_CHANGED` is `true` only when you left an edit on disk, which can only happen in `apply` mode. `DOCS_COMMENT_PATH` is the absolute or repository-relative path of the body file you wrote.
|
|
69
|
+
|
|
70
|
+
Always write all three, including on a clean verdict. A missing body file is how the runner detects that the agent did not complete, and it then posts a manual-review notice instead of your verdict.
|
|
71
|
+
|
|
72
|
+
## What the CI job does after you
|
|
73
|
+
|
|
74
|
+
The runner reads `docs-sync.env`, posts the body as the MR/PR comment behind its own marker line, and emits the metrics event. In `apply` mode it stages the documentation roots and commits with `$PIPE_COMMIT_DOCSSYNC` only when their staged diff is non-empty; `DOCS_CHANGED` only reports whether the skill left an edit on disk. It then discards anything you left outside those roots, rebases and pushes. It never asks you to run `git`. None of that is your job.
|
|
@@ -0,0 +1,65 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: e2e-test-gen
|
|
3
|
+
description: Generate or repair one focused E2E test for a runner-selected issue, using the project's configured E2E conventions. This deliberately implements only issue-driven generation; project-specific test-gap detection is intentionally omitted.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# e2e-test-gen
|
|
7
|
+
|
|
8
|
+
`e2e-test-gen` is the issue-driven half of generic E2E generation. It
|
|
9
|
+
**deliberately omits test-gap detection**: scanner-like source/package heuristics
|
|
10
|
+
are project-specific and are not a missing capability to restore.
|
|
11
|
+
|
|
12
|
+
The CI runner owns the branch, git identity, commits, push, issue labels,
|
|
13
|
+
comments, and MR/PR creation. Never perform any of those actions. Do not call a
|
|
14
|
+
forge API.
|
|
15
|
+
|
|
16
|
+
## Read inputs before editing
|
|
17
|
+
|
|
18
|
+
Read `$PIPE_CONTEXT_DIR/e2e-gen-issue-<IID>.md` for the runner context. Its
|
|
19
|
+
`Mode:` is either `generate` or `repair`; in repair mode it names the only file
|
|
20
|
+
that may be edited. Then read the project pipeline configuration at the
|
|
21
|
+
context's `Project Configuration:` path and its **E2E Tests** section.
|
|
22
|
+
|
|
23
|
+
Before any edit, require all of these configuration values to be explicit:
|
|
24
|
+
|
|
25
|
+
- E2E framework;
|
|
26
|
+
- configured test directory, matching the runner context;
|
|
27
|
+
- fixture import module;
|
|
28
|
+
- test-data prefix;
|
|
29
|
+
- required tags; and
|
|
30
|
+
- base URL, matching the runner context.
|
|
31
|
+
|
|
32
|
+
If the E2E Tests section is absent, says E2E is unused, or any required value is
|
|
33
|
+
missing, make no edit. Do not infer a framework, fixtures, credentials, command,
|
|
34
|
+
or base URL.
|
|
35
|
+
|
|
36
|
+
## Generate mode
|
|
37
|
+
|
|
38
|
+
1. Read the bounded title and description files named in the context.
|
|
39
|
+
2. Read the existing fixture and one or two nearby tests to follow the project's
|
|
40
|
+
conventions, selectors, data cleanup, and tags.
|
|
41
|
+
3. Use the shipped `e2e-test-writer` agent to draft one focused scenario.
|
|
42
|
+
4. Write exactly one new test file beneath the configured test directory. Do not
|
|
43
|
+
modify production code, configuration, fixtures, helpers, existing tests, or
|
|
44
|
+
files outside that directory.
|
|
45
|
+
5. Use the fixture module, data prefix, required tags, and cleanup convention
|
|
46
|
+
from the E2E Tests section. Do not expose credentials or invent test data
|
|
47
|
+
outside the configured prefix.
|
|
48
|
+
|
|
49
|
+
## Repair mode
|
|
50
|
+
|
|
51
|
+
Read the named generated test and the relevant existing test/fixture convention.
|
|
52
|
+
Edit only that named generated test to address the failed verification. Do not
|
|
53
|
+
create files or change any other file.
|
|
54
|
+
|
|
55
|
+
## Result
|
|
56
|
+
|
|
57
|
+
Return a marker-readable final line only after a file was actually written:
|
|
58
|
+
|
|
59
|
+
```text
|
|
60
|
+
E2E_TEST_WRITTEN=<project-relative path>
|
|
61
|
+
```
|
|
62
|
+
|
|
63
|
+
In repair mode `<project-relative path>` must be the named generated test. This
|
|
64
|
+
line is a result for the runner to read; it does not authorize git, forge, or
|
|
65
|
+
label operations.
|
|
@@ -17,7 +17,7 @@ Read `$PIPE_CONFIG_PATH` (default `.claude/pipeline-config.md`). You need the Gi
|
|
|
17
17
|
- `$PIPE_VERIFY_CMD`: verification command (compile plus unit tests) to run after a fix.
|
|
18
18
|
- `$PIPE_FIX_LOOP_CAP` (default `2`): maximum consecutive review-fix commits before escalating.
|
|
19
19
|
- `$PIPE_COMMIT_REVIEWFIX`: exact commit subject for a review fix. Its recurrence drives the cap count.
|
|
20
|
-
- `$
|
|
20
|
+
- `$PIPE_COMMIT_DOCSSYNC`: exact docs-sync commit subject. Ignore this subject while walking the fix streak so a documentation commit does not reset the cap.
|
|
21
21
|
- `$PIPE_RESULT_REVIEWFIX` (default `$PIPE_CONTEXT_DIR/review-fix-result.json`): where the coder writes its structured result.
|
|
22
22
|
- `$PIPE_CONTEXT_DIR` (default `build/pipeline`): where you write the result marker.
|
|
23
23
|
|
|
@@ -29,7 +29,7 @@ Count consecutive commits from the branch tip back toward `$PIPE_TARGET_BRANCH`
|
|
|
29
29
|
git log "origin/$PIPE_TARGET_BRANCH"..HEAD --format="%s"
|
|
30
30
|
```
|
|
31
31
|
|
|
32
|
-
|
|
32
|
+
Walk the subjects in order. Increment the count only when a subject exactly equals `$PIPE_COMMIT_REVIEWFIX`. If a subject exactly equals `$PIPE_COMMIT_DOCSSYNC`, ignore it and continue walking. Stop at every other subject. If the count is at or above `$PIPE_FIX_LOOP_CAP`, do not attempt another fix. Escalate: follow the postmortem-mr flow (spawn the `postmortem` agent with the review body as the failure text, the fix-attempt git log, and the changed files, then read `failure_category`). Write the marker with `RF_ESCALATE=true` and the `FAILURE_CATEGORY`, and stop.
|
|
33
33
|
|
|
34
34
|
## Step 4: gather prior attempts
|
|
35
35
|
|
|
@@ -20,7 +20,7 @@ Read `$PIPE_CONFIG_PATH` (default `.claude/pipeline-config.md`). You need the Bu
|
|
|
20
20
|
- `$PIPE_FIX_LOOP_CAP` (default `2`): maximum consecutive bot fix commits before escalating.
|
|
21
21
|
- `$PIPE_COMMIT_TESTFIX`: exact commit subject for a test fix. Its recurrence drives the cap count.
|
|
22
22
|
- `$PIPE_COMMIT_COVERAGE`: exact commit subject for a coverage-tests commit. Separate cap counter.
|
|
23
|
-
- `$
|
|
23
|
+
- `$PIPE_COMMIT_DOCSSYNC`: exact docs-sync commit subject. Ignore this subject while walking either fix streak so a documentation commit does not reset the cap.
|
|
24
24
|
- `$PIPE_CONTEXT_DIR` (default `build/pipeline`): where you write the result marker.
|
|
25
25
|
|
|
26
26
|
## Step 3: pick the case
|
|
@@ -40,7 +40,7 @@ Count consecutive commits from the branch tip back toward `$PIPE_TARGET_BRANCH`
|
|
|
40
40
|
git log "origin/$PIPE_TARGET_BRANCH"..HEAD --format="%s"
|
|
41
41
|
```
|
|
42
42
|
|
|
43
|
-
|
|
43
|
+
Walk the subjects in order. Increment the count only when a subject exactly equals the relevant fix subject. If a subject exactly equals `$PIPE_COMMIT_DOCSSYNC`, ignore it and continue walking. Stop at every other subject. If the count is at or above `$PIPE_FIX_LOOP_CAP`, do not attempt another fix. Escalate instead: follow the postmortem-mr flow (spawn the `postmortem` agent with the failure text, the fix-attempt git log, and the changed files, then read `failure_category` from its report). Write the marker with `FIX_ESCALATE=true` and the `FAILURE_CATEGORY`, and stop. The job applies the stuck label and posts the diagnostic.
|
|
44
44
|
|
|
45
45
|
## Step 5: delegate
|
|
46
46
|
|
|
@@ -0,0 +1,34 @@
|
|
|
1
|
+
# Memory index
|
|
2
|
+
|
|
3
|
+
<!-- Created by the ci-agent-platform install. This is tier 3 of a three-tier
|
|
4
|
+
information architecture:
|
|
5
|
+
CLAUDE.md hardwired instructions, an index into the tiers below
|
|
6
|
+
docs/ long-term patterns, how-tos, architecture
|
|
7
|
+
.claude/memory/ active project state and cross-agent signals
|
|
8
|
+
|
|
9
|
+
Keep this file an index. Each entry is one line: the file, and what an
|
|
10
|
+
agent learns from it. Files here are project-owned. The installer never
|
|
11
|
+
rewrites them. -->
|
|
12
|
+
|
|
13
|
+
## Feedback signals
|
|
14
|
+
|
|
15
|
+
<!-- One file per lesson the pipeline taught the project, named
|
|
16
|
+
feedback_<topic>.md. Written after a postmortem or a repeated review
|
|
17
|
+
finding, read by the agents whose work it constrains. -->
|
|
18
|
+
(none yet)
|
|
19
|
+
|
|
20
|
+
## Project state
|
|
21
|
+
|
|
22
|
+
<!-- One file per moving part whose current status agents need, named
|
|
23
|
+
project_<topic>.md. -->
|
|
24
|
+
(none yet)
|
|
25
|
+
|
|
26
|
+
## Reference
|
|
27
|
+
|
|
28
|
+
<!-- Lookup tables that would otherwise be re-derived every run, named
|
|
29
|
+
reference_<topic>.md. -->
|
|
30
|
+
(none yet)
|
|
31
|
+
|
|
32
|
+
## Review suppressions
|
|
33
|
+
|
|
34
|
+
- `review_suppressions.md` - patterns the reviewers must not flag, with the reason and the date. Read by every reviewer agent before it reports anything.
|
|
@@ -76,6 +76,15 @@ TODO
|
|
|
76
76
|
## Suppressions
|
|
77
77
|
- Path: .claude/memory/review_suppressions.md (sections: Code / Security / Migration / Performance)
|
|
78
78
|
|
|
79
|
+
## Metrics Conventions
|
|
80
|
+
<!-- Optional. Only needed when the project runs the metrics-snapshot job.
|
|
81
|
+
The aggregator counts post-merge defects by matching commit subjects, so
|
|
82
|
+
the patterns have to be agreed rather than guessed. Defaults below match
|
|
83
|
+
the shipped aggregator. Mark "not used" when the job is off. -->
|
|
84
|
+
- Full revert of a merged change: `Reverts !<iid>`
|
|
85
|
+
- Targeted fix for a defect introduced by a merged change: `Fixes regression from !<iid>`
|
|
86
|
+
- Re-landing work that was reverted: `Reintroduces work from !<iid>`
|
|
87
|
+
|
|
79
88
|
## Release
|
|
80
89
|
<!-- Used by the release agent. Mark "not used" if the project has no formal release flow. -->
|
|
81
90
|
- Release title format: <!-- e.g. "Update to x.y.z" --> TODO
|
|
@@ -89,7 +98,8 @@ TODO
|
|
|
89
98
|
|
|
90
99
|
## E2E Tests
|
|
91
100
|
<!-- Only for projects with a generated e2e suite. Otherwise: not used.
|
|
92
|
-
|
|
101
|
+
Required: framework, test dir, fixture import module, test-data prefix,
|
|
102
|
+
cleanup rules, required tags, and the base URL used by generated tests.
|
|
93
103
|
Also (the e2e-test-writer relies on these):
|
|
94
104
|
- frontend source dirs + selector strategy (real element IDs vs test-id attributes),
|
|
95
105
|
- spec-file naming/extension convention (one scenario per file),
|
|
@@ -11,15 +11,70 @@
|
|
|
11
11
|
## Code Review Suppressions
|
|
12
12
|
|
|
13
13
|
(none yet)
|
|
14
|
+
<!--
|
|
15
|
+
### Generated output reports a duplicate helper
|
|
16
|
+
- **Pattern:** a generated file contains a helper that resembles a hand-written helper
|
|
17
|
+
- **Why:** the generator owns this output and the apparent duplicate is not maintained independently
|
|
18
|
+
- **Date:** 2026-08-28
|
|
19
|
+
-->
|
|
20
|
+
|
|
21
|
+
<!--
|
|
22
|
+
### Explicit fallback differs from the usual convention
|
|
23
|
+
- **Pattern:** a narrowly scoped fallback uses a documented exception to the usual convention
|
|
24
|
+
- **Why:** the deviation records a deliberate compatibility decision and is reviewed with its documented context
|
|
25
|
+
- **Date:** 2026-08-28
|
|
26
|
+
-->
|
|
27
|
+
|
|
14
28
|
|
|
15
29
|
## Security Review Suppressions
|
|
16
30
|
|
|
17
31
|
(none yet)
|
|
32
|
+
<!--
|
|
33
|
+
### Generated metadata includes an unused-looking field
|
|
34
|
+
- **Pattern:** generated metadata declares a field that no direct source reference uses
|
|
35
|
+
- **Why:** the generator requires the field for a downstream consumer, so removing it would make the output incomplete
|
|
36
|
+
- **Date:** 2026-08-28
|
|
37
|
+
-->
|
|
38
|
+
|
|
39
|
+
<!--
|
|
40
|
+
### Escaped value triggers a broad input warning
|
|
41
|
+
- **Pattern:** a value is passed through a safe escaping idiom that a generic catalog flags
|
|
42
|
+
- **Why:** the idiom applies the required encoding before the value reaches its destination
|
|
43
|
+
- **Date:** 2026-08-28
|
|
44
|
+
-->
|
|
45
|
+
|
|
18
46
|
|
|
19
47
|
## Migration Review Suppressions
|
|
20
48
|
|
|
21
49
|
(none yet)
|
|
50
|
+
<!--
|
|
51
|
+
### Generated history includes a repeated declaration
|
|
52
|
+
- **Pattern:** generated change history repeats a declaration already represented elsewhere
|
|
53
|
+
- **Why:** the generator preserves both representations so independent consumers can process the history
|
|
54
|
+
- **Date:** 2026-08-28
|
|
55
|
+
-->
|
|
56
|
+
|
|
57
|
+
<!--
|
|
58
|
+
### Intentional ordering differs from the default
|
|
59
|
+
- **Pattern:** a documented ordering intentionally differs from the usual sequence
|
|
60
|
+
- **Why:** the exception preserves a required transition and is limited to the documented operation
|
|
61
|
+
- **Date:** 2026-08-28
|
|
62
|
+
-->
|
|
63
|
+
|
|
22
64
|
|
|
23
65
|
## Performance Review Suppressions
|
|
24
66
|
|
|
25
67
|
(none yet)
|
|
68
|
+
<!--
|
|
69
|
+
### Generated lookup repeats a calculation
|
|
70
|
+
- **Pattern:** generated lookup code repeats a small calculation that resembles avoidable work
|
|
71
|
+
- **Why:** the generator emits the calculation to keep each generated branch independent and correct
|
|
72
|
+
- **Date:** 2026-08-28
|
|
73
|
+
-->
|
|
74
|
+
|
|
75
|
+
<!--
|
|
76
|
+
### Guarded reuse is reported as shared-state risk
|
|
77
|
+
- **Pattern:** a safe reuse idiom protected by an explicit guard is flagged by a generic catalog
|
|
78
|
+
- **Why:** the guard establishes the required boundary before reuse occurs
|
|
79
|
+
- **Date:** 2026-08-28
|
|
80
|
+
-->
|
|
@@ -1,6 +1,11 @@
|
|
|
1
1
|
<!--
|
|
2
|
-
Spec template for issues intended for autonomous implementation
|
|
3
|
-
|
|
2
|
+
Spec template for issues intended for autonomous implementation. Use it only
|
|
3
|
+
when requirements are settled and every mandatory section can be completed
|
|
4
|
+
unambiguously.
|
|
5
|
+
|
|
6
|
+
For early ideation or product discussion, use a feature request. For a
|
|
7
|
+
reproducible defect, use a bug report. For a documentation-only change, use
|
|
8
|
+
a docs update. Promote those to a Spec once requirements are settled.
|
|
4
9
|
|
|
5
10
|
IMPORTANT: the eight section headers below (## 1. Outcome ... ## 8. Out of
|
|
6
11
|
Scope) are load-bearing. The pipeline validates decomposed sub-issues against
|
|
@@ -12,6 +17,13 @@
|
|
|
12
17
|
GitHub default .github/ISSUE_TEMPLATE/spec.md).
|
|
13
18
|
-->
|
|
14
19
|
|
|
20
|
+
## Spec Readiness
|
|
21
|
+
|
|
22
|
+
- [ ] All mandatory sections below are filled in and unambiguous.
|
|
23
|
+
- [ ] I understand that applying this project's configured ready label hands this issue to an autonomous implementation agent.
|
|
24
|
+
|
|
25
|
+
---
|
|
26
|
+
|
|
15
27
|
> Fill this into the issue DESCRIPTION, not a comment. The pipeline reads only the description for the spec check.
|
|
16
28
|
|
|
17
29
|
## 1. Outcome
|
|
@@ -31,9 +43,11 @@ TODO
|
|
|
31
43
|
## 4. Scope Envelope
|
|
32
44
|
<!-- Be honest about size. One coder session handles roughly the capacity in the
|
|
33
45
|
pipeline config (default ~15 files / ~600 lines). Larger work is decomposed. -->
|
|
34
|
-
- Estimated files touched
|
|
35
|
-
-
|
|
36
|
-
-
|
|
46
|
+
- **Estimated files touched:** TODO
|
|
47
|
+
- **Packages / layers affected:** TODO
|
|
48
|
+
- **Database migration required:** TODO (yes / no)
|
|
49
|
+
- **User-interface changes required:** TODO (yes / no)
|
|
50
|
+
- **Fits in one coder session:** TODO (yes / no; if no, decompose first)
|
|
37
51
|
|
|
38
52
|
## 5. Existing Code Anchors
|
|
39
53
|
<!-- <at least one> pointer to model the implementation after: a class, a module,
|
|
@@ -43,9 +57,10 @@ TODO
|
|
|
43
57
|
|
|
44
58
|
## 6. Risk Flags
|
|
45
59
|
<!-- Check any that apply. Each checked box may trigger a specialist reviewer and
|
|
46
|
-
changes how the coder approaches the work.
|
|
60
|
+
changes how the coder approaches the work. Projects add domain-specific
|
|
61
|
+
risk flags in pipeline-config.md. -->
|
|
47
62
|
- [ ] Schema migration
|
|
48
|
-
- [ ] Security-sensitive (
|
|
63
|
+
- [ ] Security-sensitive (authentication, authorization, permissions, secrets)
|
|
49
64
|
- [ ] Public API contract change
|
|
50
65
|
- [ ] Performance-critical (hot path, large query, batch job, N+1-prone)
|
|
51
66
|
|
|
@@ -53,11 +68,28 @@ TODO
|
|
|
53
68
|
<!-- <what will be tested> and how we will know it works. Be specific. -->
|
|
54
69
|
- New tests to add: TODO
|
|
55
70
|
- Existing tests that must still pass: TODO
|
|
71
|
+
- Manual verification (if applicable): TODO
|
|
56
72
|
|
|
57
73
|
## 8. Out of Scope
|
|
58
74
|
<!-- What this issue deliberately does NOT include. Prevents scope creep. -->
|
|
59
75
|
- TODO
|
|
60
76
|
|
|
77
|
+
## Open Questions
|
|
78
|
+
<!-- Optional. If this section has unresolved items, keep iterating before
|
|
79
|
+
applying the configured ready label. -->
|
|
80
|
+
- TODO
|
|
81
|
+
|
|
82
|
+
## Decomposition Links
|
|
83
|
+
<!-- Optional. Track the parent issue and any child issues created from this
|
|
84
|
+
specification. -->
|
|
85
|
+
- Decomposed from: TODO
|
|
86
|
+
- Decomposes into: TODO
|
|
87
|
+
|
|
88
|
+
## References
|
|
89
|
+
<!-- Optional. Link related issues, change requests, design documents, and
|
|
90
|
+
external resources. -->
|
|
91
|
+
- TODO
|
|
92
|
+
|
|
61
93
|
<!-- Dependencies (GitHub has no native issue links): to say this issue waits on
|
|
62
94
|
another, write "Blocked by #N" here. The orchestrator's release poll reads
|
|
63
95
|
it back and promotes the issue once #N closes. On GitLab, native issue
|