argus-reviewer-e2e 0.2.0 → 0.3.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -1,4 +1,4 @@
1
- # argus-reviewer
1
+ # Argus
2
2
 
3
3
  <p align="center">
4
4
  <img src="docs/assets/social.png" alt="Argus — vision-model E2E testing" width="640" />
@@ -11,95 +11,108 @@
11
11
  <a href="https://github.com/duketopceo/Argus/security/policy"><img src="https://img.shields.io/badge/security-policy-orange" alt="security policy" /></a>
12
12
  </p>
13
13
 
14
- Open-source, self-hosted vision-model E2E testing — the hundred-eyed watcher for your UI. Bring your own `OPENROUTER_API_KEY`: record a flow once, fingerprint-cache every step, replay near-free, heal on UI drift, and get results as a check + comment on the GitHub PR.
14
+ **The hundred-eyed watcher for your pull requests.** Argus reviews your diff, then goes further: it starts your real app, clicks through it like a user, and executes probes against suspected bugs — then posts a verdict on the PR with the exact dollar cost. Self-hosted, MIT-licensed, bring-your-own OpenRouter key. No SaaS middleman, no telemetry, no per-seat pricing.
15
15
 
16
- - **Vision-first**: a model looks at a screenshot and decides where to click — no selectors to write or maintain.
17
- - **Cache-first**: replay costs zero vision calls on an unchanged UI; heals re-spend only on drift and show up as reviewable cache diffs.
18
- - **Cost-explicit**: every call is metered from OpenRouter's per-call cost and rolled into a per-run dollar figure on the PR.
19
- - **Grounding specialist**: a `grounding_model` (e.g. a ui-tars-class model) can drive element location with its native coordinate output, verified against the DOM before any click executes.
20
- - **Execution-backed review**: `code-review` findings carry CI evidence, and with the opt-in sandbox lane (`sandbox: { enabled: true }`) Argus authors a test probe for unexercised findings and runs it in a hardened, network-less Docker container — a finding that fails on head and passes on base is stamped **reproduced**, not just suspected.
16
+ ## Get started in 60 seconds
21
17
 
22
18
  ```bash
23
- npm i -D argus-reviewer-e2e # or github:duketopceo/argus-reviewer
24
- npx argus-reviewer record "log in and open settings" --url https://localhost:3000
25
- npx argus-reviewer run # replays + asserts, zero-cost on cache hit
19
+ npm i -D argus-reviewer-e2e # the npm package; the command it installs is `argus-reviewer`
20
+ npx argus-reviewer init # writes config + smoke test + GitHub workflow
26
21
  ```
27
22
 
23
+ Add `OPENROUTER_API_KEY` to your environment (and repo secrets for CI). Then:
24
+
25
+ ```bash
26
+ npx argus-reviewer record "sign in and open the dashboard" --url http://localhost:3000
27
+ npx argus-reviewer run # replays + asserts — free on cache hit
28
+ ```
29
+
30
+ That's it. `init` drops a ready-to-run GitHub workflow; every PR from then on gets a review comment with findings, flow results, video evidence, and spend.
31
+
32
+ ## What lands on your PR
33
+
34
+ A sticky comment that updates on every push:
35
+
36
+ - **Verdict** — `APPROVE` / `NEEDS_CHANGES` with findings linked to concrete source lines
37
+ - **Inline comments on every severity** — one batched PR review, severity-sorted; each finding can carry a committable `suggestion` block you apply in one click
38
+ - **Blocks only on proof** — the review escalates to `REQUEST_CHANGES` only for reproduced or adjudicated blockers; everything else stays advisory. Stale request-changes reviews are dismissed automatically, and `requestChanges: false` keeps it advisory forever
39
+ - **Flow results** — which recorded user-journeys passed, healed, or broke
40
+ - **Reproduced, not suspected** — opt-in sandbox lane runs authored regression probes; a probe that fails on head and passes on base stamps the finding as *proven*
41
+ - **Cost** — every model call metered from OpenRouter's per-call pricing, totaled in dollars
42
+
43
+ That comment is **not a review**. `require_approving_reviews` reads reviews only, so on a protected branch the verdict alone leaves the gate unsatisfied. Supply `approval-token` — a GitHub App installation token, or a PAT from an account that is not the PR author — and Argus also submits a real review whose event follows the verdict. `github.token` cannot do this; GitHub refuses it outright. A token alone is **not** enough: an `APPROVE` also needs `approval-evidence` naming the test command, and `approval-check` naming a check run that has completed green on the pull request's head commit, produced by the App named in `approval-check-app` (`github-actions` by default). Details, the measured refusals, and the review-discipline rules: [`docs/approval-token.md`](docs/approval-token.md).
44
+
28
45
  <p align="center">
29
46
  <img src="docs/assets/demo.gif" alt="argus-reviewer run — live vision call, PASS, $0.0005 spend" width="900" />
30
47
  </p>
31
48
 
32
- *Real `run` output: one vision assert, `PASS`, and the exact dollar figure on the run report.*
49
+ ## Why it's different
33
50
 
34
- Configuration lives in `argus-reviewer.config.ts` (a legacy `vision-e2e.config.*` is still accepted) — see `src/config.ts` for the full shape: `model`, `grounding_model`, `escalation_model`, `provider` routing rules, `budgetUsd`, `target`, `pageSetup`, `secrets`.
51
+ | | Argus |
52
+ |---|---|
53
+ | **Selectors** | None. A vision model looks at a screenshot and decides where to click. |
54
+ | **Maintenance** | Fingerprint cache replays at zero model cost; when the UI drifts, self-healing re-grounds and the heal shows up as a reviewable diff. |
55
+ | **Review depth** | Beyond the diff: full-source evidence linkage, CI evidence, executed probes, real browser runs. |
56
+ | **Spend** | You pick the models per lane (`model`, `grounding_model`, `code_model`, `escalation_model`) and set `budgetUsd`. Cache hits cost nothing. |
57
+ | **Data** | Yours. Keys, journals, videos, and reports stay on your infra. |
35
58
 
36
- ### OpenRouter cost attribution
59
+ ## Configuration
37
60
 
38
- Add an `openrouter` block to tag every request. `trace` is sent in the request body and is the right hook for cost allocation by repo/PR/run. `headers` are sent verbatim with every OpenRouter request (useful for `HTTP-Referer` or `X-Title`).
61
+ `argus-reviewer.config.ts`:
39
62
 
40
63
  ```ts
41
- export default {
42
- openrouter: {
43
- trace: { repo: 'duketopceo/myapp', pr: '42', run: 'argus-reviewer' },
44
- headers: { 'HTTP-Referer': 'https://github.com/duketopceo/myapp' },
45
- },
46
- }
64
+ import { defineConfig } from 'argus-reviewer-e2e'
65
+
66
+ export default defineConfig({
67
+ model: 'google/gemini-2.5-flash-lite', // vision grounding + actions
68
+ code_model: 'deepseek/deepseek-v4.1-flash', // diff review
69
+ escalation_model: 'anthropic/claude-sonnet-4', // risky/complex findings
70
+ budgetUsd: 1.0,
71
+ target: { url: 'https://your-app.example.com' },
72
+ testsDir: 'e2e',
73
+ })
47
74
  ```
48
75
 
49
- The GitHub Action automatically sets `ARGUS_REVIEWER_TRACE` with the repository, PR number, commit, and run id, so every PR review is attributed in OpenRouter without extra config. You can also set `ARGUS_REVIEWER_TRACE` yourself (JSON object) to add more fields.
76
+ Point `provider.order` at fast OpenRouter backends (`cerebras`, `groq`) for sub-second review calls — speed is a routing choice, not a pricing tier. Full shape: [`src/config.ts`](src/config.ts) (a legacy `vision-e2e.config.*` is still accepted). Setup walkthrough: [`docs/quickstart.md`](docs/quickstart.md).
77
+
78
+ ## The execution ladder
79
+
80
+ Argus does more as you grant it more access — each rung is opt-in:
81
+
82
+ 1. **API review** — GitHub token only. Reviews the PR diff and posts the verdict.
83
+ 2. **Trusted checkout** — findings get verified against the full source tree.
84
+ 3. **Browser flows** — Playwright drives your real app through recorded journeys.
85
+ 4. **Sandbox probes** — suspected findings get authored regression tests, executed in a hardened container (no network, no secrets, read-only FS). Fork PRs stay behind an `argus-probe` label gate.
86
+ 5. **Agent Zero delegation** — `argus-reviewer delegate "find the checkout bug"` hands exploratory work to your own A0 instance.
87
+
88
+ ## Cost attribution
89
+
90
+ Every OpenRouter call carries a trace tag. The action auto-sets `ARGUS_REVIEWER_TRACE` (repo, PR, commit, run id) so spend is attributable per review — or set it yourself for custom fields. See [`docs/quickstart.md`](docs/quickstart.md) for the `openrouter` config block.
91
+
92
+ ## Security model
50
93
 
51
- Status: early development. See `action/` for the composite GitHub Action,
52
- `runner/` for self-hosted runner registration, `docs/quickstart.md` for
53
- setup, `SECURITY.md` for the threat model, and `CONTRIBUTING.md` to hack
54
- on it.
94
+ Reviews run against hostile input by design: untrusted checkouts never execute config code, fork PRs are label-gated, secrets are filtered from model context and comment output, and the sandbox probe lane runs network-less with a read-only filesystem. Threat model: [`SECURITY.md`](SECURITY.md).
55
95
 
56
96
  ## File structure
57
97
 
58
98
  ```text
59
99
  argus-reviewer/
60
100
  ├── action/ # GitHub Actions composite action + sticky PR comment
61
- │ ├── action.yml
62
- │ └── sticky-comment.mjs
63
101
  ├── runner/ # Self-hosted runner registration docs + script
64
- │ ├── README.md
65
- │ └── register-runner.sh
66
102
  ├── electron/ # Local observability dashboard (`npm run app`)
67
103
  ├── src/
68
- │ ├── api.ts # Test-facing `test`/`td` API + generated test file renderer
104
+ │ ├── api.ts # Test-facing `test`/`td` API + generated test renderer
69
105
  │ ├── cli.ts # record · run · code-review · delegate · cache · index · init
70
- │ ├── config.ts # `argus-reviewer.config.*` loader (legacy `vision-e2e.config.*` accepted)
71
- │ ├── cache/
72
- │ │ ├── fingerprint.ts # Per-step screenshot/a11y fingerprint + resolve
73
- │ │ └── store.ts # Flow cache read/write
74
- │ ├── driver/
75
- │ │ ├── browser.ts # Playwright browser launch (chromium/firefox/webkit) + observation capture
76
- │ │ └── target.ts # Optional local dev-server target process
77
- │ ├── engine/
78
- │ │ ├── actions.ts # Low-level page actions (click, type, scroll, …)
79
- │ │ ├── loop.ts # Vision model record/replay + healing loop
80
- │ │ └── prompts.ts # OpenRouter action/assertion prompts + JSON schemas
81
- │ ├── evidence/
82
- │ │ ├── ci.ts # PR metadata + CI check-run context for findings
83
- │ │ ├── gate.ts # Fork-PR trust gate (argus-probe label bound to head SHA)
84
- │ │ └── link.ts # Finding → evidence linkage + comment-safe sanitization
85
- │ ├── executor/
86
- │ │ ├── a0.ts # `a0 headless -p` delegation to a user's Agent Zero instance
87
- │ │ └── sandbox.ts # Hardened Docker runner for generated probes
88
- │ ├── index/ # Repo index, diff context, cache invalidation
89
- │ ├── journal/ # Per-run structured journal entries
90
- │ ├── probe/
91
- │ │ ├── author.ts # Model-authored regression probe generation + validation
92
- │ │ ├── harness.ts # vitest/jest/node:test detection + TAP classification
93
- │ │ └── queue.ts # Head-vs-merge-base probe orchestration
94
- │ ├── report/
95
- │ │ ├── comment.ts # Markdown PR comment + commit-status rendering
96
- │ │ ├── junit.ts # JUnit XML output
97
- │ │ └── run.ts # JSON run report consumed by the action
98
- │ └── vision/
99
- │ ├── cost.ts # OpenRouter cost parsing per call
100
- │ ├── ledger.ts # Per-run USD budget tracking
101
- │ └── openrouter.ts # OpenRouter chat-completion client + schema parsing
102
- └── tests/ # Unit tests + small Playwright fixture page
106
+ │ ├── config.ts # `argus-reviewer.config.*` loader
107
+ │ ├── cache/ # Per-step fingerprint + flow store
108
+ │ ├── driver/ # Playwright browser + dev-server target
109
+ │ ├── engine/ # Vision record/replay + healing loop
110
+ │ ├── evidence/ # PR/CI context, fork trust gate, finding linkage
111
+ │ ├── executor/ # Agent Zero delegation + hardened probe sandbox
112
+ │ ├── probe/ # Model-authored regression probes
113
+ │ ├── report/ # PR comment, JUnit XML, run.json
114
+ │ └── vision/ # OpenRouter client, cost parsing, budget ledger
115
+ └── tests/ # Unit tests + Playwright fixture page
103
116
  ```
104
117
 
105
- License: MIT.
118
+ License: [MIT](LICENSE).
package/action/action.yml CHANGED
@@ -5,9 +5,13 @@ inputs:
5
5
  openrouter-api-key:
6
6
  description: OpenRouter API key (BYOK). All vision calls are billed through this key.
7
7
  required: true
8
+ argus-version:
9
+ description: Optional pinned argus-reviewer-e2e npm version. Empty installs the CLI package from this action ref with lifecycle scripts disabled.
10
+ default: ''
11
+ required: false
8
12
  cli:
9
- description: Command used to run argus-reviewer (defaults to `npx --no-install argus-reviewer`). For dogfooding this repo, use `node dist/cli.js` after `npm run build`.
10
- default: npx --no-install argus-reviewer
13
+ description: Optional executable plus arguments for a trusted local CLI override. Shell operators are rejected; empty installs the CLI package from this action ref.
14
+ default: ''
11
15
  required: false
12
16
  config:
13
17
  description: Path to the argus-reviewer config file
@@ -35,11 +39,8 @@ inputs:
35
39
  description: Playwright browser engine to install and run (chromium, firefox, or webkit). Should match config.browser.
36
40
  default: 'chromium'
37
41
  required: false
38
- cache-dependency-path:
39
- description: npm lockfile path for setup-node cache (defaults to <working-directory>/package-lock.json)
40
- required: false
41
42
  report-dir:
42
- description: Report output dir, passed to both run and code-review so the post step can find run.json/code-review.json regardless of config reportDir
43
+ description: Report output dir, passed to both run and code-review so the post step can find artifacts
43
44
  default: argus-reviewer-report
44
45
  required: false
45
46
  sandbox:
@@ -50,109 +51,174 @@ inputs:
50
51
  description: Cap on inline review comments posted per run; overflow is summarized in the sticky. Overrides review.maxComments when set.
51
52
  default: ''
52
53
  required: false
54
+ install-consumer-dependencies:
55
+ description: Explicitly install the consumer project's dependencies for a trusted runtime lane. Disabled by default so code review never runs PR-controlled lifecycle scripts beside secrets.
56
+ default: 'false'
57
+ required: false
53
58
  run:
54
- description: Run the browser-flow lane (`argus-reviewer run`). Set 'false' for code-review-only consumers — repos without a per-PR web target (CLIs, libraries, infra repos). Skips the Playwright install and run steps; the sticky comment and commit status then reflect code-review alone.
59
+ description: Run the browser-flow lane (`argus-reviewer run`). Disabled by default for public code-review-only repositories. Requires an explicit target and a trusted runtime path.
60
+ default: 'false'
61
+ required: false
62
+ approval-token:
63
+ description: >-
64
+ Optional token that lets Argus submit a *formal* pull request review, which
65
+ is what `require_approving_reviews` reads. `github.token` cannot be used:
66
+ GitHub refuses it with 422 "GitHub Actions is not permitted to approve
67
+ pull requests", and no permission scope changes that. Supply a GitHub App
68
+ installation token, or a PAT belonging to an account that is not the PR
69
+ author. Empty (the default) keeps Argus to the sticky comment only. See
70
+ docs/approval-token.md.
71
+ default: ''
72
+ required: false
73
+ approval-evidence:
74
+ description: >-
75
+ The test command(s) an approval stands on, cited verbatim in the review
76
+ body — e.g. 'python -m unittest discover -s tests && ruff check .'.
77
+ Required whenever approval-token is supplied: the lane refuses to submit
78
+ an approval that cites no command, so an approver is always re-runnable.
79
+ Ignored when approval-token is empty.
80
+ default: ''
81
+ required: false
82
+ approval-check:
83
+ description: >-
84
+ Name of a check run that must have completed successfully on the head
85
+ commit before Argus will submit an APPROVE, e.g. 'test (22)' or 'lint'.
86
+ This is what makes the cited test command evidence rather than a claim:
87
+ the review cites the real run. An APPROVE with no green run of this name
88
+ is refused. Ignored on the REQUEST_CHANGES and COMMENT paths, so a
89
+ negative review still goes out while CI is red.
90
+ default: ''
91
+ required: false
92
+ approval-check-app:
93
+ description: >-
94
+ App slug that is allowed to produce the approval-check run, e.g.
95
+ 'github-actions'. A check name is not a producer: a fork can add a
96
+ workflow that reports a green run with the expected name, so the match is
97
+ pinned to one App. Set this to your own App slug if your CI is not GitHub
98
+ Actions. Empty accepts any producer and is not the default.
99
+ default: 'github-actions'
100
+ required: false
101
+ emit-review:
102
+ description: >-
103
+ Submit the formal review when `approval-token` is supplied. 'false' keeps
104
+ the token out of the review lane entirely. A `true` value without a token
105
+ is a no-op, not an error.
55
106
  default: 'true'
56
107
  required: false
57
108
  outputs:
58
109
  conclusion:
59
110
  description: 'Run conclusion: success, failure, or neutral'
60
111
  value: ${{ steps.post.outputs.conclusion }}
112
+ review-event:
113
+ description: 'GitHub review event submitted: APPROVE, REQUEST_CHANGES, COMMENT, or none'
114
+ value: ${{ steps.approve.outputs.review-event }}
115
+ review-state:
116
+ description: 'State of the submitted review, or why none was submitted'
117
+ value: ${{ steps.approve.outputs.review-state }}
118
+ stale-dismissed:
119
+ description: 'Stale approvals by this identity that were dismissed because they no longer describe the head. "unknown" if the existing reviews could not be read.'
120
+ value: ${{ steps.approve.outputs.stale-dismissed }}
61
121
  runs:
62
122
  using: composite
63
123
  steps:
64
- - uses: actions/setup-node@v4
124
+ - uses: actions/setup-node@49933ea5288caeca8642d1e84afbd3f7d6820020 # v4
65
125
  with:
66
126
  node-version: ${{ inputs.node-version }}
67
- cache: npm
68
- cache-dependency-path: ${{ inputs.cache-dependency-path || format('{0}/package-lock.json', inputs.working-directory || '.') }}
69
127
 
70
- - name: Install consumer dependencies
128
+ - name: Prepare pinned Argus runtime
129
+ id: bootstrap
71
130
  shell: bash
72
- working-directory: ${{ inputs.working-directory }}
131
+ env:
132
+ ARGUS_ACTION_PATH: ${{ github.action_path }}
133
+ ARGUS_CLI: ${{ inputs.cli }}
134
+ ARGUS_ARGUS_VERSION: ${{ inputs.argus-version }}
135
+ ARGUS_CONFIG_PATH: ${{ inputs.config }}
136
+ ARGUS_WORKING_DIRECTORY: ${{ inputs.working-directory }}
137
+ ARGUS_REPORT_DIR: ${{ inputs.report-dir }}
138
+ ARGUS_BROWSER: ${{ inputs.browser }}
139
+ ARGUS_BUDGET_USD: ${{ inputs.budget-usd }}
140
+ ARGUS_MAX_COMMENTS: ${{ inputs.max-comments }}
141
+ run: node "$ARGUS_ACTION_PATH/bootstrap.mjs"
142
+
143
+ - name: Reject executable lanes on untrusted pull requests
144
+ if: inputs.run != 'false' || inputs.install-consumer-dependencies == 'true'
145
+ shell: bash
146
+ env:
147
+ EVENT_NAME: ${{ github.event_name }}
148
+ HEAD_FORK: ${{ github.event.pull_request.head.repo.fork }}
149
+ run: |
150
+ if [ "$EVENT_NAME" = "pull_request_target" ]; then
151
+ echo "Argus runtime lanes are disabled for pull_request_target." >&2
152
+ exit 1
153
+ fi
154
+ if [ "$EVENT_NAME" = "pull_request" ] && [ "$HEAD_FORK" != "false" ]; then
155
+ echo "Argus runtime lanes require a same-repository PR or an explicitly trusted workflow." >&2
156
+ exit 1
157
+ fi
158
+ if [ "$EVENT_NAME" != "pull_request" ] && [ "$EVENT_NAME" != "" ]; then
159
+ echo "Argus runtime lanes are disabled for unlisted CI events: $EVENT_NAME" >&2
160
+ exit 1
161
+ fi
162
+
163
+ - name: Install consumer dependencies for explicit runtime lanes
164
+ if: inputs.install-consumer-dependencies == 'true'
165
+ shell: bash
166
+ working-directory: ${{ inputs.working-directory || '.' }}
73
167
  run: npm ci
74
168
 
75
- - name: Install Playwright browser
169
+ - name: Install Playwright browser for explicit runtime lanes
76
170
  if: inputs.run != 'false'
77
171
  shell: bash
78
- working-directory: ${{ inputs.working-directory }}
172
+ working-directory: ${{ inputs.working-directory || '.' }}
79
173
  env:
80
174
  ARGUS_BROWSER: ${{ inputs.browser }}
81
175
  run: |
82
- case "$ARGUS_BROWSER" in
83
- chromium|firefox|webkit) ;;
84
- *) echo "invalid browser input: '$ARGUS_BROWSER' (expected chromium, firefox, or webkit)"; exit 1 ;;
85
- esac
86
176
  npx playwright install --with-deps "$ARGUS_BROWSER" || npx playwright install "$ARGUS_BROWSER"
87
177
 
88
- - name: Stage config
89
- shell: bash
90
- working-directory: ${{ inputs.working-directory }}
91
- run: |
92
- if [ "${{ inputs.config }}" != "argus-reviewer.config.ts" ]; then
93
- ext="${{ inputs.config }}"
94
- ext="${ext##*.}"
95
- case "$ext" in
96
- ts|json) cp "${{ inputs.config }}" "argus-reviewer.config.$ext" ;;
97
- *) echo "unsupported config extension: $ext" >&2; exit 1 ;;
98
- esac
99
- fi
100
-
101
178
  - name: Build repo index
102
179
  if: inputs.index == 'true'
103
180
  shell: bash
104
- working-directory: ${{ inputs.working-directory }}
181
+ working-directory: ${{ inputs.working-directory || '.' }}
105
182
  env:
106
- # Token only feeds trust resolution on issue_comment events —
107
- # pull_request* events read fork status from the event payload.
183
+ ARGUS_ACTION_PATH: ${{ github.action_path }}
108
184
  GITHUB_TOKEN: ${{ github.token }}
109
- run: ${{ inputs.cli }} index
185
+ ARGUS_CLI_JSON: ${{ steps.bootstrap.outputs.cli-json }}
186
+ ARGUS_WORKING_DIRECTORY: ${{ steps.bootstrap.outputs.working-directory }}
187
+ run: node "$ARGUS_ACTION_PATH/cli.mjs" index
110
188
 
111
- - name: Run argus-reviewer code review
112
- id: code-review
189
+ - name: Run selected Argus lanes
190
+ id: verify
113
191
  shell: bash
114
- working-directory: ${{ inputs.working-directory }}
192
+ working-directory: ${{ inputs.working-directory || '.' }}
115
193
  continue-on-error: true
116
194
  env:
195
+ ARGUS_ACTION_PATH: ${{ github.action_path }}
117
196
  OPENROUTER_API_KEY: ${{ inputs.openrouter-api-key }}
118
197
  GITHUB_TOKEN: ${{ github.token }}
198
+ ARGUS_CLI_JSON: ${{ steps.bootstrap.outputs.cli-json }}
199
+ ARGUS_WORKING_DIRECTORY: ${{ steps.bootstrap.outputs.working-directory }}
200
+ ARGUS_REPORT_DIR: ${{ inputs.report-dir }}
201
+ ARGUS_VERIFY_FLOW: ${{ inputs.run != 'false' && '1' || '' }}
119
202
  ARGUS_SANDBOX: ${{ inputs.sandbox == 'true' && '1' || '' }}
120
203
  ARGUS_MAX_COMMENTS: ${{ inputs.max-comments }}
121
204
  ARGUS_DEBUG: '1'
122
- ARGUS_REVIEWER_TRACE: >-
123
- {"repo":"${{ github.repository }}",
124
- "pr":"${{ github.event.pull_request.number }}",
125
- "commit":"${{ github.sha }}",
126
- "run_id":"${{ github.run_id }}",
127
- "run_attempt":"${{ github.run_attempt }}",
128
- "workflow":"${{ github.workflow }}"}
129
- run: ${{ inputs.cli }} code-review --report-dir "${{ inputs.report-dir }}"
130
-
131
- - name: Run argus-reviewer
132
- id: run
133
- if: inputs.run != 'false'
134
- shell: bash
135
- working-directory: ${{ inputs.working-directory }}
136
- continue-on-error: true
137
- env:
138
- OPENROUTER_API_KEY: ${{ inputs.openrouter-api-key }}
139
- GITHUB_TOKEN: ${{ github.token }}
140
- ARGUS_DEBUG: '1'
141
205
  ARGUS_DIFF_BASE: ${{ inputs.diff-base }}
142
206
  ARGUS_BUDGET_USD: ${{ inputs.budget-usd }}
143
207
  ARGUS_REVIEWER_TRACE: >-
144
208
  {"repo":"${{ github.repository }}",
145
209
  "pr":"${{ github.event.pull_request.number }}",
146
- "commit":"${{ github.sha }}",
210
+ "commit":"${{ github.event.pull_request.head.sha || github.sha }}",
147
211
  "run_id":"${{ github.run_id }}",
148
212
  "run_attempt":"${{ github.run_attempt }}",
149
213
  "workflow":"${{ github.workflow }}"}
150
- run: ${{ inputs.cli }} run --report-dir "${{ inputs.report-dir }}"
214
+ run: node "$ARGUS_ACTION_PATH/cli.mjs" verify --report-dir "$ARGUS_REPORT_DIR"
151
215
 
152
216
  - name: Post or update sticky PR comment
153
217
  id: post
154
- uses: actions/github-script@v7
218
+ if: always()
219
+ uses: actions/github-script@f28e40c7f34bde8b3046d885e986cb6290c5673b # v7
155
220
  env:
221
+ ARGUS_ACTION_PATH: ${{ github.action_path }}
156
222
  OPENROUTER_API_KEY: ${{ inputs.openrouter-api-key }}
157
223
  VISION_E2E_WORKING_DIR: ${{ inputs.working-directory }}
158
224
  ARGUS_REPORT_DIR: ${{ inputs.report-dir }}
@@ -160,10 +226,21 @@ runs:
160
226
  with:
161
227
  github-token: ${{ github.token }}
162
228
  script: |
163
- const fs = require('fs');
164
229
  const path = require('path');
165
- const actionPath = process.env.GITHUB_ACTION_PATH;
166
- const file = path.join(actionPath, 'sticky-comment.mjs');
167
- const code = fs.readFileSync(file, 'utf8');
168
- const runFn = new Function('github', 'context', 'core', 'require', `return (async () => {\n${code}\n})()`);
169
- return await runFn(github, context, core, require);
230
+ const actionPath = process.env.ARGUS_ACTION_PATH;
231
+ const { run } = require(path.join(actionPath, 'sticky-comment.cjs'));
232
+ return await run({ github, context, core });
233
+
234
+ - name: Submit formal approval review
235
+ id: approve
236
+ if: always() && inputs.emit-review != 'false'
237
+ shell: bash
238
+ env:
239
+ ARGUS_APPROVAL_TOKEN: ${{ inputs.approval-token }}
240
+ ARGUS_APPROVAL_EVIDENCE: ${{ inputs.approval-evidence }}
241
+ ARGUS_APPROVAL_CHECK: ${{ inputs.approval-check }}
242
+ ARGUS_APPROVAL_CHECK_APP: ${{ inputs.approval-check-app }}
243
+ ARGUS_REPORT_DIR: ${{ inputs.report-dir }}
244
+ VISION_E2E_WORKING_DIR: ${{ inputs.working-directory }}
245
+ ARGUS_HEAD_SHA: ${{ github.event.pull_request.head.sha || github.sha }}
246
+ run: node "${{ github.action_path }}/emit-review.mjs"