opencode-agent-skill 13.0.0-beta.2 → 14.4.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (102) hide show
  1. package/README.md +1571 -607
  2. package/bin/ocskill.mjs +172 -24
  3. package/docs/OPENCODE-COMPAT.md +34 -97
  4. package/docs/PI-COMPAT.md +203 -0
  5. package/docs/V14-CONTEXT-MEMORY-FABRIC.md +70 -0
  6. package/docs/V14.1-QUALITY-PERFORMANCE-FABRIC.md +114 -0
  7. package/docs/V14.2-TURBO-WEAK-MODEL-RUNTIME.md +448 -0
  8. package/evals/v14/tasks.json +46 -0
  9. package/global-config/agents/executor.md +7 -0
  10. package/global-config/agents/visual-verifier.md +22 -3
  11. package/global-config/plugins/ues-router/index.js +13 -8
  12. package/global-config/plugins/ues-router/policy-runtime.js +7 -0
  13. package/global-config/plugins/ues-router/router.js +12 -2
  14. package/global-config/skills/ecommerce-engineering/SKILL.md +1 -1
  15. package/global-config/skills/file-upload-engineering/SKILL.md +1 -1
  16. package/global-config/skills/git-safety/SKILL.md +1 -1
  17. package/global-config/skills/nestjs-engineering/SKILL.md +1 -1
  18. package/global-config/skills/performance-engineering/SKILL.md +1 -1
  19. package/global-config/skills/react-native-engineering/SKILL.md +1 -1
  20. package/global-config/skills/rest-api-design/SKILL.md +1 -1
  21. package/global-config/skills/ui-ux-engineering/SKILL.md +1 -1
  22. package/lib/adaptive-context-budget.mjs +97 -0
  23. package/lib/affected-tests.mjs +260 -0
  24. package/lib/benchmark-confidence.mjs +41 -2
  25. package/lib/browser-mcp-routing.mjs +166 -0
  26. package/lib/capability-fabric.mjs +359 -0
  27. package/lib/capability-registry.mjs +9 -0
  28. package/lib/code-intelligence/edit-anchor.mjs +102 -0
  29. package/lib/code-intelligence/index.mjs +95 -0
  30. package/lib/code-intelligence/lsp-provider.mjs +189 -0
  31. package/lib/completion-auditor.mjs +82 -0
  32. package/lib/context-engine-v11.mjs +65 -1
  33. package/lib/context-graph-rank.mjs +118 -0
  34. package/lib/context-manifest.mjs +97 -18
  35. package/lib/control-center.mjs +19 -1
  36. package/lib/document-ingestion.mjs +60 -0
  37. package/lib/dynamic-workflow.mjs +3 -1
  38. package/lib/evidence-store.mjs +82 -1
  39. package/lib/fast-verification-gate.mjs +69 -0
  40. package/lib/hierarchical-context.mjs +215 -0
  41. package/lib/mcp-health.mjs +144 -0
  42. package/lib/mcp-tool-policy.mjs +19 -0
  43. package/lib/memory-engine.mjs +493 -0
  44. package/lib/model-performance.mjs +33 -8
  45. package/lib/model-policy.mjs +3 -3
  46. package/lib/orchestrator-policy.mjs +5 -209
  47. package/lib/performance-fabric.mjs +229 -0
  48. package/lib/pi-rpc-pool.mjs +433 -0
  49. package/lib/process-hang-detector.mjs +83 -0
  50. package/lib/process-supervisor.mjs +193 -0
  51. package/lib/prompt-cache.mjs +27 -11
  52. package/lib/repo-graph.mjs +53 -2
  53. package/lib/reversible-context.mjs +49 -0
  54. package/lib/runtime-config.mjs +31 -0
  55. package/lib/safety.mjs +132 -0
  56. package/lib/semantic-index.mjs +52 -3
  57. package/lib/skill-compiler.mjs +128 -0
  58. package/lib/skill-quality.mjs +48 -2
  59. package/lib/task-engine.mjs +62 -5
  60. package/lib/task-policy.mjs +291 -0
  61. package/lib/verification-broker.mjs +284 -0
  62. package/lib/verification-command.mjs +111 -0
  63. package/lib/windows-shim.mjs +35 -0
  64. package/lib/workspace-fingerprint.mjs +198 -0
  65. package/package.json +52 -42
  66. package/pi/extensions/ues-child-runtime.ts +393 -0
  67. package/pi/extensions/ues.ts +3327 -0
  68. package/pi/prompts/ues-audit.md +9 -0
  69. package/pi/prompts/ues-critique.md +9 -0
  70. package/pi/prompts/ues-debug.md +9 -0
  71. package/pi/prompts/ues-feature.md +9 -0
  72. package/pi/prompts/ues-fix.md +9 -0
  73. package/pi/prompts/ues-plan.md +9 -0
  74. package/pi/prompts/ues-research.md +9 -0
  75. package/pi/prompts/ues-resume.md +9 -0
  76. package/pi/prompts/ues-review.md +7 -0
  77. package/pi/prompts/ues-run.md +17 -0
  78. package/pi/prompts/ues-verify.md +9 -0
  79. package/scripts/check-release-consistency.mjs +119 -185
  80. package/scripts/check-runtime-exports.mjs +66 -0
  81. package/scripts/check-source-integrity.mjs +246 -0
  82. package/scripts/eval-pi.mjs +492 -0
  83. package/scripts/install.mjs +16 -0
  84. package/scripts/smoke-package-closure.mjs +110 -0
  85. package/scripts/smoke-packed-install.mjs +24 -11
  86. package/scripts/smoke-pi-extension.mjs +144 -0
  87. package/scripts/uninstall.mjs +44 -0
  88. package/CHANGELOG.md +0 -415
  89. package/docs/DETERMINISTIC-TOOLS.md +0 -105
  90. package/docs/ENGINEERING-DESIGN.md +0 -194
  91. package/docs/EVALS.md +0 -158
  92. package/docs/GITHUB-RULESET.md +0 -50
  93. package/docs/NPM-PUBLISH.md +0 -116
  94. package/docs/RESEARCH-SOURCES.md +0 -37
  95. package/docs/TRACE-SCHEMA.md +0 -122
  96. package/docs/V11-PERCEPTION-ADAPTIVE-EXECUTION.md +0 -75
  97. package/docs/V11-PERCEPTION-ADAPTIVE.md +0 -220
  98. package/docs/V12-WEAK-MODEL-INTELLIGENCE.md +0 -27
  99. package/docs/V13-PARALLEL-WEAK-MODEL-RUNTIME.md +0 -86
  100. package/docs/V7-INTELLIGENCE-RUNTIME.md +0 -166
  101. package/docs/V8-INTELLIGENCE-RELIABILITY.md +0 -206
  102. package/docs/V9-SPEED-INTELLIGENCE.md +0 -102
@@ -1,166 +0,0 @@
1
- # UES 7.7 Intelligence Runtime
2
-
3
- UES 7.7 upgrades the V6 long-horizon harness into a more observable, crash-safe and adaptive engineering runtime. The model is still the model; UES improves how work is decomposed, bounded, verified, recovered and learned from.
4
-
5
- ## V7.0 — Runtime reliability
6
-
7
- - live evals use an async process runner with periodic heartbeats
8
- - hard timeout and idle timeout are separate
9
- - Ctrl+C aborts the active OpenCode process tree instead of leaving an orphan
10
- - durable tasks carry `runId`, owner metadata, heartbeat time and lease expiry
11
- - stale running tasks can be recovered with `ocskill work recover`
12
- - `ocskill work resume` performs stale-lease recovery before reporting ready work
13
- - the V2 plugin probes actual session/runtime capabilities before fresh dispatch
14
-
15
- Example:
16
-
17
- ```bash
18
- ocskill work status checkout .
19
- ocskill work recover checkout .
20
- ocskill work heartbeat checkout T1 . --run-id <run-id>
21
- ```
22
-
23
- ## V7.1 — Structured evidence
24
-
25
- `ocskill work verify-command` executes a concrete verification command and records a receipt in `EVIDENCE.json`.
26
-
27
- A receipt contains:
28
-
29
- - command and arguments
30
- - exit code and pass/fail
31
- - start/end/duration
32
- - SHA-256 of stdout/stderr rather than full output
33
- - workspace fingerprint before and after
34
- - task/runId binding
35
-
36
- This complements narrative evidence. Task state reports whether completed work is `receipt-backed` or `narrative`.
37
-
38
- ```bash
39
- ocskill work verify-command checkout T1 . --run-id <run-id> -- npm test
40
- ```
41
-
42
- ## V7.2 — Context intelligence
43
-
44
- Fresh executor context packs now include a bounded context manifest:
45
-
46
- - declared task files
47
- - local import neighbors and reverse importers
48
- - likely related tests
49
- - top-level repository instruction/manifests
50
- - bounded source excerpts
51
- - repository graph hotspots
52
- - accepted learning items relevant to the task
53
-
54
- The goal is to reduce rediscovery cost without dumping the whole repository into one prompt.
55
-
56
- ## V7.3 — Adaptive orchestration/model policy
57
-
58
- `ocskill task-policy <text>` deterministically classifies work by complexity/risk and recommends:
59
-
60
- - `inline`, `standard` or `long-horizon`
61
- - light/standard/heavy model tier
62
- - context budget
63
- - retry budget
64
- - whether plan/integration gates are required
65
-
66
- `ocskill model-policy <role> --attempt N --text <task>` combines role defaults, task risk and retry escalation.
67
-
68
- ## V7.4 — Isolated parallel execution primitives
69
-
70
- The task graph distinguishes reads from writes. By default linked worktrees are created in a sibling `.REPO.ues-sandboxes/` directory on the same drive, avoiding nested worktrees inside the main checkout:
71
-
72
- - read/read overlap can share a safe wave
73
- - write/read and write/write conflicts serialize
74
-
75
- For parallel writers UES provides Git worktree sandboxes:
76
-
77
- ```bash
78
- ocskill sandbox create <slug> <task-id> .
79
- ocskill sandbox list .
80
- ocskill sandbox remove <worktree-path> . --force
81
- ```
82
-
83
- Sandbox creation is deterministic infrastructure. Integration/merging remains an explicit parent responsibility; UES does not silently merge branches.
84
-
85
- ## V7.5 — Evidence-gated learning loop
86
-
87
- UES can mine its own eval traces for recurring failure patterns:
88
-
89
- ```bash
90
- ocskill learn analyze . --eval-dir .ues-evals
91
- ocskill learn status .
92
- ocskill learn accept <proposal-id> .
93
- ```
94
-
95
- Accepted lessons can appear in later context packs when relevant. UES never auto-edits skills or promotes a lesson from a single run without an explicit accept step.
96
-
97
- Current deterministic proposal classes include:
98
-
99
- - hard timeout
100
- - idle timeout
101
- - agent process failure
102
- - hidden grader failure
103
- - orchestration failure
104
- - telemetry parse drift
105
-
106
- ## V7.6 — Optional Hermes bridge
107
-
108
- Hermes is treated as an optional external executor, not embedded as another runtime layer.
109
-
110
- ```bash
111
- ocskill hermes status
112
- ocskill hermes prompt <slug> <task-id> .
113
- ```
114
-
115
- The bridge emits a bounded UES delegation prompt. Hermes must not independently mutate durable UES state, merge, push, publish or deploy.
116
-
117
- ## V7.7 — Control Center
118
-
119
- Generate a zero-dependency local dashboard:
120
-
121
- ```bash
122
- ocskill dashboard .
123
- ocskill dashboard . --serve --port 4177
124
- ```
125
-
126
- When served with `--serve`, the Control Center refreshes its data every few seconds without requiring a frontend build.
127
-
128
- The Control Center summarizes:
129
-
130
- - long-horizon work items and task status
131
- - attempts, blockers and receipt count
132
- - learning proposals/accepted lessons
133
- - recent baseline/UES evaluation summaries
134
-
135
- Generated state lives under `.ues-dashboard/` and is git-ignored.
136
-
137
- ## Live benchmark observability
138
-
139
- The live harness accepts:
140
-
141
- ```bash
142
- node scripts/eval-live.mjs \
143
- --suite long \
144
- --model provider/model \
145
- --mode both \
146
- --trials 3 \
147
- --heartbeat-ms 30000 \
148
- --idle-timeout-ms 300000 \
149
- --timeout-ms 900000
150
- ```
151
-
152
- Every active trial prints a start line and periodic heartbeat rather than appearing frozen.
153
-
154
- ## Safety boundaries
155
-
156
- V7.7 deliberately keeps these actions explicit:
157
-
158
- - merging sandbox branches
159
- - forceful Git/history actions
160
- - publishing/releases
161
- - deployment
162
- - destructive database/filesystem operations
163
- - automatic promotion of learned rules
164
- - automatic execution through Hermes
165
-
166
- The goal is a stronger runtime without removing human control over irreversible side effects.
@@ -1,206 +0,0 @@
1
- # UES 8.0 Intelligence & Reliability
2
-
3
- UES 8.0 focuses on runtime reliability, stronger evidence gates, context quality, safe parallelism and measurable learning. It does not change the underlying model; it improves how engineering work is selected, executed, verified, recovered and evaluated.
4
-
5
- ## 1. Hard evidence gates
6
-
7
- Long-horizon and high-risk work uses strict evidence policy.
8
-
9
- Plan approval requires a structured `plan-verification` receipt bound to the current `PLAN.json` hash:
10
-
11
- ```bash
12
- ocskill work gate-receipt checkout plan . --verifier ues-plan-checker --evidence "plan checker PASS" --out .ues-work/checkout/reports/plan-receipt.json
13
-
14
- ocskill work approve-plan checkout . --evidence "plan checker PASS" --receipt-file .ues-work/checkout/reports/plan-receipt.json
15
- ```
16
-
17
- Task completion requires a successful verification receipt for the active `runId`. In strict mode, the receipt's `workspaceAfter` must also equal the current workspace fingerprint.
18
-
19
- Integration PASS requires an `integration-verification` receipt bound to the current workspace fingerprint. `finalize` still rejects any later workspace change.
20
-
21
- ## 2. Durable runtime journal
22
-
23
- Each long work item now includes:
24
-
25
- ```text
26
- .ues-work/<slug>/
27
- SPEC.md
28
- PLAN.json
29
- STATE.json
30
- EVIDENCE.json
31
- EVENTS.jsonl
32
- tasks/
33
- reports/
34
- ```
35
-
36
- `EVENTS.jsonl` is append-only runtime evidence for:
37
-
38
- - work initialization;
39
- - plan import and approval;
40
- - task start/session binding/heartbeat;
41
- - verification receipts;
42
- - failure and stale recovery;
43
- - task completion;
44
- - integration verification;
45
- - finalization.
46
-
47
- Read recent events with:
48
-
49
- ```bash
50
- ocskill work events <slug> . --limit 100
51
- ```
52
-
53
- ## 3. Bounded executor lifecycle
54
-
55
- The OpenCode V2 dispatcher probes capabilities instead of assuming them from a version string.
56
-
57
- A fresh executor has:
58
-
59
- - a durable `runId`;
60
- - attached OpenCode session ID;
61
- - heartbeat and lease expiry;
62
- - bounded runtime;
63
- - `session.interrupt` on timeout/cancel;
64
- - task-scoped stale recovery;
65
- - process-tree cancellation for external eval processes.
66
-
67
- On Unix, timed-out external process trees receive SIGTERM followed by SIGKILL after a bounded grace period if necessary. Windows uses `taskkill /T /F`.
68
-
69
- The V2 plugin exposes task cancellation/recovery tools when the runtime supports session interruption.
70
-
71
- ## 4. Context manifest v3
72
-
73
- Context selection now combines:
74
-
75
- - declared task files;
76
- - local imports and reverse importers;
77
- - likely related tests;
78
- - nearby repository instructions/manifests;
79
- - current Git-changed files;
80
- - multilingual task terms;
81
- - symbol hits;
82
- - TF-IDF-style content relevance;
83
- - centered source excerpts around matched terms;
84
- - accepted benchmark-validated lessons.
85
-
86
- The context remains bounded by a per-task budget rather than dumping the whole repository.
87
-
88
- ## 5. Safer parallel writes
89
-
90
- Safe-wave analysis still serializes declared read/write conflicts.
91
-
92
- Writer tasks are isolated in Git worktrees by default when the root checkout is clean (unless isolation is explicitly disabled). This keeps the first writer off the canonical root as well as later concurrent writers. Integration:
93
-
94
- - captures tracked and untracked sandbox changes;
95
- - rejects overlap with dirty files in the root checkout;
96
- - applies the patch to the root only after explicit integration;
97
- - cleans temporary UES worktree branches;
98
- - refuses to delete non-UES branches.
99
-
100
- Manual flow:
101
-
102
- ```bash
103
- ocskill sandbox create <slug> <task-id> .
104
- ocskill sandbox integrate <worktree-path> .
105
- ocskill sandbox list .
106
- ```
107
-
108
- ## 6. Learning v2
109
-
110
- Evaluation failures are clustered into recurring patterns and candidate rules.
111
-
112
- ```bash
113
- ocskill learn analyze . --eval-dir .ues-evals
114
- ocskill learn accept <proposal-id> .
115
- ```
116
-
117
- Acceptance alone does not make a shadow-required lesson active. Promotion additionally requires measured benchmark improvement:
118
-
119
- ```bash
120
- ocskill learn promote <proposal-id> . --report .ues-evals/matrix/matrix-summary-<timestamp>.json
121
- ```
122
-
123
- Promotion reads the benchmark matrix artifact itself, verifies complete/equal baseline-vs-UES coverage, hashes the artifact, and refuses caller-supplied pass-rate claims. Only promoted lessons are eligible for future context retrieval.
124
-
125
- ## 7. Benchmark matrix
126
-
127
- Run standard, long-horizon and polyglot baseline-vs-UES evaluations:
128
-
129
- ```bash
130
- npm run evals:matrix -- --model provider/model --trials 3
131
- ```
132
-
133
- The matrix checks expected baseline/UES coverage before producing a summary. Options include:
134
-
135
- ```text
136
- --long-only
137
- --standard-only
138
- --polyglot-only
139
- --without-polyglot
140
- ```
141
-
142
- The polyglot suite adds eight tasks covering Python authorization, Java money validation, .NET authorization, Next.js API error handling, React Native platform logic, safe SQL migration, monorepo dependency boundaries and generated-contract discipline.
143
-
144
- Benchmark results are evidence for the measured tasks only. They are not evidence that UES converts one model into another model.
145
-
146
- ## 8. Control Center
147
-
148
- ```bash
149
- ocskill dashboard . --serve --port 4177
150
- ```
151
-
152
- V8 adds:
153
-
154
- - runtime event visibility;
155
- - verification-receipt inspection;
156
- - stale-task recovery action;
157
- - existing work/learning/eval summaries.
158
-
159
- Executor cancellation remains a runtime operation because a standalone dashboard server cannot safely interrupt an OpenCode session it does not own.
160
-
161
- ## 9. Package migration
162
-
163
- The npm package remains:
164
-
165
- ```text
166
- opencode-agent-skill
167
- ```
168
-
169
- Users already on 7.7.0 can update normally:
170
-
171
- ```bash
172
- ocskill update
173
- ```
174
-
175
- The installer now writes:
176
-
177
- ```text
178
- <!-- managed-by: opencode-agent-skill -->
179
- ```
180
-
181
- It still recognizes the former scoped package owner and marker, then migrates them during re-sync.
182
-
183
- ## 10. Release and supply-chain checks
184
-
185
- V8 adds:
186
-
187
- - CodeQL workflow;
188
- - dependency-review workflow;
189
- - Dependabot for GitHub Actions and npm;
190
- - package/tag version consistency guard;
191
- - tag-only npm publish workflow;
192
- - OIDC-only npm Trusted Publishing permissions (no long-lived NODE_AUTH_TOKEN);
193
- - fail-closed tag/version guard;
194
- - exact plain global-install compatibility smoke in addition to packed-install smoke.
195
-
196
- Trusted Publishing still requires the npm account-side trust relationship to be configured for `laivannha0202/opencode-agent-skill-` and `publish.yml`.
197
-
198
- ## Validation before release
199
-
200
- Run:
201
-
202
- ```bash
203
- npm run ci
204
- ```
205
-
206
- Then run a real benchmark matrix with the target model. Merge/publish only after local CI is green and benchmark output has been inspected.
@@ -1,102 +0,0 @@
1
- # UES 9.0 Speed & Intelligence
2
-
3
- UES 9.0 builds on the V8 evidence and reliability gates. The goal is not to make the underlying model generate tokens faster; it is to reduce time-to-correct-result by finding the right code sooner, shrinking unnecessary context, avoiding unsafe retries and requiring evidence before completion.
4
-
5
- ## Execution profiles
6
-
7
- UES classifies work into three execution profiles:
8
-
9
- - **FAST**: small/low-risk work, up to 2 skills, 12k context budget, targeted verification, no durable worktree/critic overhead by default.
10
- - **STANDARD**: medium work, up to 4 skills, 24k context budget, semantic+Git context and targeted/affected verification.
11
- - **DEEP**: long-horizon or high-risk work, up to 5 skills, 48k context budget, durable state, worktree isolation, critic/integration verification and full CI.
12
-
13
- Risk still overrides speed. Authentication, security, payment, schema/migration, production/deploy and public-contract work remains fail-closed.
14
-
15
- ## Incremental evidence index
16
-
17
- The V9 index caches bounded source metadata under `.ues-cache/semantic-index-v1.json`. Unchanged files reuse cached evidence while changed files are reparsed. It records:
18
-
19
- - concrete source paths;
20
- - bounded symbol definitions with line numbers;
21
- - lexical identifier counts;
22
- - cache reuse/reparse statistics.
23
-
24
- This is intentionally labelled **syntax-aware lexical evidence**. It must not be presented as proof of program semantics or as a full AST/LSP call graph.
25
-
26
- Commands:
27
-
28
- ```bash
29
- ocskill index status .
30
- ocskill index build .
31
- ocskill index rebuild .
32
- ocskill aci search "checkout total" .
33
- ocskill aci refs calculateTotal .
34
- ocskill aci view src/checkout.ts . --line 120 --lines 80
35
- ocskill aci text "exact marker" .
36
- ```
37
-
38
- ## Weak-model ACI
39
-
40
- The ACI keeps search/view output bounded and evidence-first. Path traversal outside the repository is rejected. Large/binary files are refused by the bounded viewer. Reference results distinguish lexical references from concrete definition lines so an agent cannot honestly claim deeper semantic certainty than the tool measured.
41
-
42
- ## Runtime traces
43
-
44
- Operational trace events are appended under `.ues-traces/*.jsonl`. Common credentials/tokens are redacted before persistence, oversized payloads are hashed/truncated, and traces explicitly exclude hidden chain-of-thought.
45
-
46
- ```bash
47
- ocskill trace show <trace-id> .
48
- ```
49
-
50
- ## Verification sandbox
51
-
52
- When Docker or Podman is available, deterministic verification commands can run with:
53
-
54
- - network disabled by default;
55
- - Linux capabilities dropped;
56
- - no-new-privileges;
57
- - read-only container root;
58
- - bounded PIDs/memory/CPU;
59
- - only the requested workspace bind-mounted;
60
- - no host secrets forwarded by default.
61
-
62
- This protects verification commands; it does **not** claim to sandbox the OpenCode model process itself.
63
-
64
- ## Benchmark confidence gate
65
-
66
- The matrix now embeds paired baseline/UES evidence and computes:
67
-
68
- - paired wins/losses/ties;
69
- - pass-rate delta;
70
- - exact two-sided sign-test p-value;
71
- - per-suite regression checks;
72
- - mean duration ratio;
73
- - optional cost ratio.
74
-
75
- Normal reporting:
76
-
77
- ```bash
78
- npm run evals:matrix -- --model provider/model --trials 3
79
- ```
80
-
81
- Fail-closed release gate:
82
-
83
- ```bash
84
- npm run evals:matrix:gate -- --model provider/model --trials 3
85
- ```
86
-
87
- The gate requires enough paired samples, positive uplift, more UES wins than losses, statistical support, no suite regression and acceptable speed. This is evidence for the measured benchmark only; it is not evidence that UES turns a weaker model into a different model.
88
-
89
- ## Lock and Windows hardening
90
-
91
- State locking now uses owner tokens and heartbeats. Stale takeover renames the old lock before removal, and release only removes a lock whose token still belongs to the caller. Windows CLI/router invocation resolves Node-backed shims and refuses unrecognized batch shims rather than falling back to shell execution.
92
-
93
- ## Release validation
94
-
95
- Before publishing 9.0.0:
96
-
97
- ```bash
98
- npm run ci
99
- npm run evals:matrix:gate -- --model provider/model --trials 3
100
- ```
101
-
102
- If GitHub-hosted Actions still fails before the first workflow step, treat that as an infrastructure/repository-action issue rather than test evidence; local CI and the paired benchmark remain required release gates.