@plainconceptsplatform/workflows 0.27.6 → 0.28.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
|
@@ -2114,6 +2114,49 @@ if worker_installed implement; then
|
|
|
2114
2114
|
if [ "$ENDINGS_OK" -eq 1 ]; then PASS=$((PASS + 1)); else FAIL=$((FAIL + 1)); fi
|
|
2115
2115
|
fi
|
|
2116
2116
|
|
|
2117
|
+
echo "── Refine size gate ──────────────────────────────────────────────────────"
|
|
2118
|
+
|
|
2119
|
+
# The size gate added a job that skips in the happy path, and gh-aw adds every custom job to
|
|
2120
|
+
# the agent's needs. Three claims about that shape are asserted, because each one broke in
|
|
2121
|
+
# production in the same week (Pliny-Bot #320-322: sixteen runs, three parked issues, a
|
|
2122
|
+
# refusal loop on an issue the author was shrinking live):
|
|
2123
|
+
#
|
|
2124
|
+
# 1. the agent's gate carries !failure(), so a skipped refusal no longer poisons the
|
|
2125
|
+
# implicit success() and kills the agent in 0s
|
|
2126
|
+
# 2. the incomplete job excludes the refusal, so a refused issue is terminal: no attempt
|
|
2127
|
+
# counter, no x/5, no re-dispatch, no stalled
|
|
2128
|
+
# 3. the refusal itself releases bot-working, so a refused issue is not left reserved
|
|
2129
|
+
if worker_installed refine; then
|
|
2130
|
+
SIZE_GATE_OK=1
|
|
2131
|
+
REFINE_WORKER_MD="${WORKFLOWS_DIR}/agent-refine.md"
|
|
2132
|
+
|
|
2133
|
+
# The top-level if: starts at column 0 -- job-level ifs: are indented -- so anchor on that,
|
|
2134
|
+
# not on the prose comment above it. It must end with !failure().
|
|
2135
|
+
top_level_if="$(sed -n 's/^if: //p' "$REFINE_WORKER_MD" | head -1)"
|
|
2136
|
+
if [[ "$top_level_if" != *'!failure()' ]]; then
|
|
2137
|
+
SIZE_GATE_OK=0
|
|
2138
|
+
echo "FAIL: refine gates the agent without !failure(); a skipped refuse_big_issue poisons the implicit success() and skips the agent on every under-limit issue" >&2
|
|
2139
|
+
fi
|
|
2140
|
+
|
|
2141
|
+
# Extract the incomplete job's own block. A whole-file grep cannot make this claim: reserve,
|
|
2142
|
+
# refuse_big_issue and the agent gate all read size_guard, so the exclusion clause could be
|
|
2143
|
+
# deleted from incomplete while every grep still passes.
|
|
2144
|
+
incomplete_block="$(awk '/^ incomplete:/{found=1; next} found && /^ [a-z_]+:/{exit} found{print}' "$REFINE_WORKER_MD")"
|
|
2145
|
+
if ! printf '%s' "$incomplete_block" | grep -qE '^ needs: \[.*size_guard' \
|
|
2146
|
+
|| ! printf '%s' "$incomplete_block" | grep -qF "needs.size_guard.outputs.too_big != 'true'"; then
|
|
2147
|
+
SIZE_GATE_OK=0
|
|
2148
|
+
echo "FAIL: refine's incomplete does not exclude the refusal path; every refused issue also loops attempts" >&2
|
|
2149
|
+
fi
|
|
2150
|
+
|
|
2151
|
+
refuse_block="$(awk '/^ refuse_big_issue:/{found=1; next} found && /^ [a-z_]+:/{exit} found{print}' "$REFINE_WORKER_MD")"
|
|
2152
|
+
if ! printf '%s' "$refuse_block" | grep -qF 'WORKING_LABEL'; then
|
|
2153
|
+
SIZE_GATE_OK=0
|
|
2154
|
+
echo "FAIL: the refine refusal never releases bot-working; a refused issue stays reserved with no run behind it" >&2
|
|
2155
|
+
fi
|
|
2156
|
+
|
|
2157
|
+
if [ "$SIZE_GATE_OK" -eq 1 ]; then PASS=$((PASS + 1)); else FAIL=$((FAIL + 1)); fi
|
|
2158
|
+
fi
|
|
2159
|
+
|
|
2117
2160
|
echo "── Runner pools ──────────────────────────────────────────────────────────"
|
|
2118
2161
|
|
|
2119
2162
|
# Where every job runs, stated once and asserted, because GitHub gives a wrong pool no error: a
|
|
@@ -32,6 +32,14 @@ env:
|
|
|
32
32
|
REFINE_ATTEMPT_MARKER: "<!-- agent-refine-attempt -->"
|
|
33
33
|
MAX_ATTEMPTS: "5"
|
|
34
34
|
PARK_AT_ATTEMPT: "4"
|
|
35
|
+
# The size gate. An issue over either limit is refused before any model runs: a real
|
|
36
|
+
# one, Pliny-Bot #305 (13,398 chars, 92 work units), burned a full agent -- 17 minutes,
|
|
37
|
+
# 2.4M tokens -- and died at the provider's 32K output cap with zero outcome, twice.
|
|
38
|
+
# Both numbers are measured before the agent starts, so refusal is deterministic and
|
|
39
|
+
# costs ten seconds instead of half an hour.
|
|
40
|
+
REFINE_MAX_BODY_CHARS: "8000"
|
|
41
|
+
REFINE_MAX_WORK_UNITS: "24"
|
|
42
|
+
REFUSE_MARKER: "<!-- agent-refine-refused -->"
|
|
35
43
|
# Ten rather than six: the failure that prompted this took three minutes, and a slower one on a
|
|
36
44
|
# worse day would fall outside a six-minute window and park for a fault that clears by itself.
|
|
37
45
|
RETRY_UNDER_MINUTES: "10"
|
|
@@ -98,12 +106,12 @@ on:
|
|
|
98
106
|
required: false
|
|
99
107
|
type: string
|
|
100
108
|
default: first
|
|
101
|
-
# The gate
|
|
102
|
-
# activation job but gives activation no dependency on the
|
|
109
|
+
# The gate jobs that the top-level `if:` reads. gh-aw folds that `if:` into the generated
|
|
110
|
+
# activation job but gives activation no dependency on the jobs, so the reference resolves
|
|
103
111
|
# to '' and the clause is false -- the agent would never run. The package's own validator
|
|
104
112
|
# catches it after compilation; this is the line it asks for, the same one the merge gate
|
|
105
113
|
# uses for protected_changes.
|
|
106
|
-
needs: [still_open]
|
|
114
|
+
needs: [still_open, size_guard]
|
|
107
115
|
|
|
108
116
|
jobs:
|
|
109
117
|
# A route dispatched while the issue was open must not execute after it has been closed. The
|
|
@@ -130,10 +138,112 @@ jobs:
|
|
|
130
138
|
with:
|
|
131
139
|
token: ${{ github.token }}
|
|
132
140
|
issue-number: ${{ inputs.issue-number }}
|
|
133
|
-
|
|
141
|
+
# The size gate, rung 4. Measures the issue before any model starts, because an oversized
|
|
142
|
+
# body does not fail fast on its own: it produces an agent that explores for seventeen
|
|
143
|
+
# minutes and then dies mid-generation at the provider's response cap having written
|
|
144
|
+
# nothing, which reads as a green run with no outcome (Pliny-Bot #305, twice). Pure
|
|
145
|
+
# shell, no network beyond one issue read; over either limit it is a refusal, not a
|
|
146
|
+
# smaller attempt.
|
|
147
|
+
size_guard:
|
|
134
148
|
needs: [still_open]
|
|
135
149
|
if: needs.still_open.outputs.open == 'true'
|
|
136
150
|
runs-on: agents-arc
|
|
151
|
+
timeout-minutes: 5
|
|
152
|
+
permissions:
|
|
153
|
+
contents: read
|
|
154
|
+
issues: read
|
|
155
|
+
outputs:
|
|
156
|
+
too_big: ${{ steps.measure.outputs.too_big }}
|
|
157
|
+
reason: ${{ steps.measure.outputs.reason }}
|
|
158
|
+
steps:
|
|
159
|
+
- name: Measure the issue against the size limits
|
|
160
|
+
id: measure
|
|
161
|
+
env:
|
|
162
|
+
GH_TOKEN: ${{ github.token }}
|
|
163
|
+
REPO: ${{ github.repository }}
|
|
164
|
+
ISSUE_NUMBER: ${{ inputs.issue-number }}
|
|
165
|
+
MAX_BODY_CHARS: ${{ env.REFINE_MAX_BODY_CHARS }}
|
|
166
|
+
MAX_WORK_UNITS: ${{ env.REFINE_MAX_WORK_UNITS }}
|
|
167
|
+
run: |
|
|
168
|
+
set -euo pipefail
|
|
169
|
+
body=$(gh issue view "$ISSUE_NUMBER" --repo "$REPO" --json title,body --jq '.title + "\n\n" + (.body // "")')
|
|
170
|
+
chars=$(printf '%s' "$body" | wc -c)
|
|
171
|
+
# One work unit per markdown bullet, the same split the prompt's step 3 makes.
|
|
172
|
+
units=$(printf '%s' "$body" | grep -cE '^[[:space:]]*[-*] ' || true)
|
|
173
|
+
reason=""
|
|
174
|
+
too_big=false
|
|
175
|
+
if [ "$chars" -gt "$MAX_BODY_CHARS" ]; then
|
|
176
|
+
too_big=true
|
|
177
|
+
reason="body ${chars} chars > ${MAX_BODY_CHARS}"
|
|
178
|
+
fi
|
|
179
|
+
if [ "$units" -gt "$MAX_WORK_UNITS" ]; then
|
|
180
|
+
too_big=true
|
|
181
|
+
reason="${reason:+$reason; }work units ${units} > ${MAX_WORK_UNITS}"
|
|
182
|
+
fi
|
|
183
|
+
echo "too_big=$too_big" >> "$GITHUB_OUTPUT"
|
|
184
|
+
echo "reason=$reason" >> "$GITHUB_OUTPUT"
|
|
185
|
+
echo "::notice::measured $chars chars, $units work units (limits: ${MAX_BODY_CHARS} chars, ${MAX_WORK_UNITS} units) -> $too_big"
|
|
186
|
+
# The deterministic refusal. Runs only when the size gate tripped, so it costs one
|
|
187
|
+
# comment and no model. The refine label stays: a human who shrinks the body and
|
|
188
|
+
# replies re-enters rerefine through the comment route, and the classifier ignores
|
|
189
|
+
# bot comments, so this comment cannot re-trigger the worker.
|
|
190
|
+
refuse_big_issue:
|
|
191
|
+
needs: [still_open, size_guard]
|
|
192
|
+
if: needs.still_open.outputs.open == 'true' && needs.size_guard.outputs.too_big == 'true'
|
|
193
|
+
runs-on: agents-arc
|
|
194
|
+
timeout-minutes: 5
|
|
195
|
+
permissions:
|
|
196
|
+
contents: read
|
|
197
|
+
issues: write
|
|
198
|
+
steps:
|
|
199
|
+
- name: Checkout workflow actions
|
|
200
|
+
uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
|
|
201
|
+
with:
|
|
202
|
+
persist-credentials: false
|
|
203
|
+
- name: Create bot token
|
|
204
|
+
id: app-token
|
|
205
|
+
uses: actions/create-github-app-token@bcd2ba49218906704ab6c1aa796996da409d3eb1 # v3.2.0
|
|
206
|
+
with:
|
|
207
|
+
client-id: ${{ secrets.BOT_APP_ID }}
|
|
208
|
+
private-key: ${{ secrets.BOT_PRIVATE_KEY }}
|
|
209
|
+
- name: Refuse the oversized issue
|
|
210
|
+
uses: ./.github/actions/create-issue-comment
|
|
211
|
+
with:
|
|
212
|
+
token: ${{ steps.app-token.outputs.token }}
|
|
213
|
+
issue-number: ${{ inputs.issue-number }}
|
|
214
|
+
body: |
|
|
215
|
+
${{ env.REFUSE_MARKER }}
|
|
216
|
+
This issue is too large to refine automatically: ${{ needs.size_guard.outputs.reason }}.
|
|
217
|
+
A story this size does not fail fast; it burns a full agent run and dies mid-generation at the provider's response cap, producing nothing.
|
|
218
|
+
|
|
219
|
+
Split it into smaller issues, each describing one story, or edit this body down under the limits (at most ${{ env.REFINE_MAX_BODY_CHARS }} characters and ${{ env.REFINE_MAX_WORK_UNITS }} work-unit bullets), then reply here and refinement will run again.
|
|
220
|
+
- name: Flag the issue for a person
|
|
221
|
+
uses: ./.github/actions/add-issue-labels
|
|
222
|
+
with:
|
|
223
|
+
token: ${{ steps.app-token.outputs.token }}
|
|
224
|
+
issue-number: ${{ inputs.issue-number }}
|
|
225
|
+
labels: review
|
|
226
|
+
- name: Clear a stale stalled flag
|
|
227
|
+
uses: ./.github/actions/remove-issue-labels
|
|
228
|
+
with:
|
|
229
|
+
token: ${{ steps.app-token.outputs.token }}
|
|
230
|
+
issue-number: ${{ inputs.issue-number }}
|
|
231
|
+
labels: stalled
|
|
232
|
+
# A refusal must also release the reservation. authorize-bot-work adds bot-working
|
|
233
|
+
# before the router dispatches, and this job is the refused run's only terminal path:
|
|
234
|
+
# incomplete is gated off it (see there), so without this removal a refused issue
|
|
235
|
+
# stays reserved with no run behind it -- parked, invisible, not re-triggerable until
|
|
236
|
+
# the hourly reconcile sweep clears it hours later.
|
|
237
|
+
- name: Release the reservation
|
|
238
|
+
uses: ./.github/actions/remove-issue-labels
|
|
239
|
+
with:
|
|
240
|
+
token: ${{ steps.app-token.outputs.token }}
|
|
241
|
+
issue-number: ${{ inputs.issue-number }}
|
|
242
|
+
labels: ${{ env.WORKING_LABEL }}
|
|
243
|
+
reserve:
|
|
244
|
+
needs: [still_open, size_guard]
|
|
245
|
+
if: needs.still_open.outputs.open == 'true' && needs.size_guard.outputs.too_big != 'true'
|
|
246
|
+
runs-on: agents-arc
|
|
137
247
|
permissions:
|
|
138
248
|
contents: read
|
|
139
249
|
issues: write
|
|
@@ -374,9 +484,17 @@ jobs:
|
|
|
374
484
|
issue-number: ${{ inputs.issue-number }}
|
|
375
485
|
labels: ${{ env.WORKING_LABEL }}
|
|
376
486
|
incomplete:
|
|
377
|
-
needs
|
|
487
|
+
# activation for the artifact prefix the usage read needs; validate_output for its
|
|
488
|
+
# valid output, which decides whether the usage read should look for truncation at all.
|
|
489
|
+
# size_guard because a refusal is terminal: when too_big is true, refuse_big_issue has
|
|
490
|
+
# already spoken on the issue and no attempt, retry or park may follow. Without this
|
|
491
|
+
# clause every refused issue also looped here -- one human-label visit produced eight
|
|
492
|
+
# refusal comments and eight "Attempt N of 5" comments, none of which could succeed
|
|
493
|
+
# (Pliny-Bot #322).
|
|
494
|
+
needs: [activation, agent, safe_outputs, validate_output, size_guard]
|
|
378
495
|
if: >
|
|
379
496
|
always() &&
|
|
497
|
+
needs.size_guard.outputs.too_big != 'true' &&
|
|
380
498
|
(
|
|
381
499
|
needs.agent.result != 'success' ||
|
|
382
500
|
needs.safe_outputs.result != 'success' ||
|
|
@@ -407,6 +525,33 @@ jobs:
|
|
|
407
525
|
under-minutes: ${{ env.RETRY_UNDER_MINUTES }}
|
|
408
526
|
# Recorded before any label moves, so a failure in the steps below leaves a run that can be
|
|
409
527
|
# counted rather than work released with nothing to show for it.
|
|
528
|
+
# The usage read distinguishes the two very different failures that end here. A provider
|
|
529
|
+
# outage consumes almost no tokens; an output-cap truncation (Pliny-Bot #305: 32,000 output
|
|
530
|
+
# tokens, zero outcomes, run green) burns a full run and emits nothing. The comment names
|
|
531
|
+
# the second so a person does not triage it as a flaky provider. Best-effort by design:
|
|
532
|
+
# on an attempt where the agent job itself died there is no artifact, and this step must
|
|
533
|
+
# never fail the incomplete job's label release.
|
|
534
|
+
- name: Read the agent's token usage
|
|
535
|
+
id: usage
|
|
536
|
+
if: needs.agent.result == 'success' && needs.safe_outputs.result == 'success' && needs.validate_output.outputs.valid != 'true'
|
|
537
|
+
env:
|
|
538
|
+
GH_TOKEN: ${{ github.token }}
|
|
539
|
+
REPO: ${{ github.repository }}
|
|
540
|
+
RUN_ID: ${{ github.run_id }}
|
|
541
|
+
ARTIFACT: ${{ needs.activation.outputs.artifact_prefix }}agent
|
|
542
|
+
run: |
|
|
543
|
+
set -euo pipefail
|
|
544
|
+
output_tokens=""
|
|
545
|
+
if gh run download "$RUN_ID" --repo "$REPO" --name "$ARTIFACT" --dir usage-read 2>/dev/null \
|
|
546
|
+
&& [ -f usage-read/agent_usage.json ]; then
|
|
547
|
+
output_tokens=$(jq -r '.output_tokens // empty' usage-read/agent_usage.json 2>/dev/null || echo "")
|
|
548
|
+
fi
|
|
549
|
+
if [ -n "$output_tokens" ] && [ "$output_tokens" -gt 20000 ]; then
|
|
550
|
+
echo "truncated=This attempt consumed ${output_tokens} output tokens and still produced no outcome. That is what provider output truncation looks like, not an outage: the model hit the single-response cap mid-generation and every call it was writing was lost. If this repeats, split the issue or tighten its body." >> "$GITHUB_OUTPUT"
|
|
551
|
+
else
|
|
552
|
+
echo "truncated=" >> "$GITHUB_OUTPUT"
|
|
553
|
+
fi
|
|
554
|
+
echo "output_tokens=$output_tokens" >> "$GITHUB_OUTPUT"
|
|
410
555
|
- name: Report the failed attempt
|
|
411
556
|
if: steps.decide.outputs.retry == 'true'
|
|
412
557
|
uses: ./.github/actions/create-issue-comment
|
|
@@ -417,6 +562,7 @@ jobs:
|
|
|
417
562
|
${{ env.REFINE_ATTEMPT_MARKER }}
|
|
418
563
|
Attempt ${{ steps.decide.outputs.next }} of ${{ env.MAX_ATTEMPTS }} ended after ${{ steps.decide.outputs.minutes }} minutes, before the run could produce an answer.
|
|
419
564
|
${{ env.RETRY_COMMENT }}
|
|
565
|
+
${{ steps.usage.outputs.truncated }}
|
|
420
566
|
[View this workflow run](${{ github.server_url }}/${{ github.repository }}/actions/runs/${{ github.run_id }})
|
|
421
567
|
- name: Release the reservation for the retry
|
|
422
568
|
if: steps.decide.outputs.retry == 'true'
|
|
@@ -466,9 +612,18 @@ jobs:
|
|
|
466
612
|
body: |
|
|
467
613
|
${{ env.REFINE_MARKER }}
|
|
468
614
|
${{ env.INCOMPLETE_COMMENT }}
|
|
615
|
+
${{ steps.usage.outputs.truncated }}
|
|
469
616
|
[View this workflow run](${{ github.server_url }}/${{ github.repository }}/actions/runs/${{ github.run_id }})
|
|
470
617
|
|
|
471
|
-
|
|
618
|
+
# `!failure()` is load-bearing, not decoration. gh-aw adds every custom job to the agent's
|
|
619
|
+
# `needs`, including refuse_big_issue -- a job that skips precisely when refinement should
|
|
620
|
+
# run. Without a status function in this `if:` GitHub prepends an implicit `success()`,
|
|
621
|
+
# and a skipped need fails it, so every under-limit issue skipped the agent in 0s and the
|
|
622
|
+
# incomplete loop parked the issue (Pliny-Bot #320-322, 16 runs). `!failure()` suppresses
|
|
623
|
+
# the implicit success(): a refused skip no longer poisons the agent, while any real
|
|
624
|
+
# failure -- size_guard erroring, the refusal job erroring mid-refuse, reserve failing --
|
|
625
|
+
# still blocks it and lands in incomplete.
|
|
626
|
+
if: inputs.issue-number != '' && needs.still_open.outputs.open == 'true' && needs.size_guard.outputs.too_big != 'true' && !failure()
|
|
472
627
|
|
|
473
628
|
runs-on: agents-arc
|
|
474
629
|
runs-on-slim: agents-arc
|
|
@@ -623,7 +778,16 @@ timeout-minutes: 90
|
|
|
623
778
|
The visible line is for people and the marker is read by the workflow, which turns it into the
|
|
624
779
|
`sp-N` label. A body without the marker gets no estimate label at all.
|
|
625
780
|
|
|
626
|
-
7.
|
|
781
|
+
7. **One safe-output call per turn.** Never batch multiple safe-output calls into a single
|
|
782
|
+
message: `update_issue`, `create_issue` and `add_comment` each go in their own turn, with
|
|
783
|
+
nothing else in the message. The provider caps one response at a fixed size, and a batch of
|
|
784
|
+
large calls is truncated mid-JSON before any of them executes, ending the run green with
|
|
785
|
+
nothing written. On the split path, sequence `create_issue` → `create_issue` → … →
|
|
786
|
+
`update_issue` → `add_comment`, one per turn. If a single body is so large it approaches the
|
|
787
|
+
size of a very long message, tighten the body; a shorter call that lands beats a longer one
|
|
788
|
+
that is cut off.
|
|
789
|
+
|
|
790
|
+
8. Decide exactly one outcome:
|
|
627
791
|
|
|
628
792
|
Labels are workflow-owned state. Do not call `add_labels` or `remove_labels`.
|
|
629
793
|
|