@plainconceptsplatform/workflows 0.27.6 → 0.28.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -2114,6 +2114,49 @@ if worker_installed implement; then
2114
2114
  if [ "$ENDINGS_OK" -eq 1 ]; then PASS=$((PASS + 1)); else FAIL=$((FAIL + 1)); fi
2115
2115
  fi
2116
2116
 
2117
+ echo "── Refine size gate ──────────────────────────────────────────────────────"
2118
+
2119
+ # The size gate added a job that skips in the happy path, and gh-aw adds every custom job to
2120
+ # the agent's needs. Three claims about that shape are asserted, because each one broke in
2121
+ # production in the same week (Pliny-Bot #320-322: sixteen runs, three parked issues, a
2122
+ # refusal loop on an issue the author was shrinking live):
2123
+ #
2124
+ # 1. the agent's gate carries !failure(), so a skipped refusal no longer poisons the
2125
+ # implicit success() and kills the agent in 0s
2126
+ # 2. the incomplete job excludes the refusal, so a refused issue is terminal: no attempt
2127
+ # counter, no x/5, no re-dispatch, no stalled
2128
+ # 3. the refusal itself releases bot-working, so a refused issue is not left reserved
2129
+ if worker_installed refine; then
2130
+ SIZE_GATE_OK=1
2131
+ REFINE_WORKER_MD="${WORKFLOWS_DIR}/agent-refine.md"
2132
+
2133
+ # The top-level if: starts at column 0 -- job-level ifs: are indented -- so anchor on that,
2134
+ # not on the prose comment above it. It must end with !failure().
2135
+ top_level_if="$(sed -n 's/^if: //p' "$REFINE_WORKER_MD" | head -1)"
2136
+ if [[ "$top_level_if" != *'!failure()' ]]; then
2137
+ SIZE_GATE_OK=0
2138
+ echo "FAIL: refine gates the agent without !failure(); a skipped refuse_big_issue poisons the implicit success() and skips the agent on every under-limit issue" >&2
2139
+ fi
2140
+
2141
+ # Extract the incomplete job's own block. A whole-file grep cannot make this claim: reserve,
2142
+ # refuse_big_issue and the agent gate all read size_guard, so the exclusion clause could be
2143
+ # deleted from incomplete while every grep still passes.
2144
+ incomplete_block="$(awk '/^ incomplete:/{found=1; next} found && /^ [a-z_]+:/{exit} found{print}' "$REFINE_WORKER_MD")"
2145
+ if ! printf '%s' "$incomplete_block" | grep -qE '^ needs: \[.*size_guard' \
2146
+ || ! printf '%s' "$incomplete_block" | grep -qF "needs.size_guard.outputs.too_big != 'true'"; then
2147
+ SIZE_GATE_OK=0
2148
+ echo "FAIL: refine's incomplete does not exclude the refusal path; every refused issue also loops attempts" >&2
2149
+ fi
2150
+
2151
+ refuse_block="$(awk '/^ refuse_big_issue:/{found=1; next} found && /^ [a-z_]+:/{exit} found{print}' "$REFINE_WORKER_MD")"
2152
+ if ! printf '%s' "$refuse_block" | grep -qF 'WORKING_LABEL'; then
2153
+ SIZE_GATE_OK=0
2154
+ echo "FAIL: the refine refusal never releases bot-working; a refused issue stays reserved with no run behind it" >&2
2155
+ fi
2156
+
2157
+ if [ "$SIZE_GATE_OK" -eq 1 ]; then PASS=$((PASS + 1)); else FAIL=$((FAIL + 1)); fi
2158
+ fi
2159
+
2117
2160
  echo "── Runner pools ──────────────────────────────────────────────────────────"
2118
2161
 
2119
2162
  # Where every job runs, stated once and asserted, because GitHub gives a wrong pool no error: a
@@ -32,6 +32,14 @@ env:
32
32
  REFINE_ATTEMPT_MARKER: "<!-- agent-refine-attempt -->"
33
33
  MAX_ATTEMPTS: "5"
34
34
  PARK_AT_ATTEMPT: "4"
35
+ # The size gate. An issue over either limit is refused before any model runs: a real
36
+ # one, Pliny-Bot #305 (13,398 chars, 92 work units), burned a full agent -- 17 minutes,
37
+ # 2.4M tokens -- and died at the provider's 32K output cap with zero outcome, twice.
38
+ # Both numbers are measured before the agent starts, so refusal is deterministic and
39
+ # costs ten seconds instead of half an hour.
40
+ REFINE_MAX_BODY_CHARS: "8000"
41
+ REFINE_MAX_WORK_UNITS: "24"
42
+ REFUSE_MARKER: "<!-- agent-refine-refused -->"
35
43
  # Ten rather than six: the failure that prompted this took three minutes, and a slower one on a
36
44
  # worse day would fall outside a six-minute window and park for a fault that clears by itself.
37
45
  RETRY_UNDER_MINUTES: "10"
@@ -98,12 +106,12 @@ on:
98
106
  required: false
99
107
  type: string
100
108
  default: first
101
- # The gate job that the top-level `if:` reads. gh-aw folds that `if:` into the generated
102
- # activation job but gives activation no dependency on the job, so the reference resolves
109
+ # The gate jobs that the top-level `if:` reads. gh-aw folds that `if:` into the generated
110
+ # activation job but gives activation no dependency on the jobs, so the reference resolves
103
111
  # to '' and the clause is false -- the agent would never run. The package's own validator
104
112
  # catches it after compilation; this is the line it asks for, the same one the merge gate
105
113
  # uses for protected_changes.
106
- needs: [still_open]
114
+ needs: [still_open, size_guard]
107
115
 
108
116
  jobs:
109
117
  # A route dispatched while the issue was open must not execute after it has been closed. The
@@ -130,10 +138,112 @@ jobs:
130
138
  with:
131
139
  token: ${{ github.token }}
132
140
  issue-number: ${{ inputs.issue-number }}
133
- reserve:
141
+ # The size gate, rung 4. Measures the issue before any model starts, because an oversized
142
+ # body does not fail fast on its own: it produces an agent that explores for seventeen
143
+ # minutes and then dies mid-generation at the provider's response cap having written
144
+ # nothing, which reads as a green run with no outcome (Pliny-Bot #305, twice). Pure
145
+ # shell, no network beyond one issue read; over either limit it is a refusal, not a
146
+ # smaller attempt.
147
+ size_guard:
134
148
  needs: [still_open]
135
149
  if: needs.still_open.outputs.open == 'true'
136
150
  runs-on: agents-arc
151
+ timeout-minutes: 5
152
+ permissions:
153
+ contents: read
154
+ issues: read
155
+ outputs:
156
+ too_big: ${{ steps.measure.outputs.too_big }}
157
+ reason: ${{ steps.measure.outputs.reason }}
158
+ steps:
159
+ - name: Measure the issue against the size limits
160
+ id: measure
161
+ env:
162
+ GH_TOKEN: ${{ github.token }}
163
+ REPO: ${{ github.repository }}
164
+ ISSUE_NUMBER: ${{ inputs.issue-number }}
165
+ MAX_BODY_CHARS: ${{ env.REFINE_MAX_BODY_CHARS }}
166
+ MAX_WORK_UNITS: ${{ env.REFINE_MAX_WORK_UNITS }}
167
+ run: |
168
+ set -euo pipefail
169
+ body=$(gh issue view "$ISSUE_NUMBER" --repo "$REPO" --json title,body --jq '.title + "\n\n" + (.body // "")')
170
+ chars=$(printf '%s' "$body" | wc -c)
171
+ # One work unit per markdown bullet, the same split the prompt's step 3 makes.
172
+ units=$(printf '%s' "$body" | grep -cE '^[[:space:]]*[-*] ' || true)
173
+ reason=""
174
+ too_big=false
175
+ if [ "$chars" -gt "$MAX_BODY_CHARS" ]; then
176
+ too_big=true
177
+ reason="body ${chars} chars > ${MAX_BODY_CHARS}"
178
+ fi
179
+ if [ "$units" -gt "$MAX_WORK_UNITS" ]; then
180
+ too_big=true
181
+ reason="${reason:+$reason; }work units ${units} > ${MAX_WORK_UNITS}"
182
+ fi
183
+ echo "too_big=$too_big" >> "$GITHUB_OUTPUT"
184
+ echo "reason=$reason" >> "$GITHUB_OUTPUT"
185
+ echo "::notice::measured $chars chars, $units work units (limits: ${MAX_BODY_CHARS} chars, ${MAX_WORK_UNITS} units) -> $too_big"
186
+ # The deterministic refusal. Runs only when the size gate tripped, so it costs one
187
+ # comment and no model. The refine label stays: a human who shrinks the body and
188
+ # replies re-enters rerefine through the comment route, and the classifier ignores
189
+ # bot comments, so this comment cannot re-trigger the worker.
190
+ refuse_big_issue:
191
+ needs: [still_open, size_guard]
192
+ if: needs.still_open.outputs.open == 'true' && needs.size_guard.outputs.too_big == 'true'
193
+ runs-on: agents-arc
194
+ timeout-minutes: 5
195
+ permissions:
196
+ contents: read
197
+ issues: write
198
+ steps:
199
+ - name: Checkout workflow actions
200
+ uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
201
+ with:
202
+ persist-credentials: false
203
+ - name: Create bot token
204
+ id: app-token
205
+ uses: actions/create-github-app-token@bcd2ba49218906704ab6c1aa796996da409d3eb1 # v3.2.0
206
+ with:
207
+ client-id: ${{ secrets.BOT_APP_ID }}
208
+ private-key: ${{ secrets.BOT_PRIVATE_KEY }}
209
+ - name: Refuse the oversized issue
210
+ uses: ./.github/actions/create-issue-comment
211
+ with:
212
+ token: ${{ steps.app-token.outputs.token }}
213
+ issue-number: ${{ inputs.issue-number }}
214
+ body: |
215
+ ${{ env.REFUSE_MARKER }}
216
+ This issue is too large to refine automatically: ${{ needs.size_guard.outputs.reason }}.
217
+ A story this size does not fail fast; it burns a full agent run and dies mid-generation at the provider's response cap, producing nothing.
218
+
219
+ Split it into smaller issues, each describing one story, or edit this body down under the limits (at most ${{ env.REFINE_MAX_BODY_CHARS }} characters and ${{ env.REFINE_MAX_WORK_UNITS }} work-unit bullets), then reply here and refinement will run again.
220
+ - name: Flag the issue for a person
221
+ uses: ./.github/actions/add-issue-labels
222
+ with:
223
+ token: ${{ steps.app-token.outputs.token }}
224
+ issue-number: ${{ inputs.issue-number }}
225
+ labels: review
226
+ - name: Clear a stale stalled flag
227
+ uses: ./.github/actions/remove-issue-labels
228
+ with:
229
+ token: ${{ steps.app-token.outputs.token }}
230
+ issue-number: ${{ inputs.issue-number }}
231
+ labels: stalled
232
+ # A refusal must also release the reservation. authorize-bot-work adds bot-working
233
+ # before the router dispatches, and this job is the refused run's only terminal path:
234
+ # incomplete is gated off it (see there), so without this removal a refused issue
235
+ # stays reserved with no run behind it -- parked, invisible, not re-triggerable until
236
+ # the hourly reconcile sweep clears it hours later.
237
+ - name: Release the reservation
238
+ uses: ./.github/actions/remove-issue-labels
239
+ with:
240
+ token: ${{ steps.app-token.outputs.token }}
241
+ issue-number: ${{ inputs.issue-number }}
242
+ labels: ${{ env.WORKING_LABEL }}
243
+ reserve:
244
+ needs: [still_open, size_guard]
245
+ if: needs.still_open.outputs.open == 'true' && needs.size_guard.outputs.too_big != 'true'
246
+ runs-on: agents-arc
137
247
  permissions:
138
248
  contents: read
139
249
  issues: write
@@ -374,9 +484,17 @@ jobs:
374
484
  issue-number: ${{ inputs.issue-number }}
375
485
  labels: ${{ env.WORKING_LABEL }}
376
486
  incomplete:
377
- needs: [agent, safe_outputs, validate_output]
487
+ # activation for the artifact prefix the usage read needs; validate_output for its
488
+ # valid output, which decides whether the usage read should look for truncation at all.
489
+ # size_guard because a refusal is terminal: when too_big is true, refuse_big_issue has
490
+ # already spoken on the issue and no attempt, retry or park may follow. Without this
491
+ # clause every refused issue also looped here -- one human-label visit produced eight
492
+ # refusal comments and eight "Attempt N of 5" comments, none of which could succeed
493
+ # (Pliny-Bot #322).
494
+ needs: [activation, agent, safe_outputs, validate_output, size_guard]
378
495
  if: >
379
496
  always() &&
497
+ needs.size_guard.outputs.too_big != 'true' &&
380
498
  (
381
499
  needs.agent.result != 'success' ||
382
500
  needs.safe_outputs.result != 'success' ||
@@ -407,6 +525,33 @@ jobs:
407
525
  under-minutes: ${{ env.RETRY_UNDER_MINUTES }}
408
526
  # Recorded before any label moves, so a failure in the steps below leaves a run that can be
409
527
  # counted rather than work released with nothing to show for it.
528
+ # The usage read distinguishes the two very different failures that end here. A provider
529
+ # outage consumes almost no tokens; an output-cap truncation (Pliny-Bot #305: 32,000 output
530
+ # tokens, zero outcomes, run green) burns a full run and emits nothing. The comment names
531
+ # the second so a person does not triage it as a flaky provider. Best-effort by design:
532
+ # on an attempt where the agent job itself died there is no artifact, and this step must
533
+ # never fail the incomplete job's label release.
534
+ - name: Read the agent's token usage
535
+ id: usage
536
+ if: needs.agent.result == 'success' && needs.safe_outputs.result == 'success' && needs.validate_output.outputs.valid != 'true'
537
+ env:
538
+ GH_TOKEN: ${{ github.token }}
539
+ REPO: ${{ github.repository }}
540
+ RUN_ID: ${{ github.run_id }}
541
+ ARTIFACT: ${{ needs.activation.outputs.artifact_prefix }}agent
542
+ run: |
543
+ set -euo pipefail
544
+ output_tokens=""
545
+ if gh run download "$RUN_ID" --repo "$REPO" --name "$ARTIFACT" --dir usage-read 2>/dev/null \
546
+ && [ -f usage-read/agent_usage.json ]; then
547
+ output_tokens=$(jq -r '.output_tokens // empty' usage-read/agent_usage.json 2>/dev/null || echo "")
548
+ fi
549
+ if [ -n "$output_tokens" ] && [ "$output_tokens" -gt 20000 ]; then
550
+ echo "truncated=This attempt consumed ${output_tokens} output tokens and still produced no outcome. That is what provider output truncation looks like, not an outage: the model hit the single-response cap mid-generation and every call it was writing was lost. If this repeats, split the issue or tighten its body." >> "$GITHUB_OUTPUT"
551
+ else
552
+ echo "truncated=" >> "$GITHUB_OUTPUT"
553
+ fi
554
+ echo "output_tokens=$output_tokens" >> "$GITHUB_OUTPUT"
410
555
  - name: Report the failed attempt
411
556
  if: steps.decide.outputs.retry == 'true'
412
557
  uses: ./.github/actions/create-issue-comment
@@ -417,6 +562,7 @@ jobs:
417
562
  ${{ env.REFINE_ATTEMPT_MARKER }}
418
563
  Attempt ${{ steps.decide.outputs.next }} of ${{ env.MAX_ATTEMPTS }} ended after ${{ steps.decide.outputs.minutes }} minutes, before the run could produce an answer.
419
564
  ${{ env.RETRY_COMMENT }}
565
+ ${{ steps.usage.outputs.truncated }}
420
566
  [View this workflow run](${{ github.server_url }}/${{ github.repository }}/actions/runs/${{ github.run_id }})
421
567
  - name: Release the reservation for the retry
422
568
  if: steps.decide.outputs.retry == 'true'
@@ -466,9 +612,18 @@ jobs:
466
612
  body: |
467
613
  ${{ env.REFINE_MARKER }}
468
614
  ${{ env.INCOMPLETE_COMMENT }}
615
+ ${{ steps.usage.outputs.truncated }}
469
616
  [View this workflow run](${{ github.server_url }}/${{ github.repository }}/actions/runs/${{ github.run_id }})
470
617
 
471
- if: inputs.issue-number != '' && needs.still_open.outputs.open == 'true'
618
+ # `!failure()` is load-bearing, not decoration. gh-aw adds every custom job to the agent's
619
+ # `needs`, including refuse_big_issue -- a job that skips precisely when refinement should
620
+ # run. Without a status function in this `if:` GitHub prepends an implicit `success()`,
621
+ # and a skipped need fails it, so every under-limit issue skipped the agent in 0s and the
622
+ # incomplete loop parked the issue (Pliny-Bot #320-322, 16 runs). `!failure()` suppresses
623
+ # the implicit success(): a refused skip no longer poisons the agent, while any real
624
+ # failure -- size_guard erroring, the refusal job erroring mid-refuse, reserve failing --
625
+ # still blocks it and lands in incomplete.
626
+ if: inputs.issue-number != '' && needs.still_open.outputs.open == 'true' && needs.size_guard.outputs.too_big != 'true' && !failure()
472
627
 
473
628
  runs-on: agents-arc
474
629
  runs-on-slim: agents-arc
@@ -623,7 +778,16 @@ timeout-minutes: 90
623
778
  The visible line is for people and the marker is read by the workflow, which turns it into the
624
779
  `sp-N` label. A body without the marker gets no estimate label at all.
625
780
 
626
- 7. Decide exactly one outcome:
781
+ 7. **One safe-output call per turn.** Never batch multiple safe-output calls into a single
782
+ message: `update_issue`, `create_issue` and `add_comment` each go in their own turn, with
783
+ nothing else in the message. The provider caps one response at a fixed size, and a batch of
784
+ large calls is truncated mid-JSON before any of them executes, ending the run green with
785
+ nothing written. On the split path, sequence `create_issue` → `create_issue` → … →
786
+ `update_issue` → `add_comment`, one per turn. If a single body is so large it approaches the
787
+ size of a very long message, tighten the body; a shorter call that lands beats a longer one
788
+ that is cut off.
789
+
790
+ 8. Decide exactly one outcome:
627
791
 
628
792
  Labels are workflow-owned state. Do not call `add_labels` or `remove_labels`.
629
793
 
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@plainconceptsplatform/workflows",
3
- "version": "0.27.6",
3
+ "version": "0.28.1",
4
4
  "description": "Install and update Platform GitHub agentic workflows.",
5
5
  "keywords": [
6
6
  "github-actions",