pi-antiloop 1.6.1 → 1.7.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -6,13 +6,13 @@
6
6
 
7
7
  # Antiloop — Loop Detection and Break for pi
8
8
 
9
- **Antiloop watches every assistant message, tool call and thinking block, and forces the model out of reasoning loops before they eat your context and your patience.** Six simultaneous detection strategies (text similarity, tool-call sequences, thinking content, structural openings, degenerate repetition, no-progress outcome runs) find loops that humans miss — and progressive intervention (warning → force break → abort) tells the model to take a different approach, without you having to babysit it.
9
+ **Antiloop watches every assistant message, tool call and thinking block, and forces the model out of reasoning loops before they eat your context and your patience.** Seven simultaneous detection strategies (text similarity, tool-call sequences, thinking content, structural openings, degenerate repetition, block/narration repetition, no-progress outcome runs) find loops that humans miss — and progressive intervention (warning → force break → abort) tells the model to take a different approach, without you having to babysit it.
10
10
 
11
11
  ---
12
12
 
13
13
  ## Features
14
14
 
15
- - **Six detection strategies** — text repetition (trigram Jaccard + Levenshtein), tool-call sequences (name + near-identical arguments + same outcome — result-aware, so retries that make progress don't false-positive), thinking blocks, structural opening-phrase patterns, **degenerate repetition** (a single message/call stuck repeating one word hundreds of times — the `noguerol ×5145` meltdown — caught at `message_end` with no repeated peer needed, and degenerate bash commands blocked before they execute), and **no-progress outcome runs** (the NFS `test A…QQQ` class: dozens of near-identical re-runs of the same experiment, every one ending in the *same failing outcome* — args mutate so tool-loop can't see it; the repeated failure signature can)
15
+ - **Seven detection strategies** — text repetition (trigram Jaccard + Levenshtein), tool-call sequences (name + near-identical arguments + same outcome — result-aware, so retries that make progress don't false-positive), thinking blocks, structural opening-phrase patterns, **degenerate repetition** (a single message/call stuck repeating one word hundreds of times — the `noguerol ×5145` meltdown — caught at `message_end` with no repeated peer needed, and degenerate bash commands blocked before they execute), **block repetition** (ONE message replaying whole sentences/phrases — the narration loop: `Let me start by checking the environment…` ×5 — same `message_end` timing and strong turn weight), and **no-progress outcome runs** (the NFS `test A…QQQ` class: dozens of near-identical re-runs of the same experiment, every one ending in the *same failing outcome* — args mutate so tool-loop can't see it; the repeated failure signature can)
16
16
  - **Stays alive in long sessions** — tracked messages carry a monotonic sequence number, so trimming the sliding window can never make the dedupe guard skip later messages (a bug that silently blinded antiloop after ~15 tracked messages)
17
17
  - **Task-stream recognition (batch work)** — when another extension (e.g. `punched` appending lines to pi.md, or `plan` adding tasks) makes the model call the *same* tool many times with *different* content, antiloop recognizes it as N distinct tasks of one type and stays silent — no warning, no force break. A genuine loop (the *same* call repeated verbatim) is still caught
18
18
  - **Progressive intervention** — `warning` reminds the model to vary its approach; `force break` steers a real break message into the running agent before its next LLM call; `abort` stops the run entirely
@@ -100,12 +100,14 @@ Detection strategies:
100
100
  Tool loops: ✅
101
101
  Thinking loops: ✅
102
102
  Degenerate: ✅
103
+ Block repetition: ✅
103
104
  Outcome (no-progress): ✅
104
105
 
105
106
  Recent detections:
106
107
  [text] Text similarity 85% with message 3 (2m ago)
107
108
  [tool] repeated 3x: bash (5m ago)
108
109
  [degenerate] bash command: degenerate repetition — "noguerol" ×5145 (5m ago)
110
+ [block] message text: block repetition — 100% of 450 tokens replay repeated 5-word phrases (5m ago)
109
111
  [outcome] no progress: 9 near-identical bash attempts with the same failing outcome (5m ago)
110
112
  ```
111
113
 
@@ -129,7 +131,9 @@ Grouped interactive menu showing the current value in each option:
129
131
  - **🔁 call repeats** — `1 / 2 / 3` — how many times the same call must repeat before it flags (default 2)
130
132
  - **🧾 result similarity** — `95 / 80 / 60%` — how similar captured results must be to count as the *same outcome*; a repeated command that starts producing a different result is progress, not a loop (default 80%)
131
133
  - **🌀 degenerate run** — `8 / 16 / 24 / 48` — identical words in a row inside ONE message/call before it counts as stuck generation (default 16; the `noguerol ×5145` class)
132
- - **⛔ block degenerate bash** — on/off — refuse a degenerate bash command before it executes, feeding the reason back to the model (default on)
134
+ - **🔁 block repetition** — on/off — detect ONE message replaying whole sentences/phrases (the narration loop: `Let me start by checking the environment…` ×5) with no repeated peer needed
135
+ - **🔁 block share** — `70 / 85 / 95%` — share of replayed 5-word phrases in one message before it counts as a replay (default 85%)
136
+ - **⛔ block repetitive bash** — on/off — refuse a degenerate or replayed bash command before it executes, feeding the reason back to the model (default on)
133
137
  - **📉 no-progress after** — `4 / 6 / 8 / 12` — same failing outcome repeated this many times (near-identical args) before flagging (default 8; the NFS `test A…QQQ` class)
134
138
 
135
139
  **📋 Task streams** — batch work (punched_log, plan_manager, …) is N tasks of one type, not a loop
@@ -142,6 +146,7 @@ Grouped interactive menu showing the current value in each option:
142
146
  - **🔧 tools** — on/off — detect repeated tool calls
143
147
  - **🧠 thinking** — on/off — detect repeated internal reasoning
144
148
  - **🌀 degenerate** — on/off — detect one message stuck repeating a single word/token (no repeated peer needed)
149
+ - **🔁 block** — on/off — detect one message replaying whole sentences/phrases (narration loop, no repeated peer needed)
145
150
  - **📉 outcome** — on/off — detect many near-identical attempts all ending in the same failing outcome (no progress)
146
151
 
147
152
  **🧹 reset state** — clear all counters and history
@@ -175,6 +180,9 @@ degenerate legit cmd → ok (exp ok) ✅
175
180
  degenerate interleaved → flag (noguerol ×150) (exp flag — freq/share clause) ✅
176
181
  degenerate glued token → flag (noguerol ×300) (exp flag — perfect power) ✅
177
182
  degenerate turn weight → degenerate 2, text 1 (exp 2, 1) ✅
183
+ block S0 narration loop → block (message text: 100% of 450 tokens replay repeated 5-word phrases) ✅ ← v1.7: ONE message replaying sentences, no peer
184
+ block 3 cycles / 2 cycles→ flag (100%) / no (exp flag / no — needs ≥3 repeats) ✅
185
+ block legit prose/logs → ok (exp ok) ✅
178
186
  outcome fires on 9th → outcome (9 near-identical bash attempts, same failing outcome) ✅ ← v1.6.1: NFS test-A…QQQ class
179
187
  outcome needs 8 prior → silent (exp silent at 5 attempts) ✅
180
188
  outcome converging sweep → silent (exp silent — outcomes differ = progress) ✅
@@ -186,7 +194,7 @@ outcome identical OKs → silent (exp silent — success repeats ≠ loop)
186
194
 
187
195
  ### Detection pipeline
188
196
 
189
- After every assistant `message_end` event, antiloop extracts the new content (text, thinking, tool calls — including their ids) and pushes it onto a sliding window of the last `detectionWindow + 5` messages. Detection itself runs at `turn_end`, once the tool results are known: results are fingerprinted and attached to the tracked calls, then the active detection strategies run against the window. The one exception is the **degenerate** strategy: a single stuck message needs no peer and no tool result, so it is evaluated right at `message_end` — the only point before the message's own tool calls execute — and degenerate `bash` calls are also blocked at the `tool_call` hook.
197
+ After every assistant `message_end` event, antiloop extracts the new content (text, thinking, tool calls — including their ids) and pushes it onto a sliding window of the last `detectionWindow + 5` messages. Detection itself runs at `turn_end`, once the tool results are known: results are fingerprinted and attached to the tracked calls, then the active detection strategies run against the window. The exceptions are the **intra-message** strategies — **degenerate** and **block**. A single stuck message needs no peer and no tool result, so both are evaluated right at `message_end` — the only point before the message's own tool calls execute — and repetitive `bash` calls are also blocked at the `tool_call` hook.
190
198
 
191
199
  | Strategy | What it compares | Algorithm |
192
200
  |----------|------------------|-----------|
@@ -195,10 +203,11 @@ After every assistant `message_end` event, antiloop extracts the new content (te
195
203
  | Thinking | Internal reasoning/thinking blocks | Same as text |
196
204
  | Structural | First 10 words of each message | Opening-phrase similarity ≥ 90% across ≥ 3 messages |
197
205
  | Degenerate | One single message/call (no peer needed) | Run-length + frequency of identical words inside the payload: ≥ `degenerateMaxRun` (default 16) consecutive identical words, or one word ≥ `degenerateMaxFreq`× at ≥ `degenerateMaxShare` of all tokens; plus a perfect-power check for glued no-space tokens. Scanned at `message_end` — before the tool calls execute — and on every `bash` `tool_call` (blocking gate) |
206
+ | Block | One single message (no peer needed) | Sliding `blockNgram`-word n-grams over the normalized payload: a payload of ≥ `blockMinTokens` (120) words is a replay when ≥ `blockRepeatShare` (default 85%) of the n-gram positions recur AND the most repeated n-gram appears ≥ `blockMinRepeats` (3) times. Catches a message replaying whole sentences/phrases — the narration loop (`Let me start by checking the environment…` ×5) that every cross-message detector misses because there is no peer. Scanned at `message_end` and on every `bash` `tool_call` (blocking gate) |
198
207
  | Outcome | Single tool calls across the window, after the last user input | ≥ `outcomeMinRepeats` (default 8) PRIOR attempts with args ≥ `outcomeArgSimilarity` (0.85) similar AND the same *failing* outcome (failure signatures compared at ≥ `outcomeSigThreshold`, 0.7; identical OK results never count — they're the norm for batches). Catches mutated re-run loops the tool detector can't see (labels/permutations change every turn) |
199
208
  | Task stream | Same tool, many calls | When a tool appears ≥ `taskStreamMinCalls` times (default 3) in the window and *no two* calls are near-identical (`taskStreamTwinThreshold`, default 99%), the tool is an active batch: N different tasks of one type (e.g. `punched_log` appends, `plan_manager` task adds). Those calls are exempt from tool-loop detection, and text/thinking/structural patterns that only involve those batch messages are suppressed too. If even one call pair is a twin (the same task repeated), the tool is *not* a stream and detection proceeds normally |
200
209
 
201
- Each detected pair becomes a `LoopDetection { type, similarity, messageIndices, description }` and the consecutive counter increases (a degenerate turn counts `degenerateTurnWeight`, default 2 — warning on first sight).
210
+ Each detected pair becomes a `LoopDetection { type, similarity, messageIndices, description }` and the consecutive counter increases (a degenerate or block turn counts `degenerateTurnWeight`, default 2 — warning on first sight).
202
211
 
203
212
  ### Intervention levels
204
213
 
@@ -211,7 +220,7 @@ Each detected pair becomes a `LoopDetection { type, similarity, messageIndices,
211
220
 
212
221
  The level never de-escalates during an active loop; user input decays the consecutive counter naturally so a fresh prompt can break the cycle.
213
222
 
214
- **Degenerate turns are handled earlier than the ladder:** the meltdown is detected at `message_end` (the same moment the text/tool-call payload is complete, *before* pi preflights and executes its tools). A single degenerate turn already adds `degenerateTurnWeight` (2) consecutive points → warning on first sight; the second consecutive meltdown → force break steer; after the steer, further degenerate output counts against `ignoredSteerLimit` → hard stop. Degenerate `bash` calls are additionally refused by the `tool_call` gate (`blockDegenerateBash`) — the command never runs, and the block reason is fed back to the model as the tool error so it can still change approach.
223
+ **Degenerate and block turns are handled earlier than the ladder:** they are detected at `message_end` (the same moment the text/tool-call payload is complete, *before* pi preflights and executes its tools). A single such turn already adds `degenerateTurnWeight` (2) consecutive points → warning on first sight; the second consecutive one → force break steer; after the steer, further degenerate/blocked output counts against `ignoredSteerLimit` → hard stop. Repetitive `bash` calls are additionally refused by the `tool_call` gate (`blockDegenerateBash`) — the command never runs, and the block reason is fed back to the model as the tool error so it can still change approach.
215
224
 
216
225
  ### Similarity scoring
217
226
 
@@ -252,6 +261,7 @@ If the same command produced a *different* outcome, the pair is progress:
252
261
 
253
262
  Results only veto; they never trigger on their own, and calls without a
254
263
  captured result fall back to argument matching alone.
264
+ ```
255
265
 
256
266
  ### Task streams: N tasks of one type ≠ a loop
257
267
 
@@ -333,6 +343,49 @@ sweeps that motivated the tool-loop threshold). Only `bash` gets the blocking
333
343
  gate — writing a repetitive *file* (e.g. a user-requested padding fixture) is
334
344
  alerted and escalated, not refused.
335
345
 
346
+ ### Block repetition: ONE message replaying its own sentences
347
+
348
+ The degenerate detector catches a single *word* repeated hundreds of times.
349
+ There is a second intra-message failure mode, just as obvious to a human and
350
+ invisible to every cross-message detector: **the model replays whole
351
+ sentences/phrases inside one generation**. Real case (a coding session's S0
352
+ scaffold): ONE assistant message cycled ~5 times through
353
+
354
+ > Let me start by checking the environment and the current state of the
355
+ > repository, then set up a plan for the S0 slice and begin building.
356
+ > I'll run several independent checks in parallel.
357
+ > Let me begin the S0 development. First, reconnaissance of the environment
358
+ > and current repo state.
359
+
360
+ …never emitting a tool call. Text/tool/thinking/structural detection all
361
+ compare *across* messages and had nothing to compare against; the degenerate
362
+ scan only watches for one word repeated, and no single word repeated 16×.
363
+
364
+ Antiloop v1.7 detects **block repetition** on the message itself:
365
+
366
+ - the payload is read as a stream of lowercase alphanumeric words;
367
+ - a sliding window of `blockNgram` (default 5) words is built over it;
368
+ - if at least `blockRepeatShare` (default **85%**) of those n-gram positions
369
+ recur somewhere else, the most repeated n-gram appears at least
370
+ `blockMinRepeats` (default 3) times, and the payload has at least
371
+ `blockMinTokens` (default 120) words, the generation is replaying itself.
372
+
373
+ The thresholds are deliberately far apart from normal content: ordinary
374
+ prose — even long, structured documents — scores ≤ 11% coverage, while 3+
375
+ replays of a narration block reach 98–100%. Templated payloads with evolving
376
+ data (log lines, generated code) are not flagged because their n-grams carry
377
+ the varying digits and differ. A single restatement (2 cycles) does not trip
378
+ it — one echo is a summary, three are stuck generation.
379
+
380
+ Because the signal is conclusive it is handled like the degenerate one:
381
+ caught at `message_end` (before the message's tools execute), weighted
382
+ `degenerateTurnWeight` (2) so the first occurrence warns, and a replayed
383
+ `bash` command is refused by the same `blockDegenerateBash` gate. After a
384
+ force-break steer, another replayed block escalates to the hard stop.
385
+
386
+ Tunables: `detectBlockRepeats`, `blockRepeatShare` (lower = earlier),
387
+ `blockMinRepeats`, `blockNgram`, `blockMinTokens`.
388
+
336
389
  ### No-progress outcome runs: mutated re-runs, same wall
337
390
 
338
391
  A subtler meltdown than the degenerate one: the model re-runs the SAME
@@ -372,7 +425,6 @@ real session this fires at `test MM` (warn) → `NN` (steer) → `PP` (abort)
372
425
 
373
426
  Tunables: `detectOutcomeLoops`, `outcomeMinRepeats` (lower = earlier cutoff),
374
427
  `outcomeArgSimilarity`, `outcomeSigThreshold`.
375
- ```
376
428
 
377
429
  ### Sliding window
378
430
 
@@ -399,6 +451,11 @@ Persisted as JSON at `~/.pi/agent/antiloop.json`:
399
451
  "degenerateMaxShare": 0.4,
400
452
  "degenerateTurnWeight": 2,
401
453
  "blockDegenerateBash": true,
454
+ "detectBlockRepeats": true,
455
+ "blockMinTokens": 120,
456
+ "blockNgram": 5,
457
+ "blockMinRepeats": 3,
458
+ "blockRepeatShare": 0.85,
402
459
  "detectOutcomeLoops": true,
403
460
  "outcomeMinRepeats": 8,
404
461
  "outcomeArgSimilarity": 0.85,
@@ -433,7 +490,12 @@ Persisted as JSON at `~/.pi/agent/antiloop.json`:
433
490
  | `degenerateMaxFreq` | `60` | One word's total occurrences (with `degenerateMaxShare` of the payload) that flags interleaved meltdowns |
434
491
  | `degenerateMaxShare` | `0.4` | Frequency share (freq/total tokens) required together with `degenerateMaxFreq` |
435
492
  | `degenerateTurnWeight` | `2` | Consecutive-detection points added by one degenerate turn (2 = warning on first sight) |
436
- | `blockDegenerateBash` | `true` | Block a degenerate `bash` command in the `tool_call` hook before it executes; the reason is fed back to the model as the tool error |
493
+ | `blockDegenerateBash` | `true` | Block a degenerate or block-repetition `bash` command in the `tool_call` hook before it executes; the reason is fed back to the model as the tool error |
494
+ | `detectBlockRepeats` | `true` | Detect intra-message block/narration repetition — ONE message replaying whole sentences/phrases (the `Let me start by checking the environment…` ×5 class). Needs no repeated peer; scanned at `message_end` and on every bash `tool_call` |
495
+ | `blockMinTokens` | `120` | Minimum normalized words in a payload before it is scanned for block repetition (shorter payloads aren't conclusive) |
496
+ | `blockNgram` | `5` | Word window of the n-grams whose recurrence is measured |
497
+ | `blockMinRepeats` | `3` | The most repeated n-gram must occur at least this many times to flag |
498
+ | `blockRepeatShare` | `0.85` | Share of n-gram positions that must recur for the payload to count as a replay |
437
499
  | `detectOutcomeLoops` | `true` | No-progress outcome runs: ≥ `outcomeMinRepeats` near-identical attempts (args ≥ `outcomeArgSimilarity`) all ending in the *same failing outcome* — the NFS `test A…QQQ` class |
438
500
  | `outcomeMinRepeats` | `8` | Prior same-failure attempts (inside the window, after the last user input) required before the outcome detector fires |
439
501
  | `outcomeArgSimilarity` | `0.85` | How similar args must be to count as the *same experiment reshuffled* (mutations of labels/permutations stay under it — distinct tasks don't) |
@@ -460,7 +522,8 @@ Persisted as JSON at `~/.pi/agent/antiloop.json`:
460
522
  6. **Let user input clear state** — each user message decays the consecutive counter by 2, so a fresh prompt naturally resets without `/antiloop reset`.
461
523
  7. **Degenerate detector needs no tuning for most setups** — a run of ≥ 16 identical words (or one word ≥ 40% of a ≥ 50-token payload) inside a single message is conclusive stuck generation; the 46 KB `noguerol ×5145` SSH-wordlist meltdown is caught on first sight (warning), its bash never executes (`blockDegenerateBash`), and a second consecutive meltdown gets the force-break steer. If a model legitimately writes repetitive payloads, raise `degenerateMaxRun` / `degenerateMaxFreq` via `/antiloop config` — don't disable the detector.
462
524
  8. **Outcome detector catches mutated re-run loops** — a model that re-issues the same experiment with cosmetic changes (labels, permutations) while every attempt fails identically gets a warning after `outcomeMinRepeats` (8) same-failure attempts, a steer on the next, and a hard stop shortly after. Converging sweeps, evolving failures and repeated successes stay silent by design. Lower `outcomeMinRepeats` if you want earlier cutoffs.
463
- 9. **`/antiloop test`** — runs the real detection engine (text + tool-call + task-stream + degenerate + outcome regression cases) to verify calibration after any change.
525
+ 9. **Block detector catches narration loops** — a single message replaying whole sentences (the `Let me start by checking the environment…` ×5 class) warns on first sight and escalates like a meltdown. It only fires when ≥ 85% of the message's 5-word phrases recur, so ordinary prose and templated logs/code stay silent. Raise `blockRepeatShare` (or `blockMinRepeats`) if you ever see a false positive.
526
+ 10. **`/antiloop test`** — runs the real detection engine (text + tool-call + task-stream + degenerate + block + outcome regression cases) to verify calibration after any change.
464
527
 
465
528
  ## Architecture
466
529
 
@@ -485,10 +548,10 @@ Modular extension with zero external dependencies (only pi's bundled `@earendil-
485
548
 
486
549
  - **Levenshtein + trigram Jaccard** hybrid — small texts use edit distance, large texts use n-gram overlap (each is O(N) in text length)
487
550
  - **Sliding window** — only the last `detectionWindow` messages participate, capping memory at O(W × message_size)
488
- - **Early bail** — short messages and empty tool calls skip similarity computation entirely; the degenerate scan is a single linear tokenization pass
551
+ - **Early bail** — short messages and empty tool calls skip similarity computation entirely; the degenerate and block scans are single linear passes
489
552
  - **TUI integration** — uses `ctx.ui.select` for the config menu and the log viewer; `ctx.ui.notify` for state notifications; `ctx.ui.setStatus` + a custom `ctx.ui.setFooter` component for the persistent footer indicator, live level info, and the `esc+a` keyboard toggle (`ctx.ui.onTerminalInput`, never consumes input)
490
- - **Hooks** — `message_end` (track messages + tool call ids with a monotonic sequence so the sliding-window trim can never collide turn indices, and pre-handle degenerate meltdowns — the message is complete but its tools haven't executed yet), `tool_call` (block degenerate `bash` commands before they run), `turn_end` (attach result fingerprints with failure signatures, detect — including no-progress outcome runs — and intervene: steer the force break / abort the run), `input` (decay on real user messages only), `session_start` (load config + install footer + reset), `session_shutdown` (restore built-in footer)
491
- - **Intervention runs on the turn loop, not on user prompts** — escalation is decided at `turn_end` (and at `message_end` for the self-contained degenerate signal), the break is steered into the running agent before its next LLM call, and the guaranteed hard stop aborts the run (`ctx.abort`, fire-and-forget — never awaited, so the hook can't deadlock). No custom-role messages are injected into the conversation at any level (steering a real user message + aborting are the only levers; custom-role injections were removed because a model can stall on an unexpected injected message)
553
+ - **Hooks** — `message_end` (track messages + tool call ids with a monotonic sequence so the sliding-window trim can never collide turn indices, and pre-handle degenerate/block repetition — the message is complete but its tools haven't executed yet), `tool_call` (block degenerate or replayed `bash` commands before they run), `turn_end` (attach result fingerprints with failure signatures, detect — including no-progress outcome runs — and intervene: steer the force break / abort the run), `input` (decay on real user messages only), `session_start` (load config + install footer + reset), `session_shutdown` (restore built-in footer)
554
+ - **Intervention runs on the turn loop, not on user prompts** — escalation is decided at `turn_end` (and at `message_end` for the self-contained intra-message signals), the break is steered into the running agent before its next LLM call, and the guaranteed hard stop aborts the run (`ctx.abort`, fire-and-forget — never awaited, so the hook can't deadlock). No custom-role messages are injected into the conversation at any level (steering a real user message + aborting are the only levers; custom-role injections were removed because a model can stall on an unexpected injected message)
492
555
 
493
556
  ## License
494
557
 
package/package.json CHANGED
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "name": "pi-antiloop",
3
- "version": "1.6.1",
4
- "description": "Antiloop: detect reasoning loops and force a break (warn → force → abort) across text, tool, thinking, structural, degenerate-repetition, and no-progress outcome patterns. Degenerate (one message stuck repeating a word hundreds of times — noguerol ×5145) is caught at message_end and its bash blocked before executing. No-progress (the NFS test-A…QQQ class: ~90 mutated re-runs of the same experiment, every one failing identically) fires when the same failing outcome repeats ≥ outcomeMinRepeats times with near-identical args. Fixes antiloop going blind mid-session: tracked messages use a monotonic sequence so the trim of the sliding window can never collide turn indices. The force break is delivered mid-run (steer before the next LLM call) and if the model ignores it antiloop aborts the run — an autonomous tool loop always terminates. Result-aware tool-loop detection keeps sequential bash and converging sweeps quiet; task-stream recognition keeps punched/plan batch work quiet.",
3
+ "version": "1.7.0",
4
+ "description": "Antiloop: detect reasoning loops and force a break (warn → force → abort) across text, tool, thinking, structural, degenerate-repetition, block/narration-repetition, and no-progress outcome patterns. Degenerate (one message stuck repeating a word hundreds of times — noguerol ×5145) is caught at message_end and its bash blocked before executing. Block (ONE message replaying whole sentences/phrases — the 'Let me start by checking the environment…' ×5 narration loop, invisible to every cross-message detector because there is no peer) fires on the first such message with the same strong turn weight and blocks a replayed bash command. No-progress (the NFS test-A…QQQ class: ~90 mutated re-runs of the same experiment, every one failing identically) fires when the same failing outcome repeats ≥ outcomeMinRepeats times with near-identical args. Fixes antiloop going blind mid-session: tracked messages use a monotonic sequence so the trim of the sliding window can never collide turn indices. The force break is delivered mid-run (steer before the next LLM call) and if the model ignores it antiloop aborts the run — an autonomous tool loop always terminates. Result-aware tool-loop detection keeps sequential bash and converging sweeps quiet; task-stream recognition keeps punched/plan batch work quiet.",
5
5
  "keywords": [
6
6
  "pi-package",
7
7
  "antiloop",
package/src/commands.ts CHANGED
@@ -59,10 +59,11 @@ async function showStatus(ctx: ExtensionCommandContext, rt: Runtime): Promise<vo
59
59
  ` tool sim: ${(rt.config.toolSimilarityThreshold * 100).toFixed(0)}% tool repeat: ${rt.config.minToolRepeatCount}+ prior`,
60
60
  ` result sim: ${(rt.config.resultSimilarityThreshold * 100).toFixed(0)}% (same cmd + diff outcome = no loop)`,
61
61
  ` degenerate: run ≥ ${rt.config.degenerateMaxRun} same word · freq ≥ ${rt.config.degenerateMaxFreq} @ ${(rt.config.degenerateMaxShare * 100).toFixed(0)}% (≥ ${rt.config.degenerateMinTokens} tokens) · weight ${rt.config.degenerateTurnWeight} · block bash ${yn(rt.config.blockDegenerateBash)}`,
62
+ ` block repeats: ${yn(rt.config.detectBlockRepeats)} (≥ ${(rt.config.blockRepeatShare * 100).toFixed(0)}% of ≥ ${rt.config.blockMinTokens} tokens replay ${rt.config.blockNgram}-grams ×${rt.config.blockMinRepeats}+)`,
62
63
  ` outcome: same failing result ≥ ${rt.config.outcomeMinRepeats} attempts (args ≥ ${(rt.config.outcomeArgSimilarity * 100).toFixed(0)}% sim, sig ≥ ${(rt.config.outcomeSigThreshold * 100).toFixed(0)}%)`,
63
64
  ` task streams: ${yn(rt.config.detectTaskStreams)} (min ${rt.config.taskStreamMinCalls} calls, twins ≥ ${(rt.config.taskStreamTwinThreshold * 100).toFixed(0)}%)`,
64
65
  "",
65
- `detectors: text ${yn(rt.config.detectTextLoops)} · tool ${yn(rt.config.detectToolLoops)} · think ${yn(rt.config.detectThinkingLoops)} · degenerate ${yn(rt.config.detectDegenerate)} · outcome ${yn(rt.config.detectOutcomeLoops)}`,
66
+ `detectors: text ${yn(rt.config.detectTextLoops)} · tool ${yn(rt.config.detectToolLoops)} · think ${yn(rt.config.detectThinkingLoops)} · degenerate ${yn(rt.config.detectDegenerate)} · block ${yn(rt.config.detectBlockRepeats)} · outcome ${yn(rt.config.detectOutcomeLoops)}`,
66
67
  `footer: interactive ${yn(rt.config.interactiveFooter)} · toggle: ${rt.config.toggleShortcut}`,
67
68
  ];
68
69
  if (rt.state.activeTaskStreams.length) {
@@ -95,7 +96,9 @@ async function showConfigMenu(ctx: ExtensionCommandContext, rt: Runtime): Promis
95
96
  { value: "resultSim" as const, label: `🧾 result similarity: ${(c.resultSimilarityThreshold * 100).toFixed(0)}%`, description: "same command + different result = progress, not a loop" },
96
97
  // ── 🌀 Degenerate (intra-message meltdown) ───────────────────────
97
98
  { value: "degRun" as const, label: `🌀 degenerate run: ≥ ${c.degenerateMaxRun}`, description: "identical words in a row inside ONE message/call before it counts as stuck generation (noguerol ×5145 class)" },
98
- { value: "degBlock" as const, label: `⛔ block degenerate bash: ${yn(c.blockDegenerateBash)}`, description: "stop a degenerate command before it executes (default on)" },
99
+ { value: "blk" as const, label: `🔁 block repetition: ${yn(c.detectBlockRepeats)}`, description: "ONE message replaying whole sentences/phrases (the narration loop: 'Let me start by checking…' ×5)" },
100
+ { value: "blkShare" as const, label: `🔁 block share: ≥ ${(c.blockRepeatShare * 100).toFixed(0)}%`, description: "share of replayed 5-word phrases in one message before it counts as a replay" },
101
+ { value: "degBlock" as const, label: `⛔ block repetitive bash: ${yn(c.blockDegenerateBash)}`, description: "stop a degenerate or replayed command before it executes (default on)" },
99
102
  // ── 📉 Outcome (no-progress) ────────────────────────────────
100
103
  { value: "outcomeMin" as const, label: `📉 no-progress after: ${c.outcomeMinRepeats}`, description: "same failing outcome repeated this many times (mutated args ≥ 85% similar) before flagging — the NFS test-A…QQQ class" },
101
104
  // ── 📋 Task streams ─────────────────────────────────────────
@@ -227,9 +230,21 @@ async function showConfigMenu(ctx: ExtensionCommandContext, rt: Runtime): Promis
227
230
  if (v !== undefined) { c.degenerateMaxRun = v; saveConfig(c); ctx.ui.notify(`degenerate run: ≥ ${v}`, "info"); }
228
231
  break;
229
232
  }
233
+ case "blk":
234
+ c.detectBlockRepeats = !c.detectBlockRepeats; saveConfig(c);
235
+ ctx.ui.notify(`block repetition: ${yn(c.detectBlockRepeats)}`, "info"); break;
236
+ case "blkShare": {
237
+ const v = await selectFrom(ctx, "🔁 block repetition share (replayed 5-gram positions in ONE message)", [
238
+ { value: 0.7, label: "⚡ 70% (sensitive)" },
239
+ { value: 0.85, label: "🎯 85% (default)" },
240
+ { value: 0.95, label: "🐢 95% (only near-total replays)" },
241
+ ]);
242
+ if (v !== undefined) { c.blockRepeatShare = v; saveConfig(c); ctx.ui.notify(`block share: ${(v * 100).toFixed(0)}%`, "info"); }
243
+ break;
244
+ }
230
245
  case "degBlock":
231
246
  c.blockDegenerateBash = !c.blockDegenerateBash; saveConfig(c);
232
- ctx.ui.notify(`block degenerate bash: ${yn(c.blockDegenerateBash)}`, "info"); break;
247
+ ctx.ui.notify(`block repetitive bash: ${yn(c.blockDegenerateBash)}`, "info"); break;
233
248
  case "outcomeMin": {
234
249
  const v = await selectFrom(ctx, "📉 no-progress threshold (same failing outcome, near-identical args)", [
235
250
  { value: 4, label: "⚡ 4 (sensitive — long experiment series get cut early)" },
package/src/config.ts CHANGED
@@ -24,6 +24,11 @@ export const DEFAULT_CONFIG: AntiloopConfig = {
24
24
  degenerateMaxShare: 0.4,
25
25
  degenerateTurnWeight: 2,
26
26
  blockDegenerateBash: true,
27
+ detectBlockRepeats: true,
28
+ blockMinTokens: 120,
29
+ blockNgram: 5,
30
+ blockMinRepeats: 3,
31
+ blockRepeatShare: 0.85,
27
32
  detectOutcomeLoops: true,
28
33
  outcomeMinRepeats: 8,
29
34
  outcomeArgSimilarity: 0.85,
package/src/detect.ts CHANGED
@@ -185,12 +185,124 @@ export function degenerateDescription(hit: DegenerateHit): string {
185
185
  return `${hit.where}: degenerate repetition — "${hit.token}" ×${hit.freq} (${share}% of ${hit.total} tokens, longest run ${hit.maxRun})`;
186
186
  }
187
187
 
188
+ // ---------------------------------------------------------------------------
189
+ // Intra-message BLOCK repetition (v1.7).
190
+ //
191
+ // The "narration loop" class (verified against a real coding session: ONE
192
+ // assistant message replaying ~5 near-verbatim cycles of "Let me start by
193
+ // checking the environment and the current state of the repository… / I'll run
194
+ // several independent checks in parallel. / Let me begin the S0 development…").
195
+ // Every pre-existing detector stayed silent: text/tool/thinking/structural all
196
+ // compare ACROSS messages (the model produced one message, no peer), and the
197
+ // degenerate detector only watches ONE word repeated hundreds of times, not a
198
+ // whole sentence/paragraph replayed. The signal here is phrase-level: a sliding
199
+ // window of blockNgram-word n-grams over the normalized payload; if almost ALL
200
+ // of those n-grams recur and the most frequent one recurs ≥ blockMinRepeats
201
+ // times, the generation is replaying itself instead of advancing.
202
+ //
203
+ // Deliberately conservative: only a payload where ≥ blockRepeatShare (default
204
+ // 85%) of 5-gram positions repeat qualifies. Ordinary prose — even long,
205
+ // structured docs — sits far below (README/pi.md paragraphs: ≤ 0.11), while
206
+ // 3+ replays of a narration block reach 0.98–1.0. Repeat-linked code/log lines
207
+ // that share a template are NOT flagged because their n-grams carry the varying
208
+ // digits and differ.
209
+ // ---------------------------------------------------------------------------
210
+
211
+ /** Recursive n-gram coverage of one payload. */
212
+ export interface BlockRepeatInfo {
213
+ /** Recurring n-gram positions / total n-gram positions (0..1). */
214
+ ratio: number;
215
+ /** Occurrences of the MOST repeated n-gram. */
216
+ repeats: number;
217
+ /** Total normalized words in the payload. */
218
+ tokens: number;
219
+ /** The n used (words per window). */
220
+ ngram: number;
221
+ /** The most repeated n-gram, for the description. */
222
+ sample: string;
223
+ }
224
+
225
+ export interface BlockRepeatHit extends BlockRepeatInfo {
226
+ where: string;
227
+ }
228
+
229
+ /**
230
+ * Detect phrase/block-level self-repetition inside ONE payload. Returns the
231
+ * coverage ratio with the most repeated n-gram, or undefined for normal text.
232
+ *
233
+ * Reads the payload as a stream of lowercase alphanumeric words (code symbols
234
+ * and punctuation are separators, so stored-JSON escaping cannot glue tokens).
235
+ * Coverage counts an n-gram position as repeated when the SAME n-word sequence
236
+ * (digits included, so templated lines with varying numbers do not count)
237
+ * occurs somewhere else in the payload.
238
+ */
239
+ export function findRepetitiveBlock(text: string, config: AntiloopConfig): BlockRepeatInfo | undefined {
240
+ const words = String(text).toLowerCase().match(/[\p{L}\p{N}]+/gu);
241
+ if (!words) return undefined;
242
+ const total = words.length;
243
+ if (total < config.blockMinTokens) return undefined;
244
+ const n = Math.max(2, config.blockNgram);
245
+ const minRepeats = Math.max(2, config.blockMinRepeats);
246
+ if (total < n * minRepeats) return undefined;
247
+
248
+ const counts = new Map<string, number>();
249
+ const grams: string[] = [];
250
+ for (let i = 0; i + n <= total; i++) {
251
+ const g = words.slice(i, i + n).join(" ");
252
+ grams.push(g);
253
+ counts.set(g, (counts.get(g) ?? 0) + 1);
254
+ }
255
+ if (!grams.length) return undefined;
256
+
257
+ let covered = 0;
258
+ let topKey = "";
259
+ let top = 0;
260
+ for (const g of grams) {
261
+ const c = counts.get(g)!;
262
+ if (c > top) {
263
+ top = c;
264
+ topKey = g;
265
+ }
266
+ if (c >= 2) covered++;
267
+ }
268
+ // At least one n-word phrase must recur blockMinRepeats times (a single
269
+ // echo is a restatement, not a replay) and the recurrence must dominate.
270
+ if (top < minRepeats) return undefined;
271
+ const ratio = covered / grams.length;
272
+ if (ratio < config.blockRepeatShare) return undefined;
273
+ return { ratio, repeats: top, tokens: total, ngram: n, sample: topKey };
274
+ }
275
+
276
+ /** Scan one assistant message (text + each tool-call argument) for a replay. */
277
+ export function scanMessageBlock(
278
+ content: string,
279
+ toolCalls: TrackedToolCall[] | undefined,
280
+ config: AntiloopConfig,
281
+ ): BlockRepeatHit | undefined {
282
+ if (!config.detectBlockRepeats) return undefined;
283
+ if (content && content.length) {
284
+ const b = findRepetitiveBlock(content, config);
285
+ if (b) return { ...b, where: "message text" };
286
+ }
287
+ for (const tc of toolCalls ?? []) {
288
+ if (!tc.args) continue;
289
+ const b = findRepetitiveBlock(tc.args, config);
290
+ if (b) return { ...b, where: tc.name === "bash" ? "bash command" : `args(${tc.name})` };
291
+ }
292
+ return undefined;
293
+ }
294
+
295
+ export function blockRepeatDescription(hit: BlockRepeatHit): string {
296
+ const pct = Math.round(hit.ratio * 100);
297
+ return `${hit.where}: block repetition — ${pct}% of ${hit.tokens} tokens replay repeated ${hit.ngram}-word phrases ("${hit.sample}" ×${hit.repeats})`;
298
+ }
299
+
188
300
  /** Consecutive-detection weight of a detection turn. The strong
189
301
  * self-contained signals (degenerate meltdown, proven no-progress outcome run)
190
302
  * add degenerateTurnWeight (default 2) points so the FIRST one already reaches
191
303
  * the warning level and escalation is fast on repeat. */
192
304
  export function detectionTurnWeight(detections: LoopDetection[], config: AntiloopConfig): number {
193
- const strong = detections.some((d) => d.type === "degenerate" || d.type === "outcome");
305
+ const strong = detections.some((d) => d.type === "degenerate" || d.type === "block" || d.type === "outcome");
194
306
  return strong ? Math.max(1, config.degenerateTurnWeight) : 1;
195
307
  }
196
308
 
@@ -391,9 +503,10 @@ export function detectLoops(state: AntiloopState, config: AntiloopConfig): LoopD
391
503
  const win = msgs.slice(start);
392
504
  const now = Date.now();
393
505
 
394
- // Intra-message degenerate repetition fires on a SINGLE pathological message
395
- // (no peer needed) and is independent of the task-stream batch gate: a
396
- // meltdown is a meltdown even mid-batch.
506
+ // Intra-message repetition fires on a SINGLE pathological message (no peer
507
+ // needed) and is independent of the task-stream batch gate: a meltdown is a
508
+ // meltdown even mid-batch. Two flavors: degenerate (one word ×hundreds) and
509
+ // block (whole sentences/phrases replayed — the narration-loop class).
397
510
  if (config.detectDegenerate) {
398
511
  const last = win[win.length - 1];
399
512
  const hit = scanMessageDegenerate(last.content, last.toolCalls, config);
@@ -408,6 +521,22 @@ export function detectLoops(state: AntiloopState, config: AntiloopConfig): LoopD
408
521
  }
409
522
  }
410
523
 
524
+ if (config.detectBlockRepeats) {
525
+ const last = win[win.length - 1];
526
+ const hit = scanMessageBlock(last.content, last.toolCalls, config);
527
+ if (hit) {
528
+ out.push({
529
+ type: "block",
530
+ // 0.99 => isVerbatimRepeat(): after a force break, a replayed block
531
+ // is the model ignoring the break and escalates to the hard stop.
532
+ similarity: 0.99,
533
+ messageIndices: [msgs.length - 1],
534
+ description: blockRepeatDescription(hit),
535
+ timestamp: now,
536
+ });
537
+ }
538
+ }
539
+
411
540
  // Task-stream gate: if the window is a homogeneous batch (same extension
412
541
  // tool called with DISTINCT content ≥ taskStreamMinCalls times), that tool
413
542
  // is exempt from tool-loop detection, and text/thinking/structural
@@ -675,6 +804,7 @@ export function runSelfTest(): string[] {
675
804
  detectTaskStreams: true, taskStreamMinCalls: 3, taskStreamTwinThreshold: 0.99,
676
805
  detectDegenerate: true, degenerateMinTokens: 50, degenerateMaxRun: 16,
677
806
  degenerateMaxFreq: 60, degenerateMaxShare: 0.4, degenerateTurnWeight: 2, blockDegenerateBash: true,
807
+ detectBlockRepeats: true, blockMinTokens: 120, blockNgram: 5, blockMinRepeats: 3, blockRepeatShare: 0.85,
678
808
  detectOutcomeLoops: true, outcomeMinRepeats: 8, outcomeArgSimilarity: 0.85, outcomeSigThreshold: 0.7,
679
809
  };
680
810
  const asState = (recentMessages: TrackedMessage[]): AntiloopState =>
@@ -782,6 +912,56 @@ export function runSelfTest(): string[] {
782
912
  const wTxt = detectionTurnWeight(dl("text", 0.8), tcfg);
783
913
  out.push(`degenerate turn weight → degenerate ${wDeg}, text ${wTxt} (exp 2, 1) ${wDeg === 2 && wTxt === 1 ? "✅" : "❌"}`);
784
914
 
915
+ // --- v1.7: intra-message BLOCK repetition (narration loops) ---
916
+ // Regression: a real coding session where ONE assistant message replayed the
917
+ // same ~5 sentences in a loop ("Let me start by checking the environment…" /
918
+ // "I'll run several independent checks in parallel." / "Let me begin the S0
919
+ // development…"). Cross-message text/tool/thinking detectors need a peer and
920
+ // stayed silent; the degenerate scan only watches a single word. The block
921
+ // detector must fire on the message itself.
922
+ const NARRV = [
923
+ "Let me start by checking the environment and the current state of the repository, then set up a plan for the S0 slice and begin building.",
924
+ "I'll run several independent checks in parallel.",
925
+ "Let me begin the S0 development. First, reconnaissance of the environment and current repo state.",
926
+ "Let me check what's available in the environment (node, package managers, network, postgres) and the current repo state, then set up the S0 plan and start building the monorepo scaffold.",
927
+ "I'll run a batch of independent environment checks first.",
928
+ ];
929
+ const cycles = (k: number): string => {
930
+ let s = "";
931
+ for (let i = 0; i < k; i++) for (const v of NARRV) s += v + "\n\n";
932
+ return s;
933
+ };
934
+ const s0Det = detectLoops(asState([mk(cycles(5))]), tcfg);
935
+ const s0Hit = s0Det.find((d) => d.type === "block");
936
+ out.push(`block S0 narration loop → ${s0Hit ? `block (${s0Hit.description})` : "no"} (exp block — was the miss) ${s0Hit ? "✅" : "❌"}`);
937
+ out.push(`block turn weight → ${detectionTurnWeight(s0Det, tcfg)} (exp 2 — warn on first sight) ${detectionTurnWeight(s0Det, tcfg) === 2 ? "✅" : "❌"}`);
938
+ out.push(`block post-steer = ignored→ ${isVerbatimRepeat(s0Det) ? "yes" : "no"} (exp yes — replayed block after the break) ${isVerbatimRepeat(s0Det) ? "✅" : "❌"}`);
939
+
940
+ // 3 replay cycles still flag; a single restatement (2 cycles, top-gram ×2)
941
+ // does not — one echo is a summary, not stuck generation.
942
+ const c3 = findRepetitiveBlock(cycles(3), tcfg);
943
+ const c2 = findRepetitiveBlock(cycles(2), tcfg);
944
+ out.push(`block 3 cycles / 2 cycles→ ${c3 ? `flag (${Math.round(c3.ratio * 100)}%)` : "no"} / ${c2 ? "flag" : "no"} (exp flag / no — needs ≥3 repeats) ${c3 && !c2 ? "✅" : "❌"}`);
945
+
946
+ // Ordinary prose and templated (but evolving) payloads must stay silent:
947
+ // long docs score ≤ 0.11 coverage; templated logs/code carry varying digits
948
+ // so their 5-grams differ.
949
+ const prose =
950
+ "The extension watches every assistant message and tool call to decide whether the model is making progress. " +
951
+ "It stores a short window of recent turns, fingerprints tool results and compares them with earlier attempts. " +
952
+ "When a pattern repeats it escalates from a quiet warning to a forced change of approach and finally aborts. " +
953
+ "Configuration lives in a small json file next to the agent directory and every threshold can be tuned at runtime. " +
954
+ "The default values were chosen against real sessions so ordinary work never triggers a false positive. " +
955
+ "Detectors are independent: disabling one leaves the others active and the footer keeps the user informed. " +
956
+ "A clean turn cools the counter down so a recovered model is given room to finish the task. " +
957
+ "Everything is written in plain typescript with no runtime dependencies beyond the host package. " +
958
+ "New strategies should be measured against the recorded payloads before they are enabled by default. " +
959
+ "The goal is simple: catch a stuck model early and keep the context useful for the real work.";
960
+ const logLines = [...Array(30)].map((_, i) => `2026-09-24T10:${String(i).padStart(2, "0")}:00Z INFO worker ${i} processed job ${1000 + i} in ${i * 3}ms status ok`).join("\n");
961
+ const proseHit = findRepetitiveBlock(prose, tcfg);
962
+ const logHit = findRepetitiveBlock(logLines, tcfg);
963
+ out.push(`block legit prose/logs → ${proseHit || logHit ? "flag" : "ok"} (exp ok) ${!proseHit && !logHit ? "✅" : "❌"}`);
964
+
785
965
  // --- v1.6.1: no-progress outcome runs (mutated re-runs, same outcome) ---
786
966
  // Regression: the NFS session — ~90 mutated re-runs of the SAME experiment
787
967
  // (ssh exportfs/mount, labels test A…test QQQ, targets alternating Javi /
package/src/index.ts CHANGED
@@ -24,6 +24,13 @@
24
24
  * hook (blockDegenerateBash). One degenerate turn counts degenerateTurnWeight
25
25
  * (2) consecutive points: warning on first sight, force-break steer on the
26
26
  * second consecutive meltdown, hard stop shortly after if it keeps repeating.
27
+ *
28
+ * v1.7 adds the intra-message BLOCK detector for the narration-loop class:
29
+ * ONE message that replays whole sentences/phrases ("Let me start by checking
30
+ * the environment…" ×5), which every cross-message detector misses because
31
+ * there is no peer message and the degenerate scan only watches single words.
32
+ * Same message_end timing and strong turn weight; the bash gate refuses a
33
+ * command that is itself a replayed block.
27
34
  */
28
35
 
29
36
  import type { ExtensionAPI, ExtensionContext } from "@earendil-works/pi-coding-agent";
@@ -285,29 +292,55 @@ export default function antiloopExtension(pi: ExtensionAPI) {
285
292
  state.recentMessages = state.recentMessages.slice(-(config.detectionWindow + 5));
286
293
  }
287
294
 
288
- // v1.6 — degenerate meltdowns are handled HERE, at message_end, because
289
- // turn_end only fires after tool execution (too late to stop the 46 KB
290
- // "noguerol ×5000" bash from running) and an aborted generation may never
291
- // reach turn_end at all. The signal is self-contained (one pathological
292
- // payload, no peer message) so the full escalation ladder runs right now;
293
- // the turn-guard makes the later turn_end pass skip this message and no
294
- // detection is double-counted.
295
- if (!config.detectDegenerate) return;
296
- const { scanMessageDegenerate, degenerateDescription, nextLevel, isVerbatimRepeat } = await import("./detect.ts");
297
- const hit = scanMessageDegenerate(content, toolCalls, config);
298
- if (!hit) return;
295
+ // v1.6/v1.7 — intra-message repetition is handled HERE, at message_end,
296
+ // because turn_end only fires after tool execution (too late to stop the
297
+ // 46 KB "noguerol ×5000" bash from running) and an aborted generation may
298
+ // never reach turn_end at all. The signal is self-contained (one
299
+ // pathological payload, no peer message) so the full escalation ladder runs
300
+ // right now; the turn-guard makes the later turn_end pass skip this message
301
+ // and no detection is double-counted. Two flavors: degenerate (one word
302
+ // ×hundreds) and block (whole sentences/phrases replayed — the narration
303
+ // loop class where every cross-message detector is blind).
304
+ if (!config.detectDegenerate && !config.detectBlockRepeats) return;
305
+ const {
306
+ scanMessageDegenerate,
307
+ degenerateDescription,
308
+ scanMessageBlock,
309
+ blockRepeatDescription,
310
+ nextLevel,
311
+ isVerbatimRepeat,
312
+ } = await import("./detect.ts");
313
+ let det: LoopDetection | undefined;
314
+ if (config.detectDegenerate) {
315
+ const degHit = scanMessageDegenerate(content, toolCalls, config);
316
+ if (degHit) {
317
+ det = {
318
+ type: "degenerate",
319
+ similarity: 1,
320
+ messageIndices: [state.recentMessages.length - 1],
321
+ description: degenerateDescription(degHit),
322
+ timestamp: Date.now(),
323
+ };
324
+ }
325
+ }
326
+ if (!det && config.detectBlockRepeats) {
327
+ const blkHit = scanMessageBlock(content, toolCalls, config);
328
+ if (blkHit) {
329
+ det = {
330
+ type: "block",
331
+ similarity: 0.99,
332
+ messageIndices: [state.recentMessages.length - 1],
333
+ description: blockRepeatDescription(blkHit),
334
+ timestamp: Date.now(),
335
+ };
336
+ }
337
+ }
338
+ if (!det) return;
299
339
  const tracked = state.recentMessages[state.recentMessages.length - 1];
300
340
  if (tracked) state.lastDetectedTurnIndex = tracked.turnIndex;
301
341
  const prevLevel = state.currentLevel;
302
342
  state.consecutiveDetections += Math.max(1, config.degenerateTurnWeight);
303
343
  state.totalDetections++;
304
- const det: LoopDetection = {
305
- type: "degenerate",
306
- similarity: 1,
307
- messageIndices: [state.recentMessages.length - 1],
308
- description: degenerateDescription(hit),
309
- timestamp: Date.now(),
310
- };
311
344
  state.detections.push(det);
312
345
  if (state.detections.length > config.maxHistoryEntries) {
313
346
  state.detections = state.detections.slice(-config.maxHistoryEntries);
@@ -331,13 +364,13 @@ export default function antiloopExtension(pi: ExtensionAPI) {
331
364
  ctx.ui.notify(`antiloop: force break — ${det.description}`, "error");
332
365
  }
333
366
  } else if (isVerbatimRepeat([det])) {
334
- // The model produced degenerate output AGAIN after the break message:
335
- // count it; once the ignore limit is hit the run is cut for good.
367
+ // The model produced the same intra-message repetition AGAIN after the
368
+ // break message: count it; once the ignore limit is hit, cut the run.
336
369
  state.ignoredSteerCount++;
337
370
  if (state.ignoredSteerCount >= config.ignoredSteerLimit) {
338
371
  hardStop(
339
372
  ctx,
340
- `antiloop: abort — degenerate output repeated ${state.ignoredSteerCount}× after the force break — run stopped; provide new instructions`,
373
+ `antiloop: abort — repetitive output repeated ${state.ignoredSteerCount}× after the force break — run stopped; provide new instructions`,
341
374
  );
342
375
  }
343
376
  }
@@ -379,15 +412,28 @@ export default function antiloopExtension(pi: ExtensionAPI) {
379
412
  * sees why the command was refused and can change approach.
380
413
  */
381
414
  pi.on("tool_call", async (event, ctx) => {
382
- if (!config.enabled || !config.detectDegenerate || !config.blockDegenerateBash) return;
415
+ if (!config.enabled || !config.blockDegenerateBash) return;
383
416
  if (!isToolCallEventType("bash", event)) return;
384
- const { findDegenerateRepetition, degenerateDescription } = await import("./detect.ts");
385
- const hit = findDegenerateRepetition(event.input.command ?? "", config);
386
- if (!hit) return;
387
- return {
388
- block: true,
389
- reason: `[antiloop] blocked: ${degenerateDescription({ ...hit, where: "bash command" })} — this is stuck generation, not a real command. Do NOT retry it: stop and take one small, concrete step instead.`,
390
- };
417
+ const command = event.input.command ?? "";
418
+ const { findDegenerateRepetition, findRepetitiveBlock, degenerateDescription, blockRepeatDescription } = await import("./detect.ts");
419
+ if (config.detectDegenerate) {
420
+ const hit = findDegenerateRepetition(command, config);
421
+ if (hit) {
422
+ return {
423
+ block: true,
424
+ reason: `[antiloop] blocked: ${degenerateDescription({ ...hit, where: "bash command" })} — this is stuck generation, not a real command. Do NOT retry it: stop and take one small, concrete step instead.`,
425
+ };
426
+ }
427
+ }
428
+ if (config.detectBlockRepeats) {
429
+ const blk = findRepetitiveBlock(command, config);
430
+ if (blk) {
431
+ return {
432
+ block: true,
433
+ reason: `[antiloop] blocked: ${blockRepeatDescription({ ...blk, where: "bash command" })} — this is stuck generation, not a real command. Do NOT retry it: stop and take one small, concrete step instead.`,
434
+ };
435
+ }
436
+ }
391
437
  });
392
438
 
393
439
  pi.on("turn_end", async (event, ctx) => {
package/src/types.ts CHANGED
@@ -34,8 +34,30 @@ export interface AntiloopConfig {
34
34
  degenerateMaxShare: number;
35
35
  /** How many consecutive-detection points ONE degenerate turn adds (2 = warn on first sight). */
36
36
  degenerateTurnWeight: number;
37
- /** Block a bash tool call whose command shows degenerate repetition BEFORE it executes. */
37
+ /** Block a bash tool call whose command shows degenerate OR block repetition BEFORE it executes. */
38
38
  blockDegenerateBash: boolean;
39
+ /**
40
+ * Intra-message BLOCK repetition (v1.7): the "narration loop" class where ONE
41
+ * assistant message keeps replaying the same sentences/phrases (the real S0
42
+ * coding session: ~5 near-verbatim cycles of "Let me start by checking the
43
+ * environment…" inside a SINGLE generation, which every cross-message detector
44
+ * misses because it produced only one message and the degenerate detector only
45
+ * watches for ONE word repeated hundreds of times). A sliding window of
46
+ * blockNgram-word n-grams over the normalized payload is built; when at least
47
+ * blockRepeatShare of those positions recur, one n-gram appears at least
48
+ * blockMinRepeats times and the payload has at least blockMinTokens words, the
49
+ * message is a replay. Fires at message_end (before the message's tools run)
50
+ * with the same strong turn weight as the degenerate detector.
51
+ */
52
+ detectBlockRepeats: boolean;
53
+ /** Minimum normalized words in a payload before it is scanned for block repetition. */
54
+ blockMinTokens: number;
55
+ /** Word window of the n-grams whose recurrence is measured (default 5). */
56
+ blockNgram: number;
57
+ /** The MOST repeated n-gram must occur at least this many times to flag. */
58
+ blockMinRepeats: number;
59
+ /** Share of n-gram positions that must recur for the payload to count as a replay. */
60
+ blockRepeatShare: number;
39
61
  /**
40
62
  * No-progress outcome runs (v1.6.1): a model that keeps re-running the SAME
41
63
  * experiment with cosmetic mutations (labels/permutations) while the outcome
@@ -106,7 +128,7 @@ export interface AntiloopConfig {
106
128
  toggleShortcut: string;
107
129
  }
108
130
 
109
- export type LoopKind = "text" | "tool" | "thinking" | "structural" | "degenerate" | "outcome";
131
+ export type LoopKind = "text" | "tool" | "thinking" | "structural" | "degenerate" | "block" | "outcome";
110
132
 
111
133
  export interface LoopDetection {
112
134
  type: LoopKind;