pi-antiloop 1.6.1 → 1.8.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -6,15 +6,16 @@
6
6
 
7
7
  # Antiloop — Loop Detection and Break for pi
8
8
 
9
- **Antiloop watches every assistant message, tool call and thinking block, and forces the model out of reasoning loops before they eat your context and your patience.** Six simultaneous detection strategies (text similarity, tool-call sequences, thinking content, structural openings, degenerate repetition, no-progress outcome runs) find loops that humans miss — and progressive intervention (warning → force break → abort) tells the model to take a different approach, without you having to babysit it.
9
+ **Antiloop watches every assistant message, tool call and thinking block, and forces the model out of reasoning loops before they eat your context and your patience.** Seven simultaneous detection strategies (text similarity, tool-call sequences, thinking content, structural openings, degenerate repetition, block/narration repetition, no-progress outcome runs) find loops that humans miss — and progressive intervention (warning → force break → abort) tells the model to take a different approach, without you having to babysit it.
10
10
 
11
11
  ---
12
12
 
13
13
  ## Features
14
14
 
15
- - **Six detection strategies** — text repetition (trigram Jaccard + Levenshtein), tool-call sequences (name + near-identical arguments + same outcome — result-aware, so retries that make progress don't false-positive), thinking blocks, structural opening-phrase patterns, **degenerate repetition** (a single message/call stuck repeating one word hundreds of times — the `noguerol ×5145` meltdown — caught at `message_end` with no repeated peer needed, and degenerate bash commands blocked before they execute), and **no-progress outcome runs** (the NFS `test A…QQQ` class: dozens of near-identical re-runs of the same experiment, every one ending in the *same failing outcome* — args mutate so tool-loop can't see it; the repeated failure signature can)
15
+ - **Seven detection strategies** — text repetition (trigram Jaccard + Levenshtein), tool-call sequences (name + near-identical arguments + same outcome — result-aware, so retries that make progress don't false-positive), thinking blocks, structural opening-phrase patterns, **degenerate repetition** (a single message/call stuck repeating one word hundreds of times — the `noguerol ×5145` meltdown — caught at `message_end` with no repeated peer needed, and degenerate bash commands blocked before they execute), **block repetition** (ONE message replaying whole sentences/phrases — the narration loop: `Let me start by checking the environment…` ×5 — same `message_end` timing and strong turn weight), and **no-progress outcome runs** (the NFS `test A…QQQ` class: dozens of near-identical re-runs of the same experiment, every one ending in the *same failing outcome* — args mutate so tool-loop can't see it; the repeated failure signature can)
16
16
  - **Stays alive in long sessions** — tracked messages carry a monotonic sequence number, so trimming the sliding window can never make the dedupe guard skip later messages (a bug that silently blinded antiloop after ~15 tracked messages)
17
17
  - **Task-stream recognition (batch work)** — when another extension (e.g. `punched` appending lines to pi.md, or `plan` adding tasks) makes the model call the *same* tool many times with *different* content, antiloop recognizes it as N distinct tasks of one type and stays silent — no warning, no force break. A genuine loop (the *same* call repeated verbatim) is still caught
18
+ - **Snapshot/read-only tool exemption** — a coordinator closing N agents that already finished calls `trimegisto_harvest` once per agent and gets the *same* settled board every time; identical reads of a status snapshot are an idempotent serial close, not a loop, so `snapshotTools` (configurable) is exempt from tool-loop and no-progress detection — real `bash` loops next to it are still caught
18
19
  - **Progressive intervention** — `warning` reminds the model to vary its approach; `force break` steers a real break message into the running agent before its next LLM call; `abort` stops the run entirely
19
20
  - **Configurable thresholds** — independent dials for similarity cutoff, warning/force-break/abort counts, detection window, and which strategies are on
20
21
  - **Sliding window** — only the last N messages are compared, so detection is O(N) in the window size, not in the full session
@@ -100,12 +101,14 @@ Detection strategies:
100
101
  Tool loops: ✅
101
102
  Thinking loops: ✅
102
103
  Degenerate: ✅
104
+ Block repetition: ✅
103
105
  Outcome (no-progress): ✅
104
106
 
105
107
  Recent detections:
106
108
  [text] Text similarity 85% with message 3 (2m ago)
107
109
  [tool] repeated 3x: bash (5m ago)
108
110
  [degenerate] bash command: degenerate repetition — "noguerol" ×5145 (5m ago)
111
+ [block] message text: block repetition — 100% of 450 tokens replay repeated 5-word phrases (5m ago)
109
112
  [outcome] no progress: 9 near-identical bash attempts with the same failing outcome (5m ago)
110
113
  ```
111
114
 
@@ -129,7 +132,9 @@ Grouped interactive menu showing the current value in each option:
129
132
  - **🔁 call repeats** — `1 / 2 / 3` — how many times the same call must repeat before it flags (default 2)
130
133
  - **🧾 result similarity** — `95 / 80 / 60%` — how similar captured results must be to count as the *same outcome*; a repeated command that starts producing a different result is progress, not a loop (default 80%)
131
134
  - **🌀 degenerate run** — `8 / 16 / 24 / 48` — identical words in a row inside ONE message/call before it counts as stuck generation (default 16; the `noguerol ×5145` class)
132
- - **⛔ block degenerate bash** — on/off — refuse a degenerate bash command before it executes, feeding the reason back to the model (default on)
135
+ - **🔁 block repetition** — on/off — detect ONE message replaying whole sentences/phrases (the narration loop: `Let me start by checking the environment…` ×5) with no repeated peer needed
136
+ - **🔁 block share** — `70 / 85 / 95%` — share of replayed 5-word phrases in one message before it counts as a replay (default 85%)
137
+ - **⛔ block repetitive bash** — on/off — refuse a degenerate or replayed bash command before it executes, feeding the reason back to the model (default on)
133
138
  - **📉 no-progress after** — `4 / 6 / 8 / 12` — same failing outcome repeated this many times (near-identical args) before flagging (default 8; the NFS `test A…QQQ` class)
134
139
 
135
140
  **📋 Task streams** — batch work (punched_log, plan_manager, …) is N tasks of one type, not a loop
@@ -137,11 +142,15 @@ Grouped interactive menu showing the current value in each option:
137
142
  - **📋 stream min calls** — `2 / 3 / 4 / 5` — same-tool calls required before a batch is recognized (default 3)
138
143
  - **📋 twin threshold** — `99 / 95 / 90%` — calls more similar than this count as the *same task* repeated; one twin invalidates the batch and normal detection resumes (default 99%)
139
144
 
145
+ **📡 Snapshot tools** — read-only status boards (trimegisto_harvest, …) read serially are not a loop
146
+ - **📡 snapshot tools** — `trimegisto_harvest` / off — repeated identical reads of a settled board stay silent all the way through (default on; add any other status/poll tool in `antiloop.json`)
147
+
140
148
  **🔍 Detectors**
141
149
  - **📝 text** — on/off — detect repeated text messages
142
150
  - **🔧 tools** — on/off — detect repeated tool calls
143
151
  - **🧠 thinking** — on/off — detect repeated internal reasoning
144
152
  - **🌀 degenerate** — on/off — detect one message stuck repeating a single word/token (no repeated peer needed)
153
+ - **🔁 block** — on/off — detect one message replaying whole sentences/phrases (narration loop, no repeated peer needed)
145
154
  - **📉 outcome** — on/off — detect many near-identical attempts all ending in the same failing outcome (no progress)
146
155
 
147
156
  **🧹 reset state** — clear all counters and history
@@ -175,6 +184,9 @@ degenerate legit cmd → ok (exp ok) ✅
175
184
  degenerate interleaved → flag (noguerol ×150) (exp flag — freq/share clause) ✅
176
185
  degenerate glued token → flag (noguerol ×300) (exp flag — perfect power) ✅
177
186
  degenerate turn weight → degenerate 2, text 1 (exp 2, 1) ✅
187
+ block S0 narration loop → block (message text: 100% of 450 tokens replay repeated 5-word phrases) ✅ ← v1.7: ONE message replaying sentences, no peer
188
+ block 3 cycles / 2 cycles→ flag (100%) / no (exp flag / no — needs ≥3 repeats) ✅
189
+ block legit prose/logs → ok (exp ok) ✅
178
190
  outcome fires on 9th → outcome (9 near-identical bash attempts, same failing outcome) ✅ ← v1.6.1: NFS test-A…QQQ class
179
191
  outcome needs 8 prior → silent (exp silent at 5 attempts) ✅
180
192
  outcome converging sweep → silent (exp silent — outcomes differ = progress) ✅
@@ -186,7 +198,7 @@ outcome identical OKs → silent (exp silent — success repeats ≠ loop)
186
198
 
187
199
  ### Detection pipeline
188
200
 
189
- After every assistant `message_end` event, antiloop extracts the new content (text, thinking, tool calls — including their ids) and pushes it onto a sliding window of the last `detectionWindow + 5` messages. Detection itself runs at `turn_end`, once the tool results are known: results are fingerprinted and attached to the tracked calls, then the active detection strategies run against the window. The one exception is the **degenerate** strategy: a single stuck message needs no peer and no tool result, so it is evaluated right at `message_end` — the only point before the message's own tool calls execute — and degenerate `bash` calls are also blocked at the `tool_call` hook.
201
+ After every assistant `message_end` event, antiloop extracts the new content (text, thinking, tool calls — including their ids) and pushes it onto a sliding window of the last `detectionWindow + 5` messages. Detection itself runs at `turn_end`, once the tool results are known: results are fingerprinted and attached to the tracked calls, then the active detection strategies run against the window. The exceptions are the **intra-message** strategies — **degenerate** and **block**. A single stuck message needs no peer and no tool result, so both are evaluated right at `message_end` — the only point before the message's own tool calls execute — and repetitive `bash` calls are also blocked at the `tool_call` hook.
190
202
 
191
203
  | Strategy | What it compares | Algorithm |
192
204
  |----------|------------------|-----------|
@@ -195,10 +207,11 @@ After every assistant `message_end` event, antiloop extracts the new content (te
195
207
  | Thinking | Internal reasoning/thinking blocks | Same as text |
196
208
  | Structural | First 10 words of each message | Opening-phrase similarity ≥ 90% across ≥ 3 messages |
197
209
  | Degenerate | One single message/call (no peer needed) | Run-length + frequency of identical words inside the payload: ≥ `degenerateMaxRun` (default 16) consecutive identical words, or one word ≥ `degenerateMaxFreq`× at ≥ `degenerateMaxShare` of all tokens; plus a perfect-power check for glued no-space tokens. Scanned at `message_end` — before the tool calls execute — and on every `bash` `tool_call` (blocking gate) |
210
+ | Block | One single message (no peer needed) | Sliding `blockNgram`-word n-grams over the normalized payload: a payload of ≥ `blockMinTokens` (120) words is a replay when ≥ `blockRepeatShare` (default 85%) of the n-gram positions recur AND the most repeated n-gram appears ≥ `blockMinRepeats` (3) times. Catches a message replaying whole sentences/phrases — the narration loop (`Let me start by checking the environment…` ×5) that every cross-message detector misses because there is no peer. Scanned at `message_end` and on every `bash` `tool_call` (blocking gate) |
198
211
  | Outcome | Single tool calls across the window, after the last user input | ≥ `outcomeMinRepeats` (default 8) PRIOR attempts with args ≥ `outcomeArgSimilarity` (0.85) similar AND the same *failing* outcome (failure signatures compared at ≥ `outcomeSigThreshold`, 0.7; identical OK results never count — they're the norm for batches). Catches mutated re-run loops the tool detector can't see (labels/permutations change every turn) |
199
212
  | Task stream | Same tool, many calls | When a tool appears ≥ `taskStreamMinCalls` times (default 3) in the window and *no two* calls are near-identical (`taskStreamTwinThreshold`, default 99%), the tool is an active batch: N different tasks of one type (e.g. `punched_log` appends, `plan_manager` task adds). Those calls are exempt from tool-loop detection, and text/thinking/structural patterns that only involve those batch messages are suppressed too. If even one call pair is a twin (the same task repeated), the tool is *not* a stream and detection proceeds normally |
200
213
 
201
- Each detected pair becomes a `LoopDetection { type, similarity, messageIndices, description }` and the consecutive counter increases (a degenerate turn counts `degenerateTurnWeight`, default 2 — warning on first sight).
214
+ Each detected pair becomes a `LoopDetection { type, similarity, messageIndices, description }` and the consecutive counter increases (a degenerate or block turn counts `degenerateTurnWeight`, default 2 — warning on first sight).
202
215
 
203
216
  ### Intervention levels
204
217
 
@@ -211,7 +224,7 @@ Each detected pair becomes a `LoopDetection { type, similarity, messageIndices,
211
224
 
212
225
  The level never de-escalates during an active loop; user input decays the consecutive counter naturally so a fresh prompt can break the cycle.
213
226
 
214
- **Degenerate turns are handled earlier than the ladder:** the meltdown is detected at `message_end` (the same moment the text/tool-call payload is complete, *before* pi preflights and executes its tools). A single degenerate turn already adds `degenerateTurnWeight` (2) consecutive points → warning on first sight; the second consecutive meltdown → force break steer; after the steer, further degenerate output counts against `ignoredSteerLimit` → hard stop. Degenerate `bash` calls are additionally refused by the `tool_call` gate (`blockDegenerateBash`) — the command never runs, and the block reason is fed back to the model as the tool error so it can still change approach.
227
+ **Degenerate and block turns are handled earlier than the ladder:** they are detected at `message_end` (the same moment the text/tool-call payload is complete, *before* pi preflights and executes its tools). A single such turn already adds `degenerateTurnWeight` (2) consecutive points → warning on first sight; the second consecutive one → force break steer; after the steer, further degenerate/blocked output counts against `ignoredSteerLimit` → hard stop. Repetitive `bash` calls are additionally refused by the `tool_call` gate (`blockDegenerateBash`) — the command never runs, and the block reason is fed back to the model as the tool error so it can still change approach.
215
228
 
216
229
  ### Similarity scoring
217
230
 
@@ -252,6 +265,7 @@ If the same command produced a *different* outcome, the pair is progress:
252
265
 
253
266
  Results only veto; they never trigger on their own, and calls without a
254
267
  captured result fall back to argument matching alone.
268
+ ```
255
269
 
256
270
  ### Task streams: N tasks of one type ≠ a loop
257
271
 
@@ -287,6 +301,27 @@ size needed before recognition), `taskStreamTwinThreshold` (how similar args
287
301
  must be to count as *the same task* — lower it to treat near-duplicate
288
302
  entries as loops again).
289
303
 
304
+ ### Snapshot tools: the serial close of settled agents ≠ a loop
305
+
306
+ A coordinator that closes N agents which have already finished calls the
307
+ snapshot tool once per agent. When the agents are all settled, every one of
308
+ those calls returns the **same** serialized board — same arguments (`{}`),
309
+ same result. To the tool detector that is indistinguishable from a verbatim
310
+ loop: `same args + same outcome` repeated N times, so it force-broke a run
311
+ that was simply collecting finished results.
312
+
313
+ Re-reading a state snapshot is an *idempotent read*, not a stuck generation,
314
+ so tools listed in `snapshotTools` (default `["trimegisto_harvest"]`) are
315
+ exempt from tool-loop **and** no-progress outcome detection. The exemption is
316
+ deliberately narrow:
317
+
318
+ - it only applies when *every* call in the message is a snapshot tool — a
319
+ real `bash` loop re-run alongside a harvest is still flagged;
320
+ - it is name-based and configurable, so `antiloop.json` can list any other
321
+ status/poll tool (`"snapshotTools": ["trimegisto_harvest", "my_status"]`);
322
+ - clearing it (`"snapshotTools": []` or **📡 off**) restores normal detection,
323
+ where the very same 5× harvest payload is a genuine verbatim tool loop.
324
+
290
325
  ### Degenerate repetition: ONE message stuck on a word ≠ a reasoning loop
291
326
 
292
327
  All strategies above need at least two similar messages — they detect a model
@@ -333,6 +368,49 @@ sweeps that motivated the tool-loop threshold). Only `bash` gets the blocking
333
368
  gate — writing a repetitive *file* (e.g. a user-requested padding fixture) is
334
369
  alerted and escalated, not refused.
335
370
 
371
+ ### Block repetition: ONE message replaying its own sentences
372
+
373
+ The degenerate detector catches a single *word* repeated hundreds of times.
374
+ There is a second intra-message failure mode, just as obvious to a human and
375
+ invisible to every cross-message detector: **the model replays whole
376
+ sentences/phrases inside one generation**. Real case (a coding session's S0
377
+ scaffold): ONE assistant message cycled ~5 times through
378
+
379
+ > Let me start by checking the environment and the current state of the
380
+ > repository, then set up a plan for the S0 slice and begin building.
381
+ > I'll run several independent checks in parallel.
382
+ > Let me begin the S0 development. First, reconnaissance of the environment
383
+ > and current repo state.
384
+
385
+ …never emitting a tool call. Text/tool/thinking/structural detection all
386
+ compare *across* messages and had nothing to compare against; the degenerate
387
+ scan only watches for one word repeated, and no single word repeated 16×.
388
+
389
+ Antiloop v1.7 detects **block repetition** on the message itself:
390
+
391
+ - the payload is read as a stream of lowercase alphanumeric words;
392
+ - a sliding window of `blockNgram` (default 5) words is built over it;
393
+ - if at least `blockRepeatShare` (default **85%**) of those n-gram positions
394
+ recur somewhere else, the most repeated n-gram appears at least
395
+ `blockMinRepeats` (default 3) times, and the payload has at least
396
+ `blockMinTokens` (default 120) words, the generation is replaying itself.
397
+
398
+ The thresholds are deliberately far apart from normal content: ordinary
399
+ prose — even long, structured documents — scores ≤ 11% coverage, while 3+
400
+ replays of a narration block reach 98–100%. Templated payloads with evolving
401
+ data (log lines, generated code) are not flagged because their n-grams carry
402
+ the varying digits and differ. A single restatement (2 cycles) does not trip
403
+ it — one echo is a summary, three are stuck generation.
404
+
405
+ Because the signal is conclusive it is handled like the degenerate one:
406
+ caught at `message_end` (before the message's tools execute), weighted
407
+ `degenerateTurnWeight` (2) so the first occurrence warns, and a replayed
408
+ `bash` command is refused by the same `blockDegenerateBash` gate. After a
409
+ force-break steer, another replayed block escalates to the hard stop.
410
+
411
+ Tunables: `detectBlockRepeats`, `blockRepeatShare` (lower = earlier),
412
+ `blockMinRepeats`, `blockNgram`, `blockMinTokens`.
413
+
336
414
  ### No-progress outcome runs: mutated re-runs, same wall
337
415
 
338
416
  A subtler meltdown than the degenerate one: the model re-runs the SAME
@@ -372,7 +450,6 @@ real session this fires at `test MM` (warn) → `NN` (steer) → `PP` (abort)
372
450
 
373
451
  Tunables: `detectOutcomeLoops`, `outcomeMinRepeats` (lower = earlier cutoff),
374
452
  `outcomeArgSimilarity`, `outcomeSigThreshold`.
375
- ```
376
453
 
377
454
  ### Sliding window
378
455
 
@@ -399,6 +476,11 @@ Persisted as JSON at `~/.pi/agent/antiloop.json`:
399
476
  "degenerateMaxShare": 0.4,
400
477
  "degenerateTurnWeight": 2,
401
478
  "blockDegenerateBash": true,
479
+ "detectBlockRepeats": true,
480
+ "blockMinTokens": 120,
481
+ "blockNgram": 5,
482
+ "blockMinRepeats": 3,
483
+ "blockRepeatShare": 0.85,
402
484
  "detectOutcomeLoops": true,
403
485
  "outcomeMinRepeats": 8,
404
486
  "outcomeArgSimilarity": 0.85,
@@ -433,7 +515,12 @@ Persisted as JSON at `~/.pi/agent/antiloop.json`:
433
515
  | `degenerateMaxFreq` | `60` | One word's total occurrences (with `degenerateMaxShare` of the payload) that flags interleaved meltdowns |
434
516
  | `degenerateMaxShare` | `0.4` | Frequency share (freq/total tokens) required together with `degenerateMaxFreq` |
435
517
  | `degenerateTurnWeight` | `2` | Consecutive-detection points added by one degenerate turn (2 = warning on first sight) |
436
- | `blockDegenerateBash` | `true` | Block a degenerate `bash` command in the `tool_call` hook before it executes; the reason is fed back to the model as the tool error |
518
+ | `blockDegenerateBash` | `true` | Block a degenerate or block-repetition `bash` command in the `tool_call` hook before it executes; the reason is fed back to the model as the tool error |
519
+ | `detectBlockRepeats` | `true` | Detect intra-message block/narration repetition — ONE message replaying whole sentences/phrases (the `Let me start by checking the environment…` ×5 class). Needs no repeated peer; scanned at `message_end` and on every bash `tool_call` |
520
+ | `blockMinTokens` | `120` | Minimum normalized words in a payload before it is scanned for block repetition (shorter payloads aren't conclusive) |
521
+ | `blockNgram` | `5` | Word window of the n-grams whose recurrence is measured |
522
+ | `blockMinRepeats` | `3` | The most repeated n-gram must occur at least this many times to flag |
523
+ | `blockRepeatShare` | `0.85` | Share of n-gram positions that must recur for the payload to count as a replay |
437
524
  | `detectOutcomeLoops` | `true` | No-progress outcome runs: ≥ `outcomeMinRepeats` near-identical attempts (args ≥ `outcomeArgSimilarity`) all ending in the *same failing outcome* — the NFS `test A…QQQ` class |
438
525
  | `outcomeMinRepeats` | `8` | Prior same-failure attempts (inside the window, after the last user input) required before the outcome detector fires |
439
526
  | `outcomeArgSimilarity` | `0.85` | How similar args must be to count as the *same experiment reshuffled* (mutations of labels/permutations stay under it — distinct tasks don't) |
@@ -441,6 +528,7 @@ Persisted as JSON at `~/.pi/agent/antiloop.json`:
441
528
  | `detectTaskStreams` | `true` | Recognize homogeneous batch work (same tool called with distinct content — e.g. punched/plan/obsidian extensions) and stay silent; see [task streams](#task-streams-n-tasks-of-one-type--a-loop) |
442
529
  | `taskStreamMinCalls` | `3` | Same-tool calls required inside the window before a task stream is recognized |
443
530
  | `taskStreamTwinThreshold` | `0.99` | Arguments this similar (or identical) count as *the same task* — a twin invalidates the stream and re-enables normal loop detection |
531
+ | `snapshotTools` | `["trimegisto_harvest"]` | Read-only snapshot/status tools: repeated identical reads (the serial close of N settled agents) never count as a tool loop or no-progress run. Add any other status/poll tool by name |
444
532
  | `detectTextLoops` | `true` | Detect full-text repetition |
445
533
  | `detectToolLoops` | `true` | Detect tool-call sequence + argument repetition |
446
534
  | `detectThinkingLoops` | `true` | Detect repeated thinking/reasoning content |
@@ -460,7 +548,8 @@ Persisted as JSON at `~/.pi/agent/antiloop.json`:
460
548
  6. **Let user input clear state** — each user message decays the consecutive counter by 2, so a fresh prompt naturally resets without `/antiloop reset`.
461
549
  7. **Degenerate detector needs no tuning for most setups** — a run of ≥ 16 identical words (or one word ≥ 40% of a ≥ 50-token payload) inside a single message is conclusive stuck generation; the 46 KB `noguerol ×5145` SSH-wordlist meltdown is caught on first sight (warning), its bash never executes (`blockDegenerateBash`), and a second consecutive meltdown gets the force-break steer. If a model legitimately writes repetitive payloads, raise `degenerateMaxRun` / `degenerateMaxFreq` via `/antiloop config` — don't disable the detector.
462
550
  8. **Outcome detector catches mutated re-run loops** — a model that re-issues the same experiment with cosmetic changes (labels, permutations) while every attempt fails identically gets a warning after `outcomeMinRepeats` (8) same-failure attempts, a steer on the next, and a hard stop shortly after. Converging sweeps, evolving failures and repeated successes stay silent by design. Lower `outcomeMinRepeats` if you want earlier cutoffs.
463
- 9. **`/antiloop test`** — runs the real detection engine (text + tool-call + task-stream + degenerate + outcome regression cases) to verify calibration after any change.
551
+ 9. **Block detector catches narration loops** — a single message replaying whole sentences (the `Let me start by checking the environment…` ×5 class) warns on first sight and escalates like a meltdown. It only fires when ≥ 85% of the message's 5-word phrases recur, so ordinary prose and templated logs/code stay silent. Raise `blockRepeatShare` (or `blockMinRepeats`) if you ever see a false positive.
552
+ 10. **`/antiloop test`** — runs the real detection engine (text + tool-call + task-stream + degenerate + block + outcome regression cases) to verify calibration after any change.
464
553
 
465
554
  ## Architecture
466
555
 
@@ -485,10 +574,10 @@ Modular extension with zero external dependencies (only pi's bundled `@earendil-
485
574
 
486
575
  - **Levenshtein + trigram Jaccard** hybrid — small texts use edit distance, large texts use n-gram overlap (each is O(N) in text length)
487
576
  - **Sliding window** — only the last `detectionWindow` messages participate, capping memory at O(W × message_size)
488
- - **Early bail** — short messages and empty tool calls skip similarity computation entirely; the degenerate scan is a single linear tokenization pass
577
+ - **Early bail** — short messages and empty tool calls skip similarity computation entirely; the degenerate and block scans are single linear passes
489
578
  - **TUI integration** — uses `ctx.ui.select` for the config menu and the log viewer; `ctx.ui.notify` for state notifications; `ctx.ui.setStatus` + a custom `ctx.ui.setFooter` component for the persistent footer indicator, live level info, and the `esc+a` keyboard toggle (`ctx.ui.onTerminalInput`, never consumes input)
490
- - **Hooks** — `message_end` (track messages + tool call ids with a monotonic sequence so the sliding-window trim can never collide turn indices, and pre-handle degenerate meltdowns — the message is complete but its tools haven't executed yet), `tool_call` (block degenerate `bash` commands before they run), `turn_end` (attach result fingerprints with failure signatures, detect — including no-progress outcome runs — and intervene: steer the force break / abort the run), `input` (decay on real user messages only), `session_start` (load config + install footer + reset), `session_shutdown` (restore built-in footer)
491
- - **Intervention runs on the turn loop, not on user prompts** — escalation is decided at `turn_end` (and at `message_end` for the self-contained degenerate signal), the break is steered into the running agent before its next LLM call, and the guaranteed hard stop aborts the run (`ctx.abort`, fire-and-forget — never awaited, so the hook can't deadlock). No custom-role messages are injected into the conversation at any level (steering a real user message + aborting are the only levers; custom-role injections were removed because a model can stall on an unexpected injected message)
579
+ - **Hooks** — `message_end` (track messages + tool call ids with a monotonic sequence so the sliding-window trim can never collide turn indices, and pre-handle degenerate/block repetition — the message is complete but its tools haven't executed yet), `tool_call` (block degenerate or replayed `bash` commands before they run), `turn_end` (attach result fingerprints with failure signatures, detect — including no-progress outcome runs — and intervene: steer the force break / abort the run), `input` (decay on real user messages only), `session_start` (load config + install footer + reset), `session_shutdown` (restore built-in footer)
580
+ - **Intervention runs on the turn loop, not on user prompts** — escalation is decided at `turn_end` (and at `message_end` for the self-contained intra-message signals), the break is steered into the running agent before its next LLM call, and the guaranteed hard stop aborts the run (`ctx.abort`, fire-and-forget — never awaited, so the hook can't deadlock). No custom-role messages are injected into the conversation at any level (steering a real user message + aborting are the only levers; custom-role injections were removed because a model can stall on an unexpected injected message)
492
581
 
493
582
  ## License
494
583
 
package/package.json CHANGED
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "name": "pi-antiloop",
3
- "version": "1.6.1",
4
- "description": "Antiloop: detect reasoning loops and force a break (warn → force → abort) across text, tool, thinking, structural, degenerate-repetition, and no-progress outcome patterns. Degenerate (one message stuck repeating a word hundreds of times — noguerol ×5145) is caught at message_end and its bash blocked before executing. No-progress (the NFS test-A…QQQ class: ~90 mutated re-runs of the same experiment, every one failing identically) fires when the same failing outcome repeats ≥ outcomeMinRepeats times with near-identical args. Fixes antiloop going blind mid-session: tracked messages use a monotonic sequence so the trim of the sliding window can never collide turn indices. The force break is delivered mid-run (steer before the next LLM call) and if the model ignores it antiloop aborts the run — an autonomous tool loop always terminates. Result-aware tool-loop detection keeps sequential bash and converging sweeps quiet; task-stream recognition keeps punched/plan batch work quiet.",
3
+ "version": "1.8.0",
4
+ "description": "Antiloop: detect reasoning loops and force a break (warn → force → abort) across text, tool, thinking, structural, degenerate-repetition, block/narration-repetition, and no-progress outcome patterns. Degenerate (one message stuck repeating a word hundreds of times — noguerol ×5145) is caught at message_end and its bash blocked before executing. Block (ONE message replaying whole sentences/phrases — the 'Let me start by checking the environment…' ×5 narration loop, invisible to every cross-message detector because there is no peer) fires on the first such message with the same strong turn weight and blocks a replayed bash command. No-progress (the NFS test-A…QQQ class: ~90 mutated re-runs of the same experiment, every one failing identically) fires when the same failing outcome repeats ≥ outcomeMinRepeats times with near-identical args. Snapshot tools (trimegisto_harvest) are exempt: a coordinator closing N settled agents re-reads the SAME board serially, which is an idempotent read, not a loop — so isVerbatimRepeat never fires on it. Fixes antiloop going blind mid-session: tracked messages use a monotonic sequence so the trim of the sliding window can never collide turn indices. The force break is delivered mid-run (steer before the next LLM call) and if the model ignores it antiloop aborts the run — an autonomous tool loop always terminates. Result-aware tool-loop detection keeps sequential bash and converging sweeps quiet; task-stream recognition keeps punched/plan batch work quiet.",
5
5
  "keywords": [
6
6
  "pi-package",
7
7
  "antiloop",
package/src/commands.ts CHANGED
@@ -59,10 +59,12 @@ async function showStatus(ctx: ExtensionCommandContext, rt: Runtime): Promise<vo
59
59
  ` tool sim: ${(rt.config.toolSimilarityThreshold * 100).toFixed(0)}% tool repeat: ${rt.config.minToolRepeatCount}+ prior`,
60
60
  ` result sim: ${(rt.config.resultSimilarityThreshold * 100).toFixed(0)}% (same cmd + diff outcome = no loop)`,
61
61
  ` degenerate: run ≥ ${rt.config.degenerateMaxRun} same word · freq ≥ ${rt.config.degenerateMaxFreq} @ ${(rt.config.degenerateMaxShare * 100).toFixed(0)}% (≥ ${rt.config.degenerateMinTokens} tokens) · weight ${rt.config.degenerateTurnWeight} · block bash ${yn(rt.config.blockDegenerateBash)}`,
62
+ ` block repeats: ${yn(rt.config.detectBlockRepeats)} (≥ ${(rt.config.blockRepeatShare * 100).toFixed(0)}% of ≥ ${rt.config.blockMinTokens} tokens replay ${rt.config.blockNgram}-grams ×${rt.config.blockMinRepeats}+)`,
62
63
  ` outcome: same failing result ≥ ${rt.config.outcomeMinRepeats} attempts (args ≥ ${(rt.config.outcomeArgSimilarity * 100).toFixed(0)}% sim, sig ≥ ${(rt.config.outcomeSigThreshold * 100).toFixed(0)}%)`,
63
64
  ` task streams: ${yn(rt.config.detectTaskStreams)} (min ${rt.config.taskStreamMinCalls} calls, twins ≥ ${(rt.config.taskStreamTwinThreshold * 100).toFixed(0)}%)`,
65
+ ` snapshot tools: ${rt.config.snapshotTools?.length ? rt.config.snapshotTools.join(", ") : "off"} (repeated reads are not a loop)`,
64
66
  "",
65
- `detectors: text ${yn(rt.config.detectTextLoops)} · tool ${yn(rt.config.detectToolLoops)} · think ${yn(rt.config.detectThinkingLoops)} · degenerate ${yn(rt.config.detectDegenerate)} · outcome ${yn(rt.config.detectOutcomeLoops)}`,
67
+ `detectors: text ${yn(rt.config.detectTextLoops)} · tool ${yn(rt.config.detectToolLoops)} · think ${yn(rt.config.detectThinkingLoops)} · degenerate ${yn(rt.config.detectDegenerate)} · block ${yn(rt.config.detectBlockRepeats)} · outcome ${yn(rt.config.detectOutcomeLoops)}`,
66
68
  `footer: interactive ${yn(rt.config.interactiveFooter)} · toggle: ${rt.config.toggleShortcut}`,
67
69
  ];
68
70
  if (rt.state.activeTaskStreams.length) {
@@ -95,13 +97,16 @@ async function showConfigMenu(ctx: ExtensionCommandContext, rt: Runtime): Promis
95
97
  { value: "resultSim" as const, label: `🧾 result similarity: ${(c.resultSimilarityThreshold * 100).toFixed(0)}%`, description: "same command + different result = progress, not a loop" },
96
98
  // ── 🌀 Degenerate (intra-message meltdown) ───────────────────────
97
99
  { value: "degRun" as const, label: `🌀 degenerate run: ≥ ${c.degenerateMaxRun}`, description: "identical words in a row inside ONE message/call before it counts as stuck generation (noguerol ×5145 class)" },
98
- { value: "degBlock" as const, label: `⛔ block degenerate bash: ${yn(c.blockDegenerateBash)}`, description: "stop a degenerate command before it executes (default on)" },
100
+ { value: "blk" as const, label: `🔁 block repetition: ${yn(c.detectBlockRepeats)}`, description: "ONE message replaying whole sentences/phrases (the narration loop: 'Let me start by checking…' ×5)" },
101
+ { value: "blkShare" as const, label: `🔁 block share: ≥ ${(c.blockRepeatShare * 100).toFixed(0)}%`, description: "share of replayed 5-word phrases in one message before it counts as a replay" },
102
+ { value: "degBlock" as const, label: `⛔ block repetitive bash: ${yn(c.blockDegenerateBash)}`, description: "stop a degenerate or replayed command before it executes (default on)" },
99
103
  // ── 📉 Outcome (no-progress) ────────────────────────────────
100
104
  { value: "outcomeMin" as const, label: `📉 no-progress after: ${c.outcomeMinRepeats}`, description: "same failing outcome repeated this many times (mutated args ≥ 85% similar) before flagging — the NFS test-A…QQQ class" },
101
105
  // ── 📋 Task streams ─────────────────────────────────────────
102
106
  { value: "streams" as const, label: `📋 task streams: ${yn(c.detectTaskStreams)}`, description: "batch work (punched_log / plan_manager / …) is not a loop" },
103
107
  { value: "streamMin" as const, label: `📋 stream min calls: ${c.taskStreamMinCalls}`, description: "calls of the same tool before a batch is recognized" },
104
108
  { value: "streamTwin" as const, label: `📋 twin threshold: ${(c.taskStreamTwinThreshold * 100).toFixed(0)}%`, description: "calls more similar than this = the same task repeated, not a batch" },
109
+ { value: "snapshot" as const, label: `📡 snapshot tools: ${c.snapshotTools?.length ? c.snapshotTools.join(", ") : "off"}`, description: "read-only status tools (trimegisto_harvest): re-reading settled agents serially is not a loop" },
105
110
  // ── 🔍 Detectors ────────────────────────────────────────────
106
111
  { value: "text" as const, label: `📝 text: ${yn(c.detectTextLoops)}`, description: "detect repeated text messages" },
107
112
  { value: "tool" as const, label: `🔧 tools: ${yn(c.detectToolLoops)}`, description: "detect repeated tool calls" },
@@ -227,9 +232,21 @@ async function showConfigMenu(ctx: ExtensionCommandContext, rt: Runtime): Promis
227
232
  if (v !== undefined) { c.degenerateMaxRun = v; saveConfig(c); ctx.ui.notify(`degenerate run: ≥ ${v}`, "info"); }
228
233
  break;
229
234
  }
235
+ case "blk":
236
+ c.detectBlockRepeats = !c.detectBlockRepeats; saveConfig(c);
237
+ ctx.ui.notify(`block repetition: ${yn(c.detectBlockRepeats)}`, "info"); break;
238
+ case "blkShare": {
239
+ const v = await selectFrom(ctx, "🔁 block repetition share (replayed 5-gram positions in ONE message)", [
240
+ { value: 0.7, label: "⚡ 70% (sensitive)" },
241
+ { value: 0.85, label: "🎯 85% (default)" },
242
+ { value: 0.95, label: "🐢 95% (only near-total replays)" },
243
+ ]);
244
+ if (v !== undefined) { c.blockRepeatShare = v; saveConfig(c); ctx.ui.notify(`block share: ${(v * 100).toFixed(0)}%`, "info"); }
245
+ break;
246
+ }
230
247
  case "degBlock":
231
248
  c.blockDegenerateBash = !c.blockDegenerateBash; saveConfig(c);
232
- ctx.ui.notify(`block degenerate bash: ${yn(c.blockDegenerateBash)}`, "info"); break;
249
+ ctx.ui.notify(`block repetitive bash: ${yn(c.blockDegenerateBash)}`, "info"); break;
233
250
  case "outcomeMin": {
234
251
  const v = await selectFrom(ctx, "📉 no-progress threshold (same failing outcome, near-identical args)", [
235
252
  { value: 4, label: "⚡ 4 (sensitive — long experiment series get cut early)" },
@@ -262,6 +279,18 @@ async function showConfigMenu(ctx: ExtensionCommandContext, rt: Runtime): Promis
262
279
  if (v !== undefined) { c.taskStreamTwinThreshold = v; saveConfig(c); ctx.ui.notify(`twin threshold: ${(v * 100).toFixed(0)}%`, "info"); }
263
280
  break;
264
281
  }
282
+ case "snapshot": {
283
+ const v = await selectFrom(ctx, "📡 snapshot / read-only tools (repeated identical reads are never a loop)", [
284
+ { value: "default" as const, label: "🎯 trimegisto_harvest (default)", description: "serial close of settled agents: same args + same snapshot ×N stays silent" },
285
+ { value: "off" as const, label: "🚫 off", description: "treat every tool identically (harvest can loop-flag again)" },
286
+ ]);
287
+ if (v !== undefined) {
288
+ c.snapshotTools = v === "off" ? [] : ["trimegisto_harvest"];
289
+ saveConfig(c);
290
+ ctx.ui.notify(`snapshot tools: ${c.snapshotTools.length ? c.snapshotTools.join(", ") : "off"}`, "info");
291
+ }
292
+ break;
293
+ }
265
294
  case "ignoredBreak": {
266
295
  const v = await selectFrom(ctx, "🛑 stop after ignored break (identical repeats after the force break)", [
267
296
  { value: 1, label: "⚡ 1 (sensitive — one identical repeat after the break stops the run)" },
package/src/config.ts CHANGED
@@ -24,6 +24,11 @@ export const DEFAULT_CONFIG: AntiloopConfig = {
24
24
  degenerateMaxShare: 0.4,
25
25
  degenerateTurnWeight: 2,
26
26
  blockDegenerateBash: true,
27
+ detectBlockRepeats: true,
28
+ blockMinTokens: 120,
29
+ blockNgram: 5,
30
+ blockMinRepeats: 3,
31
+ blockRepeatShare: 0.85,
27
32
  detectOutcomeLoops: true,
28
33
  outcomeMinRepeats: 8,
29
34
  outcomeArgSimilarity: 0.85,
@@ -31,6 +36,7 @@ export const DEFAULT_CONFIG: AntiloopConfig = {
31
36
  detectTaskStreams: true,
32
37
  taskStreamMinCalls: 3,
33
38
  taskStreamTwinThreshold: 0.99,
39
+ snapshotTools: ["trimegisto_harvest"],
34
40
  detectToolLoops: true,
35
41
  detectThinkingLoops: true,
36
42
  detectTextLoops: true,
package/src/detect.ts CHANGED
@@ -185,12 +185,124 @@ export function degenerateDescription(hit: DegenerateHit): string {
185
185
  return `${hit.where}: degenerate repetition — "${hit.token}" ×${hit.freq} (${share}% of ${hit.total} tokens, longest run ${hit.maxRun})`;
186
186
  }
187
187
 
188
+ // ---------------------------------------------------------------------------
189
+ // Intra-message BLOCK repetition (v1.7).
190
+ //
191
+ // The "narration loop" class (verified against a real coding session: ONE
192
+ // assistant message replaying ~5 near-verbatim cycles of "Let me start by
193
+ // checking the environment and the current state of the repository… / I'll run
194
+ // several independent checks in parallel. / Let me begin the S0 development…").
195
+ // Every pre-existing detector stayed silent: text/tool/thinking/structural all
196
+ // compare ACROSS messages (the model produced one message, no peer), and the
197
+ // degenerate detector only watches ONE word repeated hundreds of times, not a
198
+ // whole sentence/paragraph replayed. The signal here is phrase-level: a sliding
199
+ // window of blockNgram-word n-grams over the normalized payload; if almost ALL
200
+ // of those n-grams recur and the most frequent one recurs ≥ blockMinRepeats
201
+ // times, the generation is replaying itself instead of advancing.
202
+ //
203
+ // Deliberately conservative: only a payload where ≥ blockRepeatShare (default
204
+ // 85%) of 5-gram positions repeat qualifies. Ordinary prose — even long,
205
+ // structured docs — sits far below (README/pi.md paragraphs: ≤ 0.11), while
206
+ // 3+ replays of a narration block reach 0.98–1.0. Repeat-linked code/log lines
207
+ // that share a template are NOT flagged because their n-grams carry the varying
208
+ // digits and differ.
209
+ // ---------------------------------------------------------------------------
210
+
211
+ /** Recursive n-gram coverage of one payload. */
212
+ export interface BlockRepeatInfo {
213
+ /** Recurring n-gram positions / total n-gram positions (0..1). */
214
+ ratio: number;
215
+ /** Occurrences of the MOST repeated n-gram. */
216
+ repeats: number;
217
+ /** Total normalized words in the payload. */
218
+ tokens: number;
219
+ /** The n used (words per window). */
220
+ ngram: number;
221
+ /** The most repeated n-gram, for the description. */
222
+ sample: string;
223
+ }
224
+
225
+ export interface BlockRepeatHit extends BlockRepeatInfo {
226
+ where: string;
227
+ }
228
+
229
+ /**
230
+ * Detect phrase/block-level self-repetition inside ONE payload. Returns the
231
+ * coverage ratio with the most repeated n-gram, or undefined for normal text.
232
+ *
233
+ * Reads the payload as a stream of lowercase alphanumeric words (code symbols
234
+ * and punctuation are separators, so stored-JSON escaping cannot glue tokens).
235
+ * Coverage counts an n-gram position as repeated when the SAME n-word sequence
236
+ * (digits included, so templated lines with varying numbers do not count)
237
+ * occurs somewhere else in the payload.
238
+ */
239
+ export function findRepetitiveBlock(text: string, config: AntiloopConfig): BlockRepeatInfo | undefined {
240
+ const words = String(text).toLowerCase().match(/[\p{L}\p{N}]+/gu);
241
+ if (!words) return undefined;
242
+ const total = words.length;
243
+ if (total < config.blockMinTokens) return undefined;
244
+ const n = Math.max(2, config.blockNgram);
245
+ const minRepeats = Math.max(2, config.blockMinRepeats);
246
+ if (total < n * minRepeats) return undefined;
247
+
248
+ const counts = new Map<string, number>();
249
+ const grams: string[] = [];
250
+ for (let i = 0; i + n <= total; i++) {
251
+ const g = words.slice(i, i + n).join(" ");
252
+ grams.push(g);
253
+ counts.set(g, (counts.get(g) ?? 0) + 1);
254
+ }
255
+ if (!grams.length) return undefined;
256
+
257
+ let covered = 0;
258
+ let topKey = "";
259
+ let top = 0;
260
+ for (const g of grams) {
261
+ const c = counts.get(g)!;
262
+ if (c > top) {
263
+ top = c;
264
+ topKey = g;
265
+ }
266
+ if (c >= 2) covered++;
267
+ }
268
+ // At least one n-word phrase must recur blockMinRepeats times (a single
269
+ // echo is a restatement, not a replay) and the recurrence must dominate.
270
+ if (top < minRepeats) return undefined;
271
+ const ratio = covered / grams.length;
272
+ if (ratio < config.blockRepeatShare) return undefined;
273
+ return { ratio, repeats: top, tokens: total, ngram: n, sample: topKey };
274
+ }
275
+
276
+ /** Scan one assistant message (text + each tool-call argument) for a replay. */
277
+ export function scanMessageBlock(
278
+ content: string,
279
+ toolCalls: TrackedToolCall[] | undefined,
280
+ config: AntiloopConfig,
281
+ ): BlockRepeatHit | undefined {
282
+ if (!config.detectBlockRepeats) return undefined;
283
+ if (content && content.length) {
284
+ const b = findRepetitiveBlock(content, config);
285
+ if (b) return { ...b, where: "message text" };
286
+ }
287
+ for (const tc of toolCalls ?? []) {
288
+ if (!tc.args) continue;
289
+ const b = findRepetitiveBlock(tc.args, config);
290
+ if (b) return { ...b, where: tc.name === "bash" ? "bash command" : `args(${tc.name})` };
291
+ }
292
+ return undefined;
293
+ }
294
+
295
+ export function blockRepeatDescription(hit: BlockRepeatHit): string {
296
+ const pct = Math.round(hit.ratio * 100);
297
+ return `${hit.where}: block repetition — ${pct}% of ${hit.tokens} tokens replay repeated ${hit.ngram}-word phrases ("${hit.sample}" ×${hit.repeats})`;
298
+ }
299
+
188
300
  /** Consecutive-detection weight of a detection turn. The strong
189
301
  * self-contained signals (degenerate meltdown, proven no-progress outcome run)
190
302
  * add degenerateTurnWeight (default 2) points so the FIRST one already reaches
191
303
  * the warning level and escalation is fast on repeat. */
192
304
  export function detectionTurnWeight(detections: LoopDetection[], config: AntiloopConfig): number {
193
- const strong = detections.some((d) => d.type === "degenerate" || d.type === "outcome");
305
+ const strong = detections.some((d) => d.type === "degenerate" || d.type === "block" || d.type === "outcome");
194
306
  return strong ? Math.max(1, config.degenerateTurnWeight) : 1;
195
307
  }
196
308
 
@@ -383,6 +495,28 @@ export function detectTaskStreams(
383
495
  return streams;
384
496
  }
385
497
 
498
+ // ---------------------------------------------------------------------------
499
+ // Snapshot / read-only tools (v1.8).
500
+ //
501
+ // A coordinator closing N agents that already finished calls the snapshot tool
502
+ // once per agent and gets the SAME serialized board every time (all settled).
503
+ // Same args + same result looks exactly like a verbatim tool loop to the
504
+ // detector, but re-reading a state snapshot is an idempotent read — the serial
505
+ // close of finished work, not a stuck generation. Tools listed in
506
+ // config.snapshotTools (default: trimegisto_harvest) are therefore exempt from
507
+ // tool-loop and outcome detection. Any other status/poll tool can be added.
508
+ // ---------------------------------------------------------------------------
509
+
510
+ /** Set of tool names whose repeated identical reads must never count as a loop. */
511
+ function snapshotToolSet(config: AntiloopConfig): Set<string> {
512
+ return new Set((config.snapshotTools ?? []).map((n) => n.trim()).filter(Boolean));
513
+ }
514
+
515
+ /** True when every call in the set is a snapshot/status read. */
516
+ function allSnapshotCalls(calls: TrackedToolCall[] | undefined, snapshots: Set<string>): boolean {
517
+ return !!calls && calls.length > 0 && calls.every((c) => snapshots.has(c.name));
518
+ }
519
+
386
520
  export function detectLoops(state: AntiloopState, config: AntiloopConfig): LoopDetection[] {
387
521
  const out: LoopDetection[] = [];
388
522
  const msgs = state.recentMessages;
@@ -391,9 +525,10 @@ export function detectLoops(state: AntiloopState, config: AntiloopConfig): LoopD
391
525
  const win = msgs.slice(start);
392
526
  const now = Date.now();
393
527
 
394
- // Intra-message degenerate repetition fires on a SINGLE pathological message
395
- // (no peer needed) and is independent of the task-stream batch gate: a
396
- // meltdown is a meltdown even mid-batch.
528
+ // Intra-message repetition fires on a SINGLE pathological message (no peer
529
+ // needed) and is independent of the task-stream batch gate: a meltdown is a
530
+ // meltdown even mid-batch. Two flavors: degenerate (one word ×hundreds) and
531
+ // block (whole sentences/phrases replayed — the narration-loop class).
397
532
  if (config.detectDegenerate) {
398
533
  const last = win[win.length - 1];
399
534
  const hit = scanMessageDegenerate(last.content, last.toolCalls, config);
@@ -408,12 +543,29 @@ export function detectLoops(state: AntiloopState, config: AntiloopConfig): LoopD
408
543
  }
409
544
  }
410
545
 
546
+ if (config.detectBlockRepeats) {
547
+ const last = win[win.length - 1];
548
+ const hit = scanMessageBlock(last.content, last.toolCalls, config);
549
+ if (hit) {
550
+ out.push({
551
+ type: "block",
552
+ // 0.99 => isVerbatimRepeat(): after a force break, a replayed block
553
+ // is the model ignoring the break and escalates to the hard stop.
554
+ similarity: 0.99,
555
+ messageIndices: [msgs.length - 1],
556
+ description: blockRepeatDescription(hit),
557
+ timestamp: now,
558
+ });
559
+ }
560
+ }
561
+
411
562
  // Task-stream gate: if the window is a homogeneous batch (same extension
412
563
  // tool called with DISTINCT content ≥ taskStreamMinCalls times), that tool
413
564
  // is exempt from tool-loop detection, and text/thinking/structural
414
565
  // detections that only involve batch messages are suppressed too — the
415
566
  // model is doing N different tasks of the same type, not looping.
416
567
  const streams = detectTaskStreams(win, config);
568
+ const snapshots = snapshotToolSet(config);
417
569
  const batchAt = new Set<number>();
418
570
  win.forEach((m, idx) => {
419
571
  const calls = m.toolCalls;
@@ -467,7 +619,10 @@ export function detectLoops(state: AntiloopState, config: AntiloopConfig): LoopD
467
619
  const lastCalls = last.toolCalls;
468
620
  // A message whose calls are all task-stream tools is batch work — skip
469
621
  // it entirely (the stream gate already proved the calls are distinct).
470
- if (lastCalls && lastCalls.length && !batchAt.has(msgs.length - 1)) {
622
+ // Snapshot/status tools (trimegisto_harvest…) are also skipped: repeated
623
+ // identical reads of a settled board are the serial close of finished
624
+ // agents, not a loop.
625
+ if (lastCalls && lastCalls.length && !batchAt.has(msgs.length - 1) && !allSnapshotCalls(lastCalls, snapshots)) {
471
626
  const matched: number[] = [];
472
627
  for (let i = 0; i < win.length - 1; i++) {
473
628
  const prev = win[i].toolCalls;
@@ -525,7 +680,9 @@ export function detectLoops(state: AntiloopState, config: AntiloopConfig): LoopD
525
680
  // produce the SAME OK outcome every call — they must never count. And
526
681
  // batch messages must NOT be skipped here: a "bash ×N stream" with
527
682
  // identical failures is exactly the no-progress loop to catch.
528
- if (lc.result && isFailResult(lc.result)) {
683
+ // Snapshot/status reads are excluded outright: an identical harvest
684
+ // snapshot is a state read, not a repeated failing experiment.
685
+ if (lc.result && isFailResult(lc.result) && !snapshots.has(lc.name)) {
529
686
  let matches = 0;
530
687
  for (let i = 0; i < win.length - 1; i++) {
531
688
  const m = win[i];
@@ -673,8 +830,10 @@ export function runSelfTest(): string[] {
673
830
  detectTextLoops: true, notifyOnDetection: true, maxHistoryEntries: 100,
674
831
  detectionWindow: 10, interactiveFooter: true, toggleShortcut: "esc+a",
675
832
  detectTaskStreams: true, taskStreamMinCalls: 3, taskStreamTwinThreshold: 0.99,
833
+ snapshotTools: ["trimegisto_harvest"],
676
834
  detectDegenerate: true, degenerateMinTokens: 50, degenerateMaxRun: 16,
677
835
  degenerateMaxFreq: 60, degenerateMaxShare: 0.4, degenerateTurnWeight: 2, blockDegenerateBash: true,
836
+ detectBlockRepeats: true, blockMinTokens: 120, blockNgram: 5, blockMinRepeats: 3, blockRepeatShare: 0.85,
678
837
  detectOutcomeLoops: true, outcomeMinRepeats: 8, outcomeArgSimilarity: 0.85, outcomeSigThreshold: 0.7,
679
838
  };
680
839
  const asState = (recentMessages: TrackedMessage[]): AntiloopState =>
@@ -782,6 +941,56 @@ export function runSelfTest(): string[] {
782
941
  const wTxt = detectionTurnWeight(dl("text", 0.8), tcfg);
783
942
  out.push(`degenerate turn weight → degenerate ${wDeg}, text ${wTxt} (exp 2, 1) ${wDeg === 2 && wTxt === 1 ? "✅" : "❌"}`);
784
943
 
944
+ // --- v1.7: intra-message BLOCK repetition (narration loops) ---
945
+ // Regression: a real coding session where ONE assistant message replayed the
946
+ // same ~5 sentences in a loop ("Let me start by checking the environment…" /
947
+ // "I'll run several independent checks in parallel." / "Let me begin the S0
948
+ // development…"). Cross-message text/tool/thinking detectors need a peer and
949
+ // stayed silent; the degenerate scan only watches a single word. The block
950
+ // detector must fire on the message itself.
951
+ const NARRV = [
952
+ "Let me start by checking the environment and the current state of the repository, then set up a plan for the S0 slice and begin building.",
953
+ "I'll run several independent checks in parallel.",
954
+ "Let me begin the S0 development. First, reconnaissance of the environment and current repo state.",
955
+ "Let me check what's available in the environment (node, package managers, network, postgres) and the current repo state, then set up the S0 plan and start building the monorepo scaffold.",
956
+ "I'll run a batch of independent environment checks first.",
957
+ ];
958
+ const cycles = (k: number): string => {
959
+ let s = "";
960
+ for (let i = 0; i < k; i++) for (const v of NARRV) s += v + "\n\n";
961
+ return s;
962
+ };
963
+ const s0Det = detectLoops(asState([mk(cycles(5))]), tcfg);
964
+ const s0Hit = s0Det.find((d) => d.type === "block");
965
+ out.push(`block S0 narration loop → ${s0Hit ? `block (${s0Hit.description})` : "no"} (exp block — was the miss) ${s0Hit ? "✅" : "❌"}`);
966
+ out.push(`block turn weight → ${detectionTurnWeight(s0Det, tcfg)} (exp 2 — warn on first sight) ${detectionTurnWeight(s0Det, tcfg) === 2 ? "✅" : "❌"}`);
967
+ out.push(`block post-steer = ignored→ ${isVerbatimRepeat(s0Det) ? "yes" : "no"} (exp yes — replayed block after the break) ${isVerbatimRepeat(s0Det) ? "✅" : "❌"}`);
968
+
969
+ // 3 replay cycles still flag; a single restatement (2 cycles, top-gram ×2)
970
+ // does not — one echo is a summary, not stuck generation.
971
+ const c3 = findRepetitiveBlock(cycles(3), tcfg);
972
+ const c2 = findRepetitiveBlock(cycles(2), tcfg);
973
+ out.push(`block 3 cycles / 2 cycles→ ${c3 ? `flag (${Math.round(c3.ratio * 100)}%)` : "no"} / ${c2 ? "flag" : "no"} (exp flag / no — needs ≥3 repeats) ${c3 && !c2 ? "✅" : "❌"}`);
974
+
975
+ // Ordinary prose and templated (but evolving) payloads must stay silent:
976
+ // long docs score ≤ 0.11 coverage; templated logs/code carry varying digits
977
+ // so their 5-grams differ.
978
+ const prose =
979
+ "The extension watches every assistant message and tool call to decide whether the model is making progress. " +
980
+ "It stores a short window of recent turns, fingerprints tool results and compares them with earlier attempts. " +
981
+ "When a pattern repeats it escalates from a quiet warning to a forced change of approach and finally aborts. " +
982
+ "Configuration lives in a small json file next to the agent directory and every threshold can be tuned at runtime. " +
983
+ "The default values were chosen against real sessions so ordinary work never triggers a false positive. " +
984
+ "Detectors are independent: disabling one leaves the others active and the footer keeps the user informed. " +
985
+ "A clean turn cools the counter down so a recovered model is given room to finish the task. " +
986
+ "Everything is written in plain typescript with no runtime dependencies beyond the host package. " +
987
+ "New strategies should be measured against the recorded payloads before they are enabled by default. " +
988
+ "The goal is simple: catch a stuck model early and keep the context useful for the real work.";
989
+ const logLines = [...Array(30)].map((_, i) => `2026-09-24T10:${String(i).padStart(2, "0")}:00Z INFO worker ${i} processed job ${1000 + i} in ${i * 3}ms status ok`).join("\n");
990
+ const proseHit = findRepetitiveBlock(prose, tcfg);
991
+ const logHit = findRepetitiveBlock(logLines, tcfg);
992
+ out.push(`block legit prose/logs → ${proseHit || logHit ? "flag" : "ok"} (exp ok) ${!proseHit && !logHit ? "✅" : "❌"}`);
993
+
785
994
  // --- v1.6.1: no-progress outcome runs (mutated re-runs, same outcome) ---
786
995
  // Regression: the NFS session — ~90 mutated re-runs of the SAME experiment
787
996
  // (ssh exportfs/mount, labels test A…test QQQ, targets alternating Javi /
@@ -872,5 +1081,32 @@ export function runSelfTest(): string[] {
872
1081
  out.push(`outcome identical OKs → ${okDet.some((d) => d.type === "outcome") ? "outcome" : "silent"} (exp silent — success repeats ≠ loop) ${!okDet.some((d) => d.type === "outcome") ? "✅" : "❌"}`);
873
1082
  out.push(`isFailResult gate → err| → ${isFailResult("err|boom") ? "fail" : "ok"}, rc32 → ${isFailResult("ok|rc32 denied") ? "fail" : "ok"}, rc0/ok → ${isFailResult("ok|rc 0 12 passed") ? "fail" : "ok"} (exp fail, fail, ok) ${isFailResult("err|boom") && isFailResult("ok|rc32 denied") && !isFailResult("ok|rc 0 12 passed") ? "✅" : "❌"}`);
874
1083
 
1084
+ // --- v1.8: snapshot/read-only tools are idempotent reads, not loops ---
1085
+ // Regression (real session /home/j 2026-09-23T21:35): a coordinator closed
1086
+ // N agents that had ALREADY settled by calling trimegisto_harvest five times
1087
+ // with `{}`; every call returned the SAME 2.3 KB cumulative snapshot (all
1088
+ // agents done). Same args + same result ×5 → the tool detector force-broke
1089
+ // the run. Re-reading a settled state board is the serial close of finished
1090
+ // agents, not a stuck generation: with the snapshot exemption it is silent.
1091
+ const harvestSnap = resultFingerprint(
1092
+ [{
1093
+ type: "text",
1094
+ text:
1095
+ "## Trimegisto harvest (instant snapshot)\n\n### ✅ t0a [Active] — done (101s)\nTask: Explora el sistema en busca de LM Studio y sus runtimes.\n\n```\n# Informe: LM Studio en el sistema\n...\n```\n*10 turns, ↑32825 ↓2429*\n\n_All agents settled._",
1096
+ }],
1097
+ false,
1098
+ )!;
1099
+ const harvestMsgs = [0, 1, 2, 3, 4].map(() =>
1100
+ mk("", [{ name: "trimegisto_harvest", args: "{}", result: harvestSnap }]),
1101
+ );
1102
+ const snapDet = detectLoops(asState(harvestMsgs), tcfg);
1103
+ out.push(`snapshot harvest ×5 → ${snapDet.length ? snapDet.map((d) => d.type).join(",") : "silent"} (exp silent — serial close of settled agents) ${snapDet.length === 0 ? "✅" : "❌"}`);
1104
+
1105
+ // Guard: the exemption is what silences it — the SAME payload with
1106
+ // snapshotTools off is a genuine verbatim tool loop.
1107
+ const noSnapCfg: AntiloopConfig = { ...tcfg, snapshotTools: [] };
1108
+ const snapLoud = detectLoops(asState(harvestMsgs), noSnapCfg);
1109
+ out.push(`snapshot off = loop → ${snapLoud.some((d) => d.type === "tool") ? "tool" : "no"} (exp tool — normal tools still loop) ${snapLoud.some((d) => d.type === "tool") ? "✅" : "❌"}`);
1110
+
875
1111
  return out;
876
1112
  }
package/src/index.ts CHANGED
@@ -24,6 +24,13 @@
24
24
  * hook (blockDegenerateBash). One degenerate turn counts degenerateTurnWeight
25
25
  * (2) consecutive points: warning on first sight, force-break steer on the
26
26
  * second consecutive meltdown, hard stop shortly after if it keeps repeating.
27
+ *
28
+ * v1.7 adds the intra-message BLOCK detector for the narration-loop class:
29
+ * ONE message that replays whole sentences/phrases ("Let me start by checking
30
+ * the environment…" ×5), which every cross-message detector misses because
31
+ * there is no peer message and the degenerate scan only watches single words.
32
+ * Same message_end timing and strong turn weight; the bash gate refuses a
33
+ * command that is itself a replayed block.
27
34
  */
28
35
 
29
36
  import type { ExtensionAPI, ExtensionContext } from "@earendil-works/pi-coding-agent";
@@ -285,29 +292,55 @@ export default function antiloopExtension(pi: ExtensionAPI) {
285
292
  state.recentMessages = state.recentMessages.slice(-(config.detectionWindow + 5));
286
293
  }
287
294
 
288
- // v1.6 — degenerate meltdowns are handled HERE, at message_end, because
289
- // turn_end only fires after tool execution (too late to stop the 46 KB
290
- // "noguerol ×5000" bash from running) and an aborted generation may never
291
- // reach turn_end at all. The signal is self-contained (one pathological
292
- // payload, no peer message) so the full escalation ladder runs right now;
293
- // the turn-guard makes the later turn_end pass skip this message and no
294
- // detection is double-counted.
295
- if (!config.detectDegenerate) return;
296
- const { scanMessageDegenerate, degenerateDescription, nextLevel, isVerbatimRepeat } = await import("./detect.ts");
297
- const hit = scanMessageDegenerate(content, toolCalls, config);
298
- if (!hit) return;
295
+ // v1.6/v1.7 — intra-message repetition is handled HERE, at message_end,
296
+ // because turn_end only fires after tool execution (too late to stop the
297
+ // 46 KB "noguerol ×5000" bash from running) and an aborted generation may
298
+ // never reach turn_end at all. The signal is self-contained (one
299
+ // pathological payload, no peer message) so the full escalation ladder runs
300
+ // right now; the turn-guard makes the later turn_end pass skip this message
301
+ // and no detection is double-counted. Two flavors: degenerate (one word
302
+ // ×hundreds) and block (whole sentences/phrases replayed — the narration
303
+ // loop class where every cross-message detector is blind).
304
+ if (!config.detectDegenerate && !config.detectBlockRepeats) return;
305
+ const {
306
+ scanMessageDegenerate,
307
+ degenerateDescription,
308
+ scanMessageBlock,
309
+ blockRepeatDescription,
310
+ nextLevel,
311
+ isVerbatimRepeat,
312
+ } = await import("./detect.ts");
313
+ let det: LoopDetection | undefined;
314
+ if (config.detectDegenerate) {
315
+ const degHit = scanMessageDegenerate(content, toolCalls, config);
316
+ if (degHit) {
317
+ det = {
318
+ type: "degenerate",
319
+ similarity: 1,
320
+ messageIndices: [state.recentMessages.length - 1],
321
+ description: degenerateDescription(degHit),
322
+ timestamp: Date.now(),
323
+ };
324
+ }
325
+ }
326
+ if (!det && config.detectBlockRepeats) {
327
+ const blkHit = scanMessageBlock(content, toolCalls, config);
328
+ if (blkHit) {
329
+ det = {
330
+ type: "block",
331
+ similarity: 0.99,
332
+ messageIndices: [state.recentMessages.length - 1],
333
+ description: blockRepeatDescription(blkHit),
334
+ timestamp: Date.now(),
335
+ };
336
+ }
337
+ }
338
+ if (!det) return;
299
339
  const tracked = state.recentMessages[state.recentMessages.length - 1];
300
340
  if (tracked) state.lastDetectedTurnIndex = tracked.turnIndex;
301
341
  const prevLevel = state.currentLevel;
302
342
  state.consecutiveDetections += Math.max(1, config.degenerateTurnWeight);
303
343
  state.totalDetections++;
304
- const det: LoopDetection = {
305
- type: "degenerate",
306
- similarity: 1,
307
- messageIndices: [state.recentMessages.length - 1],
308
- description: degenerateDescription(hit),
309
- timestamp: Date.now(),
310
- };
311
344
  state.detections.push(det);
312
345
  if (state.detections.length > config.maxHistoryEntries) {
313
346
  state.detections = state.detections.slice(-config.maxHistoryEntries);
@@ -331,13 +364,13 @@ export default function antiloopExtension(pi: ExtensionAPI) {
331
364
  ctx.ui.notify(`antiloop: force break — ${det.description}`, "error");
332
365
  }
333
366
  } else if (isVerbatimRepeat([det])) {
334
- // The model produced degenerate output AGAIN after the break message:
335
- // count it; once the ignore limit is hit the run is cut for good.
367
+ // The model produced the same intra-message repetition AGAIN after the
368
+ // break message: count it; once the ignore limit is hit, cut the run.
336
369
  state.ignoredSteerCount++;
337
370
  if (state.ignoredSteerCount >= config.ignoredSteerLimit) {
338
371
  hardStop(
339
372
  ctx,
340
- `antiloop: abort — degenerate output repeated ${state.ignoredSteerCount}× after the force break — run stopped; provide new instructions`,
373
+ `antiloop: abort — repetitive output repeated ${state.ignoredSteerCount}× after the force break — run stopped; provide new instructions`,
341
374
  );
342
375
  }
343
376
  }
@@ -379,15 +412,28 @@ export default function antiloopExtension(pi: ExtensionAPI) {
379
412
  * sees why the command was refused and can change approach.
380
413
  */
381
414
  pi.on("tool_call", async (event, ctx) => {
382
- if (!config.enabled || !config.detectDegenerate || !config.blockDegenerateBash) return;
415
+ if (!config.enabled || !config.blockDegenerateBash) return;
383
416
  if (!isToolCallEventType("bash", event)) return;
384
- const { findDegenerateRepetition, degenerateDescription } = await import("./detect.ts");
385
- const hit = findDegenerateRepetition(event.input.command ?? "", config);
386
- if (!hit) return;
387
- return {
388
- block: true,
389
- reason: `[antiloop] blocked: ${degenerateDescription({ ...hit, where: "bash command" })} — this is stuck generation, not a real command. Do NOT retry it: stop and take one small, concrete step instead.`,
390
- };
417
+ const command = event.input.command ?? "";
418
+ const { findDegenerateRepetition, findRepetitiveBlock, degenerateDescription, blockRepeatDescription } = await import("./detect.ts");
419
+ if (config.detectDegenerate) {
420
+ const hit = findDegenerateRepetition(command, config);
421
+ if (hit) {
422
+ return {
423
+ block: true,
424
+ reason: `[antiloop] blocked: ${degenerateDescription({ ...hit, where: "bash command" })} — this is stuck generation, not a real command. Do NOT retry it: stop and take one small, concrete step instead.`,
425
+ };
426
+ }
427
+ }
428
+ if (config.detectBlockRepeats) {
429
+ const blk = findRepetitiveBlock(command, config);
430
+ if (blk) {
431
+ return {
432
+ block: true,
433
+ reason: `[antiloop] blocked: ${blockRepeatDescription({ ...blk, where: "bash command" })} — this is stuck generation, not a real command. Do NOT retry it: stop and take one small, concrete step instead.`,
434
+ };
435
+ }
436
+ }
391
437
  });
392
438
 
393
439
  pi.on("turn_end", async (event, ctx) => {
package/src/types.ts CHANGED
@@ -34,8 +34,30 @@ export interface AntiloopConfig {
34
34
  degenerateMaxShare: number;
35
35
  /** How many consecutive-detection points ONE degenerate turn adds (2 = warn on first sight). */
36
36
  degenerateTurnWeight: number;
37
- /** Block a bash tool call whose command shows degenerate repetition BEFORE it executes. */
37
+ /** Block a bash tool call whose command shows degenerate OR block repetition BEFORE it executes. */
38
38
  blockDegenerateBash: boolean;
39
+ /**
40
+ * Intra-message BLOCK repetition (v1.7): the "narration loop" class where ONE
41
+ * assistant message keeps replaying the same sentences/phrases (the real S0
42
+ * coding session: ~5 near-verbatim cycles of "Let me start by checking the
43
+ * environment…" inside a SINGLE generation, which every cross-message detector
44
+ * misses because it produced only one message and the degenerate detector only
45
+ * watches for ONE word repeated hundreds of times). A sliding window of
46
+ * blockNgram-word n-grams over the normalized payload is built; when at least
47
+ * blockRepeatShare of those positions recur, one n-gram appears at least
48
+ * blockMinRepeats times and the payload has at least blockMinTokens words, the
49
+ * message is a replay. Fires at message_end (before the message's tools run)
50
+ * with the same strong turn weight as the degenerate detector.
51
+ */
52
+ detectBlockRepeats: boolean;
53
+ /** Minimum normalized words in a payload before it is scanned for block repetition. */
54
+ blockMinTokens: number;
55
+ /** Word window of the n-grams whose recurrence is measured (default 5). */
56
+ blockNgram: number;
57
+ /** The MOST repeated n-gram must occur at least this many times to flag. */
58
+ blockMinRepeats: number;
59
+ /** Share of n-gram positions that must recur for the payload to count as a replay. */
60
+ blockRepeatShare: number;
39
61
  /**
40
62
  * No-progress outcome runs (v1.6.1): a model that keeps re-running the SAME
41
63
  * experiment with cosmetic mutations (labels/permutations) while the outcome
@@ -87,6 +109,18 @@ export interface AntiloopConfig {
87
109
  detectTaskStreams: boolean;
88
110
  taskStreamMinCalls: number;
89
111
  taskStreamTwinThreshold: number;
112
+ /**
113
+ * Read-only snapshot/status tools (default: ["trimegisto_harvest"]). These
114
+ * return a view of changing state (a list of agents, a queue, a dashboard).
115
+ * Calling one again with the SAME arguments is an idempotent read — most
116
+ * strikingly when a coordinator closes N agents that already settled and the
117
+ * serialized snapshot is byte-identical every time. That is the serial close
118
+ * of finished agents, not a reasoning loop, so repeated calls to a snapshot
119
+ * tool never count as a tool loop (nor as a no-progress outcome run). Names
120
+ * are matched exactly against the tool name; add any other status/poll tool
121
+ * here to keep antiloop quiet while it is read repeatedly.
122
+ */
123
+ snapshotTools: string[];
90
124
  detectToolLoops: boolean;
91
125
  detectThinkingLoops: boolean;
92
126
  detectTextLoops: boolean;
@@ -106,7 +140,7 @@ export interface AntiloopConfig {
106
140
  toggleShortcut: string;
107
141
  }
108
142
 
109
- export type LoopKind = "text" | "tool" | "thinking" | "structural" | "degenerate" | "outcome";
143
+ export type LoopKind = "text" | "tool" | "thinking" | "structural" | "degenerate" | "block" | "outcome";
110
144
 
111
145
  export interface LoopDetection {
112
146
  type: LoopKind;