pi-antiloop 1.6.1 → 1.8.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +101 -12
- package/package.json +2 -2
- package/src/commands.ts +32 -3
- package/src/config.ts +6 -0
- package/src/detect.ts +242 -6
- package/src/index.ts +75 -29
- package/src/types.ts +36 -2
package/README.md
CHANGED
|
@@ -6,15 +6,16 @@
|
|
|
6
6
|
|
|
7
7
|
# Antiloop — Loop Detection and Break for pi
|
|
8
8
|
|
|
9
|
-
**Antiloop watches every assistant message, tool call and thinking block, and forces the model out of reasoning loops before they eat your context and your patience.**
|
|
9
|
+
**Antiloop watches every assistant message, tool call and thinking block, and forces the model out of reasoning loops before they eat your context and your patience.** Seven simultaneous detection strategies (text similarity, tool-call sequences, thinking content, structural openings, degenerate repetition, block/narration repetition, no-progress outcome runs) find loops that humans miss — and progressive intervention (warning → force break → abort) tells the model to take a different approach, without you having to babysit it.
|
|
10
10
|
|
|
11
11
|
---
|
|
12
12
|
|
|
13
13
|
## Features
|
|
14
14
|
|
|
15
|
-
- **
|
|
15
|
+
- **Seven detection strategies** — text repetition (trigram Jaccard + Levenshtein), tool-call sequences (name + near-identical arguments + same outcome — result-aware, so retries that make progress don't false-positive), thinking blocks, structural opening-phrase patterns, **degenerate repetition** (a single message/call stuck repeating one word hundreds of times — the `noguerol ×5145` meltdown — caught at `message_end` with no repeated peer needed, and degenerate bash commands blocked before they execute), **block repetition** (ONE message replaying whole sentences/phrases — the narration loop: `Let me start by checking the environment…` ×5 — same `message_end` timing and strong turn weight), and **no-progress outcome runs** (the NFS `test A…QQQ` class: dozens of near-identical re-runs of the same experiment, every one ending in the *same failing outcome* — args mutate so tool-loop can't see it; the repeated failure signature can)
|
|
16
16
|
- **Stays alive in long sessions** — tracked messages carry a monotonic sequence number, so trimming the sliding window can never make the dedupe guard skip later messages (a bug that silently blinded antiloop after ~15 tracked messages)
|
|
17
17
|
- **Task-stream recognition (batch work)** — when another extension (e.g. `punched` appending lines to pi.md, or `plan` adding tasks) makes the model call the *same* tool many times with *different* content, antiloop recognizes it as N distinct tasks of one type and stays silent — no warning, no force break. A genuine loop (the *same* call repeated verbatim) is still caught
|
|
18
|
+
- **Snapshot/read-only tool exemption** — a coordinator closing N agents that already finished calls `trimegisto_harvest` once per agent and gets the *same* settled board every time; identical reads of a status snapshot are an idempotent serial close, not a loop, so `snapshotTools` (configurable) is exempt from tool-loop and no-progress detection — real `bash` loops next to it are still caught
|
|
18
19
|
- **Progressive intervention** — `warning` reminds the model to vary its approach; `force break` steers a real break message into the running agent before its next LLM call; `abort` stops the run entirely
|
|
19
20
|
- **Configurable thresholds** — independent dials for similarity cutoff, warning/force-break/abort counts, detection window, and which strategies are on
|
|
20
21
|
- **Sliding window** — only the last N messages are compared, so detection is O(N) in the window size, not in the full session
|
|
@@ -100,12 +101,14 @@ Detection strategies:
|
|
|
100
101
|
Tool loops: ✅
|
|
101
102
|
Thinking loops: ✅
|
|
102
103
|
Degenerate: ✅
|
|
104
|
+
Block repetition: ✅
|
|
103
105
|
Outcome (no-progress): ✅
|
|
104
106
|
|
|
105
107
|
Recent detections:
|
|
106
108
|
[text] Text similarity 85% with message 3 (2m ago)
|
|
107
109
|
[tool] repeated 3x: bash (5m ago)
|
|
108
110
|
[degenerate] bash command: degenerate repetition — "noguerol" ×5145 (5m ago)
|
|
111
|
+
[block] message text: block repetition — 100% of 450 tokens replay repeated 5-word phrases (5m ago)
|
|
109
112
|
[outcome] no progress: 9 near-identical bash attempts with the same failing outcome (5m ago)
|
|
110
113
|
```
|
|
111
114
|
|
|
@@ -129,7 +132,9 @@ Grouped interactive menu showing the current value in each option:
|
|
|
129
132
|
- **🔁 call repeats** — `1 / 2 / 3` — how many times the same call must repeat before it flags (default 2)
|
|
130
133
|
- **🧾 result similarity** — `95 / 80 / 60%` — how similar captured results must be to count as the *same outcome*; a repeated command that starts producing a different result is progress, not a loop (default 80%)
|
|
131
134
|
- **🌀 degenerate run** — `8 / 16 / 24 / 48` — identical words in a row inside ONE message/call before it counts as stuck generation (default 16; the `noguerol ×5145` class)
|
|
132
|
-
-
|
|
135
|
+
- **🔁 block repetition** — on/off — detect ONE message replaying whole sentences/phrases (the narration loop: `Let me start by checking the environment…` ×5) with no repeated peer needed
|
|
136
|
+
- **🔁 block share** — `70 / 85 / 95%` — share of replayed 5-word phrases in one message before it counts as a replay (default 85%)
|
|
137
|
+
- **⛔ block repetitive bash** — on/off — refuse a degenerate or replayed bash command before it executes, feeding the reason back to the model (default on)
|
|
133
138
|
- **📉 no-progress after** — `4 / 6 / 8 / 12` — same failing outcome repeated this many times (near-identical args) before flagging (default 8; the NFS `test A…QQQ` class)
|
|
134
139
|
|
|
135
140
|
**📋 Task streams** — batch work (punched_log, plan_manager, …) is N tasks of one type, not a loop
|
|
@@ -137,11 +142,15 @@ Grouped interactive menu showing the current value in each option:
|
|
|
137
142
|
- **📋 stream min calls** — `2 / 3 / 4 / 5` — same-tool calls required before a batch is recognized (default 3)
|
|
138
143
|
- **📋 twin threshold** — `99 / 95 / 90%` — calls more similar than this count as the *same task* repeated; one twin invalidates the batch and normal detection resumes (default 99%)
|
|
139
144
|
|
|
145
|
+
**📡 Snapshot tools** — read-only status boards (trimegisto_harvest, …) read serially are not a loop
|
|
146
|
+
- **📡 snapshot tools** — `trimegisto_harvest` / off — repeated identical reads of a settled board stay silent all the way through (default on; add any other status/poll tool in `antiloop.json`)
|
|
147
|
+
|
|
140
148
|
**🔍 Detectors**
|
|
141
149
|
- **📝 text** — on/off — detect repeated text messages
|
|
142
150
|
- **🔧 tools** — on/off — detect repeated tool calls
|
|
143
151
|
- **🧠 thinking** — on/off — detect repeated internal reasoning
|
|
144
152
|
- **🌀 degenerate** — on/off — detect one message stuck repeating a single word/token (no repeated peer needed)
|
|
153
|
+
- **🔁 block** — on/off — detect one message replaying whole sentences/phrases (narration loop, no repeated peer needed)
|
|
145
154
|
- **📉 outcome** — on/off — detect many near-identical attempts all ending in the same failing outcome (no progress)
|
|
146
155
|
|
|
147
156
|
**🧹 reset state** — clear all counters and history
|
|
@@ -175,6 +184,9 @@ degenerate legit cmd → ok (exp ok) ✅
|
|
|
175
184
|
degenerate interleaved → flag (noguerol ×150) (exp flag — freq/share clause) ✅
|
|
176
185
|
degenerate glued token → flag (noguerol ×300) (exp flag — perfect power) ✅
|
|
177
186
|
degenerate turn weight → degenerate 2, text 1 (exp 2, 1) ✅
|
|
187
|
+
block S0 narration loop → block (message text: 100% of 450 tokens replay repeated 5-word phrases) ✅ ← v1.7: ONE message replaying sentences, no peer
|
|
188
|
+
block 3 cycles / 2 cycles→ flag (100%) / no (exp flag / no — needs ≥3 repeats) ✅
|
|
189
|
+
block legit prose/logs → ok (exp ok) ✅
|
|
178
190
|
outcome fires on 9th → outcome (9 near-identical bash attempts, same failing outcome) ✅ ← v1.6.1: NFS test-A…QQQ class
|
|
179
191
|
outcome needs 8 prior → silent (exp silent at 5 attempts) ✅
|
|
180
192
|
outcome converging sweep → silent (exp silent — outcomes differ = progress) ✅
|
|
@@ -186,7 +198,7 @@ outcome identical OKs → silent (exp silent — success repeats ≠ loop)
|
|
|
186
198
|
|
|
187
199
|
### Detection pipeline
|
|
188
200
|
|
|
189
|
-
After every assistant `message_end` event, antiloop extracts the new content (text, thinking, tool calls — including their ids) and pushes it onto a sliding window of the last `detectionWindow + 5` messages. Detection itself runs at `turn_end`, once the tool results are known: results are fingerprinted and attached to the tracked calls, then the active detection strategies run against the window. The
|
|
201
|
+
After every assistant `message_end` event, antiloop extracts the new content (text, thinking, tool calls — including their ids) and pushes it onto a sliding window of the last `detectionWindow + 5` messages. Detection itself runs at `turn_end`, once the tool results are known: results are fingerprinted and attached to the tracked calls, then the active detection strategies run against the window. The exceptions are the **intra-message** strategies — **degenerate** and **block**. A single stuck message needs no peer and no tool result, so both are evaluated right at `message_end` — the only point before the message's own tool calls execute — and repetitive `bash` calls are also blocked at the `tool_call` hook.
|
|
190
202
|
|
|
191
203
|
| Strategy | What it compares | Algorithm |
|
|
192
204
|
|----------|------------------|-----------|
|
|
@@ -195,10 +207,11 @@ After every assistant `message_end` event, antiloop extracts the new content (te
|
|
|
195
207
|
| Thinking | Internal reasoning/thinking blocks | Same as text |
|
|
196
208
|
| Structural | First 10 words of each message | Opening-phrase similarity ≥ 90% across ≥ 3 messages |
|
|
197
209
|
| Degenerate | One single message/call (no peer needed) | Run-length + frequency of identical words inside the payload: ≥ `degenerateMaxRun` (default 16) consecutive identical words, or one word ≥ `degenerateMaxFreq`× at ≥ `degenerateMaxShare` of all tokens; plus a perfect-power check for glued no-space tokens. Scanned at `message_end` — before the tool calls execute — and on every `bash` `tool_call` (blocking gate) |
|
|
210
|
+
| Block | One single message (no peer needed) | Sliding `blockNgram`-word n-grams over the normalized payload: a payload of ≥ `blockMinTokens` (120) words is a replay when ≥ `blockRepeatShare` (default 85%) of the n-gram positions recur AND the most repeated n-gram appears ≥ `blockMinRepeats` (3) times. Catches a message replaying whole sentences/phrases — the narration loop (`Let me start by checking the environment…` ×5) that every cross-message detector misses because there is no peer. Scanned at `message_end` and on every `bash` `tool_call` (blocking gate) |
|
|
198
211
|
| Outcome | Single tool calls across the window, after the last user input | ≥ `outcomeMinRepeats` (default 8) PRIOR attempts with args ≥ `outcomeArgSimilarity` (0.85) similar AND the same *failing* outcome (failure signatures compared at ≥ `outcomeSigThreshold`, 0.7; identical OK results never count — they're the norm for batches). Catches mutated re-run loops the tool detector can't see (labels/permutations change every turn) |
|
|
199
212
|
| Task stream | Same tool, many calls | When a tool appears ≥ `taskStreamMinCalls` times (default 3) in the window and *no two* calls are near-identical (`taskStreamTwinThreshold`, default 99%), the tool is an active batch: N different tasks of one type (e.g. `punched_log` appends, `plan_manager` task adds). Those calls are exempt from tool-loop detection, and text/thinking/structural patterns that only involve those batch messages are suppressed too. If even one call pair is a twin (the same task repeated), the tool is *not* a stream and detection proceeds normally |
|
|
200
213
|
|
|
201
|
-
Each detected pair becomes a `LoopDetection { type, similarity, messageIndices, description }` and the consecutive counter increases (a degenerate turn counts `degenerateTurnWeight`, default 2 — warning on first sight).
|
|
214
|
+
Each detected pair becomes a `LoopDetection { type, similarity, messageIndices, description }` and the consecutive counter increases (a degenerate or block turn counts `degenerateTurnWeight`, default 2 — warning on first sight).
|
|
202
215
|
|
|
203
216
|
### Intervention levels
|
|
204
217
|
|
|
@@ -211,7 +224,7 @@ Each detected pair becomes a `LoopDetection { type, similarity, messageIndices,
|
|
|
211
224
|
|
|
212
225
|
The level never de-escalates during an active loop; user input decays the consecutive counter naturally so a fresh prompt can break the cycle.
|
|
213
226
|
|
|
214
|
-
**Degenerate turns are handled earlier than the ladder:**
|
|
227
|
+
**Degenerate and block turns are handled earlier than the ladder:** they are detected at `message_end` (the same moment the text/tool-call payload is complete, *before* pi preflights and executes its tools). A single such turn already adds `degenerateTurnWeight` (2) consecutive points → warning on first sight; the second consecutive one → force break steer; after the steer, further degenerate/blocked output counts against `ignoredSteerLimit` → hard stop. Repetitive `bash` calls are additionally refused by the `tool_call` gate (`blockDegenerateBash`) — the command never runs, and the block reason is fed back to the model as the tool error so it can still change approach.
|
|
215
228
|
|
|
216
229
|
### Similarity scoring
|
|
217
230
|
|
|
@@ -252,6 +265,7 @@ If the same command produced a *different* outcome, the pair is progress:
|
|
|
252
265
|
|
|
253
266
|
Results only veto; they never trigger on their own, and calls without a
|
|
254
267
|
captured result fall back to argument matching alone.
|
|
268
|
+
```
|
|
255
269
|
|
|
256
270
|
### Task streams: N tasks of one type ≠ a loop
|
|
257
271
|
|
|
@@ -287,6 +301,27 @@ size needed before recognition), `taskStreamTwinThreshold` (how similar args
|
|
|
287
301
|
must be to count as *the same task* — lower it to treat near-duplicate
|
|
288
302
|
entries as loops again).
|
|
289
303
|
|
|
304
|
+
### Snapshot tools: the serial close of settled agents ≠ a loop
|
|
305
|
+
|
|
306
|
+
A coordinator that closes N agents which have already finished calls the
|
|
307
|
+
snapshot tool once per agent. When the agents are all settled, every one of
|
|
308
|
+
those calls returns the **same** serialized board — same arguments (`{}`),
|
|
309
|
+
same result. To the tool detector that is indistinguishable from a verbatim
|
|
310
|
+
loop: `same args + same outcome` repeated N times, so it force-broke a run
|
|
311
|
+
that was simply collecting finished results.
|
|
312
|
+
|
|
313
|
+
Re-reading a state snapshot is an *idempotent read*, not a stuck generation,
|
|
314
|
+
so tools listed in `snapshotTools` (default `["trimegisto_harvest"]`) are
|
|
315
|
+
exempt from tool-loop **and** no-progress outcome detection. The exemption is
|
|
316
|
+
deliberately narrow:
|
|
317
|
+
|
|
318
|
+
- it only applies when *every* call in the message is a snapshot tool — a
|
|
319
|
+
real `bash` loop re-run alongside a harvest is still flagged;
|
|
320
|
+
- it is name-based and configurable, so `antiloop.json` can list any other
|
|
321
|
+
status/poll tool (`"snapshotTools": ["trimegisto_harvest", "my_status"]`);
|
|
322
|
+
- clearing it (`"snapshotTools": []` or **📡 off**) restores normal detection,
|
|
323
|
+
where the very same 5× harvest payload is a genuine verbatim tool loop.
|
|
324
|
+
|
|
290
325
|
### Degenerate repetition: ONE message stuck on a word ≠ a reasoning loop
|
|
291
326
|
|
|
292
327
|
All strategies above need at least two similar messages — they detect a model
|
|
@@ -333,6 +368,49 @@ sweeps that motivated the tool-loop threshold). Only `bash` gets the blocking
|
|
|
333
368
|
gate — writing a repetitive *file* (e.g. a user-requested padding fixture) is
|
|
334
369
|
alerted and escalated, not refused.
|
|
335
370
|
|
|
371
|
+
### Block repetition: ONE message replaying its own sentences
|
|
372
|
+
|
|
373
|
+
The degenerate detector catches a single *word* repeated hundreds of times.
|
|
374
|
+
There is a second intra-message failure mode, just as obvious to a human and
|
|
375
|
+
invisible to every cross-message detector: **the model replays whole
|
|
376
|
+
sentences/phrases inside one generation**. Real case (a coding session's S0
|
|
377
|
+
scaffold): ONE assistant message cycled ~5 times through
|
|
378
|
+
|
|
379
|
+
> Let me start by checking the environment and the current state of the
|
|
380
|
+
> repository, then set up a plan for the S0 slice and begin building.
|
|
381
|
+
> I'll run several independent checks in parallel.
|
|
382
|
+
> Let me begin the S0 development. First, reconnaissance of the environment
|
|
383
|
+
> and current repo state.
|
|
384
|
+
|
|
385
|
+
…never emitting a tool call. Text/tool/thinking/structural detection all
|
|
386
|
+
compare *across* messages and had nothing to compare against; the degenerate
|
|
387
|
+
scan only watches for one word repeated, and no single word repeated 16×.
|
|
388
|
+
|
|
389
|
+
Antiloop v1.7 detects **block repetition** on the message itself:
|
|
390
|
+
|
|
391
|
+
- the payload is read as a stream of lowercase alphanumeric words;
|
|
392
|
+
- a sliding window of `blockNgram` (default 5) words is built over it;
|
|
393
|
+
- if at least `blockRepeatShare` (default **85%**) of those n-gram positions
|
|
394
|
+
recur somewhere else, the most repeated n-gram appears at least
|
|
395
|
+
`blockMinRepeats` (default 3) times, and the payload has at least
|
|
396
|
+
`blockMinTokens` (default 120) words, the generation is replaying itself.
|
|
397
|
+
|
|
398
|
+
The thresholds are deliberately far apart from normal content: ordinary
|
|
399
|
+
prose — even long, structured documents — scores ≤ 11% coverage, while 3+
|
|
400
|
+
replays of a narration block reach 98–100%. Templated payloads with evolving
|
|
401
|
+
data (log lines, generated code) are not flagged because their n-grams carry
|
|
402
|
+
the varying digits and differ. A single restatement (2 cycles) does not trip
|
|
403
|
+
it — one echo is a summary, three are stuck generation.
|
|
404
|
+
|
|
405
|
+
Because the signal is conclusive it is handled like the degenerate one:
|
|
406
|
+
caught at `message_end` (before the message's tools execute), weighted
|
|
407
|
+
`degenerateTurnWeight` (2) so the first occurrence warns, and a replayed
|
|
408
|
+
`bash` command is refused by the same `blockDegenerateBash` gate. After a
|
|
409
|
+
force-break steer, another replayed block escalates to the hard stop.
|
|
410
|
+
|
|
411
|
+
Tunables: `detectBlockRepeats`, `blockRepeatShare` (lower = earlier),
|
|
412
|
+
`blockMinRepeats`, `blockNgram`, `blockMinTokens`.
|
|
413
|
+
|
|
336
414
|
### No-progress outcome runs: mutated re-runs, same wall
|
|
337
415
|
|
|
338
416
|
A subtler meltdown than the degenerate one: the model re-runs the SAME
|
|
@@ -372,7 +450,6 @@ real session this fires at `test MM` (warn) → `NN` (steer) → `PP` (abort)
|
|
|
372
450
|
|
|
373
451
|
Tunables: `detectOutcomeLoops`, `outcomeMinRepeats` (lower = earlier cutoff),
|
|
374
452
|
`outcomeArgSimilarity`, `outcomeSigThreshold`.
|
|
375
|
-
```
|
|
376
453
|
|
|
377
454
|
### Sliding window
|
|
378
455
|
|
|
@@ -399,6 +476,11 @@ Persisted as JSON at `~/.pi/agent/antiloop.json`:
|
|
|
399
476
|
"degenerateMaxShare": 0.4,
|
|
400
477
|
"degenerateTurnWeight": 2,
|
|
401
478
|
"blockDegenerateBash": true,
|
|
479
|
+
"detectBlockRepeats": true,
|
|
480
|
+
"blockMinTokens": 120,
|
|
481
|
+
"blockNgram": 5,
|
|
482
|
+
"blockMinRepeats": 3,
|
|
483
|
+
"blockRepeatShare": 0.85,
|
|
402
484
|
"detectOutcomeLoops": true,
|
|
403
485
|
"outcomeMinRepeats": 8,
|
|
404
486
|
"outcomeArgSimilarity": 0.85,
|
|
@@ -433,7 +515,12 @@ Persisted as JSON at `~/.pi/agent/antiloop.json`:
|
|
|
433
515
|
| `degenerateMaxFreq` | `60` | One word's total occurrences (with `degenerateMaxShare` of the payload) that flags interleaved meltdowns |
|
|
434
516
|
| `degenerateMaxShare` | `0.4` | Frequency share (freq/total tokens) required together with `degenerateMaxFreq` |
|
|
435
517
|
| `degenerateTurnWeight` | `2` | Consecutive-detection points added by one degenerate turn (2 = warning on first sight) |
|
|
436
|
-
| `blockDegenerateBash` | `true` | Block a degenerate `bash` command in the `tool_call` hook before it executes; the reason is fed back to the model as the tool error |
|
|
518
|
+
| `blockDegenerateBash` | `true` | Block a degenerate or block-repetition `bash` command in the `tool_call` hook before it executes; the reason is fed back to the model as the tool error |
|
|
519
|
+
| `detectBlockRepeats` | `true` | Detect intra-message block/narration repetition — ONE message replaying whole sentences/phrases (the `Let me start by checking the environment…` ×5 class). Needs no repeated peer; scanned at `message_end` and on every bash `tool_call` |
|
|
520
|
+
| `blockMinTokens` | `120` | Minimum normalized words in a payload before it is scanned for block repetition (shorter payloads aren't conclusive) |
|
|
521
|
+
| `blockNgram` | `5` | Word window of the n-grams whose recurrence is measured |
|
|
522
|
+
| `blockMinRepeats` | `3` | The most repeated n-gram must occur at least this many times to flag |
|
|
523
|
+
| `blockRepeatShare` | `0.85` | Share of n-gram positions that must recur for the payload to count as a replay |
|
|
437
524
|
| `detectOutcomeLoops` | `true` | No-progress outcome runs: ≥ `outcomeMinRepeats` near-identical attempts (args ≥ `outcomeArgSimilarity`) all ending in the *same failing outcome* — the NFS `test A…QQQ` class |
|
|
438
525
|
| `outcomeMinRepeats` | `8` | Prior same-failure attempts (inside the window, after the last user input) required before the outcome detector fires |
|
|
439
526
|
| `outcomeArgSimilarity` | `0.85` | How similar args must be to count as the *same experiment reshuffled* (mutations of labels/permutations stay under it — distinct tasks don't) |
|
|
@@ -441,6 +528,7 @@ Persisted as JSON at `~/.pi/agent/antiloop.json`:
|
|
|
441
528
|
| `detectTaskStreams` | `true` | Recognize homogeneous batch work (same tool called with distinct content — e.g. punched/plan/obsidian extensions) and stay silent; see [task streams](#task-streams-n-tasks-of-one-type--a-loop) |
|
|
442
529
|
| `taskStreamMinCalls` | `3` | Same-tool calls required inside the window before a task stream is recognized |
|
|
443
530
|
| `taskStreamTwinThreshold` | `0.99` | Arguments this similar (or identical) count as *the same task* — a twin invalidates the stream and re-enables normal loop detection |
|
|
531
|
+
| `snapshotTools` | `["trimegisto_harvest"]` | Read-only snapshot/status tools: repeated identical reads (the serial close of N settled agents) never count as a tool loop or no-progress run. Add any other status/poll tool by name |
|
|
444
532
|
| `detectTextLoops` | `true` | Detect full-text repetition |
|
|
445
533
|
| `detectToolLoops` | `true` | Detect tool-call sequence + argument repetition |
|
|
446
534
|
| `detectThinkingLoops` | `true` | Detect repeated thinking/reasoning content |
|
|
@@ -460,7 +548,8 @@ Persisted as JSON at `~/.pi/agent/antiloop.json`:
|
|
|
460
548
|
6. **Let user input clear state** — each user message decays the consecutive counter by 2, so a fresh prompt naturally resets without `/antiloop reset`.
|
|
461
549
|
7. **Degenerate detector needs no tuning for most setups** — a run of ≥ 16 identical words (or one word ≥ 40% of a ≥ 50-token payload) inside a single message is conclusive stuck generation; the 46 KB `noguerol ×5145` SSH-wordlist meltdown is caught on first sight (warning), its bash never executes (`blockDegenerateBash`), and a second consecutive meltdown gets the force-break steer. If a model legitimately writes repetitive payloads, raise `degenerateMaxRun` / `degenerateMaxFreq` via `/antiloop config` — don't disable the detector.
|
|
462
550
|
8. **Outcome detector catches mutated re-run loops** — a model that re-issues the same experiment with cosmetic changes (labels, permutations) while every attempt fails identically gets a warning after `outcomeMinRepeats` (8) same-failure attempts, a steer on the next, and a hard stop shortly after. Converging sweeps, evolving failures and repeated successes stay silent by design. Lower `outcomeMinRepeats` if you want earlier cutoffs.
|
|
463
|
-
9.
|
|
551
|
+
9. **Block detector catches narration loops** — a single message replaying whole sentences (the `Let me start by checking the environment…` ×5 class) warns on first sight and escalates like a meltdown. It only fires when ≥ 85% of the message's 5-word phrases recur, so ordinary prose and templated logs/code stay silent. Raise `blockRepeatShare` (or `blockMinRepeats`) if you ever see a false positive.
|
|
552
|
+
10. **`/antiloop test`** — runs the real detection engine (text + tool-call + task-stream + degenerate + block + outcome regression cases) to verify calibration after any change.
|
|
464
553
|
|
|
465
554
|
## Architecture
|
|
466
555
|
|
|
@@ -485,10 +574,10 @@ Modular extension with zero external dependencies (only pi's bundled `@earendil-
|
|
|
485
574
|
|
|
486
575
|
- **Levenshtein + trigram Jaccard** hybrid — small texts use edit distance, large texts use n-gram overlap (each is O(N) in text length)
|
|
487
576
|
- **Sliding window** — only the last `detectionWindow` messages participate, capping memory at O(W × message_size)
|
|
488
|
-
- **Early bail** — short messages and empty tool calls skip similarity computation entirely; the degenerate
|
|
577
|
+
- **Early bail** — short messages and empty tool calls skip similarity computation entirely; the degenerate and block scans are single linear passes
|
|
489
578
|
- **TUI integration** — uses `ctx.ui.select` for the config menu and the log viewer; `ctx.ui.notify` for state notifications; `ctx.ui.setStatus` + a custom `ctx.ui.setFooter` component for the persistent footer indicator, live level info, and the `esc+a` keyboard toggle (`ctx.ui.onTerminalInput`, never consumes input)
|
|
490
|
-
- **Hooks** — `message_end` (track messages + tool call ids with a monotonic sequence so the sliding-window trim can never collide turn indices, and pre-handle degenerate
|
|
491
|
-
- **Intervention runs on the turn loop, not on user prompts** — escalation is decided at `turn_end` (and at `message_end` for the self-contained
|
|
579
|
+
- **Hooks** — `message_end` (track messages + tool call ids with a monotonic sequence so the sliding-window trim can never collide turn indices, and pre-handle degenerate/block repetition — the message is complete but its tools haven't executed yet), `tool_call` (block degenerate or replayed `bash` commands before they run), `turn_end` (attach result fingerprints with failure signatures, detect — including no-progress outcome runs — and intervene: steer the force break / abort the run), `input` (decay on real user messages only), `session_start` (load config + install footer + reset), `session_shutdown` (restore built-in footer)
|
|
580
|
+
- **Intervention runs on the turn loop, not on user prompts** — escalation is decided at `turn_end` (and at `message_end` for the self-contained intra-message signals), the break is steered into the running agent before its next LLM call, and the guaranteed hard stop aborts the run (`ctx.abort`, fire-and-forget — never awaited, so the hook can't deadlock). No custom-role messages are injected into the conversation at any level (steering a real user message + aborting are the only levers; custom-role injections were removed because a model can stall on an unexpected injected message)
|
|
492
581
|
|
|
493
582
|
## License
|
|
494
583
|
|
package/package.json
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "pi-antiloop",
|
|
3
|
-
"version": "1.
|
|
4
|
-
"description": "Antiloop: detect reasoning loops and force a break (warn → force → abort) across text, tool, thinking, structural, degenerate-repetition, and no-progress outcome patterns. Degenerate (one message stuck repeating a word hundreds of times — noguerol ×5145) is caught at message_end and its bash blocked before executing. No-progress (the NFS test-A…QQQ class: ~90 mutated re-runs of the same experiment, every one failing identically) fires when the same failing outcome repeats ≥ outcomeMinRepeats times with near-identical args. Fixes antiloop going blind mid-session: tracked messages use a monotonic sequence so the trim of the sliding window can never collide turn indices. The force break is delivered mid-run (steer before the next LLM call) and if the model ignores it antiloop aborts the run — an autonomous tool loop always terminates. Result-aware tool-loop detection keeps sequential bash and converging sweeps quiet; task-stream recognition keeps punched/plan batch work quiet.",
|
|
3
|
+
"version": "1.8.0",
|
|
4
|
+
"description": "Antiloop: detect reasoning loops and force a break (warn → force → abort) across text, tool, thinking, structural, degenerate-repetition, block/narration-repetition, and no-progress outcome patterns. Degenerate (one message stuck repeating a word hundreds of times — noguerol ×5145) is caught at message_end and its bash blocked before executing. Block (ONE message replaying whole sentences/phrases — the 'Let me start by checking the environment…' ×5 narration loop, invisible to every cross-message detector because there is no peer) fires on the first such message with the same strong turn weight and blocks a replayed bash command. No-progress (the NFS test-A…QQQ class: ~90 mutated re-runs of the same experiment, every one failing identically) fires when the same failing outcome repeats ≥ outcomeMinRepeats times with near-identical args. Snapshot tools (trimegisto_harvest) are exempt: a coordinator closing N settled agents re-reads the SAME board serially, which is an idempotent read, not a loop — so isVerbatimRepeat never fires on it. Fixes antiloop going blind mid-session: tracked messages use a monotonic sequence so the trim of the sliding window can never collide turn indices. The force break is delivered mid-run (steer before the next LLM call) and if the model ignores it antiloop aborts the run — an autonomous tool loop always terminates. Result-aware tool-loop detection keeps sequential bash and converging sweeps quiet; task-stream recognition keeps punched/plan batch work quiet.",
|
|
5
5
|
"keywords": [
|
|
6
6
|
"pi-package",
|
|
7
7
|
"antiloop",
|
package/src/commands.ts
CHANGED
|
@@ -59,10 +59,12 @@ async function showStatus(ctx: ExtensionCommandContext, rt: Runtime): Promise<vo
|
|
|
59
59
|
` tool sim: ${(rt.config.toolSimilarityThreshold * 100).toFixed(0)}% tool repeat: ${rt.config.minToolRepeatCount}+ prior`,
|
|
60
60
|
` result sim: ${(rt.config.resultSimilarityThreshold * 100).toFixed(0)}% (same cmd + diff outcome = no loop)`,
|
|
61
61
|
` degenerate: run ≥ ${rt.config.degenerateMaxRun} same word · freq ≥ ${rt.config.degenerateMaxFreq} @ ${(rt.config.degenerateMaxShare * 100).toFixed(0)}% (≥ ${rt.config.degenerateMinTokens} tokens) · weight ${rt.config.degenerateTurnWeight} · block bash ${yn(rt.config.blockDegenerateBash)}`,
|
|
62
|
+
` block repeats: ${yn(rt.config.detectBlockRepeats)} (≥ ${(rt.config.blockRepeatShare * 100).toFixed(0)}% of ≥ ${rt.config.blockMinTokens} tokens replay ${rt.config.blockNgram}-grams ×${rt.config.blockMinRepeats}+)`,
|
|
62
63
|
` outcome: same failing result ≥ ${rt.config.outcomeMinRepeats} attempts (args ≥ ${(rt.config.outcomeArgSimilarity * 100).toFixed(0)}% sim, sig ≥ ${(rt.config.outcomeSigThreshold * 100).toFixed(0)}%)`,
|
|
63
64
|
` task streams: ${yn(rt.config.detectTaskStreams)} (min ${rt.config.taskStreamMinCalls} calls, twins ≥ ${(rt.config.taskStreamTwinThreshold * 100).toFixed(0)}%)`,
|
|
65
|
+
` snapshot tools: ${rt.config.snapshotTools?.length ? rt.config.snapshotTools.join(", ") : "off"} (repeated reads are not a loop)`,
|
|
64
66
|
"",
|
|
65
|
-
`detectors: text ${yn(rt.config.detectTextLoops)} · tool ${yn(rt.config.detectToolLoops)} · think ${yn(rt.config.detectThinkingLoops)} · degenerate ${yn(rt.config.detectDegenerate)} · outcome ${yn(rt.config.detectOutcomeLoops)}`,
|
|
67
|
+
`detectors: text ${yn(rt.config.detectTextLoops)} · tool ${yn(rt.config.detectToolLoops)} · think ${yn(rt.config.detectThinkingLoops)} · degenerate ${yn(rt.config.detectDegenerate)} · block ${yn(rt.config.detectBlockRepeats)} · outcome ${yn(rt.config.detectOutcomeLoops)}`,
|
|
66
68
|
`footer: interactive ${yn(rt.config.interactiveFooter)} · toggle: ${rt.config.toggleShortcut}`,
|
|
67
69
|
];
|
|
68
70
|
if (rt.state.activeTaskStreams.length) {
|
|
@@ -95,13 +97,16 @@ async function showConfigMenu(ctx: ExtensionCommandContext, rt: Runtime): Promis
|
|
|
95
97
|
{ value: "resultSim" as const, label: `🧾 result similarity: ${(c.resultSimilarityThreshold * 100).toFixed(0)}%`, description: "same command + different result = progress, not a loop" },
|
|
96
98
|
// ── 🌀 Degenerate (intra-message meltdown) ───────────────────────
|
|
97
99
|
{ value: "degRun" as const, label: `🌀 degenerate run: ≥ ${c.degenerateMaxRun}`, description: "identical words in a row inside ONE message/call before it counts as stuck generation (noguerol ×5145 class)" },
|
|
98
|
-
{ value: "
|
|
100
|
+
{ value: "blk" as const, label: `🔁 block repetition: ${yn(c.detectBlockRepeats)}`, description: "ONE message replaying whole sentences/phrases (the narration loop: 'Let me start by checking…' ×5)" },
|
|
101
|
+
{ value: "blkShare" as const, label: `🔁 block share: ≥ ${(c.blockRepeatShare * 100).toFixed(0)}%`, description: "share of replayed 5-word phrases in one message before it counts as a replay" },
|
|
102
|
+
{ value: "degBlock" as const, label: `⛔ block repetitive bash: ${yn(c.blockDegenerateBash)}`, description: "stop a degenerate or replayed command before it executes (default on)" },
|
|
99
103
|
// ── 📉 Outcome (no-progress) ────────────────────────────────
|
|
100
104
|
{ value: "outcomeMin" as const, label: `📉 no-progress after: ${c.outcomeMinRepeats}`, description: "same failing outcome repeated this many times (mutated args ≥ 85% similar) before flagging — the NFS test-A…QQQ class" },
|
|
101
105
|
// ── 📋 Task streams ─────────────────────────────────────────
|
|
102
106
|
{ value: "streams" as const, label: `📋 task streams: ${yn(c.detectTaskStreams)}`, description: "batch work (punched_log / plan_manager / …) is not a loop" },
|
|
103
107
|
{ value: "streamMin" as const, label: `📋 stream min calls: ${c.taskStreamMinCalls}`, description: "calls of the same tool before a batch is recognized" },
|
|
104
108
|
{ value: "streamTwin" as const, label: `📋 twin threshold: ${(c.taskStreamTwinThreshold * 100).toFixed(0)}%`, description: "calls more similar than this = the same task repeated, not a batch" },
|
|
109
|
+
{ value: "snapshot" as const, label: `📡 snapshot tools: ${c.snapshotTools?.length ? c.snapshotTools.join(", ") : "off"}`, description: "read-only status tools (trimegisto_harvest): re-reading settled agents serially is not a loop" },
|
|
105
110
|
// ── 🔍 Detectors ────────────────────────────────────────────
|
|
106
111
|
{ value: "text" as const, label: `📝 text: ${yn(c.detectTextLoops)}`, description: "detect repeated text messages" },
|
|
107
112
|
{ value: "tool" as const, label: `🔧 tools: ${yn(c.detectToolLoops)}`, description: "detect repeated tool calls" },
|
|
@@ -227,9 +232,21 @@ async function showConfigMenu(ctx: ExtensionCommandContext, rt: Runtime): Promis
|
|
|
227
232
|
if (v !== undefined) { c.degenerateMaxRun = v; saveConfig(c); ctx.ui.notify(`degenerate run: ≥ ${v}`, "info"); }
|
|
228
233
|
break;
|
|
229
234
|
}
|
|
235
|
+
case "blk":
|
|
236
|
+
c.detectBlockRepeats = !c.detectBlockRepeats; saveConfig(c);
|
|
237
|
+
ctx.ui.notify(`block repetition: ${yn(c.detectBlockRepeats)}`, "info"); break;
|
|
238
|
+
case "blkShare": {
|
|
239
|
+
const v = await selectFrom(ctx, "🔁 block repetition share (replayed 5-gram positions in ONE message)", [
|
|
240
|
+
{ value: 0.7, label: "⚡ 70% (sensitive)" },
|
|
241
|
+
{ value: 0.85, label: "🎯 85% (default)" },
|
|
242
|
+
{ value: 0.95, label: "🐢 95% (only near-total replays)" },
|
|
243
|
+
]);
|
|
244
|
+
if (v !== undefined) { c.blockRepeatShare = v; saveConfig(c); ctx.ui.notify(`block share: ${(v * 100).toFixed(0)}%`, "info"); }
|
|
245
|
+
break;
|
|
246
|
+
}
|
|
230
247
|
case "degBlock":
|
|
231
248
|
c.blockDegenerateBash = !c.blockDegenerateBash; saveConfig(c);
|
|
232
|
-
ctx.ui.notify(`block
|
|
249
|
+
ctx.ui.notify(`block repetitive bash: ${yn(c.blockDegenerateBash)}`, "info"); break;
|
|
233
250
|
case "outcomeMin": {
|
|
234
251
|
const v = await selectFrom(ctx, "📉 no-progress threshold (same failing outcome, near-identical args)", [
|
|
235
252
|
{ value: 4, label: "⚡ 4 (sensitive — long experiment series get cut early)" },
|
|
@@ -262,6 +279,18 @@ async function showConfigMenu(ctx: ExtensionCommandContext, rt: Runtime): Promis
|
|
|
262
279
|
if (v !== undefined) { c.taskStreamTwinThreshold = v; saveConfig(c); ctx.ui.notify(`twin threshold: ${(v * 100).toFixed(0)}%`, "info"); }
|
|
263
280
|
break;
|
|
264
281
|
}
|
|
282
|
+
case "snapshot": {
|
|
283
|
+
const v = await selectFrom(ctx, "📡 snapshot / read-only tools (repeated identical reads are never a loop)", [
|
|
284
|
+
{ value: "default" as const, label: "🎯 trimegisto_harvest (default)", description: "serial close of settled agents: same args + same snapshot ×N stays silent" },
|
|
285
|
+
{ value: "off" as const, label: "🚫 off", description: "treat every tool identically (harvest can loop-flag again)" },
|
|
286
|
+
]);
|
|
287
|
+
if (v !== undefined) {
|
|
288
|
+
c.snapshotTools = v === "off" ? [] : ["trimegisto_harvest"];
|
|
289
|
+
saveConfig(c);
|
|
290
|
+
ctx.ui.notify(`snapshot tools: ${c.snapshotTools.length ? c.snapshotTools.join(", ") : "off"}`, "info");
|
|
291
|
+
}
|
|
292
|
+
break;
|
|
293
|
+
}
|
|
265
294
|
case "ignoredBreak": {
|
|
266
295
|
const v = await selectFrom(ctx, "🛑 stop after ignored break (identical repeats after the force break)", [
|
|
267
296
|
{ value: 1, label: "⚡ 1 (sensitive — one identical repeat after the break stops the run)" },
|
package/src/config.ts
CHANGED
|
@@ -24,6 +24,11 @@ export const DEFAULT_CONFIG: AntiloopConfig = {
|
|
|
24
24
|
degenerateMaxShare: 0.4,
|
|
25
25
|
degenerateTurnWeight: 2,
|
|
26
26
|
blockDegenerateBash: true,
|
|
27
|
+
detectBlockRepeats: true,
|
|
28
|
+
blockMinTokens: 120,
|
|
29
|
+
blockNgram: 5,
|
|
30
|
+
blockMinRepeats: 3,
|
|
31
|
+
blockRepeatShare: 0.85,
|
|
27
32
|
detectOutcomeLoops: true,
|
|
28
33
|
outcomeMinRepeats: 8,
|
|
29
34
|
outcomeArgSimilarity: 0.85,
|
|
@@ -31,6 +36,7 @@ export const DEFAULT_CONFIG: AntiloopConfig = {
|
|
|
31
36
|
detectTaskStreams: true,
|
|
32
37
|
taskStreamMinCalls: 3,
|
|
33
38
|
taskStreamTwinThreshold: 0.99,
|
|
39
|
+
snapshotTools: ["trimegisto_harvest"],
|
|
34
40
|
detectToolLoops: true,
|
|
35
41
|
detectThinkingLoops: true,
|
|
36
42
|
detectTextLoops: true,
|
package/src/detect.ts
CHANGED
|
@@ -185,12 +185,124 @@ export function degenerateDescription(hit: DegenerateHit): string {
|
|
|
185
185
|
return `${hit.where}: degenerate repetition — "${hit.token}" ×${hit.freq} (${share}% of ${hit.total} tokens, longest run ${hit.maxRun})`;
|
|
186
186
|
}
|
|
187
187
|
|
|
188
|
+
// ---------------------------------------------------------------------------
|
|
189
|
+
// Intra-message BLOCK repetition (v1.7).
|
|
190
|
+
//
|
|
191
|
+
// The "narration loop" class (verified against a real coding session: ONE
|
|
192
|
+
// assistant message replaying ~5 near-verbatim cycles of "Let me start by
|
|
193
|
+
// checking the environment and the current state of the repository… / I'll run
|
|
194
|
+
// several independent checks in parallel. / Let me begin the S0 development…").
|
|
195
|
+
// Every pre-existing detector stayed silent: text/tool/thinking/structural all
|
|
196
|
+
// compare ACROSS messages (the model produced one message, no peer), and the
|
|
197
|
+
// degenerate detector only watches ONE word repeated hundreds of times, not a
|
|
198
|
+
// whole sentence/paragraph replayed. The signal here is phrase-level: a sliding
|
|
199
|
+
// window of blockNgram-word n-grams over the normalized payload; if almost ALL
|
|
200
|
+
// of those n-grams recur and the most frequent one recurs ≥ blockMinRepeats
|
|
201
|
+
// times, the generation is replaying itself instead of advancing.
|
|
202
|
+
//
|
|
203
|
+
// Deliberately conservative: only a payload where ≥ blockRepeatShare (default
|
|
204
|
+
// 85%) of 5-gram positions repeat qualifies. Ordinary prose — even long,
|
|
205
|
+
// structured docs — sits far below (README/pi.md paragraphs: ≤ 0.11), while
|
|
206
|
+
// 3+ replays of a narration block reach 0.98–1.0. Repeat-linked code/log lines
|
|
207
|
+
// that share a template are NOT flagged because their n-grams carry the varying
|
|
208
|
+
// digits and differ.
|
|
209
|
+
// ---------------------------------------------------------------------------
|
|
210
|
+
|
|
211
|
+
/** Recursive n-gram coverage of one payload. */
|
|
212
|
+
export interface BlockRepeatInfo {
|
|
213
|
+
/** Recurring n-gram positions / total n-gram positions (0..1). */
|
|
214
|
+
ratio: number;
|
|
215
|
+
/** Occurrences of the MOST repeated n-gram. */
|
|
216
|
+
repeats: number;
|
|
217
|
+
/** Total normalized words in the payload. */
|
|
218
|
+
tokens: number;
|
|
219
|
+
/** The n used (words per window). */
|
|
220
|
+
ngram: number;
|
|
221
|
+
/** The most repeated n-gram, for the description. */
|
|
222
|
+
sample: string;
|
|
223
|
+
}
|
|
224
|
+
|
|
225
|
+
export interface BlockRepeatHit extends BlockRepeatInfo {
|
|
226
|
+
where: string;
|
|
227
|
+
}
|
|
228
|
+
|
|
229
|
+
/**
|
|
230
|
+
* Detect phrase/block-level self-repetition inside ONE payload. Returns the
|
|
231
|
+
* coverage ratio with the most repeated n-gram, or undefined for normal text.
|
|
232
|
+
*
|
|
233
|
+
* Reads the payload as a stream of lowercase alphanumeric words (code symbols
|
|
234
|
+
* and punctuation are separators, so stored-JSON escaping cannot glue tokens).
|
|
235
|
+
* Coverage counts an n-gram position as repeated when the SAME n-word sequence
|
|
236
|
+
* (digits included, so templated lines with varying numbers do not count)
|
|
237
|
+
* occurs somewhere else in the payload.
|
|
238
|
+
*/
|
|
239
|
+
export function findRepetitiveBlock(text: string, config: AntiloopConfig): BlockRepeatInfo | undefined {
|
|
240
|
+
const words = String(text).toLowerCase().match(/[\p{L}\p{N}]+/gu);
|
|
241
|
+
if (!words) return undefined;
|
|
242
|
+
const total = words.length;
|
|
243
|
+
if (total < config.blockMinTokens) return undefined;
|
|
244
|
+
const n = Math.max(2, config.blockNgram);
|
|
245
|
+
const minRepeats = Math.max(2, config.blockMinRepeats);
|
|
246
|
+
if (total < n * minRepeats) return undefined;
|
|
247
|
+
|
|
248
|
+
const counts = new Map<string, number>();
|
|
249
|
+
const grams: string[] = [];
|
|
250
|
+
for (let i = 0; i + n <= total; i++) {
|
|
251
|
+
const g = words.slice(i, i + n).join(" ");
|
|
252
|
+
grams.push(g);
|
|
253
|
+
counts.set(g, (counts.get(g) ?? 0) + 1);
|
|
254
|
+
}
|
|
255
|
+
if (!grams.length) return undefined;
|
|
256
|
+
|
|
257
|
+
let covered = 0;
|
|
258
|
+
let topKey = "";
|
|
259
|
+
let top = 0;
|
|
260
|
+
for (const g of grams) {
|
|
261
|
+
const c = counts.get(g)!;
|
|
262
|
+
if (c > top) {
|
|
263
|
+
top = c;
|
|
264
|
+
topKey = g;
|
|
265
|
+
}
|
|
266
|
+
if (c >= 2) covered++;
|
|
267
|
+
}
|
|
268
|
+
// At least one n-word phrase must recur blockMinRepeats times (a single
|
|
269
|
+
// echo is a restatement, not a replay) and the recurrence must dominate.
|
|
270
|
+
if (top < minRepeats) return undefined;
|
|
271
|
+
const ratio = covered / grams.length;
|
|
272
|
+
if (ratio < config.blockRepeatShare) return undefined;
|
|
273
|
+
return { ratio, repeats: top, tokens: total, ngram: n, sample: topKey };
|
|
274
|
+
}
|
|
275
|
+
|
|
276
|
+
/** Scan one assistant message (text + each tool-call argument) for a replay. */
|
|
277
|
+
export function scanMessageBlock(
|
|
278
|
+
content: string,
|
|
279
|
+
toolCalls: TrackedToolCall[] | undefined,
|
|
280
|
+
config: AntiloopConfig,
|
|
281
|
+
): BlockRepeatHit | undefined {
|
|
282
|
+
if (!config.detectBlockRepeats) return undefined;
|
|
283
|
+
if (content && content.length) {
|
|
284
|
+
const b = findRepetitiveBlock(content, config);
|
|
285
|
+
if (b) return { ...b, where: "message text" };
|
|
286
|
+
}
|
|
287
|
+
for (const tc of toolCalls ?? []) {
|
|
288
|
+
if (!tc.args) continue;
|
|
289
|
+
const b = findRepetitiveBlock(tc.args, config);
|
|
290
|
+
if (b) return { ...b, where: tc.name === "bash" ? "bash command" : `args(${tc.name})` };
|
|
291
|
+
}
|
|
292
|
+
return undefined;
|
|
293
|
+
}
|
|
294
|
+
|
|
295
|
+
export function blockRepeatDescription(hit: BlockRepeatHit): string {
|
|
296
|
+
const pct = Math.round(hit.ratio * 100);
|
|
297
|
+
return `${hit.where}: block repetition — ${pct}% of ${hit.tokens} tokens replay repeated ${hit.ngram}-word phrases ("${hit.sample}" ×${hit.repeats})`;
|
|
298
|
+
}
|
|
299
|
+
|
|
188
300
|
/** Consecutive-detection weight of a detection turn. The strong
|
|
189
301
|
* self-contained signals (degenerate meltdown, proven no-progress outcome run)
|
|
190
302
|
* add degenerateTurnWeight (default 2) points so the FIRST one already reaches
|
|
191
303
|
* the warning level and escalation is fast on repeat. */
|
|
192
304
|
export function detectionTurnWeight(detections: LoopDetection[], config: AntiloopConfig): number {
|
|
193
|
-
const strong = detections.some((d) => d.type === "degenerate" || d.type === "outcome");
|
|
305
|
+
const strong = detections.some((d) => d.type === "degenerate" || d.type === "block" || d.type === "outcome");
|
|
194
306
|
return strong ? Math.max(1, config.degenerateTurnWeight) : 1;
|
|
195
307
|
}
|
|
196
308
|
|
|
@@ -383,6 +495,28 @@ export function detectTaskStreams(
|
|
|
383
495
|
return streams;
|
|
384
496
|
}
|
|
385
497
|
|
|
498
|
+
// ---------------------------------------------------------------------------
|
|
499
|
+
// Snapshot / read-only tools (v1.8).
|
|
500
|
+
//
|
|
501
|
+
// A coordinator closing N agents that already finished calls the snapshot tool
|
|
502
|
+
// once per agent and gets the SAME serialized board every time (all settled).
|
|
503
|
+
// Same args + same result looks exactly like a verbatim tool loop to the
|
|
504
|
+
// detector, but re-reading a state snapshot is an idempotent read — the serial
|
|
505
|
+
// close of finished work, not a stuck generation. Tools listed in
|
|
506
|
+
// config.snapshotTools (default: trimegisto_harvest) are therefore exempt from
|
|
507
|
+
// tool-loop and outcome detection. Any other status/poll tool can be added.
|
|
508
|
+
// ---------------------------------------------------------------------------
|
|
509
|
+
|
|
510
|
+
/** Set of tool names whose repeated identical reads must never count as a loop. */
|
|
511
|
+
function snapshotToolSet(config: AntiloopConfig): Set<string> {
|
|
512
|
+
return new Set((config.snapshotTools ?? []).map((n) => n.trim()).filter(Boolean));
|
|
513
|
+
}
|
|
514
|
+
|
|
515
|
+
/** True when every call in the set is a snapshot/status read. */
|
|
516
|
+
function allSnapshotCalls(calls: TrackedToolCall[] | undefined, snapshots: Set<string>): boolean {
|
|
517
|
+
return !!calls && calls.length > 0 && calls.every((c) => snapshots.has(c.name));
|
|
518
|
+
}
|
|
519
|
+
|
|
386
520
|
export function detectLoops(state: AntiloopState, config: AntiloopConfig): LoopDetection[] {
|
|
387
521
|
const out: LoopDetection[] = [];
|
|
388
522
|
const msgs = state.recentMessages;
|
|
@@ -391,9 +525,10 @@ export function detectLoops(state: AntiloopState, config: AntiloopConfig): LoopD
|
|
|
391
525
|
const win = msgs.slice(start);
|
|
392
526
|
const now = Date.now();
|
|
393
527
|
|
|
394
|
-
// Intra-message
|
|
395
|
-
//
|
|
396
|
-
// meltdown
|
|
528
|
+
// Intra-message repetition fires on a SINGLE pathological message (no peer
|
|
529
|
+
// needed) and is independent of the task-stream batch gate: a meltdown is a
|
|
530
|
+
// meltdown even mid-batch. Two flavors: degenerate (one word ×hundreds) and
|
|
531
|
+
// block (whole sentences/phrases replayed — the narration-loop class).
|
|
397
532
|
if (config.detectDegenerate) {
|
|
398
533
|
const last = win[win.length - 1];
|
|
399
534
|
const hit = scanMessageDegenerate(last.content, last.toolCalls, config);
|
|
@@ -408,12 +543,29 @@ export function detectLoops(state: AntiloopState, config: AntiloopConfig): LoopD
|
|
|
408
543
|
}
|
|
409
544
|
}
|
|
410
545
|
|
|
546
|
+
if (config.detectBlockRepeats) {
|
|
547
|
+
const last = win[win.length - 1];
|
|
548
|
+
const hit = scanMessageBlock(last.content, last.toolCalls, config);
|
|
549
|
+
if (hit) {
|
|
550
|
+
out.push({
|
|
551
|
+
type: "block",
|
|
552
|
+
// 0.99 => isVerbatimRepeat(): after a force break, a replayed block
|
|
553
|
+
// is the model ignoring the break and escalates to the hard stop.
|
|
554
|
+
similarity: 0.99,
|
|
555
|
+
messageIndices: [msgs.length - 1],
|
|
556
|
+
description: blockRepeatDescription(hit),
|
|
557
|
+
timestamp: now,
|
|
558
|
+
});
|
|
559
|
+
}
|
|
560
|
+
}
|
|
561
|
+
|
|
411
562
|
// Task-stream gate: if the window is a homogeneous batch (same extension
|
|
412
563
|
// tool called with DISTINCT content ≥ taskStreamMinCalls times), that tool
|
|
413
564
|
// is exempt from tool-loop detection, and text/thinking/structural
|
|
414
565
|
// detections that only involve batch messages are suppressed too — the
|
|
415
566
|
// model is doing N different tasks of the same type, not looping.
|
|
416
567
|
const streams = detectTaskStreams(win, config);
|
|
568
|
+
const snapshots = snapshotToolSet(config);
|
|
417
569
|
const batchAt = new Set<number>();
|
|
418
570
|
win.forEach((m, idx) => {
|
|
419
571
|
const calls = m.toolCalls;
|
|
@@ -467,7 +619,10 @@ export function detectLoops(state: AntiloopState, config: AntiloopConfig): LoopD
|
|
|
467
619
|
const lastCalls = last.toolCalls;
|
|
468
620
|
// A message whose calls are all task-stream tools is batch work — skip
|
|
469
621
|
// it entirely (the stream gate already proved the calls are distinct).
|
|
470
|
-
|
|
622
|
+
// Snapshot/status tools (trimegisto_harvest…) are also skipped: repeated
|
|
623
|
+
// identical reads of a settled board are the serial close of finished
|
|
624
|
+
// agents, not a loop.
|
|
625
|
+
if (lastCalls && lastCalls.length && !batchAt.has(msgs.length - 1) && !allSnapshotCalls(lastCalls, snapshots)) {
|
|
471
626
|
const matched: number[] = [];
|
|
472
627
|
for (let i = 0; i < win.length - 1; i++) {
|
|
473
628
|
const prev = win[i].toolCalls;
|
|
@@ -525,7 +680,9 @@ export function detectLoops(state: AntiloopState, config: AntiloopConfig): LoopD
|
|
|
525
680
|
// produce the SAME OK outcome every call — they must never count. And
|
|
526
681
|
// batch messages must NOT be skipped here: a "bash ×N stream" with
|
|
527
682
|
// identical failures is exactly the no-progress loop to catch.
|
|
528
|
-
|
|
683
|
+
// Snapshot/status reads are excluded outright: an identical harvest
|
|
684
|
+
// snapshot is a state read, not a repeated failing experiment.
|
|
685
|
+
if (lc.result && isFailResult(lc.result) && !snapshots.has(lc.name)) {
|
|
529
686
|
let matches = 0;
|
|
530
687
|
for (let i = 0; i < win.length - 1; i++) {
|
|
531
688
|
const m = win[i];
|
|
@@ -673,8 +830,10 @@ export function runSelfTest(): string[] {
|
|
|
673
830
|
detectTextLoops: true, notifyOnDetection: true, maxHistoryEntries: 100,
|
|
674
831
|
detectionWindow: 10, interactiveFooter: true, toggleShortcut: "esc+a",
|
|
675
832
|
detectTaskStreams: true, taskStreamMinCalls: 3, taskStreamTwinThreshold: 0.99,
|
|
833
|
+
snapshotTools: ["trimegisto_harvest"],
|
|
676
834
|
detectDegenerate: true, degenerateMinTokens: 50, degenerateMaxRun: 16,
|
|
677
835
|
degenerateMaxFreq: 60, degenerateMaxShare: 0.4, degenerateTurnWeight: 2, blockDegenerateBash: true,
|
|
836
|
+
detectBlockRepeats: true, blockMinTokens: 120, blockNgram: 5, blockMinRepeats: 3, blockRepeatShare: 0.85,
|
|
678
837
|
detectOutcomeLoops: true, outcomeMinRepeats: 8, outcomeArgSimilarity: 0.85, outcomeSigThreshold: 0.7,
|
|
679
838
|
};
|
|
680
839
|
const asState = (recentMessages: TrackedMessage[]): AntiloopState =>
|
|
@@ -782,6 +941,56 @@ export function runSelfTest(): string[] {
|
|
|
782
941
|
const wTxt = detectionTurnWeight(dl("text", 0.8), tcfg);
|
|
783
942
|
out.push(`degenerate turn weight → degenerate ${wDeg}, text ${wTxt} (exp 2, 1) ${wDeg === 2 && wTxt === 1 ? "✅" : "❌"}`);
|
|
784
943
|
|
|
944
|
+
// --- v1.7: intra-message BLOCK repetition (narration loops) ---
|
|
945
|
+
// Regression: a real coding session where ONE assistant message replayed the
|
|
946
|
+
// same ~5 sentences in a loop ("Let me start by checking the environment…" /
|
|
947
|
+
// "I'll run several independent checks in parallel." / "Let me begin the S0
|
|
948
|
+
// development…"). Cross-message text/tool/thinking detectors need a peer and
|
|
949
|
+
// stayed silent; the degenerate scan only watches a single word. The block
|
|
950
|
+
// detector must fire on the message itself.
|
|
951
|
+
const NARRV = [
|
|
952
|
+
"Let me start by checking the environment and the current state of the repository, then set up a plan for the S0 slice and begin building.",
|
|
953
|
+
"I'll run several independent checks in parallel.",
|
|
954
|
+
"Let me begin the S0 development. First, reconnaissance of the environment and current repo state.",
|
|
955
|
+
"Let me check what's available in the environment (node, package managers, network, postgres) and the current repo state, then set up the S0 plan and start building the monorepo scaffold.",
|
|
956
|
+
"I'll run a batch of independent environment checks first.",
|
|
957
|
+
];
|
|
958
|
+
const cycles = (k: number): string => {
|
|
959
|
+
let s = "";
|
|
960
|
+
for (let i = 0; i < k; i++) for (const v of NARRV) s += v + "\n\n";
|
|
961
|
+
return s;
|
|
962
|
+
};
|
|
963
|
+
const s0Det = detectLoops(asState([mk(cycles(5))]), tcfg);
|
|
964
|
+
const s0Hit = s0Det.find((d) => d.type === "block");
|
|
965
|
+
out.push(`block S0 narration loop → ${s0Hit ? `block (${s0Hit.description})` : "no"} (exp block — was the miss) ${s0Hit ? "✅" : "❌"}`);
|
|
966
|
+
out.push(`block turn weight → ${detectionTurnWeight(s0Det, tcfg)} (exp 2 — warn on first sight) ${detectionTurnWeight(s0Det, tcfg) === 2 ? "✅" : "❌"}`);
|
|
967
|
+
out.push(`block post-steer = ignored→ ${isVerbatimRepeat(s0Det) ? "yes" : "no"} (exp yes — replayed block after the break) ${isVerbatimRepeat(s0Det) ? "✅" : "❌"}`);
|
|
968
|
+
|
|
969
|
+
// 3 replay cycles still flag; a single restatement (2 cycles, top-gram ×2)
|
|
970
|
+
// does not — one echo is a summary, not stuck generation.
|
|
971
|
+
const c3 = findRepetitiveBlock(cycles(3), tcfg);
|
|
972
|
+
const c2 = findRepetitiveBlock(cycles(2), tcfg);
|
|
973
|
+
out.push(`block 3 cycles / 2 cycles→ ${c3 ? `flag (${Math.round(c3.ratio * 100)}%)` : "no"} / ${c2 ? "flag" : "no"} (exp flag / no — needs ≥3 repeats) ${c3 && !c2 ? "✅" : "❌"}`);
|
|
974
|
+
|
|
975
|
+
// Ordinary prose and templated (but evolving) payloads must stay silent:
|
|
976
|
+
// long docs score ≤ 0.11 coverage; templated logs/code carry varying digits
|
|
977
|
+
// so their 5-grams differ.
|
|
978
|
+
const prose =
|
|
979
|
+
"The extension watches every assistant message and tool call to decide whether the model is making progress. " +
|
|
980
|
+
"It stores a short window of recent turns, fingerprints tool results and compares them with earlier attempts. " +
|
|
981
|
+
"When a pattern repeats it escalates from a quiet warning to a forced change of approach and finally aborts. " +
|
|
982
|
+
"Configuration lives in a small json file next to the agent directory and every threshold can be tuned at runtime. " +
|
|
983
|
+
"The default values were chosen against real sessions so ordinary work never triggers a false positive. " +
|
|
984
|
+
"Detectors are independent: disabling one leaves the others active and the footer keeps the user informed. " +
|
|
985
|
+
"A clean turn cools the counter down so a recovered model is given room to finish the task. " +
|
|
986
|
+
"Everything is written in plain typescript with no runtime dependencies beyond the host package. " +
|
|
987
|
+
"New strategies should be measured against the recorded payloads before they are enabled by default. " +
|
|
988
|
+
"The goal is simple: catch a stuck model early and keep the context useful for the real work.";
|
|
989
|
+
const logLines = [...Array(30)].map((_, i) => `2026-09-24T10:${String(i).padStart(2, "0")}:00Z INFO worker ${i} processed job ${1000 + i} in ${i * 3}ms status ok`).join("\n");
|
|
990
|
+
const proseHit = findRepetitiveBlock(prose, tcfg);
|
|
991
|
+
const logHit = findRepetitiveBlock(logLines, tcfg);
|
|
992
|
+
out.push(`block legit prose/logs → ${proseHit || logHit ? "flag" : "ok"} (exp ok) ${!proseHit && !logHit ? "✅" : "❌"}`);
|
|
993
|
+
|
|
785
994
|
// --- v1.6.1: no-progress outcome runs (mutated re-runs, same outcome) ---
|
|
786
995
|
// Regression: the NFS session — ~90 mutated re-runs of the SAME experiment
|
|
787
996
|
// (ssh exportfs/mount, labels test A…test QQQ, targets alternating Javi /
|
|
@@ -872,5 +1081,32 @@ export function runSelfTest(): string[] {
|
|
|
872
1081
|
out.push(`outcome identical OKs → ${okDet.some((d) => d.type === "outcome") ? "outcome" : "silent"} (exp silent — success repeats ≠ loop) ${!okDet.some((d) => d.type === "outcome") ? "✅" : "❌"}`);
|
|
873
1082
|
out.push(`isFailResult gate → err| → ${isFailResult("err|boom") ? "fail" : "ok"}, rc32 → ${isFailResult("ok|rc32 denied") ? "fail" : "ok"}, rc0/ok → ${isFailResult("ok|rc 0 12 passed") ? "fail" : "ok"} (exp fail, fail, ok) ${isFailResult("err|boom") && isFailResult("ok|rc32 denied") && !isFailResult("ok|rc 0 12 passed") ? "✅" : "❌"}`);
|
|
874
1083
|
|
|
1084
|
+
// --- v1.8: snapshot/read-only tools are idempotent reads, not loops ---
|
|
1085
|
+
// Regression (real session /home/j 2026-09-23T21:35): a coordinator closed
|
|
1086
|
+
// N agents that had ALREADY settled by calling trimegisto_harvest five times
|
|
1087
|
+
// with `{}`; every call returned the SAME 2.3 KB cumulative snapshot (all
|
|
1088
|
+
// agents done). Same args + same result ×5 → the tool detector force-broke
|
|
1089
|
+
// the run. Re-reading a settled state board is the serial close of finished
|
|
1090
|
+
// agents, not a stuck generation: with the snapshot exemption it is silent.
|
|
1091
|
+
const harvestSnap = resultFingerprint(
|
|
1092
|
+
[{
|
|
1093
|
+
type: "text",
|
|
1094
|
+
text:
|
|
1095
|
+
"## Trimegisto harvest (instant snapshot)\n\n### ✅ t0a [Active] — done (101s)\nTask: Explora el sistema en busca de LM Studio y sus runtimes.\n\n```\n# Informe: LM Studio en el sistema\n...\n```\n*10 turns, ↑32825 ↓2429*\n\n_All agents settled._",
|
|
1096
|
+
}],
|
|
1097
|
+
false,
|
|
1098
|
+
)!;
|
|
1099
|
+
const harvestMsgs = [0, 1, 2, 3, 4].map(() =>
|
|
1100
|
+
mk("", [{ name: "trimegisto_harvest", args: "{}", result: harvestSnap }]),
|
|
1101
|
+
);
|
|
1102
|
+
const snapDet = detectLoops(asState(harvestMsgs), tcfg);
|
|
1103
|
+
out.push(`snapshot harvest ×5 → ${snapDet.length ? snapDet.map((d) => d.type).join(",") : "silent"} (exp silent — serial close of settled agents) ${snapDet.length === 0 ? "✅" : "❌"}`);
|
|
1104
|
+
|
|
1105
|
+
// Guard: the exemption is what silences it — the SAME payload with
|
|
1106
|
+
// snapshotTools off is a genuine verbatim tool loop.
|
|
1107
|
+
const noSnapCfg: AntiloopConfig = { ...tcfg, snapshotTools: [] };
|
|
1108
|
+
const snapLoud = detectLoops(asState(harvestMsgs), noSnapCfg);
|
|
1109
|
+
out.push(`snapshot off = loop → ${snapLoud.some((d) => d.type === "tool") ? "tool" : "no"} (exp tool — normal tools still loop) ${snapLoud.some((d) => d.type === "tool") ? "✅" : "❌"}`);
|
|
1110
|
+
|
|
875
1111
|
return out;
|
|
876
1112
|
}
|
package/src/index.ts
CHANGED
|
@@ -24,6 +24,13 @@
|
|
|
24
24
|
* hook (blockDegenerateBash). One degenerate turn counts degenerateTurnWeight
|
|
25
25
|
* (2) consecutive points: warning on first sight, force-break steer on the
|
|
26
26
|
* second consecutive meltdown, hard stop shortly after if it keeps repeating.
|
|
27
|
+
*
|
|
28
|
+
* v1.7 adds the intra-message BLOCK detector for the narration-loop class:
|
|
29
|
+
* ONE message that replays whole sentences/phrases ("Let me start by checking
|
|
30
|
+
* the environment…" ×5), which every cross-message detector misses because
|
|
31
|
+
* there is no peer message and the degenerate scan only watches single words.
|
|
32
|
+
* Same message_end timing and strong turn weight; the bash gate refuses a
|
|
33
|
+
* command that is itself a replayed block.
|
|
27
34
|
*/
|
|
28
35
|
|
|
29
36
|
import type { ExtensionAPI, ExtensionContext } from "@earendil-works/pi-coding-agent";
|
|
@@ -285,29 +292,55 @@ export default function antiloopExtension(pi: ExtensionAPI) {
|
|
|
285
292
|
state.recentMessages = state.recentMessages.slice(-(config.detectionWindow + 5));
|
|
286
293
|
}
|
|
287
294
|
|
|
288
|
-
// v1.6 —
|
|
289
|
-
// turn_end only fires after tool execution (too late to stop the
|
|
290
|
-
// "noguerol ×5000" bash from running) and an aborted generation may
|
|
291
|
-
// reach turn_end at all. The signal is self-contained (one
|
|
292
|
-
// payload, no peer message) so the full escalation ladder runs
|
|
293
|
-
// the turn-guard makes the later turn_end pass skip this message
|
|
294
|
-
// detection is double-counted.
|
|
295
|
-
|
|
296
|
-
|
|
297
|
-
|
|
298
|
-
|
|
295
|
+
// v1.6/v1.7 — intra-message repetition is handled HERE, at message_end,
|
|
296
|
+
// because turn_end only fires after tool execution (too late to stop the
|
|
297
|
+
// 46 KB "noguerol ×5000" bash from running) and an aborted generation may
|
|
298
|
+
// never reach turn_end at all. The signal is self-contained (one
|
|
299
|
+
// pathological payload, no peer message) so the full escalation ladder runs
|
|
300
|
+
// right now; the turn-guard makes the later turn_end pass skip this message
|
|
301
|
+
// and no detection is double-counted. Two flavors: degenerate (one word
|
|
302
|
+
// ×hundreds) and block (whole sentences/phrases replayed — the narration
|
|
303
|
+
// loop class where every cross-message detector is blind).
|
|
304
|
+
if (!config.detectDegenerate && !config.detectBlockRepeats) return;
|
|
305
|
+
const {
|
|
306
|
+
scanMessageDegenerate,
|
|
307
|
+
degenerateDescription,
|
|
308
|
+
scanMessageBlock,
|
|
309
|
+
blockRepeatDescription,
|
|
310
|
+
nextLevel,
|
|
311
|
+
isVerbatimRepeat,
|
|
312
|
+
} = await import("./detect.ts");
|
|
313
|
+
let det: LoopDetection | undefined;
|
|
314
|
+
if (config.detectDegenerate) {
|
|
315
|
+
const degHit = scanMessageDegenerate(content, toolCalls, config);
|
|
316
|
+
if (degHit) {
|
|
317
|
+
det = {
|
|
318
|
+
type: "degenerate",
|
|
319
|
+
similarity: 1,
|
|
320
|
+
messageIndices: [state.recentMessages.length - 1],
|
|
321
|
+
description: degenerateDescription(degHit),
|
|
322
|
+
timestamp: Date.now(),
|
|
323
|
+
};
|
|
324
|
+
}
|
|
325
|
+
}
|
|
326
|
+
if (!det && config.detectBlockRepeats) {
|
|
327
|
+
const blkHit = scanMessageBlock(content, toolCalls, config);
|
|
328
|
+
if (blkHit) {
|
|
329
|
+
det = {
|
|
330
|
+
type: "block",
|
|
331
|
+
similarity: 0.99,
|
|
332
|
+
messageIndices: [state.recentMessages.length - 1],
|
|
333
|
+
description: blockRepeatDescription(blkHit),
|
|
334
|
+
timestamp: Date.now(),
|
|
335
|
+
};
|
|
336
|
+
}
|
|
337
|
+
}
|
|
338
|
+
if (!det) return;
|
|
299
339
|
const tracked = state.recentMessages[state.recentMessages.length - 1];
|
|
300
340
|
if (tracked) state.lastDetectedTurnIndex = tracked.turnIndex;
|
|
301
341
|
const prevLevel = state.currentLevel;
|
|
302
342
|
state.consecutiveDetections += Math.max(1, config.degenerateTurnWeight);
|
|
303
343
|
state.totalDetections++;
|
|
304
|
-
const det: LoopDetection = {
|
|
305
|
-
type: "degenerate",
|
|
306
|
-
similarity: 1,
|
|
307
|
-
messageIndices: [state.recentMessages.length - 1],
|
|
308
|
-
description: degenerateDescription(hit),
|
|
309
|
-
timestamp: Date.now(),
|
|
310
|
-
};
|
|
311
344
|
state.detections.push(det);
|
|
312
345
|
if (state.detections.length > config.maxHistoryEntries) {
|
|
313
346
|
state.detections = state.detections.slice(-config.maxHistoryEntries);
|
|
@@ -331,13 +364,13 @@ export default function antiloopExtension(pi: ExtensionAPI) {
|
|
|
331
364
|
ctx.ui.notify(`antiloop: force break — ${det.description}`, "error");
|
|
332
365
|
}
|
|
333
366
|
} else if (isVerbatimRepeat([det])) {
|
|
334
|
-
// The model produced
|
|
335
|
-
// count it; once the ignore limit is hit the run
|
|
367
|
+
// The model produced the same intra-message repetition AGAIN after the
|
|
368
|
+
// break message: count it; once the ignore limit is hit, cut the run.
|
|
336
369
|
state.ignoredSteerCount++;
|
|
337
370
|
if (state.ignoredSteerCount >= config.ignoredSteerLimit) {
|
|
338
371
|
hardStop(
|
|
339
372
|
ctx,
|
|
340
|
-
`antiloop: abort —
|
|
373
|
+
`antiloop: abort — repetitive output repeated ${state.ignoredSteerCount}× after the force break — run stopped; provide new instructions`,
|
|
341
374
|
);
|
|
342
375
|
}
|
|
343
376
|
}
|
|
@@ -379,15 +412,28 @@ export default function antiloopExtension(pi: ExtensionAPI) {
|
|
|
379
412
|
* sees why the command was refused and can change approach.
|
|
380
413
|
*/
|
|
381
414
|
pi.on("tool_call", async (event, ctx) => {
|
|
382
|
-
if (!config.enabled || !config.
|
|
415
|
+
if (!config.enabled || !config.blockDegenerateBash) return;
|
|
383
416
|
if (!isToolCallEventType("bash", event)) return;
|
|
384
|
-
const
|
|
385
|
-
const
|
|
386
|
-
if (
|
|
387
|
-
|
|
388
|
-
|
|
389
|
-
|
|
390
|
-
|
|
417
|
+
const command = event.input.command ?? "";
|
|
418
|
+
const { findDegenerateRepetition, findRepetitiveBlock, degenerateDescription, blockRepeatDescription } = await import("./detect.ts");
|
|
419
|
+
if (config.detectDegenerate) {
|
|
420
|
+
const hit = findDegenerateRepetition(command, config);
|
|
421
|
+
if (hit) {
|
|
422
|
+
return {
|
|
423
|
+
block: true,
|
|
424
|
+
reason: `[antiloop] blocked: ${degenerateDescription({ ...hit, where: "bash command" })} — this is stuck generation, not a real command. Do NOT retry it: stop and take one small, concrete step instead.`,
|
|
425
|
+
};
|
|
426
|
+
}
|
|
427
|
+
}
|
|
428
|
+
if (config.detectBlockRepeats) {
|
|
429
|
+
const blk = findRepetitiveBlock(command, config);
|
|
430
|
+
if (blk) {
|
|
431
|
+
return {
|
|
432
|
+
block: true,
|
|
433
|
+
reason: `[antiloop] blocked: ${blockRepeatDescription({ ...blk, where: "bash command" })} — this is stuck generation, not a real command. Do NOT retry it: stop and take one small, concrete step instead.`,
|
|
434
|
+
};
|
|
435
|
+
}
|
|
436
|
+
}
|
|
391
437
|
});
|
|
392
438
|
|
|
393
439
|
pi.on("turn_end", async (event, ctx) => {
|
package/src/types.ts
CHANGED
|
@@ -34,8 +34,30 @@ export interface AntiloopConfig {
|
|
|
34
34
|
degenerateMaxShare: number;
|
|
35
35
|
/** How many consecutive-detection points ONE degenerate turn adds (2 = warn on first sight). */
|
|
36
36
|
degenerateTurnWeight: number;
|
|
37
|
-
/** Block a bash tool call whose command shows degenerate repetition BEFORE it executes. */
|
|
37
|
+
/** Block a bash tool call whose command shows degenerate OR block repetition BEFORE it executes. */
|
|
38
38
|
blockDegenerateBash: boolean;
|
|
39
|
+
/**
|
|
40
|
+
* Intra-message BLOCK repetition (v1.7): the "narration loop" class where ONE
|
|
41
|
+
* assistant message keeps replaying the same sentences/phrases (the real S0
|
|
42
|
+
* coding session: ~5 near-verbatim cycles of "Let me start by checking the
|
|
43
|
+
* environment…" inside a SINGLE generation, which every cross-message detector
|
|
44
|
+
* misses because it produced only one message and the degenerate detector only
|
|
45
|
+
* watches for ONE word repeated hundreds of times). A sliding window of
|
|
46
|
+
* blockNgram-word n-grams over the normalized payload is built; when at least
|
|
47
|
+
* blockRepeatShare of those positions recur, one n-gram appears at least
|
|
48
|
+
* blockMinRepeats times and the payload has at least blockMinTokens words, the
|
|
49
|
+
* message is a replay. Fires at message_end (before the message's tools run)
|
|
50
|
+
* with the same strong turn weight as the degenerate detector.
|
|
51
|
+
*/
|
|
52
|
+
detectBlockRepeats: boolean;
|
|
53
|
+
/** Minimum normalized words in a payload before it is scanned for block repetition. */
|
|
54
|
+
blockMinTokens: number;
|
|
55
|
+
/** Word window of the n-grams whose recurrence is measured (default 5). */
|
|
56
|
+
blockNgram: number;
|
|
57
|
+
/** The MOST repeated n-gram must occur at least this many times to flag. */
|
|
58
|
+
blockMinRepeats: number;
|
|
59
|
+
/** Share of n-gram positions that must recur for the payload to count as a replay. */
|
|
60
|
+
blockRepeatShare: number;
|
|
39
61
|
/**
|
|
40
62
|
* No-progress outcome runs (v1.6.1): a model that keeps re-running the SAME
|
|
41
63
|
* experiment with cosmetic mutations (labels/permutations) while the outcome
|
|
@@ -87,6 +109,18 @@ export interface AntiloopConfig {
|
|
|
87
109
|
detectTaskStreams: boolean;
|
|
88
110
|
taskStreamMinCalls: number;
|
|
89
111
|
taskStreamTwinThreshold: number;
|
|
112
|
+
/**
|
|
113
|
+
* Read-only snapshot/status tools (default: ["trimegisto_harvest"]). These
|
|
114
|
+
* return a view of changing state (a list of agents, a queue, a dashboard).
|
|
115
|
+
* Calling one again with the SAME arguments is an idempotent read — most
|
|
116
|
+
* strikingly when a coordinator closes N agents that already settled and the
|
|
117
|
+
* serialized snapshot is byte-identical every time. That is the serial close
|
|
118
|
+
* of finished agents, not a reasoning loop, so repeated calls to a snapshot
|
|
119
|
+
* tool never count as a tool loop (nor as a no-progress outcome run). Names
|
|
120
|
+
* are matched exactly against the tool name; add any other status/poll tool
|
|
121
|
+
* here to keep antiloop quiet while it is read repeatedly.
|
|
122
|
+
*/
|
|
123
|
+
snapshotTools: string[];
|
|
90
124
|
detectToolLoops: boolean;
|
|
91
125
|
detectThinkingLoops: boolean;
|
|
92
126
|
detectTextLoops: boolean;
|
|
@@ -106,7 +140,7 @@ export interface AntiloopConfig {
|
|
|
106
140
|
toggleShortcut: string;
|
|
107
141
|
}
|
|
108
142
|
|
|
109
|
-
export type LoopKind = "text" | "tool" | "thinking" | "structural" | "degenerate" | "outcome";
|
|
143
|
+
export type LoopKind = "text" | "tool" | "thinking" | "structural" | "degenerate" | "block" | "outcome";
|
|
110
144
|
|
|
111
145
|
export interface LoopDetection {
|
|
112
146
|
type: LoopKind;
|