pi-antiloop 1.7.0 → 1.8.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/LICENSE +1 -1
- package/README.md +41 -15
- package/package.json +3 -3
- package/src/commands.ts +15 -1
- package/src/config.ts +1 -0
- package/src/detect.ts +87 -31
- package/src/index.ts +3 -3
- package/src/types.ts +13 -1
package/LICENSE
CHANGED
package/README.md
CHANGED
|
@@ -12,9 +12,10 @@
|
|
|
12
12
|
|
|
13
13
|
## Features
|
|
14
14
|
|
|
15
|
-
- **Seven detection strategies** — text repetition (trigram Jaccard + Levenshtein), tool-call sequences (name + near-identical arguments + same outcome — result-aware, so retries that make progress don't false-positive), thinking blocks, structural opening-phrase patterns, **degenerate repetition** (a single message/call stuck repeating one word hundreds of times — the `
|
|
15
|
+
- **Seven detection strategies** — text repetition (trigram Jaccard + Levenshtein), tool-call sequences (name + near-identical arguments + same outcome — result-aware, so retries that make progress don't false-positive), thinking blocks, structural opening-phrase patterns, **degenerate repetition** (a single message/call stuck repeating one word hundreds of times — the `lorem ×5145` meltdown — caught at `message_end` with no repeated peer needed, and degenerate bash commands blocked before they execute), **block repetition** (ONE message replaying whole sentences/phrases — the narration loop: `Let me start by checking the environment…` ×5 — same `message_end` timing and strong turn weight), and **no-progress outcome runs** (the NFS `test A…QQQ` class: dozens of near-identical re-runs of the same experiment, every one ending in the *same failing outcome* — args mutate so tool-loop can't see it; the repeated failure signature can)
|
|
16
16
|
- **Stays alive in long sessions** — tracked messages carry a monotonic sequence number, so trimming the sliding window can never make the dedupe guard skip later messages (a bug that silently blinded antiloop after ~15 tracked messages)
|
|
17
17
|
- **Task-stream recognition (batch work)** — when another extension (e.g. `punched` appending lines to pi.md, or `plan` adding tasks) makes the model call the *same* tool many times with *different* content, antiloop recognizes it as N distinct tasks of one type and stays silent — no warning, no force break. A genuine loop (the *same* call repeated verbatim) is still caught
|
|
18
|
+
- **Snapshot/read-only tool exemption** — a coordinator closing N agents that already finished calls `trimegisto_harvest` once per agent and gets the *same* settled board every time; identical reads of a status snapshot are an idempotent serial close, not a loop, so `snapshotTools` (configurable) is exempt from tool-loop and no-progress detection — real `bash` loops next to it are still caught
|
|
18
19
|
- **Progressive intervention** — `warning` reminds the model to vary its approach; `force break` steers a real break message into the running agent before its next LLM call; `abort` stops the run entirely
|
|
19
20
|
- **Configurable thresholds** — independent dials for similarity cutoff, warning/force-break/abort counts, detection window, and which strategies are on
|
|
20
21
|
- **Sliding window** — only the last N messages are compared, so detection is O(N) in the window size, not in the full session
|
|
@@ -106,7 +107,7 @@ Detection strategies:
|
|
|
106
107
|
Recent detections:
|
|
107
108
|
[text] Text similarity 85% with message 3 (2m ago)
|
|
108
109
|
[tool] repeated 3x: bash (5m ago)
|
|
109
|
-
[degenerate] bash command: degenerate repetition — "
|
|
110
|
+
[degenerate] bash command: degenerate repetition — "lorem" ×5145 (5m ago)
|
|
110
111
|
[block] message text: block repetition — 100% of 450 tokens replay repeated 5-word phrases (5m ago)
|
|
111
112
|
[outcome] no progress: 9 near-identical bash attempts with the same failing outcome (5m ago)
|
|
112
113
|
```
|
|
@@ -130,7 +131,7 @@ Grouped interactive menu showing the current value in each option:
|
|
|
130
131
|
- **🔧 call similarity** — `99 / 95 / 90 / 80%` — how identical tool-call *arguments* must be to count as the same call (default 95%: only near-identical repeats loop)
|
|
131
132
|
- **🔁 call repeats** — `1 / 2 / 3` — how many times the same call must repeat before it flags (default 2)
|
|
132
133
|
- **🧾 result similarity** — `95 / 80 / 60%` — how similar captured results must be to count as the *same outcome*; a repeated command that starts producing a different result is progress, not a loop (default 80%)
|
|
133
|
-
- **🌀 degenerate run** — `8 / 16 / 24 / 48` — identical words in a row inside ONE message/call before it counts as stuck generation (default 16; the `
|
|
134
|
+
- **🌀 degenerate run** — `8 / 16 / 24 / 48` — identical words in a row inside ONE message/call before it counts as stuck generation (default 16; the `lorem ×5145` class)
|
|
134
135
|
- **🔁 block repetition** — on/off — detect ONE message replaying whole sentences/phrases (the narration loop: `Let me start by checking the environment…` ×5) with no repeated peer needed
|
|
135
136
|
- **🔁 block share** — `70 / 85 / 95%` — share of replayed 5-word phrases in one message before it counts as a replay (default 85%)
|
|
136
137
|
- **⛔ block repetitive bash** — on/off — refuse a degenerate or replayed bash command before it executes, feeding the reason back to the model (default on)
|
|
@@ -141,6 +142,9 @@ Grouped interactive menu showing the current value in each option:
|
|
|
141
142
|
- **📋 stream min calls** — `2 / 3 / 4 / 5` — same-tool calls required before a batch is recognized (default 3)
|
|
142
143
|
- **📋 twin threshold** — `99 / 95 / 90%` — calls more similar than this count as the *same task* repeated; one twin invalidates the batch and normal detection resumes (default 99%)
|
|
143
144
|
|
|
145
|
+
**📡 Snapshot tools** — read-only status boards (trimegisto_harvest, …) read serially are not a loop
|
|
146
|
+
- **📡 snapshot tools** — `trimegisto_harvest` / off — repeated identical reads of a settled board stay silent all the way through (default on; add any other status/poll tool in `antiloop.json`)
|
|
147
|
+
|
|
144
148
|
**🔍 Detectors**
|
|
145
149
|
- **📝 text** — on/off — detect repeated text messages
|
|
146
150
|
- **🔧 tools** — on/off — detect repeated tool calls
|
|
@@ -175,10 +179,10 @@ batch no detections → silent (exp silent — 98.9% args would match without
|
|
|
175
179
|
loop still detected → tool (exp tool — identical repeats are NOT a stream) ✅
|
|
176
180
|
loop survives batch gate→ tool (exp tool — bash repeats are real) ✅
|
|
177
181
|
stream needs ≥3 calls → no stream (exp no stream at 2 calls) ✅
|
|
178
|
-
degenerate first sight → degenerate (bash command: "
|
|
182
|
+
degenerate first sight → degenerate (bash command: "lorem" ×402 …) ✅ ← v1.6: ONE meltdown message, no peer
|
|
179
183
|
degenerate legit cmd → ok (exp ok) ✅
|
|
180
|
-
degenerate interleaved → flag (
|
|
181
|
-
degenerate glued token → flag (
|
|
184
|
+
degenerate interleaved → flag (lorem ×150) (exp flag — freq/share clause) ✅
|
|
185
|
+
degenerate glued token → flag (lorem ×300) (exp flag — perfect power) ✅
|
|
182
186
|
degenerate turn weight → degenerate 2, text 1 (exp 2, 1) ✅
|
|
183
187
|
block S0 narration loop → block (message text: 100% of 450 tokens replay repeated 5-word phrases) ✅ ← v1.7: ONE message replaying sentences, no peer
|
|
184
188
|
block 3 cycles / 2 cycles→ flag (100%) / no (exp flag / no — needs ≥3 repeats) ✅
|
|
@@ -251,8 +255,8 @@ results are known. Each captured result becomes a normalized tail fingerprint
|
|
|
251
255
|
("err|" / "ok|" prefix + last 400 chars, so PID/timestamp noise is tolerated).
|
|
252
256
|
If the same command produced a *different* outcome, the pair is progress:
|
|
253
257
|
|
|
254
|
-
[bash("...")] → err|error: invalid argument:
|
|
255
|
-
[bash("...")] → ok|model loaded / listening on :
|
|
258
|
+
[bash("...")] → err|error: invalid argument: GPU0 (attempt 1)
|
|
259
|
+
[bash("...")] → ok|model loaded / listening on :8080 (attempt 2)
|
|
256
260
|
→ NOT a loop — the retry fixed the problem
|
|
257
261
|
|
|
258
262
|
[bash("...")] → err|failed to create context … (attempt 1)
|
|
@@ -297,6 +301,27 @@ size needed before recognition), `taskStreamTwinThreshold` (how similar args
|
|
|
297
301
|
must be to count as *the same task* — lower it to treat near-duplicate
|
|
298
302
|
entries as loops again).
|
|
299
303
|
|
|
304
|
+
### Snapshot tools: the serial close of settled agents ≠ a loop
|
|
305
|
+
|
|
306
|
+
A coordinator that closes N agents which have already finished calls the
|
|
307
|
+
snapshot tool once per agent. When the agents are all settled, every one of
|
|
308
|
+
those calls returns the **same** serialized board — same arguments (`{}`),
|
|
309
|
+
same result. To the tool detector that is indistinguishable from a verbatim
|
|
310
|
+
loop: `same args + same outcome` repeated N times, so it force-broke a run
|
|
311
|
+
that was simply collecting finished results.
|
|
312
|
+
|
|
313
|
+
Re-reading a state snapshot is an *idempotent read*, not a stuck generation,
|
|
314
|
+
so tools listed in `snapshotTools` (default `["trimegisto_harvest"]`) are
|
|
315
|
+
exempt from tool-loop **and** no-progress outcome detection. The exemption is
|
|
316
|
+
deliberately narrow:
|
|
317
|
+
|
|
318
|
+
- it only applies when *every* call in the message is a snapshot tool — a
|
|
319
|
+
real `bash` loop re-run alongside a harvest is still flagged;
|
|
320
|
+
- it is name-based and configurable, so `antiloop.json` can list any other
|
|
321
|
+
status/poll tool (`"snapshotTools": ["trimegisto_harvest", "my_status"]`);
|
|
322
|
+
- clearing it (`"snapshotTools": []` or **📡 off**) restores normal detection,
|
|
323
|
+
where the very same 5× harvest payload is a genuine verbatim tool loop.
|
|
324
|
+
|
|
300
325
|
### Degenerate repetition: ONE message stuck on a word ≠ a reasoning loop
|
|
301
326
|
|
|
302
327
|
All strategies above need at least two similar messages — they detect a model
|
|
@@ -304,8 +329,8 @@ All strategies above need at least two similar messages — they detect a model
|
|
|
304
329
|
failure mode they cannot see: the model's decoder **anchors on a token and
|
|
305
330
|
stops producing new output**, repeating the same word hundreds of times
|
|
306
331
|
*inside a single message or tool call*. Real case (session `2026-09-09T15-43`,
|
|
307
|
-
|
|
308
|
-
username wordlist repeated `
|
|
332
|
+
a local llama.cpp server): one 46 KB `bash` call whose SSH
|
|
333
|
+
username wordlist repeated `lorem` **5145 times** (a run of 5140 — 99% of
|
|
309
334
|
the payload). Every cross-message detector stayed silent (nothing to compare
|
|
310
335
|
against — it happened exactly once), and only a manual ESC stopped it.
|
|
311
336
|
|
|
@@ -319,7 +344,7 @@ occurrence, with no peer message:
|
|
|
319
344
|
with ≥ `degenerateMaxShare` (40%) of all tokens catches interleaved
|
|
320
345
|
meltdowns (`A B A B A B…`) that have no long run.
|
|
321
346
|
- **Perfect power** — a single giant token with no separators at all
|
|
322
|
-
(`
|
|
347
|
+
(`loremlorem…`) is checked for periodicity.
|
|
323
348
|
|
|
324
349
|
Tokenization uses runs of letters (unicode), so JSON stringification noise
|
|
325
350
|
(escaped `\n` in stored tool args, punctuation, digits, code symbols) never
|
|
@@ -392,7 +417,7 @@ A subtler meltdown than the degenerate one: the model re-runs the SAME
|
|
|
392
417
|
experiment over and over, mutating a cosmetic label or permutation each time
|
|
393
418
|
so no call ever repeats verbatim — while the outcome stays the SAME FAILURE.
|
|
394
419
|
Real case (same session, rows 95–249): ~90 ssh `exportfs`/`mount` tests,
|
|
395
|
-
labels `test A` … `test QQQ`, targets alternating
|
|
420
|
+
labels `test A` … `test QQQ`, targets alternating data-a/data-b — every one
|
|
396
421
|
ending `access denied` / `rc=32`, with fresh journalctl noise per attempt.
|
|
397
422
|
The tool-loop detector is blind to it BY DESIGN (args mutate every turn:
|
|
398
423
|
mean adjacent trigram similarity 0.93, but the label always changes, so the
|
|
@@ -484,7 +509,7 @@ Persisted as JSON at `~/.pi/agent/antiloop.json`:
|
|
|
484
509
|
| `toolSimilarityThreshold` | `0.95` | How close tool-call arguments must be (0.0–1.0) to count as the *same* call — see [tool loops](#how-it-works) |
|
|
485
510
|
| `minToolRepeatCount` | `2` | Prior occurrences of a near-identical call set required before a tool loop is flagged (2 = same call seen 3×) |
|
|
486
511
|
| `resultSimilarityThreshold` | `0.8` | Minimum similarity between captured result tails to still count as the *same outcome*; below this, a repeated command is treated as progress, not a loop |
|
|
487
|
-
| `detectDegenerate` | `true` | Detect intra-message degenerate repetition — one message/call stuck repeating a single word (the `
|
|
512
|
+
| `detectDegenerate` | `true` | Detect intra-message degenerate repetition — one message/call stuck repeating a single word (the `lorem ×5145` class). Needs no repeated peer message; scanned at `message_end` and on every bash `tool_call` |
|
|
488
513
|
| `degenerateMinTokens` | `50` | Minimum normalized tokens in a payload before it is scanned for degenerate repetition (shorter payloads aren't conclusive) |
|
|
489
514
|
| `degenerateMaxRun` | `16` | Consecutive identical words inside one payload that flag it as degenerate |
|
|
490
515
|
| `degenerateMaxFreq` | `60` | One word's total occurrences (with `degenerateMaxShare` of the payload) that flags interleaved meltdowns |
|
|
@@ -503,6 +528,7 @@ Persisted as JSON at `~/.pi/agent/antiloop.json`:
|
|
|
503
528
|
| `detectTaskStreams` | `true` | Recognize homogeneous batch work (same tool called with distinct content — e.g. punched/plan/obsidian extensions) and stay silent; see [task streams](#task-streams-n-tasks-of-one-type--a-loop) |
|
|
504
529
|
| `taskStreamMinCalls` | `3` | Same-tool calls required inside the window before a task stream is recognized |
|
|
505
530
|
| `taskStreamTwinThreshold` | `0.99` | Arguments this similar (or identical) count as *the same task* — a twin invalidates the stream and re-enables normal loop detection |
|
|
531
|
+
| `snapshotTools` | `["trimegisto_harvest"]` | Read-only snapshot/status tools: repeated identical reads (the serial close of N settled agents) never count as a tool loop or no-progress run. Add any other status/poll tool by name |
|
|
506
532
|
| `detectTextLoops` | `true` | Detect full-text repetition |
|
|
507
533
|
| `detectToolLoops` | `true` | Detect tool-call sequence + argument repetition |
|
|
508
534
|
| `detectThinkingLoops` | `true` | Detect repeated thinking/reasoning content |
|
|
@@ -520,7 +546,7 @@ Persisted as JSON at `~/.pi/agent/antiloop.json`:
|
|
|
520
546
|
4. **Per-strategy toggles** — if the model's reasoning legitimately repeats (e.g. it's working through a checklist), disable `thinking` detection and leave text/tool on.
|
|
521
547
|
5. **Watch the log** — `/antiloop log` shows what's actually triggering. If you see false positives, raise `similarityThreshold` instead of disabling the strategy entirely.
|
|
522
548
|
6. **Let user input clear state** — each user message decays the consecutive counter by 2, so a fresh prompt naturally resets without `/antiloop reset`.
|
|
523
|
-
7. **Degenerate detector needs no tuning for most setups** — a run of ≥ 16 identical words (or one word ≥ 40% of a ≥ 50-token payload) inside a single message is conclusive stuck generation; the 46 KB `
|
|
549
|
+
7. **Degenerate detector needs no tuning for most setups** — a run of ≥ 16 identical words (or one word ≥ 40% of a ≥ 50-token payload) inside a single message is conclusive stuck generation; the 46 KB `lorem ×5145` SSH-wordlist meltdown is caught on first sight (warning), its bash never executes (`blockDegenerateBash`), and a second consecutive meltdown gets the force-break steer. If a model legitimately writes repetitive payloads, raise `degenerateMaxRun` / `degenerateMaxFreq` via `/antiloop config` — don't disable the detector.
|
|
524
550
|
8. **Outcome detector catches mutated re-run loops** — a model that re-issues the same experiment with cosmetic changes (labels, permutations) while every attempt fails identically gets a warning after `outcomeMinRepeats` (8) same-failure attempts, a steer on the next, and a hard stop shortly after. Converging sweeps, evolving failures and repeated successes stay silent by design. Lower `outcomeMinRepeats` if you want earlier cutoffs.
|
|
525
551
|
9. **Block detector catches narration loops** — a single message replaying whole sentences (the `Let me start by checking the environment…` ×5 class) warns on first sight and escalates like a meltdown. It only fires when ≥ 85% of the message's 5-word phrases recur, so ordinary prose and templated logs/code stay silent. Raise `blockRepeatShare` (or `blockMinRepeats`) if you ever see a false positive.
|
|
526
552
|
10. **`/antiloop test`** — runs the real detection engine (text + tool-call + task-stream + degenerate + block + outcome regression cases) to verify calibration after any change.
|
|
@@ -555,4 +581,4 @@ Modular extension with zero external dependencies (only pi's bundled `@earendil-
|
|
|
555
581
|
|
|
556
582
|
## License
|
|
557
583
|
|
|
558
|
-
[MIT](LICENSE) ©
|
|
584
|
+
[MIT](LICENSE) © antiloop contributors
|
package/package.json
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "pi-antiloop",
|
|
3
|
-
"version": "1.
|
|
4
|
-
"description": "Antiloop: detect reasoning loops and force a break (warn → force → abort) across text, tool, thinking, structural, degenerate-repetition, block/narration-repetition, and no-progress outcome patterns. Degenerate (one message stuck repeating a word hundreds of times —
|
|
3
|
+
"version": "1.8.1",
|
|
4
|
+
"description": "Antiloop: detect reasoning loops and force a break (warn → force → abort) across text, tool, thinking, structural, degenerate-repetition, block/narration-repetition, and no-progress outcome patterns. Degenerate (one message stuck repeating a word hundreds of times — lorem ×5145) is caught at message_end and its bash blocked before executing. Block (ONE message replaying whole sentences/phrases — the 'Let me start by checking the environment…' ×5 narration loop, invisible to every cross-message detector because there is no peer) fires on the first such message with the same strong turn weight and blocks a replayed bash command. No-progress (the NFS test-A…QQQ class: ~90 mutated re-runs of the same experiment, every one failing identically) fires when the same failing outcome repeats ≥ outcomeMinRepeats times with near-identical args. Snapshot tools (trimegisto_harvest) are exempt: a coordinator closing N settled agents re-reads the SAME board serially, which is an idempotent read, not a loop — so isVerbatimRepeat never fires on it. Fixes antiloop going blind mid-session: tracked messages use a monotonic sequence so the trim of the sliding window can never collide turn indices. The force break is delivered mid-run (steer before the next LLM call) and if the model ignores it antiloop aborts the run — an autonomous tool loop always terminates. Result-aware tool-loop detection keeps sequential bash and converging sweeps quiet; task-stream recognition keeps punched/plan batch work quiet.",
|
|
5
5
|
"keywords": [
|
|
6
6
|
"pi-package",
|
|
7
7
|
"antiloop",
|
|
@@ -9,7 +9,7 @@
|
|
|
9
9
|
"reasoning",
|
|
10
10
|
"monitoring"
|
|
11
11
|
],
|
|
12
|
-
"author": "
|
|
12
|
+
"author": "antiloop contributors",
|
|
13
13
|
"license": "MIT",
|
|
14
14
|
"repository": {
|
|
15
15
|
"type": "git",
|
package/src/commands.ts
CHANGED
|
@@ -62,6 +62,7 @@ async function showStatus(ctx: ExtensionCommandContext, rt: Runtime): Promise<vo
|
|
|
62
62
|
` block repeats: ${yn(rt.config.detectBlockRepeats)} (≥ ${(rt.config.blockRepeatShare * 100).toFixed(0)}% of ≥ ${rt.config.blockMinTokens} tokens replay ${rt.config.blockNgram}-grams ×${rt.config.blockMinRepeats}+)`,
|
|
63
63
|
` outcome: same failing result ≥ ${rt.config.outcomeMinRepeats} attempts (args ≥ ${(rt.config.outcomeArgSimilarity * 100).toFixed(0)}% sim, sig ≥ ${(rt.config.outcomeSigThreshold * 100).toFixed(0)}%)`,
|
|
64
64
|
` task streams: ${yn(rt.config.detectTaskStreams)} (min ${rt.config.taskStreamMinCalls} calls, twins ≥ ${(rt.config.taskStreamTwinThreshold * 100).toFixed(0)}%)`,
|
|
65
|
+
` snapshot tools: ${rt.config.snapshotTools?.length ? rt.config.snapshotTools.join(", ") : "off"} (repeated reads are not a loop)`,
|
|
65
66
|
"",
|
|
66
67
|
`detectors: text ${yn(rt.config.detectTextLoops)} · tool ${yn(rt.config.detectToolLoops)} · think ${yn(rt.config.detectThinkingLoops)} · degenerate ${yn(rt.config.detectDegenerate)} · block ${yn(rt.config.detectBlockRepeats)} · outcome ${yn(rt.config.detectOutcomeLoops)}`,
|
|
67
68
|
`footer: interactive ${yn(rt.config.interactiveFooter)} · toggle: ${rt.config.toggleShortcut}`,
|
|
@@ -95,7 +96,7 @@ async function showConfigMenu(ctx: ExtensionCommandContext, rt: Runtime): Promis
|
|
|
95
96
|
{ value: "toolRepeat" as const, label: `🔁 call repeats: ${c.minToolRepeatCount}+`, description: "how many times the same call must repeat before it flags" },
|
|
96
97
|
{ value: "resultSim" as const, label: `🧾 result similarity: ${(c.resultSimilarityThreshold * 100).toFixed(0)}%`, description: "same command + different result = progress, not a loop" },
|
|
97
98
|
// ── 🌀 Degenerate (intra-message meltdown) ───────────────────────
|
|
98
|
-
{ value: "degRun" as const, label: `🌀 degenerate run: ≥ ${c.degenerateMaxRun}`, description: "identical words in a row inside ONE message/call before it counts as stuck generation (
|
|
99
|
+
{ value: "degRun" as const, label: `🌀 degenerate run: ≥ ${c.degenerateMaxRun}`, description: "identical words in a row inside ONE message/call before it counts as stuck generation (lorem ×5145 class)" },
|
|
99
100
|
{ value: "blk" as const, label: `🔁 block repetition: ${yn(c.detectBlockRepeats)}`, description: "ONE message replaying whole sentences/phrases (the narration loop: 'Let me start by checking…' ×5)" },
|
|
100
101
|
{ value: "blkShare" as const, label: `🔁 block share: ≥ ${(c.blockRepeatShare * 100).toFixed(0)}%`, description: "share of replayed 5-word phrases in one message before it counts as a replay" },
|
|
101
102
|
{ value: "degBlock" as const, label: `⛔ block repetitive bash: ${yn(c.blockDegenerateBash)}`, description: "stop a degenerate or replayed command before it executes (default on)" },
|
|
@@ -105,6 +106,7 @@ async function showConfigMenu(ctx: ExtensionCommandContext, rt: Runtime): Promis
|
|
|
105
106
|
{ value: "streams" as const, label: `📋 task streams: ${yn(c.detectTaskStreams)}`, description: "batch work (punched_log / plan_manager / …) is not a loop" },
|
|
106
107
|
{ value: "streamMin" as const, label: `📋 stream min calls: ${c.taskStreamMinCalls}`, description: "calls of the same tool before a batch is recognized" },
|
|
107
108
|
{ value: "streamTwin" as const, label: `📋 twin threshold: ${(c.taskStreamTwinThreshold * 100).toFixed(0)}%`, description: "calls more similar than this = the same task repeated, not a batch" },
|
|
109
|
+
{ value: "snapshot" as const, label: `📡 snapshot tools: ${c.snapshotTools?.length ? c.snapshotTools.join(", ") : "off"}`, description: "read-only status tools (trimegisto_harvest): re-reading settled agents serially is not a loop" },
|
|
108
110
|
// ── 🔍 Detectors ────────────────────────────────────────────
|
|
109
111
|
{ value: "text" as const, label: `📝 text: ${yn(c.detectTextLoops)}`, description: "detect repeated text messages" },
|
|
110
112
|
{ value: "tool" as const, label: `🔧 tools: ${yn(c.detectToolLoops)}`, description: "detect repeated tool calls" },
|
|
@@ -277,6 +279,18 @@ async function showConfigMenu(ctx: ExtensionCommandContext, rt: Runtime): Promis
|
|
|
277
279
|
if (v !== undefined) { c.taskStreamTwinThreshold = v; saveConfig(c); ctx.ui.notify(`twin threshold: ${(v * 100).toFixed(0)}%`, "info"); }
|
|
278
280
|
break;
|
|
279
281
|
}
|
|
282
|
+
case "snapshot": {
|
|
283
|
+
const v = await selectFrom(ctx, "📡 snapshot / read-only tools (repeated identical reads are never a loop)", [
|
|
284
|
+
{ value: "default" as const, label: "🎯 trimegisto_harvest (default)", description: "serial close of settled agents: same args + same snapshot ×N stays silent" },
|
|
285
|
+
{ value: "off" as const, label: "🚫 off", description: "treat every tool identically (harvest can loop-flag again)" },
|
|
286
|
+
]);
|
|
287
|
+
if (v !== undefined) {
|
|
288
|
+
c.snapshotTools = v === "off" ? [] : ["trimegisto_harvest"];
|
|
289
|
+
saveConfig(c);
|
|
290
|
+
ctx.ui.notify(`snapshot tools: ${c.snapshotTools.length ? c.snapshotTools.join(", ") : "off"}`, "info");
|
|
291
|
+
}
|
|
292
|
+
break;
|
|
293
|
+
}
|
|
280
294
|
case "ignoredBreak": {
|
|
281
295
|
const v = await selectFrom(ctx, "🛑 stop after ignored break (identical repeats after the force break)", [
|
|
282
296
|
{ value: 1, label: "⚡ 1 (sensitive — one identical repeat after the break stops the run)" },
|
package/src/config.ts
CHANGED
package/src/detect.ts
CHANGED
|
@@ -37,9 +37,9 @@ function opening(text: string, n = 10): string {
|
|
|
37
37
|
// ---------------------------------------------------------------------------
|
|
38
38
|
// Intra-message degenerate repetition (v1.6).
|
|
39
39
|
//
|
|
40
|
-
// The "
|
|
41
|
-
//
|
|
42
|
-
//
|
|
40
|
+
// The "lorem ×5145" class (verified against a real session: ONE 46 KB bash
|
|
41
|
+
// call whose argument list repeats a single word 5145 times, a run of 5140 —
|
|
42
|
+
// 99% of the payload). A model whose
|
|
43
43
|
// decoder anchors on a token stops producing NEW output: it repeats the same
|
|
44
44
|
// word hundreds of times INSIDE one message or tool call. The cross-message
|
|
45
45
|
// detectors (text / tool / thinking / structural) all need >= 2 similar
|
|
@@ -132,7 +132,7 @@ export function findDegenerateRepetition(
|
|
|
132
132
|
}
|
|
133
133
|
}
|
|
134
134
|
|
|
135
|
-
// No-space meltdown: one giant periodic token ("
|
|
135
|
+
// No-space meltdown: one giant periodic token ("loremlorem…" with all
|
|
136
136
|
// separators stripped) is a perfect power of a short motif. Independent of
|
|
137
137
|
// the token-count gate: a single 1200+ char token has no word runs at all.
|
|
138
138
|
if (raw.length <= 2) {
|
|
@@ -400,7 +400,7 @@ function failSig(fp: string): string | undefined {
|
|
|
400
400
|
/** Same-FAILURE comparison for the outcome detector. Prefers the embedded error
|
|
401
401
|
* signatures (they repeat across attempts of the same failure) with digits
|
|
402
402
|
* stripped — timestamps, PIDs and rc values are noise; the target words that
|
|
403
|
-
* legitimately vary between attempts (
|
|
403
|
+
* legitimately vary between attempts (e.g. distinct mount targets) survive but the
|
|
404
404
|
* threshold is looser than the veto threshold on purpose. Falls back to the
|
|
405
405
|
* full fingerprints when no signature is present. */
|
|
406
406
|
function sameFailure(a: string, b: string, threshold: number): boolean {
|
|
@@ -495,6 +495,28 @@ export function detectTaskStreams(
|
|
|
495
495
|
return streams;
|
|
496
496
|
}
|
|
497
497
|
|
|
498
|
+
// ---------------------------------------------------------------------------
|
|
499
|
+
// Snapshot / read-only tools (v1.8).
|
|
500
|
+
//
|
|
501
|
+
// A coordinator closing N agents that already finished calls the snapshot tool
|
|
502
|
+
// once per agent and gets the SAME serialized board every time (all settled).
|
|
503
|
+
// Same args + same result looks exactly like a verbatim tool loop to the
|
|
504
|
+
// detector, but re-reading a state snapshot is an idempotent read — the serial
|
|
505
|
+
// close of finished work, not a stuck generation. Tools listed in
|
|
506
|
+
// config.snapshotTools (default: trimegisto_harvest) are therefore exempt from
|
|
507
|
+
// tool-loop and outcome detection. Any other status/poll tool can be added.
|
|
508
|
+
// ---------------------------------------------------------------------------
|
|
509
|
+
|
|
510
|
+
/** Set of tool names whose repeated identical reads must never count as a loop. */
|
|
511
|
+
function snapshotToolSet(config: AntiloopConfig): Set<string> {
|
|
512
|
+
return new Set((config.snapshotTools ?? []).map((n) => n.trim()).filter(Boolean));
|
|
513
|
+
}
|
|
514
|
+
|
|
515
|
+
/** True when every call in the set is a snapshot/status read. */
|
|
516
|
+
function allSnapshotCalls(calls: TrackedToolCall[] | undefined, snapshots: Set<string>): boolean {
|
|
517
|
+
return !!calls && calls.length > 0 && calls.every((c) => snapshots.has(c.name));
|
|
518
|
+
}
|
|
519
|
+
|
|
498
520
|
export function detectLoops(state: AntiloopState, config: AntiloopConfig): LoopDetection[] {
|
|
499
521
|
const out: LoopDetection[] = [];
|
|
500
522
|
const msgs = state.recentMessages;
|
|
@@ -543,6 +565,7 @@ export function detectLoops(state: AntiloopState, config: AntiloopConfig): LoopD
|
|
|
543
565
|
// detections that only involve batch messages are suppressed too — the
|
|
544
566
|
// model is doing N different tasks of the same type, not looping.
|
|
545
567
|
const streams = detectTaskStreams(win, config);
|
|
568
|
+
const snapshots = snapshotToolSet(config);
|
|
546
569
|
const batchAt = new Set<number>();
|
|
547
570
|
win.forEach((m, idx) => {
|
|
548
571
|
const calls = m.toolCalls;
|
|
@@ -596,7 +619,10 @@ export function detectLoops(state: AntiloopState, config: AntiloopConfig): LoopD
|
|
|
596
619
|
const lastCalls = last.toolCalls;
|
|
597
620
|
// A message whose calls are all task-stream tools is batch work — skip
|
|
598
621
|
// it entirely (the stream gate already proved the calls are distinct).
|
|
599
|
-
|
|
622
|
+
// Snapshot/status tools (trimegisto_harvest…) are also skipped: repeated
|
|
623
|
+
// identical reads of a settled board are the serial close of finished
|
|
624
|
+
// agents, not a loop.
|
|
625
|
+
if (lastCalls && lastCalls.length && !batchAt.has(msgs.length - 1) && !allSnapshotCalls(lastCalls, snapshots)) {
|
|
600
626
|
const matched: number[] = [];
|
|
601
627
|
for (let i = 0; i < win.length - 1; i++) {
|
|
602
628
|
const prev = win[i].toolCalls;
|
|
@@ -622,8 +648,8 @@ export function detectLoops(state: AntiloopState, config: AntiloopConfig): LoopD
|
|
|
622
648
|
// -------------------------------------------------------------------
|
|
623
649
|
// No-progress outcome runs (v1.6.1).
|
|
624
650
|
//
|
|
625
|
-
// The NFS-test session (
|
|
626
|
-
//
|
|
651
|
+
// The NFS-test session (rows 95–249): ~90 mutated re-runs of the SAME
|
|
652
|
+
// experiment (ssh/sudo/exportfs/mount, labels
|
|
627
653
|
// "test A"…"test QQQ"), every one failing identically (rc=32 / access
|
|
628
654
|
// denied). The tool-loop detector is blind to it BY DESIGN: args mutate
|
|
629
655
|
// every turn (mean adjacent trigram similarity 0.93, but the label always
|
|
@@ -654,7 +680,9 @@ export function detectLoops(state: AntiloopState, config: AntiloopConfig): LoopD
|
|
|
654
680
|
// produce the SAME OK outcome every call — they must never count. And
|
|
655
681
|
// batch messages must NOT be skipped here: a "bash ×N stream" with
|
|
656
682
|
// identical failures is exactly the no-progress loop to catch.
|
|
657
|
-
|
|
683
|
+
// Snapshot/status reads are excluded outright: an identical harvest
|
|
684
|
+
// snapshot is a state read, not a repeated failing experiment.
|
|
685
|
+
if (lc.result && isFailResult(lc.result) && !snapshots.has(lc.name)) {
|
|
658
686
|
let matches = 0;
|
|
659
687
|
for (let i = 0; i < win.length - 1; i++) {
|
|
660
688
|
const m = win[i];
|
|
@@ -749,12 +777,12 @@ export function runSelfTest(): string[] {
|
|
|
749
777
|
|
|
750
778
|
// --- tool calls (default thresholds: 95% args similarity, 2 prior repeats) ---
|
|
751
779
|
const common =
|
|
752
|
-
"cd /
|
|
753
|
-
"setsid ./
|
|
754
|
-
"-m /
|
|
755
|
-
"-dev
|
|
780
|
+
"cd /srv/work && ulimit -l unlimited 2>/dev/null; export GPU_WORKAROUND=1 ACCELERATOR_VISIBLE=0; " +
|
|
781
|
+
"setsid ./llama.cpp/build/bin/llama-server " +
|
|
782
|
+
"-m /srv/models/example-model-q4.gguf " +
|
|
783
|
+
"-dev GPU0 -ngl 999 -fa on -c 8192 -fit off -np 1 -sm row -ub 2048 " +
|
|
756
784
|
"--spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-n-min 0 --spec-draft-p-min 0.4 " +
|
|
757
|
-
"--reasoning off --jinja --host 127.0.0.1 --port
|
|
785
|
+
"--reasoning off --jinja --host 127.0.0.1 --port 8080 --no-webui";
|
|
758
786
|
const sweepRun1 = `${common} -b 2048 -ctk q8_0 -ctv turbo4 > /tmp/sweep-turbo4.log 2>&1 & echo $!; sleep 60; grep "model loaded" /tmp/sweep-turbo4.log`;
|
|
759
787
|
const sweepRun2 = `${common} -b 8192 -ctk f16 -ctv f16 > /tmp/sweep-b8192.log 2>&1 & echo $!; sleep 70; grep "model loaded" /tmp/sweep-b8192.log`;
|
|
760
788
|
|
|
@@ -770,9 +798,9 @@ export function runSelfTest(): string[] {
|
|
|
770
798
|
out.push(`tool empty lists → ${t4 ? "match" : "no match"} (exp no match) ${!t4 ? "✅" : "❌"}`);
|
|
771
799
|
|
|
772
800
|
// --- result veto: same command, different outcome = progress, not a loop ---
|
|
773
|
-
const rErr = resultFingerprint([{ type: "text", text: "error: invalid argument:
|
|
774
|
-
const rErr2 = resultFingerprint([{ type: "text", text: "error: invalid argument:
|
|
775
|
-
const rOk = resultFingerprint([{ type: "text", text: "model loaded\nserver is listening on http://127.0.0.1:
|
|
801
|
+
const rErr = resultFingerprint([{ type: "text", text: "error: invalid argument: GPU0\nPID 74970" }], true)!;
|
|
802
|
+
const rErr2 = resultFingerprint([{ type: "text", text: "error: invalid argument: GPU0\nPID 77788" }], true)!;
|
|
803
|
+
const rOk = resultFingerprint([{ type: "text", text: "model loaded\nserver is listening on http://127.0.0.1:8080" }], false)!;
|
|
776
804
|
const sameCmdSameOut = toolCallsSimilar(
|
|
777
805
|
[{ name: "bash", args: sweepRun1, result: rErr }],
|
|
778
806
|
[{ name: "bash", args: sweepRun1, result: rErr2 }],
|
|
@@ -802,6 +830,7 @@ export function runSelfTest(): string[] {
|
|
|
802
830
|
detectTextLoops: true, notifyOnDetection: true, maxHistoryEntries: 100,
|
|
803
831
|
detectionWindow: 10, interactiveFooter: true, toggleShortcut: "esc+a",
|
|
804
832
|
detectTaskStreams: true, taskStreamMinCalls: 3, taskStreamTwinThreshold: 0.99,
|
|
833
|
+
snapshotTools: ["trimegisto_harvest"],
|
|
805
834
|
detectDegenerate: true, degenerateMinTokens: 50, degenerateMaxRun: 16,
|
|
806
835
|
degenerateMaxFreq: 60, degenerateMaxShare: 0.4, degenerateTurnWeight: 2, blockDegenerateBash: true,
|
|
807
836
|
detectBlockRepeats: true, blockMinTokens: 120, blockNgram: 5, blockMinRepeats: 3, blockRepeatShare: 0.85,
|
|
@@ -867,13 +896,13 @@ export function runSelfTest(): string[] {
|
|
|
867
896
|
out.push(`structural 0.90 → ${isVerbatimRepeat(dl("structural", 0.9)) ? "yes" : "no"} (exp no) ${!isVerbatimRepeat(dl("structural", 0.9)) ? "✅" : "❌"}`);
|
|
868
897
|
|
|
869
898
|
// --- v1.6: intra-message degenerate repetition (single-message meltdown) ---
|
|
870
|
-
// Regression:
|
|
871
|
-
//
|
|
872
|
-
// cross-message detector needs a peer message and stayed silent; the
|
|
899
|
+
// Regression: ONE 46 KB bash call whose argument list repeats "lorem" 5145
|
|
900
|
+
// times (run of 5140) — verified against a real stuck-generation session.
|
|
901
|
+
// Every cross-message detector needs a peer message and stayed silent; the
|
|
873
902
|
// degenerate scan must fire on the FIRST such message, alone in the window.
|
|
874
903
|
const argJson = (cmd: string) => JSON.stringify({ command: cmd }); // stored args form
|
|
875
904
|
const meltdownCmd =
|
|
876
|
-
`echo "=== brute usernames
|
|
905
|
+
`echo "=== brute usernames ==="; for u in alice alice@ j bob bob@ root carol ${`lorem `.repeat(400)}; do :; done`;
|
|
877
906
|
const meltMsg = mk("", [{ name: "bash", args: argJson(meltdownCmd) }]);
|
|
878
907
|
const mDet = detectLoops(asState([meltMsg]), tcfg);
|
|
879
908
|
const mHit = mDet.find((d) => d.type === "degenerate");
|
|
@@ -893,17 +922,17 @@ export function runSelfTest(): string[] {
|
|
|
893
922
|
out.push(`degenerate legit cmd → ${leg1 || leg2 || leg3 ? "flag" : "ok"} (exp ok) ${!leg1 && !leg2 && !leg3 ? "✅" : "❌"}`);
|
|
894
923
|
|
|
895
924
|
// Interleaved meltdown ("A B A B…") has no long run — caught via freq/share.
|
|
896
|
-
const inter = findDegenerateRepetition("
|
|
925
|
+
const inter = findDegenerateRepetition("lorem ipsum ".repeat(150), tcfg);
|
|
897
926
|
out.push(`degenerate interleaved → ${inter ? `flag (${inter.token} ×${inter.freq})` : "no"} (exp flag — freq/share clause) ${inter ? "✅" : "❌"}`);
|
|
898
927
|
|
|
899
928
|
// Word list written ACROSS lines: separators are the literal "\n" escapes
|
|
900
929
|
// inside the stored JSON args — must still tokenize word by word.
|
|
901
|
-
const acrossLines = argJson("for u in " + "
|
|
930
|
+
const acrossLines = argJson("for u in " + "lorem\n".repeat(120) + "done");
|
|
902
931
|
const linesHit = findDegenerateRepetition(acrossLines, tcfg);
|
|
903
932
|
out.push(`degenerate across \n → ${linesHit ? `flag (${linesHit.token} ×${linesHit.freq})` : "no"} (exp flag — escaped newlines) ${linesHit ? "✅" : "❌"}`);
|
|
904
933
|
|
|
905
|
-
// No-space giant token ("
|
|
906
|
-
const glued = findDegenerateRepetition("
|
|
934
|
+
// No-space giant token ("lorem" glued) — perfect-power clause.
|
|
935
|
+
const glued = findDegenerateRepetition("lorem".repeat(300), tcfg);
|
|
907
936
|
out.push(`degenerate glued token → ${glued ? `flag (${glued.token} ×${glued.freq})` : "no"} (exp flag — perfect power) ${glued ? "✅" : "❌"}`);
|
|
908
937
|
|
|
909
938
|
// Turn weight: one degenerate turn = 2 consecutive points (warning on first
|
|
@@ -964,15 +993,15 @@ export function runSelfTest(): string[] {
|
|
|
964
993
|
|
|
965
994
|
// --- v1.6.1: no-progress outcome runs (mutated re-runs, same outcome) ---
|
|
966
995
|
// Regression: the NFS session — ~90 mutated re-runs of the SAME experiment
|
|
967
|
-
// (ssh exportfs/mount, labels test A…test QQQ, targets alternating
|
|
968
|
-
//
|
|
996
|
+
// (ssh exportfs/mount, labels test A…test QQQ, targets alternating
|
|
997
|
+
// data-a / data-b), every one failing rc=32. Args mutate each turn (same call
|
|
969
998
|
// never recurs → tool-loop silent) and journalctl noise varies per attempt,
|
|
970
999
|
// but the FAILURE SIGNATURE repeats: that IS the loop. Fires on the 9th
|
|
971
1000
|
// attempt (8 prior same-failure matches ≥ outcomeMinRepeats).
|
|
972
1001
|
const nfsCmd = (label: string, target: string) =>
|
|
973
1002
|
argJson(
|
|
974
|
-
`sshpass -p X ssh -o ConnectTimeout=10
|
|
975
|
-
`cat > /etc/exports << EOF\n/
|
|
1003
|
+
`sshpass -p X ssh -o ConnectTimeout=10 user@host 'cd /tmp && echo X | sudo -S bash -c "echo --- test ${label}: rootdir=/volume/data, absolute paths, fsid=0 and 1, mount /${target} ---; ` +
|
|
1004
|
+
`cat > /etc/exports << EOF\n/volume/data *(rw,sync,no_subtree_check,fsid=0)\nEOF\nexportfs -ra\nsystemctl restart nfs-server\n` +
|
|
976
1005
|
`mount -t nfs4 -o vers=4.2 127.0.0.1:/${target} /tmp/nfstest 2>&1; echo rc=\$?; journalctl -u nfs-mountd | tail -4"' 2>&1`,
|
|
977
1006
|
);
|
|
978
1007
|
const nfsFailFor = (target: string, sec: number) =>
|
|
@@ -982,13 +1011,13 @@ export function runSelfTest(): string[] {
|
|
|
982
1011
|
type: "text",
|
|
983
1012
|
text:
|
|
984
1013
|
`--- mount /${target} --- | mount.nfs4: access denied by server while mounting 127.0.0.1:/${target} rc=32 | ` +
|
|
985
|
-
`Sep 09 18:${sec}
|
|
1014
|
+
`Sep 09 18:${sec} host systemd[1]: Started nfs-mountd.service (PID ${1000 + sec})`,
|
|
986
1015
|
},
|
|
987
1016
|
],
|
|
988
1017
|
false,
|
|
989
1018
|
)!;
|
|
990
1019
|
const nfsOk = resultFingerprint([{ type: "text", text: "rc=0 | TARGET SOURCE FSTYPE | /tmp/nfstest 127.0.0.1:/ nfs4 rw,relatime" }], false)!;
|
|
991
|
-
const targets = ["
|
|
1020
|
+
const targets = ["data-a", "data-b"];
|
|
992
1021
|
const nfsMsgs = [..."ABCDEFGHI"].map((l, idx) =>
|
|
993
1022
|
mk("", [{ name: "bash", args: nfsCmd(l, targets[idx % 2]), result: nfsFailFor(targets[idx % 2], 100 + idx) }]),
|
|
994
1023
|
);
|
|
@@ -1052,5 +1081,32 @@ export function runSelfTest(): string[] {
|
|
|
1052
1081
|
out.push(`outcome identical OKs → ${okDet.some((d) => d.type === "outcome") ? "outcome" : "silent"} (exp silent — success repeats ≠ loop) ${!okDet.some((d) => d.type === "outcome") ? "✅" : "❌"}`);
|
|
1053
1082
|
out.push(`isFailResult gate → err| → ${isFailResult("err|boom") ? "fail" : "ok"}, rc32 → ${isFailResult("ok|rc32 denied") ? "fail" : "ok"}, rc0/ok → ${isFailResult("ok|rc 0 12 passed") ? "fail" : "ok"} (exp fail, fail, ok) ${isFailResult("err|boom") && isFailResult("ok|rc32 denied") && !isFailResult("ok|rc 0 12 passed") ? "✅" : "❌"}`);
|
|
1054
1083
|
|
|
1084
|
+
// --- v1.8: snapshot/read-only tools are idempotent reads, not loops ---
|
|
1085
|
+
// Regression (verified against a real session): a coordinator closed
|
|
1086
|
+
// N agents that had ALREADY settled by calling trimegisto_harvest five times
|
|
1087
|
+
// with `{}`; every call returned the SAME 2.3 KB cumulative snapshot (all
|
|
1088
|
+
// agents done). Same args + same result ×5 → the tool detector force-broke
|
|
1089
|
+
// the run. Re-reading a settled state board is the serial close of finished
|
|
1090
|
+
// agents, not a stuck generation: with the snapshot exemption it is silent.
|
|
1091
|
+
const harvestSnap = resultFingerprint(
|
|
1092
|
+
[{
|
|
1093
|
+
type: "text",
|
|
1094
|
+
text:
|
|
1095
|
+
"## Trimegisto harvest (instant snapshot)\n\n### ✅ t0a [Active] — done (101s)\nTask: Explora el sistema en busca de LM Studio y sus runtimes.\n\n```\n# Informe: LM Studio en el sistema\n...\n```\n*10 turns, ↑32825 ↓2429*\n\n_All agents settled._",
|
|
1096
|
+
}],
|
|
1097
|
+
false,
|
|
1098
|
+
)!;
|
|
1099
|
+
const harvestMsgs = [0, 1, 2, 3, 4].map(() =>
|
|
1100
|
+
mk("", [{ name: "trimegisto_harvest", args: "{}", result: harvestSnap }]),
|
|
1101
|
+
);
|
|
1102
|
+
const snapDet = detectLoops(asState(harvestMsgs), tcfg);
|
|
1103
|
+
out.push(`snapshot harvest ×5 → ${snapDet.length ? snapDet.map((d) => d.type).join(",") : "silent"} (exp silent — serial close of settled agents) ${snapDet.length === 0 ? "✅" : "❌"}`);
|
|
1104
|
+
|
|
1105
|
+
// Guard: the exemption is what silences it — the SAME payload with
|
|
1106
|
+
// snapshotTools off is a genuine verbatim tool loop.
|
|
1107
|
+
const noSnapCfg: AntiloopConfig = { ...tcfg, snapshotTools: [] };
|
|
1108
|
+
const snapLoud = detectLoops(asState(harvestMsgs), noSnapCfg);
|
|
1109
|
+
out.push(`snapshot off = loop → ${snapLoud.some((d) => d.type === "tool") ? "tool" : "no"} (exp tool — normal tools still loop) ${snapLoud.some((d) => d.type === "tool") ? "✅" : "❌"}`);
|
|
1110
|
+
|
|
1055
1111
|
return out;
|
|
1056
1112
|
}
|
package/src/index.ts
CHANGED
|
@@ -18,7 +18,7 @@
|
|
|
18
18
|
*
|
|
19
19
|
* v1.6 adds the intra-message degenerate detector: a model whose decoder
|
|
20
20
|
* anchors on a token repeats it hundreds of times INSIDE one message / tool
|
|
21
|
-
* call (the real 46 KB bash "
|
|
21
|
+
* call (the real 46 KB bash "lorem ×5145" meltdown). That needs no peer
|
|
22
22
|
* message, so it is caught at message_end — BEFORE the tool calls execute —
|
|
23
23
|
* and degenerate bash commands are additionally blocked in the tool_call
|
|
24
24
|
* hook (blockDegenerateBash). One degenerate turn counts degenerateTurnWeight
|
|
@@ -294,7 +294,7 @@ export default function antiloopExtension(pi: ExtensionAPI) {
|
|
|
294
294
|
|
|
295
295
|
// v1.6/v1.7 — intra-message repetition is handled HERE, at message_end,
|
|
296
296
|
// because turn_end only fires after tool execution (too late to stop the
|
|
297
|
-
// 46 KB "
|
|
297
|
+
// 46 KB "lorem ×5000" bash from running) and an aborted generation may
|
|
298
298
|
// never reach turn_end at all. The signal is self-contained (one
|
|
299
299
|
// pathological payload, no peer message) so the full escalation ladder runs
|
|
300
300
|
// right now; the turn-guard makes the later turn_end pass skip this message
|
|
@@ -404,7 +404,7 @@ export default function antiloopExtension(pi: ExtensionAPI) {
|
|
|
404
404
|
|
|
405
405
|
/**
|
|
406
406
|
* v1.6 — degenerate bash gate. A meltdown message whose command repeats one
|
|
407
|
-
* word hundreds of times (the 46 KB "
|
|
407
|
+
* word hundreds of times (the 46 KB "lorem ×5145" argument-wordlist brute
|
|
408
408
|
* force) must NEVER execute: it is pure context burn at best, and a real
|
|
409
409
|
* brute-force / destructive repetition at worst. message_end already
|
|
410
410
|
* escalated it; here we block the actual call before it runs. The block
|
package/src/types.ts
CHANGED
|
@@ -17,7 +17,7 @@ export interface AntiloopConfig {
|
|
|
17
17
|
*/
|
|
18
18
|
ignoredSteerLimit: number;
|
|
19
19
|
/**
|
|
20
|
-
* Intra-message degenerate repetition (the "
|
|
20
|
+
* Intra-message degenerate repetition (the "lorem \u00d75145" meltdown class): a model
|
|
21
21
|
* stuck emitting the same token hundreds of times INSIDE one message / tool call.
|
|
22
22
|
* Unlike the other detectors it needs no peer message: one pathological payload is
|
|
23
23
|
* already conclusive. Fires on the first occurrence; each degenerate turn is weighted
|
|
@@ -109,6 +109,18 @@ export interface AntiloopConfig {
|
|
|
109
109
|
detectTaskStreams: boolean;
|
|
110
110
|
taskStreamMinCalls: number;
|
|
111
111
|
taskStreamTwinThreshold: number;
|
|
112
|
+
/**
|
|
113
|
+
* Read-only snapshot/status tools (default: ["trimegisto_harvest"]). These
|
|
114
|
+
* return a view of changing state (a list of agents, a queue, a dashboard).
|
|
115
|
+
* Calling one again with the SAME arguments is an idempotent read — most
|
|
116
|
+
* strikingly when a coordinator closes N agents that already settled and the
|
|
117
|
+
* serialized snapshot is byte-identical every time. That is the serial close
|
|
118
|
+
* of finished agents, not a reasoning loop, so repeated calls to a snapshot
|
|
119
|
+
* tool never count as a tool loop (nor as a no-progress outcome run). Names
|
|
120
|
+
* are matched exactly against the tool name; add any other status/poll tool
|
|
121
|
+
* here to keep antiloop quiet while it is read repeatedly.
|
|
122
|
+
*/
|
|
123
|
+
snapshotTools: string[];
|
|
112
124
|
detectToolLoops: boolean;
|
|
113
125
|
detectThinkingLoops: boolean;
|
|
114
126
|
detectTextLoops: boolean;
|