pi-antiloop 1.8.0 → 1.8.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/LICENSE +1 -1
- package/README.md +15 -15
- package/package.json +3 -3
- package/src/commands.ts +1 -1
- package/src/detect.ts +30 -30
- package/src/index.ts +3 -3
- package/src/types.ts +1 -1
package/LICENSE
CHANGED
package/README.md
CHANGED
|
@@ -12,7 +12,7 @@
|
|
|
12
12
|
|
|
13
13
|
## Features
|
|
14
14
|
|
|
15
|
-
- **Seven detection strategies** — text repetition (trigram Jaccard + Levenshtein), tool-call sequences (name + near-identical arguments + same outcome — result-aware, so retries that make progress don't false-positive), thinking blocks, structural opening-phrase patterns, **degenerate repetition** (a single message/call stuck repeating one word hundreds of times — the `
|
|
15
|
+
- **Seven detection strategies** — text repetition (trigram Jaccard + Levenshtein), tool-call sequences (name + near-identical arguments + same outcome — result-aware, so retries that make progress don't false-positive), thinking blocks, structural opening-phrase patterns, **degenerate repetition** (a single message/call stuck repeating one word hundreds of times — the `lorem ×5145` meltdown — caught at `message_end` with no repeated peer needed, and degenerate bash commands blocked before they execute), **block repetition** (ONE message replaying whole sentences/phrases — the narration loop: `Let me start by checking the environment…` ×5 — same `message_end` timing and strong turn weight), and **no-progress outcome runs** (the NFS `test A…QQQ` class: dozens of near-identical re-runs of the same experiment, every one ending in the *same failing outcome* — args mutate so tool-loop can't see it; the repeated failure signature can)
|
|
16
16
|
- **Stays alive in long sessions** — tracked messages carry a monotonic sequence number, so trimming the sliding window can never make the dedupe guard skip later messages (a bug that silently blinded antiloop after ~15 tracked messages)
|
|
17
17
|
- **Task-stream recognition (batch work)** — when another extension (e.g. `punched` appending lines to pi.md, or `plan` adding tasks) makes the model call the *same* tool many times with *different* content, antiloop recognizes it as N distinct tasks of one type and stays silent — no warning, no force break. A genuine loop (the *same* call repeated verbatim) is still caught
|
|
18
18
|
- **Snapshot/read-only tool exemption** — a coordinator closing N agents that already finished calls `trimegisto_harvest` once per agent and gets the *same* settled board every time; identical reads of a status snapshot are an idempotent serial close, not a loop, so `snapshotTools` (configurable) is exempt from tool-loop and no-progress detection — real `bash` loops next to it are still caught
|
|
@@ -107,7 +107,7 @@ Detection strategies:
|
|
|
107
107
|
Recent detections:
|
|
108
108
|
[text] Text similarity 85% with message 3 (2m ago)
|
|
109
109
|
[tool] repeated 3x: bash (5m ago)
|
|
110
|
-
[degenerate] bash command: degenerate repetition — "
|
|
110
|
+
[degenerate] bash command: degenerate repetition — "lorem" ×5145 (5m ago)
|
|
111
111
|
[block] message text: block repetition — 100% of 450 tokens replay repeated 5-word phrases (5m ago)
|
|
112
112
|
[outcome] no progress: 9 near-identical bash attempts with the same failing outcome (5m ago)
|
|
113
113
|
```
|
|
@@ -131,7 +131,7 @@ Grouped interactive menu showing the current value in each option:
|
|
|
131
131
|
- **🔧 call similarity** — `99 / 95 / 90 / 80%` — how identical tool-call *arguments* must be to count as the same call (default 95%: only near-identical repeats loop)
|
|
132
132
|
- **🔁 call repeats** — `1 / 2 / 3` — how many times the same call must repeat before it flags (default 2)
|
|
133
133
|
- **🧾 result similarity** — `95 / 80 / 60%` — how similar captured results must be to count as the *same outcome*; a repeated command that starts producing a different result is progress, not a loop (default 80%)
|
|
134
|
-
- **🌀 degenerate run** — `8 / 16 / 24 / 48` — identical words in a row inside ONE message/call before it counts as stuck generation (default 16; the `
|
|
134
|
+
- **🌀 degenerate run** — `8 / 16 / 24 / 48` — identical words in a row inside ONE message/call before it counts as stuck generation (default 16; the `lorem ×5145` class)
|
|
135
135
|
- **🔁 block repetition** — on/off — detect ONE message replaying whole sentences/phrases (the narration loop: `Let me start by checking the environment…` ×5) with no repeated peer needed
|
|
136
136
|
- **🔁 block share** — `70 / 85 / 95%` — share of replayed 5-word phrases in one message before it counts as a replay (default 85%)
|
|
137
137
|
- **⛔ block repetitive bash** — on/off — refuse a degenerate or replayed bash command before it executes, feeding the reason back to the model (default on)
|
|
@@ -179,10 +179,10 @@ batch no detections → silent (exp silent — 98.9% args would match without
|
|
|
179
179
|
loop still detected → tool (exp tool — identical repeats are NOT a stream) ✅
|
|
180
180
|
loop survives batch gate→ tool (exp tool — bash repeats are real) ✅
|
|
181
181
|
stream needs ≥3 calls → no stream (exp no stream at 2 calls) ✅
|
|
182
|
-
degenerate first sight → degenerate (bash command: "
|
|
182
|
+
degenerate first sight → degenerate (bash command: "lorem" ×402 …) ✅ ← v1.6: ONE meltdown message, no peer
|
|
183
183
|
degenerate legit cmd → ok (exp ok) ✅
|
|
184
|
-
degenerate interleaved → flag (
|
|
185
|
-
degenerate glued token → flag (
|
|
184
|
+
degenerate interleaved → flag (lorem ×150) (exp flag — freq/share clause) ✅
|
|
185
|
+
degenerate glued token → flag (lorem ×300) (exp flag — perfect power) ✅
|
|
186
186
|
degenerate turn weight → degenerate 2, text 1 (exp 2, 1) ✅
|
|
187
187
|
block S0 narration loop → block (message text: 100% of 450 tokens replay repeated 5-word phrases) ✅ ← v1.7: ONE message replaying sentences, no peer
|
|
188
188
|
block 3 cycles / 2 cycles→ flag (100%) / no (exp flag / no — needs ≥3 repeats) ✅
|
|
@@ -255,8 +255,8 @@ results are known. Each captured result becomes a normalized tail fingerprint
|
|
|
255
255
|
("err|" / "ok|" prefix + last 400 chars, so PID/timestamp noise is tolerated).
|
|
256
256
|
If the same command produced a *different* outcome, the pair is progress:
|
|
257
257
|
|
|
258
|
-
[bash("...")] → err|error: invalid argument:
|
|
259
|
-
[bash("...")] → ok|model loaded / listening on :
|
|
258
|
+
[bash("...")] → err|error: invalid argument: GPU0 (attempt 1)
|
|
259
|
+
[bash("...")] → ok|model loaded / listening on :8080 (attempt 2)
|
|
260
260
|
→ NOT a loop — the retry fixed the problem
|
|
261
261
|
|
|
262
262
|
[bash("...")] → err|failed to create context … (attempt 1)
|
|
@@ -329,8 +329,8 @@ All strategies above need at least two similar messages — they detect a model
|
|
|
329
329
|
failure mode they cannot see: the model's decoder **anchors on a token and
|
|
330
330
|
stops producing new output**, repeating the same word hundreds of times
|
|
331
331
|
*inside a single message or tool call*. Real case (session `2026-09-09T15-43`,
|
|
332
|
-
|
|
333
|
-
username wordlist repeated `
|
|
332
|
+
a local llama.cpp server): one 46 KB `bash` call whose SSH
|
|
333
|
+
username wordlist repeated `lorem` **5145 times** (a run of 5140 — 99% of
|
|
334
334
|
the payload). Every cross-message detector stayed silent (nothing to compare
|
|
335
335
|
against — it happened exactly once), and only a manual ESC stopped it.
|
|
336
336
|
|
|
@@ -344,7 +344,7 @@ occurrence, with no peer message:
|
|
|
344
344
|
with ≥ `degenerateMaxShare` (40%) of all tokens catches interleaved
|
|
345
345
|
meltdowns (`A B A B A B…`) that have no long run.
|
|
346
346
|
- **Perfect power** — a single giant token with no separators at all
|
|
347
|
-
(`
|
|
347
|
+
(`loremlorem…`) is checked for periodicity.
|
|
348
348
|
|
|
349
349
|
Tokenization uses runs of letters (unicode), so JSON stringification noise
|
|
350
350
|
(escaped `\n` in stored tool args, punctuation, digits, code symbols) never
|
|
@@ -417,7 +417,7 @@ A subtler meltdown than the degenerate one: the model re-runs the SAME
|
|
|
417
417
|
experiment over and over, mutating a cosmetic label or permutation each time
|
|
418
418
|
so no call ever repeats verbatim — while the outcome stays the SAME FAILURE.
|
|
419
419
|
Real case (same session, rows 95–249): ~90 ssh `exportfs`/`mount` tests,
|
|
420
|
-
labels `test A` … `test QQQ`, targets alternating
|
|
420
|
+
labels `test A` … `test QQQ`, targets alternating data-a/data-b — every one
|
|
421
421
|
ending `access denied` / `rc=32`, with fresh journalctl noise per attempt.
|
|
422
422
|
The tool-loop detector is blind to it BY DESIGN (args mutate every turn:
|
|
423
423
|
mean adjacent trigram similarity 0.93, but the label always changes, so the
|
|
@@ -509,7 +509,7 @@ Persisted as JSON at `~/.pi/agent/antiloop.json`:
|
|
|
509
509
|
| `toolSimilarityThreshold` | `0.95` | How close tool-call arguments must be (0.0–1.0) to count as the *same* call — see [tool loops](#how-it-works) |
|
|
510
510
|
| `minToolRepeatCount` | `2` | Prior occurrences of a near-identical call set required before a tool loop is flagged (2 = same call seen 3×) |
|
|
511
511
|
| `resultSimilarityThreshold` | `0.8` | Minimum similarity between captured result tails to still count as the *same outcome*; below this, a repeated command is treated as progress, not a loop |
|
|
512
|
-
| `detectDegenerate` | `true` | Detect intra-message degenerate repetition — one message/call stuck repeating a single word (the `
|
|
512
|
+
| `detectDegenerate` | `true` | Detect intra-message degenerate repetition — one message/call stuck repeating a single word (the `lorem ×5145` class). Needs no repeated peer message; scanned at `message_end` and on every bash `tool_call` |
|
|
513
513
|
| `degenerateMinTokens` | `50` | Minimum normalized tokens in a payload before it is scanned for degenerate repetition (shorter payloads aren't conclusive) |
|
|
514
514
|
| `degenerateMaxRun` | `16` | Consecutive identical words inside one payload that flag it as degenerate |
|
|
515
515
|
| `degenerateMaxFreq` | `60` | One word's total occurrences (with `degenerateMaxShare` of the payload) that flags interleaved meltdowns |
|
|
@@ -546,7 +546,7 @@ Persisted as JSON at `~/.pi/agent/antiloop.json`:
|
|
|
546
546
|
4. **Per-strategy toggles** — if the model's reasoning legitimately repeats (e.g. it's working through a checklist), disable `thinking` detection and leave text/tool on.
|
|
547
547
|
5. **Watch the log** — `/antiloop log` shows what's actually triggering. If you see false positives, raise `similarityThreshold` instead of disabling the strategy entirely.
|
|
548
548
|
6. **Let user input clear state** — each user message decays the consecutive counter by 2, so a fresh prompt naturally resets without `/antiloop reset`.
|
|
549
|
-
7. **Degenerate detector needs no tuning for most setups** — a run of ≥ 16 identical words (or one word ≥ 40% of a ≥ 50-token payload) inside a single message is conclusive stuck generation; the 46 KB `
|
|
549
|
+
7. **Degenerate detector needs no tuning for most setups** — a run of ≥ 16 identical words (or one word ≥ 40% of a ≥ 50-token payload) inside a single message is conclusive stuck generation; the 46 KB `lorem ×5145` SSH-wordlist meltdown is caught on first sight (warning), its bash never executes (`blockDegenerateBash`), and a second consecutive meltdown gets the force-break steer. If a model legitimately writes repetitive payloads, raise `degenerateMaxRun` / `degenerateMaxFreq` via `/antiloop config` — don't disable the detector.
|
|
550
550
|
8. **Outcome detector catches mutated re-run loops** — a model that re-issues the same experiment with cosmetic changes (labels, permutations) while every attempt fails identically gets a warning after `outcomeMinRepeats` (8) same-failure attempts, a steer on the next, and a hard stop shortly after. Converging sweeps, evolving failures and repeated successes stay silent by design. Lower `outcomeMinRepeats` if you want earlier cutoffs.
|
|
551
551
|
9. **Block detector catches narration loops** — a single message replaying whole sentences (the `Let me start by checking the environment…` ×5 class) warns on first sight and escalates like a meltdown. It only fires when ≥ 85% of the message's 5-word phrases recur, so ordinary prose and templated logs/code stay silent. Raise `blockRepeatShare` (or `blockMinRepeats`) if you ever see a false positive.
|
|
552
552
|
10. **`/antiloop test`** — runs the real detection engine (text + tool-call + task-stream + degenerate + block + outcome regression cases) to verify calibration after any change.
|
|
@@ -581,4 +581,4 @@ Modular extension with zero external dependencies (only pi's bundled `@earendil-
|
|
|
581
581
|
|
|
582
582
|
## License
|
|
583
583
|
|
|
584
|
-
[MIT](LICENSE) ©
|
|
584
|
+
[MIT](LICENSE) © antiloop contributors
|
package/package.json
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "pi-antiloop",
|
|
3
|
-
"version": "1.8.
|
|
4
|
-
"description": "Antiloop: detect reasoning loops and force a break (warn → force → abort) across text, tool, thinking, structural, degenerate-repetition, block/narration-repetition, and no-progress outcome patterns. Degenerate (one message stuck repeating a word hundreds of times —
|
|
3
|
+
"version": "1.8.1",
|
|
4
|
+
"description": "Antiloop: detect reasoning loops and force a break (warn → force → abort) across text, tool, thinking, structural, degenerate-repetition, block/narration-repetition, and no-progress outcome patterns. Degenerate (one message stuck repeating a word hundreds of times — lorem ×5145) is caught at message_end and its bash blocked before executing. Block (ONE message replaying whole sentences/phrases — the 'Let me start by checking the environment…' ×5 narration loop, invisible to every cross-message detector because there is no peer) fires on the first such message with the same strong turn weight and blocks a replayed bash command. No-progress (the NFS test-A…QQQ class: ~90 mutated re-runs of the same experiment, every one failing identically) fires when the same failing outcome repeats ≥ outcomeMinRepeats times with near-identical args. Snapshot tools (trimegisto_harvest) are exempt: a coordinator closing N settled agents re-reads the SAME board serially, which is an idempotent read, not a loop — so isVerbatimRepeat never fires on it. Fixes antiloop going blind mid-session: tracked messages use a monotonic sequence so the trim of the sliding window can never collide turn indices. The force break is delivered mid-run (steer before the next LLM call) and if the model ignores it antiloop aborts the run — an autonomous tool loop always terminates. Result-aware tool-loop detection keeps sequential bash and converging sweeps quiet; task-stream recognition keeps punched/plan batch work quiet.",
|
|
5
5
|
"keywords": [
|
|
6
6
|
"pi-package",
|
|
7
7
|
"antiloop",
|
|
@@ -9,7 +9,7 @@
|
|
|
9
9
|
"reasoning",
|
|
10
10
|
"monitoring"
|
|
11
11
|
],
|
|
12
|
-
"author": "
|
|
12
|
+
"author": "antiloop contributors",
|
|
13
13
|
"license": "MIT",
|
|
14
14
|
"repository": {
|
|
15
15
|
"type": "git",
|
package/src/commands.ts
CHANGED
|
@@ -96,7 +96,7 @@ async function showConfigMenu(ctx: ExtensionCommandContext, rt: Runtime): Promis
|
|
|
96
96
|
{ value: "toolRepeat" as const, label: `🔁 call repeats: ${c.minToolRepeatCount}+`, description: "how many times the same call must repeat before it flags" },
|
|
97
97
|
{ value: "resultSim" as const, label: `🧾 result similarity: ${(c.resultSimilarityThreshold * 100).toFixed(0)}%`, description: "same command + different result = progress, not a loop" },
|
|
98
98
|
// ── 🌀 Degenerate (intra-message meltdown) ───────────────────────
|
|
99
|
-
{ value: "degRun" as const, label: `🌀 degenerate run: ≥ ${c.degenerateMaxRun}`, description: "identical words in a row inside ONE message/call before it counts as stuck generation (
|
|
99
|
+
{ value: "degRun" as const, label: `🌀 degenerate run: ≥ ${c.degenerateMaxRun}`, description: "identical words in a row inside ONE message/call before it counts as stuck generation (lorem ×5145 class)" },
|
|
100
100
|
{ value: "blk" as const, label: `🔁 block repetition: ${yn(c.detectBlockRepeats)}`, description: "ONE message replaying whole sentences/phrases (the narration loop: 'Let me start by checking…' ×5)" },
|
|
101
101
|
{ value: "blkShare" as const, label: `🔁 block share: ≥ ${(c.blockRepeatShare * 100).toFixed(0)}%`, description: "share of replayed 5-word phrases in one message before it counts as a replay" },
|
|
102
102
|
{ value: "degBlock" as const, label: `⛔ block repetitive bash: ${yn(c.blockDegenerateBash)}`, description: "stop a degenerate or replayed command before it executes (default on)" },
|
package/src/detect.ts
CHANGED
|
@@ -37,9 +37,9 @@ function opening(text: string, n = 10): string {
|
|
|
37
37
|
// ---------------------------------------------------------------------------
|
|
38
38
|
// Intra-message degenerate repetition (v1.6).
|
|
39
39
|
//
|
|
40
|
-
// The "
|
|
41
|
-
//
|
|
42
|
-
//
|
|
40
|
+
// The "lorem ×5145" class (verified against a real session: ONE 46 KB bash
|
|
41
|
+
// call whose argument list repeats a single word 5145 times, a run of 5140 —
|
|
42
|
+
// 99% of the payload). A model whose
|
|
43
43
|
// decoder anchors on a token stops producing NEW output: it repeats the same
|
|
44
44
|
// word hundreds of times INSIDE one message or tool call. The cross-message
|
|
45
45
|
// detectors (text / tool / thinking / structural) all need >= 2 similar
|
|
@@ -132,7 +132,7 @@ export function findDegenerateRepetition(
|
|
|
132
132
|
}
|
|
133
133
|
}
|
|
134
134
|
|
|
135
|
-
// No-space meltdown: one giant periodic token ("
|
|
135
|
+
// No-space meltdown: one giant periodic token ("loremlorem…" with all
|
|
136
136
|
// separators stripped) is a perfect power of a short motif. Independent of
|
|
137
137
|
// the token-count gate: a single 1200+ char token has no word runs at all.
|
|
138
138
|
if (raw.length <= 2) {
|
|
@@ -400,7 +400,7 @@ function failSig(fp: string): string | undefined {
|
|
|
400
400
|
/** Same-FAILURE comparison for the outcome detector. Prefers the embedded error
|
|
401
401
|
* signatures (they repeat across attempts of the same failure) with digits
|
|
402
402
|
* stripped — timestamps, PIDs and rc values are noise; the target words that
|
|
403
|
-
* legitimately vary between attempts (
|
|
403
|
+
* legitimately vary between attempts (e.g. distinct mount targets) survive but the
|
|
404
404
|
* threshold is looser than the veto threshold on purpose. Falls back to the
|
|
405
405
|
* full fingerprints when no signature is present. */
|
|
406
406
|
function sameFailure(a: string, b: string, threshold: number): boolean {
|
|
@@ -648,8 +648,8 @@ export function detectLoops(state: AntiloopState, config: AntiloopConfig): LoopD
|
|
|
648
648
|
// -------------------------------------------------------------------
|
|
649
649
|
// No-progress outcome runs (v1.6.1).
|
|
650
650
|
//
|
|
651
|
-
// The NFS-test session (
|
|
652
|
-
//
|
|
651
|
+
// The NFS-test session (rows 95–249): ~90 mutated re-runs of the SAME
|
|
652
|
+
// experiment (ssh/sudo/exportfs/mount, labels
|
|
653
653
|
// "test A"…"test QQQ"), every one failing identically (rc=32 / access
|
|
654
654
|
// denied). The tool-loop detector is blind to it BY DESIGN: args mutate
|
|
655
655
|
// every turn (mean adjacent trigram similarity 0.93, but the label always
|
|
@@ -777,12 +777,12 @@ export function runSelfTest(): string[] {
|
|
|
777
777
|
|
|
778
778
|
// --- tool calls (default thresholds: 95% args similarity, 2 prior repeats) ---
|
|
779
779
|
const common =
|
|
780
|
-
"cd /
|
|
781
|
-
"setsid ./
|
|
782
|
-
"-m /
|
|
783
|
-
"-dev
|
|
780
|
+
"cd /srv/work && ulimit -l unlimited 2>/dev/null; export GPU_WORKAROUND=1 ACCELERATOR_VISIBLE=0; " +
|
|
781
|
+
"setsid ./llama.cpp/build/bin/llama-server " +
|
|
782
|
+
"-m /srv/models/example-model-q4.gguf " +
|
|
783
|
+
"-dev GPU0 -ngl 999 -fa on -c 8192 -fit off -np 1 -sm row -ub 2048 " +
|
|
784
784
|
"--spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-n-min 0 --spec-draft-p-min 0.4 " +
|
|
785
|
-
"--reasoning off --jinja --host 127.0.0.1 --port
|
|
785
|
+
"--reasoning off --jinja --host 127.0.0.1 --port 8080 --no-webui";
|
|
786
786
|
const sweepRun1 = `${common} -b 2048 -ctk q8_0 -ctv turbo4 > /tmp/sweep-turbo4.log 2>&1 & echo $!; sleep 60; grep "model loaded" /tmp/sweep-turbo4.log`;
|
|
787
787
|
const sweepRun2 = `${common} -b 8192 -ctk f16 -ctv f16 > /tmp/sweep-b8192.log 2>&1 & echo $!; sleep 70; grep "model loaded" /tmp/sweep-b8192.log`;
|
|
788
788
|
|
|
@@ -798,9 +798,9 @@ export function runSelfTest(): string[] {
|
|
|
798
798
|
out.push(`tool empty lists → ${t4 ? "match" : "no match"} (exp no match) ${!t4 ? "✅" : "❌"}`);
|
|
799
799
|
|
|
800
800
|
// --- result veto: same command, different outcome = progress, not a loop ---
|
|
801
|
-
const rErr = resultFingerprint([{ type: "text", text: "error: invalid argument:
|
|
802
|
-
const rErr2 = resultFingerprint([{ type: "text", text: "error: invalid argument:
|
|
803
|
-
const rOk = resultFingerprint([{ type: "text", text: "model loaded\nserver is listening on http://127.0.0.1:
|
|
801
|
+
const rErr = resultFingerprint([{ type: "text", text: "error: invalid argument: GPU0\nPID 74970" }], true)!;
|
|
802
|
+
const rErr2 = resultFingerprint([{ type: "text", text: "error: invalid argument: GPU0\nPID 77788" }], true)!;
|
|
803
|
+
const rOk = resultFingerprint([{ type: "text", text: "model loaded\nserver is listening on http://127.0.0.1:8080" }], false)!;
|
|
804
804
|
const sameCmdSameOut = toolCallsSimilar(
|
|
805
805
|
[{ name: "bash", args: sweepRun1, result: rErr }],
|
|
806
806
|
[{ name: "bash", args: sweepRun1, result: rErr2 }],
|
|
@@ -896,13 +896,13 @@ export function runSelfTest(): string[] {
|
|
|
896
896
|
out.push(`structural 0.90 → ${isVerbatimRepeat(dl("structural", 0.9)) ? "yes" : "no"} (exp no) ${!isVerbatimRepeat(dl("structural", 0.9)) ? "✅" : "❌"}`);
|
|
897
897
|
|
|
898
898
|
// --- v1.6: intra-message degenerate repetition (single-message meltdown) ---
|
|
899
|
-
// Regression:
|
|
900
|
-
//
|
|
901
|
-
// cross-message detector needs a peer message and stayed silent; the
|
|
899
|
+
// Regression: ONE 46 KB bash call whose argument list repeats "lorem" 5145
|
|
900
|
+
// times (run of 5140) — verified against a real stuck-generation session.
|
|
901
|
+
// Every cross-message detector needs a peer message and stayed silent; the
|
|
902
902
|
// degenerate scan must fire on the FIRST such message, alone in the window.
|
|
903
903
|
const argJson = (cmd: string) => JSON.stringify({ command: cmd }); // stored args form
|
|
904
904
|
const meltdownCmd =
|
|
905
|
-
`echo "=== brute usernames
|
|
905
|
+
`echo "=== brute usernames ==="; for u in alice alice@ j bob bob@ root carol ${`lorem `.repeat(400)}; do :; done`;
|
|
906
906
|
const meltMsg = mk("", [{ name: "bash", args: argJson(meltdownCmd) }]);
|
|
907
907
|
const mDet = detectLoops(asState([meltMsg]), tcfg);
|
|
908
908
|
const mHit = mDet.find((d) => d.type === "degenerate");
|
|
@@ -922,17 +922,17 @@ export function runSelfTest(): string[] {
|
|
|
922
922
|
out.push(`degenerate legit cmd → ${leg1 || leg2 || leg3 ? "flag" : "ok"} (exp ok) ${!leg1 && !leg2 && !leg3 ? "✅" : "❌"}`);
|
|
923
923
|
|
|
924
924
|
// Interleaved meltdown ("A B A B…") has no long run — caught via freq/share.
|
|
925
|
-
const inter = findDegenerateRepetition("
|
|
925
|
+
const inter = findDegenerateRepetition("lorem ipsum ".repeat(150), tcfg);
|
|
926
926
|
out.push(`degenerate interleaved → ${inter ? `flag (${inter.token} ×${inter.freq})` : "no"} (exp flag — freq/share clause) ${inter ? "✅" : "❌"}`);
|
|
927
927
|
|
|
928
928
|
// Word list written ACROSS lines: separators are the literal "\n" escapes
|
|
929
929
|
// inside the stored JSON args — must still tokenize word by word.
|
|
930
|
-
const acrossLines = argJson("for u in " + "
|
|
930
|
+
const acrossLines = argJson("for u in " + "lorem\n".repeat(120) + "done");
|
|
931
931
|
const linesHit = findDegenerateRepetition(acrossLines, tcfg);
|
|
932
932
|
out.push(`degenerate across \n → ${linesHit ? `flag (${linesHit.token} ×${linesHit.freq})` : "no"} (exp flag — escaped newlines) ${linesHit ? "✅" : "❌"}`);
|
|
933
933
|
|
|
934
|
-
// No-space giant token ("
|
|
935
|
-
const glued = findDegenerateRepetition("
|
|
934
|
+
// No-space giant token ("lorem" glued) — perfect-power clause.
|
|
935
|
+
const glued = findDegenerateRepetition("lorem".repeat(300), tcfg);
|
|
936
936
|
out.push(`degenerate glued token → ${glued ? `flag (${glued.token} ×${glued.freq})` : "no"} (exp flag — perfect power) ${glued ? "✅" : "❌"}`);
|
|
937
937
|
|
|
938
938
|
// Turn weight: one degenerate turn = 2 consecutive points (warning on first
|
|
@@ -993,15 +993,15 @@ export function runSelfTest(): string[] {
|
|
|
993
993
|
|
|
994
994
|
// --- v1.6.1: no-progress outcome runs (mutated re-runs, same outcome) ---
|
|
995
995
|
// Regression: the NFS session — ~90 mutated re-runs of the SAME experiment
|
|
996
|
-
// (ssh exportfs/mount, labels test A…test QQQ, targets alternating
|
|
997
|
-
//
|
|
996
|
+
// (ssh exportfs/mount, labels test A…test QQQ, targets alternating
|
|
997
|
+
// data-a / data-b), every one failing rc=32. Args mutate each turn (same call
|
|
998
998
|
// never recurs → tool-loop silent) and journalctl noise varies per attempt,
|
|
999
999
|
// but the FAILURE SIGNATURE repeats: that IS the loop. Fires on the 9th
|
|
1000
1000
|
// attempt (8 prior same-failure matches ≥ outcomeMinRepeats).
|
|
1001
1001
|
const nfsCmd = (label: string, target: string) =>
|
|
1002
1002
|
argJson(
|
|
1003
|
-
`sshpass -p X ssh -o ConnectTimeout=10
|
|
1004
|
-
`cat > /etc/exports << EOF\n/
|
|
1003
|
+
`sshpass -p X ssh -o ConnectTimeout=10 user@host 'cd /tmp && echo X | sudo -S bash -c "echo --- test ${label}: rootdir=/volume/data, absolute paths, fsid=0 and 1, mount /${target} ---; ` +
|
|
1004
|
+
`cat > /etc/exports << EOF\n/volume/data *(rw,sync,no_subtree_check,fsid=0)\nEOF\nexportfs -ra\nsystemctl restart nfs-server\n` +
|
|
1005
1005
|
`mount -t nfs4 -o vers=4.2 127.0.0.1:/${target} /tmp/nfstest 2>&1; echo rc=\$?; journalctl -u nfs-mountd | tail -4"' 2>&1`,
|
|
1006
1006
|
);
|
|
1007
1007
|
const nfsFailFor = (target: string, sec: number) =>
|
|
@@ -1011,13 +1011,13 @@ export function runSelfTest(): string[] {
|
|
|
1011
1011
|
type: "text",
|
|
1012
1012
|
text:
|
|
1013
1013
|
`--- mount /${target} --- | mount.nfs4: access denied by server while mounting 127.0.0.1:/${target} rc=32 | ` +
|
|
1014
|
-
`Sep 09 18:${sec}
|
|
1014
|
+
`Sep 09 18:${sec} host systemd[1]: Started nfs-mountd.service (PID ${1000 + sec})`,
|
|
1015
1015
|
},
|
|
1016
1016
|
],
|
|
1017
1017
|
false,
|
|
1018
1018
|
)!;
|
|
1019
1019
|
const nfsOk = resultFingerprint([{ type: "text", text: "rc=0 | TARGET SOURCE FSTYPE | /tmp/nfstest 127.0.0.1:/ nfs4 rw,relatime" }], false)!;
|
|
1020
|
-
const targets = ["
|
|
1020
|
+
const targets = ["data-a", "data-b"];
|
|
1021
1021
|
const nfsMsgs = [..."ABCDEFGHI"].map((l, idx) =>
|
|
1022
1022
|
mk("", [{ name: "bash", args: nfsCmd(l, targets[idx % 2]), result: nfsFailFor(targets[idx % 2], 100 + idx) }]),
|
|
1023
1023
|
);
|
|
@@ -1082,7 +1082,7 @@ export function runSelfTest(): string[] {
|
|
|
1082
1082
|
out.push(`isFailResult gate → err| → ${isFailResult("err|boom") ? "fail" : "ok"}, rc32 → ${isFailResult("ok|rc32 denied") ? "fail" : "ok"}, rc0/ok → ${isFailResult("ok|rc 0 12 passed") ? "fail" : "ok"} (exp fail, fail, ok) ${isFailResult("err|boom") && isFailResult("ok|rc32 denied") && !isFailResult("ok|rc 0 12 passed") ? "✅" : "❌"}`);
|
|
1083
1083
|
|
|
1084
1084
|
// --- v1.8: snapshot/read-only tools are idempotent reads, not loops ---
|
|
1085
|
-
// Regression (real session
|
|
1085
|
+
// Regression (verified against a real session): a coordinator closed
|
|
1086
1086
|
// N agents that had ALREADY settled by calling trimegisto_harvest five times
|
|
1087
1087
|
// with `{}`; every call returned the SAME 2.3 KB cumulative snapshot (all
|
|
1088
1088
|
// agents done). Same args + same result ×5 → the tool detector force-broke
|
package/src/index.ts
CHANGED
|
@@ -18,7 +18,7 @@
|
|
|
18
18
|
*
|
|
19
19
|
* v1.6 adds the intra-message degenerate detector: a model whose decoder
|
|
20
20
|
* anchors on a token repeats it hundreds of times INSIDE one message / tool
|
|
21
|
-
* call (the real 46 KB bash "
|
|
21
|
+
* call (the real 46 KB bash "lorem ×5145" meltdown). That needs no peer
|
|
22
22
|
* message, so it is caught at message_end — BEFORE the tool calls execute —
|
|
23
23
|
* and degenerate bash commands are additionally blocked in the tool_call
|
|
24
24
|
* hook (blockDegenerateBash). One degenerate turn counts degenerateTurnWeight
|
|
@@ -294,7 +294,7 @@ export default function antiloopExtension(pi: ExtensionAPI) {
|
|
|
294
294
|
|
|
295
295
|
// v1.6/v1.7 — intra-message repetition is handled HERE, at message_end,
|
|
296
296
|
// because turn_end only fires after tool execution (too late to stop the
|
|
297
|
-
// 46 KB "
|
|
297
|
+
// 46 KB "lorem ×5000" bash from running) and an aborted generation may
|
|
298
298
|
// never reach turn_end at all. The signal is self-contained (one
|
|
299
299
|
// pathological payload, no peer message) so the full escalation ladder runs
|
|
300
300
|
// right now; the turn-guard makes the later turn_end pass skip this message
|
|
@@ -404,7 +404,7 @@ export default function antiloopExtension(pi: ExtensionAPI) {
|
|
|
404
404
|
|
|
405
405
|
/**
|
|
406
406
|
* v1.6 — degenerate bash gate. A meltdown message whose command repeats one
|
|
407
|
-
* word hundreds of times (the 46 KB "
|
|
407
|
+
* word hundreds of times (the 46 KB "lorem ×5145" argument-wordlist brute
|
|
408
408
|
* force) must NEVER execute: it is pure context burn at best, and a real
|
|
409
409
|
* brute-force / destructive repetition at worst. message_end already
|
|
410
410
|
* escalated it; here we block the actual call before it runs. The block
|
package/src/types.ts
CHANGED
|
@@ -17,7 +17,7 @@ export interface AntiloopConfig {
|
|
|
17
17
|
*/
|
|
18
18
|
ignoredSteerLimit: number;
|
|
19
19
|
/**
|
|
20
|
-
* Intra-message degenerate repetition (the "
|
|
20
|
+
* Intra-message degenerate repetition (the "lorem \u00d75145" meltdown class): a model
|
|
21
21
|
* stuck emitting the same token hundreds of times INSIDE one message / tool call.
|
|
22
22
|
* Unlike the other detectors it needs no peer message: one pathological payload is
|
|
23
23
|
* already conclusive. Fires on the first occurrence; each degenerate turn is weighted
|