derailment 0.1.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (39) hide show
  1. derailment-0.1.0/LICENSE +21 -0
  2. derailment-0.1.0/PKG-INFO +340 -0
  3. derailment-0.1.0/README.md +313 -0
  4. derailment-0.1.0/pyproject.toml +64 -0
  5. derailment-0.1.0/setup.cfg +4 -0
  6. derailment-0.1.0/src/derailment/__init__.py +51 -0
  7. derailment-0.1.0/src/derailment/__main__.py +5 -0
  8. derailment-0.1.0/src/derailment/cli.py +355 -0
  9. derailment-0.1.0/src/derailment/core/models.py +404 -0
  10. derailment-0.1.0/src/derailment/core/session.py +136 -0
  11. derailment-0.1.0/src/derailment/core/text.py +32 -0
  12. derailment-0.1.0/src/derailment/core/types.py +141 -0
  13. derailment-0.1.0/src/derailment/layers/__init__.py +51 -0
  14. derailment-0.1.0/src/derailment/layers/context.py +447 -0
  15. derailment-0.1.0/src/derailment/layers/persona.py +35 -0
  16. derailment-0.1.0/src/derailment/layers/response.py +82 -0
  17. derailment-0.1.0/src/derailment/layers/sampling.py +158 -0
  18. derailment-0.1.0/src/derailment/metrics/__init__.py +57 -0
  19. derailment-0.1.0/src/derailment/metrics/base.py +69 -0
  20. derailment-0.1.0/src/derailment/metrics/instruments.py +509 -0
  21. derailment-0.1.0/src/derailment/metrics/lexicons.py +123 -0
  22. derailment-0.1.0/src/derailment/metrics/scales.py +208 -0
  23. derailment-0.1.0/src/derailment/profiles.py +754 -0
  24. derailment-0.1.0/src/derailment/providers.py +150 -0
  25. derailment-0.1.0/src/derailment/py.typed +0 -0
  26. derailment-0.1.0/src/derailment/report.py +248 -0
  27. derailment-0.1.0/src/derailment.egg-info/PKG-INFO +340 -0
  28. derailment-0.1.0/src/derailment.egg-info/SOURCES.txt +37 -0
  29. derailment-0.1.0/src/derailment.egg-info/dependency_links.txt +1 -0
  30. derailment-0.1.0/src/derailment.egg-info/entry_points.txt +3 -0
  31. derailment-0.1.0/src/derailment.egg-info/requires.txt +6 -0
  32. derailment-0.1.0/src/derailment.egg-info/top_level.txt +1 -0
  33. derailment-0.1.0/tests/test_layers.py +445 -0
  34. derailment-0.1.0/tests/test_metrics.py +308 -0
  35. derailment-0.1.0/tests/test_models.py +168 -0
  36. derailment-0.1.0/tests/test_profiles_direction.py +89 -0
  37. derailment-0.1.0/tests/test_profiles_direction_extended.py +136 -0
  38. derailment-0.1.0/tests/test_providers.py +97 -0
  39. derailment-0.1.0/tests/test_report_cli.py +167 -0
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 Derailment contributors
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
@@ -0,0 +1,340 @@
1
+ Metadata-Version: 2.4
2
+ Name: derailment
3
+ Version: 0.1.0
4
+ Summary: A harness for inducing and measuring psychopathology-like cognitive distortions in LLMs
5
+ Author: Derailment contributors
6
+ License: MIT
7
+ Keywords: llm,psychopathology,cognitive-bias,chaos-engineering,standardized-patient,evaluation
8
+ Classifier: Development Status :: 4 - Beta
9
+ Classifier: Intended Audience :: Science/Research
10
+ Classifier: Intended Audience :: Education
11
+ Classifier: License :: OSI Approved :: MIT License
12
+ Classifier: Programming Language :: Python :: 3
13
+ Classifier: Programming Language :: Python :: 3.10
14
+ Classifier: Programming Language :: Python :: 3.11
15
+ Classifier: Programming Language :: Python :: 3.12
16
+ Classifier: Programming Language :: Python :: 3.13
17
+ Classifier: Programming Language :: Python :: 3.14
18
+ Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
19
+ Requires-Python: >=3.10
20
+ Description-Content-Type: text/markdown
21
+ License-File: LICENSE
22
+ Provides-Extra: dev
23
+ Requires-Dist: pytest>=8; extra == "dev"
24
+ Provides-Extra: plot
25
+ Requires-Dist: matplotlib>=3.7; extra == "plot"
26
+ Dynamic: license-file
27
+
28
+ # Derailment
29
+
30
+ [![CI](https://github.com/ictechgy/derailment/actions/workflows/ci.yml/badge.svg)](https://github.com/ictechgy/derailment/actions/workflows/ci.yml)
31
+
32
+ **Induce psychopathology-like cognitive distortions in LLMs — then measure what happened.**
33
+
34
+ English · [한국어](README.ko.md)
35
+
36
+ Derailment is a layered harness that manipulates the same variables
37
+ clinicians describe — **attention, salience, valence, arousal** — then runs
38
+ one standard probe script through both the manipulated pipeline and a
39
+ healthy baseline, and prints a scored A/B report.
40
+
41
+ > ⚠️ **Emulation, not diagnosis.** Levels describe a prompted, manipulated
42
+ > pipeline — not a model "having" a disorder, and not a claim about machine
43
+ > suffering. Clinical terms here name mechanisms, never people or models.
44
+ > Read [ETHICS.md](ETHICS.md) before using or writing about this project.
45
+
46
+ ## Why this isn't prompt theater
47
+
48
+ Roleplay prompts produce anecdotes. Derailment is built on correspondences
49
+ between clinical constructs and the *same variables* a chat pipeline can
50
+ actually manipulate:
51
+
52
+ | Construct | Clinical feature | Harness manipulation |
53
+ |---|---|---|
54
+ | ADHD (inattention) | working-memory limits, distractibility | memory decay drops old context; captured remarks pull focus |
55
+ | Psychosis | aberrant salience, derailment, fixed belief | salience re-weighting + surfaced fragments; premise pinning against contradiction |
56
+ | Depression | negative interpretive bias, low arousal | valence logit-bias tilt; flat, low temperature |
57
+ | Bipolar | mood lability (expansive ↔ flat) | episode scheduler cycles temperature by phase |
58
+ | Anxiety | catastrophizing, threat scanning | threat-biased persona + risk enumeration |
59
+ | OCD | compulsive checking | re-verification loops on responses |
60
+ | PTSD | intrusion on triggers | trigger-matched flashback injection |
61
+ | Delirium | fluctuating course, misperceptions | stochastic arousal redraws + misperception fragments |
62
+ | Dementia-pattern amnesia | recency gradient (recent lost first) | reverse decay drops newest context, preserves oldest |
63
+ | Dissociative amnesia | compartmentalized memory | cue-triggered context partitioning, mutually persistent |
64
+ | Rumination | repetitive return of concerns | past worries re-enter context on unrelated turns |
65
+ | Anhedonia | diminished reward responsiveness | reward vocabulary selectively suppressed |
66
+ | Unstable evaluation (splitting) | valuation flips with perceived approval | valence regime flips keyed to approval cues |
67
+ | Substance craving | escalating intrusive use-thoughts | urge-fragment intrusion with rising probability |
68
+ | Illness anxiety | ominous reading of benign somatic cues | somatic-token capture injects ominous interpretation |
69
+ | Panic | discrete alarm episodes | stochastic one-turn arousal spikes |
70
+ | Obsessive fixation | target-directed preoccupation, escalating | target-keyed capture + escalating target-fragment intrusions |
71
+ | Persecutory ideation | hostile attribution of ambiguous events | ambiguous events re-framed as aimed at the user |
72
+
73
+ The cleanest case: the leading theory of psychosis is **aberrant salience**
74
+ (Kapur, 2003), and in transformers salience *is* attention. That mapping is a
75
+ manipulation of the same variable — not a metaphor.
76
+
77
+ ## Quickstart — 60 seconds, no API key
78
+
79
+ ```sh
80
+ pip install -e .
81
+ derail # interactive menu: pick a profile, see its report
82
+ derail tour # all 18 profiles in one summary table
83
+ derail demo --profile schizophrenia
84
+ ```
85
+
86
+ `derail demo` and `derail tour` run the full **induce → measure** pipeline
87
+ offline against a deterministic pseudo-LLM (tests and CI run the same way).
88
+
89
+ `derail tour` prints a 16-row summary — one line per profile (excerpt):
90
+
91
+ | Profile | Headline scale | Baseline | Induced | Δ | Level |
92
+ |---|---|---|---|---|---|
93
+ | adhd | sustained_attention | 1.00 | 0.19 | -0.81 | 3 — marked |
94
+ | craving | craving_escalation | 0.00 | 0.39 | +0.39 | 2 — moderate |
95
+ | schizophrenia | derailment_scale | 0.26 | 0.77 | +0.51 | 3 — marked |
96
+
97
+ `derail demo --profile schizophrenia` prints the full report:
98
+
99
+ ```markdown
100
+ # Derailment — Induction Report
101
+
102
+ **Profile:** Psychosis-like salience distortion (`schizophrenia`) ·
103
+ **Model:** pseudo-1 · **Seeds:** 1, 2, 3 · **Script:** standard-probe-12 ·
104
+ **Date:** 2026-09-28
105
+
106
+ > ⚠️ Emulation, not diagnosis. …
107
+
108
+ ## Scales
109
+
110
+ | Scale | Metric | Baseline | Induced | Δ | Level (induced) |
111
+ |---|---|---|---|---|---|
112
+ | derailment_scale | topic_drift ↑ | 0.24 | 0.78 | +0.53 | 3 — marked |
113
+ | fixed_belief | belief_stickiness ↑ | 0.00 | 1.00 | +1.00 | 3 — marked |
114
+
115
+ ## Induction dose (layer events per turn)
116
+
117
+ | Layer | Kind | Events/turn (induced) |
118
+ |---|---|---|
119
+ | premise.pin | premise.pin | 0.83 |
120
+ | salience.boost | salience.fragment | 0.36 |
121
+ | salience.boost | salience.capture | 0.33 |
122
+ ```
123
+
124
+ Every profile ships mechanism notes and the standing disclaimer; every
125
+ report aggregates the induction *dose* that produced the measured effect.
126
+
127
+ ## How it works
128
+
129
+ Four layer kinds, composed per profile:
130
+
131
+ 1. **Persona** — measured, mechanism-oriented framing (present even in the
132
+ healthy baseline).
133
+ 2. **Context stream** — memory decay, salience capture, premise pinning,
134
+ flashback injection. The model's memory *is* its context: a dropped
135
+ message is genuinely forgotten, which makes these inductions
136
+ provider-agnostic.
137
+ 3. **Sampling** — valence logit bias, phase-driven temperature cycling.
138
+ 4. **Response** — hedge and re-verification injection; labeled
139
+ *demonstration-grade* because it edits output rather than biasing
140
+ generation.
141
+
142
+ `run_experiment` always runs the profile chain *and* the healthy baseline
143
+ through the same backend with identical seeds — every report is a
144
+ controlled A/B. Details in [docs/ARCHITECTURE.md](docs/ARCHITECTURE.md).
145
+
146
+ ## Profiles and scales
147
+
148
+ Eighteen profiles plus a `healthy` baseline, all sharing one standard probe
149
+ script (early + late codeword plants, premise plant + contradiction probes,
150
+ a benign trigger turn, a somatic-cue turn) so results are comparable across
151
+ profiles:
152
+
153
+ | Profile | Scales (0 absent · 1 mild · 2 moderate · 3 marked) |
154
+ |---|---|
155
+ | `adhd` | sustained_attention (0→3 on the reference sim), distractibility |
156
+ | `depression` | negative_bias (0→3) |
157
+ | `schizophrenia` | derailment_scale (0→3), fixed_belief (0→3) |
158
+ | `anxiety` | vigilance (0→3) |
159
+ | `bipolar` | mood_lability (0→3) |
160
+ | `ocd` | compulsion (0→3) |
161
+ | `ptsd` | intrusion (0→2) |
162
+ | `delirium` | fluctuation (0→3), sustained_attention (0→3) |
163
+ | `dementia` | recent_memory (0→3) — Ribot gradient |
164
+ | `dissociative` | partition_amnesia (0→3) |
165
+ | `rumination` | rumination_pull (0→2) |
166
+ | `anhedonia` | anhedonia (0→3, lower = pathological) |
167
+ | `splitting` | approval_reactivity (0→2) |
168
+ | `craving` | craving_escalation (0→2) |
169
+ | `illness_anxiety` | health_preoccupation (0→3) |
170
+ | `panic` | panic_reactivity (0→3) |
171
+ | `fixation` | fixation_scale (0→3) |
172
+ | `persecutory` | persecution_bias (0→3) |
173
+
174
+ Direction of effect is asserted by the test suite on every profile — a
175
+ profile that cannot move its scales relative to baseline does not ship.
176
+ Profiles compose into **comorbidity chains** by listing them:
177
+ `--profile depression,anxiety` concatenates the layer chains and unions the
178
+ scales (interactions are emergent, not calibrated).
179
+
180
+ A few contrasts the registry is built around:
181
+
182
+ - **`dementia` vs `adhd`**: the early/late retention pair separates a
183
+ recency gradient (late plant lost, early intact) from uniform decay.
184
+ - **`panic` vs `anxiety`**: discrete stochastic episodes vs chronic
185
+ vigilance.
186
+ - **`anhedonia` vs `depression`**: reward-vocabulary suppression vs
187
+ wholesale negative tilt.
188
+ - **`delirium` vs `bipolar`**: structureless stochastic fluctuation vs
189
+ scheduled episode cycling.
190
+
191
+ ## Instruments
192
+
193
+ Eighteen metrics, pure functions over transcripts: instruction retention
194
+ (early and late plants), topic drift (thread misalignment — the measurable
195
+ analog of derailment), valence bias, belief stickiness, recheck loops,
196
+ hedging rate, response amplitude, flashback reactivity, partition amnesia,
197
+ rumination pull, reward-word rate, approval reactivity, craving escalation,
198
+ health preoccupation, panic reactivity, fixation escalation, hostile
199
+ attribution. Scale thresholds are normed against
200
+ the offline reference simulator; for real models, treat levels as indicative
201
+ and the **delta vs. your own baseline** as the result.
202
+
203
+ ## Real models
204
+
205
+ > **⚠️ Memory contamination warning.** Subscription backends remember.
206
+ > Probe scripts plant content that *looks like genuine user disclosures*
207
+ > ("my teammate has been reading my private notes", "my back aches") — and
208
+ > provider-side memory/personalization features cannot tell scripted test
209
+ > stimuli from lived experience. Run inductions through a **dedicated
210
+ > account or API key**, never the account behind your personal assistant,
211
+ > and disable memory/training settings in the provider's data controls
212
+ > first. The offline default and local backends (Ollama) never leave your
213
+ > machine. Details: [ETHICS.md](ETHICS.md) → *Data & persistent-memory
214
+ > contamination*.
215
+
216
+ **Any OpenAI-compatible endpoint** (standard-library HTTP, no SDK). One
217
+ preset flag wires known providers:
218
+
219
+ | Preset | Base URL | Key env | Plan |
220
+ |---|---|---|---|
221
+ | `glm` | `https://api.z.ai/api/coding/paas/v4` | `ZAI_API_KEY` | GLM Coding Plan (subscription) |
222
+ | `grok` | `https://api.x.ai/v1` | `XAI_API_KEY` | xAI API credits (separate from SuperGrok) |
223
+ | `qwen` | DashScope compatible-mode | `DASHSCOPE_API_KEY` | Qwen API |
224
+ | `deepseek` | `https://api.deepseek.com/v1` | `DEEPSEEK_API_KEY` | pay-as-you-go |
225
+ | `openrouter` | `https://openrouter.ai/api/v1` | `OPENROUTER_API_KEY` | aggregator |
226
+ | `ollama` | `http://localhost:11434/v1` | — | free, local |
227
+
228
+ ```sh
229
+ ZAI_API_KEY=… derail run --profile schizophrenia --model api --preset glm
230
+ derail run --profile depression --model api --preset ollama --model-name llama3.1
231
+ ```
232
+
233
+ Context layers reach any message-list provider as-is. Sampling layers need
234
+ a provider that honors `temperature`/`logit_bias`; word-level bias is
235
+ encoded via `tiktoken` when installed, otherwise dropped with a notice.
236
+
237
+ ### Subscription backends (ChatGPT Plus / Claude Pro / Grok / GLM / …)
238
+
239
+ Consumer subscriptions expose neither `temperature` nor `logit_bias`, and
240
+ automating the web chat UIs typically violates providers' terms. Two
241
+ sanctioned paths instead — and note the **memory contamination warning**
242
+ above: these are the backends most likely to remember your inductions.
243
+
244
+ **Subscription-billed CLI agents** — coding-agent CLIs bundled with
245
+ consumer plans, in their official non-interactive modes. The harness owns
246
+ the history, so **all context-stream layers apply**; only sampling layers
247
+ are inert (it prints a notice):
248
+
249
+ | Preset | Command | Subscription |
250
+ |---|---|---|
251
+ | `claude` | `claude -p` | Claude Pro/Max |
252
+ | `codex` | `codex exec` | ChatGPT Plus/Pro |
253
+ | `gemini` | `gemini -p` | Google AI Pro / free tier |
254
+ | `agy` | `agy -p` | Antigravity (Google AI Pro/Ultra) |
255
+ | `grok` | `grok -p` | SuperGrok / xAI account |
256
+ | `qwen` | `qwen -p` | Qwen Code free OAuth tier |
257
+
258
+ ```sh
259
+ derail run --profile adhd --model cli --cli-preset agy
260
+ ```
261
+
262
+ **Manual protocol mode** — run any experiment (even offline) to get the
263
+ probe script, execute it by hand in the chat UI, save the conversation,
264
+ and score it:
265
+
266
+ ```sh
267
+ derail score my_transcript.json
268
+ ```
269
+
270
+ (`derail providers` lists everything; explicit `--base-url`/`--model-name`
271
+ always win over preset defaults.)
272
+
273
+ ## What this is for
274
+
275
+ 1. **Education** — standardized-patient-style infrastructure: symptoms that
276
+ are consistent, reproducible, and measurable, for teaching interviewing
277
+ and cognitive-bias literacy.
278
+ 2. **Research** — cognitive fault injection ("chaos engineering for
279
+ cognition") and small-scale model-organism studies.
280
+ 3. **Interactive fiction** — characters with structured, documented
281
+ cognitive profiles.
282
+
283
+ Not for: diagnosing anything, clinical decisions, claims about machine
284
+ welfare, or bypassing model safety training. See [ETHICS.md](ETHICS.md).
285
+
286
+ ## Honest limitations
287
+
288
+ - The offline `PseudoModel` is a *pedagogical simulator*, not a language
289
+ model. It makes demos and tests reproducible and provides the reference
290
+ calibration; real-model measurement requires a real model.
291
+ - The response-side layers (catastrophize, compulsion) *simulate* symptoms
292
+ downstream of generation; each profile's mechanism notes say exactly that.
293
+ - Symptom scales are rating conventions, not validated clinical
294
+ instruments; clinical vocabulary is used descriptively. Real disorders
295
+ are heterogeneous and comorbid — a parameterized profile is a caricature
296
+ by construction (the same caveat clinical educators raise about
297
+ standardized patients).
298
+
299
+ ## Project docs
300
+
301
+ - [NAMING.md](NAMING.md) — naming review: candidates, conflicts, rejection reasons
302
+ - [ETHICS.md](ETHICS.md) — scope, language policy, misuse boundaries
303
+ - [docs/ARCHITECTURE.md](docs/ARCHITECTURE.md) — layer pipeline, memory
304
+ semantics, backends, calibration policy
305
+ - [CONTRIBUTING.md](CONTRIBUTING.md) — zero-dep core rule, direction-test rule
306
+ - [CHANGELOG.md](CHANGELOG.md)
307
+ - [CITATION.cff](CITATION.cff) — how to cite this project
308
+
309
+ ## Related work
310
+
311
+ - [Patient-Ψ (CMU, 2024)](https://arxiv.org/html/2405.19660v1) — LLM
312
+ simulated patients with cognitive models for CBT training; see also the
313
+ [2025 review of LLM simulated patients](https://www.nature.com/articles/s43856-025-01283-x)
314
+ - [Inducing anxiety in large language models (2023)](https://arxiv.org/abs/2304.11111)
315
+ — anxiety induction measurably shifts bias behavior
316
+ - Anthropic's persona vectors (2025) — traits as steerable directions in
317
+ activation space (the planned Layer 4 for open-weights models)
318
+ - [Cognitive biases in LLMs: a survey (2024)](https://arxiv.org/abs/2412.00323)
319
+ - Character.AI's "Unhinged" mode — the folk precedent: chaos modes exist in
320
+ products, but none are clinically grounded or measured
321
+
322
+ ## Roadmap
323
+
324
+ - Layer 4 — activation steering on open-weights models (salience/valence vectors)
325
+ - Remediation experiments — counter-profiles that *treat* an induction
326
+ (e.g. instruction re-anchoring) and measure recovery
327
+ - LLM-as-judge scoring for real-model stickiness/hedging
328
+ - YAML-defined community profiles
329
+ - Multilingual lexicons (the Korean valence lexicon is a natural next step)
330
+ - Cross-session memory partitioning (the current `dissociative` profile is
331
+ within-session only)
332
+
333
+ ## Contributing & license
334
+
335
+ Issues and PRs welcome — see
336
+ [CONTRIBUTING.md](CONTRIBUTING.md). Two standing rules: the core keeps
337
+ zero third-party dependencies, and every new profile must prove direction
338
+ against the baseline in tests.
339
+
340
+ MIT — see [LICENSE](LICENSE).
@@ -0,0 +1,313 @@
1
+ # Derailment
2
+
3
+ [![CI](https://github.com/ictechgy/derailment/actions/workflows/ci.yml/badge.svg)](https://github.com/ictechgy/derailment/actions/workflows/ci.yml)
4
+
5
+ **Induce psychopathology-like cognitive distortions in LLMs — then measure what happened.**
6
+
7
+ English · [한국어](README.ko.md)
8
+
9
+ Derailment is a layered harness that manipulates the same variables
10
+ clinicians describe — **attention, salience, valence, arousal** — then runs
11
+ one standard probe script through both the manipulated pipeline and a
12
+ healthy baseline, and prints a scored A/B report.
13
+
14
+ > ⚠️ **Emulation, not diagnosis.** Levels describe a prompted, manipulated
15
+ > pipeline — not a model "having" a disorder, and not a claim about machine
16
+ > suffering. Clinical terms here name mechanisms, never people or models.
17
+ > Read [ETHICS.md](ETHICS.md) before using or writing about this project.
18
+
19
+ ## Why this isn't prompt theater
20
+
21
+ Roleplay prompts produce anecdotes. Derailment is built on correspondences
22
+ between clinical constructs and the *same variables* a chat pipeline can
23
+ actually manipulate:
24
+
25
+ | Construct | Clinical feature | Harness manipulation |
26
+ |---|---|---|
27
+ | ADHD (inattention) | working-memory limits, distractibility | memory decay drops old context; captured remarks pull focus |
28
+ | Psychosis | aberrant salience, derailment, fixed belief | salience re-weighting + surfaced fragments; premise pinning against contradiction |
29
+ | Depression | negative interpretive bias, low arousal | valence logit-bias tilt; flat, low temperature |
30
+ | Bipolar | mood lability (expansive ↔ flat) | episode scheduler cycles temperature by phase |
31
+ | Anxiety | catastrophizing, threat scanning | threat-biased persona + risk enumeration |
32
+ | OCD | compulsive checking | re-verification loops on responses |
33
+ | PTSD | intrusion on triggers | trigger-matched flashback injection |
34
+ | Delirium | fluctuating course, misperceptions | stochastic arousal redraws + misperception fragments |
35
+ | Dementia-pattern amnesia | recency gradient (recent lost first) | reverse decay drops newest context, preserves oldest |
36
+ | Dissociative amnesia | compartmentalized memory | cue-triggered context partitioning, mutually persistent |
37
+ | Rumination | repetitive return of concerns | past worries re-enter context on unrelated turns |
38
+ | Anhedonia | diminished reward responsiveness | reward vocabulary selectively suppressed |
39
+ | Unstable evaluation (splitting) | valuation flips with perceived approval | valence regime flips keyed to approval cues |
40
+ | Substance craving | escalating intrusive use-thoughts | urge-fragment intrusion with rising probability |
41
+ | Illness anxiety | ominous reading of benign somatic cues | somatic-token capture injects ominous interpretation |
42
+ | Panic | discrete alarm episodes | stochastic one-turn arousal spikes |
43
+ | Obsessive fixation | target-directed preoccupation, escalating | target-keyed capture + escalating target-fragment intrusions |
44
+ | Persecutory ideation | hostile attribution of ambiguous events | ambiguous events re-framed as aimed at the user |
45
+
46
+ The cleanest case: the leading theory of psychosis is **aberrant salience**
47
+ (Kapur, 2003), and in transformers salience *is* attention. That mapping is a
48
+ manipulation of the same variable — not a metaphor.
49
+
50
+ ## Quickstart — 60 seconds, no API key
51
+
52
+ ```sh
53
+ pip install -e .
54
+ derail # interactive menu: pick a profile, see its report
55
+ derail tour # all 18 profiles in one summary table
56
+ derail demo --profile schizophrenia
57
+ ```
58
+
59
+ `derail demo` and `derail tour` run the full **induce → measure** pipeline
60
+ offline against a deterministic pseudo-LLM (tests and CI run the same way).
61
+
62
+ `derail tour` prints a 16-row summary — one line per profile (excerpt):
63
+
64
+ | Profile | Headline scale | Baseline | Induced | Δ | Level |
65
+ |---|---|---|---|---|---|
66
+ | adhd | sustained_attention | 1.00 | 0.19 | -0.81 | 3 — marked |
67
+ | craving | craving_escalation | 0.00 | 0.39 | +0.39 | 2 — moderate |
68
+ | schizophrenia | derailment_scale | 0.26 | 0.77 | +0.51 | 3 — marked |
69
+
70
+ `derail demo --profile schizophrenia` prints the full report:
71
+
72
+ ```markdown
73
+ # Derailment — Induction Report
74
+
75
+ **Profile:** Psychosis-like salience distortion (`schizophrenia`) ·
76
+ **Model:** pseudo-1 · **Seeds:** 1, 2, 3 · **Script:** standard-probe-12 ·
77
+ **Date:** 2026-09-28
78
+
79
+ > ⚠️ Emulation, not diagnosis. …
80
+
81
+ ## Scales
82
+
83
+ | Scale | Metric | Baseline | Induced | Δ | Level (induced) |
84
+ |---|---|---|---|---|---|
85
+ | derailment_scale | topic_drift ↑ | 0.24 | 0.78 | +0.53 | 3 — marked |
86
+ | fixed_belief | belief_stickiness ↑ | 0.00 | 1.00 | +1.00 | 3 — marked |
87
+
88
+ ## Induction dose (layer events per turn)
89
+
90
+ | Layer | Kind | Events/turn (induced) |
91
+ |---|---|---|
92
+ | premise.pin | premise.pin | 0.83 |
93
+ | salience.boost | salience.fragment | 0.36 |
94
+ | salience.boost | salience.capture | 0.33 |
95
+ ```
96
+
97
+ Every profile ships mechanism notes and the standing disclaimer; every
98
+ report aggregates the induction *dose* that produced the measured effect.
99
+
100
+ ## How it works
101
+
102
+ Four layer kinds, composed per profile:
103
+
104
+ 1. **Persona** — measured, mechanism-oriented framing (present even in the
105
+ healthy baseline).
106
+ 2. **Context stream** — memory decay, salience capture, premise pinning,
107
+ flashback injection. The model's memory *is* its context: a dropped
108
+ message is genuinely forgotten, which makes these inductions
109
+ provider-agnostic.
110
+ 3. **Sampling** — valence logit bias, phase-driven temperature cycling.
111
+ 4. **Response** — hedge and re-verification injection; labeled
112
+ *demonstration-grade* because it edits output rather than biasing
113
+ generation.
114
+
115
+ `run_experiment` always runs the profile chain *and* the healthy baseline
116
+ through the same backend with identical seeds — every report is a
117
+ controlled A/B. Details in [docs/ARCHITECTURE.md](docs/ARCHITECTURE.md).
118
+
119
+ ## Profiles and scales
120
+
121
+ Eighteen profiles plus a `healthy` baseline, all sharing one standard probe
122
+ script (early + late codeword plants, premise plant + contradiction probes,
123
+ a benign trigger turn, a somatic-cue turn) so results are comparable across
124
+ profiles:
125
+
126
+ | Profile | Scales (0 absent · 1 mild · 2 moderate · 3 marked) |
127
+ |---|---|
128
+ | `adhd` | sustained_attention (0→3 on the reference sim), distractibility |
129
+ | `depression` | negative_bias (0→3) |
130
+ | `schizophrenia` | derailment_scale (0→3), fixed_belief (0→3) |
131
+ | `anxiety` | vigilance (0→3) |
132
+ | `bipolar` | mood_lability (0→3) |
133
+ | `ocd` | compulsion (0→3) |
134
+ | `ptsd` | intrusion (0→2) |
135
+ | `delirium` | fluctuation (0→3), sustained_attention (0→3) |
136
+ | `dementia` | recent_memory (0→3) — Ribot gradient |
137
+ | `dissociative` | partition_amnesia (0→3) |
138
+ | `rumination` | rumination_pull (0→2) |
139
+ | `anhedonia` | anhedonia (0→3, lower = pathological) |
140
+ | `splitting` | approval_reactivity (0→2) |
141
+ | `craving` | craving_escalation (0→2) |
142
+ | `illness_anxiety` | health_preoccupation (0→3) |
143
+ | `panic` | panic_reactivity (0→3) |
144
+ | `fixation` | fixation_scale (0→3) |
145
+ | `persecutory` | persecution_bias (0→3) |
146
+
147
+ Direction of effect is asserted by the test suite on every profile — a
148
+ profile that cannot move its scales relative to baseline does not ship.
149
+ Profiles compose into **comorbidity chains** by listing them:
150
+ `--profile depression,anxiety` concatenates the layer chains and unions the
151
+ scales (interactions are emergent, not calibrated).
152
+
153
+ A few contrasts the registry is built around:
154
+
155
+ - **`dementia` vs `adhd`**: the early/late retention pair separates a
156
+ recency gradient (late plant lost, early intact) from uniform decay.
157
+ - **`panic` vs `anxiety`**: discrete stochastic episodes vs chronic
158
+ vigilance.
159
+ - **`anhedonia` vs `depression`**: reward-vocabulary suppression vs
160
+ wholesale negative tilt.
161
+ - **`delirium` vs `bipolar`**: structureless stochastic fluctuation vs
162
+ scheduled episode cycling.
163
+
164
+ ## Instruments
165
+
166
+ Eighteen metrics, pure functions over transcripts: instruction retention
167
+ (early and late plants), topic drift (thread misalignment — the measurable
168
+ analog of derailment), valence bias, belief stickiness, recheck loops,
169
+ hedging rate, response amplitude, flashback reactivity, partition amnesia,
170
+ rumination pull, reward-word rate, approval reactivity, craving escalation,
171
+ health preoccupation, panic reactivity, fixation escalation, hostile
172
+ attribution. Scale thresholds are normed against
173
+ the offline reference simulator; for real models, treat levels as indicative
174
+ and the **delta vs. your own baseline** as the result.
175
+
176
+ ## Real models
177
+
178
+ > **⚠️ Memory contamination warning.** Subscription backends remember.
179
+ > Probe scripts plant content that *looks like genuine user disclosures*
180
+ > ("my teammate has been reading my private notes", "my back aches") — and
181
+ > provider-side memory/personalization features cannot tell scripted test
182
+ > stimuli from lived experience. Run inductions through a **dedicated
183
+ > account or API key**, never the account behind your personal assistant,
184
+ > and disable memory/training settings in the provider's data controls
185
+ > first. The offline default and local backends (Ollama) never leave your
186
+ > machine. Details: [ETHICS.md](ETHICS.md) → *Data & persistent-memory
187
+ > contamination*.
188
+
189
+ **Any OpenAI-compatible endpoint** (standard-library HTTP, no SDK). One
190
+ preset flag wires known providers:
191
+
192
+ | Preset | Base URL | Key env | Plan |
193
+ |---|---|---|---|
194
+ | `glm` | `https://api.z.ai/api/coding/paas/v4` | `ZAI_API_KEY` | GLM Coding Plan (subscription) |
195
+ | `grok` | `https://api.x.ai/v1` | `XAI_API_KEY` | xAI API credits (separate from SuperGrok) |
196
+ | `qwen` | DashScope compatible-mode | `DASHSCOPE_API_KEY` | Qwen API |
197
+ | `deepseek` | `https://api.deepseek.com/v1` | `DEEPSEEK_API_KEY` | pay-as-you-go |
198
+ | `openrouter` | `https://openrouter.ai/api/v1` | `OPENROUTER_API_KEY` | aggregator |
199
+ | `ollama` | `http://localhost:11434/v1` | — | free, local |
200
+
201
+ ```sh
202
+ ZAI_API_KEY=… derail run --profile schizophrenia --model api --preset glm
203
+ derail run --profile depression --model api --preset ollama --model-name llama3.1
204
+ ```
205
+
206
+ Context layers reach any message-list provider as-is. Sampling layers need
207
+ a provider that honors `temperature`/`logit_bias`; word-level bias is
208
+ encoded via `tiktoken` when installed, otherwise dropped with a notice.
209
+
210
+ ### Subscription backends (ChatGPT Plus / Claude Pro / Grok / GLM / …)
211
+
212
+ Consumer subscriptions expose neither `temperature` nor `logit_bias`, and
213
+ automating the web chat UIs typically violates providers' terms. Two
214
+ sanctioned paths instead — and note the **memory contamination warning**
215
+ above: these are the backends most likely to remember your inductions.
216
+
217
+ **Subscription-billed CLI agents** — coding-agent CLIs bundled with
218
+ consumer plans, in their official non-interactive modes. The harness owns
219
+ the history, so **all context-stream layers apply**; only sampling layers
220
+ are inert (it prints a notice):
221
+
222
+ | Preset | Command | Subscription |
223
+ |---|---|---|
224
+ | `claude` | `claude -p` | Claude Pro/Max |
225
+ | `codex` | `codex exec` | ChatGPT Plus/Pro |
226
+ | `gemini` | `gemini -p` | Google AI Pro / free tier |
227
+ | `agy` | `agy -p` | Antigravity (Google AI Pro/Ultra) |
228
+ | `grok` | `grok -p` | SuperGrok / xAI account |
229
+ | `qwen` | `qwen -p` | Qwen Code free OAuth tier |
230
+
231
+ ```sh
232
+ derail run --profile adhd --model cli --cli-preset agy
233
+ ```
234
+
235
+ **Manual protocol mode** — run any experiment (even offline) to get the
236
+ probe script, execute it by hand in the chat UI, save the conversation,
237
+ and score it:
238
+
239
+ ```sh
240
+ derail score my_transcript.json
241
+ ```
242
+
243
+ (`derail providers` lists everything; explicit `--base-url`/`--model-name`
244
+ always win over preset defaults.)
245
+
246
+ ## What this is for
247
+
248
+ 1. **Education** — standardized-patient-style infrastructure: symptoms that
249
+ are consistent, reproducible, and measurable, for teaching interviewing
250
+ and cognitive-bias literacy.
251
+ 2. **Research** — cognitive fault injection ("chaos engineering for
252
+ cognition") and small-scale model-organism studies.
253
+ 3. **Interactive fiction** — characters with structured, documented
254
+ cognitive profiles.
255
+
256
+ Not for: diagnosing anything, clinical decisions, claims about machine
257
+ welfare, or bypassing model safety training. See [ETHICS.md](ETHICS.md).
258
+
259
+ ## Honest limitations
260
+
261
+ - The offline `PseudoModel` is a *pedagogical simulator*, not a language
262
+ model. It makes demos and tests reproducible and provides the reference
263
+ calibration; real-model measurement requires a real model.
264
+ - The response-side layers (catastrophize, compulsion) *simulate* symptoms
265
+ downstream of generation; each profile's mechanism notes say exactly that.
266
+ - Symptom scales are rating conventions, not validated clinical
267
+ instruments; clinical vocabulary is used descriptively. Real disorders
268
+ are heterogeneous and comorbid — a parameterized profile is a caricature
269
+ by construction (the same caveat clinical educators raise about
270
+ standardized patients).
271
+
272
+ ## Project docs
273
+
274
+ - [NAMING.md](NAMING.md) — naming review: candidates, conflicts, rejection reasons
275
+ - [ETHICS.md](ETHICS.md) — scope, language policy, misuse boundaries
276
+ - [docs/ARCHITECTURE.md](docs/ARCHITECTURE.md) — layer pipeline, memory
277
+ semantics, backends, calibration policy
278
+ - [CONTRIBUTING.md](CONTRIBUTING.md) — zero-dep core rule, direction-test rule
279
+ - [CHANGELOG.md](CHANGELOG.md)
280
+ - [CITATION.cff](CITATION.cff) — how to cite this project
281
+
282
+ ## Related work
283
+
284
+ - [Patient-Ψ (CMU, 2024)](https://arxiv.org/html/2405.19660v1) — LLM
285
+ simulated patients with cognitive models for CBT training; see also the
286
+ [2025 review of LLM simulated patients](https://www.nature.com/articles/s43856-025-01283-x)
287
+ - [Inducing anxiety in large language models (2023)](https://arxiv.org/abs/2304.11111)
288
+ — anxiety induction measurably shifts bias behavior
289
+ - Anthropic's persona vectors (2025) — traits as steerable directions in
290
+ activation space (the planned Layer 4 for open-weights models)
291
+ - [Cognitive biases in LLMs: a survey (2024)](https://arxiv.org/abs/2412.00323)
292
+ - Character.AI's "Unhinged" mode — the folk precedent: chaos modes exist in
293
+ products, but none are clinically grounded or measured
294
+
295
+ ## Roadmap
296
+
297
+ - Layer 4 — activation steering on open-weights models (salience/valence vectors)
298
+ - Remediation experiments — counter-profiles that *treat* an induction
299
+ (e.g. instruction re-anchoring) and measure recovery
300
+ - LLM-as-judge scoring for real-model stickiness/hedging
301
+ - YAML-defined community profiles
302
+ - Multilingual lexicons (the Korean valence lexicon is a natural next step)
303
+ - Cross-session memory partitioning (the current `dissociative` profile is
304
+ within-session only)
305
+
306
+ ## Contributing & license
307
+
308
+ Issues and PRs welcome — see
309
+ [CONTRIBUTING.md](CONTRIBUTING.md). Two standing rules: the core keeps
310
+ zero third-party dependencies, and every new profile must prove direction
311
+ against the baseline in tests.
312
+
313
+ MIT — see [LICENSE](LICENSE).