omp-uwu 0.1.0 → 0.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -4,10 +4,11 @@
4
4
 
5
5
  Long sessions with an agent can get a little grey: diffs, stack traces, test output, repeat. omp-uwu is an [omp](https://github.com/can1357/oh-my-pi) extension that makes the agent write its chat replies in playful uwu-speak, while everything that has to be exact stays exact.
6
6
 
7
- > I've fixed the bug in `src/parser.ts` and all 42 tests pass uwu.
8
- > The pwobwem was an off-by-one in de woop, nyow it checks `i < len` instead of `i <= len` >w<
7
+ <p align="center">
8
+ <img src="https://raw.githubusercontent.com/NaC-L/omp-uwu/main/demo/demo.gif" alt="omp streaming an uwu-speak answer about an off-by-one loop, with the corrected code block unchanged" width="784">
9
+ </p>
9
10
 
10
- The prose is uwufied, but the file path, identifiers, code and numbers are not.
11
+ The prose is uwufied, but the inline code, numbers and the fixed loop stay exactly as they should be.
11
12
 
12
13
  ## Quick start
13
14
 
@@ -19,6 +20,10 @@ Restart omp. uwu mode is **on** by default in every session.
19
20
 
20
21
  - `/uwu` toggles it
21
22
  - `/uwu on` / `/uwu off` sets it explicitly
23
+ - `/uwu rewrite` uses omp's finalized-message rewrite hook when available
24
+ - `/uwu prompt` asks the model to write in uwu style while it streams
25
+
26
+ The TUI also shows a small, theme-accent-colored `(◕ᴗ◕✿) uwu` status badge. Other clients and plain output are not colorized.
22
27
 
23
28
  ## What gets uwufied, and what doesn't
24
29
 
@@ -30,18 +35,23 @@ Restart omp. uwu mode is **on** by default in every session.
30
35
  | | Tool-call arguments and file contents the agent writes or edits |
31
36
  | | Commit messages and prompts sent to subagents |
32
37
 
33
- Style rules the agent follows: `r`/`l` → `w` in most words, `th` → `d` sometimes, `na/ne/no` → `nya/nye/nyo` sometimes, the occasional stutter (`h-hewwo`), and at most one emoticon per sentence (`uwu`, `owo`, `>w<`, `^w^`, `:3`). Meaning, numbers and warnings must stay readable.
38
+ Style rules include `r`/`l` → `w`, occasional `th` → `d`, `na/ne/no` → `nya/nye/nyo`, occasional stutter, and a broad rotating selection of kaomoji. The kaomoji selection includes examples from [kaomoji.you](https://kaomoji.you/), such as `٩(◕‿◕。)۶`, `(ฅ^•ﻌ•^ฅ)`, and `(づ。◕‿‿◕。)づ`. Meaning, numbers and warnings must stay readable.
39
+
40
+ Here's the same prompt and model, with the mode off and on:
41
+
42
+ ![Two real omp replies side by side: plain English on the left, uwu-speak on the right, with an identical code block in both](https://raw.githubusercontent.com/NaC-L/omp-uwu/main/demo/before-after.png)
43
+
44
+ These are real, unedited replies from `anthropic/claude-opus-5-5`, rendered as images. Both code blocks are identical byte for byte. The transcripts are in [`demo/`](demo).
34
45
 
35
46
  ## How it works
36
47
 
37
- omp has no extension hook for rewriting assistant text after it's generated. The `message_end` event only gets a detached copy of the message. So omp-uwu asks the model instead: on `before_agent_start` it adds a short style instruction to the end of the system prompt.
48
+ On the first turn, omp-uwu uses the prompt style until the host demonstrates support for the awaited `assistant_message` hook; after that, the default **rewrite** style deterministically uwufies finalized assistant text before it is added to history and context. This first-turn check makes the experience work on both older and newer omp builds. Only text in existing text blocks is changed; code/tool blocks and their metadata stay untouched. Rewrites are markdown-aware and preserve fenced/inline code, URLs, paths, numbers, quoted text, identifiers, and safety-critical words.
38
49
 
39
- This has some consequences:
50
+ The hook runs after streaming has finished. The TUI refreshes the current assistant message at completion, but clients that render only streamed chunks may continue showing the original text. If the hook is unavailable (including omp `18.4.3`), prompt style remains enabled as the compatibility fallback. `/uwu prompt` selects live prompt styling directly.
40
51
 
41
- - **It depends on the model.** Most models follow it well, but the style isn't guaranteed. If code ever comes back uwufied, open an issue with the model name.
42
- - **History stays clean.** Nothing is rewritten after generation, so session history, compaction and tools see exactly what the model wrote.
43
- - **Small cost.** The instruction adds roughly 200 tokens to the system prompt while the mode is on. `/uwu off` removes it from the next request.
44
- - **Session-local toggle.** The on/off state isn't saved. Each new session starts with uwu mode on.
52
+ - **Deterministic rewrite.** No style instruction/token overhead once the hook is detected.
53
+ - **Prompt style.** Model-dependent, with live uwu output while streaming.
54
+ - **Session-local toggle.** Mode and enabled state reset for each new session.
45
55
 
46
56
  ## Install options
47
57
 
@@ -63,14 +73,26 @@ omp plugin link . # use this checkout instead of the installed copy
63
73
  omp -e ./src/index.ts # or load it for a single run
64
74
  ```
65
75
 
76
+ ### Re-recording the demo
77
+
78
+ The GIF and the side-by-side image are rendered from real transcripts in `demo/`. `record.sh` captures one reply with the extension loaded and one without it. It runs in a throwaway agent dir, so your personal rules and extensions don't affect the replies. `render.py` then draws both images:
79
+
80
+ ```sh
81
+ mkdir -p /tmp/omp-demo && cp ~/.omp/agent/agent.db* /tmp/omp-demo/
82
+ demo/record.sh
83
+ uv run --with pillow python demo/render.py
84
+ ```
85
+
86
+ Model output varies between runs, so re-run `record.sh` until you get a reply that reads well.
87
+
66
88
  ### Releasing
67
89
 
68
90
  CI (`.github/workflows/ci.yml`) runs `check` and the tests on pushes to `main` and on pull requests. Pushing a `v*` tag runs them again, checks that the tag matches `version` in `package.json`, and publishes to npm with provenance (`.github/workflows/publish.yml`).
69
91
 
70
92
  ```sh
71
93
  # bump "version" in package.json, commit, then:
72
- git tag v0.1.1
73
- git push origin main v0.1.1
94
+ git tag v0.2.0
95
+ git push origin main v0.2.0
74
96
  ```
75
97
 
76
98
  The workflow uses npm [trusted publishing](https://docs.npmjs.com/trusted-publishers/), so the repository has no npm token. npm only allows trusted publishing on a package that already exists, so the first version is published by hand with `npm publish --access public`. After that, go to the package's Settings → Trusted publishing on npmjs.com and add GitHub Actions with user `NaC-L`, repository `omp-uwu`, and workflow `publish.yml`.
package/package.json CHANGED
@@ -1,14 +1,15 @@
1
1
  {
2
2
  "name": "omp-uwu",
3
- "version": "0.1.0",
4
- "description": "omp extension: the agent answers in uwu-speak, while code, paths and commands stay byte-for-byte exact",
3
+ "version": "0.2.0",
4
+ "description": "omp extension for playful uwu chat styling, markdown-safe rewrites, and kawaii kaomoji",
5
5
  "type": "module",
6
6
  "license": "MIT",
7
7
  "keywords": [
8
8
  "omp",
9
9
  "oh-my-pi",
10
10
  "omp-plugin",
11
- "uwu"
11
+ "uwu",
12
+ "kaomoji"
12
13
  ],
13
14
  "homepage": "https://github.com/NaC-L/omp-uwu#readme",
14
15
  "bugs": "https://github.com/NaC-L/omp-uwu/issues",
package/src/index.ts CHANGED
@@ -1,41 +1,106 @@
1
- /**
2
- * uwu mode — makes the agent's chat replies uwu-speak.
3
- *
4
- * Why system-prompt based: OMP gives extensions no hook to rewrite displayed
5
- * assistant text (`message_end` receives a detached clone), so the style is
6
- * requested from the model via `before_agent_start` → `systemPrompt`.
7
- *
8
- * Toggle with `/uwu` (or `/uwu on` / `/uwu off`). Enabled by default.
9
- */
10
1
  import type { ExtensionAPI } from "@oh-my-pi/pi-coding-agent";
2
+ import { uwufy } from "./uwufy.ts";
11
3
 
12
4
  export const UWU_PROMPT = `
13
5
  # uwu mode (display style only)
14
6
  Write your prose replies to the user in playful "uwu" speak:
15
7
  - replace r/l with w in most words ("weawwy", "hewwo"), "th" → "d" occasionally, "na/ne/no" → "nya/nye/nyo" sometimes
16
- - occasional stutter ("h-hewwo") and cute emoticons at sentence ends (uwu, owo, >w<, ^w^, :3), at most one per sentence
17
- - keep it readable; never let the style hide meaning, numbers, or warnings
8
+ - occasional stutter ("h-hewwo") and cute kaomoji/emoticons, varied naturally: (◕ᴗ◕✿), (o^▽^o), (´。• ω •。\u0060), ٩(◕‿◕。)۶, (✧ω✧), (๑˃ᴗ˂)ﻭ, ヽ(・∀・)ノ, (っ˘ω˘ς), (。•̀ᴗ-)✧, (づ。◕‿‿◕。)づ, (ฅ^•ﻌ•^ฅ), (๑>◡<๑), owo, >w<, ^w^, :3; at most one at a sentence end
9
+ - keep it readable; never let the style hide meaning, numbers or warnings
10
+ - colors are only for UI chrome; never insert ANSI color codes into chat text
18
11
 
19
12
  NEVER uwufy any of these — keep them exact and byte-for-byte correct:
20
13
  - code, code blocks, inline code, shell commands, file paths, URLs, identifiers, config keys, error messages you quote
21
- - tool call arguments, file contents you write or edit, commit messages, and anything sent to subagents
14
+ - tool call arguments, file contents the agent writes or edits, commit messages and prompts sent to subagents
22
15
  The style applies only to natural-language chat text shown to the user. Task quality and correctness are unchanged.
23
16
  `.trim();
24
17
 
18
+ type RewriteEvent = {
19
+ message: { role: string; content: Array<{ type: string; text?: string; [key: string]: unknown }> };
20
+ };
21
+ type MessageEndEvent = { message: { role: string; stopReason?: string } };
22
+ type UiContext = {
23
+ mode?: string;
24
+ agent?: { kind?: string };
25
+ ui?: {
26
+ notify(message: string, level: "info"): void;
27
+ setStatus?(key: string, text: string | undefined): void;
28
+ theme?: { fg(color: string, text: string): string };
29
+ };
30
+ };
31
+
25
32
  export default function uwuExtension(pi: ExtensionAPI) {
26
- let enabled = true;
27
-
28
- pi.registerCommand("uwu", {
29
- description: "Toggle uwu mode (usage: /uwu [on|off])",
30
- handler: async (args, ctx) => {
31
- const arg = String(args ?? "").trim().toLowerCase();
32
- enabled = arg === "on" ? true : arg === "off" ? false : !enabled;
33
- ctx.ui.notify(enabled ? "uwu mode enabwed! (◕ᴗ◕✿)" : "uwu mode disabled", "info");
34
- },
35
- });
36
-
37
- pi.on("before_agent_start", async event => {
38
- if (!enabled) return undefined;
39
- return { systemPrompt: [...event.systemPrompt, UWU_PROMPT] };
40
- });
33
+ let enabled = true;
34
+ let style: "rewrite" | "prompt" = "rewrite";
35
+ let hookSupport: boolean | undefined;
36
+ let promptAddedThisTurn = false;
37
+ let notifiedFallback = false;
38
+
39
+ const updateBadge = (ctx: UiContext) => {
40
+ if (ctx.mode !== "tui" || ctx.agent?.kind === "sub" || !ctx.ui?.setStatus) return;
41
+ const badge = enabled ? "(◕ᴗ◕✿) uwu" : undefined;
42
+ ctx.ui.setStatus("omp-uwu", badge && ctx.ui.theme ? ctx.ui.theme.fg("accent", badge) : badge);
43
+ };
44
+
45
+ pi.registerCommand("uwu", {
46
+ description: "UwU chat style and mode (usage: /uwu [on|off|prompt|rewrite])",
47
+ handler: async (args, rawCtx) => {
48
+ const ctx = rawCtx as UiContext;
49
+ const arg = String(args ?? "").trim().toLowerCase();
50
+ if (arg === "prompt" || arg === "rewrite") {
51
+ style = arg;
52
+ enabled = true;
53
+ } else if (arg === "on") enabled = true;
54
+ else if (arg === "off") enabled = false;
55
+ else if (!arg) enabled = !enabled;
56
+ else {
57
+ ctx.ui?.notify("Usage: /uwu [on|off|prompt|rewrite]", "info");
58
+ return;
59
+ }
60
+ updateBadge(ctx);
61
+ ctx.ui?.notify(enabled ? `uwu mode ${style === "rewrite" ? "rewrite" : "prompt"}! (◕ᴗ◕✿)` : "uwu mode off", "info");
62
+ },
63
+ });
64
+
65
+ pi.on("session_start", (_event, rawCtx) => updateBadge(rawCtx as UiContext));
66
+
67
+ pi.on("before_agent_start", async (event, rawCtx) => {
68
+ promptAddedThisTurn = false;
69
+ const ctx = rawCtx as UiContext;
70
+ if (ctx.agent?.kind === "sub") return undefined;
71
+ // Start with the prompt until this host proves it has the finalized hook.
72
+ // That styles the first reply on both old and new omp versions.
73
+ const needsPrompt = style === "prompt" || hookSupport !== true;
74
+ if (!enabled || !needsPrompt) return undefined;
75
+ promptAddedThisTurn = true;
76
+ return { systemPrompt: [...event.systemPrompt, UWU_PROMPT] };
77
+ });
78
+
79
+ // The hook is newer than the bundled 18.4.3 types, so register structurally.
80
+ // ExtensionAPI.on stores event names as strings; old hosts simply never emit it.
81
+ (pi.on as unknown as (name: string, handler: (event: RewriteEvent) => unknown) => void)(
82
+ "assistant_message",
83
+ (event) => {
84
+ hookSupport = true;
85
+ if (!enabled || style !== "rewrite" || promptAddedThisTurn || event.message.role !== "assistant") return;
86
+ let changed = false;
87
+ const content = event.message.content.map((block) => {
88
+ if (block.type !== "text" || typeof block.text !== "string") return block;
89
+ const text = uwufy(block.text);
90
+ if (text === block.text) return block;
91
+ changed = true;
92
+ return { ...block, text };
93
+ });
94
+ return changed ? { content } : undefined;
95
+ },
96
+ );
97
+
98
+ pi.on("message_end", (rawEvent) => {
99
+ const message = (rawEvent as MessageEndEvent).message;
100
+ if (message.role === "assistant" && message.stopReason !== "aborted" && message.stopReason !== "error" && hookSupport === undefined) {
101
+ hookSupport = false;
102
+ // The current response has already streamed; prompt fallback applies next turn.
103
+ if (!notifiedFallback) notifiedFallback = true;
104
+ }
105
+ });
41
106
  }
package/src/uwufy.ts ADDED
@@ -0,0 +1,331 @@
1
+ /**
2
+ * Deterministic, markdown-aware uwu rewriter for finalized assistant text.
3
+ *
4
+ * Only natural-language words are touched. Everything that has to stay exact is
5
+ * left byte-for-byte alone: fenced/indented code, math blocks, inline code,
6
+ * URLs and link targets, HTML tags, double-quoted spans, and any token that
7
+ * looks like an identifier, path, number, flag, acronym, or warning word.
8
+ *
9
+ * Every "sometimes" decision is keyed on the word's position (paragraph,
10
+ * sentence, word) rather than on its text, and no transform adds or removes a
11
+ * word token. Rewriting already-rewritten text therefore makes the same
12
+ * decisions and finds nothing left to do: `uwufy(uwufy(x)) === uwufy(x)`.
13
+ * That matters because rewritten replies go back into the model's context, and
14
+ * the model may start answering in uwu-speak on its own.
15
+ */
16
+
17
+ /** Emoticons the rewriter appends at sentence ends. */
18
+ export const EMOTICONS = [
19
+ "uwu", "owo", ">w<", "^w^", ":3", "(◕ᴗ◕✿)",
20
+ "(o^▽^o)", "(´。• ω •。`)", "٩(◕‿◕。)۶", "o(≧▽≦)o", "(✧ω✧)",
21
+ "(๑˃ᴗ˂)ﻭ", "(ᵔ◡ᵔ)", "ヽ(・∀・)ノ", "(´• ω •`)", "(。•́︿•̀。)",
22
+ "(ノ_<。)", "(っ˘ω˘ς)", "(≧◡≦)", "٩(◕‿◕)۶", "(o˘◡˘o)",
23
+ "(♡˙︶˙♡)", "ヽ(♡‿♡)ノ", "(´꒳`)♡", "(ノ´ з `)ノ", "ヾ(・ω・*)",
24
+ "(⌒ω⌒)ノ", "(ᵔ⩊ᵔ)", "(=^・ω・^=)", "(ฅ^•ﻌ•^ฅ)", "(ᵕ—ᴗ—)",
25
+ "(。•̀ᴗ-)✧", "(づ。◕‿‿◕。)づ", "(っ´▽`)っ", "( ˶ˆᗜˆ˵ )", "(๑>◡<๑)",
26
+ ] as const;
27
+
28
+ /** Tokens treated as emoticons: never counted as words, never rewritten. */
29
+ const EMOTICON_TOKENS = new Set<string>([
30
+ ...EMOTICONS,
31
+ "UwU",
32
+ "OwO",
33
+ "^_^",
34
+ "^^",
35
+ ":)",
36
+ ":D",
37
+ "<3",
38
+ "x3",
39
+ ]);
40
+
41
+ /** Words that carry meaning a reader must not miss. Kept exact. */
42
+ const KEEP_WORDS = new Set([
43
+ "no",
44
+ "not",
45
+ "nor",
46
+ "never",
47
+ "none",
48
+ "nothing",
49
+ "nowhere",
50
+ "neither",
51
+ "cannot",
52
+ "can't",
53
+ "don't",
54
+ "doesn't",
55
+ "didn't",
56
+ "won't",
57
+ "wouldn't",
58
+ "shouldn't",
59
+ "mustn't",
60
+ "isn't",
61
+ "aren't",
62
+ "wasn't",
63
+ "weren't",
64
+ "haven't",
65
+ "hasn't",
66
+ "hadn't",
67
+ "couldn't",
68
+ "warning",
69
+ "caution",
70
+ "danger",
71
+ "dangerous",
72
+ "careful",
73
+ "irreversible",
74
+ "destructive",
75
+ "error",
76
+ "errors",
77
+ "fail",
78
+ "fails",
79
+ "failed",
80
+ "failure",
81
+ ]);
82
+
83
+ const TH_WORDS = new Set(["the", "this", "that", "these", "those", "them", "then", "there", "their", "they", "than"]);
84
+
85
+ // Probabilities for the "sometimes" transforms.
86
+ const P_TH = 0.35;
87
+ const P_NY = 0.5;
88
+ const P_STUTTER = 0.12;
89
+ const P_EMOTICON = 0.3;
90
+
91
+ /**
92
+ * Inline spans that are never rewritten. Only the first alternative captures,
93
+ * so `\1` is the opening backtick run of a code span.
94
+ */
95
+ const PROTECTED_INLINE = new RegExp(
96
+ [
97
+ String.raw`(\`+)[\s\S]*?\1(?!\`)`, // code span
98
+ String.raw`\`.*$`, // unmatched backtick: keep the rest of the line
99
+ String.raw`\]\([^)\s]*(?:\s+"[^"]*")?\)`, // markdown link target
100
+ String.raw`<\/?[A-Za-z][^>\n]*>`, // HTML tag or autolink
101
+ String.raw`\b(?:[a-z][a-z0-9+.-]*:\/\/|www\.)\S*`, // bare URL
102
+ String.raw`"[^"\n]*"`, // "quoted text"
103
+ String.raw`“[^”\n]*”`, // “quoted text”
104
+ ].join("|"),
105
+ "gi",
106
+ );
107
+
108
+ const FENCE_OPEN = /^ {0,3}(`{3,}|~{3,})/;
109
+ const FENCE_CLOSE = /^ {0,3}(`{3,}|~{3,})\s*$/;
110
+ const MATH_FENCE = /^\s*\$\$/;
111
+ const INDENTED_CODE = /^(?: {4}|\t)/;
112
+ const LIST_ITEM = /^\s*(?:[-*+]|\d+[.)])\s/;
113
+ const HTML_LINE = /^\s*<[A-Za-z/!]/;
114
+ const LINK_DEFINITION = /^ {0,3}\[[^\]]+\]:\s/;
115
+ const HEADING = /^ {0,3}#{1,6}(?:\s|$)/;
116
+ const TABLE_ROW = /^\s*\|/;
117
+
118
+ const TOKEN_PARTS = /^([(\[{"'“‘*_~]*)(.*?)([)\]}"'”’*_~.,;:!?…]*)$/su;
119
+ const WORD = /^[\p{L}\p{M}][\p{L}\p{M}'’-]*$/u;
120
+ const SENTENCE_END = /[.!?…]/;
121
+ const STUTTERED = /^(\p{L})-\1/iu;
122
+
123
+ type LineKind = "code" | "blank" | "prose";
124
+
125
+ interface ProseLine {
126
+ index: number;
127
+ /** Headings and tables get word rewrites but no stutters or emoticons. */
128
+ decorate: boolean;
129
+ }
130
+
131
+ /** Rewrite the prose of a markdown string in uwu-speak. */
132
+ export function uwufy(text: string): string {
133
+ const lines = text.split("\n");
134
+ const kinds = classifyLines(lines);
135
+
136
+ const paragraphs: ProseLine[][] = [];
137
+ let current: ProseLine[] | undefined;
138
+ for (const [index, line] of lines.entries()) {
139
+ if (kinds[index] !== "prose") {
140
+ current = undefined;
141
+ continue;
142
+ }
143
+ const heading = HEADING.test(line);
144
+ // Headings, list items and table rows each start their own paragraph.
145
+ if (!current || heading || LIST_ITEM.test(line) || TABLE_ROW.test(line)) {
146
+ current = [];
147
+ paragraphs.push(current);
148
+ }
149
+ current.push({ index, decorate: !heading && !TABLE_ROW.test(line) });
150
+ if (heading) current = undefined;
151
+ }
152
+
153
+ const out = [...lines];
154
+ for (const [paragraphIndex, paragraph] of paragraphs.entries()) {
155
+ // Count once, then rewrite: the word count seeds this paragraph's decisions,
156
+ // which keeps them varied between paragraphs yet stable across re-runs.
157
+ const wordCount = walkParagraph(lines, paragraph, () => undefined);
158
+ const seed = `${paragraphIndex}|${wordCount}`;
159
+ walkParagraph(lines, paragraph, (event) => rewriteToken(event, seed), out);
160
+ }
161
+ return out.join("\n");
162
+ }
163
+
164
+ function classifyLines(lines: string[]): LineKind[] {
165
+ const kinds: LineKind[] = [];
166
+ let fence: { char: string; length: number } | undefined;
167
+ let math = false;
168
+ let previous: LineKind = "blank";
169
+ for (const line of lines) {
170
+ let kind: LineKind;
171
+ if (fence) {
172
+ const close = FENCE_CLOSE.exec(line);
173
+ if (close?.[1] && close[1][0] === fence.char && close[1].length >= fence.length) fence = undefined;
174
+ kind = "code";
175
+ } else if (math) {
176
+ if (MATH_FENCE.test(line)) math = false;
177
+ kind = "code";
178
+ } else if (FENCE_OPEN.test(line)) {
179
+ const open = FENCE_OPEN.exec(line)?.[1] ?? "```";
180
+ fence = { char: open[0] ?? "`", length: open.length };
181
+ kind = "code";
182
+ } else if (MATH_FENCE.test(line)) {
183
+ // A one-line `$$ x $$` block opens and closes on the same line.
184
+ math = !/^\s*\$\$.*\$\$\s*$/.test(line) || line.trim() === "$$";
185
+ kind = "code";
186
+ } else if (line.trim() === "") {
187
+ kind = "blank";
188
+ } else if (
189
+ (INDENTED_CODE.test(line) && !LIST_ITEM.test(line) && previous !== "prose") ||
190
+ HTML_LINE.test(line) ||
191
+ LINK_DEFINITION.test(line)
192
+ ) {
193
+ kind = "code";
194
+ } else {
195
+ kind = "prose";
196
+ }
197
+ kinds.push(kind);
198
+ previous = kind;
199
+ }
200
+ return kinds;
201
+ }
202
+
203
+ interface TokenEvent {
204
+ token: string;
205
+ /** Text after the token within its prose segment, for emoticon lookahead. */
206
+ rest: string;
207
+ /** Whether the segment ends the line (nothing protected follows it). */
208
+ lineEnd: boolean;
209
+ decorate: boolean;
210
+ sentence: number;
211
+ wordInSentence: number;
212
+ word: number;
213
+ }
214
+
215
+ /**
216
+ * Walk the word tokens of one paragraph. When `out` is given, each prose
217
+ * segment is rebuilt from the callback's replacements. Returns the word count.
218
+ */
219
+ function walkParagraph(
220
+ lines: string[],
221
+ paragraph: ProseLine[],
222
+ visit: (event: TokenEvent) => string | undefined,
223
+ out?: string[],
224
+ ): number {
225
+ let sentence = 0;
226
+ let wordInSentence = 0;
227
+ let word = 0;
228
+ for (const { index, decorate } of paragraph) {
229
+ const line = lines[index] ?? "";
230
+ let rebuilt = "";
231
+ let cursor = 0;
232
+ const segments: Array<{ prose: boolean; text: string }> = [];
233
+ for (const match of line.matchAll(PROTECTED_INLINE)) {
234
+ const start = match.index ?? 0;
235
+ if (start > cursor) segments.push({ prose: true, text: line.slice(cursor, start) });
236
+ segments.push({ prose: false, text: match[0] });
237
+ cursor = start + match[0].length;
238
+ }
239
+ if (cursor < line.length) segments.push({ prose: true, text: line.slice(cursor) });
240
+
241
+ for (const [segmentIndex, segment] of segments.entries()) {
242
+ if (!segment.prose) {
243
+ rebuilt += segment.text;
244
+ continue;
245
+ }
246
+ const lineEnd = segmentIndex === segments.length - 1;
247
+ rebuilt += segment.text.replace(/\S+/g, (token, offset: number) => {
248
+ if (EMOTICON_TOKENS.has(token)) return token;
249
+ const [, , core = "", trail = ""] = TOKEN_PARTS.exec(token) ?? [];
250
+ const isWord = WORD.test(core);
251
+ let replacement: string | undefined;
252
+ if (isWord) {
253
+ replacement = visit({
254
+ token,
255
+ rest: segment.text.slice(offset + token.length),
256
+ lineEnd,
257
+ decorate,
258
+ sentence,
259
+ wordInSentence,
260
+ word,
261
+ });
262
+ word++;
263
+ wordInSentence++;
264
+ }
265
+ if ((isWord || core === "") && SENTENCE_END.test(trail)) {
266
+ sentence++;
267
+ wordInSentence = 0;
268
+ }
269
+ return replacement ?? token;
270
+ });
271
+ }
272
+ if (out) out[index] = rebuilt;
273
+ }
274
+ return word;
275
+ }
276
+
277
+ function rewriteToken(event: TokenEvent, seed: string): string {
278
+ const [, lead = "", core = "", trail = ""] = TOKEN_PARTS.exec(event.token) ?? [];
279
+ const key = `${seed}|${event.sentence}|${event.word}`;
280
+ let result = event.token;
281
+
282
+ if (!isKeptWord(core)) {
283
+ let word = core;
284
+ if (TH_WORDS.has(word.toLowerCase()) && roll(`${key}|th`) < P_TH) {
285
+ word = (word[0] === "T" ? "D" : "d") + word.slice(2);
286
+ }
287
+ word = word.replace(/[rl]/g, "w").replace(/[RL]/g, "W");
288
+ if (roll(`${key}|ny`) < P_NY) word = word.replace(/([nN])(?=[aeo])/g, "$1y");
289
+ if (
290
+ event.decorate &&
291
+ event.wordInSentence === 0 &&
292
+ word.length >= 3 &&
293
+ !STUTTERED.test(word) &&
294
+ roll(`${key}|stutter`) < P_STUTTER
295
+ ) {
296
+ word = `${word[0]}-${word}`;
297
+ }
298
+ // Never turn a word into something that reads as an emoticon (e.g. "ulu").
299
+ if (!EMOTICON_TOKENS.has(word)) result = lead + word + trail;
300
+ }
301
+
302
+ if (
303
+ event.decorate &&
304
+ SENTENCE_END.test(trail) &&
305
+ (/^\s/.test(event.rest) || (event.rest === "" && event.lineEnd)) &&
306
+ !EMOTICON_TOKENS.has(/^\s+(\S+)/.exec(event.rest)?.[1] ?? "") &&
307
+ roll(`${seed}|${event.sentence}|emoticon`) < P_EMOTICON
308
+ ) {
309
+ result += ` ${EMOTICONS[Math.floor(roll(`${seed}|${event.sentence}|pick`) * EMOTICONS.length)]}`;
310
+ }
311
+ return result;
312
+ }
313
+
314
+ /** Acronyms, camelCase names, flags and meaning-critical words stay exact. */
315
+ function isKeptWord(core: string): boolean {
316
+ if (KEEP_WORDS.has(core.toLowerCase().replaceAll("’", "'"))) return true;
317
+ if (core.endsWith("-") || core.includes("--")) return true;
318
+ if (/\p{Ll}\p{Lu}/u.test(core)) return true; // camelCase, GitHub, iOS
319
+ const letters = core.replace(/[^\p{L}]/gu, "");
320
+ return letters.length >= 2 && letters === letters.toUpperCase() && letters !== letters.toLowerCase();
321
+ }
322
+
323
+ /** FNV-1a hash of `key`, scaled to [0, 1). */
324
+ function roll(key: string): number {
325
+ let hash = 0x811c9dc5;
326
+ for (let i = 0; i < key.length; i++) {
327
+ hash ^= key.charCodeAt(i);
328
+ hash = Math.imul(hash, 0x01000193);
329
+ }
330
+ return (hash >>> 0) / 2 ** 32;
331
+ }