omp-uwu 0.1.0 → 0.2.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +34 -12
- package/package.json +4 -3
- package/src/index.ts +92 -27
- package/src/uwufy.ts +331 -0
package/README.md
CHANGED
|
@@ -4,10 +4,11 @@
|
|
|
4
4
|
|
|
5
5
|
Long sessions with an agent can get a little grey: diffs, stack traces, test output, repeat. omp-uwu is an [omp](https://github.com/can1357/oh-my-pi) extension that makes the agent write its chat replies in playful uwu-speak, while everything that has to be exact stays exact.
|
|
6
6
|
|
|
7
|
-
>
|
|
8
|
-
|
|
7
|
+
<p align="center">
|
|
8
|
+
<img src="https://raw.githubusercontent.com/NaC-L/omp-uwu/main/demo/demo.gif" alt="omp streaming an uwu-speak answer about an off-by-one loop, with the corrected code block unchanged" width="784">
|
|
9
|
+
</p>
|
|
9
10
|
|
|
10
|
-
The prose is uwufied, but the
|
|
11
|
+
The prose is uwufied, but the inline code, numbers and the fixed loop stay exactly as they should be.
|
|
11
12
|
|
|
12
13
|
## Quick start
|
|
13
14
|
|
|
@@ -19,6 +20,10 @@ Restart omp. uwu mode is **on** by default in every session.
|
|
|
19
20
|
|
|
20
21
|
- `/uwu` toggles it
|
|
21
22
|
- `/uwu on` / `/uwu off` sets it explicitly
|
|
23
|
+
- `/uwu rewrite` uses omp's finalized-message rewrite hook when available
|
|
24
|
+
- `/uwu prompt` asks the model to write in uwu style while it streams
|
|
25
|
+
|
|
26
|
+
The TUI also shows a small, theme-accent-colored `(◕ᴗ◕✿) uwu` status badge. Other clients and plain output are not colorized.
|
|
22
27
|
|
|
23
28
|
## What gets uwufied, and what doesn't
|
|
24
29
|
|
|
@@ -30,18 +35,23 @@ Restart omp. uwu mode is **on** by default in every session.
|
|
|
30
35
|
| | Tool-call arguments and file contents the agent writes or edits |
|
|
31
36
|
| | Commit messages and prompts sent to subagents |
|
|
32
37
|
|
|
33
|
-
Style rules
|
|
38
|
+
Style rules include `r`/`l` → `w`, occasional `th` → `d`, `na/ne/no` → `nya/nye/nyo`, occasional stutter, and a broad rotating selection of kaomoji. The kaomoji selection includes examples from [kaomoji.you](https://kaomoji.you/), such as `٩(◕‿◕。)۶`, `(ฅ^•ﻌ•^ฅ)`, and `(づ。◕‿‿◕。)づ`. Meaning, numbers and warnings must stay readable.
|
|
39
|
+
|
|
40
|
+
Here's the same prompt and model, with the mode off and on:
|
|
41
|
+
|
|
42
|
+

|
|
43
|
+
|
|
44
|
+
These are real, unedited replies from `anthropic/claude-opus-5-5`, rendered as images. Both code blocks are identical byte for byte. The transcripts are in [`demo/`](demo).
|
|
34
45
|
|
|
35
46
|
## How it works
|
|
36
47
|
|
|
37
|
-
omp
|
|
48
|
+
On the first turn, omp-uwu uses the prompt style until the host demonstrates support for the awaited `assistant_message` hook; after that, the default **rewrite** style deterministically uwufies finalized assistant text before it is added to history and context. This first-turn check makes the experience work on both older and newer omp builds. Only text in existing text blocks is changed; code/tool blocks and their metadata stay untouched. Rewrites are markdown-aware and preserve fenced/inline code, URLs, paths, numbers, quoted text, identifiers, and safety-critical words.
|
|
38
49
|
|
|
39
|
-
|
|
50
|
+
The hook runs after streaming has finished. The TUI refreshes the current assistant message at completion, but clients that render only streamed chunks may continue showing the original text. If the hook is unavailable (including omp `18.4.3`), prompt style remains enabled as the compatibility fallback. `/uwu prompt` selects live prompt styling directly.
|
|
40
51
|
|
|
41
|
-
- **
|
|
42
|
-
- **
|
|
43
|
-
- **
|
|
44
|
-
- **Session-local toggle.** The on/off state isn't saved. Each new session starts with uwu mode on.
|
|
52
|
+
- **Deterministic rewrite.** No style instruction/token overhead once the hook is detected.
|
|
53
|
+
- **Prompt style.** Model-dependent, with live uwu output while streaming.
|
|
54
|
+
- **Session-local toggle.** Mode and enabled state reset for each new session.
|
|
45
55
|
|
|
46
56
|
## Install options
|
|
47
57
|
|
|
@@ -63,14 +73,26 @@ omp plugin link . # use this checkout instead of the installed copy
|
|
|
63
73
|
omp -e ./src/index.ts # or load it for a single run
|
|
64
74
|
```
|
|
65
75
|
|
|
76
|
+
### Re-recording the demo
|
|
77
|
+
|
|
78
|
+
The GIF and the side-by-side image are rendered from real transcripts in `demo/`. `record.sh` captures one reply with the extension loaded and one without it. It runs in a throwaway agent dir, so your personal rules and extensions don't affect the replies. `render.py` then draws both images:
|
|
79
|
+
|
|
80
|
+
```sh
|
|
81
|
+
mkdir -p /tmp/omp-demo && cp ~/.omp/agent/agent.db* /tmp/omp-demo/
|
|
82
|
+
demo/record.sh
|
|
83
|
+
uv run --with pillow python demo/render.py
|
|
84
|
+
```
|
|
85
|
+
|
|
86
|
+
Model output varies between runs, so re-run `record.sh` until you get a reply that reads well.
|
|
87
|
+
|
|
66
88
|
### Releasing
|
|
67
89
|
|
|
68
90
|
CI (`.github/workflows/ci.yml`) runs `check` and the tests on pushes to `main` and on pull requests. Pushing a `v*` tag runs them again, checks that the tag matches `version` in `package.json`, and publishes to npm with provenance (`.github/workflows/publish.yml`).
|
|
69
91
|
|
|
70
92
|
```sh
|
|
71
93
|
# bump "version" in package.json, commit, then:
|
|
72
|
-
git tag v0.
|
|
73
|
-
git push origin main v0.
|
|
94
|
+
git tag v0.2.0
|
|
95
|
+
git push origin main v0.2.0
|
|
74
96
|
```
|
|
75
97
|
|
|
76
98
|
The workflow uses npm [trusted publishing](https://docs.npmjs.com/trusted-publishers/), so the repository has no npm token. npm only allows trusted publishing on a package that already exists, so the first version is published by hand with `npm publish --access public`. After that, go to the package's Settings → Trusted publishing on npmjs.com and add GitHub Actions with user `NaC-L`, repository `omp-uwu`, and workflow `publish.yml`.
|
package/package.json
CHANGED
|
@@ -1,14 +1,15 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "omp-uwu",
|
|
3
|
-
"version": "0.
|
|
4
|
-
"description": "omp extension
|
|
3
|
+
"version": "0.2.0",
|
|
4
|
+
"description": "omp extension for playful uwu chat styling, markdown-safe rewrites, and kawaii kaomoji",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"license": "MIT",
|
|
7
7
|
"keywords": [
|
|
8
8
|
"omp",
|
|
9
9
|
"oh-my-pi",
|
|
10
10
|
"omp-plugin",
|
|
11
|
-
"uwu"
|
|
11
|
+
"uwu",
|
|
12
|
+
"kaomoji"
|
|
12
13
|
],
|
|
13
14
|
"homepage": "https://github.com/NaC-L/omp-uwu#readme",
|
|
14
15
|
"bugs": "https://github.com/NaC-L/omp-uwu/issues",
|
package/src/index.ts
CHANGED
|
@@ -1,41 +1,106 @@
|
|
|
1
|
-
/**
|
|
2
|
-
* uwu mode — makes the agent's chat replies uwu-speak.
|
|
3
|
-
*
|
|
4
|
-
* Why system-prompt based: OMP gives extensions no hook to rewrite displayed
|
|
5
|
-
* assistant text (`message_end` receives a detached clone), so the style is
|
|
6
|
-
* requested from the model via `before_agent_start` → `systemPrompt`.
|
|
7
|
-
*
|
|
8
|
-
* Toggle with `/uwu` (or `/uwu on` / `/uwu off`). Enabled by default.
|
|
9
|
-
*/
|
|
10
1
|
import type { ExtensionAPI } from "@oh-my-pi/pi-coding-agent";
|
|
2
|
+
import { uwufy } from "./uwufy.ts";
|
|
11
3
|
|
|
12
4
|
export const UWU_PROMPT = `
|
|
13
5
|
# uwu mode (display style only)
|
|
14
6
|
Write your prose replies to the user in playful "uwu" speak:
|
|
15
7
|
- replace r/l with w in most words ("weawwy", "hewwo"), "th" → "d" occasionally, "na/ne/no" → "nya/nye/nyo" sometimes
|
|
16
|
-
- occasional stutter ("h-hewwo") and cute emoticons
|
|
17
|
-
- keep it readable; never let the style hide meaning, numbers
|
|
8
|
+
- occasional stutter ("h-hewwo") and cute kaomoji/emoticons, varied naturally: (◕ᴗ◕✿), (o^▽^o), (´。• ω •。\u0060), ٩(◕‿◕。)۶, (✧ω✧), (๑˃ᴗ˂)ﻭ, ヽ(・∀・)ノ, (っ˘ω˘ς), (。•̀ᴗ-)✧, (づ。◕‿‿◕。)づ, (ฅ^•ﻌ•^ฅ), (๑>◡<๑), owo, >w<, ^w^, :3; at most one at a sentence end
|
|
9
|
+
- keep it readable; never let the style hide meaning, numbers or warnings
|
|
10
|
+
- colors are only for UI chrome; never insert ANSI color codes into chat text
|
|
18
11
|
|
|
19
12
|
NEVER uwufy any of these — keep them exact and byte-for-byte correct:
|
|
20
13
|
- code, code blocks, inline code, shell commands, file paths, URLs, identifiers, config keys, error messages you quote
|
|
21
|
-
- tool call arguments, file contents
|
|
14
|
+
- tool call arguments, file contents the agent writes or edits, commit messages and prompts sent to subagents
|
|
22
15
|
The style applies only to natural-language chat text shown to the user. Task quality and correctness are unchanged.
|
|
23
16
|
`.trim();
|
|
24
17
|
|
|
18
|
+
type RewriteEvent = {
|
|
19
|
+
message: { role: string; content: Array<{ type: string; text?: string; [key: string]: unknown }> };
|
|
20
|
+
};
|
|
21
|
+
type MessageEndEvent = { message: { role: string; stopReason?: string } };
|
|
22
|
+
type UiContext = {
|
|
23
|
+
mode?: string;
|
|
24
|
+
agent?: { kind?: string };
|
|
25
|
+
ui?: {
|
|
26
|
+
notify(message: string, level: "info"): void;
|
|
27
|
+
setStatus?(key: string, text: string | undefined): void;
|
|
28
|
+
theme?: { fg(color: string, text: string): string };
|
|
29
|
+
};
|
|
30
|
+
};
|
|
31
|
+
|
|
25
32
|
export default function uwuExtension(pi: ExtensionAPI) {
|
|
26
|
-
|
|
27
|
-
|
|
28
|
-
|
|
29
|
-
|
|
30
|
-
|
|
31
|
-
|
|
32
|
-
|
|
33
|
-
|
|
34
|
-
|
|
35
|
-
|
|
36
|
-
|
|
37
|
-
|
|
38
|
-
|
|
39
|
-
|
|
40
|
-
|
|
33
|
+
let enabled = true;
|
|
34
|
+
let style: "rewrite" | "prompt" = "rewrite";
|
|
35
|
+
let hookSupport: boolean | undefined;
|
|
36
|
+
let promptAddedThisTurn = false;
|
|
37
|
+
let notifiedFallback = false;
|
|
38
|
+
|
|
39
|
+
const updateBadge = (ctx: UiContext) => {
|
|
40
|
+
if (ctx.mode !== "tui" || ctx.agent?.kind === "sub" || !ctx.ui?.setStatus) return;
|
|
41
|
+
const badge = enabled ? "(◕ᴗ◕✿) uwu" : undefined;
|
|
42
|
+
ctx.ui.setStatus("omp-uwu", badge && ctx.ui.theme ? ctx.ui.theme.fg("accent", badge) : badge);
|
|
43
|
+
};
|
|
44
|
+
|
|
45
|
+
pi.registerCommand("uwu", {
|
|
46
|
+
description: "UwU chat style and mode (usage: /uwu [on|off|prompt|rewrite])",
|
|
47
|
+
handler: async (args, rawCtx) => {
|
|
48
|
+
const ctx = rawCtx as UiContext;
|
|
49
|
+
const arg = String(args ?? "").trim().toLowerCase();
|
|
50
|
+
if (arg === "prompt" || arg === "rewrite") {
|
|
51
|
+
style = arg;
|
|
52
|
+
enabled = true;
|
|
53
|
+
} else if (arg === "on") enabled = true;
|
|
54
|
+
else if (arg === "off") enabled = false;
|
|
55
|
+
else if (!arg) enabled = !enabled;
|
|
56
|
+
else {
|
|
57
|
+
ctx.ui?.notify("Usage: /uwu [on|off|prompt|rewrite]", "info");
|
|
58
|
+
return;
|
|
59
|
+
}
|
|
60
|
+
updateBadge(ctx);
|
|
61
|
+
ctx.ui?.notify(enabled ? `uwu mode ${style === "rewrite" ? "rewrite" : "prompt"}! (◕ᴗ◕✿)` : "uwu mode off", "info");
|
|
62
|
+
},
|
|
63
|
+
});
|
|
64
|
+
|
|
65
|
+
pi.on("session_start", (_event, rawCtx) => updateBadge(rawCtx as UiContext));
|
|
66
|
+
|
|
67
|
+
pi.on("before_agent_start", async (event, rawCtx) => {
|
|
68
|
+
promptAddedThisTurn = false;
|
|
69
|
+
const ctx = rawCtx as UiContext;
|
|
70
|
+
if (ctx.agent?.kind === "sub") return undefined;
|
|
71
|
+
// Start with the prompt until this host proves it has the finalized hook.
|
|
72
|
+
// That styles the first reply on both old and new omp versions.
|
|
73
|
+
const needsPrompt = style === "prompt" || hookSupport !== true;
|
|
74
|
+
if (!enabled || !needsPrompt) return undefined;
|
|
75
|
+
promptAddedThisTurn = true;
|
|
76
|
+
return { systemPrompt: [...event.systemPrompt, UWU_PROMPT] };
|
|
77
|
+
});
|
|
78
|
+
|
|
79
|
+
// The hook is newer than the bundled 18.4.3 types, so register structurally.
|
|
80
|
+
// ExtensionAPI.on stores event names as strings; old hosts simply never emit it.
|
|
81
|
+
(pi.on as unknown as (name: string, handler: (event: RewriteEvent) => unknown) => void)(
|
|
82
|
+
"assistant_message",
|
|
83
|
+
(event) => {
|
|
84
|
+
hookSupport = true;
|
|
85
|
+
if (!enabled || style !== "rewrite" || promptAddedThisTurn || event.message.role !== "assistant") return;
|
|
86
|
+
let changed = false;
|
|
87
|
+
const content = event.message.content.map((block) => {
|
|
88
|
+
if (block.type !== "text" || typeof block.text !== "string") return block;
|
|
89
|
+
const text = uwufy(block.text);
|
|
90
|
+
if (text === block.text) return block;
|
|
91
|
+
changed = true;
|
|
92
|
+
return { ...block, text };
|
|
93
|
+
});
|
|
94
|
+
return changed ? { content } : undefined;
|
|
95
|
+
},
|
|
96
|
+
);
|
|
97
|
+
|
|
98
|
+
pi.on("message_end", (rawEvent) => {
|
|
99
|
+
const message = (rawEvent as MessageEndEvent).message;
|
|
100
|
+
if (message.role === "assistant" && message.stopReason !== "aborted" && message.stopReason !== "error" && hookSupport === undefined) {
|
|
101
|
+
hookSupport = false;
|
|
102
|
+
// The current response has already streamed; prompt fallback applies next turn.
|
|
103
|
+
if (!notifiedFallback) notifiedFallback = true;
|
|
104
|
+
}
|
|
105
|
+
});
|
|
41
106
|
}
|
package/src/uwufy.ts
ADDED
|
@@ -0,0 +1,331 @@
|
|
|
1
|
+
/**
|
|
2
|
+
* Deterministic, markdown-aware uwu rewriter for finalized assistant text.
|
|
3
|
+
*
|
|
4
|
+
* Only natural-language words are touched. Everything that has to stay exact is
|
|
5
|
+
* left byte-for-byte alone: fenced/indented code, math blocks, inline code,
|
|
6
|
+
* URLs and link targets, HTML tags, double-quoted spans, and any token that
|
|
7
|
+
* looks like an identifier, path, number, flag, acronym, or warning word.
|
|
8
|
+
*
|
|
9
|
+
* Every "sometimes" decision is keyed on the word's position (paragraph,
|
|
10
|
+
* sentence, word) rather than on its text, and no transform adds or removes a
|
|
11
|
+
* word token. Rewriting already-rewritten text therefore makes the same
|
|
12
|
+
* decisions and finds nothing left to do: `uwufy(uwufy(x)) === uwufy(x)`.
|
|
13
|
+
* That matters because rewritten replies go back into the model's context, and
|
|
14
|
+
* the model may start answering in uwu-speak on its own.
|
|
15
|
+
*/
|
|
16
|
+
|
|
17
|
+
/** Emoticons the rewriter appends at sentence ends. */
|
|
18
|
+
export const EMOTICONS = [
|
|
19
|
+
"uwu", "owo", ">w<", "^w^", ":3", "(◕ᴗ◕✿)",
|
|
20
|
+
"(o^▽^o)", "(´。• ω •。`)", "٩(◕‿◕。)۶", "o(≧▽≦)o", "(✧ω✧)",
|
|
21
|
+
"(๑˃ᴗ˂)ﻭ", "(ᵔ◡ᵔ)", "ヽ(・∀・)ノ", "(´• ω •`)", "(。•́︿•̀。)",
|
|
22
|
+
"(ノ_<。)", "(っ˘ω˘ς)", "(≧◡≦)", "٩(◕‿◕)۶", "(o˘◡˘o)",
|
|
23
|
+
"(♡˙︶˙♡)", "ヽ(♡‿♡)ノ", "(´꒳`)♡", "(ノ´ з `)ノ", "ヾ(・ω・*)",
|
|
24
|
+
"(⌒ω⌒)ノ", "(ᵔ⩊ᵔ)", "(=^・ω・^=)", "(ฅ^•ﻌ•^ฅ)", "(ᵕ—ᴗ—)",
|
|
25
|
+
"(。•̀ᴗ-)✧", "(づ。◕‿‿◕。)づ", "(っ´▽`)っ", "( ˶ˆᗜˆ˵ )", "(๑>◡<๑)",
|
|
26
|
+
] as const;
|
|
27
|
+
|
|
28
|
+
/** Tokens treated as emoticons: never counted as words, never rewritten. */
|
|
29
|
+
const EMOTICON_TOKENS = new Set<string>([
|
|
30
|
+
...EMOTICONS,
|
|
31
|
+
"UwU",
|
|
32
|
+
"OwO",
|
|
33
|
+
"^_^",
|
|
34
|
+
"^^",
|
|
35
|
+
":)",
|
|
36
|
+
":D",
|
|
37
|
+
"<3",
|
|
38
|
+
"x3",
|
|
39
|
+
]);
|
|
40
|
+
|
|
41
|
+
/** Words that carry meaning a reader must not miss. Kept exact. */
|
|
42
|
+
const KEEP_WORDS = new Set([
|
|
43
|
+
"no",
|
|
44
|
+
"not",
|
|
45
|
+
"nor",
|
|
46
|
+
"never",
|
|
47
|
+
"none",
|
|
48
|
+
"nothing",
|
|
49
|
+
"nowhere",
|
|
50
|
+
"neither",
|
|
51
|
+
"cannot",
|
|
52
|
+
"can't",
|
|
53
|
+
"don't",
|
|
54
|
+
"doesn't",
|
|
55
|
+
"didn't",
|
|
56
|
+
"won't",
|
|
57
|
+
"wouldn't",
|
|
58
|
+
"shouldn't",
|
|
59
|
+
"mustn't",
|
|
60
|
+
"isn't",
|
|
61
|
+
"aren't",
|
|
62
|
+
"wasn't",
|
|
63
|
+
"weren't",
|
|
64
|
+
"haven't",
|
|
65
|
+
"hasn't",
|
|
66
|
+
"hadn't",
|
|
67
|
+
"couldn't",
|
|
68
|
+
"warning",
|
|
69
|
+
"caution",
|
|
70
|
+
"danger",
|
|
71
|
+
"dangerous",
|
|
72
|
+
"careful",
|
|
73
|
+
"irreversible",
|
|
74
|
+
"destructive",
|
|
75
|
+
"error",
|
|
76
|
+
"errors",
|
|
77
|
+
"fail",
|
|
78
|
+
"fails",
|
|
79
|
+
"failed",
|
|
80
|
+
"failure",
|
|
81
|
+
]);
|
|
82
|
+
|
|
83
|
+
const TH_WORDS = new Set(["the", "this", "that", "these", "those", "them", "then", "there", "their", "they", "than"]);
|
|
84
|
+
|
|
85
|
+
// Probabilities for the "sometimes" transforms.
|
|
86
|
+
const P_TH = 0.35;
|
|
87
|
+
const P_NY = 0.5;
|
|
88
|
+
const P_STUTTER = 0.12;
|
|
89
|
+
const P_EMOTICON = 0.3;
|
|
90
|
+
|
|
91
|
+
/**
|
|
92
|
+
* Inline spans that are never rewritten. Only the first alternative captures,
|
|
93
|
+
* so `\1` is the opening backtick run of a code span.
|
|
94
|
+
*/
|
|
95
|
+
const PROTECTED_INLINE = new RegExp(
|
|
96
|
+
[
|
|
97
|
+
String.raw`(\`+)[\s\S]*?\1(?!\`)`, // code span
|
|
98
|
+
String.raw`\`.*$`, // unmatched backtick: keep the rest of the line
|
|
99
|
+
String.raw`\]\([^)\s]*(?:\s+"[^"]*")?\)`, // markdown link target
|
|
100
|
+
String.raw`<\/?[A-Za-z][^>\n]*>`, // HTML tag or autolink
|
|
101
|
+
String.raw`\b(?:[a-z][a-z0-9+.-]*:\/\/|www\.)\S*`, // bare URL
|
|
102
|
+
String.raw`"[^"\n]*"`, // "quoted text"
|
|
103
|
+
String.raw`“[^”\n]*”`, // “quoted text”
|
|
104
|
+
].join("|"),
|
|
105
|
+
"gi",
|
|
106
|
+
);
|
|
107
|
+
|
|
108
|
+
const FENCE_OPEN = /^ {0,3}(`{3,}|~{3,})/;
|
|
109
|
+
const FENCE_CLOSE = /^ {0,3}(`{3,}|~{3,})\s*$/;
|
|
110
|
+
const MATH_FENCE = /^\s*\$\$/;
|
|
111
|
+
const INDENTED_CODE = /^(?: {4}|\t)/;
|
|
112
|
+
const LIST_ITEM = /^\s*(?:[-*+]|\d+[.)])\s/;
|
|
113
|
+
const HTML_LINE = /^\s*<[A-Za-z/!]/;
|
|
114
|
+
const LINK_DEFINITION = /^ {0,3}\[[^\]]+\]:\s/;
|
|
115
|
+
const HEADING = /^ {0,3}#{1,6}(?:\s|$)/;
|
|
116
|
+
const TABLE_ROW = /^\s*\|/;
|
|
117
|
+
|
|
118
|
+
const TOKEN_PARTS = /^([(\[{"'“‘*_~]*)(.*?)([)\]}"'”’*_~.,;:!?…]*)$/su;
|
|
119
|
+
const WORD = /^[\p{L}\p{M}][\p{L}\p{M}'’-]*$/u;
|
|
120
|
+
const SENTENCE_END = /[.!?…]/;
|
|
121
|
+
const STUTTERED = /^(\p{L})-\1/iu;
|
|
122
|
+
|
|
123
|
+
type LineKind = "code" | "blank" | "prose";
|
|
124
|
+
|
|
125
|
+
interface ProseLine {
|
|
126
|
+
index: number;
|
|
127
|
+
/** Headings and tables get word rewrites but no stutters or emoticons. */
|
|
128
|
+
decorate: boolean;
|
|
129
|
+
}
|
|
130
|
+
|
|
131
|
+
/** Rewrite the prose of a markdown string in uwu-speak. */
|
|
132
|
+
export function uwufy(text: string): string {
|
|
133
|
+
const lines = text.split("\n");
|
|
134
|
+
const kinds = classifyLines(lines);
|
|
135
|
+
|
|
136
|
+
const paragraphs: ProseLine[][] = [];
|
|
137
|
+
let current: ProseLine[] | undefined;
|
|
138
|
+
for (const [index, line] of lines.entries()) {
|
|
139
|
+
if (kinds[index] !== "prose") {
|
|
140
|
+
current = undefined;
|
|
141
|
+
continue;
|
|
142
|
+
}
|
|
143
|
+
const heading = HEADING.test(line);
|
|
144
|
+
// Headings, list items and table rows each start their own paragraph.
|
|
145
|
+
if (!current || heading || LIST_ITEM.test(line) || TABLE_ROW.test(line)) {
|
|
146
|
+
current = [];
|
|
147
|
+
paragraphs.push(current);
|
|
148
|
+
}
|
|
149
|
+
current.push({ index, decorate: !heading && !TABLE_ROW.test(line) });
|
|
150
|
+
if (heading) current = undefined;
|
|
151
|
+
}
|
|
152
|
+
|
|
153
|
+
const out = [...lines];
|
|
154
|
+
for (const [paragraphIndex, paragraph] of paragraphs.entries()) {
|
|
155
|
+
// Count once, then rewrite: the word count seeds this paragraph's decisions,
|
|
156
|
+
// which keeps them varied between paragraphs yet stable across re-runs.
|
|
157
|
+
const wordCount = walkParagraph(lines, paragraph, () => undefined);
|
|
158
|
+
const seed = `${paragraphIndex}|${wordCount}`;
|
|
159
|
+
walkParagraph(lines, paragraph, (event) => rewriteToken(event, seed), out);
|
|
160
|
+
}
|
|
161
|
+
return out.join("\n");
|
|
162
|
+
}
|
|
163
|
+
|
|
164
|
+
function classifyLines(lines: string[]): LineKind[] {
|
|
165
|
+
const kinds: LineKind[] = [];
|
|
166
|
+
let fence: { char: string; length: number } | undefined;
|
|
167
|
+
let math = false;
|
|
168
|
+
let previous: LineKind = "blank";
|
|
169
|
+
for (const line of lines) {
|
|
170
|
+
let kind: LineKind;
|
|
171
|
+
if (fence) {
|
|
172
|
+
const close = FENCE_CLOSE.exec(line);
|
|
173
|
+
if (close?.[1] && close[1][0] === fence.char && close[1].length >= fence.length) fence = undefined;
|
|
174
|
+
kind = "code";
|
|
175
|
+
} else if (math) {
|
|
176
|
+
if (MATH_FENCE.test(line)) math = false;
|
|
177
|
+
kind = "code";
|
|
178
|
+
} else if (FENCE_OPEN.test(line)) {
|
|
179
|
+
const open = FENCE_OPEN.exec(line)?.[1] ?? "```";
|
|
180
|
+
fence = { char: open[0] ?? "`", length: open.length };
|
|
181
|
+
kind = "code";
|
|
182
|
+
} else if (MATH_FENCE.test(line)) {
|
|
183
|
+
// A one-line `$$ x $$` block opens and closes on the same line.
|
|
184
|
+
math = !/^\s*\$\$.*\$\$\s*$/.test(line) || line.trim() === "$$";
|
|
185
|
+
kind = "code";
|
|
186
|
+
} else if (line.trim() === "") {
|
|
187
|
+
kind = "blank";
|
|
188
|
+
} else if (
|
|
189
|
+
(INDENTED_CODE.test(line) && !LIST_ITEM.test(line) && previous !== "prose") ||
|
|
190
|
+
HTML_LINE.test(line) ||
|
|
191
|
+
LINK_DEFINITION.test(line)
|
|
192
|
+
) {
|
|
193
|
+
kind = "code";
|
|
194
|
+
} else {
|
|
195
|
+
kind = "prose";
|
|
196
|
+
}
|
|
197
|
+
kinds.push(kind);
|
|
198
|
+
previous = kind;
|
|
199
|
+
}
|
|
200
|
+
return kinds;
|
|
201
|
+
}
|
|
202
|
+
|
|
203
|
+
interface TokenEvent {
|
|
204
|
+
token: string;
|
|
205
|
+
/** Text after the token within its prose segment, for emoticon lookahead. */
|
|
206
|
+
rest: string;
|
|
207
|
+
/** Whether the segment ends the line (nothing protected follows it). */
|
|
208
|
+
lineEnd: boolean;
|
|
209
|
+
decorate: boolean;
|
|
210
|
+
sentence: number;
|
|
211
|
+
wordInSentence: number;
|
|
212
|
+
word: number;
|
|
213
|
+
}
|
|
214
|
+
|
|
215
|
+
/**
|
|
216
|
+
* Walk the word tokens of one paragraph. When `out` is given, each prose
|
|
217
|
+
* segment is rebuilt from the callback's replacements. Returns the word count.
|
|
218
|
+
*/
|
|
219
|
+
function walkParagraph(
|
|
220
|
+
lines: string[],
|
|
221
|
+
paragraph: ProseLine[],
|
|
222
|
+
visit: (event: TokenEvent) => string | undefined,
|
|
223
|
+
out?: string[],
|
|
224
|
+
): number {
|
|
225
|
+
let sentence = 0;
|
|
226
|
+
let wordInSentence = 0;
|
|
227
|
+
let word = 0;
|
|
228
|
+
for (const { index, decorate } of paragraph) {
|
|
229
|
+
const line = lines[index] ?? "";
|
|
230
|
+
let rebuilt = "";
|
|
231
|
+
let cursor = 0;
|
|
232
|
+
const segments: Array<{ prose: boolean; text: string }> = [];
|
|
233
|
+
for (const match of line.matchAll(PROTECTED_INLINE)) {
|
|
234
|
+
const start = match.index ?? 0;
|
|
235
|
+
if (start > cursor) segments.push({ prose: true, text: line.slice(cursor, start) });
|
|
236
|
+
segments.push({ prose: false, text: match[0] });
|
|
237
|
+
cursor = start + match[0].length;
|
|
238
|
+
}
|
|
239
|
+
if (cursor < line.length) segments.push({ prose: true, text: line.slice(cursor) });
|
|
240
|
+
|
|
241
|
+
for (const [segmentIndex, segment] of segments.entries()) {
|
|
242
|
+
if (!segment.prose) {
|
|
243
|
+
rebuilt += segment.text;
|
|
244
|
+
continue;
|
|
245
|
+
}
|
|
246
|
+
const lineEnd = segmentIndex === segments.length - 1;
|
|
247
|
+
rebuilt += segment.text.replace(/\S+/g, (token, offset: number) => {
|
|
248
|
+
if (EMOTICON_TOKENS.has(token)) return token;
|
|
249
|
+
const [, , core = "", trail = ""] = TOKEN_PARTS.exec(token) ?? [];
|
|
250
|
+
const isWord = WORD.test(core);
|
|
251
|
+
let replacement: string | undefined;
|
|
252
|
+
if (isWord) {
|
|
253
|
+
replacement = visit({
|
|
254
|
+
token,
|
|
255
|
+
rest: segment.text.slice(offset + token.length),
|
|
256
|
+
lineEnd,
|
|
257
|
+
decorate,
|
|
258
|
+
sentence,
|
|
259
|
+
wordInSentence,
|
|
260
|
+
word,
|
|
261
|
+
});
|
|
262
|
+
word++;
|
|
263
|
+
wordInSentence++;
|
|
264
|
+
}
|
|
265
|
+
if ((isWord || core === "") && SENTENCE_END.test(trail)) {
|
|
266
|
+
sentence++;
|
|
267
|
+
wordInSentence = 0;
|
|
268
|
+
}
|
|
269
|
+
return replacement ?? token;
|
|
270
|
+
});
|
|
271
|
+
}
|
|
272
|
+
if (out) out[index] = rebuilt;
|
|
273
|
+
}
|
|
274
|
+
return word;
|
|
275
|
+
}
|
|
276
|
+
|
|
277
|
+
function rewriteToken(event: TokenEvent, seed: string): string {
|
|
278
|
+
const [, lead = "", core = "", trail = ""] = TOKEN_PARTS.exec(event.token) ?? [];
|
|
279
|
+
const key = `${seed}|${event.sentence}|${event.word}`;
|
|
280
|
+
let result = event.token;
|
|
281
|
+
|
|
282
|
+
if (!isKeptWord(core)) {
|
|
283
|
+
let word = core;
|
|
284
|
+
if (TH_WORDS.has(word.toLowerCase()) && roll(`${key}|th`) < P_TH) {
|
|
285
|
+
word = (word[0] === "T" ? "D" : "d") + word.slice(2);
|
|
286
|
+
}
|
|
287
|
+
word = word.replace(/[rl]/g, "w").replace(/[RL]/g, "W");
|
|
288
|
+
if (roll(`${key}|ny`) < P_NY) word = word.replace(/([nN])(?=[aeo])/g, "$1y");
|
|
289
|
+
if (
|
|
290
|
+
event.decorate &&
|
|
291
|
+
event.wordInSentence === 0 &&
|
|
292
|
+
word.length >= 3 &&
|
|
293
|
+
!STUTTERED.test(word) &&
|
|
294
|
+
roll(`${key}|stutter`) < P_STUTTER
|
|
295
|
+
) {
|
|
296
|
+
word = `${word[0]}-${word}`;
|
|
297
|
+
}
|
|
298
|
+
// Never turn a word into something that reads as an emoticon (e.g. "ulu").
|
|
299
|
+
if (!EMOTICON_TOKENS.has(word)) result = lead + word + trail;
|
|
300
|
+
}
|
|
301
|
+
|
|
302
|
+
if (
|
|
303
|
+
event.decorate &&
|
|
304
|
+
SENTENCE_END.test(trail) &&
|
|
305
|
+
(/^\s/.test(event.rest) || (event.rest === "" && event.lineEnd)) &&
|
|
306
|
+
!EMOTICON_TOKENS.has(/^\s+(\S+)/.exec(event.rest)?.[1] ?? "") &&
|
|
307
|
+
roll(`${seed}|${event.sentence}|emoticon`) < P_EMOTICON
|
|
308
|
+
) {
|
|
309
|
+
result += ` ${EMOTICONS[Math.floor(roll(`${seed}|${event.sentence}|pick`) * EMOTICONS.length)]}`;
|
|
310
|
+
}
|
|
311
|
+
return result;
|
|
312
|
+
}
|
|
313
|
+
|
|
314
|
+
/** Acronyms, camelCase names, flags and meaning-critical words stay exact. */
|
|
315
|
+
function isKeptWord(core: string): boolean {
|
|
316
|
+
if (KEEP_WORDS.has(core.toLowerCase().replaceAll("’", "'"))) return true;
|
|
317
|
+
if (core.endsWith("-") || core.includes("--")) return true;
|
|
318
|
+
if (/\p{Ll}\p{Lu}/u.test(core)) return true; // camelCase, GitHub, iOS
|
|
319
|
+
const letters = core.replace(/[^\p{L}]/gu, "");
|
|
320
|
+
return letters.length >= 2 && letters === letters.toUpperCase() && letters !== letters.toLowerCase();
|
|
321
|
+
}
|
|
322
|
+
|
|
323
|
+
/** FNV-1a hash of `key`, scaled to [0, 1). */
|
|
324
|
+
function roll(key: string): number {
|
|
325
|
+
let hash = 0x811c9dc5;
|
|
326
|
+
for (let i = 0; i < key.length; i++) {
|
|
327
|
+
hash ^= key.charCodeAt(i);
|
|
328
|
+
hash = Math.imul(hash, 0x01000193);
|
|
329
|
+
}
|
|
330
|
+
return (hash >>> 0) / 2 ** 32;
|
|
331
|
+
}
|