swearbench 0.2.1__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,18 @@
1
+ name: publish
2
+
3
+ on:
4
+ push:
5
+ tags: ["v*"]
6
+
7
+ jobs:
8
+ pypi:
9
+ runs-on: ubuntu-latest
10
+ environment: pypi
11
+ permissions:
12
+ id-token: write
13
+ contents: read
14
+ steps:
15
+ - uses: actions/checkout@v4
16
+ - uses: astral-sh/setup-uv@v6
17
+ - run: uv build
18
+ - uses: pypa/gh-action-pypi-publish@release/v1
@@ -0,0 +1,5 @@
1
+ swearbench-out/
2
+ __pycache__/
3
+ *.egg-info/
4
+ dist/
5
+ .venv/
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 AlenHay
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
@@ -0,0 +1,130 @@
1
+ Metadata-Version: 2.5
2
+ Name: swearbench
3
+ Version: 0.2.1
4
+ Summary: Rank AI coding models by how much they made you swear and how often their work shipped, from your own local chat logs.
5
+ License-Expression: MIT
6
+ License-File: LICENSE
7
+ Requires-Python: >=3.10
8
+ Description-Content-Type: text/markdown
9
+
10
+ # SwearBench
11
+
12
+ **Which AI coding model made you swear the least, and still shipped?**
13
+
14
+ Public benchmarks measure what models can do. SwearBench measures how they made *you* feel: it reads your
15
+ own local agent logs, finds every message you sent in reaction to a model's reply, has an LLM judge label how
16
+ mad you were and why, checks whether the work ended up accepted, and ranks the models.
17
+
18
+ Example from the author's own logs (~4,600 messages, July–October 2026):
19
+
20
+ ![swearing vs. result](docs/example-chart.svg)
21
+
22
+ ![leaderboard card](docs/example-card.svg)
23
+
24
+ ## Run it
25
+
26
+ ```sh
27
+ uvx swearbench
28
+ ```
29
+
30
+ (or `pipx run swearbench`, or `pip install swearbench`)
31
+
32
+ It prints a leaderboard and writes `swearbench-out/report.md`, `chart.svg` (swearing vs. result), `card.svg` and `results.json`.
33
+ Before anything leaves your machine it tells you how many messages it will send to the judge and asks.
34
+
35
+ | Flag | |
36
+ |---|---|
37
+ | `--dry-run` | show what logs were found and how many messages would be judged |
38
+ | `--judge claude-cli\|codex-cli\|anthropic\|openai\|command` | who labels your messages (default: first of `claude`, `codex`, `ANTHROPIC_API_KEY`, `OPENAI_API_KEY`) |
39
+ | `--judge-model ID` | model for the judge |
40
+ | `--judge command --judge-cmd "ollama run qwen3"` | any command that reads a prompt on stdin and prints the reply; keeps everything local |
41
+ | `--since 2026-09-01` | only recent reactions |
42
+ | `--exclude MODEL` | drop a model from the ranking |
43
+ | `--no-quotes` | leave the hall of shame out of the report |
44
+
45
+ Labels are cached in `~/.cache/swearbench/`, so re-runs only judge new messages.
46
+
47
+ ## Where it looks
48
+
49
+ | Tool | Location |
50
+ |---|---|
51
+ | Claude Code | `~/.claude/projects/**/*.jsonl` |
52
+ | Codex CLI | `~/.codex/sessions/**/*.jsonl` |
53
+ | OpenCode | `~/.local/share/opencode/opencode.db` |
54
+ | T3 Code | `~/.config/t3/userdata/state*.sqlite` (used for exact per-turn model attribution when present) |
55
+
56
+ Only messages **you** typed count. Subagent transcripts, headless runs (`claude -p`, `codex exec`) and messages
57
+ the judge flags as agent-written are skipped. Models that only ever ran as subagents are listed, never ranked.
58
+
59
+ ## How it scores
60
+
61
+ Every message you send is charged to the model whose reply you were answering. The judge labels it:
62
+
63
+ - **anger** 0–4, and **who it's aimed at**: the model, or something external (swearing at your cloud
64
+ provider is not the model's fault; "this looks fucking cool" is not anger)
65
+ - **how** you got mad, weighted by how bad it is:
66
+
67
+ | Mode | Weight | | Mode | Weight |
68
+ |---|---|---|---|---|
69
+ | fabrication (claimed success that wasn't) | 3 | | incomplete / lazy | 1.5 |
70
+ | overreach (did things you didn't ask) | 3 | | insult | 1.5 |
71
+ | giving up on it | 3 | | profanity | 1 |
72
+ | regression (broke what worked) | 2.5 | | shouting | 0.5 |
73
+ | ignored an instruction | 2 | | sarcasm | 0.5 |
74
+ | made you repeat yourself | 2 | | slow / verbose | 0.5 |
75
+ | | | | taste (critique while iterating on looks) | 0.5 |
76
+
77
+ - **satisfaction** −2 (rejects the work) to +2 (praise)
78
+ - whether it **blames earlier work** (something shipped before the last reply is broken or missing)
79
+
80
+ Rage for a message is built to match how it felt, not how often it happened:
81
+
82
+ - **Severity beats frequency.** Rage = anger² + mode weights, so one 4/4 blowup (16) outweighs four
83
+ 1/4 grumbles (4). The report counts 4/4 blowups per model.
84
+ - **Taste isn't failure.** "That looks lame" while iterating on a design, with no broken rule, lie,
85
+ regression or ignored instruction, counts a quarter.
86
+ - **Regret goes to whoever caused it.** A complaint about earlier work ("why did X disappear") is charged to
87
+ the models that worked in the same repo in the week before, split by how many turns each had there, not
88
+ to the model that happens to be fixing it.
89
+ - Each interrupt adds 1.5.
90
+
91
+ The headline is half friction, half result:
92
+
93
+ ```
94
+ Friction = 100 − 0.25 × rage per 100 turns + 10 × average satisfaction
95
+ Ships = accepted ÷ (accepted + rejected) sessions
96
+ SwearBench = (Friction + Ships) / 2
97
+ ```
98
+
99
+ Friction deliberately ignores how *often* you were annoyed (a model you use for lots of quick "merge it"
100
+ turns would look calm by volume alone); it only counts how much rage piled up per turn.
101
+ A session's verdict is your last reaction in it. Sessions that stop or move to another model without a
102
+ verdict (usage limits, "pick up the work" in a new thread) are left out of Ships rather than counted as
103
+ failures; switching away in anger still counts as a rejection.
104
+ Friction alone rewards a model that is pleasant but never finishes; Ships alone ignores what it cost you
105
+ to get there. The chart plots the two against each other.
106
+
107
+ Intervals are a 90% bootstrap over sessions; models with fewer than 40 reactions aren't ranked.
108
+ With T3 Code, the report also shows the share of sessions that ended in a merged PR. It is informational
109
+ only, since T3 records PRs only from when it started tracking them.
110
+
111
+ **Per token of work.** If the logs carry token usage, the report adds a second ranking: rage per million output
112
+ tokens the model produced in sessions you drove. A model that does twice the work per message gets credit for it.
113
+ Subagent token use is shown separately.
114
+
115
+ ## Caveats
116
+
117
+ - n = 1. It measures you, your tasks and your mood as much as the models. Models used in different months
118
+ did different work; the report tells you counts, not causes.
119
+ - The judge is a model too. If it belongs to a family being ranked, SwearBench says so; re-run with another
120
+ `--judge` and compare.
121
+ - Regret attribution is a heuristic: it blames whoever worked in the same repo during the previous week,
122
+ not the session that actually introduced the problem.
123
+ - The weights are opinions. Rage, taste and severity weights live at the top of `score.py`; change them
124
+ and re-run, labels are cached.
125
+ - Deleted or rotated logs mean missing data, especially for the per-token view.
126
+ - `report.md` quotes your own messages. Read it before you share it. The card and chart have no quotes.
127
+
128
+ ## License
129
+
130
+ MIT
@@ -0,0 +1,121 @@
1
+ # SwearBench
2
+
3
+ **Which AI coding model made you swear the least, and still shipped?**
4
+
5
+ Public benchmarks measure what models can do. SwearBench measures how they made *you* feel: it reads your
6
+ own local agent logs, finds every message you sent in reaction to a model's reply, has an LLM judge label how
7
+ mad you were and why, checks whether the work ended up accepted, and ranks the models.
8
+
9
+ Example from the author's own logs (~4,600 messages, July–October 2026):
10
+
11
+ ![swearing vs. result](docs/example-chart.svg)
12
+
13
+ ![leaderboard card](docs/example-card.svg)
14
+
15
+ ## Run it
16
+
17
+ ```sh
18
+ uvx swearbench
19
+ ```
20
+
21
+ (or `pipx run swearbench`, or `pip install swearbench`)
22
+
23
+ It prints a leaderboard and writes `swearbench-out/report.md`, `chart.svg` (swearing vs. result), `card.svg` and `results.json`.
24
+ Before anything leaves your machine it tells you how many messages it will send to the judge and asks.
25
+
26
+ | Flag | |
27
+ |---|---|
28
+ | `--dry-run` | show what logs were found and how many messages would be judged |
29
+ | `--judge claude-cli\|codex-cli\|anthropic\|openai\|command` | who labels your messages (default: first of `claude`, `codex`, `ANTHROPIC_API_KEY`, `OPENAI_API_KEY`) |
30
+ | `--judge-model ID` | model for the judge |
31
+ | `--judge command --judge-cmd "ollama run qwen3"` | any command that reads a prompt on stdin and prints the reply; keeps everything local |
32
+ | `--since 2026-09-01` | only recent reactions |
33
+ | `--exclude MODEL` | drop a model from the ranking |
34
+ | `--no-quotes` | leave the hall of shame out of the report |
35
+
36
+ Labels are cached in `~/.cache/swearbench/`, so re-runs only judge new messages.
37
+
38
+ ## Where it looks
39
+
40
+ | Tool | Location |
41
+ |---|---|
42
+ | Claude Code | `~/.claude/projects/**/*.jsonl` |
43
+ | Codex CLI | `~/.codex/sessions/**/*.jsonl` |
44
+ | OpenCode | `~/.local/share/opencode/opencode.db` |
45
+ | T3 Code | `~/.config/t3/userdata/state*.sqlite` (used for exact per-turn model attribution when present) |
46
+
47
+ Only messages **you** typed count. Subagent transcripts, headless runs (`claude -p`, `codex exec`) and messages
48
+ the judge flags as agent-written are skipped. Models that only ever ran as subagents are listed, never ranked.
49
+
50
+ ## How it scores
51
+
52
+ Every message you send is charged to the model whose reply you were answering. The judge labels it:
53
+
54
+ - **anger** 0–4, and **who it's aimed at**: the model, or something external (swearing at your cloud
55
+ provider is not the model's fault; "this looks fucking cool" is not anger)
56
+ - **how** you got mad, weighted by how bad it is:
57
+
58
+ | Mode | Weight | | Mode | Weight |
59
+ |---|---|---|---|---|
60
+ | fabrication (claimed success that wasn't) | 3 | | incomplete / lazy | 1.5 |
61
+ | overreach (did things you didn't ask) | 3 | | insult | 1.5 |
62
+ | giving up on it | 3 | | profanity | 1 |
63
+ | regression (broke what worked) | 2.5 | | shouting | 0.5 |
64
+ | ignored an instruction | 2 | | sarcasm | 0.5 |
65
+ | made you repeat yourself | 2 | | slow / verbose | 0.5 |
66
+ | | | | taste (critique while iterating on looks) | 0.5 |
67
+
68
+ - **satisfaction** −2 (rejects the work) to +2 (praise)
69
+ - whether it **blames earlier work** (something shipped before the last reply is broken or missing)
70
+
71
+ Rage for a message is built to match how it felt, not how often it happened:
72
+
73
+ - **Severity beats frequency.** Rage = anger² + mode weights, so one 4/4 blowup (16) outweighs four
74
+ 1/4 grumbles (4). The report counts 4/4 blowups per model.
75
+ - **Taste isn't failure.** "That looks lame" while iterating on a design, with no broken rule, lie,
76
+ regression or ignored instruction, counts a quarter.
77
+ - **Regret goes to whoever caused it.** A complaint about earlier work ("why did X disappear") is charged to
78
+ the models that worked in the same repo in the week before, split by how many turns each had there, not
79
+ to the model that happens to be fixing it.
80
+ - Each interrupt adds 1.5.
81
+
82
+ The headline is half friction, half result:
83
+
84
+ ```
85
+ Friction = 100 − 0.25 × rage per 100 turns + 10 × average satisfaction
86
+ Ships = accepted ÷ (accepted + rejected) sessions
87
+ SwearBench = (Friction + Ships) / 2
88
+ ```
89
+
90
+ Friction deliberately ignores how *often* you were annoyed (a model you use for lots of quick "merge it"
91
+ turns would look calm by volume alone); it only counts how much rage piled up per turn.
92
+ A session's verdict is your last reaction in it. Sessions that stop or move to another model without a
93
+ verdict (usage limits, "pick up the work" in a new thread) are left out of Ships rather than counted as
94
+ failures; switching away in anger still counts as a rejection.
95
+ Friction alone rewards a model that is pleasant but never finishes; Ships alone ignores what it cost you
96
+ to get there. The chart plots the two against each other.
97
+
98
+ Intervals are a 90% bootstrap over sessions; models with fewer than 40 reactions aren't ranked.
99
+ With T3 Code, the report also shows the share of sessions that ended in a merged PR. It is informational
100
+ only, since T3 records PRs only from when it started tracking them.
101
+
102
+ **Per token of work.** If the logs carry token usage, the report adds a second ranking: rage per million output
103
+ tokens the model produced in sessions you drove. A model that does twice the work per message gets credit for it.
104
+ Subagent token use is shown separately.
105
+
106
+ ## Caveats
107
+
108
+ - n = 1. It measures you, your tasks and your mood as much as the models. Models used in different months
109
+ did different work; the report tells you counts, not causes.
110
+ - The judge is a model too. If it belongs to a family being ranked, SwearBench says so; re-run with another
111
+ `--judge` and compare.
112
+ - Regret attribution is a heuristic: it blames whoever worked in the same repo during the previous week,
113
+ not the session that actually introduced the problem.
114
+ - The weights are opinions. Rage, taste and severity weights live at the top of `score.py`; change them
115
+ and re-run, labels are cached.
116
+ - Deleted or rotated logs mean missing data, especially for the per-token view.
117
+ - `report.md` quotes your own messages. Read it before you share it. The card and chart have no quotes.
118
+
119
+ ## License
120
+
121
+ MIT
@@ -0,0 +1,73 @@
1
+ <svg xmlns="http://www.w3.org/2000/svg" class="c" width="760" height="452" viewBox="0 0 760 452" role="img" aria-label="SwearBench leaderboard">
2
+ <style>
3
+ .c{--surface:#fcfcfb;--ink:#0b0b0b;--ink2:#52514e;--muted:#898781;--grid:#e1e0d9;--axis:#c3c2b7;
4
+ --s1:#2a78d6;--s2:#eb6834;--s3:#1baf7a}
5
+ @media (prefers-color-scheme:dark){.c{--surface:#1a1a19;--ink:#fff;--ink2:#c3c2b7;--grid:#2c2c2a;--axis:#383835;
6
+ --s1:#3987e5;--s2:#d95926;--s3:#199e70}}
7
+ text{font-family:ui-sans-serif,system-ui,-apple-system,"Segoe UI",sans-serif}
8
+ </style>
9
+ <rect width="760" height="452" rx="14" fill="var(--surface)"/>
10
+ <text x="32" y="50" font-size="20" font-weight="700" fill="var(--ink)">SwearBench</text>
11
+ <circle cx="37" cy="80" r="5" fill="var(--s1)"/><text x="47" y="84" font-size="12" fill="var(--ink2)">Claude</text>
12
+ <circle cx="116.32000000000001" cy="80" r="5" fill="var(--s2)"/><text x="126.32000000000001" y="84" font-size="12" fill="var(--ink2)">GPT</text>
13
+ <circle cx="175.48000000000002" cy="80" r="5" fill="var(--s3)"/><text x="185.48000000000002" y="84" font-size="12" fill="var(--ink2)">Other</text>
14
+ <text x="728" y="84" font-size="12" fill="var(--muted)" text-anchor="end">4585 reactions</text>
15
+ <line x1="32" x2="728" y1="120" y2="120" stroke="var(--grid)" stroke-width="1"/>
16
+ <text x="48" y="144.5" font-size="12" fill="var(--muted)" text-anchor="end">1</text>
17
+ <text x="62" y="144.5" font-size="13" font-weight="600" fill="var(--ink)">Opus 5.5</text>
18
+ <rect x="222" y="137.0" width="180" height="6" rx="3" fill="var(--grid)"/>
19
+ <rect x="222" y="137.0" width="180" height="6" rx="3" fill="var(--s1)"/>
20
+ <text x="442" y="144.5" font-size="13" font-weight="700" fill="var(--ink)" text-anchor="end" font-variant-numeric="tabular-nums">81</text>
21
+ <text x="466" y="144.5" font-size="12" fill="var(--ink2)">ships 83% · 31% angry</text>
22
+ <text x="728" y="144.5" font-size="12" fill="var(--muted)" text-anchor="end"></text>
23
+ <line x1="32" x2="728" y1="160" y2="160" stroke="var(--grid)" stroke-width="1"/>
24
+ <text x="48" y="184.5" font-size="12" fill="var(--muted)" text-anchor="end">2</text>
25
+ <text x="62" y="184.5" font-size="13" font-weight="600" fill="var(--ink)">Opus 5</text>
26
+ <rect x="222" y="177.0" width="180" height="6" rx="3" fill="var(--grid)"/>
27
+ <rect x="222" y="177.0" width="168" height="6" rx="3" fill="var(--s1)"/>
28
+ <text x="442" y="184.5" font-size="13" font-weight="700" fill="var(--ink)" text-anchor="end" font-variant-numeric="tabular-nums">75</text>
29
+ <text x="466" y="184.5" font-size="12" fill="var(--ink2)">ships 74% · 20% angry</text>
30
+ <text x="728" y="184.5" font-size="12" fill="var(--muted)" text-anchor="end">half-done</text>
31
+ <line x1="32" x2="728" y1="200" y2="200" stroke="var(--grid)" stroke-width="1"/>
32
+ <text x="48" y="224.5" font-size="12" fill="var(--muted)" text-anchor="end">3</text>
33
+ <text x="62" y="224.5" font-size="13" font-weight="600" fill="var(--ink)">Gemini 3.8 Flash</text>
34
+ <rect x="222" y="217.0" width="180" height="6" rx="3" fill="var(--grid)"/>
35
+ <rect x="222" y="217.0" width="151" height="6" rx="3" fill="var(--s3)"/>
36
+ <text x="442" y="224.5" font-size="13" font-weight="700" fill="var(--ink)" text-anchor="end" font-variant-numeric="tabular-nums">68</text>
37
+ <text x="466" y="224.5" font-size="12" fill="var(--ink2)">ships 67% · 27% angry</text>
38
+ <text x="728" y="224.5" font-size="12" fill="var(--muted)" text-anchor="end">half-done</text>
39
+ <line x1="32" x2="728" y1="240" y2="240" stroke="var(--grid)" stroke-width="1"/>
40
+ <text x="48" y="264.5" font-size="12" fill="var(--muted)" text-anchor="end">4</text>
41
+ <text x="62" y="264.5" font-size="13" font-weight="600" fill="var(--ink)">GPT-5.5</text>
42
+ <rect x="222" y="257.0" width="180" height="6" rx="3" fill="var(--grid)"/>
43
+ <rect x="222" y="257.0" width="143" height="6" rx="3" fill="var(--s2)"/>
44
+ <text x="442" y="264.5" font-size="13" font-weight="700" fill="var(--ink)" text-anchor="end" font-variant-numeric="tabular-nums">64</text>
45
+ <text x="466" y="264.5" font-size="12" fill="var(--ink2)">ships 62% · 22% angry</text>
46
+ <text x="728" y="264.5" font-size="12" fill="var(--muted)" text-anchor="end">breaks things</text>
47
+ <line x1="32" x2="728" y1="280" y2="280" stroke="var(--grid)" stroke-width="1"/>
48
+ <text x="48" y="304.5" font-size="12" fill="var(--muted)" text-anchor="end">5</text>
49
+ <text x="62" y="304.5" font-size="13" font-weight="600" fill="var(--ink)">Fable 5</text>
50
+ <rect x="222" y="297.0" width="180" height="6" rx="3" fill="var(--grid)"/>
51
+ <rect x="222" y="297.0" width="138" height="6" rx="3" fill="var(--s1)"/>
52
+ <text x="442" y="304.5" font-size="13" font-weight="700" fill="var(--ink)" text-anchor="end" font-variant-numeric="tabular-nums">62</text>
53
+ <text x="466" y="304.5" font-size="12" fill="var(--ink2)">ships 48% · 26% angry</text>
54
+ <text x="728" y="304.5" font-size="12" fill="var(--muted)" text-anchor="end">half-done</text>
55
+ <line x1="32" x2="728" y1="320" y2="320" stroke="var(--grid)" stroke-width="1"/>
56
+ <text x="48" y="344.5" font-size="12" fill="var(--muted)" text-anchor="end">6</text>
57
+ <text x="62" y="344.5" font-size="13" font-weight="600" fill="var(--ink)">GPT-5.6 Sol</text>
58
+ <rect x="222" y="337.0" width="180" height="6" rx="3" fill="var(--grid)"/>
59
+ <rect x="222" y="337.0" width="138" height="6" rx="3" fill="var(--s2)"/>
60
+ <text x="442" y="344.5" font-size="13" font-weight="700" fill="var(--ink)" text-anchor="end" font-variant-numeric="tabular-nums">62</text>
61
+ <text x="466" y="344.5" font-size="12" fill="var(--ink2)">ships 55% · 22% angry</text>
62
+ <text x="728" y="344.5" font-size="12" fill="var(--muted)" text-anchor="end">overreach</text>
63
+ <line x1="32" x2="728" y1="360" y2="360" stroke="var(--grid)" stroke-width="1"/>
64
+ <text x="48" y="384.5" font-size="12" fill="var(--muted)" text-anchor="end">7</text>
65
+ <text x="62" y="384.5" font-size="13" font-weight="600" fill="var(--ink)">Fable 5.1</text>
66
+ <rect x="222" y="377.0" width="180" height="6" rx="3" fill="var(--grid)"/>
67
+ <rect x="222" y="377.0" width="137" height="6" rx="3" fill="var(--s1)"/>
68
+ <text x="442" y="384.5" font-size="13" font-weight="700" fill="var(--ink)" text-anchor="end" font-variant-numeric="tabular-nums">61</text>
69
+ <text x="466" y="384.5" font-size="12" fill="var(--ink2)">ships 64% · 44% angry</text>
70
+ <text x="728" y="384.5" font-size="12" fill="var(--muted)" text-anchor="end">repeats</text>
71
+ <line x1="32" x2="728" y1="400" y2="400" stroke="var(--axis)" stroke-width="1"/>
72
+ <text x="32" y="420" font-size="11" fill="var(--muted)">github.com/AlenHay/swearbench</text>
73
+ </svg>
@@ -0,0 +1,78 @@
1
+ <svg xmlns="http://www.w3.org/2000/svg" class="c" width="760" height="540" viewBox="0 0 760 540" role="img" aria-labelledby="t d">
2
+ <style>
3
+ .c{--surface:#fcfcfb;--ink:#0b0b0b;--ink2:#52514e;--muted:#898781;--grid:#e1e0d9;--axis:#c3c2b7;
4
+ --s1:#2a78d6;--s2:#eb6834;--s3:#1baf7a}
5
+ @media (prefers-color-scheme:dark){.c{--surface:#1a1a19;--ink:#fff;--ink2:#c3c2b7;--grid:#2c2c2a;--axis:#383835;
6
+ --s1:#3987e5;--s2:#d95926;--s3:#199e70}}
7
+ text{font-family:ui-sans-serif,system-ui,-apple-system,"Segoe UI",sans-serif}
8
+ </style>
9
+ <title id="t">SwearBench: how much swearing gets you what result</title>
10
+ <desc id="d">claude-opus-5-5: rage 76 per 100 turns, ships 83%; claude-opus-5: rage 95 per 100 turns, ships 74%; gemini-3.8-flash-high: rage 120 per 100 turns, ships 67%; gpt-5.5: rage 131 per 100 turns, ships 62%; claude-fable-5: rage 92 per 100 turns, ships 48%; gpt-5.6-sol: rage 123 per 100 turns, ships 55%; claude-fable-5-1: rage 149 per 100 turns, ships 64%</desc>
11
+ <rect width="760" height="540" rx="14" fill="var(--surface)"/>
12
+ <text x="32" y="50" font-size="20" font-weight="700" fill="var(--ink)">Swearing vs. result</text>
13
+ <circle cx="37" cy="80" r="5" fill="var(--s1)"/><text x="47" y="84" font-size="12" fill="var(--ink2)">Claude</text>
14
+ <circle cx="116.32000000000001" cy="80" r="5" fill="var(--s2)"/><text x="126.32000000000001" y="84" font-size="12" fill="var(--ink2)">GPT</text>
15
+ <circle cx="175.48000000000002" cy="80" r="5" fill="var(--s3)"/><text x="185.48000000000002" y="84" font-size="12" fill="var(--ink2)">Other</text>
16
+ <line x1="80.0" x2="80.0" y1="120" y2="464" stroke="var(--grid)" stroke-width="1"/>
17
+ <text x="80.0" y="486" font-size="11" fill="var(--muted)" text-anchor="middle">60</text>
18
+ <line x1="209.6" x2="209.6" y1="120" y2="464" stroke="var(--grid)" stroke-width="1"/>
19
+ <text x="209.6" y="486" font-size="11" fill="var(--muted)" text-anchor="middle">80</text>
20
+ <line x1="339.2" x2="339.2" y1="120" y2="464" stroke="var(--grid)" stroke-width="1"/>
21
+ <text x="339.2" y="486" font-size="11" fill="var(--muted)" text-anchor="middle">100</text>
22
+ <line x1="468.8" x2="468.8" y1="120" y2="464" stroke="var(--grid)" stroke-width="1"/>
23
+ <text x="468.8" y="486" font-size="11" fill="var(--muted)" text-anchor="middle">120</text>
24
+ <line x1="598.4" x2="598.4" y1="120" y2="464" stroke="var(--grid)" stroke-width="1"/>
25
+ <text x="598.4" y="486" font-size="11" fill="var(--muted)" text-anchor="middle">140</text>
26
+ <line x1="728.0" x2="728.0" y1="120" y2="464" stroke="var(--grid)" stroke-width="1"/>
27
+ <text x="728.0" y="486" font-size="11" fill="var(--muted)" text-anchor="middle">160</text>
28
+ <line x1="80" x2="728" y1="464.0" y2="464.0" stroke="var(--grid)" stroke-width="1"/>
29
+ <text x="68" y="468.0" font-size="11" fill="var(--muted)" text-anchor="end">0%</text>
30
+ <line x1="80" x2="728" y1="395.2" y2="395.2" stroke="var(--grid)" stroke-width="1"/>
31
+ <text x="68" y="399.2" font-size="11" fill="var(--muted)" text-anchor="end">20%</text>
32
+ <line x1="80" x2="728" y1="326.4" y2="326.4" stroke="var(--grid)" stroke-width="1"/>
33
+ <text x="68" y="330.4" font-size="11" fill="var(--muted)" text-anchor="end">40%</text>
34
+ <line x1="80" x2="728" y1="257.6" y2="257.6" stroke="var(--grid)" stroke-width="1"/>
35
+ <text x="68" y="261.6" font-size="11" fill="var(--muted)" text-anchor="end">60%</text>
36
+ <line x1="80" x2="728" y1="188.8" y2="188.8" stroke="var(--grid)" stroke-width="1"/>
37
+ <text x="68" y="192.8" font-size="11" fill="var(--muted)" text-anchor="end">80%</text>
38
+ <line x1="80" x2="728" y1="120.0" y2="120.0" stroke="var(--grid)" stroke-width="1"/>
39
+ <text x="68" y="124.0" font-size="11" fill="var(--muted)" text-anchor="end">100%</text>
40
+ <line x1="80" x2="728" y1="464" y2="464" stroke="var(--axis)" stroke-width="1"/>
41
+ <text x="404.0" y="508" font-size="12" fill="var(--ink2)" text-anchor="middle">Swearing →</text>
42
+ <text transform="translate(36 292.0) rotate(-90)" font-size="12" fill="var(--ink2)" text-anchor="middle">Shipped →</text>
43
+ <g><title>claude-opus-5
44
+ rage 95/100 turns · 20% angry msgs
45
+ ships 74% of 121 decided sessions (90% CI 67–81%)
46
+ SwearBench 75.2</title><line x1="305.9" x2="305.9" y1="185.2" y2="233.7" stroke="var(--s1)" stroke-width="2" stroke-opacity="0.45" stroke-linecap="round"/><circle cx="305.9" cy="208.1" r="22.0" fill="transparent"/><circle cx="305.9" cy="208.1" r="14.0" fill="var(--s1)" stroke="var(--surface)" stroke-width="2"/></g>
47
+ <g><title>claude-opus-5-5
48
+ rage 76/100 turns · 31% angry msgs
49
+ ships 83% of 46 decided sessions (90% CI 73–91%)
50
+ SwearBench 80.7</title><line x1="186.1" x2="186.1" y1="152.0" y2="213.2" stroke="var(--s1)" stroke-width="2" stroke-opacity="0.45" stroke-linecap="round"/><circle cx="186.1" cy="179.8" r="17.8" fill="transparent"/><circle cx="186.1" cy="179.8" r="9.8" fill="var(--s1)" stroke="var(--surface)" stroke-width="2"/></g>
51
+ <g><title>claude-fable-5
52
+ rage 92/100 turns · 26% angry msgs
53
+ ships 48% of 25 decided sessions (90% CI 33–67%)
54
+ SwearBench 61.8</title><line x1="288.0" x2="288.0" y1="234.7" y2="349.3" stroke="var(--s1)" stroke-width="2" stroke-opacity="0.45" stroke-linecap="round"/><circle cx="288.0" cy="298.9" r="17.3" fill="transparent"/><circle cx="288.0" cy="298.9" r="9.3" fill="var(--s1)" stroke="var(--surface)" stroke-width="2"/></g>
55
+ <g><title>gpt-5.6-sol
56
+ rage 123/100 turns · 22% angry msgs
57
+ ships 55% of 20 decided sessions (90% CI 37–71%)
58
+ SwearBench 61.7</title><line x1="491.2" x2="491.2" y1="221.2" y2="337.3" stroke="var(--s2)" stroke-width="2" stroke-opacity="0.45" stroke-linecap="round"/><circle cx="491.2" cy="274.8" r="16.8" fill="transparent"/><circle cx="491.2" cy="274.8" r="8.8" fill="var(--s2)" stroke="var(--surface)" stroke-width="2"/></g>
59
+ <g><title>claude-fable-5-1
60
+ rage 149/100 turns · 44% angry msgs
61
+ ships 64% of 11 decided sessions (90% CI 38–87%)
62
+ SwearBench 61.3</title><line x1="657.2" x2="657.2" y1="165.9" y2="331.7" stroke="var(--s1)" stroke-width="2" stroke-opacity="0.45" stroke-linecap="round"/><circle cx="657.2" cy="245.1" r="15.7" fill="transparent"/><circle cx="657.2" cy="245.1" r="7.7" fill="var(--s1)" stroke="var(--surface)" stroke-width="2"/></g>
63
+ <g><title>gemini-3.8-flash-high
64
+ rage 120/100 turns · 27% angry msgs
65
+ ships 67% of 9 decided sessions (90% CI 40–90%)
66
+ SwearBench 67.8</title><line x1="469.3" x2="469.3" y1="154.4" y2="326.4" stroke="var(--s3)" stroke-width="2" stroke-opacity="0.45" stroke-linecap="round"/><circle cx="469.3" cy="234.7" r="15.3" fill="transparent"/><circle cx="469.3" cy="234.7" r="7.3" fill="var(--s3)" stroke="var(--surface)" stroke-width="2"/></g>
67
+ <g><title>gpt-5.5
68
+ rage 131/100 turns · 22% angry msgs
69
+ ships 62% of 8 decided sessions (90% CI 33–88%)
70
+ SwearBench 64.2</title><line x1="539.5" x2="539.5" y1="163.0" y2="349.3" stroke="var(--s2)" stroke-width="2" stroke-opacity="0.45" stroke-linecap="round"/><circle cx="539.5" cy="249.0" r="14.9" fill="transparent"/><circle cx="539.5" cy="249.0" r="6.9" fill="var(--s2)" stroke="var(--surface)" stroke-width="2"/></g>
71
+ <text x="327.9" y="212.1" font-size="12" fill="var(--ink)">Opus 5</text>
72
+ <text x="203.8" y="183.8" font-size="12" fill="var(--ink)">Opus 5.5</text>
73
+ <text x="305.4" y="302.9" font-size="12" fill="var(--ink)">Fable 5</text>
74
+ <text x="508.0" y="278.8" font-size="12" fill="var(--ink)">GPT-5.6 Sol</text>
75
+ <text x="581.0" y="249.1" font-size="12" fill="var(--ink)">Fable 5.1</text>
76
+ <text x="346.5" y="238.7" font-size="12" fill="var(--ink)">Gemini 3.8 Flash</text>
77
+ <text x="477.5" y="253.0" font-size="12" fill="var(--ink)">GPT-5.5</text>
78
+ </svg>
@@ -0,0 +1,15 @@
1
+ [project]
2
+ name = "swearbench"
3
+ version = "0.2.1"
4
+ description = "Rank AI coding models by how much they made you swear and how often their work shipped, from your own local chat logs."
5
+ readme = "README.md"
6
+ license = "MIT"
7
+ requires-python = ">=3.10"
8
+ dependencies = []
9
+
10
+ [project.scripts]
11
+ swearbench = "swearbench.cli:main"
12
+
13
+ [build-system]
14
+ requires = ["hatchling"]
15
+ build-backend = "hatchling.build"
@@ -0,0 +1 @@
1
+ __version__ = "0.2.1"
@@ -0,0 +1,5 @@
1
+ import sys
2
+
3
+ from .cli import main
4
+
5
+ sys.exit(main())
@@ -0,0 +1,66 @@
1
+ """Leaderboard card, styled to match chart.svg. Model names and numbers only; never quotes."""
2
+ from __future__ import annotations
3
+
4
+ from html import escape
5
+
6
+ from .chart import FAMILIES, PAD, STYLE, _text_w, family, pretty
7
+
8
+ MODE_LABEL = {
9
+ "fabrication": "lies", "overreach": "overreach", "giving_up": "gave up", "regression": "breaks things",
10
+ "ignored": "ignores me", "repeat": "repeats", "incomplete": "half-done", "insult": "insults",
11
+ "profanity": "swearing", "shouting": "CAPS", "sarcasm": "sarcasm", "slow": "slow",
12
+ }
13
+
14
+ W = 760
15
+ TOP = 120
16
+ ROW_H = 40
17
+ RANK_END = PAD + 16
18
+ NAME_X = RANK_END + 14
19
+ BAR_X = NAME_X + 160
20
+ BAR_W = 180
21
+ SCORE_END = BAR_X + BAR_W + 40
22
+ STATS_X = SCORE_END + 24
23
+
24
+
25
+ def svg(res, rows=8):
26
+ board = res["board"][:rows]
27
+ h = TOP + ROW_H * max(len(board), 1) + PAD + 20
28
+ hi = max([s["score"] for s in board] + [1])
29
+ o = [f'<svg xmlns="http://www.w3.org/2000/svg" class="c" width="{W}" height="{h}" viewBox="0 0 {W} {h}" '
30
+ 'role="img" aria-label="SwearBench leaderboard">',
31
+ f"<style>{STYLE}</style>",
32
+ f'<rect width="{W}" height="{h}" rx="14" fill="var(--surface)"/>',
33
+ f'<text x="{PAD}" y="{PAD + 18}" font-size="20" font-weight="700" fill="var(--ink)">SwearBench</text>']
34
+
35
+ lx, ly = PAD, PAD + 52
36
+ for i, (name, _) in enumerate(FAMILIES):
37
+ o.append(f'<circle cx="{lx + 5}" cy="{ly - 4}" r="5" fill="var(--s{i + 1})"/>'
38
+ f'<text x="{lx + 15}" y="{ly}" font-size="12" fill="var(--ink2)">{name}</text>')
39
+ lx += 15 + _text_w(name) + 24
40
+ o.append(f'<text x="{W - PAD}" y="{ly}" font-size="12" fill="var(--muted)" text-anchor="end">'
41
+ f'{res["n_reactions"]} reactions</text>')
42
+
43
+ for i, s in enumerate(board):
44
+ y = TOP + i * ROW_H
45
+ mid = y + ROW_H / 2
46
+ base = mid + 4.5
47
+ col = f"var(--s{family(s['model']) + 1})"
48
+ top_mode = max(s["modes"], key=s["modes"].get) if s["modes"] else None
49
+ o += [f'<line x1="{PAD}" x2="{W - PAD}" y1="{y}" y2="{y}" stroke="var(--grid)" stroke-width="1"/>',
50
+ f'<text x="{RANK_END}" y="{base}" font-size="12" fill="var(--muted)" text-anchor="end">{i + 1}</text>',
51
+ f'<text x="{NAME_X}" y="{base}" font-size="13" font-weight="600" fill="var(--ink)">'
52
+ f'{escape(pretty(s["model"]))}</text>',
53
+ f'<rect x="{BAR_X}" y="{mid - 3}" width="{BAR_W}" height="6" rx="3" fill="var(--grid)"/>',
54
+ f'<rect x="{BAR_X}" y="{mid - 3}" width="{max(6, BAR_W * max(s["score"], 0) / hi):.0f}" height="6" '
55
+ f'rx="3" fill="{col}"/>',
56
+ f'<text x="{SCORE_END}" y="{base}" font-size="13" font-weight="700" fill="var(--ink)" '
57
+ f'text-anchor="end" font-variant-numeric="tabular-nums">{s["score"]:.0f}</text>',
58
+ f'<text x="{STATS_X}" y="{base}" font-size="12" fill="var(--ink2)">'
59
+ f'ships {s["outcome"]:.0f}% · {s["angry_pct"]:.0f}% angry</text>',
60
+ f'<text x="{W - PAD}" y="{base}" font-size="12" fill="var(--muted)" text-anchor="end">'
61
+ f'{MODE_LABEL.get(top_mode, "")}</text>']
62
+ end = TOP + ROW_H * len(board)
63
+ o.append(f'<line x1="{PAD}" x2="{W - PAD}" y1="{end}" y2="{end}" stroke="var(--axis)" stroke-width="1"/>')
64
+ o.append(f'<text x="{PAD}" y="{h - PAD}" font-size="11" fill="var(--muted)">github.com/AlenHay/swearbench</text>')
65
+ o.append("</svg>")
66
+ return "\n".join(o)