diffprompt 0.1.0__py3-none-any.whl
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- diffprompt/__init__.py +2 -0
- diffprompt/cli.py +246 -0
- diffprompt/core/__init__.py +0 -0
- diffprompt/core/clusterer.py +130 -0
- diffprompt/core/embedder.py +63 -0
- diffprompt/core/generator.py +123 -0
- diffprompt/core/judge.py +94 -0
- diffprompt/core/ontology.py +160 -0
- diffprompt/core/runner.py +65 -0
- diffprompt/core/scorer.py +102 -0
- diffprompt/core/slicer.py +152 -0
- diffprompt/models/__init__.py +105 -0
- diffprompt/models/cascade.py +144 -0
- diffprompt/output/__init__.py +0 -0
- diffprompt/output/exporter.py +227 -0
- diffprompt/output/terminal.py +154 -0
- diffprompt-0.1.0.dist-info/METADATA +249 -0
- diffprompt-0.1.0.dist-info/RECORD +20 -0
- diffprompt-0.1.0.dist-info/WHEEL +4 -0
- diffprompt-0.1.0.dist-info/entry_points.txt +2 -0
|
@@ -0,0 +1,249 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: diffprompt
|
|
3
|
+
Version: 0.1.0
|
|
4
|
+
Summary: git diff for your prompt's behavior
|
|
5
|
+
Project-URL: Homepage, https://github.com/RudraDudhat2509/diffprompt
|
|
6
|
+
Project-URL: Repository, https://github.com/RudraDudhat2509/diffprompt
|
|
7
|
+
Author-email: Rudra Dudhat <contact.rdudhat@gmail.com>
|
|
8
|
+
License: MIT
|
|
9
|
+
Keywords: diff,evaluation,llm,prompt,testing
|
|
10
|
+
Classifier: Development Status :: 3 - Alpha
|
|
11
|
+
Classifier: Intended Audience :: Developers
|
|
12
|
+
Classifier: Programming Language :: Python :: 3.10
|
|
13
|
+
Classifier: Topic :: Software Development :: Testing
|
|
14
|
+
Requires-Python: >=3.10
|
|
15
|
+
Requires-Dist: click>=8.0.0
|
|
16
|
+
Requires-Dist: hdbscan>=0.8.29
|
|
17
|
+
Requires-Dist: httpx>=0.24.0
|
|
18
|
+
Requires-Dist: numpy>=1.24.0
|
|
19
|
+
Requires-Dist: pydantic>=2.0.0
|
|
20
|
+
Requires-Dist: rich>=13.0.0
|
|
21
|
+
Requires-Dist: scikit-learn>=1.3.0
|
|
22
|
+
Requires-Dist: sentence-transformers>=2.2.0
|
|
23
|
+
Requires-Dist: umap-learn>=0.5.0
|
|
24
|
+
Provides-Extra: dev
|
|
25
|
+
Requires-Dist: black>=23.0.0; extra == 'dev'
|
|
26
|
+
Requires-Dist: pytest-asyncio>=0.21.0; extra == 'dev'
|
|
27
|
+
Requires-Dist: pytest>=7.0.0; extra == 'dev'
|
|
28
|
+
Requires-Dist: ruff>=0.1.0; extra == 'dev'
|
|
29
|
+
Description-Content-Type: text/markdown
|
|
30
|
+
|
|
31
|
+
# diffprompt
|
|
32
|
+
|
|
33
|
+
> git diff for your prompt's behavior
|
|
34
|
+
|
|
35
|
+
```bash
|
|
36
|
+
pip install diffprompt
|
|
37
|
+
diffprompt diff v1.txt v2.txt --auto-generate
|
|
38
|
+
```
|
|
39
|
+
|
|
40
|
+
---
|
|
41
|
+
|
|
42
|
+
You changed one sentence in your prompt. Now you're wondering: *did that actually help?*
|
|
43
|
+
|
|
44
|
+
LangSmith tells you what happened after you shipped. LangFuse tells you what's happening right now. Neither tells you what will happen **before** you change anything.
|
|
45
|
+
|
|
46
|
+
diffprompt does.
|
|
47
|
+
|
|
48
|
+
---
|
|
49
|
+
|
|
50
|
+
## What it does
|
|
51
|
+
|
|
52
|
+
You give it two prompts. It generates test cases, runs both prompts on all of them, measures the semantic difference between every output pair, and tells you exactly where v2 works better, where it regresses, and why.
|
|
53
|
+
|
|
54
|
+
```
|
|
55
|
+
$ diffprompt diff v1.txt v2.txt --auto-generate --n 20
|
|
56
|
+
|
|
57
|
+
diffprompt v0.1.0 model: groq/llama-3.3-70b-versatile judge: local/qwen2.5:7b tests: 20
|
|
58
|
+
━━ SUMMARY
|
|
59
|
+
18.2/100 ███░░░░░░░░░░░░░░░░░ 4 improved 16 regressed 0 neutral
|
|
60
|
+
mix: 9 typical · 7 adversarial · 2 boundary · 2 format
|
|
61
|
+
|
|
62
|
+
━━ BEHAVIORAL PROFILE
|
|
63
|
+
v2 performs well when...
|
|
64
|
+
✓ user_intent:informational score 0.79 4 tests
|
|
65
|
+
v2 struggles when...
|
|
66
|
+
✗ emotional_state:frustrated score 0.43 5 tests
|
|
67
|
+
✗ request_type:specific_solutions score 0.51 11 tests
|
|
68
|
+
|
|
69
|
+
━━ KEY EXAMPLES
|
|
70
|
+
MOST IMPORTANT emotional_state:frustrated divergence 0.90 conf 0.91
|
|
71
|
+
|
|
72
|
+
input Can you help me with a math problem I'm stuck on
|
|
73
|
+
v1 I'd be happy to help you with your math problem. What kind of problem are you working on?
|
|
74
|
+
v2 What's the problem?
|
|
75
|
+
why v2's brevity instruction strips the empathetic framing that makes
|
|
76
|
+
frustrated users feel heard before the question lands.
|
|
77
|
+
|
|
78
|
+
━━ VERDICT
|
|
79
|
+
✗ DO NOT SHIP
|
|
80
|
+
Keep v1 for emotional_state:frustrated, request_type:specific_solutions.
|
|
81
|
+
Primary failure mode: CONTEXT_LOSS (6 cases).
|
|
82
|
+
```
|
|
83
|
+
|
|
84
|
+
---
|
|
85
|
+
|
|
86
|
+
## Install
|
|
87
|
+
|
|
88
|
+
```bash
|
|
89
|
+
pip install diffprompt
|
|
90
|
+
```
|
|
91
|
+
|
|
92
|
+
Requires Python 3.10+. Works fully offline with Ollama. No OpenAI key needed.
|
|
93
|
+
|
|
94
|
+
---
|
|
95
|
+
|
|
96
|
+
## Quickstart
|
|
97
|
+
|
|
98
|
+
```bash
|
|
99
|
+
# Option A — use Groq (free at console.groq.com)
|
|
100
|
+
export GROQ_API_KEY=your_key_here
|
|
101
|
+
|
|
102
|
+
diffprompt diff v1.txt v2.txt --auto-generate
|
|
103
|
+
|
|
104
|
+
# Option B — run fully offline with Ollama
|
|
105
|
+
ollama pull qwen2.5:7b
|
|
106
|
+
diffprompt diff v1.txt v2.txt --auto-generate --local-only
|
|
107
|
+
|
|
108
|
+
# Option C — bring your own test inputs
|
|
109
|
+
diffprompt diff v1.txt v2.txt --test-file inputs.jsonl
|
|
110
|
+
```
|
|
111
|
+
|
|
112
|
+
---
|
|
113
|
+
|
|
114
|
+
## How it works
|
|
115
|
+
|
|
116
|
+
### 1. Ontology inference
|
|
117
|
+
|
|
118
|
+
diffprompt reads your prompt and infers what input dimensions matter for testing it — tone, complexity, intent, emotional state, whatever's relevant. No hardcoded dimensions. Every prompt gets its own.
|
|
119
|
+
|
|
120
|
+
### 2. Test generation
|
|
121
|
+
|
|
122
|
+
Test cases are generated across four buckets: **typical** (real usage), **adversarial** (designed to find failures), **boundary** (edge cases), and **format** (unusual input styles). Each case is automatically tagged with its inferred dimensions.
|
|
123
|
+
|
|
124
|
+
### 3. Semantic diff
|
|
125
|
+
|
|
126
|
+
Both prompts run on all test cases concurrently. Outputs are compared using local embeddings (`all-MiniLM-L6-v2`) to produce a similarity score per pair. High similarity means the change didn't matter. Low similarity means something changed.
|
|
127
|
+
|
|
128
|
+
### 4. LLM judge
|
|
129
|
+
|
|
130
|
+
For every meaningfully different pair, a judge LLM evaluates direction: improvement, regression, or neutral. Confident verdicts stay local. Uncertain ones escalate to a larger model automatically.
|
|
131
|
+
|
|
132
|
+
### 5. Behavioral slicing
|
|
133
|
+
|
|
134
|
+
Results are grouped by dimension. Instead of one aggregate score, you get a score per behavioral slice — not "47/100 overall" but "works for factual, breaks for emotional."
|
|
135
|
+
|
|
136
|
+
### 6. Failure mode clustering
|
|
137
|
+
|
|
138
|
+
HDBSCAN clusters the judge's reasons automatically. Instead of 20 individual explanations, you get named failure modes: `CONTEXT_LOSS`, `TONE_SHIFT`, `REFUSAL_SHIFT`.
|
|
139
|
+
|
|
140
|
+
---
|
|
141
|
+
|
|
142
|
+
## Why not LangSmith / Langfuse?
|
|
143
|
+
|
|
144
|
+
Those tools monitor production. They tell you what happened.
|
|
145
|
+
|
|
146
|
+
diffprompt is a pre-flight check. It tells you what will happen before you touch production.
|
|
147
|
+
|
|
148
|
+
Different job. Different tool.
|
|
149
|
+
|
|
150
|
+
---
|
|
151
|
+
|
|
152
|
+
## Model cascade — zero cost by default
|
|
153
|
+
|
|
154
|
+
| Layer | Task | Default | Cost |
|
|
155
|
+
|-------|------|---------|------|
|
|
156
|
+
| Test generation | Generate inputs | qwen2.5:7b via Ollama | Free local |
|
|
157
|
+
| Embedding | Similarity | all-MiniLM-L6-v2 | Free local |
|
|
158
|
+
| Runner | Execute prompts | llama-3.3-70b via Groq | Free tier |
|
|
159
|
+
| Judge | Verdict + reason | qwen2.5:7b via Ollama | Free local |
|
|
160
|
+
| Escalation | Low confidence | llama-3.3-70b via Groq | Free tier |
|
|
161
|
+
|
|
162
|
+
Override any layer with `--model` and `--judge`.
|
|
163
|
+
|
|
164
|
+
---
|
|
165
|
+
|
|
166
|
+
## CLI reference
|
|
167
|
+
|
|
168
|
+
```
|
|
169
|
+
diffprompt diff <v1> <v2> [options]
|
|
170
|
+
|
|
171
|
+
--auto-generate Generate test cases automatically
|
|
172
|
+
--n INT Number of test cases (default: 40)
|
|
173
|
+
--test-file PATH Use existing test inputs from .jsonl file
|
|
174
|
+
--model STRING Override runner model
|
|
175
|
+
--judge STRING Override judge model
|
|
176
|
+
--local-only Never call external APIs
|
|
177
|
+
--no-judge Skip judge, similarity scores only
|
|
178
|
+
--output FORMAT terminal (default) | json | html
|
|
179
|
+
--save PATH Save report to file
|
|
180
|
+
--top-n INT Show top N key examples (default: 3)
|
|
181
|
+
--verbose Show all diffs ranked by divergence
|
|
182
|
+
--quiet Score + verdict only
|
|
183
|
+
--ci CI mode: exit 1 on regression
|
|
184
|
+
--threshold INT CI failure threshold 0-100 (default: 75)
|
|
185
|
+
```
|
|
186
|
+
|
|
187
|
+
---
|
|
188
|
+
|
|
189
|
+
## Output formats
|
|
190
|
+
|
|
191
|
+
**Terminal** — color-coded, fits in one screen.
|
|
192
|
+
|
|
193
|
+
**JSON** — full structured report for downstream processing.
|
|
194
|
+
|
|
195
|
+
**HTML** — self-contained file, open in browser.
|
|
196
|
+
|
|
197
|
+
```bash
|
|
198
|
+
diffprompt diff v1.txt v2.txt --auto-generate --output html --save report.html
|
|
199
|
+
```
|
|
200
|
+
|
|
201
|
+
---
|
|
202
|
+
|
|
203
|
+
## CI/CD integration
|
|
204
|
+
|
|
205
|
+
```yaml
|
|
206
|
+
- name: Prompt regression check
|
|
207
|
+
run: |
|
|
208
|
+
diffprompt diff prompts/v1.txt prompts/v2.txt \
|
|
209
|
+
--auto-generate \
|
|
210
|
+
--ci \
|
|
211
|
+
--threshold 75
|
|
212
|
+
```
|
|
213
|
+
|
|
214
|
+
Exits with code 1 if regression score drops below threshold. Merge blocked.
|
|
215
|
+
|
|
216
|
+
---
|
|
217
|
+
|
|
218
|
+
## Philosophy
|
|
219
|
+
|
|
220
|
+
Prompts have behavior, not just text.
|
|
221
|
+
|
|
222
|
+
When you change a prompt, you're not editing a document. You're changing how a system responds to thousands of possible inputs. Most of those inputs you've never seen. Some of them are edge cases you didn't think to test.
|
|
223
|
+
|
|
224
|
+
diffprompt makes the invisible visible. It tells you which inputs your change helped, which it hurt, and why — before any of it reaches a user.
|
|
225
|
+
|
|
226
|
+
---
|
|
227
|
+
|
|
228
|
+
## Stack
|
|
229
|
+
|
|
230
|
+
Python 3.10+ · sentence-transformers · HDBSCAN · UMAP · httpx · Click · Rich · Pydantic · Groq API · Ollama
|
|
231
|
+
|
|
232
|
+
---
|
|
233
|
+
|
|
234
|
+
## Contributing
|
|
235
|
+
|
|
236
|
+
Issues and PRs welcome.
|
|
237
|
+
|
|
238
|
+
```bash
|
|
239
|
+
git clone https://github.com/RudraDudhat2509/diffprompt
|
|
240
|
+
cd diffprompt
|
|
241
|
+
pip install -e ".[dev]"
|
|
242
|
+
pytest tests/ -v
|
|
243
|
+
```
|
|
244
|
+
|
|
245
|
+
---
|
|
246
|
+
|
|
247
|
+
## License
|
|
248
|
+
|
|
249
|
+
MIT
|
|
@@ -0,0 +1,20 @@
|
|
|
1
|
+
diffprompt/__init__.py,sha256=ihEhmM5g7DdW5NeiH8TfHQ3B1X2uZuiEUekTnWliih8,79
|
|
2
|
+
diffprompt/cli.py,sha256=fyIDvugb8PhnSP6aP0uIQ2eLOb84PJguEaiquE4DhFs,10748
|
|
3
|
+
diffprompt/core/__init__.py,sha256=47DEQpj8HBSa-_TImW-5JCeuQeRkm5NMpJWZG3hSuFU,0
|
|
4
|
+
diffprompt/core/clusterer.py,sha256=M71wv6cHKd7Q56NCAi3o9SRxQhPbCCc5AdP1cBo4cec,4722
|
|
5
|
+
diffprompt/core/embedder.py,sha256=ozFWuNb_JyZZm44op_lwqqdEt_C5vQnXzUI-NB_9Xhc,2223
|
|
6
|
+
diffprompt/core/generator.py,sha256=_PbK7s-lcnXyqTcksakO7786El-0ZOQubpDMzJshME4,4067
|
|
7
|
+
diffprompt/core/judge.py,sha256=1QIJXSn_L8TT7RPxE8I7bxaEiLc9Z5DmzcOXHSn-D3M,3530
|
|
8
|
+
diffprompt/core/ontology.py,sha256=0h2RvcWOCRCfNLVbDbENcL0c6DEWG2zyL0g-AF2Sx7E,6418
|
|
9
|
+
diffprompt/core/runner.py,sha256=d4tfXYKIZn1xZxRWxq7b90wqWkkcgQG9w3tPn4QiRkg,1845
|
|
10
|
+
diffprompt/core/scorer.py,sha256=ihw1puP-870A5Rxi8LIx0MX4JecBClc6ihRDCOT6WyA,3267
|
|
11
|
+
diffprompt/core/slicer.py,sha256=BrIWI96t4e8srSkTqo__sbhDJvLl6TSss5vA-K7-Skc,4489
|
|
12
|
+
diffprompt/models/__init__.py,sha256=w_wHHpbYqKqeWooDgqZOrXDywZCLN-eKLgJBYOYuCYE,2250
|
|
13
|
+
diffprompt/models/cascade.py,sha256=Cs3m6v-4_5LD9T5fl6V7SArqf6SEt-8H_MMBZA00_60,5182
|
|
14
|
+
diffprompt/output/__init__.py,sha256=47DEQpj8HBSa-_TImW-5JCeuQeRkm5NMpJWZG3hSuFU,0
|
|
15
|
+
diffprompt/output/exporter.py,sha256=G7V1TE8qz3d71A9el7mWDEAXolRt2WP7Ozzniu7RHD0,10198
|
|
16
|
+
diffprompt/output/terminal.py,sha256=-QM8e-VHweoSue84GrKFvRpFvqRhvBCetr4z2thzEwI,4843
|
|
17
|
+
diffprompt-0.1.0.dist-info/METADATA,sha256=QdytKis00y5EW1RLNUWmkZWi2g1JDfyvoy_2LIK2Ni0,7770
|
|
18
|
+
diffprompt-0.1.0.dist-info/WHEEL,sha256=QccIxa26bgl1E6uMy58deGWi-0aeIkkangHcxk2kWfw,87
|
|
19
|
+
diffprompt-0.1.0.dist-info/entry_points.txt,sha256=0R2O1Y5wZIT3PzVbI_8XD9dqrkpaToEcUef00adSn-s,50
|
|
20
|
+
diffprompt-0.1.0.dist-info/RECORD,,
|