diffprompt 0.1.0__py3-none-any.whl

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,249 @@
1
+ Metadata-Version: 2.4
2
+ Name: diffprompt
3
+ Version: 0.1.0
4
+ Summary: git diff for your prompt's behavior
5
+ Project-URL: Homepage, https://github.com/RudraDudhat2509/diffprompt
6
+ Project-URL: Repository, https://github.com/RudraDudhat2509/diffprompt
7
+ Author-email: Rudra Dudhat <contact.rdudhat@gmail.com>
8
+ License: MIT
9
+ Keywords: diff,evaluation,llm,prompt,testing
10
+ Classifier: Development Status :: 3 - Alpha
11
+ Classifier: Intended Audience :: Developers
12
+ Classifier: Programming Language :: Python :: 3.10
13
+ Classifier: Topic :: Software Development :: Testing
14
+ Requires-Python: >=3.10
15
+ Requires-Dist: click>=8.0.0
16
+ Requires-Dist: hdbscan>=0.8.29
17
+ Requires-Dist: httpx>=0.24.0
18
+ Requires-Dist: numpy>=1.24.0
19
+ Requires-Dist: pydantic>=2.0.0
20
+ Requires-Dist: rich>=13.0.0
21
+ Requires-Dist: scikit-learn>=1.3.0
22
+ Requires-Dist: sentence-transformers>=2.2.0
23
+ Requires-Dist: umap-learn>=0.5.0
24
+ Provides-Extra: dev
25
+ Requires-Dist: black>=23.0.0; extra == 'dev'
26
+ Requires-Dist: pytest-asyncio>=0.21.0; extra == 'dev'
27
+ Requires-Dist: pytest>=7.0.0; extra == 'dev'
28
+ Requires-Dist: ruff>=0.1.0; extra == 'dev'
29
+ Description-Content-Type: text/markdown
30
+
31
+ # diffprompt
32
+
33
+ > git diff for your prompt's behavior
34
+
35
+ ```bash
36
+ pip install diffprompt
37
+ diffprompt diff v1.txt v2.txt --auto-generate
38
+ ```
39
+
40
+ ---
41
+
42
+ You changed one sentence in your prompt. Now you're wondering: *did that actually help?*
43
+
44
+ LangSmith tells you what happened after you shipped. LangFuse tells you what's happening right now. Neither tells you what will happen **before** you change anything.
45
+
46
+ diffprompt does.
47
+
48
+ ---
49
+
50
+ ## What it does
51
+
52
+ You give it two prompts. It generates test cases, runs both prompts on all of them, measures the semantic difference between every output pair, and tells you exactly where v2 works better, where it regresses, and why.
53
+
54
+ ```
55
+ $ diffprompt diff v1.txt v2.txt --auto-generate --n 20
56
+
57
+ diffprompt v0.1.0 model: groq/llama-3.3-70b-versatile judge: local/qwen2.5:7b tests: 20
58
+ ━━ SUMMARY
59
+ 18.2/100 ███░░░░░░░░░░░░░░░░░ 4 improved 16 regressed 0 neutral
60
+ mix: 9 typical · 7 adversarial · 2 boundary · 2 format
61
+
62
+ ━━ BEHAVIORAL PROFILE
63
+ v2 performs well when...
64
+ ✓ user_intent:informational score 0.79 4 tests
65
+ v2 struggles when...
66
+ ✗ emotional_state:frustrated score 0.43 5 tests
67
+ ✗ request_type:specific_solutions score 0.51 11 tests
68
+
69
+ ━━ KEY EXAMPLES
70
+ MOST IMPORTANT emotional_state:frustrated divergence 0.90 conf 0.91
71
+
72
+ input Can you help me with a math problem I'm stuck on
73
+ v1 I'd be happy to help you with your math problem. What kind of problem are you working on?
74
+ v2 What's the problem?
75
+ why v2's brevity instruction strips the empathetic framing that makes
76
+ frustrated users feel heard before the question lands.
77
+
78
+ ━━ VERDICT
79
+ ✗ DO NOT SHIP
80
+ Keep v1 for emotional_state:frustrated, request_type:specific_solutions.
81
+ Primary failure mode: CONTEXT_LOSS (6 cases).
82
+ ```
83
+
84
+ ---
85
+
86
+ ## Install
87
+
88
+ ```bash
89
+ pip install diffprompt
90
+ ```
91
+
92
+ Requires Python 3.10+. Works fully offline with Ollama. No OpenAI key needed.
93
+
94
+ ---
95
+
96
+ ## Quickstart
97
+
98
+ ```bash
99
+ # Option A — use Groq (free at console.groq.com)
100
+ export GROQ_API_KEY=your_key_here
101
+
102
+ diffprompt diff v1.txt v2.txt --auto-generate
103
+
104
+ # Option B — run fully offline with Ollama
105
+ ollama pull qwen2.5:7b
106
+ diffprompt diff v1.txt v2.txt --auto-generate --local-only
107
+
108
+ # Option C — bring your own test inputs
109
+ diffprompt diff v1.txt v2.txt --test-file inputs.jsonl
110
+ ```
111
+
112
+ ---
113
+
114
+ ## How it works
115
+
116
+ ### 1. Ontology inference
117
+
118
+ diffprompt reads your prompt and infers what input dimensions matter for testing it — tone, complexity, intent, emotional state, whatever's relevant. No hardcoded dimensions. Every prompt gets its own.
119
+
120
+ ### 2. Test generation
121
+
122
+ Test cases are generated across four buckets: **typical** (real usage), **adversarial** (designed to find failures), **boundary** (edge cases), and **format** (unusual input styles). Each case is automatically tagged with its inferred dimensions.
123
+
124
+ ### 3. Semantic diff
125
+
126
+ Both prompts run on all test cases concurrently. Outputs are compared using local embeddings (`all-MiniLM-L6-v2`) to produce a similarity score per pair. High similarity means the change didn't matter. Low similarity means something changed.
127
+
128
+ ### 4. LLM judge
129
+
130
+ For every meaningfully different pair, a judge LLM evaluates direction: improvement, regression, or neutral. Confident verdicts stay local. Uncertain ones escalate to a larger model automatically.
131
+
132
+ ### 5. Behavioral slicing
133
+
134
+ Results are grouped by dimension. Instead of one aggregate score, you get a score per behavioral slice — not "47/100 overall" but "works for factual, breaks for emotional."
135
+
136
+ ### 6. Failure mode clustering
137
+
138
+ HDBSCAN clusters the judge's reasons automatically. Instead of 20 individual explanations, you get named failure modes: `CONTEXT_LOSS`, `TONE_SHIFT`, `REFUSAL_SHIFT`.
139
+
140
+ ---
141
+
142
+ ## Why not LangSmith / Langfuse?
143
+
144
+ Those tools monitor production. They tell you what happened.
145
+
146
+ diffprompt is a pre-flight check. It tells you what will happen before you touch production.
147
+
148
+ Different job. Different tool.
149
+
150
+ ---
151
+
152
+ ## Model cascade — zero cost by default
153
+
154
+ | Layer | Task | Default | Cost |
155
+ |-------|------|---------|------|
156
+ | Test generation | Generate inputs | qwen2.5:7b via Ollama | Free local |
157
+ | Embedding | Similarity | all-MiniLM-L6-v2 | Free local |
158
+ | Runner | Execute prompts | llama-3.3-70b via Groq | Free tier |
159
+ | Judge | Verdict + reason | qwen2.5:7b via Ollama | Free local |
160
+ | Escalation | Low confidence | llama-3.3-70b via Groq | Free tier |
161
+
162
+ Override any layer with `--model` and `--judge`.
163
+
164
+ ---
165
+
166
+ ## CLI reference
167
+
168
+ ```
169
+ diffprompt diff <v1> <v2> [options]
170
+
171
+ --auto-generate Generate test cases automatically
172
+ --n INT Number of test cases (default: 40)
173
+ --test-file PATH Use existing test inputs from .jsonl file
174
+ --model STRING Override runner model
175
+ --judge STRING Override judge model
176
+ --local-only Never call external APIs
177
+ --no-judge Skip judge, similarity scores only
178
+ --output FORMAT terminal (default) | json | html
179
+ --save PATH Save report to file
180
+ --top-n INT Show top N key examples (default: 3)
181
+ --verbose Show all diffs ranked by divergence
182
+ --quiet Score + verdict only
183
+ --ci CI mode: exit 1 on regression
184
+ --threshold INT CI failure threshold 0-100 (default: 75)
185
+ ```
186
+
187
+ ---
188
+
189
+ ## Output formats
190
+
191
+ **Terminal** — color-coded, fits in one screen.
192
+
193
+ **JSON** — full structured report for downstream processing.
194
+
195
+ **HTML** — self-contained file, open in browser.
196
+
197
+ ```bash
198
+ diffprompt diff v1.txt v2.txt --auto-generate --output html --save report.html
199
+ ```
200
+
201
+ ---
202
+
203
+ ## CI/CD integration
204
+
205
+ ```yaml
206
+ - name: Prompt regression check
207
+ run: |
208
+ diffprompt diff prompts/v1.txt prompts/v2.txt \
209
+ --auto-generate \
210
+ --ci \
211
+ --threshold 75
212
+ ```
213
+
214
+ Exits with code 1 if regression score drops below threshold. Merge blocked.
215
+
216
+ ---
217
+
218
+ ## Philosophy
219
+
220
+ Prompts have behavior, not just text.
221
+
222
+ When you change a prompt, you're not editing a document. You're changing how a system responds to thousands of possible inputs. Most of those inputs you've never seen. Some of them are edge cases you didn't think to test.
223
+
224
+ diffprompt makes the invisible visible. It tells you which inputs your change helped, which it hurt, and why — before any of it reaches a user.
225
+
226
+ ---
227
+
228
+ ## Stack
229
+
230
+ Python 3.10+ · sentence-transformers · HDBSCAN · UMAP · httpx · Click · Rich · Pydantic · Groq API · Ollama
231
+
232
+ ---
233
+
234
+ ## Contributing
235
+
236
+ Issues and PRs welcome.
237
+
238
+ ```bash
239
+ git clone https://github.com/RudraDudhat2509/diffprompt
240
+ cd diffprompt
241
+ pip install -e ".[dev]"
242
+ pytest tests/ -v
243
+ ```
244
+
245
+ ---
246
+
247
+ ## License
248
+
249
+ MIT
@@ -0,0 +1,20 @@
1
+ diffprompt/__init__.py,sha256=ihEhmM5g7DdW5NeiH8TfHQ3B1X2uZuiEUekTnWliih8,79
2
+ diffprompt/cli.py,sha256=fyIDvugb8PhnSP6aP0uIQ2eLOb84PJguEaiquE4DhFs,10748
3
+ diffprompt/core/__init__.py,sha256=47DEQpj8HBSa-_TImW-5JCeuQeRkm5NMpJWZG3hSuFU,0
4
+ diffprompt/core/clusterer.py,sha256=M71wv6cHKd7Q56NCAi3o9SRxQhPbCCc5AdP1cBo4cec,4722
5
+ diffprompt/core/embedder.py,sha256=ozFWuNb_JyZZm44op_lwqqdEt_C5vQnXzUI-NB_9Xhc,2223
6
+ diffprompt/core/generator.py,sha256=_PbK7s-lcnXyqTcksakO7786El-0ZOQubpDMzJshME4,4067
7
+ diffprompt/core/judge.py,sha256=1QIJXSn_L8TT7RPxE8I7bxaEiLc9Z5DmzcOXHSn-D3M,3530
8
+ diffprompt/core/ontology.py,sha256=0h2RvcWOCRCfNLVbDbENcL0c6DEWG2zyL0g-AF2Sx7E,6418
9
+ diffprompt/core/runner.py,sha256=d4tfXYKIZn1xZxRWxq7b90wqWkkcgQG9w3tPn4QiRkg,1845
10
+ diffprompt/core/scorer.py,sha256=ihw1puP-870A5Rxi8LIx0MX4JecBClc6ihRDCOT6WyA,3267
11
+ diffprompt/core/slicer.py,sha256=BrIWI96t4e8srSkTqo__sbhDJvLl6TSss5vA-K7-Skc,4489
12
+ diffprompt/models/__init__.py,sha256=w_wHHpbYqKqeWooDgqZOrXDywZCLN-eKLgJBYOYuCYE,2250
13
+ diffprompt/models/cascade.py,sha256=Cs3m6v-4_5LD9T5fl6V7SArqf6SEt-8H_MMBZA00_60,5182
14
+ diffprompt/output/__init__.py,sha256=47DEQpj8HBSa-_TImW-5JCeuQeRkm5NMpJWZG3hSuFU,0
15
+ diffprompt/output/exporter.py,sha256=G7V1TE8qz3d71A9el7mWDEAXolRt2WP7Ozzniu7RHD0,10198
16
+ diffprompt/output/terminal.py,sha256=-QM8e-VHweoSue84GrKFvRpFvqRhvBCetr4z2thzEwI,4843
17
+ diffprompt-0.1.0.dist-info/METADATA,sha256=QdytKis00y5EW1RLNUWmkZWi2g1JDfyvoy_2LIK2Ni0,7770
18
+ diffprompt-0.1.0.dist-info/WHEEL,sha256=QccIxa26bgl1E6uMy58deGWi-0aeIkkangHcxk2kWfw,87
19
+ diffprompt-0.1.0.dist-info/entry_points.txt,sha256=0R2O1Y5wZIT3PzVbI_8XD9dqrkpaToEcUef00adSn-s,50
20
+ diffprompt-0.1.0.dist-info/RECORD,,
@@ -0,0 +1,4 @@
1
+ Wheel-Version: 1.0
2
+ Generator: hatchling 1.29.0
3
+ Root-Is-Purelib: true
4
+ Tag: py3-none-any
@@ -0,0 +1,2 @@
1
+ [console_scripts]
2
+ diffprompt = diffprompt.cli:app