citesure 0.2.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- citesure-0.2.0/PKG-INFO +352 -0
- citesure-0.2.0/README.md +321 -0
- citesure-0.2.0/pyproject.toml +75 -0
- citesure-0.2.0/setup.cfg +4 -0
- citesure-0.2.0/src/citesure/__init__.py +70 -0
- citesure-0.2.0/src/citesure/citations.py +321 -0
- citesure-0.2.0/src/citesure/cli.py +208 -0
- citesure-0.2.0/src/citesure/fetcher.py +527 -0
- citesure-0.2.0/src/citesure/mcp_server.py +213 -0
- citesure-0.2.0/src/citesure/models.py +147 -0
- citesure-0.2.0/src/citesure/nli.py +624 -0
- citesure-0.2.0/src/citesure/overlap.py +911 -0
- citesure-0.2.0/src/citesure/reachability.py +125 -0
- citesure-0.2.0/src/citesure/report.py +128 -0
- citesure-0.2.0/src/citesure.egg-info/PKG-INFO +352 -0
- citesure-0.2.0/src/citesure.egg-info/SOURCES.txt +42 -0
- citesure-0.2.0/src/citesure.egg-info/dependency_links.txt +1 -0
- citesure-0.2.0/src/citesure.egg-info/entry_points.txt +3 -0
- citesure-0.2.0/src/citesure.egg-info/requires.txt +13 -0
- citesure-0.2.0/src/citesure.egg-info/top_level.txt +1 -0
- citesure-0.2.0/tests/test_citations.py +293 -0
- citesure-0.2.0/tests/test_cli.py +219 -0
- citesure-0.2.0/tests/test_cli_defaults.py +99 -0
- citesure-0.2.0/tests/test_compound_claim_floor.py +90 -0
- citesure-0.2.0/tests/test_concurrency.py +170 -0
- citesure-0.2.0/tests/test_context_negation.py +85 -0
- citesure-0.2.0/tests/test_contradiction.py +174 -0
- citesure-0.2.0/tests/test_corpus.py +196 -0
- citesure-0.2.0/tests/test_edge_cases.py +299 -0
- citesure-0.2.0/tests/test_eval_set_integrity.py +90 -0
- citesure-0.2.0/tests/test_fetcher.py +377 -0
- citesure-0.2.0/tests/test_fuzz_extraction.py +236 -0
- citesure-0.2.0/tests/test_fuzz_report.py +234 -0
- citesure-0.2.0/tests/test_fuzz_urls.py +157 -0
- citesure-0.2.0/tests/test_inference_cost.py +43 -0
- citesure-0.2.0/tests/test_integration.py +130 -0
- citesure-0.2.0/tests/test_live_smoke.py +50 -0
- citesure-0.2.0/tests/test_mcp_integration.py +207 -0
- citesure-0.2.0/tests/test_mcp_stress.py +228 -0
- citesure-0.2.0/tests/test_nli.py +757 -0
- citesure-0.2.0/tests/test_overlap.py +729 -0
- citesure-0.2.0/tests/test_packaging.py +43 -0
- citesure-0.2.0/tests/test_reachability.py +148 -0
- citesure-0.2.0/tests/test_topk_max.py +59 -0
citesure-0.2.0/PKG-INFO
ADDED
|
@@ -0,0 +1,352 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: citesure
|
|
3
|
+
Version: 0.2.0
|
|
4
|
+
Summary: MCP citation verifier: check whether an LLM's claims are actually supported by the sources it cites.
|
|
5
|
+
Author: DawnofGenX
|
|
6
|
+
License: MIT
|
|
7
|
+
Project-URL: Homepage, https://github.com/DawnofGenX/citesure
|
|
8
|
+
Project-URL: Repository, https://github.com/DawnofGenX/citesure
|
|
9
|
+
Keywords: citations,verification,mcp,hallucination,nli,fact-checking
|
|
10
|
+
Classifier: Development Status :: 3 - Alpha
|
|
11
|
+
Classifier: Intended Audience :: Developers
|
|
12
|
+
Classifier: Programming Language :: Python :: 3
|
|
13
|
+
Classifier: Programming Language :: Python :: 3.10
|
|
14
|
+
Classifier: Programming Language :: Python :: 3.11
|
|
15
|
+
Classifier: Programming Language :: Python :: 3.12
|
|
16
|
+
Classifier: Topic :: Software Development :: Libraries
|
|
17
|
+
Classifier: Topic :: Text Processing :: Linguistic
|
|
18
|
+
Requires-Python: >=3.10
|
|
19
|
+
Description-Content-Type: text/markdown
|
|
20
|
+
Requires-Dist: httpx>=0.27
|
|
21
|
+
Requires-Dist: trafilatura>=1.9
|
|
22
|
+
Requires-Dist: lxml_html_clean>=0.1
|
|
23
|
+
Requires-Dist: mcp>=1.0
|
|
24
|
+
Provides-Extra: nli
|
|
25
|
+
Requires-Dist: transformers>=4.40; extra == "nli"
|
|
26
|
+
Requires-Dist: torch>=2.0; extra == "nli"
|
|
27
|
+
Requires-Dist: sentence-transformers>=2.6; extra == "nli"
|
|
28
|
+
Provides-Extra: dev
|
|
29
|
+
Requires-Dist: pytest>=8.0; extra == "dev"
|
|
30
|
+
Requires-Dist: hypothesis>=6.0; extra == "dev"
|
|
31
|
+
|
|
32
|
+
<p align="center"><img src="https://raw.githubusercontent.com/DawnofGenX/citesure/main/docs/og-banner.png" alt="citesure" width="100%"/></p>
|
|
33
|
+
|
|
34
|
+
# citesure
|
|
35
|
+
|
|
36
|
+
An MCP server + CLI + library that verifies whether an LLM's claims are actually
|
|
37
|
+
supported by the sources it cites — resolves each citation, fetches the real
|
|
38
|
+
content, flags dead/retracted/paywalled links, and abstains when it can't
|
|
39
|
+
confirm.
|
|
40
|
+
|
|
41
|
+
citesure verifies *claims* (not just whether a link is alive). It needs no API
|
|
42
|
+
keys and no LLM: the fast path (tiers 1+2) is pure `httpx` + `trafilatura`, and
|
|
43
|
+
passing `--no-nli` keeps it fully offline. The NLI tier (tier 3) is on by
|
|
44
|
+
default and lazy-downloads a local cross-encoder on first use.
|
|
45
|
+
|
|
46
|
+
## Accuracy
|
|
47
|
+
|
|
48
|
+
Measured on an independent 108 case set (40 unique URLs, 13 domains), built as a three-way
|
|
49
|
+
contrast: each source fact appears three times — verbatim-supported, negation-flipped, and
|
|
50
|
+
entity-swapped — all sharing one source sentence. The A-vs-B gap isolates polarity handling and
|
|
51
|
+
A-vs-C isolates subject binding, so a drop points at a specific bug class rather than "hard cases".
|
|
52
|
+
|
|
53
|
+
| Configuration | Agreement |
|
|
54
|
+
|---|---|
|
|
55
|
+
| NLI tier on (default claim path) | **83/108 = 76.9%** |
|
|
56
|
+
| NLI tier off (deterministic tiers only) | **41/108 = 38.0%** |
|
|
57
|
+
|
|
58
|
+
The 38.9-point gap is the cross-encoder tier's entire contribution, measured rather than asserted.
|
|
59
|
+
|
|
60
|
+
**Safety metrics** — the failure modes that matter for a verifier:
|
|
61
|
+
|
|
62
|
+
| Metric | Result |
|
|
63
|
+
|---|---|
|
|
64
|
+
| Negation-flip false-supported | 4/32 = 12.5% |
|
|
65
|
+
| Entity-swap false-supported | 7/32 = 21.9% |
|
|
66
|
+
|
|
67
|
+
Both rise to ~90-100% with the NLI tier off, which is the honest cost of the fast path.
|
|
68
|
+
|
|
69
|
+
An exhaustive threshold sweep (support 0.30-0.90 x ambiguous 0.05-support) moved agreement from
|
|
70
|
+
66.7% to 69.4% — a gain of exactly one item. Threshold tuning is a dead end here; see
|
|
71
|
+
`evals/THRESHOLD_CALIBRATION.md`.
|
|
72
|
+
|
|
73
|
+
## Install
|
|
74
|
+
|
|
75
|
+
Default install stays light — no torch, no model download:
|
|
76
|
+
|
|
77
|
+
```bash
|
|
78
|
+
pip install citesure # tiers 1+2: httpx + trafilatura only
|
|
79
|
+
pip install "citesure[nli]" # adds the opt-in NLI cross-encoder tier
|
|
80
|
+
```
|
|
81
|
+
|
|
82
|
+
Or from a checkout:
|
|
83
|
+
|
|
84
|
+
```bash
|
|
85
|
+
git clone https://github.com/DawnofGenX/citesure.git && cd citesure
|
|
86
|
+
pip install . # default: offline, key-free, no torch
|
|
87
|
+
pip install ".[nli]" # adds the opt-in NLI tier
|
|
88
|
+
pip install -e ".[dev]" # editable + test deps
|
|
89
|
+
```
|
|
90
|
+
|
|
91
|
+
`requires-python >= 3.10`. The default fast path (tiers 1+2) needs only
|
|
92
|
+
`httpx` + `trafilatura`. The opt-in NLI tier pulls in the local cross-encoder
|
|
93
|
+
stack (`torch`, `transformers`, `sentence-transformers`) and lazy-downloads a
|
|
94
|
+
~425 MB model on first `--nli` use — see [NLI tier](#nli-tier-opt-in).
|
|
95
|
+
|
|
96
|
+
## Quickstart
|
|
97
|
+
|
|
98
|
+
Create two small files in any directory — a claim document and the page it
|
|
99
|
+
cites:
|
|
100
|
+
|
|
101
|
+
`source.html`
|
|
102
|
+
```html
|
|
103
|
+
<!DOCTYPE html>
|
|
104
|
+
<html>
|
|
105
|
+
<head><title>Python 3.12 release notes</title></head>
|
|
106
|
+
<body>
|
|
107
|
+
<h1>What's New In Python 3.12</h1>
|
|
108
|
+
<p>Python 3.12.0 was released on October 2, 2023. It introduces a new
|
|
109
|
+
interactive debugger, improved error messages, and faster startup times.
|
|
110
|
+
The new interactive debugger lets you step through code directly from the
|
|
111
|
+
REPL, and the interpreter now starts up roughly 5% faster than 3.11.</p>
|
|
112
|
+
</body>
|
|
113
|
+
</html>
|
|
114
|
+
```
|
|
115
|
+
|
|
116
|
+
`notes.md`
|
|
117
|
+
```markdown
|
|
118
|
+
# My notes
|
|
119
|
+
|
|
120
|
+
Python 3.12 was released on October 2, 2023, bringing a new interactive
|
|
121
|
+
debugger and faster startup times [1].
|
|
122
|
+
|
|
123
|
+
The moon is made of green cheese [2].
|
|
124
|
+
|
|
125
|
+
## Sources
|
|
126
|
+
|
|
127
|
+
1. source.html
|
|
128
|
+
2. https://nonexistent-citesure-demo.invalid/green-cheese
|
|
129
|
+
```
|
|
130
|
+
|
|
131
|
+
Then verify:
|
|
132
|
+
|
|
133
|
+
```bash
|
|
134
|
+
citesure verify notes.md
|
|
135
|
+
```
|
|
136
|
+
|
|
137
|
+
On the first run the NLI tier lazy-downloads a ~425 MB cross-encoder, so this
|
|
138
|
+
command needs network once. To stay fully offline, add `--no-nli` and use tiers
|
|
139
|
+
1+2 only (link status + content overlap):
|
|
140
|
+
|
|
141
|
+
```bash
|
|
142
|
+
citesure verify notes.md --no-nli
|
|
143
|
+
```
|
|
144
|
+
|
|
145
|
+
Real output (run from a clean venv, zero API keys, ~2 s):
|
|
146
|
+
|
|
147
|
+
```text
|
|
148
|
+
# citesure verification report — 2 citation(s) (input format: markdown)
|
|
149
|
+
|
|
150
|
+
- ✅ **[1]** PASS score=0.944 tier=2 — /tmp/demo/source.html
|
|
151
|
+
- evidence: Python 3.12.0 was released on October 2, 2023. It introduces a new interactive debugger, improved error messages, and faster startup times. The new interactive debugger lets you step through code directly from the REPL, and the interpreter now starts up roughly 5% faster than 3.11.
|
|
152
|
+
- note: page fetched; reachable with no red flags (tier 1)
|
|
153
|
+
- note: overlap tier: score 0.944 against fetched page; thresholds: >= 0.6 supported, >= 0.3 ambiguous, below unsupported
|
|
154
|
+
- ⛔ **[2]** UNVERIFIABLE tier=1 — https://nonexistent-citesure-demo.invalid/green-cheese
|
|
155
|
+
- note: ConnectError: [Errno -2] Name or service not known
|
|
156
|
+
|
|
157
|
+
## Summary
|
|
158
|
+
|
|
159
|
+
- total: 2
|
|
160
|
+
- supported: 1
|
|
161
|
+
- unsupported: 0
|
|
162
|
+
- unreachable: 1
|
|
163
|
+
- paywalled: 0
|
|
164
|
+
- ambiguous: 0
|
|
165
|
+
- unverifiable (unreachable + paywalled + ambiguous): 1
|
|
166
|
+
- pass_rate: 0.5000
|
|
167
|
+
|
|
168
|
+
**Decision: FAIL** (pass_rate 0.5000 < threshold 0.8)
|
|
169
|
+
```
|
|
170
|
+
|
|
171
|
+
Exit code is `1` because `pass_rate` (0.5) is below the default threshold
|
|
172
|
+
(0.8). Citation `[1]` is **supported** (its terms strongly overlap the cited
|
|
173
|
+
page); citation `[2]` is **unreachable** (the host doesn't resolve).
|
|
174
|
+
|
|
175
|
+
The repo also ships a larger bundled sample you can point the CLI at directly:
|
|
176
|
+
|
|
177
|
+
```bash
|
|
178
|
+
citesure verify tests/fixtures/sample.md # 5 citations, all local
|
|
179
|
+
citesure verify tests/fixtures/sample.md --json # machine-readable JSON
|
|
180
|
+
citesure verify tests/fixtures/cases.json --json --strict
|
|
181
|
+
citesure verify tests/fixtures/sample.md --md report.md # also write a .md file
|
|
182
|
+
```
|
|
183
|
+
|
|
184
|
+
## How verification works — the 3-tier pipeline
|
|
185
|
+
|
|
186
|
+
Each citation is pushed through up to three tiers. A verdict records
|
|
187
|
+
`tier_reached` so you can see how far it got.
|
|
188
|
+
|
|
189
|
+
1. **Reachability (tier 1)** — does the URL resolve and fetch? Is it dead,
|
|
190
|
+
retracted, or paywalled? Uses `httpx` + `trafilatura` (HTML → clean text),
|
|
191
|
+
5 s timeout, 2 retries with backoff, max 10 concurrent fetches, a
|
|
192
|
+
citesure-identifying User-Agent, and robots.txt respect. `file://` and bare
|
|
193
|
+
local paths are supported for offline demos/tests. No headless browser.
|
|
194
|
+
2. **Content overlap (tier 2)** — does the claim actually appear in the cited
|
|
195
|
+
region? The sentence/paragraph containing the `[n]` marker is the claim
|
|
196
|
+
unit (D4); it's matched against the top-k most relevant passages of the
|
|
197
|
+
fetched page using deterministic weighted term-coverage scoring in
|
|
198
|
+
`[0, 1]`. No LLM, no embeddings. Compound claims (joined by "and", "while",
|
|
199
|
+
"but") are split into clauses and scored independently — a claim with mixed
|
|
200
|
+
support (some clauses true, some false) returns `ambiguous`, not `supported`.
|
|
201
|
+
When all top-k passages score below 0.3, the fallback scores all passages.
|
|
202
|
+
3. **NLI entailment (tier 3, opt-in)** — a local cross-encoder scores each
|
|
203
|
+
`(claim, best-passage)` pair for entailment. Off by default; enable with
|
|
204
|
+
`--nli`. See [NLI tier](#nli-tier-opt-in).
|
|
205
|
+
|
|
206
|
+
Tiers 1+2 are the default path and run entirely offline. Tier 3 is a
|
|
207
|
+
first-class opt-in, fully built and tested — not a stub.
|
|
208
|
+
|
|
209
|
+
### Verdict statuses (D3)
|
|
210
|
+
|
|
211
|
+
| Status | Meaning |
|
|
212
|
+
|--------|---------|
|
|
213
|
+
| `supported` | Page fetched and claim terms strongly overlap the cited region (score ≥ 0.6). |
|
|
214
|
+
| `unsupported` | Page fetched but the claim is absent from the cited region (score < 0.3), or the page carries a retraction notice. |
|
|
215
|
+
| `unreachable` | DNS failure, HTTP 4xx/5xx, timeout, or robots.txt block. |
|
|
216
|
+
| `paywalled` | Login/paywall detected on the cited page. |
|
|
217
|
+
| `ambiguous` | Partial overlap (0.3 ≤ score < 0.6), or the citation marker could not be located (e.g. a JavaScript-rendered page with no extractable body). |
|
|
218
|
+
|
|
219
|
+
The overall report is `{total, supported, unsupported, unverifiable,
|
|
220
|
+
pass_rate}` where `unverifiable = unreachable + paywalled + ambiguous` and
|
|
221
|
+
`pass_rate = supported / total`.
|
|
222
|
+
|
|
223
|
+
## CLI flags
|
|
224
|
+
|
|
225
|
+
| Flag | Meaning |
|
|
226
|
+
|------|---------|
|
|
227
|
+
| `--json` | Print the full report as JSON (`report.to_dict()`). |
|
|
228
|
+
| `--md FILE` | Also write the Markdown report to `FILE` (in addition to stdout). |
|
|
229
|
+
| `--threshold F` | Minimum `pass_rate` for exit code 0 (default `0.8`). |
|
|
230
|
+
| `--strict` | CI mode: `ambiguous` counts as an explicit failure (see exit codes). |
|
|
231
|
+
| `--nli` | Enable the NLI entailment tier (tier 3). Lazy-downloads the default model (~425 MB) on first use. |
|
|
232
|
+
| `--nli-model NAME` | Cross-encoder model name (HF id) or local path; highest priority in the model-swap chain. Implies `--nli`. |
|
|
233
|
+
| `--cache-dir DIR` | Override the fetch cache directory (sets `CITECHECK_CACHE_DIR`). |
|
|
234
|
+
|
|
235
|
+
### Exit codes
|
|
236
|
+
|
|
237
|
+
* `0` — default mode: iff `pass_rate >= --threshold`. With `--strict`:
|
|
238
|
+
additionally requires **no** `ambiguous` AND **no** `unsupported` verdicts.
|
|
239
|
+
* `1` — verification ran but the rule above is not met.
|
|
240
|
+
* `2` — input error (unreadable file, malformed JSON) OR the NLI model failed
|
|
241
|
+
to load (fail fast, D6 — distinct from a verification-failure exit 1).
|
|
242
|
+
|
|
243
|
+
## Library
|
|
244
|
+
|
|
245
|
+
```python
|
|
246
|
+
import asyncio
|
|
247
|
+
from citesure import extract_citations, verify_citations, score_overlap
|
|
248
|
+
|
|
249
|
+
citations = extract_citations("The sky is blue [1].", {"1": "https://example.com"})
|
|
250
|
+
report = asyncio.run(verify_citations(citations)) # tiers 1+2 by default
|
|
251
|
+
print(report.pass_rate)
|
|
252
|
+
|
|
253
|
+
# Overlap tier in isolation (pure, deterministic):
|
|
254
|
+
score, snippet = score_overlap("Python 3.12 was released in 2023.", ["...page text..."])
|
|
255
|
+
```
|
|
256
|
+
|
|
257
|
+
## MCP server
|
|
258
|
+
|
|
259
|
+
`citesure-mcp` is a stdio MCP server (official `mcp` Python SDK) exposing two
|
|
260
|
+
tools:
|
|
261
|
+
|
|
262
|
+
* **`verify_citations(citations)`** — the structured JSON form: a list of
|
|
263
|
+
`{"claim": str, "citation": url-or-id}` objects (optional per-item
|
|
264
|
+
`"excerpt"`). Returns the D3 report: `{total, supported, unsupported,
|
|
265
|
+
unverifiable, pass_rate, verdicts[]}`.
|
|
266
|
+
* **`verify_markdown(markdown, url_map?)`** — raw markdown with inline
|
|
267
|
+
citations (`[n]` markers resolved via `url_map` or a trailing
|
|
268
|
+
`## Sources`/`## References` section, plus `[label](url)` links). Same
|
|
269
|
+
report shape.
|
|
270
|
+
|
|
271
|
+
Both tools run the full pipeline (reachability + content overlap) and return
|
|
272
|
+
clean JSON-serializable dicts; bad input comes back as `{"error": "..."}`
|
|
273
|
+
rather than a protocol error.
|
|
274
|
+
|
|
275
|
+
### Adding it to an MCP client
|
|
276
|
+
|
|
277
|
+
Generic / Claude Desktop style config (any client that launches stdio MCP
|
|
278
|
+
servers):
|
|
279
|
+
|
|
280
|
+
```json
|
|
281
|
+
{
|
|
282
|
+
"mcpServers": {
|
|
283
|
+
"citesure": {
|
|
284
|
+
"command": "citesure-mcp",
|
|
285
|
+
"args": []
|
|
286
|
+
}
|
|
287
|
+
}
|
|
288
|
+
}
|
|
289
|
+
```
|
|
290
|
+
|
|
291
|
+
If the package isn't on your `PATH`, use the absolute path to the console
|
|
292
|
+
script instead, e.g. `"command": "/path/to/venv/bin/citesure-mcp"`.
|
|
293
|
+
|
|
294
|
+
To enable the NLI tier at server start, set the env vars:
|
|
295
|
+
|
|
296
|
+
```json
|
|
297
|
+
{
|
|
298
|
+
"mcpServers": {
|
|
299
|
+
"citesure": {
|
|
300
|
+
"command": "citesure-mcp",
|
|
301
|
+
"env": {
|
|
302
|
+
"CITECHECK_NLI": "1",
|
|
303
|
+
"CITECHECK_NLI_MODEL": "cross-encoder/nli-deberta-v3-base"
|
|
304
|
+
}
|
|
305
|
+
}
|
|
306
|
+
}
|
|
307
|
+
}
|
|
308
|
+
```
|
|
309
|
+
|
|
310
|
+
`CITECHECK_NLI_MODEL` is optional (the built-in default applies); the model is
|
|
311
|
+
lazy-downloaded on first use (~425 MB) into `~/.cache/citesure/`.
|
|
312
|
+
|
|
313
|
+
## NLI tier (opt-in)
|
|
314
|
+
|
|
315
|
+
NLI entailment (tier 3) is **off by default** — the fast tiers-1+2 path runs
|
|
316
|
+
with no model download. Enable it via the CLI (`--nli`) or, for the MCP
|
|
317
|
+
server, via `CITECHECK_NLI=1` at server start.
|
|
318
|
+
|
|
319
|
+
Model selection follows a priority chain (D6): `--nli-model NAME` flag /
|
|
320
|
+
`nli_model=` library param → `CITECHECK_NLI_MODEL` env var → built-in default
|
|
321
|
+
`cross-encoder/nli-deberta-v3-base`. Any Hugging Face cross-encoder name or
|
|
322
|
+
local path is accepted; citesure fails fast with a clear error (exit 2) if it
|
|
323
|
+
won't load — it never silently falls back to another model.
|
|
324
|
+
|
|
325
|
+
When NLI is on, the report header shows which model was used, and each scorable
|
|
326
|
+
verdict records its entailment score with `tier_reached=3`.
|
|
327
|
+
|
|
328
|
+
## Cache & offline-first design
|
|
329
|
+
|
|
330
|
+
* **Fetch cache** — fetched pages are cached to disk for 24 h, keyed by
|
|
331
|
+
URL+etag, under `~/.cache/citesure/`. Override the location with
|
|
332
|
+
`CITECHECK_CACHE_DIR` (or the CLI `--cache-dir`). Re-verifying the same URLs
|
|
333
|
+
within 24 h makes no network calls.
|
|
334
|
+
* **Offline-first** — the default tiers-1+2 path needs no network beyond the
|
|
335
|
+
cited URLs themselves, no API keys, and no LLM. `file://` and bare local
|
|
336
|
+
paths let you verify entirely offline. The only thing that ever downloads is
|
|
337
|
+
the opt-in NLI model, and only on first `--nli` use.
|
|
338
|
+
* **No headless browser** — JavaScript-rendered pages yield no extractable
|
|
339
|
+
body, so their citations come back `ambiguous` rather than being silently
|
|
340
|
+
guessed at.
|
|
341
|
+
|
|
342
|
+
## Development
|
|
343
|
+
|
|
344
|
+
```bash
|
|
345
|
+
pip install -e ".[dev]" # installs pytest
|
|
346
|
+
pytest -q # offline suite (live + NLI-model tests deselected by default)
|
|
347
|
+
pytest -q -m live # run live-network smoke tests (hit the real internet)
|
|
348
|
+
pytest -q -m nli # run real DeBERTa-v3 model tests (needs the ~425 MB download)
|
|
349
|
+
```
|
|
350
|
+
|
|
351
|
+
Live-network tests (`@pytest.mark.live`) and real-model NLI tests
|
|
352
|
+
(`@pytest.mark.nli`) are skipped by default so the suite is green offline.
|
citesure-0.2.0/README.md
ADDED
|
@@ -0,0 +1,321 @@
|
|
|
1
|
+
<p align="center"><img src="https://raw.githubusercontent.com/DawnofGenX/citesure/main/docs/og-banner.png" alt="citesure" width="100%"/></p>
|
|
2
|
+
|
|
3
|
+
# citesure
|
|
4
|
+
|
|
5
|
+
An MCP server + CLI + library that verifies whether an LLM's claims are actually
|
|
6
|
+
supported by the sources it cites — resolves each citation, fetches the real
|
|
7
|
+
content, flags dead/retracted/paywalled links, and abstains when it can't
|
|
8
|
+
confirm.
|
|
9
|
+
|
|
10
|
+
citesure verifies *claims* (not just whether a link is alive). It needs no API
|
|
11
|
+
keys and no LLM: the fast path (tiers 1+2) is pure `httpx` + `trafilatura`, and
|
|
12
|
+
passing `--no-nli` keeps it fully offline. The NLI tier (tier 3) is on by
|
|
13
|
+
default and lazy-downloads a local cross-encoder on first use.
|
|
14
|
+
|
|
15
|
+
## Accuracy
|
|
16
|
+
|
|
17
|
+
Measured on an independent 108 case set (40 unique URLs, 13 domains), built as a three-way
|
|
18
|
+
contrast: each source fact appears three times — verbatim-supported, negation-flipped, and
|
|
19
|
+
entity-swapped — all sharing one source sentence. The A-vs-B gap isolates polarity handling and
|
|
20
|
+
A-vs-C isolates subject binding, so a drop points at a specific bug class rather than "hard cases".
|
|
21
|
+
|
|
22
|
+
| Configuration | Agreement |
|
|
23
|
+
|---|---|
|
|
24
|
+
| NLI tier on (default claim path) | **83/108 = 76.9%** |
|
|
25
|
+
| NLI tier off (deterministic tiers only) | **41/108 = 38.0%** |
|
|
26
|
+
|
|
27
|
+
The 38.9-point gap is the cross-encoder tier's entire contribution, measured rather than asserted.
|
|
28
|
+
|
|
29
|
+
**Safety metrics** — the failure modes that matter for a verifier:
|
|
30
|
+
|
|
31
|
+
| Metric | Result |
|
|
32
|
+
|---|---|
|
|
33
|
+
| Negation-flip false-supported | 4/32 = 12.5% |
|
|
34
|
+
| Entity-swap false-supported | 7/32 = 21.9% |
|
|
35
|
+
|
|
36
|
+
Both rise to ~90-100% with the NLI tier off, which is the honest cost of the fast path.
|
|
37
|
+
|
|
38
|
+
An exhaustive threshold sweep (support 0.30-0.90 x ambiguous 0.05-support) moved agreement from
|
|
39
|
+
66.7% to 69.4% — a gain of exactly one item. Threshold tuning is a dead end here; see
|
|
40
|
+
`evals/THRESHOLD_CALIBRATION.md`.
|
|
41
|
+
|
|
42
|
+
## Install
|
|
43
|
+
|
|
44
|
+
Default install stays light — no torch, no model download:
|
|
45
|
+
|
|
46
|
+
```bash
|
|
47
|
+
pip install citesure # tiers 1+2: httpx + trafilatura only
|
|
48
|
+
pip install "citesure[nli]" # adds the opt-in NLI cross-encoder tier
|
|
49
|
+
```
|
|
50
|
+
|
|
51
|
+
Or from a checkout:
|
|
52
|
+
|
|
53
|
+
```bash
|
|
54
|
+
git clone https://github.com/DawnofGenX/citesure.git && cd citesure
|
|
55
|
+
pip install . # default: offline, key-free, no torch
|
|
56
|
+
pip install ".[nli]" # adds the opt-in NLI tier
|
|
57
|
+
pip install -e ".[dev]" # editable + test deps
|
|
58
|
+
```
|
|
59
|
+
|
|
60
|
+
`requires-python >= 3.10`. The default fast path (tiers 1+2) needs only
|
|
61
|
+
`httpx` + `trafilatura`. The opt-in NLI tier pulls in the local cross-encoder
|
|
62
|
+
stack (`torch`, `transformers`, `sentence-transformers`) and lazy-downloads a
|
|
63
|
+
~425 MB model on first `--nli` use — see [NLI tier](#nli-tier-opt-in).
|
|
64
|
+
|
|
65
|
+
## Quickstart
|
|
66
|
+
|
|
67
|
+
Create two small files in any directory — a claim document and the page it
|
|
68
|
+
cites:
|
|
69
|
+
|
|
70
|
+
`source.html`
|
|
71
|
+
```html
|
|
72
|
+
<!DOCTYPE html>
|
|
73
|
+
<html>
|
|
74
|
+
<head><title>Python 3.12 release notes</title></head>
|
|
75
|
+
<body>
|
|
76
|
+
<h1>What's New In Python 3.12</h1>
|
|
77
|
+
<p>Python 3.12.0 was released on October 2, 2023. It introduces a new
|
|
78
|
+
interactive debugger, improved error messages, and faster startup times.
|
|
79
|
+
The new interactive debugger lets you step through code directly from the
|
|
80
|
+
REPL, and the interpreter now starts up roughly 5% faster than 3.11.</p>
|
|
81
|
+
</body>
|
|
82
|
+
</html>
|
|
83
|
+
```
|
|
84
|
+
|
|
85
|
+
`notes.md`
|
|
86
|
+
```markdown
|
|
87
|
+
# My notes
|
|
88
|
+
|
|
89
|
+
Python 3.12 was released on October 2, 2023, bringing a new interactive
|
|
90
|
+
debugger and faster startup times [1].
|
|
91
|
+
|
|
92
|
+
The moon is made of green cheese [2].
|
|
93
|
+
|
|
94
|
+
## Sources
|
|
95
|
+
|
|
96
|
+
1. source.html
|
|
97
|
+
2. https://nonexistent-citesure-demo.invalid/green-cheese
|
|
98
|
+
```
|
|
99
|
+
|
|
100
|
+
Then verify:
|
|
101
|
+
|
|
102
|
+
```bash
|
|
103
|
+
citesure verify notes.md
|
|
104
|
+
```
|
|
105
|
+
|
|
106
|
+
On the first run the NLI tier lazy-downloads a ~425 MB cross-encoder, so this
|
|
107
|
+
command needs network once. To stay fully offline, add `--no-nli` and use tiers
|
|
108
|
+
1+2 only (link status + content overlap):
|
|
109
|
+
|
|
110
|
+
```bash
|
|
111
|
+
citesure verify notes.md --no-nli
|
|
112
|
+
```
|
|
113
|
+
|
|
114
|
+
Real output (run from a clean venv, zero API keys, ~2 s):
|
|
115
|
+
|
|
116
|
+
```text
|
|
117
|
+
# citesure verification report — 2 citation(s) (input format: markdown)
|
|
118
|
+
|
|
119
|
+
- ✅ **[1]** PASS score=0.944 tier=2 — /tmp/demo/source.html
|
|
120
|
+
- evidence: Python 3.12.0 was released on October 2, 2023. It introduces a new interactive debugger, improved error messages, and faster startup times. The new interactive debugger lets you step through code directly from the REPL, and the interpreter now starts up roughly 5% faster than 3.11.
|
|
121
|
+
- note: page fetched; reachable with no red flags (tier 1)
|
|
122
|
+
- note: overlap tier: score 0.944 against fetched page; thresholds: >= 0.6 supported, >= 0.3 ambiguous, below unsupported
|
|
123
|
+
- ⛔ **[2]** UNVERIFIABLE tier=1 — https://nonexistent-citesure-demo.invalid/green-cheese
|
|
124
|
+
- note: ConnectError: [Errno -2] Name or service not known
|
|
125
|
+
|
|
126
|
+
## Summary
|
|
127
|
+
|
|
128
|
+
- total: 2
|
|
129
|
+
- supported: 1
|
|
130
|
+
- unsupported: 0
|
|
131
|
+
- unreachable: 1
|
|
132
|
+
- paywalled: 0
|
|
133
|
+
- ambiguous: 0
|
|
134
|
+
- unverifiable (unreachable + paywalled + ambiguous): 1
|
|
135
|
+
- pass_rate: 0.5000
|
|
136
|
+
|
|
137
|
+
**Decision: FAIL** (pass_rate 0.5000 < threshold 0.8)
|
|
138
|
+
```
|
|
139
|
+
|
|
140
|
+
Exit code is `1` because `pass_rate` (0.5) is below the default threshold
|
|
141
|
+
(0.8). Citation `[1]` is **supported** (its terms strongly overlap the cited
|
|
142
|
+
page); citation `[2]` is **unreachable** (the host doesn't resolve).
|
|
143
|
+
|
|
144
|
+
The repo also ships a larger bundled sample you can point the CLI at directly:
|
|
145
|
+
|
|
146
|
+
```bash
|
|
147
|
+
citesure verify tests/fixtures/sample.md # 5 citations, all local
|
|
148
|
+
citesure verify tests/fixtures/sample.md --json # machine-readable JSON
|
|
149
|
+
citesure verify tests/fixtures/cases.json --json --strict
|
|
150
|
+
citesure verify tests/fixtures/sample.md --md report.md # also write a .md file
|
|
151
|
+
```
|
|
152
|
+
|
|
153
|
+
## How verification works — the 3-tier pipeline
|
|
154
|
+
|
|
155
|
+
Each citation is pushed through up to three tiers. A verdict records
|
|
156
|
+
`tier_reached` so you can see how far it got.
|
|
157
|
+
|
|
158
|
+
1. **Reachability (tier 1)** — does the URL resolve and fetch? Is it dead,
|
|
159
|
+
retracted, or paywalled? Uses `httpx` + `trafilatura` (HTML → clean text),
|
|
160
|
+
5 s timeout, 2 retries with backoff, max 10 concurrent fetches, a
|
|
161
|
+
citesure-identifying User-Agent, and robots.txt respect. `file://` and bare
|
|
162
|
+
local paths are supported for offline demos/tests. No headless browser.
|
|
163
|
+
2. **Content overlap (tier 2)** — does the claim actually appear in the cited
|
|
164
|
+
region? The sentence/paragraph containing the `[n]` marker is the claim
|
|
165
|
+
unit (D4); it's matched against the top-k most relevant passages of the
|
|
166
|
+
fetched page using deterministic weighted term-coverage scoring in
|
|
167
|
+
`[0, 1]`. No LLM, no embeddings. Compound claims (joined by "and", "while",
|
|
168
|
+
"but") are split into clauses and scored independently — a claim with mixed
|
|
169
|
+
support (some clauses true, some false) returns `ambiguous`, not `supported`.
|
|
170
|
+
When all top-k passages score below 0.3, the fallback scores all passages.
|
|
171
|
+
3. **NLI entailment (tier 3, opt-in)** — a local cross-encoder scores each
|
|
172
|
+
`(claim, best-passage)` pair for entailment. Off by default; enable with
|
|
173
|
+
`--nli`. See [NLI tier](#nli-tier-opt-in).
|
|
174
|
+
|
|
175
|
+
Tiers 1+2 are the default path and run entirely offline. Tier 3 is a
|
|
176
|
+
first-class opt-in, fully built and tested — not a stub.
|
|
177
|
+
|
|
178
|
+
### Verdict statuses (D3)
|
|
179
|
+
|
|
180
|
+
| Status | Meaning |
|
|
181
|
+
|--------|---------|
|
|
182
|
+
| `supported` | Page fetched and claim terms strongly overlap the cited region (score ≥ 0.6). |
|
|
183
|
+
| `unsupported` | Page fetched but the claim is absent from the cited region (score < 0.3), or the page carries a retraction notice. |
|
|
184
|
+
| `unreachable` | DNS failure, HTTP 4xx/5xx, timeout, or robots.txt block. |
|
|
185
|
+
| `paywalled` | Login/paywall detected on the cited page. |
|
|
186
|
+
| `ambiguous` | Partial overlap (0.3 ≤ score < 0.6), or the citation marker could not be located (e.g. a JavaScript-rendered page with no extractable body). |
|
|
187
|
+
|
|
188
|
+
The overall report is `{total, supported, unsupported, unverifiable,
|
|
189
|
+
pass_rate}` where `unverifiable = unreachable + paywalled + ambiguous` and
|
|
190
|
+
`pass_rate = supported / total`.
|
|
191
|
+
|
|
192
|
+
## CLI flags
|
|
193
|
+
|
|
194
|
+
| Flag | Meaning |
|
|
195
|
+
|------|---------|
|
|
196
|
+
| `--json` | Print the full report as JSON (`report.to_dict()`). |
|
|
197
|
+
| `--md FILE` | Also write the Markdown report to `FILE` (in addition to stdout). |
|
|
198
|
+
| `--threshold F` | Minimum `pass_rate` for exit code 0 (default `0.8`). |
|
|
199
|
+
| `--strict` | CI mode: `ambiguous` counts as an explicit failure (see exit codes). |
|
|
200
|
+
| `--nli` | Enable the NLI entailment tier (tier 3). Lazy-downloads the default model (~425 MB) on first use. |
|
|
201
|
+
| `--nli-model NAME` | Cross-encoder model name (HF id) or local path; highest priority in the model-swap chain. Implies `--nli`. |
|
|
202
|
+
| `--cache-dir DIR` | Override the fetch cache directory (sets `CITECHECK_CACHE_DIR`). |
|
|
203
|
+
|
|
204
|
+
### Exit codes
|
|
205
|
+
|
|
206
|
+
* `0` — default mode: iff `pass_rate >= --threshold`. With `--strict`:
|
|
207
|
+
additionally requires **no** `ambiguous` AND **no** `unsupported` verdicts.
|
|
208
|
+
* `1` — verification ran but the rule above is not met.
|
|
209
|
+
* `2` — input error (unreadable file, malformed JSON) OR the NLI model failed
|
|
210
|
+
to load (fail fast, D6 — distinct from a verification-failure exit 1).
|
|
211
|
+
|
|
212
|
+
## Library
|
|
213
|
+
|
|
214
|
+
```python
|
|
215
|
+
import asyncio
|
|
216
|
+
from citesure import extract_citations, verify_citations, score_overlap
|
|
217
|
+
|
|
218
|
+
citations = extract_citations("The sky is blue [1].", {"1": "https://example.com"})
|
|
219
|
+
report = asyncio.run(verify_citations(citations)) # tiers 1+2 by default
|
|
220
|
+
print(report.pass_rate)
|
|
221
|
+
|
|
222
|
+
# Overlap tier in isolation (pure, deterministic):
|
|
223
|
+
score, snippet = score_overlap("Python 3.12 was released in 2023.", ["...page text..."])
|
|
224
|
+
```
|
|
225
|
+
|
|
226
|
+
## MCP server
|
|
227
|
+
|
|
228
|
+
`citesure-mcp` is a stdio MCP server (official `mcp` Python SDK) exposing two
|
|
229
|
+
tools:
|
|
230
|
+
|
|
231
|
+
* **`verify_citations(citations)`** — the structured JSON form: a list of
|
|
232
|
+
`{"claim": str, "citation": url-or-id}` objects (optional per-item
|
|
233
|
+
`"excerpt"`). Returns the D3 report: `{total, supported, unsupported,
|
|
234
|
+
unverifiable, pass_rate, verdicts[]}`.
|
|
235
|
+
* **`verify_markdown(markdown, url_map?)`** — raw markdown with inline
|
|
236
|
+
citations (`[n]` markers resolved via `url_map` or a trailing
|
|
237
|
+
`## Sources`/`## References` section, plus `[label](url)` links). Same
|
|
238
|
+
report shape.
|
|
239
|
+
|
|
240
|
+
Both tools run the full pipeline (reachability + content overlap) and return
|
|
241
|
+
clean JSON-serializable dicts; bad input comes back as `{"error": "..."}`
|
|
242
|
+
rather than a protocol error.
|
|
243
|
+
|
|
244
|
+
### Adding it to an MCP client
|
|
245
|
+
|
|
246
|
+
Generic / Claude Desktop style config (any client that launches stdio MCP
|
|
247
|
+
servers):
|
|
248
|
+
|
|
249
|
+
```json
|
|
250
|
+
{
|
|
251
|
+
"mcpServers": {
|
|
252
|
+
"citesure": {
|
|
253
|
+
"command": "citesure-mcp",
|
|
254
|
+
"args": []
|
|
255
|
+
}
|
|
256
|
+
}
|
|
257
|
+
}
|
|
258
|
+
```
|
|
259
|
+
|
|
260
|
+
If the package isn't on your `PATH`, use the absolute path to the console
|
|
261
|
+
script instead, e.g. `"command": "/path/to/venv/bin/citesure-mcp"`.
|
|
262
|
+
|
|
263
|
+
To enable the NLI tier at server start, set the env vars:
|
|
264
|
+
|
|
265
|
+
```json
|
|
266
|
+
{
|
|
267
|
+
"mcpServers": {
|
|
268
|
+
"citesure": {
|
|
269
|
+
"command": "citesure-mcp",
|
|
270
|
+
"env": {
|
|
271
|
+
"CITECHECK_NLI": "1",
|
|
272
|
+
"CITECHECK_NLI_MODEL": "cross-encoder/nli-deberta-v3-base"
|
|
273
|
+
}
|
|
274
|
+
}
|
|
275
|
+
}
|
|
276
|
+
}
|
|
277
|
+
```
|
|
278
|
+
|
|
279
|
+
`CITECHECK_NLI_MODEL` is optional (the built-in default applies); the model is
|
|
280
|
+
lazy-downloaded on first use (~425 MB) into `~/.cache/citesure/`.
|
|
281
|
+
|
|
282
|
+
## NLI tier (opt-in)
|
|
283
|
+
|
|
284
|
+
NLI entailment (tier 3) is **off by default** — the fast tiers-1+2 path runs
|
|
285
|
+
with no model download. Enable it via the CLI (`--nli`) or, for the MCP
|
|
286
|
+
server, via `CITECHECK_NLI=1` at server start.
|
|
287
|
+
|
|
288
|
+
Model selection follows a priority chain (D6): `--nli-model NAME` flag /
|
|
289
|
+
`nli_model=` library param → `CITECHECK_NLI_MODEL` env var → built-in default
|
|
290
|
+
`cross-encoder/nli-deberta-v3-base`. Any Hugging Face cross-encoder name or
|
|
291
|
+
local path is accepted; citesure fails fast with a clear error (exit 2) if it
|
|
292
|
+
won't load — it never silently falls back to another model.
|
|
293
|
+
|
|
294
|
+
When NLI is on, the report header shows which model was used, and each scorable
|
|
295
|
+
verdict records its entailment score with `tier_reached=3`.
|
|
296
|
+
|
|
297
|
+
## Cache & offline-first design
|
|
298
|
+
|
|
299
|
+
* **Fetch cache** — fetched pages are cached to disk for 24 h, keyed by
|
|
300
|
+
URL+etag, under `~/.cache/citesure/`. Override the location with
|
|
301
|
+
`CITECHECK_CACHE_DIR` (or the CLI `--cache-dir`). Re-verifying the same URLs
|
|
302
|
+
within 24 h makes no network calls.
|
|
303
|
+
* **Offline-first** — the default tiers-1+2 path needs no network beyond the
|
|
304
|
+
cited URLs themselves, no API keys, and no LLM. `file://` and bare local
|
|
305
|
+
paths let you verify entirely offline. The only thing that ever downloads is
|
|
306
|
+
the opt-in NLI model, and only on first `--nli` use.
|
|
307
|
+
* **No headless browser** — JavaScript-rendered pages yield no extractable
|
|
308
|
+
body, so their citations come back `ambiguous` rather than being silently
|
|
309
|
+
guessed at.
|
|
310
|
+
|
|
311
|
+
## Development
|
|
312
|
+
|
|
313
|
+
```bash
|
|
314
|
+
pip install -e ".[dev]" # installs pytest
|
|
315
|
+
pytest -q # offline suite (live + NLI-model tests deselected by default)
|
|
316
|
+
pytest -q -m live # run live-network smoke tests (hit the real internet)
|
|
317
|
+
pytest -q -m nli # run real DeBERTa-v3 model tests (needs the ~425 MB download)
|
|
318
|
+
```
|
|
319
|
+
|
|
320
|
+
Live-network tests (`@pytest.mark.live`) and real-model NLI tests
|
|
321
|
+
(`@pytest.mark.nli`) are skipped by default so the suite is green offline.
|