citesure 0.2.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (44) hide show
  1. citesure-0.2.0/PKG-INFO +352 -0
  2. citesure-0.2.0/README.md +321 -0
  3. citesure-0.2.0/pyproject.toml +75 -0
  4. citesure-0.2.0/setup.cfg +4 -0
  5. citesure-0.2.0/src/citesure/__init__.py +70 -0
  6. citesure-0.2.0/src/citesure/citations.py +321 -0
  7. citesure-0.2.0/src/citesure/cli.py +208 -0
  8. citesure-0.2.0/src/citesure/fetcher.py +527 -0
  9. citesure-0.2.0/src/citesure/mcp_server.py +213 -0
  10. citesure-0.2.0/src/citesure/models.py +147 -0
  11. citesure-0.2.0/src/citesure/nli.py +624 -0
  12. citesure-0.2.0/src/citesure/overlap.py +911 -0
  13. citesure-0.2.0/src/citesure/reachability.py +125 -0
  14. citesure-0.2.0/src/citesure/report.py +128 -0
  15. citesure-0.2.0/src/citesure.egg-info/PKG-INFO +352 -0
  16. citesure-0.2.0/src/citesure.egg-info/SOURCES.txt +42 -0
  17. citesure-0.2.0/src/citesure.egg-info/dependency_links.txt +1 -0
  18. citesure-0.2.0/src/citesure.egg-info/entry_points.txt +3 -0
  19. citesure-0.2.0/src/citesure.egg-info/requires.txt +13 -0
  20. citesure-0.2.0/src/citesure.egg-info/top_level.txt +1 -0
  21. citesure-0.2.0/tests/test_citations.py +293 -0
  22. citesure-0.2.0/tests/test_cli.py +219 -0
  23. citesure-0.2.0/tests/test_cli_defaults.py +99 -0
  24. citesure-0.2.0/tests/test_compound_claim_floor.py +90 -0
  25. citesure-0.2.0/tests/test_concurrency.py +170 -0
  26. citesure-0.2.0/tests/test_context_negation.py +85 -0
  27. citesure-0.2.0/tests/test_contradiction.py +174 -0
  28. citesure-0.2.0/tests/test_corpus.py +196 -0
  29. citesure-0.2.0/tests/test_edge_cases.py +299 -0
  30. citesure-0.2.0/tests/test_eval_set_integrity.py +90 -0
  31. citesure-0.2.0/tests/test_fetcher.py +377 -0
  32. citesure-0.2.0/tests/test_fuzz_extraction.py +236 -0
  33. citesure-0.2.0/tests/test_fuzz_report.py +234 -0
  34. citesure-0.2.0/tests/test_fuzz_urls.py +157 -0
  35. citesure-0.2.0/tests/test_inference_cost.py +43 -0
  36. citesure-0.2.0/tests/test_integration.py +130 -0
  37. citesure-0.2.0/tests/test_live_smoke.py +50 -0
  38. citesure-0.2.0/tests/test_mcp_integration.py +207 -0
  39. citesure-0.2.0/tests/test_mcp_stress.py +228 -0
  40. citesure-0.2.0/tests/test_nli.py +757 -0
  41. citesure-0.2.0/tests/test_overlap.py +729 -0
  42. citesure-0.2.0/tests/test_packaging.py +43 -0
  43. citesure-0.2.0/tests/test_reachability.py +148 -0
  44. citesure-0.2.0/tests/test_topk_max.py +59 -0
@@ -0,0 +1,352 @@
1
+ Metadata-Version: 2.4
2
+ Name: citesure
3
+ Version: 0.2.0
4
+ Summary: MCP citation verifier: check whether an LLM's claims are actually supported by the sources it cites.
5
+ Author: DawnofGenX
6
+ License: MIT
7
+ Project-URL: Homepage, https://github.com/DawnofGenX/citesure
8
+ Project-URL: Repository, https://github.com/DawnofGenX/citesure
9
+ Keywords: citations,verification,mcp,hallucination,nli,fact-checking
10
+ Classifier: Development Status :: 3 - Alpha
11
+ Classifier: Intended Audience :: Developers
12
+ Classifier: Programming Language :: Python :: 3
13
+ Classifier: Programming Language :: Python :: 3.10
14
+ Classifier: Programming Language :: Python :: 3.11
15
+ Classifier: Programming Language :: Python :: 3.12
16
+ Classifier: Topic :: Software Development :: Libraries
17
+ Classifier: Topic :: Text Processing :: Linguistic
18
+ Requires-Python: >=3.10
19
+ Description-Content-Type: text/markdown
20
+ Requires-Dist: httpx>=0.27
21
+ Requires-Dist: trafilatura>=1.9
22
+ Requires-Dist: lxml_html_clean>=0.1
23
+ Requires-Dist: mcp>=1.0
24
+ Provides-Extra: nli
25
+ Requires-Dist: transformers>=4.40; extra == "nli"
26
+ Requires-Dist: torch>=2.0; extra == "nli"
27
+ Requires-Dist: sentence-transformers>=2.6; extra == "nli"
28
+ Provides-Extra: dev
29
+ Requires-Dist: pytest>=8.0; extra == "dev"
30
+ Requires-Dist: hypothesis>=6.0; extra == "dev"
31
+
32
+ <p align="center"><img src="https://raw.githubusercontent.com/DawnofGenX/citesure/main/docs/og-banner.png" alt="citesure" width="100%"/></p>
33
+
34
+ # citesure
35
+
36
+ An MCP server + CLI + library that verifies whether an LLM's claims are actually
37
+ supported by the sources it cites — resolves each citation, fetches the real
38
+ content, flags dead/retracted/paywalled links, and abstains when it can't
39
+ confirm.
40
+
41
+ citesure verifies *claims* (not just whether a link is alive). It needs no API
42
+ keys and no LLM: the fast path (tiers 1+2) is pure `httpx` + `trafilatura`, and
43
+ passing `--no-nli` keeps it fully offline. The NLI tier (tier 3) is on by
44
+ default and lazy-downloads a local cross-encoder on first use.
45
+
46
+ ## Accuracy
47
+
48
+ Measured on an independent 108 case set (40 unique URLs, 13 domains), built as a three-way
49
+ contrast: each source fact appears three times — verbatim-supported, negation-flipped, and
50
+ entity-swapped — all sharing one source sentence. The A-vs-B gap isolates polarity handling and
51
+ A-vs-C isolates subject binding, so a drop points at a specific bug class rather than "hard cases".
52
+
53
+ | Configuration | Agreement |
54
+ |---|---|
55
+ | NLI tier on (default claim path) | **83/108 = 76.9%** |
56
+ | NLI tier off (deterministic tiers only) | **41/108 = 38.0%** |
57
+
58
+ The 38.9-point gap is the cross-encoder tier's entire contribution, measured rather than asserted.
59
+
60
+ **Safety metrics** — the failure modes that matter for a verifier:
61
+
62
+ | Metric | Result |
63
+ |---|---|
64
+ | Negation-flip false-supported | 4/32 = 12.5% |
65
+ | Entity-swap false-supported | 7/32 = 21.9% |
66
+
67
+ Both rise to ~90-100% with the NLI tier off, which is the honest cost of the fast path.
68
+
69
+ An exhaustive threshold sweep (support 0.30-0.90 x ambiguous 0.05-support) moved agreement from
70
+ 66.7% to 69.4% — a gain of exactly one item. Threshold tuning is a dead end here; see
71
+ `evals/THRESHOLD_CALIBRATION.md`.
72
+
73
+ ## Install
74
+
75
+ Default install stays light — no torch, no model download:
76
+
77
+ ```bash
78
+ pip install citesure # tiers 1+2: httpx + trafilatura only
79
+ pip install "citesure[nli]" # adds the opt-in NLI cross-encoder tier
80
+ ```
81
+
82
+ Or from a checkout:
83
+
84
+ ```bash
85
+ git clone https://github.com/DawnofGenX/citesure.git && cd citesure
86
+ pip install . # default: offline, key-free, no torch
87
+ pip install ".[nli]" # adds the opt-in NLI tier
88
+ pip install -e ".[dev]" # editable + test deps
89
+ ```
90
+
91
+ `requires-python >= 3.10`. The default fast path (tiers 1+2) needs only
92
+ `httpx` + `trafilatura`. The opt-in NLI tier pulls in the local cross-encoder
93
+ stack (`torch`, `transformers`, `sentence-transformers`) and lazy-downloads a
94
+ ~425 MB model on first `--nli` use — see [NLI tier](#nli-tier-opt-in).
95
+
96
+ ## Quickstart
97
+
98
+ Create two small files in any directory — a claim document and the page it
99
+ cites:
100
+
101
+ `source.html`
102
+ ```html
103
+ <!DOCTYPE html>
104
+ <html>
105
+ <head><title>Python 3.12 release notes</title></head>
106
+ <body>
107
+ <h1>What's New In Python 3.12</h1>
108
+ <p>Python 3.12.0 was released on October 2, 2023. It introduces a new
109
+ interactive debugger, improved error messages, and faster startup times.
110
+ The new interactive debugger lets you step through code directly from the
111
+ REPL, and the interpreter now starts up roughly 5% faster than 3.11.</p>
112
+ </body>
113
+ </html>
114
+ ```
115
+
116
+ `notes.md`
117
+ ```markdown
118
+ # My notes
119
+
120
+ Python 3.12 was released on October 2, 2023, bringing a new interactive
121
+ debugger and faster startup times [1].
122
+
123
+ The moon is made of green cheese [2].
124
+
125
+ ## Sources
126
+
127
+ 1. source.html
128
+ 2. https://nonexistent-citesure-demo.invalid/green-cheese
129
+ ```
130
+
131
+ Then verify:
132
+
133
+ ```bash
134
+ citesure verify notes.md
135
+ ```
136
+
137
+ On the first run the NLI tier lazy-downloads a ~425 MB cross-encoder, so this
138
+ command needs network once. To stay fully offline, add `--no-nli` and use tiers
139
+ 1+2 only (link status + content overlap):
140
+
141
+ ```bash
142
+ citesure verify notes.md --no-nli
143
+ ```
144
+
145
+ Real output (run from a clean venv, zero API keys, ~2 s):
146
+
147
+ ```text
148
+ # citesure verification report — 2 citation(s) (input format: markdown)
149
+
150
+ - ✅ **[1]** PASS score=0.944 tier=2 — /tmp/demo/source.html
151
+ - evidence: Python 3.12.0 was released on October 2, 2023. It introduces a new interactive debugger, improved error messages, and faster startup times. The new interactive debugger lets you step through code directly from the REPL, and the interpreter now starts up roughly 5% faster than 3.11.
152
+ - note: page fetched; reachable with no red flags (tier 1)
153
+ - note: overlap tier: score 0.944 against fetched page; thresholds: >= 0.6 supported, >= 0.3 ambiguous, below unsupported
154
+ - ⛔ **[2]** UNVERIFIABLE tier=1 — https://nonexistent-citesure-demo.invalid/green-cheese
155
+ - note: ConnectError: [Errno -2] Name or service not known
156
+
157
+ ## Summary
158
+
159
+ - total: 2
160
+ - supported: 1
161
+ - unsupported: 0
162
+ - unreachable: 1
163
+ - paywalled: 0
164
+ - ambiguous: 0
165
+ - unverifiable (unreachable + paywalled + ambiguous): 1
166
+ - pass_rate: 0.5000
167
+
168
+ **Decision: FAIL** (pass_rate 0.5000 < threshold 0.8)
169
+ ```
170
+
171
+ Exit code is `1` because `pass_rate` (0.5) is below the default threshold
172
+ (0.8). Citation `[1]` is **supported** (its terms strongly overlap the cited
173
+ page); citation `[2]` is **unreachable** (the host doesn't resolve).
174
+
175
+ The repo also ships a larger bundled sample you can point the CLI at directly:
176
+
177
+ ```bash
178
+ citesure verify tests/fixtures/sample.md # 5 citations, all local
179
+ citesure verify tests/fixtures/sample.md --json # machine-readable JSON
180
+ citesure verify tests/fixtures/cases.json --json --strict
181
+ citesure verify tests/fixtures/sample.md --md report.md # also write a .md file
182
+ ```
183
+
184
+ ## How verification works — the 3-tier pipeline
185
+
186
+ Each citation is pushed through up to three tiers. A verdict records
187
+ `tier_reached` so you can see how far it got.
188
+
189
+ 1. **Reachability (tier 1)** — does the URL resolve and fetch? Is it dead,
190
+ retracted, or paywalled? Uses `httpx` + `trafilatura` (HTML → clean text),
191
+ 5 s timeout, 2 retries with backoff, max 10 concurrent fetches, a
192
+ citesure-identifying User-Agent, and robots.txt respect. `file://` and bare
193
+ local paths are supported for offline demos/tests. No headless browser.
194
+ 2. **Content overlap (tier 2)** — does the claim actually appear in the cited
195
+ region? The sentence/paragraph containing the `[n]` marker is the claim
196
+ unit (D4); it's matched against the top-k most relevant passages of the
197
+ fetched page using deterministic weighted term-coverage scoring in
198
+ `[0, 1]`. No LLM, no embeddings. Compound claims (joined by "and", "while",
199
+ "but") are split into clauses and scored independently — a claim with mixed
200
+ support (some clauses true, some false) returns `ambiguous`, not `supported`.
201
+ When all top-k passages score below 0.3, the fallback scores all passages.
202
+ 3. **NLI entailment (tier 3, opt-in)** — a local cross-encoder scores each
203
+ `(claim, best-passage)` pair for entailment. Off by default; enable with
204
+ `--nli`. See [NLI tier](#nli-tier-opt-in).
205
+
206
+ Tiers 1+2 are the default path and run entirely offline. Tier 3 is a
207
+ first-class opt-in, fully built and tested — not a stub.
208
+
209
+ ### Verdict statuses (D3)
210
+
211
+ | Status | Meaning |
212
+ |--------|---------|
213
+ | `supported` | Page fetched and claim terms strongly overlap the cited region (score ≥ 0.6). |
214
+ | `unsupported` | Page fetched but the claim is absent from the cited region (score < 0.3), or the page carries a retraction notice. |
215
+ | `unreachable` | DNS failure, HTTP 4xx/5xx, timeout, or robots.txt block. |
216
+ | `paywalled` | Login/paywall detected on the cited page. |
217
+ | `ambiguous` | Partial overlap (0.3 ≤ score < 0.6), or the citation marker could not be located (e.g. a JavaScript-rendered page with no extractable body). |
218
+
219
+ The overall report is `{total, supported, unsupported, unverifiable,
220
+ pass_rate}` where `unverifiable = unreachable + paywalled + ambiguous` and
221
+ `pass_rate = supported / total`.
222
+
223
+ ## CLI flags
224
+
225
+ | Flag | Meaning |
226
+ |------|---------|
227
+ | `--json` | Print the full report as JSON (`report.to_dict()`). |
228
+ | `--md FILE` | Also write the Markdown report to `FILE` (in addition to stdout). |
229
+ | `--threshold F` | Minimum `pass_rate` for exit code 0 (default `0.8`). |
230
+ | `--strict` | CI mode: `ambiguous` counts as an explicit failure (see exit codes). |
231
+ | `--nli` | Enable the NLI entailment tier (tier 3). Lazy-downloads the default model (~425 MB) on first use. |
232
+ | `--nli-model NAME` | Cross-encoder model name (HF id) or local path; highest priority in the model-swap chain. Implies `--nli`. |
233
+ | `--cache-dir DIR` | Override the fetch cache directory (sets `CITECHECK_CACHE_DIR`). |
234
+
235
+ ### Exit codes
236
+
237
+ * `0` — default mode: iff `pass_rate >= --threshold`. With `--strict`:
238
+ additionally requires **no** `ambiguous` AND **no** `unsupported` verdicts.
239
+ * `1` — verification ran but the rule above is not met.
240
+ * `2` — input error (unreadable file, malformed JSON) OR the NLI model failed
241
+ to load (fail fast, D6 — distinct from a verification-failure exit 1).
242
+
243
+ ## Library
244
+
245
+ ```python
246
+ import asyncio
247
+ from citesure import extract_citations, verify_citations, score_overlap
248
+
249
+ citations = extract_citations("The sky is blue [1].", {"1": "https://example.com"})
250
+ report = asyncio.run(verify_citations(citations)) # tiers 1+2 by default
251
+ print(report.pass_rate)
252
+
253
+ # Overlap tier in isolation (pure, deterministic):
254
+ score, snippet = score_overlap("Python 3.12 was released in 2023.", ["...page text..."])
255
+ ```
256
+
257
+ ## MCP server
258
+
259
+ `citesure-mcp` is a stdio MCP server (official `mcp` Python SDK) exposing two
260
+ tools:
261
+
262
+ * **`verify_citations(citations)`** — the structured JSON form: a list of
263
+ `{"claim": str, "citation": url-or-id}` objects (optional per-item
264
+ `"excerpt"`). Returns the D3 report: `{total, supported, unsupported,
265
+ unverifiable, pass_rate, verdicts[]}`.
266
+ * **`verify_markdown(markdown, url_map?)`** — raw markdown with inline
267
+ citations (`[n]` markers resolved via `url_map` or a trailing
268
+ `## Sources`/`## References` section, plus `[label](url)` links). Same
269
+ report shape.
270
+
271
+ Both tools run the full pipeline (reachability + content overlap) and return
272
+ clean JSON-serializable dicts; bad input comes back as `{"error": "..."}`
273
+ rather than a protocol error.
274
+
275
+ ### Adding it to an MCP client
276
+
277
+ Generic / Claude Desktop style config (any client that launches stdio MCP
278
+ servers):
279
+
280
+ ```json
281
+ {
282
+ "mcpServers": {
283
+ "citesure": {
284
+ "command": "citesure-mcp",
285
+ "args": []
286
+ }
287
+ }
288
+ }
289
+ ```
290
+
291
+ If the package isn't on your `PATH`, use the absolute path to the console
292
+ script instead, e.g. `"command": "/path/to/venv/bin/citesure-mcp"`.
293
+
294
+ To enable the NLI tier at server start, set the env vars:
295
+
296
+ ```json
297
+ {
298
+ "mcpServers": {
299
+ "citesure": {
300
+ "command": "citesure-mcp",
301
+ "env": {
302
+ "CITECHECK_NLI": "1",
303
+ "CITECHECK_NLI_MODEL": "cross-encoder/nli-deberta-v3-base"
304
+ }
305
+ }
306
+ }
307
+ }
308
+ ```
309
+
310
+ `CITECHECK_NLI_MODEL` is optional (the built-in default applies); the model is
311
+ lazy-downloaded on first use (~425 MB) into `~/.cache/citesure/`.
312
+
313
+ ## NLI tier (opt-in)
314
+
315
+ NLI entailment (tier 3) is **off by default** — the fast tiers-1+2 path runs
316
+ with no model download. Enable it via the CLI (`--nli`) or, for the MCP
317
+ server, via `CITECHECK_NLI=1` at server start.
318
+
319
+ Model selection follows a priority chain (D6): `--nli-model NAME` flag /
320
+ `nli_model=` library param → `CITECHECK_NLI_MODEL` env var → built-in default
321
+ `cross-encoder/nli-deberta-v3-base`. Any Hugging Face cross-encoder name or
322
+ local path is accepted; citesure fails fast with a clear error (exit 2) if it
323
+ won't load — it never silently falls back to another model.
324
+
325
+ When NLI is on, the report header shows which model was used, and each scorable
326
+ verdict records its entailment score with `tier_reached=3`.
327
+
328
+ ## Cache & offline-first design
329
+
330
+ * **Fetch cache** — fetched pages are cached to disk for 24 h, keyed by
331
+ URL+etag, under `~/.cache/citesure/`. Override the location with
332
+ `CITECHECK_CACHE_DIR` (or the CLI `--cache-dir`). Re-verifying the same URLs
333
+ within 24 h makes no network calls.
334
+ * **Offline-first** — the default tiers-1+2 path needs no network beyond the
335
+ cited URLs themselves, no API keys, and no LLM. `file://` and bare local
336
+ paths let you verify entirely offline. The only thing that ever downloads is
337
+ the opt-in NLI model, and only on first `--nli` use.
338
+ * **No headless browser** — JavaScript-rendered pages yield no extractable
339
+ body, so their citations come back `ambiguous` rather than being silently
340
+ guessed at.
341
+
342
+ ## Development
343
+
344
+ ```bash
345
+ pip install -e ".[dev]" # installs pytest
346
+ pytest -q # offline suite (live + NLI-model tests deselected by default)
347
+ pytest -q -m live # run live-network smoke tests (hit the real internet)
348
+ pytest -q -m nli # run real DeBERTa-v3 model tests (needs the ~425 MB download)
349
+ ```
350
+
351
+ Live-network tests (`@pytest.mark.live`) and real-model NLI tests
352
+ (`@pytest.mark.nli`) are skipped by default so the suite is green offline.
@@ -0,0 +1,321 @@
1
+ <p align="center"><img src="https://raw.githubusercontent.com/DawnofGenX/citesure/main/docs/og-banner.png" alt="citesure" width="100%"/></p>
2
+
3
+ # citesure
4
+
5
+ An MCP server + CLI + library that verifies whether an LLM's claims are actually
6
+ supported by the sources it cites — resolves each citation, fetches the real
7
+ content, flags dead/retracted/paywalled links, and abstains when it can't
8
+ confirm.
9
+
10
+ citesure verifies *claims* (not just whether a link is alive). It needs no API
11
+ keys and no LLM: the fast path (tiers 1+2) is pure `httpx` + `trafilatura`, and
12
+ passing `--no-nli` keeps it fully offline. The NLI tier (tier 3) is on by
13
+ default and lazy-downloads a local cross-encoder on first use.
14
+
15
+ ## Accuracy
16
+
17
+ Measured on an independent 108 case set (40 unique URLs, 13 domains), built as a three-way
18
+ contrast: each source fact appears three times — verbatim-supported, negation-flipped, and
19
+ entity-swapped — all sharing one source sentence. The A-vs-B gap isolates polarity handling and
20
+ A-vs-C isolates subject binding, so a drop points at a specific bug class rather than "hard cases".
21
+
22
+ | Configuration | Agreement |
23
+ |---|---|
24
+ | NLI tier on (default claim path) | **83/108 = 76.9%** |
25
+ | NLI tier off (deterministic tiers only) | **41/108 = 38.0%** |
26
+
27
+ The 38.9-point gap is the cross-encoder tier's entire contribution, measured rather than asserted.
28
+
29
+ **Safety metrics** — the failure modes that matter for a verifier:
30
+
31
+ | Metric | Result |
32
+ |---|---|
33
+ | Negation-flip false-supported | 4/32 = 12.5% |
34
+ | Entity-swap false-supported | 7/32 = 21.9% |
35
+
36
+ Both rise to ~90-100% with the NLI tier off, which is the honest cost of the fast path.
37
+
38
+ An exhaustive threshold sweep (support 0.30-0.90 x ambiguous 0.05-support) moved agreement from
39
+ 66.7% to 69.4% — a gain of exactly one item. Threshold tuning is a dead end here; see
40
+ `evals/THRESHOLD_CALIBRATION.md`.
41
+
42
+ ## Install
43
+
44
+ Default install stays light — no torch, no model download:
45
+
46
+ ```bash
47
+ pip install citesure # tiers 1+2: httpx + trafilatura only
48
+ pip install "citesure[nli]" # adds the opt-in NLI cross-encoder tier
49
+ ```
50
+
51
+ Or from a checkout:
52
+
53
+ ```bash
54
+ git clone https://github.com/DawnofGenX/citesure.git && cd citesure
55
+ pip install . # default: offline, key-free, no torch
56
+ pip install ".[nli]" # adds the opt-in NLI tier
57
+ pip install -e ".[dev]" # editable + test deps
58
+ ```
59
+
60
+ `requires-python >= 3.10`. The default fast path (tiers 1+2) needs only
61
+ `httpx` + `trafilatura`. The opt-in NLI tier pulls in the local cross-encoder
62
+ stack (`torch`, `transformers`, `sentence-transformers`) and lazy-downloads a
63
+ ~425 MB model on first `--nli` use — see [NLI tier](#nli-tier-opt-in).
64
+
65
+ ## Quickstart
66
+
67
+ Create two small files in any directory — a claim document and the page it
68
+ cites:
69
+
70
+ `source.html`
71
+ ```html
72
+ <!DOCTYPE html>
73
+ <html>
74
+ <head><title>Python 3.12 release notes</title></head>
75
+ <body>
76
+ <h1>What's New In Python 3.12</h1>
77
+ <p>Python 3.12.0 was released on October 2, 2023. It introduces a new
78
+ interactive debugger, improved error messages, and faster startup times.
79
+ The new interactive debugger lets you step through code directly from the
80
+ REPL, and the interpreter now starts up roughly 5% faster than 3.11.</p>
81
+ </body>
82
+ </html>
83
+ ```
84
+
85
+ `notes.md`
86
+ ```markdown
87
+ # My notes
88
+
89
+ Python 3.12 was released on October 2, 2023, bringing a new interactive
90
+ debugger and faster startup times [1].
91
+
92
+ The moon is made of green cheese [2].
93
+
94
+ ## Sources
95
+
96
+ 1. source.html
97
+ 2. https://nonexistent-citesure-demo.invalid/green-cheese
98
+ ```
99
+
100
+ Then verify:
101
+
102
+ ```bash
103
+ citesure verify notes.md
104
+ ```
105
+
106
+ On the first run the NLI tier lazy-downloads a ~425 MB cross-encoder, so this
107
+ command needs network once. To stay fully offline, add `--no-nli` and use tiers
108
+ 1+2 only (link status + content overlap):
109
+
110
+ ```bash
111
+ citesure verify notes.md --no-nli
112
+ ```
113
+
114
+ Real output (run from a clean venv, zero API keys, ~2 s):
115
+
116
+ ```text
117
+ # citesure verification report — 2 citation(s) (input format: markdown)
118
+
119
+ - ✅ **[1]** PASS score=0.944 tier=2 — /tmp/demo/source.html
120
+ - evidence: Python 3.12.0 was released on October 2, 2023. It introduces a new interactive debugger, improved error messages, and faster startup times. The new interactive debugger lets you step through code directly from the REPL, and the interpreter now starts up roughly 5% faster than 3.11.
121
+ - note: page fetched; reachable with no red flags (tier 1)
122
+ - note: overlap tier: score 0.944 against fetched page; thresholds: >= 0.6 supported, >= 0.3 ambiguous, below unsupported
123
+ - ⛔ **[2]** UNVERIFIABLE tier=1 — https://nonexistent-citesure-demo.invalid/green-cheese
124
+ - note: ConnectError: [Errno -2] Name or service not known
125
+
126
+ ## Summary
127
+
128
+ - total: 2
129
+ - supported: 1
130
+ - unsupported: 0
131
+ - unreachable: 1
132
+ - paywalled: 0
133
+ - ambiguous: 0
134
+ - unverifiable (unreachable + paywalled + ambiguous): 1
135
+ - pass_rate: 0.5000
136
+
137
+ **Decision: FAIL** (pass_rate 0.5000 < threshold 0.8)
138
+ ```
139
+
140
+ Exit code is `1` because `pass_rate` (0.5) is below the default threshold
141
+ (0.8). Citation `[1]` is **supported** (its terms strongly overlap the cited
142
+ page); citation `[2]` is **unreachable** (the host doesn't resolve).
143
+
144
+ The repo also ships a larger bundled sample you can point the CLI at directly:
145
+
146
+ ```bash
147
+ citesure verify tests/fixtures/sample.md # 5 citations, all local
148
+ citesure verify tests/fixtures/sample.md --json # machine-readable JSON
149
+ citesure verify tests/fixtures/cases.json --json --strict
150
+ citesure verify tests/fixtures/sample.md --md report.md # also write a .md file
151
+ ```
152
+
153
+ ## How verification works — the 3-tier pipeline
154
+
155
+ Each citation is pushed through up to three tiers. A verdict records
156
+ `tier_reached` so you can see how far it got.
157
+
158
+ 1. **Reachability (tier 1)** — does the URL resolve and fetch? Is it dead,
159
+ retracted, or paywalled? Uses `httpx` + `trafilatura` (HTML → clean text),
160
+ 5 s timeout, 2 retries with backoff, max 10 concurrent fetches, a
161
+ citesure-identifying User-Agent, and robots.txt respect. `file://` and bare
162
+ local paths are supported for offline demos/tests. No headless browser.
163
+ 2. **Content overlap (tier 2)** — does the claim actually appear in the cited
164
+ region? The sentence/paragraph containing the `[n]` marker is the claim
165
+ unit (D4); it's matched against the top-k most relevant passages of the
166
+ fetched page using deterministic weighted term-coverage scoring in
167
+ `[0, 1]`. No LLM, no embeddings. Compound claims (joined by "and", "while",
168
+ "but") are split into clauses and scored independently — a claim with mixed
169
+ support (some clauses true, some false) returns `ambiguous`, not `supported`.
170
+ When all top-k passages score below 0.3, the fallback scores all passages.
171
+ 3. **NLI entailment (tier 3, opt-in)** — a local cross-encoder scores each
172
+ `(claim, best-passage)` pair for entailment. Off by default; enable with
173
+ `--nli`. See [NLI tier](#nli-tier-opt-in).
174
+
175
+ Tiers 1+2 are the default path and run entirely offline. Tier 3 is a
176
+ first-class opt-in, fully built and tested — not a stub.
177
+
178
+ ### Verdict statuses (D3)
179
+
180
+ | Status | Meaning |
181
+ |--------|---------|
182
+ | `supported` | Page fetched and claim terms strongly overlap the cited region (score ≥ 0.6). |
183
+ | `unsupported` | Page fetched but the claim is absent from the cited region (score < 0.3), or the page carries a retraction notice. |
184
+ | `unreachable` | DNS failure, HTTP 4xx/5xx, timeout, or robots.txt block. |
185
+ | `paywalled` | Login/paywall detected on the cited page. |
186
+ | `ambiguous` | Partial overlap (0.3 ≤ score < 0.6), or the citation marker could not be located (e.g. a JavaScript-rendered page with no extractable body). |
187
+
188
+ The overall report is `{total, supported, unsupported, unverifiable,
189
+ pass_rate}` where `unverifiable = unreachable + paywalled + ambiguous` and
190
+ `pass_rate = supported / total`.
191
+
192
+ ## CLI flags
193
+
194
+ | Flag | Meaning |
195
+ |------|---------|
196
+ | `--json` | Print the full report as JSON (`report.to_dict()`). |
197
+ | `--md FILE` | Also write the Markdown report to `FILE` (in addition to stdout). |
198
+ | `--threshold F` | Minimum `pass_rate` for exit code 0 (default `0.8`). |
199
+ | `--strict` | CI mode: `ambiguous` counts as an explicit failure (see exit codes). |
200
+ | `--nli` | Enable the NLI entailment tier (tier 3). Lazy-downloads the default model (~425 MB) on first use. |
201
+ | `--nli-model NAME` | Cross-encoder model name (HF id) or local path; highest priority in the model-swap chain. Implies `--nli`. |
202
+ | `--cache-dir DIR` | Override the fetch cache directory (sets `CITECHECK_CACHE_DIR`). |
203
+
204
+ ### Exit codes
205
+
206
+ * `0` — default mode: iff `pass_rate >= --threshold`. With `--strict`:
207
+ additionally requires **no** `ambiguous` AND **no** `unsupported` verdicts.
208
+ * `1` — verification ran but the rule above is not met.
209
+ * `2` — input error (unreadable file, malformed JSON) OR the NLI model failed
210
+ to load (fail fast, D6 — distinct from a verification-failure exit 1).
211
+
212
+ ## Library
213
+
214
+ ```python
215
+ import asyncio
216
+ from citesure import extract_citations, verify_citations, score_overlap
217
+
218
+ citations = extract_citations("The sky is blue [1].", {"1": "https://example.com"})
219
+ report = asyncio.run(verify_citations(citations)) # tiers 1+2 by default
220
+ print(report.pass_rate)
221
+
222
+ # Overlap tier in isolation (pure, deterministic):
223
+ score, snippet = score_overlap("Python 3.12 was released in 2023.", ["...page text..."])
224
+ ```
225
+
226
+ ## MCP server
227
+
228
+ `citesure-mcp` is a stdio MCP server (official `mcp` Python SDK) exposing two
229
+ tools:
230
+
231
+ * **`verify_citations(citations)`** — the structured JSON form: a list of
232
+ `{"claim": str, "citation": url-or-id}` objects (optional per-item
233
+ `"excerpt"`). Returns the D3 report: `{total, supported, unsupported,
234
+ unverifiable, pass_rate, verdicts[]}`.
235
+ * **`verify_markdown(markdown, url_map?)`** — raw markdown with inline
236
+ citations (`[n]` markers resolved via `url_map` or a trailing
237
+ `## Sources`/`## References` section, plus `[label](url)` links). Same
238
+ report shape.
239
+
240
+ Both tools run the full pipeline (reachability + content overlap) and return
241
+ clean JSON-serializable dicts; bad input comes back as `{"error": "..."}`
242
+ rather than a protocol error.
243
+
244
+ ### Adding it to an MCP client
245
+
246
+ Generic / Claude Desktop style config (any client that launches stdio MCP
247
+ servers):
248
+
249
+ ```json
250
+ {
251
+ "mcpServers": {
252
+ "citesure": {
253
+ "command": "citesure-mcp",
254
+ "args": []
255
+ }
256
+ }
257
+ }
258
+ ```
259
+
260
+ If the package isn't on your `PATH`, use the absolute path to the console
261
+ script instead, e.g. `"command": "/path/to/venv/bin/citesure-mcp"`.
262
+
263
+ To enable the NLI tier at server start, set the env vars:
264
+
265
+ ```json
266
+ {
267
+ "mcpServers": {
268
+ "citesure": {
269
+ "command": "citesure-mcp",
270
+ "env": {
271
+ "CITECHECK_NLI": "1",
272
+ "CITECHECK_NLI_MODEL": "cross-encoder/nli-deberta-v3-base"
273
+ }
274
+ }
275
+ }
276
+ }
277
+ ```
278
+
279
+ `CITECHECK_NLI_MODEL` is optional (the built-in default applies); the model is
280
+ lazy-downloaded on first use (~425 MB) into `~/.cache/citesure/`.
281
+
282
+ ## NLI tier (opt-in)
283
+
284
+ NLI entailment (tier 3) is **off by default** — the fast tiers-1+2 path runs
285
+ with no model download. Enable it via the CLI (`--nli`) or, for the MCP
286
+ server, via `CITECHECK_NLI=1` at server start.
287
+
288
+ Model selection follows a priority chain (D6): `--nli-model NAME` flag /
289
+ `nli_model=` library param → `CITECHECK_NLI_MODEL` env var → built-in default
290
+ `cross-encoder/nli-deberta-v3-base`. Any Hugging Face cross-encoder name or
291
+ local path is accepted; citesure fails fast with a clear error (exit 2) if it
292
+ won't load — it never silently falls back to another model.
293
+
294
+ When NLI is on, the report header shows which model was used, and each scorable
295
+ verdict records its entailment score with `tier_reached=3`.
296
+
297
+ ## Cache & offline-first design
298
+
299
+ * **Fetch cache** — fetched pages are cached to disk for 24 h, keyed by
300
+ URL+etag, under `~/.cache/citesure/`. Override the location with
301
+ `CITECHECK_CACHE_DIR` (or the CLI `--cache-dir`). Re-verifying the same URLs
302
+ within 24 h makes no network calls.
303
+ * **Offline-first** — the default tiers-1+2 path needs no network beyond the
304
+ cited URLs themselves, no API keys, and no LLM. `file://` and bare local
305
+ paths let you verify entirely offline. The only thing that ever downloads is
306
+ the opt-in NLI model, and only on first `--nli` use.
307
+ * **No headless browser** — JavaScript-rendered pages yield no extractable
308
+ body, so their citations come back `ambiguous` rather than being silently
309
+ guessed at.
310
+
311
+ ## Development
312
+
313
+ ```bash
314
+ pip install -e ".[dev]" # installs pytest
315
+ pytest -q # offline suite (live + NLI-model tests deselected by default)
316
+ pytest -q -m live # run live-network smoke tests (hit the real internet)
317
+ pytest -q -m nli # run real DeBERTa-v3 model tests (needs the ~425 MB download)
318
+ ```
319
+
320
+ Live-network tests (`@pytest.mark.live`) and real-model NLI tests
321
+ (`@pytest.mark.nli`) are skipped by default so the suite is green offline.