stillvalid 0.1.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (34) hide show
  1. stillvalid-0.1.0/LICENSE +21 -0
  2. stillvalid-0.1.0/MANIFEST.in +5 -0
  3. stillvalid-0.1.0/PKG-INFO +369 -0
  4. stillvalid-0.1.0/README.md +347 -0
  5. stillvalid-0.1.0/docs/VERIFICATION.md +307 -0
  6. stillvalid-0.1.0/examples/demo.py +163 -0
  7. stillvalid-0.1.0/pyproject.toml +38 -0
  8. stillvalid-0.1.0/setup.cfg +4 -0
  9. stillvalid-0.1.0/stillvalid/__init__.py +24 -0
  10. stillvalid-0.1.0/stillvalid/backfill.py +146 -0
  11. stillvalid-0.1.0/stillvalid/cascade.py +170 -0
  12. stillvalid-0.1.0/stillvalid/doc.py +27 -0
  13. stillvalid-0.1.0/stillvalid/history.py +142 -0
  14. stillvalid-0.1.0/stillvalid/integrations/__init__.py +5 -0
  15. stillvalid-0.1.0/stillvalid/integrations/common.py +185 -0
  16. stillvalid-0.1.0/stillvalid/integrations/langchain.py +86 -0
  17. stillvalid-0.1.0/stillvalid/integrations/llamaindex.py +81 -0
  18. stillvalid-0.1.0/stillvalid/mcp_server.py +247 -0
  19. stillvalid-0.1.0/stillvalid/probe.py +250 -0
  20. stillvalid-0.1.0/stillvalid/py.typed +0 -0
  21. stillvalid-0.1.0/stillvalid/refresh.py +124 -0
  22. stillvalid-0.1.0/stillvalid/survival.py +71 -0
  23. stillvalid-0.1.0/stillvalid/verdict.py +96 -0
  24. stillvalid-0.1.0/stillvalid.egg-info/PKG-INFO +369 -0
  25. stillvalid-0.1.0/stillvalid.egg-info/SOURCES.txt +32 -0
  26. stillvalid-0.1.0/stillvalid.egg-info/dependency_links.txt +1 -0
  27. stillvalid-0.1.0/stillvalid.egg-info/entry_points.txt +2 -0
  28. stillvalid-0.1.0/stillvalid.egg-info/requires.txt +9 -0
  29. stillvalid-0.1.0/stillvalid.egg-info/top_level.txt +1 -0
  30. stillvalid-0.1.0/tests/test_backfill.py +159 -0
  31. stillvalid-0.1.0/tests/test_integrations.py +271 -0
  32. stillvalid-0.1.0/tests/test_probe.py +225 -0
  33. stillvalid-0.1.0/tests/test_refresh.py +117 -0
  34. stillvalid-0.1.0/tests/test_stillvalid.py +213 -0
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 Junhyeok Choe
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
@@ -0,0 +1,5 @@
1
+ include LICENSE README.md
2
+ recursive-include tests *.py
3
+ recursive-include examples *.py
4
+ recursive-include docs *.md
5
+ global-exclude *.db *.db-wal *.db-shm __pycache__/*
@@ -0,0 +1,369 @@
1
+ Metadata-Version: 2.4
2
+ Name: stillvalid
3
+ Version: 0.1.0
4
+ Summary: Is this information still valid? A validity layer between retrieval and action — deterministic checks first, learned survival probability where nothing deterministic exists.
5
+ License: MIT
6
+ Project-URL: Homepage, https://github.com/junniec01-creator/stillvalid
7
+ Keywords: freshness,staleness,rag,agents,validity,cache,survival-analysis
8
+ Classifier: Development Status :: 3 - Alpha
9
+ Classifier: Intended Audience :: Developers
10
+ Classifier: Programming Language :: Python :: 3
11
+ Classifier: Topic :: Software Development :: Libraries
12
+ Requires-Python: >=3.10
13
+ Description-Content-Type: text/markdown
14
+ License-File: LICENSE
15
+ Provides-Extra: mcp
16
+ Requires-Dist: mcp>=1.0; extra == "mcp"
17
+ Provides-Extra: langchain
18
+ Requires-Dist: langchain-core>=0.3; extra == "langchain"
19
+ Provides-Extra: llamaindex
20
+ Requires-Dist: llama-index-core>=0.11; extra == "llamaindex"
21
+ Dynamic: license-file
22
+
23
+ # stillvalid
24
+
25
+ **Is this information still valid?**
26
+
27
+ A validity layer between retrieval and action. Your RAG pipeline or agent
28
+ retrieved a document — `stillvalid` tells you whether to trust it, verify
29
+ it, or refresh it, *before* the LLM acts on stale facts.
30
+
31
+ > Humans pause when they notice "this doc is from 2023". Agents don't —
32
+ > they read stale pricing and generate the quote. A tribunal has already
33
+ > held a company liable for its chatbot citing an outdated policy
34
+ > (*Moffatt v Air Canada*, 2024). Staleness is becoming an agent-safety
35
+ > problem, not a search-quality nit.
36
+
37
+ ```python
38
+ from stillvalid import Checker, Doc
39
+
40
+ sv = Checker(history_db="observations.db")
41
+
42
+ verdicts = sv.check_many(
43
+ Doc(id=d.id, last_verified_at=d.indexed_at, last_changed_at=d.modified_at)
44
+ for d in retrieved_docs
45
+ )
46
+ for d, v in zip(retrieved_docs, verdicts):
47
+ if not v.usable: # VERIFY / STALE / UNKNOWN
48
+ d = refetch(d) # re-check only the risky evidence
49
+ ```
50
+
51
+ Standard library only. Zero dependencies, zero network calls, microseconds
52
+ per check. **Experimental (v0.1) — API may change.**
53
+
54
+ ## 30-second demo
55
+
56
+ ```
57
+ python examples/demo.py
58
+ ```
59
+
60
+ A support agent indexed its knowledge base 45 days ago; the refund policy
61
+ changed upstream 10 days ago. Without a validity layer the agent quotes
62
+ the old fee *and* an expired promo. With it:
63
+
64
+ ```
65
+ document state layer reason
66
+ policy/refund STALE modified source was modified 35d after your last verification
67
+ promo/summer STALE expiry explicitly expired 5d ago
68
+ policy/baggage VALID hash live content hash equals verified hash
69
+ fees/schedule STALE survival calibrated 21% probability it is unchanged 45d after verification
70
+ guide/visa LIKELY_VALID survival calibrated 94% probability it is unchanged 45d after verification
71
+
72
+ Agent re-fetches only the risky evidence: refund, promo, fees
73
+ → 2/5 documents served from cache (no re-fetch cost);
74
+ 3/5 re-verified — the wrong answer never left the building.
75
+ ```
76
+
77
+ Every verdict also carries a one-sentence `summary` built for agents to
78
+ relay ("'policy/refund' is STALE — source was modified 35d after your
79
+ last verification. Re-fetch the source before quoting it.") — so in MCP
80
+ use, the check narrates itself in the conversation.
81
+
82
+ Note the last two rows: both are "45 days since verification", but the fee
83
+ table's learned update habit says *stale* while the visa guide's says
84
+ *still fine* — that distinction is the whole point, and no timestamp
85
+ cutoff can make it.
86
+
87
+ ## How it decides: a cascade
88
+
89
+ Cheap, certain judgments first; stop at the first layer that can decide;
90
+ abstain honestly when none can.
91
+
92
+ | # | Layer | Needs | Verdict type |
93
+ |---|-------|-------|--------------|
94
+ | 1 | Explicit expiry (`expires_at`) | nothing | deterministic |
95
+ | 2 | Live content-hash compare | a hash you just fetched | deterministic |
96
+ | 3 | Live Last-Modified compare | a timestamp you just fetched | deterministic |
97
+ | 6 | **Calibrated survival model** | gate-passed params for this doc | probability, `calibrated=True` |
98
+ | 5 | Crude change-rate estimate | ≥ 2 observed changes in local history | probability, `calibrated=False` |
99
+ | 7 | `UNKNOWN` | — | honest abstention |
100
+
101
+ Two design rules worth knowing:
102
+
103
+ - **If something deterministic exists, probability never runs.** When you
104
+ can simply look the answer up, a model is the wrong tool. The learned
105
+ layers exist for the documents where no lookup can answer — sources
106
+ without feeds, and the question "will it *still* be valid when I act?",
107
+ which no feed can answer.
108
+ - **Uncalibrated numbers say so.** Layer 5 is a memoryless estimate that
109
+ exists so the cascade is useful on day one; its verdicts carry
110
+ `calibrated=False`. Layer 6 probabilities come from a survival model
111
+ that must pass a serving gate (discrimination + ablation + calibration
112
+ checks) before a single probability is published. Documents that
113
+ haven't earned a calibrated number get `UNKNOWN`, not a guess.
114
+
115
+ ## URL sources work out of the box
116
+
117
+ For anything with a URL, the deterministic layers don't need you to wire
118
+ up anything — one conditional HEAD request (no body, no re-embedding)
119
+ fetches the signals:
120
+
121
+ ```python
122
+ from stillvalid.probe import probe, doc_from_probe
123
+
124
+ r = probe(url, etag=indexed_etag, if_modified_since=indexed_at)
125
+ verdict = sv.check(doc_from_probe(doc_id, r,
126
+ last_verified_at=indexed_at,
127
+ verified_etag=indexed_etag))
128
+ ```
129
+
130
+ A `304 Not Modified` is the cheapest deterministic VALID there is; a
131
+ changed ETag or newer `Last-Modified` is a deterministic STALE. Probe
132
+ failures degrade to UNKNOWN — a network error is not evidence of
133
+ staleness. `probe_many([...])` does batches concurrently.
134
+
135
+ In a 25-site survey (docs, APIs, dynamic pages), **74% exposed an ETag or
136
+ `Last-Modified`** — so the deterministic layers carry roughly three out of
137
+ four real-world URLs. Two field notes from that survey:
138
+
139
+ - **Use ETags, not body hashes.** Every dynamic page we measured changed
140
+ its bytes on *every* request (CSRF tokens, analytics ids) while its
141
+ content was identical. Feeding a whole-body hash into `verified_hash`
142
+ would manufacture constant false STALE; the server's own validator does
143
+ not have that problem.
144
+ - **Watch `result.age`.** Validators can arrive from a CDN cache — we saw
145
+ one up to 5.3 hours old, and `Cache-Control: no-cache` does not defeat
146
+ it (CDNs ignore it). A cached validator describes the origin as of
147
+ `age` seconds ago, not now. Pass `doc_from_probe(..., max_age_s=…)` to
148
+ drop validators older than you are willing to trust; they then fall
149
+ through to the probabilistic layers or UNKNOWN.
150
+
151
+ ### Probing is locked down by default
152
+
153
+ URLs usually arrive from a retriever — that is, from *data* — so the
154
+ probe treats them as untrusted:
155
+
156
+ - **http/https only.** Left unguarded, `urllib` happily serves `file://`
157
+ and `ftp://`, which would turn "check this document" into local file
158
+ disclosure.
159
+ - **Public addresses only.** Loopback, private ranges, link-local
160
+ (including cloud metadata at `169.254.169.254`), reserved and multicast
161
+ targets are refused.
162
+ - **Redirects re-checked at every hop**, so a public URL answering `302`
163
+ with an internal `Location` does not become an escape hatch.
164
+ - **No credentials are ever attached** — no cookies, no auth headers.
165
+
166
+ Internal sources are a legitimate case (a wiki on `10.x` is normal), so
167
+ they are opt-in, not silently allowed:
168
+
169
+ ```python
170
+ probe(url, allow_private=True) # I mean it, this host is internal
171
+ ```
172
+
173
+ Known limit: the host is resolved once for the check and again by
174
+ `urllib` for the connection, so DNS rebinding is not covered — pin your
175
+ own resolver if that is in your threat model.
176
+
177
+ **What gets stored**: the observation log keeps only `doc_id`, a
178
+ timestamp, and the content hash you passed. `doc_id` is stored verbatim,
179
+ so if you pass raw URLs, anything embedded in them (tokens, credentials,
180
+ internal paths) is stored too. Pass an opaque id if that matters.
181
+
182
+ ## Import the past instead of waiting for it
183
+
184
+ The usual deal with a freshness layer is "install it, come back in a
185
+ month". Skip that: most knowledge bases already keep their own change
186
+ history, and reading it is a one-liner.
187
+
188
+ ```python
189
+ from stillvalid.backfill import from_git, from_csv, from_rows
190
+
191
+ from_git(sv, r"C:\work\handbook") # a docs repo: commits are the log
192
+ from_csv(sv, "wiki_revisions.csv") # doc_id, ts[, rev]
193
+ from_rows(sv, confluence_revision_rows) # anything you can iterate
194
+ ```
195
+
196
+ Measured on three real repositories here: **3,228 past changes across
197
+ 1,089 documents, imported in 1.7 seconds.** Before the import, a sample of
198
+ those documents judged `UNKNOWN` across the board; afterwards the same
199
+ query returned real verdicts with probabilities spread from 0% to 88%.
200
+ That is the difference between "ask me again in a month" and "ask me now".
201
+
202
+ Git is the easiest case, but the general entry point is `from_rows` —
203
+ feed it a wiki's revision API, a CMS audit table, a warehouse query of
204
+ `updated_at` snapshots. Consecutive identical revisions read as "still the
205
+ same", which is exactly the censoring signal the learned layers want.
206
+
207
+ `summarize(sv)` tells you what the import bought: how many documents can
208
+ now be judged on update behavior, and how many have enough events (8+) to
209
+ train a calibrated model on.
210
+
211
+ ## It learns as you use it
212
+
213
+ The LangChain / LlamaIndex wrappers **auto-record** a content-hash
214
+ sighting for every document that flows through them (same-hash sightings
215
+ are deduped to one per hour; disable with `record=False`). Or log
216
+ manually whenever you fetch or reindex:
217
+
218
+ ```python
219
+ sv.record(doc_id, content_hash)
220
+ ```
221
+
222
+ Day 1: layers 1–3 and 7 carry the load. Within days, layer 5 starts
223
+ estimating from your observation log. With enough history (8+ observed
224
+ changes per document), you can train the survival model and export
225
+ parameters — then "updated 7 days ago" upgrades to "97% likely still
226
+ valid", calibrated.
227
+
228
+ One honesty note on auto-recording: it observes *your index's copy*, so
229
+ a change becomes visible at reindex time, not at source-change time.
230
+ That is exactly the fidelity the crude layer claims — calibrated
231
+ training should prefer source-side timestamps where available.
232
+
233
+ The trainer (survival analysis over your observation log, with the
234
+ serving gate) lives in the parent project; `stillvalid` only needs its
235
+ exported JSON (`fpi-qtime-params/1`).
236
+
237
+ ## Cost
238
+
239
+ Measured on one laptop (Windows, Python 3.12) — order of magnitude, not a
240
+ benchmark claim:
241
+
242
+ | Operation | Cost |
243
+ |---|---|
244
+ | Verdict, deterministic layers | ~3 µs |
245
+ | Verdict, change-rate layer | ~0.5 ms, **constant** in history size |
246
+ | Batch of 1,000 documents | ~7 ms |
247
+ | `live=True`, 50 local files | ~0.9 ms |
248
+ | `live=True`, remote URLs (batched) | **~300 ms**, one round trip |
249
+ | Recording a sighting | ~0.4 ms (deduped: ~0.02 ms) |
250
+ | Observation log on disk | ~52 bytes/row |
251
+
252
+ Two things worth planning around:
253
+
254
+ - **Remote `live=True` costs a network round trip** (~300 ms, flat for a
255
+ batch since probes run concurrently). That is fine for an agent deciding
256
+ whether to act, and probably too slow in a latency-critical search path —
257
+ there, prefer local paths, or refresh signals on a schedule rather than
258
+ per query.
259
+ - **The change-rate layer only reads the most recent sightings**
260
+ (`History.RECENT_WINDOW`, 500). Reading full history made judging cost
261
+ grow with it; recent behavior also predicts better, so the window is
262
+ both faster and more honest.
263
+
264
+ ## Verdicts
265
+
266
+ | State | Meaning | Default action |
267
+ |-------|---------|----------------|
268
+ | `VALID` | deterministically confirmed current | `use` |
269
+ | `LIKELY_VALID` | probability ≥ use-threshold (default 0.90) | `use` |
270
+ | `VERIFY` | uncertain — recheck the source | `verify` |
271
+ | `STALE` | expired / changed / probably invalid | `refresh` |
272
+ | `UNKNOWN` | no layer could judge | `verify` |
273
+
274
+ Thresholds are yours to set: `Checker(policy=Policy(use_threshold=0.95,
275
+ stale_threshold=0.70))`. Every verdict carries the deciding layer, a
276
+ human-readable reason, and the probability when one exists — evidence,
277
+ not just a verdict.
278
+
279
+ ## Why thresholds on a calibrated probability (and not doc age)?
280
+
281
+ Because age can rank documents, but it can't answer "is 30 days old fine
282
+ *for this document*?" — a HQ address and a stock level are both "30 days
283
+ old" and deserve opposite treatment. In our backtests on a public-document
284
+ corpus (300k queries, weekly-reindex scenario), threshold decisions on
285
+ calibrated probabilities left 2.4–4.8× less stale exposure than an
286
+ age-cutoff heuristic at the same re-verification budget.
287
+
288
+ ## Integrations
289
+
290
+ **LangChain** — judge every retrieved document between retrieval and the
291
+ LLM (`pip install "stillvalid[langchain]"`):
292
+
293
+ ```python
294
+ # LangChain >= 1.0 moved this retriever into langchain_classic
295
+ from langchain_classic.retrievers import ContextualCompressionRetriever
296
+ from stillvalid import Checker
297
+ from stillvalid.integrations.langchain import StillValidCompressor
298
+
299
+ retriever = ContextualCompressionRetriever(
300
+ base_compressor=StillValidCompressor(checker=Checker(...), live=True),
301
+ base_retriever=vectorstore.as_retriever(),
302
+ )
303
+ # each doc gains metadata["stillvalid"] = {state, action, layer, reason, ...}
304
+ ```
305
+
306
+ **LlamaIndex** — same idea as a node postprocessor
307
+ (`pip install "stillvalid[llamaindex]"`):
308
+
309
+ ```python
310
+ from stillvalid.integrations.llamaindex import StillValidPostprocessor
311
+
312
+ engine = index.as_query_engine(
313
+ node_postprocessors=[StillValidPostprocessor(checker=Checker(...))])
314
+ ```
315
+
316
+ Both read common metadata keys out of the box (`indexed_at`,
317
+ `last_modified`, `updated_at`, `content_hash`, LlamaIndex's
318
+ `last_modified_date`, …) and `stillvalid_*` prefixed keys always win.
319
+ `mode="filter"` drops non-usable documents but keeps `UNKNOWN` by
320
+ default — absence of evidence is not evidence of staleness
321
+ (`drop_unknown=True` if you disagree).
322
+
323
+ **Pass `live=True` unless you have a reason not to.** Index metadata is a
324
+ *past* snapshot by definition — it says when you last looked, never
325
+ whether the source has moved since — so without a current signal a fresh
326
+ install can only answer UNKNOWN. `live=True` fetches that signal at query
327
+ time, cheaply: `os.stat()` for local paths, a batched conditional HEAD for
328
+ URLs (never the body). In our end-to-end run over a real
329
+ loader → vector store → retriever pipeline, this turned *every* verdict
330
+ from UNKNOWN into a real one and flagged exactly the file that had been
331
+ edited after indexing — at **2 ms** added latency for local files.
332
+ Network sources cost one round trip; tune with `live_timeout`, and set
333
+ `allow_private=True` if your wiki is on an internal address.
334
+
335
+ **MCP server** — let the agent itself ask
336
+ (`pip install "stillvalid[mcp]"`, supports mcp 1.x and 2.x):
337
+
338
+ ```jsonc
339
+ // Claude Desktop / any MCP client
340
+ {"mcpServers": {"stillvalid": {
341
+ "command": "python", "args": ["-m", "stillvalid.mcp_server"],
342
+ "env": {"STILLVALID_HISTORY_DB": "C:/data/observations.db"}}}}
343
+ ```
344
+
345
+ Tools: `import_history` (load a corpus's past in seconds — run this
346
+ first), `check_validity` (judge one piece of evidence before acting on
347
+ it), `check_validity_batch` (shared clock), `record_observation` (log a
348
+ sighting so the cascade keeps learning).
349
+
350
+ In an agent session that looks like: *"point me at your docs repo"* →
351
+ `import_history` → *"3,228 past changes imported; 554 of 1,089 documents
352
+ can now be judged on their update behavior"* → every later answer is
353
+ checked against that.
354
+
355
+ ## Roadmap
356
+
357
+ - Cross-source verification layer (layer 4 of the design)
358
+ - Cold-start priors: borrowing update-behavior from similar sources,
359
+ gated against negative transfer
360
+ - `pip install stillvalid` (PyPI release)
361
+
362
+ ## Status & honesty
363
+
364
+ This is an extraction of a working research pipeline (survival analysis
365
+ over change histories, with calibration verification) into a standalone
366
+ layer. The calibrated path is real but requires training on your
367
+ observation log; everything else works out of the box. If your documents
368
+ all have reliable feeds or explicit expiries, you don't need the learned
369
+ layers — and this library will happily tell you so by never reaching them.