stillvalid 0.1.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- stillvalid-0.1.0/LICENSE +21 -0
- stillvalid-0.1.0/MANIFEST.in +5 -0
- stillvalid-0.1.0/PKG-INFO +369 -0
- stillvalid-0.1.0/README.md +347 -0
- stillvalid-0.1.0/docs/VERIFICATION.md +307 -0
- stillvalid-0.1.0/examples/demo.py +163 -0
- stillvalid-0.1.0/pyproject.toml +38 -0
- stillvalid-0.1.0/setup.cfg +4 -0
- stillvalid-0.1.0/stillvalid/__init__.py +24 -0
- stillvalid-0.1.0/stillvalid/backfill.py +146 -0
- stillvalid-0.1.0/stillvalid/cascade.py +170 -0
- stillvalid-0.1.0/stillvalid/doc.py +27 -0
- stillvalid-0.1.0/stillvalid/history.py +142 -0
- stillvalid-0.1.0/stillvalid/integrations/__init__.py +5 -0
- stillvalid-0.1.0/stillvalid/integrations/common.py +185 -0
- stillvalid-0.1.0/stillvalid/integrations/langchain.py +86 -0
- stillvalid-0.1.0/stillvalid/integrations/llamaindex.py +81 -0
- stillvalid-0.1.0/stillvalid/mcp_server.py +247 -0
- stillvalid-0.1.0/stillvalid/probe.py +250 -0
- stillvalid-0.1.0/stillvalid/py.typed +0 -0
- stillvalid-0.1.0/stillvalid/refresh.py +124 -0
- stillvalid-0.1.0/stillvalid/survival.py +71 -0
- stillvalid-0.1.0/stillvalid/verdict.py +96 -0
- stillvalid-0.1.0/stillvalid.egg-info/PKG-INFO +369 -0
- stillvalid-0.1.0/stillvalid.egg-info/SOURCES.txt +32 -0
- stillvalid-0.1.0/stillvalid.egg-info/dependency_links.txt +1 -0
- stillvalid-0.1.0/stillvalid.egg-info/entry_points.txt +2 -0
- stillvalid-0.1.0/stillvalid.egg-info/requires.txt +9 -0
- stillvalid-0.1.0/stillvalid.egg-info/top_level.txt +1 -0
- stillvalid-0.1.0/tests/test_backfill.py +159 -0
- stillvalid-0.1.0/tests/test_integrations.py +271 -0
- stillvalid-0.1.0/tests/test_probe.py +225 -0
- stillvalid-0.1.0/tests/test_refresh.py +117 -0
- stillvalid-0.1.0/tests/test_stillvalid.py +213 -0
stillvalid-0.1.0/LICENSE
ADDED
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 Junhyeok Choe
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|
|
@@ -0,0 +1,369 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: stillvalid
|
|
3
|
+
Version: 0.1.0
|
|
4
|
+
Summary: Is this information still valid? A validity layer between retrieval and action — deterministic checks first, learned survival probability where nothing deterministic exists.
|
|
5
|
+
License: MIT
|
|
6
|
+
Project-URL: Homepage, https://github.com/junniec01-creator/stillvalid
|
|
7
|
+
Keywords: freshness,staleness,rag,agents,validity,cache,survival-analysis
|
|
8
|
+
Classifier: Development Status :: 3 - Alpha
|
|
9
|
+
Classifier: Intended Audience :: Developers
|
|
10
|
+
Classifier: Programming Language :: Python :: 3
|
|
11
|
+
Classifier: Topic :: Software Development :: Libraries
|
|
12
|
+
Requires-Python: >=3.10
|
|
13
|
+
Description-Content-Type: text/markdown
|
|
14
|
+
License-File: LICENSE
|
|
15
|
+
Provides-Extra: mcp
|
|
16
|
+
Requires-Dist: mcp>=1.0; extra == "mcp"
|
|
17
|
+
Provides-Extra: langchain
|
|
18
|
+
Requires-Dist: langchain-core>=0.3; extra == "langchain"
|
|
19
|
+
Provides-Extra: llamaindex
|
|
20
|
+
Requires-Dist: llama-index-core>=0.11; extra == "llamaindex"
|
|
21
|
+
Dynamic: license-file
|
|
22
|
+
|
|
23
|
+
# stillvalid
|
|
24
|
+
|
|
25
|
+
**Is this information still valid?**
|
|
26
|
+
|
|
27
|
+
A validity layer between retrieval and action. Your RAG pipeline or agent
|
|
28
|
+
retrieved a document — `stillvalid` tells you whether to trust it, verify
|
|
29
|
+
it, or refresh it, *before* the LLM acts on stale facts.
|
|
30
|
+
|
|
31
|
+
> Humans pause when they notice "this doc is from 2023". Agents don't —
|
|
32
|
+
> they read stale pricing and generate the quote. A tribunal has already
|
|
33
|
+
> held a company liable for its chatbot citing an outdated policy
|
|
34
|
+
> (*Moffatt v Air Canada*, 2024). Staleness is becoming an agent-safety
|
|
35
|
+
> problem, not a search-quality nit.
|
|
36
|
+
|
|
37
|
+
```python
|
|
38
|
+
from stillvalid import Checker, Doc
|
|
39
|
+
|
|
40
|
+
sv = Checker(history_db="observations.db")
|
|
41
|
+
|
|
42
|
+
verdicts = sv.check_many(
|
|
43
|
+
Doc(id=d.id, last_verified_at=d.indexed_at, last_changed_at=d.modified_at)
|
|
44
|
+
for d in retrieved_docs
|
|
45
|
+
)
|
|
46
|
+
for d, v in zip(retrieved_docs, verdicts):
|
|
47
|
+
if not v.usable: # VERIFY / STALE / UNKNOWN
|
|
48
|
+
d = refetch(d) # re-check only the risky evidence
|
|
49
|
+
```
|
|
50
|
+
|
|
51
|
+
Standard library only. Zero dependencies, zero network calls, microseconds
|
|
52
|
+
per check. **Experimental (v0.1) — API may change.**
|
|
53
|
+
|
|
54
|
+
## 30-second demo
|
|
55
|
+
|
|
56
|
+
```
|
|
57
|
+
python examples/demo.py
|
|
58
|
+
```
|
|
59
|
+
|
|
60
|
+
A support agent indexed its knowledge base 45 days ago; the refund policy
|
|
61
|
+
changed upstream 10 days ago. Without a validity layer the agent quotes
|
|
62
|
+
the old fee *and* an expired promo. With it:
|
|
63
|
+
|
|
64
|
+
```
|
|
65
|
+
document state layer reason
|
|
66
|
+
policy/refund STALE modified source was modified 35d after your last verification
|
|
67
|
+
promo/summer STALE expiry explicitly expired 5d ago
|
|
68
|
+
policy/baggage VALID hash live content hash equals verified hash
|
|
69
|
+
fees/schedule STALE survival calibrated 21% probability it is unchanged 45d after verification
|
|
70
|
+
guide/visa LIKELY_VALID survival calibrated 94% probability it is unchanged 45d after verification
|
|
71
|
+
|
|
72
|
+
Agent re-fetches only the risky evidence: refund, promo, fees
|
|
73
|
+
→ 2/5 documents served from cache (no re-fetch cost);
|
|
74
|
+
3/5 re-verified — the wrong answer never left the building.
|
|
75
|
+
```
|
|
76
|
+
|
|
77
|
+
Every verdict also carries a one-sentence `summary` built for agents to
|
|
78
|
+
relay ("'policy/refund' is STALE — source was modified 35d after your
|
|
79
|
+
last verification. Re-fetch the source before quoting it.") — so in MCP
|
|
80
|
+
use, the check narrates itself in the conversation.
|
|
81
|
+
|
|
82
|
+
Note the last two rows: both are "45 days since verification", but the fee
|
|
83
|
+
table's learned update habit says *stale* while the visa guide's says
|
|
84
|
+
*still fine* — that distinction is the whole point, and no timestamp
|
|
85
|
+
cutoff can make it.
|
|
86
|
+
|
|
87
|
+
## How it decides: a cascade
|
|
88
|
+
|
|
89
|
+
Cheap, certain judgments first; stop at the first layer that can decide;
|
|
90
|
+
abstain honestly when none can.
|
|
91
|
+
|
|
92
|
+
| # | Layer | Needs | Verdict type |
|
|
93
|
+
|---|-------|-------|--------------|
|
|
94
|
+
| 1 | Explicit expiry (`expires_at`) | nothing | deterministic |
|
|
95
|
+
| 2 | Live content-hash compare | a hash you just fetched | deterministic |
|
|
96
|
+
| 3 | Live Last-Modified compare | a timestamp you just fetched | deterministic |
|
|
97
|
+
| 6 | **Calibrated survival model** | gate-passed params for this doc | probability, `calibrated=True` |
|
|
98
|
+
| 5 | Crude change-rate estimate | ≥ 2 observed changes in local history | probability, `calibrated=False` |
|
|
99
|
+
| 7 | `UNKNOWN` | — | honest abstention |
|
|
100
|
+
|
|
101
|
+
Two design rules worth knowing:
|
|
102
|
+
|
|
103
|
+
- **If something deterministic exists, probability never runs.** When you
|
|
104
|
+
can simply look the answer up, a model is the wrong tool. The learned
|
|
105
|
+
layers exist for the documents where no lookup can answer — sources
|
|
106
|
+
without feeds, and the question "will it *still* be valid when I act?",
|
|
107
|
+
which no feed can answer.
|
|
108
|
+
- **Uncalibrated numbers say so.** Layer 5 is a memoryless estimate that
|
|
109
|
+
exists so the cascade is useful on day one; its verdicts carry
|
|
110
|
+
`calibrated=False`. Layer 6 probabilities come from a survival model
|
|
111
|
+
that must pass a serving gate (discrimination + ablation + calibration
|
|
112
|
+
checks) before a single probability is published. Documents that
|
|
113
|
+
haven't earned a calibrated number get `UNKNOWN`, not a guess.
|
|
114
|
+
|
|
115
|
+
## URL sources work out of the box
|
|
116
|
+
|
|
117
|
+
For anything with a URL, the deterministic layers don't need you to wire
|
|
118
|
+
up anything — one conditional HEAD request (no body, no re-embedding)
|
|
119
|
+
fetches the signals:
|
|
120
|
+
|
|
121
|
+
```python
|
|
122
|
+
from stillvalid.probe import probe, doc_from_probe
|
|
123
|
+
|
|
124
|
+
r = probe(url, etag=indexed_etag, if_modified_since=indexed_at)
|
|
125
|
+
verdict = sv.check(doc_from_probe(doc_id, r,
|
|
126
|
+
last_verified_at=indexed_at,
|
|
127
|
+
verified_etag=indexed_etag))
|
|
128
|
+
```
|
|
129
|
+
|
|
130
|
+
A `304 Not Modified` is the cheapest deterministic VALID there is; a
|
|
131
|
+
changed ETag or newer `Last-Modified` is a deterministic STALE. Probe
|
|
132
|
+
failures degrade to UNKNOWN — a network error is not evidence of
|
|
133
|
+
staleness. `probe_many([...])` does batches concurrently.
|
|
134
|
+
|
|
135
|
+
In a 25-site survey (docs, APIs, dynamic pages), **74% exposed an ETag or
|
|
136
|
+
`Last-Modified`** — so the deterministic layers carry roughly three out of
|
|
137
|
+
four real-world URLs. Two field notes from that survey:
|
|
138
|
+
|
|
139
|
+
- **Use ETags, not body hashes.** Every dynamic page we measured changed
|
|
140
|
+
its bytes on *every* request (CSRF tokens, analytics ids) while its
|
|
141
|
+
content was identical. Feeding a whole-body hash into `verified_hash`
|
|
142
|
+
would manufacture constant false STALE; the server's own validator does
|
|
143
|
+
not have that problem.
|
|
144
|
+
- **Watch `result.age`.** Validators can arrive from a CDN cache — we saw
|
|
145
|
+
one up to 5.3 hours old, and `Cache-Control: no-cache` does not defeat
|
|
146
|
+
it (CDNs ignore it). A cached validator describes the origin as of
|
|
147
|
+
`age` seconds ago, not now. Pass `doc_from_probe(..., max_age_s=…)` to
|
|
148
|
+
drop validators older than you are willing to trust; they then fall
|
|
149
|
+
through to the probabilistic layers or UNKNOWN.
|
|
150
|
+
|
|
151
|
+
### Probing is locked down by default
|
|
152
|
+
|
|
153
|
+
URLs usually arrive from a retriever — that is, from *data* — so the
|
|
154
|
+
probe treats them as untrusted:
|
|
155
|
+
|
|
156
|
+
- **http/https only.** Left unguarded, `urllib` happily serves `file://`
|
|
157
|
+
and `ftp://`, which would turn "check this document" into local file
|
|
158
|
+
disclosure.
|
|
159
|
+
- **Public addresses only.** Loopback, private ranges, link-local
|
|
160
|
+
(including cloud metadata at `169.254.169.254`), reserved and multicast
|
|
161
|
+
targets are refused.
|
|
162
|
+
- **Redirects re-checked at every hop**, so a public URL answering `302`
|
|
163
|
+
with an internal `Location` does not become an escape hatch.
|
|
164
|
+
- **No credentials are ever attached** — no cookies, no auth headers.
|
|
165
|
+
|
|
166
|
+
Internal sources are a legitimate case (a wiki on `10.x` is normal), so
|
|
167
|
+
they are opt-in, not silently allowed:
|
|
168
|
+
|
|
169
|
+
```python
|
|
170
|
+
probe(url, allow_private=True) # I mean it, this host is internal
|
|
171
|
+
```
|
|
172
|
+
|
|
173
|
+
Known limit: the host is resolved once for the check and again by
|
|
174
|
+
`urllib` for the connection, so DNS rebinding is not covered — pin your
|
|
175
|
+
own resolver if that is in your threat model.
|
|
176
|
+
|
|
177
|
+
**What gets stored**: the observation log keeps only `doc_id`, a
|
|
178
|
+
timestamp, and the content hash you passed. `doc_id` is stored verbatim,
|
|
179
|
+
so if you pass raw URLs, anything embedded in them (tokens, credentials,
|
|
180
|
+
internal paths) is stored too. Pass an opaque id if that matters.
|
|
181
|
+
|
|
182
|
+
## Import the past instead of waiting for it
|
|
183
|
+
|
|
184
|
+
The usual deal with a freshness layer is "install it, come back in a
|
|
185
|
+
month". Skip that: most knowledge bases already keep their own change
|
|
186
|
+
history, and reading it is a one-liner.
|
|
187
|
+
|
|
188
|
+
```python
|
|
189
|
+
from stillvalid.backfill import from_git, from_csv, from_rows
|
|
190
|
+
|
|
191
|
+
from_git(sv, r"C:\work\handbook") # a docs repo: commits are the log
|
|
192
|
+
from_csv(sv, "wiki_revisions.csv") # doc_id, ts[, rev]
|
|
193
|
+
from_rows(sv, confluence_revision_rows) # anything you can iterate
|
|
194
|
+
```
|
|
195
|
+
|
|
196
|
+
Measured on three real repositories here: **3,228 past changes across
|
|
197
|
+
1,089 documents, imported in 1.7 seconds.** Before the import, a sample of
|
|
198
|
+
those documents judged `UNKNOWN` across the board; afterwards the same
|
|
199
|
+
query returned real verdicts with probabilities spread from 0% to 88%.
|
|
200
|
+
That is the difference between "ask me again in a month" and "ask me now".
|
|
201
|
+
|
|
202
|
+
Git is the easiest case, but the general entry point is `from_rows` —
|
|
203
|
+
feed it a wiki's revision API, a CMS audit table, a warehouse query of
|
|
204
|
+
`updated_at` snapshots. Consecutive identical revisions read as "still the
|
|
205
|
+
same", which is exactly the censoring signal the learned layers want.
|
|
206
|
+
|
|
207
|
+
`summarize(sv)` tells you what the import bought: how many documents can
|
|
208
|
+
now be judged on update behavior, and how many have enough events (8+) to
|
|
209
|
+
train a calibrated model on.
|
|
210
|
+
|
|
211
|
+
## It learns as you use it
|
|
212
|
+
|
|
213
|
+
The LangChain / LlamaIndex wrappers **auto-record** a content-hash
|
|
214
|
+
sighting for every document that flows through them (same-hash sightings
|
|
215
|
+
are deduped to one per hour; disable with `record=False`). Or log
|
|
216
|
+
manually whenever you fetch or reindex:
|
|
217
|
+
|
|
218
|
+
```python
|
|
219
|
+
sv.record(doc_id, content_hash)
|
|
220
|
+
```
|
|
221
|
+
|
|
222
|
+
Day 1: layers 1–3 and 7 carry the load. Within days, layer 5 starts
|
|
223
|
+
estimating from your observation log. With enough history (8+ observed
|
|
224
|
+
changes per document), you can train the survival model and export
|
|
225
|
+
parameters — then "updated 7 days ago" upgrades to "97% likely still
|
|
226
|
+
valid", calibrated.
|
|
227
|
+
|
|
228
|
+
One honesty note on auto-recording: it observes *your index's copy*, so
|
|
229
|
+
a change becomes visible at reindex time, not at source-change time.
|
|
230
|
+
That is exactly the fidelity the crude layer claims — calibrated
|
|
231
|
+
training should prefer source-side timestamps where available.
|
|
232
|
+
|
|
233
|
+
The trainer (survival analysis over your observation log, with the
|
|
234
|
+
serving gate) lives in the parent project; `stillvalid` only needs its
|
|
235
|
+
exported JSON (`fpi-qtime-params/1`).
|
|
236
|
+
|
|
237
|
+
## Cost
|
|
238
|
+
|
|
239
|
+
Measured on one laptop (Windows, Python 3.12) — order of magnitude, not a
|
|
240
|
+
benchmark claim:
|
|
241
|
+
|
|
242
|
+
| Operation | Cost |
|
|
243
|
+
|---|---|
|
|
244
|
+
| Verdict, deterministic layers | ~3 µs |
|
|
245
|
+
| Verdict, change-rate layer | ~0.5 ms, **constant** in history size |
|
|
246
|
+
| Batch of 1,000 documents | ~7 ms |
|
|
247
|
+
| `live=True`, 50 local files | ~0.9 ms |
|
|
248
|
+
| `live=True`, remote URLs (batched) | **~300 ms**, one round trip |
|
|
249
|
+
| Recording a sighting | ~0.4 ms (deduped: ~0.02 ms) |
|
|
250
|
+
| Observation log on disk | ~52 bytes/row |
|
|
251
|
+
|
|
252
|
+
Two things worth planning around:
|
|
253
|
+
|
|
254
|
+
- **Remote `live=True` costs a network round trip** (~300 ms, flat for a
|
|
255
|
+
batch since probes run concurrently). That is fine for an agent deciding
|
|
256
|
+
whether to act, and probably too slow in a latency-critical search path —
|
|
257
|
+
there, prefer local paths, or refresh signals on a schedule rather than
|
|
258
|
+
per query.
|
|
259
|
+
- **The change-rate layer only reads the most recent sightings**
|
|
260
|
+
(`History.RECENT_WINDOW`, 500). Reading full history made judging cost
|
|
261
|
+
grow with it; recent behavior also predicts better, so the window is
|
|
262
|
+
both faster and more honest.
|
|
263
|
+
|
|
264
|
+
## Verdicts
|
|
265
|
+
|
|
266
|
+
| State | Meaning | Default action |
|
|
267
|
+
|-------|---------|----------------|
|
|
268
|
+
| `VALID` | deterministically confirmed current | `use` |
|
|
269
|
+
| `LIKELY_VALID` | probability ≥ use-threshold (default 0.90) | `use` |
|
|
270
|
+
| `VERIFY` | uncertain — recheck the source | `verify` |
|
|
271
|
+
| `STALE` | expired / changed / probably invalid | `refresh` |
|
|
272
|
+
| `UNKNOWN` | no layer could judge | `verify` |
|
|
273
|
+
|
|
274
|
+
Thresholds are yours to set: `Checker(policy=Policy(use_threshold=0.95,
|
|
275
|
+
stale_threshold=0.70))`. Every verdict carries the deciding layer, a
|
|
276
|
+
human-readable reason, and the probability when one exists — evidence,
|
|
277
|
+
not just a verdict.
|
|
278
|
+
|
|
279
|
+
## Why thresholds on a calibrated probability (and not doc age)?
|
|
280
|
+
|
|
281
|
+
Because age can rank documents, but it can't answer "is 30 days old fine
|
|
282
|
+
*for this document*?" — a HQ address and a stock level are both "30 days
|
|
283
|
+
old" and deserve opposite treatment. In our backtests on a public-document
|
|
284
|
+
corpus (300k queries, weekly-reindex scenario), threshold decisions on
|
|
285
|
+
calibrated probabilities left 2.4–4.8× less stale exposure than an
|
|
286
|
+
age-cutoff heuristic at the same re-verification budget.
|
|
287
|
+
|
|
288
|
+
## Integrations
|
|
289
|
+
|
|
290
|
+
**LangChain** — judge every retrieved document between retrieval and the
|
|
291
|
+
LLM (`pip install "stillvalid[langchain]"`):
|
|
292
|
+
|
|
293
|
+
```python
|
|
294
|
+
# LangChain >= 1.0 moved this retriever into langchain_classic
|
|
295
|
+
from langchain_classic.retrievers import ContextualCompressionRetriever
|
|
296
|
+
from stillvalid import Checker
|
|
297
|
+
from stillvalid.integrations.langchain import StillValidCompressor
|
|
298
|
+
|
|
299
|
+
retriever = ContextualCompressionRetriever(
|
|
300
|
+
base_compressor=StillValidCompressor(checker=Checker(...), live=True),
|
|
301
|
+
base_retriever=vectorstore.as_retriever(),
|
|
302
|
+
)
|
|
303
|
+
# each doc gains metadata["stillvalid"] = {state, action, layer, reason, ...}
|
|
304
|
+
```
|
|
305
|
+
|
|
306
|
+
**LlamaIndex** — same idea as a node postprocessor
|
|
307
|
+
(`pip install "stillvalid[llamaindex]"`):
|
|
308
|
+
|
|
309
|
+
```python
|
|
310
|
+
from stillvalid.integrations.llamaindex import StillValidPostprocessor
|
|
311
|
+
|
|
312
|
+
engine = index.as_query_engine(
|
|
313
|
+
node_postprocessors=[StillValidPostprocessor(checker=Checker(...))])
|
|
314
|
+
```
|
|
315
|
+
|
|
316
|
+
Both read common metadata keys out of the box (`indexed_at`,
|
|
317
|
+
`last_modified`, `updated_at`, `content_hash`, LlamaIndex's
|
|
318
|
+
`last_modified_date`, …) and `stillvalid_*` prefixed keys always win.
|
|
319
|
+
`mode="filter"` drops non-usable documents but keeps `UNKNOWN` by
|
|
320
|
+
default — absence of evidence is not evidence of staleness
|
|
321
|
+
(`drop_unknown=True` if you disagree).
|
|
322
|
+
|
|
323
|
+
**Pass `live=True` unless you have a reason not to.** Index metadata is a
|
|
324
|
+
*past* snapshot by definition — it says when you last looked, never
|
|
325
|
+
whether the source has moved since — so without a current signal a fresh
|
|
326
|
+
install can only answer UNKNOWN. `live=True` fetches that signal at query
|
|
327
|
+
time, cheaply: `os.stat()` for local paths, a batched conditional HEAD for
|
|
328
|
+
URLs (never the body). In our end-to-end run over a real
|
|
329
|
+
loader → vector store → retriever pipeline, this turned *every* verdict
|
|
330
|
+
from UNKNOWN into a real one and flagged exactly the file that had been
|
|
331
|
+
edited after indexing — at **2 ms** added latency for local files.
|
|
332
|
+
Network sources cost one round trip; tune with `live_timeout`, and set
|
|
333
|
+
`allow_private=True` if your wiki is on an internal address.
|
|
334
|
+
|
|
335
|
+
**MCP server** — let the agent itself ask
|
|
336
|
+
(`pip install "stillvalid[mcp]"`, supports mcp 1.x and 2.x):
|
|
337
|
+
|
|
338
|
+
```jsonc
|
|
339
|
+
// Claude Desktop / any MCP client
|
|
340
|
+
{"mcpServers": {"stillvalid": {
|
|
341
|
+
"command": "python", "args": ["-m", "stillvalid.mcp_server"],
|
|
342
|
+
"env": {"STILLVALID_HISTORY_DB": "C:/data/observations.db"}}}}
|
|
343
|
+
```
|
|
344
|
+
|
|
345
|
+
Tools: `import_history` (load a corpus's past in seconds — run this
|
|
346
|
+
first), `check_validity` (judge one piece of evidence before acting on
|
|
347
|
+
it), `check_validity_batch` (shared clock), `record_observation` (log a
|
|
348
|
+
sighting so the cascade keeps learning).
|
|
349
|
+
|
|
350
|
+
In an agent session that looks like: *"point me at your docs repo"* →
|
|
351
|
+
`import_history` → *"3,228 past changes imported; 554 of 1,089 documents
|
|
352
|
+
can now be judged on their update behavior"* → every later answer is
|
|
353
|
+
checked against that.
|
|
354
|
+
|
|
355
|
+
## Roadmap
|
|
356
|
+
|
|
357
|
+
- Cross-source verification layer (layer 4 of the design)
|
|
358
|
+
- Cold-start priors: borrowing update-behavior from similar sources,
|
|
359
|
+
gated against negative transfer
|
|
360
|
+
- `pip install stillvalid` (PyPI release)
|
|
361
|
+
|
|
362
|
+
## Status & honesty
|
|
363
|
+
|
|
364
|
+
This is an extraction of a working research pipeline (survival analysis
|
|
365
|
+
over change histories, with calibration verification) into a standalone
|
|
366
|
+
layer. The calibrated path is real but requires training on your
|
|
367
|
+
observation log; everything else works out of the box. If your documents
|
|
368
|
+
all have reliable feeds or explicit expiries, you don't need the learned
|
|
369
|
+
layers — and this library will happily tell you so by never reaching them.
|