@zosmaai/pi-llm-wiki 0.11.3 → 0.11.5
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +10 -0
- package/README.de.md +8 -0
- package/README.es.md +8 -0
- package/README.fr.md +8 -0
- package/README.hi.md +8 -0
- package/README.ja.md +8 -0
- package/README.ko.md +8 -0
- package/README.md +88 -2
- package/README.pt.md +8 -0
- package/README.ru.md +8 -0
- package/README.zh.md +8 -0
- package/assets/wiki-dashboard.png +0 -0
- package/commands/wiki-digest.md +28 -0
- package/commands/wiki-discover.md +30 -0
- package/commands/wiki-ingest.md +37 -0
- package/commands/wiki-init.md +30 -0
- package/commands/wiki-lint.md +25 -0
- package/commands/wiki-query.md +37 -0
- package/commands/wiki-record.md +36 -0
- package/commands/wiki-req.md +56 -0
- package/commands/wiki-retro.md +35 -0
- package/commands/wiki-run.md +31 -0
- package/commands/wiki-skills.md +26 -0
- package/commands/wiki-status.md +16 -0
- package/dist/extensions/llm-wiki/lib/dashboard-command.js +86 -0
- package/dist/extensions/llm-wiki/lib/dashboard.js +175 -0
- package/dist/extensions/llm-wiki/lib/guardrails.js +30 -1
- package/dist/extensions/llm-wiki/lib/host.js +117 -0
- package/dist/extensions/llm-wiki/lib/ingest-worker.js +2 -1
- package/dist/extensions/llm-wiki/lib/knowledge-document.js +20 -2
- package/dist/extensions/llm-wiki/lib/knowledge-links.js +6 -3
- package/dist/extensions/llm-wiki/lib/metadata.js +1 -1
- package/dist/extensions/llm-wiki/lib/observation.js +31 -3
- package/dist/extensions/llm-wiki/lib/settings-command.js +377 -0
- package/dist/extensions/llm-wiki/lib/task-config.js +145 -43
- package/dist/extensions/llm-wiki/lib/utils.js +59 -16
- package/docs/api.md +24 -1
- package/docs/commands.md +6 -1
- package/docs/configuration.md +62 -11
- package/docs/superpowers/plans/2026-08-09-qmd-retrieval-phase-1-quality-baseline-and-compatibility.md +1520 -0
- package/docs/superpowers/roadmaps/2026-08-09-qmd-retrieval-roadmap.md +448 -0
- package/docs/superpowers/specs/2026-08-08-qmd-retrieval-design.md +806 -0
- package/extensions/llm-wiki/index.ts +48 -6
- package/extensions/llm-wiki/lib/dashboard-command.ts +106 -0
- package/extensions/llm-wiki/lib/dashboard.ts +210 -0
- package/extensions/llm-wiki/lib/guardrails.ts +26 -1
- package/extensions/llm-wiki/lib/host.ts +145 -0
- package/extensions/llm-wiki/lib/ingest-worker.ts +4 -0
- package/extensions/llm-wiki/lib/knowledge-document.ts +20 -2
- package/extensions/llm-wiki/lib/knowledge-links.ts +7 -3
- package/extensions/llm-wiki/lib/metadata.ts +1 -1
- package/extensions/llm-wiki/lib/observation.ts +37 -4
- package/extensions/llm-wiki/lib/settings-command.ts +483 -0
- package/extensions/llm-wiki/lib/task-config.ts +208 -46
- package/extensions/llm-wiki/lib/utils.ts +55 -14
- package/package.json +15 -4
- package/prompts/wiki-ingest.md +1 -0
- package/prompts/wiki-req.md +1 -0
- package/prompts/wiki-retro.md +1 -0
- package/skills/llm-wiki/SKILL.md +11 -1
|
@@ -0,0 +1,806 @@
|
|
|
1
|
+
# QMD-Backed Second-Brain Retrieval Design
|
|
2
|
+
|
|
3
|
+
**Status:** Approved design
|
|
4
|
+
**Target:** Next major release
|
|
5
|
+
**Runtime:** Node.js 22 or newer
|
|
6
|
+
**Date:** 2026-08-08
|
|
7
|
+
|
|
8
|
+
## Summary
|
|
9
|
+
|
|
10
|
+
pi-llm-wiki will replace its heuristic recall ranking with QMD as its first-class retrieval engine. QMD will own Markdown indexing, BM25 retrieval, dense retrieval, rank fusion, query expansion, and local reranking. pi-llm-wiki will continue to own the knowledge model: canonical-card prioritization, evidence and contradiction assembly, feedback, context packing, layered vault behavior, and quality evaluation.
|
|
11
|
+
|
|
12
|
+
Markdown remains authoritative. QMD's SQLite database and model artifacts are local, rebuildable implementation details. No custom database or graph database will be introduced.
|
|
13
|
+
|
|
14
|
+
The product goal is not to return pages that merely resemble a query. Recall must provide useful memory for an AI: the best reviewed conclusion first, its strongest evidence, applicable scope and freshness, and any relevant competing claim.
|
|
15
|
+
|
|
16
|
+
## Current-State Audit
|
|
17
|
+
|
|
18
|
+
The personal vault inspected during design contained:
|
|
19
|
+
|
|
20
|
+
- 701 registered pages
|
|
21
|
+
- 626 `source` pages
|
|
22
|
+
- 40 `concept` pages
|
|
23
|
+
- 26 `entity` pages
|
|
24
|
+
- 9 generic pages
|
|
25
|
+
- titles and types on every registered page
|
|
26
|
+
- tags on 63% of pages
|
|
27
|
+
- effectively no summaries, descriptions, aliases, recall triggers, or domain metadata
|
|
28
|
+
- no `meta/embeddings.json`, so live recall was lexical-only
|
|
29
|
+
|
|
30
|
+
Current recall performs weighted substring matching over metadata and heading chunks, heuristic pseudo-relevance feedback, optional page-level embedding boosts, and raw additive score fusion. Its tests establish deterministic mechanics but do not measure relevance on real queries.
|
|
31
|
+
|
|
32
|
+
The live recall examples observed during brainstorming returned unrelated accounting, webhook, shopping, and build notes for questions about retrieval architecture. This is unacceptable for automatic AI context injection.
|
|
33
|
+
|
|
34
|
+
## Goals
|
|
35
|
+
|
|
36
|
+
1. Make recall precise and useful enough to improve downstream AI work.
|
|
37
|
+
2. Return reviewed canonical knowledge before raw observations whenever possible.
|
|
38
|
+
3. Preserve source evidence, provenance, scope, freshness, and conflicting claims.
|
|
39
|
+
4. Support exact lookup, paraphrase, vague recollection, temporal questions, contradiction discovery, and synthesis.
|
|
40
|
+
5. Provide lexical, hybrid, adaptive, and maximum-quality retrieval modes.
|
|
41
|
+
6. Reindex incrementally during ordinary writes and fully on explicit request.
|
|
42
|
+
7. Learn conservatively from explicit corrections and local usage signals.
|
|
43
|
+
8. Measure retrieval quality against a versioned benchmark built from real queries.
|
|
44
|
+
9. Keep Markdown portable and Obsidian-compatible.
|
|
45
|
+
10. Fail toward no recall or a cheaper retrieval mode, never toward unrelated injected memory.
|
|
46
|
+
|
|
47
|
+
## Non-Goals
|
|
48
|
+
|
|
49
|
+
- A custom vector database
|
|
50
|
+
- A graph database
|
|
51
|
+
- Graph-first or agentic traversal as the primary retriever
|
|
52
|
+
- Automatic conversion of every observation into a canonical card
|
|
53
|
+
- Automatic deletion or silent replacement of conflicting claims
|
|
54
|
+
- Training a personalized ranking model from a small feedback corpus
|
|
55
|
+
- Supporting Node.js 18 in the new major release
|
|
56
|
+
- Maintaining two independent full retrieval engines
|
|
57
|
+
- Treating Obsidian Graph View as a ranking algorithm
|
|
58
|
+
|
|
59
|
+
## Decisions and Alternatives
|
|
60
|
+
|
|
61
|
+
### Rejected: tune the existing scorer
|
|
62
|
+
|
|
63
|
+
Adding metadata, enabling page embeddings, and adjusting weights would be the smallest change, but it would preserve the weakest parts of the current system: broad page representations, incomparable score addition, heuristic query drift, and no strong reranker.
|
|
64
|
+
|
|
65
|
+
### Rejected: graph-first retrieval
|
|
66
|
+
|
|
67
|
+
A Zettelkasten graph improves human thinking and navigation, but graph traversal cannot reliably find a seed note from an arbitrary query. The current graph is also dominated by source observations. Typed links will support bounded evidence assembly after retrieval, not replace retrieval.
|
|
68
|
+
|
|
69
|
+
### Selected: QMD retrieval plus pi-llm-wiki memory semantics
|
|
70
|
+
|
|
71
|
+
QMD already supplies the retrieval machinery this project would otherwise need to build and maintain:
|
|
72
|
+
|
|
73
|
+
- BM25 full-text retrieval
|
|
74
|
+
- dense retrieval over chunks
|
|
75
|
+
- reciprocal-rank fusion
|
|
76
|
+
- query expansion
|
|
77
|
+
- local reranking
|
|
78
|
+
- incremental indexing and embedding
|
|
79
|
+
- a typed TypeScript SDK
|
|
80
|
+
|
|
81
|
+
pi-llm-wiki will use QMD as a library rather than copy its implementation. The next major release will require Node.js 22 or newer and include QMD as a runtime dependency. Users who need Node.js 18 can remain on the previous major release.
|
|
82
|
+
|
|
83
|
+
## Zettelkasten Role
|
|
84
|
+
|
|
85
|
+
Zettelkasten is the knowledge organization model, not the search engine.
|
|
86
|
+
|
|
87
|
+
Its useful principles are:
|
|
88
|
+
|
|
89
|
+
- one durable idea per permanent note at a useful retrieval granularity
|
|
90
|
+
- source or literature notes retained as evidence
|
|
91
|
+
- explicit links that communicate meaningful relationships
|
|
92
|
+
- structure notes and backlinks that support browsing and synthesis
|
|
93
|
+
|
|
94
|
+
In pi-llm-wiki, existing `concept`, `entity`, `analysis`, `synthesis`, `skill`, and similar pages can act as permanent cards. Existing `source` and observation pages act as evidence. QMD finds candidates; Zettelkasten structure makes those candidates useful and composable.
|
|
95
|
+
|
|
96
|
+
Obsidian Graph View remains a human curation surface for finding orphans, clusters, and missing links. It visualizes existing links and tags but does not affect query relevance by itself.
|
|
97
|
+
|
|
98
|
+
## Architecture
|
|
99
|
+
|
|
100
|
+
```text
|
|
101
|
+
Markdown vaults
|
|
102
|
+
├── canonical knowledge pages
|
|
103
|
+
├── source and observation evidence
|
|
104
|
+
└── typed Markdown relationships
|
|
105
|
+
│
|
|
106
|
+
▼
|
|
107
|
+
QMD stores, one per vault
|
|
108
|
+
├── filesystem index
|
|
109
|
+
├── BM25 index
|
|
110
|
+
├── heading-aware chunks
|
|
111
|
+
├── embeddings
|
|
112
|
+
└── local reranker
|
|
113
|
+
│
|
|
114
|
+
▼
|
|
115
|
+
pi-llm-wiki recall service
|
|
116
|
+
├── layered vault merge
|
|
117
|
+
├── canonical/evidence classification
|
|
118
|
+
├── trust, scope, and freshness adjustments
|
|
119
|
+
├── feedback adjustments
|
|
120
|
+
├── evidence and contradiction expansion
|
|
121
|
+
└── diverse token-budgeted context packing
|
|
122
|
+
│
|
|
123
|
+
▼
|
|
124
|
+
Pi tool, MCP, and automatic recall adapters
|
|
125
|
+
```
|
|
126
|
+
|
|
127
|
+
### Ownership
|
|
128
|
+
|
|
129
|
+
- `.llm-wiki/wiki/**/*.md` remains authoritative, user-editable knowledge.
|
|
130
|
+
- `.llm-wiki/raw/**` remains immutable source material.
|
|
131
|
+
- `.llm-wiki/meta/qmd/**` is generated, extension-owned, and rebuildable. It contains the validated document mirror, path manifest, and complete QMD database artifacts.
|
|
132
|
+
- QMD model files remain in QMD's standard local cache.
|
|
133
|
+
- Recall feedback uses the existing authoritative `.llm-wiki/meta/events.jsonl` stream. `recall_feedback` events have `visibility: internal`, so generated human-readable logs omit them.
|
|
134
|
+
- `.llm-wiki/meta/recall-feedback.json` is a rebuildable aggregate projection of valid internal feedback events.
|
|
135
|
+
|
|
136
|
+
The existing guardrails for `meta/**` apply. Only extension code may modify QMD artifacts, feedback events, or projections. Full-vault backups retain feedback through `events.jsonl`; OKF-only exports exclude it, matching existing event-stream behavior. The implementation must update `docs/architecture.md` and the OKF ownership documentation to describe internal events and the feedback projection.
|
|
137
|
+
|
|
138
|
+
Feedback appends go through the existing `appendEvent` service as one append operation. This release upgrades that service to serialize writers with a per-vault in-process queue plus an extension-owned `meta/events.lock` acquired by exclusive file creation. The lock records writer ID and acquisition time; a writer may recover a lock older than 30 seconds only after confirming the recorded local process no longer exists. Implicit events wait at most 250 ms and may be dropped with a diagnostic; explicit corrections and conflict choices wait up to 2 seconds and return an error without mutating knowledge if the lock remains busy. Replay ignores events for deleted pages. There is no automatic retention or compaction in this release. If the event source is unreadable or malformed, feedback adjustments are disabled, lint reports the corruption, and the last generated projection is preserved rather than guessed or rewritten.
|
|
139
|
+
|
|
140
|
+
### One store per vault
|
|
141
|
+
|
|
142
|
+
Each personal or project vault gets an independent QMD store under its own `meta/qmd/` directory. This prevents a project index from copying personal content and keeps index invalidation local to the vault that changed.
|
|
143
|
+
|
|
144
|
+
Automatic recall searches only the active project vault when one exists, matching the current contamination guard. Outside a project vault, it searches the personal vault. Explicit `wiki_recall` searches both applicable vaults in parallel, deduplicates by page ID with project precedence, and combines the ranked lists using rank-based fusion rather than assuming raw scores are comparable across stores.
|
|
145
|
+
|
|
146
|
+
Every vault has a stable UUID in `.llm-wiki/config.json` (`vault_id`). Bootstrap creates it; first QMD indexing backfills it once for existing vaults while preserving all other configuration. Feedback and layered result identities use `(vault_id, page_id)`, so identical personal and project page IDs never share signals.
|
|
147
|
+
|
|
148
|
+
### Validated QMD mirror and collections
|
|
149
|
+
|
|
150
|
+
QMD must never scan authoritative `wiki/**` directly. Before `store.update`, pi-llm-wiki parses changed pages through the shared `KnowledgeDocument` layer and writes only valid documents into a generated mirror:
|
|
151
|
+
|
|
152
|
+
```text
|
|
153
|
+
meta/qmd/documents/
|
|
154
|
+
├── canonical/<folder-qualified-id>.md
|
|
155
|
+
└── evidence/<folder-qualified-id>.md
|
|
156
|
+
```
|
|
157
|
+
|
|
158
|
+
`meta/qmd/manifest.json` maps each mirror path to its original absolute path, vault ID, page ID, content hash, and role. Malformed pages and reserved generated `index.md`/`log.md` files have no mirror entry, so they cannot influence QMD candidate generation, expansion, or reranking. Removing or invalidating an authoritative page removes its mirror copy before the next QMD update.
|
|
159
|
+
|
|
160
|
+
The store defines two non-overlapping filesystem collections through QMD's supported `path` plus `pattern` configuration:
|
|
161
|
+
|
|
162
|
+
```ts
|
|
163
|
+
collections: {
|
|
164
|
+
canonical: { path: "<meta>/qmd/documents/canonical", pattern: "**/*.md" },
|
|
165
|
+
evidence: { path: "<meta>/qmd/documents/evidence", pattern: "**/*.md" },
|
|
166
|
+
}
|
|
167
|
+
```
|
|
168
|
+
|
|
169
|
+
Unknown page types default to the evidence mirror. Collection context tells QMD that canonical pages contain reusable conclusions while evidence pages contain provenance and historical observations. Every QMD result is mapped back through the manifest and parsed metadata before it enters memory assembly.
|
|
170
|
+
|
|
171
|
+
### Shared service boundary
|
|
172
|
+
|
|
173
|
+
A shared retrieval service owns all behavior. Pi tools, automatic injection, and MCP call the same service and only render its structured result.
|
|
174
|
+
|
|
175
|
+
Conceptual interfaces:
|
|
176
|
+
|
|
177
|
+
```ts
|
|
178
|
+
type RetrievalMode = "lexical" | "hybrid" | "adaptive" | "quality";
|
|
179
|
+
type ConflictRelation = "supersedes" | "qualifies" | "applies_under";
|
|
180
|
+
|
|
181
|
+
type ConflictResolution = {
|
|
182
|
+
id: string;
|
|
183
|
+
vaultId: string;
|
|
184
|
+
selectedPageId: string;
|
|
185
|
+
otherPageId: string;
|
|
186
|
+
relation: ConflictRelation;
|
|
187
|
+
scope?: string;
|
|
188
|
+
status: "active";
|
|
189
|
+
decidedBy: "user";
|
|
190
|
+
decidedAt: string;
|
|
191
|
+
replaces?: string;
|
|
192
|
+
};
|
|
193
|
+
|
|
194
|
+
type MemoryCandidate = {
|
|
195
|
+
vault: "project" | "personal";
|
|
196
|
+
vaultId: string;
|
|
197
|
+
pageId: string;
|
|
198
|
+
path: string;
|
|
199
|
+
heading?: string;
|
|
200
|
+
excerpt: string;
|
|
201
|
+
qmdRank: number;
|
|
202
|
+
qmdScore?: number;
|
|
203
|
+
score: number;
|
|
204
|
+
role: "canonical" | "evidence";
|
|
205
|
+
status: "draft" | "stable" | "deprecated";
|
|
206
|
+
matchReasons: string[];
|
|
207
|
+
};
|
|
208
|
+
|
|
209
|
+
type MemoryBundle = {
|
|
210
|
+
card?: MemoryCandidate;
|
|
211
|
+
evidence: MemoryCandidate[];
|
|
212
|
+
conflicts: MemoryCandidate[];
|
|
213
|
+
resolution?: ConflictResolution;
|
|
214
|
+
};
|
|
215
|
+
```
|
|
216
|
+
|
|
217
|
+
QMD-specific objects must not escape the QMD adapter. This keeps memory semantics testable without loading local models.
|
|
218
|
+
|
|
219
|
+
## Knowledge Model
|
|
220
|
+
|
|
221
|
+
### Card role
|
|
222
|
+
|
|
223
|
+
“Card” is a retrieval role, not a new required page type.
|
|
224
|
+
|
|
225
|
+
Canonical candidates include:
|
|
226
|
+
|
|
227
|
+
- `concept`
|
|
228
|
+
- `entity`
|
|
229
|
+
- `analysis`
|
|
230
|
+
- `synthesis`
|
|
231
|
+
- `requirement`
|
|
232
|
+
- `skill`
|
|
233
|
+
- `case`
|
|
234
|
+
|
|
235
|
+
Evidence candidates include:
|
|
236
|
+
|
|
237
|
+
- `source`
|
|
238
|
+
- observation pages
|
|
239
|
+
- `trajectory`
|
|
240
|
+
- unknown or generic pages unless explicitly promoted
|
|
241
|
+
|
|
242
|
+
Within canonical candidates, `status: stable` and current verification receive bounded preference. Draft pages remain retrievable but are labeled. Deprecated and stale pages remain retrievable when relevant and carry warnings.
|
|
243
|
+
|
|
244
|
+
### Atomic canonical pages
|
|
245
|
+
|
|
246
|
+
A canonical page should express one independently useful idea, claim, entity, requirement, or reusable procedure. Atomicity is measured by retrieval usefulness, not sentence count. The system must not split a coherent idea into artificial fragments merely to create more cards.
|
|
247
|
+
|
|
248
|
+
Recommended body sections are:
|
|
249
|
+
|
|
250
|
+
```md
|
|
251
|
+
## Claim
|
|
252
|
+
|
|
253
|
+
## Scope
|
|
254
|
+
|
|
255
|
+
## Evidence
|
|
256
|
+
|
|
257
|
+
## Related
|
|
258
|
+
```
|
|
259
|
+
|
|
260
|
+
The existing OKF-compatible metadata remains authoritative for type, title, description, sources, generation, verification, status, and staleness. pi-llm-wiki defines `relations` as an optional profile extension:
|
|
261
|
+
|
|
262
|
+
```yaml
|
|
263
|
+
relations:
|
|
264
|
+
- target: concepts/other-card
|
|
265
|
+
type: contradicts
|
|
266
|
+
```
|
|
267
|
+
|
|
268
|
+
Each relation is a mapping with exactly `target`, `type`, and optional `scope`. `target` is a normalized folder-qualified page ID without a heading or block fragment. `scope` is required for `applies_under` and forbidden for the other first-release types. Allowed types are:
|
|
269
|
+
|
|
270
|
+
- `supports`
|
|
271
|
+
- `contradicts`
|
|
272
|
+
- `qualifies`
|
|
273
|
+
- `supersedes`
|
|
274
|
+
- `applies_under`
|
|
275
|
+
- `derived_from`
|
|
276
|
+
- `related_to`
|
|
277
|
+
|
|
278
|
+
The shared parser preserves the extension, validator checks its shape, serializer round-trips it, lint resolves each target and reports duplicates, self-links, missing targets, invalid scope use, and unknown relation types. Unknown imported relation types are preserved but inert. The generated relation adjacency map deduplicates equivalent edges.
|
|
279
|
+
|
|
280
|
+
A typed relation must also have an ordinary Markdown link to the same page in the body. The Markdown link keeps Obsidian and plain OKF consumers navigable; `relations` supplies machine-readable semantics. If the body link is absent, lint reports `relation_missing_body_link` and the relation does not participate in expansion. Heading and Obsidian block anchors belong on the body link; relationship identity remains page-level. Backlinks derive one page edge from the Markdown link and attach the validated relation type without creating a duplicate edge.
|
|
281
|
+
|
|
282
|
+
Only `supports`, `contradicts`, `qualifies`, `supersedes`, `applies_under`, and `derived_from` affect evidence or conflict expansion. `related_to` remains a browsing link and receives no automatic inclusion privilege. Active user resolution records take precedence when a declared relation and a later resolution disagree, while both remain visible for audit.
|
|
283
|
+
|
|
284
|
+
### Unpromoted evidence
|
|
285
|
+
|
|
286
|
+
When no canonical card exists, recall may return the strongest source excerpt as **unpromoted evidence**. It must not be presented as an established conclusion. Promotion automation is outside this design; usage and corrections may inform a separately specified future workflow.
|
|
287
|
+
|
|
288
|
+
## Indexing and Reindexing
|
|
289
|
+
|
|
290
|
+
### Normal write path
|
|
291
|
+
|
|
292
|
+
After a successful metadata rebuild, the indexing coordinator parses changed Markdown pages and atomically updates their validated mirror copies and manifest entries. It then calls QMD's incremental update against the mirror. Stale embeddings are generated in the background. A write remains successful even if mirror or QMD indexing fails; the last valid QMD index stays available and status reports the authoritative page hash that has not yet been indexed. A malformed page never receives or retains a mirror copy for its invalid content.
|
|
293
|
+
|
|
294
|
+
### Explicit tool
|
|
295
|
+
|
|
296
|
+
The next major release adds one consolidated tool:
|
|
297
|
+
|
|
298
|
+
```text
|
|
299
|
+
wiki_reindex(
|
|
300
|
+
scope: "changed" | "all" = "changed",
|
|
301
|
+
components: ["lexical", "vectors"] = ["lexical", "vectors"],
|
|
302
|
+
force: boolean = false,
|
|
303
|
+
vault: "active" | "personal" | "project" | "all" = "active"
|
|
304
|
+
)
|
|
305
|
+
```
|
|
306
|
+
|
|
307
|
+
Semantics:
|
|
308
|
+
|
|
309
|
+
- `changed` scans files and updates added, changed, and removed documents.
|
|
310
|
+
- `all` scans the entire selected vault but still skips fresh vectors unless `force` is true.
|
|
311
|
+
- selecting only `lexical` never loads embedding or reranking models.
|
|
312
|
+
- selecting `vectors` first performs the document update required to identify stale chunks.
|
|
313
|
+
- `force` applies only to selected components.
|
|
314
|
+
- `all` vault scope processes stores independently and reports each outcome.
|
|
315
|
+
|
|
316
|
+
The tool reports:
|
|
317
|
+
|
|
318
|
+
- collections scanned
|
|
319
|
+
- pages indexed, updated, unchanged, and removed
|
|
320
|
+
- chunks needing embeddings
|
|
321
|
+
- vectors generated or skipped
|
|
322
|
+
- index and model versions
|
|
323
|
+
- elapsed time
|
|
324
|
+
- structured warnings and errors
|
|
325
|
+
|
|
326
|
+
The existing `wiki_reindex_embeddings` tool is deprecated in the new major and delegates to `wiki_reindex(components=["vectors"])` for one release before removal.
|
|
327
|
+
|
|
328
|
+
### Index lifecycle
|
|
329
|
+
|
|
330
|
+
QMD owns its SQLite schema and transactions. pi-llm-wiki does not manipulate QMD tables directly. Generated state lives as a complete artifact directory:
|
|
331
|
+
|
|
332
|
+
```text
|
|
333
|
+
meta/qmd/
|
|
334
|
+
├── current/ # index.sqlite plus any WAL/SHM/native sidecars
|
|
335
|
+
├── documents/ # validated mirror
|
|
336
|
+
├── manifest.json
|
|
337
|
+
└── swap.json # present only during recoverable replacement
|
|
338
|
+
```
|
|
339
|
+
|
|
340
|
+
A forced rebuild uses QMD-supported replacement if the pinned SDK provides it. Otherwise it builds `staging-<uuid>/`, closes both QMD stores, checkpoints through QMD when supported, validates status and document counts, and swaps the complete directory rather than a lone SQLite file. The adapter records the swap phase in `swap.json`, renames `current` to `previous`, renames staging to `current`, reopens and validates it, then removes `previous`. Startup recovery completes or rolls back any interrupted phase. A failed reopen restores `previous`. No rename occurs while either store is open.
|
|
341
|
+
|
|
342
|
+
Content hashes, QMD schema/version information, and resolved model identities determine staleness. A removed Markdown page must disappear from QMD results after the next incremental update. Historical feedback may retain its page ID but cannot boost a result that no longer exists.
|
|
343
|
+
|
|
344
|
+
## Retrieval Modes
|
|
345
|
+
|
|
346
|
+
Configuration adds one primary setting:
|
|
347
|
+
|
|
348
|
+
```json
|
|
349
|
+
{
|
|
350
|
+
"llm-wiki": {
|
|
351
|
+
"retrievalMode": "adaptive"
|
|
352
|
+
}
|
|
353
|
+
}
|
|
354
|
+
```
|
|
355
|
+
|
|
356
|
+
Invalid explicit values fail closed with a configuration diagnostic. Default is `adaptive`.
|
|
357
|
+
|
|
358
|
+
### QMD adapter mapping
|
|
359
|
+
|
|
360
|
+
| Mode | Exact QMD SDK path | Model-failure fallback |
|
|
361
|
+
|---|---|---|
|
|
362
|
+
| `lexical` | `store.searchLex(query, { limit: 40 })` | no lower mode; return diagnostic on native/store failure |
|
|
363
|
+
| `hybrid` | `store.search({ queries: [{ type: "lex", query }, { type: "vec", query }], rerank: false, candidateLimit: 40, limit: 10, explain: true })` | `searchLex` |
|
|
364
|
+
| `adaptive` initial | same typed-query call as `hybrid` | `searchLex` |
|
|
365
|
+
| `adaptive` uncertain | `store.search({ query, intent, rerank: true, candidateLimit: 40, limit: 10, explain: true })` | retain initial hybrid list |
|
|
366
|
+
| `quality` | same expanded/reranked call as adaptive-uncertain | typed hybrid, then lexical |
|
|
367
|
+
|
|
368
|
+
A typed `queries` call is mandatory for hybrid mode because plain `search({ query })` performs QMD query expansion. The adapter verifies these mappings against the pinned QMD SDK in contract tests.
|
|
369
|
+
|
|
370
|
+
### `lexical`
|
|
371
|
+
|
|
372
|
+
QMD BM25 only; no embedding, expansion, or reranking model is loaded. This mode is best for exact names, commands, identifiers, filenames, and low-resource systems.
|
|
373
|
+
|
|
374
|
+
### `hybrid`
|
|
375
|
+
|
|
376
|
+
The original query is sent as explicit lexical and vector subqueries. QMD performs rank fusion without generative expansion or final reranking.
|
|
377
|
+
|
|
378
|
+
### `adaptive`
|
|
379
|
+
|
|
380
|
+
Adaptive starts with the exact hybrid call. It invokes QMD's expanded and reranked path when any initial condition is true:
|
|
381
|
+
|
|
382
|
+
- normalized top-two score margin is below `0.08`
|
|
383
|
+
- Jaccard overlap between lexical and vector top-five page IDs is below `0.40`
|
|
384
|
+
- the normalized query begins with `who`, `what`, `when`, `where`, `why`, `how`, `which`, `compare`, `explain`, or their benchmarked multilingual equivalents
|
|
385
|
+
- the initial candidates contain a validated contradiction edge
|
|
386
|
+
|
|
387
|
+
An exact title, alias, command, filename, or page-ID match with normalized score at least `0.80` bypasses reranking.
|
|
388
|
+
|
|
389
|
+
### `quality`
|
|
390
|
+
|
|
391
|
+
Quality always uses QMD query expansion, lexical and vector retrieval, rank fusion, and local reranking of at most 40 candidates.
|
|
392
|
+
|
|
393
|
+
### Score normalization and first-release limits
|
|
394
|
+
|
|
395
|
+
QMD raw scores are used for within-store confidence only. Cross-vault merging converts each store's ordered results to reciprocal-rank score `1 / (60 + rank)`. The final bounded multiplier is clamped to `0.90–1.10`: exact identity match `+0.05`, stable canonical role `+0.03`, evidence role `-0.02`, deprecated status `-0.08`, and all implicit feedback combined at no more than `±0.02`. Ties resolve by lower QMD rank, then project vault, then page ID. These constants are changed only through benchmarked code changes.
|
|
396
|
+
|
|
397
|
+
Automatic recall requests at most 10 QMD candidates per vault, requires normalized QMD score `>= 0.50`, emits at most 3 memory bundles, and uses an 8,000-character context budget. Explicit recall uses caller `max_results` (default 5, maximum 10), a `0.25` QMD floor, and existing links-first rendering above the vault threshold.
|
|
398
|
+
|
|
399
|
+
Each automatic bundle contains at most one canonical page and two evidence excerpts of at most 800 characters each. Parent-page deduplication is mandatory. If alternatives exist, no more than two canonical cards with the same `(type, domain)` pair may occupy the three automatic slots. Evidence order is direct `supports`/`derived_from`, then QMD rank, then page ID. Lower-ranked bundles are removed before evidence attached to the top bundle.
|
|
400
|
+
|
|
401
|
+
QMD's supported environment variables remain the model-override surface. pi-llm-wiki will not duplicate every QMD model setting.
|
|
402
|
+
|
|
403
|
+
### Query and intent construction
|
|
404
|
+
|
|
405
|
+
The adapter derives `query` by applying Unicode NFKC normalization, trimming, and collapsing whitespace to the explicit `wiki_recall` query or current automatic-recall prompt. It does not remove punctuation, quoted phrases, path separators, or identifier characters before QMD receives the string. Empty queries return no result. Inputs above 2,000 characters are rejected with `recall_query_too_long` rather than silently truncated.
|
|
406
|
+
|
|
407
|
+
For expanded/reranked calls, `intent` is the concatenation of non-empty vault `topic` and `mode` values from config, capped at 256 characters. It provides collection disambiguation only and never replaces or expands the query itself. If both are absent, `intent` is omitted. The full conversational transcript is never sent to QMD.
|
|
408
|
+
|
|
409
|
+
## Recall Data Flow
|
|
410
|
+
|
|
411
|
+
1. Resolve active personal and project vaults.
|
|
412
|
+
2. Normalize the query while preserving exact identifiers and original wording.
|
|
413
|
+
3. Select automatic or explicit recall policy.
|
|
414
|
+
4. Query applicable QMD stores using the configured retrieval mode.
|
|
415
|
+
5. Merge layered result lists by rank and deduplicate project/personal page ID collisions with project precedence.
|
|
416
|
+
6. Parse each candidate through the shared knowledge-document layer.
|
|
417
|
+
7. Classify canonical versus evidence role and derive status, trust, scope, and freshness warnings.
|
|
418
|
+
8. Apply bounded feedback and exact-match adjustments.
|
|
419
|
+
9. Group excerpts under their parent pages.
|
|
420
|
+
10. Prefer reviewed canonical cards.
|
|
421
|
+
11. Follow at most one typed relationship hop for direct evidence, qualifiers, superseding claims, and contradictions.
|
|
422
|
+
12. Pack diverse memory bundles under the context budget.
|
|
423
|
+
13. Render structured context for Pi or MCP.
|
|
424
|
+
|
|
425
|
+
### Automatic versus explicit recall
|
|
426
|
+
|
|
427
|
+
Automatic recall is precision-first:
|
|
428
|
+
|
|
429
|
+
- active project only when present; personal otherwise
|
|
430
|
+
- at most 10 candidates and 3 packed bundles
|
|
431
|
+
- normalized QMD score floor `0.50`
|
|
432
|
+
- no injection when no candidate clears that floor
|
|
433
|
+
- no generic graph expansion
|
|
434
|
+
|
|
435
|
+
Explicit `wiki_recall` is recall-first:
|
|
436
|
+
|
|
437
|
+
- searches personal and project stores when both exist
|
|
438
|
+
- accepts a larger result count
|
|
439
|
+
- returns alternatives and unpromoted evidence
|
|
440
|
+
- exposes match explanations and diagnostics
|
|
441
|
+
|
|
442
|
+
“No reliable memory found” is a valid successful outcome. The system must not fill an empty slot with a weak candidate.
|
|
443
|
+
|
|
444
|
+
## Reranking and Adjustments
|
|
445
|
+
|
|
446
|
+
QMD's reranker judges direct query usefulness. It is not a candidate generator and cannot create evidence.
|
|
447
|
+
|
|
448
|
+
After QMD ranking, pi-llm-wiki may apply only bounded, explainable adjustments for:
|
|
449
|
+
|
|
450
|
+
- exact title, alias, command, or identifier match
|
|
451
|
+
- stable versus draft canonical status
|
|
452
|
+
- current human verification (`+0.02` within the existing combined clamp)
|
|
453
|
+
- applicable freshness and verification state
|
|
454
|
+
- explicit user relevance judgment
|
|
455
|
+
- weak local usage evidence
|
|
456
|
+
- project duplicate precedence
|
|
457
|
+
|
|
458
|
+
These adjustments cannot introduce a candidate that QMD did not retrieve. Trust, freshness, and popularity must not hide an otherwise relevant conflicting claim.
|
|
459
|
+
|
|
460
|
+
## Typed-Link Expansion
|
|
461
|
+
|
|
462
|
+
Graph expansion is an assembly step after strong seed retrieval.
|
|
463
|
+
|
|
464
|
+
Rules:
|
|
465
|
+
|
|
466
|
+
- at most one hop by default
|
|
467
|
+
- only `supports`, `contradicts`, `qualifies`, `supersedes`, `applies_under`, and `derived_from` can add a result
|
|
468
|
+
- a generic Markdown or `related_to` link never forces inclusion
|
|
469
|
+
- evidence expansion is capped per card
|
|
470
|
+
- high-degree pages do not receive an automatic popularity boost
|
|
471
|
+
- graph-derived items are labeled with the relation that admitted them
|
|
472
|
+
- a contradiction neighbor is kept with its seed even when diversity selection would otherwise remove it
|
|
473
|
+
|
|
474
|
+
No graph database is needed. Existing generated backlinks plus parsed typed relations are sufficient.
|
|
475
|
+
|
|
476
|
+
## Contradictions and User Resolution
|
|
477
|
+
|
|
478
|
+
Conflicting claims are preserved separately with their dates, scope, status, and evidence. Automated detection may propose a conflict, but it cannot permanently declare one without review.
|
|
479
|
+
|
|
480
|
+
When recall finds a confirmed or plausible conflict, context shows both claims and asks the user which applies. The prompt identifies exact page IDs so the response can be applied deterministically.
|
|
481
|
+
|
|
482
|
+
An unambiguous user response causes the agent to invoke:
|
|
483
|
+
|
|
484
|
+
```text
|
|
485
|
+
wiki_resolve_conflict(
|
|
486
|
+
vault: "personal" | "project",
|
|
487
|
+
selected_page_id: string,
|
|
488
|
+
other_page_id: string,
|
|
489
|
+
relation: "supersedes" | "qualifies" | "applies_under",
|
|
490
|
+
scope?: string,
|
|
491
|
+
rationale?: string,
|
|
492
|
+
replaces?: string
|
|
493
|
+
)
|
|
494
|
+
```
|
|
495
|
+
|
|
496
|
+
`selected_page_id` is the user-endorsed claim. Relation direction is always selected → other. `supersedes` means selected is current and the other claim is historical; `qualifies` means both remain stable and selected narrows the other; `applies_under` requires non-empty `scope` and means selected applies in that scope while the other remains applicable outside it. The tool rejects identical page IDs, missing pages, cross-vault pairs, invalid scope use, and any `replaces` ID that is not the active resolution for the same unordered pair.
|
|
497
|
+
|
|
498
|
+
The tool writes one atomic resolution page under `wiki/analyses/` rather than editing both claim pages:
|
|
499
|
+
|
|
500
|
+
```yaml
|
|
501
|
+
type: analysis
|
|
502
|
+
category: conflict-resolution
|
|
503
|
+
status: stable
|
|
504
|
+
resolution:
|
|
505
|
+
id: <semantic-hash>
|
|
506
|
+
selected: concepts/selected
|
|
507
|
+
other: concepts/other
|
|
508
|
+
relation: supersedes
|
|
509
|
+
scope: null
|
|
510
|
+
decided_by: user
|
|
511
|
+
decided_at: <ISO timestamp>
|
|
512
|
+
replaces: null
|
|
513
|
+
```
|
|
514
|
+
|
|
515
|
+
The body records the rationale and ordinary Markdown links to both claims. The ID and filename are a hash of vault ID, ordered page IDs, relation, normalized scope, and optional replaced resolution. Repeating the same operation is an idempotent no-op that returns the existing page. A changed choice requires `replaces` and creates a new immutable resolution record; recall follows the active replacement chain. A single temporary-file write, validation, and atomic rename commits the resolution. Metadata and QMD update occur afterward; failure leaves the valid resolution committed and reports stale derived state.
|
|
516
|
+
|
|
517
|
+
The resolution page is authoritative evidence of the user's choice. For `supersedes`, recall treats the other claim as historical without rewriting its `status`; for the other relations, both remain stable. The derived relation graph and feedback projection mark the selected interpretation as current for the specified scope. Pi and MCP expose identical parameters, validation codes, and structured results.
|
|
518
|
+
|
|
519
|
+
If the response is ambiguous, the system asks one clarifying question and makes no change. An LLM may map natural language to the explicit page IDs and relation, but it may not invent a third claim or silently edit unrelated content.
|
|
520
|
+
|
|
521
|
+
## Context Packing
|
|
522
|
+
|
|
523
|
+
Recall returns grouped memory bundles rather than a flat list of chunks.
|
|
524
|
+
|
|
525
|
+
Packing order:
|
|
526
|
+
|
|
527
|
+
1. highest final-score applicable canonical card, using the documented tie-breaks
|
|
528
|
+
2. at most two supporting evidence excerpts, ordered by typed relation and QMD rank
|
|
529
|
+
3. relevant qualifier, superseding claim, or competing claim
|
|
530
|
+
4. additional canonical cards under the `(type, domain)` diversity cap while the 8,000-character automatic budget remains
|
|
531
|
+
5. unpromoted evidence only when no suitable card covers the need
|
|
532
|
+
|
|
533
|
+
Every packed item includes:
|
|
534
|
+
|
|
535
|
+
- page ID and readable path
|
|
536
|
+
- title and type
|
|
537
|
+
- vault label
|
|
538
|
+
- exact excerpt and heading when available
|
|
539
|
+
- date, status, freshness, and scope when present
|
|
540
|
+
- match reason
|
|
541
|
+
- relationship to the canonical card
|
|
542
|
+
|
|
543
|
+
Deduplication prevents several chunks from one page or repeated observations from consuming the context budget. Contradictory claims remain grouped. Context truncation removes lower-value bundles before removing evidence from the highest-value bundle.
|
|
544
|
+
|
|
545
|
+
The existing links-first behavior remains useful for interactive expansion, but automatic AI context receives the compact canonical/evidence bundle directly when it fits the configured budget.
|
|
546
|
+
|
|
547
|
+
## Feedback
|
|
548
|
+
|
|
549
|
+
### Durable event format
|
|
550
|
+
|
|
551
|
+
Recall feedback is recorded through internal `recall_feedback` entries in the existing append-only `meta/events.jsonl`. To reduce accidental prompt retention, events store a SHA-256 hash of the normalized query rather than raw query text.
|
|
552
|
+
|
|
553
|
+
Each event includes:
|
|
554
|
+
|
|
555
|
+
- `schema: 1`
|
|
556
|
+
- timestamp
|
|
557
|
+
- `visibility: internal`
|
|
558
|
+
- stable vault ID and page ID
|
|
559
|
+
- query hash
|
|
560
|
+
- retrieval mode
|
|
561
|
+
- index and model versions
|
|
562
|
+
- shown rank
|
|
563
|
+
- action
|
|
564
|
+
|
|
565
|
+
Supported actions are `shown`, `opened`, `cited`, `shown_only`, `relevant`, `irrelevant`, `corrected`, and `conflict_selected`.
|
|
566
|
+
|
|
567
|
+
Instrumentation is exact:
|
|
568
|
+
|
|
569
|
+
- the recall service emits `shown` after context or links are successfully delivered
|
|
570
|
+
- the tool-call hook emits `opened` when a later `read` path resolves to a result shown in the active turn
|
|
571
|
+
- the turn-end hook emits `cited` when the assistant output contains the shown page's wikilink or ID
|
|
572
|
+
- the turn-end hook emits `shown_only` when no open or citation was observed; this is only an engagement proxy, not proof of irrelevance
|
|
573
|
+
- the agent calls `wiki_recall_feedback(vault_id, page_id, judgment, correction?)` for an explicit `relevant`, `irrelevant`, or `corrected` user statement; the pair must exist in the active session's shown-result set, `correction` is required only for `corrected`, and the service obtains the query hash from that set rather than accepting arbitrary query text
|
|
574
|
+
- `wiki_resolve_conflict` emits `conflict_selected`
|
|
575
|
+
|
|
576
|
+
For `corrected`, `wiki_recall_feedback` first creates a normal human-readable `source` observation containing the user's correction and a Markdown link to the corrected page, using the existing observation producer and parser validation. Its frontmatter includes `status: observation`, `category: recall-correction`, `corrected_page`, `vault_id`, and `query_hash`; this shape is how projection replay identifies it. It then appends the internal feedback event. If observation creation fails, neither event nor ranking adjustment is written. If the later event append fails, the correction page remains authoritative and the tool reports `correction_saved_feedback_pending`; a metadata rebuild can replay such correction observations into the feedback projection. `relevant` and `irrelevant` write only internal events.
|
|
577
|
+
|
|
578
|
+
`wiki_resolve_conflict` acquires the event lock before writing its resolution page, appends `conflict_selected` before releasing the lock, and aborts without a page if the lock cannot be acquired. This keeps the durable resolution and its feedback event aligned.
|
|
579
|
+
|
|
580
|
+
Copy and dwell-time signals are unavailable in Pi's current hooks and are not claimed. Explicit corrections are therefore preserved in human-readable wiki knowledge or resolution records; the event stream is not their only source of truth.
|
|
581
|
+
|
|
582
|
+
### Signal strength
|
|
583
|
+
|
|
584
|
+
- explicit correction or conflict choice: strong
|
|
585
|
+
- explicit relevant/irrelevant judgment: strong
|
|
586
|
+
- cited result: moderate
|
|
587
|
+
- opened result: weak positive
|
|
588
|
+
- `shown_only`: at most `-0.005` before decay
|
|
589
|
+
- result not shown: no signal
|
|
590
|
+
|
|
591
|
+
All boosts decay with a 90-day half-life, remain bounded to `±0.02` in final ranking, and apply only after retrieval. Implicit signals cannot rewrite Markdown, resolve contradictions, or make a non-candidate appear.
|
|
592
|
+
|
|
593
|
+
Feedback collection is local and enabled by default in the new major. A single boolean setting disables `shown`, `opened`, `cited`, and `shown_only` capture without disabling explicit corrections or conflict records.
|
|
594
|
+
|
|
595
|
+
## Error Handling
|
|
596
|
+
|
|
597
|
+
| Failure | Behavior |
|
|
598
|
+
|---|---|
|
|
599
|
+
| vectors missing or stale | continue with BM25; report stale vector status |
|
|
600
|
+
| embedding model load fails | fall back to lexical retrieval |
|
|
601
|
+
| reranker fails or times out | return fused hybrid results |
|
|
602
|
+
| QMD store cannot open | automatic recall injects nothing; explicit recall returns a structured diagnostic |
|
|
603
|
+
| incremental update fails | retain previous usable index and mark it stale |
|
|
604
|
+
| full rebuild fails | keep previous database; remove temporary replacement |
|
|
605
|
+
| malformed Markdown | exclude page from new results and report shared parser diagnostic |
|
|
606
|
+
| index/model mismatch | mark stale and recommend `wiki_reindex` |
|
|
607
|
+
| deleted page referenced by feedback | ignore feedback entry during aggregation |
|
|
608
|
+
| low confidence | return no reliable memory rather than weak results |
|
|
609
|
+
|
|
610
|
+
Model downloads and long-running indexing show visible progress. Cancellation stops new work without deleting the last usable index.
|
|
611
|
+
|
|
612
|
+
## Tools and Interfaces
|
|
613
|
+
|
|
614
|
+
### Updated
|
|
615
|
+
|
|
616
|
+
- `wiki_recall` uses the shared QMD-backed recall service.
|
|
617
|
+
- automatic `before_agent_start` recall uses the same service with precision-first policy.
|
|
618
|
+
- MCP `wiki_recall` uses the same structured operation.
|
|
619
|
+
- `wiki_status` reports QMD document, chunk, embedding, model, and stale-index state.
|
|
620
|
+
- `wiki_lint` reports missing evidence, invalid typed relations, unresolved conflict markers, and stale search state.
|
|
621
|
+
- `wiki_rebuild_meta` schedules incremental QMD update after successful projection rebuild.
|
|
622
|
+
|
|
623
|
+
### Added
|
|
624
|
+
|
|
625
|
+
- `wiki_reindex` consolidates lexical and vector reindexing.
|
|
626
|
+
- `wiki_resolve_conflict` records a user-approved resolution as one immutable analysis page.
|
|
627
|
+
- `wiki_recall_feedback` records explicit relevance, irrelevance, or correction judgments for an already shown `(vault_id, page_id)` result.
|
|
628
|
+
|
|
629
|
+
### Deprecated
|
|
630
|
+
|
|
631
|
+
- `wiki_reindex_embeddings` delegates to `wiki_reindex` for one major release cycle.
|
|
632
|
+
- the old heuristic recall and page-level embedding scorer are removed from active `wiki_recall` paths.
|
|
633
|
+
- `wiki_search` remains a fast exact registry lookup and is not presented as relevance-ranked recall.
|
|
634
|
+
|
|
635
|
+
## Configuration
|
|
636
|
+
|
|
637
|
+
First-release configuration remains narrow:
|
|
638
|
+
|
|
639
|
+
| Setting | Default | Meaning |
|
|
640
|
+
|---|---:|---|
|
|
641
|
+
| `retrievalMode` | `adaptive` | `lexical`, `hybrid`, `adaptive`, or `quality` |
|
|
642
|
+
| `recallFeedback` | `true` | capture bounded local implicit signals |
|
|
643
|
+
| `recallLinksThreshold` | existing default | switch interactive rendering to links-first |
|
|
644
|
+
| `recallSkillInlineMax` | existing default | inline recalled skills/cases |
|
|
645
|
+
|
|
646
|
+
QMD model overrides use QMD's documented environment variables. Candidate counts, adaptive thresholds, fusion constants, and feedback weights remain implementation constants until benchmark evidence justifies exposing them.
|
|
647
|
+
|
|
648
|
+
## Migration and Release
|
|
649
|
+
|
|
650
|
+
The next major release:
|
|
651
|
+
|
|
652
|
+
1. raises `engines.node` to Node.js 22 or newer
|
|
653
|
+
2. pins `@tobilu/qmd` to the exact published, contract-tested version `2.5.3`, raises the development TypeScript version to satisfy QMD's declared peer range, and requires adapter tests plus benchmark comparison before any QMD upgrade
|
|
654
|
+
3. supports the QMD package's tested native targets: Linux x64/arm64, macOS x64/arm64, and Windows x64; release CI performs clean-install smoke tests on Linux x64, macOS arm64, and Windows x64
|
|
655
|
+
4. treats native dependency installation failure as package installation failure with a clear supported-platform message; there is no runtime shim for an installation that never completed
|
|
656
|
+
5. records QMD schema version and resolved embedding, expansion, and reranker model IDs in index status
|
|
657
|
+
6. uses QMD's standard model cache and documents the approximately 2 GB first-use download for default embedding, reranking, and expansion models
|
|
658
|
+
7. verifies that `searchLex` indexes and queries without downloading or loading model files
|
|
659
|
+
8. creates QMD stores lazily per vault
|
|
660
|
+
9. keeps Markdown and existing metadata schemas readable without content migration, except for one stable `vault_id` backfill in config
|
|
661
|
+
10. ignores old page-level embedding sidecars after QMD activation
|
|
662
|
+
11. prompts users to run `wiki_reindex` for full hybrid/quality recall
|
|
663
|
+
12. supports immediate lexical recall after validated document indexing, even before embeddings finish
|
|
664
|
+
13. leaves the previous major release available for Node.js 18 users
|
|
665
|
+
|
|
666
|
+
No existing source or canonical page is deleted or rewritten merely to adopt QMD. Typed-link enrichment remains incremental and reviewable. Automated card promotion requires a separate future design.
|
|
667
|
+
|
|
668
|
+
The approved scope spans runtime migration, validated indexing, retrieval, memory assembly, feedback, and conflict resolution. It therefore requires a multi-phase implementation roadmap rather than one monolithic implementation plan.
|
|
669
|
+
|
|
670
|
+
## Evaluation
|
|
671
|
+
|
|
672
|
+
### Benchmark
|
|
673
|
+
|
|
674
|
+
Create a versioned benchmark from 50–100 real queries, with at least 20% held out from tuning. Include:
|
|
675
|
+
|
|
676
|
+
- exact note lookup
|
|
677
|
+
- vague recollection
|
|
678
|
+
- paraphrased conceptual recall
|
|
679
|
+
- entity and alias lookup
|
|
680
|
+
- “what did I conclude about” questions
|
|
681
|
+
- source and evidence requests
|
|
682
|
+
- temporal questions
|
|
683
|
+
- contradictory claims
|
|
684
|
+
- multi-note synthesis
|
|
685
|
+
- multilingual queries represented in the vault
|
|
686
|
+
- generic language that may produce dense-retrieval false positives
|
|
687
|
+
|
|
688
|
+
Each query receives graded relevance judgments for both canonical cards and evidence excerpts. Acceptable competing claims are identified for contradiction cases.
|
|
689
|
+
|
|
690
|
+
### Metrics
|
|
691
|
+
|
|
692
|
+
- candidate Recall@20
|
|
693
|
+
- MRR
|
|
694
|
+
- nDCG@5 and nDCG@10
|
|
695
|
+
- canonical-card-first success rate
|
|
696
|
+
- evidence precision and evidence recall
|
|
697
|
+
- contradiction coverage
|
|
698
|
+
- duplicate/context waste
|
|
699
|
+
- automatic-recall false-positive rate
|
|
700
|
+
- warm and cold latency by mode
|
|
701
|
+
- model download and steady-state resource cost
|
|
702
|
+
|
|
703
|
+
### Ablations
|
|
704
|
+
|
|
705
|
+
Run the same benchmark against:
|
|
706
|
+
|
|
707
|
+
1. current heuristic recall baseline
|
|
708
|
+
2. QMD lexical
|
|
709
|
+
3. QMD hybrid
|
|
710
|
+
4. adaptive reranking
|
|
711
|
+
5. quality mode
|
|
712
|
+
6. quality mode without typed-link assembly
|
|
713
|
+
7. quality mode without feedback adjustments
|
|
714
|
+
|
|
715
|
+
### Release gates
|
|
716
|
+
|
|
717
|
+
- every exact identifier/title benchmark query that the baseline places in the top three remains in the top three
|
|
718
|
+
- held-out nDCG@10 improves by at least 10% relative to the current baseline in `quality` mode
|
|
719
|
+
- held-out MRR does not decline by more than 2% in any mode intended to supersede the baseline
|
|
720
|
+
- automatic-recall false-positive rate falls by at least 25% relative to baseline
|
|
721
|
+
- contradiction coverage is 100% on judged conflict cases
|
|
722
|
+
- candidate Recall@20 does not decline by more than two percentage points
|
|
723
|
+
- malformed pages and unavailable models degrade as specified
|
|
724
|
+
- Pi and MCP return equivalent structured results
|
|
725
|
+
- every explicit correction becomes a persistent regression case
|
|
726
|
+
|
|
727
|
+
The first benchmark run records confidence intervals and hardware context. Constants may be tightened before implementation, but the release cannot weaken these gates without a new reviewed design decision.
|
|
728
|
+
|
|
729
|
+
## Testing Strategy
|
|
730
|
+
|
|
731
|
+
### Unit tests
|
|
732
|
+
|
|
733
|
+
- retrieval-mode parsing and fail-closed invalid values
|
|
734
|
+
- QMD adapter request mapping
|
|
735
|
+
- canonical/evidence classification
|
|
736
|
+
- layered rank fusion and project duplicate precedence
|
|
737
|
+
- bounded trust, freshness, and feedback adjustments
|
|
738
|
+
- typed-link admission and one-hop cap
|
|
739
|
+
- canonical/evidence bundle grouping
|
|
740
|
+
- contradiction preservation
|
|
741
|
+
- context deduplication and budget trimming
|
|
742
|
+
- query hashing and feedback aggregation
|
|
743
|
+
- stale and deleted feedback references
|
|
744
|
+
- fallback state machine
|
|
745
|
+
|
|
746
|
+
QMD adapter tests use fakes and do not load local models.
|
|
747
|
+
|
|
748
|
+
### Integration tests
|
|
749
|
+
|
|
750
|
+
- temporary Markdown vault indexed through the pinned QMD SDK
|
|
751
|
+
- incremental add, update, and delete through the validated mirror
|
|
752
|
+
- malformed or reserved pages absent from QMD candidates
|
|
753
|
+
- lexical recall before and without model downloads
|
|
754
|
+
- forced vector reindex
|
|
755
|
+
- complete SQLite artifact close/swap/reopen and interrupted-swap recovery
|
|
756
|
+
- failed rebuild preserving the prior database
|
|
757
|
+
- personal and project store isolation, including duplicate page IDs and feedback
|
|
758
|
+
- Pi/MCP parity
|
|
759
|
+
- automatic recall suppressing low-confidence results at specified defaults
|
|
760
|
+
- conflict-resolution idempotency, replacement chains, and preservation of both claims
|
|
761
|
+
- `relations` parse/serialize/lint/Markdown-link coexistence
|
|
762
|
+
- internal feedback event instrumentation and replay failure behavior
|
|
763
|
+
|
|
764
|
+
Model-heavy embedding and reranking smoke tests may use a separate CI job with cached models; ordinary unit tests must remain deterministic and network-free.
|
|
765
|
+
|
|
766
|
+
### Benchmark tests
|
|
767
|
+
|
|
768
|
+
Benchmark runs are versioned artifacts, not ordinary per-commit unit tests. Release candidates run the complete benchmark and publish mode-by-mode metrics, regressions, model versions, index versions, and hardware context.
|
|
769
|
+
|
|
770
|
+
## Risks and Mitigations
|
|
771
|
+
|
|
772
|
+
- **Generated cards distort evidence:** require exact source references and review before stable status.
|
|
773
|
+
- **Dense retrieval overmatches generic prose:** preserve BM25, use RRF and reranking, and include this failure class in the benchmark.
|
|
774
|
+
- **Reranker hides minority evidence:** add contradiction neighbors after seed ranking and test contradiction coverage.
|
|
775
|
+
- **Implicit feedback creates popularity bias:** keep it weak, bounded, decayed, and post-retrieval only.
|
|
776
|
+
- **QMD/model upgrade changes ranking:** record versions, mark affected indexes stale, and rerun benchmark before release.
|
|
777
|
+
- **Native dependency or model failure:** supported-platform clean-install CI catches native packaging defects; runtime model failures fall back to QMD lexical search or no injection with clear diagnostics.
|
|
778
|
+
- **Index contains sensitive content:** keep it under protected local `meta/`, exclude it from OKF exports, and document that full-vault backups contain derived searchable text.
|
|
779
|
+
- **Graph hubs dominate:** only typed one-hop relationships can add candidates; generic links do not boost rank.
|
|
780
|
+
- **Over-atomization harms browsing:** atomicity follows useful idea boundaries, not arbitrary size limits.
|
|
781
|
+
- **Stale sidecars:** content hashes, incremental updates, status diagnostics, and explicit reindexing keep them rebuildable.
|
|
782
|
+
|
|
783
|
+
## Success Criteria
|
|
784
|
+
|
|
785
|
+
The design succeeds when:
|
|
786
|
+
|
|
787
|
+
1. unrelated automatic recall is measurably reduced
|
|
788
|
+
2. reviewed canonical cards appear before raw observations for judged queries
|
|
789
|
+
3. correct evidence accompanies the selected card
|
|
790
|
+
4. relevant conflicts appear together and remain unresolved until user input
|
|
791
|
+
5. lexical recall works without loading local models
|
|
792
|
+
6. adaptive and quality modes materially improve held-out ranking
|
|
793
|
+
7. users can inspect status and repair search with one reindex tool
|
|
794
|
+
8. feedback improves repeated use without mutating facts
|
|
795
|
+
9. Obsidian and plain Markdown workflows remain intact
|
|
796
|
+
10. all generated search state can be rebuilt from Markdown plus durable feedback events
|
|
797
|
+
|
|
798
|
+
## Research References
|
|
799
|
+
|
|
800
|
+
- QMD repository and SDK documentation: https://github.com/tobi/qmd
|
|
801
|
+
- Obsidian Graph View documentation: https://obsidian.md/help/plugins/graph
|
|
802
|
+
- Zettelkasten introduction: https://zettelkasten.de/introduction/
|
|
803
|
+
- Zettelkasten atomicity guide: https://zettelkasten.de/atomicity/guide/
|
|
804
|
+
- Reciprocal Rank Fusion, Cormack, Clarke, and Buettcher: https://doi.org/10.1145/1571941.1572114
|
|
805
|
+
- Existing pi-llm-wiki architecture: `docs/architecture.md`
|
|
806
|
+
- Existing OKF interoperability design: `docs/superpowers/specs/2026-08-02-okf-v0.2-interoperability-design.md`
|