@wei840222/qmd 2026.9.6 → 2026.9.25
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +14 -0
- package/README.md +84 -1
- package/dist/cli/build-info.json +2 -2
- package/dist/cli/qmd.js +85 -9
- package/dist/collections.js +9 -4
- package/dist/index.d.ts +10 -0
- package/dist/index.js +14 -2
- package/dist/llm.d.ts +7 -1
- package/dist/llm.js +23 -4
- package/dist/mcp/server.js +70 -6
- package/dist/metadata-filter.d.ts +74 -0
- package/dist/metadata-filter.js +279 -0
- package/dist/metadata-store.d.ts +45 -0
- package/dist/metadata-store.js +173 -0
- package/dist/metadata.d.ts +61 -0
- package/dist/metadata.js +215 -0
- package/dist/search/zh-dict.txt +3 -0
- package/dist/store.d.ts +23 -16
- package/dist/store.js +341 -161
- package/package.json +2 -2
- package/scripts/sync-zh-dict.mjs +4 -1
- package/skills/qmd/SKILL.md +11 -0
- package/skills/release/SKILL.md +0 -141
- package/skills/release/scripts/install-hooks.sh +0 -38
- package/skills/release/scripts/release-context.sh +0 -129
package/CHANGELOG.md
CHANGED
|
@@ -2,13 +2,27 @@
|
|
|
2
2
|
|
|
3
3
|
## [Unreleased]
|
|
4
4
|
|
|
5
|
+
## [2026.9.25] - 2026-09-26
|
|
6
|
+
|
|
5
7
|
### Changed
|
|
6
8
|
|
|
7
9
|
- Replace `better-sqlite3` with Node.js built-in `node:sqlite`, preventing SQLite runtime symbol collisions when QMD is embedded alongside other `node:sqlite` consumers. The minimum supported Node.js version is now 22.16.0.
|
|
10
|
+
- Upgraded `@node-rs/jieba` to `2.0.3` and synchronized bundled Traditional Chinese dictionary assets (`zh-dict.txt`) with upstream `sysprog21/zhtw-mcp`.
|
|
11
|
+
- Relocated repository release management skill to `.agents/skills/release`.
|
|
8
12
|
|
|
9
13
|
### Added
|
|
10
14
|
|
|
15
|
+
- Added Oxlint lint fence.
|
|
16
|
+
- Document metadata and metadata filtering. Markdown documents can opt into typed metadata through a namespaced frontmatter block (`qmd.metadata` with strings, numbers, booleans, or flat homogeneous arrays), and every search surface — CLI `search`/`vsearch`/`query` via `--filter <json>`, the SDK's `filter` option on `search()`/`searchLex()`/`searchVector()`, the MCP `query` tool, and HTTP `POST /query` and `/search` — accepts one shared recursive filter AST discriminated by `operator`: `and`/`or`/`not` logical groups, `eq`/`ne`/`gt`/`gte`/`lt`/`lte` comparisons, `in`/`nin`/`all` membership, and `exists` presence. Every returned result satisfies the filter (applied before RRF fusion and reranking); like collection filtering, highly selective filters remain best-effort for top-K completeness. Frontmatter stays ordinary searchable content — no chunking, embedding, snippet, or line-number changes — and documents without `qmd.metadata` behave exactly as before. JSON/SDK/MCP/HTTP results now include each document's indexed metadata, and `qmd status` reports how many documents still need metadata extraction (a normal `qmd update` backfills existing indexes).
|
|
11
17
|
- **Disable HyDE Expansion Control**: Added `--no-hyde` CLI option for `qmd query` and `qmd vsearch`, `includeHyde` parameter to SDK (`store.search`, `store.expandQuery`) and MCP `query` tool, allowing users to disable generating hypothetical document embeddings during query expansion.
|
|
18
|
+
- Added `typesafe-ai` skill to `.agents/skills/typesafe-ai`.
|
|
19
|
+
|
|
20
|
+
### Fixed
|
|
21
|
+
|
|
22
|
+
- Embedding generation and legacy fingerprint adoption now tokenize documents
|
|
23
|
+
with the store-selected embedding model instead of the global default. This
|
|
24
|
+
keeps chunk boundaries aligned with the model that creates and verifies the
|
|
25
|
+
stored vectors without initializing an unrelated provider.
|
|
12
26
|
|
|
13
27
|
## [2026.8.23-1] - 2026-08-23
|
|
14
28
|
|
package/README.md
CHANGED
|
@@ -132,7 +132,7 @@ runs in a container and a liveness probe connects from a non-loopback address.
|
|
|
132
132
|
|
|
133
133
|
The HTTP server exposes two endpoints:
|
|
134
134
|
- `POST /mcp` — MCP Streamable HTTP (JSON responses, stateless)
|
|
135
|
-
- `POST /query` (alias `/search`) — structured search without the MCP protocol
|
|
135
|
+
- `POST /query` (alias `/search`) — structured search without the MCP protocol. Accepts the same optional `filter` object as the `query` tool (invalid filters return `400`); see [Metadata Filtering](#metadata-filtering)
|
|
136
136
|
- `GET /health` — liveness check with uptime
|
|
137
137
|
|
|
138
138
|
|
|
@@ -170,8 +170,10 @@ Point any MCP client at `http://localhost:8181/mcp` to connect.
|
|
|
170
170
|
| `query` | `searches` | array | Typed sub-queries (`lex`/`vec`/`hyde`), 1–10. Mutually exclusive with `query`; exactly one is required. First gets 2x weight. |
|
|
171
171
|
| `query` | `expansion` | string | Plain-query policy: `auto` (default), `force`, or `skip`. Ignored when `searches` is used. |
|
|
172
172
|
| `query` | `collections` | string[] | Filter by collection names (OR). **Array only** — singular `collection` is silently ignored. |
|
|
173
|
+
| `query` | `filter` | object | Metadata filter (recursive `operator`-discriminated JSON AST; see [Metadata Filtering](#metadata-filtering)) |
|
|
173
174
|
| `query` | `expansionContext` | string | Additional context used only to generate `lex` / `vec` / `hyde` query expansions. |
|
|
174
175
|
| `query` | `rerankContext` | string | Additional context used only for reranking and snippet/chunk selection. |
|
|
176
|
+
| `query` | `intent` | string | Disambiguation context (alias for expansionContext & rerankContext; does not search on its own) |
|
|
175
177
|
| `query` | `limit` | number | Max results (default 10) |
|
|
176
178
|
| `query` | `minScore` | number | Minimum relevance 0–1 (default 0) |
|
|
177
179
|
| `query` | `candidateLimit` | number | Max candidates to rerank (default 40) |
|
|
@@ -279,6 +281,20 @@ const results3 = await store.search({
|
|
|
279
281
|
|
|
280
282
|
// Skip reranking for faster results
|
|
281
283
|
const fast = await store.search({ query: "auth", rerank: false })
|
|
284
|
+
|
|
285
|
+
// Metadata filter — every returned result satisfies it (also available on
|
|
286
|
+
// searchLex() and searchVector()); results expose indexed metadata via
|
|
287
|
+
// r.metadata. See "Metadata Filtering" for the full grammar.
|
|
288
|
+
const published = await store.search({
|
|
289
|
+
query: "authentication flow",
|
|
290
|
+
filter: {
|
|
291
|
+
operator: "and",
|
|
292
|
+
operands: [
|
|
293
|
+
{ key: "topics", operator: "all", value: ["typescript"] },
|
|
294
|
+
{ key: "status", operator: "ne", value: "draft" },
|
|
295
|
+
],
|
|
296
|
+
},
|
|
297
|
+
})
|
|
282
298
|
```
|
|
283
299
|
|
|
284
300
|
For simple queries, explicit `force` or `skip` overrides `auto`. Under `auto`, CJK
|
|
@@ -979,6 +995,7 @@ and `deep-search` (→ `query`).
|
|
|
979
995
|
--full # Show full document content
|
|
980
996
|
--line-numbers # Add line numbers to output
|
|
981
997
|
--explain # Include retrieval score traces (query, JSON/CLI output)
|
|
998
|
+
--filter <json> # Metadata filter (recursive JSON AST; see Metadata Filtering)
|
|
982
999
|
--index <name> # Use named index
|
|
983
1000
|
--intent "<text>" # Legacy CLI alias for rerank context (e.g. "web page load times")
|
|
984
1001
|
--no-rerank # Skip LLM reranking (RRF scores only; faster on CPU)
|
|
@@ -1022,6 +1039,72 @@ explicitly with `-c`.
|
|
|
1022
1039
|
> lexical and vector candidate cutoff. Matching candidates from the selected
|
|
1023
1040
|
> collections are then ranked together.
|
|
1024
1041
|
|
|
1042
|
+
### Metadata Filtering
|
|
1043
|
+
|
|
1044
|
+
Documents can opt into typed metadata through a namespaced frontmatter block. A document without `qmd.metadata` behaves exactly as before, and the frontmatter stays ordinary searchable content (no chunking, embedding, or line-number changes):
|
|
1045
|
+
|
|
1046
|
+
```markdown
|
|
1047
|
+
---
|
|
1048
|
+
qmd:
|
|
1049
|
+
metadata:
|
|
1050
|
+
topics:
|
|
1051
|
+
- typescript
|
|
1052
|
+
- programming
|
|
1053
|
+
status: published
|
|
1054
|
+
priority: 3
|
|
1055
|
+
reviewed: true
|
|
1056
|
+
---
|
|
1057
|
+
|
|
1058
|
+
# Document body starts here
|
|
1059
|
+
```
|
|
1060
|
+
|
|
1061
|
+
Supported values are strings, numbers, booleans, and flat homogeneous arrays of one of those. Nested objects, nulls, empty arrays, and mixed-type arrays are rejected (the document still indexes; it is excluded from filtered search until corrected). Metadata keys are user-defined data — `tags`, `topics`, and `labels` are all ordinary keys with no special semantics.
|
|
1062
|
+
|
|
1063
|
+
Every search surface (CLI, SDK, MCP, HTTP) accepts the same recursive filter, a JSON AST discriminated by `operator`:
|
|
1064
|
+
|
|
1065
|
+
```sh
|
|
1066
|
+
# One condition
|
|
1067
|
+
qmd search "authentication" \
|
|
1068
|
+
--filter '{"key":"status","operator":"eq","value":"published"}'
|
|
1069
|
+
|
|
1070
|
+
# Composed conditions — works with search, vsearch, and query
|
|
1071
|
+
qmd query "dependency injection" --filter '{
|
|
1072
|
+
"operator": "and",
|
|
1073
|
+
"operands": [
|
|
1074
|
+
{ "key": "topics", "operator": "all", "value": ["typescript", "programming"] },
|
|
1075
|
+
{ "key": "status", "operator": "nin", "value": ["draft", "archived"] },
|
|
1076
|
+
{ "operator": "or", "operands": [
|
|
1077
|
+
{ "key": "priority", "operator": "gte", "value": 3 },
|
|
1078
|
+
{ "key": "reviewed", "operator": "eq", "value": true }
|
|
1079
|
+
] },
|
|
1080
|
+
{ "operator": "not", "operand": { "key": "audience", "operator": "eq", "value": "internal" } }
|
|
1081
|
+
]
|
|
1082
|
+
}'
|
|
1083
|
+
```
|
|
1084
|
+
|
|
1085
|
+
| Node | Shape |
|
|
1086
|
+
|------|-------|
|
|
1087
|
+
| Logical group | `{ "operator": "and" \| "or", "operands": […] }` |
|
|
1088
|
+
| Negation | `{ "operator": "not", "operand": {…} }` |
|
|
1089
|
+
| Comparison | `{ "key", "operator": "eq" \| "ne" \| "gt" \| "gte" \| "lt" \| "lte", "value" }` |
|
|
1090
|
+
| Membership | `{ "key", "operator": "in" \| "nin" \| "all", "value": […] }` |
|
|
1091
|
+
| Presence | `{ "key", "operator": "exists", "value": true \| false }` |
|
|
1092
|
+
|
|
1093
|
+
Semantics:
|
|
1094
|
+
|
|
1095
|
+
- Matching is typed and exact — no string/number/boolean coercion, and a type mismatch never matches (including `ne` and `nin`).
|
|
1096
|
+
- Array-valued metadata is a set: a condition matches when any element satisfies it, `all` requires every filter value to be present.
|
|
1097
|
+
- Missing keys do not match `ne`/`nin`; combine with `{ "operator": "exists", "value": false }` in an `or` group to include them.
|
|
1098
|
+
- Multiple conditions require an explicit `and` group — there is no implicit AND, and no `$`-prefixed shorthand.
|
|
1099
|
+
|
|
1100
|
+
Guarantees and limits:
|
|
1101
|
+
|
|
1102
|
+
- Every returned result satisfies the filter, before RRF fusion and reranking.
|
|
1103
|
+
- Like collection filtering, highly selective filters are best-effort for top-K completeness: backends over-fetch and post-filter, so a very selective filter can return fewer than `limit` results.
|
|
1104
|
+
- Filtered search only considers documents whose metadata has been extracted (run `qmd update` after upgrading; `qmd status` shows the pending count).
|
|
1105
|
+
|
|
1106
|
+
JSON output (`--format json`), the SDK, MCP structured results, and the HTTP endpoints include each result's indexed metadata.
|
|
1107
|
+
|
|
1025
1108
|
### Output Format
|
|
1026
1109
|
|
|
1027
1110
|
Default output is colorized CLI format (respects `NO_COLOR` env).
|
package/dist/cli/build-info.json
CHANGED
package/dist/cli/qmd.js
CHANGED
|
@@ -10,6 +10,8 @@ import { parseArgs } from "util";
|
|
|
10
10
|
import { readFileSync, readdirSync, realpathSync, statSync, existsSync, unlinkSync, writeFileSync, openSync, closeSync, mkdirSync, lstatSync, rmSync, symlinkSync, readlinkSync, copyFileSync } from "fs";
|
|
11
11
|
import { createInterface } from "readline/promises";
|
|
12
12
|
import { getPwd, getRealPath, isPathInsideDir, homedir, resolve, enableProductionMode, searchFTS, extractSnippet, getContextForFile, getContextForPath, listCollections, findSimilarFiles, findDocument, resolveCommaListName, matchFilesByGlob, getHashesNeedingEmbedding, clearAllEmbeddings, insertEmbedding, getStatus, hashContent, extractTitle, formatDocForEmbedding, getEmbeddingFingerprint, chunkDocumentByTokens, clearCache, getCacheKey, getCachedResult, setCachedResult, getIndexHealth, parseVirtualPath, buildVirtualPath, isVirtualPath, isDocid, resolveVirtualPath, toVirtualPath, insertContent, insertDocument, insertDocumentWithContent, findActiveDocument, findOrMigrateLegacyDocument, updateDocumentTitle, updateDocument, updateDocumentWithContent, deactivateDocument, getActiveDocumentPaths, cleanupOrphanedContent, countOrphanedVectors, previewCleanup, runCleanup, getCollectionsWithoutContext, getTopLevelPathsWithoutContext, handelize, escapeLikePattern, hybridQuery, vectorSearchQuery, structuredSearch, addLineNumbers, DEFAULT_EMBED_MODEL, DEFAULT_EMBED_MAX_BATCH_BYTES, DEFAULT_EMBED_MAX_DOCS_PER_BATCH, DEFAULT_RERANK_MODEL, DEFAULT_QUERY_MODEL, DEFAULT_GLOB, splitGlobMask, DEFAULT_MULTI_GET_MAX_BYTES, createStore, getDefaultDbPath, reindexCollection, generateEmbeddings, getPendingEmbeddingDocsReadOnly, syncConfigToDb, } from "../store.js";
|
|
13
|
+
import { syncDocumentMetadata, countDocumentsPendingMetadata } from "../metadata-store.js";
|
|
14
|
+
import { parseMetadataFilter } from "../metadata-filter.js";
|
|
13
15
|
import { disposeDefaultLlamaCpp, getDefaultLlamaCpp, setDefaultLlamaCpp, LlamaCpp, withLLMSession, pullModels, DEFAULT_MODEL_CACHE_DIR, resolveEmbedModel, resolveGenerateModel, resolveRerankModel, resolveModels, inspectGgufFile, isDarwinMetalMitigationActive } from "../llm.js";
|
|
14
16
|
import { rebuildCjkLexicalIndex } from "../search/cjk-index.js";
|
|
15
17
|
import { RemoteLLM } from "../remote-llm.js";
|
|
@@ -219,11 +221,16 @@ function mcpDaemonPaths() {
|
|
|
219
221
|
}
|
|
220
222
|
function setIndexName(name) {
|
|
221
223
|
let normalizedName = name;
|
|
222
|
-
// Normalize relative paths to prevent malformed database paths
|
|
223
|
-
|
|
224
|
-
|
|
225
|
-
|
|
226
|
-
|
|
224
|
+
// Normalize relative paths to prevent malformed database paths. Windows
|
|
225
|
+
// absolute paths (C:\..., \\server\share) are already absolute -- skip
|
|
226
|
+
// pathResolve() for those and sanitize `:` and `\` alongside `/`.
|
|
227
|
+
if (name) {
|
|
228
|
+
const isWindowsAbsolute = /^[a-zA-Z]:[\\/]/.test(name) || /^\\/.test(name);
|
|
229
|
+
if (isWindowsAbsolute || /[\\/]/.test(name)) {
|
|
230
|
+
const absolutePath = isWindowsAbsolute ? name : pathResolve(process.cwd(), name);
|
|
231
|
+
// Replace path separators with underscores to create a valid filename
|
|
232
|
+
normalizedName = absolutePath.replace(/[:\\/]+/g, '_').replace(/^_+/, '');
|
|
233
|
+
}
|
|
227
234
|
}
|
|
228
235
|
currentIndexName = normalizedName || "index";
|
|
229
236
|
storeDbPathOverride = normalizedName ? getDefaultDbPath(normalizedName) : undefined;
|
|
@@ -546,6 +553,10 @@ async function showStatus() {
|
|
|
546
553
|
if (needsEmbedding > 0) {
|
|
547
554
|
console.log(` ${c.yellow}Pending: ${needsEmbedding} need embedding${c.reset} (run 'qmd embed')`);
|
|
548
555
|
}
|
|
556
|
+
const pendingMetadata = countDocumentsPendingMetadata(db);
|
|
557
|
+
if (pendingMetadata > 0) {
|
|
558
|
+
console.log(` ${c.yellow}Metadata: ${pendingMetadata} need extraction${c.reset} (run 'qmd update'; excluded from --filter searches)`);
|
|
559
|
+
}
|
|
549
560
|
if (mostRecent.latest) {
|
|
550
561
|
const lastUpdate = new Date(mostRecent.latest);
|
|
551
562
|
console.log(` Updated: ${formatTimeAgo(lastUpdate)}`);
|
|
@@ -954,6 +965,7 @@ async function updateCollections() {
|
|
|
954
965
|
progress.clear();
|
|
955
966
|
console.log(`\nIndexed: ${result.indexed} new, ${result.updated} updated, ${result.unchanged} unchanged, ${result.removed} removed`);
|
|
956
967
|
reportSkippedReads(result.skippedFiles);
|
|
968
|
+
reportMetadataErrors(result.metadataErrors);
|
|
957
969
|
if (result.orphanedCleaned > 0) {
|
|
958
970
|
console.log(`Cleaned up ${result.orphanedCleaned} orphaned content hash(es)`);
|
|
959
971
|
}
|
|
@@ -1858,7 +1870,7 @@ async function indexFiles(pwd, globPattern = DEFAULT_GLOB, collectionName, suppr
|
|
|
1858
1870
|
console.log("No files found matching pattern.");
|
|
1859
1871
|
// Continue so the deactivation pass can mark previously indexed docs as inactive.
|
|
1860
1872
|
}
|
|
1861
|
-
let indexed = 0, updated = 0, unchanged = 0, processed = 0;
|
|
1873
|
+
let indexed = 0, updated = 0, unchanged = 0, processed = 0, metadataErrors = 0;
|
|
1862
1874
|
const skippedFiles = [];
|
|
1863
1875
|
const seenPaths = new Set();
|
|
1864
1876
|
// Literal paths of every file in this scan. Passed to the legacy-path
|
|
@@ -1896,8 +1908,12 @@ async function indexFiles(pwd, globPattern = DEFAULT_GLOB, collectionName, suppr
|
|
|
1896
1908
|
const title = extractTitle(content, relativeFile);
|
|
1897
1909
|
// Check if document exists (also migrates legacy lowercase paths)
|
|
1898
1910
|
const existing = findOrMigrateLegacyDocument(db, collectionName, path, livePaths);
|
|
1911
|
+
let documentId;
|
|
1912
|
+
let contentChanged = true;
|
|
1899
1913
|
if (existing) {
|
|
1914
|
+
documentId = existing.id;
|
|
1900
1915
|
if (existing.hash === hash) {
|
|
1916
|
+
contentChanged = false;
|
|
1901
1917
|
// Hash unchanged, but check if title needs updating
|
|
1902
1918
|
if (existing.title !== title) {
|
|
1903
1919
|
updateDocumentTitle(db, existing.id, title, now);
|
|
@@ -1918,8 +1934,12 @@ async function indexFiles(pwd, globPattern = DEFAULT_GLOB, collectionName, suppr
|
|
|
1918
1934
|
// New document - insert content and document
|
|
1919
1935
|
indexed++;
|
|
1920
1936
|
const stat = statSync(filepath);
|
|
1921
|
-
insertDocumentWithContent(db, hash, content, now, collectionName, path, title, stat ? new Date(stat.birthtime).toISOString() : now, stat ? new Date(stat.mtime).toISOString() : now);
|
|
1937
|
+
documentId = insertDocumentWithContent(db, hash, content, now, collectionName, path, title, stat ? new Date(stat.birthtime).toISOString() : now, stat ? new Date(stat.mtime).toISOString() : now);
|
|
1922
1938
|
}
|
|
1939
|
+
// Unchanged content still backfills missing or stale extraction state.
|
|
1940
|
+
const extraction = syncDocumentMetadata(db, documentId, content, path, contentChanged ? undefined : { onlyIfStale: true });
|
|
1941
|
+
if (extraction?.error)
|
|
1942
|
+
metadataErrors++;
|
|
1923
1943
|
processed++;
|
|
1924
1944
|
progress.set((processed / total) * 100);
|
|
1925
1945
|
const elapsed = (Date.now() - startTime) / 1000;
|
|
@@ -1945,6 +1965,7 @@ async function indexFiles(pwd, globPattern = DEFAULT_GLOB, collectionName, suppr
|
|
|
1945
1965
|
progress.clear();
|
|
1946
1966
|
console.log(`\nIndexed: ${indexed} new, ${updated} updated, ${unchanged} unchanged, ${removed} removed`);
|
|
1947
1967
|
reportSkippedReads(skippedFiles);
|
|
1968
|
+
reportMetadataErrors(metadataErrors);
|
|
1948
1969
|
if (orphanedContent > 0) {
|
|
1949
1970
|
console.log(`Cleaned up ${orphanedContent} orphaned content hash(es)`);
|
|
1950
1971
|
}
|
|
@@ -1962,6 +1983,11 @@ function fsErrorCode(err) {
|
|
|
1962
1983
|
}
|
|
1963
1984
|
return "ERROR";
|
|
1964
1985
|
}
|
|
1986
|
+
function reportMetadataErrors(metadataErrors) {
|
|
1987
|
+
if (metadataErrors === 0)
|
|
1988
|
+
return;
|
|
1989
|
+
console.warn(`⚠ ${metadataErrors} file(s) have invalid qmd.metadata frontmatter and are excluded from filtered search`);
|
|
1990
|
+
}
|
|
1965
1991
|
function reportSkippedReads(skippedFiles) {
|
|
1966
1992
|
if (skippedFiles.length === 0)
|
|
1967
1993
|
return;
|
|
@@ -2375,6 +2401,7 @@ function outputResults(results, query, opts) {
|
|
|
2375
2401
|
line: snippetInfo.line,
|
|
2376
2402
|
title: row.title,
|
|
2377
2403
|
...(row.context && { context: row.context }),
|
|
2404
|
+
...(row.metadata && Object.keys(row.metadata).length > 0 && { metadata: row.metadata }),
|
|
2378
2405
|
...(body && { body }),
|
|
2379
2406
|
...(snippet && { snippet }),
|
|
2380
2407
|
...(opts.explain && row.explain && { explain: row.explain }),
|
|
@@ -2624,14 +2651,46 @@ export function parseStructuredQuery(query) {
|
|
|
2624
2651
|
}
|
|
2625
2652
|
return typed.length > 0 ? { searches: typed, intent } : null;
|
|
2626
2653
|
}
|
|
2654
|
+
// Parse and validate a --filter JSON string; exits with an actionable
|
|
2655
|
+
// message on malformed JSON or an invalid filter AST.
|
|
2656
|
+
function parseCliMetadataFilter(rawFilter) {
|
|
2657
|
+
if (rawFilter === undefined)
|
|
2658
|
+
return undefined;
|
|
2659
|
+
let filterJson;
|
|
2660
|
+
try {
|
|
2661
|
+
filterJson = JSON.parse(String(rawFilter));
|
|
2662
|
+
}
|
|
2663
|
+
catch (err) {
|
|
2664
|
+
console.error(`Invalid --filter JSON: ${err instanceof Error ? err.message : String(err)}`);
|
|
2665
|
+
console.error(`Example: --filter '{"key":"status","operator":"eq","value":"published"}'`);
|
|
2666
|
+
process.exit(1);
|
|
2667
|
+
}
|
|
2668
|
+
try {
|
|
2669
|
+
return parseMetadataFilter(filterJson);
|
|
2670
|
+
}
|
|
2671
|
+
catch (err) {
|
|
2672
|
+
console.error(err instanceof Error ? err.message : String(err));
|
|
2673
|
+
process.exit(1);
|
|
2674
|
+
}
|
|
2675
|
+
}
|
|
2676
|
+
// Filtered search excludes documents without current metadata extraction;
|
|
2677
|
+
// tell the user when that makes results incomplete.
|
|
2678
|
+
function warnPendingMetadata(db) {
|
|
2679
|
+
const pendingMetadata = countDocumentsPendingMetadata(db);
|
|
2680
|
+
if (pendingMetadata === 0)
|
|
2681
|
+
return;
|
|
2682
|
+
process.stderr.write(`${c.yellow}Warning: ${pendingMetadata} document(s) lack current metadata extraction and are excluded from filtered results. Run 'qmd update'.${c.reset}\n`);
|
|
2683
|
+
}
|
|
2627
2684
|
function search(query, opts) {
|
|
2628
2685
|
const db = getDb();
|
|
2629
2686
|
// Validate collection filter (supports multiple -c flags)
|
|
2630
2687
|
// Use default collections if none specified
|
|
2631
2688
|
const collectionNames = resolveCollectionFilter(opts.collection, true);
|
|
2689
|
+
if (opts.filter)
|
|
2690
|
+
warnPendingMetadata(db);
|
|
2632
2691
|
// Use large limit for --all, otherwise fetch more than needed and let outputResults filter
|
|
2633
2692
|
const fetchLimit = opts.all ? 100000 : Math.max(50, opts.limit * 2);
|
|
2634
|
-
const results = searchFTS(db, query, fetchLimit, collectionNames);
|
|
2693
|
+
const results = searchFTS(db, query, fetchLimit, collectionSearchFilter(collectionNames), opts.filter);
|
|
2635
2694
|
// Add context to results
|
|
2636
2695
|
const resultsWithContext = results.map(r => ({
|
|
2637
2696
|
file: r.filepath,
|
|
@@ -2642,6 +2701,7 @@ function search(query, opts) {
|
|
|
2642
2701
|
context: getContextForFile(db, r.filepath),
|
|
2643
2702
|
hash: r.hash,
|
|
2644
2703
|
docid: r.docid,
|
|
2704
|
+
metadata: r.metadata,
|
|
2645
2705
|
}));
|
|
2646
2706
|
closeDb();
|
|
2647
2707
|
if (resultsWithContext.length === 0) {
|
|
@@ -2672,9 +2732,12 @@ async function vectorSearch(query, opts, _model = DEFAULT_EMBED_MODEL) {
|
|
|
2672
2732
|
// Use default collections if none specified
|
|
2673
2733
|
const collectionNames = resolveCollectionFilter(opts.collection, true);
|
|
2674
2734
|
checkIndexHealth(store.db);
|
|
2735
|
+
if (opts.filter)
|
|
2736
|
+
warnPendingMetadata(store.db);
|
|
2675
2737
|
await withLLMSession(async () => {
|
|
2676
2738
|
const results = await vectorSearchQuery(store, query, {
|
|
2677
|
-
collection: collectionNames,
|
|
2739
|
+
collection: collectionSearchFilter(collectionNames),
|
|
2740
|
+
filter: opts.filter,
|
|
2678
2741
|
limit: opts.all ? 500 : (opts.limit || 10),
|
|
2679
2742
|
minScore: opts.minScore || 0.3,
|
|
2680
2743
|
expansionContext: opts.intent,
|
|
@@ -2699,6 +2762,7 @@ async function vectorSearch(query, opts, _model = DEFAULT_EMBED_MODEL) {
|
|
|
2699
2762
|
score: r.score,
|
|
2700
2763
|
context: r.context,
|
|
2701
2764
|
docid: r.docid,
|
|
2765
|
+
metadata: r.metadata,
|
|
2702
2766
|
})), query, { ...opts, limit: results.length });
|
|
2703
2767
|
}, { maxDuration: 10 * 60 * 1000, name: 'vectorSearch' });
|
|
2704
2768
|
}
|
|
@@ -2708,6 +2772,8 @@ async function querySearch(query, opts, _embedModel = DEFAULT_EMBED_MODEL, _rera
|
|
|
2708
2772
|
// Use default collections if none specified
|
|
2709
2773
|
const collectionNames = resolveCollectionFilter(opts.collection, true);
|
|
2710
2774
|
checkIndexHealth(store.db);
|
|
2775
|
+
if (opts.filter)
|
|
2776
|
+
warnPendingMetadata(store.db);
|
|
2711
2777
|
// Check for structured query syntax (lex:/vec:/hyde:/intent: prefixes)
|
|
2712
2778
|
const parsed = parseStructuredQuery(query);
|
|
2713
2779
|
// Intent can come from --intent flag or from intent: line in query document
|
|
@@ -2735,6 +2801,7 @@ async function querySearch(query, opts, _embedModel = DEFAULT_EMBED_MODEL, _rera
|
|
|
2735
2801
|
process.stderr.write(`${c.dim}└─ Searching...${c.reset}\n`);
|
|
2736
2802
|
results = await structuredSearch(store, structuredQueries, {
|
|
2737
2803
|
collections: collectionNames.length > 0 ? collectionNames : undefined,
|
|
2804
|
+
filter: opts.filter,
|
|
2738
2805
|
limit: opts.all ? 500 : (opts.limit || 10),
|
|
2739
2806
|
minScore: opts.minScore || 0,
|
|
2740
2807
|
candidateLimit: opts.candidateLimit,
|
|
@@ -2764,6 +2831,8 @@ async function querySearch(query, opts, _embedModel = DEFAULT_EMBED_MODEL, _rera
|
|
|
2764
2831
|
// Standard hybrid query with automatic expansion
|
|
2765
2832
|
results = await hybridQuery(store, query, {
|
|
2766
2833
|
collections: collectionNames.length > 0 ? collectionNames : undefined,
|
|
2834
|
+
collection: collectionSearchFilter(collectionNames),
|
|
2835
|
+
filter: opts.filter,
|
|
2767
2836
|
limit: opts.all ? 500 : (opts.limit || 10),
|
|
2768
2837
|
minScore: opts.minScore || 0,
|
|
2769
2838
|
candidateLimit: opts.candidateLimit,
|
|
@@ -2830,6 +2899,7 @@ async function querySearch(query, opts, _embedModel = DEFAULT_EMBED_MODEL, _rera
|
|
|
2830
2899
|
score: r.score,
|
|
2831
2900
|
context: r.context,
|
|
2832
2901
|
docid: r.docid,
|
|
2902
|
+
metadata: r.metadata,
|
|
2833
2903
|
explain: r.explain,
|
|
2834
2904
|
})), displayQuery, { ...opts, limit: results.length });
|
|
2835
2905
|
}, { maxDuration: 10 * 60 * 1000, name: 'querySearch' });
|
|
@@ -2866,6 +2936,7 @@ function parseCLI() {
|
|
|
2866
2936
|
json: { type: "boolean" },
|
|
2867
2937
|
explain: { type: "boolean" },
|
|
2868
2938
|
collection: { type: "string", short: "c", multiple: true }, // Filter by collection(s)
|
|
2939
|
+
filter: { type: "string" }, // Metadata filter (JSON AST) for search/vsearch/query
|
|
2869
2940
|
// Collection options
|
|
2870
2941
|
name: { type: "string" }, // collection name
|
|
2871
2942
|
mask: { type: "string" }, // glob pattern
|
|
@@ -3462,6 +3533,8 @@ function showHelp() {
|
|
|
3462
3533
|
console.log(" --explain - Include retrieval score traces (query, CLI/--format json)");
|
|
3463
3534
|
console.log(" --format <kind> - Output format: cli (default) | json | csv | md | xml | files");
|
|
3464
3535
|
console.log(" -c, --collection <name> - Filter by one or more collections");
|
|
3536
|
+
console.log(" --filter <json> - Metadata filter (recursive JSON AST; search/vsearch/query)");
|
|
3537
|
+
console.log(" e.g. '{\"key\":\"status\",\"operator\":\"eq\",\"value\":\"published\"}'");
|
|
3465
3538
|
console.log("");
|
|
3466
3539
|
console.log("Embed/query options:");
|
|
3467
3540
|
console.log(" --chunk-strategy <auto|regex> - Chunking mode (default: regex; auto uses AST for code files)");
|
|
@@ -4532,6 +4605,7 @@ if (isMain) {
|
|
|
4532
4605
|
console.error("Usage: qmd search [options] <query>");
|
|
4533
4606
|
process.exit(1);
|
|
4534
4607
|
}
|
|
4608
|
+
cli.opts.filter = parseCliMetadataFilter(cli.values.filter);
|
|
4535
4609
|
search(cli.query, cli.opts);
|
|
4536
4610
|
break;
|
|
4537
4611
|
case "vsearch":
|
|
@@ -4544,6 +4618,7 @@ if (isMain) {
|
|
|
4544
4618
|
if (!cli.values["min-score"]) {
|
|
4545
4619
|
cli.opts.minScore = 0.3;
|
|
4546
4620
|
}
|
|
4621
|
+
cli.opts.filter = parseCliMetadataFilter(cli.values.filter);
|
|
4547
4622
|
await resolveLocalConfigTrust();
|
|
4548
4623
|
await vectorSearch(cli.query, cli.opts);
|
|
4549
4624
|
break;
|
|
@@ -4553,6 +4628,7 @@ if (isMain) {
|
|
|
4553
4628
|
console.error("Usage: qmd query [options] <query>");
|
|
4554
4629
|
process.exit(1);
|
|
4555
4630
|
}
|
|
4631
|
+
cli.opts.filter = parseCliMetadataFilter(cli.values.filter);
|
|
4556
4632
|
await resolveLocalConfigTrust();
|
|
4557
4633
|
await querySearch(cli.query, cli.opts);
|
|
4558
4634
|
break;
|
package/dist/collections.js
CHANGED
|
@@ -41,11 +41,16 @@ export function setConfigSource(source) {
|
|
|
41
41
|
* Config file will be ~/.config/qmd/{indexName}.yml
|
|
42
42
|
*/
|
|
43
43
|
export function setConfigIndexName(name) {
|
|
44
|
-
// Resolve relative paths to absolute paths and sanitize for use as filename
|
|
45
|
-
|
|
46
|
-
|
|
44
|
+
// Resolve relative paths to absolute paths and sanitize for use as filename.
|
|
45
|
+
// Windows absolute paths (C:\..., \\server\share) are already absolute --
|
|
46
|
+
// skip resolve() for those (POSIX path.resolve doesn't recognize a drive
|
|
47
|
+
// letter as absolute and would wrongly prefix it with cwd) and sanitize
|
|
48
|
+
// `:` and `\` alongside `/` so the derived filename is valid on Windows.
|
|
49
|
+
const isWindowsAbsolute = /^[a-zA-Z]:[\\/]/.test(name) || /^\\/.test(name);
|
|
50
|
+
if (isWindowsAbsolute || /[\\/]/.test(name)) {
|
|
51
|
+
const absolutePath = isWindowsAbsolute ? name : resolve(process.cwd(), name);
|
|
47
52
|
// Replace path separators with underscores to create a valid filename
|
|
48
|
-
currentIndexName = absolutePath.replace(
|
|
53
|
+
currentIndexName = absolutePath.replace(/[:\\/]+/g, '_').replace(/^_+/, '');
|
|
49
54
|
}
|
|
50
55
|
else {
|
|
51
56
|
currentIndexName = name;
|
package/dist/index.d.ts
CHANGED
|
@@ -17,9 +17,13 @@
|
|
|
17
17
|
* await store.close()
|
|
18
18
|
*/
|
|
19
19
|
import { extractSnippet, addLineNumbers, DEFAULT_MULTI_GET_MAX_BYTES, type Store as InternalStore, type DocumentResult, type DocumentNotFound, type DocumentExcludedByIgnore, type DocumentLookupError, type SearchResult, type HybridQueryResult, type HybridQueryOptions, type HybridQueryExplain, type ExpandedQuery, type StructuredSearchOptions, type MultiGetResult, type IndexStatus, type IndexHealthInfo, type SearchHooks, type ReindexProgress, type ReindexResult, type EmbedProgress, type EmbedResult, type ChunkStrategy } from "./store.js";
|
|
20
|
+
import type { DocumentMetadata, MetadataScalar, MetadataScalarArray, MetadataValue } from "./metadata.js";
|
|
21
|
+
import { parseMetadataFilter, MetadataFilterError, type MetadataFilter, type MetadataFilterGroup, type MetadataFilterNegation, type MetadataCondition } from "./metadata-filter.js";
|
|
20
22
|
import type { ExpansionMode } from "./search/query-expansion.js";
|
|
21
23
|
import { type Collection, type CollectionConfig, type NamedCollection, type ContextMap } from "./collections.js";
|
|
22
24
|
export type { DocumentResult, DocumentNotFound, DocumentExcludedByIgnore, DocumentLookupError, SearchResult, HybridQueryResult, HybridQueryOptions, HybridQueryExplain, ExpandedQuery, StructuredSearchOptions, MultiGetResult, IndexStatus, IndexHealthInfo, SearchHooks, ReindexProgress, ReindexResult, EmbedProgress, EmbedResult, Collection, CollectionConfig, NamedCollection, ContextMap, };
|
|
25
|
+
export type { DocumentMetadata, MetadataScalar, MetadataScalarArray, MetadataValue, MetadataFilter, MetadataFilterGroup, MetadataFilterNegation, MetadataCondition, };
|
|
26
|
+
export { parseMetadataFilter, MetadataFilterError };
|
|
23
27
|
export type { InternalStore };
|
|
24
28
|
export type { ExpansionMode } from "./search/query-expansion.js";
|
|
25
29
|
export { extractSnippet, addLineNumbers, DEFAULT_MULTI_GET_MAX_BYTES };
|
|
@@ -65,6 +69,8 @@ export interface SearchOptions {
|
|
|
65
69
|
collection?: string;
|
|
66
70
|
/** Filter to specific collections */
|
|
67
71
|
collections?: string[];
|
|
72
|
+
/** Metadata filter — every returned result satisfies it */
|
|
73
|
+
filter?: MetadataFilter;
|
|
68
74
|
/** Max results (default: 10) */
|
|
69
75
|
limit?: number;
|
|
70
76
|
/** Max candidates to rerank (default: 40) */
|
|
@@ -88,6 +94,8 @@ export interface SearchOptions {
|
|
|
88
94
|
export interface LexSearchOptions {
|
|
89
95
|
limit?: number;
|
|
90
96
|
collection?: string | string[];
|
|
97
|
+
/** Metadata filter — every returned result satisfies it */
|
|
98
|
+
filter?: MetadataFilter;
|
|
91
99
|
}
|
|
92
100
|
/**
|
|
93
101
|
* Options for searchVector() — vector similarity search.
|
|
@@ -95,6 +103,8 @@ export interface LexSearchOptions {
|
|
|
95
103
|
export interface VectorSearchOptions {
|
|
96
104
|
limit?: number;
|
|
97
105
|
collection?: string | string[];
|
|
106
|
+
/** Metadata filter — every returned result satisfies it */
|
|
107
|
+
filter?: MetadataFilter;
|
|
98
108
|
}
|
|
99
109
|
/**
|
|
100
110
|
* Options for expandQuery() — manual query expansion.
|
package/dist/index.js
CHANGED
|
@@ -20,6 +20,7 @@ import { existsSync } from "node:fs";
|
|
|
20
20
|
import { createStore as createStoreInternal, hybridQuery, structuredSearch, extractSnippet, addLineNumbers, DEFAULT_MULTI_GET_MAX_BYTES, reindexCollection, generateEmbeddings, listCollections as storeListCollections, syncConfigToDb, getStoreCollections, getStoreCollection, getStoreGlobalContext, getStoreContexts, upsertStoreCollection, removeCollection as removeCollectionWithDocuments, renameCollection as renameCollectionWithDocuments, updateStoreContext, removeStoreContext, setStoreGlobalContext, vacuumDatabase, cleanupOrphanedContent, cleanupOrphanedVectors, deleteLLMCache, deleteInactiveDocuments, clearAllEmbeddings, getPendingEmbeddingDocsReadOnly, getIndexHealthReadOnly, getStatusReadOnly, } from "./store.js";
|
|
21
21
|
import { DEFAULT_EMBED_MODEL_URI, LlamaCpp, waitForLLMSessionsToDrain, } from "./llm.js";
|
|
22
22
|
import { LocalEmbeddingProviderOwner } from "./embedding/local.js";
|
|
23
|
+
import { parseMetadataFilter, MetadataFilterError, } from "./metadata-filter.js";
|
|
23
24
|
import { OpenAIEmbeddingProvider, UnavailableOpenAIEmbeddingProvider, } from "./embedding/openai.js";
|
|
24
25
|
import { authorizeRemoteEmbeddingRequest, remoteEmbeddingIdentity, } from "./embedding/remote-embedding.js";
|
|
25
26
|
import { readStoredEmbeddingIdentity } from "./embedding/identity.js";
|
|
@@ -29,6 +30,7 @@ import { rebuildCjkLexicalIndex } from "./search/cjk-index.js";
|
|
|
29
30
|
import { RemoteLLM } from "./remote-llm.js";
|
|
30
31
|
import { HybridLLM } from "./hybrid-llm.js";
|
|
31
32
|
import { createCollectionConfigSource, loadConfig, addCollection as collectionsAddCollection, removeCollection as collectionsRemoveCollection, renameCollection as collectionsRenameCollection, addContext as collectionsAddContext, removeContext as collectionsRemoveContext, setGlobalContext as collectionsSetGlobalContext, } from "./collections.js";
|
|
33
|
+
export { parseMetadataFilter, MetadataFilterError };
|
|
32
34
|
// Re-export utility functions and types used by frontends
|
|
33
35
|
export { extractSnippet, addLineNumbers, DEFAULT_MULTI_GET_MAX_BYTES };
|
|
34
36
|
// Re-export getDefaultDbPath for CLI/MCP that need the default database location
|
|
@@ -208,10 +210,15 @@ export async function createStore(options) {
|
|
|
208
210
|
...(opts.collections ?? []),
|
|
209
211
|
];
|
|
210
212
|
const skipRerank = opts.rerank === false;
|
|
213
|
+
// The SDK is also a JavaScript boundary: TypeScript declarations do not
|
|
214
|
+
// protect plain-JS callers or deserialized input. Apply the same bounded,
|
|
215
|
+
// strict validation used by CLI, MCP, and HTTP before compiling SQL.
|
|
216
|
+
const filter = opts.filter === undefined ? undefined : parseMetadataFilter(opts.filter);
|
|
211
217
|
if (opts.queries) {
|
|
212
218
|
// Pre-expanded queries — use structuredSearch
|
|
213
219
|
return structuredSearch(internal, opts.queries, {
|
|
214
220
|
collections: collections.length > 0 ? collections : undefined,
|
|
221
|
+
filter,
|
|
215
222
|
limit: opts.limit,
|
|
216
223
|
minScore: opts.minScore,
|
|
217
224
|
explain: opts.explain,
|
|
@@ -225,6 +232,7 @@ export async function createStore(options) {
|
|
|
225
232
|
return hybridQuery(internal, opts.query, {
|
|
226
233
|
collections: collections.length > 0 ? collections : undefined,
|
|
227
234
|
collection: collections.length === 1 ? collections[0] : (collections.length > 0 ? collections : undefined),
|
|
235
|
+
filter,
|
|
228
236
|
limit: opts.limit,
|
|
229
237
|
minScore: opts.minScore,
|
|
230
238
|
explain: opts.explain,
|
|
@@ -238,10 +246,14 @@ export async function createStore(options) {
|
|
|
238
246
|
chunkStrategy: opts.chunkStrategy,
|
|
239
247
|
});
|
|
240
248
|
},
|
|
241
|
-
searchLex: async (q, opts) =>
|
|
249
|
+
searchLex: async (q, opts) => {
|
|
250
|
+
const filter = opts?.filter === undefined ? undefined : parseMetadataFilter(opts.filter);
|
|
251
|
+
return internal.searchFTS(q, opts?.limit, opts?.collection, filter);
|
|
252
|
+
},
|
|
242
253
|
searchVector: async (q, opts) => {
|
|
254
|
+
const filter = opts?.filter === undefined ? undefined : parseMetadataFilter(opts.filter);
|
|
243
255
|
const provider = internal.embeddingProvider;
|
|
244
|
-
return internal.searchVec(q, provider?.model ?? internal.llm?.embedModelName ?? DEFAULT_EMBED_MODEL_URI, opts?.limit, opts?.collection);
|
|
256
|
+
return internal.searchVec(q, provider?.model ?? internal.llm?.embedModelName ?? DEFAULT_EMBED_MODEL_URI, opts?.limit, opts?.collection, undefined, undefined, filter);
|
|
245
257
|
},
|
|
246
258
|
expandQuery: async (q, opts) => internal.expandQuery(q, undefined, opts?.expansionContext, {
|
|
247
259
|
includeLexical: opts?.includeLexical,
|
package/dist/llm.d.ts
CHANGED
|
@@ -199,7 +199,13 @@ export type PullResult = {
|
|
|
199
199
|
export type GgufFileInspection = {
|
|
200
200
|
exists: boolean;
|
|
201
201
|
valid: boolean;
|
|
202
|
-
|
|
202
|
+
/**
|
|
203
|
+
* "html" and "invalid" mean the header was read and is confirmed bad.
|
|
204
|
+
* "unreadable" means the read itself failed, so nothing is known about the
|
|
205
|
+
* content. The two stay distinct because only the first justifies deleting
|
|
206
|
+
* the file.
|
|
207
|
+
*/
|
|
208
|
+
kind: "missing" | "gguf" | "html" | "invalid" | "unreadable";
|
|
203
209
|
sizeBytes?: number;
|
|
204
210
|
magic?: string;
|
|
205
211
|
details: string;
|
package/dist/llm.js
CHANGED
|
@@ -235,10 +235,13 @@ export function inspectGgufFile(filePath) {
|
|
|
235
235
|
};
|
|
236
236
|
}
|
|
237
237
|
catch (error) {
|
|
238
|
+
// stat/open/read threw, so the header was never seen. Reporting this as
|
|
239
|
+
// "invalid" would claim a verdict no read supports, and `magic` stays
|
|
240
|
+
// undefined for the same reason.
|
|
238
241
|
return {
|
|
239
242
|
exists: true,
|
|
240
243
|
valid: false,
|
|
241
|
-
kind: "
|
|
244
|
+
kind: "unreadable",
|
|
242
245
|
sizeBytes,
|
|
243
246
|
details: `cannot read model file: ${error instanceof Error ? error.message : String(error)}`,
|
|
244
247
|
};
|
|
@@ -253,11 +256,25 @@ function validateGgufFile(filePath, modelUri) {
|
|
|
253
256
|
const inspection = inspectGgufFile(filePath);
|
|
254
257
|
if (!inspection.exists || inspection.valid)
|
|
255
258
|
return; // let downstream handle missing files
|
|
256
|
-
//
|
|
259
|
+
// A read that failed says nothing about the bytes on disk. Deleting here
|
|
260
|
+
// throws away a file that is usually fine (fd exhaustion, a concurrent
|
|
261
|
+
// loader, a volume that briefly went away) and re-downloading it can cost
|
|
262
|
+
// gigabytes, so surface the real error and leave the file alone.
|
|
263
|
+
if (inspection.kind === "unreadable") {
|
|
264
|
+
throw new Error(`Model file could not be read, so it could not be validated (${inspection.details}).\n` +
|
|
265
|
+
`Model: ${modelUri}\n` +
|
|
266
|
+
`Path: ${filePath}\n\n` +
|
|
267
|
+
`The file has been left in place. If this repeats, check open file limits, ` +
|
|
268
|
+
`permissions, and whether the volume holding the model cache is still mounted.`);
|
|
269
|
+
}
|
|
270
|
+
// Confirmed bad content: remove it so the next attempt re-downloads.
|
|
271
|
+
let removed = true;
|
|
257
272
|
try {
|
|
258
273
|
unlinkSync(filePath);
|
|
259
274
|
}
|
|
260
|
-
catch {
|
|
275
|
+
catch {
|
|
276
|
+
removed = false;
|
|
277
|
+
}
|
|
261
278
|
if (inspection.kind === "html") {
|
|
262
279
|
throw new Error(`Downloaded model file is an HTML page, not a GGUF model (${formatModelFileSize(inspection.sizeBytes ?? 0)}).\n` +
|
|
263
280
|
`Something is intercepting the download from huggingface.co (a proxy, firewall, or captive portal).\n\n` +
|
|
@@ -272,7 +289,9 @@ function validateGgufFile(filePath, modelUri) {
|
|
|
272
289
|
throw new Error(`Model file is not valid GGUF (expected magic "GGUF", got "${inspection.magic ?? "unknown"}", file is ${formatModelFileSize(inspection.sizeBytes ?? 0)}).\n` +
|
|
273
290
|
`Model: ${modelUri}\n` +
|
|
274
291
|
`Path: ${filePath}\n\n` +
|
|
275
|
-
|
|
292
|
+
(removed
|
|
293
|
+
? `The file has been removed. Run the command again to re-download.`
|
|
294
|
+
: `The file could NOT be removed. Delete it manually, then run the command again.`));
|
|
276
295
|
}
|
|
277
296
|
/**
|
|
278
297
|
* node-llama-cpp prints a multi-line download progress bar when the second
|