mvw-search-index 2.3.7 → 3.0.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +144 -35
- package/js/cli.js +70 -7
- package/js/href.d.ts +8 -0
- package/js/href.js +24 -0
- package/js/html.d.ts +26 -0
- package/js/html.js +71 -0
- package/js/index.d.ts +82 -9
- package/js/index.js +106 -36
- package/js/language.d.ts +8 -0
- package/js/language.js +45 -0
- package/package.json +21 -12
- package/README.Release.md +0 -5
- package/eslint.config.mjs +0 -15
- package/js/cli.js.map +0 -1
- package/js/index.js.map +0 -1
- package/vitest.config.ts +0 -8
package/README.md
CHANGED
|
@@ -8,9 +8,12 @@
|
|
|
8
8
|
|
|
9
9
|
## About
|
|
10
10
|
|
|
11
|
-
**mvw-search-index**
|
|
11
|
+
**mvw-search-index** generates a [lunr](https://lunrjs.com/) search index plus a result store from the HTML
|
|
12
|
+
files of a static website - built with Hugo, Jekyll, Gatsby, or by hand. The generated JSON file is loaded by
|
|
13
|
+
the browser, which then searches entirely client-side.
|
|
12
14
|
|
|
13
|
-
|
|
15
|
+
> **Upgrading from 2.x?** Version 3 changes the API, the CLI and what ends up in the index.
|
|
16
|
+
> See the [upgrade guide](UPGRADING.md).
|
|
14
17
|
|
|
15
18
|
---
|
|
16
19
|
|
|
@@ -18,9 +21,12 @@ It is ideal for adding fast client-side search capabilities to static websites
|
|
|
18
21
|
|
|
19
22
|
- [Installation](#installation)
|
|
20
23
|
- [Usage](#usage)
|
|
21
|
-
- [
|
|
22
|
-
- [
|
|
23
|
-
- [
|
|
24
|
+
- [CLI](#cli)
|
|
25
|
+
- [Node.js / TypeScript](#nodejs--typescript)
|
|
26
|
+
- [Options](#options)
|
|
27
|
+
- [Searching in the browser](#searching-in-the-browser)
|
|
28
|
+
- [Content in other languages](#content-in-other-languages)
|
|
29
|
+
- [What gets indexed](#what-gets-indexed)
|
|
24
30
|
- [Demo](#demo)
|
|
25
31
|
- [Releases](#releases)
|
|
26
32
|
|
|
@@ -28,7 +34,7 @@ It is ideal for adding fast client-side search capabilities to static websites
|
|
|
28
34
|
|
|
29
35
|
## Installation
|
|
30
36
|
|
|
31
|
-
|
|
37
|
+
Requires Node.js 22.12 or newer.
|
|
32
38
|
|
|
33
39
|
```bash
|
|
34
40
|
npm install --save-dev mvw-search-index
|
|
@@ -38,64 +44,164 @@ npm install --save-dev mvw-search-index
|
|
|
38
44
|
|
|
39
45
|
## Usage
|
|
40
46
|
|
|
41
|
-
|
|
47
|
+
Run the indexer **after** your site has been built, on the generated HTML.
|
|
42
48
|
|
|
43
|
-
###
|
|
49
|
+
### CLI
|
|
44
50
|
|
|
45
51
|
```bash
|
|
46
|
-
mvw-search-index <glob> <
|
|
52
|
+
mvw-search-index [options] <glob> <dest> [bodySelector]
|
|
47
53
|
```
|
|
48
54
|
|
|
49
|
-
Example
|
|
55
|
+
Example - index everything in `public/`, but only the content of `<main>`, with German stemming and
|
|
56
|
+
root-relative links:
|
|
50
57
|
|
|
51
58
|
```bash
|
|
52
|
-
mvw-search-index
|
|
59
|
+
mvw-search-index '**/*.html' public/suche/index.json main --cwd public --base-url / --language de
|
|
53
60
|
```
|
|
54
61
|
|
|
55
|
-
|
|
62
|
+
Quote the glob so your shell doesn't expand it. From an npm script:
|
|
56
63
|
|
|
57
64
|
```json
|
|
58
65
|
{
|
|
59
66
|
"scripts": {
|
|
60
|
-
"index": "mvw-search-index '
|
|
67
|
+
"index": "mvw-search-index '**/*.html' public/suche/index.json main --cwd public --base-url /"
|
|
61
68
|
}
|
|
62
69
|
}
|
|
63
70
|
```
|
|
64
71
|
|
|
65
|
-
|
|
72
|
+
| Flag | Option | Description |
|
|
73
|
+
|------------------------------|-------------------|-----------------------------------------------------------------------------|
|
|
74
|
+
| `[bodySelector]` | `bodySelector` | CSS selector of the content to index (default `body`) |
|
|
75
|
+
| `--cwd <dir>` | `cwd` | Directory the glob is resolved in; hrefs are relative to it |
|
|
76
|
+
| `-e, --exclude <selector>` | `excludeSelector` | Content to leave out (default `"nav, footer"`, `""` for none) |
|
|
77
|
+
| `-l, --language <code>` | `language` | Content language: stop words and stemmer (default `en`) |
|
|
78
|
+
| `--base-url <url>` | `baseUrl` | Prefix for every href, e.g. `/` |
|
|
79
|
+
| `--strip-index-html` | `stripIndexHtml` | Link to `dir/` instead of `dir/index.html` |
|
|
80
|
+
| `--no-noindex` | `respectNoindex` | Also index pages marked `<meta name="robots" content="noindex">` |
|
|
81
|
+
| `--allow-empty` | `allowEmpty` | Write an empty index instead of failing when the glob matches nothing |
|
|
82
|
+
| `-b, --boost <field=number>` | `boosts` | Field weight, repeatable, e.g. `-b title=10 -b body=1` |
|
|
83
|
+
| `-v, --verbose` | `logger` | List every indexed file |
|
|
66
84
|
|
|
67
|
-
|
|
85
|
+
The CLI exits with code 1 and prints `Error: …` if indexing fails, including when the glob matches no files.
|
|
86
|
+
|
|
87
|
+
`<dest>` must be inside the directory you run the command from (relative or absolute); anything outside, such as
|
|
88
|
+
`../index.json`, is rejected so a wrong argument can't overwrite unrelated files. `--cwd` only affects where the
|
|
89
|
+
HTML files are read from.
|
|
90
|
+
|
|
91
|
+
### Node.js / TypeScript
|
|
68
92
|
|
|
69
93
|
```ts
|
|
70
|
-
import {
|
|
71
|
-
import
|
|
94
|
+
import {writeFile} from "fs/promises";
|
|
95
|
+
import {SearchIndex} from "mvw-search-index";
|
|
96
|
+
|
|
97
|
+
const result = await SearchIndex.createFromGlob("**/*.html", {
|
|
98
|
+
cwd: "public",
|
|
99
|
+
bodySelector: "main",
|
|
100
|
+
language: "de",
|
|
101
|
+
baseUrl: "/",
|
|
102
|
+
});
|
|
103
|
+
await writeFile("public/suche/index.json", JSON.stringify(result));
|
|
104
|
+
```
|
|
105
|
+
|
|
106
|
+
CommonJS works the same way: `const {SearchIndex} = require("mvw-search-index");`.
|
|
107
|
+
|
|
108
|
+
There are three entry points:
|
|
72
109
|
|
|
73
|
-
|
|
74
|
-
|
|
75
|
-
|
|
76
|
-
|
|
110
|
+
| Method | Input | Returns |
|
|
111
|
+
|------------------------------------------|------------------------------------------------------------------|-------------------------------|
|
|
112
|
+
| `createFromGlob(pattern, options?)` | A glob pattern; files are read from disk | `Promise<ISearchIndexResult>` |
|
|
113
|
+
| `createFromHtml(files, options?)` | `HtmlFile[]` - `{relative: string, contents: string \| Buffer}` | `ISearchIndexResult` |
|
|
114
|
+
| `createFromInfo(files, options?)` | `IFileInformation[]` - already extracted title/body/… | `ISearchIndexResult` |
|
|
77
115
|
|
|
78
|
-
|
|
116
|
+
`ISearchIndexResult` is `{index: lunr.Index, store: {[href]: {title, description?}}}`. `JSON.stringify()` it to
|
|
117
|
+
get the file the browser loads.
|
|
118
|
+
|
|
119
|
+
### Options
|
|
120
|
+
|
|
121
|
+
| Option | Default | Applies to | Description |
|
|
122
|
+
|-------------------|--------------------------------------------------|--------------------|-------------------------------------------------------------------------------------------------------------------------------|
|
|
123
|
+
| `bodySelector` | `"body"` | Glob, HTML | CSS selector of the element(s) whose text is indexed |
|
|
124
|
+
| `excludeSelector` | `"nav, footer"` | Glob, HTML | Elements inside the body to leave out; `""` for none |
|
|
125
|
+
| `respectNoindex` | `true` | Glob, HTML | Skip pages with `<meta name="robots" content="noindex">` |
|
|
126
|
+
| `language` | `"en"` | all | Two-letter language code; other than `en` uses [lunr-languages](https://github.com/MihaiValentin/lunr-languages) (see below) |
|
|
127
|
+
| `boosts` | `{title: 5, keywords: 3, description: 2, body: 1}` | all | Relative weight of matches per field |
|
|
128
|
+
| `cwd` | `process.cwd()` | Glob | Directory the pattern is resolved in; hrefs are relative to it |
|
|
129
|
+
| `allowEmpty` | `false` | Glob | Resolve with an empty index instead of rejecting when nothing matches |
|
|
130
|
+
| `baseUrl` | `""` | Glob, HTML | Prefix for every href |
|
|
131
|
+
| `stripIndexHtml` | `false` | Glob, HTML | `foo/index.html` → `foo/` |
|
|
132
|
+
| `logger` | silent | Glob, HTML | Receives progress (`info`) and warnings (`warn`); `console` works |
|
|
133
|
+
|
|
134
|
+
---
|
|
135
|
+
|
|
136
|
+
## Searching in the browser
|
|
137
|
+
|
|
138
|
+
Load lunr (2.3.x) and the generated file, then query it. Building the query with lunr's query API - rather than
|
|
139
|
+
passing user input to `index.search()` - means no input can cause a `QueryParseError`, and every word is matched
|
|
140
|
+
both as a whole (stemmed) and as a prefix, so results appear while typing:
|
|
141
|
+
|
|
142
|
+
```html
|
|
143
|
+
<script src="https://cdnjs.cloudflare.com/ajax/libs/lunr.js/2.3.9/lunr.min.js"></script>
|
|
144
|
+
<script type="module">
|
|
145
|
+
const {index: serializedIndex, store} = await (await fetch("/suche/index.json")).json();
|
|
146
|
+
const index = lunr.Index.load(serializedIndex);
|
|
147
|
+
|
|
148
|
+
function search(input) {
|
|
149
|
+
const terms = lunr.tokenizer(input)
|
|
150
|
+
.map((token) => lunr.trimmer(token).toString())
|
|
151
|
+
.filter((term) => term.length > 0);
|
|
152
|
+
if (terms.length === 0) {
|
|
153
|
+
return [];
|
|
154
|
+
}
|
|
155
|
+
return index.query((query) => {
|
|
156
|
+
for (const term of terms) {
|
|
157
|
+
query.term(term, {boost: 10}); // whole word, stemmed like the index
|
|
158
|
+
query.term(term, {usePipeline: false, wildcard: lunr.Query.wildcard.TRAILING}); // prefix
|
|
159
|
+
}
|
|
160
|
+
}).map((result) => ({href: result.ref, ...store[result.ref]}));
|
|
161
|
+
}
|
|
162
|
+
</script>
|
|
79
163
|
```
|
|
80
164
|
|
|
81
|
-
|
|
165
|
+
Render the `title` and `description` with `textContent` (not `innerHTML`). [docs/index.html](docs/index.html) is a
|
|
166
|
+
complete, working example.
|
|
82
167
|
|
|
83
|
-
|
|
84
|
-
"use strict";
|
|
168
|
+
### Content in other languages
|
|
85
169
|
|
|
86
|
-
|
|
87
|
-
|
|
170
|
+
With `language` set to anything other than `en`, the index references that language's lunr-languages pipeline
|
|
171
|
+
functions. Load the matching scripts after lunr and **before** `lunr.Index.load()` - otherwise lunr throws
|
|
172
|
+
`Cannot load unregistered function: trimmer-de`:
|
|
88
173
|
|
|
89
|
-
|
|
90
|
-
|
|
91
|
-
|
|
174
|
+
```html
|
|
175
|
+
<script src="https://cdnjs.cloudflare.com/ajax/libs/lunr.js/2.3.9/lunr.min.js"></script>
|
|
176
|
+
<script src="https://cdn.jsdelivr.net/npm/lunr-languages@1.22.0/lunr.stemmer.support.js"></script>
|
|
177
|
+
<script src="https://cdn.jsdelivr.net/npm/lunr-languages@1.22.0/lunr.de.js"></script>
|
|
92
178
|
```
|
|
93
179
|
|
|
180
|
+
Use `lunr.de.trimmer` instead of `lunr.trimmer` in the `search()` function above, so umlauts at word boundaries are
|
|
181
|
+
kept.
|
|
182
|
+
|
|
183
|
+
---
|
|
184
|
+
|
|
185
|
+
## What gets indexed
|
|
186
|
+
|
|
187
|
+
For every page:
|
|
188
|
+
|
|
189
|
+
- **title**: `<title>`, falling back to `og:title`, then the first `<h1>`
|
|
190
|
+
- **description**: `<meta name="description">`, falling back to `og:description`
|
|
191
|
+
- **keywords**: `<meta name="keywords">` (comma separated)
|
|
192
|
+
- **body**: the text of `bodySelector`, without `excludeSelector` matches, `<script>`, `<style>`, `<noscript>` and
|
|
193
|
+
`<template>`; words in separate block elements stay separate words
|
|
194
|
+
|
|
195
|
+
Title and description also go into the result store, keyed by href. Text runs through lunr's pipeline for the
|
|
196
|
+
configured language: punctuation is trimmed, stop words are dropped and words are stemmed.
|
|
197
|
+
|
|
198
|
+
Pages with `<meta name="robots" content="noindex">` are skipped.
|
|
199
|
+
|
|
94
200
|
---
|
|
95
201
|
|
|
96
202
|
## Demo
|
|
97
203
|
|
|
98
|
-
A basic sample site is included and served from [GitHub Pages](https://
|
|
204
|
+
A basic sample site is included and served from [GitHub Pages](https://tiliavir.github.io/mvw-search-index/).
|
|
99
205
|
|
|
100
206
|
Start it locally with:
|
|
101
207
|
|
|
@@ -103,13 +209,18 @@ Start it locally with:
|
|
|
103
209
|
npm run serve
|
|
104
210
|
```
|
|
105
211
|
|
|
106
|
-
This
|
|
212
|
+
This rebuilds `docs/index.json` and serves [./docs](docs): a simple static site with a search form on `index.html`.
|
|
107
213
|
|
|
108
214
|
---
|
|
109
215
|
|
|
110
216
|
## Releases
|
|
111
217
|
|
|
112
|
-
- **
|
|
218
|
+
- **3.0.1**: Fixes all open SonarCloud issues - `node:` imports, `export … from` re-exports, `Object.hasOwn`,
|
|
219
|
+
`replaceAll`; CI installs with `--ignore-scripts` and pins third-party actions to commit SHAs; the demo page can be
|
|
220
|
+
zoomed and its search field has a label. No API or behaviour changes.
|
|
221
|
+
- **3.0.0**: Major overhaul - Promise based API, correct text extraction and lunr pipeline (stemming, stop words),
|
|
222
|
+
language support, field boosts, href options, a full-featured CLI. **Breaking** - see [UPGRADING.md](UPGRADING.md).
|
|
223
|
+
- **2.3.2 – 2.3.7**: Dependency updates.
|
|
113
224
|
- **2.3.0**: Added attribute support for metadata extraction.
|
|
114
225
|
- **2.2.10 - 2.2.16**: Dependency updates.
|
|
115
226
|
- **2.2.9**: Removed `vinyl`; introduced demo application.
|
|
@@ -130,5 +241,3 @@ This project is licensed under the [MIT License](LICENSE).
|
|
|
130
241
|
## Author
|
|
131
242
|
|
|
132
243
|
Maintained by [Tiliavir](https://github.com/Tiliavir).
|
|
133
|
-
|
|
134
|
-
---
|
package/js/cli.js
CHANGED
|
@@ -35,13 +35,76 @@ var __importStar = (this && this.__importStar) || (function () {
|
|
|
35
35
|
})();
|
|
36
36
|
Object.defineProperty(exports, "__esModule", { value: true });
|
|
37
37
|
const commander_1 = require("commander");
|
|
38
|
-
const fs = __importStar(require("fs"));
|
|
38
|
+
const fs = __importStar(require("node:fs"));
|
|
39
|
+
const path = __importStar(require("node:path"));
|
|
39
40
|
const index_1 = require("./index");
|
|
41
|
+
// read at runtime: package.json is outside of the compiled sources
|
|
42
|
+
const { version } = JSON.parse(fs.readFileSync(path.join(__dirname, "..", "package.json"), "utf8"));
|
|
43
|
+
const FIELDS = ["title", "keywords", "description", "body"];
|
|
44
|
+
function parseBoost(value, previous = {}) {
|
|
45
|
+
const match = /^(\w+)=(\d+(?:\.\d+)?)$/.exec(value);
|
|
46
|
+
if (!match || !FIELDS.includes(match[1])) {
|
|
47
|
+
throw new commander_1.InvalidArgumentError(`Expected <field>=<number> with field one of ${FIELDS.join(", ")}.`);
|
|
48
|
+
}
|
|
49
|
+
return { ...previous, [match[1]]: Number(match[2]) };
|
|
50
|
+
}
|
|
51
|
+
/**
|
|
52
|
+
* Resolves <dest> and makes sure it lies inside the current working directory, so a
|
|
53
|
+
* mistyped or injected argument (e.g. "../../somewhere") cannot overwrite arbitrary files.
|
|
54
|
+
*/
|
|
55
|
+
function resolveDestination(dest) {
|
|
56
|
+
const root = process.cwd();
|
|
57
|
+
const resolved = path.resolve(root, dest);
|
|
58
|
+
const prefix = root.endsWith(path.sep) ? root : root + path.sep;
|
|
59
|
+
if (!resolved.startsWith(prefix)) {
|
|
60
|
+
throw new Error(`<dest> must be a file inside the current directory (${root}), got "${dest}".`);
|
|
61
|
+
}
|
|
62
|
+
return resolved;
|
|
63
|
+
}
|
|
40
64
|
commander_1.program
|
|
41
|
-
.
|
|
42
|
-
.
|
|
43
|
-
.
|
|
44
|
-
|
|
65
|
+
.name("mvw-search-index")
|
|
66
|
+
.description("Generates a lunr search index and result store from HTML files.")
|
|
67
|
+
.version(version)
|
|
68
|
+
.argument("<glob>", "glob pattern of the HTML files to index (quote it, so your shell doesn't expand it)")
|
|
69
|
+
.argument("<dest>", "path of the JSON file to write; must be inside the current directory")
|
|
70
|
+
.argument("[bodySelector]", "CSS selector of the content to index", "body")
|
|
71
|
+
.option("--cwd <dir>", "directory to resolve <glob> in; hrefs are relative to it (use your site's root)")
|
|
72
|
+
.option("-e, --exclude <selector>", "CSS selector of content to leave out (\"\" for none)", index_1.DEFAULT_EXCLUDE_SELECTOR)
|
|
73
|
+
.option("-l, --language <code>", "two-letter content language, selects stemmer and stop words", index_1.DEFAULT_LANGUAGE)
|
|
74
|
+
.option("--base-url <url>", "prefix for all hrefs, e.g. \"/\"")
|
|
75
|
+
.option("--strip-index-html", "link to \"dir/\" instead of \"dir/index.html\"")
|
|
76
|
+
.option("--no-noindex", "also index pages marked <meta name=\"robots\" content=\"noindex\">")
|
|
77
|
+
.option("--allow-empty", "write an empty index instead of failing if <glob> matches no files")
|
|
78
|
+
.option("-b, --boost <field=number>", "weight of a field, repeatable (default title=5 keywords=3 description=2 body=1)", parseBoost)
|
|
79
|
+
.option("-v, --verbose", "list every indexed file")
|
|
80
|
+
.showHelpAfterError()
|
|
81
|
+
.action(async (glob, dest, bodySelector, cli) => {
|
|
82
|
+
const logger = {
|
|
83
|
+
info: (message) => cli.verbose && console.log(message),
|
|
84
|
+
warn: (message) => console.warn(`Warning: ${message}`),
|
|
85
|
+
};
|
|
86
|
+
const options = {
|
|
87
|
+
bodySelector,
|
|
88
|
+
excludeSelector: cli.exclude,
|
|
89
|
+
language: cli.language,
|
|
90
|
+
cwd: cli.cwd,
|
|
91
|
+
baseUrl: cli.baseUrl,
|
|
92
|
+
stripIndexHtml: cli.stripIndexHtml,
|
|
93
|
+
respectNoindex: cli.noindex,
|
|
94
|
+
allowEmpty: cli.allowEmpty,
|
|
95
|
+
boosts: cli.boost,
|
|
96
|
+
logger,
|
|
97
|
+
};
|
|
98
|
+
try {
|
|
99
|
+
const destination = resolveDestination(dest);
|
|
100
|
+
const index = await index_1.SearchIndex.createFromGlob(glob, options);
|
|
101
|
+
await fs.promises.mkdir(path.dirname(destination), { recursive: true });
|
|
102
|
+
await fs.promises.writeFile(destination, JSON.stringify(index));
|
|
103
|
+
console.log(`Indexed ${Object.keys(index.store).length} page(s) into ${dest}`);
|
|
104
|
+
}
|
|
105
|
+
catch (err) {
|
|
106
|
+
console.error(`Error: ${err instanceof Error ? err.message : err}`);
|
|
107
|
+
process.exitCode = 1;
|
|
108
|
+
}
|
|
45
109
|
})
|
|
46
|
-
.
|
|
47
|
-
//# sourceMappingURL=cli.js.map
|
|
110
|
+
.parseAsync(process.argv);
|
package/js/href.d.ts
ADDED
|
@@ -0,0 +1,8 @@
|
|
|
1
|
+
export declare interface HrefOptions {
|
|
2
|
+
/** Prefix for every href, e.g. `"/"` or `"https://example.org/docs/"`. Default: `""`. */
|
|
3
|
+
baseUrl?: string;
|
|
4
|
+
/** Turn `foo/index.html` into `foo/` (and `index.html` into the base URL). Default: `false`. */
|
|
5
|
+
stripIndexHtml?: boolean;
|
|
6
|
+
}
|
|
7
|
+
/** Turns the path of an indexed file into the `href` of its search result. */
|
|
8
|
+
export declare function toHref(relativePath: string, options?: HrefOptions): string;
|
package/js/href.js
ADDED
|
@@ -0,0 +1,24 @@
|
|
|
1
|
+
"use strict";
|
|
2
|
+
Object.defineProperty(exports, "__esModule", { value: true });
|
|
3
|
+
exports.toHref = toHref;
|
|
4
|
+
/** Turns the path of an indexed file into the `href` of its search result. */
|
|
5
|
+
function toHref(relativePath, options = {}) {
|
|
6
|
+
// Windows paths use backslashes, URLs never do.
|
|
7
|
+
let href = relativePath.replaceAll("\\", "/").replace(/^\.\//, "");
|
|
8
|
+
if (options.stripIndexHtml) {
|
|
9
|
+
href = href.replace(/(^|\/)index\.html?$/, "$1");
|
|
10
|
+
}
|
|
11
|
+
const baseUrl = trimTrailingSlashes(options.baseUrl ?? "");
|
|
12
|
+
if (options.baseUrl) {
|
|
13
|
+
href = baseUrl + "/" + href;
|
|
14
|
+
}
|
|
15
|
+
// an empty href would link to the search page itself
|
|
16
|
+
return href || "./";
|
|
17
|
+
}
|
|
18
|
+
function trimTrailingSlashes(value) {
|
|
19
|
+
let end = value.length;
|
|
20
|
+
while (end > 0 && value[end - 1] === "/") {
|
|
21
|
+
end--;
|
|
22
|
+
}
|
|
23
|
+
return value.slice(0, end);
|
|
24
|
+
}
|
package/js/html.d.ts
ADDED
|
@@ -0,0 +1,26 @@
|
|
|
1
|
+
import type { CheerioAPI } from "cheerio";
|
|
2
|
+
/** Collapses all whitespace runs (including single newlines and tabs) into one space. */
|
|
3
|
+
export declare function normalizeWhitespace(text: string): string;
|
|
4
|
+
export declare interface PageMetadata {
|
|
5
|
+
title: string;
|
|
6
|
+
description?: string;
|
|
7
|
+
keywords?: string;
|
|
8
|
+
}
|
|
9
|
+
/**
|
|
10
|
+
* Reads title, description and keywords, falling back to Open Graph tags and the
|
|
11
|
+
* first `<h1>` where the regular tags are missing or empty. Call this before
|
|
12
|
+
* {@link extractText}, which modifies the document.
|
|
13
|
+
*/
|
|
14
|
+
export declare function extractMetadata($: CheerioAPI): PageMetadata;
|
|
15
|
+
/** Whether the page asks search engines not to index it (`<meta name="robots" content="noindex">`). */
|
|
16
|
+
export declare function isNoindex($: CheerioAPI): boolean;
|
|
17
|
+
/** Default for the `excludeSelector` option: site-wide navigation and footers. */
|
|
18
|
+
export declare const DEFAULT_EXCLUDE_SELECTOR = "nav, footer";
|
|
19
|
+
/**
|
|
20
|
+
* Returns the visible text of all elements matching `selector`, with block-level
|
|
21
|
+
* elements and separate matches delimited by spaces. Descendants matching
|
|
22
|
+
* `excludeSelector` are left out.
|
|
23
|
+
*
|
|
24
|
+
* Note: modifies the document.
|
|
25
|
+
*/
|
|
26
|
+
export declare function extractText($: CheerioAPI, selector: string, excludeSelector?: string): string;
|
package/js/html.js
ADDED
|
@@ -0,0 +1,71 @@
|
|
|
1
|
+
"use strict";
|
|
2
|
+
Object.defineProperty(exports, "__esModule", { value: true });
|
|
3
|
+
exports.DEFAULT_EXCLUDE_SELECTOR = void 0;
|
|
4
|
+
exports.normalizeWhitespace = normalizeWhitespace;
|
|
5
|
+
exports.extractMetadata = extractMetadata;
|
|
6
|
+
exports.isNoindex = isNoindex;
|
|
7
|
+
exports.extractText = extractText;
|
|
8
|
+
/**
|
|
9
|
+
* Elements that browsers render on their own line (plus table cells). Their text
|
|
10
|
+
* must be separated from the surrounding text, otherwise `<li>a</li><li>b</li>`
|
|
11
|
+
* becomes the single word "ab".
|
|
12
|
+
*/
|
|
13
|
+
const BLOCK_ELEMENTS = [
|
|
14
|
+
"address", "article", "aside", "blockquote", "caption", "dd", "details", "dialog", "div", "dl", "dt",
|
|
15
|
+
"fieldset", "figcaption", "figure", "footer", "form", "h1", "h2", "h3", "h4", "h5", "h6", "header",
|
|
16
|
+
"hgroup", "hr", "li", "main", "nav", "ol", "option", "p", "pre", "section", "summary", "table",
|
|
17
|
+
"tbody", "td", "tfoot", "th", "thead", "tr", "ul",
|
|
18
|
+
].join(",");
|
|
19
|
+
/** Elements whose content is never visible page text. */
|
|
20
|
+
const NON_CONTENT_ELEMENTS = "script, style, noscript, template";
|
|
21
|
+
/** Collapses all whitespace runs (including single newlines and tabs) into one space. */
|
|
22
|
+
function normalizeWhitespace(text) {
|
|
23
|
+
return text.replace(/\s+/g, " ").trim();
|
|
24
|
+
}
|
|
25
|
+
function nonEmpty(value) {
|
|
26
|
+
const normalized = normalizeWhitespace(value ?? "");
|
|
27
|
+
return normalized === "" ? undefined : normalized;
|
|
28
|
+
}
|
|
29
|
+
/**
|
|
30
|
+
* Reads title, description and keywords, falling back to Open Graph tags and the
|
|
31
|
+
* first `<h1>` where the regular tags are missing or empty. Call this before
|
|
32
|
+
* {@link extractText}, which modifies the document.
|
|
33
|
+
*/
|
|
34
|
+
function extractMetadata($) {
|
|
35
|
+
const meta = (selector) => nonEmpty($(selector).first().attr("content"));
|
|
36
|
+
return {
|
|
37
|
+
title: nonEmpty($("title").first().text())
|
|
38
|
+
?? meta("meta[property='og:title']")
|
|
39
|
+
?? nonEmpty($("h1").first().text())
|
|
40
|
+
?? "",
|
|
41
|
+
description: meta("meta[name='description' i]") ?? meta("meta[property='og:description']"),
|
|
42
|
+
keywords: meta("meta[name='keywords' i]"),
|
|
43
|
+
};
|
|
44
|
+
}
|
|
45
|
+
/** Whether the page asks search engines not to index it (`<meta name="robots" content="noindex">`). */
|
|
46
|
+
function isNoindex($) {
|
|
47
|
+
return $("meta[name='robots' i]").toArray()
|
|
48
|
+
.some((meta) => /\b(noindex|none)\b/i.test($(meta).attr("content") ?? ""));
|
|
49
|
+
}
|
|
50
|
+
/** Default for the `excludeSelector` option: site-wide navigation and footers. */
|
|
51
|
+
exports.DEFAULT_EXCLUDE_SELECTOR = "nav, footer";
|
|
52
|
+
/**
|
|
53
|
+
* Returns the visible text of all elements matching `selector`, with block-level
|
|
54
|
+
* elements and separate matches delimited by spaces. Descendants matching
|
|
55
|
+
* `excludeSelector` are left out.
|
|
56
|
+
*
|
|
57
|
+
* Note: modifies the document.
|
|
58
|
+
*/
|
|
59
|
+
function extractText($, selector, excludeSelector = exports.DEFAULT_EXCLUDE_SELECTOR) {
|
|
60
|
+
$(NON_CONTENT_ELEMENTS).remove();
|
|
61
|
+
$("br").replaceWith(" ");
|
|
62
|
+
$(BLOCK_ELEMENTS).prepend(" ").append(" ");
|
|
63
|
+
// If matches are nested (e.g. selector "div"), only take the outermost ones -
|
|
64
|
+
// otherwise the inner text would be indexed twice.
|
|
65
|
+
const roots = $(selector).filter((_, el) => $(el).parents(selector).length === 0);
|
|
66
|
+
if (excludeSelector.trim()) {
|
|
67
|
+
// only descendants: a body selector that itself matches the exclusion still works
|
|
68
|
+
roots.find(excludeSelector).remove();
|
|
69
|
+
}
|
|
70
|
+
return normalizeWhitespace(roots.map((_, el) => $(el).text()).get().join(" "));
|
|
71
|
+
}
|
package/js/index.d.ts
CHANGED
|
@@ -1,32 +1,105 @@
|
|
|
1
1
|
import * as lunr from "lunr";
|
|
2
|
+
import { type HrefOptions } from "./href";
|
|
3
|
+
export { DEFAULT_EXCLUDE_SELECTOR } from "./html";
|
|
4
|
+
export type { HrefOptions } from "./href";
|
|
5
|
+
export { DEFAULT_LANGUAGE } from "./language";
|
|
2
6
|
export declare interface IResultStore {
|
|
3
7
|
[key: string]: {
|
|
4
8
|
title: string;
|
|
5
|
-
description
|
|
9
|
+
description?: string;
|
|
6
10
|
};
|
|
7
11
|
}
|
|
8
12
|
export declare interface IFileInformation {
|
|
9
13
|
body: string;
|
|
10
|
-
description
|
|
14
|
+
description?: string;
|
|
11
15
|
href: string;
|
|
12
|
-
keywords
|
|
16
|
+
keywords?: string;
|
|
13
17
|
title: string;
|
|
14
18
|
}
|
|
15
19
|
export declare interface ISearchIndexResult {
|
|
16
20
|
index: lunr.Index;
|
|
17
21
|
store: IResultStore;
|
|
18
22
|
}
|
|
19
|
-
|
|
20
|
-
|
|
23
|
+
/** An HTML document to index. */
|
|
24
|
+
export declare interface HtmlFile {
|
|
25
|
+
/** The raw HTML. */
|
|
26
|
+
contents: Buffer | string;
|
|
27
|
+
/** Path of the file; used as the `href` of the search result. */
|
|
21
28
|
relative: string;
|
|
22
29
|
}
|
|
30
|
+
/** @deprecated Use {@link HtmlFile}. */
|
|
31
|
+
export type ReadFileWithContents = HtmlFile;
|
|
32
|
+
/** Receives progress and diagnostic messages. `console` satisfies this interface. */
|
|
33
|
+
export declare interface Logger {
|
|
34
|
+
info(message: string): void;
|
|
35
|
+
warn(message: string): void;
|
|
36
|
+
}
|
|
37
|
+
/** The indexed fields. */
|
|
38
|
+
export type SearchField = "title" | "keywords" | "description" | "body";
|
|
39
|
+
/** Default for the `boosts` option. */
|
|
40
|
+
export declare const DEFAULT_BOOSTS: Readonly<Record<SearchField, number>>;
|
|
41
|
+
export declare interface SearchIndexOptions extends HrefOptions {
|
|
42
|
+
/** CSS selector of the element(s) whose text is indexed as body. Default: `"body"`. */
|
|
43
|
+
bodySelector?: string;
|
|
44
|
+
/**
|
|
45
|
+
* CSS selector of elements inside the body to leave out, e.g. navigation and footers
|
|
46
|
+
* that repeat on every page. Use `""` to exclude nothing. Default: `"nav, footer"`.
|
|
47
|
+
*/
|
|
48
|
+
excludeSelector?: string;
|
|
49
|
+
/**
|
|
50
|
+
* Skip pages with `<meta name="robots" content="noindex">` (or `none`), just like
|
|
51
|
+
* search engines do. Default: `true`.
|
|
52
|
+
*/
|
|
53
|
+
respectNoindex?: boolean;
|
|
54
|
+
/**
|
|
55
|
+
* Language of the content, as two-letter code. Selects stop words and stemmer.
|
|
56
|
+
* Anything other than `"en"` uses the matching plugin of
|
|
57
|
+
* [lunr-languages](https://github.com/MihaiValentin/lunr-languages), which the
|
|
58
|
+
* client must then load as well before calling `lunr.Index.load()`. Default: `"en"`.
|
|
59
|
+
*/
|
|
60
|
+
language?: string;
|
|
61
|
+
/**
|
|
62
|
+
* Relative weight of a match per field, so that a match in the title ranks above a
|
|
63
|
+
* match somewhere in the body text. Missing fields use the defaults. Default:
|
|
64
|
+
* `{title: 5, keywords: 3, description: 2, body: 1}`.
|
|
65
|
+
*/
|
|
66
|
+
boosts?: Partial<Record<SearchField, number>>;
|
|
67
|
+
/**
|
|
68
|
+
* `createFromGlob` only: resolve with an empty index instead of rejecting when the
|
|
69
|
+
* pattern matches no files. Default: `false`.
|
|
70
|
+
*/
|
|
71
|
+
allowEmpty?: boolean;
|
|
72
|
+
/**
|
|
73
|
+
* `createFromGlob` only: directory the pattern is resolved against. The paths of
|
|
74
|
+
* the matched files relative to it become the hrefs, so point it at the root of
|
|
75
|
+
* the built site. Default: `process.cwd()`.
|
|
76
|
+
*/
|
|
77
|
+
cwd?: string;
|
|
78
|
+
/** Where to report progress (one message per indexed file). Default: silent. */
|
|
79
|
+
logger?: Logger;
|
|
80
|
+
}
|
|
23
81
|
export declare class SearchIndex {
|
|
24
82
|
private readonly store;
|
|
25
83
|
private readonly index;
|
|
26
84
|
private constructor();
|
|
27
|
-
|
|
28
|
-
|
|
29
|
-
|
|
85
|
+
/**
|
|
86
|
+
* @param files Already extracted page information to index.
|
|
87
|
+
* @param options Only `language` and `boosts` apply here; the other options concern HTML parsing.
|
|
88
|
+
*/
|
|
89
|
+
static createFromInfo(files: IFileInformation[], options?: SearchIndexOptions): ISearchIndexResult;
|
|
90
|
+
/**
|
|
91
|
+
* @param files HTML documents to index.
|
|
92
|
+
* @param options Options, or - for backwards compatibility - just the body selector.
|
|
93
|
+
*/
|
|
94
|
+
static createFromHtml(files: HtmlFile[], options?: string | SearchIndexOptions): ISearchIndexResult;
|
|
95
|
+
/**
|
|
96
|
+
* Indexes all HTML files matching a glob pattern.
|
|
97
|
+
*
|
|
98
|
+
* @param pattern Glob pattern of the HTML files to index.
|
|
99
|
+
* @param options Options, or - for backwards compatibility - just the body selector.
|
|
100
|
+
* @returns The index and result store. Rejects if a file cannot be read.
|
|
101
|
+
*/
|
|
102
|
+
static createFromGlob(pattern: string, options?: string | SearchIndexOptions, ...legacyCallback: never[]): Promise<ISearchIndexResult>;
|
|
103
|
+
private static createFromGlobAsync;
|
|
30
104
|
private getResult;
|
|
31
105
|
}
|
|
32
|
-
export {};
|
package/js/index.js
CHANGED
|
@@ -33,60 +33,131 @@ var __importStar = (this && this.__importStar) || (function () {
|
|
|
33
33
|
};
|
|
34
34
|
})();
|
|
35
35
|
Object.defineProperty(exports, "__esModule", { value: true });
|
|
36
|
-
exports.SearchIndex = void 0;
|
|
36
|
+
exports.SearchIndex = exports.DEFAULT_BOOSTS = exports.DEFAULT_LANGUAGE = exports.DEFAULT_EXCLUDE_SELECTOR = void 0;
|
|
37
37
|
const cheerio = __importStar(require("cheerio"));
|
|
38
38
|
const glob_1 = require("glob");
|
|
39
|
-
const fs = __importStar(require("fs"));
|
|
39
|
+
const fs = __importStar(require("node:fs"));
|
|
40
|
+
const path = __importStar(require("node:path"));
|
|
40
41
|
const lunr = __importStar(require("lunr"));
|
|
42
|
+
const html_1 = require("./html");
|
|
43
|
+
const href_1 = require("./href");
|
|
44
|
+
const language_1 = require("./language");
|
|
45
|
+
var html_2 = require("./html");
|
|
46
|
+
Object.defineProperty(exports, "DEFAULT_EXCLUDE_SELECTOR", { enumerable: true, get: function () { return html_2.DEFAULT_EXCLUDE_SELECTOR; } });
|
|
47
|
+
var language_2 = require("./language");
|
|
48
|
+
Object.defineProperty(exports, "DEFAULT_LANGUAGE", { enumerable: true, get: function () { return language_2.DEFAULT_LANGUAGE; } });
|
|
49
|
+
const silentLogger = {
|
|
50
|
+
info: () => undefined,
|
|
51
|
+
warn: () => undefined,
|
|
52
|
+
};
|
|
53
|
+
/** Default for the `boosts` option. */
|
|
54
|
+
exports.DEFAULT_BOOSTS = Object.freeze({
|
|
55
|
+
title: 5,
|
|
56
|
+
keywords: 3,
|
|
57
|
+
description: 2,
|
|
58
|
+
body: 1,
|
|
59
|
+
});
|
|
60
|
+
function normalizeOptions(options) {
|
|
61
|
+
return typeof options === "string" ? { bodySelector: options } : { ...options };
|
|
62
|
+
}
|
|
41
63
|
class SearchIndex {
|
|
42
64
|
store;
|
|
43
65
|
index;
|
|
44
|
-
constructor(files) {
|
|
66
|
+
constructor(files, options) {
|
|
45
67
|
this.store = {};
|
|
46
68
|
const builder = new lunr.Builder();
|
|
47
|
-
|
|
48
|
-
|
|
49
|
-
|
|
50
|
-
|
|
69
|
+
// The same text processing lunr() sets up by default. A bare Builder has empty
|
|
70
|
+
// pipelines, which meant no stemming, no stop word removal and punctuation
|
|
71
|
+
// sticking to words ("konzert," / "page:"). The search pipeline is serialized
|
|
72
|
+
// into the index, so lunr applies the stemmer to queries on the client as well.
|
|
73
|
+
builder.pipeline.add(lunr.trimmer, lunr.stopWordFilter, lunr.stemmer);
|
|
74
|
+
builder.searchPipeline.add(lunr.stemmer);
|
|
75
|
+
const plugin = (0, language_1.languagePlugin)(options.language ?? language_1.DEFAULT_LANGUAGE);
|
|
76
|
+
if (plugin) {
|
|
77
|
+
builder.use(plugin); // replaces both pipelines with the language specific ones
|
|
78
|
+
}
|
|
79
|
+
const boosts = { ...exports.DEFAULT_BOOSTS, ...options.boosts };
|
|
80
|
+
for (const field of Object.keys(exports.DEFAULT_BOOSTS)) {
|
|
81
|
+
builder.field(field, { boost: boosts[field] });
|
|
82
|
+
}
|
|
51
83
|
builder.ref("href");
|
|
52
84
|
files.forEach((info) => {
|
|
85
|
+
if (Object.hasOwn(this.store, info.href)) {
|
|
86
|
+
throw new Error(`Duplicate href "${info.href}": every document needs a unique href.`);
|
|
87
|
+
}
|
|
53
88
|
this.store[info.href] = {
|
|
54
89
|
description: info.description,
|
|
55
90
|
title: info.title,
|
|
56
91
|
};
|
|
57
|
-
|
|
58
|
-
|
|
92
|
+
// keywords are a comma separated list, but lunr only splits on whitespace and hyphens
|
|
93
|
+
builder.add({ ...info, keywords: info.keywords?.replace(/[,;]/g, " ") });
|
|
94
|
+
});
|
|
59
95
|
this.index = builder.build();
|
|
60
96
|
}
|
|
61
|
-
|
|
62
|
-
|
|
97
|
+
/**
|
|
98
|
+
* @param files Already extracted page information to index.
|
|
99
|
+
* @param options Only `language` and `boosts` apply here; the other options concern HTML parsing.
|
|
100
|
+
*/
|
|
101
|
+
static createFromInfo(files, options) {
|
|
102
|
+
return new SearchIndex(files, normalizeOptions(options)).getResult();
|
|
63
103
|
}
|
|
64
|
-
|
|
65
|
-
|
|
66
|
-
|
|
104
|
+
/**
|
|
105
|
+
* @param files HTML documents to index.
|
|
106
|
+
* @param options Options, or - for backwards compatibility - just the body selector.
|
|
107
|
+
*/
|
|
108
|
+
static createFromHtml(files, options) {
|
|
109
|
+
const normalized = normalizeOptions(options);
|
|
110
|
+
const { bodySelector, excludeSelector, respectNoindex = true, logger = silentLogger } = normalized;
|
|
111
|
+
const infos = [];
|
|
112
|
+
for (const file of files) {
|
|
67
113
|
const dom = cheerio.load(file.contents.toString());
|
|
68
|
-
|
|
69
|
-
|
|
70
|
-
|
|
71
|
-
|
|
72
|
-
|
|
73
|
-
|
|
74
|
-
|
|
75
|
-
|
|
76
|
-
|
|
114
|
+
if (respectNoindex && (0, html_1.isNoindex)(dom)) {
|
|
115
|
+
logger.info(`Skipping ${file.relative} (robots noindex)`);
|
|
116
|
+
continue;
|
|
117
|
+
}
|
|
118
|
+
logger.info(`Indexing ${file.relative}`);
|
|
119
|
+
const metadata = (0, html_1.extractMetadata)(dom);
|
|
120
|
+
if (!metadata.title) {
|
|
121
|
+
logger.warn(`${file.relative} has no <title>, og:title or <h1> - its search result will have an empty title`);
|
|
122
|
+
}
|
|
123
|
+
infos.push({
|
|
124
|
+
...metadata,
|
|
125
|
+
body: (0, html_1.extractText)(dom, bodySelector || "body", excludeSelector),
|
|
126
|
+
href: (0, href_1.toHref)(file.relative, normalized),
|
|
127
|
+
});
|
|
128
|
+
}
|
|
129
|
+
return SearchIndex.createFromInfo(infos, normalized);
|
|
77
130
|
}
|
|
78
|
-
|
|
79
|
-
|
|
80
|
-
|
|
81
|
-
|
|
82
|
-
|
|
83
|
-
|
|
84
|
-
|
|
85
|
-
|
|
86
|
-
|
|
87
|
-
|
|
88
|
-
|
|
89
|
-
|
|
131
|
+
/**
|
|
132
|
+
* Indexes all HTML files matching a glob pattern.
|
|
133
|
+
*
|
|
134
|
+
* @param pattern Glob pattern of the HTML files to index.
|
|
135
|
+
* @param options Options, or - for backwards compatibility - just the body selector.
|
|
136
|
+
* @returns The index and result store. Rejects if a file cannot be read.
|
|
137
|
+
*/
|
|
138
|
+
static createFromGlob(pattern, options, ...legacyCallback) {
|
|
139
|
+
if (legacyCallback.length > 0) {
|
|
140
|
+
// 2.x took a callback as third argument. Silently ignoring it would mean the
|
|
141
|
+
// caller's index is simply never written - fail loudly instead.
|
|
142
|
+
throw new TypeError("SearchIndex.createFromGlob() no longer accepts a callback; it returns a Promise. "
|
|
143
|
+
+ "See https://github.com/Tiliavir/mvw-search-index/blob/main/UPGRADING.md");
|
|
144
|
+
}
|
|
145
|
+
return SearchIndex.createFromGlobAsync(pattern, normalizeOptions(options));
|
|
146
|
+
}
|
|
147
|
+
static async createFromGlobAsync(pattern, options) {
|
|
148
|
+
const cwd = path.resolve(options.cwd ?? ".");
|
|
149
|
+
// glob's result order depends on the file system - sort for reproducible output
|
|
150
|
+
const files = (await (0, glob_1.glob)(pattern, { cwd, dotRelative: false, nodir: true, posix: true }))
|
|
151
|
+
.sort((a, b) => a.localeCompare(b, "en"));
|
|
152
|
+
if (files.length === 0 && !options.allowEmpty) {
|
|
153
|
+
throw new Error(`No files match "${pattern}" in ${cwd}. `
|
|
154
|
+
+ "Set the allowEmpty option to create an empty index anyway.");
|
|
155
|
+
}
|
|
156
|
+
const readFiles = await Promise.all(files.map(async (file) => ({
|
|
157
|
+
relative: file,
|
|
158
|
+
contents: await fs.promises.readFile(path.join(cwd, file)),
|
|
159
|
+
})));
|
|
160
|
+
return SearchIndex.createFromHtml(readFiles, options);
|
|
90
161
|
}
|
|
91
162
|
getResult() {
|
|
92
163
|
return {
|
|
@@ -96,4 +167,3 @@ class SearchIndex {
|
|
|
96
167
|
}
|
|
97
168
|
}
|
|
98
169
|
exports.SearchIndex = SearchIndex;
|
|
99
|
-
//# sourceMappingURL=index.js.map
|
package/js/language.d.ts
ADDED
|
@@ -0,0 +1,8 @@
|
|
|
1
|
+
import type * as Lunr from "lunr";
|
|
2
|
+
/** lunr's built-in language; needs no plugin. */
|
|
3
|
+
export declare const DEFAULT_LANGUAGE = "en";
|
|
4
|
+
/**
|
|
5
|
+
* Returns the lunr-languages plugin for `language` (an ISO 639-1 code such as
|
|
6
|
+
* "de"), registering it with lunr on first use, or `undefined` for English.
|
|
7
|
+
*/
|
|
8
|
+
export declare function languagePlugin(language: string): Lunr.Builder.Plugin | undefined;
|
package/js/language.js
ADDED
|
@@ -0,0 +1,45 @@
|
|
|
1
|
+
"use strict";
|
|
2
|
+
Object.defineProperty(exports, "__esModule", { value: true });
|
|
3
|
+
exports.DEFAULT_LANGUAGE = void 0;
|
|
4
|
+
exports.languagePlugin = languagePlugin;
|
|
5
|
+
// lunr-languages plugins register themselves by mutating the lunr module object, so
|
|
6
|
+
// they need the actual CommonJS export - an ES namespace wrapper (as created by
|
|
7
|
+
// bundlers and test runners for `import * as`) would silently drop the additions.
|
|
8
|
+
// eslint-disable-next-line @typescript-eslint/no-require-imports
|
|
9
|
+
const lunr = require("lunr");
|
|
10
|
+
/** lunr's built-in language; needs no plugin. */
|
|
11
|
+
exports.DEFAULT_LANGUAGE = "en";
|
|
12
|
+
let stemmerSupportLoaded = false;
|
|
13
|
+
/**
|
|
14
|
+
* Returns the lunr-languages plugin for `language` (an ISO 639-1 code such as
|
|
15
|
+
* "de"), registering it with lunr on first use, or `undefined` for English.
|
|
16
|
+
*/
|
|
17
|
+
function languagePlugin(language) {
|
|
18
|
+
if (language === exports.DEFAULT_LANGUAGE) {
|
|
19
|
+
return undefined;
|
|
20
|
+
}
|
|
21
|
+
if (!/^[a-z]{2}$/.test(language)) {
|
|
22
|
+
throw new Error(`Invalid language "${language}": expected a two-letter code such as "de" or "fr".`);
|
|
23
|
+
}
|
|
24
|
+
const registry = lunr;
|
|
25
|
+
if (!registry[language]) {
|
|
26
|
+
try {
|
|
27
|
+
if (!stemmerSupportLoaded) {
|
|
28
|
+
// eslint-disable-next-line @typescript-eslint/no-require-imports -- loaded on demand, synchronously
|
|
29
|
+
require("lunr-languages/lunr.stemmer.support")(lunr);
|
|
30
|
+
stemmerSupportLoaded = true;
|
|
31
|
+
}
|
|
32
|
+
// eslint-disable-next-line @typescript-eslint/no-require-imports -- loaded on demand, synchronously
|
|
33
|
+
require(`lunr-languages/lunr.${language}`)(lunr);
|
|
34
|
+
}
|
|
35
|
+
catch (err) {
|
|
36
|
+
throw new Error(`Unsupported language "${language}": ${err instanceof Error ? err.message : err}. `
|
|
37
|
+
+ "See https://github.com/MihaiValentin/lunr-languages for the supported languages.", { cause: err });
|
|
38
|
+
}
|
|
39
|
+
}
|
|
40
|
+
const plugin = registry[language];
|
|
41
|
+
if (!plugin) {
|
|
42
|
+
throw new Error(`Unsupported language "${language}".`);
|
|
43
|
+
}
|
|
44
|
+
return plugin;
|
|
45
|
+
}
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "mvw-search-index",
|
|
3
|
-
"version": "
|
|
3
|
+
"version": "3.0.1",
|
|
4
4
|
"description": "Module to generate the search index using lunr.",
|
|
5
5
|
"keywords": [
|
|
6
6
|
"website",
|
|
@@ -22,27 +22,36 @@
|
|
|
22
22
|
},
|
|
23
23
|
"main": "js/index.js",
|
|
24
24
|
"types": "js/index.d.ts",
|
|
25
|
+
"files": [
|
|
26
|
+
"js/**/*.js",
|
|
27
|
+
"js/**/*.d.ts"
|
|
28
|
+
],
|
|
29
|
+
"engines": {
|
|
30
|
+
"node": ">=22.12.0"
|
|
31
|
+
},
|
|
25
32
|
"devDependencies": {
|
|
26
|
-
"@
|
|
33
|
+
"@eslint/js": "10.0.1",
|
|
27
34
|
"@types/lunr": "2.3.7",
|
|
28
|
-
"@types/node": "
|
|
29
|
-
"
|
|
30
|
-
"@typescript-eslint/parser": "8.58.0",
|
|
31
|
-
"eslint": "10.2.0",
|
|
35
|
+
"@types/node": "26.6.4",
|
|
36
|
+
"eslint": "10.12.0",
|
|
32
37
|
"serve": "14.2.6",
|
|
33
|
-
"typescript": "6.0.
|
|
34
|
-
"
|
|
38
|
+
"typescript": "6.0.3",
|
|
39
|
+
"typescript-eslint": "8.71.1",
|
|
40
|
+
"vitest": "5.0.3"
|
|
35
41
|
},
|
|
36
42
|
"dependencies": {
|
|
37
43
|
"cheerio": "1.2.0",
|
|
38
|
-
"commander": "
|
|
44
|
+
"commander": "15.0.0",
|
|
39
45
|
"glob": "13.0.6",
|
|
40
|
-
"lunr": "2.3.9"
|
|
46
|
+
"lunr": "2.3.9",
|
|
47
|
+
"lunr-languages": "1.22.0"
|
|
41
48
|
},
|
|
42
49
|
"scripts": {
|
|
50
|
+
"prepare": "npm run build",
|
|
43
51
|
"build": "tsc -p ./ts",
|
|
44
|
-
"
|
|
45
|
-
"
|
|
52
|
+
"lint": "eslint . && tsc -p spec",
|
|
53
|
+
"test": "npm run lint && npm run build && vitest run",
|
|
54
|
+
"preserve": "node js/cli.js '**/*.html' docs/index.json 'body.to-be-indexed' --cwd docs",
|
|
46
55
|
"serve": "serve docs"
|
|
47
56
|
},
|
|
48
57
|
"repository": {
|
package/README.Release.md
DELETED
package/eslint.config.mjs
DELETED
|
@@ -1,15 +0,0 @@
|
|
|
1
|
-
import tsEsLint from "@typescript-eslint/eslint-plugin";
|
|
2
|
-
import tsParser from "@typescript-eslint/parser";
|
|
3
|
-
|
|
4
|
-
export default [
|
|
5
|
-
{
|
|
6
|
-
files: ["**/*.ts"],
|
|
7
|
-
plugins: {
|
|
8
|
-
"@typescript-eslint": tsEsLint,
|
|
9
|
-
},
|
|
10
|
-
|
|
11
|
-
languageOptions: {
|
|
12
|
-
parser: tsParser,
|
|
13
|
-
}
|
|
14
|
-
},
|
|
15
|
-
];
|
package/js/cli.js.map
DELETED
|
@@ -1 +0,0 @@
|
|
|
1
|
-
{"version":3,"file":"cli.js","sourceRoot":"","sources":["../ts/cli.ts"],"names":[],"mappings":";;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;AAEA,yCAAoC;AACpC,uCAAyB;AAEzB,mCAAsC;AAEtC,mBAAO;KACJ,OAAO,CAAC,OAAO,CAAC;KAChB,SAAS,CAAC,8BAA8B,CAAC;KACzC,MAAM,CAAC,CAAC,IAAI,EAAE,IAAI,EAAE,YAAY,EAAE,EAAE;IACnC,mBAAW,CAAC,cAAc,CAAC,IAAI,EAAE,YAAY,EAAE,CAAC,KAAK,EAAE,EAAE,CACvD,EAAE,CAAC,aAAa,CAAC,IAAI,IAAI,cAAc,EAAE,IAAI,CAAC,SAAS,CAAC,KAAK,CAAC,CAAC,CAChE,CAAC;AACJ,CAAC,CAAC;KACD,KAAK,CAAC,OAAO,CAAC,IAAI,CAAC,CAAC"}
|
package/js/index.js.map
DELETED
|
@@ -1 +0,0 @@
|
|
|
1
|
-
{"version":3,"file":"index.js","sourceRoot":"","sources":["../ts/index.ts"],"names":[],"mappings":";;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;AAAA,iDAAmC;AACnC,+BAA0B;AAC1B,uCAAyB;AACzB,2CAA6B;AA4B7B,MAAa,WAAW;IACL,KAAK,CAAe;IACpB,KAAK,CAAa;IAEnC,YAAoB,KAAyB;QAC3C,IAAI,CAAC,KAAK,GAAG,EAAE,CAAC;QAChB,MAAM,OAAO,GAAiB,IAAI,IAAI,CAAC,OAAO,EAAE,CAAC;QACjD,OAAO,CAAC,KAAK,CAAC,OAAO,CAAC,CAAC;QACvB,OAAO,CAAC,KAAK,CAAC,UAAU,CAAC,CAAC;QAC1B,OAAO,CAAC,KAAK,CAAC,aAAa,CAAC,CAAC;QAC7B,OAAO,CAAC,KAAK,CAAC,MAAM,CAAC,CAAC;QACtB,OAAO,CAAC,GAAG,CAAC,MAAM,CAAC,CAAC;QAEpB,KAAK,CAAC,OAAO,CAAC,CAAC,IAAsB,EAAQ,EAAE;YAC7C,IAAI,CAAC,KAAK,CAAC,IAAI,CAAC,IAAI,CAAC,GAAG;gBACtB,WAAW,EAAE,IAAI,CAAC,WAAW;gBAC7B,KAAK,EAAE,IAAI,CAAC,KAAK;aAClB,CAAC;YACF,OAAO,CAAC,GAAG,CAAC,IAAI,CAAC,CAAC;QACpB,CAAC,EAAE,OAAO,CAAC,CAAC;QACZ,IAAI,CAAC,KAAK,GAAG,OAAO,CAAC,KAAK,EAAE,CAAC;IAC/B,CAAC;IAEM,MAAM,CAAC,cAAc,CAAC,KAAyB;QACpD,OAAO,IAAI,WAAW,CAAC,KAAK,CAAC,CAAC,SAAS,EAAE,CAAC;IAC5C,CAAC;IAEM,MAAM,CAAC,cAAc,CAAC,KAA6B,EAAE,eAAuB,MAAM;QACvF,MAAM,KAAK,GAAuB,KAAK,CAAC,GAAG,CAAC,CAAC,IAAI,EAAE,EAAE;YACnD,OAAO,CAAC,IAAI,CAAC,IAAI,CAAC,QAAQ,CAAC,CAAC;YAC5B,MAAM,GAAG,GAAG,OAAO,CAAC,IAAI,CAAC,IAAI,CAAC,QAAQ,CAAC,QAAQ,EAAE,CAAC,CAAC;YACnD,OAAO;gBACL,IAAI,EAAE,GAAG,CAAC,YAAY,IAAI,MAAM,CAAC,CAAC,IAAI,EAAE,CAAC,OAAO,CAAC,QAAQ,EAAE,GAAG,CAAC;gBAC/D,IAAI,EAAE,IAAI,CAAC,QAAQ;gBACnB,WAAW,EAAE,GAAG,CAAC,0BAA0B,CAAC,CAAC,IAAI,CAAC,SAAS,CAAC;gBAC5D,QAAQ,EAAE,GAAG,CAAC,uBAAuB,CAAC,CAAC,IAAI,CAAC,SAAS,CAAC;gBACtD,KAAK,EAAE,GAAG,CAAC,YAAY,CAAC,CAAC,IAAI,EAAE;aAChC,CAAC;QACJ,CAAC,CAAC,CAAC;QAEH,OAAO,WAAW,CAAC,cAAc,CAAC,KAAK,CAAC,CAAC;IAC3C,CAAC;IAEM,MAAM,CAAC,cAAc,CAAC,OAAe,EACf,YAAoB,EACpB,EAAuC;QAClE,IAAA,WAAI,EAAC,OAAO,EAAE;YACZ,WAAW,EAAE,KAAK;SACnB,CAAC,CAAC,IAAI,CAAC,KAAK,CAAC,EAAE;YACZ,MAAM,SAAS,GAA2B,KAAK,CAAC,GAAG,CAAC,CAAC,IAAI,EAAE,EAAE,CAAC,CAAC;gBAC7D,QAAQ,EAAE,IAAI;gBACd,QAAQ,EAAE,EAAE,CAAC,YAAY,CAAC,IAAI,CAAC;aAChC,CAAC,CAAC,CAAC;YACJ,EAAE,CAAC,WAAW,CAAC,cAAc,CAAC,SAAS,EAAE,YAAY,CAAC,CAAC,CAAC;QAC1D,CAAC,CACF,CAAC,KAAK,CAAC,GAAG,CAAC,EAAE;YACZ,MAAM,GAAG,CAAC;QACZ,CAAC,CAAC,CAAC;IACL,CAAC;IAEO,SAAS;QACf,OAAO;YACL,KAAK,EAAE,IAAI,CAAC,KAAK;YACjB,KAAK,EAAE,IAAI,CAAC,KAAK;SAClB,CAAC;IACJ,CAAC;CACF;AAlED,kCAkEC"}
|