pagefind 1.5.2.1 → 1.5.2.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +5 -0
- data/README.md +286 -3
- data/lib/pagefind/index.rb +162 -0
- data/lib/pagefind/service.rb +111 -0
- data/lib/pagefind/version.rb +1 -1
- data/lib/pagefind.rb +5 -0
- data/rakelib/package.rake +6 -0
- metadata +6 -4
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: 29fc9d08b459ad0bb8185f81b9b66fef23fc4e2616d1dc0472d4aa7de0754a2b
|
|
4
|
+
data.tar.gz: 89d3be5213c8b96fe48c32d6a6708344acedf2ed370db437c2ad3d988ca0fe90
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: c7a84061ed37ee139b733df2794ff830e63240ee9ffe1e7084bca1c58c5cb92c65987d995d880e8d2eda6570237dc61356625305bd8e37aca685047c095e6e72
|
|
7
|
+
data.tar.gz: c70307c54f6fc2b519c276c34e604cf5a4d4b236c1929d3132d0d2892dc4243d09e354e24fc6c0741ed4e72860d46420dafd3b65cb50c949b33970a67356084d
|
data/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,10 @@
|
|
|
1
1
|
# pagefind-ruby changelog
|
|
2
2
|
|
|
3
|
+
## v1.5.2.2
|
|
4
|
+
|
|
5
|
+
* Add `Pagefind::Index` and `Pagefind::Service` @bangseongbeom
|
|
6
|
+
* Support `PAGEFIND_EXTENDED_BINARY_PATH` and `PAGEFIND_BINARY_PATH` @bangseongbeom
|
|
7
|
+
|
|
3
8
|
## v1.5.2.1
|
|
4
9
|
|
|
5
10
|
* Ship the pagefind_extended release, as Pagefind's npm package does @bangseongbeom
|
data/README.md
CHANGED
|
@@ -4,9 +4,9 @@
|
|
|
4
4
|
[](https://github.com/standardrb/standard)
|
|
5
5
|
[](https://badge.fury.io/rb/pagefind)
|
|
6
6
|
|
|
7
|
-
A self-contained `pagefind` executable, wrapped up in a ruby gem. That's it. Nothing else.
|
|
7
|
+
A self-contained `pagefind` executable, with an indexing API, wrapped up in a ruby gem. That's it. Nothing else.
|
|
8
8
|
|
|
9
|
-
This gem is based on [tailwindcss-ruby](https://github.com/flavorjones/tailwindcss-ruby). Much of the packaging approach, code, and documentation was adapted from it.
|
|
9
|
+
This gem is based on [tailwindcss-ruby](https://github.com/flavorjones/tailwindcss-ruby). Much of the packaging approach, code, and documentation was adapted from it. The indexing API was adapted from Pagefind's [Python](https://pagefind.app/docs/py-api/) and [NodeJS](https://pagefind.app/docs/node-api/) wrappers.
|
|
10
10
|
|
|
11
11
|
|
|
12
12
|
## Installation
|
|
@@ -37,7 +37,7 @@ gem install pagefind
|
|
|
37
37
|
|
|
38
38
|
### Using a local installation of `pagefind`
|
|
39
39
|
|
|
40
|
-
If you are not able to use the vendored precompiled binaries (for example, if you're on an unsupported platform), you can use a [local installation](https://pagefind.app/docs/installation/) of the `pagefind_extended` or `pagefind` binary by setting an environment variable named `PAGEFIND_INSTALL_DIR` to the directory path containing the binary.
|
|
40
|
+
If you are not able to use the vendored precompiled binaries (for example, if you're on an unsupported platform), you can use a [local installation](https://pagefind.app/docs/installation/) of the `pagefind_extended` or `pagefind` binary by setting an environment variable named `PAGEFIND_EXTENDED_BINARY_PATH` or `PAGEFIND_BINARY_PATH` to the path of the binary, or `PAGEFIND_INSTALL_DIR` to the directory path containing the binary.
|
|
41
41
|
|
|
42
42
|
For example, if you've [built Pagefind from source](https://pagefind.app/docs/installation/#building-from-source) with `cargo install pagefind` so that the binary is found at `$HOME/.cargo/bin/pagefind`, then you should set your environment variable like so:
|
|
43
43
|
|
|
@@ -99,6 +99,289 @@ Options:
|
|
|
99
99
|
```
|
|
100
100
|
|
|
101
101
|
|
|
102
|
+
## Indexing API
|
|
103
|
+
|
|
104
|
+
This gem also provides an interface to the indexing binary as a Ruby library you can require.
|
|
105
|
+
|
|
106
|
+
There are situations where using this library is beneficial:
|
|
107
|
+
|
|
108
|
+
- Integrating Pagefind into an existing Ruby project, e.g. writing a plugin for a static site generator that can pass in-memory HTML files to Pagefind.
|
|
109
|
+
Pagefind can also return the search index in-memory, to be hosted via the dev mode alongside the files.
|
|
110
|
+
- Users looking to index their site and augment that index with extra non-HTML pages can run a standard Pagefind crawl with [`add_directory`](#indexadd_directory) and augment it with [`add_custom_record`](#indexadd_custom_record).
|
|
111
|
+
- Users looking to use Pagefind's engine for searching miscellaneous content such as PDFs or subtitles, where [`add_custom_record`](#indexadd_custom_record) can be used to build the entire index from scratch.
|
|
112
|
+
|
|
113
|
+
### Example usage
|
|
114
|
+
|
|
115
|
+
<!-- this example is copied verbatim from test/integration.rb -->
|
|
116
|
+
|
|
117
|
+
``` ruby
|
|
118
|
+
require "pagefind"
|
|
119
|
+
|
|
120
|
+
html_content = <<~HTML
|
|
121
|
+
<html>
|
|
122
|
+
<body>
|
|
123
|
+
<main>
|
|
124
|
+
<h1>Example HTML</h1>
|
|
125
|
+
<p>This is an example HTML page.</p>
|
|
126
|
+
</main>
|
|
127
|
+
</body>
|
|
128
|
+
</html>
|
|
129
|
+
HTML
|
|
130
|
+
|
|
131
|
+
Pagefind::Index.open(root_selector: "main", logfile: "index.log", output_path: "./output", verbose: true) do |index|
|
|
132
|
+
new_file = index.add_html_file(
|
|
133
|
+
content: html_content,
|
|
134
|
+
url: "https://example.com",
|
|
135
|
+
source_path: "other/example.html"
|
|
136
|
+
)
|
|
137
|
+
new_record = index.add_custom_record(
|
|
138
|
+
url: "/elephants/",
|
|
139
|
+
content: "Some testing content regarding elephants",
|
|
140
|
+
language: "en",
|
|
141
|
+
meta: {title: "Elephants"}
|
|
142
|
+
)
|
|
143
|
+
new_dir = index.add_directory("./public")
|
|
144
|
+
pp new_file, new_record, new_dir
|
|
145
|
+
|
|
146
|
+
index.get_files.each do |file|
|
|
147
|
+
puts "#{file[:content].bytesize.to_s.rjust(10)}B #{file[:path]}"
|
|
148
|
+
end
|
|
149
|
+
end
|
|
150
|
+
```
|
|
151
|
+
|
|
152
|
+
All methods are synchronous: each one blocks until the native Pagefind binary running in the background responds. Pagefind handles one request at a time, so threads sharing an index or a service take turns.
|
|
153
|
+
|
|
154
|
+
Responses are returned as hashes with symbol keys.
|
|
155
|
+
|
|
156
|
+
### Pagefind::Index
|
|
157
|
+
|
|
158
|
+
`Pagefind::Index` manages a Pagefind index.
|
|
159
|
+
|
|
160
|
+
`Pagefind::Index.open` yields the index to a block and returns the block's value.
|
|
161
|
+
Entering the block starts a backing Pagefind service and creates an in-memory index in the backing service.
|
|
162
|
+
Exiting the block writes the in-memory index to disk and then shuts down the backing Pagefind service.
|
|
163
|
+
|
|
164
|
+
``` ruby
|
|
165
|
+
Pagefind::Index.open do |index| # open the index
|
|
166
|
+
# update the index
|
|
167
|
+
end
|
|
168
|
+
# the index is closed here and files are written to disk.
|
|
169
|
+
```
|
|
170
|
+
|
|
171
|
+
Each method of `Pagefind::Index` that talks to the backing Pagefind service can raise `Pagefind::ServiceError`.
|
|
172
|
+
If an exception is raised inside `Pagefind::Index.open`'s block, the block exits without writing the index files to disk.
|
|
173
|
+
|
|
174
|
+
``` ruby
|
|
175
|
+
Pagefind::Index.open do |index| # open the index
|
|
176
|
+
index.add_directory("./public")
|
|
177
|
+
raise "not today"
|
|
178
|
+
end
|
|
179
|
+
# the index closes without writing anything to disk
|
|
180
|
+
```
|
|
181
|
+
|
|
182
|
+
`Pagefind::Index.open` optionally takes keyword arguments that can apply parts of the [Pagefind CLI config](https://pagefind.app/docs/config-options/). The options available at this level are:
|
|
183
|
+
|
|
184
|
+
``` ruby
|
|
185
|
+
Pagefind::Index.open(
|
|
186
|
+
root_selector: "main",
|
|
187
|
+
exclude_selectors: ["nav"],
|
|
188
|
+
force_language: "en",
|
|
189
|
+
include_characters: "._",
|
|
190
|
+
verbose: true,
|
|
191
|
+
logfile: "index.log",
|
|
192
|
+
keep_index_url: true,
|
|
193
|
+
write_playground: true,
|
|
194
|
+
output_path: "./output"
|
|
195
|
+
) do |index|
|
|
196
|
+
# ...
|
|
197
|
+
end
|
|
198
|
+
```
|
|
199
|
+
|
|
200
|
+
See the relevant documentation for these configuration options in the [Configuring the Pagefind CLI](https://pagefind.app/docs/config-options/) documentation.
|
|
201
|
+
|
|
202
|
+
### index.add_directory
|
|
203
|
+
|
|
204
|
+
Indexes a directory from disk using the standard Pagefind indexing behaviour.
|
|
205
|
+
This is equivalent to running the Pagefind binary with `--site <dir>`.
|
|
206
|
+
|
|
207
|
+
``` ruby
|
|
208
|
+
# Index all the HTML files in the public directory
|
|
209
|
+
indexed_dir = index.add_directory("./public")
|
|
210
|
+
page_count = indexed_dir[:page_count] # Integer
|
|
211
|
+
```
|
|
212
|
+
|
|
213
|
+
If the `path` provided is relative, it will be relative to the current working directory of your Ruby process. Both `String` and `Pathname` are accepted.
|
|
214
|
+
|
|
215
|
+
``` ruby
|
|
216
|
+
# Index files in a directory matching a given glob pattern.
|
|
217
|
+
indexed_dir = index.add_directory("./public", glob: "**/*.{html}")
|
|
218
|
+
```
|
|
219
|
+
|
|
220
|
+
Optionally, a custom `glob` can be supplied which controls which files Pagefind will consume within the directory. The default is shown, and the `glob` option can be omitted entirely.
|
|
221
|
+
See [Wax patterns documentation](https://github.com/olson-sean-k/wax#patterns) for more details.
|
|
222
|
+
|
|
223
|
+
### index.add_html_file
|
|
224
|
+
|
|
225
|
+
Adds a virtual HTML file to the Pagefind index. Useful for files that don't exist on disk, for example a static site generator that is serving files from memory.
|
|
226
|
+
|
|
227
|
+
``` ruby
|
|
228
|
+
html_content = <<~HTML
|
|
229
|
+
<html lang="en"><body>
|
|
230
|
+
<h1>A Full HTML Document</h1>
|
|
231
|
+
<p> ... </p>
|
|
232
|
+
</body></html>
|
|
233
|
+
HTML
|
|
234
|
+
|
|
235
|
+
# Index a file as if Pagefind was indexing from disk
|
|
236
|
+
new_file = index.add_html_file(
|
|
237
|
+
content: html_content,
|
|
238
|
+
source_path: "other/example.html"
|
|
239
|
+
)
|
|
240
|
+
|
|
241
|
+
# Index HTML content, giving it a specific URL
|
|
242
|
+
new_file = index.add_html_file(
|
|
243
|
+
content: html_content,
|
|
244
|
+
url: "https://example.com"
|
|
245
|
+
)
|
|
246
|
+
```
|
|
247
|
+
|
|
248
|
+
The `source_path` should represent the path of this HTML file if it were to exist on disk. Pagefind will use this path to generate the URL. It should be relative, or absolute to a path within the current working directory.
|
|
249
|
+
|
|
250
|
+
Instead of `source_path`, a `url` may be supplied to explicitly set the URL of this search result.
|
|
251
|
+
|
|
252
|
+
The `content` should be the full HTML source, including the outer `<html> </html>` tags. This will be run through Pagefind's standard HTML indexing process, and should contain any required Pagefind attributes to control behaviour.
|
|
253
|
+
|
|
254
|
+
If successful, a hash is returned containing metadata about the completed indexing.
|
|
255
|
+
|
|
256
|
+
### index.add_custom_record
|
|
257
|
+
|
|
258
|
+
Adds a direct record to the Pagefind index.
|
|
259
|
+
Useful for adding non-HTML content to the search results.
|
|
260
|
+
|
|
261
|
+
``` ruby
|
|
262
|
+
custom_record = index.add_custom_record(
|
|
263
|
+
url: "/contact/",
|
|
264
|
+
content: "My raw content to be indexed for search. " \
|
|
265
|
+
"Will be lightly processed by Pagefind.",
|
|
266
|
+
language: "en",
|
|
267
|
+
meta: {
|
|
268
|
+
title: "Contact",
|
|
269
|
+
category: "Landing Page"
|
|
270
|
+
},
|
|
271
|
+
filters: {tags: ["landing", "company"]},
|
|
272
|
+
sort: {weight: "20"}
|
|
273
|
+
)
|
|
274
|
+
|
|
275
|
+
page_word_count = custom_record[:page_word_count] # Integer
|
|
276
|
+
page_url = custom_record[:page_url] # String
|
|
277
|
+
page_meta = custom_record[:page_meta] # Hash[Symbol, String]
|
|
278
|
+
```
|
|
279
|
+
|
|
280
|
+
The `url`, `content`, and `language` keyword arguments are all required. `language` should be an [ISO 639-1 code](https://en.wikipedia.org/wiki/List_of_ISO_639-1_codes).
|
|
281
|
+
|
|
282
|
+
`meta` is optional, and is strictly a flat hash of keys to string values.
|
|
283
|
+
See the [Metadata documentation](https://pagefind.app/docs/metadata/) for semantics.
|
|
284
|
+
|
|
285
|
+
`filters` is optional, and is strictly a flat hash of keys to arrays of string values.
|
|
286
|
+
See the [Filters documentation](https://pagefind.app/docs/filtering/) for semantics.
|
|
287
|
+
|
|
288
|
+
`sort` is optional, and is strictly a flat hash of keys to string values.
|
|
289
|
+
See the [Sort documentation](https://pagefind.app/docs/sorts/) for semantics.
|
|
290
|
+
*When Pagefind is processing an index, number-like strings will be sorted numerically rather than alphabetically. As such, the value passed in should be `"20"` and not `20`*
|
|
291
|
+
|
|
292
|
+
If successful, a hash is returned containing metadata about the completed indexing.
|
|
293
|
+
|
|
294
|
+
### index.get_files
|
|
295
|
+
|
|
296
|
+
Gets raw data of all files in the Pagefind index.
|
|
297
|
+
Useful for integrating a Pagefind index into the development mode of a static site generator and hosting these files yourself.
|
|
298
|
+
|
|
299
|
+
This method returns every file at once, which can be a lot of data.
|
|
300
|
+
|
|
301
|
+
``` ruby
|
|
302
|
+
index.get_files.each do |file|
|
|
303
|
+
path = file[:path] # String
|
|
304
|
+
content = file[:content] # binary String
|
|
305
|
+
# ...
|
|
306
|
+
end
|
|
307
|
+
```
|
|
308
|
+
|
|
309
|
+
### index.write_files
|
|
310
|
+
|
|
311
|
+
Calling `index.write_files` writes the index files to disk, as they would be written when running the standard Pagefind binary directly.
|
|
312
|
+
|
|
313
|
+
Exiting `Pagefind::Index.open`'s block automatically calls `index.write_files`, so calling this method is not necessary in normal operation.
|
|
314
|
+
|
|
315
|
+
Calling this method won't prevent files being written when the block exits, which may cause duplicate files to be written.
|
|
316
|
+
If calling this method manually, you probably want to also call `index.delete_index`.
|
|
317
|
+
|
|
318
|
+
``` ruby
|
|
319
|
+
Pagefind::Index.open(output_path: "./public/pagefind") do |index|
|
|
320
|
+
# ... add content to index
|
|
321
|
+
|
|
322
|
+
# write files to the configured output path for the index:
|
|
323
|
+
index.write_files
|
|
324
|
+
|
|
325
|
+
# write files to a different output path:
|
|
326
|
+
index.write_files("./custom/pagefind")
|
|
327
|
+
|
|
328
|
+
# prevent also writing files when exiting the block:
|
|
329
|
+
index.delete_index
|
|
330
|
+
end
|
|
331
|
+
```
|
|
332
|
+
|
|
333
|
+
The output path should contain the path to the desired Pagefind bundle directory. If relative, it is relative to the current working directory of your Ruby process.
|
|
334
|
+
|
|
335
|
+
### index.delete_index
|
|
336
|
+
|
|
337
|
+
Deletes the data for the given index from its backing Pagefind service.
|
|
338
|
+
Doesn't affect any written files or data returned by `get_files`.
|
|
339
|
+
|
|
340
|
+
``` ruby
|
|
341
|
+
index.delete_index
|
|
342
|
+
```
|
|
343
|
+
|
|
344
|
+
Calling `index.get_files` or `index.write_files` doesn't consume the index, and further modifications can be made. In situations where many indexes are being created, the `delete_index` call helps clear out memory from a shared Pagefind binary service.
|
|
345
|
+
|
|
346
|
+
Reusing a `Pagefind::Index` object after calling `index.delete_index` will raise `Pagefind::ServiceError`.
|
|
347
|
+
|
|
348
|
+
Not calling this method is fine — these indexes will be cleaned up when `Pagefind::Index.open`'s block exits, its backing Pagefind service closes, or your Ruby process exits.
|
|
349
|
+
|
|
350
|
+
### Pagefind::Service
|
|
351
|
+
|
|
352
|
+
`Pagefind::Service` manages a Pagefind service running in a subprocess.
|
|
353
|
+
|
|
354
|
+
When `Pagefind::Service.open` is given a block, the backing service starts, is yielded to the block, and shuts down when the block exits. Without a block, it returns the running service, which you should shut down with `close`.
|
|
355
|
+
|
|
356
|
+
``` ruby
|
|
357
|
+
service = Pagefind::Service.open
|
|
358
|
+
# ...
|
|
359
|
+
service.close
|
|
360
|
+
|
|
361
|
+
Pagefind::Service.open do |service| # the service launches
|
|
362
|
+
# ...
|
|
363
|
+
end
|
|
364
|
+
# the service closes
|
|
365
|
+
```
|
|
366
|
+
|
|
367
|
+
You should invoke `Pagefind::Service` directly when you want to use the same backing service for many indexes. `service.create_index` takes the same options as `Pagefind::Index.open`.
|
|
368
|
+
|
|
369
|
+
Indexes created this way are not written to disk automatically, so call `write_files` yourself:
|
|
370
|
+
|
|
371
|
+
``` ruby
|
|
372
|
+
Pagefind::Service.open do |service|
|
|
373
|
+
default_index = service.create_index
|
|
374
|
+
other_index = service.create_index(output_path: "./search/nonstandard")
|
|
375
|
+
|
|
376
|
+
default_index.add_directory("./a")
|
|
377
|
+
other_index.add_directory("./b")
|
|
378
|
+
|
|
379
|
+
default_index.write_files
|
|
380
|
+
other_index.write_files
|
|
381
|
+
end
|
|
382
|
+
```
|
|
383
|
+
|
|
384
|
+
|
|
102
385
|
## Troubleshooting
|
|
103
386
|
|
|
104
387
|
### `ERROR: Cannot find the pagefind executable` for supported platform
|
|
@@ -0,0 +1,162 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
|
|
3
|
+
module Pagefind
|
|
4
|
+
# Manages a Pagefind index.
|
|
5
|
+
#
|
|
6
|
+
# Index.open yields the index to a block.
|
|
7
|
+
# Entering the block starts a backing Pagefind service and creates an in-memory index in the backing service.
|
|
8
|
+
# Exiting the block writes the in-memory index to disk and then shuts down the backing Pagefind service.
|
|
9
|
+
#
|
|
10
|
+
# Each method of Index that talks to the backing Pagefind service can raise ServiceError.
|
|
11
|
+
# If an exception is raised inside Index.open's block, the block exits without writing the index files to disk.
|
|
12
|
+
#
|
|
13
|
+
# Index optionally takes configuration options that can apply parts of the [Pagefind CLI config](https://pagefind.app/docs/config-options/). The options available at this level are listed in Index.new.
|
|
14
|
+
#
|
|
15
|
+
# See the relevant documentation for these configuration options in the
|
|
16
|
+
# [Configuring the Pagefind CLI](https://pagefind.app/docs/config-options/) documentation.
|
|
17
|
+
class Index
|
|
18
|
+
# Yields a new index to the block and returns the block's value. Takes the same options as
|
|
19
|
+
# Index.new.
|
|
20
|
+
#
|
|
21
|
+
# You should invoke Service directly when you want to use the same backing service for
|
|
22
|
+
# many indexes. See Service#create_index.
|
|
23
|
+
#
|
|
24
|
+
#: [T] (?root_selector: String?, ?exclude_selectors: Array[String]?, ?force_language: String?, ?verbose: bool?, ?logfile: String?, ?keep_index_url: bool?, ?write_playground: bool?, ?include_characters: String?, ?output_path: (String | Pathname)?) { (Index) -> T } -> T
|
|
25
|
+
def self.open(**config)
|
|
26
|
+
raise ArgumentError, "#{name}.open requires a block" unless block_given?
|
|
27
|
+
|
|
28
|
+
Service.open do |service|
|
|
29
|
+
index = new(service, **config)
|
|
30
|
+
result = yield index
|
|
31
|
+
index.write_files unless index.instance_variable_get(:@deleted)
|
|
32
|
+
result
|
|
33
|
+
end
|
|
34
|
+
end
|
|
35
|
+
|
|
36
|
+
# Creates an in-memory index in the backing Pagefind service.
|
|
37
|
+
#
|
|
38
|
+
# @rbs service: Service -- The backing Pagefind service to create the index in.
|
|
39
|
+
# @rbs root_selector: String? -- The root selector to use for the index.
|
|
40
|
+
# If not supplied, Pagefind will use the `<html>` tag.
|
|
41
|
+
# @rbs exclude_selectors: Array[String]? -- Extra element selectors that Pagefind should ignore when indexing.
|
|
42
|
+
# @rbs force_language: String? -- Ignores any detected languages and creates a single index for the entire site as the
|
|
43
|
+
# provided language. Expects an ISO 639-1 code, such as `en` or `pt`.
|
|
44
|
+
# @rbs verbose: bool? -- Prints extra logging while indexing the site. Only affects the CLI, does not impact
|
|
45
|
+
# web-facing search.
|
|
46
|
+
# @rbs logfile: String? -- A path to a file to log indexing output to in addition to stdout.
|
|
47
|
+
# The file will be created if it doesn't exist and overwritten on each run.
|
|
48
|
+
# @rbs keep_index_url: bool? -- Whether to keep `index.html` at the end of search result paths.
|
|
49
|
+
# By default, a file at `animals/cat/index.html` will be given the URL
|
|
50
|
+
# `/animals/cat/`. Setting this option to `true` will result in the URL
|
|
51
|
+
# `/animals/cat/index.html`.
|
|
52
|
+
# @rbs write_playground: bool? -- When writing or outputting files, also write the Pagefind playground to /pagefind/playground/.
|
|
53
|
+
# Defaults to false, ensuring the playground isn't available on a live site.
|
|
54
|
+
# @rbs include_characters: String? -- Include these characters when indexing and searching words.
|
|
55
|
+
# Useful for sites documenting technical topics such as programming languages.
|
|
56
|
+
# @rbs output_path: (String | Pathname)? -- The folder to output the search bundle into, relative to the working directory.
|
|
57
|
+
# Defaults to `pagefind`.
|
|
58
|
+
# @rbs return: void
|
|
59
|
+
def initialize(
|
|
60
|
+
service,
|
|
61
|
+
root_selector: nil,
|
|
62
|
+
exclude_selectors: nil,
|
|
63
|
+
force_language: nil,
|
|
64
|
+
verbose: nil,
|
|
65
|
+
logfile: nil,
|
|
66
|
+
keep_index_url: nil,
|
|
67
|
+
write_playground: nil,
|
|
68
|
+
include_characters: nil,
|
|
69
|
+
output_path: nil
|
|
70
|
+
)
|
|
71
|
+
@service = service
|
|
72
|
+
@output_path = output_path
|
|
73
|
+
|
|
74
|
+
config = {
|
|
75
|
+
root_selector:,
|
|
76
|
+
exclude_selectors:,
|
|
77
|
+
force_language:,
|
|
78
|
+
verbose:,
|
|
79
|
+
logfile:,
|
|
80
|
+
keep_index_url:,
|
|
81
|
+
write_playground:,
|
|
82
|
+
include_characters:
|
|
83
|
+
}
|
|
84
|
+
result = @service.send_message({type: "NewIndex", config:})
|
|
85
|
+
@index_id = result[:index_id]
|
|
86
|
+
end
|
|
87
|
+
|
|
88
|
+
# Adds an HTML file to the index.
|
|
89
|
+
#
|
|
90
|
+
# @rbs content: String -- The source HTML content of the file to be parsed.
|
|
91
|
+
# @rbs source_path: String? -- The source path the HTML file would have on disk.
|
|
92
|
+
# Must be a relative path, or an absolute path within the current working directory.
|
|
93
|
+
# Pagefind will compute the result URL from this path.
|
|
94
|
+
# @rbs url: String? -- An explicit URL to use, instead of having Pagefind compute the
|
|
95
|
+
# URL based on the source_path. If not supplied, source_path must be supplied.
|
|
96
|
+
# @rbs return: { type: "IndexedFile", page_word_count: Integer, page_url: String, page_meta: Hash[Symbol, String] }
|
|
97
|
+
def add_html_file(content:, source_path: nil, url: nil)
|
|
98
|
+
@service.send_message({type: "AddFile", index_id: @index_id, file_path: source_path, url:, file_contents: content})
|
|
99
|
+
end
|
|
100
|
+
|
|
101
|
+
# Adds a direct record to the Pagefind index.
|
|
102
|
+
#
|
|
103
|
+
# This method is useful for adding non-HTML content to the search results.
|
|
104
|
+
#
|
|
105
|
+
# @rbs content: String -- The raw content of this record.
|
|
106
|
+
# @rbs url: String -- The output URL of this record. Pagefind will not alter this.
|
|
107
|
+
# @rbs language: String -- ISO 639-1 code of the language this record is written in.
|
|
108
|
+
# @rbs meta: Hash[String | Symbol, String]? -- The metadata to attach to this record. Supplying a `title` is highly recommended.
|
|
109
|
+
# @rbs filters: Hash[String | Symbol, Array[String]]? -- The filters to attach to this record. Filters are used to group records together.
|
|
110
|
+
# @rbs sort: Hash[String | Symbol, String]? -- The sort keys to attach to this record.
|
|
111
|
+
# @rbs return: { type: "IndexedFile", page_word_count: Integer, page_url: String, page_meta: Hash[Symbol, String] }
|
|
112
|
+
def add_custom_record(url:, content:, language:, meta: nil, filters: nil, sort: nil)
|
|
113
|
+
@service.send_message({type: "AddRecord", index_id: @index_id, url:, content:, language:, meta:, filters:, sort:})
|
|
114
|
+
end
|
|
115
|
+
|
|
116
|
+
# Indexes a directory from disk using the standard Pagefind indexing behaviour.
|
|
117
|
+
#
|
|
118
|
+
# This is equivalent to running the Pagefind binary with `--site <dir>`.
|
|
119
|
+
#
|
|
120
|
+
# @rbs path: String | Pathname -- The path to the directory to index. If the `path` provided is relative,
|
|
121
|
+
# it will be relative to the current working directory of your Ruby process.
|
|
122
|
+
# @rbs glob: String? -- A glob pattern to filter files in the directory. If not provided, all
|
|
123
|
+
# files matching `**/*.{html}` are indexed. For more information on glob patterns,
|
|
124
|
+
# see the [Wax patterns documentation](https://github.com/olson-sean-k/wax#patterns).
|
|
125
|
+
# @rbs return: { type: "IndexedDir", page_count: Integer }
|
|
126
|
+
def add_directory(path, glob: nil)
|
|
127
|
+
@service.send_message({type: "AddDir", index_id: @index_id, path: path.to_s, glob:})
|
|
128
|
+
end
|
|
129
|
+
|
|
130
|
+
# Writes the index files to disk.
|
|
131
|
+
#
|
|
132
|
+
# If you're using Index.open, there's no need to call this method:
|
|
133
|
+
# if no error occurred, exiting the block automatically writes the index files to disk.
|
|
134
|
+
#
|
|
135
|
+
# @rbs output_path: (String | Pathname)? -- A path to override the configured output path for the index.
|
|
136
|
+
# @rbs return: { type: "WriteFiles", output_path: String }
|
|
137
|
+
def write_files(output_path = @output_path)
|
|
138
|
+
@service.send_message({type: "WriteFiles", index_id: @index_id, output_path: output_path&.to_s})
|
|
139
|
+
end
|
|
140
|
+
|
|
141
|
+
# Gets raw data of all files in the Pagefind index.
|
|
142
|
+
#
|
|
143
|
+
# This method emits all files, which can be a lot of data.
|
|
144
|
+
#
|
|
145
|
+
# @rbs return: Array[{ path: String, content: String }] -- Each file's `:content` is a binary string.
|
|
146
|
+
def get_files
|
|
147
|
+
result = @service.send_message({type: "GetFiles", index_id: @index_id})
|
|
148
|
+
result[:files].map do |file|
|
|
149
|
+
{path: file[:path], content: file[:content].unpack1("m0")}
|
|
150
|
+
end
|
|
151
|
+
end
|
|
152
|
+
|
|
153
|
+
# Deletes the data for the given index from its backing Pagefind service.
|
|
154
|
+
# Doesn't affect any written files or data returned by #get_files.
|
|
155
|
+
#
|
|
156
|
+
#: () -> void
|
|
157
|
+
def delete_index
|
|
158
|
+
@service.send_message({type: "DeleteIndex", index_id: @index_id})
|
|
159
|
+
@deleted = true
|
|
160
|
+
end
|
|
161
|
+
end
|
|
162
|
+
end
|
|
@@ -0,0 +1,111 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
|
|
3
|
+
require "json"
|
|
4
|
+
require "open3"
|
|
5
|
+
|
|
6
|
+
module Pagefind
|
|
7
|
+
# Raised when the backing Pagefind service reports an error, or when communicating with it fails.
|
|
8
|
+
class ServiceError < StandardError
|
|
9
|
+
# The original message that Pagefind failed to parse, if any.
|
|
10
|
+
attr_reader :original_message #: String?
|
|
11
|
+
|
|
12
|
+
#: (?String? message, ?original_message: String?) -> void
|
|
13
|
+
def initialize(message = nil, original_message: nil)
|
|
14
|
+
super(message)
|
|
15
|
+
@original_message = original_message
|
|
16
|
+
end
|
|
17
|
+
end
|
|
18
|
+
|
|
19
|
+
# Manages a backing Pagefind service, running as a `pagefind --service` process.
|
|
20
|
+
#
|
|
21
|
+
# Pagefind handles one request at a time, so threads sharing a service take turns.
|
|
22
|
+
class Service
|
|
23
|
+
# Launches a backing Pagefind service. If a block is given, yields the service and shuts it down
|
|
24
|
+
# when the block exits, returning the block's value.
|
|
25
|
+
#
|
|
26
|
+
#: () -> Service
|
|
27
|
+
#: [T] () { (Service) -> T } -> T
|
|
28
|
+
def self.open
|
|
29
|
+
service = new
|
|
30
|
+
return service unless block_given?
|
|
31
|
+
|
|
32
|
+
begin
|
|
33
|
+
yield service
|
|
34
|
+
ensure
|
|
35
|
+
service.close
|
|
36
|
+
end
|
|
37
|
+
end
|
|
38
|
+
|
|
39
|
+
# Launches a backing Pagefind service. Call #close to shut it down.
|
|
40
|
+
#
|
|
41
|
+
#: () -> void
|
|
42
|
+
def initialize
|
|
43
|
+
@message_id = 0
|
|
44
|
+
@mutex = Thread::Mutex.new
|
|
45
|
+
@stdin, @stdout, @wait_thread = Open3.popen2(Pagefind.executable, "--service")
|
|
46
|
+
@stdin.binmode
|
|
47
|
+
@stdout.binmode
|
|
48
|
+
end
|
|
49
|
+
|
|
50
|
+
# Creates an in-memory index in this service. Takes the same options as Index.new.
|
|
51
|
+
#
|
|
52
|
+
#: (?root_selector: String?, ?exclude_selectors: Array[String]?, ?force_language: String?, ?verbose: bool?, ?logfile: String?, ?keep_index_url: bool?, ?write_playground: bool?, ?include_characters: String?, ?output_path: (String | Pathname)?) -> Index
|
|
53
|
+
def create_index(**) = Index.new(self, **)
|
|
54
|
+
|
|
55
|
+
# Sends a request to the backing Pagefind service and returns its response.
|
|
56
|
+
#
|
|
57
|
+
# @rbs payload: Hash[Symbol, untyped] -- The request payload to send, such as `{type: "GetFiles", index_id: 0}`.
|
|
58
|
+
# @rbs return: Hash[Symbol, untyped]
|
|
59
|
+
def send_message(payload)
|
|
60
|
+
response = nil
|
|
61
|
+
@mutex.synchronize do
|
|
62
|
+
raise ServiceError, "Pagefind service is not running" if @stdin.closed?
|
|
63
|
+
|
|
64
|
+
begin
|
|
65
|
+
write_message(message_id: (@message_id += 1), payload:)
|
|
66
|
+
response = read_message
|
|
67
|
+
ensure
|
|
68
|
+
# the response wasn't read, so shut down before the next request reads it by mistake
|
|
69
|
+
@stdin.close unless response
|
|
70
|
+
end
|
|
71
|
+
end
|
|
72
|
+
result = response[:payload]
|
|
73
|
+
|
|
74
|
+
if response[:message_id].nil?
|
|
75
|
+
raise ServiceError.new("Pagefind service error when parsing a message: #{result[:message]}", original_message: result[:original_message])
|
|
76
|
+
elsif result[:type] == "Error"
|
|
77
|
+
raise ServiceError.new(result[:message], original_message: result[:original_message])
|
|
78
|
+
end
|
|
79
|
+
|
|
80
|
+
result
|
|
81
|
+
end
|
|
82
|
+
|
|
83
|
+
# Waits for any request in progress to finish, then shuts down the backing Pagefind service.
|
|
84
|
+
#
|
|
85
|
+
#: () -> nil
|
|
86
|
+
def close
|
|
87
|
+
@mutex.synchronize do
|
|
88
|
+
@stdin.close # pagefind exits when its stdin reaches EOF
|
|
89
|
+
@wait_thread.join
|
|
90
|
+
@stdout.close
|
|
91
|
+
end
|
|
92
|
+
end
|
|
93
|
+
|
|
94
|
+
private
|
|
95
|
+
|
|
96
|
+
#: (Hash[Symbol, untyped] request) -> void
|
|
97
|
+
def write_message(request)
|
|
98
|
+
@stdin.write([JSON.generate(request)].pack("m0"), ",")
|
|
99
|
+
rescue IOError, SystemCallError => e
|
|
100
|
+
raise ServiceError, "Failed to send a message to the Pagefind service: #{e.message}"
|
|
101
|
+
end
|
|
102
|
+
|
|
103
|
+
#: () -> Hash[Symbol, untyped]
|
|
104
|
+
def read_message
|
|
105
|
+
output = @stdout.gets(",", chomp: true)
|
|
106
|
+
raise ServiceError, "Pagefind service exited" if output.nil?
|
|
107
|
+
|
|
108
|
+
JSON.parse(output.unpack1("m0"), symbolize_names: true)
|
|
109
|
+
end
|
|
110
|
+
end
|
|
111
|
+
end
|
data/lib/pagefind/version.rb
CHANGED
data/lib/pagefind.rb
CHANGED
|
@@ -2,6 +2,8 @@
|
|
|
2
2
|
|
|
3
3
|
require_relative "pagefind/version"
|
|
4
4
|
require_relative "pagefind/upstream"
|
|
5
|
+
require_relative "pagefind/service"
|
|
6
|
+
require_relative "pagefind/index"
|
|
5
7
|
|
|
6
8
|
module Pagefind
|
|
7
9
|
DEFAULT_DIR = File.expand_path(File.join(__dir__, "..", "exe"))
|
|
@@ -25,6 +27,9 @@ module Pagefind
|
|
|
25
27
|
end
|
|
26
28
|
|
|
27
29
|
def executable(exe_path: DEFAULT_DIR)
|
|
30
|
+
binary_path = ENV["PAGEFIND_EXTENDED_BINARY_PATH"] || ENV["PAGEFIND_BINARY_PATH"]
|
|
31
|
+
return binary_path if binary_path
|
|
32
|
+
|
|
28
33
|
pagefind_install_dir = ENV["PAGEFIND_INSTALL_DIR"]
|
|
29
34
|
if pagefind_install_dir
|
|
30
35
|
if File.directory?(pagefind_install_dir)
|
data/rakelib/package.rake
CHANGED
|
@@ -55,6 +55,7 @@
|
|
|
55
55
|
# - rake gem # Build all the gem files
|
|
56
56
|
# - rake package # Build all the gem files (same as `gem`)
|
|
57
57
|
# - rake repackage # Force a rebuild of all the gem files
|
|
58
|
+
# - rake test # Run the tests (downloads the current platform's binary first)
|
|
58
59
|
#
|
|
59
60
|
# Note also that the binary executables will be lazily downloaded when needed, but you can
|
|
60
61
|
# explicitly download them with the `rake download` command.
|
|
@@ -156,5 +157,10 @@ end
|
|
|
156
157
|
desc "Download all pagefind binaries"
|
|
157
158
|
task "download" => [:check, *exepaths]
|
|
158
159
|
|
|
160
|
+
local_exepath = exepaths.find do |exepath|
|
|
161
|
+
Gem::Platform.match_gem?(Gem::Platform.new(File.basename(File.dirname(exepath))), PAGEFIND_RUBY_GEMSPEC.name)
|
|
162
|
+
end
|
|
163
|
+
task test: local_exepath if local_exepath
|
|
164
|
+
|
|
159
165
|
CLOBBER.add(exepaths.map { |p| File.dirname(p) })
|
|
160
166
|
CLOBBER.add(archivepaths)
|
metadata
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
--- !ruby/object:Gem::Specification
|
|
2
2
|
name: pagefind
|
|
3
3
|
version: !ruby/object:Gem::Version
|
|
4
|
-
version: 1.5.2.
|
|
4
|
+
version: 1.5.2.2
|
|
5
5
|
platform: ruby
|
|
6
6
|
authors:
|
|
7
7
|
- 방성범 (Bang Seongbeom)
|
|
@@ -9,8 +9,8 @@ bindir: exe
|
|
|
9
9
|
cert_chain: []
|
|
10
10
|
date: 1980-01-02 00:00:00.000000000 Z
|
|
11
11
|
dependencies: []
|
|
12
|
-
description: A self-contained `pagefind` executable,
|
|
13
|
-
it. Nothing else.
|
|
12
|
+
description: A self-contained `pagefind` executable, with an indexing API, wrapped
|
|
13
|
+
up in a ruby gem. That's it. Nothing else.
|
|
14
14
|
email:
|
|
15
15
|
- bangseongbeom@gmail.com
|
|
16
16
|
executables:
|
|
@@ -27,6 +27,8 @@ files:
|
|
|
27
27
|
- Rakefile
|
|
28
28
|
- exe/pagefind
|
|
29
29
|
- lib/pagefind.rb
|
|
30
|
+
- lib/pagefind/index.rb
|
|
31
|
+
- lib/pagefind/service.rb
|
|
30
32
|
- lib/pagefind/upstream.rb
|
|
31
33
|
- lib/pagefind/version.rb
|
|
32
34
|
- rakelib/package.rake
|
|
@@ -55,5 +57,5 @@ required_rubygems_version: !ruby/object:Gem::Requirement
|
|
|
55
57
|
requirements: []
|
|
56
58
|
rubygems_version: 4.0.20
|
|
57
59
|
specification_version: 4
|
|
58
|
-
summary: A self-contained `pagefind` executable.
|
|
60
|
+
summary: A self-contained `pagefind` executable, with an indexing API.
|
|
59
61
|
test_files: []
|