tablefacts 0.2.0 → 0.3.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/.env.example CHANGED
@@ -2,7 +2,7 @@
2
2
  # Bypasses row level security: keep it here, never in a browser-exposed variable.
3
3
  SUPABASE_DB_URL=SUPABASE_DB_POOLER_URL
4
4
 
5
- # Menus that are only pictures (tablefacts menu raw) are read by a vision model. Fill in
5
+ # Menus that are only pictures or a PDF (tablefacts menu raw) are read by a vision model. Fill in
6
6
  # the key of the provider you use: Claude, Gemini or Groq
7
7
  ANTHROPIC_API_KEY=
8
8
  GEMINI_API_KEY=
package/CHANGELOG.md CHANGED
@@ -1,5 +1,27 @@
1
1
  # Changelog
2
2
 
3
+ ## 0.3.0
4
+
5
+ ### Menu (`tablefacts menu raw`)
6
+
7
+ - The source now reads **PDF menus** as well as pictures. An argument ending in `.pdf` is a local file
8
+ (resolved against the project) or an http(s) URL. A page that carries its own text is transcribed from that
9
+ text (a two-column layout is read column by column); a page with no text layer (a scan) is rendered and read
10
+ like a picture. `--list` shows which each page is. New optional dependencies: `pdfjs-dist` for reading,
11
+ `@napi-rs/canvas` for rendering (loaded lazily, with an `EDEPENDENCY` error naming the install command when
12
+ missing).
13
+ - New options `--images <folder>` and `--image-base-url <url>` (`imageDir`, `imageBaseUrl` on
14
+ `importImageMenu`): the dish photos printed on a PDF page are screenshotted and saved. When the page places
15
+ each photo separately they are matched to the nearest printed dish name; when it does not (a flattened export,
16
+ vector art, a scan) the model is asked for each item's photo box instead, so photos can come from any page.
17
+ The new `--image-boxes auto|always` option controls this: `auto` (default) only reads a page as a picture when
18
+ it has no separately placed photo, while `always` reads every page as a picture so the model boxes each dish.
19
+ With a base URL each product's `image_url` is filled; without it the crops are saved and `image_url` stays
20
+ empty.
21
+ - `findPages`, `listMenuImages` and `normalizePages` now also know about PDF pages: `listMenuImages` reports
22
+ `kind` and `chars`, and `normalizePages` returns `placements` (which products each item built) so photos can
23
+ be attached.
24
+
3
25
  ## 0.2.0
4
26
 
5
27
  - The menu importers no longer write hardcoded `public.menu_*` tables. A restaurant's config (or the
package/README.md CHANGED
@@ -2,7 +2,7 @@
2
2
 
3
3
  Data extraction tools for restaurant sites, shared by every Cannario template. Give it a restaurant's name and
4
4
  place and it gathers the public facts; point it at an Instagram or TripAdvisor page and it downloads the photos;
5
- point it at a Cluvi menu (or pictures of a paper menu) and it loads the menu into Supabase.
5
+ point it at a Cluvi menu, a PDF menu, or pictures of a paper menu and it loads the menu into Supabase.
6
6
 
7
7
  Research needs no key: Google Places and OpenStreetMap run first, and a key-free web search (DuckDuckGo) fills in
8
8
  candidate links when they find no place or no website, so a bare name and place still produces a useful report.
@@ -16,7 +16,7 @@ Everything works two ways: as a **command line** (`tablefacts <tool>`) and as a
16
16
  | Instagram photos | `tablefacts photos instagram` | `downloadInstagram()` | [src/instagram](src/instagram/README.md) |
17
17
  | TripAdvisor photos | `tablefacts photos tripadvisor` | `downloadTripadvisor()` | [src/tripadvisor](src/tripadvisor/README.md) |
18
18
  | Menu from Cluvi | `tablefacts menu cluvi` | `importCluvi()` | [src/menu](src/menu/README.md) |
19
- | Menu from pictures | `tablefacts menu raw` | `importImageMenu()` | [src/menu](src/menu/README.md) |
19
+ | Menu from pictures or a PDF | `tablefacts menu raw` | `importImageMenu()` | [src/menu](src/menu/README.md) |
20
20
 
21
21
  Contents: [Requirements](#requirements) · [Install](#install) · [Quick start](#quick-start) ·
22
22
  [Command line](#command-line) · [Configuration](#configuration) · [Library](#use-it-as-a-library) ·
@@ -27,6 +27,9 @@ Contents: [Requirements](#requirements) · [Install](#install) · [Quick start](
27
27
 
28
28
  - **Node.js 24 or newer** (`engines` in `package.json`). The tools use the built-in `fetch` and `util.parseArgs`.
29
29
  - **Playwright** (optional) for the photo tools and `research --render`: `npm install --save-dev playwright`.
30
+ - **pdfjs-dist** and **@napi-rs/canvas** (optional) for PDF menus with `tablefacts menu raw`:
31
+ `npm install pdfjs-dist @napi-rs/canvas`. Only reading a PDF page's text needs `pdfjs-dist`; rendering a scanned
32
+ page or a product photo also needs the canvas package (it comes with `@napi-rs/canvas` on Windows, macOS and Linux).
30
33
  - **Microsoft Edge on Windows** for the photo tools. They attach to a normal Edge window because the sites
31
34
  guard their pages with bot checks that automated browsers fail (see the photo tool READMEs). Starting Edge
32
35
  for you only works with Edge installed in its usual `Program Files` folder; elsewhere, start it yourself with
@@ -41,7 +44,8 @@ npx tablefacts --help
41
44
  ```
42
45
 
43
46
  Playwright is an **optional** peer dependency. Without it the photo tools and `research --render` stop with an
44
- `EDEPENDENCY` error that tells you to install it; everything else works.
47
+ `EDEPENDENCY` error that tells you to install it; everything else works. `pdfjs-dist` (and `@napi-rs/canvas` for
48
+ scans and product photos) is likewise optional and only `menu raw` with a PDF needs it.
45
49
 
46
50
  ## Quick start
47
51
 
@@ -124,10 +128,19 @@ Both share these options (and replace the menu in Supabase unless `--dry-run`):
124
128
  | `--replace-all` | Replace the whole menu, not only the categories in this import. Requires `--yes` (except with `--dry-run`) |
125
129
  | `--yes` | Confirm `--replace-all`, which empties the target tables (`<prefix>menu_categories`, `<prefix>menu_sections`, `<prefix>menu_products`) |
126
130
  | `--force` | Write even if the import has fewer than half the products it replaces |
131
+ | `--images <folder>` | With a PDF, also save the dish photos printed on its pages |
132
+ | `--image-base-url <url>` | https folder those photos will be published at; fills each product's `image_url` |
133
+ | `--image-boxes <auto\|always>` | How photos are found. `auto` (default) reads the page as text and matches separately placed photos by position, using the model's boxes only when the page has none; `always` reads every page as a picture so the model boxes every dish's photo |
127
134
 
128
135
  `tablefacts menu cluvi [menu-url] [--service on_table|delivery|take_away] [--lang es]`
129
136
 
130
- `tablefacts menu raw [page-or-image-url...] [--list] [--only 1,3-5] [--provider anthropic|gemini|groq] [--model <id>] [--min-width 500] [--refresh]`
137
+ `tablefacts menu raw [page-or-image-url...] [menu.pdf...] [--list] [--only 1,3-5] [--provider anthropic|gemini|groq] [--model <id>] [--min-width 500] [--refresh] [--images <folder>] [--image-base-url <url>] [--image-boxes auto|always]`
138
+
139
+ For a PDF, an argument ending in `.pdf` is a local file (resolved against the project) or an http(s) URL. A page
140
+ with text is transcribed from that text (a two-column layout is read column by column); a scanned page is rendered
141
+ and read like a picture. `--images` also saves the printed product photos: a page that places each photo
142
+ separately is matched by dish name, and a page that does not (a flattened export, vector art, a scan) is read as a
143
+ picture so the model can box each dish — `--image-boxes always` does that on every page.
131
144
 
132
145
  Both read the restaurant-specific part (URL, category mapping, currency) from a `config.mjs` next to the tool.
133
146
  **The shipped configs hold another restaurant's values: edit them first.** See [src/menu/README.md](src/menu/README.md).
@@ -196,8 +209,8 @@ default output folder, browser profiles, the menu transcription cache, and where
196
209
  | `downloadTripadvisor({ links, out, cdp, edgeDir, max, dryRun, debug, projectDir, log })` | Downloads a restaurant page's photos. Same result | Playwright |
197
210
  | `importMenu({ menu, notes, title, dryRun, json, tablePrefix, allowUnprefixed, replaceAll, yes, force, databaseUrl, env, projectDir, log })` | Validates a menu and writes it to Supabase. Returns `{ totals, notes, written, dryRun, database }`, `database` being `{ label, tables, current: { categories, products, kept } }` or `null` when it was not reached | `SUPABASE_DB_URL` unless `dryRun` |
198
211
  | `importCluvi({ url, service, lang, config, ...importMenu options })` | Reads a Cluvi menu, then `importMenu` | `SUPABASE_DB_URL` unless `dryRun` |
199
- | `importImageMenu({ urls, only, provider, model, minWidth, refresh, apiKey, config, ...importMenu options })` | Transcribes menu pictures with a vision model, then `importMenu` | A vision key (see [Configuration](#configuration)) |
200
- | `listMenuImages({ urls, only, minWidth, config })` | Lists the menu pictures found on pages (numbered as `--list` shows) | none |
212
+ | `importImageMenu({ urls, only, provider, model, minWidth, refresh, imageDir, imageBaseUrl, imageBoxes, apiKey, config, ...importMenu options })` | Transcribes menu pictures or a PDF with a model, then `importMenu`. A PDF page uses its text when it has one, a render when it is a scan. `imageDir` also saves the printed product photos (`imageBaseUrl` fills their links; `imageBoxes` is `'auto'` or `'always'`) | A vision key (see [Configuration](#configuration)); a PDF needs `pdfjs-dist` (`@napi-rs/canvas` too for scans or photos) |
213
+ | `listMenuImages({ urls, only, minWidth, projectDir, config })` | Lists the pages found (pictures or PDF pages, numbered as `--list` shows) | none |
201
214
  | `loadEnv({ projectDir, files, env })` | Loads `.env` files into `env` (default `process.env`); returns the files it read | none |
202
215
 
203
216
  Also exported: `validateMenu`, `countMenu`, `parsePrice`, `normalizePages`, `providers`, `defaultProvider`,
@@ -233,7 +246,7 @@ Failures the library raises on purpose throw `TablefactsError`, which has a `cod
233
246
  | --- | --- |
234
247
  | `EUSAGE` | A missing or invalid argument (no `out`, no valid links, no `name`). CLIs exit with `2` |
235
248
  | `ECONFIG` | Missing configuration: database URL, API key, unknown provider, no menu URL |
236
- | `EDEPENDENCY` | An optional dependency is not installed (Playwright) |
249
+ | `EDEPENDENCY` | An optional dependency is not installed (Playwright, or `pdfjs-dist`/`@napi-rs/canvas` for a PDF menu) |
237
250
  | `EFAILED` | A source failed or was blocked, data did not validate, or an import was refused (invalid menu, no products, far fewer products than it replaces) |
238
251
 
239
252
  When the error is about one option, `err.option` names it (`'out'`, `'links'`, `'force'`...) and the message writes
@@ -263,6 +276,7 @@ folder tablefacts is installed in, so one install serves every template.
263
276
  | `.env` | `<project>/.env`, then `<project>/data/.env` (older templates) |
264
277
  | Research output | `<project>/.tablefacts/research/<slug>/`, with `latest.json` next to the folders |
265
278
  | Menu transcription cache and downloaded page images | `<project>/.tablefacts/cache/<host>/` |
279
+ | Product photos saved from a PDF | The `--images` folder, or `<project>/.tablefacts/menu-images` |
266
280
  | Fallback browser profile | `<project>/.tablefacts/instagram-profile` |
267
281
  | Photos | The `--out` folder you give |
268
282
  | Edge profile for the photo tools | `C:\ig-edge` (`--edge-dir` changes it) |
@@ -313,6 +327,12 @@ tables — is in [docs/MENU_TABLE_PREFIX.md](docs/MENU_TABLE_PREFIX.md).
313
327
  past bot protection: a blocked source is reported, not worked around.
314
328
  - Cluvi has no public API; the importer calls the two JSON endpoints its web app uses, which can change.
315
329
  - Prices read from pictures can be wrong. Read the dry run against the original before writing.
330
+ - A PDF page with text is read as text; a scan is rendered and read as a picture, so small print can still be
331
+ misread. A two-column page is read column by column when the layout is clear, but an unusual layout is
332
+ flattened into single lines, so check the dry run. Product photos are screenshots cropped from the page: a
333
+ separately placed photo is matched to the nearest dish, and when the page has none (a flattened export, or
334
+ vector art) the model is asked for each dish's photo box instead — pass `--image-boxes always` to do that on
335
+ every page too. Boxes are approximate, so check the dry run.
316
336
  - Only download photos the restaurant owns or has allowed you to use. Guest photos on TripAdvisor and
317
337
  Instagram belong to their authors.
318
338
 
@@ -330,9 +350,9 @@ and its declarations fails the tests. Document options with JSDoc in `src/lib/ty
330
350
  CI (`.github/workflows/ci.yml`) runs the tests on Ubuntu and Windows with Node 24 and checks `npm pack --dry-run`
331
351
  on every push to `main` and every pull request. Dependabot (`.github/dependabot.yml`) keeps dependencies current.
332
352
 
333
- **Release:** bump `version` in `package.json`, commit, then push a matching tag, e.g. `git tag v0.1.1 && git push origin v0.1.1`.
334
- `.github/workflows/release.yml` checks the tag equals the version, runs the tests, publishes to npm with provenance
335
- (needs an `NPM_TOKEN` repository secret) and creates a GitHub release with generated notes.
353
+ **Release:** publishing is manual. Bump `version` in `package.json`, add the entry to `CHANGELOG.md`, commit, then
354
+ run `npm publish` from a clean checkout of that commit. Publishing is not automated in CI; `prepublishOnly` builds
355
+ `types/` and runs the tests first. Optionally tag the release commit (`git tag v0.1.1 && git push origin v0.1.1`).
336
356
 
337
357
  For how the code is organised and how to add a source or a tool, see [docs/ARCHITECTURE.md](docs/ARCHITECTURE.md).
338
358
 
package/package.json CHANGED
@@ -1,10 +1,11 @@
1
1
  {
2
2
  "name": "tablefacts",
3
- "version": "0.2.0",
4
- "description": "Data extraction tools for restaurant sites: public-facts research, photos (Instagram, TripAdvisor) and menu import",
3
+ "version": "0.3.0",
4
+ "description": "Data extraction tools for restaurant sites: public-facts research, photos (Instagram, TripAdvisor) and menu import from Cluvi, pictures or PDFs",
5
5
  "keywords": [
6
6
  "restaurant",
7
7
  "menu",
8
+ "pdf",
8
9
  "scraper",
9
10
  "tripadvisor",
10
11
  "instagram",
@@ -61,15 +62,25 @@
61
62
  "pg": "^8.23.1"
62
63
  },
63
64
  "devDependencies": {
65
+ "@napi-rs/canvas": "^1.0.10",
64
66
  "@types/node": "^26.6.4",
67
+ "pdfjs-dist": "~6.2.108",
65
68
  "playwright": "^1.63.0",
66
69
  "typescript": "^7.0.2",
67
70
  "vitest": "^5.0.3"
68
71
  },
69
72
  "peerDependencies": {
73
+ "@napi-rs/canvas": "^1.0.10",
74
+ "pdfjs-dist": "~6.2.108",
70
75
  "playwright": "^1.63.0"
71
76
  },
72
77
  "peerDependenciesMeta": {
78
+ "@napi-rs/canvas": {
79
+ "optional": true
80
+ },
81
+ "pdfjs-dist": {
82
+ "optional": true
83
+ },
73
84
  "playwright": {
74
85
  "optional": true
75
86
  }
package/src/lib/types.mjs CHANGED
@@ -179,6 +179,7 @@
179
179
  * @property {string} [thousands] Thousands separator the menu prints. Default ".".
180
180
  * @property {string} [decimal] Decimal separator the menu prints. Default ",".
181
181
  * @property {number} [scale] Multiplies every price. Default 1.
182
+ * @property {number} [imageScale] Resolution a PDF page is rendered at before its product photos are screenshot. Default 2.
182
183
  * @property {{ slug: string, name: string, groups?: ('food' | 'drink')[] }[]} [categories] Category that lists each group.
183
184
  * @property {Record<string, string>} [placeIn] Section title to category slug.
184
185
  * @property {Record<string, string>} [sections] Section title to the name stored.
@@ -199,13 +200,16 @@
199
200
 
200
201
  /**
201
202
  * @typedef {object} ImageMenuReadOptions
202
- * @property {string[]} [urls] Pages or image URLs. Default: the config's url.
203
+ * @property {string[]} [urls] Pages or image URLs, or PDF file paths/URLs. Default: the config's url.
203
204
  * @property {string | number[]} [only] Pages to read, e.g. '1,3-5' or [1, 3, 4, 5].
204
205
  * @property {string} [provider] Vision provider, see `providers`. Default MENU_VISION_PROVIDER or `defaultProvider`.
205
206
  * @property {string} [model]
206
207
  * @property {number} [minWidth] Ignore images declaring a smaller width. Default 500.
207
208
  * @property {string} [apiKey] Default: the provider's key in env.
208
209
  * @property {boolean} [refresh] Read the pages again instead of using the saved transcriptions.
210
+ * @property {string} [imageDir] Also save the dish photos printed on a PDF page here (resolved against projectDir). Enables photo extraction.
211
+ * @property {string} [imageBaseUrl] https folder the saved photos will be published at; fills each product's image_url with it plus the file name.
212
+ * @property {'auto' | 'always'} [imageBoxes] How photos are found: 'auto' (default) reads a PDF page from its text and matches placed photos by position, using the model's boxes only when the page has none; 'always' reads every PDF page as a picture so the model boxes every dish's photo. Default 'auto'.
209
213
  * @property {RawConfig} [config] Default: the raw config.mjs.
210
214
  */
211
215
 
@@ -216,14 +220,17 @@
216
220
  * @property {string[]} [urls]
217
221
  * @property {string | number[]} [only]
218
222
  * @property {number} [minWidth]
223
+ * @property {string} [projectDir] Project folder. Default: TABLEFACTS_PROJECT or the current folder.
219
224
  * @property {RawConfig} [config]
220
225
  */
221
226
 
222
227
  /**
223
228
  * @typedef {object} MenuImage
224
229
  * @property {number} number Position as the CLI's --list numbers it.
225
- * @property {string} url
230
+ * @property {string} url The image URL, or "<pdf path or URL>#<page number>" for a PDF page.
226
231
  * @property {string} alt
232
+ * @property {'image' | 'pdf'} [kind] Where the page came from. Default: 'image'.
233
+ * @property {number} [chars] Characters of text a PDF page has (0 means it is a scan).
227
234
  */
228
235
 
229
236
  /**
@@ -6,6 +6,7 @@ Scripts that read a restaurant's menu from the website that hosts it and write i
6
6
  | --- | --- | --- |
7
7
  | Cluvi (`<restaurant>.cluvi.co`) | `cluvi/` | working |
8
8
  | Menu that is only pictures (one image per page) | `raw/` | working, API calls untested |
9
+ | Menu as a PDF (text layer or scan) | `raw/pdf.mjs` | working, optional `pdfjs-dist` |
9
10
 
10
11
  The same code is available as functions: `importCluvi`, `importImageMenu`, `listMenuImages` and `importMenu` (see the [main README](../../README.md#use-it-as-a-library)). Each source's `config.mjs` can be replaced with the `config` option, so a script can import a menu without editing the package:
11
12
 
@@ -23,10 +24,11 @@ const pictures = await listMenuImages({ urls: ['https://example.com/carta'], con
23
24
  ```
24
25
 
25
26
  - **Cluvi `config`**: `{ tablePrefix?, url?, categories: [{ slug, name, from[] }], sections: { 'Cluvi subcategory': 'SECTION NAME' } }`.
26
- - **Picture `config`**: `{ tablePrefix?, url?, currency, thousands, decimal, scale, categories: [{ slug, name, groups: ['food' | 'drink'] }], placeIn, sections, skipSections }`. `currency` is required.
27
+ - **Picture `config`**: `{ tablePrefix?, url?, currency, thousands, decimal, scale, imageScale?, categories: [{ slug, name, groups: ['food' | 'drink'] }], placeIn, sections, skipSections }`. `currency` is required; `imageScale` is the PDF render resolution for product photos (default 2).
28
+ - **Picture/PDF read options** (`importImageMenu`): `provider`, `model`, `minWidth`, `refresh`, `apiKey`, and the photo options `imageDir` (save the printed product photos), `imageBaseUrl` (fill their `image_url`) and `imageBoxes` (`'auto'` default, or `'always'`).
27
29
  - **Options every import takes**: `dryRun`, `json` (relative paths resolve against `projectDir`), `tablePrefix` (overrides the config's; the restaurant's table set on a shared database), `allowUnprefixed` (override the shared-database check and write the unprefixed `menu_*` tables, only when this restaurant owns them), `replaceAll`, `yes` (confirms a whole-menu `replaceAll`), `force`, `databaseUrl` (default `env.SUPABASE_DB_URL`), `env`, `projectDir` and `log(message, level)` (`'info'`, `'warn'` for notes, `'error'`).
28
30
  - **Result**: `{ totals, notes, written, dryRun, database }`; `database` is `{ label, tables, current: { categories, products, kept } }` once the database was inspected, and `null` when it was not reached (a dry run without a URL).
29
- - **Errors** are `TablefactsError`: `ECONFIG` (no URL, unknown provider, no key, bad currency, a bad `tablePrefix`, another restaurant's tables on an unprefixed import), `EFAILED` (invalid menu, no products, import refused without `force`), `EUSAGE` (bad `only`, `replaceAll` without `yes`). `err.option` names the option; the CLI shows the flag (`--force`, `--only`) and exits `2` for `EUSAGE`, `1` otherwise.
31
+ - **Errors** are `TablefactsError`: `ECONFIG` (no URL, unknown provider, no key, bad currency, a bad `tablePrefix`, another restaurant's tables on an unprefixed import), `EDEPENDENCY` (`pdfjs-dist`/`@napi-rs/canvas` missing for a PDF), `EFAILED` (invalid menu, no products, import refused without `force`), `EUSAGE` (bad `only`, a bad `imageBoxes`, `replaceAll` without `yes`). `err.option` names the option; the CLI shows the flag (`--force`, `--only`) and exits `2` for `EUSAGE`, `1` otherwise.
30
32
 
31
33
  ## Setup
32
34
 
@@ -71,9 +73,9 @@ tablefacts menu cluvi --help
71
73
 
72
74
  Edit `cluvi/config.mjs` first: the menu URL, how Cluvi's main categories fold into the site's categories, and section renames. The run prints what the site still needs (dictionary keys, `content/qr.ts`).
73
75
 
74
- ## Run (menu that is only pictures)
76
+ ## Run (menu that is only pictures or a PDF)
75
77
 
76
- For a restaurant whose site shows the menu as a gallery of page images, such as <https://www.mombasa.co/carta-restaurante-espanol/>. Each page is transcribed by a vision model from one of three providers, so it needs that provider's key in `.env`:
78
+ For a restaurant whose site shows the menu as a gallery of page images, such as <https://www.mombasa.co/carta-restaurante-espanol/>, or that offers it as a PDF. Each page is transcribed by a model from one of three providers, so it needs that provider's key in `.env`:
77
79
 
78
80
  | `--provider` | Key in `.env` | Default `--model` |
79
81
  | --- | --- | --- |
@@ -92,12 +94,32 @@ tablefacts menu raw # replace this restaurant's menu i
92
94
  tablefacts menu raw <page-or-image-url>... # another restaurant
93
95
  tablefacts menu raw --table-prefix makibar_ # override the config's table prefix for one run
94
96
  tablefacts menu raw --replace-all --yes # empty this restaurant's tables first, then write
97
+ tablefacts menu raw carta.pdf --list # pages of a PDF file (a .pdf URL works too)
98
+ tablefacts menu raw carta.pdf --dry-run # text pages from their text, scans from a render
99
+ tablefacts menu raw carta.pdf --images ./carta-fotos --image-base-url https://cdn.example.com/carta/ --dry-run
100
+ tablefacts menu raw carta.pdf --images ./carta-fotos --image-boxes always --dry-run
95
101
  ```
96
102
 
97
- Edit `raw/config.mjs` first: the URL, the currency, how the menu prints prices (`$95.000` is `thousands: "."`, `decimal: ","`) and which category each section goes to.
103
+ A PDF argument ends in `.pdf` and is a local file or an http(s) URL. A page that carries its own text is
104
+ transcribed from that text (cheaper and exact); a page with no text layer is rendered and read like a picture
105
+ (`chars` in `--list` tells them apart). Rendering needs `@napi-rs/canvas`; text-only PDFs need only `pdfjs-dist`.
106
+
107
+ `--images <folder>` also saves the dish photos printed on a PDF page, as PNG crops of the page render. Two ways
108
+ the crop finds its dish: when the page places each photo separately, each is matched to the product whose printed
109
+ name is nearest it; when it does not (a flattened export, vector art, a scan), the model is asked for each dish's
110
+ photo box instead. That fallback reads the page as a picture, so asking for photos adds a vision call to pages
111
+ with no separately placed photo. `--image-boxes always` reads **every** PDF page as a picture so the model boxes
112
+ each dish even when the page also places some photos separately (one vision call per page); `auto` (the default)
113
+ only does that when the page has no separately placed photo. `--image-base-url https://host/path/` then fills
114
+ each product's `image_url` with that URL plus the file name, so the import can use it. Without it the photos are
115
+ saved but `image_url` stays empty until you host them and re-run with the base URL.
116
+
117
+ Edit `raw/config.mjs` first: the URL (a PDF path/URL works there too), the currency, how the menu prints prices (`$95.000` is `thousands: "."`, `decimal: ","`) and which category each section goes to.
98
118
 
99
119
  - `raw/source.mjs` takes the large JPG/PNG/WebP images of the page in document order (full size, not the thumbnails; logos and icons are skipped by width) and downloads them. A site that builds its gallery with JavaScript shows up as "no images": pass the image URLs instead.
100
- - `raw/vision.mjs` has the model transcribe each page into sections and dishes with one request: a forced tool call for Anthropic, JSON held to the same schema for Gemini (`generateContent`) and Groq (strict structured output), so every provider hands back the same shape. Prices come back as printed text and `raw/normalize.mjs` parses them, so a thousands separator is never guessed by the model. Adding a provider is one more entry in `providers` there.
120
+ - `raw/pdf.mjs` reads a PDF (a local file or URL) with the optional `pdfjs-dist`: a page's text and text positions, where each printed image sits on the page, and a render of the page. A page with enough text is sent to the model as text (a two-column layout is read column by column); a scan is rendered and sent as a picture. `raw/pdfjs.mjs` loads the dependency on use and turns a missing one into an `EDEPENDENCY` error with the install command.
121
+ - `raw/vision.mjs` has the model transcribe each page into sections and dishes with one request: a forced tool call for Anthropic, JSON held to the same schema for Gemini (`generateContent`) and Groq (strict structured output), so every provider hands back the same shape. The same providers read a PDF page's text (no picture) and can return each item's photo box. Prices come back as printed text and `raw/normalize.mjs` parses them, so a thousands separator is never guessed by the model. Adding a provider is one more entry in `providers` there.
122
+ - `raw/images.mjs` ties a product photo to a dish: by the model's box (a scan, or a page with no separately placed photo), or by matching the dish name to the page's text and taking the nearest placed image. `raw/import.mjs` saves the crops and fills `image_url` when a base URL is given.
101
123
  - Groq serves a single vision model, `qwen/qwen3.8-27b`, and it is a preview one: if Groq retires it, pass its successor with `--model`.
102
124
  - Rate limits (HTTP 429) are retried after the wait the service asks for, up to a minute; for a longer wait the run stops with the service's message. Pages already read are saved (below), so run it again later.
103
125
  - The model tags each section food, drink or other; `config.categories` maps those to the site's categories. A section with no heading continues the one before it, even across pages. A dish with several price columns (bottle and glass) becomes one product per column, "Name (Botella)".
@@ -12,6 +12,9 @@ export const menuFlags = {
12
12
  provider: "--provider",
13
13
  model: "--model",
14
14
  minWidth: "--min-width",
15
+ imageDir: "--images <folder>",
16
+ imageBaseUrl: "--image-base-url <url>",
17
+ imageBoxes: "--image-boxes <auto|always>",
15
18
  tablePrefix: "--table-prefix <prefix>",
16
19
  allowUnprefixed: "--allow-unprefixed",
17
20
  yes: "--yes",
@@ -7,8 +7,8 @@ export default {
7
7
  // empty only when this restaurant owns the unprefixed menu_* tables.
8
8
  tablePrefix: "mombasa_",
9
9
 
10
- // The page that shows the menu pictures, or direct image URLs.
11
- // `tablefacts menu raw <url> [<url>...]` overrides it for one run.
10
+ // The page that shows the menu pictures, a direct image URL, or a PDF file
11
+ // path/URL. `tablefacts menu raw <url> [<url>...]` overrides it for one run.
12
12
  url: "https://www.mombasa.co/carta-restaurante-espanol/",
13
13
 
14
14
  // Currency of the prices, an ISO code. Colombian menus print pesos as
@@ -20,6 +20,10 @@ export default {
20
20
  // for 95.000).
21
21
  scale: 1,
22
22
 
23
+ // Resolution a PDF page is rendered at before its printed product photos are
24
+ // screenshotted (2 means twice the page's size). Only used with `--images`.
25
+ imageScale: 2,
26
+
23
27
  // Each section the model reads is tagged food or drink (or other, which is
24
28
  // left out), and goes into the category that lists its group. The site's
25
29
  // default categories are cocina and bar.
@@ -1,8 +1,9 @@
1
1
  #!/usr/bin/env node
2
- // Imports a restaurant's menu from pictures of its pages into the Supabase
3
- // menu tables. See src/menu/README.md. Run from the project's folder:
2
+ // Imports a restaurant's menu from pictures of its pages, or from a PDF, into
3
+ // the Supabase menu tables. See src/menu/README.md. Run from the project's folder:
4
4
  // tablefacts menu raw --list
5
5
  // tablefacts menu raw --dry-run
6
+ // tablefacts menu raw carta.pdf --images ./photos --image-base-url https://cdn.example.com/menu/ --dry-run
6
7
  import { parseArgs } from "node:util";
7
8
  import { cliLog, menuFlags, runImport } from "../lib/run.mjs";
8
9
  import { cliMessage, exitCodeFor } from "../../lib/errors.mjs";
@@ -14,25 +15,32 @@ const providerTable = Object.entries(providers)
14
15
  .map(([name, { keyName, defaultModel }]) => ` ${name.padEnd(10)} ${keyName.padEnd(18)} ${defaultModel}`)
15
16
  .join("\n");
16
17
 
17
- const usage = `Usage: tablefacts menu raw [page-or-image-url...] [options]
18
+ const usage = `Usage: tablefacts menu raw [page-or-image-url...] [pdf-file-or-url...] [options]
18
19
 
19
- Reads a menu that is only pictures (one image per page) by transcribing each
20
- page with a vision model, and replaces the menu in Supabase.
21
- The URLs are pages to scan for menu images, or direct image URLs
22
- (default: config.mjs). The key of the provider that reads the pages goes in
23
- .env:
20
+ Reads a menu that is only pictures (one image per page) or a PDF, and replaces
21
+ the menu in Supabase. A PDF page is transcribed from its text when it has one,
22
+ and from a render of the page when it is a scan. The URLs are pages to scan for
23
+ menu images, or direct image URLs; an argument ending in .pdf is a PDF file
24
+ (local path) or URL. Default: config.mjs. The key of the provider that reads the
25
+ pages goes in .env:
24
26
 
25
27
  provider key default model
26
28
  ${providerTable}
27
29
 
28
- Image menu options:
29
- --list show the pictures found and stop; nothing is read or written
30
+ Picture and PDF options:
31
+ --list show the pages found and stop; nothing is read or written
30
32
  --only <pages> only these pages, numbered as --list shows them (e.g. 1,3-5)
31
33
  --provider <name> who reads the pages: ${providerNames}
32
34
  (default: MENU_VISION_PROVIDER in .env, else ${defaultProvider})
33
35
  --model <id> model of that provider (default: see the table)
34
36
  --min-width <px> ignore images declaring a smaller width (default: 500)
35
- --refresh read the pages again instead of using the saved transcriptions`;
37
+ --refresh read the pages again instead of using the saved transcriptions
38
+ --images <folder> also save the dish photos printed on a PDF page
39
+ --image-base-url <url> https folder you will publish them at; fills image_url
40
+ --image-boxes <auto|always> auto (default) reads the page as text and matches
41
+ photos by position; always reads every page as a picture so
42
+ the model boxes every dish's photo (a vision call per page)
43
+ `;
36
44
 
37
45
  const shared = {
38
46
  "min-width": { type: "string", default: "500" },
@@ -41,10 +49,17 @@ const shared = {
41
49
 
42
50
  if (process.argv.includes("--list")) {
43
51
  try {
44
- const { values, positionals } = parseArgs({ args: process.argv.slice(2), allowPositionals: true, options: { ...shared, list: { type: "boolean" }, help: { type: "boolean", short: "h" } }, strict: false });
45
- const { pages, chosen } = await findPages({ urls: positionals, only: values.only, minWidth: values["min-width"] });
46
- console.log(`${pages.length} images found:\n`);
47
- for (const page of chosen) console.log(` ${String(page.number).padStart(2)} ${page.url}${page.alt ? `\n ${page.alt.slice(0, 100)}` : ""}`);
52
+ const { values, positionals } = parseArgs({ args: process.argv.slice(2), allowPositionals: true, options: { ...shared, list: { type: "boolean" }, help: { type: "boolean", short: "h" }, images: { type: "string" }, "image-base-url": { type: "string" }, "image-boxes": { type: "string" } }, strict: false });
53
+ const { pages, chosen, close } = await findPages({ urls: positionals, only: values.only, minWidth: values["min-width"] });
54
+ try {
55
+ console.log(`${pages.length} pages found:\n`);
56
+ for (const page of chosen) {
57
+ const detail = page.kind === "pdf" ? (page.chars >= 40 ? `${page.chars} characters of text` : "scanned; read as a picture") : page.alt?.slice(0, 100) ?? "";
58
+ console.log(` ${String(page.number).padStart(2)} ${page.url}${detail ? `\n ${detail}` : ""}`);
59
+ }
60
+ } finally {
61
+ await close();
62
+ }
48
63
  } catch (error) {
49
64
  console.error(`\n${cliMessage(error, menuFlags)}`);
50
65
  process.exitCode = exitCodeFor(error);
@@ -57,6 +72,9 @@ if (process.argv.includes("--list")) {
57
72
  provider: { type: "string" },
58
73
  model: { type: "string" },
59
74
  refresh: { type: "boolean" },
75
+ images: { type: "string" },
76
+ "image-base-url": { type: "string" },
77
+ "image-boxes": { type: "string" },
60
78
  },
61
79
  fetchMenu: ({ values, positionals }) =>
62
80
  fetchImageMenu({
@@ -66,6 +84,9 @@ if (process.argv.includes("--list")) {
66
84
  model: values.model,
67
85
  minWidth: values["min-width"],
68
86
  refresh: values.refresh,
87
+ imageDir: values.images,
88
+ imageBaseUrl: values["image-base-url"],
89
+ imageBoxes: values["image-boxes"],
69
90
  log: cliLog,
70
91
  }),
71
92
  });
@@ -0,0 +1,109 @@
1
+ // Places the product photos found on a menu page onto the products read from it.
2
+ //
3
+ // Two page kinds feed this:
4
+ // - a digital PDF page: pdf.mjs reports where each image sits on the page
5
+ // (`placed` crops with a rectangle), and the product's name is found among
6
+ // the page's text items to pick the nearest photo;
7
+ // - a scanned page: the vision model was asked for each item's photo box
8
+ // (`direct` crops already tied to an item name).
9
+ //
10
+ // Matching is pure, so it is tested with made-up pages; raw/import.mjs saves the
11
+ // crops and turns them into `image_url`s.
12
+ import { matchKey } from "../lib/menu.mjs";
13
+
14
+ const escapeRegExp = (text) => text.replace(/[.*+?^${}()|[\]\\]/g, "\\$&");
15
+
16
+ /** The gap between two rectangles, 0 when they touch or overlap. */
17
+ const gap = (a, b) => {
18
+ const dx = Math.max(a.x0 - b.x1, b.x0 - a.x1, 0);
19
+ const dy = Math.max(a.y0 - b.y1, b.y0 - a.y1, 0);
20
+ return Math.hypot(dx, dy);
21
+ };
22
+
23
+ const listShort = (names) => names.slice(0, 6).join(", ") + (names.length > 6 ? `, and ${names.length - 6} more` : "");
24
+
25
+ const textRect = (item) => ({
26
+ x0: item.x ?? 0,
27
+ y0: item.y ?? 0,
28
+ x1: (item.x ?? 0) + Math.abs(item.width ?? 0),
29
+ y1: (item.y ?? 0) + Math.abs(item.height ?? 0),
30
+ });
31
+
32
+ /**
33
+ * The page's text item that best matches `name`: an exact match first, then a
34
+ * line that contains it (preferring the shortest such line). null when no line
35
+ * holds the name, which happens when a heading was drawn as an outline.
36
+ */
37
+ export function findNameItem(items, name) {
38
+ const wanted = matchKey(name);
39
+ if (!wanted) return null;
40
+ // Whole words only, so "Ana" does not match inside "Banana".
41
+ const bounded = new RegExp(`(^|[^\\p{L}\\p{N}])${escapeRegExp(wanted)}(?=$|[^\\p{L}\\p{N}])`, "u");
42
+ let best = null;
43
+ for (const item of items ?? []) {
44
+ const key = matchKey(item.str);
45
+ const exact = key === wanted;
46
+ const contains = !exact && bounded.test(key);
47
+ if (!exact && !contains) continue;
48
+ const score = exact ? -1 : key.length;
49
+ if (!best || score < best.score) best = { item, score };
50
+ }
51
+ return best?.item ?? null;
52
+ }
53
+
54
+ /**
55
+ * One crop per product, in page order. `pages` maps a page number to
56
+ * `{ width, items, placed, direct }` (see the file comment). A crop is used
57
+ * once, so a single photo between two items goes to the closer one. Returns
58
+ * `{ matches, notes }`; matches carry the placement and the crop it won.
59
+ * @param {{ page: number, name: string, box: number[] | null, products: any[] }[]} placements
60
+ * @param {Map<number, any>} pages
61
+ * @returns {{ matches: { placement: any, crop: any }[], notes: string[] }}
62
+ */
63
+ export function matchPlacements(placements, pages) {
64
+ const notes = [];
65
+ const matches = [];
66
+ const used = new Set();
67
+ const ambiguous = [];
68
+
69
+ for (const placement of placements) {
70
+ const page = pages.get(placement.page);
71
+ if (!page) continue;
72
+ let crop = null;
73
+
74
+ // A scanned page already tied each crop to the model's item name.
75
+ const direct = (page.direct ?? []).filter((c) => !used.has(c.id));
76
+ if (direct.length) crop = direct.find((c) => matchKey(c.name) === matchKey(placement.name)) ?? null;
77
+
78
+ if (!crop && (page.placed ?? []).length) {
79
+ const item = findNameItem(page.items, placement.name);
80
+ if (item) {
81
+ const rect = textRect(item);
82
+ const limit = (page.width ?? 0) * 0.2; // a photo farther than a fifth of the page is not this item's
83
+ const ranked = (page.placed ?? [])
84
+ .filter((c) => !used.has(c.id))
85
+ .map((c) => ({ c, d: gap(rect, c.rect) }))
86
+ .sort((a, b) => a.d - b.d);
87
+ if (ranked.length && ranked[0].d <= limit) {
88
+ crop = ranked[0].c;
89
+ // A near tie means the same photo could belong to either item: say so.
90
+ if (ranked[1] && ranked[0].d > 0 && Math.abs(ranked[1].d - ranked[0].d) < (page.width ?? 0) * 0.01) {
91
+ ambiguous.push(`${placement.name} (page ${placement.page})`);
92
+ }
93
+ }
94
+ }
95
+ }
96
+
97
+ if (crop) {
98
+ used.add(crop.id);
99
+ matches.push({ placement, crop });
100
+ }
101
+ }
102
+
103
+ // A dish without a photo is normal; only a photo with no dish is worth saying.
104
+ const allCrops = [...pages.values()].flatMap((page) => [...(page.placed ?? []), ...(page.direct ?? [])]);
105
+ const unused = allCrops.filter((crop) => !used.has(crop.id));
106
+ if (unused.length) notes.push(`${unused.length} printed photo(s) could not be matched to an item and were left out.`);
107
+ if (ambiguous.length) notes.push(`A printed photo sat between items and may be attached to the wrong one: ${listShort(ambiguous)}.`);
108
+ return { matches, notes };
109
+ }