@likolabs/i18nmd 0.1.1 → 0.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,9 @@
1
1
  # Changelog
2
2
 
3
+ ## 0.2.0
4
+
5
+ - **Static websites.** `extract` reads `.html` pages: it marks each sentence, the title and description, image text and form labels with a `data-i18n` attribute and leaves the English in place, and takes `/* i18n */` strings from inline scripts. `render <site> --out dist --url …` writes a copy of every page in every language (`/haw/…`) with `lang`, `hreflang` links, `og:url`, adjusted relative links, and a language switcher wherever a page has `<nav data-i18n-languages>`. No JavaScript needed.
6
+
3
7
  ## 0.1.1
4
8
 
5
9
  - Published to npm as `@likolabs/i18nmd` (`npm install --save-dev @likolabs/i18nmd`); the command is still `i18nmd`. npm 12 refuses GitHub and tarball URLs by default, so the 0.1.0 install instructions failed there.
package/PROMPT.md CHANGED
@@ -11,6 +11,8 @@ Give these instructions to a coding agent in the application's repository. They
11
11
  npx i18nmd extract src/components --out translations/ui --in-place
12
12
  ```
13
13
 
14
+ For a site of plain HTML pages, extract the pages (`npx i18nmd extract site --in-place`) and build each language with `npx i18nmd render site --out dist`.
15
+
14
16
  3. **Review the extractor's work.** Read the diff. Then work through every diagnostic it printed:
15
17
  - *looks like a count*: make the message an ICU plural, `{count, plural, one {# item} other {# items}}`.
16
18
  - *text chosen in code*: move the wording into the message as an ICU `select` or plural, instead of passing English through a placeholder.
package/README.md CHANGED
@@ -1,6 +1,6 @@
1
1
  # i18n.md
2
2
 
3
- Keep your interface strings in Markdown: one file per language, readable and editable by translators, reviewers and LLMs. A compiler checks every language and generates typed code your app imports.
3
+ Keep your interface strings in Markdown, for apps and for plain HTML sites: one file per language, readable and editable by translators, reviewers and LLMs. A compiler checks every language and generates typed code your app imports.
4
4
 
5
5
  ````md
6
6
  # Français
@@ -22,6 +22,7 @@ The files are the source of truth. They diff in pull requests, and anyone can ha
22
22
 
23
23
  - [Install](#install)
24
24
  - [Quick start](#quick-start)
25
+ - [Static websites](#static-websites)
25
26
  - [Language files](#language-files)
26
27
  - [Divisions](#divisions)
27
28
  - [Using the generated code](#using-the-generated-code)
@@ -97,6 +98,42 @@ export function LanguagePicker() {
97
98
 
98
99
  **5. Keep it current.** After editing source text, `npx i18nmd status` shows what needs translating and `npx i18nmd translate` fills it in.
99
100
 
101
+ ## Static websites
102
+
103
+ A site made of HTML pages needs no JavaScript at all. i18nmd writes a copy of each page in each language:
104
+
105
+ ```sh
106
+ npx i18nmd extract site --in-place # site/*.html → translations/i18n-en.md
107
+ npx i18nmd --add haw # or write translations/i18n-haw.md yourself
108
+ npx i18nmd render site --out dist --url https://example.com # dist/ and dist/haw/
109
+ ```
110
+
111
+ `extract` marks each piece of text with a `data-i18n` attribute and leaves the English where it is, so the page still opens and edits as before:
112
+
113
+ ```html
114
+ <h1 data-i18n="from_first_leaf_to_full_grown">From first leaf to <em class="red">full grown.</em></h1>
115
+ ```
116
+
117
+ The message keeps the sentence whole, with its inline markup as tags: `From first leaf to <em>full grown.</em>`. Translators move the tag; its class and other attributes stay in the page. Extract also takes the page `<title>`, the description and link-preview `<meta>` tags, `alt`, `title`, `placeholder` and `aria-label` attributes, and strings in inline scripts marked `/* i18n */`:
118
+
119
+ ```js
120
+ statusEl.textContent = /* i18n */ 'Sending…';
121
+ ```
122
+
123
+ It skips addresses such as `example.com`, and anything inside an element with `translate="no"`, the standard attribute for names and code. It lists text it can't place, such as words beside a block element, and script strings that look like text people read.
124
+
125
+ The English lives in your HTML. Edit a page, run `extract` again, and the source language file follows; the other languages' versions of that text become stale for `translate` to update.
126
+
127
+ `render` writes the source language where the pages are and every other language in a directory named for it, `/haw/`, copying everything else in the site alongside. Each copy gets:
128
+
129
+ - the translated text, with untranslated text left in the source language
130
+ - `<html lang>`, and `dir="rtl"` for right-to-left languages
131
+ - `<link rel="alternate" hreflang>` to every language's copy, so search engines index each one
132
+ - with `--url`, `og:url` and canonical links pointing at that copy
133
+ - relative links adjusted for the language directory
134
+
135
+ Put `<nav data-i18n-languages></nav>` anywhere on a page, and render fills it with a link to each language, named in that language. Deploy `dist/`.
136
+
100
137
  ## Language files
101
138
 
102
139
  `translations/i18n-fr.md` holds French and nothing else. The filename gives the language code (`fr`, `pt-BR`, or a custom name such as `pirate`). The first heading names the language in that language, and each `##` heading is a token:
@@ -346,6 +383,7 @@ FormatJS and next-intl already use ICU, so conversion is lossless, and FormatJS
346
383
 
347
384
  | Command | |
348
385
  | --- | --- |
386
+ | `render <site>` | `--out dist`, `--url https://example.com` |
349
387
  | `extract <src…>` | `--in-place`, `--out translations/<division>`, `--source en`, `--runtime src/i18n/i18n`, `--dest .i18n/src`, `--locale-expr locale` (calls use `i18nmd.in(locale)`) |
350
388
  | `compile` | `--out src/i18n`, `--target ts\|js\|json\|python`, `--eager`, `--skip <division,…>` |
351
389
  | `--add <language>`, `--top <n>`, `translate` | `--only fr,de`, `--dry-run`, `--provider`, `--base-url`, `--model`, `--batch 40` |
@@ -359,7 +397,7 @@ Every command takes `--dir`, and `--source` to override the source language reco
359
397
 
360
398
  ## Limits in 0.1
361
399
 
362
- - The extractor reads JavaScript and TypeScript. Strings in object properties (`{ label: "Save" }`) and plain `.ts` files need `/* i18n */` or the extraction prompt.
400
+ - The extractor reads HTML, JavaScript and TypeScript. Strings in object properties (`{ label: "Save" }`) and plain `.ts` files need `/* i18n */` or the extraction prompt.
363
401
  - Switching language loads that whole language at once, not per division.
364
402
  - The Python target ignores tags and formats numbers without locale grouping; dates are passed through as given.
365
403
 
package/bin/i18nmd.mjs CHANGED
@@ -6,6 +6,7 @@ import { spawn } from 'node:child_process';
6
6
  import { parseCatalog, serializeCatalog, parseLanguageFiles, languageFromFilename, renameToken, textOf, divisionOf, namespaceFor, accessorFor } from '../lib/catalog.mjs';
7
7
  import { generateModule, generateModules, divisionFile, runtimeTypes } from '../lib/compiler.mjs';
8
8
  import { extractSource, missingDivisionImports } from '../lib/extractor.mjs';
9
+ import { extractHtml, renderHtml } from '../lib/html.mjs';
9
10
  import { lockPathFor, findLock, namedRoot, readLock, serializeLock, report, syncLock, markCurrent } from '../lib/lock.mjs';
10
11
  import { resolveLanguage, topLanguages, RANKED } from '../lib/languages.mjs';
11
12
  import { llmConfig, translateLanguage } from '../lib/llm.mjs';
@@ -32,6 +33,8 @@ Keep files in shape
32
33
  One Markdown file with every language, and back.
33
34
 
34
35
  Build
36
+ render <site> [--out dist] [--url https://example.com]
37
+ Static HTML pages, one copy per language (/haw/…).
35
38
  compile [--out src/i18n] [--target ts|js|json|python] [--skip division,...] [--eager]
36
39
  extract <src...> [--out translations/<division>] [--in-place] [--runtime src/i18n/i18n]
37
40
  [--source en] [--locale-expr locale]
@@ -60,14 +63,16 @@ function python(args) {
60
63
  });
61
64
  }
62
65
 
63
- async function filesAt(input) {
64
- if (!(await stat(input)).isDirectory()) return /\.(?:[cm]?[jt]sx?)$/.test(input) ? [input] : [];
66
+ // Source files under input; html: also .html pages (for extract).
67
+ async function filesAt(input, { html = false } = {}) {
68
+ const wanted = html ? /\.(?:[cm]?[jt]sx?|html?)$/ : /\.(?:[cm]?[jt]sx?)$/;
69
+ if (!(await stat(input)).isDirectory()) return wanted.test(input) ? [input] : [];
65
70
  const files = [];
66
71
  for (const entry of (await readdir(input, { withFileTypes: true })).sort((a, b) => a.name.localeCompare(b.name))) {
67
72
  if (entry.name.startsWith('.') || ['node_modules', 'dist', 'build', 'vendor', 'coverage'].includes(entry.name)) continue;
68
73
  const file = path.join(input, entry.name);
69
- if (entry.isDirectory()) files.push(...await filesAt(file));
70
- else if (entry.isFile() && /\.(?:[cm]?[jt]sx?)$/.test(file) && !/\.(test|spec)\.[jt]sx?$/.test(file) && !/\.d\.[cm]?ts$/.test(file)) files.push(file);
74
+ if (entry.isDirectory()) files.push(...await filesAt(file, { html }));
75
+ else if (entry.isFile() && wanted.test(file) && !/\.(test|spec)\.[jt]sx?$/.test(file) && !/\.d\.[cm]?ts$/.test(file)) files.push(file);
71
76
  }
72
77
  return files;
73
78
  }
@@ -218,7 +223,7 @@ async function main() {
218
223
  if (!command || ['help', '--help', '-h'].includes(command)) { console.log(help); return; }
219
224
  if (['--version', '-v', 'version'].includes(command)) { console.log(JSON.parse(await readFile(new URL('../package.json', import.meta.url), 'utf8')).version); return; }
220
225
  const aliases = { locale: 'source', language: 'locale-expr' };
221
- const valued = /^(skip|out|source|locale|language|locale-expr|dest|runtime|target|table|languages|syntax|in|from|to|model|base-url|provider|batch|only|dir)$/;
226
+ const valued = /^(url|skip|out|source|locale|language|locale-expr|dest|runtime|target|table|languages|syntax|in|from|to|model|base-url|provider|batch|only|dir)$/;
222
227
  const options = Object.create(null), positional = [];
223
228
  for (let i = 0; i < rest.length; i++) {
224
229
  const arg = rest[i];
@@ -358,6 +363,45 @@ async function main() {
358
363
  const divisions = Object.keys(files).filter(n => !/^(?:i18n|language|runtime)\.|^languages\//.test(n) && !n.endsWith('.d.mts')).length;
359
364
  console.log(`Compiled ${plural(catalog.messages.length, 'token')} into ${options.out}: ${plural(divisions, 'division module')}, ${plural(Object.keys(catalog.languages).length - 1, 'language chunk')}${options.eager ? ' (eager)' : ''}.`); return;
360
365
  }
366
+ if (command === 'render') {
367
+ // Static pages, one copy per language: the source language where the pages
368
+ // are, every other language under its code (/haw/), and every other file copied.
369
+ const site = one();
370
+ if (!site || !(await isDirectory(site))) throw new Error('Usage: i18nmd render <site directory> [--out dist] [--url https://example.com]');
371
+ const out = path.resolve(options.out || 'dist');
372
+ const root = path.resolve(site);
373
+ if (out === root) throw new Error('Render into a directory other than the site itself.');
374
+ const ws = await open(options.dir, options); ws.warn();
375
+ const { catalog } = ws;
376
+ const base = (options.url || '').replace(/\/+$/, '');
377
+ if (base && !/^https?:\/\//.test(base)) throw new Error('--url needs a full address, such as https://example.com.');
378
+ const codes = Object.keys(catalog.languages);
379
+ const tables = Object.fromEntries(codes.map(code => [code, Object.fromEntries(catalog.messages.filter(m => Object.hasOwn(m.translations, code)).map(m => [m.key, m.translations[code]]))]));
380
+ const files = [];
381
+ const walk = async dir => {
382
+ for (const entry of (await readdir(dir, { withFileTypes: true })).sort((a, b) => a.name.localeCompare(b.name))) {
383
+ const file = path.join(dir, entry.name);
384
+ if (entry.name.startsWith('.') || entry.name === 'node_modules' || file === out) continue;
385
+ if (entry.isDirectory()) await walk(file); else if (entry.isFile()) files.push(path.relative(root, file).split(path.sep).join('/'));
386
+ }
387
+ };
388
+ await walk(root);
389
+ let pages = 0;
390
+ for (const file of files) {
391
+ if (!/\.html?$/.test(file)) { await mkdir(path.dirname(path.join(out, file)), { recursive: true }); await copyFile(path.join(root, file), path.join(out, file)); continue; }
392
+ const html = await readFile(path.join(root, file), 'utf8');
393
+ const address = code => `${base}/${code === catalog.source ? '' : code + '/'}${file.replace(/(^|\/)index\.html?$/, '$1')}`;
394
+ const urls = Object.fromEntries(codes.map(code => [code, address(code)]));
395
+ for (const code of codes) {
396
+ const target = code === catalog.source ? file : `${code}/${file}`;
397
+ const relative = code === catalog.source ? '' : path.posix.relative(path.posix.dirname(target), path.posix.dirname(file)) + '/';
398
+ await save(path.join(out, target), renderHtml(html, { language: code, source: catalog.source, messages: tables[code], pages: urls, languages: catalog.languages, relative, absolute: !!base }));
399
+ pages++;
400
+ }
401
+ }
402
+ for (const row of ws.rows()) if (row.missing.length || row.stale.length) console.warn(`i18nmd: ${summary(row)}; untranslated text stays in ${catalog.languages[catalog.source]}.`);
403
+ console.log(`Rendered ${plural(pages, 'page')} in ${plural(codes.length, 'language')} into ${path.relative(process.cwd(), out) || '.'}.`); return;
404
+ }
361
405
  if (command === 'join') {
362
406
  if (!options.out || !/\.md$/.test(options.out)) throw new Error('Choose the joined Markdown file with --out, e.g. --out translations.md.');
363
407
  const ws = await open(one(), options);
@@ -495,8 +539,8 @@ async function main() {
495
539
  const outDir = path.dirname(sourceFile);
496
540
  const namespace = namespaceFor(await divisionPrefix(outDir, (await findLock(outDir)).file));
497
541
  if (languageFromFilename(sourceFile) !== locale) throw new Error('The output filename must match --source.');
498
- const files = [...new Set((await Promise.all(positional.map(filesAt))).flat())];
499
- if (!files.length) throw new Error('No JavaScript or TypeScript source files found.');
542
+ const files = [...new Set((await Promise.all(positional.map(input => filesAt(input, { html: true })))).flat())];
543
+ if (!files.length) throw new Error('No HTML, JavaScript or TypeScript files found.');
500
544
  // The deepest directory holding every input; converted copies keep their paths below it.
501
545
  const dirs = await Promise.all(positional.map(async input => (await isDirectory(input)) ? path.resolve(input) : path.dirname(path.resolve(input))));
502
546
  let root = dirs[0];
@@ -513,21 +557,26 @@ async function main() {
513
557
  // Existing tokens are always kept, so extraction can be repeated as code changes.
514
558
  let catalog = await exists(sourceFile) ? parseCatalog(await readFile(sourceFile, 'utf8'), { filename: sourceFile }) : undefined;
515
559
  const outputs = [];
516
- let replacements = 0;
560
+ let replacements = 0, updated = 0;
517
561
  for (const file of files) {
518
562
  const relative = path.relative(root, path.resolve(file));
519
563
  const output = path.join(destination, relative);
520
564
  let importPath = path.relative(path.dirname(output), runtime).split(path.sep).join('/');
521
565
  if (!importPath.startsWith('.')) importPath = './' + importPath;
522
- const result = extractSource(await readFile(file, 'utf8'), { filename: path.relative(process.cwd(), file), locale, languageExpression: options['locale-expr'], existing: catalog, importPath, namespace });
523
- catalog = result.catalog; replacements += result.replacements;
566
+ const text = await readFile(file, 'utf8'), filename = path.relative(process.cwd(), file);
567
+ const result = /\.html?$/.test(file)
568
+ ? extractHtml(text, { filename, locale, existing: catalog, namespace })
569
+ : extractSource(text, { filename, locale, languageExpression: options['locale-expr'], existing: catalog, importPath, namespace });
570
+ catalog = result.catalog; replacements += result.replacements; updated += result.updated || 0;
524
571
  result.diagnostics.forEach(message => console.warn(message));
525
572
  if (result.replacements) outputs.push([output, result.source]);
526
573
  }
527
574
  // The lock marks the top of the translations tree, so divisions keep their names.
528
575
  const lockFile = (await findLock(outDir)).file;
529
576
  if (!(await exists(lockFile))) await save(lockFile, serializeLock({ source: locale, translations: {} }));
530
- if (!replacements) { console.log(`Nothing new to extract from ${plural(files.length, 'file')}. Mark other display strings with /* i18n */.`); return; }
577
+ if (!replacements && !updated) { console.log(`Nothing new to extract from ${plural(files.length, 'file')}. Mark other display strings with /* i18n */.`); return; }
578
+ // An HTML page holds its own English: an edited sentence updates the source file.
579
+ if (!replacements) { await save(sourceFile, serializeCatalog(catalog, { locale })); console.log(`Updated ${plural(updated, 'string')} in ${sourceFile} from the pages' English; run i18nmd status to see what to translate.`); return; }
531
580
  const markdown = serializeCatalog(catalog, { locale });
532
581
  for (const [file, text] of outputs) await save(file, text);
533
582
  await save(sourceFile, markdown);
package/lib/extractor.mjs CHANGED
@@ -7,7 +7,7 @@ const sourcePlugins = ['jsx', 'typescript'];
7
7
 
8
8
  // Readable tokens from the words of the message; a numeric suffix resolves collisions.
9
9
  // Tokens never encode the text, so editing the source message keeps the token.
10
- function tokenFor(text, taken) {
10
+ export function tokenFor(text, taken) {
11
11
  const words = text.toLowerCase().replace(/\{[^}]+\}|<\/?[\w-]+>/g, ' ').normalize('NFKD').replace(/[\u0300-\u036f]/g, '').split(/[^a-z0-9]+/).filter(Boolean);
12
12
  let slug = '';
13
13
  for (const word of words) { if (slug && slug.length + word.length >= 40) break; slug += (slug ? '_' : '') + word; }
package/lib/html.mjs ADDED
@@ -0,0 +1,309 @@
1
+ // Static HTML pages: extract marks each translatable element with
2
+ // data-i18n="token" and leaves the English in place; render writes one page per
3
+ // language from the Markdown files. English is authored in the HTML.
4
+ import { posix } from 'node:path';
5
+ import { parse } from 'parse5';
6
+ import { parse as parseScript } from '@babel/parser';
7
+ import { escapeMessageLiteral, parseMessage } from './messages.mjs';
8
+ import { languageName } from './catalog.mjs';
9
+ import { tokenFor } from './extractor.mjs';
10
+
11
+ // Elements that sit inside a sentence; they become ICU tags.
12
+ const INLINE = new Set(['a', 'abbr', 'b', 'bdi', 'bdo', 'cite', 'code', 'data', 'del', 'dfn', 'em', 'i', 'ins', 'kbd', 'mark', 'q', 's', 'samp', 'small', 'span', 'strong', 'sub', 'sup', 'time', 'u', 'var']);
13
+ // Empty elements allowed at the start or end of a sentence, such as a logo before a name.
14
+ const VOID = new Set(['img', 'br', 'wbr', 'input', 'svg']);
15
+ const SKIP = new Set(['script', 'style', 'template', 'svg', 'math', 'noscript', 'pre', 'textarea', 'select', 'iframe', 'object', 'canvas']);
16
+ const ATTRIBUTES = ['alt', 'title', 'placeholder', 'aria-label', 'aria-description', 'label'];
17
+ const META = new Set(['description', 'og:title', 'og:description', 'og:image:alt', 'og:site_name', 'twitter:title', 'twitter:description', 'twitter:image:alt']);
18
+ const URL_ATTRIBUTES = ['href', 'src', 'srcset', 'action', 'poster', 'formaction'];
19
+ const WHITESPACE = /[ \t\n\r\f]+/g;
20
+
21
+ const attr = (node, name) => node.attrs?.find(a => a.name === name)?.value;
22
+ const isElement = node => !!node.tagName;
23
+ const isText = node => node.nodeName === '#text';
24
+ const letters = text => /\p{L}/u.test(text);
25
+ // Addresses read the same in every language: domains, emails and URLs.
26
+ const address = text => /^(?:[\w.+-]+@[\w-]+(?:\.[\w-]+)+|(?:https?:\/\/)?[\w-]+(?:\.[\w-]+)+(?:\/\S*)?)$/i.test(text.trim());
27
+ const translatable = text => letters(text) && !address(text);
28
+ const noTranslate = node => attr(node, 'translate') === 'no';
29
+
30
+ function describe(node) {
31
+ const tag = node.tagName;
32
+ const kind = /^h[1-6]$/.test(tag) ? 'Heading' : { p: 'Paragraph', li: 'List item', a: 'Link', button: 'Button', label: 'Form label', title: 'Page title', option: 'Choice in a list', caption: 'Table caption', figcaption: 'Figure caption', th: 'Table heading', td: 'Table cell', legend: 'Form section title', summary: 'Disclosure summary', dt: 'Term', dd: 'Definition' }[tag] || `<${tag}>`;
33
+ for (let up = node.parentNode; up; up = up.parentNode) if (isElement(up) && attr(up, 'id')) return `${kind} in #${attr(up, 'id')}`;
34
+ return kind;
35
+ }
36
+
37
+ // A sentence: an element whose content is only text and inline elements, with
38
+ // any empty elements at its edges.
39
+ function sentence(node) {
40
+ if (!isElement(node) || SKIP.has(node.tagName) || noTranslate(node)) return null;
41
+ const children = node.childNodes.filter(c => c.nodeName !== '#comment');
42
+ let start = 0, end = children.length;
43
+ const blank = c => isText(c) && !c.value.trim();
44
+ const edge = c => blank(c) || (isElement(c) && VOID.has(c.tagName));
45
+ while (start < end && edge(children[start])) start++;
46
+ while (end > start && edge(children[end - 1])) end--;
47
+ const body = children.slice(start, end);
48
+ if (!body.length) return null;
49
+ const phrasing = c => isText(c) || c.nodeName === '#comment' || (isElement(c) && INLINE.has(c.tagName) && c.childNodes.every(phrasing));
50
+ if (!body.every(phrasing)) return null;
51
+ // Links side by side are separate items, not a sentence: a sentence has words of its own.
52
+ if (body.some(isElement) && !body.some(c => isText(c) && letters(c.value))) return null;
53
+ const text = body.map(function all(c) { return isText(c) ? c.value : (c.childNodes || []).map(all).join(''); }).join('');
54
+ return translatable(text) ? { body, lead: children.slice(0, start), trail: children.slice(end) } : null;
55
+ }
56
+
57
+ // The message for a sentence, and the element each tag name stands for. Names
58
+ // come from the element (em, a, a_2) and are rebuilt the same way when rendering.
59
+ function messageFor(body) {
60
+ const tags = new Map(), counts = new Map();
61
+ const walk = nodes => nodes.map(node => {
62
+ if (isText(node)) return escapeMessageLiteral(node.value.replace(WHITESPACE, ' '));
63
+ if (!isElement(node)) return '';
64
+ const n = (counts.get(node.tagName) || 0) + 1;
65
+ counts.set(node.tagName, n);
66
+ const name = n === 1 ? node.tagName : `${node.tagName}_${n}`;
67
+ tags.set(name, node);
68
+ return `<${name}>${walk(node.childNodes)}</${name}>`;
69
+ }).join('');
70
+ return { text: walk(body).trim(), tags };
71
+ }
72
+
73
+ /** Every translatable spot in a page: sentences, attributes and annotated script strings. */
74
+ function units(html) {
75
+ const document = parse(html, { sourceCodeLocationInfo: true });
76
+ const found = [], diagnostics = [];
77
+ const attributes = node => {
78
+ for (const name of ATTRIBUTES) if (translatable(attr(node, name) || '')) found.push({ kind: 'attribute', node, name, text: escapeMessageLiteral(attr(node, name).replace(WHITESPACE, ' ').trim()), context: `${name} text of ${describe(node)}` });
79
+ };
80
+ const visit = node => {
81
+ if (!isElement(node)) { (node.childNodes || []).forEach(visit); return; }
82
+ if (noTranslate(node)) return;
83
+ const meta = node.tagName === 'meta' && (attr(node, 'name') || attr(node, 'property'));
84
+ if (meta && META.has(meta) && translatable(attr(node, 'content') || '')) found.push({ kind: 'attribute', node, name: 'content', text: escapeMessageLiteral(attr(node, 'content')), context: `Page ${meta.replace(/^(og|twitter):/, '')} for search and link previews` });
85
+ attributes(node);
86
+ if (node.tagName === 'input' && ['submit', 'button', 'reset'].includes(attr(node, 'type')) && letters(attr(node, 'value') || '')) found.push({ kind: 'attribute', node, name: 'value', text: escapeMessageLiteral(attr(node, 'value')), context: `Button in ${describe(node)}` });
87
+ if (node.tagName === 'script' && !attr(node, 'src')) { scriptStrings(node, found, diagnostics, html); return; }
88
+ if (SKIP.has(node.tagName)) return;
89
+ const unit = sentence(node);
90
+ if (unit) {
91
+ const { text, tags } = messageFor(unit.body);
92
+ found.push({ kind: 'sentence', node, ...unit, text, tags, context: describe(node) });
93
+ // Elements inside a sentence keep their own attributes, such as an icon's alt text.
94
+ const inside = n => { for (const c of n.childNodes || []) if (isElement(c) && !noTranslate(c)) { attributes(c); inside(c); } };
95
+ inside(node);
96
+ return;
97
+ }
98
+ for (const child of node.childNodes || []) {
99
+ if (isText(child) && translatable(child.value)) diagnostics.push(`line ${child.sourceCodeLocation?.startLine}: text beside other elements in <${node.tagName}> ("${child.value.trim().slice(0, 40)}"); wrap it in an element such as <span> to translate it.`);
100
+ else visit(child);
101
+ }
102
+ };
103
+ visit(document);
104
+ return { document, found, diagnostics };
105
+ }
106
+
107
+ // Script strings that read like words for people: a capitalised word and a space or
108
+ // sentence punctuation, unlike ids, selectors, URLs and JSON keys.
109
+ const shown = text => /^\s*[\p{Lu}]/u.test(text) && /[\p{L}][\s.…!?]|[.…!?]$/u.test(text.trim()) && !address(text) && !/^[\w-]+$/.test(text.trim());
110
+
111
+ // Strings in an inline script marked /* i18n */ or /* i18n:token */.
112
+ function scriptStrings(node, found, diagnostics, html) {
113
+ const text = node.childNodes[0];
114
+ if (!text?.sourceCodeLocation) return;
115
+ const offset = text.sourceCodeLocation.startOffset;
116
+ let ast;
117
+ try { ast = parseScript(text.value, { sourceType: 'unambiguous', errorRecovery: true }); } catch { return; }
118
+ const visit = n => {
119
+ if (!n || typeof n.type !== 'string') return;
120
+ const comment = n.leadingComments?.findLast(c => /^\s*i18n(?:\s*$|:)/.test(c.value));
121
+ if (comment && (n.type === 'StringLiteral' || (n.type === 'TemplateLiteral' && !n.expressions.length))) {
122
+ const value = n.type === 'StringLiteral' ? n.value : n.quasis[0].value.cooked;
123
+ found.push({ kind: 'script', start: offset + n.start, end: offset + n.end, commentStart: offset + comment.start, commentEnd: offset + comment.end, explicit: /i18n:\s*([\w.-]+)/.exec(comment.value)?.[1], text: escapeMessageLiteral(value), context: 'Text a script on the page shows' });
124
+ } else if (!comment && (n.type === 'StringLiteral' || n.type === 'TemplateLiteral') && shown(n.value ?? n.quasis.map(q => q.value.cooked).join(''))) {
125
+ diagnostics.push(`line ${node.sourceCodeLocation.startLine + (n.loc?.start.line || 1) - 1}: a script string looks like text people read ("${(n.value ?? n.quasis[0].value.cooked).slice(0, 40)}"); mark it /* i18n */ to translate it.`);
126
+ } else if (comment) diagnostics.push(`line ${node.sourceCodeLocation.startLine + (n.loc?.start.line || 1) - 1}: /* i18n */ marks a string that isn't plain; use a string without \${…}.`);
127
+ for (const [key, value] of Object.entries(n)) {
128
+ if (key === 'loc' || key.endsWith('Comments')) continue;
129
+ if (Array.isArray(value)) value.forEach(visit); else if (value && typeof value.type === 'string') visit(value);
130
+ }
131
+ };
132
+ visit(ast.program);
133
+ }
134
+
135
+ const markerOf = unit => unit.kind === 'sentence' ? 'data-i18n' : unit.kind === 'attribute' ? `data-i18n-${unit.name}` : null;
136
+
137
+ /**
138
+ * Mark a page's translatable text with data-i18n attributes (and /* i18n:token *\/
139
+ * comments in scripts), adding new messages to the catalog. An element already
140
+ * marked keeps its token, and the page's English replaces the catalog's.
141
+ */
142
+ export function extractHtml(html, { existing, locale = 'en', namespace = '', filename = 'index.html' } = {}) {
143
+ const catalog = existing ? structuredClone(existing) : { title: 'Site', source: locale, syntax: 'icu', languages: { [locale]: languageName(locale) }, messages: [] };
144
+ const byKey = new Map(catalog.messages.map(m => [m.key, m]));
145
+ const byText = new Map(catalog.messages.map(m => [m.translations[locale], m]));
146
+ const { found, diagnostics } = units(html);
147
+ const edits = [];
148
+ let marked = 0, updated = 0;
149
+ for (const unit of found) {
150
+ const marker = markerOf(unit);
151
+ const given = unit.kind === 'script' ? unit.explicit : attr(unit.node, marker);
152
+ const short = given?.startsWith(namespace) ? given.slice(namespace.length) : given;
153
+ let message = short ? byKey.get(short) : byText.get(unit.text);
154
+ if (message && short && message.translations[locale] !== unit.text) {
155
+ // The page is where the English is written: take its text.
156
+ if (byText.get(message.translations[locale]) === message) byText.delete(message.translations[locale]);
157
+ message.translations[locale] = unit.text; byText.set(unit.text, message); updated++;
158
+ }
159
+ if (!message) {
160
+ const key = short || tokenFor(unit.text, byKey);
161
+ message = { key, context: unit.context, optional: [], translations: { [locale]: unit.text } };
162
+ catalog.messages.push(message); byKey.set(key, message); byText.set(unit.text, message);
163
+ }
164
+ if (given) continue;
165
+ const token = namespace + message.key;
166
+ marked++;
167
+ if (unit.kind === 'script') edits.push({ start: unit.commentStart, end: unit.commentEnd, text: `/* i18n:${token} */` });
168
+ else {
169
+ const tag = unit.node.sourceCodeLocation.startTag;
170
+ const close = html[tag.endOffset - 2] === '/' ? tag.endOffset - 2 : tag.endOffset - 1;
171
+ edits.push({ start: close, end: close, text: ` ${marker}="${token}"` });
172
+ }
173
+ }
174
+ let source = html;
175
+ for (const edit of edits.sort((a, b) => b.start - a.start)) source = source.slice(0, edit.start) + edit.text + source.slice(edit.end);
176
+ return { source, catalog, diagnostics: diagnostics.map(d => `${filename}:${d.replace(/^line /, '')}`), replacements: marked, updated };
177
+ }
178
+
179
+ const escapeHtml = text => text.replace(/&/g, '&amp;').replace(/</g, '&lt;').replace(/>/g, '&gt;').replace(/\u00a0/g, '&nbsp;');
180
+ const escapeAttribute = text => escapeHtml(text).replace(/"/g, '&quot;');
181
+ // The text of a message without markup, for attributes and scripts.
182
+ const plain = nodes => nodes.map(n => typeof n === 'string' ? n : n.type === 'tag' ? plain(n.children) : `{${n.name}}`).join('');
183
+
184
+ /**
185
+ * One page in one language. messages maps tokens to this language's message
186
+ * text (missing ones keep the page's English). pages lists every language's URL
187
+ * for this page, for hreflang links and the language switcher. relative: the
188
+ * path from this copy's directory to the original's ("../" for a page moved into
189
+ * haw/), put before the page's relative links. absolute: pages' URLs are full
190
+ * URLs, so og:url and canonical links can name this copy.
191
+ */
192
+ export function renderHtml(html, { language, source, messages = {}, pages = {}, languages = {}, relative = '', absolute = false } = {}) {
193
+ // Two passes, since a sentence's content holds tags whose attributes change too:
194
+ // first every tag and attribute edit, then each sentence rebuilt from those tags.
195
+ const first = renderAttributes(html, { language, source, messages, pages, languages, relative, absolute });
196
+ return renderSentences(first, { language, source, messages });
197
+ }
198
+
199
+ function translationFor(language, source, messages, token) {
200
+ if (language === source || !token || !Object.hasOwn(messages, token)) return null;
201
+ try { return parseMessage(messages[token]); } catch { return null; }
202
+ }
203
+
204
+ function applyEdits(html, edits) {
205
+ // From the end; insertions at one spot keep their order.
206
+ let out = html;
207
+ edits.map((edit, index) => ({ ...edit, index })).sort((a, b) => b.start - a.start || b.end - a.end || b.index - a.index)
208
+ .forEach(edit => { out = out.slice(0, edit.start) + edit.text + out.slice(edit.end); });
209
+ return out;
210
+ }
211
+
212
+ function attributeRemover(html, edits) {
213
+ return (node, name) => {
214
+ const loc = node.sourceCodeLocation.attrs?.[name];
215
+ if (!loc) return;
216
+ let start = loc.startOffset;
217
+ while (start > 0 && /\s/.test(html[start - 1])) start--;
218
+ edits.push({ start, end: loc.endOffset, text: '' });
219
+ };
220
+ }
221
+
222
+ function renderSentences(html, { language, source, messages }) {
223
+ const { found } = units(html);
224
+ const edits = [], removeAttribute = attributeRemover(html, edits);
225
+ for (const unit of found) {
226
+ if (unit.kind !== 'sentence') continue;
227
+ const token = attr(unit.node, 'data-i18n');
228
+ if (token === undefined) continue;
229
+ removeAttribute(unit.node, 'data-i18n');
230
+ const nodes = translationFor(language, source, messages, token);
231
+ if (!nodes) continue;
232
+ const outer = node => html.slice(node.sourceCodeLocation.startOffset, node.sourceCodeLocation.endOffset);
233
+ const inner = node => html.slice(node.sourceCodeLocation.startTag.endOffset, node.sourceCodeLocation.endTag?.startOffset ?? node.sourceCodeLocation.startTag.endOffset);
234
+ const write = list => list.map(n => {
235
+ if (typeof n === 'string') return escapeHtml(n);
236
+ if (n.type !== 'tag') return escapeHtml(`{${n.name}}`);
237
+ const element = unit.tags.get(n.name);
238
+ if (!element) return write(n.children);
239
+ const loc = element.sourceCodeLocation;
240
+ // translate="no" content stays as written.
241
+ return html.slice(loc.startOffset, loc.startTag.endOffset) + (noTranslate(element) ? inner(element) : write(n.children)) + (loc.endTag ? html.slice(loc.endTag.startOffset, loc.endTag.endOffset) : '');
242
+ }).join('');
243
+ // The sentence's own spacing at its edges, such as the space after a leading icon.
244
+ const body = unit.body.map(outer).join('');
245
+ const [, before, after] = /^([ \t\n\r\f]*)[\s\S]*?([ \t\n\r\f]*)$/.exec(body);
246
+ const loc = unit.node.sourceCodeLocation;
247
+ edits.push({ start: loc.startTag.endOffset, end: loc.endTag.startOffset, text: unit.lead.map(outer).join('') + before + write(nodes) + after + unit.trail.map(outer).join('') });
248
+ }
249
+ return applyEdits(html, edits);
250
+ }
251
+
252
+ function renderAttributes(html, { language, source, messages, pages, languages, relative, absolute }) {
253
+ const { document, found } = units(html);
254
+ const edits = [];
255
+ const replace = (start, end, text) => edits.push({ start, end, text });
256
+ const removeAttribute = attributeRemover(html, edits);
257
+ for (const unit of found) {
258
+ if (unit.kind === 'sentence') continue;
259
+ const marker = markerOf(unit);
260
+ const token = unit.kind === 'script' ? unit.explicit : attr(unit.node, marker);
261
+ const nodes = translationFor(language, source, messages, token);
262
+ if (unit.kind === 'script') { if (nodes) replace(unit.start, unit.end, JSON.stringify(plain(nodes))); continue; }
263
+ if (token) removeAttribute(unit.node, marker);
264
+ if (nodes) { const loc = unit.node.sourceCodeLocation.attrs[unit.name]; replace(loc.startOffset, loc.endOffset, `${unit.name}="${escapeAttribute(plain(nodes))}"`); }
265
+ }
266
+ const root = document.childNodes.find(n => n.tagName === 'html');
267
+ const head = root?.childNodes.find(n => n.tagName === 'head');
268
+ const all = function* (node) { yield node; for (const child of node.childNodes || []) yield* all(child); };
269
+ let direction = 'ltr';
270
+ try { direction = new Intl.Locale(language).getTextInfo?.().direction || new Intl.Locale(language).textInfo?.direction || 'ltr'; } catch { /* custom languages */ }
271
+ if (root?.sourceCodeLocation?.startTag) {
272
+ const tag = root.sourceCodeLocation.startTag;
273
+ for (const name of ['lang', 'dir']) removeAttribute(root, name);
274
+ replace(tag.endOffset - 1, tag.endOffset - 1, ` lang="${language}"${direction === 'rtl' ? ' dir="rtl"' : ''}`);
275
+ }
276
+ for (const node of all(document)) {
277
+ if (!isElement(node) || !node.sourceCodeLocation) continue;
278
+ // Pages in a language directory sit deeper: relative links go up to the same files.
279
+ if (relative) for (const name of URL_ATTRIBUTES) {
280
+ const value = attr(node, name);
281
+ if (value === undefined || node.sourceCodeLocation.attrs?.[name] === undefined) continue;
282
+ const fix = part => {
283
+ if (!part || /^(?:[a-z][a-z0-9+.-]*:|\/|#|\?)/i.test(part)) return part;
284
+ const [, file, rest] = /^([^?#]*)(.*)$/.exec(part);
285
+ return posix.normalize(relative + file) + rest;
286
+ };
287
+ const next = name === 'srcset' ? value.split(',').map(s => s.trim().replace(/^\S+/, fix)).join(', ') : fix(value);
288
+ if (next !== value) { const loc = node.sourceCodeLocation.attrs[name]; replace(loc.startOffset, loc.endOffset, `${name}="${escapeAttribute(next)}"`); }
289
+ }
290
+ // Links that name this page's own address point at this language's copy.
291
+ if (absolute && ((node.tagName === 'meta' && attr(node, 'property') === 'og:url') || (node.tagName === 'link' && attr(node, 'rel') === 'canonical'))) {
292
+ const name = node.tagName === 'meta' ? 'content' : 'href';
293
+ const loc = node.sourceCodeLocation.attrs[name];
294
+ if (loc) replace(loc.startOffset, loc.endOffset, `${name}="${escapeAttribute(pages[language])}"`);
295
+ }
296
+ if (attr(node, 'data-i18n-languages') !== undefined) {
297
+ removeAttribute(node, 'data-i18n-languages');
298
+ const loc = node.sourceCodeLocation;
299
+ const links = Object.keys(pages).map(code => `<a href="${escapeAttribute(pages[code])}" hreflang="${code}" lang="${code}"${code === language ? ' aria-current="page"' : ''}>${escapeHtml(languages[code] || code)}</a>`).join(' ');
300
+ if (loc.endTag) replace(loc.startTag.endOffset, loc.endTag.startOffset, links);
301
+ }
302
+ }
303
+ if (head?.sourceCodeLocation?.endTag && Object.keys(pages).length > 1) {
304
+ const alternates = Object.keys(pages).map(code => `<link rel="alternate" hreflang="${code}" href="${escapeAttribute(pages[code])}">`);
305
+ alternates.push(`<link rel="alternate" hreflang="x-default" href="${escapeAttribute(pages[source])}">`);
306
+ replace(head.sourceCodeLocation.endTag.startOffset, head.sourceCodeLocation.endTag.startOffset, alternates.join('\n') + '\n');
307
+ }
308
+ return applyEdits(html, edits);
309
+ }
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@likolabs/i18nmd",
3
- "version": "0.1.1",
3
+ "version": "0.2.0",
4
4
  "description": "Interface strings in Markdown, one file per language. A compiler checks every language and generates typed code your app imports.",
5
5
  "author": "Liko Labs (https://likolabs.com)",
6
6
  "type": "module",
@@ -31,7 +31,8 @@
31
31
  "node": ">=22"
32
32
  },
33
33
  "dependencies": {
34
- "@babel/parser": "7.29.0"
34
+ "@babel/parser": "7.29.0",
35
+ "parse5": "^8.0.1"
35
36
  },
36
37
  "repository": {
37
38
  "type": "git",