localewarden 0.1.1 → 0.2.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +54 -9
- package/dist/checks.d.ts +19 -2
- package/dist/checks.js +94 -18
- package/dist/cli.js +24 -3
- package/dist/config.d.ts +10 -0
- package/dist/config.js +13 -0
- package/dist/files.d.ts +21 -0
- package/dist/files.js +52 -1
- package/dist/placeholders.js +2 -1
- package/dist/project.d.ts +15 -0
- package/dist/project.js +84 -4
- package/dist/prompt.d.ts +4 -0
- package/dist/prompt.js +9 -2
- package/dist/review.js +25 -5
- package/dist/scope.d.ts +15 -0
- package/dist/scope.js +26 -0
- package/dist/state.d.ts +1 -1
- package/dist/translate.js +216 -169
- package/package.json +12 -1
package/README.md
CHANGED
|
@@ -11,7 +11,7 @@ npx localewarden # translate new and changed strings
|
|
|
11
11
|
npx localewarden check # quality check, no API calls (use it in CI)
|
|
12
12
|
```
|
|
13
13
|
|
|
14
|
-
Works with i18next, react-intl / FormatJS, vue-i18n, next-intl, ngx-translate and any other setup that keeps strings in JSON files. Uses any OpenAI-compatible API (OpenAI, OpenRouter, a local Ollama, ...).
|
|
14
|
+
Works with i18next, react-intl / FormatJS, vue-i18n, next-intl, ngx-translate and any other setup that keeps strings in JSON files, with Flutter (`.arb` files) and with fastlane's App Store / Play Store metadata (`.txt` files). Uses any OpenAI-compatible API (OpenAI, OpenRouter, a local Ollama, ...).
|
|
15
15
|
|
|
16
16
|
## Why
|
|
17
17
|
|
|
@@ -22,15 +22,22 @@ Translating locale files with a language model is easy once. Keeping 20 language
|
|
|
22
22
|
- **Models are inconsistent across batches.** French screens mix "tu" and "vous", Spanish copies English Title Case ("Configure Su Cuenta"), and Polish or Russian address every user as a man.
|
|
23
23
|
- **Broken output ships silently.** A translated placeholder (`{heures}` instead of `{hours}`) shows raw braces in your app. A dropped `</strong>` breaks the layout. Stray Cyrillic letters end up in a Danish sentence.
|
|
24
24
|
|
|
25
|
-
localewarden grew out of the translation pipeline of a production app that ships in 38 languages. Every rule and check in it exists because one of these failures happened in real output.
|
|
25
|
+
localewarden grew out of the translation pipeline of a production app that ships in 38 languages. Every rule and check in it exists because one of these failures happened in real output. The checks are tuned against that app's real texts (UI, website, long-form learning content and store listings, about 165 MB) to report problems without flooding you with false alarms. On those texts they still find things that slipped through earlier pipelines: sections cut off after the English grew, stray letters from other alphabets, a trial notice left in English.
|
|
26
|
+
|
|
27
|
+
## How it compares
|
|
28
|
+
|
|
29
|
+
- **Translation platforms** (Crowdin, Lokalise, Phrase, Weblate) are hosted services with editors, translator workflows and review for teams. localewarden is a small CLI that runs in your repository and CI, with no account and no server. If you have professional translators, a platform fits better. If a model translates and people only fix the odd string, this is the lighter setup.
|
|
30
|
+
- **"Translate my JSON with GPT" scripts** usually send every string on every run and overwrite whatever is there. localewarden keeps state, so it only sends what changed, keeps human fixes, and checks the output.
|
|
31
|
+
- **Editor extensions** (such as i18n Ally) help you write and look up keys while coding. localewarden is about filling and maintaining 10 to 40 languages afterwards. The two work well together.
|
|
26
32
|
|
|
27
33
|
## What it does
|
|
28
34
|
|
|
29
35
|
- **Translates only what changed.** It remembers a hash of each source string per language. New strings are translated. Changed strings are *revised*: the model gets the existing translation and changes only what the source change requires. Removed strings are deleted from every language.
|
|
30
36
|
- **Protects hand edits.** If someone edited a translation, localewarden detects it, keeps it, and lists it for review. If the source of a hand-edited string changes later, the string is flagged instead of overwritten.
|
|
31
|
-
- **Checks every result before writing it.** Broken placeholders, injected HTML or scripts, foreign alphabets, changed links, broken HTML and echoed source text are rejected (retried once, then left for the next run). Softer problems are retried and reported.
|
|
37
|
+
- **Checks every result before writing it.** Broken placeholders, injected HTML or scripts, foreign alphabets, changed links, broken HTML and echoed source text are rejected (retried once, then left for the next run). Softer problems (too long, content missing, words left in English) are retried and reported.
|
|
32
38
|
- **Consistent style per language.** It enforces formal or informal address per language (`du`/`Sie`, `tu`/`vous`, `ты`/`вы` and 16 more), uses sentence case where the language does, avoids gendered forms for "you", and applies local typography (French spacing, `92 %` in German, CJK quotation marks).
|
|
33
39
|
- **Plural forms per language.** For i18next-style keys (`item_one`, `item_other`) it adds the forms a language needs but English lacks, such as Polish `_few` and `_many` or Arabic `_zero`, `_two`, `_few` and `_many` (CLDR plural rules).
|
|
40
|
+
- **Data files and store listings.** Fields like `id`, `type` or `image` are copied instead of translated (`ignoreKeys`), and so are URLs, email addresses and file paths. Length limits per key (`maxLength`) are passed to the model and checked: App Store names, SEO titles, buttons.
|
|
34
41
|
- **Glossary and protected names.** You choose fixed renderings ("Privacy Policy" -> "Politique de confidentialité") and names that must never be translated. The check accepts grammatical case endings.
|
|
35
42
|
- **Quality check for CI.** `localewarden check` runs all checks without any API calls and exits non-zero on errors.
|
|
36
43
|
- **Targeted repair.** `--fix-flagged` asks the model to fix only what the check flagged. The fix is accepted only if the problem is gone and little else changed.
|
|
@@ -79,7 +86,7 @@ es Mantén vivas tus plantas sin tener que pensar en ello
|
|
|
79
86
|
ja 何も考えなくても、植物を元気に保てます
|
|
80
87
|
```
|
|
81
88
|
|
|
82
|
-
German uses "du" and French "vous", as configured. Spanish and French use sentence case, not the English Title Case. French has its space before "!". The hedge "tend to" survived, and so did the placeholders, the link and the brand name. The full example is in [`examples/basic`](examples/basic).
|
|
89
|
+
German uses "du" and French "vous", as configured. Spanish and French use sentence case, not the English Title Case. French has its space before "!". The hedge "tend to" survived, and so did the placeholders, the link and the brand name. The full example is in [`examples/basic`](examples/basic). There are also examples for [Flutter ARB files](examples/flutter) and [App Store / Play Store texts with fastlane](examples/fastlane).
|
|
83
90
|
|
|
84
91
|
## Quick start
|
|
85
92
|
|
|
@@ -135,24 +142,28 @@ The same checks run in two places. Right after each model answer, a failed hard
|
|
|
135
142
|
| --- | --- | --- |
|
|
136
143
|
| `placeholder` | `{name}`, `{{count}}`, `%s`, `%1$d`, `%{x}`, `${x}`, `<0></0>` renamed, translated, added or dropped. ICU `plural`/`select` arguments are compared, while plural categories may differ per language. | error |
|
|
137
144
|
| `unsafe` | HTML tags, attributes, event handlers or `javascript:`/`data:` URLs that the source does not have. Translations are often rendered as raw HTML, so this would be a script injection. | error |
|
|
138
|
-
| `script` | Letters from an alphabet the language does not use ("刺激" in German),
|
|
145
|
+
| `script` | Letters from an alphabet the language does not use ("刺激" in German), a word that mixes Latin with Cyrillic/Greek lookalikes ("Вarda"), or Simplified characters in Traditional Chinese (`zh-TW`) and the reverse | error |
|
|
139
146
|
| `markup` | Changed link targets, different number of tags, unclosed or misnested tags, dropped list items | warning (broken tags and changed links: never written) |
|
|
140
147
|
| `years` | A year from the source missing or changed (citations, dates) | warning |
|
|
141
148
|
| `formality` | The other form of address than configured, both forms in one string, or masculine-only forms for "you" | warning |
|
|
149
|
+
| `length` | Longer than the `maxLength` configured for the key | warning |
|
|
142
150
|
| `titlecase` | English Title Case copied into a language that uses sentence case | warning |
|
|
143
151
|
| `ampersand` | "&" in languages that write the word | warning |
|
|
144
152
|
| `glossary` | A glossary rendering missing (case endings allowed), or a `doNotTranslate` name translated | warning |
|
|
145
153
|
| `untranslated` | Identical to the source (prose of 3+ words; "OK" and names are fine) | warning |
|
|
146
|
-
| `partial` | Source-language words left inside the translation, an untranslated bold lead-in,
|
|
154
|
+
| `partial` | Source-language words left inside the translation, an untranslated bold lead-in, a hedge that became certainty ("tend to" stated as fact), a translation much shorter than its source (content cut off, or the source grew after it was translated), or sibling options that got the same translation although the source differs ("Rarely" and "Occasionally" both "Selten") | warning |
|
|
147
155
|
|
|
148
156
|
```bash
|
|
149
157
|
npx localewarden check # counts per language and check
|
|
150
158
|
npx localewarden check -v # with examples
|
|
151
159
|
npx localewarden check --strict # exit 1 on warnings too
|
|
152
160
|
npx localewarden check --json # for scripts
|
|
161
|
+
npx localewarden check --fix # repair placeholders with one possible fix, no API calls
|
|
153
162
|
```
|
|
154
163
|
|
|
155
|
-
|
|
164
|
+
`--fix` repairs a translated placeholder when the source has exactly one and the translation renamed it (`{stunden}` back to `{hours}`). The file is edited in place, so its formatting stays as it is. Anything less certain is left for `--fix-flagged` or a person.
|
|
165
|
+
|
|
166
|
+
Approved strings are skipped, except for errors (placeholder, unsafe, script), which break the app either way. You can approve any string, not only hand edits: `npx localewarden review --approve de:home.title` tells the check that a person looked at it (for example a pun on a brand name that is correct without the name), and runs leave it alone.
|
|
156
167
|
|
|
157
168
|
### In CI
|
|
158
169
|
|
|
@@ -240,6 +251,9 @@ npx localewarden review --release de:home.title # hand it back: next run revis
|
|
|
240
251
|
| `glossary` | `{}` | `{"fr": {"Terms of Service": "Conditions d'utilisation"}}` |
|
|
241
252
|
| `termNotes` | `{}` | Meanings of ambiguous terms, sent only with strings that contain them: `{"snooze": "postpone a reminder"}` |
|
|
242
253
|
| `instructions` | `{}` | Extra instructions per language, `"*"` for all: `{"es": "Use neutral Latin American Spanish."}` |
|
|
254
|
+
| `ignoreKeys` | `[]` | Keys that are not text, copied from the source: `["id", "type", "**.sources.*"]`. `*` matches within a key segment, `**` across segments; a pattern without a dot matches the last segment anywhere. URLs, emails, file paths and numbers are always copied. |
|
|
255
|
+
| `exclude` | `[]` | Source files to skip: `["locales/{lang}/nav.json"]` |
|
|
256
|
+
| `maxLength` | `{}` | Character limits per key pattern: `{"**.meta.title": 60, "name": 30}`. The model is told the limit; longer results are retried once and reported by the `length` check. |
|
|
243
257
|
| `placeholders` | built-in | Regular expressions (strings) that match your placeholders. Replaces the built-in list. |
|
|
244
258
|
| `model` | `"gpt-5.4-mini"` | Any chat model your endpoint offers |
|
|
245
259
|
| `baseUrl` | `"https://api.openai.com/v1"` | Any OpenAI-compatible endpoint |
|
|
@@ -251,6 +265,37 @@ npx localewarden review --release de:home.title # hand it back: next run revis
|
|
|
251
265
|
| `batchSize` | `20` | Strings per request (smaller for scripts that need many tokens) |
|
|
252
266
|
| `stateDir` | `".localewarden"` | Where state and the review list live |
|
|
253
267
|
|
|
268
|
+
### App Store and Play Store listings (fastlane)
|
|
269
|
+
|
|
270
|
+
```json
|
|
271
|
+
{
|
|
272
|
+
"sourceLanguage": "en-US",
|
|
273
|
+
"targetLanguages": ["de-DE", "fr-FR", "ja"],
|
|
274
|
+
"files": "fastlane/metadata/{lang}/*.txt",
|
|
275
|
+
"exclude": ["fastlane/metadata/{lang}/*_url.txt"],
|
|
276
|
+
"maxLength": { "name": 30, "subtitle": 30, "keywords": 100, "promotional_text": 170, "description": 4000 },
|
|
277
|
+
"termNotes": { "keywords": "a comma-separated keyword list for store search, not a sentence" }
|
|
278
|
+
}
|
|
279
|
+
```
|
|
280
|
+
|
|
281
|
+
Each `.txt` file is one string, keyed by its file name, so the limits above apply to `name.txt`, `subtitle.txt` and so on.
|
|
282
|
+
|
|
283
|
+
### Flutter (ARB)
|
|
284
|
+
|
|
285
|
+
```json
|
|
286
|
+
{ "files": "lib/l10n/app_{lang}.arb", "targetLanguages": ["de", "fr", "pt_BR"] }
|
|
287
|
+
```
|
|
288
|
+
|
|
289
|
+
Metadata (`@@locale`, `@key` descriptions and placeholders) is copied, not translated, and `@@locale` is set to the target language. ICU plurals and selects keep their structure, and each language gets the plural categories it needs.
|
|
290
|
+
|
|
291
|
+
### Data files
|
|
292
|
+
|
|
293
|
+
For content JSON with ids, types and links, list the non-text keys:
|
|
294
|
+
|
|
295
|
+
```json
|
|
296
|
+
{ "files": "content/**/*.{lang}.json", "ignoreKeys": ["id", "type", "category", "image", "**.sources.*"] }
|
|
297
|
+
```
|
|
298
|
+
|
|
254
299
|
### Other providers
|
|
255
300
|
|
|
256
301
|
```json
|
|
@@ -268,7 +313,7 @@ Small local models make noticeably more mistakes. The checks catch the mechanica
|
|
|
268
313
|
```text
|
|
269
314
|
localewarden [translate] --dry-run --lang de,fr --fix-flagged --retranslate-all
|
|
270
315
|
--overwrite-manual --max-tokens <n> --verbose
|
|
271
|
-
localewarden check --lang de,fr --verbose --limit <n> --strict --json
|
|
316
|
+
localewarden check --lang de,fr --verbose --limit <n> --strict --json --fix
|
|
272
317
|
localewarden review --all --approve <sel>... --release <sel>...
|
|
273
318
|
localewarden init
|
|
274
319
|
Global: --config <path> --help --version
|
|
@@ -300,7 +345,7 @@ Translations are treated as untrusted: any markup the source does not have is bl
|
|
|
300
345
|
|
|
301
346
|
## Limitations
|
|
302
347
|
|
|
303
|
-
- JSON
|
|
348
|
+
- JSON (nested objects, arrays, flat keys), Flutter ARB and plain `.txt` files. YAML, PO and XLIFF are not supported yet.
|
|
304
349
|
- The checks catch mechanical problems, not every wrong meaning. Have a native speaker look at important screens, then approve their edits with `review`.
|
|
305
350
|
- Rules for form of address, gender and typography exist for the languages listed above. Other languages are translated with the general rules.
|
|
306
351
|
- A run that is interrupted keeps everything written so far. Unwritten strings are picked up on the next run.
|
package/dist/checks.d.ts
CHANGED
|
@@ -1,4 +1,5 @@
|
|
|
1
1
|
import type { Config } from './config.js';
|
|
2
|
+
import { Scope } from './scope.js';
|
|
2
3
|
/**
|
|
3
4
|
* Deterministic quality checks. No API calls, so they can run in CI on every commit.
|
|
4
5
|
*
|
|
@@ -11,14 +12,16 @@ import type { Config } from './config.js';
|
|
|
11
12
|
* years a year from the source is missing or changed (citations, dates)
|
|
12
13
|
* formality the other form of address than configured, or both mixed;
|
|
13
14
|
* masculine-only forms for "you" when genderNeutral is on
|
|
15
|
+
* length longer than the maxLength configured for the key
|
|
14
16
|
* titlecase English Title Case copied into a sentence-case language
|
|
15
17
|
* ampersand "&" in a language that writes the word
|
|
16
18
|
* glossary a glossary rendering or a doNotTranslate name is missing
|
|
17
19
|
* untranslated identical to the source (prose of 3+ words)
|
|
18
20
|
* partial source-language words left inside an otherwise translated string,
|
|
19
|
-
*
|
|
21
|
+
* a dropped hedge ("tend to" stated as certain), or a translation much
|
|
22
|
+
* shorter than its source (content missing or cut off)
|
|
20
23
|
*/
|
|
21
|
-
export type CheckName = 'placeholder' | 'unsafe' | 'script' | 'markup' | 'years' | 'formality' | 'titlecase' | 'ampersand' | 'glossary' | 'untranslated' | 'partial';
|
|
24
|
+
export type CheckName = 'placeholder' | 'unsafe' | 'script' | 'markup' | 'years' | 'formality' | 'length' | 'titlecase' | 'ampersand' | 'glossary' | 'untranslated' | 'partial';
|
|
22
25
|
export declare const CHECKS: CheckName[];
|
|
23
26
|
export declare const ERROR_CHECKS: Set<CheckName>;
|
|
24
27
|
/** Checks a targeted repair (--fix-flagged) may try to fix. */
|
|
@@ -30,7 +33,14 @@ export interface Issue {
|
|
|
30
33
|
export declare class Checker {
|
|
31
34
|
readonly config: Config;
|
|
32
35
|
readonly placeholderRe: RegExp;
|
|
36
|
+
readonly scope: Scope;
|
|
37
|
+
/** Words of termNotes and doNotTranslate: terms a translation may keep in the source language. */
|
|
38
|
+
readonly keptWords: Set<string>;
|
|
33
39
|
constructor(config: Config);
|
|
40
|
+
/** The text without doNotTranslate names, which stay the same in every language. */
|
|
41
|
+
withoutNames(text: string): string;
|
|
42
|
+
/** "62 characters, limit 60", or null. */
|
|
43
|
+
tooLong(key: string, text: string): string | null;
|
|
34
44
|
get englishSource(): boolean;
|
|
35
45
|
placeholdersMatch(key: string, source: string, text: string): boolean;
|
|
36
46
|
placeholderNote(source: string, text: string): string;
|
|
@@ -50,6 +60,8 @@ export declare class Checker {
|
|
|
50
60
|
}
|
|
51
61
|
/** Non-Latin scripts each language is written in. Latin is always allowed (names, codes). */
|
|
52
62
|
export declare const NATIVE_SCRIPTS: Record<string, string[]>;
|
|
63
|
+
/** Simplified characters in Traditional Chinese text, or the reverse (2+ distinct ones). */
|
|
64
|
+
export declare function wrongChineseScript(lang: string, text: string): string | null;
|
|
53
65
|
/** Why `text` contains letters that cannot belong to `lang`, or null. */
|
|
54
66
|
export declare function foreignScript(lang: string, text: string): string | null;
|
|
55
67
|
/**
|
|
@@ -82,3 +94,8 @@ export declare function isUnchangedProse(source: string, text: string, placehold
|
|
|
82
94
|
* allowed) copied into the translation, or null.
|
|
83
95
|
*/
|
|
84
96
|
export declare function sourceRun(source: string, text: string): string | null;
|
|
97
|
+
/**
|
|
98
|
+
* A translation with a fraction of the source's length lost content: cut off by the model, or
|
|
99
|
+
* the source grew after it was translated. Markup and placeholders are not counted.
|
|
100
|
+
*/
|
|
101
|
+
export declare function muchShorter(lang: string, source: string, text: string): string | null;
|
package/dist/checks.js
CHANGED
|
@@ -1,6 +1,7 @@
|
|
|
1
1
|
import { placeholderRegExp, placeholderSignature, placeholdersMatch } from './placeholders.js';
|
|
2
2
|
import { GENDERED_FORMS, NO_AMPERSAND_LANGUAGES, SENTENCE_CASE_LANGUAGES, formalityRule, registerFor, withoutQuotedSpeech, } from './style.js';
|
|
3
|
-
import {
|
|
3
|
+
import { Scope } from './scope.js';
|
|
4
|
+
import { baseLanguage, escapeRegExp, isTraditionalChinese } from './util.js';
|
|
4
5
|
export const CHECKS = [
|
|
5
6
|
'placeholder',
|
|
6
7
|
'unsafe',
|
|
@@ -8,6 +9,7 @@ export const CHECKS = [
|
|
|
8
9
|
'markup',
|
|
9
10
|
'years',
|
|
10
11
|
'formality',
|
|
12
|
+
'length',
|
|
11
13
|
'titlecase',
|
|
12
14
|
'ampersand',
|
|
13
15
|
'glossary',
|
|
@@ -21,6 +23,7 @@ export const FIXABLE_CHECKS = new Set([
|
|
|
21
23
|
'markup',
|
|
22
24
|
'years',
|
|
23
25
|
'formality',
|
|
26
|
+
'length',
|
|
24
27
|
'titlecase',
|
|
25
28
|
'ampersand',
|
|
26
29
|
'glossary',
|
|
@@ -29,9 +32,24 @@ export const FIXABLE_CHECKS = new Set([
|
|
|
29
32
|
export class Checker {
|
|
30
33
|
config;
|
|
31
34
|
placeholderRe;
|
|
35
|
+
scope;
|
|
36
|
+
/** Words of termNotes and doNotTranslate: terms a translation may keep in the source language. */
|
|
37
|
+
keptWords;
|
|
32
38
|
constructor(config) {
|
|
33
39
|
this.config = config;
|
|
34
40
|
this.placeholderRe = placeholderRegExp(config.placeholders);
|
|
41
|
+
this.scope = new Scope(config);
|
|
42
|
+
this.keptWords = new Set([...Object.keys(config.termNotes), ...config.doNotTranslate].flatMap(term => term.toLowerCase().split(/[^\p{L}]+/u)).filter(Boolean));
|
|
43
|
+
}
|
|
44
|
+
/** The text without doNotTranslate names, which stay the same in every language. */
|
|
45
|
+
withoutNames(text) {
|
|
46
|
+
return this.config.doNotTranslate.reduce((value, name) => value.split(name).join(' '), text);
|
|
47
|
+
}
|
|
48
|
+
/** "62 characters, limit 60", or null. */
|
|
49
|
+
tooLong(key, text) {
|
|
50
|
+
const max = this.scope.maxLength(key);
|
|
51
|
+
const length = [...text].length;
|
|
52
|
+
return max !== undefined && length > max ? `${length} characters, limit ${max}` : null;
|
|
35
53
|
}
|
|
36
54
|
get englishSource() {
|
|
37
55
|
return baseLanguage(this.config.sourceLanguage) === 'en';
|
|
@@ -88,16 +106,22 @@ export class Checker {
|
|
|
88
106
|
this.config.sentenceCase &&
|
|
89
107
|
SENTENCE_CASE_LANGUAGES.has(base) &&
|
|
90
108
|
isEnglishTitleCase(source)) {
|
|
91
|
-
const capitals = midCapitals(text, source, this.config.doNotTranslate);
|
|
109
|
+
const capitals = midCapitals(text, source, this.config.doNotTranslate, lang);
|
|
92
110
|
if (capitals.length >= (source.trim().split(/\s+/).length <= 3 ? 1 : 2)) {
|
|
93
111
|
add('titlecase', `capitalised: ${capitals.join(' ')}`);
|
|
94
112
|
}
|
|
95
113
|
}
|
|
114
|
+
const long = this.tooLong(key, text);
|
|
115
|
+
if (long)
|
|
116
|
+
add('length', long);
|
|
96
117
|
const ampersands = (value) => value.split(' & ').length - 1;
|
|
97
118
|
if (NO_AMPERSAND_LANGUAGES.has(base) && ampersands(text) > ampersands(source))
|
|
98
119
|
add('ampersand');
|
|
99
|
-
if (isUnchangedProse(source, text, this.placeholderRe))
|
|
120
|
+
if (isUnchangedProse(this.withoutNames(source), this.withoutNames(text), this.placeholderRe))
|
|
100
121
|
add('untranslated');
|
|
122
|
+
const short = muchShorter(lang, source, text);
|
|
123
|
+
if (short)
|
|
124
|
+
add('partial', short);
|
|
101
125
|
if (this.englishSource && text !== source) {
|
|
102
126
|
const copied = sourceRun(source, text);
|
|
103
127
|
if (copied)
|
|
@@ -105,7 +129,7 @@ export class Checker {
|
|
|
105
129
|
const lead = !copied ? sourceBoldLeadIn(source, text, this.config.doNotTranslate) : null;
|
|
106
130
|
if (lead)
|
|
107
131
|
add('partial', `bold lead-in still in the source language: "${lead}"`);
|
|
108
|
-
const mixed = !copied && !lead ? englishInNativeScript(lang, source, text, this.placeholderRe) : null;
|
|
132
|
+
const mixed = !copied && !lead ? englishInNativeScript(lang, source, text, this.placeholderRe, this.keptWords) : null;
|
|
109
133
|
if (mixed)
|
|
110
134
|
add('partial', mixed);
|
|
111
135
|
if (droppedHedge(lang, source, text)) {
|
|
@@ -127,7 +151,7 @@ export class Checker {
|
|
|
127
151
|
const placeholders = this.placeholdersMatch(key, source, text) ? null : `placeholder mismatch: ${this.placeholderNote(source, text)}`;
|
|
128
152
|
const foreign = foreignScript(this.config.sourceLanguage, source) ? null : foreignScript(lang, text);
|
|
129
153
|
const links = hrefSignature(source) !== hrefSignature(text) ? `links changed: [${hrefSignature(source)}] -> [${hrefSignature(text)}]` : null;
|
|
130
|
-
const echoed = isUnchangedProse(source, text, this.placeholderRe) ? 'returned the source text unchanged' : null;
|
|
154
|
+
const echoed = isUnchangedProse(this.withoutNames(source), this.withoutNames(text), this.placeholderRe) ? 'returned the source text unchanged' : null;
|
|
131
155
|
const broken = brokenMarkup(source) ? null : brokenMarkup(text);
|
|
132
156
|
const boldMarkers = (value) => (value.match(/\*\*/g) ?? []).length % 2;
|
|
133
157
|
const brokenBold = boldMarkers(text) === 1 && boldMarkers(source) === 0 ? 'unbalanced ** markers' : null;
|
|
@@ -135,7 +159,7 @@ export class Checker {
|
|
|
135
159
|
const hard = placeholders ?? unsafeAdditions(source, text) ?? foreign ?? links ?? echoed ?? broken ?? brokenBold ?? droppedBullets(source, text) ?? leaked;
|
|
136
160
|
const tags = tagCount(source) !== tagCount(text) ? `${tagCount(source)} tags in the source, got ${tagCount(text)}` : null;
|
|
137
161
|
const copied = this.englishSource && text !== source ? sourceRun(source, text) : null;
|
|
138
|
-
return { hard, soft: hard ?? tags ?? yearDifference(source, text) ?? (copied ? `source text left in: "${copied}"` : null) };
|
|
162
|
+
return { hard, soft: hard ?? this.tooLong(key, text) ?? muchShorter(lang, source, text) ?? tags ?? yearDifference(source, text) ?? (copied ? `source text left in: "${copied}"` : null) };
|
|
139
163
|
}
|
|
140
164
|
}
|
|
141
165
|
// ---------------------------------------------------------------------------------------
|
|
@@ -188,8 +212,25 @@ const LATIN_LANGUAGES = new Set([
|
|
|
188
212
|
]);
|
|
189
213
|
// Latin glued to a lookalike alphabet inside one word ("Вarda": Cyrillic В + Latin arda).
|
|
190
214
|
const HOMOGLYPH_WORD = /(?=\p{L}*\p{Script=Latin})(?=\p{L}*[\p{Script=Cyrillic}\p{Script=Greek}])\p{L}+/u;
|
|
215
|
+
// Common characters that exist in only one of the two Chinese scripts (pairs at the same index).
|
|
216
|
+
const SIMPLIFIED_ONLY = '们这说时会来对个为发过还让现实动门问间题体关点应开东头书长见认学页电话语读写买卖钱网设计习惯觉机帮爱车钟儿气无边进选择检样经验数据项结种类业务环节';
|
|
217
|
+
const TRADITIONAL_ONLY = '們這說時會來對個為發過還讓現實動門問間題體關點應開東頭書長見認學頁電話語讀寫買賣錢網設計習慣覺機幫愛車鐘兒氣無邊進選擇檢樣經驗數據項結種類業務環節';
|
|
218
|
+
/** Simplified characters in Traditional Chinese text, or the reverse (2+ distinct ones). */
|
|
219
|
+
export function wrongChineseScript(lang, text) {
|
|
220
|
+
if (baseLanguage(lang) !== 'zh')
|
|
221
|
+
return null;
|
|
222
|
+
const traditional = isTraditionalChinese(lang);
|
|
223
|
+
const wrong = traditional ? SIMPLIFIED_ONLY : TRADITIONAL_ONLY;
|
|
224
|
+
const found = [...new Set([...text].filter(ch => wrong.includes(ch)))];
|
|
225
|
+
return found.length >= 2
|
|
226
|
+
? `${traditional ? 'Simplified' : 'Traditional'} Chinese characters in ${lang}: ${found.slice(0, 6).join('')}`
|
|
227
|
+
: null;
|
|
228
|
+
}
|
|
191
229
|
/** Why `text` contains letters that cannot belong to `lang`, or null. */
|
|
192
230
|
export function foreignScript(lang, text) {
|
|
231
|
+
const chinese = wrongChineseScript(lang, text);
|
|
232
|
+
if (chinese)
|
|
233
|
+
return chinese;
|
|
193
234
|
const base = baseLanguage(lang);
|
|
194
235
|
const native = NATIVE_SCRIPTS[base] ?? (LATIN_LANGUAGES.has(base) ? [] : null);
|
|
195
236
|
if (native === null)
|
|
@@ -218,13 +259,18 @@ const unescapeEntities = (text) => text
|
|
|
218
259
|
.replace(/"|�*34;|�*22;/gi, '"')
|
|
219
260
|
.replace(/�*39;|�*27;|'/gi, "'")
|
|
220
261
|
.replace(/:|�*58;|�*3a;/gi, ':');
|
|
221
|
-
|
|
222
|
-
const
|
|
262
|
+
// HTML elements, so text in angle brackets ("<minutes>", "<your name>") is not taken for markup.
|
|
263
|
+
const HTML_ELEMENTS = new Set(('a abbr address area article aside audio b base bdi bdo blockquote body br button canvas caption cite code col colgroup data datalist dd del details dfn dialog div dl dt em embed fieldset figcaption figure footer form frame frameset h1 h2 h3 h4 h5 h6 head header hr html i iframe img input ins kbd label legend li link main map mark math meta meter nav noscript object ol optgroup option output p param picture pre progress q rp rt ruby s samp script section select slot small source span strong style sub summary sup svg table tbody td template textarea tfoot th thead time title tr track u ul var video wbr animate foreignobject use image set').split(' '));
|
|
264
|
+
/** HTML element names, in lower case. Numbered <0> tags (react-i18next) are placeholders. */
|
|
265
|
+
const tagNames = (html) => new Set([...html.matchAll(/<\/?([a-z][\w-]*)/gi)].map(m => m[1].toLowerCase()).filter(tag => HTML_ELEMENTS.has(tag)));
|
|
223
266
|
const TEXT_ATTRIBUTES = new Set(['title', 'alt', 'aria-label', 'aria-description', 'placeholder']);
|
|
224
267
|
/** Every attribute as "name=value" (value without quotes), in lower case. */
|
|
225
268
|
const attributes = (html) => {
|
|
226
269
|
const found = new Set();
|
|
227
|
-
for (const [, inner] of html.matchAll(/<[a-z][\w-]
|
|
270
|
+
for (const [, tag, inner] of html.matchAll(/<([a-z][\w-]*)\s([^>]*)>?/gi)) {
|
|
271
|
+
// "<dakika 20)" (sw: under 20 minutes) is text, not a tag.
|
|
272
|
+
if (!HTML_ELEMENTS.has(tag.toLowerCase()))
|
|
273
|
+
continue;
|
|
228
274
|
for (const [, name, v1, v2, v3] of inner.matchAll(/([^\s"'=<>\/]+)(?:\s*=\s*(?:"([^"]*)"|'([^']*)'|([^\s"'>]+)))?/g)) {
|
|
229
275
|
const attr = name.toLowerCase();
|
|
230
276
|
// Text attributes are translated along with the visible text; only their presence counts.
|
|
@@ -234,7 +280,10 @@ const attributes = (html) => {
|
|
|
234
280
|
}
|
|
235
281
|
return found;
|
|
236
282
|
};
|
|
237
|
-
|
|
283
|
+
// Script URLs in a link target or Markdown link. Plain "data:" in prose is a word (pt/it "date:").
|
|
284
|
+
const DANGEROUS_URL = /(?:(?:href|src|action|formaction|xlink:href|poster|background)\s*=\s*["']?|\]\()\s*(?:javascript|vbscript|data)\s*:/gi;
|
|
285
|
+
// Formatting a translator may add for emphasis or a title (no attributes): a markup warning, not a risk.
|
|
286
|
+
const HARMLESS_TAGS = new Set(['br', 'i', 'b', 'em', 'strong', 'u', 's', 'sub', 'sup', 'small', 'mark', 'q', 'cite']);
|
|
238
287
|
/**
|
|
239
288
|
* Markup the translation adds that the source does not have: a new tag type, a new or changed
|
|
240
289
|
* attribute, an event handler or a script URL. Translations are often rendered as raw HTML
|
|
@@ -244,7 +293,7 @@ const DANGEROUS_URL = /(?:javascript|vbscript|data)\s*:/gi;
|
|
|
244
293
|
export function unsafeAdditions(source, text) {
|
|
245
294
|
const [src, out] = [unescapeEntities(source), unescapeEntities(text)];
|
|
246
295
|
const srcTags = tagNames(src);
|
|
247
|
-
const newTags = [...tagNames(out)].filter(tag => !srcTags.has(tag) && tag
|
|
296
|
+
const newTags = [...tagNames(out)].filter(tag => !srcTags.has(tag) && !HARMLESS_TAGS.has(tag));
|
|
248
297
|
if (newTags.length > 0)
|
|
249
298
|
return `HTML tag not in the source: <${newTags.join('>, <')}>`;
|
|
250
299
|
const srcAttrs = attributes(src);
|
|
@@ -312,7 +361,8 @@ const asciiDigits = (text) => text.replace(/\p{Nd}/gu, ch => {
|
|
|
312
361
|
export function yearDifference(source, text) {
|
|
313
362
|
// Decades ("the 2020s") are written in words or with suffixes in many languages.
|
|
314
363
|
const decades = new Set(asciiDigits(source).match(/(?<!\d)(?:19|20)\d0(?=['’]?s\b)/g) ?? []);
|
|
315
|
-
|
|
364
|
+
// "1,900", "1.900" and "1 900" are numbers, not years: drop thousands separators first.
|
|
365
|
+
const years = (value) => (asciiDigits(value).replace(/(\d)[,.\u00a0\u202f ](?=\d{3}(?!\d))/g, '$1').match(/(?<!\d)(?:19|20)\d\d(?!\d)/g) ?? []).filter(y => !decades.has(y)).sort().join(',');
|
|
316
366
|
const [a, b] = [years(source), years(text)];
|
|
317
367
|
return a === b ? null : `source years [${a}], got [${b}]`;
|
|
318
368
|
}
|
|
@@ -359,20 +409,25 @@ export function keptNameMissing(name, source, text) {
|
|
|
359
409
|
// ---------------------------------------------------------------------------------------
|
|
360
410
|
// Title case
|
|
361
411
|
const COMMON_NAMES = new Set(['iOS', 'Android', 'Apple', 'Google', 'iPhone', 'iPad', 'Mac', 'Windows', 'Linux', 'AI', 'API', 'URL', 'PDF', 'FAQ', 'OK']);
|
|
362
|
-
|
|
412
|
+
// Polish capitalises "you" pronouns as a sign of respect ("W Twoim planie"): correct, not Title Case.
|
|
413
|
+
const POLISH_RESPECT = /^(?:Ty|Twój|Twoja|Twoje|Twojego|Twojej|Twoim|Twoją|Twoich|Twoimi|Ciebie|Cię|Tobie|Tobą|Wy|Wasz|Wasza|Wasze|Wam|Was|Wami)$/u;
|
|
414
|
+
function midCapitals(text, source, names, lang = '') {
|
|
363
415
|
const allowed = new Set([...COMMON_NAMES, ...names.flatMap(name => name.split(/\s+/))]);
|
|
364
416
|
const isName = (word) => allowed.has(word) ||
|
|
365
417
|
new RegExp(`(?<!\\p{L})${escapeRegExp(word)}(?!\\p{L})`, 'u').test(source) ||
|
|
366
418
|
[...allowed].some(name => name.length >= 4 && word.startsWith(name));
|
|
367
419
|
return text
|
|
368
|
-
|
|
420
|
+
// Commas and brackets start segments too: list items ("SMART: Specific, Measurable") and
|
|
421
|
+
// bracketed words ("Customize (Optional)") are capitalised in many languages.
|
|
422
|
+
.split(/[.!?:—–\n•|,;()]+/)
|
|
369
423
|
.flatMap(segment => {
|
|
370
424
|
const words = segment.trim().split(/\s+/);
|
|
371
425
|
const first = words.findIndex(word => /\p{L}/u.test(word));
|
|
372
426
|
return first === -1 ? [] : words.slice(first + 1);
|
|
373
427
|
})
|
|
374
428
|
.map(word => word.replace(/^[^\p{L}]+|[^\p{L}]+$/gu, ''))
|
|
375
|
-
.filter(word => word.length >= 3 && /^\p{Lu}\p{Ll}/u.test(word) && !isName(word))
|
|
429
|
+
.filter(word => word.length >= 3 && /^\p{Lu}\p{Ll}/u.test(word) && !isName(word))
|
|
430
|
+
.filter(word => !(baseLanguage(lang) === 'pl' && POLISH_RESPECT.test(word)));
|
|
376
431
|
}
|
|
377
432
|
function isEnglishTitleCase(source) {
|
|
378
433
|
const rest = source
|
|
@@ -430,26 +485,34 @@ function sourceBoldLeadIn(source, text, names) {
|
|
|
430
485
|
for (const lead of bold(source)) {
|
|
431
486
|
if (lead.length <= 15 || !/\p{Ll}{3}/u.test(lead) || names.some(name => lead.includes(name)))
|
|
432
487
|
continue;
|
|
488
|
+
// Only capitalised words ("Apple App Store:", "Google Play"): a name, kept on purpose.
|
|
489
|
+
if (lead.split(/\s+/).every(word => !/\p{L}/u.test(word) || /^[^\p{L}]*\p{Lu}/u.test(word)))
|
|
490
|
+
continue;
|
|
433
491
|
if (translated.has(lead))
|
|
434
492
|
return lead;
|
|
435
493
|
}
|
|
436
494
|
return null;
|
|
437
495
|
}
|
|
496
|
+
const ICU_HEADER = /\{\s*[\w.-]+\s*,\s*(?:plural|select|selectordinal)\s*,/g;
|
|
497
|
+
const ICU_HEADER_TEST = /\{\s*[\w.-]+\s*,\s*(?:plural|select|selectordinal)\s*,/;
|
|
438
498
|
// Loanwords commonly written in Latin script inside non-Latin text.
|
|
439
499
|
const LATIN_LOANWORDS = new Set(['email', 'online', 'offline', 'emoji', 'smartphone', 'podcast', 'podcasts', 'wifi', 'blog', 'login', 'like', 'likes']);
|
|
440
500
|
/** Source-language words left inside a non-Latin-script translation ("Settings → Privacy"). */
|
|
441
|
-
function englishInNativeScript(lang, source, text, placeholderRe) {
|
|
501
|
+
function englishInNativeScript(lang, source, text, placeholderRe, allowed) {
|
|
442
502
|
const base = baseLanguage(lang);
|
|
443
503
|
// Greek writes many anglicisms in Latin script; not checked.
|
|
444
504
|
if (!NATIVE_SCRIPTS[base] || base === 'el' || base === 'sr' || text === source)
|
|
445
505
|
return null;
|
|
446
506
|
const strip = (value) => withoutUrls(value.replace(/[\w.+-]+@[\w.-]+/g, ' ').replace(/<[^>]+>/g, ' '))
|
|
507
|
+
// ICU syntax ("{count, plural, one {…} other {…}}") is code, not English text.
|
|
508
|
+
.replace(ICU_HEADER, ' ')
|
|
509
|
+
.replace(ICU_HEADER_TEST.test(value) ? /(?:^|[\s}])(?:zero|one|two|few|many|other|=\d+|[\w-]+)\s*(?=\{)/g : /$^/g, ' ')
|
|
447
510
|
.replace(placeholderRe, ' ')
|
|
448
511
|
.replace(/[((][^))]*[))]/g, ' ')
|
|
449
512
|
.replace(/["“„«「『‘'][^"”“»」』’']*["”“»」』’']/g, ' ');
|
|
450
513
|
const sourceWords = new Set(strip(source).match(/\b[a-z]{4,}\b/g) ?? []);
|
|
451
514
|
const left = [
|
|
452
|
-
...new Set((strip(text).match(/(?<![\p{L}-])[a-z]{4,}(?![\p{L}])/gu) ?? []).filter(word => sourceWords.has(word) && !LATIN_LOANWORDS.has(word))),
|
|
515
|
+
...new Set((strip(text).match(/(?<![\p{L}-])[a-z]{4,}(?![\p{L}])/gu) ?? []).filter(word => sourceWords.has(word) && !LATIN_LOANWORDS.has(word) && !allowed.has(word))),
|
|
453
516
|
];
|
|
454
517
|
const menu = /\b[A-Z][a-z]+ → [A-Z][a-z]+/.exec(text);
|
|
455
518
|
if (menu && source.includes(menu[0]))
|
|
@@ -478,9 +541,22 @@ const HEDGE_MARKERS = {
|
|
|
478
541
|
cs: /tendenc|obvykle|často|zpravidla|většinou|bývá|sklon|snadno|častěji|mív/iu,
|
|
479
542
|
sk: /tendenc|obvykle|často|zvyčajne|väčšinou|býva|sklon|ľahko|častejšie|zvyk|skôr/iu,
|
|
480
543
|
el: /τείν|συχνά|συνήθως|τάση|συνήθ|εύκολα|συχνότερα/iu,
|
|
481
|
-
fi: /taipu|usein|yleensä|tapaa|tavallisesti|tuppaa|tyypillisesti|taipumus|helposti|useimmiten|herkästi/iu,
|
|
544
|
+
fi: /taipu|tapana|usein|yleensä|tapaa|tavallisesti|tuppaa|tyypillisesti|taipumus|helposti|useimmiten|herkästi/iu,
|
|
482
545
|
ca: /tendeix|tendència|sol|sovint|generalment|acostum|normalment|fàcilment|freqüent/iu,
|
|
483
546
|
};
|
|
547
|
+
// Chinese, Japanese and Korean need far fewer characters than English; Thai has no spaces.
|
|
548
|
+
const COMPACT_SCRIPTS = new Set(['zh', 'ja', 'ko']);
|
|
549
|
+
/**
|
|
550
|
+
* A translation with a fraction of the source's length lost content: cut off by the model, or
|
|
551
|
+
* the source grew after it was translated. Markup and placeholders are not counted.
|
|
552
|
+
*/
|
|
553
|
+
export function muchShorter(lang, source, text) {
|
|
554
|
+
const visible = (value) => [...value.replace(/<[^>]+>|\{\{?[^{}]*\}?\}/g, '').replace(/\s+/g, ' ').trim()].length;
|
|
555
|
+
const [a, b] = [visible(source), visible(text)];
|
|
556
|
+
// Real translations from English rarely drop below 60% (Chinese/Japanese/Korean: 25%).
|
|
557
|
+
const min = COMPACT_SCRIPTS.has(baseLanguage(lang)) ? 0.2 : 0.5;
|
|
558
|
+
return a >= 200 && b < a * min ? `translation much shorter than the source (${b} of ${a} characters): content missing?` : null;
|
|
559
|
+
}
|
|
484
560
|
function droppedHedge(lang, source, text) {
|
|
485
561
|
const markers = HEDGE_MARKERS[baseLanguage(lang)];
|
|
486
562
|
return Boolean(markers) && /\btends? to\b/i.test(source) && !markers.test(text);
|
package/dist/cli.js
CHANGED
|
@@ -5,7 +5,7 @@ import { ConfigError, CONFIG_FILE, loadConfig } from './config.js';
|
|
|
5
5
|
import { CHECKS } from './checks.js';
|
|
6
6
|
import { findSourceFiles } from './files.js';
|
|
7
7
|
import { isReasoningModel } from './llm.js';
|
|
8
|
-
import { checkProject, summaryTable } from './project.js';
|
|
8
|
+
import { checkProject, fixPlaceholders, summaryTable } from './project.js';
|
|
9
9
|
import { listReview, updateReview } from './review.js';
|
|
10
10
|
import { run } from './translate.js';
|
|
11
11
|
const HELP = `localewarden - incremental AI translation for JSON locale files
|
|
@@ -31,6 +31,7 @@ Check options:
|
|
|
31
31
|
--limit <n> findings shown per check with --verbose (default 20)
|
|
32
32
|
--strict exit 1 on warnings too (default: only placeholder/script errors)
|
|
33
33
|
--json print findings as JSON
|
|
34
|
+
--fix repair placeholders with exactly one possible fix ({heures} -> {hours})
|
|
34
35
|
|
|
35
36
|
Review options:
|
|
36
37
|
--all include approved entries
|
|
@@ -93,6 +94,8 @@ const LAYOUTS = [
|
|
|
93
94
|
'src/assets/i18n/{lang}.json',
|
|
94
95
|
'i18n/{lang}.json',
|
|
95
96
|
'lang/{lang}.json',
|
|
97
|
+
'lib/l10n/app_{lang}.arb',
|
|
98
|
+
'lib/l10n/intl_{lang}.arb',
|
|
96
99
|
];
|
|
97
100
|
function init() {
|
|
98
101
|
const file = path.resolve(CONFIG_FILE);
|
|
@@ -180,7 +183,15 @@ async function translateCommand(args) {
|
|
|
180
183
|
function checkCommand(args) {
|
|
181
184
|
const config = loadConfig(args.values.get('--config')?.[0]);
|
|
182
185
|
const languages = languagesArg(args) ?? config.targetLanguages;
|
|
183
|
-
|
|
186
|
+
let findings = checkProject(config, languages);
|
|
187
|
+
if (args.flags.has('--fix')) {
|
|
188
|
+
const fixed = fixPlaceholders(config, findings);
|
|
189
|
+
for (const f of fixed)
|
|
190
|
+
console.log(`fixed [${f.lang}] ${f.file} ${f.key}: ${f.text}`);
|
|
191
|
+
console.log(`${fixed.length} placeholder(s) repaired.\n`);
|
|
192
|
+
if (fixed.length > 0)
|
|
193
|
+
findings = checkProject(config, languages);
|
|
194
|
+
}
|
|
184
195
|
if (args.flags.has('--json')) {
|
|
185
196
|
console.log(JSON.stringify(findings, null, 2));
|
|
186
197
|
}
|
|
@@ -200,7 +211,17 @@ function checkCommand(args) {
|
|
|
200
211
|
console.log(` ... ${items.length - limit} more (--limit <n>)`);
|
|
201
212
|
}
|
|
202
213
|
}
|
|
203
|
-
|
|
214
|
+
// Errors fail CI, so show them even without --verbose.
|
|
215
|
+
const errorItems = findings.filter(f => f.severity === 'error');
|
|
216
|
+
if (!args.flags.has('--verbose') && errorItems.length > 0) {
|
|
217
|
+
console.log('\nErrors:');
|
|
218
|
+
for (const f of errorItems.slice(0, 20)) {
|
|
219
|
+
console.log(` [${f.lang}] ${f.file} ${f.key} - ${f.check}${f.note ? `: ${f.note}` : ''}`);
|
|
220
|
+
}
|
|
221
|
+
if (errorItems.length > 20)
|
|
222
|
+
console.log(` ... ${errorItems.length - 20} more (--verbose)`);
|
|
223
|
+
}
|
|
224
|
+
const errors = errorItems.length;
|
|
204
225
|
console.log(`\n${errors} error(s), ${findings.length - errors} warning(s).${findings.length && !args.flags.has('--verbose') ? ' Details: --verbose' : ''}`);
|
|
205
226
|
}
|
|
206
227
|
const errors = findings.some(f => f.severity === 'error');
|
package/dist/config.d.ts
CHANGED
|
@@ -24,6 +24,16 @@ export interface Config {
|
|
|
24
24
|
termNotes: Record<string, string>;
|
|
25
25
|
/** Extra instructions per language; "*" applies to every language. */
|
|
26
26
|
instructions: Record<string, string>;
|
|
27
|
+
/**
|
|
28
|
+
* Keys whose values are not text (ids, types, image paths). Copied from the source, never
|
|
29
|
+
* translated. "*" matches within one key segment, "**" across segments; a pattern without
|
|
30
|
+
* a dot matches the last segment anywhere ("id" matches "steps.2.id").
|
|
31
|
+
*/
|
|
32
|
+
ignoreKeys: string[];
|
|
33
|
+
/** Source files to skip, as path patterns relative to the config ("locales/{lang}/nav.json"). */
|
|
34
|
+
exclude: string[];
|
|
35
|
+
/** Maximum characters per key pattern: { "**.meta.title": 60 }. Told to the model and checked. */
|
|
36
|
+
maxLength: Record<string, number>;
|
|
27
37
|
/** Regular expressions (as strings) that match placeholders. Replaces the built-in list. */
|
|
28
38
|
placeholders?: string[];
|
|
29
39
|
model: string;
|
package/dist/config.js
CHANGED
|
@@ -10,6 +10,9 @@ export const DEFAULTS = {
|
|
|
10
10
|
glossary: {},
|
|
11
11
|
termNotes: {},
|
|
12
12
|
instructions: {},
|
|
13
|
+
ignoreKeys: [],
|
|
14
|
+
exclude: [],
|
|
15
|
+
maxLength: {},
|
|
13
16
|
model: 'gpt-5.4-mini',
|
|
14
17
|
baseUrl: 'https://api.openai.com/v1',
|
|
15
18
|
apiKeyEnv: 'OPENAI_API_KEY',
|
|
@@ -65,6 +68,16 @@ export function resolveConfig(raw, root) {
|
|
|
65
68
|
!Object.values(config.glossary).every(isStringRecord)) {
|
|
66
69
|
fail('"glossary" must map languages to { "source term": "required rendering" } objects.');
|
|
67
70
|
}
|
|
71
|
+
for (const key of ['ignoreKeys', 'exclude']) {
|
|
72
|
+
if (!Array.isArray(config[key]) || !config[key].every(p => typeof p === 'string' && p !== '')) {
|
|
73
|
+
fail(`"${key}" must be a list of patterns.`);
|
|
74
|
+
}
|
|
75
|
+
}
|
|
76
|
+
if (!config.maxLength ||
|
|
77
|
+
typeof config.maxLength !== 'object' ||
|
|
78
|
+
!Object.values(config.maxLength).every(n => Number.isInteger(n) && n > 0)) {
|
|
79
|
+
fail('"maxLength" must map key patterns to positive whole numbers, e.g. { "**.meta.title": 60 }.');
|
|
80
|
+
}
|
|
68
81
|
if (!isStringRecord(config.termNotes))
|
|
69
82
|
fail('"termNotes" must map terms to explanations.');
|
|
70
83
|
if (!isStringRecord(config.instructions))
|
package/dist/files.d.ts
CHANGED
|
@@ -10,6 +10,18 @@ export interface LocaleFile {
|
|
|
10
10
|
* Supported: {lang} (once or more), * within one path segment, and **\/ for any depth.
|
|
11
11
|
*/
|
|
12
12
|
export declare function findSourceFiles(root: string, pattern: string, sourceLanguage: string): LocaleFile[];
|
|
13
|
+
/**
|
|
14
|
+
* Key pattern -> RegExp. "*" matches within one segment, "**" any number of segments. A
|
|
15
|
+
* pattern without a dot matches the last segment anywhere ("id" matches "steps.2.id").
|
|
16
|
+
*/
|
|
17
|
+
export declare function keyPattern(pattern: string): RegExp;
|
|
18
|
+
/** Path pattern -> RegExp ("*" within a folder, "**" across folders, {lang} as written). */
|
|
19
|
+
export declare function pathPattern(pattern: string): RegExp;
|
|
20
|
+
/**
|
|
21
|
+
* Values that are not text and stay as they are in every language: URLs, email addresses,
|
|
22
|
+
* file paths and plain numbers or codes without spaces.
|
|
23
|
+
*/
|
|
24
|
+
export declare function isLiteralValue(value: string): boolean;
|
|
13
25
|
export type JsonValue = string | number | boolean | null | JsonValue[] | {
|
|
14
26
|
[key: string]: JsonValue;
|
|
15
27
|
};
|
|
@@ -48,5 +60,14 @@ export interface JsonFormat {
|
|
|
48
60
|
/** Indentation and final newline of an existing JSON text (2 spaces by default). */
|
|
49
61
|
export declare function detectFormat(text: string | null): JsonFormat;
|
|
50
62
|
export declare function serialize(value: JsonValue, format: JsonFormat): string;
|
|
63
|
+
/**
|
|
64
|
+
* Parses a locale file. A .txt file (fastlane metadata: description.txt, keywords.txt) is one
|
|
65
|
+
* string, keyed by its file name, so maxLength patterns like "keywords" apply to it.
|
|
66
|
+
*/
|
|
67
|
+
export declare function parseDoc(rel: string, text: string): JsonValue;
|
|
68
|
+
/** Flutter ARB metadata ("@@locale", "@title": { description, placeholders }): not text. */
|
|
69
|
+
export declare const isArbMetadata: (rel: string, key: string) => boolean;
|
|
70
|
+
/** Text of a locale file, or null when a .txt file has no value to write. */
|
|
71
|
+
export declare function serializeDoc(rel: string, doc: JsonValue, format: JsonFormat): string | null;
|
|
51
72
|
export declare function readText(file: string): string | null;
|
|
52
73
|
export declare function writeText(file: string, text: string): void;
|