localewarden 0.1.1 → 0.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -11,7 +11,7 @@ npx localewarden # translate new and changed strings
11
11
  npx localewarden check # quality check, no API calls (use it in CI)
12
12
  ```
13
13
 
14
- Works with i18next, react-intl / FormatJS, vue-i18n, next-intl, ngx-translate and any other setup that keeps strings in JSON files. Uses any OpenAI-compatible API (OpenAI, OpenRouter, a local Ollama, ...).
14
+ Works with i18next, react-intl / FormatJS, vue-i18n, next-intl, ngx-translate and any other setup that keeps strings in JSON files, with Flutter (`.arb` files) and with fastlane's App Store / Play Store metadata (`.txt` files). Uses any OpenAI-compatible API (OpenAI, OpenRouter, a local Ollama, ...).
15
15
 
16
16
  ## Why
17
17
 
@@ -22,15 +22,22 @@ Translating locale files with a language model is easy once. Keeping 20 language
22
22
  - **Models are inconsistent across batches.** French screens mix "tu" and "vous", Spanish copies English Title Case ("Configure Su Cuenta"), and Polish or Russian address every user as a man.
23
23
  - **Broken output ships silently.** A translated placeholder (`{heures}` instead of `{hours}`) shows raw braces in your app. A dropped `</strong>` breaks the layout. Stray Cyrillic letters end up in a Danish sentence.
24
24
 
25
- localewarden grew out of the translation pipeline of a production app that ships in 38 languages. Every rule and check in it exists because one of these failures happened in real output.
25
+ localewarden grew out of the translation pipeline of a production app that ships in 38 languages. Every rule and check in it exists because one of these failures happened in real output. The checks are tuned against that app's real texts (UI, website, long-form learning content and store listings, about 165 MB) to report problems without flooding you with false alarms. On those texts they still find things that slipped through earlier pipelines: sections cut off after the English grew, stray letters from other alphabets, a trial notice left in English.
26
+
27
+ ## How it compares
28
+
29
+ - **Translation platforms** (Crowdin, Lokalise, Phrase, Weblate) are hosted services with editors, translator workflows and review for teams. localewarden is a small CLI that runs in your repository and CI, with no account and no server. If you have professional translators, a platform fits better. If a model translates and people only fix the odd string, this is the lighter setup.
30
+ - **"Translate my JSON with GPT" scripts** usually send every string on every run and overwrite whatever is there. localewarden keeps state, so it only sends what changed, keeps human fixes, and checks the output.
31
+ - **Editor extensions** (such as i18n Ally) help you write and look up keys while coding. localewarden is about filling and maintaining 10 to 40 languages afterwards. The two work well together.
26
32
 
27
33
  ## What it does
28
34
 
29
35
  - **Translates only what changed.** It remembers a hash of each source string per language. New strings are translated. Changed strings are *revised*: the model gets the existing translation and changes only what the source change requires. Removed strings are deleted from every language.
30
36
  - **Protects hand edits.** If someone edited a translation, localewarden detects it, keeps it, and lists it for review. If the source of a hand-edited string changes later, the string is flagged instead of overwritten.
31
- - **Checks every result before writing it.** Broken placeholders, injected HTML or scripts, foreign alphabets, changed links, broken HTML and echoed source text are rejected (retried once, then left for the next run). Softer problems are retried and reported.
37
+ - **Checks every result before writing it.** Broken placeholders, injected HTML or scripts, foreign alphabets, changed links, broken HTML and echoed source text are rejected (retried once, then left for the next run). Softer problems (too long, content missing, words left in English) are retried and reported.
32
38
  - **Consistent style per language.** It enforces formal or informal address per language (`du`/`Sie`, `tu`/`vous`, `ты`/`вы` and 16 more), uses sentence case where the language does, avoids gendered forms for "you", and applies local typography (French spacing, `92 %` in German, CJK quotation marks).
33
39
  - **Plural forms per language.** For i18next-style keys (`item_one`, `item_other`) it adds the forms a language needs but English lacks, such as Polish `_few` and `_many` or Arabic `_zero`, `_two`, `_few` and `_many` (CLDR plural rules).
40
+ - **Data files and store listings.** Fields like `id`, `type` or `image` are copied instead of translated (`ignoreKeys`), and so are URLs, email addresses and file paths. Length limits per key (`maxLength`) are passed to the model and checked: App Store names, SEO titles, buttons.
34
41
  - **Glossary and protected names.** You choose fixed renderings ("Privacy Policy" -> "Politique de confidentialité") and names that must never be translated. The check accepts grammatical case endings.
35
42
  - **Quality check for CI.** `localewarden check` runs all checks without any API calls and exits non-zero on errors.
36
43
  - **Targeted repair.** `--fix-flagged` asks the model to fix only what the check flagged. The fix is accepted only if the problem is gone and little else changed.
@@ -79,7 +86,7 @@ es Mantén vivas tus plantas sin tener que pensar en ello
79
86
  ja 何も考えなくても、植物を元気に保てます
80
87
  ```
81
88
 
82
- German uses "du" and French "vous", as configured. Spanish and French use sentence case, not the English Title Case. French has its space before "!". The hedge "tend to" survived, and so did the placeholders, the link and the brand name. The full example is in [`examples/basic`](examples/basic).
89
+ German uses "du" and French "vous", as configured. Spanish and French use sentence case, not the English Title Case. French has its space before "!". The hedge "tend to" survived, and so did the placeholders, the link and the brand name. The full example is in [`examples/basic`](examples/basic). There are also examples for [Flutter ARB files](examples/flutter) and [App Store / Play Store texts with fastlane](examples/fastlane).
83
90
 
84
91
  ## Quick start
85
92
 
@@ -135,24 +142,28 @@ The same checks run in two places. Right after each model answer, a failed hard
135
142
  | --- | --- | --- |
136
143
  | `placeholder` | `{name}`, `{{count}}`, `%s`, `%1$d`, `%{x}`, `${x}`, `<0></0>` renamed, translated, added or dropped. ICU `plural`/`select` arguments are compared, while plural categories may differ per language. | error |
137
144
  | `unsafe` | HTML tags, attributes, event handlers or `javascript:`/`data:` URLs that the source does not have. Translations are often rendered as raw HTML, so this would be a script injection. | error |
138
- | `script` | Letters from an alphabet the language does not use ("刺激" in German), or a word that mixes Latin with Cyrillic/Greek lookalikes ("Вarda") | error |
145
+ | `script` | Letters from an alphabet the language does not use ("刺激" in German), a word that mixes Latin with Cyrillic/Greek lookalikes ("Вarda"), or Simplified characters in Traditional Chinese (`zh-TW`) and the reverse | error |
139
146
  | `markup` | Changed link targets, different number of tags, unclosed or misnested tags, dropped list items | warning (broken tags and changed links: never written) |
140
147
  | `years` | A year from the source missing or changed (citations, dates) | warning |
141
148
  | `formality` | The other form of address than configured, both forms in one string, or masculine-only forms for "you" | warning |
149
+ | `length` | Longer than the `maxLength` configured for the key | warning |
142
150
  | `titlecase` | English Title Case copied into a language that uses sentence case | warning |
143
151
  | `ampersand` | "&" in languages that write the word | warning |
144
152
  | `glossary` | A glossary rendering missing (case endings allowed), or a `doNotTranslate` name translated | warning |
145
153
  | `untranslated` | Identical to the source (prose of 3+ words; "OK" and names are fine) | warning |
146
- | `partial` | Source-language words left inside the translation, an untranslated bold lead-in, or a hedge that became certainty ("tend to" stated as fact) | warning |
154
+ | `partial` | Source-language words left inside the translation, an untranslated bold lead-in, a hedge that became certainty ("tend to" stated as fact), a translation much shorter than its source (content cut off, or the source grew after it was translated), or sibling options that got the same translation although the source differs ("Rarely" and "Occasionally" both "Selten") | warning |
147
155
 
148
156
  ```bash
149
157
  npx localewarden check # counts per language and check
150
158
  npx localewarden check -v # with examples
151
159
  npx localewarden check --strict # exit 1 on warnings too
152
160
  npx localewarden check --json # for scripts
161
+ npx localewarden check --fix # repair placeholders with one possible fix, no API calls
153
162
  ```
154
163
 
155
- Approved hand edits are skipped, except for errors (placeholder, unsafe, script), which break the app either way.
164
+ `--fix` repairs a translated placeholder when the source has exactly one and the translation renamed it (`{stunden}` back to `{hours}`). The file is edited in place, so its formatting stays as it is. Anything less certain is left for `--fix-flagged` or a person.
165
+
166
+ Approved strings are skipped, except for errors (placeholder, unsafe, script), which break the app either way. You can approve any string, not only hand edits: `npx localewarden review --approve de:home.title` tells the check that a person looked at it (for example a pun on a brand name that is correct without the name), and runs leave it alone.
156
167
 
157
168
  ### In CI
158
169
 
@@ -240,6 +251,9 @@ npx localewarden review --release de:home.title # hand it back: next run revis
240
251
  | `glossary` | `{}` | `{"fr": {"Terms of Service": "Conditions d'utilisation"}}` |
241
252
  | `termNotes` | `{}` | Meanings of ambiguous terms, sent only with strings that contain them: `{"snooze": "postpone a reminder"}` |
242
253
  | `instructions` | `{}` | Extra instructions per language, `"*"` for all: `{"es": "Use neutral Latin American Spanish."}` |
254
+ | `ignoreKeys` | `[]` | Keys that are not text, copied from the source: `["id", "type", "**.sources.*"]`. `*` matches within a key segment, `**` across segments; a pattern without a dot matches the last segment anywhere. URLs, emails, file paths and numbers are always copied. |
255
+ | `exclude` | `[]` | Source files to skip: `["locales/{lang}/nav.json"]` |
256
+ | `maxLength` | `{}` | Character limits per key pattern: `{"**.meta.title": 60, "name": 30}`. The model is told the limit; longer results are retried once and reported by the `length` check. |
243
257
  | `placeholders` | built-in | Regular expressions (strings) that match your placeholders. Replaces the built-in list. |
244
258
  | `model` | `"gpt-5.4-mini"` | Any chat model your endpoint offers |
245
259
  | `baseUrl` | `"https://api.openai.com/v1"` | Any OpenAI-compatible endpoint |
@@ -251,6 +265,37 @@ npx localewarden review --release de:home.title # hand it back: next run revis
251
265
  | `batchSize` | `20` | Strings per request (smaller for scripts that need many tokens) |
252
266
  | `stateDir` | `".localewarden"` | Where state and the review list live |
253
267
 
268
+ ### App Store and Play Store listings (fastlane)
269
+
270
+ ```json
271
+ {
272
+ "sourceLanguage": "en-US",
273
+ "targetLanguages": ["de-DE", "fr-FR", "ja"],
274
+ "files": "fastlane/metadata/{lang}/*.txt",
275
+ "exclude": ["fastlane/metadata/{lang}/*_url.txt"],
276
+ "maxLength": { "name": 30, "subtitle": 30, "keywords": 100, "promotional_text": 170, "description": 4000 },
277
+ "termNotes": { "keywords": "a comma-separated keyword list for store search, not a sentence" }
278
+ }
279
+ ```
280
+
281
+ Each `.txt` file is one string, keyed by its file name, so the limits above apply to `name.txt`, `subtitle.txt` and so on.
282
+
283
+ ### Flutter (ARB)
284
+
285
+ ```json
286
+ { "files": "lib/l10n/app_{lang}.arb", "targetLanguages": ["de", "fr", "pt_BR"] }
287
+ ```
288
+
289
+ Metadata (`@@locale`, `@key` descriptions and placeholders) is copied, not translated, and `@@locale` is set to the target language. ICU plurals and selects keep their structure, and each language gets the plural categories it needs.
290
+
291
+ ### Data files
292
+
293
+ For content JSON with ids, types and links, list the non-text keys:
294
+
295
+ ```json
296
+ { "files": "content/**/*.{lang}.json", "ignoreKeys": ["id", "type", "category", "image", "**.sources.*"] }
297
+ ```
298
+
254
299
  ### Other providers
255
300
 
256
301
  ```json
@@ -268,7 +313,7 @@ Small local models make noticeably more mistakes. The checks catch the mechanica
268
313
  ```text
269
314
  localewarden [translate] --dry-run --lang de,fr --fix-flagged --retranslate-all
270
315
  --overwrite-manual --max-tokens <n> --verbose
271
- localewarden check --lang de,fr --verbose --limit <n> --strict --json
316
+ localewarden check --lang de,fr --verbose --limit <n> --strict --json --fix
272
317
  localewarden review --all --approve <sel>... --release <sel>...
273
318
  localewarden init
274
319
  Global: --config <path> --help --version
@@ -300,7 +345,7 @@ Translations are treated as untrusted: any markup the source does not have is bl
300
345
 
301
346
  ## Limitations
302
347
 
303
- - JSON only (nested objects, arrays, flat keys). YAML, PO, XLIFF and ARB are not supported yet.
348
+ - JSON (nested objects, arrays, flat keys), Flutter ARB and plain `.txt` files. YAML, PO and XLIFF are not supported yet.
304
349
  - The checks catch mechanical problems, not every wrong meaning. Have a native speaker look at important screens, then approve their edits with `review`.
305
350
  - Rules for form of address, gender and typography exist for the languages listed above. Other languages are translated with the general rules.
306
351
  - A run that is interrupted keeps everything written so far. Unwritten strings are picked up on the next run.
package/dist/checks.d.ts CHANGED
@@ -1,4 +1,5 @@
1
1
  import type { Config } from './config.js';
2
+ import { Scope } from './scope.js';
2
3
  /**
3
4
  * Deterministic quality checks. No API calls, so they can run in CI on every commit.
4
5
  *
@@ -11,14 +12,16 @@ import type { Config } from './config.js';
11
12
  * years a year from the source is missing or changed (citations, dates)
12
13
  * formality the other form of address than configured, or both mixed;
13
14
  * masculine-only forms for "you" when genderNeutral is on
15
+ * length longer than the maxLength configured for the key
14
16
  * titlecase English Title Case copied into a sentence-case language
15
17
  * ampersand "&" in a language that writes the word
16
18
  * glossary a glossary rendering or a doNotTranslate name is missing
17
19
  * untranslated identical to the source (prose of 3+ words)
18
20
  * partial source-language words left inside an otherwise translated string,
19
- * or a dropped hedge ("tend to" stated as certain)
21
+ * a dropped hedge ("tend to" stated as certain), or a translation much
22
+ * shorter than its source (content missing or cut off)
20
23
  */
21
- export type CheckName = 'placeholder' | 'unsafe' | 'script' | 'markup' | 'years' | 'formality' | 'titlecase' | 'ampersand' | 'glossary' | 'untranslated' | 'partial';
24
+ export type CheckName = 'placeholder' | 'unsafe' | 'script' | 'markup' | 'years' | 'formality' | 'length' | 'titlecase' | 'ampersand' | 'glossary' | 'untranslated' | 'partial';
22
25
  export declare const CHECKS: CheckName[];
23
26
  export declare const ERROR_CHECKS: Set<CheckName>;
24
27
  /** Checks a targeted repair (--fix-flagged) may try to fix. */
@@ -30,7 +33,14 @@ export interface Issue {
30
33
  export declare class Checker {
31
34
  readonly config: Config;
32
35
  readonly placeholderRe: RegExp;
36
+ readonly scope: Scope;
37
+ /** Words of termNotes and doNotTranslate: terms a translation may keep in the source language. */
38
+ readonly keptWords: Set<string>;
33
39
  constructor(config: Config);
40
+ /** The text without doNotTranslate names, which stay the same in every language. */
41
+ withoutNames(text: string): string;
42
+ /** "62 characters, limit 60", or null. */
43
+ tooLong(key: string, text: string): string | null;
34
44
  get englishSource(): boolean;
35
45
  placeholdersMatch(key: string, source: string, text: string): boolean;
36
46
  placeholderNote(source: string, text: string): string;
@@ -50,6 +60,8 @@ export declare class Checker {
50
60
  }
51
61
  /** Non-Latin scripts each language is written in. Latin is always allowed (names, codes). */
52
62
  export declare const NATIVE_SCRIPTS: Record<string, string[]>;
63
+ /** Simplified characters in Traditional Chinese text, or the reverse (2+ distinct ones). */
64
+ export declare function wrongChineseScript(lang: string, text: string): string | null;
53
65
  /** Why `text` contains letters that cannot belong to `lang`, or null. */
54
66
  export declare function foreignScript(lang: string, text: string): string | null;
55
67
  /**
@@ -82,3 +94,8 @@ export declare function isUnchangedProse(source: string, text: string, placehold
82
94
  * allowed) copied into the translation, or null.
83
95
  */
84
96
  export declare function sourceRun(source: string, text: string): string | null;
97
+ /**
98
+ * A translation with a fraction of the source's length lost content: cut off by the model, or
99
+ * the source grew after it was translated. Markup and placeholders are not counted.
100
+ */
101
+ export declare function muchShorter(lang: string, source: string, text: string): string | null;
package/dist/checks.js CHANGED
@@ -1,6 +1,7 @@
1
1
  import { placeholderRegExp, placeholderSignature, placeholdersMatch } from './placeholders.js';
2
2
  import { GENDERED_FORMS, NO_AMPERSAND_LANGUAGES, SENTENCE_CASE_LANGUAGES, formalityRule, registerFor, withoutQuotedSpeech, } from './style.js';
3
- import { baseLanguage, escapeRegExp } from './util.js';
3
+ import { Scope } from './scope.js';
4
+ import { baseLanguage, escapeRegExp, isTraditionalChinese } from './util.js';
4
5
  export const CHECKS = [
5
6
  'placeholder',
6
7
  'unsafe',
@@ -8,6 +9,7 @@ export const CHECKS = [
8
9
  'markup',
9
10
  'years',
10
11
  'formality',
12
+ 'length',
11
13
  'titlecase',
12
14
  'ampersand',
13
15
  'glossary',
@@ -21,6 +23,7 @@ export const FIXABLE_CHECKS = new Set([
21
23
  'markup',
22
24
  'years',
23
25
  'formality',
26
+ 'length',
24
27
  'titlecase',
25
28
  'ampersand',
26
29
  'glossary',
@@ -29,9 +32,24 @@ export const FIXABLE_CHECKS = new Set([
29
32
  export class Checker {
30
33
  config;
31
34
  placeholderRe;
35
+ scope;
36
+ /** Words of termNotes and doNotTranslate: terms a translation may keep in the source language. */
37
+ keptWords;
32
38
  constructor(config) {
33
39
  this.config = config;
34
40
  this.placeholderRe = placeholderRegExp(config.placeholders);
41
+ this.scope = new Scope(config);
42
+ this.keptWords = new Set([...Object.keys(config.termNotes), ...config.doNotTranslate].flatMap(term => term.toLowerCase().split(/[^\p{L}]+/u)).filter(Boolean));
43
+ }
44
+ /** The text without doNotTranslate names, which stay the same in every language. */
45
+ withoutNames(text) {
46
+ return this.config.doNotTranslate.reduce((value, name) => value.split(name).join(' '), text);
47
+ }
48
+ /** "62 characters, limit 60", or null. */
49
+ tooLong(key, text) {
50
+ const max = this.scope.maxLength(key);
51
+ const length = [...text].length;
52
+ return max !== undefined && length > max ? `${length} characters, limit ${max}` : null;
35
53
  }
36
54
  get englishSource() {
37
55
  return baseLanguage(this.config.sourceLanguage) === 'en';
@@ -88,16 +106,22 @@ export class Checker {
88
106
  this.config.sentenceCase &&
89
107
  SENTENCE_CASE_LANGUAGES.has(base) &&
90
108
  isEnglishTitleCase(source)) {
91
- const capitals = midCapitals(text, source, this.config.doNotTranslate);
109
+ const capitals = midCapitals(text, source, this.config.doNotTranslate, lang);
92
110
  if (capitals.length >= (source.trim().split(/\s+/).length <= 3 ? 1 : 2)) {
93
111
  add('titlecase', `capitalised: ${capitals.join(' ')}`);
94
112
  }
95
113
  }
114
+ const long = this.tooLong(key, text);
115
+ if (long)
116
+ add('length', long);
96
117
  const ampersands = (value) => value.split(' & ').length - 1;
97
118
  if (NO_AMPERSAND_LANGUAGES.has(base) && ampersands(text) > ampersands(source))
98
119
  add('ampersand');
99
- if (isUnchangedProse(source, text, this.placeholderRe))
120
+ if (isUnchangedProse(this.withoutNames(source), this.withoutNames(text), this.placeholderRe))
100
121
  add('untranslated');
122
+ const short = muchShorter(lang, source, text);
123
+ if (short)
124
+ add('partial', short);
101
125
  if (this.englishSource && text !== source) {
102
126
  const copied = sourceRun(source, text);
103
127
  if (copied)
@@ -105,7 +129,7 @@ export class Checker {
105
129
  const lead = !copied ? sourceBoldLeadIn(source, text, this.config.doNotTranslate) : null;
106
130
  if (lead)
107
131
  add('partial', `bold lead-in still in the source language: "${lead}"`);
108
- const mixed = !copied && !lead ? englishInNativeScript(lang, source, text, this.placeholderRe) : null;
132
+ const mixed = !copied && !lead ? englishInNativeScript(lang, source, text, this.placeholderRe, this.keptWords) : null;
109
133
  if (mixed)
110
134
  add('partial', mixed);
111
135
  if (droppedHedge(lang, source, text)) {
@@ -127,7 +151,7 @@ export class Checker {
127
151
  const placeholders = this.placeholdersMatch(key, source, text) ? null : `placeholder mismatch: ${this.placeholderNote(source, text)}`;
128
152
  const foreign = foreignScript(this.config.sourceLanguage, source) ? null : foreignScript(lang, text);
129
153
  const links = hrefSignature(source) !== hrefSignature(text) ? `links changed: [${hrefSignature(source)}] -> [${hrefSignature(text)}]` : null;
130
- const echoed = isUnchangedProse(source, text, this.placeholderRe) ? 'returned the source text unchanged' : null;
154
+ const echoed = isUnchangedProse(this.withoutNames(source), this.withoutNames(text), this.placeholderRe) ? 'returned the source text unchanged' : null;
131
155
  const broken = brokenMarkup(source) ? null : brokenMarkup(text);
132
156
  const boldMarkers = (value) => (value.match(/\*\*/g) ?? []).length % 2;
133
157
  const brokenBold = boldMarkers(text) === 1 && boldMarkers(source) === 0 ? 'unbalanced ** markers' : null;
@@ -135,7 +159,7 @@ export class Checker {
135
159
  const hard = placeholders ?? unsafeAdditions(source, text) ?? foreign ?? links ?? echoed ?? broken ?? brokenBold ?? droppedBullets(source, text) ?? leaked;
136
160
  const tags = tagCount(source) !== tagCount(text) ? `${tagCount(source)} tags in the source, got ${tagCount(text)}` : null;
137
161
  const copied = this.englishSource && text !== source ? sourceRun(source, text) : null;
138
- return { hard, soft: hard ?? tags ?? yearDifference(source, text) ?? (copied ? `source text left in: "${copied}"` : null) };
162
+ return { hard, soft: hard ?? this.tooLong(key, text) ?? muchShorter(lang, source, text) ?? tags ?? yearDifference(source, text) ?? (copied ? `source text left in: "${copied}"` : null) };
139
163
  }
140
164
  }
141
165
  // ---------------------------------------------------------------------------------------
@@ -188,8 +212,25 @@ const LATIN_LANGUAGES = new Set([
188
212
  ]);
189
213
  // Latin glued to a lookalike alphabet inside one word ("Вarda": Cyrillic В + Latin arda).
190
214
  const HOMOGLYPH_WORD = /(?=\p{L}*\p{Script=Latin})(?=\p{L}*[\p{Script=Cyrillic}\p{Script=Greek}])\p{L}+/u;
215
+ // Common characters that exist in only one of the two Chinese scripts (pairs at the same index).
216
+ const SIMPLIFIED_ONLY = '们这说时会来对个为发过还让现实动门问间题体关点应开东头书长见认学页电话语读写买卖钱网设计习惯觉机帮爱车钟儿气无边进选择检样经验数据项结种类业务环节';
217
+ const TRADITIONAL_ONLY = '們這說時會來對個為發過還讓現實動門問間題體關點應開東頭書長見認學頁電話語讀寫買賣錢網設計習慣覺機幫愛車鐘兒氣無邊進選擇檢樣經驗數據項結種類業務環節';
218
+ /** Simplified characters in Traditional Chinese text, or the reverse (2+ distinct ones). */
219
+ export function wrongChineseScript(lang, text) {
220
+ if (baseLanguage(lang) !== 'zh')
221
+ return null;
222
+ const traditional = isTraditionalChinese(lang);
223
+ const wrong = traditional ? SIMPLIFIED_ONLY : TRADITIONAL_ONLY;
224
+ const found = [...new Set([...text].filter(ch => wrong.includes(ch)))];
225
+ return found.length >= 2
226
+ ? `${traditional ? 'Simplified' : 'Traditional'} Chinese characters in ${lang}: ${found.slice(0, 6).join('')}`
227
+ : null;
228
+ }
191
229
  /** Why `text` contains letters that cannot belong to `lang`, or null. */
192
230
  export function foreignScript(lang, text) {
231
+ const chinese = wrongChineseScript(lang, text);
232
+ if (chinese)
233
+ return chinese;
193
234
  const base = baseLanguage(lang);
194
235
  const native = NATIVE_SCRIPTS[base] ?? (LATIN_LANGUAGES.has(base) ? [] : null);
195
236
  if (native === null)
@@ -218,13 +259,18 @@ const unescapeEntities = (text) => text
218
259
  .replace(/&quot;|&#0*34;|&#x0*22;/gi, '"')
219
260
  .replace(/&#0*39;|&#x0*27;|&apos;/gi, "'")
220
261
  .replace(/&colon;|&#0*58;|&#x0*3a;/gi, ':');
221
- /** Tag names, in lower case. Numbered <0> tags (react-i18next) are placeholders, not HTML. */
222
- const tagNames = (html) => new Set([...html.matchAll(/<\/?([a-z][\w-]*)/gi)].map(m => m[1].toLowerCase()));
262
+ // HTML elements, so text in angle brackets ("<minutes>", "<your name>") is not taken for markup.
263
+ const HTML_ELEMENTS = new Set(('a abbr address area article aside audio b base bdi bdo blockquote body br button canvas caption cite code col colgroup data datalist dd del details dfn dialog div dl dt em embed fieldset figcaption figure footer form frame frameset h1 h2 h3 h4 h5 h6 head header hr html i iframe img input ins kbd label legend li link main map mark math meta meter nav noscript object ol optgroup option output p param picture pre progress q rp rt ruby s samp script section select slot small source span strong style sub summary sup svg table tbody td template textarea tfoot th thead time title tr track u ul var video wbr animate foreignobject use image set').split(' '));
264
+ /** HTML element names, in lower case. Numbered <0> tags (react-i18next) are placeholders. */
265
+ const tagNames = (html) => new Set([...html.matchAll(/<\/?([a-z][\w-]*)/gi)].map(m => m[1].toLowerCase()).filter(tag => HTML_ELEMENTS.has(tag)));
223
266
  const TEXT_ATTRIBUTES = new Set(['title', 'alt', 'aria-label', 'aria-description', 'placeholder']);
224
267
  /** Every attribute as "name=value" (value without quotes), in lower case. */
225
268
  const attributes = (html) => {
226
269
  const found = new Set();
227
- for (const [, inner] of html.matchAll(/<[a-z][\w-]*\s([^>]*)>?/gi)) {
270
+ for (const [, tag, inner] of html.matchAll(/<([a-z][\w-]*)\s([^>]*)>?/gi)) {
271
+ // "<dakika 20)" (sw: under 20 minutes) is text, not a tag.
272
+ if (!HTML_ELEMENTS.has(tag.toLowerCase()))
273
+ continue;
228
274
  for (const [, name, v1, v2, v3] of inner.matchAll(/([^\s"'=<>\/]+)(?:\s*=\s*(?:"([^"]*)"|'([^']*)'|([^\s"'>]+)))?/g)) {
229
275
  const attr = name.toLowerCase();
230
276
  // Text attributes are translated along with the visible text; only their presence counts.
@@ -234,7 +280,10 @@ const attributes = (html) => {
234
280
  }
235
281
  return found;
236
282
  };
237
- const DANGEROUS_URL = /(?:javascript|vbscript|data)\s*:/gi;
283
+ // Script URLs in a link target or Markdown link. Plain "data:" in prose is a word (pt/it "date:").
284
+ const DANGEROUS_URL = /(?:(?:href|src|action|formaction|xlink:href|poster|background)\s*=\s*["']?|\]\()\s*(?:javascript|vbscript|data)\s*:/gi;
285
+ // Formatting a translator may add for emphasis or a title (no attributes): a markup warning, not a risk.
286
+ const HARMLESS_TAGS = new Set(['br', 'i', 'b', 'em', 'strong', 'u', 's', 'sub', 'sup', 'small', 'mark', 'q', 'cite']);
238
287
  /**
239
288
  * Markup the translation adds that the source does not have: a new tag type, a new or changed
240
289
  * attribute, an event handler or a script URL. Translations are often rendered as raw HTML
@@ -244,7 +293,7 @@ const DANGEROUS_URL = /(?:javascript|vbscript|data)\s*:/gi;
244
293
  export function unsafeAdditions(source, text) {
245
294
  const [src, out] = [unescapeEntities(source), unescapeEntities(text)];
246
295
  const srcTags = tagNames(src);
247
- const newTags = [...tagNames(out)].filter(tag => !srcTags.has(tag) && tag !== 'br');
296
+ const newTags = [...tagNames(out)].filter(tag => !srcTags.has(tag) && !HARMLESS_TAGS.has(tag));
248
297
  if (newTags.length > 0)
249
298
  return `HTML tag not in the source: <${newTags.join('>, <')}>`;
250
299
  const srcAttrs = attributes(src);
@@ -312,7 +361,8 @@ const asciiDigits = (text) => text.replace(/\p{Nd}/gu, ch => {
312
361
  export function yearDifference(source, text) {
313
362
  // Decades ("the 2020s") are written in words or with suffixes in many languages.
314
363
  const decades = new Set(asciiDigits(source).match(/(?<!\d)(?:19|20)\d0(?=['’]?s\b)/g) ?? []);
315
- const years = (value) => (asciiDigits(value).match(/(?<!\d)(?:19|20)\d\d(?!\d)/g) ?? []).filter(y => !decades.has(y)).sort().join(',');
364
+ // "1,900", "1.900" and "1 900" are numbers, not years: drop thousands separators first.
365
+ const years = (value) => (asciiDigits(value).replace(/(\d)[,.\u00a0\u202f ](?=\d{3}(?!\d))/g, '$1').match(/(?<!\d)(?:19|20)\d\d(?!\d)/g) ?? []).filter(y => !decades.has(y)).sort().join(',');
316
366
  const [a, b] = [years(source), years(text)];
317
367
  return a === b ? null : `source years [${a}], got [${b}]`;
318
368
  }
@@ -359,20 +409,25 @@ export function keptNameMissing(name, source, text) {
359
409
  // ---------------------------------------------------------------------------------------
360
410
  // Title case
361
411
  const COMMON_NAMES = new Set(['iOS', 'Android', 'Apple', 'Google', 'iPhone', 'iPad', 'Mac', 'Windows', 'Linux', 'AI', 'API', 'URL', 'PDF', 'FAQ', 'OK']);
362
- function midCapitals(text, source, names) {
412
+ // Polish capitalises "you" pronouns as a sign of respect ("W Twoim planie"): correct, not Title Case.
413
+ const POLISH_RESPECT = /^(?:Ty|Twój|Twoja|Twoje|Twojego|Twojej|Twoim|Twoją|Twoich|Twoimi|Ciebie|Cię|Tobie|Tobą|Wy|Wasz|Wasza|Wasze|Wam|Was|Wami)$/u;
414
+ function midCapitals(text, source, names, lang = '') {
363
415
  const allowed = new Set([...COMMON_NAMES, ...names.flatMap(name => name.split(/\s+/))]);
364
416
  const isName = (word) => allowed.has(word) ||
365
417
  new RegExp(`(?<!\\p{L})${escapeRegExp(word)}(?!\\p{L})`, 'u').test(source) ||
366
418
  [...allowed].some(name => name.length >= 4 && word.startsWith(name));
367
419
  return text
368
- .split(/[.!?:—–\n•|]+/)
420
+ // Commas and brackets start segments too: list items ("SMART: Specific, Measurable") and
421
+ // bracketed words ("Customize (Optional)") are capitalised in many languages.
422
+ .split(/[.!?:—–\n•|,;()]+/)
369
423
  .flatMap(segment => {
370
424
  const words = segment.trim().split(/\s+/);
371
425
  const first = words.findIndex(word => /\p{L}/u.test(word));
372
426
  return first === -1 ? [] : words.slice(first + 1);
373
427
  })
374
428
  .map(word => word.replace(/^[^\p{L}]+|[^\p{L}]+$/gu, ''))
375
- .filter(word => word.length >= 3 && /^\p{Lu}\p{Ll}/u.test(word) && !isName(word));
429
+ .filter(word => word.length >= 3 && /^\p{Lu}\p{Ll}/u.test(word) && !isName(word))
430
+ .filter(word => !(baseLanguage(lang) === 'pl' && POLISH_RESPECT.test(word)));
376
431
  }
377
432
  function isEnglishTitleCase(source) {
378
433
  const rest = source
@@ -430,26 +485,34 @@ function sourceBoldLeadIn(source, text, names) {
430
485
  for (const lead of bold(source)) {
431
486
  if (lead.length <= 15 || !/\p{Ll}{3}/u.test(lead) || names.some(name => lead.includes(name)))
432
487
  continue;
488
+ // Only capitalised words ("Apple App Store:", "Google Play"): a name, kept on purpose.
489
+ if (lead.split(/\s+/).every(word => !/\p{L}/u.test(word) || /^[^\p{L}]*\p{Lu}/u.test(word)))
490
+ continue;
433
491
  if (translated.has(lead))
434
492
  return lead;
435
493
  }
436
494
  return null;
437
495
  }
496
+ const ICU_HEADER = /\{\s*[\w.-]+\s*,\s*(?:plural|select|selectordinal)\s*,/g;
497
+ const ICU_HEADER_TEST = /\{\s*[\w.-]+\s*,\s*(?:plural|select|selectordinal)\s*,/;
438
498
  // Loanwords commonly written in Latin script inside non-Latin text.
439
499
  const LATIN_LOANWORDS = new Set(['email', 'online', 'offline', 'emoji', 'smartphone', 'podcast', 'podcasts', 'wifi', 'blog', 'login', 'like', 'likes']);
440
500
  /** Source-language words left inside a non-Latin-script translation ("Settings → Privacy"). */
441
- function englishInNativeScript(lang, source, text, placeholderRe) {
501
+ function englishInNativeScript(lang, source, text, placeholderRe, allowed) {
442
502
  const base = baseLanguage(lang);
443
503
  // Greek writes many anglicisms in Latin script; not checked.
444
504
  if (!NATIVE_SCRIPTS[base] || base === 'el' || base === 'sr' || text === source)
445
505
  return null;
446
506
  const strip = (value) => withoutUrls(value.replace(/[\w.+-]+@[\w.-]+/g, ' ').replace(/<[^>]+>/g, ' '))
507
+ // ICU syntax ("{count, plural, one {…} other {…}}") is code, not English text.
508
+ .replace(ICU_HEADER, ' ')
509
+ .replace(ICU_HEADER_TEST.test(value) ? /(?:^|[\s}])(?:zero|one|two|few|many|other|=\d+|[\w-]+)\s*(?=\{)/g : /$^/g, ' ')
447
510
  .replace(placeholderRe, ' ')
448
511
  .replace(/[((][^))]*[))]/g, ' ')
449
512
  .replace(/["“„«「『‘'][^"”“»」』’']*["”“»」』’']/g, ' ');
450
513
  const sourceWords = new Set(strip(source).match(/\b[a-z]{4,}\b/g) ?? []);
451
514
  const left = [
452
- ...new Set((strip(text).match(/(?<![\p{L}-])[a-z]{4,}(?![\p{L}])/gu) ?? []).filter(word => sourceWords.has(word) && !LATIN_LOANWORDS.has(word))),
515
+ ...new Set((strip(text).match(/(?<![\p{L}-])[a-z]{4,}(?![\p{L}])/gu) ?? []).filter(word => sourceWords.has(word) && !LATIN_LOANWORDS.has(word) && !allowed.has(word))),
453
516
  ];
454
517
  const menu = /\b[A-Z][a-z]+ → [A-Z][a-z]+/.exec(text);
455
518
  if (menu && source.includes(menu[0]))
@@ -478,9 +541,22 @@ const HEDGE_MARKERS = {
478
541
  cs: /tendenc|obvykle|často|zpravidla|většinou|bývá|sklon|snadno|častěji|mív/iu,
479
542
  sk: /tendenc|obvykle|často|zvyčajne|väčšinou|býva|sklon|ľahko|častejšie|zvyk|skôr/iu,
480
543
  el: /τείν|συχνά|συνήθως|τάση|συνήθ|εύκολα|συχνότερα/iu,
481
- fi: /taipu|usein|yleensä|tapaa|tavallisesti|tuppaa|tyypillisesti|taipumus|helposti|useimmiten|herkästi/iu,
544
+ fi: /taipu|tapana|usein|yleensä|tapaa|tavallisesti|tuppaa|tyypillisesti|taipumus|helposti|useimmiten|herkästi/iu,
482
545
  ca: /tendeix|tendència|sol|sovint|generalment|acostum|normalment|fàcilment|freqüent/iu,
483
546
  };
547
+ // Chinese, Japanese and Korean need far fewer characters than English; Thai has no spaces.
548
+ const COMPACT_SCRIPTS = new Set(['zh', 'ja', 'ko']);
549
+ /**
550
+ * A translation with a fraction of the source's length lost content: cut off by the model, or
551
+ * the source grew after it was translated. Markup and placeholders are not counted.
552
+ */
553
+ export function muchShorter(lang, source, text) {
554
+ const visible = (value) => [...value.replace(/<[^>]+>|\{\{?[^{}]*\}?\}/g, '').replace(/\s+/g, ' ').trim()].length;
555
+ const [a, b] = [visible(source), visible(text)];
556
+ // Real translations from English rarely drop below 60% (Chinese/Japanese/Korean: 25%).
557
+ const min = COMPACT_SCRIPTS.has(baseLanguage(lang)) ? 0.2 : 0.5;
558
+ return a >= 200 && b < a * min ? `translation much shorter than the source (${b} of ${a} characters): content missing?` : null;
559
+ }
484
560
  function droppedHedge(lang, source, text) {
485
561
  const markers = HEDGE_MARKERS[baseLanguage(lang)];
486
562
  return Boolean(markers) && /\btends? to\b/i.test(source) && !markers.test(text);
package/dist/cli.js CHANGED
@@ -5,7 +5,7 @@ import { ConfigError, CONFIG_FILE, loadConfig } from './config.js';
5
5
  import { CHECKS } from './checks.js';
6
6
  import { findSourceFiles } from './files.js';
7
7
  import { isReasoningModel } from './llm.js';
8
- import { checkProject, summaryTable } from './project.js';
8
+ import { checkProject, fixPlaceholders, summaryTable } from './project.js';
9
9
  import { listReview, updateReview } from './review.js';
10
10
  import { run } from './translate.js';
11
11
  const HELP = `localewarden - incremental AI translation for JSON locale files
@@ -31,6 +31,7 @@ Check options:
31
31
  --limit <n> findings shown per check with --verbose (default 20)
32
32
  --strict exit 1 on warnings too (default: only placeholder/script errors)
33
33
  --json print findings as JSON
34
+ --fix repair placeholders with exactly one possible fix ({heures} -> {hours})
34
35
 
35
36
  Review options:
36
37
  --all include approved entries
@@ -93,6 +94,8 @@ const LAYOUTS = [
93
94
  'src/assets/i18n/{lang}.json',
94
95
  'i18n/{lang}.json',
95
96
  'lang/{lang}.json',
97
+ 'lib/l10n/app_{lang}.arb',
98
+ 'lib/l10n/intl_{lang}.arb',
96
99
  ];
97
100
  function init() {
98
101
  const file = path.resolve(CONFIG_FILE);
@@ -180,7 +183,15 @@ async function translateCommand(args) {
180
183
  function checkCommand(args) {
181
184
  const config = loadConfig(args.values.get('--config')?.[0]);
182
185
  const languages = languagesArg(args) ?? config.targetLanguages;
183
- const findings = checkProject(config, languages);
186
+ let findings = checkProject(config, languages);
187
+ if (args.flags.has('--fix')) {
188
+ const fixed = fixPlaceholders(config, findings);
189
+ for (const f of fixed)
190
+ console.log(`fixed [${f.lang}] ${f.file} ${f.key}: ${f.text}`);
191
+ console.log(`${fixed.length} placeholder(s) repaired.\n`);
192
+ if (fixed.length > 0)
193
+ findings = checkProject(config, languages);
194
+ }
184
195
  if (args.flags.has('--json')) {
185
196
  console.log(JSON.stringify(findings, null, 2));
186
197
  }
@@ -200,7 +211,17 @@ function checkCommand(args) {
200
211
  console.log(` ... ${items.length - limit} more (--limit <n>)`);
201
212
  }
202
213
  }
203
- const errors = findings.filter(f => f.severity === 'error').length;
214
+ // Errors fail CI, so show them even without --verbose.
215
+ const errorItems = findings.filter(f => f.severity === 'error');
216
+ if (!args.flags.has('--verbose') && errorItems.length > 0) {
217
+ console.log('\nErrors:');
218
+ for (const f of errorItems.slice(0, 20)) {
219
+ console.log(` [${f.lang}] ${f.file} ${f.key} - ${f.check}${f.note ? `: ${f.note}` : ''}`);
220
+ }
221
+ if (errorItems.length > 20)
222
+ console.log(` ... ${errorItems.length - 20} more (--verbose)`);
223
+ }
224
+ const errors = errorItems.length;
204
225
  console.log(`\n${errors} error(s), ${findings.length - errors} warning(s).${findings.length && !args.flags.has('--verbose') ? ' Details: --verbose' : ''}`);
205
226
  }
206
227
  const errors = findings.some(f => f.severity === 'error');
package/dist/config.d.ts CHANGED
@@ -24,6 +24,16 @@ export interface Config {
24
24
  termNotes: Record<string, string>;
25
25
  /** Extra instructions per language; "*" applies to every language. */
26
26
  instructions: Record<string, string>;
27
+ /**
28
+ * Keys whose values are not text (ids, types, image paths). Copied from the source, never
29
+ * translated. "*" matches within one key segment, "**" across segments; a pattern without
30
+ * a dot matches the last segment anywhere ("id" matches "steps.2.id").
31
+ */
32
+ ignoreKeys: string[];
33
+ /** Source files to skip, as path patterns relative to the config ("locales/{lang}/nav.json"). */
34
+ exclude: string[];
35
+ /** Maximum characters per key pattern: { "**.meta.title": 60 }. Told to the model and checked. */
36
+ maxLength: Record<string, number>;
27
37
  /** Regular expressions (as strings) that match placeholders. Replaces the built-in list. */
28
38
  placeholders?: string[];
29
39
  model: string;
package/dist/config.js CHANGED
@@ -10,6 +10,9 @@ export const DEFAULTS = {
10
10
  glossary: {},
11
11
  termNotes: {},
12
12
  instructions: {},
13
+ ignoreKeys: [],
14
+ exclude: [],
15
+ maxLength: {},
13
16
  model: 'gpt-5.4-mini',
14
17
  baseUrl: 'https://api.openai.com/v1',
15
18
  apiKeyEnv: 'OPENAI_API_KEY',
@@ -65,6 +68,16 @@ export function resolveConfig(raw, root) {
65
68
  !Object.values(config.glossary).every(isStringRecord)) {
66
69
  fail('"glossary" must map languages to { "source term": "required rendering" } objects.');
67
70
  }
71
+ for (const key of ['ignoreKeys', 'exclude']) {
72
+ if (!Array.isArray(config[key]) || !config[key].every(p => typeof p === 'string' && p !== '')) {
73
+ fail(`"${key}" must be a list of patterns.`);
74
+ }
75
+ }
76
+ if (!config.maxLength ||
77
+ typeof config.maxLength !== 'object' ||
78
+ !Object.values(config.maxLength).every(n => Number.isInteger(n) && n > 0)) {
79
+ fail('"maxLength" must map key patterns to positive whole numbers, e.g. { "**.meta.title": 60 }.');
80
+ }
68
81
  if (!isStringRecord(config.termNotes))
69
82
  fail('"termNotes" must map terms to explanations.');
70
83
  if (!isStringRecord(config.instructions))
package/dist/files.d.ts CHANGED
@@ -10,6 +10,18 @@ export interface LocaleFile {
10
10
  * Supported: {lang} (once or more), * within one path segment, and **\/ for any depth.
11
11
  */
12
12
  export declare function findSourceFiles(root: string, pattern: string, sourceLanguage: string): LocaleFile[];
13
+ /**
14
+ * Key pattern -> RegExp. "*" matches within one segment, "**" any number of segments. A
15
+ * pattern without a dot matches the last segment anywhere ("id" matches "steps.2.id").
16
+ */
17
+ export declare function keyPattern(pattern: string): RegExp;
18
+ /** Path pattern -> RegExp ("*" within a folder, "**" across folders, {lang} as written). */
19
+ export declare function pathPattern(pattern: string): RegExp;
20
+ /**
21
+ * Values that are not text and stay as they are in every language: URLs, email addresses,
22
+ * file paths and plain numbers or codes without spaces.
23
+ */
24
+ export declare function isLiteralValue(value: string): boolean;
13
25
  export type JsonValue = string | number | boolean | null | JsonValue[] | {
14
26
  [key: string]: JsonValue;
15
27
  };
@@ -48,5 +60,14 @@ export interface JsonFormat {
48
60
  /** Indentation and final newline of an existing JSON text (2 spaces by default). */
49
61
  export declare function detectFormat(text: string | null): JsonFormat;
50
62
  export declare function serialize(value: JsonValue, format: JsonFormat): string;
63
+ /**
64
+ * Parses a locale file. A .txt file (fastlane metadata: description.txt, keywords.txt) is one
65
+ * string, keyed by its file name, so maxLength patterns like "keywords" apply to it.
66
+ */
67
+ export declare function parseDoc(rel: string, text: string): JsonValue;
68
+ /** Flutter ARB metadata ("@@locale", "@title": { description, placeholders }): not text. */
69
+ export declare const isArbMetadata: (rel: string, key: string) => boolean;
70
+ /** Text of a locale file, or null when a .txt file has no value to write. */
71
+ export declare function serializeDoc(rel: string, doc: JsonValue, format: JsonFormat): string | null;
51
72
  export declare function readText(file: string): string | null;
52
73
  export declare function writeText(file: string, text: string): void;