champollion 0.3.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (170) hide show
  1. package/LICENSE +133 -0
  2. package/README.md +387 -0
  3. package/bin/cli.js +278 -0
  4. package/index.js +135 -0
  5. package/lib/api-key.js +127 -0
  6. package/lib/autofix.js +432 -0
  7. package/lib/bridge/method_bridge.py +430 -0
  8. package/lib/card-source-resolution.mjs +284 -0
  9. package/lib/cards/cache.js +169 -0
  10. package/lib/cards/env.js +82 -0
  11. package/lib/cards/fetch-card-child.js +38 -0
  12. package/lib/cards/reader.js +435 -0
  13. package/lib/cards/refresh.js +111 -0
  14. package/lib/cards/remote.js +387 -0
  15. package/lib/cldf-export.mjs +540 -0
  16. package/lib/cldf-terms.mjs +62 -0
  17. package/lib/command-help.js +790 -0
  18. package/lib/commands/audit.js +49 -0
  19. package/lib/commands/card.js +454 -0
  20. package/lib/commands/doctor.js +559 -0
  21. package/lib/commands/fonts.js +489 -0
  22. package/lib/commands/help.js +91 -0
  23. package/lib/commands/init.js +1259 -0
  24. package/lib/commands/integrity.js +148 -0
  25. package/lib/commands/leaderboard.js +478 -0
  26. package/lib/commands/lint.js +30 -0
  27. package/lib/commands/models.js +177 -0
  28. package/lib/commands/plugin.js +103 -0
  29. package/lib/commands/provenance.js +45 -0
  30. package/lib/commands/recommend.js +75 -0
  31. package/lib/commands/register-corpus.js +678 -0
  32. package/lib/commands/repair-script.js +42 -0
  33. package/lib/commands/seal-corpus.js +355 -0
  34. package/lib/commands/seo.js +72 -0
  35. package/lib/commands/serve.js +147 -0
  36. package/lib/commands/status.js +265 -0
  37. package/lib/commands/submit.js +332 -0
  38. package/lib/commands/sync.js +89 -0
  39. package/lib/commands/tm.js +573 -0
  40. package/lib/commands/verify.js +39 -0
  41. package/lib/commands/watch.js +20 -0
  42. package/lib/commands/wrap.js +138 -0
  43. package/lib/commands/xliff.js +327 -0
  44. package/lib/commercial-eligibility.js +235 -0
  45. package/lib/concurrent.js +87 -0
  46. package/lib/config.js +523 -0
  47. package/lib/contamination-lane.js +76 -0
  48. package/lib/content-sync.js +731 -0
  49. package/lib/content.js +733 -0
  50. package/lib/corpus-registration.mjs +608 -0
  51. package/lib/cost-report.js +346 -0
  52. package/lib/diff.js +155 -0
  53. package/lib/docusaurus-sync.js +1256 -0
  54. package/lib/flatten.js +55 -0
  55. package/lib/format.js +954 -0
  56. package/lib/hash.js +159 -0
  57. package/lib/icu.js +473 -0
  58. package/lib/integrity.js +689 -0
  59. package/lib/license-gate.mjs +478 -0
  60. package/lib/license-identify.mjs +229 -0
  61. package/lib/lint.js +629 -0
  62. package/lib/method-manifest.js +60 -0
  63. package/lib/methods/anthropic.js +140 -0
  64. package/lib/methods/apertium.js +163 -0
  65. package/lib/methods/api.js +316 -0
  66. package/lib/methods/base.js +184 -0
  67. package/lib/methods/content-separator.js +45 -0
  68. package/lib/methods/deepl.js +426 -0
  69. package/lib/methods/direct-llm.js +586 -0
  70. package/lib/methods/external.js +332 -0
  71. package/lib/methods/fetch-with-retry.js +124 -0
  72. package/lib/methods/gemini.js +147 -0
  73. package/lib/methods/google-translate.js +402 -0
  74. package/lib/methods/http-utils.js +122 -0
  75. package/lib/methods/libretranslate.js +314 -0
  76. package/lib/methods/llm-coached.js +670 -0
  77. package/lib/methods/llm.js +592 -0
  78. package/lib/methods/local.js +76 -0
  79. package/lib/methods/microsoft-translator.js +331 -0
  80. package/lib/methods/openai.js +131 -0
  81. package/lib/methods/openrouter-client.js +327 -0
  82. package/lib/methods/openrouter-pricing.js +156 -0
  83. package/lib/methods/provider-env.js +115 -0
  84. package/lib/methods/provider-pricing.js +310 -0
  85. package/lib/methods/tilde.js +150 -0
  86. package/lib/methods/translated.js +229 -0
  87. package/lib/methods/translation-error.js +80 -0
  88. package/lib/models.js +258 -0
  89. package/lib/no-translate.js +233 -0
  90. package/lib/output.js +238 -0
  91. package/lib/pairs.js +547 -0
  92. package/lib/plugins.js +447 -0
  93. package/lib/provenance.js +323 -0
  94. package/lib/recommend.js +648 -0
  95. package/lib/registers.js +1185 -0
  96. package/lib/repair-script.js +266 -0
  97. package/lib/scripts.js +994 -0
  98. package/lib/seal.mjs +464 -0
  99. package/lib/sealed-qualifier.mjs +211 -0
  100. package/lib/security.js +59 -0
  101. package/lib/segment.js +369 -0
  102. package/lib/seo.js +275 -0
  103. package/lib/serve.js +854 -0
  104. package/lib/string-classify.js +85 -0
  105. package/lib/submit.mjs +344 -0
  106. package/lib/sync.js +969 -0
  107. package/lib/tags/bcp47.js +202 -0
  108. package/lib/tags/resolve.js +314 -0
  109. package/lib/terminology.js +111 -0
  110. package/lib/tm-seed.js +294 -0
  111. package/lib/tm.js +515 -0
  112. package/lib/translate-pair.js +197 -0
  113. package/lib/translate.js +203 -0
  114. package/lib/types.js +230 -0
  115. package/lib/validate.js +510 -0
  116. package/lib/verify.js +451 -0
  117. package/lib/watch.js +145 -0
  118. package/lib/xliff.js +184 -0
  119. package/package.json +93 -0
  120. package/shared/ATTRIBUTION.md +145 -0
  121. package/shared/CORPORA-CARDS.md +288 -0
  122. package/shared/DATA-SOVEREIGNTY.md +500 -0
  123. package/shared/LANGUAGE-CARD-FIELDS.md +532 -0
  124. package/shared/card-lint-baseline.json +3189 -0
  125. package/shared/cards-fallback.json +1 -0
  126. package/shared/catalogue/card-config.json +6091 -0
  127. package/shared/catalogue/external-results.json +3888 -0
  128. package/shared/catalogue/gender-guidance.json +1038 -0
  129. package/shared/catalogue/method-coverage.json +1751 -0
  130. package/shared/catalogue/metric-coverage.json +170 -0
  131. package/shared/catalogue/metric-reliability.json +1 -0
  132. package/shared/catalogue/register-presets.json +3180 -0
  133. package/shared/catalogue/vitality-scales.json +55 -0
  134. package/shared/cldr-index.json +1115 -0
  135. package/shared/code-bridge.json +253 -0
  136. package/shared/corpora-cards-v1-reference.md +281 -0
  137. package/shared/curated-dictionary-flags.json +35 -0
  138. package/shared/curated-endonyms.json +35 -0
  139. package/shared/curated-fsts.json +51 -0
  140. package/shared/curated-orthography-conventions.json +26 -0
  141. package/shared/curated-sil-resources.json +374 -0
  142. package/shared/curated-tools.json +41 -0
  143. package/shared/docent/corpus.json +11333 -0
  144. package/shared/docent/faq.en.json +564 -0
  145. package/shared/docent/register-blocks.json +60 -0
  146. package/shared/docent/system-prompt.md +144 -0
  147. package/shared/domain-taxonomy.json +35 -0
  148. package/shared/explainers/glossary.json +2975 -0
  149. package/shared/explainers/tc-features.json +20112 -0
  150. package/shared/explainers/term-watchlist.json +147 -0
  151. package/shared/human-services.json +59 -0
  152. package/shared/license-corrections.json +261 -0
  153. package/shared/license-evidence.json +13452 -0
  154. package/shared/licenses.json +6781 -0
  155. package/shared/method-registry.json +236 -0
  156. package/shared/metric-registry.json +620 -0
  157. package/shared/model-aliases.json +7 -0
  158. package/shared/schemas/champollion-plugin.schema.json +206 -0
  159. package/shared/schemas/corpora-card.schema.json +957 -0
  160. package/shared/schemas/domain-taxonomy.schema.json +64 -0
  161. package/shared/schemas/external-results.schema.json +314 -0
  162. package/shared/schemas/human-services.schema.json +90 -0
  163. package/shared/schemas/language-card.schema.json +1308 -0
  164. package/shared/schemas/licenses.schema.json +155 -0
  165. package/shared/schemas/method-card.schema.json +412 -0
  166. package/shared/schemas/method-registry.schema.json +85 -0
  167. package/shared/schemas/metric-registry.schema.json +96 -0
  168. package/shared/schemas/metric-reliability.schema.json +178 -0
  169. package/shared/schemas/model-aliases.schema.json +27 -0
  170. package/shared/schemas/source-snapshot.schema.json +96 -0
package/lib/xliff.js ADDED
@@ -0,0 +1,184 @@
1
+ /**
2
+ * XLIFF 1.2 export/import — zero-dependency.
3
+ *
4
+ * WHY THIS EXISTS:
5
+ * XLIFF (XML Localization Interchange File Format) is the universal
6
+ * exchange format between translation tools. Supporting it lets users:
7
+ * 1. Export translations for professional review in memoQ/SDL Trados/Phrase
8
+ * 2. Import reviewed translations back into their project
9
+ * 3. Feed translations into existing CAT tool workflows
10
+ * 4. Archive translations in a standard, tool-agnostic format
11
+ *
12
+ * WHY 1.2:
13
+ * XLIFF 2.0+ is cleaner but adoption is spotty. 1.2 is universally
14
+ * supported by every CAT tool, localization platform, and TM system
15
+ * in the industry. We generate 1.2 for maximum interoperability.
16
+ *
17
+ * WHAT THIS IS NOT:
18
+ * This is NOT a full XLIFF parser. We generate clean, well-formed XLIFF
19
+ * and parse our own output back. We DON'T handle arbitrary XLIFF from
20
+ * third-party tools with extensions, inline markup, etc. If a user
21
+ * imports XLIFF from memoQ, we extract <target> text and ignore the rest.
22
+ *
23
+ * ZERO DEPENDENCIES. Uses regex-based XML parsing because:
24
+ * 1. Our XLIFF output is predictable and well-formed
25
+ * 2. We only need to extract source/target text from <trans-unit> elements
26
+ * 3. No DOM manipulation needed
27
+ * 4. Keeps the zero-dep promise
28
+ */
29
+
30
+ // -----------------------------------------------------------------
31
+ // XLIFF Export — flat key→value maps → XLIFF 1.2 XML string
32
+ // -----------------------------------------------------------------
33
+
34
+ /**
35
+ * Export a source/target pair to XLIFF 1.2 format.
36
+ *
37
+ * @param {object} options
38
+ * @param {string} options.sourceLocale - Source language code (e.g., 'en')
39
+ * @param {string} options.targetLocale - Target language code (e.g., 'fr')
40
+ * @param {object} options.sourceFlat - Source key→value map
41
+ * @param {object} options.targetFlat - Target key→value map (can be partial)
42
+ * @param {string} [options.original] - Original filename/path for metadata
43
+ * @param {import('./no-translate.js').NoTranslateMatcher} [options.noTranslate] -
44
+ * Compiled no-translate matcher. Matching units are emitted with
45
+ * translate="no", state="final", and the source pre-filled as the target.
46
+ * @returns {string} XLIFF 1.2 XML string
47
+ */
48
+ function exportXLIFF({ sourceLocale, targetLocale, sourceFlat, targetFlat, original = 'locale.json', noTranslate = null }) {
49
+ const units = [];
50
+
51
+ for (const [key, sourceValue] of Object.entries(sourceFlat)) {
52
+ if (typeof sourceValue !== 'string') continue;
53
+
54
+ // Keys the project declared no-translate carry XLIFF 1.2's own
55
+ // translate="no" and are pre-filled with the source, already final. The
56
+ // unit still ships — dropping it would make the file an incomplete key
57
+ // inventory — but a CAT tool will lock it, so nobody spends time
58
+ // "translating" a URL only for the next sync to revert it.
59
+ const exempt = !!(noTranslate && noTranslate.matches(key, sourceValue));
60
+
61
+ const targetValue = targetFlat[key];
62
+ const hasTarget = typeof targetValue === 'string' && targetValue.length > 0;
63
+
64
+ // state: "translated" if we have a target, "new" if not
65
+ const state = exempt ? 'final' : hasTarget ? 'translated' : 'new';
66
+ const body = exempt ? sourceValue : (hasTarget ? targetValue : '');
67
+
68
+ units.push(
69
+ ` <trans-unit id="${escapeXML(key)}"${exempt ? ' translate="no"' : ''} xml:space="preserve">` +
70
+ `\n <source>${escapeXML(sourceValue)}</source>` +
71
+ `\n <target state="${state}">${escapeXML(body)}</target>` +
72
+ `\n </trans-unit>`
73
+ );
74
+ }
75
+
76
+ return `<?xml version="1.0" encoding="UTF-8"?>
77
+ <xliff version="1.2" xmlns="urn:oasis:names:tc:xliff:document:1.2">
78
+ <file original="${escapeXML(original)}" source-language="${escapeXML(sourceLocale)}" target-language="${escapeXML(targetLocale)}" datatype="plaintext">
79
+ <header>
80
+ <tool tool-id="champollion" tool-name="champollion" tool-version="1.0"/>
81
+ </header>
82
+ <body>
83
+ ${units.join('\n')}
84
+ </body>
85
+ </file>
86
+ </xliff>
87
+ `;
88
+ }
89
+
90
+ // -----------------------------------------------------------------
91
+ // XLIFF Import — XLIFF 1.2 XML string → flat key→value map
92
+ // -----------------------------------------------------------------
93
+
94
+ /**
95
+ * Import translations from an XLIFF 1.2 string.
96
+ *
97
+ * Extracts <trans-unit id="..."> <target>...</target> pairs.
98
+ * Only includes units where the target has actual content.
99
+ *
100
+ * @param {string} xliffString - XLIFF 1.2 XML content
101
+ * @returns {{ translations: object, metadata: { sourceLocale: string, targetLocale: string, original: string } }}
102
+ */
103
+ function importXLIFF(xliffString) {
104
+ const translations = {};
105
+
106
+ // Extract file-level metadata
107
+ const fileMatch = xliffString.match(
108
+ /<file[^>]*source-language="([^"]*)"[^>]*target-language="([^"]*)"[^>]*(?:original="([^"]*)")?/
109
+ );
110
+ const metadata = {
111
+ sourceLocale: fileMatch ? unescapeXML(fileMatch[1]) : '',
112
+ targetLocale: fileMatch ? unescapeXML(fileMatch[2]) : '',
113
+ original: fileMatch && fileMatch[3] ? unescapeXML(fileMatch[3]) : '',
114
+ };
115
+
116
+ // Also try original before target-language (attribute order varies)
117
+ if (!metadata.original) {
118
+ const origMatch = xliffString.match(/<file[^>]*original="([^"]*)"/);
119
+ if (origMatch) metadata.original = unescapeXML(origMatch[1]);
120
+ }
121
+
122
+ // Extract trans-unit elements
123
+ // Using a non-greedy match to handle each unit individually
124
+ const unitPattern = /<trans-unit[^>]*id="([^"]*)"[^>]*>[\s\S]*?<\/trans-unit>/g;
125
+ let match;
126
+
127
+ while ((match = unitPattern.exec(xliffString)) !== null) {
128
+ const unitId = unescapeXML(match[1]);
129
+ const unitContent = match[0];
130
+
131
+ // Extract target content (if present and non-empty)
132
+ const targetMatch = unitContent.match(/<target[^>]*>([\s\S]*?)<\/target>/);
133
+ if (targetMatch && targetMatch[1].length > 0) {
134
+ translations[unitId] = unescapeXML(targetMatch[1]);
135
+ }
136
+ }
137
+
138
+ return { translations, metadata };
139
+ }
140
+
141
+ // -----------------------------------------------------------------
142
+ // XML escaping
143
+ // -----------------------------------------------------------------
144
+
145
+ /**
146
+ * Escape special XML characters in text content.
147
+ *
148
+ * @param {string} str - Raw string
149
+ * @returns {string} XML-safe string
150
+ */
151
+ function escapeXML(str) {
152
+ return str
153
+ .replace(/&/g, '&amp;')
154
+ .replace(/</g, '&lt;')
155
+ .replace(/>/g, '&gt;')
156
+ .replace(/"/g, '&quot;')
157
+ .replace(/'/g, '&apos;');
158
+ }
159
+
160
+ /**
161
+ * Unescape XML entities back to raw characters.
162
+ *
163
+ * @param {string} str - XML-escaped string
164
+ * @returns {string} Raw string
165
+ */
166
+ function unescapeXML(str) {
167
+ return str
168
+ .replace(/&apos;/g, "'")
169
+ .replace(/&quot;/g, '"')
170
+ .replace(/&gt;/g, '>')
171
+ .replace(/&lt;/g, '<')
172
+ .replace(/&amp;/g, '&');
173
+ }
174
+
175
+ // -----------------------------------------------------------------
176
+ // Exports
177
+ // -----------------------------------------------------------------
178
+
179
+ export {
180
+ exportXLIFF,
181
+ importXLIFF,
182
+ escapeXML,
183
+ unescapeXML,
184
+ };
package/package.json ADDED
@@ -0,0 +1,93 @@
1
+ {
2
+ "name": "champollion",
3
+ "version": "0.3.3",
4
+ "description": "Research-grade translation engine for i18n projects. Pluggable methods, per-pair quality tiers, and deterministic script converters. Supports JSON (next-intl, i18next), TOML, and YAML (Hugo).",
5
+ "type": "module",
6
+ "bin": {
7
+ "champollion": "bin/cli.js"
8
+ },
9
+ "exports": "./index.js",
10
+ "main": "index.js",
11
+ "scripts": {
12
+ "test": "node --test test/*.test.js",
13
+ "test:unit": "node --test test/sync.test.js",
14
+ "test:redteam": "node --test test/redteam.test.js",
15
+ "test:format": "node --test test/format.test.js",
16
+ "test:content": "node --test test/content.test.js",
17
+ "test:content-sync": "node --test test/content-sync.test.js",
18
+ "test:hugo": "node --test test/hugo-e2e.test.js",
19
+ "test:lint": "node --test test/lint.test.js",
20
+ "test:pairs": "node --test test/pairs.test.js",
21
+ "test:methods": "node --test test/methods.test.js",
22
+ "test:init-help": "node --test test/init-help.test.js",
23
+ "test:sync-pairs": "node --test test/sync-pairs.test.js",
24
+ "build:fallback": "node scripts/build-cards-fallback.mjs",
25
+ "export:cldf": "node scripts/export-cldf.mjs",
26
+ "audit:licenses": "node scripts/audit-license-gate.mjs",
27
+ "sync:shared": "node scripts/sync-shared.mjs",
28
+ "prepack": "npm run sync:shared && node scripts/build-cards-fallback.mjs",
29
+ "prepublishOnly": "npm test && npm run test:pack",
30
+ "test:pack": "node --test test/e2e/*.test.js",
31
+ "sync": "node bin/cli.js sync",
32
+ "watch": "node bin/cli.js watch",
33
+ "audit": "node bin/cli.js audit",
34
+ "lint": "node bin/cli.js lint",
35
+ "init": "node bin/cli.js init"
36
+ },
37
+ "author": "Curtis Forbes",
38
+ "repository": {
39
+ "type": "git",
40
+ "url": "git+https://github.com/gamedaysuits/Champollion.git",
41
+ "directory": "cli"
42
+ },
43
+ "homepage": "https://github.com/gamedaysuits/Champollion#readme",
44
+ "bugs": {
45
+ "url": "https://github.com/gamedaysuits/Champollion/issues"
46
+ },
47
+ "keywords": [
48
+ "i18n",
49
+ "internationalization",
50
+ "translation",
51
+ "openrouter",
52
+ "openai",
53
+ "anthropic",
54
+ "gemini",
55
+ "deepl",
56
+ "localization",
57
+ "l10n",
58
+ "auto-translate",
59
+ "multilingual",
60
+ "machine-translation",
61
+ "language-pairs",
62
+ "seo",
63
+ "hugo",
64
+ "toml",
65
+ "yaml",
66
+ "next-intl",
67
+ "i18next",
68
+ "markdown"
69
+ ],
70
+ "license": "PolyForm-Noncommercial-1.0.0",
71
+ "engines": {
72
+ "node": ">=20.11.0"
73
+ },
74
+ "files": [
75
+ "index.js",
76
+ "bin/",
77
+ "lib/",
78
+ "shared/",
79
+ "!shared/language-cards/",
80
+ "!shared/corpora-cards/",
81
+ "!shared/curated-intros/crk.md"
82
+ ],
83
+ "devDependencies": {
84
+ "better-sqlite3": "^12.10.0",
85
+ "cldr-core": "48.2.0",
86
+ "cldr-localenames-full": "^48.2.0",
87
+ "cldr-misc-full": "^48.2.0",
88
+ "language-subtag-registry": "^0.4.2"
89
+ },
90
+ "dependencies": {
91
+ "@translated/lara": "^1.12.1"
92
+ }
93
+ }
@@ -0,0 +1,145 @@
1
+ # Data Source Attribution
2
+
3
+ > This document lists all third-party data sources used in the Champollion language card system. Every auto-generated card includes a `dataSources` array identifying which sources contributed data to that card.
4
+ >
5
+ > **Champollion is committed to full license compliance.** Before ingesting any new data source, verify its license and add it to this table.
6
+
7
+ Last updated: 2026-06-08
8
+
9
+ ---
10
+
11
+ ## Data Sources
12
+
13
+ | Source | License | URL | What We Use | Our Obligations |
14
+ |--------|---------|-----|-------------|-----------------|
15
+ | **Glottolog** | CC BY 4.0 (verify per release) | [glottolog.org](https://glottolog.org) | Language classification tree, Glottocodes, family hierarchy, MED documentation depth (via OLAC integration), AES endangerment, countries, dialect counts, documentation dates. Maintained by Max Planck Institute for Evolutionary Anthropology. | Attribution + link to source + indicate modifications. Derived tree data inherits CC BY 4.0. |
16
+ | **Grambank** | CC BY 4.0 | [grambank.clld.org](https://grambank.clld.org) | 195 typological features for 2,467 languages: morphological type, word order, grammatical gender (GB103–106), politeness (GB415), case, tone, evidentiality. Skirgård et al. (2023). | Attribution + link to source. |
17
+ | **WALS** (World Atlas of Language Structures) | CC BY 4.0 | [wals.info](https://wals.info) | Typological features: word order (81A–87A), morphological type (26A), nominal categories, numeral classifiers (55A), case count (49A), gender count (30A), inflectional strategy (26A), ordinal numerals (53A), numeral-noun order (89A). Dryer & Haspelmath (2013). | Attribution + link to source. |
18
+ | **PHOIBLE** | CC BY-SA 3.0 | [phoible.org](https://phoible.org) | Phoneme inventories, consonant/vowel counts, tonal status for 2,186 languages. Moran & McCloy (2019). | Attribution + ShareAlike on derived phonological data. |
19
+ | **ElCat** (Endangered Languages Catalogue) | CC BY 4.0 | [endangeredlanguages.com](http://endangeredlanguages.com) | Endangerment status, intergenerational transmission, domains of use, speaker trends, alternate names for 3,378+ endangered languages. Campbell et al. (2022). | Attribution + link to source. |
20
+ | **LinguaMeta** | Apache 2.0 | [github.com/AetherPrior/LinguaMeta](https://github.com/AetherPrior/LinguaMeta) | Writing systems (ISO 15924), speaker estimates, endangerment status, Wikidata descriptions for 7,511 languages. Google Research (2024). | Attribution. |
21
+ | **OLAC** (Open Language Archives Community) | Metadata freely accessible | [language-archives.org](http://www.language-archives.org) | MED (Most Extensive Description) classification consumed via Glottolog integration. Also harvested directly via OAI-PMH aggregator for archive presence data across 65+ language archives (7,500+ language codes). Simons & Bird (2003). | Attribution. |
22
+ | **CLDR / Unicode** | Unicode Terms of Use | [cldr.unicode.org](https://cldr.unicode.org) | Locale data, script metadata, plural rules, typography conventions, direction, script codes, native names. | Attribution. Data is freely usable under Unicode ToS. |
23
+ | **ISO 639-3** | Free to use | [iso639-3.sil.org](https://iso639-3.sil.org) | Language codes, scope, type, macrolanguage membership. SIL International (2024). | Attribution. |
24
+ | **IANA Language Subtag Registry** | Public Domain | [iana.org](https://www.iana.org/assignments/language-subtag-registry) | BCP 47 tags, ISO 639 codes, English language names. | None required — public domain. Attributed for good practice. |
25
+ | **ISO 15924** | Public Standard | [unicode.org/iso15924](https://unicode.org/iso15924/) | Script codes, script names, code-to-name mappings. | Reference the standard. |
26
+ | **UNESCO Atlas of the World's Languages in Danger** | CC BY-SA 3.0 IGO | [unesco.org](http://www.unesco.org/languages-atlas) | Endangerment status classifications. | Attribution + ShareAlike on derived endangerment data. |
27
+ | **OPUS** (Open Parallel Corpus) | Various (per-corpus) | [opus.nlpl.eu](http://opus.nlpl.eu) | Parallel corpus availability, sentence pair counts. Tiedemann (2012). | Per-corpus license compliance. |
28
+ | **Universal Dependencies** | Various (per-treebank, mostly CC) | [universaldependencies.org](https://universaldependencies.org) | Treebank availability, sentence/token counts for 150+ languages. de Marneffe et al. (2021). | Per-treebank license compliance. |
29
+ | **Tatoeba** | CC BY 2.0 FR | [tatoeba.org](https://tatoeba.org) | Sentence counts for 428 languages. | Attribution. |
30
+ | **Wikidata** | CC0 | [wikidata.org](https://www.wikidata.org) | Native language names (P1705), official names (P1448), speaker counts (P1098), writing systems (P282). Queried via SPARQL endpoint. | None — CC0 dedication. |
31
+ | **Wikipedia / MediaWiki / Incubator** | CC BY-SA 3.0 / GFDL | [mediawiki.org](https://www.mediawiki.org/wiki/API:Siteinfo), [incubator.wikimedia.org](https://incubator.wikimedia.org) | Article counts, active editors per language edition. Incubator test wiki page counts for 719 languages without their own Wikipedia. | Attribution + ShareAlike. |
32
+ | **WOLD** (World Loanword Database) | CC BY 4.0 | [wold.clld.org](https://wold.clld.org) | Contact language influences, borrowing proportions for 41 languages. Haspelmath & Tadmor (2009). | Attribution + link to source. |
33
+ | **APiCS** (Atlas of Pidgin and Creole Language Structures) | CC BY 4.0 | [apics-online.info](https://apics-online.info) | Structural features for 76 pidgin/creole languages. Michaelis et al. (2013). | Attribution + link to source. |
34
+ | **Lexibank** | CC BY 4.0 | [lexibank.clld.org](https://lexibank.clld.org) | Lexical dataset availability for 5,478 language varieties. | Attribution + link to source. |
35
+ | **PanLex** | CC0 (Public Domain Dedication) | [panlex.org](https://panlex.org) | Cross-lingual lexical data. | None — CC0 dedication. |
36
+ | **FLORES+** | CC BY-SA 4.0 | [github.com/openlanguagedata/flores](https://github.com/openlanguagedata/flores) | **Development/debugging only.** Available for smoke testing, pipeline validation, and cross-language coverage checks. NOT used for official leaderboard evaluation — frontier LLMs have likely been trained on this data, making scores unreliable for method comparison. | Attribution + ShareAlike if distributing derived datasets. |
37
+ | **GiellaLT** | varies per language repository (lang-crk: AGPL-3.0-or-later with a §7(b) attribution rider) | [github.com/giellalt](https://github.com/giellalt) | FST availability reference only. We reference which languages have FST tools available; we do NOT bundle or distribute any FST binaries. | No GPL obligation applies — we reference availability, not distribute code. |
38
+ | **Ethnologue** | Proprietary (limited free data) | [ethnologue.com](https://www.ethnologue.com) | Speaker count estimates (free tier only). | Respect free-tier usage limits. No bulk scraping. |
39
+ | **Numeralbank** (Chan's Numeral Systems) | CC BY 4.0 | [github.com/numeralbank/channumerals](https://github.com/numeralbank/channumerals) | Numeral systems for 4,000+ languages — counting base, highest documented numeral. Chan (2019). | Attribution + link to source. |
40
+ | **D-PLACE** (Ethnographic Atlas) | CC BY-NC 4.0 | [d-place.org](https://d-place.org) | Cultural and ecological variables for 1,291 societies: subsistence type, settlement patterns, political complexity. Kirby et al. (2016). | Attribution + NonCommercial. |
41
+ | **AUTOTYP** | CC BY 4.0 | [github.com/autotyp/autotyp-data](https://github.com/autotyp/autotyp-data) | Morphological type, locus of marking (head/dependent), for 750+ languages. Bickel et al. (2023). | Attribution + link to source. |
42
+ | **SegBo** (Segment Borrowing Database) | CC BY 4.0 | [github.com/segbo-db/segbo](https://github.com/segbo-db/segbo) | Borrowed phonemes and donor languages for 500+ languages. Grossman et al. (2020). | Attribution + link to source. |
43
+ | **ABVD** (Austronesian Basic Vocabulary Database) | CC BY 4.0 | [abvd.eva.mpg.de](https://abvd.eva.mpg.de/) | Swadesh-list vocabulary for 1,500+ Austronesian languages via Lexibank. Greenhill, Blust & Gray (2008). | Attribution + link to source. |
44
+ | **NorthEuraLex** | CC BY 4.0 | [northeuralex.org](https://northeuralex.org/) | Lexical data for 107 Northern Eurasian languages (~1016 concepts). Dellert et al. (2020). | Attribution + link to source. |
45
+ | **CLICS³** (Database of Cross-Linguistic Colexifications) | CC BY 4.0 | [clics.clld.org](https://clics.clld.org) | Colexification (shared word forms across concepts) for 2,279 language varieties from 30 datasets. Rzymski et al. (2020). | Attribution + link to source. |
46
+ | **Masakhane** | Various (per-benchmark) | [github.com/masakhane-io](https://github.com/masakhane-io) | African language NLP community — MT benchmarks for 37+ languages, NER for 12 languages. Nekoto et al. (2020). | Attribution + link to project. |
47
+ | **IndicNLP** | Various (per-dataset) | [indicnlp.ai4bharat.org](https://indicnlp.ai4bharat.org) | NLP resources catalog for 14 Indic languages. AI4Bharat. | Attribution + link to project. |
48
+ | **AmericasNLP** | CC BY (shared task data) | [americasnlp.github.io](https://turing.iimas.unam.mx/americasnlp/) | MT and NLP shared tasks for 13 Indigenous American languages. Mager et al. (2021). | Attribution + link to project. |
49
+ | **PARADISEC** (Pacific And Regional Archive for Digital Sources in Endangered Cultures) | Metadata: OAI-PMH open access | [catalog.paradisec.org.au](https://catalog.paradisec.org.au) | Archive holdings metadata for 1,000+ languages via OAI-PMH/OLAC protocol. Thieberger & Barwick (2012). | Attribution + link to archive. Metadata only — no content downloaded. |
50
+ | **ASJP** (Automated Similarity Judgement Program) | Open Access | [asjp.clld.org](https://asjp.clld.org) | 40-item basic vocabulary word lists (ASJPcode) for 7,000+ doculects (11,540 entries). Wichmann, Holman & Brown (2022). | Attribution + link to source. |
51
+ | **UniMorph** (Universal Morphology) | Open Source | [github.com/unimorph](https://github.com/unimorph) | Normalized morphological inflection paradigms for 188 languages derived from Wiktionary. Batsuren et al. (2022). | Attribution + link to source. |
52
+ | **Kaikki.org / Wiktextract** | CC BY-SA (via Wiktionary) | [kaikki.org](https://kaikki.org) | Machine-readable structured extraction of Wiktionary: definitions, POS, morphology, pronunciations for 457 languages. Ylonen (2022). | Attribution + ShareAlike (inherits from Wiktionary). |
53
+ | **eWAVE** (Electronic World Atlas of Varieties of English) | CC BY | [ewave-atlas.org](https://ewave-atlas.org) | Morphosyntactic features of 77 varieties of English (pidgins, creoles, L2 varieties). Kortmann & Lunkenheimer (2013). | Attribution + link to source. |
54
+ | **Open Multilingual Wordnet** (OMW) | Various open licenses | [omwn.org](https://omwn.org) | Linked wordnets (synsets, semantic relations) across 60 languages via Collaborative Interlingual Index. Bond & Foster (2013). | Per-wordnet license compliance. |
55
+ | **Kaipuleohone** (UH Digital Language Archive) | OAI-PMH metadata: open | [scholarspace.manoa.hawaii.edu](https://scholarspace.manoa.hawaii.edu) | Archive holdings metadata for 84+ Pacific and endangered languages via OAI-PMH. University of Hawaiʻi. | Attribution + link to archive. Metadata only. |
56
+ | **Rosetta Project** (Internet Archive) | Public Domain / CC | [archive.org/details/rosettaproject](https://archive.org/details/rosettaproject) | 6,500 items across 1,305 languages — word lists, grammars, descriptions. Long Now Foundation. | Attribution + link to archive. Metadata counts only. |
57
+ | **HuggingFace Datasets** | Various (per-dataset) | [huggingface.co/datasets](https://huggingface.co/datasets) | NLP dataset availability index — dataset counts per language. Metadata pointer only, no data downloaded. | Per-dataset license compliance. |
58
+ | **WACL** (World Atlas of Classifier Languages) | CC BY 4.0 | [github.com/cldf-datasets/wacl](https://github.com/cldf-datasets/wacl) | Classifier language typology for 3,338 languages. | Attribution + link to source. |
59
+ | **SAILS** (South American Indigenous Language Structures) | CC BY 4.0 | [github.com/cldf-datasets/sails](https://github.com/cldf-datasets/sails) | Typological features for 167 South American languages. Muysken et al. | Attribution + link to source. |
60
+ | **AfBo** (Affix Borrowing Database) | CC BY 4.0 | [github.com/cldf-datasets/afbo](https://github.com/cldf-datasets/afbo) | Affix borrowing patterns across 191 languages. Seifart (2024). | Attribution + link to source. |
61
+ | **DoReCo** (Language Documentation Reference Corpus) | CC BY 4.0 | [github.com/cldf-datasets/doreco](https://github.com/cldf-datasets/doreco) | Time-aligned annotated natural speech corpora for 47 languages. Paschen et al. (2020). | Attribution + link to source. |
62
+ | **TonoDB** | CC BY 4.0 | [github.com/cldf-datasets/tonodb](https://github.com/cldf-datasets/tonodb) | Tonal system classifications for 97 languages. | Attribution + link to source. |
63
+ | **SAPhon** (South American Phonological Inventories) | CC BY 4.0 | [github.com/cldf-datasets/saphon](https://github.com/cldf-datasets/saphon) | Phonological inventories for 363 South American languages. Michael et al. | Attribution + link to source. |
64
+ | **UraTyp** | CC BY 4.0 | [github.com/cldf-datasets/uratyp](https://github.com/cldf-datasets/uratyp) | Typological features for 39 Uralic languages. | Attribution + link to source. |
65
+ | **ACD** (Austronesian Comparative Dictionary) | CC BY | [lexibank/acd](https://github.com/lexibank/acd) | Comparative lexical data for 1,064 Austronesian languages. Blust & Trussel. | Attribution + link to source. |
66
+ | **DIACL** (Diachronic Atlas of Comparative Linguistics) | CC BY 4.0 | [lexibank/diacl](https://github.com/lexibank/diacl) | Diachronic comparative wordlists for 422 languages. Carling (2024). | Attribution + link to source. |
67
+ | **IDS** (Intercontinental Dictionary Series) | CC BY 4.0 | [github.com/intercontinental-dictionary-series](https://github.com/intercontinental-dictionary-series/ids) | Concept-aligned dictionaries across 319 languages, ~1,310 concepts. Key & Comrie (2023). | Attribution + link to source. |
68
+ | **Language Atlas of the Pacific Area** | CC BY 4.0 | [cldf-datasets/languageatlasofthepacificarea](https://github.com/cldf-datasets/languageatlasofthepacificarea) | 1,873 Pacific-area languages digitized from Wurm & Hattori (1981). ECAI. | Attribution + link to source. |
69
+ | **DatSemShift** (Database of Semantic Shifts) | CC BY 4.0 | [lexibank/datsemshift](https://github.com/lexibank/datsemshift) | Semantic shift patterns across 1,629 languages. Zalizniak et al. | Attribution + link to source. |
70
+ | **UCLA Phonetics Lab Archive** | CC BY 4.0 | [cldf-datasets/uclaphoneticslabarchive](https://github.com/cldf-datasets/uclaphoneticslabarchive) | Phonetic samples for 312 languages. Ladefoged (2009). | Attribution + link to source. |
71
+ | **Hunter-Gatherer Language Database** | CC BY 4.0 | [lexibank/huntergatherer](https://github.com/lexibank/huntergatherer) | Basic vocabulary for 324 hunter-gatherer languages. | Attribution + link to source. |
72
+ | **Pama-Nyungan Lexicon** | CC BY 4.0 | [lexibank/bowernpny](https://github.com/lexibank/bowernpny) | Comparative wordlists for 190 Pama-Nyungan languages. Bowern & Atkinson. | Attribution + link to source. |
73
+ | **Vanuatu Voices** | CC BY 4.0 | [lexibank/vanuatuvoices](https://github.com/lexibank/vanuatuvoices) | Sound-comparisons data for 236 Vanuatu languages. | Attribution + link to source. |
74
+ | **ValPaL** (Valency Patterns Leipzig) | CC BY 4.0 | [lexibank/valpal](https://github.com/lexibank/valpal) | Valency patterns across 36 languages. Hartmann et al. (2013). | Attribution + link to source. |
75
+ | **Sagart Sino-Tibetan** | CC BY 4.0 | [lexibank/sagartst](https://github.com/lexibank/sagartst) | Cognate-coded Sino-Tibetan comparative data for 50 languages. Sagart et al. | Attribution + link to source. |
76
+ | **UraLex** | CC BY 4.0 | [lexibank/uralex](https://github.com/lexibank/uralex) | Basic vocabulary for 43 Uralic languages. Syrjänen et al. | Attribution + link to source. |
77
+ | **CrossAndean** | CC BY 4.0 | [lexibank/crossandean](https://github.com/lexibank/crossandean) | Andean language comparative data for 50 languages. | Attribution + link to source. |
78
+ | **Lexibank Batch** (143 family-level datasets) | CC BY 4.0 | [github.com/lexibank](https://github.com/lexibank) | Comparative lexical data from 143 individual family-level datasets in the Lexibank ecosystem. Each dataset covers a specific language family (e.g., abvdoceanic, transnewguineaorg, grollemundbantu). Individual attributions listed per-card in `dataSources`. | Attribution per individual dataset. |
79
+ | **Language Atlas of the Pacific Area** | CC BY 4.0 | [cldf-datasets/languageatlasofthepacificarea](https://github.com/cldf-datasets/languageatlasofthepacificarea) | 1,873 Pacific-area languages digitized from Wurm & Hattori (1981). ECAI. | Attribution + link to source. |
80
+ | **DatSemShift** (Database of Semantic Shifts) | CC BY 4.0 | [lexibank/datsemshift](https://github.com/lexibank/datsemshift) | Semantic shift patterns across 1,629 languages. Zalizniak et al. | Attribution + link to source. |
81
+ | **UCLA Phonetics Lab Archive** | CC BY 4.0 | [cldf-datasets/uclaphoneticslabarchive](https://github.com/cldf-datasets/uclaphoneticslabarchive) | Phonetic samples for 312 languages. Ladefoged (2009). | Attribution + link to source. |
82
+ | **Hunter-Gatherer Language Database** | CC BY 4.0 | [lexibank/huntergatherer](https://github.com/lexibank/huntergatherer) | Basic vocabulary for 324 hunter-gatherer languages. | Attribution + link to source. |
83
+ | **AILLA** (Archive of Indigenous Languages of Latin America) | Public access | [ailla.utexas.org](https://ailla.utexas.org/) | 644 indigenous American languages with ISO codes, indigenous names, alternative names, and language families. University of Texas at Austin. | Attribution + link to source. NEH/NSF funded. |
84
+
85
+ ---
86
+
87
+ ## Restricted-Access Sources
88
+
89
+ ### ELAR (Endangered Languages Archive)
90
+
91
+ **Source**: [elararchive.org](https://www.elararchive.org/) / [elar.soas.ac.uk](https://elar.soas.ac.uk/)
92
+ **License**: Custom ELAR User Agreement (NOT open-access)
93
+ **Our usage**: **Metadata citation only** — we record that ELAR holds deposits for a language. We do NOT redistribute any ELAR materials.
94
+
95
+ > **⚠️ ELAR materials may be cited but NEVER distributed.**
96
+
97
+ The following terms govern all interaction with ELAR materials and must be observed by anyone working on this project:
98
+
99
+ 1. Materials are for **non-commercial private research or educational activity only**. Any other usage requires explicit written agreement from the Collection Steward and rights holders.
100
+ 2. **AI training/fine-tuning** use must comply with **CC BY-NC-SA**: non-commercial, attributed, share-alike on derivatives. Automated extraction that circumvents access controls is prohibited.
101
+ 3. **No transfer** of materials to any other person or equipment unless explicitly permitted in writing by the Collection Steward.
102
+ 4. Individual collections may state further conditions; these must be identified and observed.
103
+ 5. Any publication or dissemination using ELAR materials must **acknowledge ELAR** and, if required, the Collection Steward per the citation instructions on the collection's landing page. A copy of the publication should be sent to ELAR where possible.
104
+ 6. Materials must not be used in any way that causes **disparagement, disrespect, damage to reputation, or harm** to any individual or group.
105
+ 7. ELAR stores user data securely under EU law and may share access records with materials owners for management purposes.
106
+
107
+ **What this means for Champollion**: We may add `archivePresence.elar = true` to language cards and note deposit counts, but we must NEVER include actual ELAR content (recordings, transcriptions, texts) in the cards or any distributed package. Users who want ELAR materials must register at elararchive.org and agree to these terms themselves.
108
+
109
+ ---
110
+
111
+ ## ShareAlike Obligations
112
+
113
+ Some sources use **ShareAlike** licenses (CC BY-SA 3.0 IGO, CC BY-SA 4.0). This means:
114
+
115
+ - **Derived data** from these sources must be distributed under the same or a compatible license.
116
+ - **Our original metadata** (register presets, formality systems, LLM prompts, method support flags) is our own creative work and is NOT subject to ShareAlike obligations.
117
+ - The `language-tree.json` classification tree, if derived from Glottolog, inherits Glottolog's license.
118
+ - Endangerment status fields derived from UNESCO Atlas inherit CC BY-SA 3.0 IGO.
119
+
120
+ ## Traceability
121
+
122
+ Every auto-generated language card includes:
123
+
124
+ ```json
125
+ {
126
+ "dataSources": ["cldr-48.2", "glottolog-5.3", "wals-2024", "iana-2026-05"],
127
+ "cldrVersion": "48.2",
128
+ "glottologVersion": "5.3",
129
+ "generatedAt": "2026-06-01T00:00:00Z"
130
+ }
131
+ ```
132
+
133
+ This allows anyone to verify what data contributed to each card and reproduce the generation process.
134
+
135
+ ---
136
+
137
+ ## Adding New Sources
138
+
139
+ Before ingesting any new data source:
140
+
141
+ 1. Verify its license
142
+ 2. Add it to the table above
143
+ 3. Confirm our obligations are met
144
+ 4. Add the source identifier to the `dataSources` array in the card schema
145
+ 5. Document the version used