pubid 2.0.0.pre.alpha.12 → 2.0.0.pre.alpha.14

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (285) hide show
  1. checksums.yaml +4 -4
  2. data/README.adoc +43 -1
  3. data/data/ieee/update_codes.yaml +17 -4
  4. data/data/nist/update_codes.yaml +7 -3
  5. data/data/parg/tables/bipm_groups.yaml +14 -0
  6. data/data/parg/tables/bipm_type_codes.yaml +5 -0
  7. data/data/parg/tables/bipm_type_names_en.yaml +6 -0
  8. data/data/parg/tables/bipm_type_names_fr.yaml +6 -0
  9. data/data/parg/tables/directives_supplements_typed_stages.yaml +3 -0
  10. data/data/parg/tables/directives_typed_stages.yaml +5 -0
  11. data/data/parg/tables/idf_typed_stages.yaml +27 -0
  12. data/data/parg/tables/idf_typed_stages_supplements.yaml +2 -0
  13. data/data/parg/tables/iec_typed_stages.yaml +130 -0
  14. data/data/parg/tables/iso_publishers.yaml +4 -0
  15. data/data/parg/tables/organizations.yaml +12 -0
  16. data/data/parg/tables/tc_types.yaml +42 -0
  17. data/data/parg/tables/typed_stages.yaml +114 -0
  18. data/data/parg/tables/typed_stages_supplements.yaml +64 -0
  19. data/data/parg/tables/wg_types.yaml +21 -0
  20. data/lib/pubid/adobe/builder.rb +2 -0
  21. data/lib/pubid/adobe/identifier.rb +11 -1
  22. data/lib/pubid/all_parts.rb +201 -0
  23. data/lib/pubid/all_parts_identifier.rb +19 -0
  24. data/lib/pubid/amca/CLAUDE.md +47 -0
  25. data/lib/pubid/amca/builder.rb +3 -5
  26. data/lib/pubid/amca/identifiers/base.rb +11 -1
  27. data/lib/pubid/amca/identifiers/publication.rb +13 -0
  28. data/lib/pubid/amca/parser.rb +2 -1
  29. data/lib/pubid/amca/renderer.rb +22 -33
  30. data/lib/pubid/amca/urn_generator.rb +21 -2
  31. data/lib/pubid/amca/urn_parser.rb +36 -10
  32. data/lib/pubid/ansi/builder.rb +6 -0
  33. data/lib/pubid/ansi/identifier.rb +1 -1
  34. data/lib/pubid/api/CLAUDE.md +23 -0
  35. data/lib/pubid/api/builder.rb +2 -0
  36. data/lib/pubid/api/identifier.rb +1 -1
  37. data/lib/pubid/api/parser.rb +8 -4
  38. data/lib/pubid/ashrae/CLAUDE.md +13 -0
  39. data/lib/pubid/ashrae/builder.rb +58 -14
  40. data/lib/pubid/ashrae/identifiers/base.rb +10 -1
  41. data/lib/pubid/ashrae/identifiers/errata.rb +14 -2
  42. data/lib/pubid/ashrae/identifiers/interpretation.rb +2 -10
  43. data/lib/pubid/ashrae/parser.rb +80 -39
  44. data/lib/pubid/ashrae/renderer.rb +32 -1
  45. data/lib/pubid/ashrae/urn_generator.rb +32 -9
  46. data/lib/pubid/asme/CLAUDE.md +25 -0
  47. data/lib/pubid/asme/builder.rb +16 -9
  48. data/lib/pubid/asme/components/code.rb +2 -0
  49. data/lib/pubid/asme/identifier.rb +1 -1
  50. data/lib/pubid/asme/identifiers/standard.rb +6 -1
  51. data/lib/pubid/asme/parser.rb +41 -14
  52. data/lib/pubid/astm/CLAUDE.md +9 -0
  53. data/lib/pubid/astm/builder.rb +2 -0
  54. data/lib/pubid/astm/components/code.rb +2 -0
  55. data/lib/pubid/astm/identifier.rb +1 -1
  56. data/lib/pubid/astm/parser.rb +4 -1
  57. data/lib/pubid/bipm/CLAUDE.md +11 -0
  58. data/lib/pubid/bipm/builder.rb +2 -0
  59. data/lib/pubid/bipm/identifier.rb +1 -1
  60. data/lib/pubid/bsi/CLAUDE.md +93 -0
  61. data/lib/pubid/bsi/builder.rb +13 -11
  62. data/lib/pubid/bsi/identifiers/addendum_document.rb +2 -0
  63. data/lib/pubid/bsi/identifiers/adopted_european_norm.rb +6 -54
  64. data/lib/pubid/bsi/identifiers/adopted_international_standard.rb +5 -22
  65. data/lib/pubid/bsi/identifiers/amendment.rb +36 -12
  66. data/lib/pubid/bsi/identifiers/bundled_identifier.rb +2 -0
  67. data/lib/pubid/bsi/identifiers/consolidated_identifier.rb +23 -26
  68. data/lib/pubid/bsi/identifiers/corrigendum.rb +29 -12
  69. data/lib/pubid/bsi/identifiers/expert_commentary.rb +6 -7
  70. data/lib/pubid/bsi/identifiers/national_annex.rb +18 -20
  71. data/lib/pubid/bsi/identifiers/root_identity.rb +31 -0
  72. data/lib/pubid/bsi/identifiers/set.rb +2 -0
  73. data/lib/pubid/bsi/identifiers/supplement_document.rb +2 -0
  74. data/lib/pubid/bsi/identifiers.rb +1 -0
  75. data/lib/pubid/bsi/parser.rb +8 -8
  76. data/lib/pubid/bsi/renderer.rb +20 -20
  77. data/lib/pubid/bsi/single_identifier.rb +1 -3
  78. data/lib/pubid/bsi/urn_generator.rb +28 -18
  79. data/lib/pubid/builder/base.rb +27 -0
  80. data/lib/pubid/calconnect/builder.rb +2 -0
  81. data/lib/pubid/calconnect/identifier.rb +5 -1
  82. data/lib/pubid/ccsds/builder.rb +2 -0
  83. data/lib/pubid/ccsds/identifier.rb +9 -1
  84. data/lib/pubid/cen_cenelec/CLAUDE.md +59 -0
  85. data/lib/pubid/cen_cenelec/builder.rb +6 -1
  86. data/lib/pubid/cen_cenelec/identifier.rb +12 -28
  87. data/lib/pubid/cen_cenelec/identifiers/amendment.rb +3 -10
  88. data/lib/pubid/cen_cenelec/identifiers/corrigendum.rb +3 -10
  89. data/lib/pubid/cen_cenelec/parser.rb +20 -5
  90. data/lib/pubid/cie/CLAUDE.md +58 -0
  91. data/lib/pubid/cie/builder.rb +2 -0
  92. data/lib/pubid/cie/components/language.rb +2 -0
  93. data/lib/pubid/cie/identifier.rb +1 -1
  94. data/lib/pubid/cie/parser.rb +9 -2
  95. data/lib/pubid/components/adoption.rb +2 -0
  96. data/lib/pubid/components/code.rb +2 -0
  97. data/lib/pubid/components/date.rb +8 -6
  98. data/lib/pubid/components/edition.rb +2 -0
  99. data/lib/pubid/components/iteration.rb +2 -0
  100. data/lib/pubid/components/language.rb +2 -0
  101. data/lib/pubid/components/locality.rb +2 -0
  102. data/lib/pubid/components/publisher.rb +2 -0
  103. data/lib/pubid/components/relationship.rb +2 -0
  104. data/lib/pubid/components/stage.rb +2 -0
  105. data/lib/pubid/components/supplement.rb +2 -0
  106. data/lib/pubid/components/type.rb +2 -0
  107. data/lib/pubid/components/typed_stage.rb +8 -0
  108. data/lib/pubid/conformance/checks.rb +1 -1
  109. data/lib/pubid/csa/CLAUDE.md +41 -0
  110. data/lib/pubid/csa/builder.rb +2 -0
  111. data/lib/pubid/csa/identifier.rb +19 -3
  112. data/lib/pubid/csa/parser.rb +25 -8
  113. data/lib/pubid/csa/renderer.rb +12 -12
  114. data/lib/pubid/csa/single_identifier.rb +17 -0
  115. data/lib/pubid/doi/builder.rb +2 -0
  116. data/lib/pubid/doi/identifier.rb +1 -1
  117. data/lib/pubid/easc/builder.rb +2 -0
  118. data/lib/pubid/easc/identifier.rb +10 -1
  119. data/lib/pubid/ecma/CLAUDE.md +28 -0
  120. data/lib/pubid/ecma/builder.rb +2 -0
  121. data/lib/pubid/ecma/identifier.rb +8 -1
  122. data/lib/pubid/etsi/CLAUDE.md +34 -0
  123. data/lib/pubid/etsi/builder.rb +2 -0
  124. data/lib/pubid/etsi/components/code.rb +6 -0
  125. data/lib/pubid/etsi/components/version.rb +2 -0
  126. data/lib/pubid/etsi/identifiers/base.rb +1 -1
  127. data/lib/pubid/etsi/identifiers/etsi_standard.rb +7 -0
  128. data/lib/pubid/evs/CLAUDE.md +58 -0
  129. data/lib/pubid/evs/builder.rb +2 -0
  130. data/lib/pubid/evs.rb +1 -1
  131. data/lib/pubid/gb/CLAUDE.md +140 -0
  132. data/lib/pubid/gb/builder.rb +7 -2
  133. data/lib/pubid/gb/identifier.rb +6 -4
  134. data/lib/pubid/gb/identifiers/all_parts.rb +17 -0
  135. data/lib/pubid/gb/identifiers.rb +1 -0
  136. data/lib/pubid/gb/renderer.rb +0 -1
  137. data/lib/pubid/gost/CLAUDE.md +64 -0
  138. data/lib/pubid/gost/builder.rb +3 -1
  139. data/lib/pubid/gost/identifier.rb +16 -1
  140. data/lib/pubid/gost/parser.rb +8 -1
  141. data/lib/pubid/iala/CLAUDE.md +82 -0
  142. data/lib/pubid/iala/builder.rb +2 -0
  143. data/lib/pubid/iala/identifier.rb +10 -1
  144. data/lib/pubid/iana/CLAUDE.md +7 -0
  145. data/lib/pubid/iana/builder.rb +2 -0
  146. data/lib/pubid/iana/identifier.rb +1 -1
  147. data/lib/pubid/identifier.rb +161 -17
  148. data/lib/pubid/idf/builder.rb +11 -1
  149. data/lib/pubid/idf/identifier.rb +5 -0
  150. data/lib/pubid/idf/identifiers/all_parts.rb +17 -0
  151. data/lib/pubid/idf/identifiers.rb +1 -0
  152. data/lib/pubid/iec/CLAUDE.md +31 -0
  153. data/lib/pubid/iec/builder.rb +7 -1
  154. data/lib/pubid/iec/components/consolidated_amendment.rb +4 -0
  155. data/lib/pubid/iec/components/sheet.rb +2 -0
  156. data/lib/pubid/iec/components/trf_info.rb +2 -0
  157. data/lib/pubid/iec/components/vap_suffix.rb +2 -0
  158. data/lib/pubid/iec/identifier.rb +8 -3
  159. data/lib/pubid/iec/identifiers/all_parts.rb +19 -0
  160. data/lib/pubid/iec/identifiers.rb +1 -0
  161. data/lib/pubid/iec/parser.rb +9 -4
  162. data/lib/pubid/iec/renderer.rb +0 -1
  163. data/lib/pubid/iec/urn_generator.rb +9 -1
  164. data/lib/pubid/iec/urn_parser.rb +3 -2
  165. data/lib/pubid/ieee/CLAUDE.md +97 -0
  166. data/lib/pubid/ieee/builder.rb +194 -27
  167. data/lib/pubid/ieee/components/code.rb +2 -0
  168. data/lib/pubid/ieee/components/draft.rb +35 -2
  169. data/lib/pubid/ieee/components/typed_stage.rb +2 -0
  170. data/lib/pubid/ieee/identifiers/base.rb +21 -1
  171. data/lib/pubid/ieee/identifiers/iec_ieee_copublished.rb +9 -0
  172. data/lib/pubid/ieee/identifiers/joint_development.rb +66 -23
  173. data/lib/pubid/ieee/identifiers/project_draft_identifier.rb +8 -1
  174. data/lib/pubid/ieee/parser.rb +156 -29
  175. data/lib/pubid/ieee/renderer.rb +42 -7
  176. data/lib/pubid/ieee/urn_generator.rb +31 -0
  177. data/lib/pubid/ietf/CLAUDE.md +7 -0
  178. data/lib/pubid/ietf/builder.rb +2 -0
  179. data/lib/pubid/ietf/identifiers/base.rb +1 -1
  180. data/lib/pubid/iho/builder.rb +2 -0
  181. data/lib/pubid/isbn/builder.rb +2 -0
  182. data/lib/pubid/isbn/identifier.rb +1 -1
  183. data/lib/pubid/iso/CLAUDE.md +47 -0
  184. data/lib/pubid/iso/builder.rb +19 -5
  185. data/lib/pubid/iso/components/publisher.rb +2 -0
  186. data/lib/pubid/iso/identifier.rb +10 -15
  187. data/lib/pubid/iso/identifiers/all_parts.rb +19 -0
  188. data/lib/pubid/iso/identifiers/directives_supplement.rb +4 -2
  189. data/lib/pubid/iso/identifiers.rb +1 -0
  190. data/lib/pubid/iso/normalizer.rb +4 -1
  191. data/lib/pubid/iso/rendering_style.rb +0 -1
  192. data/lib/pubid/itu/CLAUDE.md +115 -0
  193. data/lib/pubid/itu/builder.rb +26 -4
  194. data/lib/pubid/itu/components/code.rb +2 -0
  195. data/lib/pubid/itu/components/designation.rb +2 -0
  196. data/lib/pubid/itu/components/sector.rb +2 -0
  197. data/lib/pubid/itu/components/series.rb +2 -0
  198. data/lib/pubid/itu/identifiers/base.rb +11 -18
  199. data/lib/pubid/itu/identifiers/radio_regulations.rb +27 -0
  200. data/lib/pubid/itu/identifiers/special_publication.rb +48 -14
  201. data/lib/pubid/itu/identifiers/standard_serialization.rb +2 -0
  202. data/lib/pubid/itu/identifiers/supplement.rb +15 -0
  203. data/lib/pubid/itu/identifiers.rb +1 -0
  204. data/lib/pubid/itu/parser.rb +108 -22
  205. data/lib/pubid/itu/urn_generator.rb +9 -2
  206. data/lib/pubid/jcgm/CLAUDE.md +7 -0
  207. data/lib/pubid/jcgm/builder.rb +2 -0
  208. data/lib/pubid/jcgm/components/publisher.rb +2 -0
  209. data/lib/pubid/jcgm.rb +1 -1
  210. data/lib/pubid/jis/builder.rb +5 -1
  211. data/lib/pubid/jis/identifier.rb +6 -18
  212. data/lib/pubid/jis/identifiers/all_parts.rb +19 -0
  213. data/lib/pubid/jis/identifiers.rb +1 -0
  214. data/lib/pubid/jis/renderer.rb +0 -2
  215. data/lib/pubid/jis/urn_generator.rb +0 -1
  216. data/lib/pubid/nist/CLAUDE.md +56 -0
  217. data/lib/pubid/nist/builder.rb +3 -0
  218. data/lib/pubid/nist/components/edition.rb +2 -0
  219. data/lib/pubid/nist/components/issue_number.rb +2 -0
  220. data/lib/pubid/nist/components/part.rb +2 -0
  221. data/lib/pubid/nist/components/stage.rb +2 -0
  222. data/lib/pubid/nist/components/supplement.rb +2 -0
  223. data/lib/pubid/nist/components/translation.rb +2 -0
  224. data/lib/pubid/nist/components/update.rb +2 -0
  225. data/lib/pubid/nist/components/version.rb +2 -0
  226. data/lib/pubid/nist/components/volume.rb +2 -0
  227. data/lib/pubid/nist/identifiers/base.rb +39 -7
  228. data/lib/pubid/nist/parser.rb +24 -2
  229. data/lib/pubid/nist/preprocessor.rb +53 -2
  230. data/lib/pubid/nist/urn_parser.rb +10 -1
  231. data/lib/pubid/oasis/CLAUDE.md +19 -0
  232. data/lib/pubid/oasis/builder.rb +2 -0
  233. data/lib/pubid/oasis/identifier.rb +20 -1
  234. data/lib/pubid/ogc/CLAUDE.md +34 -0
  235. data/lib/pubid/ogc/builder.rb +2 -0
  236. data/lib/pubid/ogc/identifier.rb +12 -1
  237. data/lib/pubid/oiml/CLAUDE.md +189 -0
  238. data/lib/pubid/oiml/builder.rb +20 -0
  239. data/lib/pubid/oiml/components/code.rb +6 -0
  240. data/lib/pubid/oiml/identifier.rb +13 -0
  241. data/lib/pubid/oiml/identifiers/annex.rb +4 -0
  242. data/lib/pubid/oiml/identifiers/certification_system.rb +34 -0
  243. data/lib/pubid/oiml/identifiers/code_number.rb +8 -0
  244. data/lib/pubid/oiml/identifiers/dual_published.rb +174 -0
  245. data/lib/pubid/oiml/identifiers.rb +2 -0
  246. data/lib/pubid/oiml/parser.rb +35 -4
  247. data/lib/pubid/oiml/renderer.rb +23 -1
  248. data/lib/pubid/oiml/single_identifier.rb +4 -0
  249. data/lib/pubid/oiml/supplement_identifier.rb +7 -0
  250. data/lib/pubid/oiml/urn_generator.rb +28 -0
  251. data/lib/pubid/oiml.rb +6 -1
  252. data/lib/pubid/omg/CLAUDE.md +15 -0
  253. data/lib/pubid/omg/builder.rb +2 -0
  254. data/lib/pubid/omg/identifier.rb +1 -1
  255. data/lib/pubid/parg/artifact.rb +46 -0
  256. data/lib/pubid/parg/backend.rb +92 -0
  257. data/lib/pubid/parg.rb +8 -0
  258. data/lib/pubid/parser/grammar.rb +23 -0
  259. data/lib/pubid/pg.rb +8 -0
  260. data/lib/pubid/plateau/builder.rb +2 -0
  261. data/lib/pubid/plateau/identifiers/base.rb +4 -0
  262. data/lib/pubid/plateau/supplement_identifier.rb +14 -2
  263. data/lib/pubid/plateau/urn_generator.rb +7 -1
  264. data/lib/pubid/plateau.rb +1 -2
  265. data/lib/pubid/renderers/human_readable.rb +0 -1
  266. data/lib/pubid/sae/builder.rb +2 -0
  267. data/lib/pubid/sae/components/date.rb +2 -0
  268. data/lib/pubid/sae/components/type.rb +2 -0
  269. data/lib/pubid/sae/identifiers/base.rb +1 -1
  270. data/lib/pubid/subset_match.rb +197 -0
  271. data/lib/pubid/tgpp/CLAUDE.md +43 -0
  272. data/lib/pubid/tgpp/builder.rb +2 -0
  273. data/lib/pubid/tgpp/identifier.rb +15 -1
  274. data/lib/pubid/type_resolver.rb +14 -2
  275. data/lib/pubid/un/builder.rb +2 -0
  276. data/lib/pubid/un/identifier.rb +1 -1
  277. data/lib/pubid/version.rb +1 -1
  278. data/lib/pubid/w3c/CLAUDE.md +7 -0
  279. data/lib/pubid/w3c/builder.rb +2 -0
  280. data/lib/pubid/w3c/identifier.rb +1 -1
  281. data/lib/pubid/xsf/CLAUDE.md +11 -0
  282. data/lib/pubid/xsf/builder.rb +2 -0
  283. data/lib/pubid/xsf/identifier.rb +1 -1
  284. data/lib/pubid.rb +17 -3
  285. metadata +78 -2
@@ -128,12 +128,20 @@ module Pubid
128
128
  parts << type_str unless type_str.strip.empty?
129
129
  end
130
130
 
131
- # Code - with P prefix for projects (concatenated, not separated)
131
+ # Code - with P prefix for projects (concatenated, not separated).
132
+ # IEEE semantics (normative): P = project = the document is a
133
+ # draft; no P = it has become a standard. The P-state is
134
+ # identity-bearing, so the renderer never adds or drops it —
135
+ # whatever the source spelling carries round-trips.
132
136
  if id.code_obj
133
137
  result = id.code_obj.to_s
134
138
 
135
- # Prepend P if this is a project AND code doesn't already have P
136
- if id.typed_stage&.project_status && should_render_type && !result.start_with?("P")
139
+ # Prepend P if this is a project AND code doesn't already have P.
140
+ # A recorded project marker (the source spelled the P on a
141
+ # non-IEEE-led publisher) prints regardless of the publisher.
142
+ if (id.typed_stage&.project_status && should_render_type ||
143
+ id.is_a?(Identifiers::ProjectDraftIdentifier) &&
144
+ id.project_marker) && !result.start_with?("P")
137
145
  result = "P#{result}"
138
146
  end
139
147
 
@@ -142,6 +150,7 @@ module Pubid
142
150
  result += mark(id.code_obj.number, id.code_obj.prefix,
143
151
  publishers: publishers_of(id))
144
152
 
153
+
145
154
  # Only attach year to code if there's no edition, no month, and no draft
146
155
  result += "-#{id.year}" if id.year && !id.draft_obj && !id.edition && !id.month
147
156
 
@@ -150,9 +159,31 @@ module Pubid
150
159
  # ("P802.16Rev2/D3"). Keeps a revision distinct from its base standard.
151
160
  result += "Rev#{id.revision}" if id.revision
152
161
 
153
- # Append draft to code - with or without space based on original format
162
+ # Append draft to code - with or without space based on original
163
+ # format. An exactly-IEC/IEEE co-published reference prints only the
164
+ # draft DESIGNATOR - a comma-separated project date ("IEC/IEEE
165
+ # 61886-1/D2, May 2020" prints as "…/D2") is catalogue metadata, not
166
+ # identity. Wider sets ("IEC/ISO/IEEE P82079-1/D1, June 2016") print
167
+ # the date; a space-separated date always stays ("…/DCDV July 2015").
168
+ # IEEE-published drafts normalize the comma date to the long form
169
+ # ("D08, September, 2018").
154
170
  if id.draft_obj
155
- result += id.space_before_draft ? " #{id.draft_obj}" : id.draft_obj.to_s
171
+ printed = id.draft_obj.to_s
172
+ # The draft designator carries its date on every lead
173
+ # (docs/IEEE-DRAFT-STAGES.md §3: "the draft designator with its
174
+ # date" — the IEC/IEEE-led date-drop contradicted the DCD
175
+ # convention and the raw records).
176
+ if id.publisher == "IEEE" && id.draft_status.to_s.empty?
177
+ # The long comma form is for dated project drafts; an
178
+ # unapproved-draft render keeps its pinned single-comma form
179
+ # (pubid#318 idempotence).
180
+ printed = printed.sub(/, ([A-Z][a-z]+) (\d{4})\z/, ', \1, \2')
181
+ end
182
+ # A space-separated draft year with no print date trailing is
183
+ # IEEE's dash form ("…/D1-2006"); with a trailing print date the
184
+ # space stays ("…/D1 2004, Nov 2004").
185
+ printed = printed.sub(/\A(D\d+) (\d{4})\z/, '\1-\2')
186
+ result += id.space_before_draft ? " #{printed}" : printed
156
187
  end
157
188
 
158
189
  # Append interpretation notation (/INT)
@@ -190,8 +221,12 @@ module Pubid
190
221
  # Build the main identifier (without month yet)
191
222
  result = parts.join(" ")
192
223
 
193
- # Month/Day - append directly to avoid extra space before comma
194
- if id.month
224
+
225
+ # Month/Day - append directly to avoid extra space before comma.
226
+ # An exactly-IEC/IEEE co-published reference drops the comma date
227
+ # entirely (same catalogue-metadata rule as the draft date above):
228
+ # "IEC/IEEE P60076-16, May 2016" prints as "IEC/IEEE 60076-16".
229
+ if id.month && !(id.publisher == "IEC" && id.copublisher == ["IEEE"])
195
230
  result += ", #{id.month}"
196
231
  result += " #{id.day}" if id.day
197
232
  if id.year && !id.edition
@@ -127,6 +127,10 @@ module Pubid
127
127
  end
128
128
 
129
129
  def publisher_component
130
+ # The IEC/IEEE co-published pair is definitional to the class and
131
+ # was previously lost ("urn:ieee:ieee" — no number, no draft).
132
+ return "iec-ieee" if identifier.is_a?(Identifiers::IecIeeeCopublished)
133
+
130
134
  pub = normalized_publisher
131
135
 
132
136
  if identifier.copublisher&.any?
@@ -150,6 +154,14 @@ module Pubid
150
154
  end
151
155
 
152
156
  def type_component
157
+ # A joint stage draft carries the stage in its draft clause and the
158
+ # project P inside the code — a separate type segment would
159
+ # duplicate both.
160
+ if identifier.is_a?(Identifiers::JointDevelopment) &&
161
+ identifier.ieee_draft.to_s.start_with?("D=")
162
+ return nil
163
+ end
164
+
153
165
  return nil unless identifier.type
154
166
 
155
167
  type = identifier.type
@@ -159,6 +171,16 @@ module Pubid
159
171
  end
160
172
 
161
173
  def code_component
174
+ # The co-published code is number+parts+separators (no year — its
175
+ # own segment), e.g. "61886-1".
176
+ if identifier.is_a?(Identifiers::IecIeeeCopublished)
177
+ return nil if identifier.number.to_s.empty?
178
+
179
+ # Year-less: the publication year is its own URN segment (the
180
+ # rebuilt copublished_number would double-emit it).
181
+ return "#{identifier.number}#{identifier.parts_suffix}"
182
+ end
183
+
162
184
  return nil unless identifier.code_obj
163
185
 
164
186
  identifier.code_obj.to_s
@@ -182,6 +204,15 @@ module Pubid
182
204
  end
183
205
 
184
206
  def draft_component
207
+ if identifier.is_a?(Identifiers::JointDevelopment) &&
208
+ identifier.ieee_draft.to_s.start_with?("D=")
209
+ return "draft.#{identifier.ieee_draft}"
210
+ end
211
+
212
+ if identifier.is_a?(Identifiers::IecIeeeCopublished)
213
+ return identifier.draft_info ? "draft.#{identifier.draft_info}" : nil
214
+ end
215
+
185
216
  return nil unless identifier.draft_obj
186
217
 
187
218
  "draft.#{identifier.draft_obj}"
@@ -0,0 +1,7 @@
1
+ # IETF flavor notes
2
+
3
+ IETF readiness for the unified relaton index.
4
+
5
+ These notes were part of the root `CLAUDE.md`. Read them before you change `lib/pubid/ietf/` or `spec/pubid/ietf/`. The root file keeps the cross-flavor contract that every flavor obeys.
6
+
7
+ - **IETF readiness for the unified relaton index (`relaton-data-ietf`)**: relaton is collapsing RFCs, the RFC sub-series (BCP/STD/FYI) and Internet-Drafts into **one** index keyed on the structured pubid hash (`pubid_class: ::Pubid::Ietf::Identifier`). That index is **all-or-nothing** — `Relaton::Index::FileIO#deserialize_id` raises on the *first* unparseable row and `#load_index` then rejects the whole index — so one bad id out of 176,862 means no consumer resolves *any* IETF reference. Four changes, all verified against the published corpora (`relaton-data-{rfcs,rfcsubseries,ids}`, branch `v2`). **(1) `draft_rest` accepts `.` and `A-Z`** (`match['a-zA-Z0-9+_.\-']`): exactly 50 of the 166,740 published draft ids failed without them (21 dotted `draft-ietf-pilc-2.5g3g-12`, 29 uppercase `draft-chapin-clnp-ISO8473-00`, 0 both), and a character census confirms the class now covers the corpus **exactly** — the only chars outside `[a-z0-9-]` anywhere in it are `. _ +` and `A-Z`. Widening is safe because the top-level alternation is keyed on the leading token (`RFC`/`BCP|STD|FYI`/`draft-`), so a wider draft tail can never make another branch match; `Builder#split_draft_version` needed no change (it inspects only the last three characters). The set stays **closed** — a space, `/` or `&` in a slug is crawler junk and still fails (pinned by negative examples). **(2) Zero-padded ids parse, grammar-level**: `rfc-index.xml` writes sub-series membership as `<is-also><doc-id>STD0066</doc-id></is-also>` (369 padded doc-ids) and RFCs as `RFC0001`, and relaton#109 forbids relaton normalizing them itself. The space is `.maybe` on both the `rfc` and `subseries` rules and `Builder#unpad` strips the pad (`/\A0+(?=\d)/`, a look-ahead so an all-zero number keeps its last digit), so `STD0066`/`STD 0066`/`STD66` all render the canonical `STD 66`. `unpad` is applied unconditionally, so the **spaced** padded form is normalized too — `RFC 0001` previously parsed and rendered back verbatim, and now renders `RFC 1` with a correspondingly changed hash and URN. That is deliberate (one canonical spelling per document) and has **zero** effect on the published corpora, which contain no padded row. **This is a normalizing parse, so the padded spellings must NEVER go into the byte-exact `spec/fixtures/ietf/identifiers/pass/`** — they live in `spec/pubid/ietf/zero_padded_spec.rb`, the BIPM `CIPM/2005-06(REV)` precedent. **(3) The Internet-Draft slug moved from `name` into `number`, and `series` became derived**: relaton bsearches on `id.root.number.to_s`, so with the slug in `name` all 166,740 drafts keyed to `""` and the narrowing bought nothing (with the slug: 43,564 buckets, median 2, worst 101). `name` is **dropped entirely** (no alias); the slug — leading `draft-` included — is the key, deliberately shared by a draft's unversioned aggregator row and every one of its versions so a document and its variants cluster. `series` is likewise **not stored**: `_type: pubid:ietf:std` already encodes it, so `Bcp`/`Std`/`Fyi` expose a `SERIES` constant + a plain `#series` reader (safe — `::Pubid::Identifier` declares no `series` attribute and nothing generic reads `.series`) and `build_subseries` stops passing it. The whole serialized vocabulary is now **three keys**: `_type` + `number` always, `version` only on a versioned draft — a sub-series row (`{_type, number}`) has the same shape as an RFC's. **The attributes live on the five LEAVES, not the shared `Pubid::Ietf::Identifier`**, via the `Identifiers::Serialization` mixin (`number` + the `key_value` block; `InternetDraft` merges its own `version` mapping on top — lutaml combines a mixin-installed block with a second in-class one, the ITU precedent). This is the IEEE `number` determinism landmine: a `:string` redefinition of the parent's `Components::Code number` on a class the leaves *inherit* from resolves nondeterministically under multi-flavor load. It was invisible while drafts left `number` nil; it is load-bearing now. The base deliberately carries **no** `key_value` block — one there would be inherited-and-merged by every leaf. `spec/pubid/ietf/root_number_spec.rb` is the tripwire (asserts `number` is a `String`, the exact keys, the exact `to_hash` key sets) and **must run under the full `rake` suite** to mean anything. **(4) Round-trip is validated here because relaton structurally cannot**: `FileIO#id_supported?` **skips** its `to_hash`/`from_hash` check whenever the object is a concrete subclass (`return true unless obj.instance_of?(@pubid_class)`), and *every* IETF id is a subclass — so a lossy round-trip would pass validation silently and produce wrong lookups. Three layers: `fixtures_spec.rb` is now **zero-failure** (was a 90% threshold — wrong for an all-or-nothing index) and checks `to_s` byte-exactness, `from_hash(to_hash)` fidelity in both `to_s` and `to_hash`, **and** a non-empty `root.number` for every fixture line; the fixtures grew to ~2,500 lines (the complete 365-row sub-series corpus, a 139-id structural edge set, a deterministic 1-in-84 draft sample — each file's header records the regeneration command); and `spec/pubid/ietf/corpus_round_trip_spec.rb` runs the same four checks over all 176,862 published ids, **opt-in** via `PUBID_IETF_CORPUS=<dir with the three relaton-data checkouts> bundle exec rake test:corpus_ietf` (self-skips otherwise; ~100s; verified 0 failures). **relaton note — the indexes must be regenerated, there is no alias**: `name` and `series` are gone as serialized keys, so a pre-`number` row (`{_type, name, version}`) deserializes with **`number` nil** and no error at all — lutaml ignores unknown keys, and relaton's subclass skip means `id_supported?` won't catch it either, leaving every draft keyed under `""`. Rendering is therefore the loud failure: `Renderer#require_number!` raises `ArgumentError` on a nil/empty `number` rather than returning `nil` or `"-02"`, and the renderer's sub-series arm is an explicit `when Bcp, Std, Fyi` with an `else raise` instead of a catch-all (a bare `Pubid::Ietf::Identifier` has no `series` reader). Pinned by the "legacy `name`-shaped hash" block in `root_number_spec.rb`. **Pre-existing bug fixed in passing**: `UrnParser`'s bare `Errors::ParseError` does not resolve — `Errors` lives under `Pubid::UrnParser` (the module), which is neither in the lexical scope of `Pubid::Ietf::UrnParser` nor an ancestor of `Base` — so an invalid URN raised `NameError` instead, and the spec's `raise_error(StandardError)` accepted it. Now fully qualified as `Pubid::UrnParser::Errors::ParseError` (the adobe/easc/gost/iala convention) with the spec asserting that exact class. **`Pubid::Iec::UrnParser` has the identical bug and is untouched here.** Also: `Identifier.parse` gained the inline `MAX_INPUT_LENGTH` guard — relaton reaches that class-level funnel directly through `pubid_class:`, bypassing `Pubid::Ietf.parse` — and IETF was added to the three cross-flavor tables it was missing from (`identifier_roundtrip_spec`, `uniform_identifier_handle_spec`, `redos_guard_spec`), which is what exercises the leaf `number` under real multi-flavor load. (hand-off: ietf-index-readiness.)
@@ -76,3 +76,5 @@ module Pubid
76
76
  end
77
77
  end
78
78
  end
79
+
80
+ Pubid::Ietf::Builder.prepend(Pubid::Builder::AllPartsWrap)
@@ -64,7 +64,7 @@ module Pubid
64
64
  raise Pubid::Errors::InvalidInputError, Pubid::INPUT_TOO_LONG_MESSAGE
65
65
  end
66
66
 
67
- parsed = Parser.parse(identifier)
67
+ parsed = Pubid::Parg::Backend.parse(:ietf, identifier)
68
68
  Builder.build(parsed)
69
69
  end
70
70
  end
@@ -35,3 +35,5 @@ module Pubid
35
35
  end
36
36
  end
37
37
  end
38
+
39
+ Pubid::Iho::Builder.prepend(Pubid::Builder::AllPartsWrap)
@@ -42,3 +42,5 @@ module Pubid
42
42
  end
43
43
  end
44
44
  end
45
+
46
+ Pubid::Isbn::Builder.prepend(Pubid::Builder::AllPartsWrap)
@@ -64,7 +64,7 @@ module Pubid
64
64
 
65
65
  # @raise [Pubid::Errors::ParseError] if the string is not a valid ISBN
66
66
  def self.build_identifier(identifier)
67
- parsed = Parser.parse(identifier)
67
+ parsed = Pubid::Parg::Backend.parse(:isbn, identifier)
68
68
  Builder.build(parsed)
69
69
  rescue ArgumentError => e
70
70
  # The Builder validates length and check digit. Surface that as a parse
@@ -0,0 +1,47 @@
1
+ # ISO flavor notes
2
+
3
+ - **No identifier attribute holds a `Components::Code` any more**: `number`,
4
+ `part` and `subpart` on `Pubid::Iso::Identifier`, the committee structure on
5
+ `Identifiers::TcDocument` (`tc_type`/`tc_number`/`sc_type`/`sc_number`/
6
+ `wg_type`/`wg_number`), `Identifiers::Directives#subgroup` and the
7
+ `edition.number` the builder writes are all plain strings, and
8
+ `Pubid::Iso::Components::Code` is deleted.
9
+ **The component was degenerate everywhere it was used.** Measured over the
10
+ whole 7,613-id pass corpus: no ISO Code carries `prefix`, `part`, `subpart`
11
+ or `parts`, `parts` is never written anywhere in `lib/pubid/iso`, and no Code
12
+ renders differently from its own `value`. ISO's part and subpart are sibling
13
+ attributes of the identifier, not fields of the number — the `-` join in
14
+ `Code#to_s` was unreachable.
15
+ **`edition.number` was worse than degenerate: it leaked a live object.**
16
+ `Components::Edition#number` is typed `Lutaml::Model::Type::Value`, so lutaml
17
+ never serialized it — `to_hash` returned the `Components::Code` instance
18
+ itself, and `to_yaml` emitted `!ruby/object:Pubid::Components::Code` with its
19
+ internal ivars, which `YAML.safe_load` refuses. An edition now serializes
20
+ `{"number" => "13", "original_text" => "Ed 13"}`.
21
+ **The TC converters are gone with it**: twelve `*_to_kv`/`*_from_kv` helpers
22
+ were replaced by plain `map "tc_type", to: :tc_type` declarations (the
23
+ ETSI/OIML shape), because a String needs no converter to flatten.
24
+ **Two shared surfaces needed a change, and both are the documented shims**:
25
+ `Renderers::DirectivesRenderer` called `subgroup.render(context:)` directly
26
+ and now reads through `render_component`; `Pubid::Iso.build_code` returned a
27
+ Code and now returns the string (ASTM's `to_iso_identifier` goes through it,
28
+ which is how one ASTM example caught the `NameError` after the class was
29
+ deleted).
30
+ **Verification — replay a baseline, do not trust the suite alone.** Against
31
+ all 7,621 pass-fixture ids captured on the parent commit, `to_s`, `to_urn`,
32
+ `to_mr_string`, `root.number` and the identifier class are **byte-identical**,
33
+ and `to_hash` differs for exactly **12** ids: the 9 `subgroup` rows and the 3
34
+ editions below. The first replay also caught a real regression the specs did
35
+ not — `directives.rb` still called `part.value.downcase`, so 4 Directives
36
+ URNs raised.
37
+ **relaton note — `relaton-data-iso` needs a re-crawl, and only for one key.**
38
+ `subgroup` flattens from `{"value" => "JTC 1"}` to `"JTC 1"`, and the
39
+ published `index-v2.yaml` (branch `v2`) carries **5 such rows**, all ISO/IEC
40
+ JTC 1 directives. **No compatibility shim is kept** — a stored nested
41
+ `subgroup` will not deserialize. The edition change reaches nothing stored:
42
+ `grep -c edition` over that index returns **0**.
43
+ **Spec fallout is the shape to expect for the remaining flavors**: 592
44
+ assertions across 25 files read `.number.value` / `.part.value`. The rewrite
45
+ needs a lookbehind, because `tc_number.value` and `edition.number.value` are
46
+ different attributes — the first pass stripped `edition.number.value` too,
47
+ and two corrigendum examples caught it.
@@ -58,6 +58,12 @@ module Pubid
58
58
  end
59
59
 
60
60
  def build(parsed_hash)
61
+ # "(all parts)" and the ":ser" URN name every part of the document,
62
+ # so they build an AllPartsIdentifier around the document (the URN
63
+ # parser puts the key on the outermost identifier for the same
64
+ # reason). The document itself holds no all-parts mark.
65
+ all_parts = parsed_hash.delete(:all_parts)
66
+
61
67
  # For ISO/R legacy format, split into publisher and type
62
68
  if parsed_hash[:iso_r_prefix]
63
69
  parsed_hash[:publisher] = "ISO"
@@ -77,14 +83,20 @@ module Pubid
77
83
  end
78
84
  end
79
85
 
80
- # Instantiate the identifier based on the typed stage
81
- identifier = locate_identifier_klass(parsed_hash).new
82
-
83
- # For French GUIDE entries: "Guide ISO/CEI 37:1995"
86
+ # For French GUIDE entries: "Guide ISO/CEI 37:1995". The rename
87
+ # must happen BEFORE class selection — locate_identifier_klass
88
+ # reads :type_with_stage, and the tree of a guide-first spelling
89
+ # carries an empty :type_with_stage plus the type under
90
+ # :type_with_stage_fr, so deferring the rename selected the
91
+ # default International Standard class for every guide-first
92
+ # (and Cyrillic "Руководства ИСО …") reference.
84
93
  if type_with_stage_fr = parsed_hash.delete(:type_with_stage_fr)
85
94
  parsed_hash[:type_with_stage] = type_with_stage_fr
86
95
  end
87
96
 
97
+ # Instantiate the identifier based on the typed stage
98
+ identifier = locate_identifier_klass(parsed_hash).new
99
+
88
100
  # For DirectivesSupplement, rename :publisher to :supplement_publisher
89
101
  if identifier.is_a?(Identifiers::DirectivesSupplement) && parsed_hash[:publisher]
90
102
  parsed_hash[:supplement_publisher] = parsed_hash.delete(:publisher)
@@ -112,7 +124,7 @@ module Pubid
112
124
  identifier.type = default_typed_stage.to_type
113
125
  end
114
126
 
115
- identifier
127
+ all_parts ? identifier.to_all_parts : identifier
116
128
  end
117
129
 
118
130
  def handle_key(identifier, key, value)
@@ -313,3 +325,5 @@ module Pubid
313
325
  end
314
326
  end
315
327
  end
328
+
329
+ Pubid::Iso::Builder.prepend(Pubid::Builder::AllPartsWrap)
@@ -8,6 +8,8 @@ module Pubid
8
8
  # ISO Publisher with copublisher support
9
9
  # Examples: ISO, ISO/IEC, ISO/IEC/IEEE
10
10
  class Publisher < Lutaml::Model::Serializable
11
+ include ::Pubid::SubsetMatch
12
+
11
13
  attribute :publisher, :string, default: -> { "ISO" }
12
14
  attribute :copublisher, :string, collection: true
13
15
 
@@ -11,6 +11,12 @@ module Pubid
11
11
  attribute :copublishers, ::Pubid::Iso::Components::Publisher,
12
12
  collection: true
13
13
 
14
+ # ISO prints "(all parts)" and has a series URN, so it has its own
15
+ # all-parts class.
16
+ def self.all_parts_class
17
+ Identifiers::AllParts
18
+ end
19
+
14
20
  # The publisher implied when none is serialized. ISO for most types;
15
21
  # publisher-less types (IWA) override this to nil.
16
22
  def self.default_publisher
@@ -137,20 +143,6 @@ module Pubid
137
143
  # unique typed-stage `code` under "stage" and recompute the rest on
138
144
  # load. _type already pins the document type.
139
145
  map "stage", with: { to: :stage_to_kv, from: :stage_from_kv }
140
- # Omit the `false` default; only the meaningful `true` is serialized.
141
- map "all_parts", with: { to: :all_parts_to_kv, from: :all_parts_from_kv }
142
- end
143
-
144
- def all_parts_to_kv(model, doc)
145
- return unless model.all_parts
146
-
147
- doc.add_child(
148
- Lutaml::KeyValue::DataModel::Element.new("all_parts", true),
149
- )
150
- end
151
-
152
- def all_parts_from_kv(model, value)
153
- model.all_parts = value
154
146
  end
155
147
 
156
148
  # Serialize typed_stage as just its unique code (e.g. "is", "dis",
@@ -312,7 +304,10 @@ module Pubid
312
304
  when :mr_string
313
305
  Pubid::Parsers::MrString.parse(string)
314
306
  else
315
- parsed = Pubid::Iso::Parser.new.parse(string)
307
+ # R1 parser swap: the baked PG artifact is the identifier
308
+ # parser of record; the Builder consumes the same attribute
309
+ # hash it always has.
310
+ parsed = Pubid::Parg::Backend.parse(:iso, string)
316
311
  Pubid::Iso::Builder.new.build(parsed)
317
312
  end
318
313
  end
@@ -0,0 +1,19 @@
1
+ # frozen_string_literal: true
2
+
3
+ module Pubid
4
+ module Iso
5
+ module Identifiers
6
+ # Every part of one ISO document: "ISO 9000 (all parts)".
7
+ #
8
+ # The URN is the series URN: the document URN without its stage, plus
9
+ # the "ser" slot.
10
+ class AllParts < ::Pubid::Iso::Identifier
11
+ include ::Pubid::AllParts
12
+
13
+ def to_urn
14
+ "#{identity.exclude(:stage, :typed_stage).to_urn}:ser"
15
+ end
16
+ end
17
+ end
18
+ end
19
+ end
@@ -89,8 +89,10 @@ format: nil, stage_format_long: nil, with_date: nil, **opts)
89
89
  def to_supplement_s(lang: :en, lang_single: false, with_edition: false,
90
90
  format: nil, stage_format_long: nil, with_date: nil, **_opts)
91
91
  date_str = if date
92
- month_part = date.month ? "-#{date.month}" : ""
93
- ":#{date.render}#{month_part}"
92
+ # Components::Date#render already carries the month
93
+ # (and day) when present — appending a month_part
94
+ # here doubled it (":2016-05-05-05").
95
+ ":#{date.render}"
94
96
  else
95
97
  ""
96
98
  end
@@ -4,6 +4,7 @@ module Pubid
4
4
  module Iso
5
5
  module Identifiers
6
6
  autoload :Addendum, "#{__dir__}/identifiers/addendum"
7
+ autoload :AllParts, "#{__dir__}/identifiers/all_parts"
7
8
  autoload :Amendment, "#{__dir__}/identifiers/amendment"
8
9
  autoload :Corrigendum, "#{__dir__}/identifiers/corrigendum"
9
10
  autoload :Data, "#{__dir__}/identifiers/data"
@@ -27,7 +27,10 @@ module Pubid
27
27
  private
28
28
 
29
29
  def parse_with_builder(string)
30
- parsed = Pubid::Iso.parser.parse(string)
30
+ # R2 ingestion hook: normalized strings reach the model through
31
+ # the same parser of record as Identifier.parse (the baked PG
32
+ # artifact), so every Tier-3 normalization feeds one grammar.
33
+ parsed = Pubid::Parg::Backend.parse(:iso, string)
31
34
  Pubid::Iso.builder.build(parsed)
32
35
  end
33
36
 
@@ -46,7 +46,6 @@ module Pubid
46
46
  end
47
47
 
48
48
  # All parts notation (if applicable)
49
- result << " (all parts)" if identifier.class.attributes.key?(:all_parts) && identifier.all_parts
50
49
 
51
50
  result
52
51
  end
@@ -0,0 +1,115 @@
1
+ # ITU flavor notes
2
+
3
+ ITU grammar, versions, annexes, reports and identity surfaces.
4
+
5
+ These notes were part of the root `CLAUDE.md`. Read them before you change `lib/pubid/itu/` or `spec/pubid/itu/`. The root file keeps the cross-flavor contract that every flavor obeys.
6
+
7
+ - **ITU Questions / Handbooks / N-way combined**: `Pubid::Itu` models study-group **Questions** (`Identifiers::Question`, `_type: pubid:itu:question` — numeric `ITU-R 234-1/7:` and letter-series `ITU-R P.3/BL/7`, `ITU-R S.[4/BL/2]:`; the `/BL` segment, brackets and trailing `:` are carried as `has_bl`/`bracketed`/`has_colon` booleans + a `study_group` string, all reconstructed by `render_base`) and **Handbooks** (`Identifiers::Handbook`, `ITU-R 23.HDB`; the `.HDB` marker is implied by `_type`). Parser disambiguation is by what follows a `/`: **digits** → a Question study-group; a **series** (letters) → a combined designation — so the new rules sit before `with_series`/`without_series`. **Combined (joint) recommendations are N-way**: the primary designation stays on the base `series`/`code` (keeps `root.number` non-empty for relaton-index) and each *additional* designation is a `Components::Designation` in `CombinedIdentifier#combined` (one for a dual `G.780/Y.1351`, two+ for a triple `G.780/Y.1351/Z.1362`); the parser's `combined_designation` rule is `repeat(1)`. **`#number` at the root**: ITU stores the document number on `code`, but `Pubid::Itu::Identifier#number` delegates to `code&.number` so `id.root.number.to_s` is the key relaton-index sorts/bsearches on (non-empty for every type). The delegation is serialization-neutral (the flat block maps `number` via `number_to_kv`/`number_from_kv` off `code`, never the inherited `number` attribute); Supplement/Amendment/Corrigendum/Errata keep their own `:string` `number` ordinal (it overrides the reader), and their base document number is reached via `root.number`.
8
+
9
+ - **ITU `(V##)` versions, labelled annexes, and class-strict supplement `==`**: three ITU-T forms relaton's `relaton-data-itu` crawler needs (~930 data files had no index row because they failed to parse). **(1) `version`** — `ITU-T H.264 (V14) (08/2021)`: a plain **`attribute :version, :string`** on the shared `Pubid::Itu::Identifier` (safe — `::Pubid::Identifier` declares no `version`, so this is a *new* attribute, not the `number`/`stage` retype landmine), fed by the `version_part` rule (`space >> "(V" >> digits >> ")"`) spliced in **immediately before `date_part.maybe`** in all four document rules (`with_series`/`without_series`/`base_with_series`/`base_without_series`) — version always precedes the date in the corpus, and `date_part` requires digits after `(` so the two never compete. Rendered `" (V#{version})"` between code and date by `render_base` **and** by `CombinedIdentifier#to_s` (which overrides `to_s` and does not call `render_base` — miss it and the version is silently dropped on a joint id); mapped in `StandardSerialization`; compared in both `==` overrides. **URN and MR strings stay version-blind by design** (the ITU URN convention defines no version segment), so `(V13)`/`(V14)` share a URN. **(2) Labelled annex** — `ITU-T A.23 Annex A (06/2014)` is its **own** leaf class `Identifiers::AnnexOfRecommendation` (`_type: pubid:itu:annex-of-recommendation`), *not* an extension of `Identifiers::Annex`: that models the structurally different label-less "Annex to ITU OB No. 1000" (prefix rendering + i18n templates) whose `_type` is already persisted in index rows. Shape mirrors `Supplement`: polymorphic `base` + `:string number` (the label — `A`, `F3`, the one-off `C+`) + its own date/language, a minimal `key_value` block (never re-emitting the base's sector/series/code), `render_base` → `"<base> Annex <label>"`, `to_urn` → `"<base urn>:annex:<label>"`, `#root` walking `base` so `root.number` is the annexed document's number. Grammar: `annex_body` wraps the annexed document in **`base.as(:base)`** — load-bearing, because the base may carry its own date (`ITU-T X.692 (2002) Annex E (03/2002)`) and two `date_part`s flattened into one hash would make Parslet silently keep only the last `:year`. `annex_identifier` sits in `rule(:identifier)` **after `supplement_identifier`** (an annex can carry a trailing `Cor.`/`Err.`/`Amd.`, and PEG ordered choice never re-enters once an alternative succeeds) and **before `with_series`** (which would match the `ITU-T A.23` prefix and then fail on the unconsumed ` Annex A`); `supplement_with_base` takes `(annex_body.as(:base) | base.as(:base))` so `ITU-T G.729 Annex B (1996) Cor. 3 (03/2001)` nests annex-inside-supplement. `Builder#build_supplement`'s fallback `Components::Sector.new` is now guarded on `data[:sector]` — an annex base keeps sector/series on *its* base, and the unguarded fallback raised `Invalid sector: `. **(3) `Supplement#==`** now uses **`instance_of?(self.class)`** (so `Suppl. 2` ≠ `Amd 2` ≠ `Cor. 2`, symmetrically — `is_a?` made a Supplement equal an Amendment one-way) and compares `sector`/`series` **only when `base` is nil**. Both halves matter: the series-only form (`ITU-T A Suppl. 2`) holds its whole identity in sector+series, so ignoring them made every series' `Suppl. 2` equal (an index search returned 42 rows, and `eql?`/`hash` collapsed them in a `Set`); but a *based* supplement copies sector/series from its base while the key_value block deliberately does **not** serialize them, so comparing them unconditionally would break `parse(s) == from_hash(parse(s).to_hash)` — the very lookup this fixes. The now-identical `==` overrides on `Amendment`/`Corrigendum`/`Errata` were deleted. `language` is deliberately **not** compared (an index row without a language suffix must still match a reference that has one). Locked by `spec/pubid/itu/identifiers/{version,annex_of_recommendation,supplement}_spec.rb` + rows in `root_number_spec.rb`/`serialization_spec.rb` and the new `spec/fixtures/itu/identifiers/pass/{version,annex_of_recommendation}.txt`. **Rendering/URN/MR consequences of the annex wrapper:** `CombinedIdentifier` had its rendering in `to_s`, which the annex (composing `base.render_base`) bypassed — silently dropping the `/Y.1351` half, so `G.780/Y.1351 Annex A` and `G.780/Z.1362 Annex A` printed identically (and `to_s` is relaton's document number *and* output filename). It now renders in **`render_base`**, with the inherited `to_s` adding the language suffix and common-text twin. The annex's `to_urn` appends its **own date** after the label (`urn:itu:t:A.23:annex:a:06/2014`) — the base's date rides inside `base_urn`, so omitting the annex's would collide editions — and it defines **`mr_supplement_suffix`** (`annex.<label>[.<year>]`, `C+`→`cplus`) so the shared MrString renderer recurses into the annexed document instead of collapsing every annex onto a bare `itu.<date>`. Because the annex keeps sector/series/code on *its* base, `build_supplement` reaches identity through **`base.root`** when the base has none, so a supplement *of* an annex keys like any other supplement. **Version matching semantics:** `version` is in `==`, so a bare `ITU-T H.264` does **not** match `ITU-T H.264 (V14) …` under `ignore: %i[year month]` — version is a *separable trailing component* like ETSI's (`partial_ref_spec` now lists ITU as `omits: %i[date version]`), so callers matching a partial ref must add `:version` to `ignore`; it is deliberately **not** folded into the `:year`→`:date` alias. **Also fixed here:** `spec/pubid/itu/fixtures_spec.rb`'s glob had one `..` too many (repo-root `fixtures/`), so the whole ITU fixture round-trip spec silently iterated an empty file list. **Out of scope at the time, both since fixed:** `ITU-T V.25 ter Annex A (08/1996)` (the space-separated `ter` gap — closed by the code-suffix work below), and `Amendment#to_s`'s missing dot. (hand-off: itu-version-annex-and-supplement-matching.)
10
+
11
+ - **ITU-T unparseable print forms (the `relaton-data-itu` residue)**: 668 of the 755 unindexed `relaton-data-itu` records were ITU-T `rec_name` spellings the grammar had no rule for; all 668 now parse, render and satisfy relaton's gate (`from_hash(to_hash) == to_hash`), with **zero** change to any of the 20,728 already-indexed ids. **The single most important invariant when touching this area: every new rendering flag is a `:boolean` named for the RARE form with `default: -> { false }`**, because the canonical `to_hash` strips default-valued attributes — a flag named for the common case would add a key to every already-published index row. Ten constructs, grouped by where they live:
12
+ **(1) Supplement family.** `supplement_type` gained `Add.`/`Add` → the new leaf `Identifiers::Addendum < Supplement` (`_type: pubid:itu:addendum`), and `Builder#build_supplement`'s `case` gained an **`else raise ArgumentError`** — it previously fell through to a nil class and died with `undefined method 'new' for nil` one frame later. The near-identical `to_s` on `Supplement`/`Amendment`/`Corrigendum`/`Errata` collapsed into one **`Supplement#render_supplement(label)`**; each subclass now supplies only its label. `supplement_number` made the space optional (`ITU-T E Suppl.1`, `ITU-T D.211 Suppl.1` → `number_glued`) and the ordinal dotted (`ITU-T M Suppl. 1.1` → `number` is the whole `"1.1"`). `number_glued` is serialized but deliberately **not** in `==` — `Suppl.1` and `Suppl. 1` are one document. `chained_supplement`'s inner alternation gained `supplement_series_only` so a supplement *of* a series-only supplement parses (`ITU-T G Suppl. 39 (2006) Err. 1 (08/2006)`).
13
+ **(2) `Technical Cor.`** (158 records, the largest bucket — and the one the hand-off mislabelled as a series-code form). A `technical_marker` rule at the **head of `supplement_type`**, not folded into its alternation, so the type token stays the plain `Cor.` the builder's `case` maps; a `technical` boolean on `Corrigendum` renders and compares it. The slash-joined pair `ITU-T X.680 (1994) Amd. 1/Technical Cor. 1 (12/1997)` is a new **`slash_chained_supplement`** rule (a Corrigendum whose `base` is the Amendment) placed **before `supplement_with_base`** in the alternation — that alternative alone matches the `… Amd. 1` prefix and leaves `/Technical Cor. 1 …` unconsumed, which fails the whole parse with no re-entry.
14
+ **(3) Code suffixes — the highest-blast-radius change, since it fires on *every* recommendation parse.** A shared `code_suffixes` rule (`series_suffix_spaced.maybe >> qualifier_spaced.maybe`) spliced after `code` and **before `combined_suffixes`** (the corpus attaches the primary's suffix before the `/`: `ITU-T D.301 R/F.66`) in all four document rules. It cannot live *inside* `code`, or its leading space would also be offered to ` Suppl.`/` Annex` and have to backtrack out of a rule the wrappers depend on. `Components::Code` gained `series_suffix_spaced` (the spaced `ITU-T E.250 bis` vs the spec-locked glued `X.50bis` of issue #231 — the word itself is stored whitespace-free either way, so the serialized value stays comparable), plus `qualifier` + `qualifier_glued` for the trailing letter (`ITU-T D.200 R`, `ITU-T Q.2931 B`, `ITU-T R.38 A`, glued `ITU-T D.502R`, lowercase `ITU-T I.256.2a`). **Both guards on `qualifier_letter` are load-bearing**: the `A-R`/`a-r` cap keeps it off the `S` of ` Suppl.` and the `V` of a bare ` V2`, and the trailing `match["A-Za-z0-9"].absent?` is what stops ` Amd.`/` Add.`/` Annex A`/` App.`/` Cor.`/` Err.` being read as a qualifier. It cannot collide with `language` (dash-attached) or `date_part`/`version_part` (both need `(` after the space). Because `Code#to_s` can now contain a space, **`Code#compact_s`** (`to_s.delete(" ")`) was added and is what `urn_generator` interpolates — a space is not admissible in a URN segment, and dropping the suffix instead would collapse `D.200` and `D.200 R` onto one URN. The same slots were added inside `combined_designation` (positionally, so the suffix lands on the designation it printed against: `ITU-T D.300 R/E.282 R` sets both, `ITU-T E.211/Q.11 quater` only the last) — which required the parallel edits to `Builder#build_designations` **and** `CombinedIdentifier#combined_to_kv`/`combined_from_kv`, whose row hash is hand-built. Neither spacing flag is in `Code#==`: the two spellings of one number are one document. **`mr_number_with_part` also had to learn both suffixes** (glued to the number, `x-50bis`/`d-200r`) — it read only number/subseries/parts, so every qualified variant collapsed onto its base's MR slug and, worse, `Q.2931 B` and `Q.2931 C` onto each other. This changes the MR string of the handful of pre-existing glued-`bis` ids (`X.50bis`: `itu.t.x-50` → `itu.t.x-50bis`), which is the point — they were wrongly sharing a slug with `X.50`.
15
+ **(4) Version spellings.** `version_part` gained a bare branch, so `v.1`/`v10`/`V2` join the canonical `(V14)` in the same `version` attribute. This is the **one deliberate non-byte-exact normalisation** — all render as `(V##)` — so those strings live in `version_spec.rb` with explicit expectations and **not** in the byte-exact pass fixtures. Safe against the qualifier rule precisely because `V` is outside its `A-R` cap. relaton can now drop its own `v10` → `(V10)` normalisation.
16
+ **(5) Appendix.** `Identifiers::AppendixOfRecommendation` (`_type: pubid:itu:appendix-of-recommendation`) mirrors `AnnexOfRecommendation` one-for-one — including wrapping the appendixed document in **`base.as(:base)`** so the base's own date (`ITU-T G.722 (1988) App. IV (11/2006)`) lands a level down instead of colliding, and defining `mr_supplement_suffix` so the shared MR renderer recurses. Its alternation slot is the annex's reasoning verbatim: **after `supplement_identifier`** (an appendix can carry a trailing `Amd.`/`Err.`), **before `with_series`**. Three registrations are easy to miss and all three are needed: the `identifiers.rb` autoload, the `Builder#build` branch, and widening `urn_generator`'s `AnnexOfRecommendation` check — that check exists because a wrapper whose identity sits one level deeper otherwise emits `urn:itu:itu`. A `material` attribute carries the companion-artefact records (`App. II test vectors`, `App. I Software`), and is in the URN so they stay distinct from the appendix itself.
17
+ **(6) Series groups vs series-code documents — a corpus heuristic, not an ITU rule.** `series_group` (`E-100`, `E100-300`, `G-100`, `Q-500`) is capped at **one** leading letter; `series_code_body` (`EMC-5`, `MES-2`, `QOS-2`, `IMPL-8`, `SEC-QKD`) requires **two or more**. That split is what keeps them from shadowing each other, and it holds across the whole corpus but would misfile a future `AB-100` group or `E-QKD` document — `series_group_spec.rb` locks both sides so a change is visible. `series_group` must be tried **before** `series`, whose greedy non-backtracking `letter.repeat(1,3)` would take the `E` of `E-100` and then fail on the required space with no retry at a shorter length. `series_code_body` is a flat two-token shape (series + dash + alphanumeric number) rather than an optional middle segment, because a greedy `letter.repeat(2)` + optional `-QKD` + required dash fails at end-of-input with no backtracking into the satisfied `.maybe`; it sits **last** in `base` and after `with_series` in `identifier`, where nothing can reach it by accident. The number stays in `code.number` (not `"EMC-5"` in `series`) precisely so **`root.number`** — the key relaton-index bsearches on — is `"5"`/`"QKD"` rather than nil. `series_dash` drives the `-` vs `.` join in `render_base`; the builder's OB check gained a `series_dash.nil?` guard so a hypothetical `ITU-T OB-1` cannot be misrouted into the Operational Bulletin branch. The literal word `series` is its own `series_word` boolean and **cannot** live in the series token, because it also follows a *dotted* code (`ITU-T E.1100 series Suppl. 1`) where folding it in would corrupt the value every index row is keyed on; `Supplement` emits it only when base-less (guarded like `sector`/`series`).
18
+ **(7) Tail.** `attachment` boolean (`ITU-T H.350 attachment`), `range_end` string (`ITU-T Q.120-Q.139` — no conflict with `parts`, which needs digits after the dash), and `Identifier.parse` now collapses runs of whitespace (`ITU-T D.271 (10/2016)`) **after** the shared `Pubid::MAX_INPUT_LENGTH` guard, never before.
19
+ **The cross-cutting lesson (three review findings were all this same mistake): an identity-bearing marker must reach ALL THREE identity surfaces — `to_s`, `to_urn` and `to_mr_string` — not just `==`.** A marker that only reaches `==` still lets two distinct documents share a URN and an output slug. So `generate_base_urn` emits `series`/`attachment`/`to-<range_end>` segments and joins a series-code document with its dash (`urn:itu:t:EMC-5:2003`, not the ambiguous `EMC.5`); `generate_supplement_urn`'s base-less branch emits `series` too; and `mr_number_with_part` carries all five. `Builder#build_supplement` copies `series_word` up from the base alongside sector/series/code for the same reason — render/serialize/`==` all guard on `base.nil?`, so the copy is neutral there and only the MR slug sees it. `spec/pubid/itu/distinctness_spec.rb` is the forcing function: it asserts all four surfaces differ for every marker pair. **Two more review findings worth keeping:** `technical_marker` is bound to the **`Cor.` branch alone**, not the head of `supplement_type` — offered before any token it was accepted on `Technical Err. 1`/`Technical Amd. 1` and then silently dropped by the builder's Corrigendum-only guard, collapsing those onto a *different* document (they are now cleanly rejected); and `series_code_body` carries a `(str("OB") >> dash).absent?` guard, because `ITU-T OB-1` otherwise reached the Recommendation fallback whose `validate_ob_no_sector!` raises an `ArgumentError` that escapes `Identifier.parse`'s `Parslet::ParseFailed` rescue — turning a rejected input into a crash for callers. `Code#to_s` renders glued suffixes before spaced ones, mirroring the grammar (a fixed edition-word-first order printed `Q.11a bis` back as `Q.11 bisa`). **Two known pre-existing gaps, deliberately not fixed and both pinned by specs in `distinctness_spec.rb`:** (a) the ITU URN convention encodes no supplement *type*, so `Amd. 1`/`Cor. 1`/`Err. 1`/`Add. 1` of one base share a URN and MR slug; (b) `build_supplement` copies sector/series/code down from the base but they are deliberately not serialized, so a supplement rebuilt through `from_hash` has none of them and its **MR slug collapses to `itu.<date>`** while every other surface stays symmetric. Both are true on `main` for `Amd.`/`Cor.`/`Err.` before this change — the new types simply inherit them, and `to_s`/`to_hash`/`==` distinguish everything correctly, which is what relaton's gate uses. Both would be fixed by giving `Supplement` an `mr_supplement_suffix` so the shared MR renderer recurses into `base` (as `AnnexOfRecommendation` already does), which changes every existing ITU supplement MR string — a deliberate call, not a tag-along. Also fixed here: `render_base` gained an `elsif series` branch, so a code-less identifier no longer renders a dangling trailing dot. Locked by `spec/pubid/itu/{code_suffixes,combined_suffixes,series_group,tail_forms}_spec.rb`, `spec/pubid/itu/identifiers/{addendum,appendix_of_recommendation,supplement_spelling,technical_corrigendum}_spec.rb`, and 7 new `spec/fixtures/itu/identifiers/pass/*.txt` files drawn from the real corpus. **Still out of scope (both declared not-ours by the hand-off):** ~47 malformed ITU-R docids left by a decommissioned crawler (`ITU-R BO`, `ITU-R M.5-BL-13` — series with no number; dataset bugs, not identifiers), and the 4 space-for-dot `ITU-T G 231` spellings relaton normalises on its side. (hand-off: itu-t-unparseable-forms.)
20
+
21
+ - **ITU-R Reports (`Identifiers::Report`) and the class-strict `Identifier#==`**: ITU-R **Reports** are a publication series that numbers **independently** of Recommendations, so `ITU-R BT.2020-1` names *two* real, both-current documents (Report BT.2020-1, objective quality assessment, 2000; Recommendation BT.2020-1, UHDTV parameter values, 06/2014). 984 of 1,001 report records in `relaton-data-itu` were indexed as `pubid:itu:recommendation`, and 52 editions across 30 numbers (BT.2020, M.2083, M.2134 …) collided outright — the dataset keys files by identifier, so whichever was crawled last silently overwrote the other. ITU disambiguates with the leading word, so the identifier must too. **The leaf** `Identifiers::Report` (`_type: pubid:itu:report`) is shaped exactly like `Recommendation` — same attributes, same `StandardSerialization` flat block, so only `_type` distinguishes the two hashes — and adds just `render_base` (`"Report #{super}"`) and `mr_type` (`[super, "report"].join(".")` → `itu.r.report.bt-2020-1`). **Grammar**: a captured `report_word` marker (unlike `itu_prefix`'s decorative, *uncaptured* `"Recommendation "` literal, which stays as it is) plus a single `report_body` whose `(series >> dot).maybe` covers the with- and without-series shapes at once; `base_report` accepts **both** spellings (`Report ITU-R BT.2020-1` and the downstream `ITU-R Report BT.2020-1`), and `to_s` renders ITU's own leading form for either — so the infix spelling lives in `report_spec.rb` with an explicit expectation, **not** in the byte-exact pass fixtures. Two deliberate exclusions from `report_body`: **no `combined_suffixes`** (a joint `Report X/Y` would reach `Builder#build`'s `combined` branch, build a `CombinedIdentifier` and *silently drop the marker* — a clean parse failure is better), and a **`(str("OB") >> dot).absent?` guard** (an Operational Bulletin is cross-bureau; without it `Report ITU-T OB.1` routes to `SpecialPublication`, marker dropped, or trips `validate_ob_no_sector!`, whose `ArgumentError` escapes `Identifier.parse`'s `Parslet::ParseFailed` rescue). `base_report` is **first** in `rule(:base)` — which is what lets a supplement/annex/appendix *wrap* a Report — and order-free there because `itu_prefix` admits only `Recommendation `/`ITU` and `letter` is `[A-Z]`, so neither `series` nor `series_code_body` can consume `Report`; `report_identifier` sits in `rule(:identifier)` **after** supplement/annex/appendix (which match the longer wrapped form) and **before** `with_series`/`without_series` (which would match the infix spelling's `ITU-R` prefix and then fail on the unconsumed tail). The builder picks the class from `data[:report_marker]` in the existing Recommendation fallthrough; `build_supplement`/`build_annex_of_recommendation` need **no** change, since they recurse via `build(data[:base])`. **`Pubid::Itu::Identifier#==` is now class-strict** (`other.instance_of?(self.class)`, the rule `Supplement#==` already used) — the ITU *type* is part of the identity, not just sector/series/number. `self.class` rather than a hard-coded class keeps it **symmetric in both directions** (an `is_a?` guard would have made a Recommendation equal a Report one-way, and `#matches?` — which is `exclude(*ignore) == other.exclude(*ignore)` — resolve one to the other). Strictly stricter, so it can only turn `true` into `false`; it also closes the pre-existing one-way holes against `CombinedIdentifier` and `SpecialPublication`. **Consequences of the shared guard, both pinned by `spec/pubid/itu/type_strict_equality_spec.rb`:** comparisons that used to be *asymmetric* are now consistently false — most notably a bare primary designation vs its joint recommendation (`ITU-T G.780` vs `ITU-T G.780/Y.1351`, which was `true` one way and `false` the other on `main`, so no caller could rely on it either way). A caller matching a bare primary against a joint document must narrow on `root.number` (which still agrees) rather than on `==`. Per the cross-cutting ITU lesson, the marker reaches **all** identity surfaces: `to_s`, `generate_base_urn` (a `report` segment right after the sector — `urn:itu:r:report:BT.2020-1:2000`, which `UrnParser` also reads back), and the MR slug — including **`Supplement#mr_type`**, which is load-bearing and easy to miss: a supplement has no `mr_supplement_suffix`, so the shared MR renderer slugs it **flat** from the sector/series/code `build_supplement` copied up from its base and *never consults the base's class*, which made `Report ITU-R BT.2020-1 Suppl. 1` and `ITU-R BT.2020-1 Suppl. 1` share one slug (i.e. one output filename — the very overwrite this type prevents). It appends `report` when **`root`** is a Report, so a supplement of an *annex* of a Report is covered too. **`PREFIXES` stays `["ITU"]`** — the leading `Report` token is deliberately not registered for prefix routing (a generic English word; the mixin's ambiguous-token exclusion), and the infix spelling routes on `ITU` as before. **Backwards compatibility is exact**: a bare `ITU-R BT.2020-1` still builds a `Recommendation` with an unchanged hash/`to_s`/URN — nothing in the grammar or the builder fires without the marker, so none of the 20,728 published index rows move. Locked by `spec/pubid/itu/identifiers/report_spec.rb`, a `distinctness_spec.rb` pair, rows in `root_number_spec.rb`/`serialization_spec.rb`/`urn_parser_spec.rb`, and `spec/fixtures/itu/identifiers/pass/report.txt`. **Out of scope (relaton's side, sequenced after this):** `DataParserR`/`DataCrawlerR` emitting the new docid form, `Relaton::Itu::Pubid`'s reference parser learning a `report` rule, and the `relaton-data-itu` filename-namespace migration + re-crawl that actually recovers the 34 missing Recommendations and 18 missing Reports. (hand-off: itu-r-report-identifier-type.)
22
+
23
+ ## From the root note "Parse-failure error contract — uniform across every flavor"
24
+
25
+ **ITU's length guard was an early `return`, not a raise, and its comment blamed the wrong caller.** `Identifier.normalize_whitespace` (`lib/pubid/itu/identifiers/base.rb`) carried
26
+
27
+ ```ruby
28
+ return identifier if identifier.length > ::Pubid::MAX_INPUT_LENGTH
29
+ ```
30
+
31
+ under a note claiming "`Pubid.parse` rejects it before this point". That was wrong twice: `Pubid.parse` never routes a human-readable string to a flavor at all (`lib/pubid.rb` raises `ArgumentError: No flavor specified` for anything that is not an MR string or a URN), and ITU had no raising guard of its own — so an over-long string reached the ITU parser through both `Pubid::Itu.parse` and `Pubid::Itu::Identifier.parse`. The guard now lives in `self.parse` as the standard inline pair, and the comment says what actually protects the method.
32
+
33
+ **ITU was also the delegate in the most-visible class leak.** `Pubid::Iso.parse("ITU-T G.711")` detects an MR-shaped string, routes through `Pubid::Parsers::MrString`, and lands in `Pubid::Itu::Identifier.parse` — which raised `RuntimeError`. So `Pubid::Iso.parse` raised **two different classes** depending on its input, and a relaton-cli caller rescuing `Parslet::ParseFailed` got a raw backtrace for the ITU-shaped half. Converting ITU fixed the ISO symptom; `spec/pubid/parse_error_spec.rb` pins that exact route by name, because a generic junk string never reaches it.
34
+
35
+ - **The five supplement types rendered plain under `to_s(annotated: true)`.**
36
+ `Pubid::Itu::Identifier#to_s` annotates, but `Addendum`, `Amendment`,
37
+ `Corrigendum`, `Errata` and `Supplement` each override it and none calls
38
+ `super` — they all funnel through `Supplement#render_supplement(label)`.
39
+ The wrap cannot live in that helper, which takes a label and no options,
40
+ so each of the five `to_s` is a one-line
41
+ `annotate_plain_render(render_supplement("…"), **opts)`.
42
+
43
+ `render_supplement` starts from `base.to_s` with no options, so the base
44
+ arrives plain and the single outer annotation covers the whole string —
45
+ the base's own tokens are reached by `Annotator#emit_tokens` walking
46
+ `base`.
47
+
48
+ ## Contribution (Temporary Document) — pubid#340
49
+
50
+ `Pubid::Itu::Identifiers::Contribution` models the working documents a study
51
+ group circulates, mirroring pubid-itu 1.15's `Identifier::Contribution`
52
+ (`%{series}-C%{number}`): "ITU-R SG17-C1000", sector- and language-suffixed
53
+ per the existing rules ("ITU-T SG17-C1000-E"). The grammar entry sits between
54
+ `with_series` and `series_code_identifier` — the **-C marker (dash + "C" +
55
+ digits) is the discriminator**: a series-code document's post-dash number
56
+ starts with digits or is all letters, never "C"+digits, so the two dash
57
+ shapes cannot shadow each other in either direction ("ITU-T EMC-5" still
58
+ builds a Recommendation, locked by a spec). The builder branch carries the
59
+ `:contribution_marker`; `root.number` reaches the C-number through the shared
60
+ `code.number` reader, so relaton-index keys it normally.
61
+
62
+ **`locate_type` was broken for every leaf, not just this one**: no ITU class
63
+ ever defined a `type` hash, so `Pubid::Itu.locate_type` raised `NoMethodError`
64
+ on any call. The base now derives the key from the class name
65
+ (`Contribution` → `:contribution`, `AnnexOfRecommendation` →
66
+ `:annex_of_recommendation`) — plain `gsub` camel→snake, because ActiveSupport's
67
+ `underscore` is not a dependency. metanorma-itu constructs through this
68
+ lookup; its flavor-local `pubid_contribution.rb` render override can be
69
+ deleted once it migrates.
70
+
71
+ ## relaton's query forms — RR, OB sector, publication ids
72
+
73
+ Four forms relaton's `Relaton::Itu::Pubid` parsed and `Pubid::Itu` did not
74
+ (hand-off itu-relaton-query-forms). None of them occurs in the published
75
+ `relaton-data-itu` index — no `series: RR`, no OB row, no six-digit part — and
76
+ a replay of all 24,382 rows and every ITU pass fixture showed 0 changes.
77
+
78
+ - **Radio Regulations** — `Identifiers::RadioRegulations`
79
+ (`pubid:itu:radio-regulations`): `ITU-R RR`, `ITU-R RR (2020)`, and the URL
80
+ spelling `ITU-R RR-2020`, which used to build a *wrong* Recommendation
81
+ (series `RR`, number `2020`) with no error. "RR" is the series; there is no
82
+ code, so `#number` returns the series and `root.number` is `"RR"`. The rule
83
+ ends in `any.absent?`: it sits before `with_series`, and PEG never re-enters
84
+ the alternation, so a partial match on `ITU-R RR.1` must fail inside it.
85
+ - **Operational Bulletins keep their sector.** This reverses the old
86
+ "cross-bureau, sector must not be set" rule: `validate_ob_no_sector!` is
87
+ gone, `ITU-T OB.1096 (2016)` renders back as it is, and the sector-less
88
+ `ITU OB No. 1096` is unchanged. The long forms with a sector (`ITU-T OB No.
89
+ 1096`) now normalise to `ITU-T OB.1096`. `No.` stays in the default
90
+ render — it is how ITU's bulletin site and pubid v1 write it — and
91
+ metanorma-itu's `ITU OB 1000` / `Annex to ITU OB 1000` (its i18n template
92
+ omits `No.`) is accepted on parse (`ob_bare_body`) and rendered with it.
93
+ The sector is a spelling, not
94
+ identity: `SpecialPublication#==` skips it, the URN keeps `urn:itu:itu:…`
95
+ and `mr_type` stays nil, so both spellings are one bulletin on every
96
+ surface except `to_s`/`to_hash`. The printed date `- 15.III.2016` sets
97
+ `date.day`, which only this form does, so `render_ob_date` uses the day as
98
+ the spelling marker; `day_to_kv` emits only then.
99
+ - **`-YYYYMM` is a date, never a part.** `part` refuses a six-digit run with a
100
+ 19xx/20xx year and a 01–12 month (`yyyymm_shape`), and `id_date` reads it as
101
+ year+month. `-200313`, `-180001` stay parts. `ITU-T REC T.4` drops the
102
+ uncaptured `rec_word`; `T-REC-T.4-200307-I` is `publication_id`, last in
103
+ `identifier` (nothing else starts with a bare sector letter), and requires
104
+ the date. The trailing status letter (`I` in force, `S` superseded) is
105
+ parsed and dropped: it names the state of an edition, not the edition.
106
+ **`S` is also the Spanish language suffix**, so it is a status only inside
107
+ the full `T-REC-…` id (`id_status`); after an `ITU-T …-YYYYMM` print form
108
+ only `-I` is (`print_id_status`), and `ITU-T Z.100-199911-S` keeps language
109
+ `S`. The `-YYYYMM` date and `REC` word are also in `base_with_series`/
110
+ `base_without_series`: the part guard applies there too, so without them
111
+ `ITU-T G.989-200307 Amd 1` — a (wrong) Amendment on `main` — stopped parsing.
112
+ The day of an OB date reaches the URN (`…:15/03/2016`), since it is in `==`.
113
+ **Not done:** an RR supplement (`ITU-R RR (2020) Amd 1` fails; the
114
+ base-less `ITU-R RR Amd 1` still builds an Amendment on series `RR`), and
115
+ `ITU-R RR-E` is still a Recommendation numbered `E` — both as on `main`.