pubid 2.0.0.pre.alpha.12 → 2.0.0.pre.alpha.14
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/README.adoc +43 -1
- data/data/ieee/update_codes.yaml +17 -4
- data/data/nist/update_codes.yaml +7 -3
- data/data/parg/tables/bipm_groups.yaml +14 -0
- data/data/parg/tables/bipm_type_codes.yaml +5 -0
- data/data/parg/tables/bipm_type_names_en.yaml +6 -0
- data/data/parg/tables/bipm_type_names_fr.yaml +6 -0
- data/data/parg/tables/directives_supplements_typed_stages.yaml +3 -0
- data/data/parg/tables/directives_typed_stages.yaml +5 -0
- data/data/parg/tables/idf_typed_stages.yaml +27 -0
- data/data/parg/tables/idf_typed_stages_supplements.yaml +2 -0
- data/data/parg/tables/iec_typed_stages.yaml +130 -0
- data/data/parg/tables/iso_publishers.yaml +4 -0
- data/data/parg/tables/organizations.yaml +12 -0
- data/data/parg/tables/tc_types.yaml +42 -0
- data/data/parg/tables/typed_stages.yaml +114 -0
- data/data/parg/tables/typed_stages_supplements.yaml +64 -0
- data/data/parg/tables/wg_types.yaml +21 -0
- data/lib/pubid/adobe/builder.rb +2 -0
- data/lib/pubid/adobe/identifier.rb +11 -1
- data/lib/pubid/all_parts.rb +201 -0
- data/lib/pubid/all_parts_identifier.rb +19 -0
- data/lib/pubid/amca/CLAUDE.md +47 -0
- data/lib/pubid/amca/builder.rb +3 -5
- data/lib/pubid/amca/identifiers/base.rb +11 -1
- data/lib/pubid/amca/identifiers/publication.rb +13 -0
- data/lib/pubid/amca/parser.rb +2 -1
- data/lib/pubid/amca/renderer.rb +22 -33
- data/lib/pubid/amca/urn_generator.rb +21 -2
- data/lib/pubid/amca/urn_parser.rb +36 -10
- data/lib/pubid/ansi/builder.rb +6 -0
- data/lib/pubid/ansi/identifier.rb +1 -1
- data/lib/pubid/api/CLAUDE.md +23 -0
- data/lib/pubid/api/builder.rb +2 -0
- data/lib/pubid/api/identifier.rb +1 -1
- data/lib/pubid/api/parser.rb +8 -4
- data/lib/pubid/ashrae/CLAUDE.md +13 -0
- data/lib/pubid/ashrae/builder.rb +58 -14
- data/lib/pubid/ashrae/identifiers/base.rb +10 -1
- data/lib/pubid/ashrae/identifiers/errata.rb +14 -2
- data/lib/pubid/ashrae/identifiers/interpretation.rb +2 -10
- data/lib/pubid/ashrae/parser.rb +80 -39
- data/lib/pubid/ashrae/renderer.rb +32 -1
- data/lib/pubid/ashrae/urn_generator.rb +32 -9
- data/lib/pubid/asme/CLAUDE.md +25 -0
- data/lib/pubid/asme/builder.rb +16 -9
- data/lib/pubid/asme/components/code.rb +2 -0
- data/lib/pubid/asme/identifier.rb +1 -1
- data/lib/pubid/asme/identifiers/standard.rb +6 -1
- data/lib/pubid/asme/parser.rb +41 -14
- data/lib/pubid/astm/CLAUDE.md +9 -0
- data/lib/pubid/astm/builder.rb +2 -0
- data/lib/pubid/astm/components/code.rb +2 -0
- data/lib/pubid/astm/identifier.rb +1 -1
- data/lib/pubid/astm/parser.rb +4 -1
- data/lib/pubid/bipm/CLAUDE.md +11 -0
- data/lib/pubid/bipm/builder.rb +2 -0
- data/lib/pubid/bipm/identifier.rb +1 -1
- data/lib/pubid/bsi/CLAUDE.md +93 -0
- data/lib/pubid/bsi/builder.rb +13 -11
- data/lib/pubid/bsi/identifiers/addendum_document.rb +2 -0
- data/lib/pubid/bsi/identifiers/adopted_european_norm.rb +6 -54
- data/lib/pubid/bsi/identifiers/adopted_international_standard.rb +5 -22
- data/lib/pubid/bsi/identifiers/amendment.rb +36 -12
- data/lib/pubid/bsi/identifiers/bundled_identifier.rb +2 -0
- data/lib/pubid/bsi/identifiers/consolidated_identifier.rb +23 -26
- data/lib/pubid/bsi/identifiers/corrigendum.rb +29 -12
- data/lib/pubid/bsi/identifiers/expert_commentary.rb +6 -7
- data/lib/pubid/bsi/identifiers/national_annex.rb +18 -20
- data/lib/pubid/bsi/identifiers/root_identity.rb +31 -0
- data/lib/pubid/bsi/identifiers/set.rb +2 -0
- data/lib/pubid/bsi/identifiers/supplement_document.rb +2 -0
- data/lib/pubid/bsi/identifiers.rb +1 -0
- data/lib/pubid/bsi/parser.rb +8 -8
- data/lib/pubid/bsi/renderer.rb +20 -20
- data/lib/pubid/bsi/single_identifier.rb +1 -3
- data/lib/pubid/bsi/urn_generator.rb +28 -18
- data/lib/pubid/builder/base.rb +27 -0
- data/lib/pubid/calconnect/builder.rb +2 -0
- data/lib/pubid/calconnect/identifier.rb +5 -1
- data/lib/pubid/ccsds/builder.rb +2 -0
- data/lib/pubid/ccsds/identifier.rb +9 -1
- data/lib/pubid/cen_cenelec/CLAUDE.md +59 -0
- data/lib/pubid/cen_cenelec/builder.rb +6 -1
- data/lib/pubid/cen_cenelec/identifier.rb +12 -28
- data/lib/pubid/cen_cenelec/identifiers/amendment.rb +3 -10
- data/lib/pubid/cen_cenelec/identifiers/corrigendum.rb +3 -10
- data/lib/pubid/cen_cenelec/parser.rb +20 -5
- data/lib/pubid/cie/CLAUDE.md +58 -0
- data/lib/pubid/cie/builder.rb +2 -0
- data/lib/pubid/cie/components/language.rb +2 -0
- data/lib/pubid/cie/identifier.rb +1 -1
- data/lib/pubid/cie/parser.rb +9 -2
- data/lib/pubid/components/adoption.rb +2 -0
- data/lib/pubid/components/code.rb +2 -0
- data/lib/pubid/components/date.rb +8 -6
- data/lib/pubid/components/edition.rb +2 -0
- data/lib/pubid/components/iteration.rb +2 -0
- data/lib/pubid/components/language.rb +2 -0
- data/lib/pubid/components/locality.rb +2 -0
- data/lib/pubid/components/publisher.rb +2 -0
- data/lib/pubid/components/relationship.rb +2 -0
- data/lib/pubid/components/stage.rb +2 -0
- data/lib/pubid/components/supplement.rb +2 -0
- data/lib/pubid/components/type.rb +2 -0
- data/lib/pubid/components/typed_stage.rb +8 -0
- data/lib/pubid/conformance/checks.rb +1 -1
- data/lib/pubid/csa/CLAUDE.md +41 -0
- data/lib/pubid/csa/builder.rb +2 -0
- data/lib/pubid/csa/identifier.rb +19 -3
- data/lib/pubid/csa/parser.rb +25 -8
- data/lib/pubid/csa/renderer.rb +12 -12
- data/lib/pubid/csa/single_identifier.rb +17 -0
- data/lib/pubid/doi/builder.rb +2 -0
- data/lib/pubid/doi/identifier.rb +1 -1
- data/lib/pubid/easc/builder.rb +2 -0
- data/lib/pubid/easc/identifier.rb +10 -1
- data/lib/pubid/ecma/CLAUDE.md +28 -0
- data/lib/pubid/ecma/builder.rb +2 -0
- data/lib/pubid/ecma/identifier.rb +8 -1
- data/lib/pubid/etsi/CLAUDE.md +34 -0
- data/lib/pubid/etsi/builder.rb +2 -0
- data/lib/pubid/etsi/components/code.rb +6 -0
- data/lib/pubid/etsi/components/version.rb +2 -0
- data/lib/pubid/etsi/identifiers/base.rb +1 -1
- data/lib/pubid/etsi/identifiers/etsi_standard.rb +7 -0
- data/lib/pubid/evs/CLAUDE.md +58 -0
- data/lib/pubid/evs/builder.rb +2 -0
- data/lib/pubid/evs.rb +1 -1
- data/lib/pubid/gb/CLAUDE.md +140 -0
- data/lib/pubid/gb/builder.rb +7 -2
- data/lib/pubid/gb/identifier.rb +6 -4
- data/lib/pubid/gb/identifiers/all_parts.rb +17 -0
- data/lib/pubid/gb/identifiers.rb +1 -0
- data/lib/pubid/gb/renderer.rb +0 -1
- data/lib/pubid/gost/CLAUDE.md +64 -0
- data/lib/pubid/gost/builder.rb +3 -1
- data/lib/pubid/gost/identifier.rb +16 -1
- data/lib/pubid/gost/parser.rb +8 -1
- data/lib/pubid/iala/CLAUDE.md +82 -0
- data/lib/pubid/iala/builder.rb +2 -0
- data/lib/pubid/iala/identifier.rb +10 -1
- data/lib/pubid/iana/CLAUDE.md +7 -0
- data/lib/pubid/iana/builder.rb +2 -0
- data/lib/pubid/iana/identifier.rb +1 -1
- data/lib/pubid/identifier.rb +161 -17
- data/lib/pubid/idf/builder.rb +11 -1
- data/lib/pubid/idf/identifier.rb +5 -0
- data/lib/pubid/idf/identifiers/all_parts.rb +17 -0
- data/lib/pubid/idf/identifiers.rb +1 -0
- data/lib/pubid/iec/CLAUDE.md +31 -0
- data/lib/pubid/iec/builder.rb +7 -1
- data/lib/pubid/iec/components/consolidated_amendment.rb +4 -0
- data/lib/pubid/iec/components/sheet.rb +2 -0
- data/lib/pubid/iec/components/trf_info.rb +2 -0
- data/lib/pubid/iec/components/vap_suffix.rb +2 -0
- data/lib/pubid/iec/identifier.rb +8 -3
- data/lib/pubid/iec/identifiers/all_parts.rb +19 -0
- data/lib/pubid/iec/identifiers.rb +1 -0
- data/lib/pubid/iec/parser.rb +9 -4
- data/lib/pubid/iec/renderer.rb +0 -1
- data/lib/pubid/iec/urn_generator.rb +9 -1
- data/lib/pubid/iec/urn_parser.rb +3 -2
- data/lib/pubid/ieee/CLAUDE.md +97 -0
- data/lib/pubid/ieee/builder.rb +194 -27
- data/lib/pubid/ieee/components/code.rb +2 -0
- data/lib/pubid/ieee/components/draft.rb +35 -2
- data/lib/pubid/ieee/components/typed_stage.rb +2 -0
- data/lib/pubid/ieee/identifiers/base.rb +21 -1
- data/lib/pubid/ieee/identifiers/iec_ieee_copublished.rb +9 -0
- data/lib/pubid/ieee/identifiers/joint_development.rb +66 -23
- data/lib/pubid/ieee/identifiers/project_draft_identifier.rb +8 -1
- data/lib/pubid/ieee/parser.rb +156 -29
- data/lib/pubid/ieee/renderer.rb +42 -7
- data/lib/pubid/ieee/urn_generator.rb +31 -0
- data/lib/pubid/ietf/CLAUDE.md +7 -0
- data/lib/pubid/ietf/builder.rb +2 -0
- data/lib/pubid/ietf/identifiers/base.rb +1 -1
- data/lib/pubid/iho/builder.rb +2 -0
- data/lib/pubid/isbn/builder.rb +2 -0
- data/lib/pubid/isbn/identifier.rb +1 -1
- data/lib/pubid/iso/CLAUDE.md +47 -0
- data/lib/pubid/iso/builder.rb +19 -5
- data/lib/pubid/iso/components/publisher.rb +2 -0
- data/lib/pubid/iso/identifier.rb +10 -15
- data/lib/pubid/iso/identifiers/all_parts.rb +19 -0
- data/lib/pubid/iso/identifiers/directives_supplement.rb +4 -2
- data/lib/pubid/iso/identifiers.rb +1 -0
- data/lib/pubid/iso/normalizer.rb +4 -1
- data/lib/pubid/iso/rendering_style.rb +0 -1
- data/lib/pubid/itu/CLAUDE.md +115 -0
- data/lib/pubid/itu/builder.rb +26 -4
- data/lib/pubid/itu/components/code.rb +2 -0
- data/lib/pubid/itu/components/designation.rb +2 -0
- data/lib/pubid/itu/components/sector.rb +2 -0
- data/lib/pubid/itu/components/series.rb +2 -0
- data/lib/pubid/itu/identifiers/base.rb +11 -18
- data/lib/pubid/itu/identifiers/radio_regulations.rb +27 -0
- data/lib/pubid/itu/identifiers/special_publication.rb +48 -14
- data/lib/pubid/itu/identifiers/standard_serialization.rb +2 -0
- data/lib/pubid/itu/identifiers/supplement.rb +15 -0
- data/lib/pubid/itu/identifiers.rb +1 -0
- data/lib/pubid/itu/parser.rb +108 -22
- data/lib/pubid/itu/urn_generator.rb +9 -2
- data/lib/pubid/jcgm/CLAUDE.md +7 -0
- data/lib/pubid/jcgm/builder.rb +2 -0
- data/lib/pubid/jcgm/components/publisher.rb +2 -0
- data/lib/pubid/jcgm.rb +1 -1
- data/lib/pubid/jis/builder.rb +5 -1
- data/lib/pubid/jis/identifier.rb +6 -18
- data/lib/pubid/jis/identifiers/all_parts.rb +19 -0
- data/lib/pubid/jis/identifiers.rb +1 -0
- data/lib/pubid/jis/renderer.rb +0 -2
- data/lib/pubid/jis/urn_generator.rb +0 -1
- data/lib/pubid/nist/CLAUDE.md +56 -0
- data/lib/pubid/nist/builder.rb +3 -0
- data/lib/pubid/nist/components/edition.rb +2 -0
- data/lib/pubid/nist/components/issue_number.rb +2 -0
- data/lib/pubid/nist/components/part.rb +2 -0
- data/lib/pubid/nist/components/stage.rb +2 -0
- data/lib/pubid/nist/components/supplement.rb +2 -0
- data/lib/pubid/nist/components/translation.rb +2 -0
- data/lib/pubid/nist/components/update.rb +2 -0
- data/lib/pubid/nist/components/version.rb +2 -0
- data/lib/pubid/nist/components/volume.rb +2 -0
- data/lib/pubid/nist/identifiers/base.rb +39 -7
- data/lib/pubid/nist/parser.rb +24 -2
- data/lib/pubid/nist/preprocessor.rb +53 -2
- data/lib/pubid/nist/urn_parser.rb +10 -1
- data/lib/pubid/oasis/CLAUDE.md +19 -0
- data/lib/pubid/oasis/builder.rb +2 -0
- data/lib/pubid/oasis/identifier.rb +20 -1
- data/lib/pubid/ogc/CLAUDE.md +34 -0
- data/lib/pubid/ogc/builder.rb +2 -0
- data/lib/pubid/ogc/identifier.rb +12 -1
- data/lib/pubid/oiml/CLAUDE.md +189 -0
- data/lib/pubid/oiml/builder.rb +20 -0
- data/lib/pubid/oiml/components/code.rb +6 -0
- data/lib/pubid/oiml/identifier.rb +13 -0
- data/lib/pubid/oiml/identifiers/annex.rb +4 -0
- data/lib/pubid/oiml/identifiers/certification_system.rb +34 -0
- data/lib/pubid/oiml/identifiers/code_number.rb +8 -0
- data/lib/pubid/oiml/identifiers/dual_published.rb +174 -0
- data/lib/pubid/oiml/identifiers.rb +2 -0
- data/lib/pubid/oiml/parser.rb +35 -4
- data/lib/pubid/oiml/renderer.rb +23 -1
- data/lib/pubid/oiml/single_identifier.rb +4 -0
- data/lib/pubid/oiml/supplement_identifier.rb +7 -0
- data/lib/pubid/oiml/urn_generator.rb +28 -0
- data/lib/pubid/oiml.rb +6 -1
- data/lib/pubid/omg/CLAUDE.md +15 -0
- data/lib/pubid/omg/builder.rb +2 -0
- data/lib/pubid/omg/identifier.rb +1 -1
- data/lib/pubid/parg/artifact.rb +46 -0
- data/lib/pubid/parg/backend.rb +92 -0
- data/lib/pubid/parg.rb +8 -0
- data/lib/pubid/parser/grammar.rb +23 -0
- data/lib/pubid/pg.rb +8 -0
- data/lib/pubid/plateau/builder.rb +2 -0
- data/lib/pubid/plateau/identifiers/base.rb +4 -0
- data/lib/pubid/plateau/supplement_identifier.rb +14 -2
- data/lib/pubid/plateau/urn_generator.rb +7 -1
- data/lib/pubid/plateau.rb +1 -2
- data/lib/pubid/renderers/human_readable.rb +0 -1
- data/lib/pubid/sae/builder.rb +2 -0
- data/lib/pubid/sae/components/date.rb +2 -0
- data/lib/pubid/sae/components/type.rb +2 -0
- data/lib/pubid/sae/identifiers/base.rb +1 -1
- data/lib/pubid/subset_match.rb +197 -0
- data/lib/pubid/tgpp/CLAUDE.md +43 -0
- data/lib/pubid/tgpp/builder.rb +2 -0
- data/lib/pubid/tgpp/identifier.rb +15 -1
- data/lib/pubid/type_resolver.rb +14 -2
- data/lib/pubid/un/builder.rb +2 -0
- data/lib/pubid/un/identifier.rb +1 -1
- data/lib/pubid/version.rb +1 -1
- data/lib/pubid/w3c/CLAUDE.md +7 -0
- data/lib/pubid/w3c/builder.rb +2 -0
- data/lib/pubid/w3c/identifier.rb +1 -1
- data/lib/pubid/xsf/CLAUDE.md +11 -0
- data/lib/pubid/xsf/builder.rb +2 -0
- data/lib/pubid/xsf/identifier.rb +1 -1
- data/lib/pubid.rb +17 -3
- metadata +78 -2
data/lib/pubid/ieee/renderer.rb
CHANGED
|
@@ -128,12 +128,20 @@ module Pubid
|
|
|
128
128
|
parts << type_str unless type_str.strip.empty?
|
|
129
129
|
end
|
|
130
130
|
|
|
131
|
-
# Code - with P prefix for projects (concatenated, not separated)
|
|
131
|
+
# Code - with P prefix for projects (concatenated, not separated).
|
|
132
|
+
# IEEE semantics (normative): P = project = the document is a
|
|
133
|
+
# draft; no P = it has become a standard. The P-state is
|
|
134
|
+
# identity-bearing, so the renderer never adds or drops it —
|
|
135
|
+
# whatever the source spelling carries round-trips.
|
|
132
136
|
if id.code_obj
|
|
133
137
|
result = id.code_obj.to_s
|
|
134
138
|
|
|
135
|
-
# Prepend P if this is a project AND code doesn't already have P
|
|
136
|
-
|
|
139
|
+
# Prepend P if this is a project AND code doesn't already have P.
|
|
140
|
+
# A recorded project marker (the source spelled the P on a
|
|
141
|
+
# non-IEEE-led publisher) prints regardless of the publisher.
|
|
142
|
+
if (id.typed_stage&.project_status && should_render_type ||
|
|
143
|
+
id.is_a?(Identifiers::ProjectDraftIdentifier) &&
|
|
144
|
+
id.project_marker) && !result.start_with?("P")
|
|
137
145
|
result = "P#{result}"
|
|
138
146
|
end
|
|
139
147
|
|
|
@@ -142,6 +150,7 @@ module Pubid
|
|
|
142
150
|
result += mark(id.code_obj.number, id.code_obj.prefix,
|
|
143
151
|
publishers: publishers_of(id))
|
|
144
152
|
|
|
153
|
+
|
|
145
154
|
# Only attach year to code if there's no edition, no month, and no draft
|
|
146
155
|
result += "-#{id.year}" if id.year && !id.draft_obj && !id.edition && !id.month
|
|
147
156
|
|
|
@@ -150,9 +159,31 @@ module Pubid
|
|
|
150
159
|
# ("P802.16Rev2/D3"). Keeps a revision distinct from its base standard.
|
|
151
160
|
result += "Rev#{id.revision}" if id.revision
|
|
152
161
|
|
|
153
|
-
# Append draft to code - with or without space based on original
|
|
162
|
+
# Append draft to code - with or without space based on original
|
|
163
|
+
# format. An exactly-IEC/IEEE co-published reference prints only the
|
|
164
|
+
# draft DESIGNATOR - a comma-separated project date ("IEC/IEEE
|
|
165
|
+
# 61886-1/D2, May 2020" prints as "…/D2") is catalogue metadata, not
|
|
166
|
+
# identity. Wider sets ("IEC/ISO/IEEE P82079-1/D1, June 2016") print
|
|
167
|
+
# the date; a space-separated date always stays ("…/DCDV July 2015").
|
|
168
|
+
# IEEE-published drafts normalize the comma date to the long form
|
|
169
|
+
# ("D08, September, 2018").
|
|
154
170
|
if id.draft_obj
|
|
155
|
-
|
|
171
|
+
printed = id.draft_obj.to_s
|
|
172
|
+
# The draft designator carries its date on every lead
|
|
173
|
+
# (docs/IEEE-DRAFT-STAGES.md §3: "the draft designator with its
|
|
174
|
+
# date" — the IEC/IEEE-led date-drop contradicted the DCD
|
|
175
|
+
# convention and the raw records).
|
|
176
|
+
if id.publisher == "IEEE" && id.draft_status.to_s.empty?
|
|
177
|
+
# The long comma form is for dated project drafts; an
|
|
178
|
+
# unapproved-draft render keeps its pinned single-comma form
|
|
179
|
+
# (pubid#318 idempotence).
|
|
180
|
+
printed = printed.sub(/, ([A-Z][a-z]+) (\d{4})\z/, ', \1, \2')
|
|
181
|
+
end
|
|
182
|
+
# A space-separated draft year with no print date trailing is
|
|
183
|
+
# IEEE's dash form ("…/D1-2006"); with a trailing print date the
|
|
184
|
+
# space stays ("…/D1 2004, Nov 2004").
|
|
185
|
+
printed = printed.sub(/\A(D\d+) (\d{4})\z/, '\1-\2')
|
|
186
|
+
result += id.space_before_draft ? " #{printed}" : printed
|
|
156
187
|
end
|
|
157
188
|
|
|
158
189
|
# Append interpretation notation (/INT)
|
|
@@ -190,8 +221,12 @@ module Pubid
|
|
|
190
221
|
# Build the main identifier (without month yet)
|
|
191
222
|
result = parts.join(" ")
|
|
192
223
|
|
|
193
|
-
|
|
194
|
-
|
|
224
|
+
|
|
225
|
+
# Month/Day - append directly to avoid extra space before comma.
|
|
226
|
+
# An exactly-IEC/IEEE co-published reference drops the comma date
|
|
227
|
+
# entirely (same catalogue-metadata rule as the draft date above):
|
|
228
|
+
# "IEC/IEEE P60076-16, May 2016" prints as "IEC/IEEE 60076-16".
|
|
229
|
+
if id.month && !(id.publisher == "IEC" && id.copublisher == ["IEEE"])
|
|
195
230
|
result += ", #{id.month}"
|
|
196
231
|
result += " #{id.day}" if id.day
|
|
197
232
|
if id.year && !id.edition
|
|
@@ -127,6 +127,10 @@ module Pubid
|
|
|
127
127
|
end
|
|
128
128
|
|
|
129
129
|
def publisher_component
|
|
130
|
+
# The IEC/IEEE co-published pair is definitional to the class and
|
|
131
|
+
# was previously lost ("urn:ieee:ieee" — no number, no draft).
|
|
132
|
+
return "iec-ieee" if identifier.is_a?(Identifiers::IecIeeeCopublished)
|
|
133
|
+
|
|
130
134
|
pub = normalized_publisher
|
|
131
135
|
|
|
132
136
|
if identifier.copublisher&.any?
|
|
@@ -150,6 +154,14 @@ module Pubid
|
|
|
150
154
|
end
|
|
151
155
|
|
|
152
156
|
def type_component
|
|
157
|
+
# A joint stage draft carries the stage in its draft clause and the
|
|
158
|
+
# project P inside the code — a separate type segment would
|
|
159
|
+
# duplicate both.
|
|
160
|
+
if identifier.is_a?(Identifiers::JointDevelopment) &&
|
|
161
|
+
identifier.ieee_draft.to_s.start_with?("D=")
|
|
162
|
+
return nil
|
|
163
|
+
end
|
|
164
|
+
|
|
153
165
|
return nil unless identifier.type
|
|
154
166
|
|
|
155
167
|
type = identifier.type
|
|
@@ -159,6 +171,16 @@ module Pubid
|
|
|
159
171
|
end
|
|
160
172
|
|
|
161
173
|
def code_component
|
|
174
|
+
# The co-published code is number+parts+separators (no year — its
|
|
175
|
+
# own segment), e.g. "61886-1".
|
|
176
|
+
if identifier.is_a?(Identifiers::IecIeeeCopublished)
|
|
177
|
+
return nil if identifier.number.to_s.empty?
|
|
178
|
+
|
|
179
|
+
# Year-less: the publication year is its own URN segment (the
|
|
180
|
+
# rebuilt copublished_number would double-emit it).
|
|
181
|
+
return "#{identifier.number}#{identifier.parts_suffix}"
|
|
182
|
+
end
|
|
183
|
+
|
|
162
184
|
return nil unless identifier.code_obj
|
|
163
185
|
|
|
164
186
|
identifier.code_obj.to_s
|
|
@@ -182,6 +204,15 @@ module Pubid
|
|
|
182
204
|
end
|
|
183
205
|
|
|
184
206
|
def draft_component
|
|
207
|
+
if identifier.is_a?(Identifiers::JointDevelopment) &&
|
|
208
|
+
identifier.ieee_draft.to_s.start_with?("D=")
|
|
209
|
+
return "draft.#{identifier.ieee_draft}"
|
|
210
|
+
end
|
|
211
|
+
|
|
212
|
+
if identifier.is_a?(Identifiers::IecIeeeCopublished)
|
|
213
|
+
return identifier.draft_info ? "draft.#{identifier.draft_info}" : nil
|
|
214
|
+
end
|
|
215
|
+
|
|
185
216
|
return nil unless identifier.draft_obj
|
|
186
217
|
|
|
187
218
|
"draft.#{identifier.draft_obj}"
|
|
@@ -0,0 +1,7 @@
|
|
|
1
|
+
# IETF flavor notes
|
|
2
|
+
|
|
3
|
+
IETF readiness for the unified relaton index.
|
|
4
|
+
|
|
5
|
+
These notes were part of the root `CLAUDE.md`. Read them before you change `lib/pubid/ietf/` or `spec/pubid/ietf/`. The root file keeps the cross-flavor contract that every flavor obeys.
|
|
6
|
+
|
|
7
|
+
- **IETF readiness for the unified relaton index (`relaton-data-ietf`)**: relaton is collapsing RFCs, the RFC sub-series (BCP/STD/FYI) and Internet-Drafts into **one** index keyed on the structured pubid hash (`pubid_class: ::Pubid::Ietf::Identifier`). That index is **all-or-nothing** — `Relaton::Index::FileIO#deserialize_id` raises on the *first* unparseable row and `#load_index` then rejects the whole index — so one bad id out of 176,862 means no consumer resolves *any* IETF reference. Four changes, all verified against the published corpora (`relaton-data-{rfcs,rfcsubseries,ids}`, branch `v2`). **(1) `draft_rest` accepts `.` and `A-Z`** (`match['a-zA-Z0-9+_.\-']`): exactly 50 of the 166,740 published draft ids failed without them (21 dotted `draft-ietf-pilc-2.5g3g-12`, 29 uppercase `draft-chapin-clnp-ISO8473-00`, 0 both), and a character census confirms the class now covers the corpus **exactly** — the only chars outside `[a-z0-9-]` anywhere in it are `. _ +` and `A-Z`. Widening is safe because the top-level alternation is keyed on the leading token (`RFC`/`BCP|STD|FYI`/`draft-`), so a wider draft tail can never make another branch match; `Builder#split_draft_version` needed no change (it inspects only the last three characters). The set stays **closed** — a space, `/` or `&` in a slug is crawler junk and still fails (pinned by negative examples). **(2) Zero-padded ids parse, grammar-level**: `rfc-index.xml` writes sub-series membership as `<is-also><doc-id>STD0066</doc-id></is-also>` (369 padded doc-ids) and RFCs as `RFC0001`, and relaton#109 forbids relaton normalizing them itself. The space is `.maybe` on both the `rfc` and `subseries` rules and `Builder#unpad` strips the pad (`/\A0+(?=\d)/`, a look-ahead so an all-zero number keeps its last digit), so `STD0066`/`STD 0066`/`STD66` all render the canonical `STD 66`. `unpad` is applied unconditionally, so the **spaced** padded form is normalized too — `RFC 0001` previously parsed and rendered back verbatim, and now renders `RFC 1` with a correspondingly changed hash and URN. That is deliberate (one canonical spelling per document) and has **zero** effect on the published corpora, which contain no padded row. **This is a normalizing parse, so the padded spellings must NEVER go into the byte-exact `spec/fixtures/ietf/identifiers/pass/`** — they live in `spec/pubid/ietf/zero_padded_spec.rb`, the BIPM `CIPM/2005-06(REV)` precedent. **(3) The Internet-Draft slug moved from `name` into `number`, and `series` became derived**: relaton bsearches on `id.root.number.to_s`, so with the slug in `name` all 166,740 drafts keyed to `""` and the narrowing bought nothing (with the slug: 43,564 buckets, median 2, worst 101). `name` is **dropped entirely** (no alias); the slug — leading `draft-` included — is the key, deliberately shared by a draft's unversioned aggregator row and every one of its versions so a document and its variants cluster. `series` is likewise **not stored**: `_type: pubid:ietf:std` already encodes it, so `Bcp`/`Std`/`Fyi` expose a `SERIES` constant + a plain `#series` reader (safe — `::Pubid::Identifier` declares no `series` attribute and nothing generic reads `.series`) and `build_subseries` stops passing it. The whole serialized vocabulary is now **three keys**: `_type` + `number` always, `version` only on a versioned draft — a sub-series row (`{_type, number}`) has the same shape as an RFC's. **The attributes live on the five LEAVES, not the shared `Pubid::Ietf::Identifier`**, via the `Identifiers::Serialization` mixin (`number` + the `key_value` block; `InternetDraft` merges its own `version` mapping on top — lutaml combines a mixin-installed block with a second in-class one, the ITU precedent). This is the IEEE `number` determinism landmine: a `:string` redefinition of the parent's `Components::Code number` on a class the leaves *inherit* from resolves nondeterministically under multi-flavor load. It was invisible while drafts left `number` nil; it is load-bearing now. The base deliberately carries **no** `key_value` block — one there would be inherited-and-merged by every leaf. `spec/pubid/ietf/root_number_spec.rb` is the tripwire (asserts `number` is a `String`, the exact keys, the exact `to_hash` key sets) and **must run under the full `rake` suite** to mean anything. **(4) Round-trip is validated here because relaton structurally cannot**: `FileIO#id_supported?` **skips** its `to_hash`/`from_hash` check whenever the object is a concrete subclass (`return true unless obj.instance_of?(@pubid_class)`), and *every* IETF id is a subclass — so a lossy round-trip would pass validation silently and produce wrong lookups. Three layers: `fixtures_spec.rb` is now **zero-failure** (was a 90% threshold — wrong for an all-or-nothing index) and checks `to_s` byte-exactness, `from_hash(to_hash)` fidelity in both `to_s` and `to_hash`, **and** a non-empty `root.number` for every fixture line; the fixtures grew to ~2,500 lines (the complete 365-row sub-series corpus, a 139-id structural edge set, a deterministic 1-in-84 draft sample — each file's header records the regeneration command); and `spec/pubid/ietf/corpus_round_trip_spec.rb` runs the same four checks over all 176,862 published ids, **opt-in** via `PUBID_IETF_CORPUS=<dir with the three relaton-data checkouts> bundle exec rake test:corpus_ietf` (self-skips otherwise; ~100s; verified 0 failures). **relaton note — the indexes must be regenerated, there is no alias**: `name` and `series` are gone as serialized keys, so a pre-`number` row (`{_type, name, version}`) deserializes with **`number` nil** and no error at all — lutaml ignores unknown keys, and relaton's subclass skip means `id_supported?` won't catch it either, leaving every draft keyed under `""`. Rendering is therefore the loud failure: `Renderer#require_number!` raises `ArgumentError` on a nil/empty `number` rather than returning `nil` or `"-02"`, and the renderer's sub-series arm is an explicit `when Bcp, Std, Fyi` with an `else raise` instead of a catch-all (a bare `Pubid::Ietf::Identifier` has no `series` reader). Pinned by the "legacy `name`-shaped hash" block in `root_number_spec.rb`. **Pre-existing bug fixed in passing**: `UrnParser`'s bare `Errors::ParseError` does not resolve — `Errors` lives under `Pubid::UrnParser` (the module), which is neither in the lexical scope of `Pubid::Ietf::UrnParser` nor an ancestor of `Base` — so an invalid URN raised `NameError` instead, and the spec's `raise_error(StandardError)` accepted it. Now fully qualified as `Pubid::UrnParser::Errors::ParseError` (the adobe/easc/gost/iala convention) with the spec asserting that exact class. **`Pubid::Iec::UrnParser` has the identical bug and is untouched here.** Also: `Identifier.parse` gained the inline `MAX_INPUT_LENGTH` guard — relaton reaches that class-level funnel directly through `pubid_class:`, bypassing `Pubid::Ietf.parse` — and IETF was added to the three cross-flavor tables it was missing from (`identifier_roundtrip_spec`, `uniform_identifier_handle_spec`, `redos_guard_spec`), which is what exercises the leaf `number` under real multi-flavor load. (hand-off: ietf-index-readiness.)
|
data/lib/pubid/ietf/builder.rb
CHANGED
data/lib/pubid/iho/builder.rb
CHANGED
data/lib/pubid/isbn/builder.rb
CHANGED
|
@@ -64,7 +64,7 @@ module Pubid
|
|
|
64
64
|
|
|
65
65
|
# @raise [Pubid::Errors::ParseError] if the string is not a valid ISBN
|
|
66
66
|
def self.build_identifier(identifier)
|
|
67
|
-
parsed =
|
|
67
|
+
parsed = Pubid::Parg::Backend.parse(:isbn, identifier)
|
|
68
68
|
Builder.build(parsed)
|
|
69
69
|
rescue ArgumentError => e
|
|
70
70
|
# The Builder validates length and check digit. Surface that as a parse
|
|
@@ -0,0 +1,47 @@
|
|
|
1
|
+
# ISO flavor notes
|
|
2
|
+
|
|
3
|
+
- **No identifier attribute holds a `Components::Code` any more**: `number`,
|
|
4
|
+
`part` and `subpart` on `Pubid::Iso::Identifier`, the committee structure on
|
|
5
|
+
`Identifiers::TcDocument` (`tc_type`/`tc_number`/`sc_type`/`sc_number`/
|
|
6
|
+
`wg_type`/`wg_number`), `Identifiers::Directives#subgroup` and the
|
|
7
|
+
`edition.number` the builder writes are all plain strings, and
|
|
8
|
+
`Pubid::Iso::Components::Code` is deleted.
|
|
9
|
+
**The component was degenerate everywhere it was used.** Measured over the
|
|
10
|
+
whole 7,613-id pass corpus: no ISO Code carries `prefix`, `part`, `subpart`
|
|
11
|
+
or `parts`, `parts` is never written anywhere in `lib/pubid/iso`, and no Code
|
|
12
|
+
renders differently from its own `value`. ISO's part and subpart are sibling
|
|
13
|
+
attributes of the identifier, not fields of the number — the `-` join in
|
|
14
|
+
`Code#to_s` was unreachable.
|
|
15
|
+
**`edition.number` was worse than degenerate: it leaked a live object.**
|
|
16
|
+
`Components::Edition#number` is typed `Lutaml::Model::Type::Value`, so lutaml
|
|
17
|
+
never serialized it — `to_hash` returned the `Components::Code` instance
|
|
18
|
+
itself, and `to_yaml` emitted `!ruby/object:Pubid::Components::Code` with its
|
|
19
|
+
internal ivars, which `YAML.safe_load` refuses. An edition now serializes
|
|
20
|
+
`{"number" => "13", "original_text" => "Ed 13"}`.
|
|
21
|
+
**The TC converters are gone with it**: twelve `*_to_kv`/`*_from_kv` helpers
|
|
22
|
+
were replaced by plain `map "tc_type", to: :tc_type` declarations (the
|
|
23
|
+
ETSI/OIML shape), because a String needs no converter to flatten.
|
|
24
|
+
**Two shared surfaces needed a change, and both are the documented shims**:
|
|
25
|
+
`Renderers::DirectivesRenderer` called `subgroup.render(context:)` directly
|
|
26
|
+
and now reads through `render_component`; `Pubid::Iso.build_code` returned a
|
|
27
|
+
Code and now returns the string (ASTM's `to_iso_identifier` goes through it,
|
|
28
|
+
which is how one ASTM example caught the `NameError` after the class was
|
|
29
|
+
deleted).
|
|
30
|
+
**Verification — replay a baseline, do not trust the suite alone.** Against
|
|
31
|
+
all 7,621 pass-fixture ids captured on the parent commit, `to_s`, `to_urn`,
|
|
32
|
+
`to_mr_string`, `root.number` and the identifier class are **byte-identical**,
|
|
33
|
+
and `to_hash` differs for exactly **12** ids: the 9 `subgroup` rows and the 3
|
|
34
|
+
editions below. The first replay also caught a real regression the specs did
|
|
35
|
+
not — `directives.rb` still called `part.value.downcase`, so 4 Directives
|
|
36
|
+
URNs raised.
|
|
37
|
+
**relaton note — `relaton-data-iso` needs a re-crawl, and only for one key.**
|
|
38
|
+
`subgroup` flattens from `{"value" => "JTC 1"}` to `"JTC 1"`, and the
|
|
39
|
+
published `index-v2.yaml` (branch `v2`) carries **5 such rows**, all ISO/IEC
|
|
40
|
+
JTC 1 directives. **No compatibility shim is kept** — a stored nested
|
|
41
|
+
`subgroup` will not deserialize. The edition change reaches nothing stored:
|
|
42
|
+
`grep -c edition` over that index returns **0**.
|
|
43
|
+
**Spec fallout is the shape to expect for the remaining flavors**: 592
|
|
44
|
+
assertions across 25 files read `.number.value` / `.part.value`. The rewrite
|
|
45
|
+
needs a lookbehind, because `tc_number.value` and `edition.number.value` are
|
|
46
|
+
different attributes — the first pass stripped `edition.number.value` too,
|
|
47
|
+
and two corrigendum examples caught it.
|
data/lib/pubid/iso/builder.rb
CHANGED
|
@@ -58,6 +58,12 @@ module Pubid
|
|
|
58
58
|
end
|
|
59
59
|
|
|
60
60
|
def build(parsed_hash)
|
|
61
|
+
# "(all parts)" and the ":ser" URN name every part of the document,
|
|
62
|
+
# so they build an AllPartsIdentifier around the document (the URN
|
|
63
|
+
# parser puts the key on the outermost identifier for the same
|
|
64
|
+
# reason). The document itself holds no all-parts mark.
|
|
65
|
+
all_parts = parsed_hash.delete(:all_parts)
|
|
66
|
+
|
|
61
67
|
# For ISO/R legacy format, split into publisher and type
|
|
62
68
|
if parsed_hash[:iso_r_prefix]
|
|
63
69
|
parsed_hash[:publisher] = "ISO"
|
|
@@ -77,14 +83,20 @@ module Pubid
|
|
|
77
83
|
end
|
|
78
84
|
end
|
|
79
85
|
|
|
80
|
-
#
|
|
81
|
-
|
|
82
|
-
|
|
83
|
-
#
|
|
86
|
+
# For French GUIDE entries: "Guide ISO/CEI 37:1995". The rename
|
|
87
|
+
# must happen BEFORE class selection — locate_identifier_klass
|
|
88
|
+
# reads :type_with_stage, and the tree of a guide-first spelling
|
|
89
|
+
# carries an empty :type_with_stage plus the type under
|
|
90
|
+
# :type_with_stage_fr, so deferring the rename selected the
|
|
91
|
+
# default International Standard class for every guide-first
|
|
92
|
+
# (and Cyrillic "Руководства ИСО …") reference.
|
|
84
93
|
if type_with_stage_fr = parsed_hash.delete(:type_with_stage_fr)
|
|
85
94
|
parsed_hash[:type_with_stage] = type_with_stage_fr
|
|
86
95
|
end
|
|
87
96
|
|
|
97
|
+
# Instantiate the identifier based on the typed stage
|
|
98
|
+
identifier = locate_identifier_klass(parsed_hash).new
|
|
99
|
+
|
|
88
100
|
# For DirectivesSupplement, rename :publisher to :supplement_publisher
|
|
89
101
|
if identifier.is_a?(Identifiers::DirectivesSupplement) && parsed_hash[:publisher]
|
|
90
102
|
parsed_hash[:supplement_publisher] = parsed_hash.delete(:publisher)
|
|
@@ -112,7 +124,7 @@ module Pubid
|
|
|
112
124
|
identifier.type = default_typed_stage.to_type
|
|
113
125
|
end
|
|
114
126
|
|
|
115
|
-
identifier
|
|
127
|
+
all_parts ? identifier.to_all_parts : identifier
|
|
116
128
|
end
|
|
117
129
|
|
|
118
130
|
def handle_key(identifier, key, value)
|
|
@@ -313,3 +325,5 @@ module Pubid
|
|
|
313
325
|
end
|
|
314
326
|
end
|
|
315
327
|
end
|
|
328
|
+
|
|
329
|
+
Pubid::Iso::Builder.prepend(Pubid::Builder::AllPartsWrap)
|
|
@@ -8,6 +8,8 @@ module Pubid
|
|
|
8
8
|
# ISO Publisher with copublisher support
|
|
9
9
|
# Examples: ISO, ISO/IEC, ISO/IEC/IEEE
|
|
10
10
|
class Publisher < Lutaml::Model::Serializable
|
|
11
|
+
include ::Pubid::SubsetMatch
|
|
12
|
+
|
|
11
13
|
attribute :publisher, :string, default: -> { "ISO" }
|
|
12
14
|
attribute :copublisher, :string, collection: true
|
|
13
15
|
|
data/lib/pubid/iso/identifier.rb
CHANGED
|
@@ -11,6 +11,12 @@ module Pubid
|
|
|
11
11
|
attribute :copublishers, ::Pubid::Iso::Components::Publisher,
|
|
12
12
|
collection: true
|
|
13
13
|
|
|
14
|
+
# ISO prints "(all parts)" and has a series URN, so it has its own
|
|
15
|
+
# all-parts class.
|
|
16
|
+
def self.all_parts_class
|
|
17
|
+
Identifiers::AllParts
|
|
18
|
+
end
|
|
19
|
+
|
|
14
20
|
# The publisher implied when none is serialized. ISO for most types;
|
|
15
21
|
# publisher-less types (IWA) override this to nil.
|
|
16
22
|
def self.default_publisher
|
|
@@ -137,20 +143,6 @@ module Pubid
|
|
|
137
143
|
# unique typed-stage `code` under "stage" and recompute the rest on
|
|
138
144
|
# load. _type already pins the document type.
|
|
139
145
|
map "stage", with: { to: :stage_to_kv, from: :stage_from_kv }
|
|
140
|
-
# Omit the `false` default; only the meaningful `true` is serialized.
|
|
141
|
-
map "all_parts", with: { to: :all_parts_to_kv, from: :all_parts_from_kv }
|
|
142
|
-
end
|
|
143
|
-
|
|
144
|
-
def all_parts_to_kv(model, doc)
|
|
145
|
-
return unless model.all_parts
|
|
146
|
-
|
|
147
|
-
doc.add_child(
|
|
148
|
-
Lutaml::KeyValue::DataModel::Element.new("all_parts", true),
|
|
149
|
-
)
|
|
150
|
-
end
|
|
151
|
-
|
|
152
|
-
def all_parts_from_kv(model, value)
|
|
153
|
-
model.all_parts = value
|
|
154
146
|
end
|
|
155
147
|
|
|
156
148
|
# Serialize typed_stage as just its unique code (e.g. "is", "dis",
|
|
@@ -312,7 +304,10 @@ module Pubid
|
|
|
312
304
|
when :mr_string
|
|
313
305
|
Pubid::Parsers::MrString.parse(string)
|
|
314
306
|
else
|
|
315
|
-
|
|
307
|
+
# R1 parser swap: the baked PG artifact is the identifier
|
|
308
|
+
# parser of record; the Builder consumes the same attribute
|
|
309
|
+
# hash it always has.
|
|
310
|
+
parsed = Pubid::Parg::Backend.parse(:iso, string)
|
|
316
311
|
Pubid::Iso::Builder.new.build(parsed)
|
|
317
312
|
end
|
|
318
313
|
end
|
|
@@ -0,0 +1,19 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
|
|
3
|
+
module Pubid
|
|
4
|
+
module Iso
|
|
5
|
+
module Identifiers
|
|
6
|
+
# Every part of one ISO document: "ISO 9000 (all parts)".
|
|
7
|
+
#
|
|
8
|
+
# The URN is the series URN: the document URN without its stage, plus
|
|
9
|
+
# the "ser" slot.
|
|
10
|
+
class AllParts < ::Pubid::Iso::Identifier
|
|
11
|
+
include ::Pubid::AllParts
|
|
12
|
+
|
|
13
|
+
def to_urn
|
|
14
|
+
"#{identity.exclude(:stage, :typed_stage).to_urn}:ser"
|
|
15
|
+
end
|
|
16
|
+
end
|
|
17
|
+
end
|
|
18
|
+
end
|
|
19
|
+
end
|
|
@@ -89,8 +89,10 @@ format: nil, stage_format_long: nil, with_date: nil, **opts)
|
|
|
89
89
|
def to_supplement_s(lang: :en, lang_single: false, with_edition: false,
|
|
90
90
|
format: nil, stage_format_long: nil, with_date: nil, **_opts)
|
|
91
91
|
date_str = if date
|
|
92
|
-
|
|
93
|
-
|
|
92
|
+
# Components::Date#render already carries the month
|
|
93
|
+
# (and day) when present — appending a month_part
|
|
94
|
+
# here doubled it (":2016-05-05-05").
|
|
95
|
+
":#{date.render}"
|
|
94
96
|
else
|
|
95
97
|
""
|
|
96
98
|
end
|
|
@@ -4,6 +4,7 @@ module Pubid
|
|
|
4
4
|
module Iso
|
|
5
5
|
module Identifiers
|
|
6
6
|
autoload :Addendum, "#{__dir__}/identifiers/addendum"
|
|
7
|
+
autoload :AllParts, "#{__dir__}/identifiers/all_parts"
|
|
7
8
|
autoload :Amendment, "#{__dir__}/identifiers/amendment"
|
|
8
9
|
autoload :Corrigendum, "#{__dir__}/identifiers/corrigendum"
|
|
9
10
|
autoload :Data, "#{__dir__}/identifiers/data"
|
data/lib/pubid/iso/normalizer.rb
CHANGED
|
@@ -27,7 +27,10 @@ module Pubid
|
|
|
27
27
|
private
|
|
28
28
|
|
|
29
29
|
def parse_with_builder(string)
|
|
30
|
-
|
|
30
|
+
# R2 ingestion hook: normalized strings reach the model through
|
|
31
|
+
# the same parser of record as Identifier.parse (the baked PG
|
|
32
|
+
# artifact), so every Tier-3 normalization feeds one grammar.
|
|
33
|
+
parsed = Pubid::Parg::Backend.parse(:iso, string)
|
|
31
34
|
Pubid::Iso.builder.build(parsed)
|
|
32
35
|
end
|
|
33
36
|
|
|
@@ -0,0 +1,115 @@
|
|
|
1
|
+
# ITU flavor notes
|
|
2
|
+
|
|
3
|
+
ITU grammar, versions, annexes, reports and identity surfaces.
|
|
4
|
+
|
|
5
|
+
These notes were part of the root `CLAUDE.md`. Read them before you change `lib/pubid/itu/` or `spec/pubid/itu/`. The root file keeps the cross-flavor contract that every flavor obeys.
|
|
6
|
+
|
|
7
|
+
- **ITU Questions / Handbooks / N-way combined**: `Pubid::Itu` models study-group **Questions** (`Identifiers::Question`, `_type: pubid:itu:question` — numeric `ITU-R 234-1/7:` and letter-series `ITU-R P.3/BL/7`, `ITU-R S.[4/BL/2]:`; the `/BL` segment, brackets and trailing `:` are carried as `has_bl`/`bracketed`/`has_colon` booleans + a `study_group` string, all reconstructed by `render_base`) and **Handbooks** (`Identifiers::Handbook`, `ITU-R 23.HDB`; the `.HDB` marker is implied by `_type`). Parser disambiguation is by what follows a `/`: **digits** → a Question study-group; a **series** (letters) → a combined designation — so the new rules sit before `with_series`/`without_series`. **Combined (joint) recommendations are N-way**: the primary designation stays on the base `series`/`code` (keeps `root.number` non-empty for relaton-index) and each *additional* designation is a `Components::Designation` in `CombinedIdentifier#combined` (one for a dual `G.780/Y.1351`, two+ for a triple `G.780/Y.1351/Z.1362`); the parser's `combined_designation` rule is `repeat(1)`. **`#number` at the root**: ITU stores the document number on `code`, but `Pubid::Itu::Identifier#number` delegates to `code&.number` so `id.root.number.to_s` is the key relaton-index sorts/bsearches on (non-empty for every type). The delegation is serialization-neutral (the flat block maps `number` via `number_to_kv`/`number_from_kv` off `code`, never the inherited `number` attribute); Supplement/Amendment/Corrigendum/Errata keep their own `:string` `number` ordinal (it overrides the reader), and their base document number is reached via `root.number`.
|
|
8
|
+
|
|
9
|
+
- **ITU `(V##)` versions, labelled annexes, and class-strict supplement `==`**: three ITU-T forms relaton's `relaton-data-itu` crawler needs (~930 data files had no index row because they failed to parse). **(1) `version`** — `ITU-T H.264 (V14) (08/2021)`: a plain **`attribute :version, :string`** on the shared `Pubid::Itu::Identifier` (safe — `::Pubid::Identifier` declares no `version`, so this is a *new* attribute, not the `number`/`stage` retype landmine), fed by the `version_part` rule (`space >> "(V" >> digits >> ")"`) spliced in **immediately before `date_part.maybe`** in all four document rules (`with_series`/`without_series`/`base_with_series`/`base_without_series`) — version always precedes the date in the corpus, and `date_part` requires digits after `(` so the two never compete. Rendered `" (V#{version})"` between code and date by `render_base` **and** by `CombinedIdentifier#to_s` (which overrides `to_s` and does not call `render_base` — miss it and the version is silently dropped on a joint id); mapped in `StandardSerialization`; compared in both `==` overrides. **URN and MR strings stay version-blind by design** (the ITU URN convention defines no version segment), so `(V13)`/`(V14)` share a URN. **(2) Labelled annex** — `ITU-T A.23 Annex A (06/2014)` is its **own** leaf class `Identifiers::AnnexOfRecommendation` (`_type: pubid:itu:annex-of-recommendation`), *not* an extension of `Identifiers::Annex`: that models the structurally different label-less "Annex to ITU OB No. 1000" (prefix rendering + i18n templates) whose `_type` is already persisted in index rows. Shape mirrors `Supplement`: polymorphic `base` + `:string number` (the label — `A`, `F3`, the one-off `C+`) + its own date/language, a minimal `key_value` block (never re-emitting the base's sector/series/code), `render_base` → `"<base> Annex <label>"`, `to_urn` → `"<base urn>:annex:<label>"`, `#root` walking `base` so `root.number` is the annexed document's number. Grammar: `annex_body` wraps the annexed document in **`base.as(:base)`** — load-bearing, because the base may carry its own date (`ITU-T X.692 (2002) Annex E (03/2002)`) and two `date_part`s flattened into one hash would make Parslet silently keep only the last `:year`. `annex_identifier` sits in `rule(:identifier)` **after `supplement_identifier`** (an annex can carry a trailing `Cor.`/`Err.`/`Amd.`, and PEG ordered choice never re-enters once an alternative succeeds) and **before `with_series`** (which would match the `ITU-T A.23` prefix and then fail on the unconsumed ` Annex A`); `supplement_with_base` takes `(annex_body.as(:base) | base.as(:base))` so `ITU-T G.729 Annex B (1996) Cor. 3 (03/2001)` nests annex-inside-supplement. `Builder#build_supplement`'s fallback `Components::Sector.new` is now guarded on `data[:sector]` — an annex base keeps sector/series on *its* base, and the unguarded fallback raised `Invalid sector: `. **(3) `Supplement#==`** now uses **`instance_of?(self.class)`** (so `Suppl. 2` ≠ `Amd 2` ≠ `Cor. 2`, symmetrically — `is_a?` made a Supplement equal an Amendment one-way) and compares `sector`/`series` **only when `base` is nil**. Both halves matter: the series-only form (`ITU-T A Suppl. 2`) holds its whole identity in sector+series, so ignoring them made every series' `Suppl. 2` equal (an index search returned 42 rows, and `eql?`/`hash` collapsed them in a `Set`); but a *based* supplement copies sector/series from its base while the key_value block deliberately does **not** serialize them, so comparing them unconditionally would break `parse(s) == from_hash(parse(s).to_hash)` — the very lookup this fixes. The now-identical `==` overrides on `Amendment`/`Corrigendum`/`Errata` were deleted. `language` is deliberately **not** compared (an index row without a language suffix must still match a reference that has one). Locked by `spec/pubid/itu/identifiers/{version,annex_of_recommendation,supplement}_spec.rb` + rows in `root_number_spec.rb`/`serialization_spec.rb` and the new `spec/fixtures/itu/identifiers/pass/{version,annex_of_recommendation}.txt`. **Rendering/URN/MR consequences of the annex wrapper:** `CombinedIdentifier` had its rendering in `to_s`, which the annex (composing `base.render_base`) bypassed — silently dropping the `/Y.1351` half, so `G.780/Y.1351 Annex A` and `G.780/Z.1362 Annex A` printed identically (and `to_s` is relaton's document number *and* output filename). It now renders in **`render_base`**, with the inherited `to_s` adding the language suffix and common-text twin. The annex's `to_urn` appends its **own date** after the label (`urn:itu:t:A.23:annex:a:06/2014`) — the base's date rides inside `base_urn`, so omitting the annex's would collide editions — and it defines **`mr_supplement_suffix`** (`annex.<label>[.<year>]`, `C+`→`cplus`) so the shared MrString renderer recurses into the annexed document instead of collapsing every annex onto a bare `itu.<date>`. Because the annex keeps sector/series/code on *its* base, `build_supplement` reaches identity through **`base.root`** when the base has none, so a supplement *of* an annex keys like any other supplement. **Version matching semantics:** `version` is in `==`, so a bare `ITU-T H.264` does **not** match `ITU-T H.264 (V14) …` under `ignore: %i[year month]` — version is a *separable trailing component* like ETSI's (`partial_ref_spec` now lists ITU as `omits: %i[date version]`), so callers matching a partial ref must add `:version` to `ignore`; it is deliberately **not** folded into the `:year`→`:date` alias. **Also fixed here:** `spec/pubid/itu/fixtures_spec.rb`'s glob had one `..` too many (repo-root `fixtures/`), so the whole ITU fixture round-trip spec silently iterated an empty file list. **Out of scope at the time, both since fixed:** `ITU-T V.25 ter Annex A (08/1996)` (the space-separated `ter` gap — closed by the code-suffix work below), and `Amendment#to_s`'s missing dot. (hand-off: itu-version-annex-and-supplement-matching.)
|
|
10
|
+
|
|
11
|
+
- **ITU-T unparseable print forms (the `relaton-data-itu` residue)**: 668 of the 755 unindexed `relaton-data-itu` records were ITU-T `rec_name` spellings the grammar had no rule for; all 668 now parse, render and satisfy relaton's gate (`from_hash(to_hash) == to_hash`), with **zero** change to any of the 20,728 already-indexed ids. **The single most important invariant when touching this area: every new rendering flag is a `:boolean` named for the RARE form with `default: -> { false }`**, because the canonical `to_hash` strips default-valued attributes — a flag named for the common case would add a key to every already-published index row. Ten constructs, grouped by where they live:
|
|
12
|
+
**(1) Supplement family.** `supplement_type` gained `Add.`/`Add` → the new leaf `Identifiers::Addendum < Supplement` (`_type: pubid:itu:addendum`), and `Builder#build_supplement`'s `case` gained an **`else raise ArgumentError`** — it previously fell through to a nil class and died with `undefined method 'new' for nil` one frame later. The near-identical `to_s` on `Supplement`/`Amendment`/`Corrigendum`/`Errata` collapsed into one **`Supplement#render_supplement(label)`**; each subclass now supplies only its label. `supplement_number` made the space optional (`ITU-T E Suppl.1`, `ITU-T D.211 Suppl.1` → `number_glued`) and the ordinal dotted (`ITU-T M Suppl. 1.1` → `number` is the whole `"1.1"`). `number_glued` is serialized but deliberately **not** in `==` — `Suppl.1` and `Suppl. 1` are one document. `chained_supplement`'s inner alternation gained `supplement_series_only` so a supplement *of* a series-only supplement parses (`ITU-T G Suppl. 39 (2006) Err. 1 (08/2006)`).
|
|
13
|
+
**(2) `Technical Cor.`** (158 records, the largest bucket — and the one the hand-off mislabelled as a series-code form). A `technical_marker` rule at the **head of `supplement_type`**, not folded into its alternation, so the type token stays the plain `Cor.` the builder's `case` maps; a `technical` boolean on `Corrigendum` renders and compares it. The slash-joined pair `ITU-T X.680 (1994) Amd. 1/Technical Cor. 1 (12/1997)` is a new **`slash_chained_supplement`** rule (a Corrigendum whose `base` is the Amendment) placed **before `supplement_with_base`** in the alternation — that alternative alone matches the `… Amd. 1` prefix and leaves `/Technical Cor. 1 …` unconsumed, which fails the whole parse with no re-entry.
|
|
14
|
+
**(3) Code suffixes — the highest-blast-radius change, since it fires on *every* recommendation parse.** A shared `code_suffixes` rule (`series_suffix_spaced.maybe >> qualifier_spaced.maybe`) spliced after `code` and **before `combined_suffixes`** (the corpus attaches the primary's suffix before the `/`: `ITU-T D.301 R/F.66`) in all four document rules. It cannot live *inside* `code`, or its leading space would also be offered to ` Suppl.`/` Annex` and have to backtrack out of a rule the wrappers depend on. `Components::Code` gained `series_suffix_spaced` (the spaced `ITU-T E.250 bis` vs the spec-locked glued `X.50bis` of issue #231 — the word itself is stored whitespace-free either way, so the serialized value stays comparable), plus `qualifier` + `qualifier_glued` for the trailing letter (`ITU-T D.200 R`, `ITU-T Q.2931 B`, `ITU-T R.38 A`, glued `ITU-T D.502R`, lowercase `ITU-T I.256.2a`). **Both guards on `qualifier_letter` are load-bearing**: the `A-R`/`a-r` cap keeps it off the `S` of ` Suppl.` and the `V` of a bare ` V2`, and the trailing `match["A-Za-z0-9"].absent?` is what stops ` Amd.`/` Add.`/` Annex A`/` App.`/` Cor.`/` Err.` being read as a qualifier. It cannot collide with `language` (dash-attached) or `date_part`/`version_part` (both need `(` after the space). Because `Code#to_s` can now contain a space, **`Code#compact_s`** (`to_s.delete(" ")`) was added and is what `urn_generator` interpolates — a space is not admissible in a URN segment, and dropping the suffix instead would collapse `D.200` and `D.200 R` onto one URN. The same slots were added inside `combined_designation` (positionally, so the suffix lands on the designation it printed against: `ITU-T D.300 R/E.282 R` sets both, `ITU-T E.211/Q.11 quater` only the last) — which required the parallel edits to `Builder#build_designations` **and** `CombinedIdentifier#combined_to_kv`/`combined_from_kv`, whose row hash is hand-built. Neither spacing flag is in `Code#==`: the two spellings of one number are one document. **`mr_number_with_part` also had to learn both suffixes** (glued to the number, `x-50bis`/`d-200r`) — it read only number/subseries/parts, so every qualified variant collapsed onto its base's MR slug and, worse, `Q.2931 B` and `Q.2931 C` onto each other. This changes the MR string of the handful of pre-existing glued-`bis` ids (`X.50bis`: `itu.t.x-50` → `itu.t.x-50bis`), which is the point — they were wrongly sharing a slug with `X.50`.
|
|
15
|
+
**(4) Version spellings.** `version_part` gained a bare branch, so `v.1`/`v10`/`V2` join the canonical `(V14)` in the same `version` attribute. This is the **one deliberate non-byte-exact normalisation** — all render as `(V##)` — so those strings live in `version_spec.rb` with explicit expectations and **not** in the byte-exact pass fixtures. Safe against the qualifier rule precisely because `V` is outside its `A-R` cap. relaton can now drop its own `v10` → `(V10)` normalisation.
|
|
16
|
+
**(5) Appendix.** `Identifiers::AppendixOfRecommendation` (`_type: pubid:itu:appendix-of-recommendation`) mirrors `AnnexOfRecommendation` one-for-one — including wrapping the appendixed document in **`base.as(:base)`** so the base's own date (`ITU-T G.722 (1988) App. IV (11/2006)`) lands a level down instead of colliding, and defining `mr_supplement_suffix` so the shared MR renderer recurses. Its alternation slot is the annex's reasoning verbatim: **after `supplement_identifier`** (an appendix can carry a trailing `Amd.`/`Err.`), **before `with_series`**. Three registrations are easy to miss and all three are needed: the `identifiers.rb` autoload, the `Builder#build` branch, and widening `urn_generator`'s `AnnexOfRecommendation` check — that check exists because a wrapper whose identity sits one level deeper otherwise emits `urn:itu:itu`. A `material` attribute carries the companion-artefact records (`App. II test vectors`, `App. I Software`), and is in the URN so they stay distinct from the appendix itself.
|
|
17
|
+
**(6) Series groups vs series-code documents — a corpus heuristic, not an ITU rule.** `series_group` (`E-100`, `E100-300`, `G-100`, `Q-500`) is capped at **one** leading letter; `series_code_body` (`EMC-5`, `MES-2`, `QOS-2`, `IMPL-8`, `SEC-QKD`) requires **two or more**. That split is what keeps them from shadowing each other, and it holds across the whole corpus but would misfile a future `AB-100` group or `E-QKD` document — `series_group_spec.rb` locks both sides so a change is visible. `series_group` must be tried **before** `series`, whose greedy non-backtracking `letter.repeat(1,3)` would take the `E` of `E-100` and then fail on the required space with no retry at a shorter length. `series_code_body` is a flat two-token shape (series + dash + alphanumeric number) rather than an optional middle segment, because a greedy `letter.repeat(2)` + optional `-QKD` + required dash fails at end-of-input with no backtracking into the satisfied `.maybe`; it sits **last** in `base` and after `with_series` in `identifier`, where nothing can reach it by accident. The number stays in `code.number` (not `"EMC-5"` in `series`) precisely so **`root.number`** — the key relaton-index bsearches on — is `"5"`/`"QKD"` rather than nil. `series_dash` drives the `-` vs `.` join in `render_base`; the builder's OB check gained a `series_dash.nil?` guard so a hypothetical `ITU-T OB-1` cannot be misrouted into the Operational Bulletin branch. The literal word `series` is its own `series_word` boolean and **cannot** live in the series token, because it also follows a *dotted* code (`ITU-T E.1100 series Suppl. 1`) where folding it in would corrupt the value every index row is keyed on; `Supplement` emits it only when base-less (guarded like `sector`/`series`).
|
|
18
|
+
**(7) Tail.** `attachment` boolean (`ITU-T H.350 attachment`), `range_end` string (`ITU-T Q.120-Q.139` — no conflict with `parts`, which needs digits after the dash), and `Identifier.parse` now collapses runs of whitespace (`ITU-T D.271 (10/2016)`) **after** the shared `Pubid::MAX_INPUT_LENGTH` guard, never before.
|
|
19
|
+
**The cross-cutting lesson (three review findings were all this same mistake): an identity-bearing marker must reach ALL THREE identity surfaces — `to_s`, `to_urn` and `to_mr_string` — not just `==`.** A marker that only reaches `==` still lets two distinct documents share a URN and an output slug. So `generate_base_urn` emits `series`/`attachment`/`to-<range_end>` segments and joins a series-code document with its dash (`urn:itu:t:EMC-5:2003`, not the ambiguous `EMC.5`); `generate_supplement_urn`'s base-less branch emits `series` too; and `mr_number_with_part` carries all five. `Builder#build_supplement` copies `series_word` up from the base alongside sector/series/code for the same reason — render/serialize/`==` all guard on `base.nil?`, so the copy is neutral there and only the MR slug sees it. `spec/pubid/itu/distinctness_spec.rb` is the forcing function: it asserts all four surfaces differ for every marker pair. **Two more review findings worth keeping:** `technical_marker` is bound to the **`Cor.` branch alone**, not the head of `supplement_type` — offered before any token it was accepted on `Technical Err. 1`/`Technical Amd. 1` and then silently dropped by the builder's Corrigendum-only guard, collapsing those onto a *different* document (they are now cleanly rejected); and `series_code_body` carries a `(str("OB") >> dash).absent?` guard, because `ITU-T OB-1` otherwise reached the Recommendation fallback whose `validate_ob_no_sector!` raises an `ArgumentError` that escapes `Identifier.parse`'s `Parslet::ParseFailed` rescue — turning a rejected input into a crash for callers. `Code#to_s` renders glued suffixes before spaced ones, mirroring the grammar (a fixed edition-word-first order printed `Q.11a bis` back as `Q.11 bisa`). **Two known pre-existing gaps, deliberately not fixed and both pinned by specs in `distinctness_spec.rb`:** (a) the ITU URN convention encodes no supplement *type*, so `Amd. 1`/`Cor. 1`/`Err. 1`/`Add. 1` of one base share a URN and MR slug; (b) `build_supplement` copies sector/series/code down from the base but they are deliberately not serialized, so a supplement rebuilt through `from_hash` has none of them and its **MR slug collapses to `itu.<date>`** while every other surface stays symmetric. Both are true on `main` for `Amd.`/`Cor.`/`Err.` before this change — the new types simply inherit them, and `to_s`/`to_hash`/`==` distinguish everything correctly, which is what relaton's gate uses. Both would be fixed by giving `Supplement` an `mr_supplement_suffix` so the shared MR renderer recurses into `base` (as `AnnexOfRecommendation` already does), which changes every existing ITU supplement MR string — a deliberate call, not a tag-along. Also fixed here: `render_base` gained an `elsif series` branch, so a code-less identifier no longer renders a dangling trailing dot. Locked by `spec/pubid/itu/{code_suffixes,combined_suffixes,series_group,tail_forms}_spec.rb`, `spec/pubid/itu/identifiers/{addendum,appendix_of_recommendation,supplement_spelling,technical_corrigendum}_spec.rb`, and 7 new `spec/fixtures/itu/identifiers/pass/*.txt` files drawn from the real corpus. **Still out of scope (both declared not-ours by the hand-off):** ~47 malformed ITU-R docids left by a decommissioned crawler (`ITU-R BO`, `ITU-R M.5-BL-13` — series with no number; dataset bugs, not identifiers), and the 4 space-for-dot `ITU-T G 231` spellings relaton normalises on its side. (hand-off: itu-t-unparseable-forms.)
|
|
20
|
+
|
|
21
|
+
- **ITU-R Reports (`Identifiers::Report`) and the class-strict `Identifier#==`**: ITU-R **Reports** are a publication series that numbers **independently** of Recommendations, so `ITU-R BT.2020-1` names *two* real, both-current documents (Report BT.2020-1, objective quality assessment, 2000; Recommendation BT.2020-1, UHDTV parameter values, 06/2014). 984 of 1,001 report records in `relaton-data-itu` were indexed as `pubid:itu:recommendation`, and 52 editions across 30 numbers (BT.2020, M.2083, M.2134 …) collided outright — the dataset keys files by identifier, so whichever was crawled last silently overwrote the other. ITU disambiguates with the leading word, so the identifier must too. **The leaf** `Identifiers::Report` (`_type: pubid:itu:report`) is shaped exactly like `Recommendation` — same attributes, same `StandardSerialization` flat block, so only `_type` distinguishes the two hashes — and adds just `render_base` (`"Report #{super}"`) and `mr_type` (`[super, "report"].join(".")` → `itu.r.report.bt-2020-1`). **Grammar**: a captured `report_word` marker (unlike `itu_prefix`'s decorative, *uncaptured* `"Recommendation "` literal, which stays as it is) plus a single `report_body` whose `(series >> dot).maybe` covers the with- and without-series shapes at once; `base_report` accepts **both** spellings (`Report ITU-R BT.2020-1` and the downstream `ITU-R Report BT.2020-1`), and `to_s` renders ITU's own leading form for either — so the infix spelling lives in `report_spec.rb` with an explicit expectation, **not** in the byte-exact pass fixtures. Two deliberate exclusions from `report_body`: **no `combined_suffixes`** (a joint `Report X/Y` would reach `Builder#build`'s `combined` branch, build a `CombinedIdentifier` and *silently drop the marker* — a clean parse failure is better), and a **`(str("OB") >> dot).absent?` guard** (an Operational Bulletin is cross-bureau; without it `Report ITU-T OB.1` routes to `SpecialPublication`, marker dropped, or trips `validate_ob_no_sector!`, whose `ArgumentError` escapes `Identifier.parse`'s `Parslet::ParseFailed` rescue). `base_report` is **first** in `rule(:base)` — which is what lets a supplement/annex/appendix *wrap* a Report — and order-free there because `itu_prefix` admits only `Recommendation `/`ITU` and `letter` is `[A-Z]`, so neither `series` nor `series_code_body` can consume `Report`; `report_identifier` sits in `rule(:identifier)` **after** supplement/annex/appendix (which match the longer wrapped form) and **before** `with_series`/`without_series` (which would match the infix spelling's `ITU-R` prefix and then fail on the unconsumed tail). The builder picks the class from `data[:report_marker]` in the existing Recommendation fallthrough; `build_supplement`/`build_annex_of_recommendation` need **no** change, since they recurse via `build(data[:base])`. **`Pubid::Itu::Identifier#==` is now class-strict** (`other.instance_of?(self.class)`, the rule `Supplement#==` already used) — the ITU *type* is part of the identity, not just sector/series/number. `self.class` rather than a hard-coded class keeps it **symmetric in both directions** (an `is_a?` guard would have made a Recommendation equal a Report one-way, and `#matches?` — which is `exclude(*ignore) == other.exclude(*ignore)` — resolve one to the other). Strictly stricter, so it can only turn `true` into `false`; it also closes the pre-existing one-way holes against `CombinedIdentifier` and `SpecialPublication`. **Consequences of the shared guard, both pinned by `spec/pubid/itu/type_strict_equality_spec.rb`:** comparisons that used to be *asymmetric* are now consistently false — most notably a bare primary designation vs its joint recommendation (`ITU-T G.780` vs `ITU-T G.780/Y.1351`, which was `true` one way and `false` the other on `main`, so no caller could rely on it either way). A caller matching a bare primary against a joint document must narrow on `root.number` (which still agrees) rather than on `==`. Per the cross-cutting ITU lesson, the marker reaches **all** identity surfaces: `to_s`, `generate_base_urn` (a `report` segment right after the sector — `urn:itu:r:report:BT.2020-1:2000`, which `UrnParser` also reads back), and the MR slug — including **`Supplement#mr_type`**, which is load-bearing and easy to miss: a supplement has no `mr_supplement_suffix`, so the shared MR renderer slugs it **flat** from the sector/series/code `build_supplement` copied up from its base and *never consults the base's class*, which made `Report ITU-R BT.2020-1 Suppl. 1` and `ITU-R BT.2020-1 Suppl. 1` share one slug (i.e. one output filename — the very overwrite this type prevents). It appends `report` when **`root`** is a Report, so a supplement of an *annex* of a Report is covered too. **`PREFIXES` stays `["ITU"]`** — the leading `Report` token is deliberately not registered for prefix routing (a generic English word; the mixin's ambiguous-token exclusion), and the infix spelling routes on `ITU` as before. **Backwards compatibility is exact**: a bare `ITU-R BT.2020-1` still builds a `Recommendation` with an unchanged hash/`to_s`/URN — nothing in the grammar or the builder fires without the marker, so none of the 20,728 published index rows move. Locked by `spec/pubid/itu/identifiers/report_spec.rb`, a `distinctness_spec.rb` pair, rows in `root_number_spec.rb`/`serialization_spec.rb`/`urn_parser_spec.rb`, and `spec/fixtures/itu/identifiers/pass/report.txt`. **Out of scope (relaton's side, sequenced after this):** `DataParserR`/`DataCrawlerR` emitting the new docid form, `Relaton::Itu::Pubid`'s reference parser learning a `report` rule, and the `relaton-data-itu` filename-namespace migration + re-crawl that actually recovers the 34 missing Recommendations and 18 missing Reports. (hand-off: itu-r-report-identifier-type.)
|
|
22
|
+
|
|
23
|
+
## From the root note "Parse-failure error contract — uniform across every flavor"
|
|
24
|
+
|
|
25
|
+
**ITU's length guard was an early `return`, not a raise, and its comment blamed the wrong caller.** `Identifier.normalize_whitespace` (`lib/pubid/itu/identifiers/base.rb`) carried
|
|
26
|
+
|
|
27
|
+
```ruby
|
|
28
|
+
return identifier if identifier.length > ::Pubid::MAX_INPUT_LENGTH
|
|
29
|
+
```
|
|
30
|
+
|
|
31
|
+
under a note claiming "`Pubid.parse` rejects it before this point". That was wrong twice: `Pubid.parse` never routes a human-readable string to a flavor at all (`lib/pubid.rb` raises `ArgumentError: No flavor specified` for anything that is not an MR string or a URN), and ITU had no raising guard of its own — so an over-long string reached the ITU parser through both `Pubid::Itu.parse` and `Pubid::Itu::Identifier.parse`. The guard now lives in `self.parse` as the standard inline pair, and the comment says what actually protects the method.
|
|
32
|
+
|
|
33
|
+
**ITU was also the delegate in the most-visible class leak.** `Pubid::Iso.parse("ITU-T G.711")` detects an MR-shaped string, routes through `Pubid::Parsers::MrString`, and lands in `Pubid::Itu::Identifier.parse` — which raised `RuntimeError`. So `Pubid::Iso.parse` raised **two different classes** depending on its input, and a relaton-cli caller rescuing `Parslet::ParseFailed` got a raw backtrace for the ITU-shaped half. Converting ITU fixed the ISO symptom; `spec/pubid/parse_error_spec.rb` pins that exact route by name, because a generic junk string never reaches it.
|
|
34
|
+
|
|
35
|
+
- **The five supplement types rendered plain under `to_s(annotated: true)`.**
|
|
36
|
+
`Pubid::Itu::Identifier#to_s` annotates, but `Addendum`, `Amendment`,
|
|
37
|
+
`Corrigendum`, `Errata` and `Supplement` each override it and none calls
|
|
38
|
+
`super` — they all funnel through `Supplement#render_supplement(label)`.
|
|
39
|
+
The wrap cannot live in that helper, which takes a label and no options,
|
|
40
|
+
so each of the five `to_s` is a one-line
|
|
41
|
+
`annotate_plain_render(render_supplement("…"), **opts)`.
|
|
42
|
+
|
|
43
|
+
`render_supplement` starts from `base.to_s` with no options, so the base
|
|
44
|
+
arrives plain and the single outer annotation covers the whole string —
|
|
45
|
+
the base's own tokens are reached by `Annotator#emit_tokens` walking
|
|
46
|
+
`base`.
|
|
47
|
+
|
|
48
|
+
## Contribution (Temporary Document) — pubid#340
|
|
49
|
+
|
|
50
|
+
`Pubid::Itu::Identifiers::Contribution` models the working documents a study
|
|
51
|
+
group circulates, mirroring pubid-itu 1.15's `Identifier::Contribution`
|
|
52
|
+
(`%{series}-C%{number}`): "ITU-R SG17-C1000", sector- and language-suffixed
|
|
53
|
+
per the existing rules ("ITU-T SG17-C1000-E"). The grammar entry sits between
|
|
54
|
+
`with_series` and `series_code_identifier` — the **-C marker (dash + "C" +
|
|
55
|
+
digits) is the discriminator**: a series-code document's post-dash number
|
|
56
|
+
starts with digits or is all letters, never "C"+digits, so the two dash
|
|
57
|
+
shapes cannot shadow each other in either direction ("ITU-T EMC-5" still
|
|
58
|
+
builds a Recommendation, locked by a spec). The builder branch carries the
|
|
59
|
+
`:contribution_marker`; `root.number` reaches the C-number through the shared
|
|
60
|
+
`code.number` reader, so relaton-index keys it normally.
|
|
61
|
+
|
|
62
|
+
**`locate_type` was broken for every leaf, not just this one**: no ITU class
|
|
63
|
+
ever defined a `type` hash, so `Pubid::Itu.locate_type` raised `NoMethodError`
|
|
64
|
+
on any call. The base now derives the key from the class name
|
|
65
|
+
(`Contribution` → `:contribution`, `AnnexOfRecommendation` →
|
|
66
|
+
`:annex_of_recommendation`) — plain `gsub` camel→snake, because ActiveSupport's
|
|
67
|
+
`underscore` is not a dependency. metanorma-itu constructs through this
|
|
68
|
+
lookup; its flavor-local `pubid_contribution.rb` render override can be
|
|
69
|
+
deleted once it migrates.
|
|
70
|
+
|
|
71
|
+
## relaton's query forms — RR, OB sector, publication ids
|
|
72
|
+
|
|
73
|
+
Four forms relaton's `Relaton::Itu::Pubid` parsed and `Pubid::Itu` did not
|
|
74
|
+
(hand-off itu-relaton-query-forms). None of them occurs in the published
|
|
75
|
+
`relaton-data-itu` index — no `series: RR`, no OB row, no six-digit part — and
|
|
76
|
+
a replay of all 24,382 rows and every ITU pass fixture showed 0 changes.
|
|
77
|
+
|
|
78
|
+
- **Radio Regulations** — `Identifiers::RadioRegulations`
|
|
79
|
+
(`pubid:itu:radio-regulations`): `ITU-R RR`, `ITU-R RR (2020)`, and the URL
|
|
80
|
+
spelling `ITU-R RR-2020`, which used to build a *wrong* Recommendation
|
|
81
|
+
(series `RR`, number `2020`) with no error. "RR" is the series; there is no
|
|
82
|
+
code, so `#number` returns the series and `root.number` is `"RR"`. The rule
|
|
83
|
+
ends in `any.absent?`: it sits before `with_series`, and PEG never re-enters
|
|
84
|
+
the alternation, so a partial match on `ITU-R RR.1` must fail inside it.
|
|
85
|
+
- **Operational Bulletins keep their sector.** This reverses the old
|
|
86
|
+
"cross-bureau, sector must not be set" rule: `validate_ob_no_sector!` is
|
|
87
|
+
gone, `ITU-T OB.1096 (2016)` renders back as it is, and the sector-less
|
|
88
|
+
`ITU OB No. 1096` is unchanged. The long forms with a sector (`ITU-T OB No.
|
|
89
|
+
1096`) now normalise to `ITU-T OB.1096`. `No.` stays in the default
|
|
90
|
+
render — it is how ITU's bulletin site and pubid v1 write it — and
|
|
91
|
+
metanorma-itu's `ITU OB 1000` / `Annex to ITU OB 1000` (its i18n template
|
|
92
|
+
omits `No.`) is accepted on parse (`ob_bare_body`) and rendered with it.
|
|
93
|
+
The sector is a spelling, not
|
|
94
|
+
identity: `SpecialPublication#==` skips it, the URN keeps `urn:itu:itu:…`
|
|
95
|
+
and `mr_type` stays nil, so both spellings are one bulletin on every
|
|
96
|
+
surface except `to_s`/`to_hash`. The printed date `- 15.III.2016` sets
|
|
97
|
+
`date.day`, which only this form does, so `render_ob_date` uses the day as
|
|
98
|
+
the spelling marker; `day_to_kv` emits only then.
|
|
99
|
+
- **`-YYYYMM` is a date, never a part.** `part` refuses a six-digit run with a
|
|
100
|
+
19xx/20xx year and a 01–12 month (`yyyymm_shape`), and `id_date` reads it as
|
|
101
|
+
year+month. `-200313`, `-180001` stay parts. `ITU-T REC T.4` drops the
|
|
102
|
+
uncaptured `rec_word`; `T-REC-T.4-200307-I` is `publication_id`, last in
|
|
103
|
+
`identifier` (nothing else starts with a bare sector letter), and requires
|
|
104
|
+
the date. The trailing status letter (`I` in force, `S` superseded) is
|
|
105
|
+
parsed and dropped: it names the state of an edition, not the edition.
|
|
106
|
+
**`S` is also the Spanish language suffix**, so it is a status only inside
|
|
107
|
+
the full `T-REC-…` id (`id_status`); after an `ITU-T …-YYYYMM` print form
|
|
108
|
+
only `-I` is (`print_id_status`), and `ITU-T Z.100-199911-S` keeps language
|
|
109
|
+
`S`. The `-YYYYMM` date and `REC` word are also in `base_with_series`/
|
|
110
|
+
`base_without_series`: the part guard applies there too, so without them
|
|
111
|
+
`ITU-T G.989-200307 Amd 1` — a (wrong) Amendment on `main` — stopped parsing.
|
|
112
|
+
The day of an OB date reaches the URN (`…:15/03/2016`), since it is in `==`.
|
|
113
|
+
**Not done:** an RR supplement (`ITU-R RR (2020) Amd 1` fails; the
|
|
114
|
+
base-less `ITU-R RR Amd 1` still builds an Amendment on series `RR`), and
|
|
115
|
+
`ITU-R RR-E` is still a Recommendation numbered `E` — both as on `main`.
|