gigatoken 0.4.0 → 0.5.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: 1fd5d74bb0d8cb5be9f7cf9fa077cc9c695ab5279436f3ecf3db4374952be586
4
- data.tar.gz: 5953850c0513e14bb6aceccf4205b774c1db027f6d4d0f4508d90107f0d483ef
3
+ metadata.gz: a04a62cbb60179c04c4d839c8dd6723fe737e6e90d30962412a304f9a73973ce
4
+ data.tar.gz: b02aa32bf21824f818c8ba0e85e1b8c915a93193e6d2dad0226219f96176452e
5
5
  SHA512:
6
- metadata.gz: 9b85351d2af130d4f3b0b96c1f654b94ca5a865bd8865b6bd7d29ce2c5d297f052e8ee4a7ee40391573b88b3d81df5ba3aff8fd6edf64e6e03dee596644f1ae9
7
- data.tar.gz: c8aca888f7b6477e406fa1234633e8e334a3242d9056951153c2306a6cc5b5eda8d46636c00bfd2c515ecf309130ffaf3543ced5305cc01e00473d8cf2966a73
6
+ metadata.gz: 52af5b802c0336248dcfc2cea0e59a4db0a7af2a5e3774c498635ec68903d3f3a7cc55c8b6f7aaafeb7649be8c35105c213bf829aef7c6f206238a0400e355bb
7
+ data.tar.gz: c9be21f7f201693e31d90b2021af50c2e427b7c3cbc174f8715bda7ef469ad70882c433d49e9006e8e8f4296af434d660e23ac513de5f6cc368a8db359b35837
data/Cargo.lock CHANGED
@@ -771,7 +771,7 @@ dependencies = [
771
771
 
772
772
  [[package]]
773
773
  name = "gigatoken-rb"
774
- version = "0.4.0"
774
+ version = "0.5.0"
775
775
  dependencies = [
776
776
  "gigatoken",
777
777
  "magnus",
data/README.md CHANGED
@@ -61,9 +61,9 @@ A `.tiktoken` file holds mergeable ranks only — its pretokenization scheme and
61
61
 
62
62
  SentencePiece-BPE models (Llama, Gemma, Mistral — any `tokenizer.json` with `byte_fallback: true`) load through the same entry points and pick the right backend automatically. One difference: the SentencePiece core decodes text, so it validates input and raises `Gigatoken::InputError` on invalid UTF-8 instead of guessing.
63
63
 
64
- ### Packaged tiktoken encodings
64
+ ### Packaged encodings
65
65
 
66
- `r50k_base`, `cl100k_base`, `o200k_base`, and `o200k_harmony` are vendored directly — mergeable ranks, pretokenizer scheme, and special-token table all shipped inside the gem (`lib/gigatoken/encodings/`; see `PROVENANCE.md` there for exact source URLs and hashes) — so all four resolve by name through both entry points entirely offline: no network access, no writable cache directory. `o200k_harmony` vendors no new file at all: it reuses `o200k_base.tiktoken`'s ranks and the `o200k` scheme verbatim, differing only in its special-token table (10 named control tokens — `<|start|>`, `<|message|>`, `<|end|>`, `<|return|>`, and so on — plus 1081 reserved slots; see `PROVENANCE.md` for the exact table). It's also the one packaged encoding not checked against `tiktoken_ruby`: that gem's 0.0.17 harmony table drops `<|endofprompt|>` where `openai/tiktoken` 0.9.0 keeps it at id 200018, so the oracle is the outlier here — `spec/gigatoken/differential_spec.rb` proves harmony instead by reduction to `o200k_base` plus a pinned special-token table.
66
+ `r50k_base`, `cl100k_base`, `o200k_base`, and `o200k_harmony` are vendored directly — mergeable ranks, pretokenizer scheme, and special-token table all shipped inside the gem (`lib/gigatoken/encodings/`; see `PROVENANCE.md` there for exact source URLs and hashes) — as are `qwen35`, `qwen38` and `muse_spark`, below, so all seven resolve by name through both entry points entirely offline: no network access, no writable cache directory. `o200k_harmony` vendors no new file at all: it reuses `o200k_base.tiktoken`'s ranks and the `o200k` scheme verbatim, differing only in its special-token table (10 named control tokens — `<|start|>`, `<|message|>`, `<|end|>`, `<|return|>`, and so on — plus 1081 reserved slots; see `PROVENANCE.md` for the exact table). It's also the one packaged encoding not checked against `tiktoken_ruby`: that gem's 0.0.17 harmony table drops `<|endofprompt|>` where `openai/tiktoken` 0.9.0 keeps it at id 200018, so the oracle is the outlier here — `spec/gigatoken/differential_spec.rb` proves harmony instead by reduction to `o200k_base` plus a pinned special-token table.
67
67
 
68
68
  ```ruby
69
69
  Gigatoken::Tokenizer.from_encoding("cl100k_base")
@@ -71,6 +71,8 @@ Gigatoken::Tokenizer.load("cl100k_base") # same result — packaged nam
71
71
  # checked before the Hub-repo-id shape
72
72
  ```
73
73
 
74
+ `qwen35`, `qwen38` and `muse_spark` are whole HuggingFace `tokenizer.json` files, gzipped (`*.json.gz`, pinned to a repo revision in `PROVENANCE.md`) and loaded through `from_json`, so their special tokens come from the file. `qwen35` is the Qwen 3.5 and 3.6 tokenizer; `qwen38` is that file plus seven audio/TTS special tokens (`<|audio_start|>`, `<|audio_end|>`, `<tts_pad>`, `<tts_text_bos>`, `<tts_text_eod>`, `<tts_text_bos_single>`, `<|audio_pad|>`), so ordinary text encodes the same under both; `muse_spark` is `meta-models/Muse-Glimmer-30B`'s tokenizer, because Meta publishes no standalone Muse Spark one. Their output matches HuggingFace `tokenizers` with `add_special_tokens: false`: gigatoken applies no post-processor, so `encode` never prepends `<|begin_of_text|>` for `muse_spark`. Unlike the tiktoken entries, a registry entry for one of these is just `{json_file:}`, and each load reads, gunzips and parses its file (about 0.4 s).
75
+
74
76
  `p50k_base` and `p50k_edit` are deliberately not packaged: both load the same non-dense ranks (id 50256 is left free for `<|endoftext|>`), and the rank loader rejects non-dense ranks. Both entry points raise `Gigatoken::ModelError` explaining that, rather than `load` falling through to the Hub for a name that happens to look like a legacy repo id.
75
77
 
76
78
  `encode` on a packaged tokenizer honours its special-token table: text containing `<|endoftext|>` (or any other literal special-token string) is tokenized as that special token, not as ordinary text. That matches [`tiktoken`](https://github.com/openai/tiktoken)'s `encode_with_special_tokens`, not its plain `encode`, which treats the same literal as ordinary text — a difference worth knowing if you're tokenizing untrusted input. To get tiktoken's non-honouring default instead, build a tokenizer from the same rank file with an empty special-token table:
@@ -1,6 +1,6 @@
1
1
  [package]
2
2
  name = "gigatoken-rb"
3
- version = "0.4.0"
3
+ version = "0.5.0"
4
4
  edition = "2021"
5
5
 
6
6
  [lib]
@@ -1,4 +1,4 @@
1
- # Vendored tiktoken encodings — provenance
1
+ # Vendored encodings — provenance
2
2
 
3
3
  The `.tiktoken` files in this directory are OpenAI's published BPE mergeable-rank
4
4
  tables, vendored verbatim so `gigatoken` can resolve these encodings by name with
@@ -9,6 +9,10 @@ carries **mergeable ranks only** — the pretokenizer split regex and the specia
9
9
  tokens belong to the encoding's *definition*, not to the file, and live in code
10
10
  (see the table below).
11
11
 
12
+ The three `.json.gz` files are HuggingFace `tokenizer.json` files, not tiktoken
13
+ tables; they are covered in [their own section](#huggingface-tokenizerjson-files)
14
+ at the end.
15
+
12
16
  They live under `lib/` rather than a top-level `data/` directory because this
13
17
  repo's `.gitignore` ignores `/data/` ("downloaded test data"); these are shipped
14
18
  gem payload, not test fixtures.
@@ -108,3 +112,111 @@ The `tiktoken` project and its published encoding files are MIT licensed,
108
112
  Copyright (c) 2022 OpenAI, Shantanu Jain. See
109
113
  <https://github.com/openai/tiktoken/blob/main/LICENSE>. The files are vendored
110
114
  here unmodified.
115
+
116
+ ## HuggingFace `tokenizer.json` files
117
+
118
+ `qwen35.json.gz`, `qwen38.json.gz` and `muse_spark.json.gz` are HuggingFace
119
+ `tokenizer.json` files, vendored byte-for-byte and gzipped. Unlike a
120
+ `.tiktoken` table, a `tokenizer.json` carries the encoding's whole definition
121
+ — vocabulary, merges, normalizer, pretokenizer, added tokens — so the registry
122
+ records only the file, and `Tokenizer.from_encoding` loads it with
123
+ `Tokenizer.from_json`, reading the special tokens from the file's
124
+ `added_tokens` (those marked `"special": true`) exactly as `from_json` does.
125
+
126
+ ### Files
127
+
128
+ | File | Decompressed bytes | sha256 (decompressed) |
129
+ |------|-------------------:|-----------------------|
130
+ | `qwen35.json.gz` | 12,807,982 | `5f9e4d4901a92b997e463c1f46055088b6cca5ca61a6522d1b9f64c4bb81cb42` |
131
+ | `qwen38.json.gz` | 12,809,320 | `0997f410c57a1f4e53b09e4be8f4a172d90edd9564368fb0847030937229b9f3` |
132
+ | `muse_spark.json.gz` | 28,129,897 | `c9dbee66967b58f31a7c27f723c3760da3526ccd0427578e8905b0abb0031c4d` |
133
+
134
+ ### Source
135
+
136
+ Retrieved **2026-10-07** over HTTPS from the Hub's `resolve` endpoint, each at
137
+ a pinned commit:
138
+
139
+ | Encoding | Repo | Revision | URL |
140
+ |---|---|---|---|
141
+ | `qwen35` | `Qwen/Qwen3.5-27B` | `fc05daec18b0a78c049392ed2e771dde82bdf654` | <https://huggingface.co/Qwen/Qwen3.5-27B/resolve/fc05daec18b0a78c049392ed2e771dde82bdf654/tokenizer.json> |
142
+ | `qwen38` | `Qwen/Qwen3.8-27B` | `1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0` | <https://huggingface.co/Qwen/Qwen3.8-27B/resolve/1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0/tokenizer.json> |
143
+ | `muse_spark` | `meta-models/Muse-Glimmer-30B` | `a4e59da52a7bc87ae7251dd5545c0dd437c44b68` | <https://huggingface.co/meta-models/Muse-Glimmer-30B/resolve/a4e59da52a7bc87ae7251dd5545c0dd437c44b68/tokenizer.json> |
144
+
145
+ **`qwen35` is shared across releases.** The Hub's tree API
146
+ (`https://huggingface.co/api/models/<repo>?blobs=true`) reports the identical
147
+ LFS sha256 `5f9e4d49…` for `tokenizer.json` in `Qwen/Qwen3.5-{0.8B, 2B, 4B, 9B,
148
+ 27B, 35B-A3B, 122B-A10B, 397B-A17B}` and `Qwen/Qwen3.6-{27B, 35B-A3B}`
149
+ (queried 2026-10-07), so one file serves Qwen 3.5 and 3.6.
150
+ `Qwen/Qwen3.6-Plus` is gated and was not checked.
151
+
152
+ **`qwen38` ships its own file**, identical (`0997f410…`) across
153
+ `Qwen/Qwen3.8-{27B, Flash-Next, 2.4T-A95B}` and their `-FP8` twins. It differs
154
+ from `qwen35` only in seven more `"special": true` added tokens, ids
155
+ 248070–248076: `<|audio_start|>`, `<|audio_end|>`, `<tts_pad>`,
156
+ `<tts_text_bos>`, `<tts_text_eod>`, `<tts_text_bos_single>`, `<|audio_pad|>`.
157
+ The vocabulary, merges, normalizer, pretokenizer, decoder and post-processor
158
+ are identical, so ordinary text encodes the same under both; only those seven
159
+ literals, `vocab_size` and `special_tokens` differ. It is vendored as a second
160
+ file rather than derived from the first because the engine reads added tokens
161
+ from the JSON, and a JSON-backed tokenizer has no way to add specials after
162
+ the fact. `Qwen3.8-27B` is the repo cited; `Flash-Next` and `2.4T-A95B` carry an
163
+ "other" licence on the *model*, but their tokenizer file is byte-identical to
164
+ the Apache-2.0 repo's.
165
+
166
+ **`muse_spark` is the Muse-Glimmer-30B file.** Meta publishes no standalone
167
+ Muse Spark tokenizer (a Hub search over `meta-models` finds none, and Meta's
168
+ model-API docs at <https://dev.meta.ai/docs/muse-glimmer> point at none). The
169
+ Muse-Glimmer-30B model card states "Tokenizer: 200,000 BPE tokens + 2,048
170
+ special tokens" (vocabulary 202,048), trained from Muse Spark's outputs, and
171
+ this exact repo and revision is the one the reference application pins. As of
172
+ 2026-10-07 that revision is also `main`.
173
+
174
+ ### Authenticity
175
+
176
+ Verified three independent ways at retrieval time:
177
+
178
+ 1. **Transport** — HTTPS directly from `huggingface.co`, at the pinned
179
+ commit above.
180
+ 2. **Publisher checksum** — each file's sha256 above matches the LFS sha256 the
181
+ Hub's tree API reports for `tokenizer.json` at that revision, and so does
182
+ its size. Two independent channels agree on the bytes.
183
+ 3. **Behavioral, against the reference implementation** — loaded through
184
+ `Tokenizer.from_encoding`, every token id matched HuggingFace
185
+ [`tokenizers`](https://rubygems.org/gems/tokenizers) 0.7.0 loading the same
186
+ decompressed file with `add_special_tokens: false`, across this repo's own
187
+ source and docs (`spec/gigatoken/differential_spec.rb`). Resulting
188
+ `vocab_size`: 248070 / 248077 / 202048; `special_tokens.size`: 14 / 21 / 2048.
189
+
190
+ ### Compression
191
+
192
+ Each file was compressed with `gzip -9 -n` (no name or timestamp in the gzip
193
+ header, so the bytes are reproducible), and `gunzip -c` of it reproduces the
194
+ published file byte-for-byte — the sha256s above are of the decompressed
195
+ bytes. Compressed sizes: 3,514,616 (`qwen35`), 3,514,699 (`qwen38`), 4,475,666
196
+ (`muse_spark`).
197
+
198
+ ### What gigatoken does not take from the file
199
+
200
+ - **The post-processor.** gigatoken applies none. `muse_spark`'s is a
201
+ `TemplateProcessing` that prepends `<|begin_of_text|>`; `Tokenizer#encode`
202
+ does not, matching `tokenizers`' `add_special_tokens: false`. The Qwen files'
203
+ is a `ByteLevel` post-processor, a no-op on ids.
204
+ - **Non-special added tokens in `special_tokens`.** Qwen's 12 non-special
205
+ added tokens (`<tool_call>`, `<think>`, the `<|fim_*|>` tokens,
206
+ `<|repo_name|>`, `<|file_sep|>`, `<tool_response>` and so on) are matched
207
+ atomically by the engine, like the special ones, but only `"special": true`
208
+ tokens appear in `Tokenizer#special_tokens` — 14 for `qwen35`, 21 for
209
+ `qwen38`.
210
+ - **`tokenizer_config.json`.** Not vendored; gigatoken reads only
211
+ `tokenizer.json`.
212
+
213
+ ### Licence
214
+
215
+ All three repos are Apache-2.0, each with its own `LICENSE` at the pinned
216
+ revision:
217
+
218
+ - <https://huggingface.co/Qwen/Qwen3.5-27B/blob/fc05daec18b0a78c049392ed2e771dde82bdf654/LICENSE>
219
+ - <https://huggingface.co/Qwen/Qwen3.8-27B/blob/1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0/LICENSE>
220
+ - <https://huggingface.co/meta-models/Muse-Glimmer-30B/blob/a4e59da52a7bc87ae7251dd5545c0dd437c44b68/LICENSE>
221
+
222
+ The files are vendored here unmodified (gzipped).
@@ -1,10 +1,11 @@
1
1
  # frozen_string_literal: true
2
2
 
3
3
  module Gigatoken
4
- # The tiktoken encodings gigatoken vendors ranks for, and the pieces a
5
- # .tiktoken file doesn't carry: its pretokenizer scheme and special-token
6
- # table (see Tokenizer.from_tiktoken and lib/gigatoken/encodings/
7
- # PROVENANCE.md, the source of truth this is transcribed from).
4
+ # The encodings gigatoken vendors a file for. For the tiktoken ones, the
5
+ # pieces a .tiktoken file doesn't carry: its pretokenizer scheme and
6
+ # special-token table (see Tokenizer.from_tiktoken and lib/gigatoken/
7
+ # encodings/PROVENANCE.md, the source of truth this is transcribed from).
8
+ # The HuggingFace ones are whole tokenizer.json files, which carry theirs.
8
9
  module Encodings
9
10
  DATA_DIR = File.expand_path("encodings", __dir__)
10
11
  private_constant :DATA_DIR
@@ -40,10 +41,10 @@ module Gigatoken
40
41
  HARMONY_RESERVED_TAIL = (200013..201087).to_h { |id| ["<|reserved_#{id}|>", id] }.freeze
41
42
  private_constant :HARMONY_RESERVED_TAIL
42
43
 
43
- # Deep-frozen: entries, their `special_tokens` tables and the `rank_file`
44
- # paths. `Tokenizer#special_tokens` hands the registry's own Hash back to
45
- # callers, so anything less lets one caller's poke rewrite what every
46
- # later `from_encoding` in the process loads.
44
+ # Deep-frozen: entries, their `special_tokens` tables and the `rank_file` /
45
+ # `json_file` paths. `Tokenizer#special_tokens` hands the registry's own
46
+ # Hash back to callers, so anything less lets one caller's poke rewrite
47
+ # what every later `from_encoding` in the process loads.
47
48
  REGISTRY = {
48
49
  "r50k_base" => {
49
50
  rank_file: File.join(DATA_DIR, "r50k_base.tiktoken"),
@@ -70,7 +71,21 @@ module Gigatoken
70
71
  rank_file: File.join(DATA_DIR, "o200k_base.tiktoken"),
71
72
  pretokenizer: "o200k",
72
73
  special_tokens: HARMONY_HEAD_TOKENS.merge(HARMONY_RESERVED_TAIL).freeze
73
- }
74
+ },
75
+ # The three below are gzipped HuggingFace tokenizer.json files, vendored
76
+ # verbatim (see PROVENANCE.md for repos, revisions and hashes). Nothing
77
+ # else is registered: the pretokenizer and the special tokens are the
78
+ # file's own, read at load as Tokenizer.from_json reads them.
79
+ #
80
+ # Qwen 3.5 and 3.6 share this one file.
81
+ "qwen35" => {json_file: File.join(DATA_DIR, "qwen35.json.gz")},
82
+ # The qwen35 file plus seven audio/TTS special tokens — vendored
83
+ # separately because a JSON-backed tokenizer takes its added tokens
84
+ # from the file.
85
+ "qwen38" => {json_file: File.join(DATA_DIR, "qwen38.json.gz")},
86
+ # meta-models/Muse-Glimmer-30B's file: Meta publishes no standalone
87
+ # Muse Spark tokenizer.
88
+ "muse_spark" => {json_file: File.join(DATA_DIR, "muse_spark.json.gz")}
74
89
  }.each_value { |encoding| encoding.each_value(&:freeze).freeze }.freeze
75
90
  private_constant :REGISTRY
76
91
 
@@ -91,9 +106,9 @@ module Gigatoken
91
106
  private_constant :UNPACKABLE_REASONS
92
107
 
93
108
  class << self
94
- # The {rank_file:, pretokenizer:, special_tokens:} registered for a
95
- # packaged encoding name, or nil. Names are Strings or Symbols, as
96
- # Tokenizer.load accepts both.
109
+ # The {rank_file:, pretokenizer:, special_tokens:} (tiktoken) or
110
+ # {json_file:} (HuggingFace) registered for a packaged encoding name, or
111
+ # nil. Names are Strings or Symbols, as Tokenizer.load accepts both.
97
112
  def [](name)
98
113
  REGISTRY[name.to_s]
99
114
  end
@@ -1,6 +1,7 @@
1
1
  # frozen_string_literal: true
2
2
 
3
3
  require "json"
4
+ require "zlib"
4
5
 
5
6
  module Gigatoken
6
7
  # A tokenizer: encode, batch encode, decode, and vocabulary introspection
@@ -52,11 +53,13 @@ module Gigatoken
52
53
  new(native, special_tokens: special_tokens)
53
54
  end
54
55
 
55
- # Load one of the tiktoken encodings gigatoken vendors ranks for, by
56
- # name — see Gigatoken::Encodings::NAMES — entirely from the vendored
57
- # files: no network, no writable cache.
56
+ # Load one of the encodings gigatoken vendors a file for, by name — see
57
+ # Gigatoken::Encodings::NAMES — entirely from the vendored files: no
58
+ # network, no writable cache. A tiktoken entry loads its ranks; a
59
+ # tokenizer.json entry is gunzipped and loaded by from_json.
58
60
  def self.from_encoding(name)
59
61
  encoding = Encodings[name]
62
+ return from_json(Zlib.gunzip(File.binread(encoding[:json_file]))) if encoding&.key?(:json_file)
60
63
  return from_tiktoken(encoding[:rank_file], pretokenizer: encoding[:pretokenizer], special_tokens: encoding[:special_tokens]) if encoding
61
64
 
62
65
  reason = Encodings.unpackable_reason(name)
@@ -1,5 +1,5 @@
1
1
  # frozen_string_literal: true
2
2
 
3
3
  module Gigatoken
4
- VERSION = "0.4.0"
4
+ VERSION = "0.5.0"
5
5
  end
metadata CHANGED
@@ -1,7 +1,7 @@
1
1
  --- !ruby/object:Gem::Specification
2
2
  name: gigatoken
3
3
  version: !ruby/object:Gem::Version
4
- version: 0.4.0
4
+ version: 0.5.0
5
5
  platform: ruby
6
6
  authors:
7
7
  - Eric Jacobs
@@ -107,7 +107,10 @@ files:
107
107
  - lib/gigatoken/encodings.rb
108
108
  - lib/gigatoken/encodings/PROVENANCE.md
109
109
  - lib/gigatoken/encodings/cl100k_base.tiktoken
110
+ - lib/gigatoken/encodings/muse_spark.json.gz
110
111
  - lib/gigatoken/encodings/o200k_base.tiktoken
112
+ - lib/gigatoken/encodings/qwen35.json.gz
113
+ - lib/gigatoken/encodings/qwen38.json.gz
111
114
  - lib/gigatoken/encodings/r50k_base.tiktoken
112
115
  - lib/gigatoken/hub.rb
113
116
  - lib/gigatoken/packed_result.rb