gigatoken 0.4.0 → 0.5.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/Cargo.lock +1 -1
- data/README.md +4 -2
- data/ext/gigatoken/Cargo.toml +1 -1
- data/lib/gigatoken/encodings/PROVENANCE.md +113 -1
- data/lib/gigatoken/encodings/muse_spark.json.gz +0 -0
- data/lib/gigatoken/encodings/qwen35.json.gz +0 -0
- data/lib/gigatoken/encodings/qwen38.json.gz +0 -0
- data/lib/gigatoken/encodings.rb +27 -12
- data/lib/gigatoken/tokenizer.rb +6 -3
- data/lib/gigatoken/version.rb +1 -1
- metadata +4 -1
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: a04a62cbb60179c04c4d839c8dd6723fe737e6e90d30962412a304f9a73973ce
|
|
4
|
+
data.tar.gz: b02aa32bf21824f818c8ba0e85e1b8c915a93193e6d2dad0226219f96176452e
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: 52af5b802c0336248dcfc2cea0e59a4db0a7af2a5e3774c498635ec68903d3f3a7cc55c8b6f7aaafeb7649be8c35105c213bf829aef7c6f206238a0400e355bb
|
|
7
|
+
data.tar.gz: c9be21f7f201693e31d90b2021af50c2e427b7c3cbc174f8715bda7ef469ad70882c433d49e9006e8e8f4296af434d660e23ac513de5f6cc368a8db359b35837
|
data/Cargo.lock
CHANGED
data/README.md
CHANGED
|
@@ -61,9 +61,9 @@ A `.tiktoken` file holds mergeable ranks only — its pretokenization scheme and
|
|
|
61
61
|
|
|
62
62
|
SentencePiece-BPE models (Llama, Gemma, Mistral — any `tokenizer.json` with `byte_fallback: true`) load through the same entry points and pick the right backend automatically. One difference: the SentencePiece core decodes text, so it validates input and raises `Gigatoken::InputError` on invalid UTF-8 instead of guessing.
|
|
63
63
|
|
|
64
|
-
### Packaged
|
|
64
|
+
### Packaged encodings
|
|
65
65
|
|
|
66
|
-
`r50k_base`, `cl100k_base`, `o200k_base`, and `o200k_harmony` are vendored directly — mergeable ranks, pretokenizer scheme, and special-token table all shipped inside the gem (`lib/gigatoken/encodings/`; see `PROVENANCE.md` there for exact source URLs and hashes) — so all
|
|
66
|
+
`r50k_base`, `cl100k_base`, `o200k_base`, and `o200k_harmony` are vendored directly — mergeable ranks, pretokenizer scheme, and special-token table all shipped inside the gem (`lib/gigatoken/encodings/`; see `PROVENANCE.md` there for exact source URLs and hashes) — as are `qwen35`, `qwen38` and `muse_spark`, below, so all seven resolve by name through both entry points entirely offline: no network access, no writable cache directory. `o200k_harmony` vendors no new file at all: it reuses `o200k_base.tiktoken`'s ranks and the `o200k` scheme verbatim, differing only in its special-token table (10 named control tokens — `<|start|>`, `<|message|>`, `<|end|>`, `<|return|>`, and so on — plus 1081 reserved slots; see `PROVENANCE.md` for the exact table). It's also the one packaged encoding not checked against `tiktoken_ruby`: that gem's 0.0.17 harmony table drops `<|endofprompt|>` where `openai/tiktoken` 0.9.0 keeps it at id 200018, so the oracle is the outlier here — `spec/gigatoken/differential_spec.rb` proves harmony instead by reduction to `o200k_base` plus a pinned special-token table.
|
|
67
67
|
|
|
68
68
|
```ruby
|
|
69
69
|
Gigatoken::Tokenizer.from_encoding("cl100k_base")
|
|
@@ -71,6 +71,8 @@ Gigatoken::Tokenizer.load("cl100k_base") # same result — packaged nam
|
|
|
71
71
|
# checked before the Hub-repo-id shape
|
|
72
72
|
```
|
|
73
73
|
|
|
74
|
+
`qwen35`, `qwen38` and `muse_spark` are whole HuggingFace `tokenizer.json` files, gzipped (`*.json.gz`, pinned to a repo revision in `PROVENANCE.md`) and loaded through `from_json`, so their special tokens come from the file. `qwen35` is the Qwen 3.5 and 3.6 tokenizer; `qwen38` is that file plus seven audio/TTS special tokens (`<|audio_start|>`, `<|audio_end|>`, `<tts_pad>`, `<tts_text_bos>`, `<tts_text_eod>`, `<tts_text_bos_single>`, `<|audio_pad|>`), so ordinary text encodes the same under both; `muse_spark` is `meta-models/Muse-Glimmer-30B`'s tokenizer, because Meta publishes no standalone Muse Spark one. Their output matches HuggingFace `tokenizers` with `add_special_tokens: false`: gigatoken applies no post-processor, so `encode` never prepends `<|begin_of_text|>` for `muse_spark`. Unlike the tiktoken entries, a registry entry for one of these is just `{json_file:}`, and each load reads, gunzips and parses its file (about 0.4 s).
|
|
75
|
+
|
|
74
76
|
`p50k_base` and `p50k_edit` are deliberately not packaged: both load the same non-dense ranks (id 50256 is left free for `<|endoftext|>`), and the rank loader rejects non-dense ranks. Both entry points raise `Gigatoken::ModelError` explaining that, rather than `load` falling through to the Hub for a name that happens to look like a legacy repo id.
|
|
75
77
|
|
|
76
78
|
`encode` on a packaged tokenizer honours its special-token table: text containing `<|endoftext|>` (or any other literal special-token string) is tokenized as that special token, not as ordinary text. That matches [`tiktoken`](https://github.com/openai/tiktoken)'s `encode_with_special_tokens`, not its plain `encode`, which treats the same literal as ordinary text — a difference worth knowing if you're tokenizing untrusted input. To get tiktoken's non-honouring default instead, build a tokenizer from the same rank file with an empty special-token table:
|
data/ext/gigatoken/Cargo.toml
CHANGED
|
@@ -1,4 +1,4 @@
|
|
|
1
|
-
# Vendored
|
|
1
|
+
# Vendored encodings — provenance
|
|
2
2
|
|
|
3
3
|
The `.tiktoken` files in this directory are OpenAI's published BPE mergeable-rank
|
|
4
4
|
tables, vendored verbatim so `gigatoken` can resolve these encodings by name with
|
|
@@ -9,6 +9,10 @@ carries **mergeable ranks only** — the pretokenizer split regex and the specia
|
|
|
9
9
|
tokens belong to the encoding's *definition*, not to the file, and live in code
|
|
10
10
|
(see the table below).
|
|
11
11
|
|
|
12
|
+
The three `.json.gz` files are HuggingFace `tokenizer.json` files, not tiktoken
|
|
13
|
+
tables; they are covered in [their own section](#huggingface-tokenizerjson-files)
|
|
14
|
+
at the end.
|
|
15
|
+
|
|
12
16
|
They live under `lib/` rather than a top-level `data/` directory because this
|
|
13
17
|
repo's `.gitignore` ignores `/data/` ("downloaded test data"); these are shipped
|
|
14
18
|
gem payload, not test fixtures.
|
|
@@ -108,3 +112,111 @@ The `tiktoken` project and its published encoding files are MIT licensed,
|
|
|
108
112
|
Copyright (c) 2022 OpenAI, Shantanu Jain. See
|
|
109
113
|
<https://github.com/openai/tiktoken/blob/main/LICENSE>. The files are vendored
|
|
110
114
|
here unmodified.
|
|
115
|
+
|
|
116
|
+
## HuggingFace `tokenizer.json` files
|
|
117
|
+
|
|
118
|
+
`qwen35.json.gz`, `qwen38.json.gz` and `muse_spark.json.gz` are HuggingFace
|
|
119
|
+
`tokenizer.json` files, vendored byte-for-byte and gzipped. Unlike a
|
|
120
|
+
`.tiktoken` table, a `tokenizer.json` carries the encoding's whole definition
|
|
121
|
+
— vocabulary, merges, normalizer, pretokenizer, added tokens — so the registry
|
|
122
|
+
records only the file, and `Tokenizer.from_encoding` loads it with
|
|
123
|
+
`Tokenizer.from_json`, reading the special tokens from the file's
|
|
124
|
+
`added_tokens` (those marked `"special": true`) exactly as `from_json` does.
|
|
125
|
+
|
|
126
|
+
### Files
|
|
127
|
+
|
|
128
|
+
| File | Decompressed bytes | sha256 (decompressed) |
|
|
129
|
+
|------|-------------------:|-----------------------|
|
|
130
|
+
| `qwen35.json.gz` | 12,807,982 | `5f9e4d4901a92b997e463c1f46055088b6cca5ca61a6522d1b9f64c4bb81cb42` |
|
|
131
|
+
| `qwen38.json.gz` | 12,809,320 | `0997f410c57a1f4e53b09e4be8f4a172d90edd9564368fb0847030937229b9f3` |
|
|
132
|
+
| `muse_spark.json.gz` | 28,129,897 | `c9dbee66967b58f31a7c27f723c3760da3526ccd0427578e8905b0abb0031c4d` |
|
|
133
|
+
|
|
134
|
+
### Source
|
|
135
|
+
|
|
136
|
+
Retrieved **2026-10-07** over HTTPS from the Hub's `resolve` endpoint, each at
|
|
137
|
+
a pinned commit:
|
|
138
|
+
|
|
139
|
+
| Encoding | Repo | Revision | URL |
|
|
140
|
+
|---|---|---|---|
|
|
141
|
+
| `qwen35` | `Qwen/Qwen3.5-27B` | `fc05daec18b0a78c049392ed2e771dde82bdf654` | <https://huggingface.co/Qwen/Qwen3.5-27B/resolve/fc05daec18b0a78c049392ed2e771dde82bdf654/tokenizer.json> |
|
|
142
|
+
| `qwen38` | `Qwen/Qwen3.8-27B` | `1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0` | <https://huggingface.co/Qwen/Qwen3.8-27B/resolve/1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0/tokenizer.json> |
|
|
143
|
+
| `muse_spark` | `meta-models/Muse-Glimmer-30B` | `a4e59da52a7bc87ae7251dd5545c0dd437c44b68` | <https://huggingface.co/meta-models/Muse-Glimmer-30B/resolve/a4e59da52a7bc87ae7251dd5545c0dd437c44b68/tokenizer.json> |
|
|
144
|
+
|
|
145
|
+
**`qwen35` is shared across releases.** The Hub's tree API
|
|
146
|
+
(`https://huggingface.co/api/models/<repo>?blobs=true`) reports the identical
|
|
147
|
+
LFS sha256 `5f9e4d49…` for `tokenizer.json` in `Qwen/Qwen3.5-{0.8B, 2B, 4B, 9B,
|
|
148
|
+
27B, 35B-A3B, 122B-A10B, 397B-A17B}` and `Qwen/Qwen3.6-{27B, 35B-A3B}`
|
|
149
|
+
(queried 2026-10-07), so one file serves Qwen 3.5 and 3.6.
|
|
150
|
+
`Qwen/Qwen3.6-Plus` is gated and was not checked.
|
|
151
|
+
|
|
152
|
+
**`qwen38` ships its own file**, identical (`0997f410…`) across
|
|
153
|
+
`Qwen/Qwen3.8-{27B, Flash-Next, 2.4T-A95B}` and their `-FP8` twins. It differs
|
|
154
|
+
from `qwen35` only in seven more `"special": true` added tokens, ids
|
|
155
|
+
248070–248076: `<|audio_start|>`, `<|audio_end|>`, `<tts_pad>`,
|
|
156
|
+
`<tts_text_bos>`, `<tts_text_eod>`, `<tts_text_bos_single>`, `<|audio_pad|>`.
|
|
157
|
+
The vocabulary, merges, normalizer, pretokenizer, decoder and post-processor
|
|
158
|
+
are identical, so ordinary text encodes the same under both; only those seven
|
|
159
|
+
literals, `vocab_size` and `special_tokens` differ. It is vendored as a second
|
|
160
|
+
file rather than derived from the first because the engine reads added tokens
|
|
161
|
+
from the JSON, and a JSON-backed tokenizer has no way to add specials after
|
|
162
|
+
the fact. `Qwen3.8-27B` is the repo cited; `Flash-Next` and `2.4T-A95B` carry an
|
|
163
|
+
"other" licence on the *model*, but their tokenizer file is byte-identical to
|
|
164
|
+
the Apache-2.0 repo's.
|
|
165
|
+
|
|
166
|
+
**`muse_spark` is the Muse-Glimmer-30B file.** Meta publishes no standalone
|
|
167
|
+
Muse Spark tokenizer (a Hub search over `meta-models` finds none, and Meta's
|
|
168
|
+
model-API docs at <https://dev.meta.ai/docs/muse-glimmer> point at none). The
|
|
169
|
+
Muse-Glimmer-30B model card states "Tokenizer: 200,000 BPE tokens + 2,048
|
|
170
|
+
special tokens" (vocabulary 202,048), trained from Muse Spark's outputs, and
|
|
171
|
+
this exact repo and revision is the one the reference application pins. As of
|
|
172
|
+
2026-10-07 that revision is also `main`.
|
|
173
|
+
|
|
174
|
+
### Authenticity
|
|
175
|
+
|
|
176
|
+
Verified three independent ways at retrieval time:
|
|
177
|
+
|
|
178
|
+
1. **Transport** — HTTPS directly from `huggingface.co`, at the pinned
|
|
179
|
+
commit above.
|
|
180
|
+
2. **Publisher checksum** — each file's sha256 above matches the LFS sha256 the
|
|
181
|
+
Hub's tree API reports for `tokenizer.json` at that revision, and so does
|
|
182
|
+
its size. Two independent channels agree on the bytes.
|
|
183
|
+
3. **Behavioral, against the reference implementation** — loaded through
|
|
184
|
+
`Tokenizer.from_encoding`, every token id matched HuggingFace
|
|
185
|
+
[`tokenizers`](https://rubygems.org/gems/tokenizers) 0.7.0 loading the same
|
|
186
|
+
decompressed file with `add_special_tokens: false`, across this repo's own
|
|
187
|
+
source and docs (`spec/gigatoken/differential_spec.rb`). Resulting
|
|
188
|
+
`vocab_size`: 248070 / 248077 / 202048; `special_tokens.size`: 14 / 21 / 2048.
|
|
189
|
+
|
|
190
|
+
### Compression
|
|
191
|
+
|
|
192
|
+
Each file was compressed with `gzip -9 -n` (no name or timestamp in the gzip
|
|
193
|
+
header, so the bytes are reproducible), and `gunzip -c` of it reproduces the
|
|
194
|
+
published file byte-for-byte — the sha256s above are of the decompressed
|
|
195
|
+
bytes. Compressed sizes: 3,514,616 (`qwen35`), 3,514,699 (`qwen38`), 4,475,666
|
|
196
|
+
(`muse_spark`).
|
|
197
|
+
|
|
198
|
+
### What gigatoken does not take from the file
|
|
199
|
+
|
|
200
|
+
- **The post-processor.** gigatoken applies none. `muse_spark`'s is a
|
|
201
|
+
`TemplateProcessing` that prepends `<|begin_of_text|>`; `Tokenizer#encode`
|
|
202
|
+
does not, matching `tokenizers`' `add_special_tokens: false`. The Qwen files'
|
|
203
|
+
is a `ByteLevel` post-processor, a no-op on ids.
|
|
204
|
+
- **Non-special added tokens in `special_tokens`.** Qwen's 12 non-special
|
|
205
|
+
added tokens (`<tool_call>`, `<think>`, the `<|fim_*|>` tokens,
|
|
206
|
+
`<|repo_name|>`, `<|file_sep|>`, `<tool_response>` and so on) are matched
|
|
207
|
+
atomically by the engine, like the special ones, but only `"special": true`
|
|
208
|
+
tokens appear in `Tokenizer#special_tokens` — 14 for `qwen35`, 21 for
|
|
209
|
+
`qwen38`.
|
|
210
|
+
- **`tokenizer_config.json`.** Not vendored; gigatoken reads only
|
|
211
|
+
`tokenizer.json`.
|
|
212
|
+
|
|
213
|
+
### Licence
|
|
214
|
+
|
|
215
|
+
All three repos are Apache-2.0, each with its own `LICENSE` at the pinned
|
|
216
|
+
revision:
|
|
217
|
+
|
|
218
|
+
- <https://huggingface.co/Qwen/Qwen3.5-27B/blob/fc05daec18b0a78c049392ed2e771dde82bdf654/LICENSE>
|
|
219
|
+
- <https://huggingface.co/Qwen/Qwen3.8-27B/blob/1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0/LICENSE>
|
|
220
|
+
- <https://huggingface.co/meta-models/Muse-Glimmer-30B/blob/a4e59da52a7bc87ae7251dd5545c0dd437c44b68/LICENSE>
|
|
221
|
+
|
|
222
|
+
The files are vendored here unmodified (gzipped).
|
|
Binary file
|
|
Binary file
|
|
Binary file
|
data/lib/gigatoken/encodings.rb
CHANGED
|
@@ -1,10 +1,11 @@
|
|
|
1
1
|
# frozen_string_literal: true
|
|
2
2
|
|
|
3
3
|
module Gigatoken
|
|
4
|
-
# The
|
|
5
|
-
# .tiktoken file doesn't carry: its pretokenizer scheme and
|
|
6
|
-
# table (see Tokenizer.from_tiktoken and lib/gigatoken/
|
|
7
|
-
# PROVENANCE.md, the source of truth this is transcribed from).
|
|
4
|
+
# The encodings gigatoken vendors a file for. For the tiktoken ones, the
|
|
5
|
+
# pieces a .tiktoken file doesn't carry: its pretokenizer scheme and
|
|
6
|
+
# special-token table (see Tokenizer.from_tiktoken and lib/gigatoken/
|
|
7
|
+
# encodings/PROVENANCE.md, the source of truth this is transcribed from).
|
|
8
|
+
# The HuggingFace ones are whole tokenizer.json files, which carry theirs.
|
|
8
9
|
module Encodings
|
|
9
10
|
DATA_DIR = File.expand_path("encodings", __dir__)
|
|
10
11
|
private_constant :DATA_DIR
|
|
@@ -40,10 +41,10 @@ module Gigatoken
|
|
|
40
41
|
HARMONY_RESERVED_TAIL = (200013..201087).to_h { |id| ["<|reserved_#{id}|>", id] }.freeze
|
|
41
42
|
private_constant :HARMONY_RESERVED_TAIL
|
|
42
43
|
|
|
43
|
-
# Deep-frozen: entries, their `special_tokens` tables and the `rank_file`
|
|
44
|
-
# paths. `Tokenizer#special_tokens` hands the registry's own
|
|
45
|
-
# callers, so anything less lets one caller's poke rewrite
|
|
46
|
-
# later `from_encoding` in the process loads.
|
|
44
|
+
# Deep-frozen: entries, their `special_tokens` tables and the `rank_file` /
|
|
45
|
+
# `json_file` paths. `Tokenizer#special_tokens` hands the registry's own
|
|
46
|
+
# Hash back to callers, so anything less lets one caller's poke rewrite
|
|
47
|
+
# what every later `from_encoding` in the process loads.
|
|
47
48
|
REGISTRY = {
|
|
48
49
|
"r50k_base" => {
|
|
49
50
|
rank_file: File.join(DATA_DIR, "r50k_base.tiktoken"),
|
|
@@ -70,7 +71,21 @@ module Gigatoken
|
|
|
70
71
|
rank_file: File.join(DATA_DIR, "o200k_base.tiktoken"),
|
|
71
72
|
pretokenizer: "o200k",
|
|
72
73
|
special_tokens: HARMONY_HEAD_TOKENS.merge(HARMONY_RESERVED_TAIL).freeze
|
|
73
|
-
}
|
|
74
|
+
},
|
|
75
|
+
# The three below are gzipped HuggingFace tokenizer.json files, vendored
|
|
76
|
+
# verbatim (see PROVENANCE.md for repos, revisions and hashes). Nothing
|
|
77
|
+
# else is registered: the pretokenizer and the special tokens are the
|
|
78
|
+
# file's own, read at load as Tokenizer.from_json reads them.
|
|
79
|
+
#
|
|
80
|
+
# Qwen 3.5 and 3.6 share this one file.
|
|
81
|
+
"qwen35" => {json_file: File.join(DATA_DIR, "qwen35.json.gz")},
|
|
82
|
+
# The qwen35 file plus seven audio/TTS special tokens — vendored
|
|
83
|
+
# separately because a JSON-backed tokenizer takes its added tokens
|
|
84
|
+
# from the file.
|
|
85
|
+
"qwen38" => {json_file: File.join(DATA_DIR, "qwen38.json.gz")},
|
|
86
|
+
# meta-models/Muse-Glimmer-30B's file: Meta publishes no standalone
|
|
87
|
+
# Muse Spark tokenizer.
|
|
88
|
+
"muse_spark" => {json_file: File.join(DATA_DIR, "muse_spark.json.gz")}
|
|
74
89
|
}.each_value { |encoding| encoding.each_value(&:freeze).freeze }.freeze
|
|
75
90
|
private_constant :REGISTRY
|
|
76
91
|
|
|
@@ -91,9 +106,9 @@ module Gigatoken
|
|
|
91
106
|
private_constant :UNPACKABLE_REASONS
|
|
92
107
|
|
|
93
108
|
class << self
|
|
94
|
-
# The {rank_file:, pretokenizer:, special_tokens:}
|
|
95
|
-
#
|
|
96
|
-
# Tokenizer.load accepts both.
|
|
109
|
+
# The {rank_file:, pretokenizer:, special_tokens:} (tiktoken) or
|
|
110
|
+
# {json_file:} (HuggingFace) registered for a packaged encoding name, or
|
|
111
|
+
# nil. Names are Strings or Symbols, as Tokenizer.load accepts both.
|
|
97
112
|
def [](name)
|
|
98
113
|
REGISTRY[name.to_s]
|
|
99
114
|
end
|
data/lib/gigatoken/tokenizer.rb
CHANGED
|
@@ -1,6 +1,7 @@
|
|
|
1
1
|
# frozen_string_literal: true
|
|
2
2
|
|
|
3
3
|
require "json"
|
|
4
|
+
require "zlib"
|
|
4
5
|
|
|
5
6
|
module Gigatoken
|
|
6
7
|
# A tokenizer: encode, batch encode, decode, and vocabulary introspection
|
|
@@ -52,11 +53,13 @@ module Gigatoken
|
|
|
52
53
|
new(native, special_tokens: special_tokens)
|
|
53
54
|
end
|
|
54
55
|
|
|
55
|
-
# Load one of the
|
|
56
|
-
#
|
|
57
|
-
#
|
|
56
|
+
# Load one of the encodings gigatoken vendors a file for, by name — see
|
|
57
|
+
# Gigatoken::Encodings::NAMES — entirely from the vendored files: no
|
|
58
|
+
# network, no writable cache. A tiktoken entry loads its ranks; a
|
|
59
|
+
# tokenizer.json entry is gunzipped and loaded by from_json.
|
|
58
60
|
def self.from_encoding(name)
|
|
59
61
|
encoding = Encodings[name]
|
|
62
|
+
return from_json(Zlib.gunzip(File.binread(encoding[:json_file]))) if encoding&.key?(:json_file)
|
|
60
63
|
return from_tiktoken(encoding[:rank_file], pretokenizer: encoding[:pretokenizer], special_tokens: encoding[:special_tokens]) if encoding
|
|
61
64
|
|
|
62
65
|
reason = Encodings.unpackable_reason(name)
|
data/lib/gigatoken/version.rb
CHANGED
metadata
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
--- !ruby/object:Gem::Specification
|
|
2
2
|
name: gigatoken
|
|
3
3
|
version: !ruby/object:Gem::Version
|
|
4
|
-
version: 0.
|
|
4
|
+
version: 0.5.0
|
|
5
5
|
platform: ruby
|
|
6
6
|
authors:
|
|
7
7
|
- Eric Jacobs
|
|
@@ -107,7 +107,10 @@ files:
|
|
|
107
107
|
- lib/gigatoken/encodings.rb
|
|
108
108
|
- lib/gigatoken/encodings/PROVENANCE.md
|
|
109
109
|
- lib/gigatoken/encodings/cl100k_base.tiktoken
|
|
110
|
+
- lib/gigatoken/encodings/muse_spark.json.gz
|
|
110
111
|
- lib/gigatoken/encodings/o200k_base.tiktoken
|
|
112
|
+
- lib/gigatoken/encodings/qwen35.json.gz
|
|
113
|
+
- lib/gigatoken/encodings/qwen38.json.gz
|
|
111
114
|
- lib/gigatoken/encodings/r50k_base.tiktoken
|
|
112
115
|
- lib/gigatoken/hub.rb
|
|
113
116
|
- lib/gigatoken/packed_result.rb
|