fast_mb_chars 0.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +7 -0
- data/CHANGELOG.md +9 -0
- data/LICENSE +21 -0
- data/README.md +180 -0
- data/benchmark/compare.rb +82 -0
- data/docs/limit-cut.svg +66 -0
- data/lib/fast_mb_chars/chars.rb +130 -0
- data/lib/fast_mb_chars/limiter.rb +152 -0
- data/lib/fast_mb_chars/version.rb +5 -0
- data/lib/fast_mb_chars.rb +24 -0
- metadata +66 -0
checksums.yaml
ADDED
|
@@ -0,0 +1,7 @@
|
|
|
1
|
+
---
|
|
2
|
+
SHA256:
|
|
3
|
+
metadata.gz: 35532c3219891f1ff4d3a7fe168abb9213bafa430d7c33051eb2ac58f24ab0a5
|
|
4
|
+
data.tar.gz: 64672d7dd534037d5bda31bd52e606246f527583bd5880ca33100f4b0d435bf4
|
|
5
|
+
SHA512:
|
|
6
|
+
metadata.gz: 21021182d78d1ba4d51d518546161f95a412709dfa998f2a20068c542cd874932a6ea3de8cc981b9e6d7f77f6651a9eabef480ffdc4f6add5de179354cfb6ba1
|
|
7
|
+
data.tar.gz: e8c45c5e5aa21babbbe604151228ddd39bea12882d1dee684761ab71fdd166da801e7159e7f1ffd6161aef816a15f2d2d0e1f3d002d568ab02ea9760ed690d2e
|
data/CHANGELOG.md
ADDED
|
@@ -0,0 +1,9 @@
|
|
|
1
|
+
# Changelog
|
|
2
|
+
|
|
3
|
+
## 0.1.0
|
|
4
|
+
|
|
5
|
+
`String#mb_chars` with the usual case, split, and Unicode helpers, plus `limit(bytes)`.
|
|
6
|
+
|
|
7
|
+
The gem prepends `String#mb_chars`. ASCII and UTF-8 are supported. Other encodings need `encode('UTF-8')` first.
|
|
8
|
+
|
|
9
|
+
The byte cut stays on a character boundary. A character that crosses the limit is kept when it finishes within 128 bytes. A longer combining sequence is dropped. Invalid bytes do not raise.
|
data/LICENSE
ADDED
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 SYNORYXEL
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|
data/README.md
ADDED
|
@@ -0,0 +1,180 @@
|
|
|
1
|
+
# ⚡ fast_mb_chars
|
|
2
|
+
|
|
3
|
+
Faster `String#mb_chars` for Ruby. Same calls as `ActiveSupport::Multibyte::Chars` — `limit`, `downcase`, `upcase`, `titleize`, `reverse`, `split`, `tidy_bytes`. The byte cut does not split a character and does not raise on a broken byte.
|
|
4
|
+
|
|
5
|
+
Works on **ASCII** and **UTF-8** (Київ, café, 日本語, emoji). Other encodings need a convert first.
|
|
6
|
+
|
|
7
|
+
<p align="center">
|
|
8
|
+
<img src="docs/limit-cut.svg" alt="Byte limit lands inside я; the gem keeps the whole letter" width="880">
|
|
9
|
+
</p>
|
|
10
|
+
|
|
11
|
+
`mb_chars.limit` in Rails reads every character from the start of the string. This gem looks at the cut. On a 1 MB string that is **205 ms** vs **0.08 ms**.
|
|
12
|
+
|
|
13
|
+
| | 1 MB | 7 MB |
|
|
14
|
+
| --- | ---: | ---: |
|
|
15
|
+
| ActiveSupport-style scan | 205 ms | 1,410 ms |
|
|
16
|
+
| **fast_mb_chars** | **0.08 ms** | **0.47 ms** |
|
|
17
|
+
| | **~2,500×** | **~3,000×** |
|
|
18
|
+
|
|
19
|
+
ASCII, Ruby 4.0.5, one call. The string is longer than the limit, so both sides actually cut. Full tables are [below](#-benchmarks).
|
|
20
|
+
|
|
21
|
+
[](https://github.com/SYNORYXEL/fast_mb_chars/actions)
|
|
22
|
+
[](https://www.ruby-lang.org/)
|
|
23
|
+
[](LICENSE)
|
|
24
|
+
[](#-encodings)
|
|
25
|
+
|
|
26
|
+
Source: [github.com/SYNORYXEL/fast_mb_chars](https://github.com/SYNORYXEL/fast_mb_chars)
|
|
27
|
+
|
|
28
|
+
## 🌍 Encodings
|
|
29
|
+
|
|
30
|
+
The gem treats the string as **UTF-8 bytes**. ASCII is UTF-8. It is not a transcoding library.
|
|
31
|
+
|
|
32
|
+
| Input | What happens |
|
|
33
|
+
| --- | --- |
|
|
34
|
+
| ASCII `hello` | Fast path. Both bytes on the boundary are `< 0x80`, so the cut is exact. |
|
|
35
|
+
| UTF-8 `Київ`, `café`, `日本語`, `👨👩👧👦` | Yes. A letter or emoji that crosses the limit stays whole. |
|
|
36
|
+
| Broken / binary bytes | No `ArgumentError`. Invalid sequences are skipped or scrubbed. |
|
|
37
|
+
| Windows-1252 mixed into UTF-8 (`caf\xE9`) | `tidy_bytes` repairs it to `café`. |
|
|
38
|
+
| Windows-1251, KOI8-R, ISO-8859-1, Shift_JIS, UTF-16 | **No.** Bytes are relabeled as UTF-8, not converted. Encode first. |
|
|
39
|
+
|
|
40
|
+
```ruby
|
|
41
|
+
# already UTF-8 or ASCII — just call it
|
|
42
|
+
'Київ'.mb_chars.downcase.to_s
|
|
43
|
+
'hello'.mb_chars.limit(4).to_s
|
|
44
|
+
|
|
45
|
+
# other encodings: convert, then cut
|
|
46
|
+
cp1251.encode('UTF-8').mb_chars.limit(1_048_576).to_s
|
|
47
|
+
```
|
|
48
|
+
|
|
49
|
+
`force_encoding('UTF-8')` only changes the label. `encode('UTF-8')` changes the bytes. Use the second one when the source is not UTF-8.
|
|
50
|
+
|
|
51
|
+
## 📦 Install
|
|
52
|
+
|
|
53
|
+
Ruby 2.7 or newer. Tests run on 4.0.5. CI runs 2.7, 3.2, 3.3, 3.4, and 4.0.
|
|
54
|
+
|
|
55
|
+
```ruby
|
|
56
|
+
gem 'fast_mb_chars', github: 'SYNORYXEL/fast_mb_chars'
|
|
57
|
+
```
|
|
58
|
+
|
|
59
|
+
```ruby
|
|
60
|
+
require 'fast_mb_chars'
|
|
61
|
+
```
|
|
62
|
+
|
|
63
|
+
The gem prepends `String#mb_chars`, so existing `'text'.mb_chars.downcase` calls in a Rails app keep working. It does not patch `String#downcase`. If another library prepends `mb_chars` later, call `FastMbChars.install!` again. No runtime dependencies. Not on RubyGems yet — install from GitHub.
|
|
64
|
+
|
|
65
|
+
From a checkout: `gem build fast_mb_chars.gemspec` then `gem install ./fast_mb_chars-0.1.0.gem`. Next to an app, `gem 'fast_mb_chars', path: '../fast_mb_chars'` works too.
|
|
66
|
+
|
|
67
|
+
## ✨ Usage
|
|
68
|
+
|
|
69
|
+
```ruby
|
|
70
|
+
'ГОТІВКА'.mb_chars.downcase.to_s
|
|
71
|
+
# => "готівка"
|
|
72
|
+
|
|
73
|
+
' Київ '.mb_chars.downcase.strip.to_s
|
|
74
|
+
# => "київ"
|
|
75
|
+
|
|
76
|
+
'ÉL QUE SE ENTERÓ'.mb_chars.titleize.to_s
|
|
77
|
+
# => "Él Que Se Enteró"
|
|
78
|
+
|
|
79
|
+
'Café périferôl'.mb_chars.split(/é/).map { |part| part.upcase.to_s }
|
|
80
|
+
# => ["CAF", " P", "RIFERÔL"]
|
|
81
|
+
|
|
82
|
+
data.mb_chars.limit(1_048_576).to_s
|
|
83
|
+
data.mb_chars.limit(7_340_032).to_s
|
|
84
|
+
```
|
|
85
|
+
|
|
86
|
+
Methods on the proxy itself: `limit`, `reverse`, `reverse!`, `titleize`, `titlecase`, `compose`, `decompose`, `grapheme_length`, `tidy_bytes`, `tidy_bytes!`, `split`, `slice!`, `<=>`, `=~`, `match?`, `acts_like_string?`, `as_json`, `to_s`, `to_str`.
|
|
87
|
+
|
|
88
|
+
`downcase`, `upcase`, `capitalize`, `swapcase`, `strip`, `gsub`, `length`, and the rest of `String` go through `method_missing`. A string result comes back as `mb_chars`, so the chain does not break. A bang method that changes the string returns the same proxy; if Ruby returns `nil` (no change), this does too.
|
|
89
|
+
|
|
90
|
+
`titleize` is a Unicode-aware regex, not ActiveSupport inflections.
|
|
91
|
+
|
|
92
|
+
## 🚀 Why it is faster
|
|
93
|
+
|
|
94
|
+
Rails does this:
|
|
95
|
+
|
|
96
|
+
```ruby
|
|
97
|
+
def limit(limit)
|
|
98
|
+
chars(@wrapped_string.truncate_bytes(limit, omission: nil))
|
|
99
|
+
end
|
|
100
|
+
```
|
|
101
|
+
|
|
102
|
+
`truncate_bytes` walks graphemes from byte 0 until it hits the limit. That is fine for a button label. It is not fine for a 1 MB or 7 MB log line.
|
|
103
|
+
|
|
104
|
+
`fast_mb_chars` reads the two bytes on the boundary.
|
|
105
|
+
|
|
106
|
+
- Both ASCII, and not a CR+LF pair: one `byteslice`. Exact.
|
|
107
|
+
- A character crosses the limit and ends within 128 bytes: keep it whole. The result can be a few bytes over the limit.
|
|
108
|
+
- A long run of combining marks that does not fit in those 128 bytes: drop it. The result stays within the limit.
|
|
109
|
+
- Empty string, or a limit of 0 or less: `""`.
|
|
110
|
+
- Shorter than the limit: returned as-is, trailing spaces and punctuation included.
|
|
111
|
+
- A broken byte never reaches a regexp, so there is no `ArgumentError`.
|
|
112
|
+
|
|
113
|
+
Rails never goes past the limit, so it drops the character on the boundary. This gem keeps a short character. That is the one behavioral difference, and it is why a cut like `('a' * (limit - 1)) + 'я'` comes back one byte longer, with `я` intact.
|
|
114
|
+
|
|
115
|
+
```mermaid
|
|
116
|
+
flowchart LR
|
|
117
|
+
A[1 MB string] --> B{bytes at the limit}
|
|
118
|
+
B -->|both ASCII| C[byteslice exact]
|
|
119
|
+
B -->|UTF-8 letter or emoji| D[keep the cluster if it ends within 128 bytes]
|
|
120
|
+
B -->|broken byte| E[safe prefix no raise]
|
|
121
|
+
```
|
|
122
|
+
|
|
123
|
+
## 📊 Benchmarks
|
|
124
|
+
|
|
125
|
+
"scan" walks graphemes from the start, which is what `mb_chars.limit` does. "fast" is `FastMbChars::Limiter.limit`. Milliseconds per call, Ruby 4.0.5.
|
|
126
|
+
|
|
127
|
+
### ASCII
|
|
128
|
+
|
|
129
|
+
The cut lands on a byte. Both results are exactly the limit.
|
|
130
|
+
|
|
131
|
+
| size | scan | fast | times |
|
|
132
|
+
| --- | ---: | ---: | ---: |
|
|
133
|
+
| 1 KB | 0.198 | 0.000 | 499× |
|
|
134
|
+
| 8 KB | 1.566 | 0.001 | 1,284× |
|
|
135
|
+
| 64 KB | 12.671 | 0.010 | 1,233× |
|
|
136
|
+
| 256 KB | 50.936 | 0.028 | 1,841× |
|
|
137
|
+
| 1 MB | 205.382 | 0.083 | 2,467× |
|
|
138
|
+
| 7 MB | 1,409.907 | 0.466 | 3,028× |
|
|
139
|
+
|
|
140
|
+
### Cyrillic
|
|
141
|
+
|
|
142
|
+
The cut lands between letters. Both results are exactly the limit.
|
|
143
|
+
|
|
144
|
+
| size | scan | fast | times |
|
|
145
|
+
| --- | ---: | ---: | ---: |
|
|
146
|
+
| 1 KB | 0.145 | 0.020 | 7× |
|
|
147
|
+
| 8 KB | 1.145 | 0.022 | 52× |
|
|
148
|
+
| 64 KB | 9.282 | 0.023 | 407× |
|
|
149
|
+
| 256 KB | 36.863 | 0.040 | 931× |
|
|
150
|
+
| 1 MB | 147.932 | 0.068 | 2,164× |
|
|
151
|
+
| 7 MB | 1,032.553 | 0.194 | 5,322× |
|
|
152
|
+
|
|
153
|
+
### Character on the boundary
|
|
154
|
+
|
|
155
|
+
`я` crosses the limit. The scan stops on the limit and drops the letter. The fast cut keeps it, so the result is one byte over.
|
|
156
|
+
|
|
157
|
+
| size | scan | fast | times | fast bytes / limit |
|
|
158
|
+
| --- | ---: | ---: | ---: | ---: |
|
|
159
|
+
| 1 KB | 0.196 | 0.027 | 7× | 1025 / 1024 |
|
|
160
|
+
| 8 KB | 1.545 | 0.029 | 54× | 8193 / 8192 |
|
|
161
|
+
| 64 KB | 12.377 | 0.028 | 443× | 65537 / 65536 |
|
|
162
|
+
| 256 KB | 49.862 | 0.026 | 1,893× | 262145 / 262144 |
|
|
163
|
+
| 1 MB | 199.377 | 0.036 | 5,597× | 1048577 / 1048576 |
|
|
164
|
+
| 7 MB | 1,479.463 | 0.030 | 49,870× | 7340033 / 7340032 |
|
|
165
|
+
|
|
166
|
+
On a short string you pay for the wrapper, not for the cut. `'ГОТІВКА'.mb_chars.downcase.to_s` was 0.0006 ms. `String#downcase` was 0.0002 ms.
|
|
167
|
+
|
|
168
|
+
```bash
|
|
169
|
+
ruby benchmark/compare.rb
|
|
170
|
+
```
|
|
171
|
+
|
|
172
|
+
## ✅ Tests
|
|
173
|
+
|
|
174
|
+
```bash
|
|
175
|
+
bundle exec rake
|
|
176
|
+
```
|
|
177
|
+
|
|
178
|
+
## 📄 License
|
|
179
|
+
|
|
180
|
+
MIT. See [LICENSE](LICENSE).
|
|
@@ -0,0 +1,82 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
|
|
3
|
+
# ruby benchmark/compare.rb
|
|
4
|
+
# rake benchmark
|
|
5
|
+
|
|
6
|
+
$LOAD_PATH.unshift(File.expand_path('../lib', __dir__))
|
|
7
|
+
|
|
8
|
+
require 'fast_mb_chars'
|
|
9
|
+
|
|
10
|
+
# What ActiveSupport::Multibyte::Chars#limit does: walk every grapheme from
|
|
11
|
+
# the start and never go past the byte budget.
|
|
12
|
+
def walk_limit(text, max_bytes)
|
|
13
|
+
text = text.to_s
|
|
14
|
+
return +'' if max_bytes <= 0 || text.empty?
|
|
15
|
+
return text.dup if text.bytesize <= max_bytes
|
|
16
|
+
|
|
17
|
+
taken = 0
|
|
18
|
+
text.each_grapheme_cluster do |cluster|
|
|
19
|
+
size = cluster.bytesize
|
|
20
|
+
break if taken + size > max_bytes
|
|
21
|
+
|
|
22
|
+
taken += size
|
|
23
|
+
end
|
|
24
|
+
text.byteslice(0, taken)
|
|
25
|
+
end
|
|
26
|
+
|
|
27
|
+
SIZES = [
|
|
28
|
+
['1 KB', 1_024, 300],
|
|
29
|
+
['8 KB', 8_192, 150],
|
|
30
|
+
['64 KB', 65_536, 40],
|
|
31
|
+
['256 KB', 262_144, 15],
|
|
32
|
+
['1 MB', 1_048_576, 8],
|
|
33
|
+
['7 MB', 7_340_032, 3]
|
|
34
|
+
].freeze
|
|
35
|
+
|
|
36
|
+
# The string is longer than the limit, so both sides actually cut.
|
|
37
|
+
# split puts one 2-byte letter across the boundary.
|
|
38
|
+
def payload(kind, bytes)
|
|
39
|
+
case kind
|
|
40
|
+
when :ascii
|
|
41
|
+
['a' * (bytes + 64), bytes]
|
|
42
|
+
when :cyrillic
|
|
43
|
+
['я' * ((bytes / 2) + 16), bytes]
|
|
44
|
+
when :split
|
|
45
|
+
[(('a' * (bytes - 1)) + 'я'), bytes]
|
|
46
|
+
end
|
|
47
|
+
end
|
|
48
|
+
|
|
49
|
+
def time_ms(iterations)
|
|
50
|
+
started = Process.clock_gettime(Process::CLOCK_MONOTONIC)
|
|
51
|
+
iterations.times { yield }
|
|
52
|
+
((Process.clock_gettime(Process::CLOCK_MONOTONIC) - started) / iterations) * 1000.0
|
|
53
|
+
end
|
|
54
|
+
|
|
55
|
+
rows = []
|
|
56
|
+
|
|
57
|
+
puts "ruby #{RUBY_VERSION} fast_mb_chars #{FastMbChars::VERSION}"
|
|
58
|
+
puts "limit vs full grapheme walk (ms per call, one process)"
|
|
59
|
+
puts
|
|
60
|
+
|
|
61
|
+
%i[ascii cyrillic split].each do |kind|
|
|
62
|
+
puts kind
|
|
63
|
+
puts format('%-8s %10s %10s %8s %s', 'size', 'walk', 'fast', 'times', 'result')
|
|
64
|
+
SIZES.each do |label, bytes, iterations|
|
|
65
|
+
text, limit = payload(kind, bytes)
|
|
66
|
+
2.times { walk_limit(text, limit) }
|
|
67
|
+
2.times { FastMbChars::Limiter.limit(text, limit) }
|
|
68
|
+
slow = time_ms(iterations) { walk_limit(text, limit) }
|
|
69
|
+
fast = time_ms(iterations) { FastMbChars::Limiter.limit(text, limit) }
|
|
70
|
+
ratio = fast.positive? ? slow / fast : 0
|
|
71
|
+
cut = FastMbChars::Limiter.limit(text, limit)
|
|
72
|
+
rows << [kind, label, slow, fast, ratio, cut.bytesize, limit]
|
|
73
|
+
puts format('%-8s %10.3f %10.3f %7.0fx %d/%d bytes', label, slow, fast, ratio, cut.bytesize, limit)
|
|
74
|
+
end
|
|
75
|
+
puts
|
|
76
|
+
end
|
|
77
|
+
|
|
78
|
+
small = 'ГОТІВКА'
|
|
79
|
+
time_ms(1000) { small.mb_chars.downcase.to_s }
|
|
80
|
+
down = time_ms(20_000) { small.mb_chars.downcase.to_s }
|
|
81
|
+
plain = time_ms(20_000) { small.downcase }
|
|
82
|
+
puts format('downcase %s mb_chars %.4f ms String#downcase %.4f ms', small, down, plain)
|
data/docs/limit-cut.svg
ADDED
|
@@ -0,0 +1,66 @@
|
|
|
1
|
+
<svg xmlns="http://www.w3.org/2000/svg" width="880" height="300" viewBox="0 0 880 300" role="img" aria-labelledby="title desc">
|
|
2
|
+
<title id="title">fast_mb_chars cuts on a character boundary</title>
|
|
3
|
+
<desc id="desc">The byte limit lands inside a UTF-8 letter. The gem keeps the whole character instead of splitting it.</desc>
|
|
4
|
+
<defs>
|
|
5
|
+
<style>
|
|
6
|
+
.bg { fill: #f6f8fa; }
|
|
7
|
+
.card { fill: #ffffff; stroke: #d0d7de; }
|
|
8
|
+
.ink { fill: #1f2328; }
|
|
9
|
+
.muted { fill: #656d76; }
|
|
10
|
+
.line { stroke: #d0d7de; }
|
|
11
|
+
.keep { fill: #1a7f37; }
|
|
12
|
+
.drop { fill: #cf222e; }
|
|
13
|
+
.accent { fill: #cf222e; }
|
|
14
|
+
.chip { fill: #fff8f8; stroke: #ffcecb; }
|
|
15
|
+
.cell { fill: #ffffff; stroke: #d0d7de; }
|
|
16
|
+
.cell-keep { fill: #dafbe1; stroke: #4ac26b; }
|
|
17
|
+
.cell-cut { fill: #fff8c5; stroke: #d4a72c; }
|
|
18
|
+
@media (prefers-color-scheme: dark) {
|
|
19
|
+
.bg { fill: #0d1117; }
|
|
20
|
+
.card { fill: #161b22; stroke: #30363d; }
|
|
21
|
+
.ink { fill: #e6edf3; }
|
|
22
|
+
.muted { fill: #8b949e; }
|
|
23
|
+
.line { stroke: #30363d; }
|
|
24
|
+
.keep { fill: #3fb950; }
|
|
25
|
+
.drop { fill: #f85149; }
|
|
26
|
+
.accent { fill: #f85149; }
|
|
27
|
+
.chip { fill: #3d1619; stroke: #6e2b2f; }
|
|
28
|
+
.cell { fill: #0d1117; stroke: #30363d; }
|
|
29
|
+
.cell-keep { fill: #12261e; stroke: #238636; }
|
|
30
|
+
.cell-cut { fill: #2a2111; stroke: #9e6a03; }
|
|
31
|
+
}
|
|
32
|
+
</style>
|
|
33
|
+
</defs>
|
|
34
|
+
<rect class="bg" width="880" height="300" rx="16"/>
|
|
35
|
+
<rect class="card" x="16" y="16" width="848" height="268" rx="12"/>
|
|
36
|
+
|
|
37
|
+
<text class="ink" x="40" y="52" font-family="ui-sans-serif, system-ui, sans-serif" font-size="18" font-weight="700">limit looks at the cut, not the whole string</text>
|
|
38
|
+
<text class="muted" x="40" y="76" font-family="ui-sans-serif, system-ui, sans-serif" font-size="13">ASCII and UTF-8 — characters stay whole</text>
|
|
39
|
+
|
|
40
|
+
<g font-family="ui-monospace, SFMono-Regular, Menlo, Consolas, monospace" font-size="14">
|
|
41
|
+
<rect class="cell" x="40" y="104" width="44" height="44" rx="8"/>
|
|
42
|
+
<text class="ink" x="62" y="131" text-anchor="middle">a</text>
|
|
43
|
+
<rect class="cell" x="90" y="104" width="44" height="44" rx="8"/>
|
|
44
|
+
<text class="ink" x="112" y="131" text-anchor="middle">a</text>
|
|
45
|
+
<rect class="cell" x="140" y="104" width="44" height="44" rx="8"/>
|
|
46
|
+
<text class="ink" x="162" y="131" text-anchor="middle">a</text>
|
|
47
|
+
<rect class="cell-cut" x="190" y="104" width="44" height="44" rx="8"/>
|
|
48
|
+
<text class="ink" x="212" y="131" text-anchor="middle">a</text>
|
|
49
|
+
<rect class="cell-keep" x="240" y="104" width="92" height="44" rx="8"/>
|
|
50
|
+
<text class="keep" x="286" y="131" text-anchor="middle" font-weight="700">я</text>
|
|
51
|
+
<rect class="cell" x="338" y="104" width="44" height="44" rx="8"/>
|
|
52
|
+
<text class="muted" x="360" y="131" text-anchor="middle">…</text>
|
|
53
|
+
</g>
|
|
54
|
+
|
|
55
|
+
<path class="line" d="M234 156 v18" fill="none" stroke-width="2"/>
|
|
56
|
+
<text class="accent" x="234" y="196" text-anchor="middle" font-family="ui-sans-serif, system-ui, sans-serif" font-size="12" font-weight="700">byte limit</text>
|
|
57
|
+
<text class="muted" x="234" y="214" text-anchor="middle" font-family="ui-sans-serif, system-ui, sans-serif" font-size="11">inside the letter</text>
|
|
58
|
+
|
|
59
|
+
<rect class="chip" x="430" y="104" width="400" height="108" rx="10"/>
|
|
60
|
+
<text class="drop" x="450" y="132" font-family="ui-sans-serif, system-ui, sans-serif" font-size="13" font-weight="700">Rails scan</text>
|
|
61
|
+
<text class="muted" x="450" y="152" font-family="ui-sans-serif, system-ui, sans-serif" font-size="12">walks graphemes from byte 0 → drops я</text>
|
|
62
|
+
<text class="keep" x="450" y="180" font-family="ui-sans-serif, system-ui, sans-serif" font-size="13" font-weight="700">fast_mb_chars</text>
|
|
63
|
+
<text class="muted" x="450" y="200" font-family="ui-sans-serif, system-ui, sans-serif" font-size="12">reads the two bytes on the boundary → keeps я</text>
|
|
64
|
+
|
|
65
|
+
<text class="muted" x="40" y="256" font-family="ui-monospace, SFMono-Regular, Menlo, Consolas, monospace" font-size="12">('a' * (limit - 1)) + 'я' → 1 extra byte, letter intact</text>
|
|
66
|
+
</svg>
|
|
@@ -0,0 +1,130 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
|
|
3
|
+
module FastMbChars
|
|
4
|
+
# Proxy with the same methods as ActiveSupport::Multibyte::Chars.
|
|
5
|
+
# limit is the fast path. Everything else follows the string.
|
|
6
|
+
class Chars
|
|
7
|
+
include Comparable
|
|
8
|
+
|
|
9
|
+
attr_reader :wrapped_string
|
|
10
|
+
alias to_s wrapped_string
|
|
11
|
+
alias to_str wrapped_string
|
|
12
|
+
|
|
13
|
+
def initialize(text)
|
|
14
|
+
@wrapped_string = Limiter.utf8(text)
|
|
15
|
+
@wrapped_string = @wrapped_string.dup if @wrapped_string.frozen?
|
|
16
|
+
end
|
|
17
|
+
|
|
18
|
+
def limit(max_bytes)
|
|
19
|
+
chars(Limiter.limit(@wrapped_string, Integer(max_bytes)))
|
|
20
|
+
end
|
|
21
|
+
|
|
22
|
+
def reverse
|
|
23
|
+
chars(readable.grapheme_clusters.reverse.join)
|
|
24
|
+
end
|
|
25
|
+
|
|
26
|
+
def titleize
|
|
27
|
+
chars(readable.downcase.gsub(/\b('?\S)/u) { Regexp.last_match(1).upcase })
|
|
28
|
+
end
|
|
29
|
+
alias titlecase titleize
|
|
30
|
+
|
|
31
|
+
def decompose
|
|
32
|
+
chars(readable.unicode_normalize(:nfd))
|
|
33
|
+
end
|
|
34
|
+
|
|
35
|
+
def compose
|
|
36
|
+
chars(readable.unicode_normalize(:nfc))
|
|
37
|
+
end
|
|
38
|
+
|
|
39
|
+
def grapheme_length
|
|
40
|
+
readable.grapheme_clusters.length
|
|
41
|
+
end
|
|
42
|
+
|
|
43
|
+
def tidy_bytes(force = false)
|
|
44
|
+
chars(self.class.tidy_bytes(@wrapped_string, force))
|
|
45
|
+
end
|
|
46
|
+
|
|
47
|
+
def split(*args)
|
|
48
|
+
readable.split(*args).map { |part| self.class.new(part) }
|
|
49
|
+
end
|
|
50
|
+
|
|
51
|
+
def slice!(*args)
|
|
52
|
+
sliced = @wrapped_string.slice!(*args)
|
|
53
|
+
chars(sliced) if sliced
|
|
54
|
+
end
|
|
55
|
+
|
|
56
|
+
def reverse!
|
|
57
|
+
@wrapped_string = reverse.to_s
|
|
58
|
+
self
|
|
59
|
+
end
|
|
60
|
+
|
|
61
|
+
def tidy_bytes!(*args)
|
|
62
|
+
@wrapped_string = tidy_bytes(*args).to_s
|
|
63
|
+
self
|
|
64
|
+
end
|
|
65
|
+
|
|
66
|
+
def <=>(other)
|
|
67
|
+
to_s <=> other.to_s
|
|
68
|
+
end
|
|
69
|
+
|
|
70
|
+
def =~(other)
|
|
71
|
+
@wrapped_string =~ other
|
|
72
|
+
end
|
|
73
|
+
|
|
74
|
+
def match?(*args)
|
|
75
|
+
@wrapped_string.match?(*args)
|
|
76
|
+
end
|
|
77
|
+
|
|
78
|
+
def acts_like_string?
|
|
79
|
+
true
|
|
80
|
+
end
|
|
81
|
+
|
|
82
|
+
def as_json(options = nil)
|
|
83
|
+
to_s.respond_to?(:as_json) ? to_s.as_json(options) : to_s
|
|
84
|
+
end
|
|
85
|
+
|
|
86
|
+
def method_missing(method, ...)
|
|
87
|
+
retried = false
|
|
88
|
+
begin
|
|
89
|
+
result = @wrapped_string.__send__(method, ...)
|
|
90
|
+
if method.end_with?('!')
|
|
91
|
+
self if result
|
|
92
|
+
else
|
|
93
|
+
result.is_a?(String) ? chars(result) : result
|
|
94
|
+
end
|
|
95
|
+
rescue ArgumentError
|
|
96
|
+
raise if retried || @wrapped_string.valid_encoding?
|
|
97
|
+
|
|
98
|
+
retried = true
|
|
99
|
+
@wrapped_string = @wrapped_string.scrub
|
|
100
|
+
retry
|
|
101
|
+
end
|
|
102
|
+
end
|
|
103
|
+
|
|
104
|
+
def respond_to_missing?(method, include_private = false)
|
|
105
|
+
@wrapped_string.respond_to?(method, include_private) || super
|
|
106
|
+
end
|
|
107
|
+
|
|
108
|
+
def self.tidy_bytes(string, force = false)
|
|
109
|
+
return string if string.empty? || (string.valid_encoding? && string.ascii_only?)
|
|
110
|
+
return recode_windows1252(string) if force
|
|
111
|
+
|
|
112
|
+
string.scrub { |bad| recode_windows1252(bad) }
|
|
113
|
+
end
|
|
114
|
+
|
|
115
|
+
def self.recode_windows1252(string)
|
|
116
|
+
string.encode(Encoding::UTF_8, Encoding::Windows_1252, invalid: :replace, undef: :replace)
|
|
117
|
+
end
|
|
118
|
+
private_class_method :recode_windows1252
|
|
119
|
+
|
|
120
|
+
private
|
|
121
|
+
|
|
122
|
+
def readable
|
|
123
|
+
@wrapped_string.valid_encoding? ? @wrapped_string : @wrapped_string.scrub
|
|
124
|
+
end
|
|
125
|
+
|
|
126
|
+
def chars(string)
|
|
127
|
+
self.class.new(string)
|
|
128
|
+
end
|
|
129
|
+
end
|
|
130
|
+
end
|
|
@@ -0,0 +1,152 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
|
|
3
|
+
module FastMbChars
|
|
4
|
+
# Cut to max_bytes without splitting a character and without reading
|
|
5
|
+
# the whole string. A character that crosses the limit is kept if it ends
|
|
6
|
+
# within SLACK bytes. A longer one is dropped. Invalid bytes do not raise.
|
|
7
|
+
module Limiter
|
|
8
|
+
SLACK = 128
|
|
9
|
+
|
|
10
|
+
module_function
|
|
11
|
+
|
|
12
|
+
def limit(text, max_bytes)
|
|
13
|
+
text = utf8(text)
|
|
14
|
+
return +'' if max_bytes <= 0 || text.empty?
|
|
15
|
+
return text.dup if text.bytesize <= max_bytes
|
|
16
|
+
|
|
17
|
+
prev = text.getbyte(max_bytes - 1)
|
|
18
|
+
nxt = text.getbyte(max_bytes)
|
|
19
|
+
if ascii_boundary?(prev, nxt)
|
|
20
|
+
cut = text.byteslice(0, max_bytes)
|
|
21
|
+
return cut if cut.valid_encoding?
|
|
22
|
+
end
|
|
23
|
+
|
|
24
|
+
cut_at_grapheme(text, max_bytes)
|
|
25
|
+
rescue ArgumentError
|
|
26
|
+
safe_prefix(text, max_bytes)
|
|
27
|
+
end
|
|
28
|
+
|
|
29
|
+
def utf8(text)
|
|
30
|
+
text = text.to_s
|
|
31
|
+
return text if text.encoding == Encoding::UTF_8
|
|
32
|
+
|
|
33
|
+
text.dup.force_encoding(Encoding::UTF_8)
|
|
34
|
+
end
|
|
35
|
+
|
|
36
|
+
def ascii_boundary?(prev, nxt)
|
|
37
|
+
prev && nxt && prev < 0x80 && nxt < 0x80 && !(prev == 0x0D && nxt == 0x0A)
|
|
38
|
+
end
|
|
39
|
+
|
|
40
|
+
def cut_at_grapheme(text, max_bytes)
|
|
41
|
+
start = align_start(text, [max_bytes - SLACK, 0].max)
|
|
42
|
+
hard_end = [text.bytesize, max_bytes + SLACK].min
|
|
43
|
+
finish = align_end(text, hard_end, max_bytes + SLACK)
|
|
44
|
+
finish = start if finish < start
|
|
45
|
+
|
|
46
|
+
window = text.byteslice(start, finish - start)
|
|
47
|
+
return safe_prefix(text, max_bytes) unless window.valid_encoding?
|
|
48
|
+
|
|
49
|
+
rel = max_bytes - start
|
|
50
|
+
taken = 0
|
|
51
|
+
keep = 0
|
|
52
|
+
window.each_grapheme_cluster do |cluster|
|
|
53
|
+
size = cluster.bytesize
|
|
54
|
+
break if taken >= rel
|
|
55
|
+
|
|
56
|
+
if taken + size <= rel
|
|
57
|
+
taken += size
|
|
58
|
+
keep = taken
|
|
59
|
+
next
|
|
60
|
+
end
|
|
61
|
+
|
|
62
|
+
cluster_end = start + taken + size
|
|
63
|
+
if cluster_end <= max_bytes + SLACK && grapheme_closed?(text, start + taken, cluster_end)
|
|
64
|
+
keep = taken + size
|
|
65
|
+
elsif taken.positive?
|
|
66
|
+
keep = taken
|
|
67
|
+
else
|
|
68
|
+
return safe_prefix(text, max_bytes)
|
|
69
|
+
end
|
|
70
|
+
break
|
|
71
|
+
end
|
|
72
|
+
|
|
73
|
+
text.byteslice(0, start + keep)
|
|
74
|
+
end
|
|
75
|
+
|
|
76
|
+
def grapheme_closed?(text, cluster_start, cluster_end)
|
|
77
|
+
return true if cluster_end >= text.bytesize
|
|
78
|
+
|
|
79
|
+
probe_end = next_char_end(text, cluster_end)
|
|
80
|
+
return true if probe_end <= cluster_end
|
|
81
|
+
|
|
82
|
+
sample = text.byteslice(cluster_start, probe_end - cluster_start)
|
|
83
|
+
return false unless sample&.valid_encoding?
|
|
84
|
+
|
|
85
|
+
matched = sample[/\A\X/]
|
|
86
|
+
matched&.bytesize == cluster_end - cluster_start
|
|
87
|
+
rescue ArgumentError
|
|
88
|
+
false
|
|
89
|
+
end
|
|
90
|
+
|
|
91
|
+
def align_start(text, index)
|
|
92
|
+
index = 0 if index.negative?
|
|
93
|
+
return text.bytesize if index >= text.bytesize
|
|
94
|
+
|
|
95
|
+
steps = 0
|
|
96
|
+
while index.positive? && steps < 3 && continuation?(text.getbyte(index))
|
|
97
|
+
index -= 1
|
|
98
|
+
steps += 1
|
|
99
|
+
end
|
|
100
|
+
index
|
|
101
|
+
end
|
|
102
|
+
|
|
103
|
+
def align_end(text, index, hard_end)
|
|
104
|
+
return 0 if index <= 0
|
|
105
|
+
return text.bytesize if index >= text.bytesize
|
|
106
|
+
return index unless continuation?(text.getbyte(index))
|
|
107
|
+
|
|
108
|
+
lead = align_start(text, index)
|
|
109
|
+
length = seq_length(text.getbyte(lead))
|
|
110
|
+
char_end = length ? lead + length : lead
|
|
111
|
+
char_end <= hard_end && char_end <= text.bytesize ? char_end : lead
|
|
112
|
+
end
|
|
113
|
+
|
|
114
|
+
def next_char_end(text, index)
|
|
115
|
+
return index if index >= text.bytesize
|
|
116
|
+
|
|
117
|
+
length = seq_length(text.getbyte(index)) || 1
|
|
118
|
+
[index + length, text.bytesize].min
|
|
119
|
+
end
|
|
120
|
+
|
|
121
|
+
def seq_length(lead)
|
|
122
|
+
return 1 if lead < 0x80
|
|
123
|
+
return 2 if (lead & 0xE0) == 0xC0
|
|
124
|
+
return 3 if (lead & 0xF0) == 0xE0
|
|
125
|
+
return 4 if (lead & 0xF8) == 0xF0
|
|
126
|
+
|
|
127
|
+
nil
|
|
128
|
+
end
|
|
129
|
+
|
|
130
|
+
def continuation?(byte)
|
|
131
|
+
byte && (byte & 0xC0) == 0x80
|
|
132
|
+
end
|
|
133
|
+
|
|
134
|
+
def safe_prefix(text, max_bytes)
|
|
135
|
+
limit = [max_bytes, text.bytesize].min
|
|
136
|
+
if limit.positive? && limit < text.bytesize && continuation?(text.getbyte(limit))
|
|
137
|
+
limit = align_start(text, limit)
|
|
138
|
+
end
|
|
139
|
+
cut = text.byteslice(0, limit)
|
|
140
|
+
return cut if cut.valid_encoding?
|
|
141
|
+
|
|
142
|
+
3.times do
|
|
143
|
+
break if cut.empty?
|
|
144
|
+
|
|
145
|
+
cut = cut.byteslice(0, cut.bytesize - 1)
|
|
146
|
+
return cut if cut.valid_encoding?
|
|
147
|
+
end
|
|
148
|
+
|
|
149
|
+
text.byteslice(0, [max_bytes, text.bytesize].min).scrub('?')
|
|
150
|
+
end
|
|
151
|
+
end
|
|
152
|
+
end
|
|
@@ -0,0 +1,24 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
|
|
3
|
+
require 'fast_mb_chars/version'
|
|
4
|
+
require 'fast_mb_chars/limiter'
|
|
5
|
+
require 'fast_mb_chars/chars'
|
|
6
|
+
|
|
7
|
+
module FastMbChars
|
|
8
|
+
# Prepends a fresh module each time the method was taken over by someone
|
|
9
|
+
# else. Prepending the same module twice does not move it back to the front.
|
|
10
|
+
def self.install!
|
|
11
|
+
if String.method_defined?(:mb_chars) && String.instance_method(:mb_chars).owner == @installed
|
|
12
|
+
return
|
|
13
|
+
end
|
|
14
|
+
|
|
15
|
+
@installed = Module.new do
|
|
16
|
+
def mb_chars
|
|
17
|
+
FastMbChars::Chars.new(self)
|
|
18
|
+
end
|
|
19
|
+
end
|
|
20
|
+
String.prepend(@installed)
|
|
21
|
+
end
|
|
22
|
+
|
|
23
|
+
install!
|
|
24
|
+
end
|
metadata
ADDED
|
@@ -0,0 +1,66 @@
|
|
|
1
|
+
--- !ruby/object:Gem::Specification
|
|
2
|
+
name: fast_mb_chars
|
|
3
|
+
version: !ruby/object:Gem::Version
|
|
4
|
+
version: 0.1.0
|
|
5
|
+
platform: ruby
|
|
6
|
+
authors:
|
|
7
|
+
- SYNORYXEL
|
|
8
|
+
autorequire:
|
|
9
|
+
bindir: bin
|
|
10
|
+
cert_chain: []
|
|
11
|
+
date: 2026-10-07 00:00:00.000000000 Z
|
|
12
|
+
dependencies: []
|
|
13
|
+
description: |
|
|
14
|
+
Faster String#mb_chars. limit(bytes) cuts on a character boundary without
|
|
15
|
+
walking the whole string and without raising on a broken byte. A 1 MB cut
|
|
16
|
+
is about 2,000 times faster than the ActiveSupport grapheme scan, and a
|
|
17
|
+
7 MB cut is about 3,000 times faster.
|
|
18
|
+
|
|
19
|
+
downcase, upcase, titleize, reverse, split, slice!, compose, decompose,
|
|
20
|
+
tidy_bytes, and the other string methods are forwarded through method_missing.
|
|
21
|
+
They still return mb_chars, so existing chains keep working.
|
|
22
|
+
email:
|
|
23
|
+
- mike.pvlv@icloud.com
|
|
24
|
+
executables: []
|
|
25
|
+
extensions: []
|
|
26
|
+
extra_rdoc_files: []
|
|
27
|
+
files:
|
|
28
|
+
- CHANGELOG.md
|
|
29
|
+
- LICENSE
|
|
30
|
+
- README.md
|
|
31
|
+
- benchmark/compare.rb
|
|
32
|
+
- docs/limit-cut.svg
|
|
33
|
+
- lib/fast_mb_chars.rb
|
|
34
|
+
- lib/fast_mb_chars/chars.rb
|
|
35
|
+
- lib/fast_mb_chars/limiter.rb
|
|
36
|
+
- lib/fast_mb_chars/version.rb
|
|
37
|
+
homepage: https://github.com/SYNORYXEL/fast_mb_chars
|
|
38
|
+
licenses:
|
|
39
|
+
- MIT
|
|
40
|
+
metadata:
|
|
41
|
+
homepage_uri: https://github.com/SYNORYXEL/fast_mb_chars
|
|
42
|
+
source_code_uri: https://github.com/SYNORYXEL/fast_mb_chars
|
|
43
|
+
changelog_uri: https://github.com/SYNORYXEL/fast_mb_chars/blob/main/CHANGELOG.md
|
|
44
|
+
bug_tracker_uri: https://github.com/SYNORYXEL/fast_mb_chars/issues
|
|
45
|
+
allowed_push_host: https://rubygems.org
|
|
46
|
+
rubygems_mfa_required: 'true'
|
|
47
|
+
post_install_message:
|
|
48
|
+
rdoc_options: []
|
|
49
|
+
require_paths:
|
|
50
|
+
- lib
|
|
51
|
+
required_ruby_version: !ruby/object:Gem::Requirement
|
|
52
|
+
requirements:
|
|
53
|
+
- - ">="
|
|
54
|
+
- !ruby/object:Gem::Version
|
|
55
|
+
version: 2.7.0
|
|
56
|
+
required_rubygems_version: !ruby/object:Gem::Requirement
|
|
57
|
+
requirements:
|
|
58
|
+
- - ">="
|
|
59
|
+
- !ruby/object:Gem::Version
|
|
60
|
+
version: '0'
|
|
61
|
+
requirements: []
|
|
62
|
+
rubygems_version: 3.5.22
|
|
63
|
+
signing_key:
|
|
64
|
+
specification_version: 4
|
|
65
|
+
summary: Faster Ruby String#mb_chars. Byte limit that keeps UTF-8 characters whole
|
|
66
|
+
test_files: []
|