fast_mb_chars 0.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml ADDED
@@ -0,0 +1,7 @@
1
+ ---
2
+ SHA256:
3
+ metadata.gz: 35532c3219891f1ff4d3a7fe168abb9213bafa430d7c33051eb2ac58f24ab0a5
4
+ data.tar.gz: 64672d7dd534037d5bda31bd52e606246f527583bd5880ca33100f4b0d435bf4
5
+ SHA512:
6
+ metadata.gz: 21021182d78d1ba4d51d518546161f95a412709dfa998f2a20068c542cd874932a6ea3de8cc981b9e6d7f77f6651a9eabef480ffdc4f6add5de179354cfb6ba1
7
+ data.tar.gz: e8c45c5e5aa21babbbe604151228ddd39bea12882d1dee684761ab71fdd166da801e7159e7f1ffd6161aef816a15f2d2d0e1f3d002d568ab02ea9760ed690d2e
data/CHANGELOG.md ADDED
@@ -0,0 +1,9 @@
1
+ # Changelog
2
+
3
+ ## 0.1.0
4
+
5
+ `String#mb_chars` with the usual case, split, and Unicode helpers, plus `limit(bytes)`.
6
+
7
+ The gem prepends `String#mb_chars`. ASCII and UTF-8 are supported. Other encodings need `encode('UTF-8')` first.
8
+
9
+ The byte cut stays on a character boundary. A character that crosses the limit is kept when it finishes within 128 bytes. A longer combining sequence is dropped. Invalid bytes do not raise.
data/LICENSE ADDED
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 SYNORYXEL
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
data/README.md ADDED
@@ -0,0 +1,180 @@
1
+ # ⚡ fast_mb_chars
2
+
3
+ Faster `String#mb_chars` for Ruby. Same calls as `ActiveSupport::Multibyte::Chars` — `limit`, `downcase`, `upcase`, `titleize`, `reverse`, `split`, `tidy_bytes`. The byte cut does not split a character and does not raise on a broken byte.
4
+
5
+ Works on **ASCII** and **UTF-8** (Київ, café, 日本語, emoji). Other encodings need a convert first.
6
+
7
+ <p align="center">
8
+ <img src="docs/limit-cut.svg" alt="Byte limit lands inside я; the gem keeps the whole letter" width="880">
9
+ </p>
10
+
11
+ `mb_chars.limit` in Rails reads every character from the start of the string. This gem looks at the cut. On a 1 MB string that is **205 ms** vs **0.08 ms**.
12
+
13
+ | | 1 MB | 7 MB |
14
+ | --- | ---: | ---: |
15
+ | ActiveSupport-style scan | 205 ms | 1,410 ms |
16
+ | **fast_mb_chars** | **0.08 ms** | **0.47 ms** |
17
+ | | **~2,500×** | **~3,000×** |
18
+
19
+ ASCII, Ruby 4.0.5, one call. The string is longer than the limit, so both sides actually cut. Full tables are [below](#-benchmarks).
20
+
21
+ [![test](https://github.com/SYNORYXEL/fast_mb_chars/actions/workflows/ci.yml/badge.svg)](https://github.com/SYNORYXEL/fast_mb_chars/actions)
22
+ [![Ruby](https://img.shields.io/badge/ruby-%3E%3D%202.7-cc342d?logo=ruby&logoColor=white)](https://www.ruby-lang.org/)
23
+ [![License: MIT](https://img.shields.io/badge/license-MIT-blue.svg)](LICENSE)
24
+ [![UTF-8](https://img.shields.io/badge/encoding-UTF--8%20%2B%20ASCII-1a7f37)](#-encodings)
25
+
26
+ Source: [github.com/SYNORYXEL/fast_mb_chars](https://github.com/SYNORYXEL/fast_mb_chars)
27
+
28
+ ## 🌍 Encodings
29
+
30
+ The gem treats the string as **UTF-8 bytes**. ASCII is UTF-8. It is not a transcoding library.
31
+
32
+ | Input | What happens |
33
+ | --- | --- |
34
+ | ASCII `hello` | Fast path. Both bytes on the boundary are `< 0x80`, so the cut is exact. |
35
+ | UTF-8 `Київ`, `café`, `日本語`, `👨‍👩‍👧‍👦` | Yes. A letter or emoji that crosses the limit stays whole. |
36
+ | Broken / binary bytes | No `ArgumentError`. Invalid sequences are skipped or scrubbed. |
37
+ | Windows-1252 mixed into UTF-8 (`caf\xE9`) | `tidy_bytes` repairs it to `café`. |
38
+ | Windows-1251, KOI8-R, ISO-8859-1, Shift_JIS, UTF-16 | **No.** Bytes are relabeled as UTF-8, not converted. Encode first. |
39
+
40
+ ```ruby
41
+ # already UTF-8 or ASCII — just call it
42
+ 'Київ'.mb_chars.downcase.to_s
43
+ 'hello'.mb_chars.limit(4).to_s
44
+
45
+ # other encodings: convert, then cut
46
+ cp1251.encode('UTF-8').mb_chars.limit(1_048_576).to_s
47
+ ```
48
+
49
+ `force_encoding('UTF-8')` only changes the label. `encode('UTF-8')` changes the bytes. Use the second one when the source is not UTF-8.
50
+
51
+ ## 📦 Install
52
+
53
+ Ruby 2.7 or newer. Tests run on 4.0.5. CI runs 2.7, 3.2, 3.3, 3.4, and 4.0.
54
+
55
+ ```ruby
56
+ gem 'fast_mb_chars', github: 'SYNORYXEL/fast_mb_chars'
57
+ ```
58
+
59
+ ```ruby
60
+ require 'fast_mb_chars'
61
+ ```
62
+
63
+ The gem prepends `String#mb_chars`, so existing `'text'.mb_chars.downcase` calls in a Rails app keep working. It does not patch `String#downcase`. If another library prepends `mb_chars` later, call `FastMbChars.install!` again. No runtime dependencies. Not on RubyGems yet — install from GitHub.
64
+
65
+ From a checkout: `gem build fast_mb_chars.gemspec` then `gem install ./fast_mb_chars-0.1.0.gem`. Next to an app, `gem 'fast_mb_chars', path: '../fast_mb_chars'` works too.
66
+
67
+ ## ✨ Usage
68
+
69
+ ```ruby
70
+ 'ГОТІВКА'.mb_chars.downcase.to_s
71
+ # => "готівка"
72
+
73
+ ' Київ '.mb_chars.downcase.strip.to_s
74
+ # => "київ"
75
+
76
+ 'ÉL QUE SE ENTERÓ'.mb_chars.titleize.to_s
77
+ # => "Él Que Se Enteró"
78
+
79
+ 'Café périferôl'.mb_chars.split(/é/).map { |part| part.upcase.to_s }
80
+ # => ["CAF", " P", "RIFERÔL"]
81
+
82
+ data.mb_chars.limit(1_048_576).to_s
83
+ data.mb_chars.limit(7_340_032).to_s
84
+ ```
85
+
86
+ Methods on the proxy itself: `limit`, `reverse`, `reverse!`, `titleize`, `titlecase`, `compose`, `decompose`, `grapheme_length`, `tidy_bytes`, `tidy_bytes!`, `split`, `slice!`, `<=>`, `=~`, `match?`, `acts_like_string?`, `as_json`, `to_s`, `to_str`.
87
+
88
+ `downcase`, `upcase`, `capitalize`, `swapcase`, `strip`, `gsub`, `length`, and the rest of `String` go through `method_missing`. A string result comes back as `mb_chars`, so the chain does not break. A bang method that changes the string returns the same proxy; if Ruby returns `nil` (no change), this does too.
89
+
90
+ `titleize` is a Unicode-aware regex, not ActiveSupport inflections.
91
+
92
+ ## 🚀 Why it is faster
93
+
94
+ Rails does this:
95
+
96
+ ```ruby
97
+ def limit(limit)
98
+ chars(@wrapped_string.truncate_bytes(limit, omission: nil))
99
+ end
100
+ ```
101
+
102
+ `truncate_bytes` walks graphemes from byte 0 until it hits the limit. That is fine for a button label. It is not fine for a 1 MB or 7 MB log line.
103
+
104
+ `fast_mb_chars` reads the two bytes on the boundary.
105
+
106
+ - Both ASCII, and not a CR+LF pair: one `byteslice`. Exact.
107
+ - A character crosses the limit and ends within 128 bytes: keep it whole. The result can be a few bytes over the limit.
108
+ - A long run of combining marks that does not fit in those 128 bytes: drop it. The result stays within the limit.
109
+ - Empty string, or a limit of 0 or less: `""`.
110
+ - Shorter than the limit: returned as-is, trailing spaces and punctuation included.
111
+ - A broken byte never reaches a regexp, so there is no `ArgumentError`.
112
+
113
+ Rails never goes past the limit, so it drops the character on the boundary. This gem keeps a short character. That is the one behavioral difference, and it is why a cut like `('a' * (limit - 1)) + 'я'` comes back one byte longer, with `я` intact.
114
+
115
+ ```mermaid
116
+ flowchart LR
117
+ A[1 MB string] --> B{bytes at the limit}
118
+ B -->|both ASCII| C[byteslice exact]
119
+ B -->|UTF-8 letter or emoji| D[keep the cluster if it ends within 128 bytes]
120
+ B -->|broken byte| E[safe prefix no raise]
121
+ ```
122
+
123
+ ## 📊 Benchmarks
124
+
125
+ "scan" walks graphemes from the start, which is what `mb_chars.limit` does. "fast" is `FastMbChars::Limiter.limit`. Milliseconds per call, Ruby 4.0.5.
126
+
127
+ ### ASCII
128
+
129
+ The cut lands on a byte. Both results are exactly the limit.
130
+
131
+ | size | scan | fast | times |
132
+ | --- | ---: | ---: | ---: |
133
+ | 1 KB | 0.198 | 0.000 | 499× |
134
+ | 8 KB | 1.566 | 0.001 | 1,284× |
135
+ | 64 KB | 12.671 | 0.010 | 1,233× |
136
+ | 256 KB | 50.936 | 0.028 | 1,841× |
137
+ | 1 MB | 205.382 | 0.083 | 2,467× |
138
+ | 7 MB | 1,409.907 | 0.466 | 3,028× |
139
+
140
+ ### Cyrillic
141
+
142
+ The cut lands between letters. Both results are exactly the limit.
143
+
144
+ | size | scan | fast | times |
145
+ | --- | ---: | ---: | ---: |
146
+ | 1 KB | 0.145 | 0.020 | 7× |
147
+ | 8 KB | 1.145 | 0.022 | 52× |
148
+ | 64 KB | 9.282 | 0.023 | 407× |
149
+ | 256 KB | 36.863 | 0.040 | 931× |
150
+ | 1 MB | 147.932 | 0.068 | 2,164× |
151
+ | 7 MB | 1,032.553 | 0.194 | 5,322× |
152
+
153
+ ### Character on the boundary
154
+
155
+ `я` crosses the limit. The scan stops on the limit and drops the letter. The fast cut keeps it, so the result is one byte over.
156
+
157
+ | size | scan | fast | times | fast bytes / limit |
158
+ | --- | ---: | ---: | ---: | ---: |
159
+ | 1 KB | 0.196 | 0.027 | 7× | 1025 / 1024 |
160
+ | 8 KB | 1.545 | 0.029 | 54× | 8193 / 8192 |
161
+ | 64 KB | 12.377 | 0.028 | 443× | 65537 / 65536 |
162
+ | 256 KB | 49.862 | 0.026 | 1,893× | 262145 / 262144 |
163
+ | 1 MB | 199.377 | 0.036 | 5,597× | 1048577 / 1048576 |
164
+ | 7 MB | 1,479.463 | 0.030 | 49,870× | 7340033 / 7340032 |
165
+
166
+ On a short string you pay for the wrapper, not for the cut. `'ГОТІВКА'.mb_chars.downcase.to_s` was 0.0006 ms. `String#downcase` was 0.0002 ms.
167
+
168
+ ```bash
169
+ ruby benchmark/compare.rb
170
+ ```
171
+
172
+ ## ✅ Tests
173
+
174
+ ```bash
175
+ bundle exec rake
176
+ ```
177
+
178
+ ## 📄 License
179
+
180
+ MIT. See [LICENSE](LICENSE).
@@ -0,0 +1,82 @@
1
+ # frozen_string_literal: true
2
+
3
+ # ruby benchmark/compare.rb
4
+ # rake benchmark
5
+
6
+ $LOAD_PATH.unshift(File.expand_path('../lib', __dir__))
7
+
8
+ require 'fast_mb_chars'
9
+
10
+ # What ActiveSupport::Multibyte::Chars#limit does: walk every grapheme from
11
+ # the start and never go past the byte budget.
12
+ def walk_limit(text, max_bytes)
13
+ text = text.to_s
14
+ return +'' if max_bytes <= 0 || text.empty?
15
+ return text.dup if text.bytesize <= max_bytes
16
+
17
+ taken = 0
18
+ text.each_grapheme_cluster do |cluster|
19
+ size = cluster.bytesize
20
+ break if taken + size > max_bytes
21
+
22
+ taken += size
23
+ end
24
+ text.byteslice(0, taken)
25
+ end
26
+
27
+ SIZES = [
28
+ ['1 KB', 1_024, 300],
29
+ ['8 KB', 8_192, 150],
30
+ ['64 KB', 65_536, 40],
31
+ ['256 KB', 262_144, 15],
32
+ ['1 MB', 1_048_576, 8],
33
+ ['7 MB', 7_340_032, 3]
34
+ ].freeze
35
+
36
+ # The string is longer than the limit, so both sides actually cut.
37
+ # split puts one 2-byte letter across the boundary.
38
+ def payload(kind, bytes)
39
+ case kind
40
+ when :ascii
41
+ ['a' * (bytes + 64), bytes]
42
+ when :cyrillic
43
+ ['я' * ((bytes / 2) + 16), bytes]
44
+ when :split
45
+ [(('a' * (bytes - 1)) + 'я'), bytes]
46
+ end
47
+ end
48
+
49
+ def time_ms(iterations)
50
+ started = Process.clock_gettime(Process::CLOCK_MONOTONIC)
51
+ iterations.times { yield }
52
+ ((Process.clock_gettime(Process::CLOCK_MONOTONIC) - started) / iterations) * 1000.0
53
+ end
54
+
55
+ rows = []
56
+
57
+ puts "ruby #{RUBY_VERSION} fast_mb_chars #{FastMbChars::VERSION}"
58
+ puts "limit vs full grapheme walk (ms per call, one process)"
59
+ puts
60
+
61
+ %i[ascii cyrillic split].each do |kind|
62
+ puts kind
63
+ puts format('%-8s %10s %10s %8s %s', 'size', 'walk', 'fast', 'times', 'result')
64
+ SIZES.each do |label, bytes, iterations|
65
+ text, limit = payload(kind, bytes)
66
+ 2.times { walk_limit(text, limit) }
67
+ 2.times { FastMbChars::Limiter.limit(text, limit) }
68
+ slow = time_ms(iterations) { walk_limit(text, limit) }
69
+ fast = time_ms(iterations) { FastMbChars::Limiter.limit(text, limit) }
70
+ ratio = fast.positive? ? slow / fast : 0
71
+ cut = FastMbChars::Limiter.limit(text, limit)
72
+ rows << [kind, label, slow, fast, ratio, cut.bytesize, limit]
73
+ puts format('%-8s %10.3f %10.3f %7.0fx %d/%d bytes', label, slow, fast, ratio, cut.bytesize, limit)
74
+ end
75
+ puts
76
+ end
77
+
78
+ small = 'ГОТІВКА'
79
+ time_ms(1000) { small.mb_chars.downcase.to_s }
80
+ down = time_ms(20_000) { small.mb_chars.downcase.to_s }
81
+ plain = time_ms(20_000) { small.downcase }
82
+ puts format('downcase %s mb_chars %.4f ms String#downcase %.4f ms', small, down, plain)
@@ -0,0 +1,66 @@
1
+ <svg xmlns="http://www.w3.org/2000/svg" width="880" height="300" viewBox="0 0 880 300" role="img" aria-labelledby="title desc">
2
+ <title id="title">fast_mb_chars cuts on a character boundary</title>
3
+ <desc id="desc">The byte limit lands inside a UTF-8 letter. The gem keeps the whole character instead of splitting it.</desc>
4
+ <defs>
5
+ <style>
6
+ .bg { fill: #f6f8fa; }
7
+ .card { fill: #ffffff; stroke: #d0d7de; }
8
+ .ink { fill: #1f2328; }
9
+ .muted { fill: #656d76; }
10
+ .line { stroke: #d0d7de; }
11
+ .keep { fill: #1a7f37; }
12
+ .drop { fill: #cf222e; }
13
+ .accent { fill: #cf222e; }
14
+ .chip { fill: #fff8f8; stroke: #ffcecb; }
15
+ .cell { fill: #ffffff; stroke: #d0d7de; }
16
+ .cell-keep { fill: #dafbe1; stroke: #4ac26b; }
17
+ .cell-cut { fill: #fff8c5; stroke: #d4a72c; }
18
+ @media (prefers-color-scheme: dark) {
19
+ .bg { fill: #0d1117; }
20
+ .card { fill: #161b22; stroke: #30363d; }
21
+ .ink { fill: #e6edf3; }
22
+ .muted { fill: #8b949e; }
23
+ .line { stroke: #30363d; }
24
+ .keep { fill: #3fb950; }
25
+ .drop { fill: #f85149; }
26
+ .accent { fill: #f85149; }
27
+ .chip { fill: #3d1619; stroke: #6e2b2f; }
28
+ .cell { fill: #0d1117; stroke: #30363d; }
29
+ .cell-keep { fill: #12261e; stroke: #238636; }
30
+ .cell-cut { fill: #2a2111; stroke: #9e6a03; }
31
+ }
32
+ </style>
33
+ </defs>
34
+ <rect class="bg" width="880" height="300" rx="16"/>
35
+ <rect class="card" x="16" y="16" width="848" height="268" rx="12"/>
36
+
37
+ <text class="ink" x="40" y="52" font-family="ui-sans-serif, system-ui, sans-serif" font-size="18" font-weight="700">limit looks at the cut, not the whole string</text>
38
+ <text class="muted" x="40" y="76" font-family="ui-sans-serif, system-ui, sans-serif" font-size="13">ASCII and UTF-8 &#x2014; characters stay whole</text>
39
+
40
+ <g font-family="ui-monospace, SFMono-Regular, Menlo, Consolas, monospace" font-size="14">
41
+ <rect class="cell" x="40" y="104" width="44" height="44" rx="8"/>
42
+ <text class="ink" x="62" y="131" text-anchor="middle">a</text>
43
+ <rect class="cell" x="90" y="104" width="44" height="44" rx="8"/>
44
+ <text class="ink" x="112" y="131" text-anchor="middle">a</text>
45
+ <rect class="cell" x="140" y="104" width="44" height="44" rx="8"/>
46
+ <text class="ink" x="162" y="131" text-anchor="middle">a</text>
47
+ <rect class="cell-cut" x="190" y="104" width="44" height="44" rx="8"/>
48
+ <text class="ink" x="212" y="131" text-anchor="middle">a</text>
49
+ <rect class="cell-keep" x="240" y="104" width="92" height="44" rx="8"/>
50
+ <text class="keep" x="286" y="131" text-anchor="middle" font-weight="700">&#x44F;</text>
51
+ <rect class="cell" x="338" y="104" width="44" height="44" rx="8"/>
52
+ <text class="muted" x="360" y="131" text-anchor="middle">&#x2026;</text>
53
+ </g>
54
+
55
+ <path class="line" d="M234 156 v18" fill="none" stroke-width="2"/>
56
+ <text class="accent" x="234" y="196" text-anchor="middle" font-family="ui-sans-serif, system-ui, sans-serif" font-size="12" font-weight="700">byte limit</text>
57
+ <text class="muted" x="234" y="214" text-anchor="middle" font-family="ui-sans-serif, system-ui, sans-serif" font-size="11">inside the letter</text>
58
+
59
+ <rect class="chip" x="430" y="104" width="400" height="108" rx="10"/>
60
+ <text class="drop" x="450" y="132" font-family="ui-sans-serif, system-ui, sans-serif" font-size="13" font-weight="700">Rails scan</text>
61
+ <text class="muted" x="450" y="152" font-family="ui-sans-serif, system-ui, sans-serif" font-size="12">walks graphemes from byte 0 &#x2192; drops &#x44F;</text>
62
+ <text class="keep" x="450" y="180" font-family="ui-sans-serif, system-ui, sans-serif" font-size="13" font-weight="700">fast_mb_chars</text>
63
+ <text class="muted" x="450" y="200" font-family="ui-sans-serif, system-ui, sans-serif" font-size="12">reads the two bytes on the boundary &#x2192; keeps &#x44F;</text>
64
+
65
+ <text class="muted" x="40" y="256" font-family="ui-monospace, SFMono-Regular, Menlo, Consolas, monospace" font-size="12">('a' * (limit - 1)) + '&#x44F;' &#x2192; 1 extra byte, letter intact</text>
66
+ </svg>
@@ -0,0 +1,130 @@
1
+ # frozen_string_literal: true
2
+
3
+ module FastMbChars
4
+ # Proxy with the same methods as ActiveSupport::Multibyte::Chars.
5
+ # limit is the fast path. Everything else follows the string.
6
+ class Chars
7
+ include Comparable
8
+
9
+ attr_reader :wrapped_string
10
+ alias to_s wrapped_string
11
+ alias to_str wrapped_string
12
+
13
+ def initialize(text)
14
+ @wrapped_string = Limiter.utf8(text)
15
+ @wrapped_string = @wrapped_string.dup if @wrapped_string.frozen?
16
+ end
17
+
18
+ def limit(max_bytes)
19
+ chars(Limiter.limit(@wrapped_string, Integer(max_bytes)))
20
+ end
21
+
22
+ def reverse
23
+ chars(readable.grapheme_clusters.reverse.join)
24
+ end
25
+
26
+ def titleize
27
+ chars(readable.downcase.gsub(/\b('?\S)/u) { Regexp.last_match(1).upcase })
28
+ end
29
+ alias titlecase titleize
30
+
31
+ def decompose
32
+ chars(readable.unicode_normalize(:nfd))
33
+ end
34
+
35
+ def compose
36
+ chars(readable.unicode_normalize(:nfc))
37
+ end
38
+
39
+ def grapheme_length
40
+ readable.grapheme_clusters.length
41
+ end
42
+
43
+ def tidy_bytes(force = false)
44
+ chars(self.class.tidy_bytes(@wrapped_string, force))
45
+ end
46
+
47
+ def split(*args)
48
+ readable.split(*args).map { |part| self.class.new(part) }
49
+ end
50
+
51
+ def slice!(*args)
52
+ sliced = @wrapped_string.slice!(*args)
53
+ chars(sliced) if sliced
54
+ end
55
+
56
+ def reverse!
57
+ @wrapped_string = reverse.to_s
58
+ self
59
+ end
60
+
61
+ def tidy_bytes!(*args)
62
+ @wrapped_string = tidy_bytes(*args).to_s
63
+ self
64
+ end
65
+
66
+ def <=>(other)
67
+ to_s <=> other.to_s
68
+ end
69
+
70
+ def =~(other)
71
+ @wrapped_string =~ other
72
+ end
73
+
74
+ def match?(*args)
75
+ @wrapped_string.match?(*args)
76
+ end
77
+
78
+ def acts_like_string?
79
+ true
80
+ end
81
+
82
+ def as_json(options = nil)
83
+ to_s.respond_to?(:as_json) ? to_s.as_json(options) : to_s
84
+ end
85
+
86
+ def method_missing(method, ...)
87
+ retried = false
88
+ begin
89
+ result = @wrapped_string.__send__(method, ...)
90
+ if method.end_with?('!')
91
+ self if result
92
+ else
93
+ result.is_a?(String) ? chars(result) : result
94
+ end
95
+ rescue ArgumentError
96
+ raise if retried || @wrapped_string.valid_encoding?
97
+
98
+ retried = true
99
+ @wrapped_string = @wrapped_string.scrub
100
+ retry
101
+ end
102
+ end
103
+
104
+ def respond_to_missing?(method, include_private = false)
105
+ @wrapped_string.respond_to?(method, include_private) || super
106
+ end
107
+
108
+ def self.tidy_bytes(string, force = false)
109
+ return string if string.empty? || (string.valid_encoding? && string.ascii_only?)
110
+ return recode_windows1252(string) if force
111
+
112
+ string.scrub { |bad| recode_windows1252(bad) }
113
+ end
114
+
115
+ def self.recode_windows1252(string)
116
+ string.encode(Encoding::UTF_8, Encoding::Windows_1252, invalid: :replace, undef: :replace)
117
+ end
118
+ private_class_method :recode_windows1252
119
+
120
+ private
121
+
122
+ def readable
123
+ @wrapped_string.valid_encoding? ? @wrapped_string : @wrapped_string.scrub
124
+ end
125
+
126
+ def chars(string)
127
+ self.class.new(string)
128
+ end
129
+ end
130
+ end
@@ -0,0 +1,152 @@
1
+ # frozen_string_literal: true
2
+
3
+ module FastMbChars
4
+ # Cut to max_bytes without splitting a character and without reading
5
+ # the whole string. A character that crosses the limit is kept if it ends
6
+ # within SLACK bytes. A longer one is dropped. Invalid bytes do not raise.
7
+ module Limiter
8
+ SLACK = 128
9
+
10
+ module_function
11
+
12
+ def limit(text, max_bytes)
13
+ text = utf8(text)
14
+ return +'' if max_bytes <= 0 || text.empty?
15
+ return text.dup if text.bytesize <= max_bytes
16
+
17
+ prev = text.getbyte(max_bytes - 1)
18
+ nxt = text.getbyte(max_bytes)
19
+ if ascii_boundary?(prev, nxt)
20
+ cut = text.byteslice(0, max_bytes)
21
+ return cut if cut.valid_encoding?
22
+ end
23
+
24
+ cut_at_grapheme(text, max_bytes)
25
+ rescue ArgumentError
26
+ safe_prefix(text, max_bytes)
27
+ end
28
+
29
+ def utf8(text)
30
+ text = text.to_s
31
+ return text if text.encoding == Encoding::UTF_8
32
+
33
+ text.dup.force_encoding(Encoding::UTF_8)
34
+ end
35
+
36
+ def ascii_boundary?(prev, nxt)
37
+ prev && nxt && prev < 0x80 && nxt < 0x80 && !(prev == 0x0D && nxt == 0x0A)
38
+ end
39
+
40
+ def cut_at_grapheme(text, max_bytes)
41
+ start = align_start(text, [max_bytes - SLACK, 0].max)
42
+ hard_end = [text.bytesize, max_bytes + SLACK].min
43
+ finish = align_end(text, hard_end, max_bytes + SLACK)
44
+ finish = start if finish < start
45
+
46
+ window = text.byteslice(start, finish - start)
47
+ return safe_prefix(text, max_bytes) unless window.valid_encoding?
48
+
49
+ rel = max_bytes - start
50
+ taken = 0
51
+ keep = 0
52
+ window.each_grapheme_cluster do |cluster|
53
+ size = cluster.bytesize
54
+ break if taken >= rel
55
+
56
+ if taken + size <= rel
57
+ taken += size
58
+ keep = taken
59
+ next
60
+ end
61
+
62
+ cluster_end = start + taken + size
63
+ if cluster_end <= max_bytes + SLACK && grapheme_closed?(text, start + taken, cluster_end)
64
+ keep = taken + size
65
+ elsif taken.positive?
66
+ keep = taken
67
+ else
68
+ return safe_prefix(text, max_bytes)
69
+ end
70
+ break
71
+ end
72
+
73
+ text.byteslice(0, start + keep)
74
+ end
75
+
76
+ def grapheme_closed?(text, cluster_start, cluster_end)
77
+ return true if cluster_end >= text.bytesize
78
+
79
+ probe_end = next_char_end(text, cluster_end)
80
+ return true if probe_end <= cluster_end
81
+
82
+ sample = text.byteslice(cluster_start, probe_end - cluster_start)
83
+ return false unless sample&.valid_encoding?
84
+
85
+ matched = sample[/\A\X/]
86
+ matched&.bytesize == cluster_end - cluster_start
87
+ rescue ArgumentError
88
+ false
89
+ end
90
+
91
+ def align_start(text, index)
92
+ index = 0 if index.negative?
93
+ return text.bytesize if index >= text.bytesize
94
+
95
+ steps = 0
96
+ while index.positive? && steps < 3 && continuation?(text.getbyte(index))
97
+ index -= 1
98
+ steps += 1
99
+ end
100
+ index
101
+ end
102
+
103
+ def align_end(text, index, hard_end)
104
+ return 0 if index <= 0
105
+ return text.bytesize if index >= text.bytesize
106
+ return index unless continuation?(text.getbyte(index))
107
+
108
+ lead = align_start(text, index)
109
+ length = seq_length(text.getbyte(lead))
110
+ char_end = length ? lead + length : lead
111
+ char_end <= hard_end && char_end <= text.bytesize ? char_end : lead
112
+ end
113
+
114
+ def next_char_end(text, index)
115
+ return index if index >= text.bytesize
116
+
117
+ length = seq_length(text.getbyte(index)) || 1
118
+ [index + length, text.bytesize].min
119
+ end
120
+
121
+ def seq_length(lead)
122
+ return 1 if lead < 0x80
123
+ return 2 if (lead & 0xE0) == 0xC0
124
+ return 3 if (lead & 0xF0) == 0xE0
125
+ return 4 if (lead & 0xF8) == 0xF0
126
+
127
+ nil
128
+ end
129
+
130
+ def continuation?(byte)
131
+ byte && (byte & 0xC0) == 0x80
132
+ end
133
+
134
+ def safe_prefix(text, max_bytes)
135
+ limit = [max_bytes, text.bytesize].min
136
+ if limit.positive? && limit < text.bytesize && continuation?(text.getbyte(limit))
137
+ limit = align_start(text, limit)
138
+ end
139
+ cut = text.byteslice(0, limit)
140
+ return cut if cut.valid_encoding?
141
+
142
+ 3.times do
143
+ break if cut.empty?
144
+
145
+ cut = cut.byteslice(0, cut.bytesize - 1)
146
+ return cut if cut.valid_encoding?
147
+ end
148
+
149
+ text.byteslice(0, [max_bytes, text.bytesize].min).scrub('?')
150
+ end
151
+ end
152
+ end
@@ -0,0 +1,5 @@
1
+ # frozen_string_literal: true
2
+
3
+ module FastMbChars
4
+ VERSION = '0.1.0'
5
+ end
@@ -0,0 +1,24 @@
1
+ # frozen_string_literal: true
2
+
3
+ require 'fast_mb_chars/version'
4
+ require 'fast_mb_chars/limiter'
5
+ require 'fast_mb_chars/chars'
6
+
7
+ module FastMbChars
8
+ # Prepends a fresh module each time the method was taken over by someone
9
+ # else. Prepending the same module twice does not move it back to the front.
10
+ def self.install!
11
+ if String.method_defined?(:mb_chars) && String.instance_method(:mb_chars).owner == @installed
12
+ return
13
+ end
14
+
15
+ @installed = Module.new do
16
+ def mb_chars
17
+ FastMbChars::Chars.new(self)
18
+ end
19
+ end
20
+ String.prepend(@installed)
21
+ end
22
+
23
+ install!
24
+ end
metadata ADDED
@@ -0,0 +1,66 @@
1
+ --- !ruby/object:Gem::Specification
2
+ name: fast_mb_chars
3
+ version: !ruby/object:Gem::Version
4
+ version: 0.1.0
5
+ platform: ruby
6
+ authors:
7
+ - SYNORYXEL
8
+ autorequire:
9
+ bindir: bin
10
+ cert_chain: []
11
+ date: 2026-10-07 00:00:00.000000000 Z
12
+ dependencies: []
13
+ description: |
14
+ Faster String#mb_chars. limit(bytes) cuts on a character boundary without
15
+ walking the whole string and without raising on a broken byte. A 1 MB cut
16
+ is about 2,000 times faster than the ActiveSupport grapheme scan, and a
17
+ 7 MB cut is about 3,000 times faster.
18
+
19
+ downcase, upcase, titleize, reverse, split, slice!, compose, decompose,
20
+ tidy_bytes, and the other string methods are forwarded through method_missing.
21
+ They still return mb_chars, so existing chains keep working.
22
+ email:
23
+ - mike.pvlv@icloud.com
24
+ executables: []
25
+ extensions: []
26
+ extra_rdoc_files: []
27
+ files:
28
+ - CHANGELOG.md
29
+ - LICENSE
30
+ - README.md
31
+ - benchmark/compare.rb
32
+ - docs/limit-cut.svg
33
+ - lib/fast_mb_chars.rb
34
+ - lib/fast_mb_chars/chars.rb
35
+ - lib/fast_mb_chars/limiter.rb
36
+ - lib/fast_mb_chars/version.rb
37
+ homepage: https://github.com/SYNORYXEL/fast_mb_chars
38
+ licenses:
39
+ - MIT
40
+ metadata:
41
+ homepage_uri: https://github.com/SYNORYXEL/fast_mb_chars
42
+ source_code_uri: https://github.com/SYNORYXEL/fast_mb_chars
43
+ changelog_uri: https://github.com/SYNORYXEL/fast_mb_chars/blob/main/CHANGELOG.md
44
+ bug_tracker_uri: https://github.com/SYNORYXEL/fast_mb_chars/issues
45
+ allowed_push_host: https://rubygems.org
46
+ rubygems_mfa_required: 'true'
47
+ post_install_message:
48
+ rdoc_options: []
49
+ require_paths:
50
+ - lib
51
+ required_ruby_version: !ruby/object:Gem::Requirement
52
+ requirements:
53
+ - - ">="
54
+ - !ruby/object:Gem::Version
55
+ version: 2.7.0
56
+ required_rubygems_version: !ruby/object:Gem::Requirement
57
+ requirements:
58
+ - - ">="
59
+ - !ruby/object:Gem::Version
60
+ version: '0'
61
+ requirements: []
62
+ rubygems_version: 3.5.22
63
+ signing_key:
64
+ specification_version: 4
65
+ summary: Faster Ruby String#mb_chars. Byte limit that keeps UTF-8 characters whole
66
+ test_files: []