carray-jit 0.1.2 → 0.1.3
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +771 -3
- data/README.md +7 -6
- data/carray-jit.gemspec +1 -3
- data/docs/00_Introduction.md +4 -3
- data/docs/01_GettingStarted.md +1 -1
- data/docs/02_KernelShapes.md +93 -14
- data/docs/03_SupportedFeatures.md +582 -26
- data/docs/04_Compiling.md +33 -6
- data/docs/05_DesignNotes.md +3 -3
- data/docs/06_Cheatsheet.md +198 -5
- data/docs/07_StepByStep.ja.md +534 -0
- data/docs/07_StepByStep.md +535 -0
- data/examples/README.md +12 -0
- data/examples/applications/alarm.rb +121 -0
- data/examples/applications/collatz.rb +105 -0
- data/examples/applications/cubic_spline.rb +331 -0
- data/examples/applications/dithering.rb +144 -0
- data/examples/applications/group_stats.rb +115 -0
- data/examples/applications/lookup.rb +126 -0
- data/examples/applications/median_filter.rb +153 -0
- data/examples/applications/parcel_ascent.rb +220 -0
- data/examples/applications/point_in_polygon.rb +111 -0
- data/examples/applications/random_walk.rb +98 -0
- data/examples/applications/van_der_pol.rb +186 -0
- data/examples/applications/wet_bulb.rb +140 -0
- data/examples/features/10_complex.rb +14 -4
- data/examples/features/15_loops.rb +7 -1
- data/lib/carray/jit/access.rb +14 -0
- data/lib/carray/jit/analyzer.rb +2077 -136
- data/lib/carray/jit/block_reader.rb +37 -6
- data/lib/carray/jit/c_function.rb +613 -76
- data/lib/carray/jit/c_generator.rb +1595 -156
- data/lib/carray/jit/call.rb +68 -0
- data/lib/carray/jit/compiler.rb +75 -11
- data/lib/carray/jit/kernel.rb +369 -32
- data/lib/carray/jit/node.rb +359 -9
- data/lib/carray/jit/sorting_networks.rb +182 -0
- data/lib/carray/jit/type_assignment.rb +371 -34
- data/lib/carray/jit/version.rb +1 -1
- data/lib/carray/jit.rb +560 -64
- metadata +22 -8
- data/ext/carray_jit_access/carray_jit_access.c +0 -460
- data/ext/carray_jit_access/extconf.rb +0 -8
data/CHANGELOG.md
CHANGED
|
@@ -37,6 +37,776 @@ version you have and a newer one.
|
|
|
37
37
|
a release needs is said in the gemspec, and an entry says so only
|
|
38
38
|
when the answer changes. -->
|
|
39
39
|
|
|
40
|
+
## 0.1.3
|
|
41
|
+
|
|
42
|
+
- New: `CArray::JIT.call_provider` is asked at a `jit_call` site before
|
|
43
|
+
anything is compiled, and answering `nil` means "not mine, compile it". It
|
|
44
|
+
is handed the prototype, the block -- whose binding says which method and
|
|
45
|
+
which module the site is in -- and the names the declaration gave, and
|
|
46
|
+
answers anything that responds to `call`. The one client is carray-aot,
|
|
47
|
+
which builds these same sites into a shared object ahead of the program so
|
|
48
|
+
that a machine running that gem reaches no compiler; it is the position
|
|
49
|
+
`jit_extern` puts a function from a library in, said about a call rather
|
|
50
|
+
than about a name.
|
|
51
|
+
|
|
52
|
+
- New: `CArray.jit_call(prototype) { ... }` compiles the block as a C
|
|
53
|
+
function and calls it where it stands, with the locals around it. The
|
|
54
|
+
declaration's parameter names are the join and do the work twice: they are
|
|
55
|
+
the body's parameters, so the block declares none, and they name the locals
|
|
56
|
+
the call reads -- `CArray.jit_call("void (*)(double *out, const double
|
|
57
|
+
*values, size_t n, size_t window)")` inside a method holding `out`,
|
|
58
|
+
`values`, `n` and `window`. In C a parameter's name in a prototype is
|
|
59
|
+
decoration; here it is the whole binding, and a declared name with no local
|
|
60
|
+
behind it is refused at the call rather than read as nil inside it. What it
|
|
61
|
+
saves over `jit_function` and a `call` is the argument list, which restates
|
|
62
|
+
the declaration in an order nothing checks; what it costs is reading those
|
|
63
|
+
locals through the block's binding, about 0.3 microseconds against a call
|
|
64
|
+
that costs several. Compiled once per call site, and `clear_registry`
|
|
65
|
+
reaches those as it reaches the rest.
|
|
66
|
+
|
|
67
|
+
- New: `CArray::JIT::CFunction#c_source_as(symbol)` hands the compiled C over
|
|
68
|
+
under a symbol the caller picked, for a caller writing it into a file of
|
|
69
|
+
its own rather than letting this compile it -- the symbol this compiler
|
|
70
|
+
writes carries a digest of the body, which is right for an object in a
|
|
71
|
+
cache and wrong for one committed to a repository. A name C cannot spell
|
|
72
|
+
and a name in this compiler's own `carray_jit_` namespace are refused; a
|
|
73
|
+
name the C library already has is the caller's to avoid, since the prefix
|
|
74
|
+
that closes that hazard is what is being replaced. A function bound with
|
|
75
|
+
`jit_extern` has no C of its own and says so.
|
|
76
|
+
|
|
77
|
+
- New: a parallel assignment is in the subset -- `a, b = b, a`, and the
|
|
78
|
+
`a, b = b, a + b` a recurrence advances by. Every value on the right is
|
|
79
|
+
settled before anything on the left is written, as in Ruby, so the
|
|
80
|
+
temporary that spelling saves no longer has to be written by hand. A cell
|
|
81
|
+
is a target as much as a name is (`a[i], a[j] = a[j], a[i]`), each value
|
|
82
|
+
keeps its own type and carries its mask. The readings that take a single
|
|
83
|
+
value apart are refused by name: `a, b = f(x)`, `a, b = [1, 2]`,
|
|
84
|
+
`a, *rest = ...`, `a, (b, c) = ...`, and a count that does not match.
|
|
85
|
+
|
|
86
|
+
- New: a local array may be larger than a stack frame should hold, and its
|
|
87
|
+
shape may be written over an integer the block captured --
|
|
88
|
+
`CArray.double(n)`, which was refused. Either way the kernel allocates the
|
|
89
|
+
array once at its entry and frees it at its exit, so a constructor written
|
|
90
|
+
inside the cell loop is still one allocation; 4 KiB for one array and
|
|
91
|
+
16 KiB for one kernel's arrays together now say where an array lives rather
|
|
92
|
+
than whether it is allowed, and nothing is refused for its size. A kernel
|
|
93
|
+
whose arrays all fit in the frame emits the C it emitted before. Where the
|
|
94
|
+
length is one the kernel works out, every subscript on that axis is checked
|
|
95
|
+
where the cell is reached rather than as the block is read, and a C
|
|
96
|
+
function takes the array only through a pointer that declares no length
|
|
97
|
+
(`const double *v`, not `const double v[3]`). The same block at two lengths
|
|
98
|
+
is one compiled kernel: the length travels as an argument and is not in the
|
|
99
|
+
C. A shape that comes to zero or less raises `ArgumentError` when the
|
|
100
|
+
kernel runs, and an allocation the system refuses raises `NoMemoryError`,
|
|
101
|
+
both naming the array and what its shape came to. A `jit_function` body
|
|
102
|
+
allocates nothing -- it is called once per cell -- and refuses such an
|
|
103
|
+
array, naming the pointer parameter to take it through instead.
|
|
104
|
+
|
|
105
|
+
- New: a kernel that carries masks takes a local array, which it refused
|
|
106
|
+
before. Every local array of such a kernel is declared with a shadow of one
|
|
107
|
+
byte a cell beside its cells, and a cell carries a mask the way a plain
|
|
108
|
+
local does: what the expression written into it carried. So a window copied
|
|
109
|
+
into a workspace keeps its holes. `w[k] = UNDEF` marks a cell and
|
|
110
|
+
`w[k] == UNDEF` asks about one; the zeroed spellings clear the shadow with
|
|
111
|
+
the cells, so a cell starts every pass present, while `CArray.empty` leaves
|
|
112
|
+
both unspecified. The shadow counts against the 4 KiB an array is held to
|
|
113
|
+
and the 16 KiB a kernel is, so an array that fits by its cells alone may
|
|
114
|
+
not fit once it carries masks. Two things still refuse such an array:
|
|
115
|
+
`sum`, `min`, `max` and `sort`, because what they should do with a missing
|
|
116
|
+
cell is not decided, and a C function, because a mask travels in no C
|
|
117
|
+
declaration -- both say so where they used to be refused for having no mask
|
|
118
|
+
at all. A kernel that carries no masks emits what it emitted before, and a
|
|
119
|
+
compiled function's body carries none at all.
|
|
120
|
+
|
|
121
|
+
- New: an inner loop's range may be written over another index --
|
|
122
|
+
`3.times { |p| (p+1...3).each { |r| ... } }`, the shape a triangular loop
|
|
123
|
+
takes -- where that index is one of the loops around it. Forward
|
|
124
|
+
elimination and Neville's interpolation are the two this was wanted for,
|
|
125
|
+
and both now read as they do on paper rather than as a full loop with an
|
|
126
|
+
`if` inside it. The C was always emitted; what stood in the way was the
|
|
127
|
+
call, where each index's range is worked out so that a subscript can be
|
|
128
|
+
held to its array. A range over another index has no single pair of
|
|
129
|
+
numbers to be, so it is read as an interval, at its widest: every pass the
|
|
130
|
+
loop could take and sometimes more. Where that refuses a reach the loop
|
|
131
|
+
never makes, the message says which index the range was written over and
|
|
132
|
+
that it was read at its widest. A range over a local variable, a sibling
|
|
133
|
+
loop's index or a deeper one is still refused.
|
|
134
|
+
|
|
135
|
+
- New: a local array may have more than one axis --
|
|
136
|
+
`CArray.double(3, 4)`, `CArray.new(:float64, [2, 3, 4])` -- in any of the
|
|
137
|
+
five places one can be made. The shape is written out as it always was, so
|
|
138
|
+
the strides are constants: `m[r, c]` is `m[(r) * 4 + (c)]`, one subscript
|
|
139
|
+
per axis, and each axis is checked against its own extent. That is what
|
|
140
|
+
flattening by hand gives up -- `m[r * 4 + c]` with a column of 4 reads a
|
|
141
|
+
cell of the next row and says nothing, where `m[r, c]` is refused by name
|
|
142
|
+
and a computed column raises. The stack limits count cells rather than
|
|
143
|
+
axes, so `CArray.double(32, 32)` is 8 KiB and past the 4 KiB an array is
|
|
144
|
+
held to. Handed to a C function the array goes as the flat run of cells it
|
|
145
|
+
is, row after row, so a `double[3][4]` reaches `const double a[12]` and the
|
|
146
|
+
length is matched over every cell. The four intrinsics still take one axis.
|
|
147
|
+
|
|
148
|
+
- New: a `CArray.jit_function` body may make a local array and use
|
|
149
|
+
`sum`/`min`/`max`/`sort` over one, which completes the four entry points.
|
|
150
|
+
A body closes over nothing, so this is the only place scratch space could
|
|
151
|
+
come from other than an extra parameter -- and a signature settled
|
|
152
|
+
elsewhere, a callback's, has no room for one. The array is declared at the
|
|
153
|
+
head of the function, a recursive body gets one per call, and a body hands
|
|
154
|
+
one to another compiled function under the same rules a kernel does. The
|
|
155
|
+
helpers an intrinsic needs travel with the body: carried in its own file
|
|
156
|
+
when it is compiled alone, and merged into the kernel's preamble when it is
|
|
157
|
+
pasted, one helper per element type and length however many bodies want it.
|
|
158
|
+
The stack limits stay per function -- 4 KiB an array, 16 KiB a body -- and a
|
|
159
|
+
chain of pasted functions is not counted, so a deep chain stands as many
|
|
160
|
+
frames as it has; that is the bargain a deep recursion already takes. Still
|
|
161
|
+
excluded: more than one axis, and a contraction.
|
|
162
|
+
|
|
163
|
+
- New: a local array may be handed to a C function -- one from
|
|
164
|
+
`CArray.jit_function` or `CArray.jit_extern` -- wherever the declaration
|
|
165
|
+
takes a pointer, from any of the four entry points that make one. What the
|
|
166
|
+
declaration says is matched as the block is read rather than at the call,
|
|
167
|
+
both the element type and the length being written in the block: an exact
|
|
168
|
+
element type, and at least as many cells as a sized declarator asks for.
|
|
169
|
+
A parameter that is not `const` may be written through, and the line after
|
|
170
|
+
the call reads what the callee left; the zeroed constructors still clear at
|
|
171
|
+
the line, so nothing carries into the next cell. Passing one array to two
|
|
172
|
+
parameters is fine, the declarations carrying no `restrict`. Note that a
|
|
173
|
+
borrowed function which keeps the pointer past the call is left pointing at
|
|
174
|
+
a stack frame that has gone, and that a declaration carrying no length --
|
|
175
|
+
`const double *x` -- gives nothing to check against. Still excluded when this landed: a `jit_function` body,
|
|
176
|
+
more than one axis, and a contraction; the first two arrived later in this
|
|
177
|
+
release, and a contraction is still refused.
|
|
178
|
+
|
|
179
|
+
- New: a local array, and the four intrinsics over one, may be written in a
|
|
180
|
+
`CArray.jit_each`, `CArray.jit_map` or `CArray.jit_stencil` block as well as
|
|
181
|
+
in `CArray.jit_for`. These are the spellings that wanted one: a block with
|
|
182
|
+
no index cannot pick a row of a captured array, so a `jit_stencil` median
|
|
183
|
+
filter had nowhere to put its window -- it is now nine doubles on the cell's
|
|
184
|
+
stack and `sort(w)`. `border: :mask` takes one, the frame being marked
|
|
185
|
+
before the loop runs rather than carried through it. A name the block makes
|
|
186
|
+
an array under is refused where the block also closes over an array of that
|
|
187
|
+
name, since in these spellings an assignment writes that array's cell;
|
|
188
|
+
rename one of the two. A `jit_map` block may not end by making an array, its
|
|
189
|
+
value having to fit in a cell. Still excluded when this landed: a `jit_function` body, a
|
|
190
|
+
contraction, more than one axis, passing one to a C function, and a kernel
|
|
191
|
+
that carries masks; all but the contraction arrived later in this release.
|
|
192
|
+
|
|
193
|
+
- New: `sum(w)`, `min(w)`, `max(w)` and `sort(w)` inside a `CArray.jit_for`
|
|
194
|
+
block, over a local array of one axis. They are bare calls, the compiler's
|
|
195
|
+
own names, rather than methods on the array -- `w.sum` is refused, and
|
|
196
|
+
`sum = 0.0` beside `sum(w)` is still a local. `sum` accumulates in the
|
|
197
|
+
element's computation type in index order; `min` and `max` skip a NaN
|
|
198
|
+
wherever it stands, and answer `NaN` for an array of nothing but NaN,
|
|
199
|
+
which is what `CArray#min` and `#max` answer (they answered the type's
|
|
200
|
+
limit until a later change in this release). `sort`
|
|
201
|
+
is a statement and orders the cells ascending with every NaN after every
|
|
202
|
+
number, as `CArray#sort` does; the relative order of `-0.0` and `0.0` is
|
|
203
|
+
not promised. Up to 16 cells it emits a comparator network with no branch
|
|
204
|
+
in it, above that an insertion sort. Excluded for now: a captured array
|
|
205
|
+
(a whole-array reduction is `CArray#sum`), more than one axis, two
|
|
206
|
+
arguments, boolean for all four, Complex for `min` / `max` / `sort`, and
|
|
207
|
+
the other entry points.
|
|
208
|
+
|
|
209
|
+
- New: a `CArray` made inside a `CArray.jit_for` block is a C array on the
|
|
210
|
+
block's stack -- `w = CArray.double(9)`, `CArray.new(:int64, [256])` or
|
|
211
|
+
`CArray.empty(:float64, [9])`, read and written at a subscript. One axis,
|
|
212
|
+
with the length written out as an integer or integers joined by `+`, `-` and
|
|
213
|
+
`*`; the zeroed spellings are cleared each time the line runs, as Ruby makes
|
|
214
|
+
a fresh array there. A subscript is checked as the block is read where the
|
|
215
|
+
loop's range says it can be, and at the access otherwise, raising
|
|
216
|
+
`IndexError`. One array is held to 4 KiB of stack and one kernel's to 16 KiB.
|
|
217
|
+
Excluded for now: more than one axis, passing one to a C function, the other
|
|
218
|
+
entry points (`jit_each`, `jit_map`, `jit_stencil`, `jit_function`), a kernel
|
|
219
|
+
that carries masks, and CArray's Numo/NumPy spellings (`CArray.zeros`,
|
|
220
|
+
`CArray::Int64.empty`), which are refused with the carray spelling named.
|
|
221
|
+
Note `CArray.float` is float32 and `CArray.complex` is cmplx64.
|
|
222
|
+
|
|
223
|
+
- New: `CArray#jit_init` fills an array from a formula over its indices, with
|
|
224
|
+
the block compiled: `CArray.int32(1000, 1000).jit_init { |i, j| (i + j) % 2 }`
|
|
225
|
+
is what the constructor block `CArray.int32(n, n) { |i, j| ... }` says,
|
|
226
|
+
without the Ruby call per cell. One parameter per axis, the block's value is
|
|
227
|
+
the cell, and the receiver comes back. It is the only entry point here that
|
|
228
|
+
is an instance method, because the array being written is the receiver: the
|
|
229
|
+
extents are its shape, so neither the index space nor the target is said
|
|
230
|
+
twice, where `CArray.jit_for(n, n) { |i, j| z[i, j] = ... }` says both. A
|
|
231
|
+
block outside the compilable subset raises rather than running the slow
|
|
232
|
+
loop, as every other entry point here does. Reach for it where the formula
|
|
233
|
+
will not go through whole-array arithmetic -- where it will, that needs no
|
|
234
|
+
compiler and is faster still.
|
|
235
|
+
|
|
236
|
+
- New: a kernel can draw random numbers, from a `CArray::Rng`. `rand =
|
|
237
|
+
CArray::Rng.new(seed: 4)` and then `rand.random` in a `jit_for`, `jit_each`
|
|
238
|
+
or `jit_map` block draws one double in `[0.0, 1.0)` per cell, at about
|
|
239
|
+
0.8 ns; `random(rng: rand)` is the same draw spelled as
|
|
240
|
+
`CArray#random!(rng:)` spells it; `rand.randomn` is a standard normal, at
|
|
241
|
+
about 10 ns, and `randomn(rng: rand)` the same; and `rand.bits` is the raw
|
|
242
|
+
word a draw came from. All of them read one generator, so mixing them in a
|
|
243
|
+
kernel walks one sequence. The generator is CArray's and there is no entry point here: a
|
|
244
|
+
block closing over one is what a kernel needs. Each has its own state, so
|
|
245
|
+
two in one kernel are two sequences, and the state survives the call --
|
|
246
|
+
a second kernel carries on rather than starting again. `seed:` is data
|
|
247
|
+
rather than code, so every seed shares one compiled kernel. A sequence can
|
|
248
|
+
begin with `a.random!(rng: rand)` and continue in a kernel: those are the
|
|
249
|
+
same numbers one `random!` over both arrays would have laid down, because
|
|
250
|
+
CArray hands out the generator's C and this gem pastes it. Needs CArray
|
|
251
|
+
3.0.2 or newer, which is where `CArray::Rng` arrived; an older one is
|
|
252
|
+
refused where the block closes over the generator, and leaves every kernel
|
|
253
|
+
that draws nothing alone. A draw is refused inside `jit_function`, which
|
|
254
|
+
has nowhere to keep a state; take `int64_t state[4]` as a parameter there
|
|
255
|
+
and pass `rand.state`. Which draw lands in which cell is still the loop's
|
|
256
|
+
order, so where that has to be settled, fill an array with
|
|
257
|
+
`CArray#random!` before the call.
|
|
258
|
+
|
|
259
|
+
- New: `CFunction#watching`, and `#clear_error` / `#report_error` beside it,
|
|
260
|
+
for the window in which a compiled function's address is lent to a C
|
|
261
|
+
library. `#call` answers for one call; a library given `#pointer` calls as
|
|
262
|
+
often as it likes, and `f.watching { ... }` is what puts the flag down
|
|
263
|
+
before that and raises what the body reported after it. A failure inside
|
|
264
|
+
the block outranks the library's own complaint about the stand-in it was
|
|
265
|
+
handed. Windows nest, and a `#call` made inside one leaves it armed.
|
|
266
|
+
|
|
267
|
+
- New: `Math.gamma`, which is not `tgamma` and is not lowered to one: Ruby's
|
|
268
|
+
answer is tgamma with a table of exact values in front of it for a whole
|
|
269
|
+
number up to 23, and a `Math::DomainError` where tgamma answers a negative
|
|
270
|
+
whole number or negative infinity with a NaN. Both are reproduced -- the
|
|
271
|
+
table is written into the generated C, filled from the Ruby that compiled
|
|
272
|
+
it, as `Math::PI` is emitted as the double Ruby would have used, and the
|
|
273
|
+
error carries Ruby's class and Ruby's words. A cell with no value in it
|
|
274
|
+
does not raise. Measured over the whole numbers to 26, the halves, the
|
|
275
|
+
infinities and the overflow, a kernel and the Ruby loop agree in every bit.
|
|
276
|
+
|
|
277
|
+
- New: `Math.erf` and `Math.erfc`, which are 1:1 with math.h's `erf` and
|
|
278
|
+
`erfc` -- Ruby calls those very functions, so a kernel and the Ruby loop
|
|
279
|
+
agree to the bit, infinities included. A float32 cell is worked on narrow
|
|
280
|
+
and so gets `erff`, as it gets `sinf`. There is no postfix `x.erf`:
|
|
281
|
+
`CArray::CoreExtensions` does not provide one, and this compiles the
|
|
282
|
+
refinement's names rather than inventing them.
|
|
283
|
+
|
|
284
|
+
- New: `x.clamp(low, high)`, which answers the value or whichever bound it
|
|
285
|
+
ran past. The value and both bounds have to be one class: Ruby hands back
|
|
286
|
+
the receiver in one branch and a bound in the other, so `1.clamp(0.0, 3.0)`
|
|
287
|
+
is an Integer and `5.clamp(0.0, 3.0)` a Float, and no type settled before
|
|
288
|
+
the loop runs is both -- the refusal says which way round it is and what to
|
|
289
|
+
write. Two widths of one class are not that case, so a float32 cell keeps
|
|
290
|
+
its width. What Ruby raises `ArgumentError` for is raised here too: bounds
|
|
291
|
+
the wrong way round, and a NaN that cannot be ordered. The class is Ruby's
|
|
292
|
+
and so are the words for the first; for the NaN the message names the
|
|
293
|
+
reason rather than the value, since what comes back from a kernel is a code
|
|
294
|
+
and not a number. A cell with no value in it does not raise. The range
|
|
295
|
+
form, `x.clamp(0.0..1.0)`, is not in the subset -- two bounds are two
|
|
296
|
+
bounds.
|
|
297
|
+
|
|
298
|
+
- New: a compiled function may take and return a C99 complex --
|
|
299
|
+
`CArray.jit_function("double _Complex step(double _Complex z)") { |z| z * z
|
|
300
|
+
+ Complex(0.0, 1.0) }`. `double _Complex`, `float _Complex` and
|
|
301
|
+
`<complex.h>`'s `double complex` all read, by value or as a pointer, where
|
|
302
|
+
`double _Complex v[]` takes a `cmplx128` array as `double v[]` takes a
|
|
303
|
+
`float64` one. A kernel calls such a function as C calls it. A call from
|
|
304
|
+
Ruby goes through a second entry point compiled beside the body, since
|
|
305
|
+
Fiddle has no type to carry a complex by value in; it calls the body rather
|
|
306
|
+
than repeating it, so `f.call` and `f.block.call` stay the same body run
|
|
307
|
+
two ways. A function bound with `jit_extern` and declared with a complex is
|
|
308
|
+
callable from a kernel but not from Ruby -- there is no source here to
|
|
309
|
+
compile an entry point beside -- and says so.
|
|
310
|
+
|
|
311
|
+
- New: a `jit_function` body may take an unsigned 64-bit value by parameter --
|
|
312
|
+
`CArray.jit_function("size_t stride(size_t n, size_t width)") { |n, w| n * w }`
|
|
313
|
+
-- where before only a pointer to one could be taken. The value arrives
|
|
314
|
+
whole above 2^63 and the arithmetic on it is unsigned, wrapping at the width
|
|
315
|
+
as CArray's own `uint64` operators wrap. A kernel's captured scalars are
|
|
316
|
+
unchanged: those travel in the kernel's own buffers, which carry doubles,
|
|
317
|
+
int64s and complexes.
|
|
318
|
+
|
|
319
|
+
- New: `x.nan?` and `x.finite?`, which compile to C's `isnan` and `isfinite`
|
|
320
|
+
and answer what Ruby answers. `nan?` is a Float's question: an Integer and
|
|
321
|
+
a Complex have no method by that name and raise `NoMethodError` in Ruby, so
|
|
322
|
+
a kernel refuses both rather than answering false. `finite?` answers for an
|
|
323
|
+
Integer (true, whatever it holds) and for a Complex (both parts finite, as
|
|
324
|
+
Ruby asks it) as well as for a Float. `infinite?` is refused: Ruby answers
|
|
325
|
+
it with nil, 1 or -1 rather than true or false, and a kernel has no nil to
|
|
326
|
+
answer with -- the message names `x.abs == Float::INFINITY`, or
|
|
327
|
+
`x == Float::INFINITY` where the sign is the question.
|
|
328
|
+
|
|
329
|
+
- New: the operator assignments -- `+= -= *= /= %= **= &= |= ^= <<= >>=` --
|
|
330
|
+
on a local, on a cell (`out[i] += e`, `work[i, k] += e`, a scatter such as
|
|
331
|
+
`counts[bin[i]] += 1`), on a CScalar and through a `jit_function`'s pointer
|
|
332
|
+
parameter. Each is read as the assignment it stands for, `x = x + e`, so
|
|
333
|
+
the type rules, the mask propagation and the fold that splits an
|
|
334
|
+
accumulator into partial sums are the ones already there: a reduction
|
|
335
|
+
written `total += values[j]` is still split. `||=` and `&&=` are refused,
|
|
336
|
+
being about whether a value is nil or false rather than about arithmetic.
|
|
337
|
+
|
|
338
|
+
- New: an inner loop counts by a stride, written the way an extent writes
|
|
339
|
+
one: `(n-1).step(0, -1) { |k| ... }` is a downward sweep and
|
|
340
|
+
`(0...n).step(2) { |k| ... }` a stride of two. `step` includes the index it
|
|
341
|
+
is given, as Ruby's does, where a `...` range excludes it, and the stride
|
|
342
|
+
is a literal because it is what says which way the loop runs. This is the
|
|
343
|
+
other half of a row of workspace -- filling one and walking back down it no
|
|
344
|
+
longer needs `k = width - 1 - t`, which made the position a value the
|
|
345
|
+
kernel worked out and so put a bounds test on every cell of the sweep: 0.41
|
|
346
|
+
ns/cell against 0.24 for the same sweep written with `step`. An accumulator
|
|
347
|
+
is split into partial sums only for a loop counting by one; a stride keeps
|
|
348
|
+
the serial chain, and so keeps Ruby's order. `downto`, `upto` and
|
|
349
|
+
`reverse_each` are refused by name, naming `step` as what to write.
|
|
350
|
+
|
|
351
|
+
- New: an inner loop's index may address a write, so a cell can be given a
|
|
352
|
+
row of workspace -- `(0...width).each { |k| work[i, k] = ... }` fills it,
|
|
353
|
+
and `work[i, k]` reads it back inside the same cell. That is what an
|
|
354
|
+
algorithm needing a few numbers per cell is written with: a small dense
|
|
355
|
+
solve, a tableau, a sweep and the pass back down it, in one kernel rather
|
|
356
|
+
than in several that each pay a call. The cell is bounds-checked before the
|
|
357
|
+
kernel runs, as one addressed by `i` is. Inside the row the body may do as
|
|
358
|
+
it likes -- sort it, walk it backwards, write at a position it works out --
|
|
359
|
+
which is what a median filter needs and what no extra axis can say. Reading
|
|
360
|
+
an array the kernel writes through an inner index is still refused where
|
|
361
|
+
the read leaves the cell the outer indices picked: a read carrying an inner
|
|
362
|
+
index addresses every axis some write walks with an outer index with that
|
|
363
|
+
same index, at whatever offset, or it reaches cells another outer iteration
|
|
364
|
+
owns. Where the sweep can be said as an extra axis instead --
|
|
365
|
+
`jit_for(rows, 1...width)` -- that remains the faster form once the rows
|
|
366
|
+
are long.
|
|
367
|
+
|
|
368
|
+
- New: a function compiled with `CArray.jit_function` may call another one
|
|
369
|
+
compiled with `CArray.jit_function`, by the name the block reaches it by --
|
|
370
|
+
`hypot = CArray.jit_function("double (*)(double, double)") { |a, b|
|
|
371
|
+
root.call(a * a + b * b) }`. The called body is pasted into the caller's C
|
|
372
|
+
and reached by symbol, so what comes back is still one self-contained
|
|
373
|
+
object with one address, and a chain of any depth arrives together with the
|
|
374
|
+
messages its bodies raise. Everything else a body closes over is refused as
|
|
375
|
+
before -- a number, an array, and a function bound with `jit_extern`, which
|
|
376
|
+
is only an address and has nowhere in a compiled object to live. The name
|
|
377
|
+
may be a constant as well as a local, which is what lets a method reach
|
|
378
|
+
one: `def` closes over nothing.
|
|
379
|
+
|
|
380
|
+
- New: `CArray::JIT.contraction_of` returns the number that multiplies the
|
|
381
|
+
product as `:scale`, which is 1 where there is none, so
|
|
382
|
+
`a[i,k] * b[k,j] * 2.0` comes back as its two terms and 2.0 rather than as
|
|
383
|
+
nil. A number is not a term -- it has no indices and no cell -- but it is
|
|
384
|
+
not a reason to give up on the terms either, and a caller that takes them
|
|
385
|
+
apart puts it back. Written out or closed over is the same number; anything
|
|
386
|
+
a name holds that is not a Numeric is still nil. This shipped in 0.1.2,
|
|
387
|
+
whose entry describes the same two methods without mentioning it.
|
|
388
|
+
|
|
389
|
+
- Change: this gem no longer builds a C extension, so installing it needs no
|
|
390
|
+
compiler for itself -- one is still needed at run time to compile a kernel.
|
|
391
|
+
The addressing it used to carry is `CArray::AddressBasis`, in carray, which
|
|
392
|
+
is where the knowledge about carray's views belonged; `CArray::JIT::Access`
|
|
393
|
+
remains as an alias for this release. carray 3.0.2 or newer is required,
|
|
394
|
+
which the gemspec already asked for.
|
|
395
|
+
|
|
396
|
+
- Change: `CArray.jit_for`, `CArray.jit_each` and `CArray.jit_map` are this
|
|
397
|
+
gem's alone. From CArray 3.0.2 they are not defined until
|
|
398
|
+
`require "carray/jit"` has run, so a program that calls one without it gets
|
|
399
|
+
`NoMethodError` rather than CArray's `NotImplementedError`. With the gem
|
|
400
|
+
required nothing changes, against any CArray this gem accepts.
|
|
401
|
+
|
|
402
|
+
- Change: `CArray::JIT::CFunction#name` answers the name the declaration gave
|
|
403
|
+
-- `:square` for `CArray.jit_function("double square(double x)")` -- and
|
|
404
|
+
the new `#symbol` answers the name in the object, which for a compiled
|
|
405
|
+
function carries a digest of the body. It used to answer whichever of the
|
|
406
|
+
two its constructor was handed: the declared name for a function bound with
|
|
407
|
+
`jit_extern`, the symbol for one compiled from a block. Code reading
|
|
408
|
+
`#name` to reach or paste a function wants `#symbol`; code putting it in a
|
|
409
|
+
message wants `#name`, which is `nil` where the declaration named no
|
|
410
|
+
function. `#to_s` now prints the declaration rather than the symbol.
|
|
411
|
+
|
|
412
|
+
- Change: a compiled function whose body has failed does no more work until
|
|
413
|
+
its flag is put down again -- it returns 0 without running, and one
|
|
414
|
+
declared `void` leaves its out-parameters alone. Before, it reported the
|
|
415
|
+
first failure and then answered normally, so a library that kept calling
|
|
416
|
+
could converge on values it was no longer entitled to. Nothing changes for
|
|
417
|
+
a caller using `#call`, which puts the flag down for each call; a caller
|
|
418
|
+
holding `#pointer` opens a window with `#watching` or `#clear_error`. A
|
|
419
|
+
kernel is unaffected: it is handed its own error slot and never reaches
|
|
420
|
+
this flag.
|
|
421
|
+
|
|
422
|
+
- Change: naming a contraction's axes now replaces the convention rather than
|
|
423
|
+
adding a clause to it. `CArray.jit_contract(:i, :j) { ... }` names all of
|
|
424
|
+
the result's axes, so every index left out of the list is summed at however
|
|
425
|
+
few positions it sits -- where before, one sitting at a single position was
|
|
426
|
+
refused as a free index with nowhere to go. So `jit_contract(:i) { |k|
|
|
427
|
+
a[i,k] }` is the row sums, and `contract_terms(terms, free: [])` over a
|
|
428
|
+
product is its total; both were refused. (`a.sum(axis: 1)` remains the
|
|
429
|
+
faster way to write a reduction, being one.) With nothing named the
|
|
430
|
+
convention is unchanged: a repetition is a sum and a single position is
|
|
431
|
+
free. What this makes exact is the correspondence with einsum's two modes,
|
|
432
|
+
the argument list being the arrow's right-hand side; `"ik->i"` and
|
|
433
|
+
`"ik,kj->"` can now be said.
|
|
434
|
+
|
|
435
|
+
Where the block assigns into an array of yours, an axis on the left-hand
|
|
436
|
+
side that the list leaves out is still refused -- it is the list falling
|
|
437
|
+
short of the result rather than an index to sum, and the left-hand side is
|
|
438
|
+
the one place that can be seen. The message says so in those terms now.
|
|
439
|
+
|
|
440
|
+
- Change: `min(w)` and `max(w)` over a floating local array answer `NaN` when
|
|
441
|
+
every cell is `NaN`, where they answered `Infinity` and `-Infinity`. This
|
|
442
|
+
follows CArray 3.0.2, which made the same change to its own `min` and `max`;
|
|
443
|
+
the gemspec already asks for that version. An array holding at least one
|
|
444
|
+
number answers as before, a `NaN` still losing to any number, and an integer
|
|
445
|
+
local array is unchanged.
|
|
446
|
+
|
|
447
|
+
- Change: `CArray.jit_contract` refuses an index that stands inside a
|
|
448
|
+
subscript the kernel works out -- `a[i, idx[k]] * v[k]` -- saying that a
|
|
449
|
+
contraction counts positions and this one is a position of `idx`. It was
|
|
450
|
+
read as an index appearing once, so the sum the notation asks for did not
|
|
451
|
+
happen and the result came back a whole matrix. Write such a gather with
|
|
452
|
+
`CArray.jit_for`.
|
|
453
|
+
|
|
454
|
+
- Change: an indexed kernel (`CArray.jit_for` and the rest) refuses an
|
|
455
|
+
operand it cannot walk -- a gather, a lazy array -- that is a view of an
|
|
456
|
+
array the same kernel writes, naming `a[order[i]]` as the spelling that
|
|
457
|
+
means something. Such an operand is copied before the loop and put back
|
|
458
|
+
after it, so the loop read cells that had stopped being current and a
|
|
459
|
+
direct write to the same array was dropped at the end.
|
|
460
|
+
|
|
461
|
+
- Change: arithmetic between two booleans -- `flag[i] * flag[i]`, `+`, `-`,
|
|
462
|
+
`/`, `%` -- raises `CArray::JIT::Unsupported` where the operator stands,
|
|
463
|
+
as `true * true` raises in Ruby. Used as a condition it compiled and ran.
|
|
464
|
+
`&`, `|` and `^` on booleans are unchanged.
|
|
465
|
+
|
|
466
|
+
- Change: a captured Integer above 2**63-1 meeting an integer is refused,
|
|
467
|
+
where before it was absorbed into that integer's width and computed from a
|
|
468
|
+
wrapped value. CArray refuses the same expression whatever the array's own
|
|
469
|
+
type is -- `CArray.uint64(1) { 5 } + 2**63` raises `bignum too big to
|
|
470
|
+
convert into 'long long'` -- and this follows it: a bare Integer brings no
|
|
471
|
+
width, and the message names `CScalar.uint64() { big }`, which a kernel reads
|
|
472
|
+
as the one-cell array it is. A Float or a Complex on the other side is
|
|
473
|
+
unaffected, and so is a capture that fits an int64. A loop that adds such a
|
|
474
|
+
capture to an accumulator is refused for the same reason: state the width
|
|
475
|
+
once with a CScalar, for the seed and for the value.
|
|
476
|
+
|
|
477
|
+
- Change: a `CArray.jit_each` or `CArray.jit_map` block whose *only* array is
|
|
478
|
+
handed to a C function whole is now refused, naming the array and saying
|
|
479
|
+
that it does not settle how many cells there are to compute.
|
|
480
|
+
`CArray.jit_map { DOT4.call(w, w) }` used to answer an array as long as `w`
|
|
481
|
+
-- a length that came from an array nobody walks -- and the same block under
|
|
482
|
+
`jit_each` failed inside Fiddle with `unknown symbol "ca_call_cslab_0_r"`.
|
|
483
|
+
`CArray.jit_for` with a count takes these, as it always did.
|
|
484
|
+
|
|
485
|
+
- Change: a block parameter that names a loop index is refused when the
|
|
486
|
+
generated C already uses that name for a captured array (`p_a`, `m_a`,
|
|
487
|
+
`a_s0`, `a_ms0` or `a_n0` for an array `a`), or when it contains `__` or
|
|
488
|
+
starts with `carray_jit_`. Rename the index.
|
|
489
|
+
|
|
490
|
+
- Change: a `CArray.jit_function` body that subscripts a pointer parameter
|
|
491
|
+
with a literal outside the length its declaration gave -- `v[7]` or
|
|
492
|
+
`v[-1]` against `double v[2]` -- raises `CArray::JIT::Unsupported` as the
|
|
493
|
+
body is read. It compiled, and wrote or read past what the caller was
|
|
494
|
+
held to. A computed subscript, and any subscript on a pointer declared
|
|
495
|
+
without a length, are unchecked as before.
|
|
496
|
+
|
|
497
|
+
- Change: an extent that counts down as a `Range` -- `(n-2)..0` -- raises
|
|
498
|
+
`CArray::JIT::Unsupported` naming `(n-2).step(0, -1)`, the spelling Ruby
|
|
499
|
+
iterates backwards with. Ruby gives such a Range no elements, so the
|
|
500
|
+
kernel ran no passes and said nothing, which is what the guide says this
|
|
501
|
+
refusal exists to prevent. An empty range whose ends agree (`0...0`) is a
|
|
502
|
+
count of zero as before.
|
|
503
|
+
|
|
504
|
+
- Change: a body compiled with `CArray.jit_function` or `CArray.jit_call` may
|
|
505
|
+
call a function bound with `jit_extern`. It was refused -- a borrowed
|
|
506
|
+
function is only an address, and a compiled body has no `functions` buffer
|
|
507
|
+
to take one in -- but an address is not what a body needs from it. It needs
|
|
508
|
+
the name, which is what the declaration states and what a linker or a
|
|
509
|
+
loader resolves; `jit_extern` opened the library to find the function, so
|
|
510
|
+
the symbol is in the process by the time the body is compiled. The
|
|
511
|
+
generated file declares it (`double j0(double);`) and calls it by name. A
|
|
512
|
+
kernel is unchanged and still takes the address through its buffer, which
|
|
513
|
+
is what a kernel has and a body has not. A file built ahead of the program
|
|
514
|
+
links the library the usual way: one that is not linked by default is the
|
|
515
|
+
caller's to add.
|
|
516
|
+
|
|
517
|
+
- Change: `require "carray/jit/call"` defines `CArray.jit_call` without
|
|
518
|
+
loading the compiler, and the compiler is required at the first call site
|
|
519
|
+
no provider answers. A program whose sites were all built ahead of it --
|
|
520
|
+
by carray-jit-aot -- never loads the analyzer, the generator or the cache
|
|
521
|
+
at all, where before it loaded them and reached none of them.
|
|
522
|
+
`require "carray/jit"` loads everything as it did, and a program that
|
|
523
|
+
knows nothing about this sees no difference.
|
|
524
|
+
|
|
525
|
+
- Change: a subscript that walks with one index and adds another --
|
|
526
|
+
`a[i + r]`, `r` an inner loop's index -- is refused as the block is read,
|
|
527
|
+
with a message that says so and that the sum put in a local first
|
|
528
|
+
(`k = i + r`, then `a[k]`) is checked at each cell and runs. It was
|
|
529
|
+
refused at the call, for a reason about an inner loop's range. The
|
|
530
|
+
exception is `CArray::JIT::Unsupported`, as it was.
|
|
531
|
+
|
|
532
|
+
- Change: a kernel or `jit_function` body that reads a local after an inner
|
|
533
|
+
loop's block in which the local was first assigned, or after a `while` in
|
|
534
|
+
which it was first assigned, is refused with a message naming the local.
|
|
535
|
+
These did not compile before either, failing in the C compiler instead.
|
|
536
|
+
Give the local a value before the loop.
|
|
537
|
+
|
|
538
|
+
- Change: a local assigned before the summand of a `CArray.jit_contract`
|
|
539
|
+
block is refused saying so -- a contraction is one expression, which is
|
|
540
|
+
why it takes no local array either. It was refused as "`u` is read before
|
|
541
|
+
it is assigned", about a local assigned on the line above.
|
|
542
|
+
|
|
543
|
+
- Change: the refusal for a `CArray.jit_stencil` writing into a view of an
|
|
544
|
+
array it reads says that the two are views of one array and that the test
|
|
545
|
+
is the storage rather than the cells, so two slabs that share none are
|
|
546
|
+
refused too, and names the copy to read from. It said a stencil writes
|
|
547
|
+
into an array of its own, which of two slabs is not true.
|
|
548
|
+
|
|
549
|
+
- Change: the refusal for a block whose source cannot be read names
|
|
550
|
+
`RubyVM.keep_script_lines = true` and `CArray::JIT.compile`, which takes a
|
|
551
|
+
kernel as text. It named `source:`, which is a keyword no entry point
|
|
552
|
+
takes; the guide said the same.
|
|
553
|
+
|
|
554
|
+
- Change: `CArray#jit_init` given a block that takes a splat or an optional
|
|
555
|
+
parameter says so, where it said "the block names -1 indices".
|
|
556
|
+
|
|
557
|
+
- Change: a keyword in value position is named as it was written --
|
|
558
|
+
`unless` in `CArray.jit_map { unless a > 1.0 then ... end }` says
|
|
559
|
+
"`unless` -- write it as `if` with the condition negated", where it said
|
|
560
|
+
"unsupported expression Unless". Statement position already did.
|
|
561
|
+
|
|
562
|
+
- Change: a statement outside the subset is named as it was written --
|
|
563
|
+
``got `unless` -- write it as `if` with the condition negated`` rather than
|
|
564
|
+
`got Unless`, which was this compiler's reading of it and not anything
|
|
565
|
+
anyone typed. The list of what a body may hold was written out in two
|
|
566
|
+
places and they had come apart; it is one place now.
|
|
567
|
+
|
|
568
|
+
- Change: the three names left in Ruby's `Math` say why they are not lowered
|
|
569
|
+
rather than saying that C has no counterpart, which was untrue of all of
|
|
570
|
+
them -- `lgamma`, `frexp` and `ldexp` all exist. `Math.lgamma` and
|
|
571
|
+
`Math.frexp` each answer with a pair, and a cell holds one number;
|
|
572
|
+
`Math.ldexp` takes an exponent where a math call here computes every
|
|
573
|
+
argument in the type of its result. A name Ruby's `Math` does not have says
|
|
574
|
+
that instead of guessing at a reason.
|
|
575
|
+
|
|
576
|
+
- Change: a compiler this process cannot find -- `CARRAY_JIT_CC` naming
|
|
577
|
+
something that is not there -- raises `CArray::JIT::CompilationError`
|
|
578
|
+
saying so and naming the variable, where it raised `Errno::ENOENT` from
|
|
579
|
+
the spawn with the program name and nothing else.
|
|
580
|
+
|
|
581
|
+
- Change: a cache directory owned by another user is refused, as one other
|
|
582
|
+
users can write to already was; everything in it is `dlopen`ed. A cache
|
|
583
|
+
root that is a symlink is followed as before, and the owner of the
|
|
584
|
+
directory it lands on is what is asked about. Set `CARRAY_JIT_CACHE` to a
|
|
585
|
+
directory of your own where this refuses.
|
|
586
|
+
|
|
587
|
+
- Change: the docs now say what a compare-exchange costs under a mask. A
|
|
588
|
+
swap writes two cells where an ordinary assignment writes one, and a sort
|
|
589
|
+
runs one over every cell repeatedly, so a single missing cell spreads
|
|
590
|
+
across the row -- the ordinary rule for a branch decided by a missing
|
|
591
|
+
cell, at an unusually high gain. A sort written by hand is not refused
|
|
592
|
+
the way `sort` on a local array is, so ask `a[i] == UNDEF` first and
|
|
593
|
+
decide it.
|
|
594
|
+
|
|
595
|
+
- Change: the docs now say what a branch not taken costs under a mask. A
|
|
596
|
+
branch decided by a missing cell masks what it writes; taking the other
|
|
597
|
+
path writes nothing, so a conclusion reached that way -- `found` left at
|
|
598
|
+
-1 for a row whose only candidate was masked -- is reported as an ordinary
|
|
599
|
+
value. Ask about the mask with `a[i] == UNDEF` where that matters.
|
|
600
|
+
|
|
601
|
+
- Change: the docs now say that a negative Float to a fractional power is
|
|
602
|
+
`NaN`, as `pow` and CArray answer, where Ruby answers a Complex. This was
|
|
603
|
+
already the behaviour; a whole-number exponent or a non-negative base is
|
|
604
|
+
Ruby's number.
|
|
605
|
+
|
|
606
|
+
- Change: the docs now say that an Integer compared with a Float is
|
|
607
|
+
compared as two doubles, as CArray compares them, so past 2^53 the answer
|
|
608
|
+
can differ from Ruby's exact comparison. This was already the behaviour;
|
|
609
|
+
below 2^53 nothing differs.
|
|
610
|
+
|
|
611
|
+
- Change: the docs now say that `Math.sqrt`, `log`, `log2`, `log10`, `asin`,
|
|
612
|
+
`acos`, `acosh` and `atanh` answer `NaN` outside their domain, as `math.h`
|
|
613
|
+
does, where Ruby raises `Math::DomainError`, and that `Math.sqrt(-0.0)` is
|
|
614
|
+
`-0.0`. This was already the behaviour. `Math.gamma` still raises as Ruby
|
|
615
|
+
does. Test the argument in the block where it may leave the domain.
|
|
616
|
+
|
|
617
|
+
- Fix: a loop whose bound is an unsigned parameter -- `size_t n` and its kin
|
|
618
|
+
-- runs the passes Ruby runs. The counter is an `int64_t` and the bound was
|
|
619
|
+
compared against it uncast, so C converted the *counter* to unsigned
|
|
620
|
+
instead: `(-3...n).each` over `n = 4` ran no passes where Ruby runs seven,
|
|
621
|
+
and answered without saying anything. A count that never goes below zero
|
|
622
|
+
was unaffected, which is every `n.times`. The cast is written down now, so
|
|
623
|
+
the generated C also compiles clean where it carried `-Wsign-compare`
|
|
624
|
+
before.
|
|
625
|
+
|
|
626
|
+
- Fix: a `CArray.jit_each` or `CArray.jit_map` pass that reads a view of an
|
|
627
|
+
array it also writes -- two blocks that share cells, a transpose written
|
|
628
|
+
over itself, a gather held in `g` and written back as `src = g + 1.0` --
|
|
629
|
+
reads the values the expression started with, as Ruby and CArray's own
|
|
630
|
+
operators do. Such an operand was read as the pass went, so it came back
|
|
631
|
+
holding the kernel's own output: `hi = lo * 2.0` over overlapping blocks
|
|
632
|
+
gave all zeros, and a gather over 20000 cells was wrong in 8192 of them.
|
|
633
|
+
It is copied once before the loop instead. `a = a + 1.0`, and views that
|
|
634
|
+
share a root without sharing a cell, are untouched.
|
|
635
|
+
|
|
636
|
+
- Fix: `CArray.jit_contract` takes a summand that calls a function made with
|
|
637
|
+
`CArray.jit_function` or `CArray.jit_extern`, which the docs say it does;
|
|
638
|
+
it was refused as an unsupported method. What the function hands back
|
|
639
|
+
types the result the contraction is collected into.
|
|
640
|
+
|
|
641
|
+
- Fix: a cell of a local array written under an `if` or a `while` whose
|
|
642
|
+
condition read a missing cell is masked, as a plain local and a cell of an
|
|
643
|
+
array now are. It was left present deliberately; the rule the docs give
|
|
644
|
+
for it is the one the other two keep.
|
|
645
|
+
|
|
646
|
+
- Fix: a local assigned under an `if` or a `while` whose condition read a
|
|
647
|
+
missing cell is masked, as a cell written there already was. It came back
|
|
648
|
+
a number like any other, so `found = -1; ...; if a[i, j] > x then found =
|
|
649
|
+
j end; out[i] = found` reported a position found under a mask. A condition
|
|
650
|
+
that asks about the mask itself (`a[i] == UNDEF`) still masks nothing.
|
|
651
|
+
|
|
652
|
+
- Fix: a build no longer fails with "could not load freshly compiled" when
|
|
653
|
+
another process evicts its object in the moment between the build and the
|
|
654
|
+
load. It opens the object before publishing it to the cache. Reachable
|
|
655
|
+
where `CARRAY_JIT_CACHE_LIMIT` is smaller than the number of processes
|
|
656
|
+
building at once: with five of them, a limit of 2 lost every one.
|
|
657
|
+
|
|
658
|
+
- Fix: a build whose compile fails leaves no `.c` behind in the cache. Such
|
|
659
|
+
a file was never looked up -- a cache hit is an object -- and eviction
|
|
660
|
+
passes over it, so one stayed for good, and one more for every kernel of
|
|
661
|
+
every run against a toolchain that cannot compile.
|
|
662
|
+
|
|
663
|
+
- Fix: two threads of one process compiling at the same time no longer
|
|
664
|
+
raise `Errno::ENOENT`; three threads in four did. A build now takes a lock
|
|
665
|
+
for the length of the compile, so the second thread finds the object the
|
|
666
|
+
first one left and reuses it rather than building its own. Compiling two
|
|
667
|
+
different kernels in two threads is that much less parallel -- four of
|
|
668
|
+
them took 970 ms here against 800.
|
|
669
|
+
|
|
670
|
+
- Fix: multiplying two Complex numbers, or a real number by a Complex
|
|
671
|
+
(`x * z`), gives Ruby's answer where an infinity meets a zero:
|
|
672
|
+
`0.0 * Complex(Float::INFINITY, 0.0)` is `0.0+0.0i`, where it was
|
|
673
|
+
`NaN+NaN*i`. Finite products are unchanged, `cmplx64` included, and
|
|
674
|
+
`z * x` still scales each part as it did.
|
|
675
|
+
|
|
676
|
+
- Fix: `%` by a float zero raises `ZeroDivisionError`, as Ruby's does --
|
|
677
|
+
`x % 0.0`, `x % -0.0`, an integer cell `% 0.0` -- in a kernel and in a
|
|
678
|
+
`CArray.jit_function` alike. It answered `NaN`, which is what CArray's `%`
|
|
679
|
+
answers and not what the docs promised. A float `/` by zero is still an
|
|
680
|
+
infinity, as it is in Ruby.
|
|
681
|
+
|
|
682
|
+
- Fix: `min(w)` and `max(w)` over a floating local array keep the first of
|
|
683
|
+
two cells that compare equal, as CArray's `min` and `max` do, so `0.0`
|
|
684
|
+
and `-0.0` come back in the order they stood; which zero came back was
|
|
685
|
+
left to the C library. NaN is still skipped, and an array of nothing but
|
|
686
|
+
NaN still answers `NaN`.
|
|
687
|
+
|
|
688
|
+
- Fix: `floor`, `ceil`, `round`, `truncate` and `to_i` on a Float raise
|
|
689
|
+
`FloatDomainError` for a NaN or an infinity, as Ruby does, and `RangeError`
|
|
690
|
+
for a result past int64; they gave a clamped or arbitrary number. Written
|
|
691
|
+
straight into a float cell, or returned from a `CArray.jit_function`
|
|
692
|
+
declared `double`, the result is Ruby's however large -- `1e20.floor` is
|
|
693
|
+
`1e20`. A loop doing little but rounding runs slower for the check.
|
|
694
|
+
|
|
695
|
+
- Fix: `.abs` on an integer compiles in `CArray.jit_for`, `CArray.jit_map`,
|
|
696
|
+
`CArray.jit_function` and the other entry points; it raised
|
|
697
|
+
`CArray::JIT::CompilationError` unless the kernel also allocated a local
|
|
698
|
+
array. `(-2**63).abs` in an int64 wraps to itself, as other int64
|
|
699
|
+
overflow does. Float and complex `.abs` are unchanged.
|
|
700
|
+
|
|
701
|
+
- Fix: an inner loop stepping by more than one over a start written over
|
|
702
|
+
another index -- `2.times { |p| (p...8).step(3) { |k| ... } }` -- is held
|
|
703
|
+
to the cells every pass reaches, not only the pass from the earliest
|
|
704
|
+
start. A subscript one pass took past the end of an array was let through
|
|
705
|
+
and written; it is now refused as the call is prepared. A step of one, and
|
|
706
|
+
a start that does not move, are read as before.
|
|
707
|
+
|
|
708
|
+
- Fix: a local array whose shape is written over captured integers --
|
|
709
|
+
`CArray.double(n, m)` -- raises `ArgumentError` when its lengths each fit
|
|
710
|
+
but their product does not count in bytes. Such a product wrapped to a
|
|
711
|
+
small number, the kernel allocated that, and subscripts in range on every
|
|
712
|
+
axis wrote past it. A shape of one axis, or of lengths whose product
|
|
713
|
+
fits, is unaffected.
|
|
714
|
+
|
|
715
|
+
- Fix: `CArray::JIT.clear_registry` now forgets the functions
|
|
716
|
+
`CArray.jit_function` compiled, as it already did kernels, so the next
|
|
717
|
+
call reads them back from the cache on disk.
|
|
718
|
+
|
|
719
|
+
- Fix: a block run through `eval` -- in a console, or from code that builds
|
|
720
|
+
its kernels as text -- no longer stays in memory for the life of the
|
|
721
|
+
process once nothing refers to it. On Ruby 3.2 it still does.
|
|
722
|
+
|
|
723
|
+
- Fix: a kernel over a masked array read and wrote the wrong cells of its
|
|
724
|
+
mask, and on a large enough array wrote past the end of it, when an
|
|
725
|
+
unmasked operand whose number of axes differs from the kernel's -- a row
|
|
726
|
+
added across a grid, or a contraction's operand -- sorted by name before
|
|
727
|
+
the masked one. Kernels whose operands all share the kernel's rank, and
|
|
728
|
+
kernels over no masked array, were not affected.
|
|
729
|
+
|
|
730
|
+
- Fix: a local array indexed by an inner loop whose range is not known until
|
|
731
|
+
the kernel runs -- a bound that is a captured integer, and now a range
|
|
732
|
+
written over another index -- is checked where the cell is reached. Such a
|
|
733
|
+
subscript was checked nowhere: the check that reads the loop's range could
|
|
734
|
+
not settle it, and the check at the access did not cover an index, so a
|
|
735
|
+
write past the end of the array went into the block's stack frame. It now
|
|
736
|
+
raises `IndexError` as every other unsettled subscript does.
|
|
737
|
+
|
|
738
|
+
- Fix: in a `CArray.jit_each`, `CArray.jit_map` or `CArray.jit_stencil`
|
|
739
|
+
block, an array handed to a C function whole -- and read in no other way --
|
|
740
|
+
is no longer lined up with the operands. One whose length differed from
|
|
741
|
+
theirs used to fail with `broadcast_to: cannot broadcast axis 0`, which
|
|
742
|
+
named an axis and said nothing about the call it was written for; a set of
|
|
743
|
+
weights four cells long may now stand beside a thousand cells of operand.
|
|
744
|
+
What decides which arrays those are is the declaration, not the spelling: a
|
|
745
|
+
parameter taking a number by value reads the cell, so that array is walked
|
|
746
|
+
and lines up as before, and an array both read by cell and handed over is
|
|
747
|
+
an operand too.
|
|
748
|
+
|
|
749
|
+
- Fix: a local in a kernel or `jit_function` body, or a captured variable
|
|
750
|
+
starting `carray_jit_`, may be given a name the generated C also uses -- a
|
|
751
|
+
kernel parameter such as `error`, a C keyword such as `int`, `a_n0` beside a
|
|
752
|
+
captured array `a`, or a name containing `__`. It failed to compile, raised
|
|
753
|
+
an `IndexError` for an index in range, or shared its value with another
|
|
754
|
+
name; it is now renamed in the generated C only.
|
|
755
|
+
|
|
756
|
+
- Fix: two inner loops in one kernel or `jit_function` body may assign a
|
|
757
|
+
local of the same name, at one type or two, and a local assigned inside a
|
|
758
|
+
`while` may be read after it when it was also assigned before it. The
|
|
759
|
+
first failed in the C compiler or was refused as changing type; the second
|
|
760
|
+
failed in the C compiler.
|
|
761
|
+
|
|
762
|
+
- Fix: a loop in a `jit_function` body, and a `while` or inner loop in a
|
|
763
|
+
kernel whose body held the kernel's first way to fail, kept running after
|
|
764
|
+
an integer division by zero, a computed index out of range, `clamp` or
|
|
765
|
+
`Math.gamma` had reported a failure -- so a `while` decided by the value
|
|
766
|
+
handed back could run forever. Such a loop now leaves at the head of its
|
|
767
|
+
next pass, and the call raises as it already did.
|
|
768
|
+
|
|
769
|
+
- Fix: `CArray.jit_each` and `CArray.jit_map` no longer refuse a block that
|
|
770
|
+
hands a one-cell array to a C function. Those entries line their operands up
|
|
771
|
+
with the expression's shape, and an array passed to a pointer parameter was
|
|
772
|
+
lined up with the rest -- so a one-cell array holding state came back
|
|
773
|
+
stretched and read-only, and the copy-back after the call raised `can not
|
|
774
|
+
modify read-only array`. An array handed over by address is passed whole
|
|
775
|
+
rather than walked, so it is left as it is now. `jit_for` was never
|
|
776
|
+
affected, and a `CScalar` was affected the same way an array was.
|
|
777
|
+
|
|
778
|
+
- Fix: a captured Integer above 2**63-1 reaches a kernel whole. It was packed
|
|
779
|
+
into the int64 slot the kernel reads its integers from, which took the value
|
|
780
|
+
modulo the width and said nothing: `big / 3`, `big > 100` and `big * 1.0`
|
|
781
|
+
answered from a negative number, while `+` and `*` came out right and hid it.
|
|
782
|
+
Such a capture is a uint64 now -- the width CArray has for those values --
|
|
783
|
+
and a value neither an int64 nor a uint64 holds is refused where the capture
|
|
784
|
+
is read, naming the value.
|
|
785
|
+
|
|
786
|
+
- Fix: a C declaration written with `ptrdiff_t` compiles. `size_t` and the
|
|
787
|
+
other `<stddef.h>` names have read since 0.1.0, but the C generated for a
|
|
788
|
+
body that took one declared no such type, so `CArray.jit_function("ptrdiff_t
|
|
789
|
+
(*)(ptrdiff_t)")` failed in the C compiler with `unknown type name`.
|
|
790
|
+
`size_t` was unaffected, by luck rather than by design.
|
|
791
|
+
|
|
792
|
+
- Fix: two inner loops in one body may both be written `{ |k| ... }`. They
|
|
793
|
+
are one name in the block and were one index here, so the second loop's
|
|
794
|
+
range replaced the first's and a reach was checked against the wrong one --
|
|
795
|
+
`a[k-1]` in a loop from 1 was refused for starting at 0 once a later loop
|
|
796
|
+
started there. Each loop now counts in an identifier of its own, and the
|
|
797
|
+
messages go on speaking the name the block wrote.
|
|
798
|
+
|
|
799
|
+
- Fix: a block holding a character outside ASCII no longer raises. The file a
|
|
800
|
+
block sits in was read with `File.read`, which uses
|
|
801
|
+
`Encoding.default_external` -- a setting that has nothing to do with a
|
|
802
|
+
source's encoding: on a machine with no locale set it is US-ASCII, and the
|
|
803
|
+
file came back as its own bytes under a tag that the first `rstrip` on a
|
|
804
|
+
line holding a comment in Japanese raised on. The file is now read as bytes
|
|
805
|
+
and given the encoding the parser gave it: UTF-8, or what a `coding` magic
|
|
806
|
+
comment names on the first line or on the second where a shebang takes the
|
|
807
|
+
first.
|
|
808
|
+
|
|
809
|
+
|
|
40
810
|
## 0.1.2
|
|
41
811
|
|
|
42
812
|
- New: `CArray.jit_contract` takes the result's axes as symbols, which says
|
|
@@ -115,9 +885,7 @@ version you have and a newer one.
|
|
|
115
885
|
`CArray.jit_function("void (*)(uint64_t *, int64_t)")` reached its cells as
|
|
116
886
|
an opaque slot rather than an array, and a `uint64_t` return type was
|
|
117
887
|
refused as "no value a compiled body can produce" -- from a table written
|
|
118
|
-
before uint64 was a type a kernel computes in.
|
|
119
|
-
by value is still refused, and now says why: a value reaches a body as a
|
|
120
|
-
double, an int64 or a complex, and a uint64 fits none of them whole.
|
|
888
|
+
before uint64 was a type a kernel computes in.
|
|
121
889
|
|
|
122
890
|
- Fix: `CArray.jit_map` collects a cmplx64 value into a cmplx64 array rather
|
|
123
891
|
than refusing to allocate one. Nothing else changes type: a block whose value
|