carray-jit 0.1.1 → 0.1.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (45) hide show
  1. checksums.yaml +4 -4
  2. data/CHANGELOG.md +854 -0
  3. data/README.md +7 -6
  4. data/carray-jit.gemspec +1 -3
  5. data/docs/00_Introduction.md +4 -3
  6. data/docs/01_GettingStarted.md +1 -1
  7. data/docs/02_KernelShapes.md +150 -10
  8. data/docs/03_SupportedFeatures.md +582 -26
  9. data/docs/04_Compiling.md +44 -7
  10. data/docs/05_DesignNotes.md +3 -3
  11. data/docs/06_Cheatsheet.md +216 -11
  12. data/docs/07_StepByStep.ja.md +534 -0
  13. data/docs/07_StepByStep.md +535 -0
  14. data/examples/README.md +12 -0
  15. data/examples/applications/alarm.rb +121 -0
  16. data/examples/applications/collatz.rb +105 -0
  17. data/examples/applications/cubic_spline.rb +331 -0
  18. data/examples/applications/dithering.rb +144 -0
  19. data/examples/applications/group_stats.rb +115 -0
  20. data/examples/applications/lookup.rb +126 -0
  21. data/examples/applications/median_filter.rb +153 -0
  22. data/examples/applications/parcel_ascent.rb +220 -0
  23. data/examples/applications/point_cloud.rb +23 -11
  24. data/examples/applications/point_in_polygon.rb +111 -0
  25. data/examples/applications/random_walk.rb +98 -0
  26. data/examples/applications/van_der_pol.rb +186 -0
  27. data/examples/applications/wet_bulb.rb +140 -0
  28. data/examples/features/10_complex.rb +14 -4
  29. data/examples/features/15_loops.rb +7 -1
  30. data/lib/carray/jit/access.rb +14 -0
  31. data/lib/carray/jit/analyzer.rb +2179 -119
  32. data/lib/carray/jit/block_reader.rb +37 -6
  33. data/lib/carray/jit/c_function.rb +634 -74
  34. data/lib/carray/jit/c_generator.rb +1607 -156
  35. data/lib/carray/jit/call.rb +68 -0
  36. data/lib/carray/jit/compiler.rb +94 -12
  37. data/lib/carray/jit/kernel.rb +371 -33
  38. data/lib/carray/jit/node.rb +359 -9
  39. data/lib/carray/jit/sorting_networks.rb +182 -0
  40. data/lib/carray/jit/type_assignment.rb +386 -47
  41. data/lib/carray/jit/version.rb +1 -1
  42. data/lib/carray/jit.rb +936 -63
  43. metadata +22 -8
  44. data/ext/carray_jit_access/carray_jit_access.c +0 -460
  45. data/ext/carray_jit_access/extconf.rb +0 -8
data/CHANGELOG.md CHANGED
@@ -37,6 +37,860 @@ version you have and a newer one.
37
37
  a release needs is said in the gemspec, and an entry says so only
38
38
  when the answer changes. -->
39
39
 
40
+ ## 0.1.3
41
+
42
+ - New: `CArray::JIT.call_provider` is asked at a `jit_call` site before
43
+ anything is compiled, and answering `nil` means "not mine, compile it". It
44
+ is handed the prototype, the block -- whose binding says which method and
45
+ which module the site is in -- and the names the declaration gave, and
46
+ answers anything that responds to `call`. The one client is carray-aot,
47
+ which builds these same sites into a shared object ahead of the program so
48
+ that a machine running that gem reaches no compiler; it is the position
49
+ `jit_extern` puts a function from a library in, said about a call rather
50
+ than about a name.
51
+
52
+ - New: `CArray.jit_call(prototype) { ... }` compiles the block as a C
53
+ function and calls it where it stands, with the locals around it. The
54
+ declaration's parameter names are the join and do the work twice: they are
55
+ the body's parameters, so the block declares none, and they name the locals
56
+ the call reads -- `CArray.jit_call("void (*)(double *out, const double
57
+ *values, size_t n, size_t window)")` inside a method holding `out`,
58
+ `values`, `n` and `window`. In C a parameter's name in a prototype is
59
+ decoration; here it is the whole binding, and a declared name with no local
60
+ behind it is refused at the call rather than read as nil inside it. What it
61
+ saves over `jit_function` and a `call` is the argument list, which restates
62
+ the declaration in an order nothing checks; what it costs is reading those
63
+ locals through the block's binding, about 0.3 microseconds against a call
64
+ that costs several. Compiled once per call site, and `clear_registry`
65
+ reaches those as it reaches the rest.
66
+
67
+ - New: `CArray::JIT::CFunction#c_source_as(symbol)` hands the compiled C over
68
+ under a symbol the caller picked, for a caller writing it into a file of
69
+ its own rather than letting this compile it -- the symbol this compiler
70
+ writes carries a digest of the body, which is right for an object in a
71
+ cache and wrong for one committed to a repository. A name C cannot spell
72
+ and a name in this compiler's own `carray_jit_` namespace are refused; a
73
+ name the C library already has is the caller's to avoid, since the prefix
74
+ that closes that hazard is what is being replaced. A function bound with
75
+ `jit_extern` has no C of its own and says so.
76
+
77
+ - New: a parallel assignment is in the subset -- `a, b = b, a`, and the
78
+ `a, b = b, a + b` a recurrence advances by. Every value on the right is
79
+ settled before anything on the left is written, as in Ruby, so the
80
+ temporary that spelling saves no longer has to be written by hand. A cell
81
+ is a target as much as a name is (`a[i], a[j] = a[j], a[i]`), each value
82
+ keeps its own type and carries its mask. The readings that take a single
83
+ value apart are refused by name: `a, b = f(x)`, `a, b = [1, 2]`,
84
+ `a, *rest = ...`, `a, (b, c) = ...`, and a count that does not match.
85
+
86
+ - New: a local array may be larger than a stack frame should hold, and its
87
+ shape may be written over an integer the block captured --
88
+ `CArray.double(n)`, which was refused. Either way the kernel allocates the
89
+ array once at its entry and frees it at its exit, so a constructor written
90
+ inside the cell loop is still one allocation; 4 KiB for one array and
91
+ 16 KiB for one kernel's arrays together now say where an array lives rather
92
+ than whether it is allowed, and nothing is refused for its size. A kernel
93
+ whose arrays all fit in the frame emits the C it emitted before. Where the
94
+ length is one the kernel works out, every subscript on that axis is checked
95
+ where the cell is reached rather than as the block is read, and a C
96
+ function takes the array only through a pointer that declares no length
97
+ (`const double *v`, not `const double v[3]`). The same block at two lengths
98
+ is one compiled kernel: the length travels as an argument and is not in the
99
+ C. A shape that comes to zero or less raises `ArgumentError` when the
100
+ kernel runs, and an allocation the system refuses raises `NoMemoryError`,
101
+ both naming the array and what its shape came to. A `jit_function` body
102
+ allocates nothing -- it is called once per cell -- and refuses such an
103
+ array, naming the pointer parameter to take it through instead.
104
+
105
+ - New: a kernel that carries masks takes a local array, which it refused
106
+ before. Every local array of such a kernel is declared with a shadow of one
107
+ byte a cell beside its cells, and a cell carries a mask the way a plain
108
+ local does: what the expression written into it carried. So a window copied
109
+ into a workspace keeps its holes. `w[k] = UNDEF` marks a cell and
110
+ `w[k] == UNDEF` asks about one; the zeroed spellings clear the shadow with
111
+ the cells, so a cell starts every pass present, while `CArray.empty` leaves
112
+ both unspecified. The shadow counts against the 4 KiB an array is held to
113
+ and the 16 KiB a kernel is, so an array that fits by its cells alone may
114
+ not fit once it carries masks. Two things still refuse such an array:
115
+ `sum`, `min`, `max` and `sort`, because what they should do with a missing
116
+ cell is not decided, and a C function, because a mask travels in no C
117
+ declaration -- both say so where they used to be refused for having no mask
118
+ at all. A kernel that carries no masks emits what it emitted before, and a
119
+ compiled function's body carries none at all.
120
+
121
+ - New: an inner loop's range may be written over another index --
122
+ `3.times { |p| (p+1...3).each { |r| ... } }`, the shape a triangular loop
123
+ takes -- where that index is one of the loops around it. Forward
124
+ elimination and Neville's interpolation are the two this was wanted for,
125
+ and both now read as they do on paper rather than as a full loop with an
126
+ `if` inside it. The C was always emitted; what stood in the way was the
127
+ call, where each index's range is worked out so that a subscript can be
128
+ held to its array. A range over another index has no single pair of
129
+ numbers to be, so it is read as an interval, at its widest: every pass the
130
+ loop could take and sometimes more. Where that refuses a reach the loop
131
+ never makes, the message says which index the range was written over and
132
+ that it was read at its widest. A range over a local variable, a sibling
133
+ loop's index or a deeper one is still refused.
134
+
135
+ - New: a local array may have more than one axis --
136
+ `CArray.double(3, 4)`, `CArray.new(:float64, [2, 3, 4])` -- in any of the
137
+ five places one can be made. The shape is written out as it always was, so
138
+ the strides are constants: `m[r, c]` is `m[(r) * 4 + (c)]`, one subscript
139
+ per axis, and each axis is checked against its own extent. That is what
140
+ flattening by hand gives up -- `m[r * 4 + c]` with a column of 4 reads a
141
+ cell of the next row and says nothing, where `m[r, c]` is refused by name
142
+ and a computed column raises. The stack limits count cells rather than
143
+ axes, so `CArray.double(32, 32)` is 8 KiB and past the 4 KiB an array is
144
+ held to. Handed to a C function the array goes as the flat run of cells it
145
+ is, row after row, so a `double[3][4]` reaches `const double a[12]` and the
146
+ length is matched over every cell. The four intrinsics still take one axis.
147
+
148
+ - New: a `CArray.jit_function` body may make a local array and use
149
+ `sum`/`min`/`max`/`sort` over one, which completes the four entry points.
150
+ A body closes over nothing, so this is the only place scratch space could
151
+ come from other than an extra parameter -- and a signature settled
152
+ elsewhere, a callback's, has no room for one. The array is declared at the
153
+ head of the function, a recursive body gets one per call, and a body hands
154
+ one to another compiled function under the same rules a kernel does. The
155
+ helpers an intrinsic needs travel with the body: carried in its own file
156
+ when it is compiled alone, and merged into the kernel's preamble when it is
157
+ pasted, one helper per element type and length however many bodies want it.
158
+ The stack limits stay per function -- 4 KiB an array, 16 KiB a body -- and a
159
+ chain of pasted functions is not counted, so a deep chain stands as many
160
+ frames as it has; that is the bargain a deep recursion already takes. Still
161
+ excluded: more than one axis, and a contraction.
162
+
163
+ - New: a local array may be handed to a C function -- one from
164
+ `CArray.jit_function` or `CArray.jit_extern` -- wherever the declaration
165
+ takes a pointer, from any of the four entry points that make one. What the
166
+ declaration says is matched as the block is read rather than at the call,
167
+ both the element type and the length being written in the block: an exact
168
+ element type, and at least as many cells as a sized declarator asks for.
169
+ A parameter that is not `const` may be written through, and the line after
170
+ the call reads what the callee left; the zeroed constructors still clear at
171
+ the line, so nothing carries into the next cell. Passing one array to two
172
+ parameters is fine, the declarations carrying no `restrict`. Note that a
173
+ borrowed function which keeps the pointer past the call is left pointing at
174
+ a stack frame that has gone, and that a declaration carrying no length --
175
+ `const double *x` -- gives nothing to check against. Still excluded when this landed: a `jit_function` body,
176
+ more than one axis, and a contraction; the first two arrived later in this
177
+ release, and a contraction is still refused.
178
+
179
+ - New: a local array, and the four intrinsics over one, may be written in a
180
+ `CArray.jit_each`, `CArray.jit_map` or `CArray.jit_stencil` block as well as
181
+ in `CArray.jit_for`. These are the spellings that wanted one: a block with
182
+ no index cannot pick a row of a captured array, so a `jit_stencil` median
183
+ filter had nowhere to put its window -- it is now nine doubles on the cell's
184
+ stack and `sort(w)`. `border: :mask` takes one, the frame being marked
185
+ before the loop runs rather than carried through it. A name the block makes
186
+ an array under is refused where the block also closes over an array of that
187
+ name, since in these spellings an assignment writes that array's cell;
188
+ rename one of the two. A `jit_map` block may not end by making an array, its
189
+ value having to fit in a cell. Still excluded when this landed: a `jit_function` body, a
190
+ contraction, more than one axis, passing one to a C function, and a kernel
191
+ that carries masks; all but the contraction arrived later in this release.
192
+
193
+ - New: `sum(w)`, `min(w)`, `max(w)` and `sort(w)` inside a `CArray.jit_for`
194
+ block, over a local array of one axis. They are bare calls, the compiler's
195
+ own names, rather than methods on the array -- `w.sum` is refused, and
196
+ `sum = 0.0` beside `sum(w)` is still a local. `sum` accumulates in the
197
+ element's computation type in index order; `min` and `max` skip a NaN
198
+ wherever it stands, and answer `NaN` for an array of nothing but NaN,
199
+ which is what `CArray#min` and `#max` answer (they answered the type's
200
+ limit until a later change in this release). `sort`
201
+ is a statement and orders the cells ascending with every NaN after every
202
+ number, as `CArray#sort` does; the relative order of `-0.0` and `0.0` is
203
+ not promised. Up to 16 cells it emits a comparator network with no branch
204
+ in it, above that an insertion sort. Excluded for now: a captured array
205
+ (a whole-array reduction is `CArray#sum`), more than one axis, two
206
+ arguments, boolean for all four, Complex for `min` / `max` / `sort`, and
207
+ the other entry points.
208
+
209
+ - New: a `CArray` made inside a `CArray.jit_for` block is a C array on the
210
+ block's stack -- `w = CArray.double(9)`, `CArray.new(:int64, [256])` or
211
+ `CArray.empty(:float64, [9])`, read and written at a subscript. One axis,
212
+ with the length written out as an integer or integers joined by `+`, `-` and
213
+ `*`; the zeroed spellings are cleared each time the line runs, as Ruby makes
214
+ a fresh array there. A subscript is checked as the block is read where the
215
+ loop's range says it can be, and at the access otherwise, raising
216
+ `IndexError`. One array is held to 4 KiB of stack and one kernel's to 16 KiB.
217
+ Excluded for now: more than one axis, passing one to a C function, the other
218
+ entry points (`jit_each`, `jit_map`, `jit_stencil`, `jit_function`), a kernel
219
+ that carries masks, and CArray's Numo/NumPy spellings (`CArray.zeros`,
220
+ `CArray::Int64.empty`), which are refused with the carray spelling named.
221
+ Note `CArray.float` is float32 and `CArray.complex` is cmplx64.
222
+
223
+ - New: `CArray#jit_init` fills an array from a formula over its indices, with
224
+ the block compiled: `CArray.int32(1000, 1000).jit_init { |i, j| (i + j) % 2 }`
225
+ is what the constructor block `CArray.int32(n, n) { |i, j| ... }` says,
226
+ without the Ruby call per cell. One parameter per axis, the block's value is
227
+ the cell, and the receiver comes back. It is the only entry point here that
228
+ is an instance method, because the array being written is the receiver: the
229
+ extents are its shape, so neither the index space nor the target is said
230
+ twice, where `CArray.jit_for(n, n) { |i, j| z[i, j] = ... }` says both. A
231
+ block outside the compilable subset raises rather than running the slow
232
+ loop, as every other entry point here does. Reach for it where the formula
233
+ will not go through whole-array arithmetic -- where it will, that needs no
234
+ compiler and is faster still.
235
+
236
+ - New: a kernel can draw random numbers, from a `CArray::Rng`. `rand =
237
+ CArray::Rng.new(seed: 4)` and then `rand.random` in a `jit_for`, `jit_each`
238
+ or `jit_map` block draws one double in `[0.0, 1.0)` per cell, at about
239
+ 0.8 ns; `random(rng: rand)` is the same draw spelled as
240
+ `CArray#random!(rng:)` spells it; `rand.randomn` is a standard normal, at
241
+ about 10 ns, and `randomn(rng: rand)` the same; and `rand.bits` is the raw
242
+ word a draw came from. All of them read one generator, so mixing them in a
243
+ kernel walks one sequence. The generator is CArray's and there is no entry point here: a
244
+ block closing over one is what a kernel needs. Each has its own state, so
245
+ two in one kernel are two sequences, and the state survives the call --
246
+ a second kernel carries on rather than starting again. `seed:` is data
247
+ rather than code, so every seed shares one compiled kernel. A sequence can
248
+ begin with `a.random!(rng: rand)` and continue in a kernel: those are the
249
+ same numbers one `random!` over both arrays would have laid down, because
250
+ CArray hands out the generator's C and this gem pastes it. Needs CArray
251
+ 3.0.2 or newer, which is where `CArray::Rng` arrived; an older one is
252
+ refused where the block closes over the generator, and leaves every kernel
253
+ that draws nothing alone. A draw is refused inside `jit_function`, which
254
+ has nowhere to keep a state; take `int64_t state[4]` as a parameter there
255
+ and pass `rand.state`. Which draw lands in which cell is still the loop's
256
+ order, so where that has to be settled, fill an array with
257
+ `CArray#random!` before the call.
258
+
259
+ - New: `CFunction#watching`, and `#clear_error` / `#report_error` beside it,
260
+ for the window in which a compiled function's address is lent to a C
261
+ library. `#call` answers for one call; a library given `#pointer` calls as
262
+ often as it likes, and `f.watching { ... }` is what puts the flag down
263
+ before that and raises what the body reported after it. A failure inside
264
+ the block outranks the library's own complaint about the stand-in it was
265
+ handed. Windows nest, and a `#call` made inside one leaves it armed.
266
+
267
+ - New: `Math.gamma`, which is not `tgamma` and is not lowered to one: Ruby's
268
+ answer is tgamma with a table of exact values in front of it for a whole
269
+ number up to 23, and a `Math::DomainError` where tgamma answers a negative
270
+ whole number or negative infinity with a NaN. Both are reproduced -- the
271
+ table is written into the generated C, filled from the Ruby that compiled
272
+ it, as `Math::PI` is emitted as the double Ruby would have used, and the
273
+ error carries Ruby's class and Ruby's words. A cell with no value in it
274
+ does not raise. Measured over the whole numbers to 26, the halves, the
275
+ infinities and the overflow, a kernel and the Ruby loop agree in every bit.
276
+
277
+ - New: `Math.erf` and `Math.erfc`, which are 1:1 with math.h's `erf` and
278
+ `erfc` -- Ruby calls those very functions, so a kernel and the Ruby loop
279
+ agree to the bit, infinities included. A float32 cell is worked on narrow
280
+ and so gets `erff`, as it gets `sinf`. There is no postfix `x.erf`:
281
+ `CArray::CoreExtensions` does not provide one, and this compiles the
282
+ refinement's names rather than inventing them.
283
+
284
+ - New: `x.clamp(low, high)`, which answers the value or whichever bound it
285
+ ran past. The value and both bounds have to be one class: Ruby hands back
286
+ the receiver in one branch and a bound in the other, so `1.clamp(0.0, 3.0)`
287
+ is an Integer and `5.clamp(0.0, 3.0)` a Float, and no type settled before
288
+ the loop runs is both -- the refusal says which way round it is and what to
289
+ write. Two widths of one class are not that case, so a float32 cell keeps
290
+ its width. What Ruby raises `ArgumentError` for is raised here too: bounds
291
+ the wrong way round, and a NaN that cannot be ordered. The class is Ruby's
292
+ and so are the words for the first; for the NaN the message names the
293
+ reason rather than the value, since what comes back from a kernel is a code
294
+ and not a number. A cell with no value in it does not raise. The range
295
+ form, `x.clamp(0.0..1.0)`, is not in the subset -- two bounds are two
296
+ bounds.
297
+
298
+ - New: a compiled function may take and return a C99 complex --
299
+ `CArray.jit_function("double _Complex step(double _Complex z)") { |z| z * z
300
+ + Complex(0.0, 1.0) }`. `double _Complex`, `float _Complex` and
301
+ `<complex.h>`'s `double complex` all read, by value or as a pointer, where
302
+ `double _Complex v[]` takes a `cmplx128` array as `double v[]` takes a
303
+ `float64` one. A kernel calls such a function as C calls it. A call from
304
+ Ruby goes through a second entry point compiled beside the body, since
305
+ Fiddle has no type to carry a complex by value in; it calls the body rather
306
+ than repeating it, so `f.call` and `f.block.call` stay the same body run
307
+ two ways. A function bound with `jit_extern` and declared with a complex is
308
+ callable from a kernel but not from Ruby -- there is no source here to
309
+ compile an entry point beside -- and says so.
310
+
311
+ - New: a `jit_function` body may take an unsigned 64-bit value by parameter --
312
+ `CArray.jit_function("size_t stride(size_t n, size_t width)") { |n, w| n * w }`
313
+ -- where before only a pointer to one could be taken. The value arrives
314
+ whole above 2^63 and the arithmetic on it is unsigned, wrapping at the width
315
+ as CArray's own `uint64` operators wrap. A kernel's captured scalars are
316
+ unchanged: those travel in the kernel's own buffers, which carry doubles,
317
+ int64s and complexes.
318
+
319
+ - New: `x.nan?` and `x.finite?`, which compile to C's `isnan` and `isfinite`
320
+ and answer what Ruby answers. `nan?` is a Float's question: an Integer and
321
+ a Complex have no method by that name and raise `NoMethodError` in Ruby, so
322
+ a kernel refuses both rather than answering false. `finite?` answers for an
323
+ Integer (true, whatever it holds) and for a Complex (both parts finite, as
324
+ Ruby asks it) as well as for a Float. `infinite?` is refused: Ruby answers
325
+ it with nil, 1 or -1 rather than true or false, and a kernel has no nil to
326
+ answer with -- the message names `x.abs == Float::INFINITY`, or
327
+ `x == Float::INFINITY` where the sign is the question.
328
+
329
+ - New: the operator assignments -- `+= -= *= /= %= **= &= |= ^= <<= >>=` --
330
+ on a local, on a cell (`out[i] += e`, `work[i, k] += e`, a scatter such as
331
+ `counts[bin[i]] += 1`), on a CScalar and through a `jit_function`'s pointer
332
+ parameter. Each is read as the assignment it stands for, `x = x + e`, so
333
+ the type rules, the mask propagation and the fold that splits an
334
+ accumulator into partial sums are the ones already there: a reduction
335
+ written `total += values[j]` is still split. `||=` and `&&=` are refused,
336
+ being about whether a value is nil or false rather than about arithmetic.
337
+
338
+ - New: an inner loop counts by a stride, written the way an extent writes
339
+ one: `(n-1).step(0, -1) { |k| ... }` is a downward sweep and
340
+ `(0...n).step(2) { |k| ... }` a stride of two. `step` includes the index it
341
+ is given, as Ruby's does, where a `...` range excludes it, and the stride
342
+ is a literal because it is what says which way the loop runs. This is the
343
+ other half of a row of workspace -- filling one and walking back down it no
344
+ longer needs `k = width - 1 - t`, which made the position a value the
345
+ kernel worked out and so put a bounds test on every cell of the sweep: 0.41
346
+ ns/cell against 0.24 for the same sweep written with `step`. An accumulator
347
+ is split into partial sums only for a loop counting by one; a stride keeps
348
+ the serial chain, and so keeps Ruby's order. `downto`, `upto` and
349
+ `reverse_each` are refused by name, naming `step` as what to write.
350
+
351
+ - New: an inner loop's index may address a write, so a cell can be given a
352
+ row of workspace -- `(0...width).each { |k| work[i, k] = ... }` fills it,
353
+ and `work[i, k]` reads it back inside the same cell. That is what an
354
+ algorithm needing a few numbers per cell is written with: a small dense
355
+ solve, a tableau, a sweep and the pass back down it, in one kernel rather
356
+ than in several that each pay a call. The cell is bounds-checked before the
357
+ kernel runs, as one addressed by `i` is. Inside the row the body may do as
358
+ it likes -- sort it, walk it backwards, write at a position it works out --
359
+ which is what a median filter needs and what no extra axis can say. Reading
360
+ an array the kernel writes through an inner index is still refused where
361
+ the read leaves the cell the outer indices picked: a read carrying an inner
362
+ index addresses every axis some write walks with an outer index with that
363
+ same index, at whatever offset, or it reaches cells another outer iteration
364
+ owns. Where the sweep can be said as an extra axis instead --
365
+ `jit_for(rows, 1...width)` -- that remains the faster form once the rows
366
+ are long.
367
+
368
+ - New: a function compiled with `CArray.jit_function` may call another one
369
+ compiled with `CArray.jit_function`, by the name the block reaches it by --
370
+ `hypot = CArray.jit_function("double (*)(double, double)") { |a, b|
371
+ root.call(a * a + b * b) }`. The called body is pasted into the caller's C
372
+ and reached by symbol, so what comes back is still one self-contained
373
+ object with one address, and a chain of any depth arrives together with the
374
+ messages its bodies raise. Everything else a body closes over is refused as
375
+ before -- a number, an array, and a function bound with `jit_extern`, which
376
+ is only an address and has nowhere in a compiled object to live. The name
377
+ may be a constant as well as a local, which is what lets a method reach
378
+ one: `def` closes over nothing.
379
+
380
+ - New: `CArray::JIT.contraction_of` returns the number that multiplies the
381
+ product as `:scale`, which is 1 where there is none, so
382
+ `a[i,k] * b[k,j] * 2.0` comes back as its two terms and 2.0 rather than as
383
+ nil. A number is not a term -- it has no indices and no cell -- but it is
384
+ not a reason to give up on the terms either, and a caller that takes them
385
+ apart puts it back. Written out or closed over is the same number; anything
386
+ a name holds that is not a Numeric is still nil. This shipped in 0.1.2,
387
+ whose entry describes the same two methods without mentioning it.
388
+
389
+ - Change: this gem no longer builds a C extension, so installing it needs no
390
+ compiler for itself -- one is still needed at run time to compile a kernel.
391
+ The addressing it used to carry is `CArray::AddressBasis`, in carray, which
392
+ is where the knowledge about carray's views belonged; `CArray::JIT::Access`
393
+ remains as an alias for this release. carray 3.0.2 or newer is required,
394
+ which the gemspec already asked for.
395
+
396
+ - Change: `CArray.jit_for`, `CArray.jit_each` and `CArray.jit_map` are this
397
+ gem's alone. From CArray 3.0.2 they are not defined until
398
+ `require "carray/jit"` has run, so a program that calls one without it gets
399
+ `NoMethodError` rather than CArray's `NotImplementedError`. With the gem
400
+ required nothing changes, against any CArray this gem accepts.
401
+
402
+ - Change: `CArray::JIT::CFunction#name` answers the name the declaration gave
403
+ -- `:square` for `CArray.jit_function("double square(double x)")` -- and
404
+ the new `#symbol` answers the name in the object, which for a compiled
405
+ function carries a digest of the body. It used to answer whichever of the
406
+ two its constructor was handed: the declared name for a function bound with
407
+ `jit_extern`, the symbol for one compiled from a block. Code reading
408
+ `#name` to reach or paste a function wants `#symbol`; code putting it in a
409
+ message wants `#name`, which is `nil` where the declaration named no
410
+ function. `#to_s` now prints the declaration rather than the symbol.
411
+
412
+ - Change: a compiled function whose body has failed does no more work until
413
+ its flag is put down again -- it returns 0 without running, and one
414
+ declared `void` leaves its out-parameters alone. Before, it reported the
415
+ first failure and then answered normally, so a library that kept calling
416
+ could converge on values it was no longer entitled to. Nothing changes for
417
+ a caller using `#call`, which puts the flag down for each call; a caller
418
+ holding `#pointer` opens a window with `#watching` or `#clear_error`. A
419
+ kernel is unaffected: it is handed its own error slot and never reaches
420
+ this flag.
421
+
422
+ - Change: naming a contraction's axes now replaces the convention rather than
423
+ adding a clause to it. `CArray.jit_contract(:i, :j) { ... }` names all of
424
+ the result's axes, so every index left out of the list is summed at however
425
+ few positions it sits -- where before, one sitting at a single position was
426
+ refused as a free index with nowhere to go. So `jit_contract(:i) { |k|
427
+ a[i,k] }` is the row sums, and `contract_terms(terms, free: [])` over a
428
+ product is its total; both were refused. (`a.sum(axis: 1)` remains the
429
+ faster way to write a reduction, being one.) With nothing named the
430
+ convention is unchanged: a repetition is a sum and a single position is
431
+ free. What this makes exact is the correspondence with einsum's two modes,
432
+ the argument list being the arrow's right-hand side; `"ik->i"` and
433
+ `"ik,kj->"` can now be said.
434
+
435
+ Where the block assigns into an array of yours, an axis on the left-hand
436
+ side that the list leaves out is still refused -- it is the list falling
437
+ short of the result rather than an index to sum, and the left-hand side is
438
+ the one place that can be seen. The message says so in those terms now.
439
+
440
+ - Change: `min(w)` and `max(w)` over a floating local array answer `NaN` when
441
+ every cell is `NaN`, where they answered `Infinity` and `-Infinity`. This
442
+ follows CArray 3.0.2, which made the same change to its own `min` and `max`;
443
+ the gemspec already asks for that version. An array holding at least one
444
+ number answers as before, a `NaN` still losing to any number, and an integer
445
+ local array is unchanged.
446
+
447
+ - Change: `CArray.jit_contract` refuses an index that stands inside a
448
+ subscript the kernel works out -- `a[i, idx[k]] * v[k]` -- saying that a
449
+ contraction counts positions and this one is a position of `idx`. It was
450
+ read as an index appearing once, so the sum the notation asks for did not
451
+ happen and the result came back a whole matrix. Write such a gather with
452
+ `CArray.jit_for`.
453
+
454
+ - Change: an indexed kernel (`CArray.jit_for` and the rest) refuses an
455
+ operand it cannot walk -- a gather, a lazy array -- that is a view of an
456
+ array the same kernel writes, naming `a[order[i]]` as the spelling that
457
+ means something. Such an operand is copied before the loop and put back
458
+ after it, so the loop read cells that had stopped being current and a
459
+ direct write to the same array was dropped at the end.
460
+
461
+ - Change: arithmetic between two booleans -- `flag[i] * flag[i]`, `+`, `-`,
462
+ `/`, `%` -- raises `CArray::JIT::Unsupported` where the operator stands,
463
+ as `true * true` raises in Ruby. Used as a condition it compiled and ran.
464
+ `&`, `|` and `^` on booleans are unchanged.
465
+
466
+ - Change: a captured Integer above 2**63-1 meeting an integer is refused,
467
+ where before it was absorbed into that integer's width and computed from a
468
+ wrapped value. CArray refuses the same expression whatever the array's own
469
+ type is -- `CArray.uint64(1) { 5 } + 2**63` raises `bignum too big to
470
+ convert into 'long long'` -- and this follows it: a bare Integer brings no
471
+ width, and the message names `CScalar.uint64() { big }`, which a kernel reads
472
+ as the one-cell array it is. A Float or a Complex on the other side is
473
+ unaffected, and so is a capture that fits an int64. A loop that adds such a
474
+ capture to an accumulator is refused for the same reason: state the width
475
+ once with a CScalar, for the seed and for the value.
476
+
477
+ - Change: a `CArray.jit_each` or `CArray.jit_map` block whose *only* array is
478
+ handed to a C function whole is now refused, naming the array and saying
479
+ that it does not settle how many cells there are to compute.
480
+ `CArray.jit_map { DOT4.call(w, w) }` used to answer an array as long as `w`
481
+ -- a length that came from an array nobody walks -- and the same block under
482
+ `jit_each` failed inside Fiddle with `unknown symbol "ca_call_cslab_0_r"`.
483
+ `CArray.jit_for` with a count takes these, as it always did.
484
+
485
+ - Change: a block parameter that names a loop index is refused when the
486
+ generated C already uses that name for a captured array (`p_a`, `m_a`,
487
+ `a_s0`, `a_ms0` or `a_n0` for an array `a`), or when it contains `__` or
488
+ starts with `carray_jit_`. Rename the index.
489
+
490
+ - Change: a `CArray.jit_function` body that subscripts a pointer parameter
491
+ with a literal outside the length its declaration gave -- `v[7]` or
492
+ `v[-1]` against `double v[2]` -- raises `CArray::JIT::Unsupported` as the
493
+ body is read. It compiled, and wrote or read past what the caller was
494
+ held to. A computed subscript, and any subscript on a pointer declared
495
+ without a length, are unchecked as before.
496
+
497
+ - Change: an extent that counts down as a `Range` -- `(n-2)..0` -- raises
498
+ `CArray::JIT::Unsupported` naming `(n-2).step(0, -1)`, the spelling Ruby
499
+ iterates backwards with. Ruby gives such a Range no elements, so the
500
+ kernel ran no passes and said nothing, which is what the guide says this
501
+ refusal exists to prevent. An empty range whose ends agree (`0...0`) is a
502
+ count of zero as before.
503
+
504
+ - Change: a body compiled with `CArray.jit_function` or `CArray.jit_call` may
505
+ call a function bound with `jit_extern`. It was refused -- a borrowed
506
+ function is only an address, and a compiled body has no `functions` buffer
507
+ to take one in -- but an address is not what a body needs from it. It needs
508
+ the name, which is what the declaration states and what a linker or a
509
+ loader resolves; `jit_extern` opened the library to find the function, so
510
+ the symbol is in the process by the time the body is compiled. The
511
+ generated file declares it (`double j0(double);`) and calls it by name. A
512
+ kernel is unchanged and still takes the address through its buffer, which
513
+ is what a kernel has and a body has not. A file built ahead of the program
514
+ links the library the usual way: one that is not linked by default is the
515
+ caller's to add.
516
+
517
+ - Change: `require "carray/jit/call"` defines `CArray.jit_call` without
518
+ loading the compiler, and the compiler is required at the first call site
519
+ no provider answers. A program whose sites were all built ahead of it --
520
+ by carray-jit-aot -- never loads the analyzer, the generator or the cache
521
+ at all, where before it loaded them and reached none of them.
522
+ `require "carray/jit"` loads everything as it did, and a program that
523
+ knows nothing about this sees no difference.
524
+
525
+ - Change: a subscript that walks with one index and adds another --
526
+ `a[i + r]`, `r` an inner loop's index -- is refused as the block is read,
527
+ with a message that says so and that the sum put in a local first
528
+ (`k = i + r`, then `a[k]`) is checked at each cell and runs. It was
529
+ refused at the call, for a reason about an inner loop's range. The
530
+ exception is `CArray::JIT::Unsupported`, as it was.
531
+
532
+ - Change: a kernel or `jit_function` body that reads a local after an inner
533
+ loop's block in which the local was first assigned, or after a `while` in
534
+ which it was first assigned, is refused with a message naming the local.
535
+ These did not compile before either, failing in the C compiler instead.
536
+ Give the local a value before the loop.
537
+
538
+ - Change: a local assigned before the summand of a `CArray.jit_contract`
539
+ block is refused saying so -- a contraction is one expression, which is
540
+ why it takes no local array either. It was refused as "`u` is read before
541
+ it is assigned", about a local assigned on the line above.
542
+
543
+ - Change: the refusal for a `CArray.jit_stencil` writing into a view of an
544
+ array it reads says that the two are views of one array and that the test
545
+ is the storage rather than the cells, so two slabs that share none are
546
+ refused too, and names the copy to read from. It said a stencil writes
547
+ into an array of its own, which of two slabs is not true.
548
+
549
+ - Change: the refusal for a block whose source cannot be read names
550
+ `RubyVM.keep_script_lines = true` and `CArray::JIT.compile`, which takes a
551
+ kernel as text. It named `source:`, which is a keyword no entry point
552
+ takes; the guide said the same.
553
+
554
+ - Change: `CArray#jit_init` given a block that takes a splat or an optional
555
+ parameter says so, where it said "the block names -1 indices".
556
+
557
+ - Change: a keyword in value position is named as it was written --
558
+ `unless` in `CArray.jit_map { unless a > 1.0 then ... end }` says
559
+ "`unless` -- write it as `if` with the condition negated", where it said
560
+ "unsupported expression Unless". Statement position already did.
561
+
562
+ - Change: a statement outside the subset is named as it was written --
563
+ ``got `unless` -- write it as `if` with the condition negated`` rather than
564
+ `got Unless`, which was this compiler's reading of it and not anything
565
+ anyone typed. The list of what a body may hold was written out in two
566
+ places and they had come apart; it is one place now.
567
+
568
+ - Change: the three names left in Ruby's `Math` say why they are not lowered
569
+ rather than saying that C has no counterpart, which was untrue of all of
570
+ them -- `lgamma`, `frexp` and `ldexp` all exist. `Math.lgamma` and
571
+ `Math.frexp` each answer with a pair, and a cell holds one number;
572
+ `Math.ldexp` takes an exponent where a math call here computes every
573
+ argument in the type of its result. A name Ruby's `Math` does not have says
574
+ that instead of guessing at a reason.
575
+
576
+ - Change: a compiler this process cannot find -- `CARRAY_JIT_CC` naming
577
+ something that is not there -- raises `CArray::JIT::CompilationError`
578
+ saying so and naming the variable, where it raised `Errno::ENOENT` from
579
+ the spawn with the program name and nothing else.
580
+
581
+ - Change: a cache directory owned by another user is refused, as one other
582
+ users can write to already was; everything in it is `dlopen`ed. A cache
583
+ root that is a symlink is followed as before, and the owner of the
584
+ directory it lands on is what is asked about. Set `CARRAY_JIT_CACHE` to a
585
+ directory of your own where this refuses.
586
+
587
+ - Change: the docs now say what a compare-exchange costs under a mask. A
588
+ swap writes two cells where an ordinary assignment writes one, and a sort
589
+ runs one over every cell repeatedly, so a single missing cell spreads
590
+ across the row -- the ordinary rule for a branch decided by a missing
591
+ cell, at an unusually high gain. A sort written by hand is not refused
592
+ the way `sort` on a local array is, so ask `a[i] == UNDEF` first and
593
+ decide it.
594
+
595
+ - Change: the docs now say what a branch not taken costs under a mask. A
596
+ branch decided by a missing cell masks what it writes; taking the other
597
+ path writes nothing, so a conclusion reached that way -- `found` left at
598
+ -1 for a row whose only candidate was masked -- is reported as an ordinary
599
+ value. Ask about the mask with `a[i] == UNDEF` where that matters.
600
+
601
+ - Change: the docs now say that a negative Float to a fractional power is
602
+ `NaN`, as `pow` and CArray answer, where Ruby answers a Complex. This was
603
+ already the behaviour; a whole-number exponent or a non-negative base is
604
+ Ruby's number.
605
+
606
+ - Change: the docs now say that an Integer compared with a Float is
607
+ compared as two doubles, as CArray compares them, so past 2^53 the answer
608
+ can differ from Ruby's exact comparison. This was already the behaviour;
609
+ below 2^53 nothing differs.
610
+
611
+ - Change: the docs now say that `Math.sqrt`, `log`, `log2`, `log10`, `asin`,
612
+ `acos`, `acosh` and `atanh` answer `NaN` outside their domain, as `math.h`
613
+ does, where Ruby raises `Math::DomainError`, and that `Math.sqrt(-0.0)` is
614
+ `-0.0`. This was already the behaviour. `Math.gamma` still raises as Ruby
615
+ does. Test the argument in the block where it may leave the domain.
616
+
617
+ - Fix: a loop whose bound is an unsigned parameter -- `size_t n` and its kin
618
+ -- runs the passes Ruby runs. The counter is an `int64_t` and the bound was
619
+ compared against it uncast, so C converted the *counter* to unsigned
620
+ instead: `(-3...n).each` over `n = 4` ran no passes where Ruby runs seven,
621
+ and answered without saying anything. A count that never goes below zero
622
+ was unaffected, which is every `n.times`. The cast is written down now, so
623
+ the generated C also compiles clean where it carried `-Wsign-compare`
624
+ before.
625
+
626
+ - Fix: a `CArray.jit_each` or `CArray.jit_map` pass that reads a view of an
627
+ array it also writes -- two blocks that share cells, a transpose written
628
+ over itself, a gather held in `g` and written back as `src = g + 1.0` --
629
+ reads the values the expression started with, as Ruby and CArray's own
630
+ operators do. Such an operand was read as the pass went, so it came back
631
+ holding the kernel's own output: `hi = lo * 2.0` over overlapping blocks
632
+ gave all zeros, and a gather over 20000 cells was wrong in 8192 of them.
633
+ It is copied once before the loop instead. `a = a + 1.0`, and views that
634
+ share a root without sharing a cell, are untouched.
635
+
636
+ - Fix: `CArray.jit_contract` takes a summand that calls a function made with
637
+ `CArray.jit_function` or `CArray.jit_extern`, which the docs say it does;
638
+ it was refused as an unsupported method. What the function hands back
639
+ types the result the contraction is collected into.
640
+
641
+ - Fix: a cell of a local array written under an `if` or a `while` whose
642
+ condition read a missing cell is masked, as a plain local and a cell of an
643
+ array now are. It was left present deliberately; the rule the docs give
644
+ for it is the one the other two keep.
645
+
646
+ - Fix: a local assigned under an `if` or a `while` whose condition read a
647
+ missing cell is masked, as a cell written there already was. It came back
648
+ a number like any other, so `found = -1; ...; if a[i, j] > x then found =
649
+ j end; out[i] = found` reported a position found under a mask. A condition
650
+ that asks about the mask itself (`a[i] == UNDEF`) still masks nothing.
651
+
652
+ - Fix: a build no longer fails with "could not load freshly compiled" when
653
+ another process evicts its object in the moment between the build and the
654
+ load. It opens the object before publishing it to the cache. Reachable
655
+ where `CARRAY_JIT_CACHE_LIMIT` is smaller than the number of processes
656
+ building at once: with five of them, a limit of 2 lost every one.
657
+
658
+ - Fix: a build whose compile fails leaves no `.c` behind in the cache. Such
659
+ a file was never looked up -- a cache hit is an object -- and eviction
660
+ passes over it, so one stayed for good, and one more for every kernel of
661
+ every run against a toolchain that cannot compile.
662
+
663
+ - Fix: two threads of one process compiling at the same time no longer
664
+ raise `Errno::ENOENT`; three threads in four did. A build now takes a lock
665
+ for the length of the compile, so the second thread finds the object the
666
+ first one left and reuses it rather than building its own. Compiling two
667
+ different kernels in two threads is that much less parallel -- four of
668
+ them took 970 ms here against 800.
669
+
670
+ - Fix: multiplying two Complex numbers, or a real number by a Complex
671
+ (`x * z`), gives Ruby's answer where an infinity meets a zero:
672
+ `0.0 * Complex(Float::INFINITY, 0.0)` is `0.0+0.0i`, where it was
673
+ `NaN+NaN*i`. Finite products are unchanged, `cmplx64` included, and
674
+ `z * x` still scales each part as it did.
675
+
676
+ - Fix: `%` by a float zero raises `ZeroDivisionError`, as Ruby's does --
677
+ `x % 0.0`, `x % -0.0`, an integer cell `% 0.0` -- in a kernel and in a
678
+ `CArray.jit_function` alike. It answered `NaN`, which is what CArray's `%`
679
+ answers and not what the docs promised. A float `/` by zero is still an
680
+ infinity, as it is in Ruby.
681
+
682
+ - Fix: `min(w)` and `max(w)` over a floating local array keep the first of
683
+ two cells that compare equal, as CArray's `min` and `max` do, so `0.0`
684
+ and `-0.0` come back in the order they stood; which zero came back was
685
+ left to the C library. NaN is still skipped, and an array of nothing but
686
+ NaN still answers `NaN`.
687
+
688
+ - Fix: `floor`, `ceil`, `round`, `truncate` and `to_i` on a Float raise
689
+ `FloatDomainError` for a NaN or an infinity, as Ruby does, and `RangeError`
690
+ for a result past int64; they gave a clamped or arbitrary number. Written
691
+ straight into a float cell, or returned from a `CArray.jit_function`
692
+ declared `double`, the result is Ruby's however large -- `1e20.floor` is
693
+ `1e20`. A loop doing little but rounding runs slower for the check.
694
+
695
+ - Fix: `.abs` on an integer compiles in `CArray.jit_for`, `CArray.jit_map`,
696
+ `CArray.jit_function` and the other entry points; it raised
697
+ `CArray::JIT::CompilationError` unless the kernel also allocated a local
698
+ array. `(-2**63).abs` in an int64 wraps to itself, as other int64
699
+ overflow does. Float and complex `.abs` are unchanged.
700
+
701
+ - Fix: an inner loop stepping by more than one over a start written over
702
+ another index -- `2.times { |p| (p...8).step(3) { |k| ... } }` -- is held
703
+ to the cells every pass reaches, not only the pass from the earliest
704
+ start. A subscript one pass took past the end of an array was let through
705
+ and written; it is now refused as the call is prepared. A step of one, and
706
+ a start that does not move, are read as before.
707
+
708
+ - Fix: a local array whose shape is written over captured integers --
709
+ `CArray.double(n, m)` -- raises `ArgumentError` when its lengths each fit
710
+ but their product does not count in bytes. Such a product wrapped to a
711
+ small number, the kernel allocated that, and subscripts in range on every
712
+ axis wrote past it. A shape of one axis, or of lengths whose product
713
+ fits, is unaffected.
714
+
715
+ - Fix: `CArray::JIT.clear_registry` now forgets the functions
716
+ `CArray.jit_function` compiled, as it already did kernels, so the next
717
+ call reads them back from the cache on disk.
718
+
719
+ - Fix: a block run through `eval` -- in a console, or from code that builds
720
+ its kernels as text -- no longer stays in memory for the life of the
721
+ process once nothing refers to it. On Ruby 3.2 it still does.
722
+
723
+ - Fix: a kernel over a masked array read and wrote the wrong cells of its
724
+ mask, and on a large enough array wrote past the end of it, when an
725
+ unmasked operand whose number of axes differs from the kernel's -- a row
726
+ added across a grid, or a contraction's operand -- sorted by name before
727
+ the masked one. Kernels whose operands all share the kernel's rank, and
728
+ kernels over no masked array, were not affected.
729
+
730
+ - Fix: a local array indexed by an inner loop whose range is not known until
731
+ the kernel runs -- a bound that is a captured integer, and now a range
732
+ written over another index -- is checked where the cell is reached. Such a
733
+ subscript was checked nowhere: the check that reads the loop's range could
734
+ not settle it, and the check at the access did not cover an index, so a
735
+ write past the end of the array went into the block's stack frame. It now
736
+ raises `IndexError` as every other unsettled subscript does.
737
+
738
+ - Fix: in a `CArray.jit_each`, `CArray.jit_map` or `CArray.jit_stencil`
739
+ block, an array handed to a C function whole -- and read in no other way --
740
+ is no longer lined up with the operands. One whose length differed from
741
+ theirs used to fail with `broadcast_to: cannot broadcast axis 0`, which
742
+ named an axis and said nothing about the call it was written for; a set of
743
+ weights four cells long may now stand beside a thousand cells of operand.
744
+ What decides which arrays those are is the declaration, not the spelling: a
745
+ parameter taking a number by value reads the cell, so that array is walked
746
+ and lines up as before, and an array both read by cell and handed over is
747
+ an operand too.
748
+
749
+ - Fix: a local in a kernel or `jit_function` body, or a captured variable
750
+ starting `carray_jit_`, may be given a name the generated C also uses -- a
751
+ kernel parameter such as `error`, a C keyword such as `int`, `a_n0` beside a
752
+ captured array `a`, or a name containing `__`. It failed to compile, raised
753
+ an `IndexError` for an index in range, or shared its value with another
754
+ name; it is now renamed in the generated C only.
755
+
756
+ - Fix: two inner loops in one kernel or `jit_function` body may assign a
757
+ local of the same name, at one type or two, and a local assigned inside a
758
+ `while` may be read after it when it was also assigned before it. The
759
+ first failed in the C compiler or was refused as changing type; the second
760
+ failed in the C compiler.
761
+
762
+ - Fix: a loop in a `jit_function` body, and a `while` or inner loop in a
763
+ kernel whose body held the kernel's first way to fail, kept running after
764
+ an integer division by zero, a computed index out of range, `clamp` or
765
+ `Math.gamma` had reported a failure -- so a `while` decided by the value
766
+ handed back could run forever. Such a loop now leaves at the head of its
767
+ next pass, and the call raises as it already did.
768
+
769
+ - Fix: `CArray.jit_each` and `CArray.jit_map` no longer refuse a block that
770
+ hands a one-cell array to a C function. Those entries line their operands up
771
+ with the expression's shape, and an array passed to a pointer parameter was
772
+ lined up with the rest -- so a one-cell array holding state came back
773
+ stretched and read-only, and the copy-back after the call raised `can not
774
+ modify read-only array`. An array handed over by address is passed whole
775
+ rather than walked, so it is left as it is now. `jit_for` was never
776
+ affected, and a `CScalar` was affected the same way an array was.
777
+
778
+ - Fix: a captured Integer above 2**63-1 reaches a kernel whole. It was packed
779
+ into the int64 slot the kernel reads its integers from, which took the value
780
+ modulo the width and said nothing: `big / 3`, `big > 100` and `big * 1.0`
781
+ answered from a negative number, while `+` and `*` came out right and hid it.
782
+ Such a capture is a uint64 now -- the width CArray has for those values --
783
+ and a value neither an int64 nor a uint64 holds is refused where the capture
784
+ is read, naming the value.
785
+
786
+ - Fix: a C declaration written with `ptrdiff_t` compiles. `size_t` and the
787
+ other `<stddef.h>` names have read since 0.1.0, but the C generated for a
788
+ body that took one declared no such type, so `CArray.jit_function("ptrdiff_t
789
+ (*)(ptrdiff_t)")` failed in the C compiler with `unknown type name`.
790
+ `size_t` was unaffected, by luck rather than by design.
791
+
792
+ - Fix: two inner loops in one body may both be written `{ |k| ... }`. They
793
+ are one name in the block and were one index here, so the second loop's
794
+ range replaced the first's and a reach was checked against the wrong one --
795
+ `a[k-1]` in a loop from 1 was refused for starting at 0 once a later loop
796
+ started there. Each loop now counts in an identifier of its own, and the
797
+ messages go on speaking the name the block wrote.
798
+
799
+ - Fix: a block holding a character outside ASCII no longer raises. The file a
800
+ block sits in was read with `File.read`, which uses
801
+ `Encoding.default_external` -- a setting that has nothing to do with a
802
+ source's encoding: on a machine with no locale set it is US-ASCII, and the
803
+ file came back as its own bytes under a tag that the first `rstrip` on a
804
+ line holding a comment in Japanese raised on. The file is now read as bytes
805
+ and given the encoding the parser gave it: UTF-8, or what a `coding` magic
806
+ comment names on the first line or on the second where a shebang takes the
807
+ first.
808
+
809
+
810
+ ## 0.1.2
811
+
812
+ - New: `CArray.jit_contract` takes the result's axes as symbols, which says
813
+ which indices are free: `CArray.jit_contract(:p) { |k| x[p,k] * y[p,k] }` is
814
+ one number per point, and `CArray.jit_contract(:a) { q[a,a] }` is the
815
+ diagonal rather than the trace. A named index stays free however often it
816
+ appears, which is what an index that numbers things -- a point, a sample, a
817
+ batch -- does. What a repetition means is unchanged, so the whole rule is
818
+ that an index which repeats is summed and one that is named is free. Naming
819
+ the axes also states their order. With no arguments nothing changes.
820
+
821
+ - New: `CArray::JIT.contract_terms` runs the contraction a structure describes
822
+ rather than one a block writes --
823
+ `CArray::JIT.contract_terms([[a, [:i, :k]], [b, [:k, :j]]], free: [:i, :j])`
824
+ is a matrix product -- and `CArray::JIT.contraction_of` reads a block and
825
+ returns the terms it is a product of, or nil when it is not one. They are
826
+ `jit_contract` with the block taken out of the middle, for a caller that
827
+ rearranges a contraction before running it: the terms are compiled by the
828
+ same analyzer under the same rules, so `free:` is required and an index that
829
+ is not named must appear at more than one position.
830
+
831
+ - New: `CArray::JIT.cache_root = "path"` puts an application's compiled
832
+ kernels somewhere of its own, rather than in the cache shared under the home
833
+ directory. Say it before the first kernel is compiled; the path is expanded
834
+ where it is given. `CARRAY_JIT_CACHE` and `CARRAY_JIT_NO_CACHE` still come
835
+ first, and `nil` restores the default.
836
+
837
+ - Change: a contraction's sum is split into partial sums, as `jit_for`'s
838
+ reduction and CArray's own reduce kernels are, which makes `jit_contract` as
839
+ fast as the same loop written with `jit_for` rather than three times slower.
840
+ A floating-point contraction therefore answers what the split accumulation
841
+ answers -- usually the more accurate number, never the one a serial Ruby
842
+ loop gives. `CArray::JIT.reassociate = false`, or `CARRAY_JIT_REASSOCIATE=0`
843
+ for a whole process, asks for the serial order; `jit_contract` takes no
844
+ per-call licence. Integer contractions are unaffected.
845
+
846
+ - Change: `CArray.jit_contract` sums an index that repeats however often it
847
+ repeats, rather than refusing more than two positions. `q[i,i,i]` is the sum
848
+ along a cube's long diagonal, and `a[i,k] * b[k,k]` sums `k` at three
849
+ positions across two arrays. Nothing that compiled before compiles
850
+ differently: what changes is that these are accepted instead of raising
851
+ `CArray::JIT::Unsupported`.
852
+
853
+ - Fix: a contraction with nothing to assign into is collected into the type its
854
+ summand computes in. `CArray.jit_contract { |i, j, k| a[i,k] * b[k,j] }` over
855
+ float32 arrays came back int64 with every value truncated; over uint64 it came
856
+ back int64 and wrapped; over cmplx64 it raised. The assigned form -- the same
857
+ contraction written `c[i,j] = ...` -- was right throughout, and `jit_for`,
858
+ `jit_each` and `jit_stencil` were never affected.
859
+
860
+ - Fix: a contraction that writes an array it also reads is refused when the
861
+ block gave that array two names -- `y = x`, or a view of something being
862
+ read -- as it always was when one name was used for both. It compiled and
863
+ returned an answer that depended on the order the cells were reached in.
864
+ `CArray::JIT.contract_terms` refuses `into:` for the same reason. Write it
865
+ with `jit_for`, which is what a recurrence is for.
866
+
867
+ - Fix: `CArray.jit_stencil(source, into: source)` is refused rather than
868
+ computing a pass whose cells feed the ones after them. A window that reaches
869
+ nowhere -- one that reads only the cell it is on -- still writes in place, as
870
+ it always did.
871
+
872
+ - Fix: assigning to a loop index inside a kernel is refused. `jit_for(3) { |i|
873
+ i = 2; out[i] = ... }` assigned to the counter, so the loop walked somewhere
874
+ else -- outside the array, for a value outside its extent -- while Ruby reads
875
+ the same line as rebinding the parameter and runs the loop unchanged. Use a
876
+ local of another name; nothing that has an index only on the right changes.
877
+
878
+ - Fix: an index named after something the generated C already uses is refused
879
+ where it is written rather than by the compiler. `jit_contract { |int, j, k|
880
+ ... }` reached clang as a declaration of `int`, and an index named
881
+ `contraction` shared its identifier with the accumulator the compiler writes,
882
+ which made the sum come out zero with nothing said.
883
+
884
+ - Fix: a C function may take a `uint64_t *` and return a `uint64_t`.
885
+ `CArray.jit_function("void (*)(uint64_t *, int64_t)")` reached its cells as
886
+ an opaque slot rather than an array, and a `uint64_t` return type was
887
+ refused as "no value a compiled body can produce" -- from a table written
888
+ before uint64 was a type a kernel computes in.
889
+
890
+ - Fix: `CArray.jit_map` collects a cmplx64 value into a cmplx64 array rather
891
+ than refusing to allocate one. Nothing else changes type: a block whose value
892
+ is cmplx128 still gives cmplx128.
893
+
40
894
  ## 0.1.1
41
895
 
42
896
  - Fix: a zero divisor in a `CArray.fuse` expression no longer turns