carray-jit 0.1.2 → 0.1.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (44) hide show
  1. checksums.yaml +4 -4
  2. data/CHANGELOG.md +771 -3
  3. data/README.md +7 -6
  4. data/carray-jit.gemspec +1 -3
  5. data/docs/00_Introduction.md +4 -3
  6. data/docs/01_GettingStarted.md +1 -1
  7. data/docs/02_KernelShapes.md +93 -14
  8. data/docs/03_SupportedFeatures.md +582 -26
  9. data/docs/04_Compiling.md +33 -6
  10. data/docs/05_DesignNotes.md +3 -3
  11. data/docs/06_Cheatsheet.md +198 -5
  12. data/docs/07_StepByStep.ja.md +534 -0
  13. data/docs/07_StepByStep.md +535 -0
  14. data/examples/README.md +12 -0
  15. data/examples/applications/alarm.rb +121 -0
  16. data/examples/applications/collatz.rb +105 -0
  17. data/examples/applications/cubic_spline.rb +331 -0
  18. data/examples/applications/dithering.rb +144 -0
  19. data/examples/applications/group_stats.rb +115 -0
  20. data/examples/applications/lookup.rb +126 -0
  21. data/examples/applications/median_filter.rb +153 -0
  22. data/examples/applications/parcel_ascent.rb +220 -0
  23. data/examples/applications/point_in_polygon.rb +111 -0
  24. data/examples/applications/random_walk.rb +98 -0
  25. data/examples/applications/van_der_pol.rb +186 -0
  26. data/examples/applications/wet_bulb.rb +140 -0
  27. data/examples/features/10_complex.rb +14 -4
  28. data/examples/features/15_loops.rb +7 -1
  29. data/lib/carray/jit/access.rb +14 -0
  30. data/lib/carray/jit/analyzer.rb +2077 -136
  31. data/lib/carray/jit/block_reader.rb +37 -6
  32. data/lib/carray/jit/c_function.rb +613 -76
  33. data/lib/carray/jit/c_generator.rb +1595 -156
  34. data/lib/carray/jit/call.rb +68 -0
  35. data/lib/carray/jit/compiler.rb +75 -11
  36. data/lib/carray/jit/kernel.rb +369 -32
  37. data/lib/carray/jit/node.rb +359 -9
  38. data/lib/carray/jit/sorting_networks.rb +182 -0
  39. data/lib/carray/jit/type_assignment.rb +371 -34
  40. data/lib/carray/jit/version.rb +1 -1
  41. data/lib/carray/jit.rb +560 -64
  42. metadata +22 -8
  43. data/ext/carray_jit_access/carray_jit_access.c +0 -460
  44. data/ext/carray_jit_access/extconf.rb +0 -8
data/README.md CHANGED
@@ -8,13 +8,13 @@ The block is read with Prism, translated to C if it falls inside that subset, co
8
8
 
9
9
  ## Status
10
10
 
11
- 0.1.2 is the current release, and it still moves: behaviour can change between releases — see [CHANGELOG.md](CHANGELOG.md). A companion gem to CArray, it follows CArray's surface, which is not settled until CArray 3.1.
11
+ 0.1.3 is the current release, and it still moves: behaviour can change between releases — see [CHANGELOG.md](CHANGELOG.md). A companion gem to CArray, it follows CArray's surface, which is not settled until CArray 3.1.
12
12
 
13
13
  ## Features
14
14
 
15
15
  - **A JIT compiler for C-level loops.** A block becomes one C function over CArray's own memory, built by the system C compiler and called through Fiddle.
16
16
  - **Ordinary Ruby, and enough of it.** The source is parsed with Prism -- no DSL, no `eval` -- and every operation means what Ruby means by it, apart from the order a reduction takes its terms in. The subset is enough to state a numerical algorithm; what falls outside it is refused by name and line, not run as a Ruby loop.
17
- - **A method for each shape.** `jit_for` for recurrences and loops written out, `jit_stencil` for windows at any rank, `CArray.jit_contract` for contraction over a repeated index, or over the indices left when the result's axes are named, `jit_each` and `jit_map` for a pass that reaches no neighbour.
17
+ - **A method for each shape.** `jit_for` for recurrences and loops written out, `jit_stencil` for windows at any rank, `CArray.jit_contract` for contraction over a repeated index, or over every index the result's axes do not name, `jit_each` and `jit_map` for a pass that reaches no neighbour, `jit_init` for filling an array from its own indices.
18
18
  - **View- and mask-aware.** Columns, transposes and slices of slices are written in place without a copy, and masks propagate as CArray propagates them.
19
19
  - **Pure C functions, in and out.** `jit_extern` binds one from a library and a kernel calls it by address; `jit_function` compiles one from a block and hands back a C function pointer.
20
20
  - **The backend for `CArray.fuse`.** An array expression compiles instead of being walked a node at a time, without being asked and without changing the answer.
@@ -35,7 +35,7 @@ gem "carray-jit"
35
35
  Requires:
36
36
 
37
37
  - Ruby >= 3.2
38
- - CArray >= 3.0.1, < 3.1
38
+ - CArray >= 3.0.2, < 3.1
39
39
  - A C compiler
40
40
  - Prism and Fiddle (both ship with Ruby; Fiddle is a bundled gem)
41
41
 
@@ -61,7 +61,7 @@ legendre[0..5].to_a
61
61
  # => [1.0, 0.5, -0.125, -0.4375, -0.2890625, 0.08984375]
62
62
  ```
63
63
 
64
- `CArray.jit_for`, `CArray.jit_each` and `CArray.jit_map` are CArray's own names: without this gem they raise and point at `CArray.fuse`, which computes an array expression without a compiler. This gem is the compiler. An expression over whole arrays wants `CArray.fuse` and not these, and gets compiled anyway where this gem is installed.
64
+ The `jit_` methods are this gem's, and exist once `require "carray/jit"` has run. An expression over whole arrays wants `CArray.fuse`, which is CArray's own and needs no compiler -- and gets compiled anyway where this gem is loaded.
65
65
 
66
66
  ## Documentation
67
67
 
@@ -71,7 +71,8 @@ legendre[0..5].to_a
71
71
  * [Supported features](docs/03_SupportedFeatures.md) — locals and types, branches, raising, the types that are not just a number, calling C, and the recognized subset with what it refuses
72
72
  * [Compiling, caching and inspecting](docs/04_Compiling.md) — what the first call costs, where kernels are kept, reading the generated C, the `carray-jit` command, and what the suite checks
73
73
  * [Design notes](docs/05_DesignNotes.md) — decisions that were not obvious, and why
74
- * [Cheatsheet](docs/06_Cheatsheet.md) — the seven `jit_` methods and `CArray.fuse` on one page, to look up rather than to read
74
+ * [Cheatsheet](docs/06_Cheatsheet.md) — the eight `jit_` methods and `CArray.fuse` on one page, to look up rather than to read
75
+ * [31 exercises, with solutions](docs/07_StepByStep.md) — thirty-one small tasks in the order this guide introduces things, each with its solution and what it answers ([日本語](docs/07_StepByStep.ja.md))
75
76
 
76
77
  ## Contributing
77
78
 
@@ -81,7 +82,7 @@ Bug reports and feature requests are welcome — please open an issue.
81
82
 
82
83
  ## Credits
83
84
 
84
- carray-jit was designed and reviewed by a human developer; the implementation was produced in collaboration with AI coding tools.
85
+ carray-jit is created and maintained by himotoyoshi. The author provided the design; the implementation was written with AI coding tools and has been verified primarily through the test suite and practical use.
85
86
 
86
87
  ## License
87
88
 
data/carray-jit.gemspec CHANGED
@@ -18,7 +18,6 @@ Gem::Specification.new do |spec|
18
18
 
19
19
  spec.files = Dir[
20
20
  "lib/**/*.rb",
21
- "ext/**/*.{c,h,rb}",
22
21
  "bin/*",
23
22
  "examples/**/*.rb",
24
23
  "examples/README.md",
@@ -32,10 +31,9 @@ Gem::Specification.new do |spec|
32
31
  spec.require_paths = ["lib"]
33
32
  spec.bindir = "bin"
34
33
  spec.executables = ["carray-jit"]
35
- spec.extensions = ["ext/carray_jit_access/extconf.rb"]
36
34
 
37
35
  # A kernel reaches CArray's C by address, so the ceiling is the next minor.
38
- spec.add_dependency "carray", ">= 3.0.1", "< 3.1"
36
+ spec.add_dependency "carray", ">= 3.0.2", "< 3.1"
39
37
  # fiddle ships as a bundled gem; depend on it explicitly so Ruby 3.5+ resolves it.
40
38
  spec.add_dependency "fiddle"
41
39
  end
@@ -20,7 +20,7 @@ A block that steps outside the subset is refused, by name and with a line number
20
20
 
21
21
  A kernel reads and writes CArray arrays directly, in the memory CArray already holds them in. The arrays are the ones the block closes over, so nothing is named twice. Views are cells like any other — a column, a transpose, a slice of a slice — and are written in place without a copy. Masks propagate as CArray propagates them, and a kernel can ask whether a cell is missing.
22
22
 
23
- A kernel can also call C. `jit_extern` binds a function from any library Fiddle can open, and the kernel calls it at its address rather than reaching it per cell through Fiddle; `jit_function` compiles one from a Ruby block of your own and hands back a pure C function pointer, which a kernel can call, Ruby can call, and a C library that knows nothing about either can be given.
23
+ A kernel can also call C. `jit_extern` binds a function from any library Fiddle can open, and the kernel calls it at its address rather than reaching it per cell through Fiddle; `jit_function` compiles one from a Ruby block of your own and hands back a pure C function pointer, which a kernel can call, Ruby can call, and a C library that knows nothing about either can be given. `jit_call` compiles such a body and calls it where it stands, taking its arguments from the locals around it.
24
24
 
25
25
  Compiling costs something the first time and nothing afterwards: the shared object is cached on disk, keyed by the generated C and the compiler that built it, so a kernel is compiled once and reused by every later run.
26
26
 
@@ -33,8 +33,9 @@ Installing this gem also puts the compiler behind `CArray.fuse`. An array expres
33
33
  * [Getting started](01_GettingStarted.md) — the block, its extents, what the three methods return, and where a kernel stands beside `a + b * c` and `CArray.fuse`
34
34
  * [The shapes a kernel takes](02_KernelShapes.md) — work that reaches no neighbour, extents and subscripts, stencils, reductions, and contraction over a repeated index
35
35
  * [Supported features](03_SupportedFeatures.md) — locals and types, branches, raising, the types that are not just a number, calling C, and the recognized subset with what it refuses
36
- * [Compiling, caching and inspecting](04_Compiling.md) — what the first call costs, where kernels are kept, reading the generated C, the `carray-jit` command, and what the suite checks
36
+ * [Compiling, caching and inspecting](04_Compiling.md) — what the first call costs, where kernels are kept, reading the generated C, the `carray-jit` command, a call site answered from somewhere else, and what the suite checks
37
37
  * [Design notes](05_DesignNotes.md) — decisions that were not obvious, and why
38
- * [Cheatsheet](06_Cheatsheet.md) — the seven `jit_` methods and `CArray.fuse` on one page, to look up rather than to read
38
+ * [Cheatsheet](06_Cheatsheet.md) — the eight `jit_` methods and `CArray.fuse` on one page, to look up rather than to read
39
+ * [31 exercises, with solutions](07_StepByStep.md) — thirty-one small tasks in the order this guide introduces things, each with its solution and what it answers ([日本語](07_StepByStep.ja.md))
39
40
 
40
41
  [examples/features/](../examples/features) is a tour of the same ground, one file per feature, and [examples/applications/](../examples/applications) holds small programs that use it to do something.
@@ -16,7 +16,7 @@ CArray.jit_for(1...(rows-1), 1...(columns-1)) { |i, j|
16
16
 
17
17
  ## Where the methods come from
18
18
 
19
- `CArray.jit_for`, `CArray.jit_each` and `CArray.jit_map` are named by CArray, which defines them to raise: they say that they compile their block, that the compiler is this gem, and that it is not installed. Installing it replaces them with the ones that compile.
19
+ `CArray.jit_for`, `CArray.jit_each` and `CArray.jit_map` are this gem's: `require "carray/jit"` defines them, and before that CArray has no method by those names. `CArray.fuse` is the other way round -- CArray's own, working with nothing installed, and compiled once this gem is loaded.
20
20
 
21
21
  The name carries the rest. `jit_` says the block is read rather than run, and so has rules about what may be in it; a method called `per_cell` would not, and the subset would be something you found out about later. It also marks which methods need the compiler and which do not: an expression over whole arrays is `CArray.fuse`'s, and that one needs nothing installed.
22
22
 
@@ -28,6 +28,8 @@ An indexed kernel reaches it too, and there the missing index is the whole point
28
28
  CArray.jit_for(n) { |i| out[i] = signal[i] * gain[] }
29
29
  ```
30
30
 
31
+ `CScalar.int()` with no block is that one cell zero-filled, as `CArray.int(1)` is, which is enough for an accumulator that starts at zero; the block form says the starting value out loud, and is what to write where the value is part of what the code means.
32
+
31
33
  The two routes differ in how they get there. `jit_each` stretches it as it stretches a one-cell `CArray` -- a stride of zero -- while `jit_for` reads it where it lies. They compute the same thing, and `s[0]` keeps working in both, since it is still the one-cell array it is. Writing a wider expression into one is refused, with the shapes named, exactly as CArray refuses it.
32
34
 
33
35
  The other method is the same block with its value asked for:
@@ -40,7 +42,7 @@ The last statement is what every cell of the result gets, so there is no output
40
42
 
41
43
  An assignment is a statement with a value -- Ruby's rule, not a special case here -- so `CArray.jit_map { out = a + b }` writes `out` and hands the same value back. Which is why the two methods are split by what returns rather than by what is written: writing is the block's business either way.
42
44
 
43
- It is the same shape of block `CArray.jit_function` takes, and not the same thing. There, the loop belongs to whoever calls the function, so the body is reached through a pointer and may close over nothing. Here the loop is this compiler's and the body is inlined into it, so it costs no call and may reach a captured value, a `Math` function or a bound C function like any other kernel body.
45
+ It is the same shape of block `CArray.jit_function` takes, and not the same thing. There, the loop belongs to whoever calls the function, so the body is reached through a pointer and may close over nothing but another function compiled the same way, or one borrowed with `jit_extern`. Here the loop is this compiler's and the body is inlined into it, so it costs no call and may reach a captured value, a `Math` function or a bound C function like any other kernel body.
44
46
 
45
47
  #### Who drives the loop
46
48
 
@@ -54,7 +56,9 @@ A **strided view** is the case where there is. A column, a transpose, every othe
54
56
 
55
57
  The rank does not stand in the way, and the arrays are not reshaped to get it out of the way. CArray's acquire reads an operand's element count and element size and never looks at its shape, so a chunk is already a flat run of cells; what has to be flat is the *kernel*, and it is compiled at rank one over the same cells rather than as a nest over the axes. One such kernel then serves whatever rank it is given, because the rank has left the loop.
56
58
 
57
- Four things keep the driver here. An operand that is a **strided view**, for the reason just measured: it is walked in place here, and re-gathered there for nothing. A **stretched operand**, which is the one thing that cannot be flattened: broadcasting arrives as a stride of zero on an axis, and a flat run has no axes, so cell k of the output would stop lining up with cell k of the operand. A body that **asks about a mask**, because a chunk carries no per-cell mask to ask it of -- CArray ORs and propagates the masks, but `a[i] == UNDEF` is a question about one cell. And a CArray **without the family** at all, which is asked for by symbol rather than assumed, so an older one falls back without the caller hearing about it.
59
+ Four things keep the driver here. An operand that is a **strided view**, for the reason just measured: it is walked in place here, and re-gathered there for nothing. A **stretched operand**, which is the one thing that cannot be flattened: broadcasting arrives as a stride of zero on an axis, and a flat run has no axes, so cell k of the output would stop lining up with cell k of the operand. A body that **asks about a mask**, because a chunk carries no per-cell mask to ask it of -- CArray ORs and propagates the masks, but `a[i] == UNDEF` is a question about one cell. A CArray **without the family** at all, which is asked for by symbol rather than assumed, so an older one falls back without the caller hearing about it.
60
+
61
+ Whichever drives it, an operand that is **read while the pass writes the array it is a view of** is copied once before the loop. `a = b * 2.0` reads the whole right-hand side and then assigns; a kernel walks cell by cell, which is the same thing until a cell it reads is one it has written -- `hi = lo * 2.0` over two blocks that share six cells, a transpose written over itself. The copy is what makes those read the values the expression started with, as CArray's own operators do, and it is what lets the chunked sweep take them: what it re-gathers per chunk is then a copy nothing is writing to. Sharing a root is the test rather than overlapping cells, a view knowing what it is a view of; `a = a + 1.0` is not that case, the cell read being the cell written.
58
62
 
59
63
  `jit_for` never goes this way, and that is not a limitation: a kernel that names an index reaches neighbours, chooses an order and runs inner loops, none of which a chunked walk can offer. The two names mark the same line.
60
64
 
@@ -156,7 +160,7 @@ This is the one thing about a kernel that is not settled in advance, and it cost
156
160
 
157
161
  A view that has to be reached a box at a time still takes one, but the box becomes the whole view: a computed index could reach any cell of it. That costs what copying the view would have cost, which is what the caller would otherwise have been told to write by hand.
158
162
 
159
- One restriction remains. A write is either the cell the loop is on or a computed one -- never the cell one along, which is the cell another iteration writes.
163
+ A write may also be displaced: `out[i + 1]` reaches a distinct cell for each pass, and the extent it needs is checked before the first one. Where the array written is also the array read, that is a recurrence -- `values[i + 1] = values[i] + 1.0` -- and it computes what the same Ruby loop computes, the loop running in index order and a sweep handing the chunks over in order too.
160
164
 
161
165
  ## Stencils
162
166
 
@@ -261,7 +265,7 @@ CArray.jit_for(rows) { |i|
261
265
 
262
266
  The accumulator is split into partial sums, which is faster than the serial chain and usually the more accurate answer, and is not the order the same loop would take in Ruby. `reassociate: false` asks for that order; see "The order a reduction takes its terms in" below.
263
267
 
264
- What makes it expressible is `(from...to).each { |j| ... }` -- or `n.times { |j| ... }`, which is the same loop from zero: an inner loop whose index addresses arrays but writes nothing. The accumulator is then an ordinary block-local, which is why sum, maximum, product, count and a dot product all fall out without a primitive each -- and why a matrix multiply does too:
268
+ What makes it expressible is `(from...to).each { |j| ... }` -- or `n.times { |j| ... }`, which is the same loop from zero. The accumulator is then an ordinary block-local, which is why sum, maximum, product, count and a dot product all fall out without a primitive each -- and why a matrix multiply does too:
265
269
 
266
270
  ```ruby
267
271
  CArray.jit_for(rows, columns) { |i, j|
@@ -304,7 +308,7 @@ How many times the outer loop runs does not enter into it: both sides scale with
304
308
 
305
309
  Writing the whole thing as one kernel has no such threshold: it beats the Ruby loop at every size on this table, and by two orders of magnitude once there is real work.
306
310
 
307
- A constant subscript pins an axis: `a[i, 0]` and `a[i, j]` on the same array is ordinary. An array the kernel *writes* cannot be read through an inner index, because that would reach cells another outer iteration owns and no evaluation order settles it.
311
+ A constant subscript pins an axis: `a[i, 0]` and `a[i, j]` on the same array is ordinary.
308
312
 
309
313
  ```
310
314
  row sums over 2000 x 500
@@ -353,6 +357,70 @@ The matrix multiply carries a different subtlety -- it is a plain triple loop an
353
357
 
354
358
  So this is not a faster `sum`. It is a way to write the reduction that has no `sum`.
355
359
 
360
+ ## A row of workspace per cell
361
+
362
+ An inner index addresses a write as well as a read, which is what gives a cell a **row of workspace**: `work[i, k] = ...` fills it, and `work[i, k]` reads it back inside the same cell. That is what an algorithm needing a few numbers per cell is written with -- a small dense solve, a tableau -- and it is the alternative to splitting the body into kernels that each pay a call:
363
+
364
+ ```ruby
365
+ CArray.jit_for(rows) { |i| # a tridiagonal solve per row
366
+ swept[i, 0] = upper[i, 0] / diagonal[i, 0]
367
+ carried[i, 0] = right[i, 0] / diagonal[i, 0]
368
+ (1...width).each { |k|
369
+ denominator = diagonal[i, k] - lower[i, k] * swept[i, k-1]
370
+ swept[i, k] = upper[i, k] / denominator
371
+ carried[i, k] = (right[i, k] - lower[i, k] * carried[i, k-1]) / denominator
372
+ }
373
+ answer[i, width-1] = carried[i, width-1]
374
+ (width-2).step(0, -1) { |k| # back down the row it just filled
375
+ answer[i, k] = carried[i, k] - swept[i, k] * answer[i, k+1]
376
+ }
377
+ }
378
+ ```
379
+
380
+ An inner loop counts by the stride it was written with -- `(width-2).step(0, -1)` above is the pass back down, and it is the spelling an extent takes a direction in. The cell it reaches is bounds-checked before the kernel runs, as `i`'s is: its loop states its extent the same way an extent does. Two outer iterations landing on the same cell is not an ambiguity either -- the extent states the order, so the array holds what the same Ruby loop would have left in it.
381
+
382
+ Which cells are the cell's own is decided by the axes the **outer** indices pick it by. So a read carrying an inner index has to walk those axes with the same outer index -- at whatever offset, `work[i-1, k]` being the row before this one and a recurrence like any other displaced read. Inside the row the body may do as it likes: sort it, walk it backwards, land on a position it works out. A median filter is the shape that wants it, having no expression as an extra axis at all -- the window has to be somewhere while it is being sorted.
383
+
384
+ A row of a captured array is one place to put it, and the body makes one itself where a name will do (see [Local arrays](03_SupportedFeatures.md#local-arrays)), which is shorter and is the only way in a spelling with no index to pick a row by:
385
+
386
+ ```ruby
387
+ median = CArray.jit_stencil(image, border: :clamp) { |a|
388
+ w = CArray.double(9) # nine doubles on this cell's stack
389
+ w[0] = a[-1,-1]; w[1] = a[-1, 0]; w[2] = a[-1, 1]
390
+ w[3] = a[ 0,-1]; w[4] = a[ 0, 0]; w[5] = a[ 0, 1]
391
+ w[6] = a[ 1,-1]; w[7] = a[ 1, 0]; w[8] = a[ 1, 1]
392
+ sort(w)
393
+ w[4]
394
+ }
395
+ ```
396
+
397
+ `sort(w)` is one of the functions the compiler brings with it (see [Intrinsics](03_SupportedFeatures.md#intrinsics)); at nine cells it is a comparator network with no branch in it. Written against a row of captured workspace instead, the same filter needs `work[rows, cols, 9]` -- 288 MB over a 2000x2000 image to hold 72 bytes at a time -- and `jit_stencil` cannot use one at all, its block having no index.
398
+
399
+ **What the clearing costs.** `CArray.double(9)` starts its cells at zero every time the line runs, which for a body that writes all nine before reading any is work nobody asked for. Measured over the filter above on 2000x2000: **41.3 ns a cell, against `CArray.empty(:float64, [9])`'s 42.2** -- no difference, because the compiler sees the nine stores that follow and drops the clearing as dead. So the zeroed spelling is the one to write here, and it is the one that reads correctly.
400
+
401
+ Where the clearing is *not* dead it costs about what clearing that many bytes costs, and nothing else. Over 1000x1000, with a histogram built per cell and a body that reads back a cell it did not write:
402
+
403
+ | cells | of them read back | zeroed | `empty` | the clearing |
404
+ |---|---|---|---|---|
405
+ | 9 | 9 | 9.3 | 9.6 | none |
406
+ | 64 | 64 | 82.8 | 81.5 | none |
407
+ | 256 | 256 | 441.6 | 412.3 | 29.3 |
408
+ | 256 | 9 | 26.9 | 2.7 | 24.2 |
409
+
410
+ (ns a cell.) **The `empty` column is there to compare times and for nothing else: this body reads cells it never wrote, so with `CArray.empty` its answers are undefined -- whatever the stack held.** It is in the table because subtracting it from the column beside it is what isolates the clearing; it is not a spelling to reach for here.
411
+
412
+ Reading the last two rows together is the point: **what costs is how many cells the body reads back, not how many the array has.** A 256-cell array whose body reads nine of them is 2.7 ns a cell; the same array read all the way through is 412. Holding the cells is nearly free -- it is a stack pointer moving -- and the scan is the bill.
413
+
414
+ The clearing is the separate, smaller column: about 24 ns a cell for 2 KiB of int64, roughly the same whichever body sits on top of it, and lost in the noise at 9 and 64 cells.
415
+
416
+ So `CArray.empty(:type, [n])` is for one kind of body only: **one that writes every cell before it reads any.** Reading a cell first is out of contract, the way reading under a mask is -- and a histogram is exactly the body that breaks it, since the zeros it counts up from are values it genuinely reads.
417
+
418
+ Among bodies that do write every cell first, whether the clearing costs anything depends on how the writing is spelled. The median above writes its nine cells in nine lines, and there the clearing is dropped as dead: nothing to save, so write the plain constructor. Fill 256 cells with a loop instead and the `memset` is still in the generated C, costing **199.9 ns a cell against `empty`'s 185.8** -- 14 ns, and both spellings give the same answer because every cell really is written. That is where `CArray.empty` earns its place.
419
+
420
+ The same row of a captured array is still the right answer for a workspace that outlives the cell: a local array is the block's own and is gone when the cell is done, however large it is.
421
+
422
+ What stays refused is the read that leaves the cell: `values[i] = ...` read at `values[j]` for an inner `j` reaches cells another outer iteration owns, and no evaluation order settles that. The message names the axis and what writes it.
423
+
356
424
  ## Contraction
357
425
 
358
426
  `CArray.jit_contract` contracts over a repeated index: **an index that repeats in the term is summed**. The repetition is the notation -- it is what stands in for the sigma. How often it repeats does not enter into it: `q[i,i,i]` is one index read at three positions, and the sum runs along the cube's long diagonal.
@@ -365,6 +433,8 @@ t = CArray.jit_contract { |i| q[i,i] } # a trace
365
433
  o = CArray.jit_contract { |i, j| p[i] * r[j] } # an outer product, nothing summed
366
434
  ```
367
435
 
436
+ An index inside a subscript the kernel works out -- `a[i, idx[k]]` -- is **refused**. Counting positions is what a contraction does, and that one sits on an axis of `idx` rather than on an axis of `a`: read as a position it makes `a[i, idx[k]] * v[k]` a sum over `k`, and not read as one it makes `a[i, k] * w[idx[i]]` a sum over nothing that was asked for. Which of the two was meant is not in the notation, so the gather is written with `jit_for`, where a computed subscript is the ordinary thing it already is.
437
+
368
438
  The result is allocated and returned, its axes being the free indices in the order the block named them -- so the parameter list is where the axis order is stated, and `{ |j, i, k| ... }` gives the transpose. Assigning into an array of your own says where to put it instead:
369
439
 
370
440
  ```ruby
@@ -400,22 +470,31 @@ d = CArray.jit_contract(:a) { q[a,a] } # the diagonal
400
470
  r = CArray.jit_contract(:b, :i, :j) { |k| u[b,i,k] * v[b,k,j] } # a batch of products
401
471
  ```
402
472
 
403
- The arguments are the result's axes, in that order. What they say is which indices are free; they do not say what a repetition means, and a repetition still means a sum. So the whole of it is the convention's rule with a clause added:
473
+ The arguments are the result's axes, in that order. Naming them replaces the convention rather than adding to it, so there are two rules and which one applies is whether a list was given:
474
+
475
+ **With nothing named** -- an index that repeats is summed; one that appears once is free.
476
+
477
+ **With the axes named** -- the named indices are free, in that order, however often they appear; every other index is summed, at however few positions it sits.
478
+
479
+ The second rule's first half is what puts the per-point quantity and the diagonal inside the notation instead of outside it: `q[a,a]` is the trace under the convention and the diagonal when the axis is named. Its second half is what lets a sum along an axis be written at all --
480
+
481
+ ```ruby
482
+ CArray.jit_contract(:i) { |k| a[i,k] } # the row sums
483
+ ```
404
484
 
405
- **An index that repeats is summed; one that appears once is free; and a named one is free however often it appears.**
485
+ -- which the convention alone cannot say, since `k` sits at one position and the convention reads that as free. (`a.sum(axis: 1)` is the faster way to write this one: it is a reduction, and a compiled contraction pays for machinery it has no use for here.)
406
486
 
407
- The third clause is what puts the per-point quantity and the diagonal inside the notation instead of outside it -- `q[a,a]` is the trace under the convention and the diagonal when the axis is named.
487
+ So the list is all of the result's axes rather than some of them: name one and you have named them all. That is what makes it readable -- `jit_contract(:i)` says the result has one axis, and can be trusted to.
408
488
 
409
- A free index then needs somewhere to go, and there are three places: the argument list, the left-hand side, or -- with neither -- the result's axes, which are the free indices in the order the block's parameters named them. So a parameter at a single position is refused once the axes are named. It is free, by the second clause, and the axes are already stated:
489
+ Where the block assigns into an array of yours, the left-hand side has the result's axes written on it, and the two must agree. An axis there that the list left out is the list falling short of the result rather than an index to sum, and that is refused:
410
490
 
411
491
  ```
412
- `k` appears once, so it is free rather than summed. A contraction sums the
413
- indices that repeat; name it as an axis of the result
414
- (`CArray.jit_contract(:i, :k)`) to keep it, or use sum(axis:) to sum along
415
- the axis
492
+ `j` is an axis of the left-hand side and was not named. Naming the result's
493
+ axes names all of them, and what is left out is summed:
494
+ `CArray.jit_contract(:i, :j)`
416
495
  ```
417
496
 
418
- So the list is all of the result's axes rather than some of them: name one and you have named them all. A partial one could be given a meaning -- the axes left out would take their order from the parameter list, as they do when nothing is named -- but it would only ever produce the orders that put the named axes first, so an order like `[i, b, j]` with `b` named could not be asked for at all. The result's order would be stated in two places, and neither could state all of it.
497
+ It is the one place a short list can be caught, because it is the one place the result's axes are stated twice.
419
498
 
420
499
  With every index named there is nothing left to sum, and the block takes no parameters at all.
421
500