carray-jit 0.1.2 → 0.1.3
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +771 -3
- data/README.md +7 -6
- data/carray-jit.gemspec +1 -3
- data/docs/00_Introduction.md +4 -3
- data/docs/01_GettingStarted.md +1 -1
- data/docs/02_KernelShapes.md +93 -14
- data/docs/03_SupportedFeatures.md +582 -26
- data/docs/04_Compiling.md +33 -6
- data/docs/05_DesignNotes.md +3 -3
- data/docs/06_Cheatsheet.md +198 -5
- data/docs/07_StepByStep.ja.md +534 -0
- data/docs/07_StepByStep.md +535 -0
- data/examples/README.md +12 -0
- data/examples/applications/alarm.rb +121 -0
- data/examples/applications/collatz.rb +105 -0
- data/examples/applications/cubic_spline.rb +331 -0
- data/examples/applications/dithering.rb +144 -0
- data/examples/applications/group_stats.rb +115 -0
- data/examples/applications/lookup.rb +126 -0
- data/examples/applications/median_filter.rb +153 -0
- data/examples/applications/parcel_ascent.rb +220 -0
- data/examples/applications/point_in_polygon.rb +111 -0
- data/examples/applications/random_walk.rb +98 -0
- data/examples/applications/van_der_pol.rb +186 -0
- data/examples/applications/wet_bulb.rb +140 -0
- data/examples/features/10_complex.rb +14 -4
- data/examples/features/15_loops.rb +7 -1
- data/lib/carray/jit/access.rb +14 -0
- data/lib/carray/jit/analyzer.rb +2077 -136
- data/lib/carray/jit/block_reader.rb +37 -6
- data/lib/carray/jit/c_function.rb +613 -76
- data/lib/carray/jit/c_generator.rb +1595 -156
- data/lib/carray/jit/call.rb +68 -0
- data/lib/carray/jit/compiler.rb +75 -11
- data/lib/carray/jit/kernel.rb +369 -32
- data/lib/carray/jit/node.rb +359 -9
- data/lib/carray/jit/sorting_networks.rb +182 -0
- data/lib/carray/jit/type_assignment.rb +371 -34
- data/lib/carray/jit/version.rb +1 -1
- data/lib/carray/jit.rb +560 -64
- metadata +22 -8
- data/ext/carray_jit_access/carray_jit_access.c +0 -460
- data/ext/carray_jit_access/extconf.rb +0 -8
data/README.md
CHANGED
|
@@ -8,13 +8,13 @@ The block is read with Prism, translated to C if it falls inside that subset, co
|
|
|
8
8
|
|
|
9
9
|
## Status
|
|
10
10
|
|
|
11
|
-
0.1.
|
|
11
|
+
0.1.3 is the current release, and it still moves: behaviour can change between releases — see [CHANGELOG.md](CHANGELOG.md). A companion gem to CArray, it follows CArray's surface, which is not settled until CArray 3.1.
|
|
12
12
|
|
|
13
13
|
## Features
|
|
14
14
|
|
|
15
15
|
- **A JIT compiler for C-level loops.** A block becomes one C function over CArray's own memory, built by the system C compiler and called through Fiddle.
|
|
16
16
|
- **Ordinary Ruby, and enough of it.** The source is parsed with Prism -- no DSL, no `eval` -- and every operation means what Ruby means by it, apart from the order a reduction takes its terms in. The subset is enough to state a numerical algorithm; what falls outside it is refused by name and line, not run as a Ruby loop.
|
|
17
|
-
- **A method for each shape.** `jit_for` for recurrences and loops written out, `jit_stencil` for windows at any rank, `CArray.jit_contract` for contraction over a repeated index, or over
|
|
17
|
+
- **A method for each shape.** `jit_for` for recurrences and loops written out, `jit_stencil` for windows at any rank, `CArray.jit_contract` for contraction over a repeated index, or over every index the result's axes do not name, `jit_each` and `jit_map` for a pass that reaches no neighbour, `jit_init` for filling an array from its own indices.
|
|
18
18
|
- **View- and mask-aware.** Columns, transposes and slices of slices are written in place without a copy, and masks propagate as CArray propagates them.
|
|
19
19
|
- **Pure C functions, in and out.** `jit_extern` binds one from a library and a kernel calls it by address; `jit_function` compiles one from a block and hands back a C function pointer.
|
|
20
20
|
- **The backend for `CArray.fuse`.** An array expression compiles instead of being walked a node at a time, without being asked and without changing the answer.
|
|
@@ -35,7 +35,7 @@ gem "carray-jit"
|
|
|
35
35
|
Requires:
|
|
36
36
|
|
|
37
37
|
- Ruby >= 3.2
|
|
38
|
-
- CArray >= 3.0.
|
|
38
|
+
- CArray >= 3.0.2, < 3.1
|
|
39
39
|
- A C compiler
|
|
40
40
|
- Prism and Fiddle (both ship with Ruby; Fiddle is a bundled gem)
|
|
41
41
|
|
|
@@ -61,7 +61,7 @@ legendre[0..5].to_a
|
|
|
61
61
|
# => [1.0, 0.5, -0.125, -0.4375, -0.2890625, 0.08984375]
|
|
62
62
|
```
|
|
63
63
|
|
|
64
|
-
|
|
64
|
+
The `jit_` methods are this gem's, and exist once `require "carray/jit"` has run. An expression over whole arrays wants `CArray.fuse`, which is CArray's own and needs no compiler -- and gets compiled anyway where this gem is loaded.
|
|
65
65
|
|
|
66
66
|
## Documentation
|
|
67
67
|
|
|
@@ -71,7 +71,8 @@ legendre[0..5].to_a
|
|
|
71
71
|
* [Supported features](docs/03_SupportedFeatures.md) — locals and types, branches, raising, the types that are not just a number, calling C, and the recognized subset with what it refuses
|
|
72
72
|
* [Compiling, caching and inspecting](docs/04_Compiling.md) — what the first call costs, where kernels are kept, reading the generated C, the `carray-jit` command, and what the suite checks
|
|
73
73
|
* [Design notes](docs/05_DesignNotes.md) — decisions that were not obvious, and why
|
|
74
|
-
* [Cheatsheet](docs/06_Cheatsheet.md) — the
|
|
74
|
+
* [Cheatsheet](docs/06_Cheatsheet.md) — the eight `jit_` methods and `CArray.fuse` on one page, to look up rather than to read
|
|
75
|
+
* [31 exercises, with solutions](docs/07_StepByStep.md) — thirty-one small tasks in the order this guide introduces things, each with its solution and what it answers ([日本語](docs/07_StepByStep.ja.md))
|
|
75
76
|
|
|
76
77
|
## Contributing
|
|
77
78
|
|
|
@@ -81,7 +82,7 @@ Bug reports and feature requests are welcome — please open an issue.
|
|
|
81
82
|
|
|
82
83
|
## Credits
|
|
83
84
|
|
|
84
|
-
carray-jit
|
|
85
|
+
carray-jit is created and maintained by himotoyoshi. The author provided the design; the implementation was written with AI coding tools and has been verified primarily through the test suite and practical use.
|
|
85
86
|
|
|
86
87
|
## License
|
|
87
88
|
|
data/carray-jit.gemspec
CHANGED
|
@@ -18,7 +18,6 @@ Gem::Specification.new do |spec|
|
|
|
18
18
|
|
|
19
19
|
spec.files = Dir[
|
|
20
20
|
"lib/**/*.rb",
|
|
21
|
-
"ext/**/*.{c,h,rb}",
|
|
22
21
|
"bin/*",
|
|
23
22
|
"examples/**/*.rb",
|
|
24
23
|
"examples/README.md",
|
|
@@ -32,10 +31,9 @@ Gem::Specification.new do |spec|
|
|
|
32
31
|
spec.require_paths = ["lib"]
|
|
33
32
|
spec.bindir = "bin"
|
|
34
33
|
spec.executables = ["carray-jit"]
|
|
35
|
-
spec.extensions = ["ext/carray_jit_access/extconf.rb"]
|
|
36
34
|
|
|
37
35
|
# A kernel reaches CArray's C by address, so the ceiling is the next minor.
|
|
38
|
-
spec.add_dependency "carray", ">= 3.0.
|
|
36
|
+
spec.add_dependency "carray", ">= 3.0.2", "< 3.1"
|
|
39
37
|
# fiddle ships as a bundled gem; depend on it explicitly so Ruby 3.5+ resolves it.
|
|
40
38
|
spec.add_dependency "fiddle"
|
|
41
39
|
end
|
data/docs/00_Introduction.md
CHANGED
|
@@ -20,7 +20,7 @@ A block that steps outside the subset is refused, by name and with a line number
|
|
|
20
20
|
|
|
21
21
|
A kernel reads and writes CArray arrays directly, in the memory CArray already holds them in. The arrays are the ones the block closes over, so nothing is named twice. Views are cells like any other — a column, a transpose, a slice of a slice — and are written in place without a copy. Masks propagate as CArray propagates them, and a kernel can ask whether a cell is missing.
|
|
22
22
|
|
|
23
|
-
A kernel can also call C. `jit_extern` binds a function from any library Fiddle can open, and the kernel calls it at its address rather than reaching it per cell through Fiddle; `jit_function` compiles one from a Ruby block of your own and hands back a pure C function pointer, which a kernel can call, Ruby can call, and a C library that knows nothing about either can be given.
|
|
23
|
+
A kernel can also call C. `jit_extern` binds a function from any library Fiddle can open, and the kernel calls it at its address rather than reaching it per cell through Fiddle; `jit_function` compiles one from a Ruby block of your own and hands back a pure C function pointer, which a kernel can call, Ruby can call, and a C library that knows nothing about either can be given. `jit_call` compiles such a body and calls it where it stands, taking its arguments from the locals around it.
|
|
24
24
|
|
|
25
25
|
Compiling costs something the first time and nothing afterwards: the shared object is cached on disk, keyed by the generated C and the compiler that built it, so a kernel is compiled once and reused by every later run.
|
|
26
26
|
|
|
@@ -33,8 +33,9 @@ Installing this gem also puts the compiler behind `CArray.fuse`. An array expres
|
|
|
33
33
|
* [Getting started](01_GettingStarted.md) — the block, its extents, what the three methods return, and where a kernel stands beside `a + b * c` and `CArray.fuse`
|
|
34
34
|
* [The shapes a kernel takes](02_KernelShapes.md) — work that reaches no neighbour, extents and subscripts, stencils, reductions, and contraction over a repeated index
|
|
35
35
|
* [Supported features](03_SupportedFeatures.md) — locals and types, branches, raising, the types that are not just a number, calling C, and the recognized subset with what it refuses
|
|
36
|
-
* [Compiling, caching and inspecting](04_Compiling.md) — what the first call costs, where kernels are kept, reading the generated C, the `carray-jit` command, and what the suite checks
|
|
36
|
+
* [Compiling, caching and inspecting](04_Compiling.md) — what the first call costs, where kernels are kept, reading the generated C, the `carray-jit` command, a call site answered from somewhere else, and what the suite checks
|
|
37
37
|
* [Design notes](05_DesignNotes.md) — decisions that were not obvious, and why
|
|
38
|
-
* [Cheatsheet](06_Cheatsheet.md) — the
|
|
38
|
+
* [Cheatsheet](06_Cheatsheet.md) — the eight `jit_` methods and `CArray.fuse` on one page, to look up rather than to read
|
|
39
|
+
* [31 exercises, with solutions](07_StepByStep.md) — thirty-one small tasks in the order this guide introduces things, each with its solution and what it answers ([日本語](07_StepByStep.ja.md))
|
|
39
40
|
|
|
40
41
|
[examples/features/](../examples/features) is a tour of the same ground, one file per feature, and [examples/applications/](../examples/applications) holds small programs that use it to do something.
|
data/docs/01_GettingStarted.md
CHANGED
|
@@ -16,7 +16,7 @@ CArray.jit_for(1...(rows-1), 1...(columns-1)) { |i, j|
|
|
|
16
16
|
|
|
17
17
|
## Where the methods come from
|
|
18
18
|
|
|
19
|
-
`CArray.jit_for`, `CArray.jit_each` and `CArray.jit_map` are
|
|
19
|
+
`CArray.jit_for`, `CArray.jit_each` and `CArray.jit_map` are this gem's: `require "carray/jit"` defines them, and before that CArray has no method by those names. `CArray.fuse` is the other way round -- CArray's own, working with nothing installed, and compiled once this gem is loaded.
|
|
20
20
|
|
|
21
21
|
The name carries the rest. `jit_` says the block is read rather than run, and so has rules about what may be in it; a method called `per_cell` would not, and the subset would be something you found out about later. It also marks which methods need the compiler and which do not: an expression over whole arrays is `CArray.fuse`'s, and that one needs nothing installed.
|
|
22
22
|
|
data/docs/02_KernelShapes.md
CHANGED
|
@@ -28,6 +28,8 @@ An indexed kernel reaches it too, and there the missing index is the whole point
|
|
|
28
28
|
CArray.jit_for(n) { |i| out[i] = signal[i] * gain[] }
|
|
29
29
|
```
|
|
30
30
|
|
|
31
|
+
`CScalar.int()` with no block is that one cell zero-filled, as `CArray.int(1)` is, which is enough for an accumulator that starts at zero; the block form says the starting value out loud, and is what to write where the value is part of what the code means.
|
|
32
|
+
|
|
31
33
|
The two routes differ in how they get there. `jit_each` stretches it as it stretches a one-cell `CArray` -- a stride of zero -- while `jit_for` reads it where it lies. They compute the same thing, and `s[0]` keeps working in both, since it is still the one-cell array it is. Writing a wider expression into one is refused, with the shapes named, exactly as CArray refuses it.
|
|
32
34
|
|
|
33
35
|
The other method is the same block with its value asked for:
|
|
@@ -40,7 +42,7 @@ The last statement is what every cell of the result gets, so there is no output
|
|
|
40
42
|
|
|
41
43
|
An assignment is a statement with a value -- Ruby's rule, not a special case here -- so `CArray.jit_map { out = a + b }` writes `out` and hands the same value back. Which is why the two methods are split by what returns rather than by what is written: writing is the block's business either way.
|
|
42
44
|
|
|
43
|
-
It is the same shape of block `CArray.jit_function` takes, and not the same thing. There, the loop belongs to whoever calls the function, so the body is reached through a pointer and may close over nothing
|
|
45
|
+
It is the same shape of block `CArray.jit_function` takes, and not the same thing. There, the loop belongs to whoever calls the function, so the body is reached through a pointer and may close over nothing but another function compiled the same way, or one borrowed with `jit_extern`. Here the loop is this compiler's and the body is inlined into it, so it costs no call and may reach a captured value, a `Math` function or a bound C function like any other kernel body.
|
|
44
46
|
|
|
45
47
|
#### Who drives the loop
|
|
46
48
|
|
|
@@ -54,7 +56,9 @@ A **strided view** is the case where there is. A column, a transpose, every othe
|
|
|
54
56
|
|
|
55
57
|
The rank does not stand in the way, and the arrays are not reshaped to get it out of the way. CArray's acquire reads an operand's element count and element size and never looks at its shape, so a chunk is already a flat run of cells; what has to be flat is the *kernel*, and it is compiled at rank one over the same cells rather than as a nest over the axes. One such kernel then serves whatever rank it is given, because the rank has left the loop.
|
|
56
58
|
|
|
57
|
-
Four things keep the driver here. An operand that is a **strided view**, for the reason just measured: it is walked in place here, and re-gathered there for nothing. A **stretched operand**, which is the one thing that cannot be flattened: broadcasting arrives as a stride of zero on an axis, and a flat run has no axes, so cell k of the output would stop lining up with cell k of the operand. A body that **asks about a mask**, because a chunk carries no per-cell mask to ask it of -- CArray ORs and propagates the masks, but `a[i] == UNDEF` is a question about one cell.
|
|
59
|
+
Four things keep the driver here. An operand that is a **strided view**, for the reason just measured: it is walked in place here, and re-gathered there for nothing. A **stretched operand**, which is the one thing that cannot be flattened: broadcasting arrives as a stride of zero on an axis, and a flat run has no axes, so cell k of the output would stop lining up with cell k of the operand. A body that **asks about a mask**, because a chunk carries no per-cell mask to ask it of -- CArray ORs and propagates the masks, but `a[i] == UNDEF` is a question about one cell. A CArray **without the family** at all, which is asked for by symbol rather than assumed, so an older one falls back without the caller hearing about it.
|
|
60
|
+
|
|
61
|
+
Whichever drives it, an operand that is **read while the pass writes the array it is a view of** is copied once before the loop. `a = b * 2.0` reads the whole right-hand side and then assigns; a kernel walks cell by cell, which is the same thing until a cell it reads is one it has written -- `hi = lo * 2.0` over two blocks that share six cells, a transpose written over itself. The copy is what makes those read the values the expression started with, as CArray's own operators do, and it is what lets the chunked sweep take them: what it re-gathers per chunk is then a copy nothing is writing to. Sharing a root is the test rather than overlapping cells, a view knowing what it is a view of; `a = a + 1.0` is not that case, the cell read being the cell written.
|
|
58
62
|
|
|
59
63
|
`jit_for` never goes this way, and that is not a limitation: a kernel that names an index reaches neighbours, chooses an order and runs inner loops, none of which a chunked walk can offer. The two names mark the same line.
|
|
60
64
|
|
|
@@ -156,7 +160,7 @@ This is the one thing about a kernel that is not settled in advance, and it cost
|
|
|
156
160
|
|
|
157
161
|
A view that has to be reached a box at a time still takes one, but the box becomes the whole view: a computed index could reach any cell of it. That costs what copying the view would have cost, which is what the caller would otherwise have been told to write by hand.
|
|
158
162
|
|
|
159
|
-
|
|
163
|
+
A write may also be displaced: `out[i + 1]` reaches a distinct cell for each pass, and the extent it needs is checked before the first one. Where the array written is also the array read, that is a recurrence -- `values[i + 1] = values[i] + 1.0` -- and it computes what the same Ruby loop computes, the loop running in index order and a sweep handing the chunks over in order too.
|
|
160
164
|
|
|
161
165
|
## Stencils
|
|
162
166
|
|
|
@@ -261,7 +265,7 @@ CArray.jit_for(rows) { |i|
|
|
|
261
265
|
|
|
262
266
|
The accumulator is split into partial sums, which is faster than the serial chain and usually the more accurate answer, and is not the order the same loop would take in Ruby. `reassociate: false` asks for that order; see "The order a reduction takes its terms in" below.
|
|
263
267
|
|
|
264
|
-
What makes it expressible is `(from...to).each { |j| ... }` -- or `n.times { |j| ... }`, which is the same loop from zero
|
|
268
|
+
What makes it expressible is `(from...to).each { |j| ... }` -- or `n.times { |j| ... }`, which is the same loop from zero. The accumulator is then an ordinary block-local, which is why sum, maximum, product, count and a dot product all fall out without a primitive each -- and why a matrix multiply does too:
|
|
265
269
|
|
|
266
270
|
```ruby
|
|
267
271
|
CArray.jit_for(rows, columns) { |i, j|
|
|
@@ -304,7 +308,7 @@ How many times the outer loop runs does not enter into it: both sides scale with
|
|
|
304
308
|
|
|
305
309
|
Writing the whole thing as one kernel has no such threshold: it beats the Ruby loop at every size on this table, and by two orders of magnitude once there is real work.
|
|
306
310
|
|
|
307
|
-
A constant subscript pins an axis: `a[i, 0]` and `a[i, j]` on the same array is ordinary.
|
|
311
|
+
A constant subscript pins an axis: `a[i, 0]` and `a[i, j]` on the same array is ordinary.
|
|
308
312
|
|
|
309
313
|
```
|
|
310
314
|
row sums over 2000 x 500
|
|
@@ -353,6 +357,70 @@ The matrix multiply carries a different subtlety -- it is a plain triple loop an
|
|
|
353
357
|
|
|
354
358
|
So this is not a faster `sum`. It is a way to write the reduction that has no `sum`.
|
|
355
359
|
|
|
360
|
+
## A row of workspace per cell
|
|
361
|
+
|
|
362
|
+
An inner index addresses a write as well as a read, which is what gives a cell a **row of workspace**: `work[i, k] = ...` fills it, and `work[i, k]` reads it back inside the same cell. That is what an algorithm needing a few numbers per cell is written with -- a small dense solve, a tableau -- and it is the alternative to splitting the body into kernels that each pay a call:
|
|
363
|
+
|
|
364
|
+
```ruby
|
|
365
|
+
CArray.jit_for(rows) { |i| # a tridiagonal solve per row
|
|
366
|
+
swept[i, 0] = upper[i, 0] / diagonal[i, 0]
|
|
367
|
+
carried[i, 0] = right[i, 0] / diagonal[i, 0]
|
|
368
|
+
(1...width).each { |k|
|
|
369
|
+
denominator = diagonal[i, k] - lower[i, k] * swept[i, k-1]
|
|
370
|
+
swept[i, k] = upper[i, k] / denominator
|
|
371
|
+
carried[i, k] = (right[i, k] - lower[i, k] * carried[i, k-1]) / denominator
|
|
372
|
+
}
|
|
373
|
+
answer[i, width-1] = carried[i, width-1]
|
|
374
|
+
(width-2).step(0, -1) { |k| # back down the row it just filled
|
|
375
|
+
answer[i, k] = carried[i, k] - swept[i, k] * answer[i, k+1]
|
|
376
|
+
}
|
|
377
|
+
}
|
|
378
|
+
```
|
|
379
|
+
|
|
380
|
+
An inner loop counts by the stride it was written with -- `(width-2).step(0, -1)` above is the pass back down, and it is the spelling an extent takes a direction in. The cell it reaches is bounds-checked before the kernel runs, as `i`'s is: its loop states its extent the same way an extent does. Two outer iterations landing on the same cell is not an ambiguity either -- the extent states the order, so the array holds what the same Ruby loop would have left in it.
|
|
381
|
+
|
|
382
|
+
Which cells are the cell's own is decided by the axes the **outer** indices pick it by. So a read carrying an inner index has to walk those axes with the same outer index -- at whatever offset, `work[i-1, k]` being the row before this one and a recurrence like any other displaced read. Inside the row the body may do as it likes: sort it, walk it backwards, land on a position it works out. A median filter is the shape that wants it, having no expression as an extra axis at all -- the window has to be somewhere while it is being sorted.
|
|
383
|
+
|
|
384
|
+
A row of a captured array is one place to put it, and the body makes one itself where a name will do (see [Local arrays](03_SupportedFeatures.md#local-arrays)), which is shorter and is the only way in a spelling with no index to pick a row by:
|
|
385
|
+
|
|
386
|
+
```ruby
|
|
387
|
+
median = CArray.jit_stencil(image, border: :clamp) { |a|
|
|
388
|
+
w = CArray.double(9) # nine doubles on this cell's stack
|
|
389
|
+
w[0] = a[-1,-1]; w[1] = a[-1, 0]; w[2] = a[-1, 1]
|
|
390
|
+
w[3] = a[ 0,-1]; w[4] = a[ 0, 0]; w[5] = a[ 0, 1]
|
|
391
|
+
w[6] = a[ 1,-1]; w[7] = a[ 1, 0]; w[8] = a[ 1, 1]
|
|
392
|
+
sort(w)
|
|
393
|
+
w[4]
|
|
394
|
+
}
|
|
395
|
+
```
|
|
396
|
+
|
|
397
|
+
`sort(w)` is one of the functions the compiler brings with it (see [Intrinsics](03_SupportedFeatures.md#intrinsics)); at nine cells it is a comparator network with no branch in it. Written against a row of captured workspace instead, the same filter needs `work[rows, cols, 9]` -- 288 MB over a 2000x2000 image to hold 72 bytes at a time -- and `jit_stencil` cannot use one at all, its block having no index.
|
|
398
|
+
|
|
399
|
+
**What the clearing costs.** `CArray.double(9)` starts its cells at zero every time the line runs, which for a body that writes all nine before reading any is work nobody asked for. Measured over the filter above on 2000x2000: **41.3 ns a cell, against `CArray.empty(:float64, [9])`'s 42.2** -- no difference, because the compiler sees the nine stores that follow and drops the clearing as dead. So the zeroed spelling is the one to write here, and it is the one that reads correctly.
|
|
400
|
+
|
|
401
|
+
Where the clearing is *not* dead it costs about what clearing that many bytes costs, and nothing else. Over 1000x1000, with a histogram built per cell and a body that reads back a cell it did not write:
|
|
402
|
+
|
|
403
|
+
| cells | of them read back | zeroed | `empty` | the clearing |
|
|
404
|
+
|---|---|---|---|---|
|
|
405
|
+
| 9 | 9 | 9.3 | 9.6 | none |
|
|
406
|
+
| 64 | 64 | 82.8 | 81.5 | none |
|
|
407
|
+
| 256 | 256 | 441.6 | 412.3 | 29.3 |
|
|
408
|
+
| 256 | 9 | 26.9 | 2.7 | 24.2 |
|
|
409
|
+
|
|
410
|
+
(ns a cell.) **The `empty` column is there to compare times and for nothing else: this body reads cells it never wrote, so with `CArray.empty` its answers are undefined -- whatever the stack held.** It is in the table because subtracting it from the column beside it is what isolates the clearing; it is not a spelling to reach for here.
|
|
411
|
+
|
|
412
|
+
Reading the last two rows together is the point: **what costs is how many cells the body reads back, not how many the array has.** A 256-cell array whose body reads nine of them is 2.7 ns a cell; the same array read all the way through is 412. Holding the cells is nearly free -- it is a stack pointer moving -- and the scan is the bill.
|
|
413
|
+
|
|
414
|
+
The clearing is the separate, smaller column: about 24 ns a cell for 2 KiB of int64, roughly the same whichever body sits on top of it, and lost in the noise at 9 and 64 cells.
|
|
415
|
+
|
|
416
|
+
So `CArray.empty(:type, [n])` is for one kind of body only: **one that writes every cell before it reads any.** Reading a cell first is out of contract, the way reading under a mask is -- and a histogram is exactly the body that breaks it, since the zeros it counts up from are values it genuinely reads.
|
|
417
|
+
|
|
418
|
+
Among bodies that do write every cell first, whether the clearing costs anything depends on how the writing is spelled. The median above writes its nine cells in nine lines, and there the clearing is dropped as dead: nothing to save, so write the plain constructor. Fill 256 cells with a loop instead and the `memset` is still in the generated C, costing **199.9 ns a cell against `empty`'s 185.8** -- 14 ns, and both spellings give the same answer because every cell really is written. That is where `CArray.empty` earns its place.
|
|
419
|
+
|
|
420
|
+
The same row of a captured array is still the right answer for a workspace that outlives the cell: a local array is the block's own and is gone when the cell is done, however large it is.
|
|
421
|
+
|
|
422
|
+
What stays refused is the read that leaves the cell: `values[i] = ...` read at `values[j]` for an inner `j` reaches cells another outer iteration owns, and no evaluation order settles that. The message names the axis and what writes it.
|
|
423
|
+
|
|
356
424
|
## Contraction
|
|
357
425
|
|
|
358
426
|
`CArray.jit_contract` contracts over a repeated index: **an index that repeats in the term is summed**. The repetition is the notation -- it is what stands in for the sigma. How often it repeats does not enter into it: `q[i,i,i]` is one index read at three positions, and the sum runs along the cube's long diagonal.
|
|
@@ -365,6 +433,8 @@ t = CArray.jit_contract { |i| q[i,i] } # a trace
|
|
|
365
433
|
o = CArray.jit_contract { |i, j| p[i] * r[j] } # an outer product, nothing summed
|
|
366
434
|
```
|
|
367
435
|
|
|
436
|
+
An index inside a subscript the kernel works out -- `a[i, idx[k]]` -- is **refused**. Counting positions is what a contraction does, and that one sits on an axis of `idx` rather than on an axis of `a`: read as a position it makes `a[i, idx[k]] * v[k]` a sum over `k`, and not read as one it makes `a[i, k] * w[idx[i]]` a sum over nothing that was asked for. Which of the two was meant is not in the notation, so the gather is written with `jit_for`, where a computed subscript is the ordinary thing it already is.
|
|
437
|
+
|
|
368
438
|
The result is allocated and returned, its axes being the free indices in the order the block named them -- so the parameter list is where the axis order is stated, and `{ |j, i, k| ... }` gives the transpose. Assigning into an array of your own says where to put it instead:
|
|
369
439
|
|
|
370
440
|
```ruby
|
|
@@ -400,22 +470,31 @@ d = CArray.jit_contract(:a) { q[a,a] } # the diagonal
|
|
|
400
470
|
r = CArray.jit_contract(:b, :i, :j) { |k| u[b,i,k] * v[b,k,j] } # a batch of products
|
|
401
471
|
```
|
|
402
472
|
|
|
403
|
-
The arguments are the result's axes, in that order.
|
|
473
|
+
The arguments are the result's axes, in that order. Naming them replaces the convention rather than adding to it, so there are two rules and which one applies is whether a list was given:
|
|
474
|
+
|
|
475
|
+
**With nothing named** -- an index that repeats is summed; one that appears once is free.
|
|
476
|
+
|
|
477
|
+
**With the axes named** -- the named indices are free, in that order, however often they appear; every other index is summed, at however few positions it sits.
|
|
478
|
+
|
|
479
|
+
The second rule's first half is what puts the per-point quantity and the diagonal inside the notation instead of outside it: `q[a,a]` is the trace under the convention and the diagonal when the axis is named. Its second half is what lets a sum along an axis be written at all --
|
|
480
|
+
|
|
481
|
+
```ruby
|
|
482
|
+
CArray.jit_contract(:i) { |k| a[i,k] } # the row sums
|
|
483
|
+
```
|
|
404
484
|
|
|
405
|
-
|
|
485
|
+
-- which the convention alone cannot say, since `k` sits at one position and the convention reads that as free. (`a.sum(axis: 1)` is the faster way to write this one: it is a reduction, and a compiled contraction pays for machinery it has no use for here.)
|
|
406
486
|
|
|
407
|
-
|
|
487
|
+
So the list is all of the result's axes rather than some of them: name one and you have named them all. That is what makes it readable -- `jit_contract(:i)` says the result has one axis, and can be trusted to.
|
|
408
488
|
|
|
409
|
-
|
|
489
|
+
Where the block assigns into an array of yours, the left-hand side has the result's axes written on it, and the two must agree. An axis there that the list left out is the list falling short of the result rather than an index to sum, and that is refused:
|
|
410
490
|
|
|
411
491
|
```
|
|
412
|
-
`
|
|
413
|
-
|
|
414
|
-
|
|
415
|
-
the axis
|
|
492
|
+
`j` is an axis of the left-hand side and was not named. Naming the result's
|
|
493
|
+
axes names all of them, and what is left out is summed:
|
|
494
|
+
`CArray.jit_contract(:i, :j)`
|
|
416
495
|
```
|
|
417
496
|
|
|
418
|
-
|
|
497
|
+
It is the one place a short list can be caught, because it is the one place the result's axes are stated twice.
|
|
419
498
|
|
|
420
499
|
With every index named there is nothing left to sum, and the block takes no parameters at all.
|
|
421
500
|
|