carray-jit 0.1.2 → 0.1.3
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +771 -3
- data/README.md +7 -6
- data/carray-jit.gemspec +1 -3
- data/docs/00_Introduction.md +4 -3
- data/docs/01_GettingStarted.md +1 -1
- data/docs/02_KernelShapes.md +93 -14
- data/docs/03_SupportedFeatures.md +582 -26
- data/docs/04_Compiling.md +33 -6
- data/docs/05_DesignNotes.md +3 -3
- data/docs/06_Cheatsheet.md +198 -5
- data/docs/07_StepByStep.ja.md +534 -0
- data/docs/07_StepByStep.md +535 -0
- data/examples/README.md +12 -0
- data/examples/applications/alarm.rb +121 -0
- data/examples/applications/collatz.rb +105 -0
- data/examples/applications/cubic_spline.rb +331 -0
- data/examples/applications/dithering.rb +144 -0
- data/examples/applications/group_stats.rb +115 -0
- data/examples/applications/lookup.rb +126 -0
- data/examples/applications/median_filter.rb +153 -0
- data/examples/applications/parcel_ascent.rb +220 -0
- data/examples/applications/point_in_polygon.rb +111 -0
- data/examples/applications/random_walk.rb +98 -0
- data/examples/applications/van_der_pol.rb +186 -0
- data/examples/applications/wet_bulb.rb +140 -0
- data/examples/features/10_complex.rb +14 -4
- data/examples/features/15_loops.rb +7 -1
- data/lib/carray/jit/access.rb +14 -0
- data/lib/carray/jit/analyzer.rb +2077 -136
- data/lib/carray/jit/block_reader.rb +37 -6
- data/lib/carray/jit/c_function.rb +613 -76
- data/lib/carray/jit/c_generator.rb +1595 -156
- data/lib/carray/jit/call.rb +68 -0
- data/lib/carray/jit/compiler.rb +75 -11
- data/lib/carray/jit/kernel.rb +369 -32
- data/lib/carray/jit/node.rb +359 -9
- data/lib/carray/jit/sorting_networks.rb +182 -0
- data/lib/carray/jit/type_assignment.rb +371 -34
- data/lib/carray/jit/version.rb +1 -1
- data/lib/carray/jit.rb +560 -64
- metadata +22 -8
- data/ext/carray_jit_access/carray_jit_access.c +0 -460
- data/ext/carray_jit_access/extconf.rb +0 -8
data/docs/04_Compiling.md
CHANGED
|
@@ -67,7 +67,7 @@ In `~/.cache/carray-jit` (honouring `XDG_CACHE_HOME`), one directory per version
|
|
|
67
67
|
|
|
68
68
|
A kernel is only good for the version that generated it and the architecture it was built for, so those get their own directories. CArray's version is there for the same reason and one of its own: a kernel is handed CArray's memory, on layouts CArray decides, so one compiled against one version and run against the next would answer wrongly rather than fail to load.
|
|
69
69
|
|
|
70
|
-
The two numbers say different things and move on their own clocks: the first is the version that generated the kernel, the second the version it was generated against. This gem is not versioned with CArray, and one release of it meets more than one -- the dependency is `>= 3.0.
|
|
70
|
+
The two numbers say different things and move on their own clocks: the first is the version that generated the kernel, the second the version it was generated against. This gem is not versioned with CArray, and one release of it meets more than one -- the dependency is `>= 3.0.2, < 3.1`, so 3.0.2 and 3.0.4 both satisfy it -- and a layout the kernel reaches into can move between them. A new release does not spend its cache budget on entries nothing can reach any more, and a home directory shared between machines -- over NFS, or between Rosetta and native -- does not have one architecture evicting the other's kernels. A directory whose newest entry has not been touched in 30 days (`CARRAY_JIT_CACHE_MAX_AGE_DAYS`) is removed.
|
|
71
71
|
|
|
72
72
|
```ruby
|
|
73
73
|
CArray::JIT.cache_root #=> "/home/you/.cache/carray-jit"
|
|
@@ -107,7 +107,7 @@ Because the cache is shared across processes and across time, the key covers the
|
|
|
107
107
|
|
|
108
108
|
An entry that will not load -- a truncated write, an OS or toolchain change -- is deleted and rebuilt rather than raised. A cache that outlives the process must not be able to turn one bad write into a permanent failure of every future run. A staging file left behind by a process killed mid-compile is swept once it is old enough to be certain nothing is still writing it.
|
|
109
109
|
|
|
110
|
-
The directory is created 0700, and a cache directory other users can write to is refused rather than used -- everything in it gets `dlopen`ed, so a shared writable cache would be a way to run code as you.
|
|
110
|
+
The directory is created 0700, and a cache directory other users can write to is refused rather than used -- everything in it gets `dlopen`ed, so a shared writable cache would be a way to run code as you. So is one that belongs to another user, whatever its mode says: its owner writes to it, and a directory of theirs at the name yours would have taken is the same thing by another route. A root that is a symlink to another volume is followed as before, the question being asked of the directory it lands on.
|
|
111
111
|
|
|
112
112
|
## Inspecting a kernel
|
|
113
113
|
|
|
@@ -144,18 +144,20 @@ carray_jit_contiguous (char **pointers, int64_t *strides, int64_t *bounds, ...)
|
|
|
144
144
|
const double x = reals[0];
|
|
145
145
|
|
|
146
146
|
for (int64_t i = bounds[0]; i < bounds[1]; i++) {
|
|
147
|
-
double w
|
|
148
|
-
double wy
|
|
147
|
+
double w;
|
|
148
|
+
double wy;
|
|
149
|
+
w = x * ((double *)(p_legendre))[i - 1];
|
|
150
|
+
wy = w - ((double *)(p_legendre))[i - 2];
|
|
149
151
|
((double *)(p_legendre))[i] = wy + w - wy / (double)i;
|
|
150
152
|
}
|
|
151
153
|
}
|
|
152
154
|
```
|
|
153
155
|
|
|
154
|
-
Every kernel has that same signature, which is what lets one Fiddle::Function shape serve all of them; the per-kernel detail arrives in the buffers and is unpacked into named locals at the top, where it also reads better. Abridged above are the ones this kernel barely uses: `reals` and `integers` for the captured scalars, `functions` and `data` for the address of each C function the block called and of each array it handed to one whole, `mask_pointers` and `mask_strides` for the masks, and `error` for the one thing a cell can raise. `bounds` carries a start, a limit and a step per axis -- which is also what a chunk looks like, and is why CArray's sweep can call this kernel directly.
|
|
156
|
+
Every kernel has that same signature, which is what lets one Fiddle::Function shape serve all of them; the per-kernel detail arrives in the buffers and is unpacked into named locals at the top, where it also reads better. Abridged above are the ones this kernel barely uses: `reals` and `integers` for the captured scalars, `functions` and `data` for the address of each C function the block called and of each array it handed to one whole, `mask_pointers` and `mask_strides` for the masks, and `error` for the one thing a cell can raise. `bounds` carries a start, a limit and a step per axis -- which is also what a chunk looks like, and is why CArray's sweep can call this kernel directly. A name in the block that the generated C already has a use for -- a local called `error` or `p_legendre`, say -- is written as `carray_jit_name1_error` -- a number, then the name -- and every other name is the one the block used.
|
|
155
157
|
|
|
156
158
|
### The carray-jit command
|
|
157
159
|
|
|
158
|
-
Installing the gem provides a small command for looking after the cache. It loads only the compiler and its cache, so it works whether or not CArray
|
|
160
|
+
Installing the gem provides a small command for looking after the cache. It loads only the compiler and its cache, so it works whether or not CArray itself can be loaded -- the CArray version in the environment name is the one RubyGems says a `require` would activate, which is the one a running program would have reported itself.
|
|
159
161
|
|
|
160
162
|
```
|
|
161
163
|
$ carray-jit
|
|
@@ -220,6 +222,31 @@ carray_jit_contiguous (char **pointers, int64_t *strides, int64_t *bounds, ...)
|
|
|
220
222
|
| `CARRAY_JIT_CC` | C compiler (default `RbConfig::CONFIG["CC"]`) |
|
|
221
223
|
| `CARRAY_JIT_REASSOCIATE` | `0` makes the serial accumulator the default for the process |
|
|
222
224
|
|
|
225
|
+
## A site answered from somewhere else
|
|
226
|
+
|
|
227
|
+
A `CArray.jit_call` site (see [Compiling a body where it is called](03_SupportedFeatures.md#compiling-a-body-where-it-is-called)) names a function by its declaration and takes its arguments from the locals around it. Nothing in that says the function has to be compiled here, and the first thing asked at a site is whether something else already has one:
|
|
228
|
+
|
|
229
|
+
```ruby
|
|
230
|
+
CArray::JIT.call_provider = ->(prototype, block, names) {
|
|
231
|
+
LIBRARY.lookup(prototype, block) # or nil
|
|
232
|
+
}
|
|
233
|
+
```
|
|
234
|
+
|
|
235
|
+
What it is handed is the prototype, the block -- whose binding says which method and which module the site is in -- and the names the declaration gave. What it answers is anything that responds to `call`, or `nil`, which means "not mine, compile it" and is the position `jit_extern` puts a function from a library in, said about a call rather than about a name. The answer is kept for that site like any other, so a provider is asked once per site.
|
|
236
|
+
|
|
237
|
+
This is what `require "carray/jit/call"` is for. It defines `CArray.jit_call` and the declaration's parser and stops there; the compiler -- the analyzer, the generator, the cache -- is required at the first site no provider answers, and not before. A program whose sites were all built ahead of it therefore never loads it. `require "carray/jit"` loads everything as it always did, and a program that knows nothing about any of this sees no difference.
|
|
238
|
+
|
|
239
|
+
The other half of building ahead is getting the C out, which is `c_source_as`:
|
|
240
|
+
|
|
241
|
+
```ruby
|
|
242
|
+
square = CArray.jit_function("double sq(double x)") { |x| x * x }
|
|
243
|
+
File.write("kernels.c", square.c_source_as(:my_square))
|
|
244
|
+
```
|
|
245
|
+
|
|
246
|
+
`c_source` is the source as this compiler wrote it, under the symbol it picked -- `carray_jit_sq_1c42612898e4`, a digest of the body, which is right for an object in a cache and wrong for one committed to a repository, where the name would change every time the body was touched. `c_source_as` writes the same body under a name the caller picked. A name C cannot spell is refused, and so is one in this compiler's own `carray_jit_` namespace, where a cached kernel is entitled to the name; a name the C library already has is the caller's to avoid, the prefix that closed that hazard being what is replaced. A function bound with `jit_extern` has no C of its own and says so.
|
|
247
|
+
|
|
248
|
+
carray-jit-aot is the gem these were built for: it reads the call sites out of a program, compiles them into one shared object, and answers as the provider at run time. They are here rather than there because a gem reaching into another's method to change what it does is the arrangement that breaks quietly.
|
|
249
|
+
|
|
223
250
|
## Testing
|
|
224
251
|
|
|
225
252
|
```
|
data/docs/05_DesignNotes.md
CHANGED
|
@@ -4,7 +4,7 @@ Decisions that were not obvious, and why.
|
|
|
4
4
|
|
|
5
5
|
## The type is C's, and so is the arithmetic where the width is real
|
|
6
6
|
|
|
7
|
-
The storage type is CArray's, mapped to C exactly -- `int64_t`, `uint8_t`, `float` -- and everything the type itself decides follows from that: the width, the wrap on store, the bit patterns
|
|
7
|
+
The storage type is CArray's, mapped to C exactly -- `int64_t`, `uint8_t`, `float` -- and everything the type itself decides follows from that: the width, the wrap on store, the bit patterns. There a kernel agrees with CArray, because both are the same C. A shift is the exception that proves it: a kernel shifts an integer in the width it computes in, `int64_t`, and narrows on store, so on an `int32` cell a count past 31 gives the Ruby loop's answer where CArray, shifting in 32 bits, gives another.
|
|
8
8
|
|
|
9
9
|
The arithmetic follows the same rule where the width makes a difference to the answer, and Ruby's where it does not. That splits the types in two.
|
|
10
10
|
|
|
@@ -65,7 +65,7 @@ Ruby floors integer division and gives the remainder the sign of the divisor; C
|
|
|
65
65
|
|
|
66
66
|
The generated helper mirrors `ext/mkkernel.rb` in CArray. Dividing by a positive power of two skips the helper: an arithmetic shift already floors, and is cheaper than the truncating divide C would emit.
|
|
67
67
|
|
|
68
|
-
For the same reason `%` is **not** lowered to `fmod`, which truncates. Ruby and CArray floor it
|
|
68
|
+
For the same reason `%` is **not** lowered to `fmod`, which truncates. Ruby and CArray floor it, and so does a kernel -- through a helper for an integer, and for a float through `fmod` with the sign corrected afterwards.
|
|
69
69
|
|
|
70
70
|
## Integer division by zero
|
|
71
71
|
|
|
@@ -133,4 +133,4 @@ The columns are the part that matters. `Proc#source_location` reports only a lin
|
|
|
133
133
|
|
|
134
134
|
`RubyVM::AbstractSyntaxTree.of` looks like the obvious route and does not work: since Ruby 3.4 the default parser is Prism, and it refuses with "cannot get AST for ISEQ compiled by prism".
|
|
135
135
|
|
|
136
|
-
Two cases fall back to `
|
|
136
|
+
Two cases fall back to the text: `CArray::JIT.compile`, which takes a kernel as a string. A block defined in `eval` or in a console has no file to read -- setting `RubyVM.keep_script_lines = true` before it is defined makes `script_lines` available and handles that. And a file edited since it was loaded no longer holds the same text at that position, which is reported rather than compiled.
|
data/docs/06_Cheatsheet.md
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Cheatsheet
|
|
2
2
|
|
|
3
|
-
|
|
3
|
+
Nine entry points: the eight `jit_` methods this gem puts on `CArray`, and
|
|
4
4
|
`CArray.fuse`, which is CArray's own and gets the compiler from this gem being
|
|
5
5
|
installed. Every example here runs as written.
|
|
6
6
|
|
|
@@ -13,10 +13,14 @@ installed. Every example here runs as written.
|
|
|
13
13
|
| The same, and you want the result back | `CArray.jit_map` |
|
|
14
14
|
| A cell reads its neighbours, or the one computed before it | `CArray.jit_for` |
|
|
15
15
|
| A cell reads a window, and the edge needs a rule | `CArray.jit_stencil` |
|
|
16
|
+
| Fill an array from its own indices | `array.jit_init` |
|
|
16
17
|
| An index repeats and is summed | `CArray.jit_contract` |
|
|
17
18
|
| An index repeats and is *not* summed -- a point number, a batch | `CArray.jit_contract(:p)`, naming the result's axes |
|
|
18
19
|
| Call a C function someone else compiled | `CArray.jit_extern` |
|
|
19
20
|
| Compile a C function of your own | `CArray.jit_function` |
|
|
21
|
+
| Compile a body and call it where it stands, on the locals around it | `CArray.jit_call` |
|
|
22
|
+
| Draw random numbers inside a kernel | `CArray::Rng` + `r.random` |
|
|
23
|
+
| Bring a scalar in from outside, or accumulate into one | `CScalar` -- `s[]` is the value |
|
|
20
24
|
|
|
21
25
|
The dividing line among the first four is **what reaches what**. Element-wise
|
|
22
26
|
work reaches no neighbour, so it names no index and needs no extent. A cell
|
|
@@ -62,10 +66,112 @@ each, an Integer `n` standing for `0...n`. Naming an index is what lets a cell
|
|
|
62
66
|
reach `x[i-1]`, and reaching a cell the kernel will later write is what fixes
|
|
63
67
|
the direction the axis runs -- derived from the dependencies, not chosen.
|
|
64
68
|
|
|
69
|
+
An inner loop counts up with `(a...b).each` or `n.times`, and by a stride with
|
|
70
|
+
`a.step(b, s)` -- `(n-1).step(0, -1)` for a sweep back down a row. Its index
|
|
71
|
+
addresses writes as well as reads, which is a cell's own workspace:
|
|
72
|
+
|
|
73
|
+
```ruby
|
|
74
|
+
CArray.jit_for(rows) { |i|
|
|
75
|
+
(0...width).each { |k| work[i, k] = ... } # fill the row
|
|
76
|
+
(width-1).step(0, -1) { |k| ... work[i, k] ... } # and walk back down it
|
|
77
|
+
}
|
|
78
|
+
```
|
|
79
|
+
|
|
80
|
+
`x.nan?` and `x.finite?` are the guards; `infinite?` is refused, answering
|
|
81
|
+
nil or ±1 rather than a boolean. `x.clamp(low, high)` bounds a value, with the
|
|
82
|
+
value and both bounds one class.
|
|
83
|
+
|
|
84
|
+
An operator assignment is the assignment it stands for: `total += a[i]`,
|
|
85
|
+
`work[i, k] *= 2.0`, `counts[bin[i]] += 1`.
|
|
86
|
+
|
|
87
|
+
A parallel assignment settles every value before it writes anything, which is
|
|
88
|
+
how a recurrence advances: `a, b = b, a + b`, `a[i], a[j] = a[j], a[i]`. One
|
|
89
|
+
value per name, written out -- the readings that take a single value apart
|
|
90
|
+
(`a, b = f(x)`, `a, *rest =`, `a, (b, c) =`) are each refused by name.
|
|
91
|
+
|
|
65
92
|
`reassociate:` says whether a reduction's accumulator may be split into partial
|
|
66
93
|
sums. Default is `CArray::JIT.reassociate` (`true`). Pass `false` for the
|
|
67
94
|
serial order -- a compensated summation, or checking against the loop.
|
|
68
95
|
|
|
96
|
+
A `CArray` the block makes is a C array of the block's own -- in any block
|
|
97
|
+
but a contraction's, of any number of axes, cleared at the line each pass
|
|
98
|
+
unless it is `CArray.empty(:type, [n])`. Note `CArray.float` is float32 and
|
|
99
|
+
`CArray.complex` is cmplx64. Up to 4 KiB it stands in the frame; past that,
|
|
100
|
+
or where the shape is written over an integer the block captured
|
|
101
|
+
(`CArray.double(n)`), the kernel allocates it at its entry and frees it at
|
|
102
|
+
its exit -- and then every subscript is checked where the cell is reached,
|
|
103
|
+
and a C function takes it only through a pointer with no length. Where the kernel carries masks
|
|
104
|
+
each cell gets a mask byte beside it, so `w[k] = UNDEF` and `w[k] == UNDEF`
|
|
105
|
+
are written as a captured array's are -- but `sum`, `min`, `max`, `sort` and
|
|
106
|
+
a C function take no such array.
|
|
107
|
+
|
|
108
|
+
```ruby
|
|
109
|
+
CArray.jit_for(rows) { |i|
|
|
110
|
+
counts = CArray.int64(256) # int64_t counts[256]; + memset
|
|
111
|
+
(0...cols).each { |j| counts[labels[i, j]] += 1 }
|
|
112
|
+
...
|
|
113
|
+
}
|
|
114
|
+
```
|
|
115
|
+
|
|
116
|
+
Four functions come with the compiler, over a local array of one axis. Bare
|
|
117
|
+
calls, not methods on the array -- `w.sum` is refused.
|
|
118
|
+
|
|
119
|
+
| Written | Is | Notes |
|
|
120
|
+
|---|---|---|
|
|
121
|
+
| `sum(w)` | a value | in index order from cell 0; no partial sums. Boolean refused |
|
|
122
|
+
| `min(w)` | a value | a NaN is skipped wherever it stands; all NaN gives `NaN` |
|
|
123
|
+
| `max(w)` | a value | the same. Complex and boolean refused |
|
|
124
|
+
| `sort(w)` | a **statement** | ascending, in place, NaN last. A network up to 16 cells, an insertion sort above |
|
|
125
|
+
|
|
126
|
+
```ruby
|
|
127
|
+
median = CArray.jit_for(rows) { |i|
|
|
128
|
+
w = CArray.double(9)
|
|
129
|
+
(0...9).each { |k| w[k] = a[i, k] }
|
|
130
|
+
sort(w)
|
|
131
|
+
out[i] = w[4]
|
|
132
|
+
}
|
|
133
|
+
```
|
|
134
|
+
|
|
135
|
+
## Filling from the indices
|
|
136
|
+
|
|
137
|
+
```ruby
|
|
138
|
+
CArray.double(3, 4).jit_init { |i, j| i * 10.0 + j }
|
|
139
|
+
grid = CArray.double(n)
|
|
140
|
+
grid.jit_init { |i| Math.sin(i * step) } # fills in place
|
|
141
|
+
```
|
|
142
|
+
|
|
143
|
+
`jit_init` fills every cell of the receiver from the indices of that cell, one
|
|
144
|
+
block parameter per axis, as a constructor block does -- and hands the receiver
|
|
145
|
+
back. What is compiled is the whole fill, so the block reaches the same subset
|
|
146
|
+
every other kernel does and closes over arrays the same way. A block that names
|
|
147
|
+
no index is `jit_each`'s work, and is refused here saying so.
|
|
148
|
+
|
|
149
|
+
## A scalar from outside
|
|
150
|
+
|
|
151
|
+
```ruby
|
|
152
|
+
gain = CScalar.double() { 2.0 }
|
|
153
|
+
total = CScalar.int()
|
|
154
|
+
|
|
155
|
+
CArray.jit_each { out = signal * gain } # read at every cell
|
|
156
|
+
CArray.jit_for(a.size) { |i| total[] += a[i] } # and written like any cell
|
|
157
|
+
```
|
|
158
|
+
|
|
159
|
+
A captured Numeric has no data type of its own, so a local seeded from one has
|
|
160
|
+
none either; a `CScalar` is a value with a type, which is what makes it the way
|
|
161
|
+
to hand a kernel a scalar and the way to take one back out. `s[]` is the value
|
|
162
|
+
and `s[] = ...` puts one back. A bare `s` reads the value too, there being no
|
|
163
|
+
axis to walk; written to, a bare name is a local of the block's own, so an
|
|
164
|
+
accumulator is `s[] = s[] + ...` and not `s = s + ...`.
|
|
165
|
+
|
|
166
|
+
Every iteration writing the one cell it has is what makes it an accumulator,
|
|
167
|
+
and what is left is what the same Ruby loop leaves. `CScalar.int()` is that one
|
|
168
|
+
cell zero-filled, as `CArray.int(1)` is; write `CScalar.int() { 0 }` where the
|
|
169
|
+
starting value is part of what the code says.
|
|
170
|
+
|
|
171
|
+
See [CScalar](02_KernelShapes.md#a-cscalar-is-a-value-with-a-home) for how the
|
|
172
|
+
two routes reach it -- `jit_each` stretches it at a stride of zero, `jit_for`
|
|
173
|
+
reads it where it lies.
|
|
174
|
+
|
|
69
175
|
## Windows
|
|
70
176
|
|
|
71
177
|
```ruby
|
|
@@ -82,6 +188,9 @@ is what the cell gets.
|
|
|
82
188
|
|---|---|
|
|
83
189
|
| `border: :mask` | a cell whose window falls off is `UNDEF` -- not computed (default) |
|
|
84
190
|
| `border: :skip` | that cell is left as it was found |
|
|
191
|
+
| `border: :zero` | a read that falls off gives 0, and the cell is computed |
|
|
192
|
+
| `border: :clamp` | such a read gives the nearest cell inside |
|
|
193
|
+
| `border: :wrap` | such a read comes back the other side |
|
|
85
194
|
| `type:` | the data type to collect into; without it, the block's value's |
|
|
86
195
|
| `into:` | write into an array of yours, which then decides the type |
|
|
87
196
|
|
|
@@ -126,7 +235,88 @@ CArray.jit_each { out = j0.call(x) }
|
|
|
126
235
|
it directly rather than reaching it per cell through Fiddle. `from:` names the
|
|
127
236
|
library; `nil` searches the process. `jit_function` compiles a body of your
|
|
128
237
|
own, callable from a kernel, from Ruby, and by a C library that knows nothing
|
|
129
|
-
about either.
|
|
238
|
+
about either -- and from another `jit_function` body. That and a function
|
|
239
|
+
borrowed with `jit_extern`, which the generated file declares and calls by
|
|
240
|
+
name, are the two things such a body may reach outside its parameters.
|
|
241
|
+
|
|
242
|
+
```ruby
|
|
243
|
+
values = CArray.double(8) { |i| i }
|
|
244
|
+
out = CArray.double(8)
|
|
245
|
+
n = values.elements
|
|
246
|
+
|
|
247
|
+
CArray.jit_call("void (*)(double *out, const double *values, size_t n)") {
|
|
248
|
+
n.times { |i| out[i] = values[i] * 2.0 }
|
|
249
|
+
}
|
|
250
|
+
```
|
|
251
|
+
|
|
252
|
+
`jit_call` compiles the block and calls it there and then. The declaration's
|
|
253
|
+
parameter names are the binding: `out`, `values` and `n` are the body's
|
|
254
|
+
parameters and the locals around the call, so there is no argument list to keep
|
|
255
|
+
in step. Compiled once per call site.
|
|
256
|
+
|
|
257
|
+
A signature is written in C's own spellings: `float` and `double`, the
|
|
258
|
+
exact-width integers `int8_t`..`int64_t` and `uint8_t`..`uint64_t`, and the
|
|
259
|
+
platform's own words `size_t`, `ssize_t`, `ptrdiff_t`, `intptr_t` and
|
|
260
|
+
`uintptr_t`. A word is whatever the platform made it, so `size_t` computes as
|
|
261
|
+
a `uint64` and `size_t counts[]` takes a `uint64` array where that word is 64
|
|
262
|
+
bits; `uint64_t` is the spelling that says the width itself.
|
|
263
|
+
|
|
264
|
+
---
|
|
265
|
+
|
|
266
|
+
## Random numbers
|
|
267
|
+
|
|
268
|
+
```ruby
|
|
269
|
+
rand = CArray::Rng.new(seed: 4)
|
|
270
|
+
|
|
271
|
+
CArray.jit_for(n) { |i| out[i] = rand.random } # one draw per cell, [0.0, 1.0)
|
|
272
|
+
```
|
|
273
|
+
|
|
274
|
+
The generator is CArray's, not this gem's, and there is no entry point of
|
|
275
|
+
this gem's own: `CArray::Rng.new` makes one and a kernel draws from it.
|
|
276
|
+
Five spellings, three meanings:
|
|
277
|
+
|
|
278
|
+
| | |
|
|
279
|
+
|---|---|
|
|
280
|
+
| `r.random` | a double in `[0.0, 1.0)` |
|
|
281
|
+
| `r.randomn` | a standard normal, which costs two draws |
|
|
282
|
+
| `r.bits` | the raw word a draw came from, as a `uint64` |
|
|
283
|
+
| `random(rng: r)`, `randomn(rng: r)` | the first two, as `a.random!(rng: r)` spells them |
|
|
284
|
+
|
|
285
|
+
The keyword forms are the only keyword arguments in the subset and the only
|
|
286
|
+
bare names in it that Ruby itself does not have; both are paid for reading
|
|
287
|
+
like the array language. All of these read one generator, so mixing them in
|
|
288
|
+
a kernel walks one sequence rather than several.
|
|
289
|
+
|
|
290
|
+
A normal is two draws and keeps no spare. Box-Muller classically gives two
|
|
291
|
+
normals for two uniforms; the second would have to live in the generator's
|
|
292
|
+
state between calls, and a spare held across the boundary between an array
|
|
293
|
+
and a kernel is a second kind of state to keep in step. Two draws apiece
|
|
294
|
+
costs about a nanosecond against the ten the transform itself takes, and
|
|
295
|
+
buys a rule with nothing behind it.
|
|
296
|
+
|
|
297
|
+
Each generator has its own state, so two in one kernel are two sequences,
|
|
298
|
+
and the state survives the call -- a second kernel carries on rather than
|
|
299
|
+
starting again. `seed:` is data the kernel is handed and not part of it, so
|
|
300
|
+
every seed shares one compiled kernel.
|
|
301
|
+
|
|
302
|
+
A generator is one on both sides of the compiler, so a sequence can begin in
|
|
303
|
+
an array and continue in a kernel:
|
|
304
|
+
|
|
305
|
+
```ruby
|
|
306
|
+
a.random!(rng: rand) # CArray fills, advancing rand
|
|
307
|
+
CArray.jit_for(n) { |i| b[i] = rand.random } # the kernel takes the next draws
|
|
308
|
+
```
|
|
309
|
+
|
|
310
|
+
and those are the numbers one `random!` over both would have laid down.
|
|
311
|
+
CArray compiles the generator and hands the same text out for the kernel to
|
|
312
|
+
paste, so the two sides are one implementation rather than two that agree.
|
|
313
|
+
|
|
314
|
+
Which draw lands in which cell is the loop's order, which this compiler does
|
|
315
|
+
not fix. Where that matters -- common random numbers, antithetic variates --
|
|
316
|
+
fill an array with `CArray#random!` before the call and read a cell of it.
|
|
317
|
+
`jit_function` is the one place a draw will not go: a compiled function
|
|
318
|
+
takes everything through its parameters and has nowhere to keep a state, so
|
|
319
|
+
declare `int64_t state[4]` there and pass `rand.state`.
|
|
130
320
|
|
|
131
321
|
---
|
|
132
322
|
|
|
@@ -142,6 +332,7 @@ about either.
|
|
|
142
332
|
| `jit_contract` | a new `CArray`, or the `CompiledKernel` when the block assigns |
|
|
143
333
|
| `jit_extern` | `CFunction` |
|
|
144
334
|
| `jit_function` | `CFunction` |
|
|
335
|
+
| `jit_call` | what the function answered, or `nil` where it returns `void` |
|
|
145
336
|
|
|
146
337
|
Every `CompiledKernel` answers `#c_source` with the C that ran.
|
|
147
338
|
|
|
@@ -150,7 +341,9 @@ Every `CompiledKernel` answers `#c_source` with the C that ran.
|
|
|
150
341
|
| | Without a C compiler |
|
|
151
342
|
|---|---|
|
|
152
343
|
| `fuse` | works -- CArray walks it, same answer |
|
|
153
|
-
| the
|
|
344
|
+
| the six kernel forms and `jit_function` | `CArray::JIT::Unsupported` |
|
|
345
|
+
| `jit_call` | the same, unless a provider answers the site |
|
|
346
|
+
| `jit_extern` | works -- it compiles nothing |
|
|
154
347
|
|
|
155
348
|
That is what the prefix says. A block outside the subset is refused by name and
|
|
156
349
|
line rather than run as a Ruby loop: nobody reaches for a compiler except to
|
|
@@ -165,8 +358,8 @@ that was not asked.
|
|
|
165
358
|
| `jit_each` / `jit_map` with block parameters | an index means `jit_for` |
|
|
166
359
|
| `jit_stencil` with no array given | the arrays are arguments, not closures |
|
|
167
360
|
| `jit_stencil` with both `type:` and `into:` | `into:` already decides the type |
|
|
168
|
-
| a contraction summing an index that appears once |
|
|
169
|
-
| a
|
|
361
|
+
| a contraction summing an index that appears once | name the axes; with nothing named it is free |
|
|
362
|
+
| a left-hand axis the argument list leaves out | the list is all of the result's axes |
|
|
170
363
|
| an index whose axes disagree in extent | the shape check a contraction exists to do |
|
|
171
364
|
| an array both written and read in a contraction | a recurrence -- write it with `jit_for` |
|
|
172
365
|
| a block naming a construct outside the subset | refused by name and line |
|