thorcino 0.1.1__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- thorcino-0.1.1/.gitignore +9 -0
- thorcino-0.1.1/.python-version +1 -0
- thorcino-0.1.1/LICENSE +21 -0
- thorcino-0.1.1/PKG-INFO +310 -0
- thorcino-0.1.1/README.md +288 -0
- thorcino-0.1.1/TODO.md +10 -0
- thorcino-0.1.1/core/__init__.py +0 -0
- thorcino-0.1.1/core/activations.py +45 -0
- thorcino-0.1.1/core/autograd/__init__.py +24 -0
- thorcino-0.1.1/core/autograd/activations.py +51 -0
- thorcino-0.1.1/core/autograd/arithmetic.py +119 -0
- thorcino-0.1.1/core/autograd/base.py +19 -0
- thorcino-0.1.1/core/autograd/losses.py +42 -0
- thorcino-0.1.1/core/consts.py +2 -0
- thorcino-0.1.1/core/dataset/__init__.py +4 -0
- thorcino-0.1.1/core/dataset/dataset.py +86 -0
- thorcino-0.1.1/core/dataset/transformation.py +39 -0
- thorcino-0.1.1/core/dataset/utils.py +3 -0
- thorcino-0.1.1/core/functions.py +89 -0
- thorcino-0.1.1/core/graph.py +270 -0
- thorcino-0.1.1/core/layers.py +184 -0
- thorcino-0.1.1/core/losses.py +47 -0
- thorcino-0.1.1/core/optimizer.py +201 -0
- thorcino-0.1.1/core/tensor.py +209 -0
- thorcino-0.1.1/core/training/__init__.py +9 -0
- thorcino-0.1.1/core/training/schedulers.py +33 -0
- thorcino-0.1.1/core/training/trainer.py +189 -0
- thorcino-0.1.1/core/utils.py +11 -0
- thorcino-0.1.1/examples/linear_regression/README.md +108 -0
- thorcino-0.1.1/examples/linear_regression/cubic/arch/architecture.png +0 -0
- thorcino-0.1.1/examples/linear_regression/cubic/arch/backward.png +0 -0
- thorcino-0.1.1/examples/linear_regression/cubic/arch/forward.png +0 -0
- thorcino-0.1.1/examples/linear_regression/cubic/cubic.py +95 -0
- thorcino-0.1.1/examples/linear_regression/cubic/results.png +0 -0
- thorcino-0.1.1/examples/linear_regression/helpers/dataset.py +87 -0
- thorcino-0.1.1/examples/linear_regression/helpers/graph.py +47 -0
- thorcino-0.1.1/examples/linear_regression/helpers/train.py +59 -0
- thorcino-0.1.1/examples/linear_regression/ill-cond/README.md +230 -0
- thorcino-0.1.1/examples/linear_regression/ill-cond/images/feature_correlation.png +0 -0
- thorcino-0.1.1/examples/linear_regression/ill-cond/images/loss_curves.png +0 -0
- thorcino-0.1.1/examples/linear_regression/ill-cond/images/stability_boxplot.png +0 -0
- thorcino-0.1.1/examples/linear_regression/ill-cond/images/weights_comparison.png +0 -0
- thorcino-0.1.1/examples/linear_regression/ill-cond/main.ipynb +847 -0
- thorcino-0.1.1/examples/linear_regression/linear/README.md +135 -0
- thorcino-0.1.1/examples/linear_regression/linear/arch/WIL.md +5 -0
- thorcino-0.1.1/examples/linear_regression/linear/arch/architecture.png +0 -0
- thorcino-0.1.1/examples/linear_regression/linear/arch/backward.png +0 -0
- thorcino-0.1.1/examples/linear_regression/linear/arch/forward.png +0 -0
- thorcino-0.1.1/examples/linear_regression/linear/linear.py +104 -0
- thorcino-0.1.1/examples/linear_regression/linear/results.png +0 -0
- thorcino-0.1.1/examples/linear_regression/quadratic/README.md +144 -0
- thorcino-0.1.1/examples/linear_regression/quadratic/arch/architecture.png +0 -0
- thorcino-0.1.1/examples/linear_regression/quadratic/arch/backward.png +0 -0
- thorcino-0.1.1/examples/linear_regression/quadratic/arch/forward.png +0 -0
- thorcino-0.1.1/examples/linear_regression/quadratic/quadratic.py +95 -0
- thorcino-0.1.1/examples/linear_regression/quadratic/results.png +0 -0
- thorcino-0.1.1/examples/linear_regression/regularization/DL2/README.md +102 -0
- thorcino-0.1.1/examples/linear_regression/regularization/DL2/main.ipynb +503 -0
- thorcino-0.1.1/examples/linear_regression/regularization/DL2/test.ipynb +76 -0
- thorcino-0.1.1/examples/linear_regression/regularization/L2/README.md +128 -0
- thorcino-0.1.1/examples/linear_regression/regularization/L2/images/dataset_overview.png +0 -0
- thorcino-0.1.1/examples/linear_regression/regularization/L2/images/step1_no_regularization.png +0 -0
- thorcino-0.1.1/examples/linear_regression/regularization/L2/images/step2_weight_decay.png +0 -0
- thorcino-0.1.1/examples/linear_regression/regularization/L2/main.ipynb +436 -0
- thorcino-0.1.1/examples/linear_regression/regularization/README.md +11 -0
- thorcino-0.1.1/examples/linear_regression/variance/README.md +165 -0
- thorcino-0.1.1/examples/linear_regression/variance/images/datasets.png +0 -0
- thorcino-0.1.1/examples/linear_regression/variance/images/losses.png +0 -0
- thorcino-0.1.1/examples/linear_regression/variance/images/weights.png +0 -0
- thorcino-0.1.1/examples/linear_regression/variance/main.ipynb +268 -0
- thorcino-0.1.1/main.py +5 -0
- thorcino-0.1.1/pyproject.toml +50 -0
|
@@ -0,0 +1 @@
|
|
|
1
|
+
3.14
|
thorcino-0.1.1/LICENSE
ADDED
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 cecinuga
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|
thorcino-0.1.1/PKG-INFO
ADDED
|
@@ -0,0 +1,310 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: thorcino
|
|
3
|
+
Version: 0.1.1
|
|
4
|
+
Summary: A minimal, educational deep-learning framework built from scratch on top of NumPy
|
|
5
|
+
Project-URL: Repository, https://github.com/cecinuga/tiny-torch
|
|
6
|
+
Author-email: cecinuga <cecinuga.tdb@gmail.com>
|
|
7
|
+
License-Expression: MIT
|
|
8
|
+
License-File: LICENSE
|
|
9
|
+
Keywords: autograd,deep-learning,education,neural-network,numpy
|
|
10
|
+
Classifier: Development Status :: 3 - Alpha
|
|
11
|
+
Classifier: Intended Audience :: Education
|
|
12
|
+
Classifier: Intended Audience :: Science/Research
|
|
13
|
+
Classifier: Programming Language :: Python :: 3
|
|
14
|
+
Classifier: Programming Language :: Python :: 3.14
|
|
15
|
+
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
|
|
16
|
+
Requires-Python: >=3.14
|
|
17
|
+
Requires-Dist: graphviz>=0.21
|
|
18
|
+
Requires-Dist: numpy>=2.5.0
|
|
19
|
+
Requires-Dist: pandas>=3.0.3
|
|
20
|
+
Requires-Dist: scikit-learn>=1.9.0
|
|
21
|
+
Description-Content-Type: text/markdown
|
|
22
|
+
|
|
23
|
+
# tiny-torch
|
|
24
|
+
|
|
25
|
+
A minimal, educational deep-learning framework built from scratch on top of NumPy.
|
|
26
|
+
`tiny-torch` reimplements the essential pieces of a PyTorch-style workflow — a
|
|
27
|
+
tensor with reverse-mode automatic differentiation, a small set of layers,
|
|
28
|
+
activation and loss functions, and a data-loading pipeline — in a few hundred
|
|
29
|
+
lines of readable Python.
|
|
30
|
+
|
|
31
|
+
The goal is not performance but clarity: every gradient is computed by hand in an
|
|
32
|
+
explicit backward class, so you can read exactly how backpropagation flows through
|
|
33
|
+
the computation graph.
|
|
34
|
+
|
|
35
|
+
## Requirements
|
|
36
|
+
|
|
37
|
+
- Python >= 3.14
|
|
38
|
+
- NumPy >= 2.5.0
|
|
39
|
+
- Graphviz >= 0.21
|
|
40
|
+
|
|
41
|
+
Optional (dev): `ipykernel`, `matplotlib` (used by the examples).
|
|
42
|
+
|
|
43
|
+
## Installation
|
|
44
|
+
|
|
45
|
+
The project uses [uv](https://github.com/astral-sh/uv) for environment and
|
|
46
|
+
dependency management (`uv.lock` is committed).
|
|
47
|
+
|
|
48
|
+
```bash
|
|
49
|
+
uv sync
|
|
50
|
+
```
|
|
51
|
+
|
|
52
|
+
This creates a `.venv` and installs the runtime and dev dependencies. Prebuilt
|
|
53
|
+
artifacts for `tiny_torch-0.1.0` are also available under `dist/`.
|
|
54
|
+
|
|
55
|
+
## Architecture
|
|
56
|
+
|
|
57
|
+
The whole library lives under `core/` and is organized in a small number of
|
|
58
|
+
single-responsibility modules:
|
|
59
|
+
|
|
60
|
+
```
|
|
61
|
+
core/
|
|
62
|
+
├── tensor.py # Tensor: the numpy-backed frontend + operator overloading
|
|
63
|
+
├── functions.py # Pure numpy math: activations, softmax, loss functions
|
|
64
|
+
├── activations.py # Activation layers (ReLU, Sigmoid, Tanh, GELU, Softmax)
|
|
65
|
+
├── losses.py # Loss objects (MSE, CrossEntropy, BinaryCrossEntropy)
|
|
66
|
+
├── layers.py # Layer, Linear, Dropout, Sequential
|
|
67
|
+
├── optimizer.py # Optimizer, SGD, SGDM, Adam, AdamW
|
|
68
|
+
├── graph.py # ComputationalGraph: graphviz visualisation of a Sequential model
|
|
69
|
+
├── utils.py # unbroadcast() helper for gradient reduction
|
|
70
|
+
├── autograd/ # Reverse-mode automatic differentiation
|
|
71
|
+
│ ├── base.py # Function: base class for every backward node
|
|
72
|
+
│ ├── arithmetic.py # Add/Sub/Mul/Div/Matmul/Sum/Reshape/Transpose backward
|
|
73
|
+
│ ├── activations.py # ReLU/Sigmoid/Tanh/GELU/Softmax backward
|
|
74
|
+
│ └── losses.py # MSE/CrossEntropy/BCE backward
|
|
75
|
+
├── dataset/ # Data loading pipeline
|
|
76
|
+
│ ├── dataset.py # Dataset, TensorDataset, ImageDataset, DataLoader
|
|
77
|
+
│ ├── transformation.py # RandomHorizontalFlip, RandomCrop, Compose
|
|
78
|
+
│ └── utils.py # image loading helpers
|
|
79
|
+
└── training/ # Training loop orchestration
|
|
80
|
+
├── trainer.py # Trainer: train_epoch/eval, checkpointing, grad clipping
|
|
81
|
+
└── schedulers.py # Schedule, CosineSchedule
|
|
82
|
+
```
|
|
83
|
+
|
|
84
|
+
The design follows a clear **frontend / backend split**:
|
|
85
|
+
|
|
86
|
+
- `Tensor` is the *frontend*. It wraps a NumPy array, overloads the Python
|
|
87
|
+
operators (`+`, `-`, `*`, `/`, `@`, …) and records the operation that produced
|
|
88
|
+
it in a `_grad_fn` attribute.
|
|
89
|
+
- The `autograd` package is the *backend*. Every operation has a matching
|
|
90
|
+
`*Backward` class (a `Function`) that knows how to turn an upstream gradient
|
|
91
|
+
into the gradients of its inputs.
|
|
92
|
+
|
|
93
|
+
Layers, activations and losses are thin objects that call into `functions.py` for
|
|
94
|
+
the forward pass and attach the corresponding `Function` for the backward pass.
|
|
95
|
+
|
|
96
|
+
## Automatic differentiation (`core/autograd/`)
|
|
97
|
+
|
|
98
|
+
`tiny-torch` implements **reverse-mode autodiff** by building a dynamic graph as
|
|
99
|
+
operations execute (define-by-run), then walking it backwards to accumulate
|
|
100
|
+
gradients.
|
|
101
|
+
|
|
102
|
+
### The `Function` node
|
|
103
|
+
|
|
104
|
+
Every backward node subclasses `Function` (`core/autograd/base.py`):
|
|
105
|
+
|
|
106
|
+
```python
|
|
107
|
+
class Function:
|
|
108
|
+
def __init__(self, *tensors):
|
|
109
|
+
self.saved_tensors = tensors # inputs needed by backward
|
|
110
|
+
self.next_functions = [t._grad_fn for t in tensors] # links to parent nodes
|
|
111
|
+
|
|
112
|
+
def apply(self, grad_output):
|
|
113
|
+
"""Turn the upstream gradient into gradients for each input."""
|
|
114
|
+
raise NotImplementedError()
|
|
115
|
+
```
|
|
116
|
+
|
|
117
|
+
- `saved_tensors` holds the operands captured during the forward pass.
|
|
118
|
+
- `next_functions` records each operand's own `_grad_fn`, which is what turns the
|
|
119
|
+
set of nodes into a traversable graph.
|
|
120
|
+
- `apply(grad_output)` implements the chain rule for that specific operation and
|
|
121
|
+
returns one gradient per input.
|
|
122
|
+
|
|
123
|
+
### How the graph is built
|
|
124
|
+
|
|
125
|
+
When you write `c = a + b`, `Tensor.__add__` computes the numeric result with
|
|
126
|
+
NumPy and attaches the backward node:
|
|
127
|
+
|
|
128
|
+
```python
|
|
129
|
+
out = Tensor(self.data + other.data)
|
|
130
|
+
out._grad_fn = AddBackward(self, other)
|
|
131
|
+
```
|
|
132
|
+
|
|
133
|
+
So each output tensor remembers *how it was produced*. Chaining operations
|
|
134
|
+
produces a graph of `Function` nodes rooted at the final output.
|
|
135
|
+
|
|
136
|
+
### The backward pass
|
|
137
|
+
|
|
138
|
+
`Tensor.backward()` (`core/tensor.py`) drives backpropagation recursively:
|
|
139
|
+
|
|
140
|
+
1. If no gradient is supplied, it seeds `1.0` for a scalar output (and raises for
|
|
141
|
+
non-scalar outputs, matching PyTorch's behaviour).
|
|
142
|
+
2. It accumulates the incoming gradient into `self.grad` (gradients **add up**,
|
|
143
|
+
which is what makes shared subgraphs correct).
|
|
144
|
+
3. It calls `self._grad_fn.apply(gradient)` to get the input gradients, then
|
|
145
|
+
recurses into each input tensor that `requires_grad`.
|
|
146
|
+
|
|
147
|
+
Broadcasting is handled by `unbroadcast()` (`core/utils.py`), which sums a
|
|
148
|
+
gradient back down to the shape of the original operand so that broadcasted
|
|
149
|
+
operations (e.g. adding a bias vector to a batch) produce correctly-shaped
|
|
150
|
+
gradients.
|
|
151
|
+
|
|
152
|
+
### Managing the graph
|
|
153
|
+
|
|
154
|
+
- `Tensor.zero_grad()` resets a tensor's accumulated gradient.
|
|
155
|
+
- `Tensor.destroy_graph()` walks the graph and drops every `_grad_fn`, freeing the
|
|
156
|
+
saved tensors so the graph can be garbage-collected between iterations.
|
|
157
|
+
|
|
158
|
+
### Supported backward operations
|
|
159
|
+
|
|
160
|
+
| Category | Backward classes |
|
|
161
|
+
|-------------|------------------|
|
|
162
|
+
| Arithmetic | `AddBackward`, `SubBackward`, `MulBackward`, `DivBackward` |
|
|
163
|
+
| Linear alg. | `MatmulBackward`, `TransposeBackward` |
|
|
164
|
+
| Reductions | `SumBackward` |
|
|
165
|
+
| Shape | `ReshapeBackward` |
|
|
166
|
+
| Activations | `ReLUBackward`, `SigmoidBackward`, `TanhBackward`, `GELUBackward`, `SoftmaxBackward` |
|
|
167
|
+
| Losses | `MSELossBackward`, `CrossEntropyLossBackward`, `BCELossBackward` |
|
|
168
|
+
|
|
169
|
+
## The `Tensor` class (`core/tensor.py`)
|
|
170
|
+
|
|
171
|
+
`Tensor` is a lightweight wrapper around a `np.ndarray` (always stored as
|
|
172
|
+
`float32`). It exposes:
|
|
173
|
+
|
|
174
|
+
- **Metadata**: `data`, `shape`, `size`, `dtype`, `requires_grad`, `grad`,
|
|
175
|
+
`_grad_fn`.
|
|
176
|
+
- **Operator overloading**: `__add__`/`__radd__`, `__sub__`/`__rsub__`,
|
|
177
|
+
`__mul__`/`__rmul__`, `__truediv__`, `__matmul__`, `__pow__`, `__neg__`,
|
|
178
|
+
`__gt__`. The autograd-aware operations (`+`, `-`, `*`, `/`, `@`) attach a
|
|
179
|
+
`_grad_fn`; scalar/`ndarray` fast paths return plain results.
|
|
180
|
+
- **Tensor ops**: `matmul`, `reshape` (supports `-1` inference), `transpose`,
|
|
181
|
+
`sum`, `mean`, `max`, `min`.
|
|
182
|
+
- **Autograd control**: `backward()`, `zero_grad()`, `destroy_graph()`.
|
|
183
|
+
- **Interop**: `numpy()` returns the underlying array.
|
|
184
|
+
|
|
185
|
+
A convenience path in `__init__` lets you build a batched tensor from a list of
|
|
186
|
+
tensors — `Tensor([t1, t2, ...])` stacks their data automatically.
|
|
187
|
+
|
|
188
|
+
Note that `core.tensor` imports the backward classes at the *bottom* of the file,
|
|
189
|
+
after `Tensor` is defined, to break the circular import between the tensor
|
|
190
|
+
frontend and the autograd backend (the backward classes need `Tensor` at runtime).
|
|
191
|
+
|
|
192
|
+
## Layers (`core/layers.py`)
|
|
193
|
+
|
|
194
|
+
All layers derive from the abstract `Layer` base class, which defines `forward()`,
|
|
195
|
+
makes instances callable, and exposes a `parameters` property.
|
|
196
|
+
|
|
197
|
+
| Layer | Description |
|
|
198
|
+
|--------------|-------------|
|
|
199
|
+
| `Linear` | Fully-connected layer `y = xW + b` with Xavier weight initialization and optional bias. |
|
|
200
|
+
| `Dropout` | Inverted dropout with keep-probability scaling; a no-op when `training=False`. |
|
|
201
|
+
| `Sequential` | Chains layers and forwards through them in order; aggregates their parameters. |
|
|
202
|
+
|
|
203
|
+
`Sequential.save_graph(path, arch=True, forward=False, backward=False)` renders a
|
|
204
|
+
`.png` of the model via `core/graph.py` (needs `graphviz`): a cluster per layer
|
|
205
|
+
for the architecture, and — if requested — the forward/backward computational
|
|
206
|
+
graphs built from a synthetic input, tensors colour-coded by role
|
|
207
|
+
(input/weights/bias/hidden).
|
|
208
|
+
|
|
209
|
+
## Activation functions (`core/activations.py`)
|
|
210
|
+
|
|
211
|
+
Each activation is available both as a pure NumPy function (`core/functions.py`)
|
|
212
|
+
and as an autograd-aware `Layer`:
|
|
213
|
+
|
|
214
|
+
| Activation | Notes |
|
|
215
|
+
|------------|-------|
|
|
216
|
+
| `ReLU` | `max(0, x)` |
|
|
217
|
+
| `Sigmoid` | Numerically stable (branch on the sign of the input) |
|
|
218
|
+
| `Tanh` | `np.tanh` |
|
|
219
|
+
| `GELU` | Sigmoid approximation `x · σ(1.702·x)` |
|
|
220
|
+
| `Softmax` | Max-shifted for stability; configurable `dim` |
|
|
221
|
+
|
|
222
|
+
`functions.py` also provides a stable `log_softmax`, used internally by the
|
|
223
|
+
cross-entropy loss.
|
|
224
|
+
|
|
225
|
+
## Loss functions (`core/losses.py`)
|
|
226
|
+
|
|
227
|
+
| Loss | Input | Notes |
|
|
228
|
+
|--------------------------|-------|-------|
|
|
229
|
+
| `MSELoss` | predictions, targets | Mean squared error. |
|
|
230
|
+
| `CrossEntropyLoss` | logits, integer targets | Combines a stable `log_softmax` with negative log-likelihood; the backward is the classic `softmax(logits) − onehot(targets)`. |
|
|
231
|
+
| `BinaryCrossEntropyLoss` | probabilities, targets | Clips predictions to `[1e-7, 1 − 1e-7]` to avoid `log(0)`. |
|
|
232
|
+
|
|
233
|
+
Each loss is callable (`loss(pred, target)`) and returns a scalar `Tensor` you can
|
|
234
|
+
call `.backward()` on.
|
|
235
|
+
|
|
236
|
+
## Optimizers (`core/optimizer.py`)
|
|
237
|
+
|
|
238
|
+
| Optimizer | Notes |
|
|
239
|
+
|-----------|-------|
|
|
240
|
+
| `SGD` | Plain gradient descent with optional L2 weight decay. |
|
|
241
|
+
| `SGDM` | SGD with momentum. |
|
|
242
|
+
| `Adam` | Adaptive moments with bias correction. |
|
|
243
|
+
| `AdamW` | Adam with decoupled weight decay. |
|
|
244
|
+
|
|
245
|
+
Every optimizer takes `model.parameters` and a learning rate; `step()` updates
|
|
246
|
+
`param.data` in place, `zero_grad()` clears `param.grad`, and `get_state()`
|
|
247
|
+
returns the optimizer's hyperparameters/buffers for checkpointing.
|
|
248
|
+
|
|
249
|
+
## Training loop (`core/training/`)
|
|
250
|
+
|
|
251
|
+
- **`Trainer`** (`trainer.py`) wraps a model, loss, optimizer and optional
|
|
252
|
+
scheduler. `train_epoch(dataloader, accumulation_steps=1)` runs one epoch
|
|
253
|
+
(with gradient accumulation and optional `clip_grad_norm` clipping) and
|
|
254
|
+
`eval(dataloader)` runs a no-grad pass, both logging into `trainer.history`
|
|
255
|
+
(`train_loss`, `eval_loss`, `lr`). `save()`/`load()` (de)serialize training
|
|
256
|
+
state to a checkpoint file via `pickle`.
|
|
257
|
+
- **`Schedule`** (`schedulers.py`) is the abstract base for learning-rate
|
|
258
|
+
schedules; `CosineSchedule(max_lr, min_lr, total_epochs)` anneals the
|
|
259
|
+
learning rate from `max_lr` to `min_lr` following a cosine curve, applied by
|
|
260
|
+
`Trainer` at the end of every `train_epoch()` call.
|
|
261
|
+
|
|
262
|
+
## Data loading (`core/dataset/`)
|
|
263
|
+
|
|
264
|
+
The module mirrors the PyTorch `Dataset` / `DataLoader` pattern.
|
|
265
|
+
|
|
266
|
+
- **`Dataset`** — abstract base defining `__len__` and `__getitem__`.
|
|
267
|
+
- **`TensorDataset`** — wraps in-memory tensors and validates that they share the
|
|
268
|
+
same length along dimension 0.
|
|
269
|
+
- **`ImageDataset`** — lazily loads images from disk on access (via `load_jpeg`),
|
|
270
|
+
pairing each with its label.
|
|
271
|
+
- **`DataLoader`** — iterates a `Dataset` in mini-batches, with optional
|
|
272
|
+
shuffling, and collates each batch by stacking samples along a new leading
|
|
273
|
+
(batch) axis.
|
|
274
|
+
|
|
275
|
+
Data augmentation transforms live in `transformation.py`:
|
|
276
|
+
|
|
277
|
+
- `RandomHorizontalFlip(p)` — flips along the width axis with probability `p`.
|
|
278
|
+
- `RandomCrop(height, width, padding)` — zero-pads then crops a random window.
|
|
279
|
+
- `Compose([...])` — chains transforms into a single callable.
|
|
280
|
+
|
|
281
|
+
## Examples
|
|
282
|
+
|
|
283
|
+
See [`examples/linear_regression/`](examples/linear_regression/) for a family
|
|
284
|
+
of end-to-end regression scripts and notebooks built on `Trainer` +
|
|
285
|
+
`DataLoader`:
|
|
286
|
+
|
|
287
|
+
| Example | Description |
|
|
288
|
+
|---|---|
|
|
289
|
+
| [`linear/`](examples/linear_regression/linear/README.md) | Recovers the slope/intercept of `2·x + 5` with a single `Linear(1, 1)` layer; compares the learned fit against the closed-form `lstsq` solution. |
|
|
290
|
+
| [`quadratic/`](examples/linear_regression/quadratic/README.md) | Recovers the coefficients of `x² + 2·x + 2` via feature expansion (`[x, x²]`) fed into a `Linear(2, 1)` layer. |
|
|
291
|
+
| [`cubic/`](examples/linear_regression/cubic/cubic.py) | Same idea one degree further: recovers `1.2·x³ − 2.3·x² + 2·x + 2` with a `Linear(3, 1)` layer over `[x, x², x³]`. |
|
|
292
|
+
| [`ill-cond/`](examples/linear_regression/ill-cond/README.md) | Notebook comparing closed-form OLS vs. SGD on ill-conditioned (near-collinear) features, showing OLS's coefficients blow up while SGD's stay stable. |
|
|
293
|
+
| [`variance/`](examples/linear_regression/variance/README.md) | Notebook fitting 20 repeated `Linear(1, 1)` models at each of nine label-noise levels, showing via boxplots how the loss and learned slope/intercept drift and spread as noise grows. |
|
|
294
|
+
|
|
295
|
+
The [`linear_regression/README.md`](examples/linear_regression/README.md)
|
|
296
|
+
ties the linear/quadratic/cubic scripts together and explains why the learning
|
|
297
|
+
rate has to shrink as the polynomial degree grows.
|
|
298
|
+
|
|
299
|
+
```bash
|
|
300
|
+
uv run python examples/linear_regression/linear/linear.py
|
|
301
|
+
```
|
|
302
|
+
|
|
303
|
+
## Roadmap
|
|
304
|
+
|
|
305
|
+
Planned work is tracked in [`TODO.md`](TODO.md), and includes:
|
|
306
|
+
|
|
307
|
+
- **Loader**: parallel loading via multithreading; prefetching of the next batch.
|
|
308
|
+
- **Autograd**: move backward computation entirely onto NumPy arrays (keeping
|
|
309
|
+
`Tensor` as a pure frontend); cache forward intermediates for reuse in backward;
|
|
310
|
+
add a debug step that reports which node a backward failure occurred on.
|
thorcino-0.1.1/README.md
ADDED
|
@@ -0,0 +1,288 @@
|
|
|
1
|
+
# tiny-torch
|
|
2
|
+
|
|
3
|
+
A minimal, educational deep-learning framework built from scratch on top of NumPy.
|
|
4
|
+
`tiny-torch` reimplements the essential pieces of a PyTorch-style workflow — a
|
|
5
|
+
tensor with reverse-mode automatic differentiation, a small set of layers,
|
|
6
|
+
activation and loss functions, and a data-loading pipeline — in a few hundred
|
|
7
|
+
lines of readable Python.
|
|
8
|
+
|
|
9
|
+
The goal is not performance but clarity: every gradient is computed by hand in an
|
|
10
|
+
explicit backward class, so you can read exactly how backpropagation flows through
|
|
11
|
+
the computation graph.
|
|
12
|
+
|
|
13
|
+
## Requirements
|
|
14
|
+
|
|
15
|
+
- Python >= 3.14
|
|
16
|
+
- NumPy >= 2.5.0
|
|
17
|
+
- Graphviz >= 0.21
|
|
18
|
+
|
|
19
|
+
Optional (dev): `ipykernel`, `matplotlib` (used by the examples).
|
|
20
|
+
|
|
21
|
+
## Installation
|
|
22
|
+
|
|
23
|
+
The project uses [uv](https://github.com/astral-sh/uv) for environment and
|
|
24
|
+
dependency management (`uv.lock` is committed).
|
|
25
|
+
|
|
26
|
+
```bash
|
|
27
|
+
uv sync
|
|
28
|
+
```
|
|
29
|
+
|
|
30
|
+
This creates a `.venv` and installs the runtime and dev dependencies. Prebuilt
|
|
31
|
+
artifacts for `tiny_torch-0.1.0` are also available under `dist/`.
|
|
32
|
+
|
|
33
|
+
## Architecture
|
|
34
|
+
|
|
35
|
+
The whole library lives under `core/` and is organized in a small number of
|
|
36
|
+
single-responsibility modules:
|
|
37
|
+
|
|
38
|
+
```
|
|
39
|
+
core/
|
|
40
|
+
├── tensor.py # Tensor: the numpy-backed frontend + operator overloading
|
|
41
|
+
├── functions.py # Pure numpy math: activations, softmax, loss functions
|
|
42
|
+
├── activations.py # Activation layers (ReLU, Sigmoid, Tanh, GELU, Softmax)
|
|
43
|
+
├── losses.py # Loss objects (MSE, CrossEntropy, BinaryCrossEntropy)
|
|
44
|
+
├── layers.py # Layer, Linear, Dropout, Sequential
|
|
45
|
+
├── optimizer.py # Optimizer, SGD, SGDM, Adam, AdamW
|
|
46
|
+
├── graph.py # ComputationalGraph: graphviz visualisation of a Sequential model
|
|
47
|
+
├── utils.py # unbroadcast() helper for gradient reduction
|
|
48
|
+
├── autograd/ # Reverse-mode automatic differentiation
|
|
49
|
+
│ ├── base.py # Function: base class for every backward node
|
|
50
|
+
│ ├── arithmetic.py # Add/Sub/Mul/Div/Matmul/Sum/Reshape/Transpose backward
|
|
51
|
+
│ ├── activations.py # ReLU/Sigmoid/Tanh/GELU/Softmax backward
|
|
52
|
+
│ └── losses.py # MSE/CrossEntropy/BCE backward
|
|
53
|
+
├── dataset/ # Data loading pipeline
|
|
54
|
+
│ ├── dataset.py # Dataset, TensorDataset, ImageDataset, DataLoader
|
|
55
|
+
│ ├── transformation.py # RandomHorizontalFlip, RandomCrop, Compose
|
|
56
|
+
│ └── utils.py # image loading helpers
|
|
57
|
+
└── training/ # Training loop orchestration
|
|
58
|
+
├── trainer.py # Trainer: train_epoch/eval, checkpointing, grad clipping
|
|
59
|
+
└── schedulers.py # Schedule, CosineSchedule
|
|
60
|
+
```
|
|
61
|
+
|
|
62
|
+
The design follows a clear **frontend / backend split**:
|
|
63
|
+
|
|
64
|
+
- `Tensor` is the *frontend*. It wraps a NumPy array, overloads the Python
|
|
65
|
+
operators (`+`, `-`, `*`, `/`, `@`, …) and records the operation that produced
|
|
66
|
+
it in a `_grad_fn` attribute.
|
|
67
|
+
- The `autograd` package is the *backend*. Every operation has a matching
|
|
68
|
+
`*Backward` class (a `Function`) that knows how to turn an upstream gradient
|
|
69
|
+
into the gradients of its inputs.
|
|
70
|
+
|
|
71
|
+
Layers, activations and losses are thin objects that call into `functions.py` for
|
|
72
|
+
the forward pass and attach the corresponding `Function` for the backward pass.
|
|
73
|
+
|
|
74
|
+
## Automatic differentiation (`core/autograd/`)
|
|
75
|
+
|
|
76
|
+
`tiny-torch` implements **reverse-mode autodiff** by building a dynamic graph as
|
|
77
|
+
operations execute (define-by-run), then walking it backwards to accumulate
|
|
78
|
+
gradients.
|
|
79
|
+
|
|
80
|
+
### The `Function` node
|
|
81
|
+
|
|
82
|
+
Every backward node subclasses `Function` (`core/autograd/base.py`):
|
|
83
|
+
|
|
84
|
+
```python
|
|
85
|
+
class Function:
|
|
86
|
+
def __init__(self, *tensors):
|
|
87
|
+
self.saved_tensors = tensors # inputs needed by backward
|
|
88
|
+
self.next_functions = [t._grad_fn for t in tensors] # links to parent nodes
|
|
89
|
+
|
|
90
|
+
def apply(self, grad_output):
|
|
91
|
+
"""Turn the upstream gradient into gradients for each input."""
|
|
92
|
+
raise NotImplementedError()
|
|
93
|
+
```
|
|
94
|
+
|
|
95
|
+
- `saved_tensors` holds the operands captured during the forward pass.
|
|
96
|
+
- `next_functions` records each operand's own `_grad_fn`, which is what turns the
|
|
97
|
+
set of nodes into a traversable graph.
|
|
98
|
+
- `apply(grad_output)` implements the chain rule for that specific operation and
|
|
99
|
+
returns one gradient per input.
|
|
100
|
+
|
|
101
|
+
### How the graph is built
|
|
102
|
+
|
|
103
|
+
When you write `c = a + b`, `Tensor.__add__` computes the numeric result with
|
|
104
|
+
NumPy and attaches the backward node:
|
|
105
|
+
|
|
106
|
+
```python
|
|
107
|
+
out = Tensor(self.data + other.data)
|
|
108
|
+
out._grad_fn = AddBackward(self, other)
|
|
109
|
+
```
|
|
110
|
+
|
|
111
|
+
So each output tensor remembers *how it was produced*. Chaining operations
|
|
112
|
+
produces a graph of `Function` nodes rooted at the final output.
|
|
113
|
+
|
|
114
|
+
### The backward pass
|
|
115
|
+
|
|
116
|
+
`Tensor.backward()` (`core/tensor.py`) drives backpropagation recursively:
|
|
117
|
+
|
|
118
|
+
1. If no gradient is supplied, it seeds `1.0` for a scalar output (and raises for
|
|
119
|
+
non-scalar outputs, matching PyTorch's behaviour).
|
|
120
|
+
2. It accumulates the incoming gradient into `self.grad` (gradients **add up**,
|
|
121
|
+
which is what makes shared subgraphs correct).
|
|
122
|
+
3. It calls `self._grad_fn.apply(gradient)` to get the input gradients, then
|
|
123
|
+
recurses into each input tensor that `requires_grad`.
|
|
124
|
+
|
|
125
|
+
Broadcasting is handled by `unbroadcast()` (`core/utils.py`), which sums a
|
|
126
|
+
gradient back down to the shape of the original operand so that broadcasted
|
|
127
|
+
operations (e.g. adding a bias vector to a batch) produce correctly-shaped
|
|
128
|
+
gradients.
|
|
129
|
+
|
|
130
|
+
### Managing the graph
|
|
131
|
+
|
|
132
|
+
- `Tensor.zero_grad()` resets a tensor's accumulated gradient.
|
|
133
|
+
- `Tensor.destroy_graph()` walks the graph and drops every `_grad_fn`, freeing the
|
|
134
|
+
saved tensors so the graph can be garbage-collected between iterations.
|
|
135
|
+
|
|
136
|
+
### Supported backward operations
|
|
137
|
+
|
|
138
|
+
| Category | Backward classes |
|
|
139
|
+
|-------------|------------------|
|
|
140
|
+
| Arithmetic | `AddBackward`, `SubBackward`, `MulBackward`, `DivBackward` |
|
|
141
|
+
| Linear alg. | `MatmulBackward`, `TransposeBackward` |
|
|
142
|
+
| Reductions | `SumBackward` |
|
|
143
|
+
| Shape | `ReshapeBackward` |
|
|
144
|
+
| Activations | `ReLUBackward`, `SigmoidBackward`, `TanhBackward`, `GELUBackward`, `SoftmaxBackward` |
|
|
145
|
+
| Losses | `MSELossBackward`, `CrossEntropyLossBackward`, `BCELossBackward` |
|
|
146
|
+
|
|
147
|
+
## The `Tensor` class (`core/tensor.py`)
|
|
148
|
+
|
|
149
|
+
`Tensor` is a lightweight wrapper around a `np.ndarray` (always stored as
|
|
150
|
+
`float32`). It exposes:
|
|
151
|
+
|
|
152
|
+
- **Metadata**: `data`, `shape`, `size`, `dtype`, `requires_grad`, `grad`,
|
|
153
|
+
`_grad_fn`.
|
|
154
|
+
- **Operator overloading**: `__add__`/`__radd__`, `__sub__`/`__rsub__`,
|
|
155
|
+
`__mul__`/`__rmul__`, `__truediv__`, `__matmul__`, `__pow__`, `__neg__`,
|
|
156
|
+
`__gt__`. The autograd-aware operations (`+`, `-`, `*`, `/`, `@`) attach a
|
|
157
|
+
`_grad_fn`; scalar/`ndarray` fast paths return plain results.
|
|
158
|
+
- **Tensor ops**: `matmul`, `reshape` (supports `-1` inference), `transpose`,
|
|
159
|
+
`sum`, `mean`, `max`, `min`.
|
|
160
|
+
- **Autograd control**: `backward()`, `zero_grad()`, `destroy_graph()`.
|
|
161
|
+
- **Interop**: `numpy()` returns the underlying array.
|
|
162
|
+
|
|
163
|
+
A convenience path in `__init__` lets you build a batched tensor from a list of
|
|
164
|
+
tensors — `Tensor([t1, t2, ...])` stacks their data automatically.
|
|
165
|
+
|
|
166
|
+
Note that `core.tensor` imports the backward classes at the *bottom* of the file,
|
|
167
|
+
after `Tensor` is defined, to break the circular import between the tensor
|
|
168
|
+
frontend and the autograd backend (the backward classes need `Tensor` at runtime).
|
|
169
|
+
|
|
170
|
+
## Layers (`core/layers.py`)
|
|
171
|
+
|
|
172
|
+
All layers derive from the abstract `Layer` base class, which defines `forward()`,
|
|
173
|
+
makes instances callable, and exposes a `parameters` property.
|
|
174
|
+
|
|
175
|
+
| Layer | Description |
|
|
176
|
+
|--------------|-------------|
|
|
177
|
+
| `Linear` | Fully-connected layer `y = xW + b` with Xavier weight initialization and optional bias. |
|
|
178
|
+
| `Dropout` | Inverted dropout with keep-probability scaling; a no-op when `training=False`. |
|
|
179
|
+
| `Sequential` | Chains layers and forwards through them in order; aggregates their parameters. |
|
|
180
|
+
|
|
181
|
+
`Sequential.save_graph(path, arch=True, forward=False, backward=False)` renders a
|
|
182
|
+
`.png` of the model via `core/graph.py` (needs `graphviz`): a cluster per layer
|
|
183
|
+
for the architecture, and — if requested — the forward/backward computational
|
|
184
|
+
graphs built from a synthetic input, tensors colour-coded by role
|
|
185
|
+
(input/weights/bias/hidden).
|
|
186
|
+
|
|
187
|
+
## Activation functions (`core/activations.py`)
|
|
188
|
+
|
|
189
|
+
Each activation is available both as a pure NumPy function (`core/functions.py`)
|
|
190
|
+
and as an autograd-aware `Layer`:
|
|
191
|
+
|
|
192
|
+
| Activation | Notes |
|
|
193
|
+
|------------|-------|
|
|
194
|
+
| `ReLU` | `max(0, x)` |
|
|
195
|
+
| `Sigmoid` | Numerically stable (branch on the sign of the input) |
|
|
196
|
+
| `Tanh` | `np.tanh` |
|
|
197
|
+
| `GELU` | Sigmoid approximation `x · σ(1.702·x)` |
|
|
198
|
+
| `Softmax` | Max-shifted for stability; configurable `dim` |
|
|
199
|
+
|
|
200
|
+
`functions.py` also provides a stable `log_softmax`, used internally by the
|
|
201
|
+
cross-entropy loss.
|
|
202
|
+
|
|
203
|
+
## Loss functions (`core/losses.py`)
|
|
204
|
+
|
|
205
|
+
| Loss | Input | Notes |
|
|
206
|
+
|--------------------------|-------|-------|
|
|
207
|
+
| `MSELoss` | predictions, targets | Mean squared error. |
|
|
208
|
+
| `CrossEntropyLoss` | logits, integer targets | Combines a stable `log_softmax` with negative log-likelihood; the backward is the classic `softmax(logits) − onehot(targets)`. |
|
|
209
|
+
| `BinaryCrossEntropyLoss` | probabilities, targets | Clips predictions to `[1e-7, 1 − 1e-7]` to avoid `log(0)`. |
|
|
210
|
+
|
|
211
|
+
Each loss is callable (`loss(pred, target)`) and returns a scalar `Tensor` you can
|
|
212
|
+
call `.backward()` on.
|
|
213
|
+
|
|
214
|
+
## Optimizers (`core/optimizer.py`)
|
|
215
|
+
|
|
216
|
+
| Optimizer | Notes |
|
|
217
|
+
|-----------|-------|
|
|
218
|
+
| `SGD` | Plain gradient descent with optional L2 weight decay. |
|
|
219
|
+
| `SGDM` | SGD with momentum. |
|
|
220
|
+
| `Adam` | Adaptive moments with bias correction. |
|
|
221
|
+
| `AdamW` | Adam with decoupled weight decay. |
|
|
222
|
+
|
|
223
|
+
Every optimizer takes `model.parameters` and a learning rate; `step()` updates
|
|
224
|
+
`param.data` in place, `zero_grad()` clears `param.grad`, and `get_state()`
|
|
225
|
+
returns the optimizer's hyperparameters/buffers for checkpointing.
|
|
226
|
+
|
|
227
|
+
## Training loop (`core/training/`)
|
|
228
|
+
|
|
229
|
+
- **`Trainer`** (`trainer.py`) wraps a model, loss, optimizer and optional
|
|
230
|
+
scheduler. `train_epoch(dataloader, accumulation_steps=1)` runs one epoch
|
|
231
|
+
(with gradient accumulation and optional `clip_grad_norm` clipping) and
|
|
232
|
+
`eval(dataloader)` runs a no-grad pass, both logging into `trainer.history`
|
|
233
|
+
(`train_loss`, `eval_loss`, `lr`). `save()`/`load()` (de)serialize training
|
|
234
|
+
state to a checkpoint file via `pickle`.
|
|
235
|
+
- **`Schedule`** (`schedulers.py`) is the abstract base for learning-rate
|
|
236
|
+
schedules; `CosineSchedule(max_lr, min_lr, total_epochs)` anneals the
|
|
237
|
+
learning rate from `max_lr` to `min_lr` following a cosine curve, applied by
|
|
238
|
+
`Trainer` at the end of every `train_epoch()` call.
|
|
239
|
+
|
|
240
|
+
## Data loading (`core/dataset/`)
|
|
241
|
+
|
|
242
|
+
The module mirrors the PyTorch `Dataset` / `DataLoader` pattern.
|
|
243
|
+
|
|
244
|
+
- **`Dataset`** — abstract base defining `__len__` and `__getitem__`.
|
|
245
|
+
- **`TensorDataset`** — wraps in-memory tensors and validates that they share the
|
|
246
|
+
same length along dimension 0.
|
|
247
|
+
- **`ImageDataset`** — lazily loads images from disk on access (via `load_jpeg`),
|
|
248
|
+
pairing each with its label.
|
|
249
|
+
- **`DataLoader`** — iterates a `Dataset` in mini-batches, with optional
|
|
250
|
+
shuffling, and collates each batch by stacking samples along a new leading
|
|
251
|
+
(batch) axis.
|
|
252
|
+
|
|
253
|
+
Data augmentation transforms live in `transformation.py`:
|
|
254
|
+
|
|
255
|
+
- `RandomHorizontalFlip(p)` — flips along the width axis with probability `p`.
|
|
256
|
+
- `RandomCrop(height, width, padding)` — zero-pads then crops a random window.
|
|
257
|
+
- `Compose([...])` — chains transforms into a single callable.
|
|
258
|
+
|
|
259
|
+
## Examples
|
|
260
|
+
|
|
261
|
+
See [`examples/linear_regression/`](examples/linear_regression/) for a family
|
|
262
|
+
of end-to-end regression scripts and notebooks built on `Trainer` +
|
|
263
|
+
`DataLoader`:
|
|
264
|
+
|
|
265
|
+
| Example | Description |
|
|
266
|
+
|---|---|
|
|
267
|
+
| [`linear/`](examples/linear_regression/linear/README.md) | Recovers the slope/intercept of `2·x + 5` with a single `Linear(1, 1)` layer; compares the learned fit against the closed-form `lstsq` solution. |
|
|
268
|
+
| [`quadratic/`](examples/linear_regression/quadratic/README.md) | Recovers the coefficients of `x² + 2·x + 2` via feature expansion (`[x, x²]`) fed into a `Linear(2, 1)` layer. |
|
|
269
|
+
| [`cubic/`](examples/linear_regression/cubic/cubic.py) | Same idea one degree further: recovers `1.2·x³ − 2.3·x² + 2·x + 2` with a `Linear(3, 1)` layer over `[x, x², x³]`. |
|
|
270
|
+
| [`ill-cond/`](examples/linear_regression/ill-cond/README.md) | Notebook comparing closed-form OLS vs. SGD on ill-conditioned (near-collinear) features, showing OLS's coefficients blow up while SGD's stay stable. |
|
|
271
|
+
| [`variance/`](examples/linear_regression/variance/README.md) | Notebook fitting 20 repeated `Linear(1, 1)` models at each of nine label-noise levels, showing via boxplots how the loss and learned slope/intercept drift and spread as noise grows. |
|
|
272
|
+
|
|
273
|
+
The [`linear_regression/README.md`](examples/linear_regression/README.md)
|
|
274
|
+
ties the linear/quadratic/cubic scripts together and explains why the learning
|
|
275
|
+
rate has to shrink as the polynomial degree grows.
|
|
276
|
+
|
|
277
|
+
```bash
|
|
278
|
+
uv run python examples/linear_regression/linear/linear.py
|
|
279
|
+
```
|
|
280
|
+
|
|
281
|
+
## Roadmap
|
|
282
|
+
|
|
283
|
+
Planned work is tracked in [`TODO.md`](TODO.md), and includes:
|
|
284
|
+
|
|
285
|
+
- **Loader**: parallel loading via multithreading; prefetching of the next batch.
|
|
286
|
+
- **Autograd**: move backward computation entirely onto NumPy arrays (keeping
|
|
287
|
+
`Tensor` as a pure frontend); cache forward intermediates for reuse in backward;
|
|
288
|
+
add a debug step that reports which node a backward failure occurred on.
|
thorcino-0.1.1/TODO.md
ADDED
|
@@ -0,0 +1,10 @@
|
|
|
1
|
+
[loader](core/loader.py)
|
|
2
|
+
1) add parallel loading via multi threading
|
|
3
|
+
2) add pre fetching of N+1 batches
|
|
4
|
+
|
|
5
|
+
[autograd](core/autograd/arithmetic.py)
|
|
6
|
+
[autograd](core/autograd/activations.py)
|
|
7
|
+
[autograd](core/autograd/losses.py)
|
|
8
|
+
1) Makes backward pass computation over numpy arrays and let Tensor class be only the frontend
|
|
9
|
+
2) Add Cache: instend of recomputing base functions, cache it from forward pass and let backward pass reuse it
|
|
10
|
+
3) Add Debug step when backward fails: show on which node the failure occurred
|
|
File without changes
|