thorcino 0.1.1__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (72) hide show
  1. thorcino-0.1.1/.gitignore +9 -0
  2. thorcino-0.1.1/.python-version +1 -0
  3. thorcino-0.1.1/LICENSE +21 -0
  4. thorcino-0.1.1/PKG-INFO +310 -0
  5. thorcino-0.1.1/README.md +288 -0
  6. thorcino-0.1.1/TODO.md +10 -0
  7. thorcino-0.1.1/core/__init__.py +0 -0
  8. thorcino-0.1.1/core/activations.py +45 -0
  9. thorcino-0.1.1/core/autograd/__init__.py +24 -0
  10. thorcino-0.1.1/core/autograd/activations.py +51 -0
  11. thorcino-0.1.1/core/autograd/arithmetic.py +119 -0
  12. thorcino-0.1.1/core/autograd/base.py +19 -0
  13. thorcino-0.1.1/core/autograd/losses.py +42 -0
  14. thorcino-0.1.1/core/consts.py +2 -0
  15. thorcino-0.1.1/core/dataset/__init__.py +4 -0
  16. thorcino-0.1.1/core/dataset/dataset.py +86 -0
  17. thorcino-0.1.1/core/dataset/transformation.py +39 -0
  18. thorcino-0.1.1/core/dataset/utils.py +3 -0
  19. thorcino-0.1.1/core/functions.py +89 -0
  20. thorcino-0.1.1/core/graph.py +270 -0
  21. thorcino-0.1.1/core/layers.py +184 -0
  22. thorcino-0.1.1/core/losses.py +47 -0
  23. thorcino-0.1.1/core/optimizer.py +201 -0
  24. thorcino-0.1.1/core/tensor.py +209 -0
  25. thorcino-0.1.1/core/training/__init__.py +9 -0
  26. thorcino-0.1.1/core/training/schedulers.py +33 -0
  27. thorcino-0.1.1/core/training/trainer.py +189 -0
  28. thorcino-0.1.1/core/utils.py +11 -0
  29. thorcino-0.1.1/examples/linear_regression/README.md +108 -0
  30. thorcino-0.1.1/examples/linear_regression/cubic/arch/architecture.png +0 -0
  31. thorcino-0.1.1/examples/linear_regression/cubic/arch/backward.png +0 -0
  32. thorcino-0.1.1/examples/linear_regression/cubic/arch/forward.png +0 -0
  33. thorcino-0.1.1/examples/linear_regression/cubic/cubic.py +95 -0
  34. thorcino-0.1.1/examples/linear_regression/cubic/results.png +0 -0
  35. thorcino-0.1.1/examples/linear_regression/helpers/dataset.py +87 -0
  36. thorcino-0.1.1/examples/linear_regression/helpers/graph.py +47 -0
  37. thorcino-0.1.1/examples/linear_regression/helpers/train.py +59 -0
  38. thorcino-0.1.1/examples/linear_regression/ill-cond/README.md +230 -0
  39. thorcino-0.1.1/examples/linear_regression/ill-cond/images/feature_correlation.png +0 -0
  40. thorcino-0.1.1/examples/linear_regression/ill-cond/images/loss_curves.png +0 -0
  41. thorcino-0.1.1/examples/linear_regression/ill-cond/images/stability_boxplot.png +0 -0
  42. thorcino-0.1.1/examples/linear_regression/ill-cond/images/weights_comparison.png +0 -0
  43. thorcino-0.1.1/examples/linear_regression/ill-cond/main.ipynb +847 -0
  44. thorcino-0.1.1/examples/linear_regression/linear/README.md +135 -0
  45. thorcino-0.1.1/examples/linear_regression/linear/arch/WIL.md +5 -0
  46. thorcino-0.1.1/examples/linear_regression/linear/arch/architecture.png +0 -0
  47. thorcino-0.1.1/examples/linear_regression/linear/arch/backward.png +0 -0
  48. thorcino-0.1.1/examples/linear_regression/linear/arch/forward.png +0 -0
  49. thorcino-0.1.1/examples/linear_regression/linear/linear.py +104 -0
  50. thorcino-0.1.1/examples/linear_regression/linear/results.png +0 -0
  51. thorcino-0.1.1/examples/linear_regression/quadratic/README.md +144 -0
  52. thorcino-0.1.1/examples/linear_regression/quadratic/arch/architecture.png +0 -0
  53. thorcino-0.1.1/examples/linear_regression/quadratic/arch/backward.png +0 -0
  54. thorcino-0.1.1/examples/linear_regression/quadratic/arch/forward.png +0 -0
  55. thorcino-0.1.1/examples/linear_regression/quadratic/quadratic.py +95 -0
  56. thorcino-0.1.1/examples/linear_regression/quadratic/results.png +0 -0
  57. thorcino-0.1.1/examples/linear_regression/regularization/DL2/README.md +102 -0
  58. thorcino-0.1.1/examples/linear_regression/regularization/DL2/main.ipynb +503 -0
  59. thorcino-0.1.1/examples/linear_regression/regularization/DL2/test.ipynb +76 -0
  60. thorcino-0.1.1/examples/linear_regression/regularization/L2/README.md +128 -0
  61. thorcino-0.1.1/examples/linear_regression/regularization/L2/images/dataset_overview.png +0 -0
  62. thorcino-0.1.1/examples/linear_regression/regularization/L2/images/step1_no_regularization.png +0 -0
  63. thorcino-0.1.1/examples/linear_regression/regularization/L2/images/step2_weight_decay.png +0 -0
  64. thorcino-0.1.1/examples/linear_regression/regularization/L2/main.ipynb +436 -0
  65. thorcino-0.1.1/examples/linear_regression/regularization/README.md +11 -0
  66. thorcino-0.1.1/examples/linear_regression/variance/README.md +165 -0
  67. thorcino-0.1.1/examples/linear_regression/variance/images/datasets.png +0 -0
  68. thorcino-0.1.1/examples/linear_regression/variance/images/losses.png +0 -0
  69. thorcino-0.1.1/examples/linear_regression/variance/images/weights.png +0 -0
  70. thorcino-0.1.1/examples/linear_regression/variance/main.ipynb +268 -0
  71. thorcino-0.1.1/main.py +5 -0
  72. thorcino-0.1.1/pyproject.toml +50 -0
@@ -0,0 +1,9 @@
1
+ .venv/
2
+ .claude/
3
+ .vscode/
4
+ tmp/
5
+ dist/
6
+ uv.lock
7
+ *.pyc
8
+ __pycache__/
9
+ core/__pycache__/
@@ -0,0 +1 @@
1
+ 3.14
thorcino-0.1.1/LICENSE ADDED
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 cecinuga
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
@@ -0,0 +1,310 @@
1
+ Metadata-Version: 2.4
2
+ Name: thorcino
3
+ Version: 0.1.1
4
+ Summary: A minimal, educational deep-learning framework built from scratch on top of NumPy
5
+ Project-URL: Repository, https://github.com/cecinuga/tiny-torch
6
+ Author-email: cecinuga <cecinuga.tdb@gmail.com>
7
+ License-Expression: MIT
8
+ License-File: LICENSE
9
+ Keywords: autograd,deep-learning,education,neural-network,numpy
10
+ Classifier: Development Status :: 3 - Alpha
11
+ Classifier: Intended Audience :: Education
12
+ Classifier: Intended Audience :: Science/Research
13
+ Classifier: Programming Language :: Python :: 3
14
+ Classifier: Programming Language :: Python :: 3.14
15
+ Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
16
+ Requires-Python: >=3.14
17
+ Requires-Dist: graphviz>=0.21
18
+ Requires-Dist: numpy>=2.5.0
19
+ Requires-Dist: pandas>=3.0.3
20
+ Requires-Dist: scikit-learn>=1.9.0
21
+ Description-Content-Type: text/markdown
22
+
23
+ # tiny-torch
24
+
25
+ A minimal, educational deep-learning framework built from scratch on top of NumPy.
26
+ `tiny-torch` reimplements the essential pieces of a PyTorch-style workflow — a
27
+ tensor with reverse-mode automatic differentiation, a small set of layers,
28
+ activation and loss functions, and a data-loading pipeline — in a few hundred
29
+ lines of readable Python.
30
+
31
+ The goal is not performance but clarity: every gradient is computed by hand in an
32
+ explicit backward class, so you can read exactly how backpropagation flows through
33
+ the computation graph.
34
+
35
+ ## Requirements
36
+
37
+ - Python >= 3.14
38
+ - NumPy >= 2.5.0
39
+ - Graphviz >= 0.21
40
+
41
+ Optional (dev): `ipykernel`, `matplotlib` (used by the examples).
42
+
43
+ ## Installation
44
+
45
+ The project uses [uv](https://github.com/astral-sh/uv) for environment and
46
+ dependency management (`uv.lock` is committed).
47
+
48
+ ```bash
49
+ uv sync
50
+ ```
51
+
52
+ This creates a `.venv` and installs the runtime and dev dependencies. Prebuilt
53
+ artifacts for `tiny_torch-0.1.0` are also available under `dist/`.
54
+
55
+ ## Architecture
56
+
57
+ The whole library lives under `core/` and is organized in a small number of
58
+ single-responsibility modules:
59
+
60
+ ```
61
+ core/
62
+ ├── tensor.py # Tensor: the numpy-backed frontend + operator overloading
63
+ ├── functions.py # Pure numpy math: activations, softmax, loss functions
64
+ ├── activations.py # Activation layers (ReLU, Sigmoid, Tanh, GELU, Softmax)
65
+ ├── losses.py # Loss objects (MSE, CrossEntropy, BinaryCrossEntropy)
66
+ ├── layers.py # Layer, Linear, Dropout, Sequential
67
+ ├── optimizer.py # Optimizer, SGD, SGDM, Adam, AdamW
68
+ ├── graph.py # ComputationalGraph: graphviz visualisation of a Sequential model
69
+ ├── utils.py # unbroadcast() helper for gradient reduction
70
+ ├── autograd/ # Reverse-mode automatic differentiation
71
+ │ ├── base.py # Function: base class for every backward node
72
+ │ ├── arithmetic.py # Add/Sub/Mul/Div/Matmul/Sum/Reshape/Transpose backward
73
+ │ ├── activations.py # ReLU/Sigmoid/Tanh/GELU/Softmax backward
74
+ │ └── losses.py # MSE/CrossEntropy/BCE backward
75
+ ├── dataset/ # Data loading pipeline
76
+ │ ├── dataset.py # Dataset, TensorDataset, ImageDataset, DataLoader
77
+ │ ├── transformation.py # RandomHorizontalFlip, RandomCrop, Compose
78
+ │ └── utils.py # image loading helpers
79
+ └── training/ # Training loop orchestration
80
+ ├── trainer.py # Trainer: train_epoch/eval, checkpointing, grad clipping
81
+ └── schedulers.py # Schedule, CosineSchedule
82
+ ```
83
+
84
+ The design follows a clear **frontend / backend split**:
85
+
86
+ - `Tensor` is the *frontend*. It wraps a NumPy array, overloads the Python
87
+ operators (`+`, `-`, `*`, `/`, `@`, …) and records the operation that produced
88
+ it in a `_grad_fn` attribute.
89
+ - The `autograd` package is the *backend*. Every operation has a matching
90
+ `*Backward` class (a `Function`) that knows how to turn an upstream gradient
91
+ into the gradients of its inputs.
92
+
93
+ Layers, activations and losses are thin objects that call into `functions.py` for
94
+ the forward pass and attach the corresponding `Function` for the backward pass.
95
+
96
+ ## Automatic differentiation (`core/autograd/`)
97
+
98
+ `tiny-torch` implements **reverse-mode autodiff** by building a dynamic graph as
99
+ operations execute (define-by-run), then walking it backwards to accumulate
100
+ gradients.
101
+
102
+ ### The `Function` node
103
+
104
+ Every backward node subclasses `Function` (`core/autograd/base.py`):
105
+
106
+ ```python
107
+ class Function:
108
+ def __init__(self, *tensors):
109
+ self.saved_tensors = tensors # inputs needed by backward
110
+ self.next_functions = [t._grad_fn for t in tensors] # links to parent nodes
111
+
112
+ def apply(self, grad_output):
113
+ """Turn the upstream gradient into gradients for each input."""
114
+ raise NotImplementedError()
115
+ ```
116
+
117
+ - `saved_tensors` holds the operands captured during the forward pass.
118
+ - `next_functions` records each operand's own `_grad_fn`, which is what turns the
119
+ set of nodes into a traversable graph.
120
+ - `apply(grad_output)` implements the chain rule for that specific operation and
121
+ returns one gradient per input.
122
+
123
+ ### How the graph is built
124
+
125
+ When you write `c = a + b`, `Tensor.__add__` computes the numeric result with
126
+ NumPy and attaches the backward node:
127
+
128
+ ```python
129
+ out = Tensor(self.data + other.data)
130
+ out._grad_fn = AddBackward(self, other)
131
+ ```
132
+
133
+ So each output tensor remembers *how it was produced*. Chaining operations
134
+ produces a graph of `Function` nodes rooted at the final output.
135
+
136
+ ### The backward pass
137
+
138
+ `Tensor.backward()` (`core/tensor.py`) drives backpropagation recursively:
139
+
140
+ 1. If no gradient is supplied, it seeds `1.0` for a scalar output (and raises for
141
+ non-scalar outputs, matching PyTorch's behaviour).
142
+ 2. It accumulates the incoming gradient into `self.grad` (gradients **add up**,
143
+ which is what makes shared subgraphs correct).
144
+ 3. It calls `self._grad_fn.apply(gradient)` to get the input gradients, then
145
+ recurses into each input tensor that `requires_grad`.
146
+
147
+ Broadcasting is handled by `unbroadcast()` (`core/utils.py`), which sums a
148
+ gradient back down to the shape of the original operand so that broadcasted
149
+ operations (e.g. adding a bias vector to a batch) produce correctly-shaped
150
+ gradients.
151
+
152
+ ### Managing the graph
153
+
154
+ - `Tensor.zero_grad()` resets a tensor's accumulated gradient.
155
+ - `Tensor.destroy_graph()` walks the graph and drops every `_grad_fn`, freeing the
156
+ saved tensors so the graph can be garbage-collected between iterations.
157
+
158
+ ### Supported backward operations
159
+
160
+ | Category | Backward classes |
161
+ |-------------|------------------|
162
+ | Arithmetic | `AddBackward`, `SubBackward`, `MulBackward`, `DivBackward` |
163
+ | Linear alg. | `MatmulBackward`, `TransposeBackward` |
164
+ | Reductions | `SumBackward` |
165
+ | Shape | `ReshapeBackward` |
166
+ | Activations | `ReLUBackward`, `SigmoidBackward`, `TanhBackward`, `GELUBackward`, `SoftmaxBackward` |
167
+ | Losses | `MSELossBackward`, `CrossEntropyLossBackward`, `BCELossBackward` |
168
+
169
+ ## The `Tensor` class (`core/tensor.py`)
170
+
171
+ `Tensor` is a lightweight wrapper around a `np.ndarray` (always stored as
172
+ `float32`). It exposes:
173
+
174
+ - **Metadata**: `data`, `shape`, `size`, `dtype`, `requires_grad`, `grad`,
175
+ `_grad_fn`.
176
+ - **Operator overloading**: `__add__`/`__radd__`, `__sub__`/`__rsub__`,
177
+ `__mul__`/`__rmul__`, `__truediv__`, `__matmul__`, `__pow__`, `__neg__`,
178
+ `__gt__`. The autograd-aware operations (`+`, `-`, `*`, `/`, `@`) attach a
179
+ `_grad_fn`; scalar/`ndarray` fast paths return plain results.
180
+ - **Tensor ops**: `matmul`, `reshape` (supports `-1` inference), `transpose`,
181
+ `sum`, `mean`, `max`, `min`.
182
+ - **Autograd control**: `backward()`, `zero_grad()`, `destroy_graph()`.
183
+ - **Interop**: `numpy()` returns the underlying array.
184
+
185
+ A convenience path in `__init__` lets you build a batched tensor from a list of
186
+ tensors — `Tensor([t1, t2, ...])` stacks their data automatically.
187
+
188
+ Note that `core.tensor` imports the backward classes at the *bottom* of the file,
189
+ after `Tensor` is defined, to break the circular import between the tensor
190
+ frontend and the autograd backend (the backward classes need `Tensor` at runtime).
191
+
192
+ ## Layers (`core/layers.py`)
193
+
194
+ All layers derive from the abstract `Layer` base class, which defines `forward()`,
195
+ makes instances callable, and exposes a `parameters` property.
196
+
197
+ | Layer | Description |
198
+ |--------------|-------------|
199
+ | `Linear` | Fully-connected layer `y = xW + b` with Xavier weight initialization and optional bias. |
200
+ | `Dropout` | Inverted dropout with keep-probability scaling; a no-op when `training=False`. |
201
+ | `Sequential` | Chains layers and forwards through them in order; aggregates their parameters. |
202
+
203
+ `Sequential.save_graph(path, arch=True, forward=False, backward=False)` renders a
204
+ `.png` of the model via `core/graph.py` (needs `graphviz`): a cluster per layer
205
+ for the architecture, and — if requested — the forward/backward computational
206
+ graphs built from a synthetic input, tensors colour-coded by role
207
+ (input/weights/bias/hidden).
208
+
209
+ ## Activation functions (`core/activations.py`)
210
+
211
+ Each activation is available both as a pure NumPy function (`core/functions.py`)
212
+ and as an autograd-aware `Layer`:
213
+
214
+ | Activation | Notes |
215
+ |------------|-------|
216
+ | `ReLU` | `max(0, x)` |
217
+ | `Sigmoid` | Numerically stable (branch on the sign of the input) |
218
+ | `Tanh` | `np.tanh` |
219
+ | `GELU` | Sigmoid approximation `x · σ(1.702·x)` |
220
+ | `Softmax` | Max-shifted for stability; configurable `dim` |
221
+
222
+ `functions.py` also provides a stable `log_softmax`, used internally by the
223
+ cross-entropy loss.
224
+
225
+ ## Loss functions (`core/losses.py`)
226
+
227
+ | Loss | Input | Notes |
228
+ |--------------------------|-------|-------|
229
+ | `MSELoss` | predictions, targets | Mean squared error. |
230
+ | `CrossEntropyLoss` | logits, integer targets | Combines a stable `log_softmax` with negative log-likelihood; the backward is the classic `softmax(logits) − onehot(targets)`. |
231
+ | `BinaryCrossEntropyLoss` | probabilities, targets | Clips predictions to `[1e-7, 1 − 1e-7]` to avoid `log(0)`. |
232
+
233
+ Each loss is callable (`loss(pred, target)`) and returns a scalar `Tensor` you can
234
+ call `.backward()` on.
235
+
236
+ ## Optimizers (`core/optimizer.py`)
237
+
238
+ | Optimizer | Notes |
239
+ |-----------|-------|
240
+ | `SGD` | Plain gradient descent with optional L2 weight decay. |
241
+ | `SGDM` | SGD with momentum. |
242
+ | `Adam` | Adaptive moments with bias correction. |
243
+ | `AdamW` | Adam with decoupled weight decay. |
244
+
245
+ Every optimizer takes `model.parameters` and a learning rate; `step()` updates
246
+ `param.data` in place, `zero_grad()` clears `param.grad`, and `get_state()`
247
+ returns the optimizer's hyperparameters/buffers for checkpointing.
248
+
249
+ ## Training loop (`core/training/`)
250
+
251
+ - **`Trainer`** (`trainer.py`) wraps a model, loss, optimizer and optional
252
+ scheduler. `train_epoch(dataloader, accumulation_steps=1)` runs one epoch
253
+ (with gradient accumulation and optional `clip_grad_norm` clipping) and
254
+ `eval(dataloader)` runs a no-grad pass, both logging into `trainer.history`
255
+ (`train_loss`, `eval_loss`, `lr`). `save()`/`load()` (de)serialize training
256
+ state to a checkpoint file via `pickle`.
257
+ - **`Schedule`** (`schedulers.py`) is the abstract base for learning-rate
258
+ schedules; `CosineSchedule(max_lr, min_lr, total_epochs)` anneals the
259
+ learning rate from `max_lr` to `min_lr` following a cosine curve, applied by
260
+ `Trainer` at the end of every `train_epoch()` call.
261
+
262
+ ## Data loading (`core/dataset/`)
263
+
264
+ The module mirrors the PyTorch `Dataset` / `DataLoader` pattern.
265
+
266
+ - **`Dataset`** — abstract base defining `__len__` and `__getitem__`.
267
+ - **`TensorDataset`** — wraps in-memory tensors and validates that they share the
268
+ same length along dimension 0.
269
+ - **`ImageDataset`** — lazily loads images from disk on access (via `load_jpeg`),
270
+ pairing each with its label.
271
+ - **`DataLoader`** — iterates a `Dataset` in mini-batches, with optional
272
+ shuffling, and collates each batch by stacking samples along a new leading
273
+ (batch) axis.
274
+
275
+ Data augmentation transforms live in `transformation.py`:
276
+
277
+ - `RandomHorizontalFlip(p)` — flips along the width axis with probability `p`.
278
+ - `RandomCrop(height, width, padding)` — zero-pads then crops a random window.
279
+ - `Compose([...])` — chains transforms into a single callable.
280
+
281
+ ## Examples
282
+
283
+ See [`examples/linear_regression/`](examples/linear_regression/) for a family
284
+ of end-to-end regression scripts and notebooks built on `Trainer` +
285
+ `DataLoader`:
286
+
287
+ | Example | Description |
288
+ |---|---|
289
+ | [`linear/`](examples/linear_regression/linear/README.md) | Recovers the slope/intercept of `2·x + 5` with a single `Linear(1, 1)` layer; compares the learned fit against the closed-form `lstsq` solution. |
290
+ | [`quadratic/`](examples/linear_regression/quadratic/README.md) | Recovers the coefficients of `x² + 2·x + 2` via feature expansion (`[x, x²]`) fed into a `Linear(2, 1)` layer. |
291
+ | [`cubic/`](examples/linear_regression/cubic/cubic.py) | Same idea one degree further: recovers `1.2·x³ − 2.3·x² + 2·x + 2` with a `Linear(3, 1)` layer over `[x, x², x³]`. |
292
+ | [`ill-cond/`](examples/linear_regression/ill-cond/README.md) | Notebook comparing closed-form OLS vs. SGD on ill-conditioned (near-collinear) features, showing OLS's coefficients blow up while SGD's stay stable. |
293
+ | [`variance/`](examples/linear_regression/variance/README.md) | Notebook fitting 20 repeated `Linear(1, 1)` models at each of nine label-noise levels, showing via boxplots how the loss and learned slope/intercept drift and spread as noise grows. |
294
+
295
+ The [`linear_regression/README.md`](examples/linear_regression/README.md)
296
+ ties the linear/quadratic/cubic scripts together and explains why the learning
297
+ rate has to shrink as the polynomial degree grows.
298
+
299
+ ```bash
300
+ uv run python examples/linear_regression/linear/linear.py
301
+ ```
302
+
303
+ ## Roadmap
304
+
305
+ Planned work is tracked in [`TODO.md`](TODO.md), and includes:
306
+
307
+ - **Loader**: parallel loading via multithreading; prefetching of the next batch.
308
+ - **Autograd**: move backward computation entirely onto NumPy arrays (keeping
309
+ `Tensor` as a pure frontend); cache forward intermediates for reuse in backward;
310
+ add a debug step that reports which node a backward failure occurred on.
@@ -0,0 +1,288 @@
1
+ # tiny-torch
2
+
3
+ A minimal, educational deep-learning framework built from scratch on top of NumPy.
4
+ `tiny-torch` reimplements the essential pieces of a PyTorch-style workflow — a
5
+ tensor with reverse-mode automatic differentiation, a small set of layers,
6
+ activation and loss functions, and a data-loading pipeline — in a few hundred
7
+ lines of readable Python.
8
+
9
+ The goal is not performance but clarity: every gradient is computed by hand in an
10
+ explicit backward class, so you can read exactly how backpropagation flows through
11
+ the computation graph.
12
+
13
+ ## Requirements
14
+
15
+ - Python >= 3.14
16
+ - NumPy >= 2.5.0
17
+ - Graphviz >= 0.21
18
+
19
+ Optional (dev): `ipykernel`, `matplotlib` (used by the examples).
20
+
21
+ ## Installation
22
+
23
+ The project uses [uv](https://github.com/astral-sh/uv) for environment and
24
+ dependency management (`uv.lock` is committed).
25
+
26
+ ```bash
27
+ uv sync
28
+ ```
29
+
30
+ This creates a `.venv` and installs the runtime and dev dependencies. Prebuilt
31
+ artifacts for `tiny_torch-0.1.0` are also available under `dist/`.
32
+
33
+ ## Architecture
34
+
35
+ The whole library lives under `core/` and is organized in a small number of
36
+ single-responsibility modules:
37
+
38
+ ```
39
+ core/
40
+ ├── tensor.py # Tensor: the numpy-backed frontend + operator overloading
41
+ ├── functions.py # Pure numpy math: activations, softmax, loss functions
42
+ ├── activations.py # Activation layers (ReLU, Sigmoid, Tanh, GELU, Softmax)
43
+ ├── losses.py # Loss objects (MSE, CrossEntropy, BinaryCrossEntropy)
44
+ ├── layers.py # Layer, Linear, Dropout, Sequential
45
+ ├── optimizer.py # Optimizer, SGD, SGDM, Adam, AdamW
46
+ ├── graph.py # ComputationalGraph: graphviz visualisation of a Sequential model
47
+ ├── utils.py # unbroadcast() helper for gradient reduction
48
+ ├── autograd/ # Reverse-mode automatic differentiation
49
+ │ ├── base.py # Function: base class for every backward node
50
+ │ ├── arithmetic.py # Add/Sub/Mul/Div/Matmul/Sum/Reshape/Transpose backward
51
+ │ ├── activations.py # ReLU/Sigmoid/Tanh/GELU/Softmax backward
52
+ │ └── losses.py # MSE/CrossEntropy/BCE backward
53
+ ├── dataset/ # Data loading pipeline
54
+ │ ├── dataset.py # Dataset, TensorDataset, ImageDataset, DataLoader
55
+ │ ├── transformation.py # RandomHorizontalFlip, RandomCrop, Compose
56
+ │ └── utils.py # image loading helpers
57
+ └── training/ # Training loop orchestration
58
+ ├── trainer.py # Trainer: train_epoch/eval, checkpointing, grad clipping
59
+ └── schedulers.py # Schedule, CosineSchedule
60
+ ```
61
+
62
+ The design follows a clear **frontend / backend split**:
63
+
64
+ - `Tensor` is the *frontend*. It wraps a NumPy array, overloads the Python
65
+ operators (`+`, `-`, `*`, `/`, `@`, …) and records the operation that produced
66
+ it in a `_grad_fn` attribute.
67
+ - The `autograd` package is the *backend*. Every operation has a matching
68
+ `*Backward` class (a `Function`) that knows how to turn an upstream gradient
69
+ into the gradients of its inputs.
70
+
71
+ Layers, activations and losses are thin objects that call into `functions.py` for
72
+ the forward pass and attach the corresponding `Function` for the backward pass.
73
+
74
+ ## Automatic differentiation (`core/autograd/`)
75
+
76
+ `tiny-torch` implements **reverse-mode autodiff** by building a dynamic graph as
77
+ operations execute (define-by-run), then walking it backwards to accumulate
78
+ gradients.
79
+
80
+ ### The `Function` node
81
+
82
+ Every backward node subclasses `Function` (`core/autograd/base.py`):
83
+
84
+ ```python
85
+ class Function:
86
+ def __init__(self, *tensors):
87
+ self.saved_tensors = tensors # inputs needed by backward
88
+ self.next_functions = [t._grad_fn for t in tensors] # links to parent nodes
89
+
90
+ def apply(self, grad_output):
91
+ """Turn the upstream gradient into gradients for each input."""
92
+ raise NotImplementedError()
93
+ ```
94
+
95
+ - `saved_tensors` holds the operands captured during the forward pass.
96
+ - `next_functions` records each operand's own `_grad_fn`, which is what turns the
97
+ set of nodes into a traversable graph.
98
+ - `apply(grad_output)` implements the chain rule for that specific operation and
99
+ returns one gradient per input.
100
+
101
+ ### How the graph is built
102
+
103
+ When you write `c = a + b`, `Tensor.__add__` computes the numeric result with
104
+ NumPy and attaches the backward node:
105
+
106
+ ```python
107
+ out = Tensor(self.data + other.data)
108
+ out._grad_fn = AddBackward(self, other)
109
+ ```
110
+
111
+ So each output tensor remembers *how it was produced*. Chaining operations
112
+ produces a graph of `Function` nodes rooted at the final output.
113
+
114
+ ### The backward pass
115
+
116
+ `Tensor.backward()` (`core/tensor.py`) drives backpropagation recursively:
117
+
118
+ 1. If no gradient is supplied, it seeds `1.0` for a scalar output (and raises for
119
+ non-scalar outputs, matching PyTorch's behaviour).
120
+ 2. It accumulates the incoming gradient into `self.grad` (gradients **add up**,
121
+ which is what makes shared subgraphs correct).
122
+ 3. It calls `self._grad_fn.apply(gradient)` to get the input gradients, then
123
+ recurses into each input tensor that `requires_grad`.
124
+
125
+ Broadcasting is handled by `unbroadcast()` (`core/utils.py`), which sums a
126
+ gradient back down to the shape of the original operand so that broadcasted
127
+ operations (e.g. adding a bias vector to a batch) produce correctly-shaped
128
+ gradients.
129
+
130
+ ### Managing the graph
131
+
132
+ - `Tensor.zero_grad()` resets a tensor's accumulated gradient.
133
+ - `Tensor.destroy_graph()` walks the graph and drops every `_grad_fn`, freeing the
134
+ saved tensors so the graph can be garbage-collected between iterations.
135
+
136
+ ### Supported backward operations
137
+
138
+ | Category | Backward classes |
139
+ |-------------|------------------|
140
+ | Arithmetic | `AddBackward`, `SubBackward`, `MulBackward`, `DivBackward` |
141
+ | Linear alg. | `MatmulBackward`, `TransposeBackward` |
142
+ | Reductions | `SumBackward` |
143
+ | Shape | `ReshapeBackward` |
144
+ | Activations | `ReLUBackward`, `SigmoidBackward`, `TanhBackward`, `GELUBackward`, `SoftmaxBackward` |
145
+ | Losses | `MSELossBackward`, `CrossEntropyLossBackward`, `BCELossBackward` |
146
+
147
+ ## The `Tensor` class (`core/tensor.py`)
148
+
149
+ `Tensor` is a lightweight wrapper around a `np.ndarray` (always stored as
150
+ `float32`). It exposes:
151
+
152
+ - **Metadata**: `data`, `shape`, `size`, `dtype`, `requires_grad`, `grad`,
153
+ `_grad_fn`.
154
+ - **Operator overloading**: `__add__`/`__radd__`, `__sub__`/`__rsub__`,
155
+ `__mul__`/`__rmul__`, `__truediv__`, `__matmul__`, `__pow__`, `__neg__`,
156
+ `__gt__`. The autograd-aware operations (`+`, `-`, `*`, `/`, `@`) attach a
157
+ `_grad_fn`; scalar/`ndarray` fast paths return plain results.
158
+ - **Tensor ops**: `matmul`, `reshape` (supports `-1` inference), `transpose`,
159
+ `sum`, `mean`, `max`, `min`.
160
+ - **Autograd control**: `backward()`, `zero_grad()`, `destroy_graph()`.
161
+ - **Interop**: `numpy()` returns the underlying array.
162
+
163
+ A convenience path in `__init__` lets you build a batched tensor from a list of
164
+ tensors — `Tensor([t1, t2, ...])` stacks their data automatically.
165
+
166
+ Note that `core.tensor` imports the backward classes at the *bottom* of the file,
167
+ after `Tensor` is defined, to break the circular import between the tensor
168
+ frontend and the autograd backend (the backward classes need `Tensor` at runtime).
169
+
170
+ ## Layers (`core/layers.py`)
171
+
172
+ All layers derive from the abstract `Layer` base class, which defines `forward()`,
173
+ makes instances callable, and exposes a `parameters` property.
174
+
175
+ | Layer | Description |
176
+ |--------------|-------------|
177
+ | `Linear` | Fully-connected layer `y = xW + b` with Xavier weight initialization and optional bias. |
178
+ | `Dropout` | Inverted dropout with keep-probability scaling; a no-op when `training=False`. |
179
+ | `Sequential` | Chains layers and forwards through them in order; aggregates their parameters. |
180
+
181
+ `Sequential.save_graph(path, arch=True, forward=False, backward=False)` renders a
182
+ `.png` of the model via `core/graph.py` (needs `graphviz`): a cluster per layer
183
+ for the architecture, and — if requested — the forward/backward computational
184
+ graphs built from a synthetic input, tensors colour-coded by role
185
+ (input/weights/bias/hidden).
186
+
187
+ ## Activation functions (`core/activations.py`)
188
+
189
+ Each activation is available both as a pure NumPy function (`core/functions.py`)
190
+ and as an autograd-aware `Layer`:
191
+
192
+ | Activation | Notes |
193
+ |------------|-------|
194
+ | `ReLU` | `max(0, x)` |
195
+ | `Sigmoid` | Numerically stable (branch on the sign of the input) |
196
+ | `Tanh` | `np.tanh` |
197
+ | `GELU` | Sigmoid approximation `x · σ(1.702·x)` |
198
+ | `Softmax` | Max-shifted for stability; configurable `dim` |
199
+
200
+ `functions.py` also provides a stable `log_softmax`, used internally by the
201
+ cross-entropy loss.
202
+
203
+ ## Loss functions (`core/losses.py`)
204
+
205
+ | Loss | Input | Notes |
206
+ |--------------------------|-------|-------|
207
+ | `MSELoss` | predictions, targets | Mean squared error. |
208
+ | `CrossEntropyLoss` | logits, integer targets | Combines a stable `log_softmax` with negative log-likelihood; the backward is the classic `softmax(logits) − onehot(targets)`. |
209
+ | `BinaryCrossEntropyLoss` | probabilities, targets | Clips predictions to `[1e-7, 1 − 1e-7]` to avoid `log(0)`. |
210
+
211
+ Each loss is callable (`loss(pred, target)`) and returns a scalar `Tensor` you can
212
+ call `.backward()` on.
213
+
214
+ ## Optimizers (`core/optimizer.py`)
215
+
216
+ | Optimizer | Notes |
217
+ |-----------|-------|
218
+ | `SGD` | Plain gradient descent with optional L2 weight decay. |
219
+ | `SGDM` | SGD with momentum. |
220
+ | `Adam` | Adaptive moments with bias correction. |
221
+ | `AdamW` | Adam with decoupled weight decay. |
222
+
223
+ Every optimizer takes `model.parameters` and a learning rate; `step()` updates
224
+ `param.data` in place, `zero_grad()` clears `param.grad`, and `get_state()`
225
+ returns the optimizer's hyperparameters/buffers for checkpointing.
226
+
227
+ ## Training loop (`core/training/`)
228
+
229
+ - **`Trainer`** (`trainer.py`) wraps a model, loss, optimizer and optional
230
+ scheduler. `train_epoch(dataloader, accumulation_steps=1)` runs one epoch
231
+ (with gradient accumulation and optional `clip_grad_norm` clipping) and
232
+ `eval(dataloader)` runs a no-grad pass, both logging into `trainer.history`
233
+ (`train_loss`, `eval_loss`, `lr`). `save()`/`load()` (de)serialize training
234
+ state to a checkpoint file via `pickle`.
235
+ - **`Schedule`** (`schedulers.py`) is the abstract base for learning-rate
236
+ schedules; `CosineSchedule(max_lr, min_lr, total_epochs)` anneals the
237
+ learning rate from `max_lr` to `min_lr` following a cosine curve, applied by
238
+ `Trainer` at the end of every `train_epoch()` call.
239
+
240
+ ## Data loading (`core/dataset/`)
241
+
242
+ The module mirrors the PyTorch `Dataset` / `DataLoader` pattern.
243
+
244
+ - **`Dataset`** — abstract base defining `__len__` and `__getitem__`.
245
+ - **`TensorDataset`** — wraps in-memory tensors and validates that they share the
246
+ same length along dimension 0.
247
+ - **`ImageDataset`** — lazily loads images from disk on access (via `load_jpeg`),
248
+ pairing each with its label.
249
+ - **`DataLoader`** — iterates a `Dataset` in mini-batches, with optional
250
+ shuffling, and collates each batch by stacking samples along a new leading
251
+ (batch) axis.
252
+
253
+ Data augmentation transforms live in `transformation.py`:
254
+
255
+ - `RandomHorizontalFlip(p)` — flips along the width axis with probability `p`.
256
+ - `RandomCrop(height, width, padding)` — zero-pads then crops a random window.
257
+ - `Compose([...])` — chains transforms into a single callable.
258
+
259
+ ## Examples
260
+
261
+ See [`examples/linear_regression/`](examples/linear_regression/) for a family
262
+ of end-to-end regression scripts and notebooks built on `Trainer` +
263
+ `DataLoader`:
264
+
265
+ | Example | Description |
266
+ |---|---|
267
+ | [`linear/`](examples/linear_regression/linear/README.md) | Recovers the slope/intercept of `2·x + 5` with a single `Linear(1, 1)` layer; compares the learned fit against the closed-form `lstsq` solution. |
268
+ | [`quadratic/`](examples/linear_regression/quadratic/README.md) | Recovers the coefficients of `x² + 2·x + 2` via feature expansion (`[x, x²]`) fed into a `Linear(2, 1)` layer. |
269
+ | [`cubic/`](examples/linear_regression/cubic/cubic.py) | Same idea one degree further: recovers `1.2·x³ − 2.3·x² + 2·x + 2` with a `Linear(3, 1)` layer over `[x, x², x³]`. |
270
+ | [`ill-cond/`](examples/linear_regression/ill-cond/README.md) | Notebook comparing closed-form OLS vs. SGD on ill-conditioned (near-collinear) features, showing OLS's coefficients blow up while SGD's stay stable. |
271
+ | [`variance/`](examples/linear_regression/variance/README.md) | Notebook fitting 20 repeated `Linear(1, 1)` models at each of nine label-noise levels, showing via boxplots how the loss and learned slope/intercept drift and spread as noise grows. |
272
+
273
+ The [`linear_regression/README.md`](examples/linear_regression/README.md)
274
+ ties the linear/quadratic/cubic scripts together and explains why the learning
275
+ rate has to shrink as the polynomial degree grows.
276
+
277
+ ```bash
278
+ uv run python examples/linear_regression/linear/linear.py
279
+ ```
280
+
281
+ ## Roadmap
282
+
283
+ Planned work is tracked in [`TODO.md`](TODO.md), and includes:
284
+
285
+ - **Loader**: parallel loading via multithreading; prefetching of the next batch.
286
+ - **Autograd**: move backward computation entirely onto NumPy arrays (keeping
287
+ `Tensor` as a pure frontend); cache forward intermediates for reuse in backward;
288
+ add a debug step that reports which node a backward failure occurred on.
thorcino-0.1.1/TODO.md ADDED
@@ -0,0 +1,10 @@
1
+ [loader](core/loader.py)
2
+ 1) add parallel loading via multi threading
3
+ 2) add pre fetching of N+1 batches
4
+
5
+ [autograd](core/autograd/arithmetic.py)
6
+ [autograd](core/autograd/activations.py)
7
+ [autograd](core/autograd/losses.py)
8
+ 1) Makes backward pass computation over numpy arrays and let Tensor class be only the frontend
9
+ 2) Add Cache: instend of recomputing base functions, cache it from forward pass and let backward pass reuse it
10
+ 3) Add Debug step when backward fails: show on which node the failure occurred
File without changes