pulseml 0.2.2__tar.gz → 0.2.3__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
pulseml-0.2.3/PKG-INFO ADDED
@@ -0,0 +1,648 @@
1
+ Metadata-Version: 2.4
2
+ Name: pulseml
3
+ Version: 0.2.3
4
+ Summary: A live ML training debugger for tracking, diagnosing, verifying, and patching training runs.
5
+ Author-email: Yash Patel <codeyash09@gmail.com>
6
+ License: Proprietary
7
+ Project-URL: Homepage, https://pulsedb.netlify.app/
8
+ Project-URL: Repository, https://github.com/codeyash09/PulseML
9
+ Project-URL: Issues, https://github.com/codeyash09/PulseML/issues
10
+ Project-URL: Documentation, https://pulsedb.netlify.app/
11
+ Keywords: machine-learning,deep-learning,debugging,ml-debugging,training,pytorch,tensorflow,numpy,cupy,jax,gradients,training-debugger
12
+ Classifier: Development Status :: 3 - Alpha
13
+ Classifier: Intended Audience :: Developers
14
+ Classifier: Intended Audience :: Science/Research
15
+ Classifier: License :: Other/Proprietary License
16
+ Classifier: Operating System :: OS Independent
17
+ Classifier: Programming Language :: Python :: 3
18
+ Classifier: Programming Language :: Python :: 3 :: Only
19
+ Classifier: Programming Language :: Python :: 3.9
20
+ Classifier: Programming Language :: Python :: 3.10
21
+ Classifier: Programming Language :: Python :: 3.11
22
+ Classifier: Programming Language :: Python :: 3.12
23
+ Classifier: Programming Language :: Python :: 3.13
24
+ Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
25
+ Classifier: Topic :: Software Development :: Debuggers
26
+ Requires-Python: >=3.9
27
+ Description-Content-Type: text/markdown
28
+ License-File: LICENSE
29
+ Requires-Dist: numpy
30
+ Requires-Dist: matplotlib
31
+ Requires-Dist: pillow
32
+ Requires-Dist: litellm
33
+ Requires-Dist: fpdf2
34
+ Provides-Extra: torch
35
+ Requires-Dist: torch; extra == "torch"
36
+ Provides-Extra: tensorflow
37
+ Requires-Dist: tensorflow; extra == "tensorflow"
38
+ Provides-Extra: cupy
39
+ Requires-Dist: cupy-cuda12x; extra == "cupy"
40
+ Provides-Extra: jax
41
+ Requires-Dist: jax; extra == "jax"
42
+ Dynamic: license-file
43
+
44
+ # PulseML
45
+
46
+ **Pulse** is a live ML training debugger for CLI and headless environments. It combines runtime observability, deterministic numerical checks, ML-specific linting, and an AI debugging agent to investigate failures while training is running.
47
+
48
+ **Track → Diagnose → Verify → Patch**
49
+
50
+ [Dashboard](https://pulsedashb.netlify.app/) · [GitHub](https://github.com/codeyash09/PulseML) · [PyPI](https://pypi.org/project/pulseml/)
51
+
52
+ ---
53
+
54
+ ## Installation
55
+
56
+ Pulse is available on PyPI.
57
+
58
+ ### Requirements
59
+
60
+ - Python 3.9+
61
+ - A supported backend: NumPy, PyTorch, TensorFlow, CuPy, or JAX
62
+ - An API key for an AI provider if you want to use the AI debugging agent
63
+
64
+ ### Install
65
+
66
+ ```bash
67
+ pip install pulseml
68
+ ```
69
+
70
+ ### Verify
71
+
72
+ ```bash
73
+ python -c "import pulse; print('Pulse installed successfully')"
74
+ ```
75
+
76
+ ### Start Pulse
77
+
78
+ Import `auto_track` and call it immediately before your training loop:
79
+
80
+ ```python
81
+ from pulse import auto_track
82
+
83
+ if __name__ == "__main__":
84
+ auto_track()
85
+
86
+ # Your training loop
87
+ for epoch in range(num_epochs):
88
+ # Training logic here
89
+ pass
90
+ ```
91
+
92
+ The `__main__` guard is particularly important in multiprocessing or process-spawning environments.
93
+
94
+ Once the process is running, Pulse discovers numeric variables available to the training process and exposes them through the CLI.
95
+
96
+ ### AI provider configuration
97
+
98
+ Pulse supports cloud and local AI providers.
99
+
100
+ Common provider environment variables include:
101
+
102
+ ```text
103
+ ANTHROPIC_API_KEY
104
+ OPENAI_API_KEY
105
+ GEMINI_API_KEY
106
+ DEEPSEEK_API_KEY
107
+ MISTRAL_API_KEY
108
+ OPENROUTER_API_KEY
109
+ ```
110
+
111
+ For example:
112
+
113
+ ```bash
114
+ export GEMINI_API_KEY="your-api-key"
115
+ ```
116
+
117
+ On Windows PowerShell:
118
+
119
+ ```powershell
120
+ $env:GEMINI_API_KEY="your-api-key"
121
+ ```
122
+
123
+ Do not place API keys directly in source code or commit them to a repository.
124
+
125
+ ---
126
+
127
+ ## Why Pulse
128
+
129
+ Machine-learning failures are often diagnosed from the wrong end of the problem:
130
+
131
+ ```text
132
+ Train
133
+ ↓
134
+ Wait
135
+ ↓
136
+ Training fails
137
+ ↓
138
+ Read logs
139
+ ↓
140
+ Guess
141
+ ↓
142
+ Change code
143
+ ↓
144
+ Train again
145
+ ```
146
+
147
+ Pulse is designed to move the debugging process into the training run itself:
148
+
149
+ ```text
150
+ Train
151
+ ↓
152
+ Track
153
+ ↓
154
+ Detect
155
+ ↓
156
+ Measure
157
+ ↓
158
+ Verify
159
+ ↓
160
+ Diagnose
161
+ ↓
162
+ Patch
163
+ ```
164
+
165
+ The useful evidence behind an ML failure may be a gradient change, activation distribution, tensor shape, normalization relationship, parameter update, or the exact point at which a numerical value began to diverge.
166
+
167
+ Pulse gives the debugging agent access to that evidence instead of requiring it to reconstruct the run from logs alone.
168
+
169
+ ---
170
+
171
+ ## Benchmark
172
+
173
+ Pulse was evaluated against Claude Code on a **28-case ML debugging benchmark** designed to test whether an AI coding agent could identify and repair injected ML faults.
174
+
175
+ ### Results
176
+
177
+ | Agent | Model | Score |
178
+ |---|---|---:|
179
+ | **Pulse** | Gemini 3.5 Flash Lite | **15 / 28** |
180
+ | **Claude Code** | Opus 5 | **18 / 28** |
181
+
182
+ Pulse achieved **15/28**, compared with **18/28** for Claude Code.
183
+
184
+ The benchmark was particularly informative because the task was not simply to generate code that ran. The agent needed to identify the injected mutation, determine whether it was actually responsible for the observed behavior, and apply the necessary correction.
185
+
186
+ ### What the benchmark exposed
187
+
188
+ Pulse's largest weakness was **diagnostic precision**.
189
+
190
+ In difficult cases, Pulse could recognize that something was wrong but fail to consistently:
191
+
192
+ 1. identify the exact mutation responsible for the failure;
193
+ 2. isolate the smallest relevant code region;
194
+ 3. distinguish the root cause from surrounding symptoms; or
195
+ 4. apply only the necessary fix instead of proposing a broader change.
196
+
197
+ This matters because ML debugging is different from generic code generation.
198
+
199
+ A debugger should not rewrite a training system simply because it can. It should be able to identify the specific value, argument, operation, or line responsible for the failure and make the smallest correct change.
200
+
201
+ ### How the benchmark changed Pulse
202
+
203
+ The benchmark directly motivated upgrades to the debugging pipeline:
204
+
205
+ - **Expanded `mllint`** with additional deterministic ML-specific checks.
206
+ - **Minimal-fix-first behavior**, prioritizing single-line or small adjacent-line fixes before refactors.
207
+ - **Whitespace-tolerant patch matching**, preventing valid fixes from being rejected because of formatting differences.
208
+ - **Retry-on-mismatch patching**, allowing the agent to re-quote the exact current code when a proposed patch does not match.
209
+ - **Stronger verification**, checking both logical correctness and whether the scope of the patch is unnecessarily broad.
210
+ - **More agentic debugging workflows**, reducing the amount of manual investigation the underlying model has to perform.
211
+
212
+ The benchmark is not presented as proof that Pulse is already better than general-purpose coding agents. It provides a measurable baseline, exposes a concrete failure mode, and gives Pulse a data-driven direction for improvement.
213
+
214
+ ---
215
+
216
+ ## Key Features
217
+
218
+ ### Live ML Training Monitoring
219
+
220
+ Track losses, metrics, tensors, gradients, activations, weights, and other numerical values while training is running.
221
+
222
+ Inspect:
223
+
224
+ - shapes
225
+ - dtypes
226
+ - devices
227
+ - norms
228
+ - statistics
229
+ - NaN / Inf counts
230
+ - scalar histories
231
+ - gradient information
232
+ - activation information
233
+
234
+ ### CLI / Headless First
235
+
236
+ Pulse is built for:
237
+
238
+ - terminals
239
+ - SSH sessions
240
+ - Google Colab
241
+ - containers
242
+ - remote servers
243
+ - long-running training jobs
244
+ - cloud GPU machines
245
+
246
+ No graphical interface is required.
247
+
248
+ ### CPU-First Tracking
249
+
250
+ Pulse is intentionally conservative around GPU access.
251
+
252
+ By default:
253
+
254
+ ```text
255
+ TRACKING_MODE = CPU_DEFAULT
256
+ GPU_TRACKING = OPT-IN
257
+ DEFAULT_GPU_OVERHEAD = 0
258
+ ```
259
+
260
+ If a variable already exists on a GPU, Pulse does not automatically copy it back to the CPU on every iteration.
261
+
262
+ GPU tracking is explicitly requested:
263
+
264
+ ```text
265
+ /gputrack <variable>
266
+ ```
267
+
268
+ This can introduce device-to-host transfer overhead. That is expected when inspecting GPU-resident data.
269
+
270
+ The design principle is:
271
+
272
+ > **If nobody asks Pulse to touch the GPU, Pulse doesn't touch the GPU.**
273
+
274
+ ### Dynamic Variable Tracking
275
+
276
+ Variables can be added, removed, promoted, or demoted while training is running.
277
+
278
+ ```text
279
+ /vars
280
+ /tracked
281
+ /add <variable>
282
+ /track <variable>
283
+ /lotrack <variable>
284
+ /gputrack <variable>
285
+ /gpuuntrack <variable>
286
+ /delete <variable>
287
+ ```
288
+
289
+ ### Deterministic Numerical Verification
290
+
291
+ Pulse separates AI reasoning from exact numerical calculation.
292
+
293
+ Instead of asking an AI model to estimate whether a gradient or update is unusually large, Pulse can calculate the relevant quantity directly.
294
+
295
+ ```text
296
+ AI hypothesis
297
+ ↓
298
+ Numerical calculation
299
+ ↓
300
+ Exact result
301
+ ↓
302
+ Evidence-backed diagnosis
303
+ ```
304
+
305
+ This can be used for:
306
+
307
+ - gradient/update ratios
308
+ - scaling factors
309
+ - normalization calculations
310
+ - parameter changes
311
+ - numerical thresholds
312
+ - restricted mathematical expressions
313
+
314
+ The AI reasons about what the numbers mean. Pulse provides deterministic measurements for the arithmetic.
315
+
316
+ ### ML Linting
317
+
318
+ `mllint` provides deterministic, code-level checks for common ML failure patterns before an agent has to reason about them from scratch.
319
+
320
+ Checks include patterns involving:
321
+
322
+ - incompatible loss/metric combinations
323
+ - redundant activation/loss combinations
324
+ - `backward()` without the expected optimizer update
325
+ - training/evaluation mode issues
326
+ - suspicious learning-rate configurations
327
+ - other ML-specific static patterns
328
+
329
+ The goal is not to replace the debugging agent. It is to provide high-confidence evidence and eliminate classes of mistakes that do not require probabilistic reasoning.
330
+
331
+ ### Agentic Debugging
332
+
333
+ Pulse can combine:
334
+
335
+ - live training state
336
+ - tensor statistics
337
+ - scalar histories
338
+ - gradients
339
+ - activations
340
+ - tracebacks
341
+ - source code
342
+ - deterministic numerical checks
343
+ - static lint findings
344
+
345
+ The goal is to move from:
346
+
347
+ ```text
348
+ "What might be wrong?"
349
+ ```
350
+
351
+ toward:
352
+
353
+ ```text
354
+ "Here is the evidence.
355
+ Here is the mutation.
356
+ Here is why it caused the failure.
357
+ Here is the smallest necessary fix."
358
+ ```
359
+
360
+ ### Minimal Fixes by Default
361
+
362
+ Pulse's debugging pipeline is designed to prefer the **smallest correct change**.
363
+
364
+ The agent should first consider:
365
+
366
+ 1. a single-line change;
367
+ 2. a small adjacent-line change;
368
+ 3. a larger edit only when the root cause genuinely requires it.
369
+
370
+ Unrelated cleanup and refactoring should not be bundled into a debugging patch.
371
+
372
+ This is especially important for ML debugging because broad changes can hide whether the actual fault was correctly identified.
373
+
374
+ ### Automatic Intervention
375
+
376
+ With:
377
+
378
+ ```text
379
+ /autofix on
380
+ ```
381
+
382
+ Pulse can pause training when configured detection logic identifies serious numerical or training problems, allowing the debugging agent to investigate before additional compute is wasted.
383
+
384
+ ---
385
+
386
+ ## CLI
387
+
388
+ Useful commands include:
389
+
390
+ ```text
391
+ /help
392
+
393
+ /vars
394
+ /tracked
395
+
396
+ /add <variable>
397
+ /track <variable>
398
+ /lotrack <variable>
399
+
400
+ /gputrack <variable>
401
+ /gpuuntrack <variable>
402
+
403
+ /delete <variable>
404
+
405
+ /autofix on|off
406
+
407
+ /code
408
+ /cloud
409
+ ```
410
+
411
+ Pulse can pause training for investigation, allowing you to inspect the current state, add variables, ask the AI agent questions, and investigate a failure before continuing.
412
+
413
+ ---
414
+
415
+ ## Quickstart
416
+
417
+ After installation, import `auto_track` immediately before your training loop:
418
+
419
+ ```python
420
+ from pulse import auto_track
421
+
422
+ if __name__ == "__main__":
423
+ auto_track()
424
+
425
+ # Your training loop
426
+ for epoch in range(num_epochs):
427
+ # Training logic here
428
+ pass
429
+ ```
430
+
431
+ Pulse discovers numeric variables available to the training process and provides them through the CLI.
432
+
433
+ You can then control tracking while the run is active:
434
+
435
+ ```text
436
+ /vars
437
+ /tracked
438
+ /add <variable>
439
+ /track <variable>
440
+ /lotrack <variable>
441
+ /gputrack <variable>
442
+ /gpuuntrack <variable>
443
+ /delete <variable>
444
+ ```
445
+
446
+ ---
447
+
448
+ ## Supported Backends
449
+
450
+ | Backend | Support |
451
+ |---|---|
452
+ | NumPy | Yes |
453
+ | PyTorch | Yes |
454
+ | TensorFlow | Yes |
455
+ | CuPy | Yes |
456
+ | JAX | Yes |
457
+
458
+ Pulse uses a shared backend abstraction so the debugging workflow can remain consistent across frameworks.
459
+
460
+ ---
461
+
462
+ ## Past Debugs
463
+
464
+ Pulse has been used to investigate numerical failures in custom ML systems.
465
+
466
+ ### Residual normalization failure
467
+
468
+ A custom LLM became unstable after a 2.5× vocabulary increase. Pulse helped identify a normalization issue in which residual growth was divided by `math.sqrt(num_layers)` rather than `num_layers`.
469
+
470
+ The important part was not simply observing that training became unstable. The debugging process connected the observed activation behavior to the mathematical relationship causing the instability.
471
+
472
+ ### Attention numerical failure
473
+
474
+ A custom attention implementation produced NaN loss. Pulse traced the failure to a missing infinity check before a division operation.
475
+
476
+ These cases represent the intended workflow:
477
+
478
+ **observe the behavior → inspect the evidence → verify the math → identify the actual failure**
479
+
480
+ ---
481
+
482
+ ## Performance
483
+
484
+ A debugger should not become the training bottleneck.
485
+
486
+ Pulse is designed around:
487
+
488
+ - CPU-first inspection
489
+ - opt-in GPU probing
490
+ - selective variable tracking
491
+ - lightweight tracking modes
492
+ - separate probe cadences
493
+ - cached statistics
494
+ - host-side numerical processing where possible
495
+ - minimal intervention in the training loop
496
+
497
+ The objective is:
498
+
499
+ ```text
500
+ MORE VISIBILITY
501
+ +
502
+ LESS OVERHEAD
503
+ ```
504
+
505
+ Rather than collecting everything continuously, Pulse lets you decide what information is worth monitoring.
506
+
507
+ ---
508
+
509
+ ## Cloud Workspaces
510
+
511
+ Pulse can optionally synchronize debugging information through shared workspaces.
512
+
513
+ Depending on configuration, a workspace can provide shared access to:
514
+
515
+ - debugging sessions
516
+ - training incidents
517
+ - tracebacks
518
+ - agent conversations
519
+ - repository metadata
520
+ - team membership
521
+ - telemetry
522
+
523
+ Workspaces can be given custom names and managed from the CLI. Members can leave a workspace, while workspace administrators can delete a workspace.
524
+
525
+ Cloud synchronization is best-effort and is not intended to block the training loop.
526
+
527
+ For sensitive projects, review your cloud and telemetry configuration carefully.
528
+
529
+ Telemetry can be disabled with:
530
+
531
+ ```text
532
+ PULSE_TELEMETRY=off
533
+ ```
534
+
535
+ Using a local AI model keeps model requests local, but does not automatically disable separately enabled Pulse workspace synchronization.
536
+
537
+ ---
538
+
539
+ ## AI Providers
540
+
541
+ Pulse can use supported cloud or local AI providers for its debugging agent.
542
+
543
+ Common provider environment variables include:
544
+
545
+ ```text
546
+ ANTHROPIC_API_KEY
547
+ OPENAI_API_KEY
548
+ GEMINI_API_KEY
549
+ DEEPSEEK_API_KEY
550
+ MISTRAL_API_KEY
551
+ OPENROUTER_API_KEY
552
+ ```
553
+
554
+ Local and self-hosted models can also be used where supported.
555
+
556
+ Do not place API keys directly in source code or commit them to a repository.
557
+
558
+ ---
559
+
560
+ ## Why This Matters
561
+
562
+ The central problem Pulse is trying to solve is not code generation.
563
+
564
+ Modern coding agents are increasingly capable of writing large amounts of code. ML debugging has a different requirement: **causal precision**.
565
+
566
+ A useful ML debugger needs to connect three things:
567
+
568
+ ```text
569
+ CODE
570
+ ↕
571
+ RUNTIME BEHAVIOR
572
+ ↕
573
+ MATHEMATICS
574
+ ```
575
+
576
+ If an agent only sees code, it can miss what actually happened during training.
577
+
578
+ If it only sees runtime values, it may not know which source-level mutation produced them.
579
+
580
+ If it only reasons probabilistically, it can make confident but numerically unsupported claims.
581
+
582
+ Pulse is designed to put those pieces together.
583
+
584
+ The benchmark results reinforce why that distinction matters. At the initial 28-case benchmark, Pulse scored 15/28 versus 18/28 for Claude Code with Opus 5. Rather than treating that gap as a dead end, the failures identified a concrete engineering problem: **the system needed to make the agent better at isolating the exact mutation and applying the smallest necessary fix.**
585
+
586
+ That is the direction of Pulse's development.
587
+
588
+ ---
589
+
590
+ ## Project Direction
591
+
592
+ Pulse is being built around a debugging stack in which runtime observation, deterministic computation, static ML analysis, and AI reasoning reinforce each other:
593
+
594
+ ```text
595
+ TRAINING
596
+ ↓
597
+ OBSERVATION
598
+ ↓
599
+ STATIC ANALYSIS
600
+ ↓
601
+ NUMERICAL EVIDENCE
602
+ ↓
603
+ AI REASONING
604
+ ↓
605
+ VERIFICATION
606
+ ↓
607
+ MINIMAL PATCH
608
+ ```
609
+
610
+ The goal is not simply to report:
611
+
612
+ ```text
613
+ "Your training is broken."
614
+ ```
615
+
616
+ It is to answer:
617
+
618
+ ```text
619
+ What broke?
620
+
621
+ Why did it break?
622
+
623
+ When did it start?
624
+
625
+ Can the numbers prove it?
626
+
627
+ What exactly changed?
628
+
629
+ What is the smallest correct fix?
630
+
631
+ Can that fix be verified?
632
+ ```
633
+
634
+ ---
635
+
636
+ ## Links
637
+
638
+ - [Dashboard](https://pulsedashb.netlify.app/)
639
+ - [GitHub](https://github.com/codeyash09/PulseML)
640
+ - [PyPI](https://pypi.org/project/pulseml/)
641
+
642
+ ---
643
+
644
+ ## License
645
+
646
+ Proprietary. See `LICENSE`.
647
+
648
+ Use of this software is governed by the terms in that file. Copying, redistribution, and reverse engineering are not permitted.