computefence 0.2.0__tar.gz → 0.2.3__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,110 @@
1
+ Metadata-Version: 2.4
2
+ Name: computefence
3
+ Version: 0.2.3
4
+ Summary: Pre-flight validation for GPU training runs on rented infrastructure
5
+ License-Expression: MIT
6
+ Requires-Python: >=3.8
7
+ Description-Content-Type: text/markdown
8
+ License-File: LICENSE
9
+ Requires-Dist: rich>=13.0.0
10
+ Requires-Dist: pandas>=1.5.0
11
+ Requires-Dist: typer>=0.9.0
12
+ Requires-Dist: requests>=2.28.0
13
+ Dynamic: license-file
14
+
15
+ # ComputeFence
16
+
17
+ Pre-flight safety gate for ML training jobs on rented GPU compute. Built for RunPod, Vast.ai, Lambda Labs, and similar bare metal providers.
18
+
19
+ ## Install
20
+
21
+ ```bash
22
+ pip install computefence
23
+ ```
24
+
25
+ Or with UV:
26
+
27
+ ```bash
28
+ uvx computefence doctor
29
+ ```
30
+
31
+ ## Usage
32
+
33
+ ```bash
34
+ computefence doctor
35
+ ```
36
+
37
+ With a dataset:
38
+
39
+ ```bash
40
+ computefence doctor --dataset train.csv --input-column text --label-column label
41
+ ```
42
+
43
+ ## Example output
44
+
45
+ ComputeFence v0.2.0 — Pre-flight diagnostic
46
+ ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
47
+ 2 WARNINGS · 0 BLOCKERS · 3 PASSED
48
+
49
+ Environment
50
+ ✓ Python 3.11.4
51
+ ✓ PyTorch 2.1.0 detected
52
+ ✓ CUDA available — NVIDIA A40
53
+
54
+ Storage
55
+ ⚠ HF_HOME is not set. HuggingFace will use default local cache.
56
+ Fix: export HF_HOME=/workspace/.cache/huggingface
57
+ ⚠ Root disk (/) — 14.3 GB free of 460.4 GB (below 20 GB)
58
+ Fix: Free up disk space or attach a larger volume before launching
59
+
60
+ Dataset
61
+ ✓ No dataset path provided — skipping dataset checks
62
+
63
+ ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
64
+ 2 warning(s) found. Review before launching.
65
+
66
+
67
+ ## What it checks
68
+
69
+ - **CUDA and GPU visibility** — confirms PyTorch can see the GPU and training will not silently fall back to CPU
70
+ - **HuggingFace cache path** — confirms model files go to persistent storage not ephemeral disk
71
+ - **Accelerate GPU count** — confirms your distributed training config matches the GPUs actually on the instance
72
+ - **Disk space headroom** — checks free space on workspace volumes and root disk. Warns below 20 GB, blocks below 5 GB
73
+ - **Dataset integrity** — optional scan for duplicates, missing values, and conflicting labels
74
+
75
+ ## What it does not check
76
+
77
+ - Training script correctness
78
+ - Model architecture compatibility
79
+ - Learning rate or hyperparameter safety
80
+ - Runtime monitoring during the job
81
+ - Slow dataloader or data pipeline throughput
82
+ - Dataloader bottleneck causing low GPU utilisation
83
+
84
+ ## Why this exists
85
+
86
+ I burned approximately £1,000 on GPU training runs that failed silently. CUDA fell back to CPU with no error. 24 seconds per iteration instead of 0.4. Class weights caused loss collapse to 0.693 immediately. My dataset had 28,432 duplicate rows and 312 conflicting labels I only found after the run.
87
+
88
+ Nothing existed that caught these before the job started. So I built it.
89
+
90
+ ## Real operator results
91
+
92
+ David at Neuralic ran ComputeFence on a RunPod A100. It caught HF_HOME writing to /root/.cache and an Accelerate GPU count mismatch. He fixed both before launch.
93
+
94
+ Shahzeb Ali, a computer vision engineer running client training jobs on RunPod, confirmed the storage warning matches real pod behaviour and would not have caught the HF_HOME issue explicitly without the tool.
95
+
96
+ ## The problem it solves
97
+
98
+ A healthy GPU does not mean you are training the right job.
99
+
100
+ Docker makes environments reproducible. It does not check whether your HuggingFace cache is writing to ephemeral storage, whether your Accelerate config matches the GPUs actually on the instance, or whether your disk has enough headroom for checkpoints. ComputeFence addresses the job layer not the environment layer.
101
+
102
+ Five independent operators confirmed this independently. Docker does not solve job-specific configuration mistakes.
103
+
104
+ ## GitHub
105
+
106
+ github.com/Francisco-Booth/ComputeFence
107
+
108
+ ## PyPI
109
+
110
+ pypi.org/project/computefence
@@ -0,0 +1,96 @@
1
+ # ComputeFence
2
+
3
+ Pre-flight safety gate for ML training jobs on rented GPU compute. Built for RunPod, Vast.ai, Lambda Labs, and similar bare metal providers.
4
+
5
+ ## Install
6
+
7
+ ```bash
8
+ pip install computefence
9
+ ```
10
+
11
+ Or with UV:
12
+
13
+ ```bash
14
+ uvx computefence doctor
15
+ ```
16
+
17
+ ## Usage
18
+
19
+ ```bash
20
+ computefence doctor
21
+ ```
22
+
23
+ With a dataset:
24
+
25
+ ```bash
26
+ computefence doctor --dataset train.csv --input-column text --label-column label
27
+ ```
28
+
29
+ ## Example output
30
+
31
+ ComputeFence v0.2.0 — Pre-flight diagnostic
32
+ ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
33
+ 2 WARNINGS · 0 BLOCKERS · 3 PASSED
34
+
35
+ Environment
36
+ ✓ Python 3.11.4
37
+ ✓ PyTorch 2.1.0 detected
38
+ ✓ CUDA available — NVIDIA A40
39
+
40
+ Storage
41
+ ⚠ HF_HOME is not set. HuggingFace will use default local cache.
42
+ Fix: export HF_HOME=/workspace/.cache/huggingface
43
+ ⚠ Root disk (/) — 14.3 GB free of 460.4 GB (below 20 GB)
44
+ Fix: Free up disk space or attach a larger volume before launching
45
+
46
+ Dataset
47
+ ✓ No dataset path provided — skipping dataset checks
48
+
49
+ ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
50
+ 2 warning(s) found. Review before launching.
51
+
52
+
53
+ ## What it checks
54
+
55
+ - **CUDA and GPU visibility** — confirms PyTorch can see the GPU and training will not silently fall back to CPU
56
+ - **HuggingFace cache path** — confirms model files go to persistent storage not ephemeral disk
57
+ - **Accelerate GPU count** — confirms your distributed training config matches the GPUs actually on the instance
58
+ - **Disk space headroom** — checks free space on workspace volumes and root disk. Warns below 20 GB, blocks below 5 GB
59
+ - **Dataset integrity** — optional scan for duplicates, missing values, and conflicting labels
60
+
61
+ ## What it does not check
62
+
63
+ - Training script correctness
64
+ - Model architecture compatibility
65
+ - Learning rate or hyperparameter safety
66
+ - Runtime monitoring during the job
67
+ - Slow dataloader or data pipeline throughput
68
+ - Dataloader bottleneck causing low GPU utilisation
69
+
70
+ ## Why this exists
71
+
72
+ I burned approximately £1,000 on GPU training runs that failed silently. CUDA fell back to CPU with no error. 24 seconds per iteration instead of 0.4. Class weights caused loss collapse to 0.693 immediately. My dataset had 28,432 duplicate rows and 312 conflicting labels I only found after the run.
73
+
74
+ Nothing existed that caught these before the job started. So I built it.
75
+
76
+ ## Real operator results
77
+
78
+ David at Neuralic ran ComputeFence on a RunPod A100. It caught HF_HOME writing to /root/.cache and an Accelerate GPU count mismatch. He fixed both before launch.
79
+
80
+ Shahzeb Ali, a computer vision engineer running client training jobs on RunPod, confirmed the storage warning matches real pod behaviour and would not have caught the HF_HOME issue explicitly without the tool.
81
+
82
+ ## The problem it solves
83
+
84
+ A healthy GPU does not mean you are training the right job.
85
+
86
+ Docker makes environments reproducible. It does not check whether your HuggingFace cache is writing to ephemeral storage, whether your Accelerate config matches the GPUs actually on the instance, or whether your disk has enough headroom for checkpoints. ComputeFence addresses the job layer not the environment layer.
87
+
88
+ Five independent operators confirmed this independently. Docker does not solve job-specific configuration mistakes.
89
+
90
+ ## GitHub
91
+
92
+ github.com/Francisco-Booth/ComputeFence
93
+
94
+ ## PyPI
95
+
96
+ pypi.org/project/computefence
@@ -0,0 +1 @@
1
+ __version__ = "0.2.3"
@@ -1,19 +1,17 @@
1
1
  import sys
2
2
  from rich.console import Console
3
-
4
3
  from computefence import __version__
5
4
  from computefence.checks.environment import check_environment
6
5
  from computefence.checks.storage import check_disk_headroom, check_storage
7
6
  from computefence.checks.dataset import check_dataset
7
+ from computefence.telemetry import record_run
8
8
 
9
9
  console = Console()
10
10
 
11
-
12
11
  def print_result(result):
13
12
  status = result.get("status")
14
13
  message = result.get("message")
15
14
  fix = result.get("fix")
16
-
17
15
  if status == "pass":
18
16
  console.print(f" [green]✓[/green] {message}")
19
17
  elif status == "warn":
@@ -25,12 +23,10 @@ def print_result(result):
25
23
  if fix:
26
24
  console.print(f" [dim]Fix: {fix}[/dim]")
27
25
 
28
-
29
26
  def doctor():
30
27
  dataset = None
31
28
  input_column = None
32
29
  label_column = None
33
-
34
30
  args = sys.argv[1:]
35
31
  i = 0
36
32
  while i < len(args):
@@ -50,8 +46,6 @@ def doctor():
50
46
  console.print(f"[bold]ComputeFence v{__version__} — Pre-flight diagnostic[/bold]")
51
47
  console.print("━" * 50)
52
48
 
53
- # Run every check up front so the verdict line can be counted and printed
54
- # above the individual results.
55
49
  env_results = check_environment()
56
50
  storage_results = check_storage() + check_disk_headroom()
57
51
  dataset_results = check_dataset(
@@ -61,6 +55,7 @@ def doctor():
61
55
  )
62
56
 
63
57
  all_results = env_results + storage_results + dataset_results
58
+
64
59
  failures = [r for r in all_results if r["status"] == "fail"]
65
60
  warnings = [r for r in all_results if r["status"] == "warn"]
66
61
  passed = [r for r in all_results if r["status"] == "pass"]
@@ -71,34 +66,33 @@ def doctor():
71
66
  f"[red]{len(failures)} BLOCKERS[/red] · "
72
67
  f"[green]{len(passed)} PASSED[/green]"
73
68
  )
74
-
75
69
  console.print()
76
70
  console.print("[bold blue]Environment[/bold blue]")
77
71
  for result in env_results:
78
72
  print_result(result)
79
-
80
73
  console.print()
81
74
  console.print("[bold blue]Storage[/bold blue]")
82
75
  for result in storage_results:
83
76
  print_result(result)
84
-
85
77
  console.print()
86
78
  console.print("[bold blue]Dataset[/bold blue]")
87
79
  for result in dataset_results:
88
80
  print_result(result)
89
-
90
81
  console.print()
91
82
  console.print("━" * 50)
92
-
93
83
  if failures:
94
84
  console.print("[red]" + str(len(failures)) + " error(s) and " + str(len(warnings)) + " warning(s) found. Fix errors before launching.[/red]")
95
85
  elif warnings:
96
86
  console.print("[yellow]" + str(len(warnings)) + " warning(s) found. Review before launching.[/yellow]")
97
87
  else:
98
88
  console.print("[green]All checks passed. Safe to launch.[/green]")
99
-
89
+ console.print(
90
+ "[dim]Anonymous run stats are collected to improve ComputeFence. "
91
+ "To opt out: touch ~/.computefence_no_telemetry[/dim]"
92
+ )
100
93
  console.print()
101
94
 
95
+ record_run(all_results)
102
96
 
103
97
  def main():
104
98
  args = sys.argv[1:]
@@ -108,6 +102,5 @@ def main():
108
102
  console.print(f"[red]Unknown command: {args[0]}[/red]")
109
103
  console.print("Usage: computefence doctor")
110
104
 
111
-
112
105
  if __name__ == "__main__":
113
106
  main()
@@ -0,0 +1,65 @@
1
+ import hashlib
2
+ import platform
3
+ import sys
4
+ import threading
5
+ import uuid
6
+ from pathlib import Path
7
+
8
+ SUPABASE_URL = "https://pidpadpudbcdldyvrlog.supabase.co"
9
+ SUPABASE_ANON_KEY = "sb_publishable_A2zwbfTFiOs63jz8313GXg_elw3F5dy"
10
+ TABLE = "computefence_runs"
11
+ OPT_OUT_FILE = Path.home() / ".computefence_no_telemetry"
12
+
13
+
14
+ def _is_opted_out():
15
+ return OPT_OUT_FILE.exists()
16
+
17
+
18
+ def _send(payload: dict):
19
+ try:
20
+ import requests
21
+ requests.post(
22
+ f"{SUPABASE_URL}/rest/v1/{TABLE}",
23
+ headers={
24
+ "apikey": SUPABASE_ANON_KEY,
25
+ "Authorization": f"Bearer {SUPABASE_ANON_KEY}",
26
+ "Content-Type": "application/json",
27
+ "Prefer": "return=minimal",
28
+ },
29
+ json=payload,
30
+ timeout=3,
31
+ )
32
+ except Exception:
33
+ pass
34
+
35
+
36
+ def record_run(results: list):
37
+ if _is_opted_out():
38
+ return
39
+
40
+ warning_count = sum(1 for r in results if r.get("status") == "warn")
41
+ blocker_count = sum(1 for r in results if r.get("status") == "fail")
42
+ pass_count = sum(1 for r in results if r.get("status") == "pass")
43
+
44
+ try:
45
+ import torch
46
+ gpu_count = torch.cuda.device_count()
47
+ except Exception:
48
+ gpu_count = 0
49
+
50
+ from computefence import __version__
51
+
52
+ payload = {
53
+ "version": __version__,
54
+ "python_version": f"{sys.version_info.major}.{sys.version_info.minor}",
55
+ "platform": platform.system(),
56
+ "gpu_count": gpu_count,
57
+ "warning_count": warning_count,
58
+ "blocker_count": blocker_count,
59
+ "pass_count": pass_count,
60
+ "opted_in": True,
61
+ }
62
+
63
+ thread = threading.Thread(target=_send, args=(payload,), daemon=True)
64
+ thread.start()
65
+ thread.join(timeout=4)
@@ -0,0 +1,110 @@
1
+ Metadata-Version: 2.4
2
+ Name: computefence
3
+ Version: 0.2.3
4
+ Summary: Pre-flight validation for GPU training runs on rented infrastructure
5
+ License-Expression: MIT
6
+ Requires-Python: >=3.8
7
+ Description-Content-Type: text/markdown
8
+ License-File: LICENSE
9
+ Requires-Dist: rich>=13.0.0
10
+ Requires-Dist: pandas>=1.5.0
11
+ Requires-Dist: typer>=0.9.0
12
+ Requires-Dist: requests>=2.28.0
13
+ Dynamic: license-file
14
+
15
+ # ComputeFence
16
+
17
+ Pre-flight safety gate for ML training jobs on rented GPU compute. Built for RunPod, Vast.ai, Lambda Labs, and similar bare metal providers.
18
+
19
+ ## Install
20
+
21
+ ```bash
22
+ pip install computefence
23
+ ```
24
+
25
+ Or with UV:
26
+
27
+ ```bash
28
+ uvx computefence doctor
29
+ ```
30
+
31
+ ## Usage
32
+
33
+ ```bash
34
+ computefence doctor
35
+ ```
36
+
37
+ With a dataset:
38
+
39
+ ```bash
40
+ computefence doctor --dataset train.csv --input-column text --label-column label
41
+ ```
42
+
43
+ ## Example output
44
+
45
+ ComputeFence v0.2.0 — Pre-flight diagnostic
46
+ ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
47
+ 2 WARNINGS · 0 BLOCKERS · 3 PASSED
48
+
49
+ Environment
50
+ ✓ Python 3.11.4
51
+ ✓ PyTorch 2.1.0 detected
52
+ ✓ CUDA available — NVIDIA A40
53
+
54
+ Storage
55
+ ⚠ HF_HOME is not set. HuggingFace will use default local cache.
56
+ Fix: export HF_HOME=/workspace/.cache/huggingface
57
+ ⚠ Root disk (/) — 14.3 GB free of 460.4 GB (below 20 GB)
58
+ Fix: Free up disk space or attach a larger volume before launching
59
+
60
+ Dataset
61
+ ✓ No dataset path provided — skipping dataset checks
62
+
63
+ ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
64
+ 2 warning(s) found. Review before launching.
65
+
66
+
67
+ ## What it checks
68
+
69
+ - **CUDA and GPU visibility** — confirms PyTorch can see the GPU and training will not silently fall back to CPU
70
+ - **HuggingFace cache path** — confirms model files go to persistent storage not ephemeral disk
71
+ - **Accelerate GPU count** — confirms your distributed training config matches the GPUs actually on the instance
72
+ - **Disk space headroom** — checks free space on workspace volumes and root disk. Warns below 20 GB, blocks below 5 GB
73
+ - **Dataset integrity** — optional scan for duplicates, missing values, and conflicting labels
74
+
75
+ ## What it does not check
76
+
77
+ - Training script correctness
78
+ - Model architecture compatibility
79
+ - Learning rate or hyperparameter safety
80
+ - Runtime monitoring during the job
81
+ - Slow dataloader or data pipeline throughput
82
+ - Dataloader bottleneck causing low GPU utilisation
83
+
84
+ ## Why this exists
85
+
86
+ I burned approximately £1,000 on GPU training runs that failed silently. CUDA fell back to CPU with no error. 24 seconds per iteration instead of 0.4. Class weights caused loss collapse to 0.693 immediately. My dataset had 28,432 duplicate rows and 312 conflicting labels I only found after the run.
87
+
88
+ Nothing existed that caught these before the job started. So I built it.
89
+
90
+ ## Real operator results
91
+
92
+ David at Neuralic ran ComputeFence on a RunPod A100. It caught HF_HOME writing to /root/.cache and an Accelerate GPU count mismatch. He fixed both before launch.
93
+
94
+ Shahzeb Ali, a computer vision engineer running client training jobs on RunPod, confirmed the storage warning matches real pod behaviour and would not have caught the HF_HOME issue explicitly without the tool.
95
+
96
+ ## The problem it solves
97
+
98
+ A healthy GPU does not mean you are training the right job.
99
+
100
+ Docker makes environments reproducible. It does not check whether your HuggingFace cache is writing to ephemeral storage, whether your Accelerate config matches the GPUs actually on the instance, or whether your disk has enough headroom for checkpoints. ComputeFence addresses the job layer not the environment layer.
101
+
102
+ Five independent operators confirmed this independently. Docker does not solve job-specific configuration mistakes.
103
+
104
+ ## GitHub
105
+
106
+ github.com/Francisco-Booth/ComputeFence
107
+
108
+ ## PyPI
109
+
110
+ pypi.org/project/computefence
@@ -3,6 +3,7 @@ README.md
3
3
  pyproject.toml
4
4
  computefence/__init__.py
5
5
  computefence/cli.py
6
+ computefence/telemetry.py
6
7
  computefence.egg-info/PKG-INFO
7
8
  computefence.egg-info/SOURCES.txt
8
9
  computefence.egg-info/dependency_links.txt
@@ -1,49 +0,0 @@
1
- Metadata-Version: 2.4
2
- Name: computefence
3
- Version: 0.2.0
4
- Summary: Pre-flight validation for GPU training runs on rented infrastructure
5
- License-Expression: MIT
6
- Requires-Python: >=3.8
7
- Description-Content-Type: text/markdown
8
- License-File: LICENSE
9
- Requires-Dist: rich>=13.0.0
10
- Requires-Dist: pandas>=1.5.0
11
- Requires-Dist: typer>=0.9.0
12
- Requires-Dist: requests>=2.28.0
13
- Dynamic: license-file
14
-
15
- # ComputeFence
16
-
17
- Pre-flight validation for GPU training runs on rented infrastructure.
18
- Built specifically for RunPod, Vast.ai, Lambda, and similar providers.
19
-
20
- ## Install
21
-
22
- pip install computefence
23
-
24
- ## Usage
25
-
26
- computefence doctor
27
- computefence doctor --dataset train.csv
28
-
29
- ## What it checks
30
-
31
- - CUDA and GPU availability
32
- - HuggingFace cache volume path (catches the RunPod /root vs /workspace conflict)
33
- - Accelerate GPU count vs config
34
- - Dataset duplicates and missing values
35
-
36
- ## What it does not yet check
37
-
38
- - Training script correctness
39
- - Model architecture compatibility
40
- - Learning rate or hyperparameter safety
41
- - Runtime monitoring during the job
42
-
43
- ## Why this exists
44
-
45
- I burned ~£1,000 on GPU training runs that failed silently. CUDA fell back to
46
- CPU with no error. Class weights caused loss collapse. My dataset had 28,432
47
- duplicate rows and 312 conflicting labels I only found during the rebuild.
48
-
49
- Nothing existed that caught these before the job started. So I built it.
@@ -1,35 +0,0 @@
1
- # ComputeFence
2
-
3
- Pre-flight validation for GPU training runs on rented infrastructure.
4
- Built specifically for RunPod, Vast.ai, Lambda, and similar providers.
5
-
6
- ## Install
7
-
8
- pip install computefence
9
-
10
- ## Usage
11
-
12
- computefence doctor
13
- computefence doctor --dataset train.csv
14
-
15
- ## What it checks
16
-
17
- - CUDA and GPU availability
18
- - HuggingFace cache volume path (catches the RunPod /root vs /workspace conflict)
19
- - Accelerate GPU count vs config
20
- - Dataset duplicates and missing values
21
-
22
- ## What it does not yet check
23
-
24
- - Training script correctness
25
- - Model architecture compatibility
26
- - Learning rate or hyperparameter safety
27
- - Runtime monitoring during the job
28
-
29
- ## Why this exists
30
-
31
- I burned ~£1,000 on GPU training runs that failed silently. CUDA fell back to
32
- CPU with no error. Class weights caused loss collapse. My dataset had 28,432
33
- duplicate rows and 312 conflicting labels I only found during the rebuild.
34
-
35
- Nothing existed that caught these before the job started. So I built it.
@@ -1 +0,0 @@
1
- __version__ = "0.2.0"
@@ -1,49 +0,0 @@
1
- Metadata-Version: 2.4
2
- Name: computefence
3
- Version: 0.2.0
4
- Summary: Pre-flight validation for GPU training runs on rented infrastructure
5
- License-Expression: MIT
6
- Requires-Python: >=3.8
7
- Description-Content-Type: text/markdown
8
- License-File: LICENSE
9
- Requires-Dist: rich>=13.0.0
10
- Requires-Dist: pandas>=1.5.0
11
- Requires-Dist: typer>=0.9.0
12
- Requires-Dist: requests>=2.28.0
13
- Dynamic: license-file
14
-
15
- # ComputeFence
16
-
17
- Pre-flight validation for GPU training runs on rented infrastructure.
18
- Built specifically for RunPod, Vast.ai, Lambda, and similar providers.
19
-
20
- ## Install
21
-
22
- pip install computefence
23
-
24
- ## Usage
25
-
26
- computefence doctor
27
- computefence doctor --dataset train.csv
28
-
29
- ## What it checks
30
-
31
- - CUDA and GPU availability
32
- - HuggingFace cache volume path (catches the RunPod /root vs /workspace conflict)
33
- - Accelerate GPU count vs config
34
- - Dataset duplicates and missing values
35
-
36
- ## What it does not yet check
37
-
38
- - Training script correctness
39
- - Model architecture compatibility
40
- - Learning rate or hyperparameter safety
41
- - Runtime monitoring during the job
42
-
43
- ## Why this exists
44
-
45
- I burned ~£1,000 on GPU training runs that failed silently. CUDA fell back to
46
- CPU with no error. Class weights caused loss collapse. My dataset had 28,432
47
- duplicate rows and 312 conflicting labels I only found during the rebuild.
48
-
49
- Nothing existed that caught these before the job started. So I built it.
File without changes
File without changes