computefence 0.2.1__tar.gz → 0.2.3__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.4
2
2
  Name: computefence
3
- Version: 0.2.1
3
+ Version: 0.2.3
4
4
  Summary: Pre-flight validation for GPU training runs on rented infrastructure
5
5
  License-Expression: MIT
6
6
  Requires-Python: >=3.8
@@ -14,7 +14,7 @@ Dynamic: license-file
14
14
 
15
15
  # ComputeFence
16
16
 
17
- Pre-flight safety gate for ML training jobs on rented GPU compute. Built specifically for RunPod, Vast.ai, Lambda Labs, and similar bare metal providers.
17
+ Pre-flight safety gate for ML training jobs on rented GPU compute. Built for RunPod, Vast.ai, Lambda Labs, and similar bare metal providers.
18
18
 
19
19
  ## Install
20
20
 
@@ -22,6 +22,12 @@ Pre-flight safety gate for ML training jobs on rented GPU compute. Built specifi
22
22
  pip install computefence
23
23
  ```
24
24
 
25
+ Or with UV:
26
+
27
+ ```bash
28
+ uvx computefence doctor
29
+ ```
30
+
25
31
  ## Usage
26
32
 
27
33
  ```bash
@@ -36,7 +42,7 @@ computefence doctor --dataset train.csv --input-column text --label-column label
36
42
 
37
43
  ## Example output
38
44
 
39
- ComputeFence v0.1.2 — Pre-flight diagnostic
45
+ ComputeFence v0.2.0 — Pre-flight diagnostic
40
46
  ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
41
47
  2 WARNINGS · 0 BLOCKERS · 3 PASSED
42
48
 
@@ -47,9 +53,9 @@ Environment
47
53
 
48
54
  Storage
49
55
  ⚠ HF_HOME is not set. HuggingFace will use default local cache.
50
- Fix: Set HF_HOME to a persistent volume path before training
51
- ⚠ No mounted persistent volumes detected at /workspace, /runpod-volume, or /vast
52
- Fix: Attach a network volume before training to persist checkpoints and cache
56
+ Fix: export HF_HOME=/workspace/.cache/huggingface
57
+ ⚠ Root disk (/) — 14.3 GB free of 460.4 GB (below 20 GB)
58
+ Fix: Free up disk space or attach a larger volume before launching
53
59
 
54
60
  Dataset
55
61
  ✓ No dataset path provided — skipping dataset checks
@@ -57,11 +63,13 @@ Dataset
57
63
  ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
58
64
  2 warning(s) found. Review before launching.
59
65
 
66
+
60
67
  ## What it checks
61
68
 
62
69
  - **CUDA and GPU visibility** — confirms PyTorch can see the GPU and training will not silently fall back to CPU
63
70
  - **HuggingFace cache path** — confirms model files go to persistent storage not ephemeral disk
64
71
  - **Accelerate GPU count** — confirms your distributed training config matches the GPUs actually on the instance
72
+ - **Disk space headroom** — checks free space on workspace volumes and root disk. Warns below 20 GB, blocks below 5 GB
65
73
  - **Dataset integrity** — optional scan for duplicates, missing values, and conflicting labels
66
74
 
67
75
  ## What it does not check
@@ -71,10 +79,11 @@ Dataset
71
79
  - Learning rate or hyperparameter safety
72
80
  - Runtime monitoring during the job
73
81
  - Slow dataloader or data pipeline throughput
82
+ - Dataloader bottleneck causing low GPU utilisation
74
83
 
75
84
  ## Why this exists
76
85
 
77
- I burned approximately £1,000 on GPU training runs that failed silently. CUDA fell back to CPU with no error. Class weights caused loss collapse to 0.693 immediately. My dataset had 28,432 duplicate rows and 312 conflicting labels I only found after the run.
86
+ I burned approximately £1,000 on GPU training runs that failed silently. CUDA fell back to CPU with no error. 24 seconds per iteration instead of 0.4. Class weights caused loss collapse to 0.693 immediately. My dataset had 28,432 duplicate rows and 312 conflicting labels I only found after the run.
78
87
 
79
88
  Nothing existed that caught these before the job started. So I built it.
80
89
 
@@ -82,6 +91,16 @@ Nothing existed that caught these before the job started. So I built it.
82
91
 
83
92
  David at Neuralic ran ComputeFence on a RunPod A100. It caught HF_HOME writing to /root/.cache and an Accelerate GPU count mismatch. He fixed both before launch.
84
93
 
94
+ Shahzeb Ali, a computer vision engineer running client training jobs on RunPod, confirmed the storage warning matches real pod behaviour and would not have caught the HF_HOME issue explicitly without the tool.
95
+
96
+ ## The problem it solves
97
+
98
+ A healthy GPU does not mean you are training the right job.
99
+
100
+ Docker makes environments reproducible. It does not check whether your HuggingFace cache is writing to ephemeral storage, whether your Accelerate config matches the GPUs actually on the instance, or whether your disk has enough headroom for checkpoints. ComputeFence addresses the job layer not the environment layer.
101
+
102
+ Five independent operators confirmed this independently. Docker does not solve job-specific configuration mistakes.
103
+
85
104
  ## GitHub
86
105
 
87
106
  github.com/Francisco-Booth/ComputeFence
@@ -1,6 +1,6 @@
1
1
  # ComputeFence
2
2
 
3
- Pre-flight safety gate for ML training jobs on rented GPU compute. Built specifically for RunPod, Vast.ai, Lambda Labs, and similar bare metal providers.
3
+ Pre-flight safety gate for ML training jobs on rented GPU compute. Built for RunPod, Vast.ai, Lambda Labs, and similar bare metal providers.
4
4
 
5
5
  ## Install
6
6
 
@@ -8,6 +8,12 @@ Pre-flight safety gate for ML training jobs on rented GPU compute. Built specifi
8
8
  pip install computefence
9
9
  ```
10
10
 
11
+ Or with UV:
12
+
13
+ ```bash
14
+ uvx computefence doctor
15
+ ```
16
+
11
17
  ## Usage
12
18
 
13
19
  ```bash
@@ -22,7 +28,7 @@ computefence doctor --dataset train.csv --input-column text --label-column label
22
28
 
23
29
  ## Example output
24
30
 
25
- ComputeFence v0.1.2 — Pre-flight diagnostic
31
+ ComputeFence v0.2.0 — Pre-flight diagnostic
26
32
  ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
27
33
  2 WARNINGS · 0 BLOCKERS · 3 PASSED
28
34
 
@@ -33,9 +39,9 @@ Environment
33
39
 
34
40
  Storage
35
41
  ⚠ HF_HOME is not set. HuggingFace will use default local cache.
36
- Fix: Set HF_HOME to a persistent volume path before training
37
- ⚠ No mounted persistent volumes detected at /workspace, /runpod-volume, or /vast
38
- Fix: Attach a network volume before training to persist checkpoints and cache
42
+ Fix: export HF_HOME=/workspace/.cache/huggingface
43
+ ⚠ Root disk (/) — 14.3 GB free of 460.4 GB (below 20 GB)
44
+ Fix: Free up disk space or attach a larger volume before launching
39
45
 
40
46
  Dataset
41
47
  ✓ No dataset path provided — skipping dataset checks
@@ -43,11 +49,13 @@ Dataset
43
49
  ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
44
50
  2 warning(s) found. Review before launching.
45
51
 
52
+
46
53
  ## What it checks
47
54
 
48
55
  - **CUDA and GPU visibility** — confirms PyTorch can see the GPU and training will not silently fall back to CPU
49
56
  - **HuggingFace cache path** — confirms model files go to persistent storage not ephemeral disk
50
57
  - **Accelerate GPU count** — confirms your distributed training config matches the GPUs actually on the instance
58
+ - **Disk space headroom** — checks free space on workspace volumes and root disk. Warns below 20 GB, blocks below 5 GB
51
59
  - **Dataset integrity** — optional scan for duplicates, missing values, and conflicting labels
52
60
 
53
61
  ## What it does not check
@@ -57,10 +65,11 @@ Dataset
57
65
  - Learning rate or hyperparameter safety
58
66
  - Runtime monitoring during the job
59
67
  - Slow dataloader or data pipeline throughput
68
+ - Dataloader bottleneck causing low GPU utilisation
60
69
 
61
70
  ## Why this exists
62
71
 
63
- I burned approximately £1,000 on GPU training runs that failed silently. CUDA fell back to CPU with no error. Class weights caused loss collapse to 0.693 immediately. My dataset had 28,432 duplicate rows and 312 conflicting labels I only found after the run.
72
+ I burned approximately £1,000 on GPU training runs that failed silently. CUDA fell back to CPU with no error. 24 seconds per iteration instead of 0.4. Class weights caused loss collapse to 0.693 immediately. My dataset had 28,432 duplicate rows and 312 conflicting labels I only found after the run.
64
73
 
65
74
  Nothing existed that caught these before the job started. So I built it.
66
75
 
@@ -68,6 +77,16 @@ Nothing existed that caught these before the job started. So I built it.
68
77
 
69
78
  David at Neuralic ran ComputeFence on a RunPod A100. It caught HF_HOME writing to /root/.cache and an Accelerate GPU count mismatch. He fixed both before launch.
70
79
 
80
+ Shahzeb Ali, a computer vision engineer running client training jobs on RunPod, confirmed the storage warning matches real pod behaviour and would not have caught the HF_HOME issue explicitly without the tool.
81
+
82
+ ## The problem it solves
83
+
84
+ A healthy GPU does not mean you are training the right job.
85
+
86
+ Docker makes environments reproducible. It does not check whether your HuggingFace cache is writing to ephemeral storage, whether your Accelerate config matches the GPUs actually on the instance, or whether your disk has enough headroom for checkpoints. ComputeFence addresses the job layer not the environment layer.
87
+
88
+ Five independent operators confirmed this independently. Docker does not solve job-specific configuration mistakes.
89
+
71
90
  ## GitHub
72
91
 
73
92
  github.com/Francisco-Booth/ComputeFence
@@ -0,0 +1 @@
1
+ __version__ = "0.2.3"
@@ -56,8 +56,6 @@ def doctor():
56
56
 
57
57
  all_results = env_results + storage_results + dataset_results
58
58
 
59
- record_run(all_results)
60
-
61
59
  failures = [r for r in all_results if r["status"] == "fail"]
62
60
  warnings = [r for r in all_results if r["status"] == "warn"]
63
61
  passed = [r for r in all_results if r["status"] == "pass"]
@@ -89,11 +87,13 @@ def doctor():
89
87
  else:
90
88
  console.print("[green]All checks passed. Safe to launch.[/green]")
91
89
  console.print(
92
- "[dim]Anonymous usage stats help improve ComputeFence. "
93
- "Opt out: touch ~/.computefence_no_telemetry[/dim]"
90
+ "[dim]Anonymous run stats are collected to improve ComputeFence. "
91
+ "To opt out: touch ~/.computefence_no_telemetry[/dim]"
94
92
  )
95
93
  console.print()
96
94
 
95
+ record_run(all_results)
96
+
97
97
  def main():
98
98
  args = sys.argv[1:]
99
99
  if not args or args[0] == "doctor":
@@ -61,4 +61,5 @@ def record_run(results: list):
61
61
  }
62
62
 
63
63
  thread = threading.Thread(target=_send, args=(payload,), daemon=True)
64
- thread.start()
64
+ thread.start()
65
+ thread.join(timeout=4)
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.4
2
2
  Name: computefence
3
- Version: 0.2.1
3
+ Version: 0.2.3
4
4
  Summary: Pre-flight validation for GPU training runs on rented infrastructure
5
5
  License-Expression: MIT
6
6
  Requires-Python: >=3.8
@@ -14,7 +14,7 @@ Dynamic: license-file
14
14
 
15
15
  # ComputeFence
16
16
 
17
- Pre-flight safety gate for ML training jobs on rented GPU compute. Built specifically for RunPod, Vast.ai, Lambda Labs, and similar bare metal providers.
17
+ Pre-flight safety gate for ML training jobs on rented GPU compute. Built for RunPod, Vast.ai, Lambda Labs, and similar bare metal providers.
18
18
 
19
19
  ## Install
20
20
 
@@ -22,6 +22,12 @@ Pre-flight safety gate for ML training jobs on rented GPU compute. Built specifi
22
22
  pip install computefence
23
23
  ```
24
24
 
25
+ Or with UV:
26
+
27
+ ```bash
28
+ uvx computefence doctor
29
+ ```
30
+
25
31
  ## Usage
26
32
 
27
33
  ```bash
@@ -36,7 +42,7 @@ computefence doctor --dataset train.csv --input-column text --label-column label
36
42
 
37
43
  ## Example output
38
44
 
39
- ComputeFence v0.1.2 — Pre-flight diagnostic
45
+ ComputeFence v0.2.0 — Pre-flight diagnostic
40
46
  ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
41
47
  2 WARNINGS · 0 BLOCKERS · 3 PASSED
42
48
 
@@ -47,9 +53,9 @@ Environment
47
53
 
48
54
  Storage
49
55
  ⚠ HF_HOME is not set. HuggingFace will use default local cache.
50
- Fix: Set HF_HOME to a persistent volume path before training
51
- ⚠ No mounted persistent volumes detected at /workspace, /runpod-volume, or /vast
52
- Fix: Attach a network volume before training to persist checkpoints and cache
56
+ Fix: export HF_HOME=/workspace/.cache/huggingface
57
+ ⚠ Root disk (/) — 14.3 GB free of 460.4 GB (below 20 GB)
58
+ Fix: Free up disk space or attach a larger volume before launching
53
59
 
54
60
  Dataset
55
61
  ✓ No dataset path provided — skipping dataset checks
@@ -57,11 +63,13 @@ Dataset
57
63
  ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
58
64
  2 warning(s) found. Review before launching.
59
65
 
66
+
60
67
  ## What it checks
61
68
 
62
69
  - **CUDA and GPU visibility** — confirms PyTorch can see the GPU and training will not silently fall back to CPU
63
70
  - **HuggingFace cache path** — confirms model files go to persistent storage not ephemeral disk
64
71
  - **Accelerate GPU count** — confirms your distributed training config matches the GPUs actually on the instance
72
+ - **Disk space headroom** — checks free space on workspace volumes and root disk. Warns below 20 GB, blocks below 5 GB
65
73
  - **Dataset integrity** — optional scan for duplicates, missing values, and conflicting labels
66
74
 
67
75
  ## What it does not check
@@ -71,10 +79,11 @@ Dataset
71
79
  - Learning rate or hyperparameter safety
72
80
  - Runtime monitoring during the job
73
81
  - Slow dataloader or data pipeline throughput
82
+ - Dataloader bottleneck causing low GPU utilisation
74
83
 
75
84
  ## Why this exists
76
85
 
77
- I burned approximately £1,000 on GPU training runs that failed silently. CUDA fell back to CPU with no error. Class weights caused loss collapse to 0.693 immediately. My dataset had 28,432 duplicate rows and 312 conflicting labels I only found after the run.
86
+ I burned approximately £1,000 on GPU training runs that failed silently. CUDA fell back to CPU with no error. 24 seconds per iteration instead of 0.4. Class weights caused loss collapse to 0.693 immediately. My dataset had 28,432 duplicate rows and 312 conflicting labels I only found after the run.
78
87
 
79
88
  Nothing existed that caught these before the job started. So I built it.
80
89
 
@@ -82,6 +91,16 @@ Nothing existed that caught these before the job started. So I built it.
82
91
 
83
92
  David at Neuralic ran ComputeFence on a RunPod A100. It caught HF_HOME writing to /root/.cache and an Accelerate GPU count mismatch. He fixed both before launch.
84
93
 
94
+ Shahzeb Ali, a computer vision engineer running client training jobs on RunPod, confirmed the storage warning matches real pod behaviour and would not have caught the HF_HOME issue explicitly without the tool.
95
+
96
+ ## The problem it solves
97
+
98
+ A healthy GPU does not mean you are training the right job.
99
+
100
+ Docker makes environments reproducible. It does not check whether your HuggingFace cache is writing to ephemeral storage, whether your Accelerate config matches the GPUs actually on the instance, or whether your disk has enough headroom for checkpoints. ComputeFence addresses the job layer not the environment layer.
101
+
102
+ Five independent operators confirmed this independently. Docker does not solve job-specific configuration mistakes.
103
+
85
104
  ## GitHub
86
105
 
87
106
  github.com/Francisco-Booth/ComputeFence
@@ -1 +0,0 @@
1
- __version__ = "0.2.1"
File without changes
File without changes