computefence 0.2.4__tar.gz → 0.2.5__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- computefence-0.2.5/PKG-INFO +161 -0
- computefence-0.2.5/README.md +147 -0
- computefence-0.2.5/computefence/__init__.py +1 -0
- {computefence-0.2.4 → computefence-0.2.5}/computefence/checks/storage.py +48 -0
- {computefence-0.2.4 → computefence-0.2.5}/computefence/cli.py +6 -2
- computefence-0.2.5/computefence.egg-info/PKG-INFO +161 -0
- computefence-0.2.4/PKG-INFO +0 -110
- computefence-0.2.4/README.md +0 -96
- computefence-0.2.4/computefence/__init__.py +0 -1
- computefence-0.2.4/computefence.egg-info/PKG-INFO +0 -110
- {computefence-0.2.4 → computefence-0.2.5}/LICENSE +0 -0
- {computefence-0.2.4 → computefence-0.2.5}/computefence/checks/__init__.py +0 -0
- {computefence-0.2.4 → computefence-0.2.5}/computefence/checks/dataset.py +0 -0
- {computefence-0.2.4 → computefence-0.2.5}/computefence/checks/environment.py +0 -0
- {computefence-0.2.4 → computefence-0.2.5}/computefence/telemetry.py +0 -0
- {computefence-0.2.4 → computefence-0.2.5}/computefence.egg-info/SOURCES.txt +0 -0
- {computefence-0.2.4 → computefence-0.2.5}/computefence.egg-info/dependency_links.txt +0 -0
- {computefence-0.2.4 → computefence-0.2.5}/computefence.egg-info/entry_points.txt +0 -0
- {computefence-0.2.4 → computefence-0.2.5}/computefence.egg-info/requires.txt +0 -0
- {computefence-0.2.4 → computefence-0.2.5}/computefence.egg-info/top_level.txt +0 -0
- {computefence-0.2.4 → computefence-0.2.5}/pyproject.toml +0 -0
- {computefence-0.2.4 → computefence-0.2.5}/setup.cfg +0 -0
|
@@ -0,0 +1,161 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: computefence
|
|
3
|
+
Version: 0.2.5
|
|
4
|
+
Summary: Pre-flight validation for GPU training runs on rented infrastructure
|
|
5
|
+
License-Expression: MIT
|
|
6
|
+
Requires-Python: >=3.8
|
|
7
|
+
Description-Content-Type: text/markdown
|
|
8
|
+
License-File: LICENSE
|
|
9
|
+
Requires-Dist: rich>=13.0.0
|
|
10
|
+
Requires-Dist: pandas>=1.5.0
|
|
11
|
+
Requires-Dist: typer>=0.9.0
|
|
12
|
+
Requires-Dist: requests>=2.28.0
|
|
13
|
+
Dynamic: license-file
|
|
14
|
+
|
|
15
|
+
# ComputeFence
|
|
16
|
+
|
|
17
|
+
Pre-flight safety gate for ML training jobs on rented GPU compute. Built for RunPod, Vast.ai, Lambda Labs, CoreWeave, Paperspace, and any bare metal GPU provider.
|
|
18
|
+
|
|
19
|
+
---
|
|
20
|
+
|
|
21
|
+
## Install
|
|
22
|
+
|
|
23
|
+
```bash
|
|
24
|
+
pip install computefence
|
|
25
|
+
```
|
|
26
|
+
|
|
27
|
+
Or with UV (no install required):
|
|
28
|
+
|
|
29
|
+
```bash
|
|
30
|
+
uvx computefence doctor
|
|
31
|
+
```
|
|
32
|
+
|
|
33
|
+
---
|
|
34
|
+
|
|
35
|
+
## Usage
|
|
36
|
+
|
|
37
|
+
```bash
|
|
38
|
+
computefence doctor
|
|
39
|
+
```
|
|
40
|
+
|
|
41
|
+
With a dataset:
|
|
42
|
+
|
|
43
|
+
```bash
|
|
44
|
+
computefence doctor --dataset train.csv --input-column text --label-column label
|
|
45
|
+
```
|
|
46
|
+
|
|
47
|
+
With checkpoint output directory validation:
|
|
48
|
+
|
|
49
|
+
```bash
|
|
50
|
+
computefence doctor --output-dir ./checkpoints
|
|
51
|
+
```
|
|
52
|
+
|
|
53
|
+
Add to your pod startup script so it runs automatically before every job:
|
|
54
|
+
|
|
55
|
+
```bash
|
|
56
|
+
pip install computefence && computefence doctor && python train.py
|
|
57
|
+
```
|
|
58
|
+
|
|
59
|
+
---
|
|
60
|
+
|
|
61
|
+
## Example output
|
|
62
|
+
|
|
63
|
+
```
|
|
64
|
+
ComputeFence v0.2.4 — Pre-flight diagnostic
|
|
65
|
+
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
|
|
66
|
+
2 WARNINGS · 0 BLOCKERS · 3 PASSED
|
|
67
|
+
|
|
68
|
+
Environment
|
|
69
|
+
✓ Python 3.11.4
|
|
70
|
+
✓ PyTorch 2.1.0 detected
|
|
71
|
+
✓ CUDA available — NVIDIA A40
|
|
72
|
+
|
|
73
|
+
Storage
|
|
74
|
+
⚠ HF_HOME is not set. HuggingFace will use default local cache.
|
|
75
|
+
Fix: export HF_HOME=/workspace/.cache/huggingface
|
|
76
|
+
⚠ Root disk (/) — 14.3 GB free of 460.4 GB (below 20 GB)
|
|
77
|
+
Fix: Free up disk space or move checkpoints to a larger volume: df -h to check usage
|
|
78
|
+
|
|
79
|
+
Dataset
|
|
80
|
+
✓ No dataset path provided — skipping dataset checks
|
|
81
|
+
|
|
82
|
+
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
|
|
83
|
+
2 warning(s) found. Review before launching.
|
|
84
|
+
Anonymous run stats are collected to improve ComputeFence. To opt out: touch ~/.computefence_no_telemetry
|
|
85
|
+
```
|
|
86
|
+
|
|
87
|
+
---
|
|
88
|
+
|
|
89
|
+
## What it checks
|
|
90
|
+
|
|
91
|
+
- **CUDA and GPU visibility** — confirms PyTorch can see the GPU and training will not silently fall back to CPU
|
|
92
|
+
- **HuggingFace cache path** — confirms model weights go to persistent storage not ephemeral disk that disappears on pod stop
|
|
93
|
+
- **Accelerate GPU count** — confirms your distributed training config matches the GPUs actually on the instance
|
|
94
|
+
- **Disk space headroom** — checks free space on workspace volumes and root disk. Warns below 20 GB, blocks below 5 GB
|
|
95
|
+
- **Checkpoint output directory** — confirms your training script's output path is on persistent storage not ephemeral disk (`--output-dir`)
|
|
96
|
+
- **Dataset integrity** — optional scan for duplicates, missing values, and conflicting labels (`--dataset`)
|
|
97
|
+
|
|
98
|
+
---
|
|
99
|
+
|
|
100
|
+
## What it does not check
|
|
101
|
+
|
|
102
|
+
- Training script correctness
|
|
103
|
+
- Model architecture compatibility
|
|
104
|
+
- Learning rate or hyperparameter safety
|
|
105
|
+
- Runtime monitoring during the job
|
|
106
|
+
- Dataloader throughput or GPU utilisation during training
|
|
107
|
+
|
|
108
|
+
---
|
|
109
|
+
|
|
110
|
+
## Why this exists
|
|
111
|
+
|
|
112
|
+
I burned approximately £1,000 on GPU training runs that failed silently.
|
|
113
|
+
|
|
114
|
+
CUDA fell back to CPU with no error. 24 seconds per iteration instead of 0.4. Class weights caused loss collapse to 0.693 immediately. My dataset had 28,432 duplicate rows and 312 conflicting labels I only found after the run.
|
|
115
|
+
|
|
116
|
+
Nothing existed that caught these before the job started. So I built it.
|
|
117
|
+
|
|
118
|
+
---
|
|
119
|
+
|
|
120
|
+
## Real operator results
|
|
121
|
+
|
|
122
|
+
**David at Neuralic** ran ComputeFence on a RunPod A100. It caught HF_HOME writing to `/root/.cache` and an Accelerate GPU count mismatch. He fixed both before launch.
|
|
123
|
+
|
|
124
|
+
**Shahzeb Ali**, a computer vision engineer running client training jobs on RunPod, confirmed the storage warning matches real pod behaviour and would not have caught the HF_HOME issue explicitly without the tool.
|
|
125
|
+
|
|
126
|
+
**Seven overnight organic runs** appeared in telemetry from a Vast.ai operator running the tool six times in two minutes before a real training job — without being prompted or paid to do so.
|
|
127
|
+
|
|
128
|
+
---
|
|
129
|
+
|
|
130
|
+
## The problem it solves
|
|
131
|
+
|
|
132
|
+
A healthy GPU does not mean you are training the right job.
|
|
133
|
+
|
|
134
|
+
Docker makes environments reproducible. It does not check whether your HuggingFace cache is writing to ephemeral storage that disappears on pod stop, whether your Accelerate config matches the GPUs actually on the instance, or whether your training script's checkpoint output directory is on persistent storage. ComputeFence addresses the job configuration layer — not the environment layer.
|
|
135
|
+
|
|
136
|
+
HF_HOME and your checkpoint output directory are separate paths. Fixing one does not fix the other. Both disappear on pod restart if they point to ephemeral disk.
|
|
137
|
+
|
|
138
|
+
Thirteen independent ML engineers confirmed this problem independently across RunPod, Vast.ai, and AWS. Five confirmed that Docker does not solve job-specific configuration mistakes.
|
|
139
|
+
|
|
140
|
+
---
|
|
141
|
+
|
|
142
|
+
## Founding Design Partner pilot
|
|
143
|
+
|
|
144
|
+
Running high-cost GPU training on RunPod or Vast.ai?
|
|
145
|
+
|
|
146
|
+
We offer a hands-on **Founding Design Partner** pilot at **$99 for 3 months**:
|
|
147
|
+
|
|
148
|
+
- Personal audit of your launch templates and persistent storage configuration
|
|
149
|
+
- ComputeFence installed into your pod startup scripts so pre-flight runs automatically on every job
|
|
150
|
+
- Slack or Discord webhook alert when a check fails or blocks a launch
|
|
151
|
+
- Monthly 30-minute call where you shape what gets built next
|
|
152
|
+
- Money back if it does not catch one bad launch in 3 months
|
|
153
|
+
|
|
154
|
+
Three spots available. Email **francisco@booth.ws** to apply.
|
|
155
|
+
|
|
156
|
+
---
|
|
157
|
+
|
|
158
|
+
## Links
|
|
159
|
+
|
|
160
|
+
- **GitHub**: [github.com/Francisco-Booth/ComputeFence](https://github.com/Francisco-Booth/ComputeFence)
|
|
161
|
+
- **PyPI**: [pypi.org/project/computefence](https://pypi.org/project/computefence)
|
|
@@ -0,0 +1,147 @@
|
|
|
1
|
+
# ComputeFence
|
|
2
|
+
|
|
3
|
+
Pre-flight safety gate for ML training jobs on rented GPU compute. Built for RunPod, Vast.ai, Lambda Labs, CoreWeave, Paperspace, and any bare metal GPU provider.
|
|
4
|
+
|
|
5
|
+
---
|
|
6
|
+
|
|
7
|
+
## Install
|
|
8
|
+
|
|
9
|
+
```bash
|
|
10
|
+
pip install computefence
|
|
11
|
+
```
|
|
12
|
+
|
|
13
|
+
Or with UV (no install required):
|
|
14
|
+
|
|
15
|
+
```bash
|
|
16
|
+
uvx computefence doctor
|
|
17
|
+
```
|
|
18
|
+
|
|
19
|
+
---
|
|
20
|
+
|
|
21
|
+
## Usage
|
|
22
|
+
|
|
23
|
+
```bash
|
|
24
|
+
computefence doctor
|
|
25
|
+
```
|
|
26
|
+
|
|
27
|
+
With a dataset:
|
|
28
|
+
|
|
29
|
+
```bash
|
|
30
|
+
computefence doctor --dataset train.csv --input-column text --label-column label
|
|
31
|
+
```
|
|
32
|
+
|
|
33
|
+
With checkpoint output directory validation:
|
|
34
|
+
|
|
35
|
+
```bash
|
|
36
|
+
computefence doctor --output-dir ./checkpoints
|
|
37
|
+
```
|
|
38
|
+
|
|
39
|
+
Add to your pod startup script so it runs automatically before every job:
|
|
40
|
+
|
|
41
|
+
```bash
|
|
42
|
+
pip install computefence && computefence doctor && python train.py
|
|
43
|
+
```
|
|
44
|
+
|
|
45
|
+
---
|
|
46
|
+
|
|
47
|
+
## Example output
|
|
48
|
+
|
|
49
|
+
```
|
|
50
|
+
ComputeFence v0.2.4 — Pre-flight diagnostic
|
|
51
|
+
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
|
|
52
|
+
2 WARNINGS · 0 BLOCKERS · 3 PASSED
|
|
53
|
+
|
|
54
|
+
Environment
|
|
55
|
+
✓ Python 3.11.4
|
|
56
|
+
✓ PyTorch 2.1.0 detected
|
|
57
|
+
✓ CUDA available — NVIDIA A40
|
|
58
|
+
|
|
59
|
+
Storage
|
|
60
|
+
⚠ HF_HOME is not set. HuggingFace will use default local cache.
|
|
61
|
+
Fix: export HF_HOME=/workspace/.cache/huggingface
|
|
62
|
+
⚠ Root disk (/) — 14.3 GB free of 460.4 GB (below 20 GB)
|
|
63
|
+
Fix: Free up disk space or move checkpoints to a larger volume: df -h to check usage
|
|
64
|
+
|
|
65
|
+
Dataset
|
|
66
|
+
✓ No dataset path provided — skipping dataset checks
|
|
67
|
+
|
|
68
|
+
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
|
|
69
|
+
2 warning(s) found. Review before launching.
|
|
70
|
+
Anonymous run stats are collected to improve ComputeFence. To opt out: touch ~/.computefence_no_telemetry
|
|
71
|
+
```
|
|
72
|
+
|
|
73
|
+
---
|
|
74
|
+
|
|
75
|
+
## What it checks
|
|
76
|
+
|
|
77
|
+
- **CUDA and GPU visibility** — confirms PyTorch can see the GPU and training will not silently fall back to CPU
|
|
78
|
+
- **HuggingFace cache path** — confirms model weights go to persistent storage not ephemeral disk that disappears on pod stop
|
|
79
|
+
- **Accelerate GPU count** — confirms your distributed training config matches the GPUs actually on the instance
|
|
80
|
+
- **Disk space headroom** — checks free space on workspace volumes and root disk. Warns below 20 GB, blocks below 5 GB
|
|
81
|
+
- **Checkpoint output directory** — confirms your training script's output path is on persistent storage not ephemeral disk (`--output-dir`)
|
|
82
|
+
- **Dataset integrity** — optional scan for duplicates, missing values, and conflicting labels (`--dataset`)
|
|
83
|
+
|
|
84
|
+
---
|
|
85
|
+
|
|
86
|
+
## What it does not check
|
|
87
|
+
|
|
88
|
+
- Training script correctness
|
|
89
|
+
- Model architecture compatibility
|
|
90
|
+
- Learning rate or hyperparameter safety
|
|
91
|
+
- Runtime monitoring during the job
|
|
92
|
+
- Dataloader throughput or GPU utilisation during training
|
|
93
|
+
|
|
94
|
+
---
|
|
95
|
+
|
|
96
|
+
## Why this exists
|
|
97
|
+
|
|
98
|
+
I burned approximately £1,000 on GPU training runs that failed silently.
|
|
99
|
+
|
|
100
|
+
CUDA fell back to CPU with no error. 24 seconds per iteration instead of 0.4. Class weights caused loss collapse to 0.693 immediately. My dataset had 28,432 duplicate rows and 312 conflicting labels I only found after the run.
|
|
101
|
+
|
|
102
|
+
Nothing existed that caught these before the job started. So I built it.
|
|
103
|
+
|
|
104
|
+
---
|
|
105
|
+
|
|
106
|
+
## Real operator results
|
|
107
|
+
|
|
108
|
+
**David at Neuralic** ran ComputeFence on a RunPod A100. It caught HF_HOME writing to `/root/.cache` and an Accelerate GPU count mismatch. He fixed both before launch.
|
|
109
|
+
|
|
110
|
+
**Shahzeb Ali**, a computer vision engineer running client training jobs on RunPod, confirmed the storage warning matches real pod behaviour and would not have caught the HF_HOME issue explicitly without the tool.
|
|
111
|
+
|
|
112
|
+
**Seven overnight organic runs** appeared in telemetry from a Vast.ai operator running the tool six times in two minutes before a real training job — without being prompted or paid to do so.
|
|
113
|
+
|
|
114
|
+
---
|
|
115
|
+
|
|
116
|
+
## The problem it solves
|
|
117
|
+
|
|
118
|
+
A healthy GPU does not mean you are training the right job.
|
|
119
|
+
|
|
120
|
+
Docker makes environments reproducible. It does not check whether your HuggingFace cache is writing to ephemeral storage that disappears on pod stop, whether your Accelerate config matches the GPUs actually on the instance, or whether your training script's checkpoint output directory is on persistent storage. ComputeFence addresses the job configuration layer — not the environment layer.
|
|
121
|
+
|
|
122
|
+
HF_HOME and your checkpoint output directory are separate paths. Fixing one does not fix the other. Both disappear on pod restart if they point to ephemeral disk.
|
|
123
|
+
|
|
124
|
+
Thirteen independent ML engineers confirmed this problem independently across RunPod, Vast.ai, and AWS. Five confirmed that Docker does not solve job-specific configuration mistakes.
|
|
125
|
+
|
|
126
|
+
---
|
|
127
|
+
|
|
128
|
+
## Founding Design Partner pilot
|
|
129
|
+
|
|
130
|
+
Running high-cost GPU training on RunPod or Vast.ai?
|
|
131
|
+
|
|
132
|
+
We offer a hands-on **Founding Design Partner** pilot at **$99 for 3 months**:
|
|
133
|
+
|
|
134
|
+
- Personal audit of your launch templates and persistent storage configuration
|
|
135
|
+
- ComputeFence installed into your pod startup scripts so pre-flight runs automatically on every job
|
|
136
|
+
- Slack or Discord webhook alert when a check fails or blocks a launch
|
|
137
|
+
- Monthly 30-minute call where you shape what gets built next
|
|
138
|
+
- Money back if it does not catch one bad launch in 3 months
|
|
139
|
+
|
|
140
|
+
Three spots available. Email **francisco@booth.ws** to apply.
|
|
141
|
+
|
|
142
|
+
---
|
|
143
|
+
|
|
144
|
+
## Links
|
|
145
|
+
|
|
146
|
+
- **GitHub**: [github.com/Francisco-Booth/ComputeFence](https://github.com/Francisco-Booth/ComputeFence)
|
|
147
|
+
- **PyPI**: [pypi.org/project/computefence](https://pypi.org/project/computefence)
|
|
@@ -0,0 +1 @@
|
|
|
1
|
+
__version__ = "0.2.5"
|
|
@@ -134,3 +134,51 @@ def check_disk_headroom():
|
|
|
134
134
|
})
|
|
135
135
|
|
|
136
136
|
return results
|
|
137
|
+
|
|
138
|
+
|
|
139
|
+
# Paths that are wiped when a rented instance stops.
|
|
140
|
+
EPHEMERAL_PREFIXES = ["/root", "/tmp"]
|
|
141
|
+
# Paths that survive an instance restart.
|
|
142
|
+
PERSISTENT_PREFIXES = ["/workspace", "/runpod-volume", "/vast", "/home"]
|
|
143
|
+
OUTPUT_DIR_FIX = "Move your output_dir to /workspace/checkpoints or your mounted volume"
|
|
144
|
+
|
|
145
|
+
|
|
146
|
+
def _is_under(path, prefix):
|
|
147
|
+
"""True when path sits at or below prefix.
|
|
148
|
+
|
|
149
|
+
Compares path components rather than string prefixes so /workspace-old is
|
|
150
|
+
not mistaken for a child of /workspace. Avoids Path.is_relative_to, which
|
|
151
|
+
needs Python 3.9 while this package supports 3.8.
|
|
152
|
+
"""
|
|
153
|
+
path_parts = Path(path).parts
|
|
154
|
+
prefix_parts = Path(prefix).parts
|
|
155
|
+
return path_parts[: len(prefix_parts)] == prefix_parts
|
|
156
|
+
|
|
157
|
+
|
|
158
|
+
def check_output_dir(output_dir=None):
|
|
159
|
+
"""Flag a checkpoint directory that will not survive the instance stopping.
|
|
160
|
+
|
|
161
|
+
Uses abspath rather than resolve() so a path is judged as written: on macOS
|
|
162
|
+
/tmp is a symlink to /private/tmp, and resolving it first would hide the
|
|
163
|
+
very ephemeral prefix this check exists to catch.
|
|
164
|
+
"""
|
|
165
|
+
if output_dir is None:
|
|
166
|
+
return []
|
|
167
|
+
|
|
168
|
+
path = Path(os.path.abspath(os.path.expanduser(str(output_dir))))
|
|
169
|
+
|
|
170
|
+
is_ephemeral = any(_is_under(path, prefix) for prefix in EPHEMERAL_PREFIXES) or not any(
|
|
171
|
+
_is_under(path, prefix) for prefix in PERSISTENT_PREFIXES
|
|
172
|
+
)
|
|
173
|
+
|
|
174
|
+
if is_ephemeral:
|
|
175
|
+
return [{
|
|
176
|
+
"status": "fail",
|
|
177
|
+
"message": f"Output directory {path} is on ephemeral storage and will be lost when the instance stops",
|
|
178
|
+
"fix": OUTPUT_DIR_FIX
|
|
179
|
+
}]
|
|
180
|
+
|
|
181
|
+
return [{
|
|
182
|
+
"status": "pass",
|
|
183
|
+
"message": f"Output directory {path} is on persistent storage"
|
|
184
|
+
}]
|
|
@@ -2,7 +2,7 @@ import sys
|
|
|
2
2
|
from rich.console import Console
|
|
3
3
|
from computefence import __version__
|
|
4
4
|
from computefence.checks.environment import check_environment
|
|
5
|
-
from computefence.checks.storage import check_disk_headroom, check_storage
|
|
5
|
+
from computefence.checks.storage import check_disk_headroom, check_output_dir, check_storage
|
|
6
6
|
from computefence.checks.dataset import check_dataset
|
|
7
7
|
from computefence.telemetry import record_run
|
|
8
8
|
|
|
@@ -27,6 +27,7 @@ def doctor():
|
|
|
27
27
|
dataset = None
|
|
28
28
|
input_column = None
|
|
29
29
|
label_column = None
|
|
30
|
+
output_dir = None
|
|
30
31
|
args = sys.argv[1:]
|
|
31
32
|
i = 0
|
|
32
33
|
while i < len(args):
|
|
@@ -39,6 +40,9 @@ def doctor():
|
|
|
39
40
|
elif args[i] == "--label-column" and i + 1 < len(args):
|
|
40
41
|
label_column = args[i + 1]
|
|
41
42
|
i += 2
|
|
43
|
+
elif args[i] == "--output-dir" and i + 1 < len(args):
|
|
44
|
+
output_dir = args[i + 1]
|
|
45
|
+
i += 2
|
|
42
46
|
else:
|
|
43
47
|
i += 1
|
|
44
48
|
|
|
@@ -47,7 +51,7 @@ def doctor():
|
|
|
47
51
|
console.print("━" * 50)
|
|
48
52
|
|
|
49
53
|
env_results = check_environment()
|
|
50
|
-
storage_results = check_storage() + check_disk_headroom()
|
|
54
|
+
storage_results = check_storage() + check_disk_headroom() + check_output_dir(output_dir)
|
|
51
55
|
dataset_results = check_dataset(
|
|
52
56
|
dataset_path=dataset,
|
|
53
57
|
input_column=input_column,
|
|
@@ -0,0 +1,161 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: computefence
|
|
3
|
+
Version: 0.2.5
|
|
4
|
+
Summary: Pre-flight validation for GPU training runs on rented infrastructure
|
|
5
|
+
License-Expression: MIT
|
|
6
|
+
Requires-Python: >=3.8
|
|
7
|
+
Description-Content-Type: text/markdown
|
|
8
|
+
License-File: LICENSE
|
|
9
|
+
Requires-Dist: rich>=13.0.0
|
|
10
|
+
Requires-Dist: pandas>=1.5.0
|
|
11
|
+
Requires-Dist: typer>=0.9.0
|
|
12
|
+
Requires-Dist: requests>=2.28.0
|
|
13
|
+
Dynamic: license-file
|
|
14
|
+
|
|
15
|
+
# ComputeFence
|
|
16
|
+
|
|
17
|
+
Pre-flight safety gate for ML training jobs on rented GPU compute. Built for RunPod, Vast.ai, Lambda Labs, CoreWeave, Paperspace, and any bare metal GPU provider.
|
|
18
|
+
|
|
19
|
+
---
|
|
20
|
+
|
|
21
|
+
## Install
|
|
22
|
+
|
|
23
|
+
```bash
|
|
24
|
+
pip install computefence
|
|
25
|
+
```
|
|
26
|
+
|
|
27
|
+
Or with UV (no install required):
|
|
28
|
+
|
|
29
|
+
```bash
|
|
30
|
+
uvx computefence doctor
|
|
31
|
+
```
|
|
32
|
+
|
|
33
|
+
---
|
|
34
|
+
|
|
35
|
+
## Usage
|
|
36
|
+
|
|
37
|
+
```bash
|
|
38
|
+
computefence doctor
|
|
39
|
+
```
|
|
40
|
+
|
|
41
|
+
With a dataset:
|
|
42
|
+
|
|
43
|
+
```bash
|
|
44
|
+
computefence doctor --dataset train.csv --input-column text --label-column label
|
|
45
|
+
```
|
|
46
|
+
|
|
47
|
+
With checkpoint output directory validation:
|
|
48
|
+
|
|
49
|
+
```bash
|
|
50
|
+
computefence doctor --output-dir ./checkpoints
|
|
51
|
+
```
|
|
52
|
+
|
|
53
|
+
Add to your pod startup script so it runs automatically before every job:
|
|
54
|
+
|
|
55
|
+
```bash
|
|
56
|
+
pip install computefence && computefence doctor && python train.py
|
|
57
|
+
```
|
|
58
|
+
|
|
59
|
+
---
|
|
60
|
+
|
|
61
|
+
## Example output
|
|
62
|
+
|
|
63
|
+
```
|
|
64
|
+
ComputeFence v0.2.4 — Pre-flight diagnostic
|
|
65
|
+
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
|
|
66
|
+
2 WARNINGS · 0 BLOCKERS · 3 PASSED
|
|
67
|
+
|
|
68
|
+
Environment
|
|
69
|
+
✓ Python 3.11.4
|
|
70
|
+
✓ PyTorch 2.1.0 detected
|
|
71
|
+
✓ CUDA available — NVIDIA A40
|
|
72
|
+
|
|
73
|
+
Storage
|
|
74
|
+
⚠ HF_HOME is not set. HuggingFace will use default local cache.
|
|
75
|
+
Fix: export HF_HOME=/workspace/.cache/huggingface
|
|
76
|
+
⚠ Root disk (/) — 14.3 GB free of 460.4 GB (below 20 GB)
|
|
77
|
+
Fix: Free up disk space or move checkpoints to a larger volume: df -h to check usage
|
|
78
|
+
|
|
79
|
+
Dataset
|
|
80
|
+
✓ No dataset path provided — skipping dataset checks
|
|
81
|
+
|
|
82
|
+
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
|
|
83
|
+
2 warning(s) found. Review before launching.
|
|
84
|
+
Anonymous run stats are collected to improve ComputeFence. To opt out: touch ~/.computefence_no_telemetry
|
|
85
|
+
```
|
|
86
|
+
|
|
87
|
+
---
|
|
88
|
+
|
|
89
|
+
## What it checks
|
|
90
|
+
|
|
91
|
+
- **CUDA and GPU visibility** — confirms PyTorch can see the GPU and training will not silently fall back to CPU
|
|
92
|
+
- **HuggingFace cache path** — confirms model weights go to persistent storage not ephemeral disk that disappears on pod stop
|
|
93
|
+
- **Accelerate GPU count** — confirms your distributed training config matches the GPUs actually on the instance
|
|
94
|
+
- **Disk space headroom** — checks free space on workspace volumes and root disk. Warns below 20 GB, blocks below 5 GB
|
|
95
|
+
- **Checkpoint output directory** — confirms your training script's output path is on persistent storage not ephemeral disk (`--output-dir`)
|
|
96
|
+
- **Dataset integrity** — optional scan for duplicates, missing values, and conflicting labels (`--dataset`)
|
|
97
|
+
|
|
98
|
+
---
|
|
99
|
+
|
|
100
|
+
## What it does not check
|
|
101
|
+
|
|
102
|
+
- Training script correctness
|
|
103
|
+
- Model architecture compatibility
|
|
104
|
+
- Learning rate or hyperparameter safety
|
|
105
|
+
- Runtime monitoring during the job
|
|
106
|
+
- Dataloader throughput or GPU utilisation during training
|
|
107
|
+
|
|
108
|
+
---
|
|
109
|
+
|
|
110
|
+
## Why this exists
|
|
111
|
+
|
|
112
|
+
I burned approximately £1,000 on GPU training runs that failed silently.
|
|
113
|
+
|
|
114
|
+
CUDA fell back to CPU with no error. 24 seconds per iteration instead of 0.4. Class weights caused loss collapse to 0.693 immediately. My dataset had 28,432 duplicate rows and 312 conflicting labels I only found after the run.
|
|
115
|
+
|
|
116
|
+
Nothing existed that caught these before the job started. So I built it.
|
|
117
|
+
|
|
118
|
+
---
|
|
119
|
+
|
|
120
|
+
## Real operator results
|
|
121
|
+
|
|
122
|
+
**David at Neuralic** ran ComputeFence on a RunPod A100. It caught HF_HOME writing to `/root/.cache` and an Accelerate GPU count mismatch. He fixed both before launch.
|
|
123
|
+
|
|
124
|
+
**Shahzeb Ali**, a computer vision engineer running client training jobs on RunPod, confirmed the storage warning matches real pod behaviour and would not have caught the HF_HOME issue explicitly without the tool.
|
|
125
|
+
|
|
126
|
+
**Seven overnight organic runs** appeared in telemetry from a Vast.ai operator running the tool six times in two minutes before a real training job — without being prompted or paid to do so.
|
|
127
|
+
|
|
128
|
+
---
|
|
129
|
+
|
|
130
|
+
## The problem it solves
|
|
131
|
+
|
|
132
|
+
A healthy GPU does not mean you are training the right job.
|
|
133
|
+
|
|
134
|
+
Docker makes environments reproducible. It does not check whether your HuggingFace cache is writing to ephemeral storage that disappears on pod stop, whether your Accelerate config matches the GPUs actually on the instance, or whether your training script's checkpoint output directory is on persistent storage. ComputeFence addresses the job configuration layer — not the environment layer.
|
|
135
|
+
|
|
136
|
+
HF_HOME and your checkpoint output directory are separate paths. Fixing one does not fix the other. Both disappear on pod restart if they point to ephemeral disk.
|
|
137
|
+
|
|
138
|
+
Thirteen independent ML engineers confirmed this problem independently across RunPod, Vast.ai, and AWS. Five confirmed that Docker does not solve job-specific configuration mistakes.
|
|
139
|
+
|
|
140
|
+
---
|
|
141
|
+
|
|
142
|
+
## Founding Design Partner pilot
|
|
143
|
+
|
|
144
|
+
Running high-cost GPU training on RunPod or Vast.ai?
|
|
145
|
+
|
|
146
|
+
We offer a hands-on **Founding Design Partner** pilot at **$99 for 3 months**:
|
|
147
|
+
|
|
148
|
+
- Personal audit of your launch templates and persistent storage configuration
|
|
149
|
+
- ComputeFence installed into your pod startup scripts so pre-flight runs automatically on every job
|
|
150
|
+
- Slack or Discord webhook alert when a check fails or blocks a launch
|
|
151
|
+
- Monthly 30-minute call where you shape what gets built next
|
|
152
|
+
- Money back if it does not catch one bad launch in 3 months
|
|
153
|
+
|
|
154
|
+
Three spots available. Email **francisco@booth.ws** to apply.
|
|
155
|
+
|
|
156
|
+
---
|
|
157
|
+
|
|
158
|
+
## Links
|
|
159
|
+
|
|
160
|
+
- **GitHub**: [github.com/Francisco-Booth/ComputeFence](https://github.com/Francisco-Booth/ComputeFence)
|
|
161
|
+
- **PyPI**: [pypi.org/project/computefence](https://pypi.org/project/computefence)
|
computefence-0.2.4/PKG-INFO
DELETED
|
@@ -1,110 +0,0 @@
|
|
|
1
|
-
Metadata-Version: 2.4
|
|
2
|
-
Name: computefence
|
|
3
|
-
Version: 0.2.4
|
|
4
|
-
Summary: Pre-flight validation for GPU training runs on rented infrastructure
|
|
5
|
-
License-Expression: MIT
|
|
6
|
-
Requires-Python: >=3.8
|
|
7
|
-
Description-Content-Type: text/markdown
|
|
8
|
-
License-File: LICENSE
|
|
9
|
-
Requires-Dist: rich>=13.0.0
|
|
10
|
-
Requires-Dist: pandas>=1.5.0
|
|
11
|
-
Requires-Dist: typer>=0.9.0
|
|
12
|
-
Requires-Dist: requests>=2.28.0
|
|
13
|
-
Dynamic: license-file
|
|
14
|
-
|
|
15
|
-
# ComputeFence
|
|
16
|
-
|
|
17
|
-
Pre-flight safety gate for ML training jobs on rented GPU compute. Built for RunPod, Vast.ai, Lambda Labs, and similar bare metal providers.
|
|
18
|
-
|
|
19
|
-
## Install
|
|
20
|
-
|
|
21
|
-
```bash
|
|
22
|
-
pip install computefence
|
|
23
|
-
```
|
|
24
|
-
|
|
25
|
-
Or with UV:
|
|
26
|
-
|
|
27
|
-
```bash
|
|
28
|
-
uvx computefence doctor
|
|
29
|
-
```
|
|
30
|
-
|
|
31
|
-
## Usage
|
|
32
|
-
|
|
33
|
-
```bash
|
|
34
|
-
computefence doctor
|
|
35
|
-
```
|
|
36
|
-
|
|
37
|
-
With a dataset:
|
|
38
|
-
|
|
39
|
-
```bash
|
|
40
|
-
computefence doctor --dataset train.csv --input-column text --label-column label
|
|
41
|
-
```
|
|
42
|
-
|
|
43
|
-
## Example output
|
|
44
|
-
|
|
45
|
-
ComputeFence v0.2.0 — Pre-flight diagnostic
|
|
46
|
-
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
|
|
47
|
-
2 WARNINGS · 0 BLOCKERS · 3 PASSED
|
|
48
|
-
|
|
49
|
-
Environment
|
|
50
|
-
✓ Python 3.11.4
|
|
51
|
-
✓ PyTorch 2.1.0 detected
|
|
52
|
-
✓ CUDA available — NVIDIA A40
|
|
53
|
-
|
|
54
|
-
Storage
|
|
55
|
-
⚠ HF_HOME is not set. HuggingFace will use default local cache.
|
|
56
|
-
Fix: export HF_HOME=/workspace/.cache/huggingface
|
|
57
|
-
⚠ Root disk (/) — 14.3 GB free of 460.4 GB (below 20 GB)
|
|
58
|
-
Fix: Free up disk space or attach a larger volume before launching
|
|
59
|
-
|
|
60
|
-
Dataset
|
|
61
|
-
✓ No dataset path provided — skipping dataset checks
|
|
62
|
-
|
|
63
|
-
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
|
|
64
|
-
2 warning(s) found. Review before launching.
|
|
65
|
-
|
|
66
|
-
|
|
67
|
-
## What it checks
|
|
68
|
-
|
|
69
|
-
- **CUDA and GPU visibility** — confirms PyTorch can see the GPU and training will not silently fall back to CPU
|
|
70
|
-
- **HuggingFace cache path** — confirms model files go to persistent storage not ephemeral disk
|
|
71
|
-
- **Accelerate GPU count** — confirms your distributed training config matches the GPUs actually on the instance
|
|
72
|
-
- **Disk space headroom** — checks free space on workspace volumes and root disk. Warns below 20 GB, blocks below 5 GB
|
|
73
|
-
- **Dataset integrity** — optional scan for duplicates, missing values, and conflicting labels
|
|
74
|
-
|
|
75
|
-
## What it does not check
|
|
76
|
-
|
|
77
|
-
- Training script correctness
|
|
78
|
-
- Model architecture compatibility
|
|
79
|
-
- Learning rate or hyperparameter safety
|
|
80
|
-
- Runtime monitoring during the job
|
|
81
|
-
- Slow dataloader or data pipeline throughput
|
|
82
|
-
- Dataloader bottleneck causing low GPU utilisation
|
|
83
|
-
|
|
84
|
-
## Why this exists
|
|
85
|
-
|
|
86
|
-
I burned approximately £1,000 on GPU training runs that failed silently. CUDA fell back to CPU with no error. 24 seconds per iteration instead of 0.4. Class weights caused loss collapse to 0.693 immediately. My dataset had 28,432 duplicate rows and 312 conflicting labels I only found after the run.
|
|
87
|
-
|
|
88
|
-
Nothing existed that caught these before the job started. So I built it.
|
|
89
|
-
|
|
90
|
-
## Real operator results
|
|
91
|
-
|
|
92
|
-
David at Neuralic ran ComputeFence on a RunPod A100. It caught HF_HOME writing to /root/.cache and an Accelerate GPU count mismatch. He fixed both before launch.
|
|
93
|
-
|
|
94
|
-
Shahzeb Ali, a computer vision engineer running client training jobs on RunPod, confirmed the storage warning matches real pod behaviour and would not have caught the HF_HOME issue explicitly without the tool.
|
|
95
|
-
|
|
96
|
-
## The problem it solves
|
|
97
|
-
|
|
98
|
-
A healthy GPU does not mean you are training the right job.
|
|
99
|
-
|
|
100
|
-
Docker makes environments reproducible. It does not check whether your HuggingFace cache is writing to ephemeral storage, whether your Accelerate config matches the GPUs actually on the instance, or whether your disk has enough headroom for checkpoints. ComputeFence addresses the job layer not the environment layer.
|
|
101
|
-
|
|
102
|
-
Five independent operators confirmed this independently. Docker does not solve job-specific configuration mistakes.
|
|
103
|
-
|
|
104
|
-
## GitHub
|
|
105
|
-
|
|
106
|
-
github.com/Francisco-Booth/ComputeFence
|
|
107
|
-
|
|
108
|
-
## PyPI
|
|
109
|
-
|
|
110
|
-
pypi.org/project/computefence
|
computefence-0.2.4/README.md
DELETED
|
@@ -1,96 +0,0 @@
|
|
|
1
|
-
# ComputeFence
|
|
2
|
-
|
|
3
|
-
Pre-flight safety gate for ML training jobs on rented GPU compute. Built for RunPod, Vast.ai, Lambda Labs, and similar bare metal providers.
|
|
4
|
-
|
|
5
|
-
## Install
|
|
6
|
-
|
|
7
|
-
```bash
|
|
8
|
-
pip install computefence
|
|
9
|
-
```
|
|
10
|
-
|
|
11
|
-
Or with UV:
|
|
12
|
-
|
|
13
|
-
```bash
|
|
14
|
-
uvx computefence doctor
|
|
15
|
-
```
|
|
16
|
-
|
|
17
|
-
## Usage
|
|
18
|
-
|
|
19
|
-
```bash
|
|
20
|
-
computefence doctor
|
|
21
|
-
```
|
|
22
|
-
|
|
23
|
-
With a dataset:
|
|
24
|
-
|
|
25
|
-
```bash
|
|
26
|
-
computefence doctor --dataset train.csv --input-column text --label-column label
|
|
27
|
-
```
|
|
28
|
-
|
|
29
|
-
## Example output
|
|
30
|
-
|
|
31
|
-
ComputeFence v0.2.0 — Pre-flight diagnostic
|
|
32
|
-
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
|
|
33
|
-
2 WARNINGS · 0 BLOCKERS · 3 PASSED
|
|
34
|
-
|
|
35
|
-
Environment
|
|
36
|
-
✓ Python 3.11.4
|
|
37
|
-
✓ PyTorch 2.1.0 detected
|
|
38
|
-
✓ CUDA available — NVIDIA A40
|
|
39
|
-
|
|
40
|
-
Storage
|
|
41
|
-
⚠ HF_HOME is not set. HuggingFace will use default local cache.
|
|
42
|
-
Fix: export HF_HOME=/workspace/.cache/huggingface
|
|
43
|
-
⚠ Root disk (/) — 14.3 GB free of 460.4 GB (below 20 GB)
|
|
44
|
-
Fix: Free up disk space or attach a larger volume before launching
|
|
45
|
-
|
|
46
|
-
Dataset
|
|
47
|
-
✓ No dataset path provided — skipping dataset checks
|
|
48
|
-
|
|
49
|
-
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
|
|
50
|
-
2 warning(s) found. Review before launching.
|
|
51
|
-
|
|
52
|
-
|
|
53
|
-
## What it checks
|
|
54
|
-
|
|
55
|
-
- **CUDA and GPU visibility** — confirms PyTorch can see the GPU and training will not silently fall back to CPU
|
|
56
|
-
- **HuggingFace cache path** — confirms model files go to persistent storage not ephemeral disk
|
|
57
|
-
- **Accelerate GPU count** — confirms your distributed training config matches the GPUs actually on the instance
|
|
58
|
-
- **Disk space headroom** — checks free space on workspace volumes and root disk. Warns below 20 GB, blocks below 5 GB
|
|
59
|
-
- **Dataset integrity** — optional scan for duplicates, missing values, and conflicting labels
|
|
60
|
-
|
|
61
|
-
## What it does not check
|
|
62
|
-
|
|
63
|
-
- Training script correctness
|
|
64
|
-
- Model architecture compatibility
|
|
65
|
-
- Learning rate or hyperparameter safety
|
|
66
|
-
- Runtime monitoring during the job
|
|
67
|
-
- Slow dataloader or data pipeline throughput
|
|
68
|
-
- Dataloader bottleneck causing low GPU utilisation
|
|
69
|
-
|
|
70
|
-
## Why this exists
|
|
71
|
-
|
|
72
|
-
I burned approximately £1,000 on GPU training runs that failed silently. CUDA fell back to CPU with no error. 24 seconds per iteration instead of 0.4. Class weights caused loss collapse to 0.693 immediately. My dataset had 28,432 duplicate rows and 312 conflicting labels I only found after the run.
|
|
73
|
-
|
|
74
|
-
Nothing existed that caught these before the job started. So I built it.
|
|
75
|
-
|
|
76
|
-
## Real operator results
|
|
77
|
-
|
|
78
|
-
David at Neuralic ran ComputeFence on a RunPod A100. It caught HF_HOME writing to /root/.cache and an Accelerate GPU count mismatch. He fixed both before launch.
|
|
79
|
-
|
|
80
|
-
Shahzeb Ali, a computer vision engineer running client training jobs on RunPod, confirmed the storage warning matches real pod behaviour and would not have caught the HF_HOME issue explicitly without the tool.
|
|
81
|
-
|
|
82
|
-
## The problem it solves
|
|
83
|
-
|
|
84
|
-
A healthy GPU does not mean you are training the right job.
|
|
85
|
-
|
|
86
|
-
Docker makes environments reproducible. It does not check whether your HuggingFace cache is writing to ephemeral storage, whether your Accelerate config matches the GPUs actually on the instance, or whether your disk has enough headroom for checkpoints. ComputeFence addresses the job layer not the environment layer.
|
|
87
|
-
|
|
88
|
-
Five independent operators confirmed this independently. Docker does not solve job-specific configuration mistakes.
|
|
89
|
-
|
|
90
|
-
## GitHub
|
|
91
|
-
|
|
92
|
-
github.com/Francisco-Booth/ComputeFence
|
|
93
|
-
|
|
94
|
-
## PyPI
|
|
95
|
-
|
|
96
|
-
pypi.org/project/computefence
|
|
@@ -1 +0,0 @@
|
|
|
1
|
-
__version__ = "0.2.4"
|
|
@@ -1,110 +0,0 @@
|
|
|
1
|
-
Metadata-Version: 2.4
|
|
2
|
-
Name: computefence
|
|
3
|
-
Version: 0.2.4
|
|
4
|
-
Summary: Pre-flight validation for GPU training runs on rented infrastructure
|
|
5
|
-
License-Expression: MIT
|
|
6
|
-
Requires-Python: >=3.8
|
|
7
|
-
Description-Content-Type: text/markdown
|
|
8
|
-
License-File: LICENSE
|
|
9
|
-
Requires-Dist: rich>=13.0.0
|
|
10
|
-
Requires-Dist: pandas>=1.5.0
|
|
11
|
-
Requires-Dist: typer>=0.9.0
|
|
12
|
-
Requires-Dist: requests>=2.28.0
|
|
13
|
-
Dynamic: license-file
|
|
14
|
-
|
|
15
|
-
# ComputeFence
|
|
16
|
-
|
|
17
|
-
Pre-flight safety gate for ML training jobs on rented GPU compute. Built for RunPod, Vast.ai, Lambda Labs, and similar bare metal providers.
|
|
18
|
-
|
|
19
|
-
## Install
|
|
20
|
-
|
|
21
|
-
```bash
|
|
22
|
-
pip install computefence
|
|
23
|
-
```
|
|
24
|
-
|
|
25
|
-
Or with UV:
|
|
26
|
-
|
|
27
|
-
```bash
|
|
28
|
-
uvx computefence doctor
|
|
29
|
-
```
|
|
30
|
-
|
|
31
|
-
## Usage
|
|
32
|
-
|
|
33
|
-
```bash
|
|
34
|
-
computefence doctor
|
|
35
|
-
```
|
|
36
|
-
|
|
37
|
-
With a dataset:
|
|
38
|
-
|
|
39
|
-
```bash
|
|
40
|
-
computefence doctor --dataset train.csv --input-column text --label-column label
|
|
41
|
-
```
|
|
42
|
-
|
|
43
|
-
## Example output
|
|
44
|
-
|
|
45
|
-
ComputeFence v0.2.0 — Pre-flight diagnostic
|
|
46
|
-
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
|
|
47
|
-
2 WARNINGS · 0 BLOCKERS · 3 PASSED
|
|
48
|
-
|
|
49
|
-
Environment
|
|
50
|
-
✓ Python 3.11.4
|
|
51
|
-
✓ PyTorch 2.1.0 detected
|
|
52
|
-
✓ CUDA available — NVIDIA A40
|
|
53
|
-
|
|
54
|
-
Storage
|
|
55
|
-
⚠ HF_HOME is not set. HuggingFace will use default local cache.
|
|
56
|
-
Fix: export HF_HOME=/workspace/.cache/huggingface
|
|
57
|
-
⚠ Root disk (/) — 14.3 GB free of 460.4 GB (below 20 GB)
|
|
58
|
-
Fix: Free up disk space or attach a larger volume before launching
|
|
59
|
-
|
|
60
|
-
Dataset
|
|
61
|
-
✓ No dataset path provided — skipping dataset checks
|
|
62
|
-
|
|
63
|
-
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
|
|
64
|
-
2 warning(s) found. Review before launching.
|
|
65
|
-
|
|
66
|
-
|
|
67
|
-
## What it checks
|
|
68
|
-
|
|
69
|
-
- **CUDA and GPU visibility** — confirms PyTorch can see the GPU and training will not silently fall back to CPU
|
|
70
|
-
- **HuggingFace cache path** — confirms model files go to persistent storage not ephemeral disk
|
|
71
|
-
- **Accelerate GPU count** — confirms your distributed training config matches the GPUs actually on the instance
|
|
72
|
-
- **Disk space headroom** — checks free space on workspace volumes and root disk. Warns below 20 GB, blocks below 5 GB
|
|
73
|
-
- **Dataset integrity** — optional scan for duplicates, missing values, and conflicting labels
|
|
74
|
-
|
|
75
|
-
## What it does not check
|
|
76
|
-
|
|
77
|
-
- Training script correctness
|
|
78
|
-
- Model architecture compatibility
|
|
79
|
-
- Learning rate or hyperparameter safety
|
|
80
|
-
- Runtime monitoring during the job
|
|
81
|
-
- Slow dataloader or data pipeline throughput
|
|
82
|
-
- Dataloader bottleneck causing low GPU utilisation
|
|
83
|
-
|
|
84
|
-
## Why this exists
|
|
85
|
-
|
|
86
|
-
I burned approximately £1,000 on GPU training runs that failed silently. CUDA fell back to CPU with no error. 24 seconds per iteration instead of 0.4. Class weights caused loss collapse to 0.693 immediately. My dataset had 28,432 duplicate rows and 312 conflicting labels I only found after the run.
|
|
87
|
-
|
|
88
|
-
Nothing existed that caught these before the job started. So I built it.
|
|
89
|
-
|
|
90
|
-
## Real operator results
|
|
91
|
-
|
|
92
|
-
David at Neuralic ran ComputeFence on a RunPod A100. It caught HF_HOME writing to /root/.cache and an Accelerate GPU count mismatch. He fixed both before launch.
|
|
93
|
-
|
|
94
|
-
Shahzeb Ali, a computer vision engineer running client training jobs on RunPod, confirmed the storage warning matches real pod behaviour and would not have caught the HF_HOME issue explicitly without the tool.
|
|
95
|
-
|
|
96
|
-
## The problem it solves
|
|
97
|
-
|
|
98
|
-
A healthy GPU does not mean you are training the right job.
|
|
99
|
-
|
|
100
|
-
Docker makes environments reproducible. It does not check whether your HuggingFace cache is writing to ephemeral storage, whether your Accelerate config matches the GPUs actually on the instance, or whether your disk has enough headroom for checkpoints. ComputeFence addresses the job layer not the environment layer.
|
|
101
|
-
|
|
102
|
-
Five independent operators confirmed this independently. Docker does not solve job-specific configuration mistakes.
|
|
103
|
-
|
|
104
|
-
## GitHub
|
|
105
|
-
|
|
106
|
-
github.com/Francisco-Booth/ComputeFence
|
|
107
|
-
|
|
108
|
-
## PyPI
|
|
109
|
-
|
|
110
|
-
pypi.org/project/computefence
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|