computefence 0.2.4__tar.gz → 0.2.5__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,161 @@
1
+ Metadata-Version: 2.4
2
+ Name: computefence
3
+ Version: 0.2.5
4
+ Summary: Pre-flight validation for GPU training runs on rented infrastructure
5
+ License-Expression: MIT
6
+ Requires-Python: >=3.8
7
+ Description-Content-Type: text/markdown
8
+ License-File: LICENSE
9
+ Requires-Dist: rich>=13.0.0
10
+ Requires-Dist: pandas>=1.5.0
11
+ Requires-Dist: typer>=0.9.0
12
+ Requires-Dist: requests>=2.28.0
13
+ Dynamic: license-file
14
+
15
+ # ComputeFence
16
+
17
+ Pre-flight safety gate for ML training jobs on rented GPU compute. Built for RunPod, Vast.ai, Lambda Labs, CoreWeave, Paperspace, and any bare metal GPU provider.
18
+
19
+ ---
20
+
21
+ ## Install
22
+
23
+ ```bash
24
+ pip install computefence
25
+ ```
26
+
27
+ Or with UV (no install required):
28
+
29
+ ```bash
30
+ uvx computefence doctor
31
+ ```
32
+
33
+ ---
34
+
35
+ ## Usage
36
+
37
+ ```bash
38
+ computefence doctor
39
+ ```
40
+
41
+ With a dataset:
42
+
43
+ ```bash
44
+ computefence doctor --dataset train.csv --input-column text --label-column label
45
+ ```
46
+
47
+ With checkpoint output directory validation:
48
+
49
+ ```bash
50
+ computefence doctor --output-dir ./checkpoints
51
+ ```
52
+
53
+ Add to your pod startup script so it runs automatically before every job:
54
+
55
+ ```bash
56
+ pip install computefence && computefence doctor && python train.py
57
+ ```
58
+
59
+ ---
60
+
61
+ ## Example output
62
+
63
+ ```
64
+ ComputeFence v0.2.4 — Pre-flight diagnostic
65
+ ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
66
+ 2 WARNINGS · 0 BLOCKERS · 3 PASSED
67
+
68
+ Environment
69
+ ✓ Python 3.11.4
70
+ ✓ PyTorch 2.1.0 detected
71
+ ✓ CUDA available — NVIDIA A40
72
+
73
+ Storage
74
+ ⚠ HF_HOME is not set. HuggingFace will use default local cache.
75
+ Fix: export HF_HOME=/workspace/.cache/huggingface
76
+ ⚠ Root disk (/) — 14.3 GB free of 460.4 GB (below 20 GB)
77
+ Fix: Free up disk space or move checkpoints to a larger volume: df -h to check usage
78
+
79
+ Dataset
80
+ ✓ No dataset path provided — skipping dataset checks
81
+
82
+ ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
83
+ 2 warning(s) found. Review before launching.
84
+ Anonymous run stats are collected to improve ComputeFence. To opt out: touch ~/.computefence_no_telemetry
85
+ ```
86
+
87
+ ---
88
+
89
+ ## What it checks
90
+
91
+ - **CUDA and GPU visibility** — confirms PyTorch can see the GPU and training will not silently fall back to CPU
92
+ - **HuggingFace cache path** — confirms model weights go to persistent storage not ephemeral disk that disappears on pod stop
93
+ - **Accelerate GPU count** — confirms your distributed training config matches the GPUs actually on the instance
94
+ - **Disk space headroom** — checks free space on workspace volumes and root disk. Warns below 20 GB, blocks below 5 GB
95
+ - **Checkpoint output directory** — confirms your training script's output path is on persistent storage not ephemeral disk (`--output-dir`)
96
+ - **Dataset integrity** — optional scan for duplicates, missing values, and conflicting labels (`--dataset`)
97
+
98
+ ---
99
+
100
+ ## What it does not check
101
+
102
+ - Training script correctness
103
+ - Model architecture compatibility
104
+ - Learning rate or hyperparameter safety
105
+ - Runtime monitoring during the job
106
+ - Dataloader throughput or GPU utilisation during training
107
+
108
+ ---
109
+
110
+ ## Why this exists
111
+
112
+ I burned approximately £1,000 on GPU training runs that failed silently.
113
+
114
+ CUDA fell back to CPU with no error. 24 seconds per iteration instead of 0.4. Class weights caused loss collapse to 0.693 immediately. My dataset had 28,432 duplicate rows and 312 conflicting labels I only found after the run.
115
+
116
+ Nothing existed that caught these before the job started. So I built it.
117
+
118
+ ---
119
+
120
+ ## Real operator results
121
+
122
+ **David at Neuralic** ran ComputeFence on a RunPod A100. It caught HF_HOME writing to `/root/.cache` and an Accelerate GPU count mismatch. He fixed both before launch.
123
+
124
+ **Shahzeb Ali**, a computer vision engineer running client training jobs on RunPod, confirmed the storage warning matches real pod behaviour and would not have caught the HF_HOME issue explicitly without the tool.
125
+
126
+ **Seven overnight organic runs** appeared in telemetry from a Vast.ai operator running the tool six times in two minutes before a real training job — without being prompted or paid to do so.
127
+
128
+ ---
129
+
130
+ ## The problem it solves
131
+
132
+ A healthy GPU does not mean you are training the right job.
133
+
134
+ Docker makes environments reproducible. It does not check whether your HuggingFace cache is writing to ephemeral storage that disappears on pod stop, whether your Accelerate config matches the GPUs actually on the instance, or whether your training script's checkpoint output directory is on persistent storage. ComputeFence addresses the job configuration layer — not the environment layer.
135
+
136
+ HF_HOME and your checkpoint output directory are separate paths. Fixing one does not fix the other. Both disappear on pod restart if they point to ephemeral disk.
137
+
138
+ Thirteen independent ML engineers confirmed this problem independently across RunPod, Vast.ai, and AWS. Five confirmed that Docker does not solve job-specific configuration mistakes.
139
+
140
+ ---
141
+
142
+ ## Founding Design Partner pilot
143
+
144
+ Running high-cost GPU training on RunPod or Vast.ai?
145
+
146
+ We offer a hands-on **Founding Design Partner** pilot at **$99 for 3 months**:
147
+
148
+ - Personal audit of your launch templates and persistent storage configuration
149
+ - ComputeFence installed into your pod startup scripts so pre-flight runs automatically on every job
150
+ - Slack or Discord webhook alert when a check fails or blocks a launch
151
+ - Monthly 30-minute call where you shape what gets built next
152
+ - Money back if it does not catch one bad launch in 3 months
153
+
154
+ Three spots available. Email **francisco@booth.ws** to apply.
155
+
156
+ ---
157
+
158
+ ## Links
159
+
160
+ - **GitHub**: [github.com/Francisco-Booth/ComputeFence](https://github.com/Francisco-Booth/ComputeFence)
161
+ - **PyPI**: [pypi.org/project/computefence](https://pypi.org/project/computefence)
@@ -0,0 +1,147 @@
1
+ # ComputeFence
2
+
3
+ Pre-flight safety gate for ML training jobs on rented GPU compute. Built for RunPod, Vast.ai, Lambda Labs, CoreWeave, Paperspace, and any bare metal GPU provider.
4
+
5
+ ---
6
+
7
+ ## Install
8
+
9
+ ```bash
10
+ pip install computefence
11
+ ```
12
+
13
+ Or with UV (no install required):
14
+
15
+ ```bash
16
+ uvx computefence doctor
17
+ ```
18
+
19
+ ---
20
+
21
+ ## Usage
22
+
23
+ ```bash
24
+ computefence doctor
25
+ ```
26
+
27
+ With a dataset:
28
+
29
+ ```bash
30
+ computefence doctor --dataset train.csv --input-column text --label-column label
31
+ ```
32
+
33
+ With checkpoint output directory validation:
34
+
35
+ ```bash
36
+ computefence doctor --output-dir ./checkpoints
37
+ ```
38
+
39
+ Add to your pod startup script so it runs automatically before every job:
40
+
41
+ ```bash
42
+ pip install computefence && computefence doctor && python train.py
43
+ ```
44
+
45
+ ---
46
+
47
+ ## Example output
48
+
49
+ ```
50
+ ComputeFence v0.2.4 — Pre-flight diagnostic
51
+ ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
52
+ 2 WARNINGS · 0 BLOCKERS · 3 PASSED
53
+
54
+ Environment
55
+ ✓ Python 3.11.4
56
+ ✓ PyTorch 2.1.0 detected
57
+ ✓ CUDA available — NVIDIA A40
58
+
59
+ Storage
60
+ ⚠ HF_HOME is not set. HuggingFace will use default local cache.
61
+ Fix: export HF_HOME=/workspace/.cache/huggingface
62
+ ⚠ Root disk (/) — 14.3 GB free of 460.4 GB (below 20 GB)
63
+ Fix: Free up disk space or move checkpoints to a larger volume: df -h to check usage
64
+
65
+ Dataset
66
+ ✓ No dataset path provided — skipping dataset checks
67
+
68
+ ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
69
+ 2 warning(s) found. Review before launching.
70
+ Anonymous run stats are collected to improve ComputeFence. To opt out: touch ~/.computefence_no_telemetry
71
+ ```
72
+
73
+ ---
74
+
75
+ ## What it checks
76
+
77
+ - **CUDA and GPU visibility** — confirms PyTorch can see the GPU and training will not silently fall back to CPU
78
+ - **HuggingFace cache path** — confirms model weights go to persistent storage not ephemeral disk that disappears on pod stop
79
+ - **Accelerate GPU count** — confirms your distributed training config matches the GPUs actually on the instance
80
+ - **Disk space headroom** — checks free space on workspace volumes and root disk. Warns below 20 GB, blocks below 5 GB
81
+ - **Checkpoint output directory** — confirms your training script's output path is on persistent storage not ephemeral disk (`--output-dir`)
82
+ - **Dataset integrity** — optional scan for duplicates, missing values, and conflicting labels (`--dataset`)
83
+
84
+ ---
85
+
86
+ ## What it does not check
87
+
88
+ - Training script correctness
89
+ - Model architecture compatibility
90
+ - Learning rate or hyperparameter safety
91
+ - Runtime monitoring during the job
92
+ - Dataloader throughput or GPU utilisation during training
93
+
94
+ ---
95
+
96
+ ## Why this exists
97
+
98
+ I burned approximately £1,000 on GPU training runs that failed silently.
99
+
100
+ CUDA fell back to CPU with no error. 24 seconds per iteration instead of 0.4. Class weights caused loss collapse to 0.693 immediately. My dataset had 28,432 duplicate rows and 312 conflicting labels I only found after the run.
101
+
102
+ Nothing existed that caught these before the job started. So I built it.
103
+
104
+ ---
105
+
106
+ ## Real operator results
107
+
108
+ **David at Neuralic** ran ComputeFence on a RunPod A100. It caught HF_HOME writing to `/root/.cache` and an Accelerate GPU count mismatch. He fixed both before launch.
109
+
110
+ **Shahzeb Ali**, a computer vision engineer running client training jobs on RunPod, confirmed the storage warning matches real pod behaviour and would not have caught the HF_HOME issue explicitly without the tool.
111
+
112
+ **Seven overnight organic runs** appeared in telemetry from a Vast.ai operator running the tool six times in two minutes before a real training job — without being prompted or paid to do so.
113
+
114
+ ---
115
+
116
+ ## The problem it solves
117
+
118
+ A healthy GPU does not mean you are training the right job.
119
+
120
+ Docker makes environments reproducible. It does not check whether your HuggingFace cache is writing to ephemeral storage that disappears on pod stop, whether your Accelerate config matches the GPUs actually on the instance, or whether your training script's checkpoint output directory is on persistent storage. ComputeFence addresses the job configuration layer — not the environment layer.
121
+
122
+ HF_HOME and your checkpoint output directory are separate paths. Fixing one does not fix the other. Both disappear on pod restart if they point to ephemeral disk.
123
+
124
+ Thirteen independent ML engineers confirmed this problem independently across RunPod, Vast.ai, and AWS. Five confirmed that Docker does not solve job-specific configuration mistakes.
125
+
126
+ ---
127
+
128
+ ## Founding Design Partner pilot
129
+
130
+ Running high-cost GPU training on RunPod or Vast.ai?
131
+
132
+ We offer a hands-on **Founding Design Partner** pilot at **$99 for 3 months**:
133
+
134
+ - Personal audit of your launch templates and persistent storage configuration
135
+ - ComputeFence installed into your pod startup scripts so pre-flight runs automatically on every job
136
+ - Slack or Discord webhook alert when a check fails or blocks a launch
137
+ - Monthly 30-minute call where you shape what gets built next
138
+ - Money back if it does not catch one bad launch in 3 months
139
+
140
+ Three spots available. Email **francisco@booth.ws** to apply.
141
+
142
+ ---
143
+
144
+ ## Links
145
+
146
+ - **GitHub**: [github.com/Francisco-Booth/ComputeFence](https://github.com/Francisco-Booth/ComputeFence)
147
+ - **PyPI**: [pypi.org/project/computefence](https://pypi.org/project/computefence)
@@ -0,0 +1 @@
1
+ __version__ = "0.2.5"
@@ -134,3 +134,51 @@ def check_disk_headroom():
134
134
  })
135
135
 
136
136
  return results
137
+
138
+
139
+ # Paths that are wiped when a rented instance stops.
140
+ EPHEMERAL_PREFIXES = ["/root", "/tmp"]
141
+ # Paths that survive an instance restart.
142
+ PERSISTENT_PREFIXES = ["/workspace", "/runpod-volume", "/vast", "/home"]
143
+ OUTPUT_DIR_FIX = "Move your output_dir to /workspace/checkpoints or your mounted volume"
144
+
145
+
146
+ def _is_under(path, prefix):
147
+ """True when path sits at or below prefix.
148
+
149
+ Compares path components rather than string prefixes so /workspace-old is
150
+ not mistaken for a child of /workspace. Avoids Path.is_relative_to, which
151
+ needs Python 3.9 while this package supports 3.8.
152
+ """
153
+ path_parts = Path(path).parts
154
+ prefix_parts = Path(prefix).parts
155
+ return path_parts[: len(prefix_parts)] == prefix_parts
156
+
157
+
158
+ def check_output_dir(output_dir=None):
159
+ """Flag a checkpoint directory that will not survive the instance stopping.
160
+
161
+ Uses abspath rather than resolve() so a path is judged as written: on macOS
162
+ /tmp is a symlink to /private/tmp, and resolving it first would hide the
163
+ very ephemeral prefix this check exists to catch.
164
+ """
165
+ if output_dir is None:
166
+ return []
167
+
168
+ path = Path(os.path.abspath(os.path.expanduser(str(output_dir))))
169
+
170
+ is_ephemeral = any(_is_under(path, prefix) for prefix in EPHEMERAL_PREFIXES) or not any(
171
+ _is_under(path, prefix) for prefix in PERSISTENT_PREFIXES
172
+ )
173
+
174
+ if is_ephemeral:
175
+ return [{
176
+ "status": "fail",
177
+ "message": f"Output directory {path} is on ephemeral storage and will be lost when the instance stops",
178
+ "fix": OUTPUT_DIR_FIX
179
+ }]
180
+
181
+ return [{
182
+ "status": "pass",
183
+ "message": f"Output directory {path} is on persistent storage"
184
+ }]
@@ -2,7 +2,7 @@ import sys
2
2
  from rich.console import Console
3
3
  from computefence import __version__
4
4
  from computefence.checks.environment import check_environment
5
- from computefence.checks.storage import check_disk_headroom, check_storage
5
+ from computefence.checks.storage import check_disk_headroom, check_output_dir, check_storage
6
6
  from computefence.checks.dataset import check_dataset
7
7
  from computefence.telemetry import record_run
8
8
 
@@ -27,6 +27,7 @@ def doctor():
27
27
  dataset = None
28
28
  input_column = None
29
29
  label_column = None
30
+ output_dir = None
30
31
  args = sys.argv[1:]
31
32
  i = 0
32
33
  while i < len(args):
@@ -39,6 +40,9 @@ def doctor():
39
40
  elif args[i] == "--label-column" and i + 1 < len(args):
40
41
  label_column = args[i + 1]
41
42
  i += 2
43
+ elif args[i] == "--output-dir" and i + 1 < len(args):
44
+ output_dir = args[i + 1]
45
+ i += 2
42
46
  else:
43
47
  i += 1
44
48
 
@@ -47,7 +51,7 @@ def doctor():
47
51
  console.print("━" * 50)
48
52
 
49
53
  env_results = check_environment()
50
- storage_results = check_storage() + check_disk_headroom()
54
+ storage_results = check_storage() + check_disk_headroom() + check_output_dir(output_dir)
51
55
  dataset_results = check_dataset(
52
56
  dataset_path=dataset,
53
57
  input_column=input_column,
@@ -0,0 +1,161 @@
1
+ Metadata-Version: 2.4
2
+ Name: computefence
3
+ Version: 0.2.5
4
+ Summary: Pre-flight validation for GPU training runs on rented infrastructure
5
+ License-Expression: MIT
6
+ Requires-Python: >=3.8
7
+ Description-Content-Type: text/markdown
8
+ License-File: LICENSE
9
+ Requires-Dist: rich>=13.0.0
10
+ Requires-Dist: pandas>=1.5.0
11
+ Requires-Dist: typer>=0.9.0
12
+ Requires-Dist: requests>=2.28.0
13
+ Dynamic: license-file
14
+
15
+ # ComputeFence
16
+
17
+ Pre-flight safety gate for ML training jobs on rented GPU compute. Built for RunPod, Vast.ai, Lambda Labs, CoreWeave, Paperspace, and any bare metal GPU provider.
18
+
19
+ ---
20
+
21
+ ## Install
22
+
23
+ ```bash
24
+ pip install computefence
25
+ ```
26
+
27
+ Or with UV (no install required):
28
+
29
+ ```bash
30
+ uvx computefence doctor
31
+ ```
32
+
33
+ ---
34
+
35
+ ## Usage
36
+
37
+ ```bash
38
+ computefence doctor
39
+ ```
40
+
41
+ With a dataset:
42
+
43
+ ```bash
44
+ computefence doctor --dataset train.csv --input-column text --label-column label
45
+ ```
46
+
47
+ With checkpoint output directory validation:
48
+
49
+ ```bash
50
+ computefence doctor --output-dir ./checkpoints
51
+ ```
52
+
53
+ Add to your pod startup script so it runs automatically before every job:
54
+
55
+ ```bash
56
+ pip install computefence && computefence doctor && python train.py
57
+ ```
58
+
59
+ ---
60
+
61
+ ## Example output
62
+
63
+ ```
64
+ ComputeFence v0.2.4 — Pre-flight diagnostic
65
+ ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
66
+ 2 WARNINGS · 0 BLOCKERS · 3 PASSED
67
+
68
+ Environment
69
+ ✓ Python 3.11.4
70
+ ✓ PyTorch 2.1.0 detected
71
+ ✓ CUDA available — NVIDIA A40
72
+
73
+ Storage
74
+ ⚠ HF_HOME is not set. HuggingFace will use default local cache.
75
+ Fix: export HF_HOME=/workspace/.cache/huggingface
76
+ ⚠ Root disk (/) — 14.3 GB free of 460.4 GB (below 20 GB)
77
+ Fix: Free up disk space or move checkpoints to a larger volume: df -h to check usage
78
+
79
+ Dataset
80
+ ✓ No dataset path provided — skipping dataset checks
81
+
82
+ ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
83
+ 2 warning(s) found. Review before launching.
84
+ Anonymous run stats are collected to improve ComputeFence. To opt out: touch ~/.computefence_no_telemetry
85
+ ```
86
+
87
+ ---
88
+
89
+ ## What it checks
90
+
91
+ - **CUDA and GPU visibility** — confirms PyTorch can see the GPU and training will not silently fall back to CPU
92
+ - **HuggingFace cache path** — confirms model weights go to persistent storage not ephemeral disk that disappears on pod stop
93
+ - **Accelerate GPU count** — confirms your distributed training config matches the GPUs actually on the instance
94
+ - **Disk space headroom** — checks free space on workspace volumes and root disk. Warns below 20 GB, blocks below 5 GB
95
+ - **Checkpoint output directory** — confirms your training script's output path is on persistent storage not ephemeral disk (`--output-dir`)
96
+ - **Dataset integrity** — optional scan for duplicates, missing values, and conflicting labels (`--dataset`)
97
+
98
+ ---
99
+
100
+ ## What it does not check
101
+
102
+ - Training script correctness
103
+ - Model architecture compatibility
104
+ - Learning rate or hyperparameter safety
105
+ - Runtime monitoring during the job
106
+ - Dataloader throughput or GPU utilisation during training
107
+
108
+ ---
109
+
110
+ ## Why this exists
111
+
112
+ I burned approximately £1,000 on GPU training runs that failed silently.
113
+
114
+ CUDA fell back to CPU with no error. 24 seconds per iteration instead of 0.4. Class weights caused loss collapse to 0.693 immediately. My dataset had 28,432 duplicate rows and 312 conflicting labels I only found after the run.
115
+
116
+ Nothing existed that caught these before the job started. So I built it.
117
+
118
+ ---
119
+
120
+ ## Real operator results
121
+
122
+ **David at Neuralic** ran ComputeFence on a RunPod A100. It caught HF_HOME writing to `/root/.cache` and an Accelerate GPU count mismatch. He fixed both before launch.
123
+
124
+ **Shahzeb Ali**, a computer vision engineer running client training jobs on RunPod, confirmed the storage warning matches real pod behaviour and would not have caught the HF_HOME issue explicitly without the tool.
125
+
126
+ **Seven overnight organic runs** appeared in telemetry from a Vast.ai operator running the tool six times in two minutes before a real training job — without being prompted or paid to do so.
127
+
128
+ ---
129
+
130
+ ## The problem it solves
131
+
132
+ A healthy GPU does not mean you are training the right job.
133
+
134
+ Docker makes environments reproducible. It does not check whether your HuggingFace cache is writing to ephemeral storage that disappears on pod stop, whether your Accelerate config matches the GPUs actually on the instance, or whether your training script's checkpoint output directory is on persistent storage. ComputeFence addresses the job configuration layer — not the environment layer.
135
+
136
+ HF_HOME and your checkpoint output directory are separate paths. Fixing one does not fix the other. Both disappear on pod restart if they point to ephemeral disk.
137
+
138
+ Thirteen independent ML engineers confirmed this problem independently across RunPod, Vast.ai, and AWS. Five confirmed that Docker does not solve job-specific configuration mistakes.
139
+
140
+ ---
141
+
142
+ ## Founding Design Partner pilot
143
+
144
+ Running high-cost GPU training on RunPod or Vast.ai?
145
+
146
+ We offer a hands-on **Founding Design Partner** pilot at **$99 for 3 months**:
147
+
148
+ - Personal audit of your launch templates and persistent storage configuration
149
+ - ComputeFence installed into your pod startup scripts so pre-flight runs automatically on every job
150
+ - Slack or Discord webhook alert when a check fails or blocks a launch
151
+ - Monthly 30-minute call where you shape what gets built next
152
+ - Money back if it does not catch one bad launch in 3 months
153
+
154
+ Three spots available. Email **francisco@booth.ws** to apply.
155
+
156
+ ---
157
+
158
+ ## Links
159
+
160
+ - **GitHub**: [github.com/Francisco-Booth/ComputeFence](https://github.com/Francisco-Booth/ComputeFence)
161
+ - **PyPI**: [pypi.org/project/computefence](https://pypi.org/project/computefence)
@@ -1,110 +0,0 @@
1
- Metadata-Version: 2.4
2
- Name: computefence
3
- Version: 0.2.4
4
- Summary: Pre-flight validation for GPU training runs on rented infrastructure
5
- License-Expression: MIT
6
- Requires-Python: >=3.8
7
- Description-Content-Type: text/markdown
8
- License-File: LICENSE
9
- Requires-Dist: rich>=13.0.0
10
- Requires-Dist: pandas>=1.5.0
11
- Requires-Dist: typer>=0.9.0
12
- Requires-Dist: requests>=2.28.0
13
- Dynamic: license-file
14
-
15
- # ComputeFence
16
-
17
- Pre-flight safety gate for ML training jobs on rented GPU compute. Built for RunPod, Vast.ai, Lambda Labs, and similar bare metal providers.
18
-
19
- ## Install
20
-
21
- ```bash
22
- pip install computefence
23
- ```
24
-
25
- Or with UV:
26
-
27
- ```bash
28
- uvx computefence doctor
29
- ```
30
-
31
- ## Usage
32
-
33
- ```bash
34
- computefence doctor
35
- ```
36
-
37
- With a dataset:
38
-
39
- ```bash
40
- computefence doctor --dataset train.csv --input-column text --label-column label
41
- ```
42
-
43
- ## Example output
44
-
45
- ComputeFence v0.2.0 — Pre-flight diagnostic
46
- ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
47
- 2 WARNINGS · 0 BLOCKERS · 3 PASSED
48
-
49
- Environment
50
- ✓ Python 3.11.4
51
- ✓ PyTorch 2.1.0 detected
52
- ✓ CUDA available — NVIDIA A40
53
-
54
- Storage
55
- ⚠ HF_HOME is not set. HuggingFace will use default local cache.
56
- Fix: export HF_HOME=/workspace/.cache/huggingface
57
- ⚠ Root disk (/) — 14.3 GB free of 460.4 GB (below 20 GB)
58
- Fix: Free up disk space or attach a larger volume before launching
59
-
60
- Dataset
61
- ✓ No dataset path provided — skipping dataset checks
62
-
63
- ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
64
- 2 warning(s) found. Review before launching.
65
-
66
-
67
- ## What it checks
68
-
69
- - **CUDA and GPU visibility** — confirms PyTorch can see the GPU and training will not silently fall back to CPU
70
- - **HuggingFace cache path** — confirms model files go to persistent storage not ephemeral disk
71
- - **Accelerate GPU count** — confirms your distributed training config matches the GPUs actually on the instance
72
- - **Disk space headroom** — checks free space on workspace volumes and root disk. Warns below 20 GB, blocks below 5 GB
73
- - **Dataset integrity** — optional scan for duplicates, missing values, and conflicting labels
74
-
75
- ## What it does not check
76
-
77
- - Training script correctness
78
- - Model architecture compatibility
79
- - Learning rate or hyperparameter safety
80
- - Runtime monitoring during the job
81
- - Slow dataloader or data pipeline throughput
82
- - Dataloader bottleneck causing low GPU utilisation
83
-
84
- ## Why this exists
85
-
86
- I burned approximately £1,000 on GPU training runs that failed silently. CUDA fell back to CPU with no error. 24 seconds per iteration instead of 0.4. Class weights caused loss collapse to 0.693 immediately. My dataset had 28,432 duplicate rows and 312 conflicting labels I only found after the run.
87
-
88
- Nothing existed that caught these before the job started. So I built it.
89
-
90
- ## Real operator results
91
-
92
- David at Neuralic ran ComputeFence on a RunPod A100. It caught HF_HOME writing to /root/.cache and an Accelerate GPU count mismatch. He fixed both before launch.
93
-
94
- Shahzeb Ali, a computer vision engineer running client training jobs on RunPod, confirmed the storage warning matches real pod behaviour and would not have caught the HF_HOME issue explicitly without the tool.
95
-
96
- ## The problem it solves
97
-
98
- A healthy GPU does not mean you are training the right job.
99
-
100
- Docker makes environments reproducible. It does not check whether your HuggingFace cache is writing to ephemeral storage, whether your Accelerate config matches the GPUs actually on the instance, or whether your disk has enough headroom for checkpoints. ComputeFence addresses the job layer not the environment layer.
101
-
102
- Five independent operators confirmed this independently. Docker does not solve job-specific configuration mistakes.
103
-
104
- ## GitHub
105
-
106
- github.com/Francisco-Booth/ComputeFence
107
-
108
- ## PyPI
109
-
110
- pypi.org/project/computefence
@@ -1,96 +0,0 @@
1
- # ComputeFence
2
-
3
- Pre-flight safety gate for ML training jobs on rented GPU compute. Built for RunPod, Vast.ai, Lambda Labs, and similar bare metal providers.
4
-
5
- ## Install
6
-
7
- ```bash
8
- pip install computefence
9
- ```
10
-
11
- Or with UV:
12
-
13
- ```bash
14
- uvx computefence doctor
15
- ```
16
-
17
- ## Usage
18
-
19
- ```bash
20
- computefence doctor
21
- ```
22
-
23
- With a dataset:
24
-
25
- ```bash
26
- computefence doctor --dataset train.csv --input-column text --label-column label
27
- ```
28
-
29
- ## Example output
30
-
31
- ComputeFence v0.2.0 — Pre-flight diagnostic
32
- ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
33
- 2 WARNINGS · 0 BLOCKERS · 3 PASSED
34
-
35
- Environment
36
- ✓ Python 3.11.4
37
- ✓ PyTorch 2.1.0 detected
38
- ✓ CUDA available — NVIDIA A40
39
-
40
- Storage
41
- ⚠ HF_HOME is not set. HuggingFace will use default local cache.
42
- Fix: export HF_HOME=/workspace/.cache/huggingface
43
- ⚠ Root disk (/) — 14.3 GB free of 460.4 GB (below 20 GB)
44
- Fix: Free up disk space or attach a larger volume before launching
45
-
46
- Dataset
47
- ✓ No dataset path provided — skipping dataset checks
48
-
49
- ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
50
- 2 warning(s) found. Review before launching.
51
-
52
-
53
- ## What it checks
54
-
55
- - **CUDA and GPU visibility** — confirms PyTorch can see the GPU and training will not silently fall back to CPU
56
- - **HuggingFace cache path** — confirms model files go to persistent storage not ephemeral disk
57
- - **Accelerate GPU count** — confirms your distributed training config matches the GPUs actually on the instance
58
- - **Disk space headroom** — checks free space on workspace volumes and root disk. Warns below 20 GB, blocks below 5 GB
59
- - **Dataset integrity** — optional scan for duplicates, missing values, and conflicting labels
60
-
61
- ## What it does not check
62
-
63
- - Training script correctness
64
- - Model architecture compatibility
65
- - Learning rate or hyperparameter safety
66
- - Runtime monitoring during the job
67
- - Slow dataloader or data pipeline throughput
68
- - Dataloader bottleneck causing low GPU utilisation
69
-
70
- ## Why this exists
71
-
72
- I burned approximately £1,000 on GPU training runs that failed silently. CUDA fell back to CPU with no error. 24 seconds per iteration instead of 0.4. Class weights caused loss collapse to 0.693 immediately. My dataset had 28,432 duplicate rows and 312 conflicting labels I only found after the run.
73
-
74
- Nothing existed that caught these before the job started. So I built it.
75
-
76
- ## Real operator results
77
-
78
- David at Neuralic ran ComputeFence on a RunPod A100. It caught HF_HOME writing to /root/.cache and an Accelerate GPU count mismatch. He fixed both before launch.
79
-
80
- Shahzeb Ali, a computer vision engineer running client training jobs on RunPod, confirmed the storage warning matches real pod behaviour and would not have caught the HF_HOME issue explicitly without the tool.
81
-
82
- ## The problem it solves
83
-
84
- A healthy GPU does not mean you are training the right job.
85
-
86
- Docker makes environments reproducible. It does not check whether your HuggingFace cache is writing to ephemeral storage, whether your Accelerate config matches the GPUs actually on the instance, or whether your disk has enough headroom for checkpoints. ComputeFence addresses the job layer not the environment layer.
87
-
88
- Five independent operators confirmed this independently. Docker does not solve job-specific configuration mistakes.
89
-
90
- ## GitHub
91
-
92
- github.com/Francisco-Booth/ComputeFence
93
-
94
- ## PyPI
95
-
96
- pypi.org/project/computefence
@@ -1 +0,0 @@
1
- __version__ = "0.2.4"
@@ -1,110 +0,0 @@
1
- Metadata-Version: 2.4
2
- Name: computefence
3
- Version: 0.2.4
4
- Summary: Pre-flight validation for GPU training runs on rented infrastructure
5
- License-Expression: MIT
6
- Requires-Python: >=3.8
7
- Description-Content-Type: text/markdown
8
- License-File: LICENSE
9
- Requires-Dist: rich>=13.0.0
10
- Requires-Dist: pandas>=1.5.0
11
- Requires-Dist: typer>=0.9.0
12
- Requires-Dist: requests>=2.28.0
13
- Dynamic: license-file
14
-
15
- # ComputeFence
16
-
17
- Pre-flight safety gate for ML training jobs on rented GPU compute. Built for RunPod, Vast.ai, Lambda Labs, and similar bare metal providers.
18
-
19
- ## Install
20
-
21
- ```bash
22
- pip install computefence
23
- ```
24
-
25
- Or with UV:
26
-
27
- ```bash
28
- uvx computefence doctor
29
- ```
30
-
31
- ## Usage
32
-
33
- ```bash
34
- computefence doctor
35
- ```
36
-
37
- With a dataset:
38
-
39
- ```bash
40
- computefence doctor --dataset train.csv --input-column text --label-column label
41
- ```
42
-
43
- ## Example output
44
-
45
- ComputeFence v0.2.0 — Pre-flight diagnostic
46
- ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
47
- 2 WARNINGS · 0 BLOCKERS · 3 PASSED
48
-
49
- Environment
50
- ✓ Python 3.11.4
51
- ✓ PyTorch 2.1.0 detected
52
- ✓ CUDA available — NVIDIA A40
53
-
54
- Storage
55
- ⚠ HF_HOME is not set. HuggingFace will use default local cache.
56
- Fix: export HF_HOME=/workspace/.cache/huggingface
57
- ⚠ Root disk (/) — 14.3 GB free of 460.4 GB (below 20 GB)
58
- Fix: Free up disk space or attach a larger volume before launching
59
-
60
- Dataset
61
- ✓ No dataset path provided — skipping dataset checks
62
-
63
- ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
64
- 2 warning(s) found. Review before launching.
65
-
66
-
67
- ## What it checks
68
-
69
- - **CUDA and GPU visibility** — confirms PyTorch can see the GPU and training will not silently fall back to CPU
70
- - **HuggingFace cache path** — confirms model files go to persistent storage not ephemeral disk
71
- - **Accelerate GPU count** — confirms your distributed training config matches the GPUs actually on the instance
72
- - **Disk space headroom** — checks free space on workspace volumes and root disk. Warns below 20 GB, blocks below 5 GB
73
- - **Dataset integrity** — optional scan for duplicates, missing values, and conflicting labels
74
-
75
- ## What it does not check
76
-
77
- - Training script correctness
78
- - Model architecture compatibility
79
- - Learning rate or hyperparameter safety
80
- - Runtime monitoring during the job
81
- - Slow dataloader or data pipeline throughput
82
- - Dataloader bottleneck causing low GPU utilisation
83
-
84
- ## Why this exists
85
-
86
- I burned approximately £1,000 on GPU training runs that failed silently. CUDA fell back to CPU with no error. 24 seconds per iteration instead of 0.4. Class weights caused loss collapse to 0.693 immediately. My dataset had 28,432 duplicate rows and 312 conflicting labels I only found after the run.
87
-
88
- Nothing existed that caught these before the job started. So I built it.
89
-
90
- ## Real operator results
91
-
92
- David at Neuralic ran ComputeFence on a RunPod A100. It caught HF_HOME writing to /root/.cache and an Accelerate GPU count mismatch. He fixed both before launch.
93
-
94
- Shahzeb Ali, a computer vision engineer running client training jobs on RunPod, confirmed the storage warning matches real pod behaviour and would not have caught the HF_HOME issue explicitly without the tool.
95
-
96
- ## The problem it solves
97
-
98
- A healthy GPU does not mean you are training the right job.
99
-
100
- Docker makes environments reproducible. It does not check whether your HuggingFace cache is writing to ephemeral storage, whether your Accelerate config matches the GPUs actually on the instance, or whether your disk has enough headroom for checkpoints. ComputeFence addresses the job layer not the environment layer.
101
-
102
- Five independent operators confirmed this independently. Docker does not solve job-specific configuration mistakes.
103
-
104
- ## GitHub
105
-
106
- github.com/Francisco-Booth/ComputeFence
107
-
108
- ## PyPI
109
-
110
- pypi.org/project/computefence
File without changes
File without changes