computefence 0.2.0__tar.gz → 0.2.3__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- computefence-0.2.3/PKG-INFO +110 -0
- computefence-0.2.3/README.md +96 -0
- computefence-0.2.3/computefence/__init__.py +1 -0
- {computefence-0.2.0 → computefence-0.2.3}/computefence/cli.py +7 -14
- computefence-0.2.3/computefence/telemetry.py +65 -0
- computefence-0.2.3/computefence.egg-info/PKG-INFO +110 -0
- {computefence-0.2.0 → computefence-0.2.3}/computefence.egg-info/SOURCES.txt +1 -0
- computefence-0.2.0/PKG-INFO +0 -49
- computefence-0.2.0/README.md +0 -35
- computefence-0.2.0/computefence/__init__.py +0 -1
- computefence-0.2.0/computefence.egg-info/PKG-INFO +0 -49
- {computefence-0.2.0 → computefence-0.2.3}/LICENSE +0 -0
- {computefence-0.2.0 → computefence-0.2.3}/computefence/checks/__init__.py +0 -0
- {computefence-0.2.0 → computefence-0.2.3}/computefence/checks/dataset.py +0 -0
- {computefence-0.2.0 → computefence-0.2.3}/computefence/checks/environment.py +0 -0
- {computefence-0.2.0 → computefence-0.2.3}/computefence/checks/storage.py +0 -0
- {computefence-0.2.0 → computefence-0.2.3}/computefence.egg-info/dependency_links.txt +0 -0
- {computefence-0.2.0 → computefence-0.2.3}/computefence.egg-info/entry_points.txt +0 -0
- {computefence-0.2.0 → computefence-0.2.3}/computefence.egg-info/requires.txt +0 -0
- {computefence-0.2.0 → computefence-0.2.3}/computefence.egg-info/top_level.txt +0 -0
- {computefence-0.2.0 → computefence-0.2.3}/pyproject.toml +0 -0
- {computefence-0.2.0 → computefence-0.2.3}/setup.cfg +0 -0
|
@@ -0,0 +1,110 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: computefence
|
|
3
|
+
Version: 0.2.3
|
|
4
|
+
Summary: Pre-flight validation for GPU training runs on rented infrastructure
|
|
5
|
+
License-Expression: MIT
|
|
6
|
+
Requires-Python: >=3.8
|
|
7
|
+
Description-Content-Type: text/markdown
|
|
8
|
+
License-File: LICENSE
|
|
9
|
+
Requires-Dist: rich>=13.0.0
|
|
10
|
+
Requires-Dist: pandas>=1.5.0
|
|
11
|
+
Requires-Dist: typer>=0.9.0
|
|
12
|
+
Requires-Dist: requests>=2.28.0
|
|
13
|
+
Dynamic: license-file
|
|
14
|
+
|
|
15
|
+
# ComputeFence
|
|
16
|
+
|
|
17
|
+
Pre-flight safety gate for ML training jobs on rented GPU compute. Built for RunPod, Vast.ai, Lambda Labs, and similar bare metal providers.
|
|
18
|
+
|
|
19
|
+
## Install
|
|
20
|
+
|
|
21
|
+
```bash
|
|
22
|
+
pip install computefence
|
|
23
|
+
```
|
|
24
|
+
|
|
25
|
+
Or with UV:
|
|
26
|
+
|
|
27
|
+
```bash
|
|
28
|
+
uvx computefence doctor
|
|
29
|
+
```
|
|
30
|
+
|
|
31
|
+
## Usage
|
|
32
|
+
|
|
33
|
+
```bash
|
|
34
|
+
computefence doctor
|
|
35
|
+
```
|
|
36
|
+
|
|
37
|
+
With a dataset:
|
|
38
|
+
|
|
39
|
+
```bash
|
|
40
|
+
computefence doctor --dataset train.csv --input-column text --label-column label
|
|
41
|
+
```
|
|
42
|
+
|
|
43
|
+
## Example output
|
|
44
|
+
|
|
45
|
+
ComputeFence v0.2.0 — Pre-flight diagnostic
|
|
46
|
+
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
|
|
47
|
+
2 WARNINGS · 0 BLOCKERS · 3 PASSED
|
|
48
|
+
|
|
49
|
+
Environment
|
|
50
|
+
✓ Python 3.11.4
|
|
51
|
+
✓ PyTorch 2.1.0 detected
|
|
52
|
+
✓ CUDA available — NVIDIA A40
|
|
53
|
+
|
|
54
|
+
Storage
|
|
55
|
+
⚠ HF_HOME is not set. HuggingFace will use default local cache.
|
|
56
|
+
Fix: export HF_HOME=/workspace/.cache/huggingface
|
|
57
|
+
⚠ Root disk (/) — 14.3 GB free of 460.4 GB (below 20 GB)
|
|
58
|
+
Fix: Free up disk space or attach a larger volume before launching
|
|
59
|
+
|
|
60
|
+
Dataset
|
|
61
|
+
✓ No dataset path provided — skipping dataset checks
|
|
62
|
+
|
|
63
|
+
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
|
|
64
|
+
2 warning(s) found. Review before launching.
|
|
65
|
+
|
|
66
|
+
|
|
67
|
+
## What it checks
|
|
68
|
+
|
|
69
|
+
- **CUDA and GPU visibility** — confirms PyTorch can see the GPU and training will not silently fall back to CPU
|
|
70
|
+
- **HuggingFace cache path** — confirms model files go to persistent storage not ephemeral disk
|
|
71
|
+
- **Accelerate GPU count** — confirms your distributed training config matches the GPUs actually on the instance
|
|
72
|
+
- **Disk space headroom** — checks free space on workspace volumes and root disk. Warns below 20 GB, blocks below 5 GB
|
|
73
|
+
- **Dataset integrity** — optional scan for duplicates, missing values, and conflicting labels
|
|
74
|
+
|
|
75
|
+
## What it does not check
|
|
76
|
+
|
|
77
|
+
- Training script correctness
|
|
78
|
+
- Model architecture compatibility
|
|
79
|
+
- Learning rate or hyperparameter safety
|
|
80
|
+
- Runtime monitoring during the job
|
|
81
|
+
- Slow dataloader or data pipeline throughput
|
|
82
|
+
- Dataloader bottleneck causing low GPU utilisation
|
|
83
|
+
|
|
84
|
+
## Why this exists
|
|
85
|
+
|
|
86
|
+
I burned approximately £1,000 on GPU training runs that failed silently. CUDA fell back to CPU with no error. 24 seconds per iteration instead of 0.4. Class weights caused loss collapse to 0.693 immediately. My dataset had 28,432 duplicate rows and 312 conflicting labels I only found after the run.
|
|
87
|
+
|
|
88
|
+
Nothing existed that caught these before the job started. So I built it.
|
|
89
|
+
|
|
90
|
+
## Real operator results
|
|
91
|
+
|
|
92
|
+
David at Neuralic ran ComputeFence on a RunPod A100. It caught HF_HOME writing to /root/.cache and an Accelerate GPU count mismatch. He fixed both before launch.
|
|
93
|
+
|
|
94
|
+
Shahzeb Ali, a computer vision engineer running client training jobs on RunPod, confirmed the storage warning matches real pod behaviour and would not have caught the HF_HOME issue explicitly without the tool.
|
|
95
|
+
|
|
96
|
+
## The problem it solves
|
|
97
|
+
|
|
98
|
+
A healthy GPU does not mean you are training the right job.
|
|
99
|
+
|
|
100
|
+
Docker makes environments reproducible. It does not check whether your HuggingFace cache is writing to ephemeral storage, whether your Accelerate config matches the GPUs actually on the instance, or whether your disk has enough headroom for checkpoints. ComputeFence addresses the job layer not the environment layer.
|
|
101
|
+
|
|
102
|
+
Five independent operators confirmed this independently. Docker does not solve job-specific configuration mistakes.
|
|
103
|
+
|
|
104
|
+
## GitHub
|
|
105
|
+
|
|
106
|
+
github.com/Francisco-Booth/ComputeFence
|
|
107
|
+
|
|
108
|
+
## PyPI
|
|
109
|
+
|
|
110
|
+
pypi.org/project/computefence
|
|
@@ -0,0 +1,96 @@
|
|
|
1
|
+
# ComputeFence
|
|
2
|
+
|
|
3
|
+
Pre-flight safety gate for ML training jobs on rented GPU compute. Built for RunPod, Vast.ai, Lambda Labs, and similar bare metal providers.
|
|
4
|
+
|
|
5
|
+
## Install
|
|
6
|
+
|
|
7
|
+
```bash
|
|
8
|
+
pip install computefence
|
|
9
|
+
```
|
|
10
|
+
|
|
11
|
+
Or with UV:
|
|
12
|
+
|
|
13
|
+
```bash
|
|
14
|
+
uvx computefence doctor
|
|
15
|
+
```
|
|
16
|
+
|
|
17
|
+
## Usage
|
|
18
|
+
|
|
19
|
+
```bash
|
|
20
|
+
computefence doctor
|
|
21
|
+
```
|
|
22
|
+
|
|
23
|
+
With a dataset:
|
|
24
|
+
|
|
25
|
+
```bash
|
|
26
|
+
computefence doctor --dataset train.csv --input-column text --label-column label
|
|
27
|
+
```
|
|
28
|
+
|
|
29
|
+
## Example output
|
|
30
|
+
|
|
31
|
+
ComputeFence v0.2.0 — Pre-flight diagnostic
|
|
32
|
+
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
|
|
33
|
+
2 WARNINGS · 0 BLOCKERS · 3 PASSED
|
|
34
|
+
|
|
35
|
+
Environment
|
|
36
|
+
✓ Python 3.11.4
|
|
37
|
+
✓ PyTorch 2.1.0 detected
|
|
38
|
+
✓ CUDA available — NVIDIA A40
|
|
39
|
+
|
|
40
|
+
Storage
|
|
41
|
+
⚠ HF_HOME is not set. HuggingFace will use default local cache.
|
|
42
|
+
Fix: export HF_HOME=/workspace/.cache/huggingface
|
|
43
|
+
⚠ Root disk (/) — 14.3 GB free of 460.4 GB (below 20 GB)
|
|
44
|
+
Fix: Free up disk space or attach a larger volume before launching
|
|
45
|
+
|
|
46
|
+
Dataset
|
|
47
|
+
✓ No dataset path provided — skipping dataset checks
|
|
48
|
+
|
|
49
|
+
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
|
|
50
|
+
2 warning(s) found. Review before launching.
|
|
51
|
+
|
|
52
|
+
|
|
53
|
+
## What it checks
|
|
54
|
+
|
|
55
|
+
- **CUDA and GPU visibility** — confirms PyTorch can see the GPU and training will not silently fall back to CPU
|
|
56
|
+
- **HuggingFace cache path** — confirms model files go to persistent storage not ephemeral disk
|
|
57
|
+
- **Accelerate GPU count** — confirms your distributed training config matches the GPUs actually on the instance
|
|
58
|
+
- **Disk space headroom** — checks free space on workspace volumes and root disk. Warns below 20 GB, blocks below 5 GB
|
|
59
|
+
- **Dataset integrity** — optional scan for duplicates, missing values, and conflicting labels
|
|
60
|
+
|
|
61
|
+
## What it does not check
|
|
62
|
+
|
|
63
|
+
- Training script correctness
|
|
64
|
+
- Model architecture compatibility
|
|
65
|
+
- Learning rate or hyperparameter safety
|
|
66
|
+
- Runtime monitoring during the job
|
|
67
|
+
- Slow dataloader or data pipeline throughput
|
|
68
|
+
- Dataloader bottleneck causing low GPU utilisation
|
|
69
|
+
|
|
70
|
+
## Why this exists
|
|
71
|
+
|
|
72
|
+
I burned approximately £1,000 on GPU training runs that failed silently. CUDA fell back to CPU with no error. 24 seconds per iteration instead of 0.4. Class weights caused loss collapse to 0.693 immediately. My dataset had 28,432 duplicate rows and 312 conflicting labels I only found after the run.
|
|
73
|
+
|
|
74
|
+
Nothing existed that caught these before the job started. So I built it.
|
|
75
|
+
|
|
76
|
+
## Real operator results
|
|
77
|
+
|
|
78
|
+
David at Neuralic ran ComputeFence on a RunPod A100. It caught HF_HOME writing to /root/.cache and an Accelerate GPU count mismatch. He fixed both before launch.
|
|
79
|
+
|
|
80
|
+
Shahzeb Ali, a computer vision engineer running client training jobs on RunPod, confirmed the storage warning matches real pod behaviour and would not have caught the HF_HOME issue explicitly without the tool.
|
|
81
|
+
|
|
82
|
+
## The problem it solves
|
|
83
|
+
|
|
84
|
+
A healthy GPU does not mean you are training the right job.
|
|
85
|
+
|
|
86
|
+
Docker makes environments reproducible. It does not check whether your HuggingFace cache is writing to ephemeral storage, whether your Accelerate config matches the GPUs actually on the instance, or whether your disk has enough headroom for checkpoints. ComputeFence addresses the job layer not the environment layer.
|
|
87
|
+
|
|
88
|
+
Five independent operators confirmed this independently. Docker does not solve job-specific configuration mistakes.
|
|
89
|
+
|
|
90
|
+
## GitHub
|
|
91
|
+
|
|
92
|
+
github.com/Francisco-Booth/ComputeFence
|
|
93
|
+
|
|
94
|
+
## PyPI
|
|
95
|
+
|
|
96
|
+
pypi.org/project/computefence
|
|
@@ -0,0 +1 @@
|
|
|
1
|
+
__version__ = "0.2.3"
|
|
@@ -1,19 +1,17 @@
|
|
|
1
1
|
import sys
|
|
2
2
|
from rich.console import Console
|
|
3
|
-
|
|
4
3
|
from computefence import __version__
|
|
5
4
|
from computefence.checks.environment import check_environment
|
|
6
5
|
from computefence.checks.storage import check_disk_headroom, check_storage
|
|
7
6
|
from computefence.checks.dataset import check_dataset
|
|
7
|
+
from computefence.telemetry import record_run
|
|
8
8
|
|
|
9
9
|
console = Console()
|
|
10
10
|
|
|
11
|
-
|
|
12
11
|
def print_result(result):
|
|
13
12
|
status = result.get("status")
|
|
14
13
|
message = result.get("message")
|
|
15
14
|
fix = result.get("fix")
|
|
16
|
-
|
|
17
15
|
if status == "pass":
|
|
18
16
|
console.print(f" [green]✓[/green] {message}")
|
|
19
17
|
elif status == "warn":
|
|
@@ -25,12 +23,10 @@ def print_result(result):
|
|
|
25
23
|
if fix:
|
|
26
24
|
console.print(f" [dim]Fix: {fix}[/dim]")
|
|
27
25
|
|
|
28
|
-
|
|
29
26
|
def doctor():
|
|
30
27
|
dataset = None
|
|
31
28
|
input_column = None
|
|
32
29
|
label_column = None
|
|
33
|
-
|
|
34
30
|
args = sys.argv[1:]
|
|
35
31
|
i = 0
|
|
36
32
|
while i < len(args):
|
|
@@ -50,8 +46,6 @@ def doctor():
|
|
|
50
46
|
console.print(f"[bold]ComputeFence v{__version__} — Pre-flight diagnostic[/bold]")
|
|
51
47
|
console.print("━" * 50)
|
|
52
48
|
|
|
53
|
-
# Run every check up front so the verdict line can be counted and printed
|
|
54
|
-
# above the individual results.
|
|
55
49
|
env_results = check_environment()
|
|
56
50
|
storage_results = check_storage() + check_disk_headroom()
|
|
57
51
|
dataset_results = check_dataset(
|
|
@@ -61,6 +55,7 @@ def doctor():
|
|
|
61
55
|
)
|
|
62
56
|
|
|
63
57
|
all_results = env_results + storage_results + dataset_results
|
|
58
|
+
|
|
64
59
|
failures = [r for r in all_results if r["status"] == "fail"]
|
|
65
60
|
warnings = [r for r in all_results if r["status"] == "warn"]
|
|
66
61
|
passed = [r for r in all_results if r["status"] == "pass"]
|
|
@@ -71,34 +66,33 @@ def doctor():
|
|
|
71
66
|
f"[red]{len(failures)} BLOCKERS[/red] · "
|
|
72
67
|
f"[green]{len(passed)} PASSED[/green]"
|
|
73
68
|
)
|
|
74
|
-
|
|
75
69
|
console.print()
|
|
76
70
|
console.print("[bold blue]Environment[/bold blue]")
|
|
77
71
|
for result in env_results:
|
|
78
72
|
print_result(result)
|
|
79
|
-
|
|
80
73
|
console.print()
|
|
81
74
|
console.print("[bold blue]Storage[/bold blue]")
|
|
82
75
|
for result in storage_results:
|
|
83
76
|
print_result(result)
|
|
84
|
-
|
|
85
77
|
console.print()
|
|
86
78
|
console.print("[bold blue]Dataset[/bold blue]")
|
|
87
79
|
for result in dataset_results:
|
|
88
80
|
print_result(result)
|
|
89
|
-
|
|
90
81
|
console.print()
|
|
91
82
|
console.print("━" * 50)
|
|
92
|
-
|
|
93
83
|
if failures:
|
|
94
84
|
console.print("[red]" + str(len(failures)) + " error(s) and " + str(len(warnings)) + " warning(s) found. Fix errors before launching.[/red]")
|
|
95
85
|
elif warnings:
|
|
96
86
|
console.print("[yellow]" + str(len(warnings)) + " warning(s) found. Review before launching.[/yellow]")
|
|
97
87
|
else:
|
|
98
88
|
console.print("[green]All checks passed. Safe to launch.[/green]")
|
|
99
|
-
|
|
89
|
+
console.print(
|
|
90
|
+
"[dim]Anonymous run stats are collected to improve ComputeFence. "
|
|
91
|
+
"To opt out: touch ~/.computefence_no_telemetry[/dim]"
|
|
92
|
+
)
|
|
100
93
|
console.print()
|
|
101
94
|
|
|
95
|
+
record_run(all_results)
|
|
102
96
|
|
|
103
97
|
def main():
|
|
104
98
|
args = sys.argv[1:]
|
|
@@ -108,6 +102,5 @@ def main():
|
|
|
108
102
|
console.print(f"[red]Unknown command: {args[0]}[/red]")
|
|
109
103
|
console.print("Usage: computefence doctor")
|
|
110
104
|
|
|
111
|
-
|
|
112
105
|
if __name__ == "__main__":
|
|
113
106
|
main()
|
|
@@ -0,0 +1,65 @@
|
|
|
1
|
+
import hashlib
|
|
2
|
+
import platform
|
|
3
|
+
import sys
|
|
4
|
+
import threading
|
|
5
|
+
import uuid
|
|
6
|
+
from pathlib import Path
|
|
7
|
+
|
|
8
|
+
SUPABASE_URL = "https://pidpadpudbcdldyvrlog.supabase.co"
|
|
9
|
+
SUPABASE_ANON_KEY = "sb_publishable_A2zwbfTFiOs63jz8313GXg_elw3F5dy"
|
|
10
|
+
TABLE = "computefence_runs"
|
|
11
|
+
OPT_OUT_FILE = Path.home() / ".computefence_no_telemetry"
|
|
12
|
+
|
|
13
|
+
|
|
14
|
+
def _is_opted_out():
|
|
15
|
+
return OPT_OUT_FILE.exists()
|
|
16
|
+
|
|
17
|
+
|
|
18
|
+
def _send(payload: dict):
|
|
19
|
+
try:
|
|
20
|
+
import requests
|
|
21
|
+
requests.post(
|
|
22
|
+
f"{SUPABASE_URL}/rest/v1/{TABLE}",
|
|
23
|
+
headers={
|
|
24
|
+
"apikey": SUPABASE_ANON_KEY,
|
|
25
|
+
"Authorization": f"Bearer {SUPABASE_ANON_KEY}",
|
|
26
|
+
"Content-Type": "application/json",
|
|
27
|
+
"Prefer": "return=minimal",
|
|
28
|
+
},
|
|
29
|
+
json=payload,
|
|
30
|
+
timeout=3,
|
|
31
|
+
)
|
|
32
|
+
except Exception:
|
|
33
|
+
pass
|
|
34
|
+
|
|
35
|
+
|
|
36
|
+
def record_run(results: list):
|
|
37
|
+
if _is_opted_out():
|
|
38
|
+
return
|
|
39
|
+
|
|
40
|
+
warning_count = sum(1 for r in results if r.get("status") == "warn")
|
|
41
|
+
blocker_count = sum(1 for r in results if r.get("status") == "fail")
|
|
42
|
+
pass_count = sum(1 for r in results if r.get("status") == "pass")
|
|
43
|
+
|
|
44
|
+
try:
|
|
45
|
+
import torch
|
|
46
|
+
gpu_count = torch.cuda.device_count()
|
|
47
|
+
except Exception:
|
|
48
|
+
gpu_count = 0
|
|
49
|
+
|
|
50
|
+
from computefence import __version__
|
|
51
|
+
|
|
52
|
+
payload = {
|
|
53
|
+
"version": __version__,
|
|
54
|
+
"python_version": f"{sys.version_info.major}.{sys.version_info.minor}",
|
|
55
|
+
"platform": platform.system(),
|
|
56
|
+
"gpu_count": gpu_count,
|
|
57
|
+
"warning_count": warning_count,
|
|
58
|
+
"blocker_count": blocker_count,
|
|
59
|
+
"pass_count": pass_count,
|
|
60
|
+
"opted_in": True,
|
|
61
|
+
}
|
|
62
|
+
|
|
63
|
+
thread = threading.Thread(target=_send, args=(payload,), daemon=True)
|
|
64
|
+
thread.start()
|
|
65
|
+
thread.join(timeout=4)
|
|
@@ -0,0 +1,110 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: computefence
|
|
3
|
+
Version: 0.2.3
|
|
4
|
+
Summary: Pre-flight validation for GPU training runs on rented infrastructure
|
|
5
|
+
License-Expression: MIT
|
|
6
|
+
Requires-Python: >=3.8
|
|
7
|
+
Description-Content-Type: text/markdown
|
|
8
|
+
License-File: LICENSE
|
|
9
|
+
Requires-Dist: rich>=13.0.0
|
|
10
|
+
Requires-Dist: pandas>=1.5.0
|
|
11
|
+
Requires-Dist: typer>=0.9.0
|
|
12
|
+
Requires-Dist: requests>=2.28.0
|
|
13
|
+
Dynamic: license-file
|
|
14
|
+
|
|
15
|
+
# ComputeFence
|
|
16
|
+
|
|
17
|
+
Pre-flight safety gate for ML training jobs on rented GPU compute. Built for RunPod, Vast.ai, Lambda Labs, and similar bare metal providers.
|
|
18
|
+
|
|
19
|
+
## Install
|
|
20
|
+
|
|
21
|
+
```bash
|
|
22
|
+
pip install computefence
|
|
23
|
+
```
|
|
24
|
+
|
|
25
|
+
Or with UV:
|
|
26
|
+
|
|
27
|
+
```bash
|
|
28
|
+
uvx computefence doctor
|
|
29
|
+
```
|
|
30
|
+
|
|
31
|
+
## Usage
|
|
32
|
+
|
|
33
|
+
```bash
|
|
34
|
+
computefence doctor
|
|
35
|
+
```
|
|
36
|
+
|
|
37
|
+
With a dataset:
|
|
38
|
+
|
|
39
|
+
```bash
|
|
40
|
+
computefence doctor --dataset train.csv --input-column text --label-column label
|
|
41
|
+
```
|
|
42
|
+
|
|
43
|
+
## Example output
|
|
44
|
+
|
|
45
|
+
ComputeFence v0.2.0 — Pre-flight diagnostic
|
|
46
|
+
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
|
|
47
|
+
2 WARNINGS · 0 BLOCKERS · 3 PASSED
|
|
48
|
+
|
|
49
|
+
Environment
|
|
50
|
+
✓ Python 3.11.4
|
|
51
|
+
✓ PyTorch 2.1.0 detected
|
|
52
|
+
✓ CUDA available — NVIDIA A40
|
|
53
|
+
|
|
54
|
+
Storage
|
|
55
|
+
⚠ HF_HOME is not set. HuggingFace will use default local cache.
|
|
56
|
+
Fix: export HF_HOME=/workspace/.cache/huggingface
|
|
57
|
+
⚠ Root disk (/) — 14.3 GB free of 460.4 GB (below 20 GB)
|
|
58
|
+
Fix: Free up disk space or attach a larger volume before launching
|
|
59
|
+
|
|
60
|
+
Dataset
|
|
61
|
+
✓ No dataset path provided — skipping dataset checks
|
|
62
|
+
|
|
63
|
+
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
|
|
64
|
+
2 warning(s) found. Review before launching.
|
|
65
|
+
|
|
66
|
+
|
|
67
|
+
## What it checks
|
|
68
|
+
|
|
69
|
+
- **CUDA and GPU visibility** — confirms PyTorch can see the GPU and training will not silently fall back to CPU
|
|
70
|
+
- **HuggingFace cache path** — confirms model files go to persistent storage not ephemeral disk
|
|
71
|
+
- **Accelerate GPU count** — confirms your distributed training config matches the GPUs actually on the instance
|
|
72
|
+
- **Disk space headroom** — checks free space on workspace volumes and root disk. Warns below 20 GB, blocks below 5 GB
|
|
73
|
+
- **Dataset integrity** — optional scan for duplicates, missing values, and conflicting labels
|
|
74
|
+
|
|
75
|
+
## What it does not check
|
|
76
|
+
|
|
77
|
+
- Training script correctness
|
|
78
|
+
- Model architecture compatibility
|
|
79
|
+
- Learning rate or hyperparameter safety
|
|
80
|
+
- Runtime monitoring during the job
|
|
81
|
+
- Slow dataloader or data pipeline throughput
|
|
82
|
+
- Dataloader bottleneck causing low GPU utilisation
|
|
83
|
+
|
|
84
|
+
## Why this exists
|
|
85
|
+
|
|
86
|
+
I burned approximately £1,000 on GPU training runs that failed silently. CUDA fell back to CPU with no error. 24 seconds per iteration instead of 0.4. Class weights caused loss collapse to 0.693 immediately. My dataset had 28,432 duplicate rows and 312 conflicting labels I only found after the run.
|
|
87
|
+
|
|
88
|
+
Nothing existed that caught these before the job started. So I built it.
|
|
89
|
+
|
|
90
|
+
## Real operator results
|
|
91
|
+
|
|
92
|
+
David at Neuralic ran ComputeFence on a RunPod A100. It caught HF_HOME writing to /root/.cache and an Accelerate GPU count mismatch. He fixed both before launch.
|
|
93
|
+
|
|
94
|
+
Shahzeb Ali, a computer vision engineer running client training jobs on RunPod, confirmed the storage warning matches real pod behaviour and would not have caught the HF_HOME issue explicitly without the tool.
|
|
95
|
+
|
|
96
|
+
## The problem it solves
|
|
97
|
+
|
|
98
|
+
A healthy GPU does not mean you are training the right job.
|
|
99
|
+
|
|
100
|
+
Docker makes environments reproducible. It does not check whether your HuggingFace cache is writing to ephemeral storage, whether your Accelerate config matches the GPUs actually on the instance, or whether your disk has enough headroom for checkpoints. ComputeFence addresses the job layer not the environment layer.
|
|
101
|
+
|
|
102
|
+
Five independent operators confirmed this independently. Docker does not solve job-specific configuration mistakes.
|
|
103
|
+
|
|
104
|
+
## GitHub
|
|
105
|
+
|
|
106
|
+
github.com/Francisco-Booth/ComputeFence
|
|
107
|
+
|
|
108
|
+
## PyPI
|
|
109
|
+
|
|
110
|
+
pypi.org/project/computefence
|
computefence-0.2.0/PKG-INFO
DELETED
|
@@ -1,49 +0,0 @@
|
|
|
1
|
-
Metadata-Version: 2.4
|
|
2
|
-
Name: computefence
|
|
3
|
-
Version: 0.2.0
|
|
4
|
-
Summary: Pre-flight validation for GPU training runs on rented infrastructure
|
|
5
|
-
License-Expression: MIT
|
|
6
|
-
Requires-Python: >=3.8
|
|
7
|
-
Description-Content-Type: text/markdown
|
|
8
|
-
License-File: LICENSE
|
|
9
|
-
Requires-Dist: rich>=13.0.0
|
|
10
|
-
Requires-Dist: pandas>=1.5.0
|
|
11
|
-
Requires-Dist: typer>=0.9.0
|
|
12
|
-
Requires-Dist: requests>=2.28.0
|
|
13
|
-
Dynamic: license-file
|
|
14
|
-
|
|
15
|
-
# ComputeFence
|
|
16
|
-
|
|
17
|
-
Pre-flight validation for GPU training runs on rented infrastructure.
|
|
18
|
-
Built specifically for RunPod, Vast.ai, Lambda, and similar providers.
|
|
19
|
-
|
|
20
|
-
## Install
|
|
21
|
-
|
|
22
|
-
pip install computefence
|
|
23
|
-
|
|
24
|
-
## Usage
|
|
25
|
-
|
|
26
|
-
computefence doctor
|
|
27
|
-
computefence doctor --dataset train.csv
|
|
28
|
-
|
|
29
|
-
## What it checks
|
|
30
|
-
|
|
31
|
-
- CUDA and GPU availability
|
|
32
|
-
- HuggingFace cache volume path (catches the RunPod /root vs /workspace conflict)
|
|
33
|
-
- Accelerate GPU count vs config
|
|
34
|
-
- Dataset duplicates and missing values
|
|
35
|
-
|
|
36
|
-
## What it does not yet check
|
|
37
|
-
|
|
38
|
-
- Training script correctness
|
|
39
|
-
- Model architecture compatibility
|
|
40
|
-
- Learning rate or hyperparameter safety
|
|
41
|
-
- Runtime monitoring during the job
|
|
42
|
-
|
|
43
|
-
## Why this exists
|
|
44
|
-
|
|
45
|
-
I burned ~£1,000 on GPU training runs that failed silently. CUDA fell back to
|
|
46
|
-
CPU with no error. Class weights caused loss collapse. My dataset had 28,432
|
|
47
|
-
duplicate rows and 312 conflicting labels I only found during the rebuild.
|
|
48
|
-
|
|
49
|
-
Nothing existed that caught these before the job started. So I built it.
|
computefence-0.2.0/README.md
DELETED
|
@@ -1,35 +0,0 @@
|
|
|
1
|
-
# ComputeFence
|
|
2
|
-
|
|
3
|
-
Pre-flight validation for GPU training runs on rented infrastructure.
|
|
4
|
-
Built specifically for RunPod, Vast.ai, Lambda, and similar providers.
|
|
5
|
-
|
|
6
|
-
## Install
|
|
7
|
-
|
|
8
|
-
pip install computefence
|
|
9
|
-
|
|
10
|
-
## Usage
|
|
11
|
-
|
|
12
|
-
computefence doctor
|
|
13
|
-
computefence doctor --dataset train.csv
|
|
14
|
-
|
|
15
|
-
## What it checks
|
|
16
|
-
|
|
17
|
-
- CUDA and GPU availability
|
|
18
|
-
- HuggingFace cache volume path (catches the RunPod /root vs /workspace conflict)
|
|
19
|
-
- Accelerate GPU count vs config
|
|
20
|
-
- Dataset duplicates and missing values
|
|
21
|
-
|
|
22
|
-
## What it does not yet check
|
|
23
|
-
|
|
24
|
-
- Training script correctness
|
|
25
|
-
- Model architecture compatibility
|
|
26
|
-
- Learning rate or hyperparameter safety
|
|
27
|
-
- Runtime monitoring during the job
|
|
28
|
-
|
|
29
|
-
## Why this exists
|
|
30
|
-
|
|
31
|
-
I burned ~£1,000 on GPU training runs that failed silently. CUDA fell back to
|
|
32
|
-
CPU with no error. Class weights caused loss collapse. My dataset had 28,432
|
|
33
|
-
duplicate rows and 312 conflicting labels I only found during the rebuild.
|
|
34
|
-
|
|
35
|
-
Nothing existed that caught these before the job started. So I built it.
|
|
@@ -1 +0,0 @@
|
|
|
1
|
-
__version__ = "0.2.0"
|
|
@@ -1,49 +0,0 @@
|
|
|
1
|
-
Metadata-Version: 2.4
|
|
2
|
-
Name: computefence
|
|
3
|
-
Version: 0.2.0
|
|
4
|
-
Summary: Pre-flight validation for GPU training runs on rented infrastructure
|
|
5
|
-
License-Expression: MIT
|
|
6
|
-
Requires-Python: >=3.8
|
|
7
|
-
Description-Content-Type: text/markdown
|
|
8
|
-
License-File: LICENSE
|
|
9
|
-
Requires-Dist: rich>=13.0.0
|
|
10
|
-
Requires-Dist: pandas>=1.5.0
|
|
11
|
-
Requires-Dist: typer>=0.9.0
|
|
12
|
-
Requires-Dist: requests>=2.28.0
|
|
13
|
-
Dynamic: license-file
|
|
14
|
-
|
|
15
|
-
# ComputeFence
|
|
16
|
-
|
|
17
|
-
Pre-flight validation for GPU training runs on rented infrastructure.
|
|
18
|
-
Built specifically for RunPod, Vast.ai, Lambda, and similar providers.
|
|
19
|
-
|
|
20
|
-
## Install
|
|
21
|
-
|
|
22
|
-
pip install computefence
|
|
23
|
-
|
|
24
|
-
## Usage
|
|
25
|
-
|
|
26
|
-
computefence doctor
|
|
27
|
-
computefence doctor --dataset train.csv
|
|
28
|
-
|
|
29
|
-
## What it checks
|
|
30
|
-
|
|
31
|
-
- CUDA and GPU availability
|
|
32
|
-
- HuggingFace cache volume path (catches the RunPod /root vs /workspace conflict)
|
|
33
|
-
- Accelerate GPU count vs config
|
|
34
|
-
- Dataset duplicates and missing values
|
|
35
|
-
|
|
36
|
-
## What it does not yet check
|
|
37
|
-
|
|
38
|
-
- Training script correctness
|
|
39
|
-
- Model architecture compatibility
|
|
40
|
-
- Learning rate or hyperparameter safety
|
|
41
|
-
- Runtime monitoring during the job
|
|
42
|
-
|
|
43
|
-
## Why this exists
|
|
44
|
-
|
|
45
|
-
I burned ~£1,000 on GPU training runs that failed silently. CUDA fell back to
|
|
46
|
-
CPU with no error. Class weights caused loss collapse. My dataset had 28,432
|
|
47
|
-
duplicate rows and 312 conflicting labels I only found during the rebuild.
|
|
48
|
-
|
|
49
|
-
Nothing existed that caught these before the job started. So I built it.
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|