computefence 0.2.0__tar.gz → 0.2.1__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,91 @@
1
+ Metadata-Version: 2.4
2
+ Name: computefence
3
+ Version: 0.2.1
4
+ Summary: Pre-flight validation for GPU training runs on rented infrastructure
5
+ License-Expression: MIT
6
+ Requires-Python: >=3.8
7
+ Description-Content-Type: text/markdown
8
+ License-File: LICENSE
9
+ Requires-Dist: rich>=13.0.0
10
+ Requires-Dist: pandas>=1.5.0
11
+ Requires-Dist: typer>=0.9.0
12
+ Requires-Dist: requests>=2.28.0
13
+ Dynamic: license-file
14
+
15
+ # ComputeFence
16
+
17
+ Pre-flight safety gate for ML training jobs on rented GPU compute. Built specifically for RunPod, Vast.ai, Lambda Labs, and similar bare metal providers.
18
+
19
+ ## Install
20
+
21
+ ```bash
22
+ pip install computefence
23
+ ```
24
+
25
+ ## Usage
26
+
27
+ ```bash
28
+ computefence doctor
29
+ ```
30
+
31
+ With a dataset:
32
+
33
+ ```bash
34
+ computefence doctor --dataset train.csv --input-column text --label-column label
35
+ ```
36
+
37
+ ## Example output
38
+
39
+ ComputeFence v0.1.2 — Pre-flight diagnostic
40
+ ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
41
+ 2 WARNINGS · 0 BLOCKERS · 3 PASSED
42
+
43
+ Environment
44
+ ✓ Python 3.11.4
45
+ ✓ PyTorch 2.1.0 detected
46
+ ✓ CUDA available — NVIDIA A40
47
+
48
+ Storage
49
+ ⚠ HF_HOME is not set. HuggingFace will use default local cache.
50
+ Fix: Set HF_HOME to a persistent volume path before training
51
+ ⚠ No mounted persistent volumes detected at /workspace, /runpod-volume, or /vast
52
+ Fix: Attach a network volume before training to persist checkpoints and cache
53
+
54
+ Dataset
55
+ ✓ No dataset path provided — skipping dataset checks
56
+
57
+ ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
58
+ 2 warning(s) found. Review before launching.
59
+
60
+ ## What it checks
61
+
62
+ - **CUDA and GPU visibility** — confirms PyTorch can see the GPU and training will not silently fall back to CPU
63
+ - **HuggingFace cache path** — confirms model files go to persistent storage not ephemeral disk
64
+ - **Accelerate GPU count** — confirms your distributed training config matches the GPUs actually on the instance
65
+ - **Dataset integrity** — optional scan for duplicates, missing values, and conflicting labels
66
+
67
+ ## What it does not check
68
+
69
+ - Training script correctness
70
+ - Model architecture compatibility
71
+ - Learning rate or hyperparameter safety
72
+ - Runtime monitoring during the job
73
+ - Slow dataloader or data pipeline throughput
74
+
75
+ ## Why this exists
76
+
77
+ I burned approximately £1,000 on GPU training runs that failed silently. CUDA fell back to CPU with no error. Class weights caused loss collapse to 0.693 immediately. My dataset had 28,432 duplicate rows and 312 conflicting labels I only found after the run.
78
+
79
+ Nothing existed that caught these before the job started. So I built it.
80
+
81
+ ## Real operator results
82
+
83
+ David at Neuralic ran ComputeFence on a RunPod A100. It caught HF_HOME writing to /root/.cache and an Accelerate GPU count mismatch. He fixed both before launch.
84
+
85
+ ## GitHub
86
+
87
+ github.com/Francisco-Booth/ComputeFence
88
+
89
+ ## PyPI
90
+
91
+ pypi.org/project/computefence
@@ -0,0 +1,77 @@
1
+ # ComputeFence
2
+
3
+ Pre-flight safety gate for ML training jobs on rented GPU compute. Built specifically for RunPod, Vast.ai, Lambda Labs, and similar bare metal providers.
4
+
5
+ ## Install
6
+
7
+ ```bash
8
+ pip install computefence
9
+ ```
10
+
11
+ ## Usage
12
+
13
+ ```bash
14
+ computefence doctor
15
+ ```
16
+
17
+ With a dataset:
18
+
19
+ ```bash
20
+ computefence doctor --dataset train.csv --input-column text --label-column label
21
+ ```
22
+
23
+ ## Example output
24
+
25
+ ComputeFence v0.1.2 — Pre-flight diagnostic
26
+ ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
27
+ 2 WARNINGS · 0 BLOCKERS · 3 PASSED
28
+
29
+ Environment
30
+ ✓ Python 3.11.4
31
+ ✓ PyTorch 2.1.0 detected
32
+ ✓ CUDA available — NVIDIA A40
33
+
34
+ Storage
35
+ ⚠ HF_HOME is not set. HuggingFace will use default local cache.
36
+ Fix: Set HF_HOME to a persistent volume path before training
37
+ ⚠ No mounted persistent volumes detected at /workspace, /runpod-volume, or /vast
38
+ Fix: Attach a network volume before training to persist checkpoints and cache
39
+
40
+ Dataset
41
+ ✓ No dataset path provided — skipping dataset checks
42
+
43
+ ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
44
+ 2 warning(s) found. Review before launching.
45
+
46
+ ## What it checks
47
+
48
+ - **CUDA and GPU visibility** — confirms PyTorch can see the GPU and training will not silently fall back to CPU
49
+ - **HuggingFace cache path** — confirms model files go to persistent storage not ephemeral disk
50
+ - **Accelerate GPU count** — confirms your distributed training config matches the GPUs actually on the instance
51
+ - **Dataset integrity** — optional scan for duplicates, missing values, and conflicting labels
52
+
53
+ ## What it does not check
54
+
55
+ - Training script correctness
56
+ - Model architecture compatibility
57
+ - Learning rate or hyperparameter safety
58
+ - Runtime monitoring during the job
59
+ - Slow dataloader or data pipeline throughput
60
+
61
+ ## Why this exists
62
+
63
+ I burned approximately £1,000 on GPU training runs that failed silently. CUDA fell back to CPU with no error. Class weights caused loss collapse to 0.693 immediately. My dataset had 28,432 duplicate rows and 312 conflicting labels I only found after the run.
64
+
65
+ Nothing existed that caught these before the job started. So I built it.
66
+
67
+ ## Real operator results
68
+
69
+ David at Neuralic ran ComputeFence on a RunPod A100. It caught HF_HOME writing to /root/.cache and an Accelerate GPU count mismatch. He fixed both before launch.
70
+
71
+ ## GitHub
72
+
73
+ github.com/Francisco-Booth/ComputeFence
74
+
75
+ ## PyPI
76
+
77
+ pypi.org/project/computefence
@@ -0,0 +1 @@
1
+ __version__ = "0.2.1"
@@ -1,19 +1,17 @@
1
1
  import sys
2
2
  from rich.console import Console
3
-
4
3
  from computefence import __version__
5
4
  from computefence.checks.environment import check_environment
6
5
  from computefence.checks.storage import check_disk_headroom, check_storage
7
6
  from computefence.checks.dataset import check_dataset
7
+ from computefence.telemetry import record_run
8
8
 
9
9
  console = Console()
10
10
 
11
-
12
11
  def print_result(result):
13
12
  status = result.get("status")
14
13
  message = result.get("message")
15
14
  fix = result.get("fix")
16
-
17
15
  if status == "pass":
18
16
  console.print(f" [green]✓[/green] {message}")
19
17
  elif status == "warn":
@@ -25,12 +23,10 @@ def print_result(result):
25
23
  if fix:
26
24
  console.print(f" [dim]Fix: {fix}[/dim]")
27
25
 
28
-
29
26
  def doctor():
30
27
  dataset = None
31
28
  input_column = None
32
29
  label_column = None
33
-
34
30
  args = sys.argv[1:]
35
31
  i = 0
36
32
  while i < len(args):
@@ -50,8 +46,6 @@ def doctor():
50
46
  console.print(f"[bold]ComputeFence v{__version__} — Pre-flight diagnostic[/bold]")
51
47
  console.print("━" * 50)
52
48
 
53
- # Run every check up front so the verdict line can be counted and printed
54
- # above the individual results.
55
49
  env_results = check_environment()
56
50
  storage_results = check_storage() + check_disk_headroom()
57
51
  dataset_results = check_dataset(
@@ -61,6 +55,9 @@ def doctor():
61
55
  )
62
56
 
63
57
  all_results = env_results + storage_results + dataset_results
58
+
59
+ record_run(all_results)
60
+
64
61
  failures = [r for r in all_results if r["status"] == "fail"]
65
62
  warnings = [r for r in all_results if r["status"] == "warn"]
66
63
  passed = [r for r in all_results if r["status"] == "pass"]
@@ -71,35 +68,32 @@ def doctor():
71
68
  f"[red]{len(failures)} BLOCKERS[/red] · "
72
69
  f"[green]{len(passed)} PASSED[/green]"
73
70
  )
74
-
75
71
  console.print()
76
72
  console.print("[bold blue]Environment[/bold blue]")
77
73
  for result in env_results:
78
74
  print_result(result)
79
-
80
75
  console.print()
81
76
  console.print("[bold blue]Storage[/bold blue]")
82
77
  for result in storage_results:
83
78
  print_result(result)
84
-
85
79
  console.print()
86
80
  console.print("[bold blue]Dataset[/bold blue]")
87
81
  for result in dataset_results:
88
82
  print_result(result)
89
-
90
83
  console.print()
91
84
  console.print("━" * 50)
92
-
93
85
  if failures:
94
86
  console.print("[red]" + str(len(failures)) + " error(s) and " + str(len(warnings)) + " warning(s) found. Fix errors before launching.[/red]")
95
87
  elif warnings:
96
88
  console.print("[yellow]" + str(len(warnings)) + " warning(s) found. Review before launching.[/yellow]")
97
89
  else:
98
90
  console.print("[green]All checks passed. Safe to launch.[/green]")
99
-
91
+ console.print(
92
+ "[dim]Anonymous usage stats help improve ComputeFence. "
93
+ "Opt out: touch ~/.computefence_no_telemetry[/dim]"
94
+ )
100
95
  console.print()
101
96
 
102
-
103
97
  def main():
104
98
  args = sys.argv[1:]
105
99
  if not args or args[0] == "doctor":
@@ -108,6 +102,5 @@ def main():
108
102
  console.print(f"[red]Unknown command: {args[0]}[/red]")
109
103
  console.print("Usage: computefence doctor")
110
104
 
111
-
112
105
  if __name__ == "__main__":
113
106
  main()
@@ -0,0 +1,64 @@
1
+ import hashlib
2
+ import platform
3
+ import sys
4
+ import threading
5
+ import uuid
6
+ from pathlib import Path
7
+
8
+ SUPABASE_URL = "https://pidpadpudbcdldyvrlog.supabase.co"
9
+ SUPABASE_ANON_KEY = "sb_publishable_A2zwbfTFiOs63jz8313GXg_elw3F5dy"
10
+ TABLE = "computefence_runs"
11
+ OPT_OUT_FILE = Path.home() / ".computefence_no_telemetry"
12
+
13
+
14
+ def _is_opted_out():
15
+ return OPT_OUT_FILE.exists()
16
+
17
+
18
+ def _send(payload: dict):
19
+ try:
20
+ import requests
21
+ requests.post(
22
+ f"{SUPABASE_URL}/rest/v1/{TABLE}",
23
+ headers={
24
+ "apikey": SUPABASE_ANON_KEY,
25
+ "Authorization": f"Bearer {SUPABASE_ANON_KEY}",
26
+ "Content-Type": "application/json",
27
+ "Prefer": "return=minimal",
28
+ },
29
+ json=payload,
30
+ timeout=3,
31
+ )
32
+ except Exception:
33
+ pass
34
+
35
+
36
+ def record_run(results: list):
37
+ if _is_opted_out():
38
+ return
39
+
40
+ warning_count = sum(1 for r in results if r.get("status") == "warn")
41
+ blocker_count = sum(1 for r in results if r.get("status") == "fail")
42
+ pass_count = sum(1 for r in results if r.get("status") == "pass")
43
+
44
+ try:
45
+ import torch
46
+ gpu_count = torch.cuda.device_count()
47
+ except Exception:
48
+ gpu_count = 0
49
+
50
+ from computefence import __version__
51
+
52
+ payload = {
53
+ "version": __version__,
54
+ "python_version": f"{sys.version_info.major}.{sys.version_info.minor}",
55
+ "platform": platform.system(),
56
+ "gpu_count": gpu_count,
57
+ "warning_count": warning_count,
58
+ "blocker_count": blocker_count,
59
+ "pass_count": pass_count,
60
+ "opted_in": True,
61
+ }
62
+
63
+ thread = threading.Thread(target=_send, args=(payload,), daemon=True)
64
+ thread.start()
@@ -0,0 +1,91 @@
1
+ Metadata-Version: 2.4
2
+ Name: computefence
3
+ Version: 0.2.1
4
+ Summary: Pre-flight validation for GPU training runs on rented infrastructure
5
+ License-Expression: MIT
6
+ Requires-Python: >=3.8
7
+ Description-Content-Type: text/markdown
8
+ License-File: LICENSE
9
+ Requires-Dist: rich>=13.0.0
10
+ Requires-Dist: pandas>=1.5.0
11
+ Requires-Dist: typer>=0.9.0
12
+ Requires-Dist: requests>=2.28.0
13
+ Dynamic: license-file
14
+
15
+ # ComputeFence
16
+
17
+ Pre-flight safety gate for ML training jobs on rented GPU compute. Built specifically for RunPod, Vast.ai, Lambda Labs, and similar bare metal providers.
18
+
19
+ ## Install
20
+
21
+ ```bash
22
+ pip install computefence
23
+ ```
24
+
25
+ ## Usage
26
+
27
+ ```bash
28
+ computefence doctor
29
+ ```
30
+
31
+ With a dataset:
32
+
33
+ ```bash
34
+ computefence doctor --dataset train.csv --input-column text --label-column label
35
+ ```
36
+
37
+ ## Example output
38
+
39
+ ComputeFence v0.1.2 — Pre-flight diagnostic
40
+ ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
41
+ 2 WARNINGS · 0 BLOCKERS · 3 PASSED
42
+
43
+ Environment
44
+ ✓ Python 3.11.4
45
+ ✓ PyTorch 2.1.0 detected
46
+ ✓ CUDA available — NVIDIA A40
47
+
48
+ Storage
49
+ ⚠ HF_HOME is not set. HuggingFace will use default local cache.
50
+ Fix: Set HF_HOME to a persistent volume path before training
51
+ ⚠ No mounted persistent volumes detected at /workspace, /runpod-volume, or /vast
52
+ Fix: Attach a network volume before training to persist checkpoints and cache
53
+
54
+ Dataset
55
+ ✓ No dataset path provided — skipping dataset checks
56
+
57
+ ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
58
+ 2 warning(s) found. Review before launching.
59
+
60
+ ## What it checks
61
+
62
+ - **CUDA and GPU visibility** — confirms PyTorch can see the GPU and training will not silently fall back to CPU
63
+ - **HuggingFace cache path** — confirms model files go to persistent storage not ephemeral disk
64
+ - **Accelerate GPU count** — confirms your distributed training config matches the GPUs actually on the instance
65
+ - **Dataset integrity** — optional scan for duplicates, missing values, and conflicting labels
66
+
67
+ ## What it does not check
68
+
69
+ - Training script correctness
70
+ - Model architecture compatibility
71
+ - Learning rate or hyperparameter safety
72
+ - Runtime monitoring during the job
73
+ - Slow dataloader or data pipeline throughput
74
+
75
+ ## Why this exists
76
+
77
+ I burned approximately £1,000 on GPU training runs that failed silently. CUDA fell back to CPU with no error. Class weights caused loss collapse to 0.693 immediately. My dataset had 28,432 duplicate rows and 312 conflicting labels I only found after the run.
78
+
79
+ Nothing existed that caught these before the job started. So I built it.
80
+
81
+ ## Real operator results
82
+
83
+ David at Neuralic ran ComputeFence on a RunPod A100. It caught HF_HOME writing to /root/.cache and an Accelerate GPU count mismatch. He fixed both before launch.
84
+
85
+ ## GitHub
86
+
87
+ github.com/Francisco-Booth/ComputeFence
88
+
89
+ ## PyPI
90
+
91
+ pypi.org/project/computefence
@@ -3,6 +3,7 @@ README.md
3
3
  pyproject.toml
4
4
  computefence/__init__.py
5
5
  computefence/cli.py
6
+ computefence/telemetry.py
6
7
  computefence.egg-info/PKG-INFO
7
8
  computefence.egg-info/SOURCES.txt
8
9
  computefence.egg-info/dependency_links.txt
@@ -1,49 +0,0 @@
1
- Metadata-Version: 2.4
2
- Name: computefence
3
- Version: 0.2.0
4
- Summary: Pre-flight validation for GPU training runs on rented infrastructure
5
- License-Expression: MIT
6
- Requires-Python: >=3.8
7
- Description-Content-Type: text/markdown
8
- License-File: LICENSE
9
- Requires-Dist: rich>=13.0.0
10
- Requires-Dist: pandas>=1.5.0
11
- Requires-Dist: typer>=0.9.0
12
- Requires-Dist: requests>=2.28.0
13
- Dynamic: license-file
14
-
15
- # ComputeFence
16
-
17
- Pre-flight validation for GPU training runs on rented infrastructure.
18
- Built specifically for RunPod, Vast.ai, Lambda, and similar providers.
19
-
20
- ## Install
21
-
22
- pip install computefence
23
-
24
- ## Usage
25
-
26
- computefence doctor
27
- computefence doctor --dataset train.csv
28
-
29
- ## What it checks
30
-
31
- - CUDA and GPU availability
32
- - HuggingFace cache volume path (catches the RunPod /root vs /workspace conflict)
33
- - Accelerate GPU count vs config
34
- - Dataset duplicates and missing values
35
-
36
- ## What it does not yet check
37
-
38
- - Training script correctness
39
- - Model architecture compatibility
40
- - Learning rate or hyperparameter safety
41
- - Runtime monitoring during the job
42
-
43
- ## Why this exists
44
-
45
- I burned ~£1,000 on GPU training runs that failed silently. CUDA fell back to
46
- CPU with no error. Class weights caused loss collapse. My dataset had 28,432
47
- duplicate rows and 312 conflicting labels I only found during the rebuild.
48
-
49
- Nothing existed that caught these before the job started. So I built it.
@@ -1,35 +0,0 @@
1
- # ComputeFence
2
-
3
- Pre-flight validation for GPU training runs on rented infrastructure.
4
- Built specifically for RunPod, Vast.ai, Lambda, and similar providers.
5
-
6
- ## Install
7
-
8
- pip install computefence
9
-
10
- ## Usage
11
-
12
- computefence doctor
13
- computefence doctor --dataset train.csv
14
-
15
- ## What it checks
16
-
17
- - CUDA and GPU availability
18
- - HuggingFace cache volume path (catches the RunPod /root vs /workspace conflict)
19
- - Accelerate GPU count vs config
20
- - Dataset duplicates and missing values
21
-
22
- ## What it does not yet check
23
-
24
- - Training script correctness
25
- - Model architecture compatibility
26
- - Learning rate or hyperparameter safety
27
- - Runtime monitoring during the job
28
-
29
- ## Why this exists
30
-
31
- I burned ~£1,000 on GPU training runs that failed silently. CUDA fell back to
32
- CPU with no error. Class weights caused loss collapse. My dataset had 28,432
33
- duplicate rows and 312 conflicting labels I only found during the rebuild.
34
-
35
- Nothing existed that caught these before the job started. So I built it.
@@ -1 +0,0 @@
1
- __version__ = "0.2.0"
@@ -1,49 +0,0 @@
1
- Metadata-Version: 2.4
2
- Name: computefence
3
- Version: 0.2.0
4
- Summary: Pre-flight validation for GPU training runs on rented infrastructure
5
- License-Expression: MIT
6
- Requires-Python: >=3.8
7
- Description-Content-Type: text/markdown
8
- License-File: LICENSE
9
- Requires-Dist: rich>=13.0.0
10
- Requires-Dist: pandas>=1.5.0
11
- Requires-Dist: typer>=0.9.0
12
- Requires-Dist: requests>=2.28.0
13
- Dynamic: license-file
14
-
15
- # ComputeFence
16
-
17
- Pre-flight validation for GPU training runs on rented infrastructure.
18
- Built specifically for RunPod, Vast.ai, Lambda, and similar providers.
19
-
20
- ## Install
21
-
22
- pip install computefence
23
-
24
- ## Usage
25
-
26
- computefence doctor
27
- computefence doctor --dataset train.csv
28
-
29
- ## What it checks
30
-
31
- - CUDA and GPU availability
32
- - HuggingFace cache volume path (catches the RunPod /root vs /workspace conflict)
33
- - Accelerate GPU count vs config
34
- - Dataset duplicates and missing values
35
-
36
- ## What it does not yet check
37
-
38
- - Training script correctness
39
- - Model architecture compatibility
40
- - Learning rate or hyperparameter safety
41
- - Runtime monitoring during the job
42
-
43
- ## Why this exists
44
-
45
- I burned ~£1,000 on GPU training runs that failed silently. CUDA fell back to
46
- CPU with no error. Class weights caused loss collapse. My dataset had 28,432
47
- duplicate rows and 312 conflicting labels I only found during the rebuild.
48
-
49
- Nothing existed that caught these before the job started. So I built it.
File without changes
File without changes