trainjudge 0.2.1__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (152) hide show
  1. trainjudge-0.2.1/.claude-plugin/marketplace.json +14 -0
  2. trainjudge-0.2.1/.claude-plugin/plugin.json +9 -0
  3. trainjudge-0.2.1/.github/ISSUE_TEMPLATE/bug_report.yml +60 -0
  4. trainjudge-0.2.1/.github/ISSUE_TEMPLATE/config.yml +1 -0
  5. trainjudge-0.2.1/.github/ISSUE_TEMPLATE/domain_pack.yml +44 -0
  6. trainjudge-0.2.1/.github/ISSUE_TEMPLATE/feature_request.yml +25 -0
  7. trainjudge-0.2.1/.github/PULL_REQUEST_TEMPLATE.md +20 -0
  8. trainjudge-0.2.1/.github/dependabot.yml +11 -0
  9. trainjudge-0.2.1/.github/workflows/ci.yml +108 -0
  10. trainjudge-0.2.1/.github/workflows/release.yml +56 -0
  11. trainjudge-0.2.1/.gitignore +17 -0
  12. trainjudge-0.2.1/AGENTS.md +51 -0
  13. trainjudge-0.2.1/CHANGELOG.md +75 -0
  14. trainjudge-0.2.1/CODE_OF_CONDUCT.md +35 -0
  15. trainjudge-0.2.1/CONTRIBUTING.md +126 -0
  16. trainjudge-0.2.1/LICENSE +202 -0
  17. trainjudge-0.2.1/NOTICE +5 -0
  18. trainjudge-0.2.1/PKG-INFO +521 -0
  19. trainjudge-0.2.1/README.md +488 -0
  20. trainjudge-0.2.1/SECURITY.md +35 -0
  21. trainjudge-0.2.1/demo/README.md +37 -0
  22. trainjudge-0.2.1/demo/domains/bfsi/loan_faq/README.md +25 -0
  23. trainjudge-0.2.1/demo/domains/bfsi/loan_faq/data.jsonl +157 -0
  24. trainjudge-0.2.1/demo/domains/bfsi/loan_faq/docs/cards_and_charges.md +8 -0
  25. trainjudge-0.2.1/demo/domains/bfsi/loan_faq/docs/fixed_deposits.md +8 -0
  26. trainjudge-0.2.1/demo/domains/bfsi/loan_faq/docs/home_loans.md +8 -0
  27. trainjudge-0.2.1/demo/domains/bfsi/loan_faq/docs/kyc.md +8 -0
  28. trainjudge-0.2.1/demo/domains/bfsi/loan_faq/docs/personal_loans.md +8 -0
  29. trainjudge-0.2.1/demo/domains/bfsi/loan_faq/docs/savings_accounts.md +8 -0
  30. trainjudge-0.2.1/demo/domains/bfsi/loan_faq/generate.py +257 -0
  31. trainjudge-0.2.1/demo/domains/bfsi/transactions/README.md +35 -0
  32. trainjudge-0.2.1/demo/domains/bfsi/transactions/data.jsonl +656 -0
  33. trainjudge-0.2.1/demo/domains/bfsi/transactions/generate.py +262 -0
  34. trainjudge-0.2.1/demo/domains/customer_support/help_center_faq/README.md +19 -0
  35. trainjudge-0.2.1/demo/domains/customer_support/help_center_faq/data.jsonl +80 -0
  36. trainjudge-0.2.1/demo/domains/customer_support/help_center_faq/docs/data_policy.md +8 -0
  37. trainjudge-0.2.1/demo/domains/customer_support/help_center_faq/docs/plans.md +8 -0
  38. trainjudge-0.2.1/demo/domains/customer_support/help_center_faq/docs/support_policy.md +8 -0
  39. trainjudge-0.2.1/demo/domains/customer_support/ticket_triage/README.md +19 -0
  40. trainjudge-0.2.1/demo/domains/customer_support/ticket_triage/data.jsonl +666 -0
  41. trainjudge-0.2.1/demo/domains/ecommerce/catalog_faq/README.md +19 -0
  42. trainjudge-0.2.1/demo/domains/ecommerce/catalog_faq/data.jsonl +80 -0
  43. trainjudge-0.2.1/demo/domains/ecommerce/catalog_faq/docs/prices.md +8 -0
  44. trainjudge-0.2.1/demo/domains/ecommerce/catalog_faq/docs/promotions.md +8 -0
  45. trainjudge-0.2.1/demo/domains/ecommerce/catalog_faq/docs/stock.md +8 -0
  46. trainjudge-0.2.1/demo/domains/ecommerce/product_attributes/README.md +17 -0
  47. trainjudge-0.2.1/demo/domains/ecommerce/product_attributes/data.jsonl +666 -0
  48. trainjudge-0.2.1/demo/domains/education/course_faq/README.md +19 -0
  49. trainjudge-0.2.1/demo/domains/education/course_faq/data.jsonl +80 -0
  50. trainjudge-0.2.1/demo/domains/education/course_faq/docs/deadlines.md +8 -0
  51. trainjudge-0.2.1/demo/domains/education/course_faq/docs/exams.md +8 -0
  52. trainjudge-0.2.1/demo/domains/education/course_faq/docs/policies.md +8 -0
  53. trainjudge-0.2.1/demo/domains/education/question_tagging/README.md +17 -0
  54. trainjudge-0.2.1/demo/domains/education/question_tagging/data.jsonl +666 -0
  55. trainjudge-0.2.1/demo/domains/generate.py +590 -0
  56. trainjudge-0.2.1/demo/domains/healthcare/clinical_coding/README.md +19 -0
  57. trainjudge-0.2.1/demo/domains/healthcare/clinical_coding/data.jsonl +666 -0
  58. trainjudge-0.2.1/demo/domains/healthcare/formulary_faq/README.md +19 -0
  59. trainjudge-0.2.1/demo/domains/healthcare/formulary_faq/data.jsonl +80 -0
  60. trainjudge-0.2.1/demo/domains/healthcare/formulary_faq/docs/formulary_analgesics.md +8 -0
  61. trainjudge-0.2.1/demo/domains/healthcare/formulary_faq/docs/formulary_antibiotics.md +8 -0
  62. trainjudge-0.2.1/demo/domains/healthcare/formulary_faq/docs/formulary_policy.md +8 -0
  63. trainjudge-0.2.1/demo/domains/hr/benefits_faq/README.md +19 -0
  64. trainjudge-0.2.1/demo/domains/hr/benefits_faq/data.jsonl +80 -0
  65. trainjudge-0.2.1/demo/domains/hr/benefits_faq/docs/benefits.md +8 -0
  66. trainjudge-0.2.1/demo/domains/hr/benefits_faq/docs/leave.md +8 -0
  67. trainjudge-0.2.1/demo/domains/hr/benefits_faq/docs/pay.md +8 -0
  68. trainjudge-0.2.1/demo/domains/hr/resume_parsing/README.md +19 -0
  69. trainjudge-0.2.1/demo/domains/hr/resume_parsing/data.jsonl +666 -0
  70. trainjudge-0.2.1/demo/domains/legal/clause_extraction/README.md +19 -0
  71. trainjudge-0.2.1/demo/domains/legal/clause_extraction/data.jsonl +666 -0
  72. trainjudge-0.2.1/demo/domains/legal/statutes_faq/README.md +19 -0
  73. trainjudge-0.2.1/demo/domains/legal/statutes_faq/data.jsonl +80 -0
  74. trainjudge-0.2.1/demo/domains/legal/statutes_faq/docs/court_fees.md +8 -0
  75. trainjudge-0.2.1/demo/domains/legal/statutes_faq/docs/limitation_periods.md +8 -0
  76. trainjudge-0.2.1/demo/domains/legal/statutes_faq/docs/procedure.md +8 -0
  77. trainjudge-0.2.1/demo/policy_docs/README.md +25 -0
  78. trainjudge-0.2.1/demo/policy_docs/data.jsonl +236 -0
  79. trainjudge-0.2.1/demo/policy_docs/docs/price_matching.md +10 -0
  80. trainjudge-0.2.1/demo/policy_docs/docs/refund_policy.md +10 -0
  81. trainjudge-0.2.1/demo/policy_docs/docs/rewards_program.md +10 -0
  82. trainjudge-0.2.1/demo/policy_docs/docs/shipping_policy.md +10 -0
  83. trainjudge-0.2.1/demo/policy_docs/docs/support_policy.md +10 -0
  84. trainjudge-0.2.1/demo/policy_docs/docs/warranty_policy.md +10 -0
  85. trainjudge-0.2.1/demo/policy_docs/generate.py +326 -0
  86. trainjudge-0.2.1/demo/record-diagnose-demo.sh +31 -0
  87. trainjudge-0.2.1/demo/record-verify-demo.sh +33 -0
  88. trainjudge-0.2.1/demo/sql_generation/README.md +26 -0
  89. trainjudge-0.2.1/demo/sql_generation/data.jsonl +1830 -0
  90. trainjudge-0.2.1/demo/sql_generation/example-runs/improved-with-replay/EXPERIMENT_REPORT.md +115 -0
  91. trainjudge-0.2.1/demo/sql_generation/example-runs/improved-with-replay/MODEL_CARD.md +54 -0
  92. trainjudge-0.2.1/demo/sql_generation/example-runs/improved-with-replay/eval_results.json +60 -0
  93. trainjudge-0.2.1/demo/sql_generation/example-runs/regressed/EXPERIMENT_REPORT.md +128 -0
  94. trainjudge-0.2.1/demo/sql_generation/example-runs/regressed/MODEL_CARD.md +54 -0
  95. trainjudge-0.2.1/demo/sql_generation/example-runs/regressed/eval_results.json +62 -0
  96. trainjudge-0.2.1/demo/sql_generation/example-runs/rejected/EXPERIMENT_REPORT.md +120 -0
  97. trainjudge-0.2.1/demo/sql_generation/example-runs/rejected/MODEL_CARD.md +54 -0
  98. trainjudge-0.2.1/demo/sql_generation/example-runs/rejected/eval_results.json +60 -0
  99. trainjudge-0.2.1/demo/sql_generation/generate.py +439 -0
  100. trainjudge-0.2.1/demo/sql_generation/schema.sql +37 -0
  101. trainjudge-0.2.1/demo/sql_generation/shop.sql +3005 -0
  102. trainjudge-0.2.1/demo/trainjudge-diagnose-demo.gif +0 -0
  103. trainjudge-0.2.1/demo/trainjudge-verify-demo.gif +0 -0
  104. trainjudge-0.2.1/docs/assets/trainjudge-logo-dark.svg +8 -0
  105. trainjudge-0.2.1/docs/assets/trainjudge-logo-light.svg +8 -0
  106. trainjudge-0.2.1/notebooks/trainjudge_colab.ipynb +258 -0
  107. trainjudge-0.2.1/pyproject.toml +92 -0
  108. trainjudge-0.2.1/skills/trainjudge/SKILL.md +126 -0
  109. trainjudge-0.2.1/src/trainjudge/__init__.py +3 -0
  110. trainjudge-0.2.1/src/trainjudge/backend_base.py +160 -0
  111. trainjudge-0.2.1/src/trainjudge/backends.py +64 -0
  112. trainjudge-0.2.1/src/trainjudge/cli.py +481 -0
  113. trainjudge-0.2.1/src/trainjudge/dataset_audit.py +384 -0
  114. trainjudge-0.2.1/src/trainjudge/diagnosis.py +690 -0
  115. trainjudge-0.2.1/src/trainjudge/domains/__init__.py +109 -0
  116. trainjudge-0.2.1/src/trainjudge/domains/bfsi.py +103 -0
  117. trainjudge-0.2.1/src/trainjudge/domains/customer_support.py +76 -0
  118. trainjudge-0.2.1/src/trainjudge/domains/ecommerce.py +78 -0
  119. trainjudge-0.2.1/src/trainjudge/domains/education.py +77 -0
  120. trainjudge-0.2.1/src/trainjudge/domains/healthcare.py +89 -0
  121. trainjudge-0.2.1/src/trainjudge/domains/hr.py +83 -0
  122. trainjudge-0.2.1/src/trainjudge/domains/legal.py +79 -0
  123. trainjudge-0.2.1/src/trainjudge/eval_json.py +162 -0
  124. trainjudge-0.2.1/src/trainjudge/eval_sql.py +214 -0
  125. trainjudge-0.2.1/src/trainjudge/evaluation.py +266 -0
  126. trainjudge-0.2.1/src/trainjudge/mlx_backend.py +125 -0
  127. trainjudge-0.2.1/src/trainjudge/pii.py +294 -0
  128. trainjudge-0.2.1/src/trainjudge/regression_check.py +296 -0
  129. trainjudge-0.2.1/src/trainjudge/replay.py +83 -0
  130. trainjudge-0.2.1/src/trainjudge/reports.py +382 -0
  131. trainjudge-0.2.1/src/trainjudge/runs.py +150 -0
  132. trainjudge-0.2.1/src/trainjudge/status.py +345 -0
  133. trainjudge-0.2.1/src/trainjudge/tasks.py +87 -0
  134. trainjudge-0.2.1/src/trainjudge/textutil.py +28 -0
  135. trainjudge-0.2.1/src/trainjudge/torch_backend.py +173 -0
  136. trainjudge-0.2.1/src/trainjudge/torch_train.py +236 -0
  137. trainjudge-0.2.1/src/trainjudge/training.py +254 -0
  138. trainjudge-0.2.1/src/trainjudge/verdict.py +269 -0
  139. trainjudge-0.2.1/src/trainjudge/verification.py +138 -0
  140. trainjudge-0.2.1/tests/test_backends.py +85 -0
  141. trainjudge-0.2.1/tests/test_cli.py +17 -0
  142. trainjudge-0.2.1/tests/test_dataset_audit.py +235 -0
  143. trainjudge-0.2.1/tests/test_demo_datasets.py +87 -0
  144. trainjudge-0.2.1/tests/test_diagnosis.py +184 -0
  145. trainjudge-0.2.1/tests/test_domains.py +164 -0
  146. trainjudge-0.2.1/tests/test_eval_json.py +168 -0
  147. trainjudge-0.2.1/tests/test_eval_sql.py +227 -0
  148. trainjudge-0.2.1/tests/test_pii.py +176 -0
  149. trainjudge-0.2.1/tests/test_status.py +203 -0
  150. trainjudge-0.2.1/tests/test_torch_integration.py +74 -0
  151. trainjudge-0.2.1/tests/test_training.py +279 -0
  152. trainjudge-0.2.1/tests/test_verify.py +389 -0
@@ -0,0 +1,14 @@
1
+ {
2
+ "name": "trainjudge",
3
+ "description": "TrainJudge: decide whether to fine-tune, then verify it actually worked.",
4
+ "owner": {
5
+ "name": "Himanshu Kurrey"
6
+ },
7
+ "plugins": [
8
+ {
9
+ "name": "trainjudge",
10
+ "source": ".",
11
+ "description": "Decide whether fine-tuning is the right fix, then verify it actually worked — on task metrics, not training loss."
12
+ }
13
+ ]
14
+ }
@@ -0,0 +1,9 @@
1
+ {
2
+ "name": "trainjudge",
3
+ "description": "Decide whether fine-tuning is the right fix, then verify it actually worked — on task metrics, not training loss.",
4
+ "version": "0.2.1",
5
+ "author": {
6
+ "name": "Himanshu Kurrey"
7
+ },
8
+ "homepage": "https://github.com/Himanshukurrey/trainjudge"
9
+ }
@@ -0,0 +1,60 @@
1
+ name: Bug report
2
+ description: Something isn't working as expected
3
+ labels: ["bug"]
4
+ body:
5
+ - type: markdown
6
+ attributes:
7
+ value: |
8
+ Thanks for filing a bug. If the problem is with a training run, the fastest way to get it looked at is
9
+ to attach the run's `run.json` (and `eval_results.json` if `verify` got that far). They record the model,
10
+ settings, dataset hash and results, without the dataset itself. Check them for anything private before
11
+ attaching.
12
+
13
+ - type: textarea
14
+ id: what-happened
15
+ attributes:
16
+ label: What happened?
17
+ description: What did you expect, and what actually happened instead?
18
+ validations:
19
+ required: true
20
+
21
+ - type: input
22
+ id: command
23
+ attributes:
24
+ label: Exact command you ran
25
+ placeholder: "trainjudge verify trainjudge-runs/2026-09-24-sql_generation"
26
+ validations:
27
+ required: true
28
+
29
+ - type: textarea
30
+ id: run-files
31
+ attributes:
32
+ label: run.json / eval_results.json (if you have them)
33
+ description: Drag and drop the files here, or paste the relevant output.
34
+
35
+ - type: input
36
+ id: version
37
+ attributes:
38
+ label: TrainJudge version
39
+ description: Output of `trainjudge --version`
40
+ validations:
41
+ required: true
42
+
43
+ - type: input
44
+ id: mlx-version
45
+ attributes:
46
+ label: mlx-lm version (for train/eval/verify issues)
47
+ description: Output of `pip show mlx-lm | grep Version`
48
+
49
+ - type: dropdown
50
+ id: os
51
+ attributes:
52
+ label: Operating system
53
+ options:
54
+ - macOS (Apple Silicon)
55
+ - macOS (Intel)
56
+ - Linux
57
+ - Windows
58
+ - WSL
59
+ validations:
60
+ required: true
@@ -0,0 +1 @@
1
+ blank_issues_enabled: true
@@ -0,0 +1,44 @@
1
+ name: Domain pack request
2
+ description: Ask for TrainJudge checks tailored to an industry (healthcare, legal, retail, ...)
3
+ labels: ["domain-pack"]
4
+ body:
5
+ - type: input
6
+ id: domain
7
+ attributes:
8
+ label: Domain
9
+ placeholder: "Healthcare"
10
+ validations:
11
+ required: true
12
+
13
+ - type: textarea
14
+ id: tasks
15
+ attributes:
16
+ label: What do people fine-tune for in this domain?
17
+ description: Typical goals, e.g. "extract diagnosis codes from clinical notes" or "answer questions about drug dosing".
18
+ validations:
19
+ required: true
20
+
21
+ - type: textarea
22
+ id: changing-facts
23
+ attributes:
24
+ label: Facts that change
25
+ description: Information that goes stale and should live in retrieval, e.g. clinical guidelines or prices.
26
+
27
+ - type: textarea
28
+ id: high-stakes
29
+ attributes:
30
+ label: High-stakes decisions
31
+ description: Decisions about people that shouldn't be fully automated, e.g. triage or candidate screening.
32
+
33
+ - type: textarea
34
+ id: sensitive-data
35
+ attributes:
36
+ label: Sensitive identifiers
37
+ description: IDs or formats that appear in this domain's data and should be masked before training.
38
+
39
+ - type: checkboxes
40
+ id: contribute
41
+ attributes:
42
+ label: Contributing
43
+ options:
44
+ - label: I'd like to help build or review this pack
@@ -0,0 +1,25 @@
1
+ name: Feature request
2
+ description: Suggest an idea for TrainJudge
3
+ labels: ["enhancement"]
4
+ body:
5
+ - type: textarea
6
+ id: problem
7
+ attributes:
8
+ label: What problem would this solve?
9
+ description: What are you trying to do that TrainJudge can't help with today?
10
+ validations:
11
+ required: true
12
+
13
+ - type: textarea
14
+ id: proposal
15
+ attributes:
16
+ label: Proposed solution
17
+ description: What would you want TrainJudge to do instead?
18
+ validations:
19
+ required: true
20
+
21
+ - type: textarea
22
+ id: alternatives
23
+ attributes:
24
+ label: Alternatives considered
25
+ description: Any workarounds you're using today, or other tools that solve this.
@@ -0,0 +1,20 @@
1
+ ## What does this PR do?
2
+
3
+ <!-- One or two sentences describing the change. -->
4
+
5
+ ## Why?
6
+
7
+ <!-- What problem does this solve, or what issue does it close? Link it: "Closes #12" -->
8
+
9
+ ## How was this tested?
10
+
11
+ <!-- Commands you ran, and what you expected vs. saw. If you added tests, mention them here.
12
+ If the change affects training or evals, say whether you ran it on real hardware (and which Mac). -->
13
+
14
+ ## Checklist
15
+
16
+ - [ ] `pytest` passes locally
17
+ - [ ] `ruff check .` and `ruff format --check .` pass
18
+ - [ ] I added/updated tests for the behavior I changed (if applicable)
19
+ - [ ] I updated the README if this changes user-facing behavior
20
+ - [ ] If I changed a demo generator, I regenerated and committed its output
@@ -0,0 +1,11 @@
1
+ version: 2
2
+ updates:
3
+ - package-ecosystem: "pip"
4
+ directory: "/"
5
+ schedule:
6
+ interval: "weekly"
7
+
8
+ - package-ecosystem: "github-actions"
9
+ directory: "/"
10
+ schedule:
11
+ interval: "weekly"
@@ -0,0 +1,108 @@
1
+ name: CI
2
+
3
+ on:
4
+ push:
5
+ branches: [main]
6
+ pull_request:
7
+ workflow_dispatch:
8
+ schedule:
9
+ - cron: "0 6 * * 1" # weekly, Monday 06:00 UTC
10
+
11
+ jobs:
12
+ test:
13
+ runs-on: ${{ matrix.os }}
14
+ strategy:
15
+ fail-fast: false
16
+ matrix:
17
+ # Pushes to main run Linux only: macOS and Windows runner minutes are billed
18
+ # at 10x and 2x on private repositories. Pull requests, manual runs
19
+ # (`gh workflow run ci.yml`) and the weekly schedule run the full matrix.
20
+ os: ${{ github.event_name == 'push' && fromJSON('["ubuntu-latest"]') || fromJSON('["ubuntu-latest", "windows-latest", "macos-latest"]') }}
21
+ python-version: ["3.10", "3.12"]
22
+ steps:
23
+ - uses: actions/checkout@v7
24
+
25
+ - name: Set up Python ${{ matrix.python-version }}
26
+ uses: actions/setup-python@v7
27
+ with:
28
+ python-version: ${{ matrix.python-version }}
29
+
30
+ # The test suite never downloads models or needs mlx-lm: training and
31
+ # generation are exercised through fake backends, so CI runs everywhere.
32
+ - name: Install trainjudge and dev dependencies
33
+ run: pip install -e ".[dev]"
34
+
35
+ - name: Run tests
36
+ run: pytest -v
37
+
38
+ lint:
39
+ runs-on: ubuntu-latest
40
+ steps:
41
+ - uses: actions/checkout@v7
42
+
43
+ - name: Set up Python
44
+ uses: actions/setup-python@v7
45
+ with:
46
+ python-version: "3.12"
47
+
48
+ - name: Install trainjudge with dev tools
49
+ run: pip install -e ".[dev]"
50
+
51
+ - name: Run ruff check
52
+ run: ruff check .
53
+
54
+ - name: Run ruff format --check
55
+ run: ruff format --check .
56
+
57
+ - name: Type-check with mypy
58
+ run: mypy
59
+
60
+ demo-data:
61
+ # The committed demo datasets must match what their generators produce.
62
+ runs-on: ubuntu-latest
63
+ steps:
64
+ - uses: actions/checkout@v7
65
+
66
+ - uses: actions/setup-python@v7
67
+ with:
68
+ python-version: "3.12"
69
+
70
+ - name: Install trainjudge
71
+ run: pip install -e .
72
+
73
+ - name: Regenerate demo datasets
74
+ run: |
75
+ for demo in sql_generation policy_docs domains/bfsi/transactions domains/bfsi/loan_faq; do
76
+ python demo/$demo/generate.py
77
+ done
78
+ python demo/domains/generate.py
79
+
80
+ - name: Check nothing changed
81
+ run: git diff --exit-code -- demo/
82
+
83
+ torch-integration:
84
+ # Real LoRA training and generation with the PyTorch backend, on CPU with a
85
+ # tiny model, so a breaking transformers/peft release is caught here.
86
+ runs-on: ubuntu-latest
87
+ steps:
88
+ - uses: actions/checkout@v7
89
+
90
+ - uses: actions/setup-python@v7
91
+ with:
92
+ python-version: "3.12"
93
+
94
+ - name: Cache the tiny test model
95
+ uses: actions/cache@v4
96
+ with:
97
+ path: ~/.cache/huggingface
98
+ key: hf-smollm2-135m-instruct
99
+
100
+ - name: Install CPU-only PyTorch and trainjudge
101
+ run: |
102
+ pip install torch --index-url https://download.pytorch.org/whl/cpu
103
+ pip install -e ".[dev,cuda]"
104
+
105
+ - name: Train and generate on CPU
106
+ env:
107
+ TRAINJUDGE_TORCH_INTEGRATION: "1"
108
+ run: pytest -v tests/test_torch_integration.py
@@ -0,0 +1,56 @@
1
+ name: Publish to PyPI
2
+
3
+ # Fires when a GitHub Release is published (gh release create / the "Publish
4
+ # release" button), NOT on every tag push, so tagging alone can't trigger an
5
+ # accidental publish. Uses PyPI Trusted Publishing (OIDC): no API token is
6
+ # stored as a secret. One-time setup is required on PyPI before this works;
7
+ # see "Releasing" in CONTRIBUTING.md.
8
+
9
+ on:
10
+ release:
11
+ types: [published]
12
+
13
+ jobs:
14
+ build:
15
+ runs-on: ubuntu-latest
16
+ steps:
17
+ - uses: actions/checkout@v7
18
+
19
+ - uses: actions/setup-python@v7
20
+ with:
21
+ python-version: "3.12"
22
+
23
+ - name: Install build tooling
24
+ run: pip install build
25
+
26
+ - name: Build sdist and wheel
27
+ run: python -m build
28
+
29
+ - name: Sanity-check the built wheel actually installs and runs
30
+ run: |
31
+ python -m venv /tmp/trainjudge-release-check
32
+ /tmp/trainjudge-release-check/bin/pip install dist/*.whl
33
+ /tmp/trainjudge-release-check/bin/trainjudge --version
34
+ /tmp/trainjudge-release-check/bin/trainjudge audit demo/sql_generation/data.jsonl > /dev/null
35
+
36
+ - uses: actions/upload-artifact@v7
37
+ with:
38
+ name: dist
39
+ path: dist/
40
+
41
+ publish:
42
+ needs: build
43
+ runs-on: ubuntu-latest
44
+ environment:
45
+ name: pypi
46
+ url: https://pypi.org/project/trainjudge/
47
+ permissions:
48
+ id-token: write # required for PyPI Trusted Publishing (OIDC)
49
+ steps:
50
+ - uses: actions/download-artifact@v8
51
+ with:
52
+ name: dist
53
+ path: dist/
54
+
55
+ - name: Publish to PyPI
56
+ uses: pypa/gh-action-pypi-publish@release/v1
@@ -0,0 +1,17 @@
1
+ __pycache__/
2
+ *.py[cod]
3
+ *.egg-info/
4
+ .venv/
5
+ dist/
6
+ build/
7
+ .pytest_cache/
8
+ .ruff_cache/
9
+
10
+ # Training outputs / model weights
11
+ trainjudge-runs/
12
+ adapters/
13
+ *.safetensors
14
+ .DS_Store
15
+
16
+ # Local Claude Code settings
17
+ .claude/settings.local.json
@@ -0,0 +1,51 @@
1
+ # AGENTS.md
2
+
3
+ Instructions for coding agents (Codex and others that read `AGENTS.md`).
4
+
5
+ ## Using TrainJudge
6
+
7
+ When the user asks to fine-tune, LoRA-train or distill a model, or asks whether
8
+ a fine-tuned model is better than its base model, use the `trainjudge` CLI.
9
+ Install it with `pip install "trainjudge[mlx] @ git+https://github.com/Himanshukurrey/trainjudge"`
10
+ on an Apple Silicon Mac, or `[cuda]` on Windows/Linux with an NVIDIA GPU (install PyTorch with
11
+ CUDA first). `diagnose` and `audit` run anywhere; with no GPU, use
12
+ `notebooks/trainjudge_colab.ipynb` on Colab.
13
+
14
+ 1. **Diagnose before training.**
15
+ `trainjudge diagnose --dataset <path> --model <model> --goal "<goal>"`.
16
+ If the result is a KNOWLEDGE GAP (use retrieval) or a PROMPT-ENGINEERING GAP
17
+ (fix the prompt), tell the user TrainJudge recommends not fine-tuning and
18
+ only continue if they explicitly ask to. For a FORMAT/BEHAVIOR or
19
+ COST/LATENCY gap, ask before training, because it takes a while.
20
+ 2. **Respect sensitive-data refusals.** If `train` refuses because the audit
21
+ found card numbers, Aadhaar, PAN or similar identifiers, tell the user and
22
+ offer `--mask-sensitive`, which replaces them with placeholders like `[EMAIL]`
23
+ and trains on the masked rows. Never add `--allow-sensitive-data` unless the
24
+ user asks for it.
25
+ 3. **Train.** `trainjudge train --dataset <path> --model <model> --goal "<goal>"`.
26
+ Add `--replay 200` when general skills matter.
27
+ 4. **Verify.** `trainjudge verify <run-dir>` (add `--db <database>` for SQL tasks).
28
+ 5. **Report the verdict as-is.** IMPROVED means it worked. REGRESSED means the
29
+ task improved but general capability broke, so don't call it ready to
30
+ deploy. REJECTED means it didn't really improve, whatever the loss did.
31
+ Never claim improvement from training loss alone.
32
+
33
+ **Long jobs: keep the user informed.** `train` and `verify` take minutes, and
34
+ command output often isn't shown live. Run them in the background, then run
35
+ `trainjudge status <run-dir> --watch --milestones`, which prints one line per
36
+ milestone (stage started, 25/50/75%, stage done, finished or failed) and exits
37
+ when the job ends. Relay each milestone to the user with what's next and the
38
+ ETA, and report failures right away. Never go quiet until the result is in.
39
+
40
+ Full details: [skills/trainjudge/SKILL.md](skills/trainjudge/SKILL.md).
41
+
42
+ ## Working on this repository
43
+
44
+ - Setup: `pip install -e ".[dev]"` (add `mlx` on an Apple Silicon Mac, or `cuda` for the PyTorch backend).
45
+ - Before finishing a change, run `pytest`, `ruff check .`, `ruff format --check .` and `mypy`.
46
+ - Tests must not download models or need `mlx-lm` or `torch`, except
47
+ `tests/test_torch_integration.py` (opt-in, its own CI job). Use the fake backends in
48
+ `tests/test_training.py` and `tests/test_verify.py`.
49
+ - `demo/*/data.jsonl`, `demo/*/docs/` and `demo/sql_generation/shop.sql` are
50
+ generated. Edit the matching `generate.py`, rerun it and commit both.
51
+ - Always pass `encoding="utf-8"` when reading or writing text files; CI runs on Windows.
@@ -0,0 +1,75 @@
1
+ # Changelog
2
+
3
+ ## 0.2.1 (2026-09-26)
4
+
5
+ ### Fixes
6
+
7
+ - **GPU memory is released after every generation pass** (PyTorch backend). PyTorch's caching
8
+ allocator kept the model's memory reserved after generating replay examples, so the training
9
+ subprocess ran out of memory loading its own copy; it also built up across `verify`'s eval
10
+ passes. Found on a Colab T4 with SmolLM2-1.7B, which now trains and verifies there.
11
+
12
+ ### Internal
13
+
14
+ - Shared text helpers (`textutil.py`) replace duplicated think-block and duration code.
15
+ - `cli.py` only parses options and prints; watching, summaries and the accuracy line live with
16
+ their modules.
17
+ - mypy type-checks `src/` in CI.
18
+
19
+ ## 0.2.0 (2026-09-25)
20
+
21
+ ### Train and verify on NVIDIA GPUs, not just Macs
22
+
23
+ - **PyTorch backend** (`--backend torch`): LoRA with transformers + peft on NVIDIA GPUs via CUDA
24
+ (Windows and Linux), Apple GPUs via MPS, or CPU. `--backend auto` keeps MLX on Apple Silicon and
25
+ uses PyTorch elsewhere. Both backends render prompts the same way, train with the loss on
26
+ completions only and report the same progress. Verified end to end on an NVIDIA T4.
27
+ - **Colab notebook** (`notebooks/trainjudge_colab.ipynb`): diagnose, mask, train and verify on a
28
+ free GPU, for anyone without one.
29
+
30
+ ### Every domain, not just one
31
+
32
+ - **Domain packs** for healthcare, legal, e-commerce/retail, customer support, HR/recruiting, BFSI
33
+ and education: facts that change (pointing to retrieval), high-stakes decision warnings and
34
+ domain notes. `--domain` forces or disables one.
35
+ - **Demos for every pack**: a fine-tune case and a retrieval case per domain, 600 rows each for
36
+ the fine-tune cases so a real gain can be statistically significant.
37
+
38
+ ### Evals and verdicts
39
+
40
+ - **JSON extraction eval**: exact match plus field-level and per-field accuracy. `eval` and
41
+ `verify` detect SQL or JSON from the test split (`--task` overrides), so every domain's
42
+ fine-tune demo can be verified.
43
+ - **Baseline cache**: base-model evals are reused across runs with the same model, test split,
44
+ decoding and database. `verify --rerun` bypasses it.
45
+ - Reports include a complete reproduce command and the device training actually used.
46
+
47
+ ### Data safety
48
+
49
+ - **PII masking**: `train --mask-sensitive` and `audit --write-clean --mask-sensitive` replace
50
+ identifiers with placeholders (`[EMAIL]`, `[MRN]`, `[CARD]`…).
51
+ - **Region-aware detection**: US SSNs, UK National Insurance numbers, IBANs (mod-97),
52
+ international and US phone numbers, labelled medical record numbers and dates of birth, on top
53
+ of cards, Aadhaar, PAN, UPI IDs and emails. Compliance pointers follow what was found.
54
+
55
+ ### Progress for people and agents
56
+
57
+ - `trainjudge status --watch --milestones` prints one line per milestone, so agents can relay
58
+ progress even where command output isn't streamed (such as the Claude Code VS Code extension).
59
+
60
+ ### Fixes
61
+
62
+ - Label-like answers are split per row, so no label is missing from training; paraphrase-style
63
+ answers stay grouped.
64
+ - Status files survive Windows file locking; cleaned datasets always use LF line endings.
65
+
66
+ ### Other
67
+
68
+ - Relicensed from MIT to Apache-2.0.
69
+
70
+ ## 0.1.0 (2026-09-24)
71
+
72
+ First release: `diagnose` (knowledge / format / cost / prompt gap), `audit` (duplicates,
73
+ malformed, low-quality and sensitive rows), `train` (MLX LoRA with held-out splits and
74
+ `--replay`), `eval` (SQL execution accuracy), `verify` (IMPROVED / REGRESSED / REJECTED with a
75
+ regression suite and reports), `status`, and a Claude Code plugin plus `AGENTS.md`.
@@ -0,0 +1,35 @@
1
+ # Contributor Covenant Code of Conduct
2
+
3
+ ## Our Pledge
4
+
5
+ We as members, contributors, and maintainers pledge to make participation in our project a harassment-free experience for everyone, regardless of age, body size, visible or invisible disability, ethnicity, sex characteristics, gender identity and expression, level of experience, education, socio-economic status, nationality, personal appearance, race, religion, or sexual identity and orientation.
6
+
7
+ ## Our Standards
8
+
9
+ Examples of behavior that contributes to a positive environment:
10
+
11
+ - Being respectful of differing opinions, viewpoints, and experiences
12
+ - Giving and gracefully accepting constructive feedback
13
+ - Focusing on what is best for the community
14
+
15
+ Examples of unacceptable behavior:
16
+
17
+ - Trolling, insulting or derogatory comments, and personal or political attacks
18
+ - Public or private harassment
19
+ - Publishing others' private information without explicit permission
20
+
21
+ ## Enforcement Responsibilities
22
+
23
+ The maintainer is responsible for clarifying and enforcing standards of acceptable behavior and will take appropriate, fair corrective action in response to unacceptable behavior.
24
+
25
+ ## Scope
26
+
27
+ This Code of Conduct applies within all project spaces (issues, pull requests, discussions) and when an individual is officially representing the project elsewhere.
28
+
29
+ ## Enforcement
30
+
31
+ Instances of unacceptable behavior may be reported by opening an issue or, for sensitive matters, via [GitHub's private vulnerability reporting](https://github.com/Himanshukurrey/trainjudge/security/advisories/new) as a stand-in private contact channel. All complaints will be reviewed and investigated promptly and fairly.
32
+
33
+ ## Attribution
34
+
35
+ This Code of Conduct is adapted from the [Contributor Covenant](https://www.contributor-covenant.org/), version 2.1.
@@ -0,0 +1,126 @@
1
+ # Contributing to TrainJudge
2
+
3
+ Thanks for considering a contribution. Here's the workflow.
4
+
5
+ ## How to contribute
6
+
7
+ 1. **Fork the repo** and clone your fork locally.
8
+ 2. **Create a branch** for your change: `git checkout -b fix/short-description`.
9
+ 3. **Make your change**, with tests if you're changing behavior.
10
+ 4. **Run the test suite and linter** locally before opening a PR:
11
+ ```bash
12
+ pip install -e ".[dev]"
13
+ pytest
14
+ ruff check .
15
+ ruff format --check .
16
+ mypy
17
+ ```
18
+ CI runs the same checks on Linux, Windows and macOS, on Python 3.10 and 3.12. It also checks that the
19
+ committed demo datasets match their generators. Running these locally first saves a round trip.
20
+ 5. **Open a pull request** against `main`. Fill in the PR template: what the change does, why, and how you
21
+ tested it.
22
+ 6. A maintainer will review it. Please be patient; this is currently maintained part-time.
23
+
24
+ ## Tests don't need a GPU
25
+
26
+ The test suite never downloads a model or needs `mlx-lm` or `torch`. Training and generation go through fake backends
27
+ (see `fake_backend` in `tests/test_training.py` and `good_model` in `tests/test_verify.py`), so every test runs
28
+ on any OS in a few seconds. If you change something that only shows up with a real model, such as prompt
29
+ rendering, generation or backend flags, also run it on real hardware (an Apple Silicon Mac for MLX, or an
30
+ NVIDIA GPU, Colab included, for PyTorch) and say so in the PR:
31
+
32
+ ```bash
33
+ pip install -e ".[dev,mlx]" # or .[dev,cuda] with --backend torch
34
+ trainjudge train --dataset demo/sql_generation/data.jsonl --model Qwen3-0.6B --iters 20
35
+ trainjudge verify trainjudge-runs/<run> --db demo/sql_generation/shop.sql
36
+ ```
37
+
38
+ ## Adding a domain pack
39
+
40
+ Domain packs add industry-specific checks to `diagnose` (see the README's "Domain packs"
41
+ section). To add one, say for healthcare:
42
+
43
+ 1. Create `src/trainjudge/domains/healthcare.py` defining `PACK = DomainPack(...)`, using
44
+ `domains/bfsi.py` as the template:
45
+ - `terms`: words that identify the domain in a goal or dataset. Prefer specific terms;
46
+ generic words like "claim" appear in many domains.
47
+ - `changing_fact_terms`, `changing_facts` and `changing_fact_note`: facts that change
48
+ and belong in retrieval.
49
+ - `high_stakes_terms` and `high_stakes_note`: decisions about people that shouldn't be
50
+ fully automated.
51
+ - `closing_notes`: anything else worth saying for this domain.
52
+ 2. Register it in `_load_packs()` in `src/trainjudge/domains/__init__.py`.
53
+ 3. Add tests showing it's detected from a goal, isn't detected on the other demos, and
54
+ produces its notes.
55
+ 4. Add a `FinetuneCase` and a `RetrievalCase` for the domain in `demo/domains/generate.py`,
56
+ run it, and add both demos to the table in `tests/test_demo_datasets.py`.
57
+
58
+ Keep sensitive-data detectors in `src/trainjudge/pii.py` rather than in a pack, so every
59
+ dataset is scanned whatever its domain. Compliance pointers should name the rule, not
60
+ interpret it.
61
+
62
+ The PyTorch backend also has a real integration test that trains a tiny model on CPU (it downloads
63
+ about 270 MB the first time):
64
+
65
+ ```bash
66
+ pip install -e ".[dev,cuda]"
67
+ TRAINJUDGE_TORCH_INTEGRATION=1 pytest tests/test_torch_integration.py
68
+ ```
69
+
70
+ ## Demo datasets
71
+
72
+ The files in `demo/*/` are generated, not hand-written. To change one, edit its `generate.py`, run it, and
73
+ commit both the script and its output. Tests pin the exact number of injected duplicate, low-quality,
74
+ malformed and PII rows, so update those counts too if you change them.
75
+
76
+ ## Why PRs go through review
77
+
78
+ Every change lands through a reviewed pull request, so the project has a consistent, auditable history and a
79
+ second pair of eyes on every change.
80
+
81
+ ## What makes a good PR
82
+
83
+ - **Small and focused.** One logical change per PR is much easier to review than five unrelated fixes bundled
84
+ together.
85
+ - **Tested.** If you fixed a bug, a regression test that would have caught it is the strongest evidence the
86
+ fix is real.
87
+ - **Explained.** A one-line "fixed the bug" isn't enough. Say what was broken and why your change addresses
88
+ it.
89
+
90
+ ## Reporting bugs
91
+
92
+ Open an issue describing what you expected vs. what happened, your OS, Python and backend (`mlx-lm` or `torch`) versions, and the
93
+ exact command you ran. For problems with a run, attaching its `run.json` (and `eval_results.json`, if it got
94
+ that far) is the fastest way to get it looked at. Check them for anything private first.
95
+
96
+ ## License of contributions
97
+
98
+ TrainJudge is licensed under the [Apache License 2.0](LICENSE). By submitting a pull request, you
99
+ agree that your contribution is licensed under the same terms (section 5 of the license).
100
+
101
+ ## Code of conduct
102
+
103
+ Be respectful. Disagreements about code are fine; personal attacks aren't. See
104
+ [CODE_OF_CONDUCT.md](CODE_OF_CONDUCT.md).
105
+
106
+ ## Releasing (maintainer only)
107
+
108
+ 1. Bump `version` in `pyproject.toml`, `src/trainjudge/__init__.py`, `.claude-plugin/plugin.json` and the
109
+ version badge at the top of `README.md`, and add the release to `CHANGELOG.md`.
110
+ 2. `git tag vX.Y.Z && git push origin vX.Y.Z`
111
+ 3. `gh release create vX.Y.Z --generate-notes`
112
+
113
+ Publishing that release triggers `.github/workflows/release.yml`, which builds the package, checks that the
114
+ wheel installs and `trainjudge --version` and `trainjudge audit` run, and publishes to PyPI via
115
+ [Trusted Publishing](https://docs.pypi.org/trusted-publishers/) (no stored API token).
116
+
117
+ **One-time setup required before the first release**, done once on pypi.org by whoever owns the PyPI project:
118
+ add a "pending publisher" for a project named `trainjudge` under Account Settings → Publishing, with:
119
+
120
+ - Owner: `Himanshukurrey`
121
+ - Repository: `trainjudge`
122
+ - Workflow name: `release.yml`
123
+ - Environment name: `pypi`
124
+
125
+ This reserves the trust relationship before the package exists on PyPI, so the first `gh release create` can
126
+ publish without a manual upload.