gq-local 0.1.1__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (36) hide show
  1. gq_local-0.1.1/.github/ISSUE_TEMPLATE/bug_report.yml +58 -0
  2. gq_local-0.1.1/.github/ISSUE_TEMPLATE/config.yml +5 -0
  3. gq_local-0.1.1/.github/ISSUE_TEMPLATE/feature_request.yml +29 -0
  4. gq_local-0.1.1/.github/pull_request_template.md +18 -0
  5. gq_local-0.1.1/.github/workflows/ci.yml +63 -0
  6. gq_local-0.1.1/.github/workflows/release.yml +50 -0
  7. gq_local-0.1.1/.gitignore +10 -0
  8. gq_local-0.1.1/CHANGELOG.md +54 -0
  9. gq_local-0.1.1/CONTRIBUTING.md +123 -0
  10. gq_local-0.1.1/LICENSE +21 -0
  11. gq_local-0.1.1/PKG-INFO +179 -0
  12. gq_local-0.1.1/README.md +164 -0
  13. gq_local-0.1.1/SECURITY.md +72 -0
  14. gq_local-0.1.1/docs/architecture.md +104 -0
  15. gq_local-0.1.1/docs/troubleshooting.md +142 -0
  16. gq_local-0.1.1/docs/usage.md +207 -0
  17. gq_local-0.1.1/pyproject.toml +66 -0
  18. gq_local-0.1.1/src/gq/__init__.py +3 -0
  19. gq_local-0.1.1/src/gq/allocator.py +71 -0
  20. gq_local-0.1.1/src/gq/cli.py +550 -0
  21. gq_local-0.1.1/src/gq/daemon.py +287 -0
  22. gq_local-0.1.1/src/gq/database.py +324 -0
  23. gq_local-0.1.1/src/gq/models.py +128 -0
  24. gq_local-0.1.1/src/gq/nvml.py +89 -0
  25. gq_local-0.1.1/src/gq/paths.py +50 -0
  26. gq_local-0.1.1/src/gq/processes.py +97 -0
  27. gq_local-0.1.1/src/gq/protocol.py +48 -0
  28. gq_local-0.1.1/src/gq/scheduler.py +446 -0
  29. gq_local-0.1.1/tests/conftest.py +33 -0
  30. gq_local-0.1.1/tests/integration/test_nvml.py +70 -0
  31. gq_local-0.1.1/tests/test_allocator.py +70 -0
  32. gq_local-0.1.1/tests/test_cli.py +71 -0
  33. gq_local-0.1.1/tests/test_database.py +125 -0
  34. gq_local-0.1.1/tests/test_recovery.py +84 -0
  35. gq_local-0.1.1/tests/test_scheduler.py +209 -0
  36. gq_local-0.1.1/uv.lock +408 -0
@@ -0,0 +1,58 @@
1
+ name: Bug report
2
+ description: Something in gq behaves incorrectly
3
+ labels: ["bug"]
4
+ body:
5
+ - type: markdown
6
+ attributes:
7
+ value: |
8
+ Before filing, please check
9
+ [docs/troubleshooting.md](https://github.com/BirddWang/gq/blob/main/docs/troubleshooting.md)
10
+ — GPUs stuck in `EXTERNAL`/`UNKNOWN`, jobs stuck in `WAITING`, and daemon
11
+ startup failures are covered there.
12
+
13
+ **Do not attach `gq.sqlite3`.** It stores the environment of every job.
14
+ Skim the daemon log before pasting it: it records job commands and paths.
15
+ - type: textarea
16
+ id: what-happened
17
+ attributes:
18
+ label: What happened
19
+ description: What did you expect, and what happened instead?
20
+ validations:
21
+ required: true
22
+ - type: textarea
23
+ id: reproduce
24
+ attributes:
25
+ label: Steps to reproduce
26
+ placeholder: |
27
+ 1. gq -g 1 ...
28
+ 2. gq ps
29
+ 3. ...
30
+ validations:
31
+ required: true
32
+ - type: input
33
+ id: version
34
+ attributes:
35
+ label: gq version
36
+ description: Output of `gq --version`
37
+ validations:
38
+ required: true
39
+ - type: input
40
+ id: environment
41
+ attributes:
42
+ label: Environment
43
+ description: Linux distribution, Python version, NVIDIA driver version, and whether you run in a container
44
+ placeholder: "Ubuntu 24.04, Python 3.12, driver 560.35, inside Docker"
45
+ validations:
46
+ required: true
47
+ - type: textarea
48
+ id: state
49
+ attributes:
50
+ label: Scheduler state
51
+ description: Output of `gq gpu` and `gq ps`, plus `gq show JOB` if one job is involved
52
+ render: text
53
+ - type: textarea
54
+ id: daemon-log
55
+ attributes:
56
+ label: Daemon log
57
+ description: Relevant lines from `tail -n 100 ~/.local/share/gq/gq-daemon.log`
58
+ render: text
@@ -0,0 +1,5 @@
1
+ blank_issues_enabled: true
2
+ contact_links:
3
+ - name: Security vulnerability
4
+ url: https://github.com/BirddWang/gq/security/advisories/new
5
+ about: Please report vulnerabilities privately, not as a public issue.
@@ -0,0 +1,29 @@
1
+ name: Feature request
2
+ description: Suggest something gq should do
3
+ labels: ["enhancement"]
4
+ body:
5
+ - type: markdown
6
+ attributes:
7
+ value: |
8
+ `gq` is deliberately small, and some things are permanent non-goals:
9
+ multi-node scheduling, CPU/RAM allocation, MIG, and fractional GPU sharing.
10
+ See [CONTRIBUTING.md](https://github.com/BirddWang/gq/blob/main/CONTRIBUTING.md#scope).
11
+
12
+ Please open an issue before writing the code, so a rejected PR does not
13
+ waste your time.
14
+ - type: textarea
15
+ id: problem
16
+ attributes:
17
+ label: The problem
18
+ description: What are you trying to do that gq makes hard today? Describe the situation, not the solution.
19
+ validations:
20
+ required: true
21
+ - type: textarea
22
+ id: proposal
23
+ attributes:
24
+ label: What you have in mind
25
+ description: If you have a specific interface in mind, sketch it.
26
+ - type: textarea
27
+ id: alternatives
28
+ attributes:
29
+ label: What you do instead today
@@ -0,0 +1,18 @@
1
+ ## What this changes
2
+
3
+ <!-- And why. Link the issue it addresses, if there is one. -->
4
+
5
+ ## Checklist
6
+
7
+ - [ ] `uv run pytest` passes
8
+ - [ ] `uv run ruff check .` and `uv run mypy` pass
9
+ - [ ] Tests cover the change
10
+ - [ ] `CHANGELOG.md` updated under `## [Unreleased]`, if behavior changed
11
+ - [ ] `docs/` updated, if the interface changed
12
+ - [ ] Schema change appends a migration and bumps `SCHEMA_VERSION` (see CONTRIBUTING.md)
13
+ - [ ] IPC change bumps `PROTOCOL_VERSION`, if it is not backward compatible
14
+
15
+ ## Allocation safety
16
+
17
+ <!-- Delete if this does not touch allocation, cancellation, or recovery.
18
+ Otherwise: which invariant does this rely on, and what test shows it holds? -->
@@ -0,0 +1,63 @@
1
+ name: CI
2
+
3
+ on:
4
+ push:
5
+ branches: [main]
6
+ pull_request:
7
+ workflow_dispatch:
8
+
9
+ concurrency:
10
+ group: ci-${{ github.ref }}
11
+ cancel-in-progress: true
12
+
13
+ jobs:
14
+ test:
15
+ name: pytest (Python ${{ matrix.python-version }})
16
+ runs-on: ubuntu-latest
17
+ strategy:
18
+ fail-fast: false
19
+ matrix:
20
+ python-version: ["3.11", "3.12", "3.13"]
21
+ steps:
22
+ - uses: actions/checkout@v4
23
+ - uses: astral-sh/setup-uv@v5
24
+ with:
25
+ enable-cache: true
26
+ - name: Install
27
+ run: uv sync --all-groups --python ${{ matrix.python-version }}
28
+ - name: Test
29
+ # The suite runs against a fake NVML provider, so no GPU is needed. The
30
+ # hardware tests under tests/integration stay opt-in via GQ_RUN_GPU_TESTS.
31
+ run: uv run pytest
32
+
33
+ lint:
34
+ name: ruff and mypy
35
+ runs-on: ubuntu-latest
36
+ steps:
37
+ - uses: actions/checkout@v4
38
+ - uses: astral-sh/setup-uv@v5
39
+ with:
40
+ enable-cache: true
41
+ - name: Install
42
+ run: uv sync --all-groups
43
+ - name: Ruff
44
+ run: uv run ruff check .
45
+ - name: Ruff format check
46
+ run: uv run ruff format --check .
47
+ - name: Mypy
48
+ run: uv run mypy
49
+
50
+ build:
51
+ name: build distributions
52
+ runs-on: ubuntu-latest
53
+ steps:
54
+ - uses: actions/checkout@v4
55
+ - uses: astral-sh/setup-uv@v5
56
+ - name: Build
57
+ run: uv build
58
+ - name: Check metadata
59
+ run: uvx twine check dist/*
60
+ - uses: actions/upload-artifact@v4
61
+ with:
62
+ name: distributions
63
+ path: dist/
@@ -0,0 +1,50 @@
1
+ name: Release
2
+
3
+ # Publishes to PyPI when a version tag is pushed. Authentication uses PyPI trusted
4
+ # publishing (OIDC), so no API token is stored in the repository.
5
+ #
6
+ # One-time setup at https://pypi.org/manage/account/publishing/ :
7
+ # PyPI project name: gq-local
8
+ # Owner: BirddWang
9
+ # Repository: gq
10
+ # Workflow: release.yml
11
+ # Environment: pypi
12
+
13
+ on:
14
+ push:
15
+ tags: ["v*"]
16
+ workflow_dispatch:
17
+
18
+ jobs:
19
+ build:
20
+ runs-on: ubuntu-latest
21
+ steps:
22
+ - uses: actions/checkout@v4
23
+ - uses: astral-sh/setup-uv@v5
24
+ - run: uv build
25
+ - name: Fail if the tag and the package version disagree
26
+ run: |
27
+ version="$(uv run python -c 'import gq; print(gq.__version__)')"
28
+ tag="${GITHUB_REF_NAME#v}"
29
+ if [ "$version" != "$tag" ]; then
30
+ echo "tag $tag does not match gq.__version__ $version" >&2
31
+ exit 1
32
+ fi
33
+ if: startsWith(github.ref, 'refs/tags/')
34
+ - uses: actions/upload-artifact@v4
35
+ with:
36
+ name: distributions
37
+ path: dist/
38
+
39
+ publish:
40
+ needs: build
41
+ runs-on: ubuntu-latest
42
+ environment: pypi
43
+ permissions:
44
+ id-token: write
45
+ steps:
46
+ - uses: actions/download-artifact@v4
47
+ with:
48
+ name: distributions
49
+ path: dist/
50
+ - uses: pypa/gh-action-pypi-publish@release/v1
@@ -0,0 +1,10 @@
1
+ .venv/
2
+ .pytest_cache/
3
+ __pycache__/
4
+ *.py[cod]
5
+ dist/
6
+ build/
7
+ *.egg-info/
8
+
9
+ .ruff_cache/
10
+ .mypy_cache/
@@ -0,0 +1,54 @@
1
+ # Changelog
2
+
3
+ All notable changes to this project are documented here. This project follows
4
+ [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
5
+
6
+ ## [0.1.1] - 2026-09-10
7
+
8
+ A hardening release. No new scheduling behavior; everything here came from defects
9
+ and rough edges observed in the first week of real use.
10
+
11
+ ### Fixed
12
+
13
+ - The daemon no longer logs a traceback every time a client hangs up mid-response.
14
+ `gq ps | head` and Ctrl-C out of `gq logs -f` are ordinary events, but each one
15
+ raised `ConnectionResetError`/`BrokenPipeError` out of the connection callback.
16
+ - Running as root inside a container is now allowed without `GQ_ALLOW_ROOT=1`.
17
+ Docker and Jupyter images commonly have uid 0 as the only account, and the guard
18
+ made `gq` unusable there. It still refuses on what looks like a regular host, where
19
+ the warning is genuinely warranted.
20
+ - `PRAGMA foreign_keys = ON` is now issued as its own statement rather than inside
21
+ the schema script, so `ON DELETE CASCADE` on `job_gpus` cannot silently be a no-op.
22
+
23
+ ### Added
24
+
25
+ - **Schema migrations.** The database carries a `PRAGMA user_version` and applies
26
+ ordered migrations, each in one transaction with its version bump. Databases
27
+ created before this release are adopted in place, not rebuilt.
28
+ - **`gq rm JOB...`** deletes terminal jobs and their logs. Active jobs are refused,
29
+ and one bad id rejects the whole batch rather than partially destroying history.
30
+ - **`gq clean --older-than AGE [--state STATE] [-y]`** sweeps old terminal jobs and
31
+ their logs. Defaults to 30d and confirms before deleting.
32
+ - **`--json`** on `gq ps`, `gq gpu`, and `gq show`.
33
+ - **Protocol versioning.** Every request carries a protocol number; a mismatched
34
+ daemon reports how to restart instead of failing on a missing field. `gq daemon
35
+ status` warns when the running daemon predates the installed client.
36
+
37
+ ### Changed
38
+
39
+ - **Secret-looking environment variables are no longer persisted by default.** The
40
+ submission environment is stored verbatim in SQLite so queued jobs survive a daemon
41
+ restart, which made every captured credential a durable one. Variables matching
42
+ `TOKEN`, `SECRET`, `PASSWORD`, `CREDENTIAL`, `*_KEY`, `API_KEY` or the `AWS_`,
43
+ `AZURE_`, `GCP_`, `GOOGLE_APPLICATION_` prefixes are dropped, and the dropped names
44
+ are printed at submit time rather than being removed silently.
45
+
46
+ This is a **behavior change that can break jobs** which need a credential at
47
+ runtime — `HF_TOKEN` for a private Hugging Face dataset, for instance. Restore the
48
+ old behavior per-job with `--env-all`, globally with `GQ_ENV_ALL=1`, or for one
49
+ variable with `--env-keep HF_TOKEN`.
50
+
51
+ ## [0.1.0] - 2026-08-31
52
+
53
+ Initial MVP: queueing, atomic NVML-checked GPU allocation, `CUDA_VISIBLE_DEVICES`
54
+ assignment, detached execution, cancellation, and crash/reboot recovery.
@@ -0,0 +1,123 @@
1
+ # Contributing to gq
2
+
3
+ Thanks for taking a look. `gq` is deliberately small, and the bar for adding to it is
4
+ correspondingly high — please read "Scope" before starting on a feature.
5
+
6
+ ## Getting set up
7
+
8
+ ```bash
9
+ git clone https://github.com/BirddWang/gq
10
+ cd gq
11
+ uv sync --all-groups
12
+ uv run pytest
13
+ ```
14
+
15
+ Most of the suite runs against a fake NVML provider (`tests/conftest.py`), so you do
16
+ **not** need an NVIDIA GPU to develop or to run CI. If you do have one, the hardware
17
+ discovery tests are opt-in:
18
+
19
+ ```bash
20
+ GQ_RUN_GPU_TESTS=1 uv run pytest tests/integration
21
+ ```
22
+
23
+ Before opening a pull request:
24
+
25
+ ```bash
26
+ uv run ruff check .
27
+ uv run ruff format .
28
+ uv run mypy
29
+ uv run pytest
30
+ ```
31
+
32
+ ### Testing against a scratch instance
33
+
34
+ Never develop against your real queue. All state is redirectable, so run an isolated
35
+ daemon instead:
36
+
37
+ ```bash
38
+ export GQ_DATA_DIR=/tmp/gq-dev/data
39
+ export GQ_RUNTIME_DIR=/tmp/gq-dev/run
40
+ export GQ_RECONCILE_INTERVAL=0.5
41
+ export GQ_CANCEL_GRACE_SECONDS=2
42
+ gq daemon start
43
+ ```
44
+
45
+ A separate `GQ_RUNTIME_DIR` gets its own lock and socket, so it will not collide with
46
+ a daemon you already have running.
47
+
48
+ ## Scope
49
+
50
+ The MVP intentionally omits multi-node scheduling, CPU/RAM allocation, preemption,
51
+ MIG, containers, and fractional GPU sharing. Some of these are on the roadmap and
52
+ some are permanent non-goals. Please open an issue to discuss a feature before
53
+ writing it — a rejected PR is a waste of your time, and this project would rather
54
+ stay small than grow every reasonable idea.
55
+
56
+ ## The one rule that matters
57
+
58
+ > A GPU assigned to a managed job is never simultaneously allocatable.
59
+
60
+ Everything in `allocator.py`, `scheduler.py`, and `database.py` exists to hold that
61
+ line, and the design **fails closed**: when ownership is ambiguous, the GPU becomes
62
+ `UNKNOWN` and no one gets it. Three consequences to keep in mind when changing
63
+ allocation code:
64
+
65
+ - **Utilization is never an ownership signal.** A Jupyter kernel sitting at 0% still
66
+ owns its CUDA context. Only NVML compute-process presence and `gq`'s own durable
67
+ reservations decide ownership.
68
+ - **The reservation and the state transition are one transaction.** `Database.reserve`
69
+ moves `WAITING → STARTING` and writes the full GPU assignment inside a single
70
+ `BEGIN IMMEDIATE`. Do not introduce an `await` or a second allocation between them.
71
+ - **Process identity is checked before signalling.** Use the helpers in
72
+ `processes.py`; never signal a bare PID. A recycled PID must not be mistaken for
73
+ the original job.
74
+
75
+ If a change would weaken any of these, it needs a very good argument and a test that
76
+ demonstrates the new invariant holds.
77
+
78
+ ## Compatibility
79
+
80
+ - **Database.** Never edit the baseline schema to change an existing table. Append a
81
+ migration in `MIGRATIONS` in `database.py`, keyed by the version it produces, and
82
+ bump `SCHEMA_VERSION`. Users have real job history; a release that discards it is a
83
+ bug. Test your migration against a copy of a populated database.
84
+ - **IPC.** Adding a request type is backward compatible. Changing or removing the
85
+ meaning of an existing field is not — bump `PROTOCOL_VERSION` in `protocol.py`.
86
+ - **`--json` output** is the supported scripting interface. Adding fields is fine;
87
+ renaming or removing them requires a protocol bump. Table layouts are for humans
88
+ and may change freely.
89
+
90
+ ## Style
91
+
92
+ Match the surrounding code. It is plain, typed Python with no framework: standard
93
+ library plus `nvidia-ml-py`, and `from __future__ import annotations` everywhere.
94
+
95
+ **Adding a runtime dependency requires a strong justification.** Being installable
96
+ with one small dependency is a feature of this project, not an accident.
97
+
98
+ Comments explain *why*, especially where the reasoning is subtle — race conditions,
99
+ PID reuse, fail-closed choices. Do not add comments that restate the code.
100
+
101
+ ## Commits and pull requests
102
+
103
+ Write commit subjects in the imperative mood ("add lease reattachment", not "added"
104
+ or "adds"). Keep the PR description focused on what changed and why; if it fixes an
105
+ issue, link it. A PR that changes behavior should update `CHANGELOG.md` under an
106
+ `## [Unreleased]` heading, and one that changes the interface should update the
107
+ relevant file in `docs/`.
108
+
109
+ ## Reporting bugs
110
+
111
+ Include your `gq` version (`gq --version`), Linux distribution, Python version, and
112
+ NVIDIA driver version, plus:
113
+
114
+ ```bash
115
+ gq gpu
116
+ gq ps
117
+ gq show JOB # for a specific job
118
+ tail -n 100 ~/.local/share/gq/gq-daemon.log
119
+ ```
120
+
121
+ Please skim the daemon log before attaching it — it records job commands and file
122
+ paths. **Never attach `gq.sqlite3`**; it contains the stored environment of every
123
+ job. See [SECURITY.md](SECURITY.md).
gq_local-0.1.1/LICENSE ADDED
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 Bo-Jyun Wang
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
@@ -0,0 +1,179 @@
1
+ Metadata-Version: 2.5
2
+ Name: gq-local
3
+ Version: 0.1.1
4
+ Summary: A small, single-user NVIDIA GPU job scheduler for Linux
5
+ Project-URL: Homepage, https://github.com/BirddWang/gq
6
+ Project-URL: Repository, https://github.com/BirddWang/gq
7
+ Project-URL: Issues, https://github.com/BirddWang/gq/issues
8
+ Project-URL: Changelog, https://github.com/BirddWang/gq/blob/main/CHANGELOG.md
9
+ Author: Bo-Jyun Wang
10
+ License: MIT
11
+ License-File: LICENSE
12
+ Requires-Python: >=3.11
13
+ Requires-Dist: nvidia-ml-py>=12.560.30
14
+ Description-Content-Type: text/markdown
15
+
16
+ # gq
17
+
18
+ [![CI](https://github.com/BirddWang/gq/actions/workflows/ci.yml/badge.svg)](https://github.com/BirddWang/gq/actions/workflows/ci.yml)
19
+ [![PyPI](https://img.shields.io/pypi/v/gq-local)](https://pypi.org/project/gq-local/)
20
+ [![Python](https://img.shields.io/pypi/pyversions/gq-local)](https://pypi.org/project/gq-local/)
21
+ [![License: MIT](https://img.shields.io/badge/license-MIT-blue)](LICENSE)
22
+
23
+ A small, single-user GPU job scheduler for one Linux workstation. Queue commands,
24
+ get whole NVIDIA GPUs assigned atomically, and keep jobs running after you close the
25
+ terminal.
26
+
27
+ ```bash
28
+ gq -g 1 uv run train.py # queued, runs when a GPU is genuinely free
29
+ gq -g 2 torchrun --nproc-per-node=2 train.py
30
+ gq ps # what is queued and running
31
+ gq logs -f 12 # follow a job's output
32
+ ```
33
+
34
+ No cluster, no root, no config file. One `pip install`, one user-level daemon.
35
+
36
+ ## The part that matters
37
+
38
+ Most of the time, "is this GPU free?" is answered by looking at utilization. That
39
+ answer is wrong in the case that bites hardest on a workstation:
40
+
41
+ > Your Jupyter kernel finished a cell an hour ago. It sits at **0% utilization** and
42
+ > still holds a CUDA context and several GB of memory. Start a training run on that
43
+ > GPU and you get an OOM — or worse, two jobs quietly fighting over one card.
44
+
45
+ `gq` treats a GPU as allocatable only when **both** of these hold:
46
+
47
+ 1. `gq` itself has no reservation on it, and
48
+ 2. NVML reports no compute process on it that `gq` does not own.
49
+
50
+ Utilization is never used as an ownership signal. A process `gq` does not recognise
51
+ makes the GPU `EXTERNAL` and off-limits — `gq` routes around it and never kills it.
52
+ When ownership is ambiguous, the GPU becomes `UNKNOWN` and stays unallocated. The
53
+ design fails closed: the worst case is an idle GPU, never a double-booked one.
54
+
55
+ That care extends to the parts that usually break:
56
+
57
+ - **Reservation is atomic.** The `WAITING → STARTING` transition and the full GPU
58
+ assignment are written in one SQLite transaction, before the process launches. A
59
+ crash between the two is not possible.
60
+ - **Restarts do not lose jobs.** Stopping the daemon leaves your jobs running. On
61
+ restart it reconciles against saved PID, `/proc` start-time ticks, process group,
62
+ Linux boot ID, and current GPU UUIDs, and reattaches what it can prove is the same
63
+ process.
64
+ - **Cancellation cannot hit the wrong process.** `gq` signals a saved process group
65
+ only after checking process identity, so a recycled PID is never mistaken for your
66
+ job.
67
+
68
+ ## Install
69
+
70
+ ```bash
71
+ uv tool install gq-local # or: pipx install gq-local, pip install gq-local
72
+ ```
73
+
74
+ Requires Linux, Python 3.11+, an NVIDIA driver with NVML, and no root. The only
75
+ runtime dependency is `nvidia-ml-py`.
76
+
77
+ ## Quick start
78
+
79
+ ```bash
80
+ gq daemon start # optional; any command starts it on demand
81
+ gq gpu # what each GPU is doing, and who owns it
82
+ gq -g 1 uv run train.py # submit; returns immediately
83
+ gq ps
84
+ gq logs -f 1
85
+ ```
86
+
87
+ `gq` sets `CUDA_VISIBLE_DEVICES` for the job, so physical GPUs 2,3 appear inside your
88
+ code as `cuda:0` and `cuda:1`.
89
+
90
+ Scripts can carry their own requirements:
91
+
92
+ ```bash
93
+ #!/bin/bash
94
+ #gq --gpus=2
95
+ #gq --name=experiment
96
+
97
+ uv run train.py
98
+ ```
99
+
100
+ ```bash
101
+ gq submit experiment.sh
102
+ ```
103
+
104
+ Day to day:
105
+
106
+ | Command | |
107
+ |---|---|
108
+ | `gq -g N CMD...` / `gq run -g N CMD...` | submit a command |
109
+ | `gq submit SCRIPT` | submit a shell script with `#gq` directives |
110
+ | `gq ps [--limit N] [--json]` | list jobs |
111
+ | `gq gpu [--json]` | physical and logical GPU state |
112
+ | `gq show JOB [--json]` | everything about one job |
113
+ | `gq logs [-f] JOB` | read or follow a job's output |
114
+ | `gq cancel JOB` | `SIGTERM` the process group, then `SIGKILL` after a grace period |
115
+ | `gq rm JOB...` | delete terminal jobs and their logs |
116
+ | `gq clean --older-than 30d` | sweep old terminal jobs and their logs |
117
+ | `gq daemon start\|stop\|status` | manage the daemon |
118
+
119
+ `--json` is the supported scripting interface; table layouts may change between
120
+ releases.
121
+
122
+ ## A note on secrets
123
+
124
+ A queued job may not start for hours, so `gq` saves your environment in order to run
125
+ it later with the same context. Variables that look like credentials are dropped
126
+ rather than stored, and `gq` tells you which:
127
+
128
+ ```text
129
+ gq: not persisting 1 secret-looking variable(s): HF_TOKEN
130
+ use --env-all (or GQ_ENV_ALL=1) to keep them, or --env-keep NAME
131
+ ```
132
+
133
+ If a job genuinely needs one, `--env-keep HF_TOKEN` puts it back. Anything kept is
134
+ stored in plaintext in a `0600` SQLite database — see [SECURITY.md](SECURITY.md).
135
+
136
+ ## Is this the right tool?
137
+
138
+ | You want | Use |
139
+ |---|---|
140
+ | Many machines, many users, fair-share, accounting | Slurm, or a Kubernetes scheduler |
141
+ | One workstation, your own jobs, whole GPUs, zero setup | **gq** |
142
+ | A generic job queue with no GPU awareness | [task-spooler](https://vicerveza.homeunix.net/~viric/soft/ts/) |
143
+ | Cloud or multi-node orchestration for ML | SkyPilot, Determined |
144
+
145
+ `gq` deliberately does **not** do multi-node scheduling, multiple users, CPU/RAM
146
+ allocation, priorities, preemption, MIG, containers, or fractional GPU sharing. Some
147
+ of those are on the roadmap; see [CONTRIBUTING.md](CONTRIBUTING.md#scope) before
148
+ proposing one.
149
+
150
+ ## Documentation
151
+
152
+ - [Usage](docs/usage.md) — every command, `#gq` directives, environment handling, retention
153
+ - [Architecture](docs/architecture.md) — allocation invariant, IPC, crash and reboot recovery
154
+ - [Troubleshooting](docs/troubleshooting.md) — stuck GPUs, stuck jobs, daemon problems
155
+ - [Security](SECURITY.md) — threat model and how credentials are handled
156
+ - [Changelog](CHANGELOG.md)
157
+
158
+ ## Development
159
+
160
+ ```bash
161
+ git clone https://github.com/BirddWang/gq
162
+ cd gq
163
+ uv sync --all-groups
164
+ uv run pytest
165
+ ```
166
+
167
+ The suite runs against a fake NVML provider, so **no GPU is required** to develop or
168
+ to run CI. Hardware tests are opt-in:
169
+
170
+ ```bash
171
+ GQ_RUN_GPU_TESTS=1 uv run pytest tests/integration
172
+ ```
173
+
174
+ Contributions are welcome — please read [CONTRIBUTING.md](CONTRIBUTING.md) first,
175
+ especially the section on the allocation invariant.
176
+
177
+ ## License
178
+
179
+ MIT — see [LICENSE](LICENSE).