gq-local 0.1.1__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- gq_local-0.1.1/.github/ISSUE_TEMPLATE/bug_report.yml +58 -0
- gq_local-0.1.1/.github/ISSUE_TEMPLATE/config.yml +5 -0
- gq_local-0.1.1/.github/ISSUE_TEMPLATE/feature_request.yml +29 -0
- gq_local-0.1.1/.github/pull_request_template.md +18 -0
- gq_local-0.1.1/.github/workflows/ci.yml +63 -0
- gq_local-0.1.1/.github/workflows/release.yml +50 -0
- gq_local-0.1.1/.gitignore +10 -0
- gq_local-0.1.1/CHANGELOG.md +54 -0
- gq_local-0.1.1/CONTRIBUTING.md +123 -0
- gq_local-0.1.1/LICENSE +21 -0
- gq_local-0.1.1/PKG-INFO +179 -0
- gq_local-0.1.1/README.md +164 -0
- gq_local-0.1.1/SECURITY.md +72 -0
- gq_local-0.1.1/docs/architecture.md +104 -0
- gq_local-0.1.1/docs/troubleshooting.md +142 -0
- gq_local-0.1.1/docs/usage.md +207 -0
- gq_local-0.1.1/pyproject.toml +66 -0
- gq_local-0.1.1/src/gq/__init__.py +3 -0
- gq_local-0.1.1/src/gq/allocator.py +71 -0
- gq_local-0.1.1/src/gq/cli.py +550 -0
- gq_local-0.1.1/src/gq/daemon.py +287 -0
- gq_local-0.1.1/src/gq/database.py +324 -0
- gq_local-0.1.1/src/gq/models.py +128 -0
- gq_local-0.1.1/src/gq/nvml.py +89 -0
- gq_local-0.1.1/src/gq/paths.py +50 -0
- gq_local-0.1.1/src/gq/processes.py +97 -0
- gq_local-0.1.1/src/gq/protocol.py +48 -0
- gq_local-0.1.1/src/gq/scheduler.py +446 -0
- gq_local-0.1.1/tests/conftest.py +33 -0
- gq_local-0.1.1/tests/integration/test_nvml.py +70 -0
- gq_local-0.1.1/tests/test_allocator.py +70 -0
- gq_local-0.1.1/tests/test_cli.py +71 -0
- gq_local-0.1.1/tests/test_database.py +125 -0
- gq_local-0.1.1/tests/test_recovery.py +84 -0
- gq_local-0.1.1/tests/test_scheduler.py +209 -0
- gq_local-0.1.1/uv.lock +408 -0
|
@@ -0,0 +1,58 @@
|
|
|
1
|
+
name: Bug report
|
|
2
|
+
description: Something in gq behaves incorrectly
|
|
3
|
+
labels: ["bug"]
|
|
4
|
+
body:
|
|
5
|
+
- type: markdown
|
|
6
|
+
attributes:
|
|
7
|
+
value: |
|
|
8
|
+
Before filing, please check
|
|
9
|
+
[docs/troubleshooting.md](https://github.com/BirddWang/gq/blob/main/docs/troubleshooting.md)
|
|
10
|
+
— GPUs stuck in `EXTERNAL`/`UNKNOWN`, jobs stuck in `WAITING`, and daemon
|
|
11
|
+
startup failures are covered there.
|
|
12
|
+
|
|
13
|
+
**Do not attach `gq.sqlite3`.** It stores the environment of every job.
|
|
14
|
+
Skim the daemon log before pasting it: it records job commands and paths.
|
|
15
|
+
- type: textarea
|
|
16
|
+
id: what-happened
|
|
17
|
+
attributes:
|
|
18
|
+
label: What happened
|
|
19
|
+
description: What did you expect, and what happened instead?
|
|
20
|
+
validations:
|
|
21
|
+
required: true
|
|
22
|
+
- type: textarea
|
|
23
|
+
id: reproduce
|
|
24
|
+
attributes:
|
|
25
|
+
label: Steps to reproduce
|
|
26
|
+
placeholder: |
|
|
27
|
+
1. gq -g 1 ...
|
|
28
|
+
2. gq ps
|
|
29
|
+
3. ...
|
|
30
|
+
validations:
|
|
31
|
+
required: true
|
|
32
|
+
- type: input
|
|
33
|
+
id: version
|
|
34
|
+
attributes:
|
|
35
|
+
label: gq version
|
|
36
|
+
description: Output of `gq --version`
|
|
37
|
+
validations:
|
|
38
|
+
required: true
|
|
39
|
+
- type: input
|
|
40
|
+
id: environment
|
|
41
|
+
attributes:
|
|
42
|
+
label: Environment
|
|
43
|
+
description: Linux distribution, Python version, NVIDIA driver version, and whether you run in a container
|
|
44
|
+
placeholder: "Ubuntu 24.04, Python 3.12, driver 560.35, inside Docker"
|
|
45
|
+
validations:
|
|
46
|
+
required: true
|
|
47
|
+
- type: textarea
|
|
48
|
+
id: state
|
|
49
|
+
attributes:
|
|
50
|
+
label: Scheduler state
|
|
51
|
+
description: Output of `gq gpu` and `gq ps`, plus `gq show JOB` if one job is involved
|
|
52
|
+
render: text
|
|
53
|
+
- type: textarea
|
|
54
|
+
id: daemon-log
|
|
55
|
+
attributes:
|
|
56
|
+
label: Daemon log
|
|
57
|
+
description: Relevant lines from `tail -n 100 ~/.local/share/gq/gq-daemon.log`
|
|
58
|
+
render: text
|
|
@@ -0,0 +1,29 @@
|
|
|
1
|
+
name: Feature request
|
|
2
|
+
description: Suggest something gq should do
|
|
3
|
+
labels: ["enhancement"]
|
|
4
|
+
body:
|
|
5
|
+
- type: markdown
|
|
6
|
+
attributes:
|
|
7
|
+
value: |
|
|
8
|
+
`gq` is deliberately small, and some things are permanent non-goals:
|
|
9
|
+
multi-node scheduling, CPU/RAM allocation, MIG, and fractional GPU sharing.
|
|
10
|
+
See [CONTRIBUTING.md](https://github.com/BirddWang/gq/blob/main/CONTRIBUTING.md#scope).
|
|
11
|
+
|
|
12
|
+
Please open an issue before writing the code, so a rejected PR does not
|
|
13
|
+
waste your time.
|
|
14
|
+
- type: textarea
|
|
15
|
+
id: problem
|
|
16
|
+
attributes:
|
|
17
|
+
label: The problem
|
|
18
|
+
description: What are you trying to do that gq makes hard today? Describe the situation, not the solution.
|
|
19
|
+
validations:
|
|
20
|
+
required: true
|
|
21
|
+
- type: textarea
|
|
22
|
+
id: proposal
|
|
23
|
+
attributes:
|
|
24
|
+
label: What you have in mind
|
|
25
|
+
description: If you have a specific interface in mind, sketch it.
|
|
26
|
+
- type: textarea
|
|
27
|
+
id: alternatives
|
|
28
|
+
attributes:
|
|
29
|
+
label: What you do instead today
|
|
@@ -0,0 +1,18 @@
|
|
|
1
|
+
## What this changes
|
|
2
|
+
|
|
3
|
+
<!-- And why. Link the issue it addresses, if there is one. -->
|
|
4
|
+
|
|
5
|
+
## Checklist
|
|
6
|
+
|
|
7
|
+
- [ ] `uv run pytest` passes
|
|
8
|
+
- [ ] `uv run ruff check .` and `uv run mypy` pass
|
|
9
|
+
- [ ] Tests cover the change
|
|
10
|
+
- [ ] `CHANGELOG.md` updated under `## [Unreleased]`, if behavior changed
|
|
11
|
+
- [ ] `docs/` updated, if the interface changed
|
|
12
|
+
- [ ] Schema change appends a migration and bumps `SCHEMA_VERSION` (see CONTRIBUTING.md)
|
|
13
|
+
- [ ] IPC change bumps `PROTOCOL_VERSION`, if it is not backward compatible
|
|
14
|
+
|
|
15
|
+
## Allocation safety
|
|
16
|
+
|
|
17
|
+
<!-- Delete if this does not touch allocation, cancellation, or recovery.
|
|
18
|
+
Otherwise: which invariant does this rely on, and what test shows it holds? -->
|
|
@@ -0,0 +1,63 @@
|
|
|
1
|
+
name: CI
|
|
2
|
+
|
|
3
|
+
on:
|
|
4
|
+
push:
|
|
5
|
+
branches: [main]
|
|
6
|
+
pull_request:
|
|
7
|
+
workflow_dispatch:
|
|
8
|
+
|
|
9
|
+
concurrency:
|
|
10
|
+
group: ci-${{ github.ref }}
|
|
11
|
+
cancel-in-progress: true
|
|
12
|
+
|
|
13
|
+
jobs:
|
|
14
|
+
test:
|
|
15
|
+
name: pytest (Python ${{ matrix.python-version }})
|
|
16
|
+
runs-on: ubuntu-latest
|
|
17
|
+
strategy:
|
|
18
|
+
fail-fast: false
|
|
19
|
+
matrix:
|
|
20
|
+
python-version: ["3.11", "3.12", "3.13"]
|
|
21
|
+
steps:
|
|
22
|
+
- uses: actions/checkout@v4
|
|
23
|
+
- uses: astral-sh/setup-uv@v5
|
|
24
|
+
with:
|
|
25
|
+
enable-cache: true
|
|
26
|
+
- name: Install
|
|
27
|
+
run: uv sync --all-groups --python ${{ matrix.python-version }}
|
|
28
|
+
- name: Test
|
|
29
|
+
# The suite runs against a fake NVML provider, so no GPU is needed. The
|
|
30
|
+
# hardware tests under tests/integration stay opt-in via GQ_RUN_GPU_TESTS.
|
|
31
|
+
run: uv run pytest
|
|
32
|
+
|
|
33
|
+
lint:
|
|
34
|
+
name: ruff and mypy
|
|
35
|
+
runs-on: ubuntu-latest
|
|
36
|
+
steps:
|
|
37
|
+
- uses: actions/checkout@v4
|
|
38
|
+
- uses: astral-sh/setup-uv@v5
|
|
39
|
+
with:
|
|
40
|
+
enable-cache: true
|
|
41
|
+
- name: Install
|
|
42
|
+
run: uv sync --all-groups
|
|
43
|
+
- name: Ruff
|
|
44
|
+
run: uv run ruff check .
|
|
45
|
+
- name: Ruff format check
|
|
46
|
+
run: uv run ruff format --check .
|
|
47
|
+
- name: Mypy
|
|
48
|
+
run: uv run mypy
|
|
49
|
+
|
|
50
|
+
build:
|
|
51
|
+
name: build distributions
|
|
52
|
+
runs-on: ubuntu-latest
|
|
53
|
+
steps:
|
|
54
|
+
- uses: actions/checkout@v4
|
|
55
|
+
- uses: astral-sh/setup-uv@v5
|
|
56
|
+
- name: Build
|
|
57
|
+
run: uv build
|
|
58
|
+
- name: Check metadata
|
|
59
|
+
run: uvx twine check dist/*
|
|
60
|
+
- uses: actions/upload-artifact@v4
|
|
61
|
+
with:
|
|
62
|
+
name: distributions
|
|
63
|
+
path: dist/
|
|
@@ -0,0 +1,50 @@
|
|
|
1
|
+
name: Release
|
|
2
|
+
|
|
3
|
+
# Publishes to PyPI when a version tag is pushed. Authentication uses PyPI trusted
|
|
4
|
+
# publishing (OIDC), so no API token is stored in the repository.
|
|
5
|
+
#
|
|
6
|
+
# One-time setup at https://pypi.org/manage/account/publishing/ :
|
|
7
|
+
# PyPI project name: gq-local
|
|
8
|
+
# Owner: BirddWang
|
|
9
|
+
# Repository: gq
|
|
10
|
+
# Workflow: release.yml
|
|
11
|
+
# Environment: pypi
|
|
12
|
+
|
|
13
|
+
on:
|
|
14
|
+
push:
|
|
15
|
+
tags: ["v*"]
|
|
16
|
+
workflow_dispatch:
|
|
17
|
+
|
|
18
|
+
jobs:
|
|
19
|
+
build:
|
|
20
|
+
runs-on: ubuntu-latest
|
|
21
|
+
steps:
|
|
22
|
+
- uses: actions/checkout@v4
|
|
23
|
+
- uses: astral-sh/setup-uv@v5
|
|
24
|
+
- run: uv build
|
|
25
|
+
- name: Fail if the tag and the package version disagree
|
|
26
|
+
run: |
|
|
27
|
+
version="$(uv run python -c 'import gq; print(gq.__version__)')"
|
|
28
|
+
tag="${GITHUB_REF_NAME#v}"
|
|
29
|
+
if [ "$version" != "$tag" ]; then
|
|
30
|
+
echo "tag $tag does not match gq.__version__ $version" >&2
|
|
31
|
+
exit 1
|
|
32
|
+
fi
|
|
33
|
+
if: startsWith(github.ref, 'refs/tags/')
|
|
34
|
+
- uses: actions/upload-artifact@v4
|
|
35
|
+
with:
|
|
36
|
+
name: distributions
|
|
37
|
+
path: dist/
|
|
38
|
+
|
|
39
|
+
publish:
|
|
40
|
+
needs: build
|
|
41
|
+
runs-on: ubuntu-latest
|
|
42
|
+
environment: pypi
|
|
43
|
+
permissions:
|
|
44
|
+
id-token: write
|
|
45
|
+
steps:
|
|
46
|
+
- uses: actions/download-artifact@v4
|
|
47
|
+
with:
|
|
48
|
+
name: distributions
|
|
49
|
+
path: dist/
|
|
50
|
+
- uses: pypa/gh-action-pypi-publish@release/v1
|
|
@@ -0,0 +1,54 @@
|
|
|
1
|
+
# Changelog
|
|
2
|
+
|
|
3
|
+
All notable changes to this project are documented here. This project follows
|
|
4
|
+
[Semantic Versioning](https://semver.org/spec/v2.0.0.html).
|
|
5
|
+
|
|
6
|
+
## [0.1.1] - 2026-09-10
|
|
7
|
+
|
|
8
|
+
A hardening release. No new scheduling behavior; everything here came from defects
|
|
9
|
+
and rough edges observed in the first week of real use.
|
|
10
|
+
|
|
11
|
+
### Fixed
|
|
12
|
+
|
|
13
|
+
- The daemon no longer logs a traceback every time a client hangs up mid-response.
|
|
14
|
+
`gq ps | head` and Ctrl-C out of `gq logs -f` are ordinary events, but each one
|
|
15
|
+
raised `ConnectionResetError`/`BrokenPipeError` out of the connection callback.
|
|
16
|
+
- Running as root inside a container is now allowed without `GQ_ALLOW_ROOT=1`.
|
|
17
|
+
Docker and Jupyter images commonly have uid 0 as the only account, and the guard
|
|
18
|
+
made `gq` unusable there. It still refuses on what looks like a regular host, where
|
|
19
|
+
the warning is genuinely warranted.
|
|
20
|
+
- `PRAGMA foreign_keys = ON` is now issued as its own statement rather than inside
|
|
21
|
+
the schema script, so `ON DELETE CASCADE` on `job_gpus` cannot silently be a no-op.
|
|
22
|
+
|
|
23
|
+
### Added
|
|
24
|
+
|
|
25
|
+
- **Schema migrations.** The database carries a `PRAGMA user_version` and applies
|
|
26
|
+
ordered migrations, each in one transaction with its version bump. Databases
|
|
27
|
+
created before this release are adopted in place, not rebuilt.
|
|
28
|
+
- **`gq rm JOB...`** deletes terminal jobs and their logs. Active jobs are refused,
|
|
29
|
+
and one bad id rejects the whole batch rather than partially destroying history.
|
|
30
|
+
- **`gq clean --older-than AGE [--state STATE] [-y]`** sweeps old terminal jobs and
|
|
31
|
+
their logs. Defaults to 30d and confirms before deleting.
|
|
32
|
+
- **`--json`** on `gq ps`, `gq gpu`, and `gq show`.
|
|
33
|
+
- **Protocol versioning.** Every request carries a protocol number; a mismatched
|
|
34
|
+
daemon reports how to restart instead of failing on a missing field. `gq daemon
|
|
35
|
+
status` warns when the running daemon predates the installed client.
|
|
36
|
+
|
|
37
|
+
### Changed
|
|
38
|
+
|
|
39
|
+
- **Secret-looking environment variables are no longer persisted by default.** The
|
|
40
|
+
submission environment is stored verbatim in SQLite so queued jobs survive a daemon
|
|
41
|
+
restart, which made every captured credential a durable one. Variables matching
|
|
42
|
+
`TOKEN`, `SECRET`, `PASSWORD`, `CREDENTIAL`, `*_KEY`, `API_KEY` or the `AWS_`,
|
|
43
|
+
`AZURE_`, `GCP_`, `GOOGLE_APPLICATION_` prefixes are dropped, and the dropped names
|
|
44
|
+
are printed at submit time rather than being removed silently.
|
|
45
|
+
|
|
46
|
+
This is a **behavior change that can break jobs** which need a credential at
|
|
47
|
+
runtime — `HF_TOKEN` for a private Hugging Face dataset, for instance. Restore the
|
|
48
|
+
old behavior per-job with `--env-all`, globally with `GQ_ENV_ALL=1`, or for one
|
|
49
|
+
variable with `--env-keep HF_TOKEN`.
|
|
50
|
+
|
|
51
|
+
## [0.1.0] - 2026-08-31
|
|
52
|
+
|
|
53
|
+
Initial MVP: queueing, atomic NVML-checked GPU allocation, `CUDA_VISIBLE_DEVICES`
|
|
54
|
+
assignment, detached execution, cancellation, and crash/reboot recovery.
|
|
@@ -0,0 +1,123 @@
|
|
|
1
|
+
# Contributing to gq
|
|
2
|
+
|
|
3
|
+
Thanks for taking a look. `gq` is deliberately small, and the bar for adding to it is
|
|
4
|
+
correspondingly high — please read "Scope" before starting on a feature.
|
|
5
|
+
|
|
6
|
+
## Getting set up
|
|
7
|
+
|
|
8
|
+
```bash
|
|
9
|
+
git clone https://github.com/BirddWang/gq
|
|
10
|
+
cd gq
|
|
11
|
+
uv sync --all-groups
|
|
12
|
+
uv run pytest
|
|
13
|
+
```
|
|
14
|
+
|
|
15
|
+
Most of the suite runs against a fake NVML provider (`tests/conftest.py`), so you do
|
|
16
|
+
**not** need an NVIDIA GPU to develop or to run CI. If you do have one, the hardware
|
|
17
|
+
discovery tests are opt-in:
|
|
18
|
+
|
|
19
|
+
```bash
|
|
20
|
+
GQ_RUN_GPU_TESTS=1 uv run pytest tests/integration
|
|
21
|
+
```
|
|
22
|
+
|
|
23
|
+
Before opening a pull request:
|
|
24
|
+
|
|
25
|
+
```bash
|
|
26
|
+
uv run ruff check .
|
|
27
|
+
uv run ruff format .
|
|
28
|
+
uv run mypy
|
|
29
|
+
uv run pytest
|
|
30
|
+
```
|
|
31
|
+
|
|
32
|
+
### Testing against a scratch instance
|
|
33
|
+
|
|
34
|
+
Never develop against your real queue. All state is redirectable, so run an isolated
|
|
35
|
+
daemon instead:
|
|
36
|
+
|
|
37
|
+
```bash
|
|
38
|
+
export GQ_DATA_DIR=/tmp/gq-dev/data
|
|
39
|
+
export GQ_RUNTIME_DIR=/tmp/gq-dev/run
|
|
40
|
+
export GQ_RECONCILE_INTERVAL=0.5
|
|
41
|
+
export GQ_CANCEL_GRACE_SECONDS=2
|
|
42
|
+
gq daemon start
|
|
43
|
+
```
|
|
44
|
+
|
|
45
|
+
A separate `GQ_RUNTIME_DIR` gets its own lock and socket, so it will not collide with
|
|
46
|
+
a daemon you already have running.
|
|
47
|
+
|
|
48
|
+
## Scope
|
|
49
|
+
|
|
50
|
+
The MVP intentionally omits multi-node scheduling, CPU/RAM allocation, preemption,
|
|
51
|
+
MIG, containers, and fractional GPU sharing. Some of these are on the roadmap and
|
|
52
|
+
some are permanent non-goals. Please open an issue to discuss a feature before
|
|
53
|
+
writing it — a rejected PR is a waste of your time, and this project would rather
|
|
54
|
+
stay small than grow every reasonable idea.
|
|
55
|
+
|
|
56
|
+
## The one rule that matters
|
|
57
|
+
|
|
58
|
+
> A GPU assigned to a managed job is never simultaneously allocatable.
|
|
59
|
+
|
|
60
|
+
Everything in `allocator.py`, `scheduler.py`, and `database.py` exists to hold that
|
|
61
|
+
line, and the design **fails closed**: when ownership is ambiguous, the GPU becomes
|
|
62
|
+
`UNKNOWN` and no one gets it. Three consequences to keep in mind when changing
|
|
63
|
+
allocation code:
|
|
64
|
+
|
|
65
|
+
- **Utilization is never an ownership signal.** A Jupyter kernel sitting at 0% still
|
|
66
|
+
owns its CUDA context. Only NVML compute-process presence and `gq`'s own durable
|
|
67
|
+
reservations decide ownership.
|
|
68
|
+
- **The reservation and the state transition are one transaction.** `Database.reserve`
|
|
69
|
+
moves `WAITING → STARTING` and writes the full GPU assignment inside a single
|
|
70
|
+
`BEGIN IMMEDIATE`. Do not introduce an `await` or a second allocation between them.
|
|
71
|
+
- **Process identity is checked before signalling.** Use the helpers in
|
|
72
|
+
`processes.py`; never signal a bare PID. A recycled PID must not be mistaken for
|
|
73
|
+
the original job.
|
|
74
|
+
|
|
75
|
+
If a change would weaken any of these, it needs a very good argument and a test that
|
|
76
|
+
demonstrates the new invariant holds.
|
|
77
|
+
|
|
78
|
+
## Compatibility
|
|
79
|
+
|
|
80
|
+
- **Database.** Never edit the baseline schema to change an existing table. Append a
|
|
81
|
+
migration in `MIGRATIONS` in `database.py`, keyed by the version it produces, and
|
|
82
|
+
bump `SCHEMA_VERSION`. Users have real job history; a release that discards it is a
|
|
83
|
+
bug. Test your migration against a copy of a populated database.
|
|
84
|
+
- **IPC.** Adding a request type is backward compatible. Changing or removing the
|
|
85
|
+
meaning of an existing field is not — bump `PROTOCOL_VERSION` in `protocol.py`.
|
|
86
|
+
- **`--json` output** is the supported scripting interface. Adding fields is fine;
|
|
87
|
+
renaming or removing them requires a protocol bump. Table layouts are for humans
|
|
88
|
+
and may change freely.
|
|
89
|
+
|
|
90
|
+
## Style
|
|
91
|
+
|
|
92
|
+
Match the surrounding code. It is plain, typed Python with no framework: standard
|
|
93
|
+
library plus `nvidia-ml-py`, and `from __future__ import annotations` everywhere.
|
|
94
|
+
|
|
95
|
+
**Adding a runtime dependency requires a strong justification.** Being installable
|
|
96
|
+
with one small dependency is a feature of this project, not an accident.
|
|
97
|
+
|
|
98
|
+
Comments explain *why*, especially where the reasoning is subtle — race conditions,
|
|
99
|
+
PID reuse, fail-closed choices. Do not add comments that restate the code.
|
|
100
|
+
|
|
101
|
+
## Commits and pull requests
|
|
102
|
+
|
|
103
|
+
Write commit subjects in the imperative mood ("add lease reattachment", not "added"
|
|
104
|
+
or "adds"). Keep the PR description focused on what changed and why; if it fixes an
|
|
105
|
+
issue, link it. A PR that changes behavior should update `CHANGELOG.md` under an
|
|
106
|
+
`## [Unreleased]` heading, and one that changes the interface should update the
|
|
107
|
+
relevant file in `docs/`.
|
|
108
|
+
|
|
109
|
+
## Reporting bugs
|
|
110
|
+
|
|
111
|
+
Include your `gq` version (`gq --version`), Linux distribution, Python version, and
|
|
112
|
+
NVIDIA driver version, plus:
|
|
113
|
+
|
|
114
|
+
```bash
|
|
115
|
+
gq gpu
|
|
116
|
+
gq ps
|
|
117
|
+
gq show JOB # for a specific job
|
|
118
|
+
tail -n 100 ~/.local/share/gq/gq-daemon.log
|
|
119
|
+
```
|
|
120
|
+
|
|
121
|
+
Please skim the daemon log before attaching it — it records job commands and file
|
|
122
|
+
paths. **Never attach `gq.sqlite3`**; it contains the stored environment of every
|
|
123
|
+
job. See [SECURITY.md](SECURITY.md).
|
gq_local-0.1.1/LICENSE
ADDED
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 Bo-Jyun Wang
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|
gq_local-0.1.1/PKG-INFO
ADDED
|
@@ -0,0 +1,179 @@
|
|
|
1
|
+
Metadata-Version: 2.5
|
|
2
|
+
Name: gq-local
|
|
3
|
+
Version: 0.1.1
|
|
4
|
+
Summary: A small, single-user NVIDIA GPU job scheduler for Linux
|
|
5
|
+
Project-URL: Homepage, https://github.com/BirddWang/gq
|
|
6
|
+
Project-URL: Repository, https://github.com/BirddWang/gq
|
|
7
|
+
Project-URL: Issues, https://github.com/BirddWang/gq/issues
|
|
8
|
+
Project-URL: Changelog, https://github.com/BirddWang/gq/blob/main/CHANGELOG.md
|
|
9
|
+
Author: Bo-Jyun Wang
|
|
10
|
+
License: MIT
|
|
11
|
+
License-File: LICENSE
|
|
12
|
+
Requires-Python: >=3.11
|
|
13
|
+
Requires-Dist: nvidia-ml-py>=12.560.30
|
|
14
|
+
Description-Content-Type: text/markdown
|
|
15
|
+
|
|
16
|
+
# gq
|
|
17
|
+
|
|
18
|
+
[](https://github.com/BirddWang/gq/actions/workflows/ci.yml)
|
|
19
|
+
[](https://pypi.org/project/gq-local/)
|
|
20
|
+
[](https://pypi.org/project/gq-local/)
|
|
21
|
+
[](LICENSE)
|
|
22
|
+
|
|
23
|
+
A small, single-user GPU job scheduler for one Linux workstation. Queue commands,
|
|
24
|
+
get whole NVIDIA GPUs assigned atomically, and keep jobs running after you close the
|
|
25
|
+
terminal.
|
|
26
|
+
|
|
27
|
+
```bash
|
|
28
|
+
gq -g 1 uv run train.py # queued, runs when a GPU is genuinely free
|
|
29
|
+
gq -g 2 torchrun --nproc-per-node=2 train.py
|
|
30
|
+
gq ps # what is queued and running
|
|
31
|
+
gq logs -f 12 # follow a job's output
|
|
32
|
+
```
|
|
33
|
+
|
|
34
|
+
No cluster, no root, no config file. One `pip install`, one user-level daemon.
|
|
35
|
+
|
|
36
|
+
## The part that matters
|
|
37
|
+
|
|
38
|
+
Most of the time, "is this GPU free?" is answered by looking at utilization. That
|
|
39
|
+
answer is wrong in the case that bites hardest on a workstation:
|
|
40
|
+
|
|
41
|
+
> Your Jupyter kernel finished a cell an hour ago. It sits at **0% utilization** and
|
|
42
|
+
> still holds a CUDA context and several GB of memory. Start a training run on that
|
|
43
|
+
> GPU and you get an OOM — or worse, two jobs quietly fighting over one card.
|
|
44
|
+
|
|
45
|
+
`gq` treats a GPU as allocatable only when **both** of these hold:
|
|
46
|
+
|
|
47
|
+
1. `gq` itself has no reservation on it, and
|
|
48
|
+
2. NVML reports no compute process on it that `gq` does not own.
|
|
49
|
+
|
|
50
|
+
Utilization is never used as an ownership signal. A process `gq` does not recognise
|
|
51
|
+
makes the GPU `EXTERNAL` and off-limits — `gq` routes around it and never kills it.
|
|
52
|
+
When ownership is ambiguous, the GPU becomes `UNKNOWN` and stays unallocated. The
|
|
53
|
+
design fails closed: the worst case is an idle GPU, never a double-booked one.
|
|
54
|
+
|
|
55
|
+
That care extends to the parts that usually break:
|
|
56
|
+
|
|
57
|
+
- **Reservation is atomic.** The `WAITING → STARTING` transition and the full GPU
|
|
58
|
+
assignment are written in one SQLite transaction, before the process launches. A
|
|
59
|
+
crash between the two is not possible.
|
|
60
|
+
- **Restarts do not lose jobs.** Stopping the daemon leaves your jobs running. On
|
|
61
|
+
restart it reconciles against saved PID, `/proc` start-time ticks, process group,
|
|
62
|
+
Linux boot ID, and current GPU UUIDs, and reattaches what it can prove is the same
|
|
63
|
+
process.
|
|
64
|
+
- **Cancellation cannot hit the wrong process.** `gq` signals a saved process group
|
|
65
|
+
only after checking process identity, so a recycled PID is never mistaken for your
|
|
66
|
+
job.
|
|
67
|
+
|
|
68
|
+
## Install
|
|
69
|
+
|
|
70
|
+
```bash
|
|
71
|
+
uv tool install gq-local # or: pipx install gq-local, pip install gq-local
|
|
72
|
+
```
|
|
73
|
+
|
|
74
|
+
Requires Linux, Python 3.11+, an NVIDIA driver with NVML, and no root. The only
|
|
75
|
+
runtime dependency is `nvidia-ml-py`.
|
|
76
|
+
|
|
77
|
+
## Quick start
|
|
78
|
+
|
|
79
|
+
```bash
|
|
80
|
+
gq daemon start # optional; any command starts it on demand
|
|
81
|
+
gq gpu # what each GPU is doing, and who owns it
|
|
82
|
+
gq -g 1 uv run train.py # submit; returns immediately
|
|
83
|
+
gq ps
|
|
84
|
+
gq logs -f 1
|
|
85
|
+
```
|
|
86
|
+
|
|
87
|
+
`gq` sets `CUDA_VISIBLE_DEVICES` for the job, so physical GPUs 2,3 appear inside your
|
|
88
|
+
code as `cuda:0` and `cuda:1`.
|
|
89
|
+
|
|
90
|
+
Scripts can carry their own requirements:
|
|
91
|
+
|
|
92
|
+
```bash
|
|
93
|
+
#!/bin/bash
|
|
94
|
+
#gq --gpus=2
|
|
95
|
+
#gq --name=experiment
|
|
96
|
+
|
|
97
|
+
uv run train.py
|
|
98
|
+
```
|
|
99
|
+
|
|
100
|
+
```bash
|
|
101
|
+
gq submit experiment.sh
|
|
102
|
+
```
|
|
103
|
+
|
|
104
|
+
Day to day:
|
|
105
|
+
|
|
106
|
+
| Command | |
|
|
107
|
+
|---|---|
|
|
108
|
+
| `gq -g N CMD...` / `gq run -g N CMD...` | submit a command |
|
|
109
|
+
| `gq submit SCRIPT` | submit a shell script with `#gq` directives |
|
|
110
|
+
| `gq ps [--limit N] [--json]` | list jobs |
|
|
111
|
+
| `gq gpu [--json]` | physical and logical GPU state |
|
|
112
|
+
| `gq show JOB [--json]` | everything about one job |
|
|
113
|
+
| `gq logs [-f] JOB` | read or follow a job's output |
|
|
114
|
+
| `gq cancel JOB` | `SIGTERM` the process group, then `SIGKILL` after a grace period |
|
|
115
|
+
| `gq rm JOB...` | delete terminal jobs and their logs |
|
|
116
|
+
| `gq clean --older-than 30d` | sweep old terminal jobs and their logs |
|
|
117
|
+
| `gq daemon start\|stop\|status` | manage the daemon |
|
|
118
|
+
|
|
119
|
+
`--json` is the supported scripting interface; table layouts may change between
|
|
120
|
+
releases.
|
|
121
|
+
|
|
122
|
+
## A note on secrets
|
|
123
|
+
|
|
124
|
+
A queued job may not start for hours, so `gq` saves your environment in order to run
|
|
125
|
+
it later with the same context. Variables that look like credentials are dropped
|
|
126
|
+
rather than stored, and `gq` tells you which:
|
|
127
|
+
|
|
128
|
+
```text
|
|
129
|
+
gq: not persisting 1 secret-looking variable(s): HF_TOKEN
|
|
130
|
+
use --env-all (or GQ_ENV_ALL=1) to keep them, or --env-keep NAME
|
|
131
|
+
```
|
|
132
|
+
|
|
133
|
+
If a job genuinely needs one, `--env-keep HF_TOKEN` puts it back. Anything kept is
|
|
134
|
+
stored in plaintext in a `0600` SQLite database — see [SECURITY.md](SECURITY.md).
|
|
135
|
+
|
|
136
|
+
## Is this the right tool?
|
|
137
|
+
|
|
138
|
+
| You want | Use |
|
|
139
|
+
|---|---|
|
|
140
|
+
| Many machines, many users, fair-share, accounting | Slurm, or a Kubernetes scheduler |
|
|
141
|
+
| One workstation, your own jobs, whole GPUs, zero setup | **gq** |
|
|
142
|
+
| A generic job queue with no GPU awareness | [task-spooler](https://vicerveza.homeunix.net/~viric/soft/ts/) |
|
|
143
|
+
| Cloud or multi-node orchestration for ML | SkyPilot, Determined |
|
|
144
|
+
|
|
145
|
+
`gq` deliberately does **not** do multi-node scheduling, multiple users, CPU/RAM
|
|
146
|
+
allocation, priorities, preemption, MIG, containers, or fractional GPU sharing. Some
|
|
147
|
+
of those are on the roadmap; see [CONTRIBUTING.md](CONTRIBUTING.md#scope) before
|
|
148
|
+
proposing one.
|
|
149
|
+
|
|
150
|
+
## Documentation
|
|
151
|
+
|
|
152
|
+
- [Usage](docs/usage.md) — every command, `#gq` directives, environment handling, retention
|
|
153
|
+
- [Architecture](docs/architecture.md) — allocation invariant, IPC, crash and reboot recovery
|
|
154
|
+
- [Troubleshooting](docs/troubleshooting.md) — stuck GPUs, stuck jobs, daemon problems
|
|
155
|
+
- [Security](SECURITY.md) — threat model and how credentials are handled
|
|
156
|
+
- [Changelog](CHANGELOG.md)
|
|
157
|
+
|
|
158
|
+
## Development
|
|
159
|
+
|
|
160
|
+
```bash
|
|
161
|
+
git clone https://github.com/BirddWang/gq
|
|
162
|
+
cd gq
|
|
163
|
+
uv sync --all-groups
|
|
164
|
+
uv run pytest
|
|
165
|
+
```
|
|
166
|
+
|
|
167
|
+
The suite runs against a fake NVML provider, so **no GPU is required** to develop or
|
|
168
|
+
to run CI. Hardware tests are opt-in:
|
|
169
|
+
|
|
170
|
+
```bash
|
|
171
|
+
GQ_RUN_GPU_TESTS=1 uv run pytest tests/integration
|
|
172
|
+
```
|
|
173
|
+
|
|
174
|
+
Contributions are welcome — please read [CONTRIBUTING.md](CONTRIBUTING.md) first,
|
|
175
|
+
especially the section on the allocation invariant.
|
|
176
|
+
|
|
177
|
+
## License
|
|
178
|
+
|
|
179
|
+
MIT — see [LICENSE](LICENSE).
|