audit-findings 0.1.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,93 @@
1
+ Metadata-Version: 2.4
2
+ Name: audit-findings
3
+ Version: 0.1.0
4
+ Summary: Record audit findings from parallel agents in an append-only log and render them as PDF, HTML, and GitHub issues.
5
+ Author: Fábio Macêdo Mendes
6
+ Author-email: Fábio Macêdo Mendes <fabiomacedomendes@gmail.com>
7
+ License-Expression: MIT
8
+ Requires-Dist: filelock>=3.12
9
+ Requires-Dist: pillow>=10.0
10
+ Requires-Dist: reportlab>=4.0
11
+ Requires-Python: >=3.10
12
+ Project-URL: Repository, https://github.com/fabiommendes/audit-findings
13
+ Description-Content-Type: text/markdown
14
+
15
+ # audit-findings
16
+
17
+ Records the findings of a code or interface audit in an append-only log and
18
+ renders them into a PDF report, an HTML report, and ready-to-paste GitHub
19
+ issues. It is the tool behind the `audit-*` skills in
20
+ [fabiommendes/skills](https://github.com/fabiommendes/skills).
21
+
22
+ The log lets several agents record findings in parallel, and an audit continue
23
+ across sessions, without any of them reading or rewriting the whole report:
24
+
25
+ - every write goes through the command line, which holds a file lock, assigns
26
+ ids, and validates the record before appending it;
27
+ - a finding's location and snippet are checked against the source when it is
28
+ recorded, so mistakes surface to the agent that made them;
29
+ - `status` prints a compact summary to resume from.
30
+
31
+ ## Usage
32
+
33
+ ```bash
34
+ uvx audit-findings --help
35
+ ```
36
+
37
+ Records are JSON objects passed on stdin:
38
+
39
+ ```bash
40
+ audit-findings -a webserver init <<'EOF'
41
+ {"lang": "en", "title": "Security Audit Report", "project": "acme-api"}
42
+ EOF
43
+
44
+ audit-findings -a webserver add category <<'EOF'
45
+ [{"id": "idor", "title": "IDOR"},
46
+ {"id": "xss", "title": "XSS", "applies": false, "note": "no HTML rendering"}]
47
+ EOF
48
+
49
+ audit-findings -a webserver add finding <<'EOF'
50
+ {"category": "idor", "severity": "critical",
51
+ "title": "GET /invoices/{id} does not check the tenant",
52
+ "location": "src/api/invoices.py:42-47",
53
+ "snippet": "return db.query(Invoice).filter(Invoice.id == invoice_id).first()",
54
+ "description": "...", "impact": "...", "fix": "..."}
55
+ EOF
56
+ # prints the new id: F1
57
+
58
+ audit-findings -a webserver status
59
+ audit-findings -a webserver render
60
+ ```
61
+
62
+ The log is `docs/audits/<area>/findings.jsonl`. `--area` and `--dir` may be
63
+ omitted when the project holds a single log.
64
+
65
+ | Command | Purpose |
66
+ |---|---|
67
+ | `init` | Create the log; optional `meta` JSON on stdin. |
68
+ | `import FILE` | Create the log from an existing `findings.json`. |
69
+ | `add KIND` | Add a record or an array of records; prints the ids. |
70
+ | `update ID` | Merge a JSON object into a record; `null` deletes a field. |
71
+ | `remove ID...` | Remove records nothing else refers to. |
72
+ | `start`, `done`, `reopen` | Track progress per category. |
73
+ | `status`, `list`, `show` | Read the log. |
74
+ | `check` | Check every location and snippet against the source again. |
75
+ | `build` | Validate and write `findings.json`. |
76
+ | `render [FILE]` | Build, then write `report.pdf`, `report.html`, `issues.md`. |
77
+
78
+ Record kinds: `category`, `finding`, `strength`, `risk`, `recommendation`,
79
+ `issue`, `inventory` (a coverage table), and `row` (a row of one). The fields
80
+ are described in the `audit-report` skill.
81
+
82
+ ## Log format
83
+
84
+ Each line of `findings.jsonl` is one event:
85
+
86
+ ```json
87
+ {"op": "add", "type": "finding", "id": "F3", "data": {"...": "..."}, "agent": "a1", "ts": "2026-10-08T12:00:00+00:00"}
88
+ {"op": "update", "id": "F3", "data": {"severity": "high"}, "ts": "..."}
89
+ {"op": "remove", "id": "F3", "ts": "..."}
90
+ ```
91
+
92
+ Replaying the events gives the current records; `build` turns them into
93
+ `findings.json`. Removed ids are never reused.
@@ -0,0 +1,79 @@
1
+ # audit-findings
2
+
3
+ Records the findings of a code or interface audit in an append-only log and
4
+ renders them into a PDF report, an HTML report, and ready-to-paste GitHub
5
+ issues. It is the tool behind the `audit-*` skills in
6
+ [fabiommendes/skills](https://github.com/fabiommendes/skills).
7
+
8
+ The log lets several agents record findings in parallel, and an audit continue
9
+ across sessions, without any of them reading or rewriting the whole report:
10
+
11
+ - every write goes through the command line, which holds a file lock, assigns
12
+ ids, and validates the record before appending it;
13
+ - a finding's location and snippet are checked against the source when it is
14
+ recorded, so mistakes surface to the agent that made them;
15
+ - `status` prints a compact summary to resume from.
16
+
17
+ ## Usage
18
+
19
+ ```bash
20
+ uvx audit-findings --help
21
+ ```
22
+
23
+ Records are JSON objects passed on stdin:
24
+
25
+ ```bash
26
+ audit-findings -a webserver init <<'EOF'
27
+ {"lang": "en", "title": "Security Audit Report", "project": "acme-api"}
28
+ EOF
29
+
30
+ audit-findings -a webserver add category <<'EOF'
31
+ [{"id": "idor", "title": "IDOR"},
32
+ {"id": "xss", "title": "XSS", "applies": false, "note": "no HTML rendering"}]
33
+ EOF
34
+
35
+ audit-findings -a webserver add finding <<'EOF'
36
+ {"category": "idor", "severity": "critical",
37
+ "title": "GET /invoices/{id} does not check the tenant",
38
+ "location": "src/api/invoices.py:42-47",
39
+ "snippet": "return db.query(Invoice).filter(Invoice.id == invoice_id).first()",
40
+ "description": "...", "impact": "...", "fix": "..."}
41
+ EOF
42
+ # prints the new id: F1
43
+
44
+ audit-findings -a webserver status
45
+ audit-findings -a webserver render
46
+ ```
47
+
48
+ The log is `docs/audits/<area>/findings.jsonl`. `--area` and `--dir` may be
49
+ omitted when the project holds a single log.
50
+
51
+ | Command | Purpose |
52
+ |---|---|
53
+ | `init` | Create the log; optional `meta` JSON on stdin. |
54
+ | `import FILE` | Create the log from an existing `findings.json`. |
55
+ | `add KIND` | Add a record or an array of records; prints the ids. |
56
+ | `update ID` | Merge a JSON object into a record; `null` deletes a field. |
57
+ | `remove ID...` | Remove records nothing else refers to. |
58
+ | `start`, `done`, `reopen` | Track progress per category. |
59
+ | `status`, `list`, `show` | Read the log. |
60
+ | `check` | Check every location and snippet against the source again. |
61
+ | `build` | Validate and write `findings.json`. |
62
+ | `render [FILE]` | Build, then write `report.pdf`, `report.html`, `issues.md`. |
63
+
64
+ Record kinds: `category`, `finding`, `strength`, `risk`, `recommendation`,
65
+ `issue`, `inventory` (a coverage table), and `row` (a row of one). The fields
66
+ are described in the `audit-report` skill.
67
+
68
+ ## Log format
69
+
70
+ Each line of `findings.jsonl` is one event:
71
+
72
+ ```json
73
+ {"op": "add", "type": "finding", "id": "F3", "data": {"...": "..."}, "agent": "a1", "ts": "2026-10-08T12:00:00+00:00"}
74
+ {"op": "update", "id": "F3", "data": {"severity": "high"}, "ts": "..."}
75
+ {"op": "remove", "id": "F3", "ts": "..."}
76
+ ```
77
+
78
+ Replaying the events gives the current records; `build` turns them into
79
+ `findings.json`. Removed ids are never reused.
@@ -0,0 +1,45 @@
1
+ [project]
2
+ name = "audit-findings"
3
+ version = "0.1.0"
4
+ description = "Record audit findings from parallel agents in an append-only log and render them as PDF, HTML, and GitHub issues."
5
+ readme = "README.md"
6
+ license = "MIT"
7
+ requires-python = ">=3.10"
8
+ dependencies = [
9
+ "filelock>=3.12",
10
+ "pillow>=10.0",
11
+ "reportlab>=4.0",
12
+ ]
13
+
14
+ [[project.authors]]
15
+ name = "Fábio Macêdo Mendes"
16
+ email = "fabiomacedomendes@gmail.com"
17
+
18
+ [project.urls]
19
+ Repository = "https://github.com/fabiommendes/audit-findings"
20
+
21
+ [project.scripts]
22
+ audit-findings = "audit_findings.cli:main"
23
+
24
+ [build-system]
25
+ requires = ["uv_build>=0.11,<0.12"]
26
+ build-backend = "uv_build"
27
+
28
+ [dependency-groups]
29
+ dev = [
30
+ "pytest>=8.0",
31
+ "ruff>=0.6",
32
+ ]
33
+
34
+ [tool.ruff]
35
+ line-length = 120
36
+
37
+ [tool.ruff.lint]
38
+ select = [
39
+ "E",
40
+ "F",
41
+ "I",
42
+ "UP",
43
+ "B",
44
+ "SIM",
45
+ ]
@@ -0,0 +1,35 @@
1
+ [project]
2
+ name = "audit-findings"
3
+ version = "0.1.0"
4
+ description = "Record audit findings from parallel agents in an append-only log and render them as PDF, HTML, and GitHub issues."
5
+ readme = "README.md"
6
+ license = "MIT"
7
+ authors = [{ name = "Fábio Macêdo Mendes", email = "fabiomacedomendes@gmail.com" }]
8
+ requires-python = ">=3.10"
9
+ dependencies = [
10
+ "filelock>=3.12",
11
+ "pillow>=10.0",
12
+ "reportlab>=4.0",
13
+ ]
14
+
15
+ [project.urls]
16
+ Repository = "https://github.com/fabiommendes/audit-findings"
17
+
18
+ [project.scripts]
19
+ audit-findings = "audit_findings.cli:main"
20
+
21
+ [build-system]
22
+ requires = ["uv_build>=0.11,<0.12"]
23
+ build-backend = "uv_build"
24
+
25
+ [dependency-groups]
26
+ dev = [
27
+ "pytest>=8.0",
28
+ "ruff>=0.6",
29
+ ]
30
+
31
+ [tool.ruff]
32
+ line-length = 120
33
+
34
+ [tool.ruff.lint]
35
+ select = ["E", "F", "I", "UP", "B", "SIM"]
@@ -0,0 +1,3 @@
1
+ """Record audit findings in an append-only log and render the report."""
2
+
3
+ __version__ = "0.1.0"
@@ -0,0 +1,5 @@
1
+ import sys
2
+
3
+ from .cli import main
4
+
5
+ sys.exit(main())
@@ -0,0 +1,187 @@
1
+ """Check that the code locations and snippets cited by findings exist in the source.
2
+
3
+ Each `path:line` or `path:start-end` in a finding's `location` must name an
4
+ existing file and lines inside it, and every line of `snippet` must appear in
5
+ the cited lines. Strength `evidence` gets the same location check. Locations
6
+ that are not file paths (routes, screens) are skipped.
7
+
8
+ Snippets may elide code with `...` (whole lines or inside a line), mask secrets
9
+ with `****`, and join a statement split over up to three source lines.
10
+ Whitespace is ignored when comparing.
11
+ """
12
+
13
+ from __future__ import annotations
14
+
15
+ import re
16
+ from dataclasses import dataclass
17
+ from pathlib import Path
18
+
19
+ # Lines a snippet line may sit outside the cited range before it counts as wrong.
20
+ TOLERANCE = 2
21
+ # Source lines joined when matching a snippet line, for statements split over lines.
22
+ JOIN = 3
23
+ WILDCARDS = re.compile(r"\*\*\*\*|\.\.\.|…")
24
+ ELISIONS = {"...", "…", "# ...", "// ...", "/* ... */", "<!-- ... -->", "-- ..."}
25
+ PATH_LINE = re.compile(r"^(?P<path>[^\s:]+?):(?P<lines>\d+(?:-\d+)?(?:,\d+(?:-\d+)?)*)$")
26
+ SEPARATORS = re.compile(r"[\s;]+|,\s+")
27
+ STRIP = "`'\"()[]<>"
28
+
29
+
30
+ @dataclass
31
+ class Ref:
32
+ path: str
33
+ ranges: list[tuple[int, int]] # empty: the whole file
34
+
35
+
36
+ def parse_location(location: str) -> tuple[list[Ref], list[str]]:
37
+ """Split a location string into file references and unparsed tokens."""
38
+ refs: list[Ref] = []
39
+ skipped: list[str] = []
40
+ for raw in SEPARATORS.split(location):
41
+ token = raw.strip(STRIP).rstrip(".,")
42
+ if not token:
43
+ continue
44
+ match = PATH_LINE.match(token)
45
+ if match:
46
+ ranges = []
47
+ for part in match["lines"].split(","):
48
+ start, _, end = part.partition("-")
49
+ ranges.append((int(start), int(end or start)))
50
+ refs.append(Ref(match["path"], ranges))
51
+ elif looks_like_file(token):
52
+ refs.append(Ref(token, []))
53
+ else:
54
+ skipped.append(token)
55
+ return refs, skipped
56
+
57
+
58
+ def looks_like_file(token: str) -> bool:
59
+ """A relative path with a directory or an extension, not a route like /users/{id}."""
60
+ if token.startswith("/") or "{" in token or "://" in token:
61
+ return False
62
+ return "/" in token or bool(re.search(r"\.[A-Za-z0-9]{1,8}$", token))
63
+
64
+
65
+ def overlaps(a: str, b: str) -> str | None:
66
+ """The first file range two locations share, as `path:start-end`, or None."""
67
+ refs_b, _ = parse_location(b)
68
+ for ref_a in parse_location(a)[0]:
69
+ for ref_b in refs_b:
70
+ if ref_a.path != ref_b.path:
71
+ continue
72
+ if not ref_a.ranges or not ref_b.ranges:
73
+ return ref_a.path
74
+ for start_a, end_a in ref_a.ranges:
75
+ for start_b, end_b in ref_b.ranges:
76
+ if start_a <= end_b and start_b <= end_a:
77
+ return f"{ref_a.path}:{max(start_a, start_b)}-{min(end_a, end_b)}"
78
+ return None
79
+
80
+
81
+ def normalize(line: str) -> str:
82
+ return re.sub(r"\s+", "", line)
83
+
84
+
85
+ def line_pattern(line: str) -> re.Pattern[str]:
86
+ """Match a snippet line against source, letting **** (masked) and ... (elided) stand for any text."""
87
+ parts = [re.escape(normalize(part)) for part in WILDCARDS.split(line)]
88
+ return re.compile(".*?".join(parts))
89
+
90
+
91
+ def window(lines: list[str], number: int) -> str:
92
+ """Source line `number` (1-based) joined with the next JOIN - 1 lines, whitespace removed."""
93
+ return "".join(normalize(text) for text in lines[number - 1 : number - 1 + JOIN])
94
+
95
+
96
+ def read_lines(path: Path) -> list[str] | None:
97
+ try:
98
+ return path.read_text(encoding="utf-8", errors="replace").splitlines()
99
+ except (OSError, UnicodeError):
100
+ return None
101
+
102
+
103
+ class Checker:
104
+ def __init__(self, root: Path):
105
+ self.root = root.resolve()
106
+ self.cache: dict[str, list[str] | None] = {}
107
+ self.errors: list[str] = []
108
+ self.warnings: list[str] = []
109
+
110
+ def lines(self, path: str) -> list[str] | None:
111
+ if path not in self.cache:
112
+ file = (self.root / path).resolve()
113
+ ok = file.is_file() and file.is_relative_to(self.root)
114
+ self.cache[path] = read_lines(file) if ok else None
115
+ return self.cache[path]
116
+
117
+ def check(self, data: dict) -> None:
118
+ for finding in data.get("findings", []):
119
+ self.check_finding(f"finding {finding.get('id', '?')}", finding)
120
+ for index, strength in enumerate(data.get("strengths", []), start=1):
121
+ if isinstance(strength, dict):
122
+ self.check_strength(f"strength {index}", strength)
123
+
124
+ def check_finding(self, where: str, finding: dict) -> None:
125
+ refs, skipped = parse_location(str(finding.get("location", "")))
126
+ valid = self.check_refs(where, refs)
127
+ snippet = finding.get("snippet")
128
+ if not snippet:
129
+ return
130
+ if valid:
131
+ self.check_snippet(where, snippet, valid)
132
+ elif not refs:
133
+ self.warnings.append(f"{where}: snippet not checked; location has no file path ({' '.join(skipped)})")
134
+
135
+ def check_strength(self, where: str, strength: dict) -> None:
136
+ if strength.get("evidence"):
137
+ refs, _ = parse_location(str(strength["evidence"]))
138
+ self.check_refs(where, refs)
139
+
140
+ def check_refs(self, where: str, refs: list[Ref]) -> list[Ref]:
141
+ """Report missing files and out-of-range lines; return the valid refs."""
142
+ valid = []
143
+ for ref in refs:
144
+ lines = self.lines(ref.path)
145
+ if lines is None:
146
+ self.errors.append(f"{where}: file not found: {ref.path} (relative to {self.root})")
147
+ continue
148
+ bad = [(s, e) for s, e in ref.ranges if s < 1 or e < s or e > len(lines)]
149
+ for start, end in bad:
150
+ self.errors.append(f"{where}: {ref.path} has {len(lines)} lines; {start}-{end} is out of range")
151
+ if not bad:
152
+ valid.append(ref)
153
+ return valid
154
+
155
+ def check_snippet(self, where: str, snippet: str, refs: list[Ref]) -> None:
156
+ wanted = [line for line in snippet.splitlines() if normalize(line) and line.strip() not in ELISIONS]
157
+ for line in wanted:
158
+ pattern = line_pattern(line)
159
+ if self.found_in_range(pattern, refs):
160
+ continue
161
+ elsewhere = self.found_anywhere(pattern, refs)
162
+ text = line.strip()
163
+ if elsewhere:
164
+ self.errors.append(f"{where}: snippet line found at {elsewhere}, outside the cited lines: {text}")
165
+ else:
166
+ self.errors.append(f"{where}: snippet line not found in the cited files: {text}")
167
+
168
+ def found_in_range(self, pattern: re.Pattern[str], refs: list[Ref]) -> bool:
169
+ for ref in refs:
170
+ lines = self.lines(ref.path) or []
171
+ spans = ref.ranges or [(1, len(lines))]
172
+ for start, end in spans:
173
+ lo, hi = max(1, start - TOLERANCE), min(len(lines), end + TOLERANCE)
174
+ if any(pattern.search(window(lines, i)) for i in range(lo, hi + 1)):
175
+ return True
176
+ return False
177
+
178
+ def found_anywhere(self, pattern: re.Pattern[str], refs: list[Ref]) -> str | None:
179
+ for ref in refs:
180
+ lines = self.lines(ref.path) or []
181
+ numbers = range(1, len(lines) + 1)
182
+ # Prefer a single-line match so the reported line is exact.
183
+ number = next((n for n in numbers if pattern.search(normalize(lines[n - 1]))), None)
184
+ number = number or next((n for n in numbers if pattern.search(window(lines, n))), None)
185
+ if number:
186
+ return f"{ref.path}:{number}"
187
+ return None