org-knowledge-layer 0.1.0__tar.gz → 0.1.2__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (113) hide show
  1. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/.claude/settings.local.json +24 -1
  2. org_knowledge_layer-0.1.2/.github/ISSUE_TEMPLATE/bug_report.md +21 -0
  3. org_knowledge_layer-0.1.2/.github/ISSUE_TEMPLATE/feature_request.md +15 -0
  4. org_knowledge_layer-0.1.2/.github/pull_request_template.md +15 -0
  5. org_knowledge_layer-0.1.2/CONTRIBUTING.md +70 -0
  6. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/PKG-INFO +29 -11
  7. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/README.md +28 -10
  8. org_knowledge_layer-0.1.2/SECURITY.md +56 -0
  9. org_knowledge_layer-0.1.2/docs/ab-results-chart.png +0 -0
  10. org_knowledge_layer-0.1.2/docs/ab-results-chart.svg +27 -0
  11. org_knowledge_layer-0.1.2/docs/posts/01-memory-that-outlives-the-run.md +65 -0
  12. org_knowledge_layer-0.1.2/docs/posts/02-dont-let-a-step-grade-itself.md +45 -0
  13. org_knowledge_layer-0.1.2/docs/posts/03-enforcement-or-good-intentions.md +37 -0
  14. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/evals/REPORT.md +6 -0
  15. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/pyproject.toml +1 -1
  16. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/seed/react-defects.json +1 -1
  17. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/src/okl/scaffold/profiles/react/README.md +1 -1
  18. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/src/okl/scaffold/profiles/react/rules/frontend.md +2 -2
  19. org_knowledge_layer-0.1.0/docs/posts/01-memory-that-outlives-the-run.md +0 -43
  20. org_knowledge_layer-0.1.0/docs/posts/02-dont-let-a-step-grade-itself.md +0 -37
  21. org_knowledge_layer-0.1.0/docs/posts/03-enforcement-or-good-intentions.md +0 -33
  22. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/.claude/hooks/stop-okl-encode.sh +0 -0
  23. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/.claude/hooks/userpromptsubmit-okl-check.sh +0 -0
  24. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/.claude/settings.json +0 -0
  25. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/.github/workflows/ci.yml +0 -0
  26. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/.github/workflows/okl-verify.yml +0 -0
  27. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/.gitignore +0 -0
  28. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/AGENTS.md +0 -0
  29. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/CLAUDE.md +0 -0
  30. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/LICENSE +0 -0
  31. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/ci/okl-verify.yml +0 -0
  32. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/docs/decisions/2026-07-17-flat-retrieval-until-scale.md +0 -0
  33. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/docs/decisions/2026-07-21-subject-tags-controlled-vocabulary.md +0 -0
  34. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/docs/okl-sixth-surface.excalidraw +0 -0
  35. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/docs/okl-sixth-surface.svg +0 -0
  36. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/evals/README.md +0 -0
  37. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/evals/ab_harness.py +0 -0
  38. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/evals/results/ab-20260829-2300.json +0 -0
  39. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/evals/results/ab-20260829-2315.json +0 -0
  40. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/evals/results/ab-20260830-0003.json +0 -0
  41. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/evals/results/ab-20260830-0148.json +0 -0
  42. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/evals/results/e2e-20260830/README.md +0 -0
  43. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/evals/results/e2e-20260830/briefed-userpromptsubmit-lint.yml +0 -0
  44. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/evals/results/e2e-20260830/control-lint.yml +0 -0
  45. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/evals/results/e2e-20260830/hook.log +0 -0
  46. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/evals/results/e2e-20260830/service-record-500.log +0 -0
  47. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/evals/results/e2e-20260830/session-briefed-pretooluse.txt +0 -0
  48. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/evals/results/e2e-20260830/session-briefed-userpromptsubmit.txt +0 -0
  49. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/evals/results/e2e-20260830/session-control.txt +0 -0
  50. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/evals/tasks.jsonl +0 -0
  51. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/hooks/stop-okl-encode.sh +0 -0
  52. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/hooks/userpromptsubmit-okl-check.sh +0 -0
  53. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/seed/dotnet-canon.json +0 -0
  54. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/seed/dotnet-decisions.json +0 -0
  55. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/seed/dotnet-defects.json +0 -0
  56. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/seed/dotnet-review-surfaces.json +0 -0
  57. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/seed/frontend-canon.json +0 -0
  58. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/seed/geospatial-deeptime-defects.json +0 -0
  59. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/seed/geospatial-defects.json +0 -0
  60. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/seed/geospatial-enforcement-defects.json +0 -0
  61. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/seed/geospatial-eval-defects.json +0 -0
  62. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/seed/rag-defects.json +0 -0
  63. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/src/okl/__init__.py +0 -0
  64. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/src/okl/__main__.py +0 -0
  65. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/src/okl/bootstrap.py +0 -0
  66. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/src/okl/cli.py +0 -0
  67. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/src/okl/client.py +0 -0
  68. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/src/okl/core.py +0 -0
  69. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/src/okl/drift.py +0 -0
  70. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/src/okl/mcp_server.py +0 -0
  71. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/src/okl/scaffold/MANIFEST.md +0 -0
  72. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/src/okl/scaffold/ci/method-gates.yml +0 -0
  73. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/src/okl/scaffold/ci/okl-verify.yml +0 -0
  74. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/src/okl/scaffold/claude/agents/architecture-reviewer.md +0 -0
  75. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/src/okl/scaffold/claude/commands/check-rules.md +0 -0
  76. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/src/okl/scaffold/claude/commands/feature-spec.md +0 -0
  77. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/src/okl/scaffold/claude/rules/example-area.md +0 -0
  78. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/src/okl/scaffold/claude/skills/RECOMMENDED-COMPANIONS.md +0 -0
  79. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/src/okl/scaffold/claude/skills/encoding-loop/SKILL.md +0 -0
  80. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/src/okl/scaffold/claude/skills/verify-before-claiming/SKILL.md +0 -0
  81. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/src/okl/scaffold/evals/README.md +0 -0
  82. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/src/okl/scaffold/evals/cases.jsonl +0 -0
  83. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/src/okl/scaffold/evals/run_evals.py +0 -0
  84. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/src/okl/scaffold/gates/check-canon-size.sh +0 -0
  85. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/src/okl/scaffold/gates/check-doc-orphans.sh +0 -0
  86. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/src/okl/scaffold/gates/check-retractions.sh +0 -0
  87. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/src/okl/scaffold/gates/check-tombstones.sh +0 -0
  88. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/src/okl/scaffold/gates/run-gates.sh +0 -0
  89. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/src/okl/scaffold/hooks/hooks.json +0 -0
  90. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/src/okl/scaffold/hooks/stop-okl-encode.sh +0 -0
  91. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/src/okl/scaffold/hooks/userpromptsubmit-okl-check.sh +0 -0
  92. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/src/okl/scaffold/plugin/plugin.json +0 -0
  93. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/src/okl/scaffold/profiles/dotnet/README.md +0 -0
  94. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/src/okl/scaffold/profiles/dotnet/rules/architecture.md +0 -0
  95. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/src/okl/scaffold/profiles/dotnet/rules/messaging.md +0 -0
  96. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/src/okl/scaffold/profiles/dotnet/rules/performance-and-data.md +0 -0
  97. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/src/okl/scaffold/profiles/dotnet/rules/security.md +0 -0
  98. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/src/okl/scaffold/profiles/geospatial/README.md +0 -0
  99. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/src/okl/scaffold/profiles/geospatial/rules/geospatial-ml.md +0 -0
  100. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/src/okl/scaffold/profiles/python-rag/README.md +0 -0
  101. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/src/okl/scaffold/profiles/python-rag/rules/fastapi-backend.md +0 -0
  102. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/src/okl/scaffold/profiles/python-rag/rules/project-structure.md +0 -0
  103. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/src/okl/scaffold/profiles/python-rag/rules/rag-pipeline.md +0 -0
  104. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/src/okl/scaffold/registries/RETRACTIONS.md +0 -0
  105. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/src/okl/scaffold/registries/tombstones.txt +0 -0
  106. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/src/okl/scaffold/root/CLAUDE.md +0 -0
  107. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/src/okl/scaffold/root/METHOD.md +0 -0
  108. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/src/okl/scaffold_cmd.py +0 -0
  109. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/src/okl/seed.py +0 -0
  110. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/src/okl/service.py +0 -0
  111. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/src/okl/store.py +0 -0
  112. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/tests/test_okl.py +0 -0
  113. {org_knowledge_layer-0.1.0 → org_knowledge_layer-0.1.2}/tests/test_scaffold.py +0 -0
@@ -67,7 +67,30 @@
67
67
  "Bash(./oklbuild/bin/pip install *)",
68
68
  "Bash(./oklbuild/bin/pip show *)",
69
69
  "Bash(/tmp/oklbuild/bin/python -m build --no-isolation)",
70
- "Bash(unzip -p dist/*.whl \"*/entry_points.txt\")"
70
+ "Bash(unzip -p dist/*.whl \"*/entry_points.txt\")",
71
+ "Bash(git -c core.hooksPath=/dev/null commit -qm 'Publish as org-knowledge-layer: PyPI blocks '\\\\''okl'\\\\'' as confusable with '\\\\''oki'\\\\'' *)",
72
+ "Bash(rm -rf okltest)",
73
+ "Bash(python3 -m venv okltest)",
74
+ "Bash(./okltest/bin/pip install *)",
75
+ "Bash(./okltest/bin/okl --help)",
76
+ "Bash(./okltest/bin/python -c ' *)",
77
+ "Bash(git -c core.hooksPath=/dev/null commit -qm 'v0.1.1: lead the README with the install/command distinction *)",
78
+ "Bash(curl -s https://pypi.org/pypi/org-knowledge-layer/json)",
79
+ "Bash(curl -s https://emeraldleaf.dev/writing/enforcement-or-good-intentions/)",
80
+ "Bash(curl -s https://emeraldleaf.dev/writing/memory-that-outlives-the-run/)",
81
+ "Bash(curl -s https://emeraldleaf.dev/practice/)",
82
+ "Bash(git -C \"/Users/joshuadell/Dev/emeraldleaf-dev\" status --porcelain)",
83
+ "Bash(git -C \"/Users/joshuadell/Dev/emeraldleaf-dev\" log --oneline -1)",
84
+ "Bash(command -v rsvg-convert qlmanage)",
85
+ "Bash(node scripts/validate_palette.js \"#2a78d6,#eb6834\" --mode light)",
86
+ "Bash(rsvg-convert -w 1200 -h 720 docs/ab-results-chart.svg -o docs/ab-results-chart.png)",
87
+ "Bash(rsvg-convert -w 1200 -h 1200 docs/ab-results-chart.svg -o docs/ab-results-chart.png)",
88
+ "Bash(rsvg-convert -w 550 -h 550 docs/ab-results-chart.svg -o /tmp/chart-feedsize.png)",
89
+ "Bash(git -c core.hooksPath=/dev/null commit -qm 'Say what the store holds: rules, architecture decisions, known defects *)",
90
+ "Bash(git -c core.hooksPath=/dev/null commit -qm 'Post 1: concede the crowded category, name the narrow differentiator *)",
91
+ "Bash(git -c core.hooksPath=/dev/null commit -qm 'Public-repo hygiene: SECURITY.md, CONTRIBUTING.md, templates, provenance wording *)",
92
+ "Bash(grep -n \"^ [a-z-]*:$\\\\|^name:\\\\|runs-on\" .github/workflows/ci.yml)",
93
+ "Bash(git -C /Users/joshuadell/Dev/okl push --dry-run origin main)"
71
94
  ],
72
95
  "additionalDirectories": [
73
96
  "/Users/joshuadell/Dev/okl/e2e/scratch-briefed/.okl",
@@ -0,0 +1,21 @@
1
+ ---
2
+ name: Bug report
3
+ about: Something behaves differently than documented
4
+ labels: bug
5
+ ---
6
+
7
+ **What you ran** (the exact command, and `okl --help` version if relevant)
8
+
9
+ **What you expected**
10
+
11
+ **What actually happened** (paste the real output, not a summary)
12
+
13
+ **Environment**
14
+ - OS:
15
+ - Python:
16
+ - Installed via: `pip install org-knowledge-layer` / `pip install -e .`
17
+ - Mode: local store / shared service
18
+
19
+ **Especially useful:** if the tool reported success it had not earned (a check that
20
+ passed without running, a stamp without evidence, a hook that fired but changed
21
+ nothing), say so explicitly. That class of defect is this project's subject.
@@ -0,0 +1,15 @@
1
+ ---
2
+ name: Proposal
3
+ about: Suggest a change to how the loop works
4
+ labels: proposal
5
+ ---
6
+
7
+ **The symptom you hit** (what went wrong in real use, not the feature you want)
8
+
9
+ **Why the current design does not cover it**
10
+
11
+ **What you would change**
12
+
13
+ **How you would know it worked** — the check, measurement, or test that would show it.
14
+ Proposals that name their own evidence get adopted faster; see `evals/REPORT.md` for
15
+ the standard this repo holds itself to.
@@ -0,0 +1,15 @@
1
+ ## What changed
2
+
3
+ ## Why
4
+
5
+ ## What you ran to check it
6
+
7
+ <!-- Paste real output. `pytest -q`, `ruff check .`, and `okl drift` at minimum. -->
8
+
9
+ ## Checklist
10
+
11
+ - [ ] `pytest -q` green
12
+ - [ ] `ruff check .` clean
13
+ - [ ] `okl drift` green, or affected records re-verified with `okl verify --run`
14
+ - [ ] Mirror files still byte-identical (`ci/`, `.github/workflows/`, `hooks/` and their `src/okl/scaffold/` twins)
15
+ - [ ] Any number quoted has a committed receipt in `evals/results/`
@@ -0,0 +1,70 @@
1
+ # Contributing
2
+
3
+ Thanks for looking. This is a v0 project with one maintainer, so the most useful
4
+ contributions are small, verifiable, and self-contained.
5
+
6
+ ## Setup
7
+
8
+ ```bash
9
+ git clone https://github.com/emeraldleaf/okl && cd okl
10
+ pip install -e ".[dev]"
11
+ pytest -q # must be green before and after your change
12
+ ruff check .
13
+ ```
14
+
15
+ Try it end to end in a scratch directory before changing anything:
16
+
17
+ ```bash
18
+ mkdir /tmp/try && cd /tmp/try && git init
19
+ okl init --repo try --interests security
20
+ okl seed
21
+ okl check --task "add an endpoint that returns an order for the logged-in user"
22
+ ```
23
+
24
+ ## The repo has rules about itself
25
+
26
+ Read [CLAUDE.md](CLAUDE.md) (identical to AGENTS.md) before a first PR. It is short, and
27
+ it is the actual contract. The parts that will fail your build if you miss them:
28
+
29
+ - **Mirror files are byte-identical.** `ci/okl-verify.yml`, `.github/workflows/okl-verify.yml`
30
+ and `src/okl/scaffold/ci/okl-verify.yml` must match, as must `hooks/*.sh` and their
31
+ `src/okl/scaffold/hooks/` twins. Edit one, copy to the others in the same change.
32
+ `tests/test_scaffold.py::test_mirror_files_identical` enforces it.
33
+ - **`okl drift` goes red when you change a file a stored rule governs.** That is the
34
+ system working. Re-verify the affected records with an actual check:
35
+ `okl verify <id> --run "pytest -q -k <test>" --expect "passed"`. Do not clear drift by
36
+ re-recording with `--verified`; the whole point is that stamps come from observed runs.
37
+ - **Tags come from a closed vocabulary** (`store.KNOWN_TAGS`). Growing it is a deliberate
38
+ edit to that set plus a note in the tags ADR, not an ad-hoc string.
39
+ - **Evaluation claims need a committed receipt.** If you change the harness or quote a
40
+ number, the run that produced it belongs in `evals/results/`. Never cite a run the
41
+ harness marked RESULTS NOT USABLE.
42
+
43
+ ## What is most welcome
44
+
45
+ - **Bug reports with a reproduction.** Especially anything where the tool reports success
46
+ it did not earn; that class of defect is the project's whole subject.
47
+ - **The open work in the store.** `okl check --task "improve retrieval precision"` will
48
+ show you the recorded defects against the tool itself, including the briefing relevance
49
+ cutoff that is still unfinished.
50
+ - **Portability fixes.** Hooks, path resolution, and CI have been exercised on macOS and
51
+ GitHub Actions and nowhere else.
52
+ - **A live-Postgres test.** The ranked search path for Postgres is currently asserted at
53
+ the SQL-shape level against a fake connection; it has never run against a real server.
54
+
55
+ ## What is out of scope for now
56
+
57
+ Embedding or vector retrieval, unless one of the triggers in
58
+ [the flat-retrieval ADR](docs/decisions/2026-07-17-flat-retrieval-until-scale.md) has
59
+ actually fired and you bring the measurement showing it. The decision is falsifiable on
60
+ purpose; falsify it with data rather than preference.
61
+
62
+ ## Pull requests
63
+
64
+ Keep them single-purpose. Say what you changed and what you ran to check it, and include
65
+ the real output rather than a description of it. If your change surfaced something
66
+ non-obvious, record it in the store the way the repo records its own findings; a PR that
67
+ teaches the system something is worth more than one that only fixes code.
68
+
69
+ By contributing you agree your work is licensed under the repository's
70
+ [MIT License](LICENSE).
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.4
2
2
  Name: org-knowledge-layer
3
- Version: 0.1.0
3
+ Version: 0.1.2
4
4
  Summary: Org Knowledge Layer — an installable sixth surface that carries encoded engineering lessons across repos.
5
5
  Author: Joshua Dell
6
6
  License: MIT
@@ -37,6 +37,14 @@ Description-Content-Type: text/markdown
37
37
  > be broken — plus a command that hands the relevant ones to a coding agent (or a
38
38
  > person) **before** they start a task, so the same mistake isn't made twice.
39
39
 
40
+ ```bash
41
+ pipx install org-knowledge-layer # the distribution name on PyPI
42
+ okl --help # the command, the import package, and the repo are all `okl`
43
+ ```
44
+
45
+ *(PyPI rejects `okl` itself as confusable with the existing `oki`, so only the install
46
+ line differs — everything you type afterwards is `okl`.)*
47
+
40
48
  ## The problem it solves
41
49
 
42
50
  A team (or an AI agent) fixes a subtle bug, learns *why* it happened, and writes a
@@ -216,18 +224,18 @@ knowledge. The switch is one environment variable; none of your commands change.
216
224
  ## Install
217
225
 
218
226
  ```bash
219
- pipx install okl # once it's published; for now:
220
- pip install -e . # from this repo
227
+ pipx install org-knowledge-layer # provides the `okl` command
228
+ pip install -e . # or from a clone of this repo
221
229
  ```
222
230
 
223
231
  The core (local + client + CLI) is **stdlib-only** — zero required dependencies.
224
232
  Extras are opt-in:
225
233
 
226
234
  ```bash
227
- pip install "okl[service]" # FastAPI shared service
228
- pip install "okl[postgres]" # Postgres backend (psycopg)
229
- pip install "okl[mcp]" # MCP server for Claude Code / Cursor / Copilot
230
- pip install "okl[all]"
235
+ pip install "org-knowledge-layer[service]" # FastAPI shared service
236
+ pip install "org-knowledge-layer[postgres]" # Postgres backend (psycopg)
237
+ pip install "org-knowledge-layer[mcp]" # MCP server for Claude Code / Cursor / Copilot
238
+ pip install "org-knowledge-layer[all]"
231
239
  ```
232
240
 
233
241
  ## Wire a repo
@@ -403,19 +411,23 @@ okl serve --port 8080
403
411
  ## Run the shared service
404
412
 
405
413
  ```bash
406
- pip install "okl[service]"
414
+ pip install "org-knowledge-layer[service]"
407
415
  OKL_DATABASE_URL="sqlite:///okl.db" OKL_TOKEN="a-shared-secret" okl serve
408
416
  # repos then: okl connect https://your-host --token a-shared-secret
409
417
  ```
410
418
 
411
- `OKL_TOKEN` (optional) gates **writes**; reads stay open. Deploy the service
419
+ `OKL_TOKEN` (optional) gates **writes**; **reads stay open** — `/nodes` returns the
420
+ whole store to anyone who can reach the port. A mature store holds your defect history,
421
+ security patterns and internal architecture, so treat it as sensitive: bind it to a
422
+ private network or put authentication in front of every route before exposing it. See
423
+ [SECURITY.md](SECURITY.md). Deploy the service
412
424
  wherever you like — it's storage-agnostic by design (a `Dockerfile` and a
413
425
  `fly.toml` are the obvious next commit; not included in v0).
414
426
 
415
427
  ## Agent integration (MCP)
416
428
 
417
429
  ```bash
418
- pip install "okl[mcp]"
430
+ pip install "org-knowledge-layer[mcp]"
419
431
  okl mcp # register in your coding agent's tool config
420
432
  ```
421
433
 
@@ -466,10 +478,16 @@ tests/ # end-to-end tests
466
478
  ## Test
467
479
 
468
480
  ```bash
469
- pip install "okl[dev]"
481
+ pip install "org-knowledge-layer[dev]"
470
482
  pytest -q # full suite (one drift test self-skips where git init is unavailable)
471
483
  ```
472
484
 
473
485
  ## License
474
486
 
475
487
  MIT.
488
+
489
+ ## Contributing
490
+
491
+ See [CONTRIBUTING.md](CONTRIBUTING.md) for setup, the repo's own rules (mirror files,
492
+ drift, evidence-based verification), and where help is most useful. Security policy and
493
+ the deployment threat model: [SECURITY.md](SECURITY.md).
@@ -5,6 +5,14 @@
5
5
  > be broken — plus a command that hands the relevant ones to a coding agent (or a
6
6
  > person) **before** they start a task, so the same mistake isn't made twice.
7
7
 
8
+ ```bash
9
+ pipx install org-knowledge-layer # the distribution name on PyPI
10
+ okl --help # the command, the import package, and the repo are all `okl`
11
+ ```
12
+
13
+ *(PyPI rejects `okl` itself as confusable with the existing `oki`, so only the install
14
+ line differs — everything you type afterwards is `okl`.)*
15
+
8
16
  ## The problem it solves
9
17
 
10
18
  A team (or an AI agent) fixes a subtle bug, learns *why* it happened, and writes a
@@ -184,18 +192,18 @@ knowledge. The switch is one environment variable; none of your commands change.
184
192
  ## Install
185
193
 
186
194
  ```bash
187
- pipx install okl # once it's published; for now:
188
- pip install -e . # from this repo
195
+ pipx install org-knowledge-layer # provides the `okl` command
196
+ pip install -e . # or from a clone of this repo
189
197
  ```
190
198
 
191
199
  The core (local + client + CLI) is **stdlib-only** — zero required dependencies.
192
200
  Extras are opt-in:
193
201
 
194
202
  ```bash
195
- pip install "okl[service]" # FastAPI shared service
196
- pip install "okl[postgres]" # Postgres backend (psycopg)
197
- pip install "okl[mcp]" # MCP server for Claude Code / Cursor / Copilot
198
- pip install "okl[all]"
203
+ pip install "org-knowledge-layer[service]" # FastAPI shared service
204
+ pip install "org-knowledge-layer[postgres]" # Postgres backend (psycopg)
205
+ pip install "org-knowledge-layer[mcp]" # MCP server for Claude Code / Cursor / Copilot
206
+ pip install "org-knowledge-layer[all]"
199
207
  ```
200
208
 
201
209
  ## Wire a repo
@@ -371,19 +379,23 @@ okl serve --port 8080
371
379
  ## Run the shared service
372
380
 
373
381
  ```bash
374
- pip install "okl[service]"
382
+ pip install "org-knowledge-layer[service]"
375
383
  OKL_DATABASE_URL="sqlite:///okl.db" OKL_TOKEN="a-shared-secret" okl serve
376
384
  # repos then: okl connect https://your-host --token a-shared-secret
377
385
  ```
378
386
 
379
- `OKL_TOKEN` (optional) gates **writes**; reads stay open. Deploy the service
387
+ `OKL_TOKEN` (optional) gates **writes**; **reads stay open** — `/nodes` returns the
388
+ whole store to anyone who can reach the port. A mature store holds your defect history,
389
+ security patterns and internal architecture, so treat it as sensitive: bind it to a
390
+ private network or put authentication in front of every route before exposing it. See
391
+ [SECURITY.md](SECURITY.md). Deploy the service
380
392
  wherever you like — it's storage-agnostic by design (a `Dockerfile` and a
381
393
  `fly.toml` are the obvious next commit; not included in v0).
382
394
 
383
395
  ## Agent integration (MCP)
384
396
 
385
397
  ```bash
386
- pip install "okl[mcp]"
398
+ pip install "org-knowledge-layer[mcp]"
387
399
  okl mcp # register in your coding agent's tool config
388
400
  ```
389
401
 
@@ -434,10 +446,16 @@ tests/ # end-to-end tests
434
446
  ## Test
435
447
 
436
448
  ```bash
437
- pip install "okl[dev]"
449
+ pip install "org-knowledge-layer[dev]"
438
450
  pytest -q # full suite (one drift test self-skips where git init is unavailable)
439
451
  ```
440
452
 
441
453
  ## License
442
454
 
443
455
  MIT.
456
+
457
+ ## Contributing
458
+
459
+ See [CONTRIBUTING.md](CONTRIBUTING.md) for setup, the repo's own rules (mirror files,
460
+ drift, evidence-based verification), and where help is most useful. Security policy and
461
+ the deployment threat model: [SECURITY.md](SECURITY.md).
@@ -0,0 +1,56 @@
1
+ # Security
2
+
3
+ ## Reporting a vulnerability
4
+
5
+ Please report privately via GitHub's **Report a vulnerability** button on the
6
+ [Security tab](https://github.com/emeraldleaf/okl/security), not in a public issue.
7
+ Expect an acknowledgement within a week. This is a v0 project maintained by one person;
8
+ fixes are best-effort, and the honest expectation is a patch release or a documented
9
+ mitigation, not a same-day turnaround.
10
+
11
+ Supported: the latest `0.1.x` release. Older versions get nothing.
12
+
13
+ ## What this software does that you should know about
14
+
15
+ Three behaviors are deliberate, not bugs. If any of them is wrong for your environment,
16
+ that is a configuration decision you need to make before deploying.
17
+
18
+ ### 1. The shared service leaves reads unauthenticated by default
19
+
20
+ `okl serve` gates **writes** behind `OKL_TOKEN` when that variable is set. **Reads are
21
+ always open**: `/check`, `/search`, `/nodes`, `/metric/recurrence` and `/health` require
22
+ no credential. `/nodes` returns the entire store.
23
+
24
+ That matters because of what a mature store contains: your defect history, the security
25
+ mistakes you have already made, internal architecture decisions, retired identifiers,
26
+ and the shape of your systems. **Treat the store as sensitive** and assume anyone who
27
+ can reach the port can read all of it.
28
+
29
+ If you deploy it beyond localhost:
30
+
31
+ - put it behind a reverse proxy that enforces authentication on every route, or bind it
32
+ to a private network;
33
+ - terminate TLS upstream (the app speaks plain HTTP and has no certificate handling);
34
+ - set `OKL_TOKEN` so writes are not anonymous;
35
+ - expect no rate limiting, no audit log of reads, and no per-user access control.
36
+
37
+ ### 2. `okl verify --run` executes a shell command by design
38
+
39
+ The verification command runs whatever you pass to `--run` through a shell, reads the
40
+ real exit status, and records the command as evidence. That is the feature: verification
41
+ must come from an executed check. It also means **never pass untrusted input to
42
+ `--run`**, and never wire it to a command built from unreviewed data.
43
+
44
+ ### 3. The hooks execute scripts from your repository
45
+
46
+ `okl init` installs shell hooks into `.claude/hooks/` and registers them. They run
47
+ whenever your agent runs. Read them before installing, the same as any other hook, and
48
+ review changes to them the same way you would review CI configuration.
49
+
50
+ ## What ships in the package
51
+
52
+ The published distribution includes seed data: real engineering rules, architecture
53
+ decisions, and defect classes from the author's own projects, with project identities
54
+ genericized. It contains no credentials, no customer data, and no third-party code.
55
+ The scaffold deliberately does not vendor other people's skills; see
56
+ `RECOMMENDED-COMPANIONS.md` in the kit for what it points at instead.
@@ -0,0 +1,27 @@
1
+ <svg xmlns="http://www.w3.org/2000/svg" width="1200" height="1200" viewBox="0 0 1200 1200" font-family="Helvetica Neue, Helvetica, Arial, sans-serif">
2
+ <rect width="1200" height="1200" fill="#fcfcfb"/>
3
+
4
+ <text x="600" y="136" font-size="46" fill="#52514e" text-anchor="middle">Known defects in agent-written code</text>
5
+
6
+ <text x="288" y="430" font-size="210" font-weight="700" fill="#2a78d6" text-anchor="middle">33%</text>
7
+ <text x="600" y="414" font-size="130" fill="#a8a7a1" text-anchor="middle">→</text>
8
+ <text x="912" y="430" font-size="210" font-weight="700" fill="#eb6834" text-anchor="middle">4%</text>
9
+
10
+ <text x="288" y="504" font-size="41" fill="#0b0b0b" text-anchor="middle">rules written down</text>
11
+ <text x="288" y="554" font-size="41" fill="#0b0b0b" text-anchor="middle">but not delivered</text>
12
+ <text x="912" y="504" font-size="41" fill="#0b0b0b" text-anchor="middle">rules retrieved into</text>
13
+ <text x="912" y="554" font-size="41" fill="#0b0b0b" text-anchor="middle">context, every task</text>
14
+
15
+ <line x1="140" y1="644" x2="1060" y2="644" stroke="#e6e5e1" stroke-width="1"/>
16
+
17
+ <text x="600" y="722" font-size="43" fill="#0b0b0b" text-anchor="middle">Same model, same tasks. The only difference:</text>
18
+ <text x="600" y="776" font-size="43" fill="#0b0b0b" text-anchor="middle">whether anything put them in front of the agent.</text>
19
+
20
+ <rect x="140" y="842" width="920" height="164" rx="16" fill="#f4f3f0"/>
21
+ <text x="600" y="908" font-size="41" fill="#0b0b0b" text-anchor="middle">A budget model with the rules (12%) beat</text>
22
+ <text x="600" y="960" font-size="41" fill="#0b0b0b" text-anchor="middle">a frontier model without them (33%).</text>
23
+
24
+ <text x="600" y="1072" font-size="26" fill="#82817c" text-anchor="middle">8 tasks, 3 samples per arm, graded blind by a second, different model.</text>
25
+ <text x="600" y="1110" font-size="26" fill="#82817c" text-anchor="middle">Measures defect classes the store covers, not general code quality. Small n.</text>
26
+ <text x="600" y="1152" font-size="26" fill="#82817c" text-anchor="middle">Method and receipts: github.com/emeraldleaf/okl</text>
27
+ </svg>
@@ -0,0 +1,65 @@
1
+ # Memory that outlives the run
2
+
3
+ *Part 1 of 3: The loop that learns. How one developer's AI loop remembers, verifies, and enforces.*
4
+
5
+ An agent writing code in one of my repos has no access to what the others learned. That is the problem in one sentence, and better prompting does not solve it.
6
+
7
+ Every project I work in has an instructions file with real standards in it. But a rule in a prompt is a suggestion, not a control: sometimes the agent applies it, sometimes I restate it, and neither of those is a system. Corrections do not persist either. A fix made in review is scoped to one conversation. The next session starts clean, and the next repository starts cleaner.
8
+
9
+ ## What the audit actually found
10
+
11
+ So I audited my own setup across three codebases: a .NET microservices platform, a geospatial ML pipeline, a Python RAG service. The question was not whether the code was good, it was whether the standards I had written down were reaching the work.
12
+
13
+ Two findings.
14
+
15
+ First, several rules I had documented as enforcement were enforcement in name only. In one repo an architecture-review agent was described across six documents as a live review surface; nothing in CI or in any hook actually invoked it. Every one of those documents was accurate about the design and wrong about the system. Writing a rule down and running it are different acts, and only one of them changes what ships.
16
+
17
+ Second, and worse: the three repos had independently derived the same rules. Each had paid separately for knowledge the others already had. My standards existed, they were even written down, and they still could not cross a repository boundary. I had been thinking of my setup as "two readers, one canon": a human reads the rules, an agent reads the rules, and the two must not drift. What I actually had was N readers and N canons.
18
+
19
+ ## The fix is a layer, not a bigger prompt
20
+
21
+ The obvious move is to paste every standard into every prompt. That fails in a specific way: always-on context taxes every session whether the rule is relevant or not. My .NET repo enforces a hard size budget on its instructions file in CI, warning at 400 lines and failing the build at 500, precisely because every byte of always-on context costs both readers attention.
22
+
23
+ So the design question became: where does each piece of knowledge belong? Three questions, asked in order.
24
+
25
+ **Does a session doing a completely unrelated task still need this?** Then it belongs in the always-on instructions. Almost nothing passes. Naming conventions, the security posture, the error-handling canon. This is the most expensive real estate in the system and the size budget is what keeps it honest.
26
+
27
+ **Would I want this in front of me whenever a future task resembles it?** Then it belongs in a retrievable store. This is the large category: the rate-limiter rule that matters enormously to the three tasks a year touching rate limiting, and is noise for everything else.
28
+
29
+ **Does it stop mattering once the PR merges?** Then it stays with the task. The subtle part is extraction, because a dead plan usually contains one durable decision worth pulling out before the rest is discarded.
30
+
31
+ The middle tier is the part I had to build: a small typed store of what the projects already know. Defects with symptom, cause, and fix. Rules. Deliberate decisions, recorded so they are not silently reversed. Tombstones for identifiers that must never come back. Retractions for claims that turned out to be false. Every note carries a scope, because org-wide standards should propagate to every project while one repo's quirks should stay home. Choosing that scope is a judgment call and it is the curation step that keeps a shared layer from filling with noise.
32
+
33
+ ## Retrieved, not loaded
34
+
35
+ The reason the store can grow indefinitely is that no session ever reads it. Before a task starts, a hook queries it with the task description and injects only what survives four filters: relevance ranking, scope, the declared subject interests of the current repo, and typed routing that turns the survivors into a short action list. Fix this. Run that gate. Do not restate this retracted claim.
36
+
37
+ Each run sees a dozen relevant lines rather than the whole library. At ten times the current corpus, what reaches the agent should be the same length and better chosen.
38
+
39
+ ## None of this is a new idea
40
+
41
+ Agent memory is one of the most crowded categories in the field: mem0, Zep, Letta and Cognee on the infrastructure side, Cursor Memories and Devin Knowledge built into the coding agents, AGENTS.md and CLAUDE.md as the convention nearly everyone already uses. "Give the agent your team's knowledge" is the consensus position, not an insight, and if semantic recall over what an agent has seen is the problem you have, those tools are better at it than anything I would build.
42
+
43
+ The gap I kept hitting was invalidation. Every one of those systems accumulates: records get added, ranked, retrieved. None of them knows that a rule went stale because the code it governs moved on without it. So the bet here is narrow and specific, and it is the subject of part 2: a rule cites the source it governs, carries a verification receipt, decays on a clock, and fails CI when the code changes after the rule was last checked. Records treated like tests, not like notes.
44
+
45
+ ## Proving it does something
46
+
47
+ A tool that feels helpful and a tool that is helpful are different claims, so I built an A/B harness and committed the results.
48
+
49
+ Eight tasks, each written to invite a specific defect class the store already covers. The same model in both arms, the retrieved rules as the only variable. A second, different model grading blind, never told which arm produced the code. Failure counts printed before any score, because a metric that cannot report its own failure rate is not a metric.
50
+
51
+ With none of the rules in context, the agent reproduced a known defect class in 33 percent of runs. With them, 4 percent. On the tasks the baseline actually failed, runs with the rules in context reproduced the defect in 1 of 15. One task, an unpinned linter version in CI, failed three times out of three without the rules and never with them.
52
+
53
+ The result I did not expect: a budget model with the rules in context made roughly a third the known mistakes of a frontier model without them on identical tasks. Context bought more than the model upgrade did, which is the more useful number if you are routing work between cheap and expensive models.
54
+
55
+ ## The limits, stated
56
+
57
+ Those tasks were authored from rules already in my own store, so the experiment measures prevention of defect classes the store covers rather than general code quality. Small n. Directional, not a benchmark. The full method and every receipt are in the repository, including the runs that went the other way.
58
+
59
+ The system also found its own first defect: an early retrieval marked 28 of 33 stored notes as relevant to a single task. Ranking put the right records first, but nothing trimmed the tail. Subject tags and per-repo interest declarations cut it to 20; a proper relevance cutoff is still open work, recorded in the store as a defect against the store. The failure log lives inside the system it describes.
60
+
61
+ That is the first idea. The state in most agent-loop diagrams lives inside one execution. The compounding win is memory that outlives it: record the rule once, retrieve it into every future task that resembles it, across every project. Run N+1 should not repeat run N's mistake, and neither should the next repository.
62
+
63
+ Next: the verify step, and why a step's own report of success is not evidence.
64
+
65
+ *The model generates. The loop governs. The standard holds.*
@@ -0,0 +1,45 @@
1
+ # Don't let a step grade itself
2
+
3
+ *Part 2 of 3: The loop that learns. How one developer's AI loop remembers, verifies, and enforces.*
4
+
5
+ Never trust an exit code. This is the incident behind that rule.
6
+
7
+ A pipeline step was supposed to materialize 238 files. It produced none. Every window raised, the exceptions were swallowed inside a worker pool, and the process exited zero. Downstream steps ran happily against an empty directory. The fix was trivial once seen: count the outputs rather than read the exit status. The discipline it bought has outlasted the project: a step reporting success and the work being done are two different facts, and only one of them is evidence.
8
+
9
+ That rule matters more when an agent is writing the steps. Review catches things, but review is sampling. What makes generated code safe at volume is the checks that run whether or not anyone is paying attention, and those checks are only as good as the question they actually ask.
10
+
11
+ ## The same bug wears several costumes
12
+
13
+ A dependency's major-version upgrade compiled clean and carried three runtime breaking changes, one of which silently broke transactional atomicity. The build was green. Green answered "does this compile," which was not the question anyone cared about.
14
+
15
+ In an eval harness, an LLM judge printed a perfect 5.0 out of 5.0 while 19 of 20 cases had crashed, because it averaged the survivors. The score was real. It was also meaningless, and nothing in the output said so.
16
+
17
+ Different costumes, same failure: the verifier reported on itself, and the report was wrong.
18
+
19
+ ## The grader is usually ls, not an LLM
20
+
21
+ When I say don't let a step grade itself, people sometimes hear "put a second model in front of every step." Almost never. The correct grader for the 238-file failure was a file count. The correct grader for the upgrade was an integration test that forced a rollback and asserted no event escaped. Boring, deterministic checks, placed at the boundaries where the loop makes a decision: mark done, merge, deploy.
22
+
23
+ A second model earns its place in exactly one case: when the verify signal is itself a judgment, like scoring summary quality or reviewing generated code. Only there does judge-must-differ-from-generator apply, because a judgment shares blind spots with whoever produced it. A file count has no blind spots.
24
+
25
+ ## Four rungs of independence
26
+
27
+ In the knowledge layer from part 1, verification is not one feature. It is four rungs, each answering the same question while trusting the claimant less.
28
+
29
+ **Rung 1: assertion, quarantined.** A "verified" flag that whoever writes the record can set is a step grading itself. The flag still exists, but only for importing historical receipts. Live verification refuses it.
30
+
31
+ **Rung 2: an observed check with a stored evidence trail.** The verify command takes a record and a check to run. It runs the check, reads the real exit status, and when you name an expected success signal it requires that signal in the output, because exit zero alone is precisely what failed on those 238 files. Only an observed pass stamps the record, and the command, result, and timestamp are stored on the record itself. The claimant still chooses the check, and a weak check is still possible, but the stamp shows exactly which check ran. "Verified by: true" is visible to every future reader and invites the obvious question. Assertion hides its emptiness; evidence exposes it.
32
+
33
+ **Rung 3: an independent actor re-runs the checks.** CI runs the drift gate and the repository's mechanical gates on every pull request: a different grader, at a different time, with no stake in the original claim. When a gate proves itself by failing against real drift, the receipt is written by the job that watched it happen, not by the gate's author.
34
+
35
+ **Rung 4: time attacks every stamp.** Each rule declares the files it governs. The moment those files change after the rule was last verified, the drift detector flags it, because a stale rule is a rule nobody re-checked. Verification also decays on a clock: a stamp nobody re-earns is demoted rather than left quietly trusted. And the system is scored on an outcome it cannot flatter itself on: recurrence after arming, the count of defect classes that came back even though a catching check existed. Not "how many rules were recorded," which measures activity, but "did the mistakes stop repeating," which measures the point.
36
+
37
+ ## Where judgment stays
38
+
39
+ The remaining self-grading surface is check selection. Nothing stops someone from verifying a record against a trivially weak check. I left that to human judgment deliberately and made it inspectable instead of trying to mechanize taste: every stamp displays its evidence, so a weak check becomes a visible artifact rather than an invisible belief. Mechanize what can be mechanized, audit the rest.
40
+
41
+ One detail worth the build: the rule stating that verification comes from observed checks and never from assertion lives in the store as a record, and it was the first record verified by the verify command itself, evidence trail and all. The system's rules are subject to the system.
42
+
43
+ Next: enforcement, and why a rule that is not wired to something that runs is a preference.
44
+
45
+ *The model generates. The loop governs. The standard holds.*
@@ -0,0 +1,37 @@
1
+ # Enforcement, or good intentions
2
+
3
+ *Part 3 of 3: The loop that learns. How one developer's AI loop remembers, verifies, and enforces.*
4
+
5
+ A rule that is not wired to something that runs is a preference, not a standard.
6
+
7
+ I know that from auditing my own work. In one codebase an architecture-review agent was documented across six files as a live enforcement surface. When I audited what actually executed, meaning the CI jobs, the hooks, and the pre-commit path, that agent appeared nowhere. Every one of those documents was accurate about the design and wrong about the system.
8
+
9
+ I have now found the same shape in three codebases, which makes it a pattern rather than an oversight: if a rule only lives in the prompt, sometimes the agent applies it, sometimes I restate it, and neither of those is a control. Probabilistic compliance is still non-compliance when the point is a guarantee.
10
+
11
+ ## What mechanical looks like
12
+
13
+ In the loop from parts 1 and 2, every rule that matters is wired to something that runs without anyone deciding to run it.
14
+
15
+ **The read is enforced, fail closed.** The moment a task is submitted, a hook retrieves the relevant rules and decisions and injects them into the model's context before any work begins. If the store is unreachable, the hook blocks the task rather than proceeding, because "nothing applies here" and "I could not reach the store" look identical from the outside and only one of them is safe. Silence is never reported as safety. There is an explicit override for working offline, explicit on purpose: proceeding without the standards should be a decision someone made, not a default they fell into.
16
+
17
+ One detail that cost me a rebuild. The first version was wired to an event whose output the model never receives. The hook fired on every edit, retrieved exactly the right rules, and wrote them to a channel nothing read, while every unit test passed, because those tests proved the hook ran and produced correct text. Only a behavioral test against a bare control repository exposed it: same task, same model, defect reproduced in both arms. A hook that fires is not a hook that is heard. Test the channel, not the firing.
18
+
19
+ **The write is prompted at the only moment it can work.** What a session learns dies with it, so a stop hook blocks the first stop of any session that changed files and asks one question: did this session produce something worth keeping? Either it is recorded with a deliberate scope and subject, or the session states that nothing durable came out of it and finishes. It fires once per session and never loops. The first session running that hook was caught by it and had to record three rules it had applied but never encoded. A rule that catches its own author on day one is at least running.
20
+
21
+ **The gates run at merge time.** Retired identifiers cannot resurface. Retracted claims cannot be restated. Documentation cannot silently orphan. The instructions file cannot bloat past its budget. Each is a small script that fails CI, and none of them require anyone to remember anything.
22
+
23
+ ## Installation is where enforcement quietly dies
24
+
25
+ Auditing my own installer turned up the same failure one layer down: it did half the wiring and printed instructions for the rest. Install the hook, then "add it to your settings file." Copy this CI workflow. Register this tool. Four printed instructions, four surfaces that would never run for anyone who skipped the homework, which over time is everyone including me.
26
+
27
+ So the rule became: init must wire, not instruct. It registers the hooks in the agent's settings itself, idempotently, preserving whatever is already there. It installs the CI workflow instead of describing it. It registers the agent tools when the dependency exists and deliberately refuses when it does not, because a registration pointing at a missing dependency is a broken tool, which is worse than none. The only things it prints are the things it genuinely cannot do, and it says those loudly: this is not a git repository, so the drift gate is disabled until it is.
28
+
29
+ The same failure again at the environment layer. Hooks run in whatever process the agent harness spawns, which routinely lacks the PATH of the shell you developed in. A hook that works on the author's machine and silently no-ops on everyone else's is the "surface nobody runs" bug wearing an environment-variable costume. The fix is a layered resolver: an explicit override, then a path pinned at install time by the shell where the tool demonstrably works, then PATH, then module execution by any interpreter that can import the package. The enforcement hook fails closed with instructions when nothing resolves; the reminder hook silently disables, because a best-effort prompt must never brick someone's session. Degrade by role.
30
+
31
+ ## Where the trust boundary sits
32
+
33
+ The stop hook can be lied to. A session can answer "nothing durable" untruthfully and finish. I left that boundary deliberately, the same way a ship-moment reminder is a reminder and not a block: the mechanical part is putting the question in front of you at the right moment, every time. Answering it honestly stays a person's job. The aim is not to remove judgment, it is to guarantee judgment gets exercised at the moments that matter, with the relevant evidence already in view.
34
+
35
+ That is the series. Memory that outlives the run, verification that does not trust a step's own report, and enforcement that runs without being remembered. None of the three is exotic alone. The compounding comes from wiring them together: the enforced read puts the standard in front of the agent, the verification rungs keep the standard honest, and the recurrence metric reports whether any of it worked, in the only currency that counts: mistakes that stopped repeating.
36
+
37
+ *The model generates. The loop governs. The standard holds.*
@@ -13,6 +13,12 @@ cheaper model perform at or above an unbriefed frontier model on those defect cl
13
13
 
14
14
  ## 2. How to read these numbers
15
15
 
16
+ **On the word "briefed."** The two arms are named `baseline` and `briefed` in the
17
+ committed result files, so this report uses those names to stay aligned with the data.
18
+ Mechanically, *briefed* means: before generation, the relevant past lessons were
19
+ retrieved from the store, filtered by scope and subject, and placed in the model's
20
+ context. Nothing else differs between the arms.
21
+
16
22
  **A "run" is one complete attempt:** the generator gets one task (with or without the
17
23
  briefing) and produces code; the blind judge then answers one question — does this code
18
24
  contain the specific known mistake this task was designed to invite?
@@ -6,7 +6,7 @@ build-backend = "hatchling.build"
6
6
  # Distribution name only. The CLI command, the import package, and the repo are all `okl`;
7
7
  # PyPI rejects `okl` as too similar to the existing `oki` (l/i are treated as confusable).
8
8
  name = "org-knowledge-layer"
9
- version = "0.1.0"
9
+ version = "0.1.2"
10
10
  description = "Org Knowledge Layer — an installable sixth surface that carries encoded engineering lessons across repos."
11
11
  readme = "README.md"
12
12
  requires-python = ">=3.10"
@@ -1,5 +1,5 @@
1
1
  {
2
- "_comment": "Seed for the OKL — React/frontend-SPA lessons ported from the .NET platform's frontend canon (frontend/CLAUDE.md + the vendored vercel-react-best-practices skill, frozen 2026-06-12). React is backend-agnostic, so these are org-scoped: they propagate to any repo with a React UI. This closes the coverage gap the A/B experiment measured — react_fetch reproduced the useEffect-fetching defect 3/3 in BOTH arms because no React node existed in the store to inject.",
2
+ "_comment": "Seed for the OKL — React/frontend-SPA lessons ported from the author's own frontend canon (frontend/CLAUDE.md, frozen 2026-06-12); external guides are cited in found_by as provenance, not reproduced. React is backend-agnostic, so these are org-scoped: they propagate to any repo with a React UI. This closes the coverage gap the A/B experiment measured — react_fetch reproduced the useEffect-fetching defect 3/3 in BOTH arms because no React node existed in the store to inject.",
3
3
  "nodes": [
4
4
  {
5
5
  "key": "rx_useeffect_fetch",
@@ -14,5 +14,5 @@ Installed to `.claude/rules/`:
14
14
  (BFF trade-off documented), MSW + Playwright testing.
15
15
 
16
16
  Reference stack: Vite + React 19 + TS strict, CSR SPA, TanStack Query/Router, Zustand, Tailwind v4 +
17
- shadcn/ui, oidc-client-ts → Keycloak. Deep reference: the vendored `vercel-react-best-practices`
17
+ shadcn/ui, oidc-client-ts → Keycloak. These rules are original to this kit; pair them with any external React guide
18
18
  skill (70 rules), where the canon wins on any disagreement.
@@ -9,8 +9,8 @@ paths: ["frontend/**/*.tsx", "frontend/**/*.ts", "**/frontend/**/*.tsx", "**/fro
9
9
  > storefront) — but backend-agnostic: stack these onto ANY backend profile (dotnet, python-rag, …)
10
10
  > whose repo has a React frontend. Reference stack: Vite + React 19 + TypeScript (strict), CSR SPA;
11
11
  > TanStack Query v5 + Router; Zustand (small UI globals only); Tailwind v4 + shadcn/ui; oidc-client-ts
12
- > → Keycloak (auth-code + PKCE); React Compiler on. Deep reference: the vendored
13
- > `vercel-react-best-practices` skill (70 rules) — **but this canon wins where they disagree.**
12
+ > → Keycloak (auth-code + PKCE); React Compiler on. These rules are original to this kit;
13
+ > where you also consult an external React best-practices guide, **this canon wins where they disagree.**
14
14
  >
15
15
  > Where a rule cites "the backend" (VSA feature folders, server-controlled fields, cache-in-write-path,
16
16
  > measure-before-optimizing), the parallel holds against whatever backend you pair this with.
@@ -1,43 +0,0 @@
1
- # Memory that outlives the run
2
-
3
- *Part 1 of 3: The loop that learns. How our AI dev loop remembers, verifies, and enforces.*
4
-
5
- I noticed something embarrassing while auditing three of my own codebases: they had all learned the same lesson, separately, the hard way.
6
-
7
- One is a .NET microservices platform. One is a geospatial ML pipeline. One is a Python RAG service. In each of them, at some point, I had built a review rule or a checklist, described it in the docs as a live enforcement surface, and then discovered months later that nothing had ever run it. Three repos, three independent discoveries, one lesson: a surface nobody runs is documentation, not enforcement.
8
-
9
- The lesson wasn't the embarrassing part. The embarrassing part was the *three times*. Each repo had a careful CLAUDE.md, encoded rules, a real engineering method. And none of that knowledge could cross the repo boundary. Every project was re-deriving the org's hard-won lessons from scratch, one incident at a time.
10
-
11
- I had been thinking of my setup as "two readers, one canon": a human reads the rules, an AI agent reads the rules, and the two must not drift apart. What I actually had was N readers and N canons.
12
-
13
- ## The fix is a layer, not a bigger prompt
14
-
15
- The obvious move is to paste every lesson into every prompt. That fails fast, and it fails in a specific way: always-on context taxes every session whether the lesson is relevant or not. My .NET repo enforces a hard size budget on its instructions file in CI (warning at 400 lines, build failure at 500) precisely because every byte of always-on context is cognitive overhead for both readers.
16
-
17
- So the design question became: where does each piece of knowledge belong? I ended up with a three-way triage, decided by three questions asked in order.
18
-
19
- **Does a session doing a totally unrelated task still need this?** Then it goes in the always-on instructions. Almost nothing passes this test. Naming conventions, the security posture, the error-handling canon. This is the most expensive real estate in the system, and the CI size budget keeps it honest.
20
-
21
- **Would I want this to appear whenever a future task resembles it?** Then it goes in the knowledge store. This is the interesting category: the rate-limiter lesson that matters enormously to the three tasks a year that touch rate limiting, and is pure noise for everything else.
22
-
23
- **Does it stop mattering once the PR merges?** Then it stays with the task and is allowed to die. The subtle part is extraction: a dead plan often contains one durable decision worth pulling out before you let go of the rest.
24
-
25
- The store is the part I had to build. It's a small typed database of lessons: defects with symptom, cause, and fix. Rules. Deliberate decisions, so they don't get silently reversed. Tombstones for identifiers that must never come back. Retractions for claims that turned out to be false. Each note carries a scope: org-wide lessons propagate to every connected repo, repo-scoped quirks stay home. Choosing that scope is a human judgment, and it's what keeps a shared layer from filling up with one project's noise.
26
-
27
- ## Retrieved, not loaded
28
-
29
- The reason the store can grow forever is that no run ever reads it. Before a task starts, a hook searches the store with the task description and injects only what survives four filters: lexical relevance ranking, scope (another repo's quirks are structurally invisible), declared interests (my Python repo says it cares about eval integrity and retrieval design, so React lessons never appear no matter how well they match), and typed routing that turns the survivors into a short action list. Fix this, run that gate, don't restate this retracted claim.
30
-
31
- Each run sees a dozen relevant lines, not the whole library. At ten times the current size the briefing should be the same length, just better chosen.
32
-
33
- Does it work? I measured it the way I'd measure anything: a held-fixed A/B, same model both arms, briefing injected versus not, scored by a blind judge that was deliberately a different model from the generator. In the original experiment, defect reproduction on covered tasks dropped from 50 percent to 6 percent, and from 75 percent to 8 percent on the tasks where the baseline model actually failed. I later rebuilt that experiment as a checked-in harness anyone can re-run, and the fresh receipt (48 sampled runs, zero harness failures) tells a sharper story: today's stronger models dodge much of the bait unaided, but the baseline still reproduced known defect classes in 33 percent of runs, and the briefed arm in 4 percent. On the tasks the baseline failed at least once, briefed runs reproduced in 1 of 15. One task always failed unaided, three out of three, and never failed briefed. The value concentrates exactly where models still get things wrong, which is exactly what an org's hard-won lessons describe. And the misses in the original run were both coverage gaps, a lesson not yet in the store, so the check doubles as a detector for what to encode next.
34
-
35
- ## The honest part
36
-
37
- The first real dogfood run also found the system's first defect, in itself: the briefing for one specific task marked 28 of the 33 stored notes as relevant. Ranking put the right lessons first, but nothing cut the tail. Subject tags plus per-repo interest declarations cut it to 20 of 33, and a proper relevance cutoff is still open work, recorded in the store as a defect against the store, status: narrowed. The system's failure log lives inside the system.
38
-
39
- That's the whole first idea. The state in most loop diagrams only lives inside one execution. The compounding win is memory that outlives the run: record the lesson once, retrieve it into every future task that resembles it, across every repo. Run N+1 should never repeat run N's mistake.
40
-
41
- Next up: the verify step, and why you should never let a step grade itself.
42
-
43
- *The model generates, the loop governs, and the team learns.*
@@ -1,37 +0,0 @@
1
- # Don't let a step grade itself
2
-
3
- *Part 2 of 3: The loop that learns. How our AI dev loop remembers, verifies, and enforces.*
4
-
5
- I had a pipeline step whose job was to write 238 files. It failed on every single one, swallowed the exceptions inside a worker pool, and exited with a success code. Everything downstream ran happily on an empty folder. By the time I noticed, I was debugging the wrong end of the pipeline.
6
-
7
- That incident bought me a rule I now apply everywhere: a step saying it succeeded and the work actually being done are two different facts, and the loop has to check the second one. Exit codes, green builds, and tests passing on invented fixtures all answer a narrower question than the one the loop thinks it's asking.
8
-
9
- The same class of failure, in fancier clothes: a dependency's major-version upgrade compiled clean and carried three runtime breaking changes, one of which silently broke transactional atomicity. And in an eval harness, my LLM judge printed a perfect 5.0 out of 5.0 while 19 of 20 cases had crashed, because it averaged only the survivors. Different costumes, same bug. The verifier reported on itself, and the report was wrong.
10
-
11
- ## The grader is usually ls, not an LLM
12
-
13
- When I say don't let a step grade itself, people sometimes hear "every step needs a second model reviewing it." Almost never. The correct grader for my 238-file disaster was a file count. The correct grader for the upgrade was an integration test that forced a rollback and asserted no event escaped. Boring, deterministic checks, placed at the boundaries where the loop makes a decision: mark done, merge, deploy.
14
-
15
- A second model enters in exactly one case: when the verify signal is itself a model's judgment, like grading summary quality or reviewing generated code. Only there does judge-must-differ-from-generator apply, because a judgment shares blind spots with whoever produced it. A file count has no blind spots.
16
-
17
- ## Four rungs of independence
18
-
19
- In the knowledge layer I described in part 1, verification isn't one feature. It's four escalating rungs, each answering the same question with less trust in the claimant.
20
-
21
- **Rung 1: assertion, quarantined.** A "verified" flag that whoever writes the record can set is a step grading itself. The system still has that flag, but only for importing historical receipts. Live verification refuses it.
22
-
23
- **Rung 2: an observed check with a stored evidence trail.** The verify command takes a node and a check to run. It runs the check itself, reads the real exit code, and, if you name an expected success signal, requires that signal to appear in the output, because exit zero alone is exactly what burned me on those 238 files. Only on an observed pass does it stamp the record, and it stores the command, the result, and the timestamp on the record itself. Yes, the claimant still chooses the check, and a lazy check is still possible. But the stamp records exactly which check was run. "Verified by: true" is visible to every future reader and invites the obvious question. Assertion hides its emptiness; evidence exposes it.
24
-
25
- **Rung 3: an independent actor re-runs the checks.** CI runs the drift gate and the repo's mechanical gates on every pull request. A different grader, at a different time, with no stake in the original claim. When a gate proves itself by actually failing against real drift, the receipt is written by the job that watched it happen, not by the gate's author.
26
-
27
- **Rung 4: time attacks every stamp.** Each rule declares which files it governs. The moment those files change after the rule was last verified, the drift detector flags it: a stale rule is a rule nobody re-checked. Verification also decays on a clock; a stamp nobody re-earns demotes to stale rather than staying quietly trusted. And the whole system is scored on an outcome it cannot flatter itself on: recurrence after arming, the count of defect classes that came back even though a catching check existed. Not "how many lessons did we record," which measures activity, but "did the mistakes stop repeating," which measures the point.
28
-
29
- ## The honest part
30
-
31
- The remaining self-grading surface is check selection. Nothing stops someone from verifying a record against a trivially weak check. I chose to leave that to human judgment and make it inspectable rather than trying to mechanize taste: every stamp shows its evidence, and a weak check is now a visible, embarrassing artifact instead of an invisible belief. Mechanize what can be mechanized. Make the rest auditable.
32
-
33
- One detail I enjoy: the rule that says "verification stamps come from observed checks, never from assertion" lives in the store as a record, and it was the first record verified by the verify command itself, evidence trail and all. The system's rules are subject to the system.
34
-
35
- Next up: enforcement, and why a checklist nobody is forced to run is documentation.
36
-
37
- *The model generates, the loop governs, and the team learns.*
@@ -1,33 +0,0 @@
1
- # Enforcement, or good intentions
2
-
3
- *Part 3 of 3: The loop that learns. How our AI dev loop remembers, verifies, and enforces.*
4
-
5
- The most instructive bug in my method wasn't in code. It was in the documentation of an agent that reviewed architecture. Six separate documents described it as a live enforcement surface. It had been invoked by nothing, ever. Every one of those documents was telling the truth about the design and lying about the system.
6
-
7
- I have now learned this lesson in three different codebases, which is what finally made it stick: if a rule only lives in the prompt, sometimes the agent picks it up on its own, sometimes I have to remind it, and neither of those is a system. Probabilistic compliance from a model is still non-compliance from a control standpoint. Enforcement has to be mechanical or it's just good intentions.
8
-
9
- ## What mechanical looks like
10
-
11
- In the loop I've been describing across this series, every rule that matters is wired to something that runs without anyone deciding to run it.
12
-
13
- **The read is enforced, fail closed.** The moment a task is submitted, a hook retrieves the relevant lessons from the knowledge store and injects them into the model's context, before any work begins. If the store is unreachable, the hook blocks the task instead of proceeding, because "no lessons apply" and "I couldn't reach the lessons" look identical from the outside and only one of them is safe. Silence is never reported as safety. There's an explicit override for working offline, and it's explicit on purpose: proceeding blind should be a decision someone made, not a default someone fell into. One hard-won detail: my first version wired this to the wrong event, one whose output the model never sees. The hook fired on every edit, retrieved the right lessons, and printed them into a void, while every unit test passed. Only a behavioral test against a bare control repo exposed it. A hook that fires is not a hook that's heard. Test the channel, not the firing.
14
-
15
- **The write is prompted at the only moment it can work.** Lessons die when sessions end, so a stop hook blocks the first stop of any session that changed files and asks one question: did this session learn something worth keeping? Either the lesson gets recorded, with its scope and tags chosen deliberately, or the session states that nothing durable was learned and finishes. It fires once per session and never loops. The first session running that hook got caught by it and had to record three lessons it had fixed but never encoded. The hook worked on its own author, which is the only referee a rule like this will accept.
16
-
17
- **The gates run at merge time.** Retired identifiers can't resurface, retracted claims can't be restated, docs can't silently orphan, the instructions file can't bloat past its budget. Each is a small script that fails CI. None of them require anyone to remember anything.
18
-
19
- ## Installation is where enforcement quietly dies
20
-
21
- Here's the failure mode I only caught by auditing my own installer: it did half the wiring and printed instructions for the rest. Install the hook, then "add it to your settings file." Copy this CI workflow. Register this tool. Four printed instructions, four surfaces that would never run for anyone who didn't do the homework, which over time is everyone.
22
-
23
- So the installer's rule became: init must wire, not instruct. It registers the hooks in the agent's settings itself, idempotently, preserving whatever else is there. It installs the CI workflow instead of describing it. It registers the agent tools when the dependency exists, and deliberately refuses when it doesn't, because a registration pointing at a missing dependency is a broken tool, which is worse than none. The only things it's allowed to print are the things it genuinely cannot do, and it says those loudly: this isn't a git repository, so the drift layer is disabled until it is.
24
-
25
- One layer deeper, same lesson: hooks run in whatever environment the agent harness spawns, which routinely lacks your dev shell's PATH. A hook that works on the author's machine and silently no-ops on everyone else's is the "surface nobody runs" bug wearing an environment-variables costume. The fix is a layered resolver: an explicit override, then a path pinned at install time by the shell where the tool demonstrably works, then PATH, then module execution by any interpreter that can import the package. The enforcement hook fails closed with instructions when nothing resolves; the reminder hook silently disables, because a best-effort prompt must never brick someone's sessions. Degrade by role.
26
-
27
- ## The honest part
28
-
29
- The stop hook can be lied to. A session can answer "nothing durable learned" untruthfully and finish. I left that trust boundary in on purpose, the same way the ship-moment reminder in my .NET repo is a reminder and not a block: the mechanical part is putting the question in front of you at the right moment, every time. Answering it honestly stays a human's job. The system's aim isn't to remove judgment. It's to make sure judgment is exercised at the moments it matters, with the relevant evidence in view, instead of whenever someone happens to remember.
30
-
31
- That's the series. Memory that outlives the run, verification that doesn't trust the step's own report, and enforcement that runs without being remembered. None of the three is exotic on its own. The compounding comes from wiring them together: the enforced read injects the memory, the verify rungs keep the memory honest, and the recurrence metric tells you whether any of it is working, in the one currency that matters, mistakes that stopped repeating.
32
-
33
- *The model generates, the loop governs, and the team learns.*