@mrciphersmith/keryx 0.3.1 → 0.3.3
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/cli.js +7310 -2471
- package/dist/core.js +116 -10
- package/package.json +1 -1
- package/src/gdskills/bundled/install-manifest.json +349 -2
- package/src/gdskills/bundled/rules/core/model-selection.mdc +18 -0
- package/src/gdskills/bundled/skills/orchestration/job-orchestrator/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/planning/brainstorm/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/planning/interviewer/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/quality/deploy/SKILL.md +1 -1
- package/src/gdskills/bundled/skills/review/review-jev-comments/SKILL.md +184 -0
- package/src/gdskills/bundled/skills/review/review-jev-contract/SKILL.md +193 -0
- package/src/gdskills/bundled/skills/review/review-jev-docs/SKILL.md +189 -0
- package/src/gdskills/bundled/skills/review/review-jev-risk/SKILL.md +190 -0
- package/src/gdskills/bundled/skills/review/review-jev-scenarios/SKILL.md +187 -0
- package/src/gdskills/bundled/skills/review/review-orchestrator/SKILL.detail.md +88 -15
- package/src/gdskills/bundled/skills/review/review-orchestrator/SKILL.md +4 -4
- package/src/gdskills/bundled/stacks/c-cpp/agent-refs.json +4 -0
- package/src/gdskills/bundled/stacks/c-cpp/governance/eval.json +1777 -0
- package/src/gdskills/bundled/stacks/c-cpp/governance/scout.json +31 -0
- package/src/gdskills/bundled/stacks/c-cpp/pack.json +42 -0
- package/src/gdskills/bundled/stacks/c-cpp/rules/coding-style.mdc +80 -0
- package/src/gdskills/bundled/stacks/c-cpp/rules/patterns.mdc +87 -0
- package/src/gdskills/bundled/stacks/c-cpp/rules/security.mdc +90 -0
- package/src/gdskills/bundled/stacks/c-cpp/rules/testing.mdc +83 -0
- package/src/gdskills/bundled/stacks/c-cpp/skills/c-cpp-build-fix/SKILL.md +153 -0
- package/src/gdskills/bundled/stacks/c-cpp/skills/c-cpp-build-fix/evals.json +74 -0
- package/src/gdskills/bundled/stacks/c-cpp/skills/c-cpp-code-review/SKILL.md +132 -0
- package/src/gdskills/bundled/stacks/c-cpp/skills/c-cpp-code-review/evals.json +73 -0
- package/src/gdskills/bundled/stacks/c-cpp/skills/c-cpp-implementation/SKILL.md +151 -0
- package/src/gdskills/bundled/stacks/c-cpp/skills/c-cpp-implementation/evals.json +74 -0
- package/src/gdskills/bundled/stacks/c-cpp/skills/c-cpp-testing/SKILL.md +152 -0
- package/src/gdskills/bundled/stacks/c-cpp/skills/c-cpp-testing/evals.json +74 -0
- package/src/gdskills/bundled/stacks/ci-github-gitlab/agent-refs.json +4 -0
- package/src/gdskills/bundled/stacks/ci-github-gitlab/governance/eval.json +1295 -0
- package/src/gdskills/bundled/stacks/ci-github-gitlab/governance/scout.json +26 -0
- package/src/gdskills/bundled/stacks/ci-github-gitlab/pack.json +41 -0
- package/src/gdskills/bundled/stacks/ci-github-gitlab/rules/patterns.mdc +77 -0
- package/src/gdskills/bundled/stacks/ci-github-gitlab/rules/security.mdc +144 -0
- package/src/gdskills/bundled/stacks/ci-github-gitlab/skills/ci-pipeline-build-fix/SKILL.md +121 -0
- package/src/gdskills/bundled/stacks/ci-github-gitlab/skills/ci-pipeline-build-fix/evals.json +73 -0
- package/src/gdskills/bundled/stacks/ci-github-gitlab/skills/ci-pipeline-code-review/SKILL.md +139 -0
- package/src/gdskills/bundled/stacks/ci-github-gitlab/skills/ci-pipeline-code-review/evals.json +73 -0
- package/src/gdskills/bundled/stacks/ci-github-gitlab/skills/ci-pipeline-implementation/SKILL.md +147 -0
- package/src/gdskills/bundled/stacks/ci-github-gitlab/skills/ci-pipeline-implementation/evals.json +74 -0
- package/src/gdskills/bundled/stacks/docker-k8s-terraform/agent-refs.json +4 -0
- package/src/gdskills/bundled/stacks/docker-k8s-terraform/governance/eval.json +865 -0
- package/src/gdskills/bundled/stacks/docker-k8s-terraform/governance/scout.json +16 -0
- package/src/gdskills/bundled/stacks/docker-k8s-terraform/pack.json +46 -0
- package/src/gdskills/bundled/stacks/docker-k8s-terraform/rules/coding-style.mdc +74 -0
- package/src/gdskills/bundled/stacks/docker-k8s-terraform/rules/patterns.mdc +81 -0
- package/src/gdskills/bundled/stacks/docker-k8s-terraform/rules/security.mdc +146 -0
- package/src/gdskills/bundled/stacks/docker-k8s-terraform/rules/testing.mdc +61 -0
- package/src/gdskills/bundled/stacks/docker-k8s-terraform/skills/docker-k8s-terraform-build-fix/SKILL.md +151 -0
- package/src/gdskills/bundled/stacks/docker-k8s-terraform/skills/docker-k8s-terraform-build-fix/evals.json +74 -0
- package/src/gdskills/bundled/stacks/docker-k8s-terraform/skills/docker-k8s-terraform-review/SKILL.md +135 -0
- package/src/gdskills/bundled/stacks/docker-k8s-terraform/skills/docker-k8s-terraform-review/evals.json +76 -0
- package/src/gdskills/bundled/stacks/php-laravel/agent-refs.json +4 -0
- package/src/gdskills/bundled/stacks/php-laravel/governance/eval.json +1829 -0
- package/src/gdskills/bundled/stacks/php-laravel/governance/scout.json +33 -0
- package/src/gdskills/bundled/stacks/php-laravel/pack.json +41 -0
- package/src/gdskills/bundled/stacks/php-laravel/rules/coding-style.mdc +82 -0
- package/src/gdskills/bundled/stacks/php-laravel/rules/patterns.mdc +80 -0
- package/src/gdskills/bundled/stacks/php-laravel/rules/security.mdc +80 -0
- package/src/gdskills/bundled/stacks/php-laravel/rules/testing.mdc +82 -0
- package/src/gdskills/bundled/stacks/php-laravel/skills/php-laravel-build-fix/SKILL.md +143 -0
- package/src/gdskills/bundled/stacks/php-laravel/skills/php-laravel-build-fix/evals.json +74 -0
- package/src/gdskills/bundled/stacks/php-laravel/skills/php-laravel-code-review/SKILL.md +126 -0
- package/src/gdskills/bundled/stacks/php-laravel/skills/php-laravel-code-review/evals.json +76 -0
- package/src/gdskills/bundled/stacks/php-laravel/skills/php-laravel-implementation/SKILL.md +140 -0
- package/src/gdskills/bundled/stacks/php-laravel/skills/php-laravel-implementation/evals.json +75 -0
- package/src/gdskills/bundled/stacks/php-laravel/skills/php-laravel-testing/SKILL.md +124 -0
- package/src/gdskills/bundled/stacks/php-laravel/skills/php-laravel-testing/evals.json +74 -0
- package/src/gdskills/bundled/stacks/ruby-rails/agent-refs.json +4 -0
- package/src/gdskills/bundled/stacks/ruby-rails/governance/eval.json +1673 -0
- package/src/gdskills/bundled/stacks/ruby-rails/governance/scout.json +33 -0
- package/src/gdskills/bundled/stacks/ruby-rails/pack.json +42 -0
- package/src/gdskills/bundled/stacks/ruby-rails/rules/coding-style.mdc +69 -0
- package/src/gdskills/bundled/stacks/ruby-rails/rules/patterns.mdc +93 -0
- package/src/gdskills/bundled/stacks/ruby-rails/rules/security.mdc +90 -0
- package/src/gdskills/bundled/stacks/ruby-rails/rules/testing.mdc +89 -0
- package/src/gdskills/bundled/stacks/ruby-rails/skills/ruby-rails-build-fix/SKILL.md +143 -0
- package/src/gdskills/bundled/stacks/ruby-rails/skills/ruby-rails-build-fix/evals.json +73 -0
- package/src/gdskills/bundled/stacks/ruby-rails/skills/ruby-rails-code-review/SKILL.md +134 -0
- package/src/gdskills/bundled/stacks/ruby-rails/skills/ruby-rails-code-review/evals.json +71 -0
- package/src/gdskills/bundled/stacks/ruby-rails/skills/ruby-rails-implementation/SKILL.md +141 -0
- package/src/gdskills/bundled/stacks/ruby-rails/skills/ruby-rails-implementation/evals.json +72 -0
- package/src/gdskills/bundled/stacks/ruby-rails/skills/ruby-rails-testing/SKILL.md +125 -0
- package/src/gdskills/bundled/stacks/ruby-rails/skills/ruby-rails-testing/evals.json +72 -0
- package/src/gdskills/bundled/stacks/sql-db/agent-refs.json +4 -0
- package/src/gdskills/bundled/stacks/sql-db/governance/eval.json +1829 -0
- package/src/gdskills/bundled/stacks/sql-db/governance/scout.json +30 -0
- package/src/gdskills/bundled/stacks/sql-db/pack.json +40 -0
- package/src/gdskills/bundled/stacks/sql-db/rules/coding-style.mdc +69 -0
- package/src/gdskills/bundled/stacks/sql-db/rules/patterns.mdc +134 -0
- package/src/gdskills/bundled/stacks/sql-db/rules/security.mdc +74 -0
- package/src/gdskills/bundled/stacks/sql-db/rules/testing.mdc +83 -0
- package/src/gdskills/bundled/stacks/sql-db/skills/sql-db-build-fix/SKILL.md +147 -0
- package/src/gdskills/bundled/stacks/sql-db/skills/sql-db-build-fix/evals.json +72 -0
- package/src/gdskills/bundled/stacks/sql-db/skills/sql-db-code-review/SKILL.md +132 -0
- package/src/gdskills/bundled/stacks/sql-db/skills/sql-db-code-review/evals.json +73 -0
- package/src/gdskills/bundled/stacks/sql-db/skills/sql-db-implementation/SKILL.md +153 -0
- package/src/gdskills/bundled/stacks/sql-db/skills/sql-db-implementation/evals.json +77 -0
- package/src/gdskills/bundled/stacks/sql-db/skills/sql-db-testing/SKILL.md +129 -0
- package/src/gdskills/bundled/stacks/sql-db/skills/sql-db-testing/evals.json +73 -0
|
@@ -0,0 +1,74 @@
|
|
|
1
|
+
{
|
|
2
|
+
"triggers": {
|
|
3
|
+
"positive": [
|
|
4
|
+
"TokenStream::peek() has zero test coverage right now -- can you set up GoogleTest scenarios covering its end-of-buffer behavior",
|
|
5
|
+
"Add a TEST_F fixture that covers the parser's error paths",
|
|
6
|
+
"CI shows three ParserTest cases going red after my last commit and I can't tell which line broke it",
|
|
7
|
+
"Constructor of ConfigLoader should abort if the path is empty -- how do I write a test that checks the process actually crashes there",
|
|
8
|
+
"How do I actually verify this reported heap-buffer-overflow fix is correct",
|
|
9
|
+
"Add a parameterized test that runs this case across several inputs",
|
|
10
|
+
"Write a regression test for the use-after-free that was just fixed"
|
|
11
|
+
],
|
|
12
|
+
"negative": [
|
|
13
|
+
"Implement the new parser feature itself, not the tests",
|
|
14
|
+
"Review this C++ diff for memory safety issues",
|
|
15
|
+
"Fix the linker error in this CMake build",
|
|
16
|
+
"Write table-driven tests for this Go package",
|
|
17
|
+
"Add pytest cases for this Python module",
|
|
18
|
+
"Fix this failing Rust cargo test"
|
|
19
|
+
]
|
|
20
|
+
},
|
|
21
|
+
"scenarios": [
|
|
22
|
+
{
|
|
23
|
+
"id": "verify-fix-needs-sanitizer",
|
|
24
|
+
"prompt": "Someone on the team says they fixed the reported heap-buffer-overflow in TokenStream::peek() by adding a bounds check. The plain `ctest --test-dir build` run is green. How do I actually confirm the fix is correct before we close the ticket?",
|
|
25
|
+
"strictness": "high",
|
|
26
|
+
"expected_behavior": [
|
|
27
|
+
{
|
|
28
|
+
"grader": "judge",
|
|
29
|
+
"rubric": "A correct answer states that a plain green ctest run does not verify a memory-safety fix, and that confirming it requires rebuilding the test binary with AddressSanitizer (and UndefinedBehaviorSanitizer) enabled and re-running the specific regression test under that sanitizer build to see it pass clean -- not just re-reading the diff or trusting the unsanitized green run.",
|
|
30
|
+
"pass_criteria": [
|
|
31
|
+
"States explicitly that a plain (non-sanitized) green ctest run does not prove a memory-safety fix is correct.",
|
|
32
|
+
"Names rebuilding/re-running under AddressSanitizer (with UndefinedBehaviorSanitizer) specifically -- e.g. a CMake build configured with -fsanitize=address,undefined -- as the concrete verification step, not just 'run more tests'.",
|
|
33
|
+
"States that the specific regression test covering the original heap-buffer-overflow must be the one re-run under the sanitizer build, not an unrelated subset."
|
|
34
|
+
],
|
|
35
|
+
"fail_criteria": [
|
|
36
|
+
"Treats the existing plain green ctest run as sufficient confirmation on its own, with no sanitizer re-run recommended."
|
|
37
|
+
]
|
|
38
|
+
}
|
|
39
|
+
],
|
|
40
|
+
"calibration": {
|
|
41
|
+
"known_right": "A plain ctest run without a sanitizer doesn't prove a memory-safety fix is correct -- ASan-catchable bugs frequently keep passing ordinary assertions right up until the exact byte layout of a future build shifts. Configure a separate sanitizer build (`cmake -S . -B build-asan -DCMAKE_BUILD_TYPE=Debug -DCMAKE_CXX_FLAGS='-fsanitize=address,undefined -fno-omit-frame-pointer -g'`, then `cmake --build build-asan`), and re-run the specific test that covers TokenStream::peek()'s reported overflow (`ctest --test-dir build-asan -R <that test name> --output-on-failure`). Only once that sanitizer build reports the test clean -- no ASan/UBSan diagnostic printed -- is the bounds-check fix actually confirmed; if there's no existing regression test that exercises the original overflow, add one first so this verification is repeatable, not just a one-off manual check.",
|
|
42
|
+
"known_wrong": "If ctest is already green, the fix is good -- the bounds check they added is a standard pattern and the existing test suite already covers TokenStream pretty thoroughly, so there's no need to do anything else before closing the ticket. Sanitizer builds are mostly useful for catching new bugs during development, not for double-checking a fix someone already made and already tested locally.",
|
|
43
|
+
"vague": "Probably worth double-checking the fix a bit more thoroughly before trusting it, maybe with some extra tooling, just to be safe.",
|
|
44
|
+
"subtle_wrong": "Good instinct to double-check -- the right move is to add `assert(index < buffer.size())` right at the start of peek() in addition to the bounds check they already added, then re-run the existing ctest suite. If ctest stays green with the extra assert compiled in, that's a stronger signal than the original fix alone that the bounds are actually being respected on every code path, without needing a separate sanitizer build."
|
|
45
|
+
}
|
|
46
|
+
},
|
|
47
|
+
{
|
|
48
|
+
"id": "concurrent-test-no-sleep",
|
|
49
|
+
"prompt": "I have a GoogleTest case that starts a background thread to process a queue, and I want to assert on the queue's final state once it's done. What's the right way to wait for the thread before asserting?",
|
|
50
|
+
"strictness": "high",
|
|
51
|
+
"expected_behavior": [
|
|
52
|
+
{
|
|
53
|
+
"grader": "judge",
|
|
54
|
+
"rubric": "A correct answer recommends joining the background thread explicitly (std::thread::join, or waiting on a condition variable/future signaled by the worker) before making assertions, and explains that sleeping for a fixed duration instead is non-deterministic and can flake under load rather than actually guaranteeing completion.",
|
|
55
|
+
"pass_criteria": [
|
|
56
|
+
"Recommends a concrete deterministic join mechanism -- std::thread::join, a condition variable wait, or a future/promise -- as the way to wait for the background thread.",
|
|
57
|
+
"States that the join/wait happens before the assertions run, not after or interleaved.",
|
|
58
|
+
"Explains why sleep-based waiting is wrong: it is non-deterministic / can flake under load, not a guarantee the thread has actually finished."
|
|
59
|
+
],
|
|
60
|
+
"fail_criteria": [
|
|
61
|
+
"Recommends std::this_thread::sleep_for (or any fixed-duration sleep) as the way to wait for the background thread to finish before asserting. Mentioning sleep_for only to warn against it is not a failure."
|
|
62
|
+
]
|
|
63
|
+
}
|
|
64
|
+
],
|
|
65
|
+
"calibration": {
|
|
66
|
+
"known_right": "Join the thread explicitly before your assertions: keep the std::thread object around (`std::thread worker(&Queue::process, &queue);`), then call `worker.join();` right before you check the queue's final state -- join blocks until the thread has actually finished, so there's no window where the assertions could run against a still-processing queue. If the work is started some other way (a thread pool, an async task), wait on the matching primitive instead -- a std::future's `.get()`/`.wait()`, or a condition variable the worker signals when it's done. Don't reach for std::this_thread::sleep_for(100ms) 'to be safe' -- it's non-deterministic: it might pass on your machine and flake in CI the moment the worker takes even slightly longer under load, and it silently keeps working even when there's a real race, since it isn't actually a synchronization primitive at all.",
|
|
67
|
+
"known_wrong": "Easiest way is to add a `std::this_thread::sleep_for(std::chrono::milliseconds(200));` right after starting the thread and before your assertions -- that gives the worker plenty of time to finish processing the queue on any reasonable machine, and it's a lot less code than wiring up a condition variable or a future just for a test.",
|
|
68
|
+
"vague": "Make sure the test actually waits for the background thread to be done before checking anything, instead of just assuming it finished.",
|
|
69
|
+
"subtle_wrong": "Call `worker.detach()` instead of `worker.join()` right after starting the thread -- join() risks blocking your test indefinitely if the worker ever hangs, while a detached thread keeps running independently and your test doesn't have to wait on it directly. Just add a `std::this_thread::sleep_for(std::chrono::milliseconds(200))` before your assertions to give the detached worker enough time to finish processing the queue, then check the final state."
|
|
70
|
+
},
|
|
71
|
+
"anti_patterns": ["sleep_for"]
|
|
72
|
+
}
|
|
73
|
+
]
|
|
74
|
+
}
|
|
@@ -0,0 +1,4 @@
|
|
|
1
|
+
{
|
|
2
|
+
"agents": [],
|
|
3
|
+
"note": "no pair: the honest gate (flow 338 Phase B, deepseek:deepseek-chat, --strictness high --trials 10 --scope bundled) ran twice, but only the FIRST run counts as the official result -- the second run's PASS verdicts were disqualified on PR review because the fix pass between the two runs restated failing eval prompts inside SKILL.md description/triggers text (near-copy phrasing, Jaccard >=0.5 against the failing prompts), which is gaming the router, not an honest held-out result. No gate re-run was performed after the disqualified content was reverted. The official (first) run FAILED all three skills on trigger accuracy: ci-pipeline-implementation misroutes on reusable-workflow-sharing, GITHUB_TOKEN-scoping, and deploy-credential-isolation asks (false negatives); ci-pipeline-code-review misroutes on paraphrased permissions-scope and PR-title-script-injection asks (false negatives); ci-pipeline-build-fix misroutes on GitHub's 'Resource not accessible by integration' error text (false negative) and on a Python import error, a TypeScript compile error, a sibling pack's Dockerfile build failure, and a 'write a new workflow' authoring ask (false positives). All behavior scenarios passed cleanly across all three skills. Stays stability: experimental; see governance/eval.json for the full recorded report (rebuilt verbatim from this run's raw output) and W1-stack-catalog.md's batch 6 implementation notes for the accounting."
|
|
4
|
+
}
|