@litfamily/litgrok 1.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.grok/agents/litgrok-executor.md +33 -0
- package/.grok/agents/litgrok-korean-prose-editor.md +32 -0
- package/.grok/agents/litgrok-korean-style-analyzer.md +30 -0
- package/.grok/agents/litgrok-librarian-researcher.md +31 -0
- package/.grok/agents/litgrok-meaning-preservation-auditor.md +30 -0
- package/.grok/agents/litgrok-native-flow-reviewer.md +30 -0
- package/.grok/agents/litgrok-planner.md +31 -0
- package/.grok/agents/litgrok-polish-orchestrator.md +30 -0
- package/.grok/agents/litgrok-qa-runner.md +33 -0
- package/.grok/agents/litgrok-quality-reviewer.md +33 -0
- package/.grok/agents/litgrok-verifier.md +32 -0
- package/.grok/hooks/deliverable-hedge-guard.json +16 -0
- package/.grok/hooks/deliverable-hedge-guard.mjs +148 -0
- package/.grok/hooks/lit-mark.mjs +142 -0
- package/.grok/hooks/plan-gate.mjs +81 -0
- package/.grok/hooks/post-compact.json +15 -0
- package/.grok/hooks/post-compact.mjs +6 -0
- package/.grok/hooks/post-tool-use-failure.json +15 -0
- package/.grok/hooks/post-tool-use-failure.mjs +6 -0
- package/.grok/hooks/post-tool-use.json +16 -0
- package/.grok/hooks/post-tool-use.mjs +7 -0
- package/.grok/hooks/pre-compact.json +15 -0
- package/.grok/hooks/pre-compact.mjs +6 -0
- package/.grok/hooks/record-passive-event.mjs +198 -0
- package/.grok/hooks/session-start.json +15 -0
- package/.grok/hooks/session-start.mjs +37 -0
- package/.grok/hooks/stop-failure.json +15 -0
- package/.grok/hooks/stop-failure.mjs +6 -0
- package/.grok/hooks/stop.json +15 -0
- package/.grok/hooks/stop.mjs +36 -0
- package/.grok/hooks/subagent-start.json +15 -0
- package/.grok/hooks/subagent-start.mjs +6 -0
- package/.grok/hooks/subagent-stop.json +15 -0
- package/.grok/hooks/subagent-stop.mjs +6 -0
- package/.grok/hooks/user-prompt-submit.json +15 -0
- package/.grok/hooks/user-prompt-submit.mjs +7 -0
- package/.grok/rules/00-litgrok.md +152 -0
- package/.grok/skills/autoconference/LICENSE +21 -0
- package/.grok/skills/autoconference/PROVENANCE.md +23 -0
- package/.grok/skills/autoconference/SKILL.md +140 -0
- package/.grok/skills/autoconference/assets/conference_template.md +76 -0
- package/.grok/skills/autoconference/assets/report_template.md +56 -0
- package/.grok/skills/autoconference/assets/synthesis_template.md +40 -0
- package/.grok/skills/autoconference/references/_canonical-corpus/manifest.json +138 -0
- package/.grok/skills/autoconference/references/agent-prompts.md +30 -0
- package/.grok/skills/autoconference/references/conference-protocol.md +21 -0
- package/.grok/skills/autoconference/references/core-principles.md +13 -0
- package/.grok/skills/autoconference/references/family-contract.md +26 -0
- package/.grok/skills/autoconference/references/modes/analyze.md +11 -0
- package/.grok/skills/autoconference/references/modes/core/convergence-guide.md +41 -0
- package/.grok/skills/autoconference/references/modes/core/crash-recovery.md +11 -0
- package/.grok/skills/autoconference/references/modes/core.md +26 -0
- package/.grok/skills/autoconference/references/modes/debate.md +11 -0
- package/.grok/skills/autoconference/references/modes/plan.md +12 -0
- package/.grok/skills/autoconference/references/modes/resume.md +11 -0
- package/.grok/skills/autoconference/references/modes/ship.md +11 -0
- package/.grok/skills/autoconference/references/modes/survey.md +11 -0
- package/.grok/skills/autoconference/references/results-logging.md +12 -0
- package/.grok/skills/autoconference/references/visualization-guide.md +10 -0
- package/.grok/skills/autoconference/scripts/init_conference.py +646 -0
- package/.grok/skills/autoconference/scripts/verify-canonical-corpus.mjs +177 -0
- package/.grok/skills/autoconference/templates/code-performance.md +59 -0
- package/.grok/skills/autoconference/templates/debate-mode.md +50 -0
- package/.grok/skills/autoconference/templates/prompt-optimization.md +58 -0
- package/.grok/skills/autoconference/templates/quick-conference.md +48 -0
- package/.grok/skills/autoconference/templates/research-synthesis.md +56 -0
- package/.grok/skills/autoconference/templates/survey-mode.md +63 -0
- package/.grok/skills/autoresearch/LICENSE +21 -0
- package/.grok/skills/autoresearch/PROVENANCE.md +22 -0
- package/.grok/skills/autoresearch/SKILL.md +149 -0
- package/.grok/skills/autoresearch/assets/report_template.md +52 -0
- package/.grok/skills/autoresearch/assets/research_template.md +38 -0
- package/.grok/skills/autoresearch/assets/results_template.tsv +2 -0
- package/.grok/skills/autoresearch/references/_canonical-corpus/manifest.json +148 -0
- package/.grok/skills/autoresearch/references/core-principles.md +16 -0
- package/.grok/skills/autoresearch/references/family-contract.md +36 -0
- package/.grok/skills/autoresearch/references/modes/core/evaluator-contract.md +12 -0
- package/.grok/skills/autoresearch/references/modes/core/stuck-detection.md +11 -0
- package/.grok/skills/autoresearch/references/modes/core.md +21 -0
- package/.grok/skills/autoresearch/references/modes/debug/investigation-techniques.md +11 -0
- package/.grok/skills/autoresearch/references/modes/debug.md +11 -0
- package/.grok/skills/autoresearch/references/modes/fix.md +11 -0
- package/.grok/skills/autoresearch/references/modes/learn.md +11 -0
- package/.grok/skills/autoresearch/references/modes/plan.md +14 -0
- package/.grok/skills/autoresearch/references/modes/predict/persona-templates.md +11 -0
- package/.grok/skills/autoresearch/references/modes/predict.md +11 -0
- package/.grok/skills/autoresearch/references/modes/reason.md +11 -0
- package/.grok/skills/autoresearch/references/modes/scenario/dimensions.md +11 -0
- package/.grok/skills/autoresearch/references/modes/scenario.md +11 -0
- package/.grok/skills/autoresearch/references/modes/security/owasp-checklist.md +11 -0
- package/.grok/skills/autoresearch/references/modes/security/stride-model.md +11 -0
- package/.grok/skills/autoresearch/references/modes/security.md +12 -0
- package/.grok/skills/autoresearch/references/modes/ship/type-checklists.md +12 -0
- package/.grok/skills/autoresearch/references/modes/ship.md +11 -0
- package/.grok/skills/autoresearch/references/results-logging.md +12 -0
- package/.grok/skills/autoresearch/references/visualization-guide.md +24 -0
- package/.grok/skills/autoresearch/scripts/init_research.py +391 -0
- package/.grok/skills/autoresearch/scripts/style_presets.py +123 -0
- package/.grok/skills/autoresearch/scripts/verify-canonical-corpus.mjs +179 -0
- package/.grok/skills/browser-drive/SKILL.md +194 -0
- package/.grok/skills/browser-drive/references/snapshot-act-loop.md +61 -0
- package/.grok/skills/comment-checker/SKILL.md +194 -0
- package/.grok/skills/debugging/SKILL.md +82 -0
- package/.grok/skills/debugging/references/methodology/00-setup.md +108 -0
- package/.grok/skills/debugging/references/methodology/02-investigate.md +126 -0
- package/.grok/skills/debugging/references/methodology/04-oracle-triple.md +106 -0
- package/.grok/skills/debugging/references/methodology/05-escalate.md +69 -0
- package/.grok/skills/debugging/references/methodology/06-fix.md +116 -0
- package/.grok/skills/debugging/references/methodology/08-qa.md +94 -0
- package/.grok/skills/debugging/references/methodology/09-cleanup.md +164 -0
- package/.grok/skills/debugging/references/methodology/partial-runtime-evidence.md +228 -0
- package/.grok/skills/debugging/references/post-tool-use-failure-taxonomy.md +55 -0
- package/.grok/skills/debugging/references/reproduction-recipes.md +182 -0
- package/.grok/skills/debugging/references/runtimes/bundled-js-binary.md +415 -0
- package/.grok/skills/debugging/references/runtimes/go.md +252 -0
- package/.grok/skills/debugging/references/runtimes/native-binary.md +484 -0
- package/.grok/skills/debugging/references/runtimes/node.md +260 -0
- package/.grok/skills/debugging/references/runtimes/python.md +248 -0
- package/.grok/skills/debugging/references/runtimes/rust.md +234 -0
- package/.grok/skills/debugging/references/tools/ghidra.md +212 -0
- package/.grok/skills/debugging/references/tools/playwright-cli.md +194 -0
- package/.grok/skills/debugging/references/tools/pwndbg.md +263 -0
- package/.grok/skills/debugging/references/tools/pwntools.md +265 -0
- package/.grok/skills/deep-interview/SKILL.md +216 -0
- package/.grok/skills/frontend-ui-ux/LICENSE +21 -0
- package/.grok/skills/frontend-ui-ux/PROVENANCE.json +38 -0
- package/.grok/skills/frontend-ui-ux/SKILL.md +58 -0
- package/.grok/skills/frontend-ui-ux/SOURCE-MANIFEST.json +1060 -0
- package/.grok/skills/frontend-ui-ux/THIRD-PARTY-NOTICE.txt +14 -0
- package/.grok/skills/frontend-ui-ux/data/design-intelligence.json +1 -0
- package/.grok/skills/frontend-ui-ux/references/_canonical-corpus/legal/frontend-ATTRIBUTION.md +217 -0
- package/.grok/skills/frontend-ui-ux/references/_canonical-corpus/legal/frontend-LICENSE-Apache-2.0.txt +201 -0
- package/.grok/skills/frontend-ui-ux/references/_canonical-corpus/legal/root-LICENSE +21 -0
- package/.grok/skills/frontend-ui-ux/references/_canonical-corpus/manifest.json +873 -0
- package/.grok/skills/frontend-ui-ux/references/adaptive-layout.md +92 -0
- package/.grok/skills/frontend-ui-ux/references/brand-and-imagery.md +93 -0
- package/.grok/skills/frontend-ui-ux/references/complete-contract.md +557 -0
- package/.grok/skills/frontend-ui-ux/references/composition.md +85 -0
- package/.grok/skills/frontend-ui-ux/references/creative-directions.md +80 -0
- package/.grok/skills/frontend-ui-ux/references/design/README.md +248 -0
- package/.grok/skills/frontend-ui-ux/references/design/_INDEX.md +191 -0
- package/.grok/skills/frontend-ui-ux/references/design/airbnb.md +393 -0
- package/.grok/skills/frontend-ui-ux/references/design/airtable.md +92 -0
- package/.grok/skills/frontend-ui-ux/references/design/apple.md +250 -0
- package/.grok/skills/frontend-ui-ux/references/design/aside.md +209 -0
- package/.grok/skills/frontend-ui-ux/references/design/binance.md +348 -0
- package/.grok/skills/frontend-ui-ux/references/design/bmw.md +183 -0
- package/.grok/skills/frontend-ui-ux/references/design/brutalist-skill.md +92 -0
- package/.grok/skills/frontend-ui-ux/references/design/bugatti.md +271 -0
- package/.grok/skills/frontend-ui-ux/references/design/cal.md +262 -0
- package/.grok/skills/frontend-ui-ux/references/design/claude.md +315 -0
- package/.grok/skills/frontend-ui-ux/references/design/clay.md +307 -0
- package/.grok/skills/frontend-ui-ux/references/design/clickhouse.md +284 -0
- package/.grok/skills/frontend-ui-ux/references/design/clone-from-url.md +65 -0
- package/.grok/skills/frontend-ui-ux/references/design/cohere.md +269 -0
- package/.grok/skills/frontend-ui-ux/references/design/coinbase.md +132 -0
- package/.grok/skills/frontend-ui-ux/references/design/composio.md +310 -0
- package/.grok/skills/frontend-ui-ux/references/design/cursor.md +312 -0
- package/.grok/skills/frontend-ui-ux/references/design/design-system-architecture.md +244 -0
- package/.grok/skills/frontend-ui-ux/references/design/elevenlabs.md +268 -0
- package/.grok/skills/frontend-ui-ux/references/design/expo.md +284 -0
- package/.grok/skills/frontend-ui-ux/references/design/ferrari.md +317 -0
- package/.grok/skills/frontend-ui-ux/references/design/figma.md +223 -0
- package/.grok/skills/frontend-ui-ux/references/design/framer.md +249 -0
- package/.grok/skills/frontend-ui-ux/references/design/gpt-tasteskill.md +74 -0
- package/.grok/skills/frontend-ui-ux/references/design/hashicorp.md +281 -0
- package/.grok/skills/frontend-ui-ux/references/design/ibm.md +335 -0
- package/.grok/skills/frontend-ui-ux/references/design/image-to-code-skill.md +1228 -0
- package/.grok/skills/frontend-ui-ux/references/design/imagegen-brandkit.md +798 -0
- package/.grok/skills/frontend-ui-ux/references/design/imagegen-frontend-mobile.md +1465 -0
- package/.grok/skills/frontend-ui-ux/references/design/imagegen-frontend-web.md +987 -0
- package/.grok/skills/frontend-ui-ux/references/design/intercom.md +149 -0
- package/.grok/skills/frontend-ui-ux/references/design/kraken.md +128 -0
- package/.grok/skills/frontend-ui-ux/references/design/lamborghini.md +291 -0
- package/.grok/skills/frontend-ui-ux/references/design/layout-skill.md +107 -0
- package/.grok/skills/frontend-ui-ux/references/design/lazyweb.md +77 -0
- package/.grok/skills/frontend-ui-ux/references/design/linear.app.md +370 -0
- package/.grok/skills/frontend-ui-ux/references/design/lovable.md +301 -0
- package/.grok/skills/frontend-ui-ux/references/design/mastercard.md +368 -0
- package/.grok/skills/frontend-ui-ux/references/design/meta.md +369 -0
- package/.grok/skills/frontend-ui-ux/references/design/minimalist-skill.md +85 -0
- package/.grok/skills/frontend-ui-ux/references/design/minimax.md +260 -0
- package/.grok/skills/frontend-ui-ux/references/design/mintlify.md +329 -0
- package/.grok/skills/frontend-ui-ux/references/design/miro.md +111 -0
- package/.grok/skills/frontend-ui-ux/references/design/mistral.ai.md +264 -0
- package/.grok/skills/frontend-ui-ux/references/design/mongodb.md +269 -0
- package/.grok/skills/frontend-ui-ux/references/design/nike.md +366 -0
- package/.grok/skills/frontend-ui-ux/references/design/notion.md +312 -0
- package/.grok/skills/frontend-ui-ux/references/design/nvidia.md +296 -0
- package/.grok/skills/frontend-ui-ux/references/design/ollama.md +270 -0
- package/.grok/skills/frontend-ui-ux/references/design/opencode.ai.md +284 -0
- package/.grok/skills/frontend-ui-ux/references/design/output-skill.md +49 -0
- package/.grok/skills/frontend-ui-ux/references/design/pinterest.md +233 -0
- package/.grok/skills/frontend-ui-ux/references/design/playstation.md +367 -0
- package/.grok/skills/frontend-ui-ux/references/design/posthog.md +259 -0
- package/.grok/skills/frontend-ui-ux/references/design/raycast.md +271 -0
- package/.grok/skills/frontend-ui-ux/references/design/react-dev-tooling-skill.md +230 -0
- package/.grok/skills/frontend-ui-ux/references/design/redesign-skill.md +178 -0
- package/.grok/skills/frontend-ui-ux/references/design/renault.md +314 -0
- package/.grok/skills/frontend-ui-ux/references/design/replicate.md +264 -0
- package/.grok/skills/frontend-ui-ux/references/design/resend.md +306 -0
- package/.grok/skills/frontend-ui-ux/references/design/revolut.md +188 -0
- package/.grok/skills/frontend-ui-ux/references/design/runwayml.md +247 -0
- package/.grok/skills/frontend-ui-ux/references/design/sanity.md +360 -0
- package/.grok/skills/frontend-ui-ux/references/design/sentry.md +265 -0
- package/.grok/skills/frontend-ui-ux/references/design/shopify.md +353 -0
- package/.grok/skills/frontend-ui-ux/references/design/soft-skill.md +98 -0
- package/.grok/skills/frontend-ui-ux/references/design/spacex.md +197 -0
- package/.grok/skills/frontend-ui-ux/references/design/spotify.md +249 -0
- package/.grok/skills/frontend-ui-ux/references/design/starbucks.md +583 -0
- package/.grok/skills/frontend-ui-ux/references/design/stitch-design-example.md +121 -0
- package/.grok/skills/frontend-ui-ux/references/design/stitch-skill.md +184 -0
- package/.grok/skills/frontend-ui-ux/references/design/stripe.md +325 -0
- package/.grok/skills/frontend-ui-ux/references/design/supabase.md +258 -0
- package/.grok/skills/frontend-ui-ux/references/design/superhuman.md +255 -0
- package/.grok/skills/frontend-ui-ux/references/design/taste-skill.md +1206 -0
- package/.grok/skills/frontend-ui-ux/references/design/tesla.md +289 -0
- package/.grok/skills/frontend-ui-ux/references/design/theverge.md +342 -0
- package/.grok/skills/frontend-ui-ux/references/design/together.ai.md +266 -0
- package/.grok/skills/frontend-ui-ux/references/design/uber.md +298 -0
- package/.grok/skills/frontend-ui-ux/references/design/vercel.md +313 -0
- package/.grok/skills/frontend-ui-ux/references/design/vodafone.md +426 -0
- package/.grok/skills/frontend-ui-ux/references/design/voltagent.md +326 -0
- package/.grok/skills/frontend-ui-ux/references/design/warp.md +256 -0
- package/.grok/skills/frontend-ui-ux/references/design/webflow.md +95 -0
- package/.grok/skills/frontend-ui-ux/references/design/wired.md +281 -0
- package/.grok/skills/frontend-ui-ux/references/design/wise.md +176 -0
- package/.grok/skills/frontend-ui-ux/references/design/x.ai.md +260 -0
- package/.grok/skills/frontend-ui-ux/references/design/zapier.md +331 -0
- package/.grok/skills/frontend-ui-ux/references/designpowers/EVIDENCE.md +97 -0
- package/.grok/skills/frontend-ui-ux/references/designpowers/README.md +48 -0
- package/.grok/skills/frontend-ui-ux/references/designpowers/UPSTREAM.md +80 -0
- package/.grok/skills/frontend-ui-ux/references/designpowers/lane-a-direction.md +64 -0
- package/.grok/skills/frontend-ui-ux/references/designpowers/lane-b-execution.md +65 -0
- package/.grok/skills/frontend-ui-ux/references/designpowers/lane-c-review.md +65 -0
- package/.grok/skills/frontend-ui-ux/references/designpowers/lane-d-memory.md +83 -0
- package/.grok/skills/frontend-ui-ux/references/designpowers/orchestration.md +80 -0
- package/.grok/skills/frontend-ui-ux/references/designpowers/routing.md +79 -0
- package/.grok/skills/frontend-ui-ux/references/designpowers/vendor/LICENSE +21 -0
- package/.grok/skills/frontend-ui-ux/references/designpowers/vendor/agents/accessibility-reviewer.md +83 -0
- package/.grok/skills/frontend-ui-ux/references/designpowers/vendor/agents/content-writer.md +132 -0
- package/.grok/skills/frontend-ui-ux/references/designpowers/vendor/agents/design-builder.md +109 -0
- package/.grok/skills/frontend-ui-ux/references/designpowers/vendor/agents/design-critic.md +89 -0
- package/.grok/skills/frontend-ui-ux/references/designpowers/vendor/agents/design-lead.md +113 -0
- package/.grok/skills/frontend-ui-ux/references/designpowers/vendor/agents/design-scout.md +78 -0
- package/.grok/skills/frontend-ui-ux/references/designpowers/vendor/agents/design-strategist.md +121 -0
- package/.grok/skills/frontend-ui-ux/references/designpowers/vendor/agents/heuristic-evaluator.md +268 -0
- package/.grok/skills/frontend-ui-ux/references/designpowers/vendor/agents/inspiration-scout.md +107 -0
- package/.grok/skills/frontend-ui-ux/references/designpowers/vendor/agents/motion-designer.md +120 -0
- package/.grok/skills/frontend-ui-ux/references/designpowers/vendor/skills/accessible-content/reference.md +101 -0
- package/.grok/skills/frontend-ui-ux/references/designpowers/vendor/skills/adaptive-interfaces/reference.md +109 -0
- package/.grok/skills/frontend-ui-ux/references/designpowers/vendor/skills/cognitive-accessibility/reference.md +107 -0
- package/.grok/skills/frontend-ui-ux/references/designpowers/vendor/skills/design-debate/reference.md +199 -0
- package/.grok/skills/frontend-ui-ux/references/designpowers/vendor/skills/design-debt-tracker/reference.md +174 -0
- package/.grok/skills/frontend-ui-ux/references/designpowers/vendor/skills/design-handoff/reference.md +125 -0
- package/.grok/skills/frontend-ui-ux/references/designpowers/vendor/skills/design-md/reference.md +106 -0
- package/.grok/skills/frontend-ui-ux/references/designpowers/vendor/skills/design-retrospective/reference.md +266 -0
- package/.grok/skills/frontend-ui-ux/references/designpowers/vendor/skills/design-review/reference.md +123 -0
- package/.grok/skills/frontend-ui-ux/references/designpowers/vendor/skills/design-system-alignment/reference.md +120 -0
- package/.grok/skills/frontend-ui-ux/references/designpowers/vendor/skills/designpowers-critique/reference.md +164 -0
- package/.grok/skills/frontend-ui-ux/references/designpowers/vendor/skills/heuristic-evaluation/reference.md +85 -0
- package/.grok/skills/frontend-ui-ux/references/designpowers/vendor/skills/inclusive-personas/reference.md +98 -0
- package/.grok/skills/frontend-ui-ux/references/designpowers/vendor/skills/inspiration-scouting/reference.md +165 -0
- package/.grok/skills/frontend-ui-ux/references/designpowers/vendor/skills/interaction-design/reference.md +122 -0
- package/.grok/skills/frontend-ui-ux/references/designpowers/vendor/skills/motion-choreography/reference.md +81 -0
- package/.grok/skills/frontend-ui-ux/references/designpowers/vendor/skills/research-planning/reference.md +96 -0
- package/.grok/skills/frontend-ui-ux/references/designpowers/vendor/skills/responsive-patterns/reference.md +77 -0
- package/.grok/skills/frontend-ui-ux/references/designpowers/vendor/skills/synthetic-user-testing/reference.md +192 -0
- package/.grok/skills/frontend-ui-ux/references/designpowers/vendor/skills/taste-feedback/reference.md +165 -0
- package/.grok/skills/frontend-ui-ux/references/designpowers/vendor/skills/taste-report/reference.md +78 -0
- package/.grok/skills/frontend-ui-ux/references/designpowers/vendor/skills/token-architecture/reference.md +75 -0
- package/.grok/skills/frontend-ui-ux/references/designpowers/vendor/skills/ui-composition/reference.md +117 -0
- package/.grok/skills/frontend-ui-ux/references/designpowers/vendor/skills/usability-testing/reference.md +78 -0
- package/.grok/skills/frontend-ui-ux/references/designpowers/vendor/skills/verification-before-shipping/reference.md +125 -0
- package/.grok/skills/frontend-ui-ux/references/designpowers/vendor/skills/voice-and-tone/reference.md +79 -0
- package/.grok/skills/frontend-ui-ux/references/designpowers/vendor/skills/writing-design-plans/reference.md +119 -0
- package/.grok/skills/frontend-ui-ux/references/evidence-review.md +126 -0
- package/.grok/skills/frontend-ui-ux/references/implementation-platforms.md +109 -0
- package/.grok/skills/frontend-ui-ux/references/inclusive-interface.md +92 -0
- package/.grok/skills/frontend-ui-ux/references/interaction-motion.md +101 -0
- package/.grok/skills/frontend-ui-ux/references/operating-lanes.md +92 -0
- package/.grok/skills/frontend-ui-ux/references/perfection/README.md +160 -0
- package/.grok/skills/frontend-ui-ux/references/perfection/react-perf-tooling.md +127 -0
- package/.grok/skills/frontend-ui-ux/references/performance-delivery.md +93 -0
- package/.grok/skills/frontend-ui-ux/references/product-direction.md +84 -0
- package/.grok/skills/frontend-ui-ux/references/redesign-playbook.md +97 -0
- package/.grok/skills/frontend-ui-ux/references/system-foundations.md +84 -0
- package/.grok/skills/frontend-ui-ux/references/taste-direction.md +86 -0
- package/.grok/skills/frontend-ui-ux/references/ui-ux-db/README.md +659 -0
- package/.grok/skills/frontend-ui-ux/references/ui-ux-db/data/charts.csv +26 -0
- package/.grok/skills/frontend-ui-ux/references/ui-ux-db/data/colors.csv +162 -0
- package/.grok/skills/frontend-ui-ux/references/ui-ux-db/data/icons.csv +106 -0
- package/.grok/skills/frontend-ui-ux/references/ui-ux-db/data/landing.csv +35 -0
- package/.grok/skills/frontend-ui-ux/references/ui-ux-db/data/products.csv +162 -0
- package/.grok/skills/frontend-ui-ux/references/ui-ux-db/data/react-performance.csv +45 -0
- package/.grok/skills/frontend-ui-ux/references/ui-ux-db/data/stacks/astro.csv +54 -0
- package/.grok/skills/frontend-ui-ux/references/ui-ux-db/data/stacks/flutter.csv +53 -0
- package/.grok/skills/frontend-ui-ux/references/ui-ux-db/data/stacks/html-tailwind.csv +56 -0
- package/.grok/skills/frontend-ui-ux/references/ui-ux-db/data/stacks/jetpack-compose.csv +53 -0
- package/.grok/skills/frontend-ui-ux/references/ui-ux-db/data/stacks/nextjs.csv +53 -0
- package/.grok/skills/frontend-ui-ux/references/ui-ux-db/data/stacks/nuxt-ui.csv +51 -0
- package/.grok/skills/frontend-ui-ux/references/ui-ux-db/data/stacks/nuxtjs.csv +59 -0
- package/.grok/skills/frontend-ui-ux/references/ui-ux-db/data/stacks/react-native.csv +52 -0
- package/.grok/skills/frontend-ui-ux/references/ui-ux-db/data/stacks/react.csv +54 -0
- package/.grok/skills/frontend-ui-ux/references/ui-ux-db/data/stacks/shadcn.csv +61 -0
- package/.grok/skills/frontend-ui-ux/references/ui-ux-db/data/stacks/svelte.csv +54 -0
- package/.grok/skills/frontend-ui-ux/references/ui-ux-db/data/stacks/swiftui.csv +51 -0
- package/.grok/skills/frontend-ui-ux/references/ui-ux-db/data/stacks/vue.csv +50 -0
- package/.grok/skills/frontend-ui-ux/references/ui-ux-db/data/styles.csv +85 -0
- package/.grok/skills/frontend-ui-ux/references/ui-ux-db/data/typography.csv +74 -0
- package/.grok/skills/frontend-ui-ux/references/ui-ux-db/data/ui-reasoning.csv +162 -0
- package/.grok/skills/frontend-ui-ux/references/ui-ux-db/data/ux-guidelines.csv +100 -0
- package/.grok/skills/frontend-ui-ux/references/ui-ux-db/data/web-interface.csv +31 -0
- package/.grok/skills/frontend-ui-ux/references/ui-ux-db/scripts/core.py +262 -0
- package/.grok/skills/frontend-ui-ux/references/ui-ux-db/scripts/design_system.py +1148 -0
- package/.grok/skills/frontend-ui-ux/references/ui-ux-db/scripts/search.py +114 -0
- package/.grok/skills/frontend-ui-ux/references/visual-language.md +84 -0
- package/.grok/skills/frontend-ui-ux/references/visual-reconstruction.md +97 -0
- package/.grok/skills/frontend-ui-ux/schemas/design-contract-v1alpha1.schema.json +306 -0
- package/.grok/skills/frontend-ui-ux/schemas/design-contract-v1beta1.schema.json +291 -0
- package/.grok/skills/frontend-ui-ux/schemas/design-contract-v1beta2.schema.json +311 -0
- package/.grok/skills/frontend-ui-ux/scripts/design-contract-format.mjs +286 -0
- package/.grok/skills/frontend-ui-ux/scripts/design-contract-inventory-rules.mjs +354 -0
- package/.grok/skills/frontend-ui-ux/scripts/design-contract-rules.mjs +361 -0
- package/.grok/skills/frontend-ui-ux/scripts/design-contract-surface-rules.mjs +313 -0
- package/.grok/skills/frontend-ui-ux/scripts/design-data.mjs +68 -0
- package/.grok/skills/frontend-ui-ux/scripts/errors.mjs +12 -0
- package/.grok/skills/frontend-ui-ux/scripts/import-design-intelligence.mjs +124 -0
- package/.grok/skills/frontend-ui-ux/scripts/json-boundary.mjs +63 -0
- package/.grok/skills/frontend-ui-ux/scripts/query-design-intelligence.mjs +106 -0
- package/.grok/skills/frontend-ui-ux/scripts/search.mjs +43 -0
- package/.grok/skills/frontend-ui-ux/scripts/source-replay.mjs +153 -0
- package/.grok/skills/frontend-ui-ux/scripts/strict-json.mjs +143 -0
- package/.grok/skills/frontend-ui-ux/scripts/validate-design-contract.mjs +52 -0
- package/.grok/skills/frontend-ui-ux/scripts/verify-canonical-corpus.mjs +214 -0
- package/.grok/skills/lit-burnoff/SKILL.md +175 -0
- package/.grok/skills/lit-burnoff-file/SKILL.md +67 -0
- package/.grok/skills/lit-code/SKILL.md +531 -0
- package/.grok/skills/lit-code/references/go/README.md +90 -0
- package/.grok/skills/lit-code/references/go/backend-stack.md +641 -0
- package/.grok/skills/lit-code/references/go/bootstrap.md +328 -0
- package/.grok/skills/lit-code/references/go/bubbletea-v2.md +360 -0
- package/.grok/skills/lit-code/references/go/cobra-stack.md +468 -0
- package/.grok/skills/lit-code/references/go/concurrency.md +362 -0
- package/.grok/skills/lit-code/references/go/data-modeling.md +329 -0
- package/.grok/skills/lit-code/references/go/error-handling.md +359 -0
- package/.grok/skills/lit-code/references/go/golangci-strict.md +236 -0
- package/.grok/skills/lit-code/references/go/grpc-connect.md +375 -0
- package/.grok/skills/lit-code/references/go/libraries.md +337 -0
- package/.grok/skills/lit-code/references/go/one-liners.md +202 -0
- package/.grok/skills/lit-code/references/go/sqlc-pgx.md +471 -0
- package/.grok/skills/lit-code/references/go/testing.md +467 -0
- package/.grok/skills/lit-code/references/go/type-patterns.md +298 -0
- package/.grok/skills/lit-code/references/permission-sandbox-matrix.md +70 -0
- package/.grok/skills/lit-code/references/python/README.md +314 -0
- package/.grok/skills/lit-code/references/python/async-anyio.md +442 -0
- package/.grok/skills/lit-code/references/python/data-modeling.md +233 -0
- package/.grok/skills/lit-code/references/python/data-processing.md +133 -0
- package/.grok/skills/lit-code/references/python/error-handling.md +218 -0
- package/.grok/skills/lit-code/references/python/fastapi-stack.md +316 -0
- package/.grok/skills/lit-code/references/python/httpx2-optimization.md +360 -0
- package/.grok/skills/lit-code/references/python/libraries.md +307 -0
- package/.grok/skills/lit-code/references/python/one-liners.md +268 -0
- package/.grok/skills/lit-code/references/python/orjson-stack.md +378 -0
- package/.grok/skills/lit-code/references/python/pydantic-ai.md +285 -0
- package/.grok/skills/lit-code/references/python/pyproject-strict.md +232 -0
- package/.grok/skills/lit-code/references/python/textual-tui.md +201 -0
- package/.grok/skills/lit-code/references/python/type-patterns.md +176 -0
- package/.grok/skills/lit-code/references/rust/README.md +317 -0
- package/.grok/skills/lit-code/references/rust/async-tokio.md +299 -0
- package/.grok/skills/lit-code/references/rust/axum-stack.md +467 -0
- package/.grok/skills/lit-code/references/rust/cargo-strict.md +317 -0
- package/.grok/skills/lit-code/references/rust/clap-stack.md +409 -0
- package/.grok/skills/lit-code/references/rust/concurrency.md +375 -0
- package/.grok/skills/lit-code/references/rust/libraries.md +439 -0
- package/.grok/skills/lit-code/references/rust/one-liners.md +291 -0
- package/.grok/skills/lit-code/references/rust/proptest-insta.md +429 -0
- package/.grok/skills/lit-code/references/rust/type-state.md +354 -0
- package/.grok/skills/lit-code/references/rust/unsafe-discipline.md +250 -0
- package/.grok/skills/lit-code/references/rust/zero-cost-safety.md +527 -0
- package/.grok/skills/lit-code/references/rust-ub/README.md +289 -0
- package/.grok/skills/lit-code/references/rust-ub/miri-sanitizers-loom.md +411 -0
- package/.grok/skills/lit-code/references/rust-ub/ub-taxonomy.md +269 -0
- package/.grok/skills/lit-code/references/tool-boundaries.md +66 -0
- package/.grok/skills/lit-code/references/typescript/README.md +195 -0
- package/.grok/skills/lit-code/references/typescript/backend-hono.md +672 -0
- package/.grok/skills/lit-code/references/typescript/bootstrap.md +199 -0
- package/.grok/skills/lit-code/references/typescript/data-modeling.md +202 -0
- package/.grok/skills/lit-code/references/typescript/error-handling.md +169 -0
- package/.grok/skills/lit-code/references/typescript/tsconfig-strict.md +152 -0
- package/.grok/skills/lit-code/references/typescript/type-patterns.md +196 -0
- package/.grok/skills/lit-code/references/worked-cases.md +190 -0
- package/.grok/skills/lit-commit/SKILL.md +218 -0
- package/.grok/skills/lit-comprehend/SKILL.md +63 -0
- package/.grok/skills/lit-comprehend/references/artifact-format.md +58 -0
- package/.grok/skills/lit-comprehend/references/artifact-template.md +219 -0
- package/.grok/skills/lit-comprehend/references/honesty-ledger-contract.md +51 -0
- package/.grok/skills/lit-comprehend/references/micro-worlds.md +195 -0
- package/.grok/skills/lit-comprehend/references/worked-explainer.md +51 -0
- package/.grok/skills/lit-crucible/SKILL.md +231 -0
- package/.grok/skills/lit-handoff/SKILL.md +159 -0
- package/.grok/skills/lit-handoff/evals/evals.json +154 -0
- package/.grok/skills/lit-handoff/examples/HANDOFF-example-generic-auth-refactor.md +97 -0
- package/.grok/skills/lit-handoff/references/_canonical-corpus/manifest.json +17 -0
- package/.grok/skills/lit-handoff/references/source-pointer.md +31 -0
- package/.grok/skills/lit-handoff/scripts/verify-canonical-corpus.mjs +137 -0
- package/.grok/skills/lit-handoff/templates/HANDOFF.md +121 -0
- package/.grok/skills/lit-init/SKILL.md +244 -0
- package/.grok/skills/lit-korean/SKILL.md +176 -0
- package/.grok/skills/lit-plan/SKILL.md +71 -0
- package/.grok/skills/lit-plan/references/plan-schema.md +87 -0
- package/.grok/skills/lit-plan/references/start-work-handoff-contract.md +59 -0
- package/.grok/skills/lit-plan/scripts/scaffold-plan.mjs +259 -0
- package/.grok/skills/lit-plan/scripts/validate-plan.mjs +89 -0
- package/.grok/skills/lit-recap/SKILL.md +57 -0
- package/.grok/skills/lit-scientific-visualization/SKILL.md +213 -0
- package/.grok/skills/lit-scientific-visualization/scripts/verify-canonical-corpus.mjs +181 -0
- package/.grok/skills/lit-team/SKILL.md +73 -0
- package/.grok/skills/lit-team/references/explore-packet.md +70 -0
- package/.grok/skills/lit-team/references/general-purpose-packet.md +75 -0
- package/.grok/skills/lit-team/references/plan-packet.md +73 -0
- package/.grok/skills/litgoal/SKILL.md +98 -0
- package/.grok/skills/litgrok/SKILL.md +77 -0
- package/.grok/skills/litresearch/SKILL.md +60 -0
- package/.grok/skills/litresearch/references/mcp-tool-use-patterns.md +83 -0
- package/.grok/skills/litresearch/references/source-verdict-taxonomy.md +40 -0
- package/.grok/skills/litwork/SKILL.md +109 -0
- package/.grok/skills/lsp/SKILL.md +56 -0
- package/.grok/skills/lsp/references/built-in-lsp-contract.md +71 -0
- package/.grok/skills/lsp-setup/SKILL.md +82 -0
- package/.grok/skills/lsp-setup/references/bash/README.md +54 -0
- package/.grok/skills/lsp-setup/references/c-cpp/README.md +58 -0
- package/.grok/skills/lsp-setup/references/csharp/README.md +64 -0
- package/.grok/skills/lsp-setup/references/dart/README.md +48 -0
- package/.grok/skills/lsp-setup/references/elixir/README.md +51 -0
- package/.grok/skills/lsp-setup/references/go/README.md +53 -0
- package/.grok/skills/lsp-setup/references/haskell/README.md +57 -0
- package/.grok/skills/lsp-setup/references/java/README.md +55 -0
- package/.grok/skills/lsp-setup/references/julia/README.md +56 -0
- package/.grok/skills/lsp-setup/references/kotlin/README.md +58 -0
- package/.grok/skills/lsp-setup/references/lua/README.md +48 -0
- package/.grok/skills/lsp-setup/references/php/README.md +49 -0
- package/.grok/skills/lsp-setup/references/python/README.md +60 -0
- package/.grok/skills/lsp-setup/references/ruby/README.md +53 -0
- package/.grok/skills/lsp-setup/references/rust/README.md +55 -0
- package/.grok/skills/lsp-setup/references/swift/README.md +52 -0
- package/.grok/skills/lsp-setup/references/terraform/README.md +50 -0
- package/.grok/skills/lsp-setup/references/typescript/README.md +63 -0
- package/.grok/skills/lsp-setup/references/yaml/README.md +47 -0
- package/.grok/skills/lsp-setup/references/zig/README.md +49 -0
- package/.grok/skills/lsp-setup/scripts/detect-lsp.mjs +40 -0
- package/.grok/skills/lsp-setup/scripts/lsp-server-table.mjs +309 -0
- package/.grok/skills/lsp-setup/scripts/verify-lsp.mjs +54 -0
- package/.grok/skills/refactor/SKILL.md +210 -0
- package/.grok/skills/review-work/SKILL.md +515 -0
- package/.grok/skills/review-work/references/behavior-lane-contract.md +48 -0
- package/.grok/skills/review-work/references/documentation-lane-contract.md +49 -0
- package/.grok/skills/review-work/references/integration-lane-contract.md +51 -0
- package/.grok/skills/review-work/references/regression-lane-contract.md +48 -0
- package/.grok/skills/review-work/references/safety-lane-contract.md +48 -0
- package/.grok/skills/review-work/references/test-lane-contract.md +48 -0
- package/.grok/skills/review-work/scripts/check-lanes.mjs +63 -0
- package/.grok/skills/rules/SKILL.md +51 -0
- package/.grok/skills/rules/references/loading-order-contract.md +66 -0
- package/.grok/skills/rules/scripts/resolve-guidance.mjs +180 -0
- package/.grok/skills/skill-observer/SKILL.md +78 -0
- package/.grok/skills/skill-observer/references/review-contract.md +77 -0
- package/.grok/skills/skill-observer/scripts/curator.mjs +121 -0
- package/.grok/skills/skill-observer/scripts/review.mjs +347 -0
- package/.grok/skills/skill-observer/scripts/skill-loop.mjs +2087 -0
- package/.grok/skills/skill-observer/scripts/validate-skills.mjs +83 -0
- package/.grok/skills/start-work/SKILL.md +396 -0
- package/.grok/skills/structural-search/SKILL.md +195 -0
- package/.grok/skills/visual-qa/SKILL.md +50 -0
- package/.grok/skills/visual-qa/references/capture-playbook.md +47 -0
- package/.grok/skills/visual-qa/references/complete-contract.md +721 -0
- package/.grok/skills/visual-qa/references/verdict-taxonomy.md +34 -0
- package/.grok/skills/visual-qa/scripts/verify-evidence-manifest.mjs +91 -0
- package/.grok/skills/wikify/SKILL.md +65 -0
- package/.grok/skills/wikify/references/page-format.md +81 -0
- package/.grok/skills/wikify/references/provenance-contract.md +52 -0
- package/.grok/vendor/NOTICE.md +16 -0
- package/.grok/vendor/licenses/045_scientific-visualization-MIT.txt +21 -0
- package/.grok/vendor/provenance/045_scientific-visualization.md +37 -0
- package/.grok/vendor/scientific-visualization/assets/color_palettes.py +197 -0
- package/.grok/vendor/scientific-visualization/assets/nature.mplstyle +75 -0
- package/.grok/vendor/scientific-visualization/assets/presentation.mplstyle +74 -0
- package/.grok/vendor/scientific-visualization/assets/publication.mplstyle +78 -0
- package/.grok/vendor/scientific-visualization/evals/evals.json +158 -0
- package/.grok/vendor/scientific-visualization/references/_canonical-corpus/manifest.json +38 -0
- package/.grok/vendor/scientific-visualization/references/color_palettes.md +380 -0
- package/.grok/vendor/scientific-visualization/references/journal_requirements.md +359 -0
- package/.grok/vendor/scientific-visualization/references/matplotlib_examples.md +608 -0
- package/.grok/vendor/scientific-visualization/references/mdanalysis_martini_visualization.md +85 -0
- package/.grok/vendor/scientific-visualization/references/publication_guidelines.md +217 -0
- package/.grok/vendor/scientific-visualization/references/seaborn_for_publications.md +293 -0
- package/.grok/vendor/scientific-visualization/scripts/figure_export.py +238 -0
- package/.grok/vendor/scientific-visualization/scripts/style_presets.py +467 -0
- package/.grok/vendor/scientific-visualization/tests/test_figure_export.py +51 -0
- package/.grok/vendor/scientific-visualization/tests/test_style_presets.py +114 -0
- package/CHANGELOG.md +131 -0
- package/CODE_OF_CONDUCT.md +9 -0
- package/CONTRIBUTING.md +22 -0
- package/LICENSE +21 -0
- package/README.md +228 -0
- package/README_ko-KR.md +228 -0
- package/SECURITY.md +11 -0
- package/SUPPORT.md +9 -0
- package/bin/litgrok.mjs +860 -0
- package/docs/assets/cover.webp +0 -0
- package/docs/assets/litgrok-clay-icon.png +0 -0
- package/docs/assets/litgrok-continuity-1600.webp +0 -0
- package/docs/assets/litgrok-ignition-1600.webp +0 -0
- package/docs/assets/litgrok-wordmark.svg +5 -0
- package/docs/assets/readme/README.md +31 -0
- package/docs/assets/readme/badge-license.svg +1 -0
- package/docs/assets/readme/badge-version.svg +1 -0
- package/docs/privacy.md +13 -0
- package/docs/reference.md +271 -0
- package/docs/reference_ko-KR.md +267 -0
- package/package.json +50 -0
- package/plugin.json +27 -0
|
@@ -0,0 +1,94 @@
|
|
|
1
|
+
# Phase 8 — Manual QA by Actually Using It
|
|
2
|
+
|
|
3
|
+
Tests cover cases you thought of. Real usage covers the ones you didn't.
|
|
4
|
+
|
|
5
|
+
The single fastest way to ship a broken fix is to stop at "tests pass". Manual QA means interacting with the running system the way the user does, then comparing observed behavior to the original bug report.
|
|
6
|
+
|
|
7
|
+
---
|
|
8
|
+
|
|
9
|
+
## Product-type playbook
|
|
10
|
+
|
|
11
|
+
Pick the row that matches the product. Do what it says. Do not substitute.
|
|
12
|
+
|
|
13
|
+
| Product type | QA means… |
|
|
14
|
+
|---|---|
|
|
15
|
+
| **CLI tool** | Open `tmux`, run the actual command end-to-end, capture output. Paste the session transcript into the journal. Include exit code, stdout, stderr, side-effect check (files created/modified). |
|
|
16
|
+
| **HTTP API** | Start the real server, hit endpoints with `curl` or `httpie`, inspect response status + body + headers. Hit the specific endpoint that reproduced the bug. If there's auth, use real auth. |
|
|
17
|
+
| **Browser-served web app** | **Drive a real browser via Playwright CLI.** See [tools/playwright-cli.md](../tools/playwright-cli.md). Navigate the exact page/flow that reproduced the bug. Capture screenshot + DOM + network evidence. **Do not substitute with curl** — browsers have state (cookies, localStorage, service workers, client-side JS, viewport-dependent CSS) that curl does not have. |
|
|
18
|
+
| **Agent / LLM pipeline** | Run the same user prompt that originally failed. Capture the full turn — tool calls, messages, usage counters. **Confirm non-zero usage** (zero usage = still failing silently, see silent-failure check below). |
|
|
19
|
+
| **Background worker / job queue** | Trigger the job through the normal entry point (API call, cron tick, message publish), tail the worker logs, observe completion state in the queue or DB. Don't just call the worker function directly — the trigger path matters. |
|
|
20
|
+
| **MCP server** | Invoke the tool via its actual client (Claude Desktop, Cursor, etc. if available) or `mcp-cli`, not just the HTTP probe endpoint. The MCP handshake itself is sometimes where bugs live. |
|
|
21
|
+
| **Native binary** | Re-run the exact command that crashed / misbehaved. If the input was a file, use the same file. If the bug was exploitable, confirm the exploit repro via pwntools (see [tools/pwntools.md](../tools/pwntools.md)). Capture exit code, signal if any, core dump if generated. |
|
|
22
|
+
| **Bundled-app binary** (Bun SEA, Node SEA, Electron, etc.) | Re-run the exact command. If the operation requires paid quota / blocked network, capture the **app's debug log** (`APP_DEBUG=1 APP_LOG_LEVEL=debug APP_LOG_FILE=/tmp/trace.log`) which usually emits the assembled request before sending. See [methodology/partial-runtime-evidence.md](partial-runtime-evidence.md) for combining partial signals into a defensible verification. |
|
|
23
|
+
| **Long-running daemon** | Start fresh, let it run for the amount of time the bug originally took to manifest (not less), capture resource usage (memory, fd, cpu) throughout. Short-running QA misses resource leaks and cumulative state bugs. |
|
|
24
|
+
|
|
25
|
+
---
|
|
26
|
+
|
|
27
|
+
## Journal format
|
|
28
|
+
|
|
29
|
+
Every QA run goes in the journal under "Findings":
|
|
30
|
+
|
|
31
|
+
```markdown
|
|
32
|
+
### Manual QA — <product type> (<ISO timestamp>)
|
|
33
|
+
- Scenario: <one line describing what you did>
|
|
34
|
+
- Command: `<exact invocation>`
|
|
35
|
+
- Observed output:
|
|
36
|
+
```
|
|
37
|
+
<verbatim output, trimmed to relevant section>
|
|
38
|
+
```
|
|
39
|
+
- Expected output: <what correct behavior looks like>
|
|
40
|
+
- Fix verified: yes / no / partial — <details>
|
|
41
|
+
```
|
|
42
|
+
|
|
43
|
+
If any QA step shows **partial or regressed behavior**, this is not "mostly done" — it's incomplete. Return to Phase 6.
|
|
44
|
+
|
|
45
|
+
---
|
|
46
|
+
|
|
47
|
+
## The silent-failure check (always run)
|
|
48
|
+
|
|
49
|
+
Regardless of product type, audit the fix against these silent-failure patterns. If the original bug was a silent failure, the same pattern may exist in adjacent code that you haven't tested yet.
|
|
50
|
+
|
|
51
|
+
### Universal silent-failure signals
|
|
52
|
+
|
|
53
|
+
- HTTP 2xx with empty or default body
|
|
54
|
+
- Response `ok: true` but a sub-field contains an error token (e.g. `stopReason: "error"`, `status: "failed"`)
|
|
55
|
+
- `usage.totalTokens === 0` on an LLM response
|
|
56
|
+
- Process exit code 0 but stderr contains an exception traceback
|
|
57
|
+
- Panic recovered and logged but ignored
|
|
58
|
+
- Goroutine / task / promise rejection with no top-level handler
|
|
59
|
+
- `try { ... } catch { /* swallowed */ }` or `except: pass`
|
|
60
|
+
- Success response shape but semantic field indicates failure (e.g. `error: null` actually being `error: "..."` with falsy check)
|
|
61
|
+
- Write returned success but read-back shows stale data
|
|
62
|
+
- Job marked complete but side-effect did not happen
|
|
63
|
+
- Cache hit path returned stale data and no refresh was triggered
|
|
64
|
+
|
|
65
|
+
### Language-specific silent-failure signals
|
|
66
|
+
|
|
67
|
+
Check the runtime reference for additional patterns:
|
|
68
|
+
|
|
69
|
+
- [runtimes/python.md](../runtimes/python.md) — asyncio task exceptions, bare `except`, `logging.exception` that goes nowhere
|
|
70
|
+
- [runtimes/node.md](../runtimes/node.md) — unhandled promise rejections, `void` on async, swallowed `.catch(() => {})`
|
|
71
|
+
- [runtimes/rust.md](../runtimes/rust.md) — `.unwrap_or_default()`, `let _ = result`, error variants discarded
|
|
72
|
+
- [runtimes/go.md](../runtimes/go.md) — `if err != nil { return err }` that never reaches user output, recovered panics, buffered channels that block silently
|
|
73
|
+
- [runtimes/native-binary.md](../runtimes/native-binary.md) — ignored return codes from libc, missing `perror`, `alarm()` / signal masks
|
|
74
|
+
- [runtimes/bundled-js-binary.md](../runtimes/bundled-js-binary.md) — `process.env.X` baked at build time, dead code from tree-shaking failures, worker sub-bundles diverging from main bundle
|
|
75
|
+
|
|
76
|
+
### What to do when you find another silent-failure spot
|
|
77
|
+
|
|
78
|
+
Don't fix it. This is out of scope for the current bug.
|
|
79
|
+
|
|
80
|
+
Note it in the journal under a "Follow-ups" section with:
|
|
81
|
+
- File:line
|
|
82
|
+
- Pattern matched
|
|
83
|
+
- Proposed fix sketch (one line)
|
|
84
|
+
- Risk level (what happens if left unfixed)
|
|
85
|
+
|
|
86
|
+
Surface these to the user in the final message under "Next steps I didn't take".
|
|
87
|
+
|
|
88
|
+
---
|
|
89
|
+
|
|
90
|
+
## The "fix verified" bar
|
|
91
|
+
|
|
92
|
+
"Fix verified" means: the exact original failing scenario, re-run, now produces the correct output. Not a similar scenario. Not a unit test of the fix. The original scenario.
|
|
93
|
+
|
|
94
|
+
If you can't re-run the original scenario (e.g. it required a specific data state that's gone), construct the closest equivalent and document the difference in the journal. Escalate to the user if the equivalent is materially different.
|
|
@@ -0,0 +1,164 @@
|
|
|
1
|
+
# Phase 9 + 10 — Cleanup & Final Verification
|
|
2
|
+
|
|
3
|
+
The working tree after the session must differ from before only by the real fix and its test. Anything else is a process failure.
|
|
4
|
+
|
|
5
|
+
---
|
|
6
|
+
|
|
7
|
+
## Phase 9 — Cleanup & Revert
|
|
8
|
+
|
|
9
|
+
### The walk
|
|
10
|
+
|
|
11
|
+
Open the journal's "Artifacts to revert" list. Walk it top to bottom. Check each box only after the revert command succeeds and produces no error.
|
|
12
|
+
|
|
13
|
+
### Standard revert operations
|
|
14
|
+
|
|
15
|
+
Most sessions create some combination of these artifacts. The commands below are the defaults — your journal should have the exact commands for this session.
|
|
16
|
+
|
|
17
|
+
```bash
|
|
18
|
+
# --- Temporary source edits (instrumentation statements, debug prints) ---
|
|
19
|
+
git checkout <file> # reverts only that file
|
|
20
|
+
git diff <file> # verify clean
|
|
21
|
+
|
|
22
|
+
# --- tmux sessions ---
|
|
23
|
+
tmux kill-session -t <session-name>
|
|
24
|
+
tmux ls # confirm gone
|
|
25
|
+
|
|
26
|
+
# --- Temp fixtures / scratch scripts ---
|
|
27
|
+
rm -f /tmp/debug-*.*
|
|
28
|
+
ls /tmp/debug-*.* 2>/dev/null # confirm gone (ls returns non-zero when no match)
|
|
29
|
+
|
|
30
|
+
# --- Background processes (debugger-attached runtimes) ---
|
|
31
|
+
pkill -f 'node --inspect' || true
|
|
32
|
+
pkill -f 'python -m pdb' || true
|
|
33
|
+
pkill -f 'debugpy' || true
|
|
34
|
+
pkill -f 'dlv' || true
|
|
35
|
+
pkill -f 'gdb' || true
|
|
36
|
+
pkill -f 'lldb' || true
|
|
37
|
+
|
|
38
|
+
# --- Debug-relevant ports confirmed free ---
|
|
39
|
+
lsof -iTCP:9229 -sTCP:LISTEN -nP 2>/dev/null # Node inspector default
|
|
40
|
+
lsof -iTCP:5678 -sTCP:LISTEN -nP 2>/dev/null # debugpy default
|
|
41
|
+
lsof -iTCP:2345 -sTCP:LISTEN -nP 2>/dev/null # dlv default
|
|
42
|
+
lsof -iTCP:9999 -sTCP:LISTEN -nP 2>/dev/null # pwndbg/gdb-server default
|
|
43
|
+
|
|
44
|
+
# --- Env var overrides in current shell ---
|
|
45
|
+
unset DEBUG_OVERRIDE_FOO
|
|
46
|
+
unset PYTHONBREAKPOINT
|
|
47
|
+
unset RUST_LOG
|
|
48
|
+
unset DEBUG
|
|
49
|
+
|
|
50
|
+
# --- Ghidra scratch projects (if created just for this session) ---
|
|
51
|
+
# rm -rf ~/ghidra-projects/debug-scratch
|
|
52
|
+
|
|
53
|
+
# --- Core dumps from debugging (if any) ---
|
|
54
|
+
rm -f ./core ./core.* ~/core.*
|
|
55
|
+
|
|
56
|
+
# --- Playwright trace files ---
|
|
57
|
+
rm -rf playwright-report/ test-results/
|
|
58
|
+
```
|
|
59
|
+
|
|
60
|
+
### The verify command
|
|
61
|
+
|
|
62
|
+
This is the single most important check of the whole skill:
|
|
63
|
+
|
|
64
|
+
```bash
|
|
65
|
+
git status
|
|
66
|
+
git diff --stat
|
|
67
|
+
```
|
|
68
|
+
|
|
69
|
+
The diff must contain **only**:
|
|
70
|
+
|
|
71
|
+
1. The real fix.
|
|
72
|
+
2. The new failing-first test.
|
|
73
|
+
3. Nothing else.
|
|
74
|
+
|
|
75
|
+
### Detector checklist — scan the diff for these
|
|
76
|
+
|
|
77
|
+
If `git status` shows any untracked debug file, or `git diff` shows any of the patterns below, **you are not done**. Clean it.
|
|
78
|
+
|
|
79
|
+
| Pattern | Usually means |
|
|
80
|
+
|---|---|
|
|
81
|
+
| `debugger;` | Node debug statement left behind |
|
|
82
|
+
| `breakpoint()` | Python debug statement left behind |
|
|
83
|
+
| `dbg!(...)` | Rust debug macro left behind |
|
|
84
|
+
| `fmt.Println("DEBUG: ...")` | Go ad-hoc print |
|
|
85
|
+
| `console.log("[DEBUG]` | Node ad-hoc log |
|
|
86
|
+
| `print(f"DEBUG: ` | Python ad-hoc print |
|
|
87
|
+
| `// TODO DEBUG`, `// HACK`, `// XXX` | Stale debug marker |
|
|
88
|
+
| `// <PROJECT>-DEBUG` | Session-specific marker from this skill's edits |
|
|
89
|
+
| Commented-out code blocks near the fix | Dead code from trial fixes |
|
|
90
|
+
| Reordered imports or formatting in unrelated files | Drift from your editor's autoformat during the session |
|
|
91
|
+
|
|
92
|
+
### Remove the journal
|
|
93
|
+
|
|
94
|
+
Only once the git check is clean:
|
|
95
|
+
|
|
96
|
+
```bash
|
|
97
|
+
rm .debug-journal.md
|
|
98
|
+
sed -i.bak '/^\.debug-journal\.md$/d' .git/info/exclude && rm -f .git/info/exclude.bak
|
|
99
|
+
```
|
|
100
|
+
|
|
101
|
+
The journal is not part of the fix; it doesn't belong in the commit or in the git exclude list.
|
|
102
|
+
|
|
103
|
+
---
|
|
104
|
+
|
|
105
|
+
## Phase 10 — Final Verification
|
|
106
|
+
|
|
107
|
+
Last gate before reporting done. All four gates must be true, and all four must have **evidence in your final message** to the user. Passing a gate without evidence is the same as failing it.
|
|
108
|
+
|
|
109
|
+
### The four gates
|
|
110
|
+
|
|
111
|
+
1. **Red→green toggle confirmed** — show the failing test output from before the fix and passing output after. Both outputs visible in the reply or the journal.
|
|
112
|
+
|
|
113
|
+
2. **Full test suite green** — show the suite's final pass line (e.g. `42 passed in 3.14s`). Not just the new test.
|
|
114
|
+
|
|
115
|
+
3. **Manual QA reproduced the fix** — show the command or scenario that originally failed and its now-correct output. Verbatim, not paraphrased.
|
|
116
|
+
|
|
117
|
+
4. **Working tree clean of debug artifacts** — show `git diff --stat` output containing only fix + test, plus `git status` clean of untracked debug files.
|
|
118
|
+
|
|
119
|
+
If any of the four lacks evidence, you have not finished — return to the appropriate phase.
|
|
120
|
+
|
|
121
|
+
### Final message template
|
|
122
|
+
|
|
123
|
+
Keep it short. Evidence-dense. The user should be able to skim it in 30 seconds.
|
|
124
|
+
|
|
125
|
+
```markdown
|
|
126
|
+
Fixed.
|
|
127
|
+
|
|
128
|
+
**Root cause**: <one sentence — the mechanism, not the symptom>
|
|
129
|
+
**Fix**: `<file:line>` — <two words>
|
|
130
|
+
**Test**: `<test file>::<test name>` — red without fix, green with fix
|
|
131
|
+
**QA**: <one line describing what you ran and what you saw>
|
|
132
|
+
|
|
133
|
+
Diff:
|
|
134
|
+
```
|
|
135
|
+
<git diff --stat output — should be tiny>
|
|
136
|
+
```
|
|
137
|
+
|
|
138
|
+
**Next steps I didn't take** (awaiting your decision):
|
|
139
|
+
- <follow-up 1, if any — from QA silent-failure scan or refactor opportunities noted during Phase 7>
|
|
140
|
+
- <follow-up 2 — or "none" if nothing else surfaced>
|
|
141
|
+
```
|
|
142
|
+
|
|
143
|
+
### Example (from a real session)
|
|
144
|
+
|
|
145
|
+
```markdown
|
|
146
|
+
Fixed.
|
|
147
|
+
|
|
148
|
+
**Root cause**: pi-mono Agent's `model.baseUrl` was hardcoded to `api.anthropic.com`, so the `ANTHROPIC_BASE_URL` env var was silently ignored. The proxy API key was rejected by the real Anthropic API with 401, but pi-mono packaged the error into the assistant message's `errorMessage` field instead of throwing, so the route's try/catch never fired and the client received HTTP 200 with empty content.
|
|
149
|
+
|
|
150
|
+
**Fix**: `core/pi-bridge/modelResolver.ts:117` — override baseUrl
|
|
151
|
+
**Test**: `__tests__/core/modelResolver.test.ts::resolves_env_override` — red without fix, green with fix
|
|
152
|
+
**QA**: `curl -X POST /api/refinement/chat` with proxy env set, observed non-zero usage and non-empty content
|
|
153
|
+
|
|
154
|
+
Diff:
|
|
155
|
+
```
|
|
156
|
+
core/pi-bridge/modelResolver.ts | 3 +++
|
|
157
|
+
__tests__/core/modelResolver.test.ts | 42 ++++++++++++++++++++++
|
|
158
|
+
2 files changed, 45 insertions(+)
|
|
159
|
+
```
|
|
160
|
+
|
|
161
|
+
**Next steps I didn't take** (awaiting your decision):
|
|
162
|
+
- pi-mono itself silently swallows LLM errors into `errorMessage`; adding a throw-on-error wrapper at our orchestrator layer would surface these upstream
|
|
163
|
+
- Same silent-failure pattern exists in the planning route — likely the same fix applies
|
|
164
|
+
```
|
|
@@ -0,0 +1,228 @@
|
|
|
1
|
+
# Partial Runtime Evidence — When You Cannot Execute the Real Operation
|
|
2
|
+
|
|
3
|
+
Read this when **runtime truth beats code reading** is in conflict with **you cannot run the actual operation**.
|
|
4
|
+
|
|
5
|
+
The skill's first invariant is "runtime state is the only source of truth." But sometimes the only state you can produce is a *partial* observation — the real call requires paid credits, a hardware device you don't have, network access through a corporate proxy, a production secret, or a customer dataset.
|
|
6
|
+
|
|
7
|
+
**Partial runtime evidence is still runtime evidence.** This reference tells you which partial signals to harvest and how to combine them so the conclusion is defensible.
|
|
8
|
+
|
|
9
|
+
---
|
|
10
|
+
|
|
11
|
+
## When this applies
|
|
12
|
+
|
|
13
|
+
Use this reference when ALL are true:
|
|
14
|
+
|
|
15
|
+
1. The bug or extraction question requires runtime confirmation (per skill invariant #1).
|
|
16
|
+
2. You attempted the obvious "just run it" path and it failed for reasons unrelated to the bug:
|
|
17
|
+
- 401/402/403 from a paid API
|
|
18
|
+
- "device not found" / "permission denied" / SIP block
|
|
19
|
+
- Production-only credentials
|
|
20
|
+
- Network isolation (air-gapped, behind VPN you don't have)
|
|
21
|
+
- Time-of-day or quota limits
|
|
22
|
+
3. **Mocking the entire system** would defeat the verification — you specifically need evidence about how the *real* code behaves, not a stub.
|
|
23
|
+
|
|
24
|
+
If only #1 and #2 are true and you can mock cleanly, just mock and proceed. This file is for cases where mocking would invalidate the answer.
|
|
25
|
+
|
|
26
|
+
---
|
|
27
|
+
|
|
28
|
+
## The hierarchy of partial evidence (strongest first)
|
|
29
|
+
|
|
30
|
+
When you cannot capture the full outbound payload + full response, capture as much as possible from this list. **Evidence further down the list has more inference; evidence higher up is closer to ground truth.**
|
|
31
|
+
|
|
32
|
+
### Tier 1 — Pre-send / post-receive logs (best partial evidence)
|
|
33
|
+
|
|
34
|
+
The system you're investigating builds a request, then sends it. If the build step logs the assembled request **before** transmission, that log is ground truth for everything except the wire-level bytes (TLS, headers added by HTTP library, etc.).
|
|
35
|
+
|
|
36
|
+
```bash
|
|
37
|
+
# Maximize debug logging
|
|
38
|
+
APP_DEBUG=1 APP_LOG_LEVEL=debug APP_LOG_FILE=/tmp/trace.log ./target -x "minimal valid input" 2>&1 | head -200
|
|
39
|
+
```
|
|
40
|
+
|
|
41
|
+
Look for log lines like:
|
|
42
|
+
- `Building request: model=X, params={...}`
|
|
43
|
+
- `[provider] payload: {...}`
|
|
44
|
+
- `Sending to <url>: <serialized body>`
|
|
45
|
+
|
|
46
|
+
**Strength**: 95% of ground truth. Missing only wire-level transformations.
|
|
47
|
+
|
|
48
|
+
### Tier 2 — Local interception via proxy / shim
|
|
49
|
+
|
|
50
|
+
Run the real binary against a local proxy that records and (optionally) returns a canned response.
|
|
51
|
+
|
|
52
|
+
```bash
|
|
53
|
+
# mitmproxy approach
|
|
54
|
+
mitmproxy --listen-host 127.0.0.1 --listen-port 8888 --mode regular &
|
|
55
|
+
HTTPS_PROXY=http://127.0.0.1:8888 SSL_CERT_FILE=~/.mitmproxy/mitmproxy-ca-cert.pem ./target ...
|
|
56
|
+
# Now mitmproxy logs the actual TLS-decrypted request
|
|
57
|
+
```
|
|
58
|
+
|
|
59
|
+
```bash
|
|
60
|
+
# DYLD_INSERT_LIBRARIES / LD_PRELOAD shim approach
|
|
61
|
+
# Wrap the network call to log payload, return a fake 200
|
|
62
|
+
# See pwntools.md for shim examples
|
|
63
|
+
```
|
|
64
|
+
|
|
65
|
+
**Strength**: Wire-level ground truth, but requires the target to honor your proxy / preload.
|
|
66
|
+
|
|
67
|
+
### Tier 3 — Static extraction × runtime fingerprint cross-check
|
|
68
|
+
|
|
69
|
+
When you cannot send a request at all, you can still cross-check static analysis with whatever the binary does that *doesn't* require the real call:
|
|
70
|
+
|
|
71
|
+
- The binary builds the request — even if sending fails, the build step ran. Trace it (Tier 1).
|
|
72
|
+
- The binary writes a state file or cache — read it.
|
|
73
|
+
- The binary emits version-specific User-Agent strings; verify they match your static extraction.
|
|
74
|
+
- The binary's `--help` or `--version` output reveals build metadata; verify model lists / feature flags.
|
|
75
|
+
|
|
76
|
+
**Strength**: Disjoint evidence sources confirming the same fact. Two independent partial signals that agree are nearly as strong as one full observation.
|
|
77
|
+
|
|
78
|
+
### Tier 4 — Contrastive runtime under different inputs
|
|
79
|
+
|
|
80
|
+
If you can run with input variant A but not B, run A and reason about B from code:
|
|
81
|
+
|
|
82
|
+
```bash
|
|
83
|
+
# A: minimal trial input — works for free tier
|
|
84
|
+
./target --action=read --resource=local-file
|
|
85
|
+
# B: full inference call — paid tier required, blocked
|
|
86
|
+
# But the request-building code is shared between A and B!
|
|
87
|
+
# Capture A's logs, then inspect the code path for B and verify only the model/endpoint diff.
|
|
88
|
+
```
|
|
89
|
+
|
|
90
|
+
**Strength**: Confirms shared code paths; remaining gap is only the difference between A and B.
|
|
91
|
+
|
|
92
|
+
### Tier 5 — Vendor-published API logs / dashboard
|
|
93
|
+
|
|
94
|
+
If the operation succeeded earlier (before quota ran out, before access was revoked), the vendor's dashboard / audit log may show the request. Lower fidelity but still observed behavior.
|
|
95
|
+
|
|
96
|
+
**Strength**: Real wire data, but often summarized — token counts, status codes, no payload bodies.
|
|
97
|
+
|
|
98
|
+
### Tier 6 — Pure code reading with peer review
|
|
99
|
+
|
|
100
|
+
If literally none of the above is available, read the code carefully and submit it to **one Oracle for skeptical review** (see "Verification Oracle" below). This is the weakest tier and you must explicitly mark conclusions as "unverified" in the journal.
|
|
101
|
+
|
|
102
|
+
---
|
|
103
|
+
|
|
104
|
+
## How to combine partial signals
|
|
105
|
+
|
|
106
|
+
A defensible conclusion **prefers two independent signals from different tiers**, with one exception: a complete Tier 2 wire-level capture is wire-level ground truth and can stand alone for request-shape claims (because the wire bytes are exactly what the remote received). For *behavioral* claims (what the system does next, what state it stores, what side effects it produces), still combine with another signal.
|
|
107
|
+
|
|
108
|
+
| Available evidence | Defensibility |
|
|
109
|
+
|---|---|
|
|
110
|
+
| Tier 1 + Tier 1 (same log, different lines) | weak — single source |
|
|
111
|
+
| Tier 1 + Tier 2 (debug log + proxy capture) | **strong** — independent confirmation |
|
|
112
|
+
| Tier 1 + Tier 3 (debug log + version output cross-check) | **strong** — disjoint sources |
|
|
113
|
+
| Tier 2 alone (full proxy capture) | strong **for request-shape claims only** — stands alone for "what bytes were sent". Add a second signal for response-handling or state claims. |
|
|
114
|
+
| Tier 3 + Tier 4 (cross-check + contrastive run) | medium — both partial |
|
|
115
|
+
| Tier 6 alone (code reading only) | **insufficient** — escalate or mark unverified |
|
|
116
|
+
|
|
117
|
+
Record in the journal:
|
|
118
|
+
|
|
119
|
+
```markdown
|
|
120
|
+
## Partial runtime evidence
|
|
121
|
+
### Question being verified
|
|
122
|
+
<the specific claim, e.g. "Opus 4.7 default effort is 'high'">
|
|
123
|
+
|
|
124
|
+
### Available signals
|
|
125
|
+
- Tier 1: debug log /tmp/trace.log line 47-49 shows `effort: "high"` ✓
|
|
126
|
+
- Tier 3: static extraction of m5T() function returns "high" for smart mode ✓
|
|
127
|
+
- Tier 6: code path verified by reading prompt-builder.js ✓
|
|
128
|
+
|
|
129
|
+
### Independence assessment
|
|
130
|
+
Tier 1 and Tier 3 are independent — the log was emitted by a different
|
|
131
|
+
code path than m5T() and would diverge if the static reading were wrong.
|
|
132
|
+
|
|
133
|
+
### Conclusion
|
|
134
|
+
VERIFIED via Tier 1 + Tier 3 agreement. No need to escalate.
|
|
135
|
+
```
|
|
136
|
+
|
|
137
|
+
If you cannot achieve a complete Tier 2 capture **or** two independent non-Tier-6 signals from the table above, **write an explicit note in the deliverable**:
|
|
138
|
+
|
|
139
|
+
> ⚠️ Partial-evidence finding. The full outbound payload could not be captured because [reason]. The conclusion rests on:
|
|
140
|
+
> - [signal A — tier and source]
|
|
141
|
+
> - [signal B — tier and source]
|
|
142
|
+
> A future verification should attempt [the missing tier] when [condition].
|
|
143
|
+
|
|
144
|
+
---
|
|
145
|
+
|
|
146
|
+
## Verification Oracle pattern (for non-debug tasks)
|
|
147
|
+
|
|
148
|
+
The skill's main Oracle Triple (`04-oracle-triple.md`) is for **stuck debugging** — 2 failed rounds, mental box, three orthogonal framings to break out.
|
|
149
|
+
|
|
150
|
+
For tasks where the deliverable is an **artifact, not a bug fix** (reverse engineering, extraction, audit, compliance documentation), use a different pattern: **single Oracle, late, skeptical, with the deliverable in hand**.
|
|
151
|
+
|
|
152
|
+
### When to invoke
|
|
153
|
+
|
|
154
|
+
- Right before declaring an extraction/audit task "done"
|
|
155
|
+
- After every significant revision of the deliverable (not after every small edit)
|
|
156
|
+
- Maximum 3-4 iterations before escalating to user
|
|
157
|
+
|
|
158
|
+
### Pattern
|
|
159
|
+
|
|
160
|
+
Use a Grok Build `explore` lane or verifier subagent with this prompt shape:
|
|
161
|
+
|
|
162
|
+
```text
|
|
163
|
+
SKEPTICAL FINAL VERIFICATION — be critical, look for reasons the task is incomplete or wrong.
|
|
164
|
+
|
|
165
|
+
## Original task
|
|
166
|
+
<verbatim user request>
|
|
167
|
+
|
|
168
|
+
## What I produced
|
|
169
|
+
<list of artifacts with paths and brief descriptions>
|
|
170
|
+
|
|
171
|
+
## Specific claims to verify
|
|
172
|
+
<bullet list of every concrete claim in the deliverable>
|
|
173
|
+
|
|
174
|
+
## Where to look
|
|
175
|
+
<paths the Oracle should Read / Bash to verify>
|
|
176
|
+
|
|
177
|
+
## Your job
|
|
178
|
+
1. Read the deliverables.
|
|
179
|
+
2. Spot-check each claim against the source/evidence the deliverable cites.
|
|
180
|
+
3. Identify any unsubstantiated claims, missing pieces, or factual errors.
|
|
181
|
+
4. End with PASS / FAIL / PARTIAL with specific gaps.
|
|
182
|
+
Be skeptical. Don't rubber-stamp.
|
|
183
|
+
```
|
|
184
|
+
|
|
185
|
+
### Why this differs from the Oracle Triple
|
|
186
|
+
|
|
187
|
+
| | Oracle Triple (debug) | Verification Oracle (artifact) |
|
|
188
|
+
|---|---|---|
|
|
189
|
+
| Trigger | 2 failed hypothesis rounds | About to declare "done" |
|
|
190
|
+
| Count | 3 in parallel, orthogonal framings | 1 sequential, focused review |
|
|
191
|
+
| Goal | Break out of mental box | Catch unsubstantiated claims |
|
|
192
|
+
| Tone of prompt | Brainstorm wide alternatives | Skeptical audit |
|
|
193
|
+
| Iteration | Reset hypothesis set after | Fix gaps, re-invoke until PASS |
|
|
194
|
+
|
|
195
|
+
### Don't conflate them
|
|
196
|
+
|
|
197
|
+
If you're stuck debugging, do the Triple. If you have a deliverable and need it audited, do the Verification Oracle. Doing the Triple on a finished extraction will return three diverging "what if you tried…" tangents that are not what you need. Doing the Verification Oracle on a stuck debugging session will return a polite "the evidence is incomplete" that you already knew.
|
|
198
|
+
|
|
199
|
+
---
|
|
200
|
+
|
|
201
|
+
## Common partial-evidence anti-patterns
|
|
202
|
+
|
|
203
|
+
| Anti-pattern | Why it fails | Replacement |
|
|
204
|
+
|---|---|---|
|
|
205
|
+
| "It looks right in the code, so it works" | Tier 6 alone, unverified | Add at least one Tier 1-3 signal |
|
|
206
|
+
| "I ran it once, didn't error, so it's correct" | Absence of error ≠ presence of correctness | Capture the actual output and verify content |
|
|
207
|
+
| "The mock returns the value I wrote, so the code is fine" | Tautology — mock loops back your assumption | Use Tier 2 (proxy) instead, or cross-check with Tier 3 |
|
|
208
|
+
| "The vendor's dashboard shows my call worked" | Dashboard often only shows status code, not behavior | Combine with Tier 1 if available |
|
|
209
|
+
| "I'll trust the most-recent stack overflow answer" | Code from a different version / context | Verify against the actual binary you have |
|
|
210
|
+
|
|
211
|
+
---
|
|
212
|
+
|
|
213
|
+
## Cleanup additions for partial-evidence work
|
|
214
|
+
|
|
215
|
+
```bash
|
|
216
|
+
# Proxy artifacts
|
|
217
|
+
pkill -f mitmproxy 2>/dev/null
|
|
218
|
+
rm -f ~/.mitmproxy/cache_* 2>/dev/null
|
|
219
|
+
|
|
220
|
+
# Debug log files
|
|
221
|
+
rm -f /tmp/trace.log /tmp/*-debug-trace.log
|
|
222
|
+
|
|
223
|
+
# DYLD_INSERT / LD_PRELOAD shim libraries
|
|
224
|
+
rm -f /tmp/*.dylib /tmp/*.so
|
|
225
|
+
|
|
226
|
+
# Verify env vars set in your shell are not persisted
|
|
227
|
+
unset HTTPS_PROXY APP_DEBUG APP_LOG_LEVEL APP_LOG_FILE 2>/dev/null
|
|
228
|
+
```
|
|
@@ -0,0 +1,55 @@
|
|
|
1
|
+
# PostToolUseFailure evidence taxonomy
|
|
2
|
+
|
|
3
|
+
Load this taxonomy after a failed tool call when the visible failure and the session ledger must be classified before a fix is considered. It is keyed to Grok Build's documented PostToolUseFailure event without inventing an error-result field.
|
|
4
|
+
|
|
5
|
+
## Documented evidence envelope
|
|
6
|
+
|
|
7
|
+
A tool event can provide hookEventName, sessionId, cwd, workspaceRoot, toolName, and toolInput on stdin. The hook environment provides the event, hook name, session id, and workspace root. The documented envelope does not name an error message, exit code, stderr, duration, subagent id, or result object.
|
|
8
|
+
|
|
9
|
+
Therefore the event proves that the host classified a tool call as failed and identifies its invocation context. The visible tool transcript, command status, generated artifacts, and passive session ledger must supply any additional evidence. Absence of a ledger record does not prove success: passive hooks can be missing, untrusted, timed out, crashed, or disabled.
|
|
10
|
+
|
|
11
|
+
## Failure taxonomy
|
|
12
|
+
|
|
13
|
+
| Failure class | Presentation | PostToolUseFailure correlation | Next discriminating action | Do not conclude |
|
|
14
|
+
| --- | --- | --- | --- | --- |
|
|
15
|
+
| input validation | tool rejects empty, malformed, conflicting, or out-of-range input | toolName and inert toolInput identify the attempted invocation | reduce to the smallest invalid field and compare with documented or tested input | the implementation is broken before valid input is tried |
|
|
16
|
+
| discovery | path, command, rule, skill, hook, or config cannot be found | event cwd/workspace may reveal lookup root | resolve actual cwd, inspect discovery rules, and test exact path existence | missing at one lookup root means missing everywhere |
|
|
17
|
+
| permission denial | permission system prevents the tool from running | failure event may exist, but permission UI or denial text is primary | identify matching deny/ask rule and whether the call was attempted | sandbox or program logic caused the denial |
|
|
18
|
+
| sandbox denial | approved process crosses filesystem or child-network boundary | tool invocation is visible; host event lacks the blocked path unless present in input | record active profile and reproduce with the same permitted target | disabling the sandbox is a program fix |
|
|
19
|
+
| executable missing | shell or command cannot locate a program | toolName names the shell-like tool, not the missing executable | inspect PATH and resolve the named command without installing anything | a similarly named binary is the required dependency |
|
|
20
|
+
| working-directory mismatch | relative path or package command resolves against the wrong root | compare event cwd and workspaceRoot with intended product root | run a read-only root/status check, then repeat the exact command from documented cwd | the file itself is absent or corrupt |
|
|
21
|
+
| program exit | process starts and returns a non-zero status | event establishes failed tool call; transcript supplies status/stderr | preserve exit and reduce program input while holding environment constant | every non-zero exit is the same failure class |
|
|
22
|
+
| assertion mismatch | test runs but observed value differs from expected | correlate event with the precise test command and transcript | run the narrow test and identify first semantic mismatch | the entire suite is invalid |
|
|
23
|
+
| timeout or hang | foreground result does not arrive within the bounded interval | event may appear only after the host classifies failure | distinguish timeout, background continuation, deadlock, and slow completion | elapsed time alone proves termination |
|
|
24
|
+
| cancellation | user or parent cancels a running action | ledger ordering can show the turn boundary but not invented cancellation details | inspect task state and any partial artifacts before retry | cancellation preserved atomicity |
|
|
25
|
+
| partial mutation | call fails after creating or changing some state | event input names intended operation, not completed writes | inventory target bytes, temp files, locks, and external effects before recovery | failure means nothing changed |
|
|
26
|
+
| environment skew | version, variable, executable, architecture, or cwd differs from expected baseline | session and workspace fields help bind the observation | record exact version and environment fact, then compare one dimension | a clean-HEAD run automatically matches the dirty environment |
|
|
27
|
+
| hook recorder failure | passive recorder itself exits non-zero or never records | visible tool failure may exist without a usable ledger line | drive the hook directly with documented JSON and inspect stderr/state | no record means no tool failure |
|
|
28
|
+
| evidence ambiguity | transcript is truncated, overwritten, stale, or cannot be attributed | event envelope may identify session/cwd but not missing result | rerun a smaller safe reproducer with bounded output | the most plausible explanation is proven |
|
|
29
|
+
|
|
30
|
+
## Correlation rules
|
|
31
|
+
|
|
32
|
+
Correlate on more than the event name:
|
|
33
|
+
|
|
34
|
+
1. sessionId must name the active investigation.
|
|
35
|
+
2. cwd must match the directory from which the failure was attempted.
|
|
36
|
+
3. workspaceRoot must match the project being diagnosed.
|
|
37
|
+
4. toolName must match the visible failed tool call.
|
|
38
|
+
5. toolInput must be treated as sensitive inert data and compared only as needed.
|
|
39
|
+
6. Ledger ordering must not be mistaken for a host timestamp contract unless a product-owned recorder added its own timestamp.
|
|
40
|
+
|
|
41
|
+
If any field disagrees, record a correlation conflict rather than merging the evidence. Multiple failures in one session require ordering and a stable identifier from product-owned evidence; do not invent one in the host event.
|
|
42
|
+
|
|
43
|
+
## Classification procedure
|
|
44
|
+
|
|
45
|
+
Start at the earliest failed boundary. Ask whether the tool was invoked. If not, classify permission or discovery first. If invoked, ask whether the operating boundary rejected access. If access was possible, inspect process start, program exit, assertions, and timing. After any mutating call, check partial state before retrying.
|
|
46
|
+
|
|
47
|
+
Choose one primary class and any contributing classes. For example, “program exit caused by working-directory mismatch” is more actionable than two unrelated labels. Each hypothesis must predict a probe result. Keep the same input while changing one environmental dimension, or keep the same environment while reducing one input dimension.
|
|
48
|
+
|
|
49
|
+
## Evidence packet
|
|
50
|
+
|
|
51
|
+
For each classified failure, record the visible symptom and expected result, exact invocation and cwd, primary and contributing classes, PostToolUseFailure fields actually observed, permission and sandbox state, smallest reproducer and status, partial-state inventory, supported and contradicted hypotheses, next safe probe, and inline limitation where evidence is missing.
|
|
52
|
+
|
|
53
|
+
## Fix gate
|
|
54
|
+
|
|
55
|
+
A code fix is justified only when the reproducer fails for the predicted reason, the proposed cause explains that observation, and a focused regression can turn GREEN without weakening permissions or sandbox settings. Configuration, environment, and authorization failures need their own remedies; rewriting program code can hide them without solving them.
|