@kensaurus/skills 0.0.0-stage → 2.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude-plugin/marketplace.json +53 -0
- package/.claude-plugin/plugin.json +40 -0
- package/.cursor-plugin/plugin.json +38 -0
- package/.mcp.json +28 -0
- package/CHANGELOG.md +1761 -0
- package/LICENSE +21 -0
- package/NOTICE +13 -0
- package/README.md +820 -2
- package/SECURITY.md +55 -0
- package/agents/code-reviewer.md +60 -0
- package/agents/completion-judge.md +89 -0
- package/agents/db-migrator.md +125 -0
- package/agents/debugger.md +47 -0
- package/agents/deploy-checker.md +100 -0
- package/agents/perf-monitor.md +74 -0
- package/assets/favicon.png +0 -0
- package/assets/logo-light.png +0 -0
- package/assets/logo.png +0 -0
- package/assets/logo.svg +6 -0
- package/assets/og.png +0 -0
- package/bin/install.mjs +1127 -0
- package/bin/kenji.js +2 -0
- package/commands/adr.md +17 -0
- package/commands/aeo-plan.md +18 -0
- package/commands/arch-boundaries.md +17 -0
- package/commands/aso-plan.md +18 -0
- package/commands/auth-flows.md +19 -0
- package/commands/backup-plan.md +17 -0
- package/commands/burndown-full.md +25 -0
- package/commands/capacitor-plan.md +18 -0
- package/commands/codemod-safety.md +22 -0
- package/commands/commit.md +18 -0
- package/commands/complete-everything.md +40 -0
- package/commands/cost-plan.md +19 -0
- package/commands/deadcode-plan.md +26 -0
- package/commands/deadcode.md +32 -0
- package/commands/debug-issue.md +17 -0
- package/commands/deps-plan.md +18 -0
- package/commands/docs-plan.md +17 -0
- package/commands/doctrine.md +19 -0
- package/commands/error-plan.md +19 -0
- package/commands/feedback-to-closure.md +36 -0
- package/commands/fix-issue.md +76 -0
- package/commands/gate-logic.md +26 -0
- package/commands/green-repo.md +36 -0
- package/commands/grill-me.md +19 -0
- package/commands/gtm-plan.md +21 -0
- package/commands/gtm-weekly.md +17 -0
- package/commands/gtm.md +22 -0
- package/commands/handoff.md +15 -0
- package/commands/housekeep-backlog.md +18 -0
- package/commands/housekeep-files.md +22 -0
- package/commands/housekeep-gates.md +18 -0
- package/commands/instant-nav.md +11 -0
- package/commands/integrity-plan.md +19 -0
- package/commands/launch-kit.md +16 -0
- package/commands/mcp-guide.md +40 -0
- package/commands/mobile-plan.md +19 -0
- package/commands/native-rn-monorepo/README.md +78 -0
- package/commands/native-rn-monorepo/android-build.md +26 -0
- package/commands/native-rn-monorepo/android-install.md +32 -0
- package/commands/native-rn-monorepo/android-logcat.md +37 -0
- package/commands/native-rn-monorepo/ios-ci-logs.md +56 -0
- package/commands/native-rn-monorepo/ios-ci-status.md +52 -0
- package/commands/native-rn-monorepo/ios-ci-trigger.md +59 -0
- package/commands/native-rn-monorepo/rn-reset.md +50 -0
- package/commands/native-rn-monorepo/rn-ship-ios.md +67 -0
- package/commands/native-rn-monorepo/rn-verify.md +53 -0
- package/commands/perf-plan.md +18 -0
- package/commands/plan-mode.md +74 -0
- package/commands/pr.md +16 -0
- package/commands/pricing-plan.md +20 -0
- package/commands/privacy-plan.md +18 -0
- package/commands/readability.md +12 -0
- package/commands/readme.md +15 -0
- package/commands/refactor.md +15 -0
- package/commands/release-prep.md +17 -0
- package/commands/research.md +25 -0
- package/commands/responsive-audit.md +21 -0
- package/commands/review-code.md +18 -0
- package/commands/rls-plan.md +18 -0
- package/commands/secrets-plan.md +18 -0
- package/commands/security-plan.md +19 -0
- package/commands/ship-and-observe.md +36 -0
- package/commands/skill-conflicts.md +19 -0
- package/commands/slop-plan.md +18 -0
- package/commands/stub-plan.md +18 -0
- package/commands/test-mutation.md +16 -0
- package/commands/test-plan.md +17 -0
- package/commands/test.md +29 -0
- package/commands/thirdparty-web-interface-guidelines.md +185 -0
- package/commands/uiux-plan.md +18 -0
- package/commands/uiux.md +45 -0
- package/commands/update-deps.md +21 -0
- package/commands/validation-plan.md +19 -0
- package/commands-portable/fix-issue.md +72 -0
- package/commands-portable/plan-mode.md +92 -0
- package/commands-portable/research.md +91 -0
- package/docs/screenshots/README.md +5 -0
- package/docs/screenshots/audit-dark.png +0 -0
- package/docs/screenshots/build-dark.png +0 -0
- package/docs/screenshots/grill-dark.png +0 -0
- package/docs/screenshots/hero-dark.png +0 -0
- package/docs/screenshots/hero-light.png +0 -0
- package/docs/screenshots/ship-dark.png +0 -0
- package/docs/screenshots/src/showcase.html +320 -0
- package/hooks/completion-gate.mjs +258 -0
- package/hooks/cursor-hooks.json +13 -0
- package/hooks/hooks.json +15 -0
- package/install.sh +21 -0
- package/llms.txt +40 -0
- package/mcp/README.md +266 -0
- package/mcp/VERSIONS.md +41 -0
- package/mcp/mcp-full.json.template +124 -0
- package/mcp/mcp.json.template +29 -0
- package/mcp/pinned-versions.json +27 -0
- package/package.json +93 -4
- package/rules/approved-plan-execution.mdc +65 -0
- package/rules/full-stack-ship-discipline.mdc +37 -0
- package/rules/native-rn-monorepo/README.md +63 -0
- package/rules/native-rn-monorepo/_project.mdc +69 -0
- package/rules/native-rn-monorepo/native-android.mdc +72 -0
- package/rules/native-rn-monorepo/native-ios.mdc +61 -0
- package/rules/native-rn-monorepo/react-native-js.mdc +78 -0
- package/rules/native-rn-monorepo/web.mdc +60 -0
- package/rules/project-starter/components.mdc +54 -0
- package/rules/project-starter/data-fetching.mdc +77 -0
- package/rules/project-starter/git.mdc +41 -0
- package/rules/project-starter/supabase.mdc +37 -0
- package/rules/project-starter/tailwind.mdc +48 -0
- package/rules/project-starter/typescript.mdc +36 -0
- package/rules/project-starter/web-performance.mdc +42 -0
- package/rules/senior-engineer.mdc +30 -0
- package/rules/shell-first-search.mdc +19 -0
- package/rules/skill-workflows.mdc +35 -0
- package/rules/verification-before-completion.mdc +57 -0
- package/skills/audit-accessibility/SKILL.md +441 -0
- package/skills/audit-agent-speed/SKILL.md +181 -0
- package/skills/audit-agent-speed/scripts/stop-typecheck.mjs +151 -0
- package/skills/audit-analytics/SKILL.md +138 -0
- package/skills/audit-auth-flows/SKILL.md +267 -0
- package/skills/audit-backend-architecture/SKILL.md +266 -0
- package/skills/audit-backend-architecture/references/patterns.md +386 -0
- package/skills/audit-bundle-size/SKILL.md +296 -0
- package/skills/audit-cicd/SKILL.md +218 -0
- package/skills/audit-code-quality/SKILL.md +314 -0
- package/skills/audit-code-review/SKILL.md +289 -0
- package/skills/audit-codemod-safety/SKILL.md +159 -0
- package/skills/audit-db-schema/SKILL.md +465 -0
- package/skills/audit-db-schema/references/details.md +110 -0
- package/skills/audit-doctrine/SKILL.md +189 -0
- package/skills/audit-env-parity/SKILL.md +133 -0
- package/skills/audit-fe-api/SKILL.md +458 -0
- package/skills/audit-gate-logic/SKILL.md +219 -0
- package/skills/audit-i18n/SKILL.md +339 -0
- package/skills/audit-infra-cost/SKILL.md +142 -0
- package/skills/audit-langfuse-llm/SKILL.md +468 -0
- package/skills/audit-langfuse-llm/references/details.md +226 -0
- package/skills/audit-llm-security/SKILL.md +147 -0
- package/skills/audit-monetization-iap/SKILL.md +137 -0
- package/skills/audit-payment-system/SKILL.md +268 -0
- package/skills/audit-payment-system/references/checklist.md +283 -0
- package/skills/audit-performance/SKILL.md +383 -0
- package/skills/audit-performance/references/loading-priority-2026.md +81 -0
- package/skills/audit-realworld/SKILL.md +287 -0
- package/skills/audit-registry-listing/SKILL.md +122 -0
- package/skills/audit-resilience/SKILL.md +154 -0
- package/skills/audit-responsive/SKILL.md +221 -0
- package/skills/audit-responsive/references/checklist.md +166 -0
- package/skills/audit-security/SKILL.md +289 -0
- package/skills/audit-skill-conflicts/SKILL.md +178 -0
- package/skills/audit-ui-states/SKILL.md +146 -0
- package/skills/audit-uiux-design-system/SKILL.md +475 -0
- package/skills/audit-uiux-design-system/references/details.md +71 -0
- package/skills/audit-ux/SKILL.md +379 -0
- package/skills/audit-ux/references/details.md +245 -0
- package/skills/audit-ux-journeys/SKILL.md +215 -0
- package/skills/audit-ux-journeys/references/checklist.md +179 -0
- package/skills/backend-db-performance/SKILL.md +441 -0
- package/skills/backend-error-handling/SKILL.md +489 -0
- package/skills/backend-error-handling/references/details.md +58 -0
- package/skills/backend-observability/SKILL.md +88 -0
- package/skills/backend-patterns/SKILL.md +499 -0
- package/skills/backend-patterns/references/architecture-patterns.md +298 -0
- package/skills/backend-realtime/SKILL.md +403 -0
- package/skills/backend-realtime/references/patterns.md +74 -0
- package/skills/burndown-full/SKILL.md +174 -0
- package/skills/complete-everything/SKILL.md +295 -0
- package/skills/data-pipeline/SKILL.md +109 -0
- package/skills/data-visualization/SKILL.md +488 -0
- package/skills/debug-error/SKILL.md +322 -0
- package/skills/debug-fe-be-integration/SKILL.md +459 -0
- package/skills/debug-sentry-monitor/SKILL.md +497 -0
- package/skills/debug-sentry-monitor/references/details.md +165 -0
- package/skills/deploy-npm/SKILL.md +394 -0
- package/skills/deploy-npm/references/example-mushi-mushi.md +52 -0
- package/skills/deploy-verify/SKILL.md +489 -0
- package/skills/design-api/SKILL.md +379 -0
- package/skills/design-canvas/SKILL.md +155 -0
- package/skills/design-email/SKILL.md +370 -0
- package/skills/design-frontend/SKILL.md +143 -0
- package/skills/design-generative-art/SKILL.md +474 -0
- package/skills/design-mobile-first/SKILL.md +506 -0
- package/skills/design-motion/SKILL.md +333 -0
- package/skills/design-motion/references/delight-interactions.md +191 -0
- package/skills/design-prd/SKILL.md +443 -0
- package/skills/design-system/SKILL.md +457 -0
- package/skills/design-theme/SKILL.md +226 -0
- package/skills/design-theme/themes/tsumagoi-ranch.md +150 -0
- package/skills/docs-adr/SKILL.md +168 -0
- package/skills/docs-coauthor/SKILL.md +368 -0
- package/skills/docs-comparison-pages/SKILL.md +117 -0
- package/skills/docs-domain-modeling/SKILL.md +97 -0
- package/skills/docs-launch-kit/SKILL.md +139 -0
- package/skills/docs-writer/SKILL.md +469 -0
- package/skills/enhance-agent-guardrails/SKILL.md +164 -0
- package/skills/enhance-arch-boundaries/SKILL.md +154 -0
- package/skills/enhance-capacitor-ui/SKILL.md +463 -0
- package/skills/enhance-capacitor-ui/references/details.md +750 -0
- package/skills/enhance-email-deliverability/SKILL.md +143 -0
- package/skills/enhance-growth-loops/SKILL.md +122 -0
- package/skills/enhance-lifecycle-email/SKILL.md +130 -0
- package/skills/enhance-motion/SKILL.md +193 -0
- package/skills/enhance-onboarding/SKILL.md +148 -0
- package/skills/enhance-pwa/SKILL.md +304 -0
- package/skills/enhance-readability/SKILL.md +146 -0
- package/skills/enhance-readme/SKILL.md +496 -0
- package/skills/enhance-readme/package-lock.json +187 -0
- package/skills/enhance-readme/package.json +17 -0
- package/skills/enhance-readme/scripts/generate-readme-blocks.mjs +199 -0
- package/skills/enhance-readme/scripts/record-readme-tour.mjs +442 -0
- package/skills/enhance-skill-prompts/SKILL.md +167 -0
- package/skills/enhance-skill-prompts/references/exemplar-audit-auth-flows.md +311 -0
- package/skills/enhance-web-conversion/SKILL.md +155 -0
- package/skills/enhance-web-forms/SKILL.md +154 -0
- package/skills/enhance-web-instant-nav/SKILL.md +138 -0
- package/skills/enhance-web-instant-nav/references/bfcache-blockers.md +23 -0
- package/skills/enhance-web-instant-nav/references/early-hints.md +33 -0
- package/skills/enhance-web-instant-nav/references/speculation-rules.md +44 -0
- package/skills/enhance-web-landing/SKILL.md +459 -0
- package/skills/enhance-web-landing/references/details.md +773 -0
- package/skills/enhance-web-redesign/SKILL.md +228 -0
- package/skills/enhance-web-seo/SKILL.md +276 -0
- package/skills/enhance-web-ui/SKILL.md +473 -0
- package/skills/enhance-web-ui/references/details.md +674 -0
- package/skills/enhance-web-ux/HEURISTICS.md +242 -0
- package/skills/enhance-web-ux/PATTERNS.md +375 -0
- package/skills/enhance-web-ux/SKILL.md +464 -0
- package/skills/enhance-web-ux/examples.md +222 -0
- package/skills/enhance-web-ux/references/details.md +406 -0
- package/skills/enhance-web-web3d/SKILL.md +397 -0
- package/skills/enhance-web-web3d/references/css-canvas-effects.md +180 -0
- package/skills/handoff/SKILL.md +66 -0
- package/skills/housekeep-backlog/SKILL.md +149 -0
- package/skills/housekeep-dead-code/SKILL.md +387 -0
- package/skills/housekeep-dead-code/references/ratchet-ci.md +205 -0
- package/skills/housekeep-dead-code/references/supabase-hygiene.md +152 -0
- package/skills/housekeep-design/SKILL.md +207 -0
- package/skills/housekeep-files/SKILL.md +220 -0
- package/skills/housekeep-files/references/naming-and-catalog.md +86 -0
- package/skills/housekeep-files/scripts/housekeep-files.ps1 +360 -0
- package/skills/housekeep-files/scripts/housekeep-files.sh +238 -0
- package/skills/housekeep-gates/SKILL.md +174 -0
- package/skills/iterate-agent-harness/SKILL.md +137 -0
- package/skills/iterate-gtm-weekly/SKILL.md +103 -0
- package/skills/iterate-post-launch/SKILL.md +292 -0
- package/skills/meta-mcp-builder/SKILL.md +313 -0
- package/skills/meta-skill-creator/SKILL.md +304 -0
- package/skills/mobile-capacitor-platform/SKILL.md +104 -0
- package/skills/mobile-emulator-start/SKILL.md +296 -0
- package/skills/mobile-emulator-test/SKILL.md +491 -0
- package/skills/mobile-emulator-test/references/details.md +478 -0
- package/skills/mobile-rn-performance/SKILL.md +107 -0
- package/skills/mobile-rn-screen/SKILL.md +476 -0
- package/skills/mobile-rn-screen/references/details.md +785 -0
- package/skills/mushi-health/SKILL.md +206 -0
- package/skills/mushi-integration/SKILL.md +257 -0
- package/skills/plan-aeo-readiness/SKILL.md +166 -0
- package/skills/plan-antislop/SKILL.md +281 -0
- package/skills/plan-aso/SKILL.md +149 -0
- package/skills/plan-backup-dr/SKILL.md +131 -0
- package/skills/plan-capacitor-hardening/SKILL.md +217 -0
- package/skills/plan-data-integrity/SKILL.md +187 -0
- package/skills/plan-dead-code/SKILL.md +386 -0
- package/skills/plan-dead-code/references/knip-config.md +214 -0
- package/skills/plan-dead-code/references/output-templates.md +133 -0
- package/skills/plan-dead-code/references/preservation-contract.md +50 -0
- package/skills/plan-dead-code/references/residue-greps.md +84 -0
- package/skills/plan-dependency-provenance/SKILL.md +200 -0
- package/skills/plan-docs-sync/SKILL.md +143 -0
- package/skills/plan-docs-sync/references/drift-taxonomy.md +43 -0
- package/skills/plan-docs-sync/references/output-templates.md +33 -0
- package/skills/plan-docs-sync/references/preservation-contract.md +17 -0
- package/skills/plan-error-handling/SKILL.md +205 -0
- package/skills/plan-gtm/SKILL.md +276 -0
- package/skills/plan-gtm/references/benchmarks-2026.md +183 -0
- package/skills/plan-input-validation/SKILL.md +179 -0
- package/skills/plan-llm-cost-guardrails/SKILL.md +176 -0
- package/skills/plan-mobile-readiness/SKILL.md +171 -0
- package/skills/plan-perf-audit/SKILL.md +145 -0
- package/skills/plan-perf-audit/references/audit-scope.md +51 -0
- package/skills/plan-perf-audit/references/output-templates.md +33 -0
- package/skills/plan-perf-audit/references/preservation-contract.md +13 -0
- package/skills/plan-pricing/SKILL.md +173 -0
- package/skills/plan-privacy-compliance/SKILL.md +148 -0
- package/skills/plan-rls-audit/SKILL.md +231 -0
- package/skills/plan-secrets-audit/SKILL.md +181 -0
- package/skills/plan-security-audit/SKILL.md +168 -0
- package/skills/plan-security-audit/references/output-templates.md +36 -0
- package/skills/plan-security-audit/references/owasp-supabase-scope.md +55 -0
- package/skills/plan-security-audit/references/preservation-contract.md +18 -0
- package/skills/plan-stub-checker/SKILL.md +216 -0
- package/skills/plan-stub-checker/references/detection-methodology.md +75 -0
- package/skills/plan-stub-checker/references/detection-taxonomy.md +34 -0
- package/skills/plan-stub-checker/references/output-templates.md +63 -0
- package/skills/plan-stub-checker/references/preservation-contract.md +24 -0
- package/skills/plan-test-coverage/SKILL.md +170 -0
- package/skills/plan-test-coverage/references/methodology.md +54 -0
- package/skills/plan-test-coverage/references/output-templates.md +34 -0
- package/skills/plan-test-coverage/references/preservation-contract.md +15 -0
- package/skills/plan-uiux-unification/SKILL.md +230 -0
- package/skills/plan-uiux-unification/references/output-templates.md +67 -0
- package/skills/plan-uiux-unification/references/phase-workbook.md +85 -0
- package/skills/plan-uiux-unification/references/preservation-contract.md +24 -0
- package/skills/protocol-browser-anti-stall/SKILL.md +211 -0
- package/skills/protocol-browser-anti-stall/references/mcp-to-cli-map.md +113 -0
- package/skills/protocol-browser-anti-stall/references/playwright-session-coordination.md +170 -0
- package/skills/research/SKILL.md +422 -0
- package/skills/test-exploratory/SKILL.md +165 -0
- package/skills/test-exploratory/references/charter-template.md +29 -0
- package/skills/test-load/SKILL.md +126 -0
- package/skills/test-mutation/SKILL.md +160 -0
- package/skills/test-playwright/SKILL.md +354 -0
- package/skills/test-qa/SKILL.md +364 -0
- package/skills/test-qa/references/details.md +268 -0
- package/skills/test-red-team/SKILL.md +387 -0
- package/skills/test-red-team/references/owasp-attack-checklist.md +193 -0
- package/skills/test-unit/SKILL.md +259 -0
- package/skills/test-unit/references/details.md +267 -0
- package/skills/test-visual-regression/SKILL.md +132 -0
- package/skills/thirdparty-emil-design-eng/ATTRIBUTION.md +20 -0
- package/skills/thirdparty-emil-design-eng/SKILL.md +21 -0
- package/skills/thirdparty-emil-design-eng/references/emil-design-eng.md +676 -0
- package/skills/thirdparty-ui-ux-pro-max/ATTRIBUTION.md +22 -0
- package/skills/thirdparty-ui-ux-pro-max/SKILL.md +304 -0
- package/skills/thirdparty-ui-ux-pro-max/data/charts.csv +26 -0
- package/skills/thirdparty-ui-ux-pro-max/data/colors.csv +97 -0
- package/skills/thirdparty-ui-ux-pro-max/data/icons.csv +101 -0
- package/skills/thirdparty-ui-ux-pro-max/data/landing.csv +31 -0
- package/skills/thirdparty-ui-ux-pro-max/data/products.csv +97 -0
- package/skills/thirdparty-ui-ux-pro-max/data/react-performance.csv +45 -0
- package/skills/thirdparty-ui-ux-pro-max/data/stacks/astro.csv +54 -0
- package/skills/thirdparty-ui-ux-pro-max/data/stacks/flutter.csv +53 -0
- package/skills/thirdparty-ui-ux-pro-max/data/stacks/html-tailwind.csv +56 -0
- package/skills/thirdparty-ui-ux-pro-max/data/stacks/jetpack-compose.csv +53 -0
- package/skills/thirdparty-ui-ux-pro-max/data/stacks/nextjs.csv +53 -0
- package/skills/thirdparty-ui-ux-pro-max/data/stacks/nuxt-ui.csv +51 -0
- package/skills/thirdparty-ui-ux-pro-max/data/stacks/nuxtjs.csv +59 -0
- package/skills/thirdparty-ui-ux-pro-max/data/stacks/react-native.csv +52 -0
- package/skills/thirdparty-ui-ux-pro-max/data/stacks/react.csv +54 -0
- package/skills/thirdparty-ui-ux-pro-max/data/stacks/shadcn.csv +61 -0
- package/skills/thirdparty-ui-ux-pro-max/data/stacks/svelte.csv +54 -0
- package/skills/thirdparty-ui-ux-pro-max/data/stacks/swiftui.csv +51 -0
- package/skills/thirdparty-ui-ux-pro-max/data/stacks/vue.csv +50 -0
- package/skills/thirdparty-ui-ux-pro-max/data/styles.csv +68 -0
- package/skills/thirdparty-ui-ux-pro-max/data/typography.csv +58 -0
- package/skills/thirdparty-ui-ux-pro-max/data/ui-reasoning.csv +101 -0
- package/skills/thirdparty-ui-ux-pro-max/data/ux-guidelines.csv +100 -0
- package/skills/thirdparty-ui-ux-pro-max/data/web-interface.csv +31 -0
- package/skills/thirdparty-ui-ux-pro-max/scripts/core.py +253 -0
- package/skills/thirdparty-ui-ux-pro-max/scripts/design_system.py +1067 -0
- package/skills/thirdparty-ui-ux-pro-max/scripts/search.py +114 -0
- package/skills/thirdparty-web-interface-guidelines/ATTRIBUTION.md +23 -0
- package/skills/thirdparty-web-interface-guidelines/SKILL.md +190 -0
- package/skills/workflow-build-feature/SKILL.md +118 -0
- package/skills/workflow-coding-discipline/SKILL.md +140 -0
- package/skills/workflow-environment-ready/SKILL.md +128 -0
- package/skills/workflow-feature-flag/SKILL.md +262 -0
- package/skills/workflow-feedback-to-closure/SKILL.md +165 -0
- package/skills/workflow-fix-and-ship/SKILL.md +136 -0
- package/skills/workflow-git-commit/SKILL.md +200 -0
- package/skills/workflow-green-repo/SKILL.md +166 -0
- package/skills/workflow-grilling/SKILL.md +73 -0
- package/skills/workflow-gtm/SKILL.md +153 -0
- package/skills/workflow-housekeep/SKILL.md +453 -0
- package/skills/workflow-housekeep/references/templates.md +109 -0
- package/skills/workflow-launch-ready/SKILL.md +145 -0
- package/skills/workflow-merge-conflicts/SKILL.md +62 -0
- package/skills/workflow-onboard/SKILL.md +99 -0
- package/skills/workflow-parallel-agents/SKILL.md +164 -0
- package/skills/workflow-pr/SKILL.md +197 -0
- package/skills/workflow-quality-gate/SKILL.md +147 -0
- package/skills/workflow-refactor/SKILL.md +274 -0
- package/skills/workflow-release-prep/SKILL.md +207 -0
- package/skills/workflow-ship-and-observe/SKILL.md +164 -0
- package/skills/workflow-spec-tdd/SKILL.md +141 -0
- package/skills/workflow-spec-tdd/references/spec-template.md +126 -0
- package/skills/workflow-spec-tdd/references/tdd-patterns.md +167 -0
- package/skills-cursor/babysit/SKILL.md +17 -0
- package/skills-cursor/canvas/SKILL.md +142 -0
- package/skills-cursor/canvas/sdk/canvas-tokens.d.ts +235 -0
- package/skills-cursor/canvas/sdk/chart-primitives.d.ts +200 -0
- package/skills-cursor/canvas/sdk/dag-layout.d.ts +102 -0
- package/skills-cursor/canvas/sdk/diff-view.d.ts +130 -0
- package/skills-cursor/canvas/sdk/form-primitives.d.ts +194 -0
- package/skills-cursor/canvas/sdk/hooks.d.ts +117 -0
- package/skills-cursor/canvas/sdk/index.d.ts +47 -0
- package/skills-cursor/canvas/sdk/theme.d.ts +61 -0
- package/skills-cursor/canvas/sdk/todo-list.d.ts +49 -0
- package/skills-cursor/canvas/sdk/ui-primitives.d.ts +549 -0
- package/skills-cursor/canvas/sdk/ui-primitives.test.d.ts +2 -0
- package/skills-cursor/create-hook/SKILL.md +238 -0
- package/skills-cursor/create-rule/SKILL.md +185 -0
- package/skills-cursor/create-skill/SKILL.md +269 -0
- package/skills-cursor/create-skill/references/authoring-guide.md +182 -0
- package/skills-cursor/create-subagent/SKILL.md +228 -0
- package/skills-cursor/migrate-to-skills/SKILL.md +121 -0
- package/skills-cursor/shell/SKILL.md +22 -0
- package/skills-cursor/split-to-prs/SKILL.md +47 -0
- package/skills-cursor/statusline/SKILL.md +193 -0
- package/skills-cursor/update-cli-config/SKILL.md +85 -0
- package/skills-cursor/update-cursor-settings/SKILL.md +137 -0
- package/skills.sh.json +296 -0
|
@@ -0,0 +1,468 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: audit-langfuse-llm
|
|
3
|
+
description: >
|
|
4
|
+
PDCA quality audit of LLM features: traces, prompts, costs, evals,
|
|
5
|
+
grounding, hallucination. Use for "audit LLM quality", "check Langfuse",
|
|
6
|
+
or "audit AI costs". Jailbreaks → audit-llm-security. Cost caps →
|
|
7
|
+
plan-llm-cost-guardrails.
|
|
8
|
+
license: MIT
|
|
9
|
+
effort: high
|
|
10
|
+
---
|
|
11
|
+
|
|
12
|
+
# Langfuse LLM Quality Audit
|
|
13
|
+
|
|
14
|
+
**Degree of freedom: MIXED** — Phases 0–1, 4 `[HIGH freedom]`; Phases 2–3
|
|
15
|
+
CLI traces and playwright `[LOW freedom — run exactly]`. Read
|
|
16
|
+
`protocol-browser-anti-stall` before any browser step. Phase 0 first: Phases 2–3 verify the feature map it produces.
|
|
17
|
+
|
|
18
|
+
## How to reason
|
|
19
|
+
|
|
20
|
+
1. **Observe** — quote the trace, prompt version, token/cost, or live output
|
|
21
|
+
2. **Interpret** — is quality, cost, or the pipeline actually broken?
|
|
22
|
+
3. **Classify** — missing-trace / prompt / cost / eval / grounding / correct
|
|
23
|
+
4. **Severity** — pipeline break or untraced prod feature = P0; cost/eval gap = P1
|
|
24
|
+
|
|
25
|
+
## Worked example
|
|
26
|
+
|
|
27
|
+
> **Observe:** after Playwright chat, `langfuse-cli api traces list` shows no new
|
|
28
|
+
> row; `app/api/chat/route.ts` calls `openai.chat` with no Langfuse wrap.
|
|
29
|
+
> **Interpret:** the live path is uninstrumented — the static map was wrong.
|
|
30
|
+
> **Classify:** missing instrumentation (pipeline break).
|
|
31
|
+
> **Severity:** P0 — prod chat is invisible.
|
|
32
|
+
> **Finding:** chat | P0 | no trace after live send | wrap the SDK call.
|
|
33
|
+
|
|
34
|
+
## Self-critique before reporting [LOW freedom — do not skip]
|
|
35
|
+
|
|
36
|
+
1. **Concrete numbers** — model + $/call or tokens, not "costs seem high"
|
|
37
|
+
2. **Live, not static** — trigger the feature; missing post-trigger trace = P0
|
|
38
|
+
3. **Severity justified** — P0 = pipeline break or untraced user-facing call
|
|
39
|
+
4. **Right owner** — jailbreak/OWASP → `audit-llm-security`; token caps → `plan-llm-cost-guardrails`
|
|
40
|
+
5. **Keys stay in env** — report host + presence, never secret values
|
|
41
|
+
|
|
42
|
+
---
|
|
43
|
+
|
|
44
|
+
## Phase 0: Auto-Detect Langfuse Integration
|
|
45
|
+
|
|
46
|
+
### 0a. Find Langfuse Configuration
|
|
47
|
+
|
|
48
|
+
Search for environment variables and config files (in order):
|
|
49
|
+
|
|
50
|
+
1. `.env`, `.env.local`, `.env.production` — look for `LANGFUSE_PUBLIC_KEY`, `LANGFUSE_SECRET_KEY`, `LANGFUSE_BASE_URL`, `LANGFUSE_HOST`
|
|
51
|
+
2. `langfuse.config.ts`, `langfuse.config.js` — dedicated config files
|
|
52
|
+
3. `instrumentation.ts` / `instrumentation.js` — Next.js instrumentation with Langfuse
|
|
53
|
+
4. Supabase Edge Functions — `Glob("**/supabase/functions/**/index.ts")` and search for `Langfuse` imports
|
|
54
|
+
|
|
55
|
+
```
|
|
56
|
+
Grep(pattern: "LANGFUSE_PUBLIC_KEY|LANGFUSE_SECRET_KEY|LANGFUSE_BASE_URL|LANGFUSE_HOST", glob: ".env*")
|
|
57
|
+
Grep(pattern: "langfuse|Langfuse|@langfuse", glob: "*.{ts,js,tsx,jsx,py,rb,go}")
|
|
58
|
+
```
|
|
59
|
+
|
|
60
|
+
Record: `LANGFUSE_HOST`, public-key present (never the secret), import sites, CLI env ready.
|
|
61
|
+
|
|
62
|
+
### 0b. Detect LLM Framework and Provider
|
|
63
|
+
|
|
64
|
+
```
|
|
65
|
+
Grep(pattern: "openai|OpenAI|anthropic|Anthropic|@google/generative-ai|gemini|cohere|mistral|groq|together|replicate", glob: "*.{ts,js,py}")
|
|
66
|
+
Grep(pattern: "langchain|LangChain|@langchain|vercel/ai|ai/core|createOpenAI|createAnthropic", glob: "*.{ts,js,py}")
|
|
67
|
+
```
|
|
68
|
+
|
|
69
|
+
Record: providers, frameworks, and the exact model ID strings in use, with each one's tier (frontier, mid, or small).
|
|
70
|
+
|
|
71
|
+
### 0c. Map AI Features
|
|
72
|
+
|
|
73
|
+
```
|
|
74
|
+
SemanticSearch(query: "Where are LLM/AI features called in the codebase?", target_directories: [])
|
|
75
|
+
```
|
|
76
|
+
|
|
77
|
+
Build a feature map:
|
|
78
|
+
| Feature | File(s) | Provider | Model | Traced? |
|
|
79
|
+
|---------|---------|----------|-------|---------|
|
|
80
|
+
| _e.g. Chat_ | `app/api/chat/route.ts` | OpenAI | `<model id>` (frontier) | Yes |
|
|
81
|
+
|
|
82
|
+
### 0d. Detect Eval and Prompt Management Setup
|
|
83
|
+
|
|
84
|
+
```
|
|
85
|
+
Grep(pattern: "createScore|langfuse.score|annotation|eval|judge|dataset", glob: "*.{ts,js,py}")
|
|
86
|
+
Grep(pattern: "getPrompt|langfuse.prompt|fetchPrompt|compilePrompt", glob: "*.{ts,js,py}")
|
|
87
|
+
```
|
|
88
|
+
|
|
89
|
+
Record: prompt source (Langfuse vs hardcoded vs config), eval setup, version/label.
|
|
90
|
+
|
|
91
|
+
---
|
|
92
|
+
|
|
93
|
+
## Phase 1: Research LLM Best Practices
|
|
94
|
+
|
|
95
|
+
### 1a. Firecrawl Research
|
|
96
|
+
|
|
97
|
+
```json
|
|
98
|
+
firecrawl:firecrawl_search
|
|
99
|
+
{
|
|
100
|
+
"query": "LLM observability best practices production monitoring [current year]",
|
|
101
|
+
"limit": 5
|
|
102
|
+
}
|
|
103
|
+
```
|
|
104
|
+
|
|
105
|
+
```json
|
|
106
|
+
firecrawl:firecrawl_search
|
|
107
|
+
{
|
|
108
|
+
"query": "prompt engineering evaluation scoring hallucination detection [current year]",
|
|
109
|
+
"limit": 5
|
|
110
|
+
}
|
|
111
|
+
```
|
|
112
|
+
|
|
113
|
+
```json
|
|
114
|
+
firecrawl:firecrawl_search
|
|
115
|
+
{
|
|
116
|
+
"query": "LLM cost optimization token usage model selection production [current year]",
|
|
117
|
+
"limit": 5
|
|
118
|
+
}
|
|
119
|
+
```
|
|
120
|
+
|
|
121
|
+
Scrape the top 2–3 results:
|
|
122
|
+
|
|
123
|
+
```json
|
|
124
|
+
firecrawl:firecrawl_scrape
|
|
125
|
+
{
|
|
126
|
+
"url": "<BEST_RESULT_URL>",
|
|
127
|
+
"formats": ["markdown"]
|
|
128
|
+
}
|
|
129
|
+
```
|
|
130
|
+
|
|
131
|
+
### 1b. Langfuse Documentation
|
|
132
|
+
|
|
133
|
+
Langfuse docs for the detected setup:
|
|
134
|
+
|
|
135
|
+
```json
|
|
136
|
+
firecrawl:firecrawl_search
|
|
137
|
+
{
|
|
138
|
+
"query": "site:langfuse.com docs tracing prompts evaluation scores",
|
|
139
|
+
"limit": 5
|
|
140
|
+
}
|
|
141
|
+
```
|
|
142
|
+
|
|
143
|
+
### 1c. Context7 for LLM Framework Docs
|
|
144
|
+
|
|
145
|
+
|
|
146
|
+
```json
|
|
147
|
+
context7:resolve-library-id
|
|
148
|
+
{
|
|
149
|
+
"libraryName": "<DETECTED_FRAMEWORK e.g. langchain or vercel-ai>"
|
|
150
|
+
}
|
|
151
|
+
```
|
|
152
|
+
|
|
153
|
+
```json
|
|
154
|
+
context7:query-docs
|
|
155
|
+
{
|
|
156
|
+
"libraryId": "<RESOLVED_ID>",
|
|
157
|
+
"query": "Langfuse integration tracing observability"
|
|
158
|
+
}
|
|
159
|
+
```
|
|
160
|
+
|
|
161
|
+
---
|
|
162
|
+
|
|
163
|
+
## Phase 2: Audit via Langfuse CLI
|
|
164
|
+
|
|
165
|
+
Shell + Langfuse CLI. Require `LANGFUSE_PUBLIC_KEY` and `LANGFUSE_SECRET_KEY` in the env.
|
|
166
|
+
|
|
167
|
+
### 2a. Trace Completeness
|
|
168
|
+
|
|
169
|
+
```bash
|
|
170
|
+
npx langfuse-cli api traces list --limit 50
|
|
171
|
+
```
|
|
172
|
+
|
|
173
|
+
For each AI feature identified in Phase 0c, verify:
|
|
174
|
+
- [ ] Trace exists with a matching name/metadata
|
|
175
|
+
- [ ] Trace has spans/generations (not just a top-level trace with no children)
|
|
176
|
+
- [ ] Trace includes input/output (not empty)
|
|
177
|
+
- [ ] Trace has proper metadata (userId, sessionId, tags)
|
|
178
|
+
- [ ] Latency is recorded
|
|
179
|
+
|
|
180
|
+
**Red flags:**
|
|
181
|
+
- AI feature exists in code but produces no traces → **missing instrumentation**
|
|
182
|
+
- Traces exist but have no generations → **incomplete tracing** (wrapper created but LLM call not captured)
|
|
183
|
+
- Traces with empty output → **output not being captured** (fire-and-forget pattern)
|
|
184
|
+
|
|
185
|
+
### 2b. Prompt Quality Audit
|
|
186
|
+
|
|
187
|
+
```bash
|
|
188
|
+
npx langfuse-cli api prompts list
|
|
189
|
+
```
|
|
190
|
+
|
|
191
|
+
For each prompt:
|
|
192
|
+
|
|
193
|
+
```bash
|
|
194
|
+
npx langfuse-cli api prompts get --name "<PROMPT_NAME>"
|
|
195
|
+
```
|
|
196
|
+
|
|
197
|
+
Evaluate:
|
|
198
|
+
- [ ] **Versioning**: Are prompts versioned (v1, v2, v3+) or stuck at v1?
|
|
199
|
+
- [ ] **Labels**: Is there a `production` label? Are there `staging`/`experiment` labels for A/B testing?
|
|
200
|
+
- [ ] **System message quality**: Clear role definition, constraints, output format instructions?
|
|
201
|
+
- [ ] **Few-shot examples**: Does the prompt include examples for complex tasks?
|
|
202
|
+
- [ ] **Guardrails**: Does the prompt include instructions to refuse off-topic/harmful requests?
|
|
203
|
+
- [ ] **Variables**: Are dynamic parts properly templated with `{{variables}}` not string concatenation?
|
|
204
|
+
- [ ] **Freshness**: When was the prompt last updated? Stale prompts may not use newer model capabilities.
|
|
205
|
+
|
|
206
|
+
If prompts are hardcoded in source code instead of managed via Langfuse:
|
|
207
|
+
|
|
208
|
+
```
|
|
209
|
+
Grep(pattern: "You are|system.*message|systemPrompt|SYSTEM_PROMPT", glob: "*.{ts,js,py}")
|
|
210
|
+
```
|
|
211
|
+
|
|
212
|
+
Flag hardcoded prompts as a finding — they should be migrated to Langfuse for versioning and A/B testing.
|
|
213
|
+
|
|
214
|
+
### 2c. Model and Cost Efficiency
|
|
215
|
+
|
|
216
|
+
From trace data, analyze:
|
|
217
|
+
|
|
218
|
+
```bash
|
|
219
|
+
npx langfuse-cli api traces list --limit 50
|
|
220
|
+
```
|
|
221
|
+
|
|
222
|
+
Build a cost table from generation details (model, tokens, latency, cost):
|
|
223
|
+
| Feature | Model | Avg Input Tokens | Avg Output Tokens | Avg Latency | Est. Cost/Call |
|
|
224
|
+
|---------|-------|------------------|-------------------|-------------|----------------|
|
|
225
|
+
|
|
226
|
+
**Red flags:**
|
|
227
|
+
- Frontier-tier model used for simple classification/extraction → **recommend the provider's small tier**
|
|
228
|
+
- High input token counts → **check for unnecessary context stuffing**
|
|
229
|
+
- Output tokens much larger than needed → **add an output-token cap or response format constraints**
|
|
230
|
+
- High latency on user-facing features → **consider streaming, caching, or smaller model**
|
|
231
|
+
- Same content sent repeatedly → **implement semantic caching**
|
|
232
|
+
|
|
233
|
+
### 2d. Eval Score Health
|
|
234
|
+
|
|
235
|
+
```bash
|
|
236
|
+
npx langfuse-cli api scores list --limit 50
|
|
237
|
+
```
|
|
238
|
+
|
|
239
|
+
Evaluate:
|
|
240
|
+
- [ ] **Score existence**: Are evals running at all?
|
|
241
|
+
- [ ] **Score types**: What's being measured (relevance, faithfulness, toxicity, custom)?
|
|
242
|
+
- [ ] **Score distribution**: Are scores clustered (all 1.0 = useless eval) or distributed?
|
|
243
|
+
- [ ] **Annotation queues**: Are humans reviewing AI outputs?
|
|
244
|
+
- [ ] **Judge LLM**: If using LLM-as-judge, which model? Is the judge prompt well-designed?
|
|
245
|
+
|
|
246
|
+
**Red flags:**
|
|
247
|
+
- No scores at all → **no quality feedback loop**
|
|
248
|
+
- Only manual scores, no automated → **quality is not continuously monitored**
|
|
249
|
+
- All scores are identical → **eval criteria too loose or rubric too vague**
|
|
250
|
+
- Scores declining over time → **model degradation or prompt drift**
|
|
251
|
+
|
|
252
|
+
### 2e. Session and User Attribution
|
|
253
|
+
|
|
254
|
+
```bash
|
|
255
|
+
npx langfuse-cli api sessions list --limit 20
|
|
256
|
+
```
|
|
257
|
+
|
|
258
|
+
Verify:
|
|
259
|
+
- [ ] Sessions group related interactions (multi-turn conversations have one session ID)
|
|
260
|
+
- [ ] User IDs are attributed (not all anonymous)
|
|
261
|
+
- [ ] Session metadata is useful (page, feature, user segment)
|
|
262
|
+
|
|
263
|
+
### 2f. Dataset Health
|
|
264
|
+
|
|
265
|
+
```bash
|
|
266
|
+
npx langfuse-cli api datasets list
|
|
267
|
+
```
|
|
268
|
+
|
|
269
|
+
Evaluate:
|
|
270
|
+
- [ ] **Datasets exist**: Are there regression test datasets?
|
|
271
|
+
- [ ] **Dataset freshness**: When were items last added?
|
|
272
|
+
- [ ] **Coverage**: Do datasets cover all AI features or just one?
|
|
273
|
+
- [ ] **Expected outputs**: Do dataset items have expected outputs for automated comparison?
|
|
274
|
+
|
|
275
|
+
---
|
|
276
|
+
|
|
277
|
+
## Phase 3: Live Verification
|
|
278
|
+
|
|
279
|
+
### 3a. Trigger AI Features via Playwright
|
|
280
|
+
|
|
281
|
+
For each Phase 0c feature, trigger it live. Timeouts 15s; `sleep 2` → `snapshot` (not one long block).
|
|
282
|
+
|
|
283
|
+
```bash
|
|
284
|
+
PW="npx --yes @playwright/cli@latest"
|
|
285
|
+
$PW -s=langfuse-audit open --headed "<APP_URL>"
|
|
286
|
+
```
|
|
287
|
+
|
|
288
|
+
Navigate to the feature, interact with it (fill form, click button, send message), and capture:
|
|
289
|
+
- The AI-generated response (via `snapshot`)
|
|
290
|
+
- Console messages (via `console`) — look for errors
|
|
291
|
+
- Network requests (via `requests`) — look for failed API calls
|
|
292
|
+
|
|
293
|
+
### 3b. Verify Trace Pipeline
|
|
294
|
+
|
|
295
|
+
After triggering each feature, wait 5-10 seconds, then verify the trace landed:
|
|
296
|
+
|
|
297
|
+
```bash
|
|
298
|
+
npx langfuse-cli api traces list --limit 5
|
|
299
|
+
```
|
|
300
|
+
|
|
301
|
+
Check:
|
|
302
|
+
- [ ] New trace appeared with correct name
|
|
303
|
+
- [ ] Trace has generations with model and token data
|
|
304
|
+
- [ ] Trace latency matches observed UX latency
|
|
305
|
+
- [ ] Input/output captured correctly
|
|
306
|
+
|
|
307
|
+
### 3c. Cross-Check with Sentry
|
|
308
|
+
|
|
309
|
+
```json
|
|
310
|
+
sentry:search_issues
|
|
311
|
+
{
|
|
312
|
+
"organizationSlug": "<ORG_SLUG>",
|
|
313
|
+
"projectSlugOrId": "<PROJECT_SLUG>",
|
|
314
|
+
"query": "is:unresolved ai OR llm OR openai OR anthropic OR langfuse OR completion OR embedding"
|
|
315
|
+
}
|
|
316
|
+
```
|
|
317
|
+
|
|
318
|
+
Check for:
|
|
319
|
+
- LLM timeout errors
|
|
320
|
+
- Rate limiting (429) errors
|
|
321
|
+
- Token limit exceeded errors
|
|
322
|
+
- Langfuse SDK errors (failed to send trace)
|
|
323
|
+
- JSON parse errors on LLM responses
|
|
324
|
+
|
|
325
|
+
### 3d. Cross-Check with Supabase (if AI results stored in DB)
|
|
326
|
+
|
|
327
|
+
If the project stores AI outputs in the database:
|
|
328
|
+
|
|
329
|
+
```json
|
|
330
|
+
supabase:list_tables
|
|
331
|
+
{
|
|
332
|
+
"schemas": ["public"]
|
|
333
|
+
}
|
|
334
|
+
```
|
|
335
|
+
|
|
336
|
+
Find tables that store AI outputs and verify data landed:
|
|
337
|
+
|
|
338
|
+
```json
|
|
339
|
+
supabase:execute_sql
|
|
340
|
+
{
|
|
341
|
+
"query": "SELECT id, created_at, <ai_output_column> FROM <table> ORDER BY created_at DESC LIMIT 5"
|
|
342
|
+
}
|
|
343
|
+
```
|
|
344
|
+
|
|
345
|
+
### 3e. Grounding / Hallucination Check
|
|
346
|
+
|
|
347
|
+
For features where the AI should reference source data (RAG, summarization, data extraction):
|
|
348
|
+
|
|
349
|
+
1. Get the source data from the database (Supabase `execute_sql`)
|
|
350
|
+
2. Trigger the AI feature via Playwright
|
|
351
|
+
3. Compare the AI output against the source data
|
|
352
|
+
|
|
353
|
+
**Red flags:**
|
|
354
|
+
- AI mentions facts not in the source data → **hallucination**
|
|
355
|
+
- AI omits critical facts from the source data → **incomplete extraction**
|
|
356
|
+
- AI contradicts the source data → **grounding failure**
|
|
357
|
+
- AI generates plausible but wrong numbers → **numerical hallucination**
|
|
358
|
+
|
|
359
|
+
---
|
|
360
|
+
|
|
361
|
+
## Phase 4: Report
|
|
362
|
+
|
|
363
|
+
Generate a structured report with the following sections.
|
|
364
|
+
|
|
365
|
+
```
|
|
366
|
+
═══════════════════════════════════════════════════════
|
|
367
|
+
LANGFUSE LLM QUALITY AUDIT REPORT
|
|
368
|
+
Project: <PROJECT_NAME>
|
|
369
|
+
Date: <DATE>
|
|
370
|
+
Langfuse Host: <HOST_URL>
|
|
371
|
+
═══════════════════════════════════════════════════════
|
|
372
|
+
|
|
373
|
+
## 1. TRACE COVERAGE
|
|
374
|
+
|
|
375
|
+
| Feature | Traced? | Generations? | Input/Output? | Metadata? | Status |
|
|
376
|
+
|---------|---------|-------------|---------------|-----------|--------|
|
|
377
|
+
| ... | ... | ... | ... | ... | ✅/❌ |
|
|
378
|
+
|
|
379
|
+
Coverage: X/Y features traced (Z%)
|
|
380
|
+
Missing instrumentation: [list features with no traces]
|
|
381
|
+
|
|
382
|
+
## 2. PROMPT QUALITY
|
|
383
|
+
|
|
384
|
+
| Prompt | Version | Label | System Msg | Few-Shot | Guardrails | Variables | Score |
|
|
385
|
+
|--------|---------|-------|------------|----------|------------|-----------|-------|
|
|
386
|
+
| ... | ... | ... | ... | ... | ... | ... | A-F |
|
|
387
|
+
|
|
388
|
+
Hardcoded prompts found: [list files with inline prompts]
|
|
389
|
+
Recommendations: [specific improvements per prompt]
|
|
390
|
+
|
|
391
|
+
## 3. COST EFFICIENCY
|
|
392
|
+
|
|
393
|
+
| Feature | Model | Avg Tokens (in/out) | Avg Latency | Est. Cost/Call | Recommendation |
|
|
394
|
+
|---------|-------|---------------------|-------------|----------------|----------------|
|
|
395
|
+
| ... | ... | ... | ... | ... | ... |
|
|
396
|
+
|
|
397
|
+
Monthly estimate: $X (at current usage rate)
|
|
398
|
+
Savings opportunity: $Y (by implementing recommendations)
|
|
399
|
+
|
|
400
|
+
## 4. EVAL HEALTH
|
|
401
|
+
|
|
402
|
+
| Metric | Status | Details |
|
|
403
|
+
|------------------|-----------|----------------------------------|
|
|
404
|
+
| Automated evals | ✅/❌ | [count and types] |
|
|
405
|
+
| Manual reviews | ✅/❌ | [annotation queue status] |
|
|
406
|
+
| Score distribution| ✅/❌ | [healthy spread vs clustered] |
|
|
407
|
+
| Datasets | ✅/❌ | [count, freshness, coverage] |
|
|
408
|
+
| Regression tests | ✅/❌ | [dataset run frequency] |
|
|
409
|
+
|
|
410
|
+
## 5. PIPELINE INTEGRITY
|
|
411
|
+
|
|
412
|
+
| Step | Status | Evidence |
|
|
413
|
+
|-------------------------|--------|-------------------------------------|
|
|
414
|
+
| FE triggers AI feature | ✅/❌ | [Playwright observation] |
|
|
415
|
+
| API receives request | ✅/❌ | [network request captured] |
|
|
416
|
+
| LLM call executes | ✅/❌ | [trace generation exists] |
|
|
417
|
+
| Trace lands in Langfuse | ✅/❌ | [CLI verification] |
|
|
418
|
+
| Result stored in DB | ✅/❌ | [Supabase query result] |
|
|
419
|
+
| Result displayed in FE | ✅/❌ | [Playwright snapshot] |
|
|
420
|
+
| Eval score recorded | ✅/❌ | [score attached to trace] |
|
|
421
|
+
|
|
422
|
+
## 6. GROUNDING & HALLUCINATION
|
|
423
|
+
|
|
424
|
+
| Feature | Source Data | AI Output Match | Hallucinations | Score |
|
|
425
|
+
|---------|-------------|-----------------|----------------|-------|
|
|
426
|
+
| ... | ... | ... | ... | A-F |
|
|
427
|
+
|
|
428
|
+
## 7. SENTRY LLM ERRORS
|
|
429
|
+
|
|
430
|
+
| Issue | Error Type | Events | Impact | Fix Needed |
|
|
431
|
+
|-------|------------|--------|--------|------------|
|
|
432
|
+
| ... | ... | ... | ... | ... |
|
|
433
|
+
|
|
434
|
+
## 8. CRITICAL FINDINGS (Action Required)
|
|
435
|
+
|
|
436
|
+
P0 — Must fix immediately:
|
|
437
|
+
1. [finding with evidence]
|
|
438
|
+
|
|
439
|
+
P1 — Should fix this sprint:
|
|
440
|
+
1. [finding with evidence]
|
|
441
|
+
|
|
442
|
+
P2 — Improvement opportunity:
|
|
443
|
+
1. [finding with evidence]
|
|
444
|
+
|
|
445
|
+
## 9. RECOMMENDATIONS
|
|
446
|
+
|
|
447
|
+
| # | Category | Current State | Recommended State | Effort | Impact |
|
|
448
|
+
|---|----------|---------------|-------------------|--------|--------|
|
|
449
|
+
| 1 | ... | ... | ... | S/M/L | S/M/L |
|
|
450
|
+
|
|
451
|
+
## 10. PDCA IMPROVEMENT RESULTS
|
|
452
|
+
|
|
453
|
+
| Prompt | Baseline Score | Iter 1 Score | Iter 2 Score | Iter 3 Score | Final Score | Action Taken |
|
|
454
|
+
|--------|---------------|-------------|-------------|-------------|-------------|--------------|
|
|
455
|
+
| ... | ... | ... | ... | ... | ... | Promoted / Rolled back / Needs manual |
|
|
456
|
+
```
|
|
457
|
+
|
|
458
|
+
## Further reading
|
|
459
|
+
|
|
460
|
+
- [Improvement Details and more](references/details.md)
|
|
461
|
+
|
|
462
|
+
## Related
|
|
463
|
+
|
|
464
|
+
- `audit-llm-security` — OWASP LLM Top 10 (injection, agency, leaks) — not quality/evals
|
|
465
|
+
- `plan-llm-cost-guardrails` — token budgets and quota abuse
|
|
466
|
+
- `plan-privacy-compliance` — PII in traces
|
|
467
|
+
- `backend-observability` / `debug-sentry-monitor` — non-LLM telemetry
|
|
468
|
+
|
|
@@ -0,0 +1,226 @@
|
|
|
1
|
+
```
|
|
2
|
+
### Improvement Details
|
|
3
|
+
|
|
4
|
+
**[Prompt Name] — Iteration 1:**
|
|
5
|
+
- Change: [what was changed and why]
|
|
6
|
+
- Result: Score [X] -> [Y] ([+/-Z%]), Latency [A]ms -> [B]ms, Cost $[C] -> $[D]
|
|
7
|
+
- Evidence: Langfuse trace ID [ID], Playwright screenshot [ref]
|
|
8
|
+
|
|
9
|
+
**[Prompt Name] — Iteration 2:**
|
|
10
|
+
- Change: [next improvement attempt]
|
|
11
|
+
- Result: ...
|
|
12
|
+
|
|
13
|
+
═══════════════════════════════════════════════════════
|
|
14
|
+
```
|
|
15
|
+
|
|
16
|
+
---
|
|
17
|
+
|
|
18
|
+
## Phase 5: Prompt Improvement Cycle (PDCA)
|
|
19
|
+
|
|
20
|
+
After the audit report identifies prompt weaknesses, run an iterative improvement cycle for each
|
|
21
|
+
flagged prompt. This is the **Do-Check-Act** loop that turns audit findings into measurable improvements.
|
|
22
|
+
|
|
23
|
+
> **Max 3 iterations per prompt.** If no improvement after 3 attempts, log as "needs manual prompt engineering" and move on.
|
|
24
|
+
|
|
25
|
+
> **Always capture baseline BEFORE making changes.** Without a baseline, you cannot measure improvement.
|
|
26
|
+
|
|
27
|
+
> **Never promote a prompt version that scores worse than baseline.** The goal is monotonic improvement.
|
|
28
|
+
|
|
29
|
+
### 5a. Capture Baseline
|
|
30
|
+
|
|
31
|
+
For each prompt flagged in Phase 2b (Prompt Quality Audit), record the current state:
|
|
32
|
+
|
|
33
|
+
```bash
|
|
34
|
+
npx langfuse-cli api prompts get --name "<PROMPT_NAME>"
|
|
35
|
+
```
|
|
36
|
+
|
|
37
|
+
Record:
|
|
38
|
+
- **Current version number** and **label** (e.g., v3, label: `production`)
|
|
39
|
+
- **Current eval scores** from Phase 2d (average score, score distribution)
|
|
40
|
+
- **Current cost metrics** from Phase 2c (model, avg tokens, avg latency, est. cost/call)
|
|
41
|
+
- **Specific weaknesses** identified in the audit (missing guardrails, no few-shot, vague instructions, etc.)
|
|
42
|
+
|
|
43
|
+
Build a baseline table:
|
|
44
|
+
|
|
45
|
+
| Prompt | Version | Avg Score | Avg Latency | Avg Cost | Weaknesses |
|
|
46
|
+
|--------|---------|-----------|-------------|----------|------------|
|
|
47
|
+
| ... | ... | ... | ... | ... | ... |
|
|
48
|
+
|
|
49
|
+
### 5b. DO: Improve the Prompt
|
|
50
|
+
|
|
51
|
+
For each flagged prompt, create an improved version based on the audit findings.
|
|
52
|
+
|
|
53
|
+
**Step 1: Research the specific improvement needed**
|
|
54
|
+
|
|
55
|
+
```json
|
|
56
|
+
firecrawl:firecrawl_search
|
|
57
|
+
{
|
|
58
|
+
"query": "<WEAKNESS_TYPE> prompt engineering best practices [current year]",
|
|
59
|
+
"limit": 5
|
|
60
|
+
}
|
|
61
|
+
```
|
|
62
|
+
|
|
63
|
+
Example queries by weakness type:
|
|
64
|
+
- Missing guardrails: `"LLM safety guardrails system prompt best practices [current year]"`
|
|
65
|
+
- No few-shot: `"few-shot prompting examples for <TASK_TYPE> [current year]"`
|
|
66
|
+
- Vague instructions: `"prompt engineering specific instructions examples structured output [current year]"`
|
|
67
|
+
- Wrong model: `"<TASK_TYPE> model selection small tier vs frontier tier <PROVIDER> [current year]"`
|
|
68
|
+
|
|
69
|
+
Scrape the top 1-2 results for concrete patterns:
|
|
70
|
+
|
|
71
|
+
```json
|
|
72
|
+
firecrawl:firecrawl_scrape
|
|
73
|
+
{
|
|
74
|
+
"url": "<BEST_RESULT_URL>",
|
|
75
|
+
"formats": ["markdown"],
|
|
76
|
+
"onlyMainContent": true
|
|
77
|
+
}
|
|
78
|
+
```
|
|
79
|
+
|
|
80
|
+
**Step 2: Draft the improved prompt**
|
|
81
|
+
|
|
82
|
+
Apply the researched improvements. Common enhancement patterns:
|
|
83
|
+
- **Add role definition**: "You are a [specific role] with expertise in [domain]."
|
|
84
|
+
- **Add output format**: "Respond in JSON with the following schema: {...}"
|
|
85
|
+
- **Add guardrails**: "If the user asks about [off-topic], respond with [refusal]."
|
|
86
|
+
- **Add few-shot examples**: Include 2-3 input/output pairs for complex tasks
|
|
87
|
+
- **Use native reasoning**: turn on the provider's reasoning mode (Anthropic adaptive thinking plus `effort`, OpenAI reasoning models) instead of adding "Think step by step" to the prompt — reasoning models already do it, and on the rest the phrase mostly lengthens output.
|
|
88
|
+
- **Shape the output**: name the audience and the format; cap length with the request's output-token limit (`max_tokens` on Anthropic, `max_completion_tokens` or `max_output_tokens` on OpenAI) rather than a word count in the prompt, which starves reasoning on hard inputs.
|
|
89
|
+
- **Add grounding instructions**: "Only use information from the provided context. If unsure, say so."
|
|
90
|
+
|
|
91
|
+
**Step 3: Create the new version**
|
|
92
|
+
|
|
93
|
+
If prompts are managed in Langfuse:
|
|
94
|
+
|
|
95
|
+
```bash
|
|
96
|
+
npx langfuse-cli api prompts create --name "<PROMPT_NAME>" --prompt "<IMPROVED_PROMPT_TEXT>" --labels experiment
|
|
97
|
+
```
|
|
98
|
+
|
|
99
|
+
This creates a new version with the `experiment` label (not `production` — we test first).
|
|
100
|
+
|
|
101
|
+
If prompts are hardcoded in source code:
|
|
102
|
+
1. Read the file containing the prompt
|
|
103
|
+
2. Edit the prompt text in place (StrReplace in Cursor, Edit in Claude Code)
|
|
104
|
+
3. Record the file path and the change for potential rollback
|
|
105
|
+
|
|
106
|
+
### 5c. CHECK: Trigger and Verify via Playwright
|
|
107
|
+
|
|
108
|
+
**Important**: Apply the `protocol-browser-anti-stall` protocol for all browser interactions.
|
|
109
|
+
|
|
110
|
+
**Step 1: Trigger the AI feature**
|
|
111
|
+
|
|
112
|
+
```bash
|
|
113
|
+
PW="npx --yes @playwright/cli@latest"
|
|
114
|
+
$PW -s=langfuse-audit open --headed "<APP_URL>"
|
|
115
|
+
```
|
|
116
|
+
|
|
117
|
+
Navigate to the feature that uses this prompt. Interact with it (fill form, click button, send message).
|
|
118
|
+
|
|
119
|
+
Capture:
|
|
120
|
+
- AI-generated response via `snapshot`
|
|
121
|
+
- Console errors via `console`
|
|
122
|
+
- Network requests via `requests`
|
|
123
|
+
- Screenshot via `screenshot` (visual evidence for the report)
|
|
124
|
+
|
|
125
|
+
**Step 2: Repeat N=3 times minimum**
|
|
126
|
+
|
|
127
|
+
Trigger the feature at least 3 times with different inputs to get a representative sample.
|
|
128
|
+
Vary the inputs to test edge cases relevant to the improvement (e.g., if you added guardrails,
|
|
129
|
+
test with an off-topic input).
|
|
130
|
+
|
|
131
|
+
**Step 3: Verify traces landed in Langfuse**
|
|
132
|
+
|
|
133
|
+
Wait 5-10 seconds after each trigger, then:
|
|
134
|
+
|
|
135
|
+
```bash
|
|
136
|
+
npx langfuse-cli api traces list --limit 10
|
|
137
|
+
```
|
|
138
|
+
|
|
139
|
+
For each new trace, verify:
|
|
140
|
+
- [ ] Trace uses the new prompt version (check metadata or generation details)
|
|
141
|
+
- [ ] Input/output captured correctly
|
|
142
|
+
- [ ] Generation includes model, token usage, latency
|
|
143
|
+
|
|
144
|
+
### 5d. CHECK: Score Comparison
|
|
145
|
+
|
|
146
|
+
**Option A: If Langfuse datasets exist** (preferred — most rigorous)
|
|
147
|
+
|
|
148
|
+
Run the dataset experiment with the new prompt version vs the old:
|
|
149
|
+
|
|
150
|
+
```bash
|
|
151
|
+
npx langfuse-cli api datasets list
|
|
152
|
+
```
|
|
153
|
+
|
|
154
|
+
If a relevant dataset exists, use Langfuse's Experiment feature (via UI or SDK) to run both
|
|
155
|
+
prompt versions against the same dataset and compare scores side-by-side.
|
|
156
|
+
|
|
157
|
+
**Option B: If LLM-as-a-Judge evaluators are configured**
|
|
158
|
+
|
|
159
|
+
Check if automated eval scores have been generated for the new traces:
|
|
160
|
+
|
|
161
|
+
```bash
|
|
162
|
+
npx langfuse-cli api scores list --limit 20
|
|
163
|
+
```
|
|
164
|
+
|
|
165
|
+
Compare the scores attached to new traces (experiment version) vs old traces (production version).
|
|
166
|
+
|
|
167
|
+
**Option C: Live trace comparison (minimum viable)**
|
|
168
|
+
|
|
169
|
+
From the N=3+ triggers in Step 5c, compare the new traces against baseline:
|
|
170
|
+
|
|
171
|
+
| Metric | Baseline (avg) | New Version (avg) | Delta | Verdict |
|
|
172
|
+
|--------|---------------|-------------------|-------|---------|
|
|
173
|
+
| Eval score | ... | ... | +/-% | Better / Worse / Same |
|
|
174
|
+
| Latency | ...ms | ...ms | +/-% | Better / Worse / Same |
|
|
175
|
+
| Input tokens | ... | ... | +/-% | Better / Worse / Same |
|
|
176
|
+
| Output tokens | ... | ... | +/-% | Better / Worse / Same |
|
|
177
|
+
| Est. cost/call | $... | $... | +/-% | Better / Worse / Same |
|
|
178
|
+
| Grounding accuracy | ... | ... | +/-% | Better / Worse / Same |
|
|
179
|
+
|
|
180
|
+
### 5e. ACT: Promote or Rollback
|
|
181
|
+
|
|
182
|
+
Based on the score comparison:
|
|
183
|
+
|
|
184
|
+
**Score improved (any eval metric improved, no metric degraded):**
|
|
185
|
+
- If Langfuse-managed: promote the new version to `production` label
|
|
186
|
+
- If hardcoded: keep the source code change, commit it
|
|
187
|
+
- Log the improvement in the report
|
|
188
|
+
|
|
189
|
+
**Score unchanged (no significant difference):**
|
|
190
|
+
- Keep the new version as `staging` for further observation
|
|
191
|
+
- Document why the improvement didn't move the needle
|
|
192
|
+
- Try a different improvement strategy in the next iteration
|
|
193
|
+
|
|
194
|
+
**Score degraded (any eval metric worsened):**
|
|
195
|
+
- If Langfuse-managed: do NOT change labels — the `production` label stays on the old version
|
|
196
|
+
- If hardcoded: revert the source code change with the same edit tool, using the path and text recorded in 5b step 3
|
|
197
|
+
- Document what went wrong and why
|
|
198
|
+
|
|
199
|
+
**Repeat the cycle** from Step 5b with a different strategy if the score target has not been met
|
|
200
|
+
and iterations remain (max 3).
|
|
201
|
+
|
|
202
|
+
### 5f. Iteration Log
|
|
203
|
+
|
|
204
|
+
Track all iterations in a structured log:
|
|
205
|
+
|
|
206
|
+
```
|
|
207
|
+
PDCA ITERATION LOG — [Prompt Name]
|
|
208
|
+
|
|
209
|
+
Baseline: v[N], score [X], latency [Y]ms, cost $[Z]/call
|
|
210
|
+
|
|
211
|
+
--- Iteration 1 ---
|
|
212
|
+
Strategy: [what was changed and why]
|
|
213
|
+
Changes: [specific edits to the prompt]
|
|
214
|
+
Result: score [X1], latency [Y1]ms, cost $[Z1]/call
|
|
215
|
+
Delta: score [+/-A%], latency [+/-B%], cost [+/-C%]
|
|
216
|
+
Verdict: [PROMOTE / CONTINUE / ROLLBACK]
|
|
217
|
+
|
|
218
|
+
--- Iteration 2 ---
|
|
219
|
+
Strategy: [different approach based on Iteration 1 results]
|
|
220
|
+
...
|
|
221
|
+
|
|
222
|
+
--- Iteration 3 (final) ---
|
|
223
|
+
...
|
|
224
|
+
|
|
225
|
+
FINAL OUTCOME: [Promoted vN+K to production / No improvement — needs manual engineering]
|
|
226
|
+
```
|