opencode-skills-collection 4.0.68 → 4.0.70
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/bundled-skills/.antigravity-install-manifest.json +304 -1
- package/bundled-skills/access-review/SKILL.md +394 -0
- package/bundled-skills/access-review/references/details.md +121 -0
- package/bundled-skills/agent-evals/SKILL.md +420 -0
- package/bundled-skills/agent-observability/SKILL.md +346 -0
- package/bundled-skills/agent-observability/references/details.md +786 -0
- package/bundled-skills/ai-agent-security/SKILL.md +393 -0
- package/bundled-skills/ai-agent-security/references/details.md +912 -0
- package/bundled-skills/ai-coding-agent-guardrails/SKILL.md +442 -0
- package/bundled-skills/ai-coding-agent-guardrails/references/details.md +753 -0
- package/bundled-skills/ai-inference-service-mesh/SKILL.md +449 -0
- package/bundled-skills/ai-pipeline-orchestration/SKILL.md +287 -0
- package/bundled-skills/ai-red-teaming/SKILL.md +409 -0
- package/bundled-skills/ai-security-hardening/SKILL.md +343 -0
- package/bundled-skills/ai-sre-incident-response/SKILL.md +336 -0
- package/bundled-skills/alerting-oncall/SKILL.md +458 -0
- package/bundled-skills/alerting-oncall/references/details.md +84 -0
- package/bundled-skills/api-integration-architect/SKILL.md +241 -0
- package/bundled-skills/apify-generate-output-schema/SKILL.md +438 -0
- package/bundled-skills/apify-integration-development/SKILL.md +168 -0
- package/bundled-skills/apify-integration-development/references/ai-framework-package.md +158 -0
- package/bundled-skills/apify-integration-development/references/ai-harness-plugin.md +192 -0
- package/bundled-skills/apify-integration-development/references/sdk-integration.md +236 -0
- package/bundled-skills/apify-integration-development/references/workflow-automation.md +163 -0
- package/bundled-skills/apk-redteam-pipeline/SKILL.md +446 -0
- package/bundled-skills/architecture-review/README.md +42 -0
- package/bundled-skills/architecture-review/SKILL.md +77 -0
- package/bundled-skills/architecture-review/examples.md +11 -0
- package/bundled-skills/architecture-review/reference/best-practices.md +7 -0
- package/bundled-skills/architecture-review/reference/capabilities.md +20 -0
- package/bundled-skills/architecture-review/reference/fallbacks.md +11 -0
- package/bundled-skills/architecture-review/reference/graph.md +15 -0
- package/bundled-skills/architecture-review/reference/mcp.md +14 -0
- package/bundled-skills/architecture-review/reference/workflow.md +15 -0
- package/bundled-skills/architecture-review/templates/architecture-review.md +21 -0
- package/bundled-skills/argocd-gitops/SKILL.md +469 -0
- package/bundled-skills/arm-templates/SKILL.md +438 -0
- package/bundled-skills/arm-templates/references/details.md +64 -0
- package/bundled-skills/asset-inventory/SKILL.md +412 -0
- package/bundled-skills/asset-inventory/references/details.md +127 -0
- package/bundled-skills/audit-logging/SKILL.md +476 -0
- package/bundled-skills/aws-cloudtrail/SKILL.md +486 -0
- package/bundled-skills/aws-cost-optimization/SKILL.md +331 -0
- package/bundled-skills/aws-ec2/SKILL.md +426 -0
- package/bundled-skills/aws-ecs-fargate/SKILL.md +388 -0
- package/bundled-skills/aws-iam/SKILL.md +463 -0
- package/bundled-skills/aws-lambda/SKILL.md +428 -0
- package/bundled-skills/aws-rds/SKILL.md +380 -0
- package/bundled-skills/aws-s3/SKILL.md +434 -0
- package/bundled-skills/aws-secrets-manager/SKILL.md +486 -0
- package/bundled-skills/aws-vpc/SKILL.md +436 -0
- package/bundled-skills/azure-ai-document-intelligence-ts/SKILL.md +1 -1
- package/bundled-skills/azure-aks/SKILL.md +423 -0
- package/bundled-skills/azure-devops/SKILL.md +457 -0
- package/bundled-skills/azure-functions-devsec/SKILL.md +436 -0
- package/bundled-skills/azure-keyvault/SKILL.md +455 -0
- package/bundled-skills/azure-keyvault/references/details.md +83 -0
- package/bundled-skills/azure-monitor-audit/SKILL.md +379 -0
- package/bundled-skills/azure-networking/SKILL.md +448 -0
- package/bundled-skills/azure-networking/references/details.md +135 -0
- package/bundled-skills/azure-sql/SKILL.md +413 -0
- package/bundled-skills/azure-sql/references/details.md +113 -0
- package/bundled-skills/azure-vms/SKILL.md +402 -0
- package/bundled-skills/azure-vms/references/details.md +134 -0
- package/bundled-skills/backup-recovery/SKILL.md +388 -0
- package/bundled-skills/bb-methodology/SKILL.md +451 -0
- package/bundled-skills/bb-methodology/references/details.md +120 -0
- package/bundled-skills/block-storage/SKILL.md +371 -0
- package/bundled-skills/blue-green-deploy/SKILL.md +453 -0
- package/bundled-skills/blue-green-deploy/references/details.md +90 -0
- package/bundled-skills/bug-bounty/SKILL.md +447 -0
- package/bundled-skills/bug-bounty/references/details.md +1316 -0
- package/bundled-skills/bugcrowd-reporting/SKILL.md +351 -0
- package/bundled-skills/business-continuity/SKILL.md +463 -0
- package/bundled-skills/career-ops/SKILL.md +186 -0
- package/bundled-skills/cdn-setup/SKILL.md +374 -0
- package/bundled-skills/change-management/SKILL.md +438 -0
- package/bundled-skills/change-management/references/details.md +105 -0
- package/bundled-skills/circleci/SKILL.md +475 -0
- package/bundled-skills/cis-benchmarks/SKILL.md +150 -0
- package/bundled-skills/cloudflare-pages/SKILL.md +318 -0
- package/bundled-skills/cloudflare-r2/SKILL.md +353 -0
- package/bundled-skills/cloudflare-workers/SKILL.md +415 -0
- package/bundled-skills/cloudflare-zero-trust/SKILL.md +361 -0
- package/bundled-skills/cloudformation/SKILL.md +461 -0
- package/bundled-skills/code-review-sensei/SKILL.md +177 -0
- package/bundled-skills/codebase-onboarding/README.md +42 -0
- package/bundled-skills/codebase-onboarding/SKILL.md +77 -0
- package/bundled-skills/codebase-onboarding/examples.md +11 -0
- package/bundled-skills/codebase-onboarding/reference/best-practices.md +7 -0
- package/bundled-skills/codebase-onboarding/reference/capabilities.md +20 -0
- package/bundled-skills/codebase-onboarding/reference/fallbacks.md +11 -0
- package/bundled-skills/codebase-onboarding/reference/graph.md +15 -0
- package/bundled-skills/codebase-onboarding/reference/mcp.md +14 -0
- package/bundled-skills/codebase-onboarding/reference/workflow.md +15 -0
- package/bundled-skills/codebase-onboarding/templates/repository-onboarding.md +21 -0
- package/bundled-skills/connection-auth-rules/SKILL.md +199 -0
- package/bundled-skills/connection-auth-rules/fetch_schema.py +320 -0
- package/bundled-skills/constraint-driven-development/SKILL.md +335 -0
- package/bundled-skills/constraint-driven-development/references/floor-guard.md +99 -0
- package/bundled-skills/container-hardening/SKILL.md +126 -0
- package/bundled-skills/container-registries/SKILL.md +435 -0
- package/bundled-skills/container-scanning/SKILL.md +416 -0
- package/bundled-skills/convex-backend/SKILL.md +338 -0
- package/bundled-skills/dast-scanning/SKILL.md +437 -0
- package/bundled-skills/database-backups/SKILL.md +425 -0
- package/bundled-skills/datadog/SKILL.md +487 -0
- package/bundled-skills/dependency-analysis/README.md +42 -0
- package/bundled-skills/dependency-analysis/SKILL.md +76 -0
- package/bundled-skills/dependency-analysis/examples.md +11 -0
- package/bundled-skills/dependency-analysis/reference/best-practices.md +7 -0
- package/bundled-skills/dependency-analysis/reference/capabilities.md +20 -0
- package/bundled-skills/dependency-analysis/reference/fallbacks.md +11 -0
- package/bundled-skills/dependency-analysis/reference/graph.md +15 -0
- package/bundled-skills/dependency-analysis/reference/mcp.md +14 -0
- package/bundled-skills/dependency-analysis/reference/workflow.md +15 -0
- package/bundled-skills/dependency-analysis/templates/dependency-review.md +21 -0
- package/bundled-skills/dependency-scanning/SKILL.md +457 -0
- package/bundled-skills/devcontainers-nix/SKILL.md +416 -0
- package/bundled-skills/devops-pipeline-builder/SKILL.md +200 -0
- package/bundled-skills/disaster-recovery/SKILL.md +374 -0
- package/bundled-skills/disaster-recovery/references/details.md +219 -0
- package/bundled-skills/dns-management/SKILL.md +375 -0
- package/bundled-skills/docker-compose/SKILL.md +482 -0
- package/bundled-skills/docker-management/SKILL.md +426 -0
- package/bundled-skills/eas-app-stores/SKILL.md +197 -0
- package/bundled-skills/eas-app-stores/agents/openai.yaml +4 -0
- package/bundled-skills/eas-app-stores/references/app-store-metadata.md +497 -0
- package/bundled-skills/eas-app-stores/references/ios-app-store.md +376 -0
- package/bundled-skills/eas-app-stores/references/native-ios.md +167 -0
- package/bundled-skills/eas-app-stores/references/play-store.md +244 -0
- package/bundled-skills/eas-app-stores/references/testflight.md +62 -0
- package/bundled-skills/eas-app-stores/references/workflows.md +120 -0
- package/bundled-skills/eas-hosting/SKILL.md +448 -0
- package/bundled-skills/eas-hosting/agents/openai.yaml +4 -0
- package/bundled-skills/eas-observe/SKILL.md +75 -0
- package/bundled-skills/eas-observe/agents/openai.yaml +4 -0
- package/bundled-skills/eas-observe/references/metrics.md +98 -0
- package/bundled-skills/eas-observe/references/queries.md +403 -0
- package/bundled-skills/eas-observe/references/setup.md +476 -0
- package/bundled-skills/eas-observe/references/third-party.md +136 -0
- package/bundled-skills/eas-simulator/SKILL.md +251 -0
- package/bundled-skills/eas-simulator/agents/openai.yaml +4 -0
- package/bundled-skills/eas-simulator/references/controllers.md +135 -0
- package/bundled-skills/eas-simulator/references/run-your-app.md +240 -0
- package/bundled-skills/eas-simulator/references/troubleshooting.md +47 -0
- package/bundled-skills/eas-workflows/SKILL.md +119 -0
- package/bundled-skills/eas-workflows/agents/openai.yaml +4 -0
- package/bundled-skills/eas-workflows/scripts/fetch.js +109 -0
- package/bundled-skills/ebpf-observability/SKILL.md +436 -0
- package/bundled-skills/ebpf-observability/references/details.md +542 -0
- package/bundled-skills/elk-stack/SKILL.md +487 -0
- package/bundled-skills/enterprise-vpn-attack/SKILL.md +395 -0
- package/bundled-skills/evidence-hygiene/SKILL.md +404 -0
- package/bundled-skills/expo-animation/LICENSE +21 -0
- package/bundled-skills/expo-animation/RECIPES.md +385 -0
- package/bundled-skills/expo-animation/SKILL.md +295 -0
- package/bundled-skills/expo-animation/agents/openai.yaml +4 -0
- package/bundled-skills/fact-check-x-unified/SKILL.md +178 -0
- package/bundled-skills/fact-check-x-unified/agents/openai.yaml +4 -0
- package/bundled-skills/fact-check-x-unified/references/acceptance-criteria.md +44 -0
- package/bundled-skills/fact-check-x-unified/references/contracts.md +39 -0
- package/bundled-skills/fact-check-x-unified/scripts/common.py +31 -0
- package/bundled-skills/fact-check-x-unified/scripts/fact_check_x.py +1832 -0
- package/bundled-skills/fact-check-x-unified/scripts/trusted_search_config.py +324 -0
- package/bundled-skills/fact-check-x-unified/tests/anchor_downgrade_test.py +90 -0
- package/bundled-skills/fact-check-x-unified/tests/multi_platform_test.py +369 -0
- package/bundled-skills/fact-check-x-unified/tests/smoke_test.py +740 -0
- package/bundled-skills/fact-check-x-unified/tests/stage_checkpoint_test.py +103 -0
- package/bundled-skills/fact-check-x-unified/tests/trusted_search_config_test.py +156 -0
- package/bundled-skills/feature-flags/SKILL.md +426 -0
- package/bundled-skills/feature-flags/references/details.md +86 -0
- package/bundled-skills/fedramp-compliance/SKILL.md +453 -0
- package/bundled-skills/firebase-app-platform/SKILL.md +381 -0
- package/bundled-skills/firewall-config/SKILL.md +479 -0
- package/bundled-skills/gcp-audit-logs/SKILL.md +452 -0
- package/bundled-skills/gcp-audit-logs/references/details.md +56 -0
- package/bundled-skills/gcp-cloud-functions/SKILL.md +284 -0
- package/bundled-skills/gcp-cloud-sql/SKILL.md +277 -0
- package/bundled-skills/gcp-compute/SKILL.md +319 -0
- package/bundled-skills/gcp-gke/SKILL.md +307 -0
- package/bundled-skills/gcp-networking/SKILL.md +293 -0
- package/bundled-skills/gcp-secret-manager/SKILL.md +421 -0
- package/bundled-skills/gcp-secret-manager/references/details.md +131 -0
- package/bundled-skills/gdpr-compliance/SKILL.md +451 -0
- package/bundled-skills/gdpr-compliance/references/details.md +145 -0
- package/bundled-skills/geo-audit/SKILL.md +368 -0
- package/bundled-skills/geo-brand-mentions/SKILL.md +68 -0
- package/bundled-skills/geo-brand-mentions/references/details.md +471 -0
- package/bundled-skills/geo-citability/SKILL.md +350 -0
- package/bundled-skills/geo-compare/SKILL.md +340 -0
- package/bundled-skills/geo-content/SKILL.md +383 -0
- package/bundled-skills/geo-crawlers/SKILL.md +408 -0
- package/bundled-skills/geo-llmstxt/SKILL.md +464 -0
- package/bundled-skills/geo-platform-optimizer/SKILL.md +314 -0
- package/bundled-skills/geo-proposal/SKILL.md +378 -0
- package/bundled-skills/geo-prospect/SKILL.md +225 -0
- package/bundled-skills/geo-report/SKILL.md +436 -0
- package/bundled-skills/geo-report-pdf/SKILL.md +157 -0
- package/bundled-skills/geo-schema/SKILL.md +408 -0
- package/bundled-skills/geo-technical/SKILL.md +78 -0
- package/bundled-skills/geo-technical/references/details.md +543 -0
- package/bundled-skills/git-workflow/SKILL.md +460 -0
- package/bundled-skills/github-actions/SKILL.md +368 -0
- package/bundled-skills/gitlab-ci/SKILL.md +340 -0
- package/bundled-skills/gpt-taste/SKILL.md +8 -1
- package/bundled-skills/gpu-kubernetes-operations/SKILL.md +468 -0
- package/bundled-skills/gpu-server-management/SKILL.md +236 -0
- package/bundled-skills/hashicorp-vault/SKILL.md +408 -0
- package/bundled-skills/helm-charts/SKILL.md +469 -0
- package/bundled-skills/hf-cli/SKILL.md +263 -0
- package/bundled-skills/hipaa-compliance/SKILL.md +451 -0
- package/bundled-skills/huggingface-community-evals/SKILL.md +228 -0
- package/bundled-skills/huggingface-community-evals/examples/.env.example +3 -0
- package/bundled-skills/huggingface-community-evals/examples/USAGE_EXAMPLES.md +101 -0
- package/bundled-skills/huggingface-community-evals/scripts/inspect_eval_uv.py +104 -0
- package/bundled-skills/huggingface-community-evals/scripts/inspect_vllm_uv.py +306 -0
- package/bundled-skills/huggingface-community-evals/scripts/lighteval_vllm_uv.py +297 -0
- package/bundled-skills/huggingface-datasets/SKILL.md +130 -0
- package/bundled-skills/hunt-aspnet/SKILL.md +321 -0
- package/bundled-skills/hunt-ato/SKILL.md +184 -0
- package/bundled-skills/hunt-auth-bypass/SKILL.md +426 -0
- package/bundled-skills/hunt-auth-bypass/references/details.md +80 -0
- package/bundled-skills/hunt-brute-force/SKILL.md +341 -0
- package/bundled-skills/hunt-business-logic/SKILL.md +281 -0
- package/bundled-skills/hunt-cache-poison/SKILL.md +382 -0
- package/bundled-skills/hunt-captcha-bypass/SKILL.md +136 -0
- package/bundled-skills/hunt-cicd/SKILL.md +311 -0
- package/bundled-skills/hunt-clickjacking/SKILL.md +110 -0
- package/bundled-skills/hunt-cors/SKILL.md +335 -0
- package/bundled-skills/hunt-dom/SKILL.md +323 -0
- package/bundled-skills/hunt-exceptional-conditions/SKILL.md +111 -0
- package/bundled-skills/hunt-file-upload/SKILL.md +202 -0
- package/bundled-skills/hunt-fintech-graphql/SKILL.md +289 -0
- package/bundled-skills/hunt-forgot-password/SKILL.md +114 -0
- package/bundled-skills/hunt-grpc/SKILL.md +317 -0
- package/bundled-skills/hunt-host-header/SKILL.md +309 -0
- package/bundled-skills/hunt-html-injection/SKILL.md +106 -0
- package/bundled-skills/hunt-http-smuggling/SKILL.md +129 -0
- package/bundled-skills/hunt-http-smuggling/references/phase2h-smuggling-cachepoison.md +177 -0
- package/bundled-skills/hunt-idor/SKILL.md +434 -0
- package/bundled-skills/hunt-jwt-crypto/SKILL.md +221 -0
- package/bundled-skills/hunt-k8s/SKILL.md +337 -0
- package/bundled-skills/hunt-laravel/SKILL.md +255 -0
- package/bundled-skills/hunt-ldap/SKILL.md +351 -0
- package/bundled-skills/hunt-lfi/SKILL.md +311 -0
- package/bundled-skills/hunt-llm-ai/SKILL.md +289 -0
- package/bundled-skills/hunt-mfa-bypass/SKILL.md +177 -0
- package/bundled-skills/hunt-misc/SKILL.md +378 -0
- package/bundled-skills/hunt-nextjs/SKILL.md +299 -0
- package/bundled-skills/hunt-nodejs/SKILL.md +263 -0
- package/bundled-skills/hunt-nosqli/SKILL.md +210 -0
- package/bundled-skills/hunt-ntlm-info/SKILL.md +314 -0
- package/bundled-skills/hunt-oauth/SKILL.md +459 -0
- package/bundled-skills/hunt-open-redirect/SKILL.md +223 -0
- package/bundled-skills/hunt-race-condition/SKILL.md +381 -0
- package/bundled-skills/hunt-race-condition/references/details.md +159 -0
- package/bundled-skills/hunt-rag-vector/SKILL.md +212 -0
- package/bundled-skills/hunt-rce/SKILL.md +444 -0
- package/bundled-skills/hunt-rce/references/details.md +110 -0
- package/bundled-skills/hunt-saml/SKILL.md +156 -0
- package/bundled-skills/hunt-session/SKILL.md +342 -0
- package/bundled-skills/hunt-shadow-api/SKILL.md +198 -0
- package/bundled-skills/hunt-source-leak/SKILL.md +345 -0
- package/bundled-skills/hunt-spa-api/SKILL.md +163 -0
- package/bundled-skills/hunt-springboot/SKILL.md +285 -0
- package/bundled-skills/hunt-sqli/SKILL.md +466 -0
- package/bundled-skills/hunt-ssrf/SKILL.md +396 -0
- package/bundled-skills/hunt-ssrf/references/details.md +179 -0
- package/bundled-skills/hunt-ssti/SKILL.md +163 -0
- package/bundled-skills/hunt-subdomain/SKILL.md +379 -0
- package/bundled-skills/hunt-tls-network/SKILL.md +399 -0
- package/bundled-skills/hunt-xxe/SKILL.md +466 -0
- package/bundled-skills/i-have-adhd/SKILL.md +170 -0
- package/bundled-skills/identity-access-management/SKILL.md +382 -0
- package/bundled-skills/identity-access-management/references/details.md +524 -0
- package/bundled-skills/incident-management/SKILL.md +484 -0
- package/bundled-skills/incident-response/SKILL.md +448 -0
- package/bundled-skills/incident-response/references/details.md +113 -0
- package/bundled-skills/interview-me/SKILL.md +248 -0
- package/bundled-skills/iso27001-compliance/SKILL.md +460 -0
- package/bundled-skills/jenkins/SKILL.md +462 -0
- package/bundled-skills/jev-social/SKILL.md +182 -0
- package/bundled-skills/jev-use/SKILL.md +158 -0
- package/bundled-skills/kubernetes-hardening/SKILL.md +154 -0
- package/bundled-skills/kubernetes-ops/SKILL.md +449 -0
- package/bundled-skills/kubernetes-ops/references/details.md +108 -0
- package/bundled-skills/kustomize/SKILL.md +478 -0
- package/bundled-skills/linux-administration/SKILL.md +367 -0
- package/bundled-skills/linux-hardening/SKILL.md +154 -0
- package/bundled-skills/llm-app-security/SKILL.md +389 -0
- package/bundled-skills/llm-app-security/references/details.md +674 -0
- package/bundled-skills/llm-caching/SKILL.md +334 -0
- package/bundled-skills/llm-cost-optimization/SKILL.md +311 -0
- package/bundled-skills/llm-fine-tuning/SKILL.md +329 -0
- package/bundled-skills/llm-gateway/SKILL.md +282 -0
- package/bundled-skills/llm-inference-scaling/SKILL.md +286 -0
- package/bundled-skills/llmops-platform-engineering/SKILL.md +472 -0
- package/bundled-skills/load-balancing/SKILL.md +403 -0
- package/bundled-skills/loki-logging/SKILL.md +479 -0
- package/bundled-skills/longbridge-derivatives/SKILL.md +117 -0
- package/bundled-skills/longbridge-derivatives/references/option.md +36 -0
- package/bundled-skills/longbridge-derivatives/references/options-advanced.md +101 -0
- package/bundled-skills/longbridge-derivatives/references/options-pnl.md +74 -0
- package/bundled-skills/longbridge-derivatives/references/options-strategy.md +82 -0
- package/bundled-skills/longbridge-derivatives/references/options-volatility.md +70 -0
- package/bundled-skills/longbridge-derivatives/references/warrant.md +12 -0
- package/bundled-skills/longbridge-quant/SKILL.md +151 -0
- package/bundled-skills/longbridge-quant/references/correlation.md +51 -0
- package/bundled-skills/longbridge-quant/references/execution-model.md +68 -0
- package/bundled-skills/longbridge-quant/references/factor-research.md +95 -0
- package/bundled-skills/longbridge-quant/references/factor-screen.md +101 -0
- package/bundled-skills/longbridge-quant/references/hedging.md +136 -0
- package/bundled-skills/longbridge-quant/references/ml-strategy.md +77 -0
- package/bundled-skills/longbridge-quant/references/multifactor.md +68 -0
- package/bundled-skills/longbridge-quant/references/pairs-trading.md +61 -0
- package/bundled-skills/longbridge-quant/references/quant-cli.md +133 -0
- package/bundled-skills/longbridge-quant/references/quant-stats.md +150 -0
- package/bundled-skills/longbridge-quant/references/seasonality.md +50 -0
- package/bundled-skills/longbridge-quant/references/strategy-optimizer.md +68 -0
- package/bundled-skills/longbridge-quant/references/volatility-strategy.md +52 -0
- package/bundled-skills/longbridge-research/SKILL.md +187 -0
- package/bundled-skills/longbridge-research/references/company-profile.md +96 -0
- package/bundled-skills/longbridge-research/references/company-tearsheet.md +82 -0
- package/bundled-skills/longbridge-research/references/competitive-analysis.md +81 -0
- package/bundled-skills/longbridge-research/references/consensus.md +92 -0
- package/bundled-skills/longbridge-research/references/coverage-initiation.md +76 -0
- package/bundled-skills/longbridge-research/references/defi-yield.md +60 -0
- package/bundled-skills/longbridge-research/references/finance-calendar.md +165 -0
- package/bundled-skills/longbridge-research/references/financial-planning.md +77 -0
- package/bundled-skills/longbridge-research/references/forecast-eps.md +39 -0
- package/bundled-skills/longbridge-research/references/fund-holder.md +44 -0
- package/bundled-skills/longbridge-research/references/hkipo-analysis.md +101 -0
- package/bundled-skills/longbridge-research/references/industry-peers.md +46 -0
- package/bundled-skills/longbridge-research/references/industry-rank.md +62 -0
- package/bundled-skills/longbridge-research/references/insider-trades.md +48 -0
- package/bundled-skills/longbridge-research/references/institution-rating.md +62 -0
- package/bundled-skills/longbridge-research/references/investment-ideas.md +69 -0
- package/bundled-skills/longbridge-research/references/investment-proposal.md +95 -0
- package/bundled-skills/longbridge-research/references/investors.md +87 -0
- package/bundled-skills/longbridge-research/references/onchain.md +70 -0
- package/bundled-skills/longbridge-research/references/post-investment.md +76 -0
- package/bundled-skills/longbridge-research/references/shareholder.md +72 -0
- package/bundled-skills/longbridge-research/references/short-positions.md +50 -0
- package/bundled-skills/longbridge-research/references/short-trades.md +50 -0
- package/bundled-skills/longbridge-research/references/stock-research.md +61 -0
- package/bundled-skills/longbridge-research/references/thesis-tracker.md +64 -0
- package/bundled-skills/m365-entra-attack/SKILL.md +423 -0
- package/bundled-skills/mac-mini-llm-lab/SKILL.md +350 -0
- package/bundled-skills/makepad-2-0-animation/SKILL.md +318 -0
- package/bundled-skills/makepad-2-0-animation/references/animator-reference.md +433 -0
- package/bundled-skills/makepad-2-0-dsl/SKILL.md +492 -0
- package/bundled-skills/makepad-2-0-dsl/references/dsl-syntax-reference.md +511 -0
- package/bundled-skills/makepad-2-0-dsl/references/extended-guide.md +56 -0
- package/bundled-skills/makepad-2-0-dsl/references/property-system.md +757 -0
- package/bundled-skills/makepad-2-0-events/SKILL.md +497 -0
- package/bundled-skills/makepad-2-0-events/references/event-patterns.md +802 -0
- package/bundled-skills/makepad-2-0-events/references/extended-guide.md +590 -0
- package/bundled-skills/makepad-2-0-layout/SKILL.md +499 -0
- package/bundled-skills/makepad-2-0-layout/references/extended-guide.md +243 -0
- package/bundled-skills/makepad-2-0-layout/references/layout-patterns.md +881 -0
- package/bundled-skills/makepad-2-0-widgets/SKILL.md +261 -0
- package/bundled-skills/makepad-2-0-widgets/references/widget-advanced.md +648 -0
- package/bundled-skills/makepad-2-0-widgets/references/widget-catalog.md +547 -0
- package/bundled-skills/mcp-server-security/SKILL.md +356 -0
- package/bundled-skills/mcp-server-security/references/details.md +745 -0
- package/bundled-skills/mdm-device-management/SKILL.md +404 -0
- package/bundled-skills/mdm-device-management/references/details.md +410 -0
- package/bundled-skills/meeting-distiller-pro/SKILL.md +120 -0
- package/bundled-skills/meme-coin-audit/SKILL.md +402 -0
- package/bundled-skills/mid-engagement-ir-detection/SKILL.md +377 -0
- package/bundled-skills/model-registry-governance/SKILL.md +452 -0
- package/bundled-skills/model-serving-kubernetes/SKILL.md +339 -0
- package/bundled-skills/model-supply-chain-security/SKILL.md +427 -0
- package/bundled-skills/mongodb/SKILL.md +436 -0
- package/bundled-skills/monte-carlo-analyze-root-cause/SKILL.md +12 -1
- package/bundled-skills/monte-carlo-asset-health/SKILL.md +12 -1
- package/bundled-skills/monte-carlo-context-detection/SKILL.md +170 -0
- package/bundled-skills/monte-carlo-context-detection/references/signal-definitions.md +46 -0
- package/bundled-skills/multi-tenant-llm-hosting/SKILL.md +435 -0
- package/bundled-skills/multi-tenant-llm-hosting/references/details.md +211 -0
- package/bundled-skills/mysql/SKILL.md +390 -0
- package/bundled-skills/new-relic/SKILL.md +472 -0
- package/bundled-skills/nfs-storage/SKILL.md +356 -0
- package/bundled-skills/object-storage/SKILL.md +378 -0
- package/bundled-skills/offensive-osint/SKILL.md +443 -0
- package/bundled-skills/okta-attack/SKILL.md +436 -0
- package/bundled-skills/ollama-stack/SKILL.md +379 -0
- package/bundled-skills/openclaw-deployment-hardening/SKILL.md +135 -0
- package/bundled-skills/openclaw-local-mac-mini/SKILL.md +426 -0
- package/bundled-skills/openclaw-local-mac-mini/references/details.md +221 -0
- package/bundled-skills/openclaw-security-hardening/SKILL.md +135 -0
- package/bundled-skills/openshift/SKILL.md +485 -0
- package/bundled-skills/opentelemetry/SKILL.md +438 -0
- package/bundled-skills/opentelemetry/references/details.md +78 -0
- package/bundled-skills/opentofu-migration/SKILL.md +349 -0
- package/bundled-skills/osint-methodology/SKILL.md +460 -0
- package/bundled-skills/osint-methodology/references/details.md +1350 -0
- package/bundled-skills/pci-dss-compliance/SKILL.md +446 -0
- package/bundled-skills/penetration-testing/SKILL.md +152 -0
- package/bundled-skills/performance-tuning/SKILL.md +381 -0
- package/bundled-skills/planetscale/SKILL.md +297 -0
- package/bundled-skills/platform-engineering/SKILL.md +348 -0
- package/bundled-skills/platform-engineering/references/details.md +944 -0
- package/bundled-skills/podman/SKILL.md +405 -0
- package/bundled-skills/policy-as-code/SKILL.md +434 -0
- package/bundled-skills/policy-as-code/references/details.md +204 -0
- package/bundled-skills/postgresql-devsec/SKILL.md +378 -0
- package/bundled-skills/prometheus-grafana/SKILL.md +469 -0
- package/bundled-skills/prompt-injection-defense/SKILL.md +483 -0
- package/bundled-skills/rag-infrastructure/SKILL.md +269 -0
- package/bundled-skills/rag-observability-evals/SKILL.md +444 -0
- package/bundled-skills/rag-observability-evals/references/details.md +92 -0
- package/bundled-skills/recon-scope-triage/SKILL.md +128 -0
- package/bundled-skills/redis/SKILL.md +421 -0
- package/bundled-skills/redteam-report-template/SKILL.md +370 -0
- package/bundled-skills/remotion-captions/SKILL.md +57 -0
- package/bundled-skills/remotion-captions/agents/openai.yaml +7 -0
- package/bundled-skills/remotion-captions/assets/remotion-icon.svg +4 -0
- package/bundled-skills/remotion-captions/display-captions.md +190 -0
- package/bundled-skills/remotion-captions/import-srt-captions.md +73 -0
- package/bundled-skills/remotion-captions/transcribe-captions.md +70 -0
- package/bundled-skills/remotion-create/SKILL.md +106 -0
- package/bundled-skills/remotion-create/agents/openai.yaml +7 -0
- package/bundled-skills/remotion-create/assets/remotion-icon.svg +4 -0
- package/bundled-skills/remotion-create/tailwind.md +11 -0
- package/bundled-skills/remotion-create/video-layout.md +9 -0
- package/bundled-skills/remotion-docs/SKILL.md +67 -0
- package/bundled-skills/remotion-docs/agents/openai.yaml +7 -0
- package/bundled-skills/remotion-docs/assets/remotion-icon.svg +4 -0
- package/bundled-skills/remotion-interactivity/SKILL.md +270 -0
- package/bundled-skills/remotion-interactivity/agents/openai.yaml +7 -0
- package/bundled-skills/remotion-interactivity/assets/remotion-icon.svg +4 -0
- package/bundled-skills/remotion-render/SKILL.md +48 -0
- package/bundled-skills/remotion-render/agents/openai.yaml +7 -0
- package/bundled-skills/remotion-render/assets/remotion-icon.svg +4 -0
- package/bundled-skills/remotion-render/transparent-videos.md +106 -0
- package/bundled-skills/report-writing/SKILL.md +426 -0
- package/bundled-skills/report-writing/references/details.md +187 -0
- package/bundled-skills/reverse-proxy/SKILL.md +420 -0
- package/bundled-skills/runbook-creation/SKILL.md +438 -0
- package/bundled-skills/runbook-creation/references/details.md +71 -0
- package/bundled-skills/saas-pricing-strategist/SKILL.md +169 -0
- package/bundled-skills/saas-security-posture/SKILL.md +415 -0
- package/bundled-skills/sast-scanning/SKILL.md +444 -0
- package/bundled-skills/sbom-supply-chain/SKILL.md +433 -0
- package/bundled-skills/score-eval/SKILL.md +35 -0
- package/bundled-skills/security-arsenal/SKILL.md +446 -0
- package/bundled-skills/security-arsenal/references/details.md +540 -0
- package/bundled-skills/security-automation/SKILL.md +146 -0
- package/bundled-skills/semantic-versioning/SKILL.md +434 -0
- package/bundled-skills/semantic-versioning/references/details.md +83 -0
- package/bundled-skills/service-mesh/SKILL.md +422 -0
- package/bundled-skills/soc2-compliance/SKILL.md +409 -0
- package/bundled-skills/sops-encryption/SKILL.md +124 -0
- package/bundled-skills/sre-dashboards/SKILL.md +143 -0
- package/bundled-skills/ssh-configuration/SKILL.md +324 -0
- package/bundled-skills/ssl-tls-management/SKILL.md +428 -0
- package/bundled-skills/ssl-tls-management/references/details.md +99 -0
- package/bundled-skills/startup-it-troubleshooting/SKILL.md +415 -0
- package/bundled-skills/supply-chain-attack-recon/SKILL.md +453 -0
- package/bundled-skills/supply-chain-attack-recon/references/details.md +258 -0
- package/bundled-skills/systemd-services/SKILL.md +379 -0
- package/bundled-skills/terraform-aws/SKILL.md +125 -0
- package/bundled-skills/terraform-azure/SKILL.md +415 -0
- package/bundled-skills/terraform-azure/references/details.md +231 -0
- package/bundled-skills/terraform-gcp/SKILL.md +369 -0
- package/bundled-skills/threat-modeling/SKILL.md +487 -0
- package/bundled-skills/user-management/SKILL.md +383 -0
- package/bundled-skills/using-agent-skills/SKILL.md +220 -0
- package/bundled-skills/vector-database-ops/SKILL.md +300 -0
- package/bundled-skills/vendor-management/SKILL.md +439 -0
- package/bundled-skills/vendor-management/references/details.md +109 -0
- package/bundled-skills/vercel-deployments/SKILL.md +296 -0
- package/bundled-skills/vllm-server/SKILL.md +236 -0
- package/bundled-skills/vmware-vcenter-attack/SKILL.md +412 -0
- package/bundled-skills/vpn-setup/SKILL.md +452 -0
- package/bundled-skills/vulnerability-scanning/SKILL.md +448 -0
- package/bundled-skills/waf-setup/SKILL.md +354 -0
- package/bundled-skills/waf-setup/references/details.md +211 -0
- package/bundled-skills/web2-recon/SKILL.md +440 -0
- package/bundled-skills/web2-recon/references/details.md +319 -0
- package/bundled-skills/web3-audit/SKILL.md +445 -0
- package/bundled-skills/web3-audit/references/details.md +224 -0
- package/bundled-skills/windows-hardening/SKILL.md +454 -0
- package/bundled-skills/windows-hardening/references/details.md +204 -0
- package/bundled-skills/windows-server/SKILL.md +318 -0
- package/bundled-skills/writing-guidelines/SKILL.md +60 -0
- package/bundled-skills/zero-trust/SKILL.md +461 -0
- package/package.json +1 -1
- package/skills_index.json +7874 -277
|
@@ -0,0 +1,334 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: llm-caching
|
|
3
|
+
description: Implement multi-layer LLM caching with exact match, semantic similarity,
|
|
4
|
+
and provider-side prompt caching.
|
|
5
|
+
category: devops
|
|
6
|
+
risk: critical
|
|
7
|
+
source: https://github.com/BagelHole/DevOps-Security-Agent-Skills
|
|
8
|
+
source_repo: BagelHole/DevOps-Security-Agent-Skills
|
|
9
|
+
source_type: community
|
|
10
|
+
date_added: '2026-09-20'
|
|
11
|
+
license: MIT
|
|
12
|
+
license_source: https://github.com/BagelHole/DevOps-Security-Agent-Skills/blob/main/LICENSE
|
|
13
|
+
compatibility: Requires the relevant platform CLIs (kubectl, helm, terraform, git,
|
|
14
|
+
CI runners) and authorized access to the target environment. Docs-only; helper scripts
|
|
15
|
+
and templates not bundled.
|
|
16
|
+
metadata:
|
|
17
|
+
author: devops-skills
|
|
18
|
+
version: '1.0'
|
|
19
|
+
---
|
|
20
|
+
|
|
21
|
+
# LLM Caching
|
|
22
|
+
|
|
23
|
+
Cut LLM costs and latency with exact match, semantic, and provider-side caching layers.
|
|
24
|
+
|
|
25
|
+
## When to Use This Skill
|
|
26
|
+
|
|
27
|
+
Use this skill when:
|
|
28
|
+
- The same or similar queries are asked repeatedly (FAQ bots, support tools)
|
|
29
|
+
- LLM API costs are growing and you need immediate savings
|
|
30
|
+
- Serving high request volumes where repeated queries cause bottlenecks
|
|
31
|
+
- Implementing prompt caching for long system prompts (Anthropic/OpenAI)
|
|
32
|
+
- Building offline-capable AI features that need response persistence
|
|
33
|
+
|
|
34
|
+
## Caching Layers
|
|
35
|
+
|
|
36
|
+
```
|
|
37
|
+
Request → Exact Cache → Semantic Cache → Provider Cache → LLM API
|
|
38
|
+
↓ hit ↓ hit ↓ hit
|
|
39
|
+
instant ~5ms 50-80% cheaper
|
|
40
|
+
```
|
|
41
|
+
|
|
42
|
+
## Layer 1: Exact Match Cache (Redis)
|
|
43
|
+
|
|
44
|
+
```python
|
|
45
|
+
import hashlib
|
|
46
|
+
import json
|
|
47
|
+
import redis
|
|
48
|
+
from openai import OpenAI
|
|
49
|
+
|
|
50
|
+
r = redis.Redis(host="localhost", port=6379, decode_responses=True)
|
|
51
|
+
client = OpenAI()
|
|
52
|
+
|
|
53
|
+
def build_cache_key(model: str, messages: list, temperature: float) -> str:
|
|
54
|
+
"""Deterministic key from request parameters."""
|
|
55
|
+
payload = json.dumps({
|
|
56
|
+
"model": model,
|
|
57
|
+
"messages": messages,
|
|
58
|
+
"temperature": temperature,
|
|
59
|
+
}, sort_keys=True)
|
|
60
|
+
return f"llm:exact:{hashlib.sha256(payload.encode()).hexdigest()}"
|
|
61
|
+
|
|
62
|
+
def cached_completion(model: str, messages: list, temperature: float = 0.0,
|
|
63
|
+
ttl: int = 3600) -> dict:
|
|
64
|
+
key = build_cache_key(model, messages, temperature)
|
|
65
|
+
|
|
66
|
+
# Check cache
|
|
67
|
+
if cached := r.get(key):
|
|
68
|
+
return json.loads(cached)
|
|
69
|
+
|
|
70
|
+
# Call API
|
|
71
|
+
response = client.chat.completions.create(
|
|
72
|
+
model=model, messages=messages, temperature=temperature
|
|
73
|
+
)
|
|
74
|
+
result = response.model_dump()
|
|
75
|
+
|
|
76
|
+
# Cache result (only cache deterministic responses)
|
|
77
|
+
if temperature == 0.0:
|
|
78
|
+
r.setex(key, ttl, json.dumps(result))
|
|
79
|
+
|
|
80
|
+
return result
|
|
81
|
+
```
|
|
82
|
+
|
|
83
|
+
## Layer 2: Semantic Cache (GPTCache)
|
|
84
|
+
|
|
85
|
+
```python
|
|
86
|
+
from gptcache import cache, Config
|
|
87
|
+
from gptcache.adapter import openai
|
|
88
|
+
from gptcache.embedding import Onnx
|
|
89
|
+
from gptcache.manager import CacheBase, VectorBase, get_data_manager
|
|
90
|
+
from gptcache.similarity_evaluation.distance import SearchDistanceEvaluation
|
|
91
|
+
|
|
92
|
+
# Configure GPTCache with Qdrant backend
|
|
93
|
+
def init_gptcache(cache_obj, llm: str):
|
|
94
|
+
onnx = Onnx() # local embedding model
|
|
95
|
+
data_manager = get_data_manager(
|
|
96
|
+
CacheBase("redis"), # metadata store
|
|
97
|
+
VectorBase("qdrant",
|
|
98
|
+
host="localhost",
|
|
99
|
+
port=6333,
|
|
100
|
+
collection_name=f"llm-cache-{llm}",
|
|
101
|
+
dimension=onnx.dimension),
|
|
102
|
+
)
|
|
103
|
+
cache_obj.init(
|
|
104
|
+
embedding_func=onnx.to_embeddings,
|
|
105
|
+
data_manager=data_manager,
|
|
106
|
+
similarity_evaluation=SearchDistanceEvaluation(),
|
|
107
|
+
config=Config(similarity_threshold=0.80), # 80% similarity = cache hit
|
|
108
|
+
)
|
|
109
|
+
|
|
110
|
+
cache.set_openai_key()
|
|
111
|
+
init_gptcache(cache, "gpt-4o-mini")
|
|
112
|
+
|
|
113
|
+
# Now openai calls are automatically cached
|
|
114
|
+
response = openai.ChatCompletion.create(
|
|
115
|
+
model="gpt-4o-mini",
|
|
116
|
+
messages=[{"role": "user", "content": "What is machine learning?"}],
|
|
117
|
+
)
|
|
118
|
+
# Second call with similar question ("Explain machine learning") → cache hit
|
|
119
|
+
```
|
|
120
|
+
|
|
121
|
+
## Custom Semantic Cache (Production-Grade)
|
|
122
|
+
|
|
123
|
+
```python
|
|
124
|
+
from sentence_transformers import SentenceTransformer
|
|
125
|
+
from qdrant_client import QdrantClient
|
|
126
|
+
from qdrant_client.models import Distance, VectorParams, PointStruct, Filter, FieldCondition, Range
|
|
127
|
+
import numpy as np
|
|
128
|
+
import uuid
|
|
129
|
+
import time
|
|
130
|
+
|
|
131
|
+
embed_model = SentenceTransformer("BAAI/bge-small-en-v1.5") # fast, 33M params
|
|
132
|
+
qdrant = QdrantClient("http://localhost:6333")
|
|
133
|
+
|
|
134
|
+
CACHE_COLLECTION = "semantic-cache"
|
|
135
|
+
SIMILARITY_THRESHOLD = 0.88
|
|
136
|
+
CACHE_TTL_SECONDS = 86400 # 24h
|
|
137
|
+
|
|
138
|
+
# Create collection once
|
|
139
|
+
qdrant.create_collection(
|
|
140
|
+
collection_name=CACHE_COLLECTION,
|
|
141
|
+
vectors_config=VectorParams(size=384, distance=Distance.COSINE),
|
|
142
|
+
on_disk_payload=True,
|
|
143
|
+
)
|
|
144
|
+
|
|
145
|
+
def semantic_cache_lookup(query: str, model: str) -> str | None:
|
|
146
|
+
embedding = embed_model.encode(query).tolist()
|
|
147
|
+
results = qdrant.query_points(
|
|
148
|
+
collection_name=CACHE_COLLECTION,
|
|
149
|
+
query=embedding,
|
|
150
|
+
query_filter=Filter(must=[
|
|
151
|
+
FieldCondition(key="model", match={"value": model}),
|
|
152
|
+
FieldCondition(key="expires_at", range=Range(gte=time.time())),
|
|
153
|
+
]),
|
|
154
|
+
limit=1,
|
|
155
|
+
score_threshold=SIMILARITY_THRESHOLD,
|
|
156
|
+
)
|
|
157
|
+
if results.points:
|
|
158
|
+
return results.points[0].payload["response"]
|
|
159
|
+
return None
|
|
160
|
+
|
|
161
|
+
def semantic_cache_store(query: str, response: str, model: str):
|
|
162
|
+
embedding = embed_model.encode(query).tolist()
|
|
163
|
+
qdrant.upsert(
|
|
164
|
+
collection_name=CACHE_COLLECTION,
|
|
165
|
+
points=[PointStruct(
|
|
166
|
+
id=str(uuid.uuid4()),
|
|
167
|
+
vector=embedding,
|
|
168
|
+
payload={
|
|
169
|
+
"query": query,
|
|
170
|
+
"response": response,
|
|
171
|
+
"model": model,
|
|
172
|
+
"created_at": time.time(),
|
|
173
|
+
"expires_at": time.time() + CACHE_TTL_SECONDS,
|
|
174
|
+
},
|
|
175
|
+
)],
|
|
176
|
+
)
|
|
177
|
+
|
|
178
|
+
def smart_llm_call(query: str, model: str = "gpt-4o-mini") -> dict:
|
|
179
|
+
# 1. Semantic lookup
|
|
180
|
+
if cached_response := semantic_cache_lookup(query, model):
|
|
181
|
+
return {"response": cached_response, "source": "semantic_cache", "cost": 0}
|
|
182
|
+
|
|
183
|
+
# 2. LLM call
|
|
184
|
+
response = client.chat.completions.create(
|
|
185
|
+
model=model,
|
|
186
|
+
messages=[{"role": "user", "content": query}],
|
|
187
|
+
)
|
|
188
|
+
text = response.choices[0].message.content
|
|
189
|
+
cost = litellm.completion_cost(response)
|
|
190
|
+
|
|
191
|
+
# 3. Store in cache
|
|
192
|
+
semantic_cache_store(query, text, model)
|
|
193
|
+
|
|
194
|
+
return {"response": text, "source": "llm_api", "cost": cost}
|
|
195
|
+
```
|
|
196
|
+
|
|
197
|
+
## Layer 3: Provider-Side Prompt Caching
|
|
198
|
+
|
|
199
|
+
```python
|
|
200
|
+
# Anthropic — cache long system prompts (saves 90% on cached input tokens)
|
|
201
|
+
import anthropic
|
|
202
|
+
|
|
203
|
+
client = anthropic.Anthropic()
|
|
204
|
+
|
|
205
|
+
# Long system prompt — mark for caching
|
|
206
|
+
SYSTEM_PROMPT = open("knowledge-base.txt").read() # e.g., 50k tokens
|
|
207
|
+
|
|
208
|
+
def call_with_prompt_cache(user_question: str) -> str:
|
|
209
|
+
response = client.messages.create(
|
|
210
|
+
model="claude-sonnet-4-6",
|
|
211
|
+
max_tokens=1024,
|
|
212
|
+
system=[
|
|
213
|
+
{"type": "text", "text": "You are a helpful assistant."},
|
|
214
|
+
{
|
|
215
|
+
"type": "text",
|
|
216
|
+
"text": SYSTEM_PROMPT,
|
|
217
|
+
"cache_control": {"type": "ephemeral"}, # cache this block
|
|
218
|
+
}
|
|
219
|
+
],
|
|
220
|
+
messages=[{"role": "user", "content": user_question}],
|
|
221
|
+
)
|
|
222
|
+
# Log cache efficiency
|
|
223
|
+
usage = response.usage
|
|
224
|
+
cache_savings = usage.cache_read_input_tokens * 0.9 # 90% discount on cached
|
|
225
|
+
print(f"Cache hits: {usage.cache_read_input_tokens} tokens "
|
|
226
|
+
f"(saved ~${cache_savings * 3.0 / 1_000_000:.4f})")
|
|
227
|
+
return response.content[0].text
|
|
228
|
+
|
|
229
|
+
# OpenAI — automatic for repeated prefixes (≥1,024 tokens)
|
|
230
|
+
# No code change needed; cached tokens appear in usage.prompt_tokens_details
|
|
231
|
+
response = client.chat.completions.create(
|
|
232
|
+
model="gpt-4o-mini",
|
|
233
|
+
messages=[
|
|
234
|
+
{"role": "system", "content": LONG_SYSTEM_PROMPT}, # auto-cached
|
|
235
|
+
{"role": "user", "content": user_question},
|
|
236
|
+
]
|
|
237
|
+
)
|
|
238
|
+
cached = response.usage.prompt_tokens_details.cached_tokens
|
|
239
|
+
print(f"OpenAI cached {cached} tokens")
|
|
240
|
+
```
|
|
241
|
+
|
|
242
|
+
## Cache Warming
|
|
243
|
+
|
|
244
|
+
```python
|
|
245
|
+
async def warm_cache(common_queries: list[str], model: str):
|
|
246
|
+
"""Pre-populate cache with known frequent queries."""
|
|
247
|
+
import asyncio
|
|
248
|
+
from openai import AsyncOpenAI
|
|
249
|
+
|
|
250
|
+
aclient = AsyncOpenAI()
|
|
251
|
+
|
|
252
|
+
async def warm_single(query: str):
|
|
253
|
+
if not semantic_cache_lookup(query, model):
|
|
254
|
+
response = await aclient.chat.completions.create(
|
|
255
|
+
model=model,
|
|
256
|
+
messages=[{"role": "user", "content": query}],
|
|
257
|
+
)
|
|
258
|
+
text = response.choices[0].message.content
|
|
259
|
+
semantic_cache_store(query, text, model)
|
|
260
|
+
print(f"Warmed: {query[:50]}...")
|
|
261
|
+
|
|
262
|
+
await asyncio.gather(*[warm_single(q) for q in common_queries])
|
|
263
|
+
|
|
264
|
+
# Warm on startup
|
|
265
|
+
import asyncio
|
|
266
|
+
asyncio.run(warm_cache(FREQUENT_QUERIES, "gpt-4o-mini"))
|
|
267
|
+
```
|
|
268
|
+
|
|
269
|
+
## Cache Metrics
|
|
270
|
+
|
|
271
|
+
```python
|
|
272
|
+
from prometheus_client import Counter, Histogram
|
|
273
|
+
|
|
274
|
+
cache_hits = Counter("llm_cache_hits_total", "Cache hits", ["cache_layer", "model"])
|
|
275
|
+
cache_misses = Counter("llm_cache_misses_total", "Cache misses", ["model"])
|
|
276
|
+
cache_savings_usd = Counter("llm_cache_savings_usd_total", "USD saved by cache", ["model"])
|
|
277
|
+
|
|
278
|
+
# Use in your smart_llm_call function
|
|
279
|
+
if source == "semantic_cache":
|
|
280
|
+
cache_hits.labels(cache_layer="semantic", model=model).inc()
|
|
281
|
+
cache_savings_usd.labels(model=model).inc(estimated_cost)
|
|
282
|
+
else:
|
|
283
|
+
cache_misses.labels(model=model).inc()
|
|
284
|
+
```
|
|
285
|
+
|
|
286
|
+
## Redis Configuration for LLM Caching
|
|
287
|
+
|
|
288
|
+
```bash
|
|
289
|
+
# redis.conf tuning for LLM cache workload
|
|
290
|
+
maxmemory 8gb
|
|
291
|
+
maxmemory-policy allkeys-lru # evict least-recently-used when full
|
|
292
|
+
save "" # disable persistence (cache is ephemeral)
|
|
293
|
+
appendonly no
|
|
294
|
+
tcp-keepalive 60
|
|
295
|
+
```
|
|
296
|
+
|
|
297
|
+
## Common Issues
|
|
298
|
+
|
|
299
|
+
| Issue | Cause | Fix |
|
|
300
|
+
|-------|-------|-----|
|
|
301
|
+
| Low cache hit rate | Threshold too strict | Lower `SIMILARITY_THRESHOLD` to 0.82–0.85 |
|
|
302
|
+
| Stale cached responses | Long TTL | Use topic-specific TTLs; invalidate on data updates |
|
|
303
|
+
| Cache serving wrong answers | Threshold too loose | Raise threshold or add model-name filtering |
|
|
304
|
+
| Redis OOM | No eviction policy | Set `maxmemory` + `allkeys-lru` |
|
|
305
|
+
| Slow semantic lookup | Large cache collection | Add payload index on `model` + `expires_at` |
|
|
306
|
+
|
|
307
|
+
## Best Practices
|
|
308
|
+
|
|
309
|
+
- Start with exact cache — zero cost, instant wins for identical queries.
|
|
310
|
+
- Semantic threshold of 0.88–0.92 balances hit rate vs. accuracy; tune with your data.
|
|
311
|
+
- Set per-model TTLs: longer for stable knowledge (1 week), shorter for news/events (1 hour).
|
|
312
|
+
- Always filter by model name in semantic cache — different models give different answers.
|
|
313
|
+
- Log cache hit rate as a KPI; target 30%+ for FAQ-style applications.
|
|
314
|
+
|
|
315
|
+
## Related Skills
|
|
316
|
+
|
|
317
|
+
- llm-cost-optimization (`llm-cost-optimization`) - Full cost strategy
|
|
318
|
+
- llm-gateway (`llm-gateway`) - Gateway-level caching
|
|
319
|
+
- vector-database-ops (`vector-database-ops`) - Qdrant setup
|
|
320
|
+
- agent-observability (`agent-observability`) - Cache metrics dashboards
|
|
321
|
+
|
|
322
|
+
## Limitations
|
|
323
|
+
|
|
324
|
+
- Guidance executes against real environments: confirm target, blast radius, and rollback plan before applying anything.
|
|
325
|
+
- Never deploy to production without explicit approval. Docs-only import: upstream scripts and templates not bundled.
|
|
326
|
+
|
|
327
|
+
### Example
|
|
328
|
+
|
|
329
|
+
```bash
|
|
330
|
+
git status && git diff --stat
|
|
331
|
+
kubectl diff -f manifest.yaml
|
|
332
|
+
```
|
|
333
|
+
|
|
334
|
+
> Adapted from [BagelHole/DevOps-Security-Agent-Skills](https://github.com/BagelHole/DevOps-Security-Agent-Skills) (MIT); frontmatter, When to Use/Limitations, and safety boundaries added for upstream compliance. Docs-only import: helper scripts and templates not bundled.
|
|
@@ -0,0 +1,311 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: llm-cost-optimization
|
|
3
|
+
description: Reduce LLM API and infrastructure costs through model selection, prompt
|
|
4
|
+
caching, batching, caching, quantization, and self-hosting strategies.
|
|
5
|
+
category: devops
|
|
6
|
+
risk: critical
|
|
7
|
+
source: https://github.com/BagelHole/DevOps-Security-Agent-Skills
|
|
8
|
+
source_repo: BagelHole/DevOps-Security-Agent-Skills
|
|
9
|
+
source_type: community
|
|
10
|
+
date_added: '2026-09-20'
|
|
11
|
+
license: MIT
|
|
12
|
+
license_source: https://github.com/BagelHole/DevOps-Security-Agent-Skills/blob/main/LICENSE
|
|
13
|
+
compatibility: Requires the relevant platform CLIs (kubectl, helm, terraform, git,
|
|
14
|
+
CI runners) and authorized access to the target environment. Docs-only; helper scripts
|
|
15
|
+
and templates not bundled.
|
|
16
|
+
metadata:
|
|
17
|
+
author: devops-skills
|
|
18
|
+
version: '1.0'
|
|
19
|
+
---
|
|
20
|
+
|
|
21
|
+
# LLM Cost Optimization
|
|
22
|
+
|
|
23
|
+
Cut LLM costs by 50–90% with the right combination of caching, model selection, prompt optimization, and self-hosting.
|
|
24
|
+
|
|
25
|
+
## When to Use This Skill
|
|
26
|
+
|
|
27
|
+
Use this skill when:
|
|
28
|
+
- LLM API spend is growing faster than revenue
|
|
29
|
+
- You need to attribute AI costs to teams, products, or customers
|
|
30
|
+
- Implementing caching to avoid redundant LLM calls
|
|
31
|
+
- Deciding when to switch from API providers to self-hosted models
|
|
32
|
+
- Optimizing prompt length without sacrificing quality
|
|
33
|
+
|
|
34
|
+
## Cost Levers by Impact
|
|
35
|
+
|
|
36
|
+
| Strategy | Typical Savings | Effort |
|
|
37
|
+
|----------|-----------------|--------|
|
|
38
|
+
| Semantic caching | 20–50% | Low |
|
|
39
|
+
| Model right-sizing | 30–70% | Low |
|
|
40
|
+
| Prompt compression | 10–30% | Medium |
|
|
41
|
+
| Provider caching (prompt cache) | 10–25% | Low |
|
|
42
|
+
| Batching offline workloads | 50% (Batch API) | Medium |
|
|
43
|
+
| Self-hosting 7–8B models | 80–95% at scale | High |
|
|
44
|
+
| Quantization | 30–50% VRAM cost | Medium |
|
|
45
|
+
|
|
46
|
+
## Track Costs First
|
|
47
|
+
|
|
48
|
+
```python
|
|
49
|
+
# Use LiteLLM's cost tracking (automatic per-model pricing)
|
|
50
|
+
import litellm
|
|
51
|
+
|
|
52
|
+
response = litellm.completion(
|
|
53
|
+
model="gpt-4o-mini",
|
|
54
|
+
messages=[{"role": "user", "content": "Hello"}],
|
|
55
|
+
)
|
|
56
|
+
cost = litellm.completion_cost(response)
|
|
57
|
+
print(f"Cost: ${cost:.6f}")
|
|
58
|
+
|
|
59
|
+
# Add custom cost callbacks
|
|
60
|
+
def log_cost(kwargs, completion_response, start_time, end_time):
|
|
61
|
+
cost = kwargs.get("response_cost", 0)
|
|
62
|
+
model = kwargs.get("model")
|
|
63
|
+
user = kwargs.get("user")
|
|
64
|
+
# Send to your analytics DB
|
|
65
|
+
db.record_cost(user=user, model=model, cost=cost)
|
|
66
|
+
|
|
67
|
+
litellm.success_callback = [log_cost]
|
|
68
|
+
```
|
|
69
|
+
|
|
70
|
+
## Model Right-Sizing
|
|
71
|
+
|
|
72
|
+
```python
|
|
73
|
+
# Route by task complexity — don't use GPT-4o for everything
|
|
74
|
+
def get_model_for_task(task_type: str) -> str:
|
|
75
|
+
routing = {
|
|
76
|
+
"classification": "gpt-4o-mini", # ~30× cheaper than gpt-4o
|
|
77
|
+
"summarization": "gpt-4o-mini",
|
|
78
|
+
"extraction": "gpt-4o-mini",
|
|
79
|
+
"simple_qa": "gpt-4o-mini",
|
|
80
|
+
"complex_reasoning": "gpt-4o",
|
|
81
|
+
"code_generation": "claude-sonnet-4-6",
|
|
82
|
+
"creative_writing": "claude-opus-4-6",
|
|
83
|
+
}
|
|
84
|
+
return routing.get(task_type, "gpt-4o-mini")
|
|
85
|
+
|
|
86
|
+
# Cost comparison (per 1M tokens, 2025 approx.)
|
|
87
|
+
# gpt-4o-mini: input $0.15 / output $0.60
|
|
88
|
+
# gpt-4o: input $2.50 / output $10.00
|
|
89
|
+
# claude-sonnet-4-6: input $3.00 / output $15.00
|
|
90
|
+
# llama-3.1-8b (self): ~$0.05–0.10 all-in (GPU amortized)
|
|
91
|
+
```
|
|
92
|
+
|
|
93
|
+
## Prompt Caching (Provider-Side)
|
|
94
|
+
|
|
95
|
+
```python
|
|
96
|
+
# Anthropic — cache long system prompts (saves 90% on cached tokens)
|
|
97
|
+
import anthropic
|
|
98
|
+
|
|
99
|
+
client = anthropic.Anthropic()
|
|
100
|
+
|
|
101
|
+
response = client.messages.create(
|
|
102
|
+
model="claude-sonnet-4-6",
|
|
103
|
+
max_tokens=1024,
|
|
104
|
+
system=[
|
|
105
|
+
{
|
|
106
|
+
"type": "text",
|
|
107
|
+
"text": "You are a helpful assistant.",
|
|
108
|
+
},
|
|
109
|
+
{
|
|
110
|
+
"type": "text",
|
|
111
|
+
"text": open("large-context.txt").read(), # large doc
|
|
112
|
+
"cache_control": {"type": "ephemeral"}, # cache this!
|
|
113
|
+
}
|
|
114
|
+
],
|
|
115
|
+
messages=[{"role": "user", "content": "Summarize the key points."}],
|
|
116
|
+
)
|
|
117
|
+
# First call: full price. Subsequent calls: 90% discount on cached part.
|
|
118
|
+
print(f"Cache read tokens: {response.usage.cache_read_input_tokens}")
|
|
119
|
+
|
|
120
|
+
# OpenAI — prompt caching is automatic for repeated prefixes >1024 tokens
|
|
121
|
+
# No code change needed; check usage.prompt_tokens_details.cached_tokens
|
|
122
|
+
```
|
|
123
|
+
|
|
124
|
+
## Batching with OpenAI Batch API (50% Discount)
|
|
125
|
+
|
|
126
|
+
```python
|
|
127
|
+
import json
|
|
128
|
+
from openai import OpenAI
|
|
129
|
+
|
|
130
|
+
client = OpenAI()
|
|
131
|
+
|
|
132
|
+
# Prepare batch requests
|
|
133
|
+
requests = [
|
|
134
|
+
{
|
|
135
|
+
"custom_id": f"task-{i}",
|
|
136
|
+
"method": "POST",
|
|
137
|
+
"url": "/v1/chat/completions",
|
|
138
|
+
"body": {
|
|
139
|
+
"model": "gpt-4o-mini",
|
|
140
|
+
"messages": [{"role": "user", "content": f"Classify: {text}"}],
|
|
141
|
+
"max_tokens": 50,
|
|
142
|
+
}
|
|
143
|
+
}
|
|
144
|
+
for i, text in enumerate(texts)
|
|
145
|
+
]
|
|
146
|
+
|
|
147
|
+
# Write JSONL file
|
|
148
|
+
with open("batch.jsonl", "w") as f:
|
|
149
|
+
for req in requests:
|
|
150
|
+
f.write(json.dumps(req) + "\n")
|
|
151
|
+
|
|
152
|
+
# Upload and create batch
|
|
153
|
+
batch_file = client.files.create(file=open("batch.jsonl", "rb"), purpose="batch")
|
|
154
|
+
batch = client.batches.create(
|
|
155
|
+
input_file_id=batch_file.id,
|
|
156
|
+
endpoint="/v1/chat/completions",
|
|
157
|
+
completion_window="24h",
|
|
158
|
+
)
|
|
159
|
+
print(f"Batch ID: {batch.id}") # poll status with client.batches.retrieve(batch.id)
|
|
160
|
+
```
|
|
161
|
+
|
|
162
|
+
## Semantic Caching
|
|
163
|
+
|
|
164
|
+
```python
|
|
165
|
+
import hashlib
|
|
166
|
+
import json
|
|
167
|
+
import redis
|
|
168
|
+
import numpy as np
|
|
169
|
+
from sentence_transformers import SentenceTransformer
|
|
170
|
+
|
|
171
|
+
r = redis.Redis(host="localhost", port=6379)
|
|
172
|
+
embed_model = SentenceTransformer("BAAI/bge-small-en-v1.5")
|
|
173
|
+
|
|
174
|
+
SIMILARITY_THRESHOLD = 0.92
|
|
175
|
+
CACHE_TTL = 3600 * 24 # 24 hours
|
|
176
|
+
|
|
177
|
+
def cached_llm_call(prompt: str, llm_fn) -> str:
|
|
178
|
+
# 1. Exact match (free)
|
|
179
|
+
exact_key = f"exact:{hashlib.sha256(prompt.encode()).hexdigest()}"
|
|
180
|
+
if cached := r.get(exact_key):
|
|
181
|
+
return cached.decode()
|
|
182
|
+
|
|
183
|
+
# 2. Semantic match
|
|
184
|
+
query_vec = embed_model.encode(prompt)
|
|
185
|
+
cached_keys = r.keys("sem:*")
|
|
186
|
+
for key in cached_keys:
|
|
187
|
+
data = json.loads(r.get(key))
|
|
188
|
+
similarity = np.dot(query_vec, data["embedding"]) / (
|
|
189
|
+
np.linalg.norm(query_vec) * np.linalg.norm(data["embedding"])
|
|
190
|
+
)
|
|
191
|
+
if similarity >= SIMILARITY_THRESHOLD:
|
|
192
|
+
return data["response"]
|
|
193
|
+
|
|
194
|
+
# 3. Cache miss — call LLM
|
|
195
|
+
response = llm_fn(prompt)
|
|
196
|
+
|
|
197
|
+
# Store exact match
|
|
198
|
+
r.setex(exact_key, CACHE_TTL, response)
|
|
199
|
+
|
|
200
|
+
# Store semantic embedding
|
|
201
|
+
sem_key = f"sem:{hashlib.sha256(prompt.encode()).hexdigest()}"
|
|
202
|
+
r.setex(sem_key, CACHE_TTL, json.dumps({
|
|
203
|
+
"embedding": query_vec.tolist(),
|
|
204
|
+
"response": response,
|
|
205
|
+
"prompt": prompt,
|
|
206
|
+
}))
|
|
207
|
+
return response
|
|
208
|
+
```
|
|
209
|
+
|
|
210
|
+
## Prompt Compression
|
|
211
|
+
|
|
212
|
+
```python
|
|
213
|
+
# LLMLingua — compress long prompts by 3–20× with minimal quality loss
|
|
214
|
+
from llmlingua import PromptCompressor
|
|
215
|
+
|
|
216
|
+
compressor = PromptCompressor(
|
|
217
|
+
model_name="microsoft/llmlingua-2-bert-base-multilingual-cased-meetingbank",
|
|
218
|
+
device_map="cpu",
|
|
219
|
+
)
|
|
220
|
+
|
|
221
|
+
compressed = compressor.compress_prompt(
|
|
222
|
+
long_context,
|
|
223
|
+
ratio=0.5, # keep 50% of tokens
|
|
224
|
+
rank_method="longllmlingua",
|
|
225
|
+
)
|
|
226
|
+
print(f"Original: {len(long_context.split())} words")
|
|
227
|
+
print(f"Compressed: {len(compressed['compressed_prompt'].split())} words")
|
|
228
|
+
print(f"Savings: {compressed['saving']}")
|
|
229
|
+
```
|
|
230
|
+
|
|
231
|
+
## Self-Hosting Break-Even Calculator
|
|
232
|
+
|
|
233
|
+
```python
|
|
234
|
+
def break_even_analysis(
|
|
235
|
+
monthly_api_spend_usd: float,
|
|
236
|
+
gpu_cost_per_hour_usd: float = 2.50, # e.g., A10G on AWS
|
|
237
|
+
utilization: float = 0.70, # 70% GPU utilization
|
|
238
|
+
) -> dict:
|
|
239
|
+
monthly_gpu_cost = gpu_cost_per_hour_usd * 24 * 30 * utilization
|
|
240
|
+
break_even = monthly_gpu_cost / monthly_api_spend_usd
|
|
241
|
+
recommendation = (
|
|
242
|
+
"Self-host now — strong ROI" if break_even < 0.5 else
|
|
243
|
+
"Self-host if traffic grows 2×" if break_even < 0.8 else
|
|
244
|
+
"Stick with API — not enough scale yet"
|
|
245
|
+
)
|
|
246
|
+
return {
|
|
247
|
+
"monthly_gpu_cost": f"${monthly_gpu_cost:.0f}",
|
|
248
|
+
"monthly_api_spend": f"${monthly_api_spend_usd:.0f}",
|
|
249
|
+
"gpu_as_pct_of_api": f"{break_even*100:.0f}%",
|
|
250
|
+
"recommendation": recommendation,
|
|
251
|
+
}
|
|
252
|
+
|
|
253
|
+
# Example: $5k/month on OpenAI, $2.50/hr A10G
|
|
254
|
+
print(break_even_analysis(5000))
|
|
255
|
+
# → gpu_cost ~$1,260/mo = 25% of API spend → self-host now
|
|
256
|
+
```
|
|
257
|
+
|
|
258
|
+
## Cost Dashboard (Grafana)
|
|
259
|
+
|
|
260
|
+
```python
|
|
261
|
+
# Emit cost metrics to Prometheus
|
|
262
|
+
from prometheus_client import Counter, Histogram
|
|
263
|
+
|
|
264
|
+
llm_cost_total = Counter(
|
|
265
|
+
"llm_cost_usd_total",
|
|
266
|
+
"Total LLM spend in USD",
|
|
267
|
+
["model", "team", "task_type"],
|
|
268
|
+
)
|
|
269
|
+
llm_tokens_total = Counter(
|
|
270
|
+
"llm_tokens_total",
|
|
271
|
+
"Total tokens used",
|
|
272
|
+
["model", "token_type"], # token_type: prompt, completion, cached
|
|
273
|
+
)
|
|
274
|
+
|
|
275
|
+
def track_call(model, team, task_type, response):
|
|
276
|
+
cost = calculate_cost(model, response.usage)
|
|
277
|
+
llm_cost_total.labels(model=model, team=team, task_type=task_type).inc(cost)
|
|
278
|
+
llm_tokens_total.labels(model=model, token_type="prompt").inc(
|
|
279
|
+
response.usage.prompt_tokens)
|
|
280
|
+
llm_tokens_total.labels(model=model, token_type="completion").inc(
|
|
281
|
+
response.usage.completion_tokens)
|
|
282
|
+
```
|
|
283
|
+
|
|
284
|
+
## Best Practices
|
|
285
|
+
|
|
286
|
+
- Use `gpt-4o-mini` or `claude-haiku` for 80% of tasks — they're 10–30× cheaper.
|
|
287
|
+
- Enable prompt caching for system prompts >1,024 tokens (Anthropic) or >1,024 tokens (OpenAI).
|
|
288
|
+
- Audit your top 5 prompts by token count — compress or cache them.
|
|
289
|
+
- Set hard budget limits with LiteLLM virtual keys before costs spiral.
|
|
290
|
+
- Self-host 7B–8B models when monthly API spend exceeds $2k/month.
|
|
291
|
+
|
|
292
|
+
## Related Skills
|
|
293
|
+
|
|
294
|
+
- llm-gateway (`llm-gateway`) - Centralized cost control
|
|
295
|
+
- llm-caching (`llm-caching`) - Semantic caching patterns
|
|
296
|
+
- vllm-server (`vllm-server`) - Self-hosted inference
|
|
297
|
+
- agent-observability (`agent-observability`) - Token and cost telemetry
|
|
298
|
+
|
|
299
|
+
## Limitations
|
|
300
|
+
|
|
301
|
+
- Guidance executes against real environments: confirm target, blast radius, and rollback plan before applying anything.
|
|
302
|
+
- Never deploy to production without explicit approval. Docs-only import: upstream scripts and templates not bundled.
|
|
303
|
+
|
|
304
|
+
### Example
|
|
305
|
+
|
|
306
|
+
```bash
|
|
307
|
+
git status && git diff --stat
|
|
308
|
+
kubectl diff -f manifest.yaml
|
|
309
|
+
```
|
|
310
|
+
|
|
311
|
+
> Adapted from [BagelHole/DevOps-Security-Agent-Skills](https://github.com/BagelHole/DevOps-Security-Agent-Skills) (MIT); frontmatter, When to Use/Limitations, and safety boundaries added for upstream compliance. Docs-only import: helper scripts and templates not bundled.
|