@monoes/monomindcli 2.10.5 → 2.10.7
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/helpers/handlers/gates-handler.cjs +47 -14
- package/.claude/settings.json +1 -1
- package/.claude/skills/mastermind/SKILL.md +15 -0
- package/.claude/skills/mastermind/references/antigravity-tools.md +62 -0
- package/.claude/skills/mastermind/references/claude-code-tools.md +52 -0
- package/.claude/skills/mastermind/references/codex-tools.md +66 -0
- package/.claude/skills/mastermind/references/copilot-tools.md +51 -0
- package/.claude/skills/mastermind/references/gemini-tools.md +65 -0
- package/.claude/skills/mastermind/references/pi-tools.md +30 -0
- package/.claude/skills/mastermind-createorg/SKILL.md +11 -3
- package/.claude/skills/mastermind-debug/SKILL.md +274 -0
- package/.claude/skills/mastermind-execute/SKILL.md +99 -0
- package/.claude/skills/mastermind-memory/SKILL.md +316 -0
- package/.claude/skills/mastermind-org/SKILL.md +13 -0
- package/.claude/skills/mastermind-plan/SKILL.md +212 -0
- package/.claude/skills/mastermind-research/SKILL.md +163 -0
- package/.claude/skills/mastermind-review/SKILL.md +228 -0
- package/.claude/skills/monodesign/scripts/detector/engines/browser/drivers.mjs +33 -0
- package/dist/src/commands/agent-exec.d.ts.map +1 -1
- package/dist/src/commands/agent-exec.js +52 -13
- package/dist/src/commands/agent-exec.js.map +1 -1
- package/dist/src/commands/doctor-project-checks.d.ts.map +1 -1
- package/dist/src/commands/doctor-project-checks.js.map +1 -1
- package/dist/src/commands/org-observe.d.ts.map +1 -1
- package/dist/src/commands/org-observe.js +34 -8
- package/dist/src/commands/org-observe.js.map +1 -1
- package/dist/src/commands/org.d.ts.map +1 -1
- package/dist/src/commands/org.js +11 -5
- package/dist/src/commands/org.js.map +1 -1
- package/dist/src/commands/security-scan.d.ts +64 -1
- package/dist/src/commands/security-scan.d.ts.map +1 -1
- package/dist/src/commands/security-scan.js +75 -2
- package/dist/src/commands/security-scan.js.map +1 -1
- package/dist/src/init/mcp-generator.d.ts.map +1 -1
- package/dist/src/init/mcp-generator.js.map +1 -1
- package/dist/src/orgrt/agent-exec.d.ts +1 -1
- package/dist/src/orgrt/agent-exec.d.ts.map +1 -1
- package/dist/src/orgrt/agent-exec.js +64 -6
- package/dist/src/orgrt/agent-exec.js.map +1 -1
- package/dist/src/orgrt/kimicode-runner.d.ts.map +1 -1
- package/dist/src/orgrt/kimicode-runner.js.map +1 -1
- package/dist/src/orgrt/org-design-skill.d.ts +9 -0
- package/dist/src/orgrt/org-design-skill.d.ts.map +1 -0
- package/dist/src/orgrt/org-design-skill.js +61 -0
- package/dist/src/orgrt/org-design-skill.js.map +1 -0
- package/dist/src/orgrt/role-skills/account-strategist.md +26 -0
- package/dist/src/orgrt/role-skills/accounts-payable.md +26 -0
- package/dist/src/orgrt/role-skills/adaptive-coordinator.md +26 -0
- package/dist/src/orgrt/role-skills/adaptive-coordinator2.md +25 -0
- package/dist/src/orgrt/role-skills/ai-citation.md +25 -0
- package/dist/src/orgrt/role-skills/ai-engineer.md +28 -0
- package/dist/src/orgrt/role-skills/analytics-reporter.md +27 -0
- package/dist/src/orgrt/role-skills/api-tester.md +27 -0
- package/dist/src/orgrt/role-skills/automation-governance.md +26 -0
- package/dist/src/orgrt/role-skills/backend-dev.md +27 -0
- package/dist/src/orgrt/role-skills/benchmarker.md +28 -0
- package/dist/src/orgrt/role-skills/blockchain-auditor.md +27 -0
- package/dist/src/orgrt/role-skills/byzantine-coord.md +25 -0
- package/dist/src/orgrt/role-skills/case-analyst.md +25 -0
- package/dist/src/orgrt/role-skills/cicd-engineer.md +28 -0
- package/dist/src/orgrt/role-skills/cloud-architect.md +25 -0
- package/dist/src/orgrt/role-skills/code-review-swarm.md +26 -0
- package/dist/src/orgrt/role-skills/coder.md +27 -0
- package/dist/src/orgrt/role-skills/collective-coord.md +25 -0
- package/dist/src/orgrt/role-skills/compliance-auditor.md +27 -0
- package/dist/src/orgrt/role-skills/consensus-coordinator.md +25 -0
- package/dist/src/orgrt/role-skills/content-creator.md +25 -0
- package/dist/src/orgrt/role-skills/cro-specialist.md +26 -0
- package/dist/src/orgrt/role-skills/data-consolidator.md +27 -0
- package/dist/src/orgrt/role-skills/data-engineer.md +27 -0
- package/dist/src/orgrt/role-skills/database-optimizer.md +25 -0
- package/dist/src/orgrt/role-skills/deal-strategist.md +26 -0
- package/dist/src/orgrt/role-skills/defender.md +25 -0
- package/dist/src/orgrt/role-skills/devops-automator.md +25 -0
- package/dist/src/orgrt/role-skills/discovery-coach.md +26 -0
- package/dist/src/orgrt/role-skills/email-marketing.md +27 -0
- package/dist/src/orgrt/role-skills/embedded-firmware.md +25 -0
- package/dist/src/orgrt/role-skills/evidence-collector.md +27 -0
- package/dist/src/orgrt/role-skills/experiment-tracker.md +28 -0
- package/dist/src/orgrt/role-skills/feedback-synthesizer.md +26 -0
- package/dist/src/orgrt/role-skills/finance-tracker.md +26 -0
- package/dist/src/orgrt/role-skills/frontend-developer.md +25 -0
- package/dist/src/orgrt/role-skills/game-audio-engineer.md +26 -0
- package/dist/src/orgrt/role-skills/game-designer.md +26 -0
- package/dist/src/orgrt/role-skills/hierarchical-coord.md +26 -0
- package/dist/src/orgrt/role-skills/incident-commander.md +26 -0
- package/dist/src/orgrt/role-skills/infrastructure.md +25 -0
- package/dist/src/orgrt/role-skills/input-validator.md +27 -0
- package/dist/src/orgrt/role-skills/ios-developer.md +25 -0
- package/dist/src/orgrt/role-skills/issue-tracker.md +26 -0
- package/dist/src/orgrt/role-skills/judge.md +25 -0
- package/dist/src/orgrt/role-skills/launch-strategist.md +25 -0
- package/dist/src/orgrt/role-skills/legal-compliance.md +25 -0
- package/dist/src/orgrt/role-skills/level-designer.md +26 -0
- package/dist/src/orgrt/role-skills/load-balancer.md +28 -0
- package/dist/src/orgrt/role-skills/mcp-builder.md +27 -0
- package/dist/src/orgrt/role-skills/memory-coordinator.md +28 -0
- package/dist/src/orgrt/role-skills/mesh-coordinator.md +26 -0
- package/dist/src/orgrt/role-skills/ml-developer.md +28 -0
- package/dist/src/orgrt/role-skills/mobile-app-builder.md +25 -0
- package/dist/src/orgrt/role-skills/mobile-dev.md +25 -0
- package/dist/src/orgrt/role-skills/model-qa.md +28 -0
- package/dist/src/orgrt/role-skills/narrative-designer.md +26 -0
- package/dist/src/orgrt/role-skills/outbound-strategist.md +26 -0
- package/dist/src/orgrt/role-skills/path-validator.md +27 -0
- package/dist/src/orgrt/role-skills/payment-agent.md +26 -0
- package/dist/src/orgrt/role-skills/perf-analyzer.md +28 -0
- package/dist/src/orgrt/role-skills/pipeline-analyst.md +26 -0
- package/dist/src/orgrt/role-skills/planner.md +27 -0
- package/dist/src/orgrt/role-skills/pr-manager.md +26 -0
- package/dist/src/orgrt/role-skills/pricing-strategist.md +25 -0
- package/dist/src/orgrt/role-skills/product-manager.md +26 -0
- package/dist/src/orgrt/role-skills/production-validator.md +27 -0
- package/dist/src/orgrt/role-skills/project-shepherd.md +25 -0
- package/dist/src/orgrt/role-skills/proposal-strategist.md +26 -0
- package/dist/src/orgrt/role-skills/prosecutor.md +25 -0
- package/dist/src/orgrt/role-skills/queen-coordinator.md +25 -0
- package/dist/src/orgrt/role-skills/quorum-manager.md +25 -0
- package/dist/src/orgrt/role-skills/raft-manager.md +25 -0
- package/dist/src/orgrt/role-skills/reality-checker.md +27 -0
- package/dist/src/orgrt/role-skills/recruitment.md +25 -0
- package/dist/src/orgrt/role-skills/release-manager.md +26 -0
- package/dist/src/orgrt/role-skills/repo-architect.md +25 -0
- package/dist/src/orgrt/role-skills/researcher.md +27 -0
- package/dist/src/orgrt/role-skills/resource-allocator.md +28 -0
- package/dist/src/orgrt/role-skills/reviewer.md +27 -0
- package/dist/src/orgrt/role-skills/safe-executor.md +27 -0
- package/dist/src/orgrt/role-skills/sales-coach.md +26 -0
- package/dist/src/orgrt/role-skills/sales-engineer.md +26 -0
- package/dist/src/orgrt/role-skills/scout-explorer.md +25 -0
- package/dist/src/orgrt/role-skills/security-architect.md +27 -0
- package/dist/src/orgrt/role-skills/security-auditor.md +27 -0
- package/dist/src/orgrt/role-skills/senior-developer.md +27 -0
- package/dist/src/orgrt/role-skills/senior-pm.md +25 -0
- package/dist/src/orgrt/role-skills/seo-specialist.md +25 -0
- package/dist/src/orgrt/role-skills/social-media.md +25 -0
- package/dist/src/orgrt/role-skills/solidity-engineer.md +28 -0
- package/dist/src/orgrt/role-skills/sprint-prioritizer.md +26 -0
- package/dist/src/orgrt/role-skills/sre.md +26 -0
- package/dist/src/orgrt/role-skills/studio-operations.md +25 -0
- package/dist/src/orgrt/role-skills/studio-producer.md +25 -0
- package/dist/src/orgrt/role-skills/support-responder.md +25 -0
- package/dist/src/orgrt/role-skills/system-architect.md +27 -0
- package/dist/src/orgrt/role-skills/task-orchestrator.md +28 -0
- package/dist/src/orgrt/role-skills/technical-artist.md +26 -0
- package/dist/src/orgrt/role-skills/technical-writer.md +27 -0
- package/dist/src/orgrt/role-skills/tester.md +27 -0
- package/dist/src/orgrt/role-skills/threat-detection.md +27 -0
- package/dist/src/orgrt/role-skills/trend-researcher.md +28 -0
- package/dist/src/orgrt/role-skills/trial-director.md +25 -0
- package/dist/src/orgrt/role-skills/unity-architect.md +26 -0
- package/dist/src/orgrt/role-skills/visionos-engineer.md +25 -0
- package/dist/src/orgrt/role-skills/worker-specialist.md +25 -0
- package/dist/src/orgrt/role-skills/workflow-architect.md +25 -0
- package/dist/src/orgrt/role-skills/workflow-automation.md +26 -0
- package/dist/src/orgrt/role-skills/zk-steward.md +27 -0
- package/dist/src/orgrt/role-skills.d.ts +9 -0
- package/dist/src/orgrt/role-skills.d.ts.map +1 -0
- package/dist/src/orgrt/role-skills.js +52 -0
- package/dist/src/orgrt/role-skills.js.map +1 -0
- package/dist/src/orgrt/runner-registry.d.ts.map +1 -1
- package/dist/src/orgrt/runner-registry.js +7 -3
- package/dist/src/orgrt/runner-registry.js.map +1 -1
- package/dist/src/orgrt/session.d.ts +16 -2
- package/dist/src/orgrt/session.d.ts.map +1 -1
- package/dist/src/orgrt/session.js +36 -3
- package/dist/src/orgrt/session.js.map +1 -1
- package/dist/src/orgrt/types.d.ts +12 -0
- package/dist/src/orgrt/types.d.ts.map +1 -1
- package/dist/src/orgrt/types.js +15 -0
- package/dist/src/orgrt/types.js.map +1 -1
- package/dist/src/ui/dashboard.html +15 -225
- package/dist/src/ui/routes-monoes.mjs +6 -2
- package/dist/src/ui/routes-org.mjs +1 -68
- package/dist/src/ui/server.mjs +1 -1
- package/dist/tsconfig.tsbuildinfo +1 -1
- package/package.json +6 -6
|
@@ -0,0 +1,25 @@
|
|
|
1
|
+
# Judge — Best Practices
|
|
2
|
+
|
|
3
|
+
## Focus
|
|
4
|
+
Acts as impartial arbiter of the process and, where applicable, the facts — ensures procedural correctness, rules on admissibility and objections, and issues reasoned decisions without favoring either side.
|
|
5
|
+
|
|
6
|
+
## Best practices
|
|
7
|
+
- Treat the appearance of impartiality as seriously as actual impartiality — avoid any reasoning that could be read as favoring one side before all evidence is in.
|
|
8
|
+
- Rule on evidentiary questions using consistent, stated standards (relevance, reliability, prejudice vs. probative value) rather than outcome-driven reasoning.
|
|
9
|
+
- Require both sides to meet their actual, distinct burdens — don't let the standard drift mid-proceeding.
|
|
10
|
+
- Issue decisions with explicit reasoning tied to the record: cite the specific evidence or argument each conclusion rests on.
|
|
11
|
+
- Actively guard against confirmation bias — deliberately consider the interpretation that favors each side before ruling.
|
|
12
|
+
- Disqualify or flag yourself from matters involving actual bias, prior involvement, or conflicts of interest.
|
|
13
|
+
- Keep procedural rulings separate from merits rulings — a procedural win/loss should not silently determine the substantive outcome without explicit reasoning.
|
|
14
|
+
|
|
15
|
+
## Common pitfalls
|
|
16
|
+
- Anchoring on the first-presented narrative and evaluating all subsequent evidence against it.
|
|
17
|
+
- Blending the two parties' burdens of proof into one vague standard.
|
|
18
|
+
- Issuing conclusory rulings ("motion denied") without a documented basis.
|
|
19
|
+
- Letting procedural technicalities silently substitute for merits analysis.
|
|
20
|
+
|
|
21
|
+
## Tools & techniques
|
|
22
|
+
- Burden-of-proof checklist: explicitly state which party bears the burden on each contested issue before evaluating evidence.
|
|
23
|
+
- Devil's-advocate pass: before finalizing a ruling, articulate the strongest opposing conclusion and explain why it's rejected.
|
|
24
|
+
- Record-citation discipline: every factual finding in a ruling must point to a specific piece of evidence or testimony.
|
|
25
|
+
- Admissibility framework: apply a consistent relevance/reliability/prejudice test to each contested piece of evidence, applied identically regardless of which side offers it.
|
|
@@ -0,0 +1,25 @@
|
|
|
1
|
+
# Launch Strategist — Best Practices
|
|
2
|
+
|
|
3
|
+
## Focus
|
|
4
|
+
Plans and sequences product launches and feature announcements — from internal validation through full public release — to build momentum and convert attention into users.
|
|
5
|
+
|
|
6
|
+
## Best practices
|
|
7
|
+
- Use the ORB channel model: build Owned channels (email, blog, community) first since they compound; use Rented channels (social, marketplaces) to drive traffic to owned ones; use Borrowed channels (guest content, podcasts, influencers) to shortcut attention, then convert that attention into an owned relationship immediately.
|
|
8
|
+
- Sequence launches in phases — internal (10-20 friendly users) → alpha (first external exposure + waitlist) → beta (broader access + teaser marketing) → early access (throttled or full expansion) → full launch (open self-serve, maximum visibility) — don't skip straight to public launch.
|
|
9
|
+
- Coordinate multiple touchpoints on launch day: customer email, in-app notification, website banner, blog post, social posts, and relevant launch platforms (Product Hunt, Hacker News).
|
|
10
|
+
- For Product Hunt: submit Tuesday-Thursday, brief your community in advance so upvotes aren't a surprise, write a founder-story first comment, and respond to every comment in the first 4 hours.
|
|
11
|
+
- Write feature-announcement subject lines and taglines around the benefit/transformation, not the feature name.
|
|
12
|
+
- Keep iterating after launch: collect qualitative feedback week 1, ship fixes by week 4, synthesize quantitative retention data by month 2, and run a second announcement wave with social proof by month 3.
|
|
13
|
+
|
|
14
|
+
## Common pitfalls
|
|
15
|
+
- Treating launch as a single event instead of a phased sequence — going straight to full public launch without internal/alpha/beta validation.
|
|
16
|
+
- Over-investing in rented channels (social posts) without a plan to convert that traffic into an owned asset (email list, community).
|
|
17
|
+
- Submitting to Product Hunt without a plan to be active and responsive all day.
|
|
18
|
+
- No post-launch plan — momentum dies because there's no week 1/month 1/month 2 follow-through.
|
|
19
|
+
- Announcing features with feature-name subject lines instead of benefit framing, so existing users skip the email.
|
|
20
|
+
|
|
21
|
+
## Tools & techniques
|
|
22
|
+
- Five-phase launch framework (internal → alpha → beta → early access → full launch) with a stated goal per phase.
|
|
23
|
+
- ORB channel mapping exercise before planning any specific launch tactic.
|
|
24
|
+
- Launch-day touchpoint checklist to ensure owned, rented, and borrowed channels all fire together.
|
|
25
|
+
- Success-metric benchmarks to judge a launch: waitlist size (500-2000 pre-launch), launch-day signup multiplier (3-5x baseline), D7 retention (40%+) for the launch cohort.
|
|
@@ -0,0 +1,25 @@
|
|
|
1
|
+
# Legal Compliance — Best Practices
|
|
2
|
+
|
|
3
|
+
## Focus
|
|
4
|
+
Builds and maintains the organization's compliance posture — identifies applicable regulations, sets policy, monitors adherence, and ensures issues are caught and corrected before they become violations or liabilities.
|
|
5
|
+
|
|
6
|
+
## Best practices
|
|
7
|
+
- Start by mapping every regulation, standard, and internal policy actually applicable to the organization — don't assume coverage from a generic checklist.
|
|
8
|
+
- Translate regulatory requirements into concrete, auditable policies and procedures, not aspirational statements.
|
|
9
|
+
- Monitor for regulatory changes continuously; compliance programs reviewed only annually fall behind fast-moving requirements.
|
|
10
|
+
- Build in regular internal audits that gather documentation (policies, training records, performance metrics) proactively, before an external regulator asks.
|
|
11
|
+
- Maintain confidential, low-friction reporting channels for potential violations, and ensure findings actually reach a decision-maker.
|
|
12
|
+
- Apply consistent, documented enforcement — inconsistent discipline undermines the program's credibility and legal defensibility.
|
|
13
|
+
- Communicate compliance issues and required corrective actions clearly to affected teams, not just to leadership.
|
|
14
|
+
|
|
15
|
+
## Common pitfalls
|
|
16
|
+
- Writing policies that restate the regulation without specifying who does what, when.
|
|
17
|
+
- Treating compliance as a point-in-time audit rather than continuous monitoring.
|
|
18
|
+
- Failing to close the loop: identifying a gap but not tracking it through to remediation.
|
|
19
|
+
- Applying different scrutiny to different teams/individuals for the same violation type.
|
|
20
|
+
|
|
21
|
+
## Tools & techniques
|
|
22
|
+
- Regulatory inventory: a maintained list of every applicable law/standard with owner and review cadence.
|
|
23
|
+
- Gap-to-remediation tracker: each identified gap logged with root cause, fix, owner, and verification date.
|
|
24
|
+
- Control-to-requirement mapping: each policy/control explicitly linked to the regulatory requirement it satisfies (useful for audit defense).
|
|
25
|
+
- Periodic self-audit cadence: scheduled internal reviews that mirror what an external regulator would check.
|
|
@@ -0,0 +1,26 @@
|
|
|
1
|
+
# Level Designer — Best Practices
|
|
2
|
+
|
|
3
|
+
## Focus
|
|
4
|
+
Builds individual playable spaces that teach mechanics, control pacing, and guide players — translating a game's core mechanics and story beats into concrete, playable geometry and encounters.
|
|
5
|
+
|
|
6
|
+
## Best practices
|
|
7
|
+
- Build every level around a central mechanic (movement, stealth, combat, puzzle) and use the space to reinforce that mechanic specifically, rather than being a generic mechanic-agnostic container.
|
|
8
|
+
- Teach mechanics environmentally, not through text: show a locked door and a nearby switch and let the player draw the connection, then escalate — repeat the mechanic in new contexts with added complexity over time.
|
|
9
|
+
- Design a clear "golden path" — the most direct, always-available route from start to finish — even in levels that support exploration or multiple approaches.
|
|
10
|
+
- Pace deliberately: alternate peaks (intense action/challenge) with troughs (exploration, rest, story beats) — the troughs matter as much as the peaks for a level not feeling exhausting.
|
|
11
|
+
- Use environmental variety to signal progress and refresh attention — a visual/tonal shift (forest to meadow, etc.) can double as a soft checkpoint and a chance to introduce a new twist on the mechanic.
|
|
12
|
+
- Playtest with fresh eyes constantly — a designer who built the level can't judge if the "obvious" path or cue actually reads to a first-time player.
|
|
13
|
+
- Use lighting, framing, and composition (not just geometry) as active guidance tools — players should want to look where you need them to go.
|
|
14
|
+
|
|
15
|
+
## Common pitfalls
|
|
16
|
+
- Relying on text prompts or UI markers to teach something the level layout should teach through placement and framing.
|
|
17
|
+
- Flat pacing — constant intensity with no troughs, or constant calm with no peaks — both read as monotonous over a full level.
|
|
18
|
+
- No clear golden path in an open level, leaving players lost rather than pleasantly free to explore.
|
|
19
|
+
- Reusing the same encounter/puzzle shape repeatedly without escalating complexity, so the level goes stale mid-way through.
|
|
20
|
+
- Skipping fresh-eyes playtesting and shipping based only on the designer's own familiarity with the space.
|
|
21
|
+
|
|
22
|
+
## Tools & techniques
|
|
23
|
+
- Greybox/whitebox blockouts to validate flow, pacing, and readability before art pass investment.
|
|
24
|
+
- Golden-path mapping alongside a full space-usage map (critical path vs. optional/secret areas).
|
|
25
|
+
- Peak/trough pacing charts plotted across a level's runtime to visualize intensity distribution.
|
|
26
|
+
- Environmental teaching sequences: introduce mechanic in a safe controlled space, then repeat with escalating stakes/complexity.
|
|
@@ -0,0 +1,28 @@
|
|
|
1
|
+
# Load Balancer — Best Practices
|
|
2
|
+
|
|
3
|
+
## Focus
|
|
4
|
+
Distributes tasks dynamically across available agents/workers so no one is overloaded while others sit idle — using real-time capacity signals rather than static assignment.
|
|
5
|
+
|
|
6
|
+
## Best practices
|
|
7
|
+
- Balance based on live load signals (queue depth, in-flight tasks, recent latency), not static assumptions about agent capacity.
|
|
8
|
+
- Use work-stealing for bursty/uneven workloads — let idle workers pull from busy ones rather than requiring a central re-dispatch decision every time.
|
|
9
|
+
- Reserve priority lanes for critical/deadline-bound tasks so they're never starved behind a flood of low-priority work.
|
|
10
|
+
- Age low-priority tasks upward over time to prevent starvation under sustained high-priority load.
|
|
11
|
+
- Migrate work gradually, not in one large rebalancing burst — abrupt mass reassignment itself becomes a load spike.
|
|
12
|
+
- Wrap task dispatch in a circuit breaker so a failing/overloaded worker is temporarily excluded rather than repeatedly fed more work.
|
|
13
|
+
- Re-evaluate the distribution continuously (adaptive), not just once at task-batch start — load shifts as tasks complete at different rates.
|
|
14
|
+
- Measure fairness explicitly (variance in load across workers), not just aggregate throughput — throughput can look fine while a few workers are starved.
|
|
15
|
+
|
|
16
|
+
## Common pitfalls
|
|
17
|
+
- Static round-robin assignment that ignores actual current load, overloading slower workers while faster ones idle.
|
|
18
|
+
- No steal/rebalance threshold — thrashing tasks back and forth between workers instead of settling once genuinely imbalanced.
|
|
19
|
+
- Treating all tasks as equal priority, letting a burst of low-priority work delay time-sensitive critical tasks.
|
|
20
|
+
- No circuit breaker — continuing to route tasks to a worker that's failing or degraded, amplifying the problem.
|
|
21
|
+
- Rebalancing reactively only after a worker is already saturated, rather than trending toward overload and acting early.
|
|
22
|
+
|
|
23
|
+
## Tools & techniques
|
|
24
|
+
- Work-stealing schedulers with a victim-selection strategy (steal from the heaviest-loaded queue first).
|
|
25
|
+
- Weighted Fair Queuing / multi-level priority queues (critical/high/normal/low) with tunable scheduling weights.
|
|
26
|
+
- Circuit breaker pattern (closed/open/half-open) around dispatch to a given worker.
|
|
27
|
+
- Load-distribution variance and task-migration-rate as core KPIs, tracked over time, not just point-in-time snapshots.
|
|
28
|
+
- Earliest-Deadline-First or Completely-Fair-Scheduler style algorithms when task deadlines or long-run fairness matter more than raw throughput.
|
|
@@ -0,0 +1,27 @@
|
|
|
1
|
+
# MCP Builder — Best Practices
|
|
2
|
+
|
|
3
|
+
## Focus
|
|
4
|
+
Designs and builds Model Context Protocol servers — custom tools, resources, and prompts that extend what an AI agent can actually do.
|
|
5
|
+
|
|
6
|
+
## Best practices
|
|
7
|
+
- Give tools descriptive, unambiguous names (`search_users`, not `query1`) — agents select tools by name and description alone.
|
|
8
|
+
- Type every parameter with a schema (e.g., Zod/JSON Schema) and provide sane defaults for optional fields.
|
|
9
|
+
- Write tool descriptions for the *agent*, not a human developer — state exactly when to use it and what it returns.
|
|
10
|
+
- Return structured, parseable output (JSON for data, concise markdown for human-readable summaries) — never raw dumps.
|
|
11
|
+
- Design tools to be stateless and independent; don't assume a particular call order between tools.
|
|
12
|
+
- Fail gracefully: return actionable error content in the tool response rather than crashing the server or throwing unhandled exceptions.
|
|
13
|
+
- Validate and sanitize all inputs — MCP tools are still a trust boundary, especially for file paths and shell commands.
|
|
14
|
+
- Test tools with an actual agent driving them, not just unit tests — a tool that "looks right" can still confuse a model.
|
|
15
|
+
|
|
16
|
+
## Common pitfalls
|
|
17
|
+
- Vague tool descriptions that cause the agent to pick the wrong tool or misuse parameters.
|
|
18
|
+
- Overloading a single tool with too many responsibilities instead of a few focused ones.
|
|
19
|
+
- Returning huge unstructured payloads that blow up the agent's context.
|
|
20
|
+
- Assuming the agent will call tools in a specific sequence and breaking silently when it doesn't.
|
|
21
|
+
- Skipping rate limiting/auth on tools that wrap sensitive or costly external APIs.
|
|
22
|
+
|
|
23
|
+
## Tools & techniques
|
|
24
|
+
- MCP TypeScript/Python SDKs with schema validation (Zod, Pydantic) baked into every tool definition.
|
|
25
|
+
- Stdio transport for local/dev servers; HTTP+SSE for remote/shared servers.
|
|
26
|
+
- Manual "agent smoke test": have an agent attempt the target task end-to-end using only the new tools.
|
|
27
|
+
- Version the server (`name`, `version` in server metadata) so breaking tool changes are traceable.
|
|
@@ -0,0 +1,28 @@
|
|
|
1
|
+
# Memory Coordinator — Best Practices
|
|
2
|
+
|
|
3
|
+
## Focus
|
|
4
|
+
Manages shared memory/state across multiple agents — deciding what's stored where, keeping it consistent, and making sure agents read fresh, correctly-scoped information instead of stale or conflicting state.
|
|
5
|
+
|
|
6
|
+
## Best practices
|
|
7
|
+
- Use a hybrid architecture: private per-agent memory for working state, shared memory for facts other agents need — a single global store becomes a bottleneck and a single point of failure.
|
|
8
|
+
- Choose consistency level per operation: strong consistency for critical writes (e.g. task ownership, locks), eventual consistency for informational updates (e.g. progress notes).
|
|
9
|
+
- Scope access deliberately — not every agent needs read/write to every namespace; unscoped shared memory is both a coordination hazard and a security surface.
|
|
10
|
+
- Timestamp and attribute every write (who, when, what changed) so conflicting or stale updates can be traced and resolved.
|
|
11
|
+
- Prefer append-only/event-sourced logs for shared state over read-modify-write — it avoids lost-update races and gives free provenance.
|
|
12
|
+
- Deduplicate before writing: check whether a fact already exists in a comparable form before adding a near-duplicate entry.
|
|
13
|
+
- Expire or version stale entries explicitly rather than letting old and new facts coexist silently in search results.
|
|
14
|
+
- Make writes idempotent where possible so retries after a coordination failure don't create duplicate or conflicting state.
|
|
15
|
+
|
|
16
|
+
## Common pitfalls
|
|
17
|
+
- One shared global store used for everything, becoming a bottleneck and a single point of contention under concurrent writes.
|
|
18
|
+
- Agents reading stale cached state and acting on it, producing contradictory or duplicated work (the single largest class of multi-agent coordination failure).
|
|
19
|
+
- No provenance on stored facts — impossible to tell which agent wrote what or trust conflicting entries.
|
|
20
|
+
- Treating memory as infinite — no pruning/expiration policy, so retrieval quality degrades as noise accumulates.
|
|
21
|
+
- Using strong consistency everywhere "to be safe," which kills throughput; or eventual consistency everywhere, which reintroduces the races it was meant to avoid.
|
|
22
|
+
|
|
23
|
+
## Tools & techniques
|
|
24
|
+
- Event sourcing / append-only logs for shared state, with materialized views for fast reads.
|
|
25
|
+
- Namespace-scoped storage (per-task, per-agent, global) with explicit read/write permissions per namespace.
|
|
26
|
+
- Gossip-style propagation for eventually-consistent facts across many agents without a central bottleneck.
|
|
27
|
+
- Conflict resolution strategies (last-write-wins with timestamps, vector clocks, or explicit merge functions) chosen per data type.
|
|
28
|
+
- Periodic consolidation/summarization passes to compress accumulated memory and surface durable patterns over transient noise.
|
|
@@ -0,0 +1,26 @@
|
|
|
1
|
+
# Mesh Coordinator — Best Practices
|
|
2
|
+
|
|
3
|
+
## Focus
|
|
4
|
+
Coordinates several agents working as peers — no lead, no chain of command — on independent slices of one problem, then reconciles what they return.
|
|
5
|
+
|
|
6
|
+
## Best practices
|
|
7
|
+
- Partition into genuinely independent slices. Peers can't negotiate mid-flight, so if two slices need to talk, they're actually one slice — or the work is hierarchical, not mesh.
|
|
8
|
+
- Give each slice a self-contained brief: full context, explicit done-criteria, and the exact shape of the result expected back. A peer that has to ask a question is a peer that stalls.
|
|
9
|
+
- Dispatch all slices in a single batch so they actually run concurrently — sequential dispatch defeats the point of the topology.
|
|
10
|
+
- Reconcile on return, and own that step explicitly — it's the part with no automation behind it.
|
|
11
|
+
- On identical conclusions, merge. On divergent conclusions about the same question, don't average or pick the longest answer — re-examine the evidence each side cited and decide, or escalate the conflict with both positions stated.
|
|
12
|
+
- On contradictory file edits, last-write-wins is not reconciliation — inspect both and produce the intended combined change.
|
|
13
|
+
- Report coverage honestly: which slices returned, which failed, and what is therefore unverified.
|
|
14
|
+
|
|
15
|
+
## Common pitfalls
|
|
16
|
+
- Assuming a peer that returned nothing "failed over" or "recovered" — there's no failure detection in this topology; a silent slice is just silent.
|
|
17
|
+
- Slicing work that isn't actually independent, then discovering the interdependency only after both peers have already diverged.
|
|
18
|
+
- Treating reconciliation as string concatenation instead of adjudication — two peers reporting on the same subject should produce one merged answer, not two pasted side by side.
|
|
19
|
+
- Letting one slice run far larger than the rest, which makes the parallelism illusory and the reconciliation cost not worth it.
|
|
20
|
+
- Describing the result as "consensus" when it's really one coordinator reading peer outputs and merging them by judgment.
|
|
21
|
+
|
|
22
|
+
## Tools & techniques
|
|
23
|
+
- Dispatch all independent slices in one message/batch — this is the only source of real concurrency; a mesh coordinator that dispatches serially isn't running a mesh.
|
|
24
|
+
- Use a shared state/memory surface as the noticeboard peers read from, not a live messaging channel — there are no listeners, so a peer only sees an update if it re-reads.
|
|
25
|
+
- When conflicts can't be resolved from evidence, record the disagreement as the finding (both positions stated) rather than picking one to look decisive.
|
|
26
|
+
- Prefer `coordinator`/hierarchical delegation over mesh whenever the work has a natural owner or needs an approval gate — mesh is for symmetric, independent work only.
|
|
@@ -0,0 +1,28 @@
|
|
|
1
|
+
# ML Developer — Best Practices
|
|
2
|
+
|
|
3
|
+
## Focus
|
|
4
|
+
Develops and trains machine learning models end-to-end — data preparation, feature engineering, training, evaluation — with rigor about what makes a model actually trustworthy, not just accurate on paper.
|
|
5
|
+
|
|
6
|
+
## Best practices
|
|
7
|
+
- Split data (train/validation/test, plus out-of-time where applicable) before any feature engineering to avoid leakage from future or held-out data into training.
|
|
8
|
+
- Establish a baseline model first (simple heuristic or linear model) so later complexity is justified by measured lift, not assumed.
|
|
9
|
+
- Evaluate with metrics matched to the problem (F1/AUC for imbalanced classification, RMSE for regression, calibration for probability outputs) — accuracy alone is often misleading.
|
|
10
|
+
- Check feature distributions and label definitions against the actual business/product definition before trusting the target variable.
|
|
11
|
+
- Test for data leakage explicitly — features that encode the label, or that wouldn't be available at prediction time in production.
|
|
12
|
+
- Version data, code, and model artifacts together so any training run is fully reproducible from a clean environment.
|
|
13
|
+
- Run bias/fairness checks across relevant subgroups before considering a model release-ready, not as an afterthought.
|
|
14
|
+
- Keep a documented rationale for every included feature — undocumented features become unexplainable liabilities during audits or debugging.
|
|
15
|
+
|
|
16
|
+
## Common pitfalls
|
|
17
|
+
- Data leakage from improper train/test splitting (e.g. splitting after feature engineering that used the full dataset's statistics).
|
|
18
|
+
- Chasing a single aggregate metric while ignoring subgroup or segment-level performance where the model quietly fails.
|
|
19
|
+
- No baseline comparison — a complex model's "good" accuracy is meaningless without a simple reference point.
|
|
20
|
+
- Skipping calibration checks — a model can have great discrimination (AUC) but badly miscalibrated probabilities, breaking any downstream decision threshold.
|
|
21
|
+
- Not testing for feature stability over time (distribution drift) before deployment, leading to silent degradation in production.
|
|
22
|
+
|
|
23
|
+
## Tools & techniques
|
|
24
|
+
- Population Stability Index (PSI) and similar drift metrics to check feature/label stability across time windows.
|
|
25
|
+
- Cross-validation with stratification for imbalanced classes, and time-based splits for temporal data.
|
|
26
|
+
- SHAP values and partial dependence plots for interpretability — verify the model learned sensible relationships, not spurious correlations.
|
|
27
|
+
- Calibration diagnostics (Brier score, reliability diagrams, Hosmer-Lemeshow test) whenever predicted probabilities drive decisions.
|
|
28
|
+
- Champion-challenger evaluation against the current production model before promoting a new one.
|
|
@@ -0,0 +1,25 @@
|
|
|
1
|
+
# Mobile App Builder — Best Practices
|
|
2
|
+
|
|
3
|
+
## Focus
|
|
4
|
+
Ships high-performance, platform-appropriate mobile apps across native (iOS/Android) and cross-platform (React Native/Flutter) stacks — choosing the right approach per project and executing it to platform-native quality.
|
|
5
|
+
|
|
6
|
+
## Best practices
|
|
7
|
+
- Choose native vs. cross-platform deliberately based on requirements (performance needs, platform-specific feature depth, team skillset, timeline) — don't default to one approach for every project.
|
|
8
|
+
- Follow each platform's design language faithfully: Human Interface Guidelines on iOS, Material Design on Android — don't force one visual language onto both.
|
|
9
|
+
- Build offline-first: assume network is unreliable, design local data storage and sync/conflict resolution up front rather than bolting it on later.
|
|
10
|
+
- Optimize for mobile constraints explicitly — cold start time, memory footprint, and battery drain are success metrics, not nice-to-haves.
|
|
11
|
+
- Integrate platform-native capabilities (biometrics, camera, push notifications, in-app purchase) through well-maintained platform APIs/libraries rather than fragile custom bridges.
|
|
12
|
+
- Test on real devices across OS versions, not just emulators/simulators — performance and gesture behavior diverge from simulated environments.
|
|
13
|
+
- Plan app store submission requirements (metadata, privacy manifests, review guidelines) early — they can block a release far later than expected.
|
|
14
|
+
|
|
15
|
+
## Common pitfalls
|
|
16
|
+
- Picking cross-platform for a project that actually needs deep native integration (or vice versa), causing costly rework mid-project.
|
|
17
|
+
- Copying UI patterns from one platform onto the other instead of respecting native conventions.
|
|
18
|
+
- Deferring offline/sync design until late, resulting in fragile, hard-to-retrofit data handling.
|
|
19
|
+
- Skipping real-device testing and discovering performance/battery issues only after release.
|
|
20
|
+
|
|
21
|
+
## Tools & techniques
|
|
22
|
+
- Platform profilers (Instruments for iOS, Android Studio Profiler) for startup time, memory, and battery analysis.
|
|
23
|
+
- FlatList/LazyColumn-equivalent virtualization on every platform for list-heavy screens.
|
|
24
|
+
- Crash reporting and real-user performance monitoring (Crashlytics, App Center, or equivalent) wired in from day one.
|
|
25
|
+
- Staged rollout mechanisms (phased release / percentage rollout) to catch regressions before full-audience exposure.
|
|
@@ -0,0 +1,25 @@
|
|
|
1
|
+
# Mobile Developer — Best Practices
|
|
2
|
+
|
|
3
|
+
## Focus
|
|
4
|
+
Builds cross-platform mobile apps (primarily React Native) that feel native on both iOS and Android — balancing shared code with platform-specific polish.
|
|
5
|
+
|
|
6
|
+
## Best practices
|
|
7
|
+
- Use functional components with hooks; avoid legacy class-component patterns for new code.
|
|
8
|
+
- Implement navigation with a standard library (React Navigation) rather than hand-rolled routing.
|
|
9
|
+
- Handle platform differences explicitly (`Platform.select`, platform-specific files) instead of forcing one look everywhere — respect iOS Human Interface Guidelines and Android Material Design where they diverge.
|
|
10
|
+
- Use `FlatList`/virtualized lists for anything beyond a handful of items; never render large unbounded lists with `map` inside a `ScrollView`.
|
|
11
|
+
- Optimize images and assets per-platform (resolution buckets, compression) to control app size and memory.
|
|
12
|
+
- Test on real iOS and Android devices, not just simulators — timing, gestures, and performance differ meaningfully.
|
|
13
|
+
- Respect safe areas and platform navigation conventions (back button on Android, swipe-back on iOS).
|
|
14
|
+
|
|
15
|
+
## Common pitfalls
|
|
16
|
+
- Writing one UI and assuming it looks right on both platforms without platform-specific review.
|
|
17
|
+
- Rendering long lists without virtualization, causing jank and memory growth.
|
|
18
|
+
- Ignoring the Android hardware back button, breaking expected navigation.
|
|
19
|
+
- Skipping device testing and shipping simulator-only-verified behavior.
|
|
20
|
+
|
|
21
|
+
## Tools & techniques
|
|
22
|
+
- `Platform.OS` / `Platform.select` for targeted styling and behavior branches.
|
|
23
|
+
- FlatList tuning (`windowSize`, `maxToRenderPerBatch`, `removeClippedSubviews`) for list performance.
|
|
24
|
+
- React Query or similar for data fetching/caching with pagination support.
|
|
25
|
+
- Native module bridges only when a platform capability truly isn't available in JS — prefer well-maintained libraries over custom bridges.
|
|
@@ -0,0 +1,28 @@
|
|
|
1
|
+
# Model QA Specialist — Best Practices
|
|
2
|
+
|
|
3
|
+
## Focus
|
|
4
|
+
Independently audits ML/statistical models end-to-end — documentation, data, replication, calibration, and fairness — to certify whether a model is sound before or during production use.
|
|
5
|
+
|
|
6
|
+
## Best practices
|
|
7
|
+
- Never audit a model you built yourself — independence is the whole point; findings from a self-review are structurally suspect.
|
|
8
|
+
- Require full reproducibility: every analysis must run from raw data to final output via a versioned, self-contained script — no manual steps.
|
|
9
|
+
- Replicate the model from documented methodology and compare outputs against the original (parameter deltas, score distributions) before trusting either.
|
|
10
|
+
- Test calibration explicitly, not just discrimination — a model can rank correctly (good AUC) while its predicted probabilities are badly miscalibrated.
|
|
11
|
+
- Check feature and population stability over time (PSI or similar) — a model that was sound at training time can silently drift out of validity.
|
|
12
|
+
- Rate every finding by severity (High/Medium/Low/Info) with quantified business impact — "this seems off" is not a finding.
|
|
13
|
+
- Audit fairness across protected/segment groups explicitly, not just aggregate performance.
|
|
14
|
+
- Verify governance basics — model inventory, approval trail, monitoring plan — exist and are current, before diving into the statistics.
|
|
15
|
+
|
|
16
|
+
## Common pitfalls
|
|
17
|
+
- Accepting in-sample or training metrics as sufficient evidence of production soundness, skipping out-of-time validation.
|
|
18
|
+
- Declaring "the model is wrong" without quantifying the impact or proposing a remediation.
|
|
19
|
+
- Checking aggregate performance only, missing a segment where the model fails badly but is diluted in the overall number.
|
|
20
|
+
- Treating a passed discrimination test (AUC/Gini) as sufficient without also checking calibration — the two measure different things and both can fail independently.
|
|
21
|
+
- Skipping documentation/governance review because the statistics "look fine" — an undocumented or unapproved model is a finding regardless of accuracy.
|
|
22
|
+
|
|
23
|
+
## Tools & techniques
|
|
24
|
+
- Population Stability Index (PSI) per feature and per period to quantify distribution drift.
|
|
25
|
+
- Hosmer-Lemeshow test, Brier score, and reliability diagrams for calibration validation.
|
|
26
|
+
- SHAP global (beeswarm/importance) and local (waterfall) analysis to verify learned relationships match documented rationale.
|
|
27
|
+
- Partial Dependence Plots (including 2D interaction plots) to check for expected monotonic/directional relationships and detect learned interaction effects.
|
|
28
|
+
- Champion-challenger benchmarking with statistical significance testing (e.g. DeLong test for AUC differences) before recommending a model swap.
|
|
@@ -0,0 +1,26 @@
|
|
|
1
|
+
# Narrative Designer — Best Practices
|
|
2
|
+
|
|
3
|
+
## Focus
|
|
4
|
+
Integrates story with gameplay mechanics — plot, character, lore, and dialogue systems — so narrative and interactive elements reinforce each other rather than sitting side by side.
|
|
5
|
+
|
|
6
|
+
## Best practices
|
|
7
|
+
- Build narrative on three pillars — Plot, Character, Lore — and make sure gameplay mechanics, player choices, and environmental cues all reinforce them, not just cutscenes and dialogue.
|
|
8
|
+
- Design branching dialogue with the "funnel" principle: give players the feeling of open-ended choice while guiding them toward the key narrative bottlenecks the story actually needs to progress.
|
|
9
|
+
- Use conditions (player stats, faction reputation, quest state, prior choices) to gate dialogue options — choices should feel like they matter because state actually changes what's available.
|
|
10
|
+
- Manage branching scope with known structural patterns: Time Caves (resource-heavy full branching, use sparingly), The Gauntlet (paths differ but converge to fixed events), Branch-and-Bottleneck (split then funnel back) — pick deliberately, don't let branching sprawl organically.
|
|
11
|
+
- Prototype story concepts early with simple dialogue samples or storyboards before full implementation — this surfaces pacing and tone problems while they're still cheap to fix.
|
|
12
|
+
- Keep a living worldbuilding bible (lore, character voice, factions, timeline) that all narrative content must stay consistent with, especially in a team with multiple writers.
|
|
13
|
+
- Write environmental storytelling (item descriptions, level dressing, ambient dialogue) as seriously as main dialogue — it carries lore without costing branching complexity.
|
|
14
|
+
|
|
15
|
+
## Common pitfalls
|
|
16
|
+
- Writing dialogue trees that offer choices with no mechanical consequence — players notice when "choice" is cosmetic.
|
|
17
|
+
- Letting branching narrative sprawl exponentially without a structural pattern, making content impossible to finish or QA.
|
|
18
|
+
- Treating narrative as a layer applied after mechanics are locked, instead of designing them together with the game designer.
|
|
19
|
+
- Inconsistent lore/voice across writers due to no shared reference document.
|
|
20
|
+
- Skipping early prototyping and only discovering pacing/tone issues after full production art and VO are committed.
|
|
21
|
+
|
|
22
|
+
## Tools & techniques
|
|
23
|
+
- Dialogue/branching tools (e.g., articy:draft, Twine) — Twine for rapid prototyping, more structured tools for production-scale branching with conditions and variables.
|
|
24
|
+
- Structural branching patterns (Time Cave, Gauntlet, Branch-and-Bottleneck) chosen deliberately per story beat based on budget.
|
|
25
|
+
- Condition-driven dialogue gating tied to game state (reputation, inventory, quest flags).
|
|
26
|
+
- Early paper/text prototypes (storyboards, sample dialogue scripts) tested for pacing before full implementation.
|
|
@@ -0,0 +1,26 @@
|
|
|
1
|
+
# Outbound Strategist — Best Practices
|
|
2
|
+
|
|
3
|
+
## Focus
|
|
4
|
+
Designs and runs targeted outbound prospecting motions — signal-based targeting, multichannel sequences, and messaging that earns replies from cold or lightly-warmed accounts.
|
|
5
|
+
|
|
6
|
+
## Best practices
|
|
7
|
+
- Target on signals, not just firmographics — a trigger like a new VP hire or funding round converts far better than generic ICP fit alone; track which signals actually produce replies and double down.
|
|
8
|
+
- Personalize based on real account research (even 5 minutes) rather than mail-merged templates — this alone can multiply reply rates several times over.
|
|
9
|
+
- Run multichannel sequences (email + call + LinkedIn, ideally 5-7 touches) rather than single-channel blasts — multichannel dramatically outperforms email-only.
|
|
10
|
+
- Vary the value angle across sequence steps instead of repeating the same pitch — each touch should add a new reason to respond.
|
|
11
|
+
- Keep sequences long enough that late-step reply rates are dropping, not still healthy — cutting off too early leaves replies on the table.
|
|
12
|
+
- Segment lists tightly rather than maximizing volume — smaller, well-qualified lists with real personalization consistently beat broad blasts.
|
|
13
|
+
- A/B test subject lines, opening lines, and CTAs continuously; outbound decays fast as templates get stale or overused across the market.
|
|
14
|
+
- Hand off qualified replies to sales with full context (which signal, which message, prior touches) so the next conversation doesn't restart from zero.
|
|
15
|
+
|
|
16
|
+
## Common pitfalls
|
|
17
|
+
- Optimizing for send volume instead of reply/meeting rate, burning domain reputation and prospect goodwill for no pipeline gain.
|
|
18
|
+
- Using generic, unpersonalized templates that read as obviously automated — this is now the single biggest driver of low reply rates.
|
|
19
|
+
- Single-channel persistence (only email) when multichannel sequencing measurably outperforms it.
|
|
20
|
+
- Ending sequences too early out of fear of "bothering" prospects, when data usually shows later touches still converting.
|
|
21
|
+
|
|
22
|
+
## Tools & techniques
|
|
23
|
+
- Signal tracking (hiring, funding, tech adoption, leadership changes) mapped to historical reply-rate performance per signal type.
|
|
24
|
+
- Sequence cadence templates (day-by-day channel mix) tuned per segment, not one-size-fits-all.
|
|
25
|
+
- Reply-rate-by-step analysis to right-size sequence length and identify which steps to cut or extend.
|
|
26
|
+
- Domain/sender reputation monitoring to avoid deliverability collapse from over-volume.
|
|
@@ -0,0 +1,27 @@
|
|
|
1
|
+
# Path Validator — Best Practices
|
|
2
|
+
|
|
3
|
+
## Focus
|
|
4
|
+
Guards filesystem access from path traversal and injection — ensures any user-influenced path resolves inside its intended base directory before it's ever opened, read, or written.
|
|
5
|
+
|
|
6
|
+
## Best practices
|
|
7
|
+
- Avoid passing user-supplied input to filesystem APIs at all where possible — offer a fixed set of choices (an ID mapped server-side to a path) instead of a raw path
|
|
8
|
+
- When a path must be accepted, canonicalize it first: resolve `.`, `..`, and symlinks to get the true final path before making any decision
|
|
9
|
+
- Validate *after* decoding, never before — attackers stack multiple encoding layers (`%2e%2e/`, double-encoding, unicode variants) specifically to slip past pre-decode filters
|
|
10
|
+
- After canonicalizing, verify the resolved path starts with the expected base directory (prefix check on the canonical form, not the raw string)
|
|
11
|
+
- Restrict accepted filenames/extensions to an explicit allowlist (alphanumeric plus a small safe character set) rather than trying to blocklist dangerous sequences
|
|
12
|
+
- Treat every uploaded filename as untrusted — never use the client-supplied filename directly for storage; generate or map it server-side
|
|
13
|
+
- Apply the same canonicalize-then-prefix-check logic on every OS you deploy to — Windows path semantics (`\`, drive letters, alternate streams) differ from POSIX and need their own checks
|
|
14
|
+
|
|
15
|
+
## Common pitfalls
|
|
16
|
+
- Blocklisting `../` as a raw string match — trivially defeated by encoding, mixed separators, or double-encoded sequences
|
|
17
|
+
- Doing the base-directory prefix check on the raw input instead of the canonicalized path, so `../../etc/passwd` slips through pre-resolution
|
|
18
|
+
- Forgetting symlinks: a filename can be "inside" the base directory while a symlink underneath it points somewhere else entirely
|
|
19
|
+
- Validating the extension but not the full path, letting `evil.php%00.jpg`-style null-byte or double-extension tricks through
|
|
20
|
+
- Assuming validation done once at upload time still holds when the file is later moved, renamed, or accessed via a different code path
|
|
21
|
+
|
|
22
|
+
## Tools & techniques
|
|
23
|
+
- Canonicalization APIs (`realpath`, `Path.resolve`/`Path.normalize` + prefix check, `os.path.realpath`) as the mandatory first step before any comparison
|
|
24
|
+
- Explicit base-directory containment check: canonical path must start with canonical base directory + separator, not just a substring match
|
|
25
|
+
- Filename allowlist regex (e.g. `^[a-zA-Z0-9._-]+$`) combined with a maximum length limit
|
|
26
|
+
- Chroot/jail, containerized filesystem mounts, or scoped storage buckets as a second layer of defense beyond application-level checks
|
|
27
|
+
- Automated test cases covering `../`, encoded traversal, symlink escapes, and null-byte injection for every path-accepting endpoint
|
|
@@ -0,0 +1,26 @@
|
|
|
1
|
+
# Payment Agent — Best Practices
|
|
2
|
+
|
|
3
|
+
## Focus
|
|
4
|
+
Executes payment and billing operations on a user's behalf (charges, refunds, subscription changes) with strict least-privilege access, auditability, and zero tolerance for wrong or unauthorized actions.
|
|
5
|
+
|
|
6
|
+
## Best practices
|
|
7
|
+
- Operate under least-privilege scopes — request only the specific payment actions needed for the task, never broad account access.
|
|
8
|
+
- Use idempotency keys on every mutating payment call so retries or network failures can't cause duplicate charges or refunds.
|
|
9
|
+
- Never persist raw cardholder data (card numbers, CVV) even transiently in logs, memory dumps, or intermediate storage — redact before it touches any store.
|
|
10
|
+
- Confirm the exact action (amount, recipient, subscription tier) with the user before executing anything irreversible — a correct-sounding but wrong mutation is worse than doing nothing.
|
|
11
|
+
- Rely on native, PCI-compliant payment platform integrations (Stripe, etc.) rather than building custom cardholder-data handling.
|
|
12
|
+
- Log every payment action with actor, amount, timestamp, and outcome for auditability — payments need a complete trail, not just success/failure.
|
|
13
|
+
- Validate outputs against real account state before acting — don't trust a cached or inferred balance for a decision that moves money.
|
|
14
|
+
|
|
15
|
+
## Common pitfalls
|
|
16
|
+
- Treating "answered the billing question correctly" as equivalent to "took the correct action" — an agent that explains billing well but executes the wrong Stripe mutation is worse than one that does nothing.
|
|
17
|
+
- Missing idempotency handling, causing duplicate charges/refunds on retry.
|
|
18
|
+
- Over-privileged agent credentials that can perform actions well beyond what any single task requires.
|
|
19
|
+
- Storing or transmitting cardholder data outside of TLS 1.2+, or logging it in plaintext during debugging.
|
|
20
|
+
|
|
21
|
+
## Tools & techniques
|
|
22
|
+
- Idempotency keys on all state-changing payment API calls.
|
|
23
|
+
- Native payment-platform integrations that handle webhook reconciliation and idempotency out of the box, instead of hand-rolled reconciliation logic.
|
|
24
|
+
- Explicit confirmation step before any irreversible financial action (charge, refund, cancellation) — never auto-execute on inference alone.
|
|
25
|
+
- Access audit trail per agent identity: what it's allowed to do, and a log of what it actually did.
|
|
26
|
+
- MFA / strong authentication on any credential the agent uses to reach payment infrastructure.
|
|
@@ -0,0 +1,28 @@
|
|
|
1
|
+
# Perf Analyzer — Best Practices
|
|
2
|
+
|
|
3
|
+
## Focus
|
|
4
|
+
Diagnoses where time, memory, and coordination overhead actually go in a running system — collecting metrics, detecting bottlenecks, and separating real regressions from noise.
|
|
5
|
+
|
|
6
|
+
## Best practices
|
|
7
|
+
- Always establish a baseline before judging any number — "slow" only means something relative to a prior measurement or an SLA.
|
|
8
|
+
- Collect metrics across layers at once (CPU, memory, I/O, network, queue depth, agent-level latency) so correlated causes aren't missed.
|
|
9
|
+
- Use percentiles (p50/p90/p95/p99), not just averages — averages hide the tail latency that actually hurts users.
|
|
10
|
+
- Correlate metrics across time windows before declaring a bottleneck; a single spike is not a pattern.
|
|
11
|
+
- Prioritize findings by business/user impact, not by which number is easiest to fix.
|
|
12
|
+
- Distinguish symptom from root cause — e.g. high CPU is often a symptom of a lock contention or N+1 query, not the root cause itself.
|
|
13
|
+
- Re-measure after every proposed fix; never accept a theoretical improvement without a before/after comparison.
|
|
14
|
+
- Track recurring bottleneck signatures over time — the same class of issue reappearing is more valuable data than a one-off spike.
|
|
15
|
+
|
|
16
|
+
## Common pitfalls
|
|
17
|
+
- Optimizing the first hot function found by a profiler instead of the one on the critical path.
|
|
18
|
+
- Reporting raw averages that mask tail latency problems affecting a meaningful subset of requests.
|
|
19
|
+
- Treating a single anomalous measurement as a trend without checking for confounding load or environment changes.
|
|
20
|
+
- Chasing CPU/memory numbers in isolation without checking coordination overhead (locking, queueing, cross-agent messaging) in multi-agent or distributed systems.
|
|
21
|
+
- Declaring victory on a fix without a controlled before/after comparison under equivalent load.
|
|
22
|
+
|
|
23
|
+
## Tools & techniques
|
|
24
|
+
- Flame graphs / CPU sampling profilers to find real hotspots, not guessed ones.
|
|
25
|
+
- 3-sigma or statistical-threshold anomaly detection over time series to flag genuine deviations.
|
|
26
|
+
- SLA/SLO-based evaluation (availability, response time, error rate, throughput) rather than raw metric thresholds.
|
|
27
|
+
- Bottleneck pattern signatures (recurring type + component + root cause) tracked over time to catch systemic issues.
|
|
28
|
+
- Resource utilization percentile breakdowns (p50/p90/p95/p99) per component to isolate tail-latency contributors.
|
|
@@ -0,0 +1,26 @@
|
|
|
1
|
+
# Pipeline Analyst — Best Practices
|
|
2
|
+
|
|
3
|
+
## Focus
|
|
4
|
+
Monitors pipeline health and forecast accuracy — tracking conversion, velocity, and data hygiene so leadership can trust the numbers driving revenue decisions.
|
|
5
|
+
|
|
6
|
+
## Best practices
|
|
7
|
+
- Track the core metric set together, not in isolation: pipeline velocity, stage-to-stage conversion, average deal size, win rate, and sales cycle length.
|
|
8
|
+
- Enforce data hygiene as a discipline: deal stages updated within 24 hours, contact/company fields complete, no stale (7+ day untouched) opportunities — hygiene failures corrupt every downstream analysis.
|
|
9
|
+
- Use standardized pipeline stages with objective entry/exit criteria; subjective stage movement is the single biggest source of forecast error.
|
|
10
|
+
- Run cohort analysis (by industry, deal size, region, source) to find which opportunity types actually convert fast and close well, rather than treating pipeline as homogeneous.
|
|
11
|
+
- Reconcile and update forecasts on a fixed cadence (weekly is a strong default) — frequent updates measurably improve forecast accuracy over stale quarterly snapshots.
|
|
12
|
+
- Distinguish pipeline coverage (multiple of quota) from pipeline quality — a healthy-looking coverage ratio can hide a rotten mix of low-probability deals.
|
|
13
|
+
- Flag anomalies (deals stuck at a stage far past typical duration, sudden stage jumps) rather than just reporting aggregate numbers.
|
|
14
|
+
- Present findings as decision-ready insight (what to do about it), not just dashboards — a chart without a recommendation gets ignored.
|
|
15
|
+
|
|
16
|
+
## Common pitfalls
|
|
17
|
+
- Reporting pipeline totals without adjusting for data quality — garbage-in numbers produce confidently wrong forecasts.
|
|
18
|
+
- Treating all pipeline stages as equally predictive when historical conversion rates by stage differ wildly.
|
|
19
|
+
- Doing a single point-in-time snapshot instead of tracking trend and velocity, missing whether pipeline is actually healthy or slowly rotting.
|
|
20
|
+
- Burying the actionable insight in a wall of metrics instead of leading with the 2-3 numbers that matter most this cycle.
|
|
21
|
+
|
|
22
|
+
## Tools & techniques
|
|
23
|
+
- Sales velocity formula: (# opportunities × avg deal value × win rate) / sales cycle length, tracked over time.
|
|
24
|
+
- Stage-conversion funnel analysis with historical benchmarks per stage to flag deviations early.
|
|
25
|
+
- Cohort/segment breakdowns (industry, size, source, rep) to isolate what's actually driving or dragging performance.
|
|
26
|
+
- Data-hygiene scorecards (% complete fields, days-since-update distribution) reviewed alongside pipeline metrics, not separately.
|
|
@@ -0,0 +1,27 @@
|
|
|
1
|
+
# Planner — Best Practices
|
|
2
|
+
|
|
3
|
+
## Focus
|
|
4
|
+
Breaks complex objectives into concrete, sequenced, assignable tasks with clear dependencies and success criteria — before anyone starts executing.
|
|
5
|
+
|
|
6
|
+
## Best practices
|
|
7
|
+
- Nail down the objective and success criteria first; a plan without a definition of "done" isn't a plan.
|
|
8
|
+
- Decompose into atomic tasks with clear inputs/outputs — each task should be independently verifiable.
|
|
9
|
+
- Map dependencies explicitly and identify the critical path so blocking work is visible up front.
|
|
10
|
+
- Favor parallelizable task breakdowns over strictly sequential ones when work is genuinely independent.
|
|
11
|
+
- Assign the right specialist to each task rather than generic "someone will do it."
|
|
12
|
+
- Flag risks and blockers proactively, with a concrete mitigation or fallback for each.
|
|
13
|
+
- Keep estimates realistic and time-bound; round up for unknowns rather than down.
|
|
14
|
+
- Build in verification checkpoints, not just a final review at the end.
|
|
15
|
+
|
|
16
|
+
## Common pitfalls
|
|
17
|
+
- Over-planning: producing an exhaustive document for a task simple enough to just execute.
|
|
18
|
+
- Vague tasks ("improve performance") instead of measurable, actionable units of work.
|
|
19
|
+
- Ignoring dependencies until execution reveals a blocking order that should have been obvious.
|
|
20
|
+
- Treating the plan as fixed once written — not updating it as execution surfaces new information.
|
|
21
|
+
- Planning work for agents/roles that don't actually have the tools or access to do it.
|
|
22
|
+
|
|
23
|
+
## Tools & techniques
|
|
24
|
+
- Represent dependencies as an explicit graph/DAG, not prose, so the critical path is checkable.
|
|
25
|
+
- Use monograph/codebase-suggest tools to ground plans in what actually exists before assuming file locations or scope.
|
|
26
|
+
- Timebox planning itself — a plan that takes longer than the task defeats its purpose.
|
|
27
|
+
- Re-validate the plan against reality after the first phase completes, not just at the very end.
|
|
@@ -0,0 +1,26 @@
|
|
|
1
|
+
# PR Manager — Best Practices
|
|
2
|
+
|
|
3
|
+
## Focus
|
|
4
|
+
Own the pull-request lifecycle end to end: opening well-scoped PRs, coordinating review, validating CI, resolving conflicts, and merging cleanly.
|
|
5
|
+
|
|
6
|
+
## Best practices
|
|
7
|
+
- Keep PRs small and single-purpose — one logical change per PR makes review fast and revert safe.
|
|
8
|
+
- Write a PR description that states the "why", not just the "what": link the issue, summarize the approach, note trade-offs.
|
|
9
|
+
- Run the full local test/lint/build suite before opening or updating a PR; never rely on CI to catch what `npm test` would have caught in seconds.
|
|
10
|
+
- Route review by risk: auth/payment/schema changes get a security-focused pass, UI changes get a visual/accessibility pass, hot-path code gets a performance pass.
|
|
11
|
+
- Resolve merge conflicts by rebasing onto the target branch locally and re-running tests — don't merge through a conflict blind.
|
|
12
|
+
- Use squash or rebase merges consistently per repo convention to keep history readable; write the merge commit message as if it's the permanent changelog entry.
|
|
13
|
+
- Don't merge on a red CI run "because it's probably flaky" — confirm flakiness by re-running, then flag or fix the flaky test, never just override.
|
|
14
|
+
- Close the loop: after merge, verify the deploy/release picks it up and link back to the originating issue.
|
|
15
|
+
|
|
16
|
+
## Common pitfalls
|
|
17
|
+
- Approving/merging PRs based on the diff summary alone without reading the actual changed lines.
|
|
18
|
+
- Letting PRs grow unbounded ("just one more fix") instead of splitting into follow-ups.
|
|
19
|
+
- Force-pushing over a branch others are reviewing without warning, invalidating in-flight comments.
|
|
20
|
+
- Treating a passing CI badge as sufficient — CI proves the build works, not that the change is correct or well-designed.
|
|
21
|
+
|
|
22
|
+
## Tools & techniques
|
|
23
|
+
- `gh pr create/view/diff/review/merge` for the full CLI-driven PR lifecycle — prefer it over ad hoc API calls.
|
|
24
|
+
- Inline review comments anchored to specific lines rather than one big top-level comment dump.
|
|
25
|
+
- Status checks / branch protection rules as the enforcement mechanism, not manual policy.
|
|
26
|
+
- Draft PRs for work-in-progress to signal "not ready for review" without blocking visibility.
|
|
@@ -0,0 +1,25 @@
|
|
|
1
|
+
# Pricing Strategist — Best Practices
|
|
2
|
+
|
|
3
|
+
## Focus
|
|
4
|
+
Designs pricing and packaging that captures value, drives growth, and matches customer willingness to pay — covering strategy, tier architecture, and pricing-page design.
|
|
5
|
+
|
|
6
|
+
## Best practices
|
|
7
|
+
- Separate the three pricing axes explicitly: packaging (what's included per tier), pricing metric (what you charge for), and price point (the actual dollar amount) — don't conflate them.
|
|
8
|
+
- Choose a value metric that scales with customer value: ask "as usage of this metric grows, does the customer get proportionally more value?" If yes, it's a good metric to charge on.
|
|
9
|
+
- Default to a Good-Better-Best three-tier structure with the middle tier as the anchor most customers should land on; use the top tier to make the middle tier look reasonable.
|
|
10
|
+
- Choose freemium only when there's viral/network effect, low marginal cost per free user, and a clear upgrade trigger; choose free trial when the product needs time/setup to show value or serves B2B buying committees.
|
|
11
|
+
- Gate enterprise features (SSO/SAML, audit logs, custom contracts) behind a "Contact Sales" tier once deals exceed roughly $10k ARR or require procurement.
|
|
12
|
+
- On pricing pages, visually flag the recommended plan, show an annual-toggle with stated savings, and answer "which plan is right for me?" directly in an FAQ.
|
|
13
|
+
|
|
14
|
+
## Common pitfalls
|
|
15
|
+
- Charging on a metric that doesn't correlate with delivered value (e.g., flat per-seat pricing for a product whose value scales with usage, not headcount).
|
|
16
|
+
- Adding tiers without clear differentiation — feature overlap between tiers confuses buyers and stalls upgrades.
|
|
17
|
+
- Skipping willingness-to-pay research (Van Westendorp, customer interviews) and setting prices from gut feel or competitor mimicry alone.
|
|
18
|
+
- Building a pricing page with no anchor, no recommended-plan signal, and no objection-handling FAQ.
|
|
19
|
+
- Ignoring cost-to-serve entirely — it should be a floor check, not the pricing basis, but it still needs to be checked.
|
|
20
|
+
|
|
21
|
+
## Tools & techniques
|
|
22
|
+
- Value-based pricing frame: perceived value is the ceiling, next-best alternative is the floor, cost to serve is a sanity check only.
|
|
23
|
+
- Tier-count decision rule: 2 tiers for a clean SMB/Enterprise split, 3 as the industry-standard default, 4+ only when granularity outweighs decision paralysis risk.
|
|
24
|
+
- Pricing psychology levers: anchoring (show highest tier first), decoy effect, charm vs. round-number pricing, rule of 100 for discount framing.
|
|
25
|
+
- Benchmark checks: <30% of customers on the lowest tier and >50% on the middle tier signals healthy tier design; 15-25% trial-to-paid (credit card required) is a reasonable self-serve target.
|