opencode-skills-collection 4.0.4 → 4.0.6
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/bundled-skills/.antigravity-install-manifest.json +5 -1
- package/bundled-skills/antigravity-maintainer-batch-release/SKILL.md +151 -0
- package/bundled-skills/create-pr/SKILL.md +4 -4
- package/bundled-skills/docs/README.md +1 -0
- package/bundled-skills/docs/contributors/quality-bar.md +1 -2
- package/bundled-skills/docs/integrations/jetski-cortex.md +5 -3
- package/bundled-skills/docs/integrations/jetski-gemini-loader/README.md +3 -1
- package/bundled-skills/docs/maintainers/aas-agent-first-control-plane-preview-profile.md +63 -0
- package/bundled-skills/docs/maintainers/aas-agent-first-control-plane-v1-design.md +304 -0
- package/bundled-skills/docs/maintainers/aas-agent-first-control-plane-v1-goal.md +171 -0
- package/bundled-skills/docs/maintainers/aas-agent-first-control-plane-v1-worklog.md +104 -0
- package/bundled-skills/docs/maintainers/audit.md +3 -8
- package/bundled-skills/docs/maintainers/legacy-redirect-bridge.md +46 -0
- package/bundled-skills/docs/maintainers/merge-batch.md +3 -7
- package/bundled-skills/docs/maintainers/merging-prs.md +1 -1
- package/bundled-skills/docs/maintainers/pr-autonomy.md +2 -3
- package/bundled-skills/docs/maintainers/release-process.md +12 -5
- package/bundled-skills/docs/maintainers/repo-growth-seo.md +16 -14
- package/bundled-skills/docs/maintainers/skills-update-guide.md +1 -1
- package/bundled-skills/docs/users/aas-core.md +182 -0
- package/bundled-skills/docs/users/agentic-awesome-skills-vs-awesome-claude-skills.md +10 -9
- package/bundled-skills/docs/users/ai-agent-skills.md +22 -19
- package/bundled-skills/docs/users/best-claude-code-skills-github.md +8 -8
- package/bundled-skills/docs/users/best-cursor-skills-github.md +5 -5
- package/bundled-skills/docs/users/bundles.md +3 -1
- package/bundled-skills/docs/users/claude-code-skills.md +23 -7
- package/bundled-skills/docs/users/codex-cli-skills.md +30 -13
- package/bundled-skills/docs/users/discovery-manifest.md +6 -3
- package/bundled-skills/docs/users/faq.md +37 -10
- package/bundled-skills/docs/users/gemini-cli-skills.md +1 -1
- package/bundled-skills/docs/users/getting-started.md +24 -10
- package/bundled-skills/docs/users/kiro-integration.md +1 -1
- package/bundled-skills/docs/users/plugins.md +15 -3
- package/bundled-skills/docs/users/skills-vs-mcp-tools.md +68 -49
- package/bundled-skills/docs/users/usage.md +50 -25
- package/bundled-skills/docs/users/visual-guide.md +28 -22
- package/bundled-skills/docs/vietnamese/AAS_CORE.vi.md +28 -0
- package/bundled-skills/docs/vietnamese/README.vi.md +10 -6
- package/bundled-skills/game-development/2d-games/SKILL.md +22 -19
- package/bundled-skills/game-development/SKILL.md +27 -35
- package/bundled-skills/game-development/engine-selection/SKILL.md +115 -0
- package/bundled-skills/game-development/web-games/SKILL.md +47 -45
- package/bundled-skills/git-advanced-workflows/SKILL.md +1 -11
- package/bundled-skills/git-pushing/SKILL.md +4 -4
- package/bundled-skills/markstream-install/SKILL.md +187 -0
- package/bundled-skills/markstream-install/references/scenarios.md +48 -0
- package/bundled-skills/modellix/SKILL.md +80 -0
- package/bundled-skills/pptx-deck-creation/SKILL.md +15 -7
- package/bundled-skills/pptx-deck-creation/references/audit-checklist.md +7 -3
- package/bundled-skills/pptx-deck-creation/references/ooxml-parsing.md +58 -0
- package/bundled-skills/pptx-deck-creation/references/{python-snippets.md → reference-deck-analysis-patterns.md} +1 -1
- package/bundled-skills/pptx-deck-creation/references/reference-deck-analysis.md +30 -8
- package/bundled-skills/pptx-deck-creation/references/visual-asset-adapters.md +23 -13
- package/bundled-skills/pr-writer/SKILL.md +2 -2
- package/bundled-skills/repo-maintainer/SKILL.md +4 -2
- package/bundled-skills/tools-page-seo-optimizer/SKILL.md +6 -6
- package/package.json +1 -1
- package/skills_index.json +121 -7
- package/bundled-skills/docs/contributors/skill-scoring.md +0 -235
|
@@ -0,0 +1,304 @@
|
|
|
1
|
+
# AAS Agent-First Control Plane v1 Design
|
|
2
|
+
|
|
3
|
+
> **Historical design:** deterministic Core recommendation, metadata eligibility, and selection policy described below were superseded on 2026-07-19 by the active [Agent-Owned Selection Profile](aas-agent-first-control-plane-preview-profile.md). Retained for architecture history; not current product policy.
|
|
4
|
+
|
|
5
|
+
Status: frozen and approved for implementation
|
|
6
|
+
Date: 2026-07-17
|
|
7
|
+
|
|
8
|
+
> **Historical target design:** This frozen document records the stronger certified-v1 target and preserves the terminology approved at that time. It is not a statement of current public guarantees. The supported public preview stops after plan review; see [`aas-agent-first-control-plane-preview-profile.md`](aas-agent-first-control-plane-preview-profile.md). In current product documentation, AAS Core is the product and `aas-stack.json` plus the plan are its durable artifacts.
|
|
9
|
+
|
|
10
|
+
## Product statement
|
|
11
|
+
|
|
12
|
+
> L'agente compone. Tu controlli. AAS mantiene lo stack aggiornato.
|
|
13
|
+
|
|
14
|
+
AAS finds, installs, and maintains the right set of skills for each project and agent. The durable product is the approved stack; CLI and MCP are its operational interfaces, while Workbench is the review surface.
|
|
15
|
+
|
|
16
|
+
## Understanding summary
|
|
17
|
+
|
|
18
|
+
- Users will ask Codex, Claude Code, or another agent to select skills instead of manually choosing among almost 2,000 entries.
|
|
19
|
+
- The agent inspects the project, sends an allowlisted synthetic profile to a local AAS MCP process, and explains the deterministic AAS result.
|
|
20
|
+
- A minimal `aas-stack.json` stores approved intent, policy, targets, catalog identity, and exact skill IDs.
|
|
21
|
+
- A human approves an immutable plan before the CLI writes anything.
|
|
22
|
+
- MCP is local, stdio, process-per-session, read-only, offline-capable, and contains no model or API credentials.
|
|
23
|
+
- One versioned core powers CLI, MCP, and Workbench projections.
|
|
24
|
+
- Missing metadata is reported as `unknown`; it is never presented as algorithmic certainty.
|
|
25
|
+
|
|
26
|
+
## Baseline
|
|
27
|
+
|
|
28
|
+
The published `agentic-awesome-skills` package is version 14.6.0 and exposes the legacy `agentic-awesome-skills` installer entrypoint. The installer already supports exact skill IDs, multiple targets, release pinning, dry-run, managed state, atomic preflight, and symlink safety. It does not yet expose stack lifecycle commands, a JSON recommendation API, or an MCP server.
|
|
29
|
+
|
|
30
|
+
The current catalog contains 1,965 skills. All entries have basic identity, category, source, risk, compatibility, and setup fields, but evidence coverage is uneven: 936 risks are `unknown`; only 361 entries have non-empty tags; `source_type`, source repository, and license are present for roughly one quarter of the catalog; structured test, review, and quality evidence fields are absent. Existing compatibility and setup values require provenance auditing before they can be treated as strong evidence.
|
|
31
|
+
|
|
32
|
+
## Scope
|
|
33
|
+
|
|
34
|
+
### Included
|
|
35
|
+
|
|
36
|
+
- One npm package with a shared deterministic core.
|
|
37
|
+
- Public entrypoints `aas`, `aas-mcp`, and the compatible legacy alias `agentic-awesome-skills`.
|
|
38
|
+
- Minimal stack manifest and versioned JSON schemas.
|
|
39
|
+
- Explicit catalog update/status and content-addressed local cache.
|
|
40
|
+
- Deterministic search, skill inspection, recommendation, eligibility, diff, validation, planning, apply, doctor, and recovery.
|
|
41
|
+
- User- or project-scoped MCP configuration adapters for Codex and Claude in v1.
|
|
42
|
+
- Minimal Workbench import/review for manifests and plans through user-mediated paste/upload; no ambient filesystem access or browser-side installation.
|
|
43
|
+
- Versioned benchmark, hostile-input corpus, independent package verifier, and release evidence bundle.
|
|
44
|
+
|
|
45
|
+
### Explicit non-goals
|
|
46
|
+
|
|
47
|
+
- Remote or hosted MCP, AAS accounts, login, cloud sync, or hosted API.
|
|
48
|
+
- Repository profile upload, remote telemetry, or analytics by default.
|
|
49
|
+
- Internal models, embeddings, remote ranking, or free inference from skill prose.
|
|
50
|
+
- Marketplace, public stack publishing, sharing, remixing, or community registry.
|
|
51
|
+
- Resident daemon/socket, implicit auto-update, or mandatory global installation.
|
|
52
|
+
- Runtime copied into every repository, multi-package split, or plugin system for the core.
|
|
53
|
+
- Full Workbench editor, browser filesystem access, or browser-triggered installation.
|
|
54
|
+
- Native configuration adapters for every agent host.
|
|
55
|
+
- Enrichment of all catalog entries as a prerequisite for preview.
|
|
56
|
+
- An OS sandbox guarantee or protection from a machine already compromised by a same-user process.
|
|
57
|
+
|
|
58
|
+
## Architecture
|
|
59
|
+
|
|
60
|
+
```text
|
|
61
|
+
User
|
|
62
|
+
-> Codex / Claude
|
|
63
|
+
-> aas-mcp (local stdio, one process per host session)
|
|
64
|
+
-> deterministic core
|
|
65
|
+
-> bundled or verified cached catalog
|
|
66
|
+
-> proposed aas-stack.json
|
|
67
|
+
-> human approval
|
|
68
|
+
-> aas stack plan --out .aas/plan.json
|
|
69
|
+
-> aas stack apply --plan .aas/plan.json
|
|
70
|
+
-> Workbench import/review
|
|
71
|
+
```
|
|
72
|
+
|
|
73
|
+
The npm package ships one internal core and two implementations, plus the legacy entrypoint alias. A user-local installation is the default for personal and non-JavaScript work. A project-local pinned devDependency is optional for JavaScript teams. Explicitly pinned `npx` is the bootstrap and CI path.
|
|
74
|
+
|
|
75
|
+
The first user-local bootstrap is an explicit trust-on-first-use boundary: a pinned `npx agentic-awesome-skills@<version> mcp configure --scope user` invocation runs with package lifecycle scripts disabled, records npm SRI, and installs the exact runtime closure into the content-addressed cache. Later runtime changes are performed by an already installed CLI through the same explicit `mcp configure` flow, with preview, integrity verification, atomic promotion, and no lifecycle execution. Catalog update remains data-only and cannot populate or execute the runtime cache.
|
|
76
|
+
|
|
77
|
+
The user-local cache separates runtime and catalog identities:
|
|
78
|
+
|
|
79
|
+
```text
|
|
80
|
+
runtimes/<package-version>/<filesystem-safe-integrity-key>/
|
|
81
|
+
catalogs/<package-version>/<catalog-digest>/
|
|
82
|
+
```
|
|
83
|
+
|
|
84
|
+
The runtime directory key is a canonical filesystem-safe encoding of the original npm SRI digest; the unmodified `dist.integrity` remains the value verified and recorded in plans and configuration. No manifest stores a local runtime path. MCP configuration binds to an exact runtime version and integrity. Upgrades are explicit.
|
|
85
|
+
|
|
86
|
+
## Stack manifest
|
|
87
|
+
|
|
88
|
+
The public v1 manifest stores desired state, not derived repository observations or natural-language reasoning:
|
|
89
|
+
|
|
90
|
+
```json
|
|
91
|
+
{
|
|
92
|
+
"schemaVersion": 1,
|
|
93
|
+
"name": "react-vite-production",
|
|
94
|
+
"catalog": {
|
|
95
|
+
"package": "agentic-awesome-skills",
|
|
96
|
+
"version": "14.6.0",
|
|
97
|
+
"integrity": "sha256-..."
|
|
98
|
+
},
|
|
99
|
+
"targets": [
|
|
100
|
+
{ "host": "codex", "scope": "project" }
|
|
101
|
+
],
|
|
102
|
+
"intent": {
|
|
103
|
+
"goals": ["build", "test", "deploy"]
|
|
104
|
+
},
|
|
105
|
+
"policy": {
|
|
106
|
+
"allowedRisk": ["none", "safe"],
|
|
107
|
+
"requireKnownSource": true,
|
|
108
|
+
"allowManualSetup": false
|
|
109
|
+
},
|
|
110
|
+
"skills": [
|
|
111
|
+
{ "id": "react-best-practices" },
|
|
112
|
+
{ "id": "playwright-skill" }
|
|
113
|
+
]
|
|
114
|
+
}
|
|
115
|
+
```
|
|
116
|
+
|
|
117
|
+
Detected languages/frameworks, excluded alternatives, evidence breakdown, and natural-language explanations remain in recommendation and plan output. The installer's internal managed-state manifest remains separate.
|
|
118
|
+
|
|
119
|
+
## MCP v1 contract
|
|
120
|
+
|
|
121
|
+
The server exposes only:
|
|
122
|
+
|
|
123
|
+
- `search_skills`
|
|
124
|
+
- `get_skill`
|
|
125
|
+
- `recommend_stack`
|
|
126
|
+
- `inspect_stack`
|
|
127
|
+
- `diff_stack`
|
|
128
|
+
- resource template `aas://skills/{id}`
|
|
129
|
+
|
|
130
|
+
`recommend_stack` applies deterministic rules and returns structured factors and evidence. It does not call another model or create subjective reasoning. `diff_stack` uses only verified catalogs already present locally. No MCP tool installs, removes, applies, updates catalogs, scans a repository, or changes configuration.
|
|
131
|
+
|
|
132
|
+
Full skill text is returned only on request and is separated as `untrustedContent`; metadata and prose cannot acquire instruction authority. The server can signal this trust boundary but cannot guarantee how an external model will behave.
|
|
133
|
+
|
|
134
|
+
Every structured response declares `protocolVersion`, `coreVersion`, `metadataSchemaVersion`, `scorerVersion`, and catalog digest. Incompatible versions fail explicitly.
|
|
135
|
+
|
|
136
|
+
## CLI lifecycle
|
|
137
|
+
|
|
138
|
+
```text
|
|
139
|
+
catalog update/status
|
|
140
|
+
-> stack init
|
|
141
|
+
-> agent/MCP recommend
|
|
142
|
+
-> stack validate
|
|
143
|
+
-> stack plan --out .aas/plan.json
|
|
144
|
+
-> human approval
|
|
145
|
+
-> stack apply --plan .aas/plan.json
|
|
146
|
+
-> stack doctor
|
|
147
|
+
```
|
|
148
|
+
|
|
149
|
+
- `stack init` creates targets and policy only.
|
|
150
|
+
- `stack recommend` is a deterministic fallback over an explicit profile file; it does not inspect the repository.
|
|
151
|
+
- The plan binds manifest hash, runtime identity/integrity, catalog digest, protocol/core/schema/scorer versions, installed managed state, and exact logical operations.
|
|
152
|
+
- Each immutable v1 apply plan binds exactly one target and one filesystem transaction. A multi-target manifest produces independently approved per-target plans; AAS does not claim impossible crash-atomic commit across unrelated filesystems.
|
|
153
|
+
- `apply` never recalculates an approved plan.
|
|
154
|
+
- Unmanaged skills are never overwritten or removed.
|
|
155
|
+
- Managed local modifications block apply unless an explicit override and approved backup are present.
|
|
156
|
+
- `stack doctor` is read-only. `stack recover --id ... --action rollback|cleanup` is a separate, approved, revalidated write.
|
|
157
|
+
- Reapplying a completed plan returns `alreadyApplied` without writing. A partial plan cannot be reused until recovery closes its journal.
|
|
158
|
+
- Interactive apply/recovery displays and confirms the exact plan or recovery digest. Non-interactive execution requires an explicit approval value bound to that digest; absence or mismatch blocks. This is an audit marker within the stated same-user threat boundary, not a cryptographic proof against a compromised machine.
|
|
159
|
+
- The legacy installer remains compatible and never creates `aas-stack.json` implicitly.
|
|
160
|
+
|
|
161
|
+
## Metadata and deterministic recommendation
|
|
162
|
+
|
|
163
|
+
The versioned metadata contract represents capability, target compatibility, risk, provenance/license, setup, dependencies/conflicts, validation, tests, and reviews. Each judgment can be `known`, `unknown`, or `notApplicable` and carries evidence references when known. The engine does not infer authoritative values from free-form skill text.
|
|
164
|
+
|
|
165
|
+
Eligibility is computed from public rules, never a hidden ID whitelist:
|
|
166
|
+
|
|
167
|
+
```text
|
|
168
|
+
eligibleForRecommendation: true | false
|
|
169
|
+
eligibilityReasonCodes: [...]
|
|
170
|
+
evidenceLevel: ...
|
|
171
|
+
unknownFields: [...]
|
|
172
|
+
```
|
|
173
|
+
|
|
174
|
+
All skills remain searchable. `recommend_stack` returns:
|
|
175
|
+
|
|
176
|
+
- `recommended`: candidates with sufficient structured evidence;
|
|
177
|
+
- `discoveryCandidates`: potentially relevant candidates with material unknowns.
|
|
178
|
+
|
|
179
|
+
An agent may promote a discovery candidate only through a visible plan override. Validation and apply still enforce the approved policy.
|
|
180
|
+
|
|
181
|
+
Ranking uses versioned aliases/ontology, explicit normalization, BM25-style lexical retrieval, fixed-point integer factors, and stable skill-ID tie-breaking. Composition is a declared greedy set-cover algorithm with skill budget, dependency/conflict handling, overlap penalty, minimum-value threshold, and permission to leave goals uncovered instead of adding a weak skill.
|
|
182
|
+
|
|
183
|
+
The canonical output includes scorer version, catalog digest, normalized input, factor breakdown, goal-capability matrix, exclusions with reason codes, material unknowns, and proposed stack. Identical input, catalog digest, schema, and scorer produce byte-identical canonical JSON. Timestamps, correlation IDs, localized messages, and diagnostics are outside that payload.
|
|
184
|
+
|
|
185
|
+
Coverage is reported with separate `goalCoverage`, `metadataCompleteness`, and `evidenceStrength`; no opaque aggregate confidence is required. Quality evidence is limited to auditable validation, provenance, metadata completeness, tests, and recorded reviews. Popularity, stars, and opaque editorial scores are excluded.
|
|
186
|
+
|
|
187
|
+
## Result and error model
|
|
188
|
+
|
|
189
|
+
`unknown` and `insufficientCoverage` are valid recommendation outcomes:
|
|
190
|
+
|
|
191
|
+
```json
|
|
192
|
+
{
|
|
193
|
+
"ok": true,
|
|
194
|
+
"status": "insufficientCoverage",
|
|
195
|
+
"reasonCodes": [],
|
|
196
|
+
"unknown": []
|
|
197
|
+
}
|
|
198
|
+
```
|
|
199
|
+
|
|
200
|
+
`ok: false` is reserved for invalid input, a policy-blocked requested operation, or execution failure. Machine clients depend only on versioned schemas, namespaced append-only codes, categories, and structured details. Human messages can change or be localized. Remediation is expressed as structured actions and arguments, never executable shell strings.
|
|
201
|
+
|
|
202
|
+
Mutating operations use OS-enforced exclusive lock creation plus an informational crash-safe record, same-filesystem staging, journal, revalidation immediately before writes, atomic rename where possible, fsync at commit points, rollback only when current bytes match AAS-written bytes, and internal-manifest update as the final commit. Paths are logical in plans and are derived from allowlisted host adapters. File handles, file type/ownership checks, `realpath`, and target identity are revalidated at write time to reduce swap/TOCTOU risk. Node cannot provide portable descriptor-relative `openat`/`renameat`; resistance to a malicious same-user process remains outside the v1 boundary rather than being overstated. Recovery uses a `recoveryId`, target identity, hashes, and explicit approval.
|
|
203
|
+
|
|
204
|
+
Diagnostics are redacted. Correlation IDs are allowed; stack traces require explicit redacted debug mode.
|
|
205
|
+
|
|
206
|
+
## Privacy and threat model
|
|
207
|
+
|
|
208
|
+
Trust boundaries are npm registry to cache, cache to runtime, host/agent to MCP, recommendation to human approval, approved plan to CLI/filesystem, catalog content to the interpreting agent, and host configuration to its adapter/backup.
|
|
209
|
+
|
|
210
|
+
Key invariants:
|
|
211
|
+
|
|
212
|
+
- The updater downloads a tarball, verifies registry `dist.integrity`, extracts only allowlisted data assets, and verifies the internal catalog digest. It never executes release lifecycle scripts, binaries, modules, templates, dynamic imports, or catalog code. Registry integrity proves byte correspondence, not that the publisher account was uncompromised; provenance/attestations are recorded separately.
|
|
213
|
+
- Archive extraction rejects absolute/traversal paths, symlinks, hardlinks, devices/FIFOs, duplicate files, case/Unicode collisions, anomalous permissions, excessive file counts or sizes, and decompression bombs.
|
|
214
|
+
- MCP contains no network calls, updater, model, credentials, or telemetry and works offline. Tests must observe zero network attempts. OS-level network denial is optional external hardening.
|
|
215
|
+
- MCP accepts only bounded, allowlisted structured profiles. It does not persist profiles or log source, secrets, raw files, or absolute paths by default.
|
|
216
|
+
- MCP enforces byte, JSON-depth, query-length, result-count, memory, and timeout limits.
|
|
217
|
+
- Apply is transactional and fail-closed against drift, target swap, path traversal, symlink races, incompatible producers, and hard-policy violations.
|
|
218
|
+
- Host adapters reject symlinks/non-regular files and wrong ownership, preserve mode/owner, lock and patch atomically, redact secrets from diffs/logs, and create user-only backups with explicit retention and cleanup.
|
|
219
|
+
- A same-user malicious process or compromised machine is outside the v1 protection boundary.
|
|
220
|
+
|
|
221
|
+
## Benchmark and acceptance gates
|
|
222
|
+
|
|
223
|
+
Before product implementation, a separate bootstrap phase creates the reference evaluator, schemas, tuning set, held-out set, hostile corpus, legacy command corpus, and verifier. Their initial versions and digests require review by two named reviewers, at least one of whom does not implement the scorer. After that freeze, the product change cannot modify them; any separately approved revision invalidates prior evidence.
|
|
224
|
+
|
|
225
|
+
The v1 public benchmark freezes tuning data separately from held-out data. Each supported intent has at least 30 held-out cases distributed across declared sub-intents and project archetypes, deduplicated by task/project family. Gold sets allow multiple equivalent solutions, record provenance/version, and require two reviewers for ambiguous cases. Labels are frozen before execution and cannot be reclassified after observing a result.
|
|
226
|
+
|
|
227
|
+
For every supported intent and in macro-average:
|
|
228
|
+
|
|
229
|
+
- verified recommendation coverage is at least 80%;
|
|
230
|
+
- inclusion precision is at least 90%;
|
|
231
|
+
- explicitly out-of-coverage cases abstain 100%;
|
|
232
|
+
- hard-policy violations are zero;
|
|
233
|
+
- critical goals are fully covered;
|
|
234
|
+
- declared minimum coverage for non-critical goals is met;
|
|
235
|
+
- discovery promotions always have visible overrides.
|
|
236
|
+
|
|
237
|
+
A verified recommendation must satisfy all of those conditions, not merely produce a stack. The minimum non-critical-goal coverage is 80%. Verified coverage uses every frozen in-scope case as its denominator. Inclusion precision is computed per stack and macro-averaged per intent over every included skill, using the accepted-equivalent sets; an empty or partial in-scope stack fails verified coverage and cannot disappear from the denominator. Out-of-coverage cases are measured only in the separately frozen abstention set. Candidate diversity is reported; three eligible candidates are preferred when the ecosystem genuinely offers them, but weak candidates are never added to satisfy a quota.
|
|
238
|
+
|
|
239
|
+
Hard-policy violations must also remain zero across the independently approved generative, property, fuzz, and hostile-input corpus. Minimum budgets are 100,000 stratified generated/property cases for policy and eligibility plus 50,000 bounded parser/MCP fuzz inputs, with all seeds and distributions frozen before scorer implementation. The hostile corpus contains at least one exploit and one boundary-adjacent valid control for every declared archive/input class. The canonical core payload must be byte-identical across the supported OS/Node matrix.
|
|
240
|
+
|
|
241
|
+
The initial supported intents are:
|
|
242
|
+
|
|
243
|
+
1. web application delivery;
|
|
244
|
+
2. API/backend delivery;
|
|
245
|
+
3. test and QA automation;
|
|
246
|
+
4. security review and hardening;
|
|
247
|
+
5. deployment and DevOps;
|
|
248
|
+
6. agent and MCP development.
|
|
249
|
+
|
|
250
|
+
An intent failing any gate remains `preview` or unsupported. The supported runtime matrix is Node v22 and v24 on Linux, macOS, and Windows; Node v20 is EOL and excluded. The verifier manifest freezes exact Node patch versions, runner/image identities, filesystem assumptions, architecture, and all required jobs before execution. Skips and `continue-on-error` fail the gate.
|
|
251
|
+
|
|
252
|
+
## Independent completion verifier
|
|
253
|
+
|
|
254
|
+
Repository tests are supporting evidence, not the final verifier. The strongest check is an independently controlled black-box harness:
|
|
255
|
+
|
|
256
|
+
```text
|
|
257
|
+
candidate commit
|
|
258
|
+
-> npm pack
|
|
259
|
+
-> content-addressed tarball
|
|
260
|
+
-> clean install outside the checkout with --ignore-scripts
|
|
261
|
+
-> independent verifier and evidence bundle
|
|
262
|
+
-> benchmark/security/release approvals
|
|
263
|
+
-> protected publish
|
|
264
|
+
-> registry re-download and integrity/behavior comparison
|
|
265
|
+
```
|
|
266
|
+
|
|
267
|
+
The verifier checks packaging, all entrypoints, offline catalog access, MCP protocol and resource limits, zero persistent MCP writes anywhere and zero network attempts, hostile archives, tampered plans, fault injection and crash recovery, canary-secret leakage, adapter fixtures, legacy CLI differential behavior against the integrity-pinned 14.6.0 package, and OS/Node behavior. Fault injection acts on the production binary using OS/process/filesystem observation and kills or swaps at every observed mutating boundary; mock-only or test-mode coverage is insufficient. The frozen legacy corpus enumerates every public flag, target, representative combination, filesystem result, exit code, and explicitly allowed difference.
|
|
268
|
+
|
|
269
|
+
Eligibility and scoring also undergo metamorphic tests that permute catalog order and replace IDs consistently; results must remain invariant except for the documented stable tie-break and returned renamed identity. This prevents public rules from becoming an indirect hardcoded whitelist.
|
|
270
|
+
|
|
271
|
+
The verifier runs in protected CI under declared ownership and cannot be self-approved by the product implementer. It produces a content-addressed evidence bundle bound to CI run identity, verifier hash, commit, tarball, approvals, and retained protected artifact storage. The bundle contains all versions/digests, full denominators, seeds/budgets, system traces, transaction matrices, redacted logs, and reviewer approvals.
|
|
272
|
+
|
|
273
|
+
Workbench accepts only size/depth-bounded, user-mediated paste or file upload held in memory. It performs schema validation and text-only/XSS-safe rendering; it does not use ambient filesystem APIs or persist imported content. Completion requires live GitHub Pages deployment and readback of the exact reviewed version, behind the publication approval gate.
|
|
274
|
+
|
|
275
|
+
`implementationVerified`, `releaseReady`, and `released` are distinct states. Publication requires explicit user approval and the repository's protected maintainer release workflow.
|
|
276
|
+
|
|
277
|
+
## Decision log
|
|
278
|
+
|
|
279
|
+
| Decision | Alternatives considered | Rationale |
|
|
280
|
+
| --- | --- | --- |
|
|
281
|
+
| Stack + CLI + MCP + Workbench | Site-only stack, new installer CLI, MCP wrapper | The stack is durable state; interfaces alone do not create recurring product value. |
|
|
282
|
+
| Agent-first, human-approved | Manual Workbench composition | Users delegate selection; humans need review and control, not 2,000 checkboxes. |
|
|
283
|
+
| Local stdio MCP | Hosted MCP/API, resident daemon | Preserves privacy, works offline, and keeps process isolation simple. |
|
|
284
|
+
| MCP read-only; CLI owns writes | MCP apply/install tools | Preserves an explicit approval boundary and reduces host-agent blast radius. |
|
|
285
|
+
| One npm package and one core | Multiple packages or duplicated logic | Minimizes version skew while preserving later split options. |
|
|
286
|
+
| Hybrid runtime placement | Mandatory project install or global npm install | User-local works across languages; project-local remains available for team pinning. |
|
|
287
|
+
| Minimal manifest | Persist detected profile and prose reasoning | Derived data drifts and causes noisy diffs; approved intent and IDs are durable. |
|
|
288
|
+
| Catalog identity includes integrity | Version string only or immediate lockfile | Binds desired state to verified bytes without adding a second public state file. |
|
|
289
|
+
| Public eligibility rules and two result lanes | Hidden enriched whitelist | Keeps the whole catalog visible and makes incomplete evidence explicit. |
|
|
290
|
+
| Unknown is first-class | Treat unknown as incompatible or infer from prose | Prevents false certainty while allowing policy-controlled caution. |
|
|
291
|
+
| Lexical fixed-point deterministic ranking | Embeddings, remote scoring, model ranking | Reproducible across CLI, MCP, and Workbench with auditable factors. |
|
|
292
|
+
| Separate coverage/evidence measures | Single confidence percentage | Avoids presenting missing metadata as certainty. |
|
|
293
|
+
| Immutable plan and transactional apply | Recompute on apply or direct install | Human approval must bind the exact operation and survive drift/failure safely. |
|
|
294
|
+
| One target/filesystem per mutating plan | Claim whole-plan atomicity across multiple host filesystems | Preserves real crash-atomic semantics; multi-host manifests remain portable through separate approved plans. |
|
|
295
|
+
| Pure Node write-time containment | Native addon or overstated portable `openat` guarantee | Matches the same-user threat boundary and avoids a new native distribution surface while documenting residual TOCTOU risk. |
|
|
296
|
+
| Registry integrity plus internal digest | Digest from same untrusted file alone | Verifies published bytes and catalog consistency while documenting publisher compromise as residual risk. |
|
|
297
|
+
| Valid abstention is `ok: true` | Treat insufficient coverage as error | Clients and metrics must distinguish safe abstention from system failure. |
|
|
298
|
+
| Independent black-box tarball verifier | Repository test suite alone | Prevents checkout-only success and proves the published artifact boundary. |
|
|
299
|
+
| Per-intent 80/90/100 gates | Global average or raw skill-count gate | Prevents strong categories from hiding weak ones and tests correct abstention. |
|
|
300
|
+
| Launch parameters frozen before activation | Leave intent, adapter, benchmark-size, runtime, and Workbench scope open | Product-owner approval fixes the cost and finish line before implementation begins. |
|
|
301
|
+
|
|
302
|
+
## Maintenance ownership
|
|
303
|
+
|
|
304
|
+
The scorer, metadata schema, benchmark, hostile corpus, and verifier are versioned public maintenance surfaces. Their initial baseline is created and independently approved before scorer implementation. Changes to held-out labels, security corpus, intent set, denominators, fuzz budgets, or verifier require separate review from product implementation and invalidate previous evidence. Adapter fixtures record provenance, host version, and validation date and require an isolated smoke test against the current version before release. Release artifacts must use the exact verified tarball; rebuilding requires a new verification cycle.
|
|
@@ -0,0 +1,171 @@
|
|
|
1
|
+
# AAS Agent-First Control Plane v1 Goal
|
|
2
|
+
|
|
3
|
+
> **Historical goal:** recommendation-quality, metadata, and policy gates described below were superseded on 2026-07-19 by the active [Agent-Owned Selection Profile](aas-agent-first-control-plane-preview-profile.md). Retained for decision history; not a current release gate.
|
|
4
|
+
|
|
5
|
+
Status: approved source packet for the active Codex goal
|
|
6
|
+
Design source: `docs/maintainers/aas-agent-first-control-plane-v1-design.md`
|
|
7
|
+
|
|
8
|
+
> **Historical goal packet:** This file preserves the original certified-v1 finish line, including apply/recovery and independent verification work. It is not the current public preview contract or a claim that those gates shipped. The supported public preview stops after plan review; see [`aas-agent-first-control-plane-preview-profile.md`](aas-agent-first-control-plane-preview-profile.md).
|
|
9
|
+
|
|
10
|
+
## Fit
|
|
11
|
+
|
|
12
|
+
Use a durable goal. The work crosses catalog schema, deterministic recommendation, CLI lifecycle, local MCP, transactional filesystem behavior, host adapters, Workbench review, cross-platform packaging, benchmark construction, security abuse testing, and protected release verification. It needs repeated implementation/verification loops and has an independent observable finish line.
|
|
13
|
+
|
|
14
|
+
## Outcome
|
|
15
|
+
|
|
16
|
+
Ship and independently verify the AAS v1 agent-first control plane exactly within the frozen design: a user can ask Codex or Claude to obtain a deterministic local recommendation, approve a minimal stack and immutable plan, safely apply it through the CLI, diagnose/recover failures, and review the result in Workbench without repository data leaving the machine by default.
|
|
17
|
+
|
|
18
|
+
## Baseline
|
|
19
|
+
|
|
20
|
+
- Package 14.6.0 exposes only the legacy installer entrypoint.
|
|
21
|
+
- The installer already provides exact-set, multi-target, pinning, dry-run, managed-state, atomic-preflight, and symlink-safety foundations.
|
|
22
|
+
- There is no AAS MCP server, stack lifecycle, deterministic recommendation API, verified catalog cache lifecycle, benchmark, or independent package verifier.
|
|
23
|
+
- Catalog metadata is incomplete and cannot yet support strong recommendations across the full catalog.
|
|
24
|
+
- The working tree contained unrelated local-skill-reviewer work when this goal was drafted; it must remain isolated from this goal.
|
|
25
|
+
|
|
26
|
+
## Fixed scope and launch parameters
|
|
27
|
+
|
|
28
|
+
Implement only the surfaces frozen in the design document. Initial supported intents are web application delivery, API/backend delivery, test/QA automation, security review/hardening, deployment/DevOps, and agent/MCP development. Each requires at least 30 diversified and deduplicated held-out cases. Initial configuration adapters are Codex and Claude. Workbench scope is read-only, in-memory, user-mediated paste/upload and review of stack and plan artifacts, including a verified live Pages deployment.
|
|
29
|
+
|
|
30
|
+
The runtime matrix is Node majors 22 and 24 on Linux, macOS, and Windows, with exact patch, runner/image, architecture, and filesystem identities frozen in the verifier manifest. The npm package exposes `aas`, `aas-mcp`, and the compatible `agentic-awesome-skills` alias.
|
|
31
|
+
|
|
32
|
+
## Non-goals
|
|
33
|
+
|
|
34
|
+
No hosted service, remote MCP, account system, telemetry, repository upload, embeddings/model in the core, marketplace, public stack sharing, daemon, auto-update, multi-package split, full Workbench editor, browser installation, all-host adapter program, or mandatory enrichment of all skills.
|
|
35
|
+
|
|
36
|
+
## Primary verifier
|
|
37
|
+
|
|
38
|
+
Before scorer/product implementation, a separate baseline phase creates the reference evaluator, metric schema, benchmark/gold data, hostile corpus, legacy command corpus, and black-box verifier. Two named reviewers, including at least one non-implementer of the scorer, approve and freeze their versions and digests. The product implementation cannot modify these surfaces.
|
|
39
|
+
|
|
40
|
+
That independently controlled black-box harness must verify the content-addressed `npm pack` tarball from a clean directory outside the repository, installed with lifecycle scripts disabled and without checkout or devDependency access. Protected CI, with declared ownership and no self-approval by the product implementer, produces the evidence bundle and distinguishes:
|
|
41
|
+
|
|
42
|
+
- `implementationVerified`: all technical gates pass on the candidate tarball;
|
|
43
|
+
- `releaseReady`: benchmark and independent security/release review are approved;
|
|
44
|
+
- `released`: the protected tag/npm release exists and the registry tarball is re-downloaded and proven identical in integrity and behavior.
|
|
45
|
+
|
|
46
|
+
Repository tests alone cannot complete the goal.
|
|
47
|
+
|
|
48
|
+
## Required gates
|
|
49
|
+
|
|
50
|
+
### Recommendation and metadata
|
|
51
|
+
|
|
52
|
+
- Versioned metadata schema with explicit unknowns and evidence.
|
|
53
|
+
- Entire catalog searchable; eligibility derived only from public rules; no ID whitelist.
|
|
54
|
+
- `recommended` and `discoveryCandidates` are both visible with structured reasons.
|
|
55
|
+
- Deterministic core output is byte-identical for identical input/digests/versions across the matrix.
|
|
56
|
+
- For every supported intent and in macro-average: verified coverage >=80%, per-stack inclusion precision macro-averaged per intent >=90%, explicit out-of-coverage abstention 100%, critical-goal coverage 100%, non-critical-goal coverage >=80%, zero hard-policy violations.
|
|
57
|
+
- Verified coverage counts every frozen in-scope case; empty, partial, crashed, timed-out, or missing results fail and remain in the denominator. Inclusion precision evaluates every included skill against accepted-equivalent gold sets. The separately frozen abstention set cannot be relabeled after results are known.
|
|
58
|
+
- At least 30 frozen held-out cases per intent, distributed across sub-intents and project archetypes, deduplicated by task/project family, with multiple equivalent gold solutions, provenance/version, and double review for ambiguous cases.
|
|
59
|
+
- Zero hard-policy violations across at least 100,000 independently frozen stratified property/generative cases and 50,000 bounded parser/MCP fuzz inputs, plus the hostile corpus with exploit and valid boundary controls per class.
|
|
60
|
+
- Metamorphic catalog-order and consistent-ID-permutation tests prove eligibility/scoring do not encode a direct or indirect ID whitelist.
|
|
61
|
+
|
|
62
|
+
### Package, CLI, and MCP
|
|
63
|
+
|
|
64
|
+
- Clean tarball contains only allowlisted assets and no sensitive or checkout-only dependencies.
|
|
65
|
+
- `aas`, `aas-mcp`, and the legacy alias pass black-box smoke tests.
|
|
66
|
+
- Legacy behavior is differentially tested against the integrity-pinned 14.6.0 package using a frozen corpus of every public flag, target, representative combination, output tree, exit code, and allowed difference; it does not create stack state implicitly.
|
|
67
|
+
- MCP exposes only the frozen five tools and resource template, performs no mutations or updates, works completely offline, produces zero observed network attempts, and produces zero persistent filesystem writes anywhere. HOME, project, cache, and TMP are isolated and observed; stdout/stderr are process streams, not write exceptions.
|
|
68
|
+
- MCP keeps protocol-only stdout, redacted diagnostics, untrusted skill-content separation, and resource limits.
|
|
69
|
+
|
|
70
|
+
### Catalog and supply chain
|
|
71
|
+
|
|
72
|
+
- Runtime/catalog caches are content-addressed and atomically promoted.
|
|
73
|
+
- First bootstrap uses explicitly pinned `npx ... mcp configure --scope user` with lifecycle scripts disabled and records the npm SRI; later runtime changes use the installed CLI with preview and atomic integrity-verified promotion. Catalog update remains data-only.
|
|
74
|
+
- Updater verifies npm `dist.integrity`, extracts only allowlisted data assets without executing code, and verifies the internal catalog digest.
|
|
75
|
+
- The hostile archive corpus covers traversal, absolute paths, links, special files, duplicates, Unicode/case collisions, permissions, count/size limits, and decompression bombs.
|
|
76
|
+
- Every hostile case leaves the cache unchanged and launches no child process or code.
|
|
77
|
+
|
|
78
|
+
### Plan, apply, and recovery
|
|
79
|
+
|
|
80
|
+
- Plan binds every approved input and never gets recomputed by apply.
|
|
81
|
+
- Every mutating plan binds exactly one target/filesystem. Multi-target manifests produce separate plans and approvals; cross-filesystem atomicity is not claimed.
|
|
82
|
+
- Apply uses OS-enforced exclusive lock creation plus crash-safe identity records, and is transactional, same-filesystem staged, crash-safe, fail-closed, and idempotent.
|
|
83
|
+
- Black-box fault injection against the production binary, driven by OS/process/filesystem observation rather than mocks or test-mode branches, covers every observed lock, journal, backup, write, fsync, rename, and commit boundary, including kill, concurrency, drift, symlink/target swap, corrupted journal, and recovery races.
|
|
84
|
+
- Unmanaged bytes never change. Managed local edits block unless separately overridden and backed up.
|
|
85
|
+
- Final state is entirely previous or entirely new; never hybrid. Internal managed state commits last.
|
|
86
|
+
- Final-state atomicity is evaluated per approved target transaction after successful apply or completed recovery. Write-time file-handle/type/ownership/realpath/identity checks reduce swap risk; a malicious same-user process remains outside the v1 boundary.
|
|
87
|
+
- Doctor is read-only. Recovery uses an approved ID/action and revalidates target and hashes.
|
|
88
|
+
- Interactive writes require confirmation of the exact plan/recovery digest. Non-interactive writes require an explicit approval value bound to that digest; absence or mismatch blocks.
|
|
89
|
+
|
|
90
|
+
### Host configuration and Workbench
|
|
91
|
+
|
|
92
|
+
- Codex and Claude adapters use real anonymized fixtures with provenance, host version, and validation date, plus isolated current-client smoke tests. They preserve unknown fields, apply minimal atomic patches, preserve ownership/mode, create user-only retained/cleanable backups, lock, redact secrets, and reject unsafe file types or ownership.
|
|
93
|
+
- Workbench accepts only size/depth-bounded user paste/upload held in memory and renders stack/plan evidence through schema-validated, text-only/XSS-safe views. It does not install, persist imports, use ambient filesystem APIs, or read local files without the user's explicit selection. Release proof includes the live Pages URL, deployed version, and readback after publication approval.
|
|
94
|
+
|
|
95
|
+
## Iteration loop
|
|
96
|
+
|
|
97
|
+
1. Treat the approved intent, held-out, adapter, runtime-matrix, Workbench, and fuzz-budget parameters as frozen.
|
|
98
|
+
2. Create and independently approve the initial reference evaluator, benchmark/gold, hostile corpus, legacy corpus, and verifier baseline; record their protected digests before product implementation.
|
|
99
|
+
3. Inspect the current design, goal, worklog, repository state, and active goal state.
|
|
100
|
+
4. Select one bounded vertical slice that advances the black-box verifier.
|
|
101
|
+
5. Implement it without changing frozen scope or benchmark labels.
|
|
102
|
+
6. Run targeted tests, then the relevant repository gates.
|
|
103
|
+
7. Run the independent verifier or its currently available slice against an actual tarball.
|
|
104
|
+
8. Record commands, hashes, failures, and the next smallest corrective action in a durable worklog.
|
|
105
|
+
9. Repeat until every gate passes from a clean state on the full frozen matrix.
|
|
106
|
+
10. Stop before public/tag/npm/Pages actions and request explicit approval.
|
|
107
|
+
|
|
108
|
+
## Anti-cheating rules
|
|
109
|
+
|
|
110
|
+
- Do not weaken, skip, quarantine, relabel, post-filter, or retry-until-green any gate.
|
|
111
|
+
- Initial benchmark/corpus/verifier baselines require independent approval before scorer implementation. The product PR cannot modify them or self-approve their workflow/ownership controls. Changes require a separate approval and invalidate prior evidence.
|
|
112
|
+
- Freeze intents, formulas, denominators, outside-coverage labels, minimum fuzz seeds/budgets/distributions, exact runners, and OS/Node matrix before scorer implementation.
|
|
113
|
+
- Crashes, timeouts, and missing outputs count as failures.
|
|
114
|
+
- No test-mode branch based on case IDs, fixture names, or environment markers.
|
|
115
|
+
- Canonical comparison exclusions must be enumerated fields, never a generic metadata exclusion.
|
|
116
|
+
- Do not replace real tarball, network/filesystem observation, or fault injection with mocks for completion proof.
|
|
117
|
+
- Do not use skipped/allowed-failure matrix jobs or a fault-injection list that omits an observed production mutation boundary.
|
|
118
|
+
- Do not touch or absorb unrelated dirty work.
|
|
119
|
+
- Do not call the goal complete at `implementationVerified` or `releaseReady`; completion requires the separately approved released state.
|
|
120
|
+
|
|
121
|
+
## Approval gates
|
|
122
|
+
|
|
123
|
+
Explicit user approval is required before:
|
|
124
|
+
|
|
125
|
+
- modifying real user-level Codex or Claude configuration outside isolated fixtures;
|
|
126
|
+
- changing frozen intent, benchmark, corpus, verifier, security boundary, or v1 scope;
|
|
127
|
+
- publishing a tag, GitHub Release, npm package, or public product announcement;
|
|
128
|
+
- deploying the Workbench changes publicly to Pages;
|
|
129
|
+
- any migration or cleanup that could remove existing user state.
|
|
130
|
+
|
|
131
|
+
The protected maintainer release workflow remains mandatory for publication.
|
|
132
|
+
|
|
133
|
+
## Blocker standard
|
|
134
|
+
|
|
135
|
+
Difficulty, test failure, incomplete metadata, or a long implementation is not a blocker. Report blocked only after the same external dependency or missing approval prevents meaningful progress for the required repeated goal turns. Preserve partial artifacts and state the smallest user or external action that would unblock the verifier.
|
|
136
|
+
|
|
137
|
+
## Completion proof
|
|
138
|
+
|
|
139
|
+
Completion requires:
|
|
140
|
+
|
|
141
|
+
1. exact candidate commit and clean-scope proof;
|
|
142
|
+
2. tarball SHA-512, npm pack manifest, and registry `dist.integrity`;
|
|
143
|
+
3. protocol/core/schema/scorer versions and catalog digest;
|
|
144
|
+
4. per-intent reports with full denominators and reviewer provenance;
|
|
145
|
+
5. canonical output hashes plus exact Node, runner/image, architecture, and filesystem identities for the complete matrix, with no skipped or allowed-failure job;
|
|
146
|
+
6. fuzz seeds/budgets, hostile-corpus results, and zero-policy-violation report;
|
|
147
|
+
7. observed MCP network/filesystem/process traces;
|
|
148
|
+
8. updater archive matrix and unchanged-cache proofs;
|
|
149
|
+
9. fault-injection, crash/recovery, and pre/post filesystem snapshots;
|
|
150
|
+
10. frozen legacy command corpus, baseline 14.6.0 integrity, explicit allowed-difference list, and differential report;
|
|
151
|
+
11. Codex/Claude fixture provenance/current-client adapter reports and canary-secret scans;
|
|
152
|
+
12. Workbench import-security tests plus approved live Pages URL/version/readback;
|
|
153
|
+
13. protected CI run identity, verifier owner/version/hash, reviewer identities, attestation, retention location, and content-addressed evidence bundle bound to commit and tarball;
|
|
154
|
+
14. benchmark, security, and release approvals;
|
|
155
|
+
15. protected release evidence plus registry re-download integrity and behavior comparison.
|
|
156
|
+
|
|
157
|
+
Only after all evidence exists and no required work remains may the active goal be marked complete.
|
|
158
|
+
|
|
159
|
+
## Delegation map
|
|
160
|
+
|
|
161
|
+
The primary agent owns scope, integration, repository changes, conflict resolution, and completion. Bounded subagents may independently handle metadata/benchmark audit, CLI/MCP contract tests, security corpus review, cross-platform verification, Workbench review validation, or release evidence review. They may not change the frozen benchmark, approve their own implementation, publish, or declare the parent goal complete.
|
|
162
|
+
|
|
163
|
+
## Exact activation objective
|
|
164
|
+
|
|
165
|
+
```text
|
|
166
|
+
Implement, independently verify, and—only after explicit publication approval—release the AAS agent-first control plane v1 defined in /Users/nicco/Projects/antigravity-awesome-skills/docs/maintainers/aas-agent-first-control-plane-v1-design.md, satisfying every gate and completion proof in /Users/nicco/Projects/antigravity-awesome-skills/docs/maintainers/aas-agent-first-control-plane-v1-goal.md without expanding the frozen v1 scope or touching unrelated dirty work.
|
|
167
|
+
```
|
|
168
|
+
|
|
169
|
+
## Activation state
|
|
170
|
+
|
|
171
|
+
This packet is the source of truth for the active Codex goal. Live activation status is maintained by Codex goal state.
|
|
@@ -0,0 +1,104 @@
|
|
|
1
|
+
# AAS Agent-First Control Plane v1 Worklog
|
|
2
|
+
|
|
3
|
+
- 2026-07-19: Semantic skill selection moved to Codex and Claude. Core now exposes the complete catalog and validates/pins exact agent-selected IDs through `compose_stack`; selection policy and metadata eligibility gates were retired. Every canonical skill must remain searchable, readable, selectable, and usable. Earlier recommendation entries below are historical.
|
|
4
|
+
|
|
5
|
+
## 2026-07-18 — Baseline workflow retired
|
|
6
|
+
|
|
7
|
+
- The standalone `aas-v1-baseline` pull-request workflow and required status check were retired by maintainer decision. The obsolete verifier corpus, harness, tuning runner, and paused apply/optimize workflows were subsequently removed. The protected `pr-policy`, `pr-evidence`, `source-validation`, and `artifact-preview` gates remain required.
|
|
8
|
+
- Entries below this point are a historical construction log. References to frozen verifier assets, matrices, paths, or pending certification gates describe the state at that date and are not current repository policy.
|
|
9
|
+
|
|
10
|
+
## 2026-07-17 — Goal activation and clean baseline
|
|
11
|
+
|
|
12
|
+
- Active objective is defined by the approved design and goal documents.
|
|
13
|
+
- Original worktree `/Users/nicco/Projects/antigravity-awesome-skills` contains unrelated local-skill-reviewer work and remains untouched.
|
|
14
|
+
- Created isolated worktree `/private/tmp/aas-agent-first-control-plane-v1` on branch `codex/aas-agent-first-control-plane-v1` from `origin/main` commit `4101f32402448f4fdd96b3cf166a81b8cee8b557`.
|
|
15
|
+
- Live source baseline: `origin/main` current; latest main CI, CodeQL, Actionlint, and Pages succeeded; no open Dependabot, code-scanning, secret-scanning, or runtime npm-audit findings.
|
|
16
|
+
- Open PRs 867 and 871 are unrelated skill contributions and remain out of scope.
|
|
17
|
+
- Copied the approved design and goal packet into the isolated branch.
|
|
18
|
+
|
|
19
|
+
## Phase 0 status
|
|
20
|
+
|
|
21
|
+
- Freeze-ready baseline complete: 11 public schemas, exact metric formulas, fixed 100,000 property/generative and 50,000 parser/MCP fuzz budgets, six exact OS/Node jobs, and zero pending requirements.
|
|
22
|
+
- Benchmark corpus: 180 held-out cases and accepted-equivalent gold sets, 60 disjoint tuning cases, and 30 disjoint explicit out-of-coverage abstention cases. The six intents each retain a denominator of 30 held-out cases.
|
|
23
|
+
- Independent review: `codex-independent-alpha` and `codex-independent-beta` each approved all 270 case/gold or case/label pairs. Their content-addressed reports are bound by `verification/aas-v1/ownership.v1.json`; both reviewers are independent of the future scorer implementation.
|
|
24
|
+
- Hostile corpus: 32 exploit/control classes, 64 hash-verified fixtures, zero extracted archives, and zero special filesystem entries.
|
|
25
|
+
- Registry legacy baseline: `agentic-awesome-skills@14.6.0`, SRI `sha512-VTOb3O9PSYKCDO99i3h0vOn7vHQlGtO/+jSErR80g6OGaDJoBzg3q2GE9Nu890en1/Z54hBEYiVQj/1Rl95xEg==`, tarball SHA-256 `98f8cbb399613621598ac6aeca619fc7c454530895b4e237eee695d82fbdf0cb`, tag commit `ab5f6c205a548d2f4bec411728c79b9c156fc696`.
|
|
26
|
+
- Legacy replay: 41 command cases pass with fixture tree `sha256-80c220b08a221685c26a23e2cd7c1b06628bfec7a542fa91d8ee19a5d3e035f8` and aggregate fake-Git trace `sha256-b36ea8cfd26d642235ce061bb5dcc925bec4b5bf63137cc17edb702b62560384`. Two consecutive complete replays produced the same 64-file corpus aggregate `3797c2c61c6334aa8081380be11958cd555d277baa26114e0cbe35fd62e099f6`.
|
|
27
|
+
- The legacy harness binds the exact dependency closure, runtime tree, and entrypoint; validates pre/post filesystem evidence; constrains case, target, and fake-Git paths; observes and denies Node networking with a sentinel self-test; and records zero network attempts. Full OS-syscall observation remains a separate black-box product acceptance gate and is not claimed by this corpus.
|
|
28
|
+
- Live `main` protection now requires `aas-v1-baseline`, while retaining `pr-policy`, `pr-evidence`, `source-validation`, and `artifact-preview`; admin enforcement remains enabled and force-pushes/deletions remain disabled.
|
|
29
|
+
- Local gate: 9/9 verifier tests, schema validation, structure validation, frozen benchmark and secondary corpora, hostile fixtures, legacy snapshots, and freeze readiness all pass.
|
|
30
|
+
- Content-addressed freeze manifest: 712 files, root digest `sha256-c7a4d4b3efa9f2bdf5a126fc3c384680d39534880a3900302785397e4ddd451c`; an immediate independent `freeze:check` reproduced it exactly. GitHub's Linux replay exposed zlib-version variation in valid DEFLATE streams; gzip fixtures now require the frozen compressed digest plus deterministic expanded USTAR bytes and preserve the canonical committed stream during regeneration.
|
|
31
|
+
|
|
32
|
+
## Feasibility decisions
|
|
33
|
+
|
|
34
|
+
1. A v1 mutating plan is single-target and single-filesystem. Multi-target manifests generate independently approved plans. This avoids claiming impossible crash-atomic commit across unrelated filesystems.
|
|
35
|
+
2. The pure-Node implementation uses atomic exclusive lock creation, file handles, type/ownership/realpath/identity revalidation, same-filesystem staging, journal, fsync, and atomic rename. It documents that portable Node lacks descriptor-relative `openat`/`renameat`; a malicious same-user process is already outside the frozen threat boundary. No native addon is added to v1.
|
|
36
|
+
3. Legacy `install.js` remains isolated. The new stack transaction engine will not reuse its per-entry mutation path.
|
|
37
|
+
|
|
38
|
+
## Next evidence gate
|
|
39
|
+
|
|
40
|
+
- Baseline landed through protected PR [#878](https://github.com/sickn33/agentic-awesome-skills/pull/878) as `09f2d8612d2e68a087377685e9c876bd737e1782`. Required `aas-v1-baseline`, source validation, and artifact preview passed from committed bytes.
|
|
41
|
+
- Product implementation now runs in isolated worktree `/private/tmp/aas-agent-first-product` on `codex/aas-agent-first-control-plane-v1-product`, based on that protected commit. `verification/aas-v1` remains unchanged.
|
|
42
|
+
|
|
43
|
+
## Phase 1 — Deterministic core and immutable state contracts
|
|
44
|
+
|
|
45
|
+
- Added a pure CommonJS deterministic core with canonical JSON, explicit version handshake, allowlisted input normalization, public metadata judgments, complete-catalog search, deterministic recommendation/discovery lanes, fixed-point lexical factors, and greedy set-cover composition with skill budget, overlap penalty, dependencies, and conflicts.
|
|
46
|
+
- Canonical registry identity uses `canonical_id`, exposing all 1,965 skills exactly once. Generated compatibility defaults are not treated as authoritative support evidence; absent reviewed metadata remains `unknown`.
|
|
47
|
+
- The first tuning-only metadata overlay covered the 85 candidate IDs referenced by tuning gold. Independent audit rejected that approach as an indirect gold-derived whitelist: tuning/gold review is not proof of skill quality or provenance. This overlay is development history only and must be replaced before product acceptance. No held-out input was read.
|
|
48
|
+
- Added 11 public v1 schemas, strict minimal-manifest validation, exact version handshake, and immutable single-target plan envelopes. Plans bind manifest, catalog, runtime, target identity, installed state, logical operations, overrides, and a final state commit without physical destination paths.
|
|
49
|
+
- Added a tuning-only diagnostic runner. After conservative removal of unsupported capability claims, the honest baseline was macro verified coverage `0.366667` and macro inclusion precision `0.911859`. Coverage by intent was `0.4/0.2/0.4/0.4/0.4/0.4`; this is a development diagnostic, not held-out evidence or a release claim.
|
|
50
|
+
- Corrected the compositor to be lexicographically coverage-first, forbid non-critical-only additions while critical goals remain uncovered, use exact versioned capability matching, and keep search-only ID tokens out of recommendation BM25. The same still-gold-derived development overlay then measured `0.366667` coverage and `0.928030` macro precision; it remains rejected as product evidence.
|
|
51
|
+
- Added a benchmark-independent public review queue over all 1,965 catalog entries and 120 versioned capabilities. Three independent semantic audits selected and reviewed 121 unique candidates across the six intents without reading tuning gold or held-out data. Field-level risk, provenance, setup, dependency, and conflict evidence is being audited separately; unknown values remain unknown.
|
|
52
|
+
|
|
53
|
+
## Phase 2 — Local runtime, CLI/MCP, transaction, and review UI slices
|
|
54
|
+
|
|
55
|
+
- Added content-addressed catalog and runtime caches, canonical identity records, npm `dist.integrity` verification, data-only catalog update, a hardened archive parser, and atomic cache promotion. The hostile archive matrix passes 34/34; updater/cache/archive targeted tests pass 44/44.
|
|
56
|
+
- A fresh current-worktree npm tarball was parsed and promoted as a runtime without lifecycle execution: 6,455 files, 99,652,053 expanded bytes, SRI `sha512-tulK12nYIjp0JNrwAUQgDnluU5X8BXuhzrJ5lBWeZvKBOA6ENs2/QuxKpgv4hR3hU/GaTQDdNSYTV7YLnJfERg==`, closure `sha256-5c8c616cea0c9fb5d62fc707647a71e554001bbea0dea915a4865bb92345f22c`; every cached byte verified.
|
|
57
|
+
- Added the `aas` lifecycle for catalog status/update, stack init/recommend/validate/plan/apply/doctor/recover, exact plan/recovery approvals, and legacy argument dispatch isolation. Added Codex/Claude host adapters plus explicit `mcp configure` preview/apply and backup cleanup; real user configuration remains untouched.
|
|
58
|
+
- Added the five-tool read-only local MCP server and resource template with strict bounded JSON-lines parsing, offline catalog resolution, no updater, no model, and no write surface. Unit/static evidence does not replace the required external syscall observation.
|
|
59
|
+
- Added transactional apply with same-filesystem staging, exclusive lock, chained journal, backup, state-last commit, idempotence, and explicit recovery. Unit fault hooks do not replace the required production-binary kill/swap/race matrix.
|
|
60
|
+
- Replaced the manual Workbench selector with in-memory paste/upload review of stack and plan artifacts. The route is isolated from catalog/Supabase fetch, renders text-only evidence, and exposes no install/apply/share/persistence surface. App tests pass 152/152 and the production build/prerender passes locally.
|
|
61
|
+
- Current integrated AAS v1 unit suite passed 87/87 before the latest runtime and metadata-hardening additions; a fresh full rerun is required after integration. Frozen `verification/aas-v1` remains unchanged from protected baseline commit `09f2d8612d2e68a087377685e9c876bd737e1782`.
|
|
62
|
+
|
|
63
|
+
## Phase 2 integration update — current product tree
|
|
64
|
+
|
|
65
|
+
- Replaced the rejected tuning-derived overlay with committed, benchmark-independent review sources. The pipeline deterministically validates 120 public capability queues, imports 121 independent metadata reviews, builds 121 overrides, and emits all 1,965 catalog records at catalog digest `sha256-cf14b1b22b826b5aed3f82f7ea3c8895b4b5ae1af372bd310b7154db80bb4628`. Dependency-path tests forbid frozen verification, held-out, and gold inputs. No held-out case has been read by the product implementation.
|
|
66
|
+
- Real Draft 2020-12 instance validation now covers recommendation input, catalog manifest, plans, recovery plans, journal records, managed state, doctor output, and CLI success/error envelopes. Recommendation validation occurs before normalization, so malformed booleans, targets, policy, and unknown fields fail closed.
|
|
67
|
+
- Runtime packaging now bundles the declared npm dependency closure and promotes only allowlisted runtime assets. The current commit-bound tarball is 42,307,013 bytes with SHA-256 `3a1478e84f02692313a72656041b0d114dcc57d1906c53f0da83243dfdc2c9f1` and npm integrity `sha512-T2kL318lXJBms8GeFIbeMcCileDWJ2gqCjpckQa4llWsWpA7jFdRNeuD2EfN0Ce/7K2LkBuFNDD0GdM2jegbXg==`; its 7,237 selected runtime assets total 100,632,903 bytes and bind closure digest `sha256-bb527d301fa696ba319ccd6b1a745410da1861b5f86f328ceb46e4ca6e085b5c`. The tarball is installed into an isolated content-addressed cache, reverified byte-for-byte, and launches `aas-mcp` successfully with `NODE_PATH` cleared. The published `aas` and `aas-mcp` bins are asserted executable on POSIX and smoke-tested from the installed tarball.
|
|
68
|
+
- Transaction recovery now uses a root-anchored WAL plus bootstrap evidence, crash-safe pending publication, immutable checkpoints, mutation intent/completion records, target locking, state-last commit, and retryable rollback/cleanup. Explicit fault tests cover bootstrap rename before fsync, layout publication before fsync, tombstone cleanup interruption, torn WAL tails, interrupted mutation recovery, post-commit cleanup, empty prior state, live locks, stale recovery guards, and symlink/containment drift.
|
|
69
|
+
- Windows-specific source paths reject NTFS streams/device aliases/reserved names, harden backup and staging ACLs before sensitive bytes, preserve existing config ACLs, and request a `GENERIC_WRITE` directory handle for `FlushFileBuffers`. These paths fail closed but are not claimed as executed evidence until the protected Windows matrix runs.
|
|
70
|
+
- Recommendation output now keeps full deterministic scoring while returning compact evidence summaries: at most 25 candidates per lane, 25 detailed exclusions, total/returned counts, reason-code counts, and at most eight content-addressed evidence references per candidate. It no longer embeds a second copy of itself. Across all 60 tuning cases, the largest serialized MCP tool result is 203,884 bytes. The frozen hostile regex-like search query is rejected before tokenization while normal technical queries such as `c++ node.js api/v1` remain valid.
|
|
71
|
+
- `mcp configure` now has two explicit and non-ambiguous runtime paths. Registry bootstrap/update continues to preview and bind the exact npm SRI before download. Supplying both `--runtime-integrity` and `--runtime-closure-digest` selects an already verified content-addressed runtime without any registry access; preview binds the complete identity and apply reverifies every cached byte immediately before the atomic host-config write. Supplying only one identity component, a missing runtime, a closure mismatch, or post-preview tampering fails closed without a config write.
|
|
72
|
+
- Current local gates: skill validation for 1,965 skills, docs-security checks, deterministic catalog checks, 128/128 AAS v1 tests, the complete repository test suite, 152/152 Workbench tests, production Vite build/prerender, package-content checks, isolated CLI/MCP tarball smoke, and `git diff --check` pass. This supersedes the earlier stale 87-test snapshot above.
|
|
73
|
+
- The canonical `repo-maintainer` skill now preserves the independent product/verifier/gold boundary. Its packaged mirrors remain owned by the protected canonical-sync flow and will be regenerated only after the source PR lands. The installed private `antigravity-maintainer-batch-release` skill and its UI metadata were also updated and validated so future sweeps retain the same frozen-matrix, black-box, and publication-approval gates.
|
|
74
|
+
- Honest tuning-only diagnostic remains below release thresholds: macro verified coverage `0.366667`, macro inclusion precision `0.563889`, critical-goal coverage `0.616667`, and non-critical-goal coverage `0.533333`, with zero hard-policy violations. Inspection showed semantically valid alternatives omitted from the current frozen tuning gold. Changing that gold is prohibited without a separate independent equivalence audit, explicit approval, protected PR, and baseline re-freeze.
|
|
75
|
+
|
|
76
|
+
## Remaining approval gates
|
|
77
|
+
|
|
78
|
+
- Explicit approval was granted for the independently owned product black-box verifier and six-job Node 22/24 acceptance workflow for Linux, macOS, and Windows. That work remains isolated from the product branch and must land through its own protected PR before the baseline is re-frozen.
|
|
79
|
+
- Explicit approval was granted for an independent tuning-gold equivalence audit. Its accepted equivalences, review evidence, and updated freeze remain isolated on their own branch and must land through a separate protected PR. Product code has not tuned around the known omissions or read held-out inputs.
|
|
80
|
+
- Only after those two gates pass may the clean-tarball held-out/abstention evaluator, 100,000 property cases, 50,000 parser/MCP fuzz cases, observed zero-network/zero-write MCP traces, production-binary fault/race matrix, legacy differential, host smoke tests, and exact evidence bundle be treated as release acceptance.
|
|
81
|
+
- Tag, npm publication, Pages deployment, or writes to real user MCP configuration remain separately approval-gated and have not occurred.
|
|
82
|
+
|
|
83
|
+
## 2026-07-17 — Additive agent-first preview profile
|
|
84
|
+
|
|
85
|
+
- The frozen certified-v1 design, benchmark, hostile corpus, verifier, and goal
|
|
86
|
+
remain unchanged. The new preview profile is an intermediate product-learning
|
|
87
|
+
gate and cannot complete the active certified-v1 goal.
|
|
88
|
+
- `stack apply` and `stack recover` are now disabled by default in the preview.
|
|
89
|
+
Controlled development requires `--experimental-apply` or
|
|
90
|
+
`--experimental-recovery`; successful writes are labelled experimental.
|
|
91
|
+
- Added a packed-product functional runner for `init -> recommend -> validate ->
|
|
92
|
+
plan -> doctor`, deterministic explanation output, all five local stdio MCP
|
|
93
|
+
tools, the skill resource template, read-only project/cache snapshots, legacy
|
|
94
|
+
isolation, and default write guards.
|
|
95
|
+
- Added a six-job Node 22/24 matrix for Linux, macOS, and Windows, plus Workbench
|
|
96
|
+
tests/build and a fail-closed receipt aggregator. A passing bundle must say
|
|
97
|
+
`previewQualified: true` and `certifiedV1: false` and must enumerate native
|
|
98
|
+
syscall observation, transactional crash/race certification, the 80/90/100
|
|
99
|
+
benchmark, real host writes, and public release as not evaluated.
|
|
100
|
+
- The preview contract and evidence bundle preserve this boundary. Any further
|
|
101
|
+
canonical maintainer-skill wording change remains a separate source/sync
|
|
102
|
+
maintenance action, so this product PR does not create registry drift.
|
|
103
|
+
- No npm package, GitHub release, Pages deployment, announcement, or real MCP
|
|
104
|
+
configuration write is authorized by this profile.
|