code-foundry 1.20.2 → 1.25.3
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.github/CONTRIBUTING.md +5 -2
- package/.github/code-foundry.yml +1 -0
- package/.github/release-please-foundry.json +56 -0
- package/.github/workflows/cloudflare-delivery.yml +9 -0
- package/.github/workflows/cloudflare-deploy.yml +8 -1
- package/.github/workflows/eval.yml +116 -0
- package/.github/workflows/opencode-security_self-ci.yml +1 -1
- package/.github/workflows/qualified-foundry-publish.yml +6 -7
- package/.github/workflows/release.yml +58 -10
- package/.github/workflows/release_self-ci.yml +211 -8
- package/.github/workflows/validation-no-codeql.yml +34 -6
- package/.github/workflows/validation.yml +32 -5
- package/.github/workflows/validation_audit_self-ci.yml +1 -0
- package/.github/workflows/validation_self-ci.yml +1 -0
- package/.gitignore +1 -1
- package/AGENTS.md +5 -1
- package/CHANGELOG.md +95 -0
- package/README.md +23 -17
- package/docs/CONFIGURATION.md +189 -152
- package/docs/EVALS.md +139 -0
- package/docs/EXTENSIONS.md +28 -7
- package/docs/INITIALIZATION.md +13 -9
- package/docs/PERFORMANCE.md +67 -59
- package/docs/PUBLISHING.md +39 -14
- package/docs/README.md +37 -22
- package/docs/RELEASES.md +18 -9
- package/docs/WORKFLOWS.md +23 -4
- package/docs/agent-validation.md +6 -5
- package/docs/cloudflare-delivery.md +20 -8
- package/docs/consumer-qualification.md +11 -9
- package/docs/fleet-release-eligibility.md +14 -15
- package/docs/fleet-rollouts.md +3 -3
- package/docs/merge-queues.md +21 -21
- package/docs/product-quality.md +9 -10
- package/docs/qualified-publication.md +166 -93
- package/docs/release-integrity.md +7 -6
- package/docs/required-capabilities.md +24 -15
- package/package.json +1 -1
- package/src/commands/cloudflare-delivery.mjs +6 -1
- package/src/commands/qualified-publication.mjs +103 -12
- package/src/commands/release-integrity.mjs +46 -14
- package/src/commands/sync.mjs +30 -24
- package/src/lib/eval-envelope.mjs +234 -0
- package/src/lib/merge-queue.mjs +1 -0
- package/src/lib/product-quality.mjs +220 -16
- package/src/lib/task-policy.mjs +7 -0
- package/src/lib/validation-policy.mjs +10 -6
- package/src/runtime-core.mjs +166 -8
- package/src/runtime.mjs +3 -1
- package/src/templates/gitignore +3 -1
package/docs/CONFIGURATION.md
CHANGED
|
@@ -1,19 +1,20 @@
|
|
|
1
1
|
# Configuration reference
|
|
2
2
|
|
|
3
|
-
Code Foundry has one repository-owned control plane:
|
|
3
|
+
Code Foundry has one repository-owned control plane:
|
|
4
|
+
`.github/code-foundry.yml`. `init` creates it, `sync` renders the selected
|
|
5
|
+
baseline, and `doctor` checks local and GitHub-facing prerequisites.
|
|
4
6
|
|
|
5
|
-
```
|
|
7
|
+
```sh
|
|
6
8
|
npx code-foundry init
|
|
9
|
+
# edit .github/code-foundry.yml
|
|
10
|
+
npx code-foundry sync
|
|
11
|
+
npx code-foundry doctor
|
|
7
12
|
```
|
|
8
13
|
|
|
9
|
-
|
|
10
|
-
|
|
14
|
+
The generated file is deliberately explicit. Keep it under version control and
|
|
15
|
+
change it directly rather than passing one-off flags to `sync`.
|
|
11
16
|
|
|
12
|
-
|
|
13
|
-
selects native setup for the detected languages. Use `toolchain: native` to
|
|
14
|
-
prohibit mise or `toolchain: mise` to require it.
|
|
15
|
-
|
|
16
|
-
## Configuration flow
|
|
17
|
+
## How configuration is applied
|
|
17
18
|
|
|
18
19
|
```text
|
|
19
20
|
repository manifests and source
|
|
@@ -21,181 +22,217 @@ repository manifests and source
|
|
|
21
22
|
v
|
|
22
23
|
.github/code-foundry.yml
|
|
23
24
|
|
|
|
24
|
-
+-->
|
|
25
|
+
+--> detected language and package-manager setup
|
|
25
26
|
+--> standard workflow callers
|
|
26
27
|
+--> runtime repository and version
|
|
27
|
-
+--> release, license,
|
|
28
|
+
+--> validation, release, license, and cache policy
|
|
28
29
|
```
|
|
29
30
|
|
|
30
|
-
|
|
31
|
-
|
|
32
|
-
|
|
33
|
-
|
|
34
|
-
|
|
35
|
-
|
|
36
|
-
|
|
|
37
|
-
|
|
|
38
|
-
| `
|
|
39
|
-
| `
|
|
40
|
-
| `
|
|
41
|
-
| `
|
|
42
|
-
| `
|
|
43
|
-
| `
|
|
44
|
-
| `
|
|
45
|
-
| `
|
|
46
|
-
| `
|
|
47
|
-
| `
|
|
48
|
-
| `
|
|
49
|
-
| `
|
|
50
|
-
|
|
51
|
-
|
|
52
|
-
|
|
53
|
-
|
|
54
|
-
|
|
55
|
-
|
|
56
|
-
|
|
57
|
-
|
|
58
|
-
|
|
59
|
-
|
|
60
|
-
|
|
61
|
-
|
|
62
|
-
|
|
63
|
-
|
|
64
|
-
|
|
65
|
-
|
|
66
|
-
`
|
|
67
|
-
|
|
68
|
-
|
|
69
|
-
|
|
70
|
-
|
|
71
|
-
|
|
72
|
-
|
|
73
|
-
|
|
31
|
+
`toolchain: auto` reuses an existing `.mise.toml`; otherwise it uses native
|
|
32
|
+
language tooling. Set `toolchain: native` to prohibit mise or `toolchain: mise`
|
|
33
|
+
to require an existing mise configuration.
|
|
34
|
+
|
|
35
|
+
## Repository and runtime
|
|
36
|
+
|
|
37
|
+
| Key | Values | Notes |
|
|
38
|
+
| ----------------------- | -------------------------------------------- | -------------------------------------------------------------------- |
|
|
39
|
+
| `version` | `1` | Configuration schema version. |
|
|
40
|
+
| `profile` | `auto`, `application`, `monorepo`, `minimal` | Repository shape; `auto` detects it. |
|
|
41
|
+
| `languages` | comma-separated language names | Supported values are `typescript`, `rust`, `python`, and `solidity`. |
|
|
42
|
+
| `package_manager` | `bun`, `pnpm`, `yarn`, `npm`, `none` | JavaScript package-manager policy. |
|
|
43
|
+
| `toolchain` | `auto`, `native`, `mise` | Environment setup policy. |
|
|
44
|
+
| `runtime_repository` | `OWNER/REPO` | Source of reusable workflows and runtime code. |
|
|
45
|
+
| `runtime_ref` | tag or commit | Runtime version used by generated callers. |
|
|
46
|
+
| `features` | `all` or a list | See [Feature selection](#feature-selection). |
|
|
47
|
+
| `draft_protection` | `true`, `false` | Skip generated runner-heavy gates for draft PRs when false. |
|
|
48
|
+
| `codeql` | `auto`, `true`, `false` | Enable CodeQL when the repository and GitHub plan support it. |
|
|
49
|
+
| `dependency_review` | `auto`, `true`, `false` | Enable Dependency Review when supported. |
|
|
50
|
+
| `runner` and `*_runner` | GitHub runner labels | Override the default runner per workflow. |
|
|
51
|
+
|
|
52
|
+
For `codeql: auto` and `dependency_review: auto`, public repositories use the
|
|
53
|
+
available GitHub security checks and private repositories require the relevant
|
|
54
|
+
capability. Set either key to `false` when the check is unavailable or not
|
|
55
|
+
wanted. CodeQL is omitted from the generated validation caller; Dependency
|
|
56
|
+
Review remains a conditional step inside Security rather than a separate check.
|
|
57
|
+
|
|
58
|
+
Code Foundry runner-heavy validation, security, qualification, and Cloudflare
|
|
59
|
+
Deployment jobs protect draft pull requests by default. The `draft_protection`
|
|
60
|
+
configuration key defaults to `true`; set it to `false` only when the repository
|
|
61
|
+
intentionally runs generated gates for draft PRs. Cloudflare reusable-workflow
|
|
62
|
+
callers use their equivalent `draft-protection` input. These opt-outs affect
|
|
63
|
+
CI/deployment gates only and do not disable Draft Guard or draft-PR automation.
|
|
64
|
+
|
|
65
|
+
## Feature selection
|
|
66
|
+
|
|
67
|
+
Use `features: all` or a comma/space-separated list. The canonical validation
|
|
68
|
+
feature is `validation`; the legacy names `ci`, `test`, `security`, and `codeql`
|
|
69
|
+
remain aliases for compatibility. Other selectable features are:
|
|
70
|
+
|
|
71
|
+
- `draft-pr` — create or update development pull requests.
|
|
72
|
+
- `release-pr` — promote `staging` into `main` in the `staging-release` topology.
|
|
73
|
+
- `release` — run Release Please and optional package publication.
|
|
74
|
+
- `dependabot` — install the language-aware Dependabot configuration.
|
|
75
|
+
|
|
76
|
+
The release-integrity and OpenCode Security callers are installed independently
|
|
77
|
+
of feature selection. OpenCode Security is disabled unless the repository or
|
|
78
|
+
organization variable `OPENCODE_SECURITY` is `true` and the
|
|
79
|
+
`OPENCODE_API_KEY` secret exists. `opencode_security_model` optionally replaces
|
|
80
|
+
the generated scanner model.
|
|
81
|
+
|
|
82
|
+
Merge queues are a separate opt-in because they need a stable runtime pin:
|
|
74
83
|
|
|
75
84
|
```yaml
|
|
76
|
-
|
|
77
|
-
performance_command: '["python3","scripts/performance_audit.py","--check"]'
|
|
78
|
-
performance_runner: ubuntu-latest
|
|
85
|
+
merge_queue: true
|
|
79
86
|
```
|
|
80
87
|
|
|
81
|
-
|
|
88
|
+
See [Merge queue validation](merge-queues.md) before enabling it.
|
|
89
|
+
|
|
90
|
+
## Validation and quality
|
|
91
|
+
|
|
92
|
+
| Key | Values | Purpose |
|
|
93
|
+
| ------------------------- | ---------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
94
|
+
| `performance` | `auto`, `true`, `false` | Discover, require, or disable performance checks. |
|
|
95
|
+
| `performance_command` | JSON argv array or array of argv arrays | Ordered commands for a non-package performance harness. |
|
|
96
|
+
| `performance_profile` | empty or `node-package` | Shared package import, memory, archive, and dependency audit. |
|
|
97
|
+
| `performance_budget_file` | repository-relative path | Budget file for `node-package`; defaults to `performance-package-budgets.json`. |
|
|
98
|
+
| `eval` | `auto`, `true`, `false` | Discover, require, or disable the deterministic eval tier. |
|
|
99
|
+
| `eval_command` | JSON argv array | Explicit harness command when there is no `eval` package script. |
|
|
100
|
+
| `eval_report_file` | repository-relative path | Report validated against the eval contract; defaults to `eval-results/result.json`. |
|
|
101
|
+
| `eval_budget_file` | repository-relative path | Optional budget file; defaults to `eval-budgets.json`. |
|
|
102
|
+
| `eval_runner` | runner label | Runner for the `Validation / Eval` lane; defaults to the repository runner. Browser evals need a Chrome-capable runner such as `ubuntu-latest`. |
|
|
103
|
+
| `required_capabilities` | comma-separated task names | Fail closed when a required task or coverage evidence is unavailable. |
|
|
104
|
+
| `coverage_enforcement` | `auto`, `required`, `off` | Shared coverage-report policy. |
|
|
105
|
+
| `coverage_minimum` | `0`–`100` | Minimum percentage; defaults to `80`. |
|
|
106
|
+
| `coverage_metrics` | `lines`, `functions`, `branches`, `statements` | Metrics checked by the coverage gate. |
|
|
107
|
+
| `coverage_report` | comma-separated repository paths | Istanbul JSON summary or LCOV evidence files. |
|
|
108
|
+
|
|
109
|
+
Supported task capabilities are `format`, `lint`, `type_check`, `build`, `unit`,
|
|
110
|
+
`integration`, `e2e`, `smoke`, `eval`, and `performance`. `coverage` is a policy
|
|
111
|
+
capability that also requires unit tests. See [Required capabilities and task
|
|
112
|
+
evidence](required-capabilities.md). The eval tier runs the repository's own
|
|
113
|
+
harness against the shared report contract and optional budgets; see
|
|
114
|
+
[Evals](EVALS.md).
|
|
115
|
+
|
|
116
|
+
The shared performance job discovers `performance:check`, then `perf:check`,
|
|
117
|
+
in JavaScript repositories. Other repositories can provide one command or an
|
|
118
|
+
ordered list of argv arrays without shell interpolation:
|
|
82
119
|
|
|
83
120
|
```yaml
|
|
84
121
|
performance: true
|
|
85
|
-
|
|
86
|
-
|
|
87
|
-
```
|
|
88
|
-
|
|
89
|
-
```json
|
|
90
|
-
{
|
|
91
|
-
"schemaVersion": 1,
|
|
92
|
-
"importTarget": "./dist/index.js",
|
|
93
|
-
"controlImport": "typebox",
|
|
94
|
-
"samples": 7,
|
|
95
|
-
"budgets": {
|
|
96
|
-
"coldImportP95Ms": 150,
|
|
97
|
-
"coldImportRssMaxBytes": 25000000,
|
|
98
|
-
"packedBytes": 500000,
|
|
99
|
-
"productionDependencyCount": 20
|
|
100
|
-
}
|
|
101
|
-
}
|
|
122
|
+
performance_command: '["python3", "scripts/performance_audit.py", "--check"]'
|
|
123
|
+
performance_runner: ubuntu-latest
|
|
102
124
|
```
|
|
103
125
|
|
|
104
|
-
|
|
105
|
-
|
|
106
|
-
`
|
|
107
|
-
`
|
|
108
|
-
`
|
|
109
|
-
`performance-results
|
|
110
|
-
|
|
111
|
-
|
|
112
|
-
|
|
113
|
-
|
|
114
|
-
|
|
115
|
-
|
|
116
|
-
|
|
117
|
-
|
|
118
|
-
|
|
119
|
-
|
|
120
|
-
|
|
121
|
-
|
|
122
|
-
`
|
|
123
|
-
|
|
124
|
-
|
|
125
|
-
|
|
126
|
-
|
|
127
|
-
|
|
128
|
-
|
|
129
|
-
|
|
130
|
-
|
|
131
|
-
|
|
132
|
-
|
|
133
|
-
|
|
134
|
-
`git_workflow` selects the branch topology:
|
|
135
|
-
|
|
136
|
-
- `direct` (default): feature branches open pull requests directly into
|
|
137
|
-
`main`. Validation and security scans run on every PR. No `staging` branch
|
|
138
|
-
exists, no promotion caller is generated, and `merge_strategy` must be
|
|
139
|
-
`squash`. Release Please version PRs squash into `main`
|
|
140
|
-
(`release_merge_strategy: squash`). Dependabot updates target `main`. This is
|
|
141
|
-
the right choice when a repository has no preview or staging environment.
|
|
142
|
-
- `staging-release` (opt-in): feature branches squash into `staging`, a
|
|
143
|
-
promotion PR rebases validated changes into `main` (`merge_strategy:
|
|
144
|
-
rebase`), and Release Please version PRs rebase into `main`
|
|
145
|
-
(`release_merge_strategy: rebase`). Choose this only when the repository
|
|
146
|
-
maintains a preview/staging environment that needs validated integration
|
|
147
|
-
before release.
|
|
126
|
+
The `node-package` profile supports cold-import, memory, package-size, file-count,
|
|
127
|
+
and production-dependency budgets. Supported budget names are
|
|
128
|
+
`coldImportP50Ms`, `coldImportP95Ms`, `coldImportRssMaxBytes`,
|
|
129
|
+
`coldImportRelativeP50`, `packedBytes`, `unpackedBytes`, `packageFileCount`,
|
|
130
|
+
`packageMapFileCount`, and `productionDependencyCount`. Reports are written under
|
|
131
|
+
`performance-results/` and are uploaded when present.
|
|
132
|
+
|
|
133
|
+
`performance: true` makes the performance task required. Use `performance: auto`
|
|
134
|
+
to keep discovery optional. Product-quality profiles are repository-owned
|
|
135
|
+
manifests invoked by existing build or E2E commands; they are not activated by a
|
|
136
|
+
configuration key. See [Product quality profiles](product-quality.md).
|
|
137
|
+
|
|
138
|
+
## Release and branch policy
|
|
139
|
+
|
|
140
|
+
| Key | Values | Purpose |
|
|
141
|
+
| ------------------------- | -------------------------------------------------------------------------------- | ----------------------------------------------------------------------- |
|
|
142
|
+
| `release_type` | `auto`, `node`, `python`, `rust`, `simple`, `none` | Select a release manifest; `auto` detects one. |
|
|
143
|
+
| `npm_publish` | `true`, `false` | Opt into npm publication. |
|
|
144
|
+
| `license` | `gpl-3.0-or-later`, `agpl-3.0-or-later`, `apache-2.0`, `mit`, `preserve`, `none` | License policy for initialized repositories. |
|
|
145
|
+
| `git_workflow` | `direct`, `staging-release` | Choose the branch topology. |
|
|
146
|
+
| `merge_strategy` | `squash` or `rebase` | Required topology-specific merge method. |
|
|
147
|
+
| `release_merge_strategy` | `squash` or `rebase` | Required method for Release Please version PRs. |
|
|
148
|
+
| `staging_validation_mode` | `fast`, `audit` | Validation tier for pull requests into `staging`; staging-release only. |
|
|
149
|
+
|
|
150
|
+
`direct` is the default: feature branches and release PRs target `main`, and
|
|
151
|
+
both merge with squash. `staging-release` sends feature branches to `staging`,
|
|
152
|
+
uses rebase for the `staging` → `main` promotion and Release Please PR, and
|
|
153
|
+
keeps feature PRs into `staging` on squash. `sync`, `doctor`, and release
|
|
154
|
+
automation reject a strategy that does not match the selected topology.
|
|
148
155
|
|
|
149
156
|
```yaml
|
|
150
|
-
#
|
|
157
|
+
# Preview/staging environment
|
|
158
|
+
release_type: auto
|
|
151
159
|
git_workflow: staging-release
|
|
160
|
+
staging_validation_mode: fast
|
|
161
|
+
merge_strategy: rebase
|
|
162
|
+
release_merge_strategy: rebase
|
|
152
163
|
```
|
|
153
164
|
|
|
154
|
-
|
|
165
|
+
Use `simple` with `version.txt` when no package manifest exists. Use `none` to
|
|
166
|
+
skip automated releases. `npm_publish` affects generated consumer release
|
|
167
|
+
callers; Code Foundry's own repository uses the qualified publication path
|
|
168
|
+
described in [Qualified publication](qualified-publication.md).
|
|
169
|
+
|
|
170
|
+
## Synchronization and extensions
|
|
171
|
+
|
|
172
|
+
| Key | Values | Purpose |
|
|
173
|
+
| ----------------------- | -------------------------------------------------------- | ----------------------------------------------------------------- |
|
|
174
|
+
| `sync_mode` | `overlay`, `strict` | Synchronization policy; `overlay` is the default. |
|
|
175
|
+
| `custom_workflows` | `preserve` | Custom workflows are always preserved; other values are rejected. |
|
|
176
|
+
| `post_release` | `true`, `auto`, `false` | Enable a post-release delivery hook. |
|
|
177
|
+
| `post_release_workflow` | workflow filename | Workflow dispatched by the post-release hook. |
|
|
178
|
+
| `post_release_mode` | `auto`, `workflow-dispatch`, `release-event`, `disabled` | Select the hook delivery mechanism. |
|
|
179
|
+
|
|
180
|
+
`sync_mode` accepts `overlay` (the default) or `strict`; `sync` validates the
|
|
181
|
+
selected value before writing. Custom workflows remain preserved in either mode,
|
|
182
|
+
and `custom_workflows` must remain `preserve`. See [Extension points](EXTENSIONS.md).
|
|
183
|
+
|
|
184
|
+
## Caching and remote caching
|
|
185
|
+
|
|
186
|
+
The standard workflows use lockfile- and configuration-keyed caches. Their
|
|
187
|
+
repository variables, rather than application source files, control cache
|
|
188
|
+
behavior:
|
|
155
189
|
|
|
156
|
-
|
|
190
|
+
- `REPO_FOUNDRY_CACHE_PACKAGES` controls package-store caching.
|
|
191
|
+
- `REPO_FOUNDRY_CACHE_BUILD` controls build-cache reuse.
|
|
192
|
+
- `turbo_remote: auto`, `true`, or `false` declares the remote-cache policy;
|
|
193
|
+
`doctor --github` warns when enabled remote caching lacks `TURBO_TOKEN` or
|
|
194
|
+
`TURBO_TEAM`.
|
|
195
|
+
- `TURBO_TOKEN` and `TURBO_TEAM` provide the Turborepo remote-cache
|
|
196
|
+
credentials.
|
|
157
197
|
|
|
158
|
-
|
|
159
|
-
|
|
160
|
-
short and replaceable; custom workflows and project documentation are kept.
|
|
198
|
+
Use these controls only after measuring a repeatable benefit. See [Caching and
|
|
199
|
+
remote caching](CACHING.md).
|
|
161
200
|
|
|
162
|
-
|
|
163
|
-
understand the repository without memorizing flags or environment variables.
|
|
201
|
+
## Rust CodeQL tuning
|
|
164
202
|
|
|
165
|
-
Rust CodeQL defaults to one full scan with one worker.
|
|
166
|
-
repositories can opt into bounded parallelism
|
|
203
|
+
Rust CodeQL defaults to one full scan with one worker. Larger multi-crate
|
|
204
|
+
repositories can opt into bounded parallelism:
|
|
167
205
|
|
|
168
206
|
```yaml
|
|
169
|
-
codeql_rust_shards: '["crates/api","crates/worker"]'
|
|
207
|
+
codeql_rust_shards: '["crates/api", "crates/worker"]'
|
|
170
208
|
codeql_rust_threads: 2
|
|
171
209
|
codeql_rust_max_parallel: 2
|
|
172
210
|
```
|
|
173
211
|
|
|
174
|
-
Each
|
|
175
|
-
|
|
176
|
-
shards. Do not split a single crate by arbitrary non-Rust directories: use
|
|
212
|
+
Each shard must contain tracked Rust source. Absolute paths, parent traversal,
|
|
213
|
+
duplicates, empty scopes, and more than eight shards are rejected. Use
|
|
177
214
|
`["all"]` when complete, non-overlapping source scopes are not available.
|
|
178
215
|
|
|
179
|
-
## Cloudflare Workers
|
|
216
|
+
## Cloudflare Workers
|
|
180
217
|
|
|
181
|
-
Repositories that deploy to Cloudflare Workers can
|
|
182
|
-
|
|
183
|
-
verification, and version-identity promotion
|
|
184
|
-
|
|
185
|
-
|
|
186
|
-
|
|
187
|
-
`Production` environment. The workflow requires the
|
|
188
|
-
`CLOUDFLARE_API_TOKEN` and `CLOUDFLARE_ACCOUNT_ID` secrets in the consumer
|
|
189
|
-
repository. See [Verified Cloudflare delivery](./cloudflare-delivery.md) for
|
|
190
|
-
binding policy, canary, rollback, and evidence requirements.
|
|
218
|
+
Repositories that deploy to Cloudflare Workers can use the opt-in verified
|
|
219
|
+
delivery workflow with fixed `Preview` and `Production` environments,
|
|
220
|
+
candidate verification, and version-identity promotion. Pin both the reusable
|
|
221
|
+
workflow reference and `runtime-ref` to the same reviewed 40-character commit
|
|
222
|
+
SHA. Provide `CLOUDFLARE_API_TOKEN` and `CLOUDFLARE_ACCOUNT_ID` in the consumer
|
|
223
|
+
repository and configure environment reviewers separately.
|
|
191
224
|
|
|
192
225
|
The legacy `cloudflare-deploy.yml` workflow remains available for direct
|
|
193
226
|
(unverified) deployments. It runs `wrangler versions upload` for previews and
|
|
194
227
|
`wrangler deploy` for production, records a GitHub deployment plus status, and
|
|
195
|
-
respects `CI_BILLING_PAUSED`.
|
|
196
|
-
|
|
197
|
-
|
|
198
|
-
|
|
199
|
-
|
|
200
|
-
|
|
201
|
-
|
|
228
|
+
respects `CI_BILLING_PAUSED`. Preview deployment records use the pull request
|
|
229
|
+
head SHA when called from a PR, which lets GitHub show the completed preview in
|
|
230
|
+
the PR's Deployments section; direct pushes use the workflow SHA. Its
|
|
231
|
+
legacy-compatible Wrangler default is `latest`; callers should prefer `local` or
|
|
232
|
+
provide an exact `wrangler-version` for reproducibility. Bun consumers may pass
|
|
233
|
+
`build-script`, `install-working-directory`, and `bun-version`; the runtime
|
|
234
|
+
installs the frozen lockfile and builds the Worker before invoking Wrangler.
|
|
235
|
+
Bun-backed callers invoke Wrangler through `bunx` so OpenNext's production
|
|
236
|
+
delegation resolves the workspace-local `opennextjs-cloudflare` binary; callers
|
|
237
|
+
without `build-script` retain the npm/npx path. See [Verified Cloudflare
|
|
238
|
+
delivery](cloudflare-delivery.md).
|
package/docs/EVALS.md
ADDED
|
@@ -0,0 +1,139 @@
|
|
|
1
|
+
# Evals
|
|
2
|
+
|
|
3
|
+
`ci eval` is the repository's deterministic behavior-evaluation tier. It runs the
|
|
4
|
+
repository's own eval harness and consumes its result against a shared contract,
|
|
5
|
+
so task outcomes stay comparable across revisions and — later — across
|
|
6
|
+
executors (deterministic today, model-agent when a repository adopts one).
|
|
7
|
+
|
|
8
|
+
Evals are tests' measurement-oriented sibling: tests verify a specified
|
|
9
|
+
behavior (pass/fail); evals record how well the system does at repeated probes
|
|
10
|
+
(success rates, timing percentiles, bounded evidence) and compare those numbers
|
|
11
|
+
against pinned baselines. The `performance` tier is the resource-metric
|
|
12
|
+
special case of the same pattern.
|
|
13
|
+
|
|
14
|
+
## How the runtime finds your harness
|
|
15
|
+
|
|
16
|
+
`ci eval` discovers its subject in this order and skips when neither exists:
|
|
17
|
+
|
|
18
|
+
1. A package script named `eval`.
|
|
19
|
+
2. An explicit `eval_command` (JSON argv array) in `.github/code-foundry.yml`.
|
|
20
|
+
|
|
21
|
+
Both mechanisms use `eval: auto` by default: the tier runs when a subject is
|
|
22
|
+
present and skips cleanly otherwise. Set `eval: true` to require it (the
|
|
23
|
+
validation policy then treats a missing subject as an error), or `eval: false`
|
|
24
|
+
to disable discovery.
|
|
25
|
+
|
|
26
|
+
In managed validation the eval tier runs as its own `Validation / Eval` lane
|
|
27
|
+
during audit-mode runs, with the receipt retained as a task artifact. Browser
|
|
28
|
+
eval harnesses need a Chrome-capable runner: set `eval_runner: ubuntu-latest`
|
|
29
|
+
in `.github/code-foundry.yml` (the default when unset is the repository's
|
|
30
|
+
default runner, which may not ship a browser). The lane uploads the eval
|
|
31
|
+
report alongside the task receipt whenever a run fails, so reviewers get the
|
|
32
|
+
measured numbers with the red check. Performance's task profile and gating
|
|
33
|
+
rules apply unchanged; evals are an optional, non-gating tier like
|
|
34
|
+
performance.
|
|
35
|
+
|
|
36
|
+
| Config key | Values | Meaning |
|
|
37
|
+
| ------------------ | --------------------------------- | ---------------------------------------------------------------- |
|
|
38
|
+
| `eval` | `auto` (default), `true`, `false` | Enable, require, or disable the eval tier. |
|
|
39
|
+
| `eval_command` | JSON argv array | Explicit harness command when there is no `eval` package script. |
|
|
40
|
+
| `eval_report_file` | repository-relative path | The report to validate; defaults to `eval-results/result.json`. |
|
|
41
|
+
| `eval_budget_file` | repository-relative path | Optional budget file; defaults to `eval-budgets.json`. |
|
|
42
|
+
|
|
43
|
+
## The report contract
|
|
44
|
+
|
|
45
|
+
Your harness writes `eval-results/result.json`. The runtime validates the
|
|
46
|
+
envelope before budgets; a report that violates the contract fails the tier
|
|
47
|
+
with every violation listed (fail closed, typos can never pass silently):
|
|
48
|
+
|
|
49
|
+
```jsonc
|
|
50
|
+
{
|
|
51
|
+
"schemaVersion": 1, // contract version, currently 1
|
|
52
|
+
"revision": "<git sha>", // optional but strongly recommended
|
|
53
|
+
"dependencyHash": "<sha256>", // optional; pins the dependency set
|
|
54
|
+
"summary": {
|
|
55
|
+
"taskCount": 7,
|
|
56
|
+
"attempts": 7,
|
|
57
|
+
"passed": 7,
|
|
58
|
+
"failed": 0,
|
|
59
|
+
"harnessFailures": 0, // environment broke; not a task regression
|
|
60
|
+
"successRate": 1.0,
|
|
61
|
+
"toolCalls": 31,
|
|
62
|
+
"evidenceErrors": 0,
|
|
63
|
+
"taskDurationMs": { "count": 7, "mean": 2.1, "p50": 1.8, "p95": 3.0, "max": 3.4 },
|
|
64
|
+
"startupMs": { "count": 7, "mean": 0.3, "p50": 0.3, "p95": 0.4, "max": 0.4 },
|
|
65
|
+
"stepDurationMs": { "count": 31, "mean": 0.2, "p50": 0.1, "p95": 0.6, "max": 0.6 },
|
|
66
|
+
},
|
|
67
|
+
"tasks": [
|
|
68
|
+
{
|
|
69
|
+
"id": "form-submit",
|
|
70
|
+
"attempts": [
|
|
71
|
+
{
|
|
72
|
+
"iteration": 1,
|
|
73
|
+
"status": "passed",
|
|
74
|
+
"durationMs": 2.1,
|
|
75
|
+
"startupMs": 0.3,
|
|
76
|
+
"steps": [{ "tool": "fill_form", "status": "passed", "durationMs": 0.4 }],
|
|
77
|
+
"checks": ["fill_form completed"],
|
|
78
|
+
"metrics": { "formFields": 2 },
|
|
79
|
+
},
|
|
80
|
+
],
|
|
81
|
+
},
|
|
82
|
+
],
|
|
83
|
+
}
|
|
84
|
+
```
|
|
85
|
+
|
|
86
|
+
Rules the validator enforces:
|
|
87
|
+
|
|
88
|
+
- `schemaVersion` must match the contract version; `revision`/`dependencyHash`
|
|
89
|
+
must be strings when present.
|
|
90
|
+
- Summary counts are non-negative integers; `passed + failed` cannot exceed
|
|
91
|
+
`attempts`; `successRate` is between 0 and 1.
|
|
92
|
+
- `harnessFailures` counts attempts where the environment itself broke (browser
|
|
93
|
+
never started, zero steps succeeded, cleanup failed). It must not exceed
|
|
94
|
+
`attempts`. Fix the environment; do not treat these as task regressions.
|
|
95
|
+
- Failed attempts carry a bounded `failure` (non-empty `message`) plus a
|
|
96
|
+
`failureClass` of `harness` or `task`.
|
|
97
|
+
- `taskDurationMs`, `startupMs`, and `stepDurationMs` are stats objects with
|
|
98
|
+
`count`, `mean`, `p50`, `p95`, `max` (or bare `{ "count": 0 }`).
|
|
99
|
+
|
|
100
|
+
A reference implementation ships in
|
|
101
|
+
[`pi-browser-use`](https://github.com/0xPlayerOne/pi-browser-use)
|
|
102
|
+
(`scripts/eval.mjs` + `docs/eval-results.md`): deterministic browser-tool tasks
|
|
103
|
+
that emit this exact envelope.
|
|
104
|
+
|
|
105
|
+
## Budgets
|
|
106
|
+
|
|
107
|
+
Create `eval-budgets.json` (or point `eval_budget_file` at another file) to
|
|
108
|
+
gate on the measured numbers:
|
|
109
|
+
|
|
110
|
+
```json
|
|
111
|
+
{
|
|
112
|
+
"successRate": 1.0,
|
|
113
|
+
"stepP95Ms": 1000,
|
|
114
|
+
"taskP95Ms": 45000,
|
|
115
|
+
"maxHarnessFailures": 0
|
|
116
|
+
}
|
|
117
|
+
```
|
|
118
|
+
|
|
119
|
+
Supported budgets: `successRate`, `taskP95Ms`, `startupP95Ms`, `stepP95Ms`,
|
|
120
|
+
`maxHarnessFailures`, `maxEvidenceErrors`, `maxToolCalls`. Unknown keys fail
|
|
121
|
+
closed. When the budget file is absent the tier validates the contract only.
|
|
122
|
+
Budget files are committed configuration; only `eval-results/` is a local
|
|
123
|
+
artifact.
|
|
124
|
+
|
|
125
|
+
The runtime writes `eval-results/summary.json` with the executed commands, the
|
|
126
|
+
budget outcome, and the artifact list, mirroring the performance summary.
|
|
127
|
+
|
|
128
|
+
## Determinism rules for eval tasks
|
|
129
|
+
|
|
130
|
+
- Fixed fixtures and no network. A task's outcome must depend only on the code
|
|
131
|
+
under test.
|
|
132
|
+
- Bounded evidence: failure text and screenshots are capped; a failing task
|
|
133
|
+
never produces unbounded output.
|
|
134
|
+
- Cleanup failures are surfaced, never swallowed silently.
|
|
135
|
+
- Every attempt records the revision and dependency hash it ran against.
|
|
136
|
+
- Keep model-in-the-loop runs out of this tier. Reuse the same task IDs in a
|
|
137
|
+
separate scheduled harness with frozen model/reasoning/judge baselines when
|
|
138
|
+
you need agent-capability measurement; its variance is why it must never
|
|
139
|
+
gate pull requests.
|
package/docs/EXTENSIONS.md
CHANGED
|
@@ -1,19 +1,40 @@
|
|
|
1
1
|
# Code Foundry extension points
|
|
2
2
|
|
|
3
|
-
Code Foundry uses an overlay model
|
|
3
|
+
Code Foundry uses an overlay model: `sync` refreshes the documented baseline while
|
|
4
|
+
repository-owned behavior stays in separate files.
|
|
4
5
|
|
|
5
6
|
## Managed files
|
|
6
7
|
|
|
7
|
-
The standard workflows, hooks, governance documents, language configuration, and
|
|
8
|
+
The standard workflows, hooks, governance documents, language configuration, and
|
|
9
|
+
release configuration are refreshed from the configured runtime. Keep
|
|
10
|
+
repository-specific behavior outside those managed paths.
|
|
8
11
|
|
|
9
12
|
## Custom workflows
|
|
10
13
|
|
|
11
|
-
Any workflow not named by the baseline is preserved automatically. This is the
|
|
14
|
+
Any workflow not named by the baseline is preserved automatically. This is the
|
|
15
|
+
supported place for project-specific workflows such as Slither, search indexing,
|
|
16
|
+
deployment, Docker publishing, or Vercel tasks.
|
|
12
17
|
|
|
13
|
-
|
|
18
|
+
`custom_workflows: preserve` is the default and the only supported value. Code
|
|
19
|
+
Foundry intentionally has no prune mode for custom workflows; remove those files
|
|
20
|
+
explicitly when they are no longer needed.
|
|
14
21
|
|
|
15
|
-
##
|
|
22
|
+
## Post-release delivery
|
|
16
23
|
|
|
17
|
-
Use `post_release`, `post_release_workflow`, and `post_release_mode` for a
|
|
24
|
+
Use `post_release`, `post_release_workflow`, and `post_release_mode` for a
|
|
25
|
+
post-release artifact workflow:
|
|
18
26
|
|
|
19
|
-
|
|
27
|
+
```yaml
|
|
28
|
+
post_release: true
|
|
29
|
+
post_release_workflow: deploy.yml
|
|
30
|
+
post_release_mode: auto
|
|
31
|
+
```
|
|
32
|
+
|
|
33
|
+
`auto` uses a single workflow dispatch when `CODE_FOUNDRY_TOKEN` is available;
|
|
34
|
+
otherwise it uses the published-release event when that path is enabled. The
|
|
35
|
+
workflow receives `release-tag` and a deterministic `delivery-key`. Delivery is
|
|
36
|
+
at most once per tag, so retries must be explicit and idempotent.
|
|
37
|
+
|
|
38
|
+
Keep deployment credentials, environment files, and project-specific secrets in
|
|
39
|
+
repository or organization configuration. Code Foundry never copies secret
|
|
40
|
+
values or overwrites custom workflows in overlay mode.
|