flameox 0.2.2__tar.gz → 0.2.3__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- {flameox-0.2.2 → flameox-0.2.3}/CHANGELOG.md +22 -0
- {flameox-0.2.2 → flameox-0.2.3}/CONTRIBUTING.md +6 -6
- {flameox-0.2.2 → flameox-0.2.3}/PKG-INFO +48 -21
- {flameox-0.2.2 → flameox-0.2.3}/README.md +47 -20
- {flameox-0.2.2 → flameox-0.2.3}/README.zh-CN.md +5 -4
- {flameox-0.2.2 → flameox-0.2.3}/docs/adapters.md +30 -12
- {flameox-0.2.2 → flameox-0.2.3}/docs/architecture.md +14 -9
- flameox-0.2.3/docs/interfaces.md +173 -0
- {flameox-0.2.2 → flameox-0.2.3}/docs/investigations.md +16 -4
- {flameox-0.2.2 → flameox-0.2.3}/docs/runtime-safety.md +8 -2
- {flameox-0.2.2 → flameox-0.2.3}/docs/storage-and-evidence.md +7 -3
- {flameox-0.2.2 → flameox-0.2.3}/docs/testing.md +3 -3
- {flameox-0.2.2 → flameox-0.2.3}/examples/semantic-matrix/README.md +2 -2
- flameox-0.2.3/npm/README.md +17 -0
- {flameox-0.2.2 → flameox-0.2.3}/npm/package-lock.json +2 -2
- {flameox-0.2.2 → flameox-0.2.3}/npm/package.json +1 -1
- {flameox-0.2.2 → flameox-0.2.3}/pyproject.toml +1 -1
- {flameox-0.2.2 → flameox-0.2.3}/server.json +2 -2
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/cli.py +29 -58
- flameox-0.2.3/src/flameox/mcp/capability_tools.py +79 -0
- flameox-0.2.3/src/flameox/mcp/server.py +459 -0
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/providers/nsight_compute.py +5 -3
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/providers/reliability.py +86 -11
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/providers/structured_workers.py +42 -1
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/pytest_capture.py +14 -3
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/runtime_contracts.py +185 -162
- flameox-0.2.3/src/flameox/setup.py +164 -0
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/stateless.py +472 -679
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/workers/protocol.py +1 -0
- flameox-0.2.3/src/flameox/workers/pstats.py +58 -0
- flameox-0.2.3/src/flameox/workers/pstats_contract.py +40 -0
- flameox-0.2.3/src/flameox/workers/pyperf_target.py +25 -0
- {flameox-0.2.2 → flameox-0.2.3}/tests/providers/test_cpu.py +29 -0
- {flameox-0.2.2 → flameox-0.2.3}/tests/providers/test_nsight_compute.py +44 -0
- {flameox-0.2.2 → flameox-0.2.3}/tests/test_cli_stateless.py +54 -10
- flameox-0.2.3/tests/test_setup_stateless.py +182 -0
- {flameox-0.2.2 → flameox-0.2.3}/tests/test_stateless.py +442 -308
- {flameox-0.2.2 → flameox-0.2.3}/uv.lock +1 -1
- flameox-0.2.2/docs/interfaces.md +0 -108
- flameox-0.2.2/npm/README.md +0 -14
- flameox-0.2.2/src/flameox/mcp/server.py +0 -271
- flameox-0.2.2/src/flameox/setup.py +0 -141
- flameox-0.2.2/tests/test_setup_stateless.py +0 -81
- {flameox-0.2.2 → flameox-0.2.3}/.devcontainer/devcontainer.json +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/.env.example +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/.github/CODEOWNERS +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/.github/ISSUE_TEMPLATE/bug_report.yml +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/.github/ISSUE_TEMPLATE/config.yml +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/.github/ISSUE_TEMPLATE/design_discussion.yml +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/.github/ISSUE_TEMPLATE/feature_request.yml +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/.github/labels.yml +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/.github/pull_request_template.md +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/.github/renovate.json +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/.github/workflows/ci.yml +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/.github/workflows/labels.yml +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/.github/workflows/release-please.yml +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/.github/workflows/release.yml +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/.gitignore +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/.jscpd.json +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/.pre-commit-config.yaml +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/.python-version +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/AGENTS.md +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/LICENSE +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/SECURITY.md +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/cliff.toml +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/docs/assets/flameox-mascot-flamegraph.png +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/examples/semantic-matrix/semantic_workload.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/mcp-publisher +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/mcp-publisher.tar.gz +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/npm/.oxlintrc.json +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/npm/bin/flameox.cjs +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/npm/test/bootstrap.test.cjs +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/__init__.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/__main__.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/adapters/__init__.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/adapters/sarif.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/adapters/torch_benchmark.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/atomic.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/benchmark_samples.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/canonical.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/command_binding.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/environment_policy.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/executable_models.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/execution.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/filesystem.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/mcp/__init__.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/models.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/nsight_compute.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/process_models.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/providers/__init__.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/providers/aiperf.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/providers/benchmarks.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/providers/contracts.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/providers/cpu.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/providers/inference_exports.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/providers/kernel_evidence.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/providers/memray.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/providers/nsight_systems.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/providers/nvbench.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/providers/otlp.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/providers/perfetto.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/providers/source_evidence.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/providers/xctrace.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/py.typed +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/pytest_runner.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/repository.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/runtime_errors.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/sdk.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/workers/__init__.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/workers/aiperf.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/workers/aiperf_contract.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/workers/benchmark_samples.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/workers/benchmark_samples_contract.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/workers/compute_sanitizer.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/workers/compute_sanitizer_contract.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/workers/coverage.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/workers/coverage_contract.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/workers/harness.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/workers/memray.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/workers/memray_contract.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/workers/nsight_compute.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/workers/nsight_compute_contract.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/workers/nsight_systems.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/workers/nsight_systems_contract.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/workers/otlp.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/workers/otlp_contract.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/workers/otlp_parser.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/workers/parquet_schemas.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/workers/perfetto.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/workers/perfetto_contract.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/workers/pyperf.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/workers/pyperf_contract.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/workers/v8_profiles.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/src/flameox/workers/v8_profiles_contract.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/tests/__init__.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/tests/adapters/test_memray_aggregation.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/tests/adapters/test_torch_benchmark.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/tests/adapters/test_torch_benchmark_live.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/tests/conftest.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/tests/execution/test_bound_execution.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/tests/execution/test_broker.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/tests/execution/test_command_binding.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/tests/fixtures/compute_sanitizer/kernel_probe.cu +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/tests/fixtures/coverage/README.md +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/tests/fixtures/coverage/coverage-7.13.sqlite +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/tests/fixtures/coverage/coverage-7.14.sqlite +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/tests/fixtures/coverage/coverage-7.15.sqlite +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/tests/fixtures/coverage/workload.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/tests/fixtures/kernel_validation/pass.json +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/tests/fixtures/nsight_compute/basic.py.txt +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/tests/fixtures/nsight_compute/corrupt.py.txt +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/tests/fixtures/nsight_compute/invalid_uint64.py.txt +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/tests/fixtures/nsight_compute/multi_range.py.txt +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/tests/fixtures/nsight_compute/source_files.py.txt +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/tests/fixtures/nsight_compute/source_markers.py.txt +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/tests/fixtures/nsight_compute/uint_boundaries.py.txt +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/tests/fixtures/pyspy/chrometrace-0.4.2.json +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/tests/fixtures/rocprofv3/project-owned-rocm-shaped-perfetto.json +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/tests/fixtures/triton/vector_add.py.txt +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/tests/performance/test_stateless_scale.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/tests/providers/test_aiperf.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/tests/providers/test_compute_sanitizer.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/tests/providers/test_contracts.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/tests/providers/test_inference_exports.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/tests/providers/test_kernel_evidence.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/tests/providers/test_nsight_systems.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/tests/providers/test_nvbench.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/tests/providers/test_rocprof.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/tests/providers/test_xctrace.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/tests/storage/test_bounded_filesystem.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/tests/support/__init__.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/tests/support/execution.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/tests/support/providers.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/tests/test_canonical.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/tests/test_release_metadata.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/tests/test_runtime_contracts.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/tests/test_sdk.py +0 -0
- {flameox-0.2.2 → flameox-0.2.3}/tools/__init__.py +0 -0
|
@@ -2,6 +2,28 @@
|
|
|
2
2
|
|
|
3
3
|
All notable changes to flameox are documented in this file.
|
|
4
4
|
|
|
5
|
+
## [0.2.3] - 2026-09-02
|
|
6
|
+
|
|
7
|
+
### Bug Fixes
|
|
8
|
+
|
|
9
|
+
- **setup:** Pin generated MCP launchers
|
|
10
|
+
|
|
11
|
+
### Documentation
|
|
12
|
+
|
|
13
|
+
- **mcp:** Clarify provider and comparison workflows
|
|
14
|
+
- **setup:** Clarify uvx preparation lifecycle (#421) ([#421](https://github.com/morluto/flameox/pull/421))
|
|
15
|
+
|
|
16
|
+
### Features
|
|
17
|
+
|
|
18
|
+
- **mcp:** Return provider reconnection action
|
|
19
|
+
- **mcp:** Expose typed capability tools (#422) ([#422](https://github.com/morluto/flameox/pull/422))
|
|
20
|
+
- **mcp:** Restore explicit provider preparation (#419) ([#419](https://github.com/morluto/flameox/pull/419))
|
|
21
|
+
- **runtime:** Unify bounded capture contracts (#417) ([#417](https://github.com/morluto/flameox/pull/417))
|
|
22
|
+
|
|
23
|
+
### Refactoring
|
|
24
|
+
|
|
25
|
+
- **setup:** Prepare the returned uvx environment (#420) ([#420](https://github.com/morluto/flameox/pull/420))
|
|
26
|
+
- **runtime:** Validate captures atomically (#418) ([#418](https://github.com/morluto/flameox/pull/418))
|
|
5
27
|
## [0.2.2] - 2026-09-01
|
|
6
28
|
|
|
7
29
|
### Bug Fixes
|
|
@@ -45,11 +45,11 @@ uv sync --extra dev --extra memory --extra trace --extra cpu --extra torch
|
|
|
45
45
|
|
|
46
46
|
## Understand the contract you are changing
|
|
47
47
|
|
|
48
|
-
Production code uses a `src/` layout.
|
|
49
|
-
`
|
|
50
|
-
`
|
|
51
|
-
|
|
52
|
-
and `
|
|
48
|
+
Production code uses a `src/` layout. `runtime_contracts.py` owns public contracts and registries,
|
|
49
|
+
`stateless.py` owns request-local orchestration, `repository.py` owns optional immutable
|
|
50
|
+
preservation, and `execution.py` owns bounded subprocess work. Provider integrations live in
|
|
51
|
+
`providers/`, reusable native-format parsing in `adapters/`, isolated protocols in `workers/`, and
|
|
52
|
+
transport code in `cli.py` and `mcp/`. Tests mirror these semantic owners under `tests/`.
|
|
53
53
|
|
|
54
54
|
Read the contract that owns the behavior before editing it:
|
|
55
55
|
|
|
@@ -58,7 +58,7 @@ Read the contract that owns the behavior before editing it:
|
|
|
58
58
|
| Process model, dependencies, and package boundaries | [Architecture](docs/architecture.md) |
|
|
59
59
|
| Storage, provenance, publication, and schemas | [Storage and evidence](docs/storage-and-evidence.md) |
|
|
60
60
|
| Experiments, comparisons, statistics, and evidence quality | [Investigations](docs/investigations.md) |
|
|
61
|
-
| Profiler integrations, compatibility, and
|
|
61
|
+
| Profiler integrations, compatibility, and adapter policy | [Adapters](docs/adapters.md) |
|
|
62
62
|
| Concurrency, recovery, integrity, security, and privacy | [Runtime safety](docs/runtime-safety.md) |
|
|
63
63
|
| CLI and MCP behavior and trust boundaries | [Interfaces](docs/interfaces.md) |
|
|
64
64
|
| Test markers, provider requirements, and CI | [Testing](docs/testing.md) |
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
Metadata-Version: 2.5
|
|
2
2
|
Name: flameox
|
|
3
|
-
Version: 0.2.
|
|
3
|
+
Version: 0.2.3
|
|
4
4
|
Summary: Runtime evidence that helps agents trace, profile, and burn down hotspots in application and native code, GPU kernels, and inference stacks.
|
|
5
5
|
Project-URL: Homepage, https://github.com/morluto/flameox
|
|
6
6
|
Project-URL: Repository, https://github.com/morluto/flameox
|
|
@@ -77,7 +77,7 @@ paths and formats to `analyze`; old `.diagnostics` state is not migrated.
|
|
|
77
77
|
|
|
78
78
|
```console
|
|
79
79
|
uv sync --extra dev --extra memory --extra trace --extra cpu
|
|
80
|
-
uv run flameox
|
|
80
|
+
uv run flameox mcp inspect
|
|
81
81
|
uv run flameox analyze artifact.preview /absolute/path/to/artifact.json
|
|
82
82
|
uv run flameox capture --provider direct -- python benchmark.py
|
|
83
83
|
```
|
|
@@ -89,11 +89,17 @@ the startup working directory:
|
|
|
89
89
|
uv run flameox mcp serve --project-root "$PWD"
|
|
90
90
|
```
|
|
91
91
|
|
|
92
|
-
`flameox setup` prints
|
|
93
|
-
|
|
94
|
-
|
|
95
|
-
|
|
96
|
-
|
|
92
|
+
`flameox setup` prints a version-bound stdio client configuration using Python
|
|
93
|
+
3.12 and the exact running Flameox release. It does not change an MCP client
|
|
94
|
+
registration; apply the printed command through the client's supported MCP
|
|
95
|
+
management interface. Explicit
|
|
96
|
+
`--provider` selections prepare the exact version-pinned `uvx` environment in
|
|
97
|
+
the printed launcher by resolving it once into uvx's cache; they do not create a
|
|
98
|
+
persistent global `uv tool` installation. Each invocation declares the complete
|
|
99
|
+
managed provider set for that launcher rather than adding to remembered state.
|
|
100
|
+
Use `--timeout-seconds` for a slow cold resolution. System and vendor tools are
|
|
101
|
+
diagnosed with external install guidance. Setup never initializes the project
|
|
102
|
+
or creates `.flameox`.
|
|
97
103
|
|
|
98
104
|
## Authority model
|
|
99
105
|
|
|
@@ -124,14 +130,23 @@ coverage, truncation, limitations, and optional immutable preservation.
|
|
|
124
130
|
|
|
125
131
|
## MCP interface
|
|
126
132
|
|
|
127
|
-
The server exposes
|
|
133
|
+
The server exposes actual evidence operations for client-side tool search instead of hiding its
|
|
134
|
+
capabilities behind `discover`, `inspect`, or generic `analyze(capability_id, arguments)` calls.
|
|
135
|
+
There are 24 read-only analysis tools, 17 executing capture tools, and three lifecycle tools. For
|
|
136
|
+
example:
|
|
128
137
|
|
|
129
|
-
|
|
130
|
-
|
|
131
|
-
|
|
132
|
-
|
|
133
|
-
|
|
134
|
-
|
|
138
|
+
```text
|
|
139
|
+
analyze_cpu_hotspots capture_cpu_hotspots
|
|
140
|
+
analyze_gpu_launches capture_gpu_launches
|
|
141
|
+
analyze_benchmark_compare capture_benchmark_summary
|
|
142
|
+
analyze_kernel_validation capture_sanitizer_failures
|
|
143
|
+
prepare_providers preserve_evidence query_evidence
|
|
144
|
+
```
|
|
145
|
+
|
|
146
|
+
Each tool advertises its capability-specific options and compatible providers in its input schema.
|
|
147
|
+
Analysis and capture have separate names and annotations because reading an artifact and executing a
|
|
148
|
+
target are materially different effects. Tool search happens in the MCP client; Flameox does not
|
|
149
|
+
require an additional catalog-search call.
|
|
135
150
|
|
|
136
151
|
It exposes one resource template, `flameox://evidence/{evidence_id}`, for the
|
|
137
152
|
digest-bound, redacted projection of the canonical immutable manifest. Full
|
|
@@ -139,16 +154,28 @@ argv, environment values, working directories, and host paths remain available
|
|
|
139
154
|
only through explicit local manifest inspection. Native artifact bytes are
|
|
140
155
|
deliberately not available as MCP resources.
|
|
141
156
|
|
|
142
|
-
Direct capture accepts an argv array, a project-contained cwd, bounded
|
|
143
|
-
|
|
144
|
-
|
|
145
|
-
|
|
146
|
-
restart-surviving tasks.
|
|
157
|
+
Direct capture accepts an argv array, a project-contained cwd, bounded environment overrides, a
|
|
158
|
+
typed compatible-provider variant, capability-specific options, an explicit single/experiment
|
|
159
|
+
choice, and limits as top-level tool arguments. There is no generic request or arguments envelope.
|
|
160
|
+
Shell strings are never accepted. Work remains owned by the live MCP request, so SDK progress and
|
|
161
|
+
cancellation apply directly; there are no detached or restart-surviving tasks.
|
|
147
162
|
|
|
148
|
-
Managed external collectors such as py-spy execute from Flameox's
|
|
163
|
+
Managed external collectors such as py-spy execute from Flameox's uvx
|
|
149
164
|
environment. In-process collectors such as coverage.py and Memray are verified
|
|
150
165
|
in, and run with, the workload's declared Python interpreter. Flameox does not
|
|
151
|
-
substitute one Python runtime for the other.
|
|
166
|
+
substitute one Python runtime for the other. When a capture reports a missing managed provider,
|
|
167
|
+
`prepare_providers` prepares its version-pinned uvx environment and returns that same launcher for
|
|
168
|
+
reconnection. The agent supplies the complete provider list it wants in that launcher; Flameox does
|
|
169
|
+
not merge it with prior calls. Preparation does not modify the running MCP process. When the client
|
|
170
|
+
must reconnect, the result returns a typed `next_action` with `kind: "reconnect_mcp"`, an agent-facing
|
|
171
|
+
message, and the launcher to use. The managed provider IDs are `aiperf`, `memray`, `otlp`, `perfetto`,
|
|
172
|
+
`py-spy`, and `torch`. Host tools, drivers, and permissions are never installed or changed; the same
|
|
173
|
+
result reports their setup guidance.
|
|
174
|
+
|
|
175
|
+
Comparison is intentionally a two-stage workflow. Flameox captures representative baseline and
|
|
176
|
+
candidate summaries separately, optionally preserves them, and then passes both artifacts to an
|
|
177
|
+
`analyze_*_compare` tool. There are no `capture_*_compare` tools: experiment capture measures cases
|
|
178
|
+
and reports an effect, but it is not a substitute for comparing explicit native artifacts.
|
|
152
179
|
|
|
153
180
|
## Evidence quality
|
|
154
181
|
|
|
@@ -17,7 +17,7 @@ paths and formats to `analyze`; old `.diagnostics` state is not migrated.
|
|
|
17
17
|
|
|
18
18
|
```console
|
|
19
19
|
uv sync --extra dev --extra memory --extra trace --extra cpu
|
|
20
|
-
uv run flameox
|
|
20
|
+
uv run flameox mcp inspect
|
|
21
21
|
uv run flameox analyze artifact.preview /absolute/path/to/artifact.json
|
|
22
22
|
uv run flameox capture --provider direct -- python benchmark.py
|
|
23
23
|
```
|
|
@@ -29,11 +29,17 @@ the startup working directory:
|
|
|
29
29
|
uv run flameox mcp serve --project-root "$PWD"
|
|
30
30
|
```
|
|
31
31
|
|
|
32
|
-
`flameox setup` prints
|
|
33
|
-
|
|
34
|
-
|
|
35
|
-
|
|
36
|
-
|
|
32
|
+
`flameox setup` prints a version-bound stdio client configuration using Python
|
|
33
|
+
3.12 and the exact running Flameox release. It does not change an MCP client
|
|
34
|
+
registration; apply the printed command through the client's supported MCP
|
|
35
|
+
management interface. Explicit
|
|
36
|
+
`--provider` selections prepare the exact version-pinned `uvx` environment in
|
|
37
|
+
the printed launcher by resolving it once into uvx's cache; they do not create a
|
|
38
|
+
persistent global `uv tool` installation. Each invocation declares the complete
|
|
39
|
+
managed provider set for that launcher rather than adding to remembered state.
|
|
40
|
+
Use `--timeout-seconds` for a slow cold resolution. System and vendor tools are
|
|
41
|
+
diagnosed with external install guidance. Setup never initializes the project
|
|
42
|
+
or creates `.flameox`.
|
|
37
43
|
|
|
38
44
|
## Authority model
|
|
39
45
|
|
|
@@ -64,14 +70,23 @@ coverage, truncation, limitations, and optional immutable preservation.
|
|
|
64
70
|
|
|
65
71
|
## MCP interface
|
|
66
72
|
|
|
67
|
-
The server exposes
|
|
73
|
+
The server exposes actual evidence operations for client-side tool search instead of hiding its
|
|
74
|
+
capabilities behind `discover`, `inspect`, or generic `analyze(capability_id, arguments)` calls.
|
|
75
|
+
There are 24 read-only analysis tools, 17 executing capture tools, and three lifecycle tools. For
|
|
76
|
+
example:
|
|
68
77
|
|
|
69
|
-
|
|
70
|
-
|
|
71
|
-
|
|
72
|
-
|
|
73
|
-
|
|
74
|
-
|
|
78
|
+
```text
|
|
79
|
+
analyze_cpu_hotspots capture_cpu_hotspots
|
|
80
|
+
analyze_gpu_launches capture_gpu_launches
|
|
81
|
+
analyze_benchmark_compare capture_benchmark_summary
|
|
82
|
+
analyze_kernel_validation capture_sanitizer_failures
|
|
83
|
+
prepare_providers preserve_evidence query_evidence
|
|
84
|
+
```
|
|
85
|
+
|
|
86
|
+
Each tool advertises its capability-specific options and compatible providers in its input schema.
|
|
87
|
+
Analysis and capture have separate names and annotations because reading an artifact and executing a
|
|
88
|
+
target are materially different effects. Tool search happens in the MCP client; Flameox does not
|
|
89
|
+
require an additional catalog-search call.
|
|
75
90
|
|
|
76
91
|
It exposes one resource template, `flameox://evidence/{evidence_id}`, for the
|
|
77
92
|
digest-bound, redacted projection of the canonical immutable manifest. Full
|
|
@@ -79,16 +94,28 @@ argv, environment values, working directories, and host paths remain available
|
|
|
79
94
|
only through explicit local manifest inspection. Native artifact bytes are
|
|
80
95
|
deliberately not available as MCP resources.
|
|
81
96
|
|
|
82
|
-
Direct capture accepts an argv array, a project-contained cwd, bounded
|
|
83
|
-
|
|
84
|
-
|
|
85
|
-
|
|
86
|
-
restart-surviving tasks.
|
|
97
|
+
Direct capture accepts an argv array, a project-contained cwd, bounded environment overrides, a
|
|
98
|
+
typed compatible-provider variant, capability-specific options, an explicit single/experiment
|
|
99
|
+
choice, and limits as top-level tool arguments. There is no generic request or arguments envelope.
|
|
100
|
+
Shell strings are never accepted. Work remains owned by the live MCP request, so SDK progress and
|
|
101
|
+
cancellation apply directly; there are no detached or restart-surviving tasks.
|
|
87
102
|
|
|
88
|
-
Managed external collectors such as py-spy execute from Flameox's
|
|
103
|
+
Managed external collectors such as py-spy execute from Flameox's uvx
|
|
89
104
|
environment. In-process collectors such as coverage.py and Memray are verified
|
|
90
105
|
in, and run with, the workload's declared Python interpreter. Flameox does not
|
|
91
|
-
substitute one Python runtime for the other.
|
|
106
|
+
substitute one Python runtime for the other. When a capture reports a missing managed provider,
|
|
107
|
+
`prepare_providers` prepares its version-pinned uvx environment and returns that same launcher for
|
|
108
|
+
reconnection. The agent supplies the complete provider list it wants in that launcher; Flameox does
|
|
109
|
+
not merge it with prior calls. Preparation does not modify the running MCP process. When the client
|
|
110
|
+
must reconnect, the result returns a typed `next_action` with `kind: "reconnect_mcp"`, an agent-facing
|
|
111
|
+
message, and the launcher to use. The managed provider IDs are `aiperf`, `memray`, `otlp`, `perfetto`,
|
|
112
|
+
`py-spy`, and `torch`. Host tools, drivers, and permissions are never installed or changed; the same
|
|
113
|
+
result reports their setup guidance.
|
|
114
|
+
|
|
115
|
+
Comparison is intentionally a two-stage workflow. Flameox captures representative baseline and
|
|
116
|
+
candidate summaries separately, optionally preserves them, and then passes both artifacts to an
|
|
117
|
+
`analyze_*_compare` tool. There are no `capture_*_compare` tools: experiment capture measures cases
|
|
118
|
+
and reports an effect, but it is not a substitute for comparing explicit native artifacts.
|
|
92
119
|
|
|
93
120
|
## Evidence quality
|
|
94
121
|
|
|
@@ -8,7 +8,7 @@ Flameox 是面向编码代理的本地、有界运行时证据层。它协调分
|
|
|
8
8
|
包含 argv、项目内 cwd、环境覆盖、提供方参数和限制的类型化目标。
|
|
9
9
|
|
|
10
10
|
```console
|
|
11
|
-
uv run flameox
|
|
11
|
+
uv run flameox mcp inspect
|
|
12
12
|
uv run flameox analyze artifact.preview /absolute/path/to/artifact.json
|
|
13
13
|
uv run flameox capture --provider direct -- python benchmark.py
|
|
14
14
|
uv run flameox mcp serve --project-root "$PWD"
|
|
@@ -23,9 +23,10 @@ Python 扩展安装到持久的 uv 工具环境;NVIDIA 等系统或厂商工
|
|
|
23
23
|
`<project>/.flameox`。原生字节和证据清单按 SHA-256 寻址,并通过同一文件系统
|
|
24
24
|
上的暂存、校验、fsync 和原子重命名发布。
|
|
25
25
|
|
|
26
|
-
MCP
|
|
27
|
-
`
|
|
28
|
-
`query_evidence
|
|
26
|
+
MCP 公开面向任务的类型化分析和采集工具,例如 `analyze_cpu_hotspots`、
|
|
27
|
+
`capture_gpu_launches` 和 `analyze_benchmark_compare`,以及
|
|
28
|
+
`prepare_providers`、`preserve_evidence` 和 `query_evidence`。工具搜索由 MCP
|
|
29
|
+
客户端负责。唯一资源模板是
|
|
29
30
|
`flameox://evidence/{evidence_id}`,只返回不可变规范清单,不公开原生载荷。
|
|
30
31
|
|
|
31
32
|
`analysis_id` 仅在当前服务进程内有效;重启后过期。`evidence_id` 是持久的内容
|
|
@@ -11,14 +11,19 @@ Each capability registry entry owns:
|
|
|
11
11
|
- the capability descriptor and stable ID;
|
|
12
12
|
- accepted source modes and formats;
|
|
13
13
|
- a strict validation-equivalent argument model;
|
|
14
|
-
-
|
|
14
|
+
- compatibility rules derived from declared artifact formats;
|
|
15
15
|
- capture and analysis handlers;
|
|
16
|
-
- bounded examples, limits, overhead, and limitations.
|
|
16
|
+
- model-visible selection guidance, bounded examples, limits, overhead, and limitations.
|
|
17
17
|
|
|
18
|
-
|
|
19
|
-
and
|
|
20
|
-
|
|
21
|
-
|
|
18
|
+
One immutable capture-provider contract owns each provider's argument model, declared artifact
|
|
19
|
+
roles and formats, and selection description. Generated MCP schemas, pre-execution compatibility
|
|
20
|
+
checks, and capture argument validation consume that same contract so their format claims cannot
|
|
21
|
+
drift.
|
|
22
|
+
|
|
23
|
+
Capability tools remain discoverable when a provider is absent. Capture validates the selected
|
|
24
|
+
provider's package, executable, platform, version, permission, and required external resources as
|
|
25
|
+
part of the attempted operation, before workload execution. An unavailable provider returns a typed
|
|
26
|
+
error with either an exact `prepare_providers` retry action or external host guidance.
|
|
22
27
|
|
|
23
28
|
## Evidence and capture support
|
|
24
29
|
|
|
@@ -29,7 +34,7 @@ Artifact analysis and typed capture are separate contracts:
|
|
|
29
34
|
|
|
30
35
|
| Evidence family | Explicit artifact analysis | Typed capture provider |
|
|
31
36
|
| --- | --- | --- |
|
|
32
|
-
| CPU profiles | Node/V8, py-spy Speedscope, perf collapsed stacks, perf data | `node-cpu-profile`, `py-spy`, `perf` |
|
|
37
|
+
| CPU profiles | Node/V8, cProfile pstats, py-spy Speedscope, perf collapsed stacks, perf data | `node-cpu-profile`, `py-spy`, `perf` |
|
|
33
38
|
| Memory profiles | Memray | `memray` |
|
|
34
39
|
| Benchmarks | pyperf, benchmark samples, PyTorch samples, NVBench | `pyperf`, `benchmark-samples`, `nvbench` |
|
|
35
40
|
| Execution traces | Perfetto/Chrome, PyTorch, OTLP, ROCprof PFTrace, xctrace, Nsight Systems | `torch-profiler`, `rocprofv3`, `xctrace`, `nsight-systems` |
|
|
@@ -37,7 +42,7 @@ Artifact analysis and typed capture are separate contracts:
|
|
|
37
42
|
| Reliability | pytest events, observations, coverage.py | `pytest`, `observations`, `coverage` |
|
|
38
43
|
| Static candidates | SARIF | — |
|
|
39
44
|
| Inference exports | AIPerf, vLLM, SGLang, Mooncake | — |
|
|
40
|
-
| Generic previews | JSON, JSONL, CSV, text, Parquet |
|
|
45
|
+
| Generic previews | JSON, JSONL, CSV, text, Parquet | `direct` process output |
|
|
41
46
|
|
|
42
47
|
An em dash means callers provide an explicit artifact path; it does not mean
|
|
43
48
|
the format is unsupported. The offline inference readers omit prompts,
|
|
@@ -45,6 +50,11 @@ generations, error text, endpoints, tools, payloads, and prefix-hash values.
|
|
|
45
50
|
Pytest capture runs an explicit `python -m pytest` target with a request-bound,
|
|
46
51
|
bounded event plugin.
|
|
47
52
|
|
|
53
|
+
`failures.summary` scans the complete bounded pytest event artifact for aggregate outcomes, but its
|
|
54
|
+
table projects only failed, errored, interrupted, and unexecuted identities. Failed collection
|
|
55
|
+
reports retain their collector identity; successful collection events, passing tests, and skipped
|
|
56
|
+
tests cannot consume the diagnostic row budget.
|
|
57
|
+
|
|
48
58
|
Support is honest rather than substitutive. A missing Trace Processor does not
|
|
49
59
|
turn a Perfetto request into a JSON preview; a missing `ncu` does not become an
|
|
50
60
|
empty kernel report.
|
|
@@ -82,14 +92,22 @@ session-local in-memory DuckDB; it does not create a temporary SQLite database.
|
|
|
82
92
|
|
|
83
93
|
## Provider setup boundary
|
|
84
94
|
|
|
85
|
-
An explicit `flameox setup --provider ...` invocation
|
|
86
|
-
|
|
87
|
-
|
|
95
|
+
An explicit `flameox setup --provider ...` invocation prepares the exact
|
|
96
|
+
version-pinned uvx environment returned in its launcher. The invocation's
|
|
97
|
+
provider list is complete; Flameox retains no provider inventory or setup receipt.
|
|
88
98
|
System and vendor packages remain externally installed. Setup creates no
|
|
89
99
|
project state and is not a prerequisite for explicit-path analysis.
|
|
100
|
+
Nsight Compute capture requires both the `ncu` executable and NVIDIA's vendor-shipped
|
|
101
|
+
`extras/python/ncu_report.py` reader. Flameox resolves both before workload execution and returns
|
|
102
|
+
external setup guidance if the installation is incomplete.
|
|
90
103
|
|
|
91
104
|
Provider ownership is explicit during capture. External collectors installed
|
|
92
|
-
with Flameox, such as py-spy, execute from the
|
|
105
|
+
with Flameox, such as py-spy, execute from the launched uvx environment rather
|
|
93
106
|
than ambient request `PATH`. In-process collectors, including coverage.py and
|
|
94
107
|
Memray, must be installed in the declared workload interpreter; capture probes
|
|
95
108
|
that interpreter before execution and never substitutes Flameox's interpreter.
|
|
109
|
+
|
|
110
|
+
Pyperf command capture preserves multiline typed argv through a metadata-safe launcher because
|
|
111
|
+
pyperf rejects newline characters in its display metadata. The launcher executes the original argv
|
|
112
|
+
without shell parsing, but its startup is included in each command measurement; exact Flameox
|
|
113
|
+
capture provenance remains authoritative for the requested argv.
|
|
@@ -36,8 +36,8 @@ durable SHA-256 identity derived from the canonical manifest body.
|
|
|
36
36
|
|
|
37
37
|
## Package boundaries
|
|
38
38
|
|
|
39
|
-
- `
|
|
40
|
-
|
|
39
|
+
- `runtime_contracts.py` owns public models and the capability/capture-provider registries.
|
|
40
|
+
- `stateless.py` owns bounded analysis, capture orchestration, scratch, and the session cache.
|
|
41
41
|
- `repository.py` owns lazy repository creation, validation, publication,
|
|
42
42
|
inventory queries, and immutable resource reads.
|
|
43
43
|
- `execution.py` and `command_binding.py` own executable binding, subprocess
|
|
@@ -51,13 +51,18 @@ catalog. Flameox production code must not create or depend on SQLite state.
|
|
|
51
51
|
|
|
52
52
|
## Capability boundary
|
|
53
53
|
|
|
54
|
-
One registry entry owns a capability descriptor, strict argument model,
|
|
55
|
-
|
|
56
|
-
|
|
57
|
-
|
|
58
|
-
|
|
59
|
-
|
|
60
|
-
|
|
54
|
+
One registry entry owns a capability descriptor, strict argument model, accepted formats,
|
|
55
|
+
capture/analysis semantics, and model-visible selection guidance. MCP projects each entry
|
|
56
|
+
into a read-only analysis tool and, when a compatible capture provider exists, a separate executing
|
|
57
|
+
capture tool. The generated callable closes over the stable capability ID; agents never pass a
|
|
58
|
+
capability selector or a free-form analysis argument object.
|
|
59
|
+
|
|
60
|
+
Capture-provider contracts supply the discriminated provider variants for each compatible capture
|
|
61
|
+
tool. Missing packages, executables, permissions, versions, or platforms do not change the catalog;
|
|
62
|
+
the attempted tool returns typed remediation. The separately invoked CLI setup command or MCP
|
|
63
|
+
`prepare_providers` tool may prepare the exact version-pinned uvx environment named by an explicit
|
|
64
|
+
Python provider set; neither creates project state nor owns a durable operation or provider
|
|
65
|
+
inventory. Host profilers, drivers, and permissions remain external and receive guidance only.
|
|
61
66
|
|
|
62
67
|
Direct capture is trusted local execution, not containment. Typed argv prevents
|
|
63
68
|
shell interpretation, while the broker provides process-group cleanup, bounded
|
|
@@ -0,0 +1,173 @@
|
|
|
1
|
+
# Interfaces
|
|
2
|
+
|
|
3
|
+
CLI and MCP are thin transports over `AnalysisRuntime`. They do not own storage,
|
|
4
|
+
provider behavior, or lifecycle state.
|
|
5
|
+
|
|
6
|
+
## MCP catalog
|
|
7
|
+
|
|
8
|
+
The catalog exposes task-shaped tools for client-side tool search. Flameox does not add a second
|
|
9
|
+
search/inspect protocol in front of its operations. A caller that knows the evidence question can
|
|
10
|
+
invoke its tool directly; an unfamiliar caller relies on the MCP client's ordinary tool search and
|
|
11
|
+
then receives the selected tool's complete schema.
|
|
12
|
+
|
|
13
|
+
There are exactly 44 tools:
|
|
14
|
+
|
|
15
|
+
| Group | Count | Examples | Effect |
|
|
16
|
+
| --- | ---: | --- | --- |
|
|
17
|
+
| Existing-artifact analysis | 24 | `analyze_cpu_hotspots`, `analyze_gpu_launches`, `analyze_benchmark_compare`, `analyze_kernel_validation`, `preview_artifact` | Read-only and idempotent. |
|
|
18
|
+
| Capture and immediate analysis | 17 | `capture_cpu_hotspots`, `capture_gpu_launches`, `capture_benchmark_summary`, `capture_sanitizer_failures`, `capture_process_output` | Executes typed argv; not read-only or idempotent. |
|
|
19
|
+
| Evidence lifecycle | 3 | `prepare_providers`, `preserve_evidence`, `query_evidence` | Prepare an explicit uvx environment or manage immutable evidence. |
|
|
20
|
+
|
|
21
|
+
The capability registry generates the analysis and capture tools through the Python MCP SDK 2.0
|
|
22
|
+
registration API. The SDK derives each top-level input schema directly from the registered callable.
|
|
23
|
+
A generated analysis tool has `sources`, capability-specific typed `options`, optional lowered
|
|
24
|
+
`limits`, and an optional `continuation`. A generated capture tool has `target`, a discriminated
|
|
25
|
+
`provider` union containing only compatible capture providers, capability-specific typed `options`,
|
|
26
|
+
an explicit `single` or `experiment` execution union, optional lowered `limits`, and optional
|
|
27
|
+
`preserve`. There is no extra request envelope and there are no free-form provider or analysis
|
|
28
|
+
argument objects.
|
|
29
|
+
|
|
30
|
+
For example, a single Nsight Compute capture for kernel metrics has this argument shape:
|
|
31
|
+
|
|
32
|
+
```json
|
|
33
|
+
{
|
|
34
|
+
"target": {"argv": ["python", "kernel.py"]},
|
|
35
|
+
"provider": {"kind": "nsight-compute", "options": {"launch_count": 1}},
|
|
36
|
+
"options": {},
|
|
37
|
+
"execution": {"kind": "single"},
|
|
38
|
+
"preserve": true
|
|
39
|
+
}
|
|
40
|
+
```
|
|
41
|
+
|
|
42
|
+
Analysis and capture remain separate tools even when they return the same evidence envelope. MCP
|
|
43
|
+
annotations describe a whole tool, so combining read-only artifact analysis and target execution
|
|
44
|
+
behind a mode flag would conceal a material effect change. Provider choice stays inside a capture
|
|
45
|
+
tool because it is a typed implementation choice for one evidence question; incompatible providers
|
|
46
|
+
cannot be represented by that tool's schema.
|
|
47
|
+
|
|
48
|
+
There is one resource template:
|
|
49
|
+
|
|
50
|
+
```text
|
|
51
|
+
flameox://evidence/{evidence_id}
|
|
52
|
+
```
|
|
53
|
+
|
|
54
|
+
`resources/list` is empty. `resources/read` returns a redacted, digest-bound projection of the
|
|
55
|
+
canonical manifest with its own versioned media type. It omits argv, environment values, working
|
|
56
|
+
directories, and host paths. The local CLI `evidence show` command is the explicit full-provenance
|
|
57
|
+
view. A missing resource is a protocol error.
|
|
58
|
+
|
|
59
|
+
Every tool advertises a compact output schema for its stable result envelope. Provider-specific
|
|
60
|
+
metrics and rows remain open JSON values. Success uses structured content directly, without an
|
|
61
|
+
`ok/result/error` wrapper. Tool failures set `isError=true` and carry a stable code, message, and
|
|
62
|
+
details. MCP SDK argument-validation errors occur before tool execution and therefore use the
|
|
63
|
+
protocol error shape rather than the tool's output schema.
|
|
64
|
+
|
|
65
|
+
`prepare_providers` is the only open-world tool. It resolves the exact package requirement through
|
|
66
|
+
`uvx`, verifies that environment by running Flameox's version command, and returns the same
|
|
67
|
+
requirement in a project-bound MCP launcher. The request names the complete managed provider set;
|
|
68
|
+
Flameox keeps no installed-provider inventory or setup receipt. System profilers, drivers, device
|
|
69
|
+
access, and OS permissions remain external requirements; Flameox returns guidance for them but does
|
|
70
|
+
not invoke a system package manager or elevate privileges. Preparation creates no project state,
|
|
71
|
+
durable job, or plan, and it cannot add packages to the currently running server. A result with
|
|
72
|
+
a non-null `next_action` identifies the typed `reconnect_mcp` handoff, explains that the current
|
|
73
|
+
server is unchanged, and directs the agent to reconnect with the returned launcher before retrying
|
|
74
|
+
the capture. Host-only preparation returns `next_action: null`. Managed provider IDs are `aiperf`,
|
|
75
|
+
`memray`, `otlp`, `perfetto`, `py-spy`, and `torch`; every call declares the complete desired set. A
|
|
76
|
+
provider such as Perfetto may be both prepared Python support and an external host Trace Processor
|
|
77
|
+
requirement. Preparation waits up to 1,800 seconds by default; callers may set `timeout_seconds`
|
|
78
|
+
from 1 through 3,600. A uvx failure returns its complete stderr in `SETUP_FAILURE`.
|
|
79
|
+
|
|
80
|
+
## Sources and limits
|
|
81
|
+
|
|
82
|
+
The strict source union is:
|
|
83
|
+
|
|
84
|
+
```text
|
|
85
|
+
PathSource {kind: "path", path, format?, producer?, expected_sha256?}
|
|
86
|
+
EvidenceSource {kind: "evidence", evidence_id, artifact_role?}
|
|
87
|
+
```
|
|
88
|
+
|
|
89
|
+
Continuations are opaque integrity cursors bound to the request and exact input
|
|
90
|
+
digests. They contain no authority, credentials, or artifact data and are not
|
|
91
|
+
an authentication boundary: a caller already authorized to submit the analysis
|
|
92
|
+
can choose which of its rows to request. They can cross process boundaries, so
|
|
93
|
+
a CLI invocation can resume a previous page. A changed input cannot reuse a
|
|
94
|
+
continuation.
|
|
95
|
+
|
|
96
|
+
Projection providers may expose a bounded prefix when their native reader cannot
|
|
97
|
+
resume safely. Such results keep `coverage.complete=false`, identify
|
|
98
|
+
`truncation.reason=provider_limit`, and do not emit a continuation after the last
|
|
99
|
+
retrievable row. A continuation therefore always names a consumable next page;
|
|
100
|
+
it never promises access beyond a provider's declared projection bound.
|
|
101
|
+
|
|
102
|
+
Requests may lower startup row, result-byte, timeout, output-byte, and durable
|
|
103
|
+
provenance-byte limits. Durable provenance bounds the captured argv and execution
|
|
104
|
+
metadata retained for explicit preservation.
|
|
105
|
+
They cannot raise them.
|
|
106
|
+
|
|
107
|
+
## Capture
|
|
108
|
+
|
|
109
|
+
A direct target contains an argv array, project-contained cwd, and at most 32 bounded environment
|
|
110
|
+
overrides after experiment-case overrides are merged. Provider fields live in the capture tool's
|
|
111
|
+
typed provider union, and analysis fields live in its capability-specific `options` model. Shell
|
|
112
|
+
command strings are not accepted.
|
|
113
|
+
|
|
114
|
+
Provider output formats are compared with the requested capability before scratch creation or
|
|
115
|
+
execution. Statically incompatible pairs fail with the declared formats and compatible capture
|
|
116
|
+
providers. Capture also resolves the cwd and executable, checks aggregate scratch and durable
|
|
117
|
+
provenance capacity, and validates every experiment case before creating request scratch. There is
|
|
118
|
+
no separate plan or preflight authority: the typed capture request is validated atomically by the
|
|
119
|
+
operation that executes it.
|
|
120
|
+
|
|
121
|
+
Callers may request preservation as part of capture. Once native collection succeeds, requested
|
|
122
|
+
preservation publishes the native artifacts even if immediate analysis fails; the result then
|
|
123
|
+
reports separate capture execution state and a typed `analysis_failure`. An analysis failure is
|
|
124
|
+
never converted into empty successful evidence.
|
|
125
|
+
|
|
126
|
+
Experiment mode adds 2-16 cases, 1-100 blocks, a seed, metric, estimand,
|
|
127
|
+
practical threshold, and optional semantic-oracle argv. Version 0.2 evaluates
|
|
128
|
+
`wall_time_ns` with a paired `median_difference` or `mean_difference`, reports
|
|
129
|
+
eligible blocks and a deterministic percentile interval when at least three
|
|
130
|
+
blocks survive capture/oracle validation, and classifies the effect against the
|
|
131
|
+
declared threshold. Work is not detached; the request receives progress and owns
|
|
132
|
+
cancellation.
|
|
133
|
+
|
|
134
|
+
Comparison tools consume explicit artifacts; they do not capture their inputs. A caller captures
|
|
135
|
+
representative baseline and candidate summaries separately, preserves them when durable provenance
|
|
136
|
+
is needed, and supplies at least two sources to `analyze_benchmark_compare`,
|
|
137
|
+
`analyze_inference_compare`, or `analyze_kernel_compare`. Flameox does not advertise
|
|
138
|
+
`capture_*_compare`: experiment capture reports the declared cases' effect but does not create the
|
|
139
|
+
case-grouped native inputs required by artifact comparison.
|
|
140
|
+
|
|
141
|
+
## Stable failure codes
|
|
142
|
+
|
|
143
|
+
The transport distinguishes invalid input, unavailable providers,
|
|
144
|
+
missing or changed input, unsupported format, decode failure, execution failure,
|
|
145
|
+
cancellation, limit exceeded, expired session analysis, missing evidence,
|
|
146
|
+
repository I/O failure, repository corruption, and unsupported repository
|
|
147
|
+
format. An unavailable managed provider identifies `prepare_providers` and the exact provider list
|
|
148
|
+
needed for a retry; an unavailable system provider returns external setup guidance.
|
|
149
|
+
|
|
150
|
+
## CLI
|
|
151
|
+
|
|
152
|
+
The retained surface is:
|
|
153
|
+
|
|
154
|
+
```text
|
|
155
|
+
flameox setup
|
|
156
|
+
flameox mcp serve|inspect
|
|
157
|
+
flameox analyze [--continuation TOKEN] [--preserve]
|
|
158
|
+
flameox capture [--preserve] -- <argv...>
|
|
159
|
+
flameox evidence query|show
|
|
160
|
+
```
|
|
161
|
+
|
|
162
|
+
`setup` prints stdio configuration that launches the exact running Flameox
|
|
163
|
+
release through `uvx` on Python 3.12. It reports that no MCP client registration
|
|
164
|
+
was changed; client-owned management interfaces apply the returned command.
|
|
165
|
+
Repeated `--provider` options declare the complete Python provider set for the
|
|
166
|
+
exact version-pinned uvx environment used by the returned launcher.
|
|
167
|
+
`--timeout-seconds` accepts 1 through 3,600 and defaults to 1,800. Resolver,
|
|
168
|
+
download, and compatibility failures retain uvx's complete stderr. System and
|
|
169
|
+
vendor providers receive external installation guidance. Setup does not create
|
|
170
|
+
a persistent global tool, project repository, durable operation, or MCP setup
|
|
171
|
+
endpoint.
|
|
172
|
+
Other CLI commands construct the same runtime and project-root rules used by
|
|
173
|
+
MCP.
|
|
@@ -33,10 +33,16 @@ containment must not be silently promoted to complete evidence.
|
|
|
33
33
|
|
|
34
34
|
## Experiments
|
|
35
35
|
|
|
36
|
-
`
|
|
37
|
-
declares cases, blocks, seed, metric, estimand, practical threshold, and an
|
|
38
|
-
optional semantic oracle. Cases are bounded and execute through the same broker
|
|
39
|
-
|
|
36
|
+
Every `capture_*` capability tool accepts a discriminated `single` or `experiment` execution
|
|
37
|
+
request. An experiment declares cases, blocks, seed, metric, estimand, practical threshold, and an
|
|
38
|
+
optional semantic oracle. Cases are bounded and execute through the same broker as a single capture.
|
|
39
|
+
|
|
40
|
+
For GPU kernel work, the agent normally compiles and edits with its native coding tools, records
|
|
41
|
+
correctness through `analyze_kernel_validation` and `analyze_kernel_compare`, checks hazards with
|
|
42
|
+
`capture_sanitizer_failures`, measures representative baseline/candidate cases with
|
|
43
|
+
`capture_benchmark_summary` in experiment mode, and profiles only the remaining uncertainty with
|
|
44
|
+
`capture_gpu_launches` or `capture_gpu_kernel_metrics`. Flameox preserves the verification evidence;
|
|
45
|
+
it does not generate kernels, wrap compilers, or decide which optimization to implement.
|
|
40
46
|
|
|
41
47
|
The 0.2 runtime accepts `wall_time_ns` and paired `median_difference` or
|
|
42
48
|
`mean_difference`. Each non-baseline case is compared with the first declared
|
|
@@ -45,6 +51,12 @@ reported as limitations; fewer than three eligible pairs produce a descriptive
|
|
|
45
51
|
estimate without a confidence interval. The practical-threshold decision is
|
|
46
52
|
returned as typed comparison evidence, not retained only as request metadata.
|
|
47
53
|
|
|
54
|
+
Artifact comparison is separate from that experiment result. Capture the representative baseline
|
|
55
|
+
and candidate summaries independently, preserve them if they must survive the session, then submit
|
|
56
|
+
both sources to the matching `analyze_*_compare` tool. There is no `capture_*_compare` shortcut:
|
|
57
|
+
comparison requires explicit artifact identity, while experiment mode owns randomized case order
|
|
58
|
+
and repeated measurements within one request.
|
|
59
|
+
|
|
48
60
|
Randomization and blocking reduce ordering and environmental bias; they do not
|
|
49
61
|
make an unrepresentative workload representative. Failed and partial trials are
|
|
50
62
|
evidence and must stay visible in the returned episode. A semantic oracle checks
|