skillstate-kit 0.1.1__tar.gz → 0.2.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- skillstate_kit-0.2.0/CHANGELOG.md +40 -0
- {skillstate_kit-0.1.1 → skillstate_kit-0.2.0}/CITATION.cff +1 -1
- {skillstate_kit-0.1.1 → skillstate_kit-0.2.0}/PKG-INFO +36 -7
- skillstate_kit-0.2.0/README.md +264 -0
- {skillstate_kit-0.1.1 → skillstate_kit-0.2.0}/SECURITY.md +1 -1
- {skillstate_kit-0.1.1 → skillstate_kit-0.2.0}/docs/PYPI_README.md +35 -6
- skillstate_kit-0.2.0/docs/README_TR.md +67 -0
- skillstate_kit-0.2.0/docs/architecture-audit.md +40 -0
- {skillstate_kit-0.1.1 → skillstate_kit-0.2.0}/docs/architecture.md +18 -0
- skillstate_kit-0.2.0/docs/assets/logo.svg +27 -0
- skillstate_kit-0.2.0/docs/assets/mark.svg +10 -0
- skillstate_kit-0.2.0/docs/benchmarks/coding-agent.md +105 -0
- {skillstate_kit-0.1.1 → skillstate_kit-0.2.0}/docs/cli.md +63 -1
- {skillstate_kit-0.1.1 → skillstate_kit-0.2.0}/docs/compatibility.md +13 -0
- skillstate_kit-0.2.0/docs/evidence/coding-agent-2026-09-07.json +409 -0
- skillstate_kit-0.2.0/docs/evidence/smolagents-2026-09-07.json +92 -0
- skillstate_kit-0.2.0/docs/execution-state.md +64 -0
- skillstate_kit-0.2.0/docs/host-acceptance.md +110 -0
- skillstate_kit-0.2.0/docs/integrations.md +109 -0
- skillstate_kit-0.2.0/docs/releases/0.1.1.md +28 -0
- skillstate_kit-0.2.0/docs/releases/0.1.2.md +35 -0
- skillstate_kit-0.2.0/docs/releases/0.2.0.md +43 -0
- {skillstate_kit-0.1.1 → skillstate_kit-0.2.0}/docs/releasing.md +4 -4
- skillstate_kit-0.2.0/docs/smolagents-acceptance.md +124 -0
- {skillstate_kit-0.1.1 → skillstate_kit-0.2.0}/docs/validation.md +29 -1
- skillstate_kit-0.2.0/examples/coding_agent/README.md +118 -0
- skillstate_kit-0.2.0/examples/coding_agent/TASK.md +30 -0
- skillstate_kit-0.2.0/examples/coding_agent/evaluate.py +71 -0
- skillstate_kit-0.2.0/examples/coding_agent/fixture/README.md +5 -0
- skillstate_kit-0.2.0/examples/coding_agent/fixture/tests/test_existing.py +5 -0
- skillstate_kit-0.2.0/examples/coding_agent/fixture/timebox/__init__.py +9 -0
- skillstate_kit-0.2.0/examples/coding_agent/run.py +309 -0
- skillstate_kit-0.2.0/examples/coding_agent/verify_handoff.py +59 -0
- skillstate_kit-0.2.0/examples/desktop_acceptance/README.md +49 -0
- skillstate_kit-0.2.0/examples/desktop_acceptance/SKILL.md +16 -0
- skillstate_kit-0.2.0/examples/smolagents_acceptance/README.md +93 -0
- skillstate_kit-0.2.0/examples/smolagents_acceptance/run.py +361 -0
- skillstate_kit-0.2.0/examples/smolagents_acceptance/verify.py +133 -0
- skillstate_kit-0.2.0/integrations/README.md +40 -0
- skillstate_kit-0.2.0/integrations/claude/skillstate-kit/.claude-plugin/plugin.json +11 -0
- skillstate_kit-0.2.0/integrations/claude/skillstate-kit/skills/skillstate-task/SKILL.md +18 -0
- skillstate_kit-0.2.0/integrations/codex/skillstate-kit/.codex-plugin/plugin.json +25 -0
- skillstate_kit-0.2.0/integrations/codex/skillstate-kit/skills/skillstate-task/SKILL.md +18 -0
- {skillstate_kit-0.1.1 → skillstate_kit-0.2.0}/pyproject.toml +2 -2
- skillstate_kit-0.2.0/scripts/build_integrations.py +39 -0
- {skillstate_kit-0.1.1 → skillstate_kit-0.2.0}/scripts/check_wheel.py +17 -1
- {skillstate_kit-0.1.1 → skillstate_kit-0.2.0}/src/skillstate/__init__.py +3 -1
- skillstate_kit-0.2.0/src/skillstate/assets/task-SKILL.md +18 -0
- {skillstate_kit-0.1.1 → skillstate_kit-0.2.0}/src/skillstate/cli.py +72 -5
- {skillstate_kit-0.1.1 → skillstate_kit-0.2.0}/src/skillstate/compiler.py +7 -1
- skillstate_kit-0.2.0/src/skillstate/desktop.py +142 -0
- skillstate_kit-0.2.0/src/skillstate/host_registry.py +161 -0
- {skillstate_kit-0.1.1 → skillstate_kit-0.2.0}/src/skillstate/hosts.py +107 -1
- skillstate_kit-0.2.0/src/skillstate/lifecycle.py +218 -0
- {skillstate_kit-0.1.1 → skillstate_kit-0.2.0}/src/skillstate/mcp_server.py +32 -2
- skillstate_kit-0.2.0/src/skillstate/service.py +167 -0
- {skillstate_kit-0.1.1 → skillstate_kit-0.2.0}/tests/test_cli_mcp_provider.py +56 -1
- skillstate_kit-0.2.0/tests/test_coding_benchmark.py +48 -0
- {skillstate_kit-0.1.1 → skillstate_kit-0.2.0}/tests/test_compiler_hosts.py +8 -0
- skillstate_kit-0.2.0/tests/test_desktop.py +120 -0
- skillstate_kit-0.2.0/tests/test_host_registry.py +86 -0
- skillstate_kit-0.2.0/tests/test_lifecycle.py +207 -0
- {skillstate_kit-0.1.1 → skillstate_kit-0.2.0}/uv.lock +1 -1
- skillstate_kit-0.1.1/CHANGELOG.md +0 -21
- skillstate_kit-0.1.1/README.md +0 -167
- skillstate_kit-0.1.1/docs/README_TR.md +0 -43
- skillstate_kit-0.1.1/src/skillstate/service.py +0 -59
- {skillstate_kit-0.1.1 → skillstate_kit-0.2.0}/.gitignore +0 -0
- {skillstate_kit-0.1.1 → skillstate_kit-0.2.0}/CONTRIBUTING.md +0 -0
- {skillstate_kit-0.1.1 → skillstate_kit-0.2.0}/LICENSE +0 -0
- {skillstate_kit-0.1.1 → skillstate_kit-0.2.0}/docs/research/DESIGN_TR.md +0 -0
- {skillstate_kit-0.1.1 → skillstate_kit-0.2.0}/docs/research/REPO_INCELEME_TR.md +0 -0
- {skillstate_kit-0.1.1 → skillstate_kit-0.2.0}/docs/research/URUN_MIMARI_V2_TR.md +0 -0
- {skillstate_kit-0.1.1 → skillstate_kit-0.2.0}/examples/managed_runtime.py +0 -0
- {skillstate_kit-0.1.1 → skillstate_kit-0.2.0}/examples/qa/SKILL.md +0 -0
- {skillstate_kit-0.1.1 → skillstate_kit-0.2.0}/scripts/verify_index.py +0 -0
- {skillstate_kit-0.1.1 → skillstate_kit-0.2.0}/src/skillstate/__main__.py +0 -0
- {skillstate_kit-0.1.1 → skillstate_kit-0.2.0}/src/skillstate/artifacts.py +0 -0
- {skillstate_kit-0.1.1 → skillstate_kit-0.2.0}/src/skillstate/demo.py +0 -0
- {skillstate_kit-0.1.1 → skillstate_kit-0.2.0}/src/skillstate/errors.py +0 -0
- {skillstate_kit-0.1.1 → skillstate_kit-0.2.0}/src/skillstate/jsonio.py +0 -0
- {skillstate_kit-0.1.1 → skillstate_kit-0.2.0}/src/skillstate/models.py +0 -0
- {skillstate_kit-0.1.1 → skillstate_kit-0.2.0}/src/skillstate/providers.py +0 -0
- {skillstate_kit-0.1.1 → skillstate_kit-0.2.0}/src/skillstate/py.typed +0 -0
- {skillstate_kit-0.1.1 → skillstate_kit-0.2.0}/src/skillstate/runtime.py +0 -0
- {skillstate_kit-0.1.1 → skillstate_kit-0.2.0}/src/skillstate/schema.py +0 -0
- {skillstate_kit-0.1.1 → skillstate_kit-0.2.0}/src/skillstate/store.py +0 -0
- {skillstate_kit-0.1.1 → skillstate_kit-0.2.0}/tests/conftest.py +0 -0
- {skillstate_kit-0.1.1 → skillstate_kit-0.2.0}/tests/test_json_schema.py +0 -0
- {skillstate_kit-0.1.1 → skillstate_kit-0.2.0}/tests/test_publish_verification.py +0 -0
- {skillstate_kit-0.1.1 → skillstate_kit-0.2.0}/tests/test_runtime.py +0 -0
- {skillstate_kit-0.1.1 → skillstate_kit-0.2.0}/tests/test_store.py +0 -0
|
@@ -0,0 +1,40 @@
|
|
|
1
|
+
# Changelog
|
|
2
|
+
|
|
3
|
+
## 0.2.0 — portable task execution state
|
|
4
|
+
|
|
5
|
+
- Add an optional evidence-backed task profile with ordered milestones, explicit revalidation, per-resource freshness checks and validated completion.
|
|
6
|
+
- Add host discovery and a shared adapter contract; install project lifecycle instructions automatically without replacing unrelated AGENTS.md or CLAUDE.md content.
|
|
7
|
+
- Add selective project host connection/removal and thin Codex/Claude plugin assets using the same Python runtime.
|
|
8
|
+
- Add `run find`, `task start/checkpoint/complete` and four additive MCP tools. Preserve the original twelve MCP tools, public SDK imports, semantic skill bundles and SQLite v1 stores.
|
|
9
|
+
- Add a real, paired Codex coding experiment with fresh sessions after stages 2/4/6, independent correctness checks, actual host-reported token usage and explicit measurement limitations.
|
|
10
|
+
- Preserve the existing external smolagents acceptance. Claude Code live acceptance remains unexecuted; a reproducible manual procedure is provided.
|
|
11
|
+
- Execution-state persistence does not control native host transcripts or guarantee token, latency or cost savings.
|
|
12
|
+
|
|
13
|
+
## 0.1.2 — desktop connection and host acceptance
|
|
14
|
+
|
|
15
|
+
- Explicit `connect claude-desktop` / `disconnect claude-desktop` commands, with Microsoft Store and standalone Windows detection, macOS defaults and explicit config paths.
|
|
16
|
+
- Preserve unrelated app settings, isolate server names per project, serialize connection writes, refuse user-modified entries, and roll back configuration on receipt write failure.
|
|
17
|
+
- Include desktop connections in `doctor`; retain state when disconnecting.
|
|
18
|
+
- Disambiguate the source resource directory for root-level SKILL.md files.
|
|
19
|
+
- Add a reproducible synthetic invoice acceptance exercise and report actual host successes and blockers separately.
|
|
20
|
+
- Codex completed a real MCP review and handoff. Follow-up verification confirmed that Claude Desktop resumed the same run after the initial permission/timeout interruption, saved second-review evidence and completed at revision 4. Antigravity desktop reported tools unavailable; IDE login was not configured. These are not three passing end-to-end host tests.
|
|
21
|
+
|
|
22
|
+
## 0.1.1 — public distribution
|
|
23
|
+
|
|
24
|
+
- Self-contained PyPI package description with installation and SDK examples.
|
|
25
|
+
- GitHub OIDC publishing workflow: full test matrix, clean installation, optional TestPyPI rehearsal, PyPI and downloaded package verification.
|
|
26
|
+
- Verify registry file hashes against the tested artifacts before installation.
|
|
27
|
+
- No changes to the runtime or storage contract from 0.1.0.
|
|
28
|
+
|
|
29
|
+
## 0.1.0 — initial alpha
|
|
30
|
+
|
|
31
|
+
- Source-preserving SKILL.md conversion and Python test tracking profile.
|
|
32
|
+
- Source-bound semantic proposal preparation/validation and optional stateless JSON model adapter.
|
|
33
|
+
- Immutable skill definitions and source drift detection.
|
|
34
|
+
- SQLite revisions, ownership handoff, operation intents and explicit reconciliation.
|
|
35
|
+
- Managed async runtime with bounded context and tool argument validation.
|
|
36
|
+
- Project-scoped Codex, Claude Code and Antigravity skill/MCP installers.
|
|
37
|
+
- STDIO MCP server, protocol smoke test, CLI and text artifact store.
|
|
38
|
+
- Offline examples, cross-platform CI and package checks.
|
|
39
|
+
|
|
40
|
+
Not included: native transcript replacement, arbitrary code rewriting, live certification of every host version, distributed multi-tenant storage, automatic external-effect rollback or reproduced paper benchmarks.
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
Metadata-Version: 2.4
|
|
2
2
|
Name: skillstate-kit
|
|
3
|
-
Version: 0.
|
|
3
|
+
Version: 0.2.0
|
|
4
4
|
Summary: Compile existing agent skills into portable, validated execution state.
|
|
5
5
|
Project-URL: Homepage, https://github.com/Atakan-Emre/skillstate-kit
|
|
6
6
|
Project-URL: Repository, https://github.com/Atakan-Emre/skillstate-kit
|
|
@@ -39,26 +39,42 @@ Description-Content-Type: text/markdown
|
|
|
39
39
|
|
|
40
40
|
# skillstate-kit
|
|
41
41
|
|
|
42
|
-
**
|
|
42
|
+
**A portable, validated execution-state layer for long-running AI agents.**
|
|
43
43
|
|
|
44
|
-
|
|
44
|
+
Maintain task progress, observations, artifacts and operation state outside conversational history. Use your existing coding agent through CLI/MCP and project integrations, or embed the same canonical engine in Python. Semantic generation from existing skills remains supported. No custom Python agent is required for host integration.
|
|
45
45
|
|
|
46
|
-
Python 3.11+ · MIT license ·
|
|
46
|
+
Python 3.11+ · MIT license · Alpha
|
|
47
47
|
|
|
48
48
|
## Install
|
|
49
49
|
|
|
50
50
|
```bash
|
|
51
|
-
python -m pip install "skillstate-kit[mcp
|
|
52
|
-
skillstate
|
|
51
|
+
python -m pip install "skillstate-kit[mcp]"
|
|
52
|
+
skillstate init
|
|
53
|
+
skillstate doctor --mcp
|
|
53
54
|
```
|
|
54
55
|
|
|
55
|
-
|
|
56
|
+
Run setup in your project. Host detection installs the applicable project lifecycle and generator skills; unrelated instructions/configuration are preserved. After reloading discovery, use your agent normally. The integration guides it to inspect compatible runs, preserve completed work, record evidence and validate completion.
|
|
57
|
+
|
|
58
|
+
The base package also supports `pip install skillstate-kit` and `skillstate init` through CLI instructions. MCP is optional. The core requires no model account; your host supplies its model. `skillstate demo` is an explicitly scripted offline example.
|
|
56
59
|
|
|
57
60
|
Optional extras:
|
|
58
61
|
|
|
59
62
|
- `mcp`: a project-scoped STDIO MCP server and connection diagnostics.
|
|
60
63
|
- `http`: a stateless adapter for a configured JSON-compatible chat-completions endpoint.
|
|
61
64
|
|
|
65
|
+
## Task lifecycle
|
|
66
|
+
|
|
67
|
+
The optional task profile provides `task_start`, `task_checkpoint` and `task_complete` over CLI/MCP. Milestones reference artifacts, completion time and resource fingerprints; repeated milestones require an explicit revalidation reason. Completion validates required steps, evidence integrity, blockers and freshness. Agent-reported evidence does not independently prove business truth. Existing semantic state schemas and run APIs remain supported.
|
|
68
|
+
|
|
69
|
+
`skillstate hosts detect` lists discovery evidence. `skillstate init --host codex --mcp` or `--host claude-code --mcp` selects a project integration. `skillstate doctor codex` checks that integration. `skillstate disconnect codex` reverses its owned setup while retaining state.
|
|
70
|
+
|
|
71
|
+
Execution-state optimization does not remove a host's native transcript. Token, latency and cost improvements must be measured independently; smaller application context is not proof of provider savings.
|
|
72
|
+
|
|
73
|
+
A real paired Codex coding trial passed 29 independent checks in both modes and
|
|
74
|
+
preserved the integrated run across four fresh sessions. It used more tokens and
|
|
75
|
+
wall time with SkillState on that small task. See the
|
|
76
|
+
[measurement report](https://github.com/Atakan-Emre/skillstate-kit/blob/main/docs/benchmarks/coding-agent.md).
|
|
77
|
+
|
|
62
78
|
## Use an existing skill
|
|
63
79
|
|
|
64
80
|
Run these commands in your target project. Replace `skills/qa/SKILL.md` with your existing skill file:
|
|
@@ -83,6 +99,19 @@ The preparation response contains source text, its fingerprint and the proposal
|
|
|
83
99
|
|
|
84
100
|
You can also supply an explicitly configured endpoint with `--base-url` and `--model`. Terminal commands do not borrow your IDE's model credentials.
|
|
85
101
|
|
|
102
|
+
## Claude Desktop Chat
|
|
103
|
+
|
|
104
|
+
Version 0.1.2 adds an explicit local connection, separate from Claude Code:
|
|
105
|
+
|
|
106
|
+
```console
|
|
107
|
+
skillstate connect claude-desktop
|
|
108
|
+
skillstate doctor --mcp
|
|
109
|
+
```
|
|
110
|
+
|
|
111
|
+
Fully quit and reopen Claude Desktop, use **Chat**, and approve the tool calls you intend to allow. The installer preserves unrelated settings and uses a separate server name per project. Windows standalone/Microsoft Store and macOS paths are supported; `--config PATH` selects a custom location. If two Windows configs exist, choose explicitly. Disconnect with `skillstate disconnect claude-desktop`; run data remains intact.
|
|
112
|
+
|
|
113
|
+
Live acceptance evidence: Codex used eight real MCP calls to record a synthetic invoice review and hand it to Claude Desktop. After an initial permission/timeout interruption, Desktop resumed the same run, saved a second-review artifact and completed at revision 4. A separate process verified the stored state, both artifacts and the event history. This representative Codex-to-Claude Desktop Chat success does not establish universal desktop compatibility.
|
|
114
|
+
|
|
86
115
|
## Python integration
|
|
87
116
|
|
|
88
117
|
This complete example uses a scripted model. Replace `model` and `record` with your own model and service functions:
|
|
@@ -0,0 +1,264 @@
|
|
|
1
|
+
<p align="center">
|
|
2
|
+
<img src="docs/assets/logo.svg" alt="skillstate-kit — State that survives the next session." width="1040">
|
|
3
|
+
</p>
|
|
4
|
+
|
|
5
|
+
<p align="center">
|
|
6
|
+
<a href="https://pypi.org/project/skillstate-kit/"><img src="https://img.shields.io/pypi/v/skillstate-kit?color=246b52" alt="PyPI version"></a>
|
|
7
|
+
<img src="https://img.shields.io/badge/Python-3.11%2B-246b52" alt="Python 3.11 or newer">
|
|
8
|
+
<a href="LICENSE"><img src="https://img.shields.io/badge/License-MIT-246b52" alt="MIT license"></a>
|
|
9
|
+
<img src="https://img.shields.io/badge/Status-Alpha-b76538" alt="Alpha release">
|
|
10
|
+
</p>
|
|
11
|
+
|
|
12
|
+
<p align="center">
|
|
13
|
+
<a href="https://pypi.org/project/skillstate-kit/">PyPI</a> ·
|
|
14
|
+
<a href="#quick-start">Quick start</a> ·
|
|
15
|
+
<a href="#python-integration">Python example</a> ·
|
|
16
|
+
<a href="docs/README_TR.md">Türkçe</a> ·
|
|
17
|
+
<a href="docs/host-acceptance.md">Verified integrations</a>
|
|
18
|
+
</p>
|
|
19
|
+
|
|
20
|
+
# skillstate-kit
|
|
21
|
+
|
|
22
|
+
**A portable, validated execution-state layer for long-running AI agents.**
|
|
23
|
+
|
|
24
|
+
Keep task progress, observations, artifacts and operation state outside conversational history. Use your existing agent through project skills, CLI or MCP, or embed the canonical runtime in Python. Task-specific semantic state remains supported; no custom Python agent is required for host integration.
|
|
25
|
+
|
|
26
|
+
By reducing redundant work and providing bounded application state, skillstate-kit is designed to improve long-horizon execution efficiency. Provider-level token, latency and cost effects depend on the host and require separate measurement.
|
|
27
|
+
|
|
28
|
+
**Available on [PyPI](https://pypi.org/project/skillstate-kit/), open source under the [MIT license](LICENSE).** The project is in alpha; this checkout documents the `0.2.0` release candidate. Install the package from PyPI or explore and contribute to this repository.
|
|
29
|
+
|
|
30
|
+
**External-project validation:** the existing Hugging Face smolagents SQL agent
|
|
31
|
+
produced the same five correct results on 1,000 synthetic receipts before and
|
|
32
|
+
after integration. The managed run resumed in a new process after an injected
|
|
33
|
+
interruption, without repeating completed queries.
|
|
34
|
+
[Read the experiment and its scope](docs/smolagents-acceptance.md).
|
|
35
|
+
|
|
36
|
+
**Coding-agent measurement:** four fresh Codex sessions completed the same persisted
|
|
37
|
+
task and passed 29 independent checks. This small paired trial used **more**, not
|
|
38
|
+
fewer, tokens and wall time with SkillState. [Read the measured results and
|
|
39
|
+
limitations](docs/benchmarks/coding-agent.md) before making performance claims.
|
|
40
|
+
|
|
41
|
+
## Why skillstate-kit?
|
|
42
|
+
|
|
43
|
+
| Capability | What it gives you |
|
|
44
|
+
|---|---|
|
|
45
|
+
| Source-preserving generation | Keep your instructions and add structured progress tracking. |
|
|
46
|
+
| Domain-specific state | Let your agent propose fields and steps, then validate the proposal against the source. |
|
|
47
|
+
| Durable checkpoints | SQLite transactions, revision checks, evidence artifacts and explicit ownership handoff. |
|
|
48
|
+
| Controlled operations | Validate decisions before execution and retain uncertain outcomes for reconciliation. |
|
|
49
|
+
| Python, CLI and MCP | Use one state engine across application code and supported agent integrations. |
|
|
50
|
+
|
|
51
|
+
## Quick start
|
|
52
|
+
|
|
53
|
+
### 1. Install
|
|
54
|
+
|
|
55
|
+
Use Python **3.11 or newer**, preferably in your project's virtual environment:
|
|
56
|
+
|
|
57
|
+
```bash
|
|
58
|
+
python -m pip install "skillstate-kit[mcp]"
|
|
59
|
+
skillstate init
|
|
60
|
+
skillstate doctor --mcp
|
|
61
|
+
```
|
|
62
|
+
|
|
63
|
+
Run setup inside your target project. `init` detects supported project hosts, installs lifecycle/generator skills and preserves unrelated instructions/configuration. Reload host discovery, then work normally—for example: **“Implement duration parsing for this project.”** The integration guides your agent to inspect existing runs, record evidence-backed milestones and validate completion.
|
|
64
|
+
|
|
65
|
+
The base `pip install skillstate-kit` also works with `skillstate init` through CLI instructions. MCP is optional. `skillstate demo` remains an explicitly scripted offline example. See [host setup and distribution](docs/integrations.md) and [execution state versus chat history](docs/execution-state.md).
|
|
66
|
+
|
|
67
|
+
For Python-only use, install `skillstate-kit`. Add the `http` extra when you need a configured JSON-compatible model endpoint: `python -m pip install "skillstate-kit[mcp,http]"`.
|
|
68
|
+
|
|
69
|
+
### 2. Convert your existing skill
|
|
70
|
+
|
|
71
|
+
Run these commands **inside the project you want to integrate**. Replace `skills/qa/SKILL.md` with your existing skill file:
|
|
72
|
+
|
|
73
|
+
```bash
|
|
74
|
+
skillstate init --host codex --mcp
|
|
75
|
+
skillstate generate skills/qa/SKILL.md --name qa-state --install
|
|
76
|
+
skillstate validate qa-state
|
|
77
|
+
skillstate doctor --mcp
|
|
78
|
+
```
|
|
79
|
+
|
|
80
|
+
`init` sets up the selected host; `generate --install` writes the generated skill wrappers for the three project adapters. Original source files remain unchanged. To target another directory, place `--project PATH` before the subcommand.
|
|
81
|
+
|
|
82
|
+
**No existing skill file?** In a Python project, start with the built-in test workflow:
|
|
83
|
+
|
|
84
|
+
```bash
|
|
85
|
+
skillstate generate . --profile python-tests --name tests-state --install
|
|
86
|
+
skillstate validate tests-state
|
|
87
|
+
```
|
|
88
|
+
|
|
89
|
+
This produces test-oriented tracking instructions from a bounded project inventory. It does not infer every business rule in your application.
|
|
90
|
+
|
|
91
|
+
### 3. Ask your agent to use it
|
|
92
|
+
|
|
93
|
+
After your host discovers the installed skills and MCP server, ask:
|
|
94
|
+
|
|
95
|
+
> Use qa-state for this task. Open a run, record progress and evidence, and leave a checkpoint if the work is interrupted.
|
|
96
|
+
|
|
97
|
+
For agent-assisted **domain-specific generation**, use the installed `generate-skill-state` skill:
|
|
98
|
+
|
|
99
|
+
> Generate skill state from skills/qa/SKILL.md. Propose the fields and steps this procedure needs, validate the proposal, and install the result.
|
|
100
|
+
|
|
101
|
+
Direct CLI generation creates generic tracking state. Agent-assisted generation can propose domain-specific state. **Generating a definition does not execute the task.** The host still runs the procedure and records its progress.
|
|
102
|
+
|
|
103
|
+
## Choose your environment
|
|
104
|
+
|
|
105
|
+
Use `skillstate hosts detect` to inspect discovery evidence and `skillstate doctor codex` or `skillstate doctor claude-code` for focused checks. Detection is not a live execution test. Project instruction blocks encourage automatic lifecycle use; host/model selection and permissions still apply.
|
|
106
|
+
|
|
107
|
+
| Environment | Setup inside your project | Verified scope |
|
|
108
|
+
|---|---|---|
|
|
109
|
+
| Codex | `skillstate init --host codex --mcp` | Real MCP generation and continuation across two model sessions passed. |
|
|
110
|
+
| Claude Desktop **Chat** | `skillstate connect claude-desktop` | Codex → Claude Desktop handoff and independent invoice review passed. |
|
|
111
|
+
| Claude Code | `skillstate init --host claude-code --mcp` | Adapter available; a live Claude Code run has not been verified. |
|
|
112
|
+
| Antigravity | `skillstate init --host antigravity --mcp` | Project skill and MCP configuration adapter available. |
|
|
113
|
+
|
|
114
|
+
For Claude Desktop, **fully quit and reopen the app after connecting**, use Chat, and approve the requested tools in the app. This connection is separate from Claude Code and Cowork. Run `skillstate disconnect claude-desktop` to remove the connection while retaining run data.
|
|
115
|
+
|
|
116
|
+
MCP configurations contain this machine's Python and project paths. Keep them local and repeat setup when changing machines or virtual environments. `doctor --mcp` checks the local server; it does not prove that a host's model can call it.
|
|
117
|
+
|
|
118
|
+
See the [installation guide](docs/compatibility.md), [actual acceptance evidence](docs/host-acceptance.md), and [repeatable two-reviewer exercise](examples/desktop_acceptance/README.md).
|
|
119
|
+
|
|
120
|
+
## Evidence-backed task lifecycle
|
|
121
|
+
|
|
122
|
+
For a normal multi-step task, the agent can use `task_start`, `task_checkpoint` and `task_complete`. Identical goal/plan startup reuses its deterministic run ID; an intentionally new task can supply a new ID. A checkpoint references real artifacts and relevant resource fingerprints. Repeating a completed milestone requires a stated revalidation reason. Completion checks evidence integrity, remaining work, blockers and freshness—not the truth of arbitrary business claims.
|
|
123
|
+
|
|
124
|
+
Existing semantic skills retain `run_open`/`run_update` and their own domain schema. All paths share one SQLite store and the existing pending/unknown operation protocol. See the [CLI reference](docs/cli.md).
|
|
125
|
+
|
|
126
|
+
## Inspect a checkpoint
|
|
127
|
+
|
|
128
|
+
Open a run for a generated definition, then inspect its state and audit trail:
|
|
129
|
+
|
|
130
|
+
```bash
|
|
131
|
+
skillstate run open qa-state --owner codex --id qa-001
|
|
132
|
+
skillstate run context qa-001
|
|
133
|
+
skillstate run events qa-001
|
|
134
|
+
skillstate status
|
|
135
|
+
```
|
|
136
|
+
|
|
137
|
+
These commands create and inspect a run; they do not execute its steps. Use a new run ID for each task. To continue an existing task, read its context and resume that run rather than opening it again. Ownership changes use an explicit handoff with the current revision; see the [CLI lifecycle guide](docs/cli.md).
|
|
138
|
+
|
|
139
|
+
## Python integration
|
|
140
|
+
|
|
141
|
+
Save the following as `example.py`, then run `python example.py`. This complete example uses a **scripted model and a simulated tool**, so it works without credentials:
|
|
142
|
+
|
|
143
|
+
```python
|
|
144
|
+
import asyncio
|
|
145
|
+
|
|
146
|
+
from skillstate import Skill, SkillRuntime, SQLiteStore, Tool, ToolResult
|
|
147
|
+
|
|
148
|
+
async def main():
|
|
149
|
+
skill = Skill(
|
|
150
|
+
"record-job",
|
|
151
|
+
"Record the job; finish after its result is confirmed.",
|
|
152
|
+
{
|
|
153
|
+
"type": "object",
|
|
154
|
+
"properties": {"recorded": {"type": "boolean"}},
|
|
155
|
+
"required": ["recorded"],
|
|
156
|
+
"additionalProperties": False,
|
|
157
|
+
},
|
|
158
|
+
{"recorded": False},
|
|
159
|
+
)
|
|
160
|
+
|
|
161
|
+
def model(context):
|
|
162
|
+
if context["state"]["recorded"]:
|
|
163
|
+
return {"patch": [], "action": None, "done": True}
|
|
164
|
+
return {
|
|
165
|
+
"patch": [{"op": "set", "path": "/recorded", "value": True}],
|
|
166
|
+
"action": {"name": "record", "arguments": {}},
|
|
167
|
+
"done": False,
|
|
168
|
+
}
|
|
169
|
+
|
|
170
|
+
def record(arguments, operation_id):
|
|
171
|
+
return ToolResult(True, {"recorded": True, "operation_id": operation_id})
|
|
172
|
+
|
|
173
|
+
tool = Tool("record", "Record a job", {"type": "object", "additionalProperties": False}, record)
|
|
174
|
+
with SQLiteStore() as store:
|
|
175
|
+
store.create("example", skill, "worker", {"job": "example"})
|
|
176
|
+
runtime = SkillRuntime(store, model, [tool], completion_check=lambda s: s["recorded"])
|
|
177
|
+
result = await runtime.run("example", "worker")
|
|
178
|
+
print(result["status"], result["state"])
|
|
179
|
+
|
|
180
|
+
if __name__ == "__main__":
|
|
181
|
+
asyncio.run(main())
|
|
182
|
+
```
|
|
183
|
+
|
|
184
|
+
Expected output:
|
|
185
|
+
|
|
186
|
+
```text
|
|
187
|
+
completed {'recorded': True}
|
|
188
|
+
```
|
|
189
|
+
|
|
190
|
+
Replace `model(context)` with your synchronous or asynchronous model callback, and `record` with a real tool that checks its result. Forward `operation_id` to the external service's idempotency facility when supported.
|
|
191
|
+
|
|
192
|
+
The example uses an in-memory store for easy reruns. Use `SQLiteStore("jobs.sqlite3")` for durable storage. Create each run once; resume by reopening that database and calling `runtime.run` with the existing run ID and owner. See the [runnable source](examples/managed_runtime.py) and [runtime contract](docs/architecture.md).
|
|
193
|
+
|
|
194
|
+
### Model decision contract
|
|
195
|
+
|
|
196
|
+
```json
|
|
197
|
+
{
|
|
198
|
+
"patch": [{"op": "set", "path": "/recorded", "value": true}],
|
|
199
|
+
"action": {"name": "record", "arguments": {}},
|
|
200
|
+
"done": false
|
|
201
|
+
}
|
|
202
|
+
```
|
|
203
|
+
|
|
204
|
+
Finish with `{"patch": [], "action": null, "done": true}`. State updates use explicit `set`/`delete` operations and JSON Pointer paths. Setting `null` does not delete a key. No reasoning-trace field is accepted.
|
|
205
|
+
|
|
206
|
+
## How it fits together
|
|
207
|
+
|
|
208
|
+
```mermaid
|
|
209
|
+
flowchart LR
|
|
210
|
+
Source[Existing skill or Python project] --> Generate[Tracking state or semantic proposal]
|
|
211
|
+
Generate --> Validate[Schema and source validation]
|
|
212
|
+
Validate --> Bundle[Versioned skill bundle]
|
|
213
|
+
Bundle --> Host[Agent via CLI or MCP]
|
|
214
|
+
Bundle --> Runtime[Python managed runtime]
|
|
215
|
+
Host <--> Store[(State, evidence and operation journal)]
|
|
216
|
+
Runtime <--> Store
|
|
217
|
+
```
|
|
218
|
+
|
|
219
|
+
| Execution mode | What skillstate-kit controls |
|
|
220
|
+
|---|---|
|
|
221
|
+
| Native agent integration | Validated state, checkpoints, explicit operation records and bounded returned context. The host owns its model and tools. |
|
|
222
|
+
| Python managed runtime | Model context, decision validation, registered tool execution, retries and recorded outcomes. |
|
|
223
|
+
|
|
224
|
+
Native integrations do not replace a host's conversation history or intercept every native tool. For the paper's history-free model-input pattern, use the managed runtime with a stateless model adapter. A terminal process cannot borrow your IDE's model subscription automatically.
|
|
225
|
+
|
|
226
|
+
Definitions live in `.skillstate/definitions/`; private run data and artifacts live in `.skillstate/local/`. Source changes are detected before new runs. Existing runs retain their pinned definition and expose source drift.
|
|
227
|
+
|
|
228
|
+
## Failure behavior
|
|
229
|
+
|
|
230
|
+
| Event | Behavior |
|
|
231
|
+
|---|---|
|
|
232
|
+
| Invalid JSON, state or tool arguments | Reject before tool execution |
|
|
233
|
+
| Stale revision or wrong owner | Reject; require a fresh read/handoff |
|
|
234
|
+
| Tool reports failure | Keep prior state and record its observation |
|
|
235
|
+
| Timeout, exception or invalid tool result | Mark `unknown`; require reconciliation |
|
|
236
|
+
| Process exits after reserving an operation | Retain the pending intent; do not replay automatically |
|
|
237
|
+
| Event write fails during result commit | Roll back state, operation outcome and event together |
|
|
238
|
+
| User edits installed files/configuration | Preserve changes and report a conflict |
|
|
239
|
+
|
|
240
|
+
SQLite does not create a distributed transaction with external APIs. A checkpoint cannot undo an external side effect. Owner IDs coordinate local clients; they are not an authentication boundary. Read [the execution contract](docs/architecture.md) before using mutating tools.
|
|
241
|
+
|
|
242
|
+
## Development
|
|
243
|
+
|
|
244
|
+
```bash
|
|
245
|
+
python -m pip install uv
|
|
246
|
+
uv sync --locked --extra dev
|
|
247
|
+
uv run ruff check src tests examples
|
|
248
|
+
uv run ruff format --check src tests examples
|
|
249
|
+
uv run pytest --cov=skillstate --cov-fail-under=85
|
|
250
|
+
uv run python -m build
|
|
251
|
+
uv run twine check dist/*
|
|
252
|
+
```
|
|
253
|
+
|
|
254
|
+
CI tests the Python/OS matrix, builds distributions and checks wheel installation. Tests include real MCP sessions, competing connections, crash recovery, configuration preservation, invalid schemas and source drift. Coverage is a regression signal, not proof of correctness.
|
|
255
|
+
|
|
256
|
+
## Research and attribution
|
|
257
|
+
|
|
258
|
+
An **independent implementation** inspired by [SKILL.state: Scalable Long-Horizon Agent Skills](https://arxiv.org/abs/2608.26263), by Sanket Badhe, Priyanka Tiwari and Jonghyun Chung. Not affiliated with or endorsed by the authors or their institutions.
|
|
259
|
+
|
|
260
|
+
The managed runtime follows the explicit-state input pattern. Our compiler, host adapters, explicit patch format and operation journal are engineering extensions. No paper benchmark scores are claimed. [skill-state-minimal](https://github.com/kissishka/skill-state-minimal) was inspected as a comparison; its source is not bundled here.
|
|
261
|
+
|
|
262
|
+
## License and contributing
|
|
263
|
+
|
|
264
|
+
[MIT](LICENSE). See [CONTRIBUTING.md](CONTRIBUTING.md), [SECURITY.md](SECURITY.md), [CHANGELOG.md](CHANGELOG.md) and [release guidance](docs/releasing.md).
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
## Reporting
|
|
4
4
|
|
|
5
|
-
Use GitHub private vulnerability reporting
|
|
5
|
+
Use [GitHub private vulnerability reporting](https://github.com/Atakan-Emre/skillstate-kit/security/advisories/new) for security issues. Do not put credentials or exploitable production details in a public issue or test another person's services.
|
|
6
6
|
|
|
7
7
|
## Scope
|
|
8
8
|
|
|
@@ -1,25 +1,41 @@
|
|
|
1
1
|
# skillstate-kit
|
|
2
2
|
|
|
3
|
-
**
|
|
3
|
+
**A portable, validated execution-state layer for long-running AI agents.**
|
|
4
4
|
|
|
5
|
-
|
|
5
|
+
Maintain task progress, observations, artifacts and operation state outside conversational history. Use your existing coding agent through CLI/MCP and project integrations, or embed the same canonical engine in Python. Semantic generation from existing skills remains supported. No custom Python agent is required for host integration.
|
|
6
6
|
|
|
7
|
-
Python 3.11+ · MIT license ·
|
|
7
|
+
Python 3.11+ · MIT license · Alpha
|
|
8
8
|
|
|
9
9
|
## Install
|
|
10
10
|
|
|
11
11
|
```bash
|
|
12
|
-
python -m pip install "skillstate-kit[mcp
|
|
13
|
-
skillstate
|
|
12
|
+
python -m pip install "skillstate-kit[mcp]"
|
|
13
|
+
skillstate init
|
|
14
|
+
skillstate doctor --mcp
|
|
14
15
|
```
|
|
15
16
|
|
|
16
|
-
|
|
17
|
+
Run setup in your project. Host detection installs the applicable project lifecycle and generator skills; unrelated instructions/configuration are preserved. After reloading discovery, use your agent normally. The integration guides it to inspect compatible runs, preserve completed work, record evidence and validate completion.
|
|
18
|
+
|
|
19
|
+
The base package also supports `pip install skillstate-kit` and `skillstate init` through CLI instructions. MCP is optional. The core requires no model account; your host supplies its model. `skillstate demo` is an explicitly scripted offline example.
|
|
17
20
|
|
|
18
21
|
Optional extras:
|
|
19
22
|
|
|
20
23
|
- `mcp`: a project-scoped STDIO MCP server and connection diagnostics.
|
|
21
24
|
- `http`: a stateless adapter for a configured JSON-compatible chat-completions endpoint.
|
|
22
25
|
|
|
26
|
+
## Task lifecycle
|
|
27
|
+
|
|
28
|
+
The optional task profile provides `task_start`, `task_checkpoint` and `task_complete` over CLI/MCP. Milestones reference artifacts, completion time and resource fingerprints; repeated milestones require an explicit revalidation reason. Completion validates required steps, evidence integrity, blockers and freshness. Agent-reported evidence does not independently prove business truth. Existing semantic state schemas and run APIs remain supported.
|
|
29
|
+
|
|
30
|
+
`skillstate hosts detect` lists discovery evidence. `skillstate init --host codex --mcp` or `--host claude-code --mcp` selects a project integration. `skillstate doctor codex` checks that integration. `skillstate disconnect codex` reverses its owned setup while retaining state.
|
|
31
|
+
|
|
32
|
+
Execution-state optimization does not remove a host's native transcript. Token, latency and cost improvements must be measured independently; smaller application context is not proof of provider savings.
|
|
33
|
+
|
|
34
|
+
A real paired Codex coding trial passed 29 independent checks in both modes and
|
|
35
|
+
preserved the integrated run across four fresh sessions. It used more tokens and
|
|
36
|
+
wall time with SkillState on that small task. See the
|
|
37
|
+
[measurement report](https://github.com/Atakan-Emre/skillstate-kit/blob/main/docs/benchmarks/coding-agent.md).
|
|
38
|
+
|
|
23
39
|
## Use an existing skill
|
|
24
40
|
|
|
25
41
|
Run these commands in your target project. Replace `skills/qa/SKILL.md` with your existing skill file:
|
|
@@ -44,6 +60,19 @@ The preparation response contains source text, its fingerprint and the proposal
|
|
|
44
60
|
|
|
45
61
|
You can also supply an explicitly configured endpoint with `--base-url` and `--model`. Terminal commands do not borrow your IDE's model credentials.
|
|
46
62
|
|
|
63
|
+
## Claude Desktop Chat
|
|
64
|
+
|
|
65
|
+
Version 0.1.2 adds an explicit local connection, separate from Claude Code:
|
|
66
|
+
|
|
67
|
+
```console
|
|
68
|
+
skillstate connect claude-desktop
|
|
69
|
+
skillstate doctor --mcp
|
|
70
|
+
```
|
|
71
|
+
|
|
72
|
+
Fully quit and reopen Claude Desktop, use **Chat**, and approve the tool calls you intend to allow. The installer preserves unrelated settings and uses a separate server name per project. Windows standalone/Microsoft Store and macOS paths are supported; `--config PATH` selects a custom location. If two Windows configs exist, choose explicitly. Disconnect with `skillstate disconnect claude-desktop`; run data remains intact.
|
|
73
|
+
|
|
74
|
+
Live acceptance evidence: Codex used eight real MCP calls to record a synthetic invoice review and hand it to Claude Desktop. After an initial permission/timeout interruption, Desktop resumed the same run, saved a second-review artifact and completed at revision 4. A separate process verified the stored state, both artifacts and the event history. This representative Codex-to-Claude Desktop Chat success does not establish universal desktop compatibility.
|
|
75
|
+
|
|
47
76
|
## Python integration
|
|
48
77
|
|
|
49
78
|
This complete example uses a scripted model. Replace `model` and `record` with your own model and service functions:
|
|
@@ -0,0 +1,67 @@
|
|
|
1
|
+

|
|
2
|
+
|
|
3
|
+
# Türkçe başlangıç
|
|
4
|
+
|
|
5
|
+
[PyPI paketi](https://pypi.org/project/skillstate-kit/) · [Çalışan Python örneği](../README.md#python-integration) · [Kabul testleri](host-acceptance.md)
|
|
6
|
+
|
|
7
|
+
skillstate-kit, AI agent'lar için taşınabilir bir görev yürütme durumu katmanıdır. İlerleme, kanıtlar ve işlem durumu konuşma geçmişinden ayrı saklanır. Python, CLI ve MCP aynı çekirdeği kullanır; özel bir Python agent yazmanız gerekmez. MIT lisansıyla açık kaynak olarak yayımlanır. Sürüm alpha durumundadır.
|
|
8
|
+
|
|
9
|
+
**Dış proje doğrulaması:** Hugging Face smolagents'ın mevcut SQL agent'ı, 1.000 sentetik kayıt üzerindeki beş kontrolde entegrasyon öncesi ve sonrası aynı doğru sonuçları üretti. Skillstate ile çalışan süreç zorla kapatıldıktan sonra yeni süreç kalan iki sorguyla tamamlandı; bitmiş sorgular tekrarlanmadı. [Deneyin kapsamı, sonuçları ve tekrar çalıştırma adımları](smolagents-acceptance.md).
|
|
10
|
+
|
|
11
|
+
## Kurulum
|
|
12
|
+
|
|
13
|
+
Python 3.11 veya üzeri bir ortamda:
|
|
14
|
+
|
|
15
|
+
```text
|
|
16
|
+
python -m pip install "skillstate-kit[mcp]"
|
|
17
|
+
```
|
|
18
|
+
|
|
19
|
+
Ardından kullanacağınız projenin klasöründe:
|
|
20
|
+
|
|
21
|
+
```text
|
|
22
|
+
skillstate init
|
|
23
|
+
skillstate doctor --mcp
|
|
24
|
+
```
|
|
25
|
+
|
|
26
|
+
0.2.0 ile `init`, algılanan proje ortamlarını seçer ve MCP kuruluysa yapılandırır. `skillstate hosts detect` seçim gerekçelerini gösterir. Yalnızca `pip install skillstate-kit` de yeterlidir: temel paket CLI üzerinden çalışır. Belirli bir ortam için `skillstate init --host codex --mcp` kullanabilirsiniz. MCP ayarları bu makineye ait Python/proje yollarını içerir.
|
|
27
|
+
|
|
28
|
+
Agent oturumunu yeniledikten sonra normal görevinizi verin: “Bu projeye JWT doğrulaması ekle ve testlerini çalıştır.” Kurulan yönergeler agent'ı mevcut görevi bulmaya, uygun kaydı sürdürmeye ve kanıtlı adımlar kaydetmeye yönlendirir. Host'un yönergeleri izlemesi gerekir; paket her araç çağrısını zorla denetlemez.
|
|
29
|
+
|
|
30
|
+
`skillstate run find` ile kayıtları, `skillstate run context RUN_ID` ile güncel durumu görebilirsiniz. Tamamlanan adımı yeniden kaydetmek açık bir yeniden doğrulama gerekçesi ister. Kaynak dosyası değişirse yalnızca o dosyaya bağlı kanıtlar eski olarak işaretlenir. [Durum ve doğrulama sözleşmesi](execution-state.md).
|
|
31
|
+
|
|
32
|
+
**Ölçüm:** Yeni [kodlama deneyi](benchmarks/coding-agent.md) gerçek token ve süreyi ayrı raporlar. Kalıcı durumun çalışması otomatik maliyet tasarrufu anlamına gelmez; host'un konuşma geçmişini bu paket silemez.
|
|
33
|
+
|
|
34
|
+
## Otomatik üretim
|
|
35
|
+
|
|
36
|
+
Claude Desktop'ın **Chat** bölümünü kullanıyorsanız ayrıca `skillstate connect claude-desktop` çalıştırıp uygulamadan tamamen çıkın ve yeniden açın. Bu bağlantı Claude Code kurulumundan ayrıdır. Araç izinlerini uygulama içinde siz verirsiniz. `skillstate disconnect claude-desktop` bağlantıyı kaldırır, görev durumunu korur. Windows Store sürümü de desteklenir; birden fazla ayar dosyası bulunursa `--config DOSYA` ile seçilir.
|
|
37
|
+
|
|
38
|
+
Gerçek uygulama testlerinin sonucu [kabul raporunda](host-acceptance.md), tekrar edilebilir örnek ise [örnek projede](../examples/desktop_acceptance/README.md). Codex'te gerçek MCP üretimi ve ayrı oturumdan devam etme; ayrıca Codex → Claude Desktop Chat devri ve ikinci incelemeyle tamamlama doğrulandı. Claude Code için canlı uygulama testi yapılmadı.
|
|
39
|
+
|
|
40
|
+
Kurulumdan sonra agent'a “Generate skill state; bu skill'i durum yapısına dönüştür” diyebilirsiniz. Üretici kaynak envanterini hazırlayıp agent'ın önerdiği şemayı doğrular.
|
|
41
|
+
|
|
42
|
+
Doğrudan `generate`, yönergeleri koruyan sınırlı bir genel ilerleme şeması oluşturur. Alana özel dönüşüm için agent destekli `--prepare`/`--proposal` akışı veya yapılandırılmış `--base-url`/`--model` kullanılır. Terminal, IDE model hesabını otomatik kullanmaz.
|
|
43
|
+
|
|
44
|
+
## Çalıştırma ve devam
|
|
45
|
+
|
|
46
|
+
Üretim komutu görevi çalıştırmaz. Agent'a “qa-state skill'ini bu görevde kullan, ilerlemeyi ve kanıtları kaydet” diyerek çalışmayı başlatın. Durumu terminalden incelemek için:
|
|
47
|
+
|
|
48
|
+
```text
|
|
49
|
+
skillstate run open qa-state --owner codex --id qa-001
|
|
50
|
+
skillstate run context qa-001
|
|
51
|
+
skillstate run events qa-001
|
|
52
|
+
skillstate status
|
|
53
|
+
```
|
|
54
|
+
|
|
55
|
+
Her yeni görev için farklı bir run ID kullanın. Var olan göreve devam ederken yeniden açmak yerine mevcut durumu okuyun. Python ile doğrudan kullanım için [kopyalanıp çalıştırılabilen örneğe](../README.md#python-integration) bakın.
|
|
56
|
+
|
|
57
|
+
Üretilen skill, state açma, okuma, güncelleme, işlem sonucu kaydetme ve ortamlar arasında devretme adımlarını içerir. Python projeleri `SkillRuntime` ile bütün model/araç döngüsünü yönetebilir.
|
|
58
|
+
|
|
59
|
+
Native skill host'un konuşma geçmişini silmez. Makaledeki geçmiş taşımayan model girdisini kurmak için stateless adapter ile managed runtime kullanın.
|
|
60
|
+
|
|
61
|
+
`pending` veya `unknown` bir dış işlemin sonucunun belirsiz olduğunu gösterir. Tekrarlamadan önce dış sistemden sonucu kontrol edin ve açıkça uzlaştırın. Checkpoint dış dünyadaki işlemi geri almaz.
|
|
62
|
+
|
|
63
|
+
## Doğrulama
|
|
64
|
+
|
|
65
|
+
`skillstate demo` model hesabı gerektirmeyen kontrollü örnektir. `doctor --mcp` gerçek yerel MCP bağlantısı kurar. Bunlar bütün IDE sürümlerinin davranışını veya makalenin performans sonuçlarını doğrulamaz.
|
|
66
|
+
|
|
67
|
+
Ayrıntılar: [CLI](cli.md), [mimari](architecture.md), [ortam desteği](compatibility.md), [test matrisi](validation.md).
|
|
@@ -0,0 +1,40 @@
|
|
|
1
|
+
# Architecture audit and implementation contract — 2026-09-07
|
|
2
|
+
|
|
3
|
+
The pre-change architecture is intentionally retained: models/schema/jsonio → store/runtime → ProjectService → CLI/MCP → hosts/desktop. Flat modules already enforce useful boundaries; moving them into packages would break imports without product benefit.
|
|
4
|
+
|
|
5
|
+
| Requested capability | Initial classification | Decision |
|
|
6
|
+
|---|---|---|
|
|
7
|
+
| Python SDK, transactional SQLite, revisions, ownership | ALREADY EXISTS | Preserve public contracts and database v1 |
|
|
8
|
+
| Pending/unknown operation reconciliation | ALREADY EXISTS | Keep existing reservation protocol |
|
|
9
|
+
| Semantic generation and immutable bundles | ALREADY EXISTS | Do not force task schemas onto old bundles |
|
|
10
|
+
| Artifact integrity and bounded runtime context | ALREADY EXISTS | Reuse, keep audit outside model context |
|
|
11
|
+
| Project installers, MCP, host-specific wrappers | PARTIAL | Add lifecycle discovery and isolated adapter registry |
|
|
12
|
+
| Host detection and minimal init | MISSING | Detect without reading credentials; retain explicit host flags |
|
|
13
|
+
| Evidence-backed task milestones and granular freshness | MISSING | Optional task profile via shared service, existing store |
|
|
14
|
+
| Natural task startup and compatible-run discovery | PARTIAL | Lifecycle skill + managed instruction block; no guessing task identity |
|
|
15
|
+
| Codex/Claude skill/plugin distribution | PARTIAL | Thin assets depending on Python CLI/MCP |
|
|
16
|
+
| Source drift | PARTIAL | Existing bundle drift plus milestone resource fingerprints |
|
|
17
|
+
| Cross-host handoff | ALREADY EXISTS | Extend acceptance, never reset state |
|
|
18
|
+
| smolagents correctness/restart acceptance | ALREADY EXISTS | Preserve as a correctness experiment |
|
|
19
|
+
| Coding-agent benchmark and real host measurements | MISSING | Deterministic feature fixture, baseline/integrated + boundaries 2/4/6 |
|
|
20
|
+
| Claude Code live acceptance | MISSING | Execute if available, otherwise reproducible manual procedure |
|
|
21
|
+
| Native transcript compression or universal token savings | SHOULD NOT IMPLEMENT | State explicit host limitations |
|
|
22
|
+
| Command-string deduplication, second state engine, scheduler | SHOULD NOT IMPLEMENT | Use state/evidence/operation semantics |
|
|
23
|
+
|
|
24
|
+
## Changes and compatibility
|
|
25
|
+
|
|
26
|
+
Evolve hosts.py, service.py, cli.py and mcp_server.py; introduce small lifecycle and host-registry modules; add thin plugin assets and benchmark fixtures. Preserve store tables, Skill/Tool/SkillRuntime imports, existing CLI operations, all existing MCP tools, historical bundles and acceptance evidence. New lifecycle APIs are opt-in and cannot retroactively add completion requirements to arbitrary schemas.
|
|
27
|
+
|
|
28
|
+
Risks: preserving edited AGENTS.md/CLAUDE.md sections; shared Codex/Antigravity skill paths; misleading host detection; stale evidence; interpreting repeated calls as unnecessary; assuming usage statistics equal billed cost. Address these with marked-block ownership, conflict-aware rollback, explicit detection evidence, per-resource hashes, reported repetition categories, and unavailable metrics rather than guessed numbers.
|
|
29
|
+
|
|
30
|
+
## Phases and acceptance
|
|
31
|
+
|
|
32
|
+
0. Audit and regression baseline: 113 passed / 1 Windows symlink privilege skip; 86.84% coverage.
|
|
33
|
+
1. Registry/detection and reversible installation; explicit host flags remain supported.
|
|
34
|
+
2. Codex lifecycle instructions and bounded compatible-run summaries.
|
|
35
|
+
3. Claude Code lifecycle and thin plugin distribution, same Python engine.
|
|
36
|
+
4. Evidence-backed milestones, explicit completion and cross-host continuation.
|
|
37
|
+
5. Real coding fixture and baseline/integrated measurements. Session boundaries after stages 2/4/6. Independent correctness, repeated calls, search/read counts, context bytes, available host usage, wall time; cost only from explicitly configured pricing.
|
|
38
|
+
6. Documentation, package smoke checks and full tests with >=85% coverage.
|
|
39
|
+
|
|
40
|
+
Acceptance: old APIs/stores remain readable; no blind operation replay; instructions/config preserve unrelated edits; completed work requires evidence and justified revalidation; one canonical run survives fresh sessions/handoff; MCP and CLI use the same service; smolagents evidence stays valid; unexecuted host paths are never marked passing. Benchmark reports separate observed, measured, inferred and unavailable values.
|
|
@@ -61,3 +61,21 @@ Installers preflight conflicts, atomically replace individual files and preserve
|
|
|
61
61
|
## Deliberate limits
|
|
62
62
|
|
|
63
63
|
No automatic native transcript replacement, full tool interception, arbitrary code rewriting, multi-machine run transfer, lease expiration, automatic compensation or exactly-once external effect guarantee. A sync handler can continue after an async timeout; callbacks are not sandboxed. These limits are part of the contract, not features implied by the paper.
|
|
64
|
+
|
|
65
|
+
## Host registry and optional task lifecycle (0.2)
|
|
66
|
+
|
|
67
|
+
`host_registry` implements a small HostAdapter protocol over the existing hosts
|
|
68
|
+
and desktop installers. It provides detect/install/uninstall/doctor/manifest
|
|
69
|
+
operations. No host name enters the SQLite schema or managed runtime. Core state,
|
|
70
|
+
operations, evidence and bounded model inputs retain their existing boundaries.
|
|
71
|
+
|
|
72
|
+
`lifecycle` defines an optional task profile and evidence/freshness validation.
|
|
73
|
+
`ProjectService` exposes start_task, checkpoint_task, complete_task and find_runs
|
|
74
|
+
in addition to the existing semantic run operations. CLI and the four additive
|
|
75
|
+
MCP tools call this same service. SQLite remains schema v1; old imports and all
|
|
76
|
+
twelve original MCP tools are preserved.
|
|
77
|
+
|
|
78
|
+
Milestone records live in the run's validated state, not a second database.
|
|
79
|
+
Artifact bodies and the event audit remain separate from working context. The
|
|
80
|
+
profile is not forced onto domain-specific schemas. See [state semantics](execution-state.md)
|
|
81
|
+
and [host installation](integrations.md) for boundaries and downgrade guidance.
|