veract 0.3.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (35) hide show
  1. veract-0.3.0/LICENSE +9 -0
  2. veract-0.3.0/PKG-INFO +384 -0
  3. veract-0.3.0/README.md +361 -0
  4. veract-0.3.0/agent_runtime/__init__.py +3 -0
  5. veract-0.3.0/agent_runtime/capabilities.py +221 -0
  6. veract-0.3.0/agent_runtime/checkpoint.py +34 -0
  7. veract-0.3.0/agent_runtime/cli.py +132 -0
  8. veract-0.3.0/agent_runtime/compiler.py +82 -0
  9. veract-0.3.0/agent_runtime/config.py +105 -0
  10. veract-0.3.0/agent_runtime/contract.py +153 -0
  11. veract-0.3.0/agent_runtime/engines.py +614 -0
  12. veract-0.3.0/agent_runtime/executor.py +77 -0
  13. veract-0.3.0/agent_runtime/guard.py +95 -0
  14. veract-0.3.0/agent_runtime/llm.py +126 -0
  15. veract-0.3.0/agent_runtime/mcp.py +147 -0
  16. veract-0.3.0/agent_runtime/memory.py +78 -0
  17. veract-0.3.0/agent_runtime/models.py +110 -0
  18. veract-0.3.0/agent_runtime/planner.py +111 -0
  19. veract-0.3.0/agent_runtime/recovery.py +39 -0
  20. veract-0.3.0/agent_runtime/runtime.py +339 -0
  21. veract-0.3.0/agent_runtime/scaffold.py +45 -0
  22. veract-0.3.0/agent_runtime/tools.py +459 -0
  23. veract-0.3.0/agent_runtime/verifier.py +247 -0
  24. veract-0.3.0/pyproject.toml +52 -0
  25. veract-0.3.0/setup.cfg +4 -0
  26. veract-0.3.0/tests/test_engines.py +191 -0
  27. veract-0.3.0/tests/test_llm_paths.py +144 -0
  28. veract-0.3.0/tests/test_runtime.py +158 -0
  29. veract-0.3.0/tests/test_v02.py +276 -0
  30. veract-0.3.0/tests/test_v03.py +263 -0
  31. veract-0.3.0/veract.egg-info/PKG-INFO +384 -0
  32. veract-0.3.0/veract.egg-info/SOURCES.txt +33 -0
  33. veract-0.3.0/veract.egg-info/dependency_links.txt +1 -0
  34. veract-0.3.0/veract.egg-info/entry_points.txt +3 -0
  35. veract-0.3.0/veract.egg-info/top_level.txt +1 -0
veract-0.3.0/LICENSE ADDED
@@ -0,0 +1,9 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 G Sunil Kumar and agent-runtime contributors
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files (the "Software"), to deal in the Software without restriction, including without limitation the rights to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software, and to permit persons to whom the Software is furnished to do so, subject to the following conditions:
6
+
7
+ The above copyright notice and this permission notice shall be included in all copies or substantial portions of the Software.
8
+
9
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND.
veract-0.3.0/PKG-INFO ADDED
@@ -0,0 +1,384 @@
1
+ Metadata-Version: 2.4
2
+ Name: veract
3
+ Version: 0.3.0
4
+ Summary: Local-first verifiable agent runtime with contract floors and zero dependencies.
5
+ Author: G Sunil Kumar, Divakar
6
+ License: MIT
7
+ Project-URL: Homepage, https://github.com/Diwakarsrd/Veract
8
+ Project-URL: Issues, https://github.com/Diwakarsrd/Veract/issues
9
+ Keywords: veract,agent,llm,local-first,verification,ollama,mcp
10
+ Classifier: Development Status :: 4 - Beta
11
+ Classifier: License :: OSI Approved :: MIT License
12
+ Classifier: Operating System :: OS Independent
13
+ Classifier: Programming Language :: Python :: 3 :: Only
14
+ Classifier: Programming Language :: Python :: 3.10
15
+ Classifier: Programming Language :: Python :: 3.11
16
+ Classifier: Programming Language :: Python :: 3.12
17
+ Classifier: Programming Language :: Python :: 3.13
18
+ Classifier: Topic :: Software Development :: Libraries
19
+ Requires-Python: >=3.10
20
+ Description-Content-Type: text/markdown
21
+ License-File: LICENSE
22
+ Dynamic: license-file
23
+
24
+ # Veract
25
+
26
+ ### Give your agent a mission — not a prompt.
27
+
28
+ An open-source, local-first runtime for autonomous AI agents that **plan, execute, verify, recover, and deliver**.
29
+
30
+ > *Agents shouldn't decide whether they succeeded.*
31
+ > *The runtime should prove it.*
32
+
33
+ [Quick Start](#quick-start)  ·  [Why Veract?](#why-veract)  ·  [Architecture](#architecture)  ·  [Benchmarks](#benchmarks)  ·  [Security](#security-by-default)  ·  [Documentation](#cli-reference)
34
+
35
+ [![PyPI version](https://img.shields.io/badge/pypi-v0.3.0-blue.svg)](https://pypi.org/project/veract/)
36
+ [![Python](https://img.shields.io/badge/python-3.10%20%7C%203.11%20%7C%203.12%20%7C%203.13-blue.svg)](https://www.python.org/)
37
+ [![CI](https://github.com/Diwakarsrd/Veract/actions/workflows/ci.yml/badge.svg)](https://github.com/Diwakarsrd/Veract/actions)
38
+ [![License](https://img.shields.io/badge/license-MIT-green.svg)](LICENSE)
39
+ [![Tests](https://img.shields.io/badge/tests-65%20passed-success.svg)](tests/)
40
+
41
+ ---
42
+
43
+ ## 30-Second Demo
44
+
45
+ ```bash
46
+ pip install veract
47
+ export AGENT_LLM_BASE_URL=http://localhost:11434/v1 # Ollama, vLLM, LM Studio, or OpenAI
48
+ export AGENT_LLM_MODEL=llama3.2
49
+
50
+ veract run "Research the 10 best open-source vector databases and save a sourced comparison to dbs.md"
51
+ ```
52
+
53
+ ```text
54
+ ✓ Mission compiled: "10 best open-source vector databases"
55
+ ✓ Success contract created: [min_items: 10, unique: true, sourced: true, live_urls: true]
56
+ ✓ Execution DAG generated: 4 parallel waves
57
+ ✓ 10/10 vector databases extracted & grounded to source text
58
+ ✓ Independent verification: 10/10 URLs resolve, 0 hallucinations
59
+ ✓ Artifact sealed: dbs.md (SHA-256 verified)
60
+ ✓ Evidence checkpoint saved: .agent/missions/20261006-033120-0db756
61
+
62
+ MISSION PASSED (0 replans, 13.3s wall time)
63
+ ```
64
+
65
+ Veract doesn't ask an LLM if it finished. It requires independent, reproducible proof against an explicit contract floor.
66
+
67
+ ---
68
+
69
+ ## Why Veract?
70
+
71
+ Most agent frameworks optimize for **getting an answer**.
72
+ Veract optimizes for **proving the answer is correct**.
73
+
74
+ | Traditional Agent | Veract |
75
+ |---|---|
76
+ | **Prompt → response** | **Mission → explicit success contract** |
77
+ | Model decides whether it succeeded | Independent verifier validates proof from disk and live APIs |
78
+ | Blind tool execution | Capability-brokered tools with audit logs and approval gates |
79
+ | Infinite unguided retry loops | Failure-classified recovery (*fix → retry → replan*) |
80
+ | Lossy conversational history | SQLite typed memory with provenance, confidence decay, and contradictions |
81
+ | *"Looks complete to me"* | Cryptographically hashed evidence checkpoint |
82
+ | Silent hallucinated success | Honest terminal states: `passed`, `failed`, `refused`, `partial` |
83
+
84
+ ---
85
+
86
+ ## What Veract Can Do
87
+
88
+ - **🔎 Research** — Search candidate pages, fetch sources, extract grounded entities, verify live registry URLs, and format citations.
89
+ - **💻 Coding** — Run tests, isolate failures, patch files with exact-once verification, re-test, and rollback if tests degrade.
90
+ - **📊 Data Extraction** — Ingest structured files (CSV, JSON), calculate aggregations, validate anomalies, and verify against truth data.
91
+ - **🔐 Security Guardrails** — Prevent prompt injection, block out-of-workspace file traversal, deny private IP SSRF, and redact credentials.
92
+ - **♻️ Self-Healing Recovery** — Classify failures into transient, permission, syntax, or logic errors and apply deterministic fixes before replanning.
93
+ - **🧠 Long-Term Memory** — Store verified facts with provenance tags, confidence scores, and automatic contradiction invalidation.
94
+
95
+ ---
96
+
97
+ ## Architecture
98
+
99
+ Veract separates execution from evaluation. The agent never grades its own work:
100
+
101
+ ```text
102
+ ┌─────────────┐
103
+ │ MISSION │
104
+ └──────┬──────┘
105
+ ↓
106
+ ┌────────────────────┐
107
+ │ INTENT COMPILER │
108
+ └─────────┬──────────┘
109
+ ↓
110
+ ┌────────────────────┐
111
+ │ SUCCESS CONTRACT │
112
+ └─────────┬──────────┘
113
+ ↓
114
+ ┌────────────────────┐
115
+ │ PLAN / DAG │
116
+ └─────────┬──────────┘
117
+ ↓
118
+ ┌─────────────────────────────┐
119
+ │ BROKERED EXECUTOR │ ⇄ [ Capability Broker ]
120
+ └──────────────┬──────────────┘
121
+ ↓
122
+ ┌────────────┐
123
+ │ EVIDENCE │
124
+ └─────┬──────┘
125
+ ↓
126
+ ┌─────────────────┐
127
+ │ VERIFIER │ ⇄ [ Independent Checkers / Live APIs ]
128
+ └───────┬─────────┘
129
+ │
130
+ ┌──────┴──────┐
131
+ │ │
132
+ PASS FAIL
133
+ │ │
134
+ ↓ ↓
135
+ DELIVER RECOVERY
136
+ │
137
+ ↓
138
+ REPLAN
139
+ ```
140
+
141
+ ### The Execution Model
142
+
143
+ ```text
144
+ PROMPT-DRIVEN AGENTS VERACT RUNTIME
145
+
146
+ Prompt Mission
147
+ ↓ ↓
148
+ Model Success Contract
149
+ ↓ ↓
150
+ Tool Call Execution DAG
151
+ ↓ ↓
152
+ Answer Sandboxed Execution
153
+ ↓ ↓
154
+ (Self-Judged: "Looks good") Evidence Collection
155
+ ↓
156
+ Independent Verification
157
+ ↓
158
+ ┌──────┴──────┐
159
+ PASS FAIL
160
+ ↓ ↓
161
+ Deliver Recover → Replan
162
+ ```
163
+
164
+ ---
165
+
166
+ ## Security by Default
167
+
168
+ Untrusted models and third-party tools cannot be given raw system access. Veract implements least privilege across the entire lifecycle:
169
+
170
+ ```text
171
+ ┌──────────────────────────────────────────────────────────────┐
172
+ │ VERACT RUNTIME │
173
+ │ │
174
+ │ Mission Input │
175
+ │ ↓ │
176
+ │ Scope Guard (Refuses out-of-workspace paths) │
177
+ │ ↓ │
178
+ │ Capability Broker (Enforces filesystem & network ACLs) │
179
+ │ ↓ │
180
+ │ Approval Engine (Interactive prompts for new code) │
181
+ │ ↓ │
182
+ │ Container Sandbox (Read-only root, memory/CPU caps) │
183
+ │ ↓ │
184
+ │ Audit Trail (Append-only JSONL event log) │
185
+ └──────────────────────────────────────────────────────────────┘
186
+ ```
187
+
188
+ 1. **Scope Guard**: Any mission referencing paths outside the workspace (e.g. `C:\Windows\win.ini` or `/etc/passwd`) is refused immediately before calling the LLM or touching tools.
189
+ 2. **Capability Broker**: Filesystem reads and writes are restricted to workspace roots. `.agent/` and `.git/` are immutable to the agent. Outbound network traffic is limited to HTTP/HTTPS, blocking private subnets (`127.0.0.1`, `10.0.0.0/8`, `169.254.0.0/16`).
190
+ 3. **Execution Rules**: The agent cannot execute arbitrary shell scripts or code it just wrote without human approval. Banned flags (`python -c`, `pytest -p evil_plugin`) are blocked at the argv parser.
191
+ 4. **Secret Scrubbing**: API keys (`sk-...`, `AKIA...`, `ghp_...`) and email addresses are automatically stripped from web search queries and redacted from disk deliverables.
192
+ 5. **Container Sandboxing**: For untrusted code, Docker mode enforces `--network none`, `--read-only` root, and hard memory/CPU limits.
193
+
194
+ ---
195
+
196
+ ## Benchmarks
197
+
198
+ Veract has been evaluated across 240 controlled offline trials and live head-to-head runs on small local models (`llama3.2-3B`, 16k context):
199
+
200
+ ### 1. Controlled Held-Out Evaluations (Offline)
201
+ Tested against a simulated weak model reproducing small-model failure modes (dropped args, hallucinated citations, prompt injection, and weak contracts):
202
+
203
+ | Metric | Set 1 (Unhardened Baseline) | Set 1 (Veract Hardened) | Set 2 (Frozen Held-Out) | 5 Repeated Seeds (150 trials) |
204
+ |---|:---:|:---:|:---:|:---:|
205
+ | **Solved Tasks** | 40 / 60 | **57 / 60** | **30 / 30** | **147 / 150 (98.0%)** |
206
+ | **False Passes** | 11 | **0** | **0** | **0 (0.0%)** |
207
+ | **Canary / Secret Leaks** | 0 | **0** | **0** | **0 (0.0%)** |
208
+ | **Path Leaks in Search** | 6 | **0** | **0** | **0 (0.0%)** |
209
+ | **Runtime Crashes** | 0 | **0** | **0** | **0 (0.0%)** |
210
+
211
+ > *Results are empirical measurements from our reproducible test harness (`bench/heldout.py`). Full methodology: [`bench/HELDOUT.md`](bench/HELDOUT.md). Raw data: [`bench/heldout_runs/all-5seeds.json`](bench/heldout_runs/all-5seeds.json).*
212
+
213
+ ### 2. Live Head-to-Head Comparison (Ollama `llama3.2-16k`)
214
+ Evaluated across 5 complex real-world tasks (PyPI registry verification, math debugging, data aggregation, multi-source research) against leading agent runtimes under a 180s–400s deadline:
215
+
216
+ | Agent Runtime | Tasks Passed | False Passes | Average Time | Failure Mode |
217
+ |---|:---:|:---:|:---:|---|
218
+ | **Veract** | **4 / 5** | **0** | **~13.3s** | Clean, verified deliverables; 1 honest timeout on dead upstream API |
219
+ | **Hermes Agent** | 0 / 5 | 0 | 180s+ (Timeout) | Stalled on 16k system prompt context; dropped tool arguments |
220
+ | **OpenClaw** | 0 / 5 | 0 | 300s+ (Timeout) | Context overflow on large prompts; crashed on Windows file locks |
221
+
222
+ *Raw execution logs: [`bench/logs/`](bench/logs/)  ·  Detailed report: [`bench/REPORT.md`](bench/REPORT.md)*
223
+
224
+ ---
225
+
226
+ ## Engineering Principles
227
+
228
+ Veract's architecture was shaped by analyzing empirical failure logs in existing agents:
229
+
230
+ - **Independent Contract Floor**: The agent cannot negotiate away its success criteria. A research task must prove entity count, uniqueness, source citation, and live URL resolution. The presence of an empty file will never satisfy the contract.
231
+ - **Query Sanitization**: Agents frequently leak private local filesystem paths into public search engines when trying to understand a user request. Veract strips paths, emails, and credentials prior to sending web search requests.
232
+ - **Exact-Once Patching**: Code edits require verified single-instance matches with immediate read-back verification. Ambiguous edits or partial deletes are aborted and sent back to the recovery engine.
233
+ - **Deadline-Aware Checkpoints**: If an agent runs out of time, Veract saves partial verified progress, stores an inspectable audit trace, and returns an honest `partial` state rather than hanging indefinitely.
234
+
235
+ ---
236
+
237
+ ## Quick Start
238
+
239
+ ### Installation
240
+
241
+ ```bash
242
+ pip install veract
243
+ ```
244
+
245
+ ### Environment Configuration
246
+
247
+ Veract works with any OpenAI-compatible server:
248
+
249
+ ```bash
250
+ # Local Ollama (Recommended)
251
+ export AGENT_LLM_BASE_URL=http://localhost:11434/v1
252
+ export AGENT_LLM_MODEL=llama3.2
253
+ export AGENT_CTX_TOKENS=16384
254
+
255
+ # Or vLLM / LM Studio / OpenAI
256
+ export AGENT_LLM_BASE_URL=https://api.openai.com/v1
257
+ export AGENT_LLM_API_KEY=sk-...
258
+ export AGENT_LLM_MODEL=gpt-4o-mini
259
+ ```
260
+
261
+ ### Basic Commands
262
+
263
+ ```bash
264
+ # Execute a mission
265
+ veract run "Fix failing tests in tests/test_calc.py without modifying the test file"
266
+
267
+ # Check mission status across workspace
268
+ veract status
269
+
270
+ # Inspect full capability audit log
271
+ veract audit <mission-id>
272
+
273
+ # Resume an interrupted mission from last checkpoint
274
+ veract resume <mission-id>
275
+
276
+ # Query long-term memory
277
+ veract memory "vector databases"
278
+
279
+ # Inspect mission evidence and contract
280
+ veract show <mission-id>
281
+ ```
282
+
283
+ ---
284
+
285
+ ## Configuration (`.agent/policy.json`)
286
+
287
+ Configure permissions, sandboxing, and MCP servers per workspace:
288
+
289
+ ```json
290
+ {
291
+ "net_domains": ["*"],
292
+ "exec_new_code": "approve",
293
+ "approve": [
294
+ {"action": "execute", "pattern": "python -m pytest*"}
295
+ ],
296
+ "search": {
297
+ "provider": "searxng",
298
+ "url": "http://127.0.0.1:8888"
299
+ },
300
+ "sandbox": {
301
+ "mode": "docker",
302
+ "image": "agent-runtime-sandbox",
303
+ "memory_mb": 2048
304
+ },
305
+ "mcp": {
306
+ "filesystem": {
307
+ "command": ["npx", "-y", "@modelcontextprotocol/server-filesystem", "."]
308
+ }
309
+ },
310
+ "mcp_allow": ["filesystem.read_*"]
311
+ }
312
+ ```
313
+
314
+ - **Sandbox Modes**:
315
+ - `docker` *(Recommended for untrusted code)*: Network-isolated container with read-only root and memory/CPU limits.
316
+ - `limits` *(POSIX environments)*: Resource limits via `setrlimit` (CPU, memory, file size).
317
+ - `none`: Direct host execution with broker auditing.
318
+ - **Model Context Protocol (MCP)**: Native stdio client. Tools are registered as `mcp.<server>.<tool>` and subject to capability gating.
319
+ - **Interactive Approval**: When an agent attempts an ungranted sensitive action, you receive an interactive `y / N / a` (always) terminal prompt.
320
+
321
+ ---
322
+
323
+ ## Honest Status Values
324
+
325
+ Veract enforces rigorous status distinctions:
326
+
327
+ | Status | Meaning |
328
+ |---|---|
329
+ | `passed` | Every criterion was independently verified with reproducible proof. |
330
+ | `failed` | Mission could not be satisfied within the replan/retry budget. Failing checks are detailed in the artifact. |
331
+ | `refused` | The mission violates policy (e.g., path traversal). Refused before LLM invocation or tool execution. |
332
+ | `partial` | Mission was terminated by deadline; only verified evidence was committed. |
333
+ | `unverified` | Supported only by model self-judgment. **Never reported as passed**. |
334
+
335
+ ---
336
+
337
+ ## Technical Internals
338
+
339
+ For contributors and developers building on Veract:
340
+
341
+ | Subsystem | File | Responsibility |
342
+ |---|---|---|
343
+ | **Scope Guard** | [`guard.py`](agent_runtime/guard.py) | Pre-execution path inspection, query sanitization, and secret redaction. |
344
+ | **Contract Engine** | [`contract.py`](agent_runtime/contract.py) | Success contract floor and deterministic verification rules. |
345
+ | **Compiler** | [`compiler.py`](agent_runtime/compiler.py) | Converts natural language requests into objectives and criteria graphs. |
346
+ | **Planner** | [`planner.py`](agent_runtime/planner.py) | Generates dependency-aware DAG execution plans. |
347
+ | **Executor** | [`executor.py`](agent_runtime/executor.py) | Wave-based concurrent tool execution and deadline management. |
348
+ | **Verifier** | [`verifier.py`](agent_runtime/verifier.py) | Independent claim verification, content re-derivation, and live checks. |
349
+ | **Recovery** | [`recovery.py`](agent_runtime/recovery.py) | Failure classification and self-healing repair strategies. |
350
+ | **Capability Broker** | [`capabilities.py`](agent_runtime/capabilities.py) | Permission enforcement, filesystem isolation, and JSONL audit logging. |
351
+ | **Engines** | [`engines.py`](agent_runtime/engines.py) | Deterministic fast-paths for research, coding, and tabular data. |
352
+ | **Memory** | [`memory.py`](agent_runtime/memory.py) | SQLite-backed episodic and semantic memory with confidence decay. |
353
+ | **MCP Client** | [`mcp.py`](agent_runtime/mcp.py) | Model Context Protocol client with broker gating. |
354
+
355
+ ---
356
+
357
+ ## What Veract Is Not
358
+
359
+ - **Not a hosted SaaS**: Veract is a local-first Python library and CLI. Your data, code, and keys stay on your machine.
360
+ - **Not locked to proprietary models**: Designed specifically to make small, local open-weights models (3B–14B) reliable.
361
+ - **Not an unconstrained auto-coder**: Veract does not rewrite its own codebase or bypass approval boundaries.
362
+ - **Not a replacement for virtualization**: In `limits` mode, resource limits apply, but full system isolation requires `mode: "docker"`.
363
+
364
+ ---
365
+
366
+ ## Contributing
367
+
368
+ We welcome contributions to Veract!
369
+
370
+ ```bash
371
+ git clone https://github.com/Diwakarsrd/Veract.git
372
+ cd Veract
373
+ pip install -e .
374
+ python -m unittest discover -s tests -v
375
+ ruff check .
376
+ ```
377
+
378
+ Please review [`CONTRIBUTING.md`](CONTRIBUTING.md) and [`SECURITY.md`](SECURITY.md) before submitting pull requests.
379
+
380
+ ---
381
+
382
+ ## License
383
+
384
+ Licensed under the [MIT License](LICENSE).