@akinet/akidevrule 3.3.0 → 3.3.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +5 -1
- package/package.json +1 -1
- package/payload/RULE-agent-behavior.md +1 -1
- package/payload/RULE-docs.md +2 -2
package/CHANGELOG.md
CHANGED
|
@@ -1,6 +1,10 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
-
## [
|
|
3
|
+
## [3.3.1] - 2026-09-22
|
|
4
|
+
|
|
5
|
+
### Changed
|
|
6
|
+
- **`docs.B2` leads with the in-place-edit path instead of the freeze command.** Evidence: models repeatedly read the opening line — *"Its body is frozen: never rewrite a claim, a number, or a verification status in place"* — and stopped there, then either refused to correct a research doc at all or spawned a whole new doc for a fix the Decision field survived, inflating the tree with near-duplicate files. Root cause: the rule stated the prohibition first and the flexibility (Cosmetic / Erratum-via-Amendments) only after, so a model that pattern-matches the first strong sentence never reaches the branch that authorizes an in-place amendment. The mechanism the rule already contained was correct; its presentation order taught the opposite of its intent. Mechanism: the opening now states that frozen is the default not a wall, that most corrections land in place, and that one mechanical read of the **Decision** field is the discriminator (Decision would change → successor doc; otherwise edit in place) — and it names the refuse-to-touch / needless-new-doc behavior as the exact failure being prevented, before narrowing "frozen" to only a claim/number/status the Decision still rests on. The four correction classes and the successor-chain naming are unchanged. Rejected: rewriting B2 wholesale (wasteful and self-violating — the fix is a reorder plus one clause, not a new rule). Tradeoff: the opening sentence is longer; that is the load-bearing detail that was previously buried.
|
|
7
|
+
- **`agent.A3` adds a fourth kill-test — Self-sufficiency — before any question reaches the user.** Evidence: the owner repeatedly had to paste the same reminder ("your original request already covers this; `/akirule` has every principle; `/akithink` 4–6 rounds is enough deep-think; the verbatim prompt and history are enough to continue") to unstick a model that had halted to ask something it already had the means to answer. Root cause: A3's three existing kill-tests (Impact, Already-authorized, Silence-is-not-contradiction) plus reversibility screen out *unnecessary* questions, but none catch the *amnesia* question — one whose answer is already sitting in the rule corpus, the deep-think budget, or the original request and conversation history, sources the model is expected to have exhausted before interrupting. Mechanism: a fourth test requires confirming the answer is not already in those three sources before the question leaves, names re-reading them as cheaper than an interrupt and not optional, and is explicitly fenced so it only redirects a self-answerable question — it never overrides the escalation floor (a real one-way door with outward effect is still asked). Rejected: encoding the owner's "4–6 rounds" figure into the rule (that count is an owner instruction at command time, and a fixed round count would duplicate and contradict `think.A2`/`METHOD-deep-think.md`, which scale depth to difficulty); promoting this to a mechanically-detectable penalty card (it is a pre-question judgment, not a locally-detectable output defect — reopen to a card if prose proves insufficient). Tradeoff: the kill-test list is longer and the reversibility note renumbers from fourth to fifth.
|
|
4
8
|
|
|
5
9
|
## [3.3.0] - 2026-09-19
|
|
6
10
|
|
package/package.json
CHANGED
|
@@ -36,7 +36,7 @@ Classify every turn before acting: is it **communication** (a question, discussi
|
|
|
36
36
|
- **Communication → answer, do not act.** Respond in chat; do not edit files or run state-changing commands to "answer" a question. "Can we X?" / "Should we X?" is a question, not permission to do X. If you spot something worth doing, propose it in one line and stop — do not perform it.
|
|
37
37
|
- **Task → execute, do not stall.** Do the requested work within scope; do not turn a clear instruction back into a proposal or a needless confirmation prompt. Report when done, then stop.
|
|
38
38
|
- **Calibrate autonomy by reversibility, not by asking-always.** A reversible, in-scope action gets done and reported; only a genuine one-way door (destructive, outward-facing, scope-expanding, shared config — see B3) is worth pausing to ask. Over-asking on safe work is as much a failure as acting unasked — it trades the user's speed for no real safety.
|
|
39
|
-
- **
|
|
39
|
+
- **Four kill-tests before any question reaches the user — failing one means answer it yourself and record the answer.** Reversibility (above) is the fifth. **Impact:** if the user answers against your default, does any artifact change? "The conclusion holds either way" is a default to write down, never a question to ask. **Already authorized:** the request may have settled it — asking the user to re-confirm a course they just ordered charges them twice for one decision. **Silence is not contradiction:** a doc that does not mention X does not conflict with X; that is a one-line gap to close, i.e. a work item, not a question. **Self-sufficiency:** before the question leaves, confirm the answer is not already sitting in the three sources you are expected to have exhausted — the rule corpus (`akirule` routes it; the core files are already in context), the deep-think budget (`METHOD-deep-think.md`, run to convergence, not one shallow pass), and the owner's original request read verbatim plus the conversation history. A question those three already answer is amnesia, not a genuine unknown; re-reading them is cheaper than an interrupt and is not optional. This test only redirects a question you could answer yourself — it never overrides the escalation floor below: a real one-way door with outward effect is still asked, no matter how self-sufficient the reasoning feels. A question dressed as a "decision with a recommendation" still costs a read and an answer — the shape does not exempt it from these tests.
|
|
40
40
|
- **Deep-think trigger (mandatory, self-driven, non-interactive).** When any holds, Read `METHOD-deep-think.md` and run it before acting or asking: (a) about to ask or escalate to the owner; (b) about to take a one-way-door action; (c) the same fix failed a second time, or a third patch lands on one transition (`pattern.B2`); (d) two rules or instructions conflict; (e) the owner's wording admits readings that produce different artifacts; (f) the change touches documented design or goals. Depth scales with difficulty — repeat goal chain → first principles → critique → pre-mortem until the answer converges, never a fixed round count. Trivial reversible work triggers nothing.
|
|
41
41
|
- **Outcome — converged: act, report the decision.** Self-answer what the analysis settles and act; for important or hard calls report one block — `Decided: X · because Y · rejected Z (why) · reopen if W` — so the owner can overrule after the fact instead of being asked before.
|
|
42
42
|
- **Outcome — escalate only when:** a one-way door with outward effect (publish, tag, destructive data change); contradiction with documented design (`B3`); the `coding.C4` security/money/auth floor; the owner's own wording is ambiguous AND the readings lead to different irreversible artifacts; or deep-think does not converge (name exactly where it is stuck). What survives is asked in a presentation the user can absorb at a glance: everyday wording, jargon glossed, each option carrying its concrete consequence, plus the analysis and one recommendation. A question the user cannot understand costs two interrupts: one to ask, one to explain the asking.
|
package/payload/RULE-docs.md
CHANGED
|
@@ -62,13 +62,13 @@ The stamp is what makes drift mechanically visible: a `docs/arch/` file stamped
|
|
|
62
62
|
|
|
63
63
|
### B2. Research doc structure (`docs/research/`)
|
|
64
64
|
|
|
65
|
-
A research doc is an **event record** — the reasoning as it stood when written (the current-state vs. history split is defined in A2).
|
|
65
|
+
A research doc is an **event record** — the reasoning as it stood when written (the current-state vs. history split is defined in A2). Frozen is the default, not a wall: most corrections land **in place** here, and the discriminator is one mechanical read of the **Decision** field — if applying the fix would change Decision, write a successor doc; if not, edit in place (amendment or cosmetic). Refusing to touch a research doc, or spawning a new doc for a fix the Decision survives, is the failure this rule exists to prevent — it inflates the tree and buries the correction. What is actually frozen is narrow: a **claim, number, or verification status that the Decision still rests on** — never rewrite that in place, because the record of what was believed is the doc's whole value; append an amendment instead. Correction classes:
|
|
66
66
|
- **Cosmetic** (typo, broken link, a path after a rename) — edit in place, no marker.
|
|
67
67
|
- **Erratum on a claim** — a fact turned out wrong, a number was re-measured, an unverified claim was later verified or contradicted, but the **Decision** field still stands: append a dated entry to a closing `## Amendments` section (`- 2026-08-02 · § R9: verified on kiro-cli 2.16.0; the row above was written unverified`), naming the section it corrects and stating only the corrected fact (no story of finding it — `agent.C2`), and add `Status: amended <date>` under the H1 so a reader is warned before reaching the stale claim. The original text stays.
|
|
68
68
|
- **Decision changes** — applying the correction would alter the Decision field: create a **new** research doc and add `Status: superseded by <path>` at the top of the old one. Name the chain with a sequential numeric suffix, ADR-style: `db-engine-choice.md` → `db-engine-choice-2.md` → `db-engine-choice-3.md`, each `superseded by` pointing only at its immediate successor so the chain can be walked backward.
|
|
69
69
|
- **Decision-field links and cross-references** — an Action link to where the result landed, a new cross-ref: edit in place; those fields describe where the event's consequences live, not the event.
|
|
70
70
|
|
|
71
|
-
|
|
71
|
+
Anything outside research that links to a chain (`arch/feat/biz`) points at the latest number and gets updated each time the chain grows — that edit is allowed because those docs hold current state, not history.
|
|
72
72
|
|
|
73
73
|
Required fields, in order:
|
|
74
74
|
1. **Start time** — when the research began
|