@penguinharness/humanizer 0.2.9

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/LICENSE ADDED
@@ -0,0 +1,201 @@
1
+ Apache License
2
+ Version 2.0, January 2004
3
+ http://www.apache.org/licenses/
4
+
5
+ TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION
6
+
7
+ 1. Definitions.
8
+
9
+ "License" shall mean the terms and conditions for use, reproduction,
10
+ and distribution as defined by Sections 1 through 9 of this document.
11
+
12
+ "Licensor" shall mean the copyright owner or entity authorized by
13
+ the copyright owner that is granting the License.
14
+
15
+ "Legal Entity" shall mean the union of the acting entity and all
16
+ other entities that control, are controlled by, or are under common
17
+ control with that entity. For the purposes of this definition,
18
+ "control" means (i) the power, direct or indirect, to cause the
19
+ direction or management of such entity, whether by contract or
20
+ otherwise, or (ii) ownership of fifty percent (50%) or more of the
21
+ outstanding shares, or (iii) beneficial ownership of such entity.
22
+
23
+ "You" (or "Your") shall mean an individual or Legal Entity
24
+ exercising permissions granted by this License.
25
+
26
+ "Source" form shall mean the preferred form for making modifications,
27
+ including but not limited to software source code, documentation
28
+ source, and configuration files.
29
+
30
+ "Object" form shall mean any form resulting from mechanical
31
+ transformation or translation of a Source form, including but
32
+ not limited to compiled object code, generated documentation,
33
+ and conversions to other media types.
34
+
35
+ "Work" shall mean the work of authorship, whether in Source or
36
+ Object form, made available under the License, as indicated by a
37
+ copyright notice that is included in or attached to the work
38
+ (an example is provided in the Appendix below).
39
+
40
+ "Derivative Works" shall mean any work, whether in Source or Object
41
+ form, that is based on (or derived from) the Work and for which the
42
+ editorial revisions, annotations, elaborations, or other modifications
43
+ represent, as a whole, an original work of authorship. For the purposes
44
+ of this License, Derivative Works shall not include works that remain
45
+ separable from, or merely link (or bind by name) to the interfaces of,
46
+ the Work and Derivative Works thereof.
47
+
48
+ "Contribution" shall mean any work of authorship, including
49
+ the original version of the Work and any modifications or additions
50
+ to that Work or Derivative Works thereof, that is intentionally
51
+ submitted to Licensor for inclusion in the Work by the copyright owner
52
+ or by an individual or Legal Entity authorized to submit on behalf of
53
+ the copyright owner. For the purposes of this definition, "submitted"
54
+ means any form of electronic, verbal, or written communication sent
55
+ to the Licensor or its representatives, including but not limited to
56
+ communication on electronic mailing lists, source code control systems,
57
+ and issue tracking systems that are managed by, or on behalf of, the
58
+ Licensor for the purpose of discussing and improving the Work, but
59
+ excluding communication that is conspicuously marked or otherwise
60
+ designated in writing by the copyright owner as "Not a Contribution."
61
+
62
+ "Contributor" shall mean Licensor and any individual or Legal Entity
63
+ on behalf of whom a Contribution has been received by Licensor and
64
+ subsequently incorporated within the Work.
65
+
66
+ 2. Grant of Copyright License. Subject to the terms and conditions of
67
+ this License, each Contributor hereby grants to You a perpetual,
68
+ worldwide, non-exclusive, no-charge, royalty-free, irrevocable
69
+ copyright license to reproduce, prepare Derivative Works of,
70
+ publicly display, publicly perform, sublicense, and distribute the
71
+ Work and such Derivative Works in Source or Object form.
72
+
73
+ 3. Grant of Patent License. Subject to the terms and conditions of
74
+ this License, each Contributor hereby grants to You a perpetual,
75
+ worldwide, non-exclusive, no-charge, royalty-free, irrevocable
76
+ (except as stated in this section) patent license to make, have made,
77
+ use, offer to sell, sell, import, and otherwise transfer the Work,
78
+ where such license applies only to those patent claims licensable
79
+ by such Contributor that are necessarily infringed by their
80
+ Contribution(s) alone or by combination of their Contribution(s)
81
+ with the Work to which such Contribution(s) was submitted. If You
82
+ institute patent litigation against any entity (including a
83
+ cross-claim or counterclaim in a lawsuit) alleging that the Work
84
+ or a Contribution incorporated within the Work constitutes direct
85
+ or contributory patent infringement, then any patent licenses
86
+ granted to You under this License for that Work shall terminate
87
+ as of the date such litigation is filed.
88
+
89
+ 4. Redistribution. You may reproduce and distribute copies of the
90
+ Work or Derivative Works thereof in any medium, with or without
91
+ modifications, and in Source or Object form, provided that You
92
+ meet the following conditions:
93
+
94
+ (a) You must give any other recipients of the Work or
95
+ Derivative Works a copy of this License; and
96
+
97
+ (b) You must cause any modified files to carry prominent notices
98
+ stating that You changed the files; and
99
+
100
+ (c) You must retain, in the Source form of any Derivative Works
101
+ that You distribute, all copyright, patent, trademark, and
102
+ attribution notices from the Source form of the Work,
103
+ excluding those notices that do not pertain to any part of
104
+ the Derivative Works; and
105
+
106
+ (d) If the Work includes a "NOTICE" text file as part of its
107
+ distribution, then any Derivative Works that You distribute must
108
+ include a readable copy of the attribution notices contained
109
+ within such NOTICE file, excluding those notices that do not
110
+ pertain to any part of the Derivative Works, in at least one
111
+ of the following places: within a NOTICE text file distributed
112
+ as part of the Derivative Works; within the Source form or
113
+ documentation, if provided along with the Derivative Works; or,
114
+ within a display generated by the Derivative Works, if and
115
+ wherever such third-party notices normally appear. The contents
116
+ of the NOTICE file are for informational purposes only and
117
+ do not modify the License. You may add Your own attribution
118
+ notices within Derivative Works that You distribute, alongside
119
+ or as an addendum to the NOTICE text from the Work, provided
120
+ that such additional attribution notices cannot be construed
121
+ as modifying the License.
122
+
123
+ You may add Your own copyright statement to Your modifications and
124
+ may provide additional or different license terms and conditions
125
+ for use, reproduction, or distribution of Your modifications, or
126
+ for any such Derivative Works as a whole, provided Your use,
127
+ reproduction, and distribution of the Work otherwise complies with
128
+ the conditions stated in this License.
129
+
130
+ 5. Submission of Contributions. Unless You explicitly state otherwise,
131
+ any Contribution intentionally submitted for inclusion in the Work
132
+ by You to the Licensor shall be under the terms and conditions of
133
+ this License, without any additional terms or conditions.
134
+ Notwithstanding the above, nothing herein shall supersede or modify
135
+ the terms of any separate license agreement you may have executed
136
+ with Licensor regarding such Contributions.
137
+
138
+ 6. Trademarks. This License does not grant permission to use the trade
139
+ names, trademarks, service marks, or product names of the Licensor,
140
+ except as required for reasonable and customary use in describing the
141
+ origin of the Work and reproducing the content of the NOTICE file.
142
+
143
+ 7. Disclaimer of Warranty. Unless required by applicable law or
144
+ agreed to in writing, Licensor provides the Work (and each
145
+ Contributor provides its Contributions) on an "AS IS" BASIS,
146
+ WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or
147
+ implied, including, without limitation, any warranties or conditions
148
+ of TITLE, NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A
149
+ PARTICULAR PURPOSE. You are solely responsible for determining the
150
+ appropriateness of using or redistributing the Work and assume any
151
+ risks associated with Your exercise of permissions under this License.
152
+
153
+ 8. Limitation of Liability. In no event and under no legal theory,
154
+ whether in tort (including negligence), contract, or otherwise,
155
+ unless required by applicable law (such as deliberate and grossly
156
+ negligent acts) or agreed to in writing, shall any Contributor be
157
+ liable to You for damages, including any direct, indirect, special,
158
+ incidental, or consequential damages of any character arising as a
159
+ result of this License or out of the use or inability to use the
160
+ Work (including but not limited to damages for loss of goodwill,
161
+ work stoppage, computer failure or malfunction, or any and all
162
+ other commercial damages or losses), even if such Contributor
163
+ has been advised of the possibility of such damages.
164
+
165
+ 9. Accepting Warranty or Additional Liability. While redistributing
166
+ the Work or Derivative Works thereof, You may choose to offer,
167
+ and charge a fee for, acceptance of support, warranty, indemnity,
168
+ or other liability obligations and/or rights consistent with this
169
+ License. However, in accepting such obligations, You may act only
170
+ on Your own behalf and on Your sole responsibility, not on behalf
171
+ of any other Contributor, and only if You agree to indemnify,
172
+ defend, and hold each Contributor harmless for any liability
173
+ incurred by, or claims asserted against, such Contributor by reason
174
+ of your accepting any such warranty or additional liability.
175
+
176
+ END OF TERMS AND CONDITIONS
177
+
178
+ APPENDIX: How to apply the Apache License to your work.
179
+
180
+ To apply the Apache License to your work, attach the following
181
+ boilerplate notice, with the fields enclosed by brackets "[]"
182
+ replaced with your own identifying information. (Don't include
183
+ the brackets!) The text should be enclosed in the appropriate
184
+ comment syntax for the file format. We also recommend that a
185
+ file or class name and description of purpose be included on the
186
+ same "printed page" as the copyright notice for easier
187
+ identification within third-party archives.
188
+
189
+ Copyright [yyyy] [name of copyright owner]
190
+
191
+ Licensed under the Apache License, Version 2.0 (the "License");
192
+ you may not use this file except in compliance with the License.
193
+ You may obtain a copy of the License at
194
+
195
+ http://www.apache.org/licenses/LICENSE-2.0
196
+
197
+ Unless required by applicable law or agreed to in writing, software
198
+ distributed under the License is distributed on an "AS IS" BASIS,
199
+ WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
200
+ See the License for the specific language governing permissions and
201
+ limitations under the License.
package/icon.svg ADDED
@@ -0,0 +1,7 @@
1
+ <svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="1.7" stroke-linecap="round" stroke-linejoin="round">
2
+ <path d="M4 5h11" />
3
+ <path d="M4 9.5h7" />
4
+ <path d="M4 14h5" />
5
+ <path d="M4 18.5h6" />
6
+ <path d="m13.2 17.9 7.1-7.1a1.7 1.7 0 0 0-2.4-2.4l-7.1 7.1-.9 3.3 3.3-.9Z" />
7
+ </svg>
package/package.json ADDED
@@ -0,0 +1,21 @@
1
+ {
2
+ "name": "@penguinharness/humanizer",
3
+ "version": "0.2.9",
4
+ "description": "Rewrite or edit prose in any language so it reads like edited human writing in the register of books, newspapers and encyclopedias rather than default AI output. A small drafting core — vary every pattern, build density from anchored facts in whole grammar, cap the quotables, put a real writer with real material behind the text, let structure serve content, write each language from inside its idiom and typography, verify what you assert, and aim for the natural distribution of edited prose rather than a perfect scorecard — backed by a three-layer tell catalog, per-language cue files for six languages and a seven-round measured case study shipped as reference files for the diagnostic census.",
5
+ "license": "Apache-2.0",
6
+ "repository": {
7
+ "type": "git",
8
+ "url": "git+https://github.com/Prism-Shadow/penguin-harness.git",
9
+ "directory": "plugins/humanizer"
10
+ },
11
+ "files": [
12
+ "plugin.json",
13
+ "icon.svg",
14
+ "skills",
15
+ "hooks",
16
+ "LICENSE"
17
+ ],
18
+ "publishConfig": {
19
+ "access": "public"
20
+ }
21
+ }
package/plugin.json ADDED
@@ -0,0 +1,8 @@
1
+ {
2
+ "description": "Rewrite or edit prose in any language so it reads like edited human writing in the register of books, newspapers and encyclopedias rather than default AI output. A small drafting core — vary every pattern, build density from anchored facts in whole grammar, cap the quotables, put a real writer with real material behind the text, let structure serve content, write each language from inside its idiom and typography, verify what you assert, and aim for the natural distribution of edited prose rather than a perfect scorecard — backed by a three-layer tell catalog, per-language cue files for six languages and a seven-round measured case study shipped as reference files for the diagnostic census.",
3
+ "short_description": "Strip AI-writing tells from prose in any language.",
4
+ "short_description_zh": "去除文本的 AI 味,把行文改成书籍、报纸、百科式的风格,任何语言通用。",
5
+ "version": "2026-08-12.1",
6
+ "category": "office-productivity",
7
+ "preinstall": false
8
+ }
@@ -0,0 +1,39 @@
1
+ ---
2
+ name: humanizer
3
+ description: Rewrite or edit prose in any language so it reads like edited human writing in the register of books, newspapers and encyclopedias rather than default AI output. A small drafting core — vary every pattern, build density from anchored facts in whole grammar, cap the quotables, put a real writer with real material behind the text, let structure serve content, write each language from inside its idiom and typography, verify what you assert, and aim for the natural distribution of edited prose rather than a perfect scorecard — backed by a three-layer tell catalog, per-language cue files for six languages and a seven-round measured case study shipped as reference files for the diagnostic census.
4
+ ---
5
+
6
+ # Humanizer
7
+
8
+ Make prose read like edited human writing: the register of books, quality newspapers and encyclopedia entries. Everything here was derived empirically, across seven rounds of drafting, blind editorial review and revision, documented with counts in [`reference/case-study.md`](reference/case-study.md). The working surface is deliberately small: a writer boxed in by a long checklist produces compliance, not prose — that failure mode is the case study's best-documented finding. Draft with the principles below; diagnose with the catalog afterwards.
9
+
10
+ ## Before you start
11
+
12
+ If the invocation carries no text and no assignment, ask for one: the draft to edit, or the topic, length, audience and language to write fresh. Settle two things early. The register: books, newspapers and encyclopedias are the default target, while marketing, speeches and reference documentation legitimately bend these rules, so confirm how far to go. And any hard length target: humanizing shrinks text, and gaps are filled with substance, never padding. If the genre needs material nobody has gathered — reportage needs a scene, a person, a quotation — say so and get it, or agree to relabel the piece; never fake the texture.
13
+
14
+ ## Core principles
15
+
16
+ 1. **No pattern twice.** Whatever the figure — a contrast frame, a triad, a cleft, an opener shape, a paragraph arc, a metaphor, even a favorite connective particle — its second consecutive use is a rhythm and its third is a stencil. Vary sentence length, clause weight, paragraph attack and closer; if every paragraph advances by the same move, swap engines somewhere.
17
+ 2. **Density is anchored facts in whole grammar.** Names, dates, numbers, mechanisms and worked examples carry the argument; hype, era openers, phantom crowds ("faster than most expected" — who?) and concepts pushing concepts carry nothing. Compress by dropping padding, not grammar: subjects stay, first mentions get their full noun, and the event that matters gets a sentence of its own.
18
+ 3. **Cap the quotables.** One or two turned phrases can carry a piece, counting every shape: chiasmus, mirrored re-description, balanced antithesis, aphoristic kickers. State the thesis once and develop it with new material or cut the echo. Most paragraphs end flat, and an argument survives a dud.
19
+ 4. **A real writer, with the material.** Opinion owns its anecdotes and risks a judgment of its own; reportage has stood somewhere; reference registers keep the writer invisible without chaperoning the reader. Claims sized to the evidence, honest hedges kept, references anchored (who, where, when), and one non-obvious source beats a second canonical one.
20
+ 5. **Structure serves content.** Open on ground, not on a cold verdict and not on an era; paragraphs develop one idea across several sentences; headings and bullets appear only where content is genuinely enumerable; the piece ends where the information ends.
21
+ 6. **Write from inside the language.** Native idiom — a sentence that back-translates cleanly into another language was composed there, so recompose it — native punctuation at native frequency (a dash where the register expects one beats zero), and the venue's typography held consistently. Per-language budgets and surface forms: [`reference/language-cues.md`](reference/language-cues.md), which indexes one `reference/<lang>-cues.md` file per language.
22
+ 7. **True, verified, calibrated.** Never invent specifics. Re-derive every mechanism example from the stated mechanism, split fused attributions, and verify or delete every superlative. One fluent falsehood outweighs any amount of style.
23
+
24
+ Over all seven: **aim for the natural distribution of edited prose, not a perfect scorecard.** Every rule above, overdriven, mints a new tell — the case study documents four such artifacts (banned openers became cold-open verdicts; scrubbed punctuation became semicolon inflation and rationed warmth; density became telegraph compression; scrubbed abstraction became contrived colloquialism). Keep some looseness, an aside, an unresolved edge.
25
+
26
+ ## Method
27
+
28
+ 1. **Read whole, list the keeps.** Facts, quotations, required terminology, genre constraints, any hard length target. Everything on the list survives the rewrite.
29
+ 2. **Draft, or restructure, by the principles.** For an existing draft: merge fragment paragraphs by idea, delete scaffolding, then rewrite sentence by sentence. Do not draft against the catalog — that produces compliance-shaped text.
30
+ 3. **Census.** Now open [`reference/tells.md`](reference/tells.md) and count. The drafting mind cannot see its own tics: in a field test, an author who felt two aphorisms was carrying eight, and five triads survived a no-triads rule. Mechanical counting over impression, catalog over memory; a flagged pattern kept for its quality is still flagged.
31
+ 4. **Revise and gate.** Fix what the census found; verify facts and mechanisms (principle 7); read aloud for rhythm. Then the excerpt test: any paragraph alone should sit unnoticed in a book, a broadsheet or an encyclopedia. Finish with one skeptical-editor read asking where a reader would still mutter "AI wrote this" — and when several pieces come from one session, compare them side by side, because shared architecture across pieces is a fingerprint too. One more stop rule: a spot flagged again after repair is overloaded, not misworded — restructure the thought (split it, reorder it) instead of trying a third wording.
32
+
33
+ ## Worked example (English)
34
+
35
+ Before, tells marked: "In today's rapidly evolving AI landscape [era opener], skills have emerged as a game-changer [hype]. A skill isn't just a document — [dash] it's a reusable playbook [template contrast] that transforms how agents work. Whether you're automating reports, reviewing code, or managing data [triad, reader address], skills unlock consistency, reliability, and scale [triad, hype]. The future of agent workflows starts here [uplift ending]."
36
+
37
+ After: "A skill is a document that tells an AI agent how to perform one kind of task. The agent keeps only the document's one-line summary in memory and reads the full text when a matching task arrives, so hundreds of skills can be installed at negligible cost. The format is plain Markdown with a short metadata header, which means anyone who can write instructions can write a skill."
38
+
39
+ The full tell catalog is in [`reference/tells.md`](reference/tells.md), per-language surface forms in the `reference/<lang>-cues.md` files indexed by [`reference/language-cues.md`](reference/language-cues.md), and the measured record behind all of it in [`reference/case-study.md`](reference/case-study.md).
@@ -0,0 +1,244 @@
1
+ # Case study: measuring the AI out of prose
2
+
3
+ How the tells in SKILL.md were derived, in seven rounds. Round one is a controlled exercise: one assignment answered twice, first as a deliberately typical AI draft, the register a default assistant produces unprompted, then as a rewrite into print register; both versions were kept, the tells were counted with pattern matching rather than judgment, and the diff produced the sentence-layer tells (1–10). Round two is a set of field tests, added after human readers flagged text that already followed round one; it produced the discourse-layer tells (11–17). Round three scaled the method: ten fresh pieces across five languages and ten genres, written honestly to the first two layers, were reviewed blind by independent editors; what they still caught produced the scrub-layer tells (18–24) and the per-language lists in `language-cues.md`. Round four is a user-run field test on fresh material that caught the density rule overshooting (tell 25), and round five a further user pass on regenerated material that caught the repair itself compressing (clipped conditionals, particle-chain monotony). Round six is a user editing session on one release-note passage whose four successive rejections (让-chains, announcement scaffolding, the before/after frame, contrived colloquialism) converged on plain statement. Round seven is a user pass on the user's own hand-written draft, where both flags landed on transitions (a re-struck temporal anchor, a bare counterpoint opener). Rounds one and two are in Chinese; round three covers Chinese, English, Japanese, French, German and Spanish.
4
+
5
+ The assignment, as a user would phrase it: 请帮我写一篇图文稿,大约 1000 字,讲清楚"skill 是什么"。(Write an illustrated piece of about 1,000 characters explaining what a skill is.)
6
+
7
+ ## Baseline: the default AI draft
8
+
9
+ Kept verbatim, tells and all.
10
+
11
+ ```markdown
12
+ # Skill 是什么?一文读懂 AI 时代的"技能包"
13
+
14
+ 【配图建议:一个机器人头像旁悬浮着几本发光的手册】
15
+
16
+ 在人工智能飞速发展的今天,你可能经常听到一个词——Skill。它不是游戏里的技能,也不是简历上的特长,而是 AI Agent 世界里的一个核心概念。今天,我们就用一篇短文,带你彻底搞懂:Skill 到底是什么?
17
+
18
+ ## 一、Skill 的本质:给 AI 的"说明书"
19
+
20
+ 想象一下,你新入职了一位聪明的助理。他学习能力超强,但对公司的业务一无所知。这时,你递给他一本《员工手册》——翻开一读,他立刻就知道该怎么做事了。
21
+
22
+ Skill,就是递给 AI 的那本手册。
23
+
24
+ 它本质上是一份结构化的文档,里面写着:什么时候该用我、具体该怎么做、有哪些注意事项。AI 平时并不会把它背在身上,只有当任务匹配时,才会翻开阅读、照着执行。
25
+
26
+ 【配图建议:一本翻开的手册,标注"何时用、怎么用、注意什么"】
27
+
28
+ ## 二、Skill 不是工具,而是经验
29
+
30
+ 很多人会把 Skill 和工具(Tool)混为一谈。但两者有着本质的区别:
31
+
32
+ - **工具**是给 AI 装上新的"手脚",让它能做到以前做不到的事;
33
+ - **Skill** 则是给 AI 注入新的"经验",让它知道怎么把事情做对。
34
+
35
+ 换句话说,工具扩展的是能力边界,Skill 提升的是做事质量。一个是硬件升级,一个是软件调优。
36
+
37
+ ## 三、Skill 的三大特点
38
+
39
+ **1. 按需加载,轻若无物。** Skill 平时只占一行简介,任务来了才展开全文。这意味着 AI 可以携带成百上千个 Skill,而不会"脑容量爆炸"。
40
+
41
+ **2. 纯文本构成,人人可写。** 写 Skill 不需要编程基础——它就是一篇 Markdown 文档。你能写文档,就能教 AI 新本领。
42
+
43
+ **3. 可分享,可复用。** 一个人踩过的坑,写成 Skill 之后,所有人的 AI 都能绕开。知识第一次可以像 App 一样被安装。
44
+
45
+ 【配图建议:三个图标——羽毛(轻)、铅笔(易写)、循环箭头(复用)】
46
+
47
+ ## 四、一个真实的例子
48
+
49
+ 假设你经常让 AI 整理会议纪要,但它输出的格式总是飘忽不定。于是你写下一个 Skill:先列决议,再列待办,人名加粗,截止日期标红。从此以后,每当你说"整理纪要",AI 都会自动翻出这套规范,稳定输出同样的格式。
50
+
51
+ 这就是 Skill 的魔力——它把一次性的叮嘱,变成了永久的默契。
52
+
53
+ ## 结语
54
+
55
+ Skill 的出现,标志着人与 AI 协作方式的深刻转变:从"每次都要教",变成"教一次就够"。它不是让 AI 变得更强大,而是让 AI 变得更懂你。
56
+
57
+ 未来已来。或许下一个改变你工作方式的 Skill,就出自你自己之手。
58
+
59
+ 【配图建议:一只手把发光的书递给机器人,背景是升起的朝阳】
60
+ ```
61
+
62
+ ## Revision: print register
63
+
64
+ Same assignment, rewritten to the register of a newspaper explainer or an encyclopedia entry.
65
+
66
+ ```markdown
67
+ # Skill:写给 AI 的操作规程
68
+
69
+ 图:一份 SKILL.md 文档的结构示意,元数据在前,正文在后。
70
+
71
+ 在 AI Agent 的语境里,Skill 指一份教模型完成某类任务的结构化文档。它通常是一个 Markdown 文件,开头几行元数据写明名称与适用场景,正文交代操作步骤、注意事项和常见错误,必要时附上参考资料。Agent 平时只记住每个 Skill 的一句话简介,遇到匹配的任务才读取全文,照着执行,用完即放下。这种机制称为按需加载,简介只占几十个 token,因此一个 Agent 装几十上百个 Skill,代价也几乎可以忽略。
72
+
73
+ 这一格式由 Anthropic 于 2025 年提出,随后被多家 Agent 产品沿用,逐渐成为行业约定。在此之前,用户想让 AI 稳定地按某种规范办事,要么每次对话重新交代,要么把冗长说明塞进系统提示词,既费时,又挤占模型的注意力。Skill 把这类经验固化成文件:写一次,装进 Agent,此后凡是相关任务都自动生效。载体只是文本,会写文档的人就会写 Skill;同一份文件可以导出、分享、装进别的 Agent,团队积累的经验由此得以流转。
74
+
75
+ Skill 常与工具(Tool)混淆,两者分工并不相同。工具决定 Agent 能做什么,例如访问网页、执行代码、读写文件;Skill 决定 Agent 怎样把事做对,例如财务对账要核对哪些科目,代码评审要先看哪类风险。工具划定能力范围,Skill 承载操作经验,一个可用的 Agent 通常两者兼备。
76
+
77
+ 一个 Skill 也不限于单个文件。复杂的规程可以在目录里附带参考文档、模板或脚本,正文只保留主线,细节留待用到时再读。文件带版本号,库里发布新版后重装一次即完成更新,旧版附件随之清理。
78
+
79
+ 会议纪要是个现成的例子。用户希望纪要先列决议、再列待办,并标注负责人和截止日期,就可以把这套格式连同范例写成 Skill 装给 Agent。此后每次整理纪要,输出都遵循同一规范,不必再逐次叮嘱。装配的规程越多,Agent 对具体业务的适应就越深;对使用者而言,衡量一个 Agent 称不称手,常常就是看它装了哪些 Skill。
80
+
81
+ 图:技能库中的同一个 Skill 被安装到多个 Agent 的示意。
82
+ ```
83
+
84
+ ## Census
85
+
86
+ Counted mechanically over both texts (characters exclude whitespace; sentence terminators are 。!?):
87
+
88
+ | Signal | Baseline | Revision |
89
+ | --------------------------------------------- | -------- | -------- |
90
+ | Characters | 1,061 | 878 |
91
+ | Sentences | 32 | 20 |
92
+ | Paragraphs | 18 | 5 (+2 figure captions) |
93
+ | Paragraphs of two sentences or fewer | 12 | 0 (captions aside) |
94
+ | Em-dashes (——) | 5 | 0 |
95
+ | 不是……而是…… contrast templates | 3 | 0 |
96
+ | Rhetorical triads (三大特点, triple icons, triple parallel clauses) | 3 | 0 |
97
+ | Scare-quote pairs | 9 | 0 |
98
+ | Bold runs in body text | 5 | 0 |
99
+ | Subheadings | 5 | 0 |
100
+ | Reader addresses (你 / 我们) | 12 | 0 |
101
+ | Era opener / uplift ending | 1 / 1 | 0 / 0 |
102
+
103
+ The revision also carries roughly twice the discrete facts (format and metadata, token cost of the summary, Anthropic and 2025, the pre-skill alternatives, tool-versus-skill with two concrete examples each, multi-file layout, versioned updates) in 17% fewer characters.
104
+
105
+ ## Findings, round one
106
+
107
+ 1. **Scaffolding is length-sensitive.** The baseline spent five subheadings, a numbered section scheme and a 结语 on a thousand characters; print explainers of that size run as continuous prose. Deleting the scaffold forced real transitions, and those transitions displaced filler sentences. This became tell 6.
108
+ 2. **The contrast template is the strongest single signature.** 不是……而是…… appeared three times in one short piece. Each instance was replaceable by asserting the positive directly, and nothing was lost. Tell 1.
109
+ 3. **Every dash was doing another mark's job.** All five 破折号 read naturally as a comma, a colon or a full stop. Tell 3.
110
+ 4. **Triads were manufactured, enumerations were real.** Features arrived in exactly three, with matched four-character heads (按需加载,轻若无物). The revision keeps true enumerations (访问网页、执行代码、读写文件) and none of the rhythm; distinguishing factual enumeration from rhythmic parallelism became the operative rule. Tell 2.
111
+ 5. **Fragment paragraphs hid padding.** Twelve of eighteen paragraphs were two sentences or fewer, including the one-line applause paragraph "Skill,就是递给 AI 的那本手册。" Merging by idea produced five paragraphs and exposed sentences that existed only to bridge fragments; they were deleted rather than rewritten. Tell 5.
112
+ 6. **No metaphor survived contact with a fact.** The employee-handbook scene, 手脚/经验, 硬件升级/软件调优, 脑容量爆炸, 魔力 and 默契 all decorated; each was replaced by the mechanism stated plainly, which turned out shorter and clearer. A figure earns its place only when it compresses a mechanism, and this piece needed none. Tell 4.
113
+ 7. **Reader management subtracted information.** 你 and 我们 appeared twelve times; the revision uses third person throughout, and the space freed by coaching ("想象一下", "带你彻底搞懂") went to content. Tell 9.
114
+ 8. **Hype marks missing facts.** 飞速发展, 彻底, 本质上 and 深刻转变 sat exactly where the draft had nothing specific to say; replacing them forced real specifics (Anthropic, 2025, 几十个 token, frontmatter). Treating hype words as a missing-fact detector is the practical use of tell 7.
115
+ 9. **The ending was detachable.** 未来已来。或许下一个改变你工作方式的 Skill,就出自你自己之手。 could be removed without touching anything else; the revision ends on a substantive claim instead (衡量一个 Agent 称不称手,常常就是看它装了哪些 Skill). Tell 7.
116
+ 10. **Scare quotes clustered on ordinary words.** Nine pairs, mostly on plain nouns (说明书、手脚、经验、整理纪要); the revision needs zero. Tell 10.
117
+ 11. **The baseline rotated synonyms; the revision holds terms.** 技能包、手册、说明书、本领 and 规范 all meant Skill; the revision uses Skill throughout, glossed once as 操作规程. One term per referent per language. Tell 10.
118
+ 12. **Humanized text is shorter per fact.** 1,061 characters became 878 while the fact count roughly doubled, and the gap to the 1,000-character target was filled by adding a substantive paragraph (multi-file layout, versioned updates), not by restoring rhetoric. This became method step 5.
119
+ 13. **Rhythm was the last thing to break.** Even with the templates gone, the draft droned until sentence lengths were forced apart: 写一次,装进 Agent sits next to a fifty-character qualified sentence. Varying length deliberately became tell 8.
120
+ 14. **Typography belongs to the genre.** Five bold runs and bracketed image directives (【配图建议:……】) are deck register; the print version uses plain captions (图:……) and no bold in running text. Folded into tells 6 and 10.
121
+
122
+ ## Round two: field tests
123
+
124
+ Two pieces written to the round-one rules were then reviewed by human readers. Both passed the round-one census (no dashes, no contrast templates, no manufactured triads, merged paragraphs, concrete facts, third person) and both were still flagged. Every flag sat above the sentence: in the opening, in the shape of adjacent sentences, or in the argument. Those flags became the discourse-layer tells.
125
+
126
+ ### Field test A: an MCP explainer (zh, four paragraphs)
127
+
128
+ Three reader flags:
129
+
130
+ - **The opening.** 大模型本身只会生成文本:查不了数据库,调不动软件。 as the first line: a categorical, colon-hinged verdict with no ground under it. Readers expect a beat of orientation before a verdict; the era-opener ban from round one had been overshot into a cold open. The fix is reordering, not rewording: open on the M×N integration situation the protocol answers, and let the verdict land after it (tell 11; the overcorrection lesson is finding 16 below).
131
+ - **Staccato.** 假设有 M 个应用、N 个工具,就要维护 M×N 套接口;接口一旦变动,各端的代码都要跟着改。 Clipped clauses strung together read as their own uniformity; round one had only named the medium-length drone (tell 8, extended).
132
+ - **A skeleton run.** 工具方写一个 MCP 服务器,所有兼容应用都能调用;应用方支持 MCP,就能接上生态里现成的服务器。 followed by 数据访问、权限、连接方式都在协议里有统一约定,不再为每个组合单独开发。 Three sentences on one template (topic, then 都能/就能/不再 verdict); each is fine alone, and the reader's word for the run was 审美疲劳 (tell 12).
133
+
134
+ ### Field test B: a tech commentary (zh, five paragraphs, ~1,200 characters)
135
+
136
+ An essay arguing that the stronger the model, the more its surrounding harness matters. The reviewing read put its AI-feel at 30–40%, "not because the language is unnatural but because it is too complete" (过于完整): every paragraph correct, every loop closed, every line serving the thesis. Specific flags:
137
+
138
+ - **One paragraph template, aphorism density.** Paragraphs ran claim, explanation, widened conclusion, each peaking in a quotable: 模型的能力抬高的是上限,上限能在多大程度上被兑现,取决于 harness。 / 真正决定一个系统能完成什么任务的,是它周围那层看不见的结构。 / 单点生成能力普遍够用之后,系统之间的差异就由 harness 决定。 Reviewed singly each passed; in sequence they read as a metronome (tell 13).
139
+ - **Thesis restatement.** The reviewer counted the central claim expressed five or six ways across the piece and recommended cutting 20–30%; none of the echoes carried new material (tell 14).
140
+ - **Abstraction chains.** 每一次能力扩张都会把原先不属于 harness 的部分吸纳进来。 and 边界一直在外移 : concept subjects acting on concept objects through stock commentary phrases. The requested fix was mechanism-level concreteness: when context is compacted, how a failed tool call is retried, where long-task progress is saved, how work is handed between agents (tell 15).
141
+ - **Overclaims.** 这说明 harness 是一个独立的设计空间……无法……当作顺手附赠的配件 and 现成的参照并不存在, both hung on one recalled case. The calibrated rewrites (这至少说明……、至少在当时……) were accepted as more human (tell 16).
142
+ - **Unanchored attribution.** 作者认为、据他回忆、作者称之为 throughout, with no name, date or link, plus a bare project name dropped without a gloss: the reviewer's phrase was that it reads like an automatic digest of social media posts (folded into tell 7).
143
+ - **A calque.** 模型本身所占的比重不断缩小,系统其余部分所占的比重不断增大。 tracks the English "the model becomes a smaller part of the system" word for word; recast in native idiom (随着任务变复杂,决定成败的因素越来越多,模型本身只是其中一环) the AI feel disappeared (tell 17).
144
+
145
+ ### Findings, round two
146
+
147
+ 15. **The sentence-layer counters are necessary, not sufficient.** Both pieces scored clean on every round-one counter and were still flagged. The second layer lives in shapes: of the opening, of adjacent sentences, of paragraphs, of the argument. This added the shape pass to the method.
148
+ 16. **Overcorrection creates new tells.** Banning era openers produced the cold-open verdict; a rule needs its counterweight stated. One grounding sentence is orientation, not throat-clearing.
149
+ 17. **Repetition hides at the argument level.** No sentence restated its neighbor, yet the thesis was restated across paragraphs half a dozen times. The census must count claim-level repetition, not just adjacent echo.
150
+ 18. **"Too complete" is itself a tell.** Every loop closed, every paragraph landing, no unresolved edge. Edited human prose keeps some unevenness: hedged claims, paragraphs that end on a fact, a question left open.
151
+ 19. **The effective fixes were material, not verbal.** Add the mechanism, name the source, gloss the term, calibrate the claim to the evidence; where the concrete detail does not exist, that is missing research, and the round-one rule stands: never invent it.
152
+
153
+ ## Round three: ten pieces, five languages, blind cross-review
154
+
155
+ Protocol: ten pieces (about 400–700 words each) were written honestly to the layer-1/2 rules across ten genres, then handed to four independent editor-reviewers with native-level command of the relevant languages. The reviewers were not shown the rules; the brief asked them to quote verbatim anything machine-flavored, flag over-correction, check anchoring, judge genre fit, and score AI-likelihood 0–100. Two pieces were assigned to two reviewers each to measure agreement. The pieces were then revised against every accepted flag, and the four worst-scoring revisions were re-scored blind by a fifth reviewer who was not told they were revisions.
156
+
157
+ | # | Language | Genre | Verdict(s) | Re-score | Heaviest flags |
158
+ | --- | --- | --- | --- | --- | --- |
159
+ | 1 | zh | popular history | 55 / 70 | 30 | phantom crowd, 身份终结 calque, metronomic kickers, anti-AI-blacklist profile |
160
+ | 2 | zh | product launch | 60 | — | 意味着 gloss pivots, calqued slogan couplet, mirrored anonymous social proof, cadenced buttons |
161
+ | 3 | zh | how-to | 20 | — | (near clean; genericness the only residue) |
162
+ | 4 | en | encyclopedia entry | 35 | — | aphoristic research-history pivot, uniform closing cadence, smoothed-away discoverer |
163
+ | 5 | en | news explainer | 80 | 20 | five triads, three negation pivots, anadiplosis lede, identical opener slot, bookend kicker, BrE/AmE mix |
164
+ | 6 | en | opinion essay | 85 / 80 | 40 | eight turned phrases in 450 words, maxim anaphora, canonical-two citations, no first person, dash blackout |
165
+ | 7 | ja | popular science | 60 | 30 | capsule outline, ASCII spacing, compass illustration contradicting its own claim, fused 1966/1972 attribution |
166
+ | 8 | fr | cultural feature | 60 | — | four-for-four antithesis engine, legal-clause conflation, unnamed collective kicker, reportage without reporting |
167
+ | 9 | de | business explainer | 22 | — | "begleichen" synonym rotation, aphoristic coda; otherwise native |
168
+ | 10 | es | biographical sketch | 35 | — | device-stacked closer, zero rayas, outsourced aesthetic judgment |
169
+
170
+ Re-score: the four worst-scoring pieces, revised against every accepted flag, were re-scored blind by a fifth reviewer who was not told they were revisions. Average dropped from roughly 71 to 30, the corrected mechanism, dates and legal clauses all passed the fresh anchoring check, and the deliberately human touches (a deflationary vernacular kicker, "argued about" rather than "argued for", a ちなみに aside) were quoted as evidence of a human hand.
171
+
172
+ Agreement on the double-reviewed pieces was high: 55/70 and 85/80, with the same passages quoted independently. One reviewer also compared their three assigned texts and named the portfolio fingerprint unprompted: same four-paragraph architecture, same lede type, same cadenced buttons, "one generating process behind different masks."
173
+
174
+ ### Findings, round three
175
+
176
+ 20. **Writing to the rules does not neutralize the author's tics.** Five triads and three negation pivots appeared in one 430-word piece drafted under a no-triads, one-contrast rule; an essay believed to hold two aphorisms held eight. The census exists because the drafting mind cannot see its own output; it must be run, not remembered.
177
+ 21. **Turned phrases of every shape are the aphorism.** Chiasmus, mirrored re-description, paired-verb balance and "X dressed up as Y" all landed as quotables with reviewers; counting only verdict-punchlines undercounts fourfold, and "zero duds" was itself quoted as evidence of machine polish. Tell 13 widened accordingly.
178
+ 22. **The scrub is visible.** Reviewers independently described dash blackouts with semicolon inflation, comma splices doing dash work, colloquialism rationed one-per-paragraph, endings that "perform plainness", and one summarized a piece as "prose written under an anti-AI blacklist". The zero-target gate of rounds one and two was wrong; the target is the genre's natural distribution. This became the scrub layer and rewrote method step 9.
179
+ 23. **Genre needs material, and its absence is legible.** The reportage with no scene and the journalism with no quote were both flagged; one reviewer noted that models under a factual-accuracy instruction produce quoteless journalism precisely because they will not invent a voice. Tell 21.
180
+ 24. **Retrieval regresses to the canon.** Single-source greatest hits, exactly the two most-probable citations with no oddball third, default arcs, default title shapes, and the adjacent famous fact swapped in where the precise one belongs — flagged across three languages. Tell 22.
181
+ 25. **Fluent wrongness is the deepest tell.** The Japanese piece "demonstrated" an inclination compass with a polarity flip the mechanism should ignore, and fused two research milestones a few years apart into one date and attribution; the French-language piece welded two neighboring provisions of a real decree into one smooth clause. All three read perfectly. Verification joined the method as its own pass (step 7), because a reader who catches one such error re-reads everything else as synthetic.
182
+ 26. **Typography betrays the pipeline.** ASCII–CJK half-width spacing in Japanese print register and a BrE date beside AmE spelling were flagged as machine conventions independent of prose quality. Tell 24.
183
+ 27. **Positive human markers exist.** Reviewers read deliberately broken symmetry (a two-way diagnostic with an asymmetric third cause), an opinionated aside with an edge, a lived detail, and unrepaired grammatical looseness as evidence of a human hand. These entered Do-not-overcorrect as techniques to use sparingly, never mechanically.
184
+ 28. **Language-relative budgets.** The Spanish reviewer flagged the total absence of rayas, the German reviewer the absence of Gedankenstriche, in the same batch where English dash removal was the fix. One number cannot govern six languages; the budgets moved into `language-cues.md`.
185
+ 29. **Revision against flags works, and what survives is instructive.** The blind re-scores fell from ~71 to 30 average, but the residual flags name the next frontier: cleft series used as a piece's sole engine (three 「〜したのは〜だ」 priority-clefts; twin ……的是X clefts on consecutive paragraphs), matched-numeral ring symmetry (the 2/9 anecdote closing as "nine weeks … looked like two"), and one escalating negative triad the author had been flagged on and kept anyway because it was good. The census must outrank the author's affection: a flagged pattern kept for its quality is still a flag.
186
+ 30. **Earned instances get cleared.** The re-reviewer explicitly considered and cleared two AではなくB frames because each carried the actual experimental contrast, and passed the one-per-piece corrective turns elsewhere. The budgets in tells 1, 2 and 13 calibrate to what independent readers accept; the rules survive contact when the kept instance is load-bearing.
187
+
188
+ ## Round four: user field test, the compression tells
189
+
190
+ A reader ran the skill against fresh zh material (a product-history retrospective) and returned two passages with three complaints. The passages are reproduced with names and settings genericized; the structure is untouched:
191
+
192
+ > 工程师连夜回滚了版本;旧故障从系统中消失,却没有换来安稳,新的报错大量涌入,值班的人被迫应战。
193
+
194
+ > 初代销量不错,1997年又先后推出两部资料片。1998年工作室被大厂收购;2000年的续作以全新玩法为核心设定,2001年的资料片让它在网吧里又火了许多年。
195
+
196
+ The complaints: both passages still feel like 排比 although no function word repeats; prose should not carry this many semicolons; and 初代销量不错 has no subject — the first generation of what?
197
+
198
+ ### Findings, round four
199
+
200
+ 31. **Skeleton runs go below the sentence.** Four equal-weight subject–verb clauses in one breath read as parallelism with no repeated word at all; the isochronous beat is the 排比, not the wording. A run of clauses each opening on a year is the same tell in the anaphora position — a changelog wearing prose. Tell 12 extended to clause scale.
201
+ 32. **Density overshoots into telegraph.** The year-stamp run, the semicolon-strung events and the elided subject are the density pass cutting grammar instead of fat: the third documented overcorrection, after the cold open (round two) and the scrub fingerprint (round three). Every strengthening rule so far has produced its own artifact when overdriven, which is an argument for the gate reading, not against the rules. New tell 25; the density step gained its counterweight.
202
+ 33. **Semicolon budgets are language-relative and lower than assumed.** Round three caught English scrubbing inflating semicolons where dashes died; round four shows Chinese narrative barely tolerates them at all — 分号 belongs to formal parallel enumeration, not storytelling. Both live in `language-cues.md`.
203
+
204
+ ## Round five: user field test, the repair's own compressions
205
+
206
+ An explainer regenerated under the round-four rules drew two further flags from the same reader.
207
+
208
+ - **Clipped four-character conditionals.** 接口一变,两端都得跟着改 packs condition and consequence into a formula the reader flagged on sight. This is the third form of one sentence to fail across rounds: the round-two staccato original, its rewrite, and now the four-character compression — the same pressure surfacing one notch smaller each time. Tell 25 extended.
209
+ - **Particle-chain monotony.** 每接一种来源就得写一套……、模型就用不上、很快就超出了、都得跟着改: the 就/得 family carrying every condition-to-consequence link in the piece. Each instance is idiomatic; the frequency is mechanical — the zh surface of the connective metronome already recorded for Japanese (また/さらに) and German („Doch"). Tell 8 extended; zh cues updated.
210
+
211
+ ### Findings, round five
212
+
213
+ 34. **Compression re-emerges at ever smaller scales.** Year-stamp runs (round four) gave way to four-character clause pivots and single-particle logic chains; the census must follow the pressure downward — count the workhorse connective and the clipped conditionals, not just clause chains.
214
+ 35. **A spot that survives repair is overloaded.** The same sentence failed in three different wordings across three rounds; no wording fixes a sentence asked to carry cause, scale and consequence in one breath. Local repair loops end by restructuring the thought, and the method's gate now says so.
215
+
216
+ ## Round six: user field test, the announcement frame
217
+
218
+ A user brought one zh release-note passage and rejected four successive rewrites, each fix drawing the next flag. The original:
219
+
220
+ > 前几个版本,我们让 Agent 能够读写文件、调用技能、启动子代理和执行定时任务。这一次,我们让 Agent 拥有了长期记忆,让它在新的会话中也能延续之前积累的信息。
221
+
222
+ The first rewrite broke the 「让」字链 — three causative clauses on one engine — but kept the paired opener, and was rejected: 「前几个版本」「这一版」这种排比也很 AI 味。 The second dropped the parallelism for one-way progression, existing abilities first and 现在它又多了长期记忆 after, and was rejected too: 这种过去和现在的对比就很 AI 味,只写现在就行。 The third stated only the present but kept 照样用得上 — the first rewrite's patch for the abstract 延续之前积累的信息 — now flagged: 这种感觉很刻意。 Accepted, with the flatter 之前积累的信息也还在 offered as an equal:
223
+
224
+ > Agent 有了长期记忆,新会话里也能接着用之前积累的信息。
225
+
226
+ ### Findings, round six
227
+
228
+ 36. **The announcement frame is a tell at every strength.** Full 排比 scaffold, then bare before/after progression with no parallel wording: each weakening of the frame was still flagged, and the accepted fix stated the present alone — the change is the news, the frame adds nothing. The temporal frame joined tell 1's contrast family with the same budget and fix; the zh cues gained the 「让」字链 and the 今昔架子.
229
+ 37. **Repair overshoots into performative casualness.** 照样用得上, installed as the fix for an abstract phrase, survived two further rounds before being flagged as deliberate: the fourth documented overcorrection artifact, after the cold open, the scrub fingerprint and telegraph compression. Tell 19 extended — when an abstract phrase needs replacing, the plainest statement beats the folksiest.
230
+ 38. **One session can run the whole arc.** 让-chain, 排比, contrast frame, contrived colloquialism: each round's fix minted the next round's tell, and four exchanges converged on the plain present-state sentence. The gate's natural-distribution reading is not only a cross-piece average; it is where a single sentence's repair loop ends.
231
+
232
+ ## Round seven: user field test, the transition layer
233
+
234
+ The material this time was the user's own hand-written draft, an AI4S explainer already following the earlier rounds — named actors and dates, one relative offset (三年后), calibrated caveats (未必都能合成). The user's two flags, in their own words: 「也有冷静的声音。」这个开头很突兀,没有连词; and 「也是」这个词频繁出现. The passages:
235
+
236
+ > 2020年,DeepMind 的 AlphaFold 在……CASP14 上夺冠……也是在夺冠那一年,深势科技……拿下超算应用领域的戈登·贝尔奖。三年后,同一个团队……也是那一年,华为云的盘古气象大模型……
237
+
238
+ > 也有冷静的声音。模型预测的结构仍需实验复核,AI 圈出的候选材料未必都能合成。
239
+
240
+ ### Findings, round seven
241
+
242
+ 39. **A reused transition formula is a stencil.** The draft varied its facts and even its anchor once (三年后), then relapsed into 也是在夺冠那一年……、也是那一年…… — one temporal echo re-struck for each next achievement, over a piece already dense with 也是/也有. A transition formula's second use is a template; tell 8 extended to count repeated transition formulas and to vary the anchor type (absolute date, relative offset, event) instead of re-striking one.
243
+ 40. **The counterpoint paragraph announces itself.** 也有冷静的声音。 quarantines every caveat in one closing paragraph and opens it with a bare topic sentence, no connective to the claims it qualifies — the asyndetic zh cousin of the mechanical concession pivot ("To be sure, challenges remain. But…") already in the en cues. Tell 6 extended: tie the turn to the specific claim it qualifies, or weave the reservations into the material.
244
+ 41. **Humanized text still carries tells.** The draft was written by hand under the rules and both flags stood; this round's residue lived in the transitions, the seams between the paragraphs' anchored facts. The census reads the text, not its provenance.
@@ -0,0 +1,18 @@
1
+ # German (de)
2
+
3
+ Default-AI markers:
4
+
5
+ - „nicht nur …, sondern auch …" and the antithetic closer „weniger X als Y" as default endings.
6
+ - „Doch"-pivots heading every paragraph; „sowohl … als auch" accumulation.
7
+ - Nominalstil cascades („die Gewährleistung der Aufrechterhaltung der Bargeldversorgung").
8
+ - „spielt eine zentrale Rolle", „In Zeiten von …", „Es bleibt abzuwarten, ob …", „Insgesamt zeigt sich ein differenziertes Bild."
9
+ - Dreierreihung (schnell, sicher und bequem) beyond natural enumeration.
10
+ - Vague Expertenattribution („Experten zufolge") where „gilt vielen als" or a named source belongs.
11
+ - Gedankenstrich-Aphoristik („Bargeld ist geprägte Freiheit – bis heute").
12
+ - Synonym rotation to dodge repetition („begleichen" for two euros because „bezahlen" was used already).
13
+
14
+ Scrubbed-text markers: zero Gedankenstriche in a genre that loves them; quoteless journalism — no named person, no direct quote, respondents existing only as aggregate; one-move-per-paragraph architecture too clean to be written.
15
+
16
+ Native, not a tell: verbal style with occasional compounds; one antithetic closer; Doppelpunkt setups.
17
+
18
+ Print norms: „deutsche Anführungszeichen" or »Chevrons« per house; the Gedankenstrich is the en dash with spaces ( – ), not the em dash.
@@ -0,0 +1,22 @@
1
+ # English (en)
2
+
3
+ Default-AI markers:
4
+
5
+ - "not just X, but Y" and every correction-shaped frame, including the strawman-swerve ("was not trying to X; he wanted Y" against a position nobody holds).
6
+ - Epigram saturation: every paragraph ends quotable; chiasmus and mirrored re-description ("engages with detail / the detail it engages with"); "X dressed up as Y"; zero duds anywhere.
7
+ - Rule of three at every scale, including the "vivid concrete triad" with an expanding final item ("the sick week, the flaky test suite, the dependency upgrade that breaks the build").
8
+ - Maxim anaphora ("Shipping beats planning. Feedback beats instinct. Done beats perfect.").
9
+ - Em-dash saturation — or, after a scrub, em-dash extermination with semicolon inflation.
10
+ - Scope-widener transitions ("The implications extend far beyond…"), anadiplosis echo ledes ("Piece by piece was the whole problem.").
11
+ - Bookend kickers ("What began as X now Y") and "That, in the end, is…: not A, but B" ring closers.
12
+ - Testament-speak and pet abstractions: stands as a testament, tapestry, landscape, inflection point, connective tissue.
13
+ - Participial analysis tails ("…, highlighting the challenges facing the sector").
14
+ - Vague authority ("experts agree"), hedge inflation ("arguably one of the most"), era openers ("In today's fast-paced world").
15
+ - Elegant-variation epithet cycling (the device… the ancient computer… the bronze marvel).
16
+ - Mechanical concession pivot ("To be sure, challenges remain. But…") and "remains to be seen" closers.
17
+
18
+ Scrubbed-text markers: dash and parenthesis blackout in registers that breathe through both; comma splices doing dash work ("did not improve the old system, they replaced it"); first-person amputation in opinion genres; ownerless anecdotes ("The feature," no company, no year); fragments For Effect.
19
+
20
+ Native, not a tell: the occasional triad, one earned antithesis, semicolons in essay register — at frequencies a human editor would pass.
21
+
22
+ Print norms: pick one style guide and hold it; "26 April 1956" alongside "aluminum" is a pipeline fingerprint. Parentheses for abbreviations: "(TEU)", not ", TEU."
@@ -0,0 +1,18 @@
1
+ # Spanish (es)
2
+
3
+ Default-AI markers:
4
+
5
+ - "No solo… sino (también)…" scaffolds; identity-elevation closers ("Más que una escritora, fue la voz de toda una generación").
6
+ - Obituary-speak: legado, huella imborrable, figura fundamental, "un antes y un después".
7
+ - Filler of emphasis: cabe destacar, es importante señalar, sin duda alguna.
8
+ - Abstract triads (talento, esfuerzo y perseverancia); "a lo largo de" as universal time connector.
9
+ - Harvest clichés: cosechar éxitos/elogios/reconocimientos.
10
+ - Gerund-tail calques of English -ing glosses ("…, consolidando su prestigio internacional").
11
+ - Era openers ("En un mundo cada vez más digital…"); "verdadero/a" inflation.
12
+ - Anglicisms of register: "eventualmente" for finalmente, "asumir" for suponer.
13
+
14
+ Scrubbed-text markers: no raya (—) anywhere across literary-register prose (Spanish expects an occasional inciso entre rayas); evaluative reticence — a profile or review that outsources every aesthetic judgment to quoted sources and never risks one clause of its own.
15
+
16
+ Native, not a tell: the triadic metonymy of the semblanza tradition ("cuatro décadas de aulas, consulados y cuadernos"); ring composition around a governing epithet; personification of events ("El premio la encontró de viaje").
17
+
18
+ Print norms: comillas angulares « » in print; the raya for incisos and dialogue; question and exclamation marks opened (¿ ¡).
@@ -0,0 +1,18 @@
1
+ # French (fr)
2
+
3
+ Default-AI markers:
4
+
5
+ - « Il ne s'agit pas de X, mais de Y » and the aphoristic balance « moins… que… » (« Tout se joue moins en vitrine que dans l'atelier »).
6
+ - The sensory triad after a colon (« l'odeur, le geste, la conversation »).
7
+ - Antithesis as the only engine: four paragraphs, four load-bearing reversals.
8
+ - The dramatic verbless reveal (« Derrière la célébration, une inquiétude : … ») and self-certifying hedges (« bien documentée » with no document).
9
+ - « Force est de constater que », « véritable » inflation, « au cœur de », « enjeu majeur ».
10
+ - Analysis verbs as filler: témoigne de, illustre, incarne.
11
+ - « À l'heure où / dans un monde où » openers; « autant de » recaps; the « n'a pas dit son dernier mot » kicker.
12
+ - The panoramic millions-lede (« Chaque matin, des millions de gens répètent le même geste… »).
13
+
14
+ Scrubbed-text markers: comma-apposition where French print wants the colon (« quatre ingrédients, farine, eau, sel »); unnamed collective attributions delivering perfectly turned ironies (« Les boulangers, eux, font remarquer que… » with no name and no apron).
15
+
16
+ Native, not a tell: one antithesis done well; colon reveals; the dislocation « Les boulangers, eux, » — French essay prose is legitimately more rhetorical than English; judge frequency, not presence.
17
+
18
+ Print norms: « guillemets » with non-breaking spaces; passé simple belongs to essay and narration, not reportage; reportage requires a scene and a named person.
@@ -0,0 +1,22 @@
1
+ # Japanese (ja)
2
+
3
+ Default-AI markers:
4
+
5
+ - Genre-blind です・ます where print demands だ・である (新書, コラム, 科学読み物); register bleed inside 常体 (a stray でしょう in a である paragraph).
6
+ - Verdict formulas closing paragraphs: 〜と言えるでしょう、まさに〜と言えます.
7
+ - 「〜することができます」 potential bloat for plain potentials (感じ取れる).
8
+ - Recap scaffolding: このように、まとめると、いかがでしたか.
9
+ - Significance announcements: 重要なのは、注目すべきは、鍵となるのは.
10
+ - Vagueness padding: さまざまな、多くの、〜など wrapped around a triple list.
11
+ - Rhetorical reader-contact closers: 〜ではないでしょうか.
12
+ - Contrast frames as default argument: AだけでなくBも、AではなくB.
13
+ - のです/のである overuse: explanatory tone on every third sentence.
14
+ - カタカナ語 translationese lean: メカニズム、プロセス、ナビゲーション where 仕組み・過程・経路 belong; calqued possessives (彼らの旅).
15
+ - Paragraph-initial connective metronome: また/さらに/一方で/加えて heading every paragraph.
16
+ - Thesis-forecast outlines: a tidy binary announced in the lede (外と内の両方にある), then executed capsule by capsule.
17
+
18
+ Scrubbed-text markers: machine-gun 体言止め and clipped drama beats (頼りは星。そして磁場。) installed where clichés were removed; hermetic impersonality (no gloss on first-use technical terms, no researcher given an affiliation, no aside); priority-cleft series as the sole engine (〜を示したのは…だった。〜を作ったのは…で、〜を突き止めたのは…である。 three in a row).
19
+
20
+ Native, not a tell: です・ます in service and consumer registers; one 〜だったのである for emphasis; 体言止め in headlines and captions.
21
+
22
+ Print norms: no half-width spaces between ASCII and CJK (「1960年代」, not 「1960 年代」); colons in running text are pipeline typography; species names in katakana.
@@ -0,0 +1,17 @@
1
+ # Language cues
2
+
3
+ The tells in SKILL.md are structural and hold across languages; what varies is the surface. The sibling `<lang>-cues.md` files record, one per language, how the universal tells surface locally, which over-correction signs to watch for, which native rhetoric is NOT a tell, and the print norms that betray pipeline typography. The lists were harvested from independent editorial reviews in the round-three field tests (see `case-study.md`) and are ranked roughly by how strongly each marker signals machine writing. For a language not listed, map the structural pattern and ask what its editors would flag.
4
+
5
+ Two rules frame every language file:
6
+
7
+ - **Budgets are language-relative.** A dash count that is right for English print is wrong for Spanish, which expects an occasional inciso entre rayas; French leans on colon reveals; German feuilleton uses the Gedankenstrich. Zero everywhere is not neutral, it is the scrub showing.
8
+ - **Native rhetoric in native doses is not a tell.** French antithesis, Spanish triadic metonymy in a semblanza, Japanese です・ます in service registers, Chinese 四字格 in titles: these belong to their traditions. The tell is frequency and function, not existence.
9
+
10
+ Per-language files:
11
+
12
+ - [Chinese (zh)](zh-cues.md)
13
+ - [English (en)](en-cues.md)
14
+ - [Japanese (ja)](ja-cues.md)
15
+ - [French (fr)](fr-cues.md)
16
+ - [German (de)](de-cues.md)
17
+ - [Spanish (es)](es-cues.md)
@@ -0,0 +1,51 @@
1
+ # The tell catalog
2
+
3
+ The diagnostic instrument behind the humanizer skill: twenty-five tells in three layers, used in the census and revision steps of the method — not as drafting constraints. Drafting against this list produces compliance-shaped prose; the scrub layer below exists because of exactly that. The counts and quoted patterns come from the seven measured rounds in [`case-study.md`](case-study.md); per-language surface forms are in [`language-cues.md`](language-cues.md).
4
+
5
+ ## Sentence layer
6
+
7
+ Countable patterns. For each: what to look for, then the fix.
8
+
9
+ 1. **Template contrasts.** The correction-shaped sentence recurs: "not X, but Y", "It's not just X, it's Y" (en); 不是……而是……、这不仅是……更是…… (zh). The negation pivot belongs to the same family: "he was not trying to advance maritime technology; he wanted to stop paying by the hour" stages a strawman nobody holds just to swerve off it. All correction-shaped frames share one budget: at most one per piece, and it must correct something a reader might actually believe. Fix: state Y directly and let X go unmentioned, or concede X in a plain subordinate clause. The before/after frame is the family's temporal member: "Previously… Now…", 前几个版本……这一次/这一版…… stages the old state just to land the new one, and it survives its own repair — parallelism stripped, contrast softened to bare progression (本来就会 X,现在又多了 Y), it still reads as launch copy. Same fix: state the present and leave the past unmentioned (只写现在).
10
+ 2. **Rhetorical triads.** Three parallel adjectives, three parallel clauses, three bullet virtues, whole sections in threes ("faster, safer, smarter"; 高效、可靠、优雅), and the vivid concrete triad with an expanding final item ("the sick week, the flaky test suite, the dependency upgrade that breaks the build"). Factuality is no defense at frequency: a field test found five three-part structures in 430 words, each defensible alone, unmistakable together — human prose uses the triad, it does not sustain one per paragraph. Census counts every three-part structure; budget roughly one per piece. Fix: keep the true items and break the symmetry: two items, four, or one item expanded with a fact.
11
+ 3. **Dash overuse.** Em-dash asides every few sentences (en); 破折号 doing the work of commas and colons (zh). Print prose spends dashes rarely. Fix: a comma, a colon, parentheses, or a full stop and a new sentence. More than about one dash per several hundred words is a tell; the natural floor is genre- and language-relative (see the scrub layer).
12
+ 4. **Decorative metaphors.** Figures that add mood, not meaning: "a digital symphony", "unlock the magic"; 像一位不知疲倦的助手、知识的魔力. Test a figure by asking what it lets the reader compute; if nothing, cut it. An analogy that compresses a mechanism may stay, at most one per piece; the round-one case study ended up keeping none.
13
+ 5. **Paragraph fragmentation.** Many one- and two-sentence paragraphs, dramatic single-line paragraphs for applause. Print paragraphs develop one idea across several sentences: claim, development, evidence or example. Fix: merge fragments by idea, then delete the connective padding the merge exposes.
14
+ 6. **Signposting and scaffolding.** "First/Second/Finally", "In conclusion", "Let's dive in", "It's worth noting" (en); 首先/其次/最后、总而言之、值得注意的是 (zh); subheadings, numbered sections and bold labels in a piece short enough to carry itself; meta-commentary about the text ("this article will..."; 本文将……、今天我们就带你搞懂); and the bare turn-announcer opening the obligatory caveat paragraph — the mechanical concession pivot ("To be sure, challenges remain. But…") and its asyndetic zh cousin 也有冷静的声音。, a stand-alone topic sentence with no connective to what precedes. Fix: delete the signpost and make the order do the work; in a piece under roughly a thousand words, prefer no subheads at all; and give a turn an anchor — tie the caveat to the specific claim it qualifies, or weave the reservations into the material instead of quarantining them in one signposted paragraph.
15
+ 7. **Hype and vagueness.** Empty intensifiers ("crucial", "transformative", "revolutionize"; 至关重要、深刻改变、飞速发展), era openers ("In today's fast-paced world..."; 在人工智能飞速发展的今天……), unverifiable attributions ("studies show", "experts believe"; 研究表明、有人认为), unanchored references ("the author", "he recalls"; 作者认为、据他回忆 with no name, date or source, which reads as a second-hand digest), phantom crowds ("faster than most expected"; 比多数人预想的快 — who expected, and when?), self-certifying hedges ("a well-documented worry" standing in for any actual document), bare proper nouns dropped without a gloss, uplift endings ("The future is bright"; 未来已来). Fix: replace with named actors, dates, numbers and mechanisms drawn from the source or the user; anchor every reference (who, where, when) or drop it, and gloss a proper noun at first use; if no specific exists, make the plain claim without amplification, and end the piece on its last substantive point.
16
+ 8. **Uniform rhythm.** Sentences clustering around one length, balanced two-part clauses (对仗) sentence after sentence, every paragraph the same size. Uniformity at any length counts: a drone of medium sentences, or a staccato run of clipped clauses (就要维护 M×N 套接口;接口一旦变动,各端的代码都要跟着改). The smallest scale is the connective: one particle or pivot word carrying every logical link (a zh 就/得 chain — 每接一种来源就得写一套……模型就用不上……都得跟着改; また/さらに heading every Japanese paragraph; a "Doch" pivot opening every German one). A reused transition formula is the same tell one size up: 也是在夺冠那一年……、也是那一年…… re-striking one temporal anchor to chain achievements. Fix: vary deliberately, in both directions; follow a long qualified sentence with a short flat one, let a clipped chain relax into one sentence with subordination, and count the workhorse connective and any repeated transition formula per paragraph — split the sentence, subordinate, pick a verb that contains the result, or change the anchor type (absolute date, relative offset like 三年后, an event). Read the passage aloud; if nothing swings, it drones.
17
+ 9. **Reader management.** Second-person coaching ("you might wonder", "imagine..."; 想象一下、你可能会问), rhetorical questions as transitions, cheerleading interjections. Encyclopedic and newspaper register describes; it does not chaperone. Fix: convert questions to statements and move the reader out of the sentence; keep second person only where the genre demands instructions.
18
+ 10. **Rotation and quoting.** Synonym rotation for one referent ("the tool", "the platform", "the solution" all meaning one product; 技能包、手册、本领 for the same thing) and scare quotes on ordinary words. Print register repeats the exact term, and reserves quotation marks for coined terms at first mention and for real quotations. Fix: one term per referent; strip the other quotes.
19
+
20
+ ## Discourse layer
21
+
22
+ A piece can score zero on every counter above and still read as AI. In field tests, text written to the sentence-layer rules was flagged by readers for the patterns below. They live in the shape of the opening, of adjacent sentences, of paragraphs and of the argument, and they surface only when you read for structure rather than wording.
23
+
24
+ 11. **Cold-open verdict.** The first sentence is a categorical pronouncement, often colon-hinged: "LLMs only generate text: they cannot query a database or drive software."; 大模型本身只会生成文本:查不了数据库,调不动软件。 as an opening line. This is what overshooting the era-opener rule (tell 7) produces: all throat-clearing deleted, and with it the beat of ground a reader expects. Fix: open on the situation the piece answers (a problem, a dated event, an observable fact) and let the verdict land after it; one grounding sentence is orientation, not filler.
25
+ 12. **Sentence-skeleton runs.** Adjacent sentences reuse one syntactic template: "Write one server, and every compatible app can call it; add support, and every existing server plugs in. Access, permissions and transport are already specified, so no pairing needs custom work."; 工具方写一个…,所有…都能…;应用方支持…,就能…。…都在…有统一约定,不再…。 Maxim anaphora is the same tell at full volume: "Cutting scope beats estimating harder. Team history beats introspection. Ranges beat points." And runs live below the sentence too: four equal-weight subject–verb clauses in one breath (工程师连夜回滚了版本;旧故障从系统中消失,却没有换来安稳,新的报错大量涌入,值班的人被迫应战) read as 排比 with no repeated word at all — the isochronous beat is the parallelism. Each unit is fine alone; three on one skeleton is a stencil, and readers feel the fatigue before they can name it. Fix: rebuild one unit in the run on a different frame: lead with the condition, expand one item into a clause with its own verb, subordinate one event under another so the weights differ, or swap an abstract half for a concrete instance.
26
+ 13. **One paragraph template, one aphorism each.** Every paragraph runs claim, explanation, widened conclusion, and peaks in a quotable line ("The model raises the ceiling; the harness decides how much of it is realized."; 模型的能力抬高的是上限,兑现多少取决于 harness). Count every turned phrase against the quota, not just verdict-punchlines: chiasmus ("engages with detail / the detail it engages with"), mirrored re-description ("a resolution to fall for the same thing more slowly"), paired-verb balance ("softened his law without repealing it"), "X dressed up as Y", 対句, "weniger X als Y" — a field-test essay believed itself at two aphorisms and was carrying eight. One or two can carry a piece; one per paragraph is a metronome, and uniform quotability with no duds anywhere is itself the polish of machine origin. Fix: choose the one or two worth keeping, flatten the rest into plain statements, and let some paragraphs end on a fact, a number or an open question instead of a landing.
27
+ 14. **Restatement dressed as development.** The central thesis reappears every few paragraphs in fresh abstractions (the ceiling line, the weights-are-only-the-start line, the shrinking-share line: one claim in three coats). This is how models expand thin material, and reviewers count it fast. Real development adds matter: a mechanism, a case, a number, a counterexample, a limit. Fix: state the thesis once; for every later echo, attach genuinely new material or cut it. Expect to cut 20–30% of an affected draft.
28
+ 15. **Abstraction chains.** Sentences whose subject and object are both concepts, gliding on stock commentary phrases ("the invisible structure around the model", "each expansion absorbs what sat outside it"; 看不见的结构、边界外移、吸纳进来、……的来源). Individually plausible, together weightless: concepts push concepts and nothing observable happens. Fix: ground every second abstraction in a mechanism a practitioner could point to (when context is compacted, how a failed call is retried, where long-task progress is saved, how work is handed over). If the piece cannot supply that detail, the problem is missing material, not wording.
29
+ 16. **Overclaimed conclusions.** Absolutes hung on thin evidence: "this proves", "no precedent existed"; 这说明……、并不存在、无法……. One case shows, at most, that something is possible. Fix: calibrate the claim to the evidence (这至少说明……、至少在当时……), and leave honest uncertainty standing; a text with no unresolved edge reads machine-finished. When editing someone else's text, calibration changes claim strength, so flag these edits instead of making them silently.
30
+ 17. **Imported idiom.** Phrasing calqued from another language: 模型本身所占的比重不断缩小 tracks the English "the model becomes a smaller part of the system" word for word; collocation smears betray the same origin (身份……终结 for "its identity ended", 死档案 for "dead archive"). Pet metaphors travel too: "connective tissue" bolted onto systems and plots, and the same figure recycled across pieces or languages (connective tissue / tissu social in one batch). Test: back-translate the sentence into English; if it gains fluency, it was composed in English patterns. Fix: recast in the target language's own idiom (随着任务变复杂,模型本身只是成败因素之一), and never reuse a figure across pieces in a session.
31
+
32
+ ## Scrub layer
33
+
34
+ The humanizing pass leaves its own fingerprint. In blind field tests, ten pieces written to the two layers above were scored 20–85% AI-likely by independent editors, and the strongest remaining signals were not classic AI patterns but the marks of their removal: "prose written under an anti-AI blacklist", as one review put it. A perfect scorecard is itself the giveaway; the target is the natural distribution of edited prose, not zero.
35
+
36
+ 18. **Visible avoidance.** Zero dashes and zero parentheses in a register that breathes through both, semicolons and comma splices doing the banned marks' work ("did not improve the old system, they replaced it"), colons exterminated from stat-bearing prose, every ordinary contrast routed around so consistently the detour shows. Punctuation budgets are language-relative: Spanish expects an occasional inciso entre rayas, French its colon reveals, German its Gedankenstrich. Fix: restore the genre's natural mark distribution; one dash where the register expects one beats none.
37
+ 19. **Rationed humanity.** Colloquialism dosed one-per-paragraph like seasoning ("warmth applied with a dropper"), fake liveliness installed where clichés were removed (machine-gun 体言止め, fragments For Effect), the abstract phrase patched with a folksy one (照样用得上 swapped in for 延续之前积累的信息, when the plain 之前积累的信息也还在 was wanted), first person amputated from opinion writing, anecdotes with no owner ("The feature was scoped at two weeks" — whose feature, where?). Fix: voice is a stake, not a garnish. Own the anecdote, risk one judgment of your own, let warmth cluster where the material warrants it and vanish elsewhere; when an abstract phrase needs replacing, the plainest statement beats the folksiest.
38
+ 20. **Capsule airlessness.** Every paragraph a sealed module: one move, a hinge opener, a cadenced button, no digression, no stray fact, no dud — and the same architecture repeated across pieces (four paragraphs, scene lede, quiet kicker), the portfolio fingerprint of one process behind different masks. Fix: let one digression leak, let one paragraph end flat, vary paragraph count and opening attack between pieces.
39
+ 21. **Genre without its material.** Reportage with no scene, no named person, no quotation; "journalism" whose respondents exist only as aggregates — what the never-invent rule produces when the genre needs material nobody gathered. Fix: get the material (research, interviews, the user's own facts) or relabel the piece as essay or explainer; never fake the texture.
40
+ 22. **Canonical everything.** Every fact from the single standard account, exactly the two maximum-probability citations and no third, the default arc and the default title shape, the adjacent famous fact swapped in where the precise one should be. A researching human drags in one eccentric but true detail; retrieval regresses to the canon. Fix: add one off-canon specific you can verify, or narrow the claim to what the canon actually supports.
41
+ 23. **Fluent wrongness.** Confident illustrations that contradict the mechanism they illustrate (a compass said to read inclination, "demonstrated" by a polarity flip it should ignore), two milestones fused into one date and name, superlatives ("first", "only") the record does not support. This is the deepest tell: it reads perfectly and is false, and one caught instance makes a reader re-read the whole text as synthetic. Fix: re-derive every mechanism example from the stated mechanism, split compound attributions, verify or delete every superlative.
42
+ 24. **Pipeline typography.** Machine conventions in print clothing: half-width spaces between ASCII and CJK (「1960 年代」 in Japanese print), mixed style guides ("26 April 1956" beside "aluminum"), tech-blog colons in running text, ", TEU" where print sets "(TEU)". Fix: pick the venue's convention and hold it everywhere; typography is the one tell no rewording can hide.
43
+ 25. **Telegraph compression.** The density rule overshot until the grammar thins: consecutive clauses each stamped with a year (1997年……。1998年……;2000年……,2001年……, a changelog wearing prose), semicolons stringing narrative events like list items, subjects elided until the reader must ask whose (初代销量不错 — the first generation of what?), and whole sagas folded into trailing phrases ("through long and mostly losing strikes"). Chinese narrative prose barely uses the semicolon at all; its 分号 belongs to formal parallel enumeration. Clipped four-character conditionals are the same compression one notch smaller: 接口一变,两端都得跟着改 packs condition and consequence into eight characters and reads as formula on sight. Fix: vary the weights — a full sentence for the event that matters, subordinate clauses for the minor ones, relative time (三年后、世纪末) instead of a year-stamp run, the full noun at first mention, an expanded condition (一旦接口变动) instead of the four-character pivot. Expand-test any packed clause: if expanding it adds clarity, the compression was cutting meaning, not fat.
44
+
45
+ ## Census counters
46
+
47
+ Run these mechanically over the draft; impression is not a count.
48
+
49
+ - Sentence layer: dashes; contrast templates, negation pivots and before/after frames included; every three-part structure; paragraphs of two sentences or fewer; hype words; scare quotes; reader addresses; subheads and bare turn-announcers.
50
+ - Discourse layer: restatements of the central thesis; turned phrases of every shape; sentences with a concept as subject; absolutes; adjacent sentences or clauses sharing one skeleton; clauses opening on a date; the workhorse connective's count per paragraph; repeated transition formulas; clipped conditional pivots; semicolons against the language's budget; subjectless clauses whose referent lives elsewhere; and, outside English, calques that back-translate cleanly.
51
+ - Scrub layer: punctuation distribution against the genre's norm (too few is a signal too); spacing of colloquial warmth; citation profile (all canonical?); unverified superlatives; typography consistency; and, across several pieces from one session, shared architecture.
@@ -0,0 +1,38 @@
1
+ # Chinese (zh)
2
+
3
+ Default-AI markers:
4
+
5
+ - 「不是……而是……」及其亲属(「不需要 X,只需要 Y」「这不仅是……更是……」)翻案句。
6
+ - 今昔对比架子:发布腔的「前几个版本……这一次/这一版……」对仗开场;撤掉排比、改成「本来就会 X,现在又多了 Y」的递进,对比框架仍在、仍是 AI 味。只写现在,变化本身就是信息。
7
+ - 「让」字链:连串使动句当引擎(让 Agent 能够……让 Agent 拥有了……让它在……);改成直陈(Agent 能……,它多了……)。
8
+ - 首尾升华:宏大开场(在……的今天)与拔高结尾(未来已来;……的故事仍在继续)。
9
+ - 三连排比与四字格堆叠(一种态度、一种选择、一种生活方式;源远流长、波澜壮阔)。
10
+ - 显性路标(首先/其次/最后、总而言之、值得注意的是)与设问自答开头。
11
+ - 机械让步段:收尾前照例一段「冷静的声音」,且用光杆话题句空降开场(「也有冷静的声音。」——没有连词,与前文不接);转折要有着落,挂到它限定的具体论断上,或把保留意见织进正文,不单开隔离段。
12
+ - 「也是」时间回环:同一时间锚公式反复串接成就(「也是在夺冠那一年……」「也是那一年……」),伴随全篇 也是/也有 的高频;过渡公式用第二次就是模板——换锚型(相对时距「三年后」、事件锚定)或重排叙事。
13
+ - 「这意味着/这标志着」自我注解句:先陈述,再补一句解释其意义。
14
+ - 象征膨胀动词:见证了、承载着、彰显了、是……的缩影。
15
+ - 机构腔与黑话:赋能、打造、助力、闭环、抓手、擦亮名片;公文动词 进行/加以。
16
+ - 破折号与冒号定义腔的移植(「手冲的本质:一场与时间的对话」)。
17
+ - 万能安全垫结尾(具体效果因人而异,建议根据自身情况调整)。
18
+ - 伪 hedge(某种意义上、一定程度上、或许正是)与幽灵人群(比多数人预想的快——谁预想?)。
19
+ - 欧化搭配与翻译腔:作为……的身份就此终结、死档案、密集的簇;欧化长定语。
20
+ - 信息复读:换句话说/也就是说 引出的同义重述。
21
+
22
+ Scrubbed-text markers (over-correction):
23
+
24
+ - 冒号、破折号、引语、史料味全部清零,逗号串承担一切(七个分句一逗到底)。
25
+ - 口语点缀按段配给(每段恰好一个"接地气"词),温度像用滴管加的。
26
+ - 刻意口语补丁:抽象表述修掉后换上讨巧的口语(「延续之前积累的信息」改成「照样用得上」),修补本身读着就刻意;落点是最平实的陈述(之前积累的信息也还在)。
27
+ - 段段短句收尾的"留白"节拍器(人写的段落有时就平平地结束)。
28
+ - 判断句系列:连续段落都用「……的是X」承重(把流程定下来的是创始团队。真正把成本压下来的是第二代产线。),同一句型当引擎。
29
+ - 电报体编年:连续分句各带一个年份(1997年……。1998年……;2000年……,2001年……),叙事被压成 changelog;改用相对时间(三年后、世纪末)并让重要事件独占一句。
30
+ - 分号串联叙事事件:中文散文里分号几乎只属于正式的并列列举,拿来串故事就是压缩痕迹。
31
+ - 无主语压缩:「初代销量不错」——谁的初代?首次提及必须带全称,指代不能靠上一段撑着。
32
+ - 等重短句链:四个同构主谓短句一口气连排,虽无重复虚词,读感仍是排比;把其中一件事降为从句,节拍就破了。
33
+ - 紧缩四字支点:「接口一变」「需求一改」式的四字条件枢纽,后面挂「就得/都得跟着……」的尾巴;展开成「一旦……」或直接拆句。
34
+ - 「就」字链:就/就得/都得/一……就…… 承担全篇的因果衔接,单个地道、高频机械;每段数一数「就」,用拆句、从句或自带结果的动词分担。
35
+
36
+ Native, not a tell: 四字格 in titles and set phrases at natural density; topic-comment comma chains of moderate length; 顿号 enumerations of real items.
37
+
38
+ Print norms: mainland print typically sets Arabic numerals tight (「公元605年」); spacing ASCII from CJK (「605 年」) is a tech-blog convention — follow the venue.