@qvac/skills 0.1.10 → 0.1.12

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (33) hide show
  1. package/bundled.js +11 -29
  2. package/hash.js +1 -1
  3. package/package.json +1 -1
  4. package/skills/excel/SKILL.md +146 -111
  5. package/skills/gmail/SKILL.md +7 -4
  6. package/skills/google-calendar/SKILL.md +7 -4
  7. package/skills/google-docs/SKILL.md +7 -5
  8. package/skills/google-drive/SKILL.md +7 -4
  9. package/skills/google-sheets/SKILL.md +7 -5
  10. package/skills/music-generation/SKILL.md +15 -0
  11. package/skills/pdf/SKILL.md +146 -96
  12. package/skills/presentations/SKILL.md +133 -109
  13. package/skills/sheet-music/SKILL.md +93 -0
  14. package/skills/word/SKILL.md +144 -124
  15. package/skills/excel/references/create.md +0 -374
  16. package/skills/excel/references/edit.md +0 -357
  17. package/skills/excel/references/read.md +0 -99
  18. package/skills/pdf/references/create.md +0 -169
  19. package/skills/pdf/references/transform.md +0 -270
  20. package/skills/pdf/scripts/decrypt.py +0 -26
  21. package/skills/pdf/scripts/encrypt.py +0 -25
  22. package/skills/pdf/scripts/extract_text.py +0 -25
  23. package/skills/pdf/scripts/merge.py +0 -21
  24. package/skills/pdf/scripts/rotate.py +0 -27
  25. package/skills/presentations/references/create.md +0 -399
  26. package/skills/presentations/references/edit.md +0 -314
  27. package/skills/presentations/references/read.md +0 -127
  28. package/skills/word/references/create.md +0 -400
  29. package/skills/word/references/paragraphs.md +0 -125
  30. package/skills/word/references/read.md +0 -141
  31. package/skills/word/references/rework.md +0 -767
  32. package/skills/word/scripts/list_paragraphs.py +0 -25
  33. package/skills/word/scripts/replace_paragraphs.py +0 -58
@@ -1,767 +0,0 @@
1
- # Reworking an Existing Word Document (python-docx)
2
-
3
- Change, replace, extend, trim, or rework a `.docx` that is already in this
4
- chat by running Python through the `exec` tool: stage it as an input, modify
5
- paragraphs and tables, and save a **new** output such as
6
- `existing_revised.docx`. Never overwrite the staged input.
7
-
8
- ## Replacing Some Facts, Points, Bullets or Paragraphs: Two Script Calls
9
-
10
- "Replace the first 10 facts", "change fact 3", "swap these bullets for those",
11
- "reword paragraph 7" — any request that changes some of the paragraphs and
12
- keeps the rest — is two `exec` calls that run scripts bundled with this skill.
13
- **Write no Python for it.** Nothing else in this file applies to that job: no
14
- `command`, no `Document(...)`, no fingerprint, no loop. `references/paragraphs.md`
15
- is this same recipe with its error table.
16
-
17
- The document's `attachmentId` is already in the chat — in the `attachments` of
18
- the `exec` result that produced it, or on the user's `[Attached file …]` line.
19
- Copy it into `inputs`; an entry with only a `path` stages an image, not a
20
- document.
21
-
22
- Step 1 — list the paragraphs (no `outputs`):
23
-
24
- ```json
25
- {
26
- "language": "python",
27
- "packages": ["python-docx==1.2.0"],
28
- "inputs": [{ "attachmentId": "<real id>", "path": "existing.docx" }],
29
- "skill": "word",
30
- "script": "scripts/list_paragraphs.py",
31
- "scriptArgs": ["existing.docx"]
32
- }
33
- ```
34
-
35
- It prints one line per paragraph — `index`, style, text — then a count line.
36
- Only body paragraphs (`Normal`, `List Bullet`, `List Number`) are facts, points,
37
- or bullets; `Title`, `Heading N`, and an intro sentence are never counted as
38
- one. "The first 10 facts" = the first 10 body-paragraph indexes after the
39
- heading or sentence that introduces them — not indexes 0–9.
40
-
41
- Step 2 — replace exactly those paragraphs (one `outputs` entry). `scriptArgs`
42
- is the input name, the output name, then one `index, new text` pair per
43
- replaced paragraph — "the first 10 facts" is 10 pairs, each a different
44
- complete sentence:
45
-
46
- ```json
47
- {
48
- "language": "python",
49
- "packages": ["python-docx==1.2.0"],
50
- "inputs": [{ "attachmentId": "<real id>", "path": "existing.docx" }],
51
- "outputs": ["existing_revised.docx"],
52
- "skill": "word",
53
- "script": "scripts/replace_paragraphs.py",
54
- "scriptArgs": [
55
- "existing.docx", "existing_revised.docx",
56
- "3", "Dogs have about 1,700 taste buds.",
57
- "4", "A dog's nose print is unique, like a fingerprint."
58
- ]
59
- }
60
- ```
61
-
62
- Wrong, for this job — a `command` instead of a `script`:
63
-
64
- ```json
65
- { "command": "from docx import Document\ndoc = Document(\"existing.docx\")\nfor i in range(1, 11): ..." }
66
- ```
67
-
68
- `exitCode 0` plus an attachment = done: reply with the file name and the
69
- printed `K of N paragraphs replaced`, and do not call `exec` again. A script
70
- error names the fix (a heading index, an index out of range, unchanged text,
71
- missing pairs, a missing `outputs` for the revised name); correct the
72
- arguments and rerun the **same script**.
73
-
74
- Everything below is for the other edits: extending a document, trimming it,
75
- rewriting a whole section, resizing its text, embedding an image.
76
-
77
- ## Staging the Document
78
-
79
- Stage the document as an input **by its `attachmentId`** and open it with
80
- `Document("existing.docx")`. The id comes from wherever the document entered
81
- the chat:
82
-
83
- - **Produced earlier in this chat** — the `attachmentId` is in that `exec` result.
84
- - **Uploaded by the user** — the `[Attached file …]` line on their message names
85
- it, when the message carries one:
86
-
87
- ```
88
- [Attached file "report.docx" (application/vnd.openxmlformats-officedocument.wordprocessingml.document) — attachmentId: 4f9c2ab1]
89
- ```
90
-
91
- Copy the id verbatim — never placeholders like `att_doc`, `att_image`, or any
92
- id you made up. A `.docx` is **never** staged id-less: an id-less input
93
- resolves to an uploaded _image_, so it can never reach a document. If no
94
- `attachmentId` for the document appears anywhere in the chat, say you cannot
95
- open that file for editing and ask the user to attach it again — do not invent
96
- an id, do not stage it id-less, and do not retry. An attachment from an
97
- earlier turn can be used when its attachment id is available in the
98
- conversation.
99
-
100
- **The file only exists if this same `exec` call stages it.** The working
101
- directory starts empty on every call, so an edit needs an `inputs` entry
102
- naming the attachment, and `Document("existing.docx")` must use that entry's
103
- exact `path`. Opening a name that was never staged raises
104
- `PackageNotFoundError: Package not found at '…'` — the fix is the missing
105
- `inputs`, never a different file name.
106
-
107
- ## The exec call
108
-
109
- Images can be staged alongside the document. Tool-produced images
110
- (`generate_image` output) take the exact `attachmentId` from the tool result;
111
- an image the user uploaded is staged with `path` only and **no
112
- `attachmentId` key** — the first id-less entry is the first image of the
113
- user's latest message, and so on. Id-less entries resolve _images only_.
114
-
115
- ```json
116
- {
117
- "language": "python",
118
- "packages": ["python-docx==1.2.0"],
119
- "inputs": [
120
- {
121
- "attachmentId": "<id from the exec result or the [Attached file …] line>",
122
- "path": "existing.docx"
123
- },
124
- { "path": "photo.png" }
125
- ],
126
- "outputs": ["existing_revised.docx"],
127
- "command": "..."
128
- }
129
- ```
130
-
131
- - `packages` — `["python-docx==1.2.0"]` on every call. The PyPI package is
132
- `python-docx` but the import is `docx`; never list `docx` as the package —
133
- that resolves a different, abandoned library. Pin the version; this exact
134
- version ships with the app and installs with no network; any other version
135
- has to be downloaded, which fails on a device that is offline.
136
- - `inputs` — staged files land in the working directory under the bare `path`
137
- names — reference `Document("existing.docx")` /
138
- `doc.add_picture("photo.png", …)` by that name only. Paths must be unique
139
- bare filenames.
140
- - `outputs` — the new file to deliver; a file you write but do not declare
141
- here is discarded. Never the staged input's name. **Name it after the
142
- document you edited**, not after the change: keep the input's stem and add a
143
- marker — `report.docx` edited is `report_revised.docx`. A fresh name picked
144
- from the new content (`cats.docx` for an edit of `parrot_facts.docx`) reads
145
- as a second, unrelated document and hides the fact that an edit happened at
146
- all.
147
- - `command` — the multi-line Python source, with real newline characters.
148
- Never collapse it to one line joined by `;` — a `for`/`if`/`with` after a
149
- semicolon is a `SyntaxError`. There is no shell and no installer — packages
150
- are declared in `packages`.
151
-
152
- If the image was staged in `inputs`, embed it in **that** single build with
153
- `doc.add_picture("photo.png", width=Inches(5.5))` — never deliver a document
154
- and then rebuild to add the image. **Do not soft-fail images or imports**:
155
- never wrap `add_picture` or an import in `try`/`except` that prints a warning
156
- and continues — a document saved without the requested image is a failed
157
- turn, not a success. Pass only one of `width`/`height`; passing both distorts
158
- the picture.
159
-
160
- **Image URLs do not work — never download.** Your Python code has **no
161
- network access**: `requests`, `urllib`, and `socket` all fail with a network
162
- error, and `http_request` returns truncated text, never image bytes. Say the
163
- link cannot be downloaded and ask the user to attach the image itself, or
164
- offer `generate_image` for a similar visual.
165
-
166
- ## Editing: Work the Objects, Save a New Name
167
-
168
- One call does the whole edit: open, change, verify the document actually
169
- changed, save. Keep the fingerprint lines exactly as written — they are what
170
- stops an edit that silently matched nothing (or a read that only inspected)
171
- from delivering an unchanged copy of the user's document at `exitCode 0`. The
172
- fingerprint covers the body **and** the styles part, so a style-only change —
173
- the resize recipe below — counts as a change too:
174
-
175
- ```python
176
- import hashlib
177
- from docx import Document
178
-
179
- doc = Document("existing.docx")
180
- fingerprint = hashlib.md5((doc.element.xml + doc.styles.element.xml).encode()).hexdigest()
181
-
182
- for paragraph in doc.paragraphs:
183
- if paragraph.text == "Prepared by the finance team.":
184
- paragraph.text = "Prepared by the finance team. Revised after board review."
185
-
186
- table = doc.tables[0]
187
- row = table.add_row().cells
188
- row[0].text = "Margin"
189
- row[1].text = "25%"
190
- row[2].text = "36%"
191
-
192
- doc.add_heading("Appendix", level=1)
193
- doc.add_paragraph("Margins recovered as one-off costs rolled out of the base.")
194
-
195
- assert (
196
- hashlib.md5((doc.element.xml + doc.styles.element.xml).encode()).hexdigest() != fingerprint
197
- ), "nothing changed — the edit matched nothing or never ran; fix it, never deliver an unchanged copy"
198
- doc.save("existing_revised.docx") # a NEW name — never the staged input
199
- print(f"{len(doc.paragraphs)} paragraphs, {len(doc.tables)} table(s)")
200
- ```
201
-
202
- **An assert that fires is a failed turn to diagnose, not a document to
203
- deliver**: the usual cause is a paragraph match on text that is not exactly
204
- there — print the real `.text` values in the rerun, fix the match, and never
205
- delete the assert to get a file out.
206
-
207
- **Keep it an `assert`, never a `print` or an `if`.** Two `print` lines showing
208
- the old and new hashes let a no-op save and deliver anyway, which is the one
209
- thing the assert exists to stop. They also invite a second miscoding: taking
210
- both hashes together, before the change. Then they match whatever the edit did,
211
- and the run reports "nothing changed" over a document that changed correctly.
212
- Take the second hash after the last mutation and before `save`, and let the
213
- assert raise.
214
-
215
- **A printed line is never a reason to call `exec` again.** The guard is the
216
- assert: if it did not fire and the result carries an attachment, the document is
217
- delivered and the turn is over, whatever stdout says about it. Re-running to
218
- check saves the same edit under a second name, and the user gets two documents
219
- for one request.
220
-
221
- `doc.paragraphs` walks only the document body — text inside tables, headers, and
222
- footers is **not** in it. Table text is reached through `doc.tables`; match
223
- paragraphs by their exact `.text` before rewriting them, and remember the
224
- formatting-loss rule below.
225
-
226
- The `add_heading`/`add_paragraph` pair above appends an **Appendix** because
227
- that is what the sample edit asks for. Copy that shape only when the user
228
- genuinely wants new content at the end. Substituting content that is already
229
- in the document is a different job with its own guards: some of the points —
230
- "change the first five points" — is the two script calls at the top of this
231
- file; a whole section — "rewrite section 2" — is Rewriting Whole Sections.
232
-
233
- When an edit adds substantial new content — new sections, formatted runs,
234
- bulleted lists, whole tables — the writing rules apply unchanged: load
235
- `references/create.md` too and copy its shapes (no markdown characters in
236
- strings, bold/italic on runs never paragraphs, one string per `add_paragraph`,
237
- tables built from a header row with `rows=1`).
238
-
239
- ### Setting `paragraph.text` erases formatting
240
-
241
- Assigning `paragraph.text = "…"` replaces **all** runs with one plain run: every
242
- bold, italic, size, and color in that paragraph is gone. Fine for plain
243
- paragraphs; on a formatted paragraph edit the runs instead, or accept the loss
244
- deliberately. This is the top footgun when editing an uploaded document.
245
-
246
- ### Styles must exist in the document
247
-
248
- `style="List Bullet"` names a style **inside the document**. A missing name
249
- raises `KeyError: "no style with name 'List Bullet'"` at `add_paragraph` time.
250
- An **uploaded** document carries only its own styles — one written by another
251
- tool may lack even `List Bullet`. When editing, guard once and fall back:
252
-
253
- ```python
254
- names = [s.name for s in doc.styles]
255
- bullet = "List Bullet" if "List Bullet" in names else None
256
- doc.add_paragraph("point one", style=bullet) # style=None → Normal
257
- ```
258
-
259
- ## Rewriting Whole Sections
260
-
261
- This recipe is for whole _sections_ (a heading plus its body). Changing some
262
- of the facts, points, bullets, or paragraphs is the two script calls at the
263
- top of this file — never hand-write a loop for that.
264
-
265
- `add_paragraph`, `add_heading`, and `add_picture` **always append at the end of
266
- the document.** None of them takes a position. "Change the first five points",
267
- "rewrite section 2", "swap these facts for those" are all *replacements*, and
268
- reaching for `add_*` silently turns them into an append: the original content
269
- stays where it is, the new content lands after the closing line, and the
270
- document comes back longer than it started with both versions in it. That is a
271
- failed turn, not a partial success — it is the most common way this skill goes
272
- wrong.
273
-
274
- **A section is a heading plus everything under it, up to the next heading of
275
- the same or higher rank** — which may be one paragraph, or six bullets, or a
276
- whole subsection, or nothing at all. Never assume it is exactly one paragraph:
277
- rewriting the heading and the single paragraph after it leaves the rest of the
278
- old section sitting under its new title, which is the same contradiction an
279
- append produces and is just as invisible in the result. Work out where each
280
- section ends before changing anything.
281
-
282
- Rank matters as much as position. `Heading 2` under a `Heading 1` is a
283
- subsection, not the next section, so "replace the first two sections" on a
284
- document with subheadings must not consume the parent's own subheading as
285
- section two — the same rule the removal recipe below follows. `rank()` reads
286
- the level off the style name, and only the shallowest rank counts as a section
287
- start.
288
-
289
- **A heading shallower than every other heading is the document's title, not its
290
- first section.** A document headed `Heading 1` and sectioned `Heading 2` — the
291
- shape most attached documents have — would otherwise have exactly one
292
- "section": the title, spanning everything under it. Replacing that section
293
- replaces the entire document, and nothing about the result says so. The `if`
294
- drops such a heading before sections are picked, and the whole-document assert
295
- refuses the span even if one is somehow selected.
296
-
297
- **One heading is dropped, never a chain of them.** It is an `if`, not a
298
- `while`: a document is titled once. Stripping repeatedly walks down the
299
- outline — on a `Heading 1` title over a `Heading 2` phase holding `Heading 3`
300
- weeks it drops the title, then the phase, and the weeks become the "sections",
301
- so replacing the first two rewrites the weeks and leaves the phase untouched.
302
- That is the silent wrong target this section exists to prevent. Stopping after
303
- one leaves the phase as the only section, and asking for a second raises an
304
- error that says so.
305
-
306
- Take one snapshot of `doc.paragraphs` and index into it. **Every string in
307
- `NEW` is a placeholder** — the sample fills it with report sections so the
308
- shape is readable, and you replace all of it with the content this request asks
309
- for. Shipping a sample string in the user's document is a failed turn:
310
-
311
- ```python
312
- from docx import Document
313
-
314
- doc = Document("existing.docx")
315
- paras = doc.paragraphs # one snapshot — index into THIS list
316
- blocks = list(doc.element.body) # paragraphs AND tables, in document order
317
-
318
- NEW = [ # placeholders — you write every string here
319
- ("Regional Performance", "Revenue grew in every region except EMEA, where the quarter closed flat."),
320
- ("Cost Base", "Headcount costs fell as the contractor pool wound down, and the saving held."),
321
- ]
322
- TARGET = range(len(NEW)) # which sections to replace — here the first len(NEW)
323
-
324
- def rank(paragraph): # "Heading 2" -> 2; a bare "Heading" is rank 1
325
- tail = paragraph.style.name.split()[-1]
326
- return int(tail) if tail.isdigit() else 1
327
-
328
- heads = [i for i, p in enumerate(paras) if p.style.name.startswith("Heading")]
329
- if len(heads) > 1 and all(rank(paras[heads[0]]) < rank(paras[i]) for i in heads[1:]):
330
- heads = heads[1:] # a lone heading above all the rest is the title
331
- top = min((rank(paras[i]) for i in heads), default=1)
332
- starts = [i for i in heads if rank(paras[i]) == top] # sections, never their subsections
333
- ends = [next((j for j in heads if j > i and rank(paras[j]) <= top), len(paras)) for i in starts]
334
-
335
- assert NEW, "NEW is empty — write the replacement content before running the edit"
336
- assert len(TARGET) == len(NEW), f"TARGET names {len(TARGET)} sections but NEW has {len(NEW)} items"
337
- assert len(starts) > max(TARGET), f"TARGET reaches section {max(TARGET) + 1}, but the document has {len(starts)}"
338
- at = [blocks.index(p._element) for p in paras] # where each paragraph sits among the blocks
339
- for k in TARGET:
340
- assert ends[k] > starts[k] + 1, f"section {paras[starts[k]].text!r} has no body paragraph to replace"
341
- assert (starts[k], ends[k]) != (heads[0], len(paras)), f"section {paras[starts[k]].text!r} spans the whole document — that is a rewrite, not a section replacement"
342
- span = blocks[at[starts[k]] : at[ends[k]] if ends[k] < len(paras) else len(blocks)]
343
- assert not any(el.tag.endswith("}tbl") for el in span), f"section {paras[starts[k]].text!r} holds a table — this recipe replaces paragraphs only"
344
-
345
- # measured from the document, before anything changes — never from what the loop below does
346
- before = len(paras)
347
- old_body = [p.text for k in TARGET for p in paras[starts[k] + 1 : ends[k]]]
348
- doomed = [(p.text, p._element) for k in TARGET for p in paras[starts[k] + 2 : ends[k]]]
349
- expected = before - len(old_body) + len(NEW) # each replaced section keeps exactly one body paragraph
350
-
351
- for k, (title, body) in zip(TARGET, NEW):
352
- paras[starts[k]].text = title # the heading keeps its own style
353
- paras[starts[k] + 1].text = body
354
- paras[starts[k] + 1].style = doc.styles["Normal"] # the reused paragraph may have been a bullet
355
- for p in paras[starts[k] + 2 : ends[k]]: # whatever else the section held
356
- p._element.getparent().remove(p._element)
357
-
358
- assert len(doc.paragraphs) == expected, f"expected {expected} paragraphs, got {len(doc.paragraphs)} — an old section was not fully replaced, or content was appended"
359
- for text, el in doomed:
360
- assert el.getparent() is None, f"an old paragraph is still in the document: {text[:40]!r}"
361
- doc.save("existing_revised.docx") # a NEW name — never the staged input
362
- print(f"{len(NEW)} of {len(starts)} sections replaced, {before} -> {len(doc.paragraphs)} paragraphs")
363
- ```
364
-
365
- `TARGET` names the sections to replace, once, and `zip` pairs each new item with
366
- the section it overwrites. Replacing a different range is a change to that one
367
- line — `TARGET = range(2, 5)` for "sections 3 through 5", with three items in
368
- `NEW` to match. Keep it bound in a single place: a range written twice drifts
369
- apart the moment one copy is edited, and every guard below reads `TARGET`
370
- rather than assuming the range starts at zero.
371
-
372
- Assigning `paras[head].text` keeps that paragraph's style, because the style
373
- lives on the paragraph and not on its runs: a `Heading 2` stays a `Heading 2`.
374
- Only the run-level formatting inside it is lost, per the rule above. The body
375
- paragraph is the opposite case — it is reused, so it arrives carrying whatever
376
- style the old body had, which is why the sample sets it back to `Normal`. Set it
377
- to something else when the new body should be a bullet or a quote, and guard the
378
- name as shown under Styles.
379
-
380
- **Keep the document's own numbering.** The sample titles carry no `1.`, `2.`
381
- prefix because the document it edits does not number itself, and a typed prefix
382
- on a `List Number` paragraph double-numbers. When the headings you are
383
- overwriting *do* carry manual numbers, take each number from the position being
384
- overwritten so the sequence continues — replacing sections 3 through 5 writes
385
- `3.`, `4.`, `5.`, never restarting at `1.`
386
-
387
- **Every assert, exactly as written — and measured before the loop runs.**
388
- `old_body` and `expected` come from the document's own structure, never from
389
- what the loop reports about itself. That is the whole point: a loop that
390
- rewrites only the paragraph after each heading, the mistake this recipe exists
391
- to prevent, would tally its own work as complete. Derived up front, the numbers
392
- contradict it. None of these failures is distinguishable from success by
393
- `exitCode 0` plus an attachment:
394
-
395
- - **`assert NEW`** catches an empty content list. Without it `max(TARGET)`
396
- raises a bare `ValueError`, and were it not for that the run would save an
397
- untouched copy of the user's document at `exitCode 0`.
398
- - **`len(TARGET) == len(NEW)`** catches a target range and a content list that
399
- drifted apart. `zip` would silently pair only the shorter of the two.
400
- - **`len(starts) > max(TARGET)`** catches a range reaching past the last
401
- section. It reads `TARGET`, not `len(NEW)`, because the range need not start
402
- at zero — a `len(NEW)` check passes on `range(2, 5)` over four sections and
403
- the run then dies on an `IndexError` that names nothing.
404
- - **the whole-document assert** refuses a section running from the first
405
- heading to the last paragraph. That is not a replacement, it is a rewrite:
406
- every other check passes while the document is emptied down to one heading
407
- and one paragraph. It anchors on `heads[0]`, not paragraph 0 — a document
408
- whose only heading sits under a draft notice, a date, or a byline still has
409
- exactly one section, and anchoring on index 0 would wave it through.
410
- - **`ends[k] > starts[k] + 1`** catches a section with no body paragraph — a
411
- heading followed straight by a table, or the last heading in the document.
412
- There is nothing under it to rewrite. It runs before any mutation, so a bad
413
- target changes nothing.
414
- - **the `}tbl` assert** catches a table inside a section being replaced.
415
- `doc.paragraphs` does not see tables, so the loop below cannot remove one:
416
- without this the old table survives under the new heading with every other
417
- check passing. Say the table has to be rebuilt, or target a different section.
418
- - **the `expected` assert** catches an old section left partly in place *and*
419
- new content appended, because `expected` is what the paragraph count must be
420
- once each replaced section holds exactly one body paragraph.
421
- - **the `getparent() is None` assert** catches an old paragraph the loop was
422
- supposed to drop but left attached. Compare **elements, not text**: text
423
- comparison cannot tell a paragraph that survived from an identical one
424
- standing legitimately elsewhere, and a document that repeats a line — three
425
- status sections each reading `Nothing to report.` — would fail a correct edit
426
- with no way to satisfy the assert. Identity has no such collision, and it
427
- needs no special case for blank paragraphs.
428
-
429
- An assert that fires is a failed turn to diagnose, never a document to deliver.
430
-
431
- ### When the replacement needs more than one paragraph
432
-
433
- The loop above reuses one paragraph per section and drops the rest. When a
434
- replacement needs an **extra** paragraph, insert it before the paragraph that
435
- should follow it. `insert_paragraph_before` is the only insert there is, and it
436
- is a method on the paragraph you want to push down:
437
-
438
- ```python
439
- anchor = paras[ends[k]] # the next section's heading
440
- extra = anchor.insert_paragraph_before("A second body paragraph.", style="Normal")
441
- ```
442
-
443
- It takes the same style names as `add_paragraph` (`"Heading 2"`, `"List
444
- Bullet"`, `None` for Normal) and returns the new paragraph, so runs can be
445
- formatted on it. Inserting a whole new section is this call once per paragraph,
446
- each against the heading it goes above. A section at the very end of the
447
- document has no next heading to anchor to — `ends[k]` is `len(paras)` — so
448
- append there with `doc.add_paragraph`, the one case where appending is right.
449
-
450
- Inserting does not disturb the `paras` snapshot: it is a plain Python list
451
- holding the paragraphs that already existed, so every index taken before the
452
- insert still points at the same paragraph afterwards. Only a fresh
453
- `doc.paragraphs` shifts.
454
-
455
- Count what you insert and fold it into `expected` rather than dropping the
456
- guard — `expected = before - len(old_body) + len(NEW) + added` — so an
457
- accidental append is still caught.
458
-
459
- ## Removing Content
460
-
461
- python-docx has **no delete API.** There is no `doc.remove_paragraph` and no
462
- `paragraph.delete`, and `doc.paragraphs` is rebuilt on every access, so
463
- `doc.paragraphs.remove(p)` edits a throwaway list and changes nothing in the file.
464
- Removing anything means dropping its XML element from the parent — this one line
465
- is the whole technique, and there is no alternative to it:
466
-
467
- ```python
468
- p._element.getparent().remove(p._element)
469
- ```
470
-
471
- Code that finds the paragraphs and never runs that line — a `for`/`if` that
472
- matches the text and falls through, or a comment like
473
- `# Find and remove paragraphs containing "Conclusion"` standing in for the
474
- removal — saves a document byte-identical to the input at `exitCode 0`, with an
475
- attachment that looks like a success. Nothing in the result says the edit was a
476
- no-op, which is why the sample below asserts the count changed before it saves.
477
-
478
- Because `doc.paragraphs` is a fresh list each time, `for p in doc.paragraphs:`
479
- walks a snapshot and removing inside the loop is safe.
480
-
481
- **A whole section** — a heading plus everything under it, up to the next heading
482
- of the same or higher rank — is that line plus a flag. Track the heading's level,
483
- or a sub-heading inside the section ends the removal early and orphans the
484
- paragraphs below it:
485
-
486
- ```python
487
- from docx import Document
488
-
489
- doc = Document("existing.docx")
490
-
491
- TARGET = "Conclusion" # the heading text that opens the section
492
-
493
- before = len(doc.paragraphs)
494
- depth = None # the target heading's level while removing
495
- for p in doc.paragraphs:
496
- style = p.style.name # a style object — compare through .name
497
- if style.startswith("Heading"):
498
- tail = style.split()[-1]
499
- level = int(tail) if tail.isdigit() else 1 # "Heading 2" -> 2
500
- if depth is not None and level <= depth:
501
- depth = None # a sibling heading closes the section
502
- if p.text.strip() == TARGET:
503
- depth = level
504
- if depth is not None:
505
- p._element.getparent().remove(p._element)
506
-
507
- after = len(doc.paragraphs)
508
- assert after < before, f"removed nothing ({before} -> {after}) — the match never fired"
509
- doc.save("existing_revised.docx") # a NEW name — never the staged input
510
- print(f"{before} -> {after} paragraphs")
511
- ```
512
-
513
- Find the heading through `p.style.name`, never the text alone — a body paragraph
514
- that mentions "Conclusion" is not the section heading. Removing individual
515
- paragraphs is the same loop without the flag: match them, and call the removal
516
- line on each one.
517
-
518
- **Assert the count changed, before you save.** Keep the
519
- `assert after < before` line exactly where the sample puts it — between the loop
520
- and `doc.save(...)` — and do not soften it to a `print` or an `if`. It is what
521
- makes a no-op impossible to deliver: the assert raises, `save` never runs, so no
522
- file is written and the result comes back with `missingOutputs` instead of an
523
- attachment. Without it a removal that never fired still saves the unchanged
524
- document, and the run is indistinguishable from a real edit — `exitCode 0`, an
525
- attachment, and nothing anywhere saying the document is a copy of the input.
526
-
527
- An assert that fires is a **failed turn to diagnose**, never a result to report.
528
- It means the match did not fire: wrong heading text, a heading style the document
529
- does not use, or text living in a table, header, or footer, which
530
- `doc.paragraphs` never walks. Fix the match and rerun — do not delete the assert
531
- to get a file out.
532
-
533
- The printed `before -> after` line is then just the reply line (`16 -> 12
534
- paragraphs`), not the check. Both live inside the build, so this takes no extra
535
- call: the assert and the `print` are in the same `exec` that does the removal.
536
-
537
- This assert is also why "Success = stop" needs no second call on a destructive
538
- edit: `exitCode: 0` plus an attachment cannot on its own tell a real edit from
539
- a copy of the input, because a removal that never fired produces both. The
540
- build itself closes that gap — it asserts the count changed before `save`, so
541
- a no-op returns `missingOutputs` rather than a convincing attachment. A
542
- delivered document is still never reopened to "verify" it; the fix for a
543
- failed assert is a corrected build, never an `exec` opened to inspect what was
544
- already delivered.
545
-
546
- **Table rows** have no delete API either, and take the same idiom on the row's own
547
- element. `table.rows` iterates a snapshot just as `doc.paragraphs` does, so
548
- removing inside the loop is safe — and the count gets the same assert, because a
549
- row matched by its cell text can miss exactly the way a paragraph can:
550
-
551
- ```python
552
- table = doc.tables[0]
553
-
554
- before = len(table.rows)
555
- for row in table.rows:
556
- if row.cells[0].text == "Discontinued":
557
- row._element.getparent().remove(row._element)
558
- assert len(table.rows) < before, f"no row matched ({before} rows unchanged)"
559
- ```
560
-
561
- Removing a row by position needs no assert — `table.rows[9]` on a 4-row table
562
- raises `IndexError` rather than quietly doing nothing:
563
-
564
- ```python
565
- row = table.rows[2]
566
- row._element.getparent().remove(row._element)
567
- ```
568
-
569
- Rows only grow at the bottom: there is no insert-at, and no delete either
570
- outside this idiom. Address cells as `table.cell(row, col)` or
571
- `table.rows[r].cells[c]` — they are the same cell; for formatting inside a
572
- cell go through `cell.paragraphs[0]` and its runs like any other paragraph.
573
-
574
- **Table columns cannot be removed.** A column is not one element — it is an entry
575
- in the table grid plus one cell in every row — and a horizontally merged cell is a
576
- single `<w:tc>` shared across two grid positions, so removing "the second cell" of
577
- every row deletes that merged cell whole and leaves its row a column short. The
578
- document opens visibly ragged and nothing raises. Rebuild the table with the
579
- columns you want instead, or say the column has to be dropped in Word.
580
-
581
- ## Reading What Is Already in the Document
582
-
583
- A document exposes exactly two collections — `doc.paragraphs` and `doc.tables`.
584
- Everything else is derived by filtering them; there is no `doc.headings`, no
585
- `doc.sections_by_title`, no `doc.text`. A `Paragraph` has `.text`, `.style` and
586
- `.runs`, and no `.paragraphs` of its own.
587
-
588
- **`paragraph.style` is a style object, not a string** — compare through
589
- `.name`, or you get
590
- `AttributeError: 'ParagraphStyle' object has no attribute 'startswith'`:
591
-
592
- ```python
593
- headings = [p for p in doc.paragraphs if p.style.name.startswith("Heading")]
594
- body = [p for p in doc.paragraphs if p.style.name == "Normal"]
595
- ```
596
-
597
- Most edits need no inspection at all — go straight to the change. When a look
598
- is genuinely needed first (an exact `.text` to match, a style name), that call
599
- only prints: **an inspection never saves and declares no `outputs`** — a save
600
- without the change delivers a stale copy of the user's document. The new file
601
- comes only from the one call that changes it. Reading for the _user_ — a
602
- summary or an answer delivered as chat text — is its own flow with its own
603
- call shape: load `references/read.md`.
604
-
605
- ## Resizing Text: Set the Styles, Never Scale `run.font.size`
606
-
607
- `run.font.size` is `None` whenever the size comes from the paragraph's style,
608
- which is the normal case for a document you did not hand-size. **`None` does not
609
- mean zero.** Reading it as a number and scaling it writes a 0pt font, and 0pt
610
- text is invisible in Word and Pages — the document opens looking blank, with no
611
- error anywhere to tell you why:
612
-
613
- ```python
614
- size = run.font.size.pt if run.font.size else 0
615
- run.font.size = Pt(size * 1.5) # WRONG: 0 * 1.5 = 0pt, invisible text
616
- ```
617
-
618
- "Make the font bigger" is a change to the **styles**, because every run without
619
- its own size inherits from them. Set absolute point sizes on the styles the
620
- document actually uses, and the whole document — body, tables, headers — follows
621
- in four lines:
622
-
623
- ```python
624
- from docx import Document
625
- from docx.shared import Pt
626
-
627
- doc = Document("existing.docx")
628
-
629
- doc.styles["Normal"].font.size = Pt(14) # body text; 11pt is the default
630
- doc.styles["List Bullet"].font.size = Pt(14)
631
- doc.styles["Heading 1"].font.size = Pt(20)
632
- doc.styles["Title"].font.size = Pt(32)
633
-
634
- doc.save("larger.docx")
635
- print(f"{len(doc.paragraphs)} paragraphs resized")
636
- ```
637
-
638
- **Keep the hierarchy.** Raise every style you touch, not one size for all of
639
- them — a title and a heading set to the body size read as unstyled text. Body
640
- around 14pt pairs with roughly 20pt headings and a 32pt title, and the same
641
- ratios hold at any size the user asks for.
642
-
643
- Only touch a style the document has — guard with the `doc.styles` check above
644
- when unsure. If a specific run really must be sized on its own, assign an
645
- absolute `Pt(...)` value; never one derived from the size you read back.
646
-
647
- ## Errors
648
-
649
- - `PackageNotFoundError: Package not found at '…'` — the document was never
650
- staged, or an id-less entry staged an image under a `.docx` path. Add
651
- `inputs: [{ "attachmentId": "<id from the exec result or the [Attached file …]
652
- line>", "path": "existing.docx" }]` and open that exact path. If no id is
653
- available, ask the user to attach the file again rather than guessing a name.
654
- - `attachment … not found in this chat` means `inputs` listed an id that is not
655
- in this chat (often a copied placeholder like `att_doc`). Re-copy the exact id
656
- from the `exec` result or the `[Attached file …]` line that names the file;
657
- if no id appears anywhere in the chat, ask the user to re-attach.
658
- - `ModuleNotFoundError: No module named 'docx'` means `packages` was missing or
659
- wrong — add `["python-docx==1.2.0"]` and rerun. Never try to install it, and
660
- never "fix" it by importing `python_docx`; the import stays `docx`.
661
- - `KeyError: "no style with name '…'"` — the style is not in this document.
662
- Check `doc.styles` and fall back as shown under Styles.
663
- - `NameError: name 'Pt' is not defined` (or `Inches`, `RGBColor`) — the import
664
- line is missing that name; they all import from `docx.shared`.
665
- - `AssertionError: removed nothing (16 -> 16)` means the removal matched nothing:
666
- either the loop never fired or it never called
667
- `p._element.getparent().remove(p._element)`; python-docx has no delete method to
668
- reach for instead. Diagnose the match and rerun — never delete the assert to get
669
- a file out, since the file it would produce is a copy of the input.
670
- - `AssertionError: expected 7 paragraphs, got 9` means the document did not end
671
- up the shape a replacement makes. Two causes: the old sections were only
672
- partly replaced — the loop rewrote the paragraph after each heading and left
673
- the rest of the section standing — or new content was appended with
674
- `add_paragraph`/`add_heading`, which only ever append. `expected` is derived
675
- from the section bounds before anything changes, so it is right and the
676
- document is wrong: replace each section through to the next heading.
677
- - `AssertionError: an old paragraph is still in the document: '…'` means the
678
- loop was adapted and no longer drops everything past the paragraph it reuses.
679
- Every paragraph from `starts[k] + 2` to `ends[k]` has to go; the removal idiom
680
- below is the only thing that removes one. This compares elements, so it never
681
- fires because the document happens to repeat a line elsewhere.
682
- - `AssertionError: section '…' has no body paragraph to replace` means that
683
- heading is followed straight by a table, or is the last paragraph in the
684
- document. There is nothing under it to rewrite: target a different section, or
685
- insert the body with `insert_paragraph_before` before adding to it.
686
- - `IndexError: list index out of range` while walking sections means an index
687
- ran past the end of `starts` or of `paras`: a `TARGET` reaching past the last
688
- section, or `paras[i + 1]` on a document whose final paragraph is a heading.
689
- Guard the range with `len(starts) > max(TARGET)` — not against `len(NEW)`,
690
- which says nothing when the range does not start at zero — and take section
691
- ends from the next heading of the same or higher rank, with `len(paras)`
692
- closing the last one.
693
- - `AssertionError: section '…' holds a table` means the section being replaced
694
- contains a table. `doc.paragraphs` never sees tables, so the loop cannot
695
- remove one and it would survive under the new heading. Rebuild the table
696
- explicitly, or tell the user that section has to be replaced by hand.
697
- - `AssertionError: TARGET names 2 sections but NEW has 3 items` means the range
698
- and the content list drifted apart. Fix whichever is wrong; do not let `zip`
699
- quietly use the shorter.
700
- - `AssertionError: TARGET reaches section 2, but the document has 1` on a
701
- document that plainly has several usually means its sections are `Heading 2`
702
- under a `Heading 1` title. The title is dropped before sections are picked,
703
- so check `rank()` is reading the style names this document actually uses —
704
- print `[p.style.name for p in doc.paragraphs]` — rather than lowering
705
- `TARGET` until the assert passes. Section 1 of a title-only document is the
706
- whole document.
707
- - `AssertionError: section '…' spans the whole document` means the heading
708
- selected covers every paragraph, so replacing it would empty the document.
709
- It is a title being treated as a section, or a request to rewrite rather than
710
- edit — build a new document with `references/create.md` if that is what the
711
- user wants.
712
- - `AttributeError: 'Document' object has no attribute 'insert_paragraph'` means
713
- the code guessed an insert API on the document. There is none. The only insert
714
- is `paragraph.insert_paragraph_before(text, style)`, on the paragraph the new
715
- one goes above.
716
- - `TypeError: Document.add_paragraph() takes from 1 to 3 positional arguments
717
- but 4 were given` means a position was passed to `add_paragraph`. It has no
718
- position parameter and always appends; use `insert_paragraph_before`.
719
- - Identical old and new fingerprints on a run that saved anyway means the
720
- assert was softened into `print` lines and both hashes were taken before the
721
- change. It is not evidence the edit failed, and it is not grounds for another
722
- `exec`: restore the assert and take the second hash after the mutation.
723
- - A delivered document identical to the one you opened means an edit ran
724
- without the fingerprint assert — an edit that matched nothing, or an
725
- inspection that saved. Add the assert before `save` and rerun the actual
726
- change.
727
- - `AssertionError: nothing changed — the edit matched nothing or never ran`
728
- means exactly that: the paragraph match found no text, or no mutation
729
- happened before `save`. Print the real `.text` values, fix the match, rerun
730
- — never remove the assert.
731
- - `TypeError: 'Table' object is not subscriptable` — a table was indexed
732
- directly (`table[0]`). Cells are reached through `table.rows[r].cells[c]` or
733
- `table.cell(r, c)`; a whole row of cells is `table.add_row().cells`.
734
- - `AttributeError: 'Document' object has no attribute 'remove_paragraph'` (or
735
- `'Paragraph' object has no attribute 'delete'`) means the code guessed a delete
736
- API. There is none; drop the XML element instead.
737
- - A resize that "worked" but left the document blank means a 0pt font: something
738
- scaled `run.font.size` while it was `None`. Set absolute sizes on the styles
739
- instead — see Resizing Text.
740
- - `SyntaxError: invalid syntax` on a one-line `for`/`if` means the source was
741
- collapsed — restore multi-line newlines from the sample and rerun. Underscores
742
- in names (`add_paragraph`, not `addparagraph`) must stay. Do not switch to
743
- `python -c` or change the package pin.
744
- - On an `AttributeError` from python-docx the API name is wrong; on a `TypeError`
745
- about positional arguments the call passes the wrong number of them — usually
746
- several strings where one is allowed. Fix either against this file's examples,
747
- reading the line number in the traceback. Do not retry the same call, and do
748
- not switch to a shell.
749
- - If the result has `missingOutputs`, the file was never written. Read stderr
750
- first: an `AssertionError` there means a guard stopped the save on purpose
751
- and its message names what to fix — rerunning the same code fails the same
752
- way. Only when stderr is clean is this a naming problem: check the `save()`
753
- name matches the declared output and rerun once.
754
- - Never print the document's bytes or base64 — stdout is capped and the file
755
- travels through `outputs`. A build call prints exactly one line (e.g. `9
756
- paragraphs, 1 table(s)`).
757
- - Never pass an absolute path to `save()`.
758
-
759
- ## Finish
760
-
761
- When `exitCode` is `0` and `attachments` lists the `.docx`, the edit is done —
762
- the `exec` result carries
763
- `attachments: [{ attachmentId, fileName, byteLength }]` and the file is already
764
- attached to the chat for the user to open or save. Stop tool use and reply
765
- with a single line: file name + the count line from stdout. Exactly one
766
- successful `exec` per request; never reopen a delivered document to "verify"
767
- it.