@qvac/skills 0.1.11 → 0.1.12

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,400 +0,0 @@
1
- # Creating a Word Document (python-docx)
2
-
3
- Create a new `.docx` from scratch by running Python through the `exec` tool.
4
- A new document needs **no** `inputs` — do not invent attachment ids — unless
5
- it embeds an image (see Embedding Images). **Exactly one** `exec` call per
6
- user request when that call succeeds.
7
-
8
- **A document that already exists in this chat is never rebuilt here.** "Replace
9
- the first 10 facts", "reword this", "add a section" — any request that starts
10
- from an existing `.docx` is a change to that document: some of its paragraphs
11
- is `references/paragraphs.md`, anything else is `references/rework.md`, and
12
- either one stages the document by its `attachmentId`. Building a fresh
13
- document for such a request throws away everything the user already has.
14
-
15
- ## The exec call
16
-
17
- ```json
18
- {
19
- "language": "python",
20
- "packages": ["python-docx==1.2.0"],
21
- "outputs": ["report.docx"],
22
- "command": "..."
23
- }
24
- ```
25
-
26
- - `language` — always `"python"`.
27
- - `packages` — `["python-docx==1.2.0"]` on every call. The PyPI package is
28
- `python-docx` but the import is `docx`; never list `docx` as the package —
29
- that resolves a different, abandoned library. Pin the version; an unpinned
30
- install resolves a potentially different library version. This exact version
31
- ships with the app and installs with no network; any other version has to be
32
- downloaded, which fails on a device that is offline.
33
- - `outputs` — `["report.docx"]`. `save("report.docx")` must match the declared
34
- output name. A file you write but do not declare here is discarded. A `.doc`
35
- output name is rejected — name it `.docx`.
36
- - `command` — the multi-line Python source, with real newline characters.
37
- Never collapse it to one line joined by `;` — a `for`/`if`/`with` after a
38
- semicolon is a `SyntaxError`. Its first line is the first line of Python
39
- that runs: there is no shell and no interpreter to invoke, and no
40
- installer — packages are declared in `packages`.
41
-
42
- ## Embedding Images
43
-
44
- Two kinds of image input, told apart by where the file came from:
45
-
46
- **Tool-produced images** (`generate_image` output): stage them with the exact
47
- `attachmentId` from the tool result — never placeholders like `att_image` or
48
- any id you made up.
49
-
50
- **Images the user uploaded** ("use this photo"): there is no id to copy — an
51
- uploaded image never shows one. Stage it with `path` only and **no
52
- `attachmentId` key**; the first id-less entry is the first image of the user's
53
- latest message, the second is its second image, and so on. Id-less entries
54
- resolve _images only_.
55
-
56
- ```json
57
- {
58
- "language": "python",
59
- "packages": ["python-docx==1.2.0"],
60
- "inputs": [{ "path": "photo.png" }],
61
- "outputs": ["report.docx"],
62
- "command": "..."
63
- }
64
- ```
65
-
66
- Staged files land in the working directory under the bare `path` names —
67
- reference `doc.add_picture("photo.png", …)` by that name only. Paths must be
68
- unique bare filenames. `attachment … not found in this chat` means you
69
- invented an id or the file is not attached: re-copy the exact id from the tool
70
- result, or for a document with no image drop `inputs` entirely.
71
-
72
- If the image was staged in `inputs`, embed it in **that** single build with
73
- `doc.add_picture` — never deliver a document and then rebuild to add the
74
- image. Soft-failing (`try`/`except` around the picture) and saving without it
75
- is a failed turn, not a success.
76
-
77
- **Image URLs do not work — never download.** Your Python code has **no
78
- network access**: `requests`, `urllib`, and `socket` all fail with a network
79
- error, and `http_request` returns truncated text, never image bytes. When the
80
- user gives an image URL, do not try to fetch it from Python and do not retry
81
- through other tools — that is a dead end. Say the link cannot be downloaded
82
- and ask the user to attach the image itself, or offer `generate_image` for a
83
- similar visual. Then build the document with the staged attachment as above.
84
-
85
- ## Which Shape
86
-
87
- | The user asks for | Shape |
88
- | ---------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------- |
89
- | "30 fun facts about cats", "10 tips for …", "a list of …", any number of items or points | **List** — a title, then exactly N `List Bullet` paragraphs and nothing else: no intro sentence, no section headings, no numbers typed into the text |
90
- | a report, memo, letter, plan — anything with sections | **Report** — the recipe under The Recipe below |
91
-
92
- ### The list shape
93
-
94
- ```python
95
- from docx import Document
96
-
97
- doc = Document()
98
- doc.add_heading("30 Fun Facts About Cats", level=0)
99
- facts = [
100
- "Cats sleep for about 70 percent of their lives.",
101
- "A group of cats is called a clowder.",
102
- "A cat's nose print is unique, like a fingerprint.",
103
- ] # one plain string per item — write all N here
104
- for fact in facts:
105
- doc.add_paragraph(fact, style="List Bullet")
106
- doc.save("cat_facts.docx") # must match the declared output exactly
107
- print(f"{len(facts)} items, {len(doc.paragraphs)} paragraphs, {len(doc.tables)} table(s)")
108
- ```
109
-
110
- One string per item, as many as the user asked for. No headings between
111
- groups of items and no introductory sentence: each of those is a paragraph the
112
- user did not ask for, and a later "change the first 10 items" then lands on the
113
- wrong lines. The count line it prints is the reply — a delivered list is done,
114
- whatever the count says; never rebuild it to fix the number.
115
-
116
- ## The Recipe
117
-
118
- Start from this for a report. It is a complete, working document — a title, headings,
119
- paragraphs with bold and italic runs, a bulleted list, and a table — saved
120
- under the declared output name. Copy it and change the content; do not
121
- assemble a document from memory.
122
-
123
- **Keep the source multi-line.** A `for`/`if`/`with` after a semicolon is a
124
- `SyntaxError` — paste the block with real newlines, not `stmt; for x in y: …`.
125
-
126
- **Hold content in plain lists of strings, and walk them.** Every list of bullets
127
- is a flat `["…", "…"]`, and every table is a list of row lists. Do not reach for
128
- a dict, a tuple of mixed widths, or a nested comprehension to hold document
129
- content — those are where a `SyntaxError` or a
130
- `ValueError: too many values to unpack` comes from, and they buy nothing here.
131
-
132
- **Keep every underscore in API names.** `add_heading`, `add_paragraph`,
133
- `add_run`, `add_table`, `add_row`, `add_picture`, `add_page_break` — stripping
134
- them to `addheading` / `addparagraph` fails. Copy identifiers exactly as written
135
- below:
136
-
137
- ```python
138
- from docx import Document
139
- from docx.shared import Inches, Pt, RGBColor # one import line covers sizes, widths, colors
140
-
141
- doc = Document()
142
-
143
- doc.add_heading("Quarterly Report", level=0)
144
- doc.add_paragraph("Prepared by the finance team.")
145
-
146
- doc.add_heading("Summary", level=1)
147
- p = doc.add_paragraph("Revenue grew ")
148
- strong = p.add_run("18 percent")
149
- strong.bold = True
150
- p.add_run(" against a ")
151
- emphasis = p.add_run("flat")
152
- emphasis.italic = True
153
- p.add_run(" cost base.")
154
-
155
- doc.add_heading("Highlights", level=1)
156
- for point in [
157
- "New retail partners in two regions",
158
- "Churn down for the third quarter",
159
- "Support backlog cleared",
160
- ]:
161
- doc.add_paragraph(point, style="List Bullet")
162
-
163
- doc.add_heading("Key Figures", level=1)
164
- figures = [
165
- ["Metric", "Q3", "Q4"], # first list is the header row
166
- ["Revenue", "$1.2M", "$1.4M"],
167
- ["Costs", "$0.9M", "$0.9M"],
168
- ]
169
- table = doc.add_table(rows=1, cols=len(figures[0]))
170
- table.style = "Table Grid"
171
- for index, cells in enumerate(figures):
172
- row = table.rows[0].cells if index == 0 else table.add_row().cells
173
- for column, value in enumerate(cells):
174
- row[column].text = value
175
-
176
- doc.save("report.docx") # must match the declared output exactly
177
- print(f"{len(doc.paragraphs)} paragraphs, {len(doc.tables)} table(s)")
178
- ```
179
-
180
- ## Write a document, not markdown
181
-
182
- A `.docx` carries real styles, so the structure is the style — never the
183
- punctuation. Markdown written into text stays there verbatim and reads as a
184
- typo in the finished document:
185
-
186
- - **No markdown characters in any string.** `#`, `##`, `-`, `*`, `1.`, `**bold**`
187
- and backticks all render literally. `add_heading("Security", level=2)` — never
188
- `add_heading("- Security", level=2)` or `"## Security"`. A numbered list is
189
- `style="List Number"`, which numbers itself; a typed `"1. "` prefix double-numbers.
190
- - **No typed rules or line breaks.** A row of dashes or underscores as a section
191
- divider is just those characters on the page, and a leading `"\n"` is a blank
192
- line inside the paragraph. Headings already separate sections.
193
- - **Every section title is a heading.** A first section called "Introduction" or
194
- "Overview" goes through `add_heading(..., level=1)` like every other one; as a
195
- plain `add_paragraph` it renders as body text and the document looks unstructured.
196
- - **No blank paragraphs for spacing.** `add_paragraph("")` leaves a visible gap —
197
- the heading and body styles already carry their own space before and after.
198
- - **Bold is for a few words, not a sentence.** A fully bold paragraph reads as a
199
- formatting mistake; bold the term, then continue in a normal run.
200
-
201
- ## One paragraph, one string
202
-
203
- `add_paragraph` takes a single text string, optionally with `style=` — nothing
204
- else. Several sentences passed positionally raise
205
- `TypeError: Document.add_paragraph() takes from 1 to 3 positional arguments but 4
206
- were given`. Join them into one string, or open the paragraph with the first
207
- piece and add the rest as runs:
208
-
209
- ```python
210
- p = doc.add_paragraph("As of 2026, Bitcoin is widely held. ")
211
- p.add_run("Adoption keeps growing.")
212
- ```
213
-
214
- **The text you pass to `add_paragraph` is already the paragraph's first run.** A
215
- run added afterwards _appends_ — repeating any of those words writes them twice
216
- into the document (`"…finite supplyfinite supply"`). Each run carries the next
217
- words and only those, so give a mixed-format paragraph an empty start and add
218
- every piece as its own run:
219
-
220
- ```python
221
- p = doc.add_paragraph()
222
- p.add_run("Digital scarcity ")
223
- tail = p.add_run("and a finite supply")
224
- tail.italic = True
225
- ```
226
-
227
- ## Bold and italic live on runs, never on paragraphs
228
-
229
- `paragraph.bold = True` raises no error and changes **nothing** in the file — a
230
- paragraph has no bold; the assignment lands on the Python object and is silently
231
- discarded on save. Formatting belongs to runs:
232
-
233
- ```python
234
- p = doc.add_paragraph("normal, then ")
235
- strong = p.add_run("bold")
236
- strong.bold = True
237
- p.add_run(" and ")
238
- emphasis = p.add_run("italic")
239
- emphasis.italic = True
240
- ```
241
-
242
- Two rules make that shape the only one to write:
243
-
244
- - **`add_run` takes the text and nothing else.** `p.add_run("x", bold=True)`
245
- raises `TypeError: Paragraph.add_run() got an unexpected keyword argument
246
- 'bold'` — create the run, then set the attribute.
247
- - **Never chain an attribute onto the `add_run(...)` call.** Name the run on one
248
- line and format it on the next, as above. A run that needs no formatting is a
249
- bare `p.add_run("plain text")` and the line ends there — a trailing `.` left
250
- over from a half-written chain is `SyntaxError: invalid syntax`.
251
- - **Runs join with no gap between them.** The next run starts exactly where the
252
- last one ended, so the separating space belongs inside one of the strings —
253
- `"…without intermediaries. "` then `"It was invented"`, never
254
- `"…intermediaries."` followed by `"It was invented"`.
255
-
256
- **`add_run` belongs to the paragraph, not to a run.** Keep the paragraph in a
257
- variable and call `p.add_run(...)` for every run in it — chaining a second run off
258
- the first raises `AttributeError: 'Run' object has no attribute 'add_run'`. A run
259
- owns `.text`, `.bold`, `.italic` and `.font`, and nothing else: it has no
260
- `add_run`, no `add_paragraph`, and no `.style`.
261
-
262
- A run is also not a string: `p.add_run(" ") * 2` raises
263
- `TypeError: unsupported operand type(s) for *: 'Run' and 'int'`. Put any repeated
264
- text inside the string itself — and reach for neither, since spacing is the
265
- style's job, not padding you type.
266
-
267
- Character detail goes through `run.font` — size, color:
268
-
269
- ```python
270
- from docx.shared import Pt, RGBColor
271
-
272
- p = doc.add_paragraph()
273
- run = p.add_run("Key finding")
274
- run.font.size = Pt(14)
275
- run.font.color.rgb = RGBColor(0x1A, 0x73, 0xE8) # RGB in all caps
276
- ```
277
-
278
- `Pt`, `Inches`, and `RGBColor` all import from `docx.shared` — there is no
279
- `docx.util` and no `docx.dml.color`; those are python-pptx paths and fail here.
280
-
281
- ## Styles must exist in the document
282
-
283
- `style="List Bullet"` names a style **inside the document**. A missing name
284
- raises `KeyError: "no style with name 'List Bullet'"` at `add_paragraph` time.
285
-
286
- A **new** `Document()` ships these styles — safe to use without checking:
287
- `Title`, `Heading 1` … `Heading 9`, `Normal`, `List Bullet` (+ ` 2`, ` 3`),
288
- `List Number` (+ ` 2`, ` 3`), `Intense Quote`, and the table style `Table Grid`.
289
- Do not invent other names for a new document. (An uploaded document carries
290
- only its own styles — when editing one, load `references/rework.md` for the
291
- guard.)
292
-
293
- ## Headings and lists
294
-
295
- - `doc.add_heading(text, level=N)` — level `0` is the document title style,
296
- `1`–`9` map to `Heading 1`–`Heading 9`. Any other level raises
297
- `ValueError: level must be in range 0-9`.
298
- - Bullets: one `add_paragraph(point, style="List Bullet")` per point, over a flat
299
- list of plain strings. Never pack several points into one paragraph with `\n` —
300
- a `\n` is a soft line break inside the same list item, not a new bullet. A
301
- bullet that needs a label and a detail is one string (`"Limited supply — 21
302
- million coins"`), never a dict entry or a tuple.
303
- - Numbered lists: `style="List Number"`. Indent a level with `List Bullet 2` /
304
- `List Number 2`.
305
-
306
- ## Tables
307
-
308
- Write the whole table as a list of row lists — header first — then let the code
309
- above derive everything from it. **Always `rows=1` and `cols=len(rows[0])`**:
310
-
311
- ```python
312
- rows = [
313
- ["Item", "Status"], # header
314
- ["Search", "Shipped"],
315
- ["Export", "In review"],
316
- ]
317
- table = doc.add_table(rows=1, cols=len(rows[0]))
318
- table.style = "Table Grid" # borders; omit for invisible grid
319
- for index, cells in enumerate(rows):
320
- row = table.rows[0].cells if index == 0 else table.add_row().cells
321
- for column, value in enumerate(cells):
322
- row[column].text = value
323
- ```
324
-
325
- That shape exists because the two hand-written alternatives both fail:
326
-
327
- - **`rows=` is a count of blank rows created immediately, not a maximum.**
328
- `add_table(rows=4, …)` followed by `add_row()` per entry leaves three empty
329
- rows sitting between the header and the data, plainly visible in the finished
330
- document. `rows=1` is the header; every other row comes from `add_row()`.
331
- - **Unpacking a row into fixed names breaks the moment a row is a different
332
- width.** `for name, q3, q4 in data:` raises
333
- `ValueError: too many values to unpack (expected 3, got 4)`, and hand-counting
334
- `cols=` against the data is the same mistake one step earlier. Index the cells
335
- instead, and take the column count from the header.
336
-
337
- Address cells as `table.cell(row, col)` or `table.rows[r].cells[c]` — they are
338
- the same cell. Rows only grow at the bottom: there is no insert-at.
339
- `table.rows[9]` on a 4-row table raises `IndexError`. Write text with
340
- `cell.text = "…"`; for formatting inside a cell go through `cell.paragraphs[0]`
341
- and its runs like any other paragraph.
342
-
343
- ## Images and page breaks
344
-
345
- `doc.add_picture(name, width=…)` appends the image in its own paragraph. Pass
346
- only one of `width`/`height`; passing both distorts the picture.
347
-
348
- ```python
349
- from docx.shared import Inches
350
-
351
- doc.add_picture("figure1.png", width=Inches(5.5))
352
- doc.add_page_break()
353
- ```
354
-
355
- **Do not soft-fail images or imports.** Never wrap `add_picture` or an import in
356
- `try`/`except` that prints a warning and continues. A missing file must raise so
357
- you fix it and rerun — a document saved without the requested image is a failed
358
- turn, not a success.
359
-
360
- ## Errors
361
-
362
- - `ModuleNotFoundError: No module named 'docx'` means `packages` was missing or
363
- wrong — add `["python-docx==1.2.0"]` and rerun. Never try to install it, and
364
- never "fix" it by importing `python_docx`; the import stays `docx`.
365
- - `TypeError: 'Table' object is not subscriptable` — a table was indexed
366
- directly (`table[0]`). Cells are reached through `table.rows[r].cells[c]` or
367
- `table.cell(r, c)`; a whole row of cells is `table.add_row().cells`.
368
- - `KeyError: "no style with name '…'"` — the style is not in this document. For
369
- a new document use only the names listed under Styles.
370
- - `NameError: name 'RGBColor' is not defined` (or `Pt`, `Inches`) — the import
371
- line is missing that name. Keep the sample's single
372
- `from docx.shared import Inches, Pt, RGBColor` rather than importing one at a time.
373
- - `SyntaxError: invalid syntax` on a one-line `for`/`if` means the source was
374
- collapsed — restore multi-line newlines from the sample and rerun. Underscores
375
- in names (`add_paragraph`, not `addparagraph`) must stay. Do not switch to
376
- `python -c` or change the package pin.
377
- - On an `AttributeError` from python-docx the API name is wrong; on a `TypeError`
378
- about positional arguments the call passes the wrong number of them — usually
379
- several strings where one is allowed. Fix either against this file's examples,
380
- reading the line number in the traceback. Do not retry the same call, and do
381
- not switch to a shell.
382
- - `attachment … not found in this chat` means `inputs` listed an id that is not
383
- in this chat (often a copied placeholder like `att_doc`). For a new document,
384
- omit `inputs` entirely and rerun. Only stage real ids from prior tool results.
385
- - Never print the document's bytes or base64 — stdout is capped and the file
386
- travels through `outputs`. A build call prints exactly one line (e.g. `9
387
- paragraphs, 1 table(s)`).
388
- - Never pass an absolute path to `save()`.
389
-
390
- ## Finish
391
-
392
- When `exitCode` is `0` and `attachments` lists the `.docx`, the document is
393
- done — the `exec` result carries
394
- `attachments: [{ attachmentId, fileName, byteLength }]` and the file is already
395
- attached to the chat for the user to open or save, exactly like a
396
- `generate_image` result. Stop tool use and reply with a single line: file name
397
-
398
- - the count line from stdout. Exactly one successful `exec` per request. If
399
- the result has `missingOutputs` instead, the file was never written: check the
400
- `save()` name matches the declared output and rerun once.
@@ -1,125 +0,0 @@
1
- # Changing Some Paragraphs of a Document (bundled scripts)
2
-
3
- "Replace the first 10 facts", "change fact 3", "swap these bullets for those",
4
- "reword paragraph 7": two `exec` calls, both running a script bundled with this
5
- skill. **Write no Python.** There is no `command` in this job — a call with
6
- `command` is the wrong call. Pass `skill`, `script`, and `scriptArgs` exactly as
7
- shown, with `inputs` staging the document by its `attachmentId` (from the
8
- earlier `exec` result or the `[Attached file …]` line — copy it verbatim, never
9
- invent one).
10
-
11
- | Step | The `exec` call |
12
- | --------------------------------------- | ---------------------------------------------------------------- |
13
- | 1. see the paragraphs and their indexes | `scripts/list_paragraphs.py`, `inputs` staged, no `outputs` |
14
- | 2. replace exactly the chosen indexes | `scripts/replace_paragraphs.py`, `inputs` staged, one `outputs` |
15
-
16
- ## Step 0 — find the document's `attachmentId`
17
-
18
- The id is in the chat already, never invented: a document built earlier in
19
- this chat has it in the `attachments` of the `exec` result that produced it —
20
- `{"attachmentId":"922bd4e17517b90593be1c5ae4f12fbd","fileName":"cat_facts.docx"}`
21
- — and a document the user uploaded has it on the `[Attached file …]` line of
22
- their message. Copy that exact id into `inputs`. An `inputs` entry with a
23
- `path` and no `attachmentId` is an _image_ upload and is refused for a
24
- document:
25
-
26
- ```json
27
- { "inputs": [{ "path": "existing.docx" }] }
28
- ```
29
-
30
- ## Step 1 — list the paragraphs (no `outputs`)
31
-
32
- ```json
33
- {
34
- "language": "python",
35
- "packages": ["python-docx==1.2.0"],
36
- "inputs": [{ "attachmentId": "<real id>", "path": "existing.docx" }],
37
- "skill": "word",
38
- "script": "scripts/list_paragraphs.py",
39
- "scriptArgs": ["existing.docx"]
40
- }
41
- ```
42
-
43
- Every key above is required — `inputs` with the document's real
44
- `attachmentId`, `skill`, `script`, `scriptArgs`. A call missing `skill` or
45
- `inputs` is refused.
46
-
47
- It prints one line per paragraph — `index`, style, text — then a count line.
48
- Pick the indexes to replace from that list:
49
-
50
- - Only body paragraphs (`Normal`, `List Bullet`, `List Number`) are facts,
51
- points, or bullets. `Title`, `Heading N`, and an intro sentence are never
52
- counted as one.
53
- - "The first 10 facts" = the first 10 body-paragraph indexes after the heading
54
- or sentence that introduces them — not indexes 0–9.
55
- - Fewer facts in the document than asked for: replace the ones that exist and
56
- say so in the reply.
57
-
58
- ## Step 2 — replace exactly those paragraphs (one `outputs` entry)
59
-
60
- `scriptArgs` is: input name, output name, then `index, new text` pairs — one
61
- pair per replaced paragraph, as many pairs as facts requested. The output name
62
- keeps the input's stem plus `_revised`.
63
-
64
- Count the pairs before sending: "the first 10 facts" is 10 pairs — 20 strings
65
- after the two file names, 10 different sentences, the last index being
66
- start + 9.
67
-
68
- ```json
69
- {
70
- "language": "python",
71
- "packages": ["python-docx==1.2.0"],
72
- "inputs": [{ "attachmentId": "<real id>", "path": "existing.docx" }],
73
- "outputs": ["existing_revised.docx"],
74
- "skill": "word",
75
- "script": "scripts/replace_paragraphs.py",
76
- "scriptArgs": [
77
- "existing.docx", "existing_revised.docx",
78
- "3", "Dogs have about 1,700 taste buds.",
79
- "4", "A dog's nose print is unique, like a fingerprint."
80
- ]
81
- }
82
- ```
83
-
84
- Wrong, for this job — a `command` instead of a `script`:
85
-
86
- ```json
87
- { "command": "from docx import Document\ndoc = Document(\"existing.docx\")\nfor i in range(1, 11): ..." }
88
- ```
89
-
90
- Each new text is one complete plain sentence, no markdown, each different. The
91
- script keeps each paragraph's paragraph style, refuses a heading index,
92
- refuses text that already reads the same, and prints
93
- `K of N paragraphs replaced`.
94
-
95
- ## Finish
96
-
97
- `exitCode 0` plus an attachment = done. Reply with one line: the file name and
98
- the printed count line. Do not call `exec` again for this request.
99
-
100
- ## Errors
101
-
102
- The script stops with a message that names the fix; correct the arguments and
103
- rerun the **same script** — never switch to writing Python.
104
-
105
- - `usage: replace_paragraphs.py …` — the pairs are incomplete: after the two
106
- file names, arguments alternate `index`, `text`.
107
- - `scriptArgs name "existing_revised.docx" but the working directory starts
108
- empty` — the call has no `outputs`; add `"outputs": ["existing_revised.docx"]`
109
- (the same name as in `scriptArgs`) and rerun the same script.
110
- - `script runs need the owning skill name in skill` — add `"skill": "word"`.
111
- - `index N is the heading '…'` — that paragraph is a heading, not a fact. Pick
112
- body indexes from the Step 1 list.
113
- - `index N is outside the document's M paragraphs` — re-read the Step 1 list;
114
- indexes run from 0 to M-1.
115
- - `index N already reads exactly that` — the new text equals the old one; write
116
- a different sentence.
117
- - `output … must be a new name` — the output name equals the input's; use
118
- `existing_revised.docx`.
119
- - `an id-less input stages an uploaded image` — the `inputs` entry has no
120
- `attachmentId`; add the document's id from Step 0 and rerun the same script.
121
- - `usage: list_paragraphs.py <input.docx>` — `scriptArgs` was left out; pass
122
- the staged path, `["existing.docx"]`.
123
- - `PackageNotFoundError` / `attachment … not found` / `does not exist in the
124
- working directory` — `inputs` is missing or carries an invented id; stage the
125
- document by its real `attachmentId`.
@@ -1,141 +0,0 @@
1
- # Reading a Word Document to Answer in Chat
2
-
3
- When the user asks what an attached `.docx` _says_ — a summary, a question
4
- answered, specific content pulled out — the deliverable is your reply in the
5
- chat, not a file. This is a **read request**: exactly one `exec` call, staging
6
- the document in `inputs` and declaring **no `outputs`**, whose whole job is to
7
- print the document's text so you can read it in the result.
8
-
9
- ## Staging the Document
10
-
11
- Stage the document **by its `attachmentId`**. The id comes from wherever the
12
- document entered the chat:
13
-
14
- - **Produced earlier in this chat** — the `attachmentId` is in that `exec` result.
15
- - **Uploaded by the user** — the `[Attached file …]` line on their message names
16
- it, when the message carries one:
17
-
18
- ```
19
- [Attached file "report.docx" (application/vnd.openxmlformats-officedocument.wordprocessingml.document) — attachmentId: 4f9c2ab1]
20
- ```
21
-
22
- Copy the id verbatim — never invent one, never stage a document id-less: an
23
- id-less input resolves to an uploaded _image_, so it can never reach a
24
- document. If no `attachmentId` for the document appears anywhere in the chat,
25
- say you cannot open that file and ask the user to attach it again — do not
26
- retry. An attachment from an earlier turn can be used when its attachment id
27
- is available in the conversation.
28
-
29
- ## The exec call
30
-
31
- ```json
32
- {
33
- "language": "python",
34
- "packages": ["python-docx==1.2.0"],
35
- "inputs": [{ "attachmentId": "<id from the [Attached file …] line>", "path": "existing.docx" }],
36
- "maxOutputChars": 24000,
37
- "command": "..."
38
- }
39
- ```
40
-
41
- - `packages` — `["python-docx==1.2.0"]` on every call. The PyPI package is
42
- `python-docx` but the import is `docx`; never list `docx` as the package —
43
- that resolves a different, abandoned library. This exact version ships with
44
- the app and installs with no network.
45
- - `inputs` — the staged document lands in the working directory under the bare
46
- `path` name; open `Document("existing.docx")` by that name only. The working
47
- directory starts empty on every call.
48
- - No `outputs` — a read builds nothing.
49
- - `maxOutputChars` — stdout cap in characters (default 8192, max 65536). Set
50
- it only on a read call, where the document text must fit in one result —
51
- keep the sample's 24000. A build call prints one line and never needs it.
52
- - `command` — the multi-line Python source, with real newline characters.
53
- Never collapse it to one line joined by `;` — a `for`/`if`/`with` after a
54
- semicolon is a `SyntaxError`.
55
-
56
- ## The Read Program
57
-
58
- The read program prints every paragraph under its style name — the style names
59
- are the document's structure — then every table:
60
-
61
- ```python
62
- from docx import Document
63
-
64
- doc = Document("existing.docx")
65
- for p in doc.paragraphs:
66
- text = p.text.replace("\n", " ").strip()
67
- if text:
68
- print(f"[{p.style.name}] {text}")
69
- for i, t in enumerate(doc.tables):
70
- print(f"[Table {i + 1}]")
71
- for row in t.rows:
72
- print(" | ".join(c.text.replace("\n", " ") for c in row.cells))
73
- ```
74
-
75
- The `replace` calls are load-bearing: a multi-paragraph cell and a soft line
76
- break both embed `"\n"` in `.text`, and an embedded newline would split one
77
- table row — or one paragraph — across two printed lines. Flattened, every line
78
- starts with its `[...]` marker and every table row is exactly one line.
79
-
80
- Tables print after the body text — python-docx does not expose their position
81
- between paragraphs. When that order matters to the answer, say the tables are
82
- listed separately. Note `p.style.name` in the sample: `paragraph.style` is a
83
- style object, not a string — go through `.name` to print or compare it.
84
-
85
- **You cannot summarize in the call that reads.** The words in `command` are
86
- fixed before the program runs, so one call cannot inform itself: any summary
87
- written into it was written blind — recalled or invented, not read. Python only
88
- _transports_ the text; the summarizing happens in your reply, after the result
89
- comes back.
90
-
91
- ## Answering
92
-
93
- **A successful read ends tool use.** When the result prints the document, reply
94
- with the summary or the answer as chat text. **Scale the reply to the
95
- document**: a summary is much shorter than what it summarizes — a page or two
96
- of source earns three to five sentences, and only a long document earns
97
- sections. Restating the document near its full length is not a summary. Do
98
- **not**:
99
-
100
- - call `exec` again to "re-check", "read more", or read the same document a
101
- second time;
102
- - build a summary `.docx` the user never asked for — an unrequested file is a
103
- failed turn, not a bonus.
104
-
105
- If stdout ends with `… [truncated]`, the document is longer than the cap:
106
- answer from what came back and say the answer covers the document up to that
107
- point. Do not rerun the read — it prints the same beginning again.
108
-
109
- If the user asks for the summary **as a file**, that is a read followed by a
110
- build: the read call above first, then one build call that writes the new
111
- document from the text you actually read (load `references/create.md` for the
112
- build). The read still declares no `outputs`, delivers nothing, and does not
113
- count against the one successful build `exec` per document request — but it
114
- belongs before the build, never after it.
115
-
116
- ## Errors
117
-
118
- - `PackageNotFoundError: Package not found at '…'` — the document was never
119
- staged, or an id-less entry staged an image under a `.docx` path. Add
120
- `inputs: [{ "attachmentId": "<id from the exec result or the [Attached file …]
121
- line>", "path": "existing.docx" }]` and open that exact path. If no id is
122
- available, ask the user to attach the file again rather than guessing a name.
123
- - `attachment … not found in this chat` means `inputs` listed an id that is not
124
- in this chat (often a copied placeholder like `att_doc`). Only stage real ids
125
- from prior tool results or `[Attached file …]` lines; if none exists, ask the
126
- user to re-attach.
127
- - `ModuleNotFoundError: No module named 'docx'` means `packages` was missing or
128
- wrong — add `["python-docx==1.2.0"]` and rerun. Never try to install it, and
129
- never "fix" it by importing `python_docx`; the import stays `docx`.
130
- - `SyntaxError: invalid syntax` on a one-line `for`/`if` means the source was
131
- collapsed — restore multi-line newlines from the sample and rerun. Do not
132
- switch to `python -c` or change the package pin.
133
- - A read result ending in `… [truncated]` means the document outgrew the cap:
134
- answer from what came back and say the answer covers the document up to that
135
- point. Do not rerun the read — it prints the same beginning again.
136
-
137
- ## Finish
138
-
139
- When the read result prints the document, stop tool use and answer the user in
140
- the chat. Exactly one read `exec` per request; only a read call prints
141
- document text.