cpu-perf 0.1.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (181) hide show
  1. cpu_perf-0.1.0/.gitignore +23 -0
  2. cpu_perf-0.1.0/PKG-INFO +420 -0
  3. cpu_perf-0.1.0/README.md +399 -0
  4. cpu_perf-0.1.0/hatch_build.py +68 -0
  5. cpu_perf-0.1.0/pyproject.toml +46 -0
  6. cpu_perf-0.1.0/src/cpu_perf/__init__.py +15 -0
  7. cpu_perf-0.1.0/src/cpu_perf/__main__.py +360 -0
  8. cpu_perf-0.1.0/src/cpu_perf/brain.py +1149 -0
  9. cpu_perf-0.1.0/src/cpu_perf/context.py +1115 -0
  10. cpu_perf-0.1.0/src/cpu_perf/corpus/.github/ISSUE_TEMPLATE/add-entry.md +20 -0
  11. cpu_perf-0.1.0/src/cpu_perf/corpus/.github/ISSUE_TEMPLATE/dead-or-moved-link.md +11 -0
  12. cpu_perf-0.1.0/src/cpu_perf/corpus/.github/ISSUE_TEMPLATE/evidence-concern.md +11 -0
  13. cpu_perf-0.1.0/src/cpu_perf/corpus/.github/pull_request_template.md +25 -0
  14. cpu_perf-0.1.0/src/cpu_perf/corpus/CONTRIBUTING.md +125 -0
  15. cpu_perf-0.1.0/src/cpu_perf/corpus/LICENSE +21 -0
  16. cpu_perf-0.1.0/src/cpu_perf/corpus/README.md +743 -0
  17. cpu_perf-0.1.0/src/cpu_perf/corpus/_build_info.json +259 -0
  18. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/README.md +19 -0
  19. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/02-branch-misprediction/README.md +130 -0
  20. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/02-branch-misprediction/bench.c +204 -0
  21. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/02-branch-misprediction/build.sh +12 -0
  22. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/02-branch-misprediction/results/raw.txt +86 -0
  23. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/02-branch-misprediction/results/summary.md +31 -0
  24. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/02-branch-misprediction/run.sh +101 -0
  25. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/03-latency-vs-throughput/README.md +208 -0
  26. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/03-latency-vs-throughput/bench.c +446 -0
  27. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/03-latency-vs-throughput/build.sh +14 -0
  28. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/03-latency-vs-throughput/results/raw.txt +141 -0
  29. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/03-latency-vs-throughput/results/summary.md +42 -0
  30. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/03-latency-vs-throughput/run.sh +255 -0
  31. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/04-cache-latency/README.md +116 -0
  32. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/04-cache-latency/bench.c +243 -0
  33. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/04-cache-latency/build.sh +12 -0
  34. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/04-cache-latency/results/raw.txt +307 -0
  35. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/04-cache-latency/results/summary.md +52 -0
  36. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/04-cache-latency/run.sh +145 -0
  37. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/05-measurement-pitfalls/README.md +135 -0
  38. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/05-measurement-pitfalls/bench.c +243 -0
  39. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/05-measurement-pitfalls/build.sh +14 -0
  40. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/05-measurement-pitfalls/results/raw.txt +505 -0
  41. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/05-measurement-pitfalls/results/summary.md +40 -0
  42. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/05-measurement-pitfalls/run.sh +252 -0
  43. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/06-roofline/README.md +246 -0
  44. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/06-roofline/bench.c +807 -0
  45. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/06-roofline/build.sh +12 -0
  46. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/06-roofline/results/raw.txt +477 -0
  47. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/06-roofline/results/summary.md +69 -0
  48. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/06-roofline/run.sh +126 -0
  49. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/07-aos-vs-soa-simd/README.md +286 -0
  50. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/07-aos-vs-soa-simd/bench.c +498 -0
  51. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/07-aos-vs-soa-simd/build.sh +42 -0
  52. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/07-aos-vs-soa-simd/results/raw.txt +162 -0
  53. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/07-aos-vs-soa-simd/results/summary.md +39 -0
  54. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/07-aos-vs-soa-simd/run.sh +123 -0
  55. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/08-autovectorization-aliasing/README.md +341 -0
  56. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/08-autovectorization-aliasing/bench.c +483 -0
  57. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/08-autovectorization-aliasing/build.sh +44 -0
  58. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/08-autovectorization-aliasing/results/raw.txt +509 -0
  59. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/08-autovectorization-aliasing/results/summary.md +64 -0
  60. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/08-autovectorization-aliasing/run.sh +189 -0
  61. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/09-false-sharing/README.md +149 -0
  62. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/09-false-sharing/bench.c +354 -0
  63. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/09-false-sharing/build.sh +12 -0
  64. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/09-false-sharing/results/raw.txt +257 -0
  65. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/09-false-sharing/results/summary.md +50 -0
  66. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/09-false-sharing/run.sh +144 -0
  67. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/10-first-touch/README.md +161 -0
  68. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/10-first-touch/bench.c +482 -0
  69. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/10-first-touch/build.sh +12 -0
  70. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/10-first-touch/results/raw.txt +168 -0
  71. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/10-first-touch/results/summary.md +52 -0
  72. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/10-first-touch/run.sh +146 -0
  73. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/11-syscall-cost/README.md +151 -0
  74. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/11-syscall-cost/bench.c +315 -0
  75. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/11-syscall-cost/build.sh +12 -0
  76. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/11-syscall-cost/results/raw.txt +134 -0
  77. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/11-syscall-cost/results/summary.md +42 -0
  78. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/11-syscall-cost/run.sh +137 -0
  79. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/12-coordinated-omission/README.md +179 -0
  80. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/12-coordinated-omission/bench.c +460 -0
  81. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/12-coordinated-omission/build.sh +12 -0
  82. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/12-coordinated-omission/results/raw.txt +150 -0
  83. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/12-coordinated-omission/results/summary.md +34 -0
  84. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/12-coordinated-omission/run.sh +105 -0
  85. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/13-sgemm-naive-vs-blas/README.md +169 -0
  86. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/13-sgemm-naive-vs-blas/bench.c +636 -0
  87. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/13-sgemm-naive-vs-blas/build.sh +23 -0
  88. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/13-sgemm-naive-vs-blas/results/raw.txt +102 -0
  89. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/13-sgemm-naive-vs-blas/results/summary.md +51 -0
  90. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/13-sgemm-naive-vs-blas/run.sh +158 -0
  91. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/14-pcore-vs-ecore/README.md +150 -0
  92. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/14-pcore-vs-ecore/bench.c +612 -0
  93. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/14-pcore-vs-ecore/build.sh +13 -0
  94. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/14-pcore-vs-ecore/results/raw.txt +194 -0
  95. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/14-pcore-vs-ecore/results/summary.md +35 -0
  96. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/14-pcore-vs-ecore/run.sh +187 -0
  97. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/15-stream-bandwidth/README.md +137 -0
  98. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/15-stream-bandwidth/bench.c +426 -0
  99. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/15-stream-bandwidth/build.sh +12 -0
  100. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/15-stream-bandwidth/results/raw.txt +189 -0
  101. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/15-stream-bandwidth/results/summary.md +29 -0
  102. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/15-stream-bandwidth/run.sh +140 -0
  103. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/README.md +72 -0
  104. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/common/clock_estimate.c +46 -0
  105. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/common/machine.sh +30 -0
  106. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/common/refresh_readme.py +28 -0
  107. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/common/timing.h +75 -0
  108. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/compile_all.sh +18 -0
  109. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/benchmarks/run_all.sh +19 -0
  110. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/notes/banner/README.md +63 -0
  111. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/notes/benchmark-brief.md +97 -0
  112. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/notes/changelog.md +881 -0
  113. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/notes/conventions.md +126 -0
  114. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/notes/owner-brief.md +102 -0
  115. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/notes/sections/01-start-here.md +78 -0
  116. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/notes/sections/02-one-instruction.md +99 -0
  117. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/notes/sections/03-microarchitecture.md +114 -0
  118. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/notes/sections/04-memory-hierarchy.md +101 -0
  119. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/notes/sections/05-measurement.md +95 -0
  120. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/notes/sections/06-models.md +119 -0
  121. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/notes/sections/07-single-thread.md +118 -0
  122. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/notes/sections/08-compilers.md +113 -0
  123. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/notes/sections/09-concurrency.md +140 -0
  124. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/notes/sections/10-numa.md +104 -0
  125. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/notes/sections/11-os-and-io.md +127 -0
  126. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/notes/sections/12-tail-latency.md +129 -0
  127. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/notes/sections/13-inference-on-cpu.md +145 -0
  128. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/notes/sections/14-hardware-generations.md +140 -0
  129. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/notes/sections/15-benchmarks.md +109 -0
  130. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/notes/sections/16-watchlist.md +132 -0
  131. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/notes/status.md +183 -0
  132. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/notes/voice.md +52 -0
  133. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/scripts/build_changelog.py +86 -0
  134. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/scripts/check_format.py +256 -0
  135. cpu_perf-0.1.0/src/cpu_perf/corpus/misc/scripts/check_links.py +287 -0
  136. cpu_perf-0.1.0/src/cpu_perf/corpus.py +344 -0
  137. cpu_perf-0.1.0/src/cpu_perf/corpus_update.py +286 -0
  138. cpu_perf-0.1.0/src/cpu_perf/evidence.py +163 -0
  139. cpu_perf-0.1.0/src/cpu_perf/grammar.py +70 -0
  140. cpu_perf-0.1.0/src/cpu_perf/library/__init__.py +2 -0
  141. cpu_perf-0.1.0/src/cpu_perf/library/chunk.py +85 -0
  142. cpu_perf-0.1.0/src/cpu_perf/library/crawler.py +475 -0
  143. cpu_perf-0.1.0/src/cpu_perf/library/embed.py +104 -0
  144. cpu_perf-0.1.0/src/cpu_perf/library/extract.py +516 -0
  145. cpu_perf-0.1.0/src/cpu_perf/library/resolve.py +172 -0
  146. cpu_perf-0.1.0/src/cpu_perf/library/retrieve.py +409 -0
  147. cpu_perf-0.1.0/src/cpu_perf/library/service.py +445 -0
  148. cpu_perf-0.1.0/src/cpu_perf/library/store.py +601 -0
  149. cpu_perf-0.1.0/src/cpu_perf/locate.py +180 -0
  150. cpu_perf-0.1.0/src/cpu_perf/manifest.py +180 -0
  151. cpu_perf-0.1.0/src/cpu_perf/models.py +316 -0
  152. cpu_perf-0.1.0/src/cpu_perf/net.py +405 -0
  153. cpu_perf-0.1.0/src/cpu_perf/parse_benchmarks.py +165 -0
  154. cpu_perf-0.1.0/src/cpu_perf/parse_readme.py +237 -0
  155. cpu_perf-0.1.0/src/cpu_perf/parse_record.py +257 -0
  156. cpu_perf-0.1.0/src/cpu_perf/prompts.py +124 -0
  157. cpu_perf-0.1.0/src/cpu_perf/render.py +339 -0
  158. cpu_perf-0.1.0/src/cpu_perf/resolve.py +166 -0
  159. cpu_perf-0.1.0/src/cpu_perf/schema.py +162 -0
  160. cpu_perf-0.1.0/src/cpu_perf/search.py +367 -0
  161. cpu_perf-0.1.0/src/cpu_perf/server.py +471 -0
  162. cpu_perf-0.1.0/src/cpu_perf/snippets.py +76 -0
  163. cpu_perf-0.1.0/src/cpu_perf/text.py +260 -0
  164. cpu_perf-0.1.0/tests/conftest.py +43 -0
  165. cpu_perf-0.1.0/tests/context_fixtures.py +407 -0
  166. cpu_perf-0.1.0/tests/eval_queries.py +139 -0
  167. cpu_perf-0.1.0/tests/fixture_site.py +172 -0
  168. cpu_perf-0.1.0/tests/test_cli.py +90 -0
  169. cpu_perf-0.1.0/tests/test_context.py +281 -0
  170. cpu_perf-0.1.0/tests/test_eval.py +34 -0
  171. cpu_perf-0.1.0/tests/test_heavy_use.py +98 -0
  172. cpu_perf-0.1.0/tests/test_library.py +306 -0
  173. cpu_perf-0.1.0/tests/test_repo_knowledge.py +151 -0
  174. cpu_perf-0.1.0/tests/test_retrieval.py +141 -0
  175. cpu_perf-0.1.0/tests/test_search_and_resolve.py +115 -0
  176. cpu_perf-0.1.0/tests/test_server.py +435 -0
  177. cpu_perf-0.1.0/tests/test_soak.py +107 -0
  178. cpu_perf-0.1.0/tests/test_store.py +209 -0
  179. cpu_perf-0.1.0/tests/test_swap.py +83 -0
  180. cpu_perf-0.1.0/tests/test_tokens.py +91 -0
  181. cpu_perf-0.1.0/tests/test_update.py +221 -0
@@ -0,0 +1,23 @@
1
+ # compiled benchmarks
2
+ misc/benchmarks/*/bench
3
+ misc/benchmarks/*/bench_*
4
+ misc/benchmarks/common/clock_estimate
5
+ *.o
6
+ *.dSYM/
7
+ .DS_Store
8
+ # link-check output
9
+ link-report.md
10
+ links.json
11
+ misc/benchmarks/*/results/quick-*
12
+ misc/benchmarks/*/*.s
13
+ __pycache__/
14
+ # MCP server environments and builds
15
+ misc/mcp/.venv/
16
+ misc/mcp/dist/
17
+ misc/mcp/uv.lock
18
+ .pytest_cache/
19
+ # website build output
20
+ misc/site/build/
21
+ misc/site/dist/
22
+ misc/site/test-results/
23
+ misc/site/*.egg-info/
@@ -0,0 +1,420 @@
1
+ Metadata-Version: 2.5
2
+ Name: cpu-perf
3
+ Version: 0.1.0
4
+ Summary: MCP server for the CPU Performance Engineering list: the reading list, its editorial record, its benchmarks and a searchable local index of every linked source.
5
+ Project-URL: Homepage, https://github.com/usamahz/cpu-performance-engineering
6
+ Project-URL: Source, https://github.com/usamahz/cpu-performance-engineering/tree/main/misc/mcp
7
+ Author: usamahz
8
+ License-Expression: MIT
9
+ Keywords: cpu,mcp,model-context-protocol,performance,rag
10
+ Classifier: Operating System :: OS Independent
11
+ Classifier: Programming Language :: Python :: 3
12
+ Classifier: Topic :: Software Development :: Libraries
13
+ Requires-Python: >=3.10
14
+ Requires-Dist: huggingface-hub>=0.20
15
+ Requires-Dist: mcp<3,>=2.2
16
+ Requires-Dist: model2vec>=0.6
17
+ Requires-Dist: pypdfium2>=4
18
+ Provides-Extra: test
19
+ Requires-Dist: pytest>=8; extra == 'test'
20
+ Description-Content-Type: text/markdown
21
+
22
+ # cpu-perf
23
+
24
+ An [MCP](https://modelcontextprotocol.io) server that gives any AI client the
25
+ whole [CPU Performance Engineering](https://github.com/usamahz/cpu-performance-engineering)
26
+ list as something it can search and reason over, instead of a page it has to
27
+ be pasted. Connect it to Claude Code, Codex, Claude Desktop, Cursor or VS
28
+ Code and use it for your own performance work: ask questions, paste `perf stat` or compiler
29
+ output, and the client's own model writes the answer from what the server
30
+ returns, with citations back to the sources.
31
+
32
+ It knows two things.
33
+
34
+ 1. **The repository.** Every entry and its reason, in reading order; the
35
+ watchlist and what would promote each line; the editorial record behind the
36
+ list (every candidate that was considered and left out, with the rule it
37
+ failed; every performance number examined against the seven-field rule,
38
+ with its verdict; the link-verification notes); the fourteen benchmarks
39
+ with their claims, machines, results, analysis, code and raw output; and
40
+ the house rules. This is bundled with the server and loads in a fraction
41
+ of a second.
42
+ 2. **The sources themselves.** On first run the server reads every source the
43
+ README links, plus the main document behind each link (the PDF behind an
44
+ arXiv abstract, the manual behind a vendor landing page, a repository's
45
+ README, a pull request's description), extracts the text with page numbers,
46
+ and builds a local full-text and semantic index of it. Questions are then
47
+ answered from the papers, manuals and documentation, not from memory.
48
+
49
+ Nothing is invented on top: the server reads the README, the section drafts
50
+ and the benchmarks at start-up, the README stays the product, and no generated
51
+ index is committed anywhere.
52
+
53
+ ## Connect it
54
+
55
+ Your AI client starts the server itself, with [uv](https://docs.astral.sh/uv/):
56
+ it runs `uvx cpu-perf`, which fetches the release and starts it. Add that
57
+ command to your client once, as below; run in a terminal, it only waits for
58
+ a client. (`pip install cpu-perf` works too and gives the same `cpu-perf`
59
+ command.)
60
+
61
+ The first start downloads its dependencies, which can take longer than some
62
+ clients wait for a new server. Run `uvx cpu-perf --version` once in a
63
+ terminal first, and every client after that starts it in about a second.
64
+
65
+ | Client | Setup |
66
+ |---|---|
67
+ | Claude Code | one command |
68
+ | Claude Desktop | a few lines of config |
69
+ | Codex (CLI, IDE extension, ChatGPT desktop app) | one command |
70
+ | Cursor, VS Code | a few lines of config |
71
+ | Claude.ai, Claude mobile, ChatGPT on the web | not yet |
72
+
73
+ Every client above starts the server on your own machine; there is nothing
74
+ to host. Claude.ai, the mobile apps and ChatGPT on the web connect only to
75
+ servers on the internet, not to a program on your computer, so they cannot
76
+ use it yet.
77
+
78
+ ### Claude
79
+
80
+ **Claude Code**
81
+
82
+ claude mcp add --scope user cpu-perf -- uvx cpu-perf
83
+
84
+ **Claude Desktop**: Settings, Developer, Edit Config, then add to
85
+ `claude_desktop_config.json`. Desktop does not always see your shell's
86
+ `PATH`, so give the full path that `which uvx` prints:
87
+
88
+ ```json
89
+ {
90
+ "mcpServers": {
91
+ "cpu-perf": { "command": "/full/path/to/uvx", "args": ["cpu-perf"] }
92
+ }
93
+ }
94
+ ```
95
+
96
+ ### ChatGPT
97
+
98
+ The ChatGPT desktop app runs local servers through its Codex host,
99
+ configured as below.
100
+
101
+ ### Codex
102
+
103
+ codex mcp add cpu-perf -- uvx cpu-perf
104
+
105
+ or, in `~/.codex/config.toml` (shared by the CLI, the IDE extension and the
106
+ ChatGPT desktop app):
107
+
108
+ ```toml
109
+ [mcp_servers.cpu-perf]
110
+ command = "uvx"
111
+ args = ["cpu-perf"]
112
+ startup_timeout_sec = 120 # room for the first start's downloads
113
+ # Codex starts servers with a minimal environment; behind a proxy, pass it on:
114
+ # env_vars = ["HTTPS_PROXY", "HTTP_PROXY", "NO_PROXY"]
115
+ ```
116
+
117
+ ### Cursor and VS Code
118
+
119
+ **Cursor** (`~/.cursor/mcp.json`) and **VS Code** (`.vscode/mcp.json`,
120
+ which names the key `servers` and adds `"type": "stdio"`):
121
+
122
+ ```json
123
+ {
124
+ "mcpServers": {
125
+ "cpu-perf": { "command": "uvx", "args": ["cpu-perf"] }
126
+ }
127
+ }
128
+ ```
129
+
130
+ ### What every client sees
131
+
132
+ The answers are markdown written for a model to read. Claude Code and Codex
133
+ show the model only a tool's structured data when a tool returns any, so the
134
+ server returns none by default and every client reads the same text
135
+ (`--structured-output` adds it back for programmatic use). Every tool is
136
+ marked read-only, so no client asks for approval on each call, and slow
137
+ reads of large documents return within a minute, finishing in the
138
+ background.
139
+
140
+ Then ask, for example:
141
+
142
+ - Why does my multithreaded counter stop scaling past two threads?
143
+ - Here is my `perf stat` output; where is the time going?
144
+ - gcc says "not vectorized: complicated access pattern"; what do I change?
145
+ - What does `cycle_activity.stalls_l3_miss` count?
146
+ - What does the Intel optimisation manual say about store forwarding?
147
+ - Give me a reading path for NUMA, ending with something I can run.
148
+ - Why is cppreference not in the list?
149
+ - Is "AVX-512 gives 2x on Zen 5" a claim the list would quote?
150
+
151
+ ## Use it for your own work
152
+
153
+ `ask` takes the question and, optionally, whatever the user pasted as
154
+ `context`. The server reads that output itself, with no model involved:
155
+
156
+ - **`perf stat`** in its plain, `-x` and `-j` forms, per-CPU and interval
157
+ output included: the counters as read, and the ratios computed from them
158
+ (instructions per cycle, frequency, branch, cache and TLB miss rates,
159
+ misses per thousand instructions, stalled-cycle shares, faults and context
160
+ switches per second), each with its formula, computed as perf computes its
161
+ own columns. P-core and E-core counts on hybrid parts are never divided by
162
+ each other.
163
+ - **Top-down** level 1 from `perf stat --topdown`, `-M TopdownL1`, the AMD
164
+ `PipelineL1` group, or toplev, and level 2 from `-M TopdownL2`. A level is
165
+ flagged only against a threshold a listed source states: Intel's own values
166
+ from its TMA metrics sheet, applied only to Intel P-cores and cited with
167
+ every flag. Everything else is reported as measured, without a verdict. A
168
+ flagged level sends the answer to the part of the list about it: a
169
+ memory-bound run to the memory hierarchy, a front-end-bound one to fetch
170
+ and decode.
171
+ - **The machine:** when the PMUs and event names show Intel, AMD or Arm,
172
+ sources about the other vendors' hardware and tools are left out.
173
+ - **How far to trust it:** multiplexed counters (and the lowest running
174
+ share, metric groups included), events that were not counted or not
175
+ supported, how many `-I` intervals were summed, and lines the parser could
176
+ not read, which are listed rather than guessed.
177
+ - **perf's own errors:** a missing metric group, `perf_event_paranoid`
178
+ refusals, unknown or unsupported events, the NMI watchdog: each restated
179
+ with what perf itself says to do, instead of being searched for word by
180
+ word.
181
+ - **gcc `-fopt-info` and clang `-Rpass` remarks:** why each loop was left
182
+ scalar, counted per loop, routed to the list's auto-vectorisation sources and
183
+ benchmark.
184
+ - **Assembly and code:** `objdump -d` with or without the opcode bytes, gdb's
185
+ `disassemble`, `perf annotate`, and source code. The instructions and
186
+ identifiers that matter (gathers, atomics, fences, intrinsics, `alignas`,
187
+ `restrict`) are routed to the matching sections, and so is what a loop does:
188
+ a float sum carried across iterations, which stays one serial chain of adds
189
+ without `-fassociative-math` even when a remark says the loop was vectorised
190
+ (and packed multiplies feeding a run of scalar adds, its shape in assembly),
191
+ an early exit, or fields read from an array of structs.
192
+
193
+ The event names, remarks and identifiers then steer the search, so the
194
+ passages that come back are about the pasted output, not just the question.
195
+
196
+ Every answer says which of the list's sources it carries text from and which
197
+ it does not. Each entry is marked as quoted (with the passages and pages),
198
+ in the library but without a matching passage (with the `read_source` call that
199
+ looks inside), or not read on this machine (with the reason: blocked, refused
200
+ by a proxy, not fetched yet, a talk with only its description). The model is
201
+ told to attribute a claim to a source only through a passage it was given, and
202
+ to offer an unread source as further reading, never as a citation. A listed
203
+ paper the question is about comes with its abstract, and every passage carries
204
+ the list's title for its source rather than the document's own, which is often
205
+ a placeholder such as "Untitled Document".
206
+
207
+ Answers are brief by default: passages are trimmed to the part that matches,
208
+ the benchmark and the editorial record come only when they are relevant, and
209
+ each passage carries an id that `read_source(ref, passage=id)` expands in
210
+ full. `detail="full"` returns whole passages and everything related.
211
+ Repeated questions are answered from a cache until the library changes.
212
+
213
+ ## The first run: building the library
214
+
215
+ The server answers from the repository immediately. In the background it
216
+ fetches the linked sources into a local library:
217
+
218
+ - where: `~/.local/share/cpu-perf` on Linux,
219
+ `~/Library/Application Support/cpu-perf` on macOS,
220
+ `%LOCALAPPDATA%\cpu-perf` on Windows, or `CPU_PERF_DATA_DIR`;
221
+ - how long: a few minutes to a quarter of an hour, depending on the network
222
+ and the large manuals; the crawl resumes where it stopped if the client
223
+ closes the server;
224
+ - how big: one SQLite file of passages, a keyword index and embeddings, a few
225
+ hundred megabytes at most; the downloaded files themselves are not kept;
226
+ - what it skips: very large PDFs are indexed up to a page cap and the rest is
227
+ read on demand; scanned PDFs, compressed PostScript and videos have no text
228
+ to index (videos keep their title and description).
229
+
230
+ To build it in the foreground with progress, run
231
+ `cpu-perf index`; `cpu-perf status --detail` lists every source with
232
+ its state.
233
+
234
+ Several MCP clients (Claude Desktop, Claude Code and Cursor at once, say)
235
+ share one library. Every few minutes one of them, whichever holds an
236
+ operating-system lock on the data folder, does the upkeep: it fetches sources
237
+ that are new, due for a refresh or due for a retry, embeds passages that have
238
+ no vector, and brings a library built by an older release up to date in
239
+ place; a source is fetched again only when a release improves how its kind of
240
+ document is read (this one rejoins words PDFs hyphenate across lines). A
241
+ refresh that fails keeps the copy already in the library, unless the document
242
+ is gone. The lock is released by the system if that client exits or crashes,
243
+ and the work pauses while requests arrive.
244
+
245
+ Some publishers (ACM, IEEE, parts of the Intel and Arm portals) refuse
246
+ automated clients or render their documents only in a browser. Those sources
247
+ are reported as blocked or partial, with the list's own link notes on why, and
248
+ the answer points the reader to the link instead. Coverage is reported as it
249
+ is, never padded.
250
+
251
+ Semantic search uses the small static embedding model
252
+ [potion-base-8M](https://huggingface.co/minishlab/potion-base-8M), downloaded
253
+ once. Without it (offline, or `CPU_PERF_EMBED_MODEL=none`) the library
254
+ falls back to keyword search alone.
255
+
256
+ ### Copyright and politeness
257
+
258
+ The library is built on the user's own machine, or the operator's own server,
259
+ from the public URLs the list links; nothing crawled is committed, published
260
+ or shipped in the package or the container image. The crawler fetches only
261
+ those documents, identifies itself, waits between requests to one host and
262
+ backs off on rate limits. It reads each listed link the way a reader opening
263
+ it would, so it does not consult `robots.txt` unless asked to with
264
+ `--respect-robots` (`CPU_PERF_RESPECT_ROBOTS=1`); sources a site then
265
+ disallows are reported as blocked.
266
+
267
+ ## Tools
268
+
269
+ | Tool | What it answers |
270
+ |---|---|
271
+ | `ask` | The evidence for a question: the best passages from the linked sources (with page numbers), the list's entries and reasons, and, when relevant, the matching benchmark and the editorial record. With `context`, the pasted output's metrics and notes too. Call it first. |
272
+ | `lookup` | Ranked lookup over entries, sections, benchmarks, the record, notes and benchmark code, or over the sources' text (`scope="sources"`). |
273
+ | `search`, `fetch` | Find documents (entries, sections, benchmarks, record items, source passages) by id and read one in full: the pair ChatGPT deep research and company knowledge use. |
274
+ | `get_section` | The table of contents, or one section in dependency order with its benchmarks and, for the watchlist, the promotion conditions. |
275
+ | `get_entry` | One entry in context: why it is listed, what to read before and after, other places it is listed, its benchmark, the numbers examined in it, alternatives left out, link notes. |
276
+ | `reading_path` | What to read, in order, for a topic, ending with the benchmark that reproduces it. |
277
+ | `get_benchmark` | A benchmark's claim, machine (the seven fields), results, analysis, limits, code, build and run scripts, raw output and metrics. |
278
+ | `editorial_record` | Why something is or is not listed: rejections by rule, claims by verdict, link notes. |
279
+ | `check_evidence` | A seven-field audit of a performance claim, with how the list judged similar numbers. |
280
+ | `read_source` | The text of one linked source, from the library or fetched now; with `query`, only the matching passages; with `passage`, one passage in full. |
281
+ | `read_file` | Any repository file the server carries. |
282
+ | `library_status` | Which copy of the list is served, the daily update's state, how much of the linked material is indexed, and what is blocked and why. |
283
+
284
+ **Resources:** `cpuperf://readme`, `cpuperf://contents`, `cpuperf://rules`,
285
+ `cpuperf://watchlist`, `cpuperf://benchmarks`, and the templates
286
+ `cpuperf://section/{number}`, `cpuperf://entry/{id}`,
287
+ `cpuperf://benchmark/{slug}`, `cpuperf://file/{+path}` and
288
+ `cpuperf://source/{id}`.
289
+
290
+ **Prompts:** `ask_the_list`, `study_plan`, `diagnose` (the list's own method:
291
+ the USE method, counters that work, top-down, roofline, then the mechanism),
292
+ `audit_claim`, `reproduce_benchmark` and `review_candidate` (pre-screens a
293
+ proposed entry against CONTRIBUTING.md).
294
+
295
+ Entry ids (`4.3.5`, Start here `1.7`) follow the README and change when it
296
+ does; every output also carries the URL, which does not.
297
+
298
+ ## Keeping the list current
299
+
300
+ Once a day (one request shared by every client on the machine) the server
301
+ asks GitHub for the newest commit of the list. When there is a newer one it
302
+ downloads that commit, keeps only the list's own files (the same set the
303
+ package bundles), checks that they parse into a list no smaller than the one
304
+ it is serving, and switches to it between two calls. Links the new list adds
305
+ are fetched by the next upkeep pass; links it drops leave the answers. An
306
+ entry id that now names a different source is flagged in `get_entry`.
307
+
308
+ Downloaded files are read as data. Nothing from them is imported or run, file
309
+ sizes and paths are checked before anything is written, and a copy that does
310
+ not parse is kept off with a note in `library_status` to upgrade the server.
311
+ A checkout (`--repo`, or running from the repository) is never updated: it is
312
+ the copy being edited. `CPU_PERF_AUTO_UPDATE=0` turns the check off;
313
+ `CPU_PERF_UPSTREAM=owner/repo` follows a fork instead.
314
+
315
+ ## How it stays honest
316
+
317
+ - It reads the README with the same grammar as `misc/scripts/check_format.py`;
318
+ a test fails if the two drift apart, and another fails if the parsed counts
319
+ disagree with the README's badges or the changelog's totals.
320
+ - The README is authoritative. The section drafts contribute only their
321
+ Rejected, Claims, Link notes and Benchmark proposal blocks, joined by file
322
+ number.
323
+ - Benchmark numbers come from one Apple M4 Pro; outputs say so, and the
324
+ server tells the client to quote a number only with all seven fields.
325
+ - Text from sources is fenced and labelled as untrusted data.
326
+ - Metrics from pasted output are computed exactly as the output shows them,
327
+ and the only thresholds applied are the ones Intel publishes for top-down
328
+ level 1, cited each time.
329
+ - A retrieval test set of everyday questions guards the ranking: every change
330
+ must keep its recall (`python tests/eval_queries.py path/to/library.sqlite`
331
+ prints the report).
332
+
333
+ ## Safety of fetching
334
+
335
+ Only URLs that appear in the repository are read. Every request and every
336
+ redirect is checked: http and https on their default ports only, and the host
337
+ must resolve to public addresses (no loopback, private, link-local or cloud
338
+ metadata addresses); the connection is pinned to the address that was checked.
339
+ Downloads and decompression are size-capped and time-boxed. `HTTPS_PROXY` and
340
+ `NO_PROXY` are honoured.
341
+
342
+ ## Speed
343
+
344
+ Everything about the repository is held in memory and every lookup is a
345
+ dictionary walk; passages are served from SQLite FTS5 and a matrix of
346
+ quantised embeddings. Run `cpu-perf --selftest` to see load, index and
347
+ query times on your own machine.
348
+
349
+ ## Configuration
350
+
351
+ | Variable (flag) | Meaning |
352
+ |---|---|
353
+ | `CPU_PERF_DATA_DIR` (`--data-dir`) | Where the library lives. |
354
+ | `CPU_PERF_EMBED_MODEL` (`--embed-model`) | A model2vec model id, or `none` for keyword search only. |
355
+ | `CPU_PERF_AUTO_INDEX=0` (`--no-auto-index`) | Do not build the library in the background. |
356
+ | `CPU_PERF_LIVE_FETCH=0` (`--no-live-fetch`) | `read_source` serves only what is indexed. |
357
+ | `CPU_PERF_LIBRARY=0` (`--no-library`) | Repository knowledge only. |
358
+ | `CPU_PERF_RESPECT_ROBOTS=1` (`--respect-robots`) | Skip what `robots.txt` disallows. By default every listed link is read. |
359
+ | `CPU_PERF_REPO` (`--repo`) | Serve a checkout instead of the bundled copy. |
360
+ | `CPU_PERF_AUTO_UPDATE=0` (`--no-auto-update`) | Serve the installed copy of the list; no daily check. |
361
+ | `CPU_PERF_UPSTREAM` | The `owner/repo` the daily check follows (a fork, say). |
362
+ | `CPU_PERF_MAINTENANCE_SECONDS` | How often the upkeep pass runs (default 300). |
363
+ | `CPU_PERF_STRUCTURED_OUTPUT=1` (`--structured-output`) | Also return structured data and output schemas, for programmatic clients. |
364
+ | `CPU_PERF_LIVE_WAIT_SECONDS` | How long a call waits for a live read before answering "still fetching" (default 40). |
365
+ | `CPU_PERF_TRANSPORT`, `_HOST`, `_PORT`, `_PATH` | HTTP serving (`--transport http`). |
366
+ | `CPU_PERF_ALLOWED_HOSTS` (`--allowed-host`) | Host names accepted over HTTP. |
367
+ | `GITHUB_TOKEN` | Not needed; pull request descriptions come from the public API. |
368
+
369
+ ## Serving over HTTP
370
+
371
+ Nobody using cpu-perf needs this: every client above starts it locally. It
372
+ is for running one shared instance. The same server speaks Streamable HTTP;
373
+ from the repository root:
374
+
375
+ docker build -f misc/mcp/Dockerfile -t cpu-perf .
376
+ docker run -p 8000:8000 -v cpu-perf-data:/data -e CPU_PERF_ALLOWED_HOSTS=your-host cpu-perf
377
+
378
+ The endpoint is `/mcp`, with `/healthz` for health checks, and goes behind
379
+ HTTPS. Without an allowed host the server refuses requests addressed to any
380
+ host but localhost, which is what protects it from DNS rebinding. There is
381
+ no authentication, so anyone with the URL can call the tools, all of which
382
+ only read. The container builds its library into the `/data` volume on first
383
+ start; the image itself carries no crawled text. A public instance serves
384
+ passages of other people's work alongside their links, much as a search
385
+ engine shows snippets.
386
+
387
+ ## Development
388
+
389
+ cd misc/mcp
390
+ pip install -e ".[test]"
391
+ pytest
392
+ cpu-perf --selftest
393
+
394
+ An editable install reads the checkout, so a README edit shows up on the next
395
+ start. The tests need no network: the crawler runs against a local fixture
396
+ site and a deterministic embedder, and the daily update against a local
397
+ stand-in for GitHub. With `CPU_PERF_EVAL_DB` pointing at a crawled
398
+ `library.sqlite`, the retrieval tests also run against real sources.
399
+
400
+ ## Releasing
401
+
402
+ Set the version in `pyproject.toml`, merge, then tag the merge commit on
403
+ `main` with the same version:
404
+
405
+ git tag mcp-v0.1.1 && git push origin mcp-v0.1.1
406
+
407
+ `.github/workflows/mcp-release.yml` builds, tests and publishes to PyPI with
408
+ trusted publishing. Once, before the first release: on PyPI add a pending
409
+ publisher for project `cpu-perf`, owner `usamahz`, repository
410
+ `cpu-performance-engineering`, workflow `mcp-release.yml`, environment
411
+ `pypi`. GitHub creates the `pypi` environment on the first run.
412
+
413
+ A release candidate (`0.1.1rc1`, say) is tagged the same way; it installs
414
+ only when asked for by version, with `uvx cpu-perf@0.1.1rc1`, so testing one
415
+ never reaches people on the latest release.
416
+
417
+ ## Licence
418
+
419
+ MIT, as the repository. The wheel carries the repository's files and its
420
+ LICENSE.