MCP · cpu-perf
The whole list, as something an AI client can search and reason over.
An MCP server that gives any AI client the whole CPU Performance Engineering list as something it can search and reason over, instead of a page it has to be pasted. Connect it to Claude Code, Codex, Claude Desktop, Cursor or VS Code and use it for your own performance work: ask questions, paste perf stat or compiler output, and the client's own model writes the answer from what the server returns, with citations back to the sources.
- Version
- 0.1.1
- Protocol
- 2026-07-28
- tools
- 13
- prompts
- 6
- resources
- 44
- Templates
- 5
claude mcp add --scope user cpu-perf -- uvx cpu-perf
codex mcp add cpu-perf -- uvx cpu-perf
What the server knows
What the server knows
It knows two things.
- The repository. Every entry and its reason, in reading order; the watchlist and what would promote each line; the editorial record behind the list (every candidate that was considered and left out, with the rule it failed; every performance number examined against the seven-field rule, with its verdict; the link-verification notes); the fourteen benchmarks with their claims, machines, results, analysis, code and raw output; and the house rules. This is bundled with the server and loads in a fraction of a second.
- The sources themselves. On first run the server reads every source the README links, plus the main document behind each link (the PDF behind an arXiv abstract, the manual behind a vendor landing page, a repository's README, a pull request's description), extracts the text with page numbers, and builds a local full-text and semantic index of it. Questions are then answered from the papers, manuals and documentation, not from memory.
Nothing is invented on top: the server reads the README, the section drafts and the benchmarks at start-up, the README stays the product, and no generated index is committed anywhere.
Path of a question
What the server's ask tool returns
- 01
Pasted output
perf stat, top-down, compiler remarks, assembly
- 02
Metrics
computed as perf computes them, each with its formula
- 03
Entries and reasons
the list's own sources, in reading order
- 04
Source passages
from the linked papers and manuals, with page numbers
- 05
Benchmark
the matching committed experiment
- 06
Editorial record
what was left out, and the rule it failed
Questions
What to ask it
The example questions in the server's README, with the first results the server's own search tool returns for each at this commit.
-
Why does my multithreaded counter stop scaling past two threads?
- 09-false-sharingFalse sharing
- §9Thread pools and work stealing
- §9Locks, contention and allocators
-
Here is my
perf statoutput; where is the time going? -
gcc says "not vectorized: complicated access pattern"; what do I change?
-
What does
cycle_activity.stalls_l3_misscount? -
What does the Intel optimisation manual say about store forwarding?
-
Give me a reading path for NUMA, ending with something I can run.
- 10-first-touchFirst touch
- §10NUMA and multi-socket
- §10NUMA (Non-Uniform Memory Access): An Overview
-
Why is cppreference not in the list?
-
Is "AVX-512 gives 2x on Zen 5" a claim the list would quote?