---
title: "Connect the server"
url: https://cpuperf.com/mcp/quickstart/
source: https://github.com/usamahz/cpu-performance-engineering/blob/deb5a0bac46760503b6f4a2608bdfed470c8532e/misc/mcp/README.md
commit: deb5a0bac46760503b6f4a2608bdfed470c8532e
---
# Connect the server

## Connect it

Your AI client starts the server itself, with [uv](https://docs.astral.sh/uv/):
it runs `uvx cpu-perf`, which fetches the release and starts it. Add that
command to your client once, as below; run in a terminal, it only waits for
a client. (`pip install cpu-perf` works too and gives the same `cpu-perf`
command.)

The first start downloads its dependencies, which can take longer than some
clients wait for a new server. Run `uvx cpu-perf --version` once in a
terminal first, and every client after that starts it in about a second.

| Client | Setup |
|---|---|
| Claude Code | one command |
| Claude Desktop | a few lines of config |
| Codex (CLI, IDE extension, ChatGPT desktop app) | one command |
| Cursor, VS Code | a few lines of config |
| Claude.ai, Claude mobile, ChatGPT on the web | not yet |

Every client above starts the server on your own machine; there is nothing
to host. Claude.ai, the mobile apps and ChatGPT on the web connect only to
servers on the internet, not to a program on your computer, so they cannot
use it yet.

### Claude

**Claude Code**

    claude mcp add --scope user cpu-perf -- uvx cpu-perf

**Claude Desktop**: Settings, Developer, Edit Config, then add to
`claude_desktop_config.json`. Desktop does not always see your shell's
`PATH`, so give the full path that `which uvx` prints:

```json
{
  "mcpServers": {
    "cpu-perf": { "command": "/full/path/to/uvx", "args": ["cpu-perf"] }
  }
}
```

### ChatGPT

The ChatGPT desktop app runs local servers through its Codex host,
configured as below.

### Codex

    codex mcp add cpu-perf -- uvx cpu-perf

or, in `~/.codex/config.toml` (shared by the CLI, the IDE extension and the
ChatGPT desktop app):

```toml
[mcp_servers.cpu-perf]
command = "uvx"
args = ["cpu-perf"]
startup_timeout_sec = 120   # room for the first start's downloads
# Codex starts servers with a minimal environment; behind a proxy, pass it on:
# env_vars = ["HTTPS_PROXY", "HTTP_PROXY", "NO_PROXY"]
```

### Cursor and VS Code

**Cursor** (`~/.cursor/mcp.json`) and **VS Code** (`.vscode/mcp.json`,
which names the key `servers` and adds `"type": "stdio"`):

```json
{
  "mcpServers": {
    "cpu-perf": { "command": "uvx", "args": ["cpu-perf"] }
  }
}
```

### What every client sees

The answers are markdown written for a model to read. Claude Code and Codex
show the model only a tool's structured data when a tool returns any, so the
server returns none by default and every client reads the same text
(`--structured-output` adds it back for programmatic use). Every tool is
marked read-only, so no client asks for approval on each call, and slow
reads of large documents return within a minute, finishing in the
background.

Then ask, for example:

- Why does my multithreaded counter stop scaling past two threads?
- Here is my `perf stat` output; where is the time going?
- gcc says "not vectorized: complicated access pattern"; what do I change?
- What does `cycle_activity.stalls_l3_miss` count?
- What does the Intel optimisation manual say about store forwarding?
- Give me a reading path for NUMA, ending with something I can run.
- Why is cppreference not in the list?
- Is "AVX-512 gives 2x on Zen 5" a claim the list would quote?

## The first run: building the library

The server answers from the repository immediately. In the background it
fetches the linked sources into a local library:

- where: `~/.local/share/cpu-perf` on Linux,
  `~/Library/Application Support/cpu-perf` on macOS,
  `%LOCALAPPDATA%\cpu-perf` on Windows, or `CPU_PERF_DATA_DIR`;
- how long: a few minutes to a quarter of an hour, depending on the network
  and the large manuals; the crawl resumes where it stopped if the client
  closes the server;
- how big: one SQLite file of passages, a keyword index and embeddings, a few
  hundred megabytes at most; the downloaded files themselves are not kept;
- what it skips: very large PDFs are indexed up to a page cap and the rest is
  read on demand; scanned PDFs, compressed PostScript and videos have no text
  to index (videos keep their title and description).

To build it in the foreground with progress, run
`cpu-perf index`; `cpu-perf status --detail` lists every source with
its state.

Several MCP clients (Claude Desktop, Claude Code and Cursor at once, say)
share one library. Every few minutes one of them, whichever holds an
operating-system lock on the data folder, does the upkeep: it fetches sources
that are new, due for a refresh or due for a retry, embeds passages that have
no vector, and brings a library built by an older release up to date in
place; a source is fetched again only when a release improves how its kind of
document is read (this one rejoins words PDFs hyphenate across lines). A
refresh that fails keeps the copy already in the library, unless the document
is gone. The lock is released by the system if that client exits or crashes,
and the work pauses while requests arrive.

Some publishers (ACM, IEEE, parts of the Intel and Arm portals) refuse
automated clients or render their documents only in a browser. Those sources
are reported as blocked or partial, with the list's own link notes on why, and
the answer points the reader to the link instead. Coverage is reported as it
is, never padded.

Semantic search uses the small static embedding model
[potion-base-8M](https://huggingface.co/minishlab/potion-base-8M), downloaded
once. Without it (offline, or `CPU_PERF_EMBED_MODEL=none`) the library
falls back to keyword search alone.

## Configuration

| Variable (flag) | Meaning |
|---|---|
| `CPU_PERF_DATA_DIR` (`--data-dir`) | Where the library lives. |
| `CPU_PERF_EMBED_MODEL` (`--embed-model`) | A model2vec model id, or `none` for keyword search only. |
| `CPU_PERF_AUTO_INDEX=0` (`--no-auto-index`) | Do not build the library in the background. |
| `CPU_PERF_LIVE_FETCH=0` (`--no-live-fetch`) | `read_source` serves only what is indexed. |
| `CPU_PERF_LIBRARY=0` (`--no-library`) | Repository knowledge only. |
| `CPU_PERF_RESPECT_ROBOTS=1` (`--respect-robots`) | Skip what `robots.txt` disallows. By default every listed link is read. |
| `CPU_PERF_REPO` (`--repo`) | Serve a checkout instead of the bundled copy. |
| `CPU_PERF_AUTO_UPDATE=0` (`--no-auto-update`) | Serve the installed copy of the list; no daily check. |
| `CPU_PERF_UPSTREAM` | The `owner/repo` the daily check follows (a fork, say). |
| `CPU_PERF_MAINTENANCE_SECONDS` | How often the upkeep pass runs (default 300). |
| `CPU_PERF_STRUCTURED_OUTPUT=1` (`--structured-output`) | Also return structured data and output schemas, for programmatic clients. |
| `CPU_PERF_LIVE_WAIT_SECONDS` | How long a call waits for a live read before answering "still fetching" (default 40). |
| `CPU_PERF_TRANSPORT`, `_HOST`, `_PORT`, `_PATH` | HTTP serving (`--transport http`). |
| `CPU_PERF_ALLOWED_HOSTS` (`--allowed-host`) | Host names accepted over HTTP. |
| `GITHUB_TOKEN` | Not needed; pull request descriptions come from the public API. |