MCP server / Security
Permissions and boundaries
How it stays honest
- It reads the README with the same grammar as
misc/scripts/check_format.py; a test fails if the two drift apart, and another fails if the parsed counts disagree with the README's badges or the changelog's totals. - The README is authoritative. The section drafts contribute only their Rejected, Claims, Link notes and Benchmark proposal blocks, joined by file number.
- Benchmark numbers come from one Apple M4 Pro; outputs say so, and the server tells the client to quote a number only with all seven fields.
- Text from sources is fenced and labelled as untrusted data.
- Metrics from pasted output are computed exactly as the output shows them, and the only thresholds applied are the ones Intel publishes for top-down level 1, cited each time.
- A retrieval test set of everyday questions guards the ranking: every change
must keep its recall (
python tests/eval_queries.py path/to/library.sqliteprints the report).
Safety of fetching
Only URLs that appear in the repository are read. Every request and every
redirect is checked: http and https on their default ports only, and the host
must resolve to public addresses (no loopback, private, link-local or cloud
metadata addresses); the connection is pinned to the address that was checked.
Downloads and decompression are size-capped and time-boxed. HTTPS_PROXY and
NO_PROXY are honoured.
Copyright and politeness
The library is built on the user's own machine, or the operator's own server,
from the public URLs the list links; nothing crawled is committed, published
or shipped in the package or the container image. The crawler fetches only
those documents, identifies itself, waits between requests to one host and
backs off on rate limits. It reads each listed link the way a reader opening
it would, so it does not consult robots.txt unless asked to with
--respect-robots (CPU_PERF_RESPECT_ROBOTS=1); sources a site then
disallows are reported as blocked.
Keeping the list current
Once a day (one request shared by every client on the machine) the server
asks GitHub for the newest commit of the list. When there is a newer one it
downloads that commit, keeps only the list's own files (the same set the
package bundles), checks that they parse into a list no smaller than the one
it is serving, and switches to it between two calls. Links the new list adds
are fetched by the next upkeep pass; links it drops leave the answers. An
entry id that now names a different source is flagged in get_entry.
Downloaded files are read as data. Nothing from them is imported or run, file
sizes and paths are checked before anything is written, and a copy that does
not parse is kept off with a note in library_status to upgrade the server.
A checkout (--repo, or running from the repository) is never updated: it is
the copy being edited. CPU_PERF_AUTO_UPDATE=0 turns the check off;
CPU_PERF_UPSTREAM=owner/repo follows a fork instead.
Serving over HTTP
Nobody using cpu-perf needs this: every client above starts it locally. It is for running one shared instance. The same server speaks Streamable HTTP; from the repository root:
docker build -f misc/mcp/Dockerfile -t cpu-perf .
docker run -p 8000:8000 -v cpu-perf-data:/data -e CPU_PERF_ALLOWED_HOSTS=your-host cpu-perf
The endpoint is /mcp, with /healthz for health checks, and goes behind
HTTPS. Without an allowed host the server refuses requests addressed to any
host but localhost, which is what protects it from DNS rebinding. There is
no authentication, so anyone with the URL can call the tools, all of which
only read. The container builds its library into the /data volume on first
start; the image itself carries no crawled text. A public instance serves
passages of other people's work alongside their links, much as a search
engine shows snippets.
Tool annotations
| Tool | Read-only | Reaches the web | Idempotent | Destructive |
|---|---|---|---|---|
ask |
yes | no | yes | no |
check_evidence |
yes | no | yes | no |
editorial_record |
yes | no | yes | no |
fetch |
yes | no | yes | no |
get_benchmark |
yes | no | yes | no |
get_entry |
yes | no | yes | no |
get_section |
yes | no | yes | no |
library_status |
yes | no | yes | no |
lookup |
yes | no | yes | no |
read_file |
yes | no | yes | no |
read_source |
yes | yes | no | no |
reading_path |
yes | no | yes | no |
search |
yes | no | yes | no |