---
title: "Prompts and workflows"
url: https://cpuperf.com/mcp/workflows/
source: https://github.com/usamahz/cpu-performance-engineering/blob/deb5a0bac46760503b6f4a2608bdfed470c8532e/misc/mcp/README.md
commit: deb5a0bac46760503b6f4a2608bdfed470c8532e
---

# Prompts and workflows

The prompts the server ships, with their arguments, and how it reads pasted output.

## ask_the_list

Ask the list: Answer a CPU performance question from the list and its sources, with citations.

- `question` (required)

## audit_claim

Audit a performance claim: Check a number against the seven-field rule and the list's own record.

- `claim` (required)
- `source_url`

## diagnose

Diagnose a performance problem: Work a performance problem through the list's method: USE, counters, top-down, roofline, mechanism.

- `symptom` (required)
- `platform`

## reproduce_benchmark

Reproduce a benchmark: Build and run one of the repository's benchmarks and record the seven fields.

- `slug` (required)
- `machine`

## review_candidate

Review a candidate source: Pre-screen a proposed entry against CONTRIBUTING.md, the owner rules and the record.

- `url` (required)
- `title`
- `section`

## study_plan

Study plan: A study plan on a topic, in the list's dependency order, ending with a benchmark to reproduce.

- `topic` (required)
- `weeks`

## Use it for your own work

`ask` takes the question and, optionally, whatever the user pasted as
`context`. The server reads that output itself, with no model involved:

- **`perf stat`** in its plain, `-x` and `-j` forms, per-CPU and interval
  output included: the counters as read, and the ratios computed from them
  (instructions per cycle, frequency, branch, cache and TLB miss rates,
  misses per thousand instructions, stalled-cycle shares, faults and context
  switches per second), each with its formula, computed as perf computes its
  own columns. P-core and E-core counts on hybrid parts are never divided by
  each other.
- **Top-down** level 1 from `perf stat --topdown`, `-M TopdownL1`, the AMD
  `PipelineL1` group, or toplev, and level 2 from `-M TopdownL2`. A level is
  flagged only against a threshold a listed source states: Intel's own values
  from its TMA metrics sheet, applied only to Intel P-cores and cited with
  every flag. Everything else is reported as measured, without a verdict. A
  flagged level sends the answer to the part of the list about it: a
  memory-bound run to the memory hierarchy, a front-end-bound one to fetch
  and decode.
- **The machine:** when the PMUs and event names show Intel, AMD or Arm,
  sources about the other vendors' hardware and tools are left out.
- **How far to trust it:** multiplexed counters (and the lowest running
  share, metric groups included), events that were not counted or not
  supported, how many `-I` intervals were summed, and lines the parser could
  not read, which are listed rather than guessed.
- **perf's own errors:** a missing metric group, `perf_event_paranoid`
  refusals, unknown or unsupported events, the NMI watchdog: each restated
  with what perf itself says to do, instead of being searched for word by
  word.
- **gcc `-fopt-info` and clang `-Rpass` remarks:** why each loop was left
  scalar, counted per loop, routed to the list's auto-vectorisation sources and
  benchmark.
- **Assembly and code:** `objdump -d` with or without the opcode bytes, gdb's
  `disassemble`, `perf annotate`, and source code. The instructions and
  identifiers that matter (gathers, atomics, fences, intrinsics, `alignas`,
  `restrict`) are routed to the matching sections, and so is what a loop does:
  a float sum carried across iterations, which stays one serial chain of adds
  without `-fassociative-math` even when a remark says the loop was vectorised
  (and packed multiplies feeding a run of scalar adds, its shape in assembly),
  an early exit, or fields read from an array of structs.

The event names, remarks and identifiers then steer the search, so the
passages that come back are about the pasted output, not just the question.

Every answer says which of the list's sources it carries text from and which
it does not. Each entry is marked as quoted (with the passages and pages),
in the library but without a matching passage (with the `read_source` call that
looks inside), or not read on this machine (with the reason: blocked, refused
by a proxy, not fetched yet, a talk with only its description). The model is
told to attribute a claim to a source only through a passage it was given, and
to offer an unread source as further reading, never as a citation. A listed
paper the question is about comes with its abstract, and every passage carries
the list's title for its source rather than the document's own, which is often
a placeholder such as "Untitled Document".

Answers are brief by default: passages are trimmed to the part that matches,
the benchmark and the editorial record come only when they are relevant, and
each passage carries an id that `read_source(ref, passage=id)` expands in
full. `detail="full"` returns whole passages and everything related.
Repeated questions are answered from a cache until the library changes.
