Learn / Down
One instruction, end to end
Read first
Vendors name the same structures differently (Intel's decoded ICache is AMD's op cache, and AMD's macro-op is Arm's MOP), so the stage names below are generic.
- Sources
- 18 entries in 4 parts
- Reproduce it
- 02-branch-misprediction
- Related sections
- §3 Microarchitecture, §14 Hardware generations, §15 Benchmarks
- In the MCP server
cpuperf://section/2
Fetch and decode
-
01
The origin of the decoupled front end, where the predictor runs ahead of fetch and drives instruction prefetch.
-
02
Introduces the decoded micro-op cache and its power case, the structure both x86 vendors later built.
-
03
Where AMD states when the op cache feeds micro-ops, and the fusion and alignment rules for hot loops.
-
04
Measures rather than quotes each x86 core's misprediction penalty, micro-op cache and loop buffer behaviour.
-
05
States what a microcode fix evicts from the decoded cache and which counters show the fall back to legacy decode.
Reproduce it
the cost of a mispredicted branch, sorted against unsorted against branchless.
Reproduce it · 02-branch-misprediction Branch misprediction A mispredicted branch costs on the order of the pipeline depth; sorted or branchless data removes itRename and issue
-
01
The origin of tag-based renaming and reservation stations, from which every out-of-order issue queue descends.
-
02
Shows which chains renaming cannot break, from partial registers to flags, and the zero-cost idioms that do.
-
03
Sets the user-space method that measures the window and register files, and which idioms take no physical register.
Execute
-
01
Defines the dependence graph of an out-of-order core and the critical path through it that sets run time.
-
02
Measures latency, reciprocal throughput and port assignment per instruction, the weights a dependence graph needs.
-
03
States every stage of one Arm core from fetch to issue, then per-instruction latency, throughput and pipe assignment.
-
04
Predicts a real decoder loop's cycle count from its carried chain and issue slots, then measures the prediction.
-
05
Predicts a decoder's refill chain from Arm core latencies, fits extra work under it, and measures the gain.
Memory access and retire
-
01
The origin of memory dependence prediction, which lets a load pass older stores and flushes on a wrong guess.
-
02
Measures which store and load size and offset pairs forward or stall, and which cores predict memory dependences.
-
03
Defines memory-level parallelism as the misses one window overlaps, and shows which core limits cap it.
-
04
Introduces the reorder buffer and defines a precise exception as in-order commit of out-of-order results.
-
05 Where Do Interrupts Happen? report
Measures that an interrupt lands on the oldest unretired instruction, which decides what a sampling profiler blames.