跳到主内容
@wquguru
精选70Hacker News Best(web_list)技巧与观点

汇编语言耻辱堂:单指令性能下限挑战

汇编语言耻辱堂:反模式代码集锦

原文
发到 X

Assembly Hall of Shame

Overview

Instruction latency analysis usually focuses on performance optimization—making code run as fast as possible. The Assembly Hall of Shame takes the opposite approach: searching for the absolute floor of single-instruction performance.

🏆 Current Champions 🏆

x86: fxrstor64

Strategy: Use fxrstor64 to load 512-byte FPU/MMX/XMM state from a high-latency MMIO region in the PCIe fabric, then starve the fabric while the load is in flight — a fleet of hammer cores pounds a different high-latency MMIO register with tight 4-byte reads, saturating the PCIe root complex and endpoint with non-posted transactions, so CPU 0's 512-byte fxrstor64 must queue behind all that contending traffic.

Contender: AMD Ryzen 7 5800H

代码 · 5
; CPU 0 — timed instruction
movl $0xfcc68830, %rsi
fxrstor64 %rsi
; CPUs 1..N — hammer loop against a different high-latency location
movl 0xfcc68858, %eax

🏆 Score: 198,002,498,236 cycles

🏆 Time: 62 seconds

Honorable Mentions

A spec-violating unaligned ymm0 load that forced non-posted dword transactions from stalled GPU registers was used to break the fundamental design of System Management Mode in smiiiiiiiiiiiiiiii.

代码 · 1
vmovdqu 0xfcc003b1, %ymm0

Rules

  • Instructions may use whatever setup is necessary, but only a single instruction is eligible to be scored.
  • Trapped/emulated/virtualized instructions may only time the trap, not the handler.
  • Instructions must not be interruptible. rep movs, pause, etc. are disqualified.
  • Times are normalized based on the CPU base clock frequency.
  • All platforms must be in their factory stock configurations - no hardware modifications.

x86 Leaderboard

27. nop

Strategy: nop does nothing. It opens the leaderboard accordingly.

Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz

代码 · 1
nop

Score: 1 cycles

Time: 0 nanoseconds

26. nop16

Strategy: Regular nop was too short, but how do we make nothing take longer? Try a lonnnnnng nop.

Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz

代码 · 1
data16 data16 data16 data16 data16 data16 data16 nopl 0x00000000(%%eax,%%eax,1)

Score: 20 cycles

Time: 7 nanoseconds

25. rdtsc

Strategy: Just a reference instruction to get our bearings.

Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz

代码 · 1
rdtsc

Score: 49 cycles

Time: 18 nanoseconds

24. idiv

Strategy: Use 128-bit dividend (rdx:rax=2:0) with small divisor to push the quotient above the ceiling imposed by sign-extension, driving the longest path through the divider microcode.

Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz

代码 · 4
xorq %rax, %rax   ; rax = 0  (low 64 bits of dividend)
movq $2, %rdx     ; rdx = 2  (high 64 bits: full dividend = 2^65)
movq $5, %rbx     ; divisor → quotient = 2^65/5 ≈ 7.4×10^18
idivq %rbx

Score: 77 cycles

Time: 28 nanoseconds

23. enter

Strategy: Use maximum nesting depth (31) to force 30 display-pointer loads and pushes through the microcode display-walk path.

Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz

代码 · 1
enter $0, $31       ; 0 bytes allocated, nesting depth 31 (maximum)

Score: 112 cycles

Time: 41 nanoseconds

22. fldl

Strategy: Try a small denormal to trigger an FP microcode assist.

Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz

代码 · 3
    movabsq $0x0000000000000001, %rax
    movq    %rax, -8(%rsp)
    fldl    -8(%rsp)

Score: 133 cycles

Time: 49 nanoseconds

21. clflush

Strategy: Just ensure the cache line is dirty.

Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz

代码 · 1
clflush (%rax)          ; rax -> dirty cache line resident in L3

Score: 165 cycles

Time: 60 nanoseconds

20. fsin

Strategy: Use exponent 0x7ff to reach 'special value' processing in microcode; positive/negative, NaN/inf doesn't seem to make a difference, go with QNaN.

Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz

代码 · 4
    movabsq $0x7fffffffffffffff, %rax
    movq    %rax, -8(%rsp)
    fldl    -8(%rsp)
    fsin

Score: 257 cycles

Time: 94 nanoseconds

19. mfence

Strategy: Saturate all write-combining line-fill buffers with movnti stores to distinct cache lines, forcing mfence to drain the full LFB write path to the uncore before retiring.

Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz

代码 · 4
movnti %r9,  0*64(%rdi)   ; ×16 distinct cache lines — saturate the write-combining LFBs
; …
movnti %r9, 15*64(%rdi)
mfence                     ; must drain all pending LFB writes before retiring

Score: 326 cycles

Time: 120 nanoseconds

18. mov cr3

Strategy: Nothing for now, just check how long it takes to invalidate the TLB.

Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)

代码 · 1
mov %rax, %cr3

Score: 352 cycles

Time: 110 nanoseconds

17. fadd

Strategy: Hit x87 FP microcode assist path by using denormal source operand.

Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz

代码 · 2
fldl   subnorm    ; 1e-310: value < DBL_MIN, biased exponent = 0
faddl  subnorm    ; source is subnormal → FP microcode assist

Score: 677 cycles

Time: 249 nanoseconds

16. split lock

Strategy: Align lock-prefixed operand to straddle cache-line boundary, forcing CPU to assert the external bus lock rather than using the fast MESI cache-coherence path.

Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz

代码 · 2
; split_ptr % 64 == 63 — dword spans bytes 63 (line N) and 64–66 (line N+1)
lock xaddl %r9d, (%rdi)

Score: 865 cycles

Time: 319 nanoseconds

15. fdiv -

Strategy: Use subnormal divisor, hardware hands control to microcode assist, assist normalizes operand, performs the division, then restores architectural state.

Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz

代码 · 6
    movabsq $0x3ff0000000000000, %rax   ; 1.0 (normal dividend)
    movq    %rax, -8(%rsp)
    fldl    -8(%rsp)                     ; ST(0) = 1.0
    movabsq $0x0000002000000000, %rax   ; 6.79e-313 (subnormal divisor)
    movq    %rax, -8(%rsp)
    fdivl   -8(%rsp)                     ; ST(0) = 1.0 / subnormal → FP assist

Score: 883 cycles

Time: 325 nanoseconds

14. cpuid

Strategy: Use rakefield to find the highest latency CPUID leaves.

Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz

代码 · 2
movl $6, %eax
cpuid

Score: 1248 cycles

Time: 460 nanoseconds

13. rdrand

Strategy: Execute in a tight loop to deplete the hardware entropy pool faster than it can be refilled, forcing subsequent calls to stall while the entropy source recovers.

Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz

代码 · 1
rdrand %rax

Score: 5,579 cycles

Time: 2.057 microseconds

12. wrmsr

Strategy: Use project:nightshyft to identify high latency MSRs. MCG_CTL on Zen look like a winner: may be a microcode quiesce and synchronize on MCA error banks across hardware units, some potentially off-die, requiring fabric-level communication rather than a simple local register write.

Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)

代码 · 2
movl $0x17b, %ecx       ; MCG_CTL
wrmsr

Score: 34,304 cycles

Time: 10.742 microseconds

11. out

Strategy: Target an I/O port that straddles a NIC device register boundary, triggering the device to quiesce its TX DMA engine on each write.

Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)

代码 · 2
mov $0xf019, %dx
outl %eax, %dx

Score: 49,857 cycles

Time: 15.580 microseconds

10. rdmsr

Strategy: Use project:nightshyft to identify high latency model-specific-registers: VIA uses an undocumented register at 0x133 that gives wildly high response time. No idea what it does.

Contender: VIA Eden Processor 800MHz

代码 · 2
movl $0x133, %ecx ; undocumented MSR
rdmsr

Score: 161,602 cycles

Time: 202.004 microseconds

9. wbinvd

Strategy: Fully load L1/L2/L3 caches with dirty lines to force DRAM writeback of entire hierarchy.

Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)

代码 · 1
wbinvd

Score: 1,616,480 cycles

Time: 506.165 microseconds

8. in

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近