Lets continue our methodical dissection of PTX/SASS code generated by different compilers
Last week I implemented on CUDA count-sketch algorithm. It requires d = 2 * t + 1 calculating of 2 different family of hashes. As you can see d is always odd and don't fit on warpsize 32, so I used reverse approach - warp processing 32 items from stream and each thread calculates it's own hashes in loop. The code is terrible, consists of nothing but dirty hacks and optimized for specific structures, so that’s why I’m ashamed to show it in its entirety. Nsight shows that code for hashes calculating is where most of the time is spent (~80%), so this code will be our test data. As you can see it's just plain boring integer arith - no fashionable mma or even warp reduce operations
For tests I used clang-21 and nvcc/ptxas from CUDA SDK 13.1. Output data located in 2 sub-dirs:
clang-21, options for clang --target=nvptx64-nvidia-cuda --cuda-device-only -fomit-frame-pointer -O --offload-arch=sm_XX -S
nvcc, see details how to dump LLVM bc here. Btw cicc from sdk 13.1 is based on llvm21 too, likestrings ../../../13.1/cicc/cicc | grep ^LLVM
LLVM21.0.0
As final stage the same ptxas was used to produce object files with same set of options -c -arch sm_XX
Results
- size of object file from nvcc 19200 bytes vs 16384 from clang
- nvcc version surpasses in speed by ~15-18%
Why? Lets check SASS with nvdisasm
clang version has SHI_REGISTERS=38 for function _Z9BOBHash32PKcjj and SHI_REGISTERS=24 for _Z18MurmurHash3_x86_32PKvij
nvcc: SHI_REGISTERS=60 for _Z9BOBHash32PKcjj and SHI_REGISTERS=59 for _Z18MurmurHash3_x86_32PKvij
So nvcc just applied more aggressive optimization, used more registers and got significant speed-up
clang generates odd functions
nvdisasm also shows strange set of stubs for functions from CUDA runtime API - like cudaMalloc. I don't know why they are in resulting object file despite --cuda-device-only flag. Seems that removing them in clang is not easy task - AST considered as read-only and I am too lazy to write my own IR plugin
So as usually I just wrote quick & dirty perl script to remove them from PTX. New object file has size 12032 bytes (and only 21 section vs 33 in original file)
optimization with dg2.pl
Next I applied stall counts reduction (see details here) for both object files. Results:
clang
11 instrs with gain: 0.055838 nvcc
55 instrs with gain: 0.0851392reduction ~8.5% - it’s not surprising: the larger codebase, the higher the likelihood of finding redundancies. Overall speed-up was ~5% - actually this is very good result
holes in allocated registers
Finally let's check how tight were registers allocated
clang
2 public syms, avg len 13.500000 blocks
; 118 regs, avg 4.370370 per block
; 8 preds, avg 0.296296 per block
; 10 BD, avg 0.370370 per block
; 23 holes in regs (56), 0.410714
; 2 functions with holes (1.000000 from total), can reduce 4 holes 0.173913 from 23 nvcc
2 public syms, avg len 19.500000 blocks
; 200 regs, avg 5.128205 per block
; 23 preds, avg 0.589744 per block
; 32 BD, avg 0.820513 per block
; 58 holes in regs (113), 0.513274
; 2 functions with holes (1.000000 from total), can reduce 25 holes 0.431034 from 58
As you can see code from nvcc used more predicates, barriers and allocated registers contains much more (more than twofold) unused holes
Conclusion
| Metric / Attribute | Clang 21 (NVPTX) | NVCC 13.1 (cicc) |
Difference / Impact |
|---|---|---|---|
| Object File Size | 16,384 bytes (12,032 scrubbed) | 19,200 bytes | NVCC +17% larger |
| Execution Time | Baseline | ~15–18% faster | NVCC wins |
| BOBHash32 regs | 38 registers | 60 registers | NVCC uses +58% more registers |
| MurmurHash3 regs | 24 registers | 59 registers | NVCC uses +145% more registers |
| dg2.pl Speedup | +3% (5% stall reduction) | +5% (8.5% stall reduction) | Higher gain on larger codebase |
| Register Holes | 23 holes (118 regs, 4 fixable) | 58 holes (200 regs, 25 fixable) | NVCC leaves 2.5x more holes |
Nvidia added many aggressive custom optimization passes which results in the generation of significantly faster code
On other side almost all MLIR based compilers require vanilla LLVM NVPTX backend and so (despite on sophisticated optimizations with tensor/linalg/tosa dialects) produce more slow code :-(
Комментариев нет:
Отправить комментарий