среда, 7 октября 2026 г.

comparison of code from nvcc vs clang

Lets continue our methodical dissection of PTX/SASS code generated by different compilers

Last week I implemented on CUDA count-sketch algorithm. It requires d = 2 * t + 1 calculating of 2 different family of hashes. As you can see d is always odd and don't fit on warpsize 32, so I used reverse approach - warp processing 32 items from stream and each thread calculates it's own hashes in loop. The code is terrible, consists of nothing but dirty hacks and optimized for specific structures, so that’s why I’m ashamed to show it in its entirety. Nsight shows that code for hashes calculating is where most of the time is spent (~80%), so this code will be our test data. As you can see it's just plain boring integer arith - no fashionable mma or even warp reduce operations

For tests I used clang-21 and nvcc/ptxas from CUDA SDK 13.1. Output data located in 2 sub-dirs:

clang-21, options for clang --target=nvptx64-nvidia-cuda --cuda-device-only -fomit-frame-pointer -O --offload-arch=sm_XX -S

nvcc, see details how to dump LLVM bc here. Btw cicc from sdk 13.1 is based on llvm21 too, like
strings ../../../13.1/cicc/cicc | grep ^LLVM
LLVM21.0.0

As final stage the same ptxas was used to produce object files with same set of options -c -arch sm_XX

Results

  1. size of object file from nvcc 19200 bytes vs 16384 from clang
  2. nvcc version surpasses in speed by ~15-18%

Why? Lets check SASS with nvdisasm

clang version has SHI_REGISTERS=38 for function _Z9BOBHash32PKcjj and SHI_REGISTERS=24 for _Z18MurmurHash3_x86_32PKvij

nvcc: SHI_REGISTERS=60 for _Z9BOBHash32PKcjj and SHI_REGISTERS=59 for _Z18MurmurHash3_x86_32PKvij

So nvcc just applied more aggressive optimization, used more registers and got significant speed-up 

clang generates odd functions

nvdisasm also shows strange set of stubs for functions from CUDA runtime API - like cudaMalloc. I don't know why they are in resulting object file despite --cuda-device-only flag. Seems that removing them in clang is not easy task - AST considered as read-only and I am too lazy to write my own IR plugin

So as usually I just wrote quick & dirty perl script to remove them from PTX. New object file has size 12032 bytes (and only 21 section vs 33 in original file)

optimization with dg2.pl

Next I applied stall counts reduction (see details here) for both object files. Results:

clang

11 instrs with gain: 0.055838
reduction ~5%, but overall speed-up was ~3%

nvcc

55 instrs with gain: 0.0851392
reduction ~8.5% - it’s not surprising: the larger codebase, the higher the likelihood of finding redundancies. Overall speed-up was ~5% - actually this is very good result
 

holes in allocated registers

Finally let's check how tight were registers allocated

clang

2 public syms, avg len 13.500000 blocks
; 118 regs, avg 4.370370 per block
; 8 preds, avg 0.296296 per block
; 10 BD, avg 0.370370 per block
; 23 holes in regs (56), 0.410714
; 2 functions with holes (1.000000 from total), can reduce 4 holes 0.173913 from 23

nvcc

2 public syms, avg len 19.500000 blocks
; 200 regs, avg 5.128205 per block
; 23 preds, avg 0.589744 per block
; 32 BD, avg 0.820513 per block
; 58 holes in regs (113), 0.513274
; 2 functions with holes (1.000000 from total), can reduce 25 holes 0.431034 from 58

As you can see code from nvcc used more predicates, barriers and allocated registers contains much more (more than twofold) unused holes

Conclusion

Metric / Attribute Clang 21 (NVPTX) NVCC 13.1 (cicc) Difference / Impact
Object File Size 16,384 bytes (12,032 scrubbed) 19,200 bytes NVCC +17% larger
Execution Time Baseline ~15–18% faster NVCC wins
BOBHash32 regs 38 registers 60 registers NVCC uses +58% more registers
MurmurHash3 regs 24 registers 59 registers NVCC uses +145% more registers
dg2.pl Speedup +3% (5% stall reduction) +5% (8.5% stall reduction) Higher gain on larger codebase
Register Holes 23 holes (118 regs, 4 fixable) 58 holes (200 regs, 25 fixable) NVCC leaves 2.5x more holes

Nvidia added many aggressive custom optimization passes which results in the generation of significantly faster code

On other side almost all MLIR based compilers require vanilla LLVM NVPTX backend and so (despite on sophisticated optimizations with tensor/linalg/tosa dialects) produce more slow code :-(

Комментариев нет:

Отправить комментарий