LLMQRT: quantized inference runtime
A PyTorch Extension runtime for two-H200, 32B quantized inference and tensor parallelism.
Problem & contribution
Deploying a 32B quantized model on two H200 GPUs required integration across quantization backends, attention, KV cache and collectives, plus tail-group memory-safety fixes.
Adapted W4A16 / W8A8 / FP8 backends and integrated FlashAttention-2/4, GQA, softcap and SDPA. Used Compute Sanitizer to isolate gemv_kernel_g128 tail-group out-of-bounds accesses; added guards, packed-AWQ TP=2 column/row sharding and NCCL all-reduce.
Validation
| Condition / metric | Result |
|---|---|
| Qwen2.5-Coder-32B W4A16 checkpoint size | −70.5% |
| Peak GPU memory | −66.8% |
| End-to-end output throughput | +76.8% |
| Four sanitizer checks on representative shapes | 0 error / 0 hazard |
December 2025–present. Passed numerical, cross-rank token and interactive acceptance checks for the two-H200 deployment.
Source: the author’s September 2026 resume and engineering summary; results apply to the stated test conditions.
Percentages describe the project’s paired Qwen2.5-Coder-32B deployment comparison and depend on its baseline and configuration. They are not general gains for every model or attention backend.