runtime / Engineering

LLMQRT: quantized inference runtime

A PyTorch Extension runtime for two-H200, 32B quantized inference and tensor parallelism.

Problem & contribution

Deploying a 32B quantized model on two H200 GPUs required integration across quantization backends, attention, KV cache and collectives, plus tail-group memory-safety fixes.

Adapted W4A16 / W8A8 / FP8 backends and integrated FlashAttention-2/4, GQA, softcap and SDPA. Used Compute Sanitizer to isolate gemv_kernel_g128 tail-group out-of-bounds accesses; added guards, packed-AWQ TP=2 column/row sharding and NCCL all-reduce.

Validation

Condition / metricResult
Qwen2.5-Coder-32B W4A16 checkpoint size−70.5%
Peak GPU memory−66.8%
End-to-end output throughput+76.8%
Four sanitizer checks on representative shapes0 error / 0 hazard

December 2025–present. Passed numerical, cross-rank token and interactive acceptance checks for the two-H200 deployment.

Source: the author’s September 2026 resume and engineering summary; results apply to the stated test conditions.

Percentages describe the project’s paired Qwen2.5-Coder-32B deployment comparison and depend on its baseline and configuration. They are not general gains for every model or attention backend.