Projects/DeepGEMM offline cubin and wheel delivery
Inference Delivery · JIT · Offline Runtime

DeepGEMM offline cubin and wheel delivery

Converting runtime compilation artifacts into a complete, verifiable bundle for isolated production environments.

109→359Bundled kernels
0Cold compile in 8 models
TP=2Representative coverage
7.7×Precompiled load

Problem

JIT and isolated environments do not mix

Runtime compilation can fail or add unpredictable startup cost in environments without compilers, network access, or writable caches.

Pipeline

Collect, union, verify, package

Required shapes were collected from model execution, cubins were unioned and deduplicated, integrity metadata was generated, and the result was packaged into an offline-installable wheel/bundle.

Coverage

Single-GPU and tensor parallel

Eight single-GPU model configurations and a representative TP=2 configuration loaded without cold compilation in the tested scope.

Claim boundary

Startup result, not hot GEMM speed

The 7.7× figure refers to cold/precompiled loading behavior. It is not a claim that steady-state GEMM TFLOPS improved by 7.7×.