A GPTQ INT4 decode megakernel on a Quadro RTX 4000 that runs 1.09–1.64× faster than llama.cpp Q4_0 on Qwen3-0.6B and Qwen3-1.7B.
The abandoned lattice-coding approach and every other dead end are documented in the repository.
Cite this entry
@misc{tbc_latticemk,
title = {A batch-1 INT4 decode megakernel on a 2018 Turing GPU},
author = {{TBC Research}},
note = {F-004, catalogue version 1},
howpublished = {\url{https://tbcresearch.org/research/latticemk/}}
}