tbcresearch.
Menu
← Back to the research index
F-004 / Result

A batch-1 INT4 decode megakernel on a 2018 Turing GPU

A GPTQ INT4 decode megakernel on a Quadro RTX 4000 running 1.09–1.64× faster than llama.cpp Q4_0 on Qwen3-0.6B and 1.7B.

A GPTQ INT4 decode megakernel on a Quadro RTX 4000 that runs 1.09–1.64× faster than llama.cpp Q4_0 on Qwen3-0.6B and Qwen3-1.7B.

The abandoned lattice-coding approach and every other dead end are documented in the repository.

Cite this entry

@misc{tbc_latticemk,
  title = {A batch-1 INT4 decode megakernel on a 2018 Turing GPU},
  author = {{TBC Research}},
  note = {F-004, catalogue version 1},
  howpublished = {\url{https://tbcresearch.org/research/latticemk/}}
}
← All research