A from-scratch engine that runs Qwen3.5-2B across three compute units:
| Stage | Unit |
|---|---|
| Prefill | Apple Neural Engine (private APIs) |
| Decode | Metal GPU, custom shaders |
| Glue | CPU |
Metal decode matches llama.cpp at about 32 tok/s. No CoreML, Python, or MLX.
Cite this entry
@misc{tbc_ane_infer,
title = {Hybrid Neural Engine, GPU, and CPU LLM inference on Apple Silicon},
author = {{TBC Research}},
note = {F-003, catalogue version 1},
howpublished = {\url{https://tbcresearch.org/research/ane-infer/}}
}