Tests whether SSD-style parallel speculation can beat plain MLX speculative decoding on Apple Silicon.
Answer: no. 66.9 tok/s is the practical ceiling, set by memory bandwidth.
Along the way it measures real ANE and GPU parallelism and a fast CoreML-to-MLX KV-cache handoff.
Cite this entry
@misc{tbc_ssd_apple_silicon,
title = {Speculative speculative decoding on an M4: a documented ceiling},
author = {{TBC Research}},
note = {F-005, catalogue version 1},
howpublished = {\url{https://tbcresearch.org/research/ssd-apple-silicon/}}
}