tbcresearch.
Menu
← Back to the research index
F-005 / Negative result

Speculative speculative decoding on an M4: a documented ceiling

Can SSD-style parallel speculation beat plain MLX speculative decoding on Apple Silicon? No: 66.9 tok/s is the practical ceiling, set by memory bandwidth.

Tests whether SSD-style parallel speculation can beat plain MLX speculative decoding on Apple Silicon.

Answer: no. 66.9 tok/s is the practical ceiling, set by memory bandwidth.

Along the way it measures real ANE and GPU parallelism and a fast CoreML-to-MLX KV-cache handoff.

Cite this entry

@misc{tbc_ssd_apple_silicon,
  title = {Speculative speculative decoding on an M4: a documented ceiling},
  author = {{TBC Research}},
  note = {F-005, catalogue version 1},
  howpublished = {\url{https://tbcresearch.org/research/ssd-apple-silicon/}}
}
← All research