最終更新:2026-07-02 (木) 03:54:55 (14d)
LLM/MTP
Multi-Token Prediction?
Ollama
Ollama 0.31? (2026/06/29)
https://ollama.com/blog/faster-gemma-4-mlx-mtp
- MLXでMTP対応して高速化
- Improved Gemma 4 multi-token prediction (MTP) performance
- Gemma 4 12B@Apple M5 Max
with MTP 95tok/s without MTP 50.2tok/s
Ollama 0.23.1? (2026/05/06)
- Gemma 4 MTP speculative decoding is now supported on Mac
LM Studio
2026/05/22 LM Studio 0.4.14 (Build 4) Stable release of MTP Speculative Decoding Qwen3.6 2026/05/20 LM Studio 0.4.14 (Build 2) Beta release of MTP Speculative Decoding

