Development·EASYHUB JOURNAL
llama.cpp 0.6.0 adds an extended batch API, /v1/systemone decision-model serving, GLM-5.3-Flash support and a new model-download UI

What changed
llama.cpp released 0.6.0 on October 5. It introduces llama_batch_ext / llama_process() for mixed token-and-embedding batches plus MTP/deepstack state embeddings, adds a /v1/systemone server API for decision models, and expands support for GLM-5.3-Flash, Clef, Ling 3.0 VL and other models. Qwen4Exp gains MTP speculative decoding, with the project reporting roughly 1.5× decode speedup on DGX Spark; new Apple Metal F16-KV flash attention and few-row MMA matmul kernels are reported at up to about 3× in some project tests. The Web UI also gets a Hugging Face Hub data layer, model-download pipeline and memory-fit estimation.
- Original title
- v0.6.0
- Source
- GitHub · github.com
- Topic
- Development
- Source month
- 2026-10
This is a concise EasyHub summary of the linked source, not the full report or original reporting. Availability and preview conditions are described in the summary and original.
Summary page published · Editorial information