Development·EASYHUB JOURNAL
SGLang v0.5.21 adds Decisions and Score APIs, dynamic prefill/decode role switching and a Rust prefix-cache core

What changed
SGLang released v0.5.21 on October 2 with 779 PRs from 227 contributors. The release adds /v1/decisions and /v1/score for using LLMs/VLMs as low-latency classifiers and scorers, lets prefill/decode instances switch roles without a restart, and moves the prefix cache to a Rust core by default. It also adds support for models including DeepSeek-V4.1 Flash, MiMo-V2.6, Qwen-Image 2.1 and FLUX 3 Action. Project benchmarks report up to 22% faster first-token latency for long DeepSeek-V4.1 prompts and 20.6% higher Kimi K3 prefill throughput in PD serving; those figures are project measurements rather than independent benchmarks.
- Original title
- v0.5.21
- Source
- GitHub · github.com
- Topic
- Development
- Source month
- 2026-10
This is a concise EasyHub summary of the linked source, not the full report or original reporting. Availability and preview conditions are described in the summary and original.
Summary page published · Updated · Editorial information