ShawnY112358
|
007c3e234c
|
[feat] support in-flight weight update (#10071)
Co-authored-by: 赵晨阳 <zhaochen20@outlook.com>
|
2025-11-25 22:03:13 -08:00 |
|
fzyzcjy
|
7130ad3a29
|
Fix SGLANG_ENABLE_HEALTH_ENDPOINT_GENERATION not working (#13961)
|
2025-11-25 21:58:55 -08:00 |
|
sglang-bot
|
c53e729d45
|
chore: bump sgl-kernel version to 0.3.18.post1 (#13951)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
|
2025-11-25 18:14:28 -08:00 |
|
Lianmin Zheng
|
1ab6ce0e62
|
[Auto Sync] Improve profilers and simplify bench_one_batch_server.py (#13866)
|
2025-11-25 12:13:31 -08:00 |
|
Zhi Yiliu
|
a95a38078b
|
[Fix] Fix uvloop get_event_loop() is not suitable for 0.22.x (#13612)
Signed-off-by: lzy <tomlzy213@gmail.com>
Co-authored-by: lzy <tomlzy213@gmail.com>
|
2025-11-25 01:20:00 +08:00 |
|
Baizhou Zhang
|
04b52fa8d6
|
[chore]Upgrade flashinfer to 0.5.3 (#13751)
|
2025-11-23 23:38:36 -08:00 |
|
Lianmin Zheng
|
90a0133515
|
[Auto Sync] Update http_server.py, io_struct.py, scheduler_... (20251120) (#13679)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Zhuqi Li <zhli@x.ai>
|
2025-11-20 23:17:15 -08:00 |
|
ishandhanani
|
b537ac0d1f
|
Fix ZMQ bind error on non-zero rank nodes when using SGLANG_BLOCK_NONZERO_RANK_CHILDREN=0 (#13686)
|
2025-11-20 22:54:41 -08:00 |
|
sglang-bot
|
bfaf0b8607
|
chore: bump sgl-kernel version to 0.3.17.post2 (#13570)
|
2025-11-19 14:02:57 -08:00 |
|
Binyao Jiang
|
26ca07469b
|
[GLM4.6v] Relax the constraint of non-user role chat completion message schema for new GLM-v release (#13258)
|
2025-11-18 10:58:17 +08:00 |
|
Binyao Jiang
|
90c18a16cb
|
[GLM4.6v] Required changes for bumping up to transformer 5.x (#13229)
|
2025-11-18 10:58:00 +08:00 |
|
Simo Lin
|
25acbbc612
|
Remove verbs from GET endpoint paths to follow REST standards (#13273)
|
2025-11-17 12:34:36 -08:00 |
|
kebyn
|
15db5497d3
|
[feature] Custom base path on FastAPI server (#5879)
Co-authored-by: lianhu.yin <lianhu.yin@nio.com>
Co-authored-by: kebyn <kebyn@kebyn.cc>
Co-authored-by: Liangsheng Yin <lsyincs@gmail.com>
|
2025-11-17 22:40:09 +08:00 |
|
sglang-bot
|
1ca205f6da
|
chore: bump sgl-kernel version to 0.3.17.post1 (#13358)
|
2025-11-15 19:11:41 -08:00 |
|
Lianmin Zheng
|
db7299aa30
|
[Fix] Register custom ops only if they exist (#13321)
|
2025-11-15 17:48:17 -08:00 |
|
Zilin Zhu
|
f0b5ccf5f5
|
[RL] Allow bypassing /health check (#13320)
|
2025-11-15 16:33:54 +08:00 |
|
Yineng Zhang
|
f8d3d80f63
|
chore: bump flashinfer v0.5.2 (#13242)
|
2025-11-14 02:47:09 -08:00 |
|
fzyzcjy
|
ed1d18d472
|
Tiny fix update version logic location (#12620)
|
2025-11-14 17:33:36 +08:00 |
|
billishyahao
|
ee3e337c62
|
[feat] make warmup timeout configurable through SGLANG_WARMUP_TIMEOUT (#13243)
|
2025-11-14 10:14:37 +08:00 |
|
Binyao Jiang
|
9b41f31a66
|
Use 32x32 black image for VLM server warmup and bring glm4.1v back to UT (#13222)
|
2025-11-13 14:21:43 -08:00 |
|
ishandhanani
|
1a5c313f97
|
feat(engine): add rid parameter to methods in Engine class (#13095)
|
2025-11-11 22:37:54 -08:00 |
|
Frank Minors
|
7f5055ed5e
|
Fix cached tokens usage bug (#12814)
Co-authored-by: FrankMinions <liuchen@shinemo.com>
|
2025-11-12 01:02:43 +08:00 |
|
sglang-bot
|
37c40a87a8
|
chore: bump sgl-kernel version to 0.3.17 (#12966)
|
2025-11-10 21:50:58 +08:00 |
|
Mick
|
95876d75cb
|
chore: include a minimum image for vlms when warming-up (#9528)
|
2025-11-10 14:56:59 +08:00 |
|
Lianmin Zheng
|
ae62279071
|
Fix data parallel controller launch for num nodes > 2 (#12822)
|
2025-11-07 12:13:21 -08:00 |
|
yinghui
|
58095cb00a
|
Add timing metrics for requests (#12646)
Co-authored-by: Scott Lee <scottjlee@users.noreply.github.com>
|
2025-11-05 23:07:16 -08:00 |
|
gongwei-130
|
97be66c358
|
fix sgl-kernel version (#12723)
|
2025-11-05 19:01:03 -08:00 |
|
Lianmin Zheng
|
c7d57d5bb3
|
Fix CI and style (#12658)
|
2025-11-05 15:08:15 -08:00 |
|
Lianmin Zheng
|
7a21d8b276
|
Reduce the overhead of nccl symmetric memory (#12524)
Co-authored-by: Nicolas Castet <ncastet@nvidia.com>
|
2025-11-03 11:56:27 -08:00 |
|
Yineng Zhang
|
0c3543d7d5
|
chore: upgrade flashinfer 0.5.0 (#12523)
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
|
2025-11-02 20:54:12 -08:00 |
|
fzyzcjy
|
30ad107028
|
Try to allow NCCL cumem for multi node nvlink case (#11987)
|
2025-10-31 12:48:25 -07:00 |
|
Lianmin Zheng
|
c0652d907b
|
Clean up sgl kernel (#12413)
Co-authored-by: Byron Hsu <byronhsu1230@gmail.com>
|
2025-10-31 01:13:34 -07:00 |
|
Xiaoyu Zhang
|
334543ff3b
|
Add continuous_usage_stats support for streaming responses (#12241)
|
2025-10-29 10:01:23 +08:00 |
|
Feng Su
|
ea96106000
|
[Feature] Sglang Tracing: Fine-Grained Tracking for Request Latency - Part 2 (#10804)
Signed-off-by: Feng Su <sufeng@linux.alibaba.com>
|
2025-10-28 01:25:46 -07:00 |
|
Chenxi Li
|
cc7b04a29c
|
Feature/Add GET endpoint to query loaded LoRA adapters (#12229)
|
2025-10-28 01:22:00 -07:00 |
|
satyamk7054
|
9fc3e8aac7
|
Add support for Matryoshka embeddings (#126) (#11142)
Co-authored-by: Satyam Kumar <satyamk@linkedin.com>
|
2025-10-28 02:49:36 +08:00 |
|
Chang Su
|
94aad0de99
|
[misc][grpc] Remove duplicate log (#12168)
|
2025-10-26 14:06:59 -07:00 |
|
Lianmin Zheng
|
7b36c47b3b
|
Clean up attention backend selection code & Other minor rename (#12136)
|
2025-10-25 23:50:12 -07:00 |
|
Lianmin Zheng
|
8e70064c37
|
Clean up server launch code and multi tokenizer (#12132)
|
2025-10-25 16:40:27 -07:00 |
|
Baizhou Zhang
|
4b0ac1d52a
|
Update sgl-kernel version to 0.3.16.post4 (#12125)
|
2025-10-25 14:33:33 -07:00 |
|
Teng Ma
|
96a5e4dd79
|
[Feature] Support loading weights from ckpt engine worker (#11755)
Signed-off-by: Yang Kaiyong <yangkaiyong.yky@antgroup.com>
Signed-off-by: Cruz Zhao <CruzZhao@linux.alibaba.com>
Signed-off-by: Xuchun Shang <xuchun.shang@gmail.com>
Co-authored-by: Yang Kaiyong <yangkaiyong.yky@antgroup.com>
Co-authored-by: Cruz Zhao <CruzZhao@linux.alibaba.com>
Co-authored-by: Xuchun Shang <xuchun.shang@gmail.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>
|
2025-10-23 09:23:30 -07:00 |
|
Chang Su
|
6ade6a02d4
|
[grpc] Support gRPC standard health check (#11955)
|
2025-10-22 16:59:09 -07:00 |
|
Liangsheng Yin
|
9d61205dac
|
[lint] improve ruff check (#11922)
Co-authored-by: Xiaoyu Zhang <35585791+BBuf@users.noreply.github.com>
|
2025-10-22 11:32:50 +08:00 |
|
Chang Su
|
70f6309cd4
|
[router][grpc] Support v1/responses API (#11926)
|
2025-10-21 17:41:48 -07:00 |
|
Yineng Zhang
|
9792b9d7e3
|
chore: upgrade flashinfer 0.4.1 (#11933)
|
2025-10-21 14:46:31 -07:00 |
|
Baizhou Zhang
|
ebff4ee648
|
Update sgl-kernel and remove fast hadamard depedency (#11844)
|
2025-10-21 13:13:54 -07:00 |
|
Zhengke Zhou
|
260fe755b6
|
Simplify multi-tokenizer (#11295)
Signed-off-by: zhengkezhou1 <madzhou1@gmail.com>
Co-authored-by: Liangsheng Yin <lsyincs@gmail.com>
|
2025-10-21 16:33:29 +08:00 |
|
ybyang
|
dbb16bedd5
|
Support Thinking Budget (via custom_logit_processor for OpenAI API) [Fix #6572] (#11416)
Signed-off-by: ybyang <ybyang7@iflytek.com>
Co-authored-by: YorkSu <york_su@qq.com>
|
2025-10-21 16:27:56 +08:00 |
|
Neelabh Sinha
|
852c0578fd
|
[FEATURE] Add OpenAI-Compatible LoRA Adapter Selection (#11570)
|
2025-10-21 15:44:33 +08:00 |
|
Chang Su
|
9c0b1eb5ad
|
[router][grpc] Fix wram-up random token ids for small models (#11887)
|
2025-10-20 19:22:17 -07:00 |
|