Minglei Zhu
|
fafaa2ccea
|
[BugFix] fix outplace_fused_experts missing is_gated (#13864)
|
2025-11-24 12:36:14 -08:00 |
|
Simo Lin
|
9b4b344115
|
[model-gateway] add grpc server code owner (#13865)
|
2025-11-24 12:29:57 -08:00 |
|
fzyzcjy
|
94216a9cc4
|
Fix quantized moe checker fail for Qwen3 dense fp8 model (#13853)
|
2025-11-24 11:16:50 -08:00 |
|
Yuhao Yao
|
9535015d05
|
[Perf] Optimize DeepSeek-R1 w4afp8 glue kernels (#10027)
Co-authored-by: Fan Yin <1106310035@qq.com>
|
2025-11-24 11:05:38 -08:00 |
|
Xinyue Zhang
|
a3b578fc60
|
[model-gateway] Refactor router e2e responses tests (#13745)
Co-authored-by: Chang Su <chang.s.su@oracle.com>
Co-authored-by: Simo Lin <linsimo.mark@gmail.com>
|
2025-11-24 10:58:53 -08:00 |
|
Liangsheng Yin
|
b60e769d0e
|
Tiny unpin uvloop for other backends (#13858)
|
2025-11-25 01:34:57 +08:00 |
|
Zhi Yiliu
|
a95a38078b
|
[Fix] Fix uvloop get_event_loop() is not suitable for 0.22.x (#13612)
Signed-off-by: lzy <tomlzy213@gmail.com>
Co-authored-by: lzy <tomlzy213@gmail.com>
|
2025-11-25 01:20:00 +08:00 |
|
YAMY
|
98b38de3f2
|
Fix: Safe RoPE Cache Expansion to Prevent Position-ID Out-of-Bounds in EAGLE + Long-Sequence Workloads (#11871)
|
2025-11-25 01:19:06 +08:00 |
|
ant-yy
|
a146f833f1
|
[Fix]: Adjust FutureMap's token_id_bufs Size to Prevent ChunkedPrefill's next_token_ids from Overwriting Previous Prefill Requests' next_token_id (#13713)
Signed-off-by: vito.yy <vito.yy@antgroup.com>
|
2025-11-25 01:08:52 +08:00 |
|
StonyPort
|
1dd9a6ae4d
|
Fix TorchAO quant in VLM (#13508)
Co-authored-by: qiuxuan.lzw <qiuxuan.lzw@alibaba-inc.com>
|
2025-11-24 22:15:40 +08:00 |
|
Yuan Luo
|
8ef11569a2
|
[VLM] Revise InternVL Piecewise CUDA Graph Supporting (#13846)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
|
2025-11-24 22:15:10 +08:00 |
|
Xiaoyu Zhang
|
ecefc7904f
|
[sgl-kernel Code Clean] Remove useless lightning_attention kernel (#13819)
|
2025-11-24 18:26:25 +08:00 |
|
gaopengff
|
aeac622058
|
[Intel XPU]support xgrammar backend for intel xpu (#13245)
|
2025-11-24 16:48:00 +08:00 |
|
Baizhou Zhang
|
04b52fa8d6
|
[chore]Upgrade flashinfer to 0.5.3 (#13751)
|
2025-11-23 23:38:36 -08:00 |
|
Yuhao Yang
|
e5c0f59133
|
[diffusion] CI: send nightly-test outputs of diffusion to slack for correctness monitoring (#13833)
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2025-11-24 15:22:26 +08:00 |
|
Liangsheng Yin
|
981ca8313f
|
[misc] Rename minilb install env & remove files & fix lint (#13831)
|
2025-11-24 13:50:00 +08:00 |
|
Mick
|
414248e0d0
|
[diffusion] doc: minor update contributing.md with test section (#13792)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2025-11-24 13:00:08 +08:00 |
|
Yuan Luo
|
f56b9b42e6
|
[Bugfix] Add jit kernel files in packaging (#13829)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
Co-authored-by: Xu Yongfei <xuyongfei.xyf@antgroup.com>
|
2025-11-24 12:32:16 +08:00 |
|
Liangsheng Yin
|
b2f7b08c49
|
Refactor cache init logic (#13800)
|
2025-11-24 11:41:46 +08:00 |
|
Tiance Wang
|
75222bfed9
|
Update MindSpore documentation (#13656)
Co-authored-by: wangtiance <tiancew@qq.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2025-11-24 11:20:51 +08:00 |
|
Xiaoyu Zhang
|
9ea1953331
|
[Doc] Refine fused_moe_triton configs doc (#13820)
|
2025-11-23 19:09:41 -08:00 |
|
liuhuijiayou
|
dbf22152d6
|
Fix bug: Incorrect variable used in rem_total_token_offset calculatio… (#13201)
|
2025-11-24 11:04:34 +08:00 |
|
Swipe4057
|
d5e0346847
|
xgrammar up version to 0.1.27 (#13650)
|
2025-11-24 10:53:45 +08:00 |
|
sglang-bot
|
a22104a676
|
chore: bump sgl-kernel version to 0.3.18 (#13816)
|
2025-11-23 17:24:54 -08:00 |
|
Baizhou Zhang
|
4683e244fe
|
[1/2] Refactor DeepGeem requant for FP8 Linear on Blackwell (#13601)
Co-authored-by: fy1214
|
2025-11-23 16:07:56 -08:00 |
|
Simo Lin
|
9054e844ea
|
remove package json which is not used (#13810)
|
2025-11-23 15:00:30 -08:00 |
|
yinghui
|
18403f6bfe
|
make trtllm attn backend's init_forward_metadat non blocking (#13802)
|
2025-11-23 13:35:40 -08:00 |
|
Qiaolin Yu
|
2892265d4c
|
Tune fp8_w8a8 fused triton moe for GLM-4.6-FP8 (#13815)
|
2025-11-23 13:29:54 -08:00 |
|
Baizhou Zhang
|
c9bd1aca32
|
[CI] Tiny refactoring sgl-kernel tests (#13813)
|
2025-11-23 12:45:17 -08:00 |
|
hlu1
|
618ca23802
|
[Deepseek] Refactor deepseek server_args _handle_model_specific_adjustments (#13687)
|
2025-11-23 12:41:14 -08:00 |
|
Liangsheng Yin
|
5c2915494c
|
[Scheduler] Tiny organize code style (#13806)
|
2025-11-23 23:41:34 +08:00 |
|
alisonshao
|
aaa40a9b3b
|
Fix pagination bug in CI monitor preventing performance-test-2-gpu data collection (#13781)
|
2025-11-23 22:02:30 +08:00 |
|
Mick
|
dd70cf99c1
|
[diffusion] CI: add run_suite to multimodal_gen CI (#13791)
|
2025-11-23 21:27:33 +08:00 |
|
Mick
|
d4593964fe
|
[diffusion] feat: support sp for image models (#13180)
|
2025-11-23 18:11:42 +08:00 |
|
YAMY
|
53fffefd5d
|
Upgrade flashmla kernel for NSA tp support (#13718)
|
2025-11-23 01:36:49 -08:00 |
|
Xiaoyu Zhang
|
b964ce61d6
|
[DeepEP] Add SGLANG_DEEPEP_BF16_DISPATCH env var in Normal mode (#13787)
|
2025-11-23 17:32:43 +08:00 |
|
DarkSharpness
|
ac5505b04c
|
[Feature] HiCache JIT kernel (once again) (#13764)
|
2025-11-22 22:19:16 -08:00 |
|
Peiqi Yin
|
a90435c059
|
Fix typo in docs (#13709)
|
2025-11-23 10:49:49 +08:00 |
|
Simo Lin
|
5354d7b7fd
|
[model-gateway] clean up router manager function order (#13776)
|
2025-11-22 16:16:20 -08:00 |
|
Simo Lin
|
e01486778a
|
[model-gateway] update smg code owner (#13777)
|
2025-11-22 15:50:53 -08:00 |
|
ErsongWang
|
047935080d
|
Revert "Fix RMSNorm API CALL mismatch issue. (#10032)" (#13727)
|
2025-11-22 14:25:52 -08:00 |
|
cctry
|
cad7878964
|
Gather static input buffers for cuda graph (#13676)
|
2025-11-22 13:01:50 -08:00 |
|
Binyao Jiang
|
b29769f3b6
|
Move unnecessary input_addr capture under debug mode flag for speed-up (#13690)
|
2025-11-22 11:42:26 -08:00 |
|
Ho-Ren (Jack) Chuang
|
3990b84bd3
|
Refactor MHA & MLA KV caches to support FP4 (#13547)
Signed-off-by: Ho-Ren (Jack) Chuang <horenchuang@bytedance.com>
|
2025-11-22 11:13:43 -08:00 |
|
yinghui
|
5a4394a342
|
align code style eagle draft&draft_extend cuda graph runner (#13533)
|
2025-11-23 01:55:16 +08:00 |
|
Yuhao Yang
|
dd303614e0
|
[diffusion] CI: tinyfix diffusion ci (#13769)
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2025-11-23 01:03:36 +08:00 |
|
Liangsheng Yin
|
863124684c
|
[Spec v2] Remove allocate_lens and enable over-allocation (#13478)
|
2025-11-22 22:49:10 +08:00 |
|
Yuan Luo
|
5625e32cae
|
[VLM] Replace torch.repeat_interleave with faster np.repeat for Qwen-VL series (#13736)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
|
2025-11-22 22:45:32 +08:00 |
|
Mick
|
ca548d8324
|
[diffusion] refactor: refactor sampling params (#13706)
|
2025-11-22 22:43:04 +08:00 |
|
Liangsheng Yin
|
3397bcee84
|
Tiny support different prompts in send_one.py (#13768)
|
2025-11-22 21:59:58 +08:00 |
|