jacky.cheng
|
eff7df6d0a
|
[AMD] Enable fused shared expert append and flatten quant for fp8 deepseekR1 model (#13705)
Co-authored-by: yctseng0211 <yctseng@amd.com>
|
2025-11-21 02:48:28 -08:00 |
|
Mick
|
5e7f91d451
|
[diffusion] profile: support performance metric dumping and comparison (#13630)
|
2025-11-21 18:47:16 +08:00 |
|
Xiaoyu Zhang
|
a34d3abb54
|
[Clean code] Compressed_tensors_moe code clean (#13719)
|
2025-11-21 18:15:46 +08:00 |
|
Cheng Wan
|
6d0e0b9bfc
|
[11/N] MoE Refactor: Simplifying SBO Implementation with Dispatcher Hooks (#13327)
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
|
2025-11-21 01:11:37 -08:00 |
|
Even Zhou
|
589d9ad55b
|
[NPU] chore: bump to CANN 8.3.RC1 and Pytorch 2.8.0 (#13647)
|
2025-11-21 17:07:08 +08:00 |
|
Ziming Huang
|
a244c0309c
|
[Fix] Qwen3Next lmhead dtype (#13708)
|
2025-11-21 16:01:33 +08:00 |
|
Teng Ma
|
1bb063aac8
|
[HiCache] fix unit test with changed new APIs (#13498)
|
2025-11-21 15:40:31 +08:00 |
|
Michele Marzollo
|
b30f63c40f
|
[Bugfix] Fix hidden state size in EAGLE PD disaggregation buffers (#13590)
Co-authored-by: ZeldaHuang <hzm414167@alibaba-inc.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>
|
2025-11-21 15:32:40 +08:00 |
|
Lianmin Zheng
|
90a0133515
|
[Auto Sync] Update http_server.py, io_struct.py, scheduler_... (20251120) (#13679)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Zhuqi Li <zhli@x.ai>
|
2025-11-20 23:17:15 -08:00 |
|
zyksir
|
4360279036
|
[diffusion] server: use meta to avoid Linear init for TextEncoder (#13564)
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2025-11-21 15:14:24 +08:00 |
|
ishandhanani
|
b537ac0d1f
|
Fix ZMQ bind error on non-zero rank nodes when using SGLANG_BLOCK_NONZERO_RANK_CHILDREN=0 (#13686)
|
2025-11-20 22:54:41 -08:00 |
|
Mick
|
eda2f70033
|
[diffusion] doc: minor update docs (#13177)
|
2025-11-21 14:35:29 +08:00 |
|
Yuhao Yang
|
8c212a2029
|
[difusion] CI: speed up multimodal_gen ci (#13665)
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2025-11-21 14:26:33 +08:00 |
|
Stefan He
|
d754ce973e
|
[Piecewise Cuda Graph] rename, refactor and add more logging (#13675)
Co-authored-by: Minglei Zhu <mingleizhu1122@gmail.com>
Co-authored-by: Ke Bao <ISPObaoke@163.com>
Co-authored-by: Oasis-Git <ayw.sirius19@gmail.com>
|
2025-11-21 13:28:38 +08:00 |
|
Yuan Luo
|
475962a139
|
[VLM] Support Piecewise CUDA Graph for InternVL (#13640)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
|
2025-11-21 13:21:51 +08:00 |
|
Xiaoyu Zhang
|
bfcf15a129
|
[opt kimi k2 4 / n] Delete useless pad kernel in sgl_moe_align_block_size (#13587)
|
2025-11-21 13:16:42 +08:00 |
|
Xiaoyu Zhang
|
fb04d43428
|
[kimi k2 thinking] Avoid useless torch.zeros_ (#13596)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
|
2025-11-21 13:15:27 +08:00 |
|
Mike Qiu
|
6be65ae462
|
Fix target MLA with eagle3 support for PD disaggregation (#13555)
Signed-off-by: Michael Qiu <qiudayu.qdy@antgroup.com>
Co-authored-by: Mike_Qiu <qiudayu.qdy@antgroup.com>
|
2025-11-21 12:16:05 +08:00 |
|
Glen Liu
|
750084ae08
|
remove unnecessary starvation check (#13619)
|
2025-11-20 19:10:51 -08:00 |
|
alisonshao
|
64480ec712
|
Add sgl-kernel CI test for Blackwell (B200) (#13301)
|
2025-11-20 19:02:42 -08:00 |
|
Kaixi Hou
|
db2d362d04
|
[NVIDIA] Add cutedsl e2e test to GB200 CI (#12672)
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
|
2025-11-20 18:30:12 -08:00 |
|
alisonshao
|
81e86992cd
|
[CI] Move nightly tests to test/nightly/ (#13683)
|
2025-11-20 18:00:02 -08:00 |
|
Chang Su
|
c4db77f8a9
|
[model-gateway] fix gateway cli arg parser to not use = (#13685)
|
2025-11-20 17:57:17 -08:00 |
|
Mick
|
c0a2513b07
|
[diffusion] CI: improve validation method (#13627)
|
2025-11-21 09:13:13 +08:00 |
|
Simo Lin
|
3ae664d786
|
[model-gateway] add both python and rust cli alias (#13678)
|
2025-11-20 17:00:28 -08:00 |
|
赵晨阳
|
c56fc42430
|
Update quantization.md with new model resources (#13677)
|
2025-11-20 15:50:16 -08:00 |
|
fzyzcjy
|
3f1cfd87b6
|
Super tiny remove unused MiniMaxM2MLP class (#13659)
|
2025-11-21 07:35:12 +08:00 |
|
Stefan He
|
b5344b31b8
|
[Piecewise CUDA Graph] Fix recompile issue for Mixtral and Grok2 (#13667)
Co-authored-by: Minglei Zhu <mingleizhu1122@gmail.com>
Co-authored-by: Ke Bao <ISPObaoke@163.com>
Co-authored-by: Oasis-Git <ayw.sirius19@gmail.com>
|
2025-11-20 14:20:11 -08:00 |
|
alisonshao
|
6b262ac839
|
Test reorganization: Move tests to manual/ (#13610)
|
2025-11-20 13:41:58 -08:00 |
|
Glen Liu
|
ada8ce1fd0
|
allow loras to be implicitly evicted and loaded based on max_loaded_loras (#11526)
|
2025-11-20 13:34:32 -08:00 |
|
alisonshao
|
5a2c70396e
|
Add nightly test CI monitor workflow (#13038)
|
2025-11-20 12:55:13 -08:00 |
|
b8zhong
|
42028af614
|
enable csgmv automatically on cuda (#13600)
|
2025-11-20 12:53:02 -08:00 |
|
hlu1
|
7291c72e57
|
[Deepseek V3.2] Change indexer weights_proj to fp32 (#13459)
|
2025-11-20 12:24:10 -08:00 |
|
Vincent Zhong
|
6bc3062894
|
Fix launch of Olmo3 (#13666)
Signed-off-by: Vincent Zhong <207368749+vincentzed@users.noreply.github.com>
|
2025-11-20 12:04:30 -08:00 |
|
YAMY
|
fa92441027
|
[DeepseekV3.2] Deepseek fp8 support for MHA path (#12964)
|
2025-11-20 11:13:36 -08:00 |
|
Douglas Yang
|
fc9efdcb98
|
Adding nightly tests as release guard for bot bump workflows (#13655)
|
2025-11-20 10:04:36 -08:00 |
|
Fan Yin
|
2dec555d36
|
[10/n] decouple quantization impl from vllm dependency - fix import (#13524)
|
2025-11-21 01:45:51 +08:00 |
|
YanbingJiang
|
acde21d8d5
|
Add fused_rmsnorm_gated_cpu kernel for CPU to support Qwen3-Next (#11577)
|
2025-11-21 01:33:31 +08:00 |
|
Liangsheng Yin
|
4528cb7d40
|
[CI] apply pr-gate for XPU (#13663)
|
2025-11-20 23:07:04 +08:00 |
|
Lianmin Zheng
|
a352e833c4
|
CI: Kill zombie diffusion processes in CI & minor code style fix on rotary embedding fallback (#13637)
|
2025-11-20 21:57:01 +08:00 |
|
Liangsheng Yin
|
852eb6ce2a
|
[CI] optimize CI workflow info (#13634)
|
2025-11-20 18:12:51 +08:00 |
|
Lianmin Zheng
|
7af9b88c6c
|
Revert "[Feature] Introduce JIT Kernel in sglang (with hicache JIT kernel)" (#13644)
|
2025-11-20 02:11:12 -08:00 |
|
Lzhang-hub
|
2847e5c4b4
|
fix bench_speculative bug (#13197)
|
2025-11-20 17:09:04 +08:00 |
|
yctseng0211
|
c8ede0e93c
|
[ROCM] Optimized deepseek-r1 fp8 model with + triton_gemm_a8w8 + batch_gemm_a8w8 + fused set_mla_kv_buffer kernel (#13617)
Co-authored-by: root <root@smci355-ccs-aus-m12-17.cs-aus.dcgpu>
Co-authored-by: jacky.cheng <yichiche@amd.com>
|
2025-11-20 00:29:56 -08:00 |
|
Liangsheng Yin
|
19729f723e
|
[CI] Align metric units for CI rate limit (#13633)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2025-11-20 16:25:57 +08:00 |
|
DarkSharpness
|
b51f9bbee7
|
[Feature] Introduce JIT Kernel in sglang (with hicache JIT kernel) (#13453)
|
2025-11-20 00:03:32 -08:00 |
|
Mick
|
4a8442af1b
|
diffusion: improve baseline performance monitor (#13614)
|
2025-11-20 14:29:56 +08:00 |
|
Thomas Wang
|
7dcf910d53
|
Add support for new aiter version (AR accuracy, is_shuffled PR) (#13554)
Co-authored-by: sogalin <39478626+sogalin@users.noreply.github.com>
|
2025-11-19 22:17:57 -08:00 |
|
joesun
|
c7b37b7074
|
[diffusion] fix: remove multimodal_gen redundant get_bool_env_var func (#13583)
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2025-11-20 14:05:29 +08:00 |
|
iLeGend
|
10e0b83a4c
|
Add FP32 dtype support for RoPE - Part2 (#13328)
|
2025-11-19 21:19:53 -08:00 |
|