jiashaokun-1
|
0fe282543f
|
[NPU] support the Enable return routed experts (#17025)
|
2026-02-01 18:31:39 +08:00 |
|
Estrella-xx
|
27bec34203
|
[NPU] disaggregation_decode_enable_fake_auto parameter adaptation (#17811)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
Co-authored-by: Even Zhou <even.y.zhou@outlook.com>
|
2026-02-01 18:14:26 +08:00 |
|
Ke Bao
|
d9050b4a9c
|
Reset evict swa status when retract (#18059)
|
2026-02-01 17:17:37 +08:00 |
|
Alison Shao
|
56907cbcb1
|
Move deleted 8-GPU tests to test/manual/ (#18060)
|
2026-02-01 00:21:56 -08:00 |
|
sunxxuns
|
47592a23c7
|
[CI] Fix AMD CI by inlining dummy_grok config (#18044)
Co-authored-by: root <root@mi300x8-005.atl1.do.cpe.ice.amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
|
2026-02-01 00:20:57 -08:00 |
|
yefei12
|
855dd0546c
|
feat: Add Ling Flash v2.0 support for Eagle3 (#15119)
Co-authored-by: chenyefei.cyf <chenyefei.cyf@U-9V5T77LW-2356.local>
Co-authored-by: GeLee-Q <865038696@qq.com>
Co-authored-by: Gao016 <yngao016@163.com>
Co-authored-by: Shenggui Li <somerlee.9@gmail.com>
Co-authored-by: Yineng Zhang <me@zhyncs.com>
|
2026-01-31 23:57:45 -08:00 |
|
lukec
|
3ca29dffc7
|
support qwen3-next eagle3 (#14607)
|
2026-01-31 23:45:23 -08:00 |
|
Praneth Paruchuri
|
9bb1260558
|
[Feature] Support file:// URL format for multimodal inputs (#14490)
Co-authored-by: Yuhao Yang <47235274+yhyang201@users.noreply.github.com>
|
2026-01-31 23:44:44 -08:00 |
|
ZhenshengWu
|
71babdef51
|
Fix CUDA 12 dependency when importing Mooncake in official CUDA 13.x image (#17540)
Co-authored-by: wuzhensheng01 <wuzhensheng01@baidu.com>
|
2026-01-31 23:41:21 -08:00 |
|
Roger Young
|
486c7de39f
|
Optimizing all_reduce in RMSNormTP in minimax_m2 (#16483)
|
2026-01-31 23:39:37 -08:00 |
|
Kangyan-Zhou
|
9c168fcac7
|
Fix Diffusion Request Validation to allow missing input artifacts if the input only contains text (#16610)
|
2026-01-31 23:38:40 -08:00 |
|
tc-mb
|
4d28cda007
|
[model] Support MiniCPM-V 4.5 (#9610)
Signed-off-by: tc-mb <caitianchi@modelbest.cn>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
|
2026-01-31 23:37:36 -08:00 |
|
linhaifeng
|
2c036f1eb1
|
[Bugfix] fix the display error (inconsistent context) (#17699)
Signed-off-by: linhaifeng <1371675203@qq.com>
|
2026-01-31 23:35:11 -08:00 |
|
Kangyan-Zhou
|
e5ac6229e1
|
Fix installation script for H200 runners (#18050)
|
2026-01-31 23:30:51 -08:00 |
|
Alison Shao
|
a0bae4c343
|
Migrate 4-GPU/8-GPU workflow jobs to stage-c and add CI registry decorators (#17299)
|
2026-01-31 22:37:22 -08:00 |
|
Alison Shao
|
95180484e9
|
Disable test_mla_int8_deepseek_v3.py temporarily (#18057)
|
2026-01-31 22:33:43 -08:00 |
|
Ke Bao
|
d396650bd2
|
Fix swa kv cache memory allocation (#18039)
|
2026-02-01 14:26:51 +08:00 |
|
Minglei Zhu
|
38d275a9fd
|
[BugFix] fix gpt-oss accuracy issue when enabling piecewise cuda graph (#18013)
|
2026-02-01 14:26:26 +08:00 |
|
Yinghai Lu
|
11892599f1
|
[metric] Optional extra metric labels (#18049)
|
2026-02-01 14:25:59 +08:00 |
|
Baizhou Zhang
|
c7d53fa26a
|
Set torch url index in pyproject.toml (#16802)
|
2026-02-01 13:23:52 +08:00 |
|
khalilzhk
|
429ef988bc
|
[BugFix] Fix draft model specified config file (#17815)
|
2026-01-31 20:45:36 -08:00 |
|
Kangyan-Zhou
|
e884b17632
|
Fix rerun stage command with merged commit history (#17960)
|
2026-01-31 20:37:55 -08:00 |
|
Kaixi
|
2b2515423a
|
Skipped warning on sm100 (#18000)
|
2026-01-31 20:21:03 -08:00 |
|
Kangyan-Zhou
|
d443d2d2ae
|
Improve error output in tnightly tets (#18053)
|
2026-01-31 19:26:19 -08:00 |
|
Yingchun Lai
|
0f2df9370a
|
feat: validate ib devices in server args (#17598)
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
Co-authored-by: Kangyan Zhou <zky314343421@gmail.com>
|
2026-01-31 17:50:42 -08:00 |
|
b8zhong
|
398d13a189
|
[Perf] Add Flashinfer DeepGEMM SM90 for SwapAB Optimization (#15514)
Co-authored-by: Brayden Zhong <b8zhong@users.noreply.github.com>
|
2026-02-01 08:56:23 +08:00 |
|
Chongchong Tian
|
9951a1ae07
|
Fix: Remove duplicate assignment for use_w4afp8 (#17858)
|
2026-01-31 16:04:08 -08:00 |
|
Yi Zhong
|
c2ab3713e9
|
[Performance] Optimize Mllama LayerNorm -> Upd (#9725)
|
2026-01-31 16:02:57 -08:00 |
|
Hao Jin
|
6aaea09b3d
|
Update python/sglang/README.md (#18045)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2026-01-31 13:20:41 -08:00 |
|
Zaili Wang
|
97593c9f41
|
[CPU] toml file update (#17861)
|
2026-01-31 13:16:06 -08:00 |
|
Mohammad Miadh Angkad
|
9ac4dcada4
|
[Tiny] Fix grammar in shared experts fusion log messages (#18043)
|
2026-01-31 13:14:25 -08:00 |
|
R0CKSTAR
|
46095f0551
|
[MUSA] Update 3rd party dir to build/_deps (#18035)
Signed-off-by: Xiaodong Ye <xiaodong.ye@mthreads.com>
|
2026-01-31 12:02:39 -08:00 |
|
b8zhong
|
ef134d407d
|
[Fix] Revert back to using CUTLASS mm_fp4 backend (#17369)
|
2026-01-31 23:01:29 +08:00 |
|
Mick
|
1a006c2a0d
|
[diffusion] refactor: split component_loader into component-wise files (#17820)
|
2026-01-31 20:22:31 +08:00 |
|
Lianmin Zheng
|
7412ceb4eb
|
[Auto Sync] Update linear.py to assert shapes (20260130) (#17966)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Byron Hsu <byronhsu1230@gmail.com>
|
2026-01-31 01:01:55 -08:00 |
|
Lianmin Zheng
|
0e184609d3
|
Add launch_command assignment in crash dump (#17967)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Archit Patke <apatke@x.ai>
|
2026-01-31 01:00:40 -08:00 |
|
b8zhong
|
22498e10c0
|
[Fix] Triton TP MoE Dpsk V3/Qwen3 Coder with SwapAB (#17965)
|
2026-01-31 15:56:26 +08:00 |
|
Zheng Wengang
|
a4df95c15f
|
[EPD][Perf] parallelize ZMQ send for encode server (#16487)
Co-authored-by: siyu <liusy58@linux.alibaba.com>
|
2026-01-31 14:30:11 +08:00 |
|
jeff
|
04efd03dbf
|
Fix OOM in DeepSeek weight loading by deferring dict(weights) materialization (#17744)
|
2026-01-31 13:59:00 +08:00 |
|
Yifan Cui
|
45fe51a28e
|
Reduce topk kernel shared memory from 128KB to 32KB for better occupancy (#17747)
Co-authored-by: Claude <noreply@anthropic.com>
|
2026-01-30 21:42:21 -08:00 |
|
Hudson Xing
|
c72bf50706
|
add reasoning_tokens usage test for tool call (#18022)
|
2026-01-30 21:09:23 -08:00 |
|
husf
|
c52578c7fd
|
【docs】【NPU】Update Expert Parallelism docs for Ascend NPU (#17940)
|
2026-01-30 23:42:11 -05:00 |
|
R0CKSTAR
|
dc77defdd0
|
Fix .gitignore may ignore files like core_attention.py (#18021)
Signed-off-by: Xiaodong Ye <xiaodong.ye@mthreads.com>
|
2026-01-30 20:33:32 -08:00 |
|
Mohammad Miadh Angkad
|
d0d9cecd1b
|
Fix cuBLAS >=12.9 detection for cu12/cu13 package naming (#17766)
|
2026-01-31 12:01:52 +08:00 |
|
Xiaoyu Zhang
|
22aad4e2c4
|
[Diffusion] Fix FLUX.1-schnell time embedding argument mismatch (#17988)
|
2026-01-31 11:47:27 +08:00 |
|
kk
|
ec76c390c9
|
Add ROCm + Mori docker build instructions in rocm.Dockerfile (#18018)
Co-authored-by: wunhuang <wunhuang@amd.com>
|
2026-01-30 19:12:07 -08:00 |
|
22dimensions
|
ee3058c6e8
|
[NPU] fix sgl-kernel-npu package url error in npu.Dockerfile (#18017)
Signed-off-by: 22dimensions <waitingwind@foxmail.com>
|
2026-01-31 11:06:58 +08:00 |
|
Tiance Wang
|
f6a4ff718f
|
doc update for CANN version (#18014)
Co-authored-by: wangtiance <tiancew@qq.com>
|
2026-01-30 21:10:23 -05:00 |
|
Bi Xue
|
5d00150e99
|
[sglang] fix mm token padded value overlap with text token id (#17781)
|
2026-01-30 17:09:13 -08:00 |
|
JiaruiChang5268
|
e86476acfc
|
[NPU] support llama-3.2-11B-vision-instruct mode for NPU (#17492)
Co-authored-by: McZyWu <zhuoyun.wu.23@ucl.ac.uk>
Co-authored-by: chenyang08056032 <chenyang08056032@163.com>
Co-authored-by: Hexq0210 <893781835@qq.com>
|
2026-01-31 08:49:38 +08:00 |
|