RoyWang
|
a1ef8e2cc0
|
[AMD] optimize Kimi K2.5 fused_moe_triton performance by tuning (#19228)
|
2026-02-26 11:50:13 -08:00 |
|
Shangming Cai
|
288300aafd
|
[PD] Tiny code cleanup for prefill info registering (#19414)
Signed-off-by: Shangming Cai <csmthu@gmail.com>
|
2026-02-27 01:35:26 +08:00 |
|
zhangheng
|
e4b708d3e9
|
[Spec V2] Support specV2 for mamba hybrid attention (#18808)
Co-authored-by: Yi Zhong <207368749+vincentzed@users.noreply.github.com>
Co-authored-by: yizhang2077 <1109276519@qq.com>
Co-authored-by: Hanming Lu <hanming@x.ai>
|
2026-02-27 00:36:01 +08:00 |
|
DefTruth
|
78d6674c45
|
[diffusion] feat: support hybrid parallelism for diffusers backend (#19405)
|
2026-02-27 00:06:08 +08:00 |
|
johnnycxm
|
5939b8912a
|
[MUSA][11/N] ci: add MUSA 4.3 kernel build and release pipeline (#18537)
Co-authored-by: ximin.chen <ximin.chen@mthreads.com>
|
2026-02-26 07:45:12 -08:00 |
|
Shangming Cai
|
e55e65535e
|
[Bugfix] Add rids to the batch filtering for two batch overlap (#19418)
|
2026-02-26 06:57:25 -08:00 |
|
Shangming Cai
|
97f1fa5e6b
|
[NPU] Fix disaggregation metadata buffer bootstrap_room_dtype for npu backend (#19423)
|
2026-02-26 21:10:50 +08:00 |
|
khalilzhk
|
86eb80007e
|
[NPU] support Kimi-K2.5 on NPU (#19331)
|
2026-02-26 20:41:44 +08:00 |
|
AlfredYong
|
bdc1e46e5a
|
[Qwen3.5] Qwen3.5-27B inference repeat bug fix (#19411)
|
2026-02-26 20:11:29 +08:00 |
|
Xiaoyu Zhang
|
74c8e7b215
|
refactor(jit_kernel): reduce duplication and separate test code (#19323)
|
2026-02-26 18:30:49 +08:00 |
|
Junhao Liu
|
a7152df2e3
|
[diffusion ] CLI: Fix typo in CLI usage doc string (#19316)
|
2026-02-26 13:24:14 +03:00 |
|
Shangming Cai
|
27fd014726
|
[PD] Add kv_cache_dtype consistency check for PD Disaggregation (#19407)
Signed-off-by: Shangming Cai <csmthu@gmail.com>
|
2026-02-26 17:15:58 +08:00 |
|
Alison Shao
|
e14fd4accb
|
Fix nightly Mistral-Large-3 NVFP4 accuracy threshold (#19402)
Co-authored-by: Alison Shao <alisonshao@MacBook-Pro-D2W773R9CD.local>
|
2026-02-26 15:59:11 +08:00 |
|
Feng Su
|
d2b3c7fb14
|
[Tracing] update script for converting otel tracing data to perfetto format (#19396)
|
2026-02-26 14:03:22 +08:00 |
|
Yilong Zhao
|
de3d1e7669
|
[misc] use ORJSONResponse in http-server generate (#19191)
|
2026-02-25 21:26:25 -08:00 |
|
Alison Shao
|
0fd44ff342
|
Fix NSA CP positions mismatch in eagle NextN model (#19367)
|
2026-02-25 20:14:33 -08:00 |
|
Xinyu Zhang
|
119c91cb8b
|
Skip signal handler registration when not on main thread (#18752)
|
2026-02-25 19:30:05 -08:00 |
|
Qi Yuhang
|
88ad3b894a
|
[sgl-kernel][Feat][B200][2/N] Support MXFP8 Grouped GEMM in Blackwell (#14640)
Co-authored-by: Xiaoyu Zhang <35585791+BBuf@users.noreply.github.com>
|
2026-02-26 11:23:37 +08:00 |
|
Minglei Zhu
|
b3202fe6d0
|
[PCG] fix piecewise cuda graph for Qwen3.5 (#19220)
|
2026-02-26 11:16:52 +08:00 |
|
Alison Shao
|
a0a8f1473c
|
[Benchmark] Fix generated_shared_prefix attribute naming and remove args dependency (#19363)
Co-authored-by: Alison Shao <alisonshao@Mac.attlocal.net>
Co-authored-by: sglang-bot <sglangbot@gmail.com>
|
2026-02-25 18:45:54 -08:00 |
|
sglang-bot
|
6e82183f5a
|
[Disagg] Route disagg prefill results through process_batch_result (#19364)
|
2026-02-25 18:38:39 -08:00 |
|
Xiaoyu Zhang
|
914ed34757
|
update jit_kernel codeowners (#19385)
|
2026-02-26 10:36:34 +08:00 |
|
fzyzcjy
|
265eb56d44
|
Support multi-step alignment and pipeline integration in dump comparator (#19378)
|
2026-02-26 10:23:22 +08:00 |
|
Yuan Luo
|
4e843f1216
|
[DeepSeek-V3.2][JIT-kernel] Support nsa fuse store indexer k cache (#19148)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
Co-authored-by: DarkSharpness <76582120+darksharpness@users.noreply.github.com>
|
2026-02-26 10:23:10 +08:00 |
|
Michael
|
f230967e65
|
[AMD] Fix ROCm Docker builds, update apache-tvm-ffi (#19359)
|
2026-02-26 10:16:28 +08:00 |
|
fzyzcjy
|
f9a2f0398f
|
Support token aligner planning and execution in dump comparator (#19377)
|
2026-02-26 10:04:33 +08:00 |
|
fzyzcjy
|
d34d5aca07
|
Support loading token aligner data in dump comparator (#19376)
|
2026-02-26 10:03:56 +08:00 |
|
fzyzcjy
|
e8dd14519d
|
Add aligner entrypoint and bundle handler in dump comparator (#19375)
|
2026-02-26 10:03:22 +08:00 |
|
pansicheng
|
2ad475b4ed
|
use flashinfer.sampling (#18696)
|
2026-02-26 10:02:38 +08:00 |
|
fzyzcjy
|
2739d7df62
|
Reorganize modules and pipeline in dump comparator (#19374)
|
2026-02-26 10:00:13 +08:00 |
|
fzyzcjy
|
508b8e3387
|
Handle warnings via sink for structured output and add pair in dump comparator (#19373)
|
2026-02-26 09:59:15 +08:00 |
|
fzyzcjy
|
46321ee70e
|
Support dumping rid for correlation across passes in dump comparator (#19372)
|
2026-02-26 09:57:57 +08:00 |
|
Yuan Luo
|
7c9e8e2def
|
[Re-land][jit kernel] Support per_token_group_quant_8bit jit kernel (#19140)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
Co-authored-by: Mohammad Miadh Angkad <mangkad.bsdsba2027@aim.edu>
|
2026-02-26 09:53:57 +08:00 |
|
Linyu Wu
|
beabaa8d37
|
[Kernel Slimming] Migrate marlin moe kernel to JIT (#19181)
Co-authored-by: Xiaoyu Zhang <35585791+BBuf@users.noreply.github.com>
|
2026-02-26 09:05:13 +08:00 |
|
Daniel Cámpora
|
350190487b
|
Flashinfer MOE FP8 support for Mistral Large 3. (#15422)
Co-authored-by: elvischenv <219235043+elvischenv@users.noreply.github.com>
|
2026-02-25 15:00:37 -08:00 |
|
Liangsheng Yin
|
c60dcc40bb
|
[Logging] Guard log_prefill_stats against idle batches in disagg prefill (#19361)
|
2026-02-25 13:31:52 -08:00 |
|
YAMY
|
08957c88ea
|
[Logging] Fix prefill side logging in pd disagg (#19350)
|
2026-02-25 12:42:18 -08:00 |
|
Kangyan-Zhou
|
306c552639
|
Revert "Fix HybridAttnBackend forward for linear attention" (#19356)
|
2026-02-25 11:49:50 -08:00 |
|
HAI
|
a0f3361023
|
Update (#19351)
|
2026-02-25 11:37:07 -08:00 |
|
jacky.cheng
|
b2c46fc60b
|
[AMD] Support Qwen3-Coder-Next on AMD platform (#18355)
Co-authored-by: yichiche@amd.com <jacky.cheng>
|
2026-02-25 11:06:22 -08:00 |
|
Alison Shao
|
cc1ca61c81
|
fix: add --cuda-graph-max-bs to DSV3 FA3 FP8 KV cache test (#19307)
|
2026-02-25 10:27:31 -08:00 |
|
Makcum888e
|
0217e82a08
|
[diffusion] Clean code (#19325)
|
2026-02-25 21:16:03 +03:00 |
|
Even Zhou
|
2fb239450e
|
Revert "bugfix: prioritize init_npu_backend to fix various initialization bugs" (#19343)
|
2026-02-25 23:04:30 +08:00 |
|
Yuhao Yang
|
c7c4a1cbbd
|
refactor linear attention backend (#18622)
Co-authored-by: yizhang2077 <1109276519@qq.com>
|
2026-02-25 23:02:44 +08:00 |
|
Mick
|
471acd98b9
|
[diffusion] logging: improve logging (#19312)
|
2026-02-25 23:00:35 +08:00 |
|
Hexq0210
|
56891e46bc
|
[Ascend ] Add qwen3.5 122B/35B/27B deployment examples on doc (#19339)
Co-authored-by: sglang-npu-bot <sglangnpu@163.com>
|
2026-02-25 21:09:22 +08:00 |
|
Qingfu Wen
|
59b9d1e86d
|
[diffusion] improve: improve fuse_scale_shift_kernel with non-blocking op (#18710)
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2026-02-25 21:04:20 +08:00 |
|
akhilg-nv
|
c144e55462
|
Fix HybridAttnBackend forward for linear attention (#19006)
|
2026-02-25 21:02:37 +08:00 |
|
Zheng Li
|
d38c0e537d
|
fix(dense): fix Qwen3.5 dense model precision bug in TP_SIZE>1 (#19070)
|
2026-02-25 20:54:42 +08:00 |
|
Even Zhou
|
cdc411160b
|
[NPU] Fix a corner case where FusedMoE.top_k is not explicitly declared (#19287)
|
2026-02-25 20:49:59 +08:00 |
|