Alison Shao
|
bedade1ef0
|
Merge stage-c-test-large-4-gpu suites into partitioned suites (#18325)
|
2026-02-06 15:32:33 -08:00 |
|
Neal Vaidya
|
f1ff697494
|
add hybrid model PD to NIXL connector (#16229)
Signed-off-by: Neal Vaidya <nealv@nvidia.com>
Co-authored-by: ishandhanani <82981111+ishandhanani@users.noreply.github.com>
|
2026-02-06 15:05:36 -08:00 |
|
Rishit Shivam
|
c850a8a41a
|
[Docs] Add Falcon H1, Hunyuan-Large, Qwen3-Omni support and update Diffusion usage (#17888)
Co-authored-by: Rishitshivam <164783543+Rishitshivam@users.noreply.github.com>
Co-authored-by: Ratish P <114130421+Ratish1@users.noreply.github.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Adarsh Shirawalmath <114558126+adarshxs@users.noreply.github.com>
Co-authored-by: zhaochenyang20 <zhaochen20@outlook.com>
|
2026-02-06 13:17:51 -08:00 |
|
Baizhou Zhang
|
f2e0048d06
|
Add CI permission for Shunkangz, dongjiyingdjy, samuellees (#18377)
|
2026-02-07 01:19:02 +08:00 |
|
Prozac614
|
e13b727e92
|
[diffusion] CI: update perf baseline (#17512)
|
2026-02-07 00:28:44 +08:00 |
|
shaharmor98
|
c6aa1863be
|
Add Nemotron 3 Nano tests (#18119)
Signed-off-by: Shahar Mor <smor@nvidia.com>
|
2026-02-06 23:55:42 +08:00 |
|
xiaoye
|
79d409f210
|
[diffusion] fix: offload text encoder model in image encoding stage (#18317)
|
2026-02-06 22:55:56 +08:00 |
|
Xuchun Shang
|
3d68bd9d9b
|
add hicache jit test (#17847)
Signed-off-by: Xuchun Shang <xuchun.shang@linux.alibaba.com>
|
2026-02-06 16:54:33 +08:00 |
|
陈一涵
|
f798ab9775
|
[diffusion] fix: fix torch.compile graph break caused by torch._dynamo.disable (#18336)
|
2026-02-06 14:48:09 +08:00 |
|
Alison Shao
|
d0c39bc219
|
Fix cross-container HF download race condition in CI (#18328)
|
2026-02-05 21:01:41 -08:00 |
|
amote-i
|
92b8bd6833
|
fix npu best practice (#18330)
|
2026-02-05 21:14:46 -05:00 |
|
Linyu Wu
|
aa390d2762
|
[Kernel] Migrate GPTQ-Marlin GEMM kernel to JIT (#18067)
|
2026-02-06 08:31:42 +08:00 |
|
shuwenn
|
ef1d0ea885
|
[Doc] add a summary section for spec decode document (#18323)
|
2026-02-05 16:34:31 -05:00 |
|
shuwenn
|
8b21dd4b77
|
[Doc] refine spec decode docs for SpecV2/STANDALONE/NGRAM (#18321)
|
2026-02-05 15:12:33 -05:00 |
|
aaaandychen
|
6a4b81e2d9
|
Refactor(qwen3-vl) optimize position encoding interpolation (#16781)
Signed-off-by: chenzhenyang <andy271828@163.com>
Signed-off-by: chenzhenyang <chenzhenyang@moonshot.cn>
Co-authored-by: chenzhenyang <chenzhenyang@moonshot.cn>
Co-authored-by: zhaochenyang20 <zhaochen20@outlook.com>
|
2026-02-05 10:26:35 -08:00 |
|
Alison Shao
|
d22163eb8c
|
Fix flaky test_frequency_penalty_reduces_word_repetition by using deterministic seeds (#18285)
|
2026-02-05 10:24:18 -08:00 |
|
ovidiusm
|
498d8d0680
|
NixlKVManager optimizations (#17654)
Signed-off-by: Ovidiu Mara <ovidium@nvidia.com>
|
2026-02-06 00:25:23 +08:00 |
|
wxy
|
b639779dd8
|
[diffusion] feat: allow T5's TP Group to reuse the transformer's SP Group (#17818)
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2026-02-06 00:12:19 +08:00 |
|
Glen Liu
|
3f32a5831d
|
throw error if got adapter with added_tokens (#18046)
|
2026-02-05 23:55:43 +08:00 |
|
pansicheng
|
2eb4359ada
|
[Kernel] Add JIT apply_rope_with_cos_sin_cache_inplace (#18155)
|
2026-02-05 21:49:37 +08:00 |
|
陈一涵
|
4aa03d91fd
|
[diffusion] fix: fix accuracy bug caused by #14717 (#18296)
|
2026-02-05 20:36:18 +08:00 |
|
ishandhanani
|
8f8c1724ae
|
docker: add patch to increase GPU deepep timeout (#18298)
|
2026-02-05 18:26:15 +08:00 |
|
Shangming Cai
|
afae4c7178
|
[PD] Minor code cleanup for mooncake backend (#18279)
Signed-off-by: Shangming Cai <csmthu@gmail.com>
|
2026-02-05 17:38:09 +08:00 |
|
zhangheng
|
079fc8f3c5
|
[piecewise graph]: support MiniMax-M2 (#18217)
|
2026-02-04 23:24:38 -08:00 |
|
danielafrimi
|
3f1df322f9
|
[FIX] Always support TP > 4 for FP4 Gemm (#17300)
|
2026-02-05 15:10:26 +08:00 |
|
Meng, Hengyu
|
368936a62b
|
[XPU] Integrate MoE and minor improvements in XPU attention backend (#13561)
|
2026-02-04 23:09:59 -08:00 |
|
Xiaoyu Zhang
|
dff3ba202a
|
[Diffusion] Support layerwise offload for mova (#18272)
|
2026-02-05 13:16:07 +08:00 |
|
Alison Shao
|
c910829708
|
Fix test_return_routed_experts to use response-level sglext (#18274)
|
2026-02-04 20:16:01 -08:00 |
|
Kun Lin
|
e616d35847
|
Support Markdown/Notebook-Friendly Documentation Export for Downstream Integration(convert rat files to md files and save) (#18278)
|
2026-02-04 19:59:40 -08:00 |
|
rinbaro
|
de6a03260f
|
[docs] fix misspellings & typos (#18276)
|
2026-02-05 03:35:29 +00:00 |
|
Teng Ma
|
c8212b9fac
|
[PD] doc: Document SGLANG_MOONCAKE_CUSTOM_MEM_POOL and supported values (#18259)
Co-authored-by: Shangming Cai <csmthu@gmail.com>
|
2026-02-05 03:03:33 +00:00 |
|
Ch3ngY1
|
f730c18679
|
[PD] improve kv offset calculation for MHA model with different tp size (#18163)
Co-authored-by: Shangming Cai <csmthu@gmail.com>
|
2026-02-05 10:43:23 +08:00 |
|
Mick
|
f218234e4f
|
[diffusion] chore: prohibit Chinese characters usage (#18249)
|
2026-02-05 09:22:26 +08:00 |
|
yinghui
|
599c5f4922
|
fix kimi k2.5's moe gemm config init (#18064)
|
2026-02-04 16:59:01 -08:00 |
|
zwang86
|
bdaf3de9b3
|
fix: add SGLANG_IS_IN_CI env var to release-docs workflow (#18225)
Co-authored-by: Zeyu Wang <zeyu.wang@yahooinc.com>
Co-authored-by: edwingao28 <edwingao28@users.noreply.github.com>
|
2026-02-04 15:49:41 -08:00 |
|
linhaifeng
|
c1d5cc3b24
|
[Bugfix] fix a obvious logic error (#18254)
|
2026-02-04 13:59:58 -08:00 |
|
Mohammad Miadh Angkad
|
efbf39583e
|
Add MoE fused config for Qwen3-Coder-Next-FP8 on H100 TP=2 (#18195)
|
2026-02-04 13:36:35 -08:00 |
|
Zack Yu
|
2e87c2bd5e
|
fix: fix MockModelRunner in attention tests (#18240)
|
2026-02-04 13:18:02 -08:00 |
|
Michael
|
6fd878b41d
|
[AMD] Add kimi mi35x nightly test, folder organization and several stability fixes (#17895)
|
2026-02-04 12:03:57 -08:00 |
|
Mick
|
36a3e78af9
|
[diffusion] refactor: move model_stages into stages folder (#18248)
|
2026-02-05 00:23:31 +08:00 |
|
RunningLeon
|
3e7ecb78a6
|
model: support interns1-pro (#18145)
Co-authored-by: Ke Bao <ispobaoke@gmail.com>
|
2026-02-05 00:22:44 +08:00 |
|
RunningLeon
|
a6f53cc5e3
|
entrypoint: support passing spaces_between_special_tokens per request (#17939)
|
2026-02-04 22:18:36 +08:00 |
|
wxy
|
4c403045ec
|
[diffusion] fix: fix the bug of redundant memory usage on GPU-0 (#18221)
|
2026-02-04 21:25:23 +08:00 |
|
Zhang Yiyang (SII)
|
0c9a0adc53
|
[diffusion] chore: clean MOVA codes (#18107)
|
2026-02-04 21:23:41 +08:00 |
|
BingjiaWang
|
760ae933bb
|
optimize get_topk_ragged by fusing get k and k_scale triton kernel (#16043)
Co-authored-by: abing <wangbingjia.wbj@alibaba-inc.com>
|
2026-02-04 19:59:41 +08:00 |
|
Nicolas Castet
|
315306d8a9
|
Make sure we always disable symm memory without dp padding (#18129)
|
2026-02-04 19:58:28 +08:00 |
|
Jincong Chen
|
a72f4f839c
|
Tiny fix for fp8 moe backend flashinfer_trtllm naming (#18243)
|
2026-02-04 19:58:04 +08:00 |
|
Evrard-Nil
|
ce02df8592
|
[diffusion] logging: downgrade default prompt log from info to debug (#17813)
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2026-02-04 19:19:02 +08:00 |
|
Xiaoyu Zhang
|
2e9d0442e2
|
[diffusion] update code owner (#18247)
|
2026-02-04 19:12:32 +08:00 |
|
Cheng Wan
|
84c09913eb
|
Moving _alloc_extend_naive out of npu allocator (#18200)
|
2026-02-04 02:09:55 -08:00 |
|