Baizhou Zhang
|
0dfe46dafb
|
[Docker] Install cudnn==9.16 for cuda 13 image to avoid check error (#17668)
|
2026-01-24 11:27:03 +08:00 |
|
Lianmin Zheng
|
bc6f0b5ce7
|
[Auto Sync] Update logits_processor.py, test_logprobs.py (20260124) (#17664)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: yehu-ux <yehu@x.ai>
|
2026-01-23 17:57:41 -08:00 |
|
R0CKSTAR
|
e1833c4f5a
|
Add yeahdongcn to CI permissions (#17667)
Signed-off-by: Xiaodong Ye <xiaodong.ye@mthreads.com>
|
2026-01-23 17:56:09 -08:00 |
|
McZyWu
|
b4a611fb33
|
[NPU] solve accuracy problem for stablelm-2-1-6b for npu (#17470)
|
2026-01-24 08:27:38 +08:00 |
|
McZyWu
|
8a5ed2434f
|
[NPU]support model MiniCPM3-4B for npu (#16866)
|
2026-01-24 08:25:12 +08:00 |
|
Douglas Yang
|
4c7136bb36
|
feature: adding openai compatible API request to bench_serving (#17219)
|
2026-01-23 16:04:28 -08:00 |
|
Nan Jiang
|
ad05782160
|
fix post_residual_addition more generally (#17286)
|
2026-01-23 15:43:37 -08:00 |
|
R0CKSTAR
|
628ab5d57b
|
[MUSA][2/N] sgl-kernel build (#17053)
Signed-off-by: Xiaodong Ye <xiaodong.ye@mthreads.com>
|
2026-01-23 14:41:47 -08:00 |
|
R0CKSTAR
|
a77729a276
|
[MUSA][1/N] sglang.check_env (#16959)
Signed-off-by: Xiaodong Ye <xiaodong.ye@mthreads.com>
|
2026-01-23 14:41:17 -08:00 |
|
Mansoor
|
bdaa3de075
|
Add return routed experts to the completions and chat/completions endpoints (#17434)
|
2026-01-23 12:12:36 -08:00 |
|
Tiwei Bie
|
5438cd20ce
|
[DLLM] Remove cuda graph batch size limitation (#17458)
|
2026-01-23 09:52:39 -08:00 |
|
Jerry Ji
|
010c17a133
|
[Refactor] Algebraic data type for nextn config + some basic refactors (#17347)
|
2026-01-24 01:16:55 +08:00 |
|
Yi Zhong
|
08fcda2f63
|
add the fa4 mm backend and varlen func (#13539)
Signed-off-by: vincentzed <207368749+vincentzed@users.noreply.github.com>
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
|
2026-01-23 23:12:06 +08:00 |
|
akhilg-nv
|
2fb328109f
|
[DeepSeek V3.2] Enable trtllm NSA with bf16 kvcache (#16758)
Co-authored-by: DarkSharpness <76582120+DarkSharpness@users.noreply.github.com>
|
2026-01-23 20:26:21 +08:00 |
|
Nicolas Castet
|
48e9daadff
|
Support symmetric memory pre-allocation to avoid fragmentation (#17089)
|
2026-01-23 17:57:04 +08:00 |
|
Alison Shao
|
b23470e95a
|
Fix CI install failure when rerunning tests via workflow_dispatch (#17612)
|
2026-01-23 00:04:16 -08:00 |
|
siyu
|
17349168bb
|
set cooldown_interval_minutes to 0 for liusy58 (#17637)
|
2026-01-22 23:36:05 -08:00 |
|
Yuzhen Zhou
|
2169025b77
|
turn off dit_layerwise_offload for wan on rocm (#17569)
|
2026-01-23 15:22:42 +08:00 |
|
Hexq0210
|
69470dbc1f
|
[NPU] update doc for Ascend NPU (#17621)
|
2026-01-23 15:21:00 +08:00 |
|
Even Zhou
|
69ac8b58f7
|
[NPU] [CI] temporarily disable mtp test (#17614)
|
2026-01-23 15:17:31 +08:00 |
|
Minglei Zhu
|
2b2f317383
|
fix gpt-oss launch failure with piecewise cuda graph (#17532)
|
2026-01-22 22:41:38 -08:00 |
|
Alison Shao
|
d7dd0b8832
|
Re-enable unit-test-deepep-8-gpu and unit-test-backend-4-gpu-gb200 (#17438)
|
2026-01-23 14:31:44 +08:00 |
|
Lianmin Zheng
|
56e6652d1d
|
Lazy import torchao (#17626)
|
2026-01-22 22:04:51 -08:00 |
|
Bingxu Chen
|
50a2e4345a
|
[AMD CI] Add 2-GPU sgl-kernel Tests (#17555)
Co-authored-by: YC Tseng <yctseng@amd.com>
|
2026-01-22 21:48:52 -08:00 |
|
JiaruiChang5268
|
c0b5a180fe
|
[NPU]bugfix: fix for dsv3.2 and dsvl2 (#17007)
Co-authored-by: Hexq0210 <893781835@qq.com>
Co-authored-by: liupeng374 <782420244@qq.com>
Co-authored-by: cy <chenyang08056032@163.com>
|
2026-01-23 11:15:15 +08:00 |
|
Ke Bao
|
7ace64d1d8
|
Update mamba env setting (#17566)
|
2026-01-23 11:02:32 +08:00 |
|
siyu
|
62e6a749b0
|
Skip mm feature pool init to avoid EPD OOM (#16388)
|
2026-01-23 10:53:45 +08:00 |
|
YC Tseng
|
04a10c9bc2
|
[AMD] CI - migrate perf test and fix stage-b-test-1-gpu-amd (#17340)
Co-authored-by: Bingxu Chen <bingxche@amd.com>
Co-authored-by: bingxche <Bingxu.Chen@amd.com>
Co-authored-by: michaelzhang-ai <michaelzhang.ai@users.noreply.github.com>
|
2026-01-22 18:45:05 -08:00 |
|
amote-i
|
0fec8820d1
|
update dependence docs of npu (#17573)
|
2026-01-23 09:13:36 +08:00 |
|
Alison Shao
|
1e8e0cca2c
|
Update test README with CI registry documentation and 5090/H100 guidance (#17368)
|
2026-01-22 16:29:08 -08:00 |
|
MMuzzammil1
|
2399af5557
|
Bugfix: Writing to storage when write-back method is chosen (#14718)
|
2026-01-22 15:08:25 -08:00 |
|
hxie
|
13f88045b3
|
configuration file support and nixl integration augmentation for hicache-storage-backend-extra-config (#16602)
|
2026-01-22 14:31:48 -08:00 |
|
Jacob Gordon
|
a296c99ff4
|
refactor(benchmark): prevents variable shadowing (#17607)
|
2026-01-22 17:00:11 -05:00 |
|
wufann
|
a921029b97
|
[AMD] Support ds3.2 on gfx942 platform (#17504)
Co-authored-by: Hubert Lu <55214931+hubertlu-tw@users.noreply.github.com>
|
2026-01-22 13:57:08 -08:00 |
|
Jacob Gordon
|
15b511771d
|
refactor(codespell): corrects typos covered up by whitelist (#17601)
|
2026-01-22 11:45:47 -08:00 |
|
Jacob Gordon
|
2c86a065db
|
types: preps SessionReqNode for IDE-driven renames (#17602)
|
2026-01-22 11:36:23 -08:00 |
|
Baizhou Zhang
|
283a2daeaa
|
[hotfix] Reenable all reduce fusion on sm100 (#17591)
|
2026-01-22 23:36:38 +08:00 |
|
Yuhao Yang
|
f7a0bcda1e
|
model: step3-vl-10b (#17513)
|
2026-01-22 23:15:08 +08:00 |
|
Jincong Chen
|
72f790bf6f
|
[BUGFIX] Skip Mamba Cache Slot 0 to Avoid Using Dummy Cache (#17404)
|
2026-01-22 23:08:38 +08:00 |
|
Jacob Gordon
|
42523d0364
|
ci(pre-commit): avoids extraneous codespell exclusions (#17590)
|
2026-01-22 09:55:23 -05:00 |
|
Xiaoyu Zhang
|
5324027007
|
[Diffusion] Make the apply_qknorm function easier to use (#17537)
|
2026-01-22 22:32:15 +08:00 |
|
triple-mu
|
3705f90629
|
[diffusion] model: optimize torch.compile (#17472)
|
2026-01-22 22:05:47 +08:00 |
|
chenxu214
|
5d299c25c0
|
[NPU] bugfix with Kimi-k2 and bge-reranker-v2 model (#17478)
Co-authored-by: amote-i <49533125+amote-i@users.noreply.github.com>
Co-authored-by: cy <chenyang08056032@163.com>
|
2026-01-22 22:02:05 +08:00 |
|
chenxu214
|
a4dc432587
|
Change naming for graph mode on multiplatform (#17469)
|
2026-01-22 19:25:06 +08:00 |
|
Minglei Zhu
|
419bbcee10
|
refactor Qwen3-Next with a new RadixLinearAttention (#17373)
|
2026-01-22 17:42:06 +08:00 |
|
zhangheng
|
f33022d039
|
[RadixTree][3/N Refactor]:Support unified insert/evict params (#17401)
|
2026-01-22 17:36:31 +08:00 |
|
Shangming Cai
|
2262c5c9b5
|
Add ZhengWG to CI_Permission (#17572)
|
2026-01-22 17:24:06 +08:00 |
|
Baizhou Zhang
|
8dae6ec03c
|
Add xyjixyjixyji to CI_Permission (#17559)
|
2026-01-21 23:27:25 -08:00 |
|
Chandrakant Khandelwal
|
61abff66c1
|
[NPU] [Bug Fix] Fix typo in npu device check in gpt_oss.py (#17553)
|
2026-01-21 23:26:35 -08:00 |
|
Zaili Wang
|
6a8f68b6d2
|
[Fix] fix device orientation for image processor (#15859)
|
2026-01-22 15:03:10 +08:00 |
|