Yineng Zhang
|
417e75a60f
|
fix: use 2-gpu-runner for pp 2 NemotronH ut (#16239)
|
2025-12-31 14:51:33 -08:00 |
|
ishandhanani
|
0500fea965
|
fix editable install (#16241)
|
2025-12-31 14:34:54 -08:00 |
|
Cheng Wan
|
2b461c15b4
|
Update logprob_start_len handling in scheduler (#16240)
|
2025-12-31 14:11:24 -08:00 |
|
Yineng Zhang
|
5595ae142c
|
docs: fix markdown preview (#16236)
|
2025-12-31 12:43:57 -08:00 |
|
Simo Lin
|
ace6f3005f
|
ci(benchmark): disable sccache summary annotations and fix tree output (#16233)
|
2025-12-31 11:05:21 -08:00 |
|
Kangyan-Zhou
|
2da49eec50
|
Fix XEON docker release workflow (#16107)
|
2025-12-31 09:47:01 -08:00 |
|
Kangyan-Zhou
|
d65ae0ec7a
|
Use a different concurrency group for release branch testing (#16202)
|
2025-12-31 09:46:43 -08:00 |
|
Simo Lin
|
60d7279c46
|
test(tree): add comprehensive unit tests and fix input_char_count bug (#16228)
|
2025-12-31 07:31:39 -08:00 |
|
siyu
|
abdf65d4f3
|
Fix OOM by offloading multimodal features to CPU after embedding (#16018)
|
2025-12-31 23:02:34 +08:00 |
|
Huaixin Chang
|
c1dfbc777b
|
deprecate prefill-round-robin-balance (#16195)
Signed-off-by: Chang Huaixin (OpenAnolis) <changhuaixin@linux.alibaba.com>
Co-authored-by: Liangsheng Yin <lsyincs@gmail.com>
|
2025-12-31 22:25:33 +08:00 |
|
Even Zhou
|
3b3c5a05c1
|
[CI/NPU] fix multiple NPU CI issue (#16111)
|
2025-12-31 21:14:04 +08:00 |
|
Muqi Li
|
2667c857a7
|
Fix DeepSeekV31's structural tag trigger (#13394)
|
2025-12-31 21:13:52 +08:00 |
|
Kangyan-Zhou
|
fc643ffbc9
|
Download missing shards in model weights files when not in CI (#16211)
|
2025-12-31 20:42:34 +08:00 |
|
Yuhao Yang
|
4280a18a13
|
[diffusion] CI: add test for cache-dit (#16204)
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2025-12-31 19:58:10 +08:00 |
|
Hexq0210
|
386e541520
|
Update document for Ascend NPU (#16214)
|
2025-12-31 19:03:53 +08:00 |
|
Ke Bao
|
b6267de5ae
|
Upgrade python version in lint ci (#16221)
|
2025-12-31 18:22:53 +08:00 |
|
Baizhou Zhang
|
e47afa0237
|
[DP]Fix sync bubble in adjust_num_token_non_padded_for_attn_tp (#16178)
|
2025-12-31 17:58:18 +08:00 |
|
Ethan (Yusheng) Su
|
5ed384d07b
|
[CI/CD] re-enable lora test (#16187)
|
2025-12-31 16:32:20 +08:00 |
|
Simo Lin
|
b4ce7a6d71
|
[model-gateway] cache_aware eliminate String allocations in hot path (#16209)
|
2025-12-31 00:26:10 -08:00 |
|
cen121212
|
25b48564c3
|
[NPU][Bugfix] fix Qwen3-VL-30B-A3B-Instruct accuracy loss (#15597)
|
2025-12-31 15:57:38 +08:00 |
|
Bingxu Chen
|
1c360bf753
|
[AMD CI] add testcases to unit-test-backend-1-gpu (#16117)
|
2025-12-30 23:08:45 -08:00 |
|
chhnb
|
3619ec61b4
|
[diffusion] feat: support multi-frame image output (#15878)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2025-12-31 13:50:02 +08:00 |
|
Liangsheng Yin
|
db6b51a838
|
Add CI_PERMISSION sort hook. (#16200)
|
2025-12-31 12:15:15 +08:00 |
|
Hudson Xing
|
8a9ca41fda
|
Move log_prefill_stats_late to correct location in PP mode (#15946)
|
2025-12-31 12:00:24 +08:00 |
|
shuwenn
|
ac11e6a7c5
|
Add alphabetc1 into CI_PERMISSION (#16198)
|
2025-12-31 11:31:47 +08:00 |
|
roikoren755
|
47a660d5b9
|
[NemotronH] PP support (#16172)
Signed-off-by: Roi Koren <roik@nvidia.com>
|
2025-12-31 11:16:15 +08:00 |
|
shuwenn
|
c0fc7a89e7
|
[sgl-kernel] fix: make sgl-kernel build respect MAX_JOBS (#15575)
|
2025-12-31 10:44:45 +08:00 |
|
Mick
|
5bf0d862dd
|
[diffusion] CI: fix generate mode and add cli test (#16174)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2025-12-31 09:51:48 +08:00 |
|
liupeng374
|
75b72eb8b2
|
[cp] assert dsv3.2 cp in pd decode mode (#16156)
|
2025-12-31 09:24:46 +08:00 |
|
Leoyzen
|
bc8b526eda
|
Fix: Handle empty func_name and None values in GLM MoE detectors (#15754)
Signed-off-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
|
2025-12-30 14:32:35 -08:00 |
|
Izzy Putterman
|
3dfff6ae3c
|
Eagle: GPT-OSS Eagle v2 support (#14920)
Signed-off-by: Izzy Putterman <iputterman@nvidia.com>
|
2025-12-30 14:23:07 -08:00 |
|
EkiRui
|
ad2c1ee352
|
doc: mooncake store add dummy client support (#16050)
Signed-off-by: Xingrui Yi <yixingrui@linux.alibaba.com>
|
2025-12-30 12:37:57 -08:00 |
|
Simo Lin
|
9940c6f592
|
[model-gateway]: ASCII byte comparison and probabilistic timestamp updates (#16181)
|
2025-12-30 12:28:01 -08:00 |
|
Yineng Zhang
|
00e607111a
|
[Auto Sync] Update request_metrics_exporter.py (20251230) (#16183)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Scott Lee <scottjlee@users.noreply.github.com>
|
2025-12-30 11:09:41 -08:00 |
|
Simo Lin
|
4bc2f2e0bc
|
[model-gateway][docs] Add Classification API documentation (#16182)
Co-authored-by: Chang Su <chang.s.su@oracle.com>
|
2025-12-30 10:53:31 -08:00 |
|
Roger Young
|
d17b9e6392
|
Fusing RMSNormTP in minimax_m2 (#14416)
Co-authored-by: xuebi <xuebi@minimaxi.com>
|
2025-12-30 10:22:07 -08:00 |
|
Liangsheng Yin
|
ba67e006a7
|
Refactor speculative algorithm registry. (#16168)
|
2025-12-31 01:24:22 +08:00 |
|
DarkSharpness
|
45f3ad2f52
|
[Refactor] Rename CustomOp -> MultiPlatformOp (#16175)
|
2025-12-31 01:16:32 +08:00 |
|
Baizhou Zhang
|
f35b5da521
|
[CI] Append test variant name to markdown report header in nightly test (#16166)
|
2025-12-31 00:09:24 +08:00 |
|
Xiaoyu Zhang
|
733a0c1a37
|
[Diffusion] Zimage opt with qknorm and flashinfer rope (#16161)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2025-12-30 23:39:32 +08:00 |
|
Xiaoyu Zhang
|
b369aaa23f
|
[Diffusion] Refine diffusion profling doc (#16163)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
|
2025-12-30 23:38:46 +08:00 |
|
Ke Bao
|
b973202526
|
Split tp model worker init (#16165)
|
2025-12-30 23:36:03 +08:00 |
|
Mufeez Amjad
|
cbff7ad985
|
dp-attention: add follow_bootstrap_room + auto load-balance; drop decode_round_robin (#16110)
|
2025-12-30 22:33:06 +08:00 |
|
Liangsheng Yin
|
4de59d83a1
|
Reduce stages of pr-test from 4 to 3. (#16173)
|
2025-12-30 22:21:48 +08:00 |
|
Yuhao Yang
|
39ca57cd28
|
[diffusion] chore: tiny fix model config (#16159)
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2025-12-30 22:11:38 +08:00 |
|
Mick
|
3449806727
|
[diffusion] feat: generalize layer-wise-offload to all supported models (#16150)
|
2025-12-30 22:06:57 +08:00 |
|
Ke Bao
|
b3817fa93b
|
Split model_worker init function (#16160)
|
2025-12-30 21:39:11 +08:00 |
|
Ke Bao
|
059428bd8a
|
Tiny remove additional args in init_memory_pool (#16158)
|
2025-12-30 21:38:06 +08:00 |
|
jiapingW
|
5d200dd8d9
|
[diffusion] bench: distinguish between video generation and image generation in the bench_serving (#16149)
|
2025-12-30 21:29:18 +08:00 |
|
husf
|
7f9a3d0609
|
[docs][NPU]Update model and feature docs support (#16124)
|
2025-12-30 20:05:40 +08:00 |
|