Commit Graph

11116 Commits

Author SHA1 Message Date
陈一涵
647428d8d6 [diffusion] perf: apply mul add fusion for Qwen-Image (#16299)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-01-28 09:40:13 +08:00
Yashika Gandhi - Google
32ea7bcdd8 [diffusion] endpoint: fix vertex generate (#17611) 2026-01-28 09:38:56 +08:00
Mick
88fcd8535f [diffusion] feat: add an arg for controlling the number of prefetched layers in layerwise-offload (#17693) 2026-01-28 09:34:27 +08:00
Mick
1507dc6cdf [diffusion] fix: fix suppressing error log on non-main ranks (#17712) 2026-01-28 09:29:19 +08:00
Xiaoyu Zhang
331a22427c [Diffusion] glm-image apply flashinfer rope (#17689) 2026-01-28 08:51:37 +08:00
Hubert Lu
93423ff780 [AMD] Deprecate ROCm 6.3 artifacts and standardize gfx942 on ROCm 7 (#17785) 2026-01-27 15:58:49 -08:00
Liangsheng Yin
8278ef0e68 Pass GPU ids to kill specified devices in script. (#17840)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-01-27 13:52:55 -08:00
siyu
4d00bd17a3 use shared memory for multimodal feature transport between Tokenizer and Scheduler (#16402)
Co-authored-by: Yuhao Yang <47235274+yhyang201@users.noreply.github.com>
2026-01-27 11:01:08 -08:00
Minglei Zhu
d90c0837e5 [hybrid-model] clean up and consolidate redundant fields in RadixLinearAttention (#17660) 2026-01-27 10:37:58 -08:00
Yi Zhong
8acd4d7d7e Make flashMLA work on: Cu13, B300 (#17600)
Signed-off-by: vincentzed <207368749+vincentzed@users.noreply.github.com>
2026-01-28 00:12:47 +08:00
fsygd
547e2d037e [diffusion] refactor: add arg to control the precision of dit (#17751) 2026-01-27 23:01:23 +08:00
Baizhou Zhang
1d942e4eef [DeepSeek] Update tests and document for DeepSeek V3.2 NVFP4 checkpoint (#17657) 2026-01-27 22:10:57 +08:00
monkeyLoveding
d578b41bad [NPU] Adapt cann 8.5: use sfa and lightning indexer op from cann and CI update (#17615)
Co-authored-by: Kelon <kelonlu@163.com>
2026-01-27 19:03:53 +08:00
Baizhou Zhang
832c756549 [Doc] Tiny update description on torch compile (#17819) 2026-01-27 18:59:04 +08:00
MikkoParkkola
c56d19b977 fix(quantization): add sgl_kernel fallback for FP4 quantize on Blackwell GPUs (#17816) 2026-01-27 18:43:17 +08:00
Xuchun Shang
dba264ac73 [PP] fix wrong weight logic for tie_word_embeddings model (#15890)
Signed-off-by: Xuchun Shang <xuchun.shang@gmail.com>
2026-01-27 17:41:17 +08:00
Yuxuan Zhang
7106f6c8e1 [GLM-OCR] Support GLM-OCR Model (#17582)
Signed-off-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
2026-01-26 22:24:00 -08:00
Taemin Jung
81c0f5c5ad [Model] Add support for EXAONE-4.0 Model (#8205)
Signed-off-by: BoxBy <lute7071@gmail.com>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
2026-01-27 14:08:24 +08:00
laixin
6c9b054ab7 [Bug Fix] Fix reasoning parser when continue_final_message=true (#17065)
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
2026-01-27 14:04:44 +08:00
Hubert Lu
df42f4d386 [AMD] Update dsv3.2 AMD GPU docs and unify ROCm TileLang build (#17783)
Co-authored-by: wufann <715544327@qq.com>
2026-01-26 21:10:32 -08:00
shuwenn
a723d1c5ef [model-gateway] ignore error for embeddings/classify in PD router (#15931) 2026-01-26 22:49:14 -05:00
shuwenn
57e432d951 fix: preserve disconnect events in api key middleware (#17253) 2026-01-26 22:48:24 -05:00
shuwenn
fd3b179ffd [HiCache][HA 1/N] Support HiCache storage runtime attach/detach (#15892) 2026-01-26 19:33:19 -08:00
Zhongdongming Dai
1b56a886bb [chore]: improve time tracing of model loading process (#15426)
Co-authored-by: Michael Shin <mmshin@nvidia.com>
Co-authored-by: ishandhanani <82981111+ishandhanani@users.noreply.github.com>
2026-01-26 19:04:25 -08:00
Shangming Cai
3ad3268e06 [CI] Skip PD hybrid attention test with different TP temporarily (#17791) 2026-01-27 10:59:33 +08:00
Yuhao Yang
479ab7a4e7 model: support Kimi-K2.5 (#17789)
Co-authored-by: Mick <mickjagger19@icloud.com>
2026-01-27 10:57:00 +08:00
WenhaoZhang
0519b0935f [diffusion] comfyui: support Qwen-Image, Multi-GPU Z-Image, and Enhanced ComfyUI Integration (#17678)
Co-authored-by: niehen6174 <niehen.6174@gmail.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-01-27 10:06:42 +08:00
FlyPanda
2d8c22a15e [bugfix] Internal processing of hf3fs crash # 16614 (#16938) 2026-01-26 18:01:50 -08:00
Mahdi-CV
539924037f fix(processor): support InternS1 text_config in InternVL processor (#17040)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-01-26 13:02:54 -08:00
zijiexia
dd97e1fe38 [Docs] Add RL documentation (#17663)
Co-authored-by: JD <jaedon.guo@gmail.com>
2026-01-26 12:16:54 -08:00
ybyang
5ab76ff220 Special logic for healthcheck (#17734)
Co-authored-by: Liangsheng Yin <lsyincs@gmail.com>
2026-01-26 10:26:40 -08:00
Liangsheng Yin
85d077f44d Introduce global alloc_len_per_decode & clean check decode memory (#15115) 2026-01-26 10:26:20 -08:00
Douglas Yang
8643fb2f52 fix: remove truncation for test and job names in ci failure monitor (#17765) 2026-01-26 09:46:33 -08:00
Makcum888e
bba6e38ff8 [NPU] Split pyproject npu from pyproject other (#17641) 2026-01-26 09:45:44 -08:00
Douglas Yang
51d139b867 fix: move nightly whl to cuda version folder (#17762) 2026-01-27 00:13:46 +08:00
Yuan Luo
7bb41989fa [1/N] Optimize All Reduce - Benchmark different AR operations (#13797)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
2026-01-26 22:44:13 +08:00
Alison Shao
6c0f9b4824 Add test_gpt_oss_4gpu.py to B200 test suite (#17743) 2026-01-26 21:57:06 +08:00
lawtherWu
b56366f827 [NPU]DeepSeek-V3.2 support npu mlaprolog (#15381)
Co-authored-by: Zhengda Qin <zhengdqin@gmail.com>
Co-authored-by: richhuan <huan_rz@qq.com>
2026-01-26 20:42:37 +08:00
sogalin
738b1ac988 [AMD CI] Add moonshotai/Kimi-K2-Instruct-0905 testcases (#17656) 2026-01-26 02:12:34 -08:00
Yi Zhang
5844cb2fd8 refactor mamba radix cache logic in server_args (#17645) 2026-01-26 17:02:49 +08:00
shaharmor98
f6f1b6d000 Bump FI version (#17700)
Signed-off-by: Shahar Mor <smor@nvidia.com>
Co-authored-by: b8zhong <b8zhong@uwaterloo.ca>
2026-01-26 16:50:06 +08:00
McZyWu
2734b23481 accuracy enhancement for baichuan2-13B for npu (#16868)
Co-authored-by: cy <chenyang08056032@163.com>
2026-01-26 16:14:35 +08:00
Simo Lin
6756cf18c6 remove multimodal as this is completely dead code (#17750) 2026-01-26 00:07:40 -08:00
Prozac614
12f794e516 [diffusion] fix: fix missing backend argument in pipelines_core initialization (#17343) 2026-01-26 15:47:10 +08:00
Simo Lin
d4adff31aa update wasm endpoint (#17748) 2026-01-25 23:43:29 -08:00
Kangyan-Zhou
48f4340b14 Exclude some diffusion package for ARM in docker release (#17745) 2026-01-25 23:32:39 -08:00
Simo Lin
ed75136e85 remove self managed wasm as it has been replaced with official smg wa… (#17746) 2026-01-25 22:56:33 -08:00
Alison Shao
30b3192039 Merge performance/accuracy test suites into regular stage-b suites (#17609) 2026-01-25 22:49:19 -08:00
Praneth Paruchuri
8c2d8b51e9 [model-gateway] fix wasm example2 (#17244) 2026-01-25 22:45:36 -08:00
Praneth Paruchuri
02c1dabf5d [model-gateway] fix wasm example3 (#17277) 2026-01-25 22:45:21 -08:00