Commit Graph
6169 Commits
Author SHA1 Message Date
Yashika Gandhi - Google 32ea7bcdd8 [diffusion] endpoint: fix vertex generate (#17611) 2026-01-28 09:38:56 +08:00
Mick 88fcd8535f [diffusion] feat: add an arg for controlling the number of prefetched layers in layerwise-offload (#17693) 2026-01-28 09:34:27 +08:00
Mick 1507dc6cdf [diffusion] fix: fix suppressing error log on non-main ranks (#17712) 2026-01-28 09:29:19 +08:00
Xiaoyu Zhang 331a22427c [Diffusion] glm-image apply flashinfer rope (#17689) 2026-01-28 08:51:37 +08:00
siyuandYuhao Yang 4d00bd17a3 use shared memory for multimodal feature transport between Tokenizer and Scheduler (#16402)
Co-authored-by: Yuhao Yang <47235274+yhyang201@users.noreply.github.com>
2026-01-27 11:01:08 -08:00
Minglei Zhu d90c0837e5 [hybrid-model] clean up and consolidate redundant fields in RadixLinearAttention (#17660) 2026-01-27 10:37:58 -08:00
fsygd 547e2d037e [diffusion] refactor: add arg to control the precision of dit (#17751) 2026-01-27 23:01:23 +08:00
monkeyLovedingandKelon d578b41bad [NPU] Adapt cann 8.5: use sfa and lightning indexer op from cann and CI update (#17615)
Co-authored-by: Kelon <kelonlu@163.com>
2026-01-27 19:03:53 +08:00
MikkoParkkola c56d19b977 fix(quantization): add sgl_kernel fallback for FP4 quantize on Blackwell GPUs (#17816) 2026-01-27 18:43:17 +08:00
Xuchun Shang dba264ac73 [PP] fix wrong weight logic for tie_word_embeddings model (#15890)
Signed-off-by: Xuchun Shang <xuchun.shang@gmail.com>
2026-01-27 17:41:17 +08:00
7106f6c8e1 [GLM-OCR] Support GLM-OCR Model (#17582)
Signed-off-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
2026-01-26 22:24:00 -08:00
Taemin JungandXinyuan Tong 81c0f5c5ad [Model] Add support for EXAONE-4.0 Model (#8205)
Signed-off-by: BoxBy <lute7071@gmail.com>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
2026-01-27 14:08:24 +08:00
laixinandXinyuan Tong 6c9b054ab7 [Bug Fix] Fix reasoning parser when continue_final_message=true (#17065)
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
2026-01-27 14:04:44 +08:00
shuwenn 57e432d951 fix: preserve disconnect events in api key middleware (#17253) 2026-01-26 22:48:24 -05:00
shuwenn fd3b179ffd [HiCache][HA 1/N] Support HiCache storage runtime attach/detach (#15892) 2026-01-26 19:33:19 -08:00
1b56a886bb [chore]: improve time tracing of model loading process (#15426)
Co-authored-by: Michael Shin <mmshin@nvidia.com>
Co-authored-by: ishandhanani <82981111+ishandhanani@users.noreply.github.com>
2026-01-26 19:04:25 -08:00
Yuhao YangandMick 479ab7a4e7 model: support Kimi-K2.5 (#17789)
Co-authored-by: Mick <mickjagger19@icloud.com>
2026-01-27 10:57:00 +08:00
WenhaoZhangniehen6174gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
0519b0935f [diffusion] comfyui: support Qwen-Image, Multi-GPU Z-Image, and Enhanced ComfyUI Integration (#17678)
Co-authored-by: niehen6174 <niehen.6174@gmail.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-01-27 10:06:42 +08:00
FlyPanda 2d8c22a15e [bugfix] Internal processing of hf3fs crash # 16614 (#16938) 2026-01-26 18:01:50 -08:00
Mahdi-CVandgemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com> 539924037f fix(processor): support InternS1 text_config in InternVL processor (#17040)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-01-26 13:02:54 -08:00
ybyangandLiangsheng Yin 5ab76ff220 Special logic for healthcheck (#17734)
Co-authored-by: Liangsheng Yin <lsyincs@gmail.com>
2026-01-26 10:26:40 -08:00
Liangsheng Yin 85d077f44d Introduce global alloc_len_per_decode & clean check decode memory (#15115) 2026-01-26 10:26:20 -08:00
Makcum888e bba6e38ff8 [NPU] Split pyproject npu from pyproject other (#17641) 2026-01-26 09:45:44 -08:00
Yuan Luoandluoyuan.luo 7bb41989fa [1/N] Optimize All Reduce - Benchmark different AR operations (#13797)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
2026-01-26 22:44:13 +08:00
b56366f827 [NPU]DeepSeek-V3.2 support npu mlaprolog (#15381)
Co-authored-by: Zhengda Qin <zhengdqin@gmail.com>
Co-authored-by: richhuan <huan_rz@qq.com>
2026-01-26 20:42:37 +08:00
Yi Zhang 5844cb2fd8 refactor mamba radix cache logic in server_args (#17645) 2026-01-26 17:02:49 +08:00
shaharmor98andb8zhong f6f1b6d000 Bump FI version (#17700)
Signed-off-by: Shahar Mor <smor@nvidia.com>
Co-authored-by: b8zhong <b8zhong@uwaterloo.ca>
2026-01-26 16:50:06 +08:00
McZyWuandcy 2734b23481 accuracy enhancement for baichuan2-13B for npu (#16868)
Co-authored-by: cy <chenyang08056032@163.com>
2026-01-26 16:14:35 +08:00
Prozac614 12f794e516 [diffusion] fix: fix missing backend argument in pipelines_core initialization (#17343) 2026-01-26 15:47:10 +08:00
Kangyan-Zhou 48f4340b14 Exclude some diffusion package for ARM in docker release (#17745) 2026-01-25 23:32:39 -08:00
Alison Shao 30b3192039 Merge performance/accuracy test suites into regular stage-b suites (#17609) 2026-01-25 22:49:19 -08:00
CSWYF3634076 1a19b3987d [Model] Add Ernie4.5 VL model support (#15679)
Signed-off-by: CSWYF3634076 <wangyafeng@baidu.com>
Signed-off-by: wangyafeng <wangyafeng@baidu.com>
2026-01-25 22:36:29 -08:00
zackyoray d275d47973 [NIXL] Add custom NIXL backend selection for KVManager (#17146)
Signed-off-by: Yoray Zack <yorayz@nvidia.com>
2026-01-26 14:35:38 +08:00
Yuan Luo 1e8db18290 [Kimi-Linear] Remove duplicated code in kimi-linear (#17731) 2026-01-26 14:20:24 +08:00
chenxu214 444b9521e4 [Bugfix]Repeated add modelslim quant_config and bugfix with "enable-piecewise-cuda-graph" on NPU (#17511) 2026-01-26 09:51:07 +08:00
Kangyan-ZhouandXinyuan Tong 592603d77b Fix flaky streaming logprobs test by handling detokenizer text buffering (#17687)
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
2026-01-25 15:09:06 -08:00
Kangyan-Zhou 344eeaee90 Upload nightly test metrics to GH artifacts (#17696) 2026-01-25 14:35:14 -08:00
Kangyan-Zhou 8d3e1ac0c8 Add an all type in pyproject.tml to include diffusion support (#17697) 2026-01-25 12:52:13 -08:00
Kangyan-Zhou 9123491430 A few updates to the night tests (#17694) 2026-01-25 11:20:17 -08:00
HandH1998 a883906a24 Support mxint4 flashinfer_trtllm moe gemm (#16892) 2026-01-26 00:15:53 +08:00
Mick b105dad5da [diffusion] refactor: remove useless lazy-import cache-dit codes (#17659) 2026-01-25 22:43:22 +08:00
Zhengbo Wang fb61164f27 [Refactor] Use is_in_ci() utility in JIT kernel benchmarks (#17118) 2026-01-25 20:40:47 +08:00
xjx471258437andxjx392321 9bd92ba0f6 Support PD disaggregation with different TP/DP size for Qwen3-Next (#16056)
Co-authored-by: xjx392321 <xjx392321@alibaba-inc.com>
2026-01-25 15:34:02 +08:00
Ke Bao 30ece5e1d6 Fix swa memory pool size with spec (#17630) 2026-01-25 14:10:43 +08:00
Mohammad Miadh Angkad 1674b9ef44 [DeepSeek-V3.2] Fix TRT-LLM NSA in target_verify/draft_extend (#17662) 2026-01-25 13:10:14 +08:00
Alison Shao 9121f22656 Add PyTorch .bin file validation to CI weight validation (#17533) 2026-01-24 19:18:15 -08:00
59f027a8c8 [diffusion]: Fix ZImage SP sharding for caption and latent (#17301)
Co-authored-by: rhyshen <rhyshen@tencent.com>
Co-authored-by: florianzhao <florianzhao@tencent.com>
2026-01-25 10:10:48 +08:00
Xinyuan Tong 37c04c2245 fix: Refactor register_image_processor to use kwarg instead of positional arg (#17685)
Signed-off-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
2026-01-24 15:31:01 -08:00
Trevor Morris 2c2c4e446b [NVIDIA] Add flashinfer all-to-all MOE dispatcher (#14668) 2026-01-24 22:59:55 +08:00
TMC 458a43d4ac [NPU] torch_npu profiler tensorboard path type fix (#17545) 2026-01-24 22:55:49 +08:00