Commit Graph
7033 Commits
Author SHA1 Message Date
Baizhou Zhang be63f982b7 [V32/GLM5] Control the threshold of applying dense attention with an environ (#20062) 2026-03-09 14:36:10 -07:00
d39ed074cf fix: default FP4 GEMM backend to flashinfer_cudnn on SM120 (Blackwell) (#20047)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
2026-03-09 14:13:08 -07:00
Baizhou Zhang 61d530e8ac [CI] Fix lint (#20209) 2026-03-09 14:09:59 -07:00
ybyang 3e8abc71ca [Disagg] Skip health check enqueue when PD disagg queues have backlog (#20191) 2026-03-09 12:58:10 -07:00
AMD-yanfeiwangandDuyi-Wang f0153ad225 [AMD][Feature] support fp4 dispatch and fp8 combine in moriep (#19757)
Co-authored-by: Duyi-Wang <duyi.wang@amd.com>
2026-03-09 12:52:05 -07:00
Liangsheng Yin ffb4b6f4c1 [Core] Replace server_args mutation hack with explicit MemoryPoolConfig for draft worker init (#20183) 2026-03-09 11:45:54 -07:00
Yuhao Yang ecca8c553d [diffusion] fix: fix diffusers backend issues in diffusion ci gt workflow (#20173) 2026-03-10 00:51:48 +08:00
Ke Bao 2e444bdced Move stop words to args in send one (#20193) 2026-03-09 23:05:32 +08:00
eb4ba1bde2 Feature/support longcat flash lite (#17838)
Co-authored-by: sunjiaqi11 <sunjiaqi11@meituan.com>
Co-authored-by: ispobock <ispobaoke@gmail.com>
2026-03-09 23:00:11 +08:00
11b76d24dc [NPU] [DLLM]DLLM LLaDA2.x graph mode support with NPU speedup modifications (#18485)
Co-authored-by: Zhang-Xiaoxue <xiaoxuezhang17@outlook.com>
Co-authored-by: dawncc <dawn.cc022@gmail.com>
Co-authored-by: lixinqi7 <li_xinqi7@163.com>
Co-authored-by: rangejay <rangejay1st@163.com>
2026-03-09 22:41:05 +08:00
Xinyuan Tong d116a8cd94 [Bugfix] Fix load_audio: mono before resample + use torchaudio (#20054) 2026-03-09 19:24:20 +08:00
4a757990a1 [VLM] Replace decord with torchcodec for video decoding (#20055)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: BakerBunker <17872844+BakerBunker@users.noreply.github.com>
2026-03-09 19:23:49 +08:00
Yuzhen Zhou b719219de9 [ROCm] Use unreg path for aiter custom all-reduce during CUDA graph capture (#20155) 2026-03-09 01:09:04 -07:00
luoyuyan cabe171b6c Fix qwen3.5 mtp eplb related issues (#19767) 2026-03-09 16:05:32 +08:00
roikoren755 c76251f70c Return intermediate Mamba states (#19716) 2026-03-09 16:04:36 +08:00
Mike Qiu 96724f490c Add auto bind numa node (#15678)
Signed-off-by: Michael Qiu <qiudayu.qdy@antgroup.com>
2026-03-08 23:46:09 -07:00
Liangsheng Yin 2ef00383ab [Core] Refactor init_memory_pool into composable resolution helpers (#20142) 2026-03-08 21:46:27 -07:00
siyu c6184b7dc0 Fix EPD OOM by offloading precomputed_embeddings during chunked prefill (#16503) 2026-03-08 20:10:40 -07:00
Yuhao Yang 1cb86f5171 [diffusion] CI: fix CI script path and missing server arg in perf baseline generator (#20138) 2026-03-09 10:35:21 +08:00
cen121212 fc543df289 [NPU] qwen3_vl encoder support graph 2026-03-09 10:13:35 +08:00
Kaixi 8c5ca37aef Batch copy_ with torch._foreach_copy_ (#18558) 2026-03-08 19:09:02 -07:00
Yuhao Yang 57f28fda90 [diffusion] chore: add diffusion new model skill (#19605) 2026-03-09 09:45:23 +08:00
Simo LinandChang Su 3f3eb206fa feat(grpc): add SubscribeKvEvents RPC for KV cache event streaming (#20112)
Co-authored-by: Chang Su <chang.s.su@oracle.com>
2026-03-08 16:00:29 -07:00
Liangsheng YinandClaude Opus 4.6 7105bf3782 [Bug] Fix missing TTFT histogram for single-batch requests (#20122)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-08 15:18:51 -07:00
Junhao Liu 7662b8b919 [diffusion] feat: implement upscaling (#19723) 2026-03-09 02:06:40 +08:00
xingsy97 b77dd41db0 [diffusion] fix: fix temporary resolution workaround (#20046) 2026-03-09 02:05:35 +08:00
hzh0425 0ac6c63ae4 [SpecV2-Mamba]: Refactor additional_ratio calculation when init mamba pool (#19660) 2026-03-09 00:39:26 +08:00
Ke Bao 07359efce9 Fix missing clone in hicache (#20130) 2026-03-08 23:21:18 +08:00
Junhao Liuandronnie_zheng 051427c0a3 [diffusion] benchmark: add SLO metric forinbench_serving (#18907)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-03-08 22:35:57 +08:00
liubiyonggeandzhangheng cc73355a1f [Feature] Add SLRU eviction policy & fix RadixCache hit_count bug (#18843)
Co-authored-by: zhangheng <hzh0425@apache.org>
2026-03-08 21:30:55 +08:00
Mick 2c183350be [diffusion] fix: fix wrong dit config for qwen-image-edit-plus-2511 (#20123) 2026-03-08 20:08:36 +08:00
Ratish P ab9de886c5 [diffusion] reduce LayerwiseOffloadManager reserved GPU memory (#20042) 2026-03-08 19:26:17 +08:00
Liangsheng Yinandgemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com> 29f3a5396e [Minor] Add SessionSlot.is_holding_kv property for readability (#20120)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-03-08 03:25:13 -07:00
36b557d2c9 Fix streaming session with paged KV cache (SWA/MLA) (#20070)
Co-authored-by: Yilong Zhao <74357408+happierpig@users.noreply.github.com>
Co-authored-by: Aurick Qiao <6137920+aurickq@users.noreply.github.com>
2026-03-08 03:00:32 -07:00
yuyu5333andzhangheng 230fb55899 [Performance] Decode Offload improves the long texts performance 100% through dynamic block offload. (#17216)
Co-authored-by: zhangheng <hzh0425@apache.org>
2026-03-08 17:16:53 +08:00
Yuan Luoandluoyuan.luo 97a2a9be0f [VLM] Replace conv3d proj with linear for GLM4V (#20033)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
2026-03-07 22:50:47 -08:00
Fan LinandYihan Chen 7fb282a96f [diffusion] fix: fix bug of copy_if (#20094)
Co-authored-by: Yihan Chen <yingluosanqian@gmail.com>
2026-03-08 14:27:58 +08:00
xingsy97 7f9f85d4c8 [diffusion] feat: make QwenImageLayered resolution configurable (#20044) 2026-03-08 14:26:05 +08:00
Lancer a73369c39f [diffusion] chore: ensure CFG Zero Star numerical stability for Helios model (#20091)
Signed-off-by: Lancer <maruixiang6688@gmail.com>
2026-03-08 14:25:14 +08:00
shuwenn 72f6dfcc31 fix: add ModelScope cache lookup and speculative path support (#20098) 2026-03-07 22:23:16 -08:00
Liangsheng Yin d02c515ee8 Decouple scheduler log printing from metrics collection (#20107) 2026-03-07 22:09:10 -08:00
Baizhou Zhang d28f35240a [V32/GLM5] Change default setting of V32 nvfp4 on TP4 (#20086) 2026-03-07 15:13:25 -08:00
Alison ShaoandAlison Shao 0f62da6953 [CI] Show test partition assignments after checkout (#20085)
Co-authored-by: Alison Shao <alisonshao@mac.lan>
2026-03-07 13:50:49 -08:00
VDV1985gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>Egor Filimonovronnie_zheng
45bd30e29d [NPU] make torch_native lora backend a little bit faster (#17228)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Egor Filimonov <44640852+ssshinigami@users.noreply.github.com>
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-03-07 20:14:46 +03:00
Ke Baoandhzh0425 5867c3fa80 Support HiCache for MambaRadixCache (#19663)
Co-authored-by: hzh0425 <hzh0425@apache.org>
2026-03-08 00:36:25 +08:00
Bingxu Chen 17721b00fd [AMD] Fix Tensor Memory Aliasing (#19928) 2026-03-07 08:06:10 -08:00
Yuan Luoandluoyuan.luo 7da590d4d0 [Qwen3.5] Support Qwen3.5 Pipeline Parallelism (#19670)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
2026-03-07 23:34:08 +08:00
13bdc7bf4a [Feature][NPU]: add runtime support for AutoRound quantized models (#16699)
Co-authored-by: root <root@localhost.localdomain>
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-03-07 18:03:55 +03:00
Артем Савкинandronnie_zheng 5297b02c88 [Diffusion] [NPU] Wan2.2-T2V-A14B-Diffusers modelslim quantization support (#17996)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-03-07 17:26:44 +03:00
xingsy97 f8d4eb7022 [Docs] Add docstrings to JIT kernel include headers (#19770) 2026-03-07 20:48:00 +08:00