Commit Graph

  • 62361e8fd5 Temporarily adjust the scheulde for pr-test.yml to 12 hours (#20951) Kangyan-Zhou 2026-03-19 14:48:10 -07:00
  • b42b9f6e1a Support CuteDSL mm_fp4 backend (#18801) Brayden Zhong 2026-03-19 17:20:01 -04:00
  • d8ece7fb22 [Tiny Fix] Filter lru related warning with pcg (#20940) Yuwei An 2026-03-19 13:20:49 -07:00
  • 0949b138af Simplify server startup output (#20885) Lianmin Zheng 2026-03-19 13:11:37 -07:00
  • a02cff7f2b [Fix] Patch is_flash_attn_2_available for flash-attn-4 in VLM input format test (#20946) Xinyuan Tong 2026-03-19 20:00:51 +00:00
  • c562e0d13b [feat] Enhance Kimi-K2/K2.5 function call and reasoning detection (#19552) AlfredYong 2026-03-20 03:57:57 +08:00
  • 2ee7f41e25 [AMD] add mori ep normal tbo unittest (#20941) billishyahao 2026-03-20 02:52:25 +08:00
  • 9e629d31fd [AMD] CI - Fix AMD CI (multimodal test, move flaky test to non-deterministic group) (#20815) YC Yen-Ching Tseng 2026-03-20 02:50:19 +08:00
  • 3b89c6d3ea simplify and improve pr-test.yml (#20911) Lianmin Zheng 2026-03-19 11:01:06 -07:00
  • 29ced9c162 [UX] Suppress noisy httpx/httpcore INFO logs (#20944) Mohammad Miadh Angkad 2026-03-20 01:58:41 +08:00
  • 319bb4974c [Fix] RayEngine multi-node: co-locate rank0 scheduler with Engine and fix CUDA device setting (#20722) Xinyu Zhang 2026-03-19 10:27:16 -07:00
  • 274581fb77 Add support for more batch sizes in cpu_graph_runner (#13881) Cao E 2026-03-20 00:50:56 +08:00
  • 4c52b7fcc6 [CI] Improve PP consistency check success rate (#20838) Shangming Cai 2026-03-19 19:12:50 +08:00
  • 671bd266c1 [AMD] CI - update amd nightly cronjob (#20928) YC Yen-Ching Tseng 2026-03-19 18:10:17 +08:00
  • c8f0122acf Fix gpu-fault issue when run deepseek-r1 and enable dp (#20841) kk 2026-03-19 17:36:12 +08:00
  • cb8105fe28 [sgl-kernel][6/7]Support Expert Specialization Grouped GEMM (#15471) Qi Yuhang 2026-03-19 15:39:52 +08:00
  • 574572b21b [BugFix] bug fix for DeepSeek eagle3 in Attn-DP mode (#20492) khalilzhk 2026-03-19 14:48:46 +08:00
  • fd05532da1 Add logging for BootstrapServer for CI diagnosis (#20844) Shangming Cai 2026-03-19 14:42:12 +08:00
  • a98b456c70 [CPU] Add frontend support for Gemma (#12590) blzheng 2026-03-19 14:02:26 +08:00
  • 8d4fcf2f7b [CPU] Fix MoE layer support for DeepSeek-OCR models (#12555) jianan-gu 2026-03-19 13:57:55 +08:00
  • 85fe8c6793 [AMD] Use aiter_dsv3_router_gemm kernel if number of experts <= 256. (#18451) Matti Varjokallio 2026-03-19 07:40:48 +02:00
  • 126cd5cfae gpt-oss decode performance optimization (#20392) kk 2026-03-19 13:30:03 +08:00
  • cd22aa27a9 [CPU] Add FP8 Bmm support (#9744) blzheng 2026-03-19 13:19:48 +08:00
  • c2b01bd2fc [CPU] fix bug in AVX512 implementation of flash_attn_softmax (#20220) blzheng 2026-03-19 13:18:47 +08:00
  • 687d9eb66f [CPU] Optimize image preprocessor performance for Qwen2VLImageProcessorFast (#15168) Ma Mingfei 2026-03-19 13:18:15 +08:00
  • 62d7454976 optimize conv3d used in patch embedding (#16040) Ma Mingfei 2026-03-19 13:17:53 +08:00
  • cbea9f6909 [CPU] improve numa memory binding (#19666) blzheng 2026-03-19 13:15:50 +08:00
  • 2f4babe32b [CPU] support LayerNorm with 3D shape (#15075) Zaili Wang 2026-03-19 13:15:24 +08:00
  • dc6aa26ce9 [CPU] Add mrope kernel for Qwen3-vl (#12531) blzheng 2026-03-19 13:12:48 +08:00
  • 4052b53227 fix scheduler for non-cuda devices and disable piecewise cuda graph f… (#19992) Juan Muneton 2026-03-18 21:54:19 -07:00
  • f85455ab24 [Bugfix] fix qwen3vl hang when --mm-enable-dp-encoder is enable (#20759) Ling Zhang 2026-03-19 12:51:39 +08:00
  • 7f6f1a3ab1 [LoRA][II] Add fused MOE LoRA Triton kernel and tests (#19711) Ethan (Yusheng) Su 2026-03-18 19:58:14 -07:00
  • 7553b7dcb0 chore: extract diffusion_common in python/pyproject_other.toml (#20803) R0CKSTAR 2026-03-19 10:39:16 +08:00
  • eea9e19c13 fix lint introduced in #20708 (#20886) Qiaolin Yu 2026-03-18 15:38:52 -07:00
  • 0d23a461a0 feat(mm)(grpc): compute M-RoPE positions for preprocessed VL inputs (#19973) Chang Su 2026-03-18 15:34:50 -07:00
  • 8b9482e665 fix(dp-attn): consistent overlap disable decision across DP ranks (#20853) Liangsheng Yin 2026-03-18 15:16:39 -07:00
  • 4e8829e4cd Replace topk_ids with curr_topk_ids in fused_moe.py (#20302) maocheng23 2026-03-18 14:57:05 -07:00
  • a3196d08b8 [MiniMax M2] Fix KV cache scale loading (#20870) Chad Voegele 2026-03-18 16:54:43 -05:00
  • 6b8a6545b2 Add Mistral Small 4 (Pixtral) support (#20708) Xinyuan Tong 2026-03-18 21:15:32 +00:00
  • df1d046de2 Add packed_modules_mapping for MiniMax-M2 (#19995) Trevor Morris 2026-03-18 14:10:01 -07:00
  • d1e95af282 Upgrade transformers==5.3.0 (#17784) Xinyuan Tong 2026-03-18 20:50:43 +00:00
  • e5750a572c Support TP for lora lm_head layer (#18511) Bruce Wu 2026-03-18 13:48:03 -07:00
  • 8f0f36c64b [1/2] Add ModelExpress coordination for remote instance weight loading - matching TP (#19920) ishandhanani 2026-03-18 15:38:32 -05:00
  • 6cca5b9b97 [NPU] [BUGFIX] Test ascend memory consumption.py fix (#17995) Артем Савкин 2026-03-18 22:47:43 +03:00
  • 46a392658e Refine RL & Post-Training description in README (#20877) Lianmin Zheng 2026-03-18 12:43:42 -07:00
  • c7a71740a5 [NPU][diffusion] npu support enable_torch_compile for torchair backend on diffusion models (#20687) Yaochen Han 2026-03-19 03:40:35 +08:00
  • 39e83d1401 [CI] Eliminate per-arch Docker tags by using push-by-digest (#20793) Lianmin Zheng 2026-03-18 12:35:06 -07:00
  • b9dba851a0 Fix streaming token ids data loss under load (#19977) Vladislav Nosivskoy 2026-03-18 22:23:45 +03:00
  • 70876ae93b fix: guard configure_deep_gemm_num_sms when JIT disabled (#20868) Gabriel Wu 2026-03-19 02:15:20 +08:00
  • a6c7bb54eb [Perf]Optimize waiting queue update with set usage (#20503) Jackie 2026-03-19 00:56:24 +08:00
  • 20a23e3173 [SKILL] Refine kernel authoring docs and validate add-jit-kernel / add-sgl-kernel end to end with Codex (#20867) Xiaoyu Zhang 2026-03-18 23:00:33 +08:00
  • 21c4fc6334 [DP encoder] Fix pos_emb layer TP issue when DP encoder enabled for Qwen3 VL (#20788) jianan-gu 2026-03-18 17:14:47 +08:00
  • 05c00088e3 Cleanup broken dist-info directories in ci deps install (#20817) Ke Bao 2026-03-18 17:09:23 +08:00
  • c0a4408f78 [AMD] Fix dpsk-v32 accuracy issue on mi355 (#20840) Thomas Wang 2026-03-18 17:06:15 +08:00
  • f0d7a3f427 [AMD][TBO] Fix mori ep dual stream accuracy (#19888) billishyahao 2026-03-18 17:00:55 +08:00
  • 8b46f1f4ec [PD] Add retry interval in ensure_prefill_info (#20832) Shangming Cai 2026-03-18 16:02:20 +08:00
  • 93422f27d6 [AMD][AITER] Guard _use_mla_ps_kernel with self.use_mla in draft_extend_v2 paths (#20409) Chuan (Richard) Li 2026-03-18 00:45:22 -07:00
  • ead9d7aa43 [diffusion] fix: fix vae model offload on mps(#20607) R0CKSTAR 2026-03-18 15:44:59 +08:00
  • 532470bcca [NPU] add new fusion operator DispatchFFNCombine (#20245) chenxu214 2026-03-18 15:22:04 +08:00
  • ae15fca192 [Bugfix] fix hicache mooncake backend extra config loading (#16808) jinke 2026-03-18 15:07:39 +08:00
  • d20e9a20fa [JIT] Inject target architecture flag into JIT compilation (#20103) xingsy97 2026-03-18 14:16:49 +08:00
  • f78d5c3b3c [JIT Kernel] Add hadamard kernel test and benchmark (#20030) xingsy97 2026-03-18 14:16:35 +08:00
  • c64681f162 [Bugfix] [diffusion] Fix cache-dit with sp-degree only (#19965) Артем Савкин 2026-03-18 09:05:12 +03:00
  • b6055e59cd [HiCache] Reduce per-request backup log noise (#20813) Kangyan-Zhou 2026-03-17 22:47:14 -07:00
  • 30a35ecd90 Add gigachat3.1 parser (#19886) Viacheslav 2026-03-18 08:45:01 +03:00
  • 2e860233ca rocm: fix oom when loading fp8 weights close to size of available vram (#19941) Evgueni Petrov 2026-03-18 12:44:19 +07:00
  • 0acc1d3c9a fix: change qwen 3.5 linear attention a_log to fp32 (#19961) shiyu7 2026-03-18 13:42:06 +08:00
  • 88c40ec16d Use Flashinfer for target_verify in GDN model for SM120 (#20604) Brayden Zhong 2026-03-18 01:40:56 -04:00
  • 97d5386a21 Use TRTLLM allreduce fusion for Qwen 3.5 (#19889) Brayden Zhong 2026-03-18 01:40:22 -04:00
  • 1b690836fa Add ut coverage workflow (#20804) Ke Bao 2026-03-18 13:25:53 +08:00
  • c42da50289 Update test guide to contribution guide (#20805) Ke Bao 2026-03-18 13:25:16 +08:00
  • 9c87e137ee [GDN] Support GDN packed decode (#20627) Yuan Luo 2026-03-18 13:20:07 +08:00
  • 4cc19862ef [NVIDIA] Integrate FlashInfer decode kernel (Blackwell) for Qwen3.5 (#19150) Kaixi Hou 2026-03-17 22:11:18 -07:00
  • da24478913 [AMD] Fix CI: update transformers in Qwen 3.5 and GLM-5 nightly tests (#20750) Michael 2026-03-17 22:07:07 -07:00
  • c43d495dd5 [RadixTree][9/N Refactor]: Support unified init_load_back params (#20590) hzh0425 2026-03-18 11:19:52 +08:00
  • f15b3338c9 Revert "[Bugfix] Fix GLM-4.6V vision regression in glm4v_moe and glm_ocr" (#20740) Mick 2026-03-18 10:09:50 +08:00
  • 944355c66f [Bugfix] Fix model output corruption caused by EPLB rebalance (Eager and CUDA Graph modes) (#18213) lviy 2026-03-18 09:30:24 +08:00
  • 41c8651c0d Revert "[AMD] CI - skip tests on mi325 (#20660)" (#20802) YC Yen-Ching Tseng 2026-03-18 09:01:51 +08:00
  • 4d3976b6c5 [HiCache] Check in-flight async ops in is_fully_idle() before attach/detach (#20746) Liangsheng Yin 2026-03-17 17:28:26 -07:00
  • c5d2528bff Revert "[AMD][MORI] Fix MTP crash with FP4/FP8 dispatch and add NEXTN dispatch env vars." (#20797) Qiaolin Yu 2026-03-17 17:28:09 -07:00
  • 2acb20f53b [Disagg] Non-blocking try_ensure_parallel_info in pending queue, consolidate rank mapping into PrefillServerInfo (#20785) Shangming Cai 2026-03-18 08:26:18 +08:00
  • cb1e63aba4 bump fa4 to official released fa4 pkg (#20303) Rain Jiang 2026-03-17 17:22:56 -07:00
  • c77d7c629e [Bugfix] Fix MTP prefill cuda graph logging (#20279) Jincong Chen 2026-03-18 07:36:52 +08:00
  • 744b1c9e6f Added fallback to individual copy_ (#20683) Kaixi 2026-03-17 22:44:38 +01:00
  • a515cc38a7 Delete .editorconfig (#20795) Lianmin Zheng 2026-03-17 13:02:04 -07:00
  • dc1ab5d2b2 Delete CODE_OF_CONDUCT.md (#20794) Lianmin Zheng 2026-03-17 13:01:51 -07:00
  • 3d8fc9a0ca Revert "[Nvidia] Add trtllm mnnvl allreduce with unified flashinfer allreduce fusion api" (#20792) Kangyan-Zhou 2026-03-17 11:59:02 -07:00
  • 666b5e4852 [NPU] Update torch and torch_npu version for NPU (#20013) Makcum888e 2026-03-17 21:25:02 +03:00
  • 09f5097fe4 [NPU] [Bugfix] [diffusion] Fix NZ performance bug for diffusion models (#20684) Артем Савкин 2026-03-17 21:23:09 +03:00
  • d35fea1b2b [Nvidia] Add trtllm mnnvl allreduce with unified flashinfer allreduce fusion api (#12787) Shu Wang 2026-03-17 10:02:45 -07:00
  • 17031120b8 [DeepSeek v3.2][Bugfix] get_index_k_scale_buffer support cp (#18280) Yongfei Xu 2026-03-18 00:54:54 +08:00
  • 466ff20e51 [Model] Fix NemotronH OOM on unified-mem systems: stream weights + safetensors cleanup (#20580) Serge Panev 2026-03-17 09:47:58 -07:00
  • 24a27d5320 vlm: support piecewise cuda graph for Kimi-K2.5 (#20747) Yuhao Yang 2026-03-18 00:32:07 +08:00
  • 7f99319c56 Add evict policy ut (#20787) Ke Bao 2026-03-17 22:06:14 +08:00
  • b5f3eaecbc [NPU] Support dequant_swiglu_quant & moe_init_routing_v2 & npu_moe_token_unpermute for W8A8 MoE decode (#19913) heziiop 2026-03-17 21:39:29 +08:00
  • 5717834f1f [diffusion] refactor: cleanup parallel_state.py (#20760) Mick 2026-03-17 21:21:42 +08:00
  • 17c81a3e07 Revert "[PD] Make pending reqs resolving more robust" (#20779) Shangming Cai 2026-03-17 20:31:12 +08:00
  • cfead25bbf [Qwen3.5] mamba slice fix (Prefill TP != Decode TP & decode TP size>1) (#20655) YAMY 2026-03-17 04:30:58 -07:00
  • 966ae87d02 [AMD] avoid correction_bias_dtype dtype convert (#20692) AMD-yanfeiwang 2026-03-17 17:55:05 +08:00
  • 5270a06488 [Disagg] Fix health check false-positive in disagg is_fully_idle (#20756) Liangsheng Yin 2026-03-17 02:18:54 -07:00