Commit Graph

  • ba9f6d8f26 [Refactor] Clean up JIT kernel utilites (#16884) DarkSharpness 2026-01-13 17:54:16 +08:00
  • 740d3c0b39 [Diffusion] Remove useless dependency in diffusion (#16967) Xiaoyu Zhang 2026-01-13 17:25:53 +08:00
  • 8716589826 [AMD][Diffusion] support timestep embedding kernel for AMD GPUs (#16766) Hubert Lu 2026-01-12 22:17:07 -08:00
  • ff3ddb9d9b Support min num routing keys in key-based load balancing policy (#16564) fzyzcjy 2026-01-13 13:38:03 +08:00
  • 9d3018f484 Support min load besides random routing key assignment policy in ManualPolicy (#16767) fzyzcjy 2026-01-13 10:31:18 +08:00
  • a83484275d [diffusion] perf: optimize linear calculation in SLA (#16648) HuangJi 2026-01-13 09:38:35 +08:00
  • 47d485f35f [diffusion] fix: fix not respecting dit_layerwise_offload server arg (#16252) Mick 2026-01-13 09:29:07 +08:00
  • 2b42309955 [diffusion] UX: provide solutions for OOM (#16940) Mick 2026-01-13 09:25:27 +08:00
  • ae0baefb94 [NPU] upgrade npu mf_apater plugin (#15853) James 2026-01-13 09:02:10 +08:00
  • 1f0e3d7fd8 Support tracking worker routing key loads in gateway (#16765) fzyzcjy 2026-01-13 08:07:17 +08:00
  • d3c08fb07c [layers] support zero-dim rmsnorm (#16978) Yinghai Lu 2026-01-12 15:53:19 -08:00
  • c6a64e9f69 [smg] fix type complexity for workflow run_if (#16981) Simo Lin 2026-01-12 15:41:52 -08:00
  • e0ac559ae1 feat(workflow): add scheduled/delayed steps and conditional branching (#16980) Simo Lin 2026-01-12 15:06:12 -08:00
  • 6620548fd8 [model-gateway] make StateStore trait async for external persistence (#16979) Simo Lin 2026-01-12 12:55:59 -08:00
  • 6e158e55b4 [model-gateway] improve workflow engine code quality (#16977) Simo Lin 2026-01-12 12:14:46 -08:00
  • ed729d22b3 [model-gateway] refactor workflow engine from type erasure to typed engines (#16973) Simo Lin 2026-01-12 10:47:00 -08:00
  • fa51b85466 [model-gateway] convert workflow system to type-safe workflow data (#16970) Simo Lin 2026-01-12 10:06:36 -08:00
  • 559ff9ecaf Bug: fixed multi_chain_reasoning test (#16192) Bhavneek Singh 2026-01-13 02:06:41 +09:00
  • 7b682de870 [Model] Support IQuest-Coder-40B-Loop (#16348) Gaoji Liu 2026-01-12 23:44:45 +08:00
  • d0092decb1 [model-gateway] Fix workflow engine race conditions and add graceful shutdown (#16963) Simo Lin 2026-01-12 06:37:39 -08:00
  • 76f69b7753 [diffusion] app: add ComfyUI plugin support for SGLang-Diffusion (#15271) WenhaoZhang 2026-01-12 21:58:16 +08:00
  • 9a628744fc [CI] fix piecewise graph test case on ascend (#16933) khalilzhk 2026-01-12 20:16:30 +08:00
  • 53dca74f47 Bugfix: EagleDraftWorker has not attribute "eagle_use_aux_hidden_state" (#16480) chenxu214 2026-01-12 20:14:35 +08:00
  • 2dadf63562 [diffusion] Support I2I/TI2I/I2V/TI2V warmup && T2I/T2V warmup bug fix (#16922) HuangJi 2026-01-12 19:40:37 +08:00
  • aab640c99f add doc for dsv32 cp+pp (#16916) ybyang 2026-01-12 19:14:07 +08:00
  • 2b3791ed37 Fix wrong kernel selection for int32/int64 indices (#16912) Liangsheng Yin 2026-01-12 17:26:57 +08:00
  • 9f5cd80a8d Re-introduce the unit test of test_mooncake_ep_small (#16019) Xun Sun 2026-01-12 17:01:24 +08:00
  • b1ee75ae7b [AMD] CI - enable test case for amd ci : triton_attention_kernels , torch_compile_moe (#16559) YC Tseng 2026-01-12 16:06:31 +08:00
  • f44c63eef7 [diffusion] chore: validate sampling params (#16677) Hu Chong 2026-01-12 16:00:10 +08:00
  • c54c70ab62 [model-gateway] Improve Health Check Logging (#16930) Wenyi Xu 2026-01-12 15:01:29 +08:00
  • aab906a3d4 [docs] sync diffusion docs to main docs (#16932) Adarsh Shirawalmath 2026-01-12 12:19:55 +05:30
  • feb39f7768 [diffusion] model: Support TurboWan2.2-I2V SLA && add CI test for TurboWan (#16536) HuangJi 2026-01-12 13:55:38 +08:00
  • 38b30c7b56 Tiny refactor age computation in router (#16850) fzyzcjy 2026-01-12 13:42:43 +08:00
  • c581b5ed79 [NPU] update feature supported on ascend NPU (#16915) Hexq0210 2026-01-12 11:47:58 +08:00
  • a1c48943d7 Tiny fix NoAvailableWorkers being a RetryError (#16896) fzyzcjy 2026-01-12 10:53:17 +08:00
  • 5b7bed7ca4 Decouple grammar logic out of scheduler. (#16820) Liangsheng Yin 2026-01-12 10:52:42 +08:00
  • 38a88479c6 llama model and llama eagle3 model support dp-attn (#15268) chenxu140 2026-01-12 08:54:56 +08:00
  • 503c3d9566 [model-gateway]: add qwen coder tool parser support xml format for qwen3 coder and microthinker (#12909) ybyang 2026-01-12 08:09:41 +08:00
  • 934ae89abe Tiny fix typo minxin -> mixin (#16908) Liangsheng Yin 2026-01-12 00:36:45 +08:00
  • cf1426a7b7 [CI] reapply max-parallel in stage-b-test-small-1-gpu (#16906) Liangsheng Yin 2026-01-11 23:17:38 +08:00
  • 17cb3c8e49 Enable /rerun-stage workflow URL lookup for fork PRs (#16851) Alison Shao 2026-01-11 07:05:37 -08:00
  • 2f4a6addf3 [cpu/arm64] support run sglang on arm64 cpu (#14867) Yibo Cai 2026-01-11 20:27:19 +08:00
  • f9fc50acd6 [Tiny] Rename test_sparse_flash_attn.py to fix CI (#16895) Baizhou Zhang 2026-01-11 18:18:29 +08:00
  • 3c16c58619 [model-gateway] Add Redis support as a history backend (#16300) Wenyi Xu 2026-01-11 17:03:00 +08:00
  • 7b089ae4e0 [Diffusion] Docs for Diffusers backend (#16864) Adarsh Shirawalmath 2026-01-11 13:38:26 +05:30
  • 8b5d426340 [CI]Move fa4 e2e test to 4-gpu-b200 runner (#16889) Baizhou Zhang 2026-01-11 15:53:38 +08:00
  • b5493f65be [NVIDIA] upstream FA4 (#15182) Johnny 2026-01-11 08:31:28 +01:00
  • 09e2571e2e Clarify the meaning of cpu_group / entry_rank when dp + tp is enabled. (#16876) Liangsheng Yin 2026-01-11 13:04:43 +08:00
  • c0248d6f37 [dpc]: unify DP controller load balancing and simplify dispatch logic (#16258) Ratish P 2026-01-11 10:08:03 +05:30
  • cc25f9df50 Update est_time for stage-b-test-small-1-gpu tests (#16835) Alison Shao 2026-01-10 20:03:43 -08:00
  • cf14feba4d Fix parallel tool call parsing bug when tool parameters contain arrays (#16345) Leoyzen 2026-01-11 11:31:38 +08:00
  • 7c25687c9b fix(gateway): rewrite gauge_histogram.rs for zero-allocation hot path (#16878) Simo Lin 2026-01-10 18:15:21 -08:00
  • ff97814232 Tiny fix hicache kernel backend comparison (#16867) Mohammad Miadh Angkad 2026-01-11 10:10:42 +08:00
  • d112f6a25b [Feature] Support JIT set kv cache (#16273) DarkSharpness 2026-01-11 09:34:09 +08:00
  • a2c2c09d7d [BugFix] fix gpt-oss-120b launch failure with --enable-piecewise-cuda-graph (#16757) Minglei Zhu 2026-01-10 17:19:59 -08:00
  • 2a9344d320 [tiny remove] remove torch_compile in parallel_state (#16865) Yuwei An 2026-01-10 16:04:25 -08:00
  • 78c41758ad [hot fix ci] Hot fix for the unregistered. (#16877) Liangsheng Yin 2026-01-11 01:24:32 +08:00
  • 206db66f5c tiny refactor pcg split op registration (#16863) Qiaolin Yu 2026-01-10 07:45:28 -08:00
  • 5c72be1e51 [diffusion] feat: support multiple LoRA adapters loading and application (#16667) WenhaoZhang 2026-01-10 23:32:45 +08:00
  • 76d4881794 [diffusion] improve: apply tp optim to cross-attn for wan2.2 (#16788) wxy 2026-01-10 21:25:59 +08:00
  • d1ec93e3ac Optimize layernorm_gated for Qwen3-Next (#16397) Yuan Luo 2026-01-10 20:55:31 +08:00
  • bdb76b34db [diffusion] fix: fix LoRA weight merging when using layerwise offload (#16737) WenhaoZhang 2026-01-10 20:17:35 +08:00
  • dae6a4092a Tiny add scheduler status logging (#16872) fzyzcjy 2026-01-10 20:12:24 +08:00
  • a0899bdbd8 Fix log_decode_stats_every_iteration when having TP in attention (#16871) fzyzcjy 2026-01-10 20:10:23 +08:00
  • 641830c1c2 Tiny extract file logging utils (#16870) fzyzcjy 2026-01-10 20:02:37 +08:00
  • 3fd88ea9b5 [MTP][spec_v2] Fix TRTLLM MLA backend crash in EAGLE draft_extend mode (#15790) YAMY 2026-01-10 03:58:23 -08:00
  • 145bd54f1b Piecewise Cuda Graph Memory Usage (#15927) Yuwei An 2026-01-10 03:29:13 -08:00
  • 2d088b85d9 [IDLE FORWARD][Indexer] Fix forward_idle bs mismatch issue in DeepseekV3.2's NSAIndexer (#15227) YAMY 2026-01-10 02:14:30 -08:00
  • 3a8b44fe89 Update LoRA Weights via Tensor (#16226) lg(x) 2026-01-10 17:36:43 +08:00
  • aeb480c11f Add top-p to run_eval.py (#16844) hlu1 2026-01-10 01:10:37 -08:00
  • 9fd2358cc2 Update Cutedsl version and pin cuda-python version (#16838) Baizhou Zhang 2026-01-10 17:08:43 +08:00
  • 3c358736d1 Enhance test for dp-attention + constrained decoding. (#16849) Liangsheng Yin 2026-01-10 16:49:10 +08:00
  • 6327dff242 enhance LoRA tests and fix base model LoRA eviction in Scheduler (#16333) Glen Liu 2026-01-10 03:49:00 -05:00
  • 675acecec6 Attention backend selection bug fix for hicache (#16779) Zhiqiang Xie 2026-01-10 00:22:41 -08:00
  • ad20127359 [CI] Remove duplicate code in test_mamba_ut (#16854) Yuan Luo 2026-01-10 16:16:56 +08:00
  • 7f393d9512 [Docker] Add nightly dev docker for Cuda 13 (#16862) Baizhou Zhang 2026-01-10 14:56:53 +08:00
  • 4b14f622e1 [CI] Add PD Disaggregation aarch64 test (#16572) Shangming Cai 2026-01-10 14:44:54 +08:00
  • 94fc26aad8 [Doc]Update note for Cuda 13 container usage (#16805) Baizhou Zhang 2026-01-10 14:03:19 +08:00
  • 67b61a4e8d [Rework] Add SwapAB Optimization for triton fused_moe_kernel on SM90. (#16723) Insideyyy 2026-01-10 13:57:44 +08:00
  • d27f16f38a Fix EPLB + FP4 Quantization Compatibility Issue (#13715) Shifang Xu 2026-01-10 13:38:19 +08:00
  • c89949bbaf Tiny let soft watchdog cover initialization phase (#16853) fzyzcjy 2026-01-10 13:20:08 +08:00
  • 20abaee26c [DSv32] Overlap indexer weights_proj during dual_stream decode (#16637) Ziang Li 2026-01-09 21:06:44 -08:00
  • 32a569fb77 Tiny add CPU resource monitoring for overload diagnosis (#16852) fzyzcjy 2026-01-10 12:48:18 +08:00
  • 3ed3b7ef7c Tiny add routing key distribution metrics (#16847) fzyzcjy 2026-01-10 12:06:40 +08:00
  • 1f9d4795a9 Tiny add gauge histogram abstraction for engine and router (#16848) fzyzcjy 2026-01-10 11:45:25 +08:00
  • e6d40bff81 Revert "feat: reduce constrained-decoding overhead in TP" (#16845) Liangsheng Yin 2026-01-10 11:39:38 +08:00
  • fbc128a32e fix(function_call): group batch decode by options instead of fallback (#16698) 若可 2026-01-10 11:38:13 +08:00
  • 9c64a15ad4 feat: add workflow run URL to /rerun-stage comment (#16825) Alison Shao 2026-01-09 18:41:20 -08:00
  • e91a717632 [llama] Allow passing tp_rank and tp_size into llama mlp (#16837) Yinghai Lu 2026-01-09 18:32:05 -08:00
  • 1f0ea4f958 Add routing key based schedule policy (#16840) fzyzcjy 2026-01-10 10:16:23 +08:00
  • 15da3061bc Tiny pass routing key to scheduler processes (#16839) fzyzcjy 2026-01-10 10:13:00 +08:00
  • 6406a5969b Tiny support customizing prometheus buckets for prefill delayer (#16831) fzyzcjy 2026-01-10 08:17:37 +08:00
  • cec19b56c6 Tiny add command line args for prefill delayer and unify names (#16830) fzyzcjy 2026-01-10 07:54:25 +08:00
  • ef35d8fe4e Migrate VLM tests and remove unit-test-backend-1-gpu job (#16679) Alison Shao 2026-01-09 15:24:26 -08:00
  • 08636f72b5 [Fix CI] Fix test_mamba_unittest.py (#16810) Yuan Luo 2026-01-10 07:03:47 +08:00
  • a6c29d4cbd [AMD CI] Temporarily disable docker caching. (#16783) Sai Enduri 2026-01-09 14:35:49 -08:00
  • 84ab32a2a0 Reduce some small cpu overhead in stream fetch (#16587) Yi Zhong 2026-01-09 16:53:06 -05:00
  • 7066711529 [ci hot fix] fix global server args init for mamba ut (#16821) Liangsheng Yin 2026-01-10 02:03:32 +08:00
  • bd1afeb568 [model-gateway] Restore response streaming by optimizing WASM middleware buffering (#16804) Praneth Paruchuri 2026-01-09 22:59:22 +05:30
  • 8ef5b90528 Fix GLM-4.7 MoE Detector complex JSON Schema type parsing (#15753) Leoyzen 2026-01-10 01:21:10 +08:00