Commit Graph

  • 9325f94534 [model-gateway] add ModelCard and ProviderType for model configuration (#14237) Simo Lin 2025-12-01 09:47:51 -08:00
  • 41b7aab848 Disable Deepep 8 GPU tests (#14152) Kangyan-Zhou 2025-12-01 09:37:38 -08:00
  • 491f4fe8e1 Modify git tag for DeepGemm in sgl-kernel. (#14179) Sulfur6-L8972 2025-12-02 01:01:22 +08:00
  • 34035d8cd9 chore: bump sgl-kernel version to 0.3.18.post2 (#14229) sglang-bot 2025-12-01 08:50:40 -08:00
  • 92ca629582 [model-gateway] add ModelType bitflags and Endpoint enum for worker (#14230) Simo Lin 2025-12-01 08:14:44 -08:00
  • ec92d7f14e [model-gateway] fix v1/models response format to be oai compatible (#13693) Chang Su 2025-12-01 07:20:38 -08:00
  • 79b389da6b [model-gateway] refactor oai router 1/n (#14228) Simo Lin 2025-12-01 07:16:35 -08:00
  • 3de09aadbc Add new moe wna16 marlin gemm (#14122) Xiaoyu Zhang 2025-12-01 23:07:53 +08:00
  • 2e8f54e61e [spec-overlap] bugfix for pd disaggregation and npu (#14088) liupeng374 2025-12-01 22:58:20 +08:00
  • e55731b6e8 [model-gateway] Avoid logging MCP connection token (#13887) Wenyi Xu 2025-12-01 20:44:31 +08:00
  • 45264554f3 Super tiny fix typo (#14219) fzyzcjy 2025-12-01 20:19:17 +08:00
  • c4293f59ac Change PR test schedule to run every 6 hours (#14218) Lianmin Zheng 2025-12-01 03:50:58 -08:00
  • a2423052f6 Add cuda event based on waiting value (#14214) Liangsheng Yin 2025-12-01 18:51:44 +08:00
  • bc3d2a85af [Minor] update docs (#14212) Lianmin Zheng 2025-12-01 02:33:58 -08:00
  • d815d00248 Tiny call cudaProfilerStart only on first rank in node (#14211) fzyzcjy 2025-12-01 18:18:45 +08:00
  • fa9021b21f fix: Increase FlashInfer workspace size for Qwen3VL models (#14173) Xiaoyu Zhang 2025-12-01 17:54:23 +08:00
  • 9c80072845 Add peak output tokens per second in bench_serving (#14165) Xiaoyu Zhang 2025-12-01 17:47:54 +08:00
  • 630a693081 [VLM] Boost Memory Pool based CUDA IPC (#14123) Yuan Luo 2025-12-01 17:17:46 +08:00
  • 7ce8faae28 [diffusion] refactor: remove hard-code of instanceof on PipelineConfig (#14186) Mick 2025-12-01 16:35:34 +08:00
  • de153cf76a Fix speculative decoding error when retracting (#14180) fzyzcjy 2025-12-01 15:30:13 +08:00
  • f4a0c5c76b Try to remove wrong logic about max total token in spec decoding (#14167) fzyzcjy 2025-12-01 15:29:58 +08:00
  • 0f8e53947d [Piecewise] Use same global graph memory pool as the main cuda graph … (#14044) Binyao Jiang 2025-11-30 23:04:10 -08:00
  • e8ba5a668c Support profiling only prefill or decode without the other (#14182) fzyzcjy 2025-12-01 14:46:30 +08:00
  • a2960bdd6b Super tiny allow millisecond precision in logging (#14183) fzyzcjy 2025-12-01 14:46:09 +08:00
  • 487c8d4df3 Tiny add several args to bench serving (#14181) fzyzcjy 2025-12-01 14:45:47 +08:00
  • f87b8eab23 Tiny fix transform_scale_ue8m0 wrong output in some scenarios (#14003) fzyzcjy 2025-12-01 14:45:27 +08:00
  • e8542db558 [piecewise] move piecewise_cuda_graph_runner init to model_runner initialize (#14034) Minglei Zhu 2025-11-30 22:16:04 -08:00
  • 6df1e8d628 [Auto Sync] Update backend.py (20251130) (#14153) Lianmin Zheng 2025-11-30 22:15:02 -08:00
  • 4addb60274 Pull Request Instructions: RL and Training Framework Integrations (#14187) Richard Chen 2025-12-01 00:58:54 -05:00
  • bd0e690857 [Feature] Enable PTPC FP8 for compressed tensors moe (aiter kernel) (#12181) qichu-yun 2025-12-01 13:54:28 +08:00
  • 0825d7f4c6 [piecewise] Refactor VLM to support input embed buffer and remove external embedder hack (#14155) Byron Hsu 2025-11-30 21:43:09 -08:00
  • 0b9dbea593 [diffusion] chore: improve z-image (#14104) Yuhao Yang 2025-12-01 12:26:17 +08:00
  • 982db4ebac Feat: GLM-4.6 supports shared experts fusion (#13873) Uranus 2025-12-01 11:33:18 +08:00
  • f5f3a5d98c [PD] Support json file configuration for Transfer Engine (#14059) Teng Ma 2025-12-01 10:47:33 +08:00
  • f138ae5789 [model-gateway] support VL models in router (#14140) ooapex 2025-12-01 09:50:22 +08:00
  • decb48965d [DeepSeekV3.2] Enable pure TP & Partial DP Attention (#13646) YAMY 2025-11-30 15:59:23 -08:00
  • c72f0756d2 Fix: fix flashmla fp8 kv cache acc error (#13841) Fan Yin 2025-12-01 05:38:19 +08:00
  • f1115cf58d Revert "[Minor]Raise Error when deepep num dispatch token per rank is smaller than cuda graph bs" (#14171) Baizhou Zhang 2025-11-30 13:49:46 -07:00
  • 412160f4c1 [sgl-kernel] fix b200 kernel ci (#13907) Fan Yin 2025-12-01 02:15:37 +08:00
  • 7b03cc6482 [Minor]Raise Error when deepep num dispatch token per rank is smaller than cuda graph bs (#14065) Baizhou Zhang 2025-11-30 11:11:42 -07:00
  • 9872a677b4 [ci]fix deepep import error on H20 action (#14166) Hank Han 2025-12-01 01:42:57 +08:00
  • c15c864b6f Fix LMCache unit test and init bug (#14005) Dongjoo Seo 2025-11-30 07:57:32 -08:00
  • dc7bdc7329 bugfix[schedule]: Excessive preemption occurs when preempting running requests to schedule new prefill requests. (#12494) PiteXChen 2025-11-30 22:29:26 +08:00
  • 0a9d64530d Support grammar + spec + reasoning (#14163) Liangsheng Yin 2025-11-30 21:19:57 +08:00
  • 340c613ab5 Support numactl bind for CPU and memory before process starts (#14156) fzyzcjy 2025-11-30 17:00:33 +08:00
  • 36b729c2b8 Implement profiler v2 and fix stage mixture bug (#14148) fzyzcjy 2025-11-30 16:59:52 +08:00
  • 67e6ef4b2d feat: longcat flash add aux layers capture for eagle3 (#14161) Tianhao Zhou 2025-11-30 00:50:55 -08:00
  • 65ba5ab8b1 add cpp files for cpp_radix_tree to pyproject.toml. (#14052) strgrb 2025-11-30 13:05:04 +08:00
  • 990023e59b [diffusion] lora: Fix LoRA weight merging for torch.nn.Linear layers from diffusers modules (#14150) WenhaoZhang 2025-11-30 12:44:12 +08:00
  • 5ddd2f6b6b Always run model evaluation even if the trace upload step fails (#14157) Kangyan-Zhou 2025-11-29 20:13:59 -08:00
  • 0ae4b1ad81 Show errors when misusing env variables (#14154) fzyzcjy 2025-11-30 10:57:35 +08:00
  • 94cd64a7b0 Support checking fp8 params in weight_checker (#14147) fzyzcjy 2025-11-30 09:08:59 +08:00
  • b870271a50 Fix spec v2 does not support RL update weights from tensor (#14146) fzyzcjy 2025-11-30 09:08:05 +08:00
  • 22ee9b0111 Super tiny add more info in dumper (#14145) fzyzcjy 2025-11-30 09:07:39 +08:00
  • 9d0e5f1f74 Tiny fix DeepGEMM precompile rank check (#14136) fzyzcjy 2025-11-30 09:07:17 +08:00
  • 1d3d8b3418 Fix Minimax M2 loading issue (#13956) Kangyan-Zhou 2025-11-29 14:07:19 -08:00
  • 155a9e7237 Fix condition for streaming output_ids in tokenizer manager (#13759) Lianmin Zheng 2025-11-29 13:56:15 -08:00
  • d7cb08c5be Always run all stages in cron based PR tests (#14151) Kangyan-Zhou 2025-11-29 12:32:06 -08:00
  • 3339c81072 fix RuntimeError: RMSNorm failed with error code an illegal memory access was encountered (#14135) gongwei-130 2025-11-29 12:17:41 -08:00
  • d6c88d519d Trigger PR test on main every 3 hours instead of push event (#14130) Kangyan-Zhou 2025-11-29 10:12:50 -08:00
  • f03ea34a3d add runtime check for PyTorch 2.9.1 + CuDNN < 9.15 to prevent Conv3d performance issues (#14119) Yuhao Yang 2025-11-29 23:05:54 +08:00
  • 4cafc835d3 Super tiny fix typo (#14131) fzyzcjy 2025-11-29 21:08:31 +08:00
  • c6a52f4411 [diffusion] chore: add resolution shortcuts for sampling params (#14129) Mick 2025-11-29 18:00:21 +08:00
  • c6d34a0688 Move piecewise cuda graph test to manual dir to fix CI (#14121) Shangming Cai 2025-11-29 16:52:01 +08:00
  • 848ee57067 feat: support flashinfer kernel autotune (#12306) elvischenv 2025-11-29 16:05:37 +08:00
  • ce6b7dfce7 Add auto-tune workflow (#14124) Lianmin Zheng 2025-11-28 23:23:53 -08:00
  • 0fe74af563 Remove incorrect deep_gemm assertions from server_args.py (#14113) Cheng Wan 2025-11-28 20:25:39 -08:00
  • 0a362d653f [diffusion] log: unify generation performance logging (#14117) Mick 2025-11-29 12:21:59 +08:00
  • 143b57b805 enable piecewise cuda graph for prefill server (#13377) fjybiocs 2025-11-29 12:09:26 +08:00
  • f446b51c41 fix: malformed KV events for NVIDIA Dynamo (#13488) Yan Ru Pei 2025-11-28 14:55:20 -08:00
  • a102a0507a Disable Deepep 2 GPU tests (#14111) Kangyan-Zhou 2025-11-28 13:56:34 -08:00
  • 6bad6a3655 [model-gateway] Add version command support to SMG (#12558) Tony Lu 2025-11-29 04:34:19 +08:00
  • 11b6217aee Fix NIXL OBJ desciptors (#10712) Tomer Shmilovich 2025-11-28 21:32:07 +02:00
  • 0b0b2607ca [CPU] Apply uv as package manager (#14106) Zaili Wang 2025-11-29 02:36:53 +08:00
  • 841eb29d3d [diffusion] model: support z-image (#14067) Yuhao Yang 2025-11-28 21:48:31 +08:00
  • 45cf575852 Fix overlap scheduler not take effect when outputing logprobs (#14096) fzyzcjy 2025-11-28 18:15:56 +08:00
  • 0e8ce1e832 [diffusion] refactor: clean useless files (#14094) Mick 2025-11-28 18:14:00 +08:00
  • ea1e9f6b3c feat: support qwen3_vl vision model dp (#13724) Lzhang-hub 2025-11-28 17:29:07 +08:00
  • f6e37d3edb [Bugfix] qwen2.5-vl spec decode accept_len low (#13904) Lzhang-hub 2025-11-28 17:26:32 +08:00
  • ab9a46d462 Support configuring the request limit per receiving poll (#14076) vipwangerxiao 2025-11-28 16:14:21 +08:00
  • 621061f017 [Bugfix] input prompt was not logged (#13936) shuwenn 2025-11-28 16:00:51 +08:00
  • 7daddcdb58 Fix structural_tag tool call with null schema (#14006) Aleksandr Krotov 2025-11-28 10:04:16 +03:00
  • 91d249cd9b fix: small changes to enable test_mrope.py (#14082) Raayan Dhar 2025-11-27 20:58:17 -10:00
  • 951028968c [diffusion] refactor: refactor ComponentLoader and support loading native models from diffusers and transformers (#13205) Mick 2025-11-28 14:17:32 +08:00
  • 3543a04a48 [diffusion] refactor: refactor condition image resize logic (#14079) Mick 2025-11-28 14:06:34 +08:00
  • e12c78aab6 [sgl-kernel][1/2] Fused qk_norm_rope for Qwen3-MoE (#14036) Yuan Luo 2025-11-28 12:25:15 +08:00
  • bce40fa217 Fix utils import issue for nightly tests (#13944) alisonshao 2025-11-27 15:24:27 -08:00
  • 051ad83347 [chore] Arrange NV packages in Dockerfile (#13749) Baizhou Zhang 2025-11-27 18:08:27 -05:00
  • 4c9f7c97d3 Temporarily disabled test (#14069) Kangyan-Zhou 2025-11-27 14:58:46 -08:00
  • 63b056213f Remove disused B300 Dockerfile (#13946) Mohammad Miadh Angkad 2025-11-27 22:48:00 +08:00
  • 21af8e73ad Super tiny add comments to SGLANG_DEEPEP_NUM_MAX_DISPATCH_TOKENS_PER_RANK (#14048) fzyzcjy 2025-11-27 22:16:43 +08:00
  • 7ab548ef64 [2/2] Refactor DeepGeem requant for FP8 FusedMoE on Blackwell (#13960) Baizhou Zhang 2025-11-27 09:00:26 -05:00
  • bab033b970 Adjust max-parallel for CUDA CI (#14057) Liangsheng Yin 2025-11-27 20:48:37 +08:00
  • 6350042696 feat: Naive support Spec V2 + Constrained Decoding (#13425) Yixin Dong 2025-11-27 04:31:46 -08:00
  • 25758647b1 Support sanity checking weight consistency especially for RL (#13854) fzyzcjy 2025-11-27 20:25:12 +08:00
  • 2bc8ee8b74 Tiny support 3D tensors in inverse_transform_scale_ue8m0 (#14002) fzyzcjy 2025-11-27 20:20:45 +08:00
  • ab843ced31 [Feat]Add scheduler recv skipper weights to environment configuration (#13855) Jimmy 2025-11-27 18:16:11 +08:00
  • 6edffc6391 [diffusion] perf: improve black-forest-labs/FLUX.2-dev (#14040) Mick 2025-11-27 14:49:52 +08:00
  • 077ca70ee4 [Intel XPU]Add xpu support for get_device_memory_capacity (#13895) gaopengff 2025-11-27 12:55:52 +08:00
  • 7cb04dc0e5 Use trtllm mha decode kernel for target_verify in speculative decoding (#13976) Qiaolin Yu 2025-11-26 20:40:34 -08:00