Commit Graph

  • 079ac237da [diffusion] fix: fix gen video doc (#14409) R0CKSTAR 2025-12-04 22:05:38 +08:00
  • 8428078436 Add Mistral Large 3 support. (#14213) Daniel Cámpora 2025-12-04 13:00:05 +01:00
  • af35023e65 [bug fix] fix ima with get_mla_kv_buffer_kernel overflow (#14224) Xuchun Shang 2025-12-04 17:20:11 +08:00
  • cb8df87fc1 [1/2] Add rope kernel in sgl-kernel (#14334) Qiaolin Yu 2025-12-04 00:45:44 -08:00
  • e3ab23c1a6 Try to fix B200 DeepEP error (#14399) fzyzcjy 2025-12-04 16:41:24 +08:00
  • 70d2587324 [CPU] Optimize small oc GEMM for Qwen3-next on CPU (#12446) jianan-gu 2025-12-04 16:38:47 +08:00
  • 894c0dc57c [NPU][1/N] NPU basic functions refactor and new modelslim quant type (#13359) Even Zhou 2025-12-04 16:15:31 +08:00
  • d6c490192d [AMD] fix the regression issue for DeepseekV3 on MI300 (#14383) yctseng0211 2025-12-04 15:30:11 +08:00
  • fa78c44a3d [model-gateway] introduce provider in openai router (#14394) Simo Lin 2025-12-03 22:47:20 -08:00
  • 654a78f9c8 [model-gateway] add phi3 vision image processor (#14381) Simo Lin 2025-12-03 22:32:40 -08:00
  • f90b400431 [CPU] add support for mamba causal conv1d for qwen3-next (#12309) Ma Mingfei 2025-12-04 13:41:42 +08:00
  • 78647e08dd [model-gateway][doc] Add STDIO Explicitly to Example in README (#14393) Wenyi Xu 2025-12-04 12:36:44 +08:00
  • 46f21a5956 use faster covnersion from float8_e4m3fn to bfloat16 (#12316) Ma Mingfei 2025-12-04 12:34:05 +08:00
  • b2b09f5f24 [VLM] Introduce Cache for positional embedding ids for Qwen-VL family (#14292) Yuan Luo 2025-12-04 12:32:00 +08:00
  • 04df80a9a1 Support PP x PD decode with nixl backend (#14392) Kevin Li 2025-12-03 20:22:17 -08:00
  • 4f73e53dcc [CPU] document updates (#14272) Zaili Wang 2025-12-04 11:56:06 +08:00
  • 16ff892c18 [sgl-kernel][Feat][B200][1/N] Support MXFP8 Grouped GEMM in Blackwell (#13731) Qi Yuhang 2025-12-04 10:09:09 +08:00
  • df026bb110 Fix sgl-router silently parse selector wrongly causing OME fail to discover pods (#14359) fzyzcjy 2025-12-04 09:56:37 +08:00
  • d42c167bfd [model-gateway] add qwen3_vl model image processor (#14377) Simo Lin 2025-12-03 17:16:37 -08:00
  • 388151053d [model-gateway] use worker crate in openai router (#14330) Simo Lin 2025-12-03 13:36:32 -08:00
  • 9d82340298 Revert "Revert "enable csgmv automatically on cuda"" (#14277) b8zhong 2025-12-03 13:12:30 -08:00
  • 8ab5d8b40d [model-gateway] add qwen2.5_vl model image processor (#14375) Simo Lin 2025-12-03 12:51:30 -08:00
  • 03575ce38d [model-gateway] add qwen2_vl model image processor and tests (#14374) Simo Lin 2025-12-03 12:34:06 -08:00
  • 80518bea65 Fix validation to detect missing model files before loading (#14253) alisonshao 2025-12-03 11:36:07 -08:00
  • 7e78825d5a [Tiny]Small fixes in deepseek v32 doc (#14372) Baizhou Zhang 2025-12-03 11:35:40 -08:00
  • 5bbd83a2c8 ci: Migrate AMD workflows to new MI325 runners; temporarily disabled failed CI's to be added back (#14226) sunxxuns 2025-12-03 14:33:27 -05:00
  • abf6272bdd [model-gateway] add llava model image processor and tests (#14371) Simo Lin 2025-12-03 11:18:00 -08:00
  • 46d7b35ec7 Move custom_ops under layers; move _custom_ops.py → custom_all_reduce_ops.py (#14326) Lianmin Zheng 2025-12-03 10:33:37 -08:00
  • 20aad5b5ab Single Batch Overlap for MoE Models (#9660) Sulfur6-L8972 2025-12-04 02:07:42 +08:00
  • 974c562a25 [CPU] add fused_qkvzba_split_reshape_cat kernel for Qwen3-next (#12330) blzheng 2025-12-03 23:46:08 +08:00
  • aca0d01d3f [diffusion] doc: add vae path to cli doc#14004 (#14355) Dongjie Zou 2025-12-03 10:13:25 -05:00
  • dc1635023f [NPU][Doc] updated installation guide for Ascend NPU (#13585) VDV1985 2025-12-03 16:58:49 +03:00
  • 24903b88ba Tiny adjust CI testcases (#14362) Liangsheng Yin 2025-12-03 21:19:51 +08:00
  • 443d7bcd83 [Ascend] fix AscendAttnMaskBuilder bug to support float16 models (#14271) Michelle Wu 2025-12-03 19:37:16 +08:00
  • 16d8de2284 [bugfix] NpuFuseEPMoE miss initialization parameters (#14295) chenxu140 2025-12-03 19:36:41 +08:00
  • d122e32467 [NPU] bug fix: w_vc need contiguous for NPU batch_matmul_transpose ops (#13980) ZhengdQin 2025-12-03 19:35:18 +08:00
  • 96cc10834a [CI] update estimated elapsed time of some unittests (#14347) Cheng Wan 2025-12-03 01:21:40 -08:00
  • 77512ae0d7 [bugfix] Fix prefill tbo disabled when --deepep-mode=auto (#14333) Yuhao Yao 2025-12-03 17:20:33 +08:00
  • c233e9d7a9 [CPU] Support chunk_gated_delta_rule kernel for Qwen3-Next (#12441) Xuan Liao 2025-12-03 17:03:48 +08:00
  • 58ac3f317b [model-gateway] add image processor and transformer structure (#14344) Simo Lin 2025-12-03 00:04:06 -08:00
  • 93452a8252 [PD] Support decode pp for PD disaggregation (#14265) Shangming Cai 2025-12-03 14:35:29 +08:00
  • 65c8568c4a sync attention, deepseek doc (#14335) b8zhong 2025-12-02 21:19:40 -08:00
  • 4bcc5879af [Doc] Fix DeepSeek V32 Doc (#14336) Baizhou Zhang 2025-12-02 21:06:55 -08:00
  • d5ea8c7176 [model-gateway] multimodality initialization (#13350) Simo Lin 2025-12-02 19:41:40 -08:00
  • 42271376d1 [bug fix] use npu phy id in container env (#14266) Vikram 2025-12-03 11:33:43 +08:00
  • 84e0abb7b1 Update CODEOWNERS for multimodal (#14329) Mick 2025-12-03 11:04:40 +08:00
  • 043f13171f [Performance] Optimize NSA Indexer K/S Buffer Access with Fused Triton Kernels (#13812) Johnsonms 2025-12-02 18:53:06 -08:00
  • f764c6910d [diffusion] feat: support distilled vae generic (#14195) Dongjie Zou 2025-12-02 21:27:31 -05:00
  • 922054079c [Doc] Update DeepSeek-V3.2 document (#14321) Baizhou Zhang 2025-12-02 18:19:39 -08:00
  • 7d1a130cde Refactor custom allreduce logics (#13710) Even Zhou 2025-12-03 09:20:05 +08:00
  • 7ae368efde chore: bump SGLang version to 0.5.6 (#14316) sglang-bot 2025-12-02 17:17:13 -08:00
  • c4e20caded ci: Add zyzshishui to CI permissions (#14324) sunxxuns 2025-12-02 20:00:31 -05:00
  • 5dad1ff1b6 [model-gateway] add workflow for external model providers (#14323) Simo Lin 2025-12-02 16:47:06 -08:00
  • ca52ed425f Clean up imports and move files (#14317) Lianmin Zheng 2025-12-02 16:31:54 -08:00
  • 5c03aa3e9d Adding section for scheduled PR test runs on main (#14309) Douglas Yang 2025-12-02 14:52:44 -08:00
  • 253be18e52 Fix nonetype error for ci failure monitor (#14319) Douglas Yang 2025-12-02 14:24:25 -08:00
  • 084b06e79d Add /rerun-stage slash command to rerun specific PR test stages (#14262) alisonshao 2025-12-02 14:23:38 -08:00
  • fc6fb550ec [Minor] update docs on CI (#14315) Lianmin Zheng 2025-12-02 12:06:54 -08:00
  • 7c38eca1e4 feat: DeepSeek new v3.2 encoding (#14249) Eva20150932-atlascloud 2025-12-02 19:41:05 +00:00
  • 427b08e24d Init TBO with dp_padded batch (#11423) Quanfeng Li 2025-12-03 02:34:26 +08:00
  • 0141ca370f Revert PR #14044: Restore separate memory pool for piecewise CUDA graph (#14278) alisonshao 2025-12-02 09:53:16 -08:00
  • 25a6be4930 Fix duplicate download log messages in multi-process environment (#14299) alisonshao 2025-12-02 09:33:18 -08:00
  • df1f31241b [sgl-kernel] fix runtime error while preloading CUDA runtime (#13089) Antonin Vidon 2025-12-02 10:03:51 -05:00
  • c5947ecd85 Opt moe align block size kernel (#14133) Xiaoyu Zhang 2025-12-02 19:13:55 +08:00
  • 9530b76630 [diffusion] refactor: simplify DmdDenoisingStage (#14269) Mick 2025-12-02 18:59:40 +08:00
  • 3067b3f050 [diffusion] chore: improve model info registration and searching strategy (#14281) Jinyan Chen 2025-12-02 18:28:59 +08:00
  • e0ec42c710 [CI] Fix 4-GPU test timeout by using 3 partitions (#14287) alisonshao 2025-12-02 00:47:26 -08:00
  • 9c9d7091cb Update CODEOWNERS for multimodal_gen (#14286) Mick 2025-12-02 16:15:19 +08:00
  • 51a86ce69e [model-gateway] change rust package name to sgl-model-gateway instead (#14283) Simo Lin 2025-12-02 00:06:54 -08:00
  • 64092c8b55 [Auto Sync] Rename is_hybrid to is_hybrid_swa (#14252) Lianmin Zheng 2025-12-01 23:24:24 -08:00
  • 63b9300f00 chore: bump sgl-kernel version to 0.3.18.post2 (#14244) sglang-bot 2025-12-01 23:14:12 -08:00
  • 21ec99beff [VLM][Doc] Document for VLM DP Encoder (#14279) Yuan Luo 2025-12-02 15:08:37 +08:00
  • 383689e3ad [model-gateway] fix version output (#14276) Simo Lin 2025-12-01 22:51:57 -08:00
  • 236a7c2370 fix trtllm mla spec (#13738) b8zhong 2025-12-01 22:16:25 -08:00
  • 3dabd609fb Optimize topk sigmoid in minimax_m2 (#14047) Roger Young 2025-12-02 14:07:12 +08:00
  • 73df525382 [model-gateway] include smg version command in py binding (#14274) Simo Lin 2025-12-01 21:21:02 -08:00
  • e6420100ee sync attention doc and ep doc to doctree (#14257) b8zhong 2025-12-01 21:15:22 -08:00
  • c9e2090101 fix: Support PP for Mistral Small 3.1 (#14254) Kevin Li 2025-12-01 21:04:14 -08:00
  • 106df4eac5 Fix mrope_positions size when req is retracted (#13700) kun-llfl 2025-12-02 11:38:20 +08:00
  • 427a19b643 Remove cargo config also in .zshenv (#14267) Liangsheng Yin 2025-12-02 11:31:53 +08:00
  • 1f930cd23d [diffusion] CI: add testcase-wise retry mechanism (#14261) Mick 2025-12-02 11:06:12 +08:00
  • 11ce05163d Fix NIXL exception message (#14172) Kartik Ramesh 2025-12-01 20:39:45 -06:00
  • cd4151abc7 [model-gateway] add audio and moderation in model card (#14263) Simo Lin 2025-12-01 18:36:38 -08:00
  • 8fe8b63576 Revert "Try to remove wrong logic about max total token in spec decoding" (#14259) Stefan He 2025-12-01 18:18:03 -08:00
  • 8a7b1b8301 [Docs] Update CI docs (#14260) Lianmin Zheng 2025-12-01 18:15:03 -08:00
  • 26aebf83d3 [VLM] Support Piecewise CUDA Graph for Qwen3-Omni-MOE (#14222) Yuan Luo 2025-12-02 10:12:10 +08:00
  • 3ab8ae6847 [diffusion] fix: fix Flux.2 condition image resize (#14232) Mick 2025-12-02 10:05:44 +08:00
  • 03888b9de5 [Minor] Upgrade cutedsl version in Dockerfile (#13968) Baizhou Zhang 2025-12-01 17:15:26 -08:00
  • 796d82b107 [Auto Sync] Add max_total_num_tokens metric: Update scheduler_metrics_mixin.py, collector.py (20251202) (#14256) Lianmin Zheng 2025-12-01 16:34:34 -08:00
  • 1da59e8304 [Auto Sync] optionally disable fake register in Update fp8_kernel.py (20251202) (#14255) Lianmin Zheng 2025-12-01 16:11:12 -08:00
  • 1d66a14c2e [model-gateway] Add e2e tests of streaming events and tool choice for response api (#13880) Xinyue Zhang 2025-12-01 15:27:12 -08:00
  • 02af51e4fc Support fp4 fp8 non gated moe (#13794) TomerBN-Nvidia 2025-12-02 01:26:28 +02:00
  • eb5008846a [CI] Fix test_deepep_large.py (#14247) Baizhou Zhang 2025-12-01 15:18:48 -08:00
  • 079b173853 Fix a distributed initialization error (#13843) Zhiyu 2025-12-01 15:10:05 -08:00
  • 57f933fd7d [model-gateway] Migrate Worker trait to model-aware methods (#14250) Simo Lin 2025-12-01 14:32:09 -08:00
  • 1f2b84d28d Fix NSA Bug in Centralize NSA Dispatch Logic (#14245) YAMY 2025-12-01 13:18:18 -08:00
  • e7d6027e4a [model-gateway] add ModelCard support to WorkerMetadata (#14243) Simo Lin 2025-12-01 13:04:09 -08:00
  • 07821352fb Revert "Skip weight loading in deepgemm compilation" (#14241) ishandhanani 2025-12-01 12:59:09 -08:00
  • edbeaf3b88 [MM][style] rename inputs_embeds to input_embeds for consistency (#14240) Byron Hsu 2025-12-01 11:36:51 -08:00
  • d9dca28247 Update pr-test.yml to fix unknown job name deepep-8-gpu Kangyan-Zhou 2025-12-01 10:02:29 -08:00