Commit Graph

  • c56fc42430 Update quantization.md with new model resources (#13677) 赵晨阳 2025-11-20 15:50:16 -08:00
  • 3f1cfd87b6 Super tiny remove unused MiniMaxM2MLP class (#13659) fzyzcjy 2025-11-21 07:35:12 +08:00
  • b5344b31b8 [Piecewise CUDA Graph] Fix recompile issue for Mixtral and Grok2 (#13667) Stefan He 2025-11-20 14:20:11 -08:00
  • 6b262ac839 Test reorganization: Move tests to manual/ (#13610) alisonshao 2025-11-20 13:41:58 -08:00
  • ada8ce1fd0 allow loras to be implicitly evicted and loaded based on max_loaded_loras (#11526) Glen Liu 2025-11-20 16:34:32 -05:00
  • 5a2c70396e Add nightly test CI monitor workflow (#13038) alisonshao 2025-11-20 12:55:13 -08:00
  • 42028af614 enable csgmv automatically on cuda (#13600) b8zhong 2025-11-20 12:53:02 -08:00
  • 7291c72e57 [Deepseek V3.2] Change indexer weights_proj to fp32 (#13459) hlu1 2025-11-20 12:24:10 -08:00
  • 6bc3062894 Fix launch of Olmo3 (#13666) Vincent Zhong 2025-11-20 15:04:30 -05:00
  • fa92441027 [DeepseekV3.2] Deepseek fp8 support for MHA path (#12964) YAMY 2025-11-20 11:13:36 -08:00
  • fc9efdcb98 Adding nightly tests as release guard for bot bump workflows (#13655) Douglas Yang 2025-11-20 10:04:36 -08:00
  • 2dec555d36 [10/n] decouple quantization impl from vllm dependency - fix import (#13524) Fan Yin 2025-11-21 01:45:51 +08:00
  • acde21d8d5 Add fused_rmsnorm_gated_cpu kernel for CPU to support Qwen3-Next (#11577) YanbingJiang 2025-11-21 01:33:31 +08:00
  • 4528cb7d40 [CI] apply pr-gate for XPU (#13663) Liangsheng Yin 2025-11-20 23:07:04 +08:00
  • a352e833c4 CI: Kill zombie diffusion processes in CI & minor code style fix on rotary embedding fallback (#13637) Lianmin Zheng 2025-11-20 05:57:01 -08:00
  • 852eb6ce2a [CI] optimize CI workflow info (#13634) Liangsheng Yin 2025-11-20 18:12:51 +08:00
  • 7af9b88c6c Revert "[Feature] Introduce JIT Kernel in sglang (with hicache JIT kernel)" (#13644) Lianmin Zheng 2025-11-20 02:11:12 -08:00
  • 2847e5c4b4 fix bench_speculative bug (#13197) Lzhang-hub 2025-11-20 17:09:04 +08:00
  • c8ede0e93c [ROCM] Optimized deepseek-r1 fp8 model with + triton_gemm_a8w8 + batch_gemm_a8w8 + fused set_mla_kv_buffer kernel (#13617) yctseng0211 2025-11-20 16:29:56 +08:00
  • 19729f723e [CI] Align metric units for CI rate limit (#13633) Liangsheng Yin 2025-11-20 16:25:57 +08:00
  • b51f9bbee7 [Feature] Introduce JIT Kernel in sglang (with hicache JIT kernel) (#13453) DarkSharpness 2025-11-20 16:03:32 +08:00
  • 4a8442af1b diffusion: improve baseline performance monitor (#13614) Mick 2025-11-20 14:29:56 +08:00
  • 7dcf910d53 Add support for new aiter version (AR accuracy, is_shuffled PR) (#13554) Thomas Wang 2025-11-20 14:17:57 +08:00
  • c7b37b7074 [diffusion] fix: remove multimodal_gen redundant get_bool_env_var func (#13583) joesun 2025-11-20 14:05:29 +08:00
  • 10e0b83a4c Add FP32 dtype support for RoPE - Part2 (#13328) iLeGend 2025-11-20 13:19:53 +08:00
  • bc42c8c415 [diffusion] refactor: refactor pipeline folders (#13253) Mick 2025-11-20 12:56:51 +08:00
  • 2e3a69ae05 [diffusion] refactor: remove PreprocessorConfig (#13248) Mick 2025-11-20 12:53:26 +08:00
  • 21370ef7b5 [UT] Destroy process group after broadcast to resolve port occupation issues in multi-server tests (#12379) Zeyu Li 2025-11-20 12:51:58 +08:00
  • c3c4da71fb [NVIDIA] Add fp8 gemm benchmark on blackwell (#13528) Kaixi Hou 2025-11-19 19:35:00 -08:00
  • dc69462456 [CI fix] Fix image download failures in VLM CI tests (#13613) Xiaoyu Zhang 2025-11-20 11:18:06 +08:00
  • 48ca9f7518 feat: support external custom models (#13429) StonyPort 2025-11-20 11:16:58 +08:00
  • 127d59cd2c [diffusion] CI: improve diffusion CI (#13562) Mick 2025-11-20 10:54:13 +08:00
  • af6bcadcf7 [VLM] Support Piecewise CUDA Graph for Qwen2.5-VL (#13055) Yuan Luo 2025-11-20 10:23:44 +08:00
  • 67fca6b297 [GDN] Remove unnecessary conv state clone (#13603) Binyao Jiang 2025-11-19 16:09:12 -08:00
  • f88b2aa6af [GDN] Remove unnecessary contiguous() (#13604) Binyao Jiang 2025-11-19 16:07:49 -08:00
  • 17b24aca6e [Auto Sync] Update base_grammar_backend.py, collector.py (20251116) (#13357) Lianmin Zheng 2025-11-19 16:05:50 -08:00
  • bfaf0b8607 chore: bump sgl-kernel version to 0.3.17.post2 (#13570) sglang-bot 2025-11-19 14:02:57 -08:00
  • 83756a4b33 add https://github.com/netanel-haber to CI_PERMISSIONS.json (#13577) Netanel Haber 2025-11-19 23:20:30 +02:00
  • a355794905 Expend compatibility check for all quantized MoE models (#13465) Xinyuan Tong 2025-11-19 17:24:27 +00:00
  • e72cf13693 Support moe topk sigmoid kernel (#13049) Roger Young 2025-11-20 00:24:37 +08:00
  • 196b940aed [3/N] CI refactor: move some manually triggered tests. (#13448) Liangsheng Yin 2025-11-19 23:06:53 +08:00
  • a1e1e533b9 Tiny enhance test suites sanity check (#13589) Liangsheng Yin 2025-11-19 22:43:58 +08:00
  • d4a4dcdfb3 [NPU] Adapt pr-gate for pr-test workflow & workflows refresh (#13567) Even Zhou 2025-11-19 21:55:17 +08:00
  • 97ba2c2d1e [CI] Fix CUDA workflow's dependency. (#13568) Liangsheng Yin 2025-11-19 19:06:42 +08:00
  • e197bef5ce [RL] Allow passing tensors of different dtypes for FlattenedTensorBucket (#13413) Zilin Zhu 2025-11-19 17:24:48 +08:00
  • 8900f996aa CI Failure Monitor Improvements (#13558) Douglas Yang 2025-11-18 22:55:38 -08:00
  • ba9102f924 Add lmsys/gpt-oss-20b-bf16 to model validation check (#13557) Liangsheng Yin 2025-11-19 14:23:01 +08:00
  • b638abbae6 chore: bump sgl-kernel version to 0.3.17.post2 (#13542) sglang-bot 2025-11-18 22:08:23 -08:00
  • 3798055929 purge unnecessary env variable set in deterministic test (#13481) Minglei Zhu 2025-11-18 20:43:51 -08:00
  • 075ba74dd4 [Doc] Update HiCache and Mooncake docs & Mooncake Setup Error Checking (#12740) ykwd 2025-11-19 12:03:27 +08:00
  • f7be98e113 [CI] fix amd 1 gpu basic test (#13551) Liangsheng Yin 2025-11-19 11:47:41 +08:00
  • 9a1a9a4209 [AMD CI] Local cache fallback. (#13452) Sai Enduri 2025-11-18 19:23:45 -08:00
  • 6c2e5fcd91 [feat][Ascend][Mindspore]: support model-impl of mindspore (#9234) Chen Haozhe 2025-11-19 04:17:47 +03:00
  • 0d2d687812 [router][grpc] Support num_reasoning_tokens in haromy models (#13047) Chang Su 2025-11-18 16:40:32 -08:00
  • 9f59194f29 [Fix] Fix DeepSeek V3 MTP on B200 (#13548) Baizhou Zhang 2025-11-18 16:30:26 -08:00
  • cf1f0166b6 Add Qwen/Qwen1.5-MoE-A2.7B to model list (#13543) Kangyan-Zhou 2025-11-18 14:43:38 -08:00
  • 10969ae4be [chore] Disable ccache for sgl-kernel release (#13541) Baizhou Zhang 2025-11-18 14:28:58 -08:00
  • 92ad2ff9ce Flashinfer TRTLLM-GEN-MoE + Qwen3 (#13489) b8zhong 2025-11-18 14:18:29 -08:00
  • 9b64f6f359 [fix] Fixes accuracy issues caused by incorrect use of rope (#13495) Baidu-AIAK 2025-11-19 06:17:33 +08:00
  • c0d1a3383b Remove jet-ai/Jet-Nemotron-2B in nightly text tests as this is constantly failing (#13540) Kangyan-Zhou 2025-11-18 13:58:16 -08:00
  • a9d22b751d fix lora test (#13537) gongwei-130 2025-11-18 12:40:35 -08:00
  • 6b9459e86c [docker] fix dockerfile naming for diffusion (#13534) Simo Lin 2025-11-18 12:18:32 -08:00
  • 3a6ec47b0b [model-gateway] limit opened files in docker build to fix edge case (#13536) Chang Su 2025-11-18 12:17:38 -08:00
  • b8e32e7965 [model-gateway] fix gateway docker build due to recent py code change (#13532) Chang Su 2025-11-18 11:24:25 -08:00
  • f5566acc2a [CI] fix CI skipped on main (#13527) Liangsheng Yin 2025-11-19 01:19:08 +08:00
  • d79e12941c Small cleanups related to LoRA weight loading (#13474) Glen Liu 2025-11-18 12:14:34 -05:00
  • 6e9b154981 [CI] fix skipping pr-gate on main (#13525) Liangsheng Yin 2025-11-19 00:57:22 +08:00
  • 109f27ba3a [CI] update pr-gate to be compatible with new slash triggering mananer. (#13522) Liangsheng Yin 2025-11-19 00:49:13 +08:00
  • c1a30aa765 Add /tag-and-rerun-ci (#13521) sglang-bot 2025-11-18 06:53:53 -08:00
  • 2e1dbdb258 Update docs (#13519) Lianmin Zheng 2025-11-18 06:24:58 -08:00
  • 6d025fd35b Trigger CI retry with edit (#13516) Lianmin Zheng 2025-11-18 05:42:31 -08:00
  • 518467be64 [Feature] Re:Enable hybrid mem saver (#12962) Junrong Lin 2025-11-18 21:29:15 +08:00
  • 63807079b9 Add docs on trigger ci (#13513) Lianmin Zheng 2025-11-18 05:23:05 -08:00
  • f6cfe9f197 Use slash command to trigger CI (#13512) Lianmin Zheng 2025-11-18 04:39:58 -08:00
  • 7bc99d4164 README.md -> FOLDER_README.md (#13510) Lianmin Zheng 2025-11-18 04:10:04 -08:00
  • e2d6746808 Add .github/CI_PERMISSIONS.json to define the CI permissions (#13509) Lianmin Zheng 2025-11-18 04:00:15 -08:00
  • 6beb6e991f fix: create git tags directly instead of temporary branches (#13168) alisonshao 2025-11-18 03:23:27 -08:00
  • 7e88b9c128 [BugFix] Accuracy and function Issue when run ptpc quant model (#13157) Morpheus Guo 2025-11-18 17:54:58 +08:00
  • cfcf2758aa [HiCache] Critical fix to host memory double free (#13501) Zhiqiang Xie 2025-11-18 01:34:23 -08:00
  • ac81db66c2 [VLM][feat] Support encoder DP for Qwen2.5-VL (#13126) Nicholas 2025-11-18 16:13:18 +08:00
  • 33905005ee [CI] fix lint yml (syntax error) (#13496) Liangsheng Yin 2025-11-18 15:53:33 +08:00
  • 820e13c9c1 [opt kimi k2 3/n] opt kimi_k2 moe_fused_gate kernel (#13374) Xiaoyu Zhang 2025-11-18 15:36:31 +08:00
  • 595adf6d93 [CI] Add input for pr-gate (#13491) Liangsheng Yin 2025-11-18 15:35:49 +08:00
  • a5ad0069b2 fix: change performance log directory to cache path (#13482) Cheng Wan 2025-11-17 23:18:43 -08:00
  • 4e41edcb9c [CI] remove auto-labeling run-ci label. (#13486) Liangsheng Yin 2025-11-18 14:59:46 +08:00
  • 67071f55a8 [CI] fix triggered by a non-run-ci label (#13393) Liangsheng Yin 2025-11-18 14:24:58 +08:00
  • 0c96677902 CI: fix NFS EBUSY error in PR test workflow (#13460) alisonshao 2025-11-17 21:25:29 -08:00
  • f33860777c [Piecewise CUDA Graph] Support ModelOpt FP8 (#13094) b8zhong 2025-11-17 20:46:24 -08:00
  • 4ce8fb3cc2 Fix lora test (#13479) Liangsheng Yin 2025-11-18 12:21:57 +08:00
  • 4c3573e493 [PD] Clarify init method docstrings for kvsender and kvreceiver (#13476) Shangming Cai 2025-11-18 12:20:19 +08:00
  • f1be8aa0f2 chore: add an unified server arg for multimodal inputs preprocess config(#12149) wingedge 2025-11-18 12:18:50 +08:00
  • aa8ecbda7a model: support JetVLM (#13289) Zijian Zhang 2025-11-18 12:02:03 +08:00
  • 9188feccca [model-gateway] use worker startup time out for worker registration (#13473) Simo Lin 2025-11-17 19:28:21 -08:00
  • 26ca07469b [GLM4.6v] Relax the constraint of non-user role chat completion message schema for new GLM-v release (#13258) Binyao Jiang 2025-11-17 18:58:17 -08:00
  • 90c18a16cb [GLM4.6v] Required changes for bumping up to transformer 5.x (#13229) Binyao Jiang 2025-11-17 18:58:00 -08:00
  • d7984f3125 Adding CI Monitor Improvements (#13462) Douglas Yang 2025-11-17 18:42:35 -08:00
  • 7119d188f6 [CI] re-enable test_vision_openai_server_a ci (#13444) Yuhao Yang 2025-11-18 10:32:59 +08:00
  • fe3bbfb40c [AMD CI] Update sgl-router python path in dockerfile. (#13458) Sai Enduri 2025-11-17 18:13:24 -08:00
  • 9846f8ed72 fix MambaPool clear method after refactoring (#13449) Minglei Zhu 2025-11-17 18:03:45 -08:00
  • 85ae508e8b Add bfloat16 tuned fused moe config for Dpsk-MTP layer on B200 (#13455) Baizhou Zhang 2025-11-17 17:44:31 -08:00