Commit Graph

  • ae004e15c9 fix: ensuring nightly whls are tagged with latest commit (#18204) Douglas Yang 2026-02-03 15:54:41 -08:00
  • 793bf9fc06 Update weight rename check for Qwen3 Embeddings (#17535) satyamk7054 2026-02-03 13:55:11 -08:00
  • e867040fc6 add streaming parallel tool call test case (#18097) Hudson Xing 2026-02-03 12:46:01 -08:00
  • 7de650c83c [diffusion] hardware: support diffusion models on MTGPU (doc, 6/N) (#17346) R0CKSTAR 2026-02-04 04:44:57 +08:00
  • ec2461bc16 [diffusion] hardware: support diffusion models on MTGPU (multi-GPU, 5/N) (#17318) R0CKSTAR 2026-02-04 04:44:22 +08:00
  • acf724b036 [Diffusion] Only import sgl_kernel in custom op cuda path (SiluAndMul and RMSNorm) (#15592) R0CKSTAR 2026-02-04 04:42:58 +08:00
  • e166ca8758 [HiCache] feat: Add detailed cache hit breakdown for HiCache in sglext and Prometheus metrics (#17648) Vladislav Nosivskoy 2026-02-03 22:45:35 +03:00
  • d48bbe3bed [CI][NPU] Bugfix import sgl-kernel error (#18173) Even Zhou 2026-02-04 03:39:38 +08:00
  • 495290aefd enable ut test for xpu devices (#11712) DiweiSun 2026-02-04 03:15:14 +08:00
  • 0a6925639b ci: improve docker for cu13 builds (#18194) ishandhanani 2026-02-03 13:09:38 -06:00
  • 0db6fd4dbe Revert broken sgl_kernel exclusion patterns in paths-filter (#18193) Kangyan-Zhou 2026-02-03 10:56:44 -08:00
  • 820df545f2 fix: add cu13 dev container to our release (#18192) ishandhanani 2026-02-03 12:42:05 -06:00
  • 99fab2ce67 [Bugfix] Fix Mistral Large 3 NVFP4 TRTLLM MoE (#18065) elvischenv 2026-02-03 20:32:49 +08:00
  • a45647bce1 [PD] feat: support mooncake intra-node nvlink kv transfer (#17866) Lewis 2026-02-03 17:47:52 +08:00
  • cc69ac9e7a Warmup before profiling prefill latency for dynamic chunk sizing (#17198) Xiaowei Wang 2026-02-03 17:45:23 +08:00
  • 8e933e1914 AMD PD/D PR ci (#17183) Zhaoyi Li 2026-02-03 01:29:14 -06:00
  • 25508d11c0 [Docker] Remove hardcoded America/Los_Angeles timezone, default to UTC (#18121) Mohammad Miadh Angkad 2026-02-03 15:22:15 +08:00
  • 6f6b9c6e42 [Perf] Use safetensors load_file in multithread loader (#18124) Mohammad Miadh Angkad 2026-02-03 15:21:13 +08:00
  • 7a9d9c79d1 [HiCache] fix: apply extra_backend_tag in Mooncake batch_exists (#17265) fatSheep 2026-02-03 14:54:56 +08:00
  • 74f716dbd7 Gigachat 3 tool parser and tests (#14765) Viacheslav 2026-02-03 09:28:34 +03:00
  • 4181290efd [NVIDIA] Add --top-k argument to run_eval.py (#18025) Kaixi Hou 2026-02-02 22:17:53 -08:00
  • fe57a887b1 [TestFix] use unit tests for LoRA overlap loading tests (#18140) Glen Liu 2026-02-03 01:06:50 -05:00
  • f032c4f3d6 Support Markdown/Notebook-Friendly Documentation Export for Downstream Integration (#18131) Kun Lin 2026-02-03 00:43:20 -05:00
  • 78bf13db44 MoE Refactor: Refactor modelopt_quant.py -> flashinfer_trllm.py (#16685) b8zhong 2026-02-02 23:45:14 -05:00
  • eedd472025 [Diffusion] fix serving image_edit get input image bug (#18109) Xiaoyu Zhang 2026-02-03 12:17:16 +08:00
  • e484c90cc7 Add triton_fused_moe config for GLM-4.7-FP8 tp8 H20 H20-3e (#18091) Hank Han 2026-02-03 12:08:23 +08:00
  • 9b1619c148 [Move sgl-kernel Kernel to JIT] Add JIT concat MLA kernels (#17889) Linyu Wu 2026-02-03 10:49:17 +08:00
  • 62004fd2be [diffusion] UX: improve logging (#18122) Mick 2026-02-03 10:35:05 +08:00
  • 180594358b [HiCache]: Support DeepSeek v32 cpu offloading (#17415) zhangheng 2026-02-03 10:07:37 +08:00
  • a1bbc892af [Diffsuion & JIT_kernel] QKNorm cross heads kernel (#18073) Xiaoyu Zhang 2026-02-03 10:03:17 +08:00
  • fd983b09b6 [Performance] Optimize radix cache eviction performance (#14339) EkiRui 2026-02-03 09:44:20 +08:00
  • c8da307d7e feature: adding gpt-oss 120b nightly test (#18134) Douglas Yang 2026-02-02 17:11:28 -08:00
  • 28e2340725 Fix HF hub race condition in CI by coordinating model downloads across TP ranks (#17787) Alison Shao 2026-02-02 14:57:45 -08:00
  • 812fd47cb4 Re-enable test_mla_int8_deepseek_v3.py after HF token fix (#18123) Alison Shao 2026-02-02 14:38:42 -08:00
  • 027f314050 [Fix] data race in req_to_token pool (#17850) cctry 2026-02-02 14:38:15 -08:00
  • cbf1500390 [MiMoV2Flash] [feat]: support two batch overlap (#17634) TZHelloWorld 2026-02-03 06:03:57 +08:00
  • b0a6d5244c [NPU] support dsv32 radixcache on ascend (#17964) khalilzhk 2026-02-03 03:34:12 +08:00
  • 677f3c49da [DeepSeek V3.2] [Bugfix] slice indexer and padding fa3 when can not run cuda graph (#17076) Yongfei Xu 2026-02-03 01:32:20 +08:00
  • 980d2936cd model: support Step-3.5-Flash (#18084) Yuhao Yang 2026-02-03 00:40:07 +08:00
  • c781db0f6c [NPU] update nightly tests (#17952) Sugar920 2026-02-03 00:13:30 +08:00
  • c971852ffc docs: move deepseek_ocr to popular model usage and add cookbook reference (#18120) sglang-bot 2026-02-02 05:45:41 -08:00
  • 5636d16dec [CI] Add logs to debug TestOpenAIServer.test_completion_stream (#17471) Byron Hsu 2026-02-02 02:38:08 -08:00
  • 0537232b05 [NPU]mindspore model support moe (#15363) Xuhao Zhang 2026-02-02 17:52:49 +08:00
  • aa780a6258 [diffusion] fix: remove accelerate dependency for device mapping (#18026) CHEN Xi 2026-02-02 17:24:19 +08:00
  • e3021b65fe support smem in per_token_quant_fp8 kernel (#16725) zhangxin81 2026-02-02 17:18:50 +08:00
  • a0757c9624 [Diffusion] Fix Ring Parallel bug with FA4 (#18062) Xiaoyu Zhang 2026-02-02 17:06:51 +08:00
  • 750ad0d290 [AMD] enable MoRI to release and nightly builds (#18101) HAI 2026-02-02 00:28:30 -08:00
  • 86117dfe0e [diffusion] CI: deprecate WarmupRunner in CI (#18038) 陈一涵 2026-02-02 15:16:08 +08:00
  • 522e13b4d2 fix: correct weight loading prefix mapping for Qwen3-VL (#18024) Xiaoming Liu 2026-02-02 13:50:21 +08:00
  • a480ca7ead fix: zmq_to_tokenizer encoder transfer when host listens to 0.0.0.0 (#17929) RangerCD 2026-02-02 13:27:27 +08:00
  • cd31540fd7 Improve Per Commit Test job filtering for sglang-kernel (#18054) Kangyan-Zhou 2026-02-01 21:15:23 -08:00
  • f1824a957b [EPD][refactor]: introduce BaseMMReceiver for gRPC transport integration (#17921) siyu 2026-02-02 11:37:32 +08:00
  • ab8b99eb23 Refine logprob logic for request handling (#17986) Cheng Wan 2026-02-01 19:11:52 -08:00
  • ea04bc1dd6 [AMD] Fix aiter version in rocm image (#18076) YC Tseng 2026-02-02 11:00:38 +08:00
  • 8ed35df204 Add bootstrap_room validation to detect metadata corruption in PD disaggregation (#17430) Simon (Jiyou) Li 2026-02-02 10:43:50 +08:00
  • c84cd4b5ff [diffusion] fix: fix missing component names for VAELoader (#18069) Mick 2026-02-02 09:48:17 +08:00
  • 977096ae03 [diffusion] cli: introduce generic attention backend configuration in ServerArgs (#18036) Mick 2026-02-02 09:47:40 +08:00
  • d11ccc0a0a fix: avoid double reduce in VLM dp attention (#17991) Yuhao Yang 2026-02-02 09:44:32 +08:00
  • 9227d4f748 [Fix] Remove no use code in MiMo-V2-Flash (#18051) Yuan Luo 2026-02-02 07:34:09 +08:00
  • 99dad105fd [TestFix] rewrite LoRA overlap loading tests (#18047) Glen Liu 2026-02-01 17:52:08 -05:00
  • 993ec178ef [BUGFIX]: using language-only should not reserve space for the vision encoder (#18011) Koushik Dutta 2026-02-01 14:49:00 -08:00
  • 1acae30806 [Auto Sync] Update test_deterministic.py (20260131) (#18034) Lianmin Zheng 2026-02-01 07:40:07 -08:00
  • fb609669ca [Auto Sync] Update elementwise.py (20260131) (#18033) Lianmin Zheng 2026-02-01 07:39:37 -08:00
  • afebb7ab78 Optimize custom-all-reduce (#17674) Yuan Luo 2026-02-01 18:59:31 +08:00
  • 4ea4f2a20c [VLM] Optimize get_rope_index for GLM4v (#17420) Yuan Luo 2026-02-01 18:59:15 +08:00
  • 0fe282543f [NPU] support the Enable return routed experts (#17025) jiashaokun-1 2026-02-01 18:31:39 +08:00
  • 27bec34203 [NPU] disaggregation_decode_enable_fake_auto parameter adaptation (#17811) Estrella-xx 2026-02-01 18:14:26 +08:00
  • d9050b4a9c Reset evict swa status when retract (#18059) Ke Bao 2026-02-01 17:17:37 +08:00
  • 56907cbcb1 Move deleted 8-GPU tests to test/manual/ (#18060) Alison Shao 2026-02-01 00:21:56 -08:00
  • 47592a23c7 [CI] Fix AMD CI by inlining dummy_grok config (#18044) sunxxuns 2026-02-01 00:20:57 -08:00
  • 855dd0546c feat: Add Ling Flash v2.0 support for Eagle3 (#15119) yefei12 2026-02-01 15:57:45 +08:00
  • 3ca29dffc7 support qwen3-next eagle3 (#14607) lukec 2026-02-01 15:45:23 +08:00
  • 9bb1260558 [Feature] Support file:// URL format for multimodal inputs (#14490) Praneth Paruchuri 2026-02-01 13:14:44 +05:30
  • 71babdef51 Fix CUDA 12 dependency when importing Mooncake in official CUDA 13.x image (#17540) ZhenshengWu 2026-02-01 15:41:21 +08:00
  • 486c7de39f Optimizing all_reduce in RMSNormTP in minimax_m2 (#16483) Roger Young 2026-02-01 15:39:37 +08:00
  • 9c168fcac7 Fix Diffusion Request Validation to allow missing input artifacts if the input only contains text (#16610) Kangyan-Zhou 2026-01-31 23:38:40 -08:00
  • 4d28cda007 [model] Support MiniCPM-V 4.5 (#9610) tc-mb 2026-02-01 15:37:36 +08:00
  • 2c036f1eb1 [Bugfix] fix the display error (inconsistent context) (#17699) linhaifeng 2026-02-01 15:35:11 +08:00
  • e5ac6229e1 Fix installation script for H200 runners (#18050) Kangyan-Zhou 2026-01-31 23:30:51 -08:00
  • a0bae4c343 Migrate 4-GPU/8-GPU workflow jobs to stage-c and add CI registry decorators (#17299) Alison Shao 2026-01-31 22:37:22 -08:00
  • 95180484e9 Disable test_mla_int8_deepseek_v3.py temporarily (#18057) Alison Shao 2026-01-31 22:33:43 -08:00
  • d396650bd2 Fix swa kv cache memory allocation (#18039) Ke Bao 2026-02-01 14:26:51 +08:00
  • 38d275a9fd [BugFix] fix gpt-oss accuracy issue when enabling piecewise cuda graph (#18013) Minglei Zhu 2026-01-31 22:26:26 -08:00
  • 11892599f1 [metric] Optional extra metric labels (#18049) Yinghai Lu 2026-01-31 22:25:59 -08:00
  • c7d53fa26a Set torch url index in pyproject.toml (#16802) Baizhou Zhang 2026-02-01 13:23:52 +08:00
  • 429ef988bc [BugFix] Fix draft model specified config file (#17815) khalilzhk 2026-02-01 12:45:36 +08:00
  • e884b17632 Fix rerun stage command with merged commit history (#17960) Kangyan-Zhou 2026-01-31 20:37:55 -08:00
  • 2b2515423a Skipped warning on sm100 (#18000) Kaixi 2026-02-01 05:21:03 +01:00
  • d443d2d2ae Improve error output in tnightly tets (#18053) Kangyan-Zhou 2026-01-31 19:26:19 -08:00
  • 0f2df9370a feat: validate ib devices in server args (#17598) Yingchun Lai 2026-02-01 09:50:42 +08:00
  • 398d13a189 [Perf] Add Flashinfer DeepGEMM SM90 for SwapAB Optimization (#15514) b8zhong 2026-01-31 19:56:23 -05:00
  • 9951a1ae07 Fix: Remove duplicate assignment for use_w4afp8 (#17858) Chongchong Tian 2026-02-01 08:04:08 +08:00
  • c2ab3713e9 [Performance] Optimize Mllama LayerNorm -> Upd (#9725) Yi Zhong 2026-01-31 19:02:57 -05:00
  • 6aaea09b3d Update python/sglang/README.md (#18045) Hao Jin 2026-01-31 13:20:41 -08:00
  • 97593c9f41 [CPU] toml file update (#17861) Zaili Wang 2026-02-01 05:16:06 +08:00
  • 9ac4dcada4 [Tiny] Fix grammar in shared experts fusion log messages (#18043) Mohammad Miadh Angkad 2026-02-01 05:14:25 +08:00
  • 46095f0551 [MUSA] Update 3rd party dir to build/_deps (#18035) R0CKSTAR 2026-02-01 04:02:39 +08:00
  • ef134d407d [Fix] Revert back to using CUTLASS mm_fp4 backend (#17369) b8zhong 2026-01-31 10:01:29 -05:00
  • 1a006c2a0d [diffusion] refactor: split component_loader into component-wise files (#17820) Mick 2026-01-31 20:22:31 +08:00
  • 7412ceb4eb [Auto Sync] Update linear.py to assert shapes (20260130) (#17966) Lianmin Zheng 2026-01-31 01:01:55 -08:00