Commit Graph

  • 0e184609d3 Add launch_command assignment in crash dump (#17967) Lianmin Zheng 2026-01-31 01:00:40 -08:00
  • 22498e10c0 [Fix] Triton TP MoE Dpsk V3/Qwen3 Coder with SwapAB (#17965) b8zhong 2026-01-31 02:56:26 -05:00
  • a4df95c15f [EPD][Perf] parallelize ZMQ send for encode server (#16487) Zheng Wengang 2026-01-31 14:30:11 +08:00
  • 04efd03dbf Fix OOM in DeepSeek weight loading by deferring dict(weights) materialization (#17744) jeff 2026-01-31 13:59:00 +08:00
  • 45fe51a28e Reduce topk kernel shared memory from 128KB to 32KB for better occupancy (#17747) Yifan Cui 2026-01-31 13:42:21 +08:00
  • c72bf50706 add reasoning_tokens usage test for tool call (#18022) Hudson Xing 2026-01-31 14:09:23 +09:00
  • c52578c7fd 【docs】【NPU】Update Expert Parallelism docs for Ascend NPU (#17940) husf 2026-01-31 12:42:11 +08:00
  • dc77defdd0 Fix .gitignore may ignore files like core_attention.py (#18021) R0CKSTAR 2026-01-31 12:33:32 +08:00
  • d0d9cecd1b Fix cuBLAS >=12.9 detection for cu12/cu13 package naming (#17766) Mohammad Miadh Angkad 2026-01-31 12:01:52 +08:00
  • 22aad4e2c4 [Diffusion] Fix FLUX.1-schnell time embedding argument mismatch (#17988) Xiaoyu Zhang 2026-01-31 11:47:27 +08:00
  • ec76c390c9 Add ROCm + Mori docker build instructions in rocm.Dockerfile (#18018) kk 2026-01-31 11:12:07 +08:00
  • ee3058c6e8 [NPU] fix sgl-kernel-npu package url error in npu.Dockerfile (#18017) 22dimensions 2026-01-31 11:06:58 +08:00
  • f6a4ff718f doc update for CANN version (#18014) Tiance Wang 2026-01-31 10:10:23 +08:00
  • 5d00150e99 [sglang] fix mm token padded value overlap with text token id (#17781) Bi Xue 2026-01-30 17:09:13 -08:00
  • e86476acfc [NPU] support llama-3.2-11B-vision-instruct mode for NPU (#17492) JiaruiChang5268 2026-01-31 08:49:38 +08:00
  • 578b119bc6 [BugFix] Fix server crashes when req.grammar and ngram spec are enabled (#17585) Siyuan Chen 2026-01-31 03:57:42 +08:00
  • 81449b4bee Optimize GDN decode for Qwen3 Next (#17094) Sam (Kesen Li) 2026-01-31 01:02:12 +08:00
  • abf13ccc11 [Diffusion] Fix lora default lora_scale bug (#17982) Xiaoyu Zhang 2026-01-30 22:04:54 +08:00
  • 0c5a81acb8 [BUGFIX] Fix dp size > 1 for qwen3 vl model (#17624) Zheng Li 2026-01-30 20:44:25 +08:00
  • c04efe030a [Model] Add K-EXAONE model support (#16294) Changhun Lee 2026-01-30 21:01:14 +09:00
  • 9c2a468e2c update ascend docs (#17987) amote-i 2026-01-30 18:03:55 +08:00
  • 8ce9609fa2 fix: fix SHM pointer re-serialization in DP attention (#17930) Fan Yin 2026-01-30 17:03:30 +08:00
  • 77a27e728c Add cuda graph status to prefill log (#17836) Ke Bao 2026-01-30 16:56:53 +08:00
  • c8dc543dc5 SGLang Tracing: Improve root span attributes (#17008) Haotong Zhang 2026-01-30 16:02:05 +08:00
  • a16aba8340 Increase install dependency for gb200 (#17977) Kangyan-Zhou 2026-01-29 23:55:25 -08:00
  • 96584ab692 adapt MODELSCOPE download (#17922) jianzhao-xu 2026-01-30 15:26:54 +08:00
  • 70db3398d1 [NPU] enhance accuracy for model kimi-vl-a3b-instruct (#17480) McZyWu 2026-01-30 15:19:42 +08:00
  • 33c053c50c fix(benchmark): add missing args for speculative decoding benchmark (#17974) cswuyg 2026-01-30 15:05:42 +08:00
  • c35aa0238c [CPU][INT4] Add INT4 kernels for CPU (#8226) jianan-gu 2026-01-30 14:30:13 +08:00
  • 88f7759402 [CPU] optimize flash_attn_varlen_func (#15708) Ma Mingfei 2026-01-30 14:07:05 +08:00
  • 336dc4579e [CPU] Optimize Qwen3-next model on CPU (#12525) jianan-gu 2026-01-30 14:03:58 +08:00
  • 71e4d3b6bc [Intel GPU] fix import error to run DeepSeek-V2-Lite model with BF16 on XPU (#10858) Polisetty V R K Jyothendra Varma 2026-01-30 11:23:53 +05:30
  • 7541da15d2 Fix prefill latency performance drop of bench serving (#14592) gaopengff 2026-01-30 13:28:17 +08:00
  • 858dc80aff [Intel GPU] fix device in DeepseekScalingRotaryEmbedding to run DeepSeek-V2-Lite BF16 on XPU (#10021) Polisetty V R K Jyothendra Varma 2026-01-30 10:51:38 +05:30
  • 0e4d9ddbd6 Fix the scenario where eh_proj is quantized in the bailing moe nextn weights (#17808) LHXuuu 2026-01-30 13:11:47 +08:00
  • 606ff09ef8 [Fix] Remove unused Type import in gpt_j.py (#17975) Kangyan-Zhou 2026-01-29 21:11:11 -08:00
  • 632c7afa8c [Fix] add block size logic for sm120 smem size (#14311) Koushik Dutta 2026-01-29 21:01:57 -08:00
  • 046b29be16 GPTJForCausalLM Support (#7839) Wenchen Lo 2026-01-29 21:00:04 -08:00
  • 22df62d586 add weightless qk norm to RMSNorm interface for Llama 4 (#12813) b8zhong 2026-01-29 22:09:55 -05:00
  • 84ab611af8 model: support DeepSeek-OCR-2 (#17897) baonudesifeizhai 2026-01-29 20:49:51 -05:00
  • 2cd2c3118d Add concurrency tracking to runner utilization report (#17963) Kangyan-Zhou 2026-01-29 17:31:55 -08:00
  • b7b1ba329e [AMD] fix pip sglang version (#17950) YC Tseng 2026-01-30 09:19:57 +08:00
  • 2b3408ff14 feat: add forward timeout (#17831) StonyPort 2026-01-30 08:52:29 +08:00
  • 6a6b36367e Fix logprob_start_len handling for prefill-only requests (#17395) Cheng Wan 2026-01-29 15:14:43 -08:00
  • c3bf53c7c1 Fix ci weight validation logic to check the safetensor completeness (#17917) Kangyan-Zhou 2026-01-29 13:00:42 -08:00
  • 1f75c2af4d Fix /tag-and-rerun-ci to do full rerun when PR has sgl-kernel changes (#17729) Alison Shao 2026-01-29 12:54:30 -08:00
  • a416af4be7 Fix capture_sizes range for pcg (#17956) Cheng Wan 2026-01-29 12:46:35 -08:00
  • 1b6798a6a4 Fix torch.__version__ for PEP440 (#15682) EduardDurech 2026-01-29 20:55:13 +01:00
  • 4f2b73baf6 [MUSA] Add labeler config (#17923) R0CKSTAR 2026-01-30 01:58:54 +08:00
  • d417c6809e Add tool call tests for DeepSeek V3.2 in nightly CI (#17951) Hudson Xing 2026-01-30 01:50:54 +08:00
  • 88fb927cc9 [diffusion]: add dummy device attribute to fix AttributeError (#17949) Ratish P 2026-01-29 23:05:12 +05:30
  • 0769de9b0f Support LightOnOCR-2-1B (#17806) Shivam jindal 2026-01-29 20:33:41 +05:30
  • 3c9cc44ff5 Add mxfp8 support for online quantization, Triton dense linear, and CUTLASS MoE (#17449) Ziang Li 2026-01-29 05:33:57 -08:00
  • cfa09d311c [NPU] add support for npu x86_64 image release (#14127) 22dimensions 2026-01-29 21:23:56 +08:00
  • 3c2f4c7bbe [diffusion] model: sync with upstream z-Image (#17822) Yuhao Yang 2026-01-29 21:10:11 +08:00
  • 30adf78f82 [diffusion]: align sglang diffusion AMD pyproject_other.toml diffusion dependency with pyproject.toml (#16225) RoyWang 2026-01-29 17:50:57 +08:00
  • ef1c512754 Add aiter bias moe support in gpt-oss mxfp4 model (#17735) kk 2026-01-29 17:50:11 +08:00
  • 319f6886fe [diffusion] model: move tp_rmsnorm check to WanTransformerBlock (#17792) triple-mu 2026-01-29 17:39:00 +09:00
  • cdedbf1486 [diffusion] fix: resolve library mismatch in scheduler and update dit offload method name (#17916) Zhang Yiyang (SII) 2026-01-29 15:54:36 +08:00
  • 7b79326751 [NPU] support GPTQ quantization on npu (#15203) 22dimensions 2026-01-29 15:48:18 +08:00
  • cbf90d70ff [PD] Support KV transfer with MORI-IO (#14626) Niko Ma 2026-01-29 15:22:41 +08:00
  • d3cdee0a04 [MUSA][4/N] Add common device utilities, distributed backend, and custom op wiring (#17246) R0CKSTAR 2026-01-29 15:13:24 +08:00
  • 9409c43593 Fix flaky tool calls in the Kimi K2.5 model (#17914) Xinyuan Tong 2026-01-28 23:58:16 -05:00
  • 3c34d2c3eb [FIX] kimi_k2 reasoning parser (#17901) Xinyuan Tong 2026-01-28 22:47:09 -05:00
  • 1b22f2ee1c update ascend docs (#17741) amote-i 2026-01-29 11:43:48 +08:00
  • 0ff0d181ca feat: add custom request header logging (#17786) Joe Redmond 2026-01-28 19:33:08 -08:00
  • f1384f5293 Integration mori backend for EP a2a data communication (#17012) kk 2026-01-29 11:07:34 +08:00
  • 673dc09d9b [Fix][trtllm-mha] Canonicalize the strides when num_head = 1 (#17732) Jerry Ji 2026-01-28 18:11:18 -08:00
  • ac16d4450e [smg][mesh] extract mesh to mesh crate to reduce compile time (#17907) Simo Lin 2026-01-29 10:58:42 +09:00
  • 0368ddf9ea [JIT Kernel]Support fused_add_rmsnorm in JIT Kernel (#17677) Qi Yuhang 2026-01-29 09:29:59 +08:00
  • 09a9147f59 [diffusion] model: support MOVA (#17704) Zhang Yiyang (SII) 2026-01-29 09:12:08 +08:00
  • 3fcda00e8c [CI] Fix CI timeouts by upgrading runai_model_streamer (related to #16937) (#17636) Prozac614 2026-01-29 09:09:45 +08:00
  • d4180815a4 Make the functions in logits_processor.py and sampler.py more modular (#17885) Lianmin Zheng 2026-01-28 16:24:23 -08:00
  • 0998de088b [Perf] Tune Llama-4-Scout-17B-16E-Instruct fused moe kernel (#17891) jackey hua 2026-01-28 17:06:46 -05:00
  • e9d727cb92 [MUSA][7/N] Enhance CUDA / PyNccl wrapper to support MTLink connectivity detection (#17499) gingerXue 2026-01-29 03:36:30 +08:00
  • b77b0ffd60 [NPU] NZ for non-quantized MOE, Qwen3 MOE double memory consumption fix (#15904) Артем Савкин 2026-01-28 19:55:08 +03:00
  • 1953efb60e [AMD] ROCm: route W4A16 MoE to Triton and fix packed-weight loading (#17863) Jinn 2026-01-28 10:20:23 -06:00
  • 1d1e72e516 [diffusion] fix: fix comfyui import typo (#17834) triple-mu 2026-01-29 00:49:55 +09:00
  • f8636fbb25 [AMD] Add Kimi-K2, DeepSeek-V3.2 tests to nightly CI (#17523) Michael 2026-01-28 00:55:46 -08:00
  • 6077de1237 [CI] [NPU] npu ci use existing modelscope model (#17868) Even Zhou 2026-01-28 16:51:41 +08:00
  • c08b54a575 [JIT kernel] Update jit_kernel cache and develop doc (#17842) Xiaoyu Zhang 2026-01-28 15:09:47 +08:00
  • 2573a262af [diffusion] doc: fix wrong docker run command (#17856) Mick 2026-01-28 14:52:33 +08:00
  • 6f009961bb [model-gateway] Optimize HashRing construction to reduce heap allocations (#17575) Praneth Paruchuri 2026-01-28 12:16:40 +05:30
  • 897c35b457 [model-gateway] Optimize consistent hashing hot path to eliminate allocations (#17467) Praneth Paruchuri 2026-01-28 12:16:11 +05:30
  • c0b4dd68a2 Add a performance dashboard server and frontend for nightly CUDA tests (#17725) Kangyan-Zhou 2026-01-27 22:22:33 -08:00
  • fb74e43707 [Diffusion] Delete sgl-kernel outdated time_embedding kernel (#17278) Xiaoyu Zhang 2026-01-28 14:18:53 +08:00
  • 67fb492c9a [CI] Fix test_moe_fused_gate error (#17844) Xiaoyu Zhang 2026-01-28 12:03:17 +08:00
  • a8dda2aa57 [DSv32] Overlap indexer qk projection and activation quant (#17688) Ziang Li 2026-01-27 19:46:49 -08:00
  • 52bca42870 [AMD] CI - enable deepseekv3.2 on MI325-8gpu and merge perf/accuracy test suites into stage-b suites (#17633) YC Tseng 2026-01-28 10:54:36 +08:00
  • 1c4616a034 fix: add bias when enable mm fallback variant (#17690) Yisheng Gong 2026-01-27 17:50:49 -08:00
  • 647428d8d6 [diffusion] perf: apply mul add fusion for Qwen-Image (#16299) 陈一涵 2026-01-28 09:40:13 +08:00
  • 32ea7bcdd8 [diffusion] endpoint: fix vertex generate (#17611) Yashika Gandhi - Google 2026-01-27 17:38:56 -08:00
  • 88fcd8535f [diffusion] feat: add an arg for controlling the number of prefetched layers in layerwise-offload (#17693) Mick 2026-01-28 09:34:27 +08:00
  • 1507dc6cdf [diffusion] fix: fix suppressing error log on non-main ranks (#17712) Mick 2026-01-28 09:29:19 +08:00
  • 331a22427c [Diffusion] glm-image apply flashinfer rope (#17689) Xiaoyu Zhang 2026-01-28 08:51:37 +08:00
  • 93423ff780 [AMD] Deprecate ROCm 6.3 artifacts and standardize gfx942 on ROCm 7 (#17785) Hubert Lu 2026-01-27 15:58:49 -08:00
  • 8278ef0e68 Pass GPU ids to kill specified devices in script. (#17840) Liangsheng Yin 2026-01-27 13:52:55 -08:00
  • 4d00bd17a3 use shared memory for multimodal feature transport between Tokenizer and Scheduler (#16402) siyu 2026-01-28 03:01:08 +08:00
  • d90c0837e5 [hybrid-model] clean up and consolidate redundant fields in RadixLinearAttention (#17660) Minglei Zhu 2026-01-27 10:37:58 -08:00
  • 8acd4d7d7e Make flashMLA work on: Cu13, B300 (#17600) Yi Zhong 2026-01-27 11:12:47 -05:00