Commit Graph
9775 Commits
Author SHA1 Message Date
Makcum888e 2aa0db7d9c [Diffusion] [NPU] Fix CI run (#18921) 2026-02-17 16:54:19 +03:00
HAI 8bb1037796 ROCm use rotary_embedding from sgl-kernel (#18920) 2026-02-17 03:00:37 -08:00
Makcum888e 5f81ec1ad5 [Diffusion] Fix get model name when model local path end with "/" (#18918) 2026-02-17 13:19:54 +03:00
Ratish Pandronnie_zheng f6cc02489f [diffusion]: fix sparse video gen 2 backend being applied to cross-attention (#18900)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-02-17 13:17:46 +03:00
HAI b158f5d4a2 Revert "[AMD] Fix RotaryEmbedding crash on AMD/ROCm (regression from #17934)" (#18922) 2026-02-17 01:07:50 -08:00
Makcum888eandgemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com> 14c95d255c [Diffusion] [NPU] [Doc] Add NPU documentation for sglang-diffusion (#18894)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-02-17 10:12:20 +03:00
billishyahao 899e2be7d0 [TBO] fix cuda graph intermittently becomes disabled bug (#18320) 2026-02-16 22:18:57 -08:00
Michaelandmichaelzhang-ai 5e3103a787 [AMD] Fix RotaryEmbedding crash on AMD/ROCm (regression from #17934) (#18903)
Co-authored-by: michaelzhang-ai <michaelzhang-ai@users.noreply.github.com>
2026-02-17 12:59:40 +08:00
Mohammad Miadh Angkad 90a0d66e1e [Tiny] Fix assert syntax warning in compressed_tensors_w4a4_mxint4_moe.py (#18899) 2026-02-17 12:54:30 +08:00
Alison Shao 7e41ac6c8d Skip flaky test_tool_choice_required_non_streaming for Mistral (#18889) 2026-02-17 12:50:55 +08:00
Yilong Zhao d5307ce022 [misc] adding metadata field in UpdateWeightFromDiskReqInput (#18821) 2026-02-17 12:14:15 +08:00
triple-muandgemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com> 26b2c63d03 [diffusion] operator: unify rotary embedding impl (#18164)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-02-17 12:02:48 +08:00
pansicheng b21390f8f3 Adapt the Qwen2Model._update_causal_mask for transformers==4.57.1 (#18774) 2026-02-17 10:20:41 +08:00
Ratish PandXiaoyu Zhang 50ca24aebb [diffusion]: fix scheduler crash on ZMQ messages with unexpected frame counts (#17890)
Co-authored-by: Xiaoyu Zhang <35585791+BBuf@users.noreply.github.com>
2026-02-17 09:45:05 +08:00
Alison ShaoandLiangsheng Yin f9c3def7fe Fix CI: add flashinfer --download-cubin to install dependencies (#18887)
Co-authored-by: Liangsheng Yin <lsyincs@gmail.com>
2026-02-16 13:50:10 -08:00
1b659bcb08 Fix GLM-5 fused shared expert (#18804)
Co-authored-by: FrankMinions <liuchen@shinemo.com>
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
Co-authored-by: Kangyan-Zhou <zky314343421@gmail.com>
2026-02-16 19:50:39 +00:00
danielafrimi 0ff24159a5 Fix modelopt FP8 create weights (#18447)
Signed-off-by: root <dafrimi@nvidia.com>
2026-02-17 00:59:50 +08:00
Tamir Baydasovgemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>ronnie_zhengPeng Zhang
eba6af385d [2/N] Quantization Refactor: Compressed tensors MoE schemes (#17503)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: ronnie_zheng <zl19940307@163.com>
Co-authored-by: Peng Zhang <aniz1905@gmail.com>
2026-02-16 18:03:51 +03:00
Estrella-xx 1b3513a7e4 refactor FAKE transfer backend and remove --disaggregation-decode-enable-fake-auto parameter (#18345) 2026-02-16 17:27:02 +03:00
Ratish P c1d1337afc [diffusion][Wan]: fix sparse attention backends being applied to cross-attention (#17596) 2026-02-16 21:57:58 +08:00
Mohammad Miadh Angkad b86c6491fa [Perf] ~9.5x faster Blackwell MXFP4 MoE weight loading (#18858) 2026-02-16 19:47:09 +08:00
Shivam jindalandyes-its-shivam 4f0409f8aa [Model] Add Qwen3ForRewardModel and fix Qwen3ForSequenceClassification (#17992)
Co-authored-by: yes-its-shivam <yes-its-shivam@users.noreply.github.com>
2026-02-16 19:44:41 +08:00
Mick de833f9e8e Revert "[diffusion]: Improve layerwise offload buffer reuse and shared-storage handling" (#18866) 2026-02-16 18:00:58 +08:00
Mick d0c94e136a [diffusion] logging: improve peak vram logging (#18865) 2026-02-16 16:44:37 +08:00
Yi Zhong ed22720c07 [JIT kernel] hd=512,1024 in JIT QK norm (cta based) (#17515)
Signed-off-by: vincentzed <207368749+vincentzed@users.noreply.github.com>
2026-02-16 16:07:24 +08:00
Alison Shao 206accd15d Fix GLM-4V processor registration when glm_ocr is unavailable (#18885) 2026-02-16 16:02:31 +08:00
Changyi Yang 61da34ad0b [diffusion] fix: fix LoRA weight snapshot aliasing in unmerge logic (#18883) 2026-02-16 15:39:45 +08:00
RAY d85884ca57 Update ascend_npu_qwen3_5_examples.md (#18888) 2026-02-16 10:24:27 +03:00
Alison Shao 86c181e335 Fix test_lora_qwen3 nightly failure: replace adapter with added_tokens (#18884) 2026-02-16 14:35:06 +08:00
Douglas Yang f1efb46bdd fix: adding performance logging for nightly diffusion (#18023) 2026-02-16 14:09:00 +08:00
Douglas Yang 2050875424 fix: unifying docker image build pipeline (#18814) 2026-02-16 14:05:55 +08:00
fzyzcjyandYueming Yuan f554b3c27b Support dumping gradients, parameters, lazy values (#18881)
Co-authored-by: Yueming Yuan <112649537+yueming-yuan@users.noreply.github.com>
2026-02-16 13:34:06 +08:00
fzyzcjy 9a7d8d5eb0 Collect upper level metadata to dump output (#18880) 2026-02-16 13:31:19 +08:00
fzyzcjy 949792d0c6 Change dump output format to dict with value and metadata (#18879) 2026-02-16 13:30:47 +08:00
fzyzcjy 02816abc0d Flip dumper to disable by default and refactor environment handling (#18878) 2026-02-16 13:29:32 +08:00
Duyi-WangandHAI 5ddc84e33e [AMD] MORI-EP inter kernel type switch (#18437)
Co-authored-by: HAI <hixiao@gmail.com>
2026-02-15 20:59:39 -08:00
Johnsonmsandgemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com> bc79a64d3a [Diff]: support SGLANG_TORCH_PROFILER_DIR environment variable for profiler log directory (#18454)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-02-16 12:47:29 +08:00
Mick 0af9dcc407 [diffusion] refactor: refactor server_args adjust and validate logics (#18863) 2026-02-16 11:49:06 +08:00
Mick 78b4c9e248 [diffusion] fix: avoid saving output for warmup requests (#18867) 2026-02-16 11:48:28 +08:00
Yuan Luoandluoyuan.luo 8a82c70297 [VLM] Optimize Ernie4.5-VL rotary embedding with fused triton kernel (#18856)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
2026-02-16 11:19:44 +08:00
Douglas Yang 45715af50c fix: nightly whl dev date suffix (#18873) 2026-02-16 10:57:37 +08:00
Rain Jiang 0ffd0a3995 Nsa trtllm mla sparse fp8 support with Deepseek v3.2 NVFP4 (#18389) 2026-02-16 09:29:54 +08:00
Mohammad Miadh Angkad 8290171f52 [CI] Remove --mem-fraction-static 0.93 from gpt-oss test (#18869) 2026-02-16 09:24:11 +08:00
Xiaoyu Zhang d3bae71e3f Add claude skills for sgl-kernel and jit-kernel (#18855) 2026-02-15 15:14:03 -08:00
chenxu214 fd5a45d5cf Update ascend_npu_support.rst (#18868) 2026-02-16 01:41:38 +08:00
chenxu214 f2d72866e9 Create ascend_npu_qwen3_5_examples.md (#18864) 2026-02-16 01:15:20 +08:00
blake-sncandClaude Opus 4.6 0d30896015 fix(sgl-kernel): use >= 120 for SM12x CUDA kernel dispatch (#18750)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-16 00:44:47 +08:00
blake-sncandClaude Opus 4.6 5fc328465a fix(sgl-kernel): support CUDA 13 runtime preloading for DGX Spark (#18747)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-16 00:43:04 +08:00
b79808bee2 Fix libnuma.so does not exsit (#15355)
Signed-off-by: Michael Qiu <qiudayu.qdy@antgroup.com>
Co-authored-by: Mike_Qiu <qiudayu.qdy@antgroup.com>
Co-authored-by: fzyzcjy <5236035+fzyzcjy@users.noreply.github.com>
Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu>
2026-02-16 00:37:50 +08:00
akhilg-nv 48eac1b62d Improve profiler options for bench_serving (#16991) 2026-02-16 00:36:01 +08:00