Cheng Wan
|
84c09913eb
|
Moving _alloc_extend_naive out of npu allocator (#18200)
|
2026-02-04 02:09:55 -08:00 |
|
zhangheng
|
be557cbc5f
|
[RadixTree][5/N Refactor]: Introduce pre and post-processing methods for key matching (#18147)
|
2026-02-04 17:10:46 +08:00 |
|
Baizhou Zhang
|
d279520ba5
|
[DeepGemm] Add a flag for fast warmup (#18111)
|
2026-02-04 14:12:13 +08:00 |
|
Jianying
|
4739f2e8d5
|
[diffusion] kernel: gated residual layernorm scale shift and layernorm scale shift kernel fusion for Qwen-Image, WAN and HunyuanVideo (#14717)
Co-authored-by: AichenF <aichenf@nvidia.com>
Co-authored-by: jianyingzhu <joeyzhu@nvidia.com>
Co-authored-by: root <root@a4u8g-0120.ipp2a2.colossus.nvidia.com>
Co-authored-by: Yihan Chen <yingluosanqian@example.com>
Co-authored-by: 陈一涵 <yingluosanqian@gmail.com>
Co-authored-by: Xiaoyu Zhang <35585791+BBuf@users.noreply.github.com>
|
2026-02-04 13:46:20 +08:00 |
|
Kun Lin
|
669a9bd180
|
Support Markdown/Notebook-Friendly Documentation Export for Downstream Integration (copy all markdown and rst files) (#18223)
|
2026-02-03 20:53:17 -08:00 |
|
Douglas Yang
|
b7c1dfc602
|
fix: bumping nightly whl version (#18212)
|
2026-02-03 20:43:38 -08:00 |
|
strgrb
|
37c33cc0aa
|
fuse qkvbfg linear into one gemm and f_b g_b into batched gemm. (#17801)
|
2026-02-04 11:41:26 +08:00 |
|
Aurick Qiao
|
c1d529c196
|
Fix Session for multimodal and expose it through Engine (#18152)
|
2026-02-04 10:33:27 +08:00 |
|
Qi Jia
|
1f72f66c6d
|
[Docs] fix readme typo (#18207)
|
2026-02-03 17:37:28 -08:00 |
|
wxy
|
da758ed601
|
[diffusion] fix: fix server cache-dit bug under continuous dynamic requests (#17140)
|
2026-02-04 09:03:37 +08:00 |
|
Douglas Yang
|
ae004e15c9
|
fix: ensuring nightly whls are tagged with latest commit (#18204)
|
2026-02-03 15:54:41 -08:00 |
|
satyamk7054
|
793bf9fc06
|
Update weight rename check for Qwen3 Embeddings (#17535)
|
2026-02-03 13:55:11 -08:00 |
|
Hudson Xing
|
e867040fc6
|
add streaming parallel tool call test case (#18097)
|
2026-02-03 12:46:01 -08:00 |
|
R0CKSTAR
|
7de650c83c
|
[diffusion] hardware: support diffusion models on MTGPU (doc, 6/N) (#17346)
Signed-off-by: Xiaodong Ye <xiaodong.ye@mthreads.com>
|
2026-02-03 12:44:57 -08:00 |
|
R0CKSTAR
|
ec2461bc16
|
[diffusion] hardware: support diffusion models on MTGPU (multi-GPU, 5/N) (#17318)
Signed-off-by: Xiaodong Ye <xiaodong.ye@mthreads.com>
|
2026-02-03 12:44:22 -08:00 |
|
R0CKSTAR
|
acf724b036
|
[Diffusion] Only import sgl_kernel in custom op cuda path (SiluAndMul and RMSNorm) (#15592)
Signed-off-by: Xiaodong Ye <xiaodong.ye@mthreads.com>
|
2026-02-03 12:42:58 -08:00 |
|
Vladislav Nosivskoy
|
e166ca8758
|
[HiCache] feat: Add detailed cache hit breakdown for HiCache in sglext and Prometheus metrics (#17648)
Signed-off-by: Vladislav Nosivskoy <vladnosiv@gmail.com>
Co-authored-by: ishandhanani <82981111+ishandhanani@users.noreply.github.com>
|
2026-02-03 11:45:35 -08:00 |
|
Even Zhou
|
d48bbe3bed
|
[CI][NPU] Bugfix import sgl-kernel error (#18173)
|
2026-02-03 11:39:38 -08:00 |
|
DiweiSun
|
495290aefd
|
enable ut test for xpu devices (#11712)
Co-authored-by: jundu <jun.du@intel.com>
Co-authored-by: Gao, Pengfei <pengfei.gao@intel.com>
|
2026-02-03 11:15:14 -08:00 |
|
ishandhanani
|
0a6925639b
|
ci: improve docker for cu13 builds (#18194)
|
2026-02-03 11:09:38 -08:00 |
|
Kangyan-Zhou
|
0db6fd4dbe
|
Revert broken sgl_kernel exclusion patterns in paths-filter (#18193)
|
2026-02-03 10:56:44 -08:00 |
|
ishandhanani
|
820df545f2
|
fix: add cu13 dev container to our release (#18192)
|
2026-02-03 10:42:05 -08:00 |
|
elvischenv
|
99fab2ce67
|
[Bugfix] Fix Mistral Large 3 NVFP4 TRTLLM MoE (#18065)
|
2026-02-03 20:32:49 +08:00 |
|
Lewis
|
a45647bce1
|
[PD] feat: support mooncake intra-node nvlink kv transfer (#17866)
Co-authored-by: 百麒 <yaozhong.lyz@alibaba-inc.com>
Co-authored-by: Teng Ma <teng-ma@linux.alibaba.com>
|
2026-02-03 17:47:52 +08:00 |
|
Xiaowei Wang
|
cc69ac9e7a
|
Warmup before profiling prefill latency for dynamic chunk sizing (#17198)
Co-authored-by: Shangming Cai <csmthu@gmail.com>
|
2026-02-03 17:45:23 +08:00 |
|
Zhaoyi Li
|
8e933e1914
|
AMD PD/D PR ci (#17183)
Co-authored-by: YC Tseng <yctseng@amd.com>
Co-authored-by: Bingxu Chen <bingxche@amd.com>
Co-authored-by: bingxche <Bingxu.Chen@amd.com>
|
2026-02-02 23:29:14 -08:00 |
|
Mohammad Miadh Angkad
|
25508d11c0
|
[Docker] Remove hardcoded America/Los_Angeles timezone, default to UTC (#18121)
|
2026-02-02 23:22:15 -08:00 |
|
Mohammad Miadh Angkad
|
6f6b9c6e42
|
[Perf] Use safetensors load_file in multithread loader (#18124)
|
2026-02-02 23:21:13 -08:00 |
|
fatSheep
|
7a9d9c79d1
|
[HiCache] fix: apply extra_backend_tag in Mooncake batch_exists (#17265)
|
2026-02-02 22:54:56 -08:00 |
|
Viacheslav
|
74f716dbd7
|
Gigachat 3 tool parser and tests (#14765)
|
2026-02-02 22:28:34 -08:00 |
|
Kaixi Hou
|
4181290efd
|
[NVIDIA] Add --top-k argument to run_eval.py (#18025)
Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>
|
2026-02-02 22:17:53 -08:00 |
|
Glen Liu
|
fe57a887b1
|
[TestFix] use unit tests for LoRA overlap loading tests (#18140)
|
2026-02-02 22:06:50 -08:00 |
|
Kun Lin
|
f032c4f3d6
|
Support Markdown/Notebook-Friendly Documentation Export for Downstream Integration (#18131)
Co-authored-by: 赵晨阳 <zhaochen20@outlook.com>
|
2026-02-02 21:43:20 -08:00 |
|
b8zhong
|
78bf13db44
|
MoE Refactor: Refactor modelopt_quant.py -> flashinfer_trllm.py (#16685)
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
|
2026-02-02 20:45:14 -08:00 |
|
Xiaoyu Zhang
|
eedd472025
|
[Diffusion] fix serving image_edit get input image bug (#18109)
|
2026-02-03 12:17:16 +08:00 |
|
Hank Han
|
e484c90cc7
|
Add triton_fused_moe config for GLM-4.7-FP8 tp8 H20 H20-3e (#18091)
|
2026-02-03 12:08:23 +08:00 |
|
Linyu Wu
|
9b1619c148
|
[Move sgl-kernel Kernel to JIT] Add JIT concat MLA kernels (#17889)
|
2026-02-03 10:49:17 +08:00 |
|
Mick
|
62004fd2be
|
[diffusion] UX: improve logging (#18122)
|
2026-02-03 10:35:05 +08:00 |
|
zhangheng
|
180594358b
|
[HiCache]: Support DeepSeek v32 cpu offloading (#17415)
Co-authored-by: 晟海 <huangtingwei.htw@antgroup.com>
|
2026-02-02 18:07:37 -08:00 |
|
Xiaoyu Zhang
|
a1bbc892af
|
[Diffsuion & JIT_kernel] QKNorm cross heads kernel (#18073)
|
2026-02-03 10:03:17 +08:00 |
|
EkiRui
|
fd983b09b6
|
[Performance] Optimize radix cache eviction performance (#14339)
Signed-off-by: Xingrui Yi <yixingrui@linux.alibaba.com>
Co-authored-by: Xuchun Shang <xuchun.shang@gmail.com>
|
2026-02-03 09:44:20 +08:00 |
|
Douglas Yang
|
c8da307d7e
|
feature: adding gpt-oss 120b nightly test (#18134)
|
2026-02-02 17:11:28 -08:00 |
|
Alison Shao
|
28e2340725
|
Fix HF hub race condition in CI by coordinating model downloads across TP ranks (#17787)
Co-authored-by: Kangyan-Zhou <zky314343421@gmail.com>
|
2026-02-02 14:57:45 -08:00 |
|
Alison Shao
|
812fd47cb4
|
Re-enable test_mla_int8_deepseek_v3.py after HF token fix (#18123)
|
2026-02-02 14:38:42 -08:00 |
|
cctry
|
027f314050
|
[Fix] data race in req_to_token pool (#17850)
|
2026-02-02 14:38:15 -08:00 |
|
TZHelloWorld
|
cbf1500390
|
[MiMoV2Flash] [feat]: support two batch overlap (#17634)
|
2026-02-02 14:03:57 -08:00 |
|
khalilzhk
|
b0a6d5244c
|
[NPU] support dsv32 radixcache on ascend (#17964)
|
2026-02-03 03:34:12 +08:00 |
|
Yongfei Xu
|
677f3c49da
|
[DeepSeek V3.2] [Bugfix] slice indexer and padding fa3 when can not run cuda graph (#17076)
|
2026-02-03 01:32:20 +08:00 |
|
Yuhao Yang
|
980d2936cd
|
model: support Step-3.5-Flash (#18084)
Co-authored-by: ltd0924 <ltd0924@sina.com>
|
2026-02-03 00:40:07 +08:00 |
|
Sugar920
|
c781db0f6c
|
[NPU] update nightly tests (#17952)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: cy <chenyang08056032@163.com>
|
2026-02-03 00:13:30 +08:00 |
|