Ke Bao
|
ce8a6ac690
|
Evict swa kv cache during decoding (#17220)
|
2026-01-19 22:36:52 +08:00 |
|
ybyang
|
ebca5879a1
|
Fix v32 continue_final_message not work (#16567)
|
2026-01-19 21:45:39 +08:00 |
|
Xiaoyu Zhang
|
cc410a1088
|
[Diffusion] Apply qknorm to flux2 and apply lightx2v rms_norm_one_pass kernel(without residual) (#17305)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2026-01-19 21:25:33 +08:00 |
|
b8zhong
|
f374623fa9
|
[Refactor] Set fp4-gemm-backend=auto on SM100 and rename fp4-gemm-backend with flashinfer_ prefix (#17309)
|
2026-01-19 20:09:07 +08:00 |
|
Shu Wang
|
5c02217746
|
Inclusion of nvfp4 blockscale in EPLB Rebalance (#17158)
|
2026-01-19 17:45:27 +08:00 |
|
Alison Shao
|
8916b9d080
|
Migrate performance, accuracy, and quantization tests to CI registry (#17177)
Co-authored-by: Kangyan-Zhou <zky314343421@gmail.com>
|
2026-01-18 23:25:24 -08:00 |
|
Xiaoyu Zhang
|
a3d9a21882
|
Revert "[Perf] fuse q, k norm for Flux2Attention (#17241)" (#17332)
|
2026-01-19 15:24:11 +08:00 |
|
Alison Shao
|
fb88fb672e
|
fix(ci): rate limit and permission errors in trace publishing (#17238)
|
2026-01-18 23:20:22 -08:00 |
|
Alison Shao
|
2d72e168fd
|
[CI] Add partition to stage-b-test-large-1-gpu (11->12) (#17245)
|
2026-01-18 23:19:39 -08:00 |
|
Minglei Zhu
|
64946679b5
|
[Perf] fuse q, k norm for Flux2Attention (#17241)
Co-authored-by: Minglei Zhu <zminglei@linkedin.com>
|
2026-01-19 14:33:34 +08:00 |
|
Kartik Ramesh
|
5836324c55
|
KV Cache Events with Attention DP bug fix (#16030) (#16412)
|
2026-01-19 14:13:48 +08:00 |
|
yudian0504
|
9fe56cd0fb
|
Fix kernel selection in biased_grouped_topk_gpu (#17325)
|
2026-01-19 14:06:11 +08:00 |
|
Gaoji Liu
|
858a4d659b
|
support new qwen3_coder_detector (#16744)
Co-authored-by: liugaoji.lgj <liugaoji.lgj@alibaba-inc.com>
|
2026-01-18 21:16:27 -08:00 |
|
Lianmin Zheng
|
e619f53113
|
[Auto Sync] Update tokenizer_manager.py (20260119) (#17317)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
|
2026-01-18 20:32:19 -08:00 |
|
Lianmin Zheng
|
fc4b932f4e
|
Update code sync scripts (#17319)
|
2026-01-18 19:57:28 -08:00 |
|
Yongfei Xu
|
d2105d4abd
|
[DeepSeek v3.2] Opt MTP decode cuda batch sizes and nsa implementation (#16961)
|
2026-01-19 11:54:11 +08:00 |
|
Lee Nau
|
84c8390514
|
Use dsv3 optimized routing fused_topk_deepseek instead of moe_fused_gate (#15347)
|
2026-01-19 11:50:16 +08:00 |
|
Baizhou Zhang
|
ea879c7739
|
[Minor] Correct sglang version when installing from source (#17315)
|
2026-01-18 19:36:16 -08:00 |
|
Shangming Cai
|
0227db8926
|
[PD] Optimize MHA models pp util calculation logic (#17306)
|
2026-01-19 11:23:15 +08:00 |
|
Glen Liu
|
ad1b4e4728
|
[Feature] overlap LoRA weight loading with compute (#15512)
|
2026-01-19 10:43:17 +08:00 |
|
Mick
|
51f147ada3
|
Update CODEOWNERS for multimodal_gen (#17308)
Co-authored-by: Xiaoyu Zhang <35585791+BBuf@users.noreply.github.com>
|
2026-01-19 08:39:20 +08:00 |
|
Koushik Dutta
|
d3eafc7357
|
[GLM 4.7] Add RTX 6000 Pro aka sm120 (#17235)
Co-authored-by: root <root@ubuntu-nvidia.localdomain>
|
2026-01-18 13:33:19 -08:00 |
|
Jinyan Chen
|
e00b43442d
|
[jit-kernel] Add CuTe DSL GDN Decode Kernel (#15631)
Co-authored-by: Jinyan Chen <jinyanc@nvidia.com>
|
2026-01-18 12:54:36 -08:00 |
|
Todobe
|
733de6be31
|
[NPU]Support GPT-OSS for NPU (#14197)
|
2026-01-19 04:13:41 +08:00 |
|
Jerry Ji
|
93433726eb
|
[Refactor] Split out deepseek v2 weight loader function into mixin (#16649)
|
2026-01-19 00:41:59 +08:00 |
|
Xiaoyu Zhang
|
330605cc88
|
[Diffusion] Apply jit qk_norm to flux1 (#17296)
|
2026-01-19 00:28:36 +08:00 |
|
Ch3ngY1
|
bb6055b43c
|
[Bugfix] Fix PD accuracy when MTP is not configured on the prefill node (#17212)
Co-authored-by: Shangming Cai <csmthu@gmail.com>
|
2026-01-18 23:37:58 +08:00 |
|
b8zhong
|
4df74eb576
|
[Refactor] Add -fp4-gemm-backend to replace SGLANG_FLASHINFER_FP4_GEMM_BACKEND (#16534)
Co-authored-by: Vincent Zhong <207368749+vincentzed@users.noreply.github.com>
|
2026-01-18 23:25:46 +08:00 |
|
Ke Bao
|
f3a7c7dcd9
|
Move radix cache related tests (#17295)
|
2026-01-18 17:47:32 +08:00 |
|
Xinyuan Tong
|
2069050d3f
|
fix: Handle multiple named chat templates in HuggingFace tokenizers (#17236)
Signed-off-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
|
2026-01-18 17:20:04 +08:00 |
|
Ke Bao
|
1fe0c82f67
|
Add kl test for swa radix cache (#17281)
|
2026-01-18 16:19:57 +08:00 |
|
YAMY
|
a45e0e5df4
|
[SPEC_V2] Enable cudagraph draft_extend for trtllm_mla_backend and Acclen Fix for DP under cudagraph mode (#16974)
|
2026-01-18 15:56:21 +08:00 |
|
Ke Bao
|
f78201f3a9
|
Tiny fix comment typo (#17287)
|
2026-01-18 15:06:33 +08:00 |
|
Changyi Yang
|
8fd3399880
|
[diffusion] fix: set guidance_scale default to None (#17182)
|
2026-01-18 15:04:11 +08:00 |
|
Mohammad Miadh Angkad
|
088758c1c1
|
[Tiny] Improve docs (#17264)
|
2026-01-18 14:57:01 +08:00 |
|
Yuan Luo
|
6d29d8ab16
|
[VLM][Reland] Refactor load_mm_data to improve performance (#16152)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
|
2026-01-18 14:11:17 +08:00 |
|
Ke Bao
|
e499258e97
|
Use swa radix cache and memory pool for gpt-oss model (#17261)
|
2026-01-18 13:47:07 +08:00 |
|
Mick
|
09491a9bcd
|
cli: support sglang version (#17250)
|
2026-01-18 13:20:24 +08:00 |
|
Alison Shao
|
7edb06158e
|
Add runner utilization report workflow (#17234)
|
2026-01-17 19:28:05 -08:00 |
|
Lancer
|
e486a4dac1
|
[diffusion] feat: support default 4-step inference for Flux2-Klein distilled models (#17225)
Signed-off-by: Lancer <maruixiang6688@gmail.com>
|
2026-01-18 10:15:50 +08:00 |
|
Hudson Xing
|
90399cbc07
|
fix(ci): recover from corrupted MMMU parquet cache (#17256)
|
2026-01-17 17:32:01 -08:00 |
|
Michael
|
53609e5e5b
|
Revert "[Diffusion] Move diffusion time embedding to jit kernel" (#17257)
|
2026-01-17 21:29:22 +08:00 |
|
fzyzcjy
|
9c2530642c
|
Change routing policy API to be async to support more policies (#17048)
|
2026-01-17 17:12:51 +08:00 |
|
Zehuan Li
|
d2c863878c
|
[DLLM] Implement initial dynamic batching for diffusion LLM (#14883)
|
2026-01-17 16:48:15 +08:00 |
|
Yi Zhang
|
737a1183d6
|
[BUGFIX] fix radix cache memory consumption to avoid OOM (#17191)
Co-authored-by: Liangsheng Yin <lsyincs@gmail.com>
|
2026-01-17 16:47:37 +08:00 |
|
Mick
|
dc743fe4ba
|
[diffusion] chore: clean srt imports (#17252)
|
2026-01-17 15:47:49 +08:00 |
|
Nan Jiang
|
dd99f818e0
|
fix: fix regression and unclear pattern (#16561)
|
2026-01-16 23:21:42 -08:00 |
|
Hudson Xing
|
8ce64aa155
|
fix(ci): skip offline mode for LoRA scenarios (#17248)
|
2026-01-16 23:06:20 -08:00 |
|
Mick
|
eb768189ee
|
[diffusion] doc: add instruction for adding performance baseline of new model (#17249)
|
2026-01-17 12:24:59 +08:00 |
|
Xiaoyu Zhang
|
2cdd4370bc
|
[Diffusion] Move diffusion time embedding to jit kernel (#16879)
|
2026-01-17 12:21:22 +08:00 |
|