fzyzcjy
|
02816abc0d
|
Flip dumper to disable by default and refactor environment handling (#18878)
|
2026-02-16 13:29:32 +08:00 |
|
Rain Jiang
|
0ffd0a3995
|
Nsa trtllm mla sparse fp8 support with Deepseek v3.2 NVFP4 (#18389)
|
2026-02-16 09:29:54 +08:00 |
|
Mohammad Miadh Angkad
|
8290171f52
|
[CI] Remove --mem-fraction-static 0.93 from gpt-oss test (#18869)
|
2026-02-16 09:24:11 +08:00 |
|
Chanh Nguyen
|
597d17dd18
|
Use ephemeral nccl port via get_free_port() (#18009)
Co-authored-by: Chanh Nguyen <cnguyen@linkedin.com>
|
2026-02-16 00:32:47 +08:00 |
|
Zack Yu
|
536ed3143b
|
test: add test for Modelopt FP8 on SM90 (#18463)
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2026-02-16 00:29:37 +08:00 |
|
SoluMilken
|
07a24f1a38
|
update pre-commit config (#18860)
|
2026-02-16 00:18:31 +08:00 |
|
Michael
|
88010e9601
|
[AMD] Fix nightly 1-GPU test failures and bench_serving regression (#18761)
Co-authored-by: michaelzhang-ai <michaelzhang-ai@users.noreply.github.com>
|
2026-02-15 20:36:47 +08:00 |
|
fzyzcjy
|
90555a0228
|
Add missing dumper tests (#18859)
|
2026-02-15 18:42:57 +08:00 |
|
fzyzcjy
|
4c7f986c6b
|
Extract dumper and prefill delayer tests common utils (#18857)
|
2026-02-15 18:33:23 +08:00 |
|
Bhavneek Singh
|
1ce3420784
|
Model: Support IBM Granite (Dense/Mamba + MoE) (#18040)
|
2026-02-15 11:24:41 +08:00 |
|
Xiaoyu Zhang
|
c29394e3c8
|
[kernel slimming] Move fast_hadamard_transform to jit_kernel (#18475)
|
2026-02-14 23:06:21 +08:00 |
|
Kangyan-Zhou
|
ae95869292
|
Enable SGLANG_ENABLE_SPEC_V2 for nightly speculative decoding tests (#18719)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
|
2026-02-14 23:00:33 +08:00 |
|
Raayan Dhar
|
92cdd398cd
|
feat: Support mrope_section with rope_type: "yarn" (#13313)
Signed-off-by: Raayan Dhar raayan.dhar@gmail.com <raayan.dhar@gmail.com>
Signed-off-by: raayandhar <raayan.dhar@gmail.com>
|
2026-02-14 22:51:44 +08:00 |
|
Ke Bao
|
f51e9d9ca1
|
Add ci test for ring model (#18829)
|
2026-02-14 22:20:23 +08:00 |
|
JD
|
f6c18c3a85
|
Fix/partial gen from waiting queue miss metadata (#17610)
|
2026-02-13 19:04:08 -08:00 |
|
Liangsheng Yin
|
dcea74d63f
|
Add timeout abort kits for normal / eagle. (#18815)
|
2026-02-13 17:57:30 -08:00 |
|
Minglei Zhu
|
8be18c655d
|
[Perf] refactor piecewise cuda graph support of Qwen3-Next (#17613)
|
2026-02-14 09:30:50 +08:00 |
|
shuwenn
|
3299c4f9c1
|
[CI] feat: add early exit to wait_for_server when process dies (#18602)
|
2026-02-13 16:46:09 -08:00 |
|
JD
|
191d354f53
|
fix double-free kv cache for requests that have already finished and been freed during preemption (#18694)
|
2026-02-13 13:17:44 -08:00 |
|
dongjiyingdjy
|
8b4c364960
|
refactor context parallel state (#17213)
Co-authored-by: Shunkang <182541032+Shunkangz@users.noreply.github.co>
|
2026-02-13 23:18:17 +08:00 |
|
Baizhou Zhang
|
9a32f8ccb9
|
[CI] Move test_load_lora_from_tensor test to H100 (#18797)
|
2026-02-13 21:28:00 +08:00 |
|
Liangsheng Yin
|
e6f7a372ef
|
Rename request timeout env vars for waiting/running stages (#18766)
|
2026-02-12 22:58:40 -08:00 |
|
Alison Shao
|
0abe4a22c6
|
Fix flaky penalty tests by using higher temperature for effect comparison (#18380)
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
|
2026-02-12 21:08:37 +08:00 |
|
YC Tseng
|
20554a0a4f
|
[AMD] rocm 7.2 image release, PR test, Nightly Test (#17799)
Co-authored-by: Alan Kao <akao@amd.com>
Co-authored-by: bingxche <Bingxu.Chen@amd.com>
Co-authored-by: Michael <13900043+michaelzhang-ai@users.noreply.github.com>
|
2026-02-11 21:29:25 -08:00 |
|
Piotr Mazurek
|
ded068a76e
|
Add LMF2 MoE model architecture (#17997)
|
2026-02-12 01:03:43 +08:00 |
|
Vedant V Jhaveri
|
98b5013d59
|
add support to enable lora with embedding models (#17780)
Co-authored-by: Vedant Jhaveri <vjhaveri@linkedin.com>
|
2026-02-11 23:19:40 +08:00 |
|
McZyWu
|
4f7422f7ba
|
[NPU] support model skywork-reward-gemma2-2-27B-v0.2 (#16947)
Co-authored-by: cy <chenyang08056032@163.com>
|
2026-02-11 15:34:53 +08:00 |
|
Michael
|
d84d2063d3
|
[AMD] Fix Janus-Pro crash and add Kimi-K2.5 nightly test (#18269)
|
2026-02-10 22:33:13 -08:00 |
|
Liangsheng Yin
|
cd90346a2b
|
Add cache hit rate UT (#18566)
|
2026-02-10 21:27:41 -08:00 |
|
Liangsheng Yin
|
50f74285e9
|
Tiny fix regex warning (#18592)
|
2026-02-10 21:07:53 -08:00 |
|
cutetocute
|
8d2892330c
|
chore: fix some typos (#18577)
Co-authored-by: Liangsheng Yin <lsyincs@gmail.com>
|
2026-02-10 20:47:41 -08:00 |
|
Zehuan Li
|
26f2b3798d
|
[DLLM] Basic dLLM scheduling strategy and implementation (#17484)
Signed-off-by: Zehuan Li <lizehuan.lzh@antgroup.com>
|
2026-02-10 16:54:15 +08:00 |
|
Liangsheng Yin
|
2825f5d8e0
|
Tiny fix wrong metric collect key name forward_prefill -> forward_extend (#18506)
|
2026-02-09 16:23:41 -08:00 |
|
Bingxu Chen
|
3f3c201243
|
[AMD] Update aiter to v0.1.10.post2 (#18423)
Co-authored-by: kkHuang-amd <wunhuang@amd.com>
Co-authored-by: YC Tseng <yctseng@amd.com>
|
2026-02-08 22:08:24 -08:00 |
|
Shangming Cai
|
52401bec1d
|
chore: bump mooncake version to 0.3.9 (#18316)
Signed-off-by: Shangming Cai <csmthu@gmail.com>
|
2026-02-07 17:30:01 +08:00 |
|
Alison Shao
|
bedade1ef0
|
Merge stage-c-test-large-4-gpu suites into partitioned suites (#18325)
|
2026-02-06 15:32:33 -08:00 |
|
shaharmor98
|
c6aa1863be
|
Add Nemotron 3 Nano tests (#18119)
Signed-off-by: Shahar Mor <smor@nvidia.com>
|
2026-02-06 23:55:42 +08:00 |
|
Alison Shao
|
d22163eb8c
|
Fix flaky test_frequency_penalty_reduces_word_repetition by using deterministic seeds (#18285)
|
2026-02-05 10:24:18 -08:00 |
|
Alison Shao
|
c910829708
|
Fix test_return_routed_experts to use response-level sglext (#18274)
|
2026-02-04 20:16:01 -08:00 |
|
Michael
|
6fd878b41d
|
[AMD] Add kimi mi35x nightly test, folder organization and several stability fixes (#17895)
|
2026-02-04 12:03:57 -08:00 |
|
DiweiSun
|
495290aefd
|
enable ut test for xpu devices (#11712)
Co-authored-by: jundu <jun.du@intel.com>
Co-authored-by: Gao, Pengfei <pengfei.gao@intel.com>
|
2026-02-03 11:15:14 -08:00 |
|
elvischenv
|
99fab2ce67
|
[Bugfix] Fix Mistral Large 3 NVFP4 TRTLLM MoE (#18065)
|
2026-02-03 20:32:49 +08:00 |
|
Zhaoyi Li
|
8e933e1914
|
AMD PD/D PR ci (#17183)
Co-authored-by: YC Tseng <yctseng@amd.com>
Co-authored-by: Bingxu Chen <bingxche@amd.com>
Co-authored-by: bingxche <Bingxu.Chen@amd.com>
|
2026-02-02 23:29:14 -08:00 |
|
Viacheslav
|
74f716dbd7
|
Gigachat 3 tool parser and tests (#14765)
|
2026-02-02 22:28:34 -08:00 |
|
Glen Liu
|
fe57a887b1
|
[TestFix] use unit tests for LoRA overlap loading tests (#18140)
|
2026-02-02 22:06:50 -08:00 |
|
zhangheng
|
180594358b
|
[HiCache]: Support DeepSeek v32 cpu offloading (#17415)
Co-authored-by: 晟海 <huangtingwei.htw@antgroup.com>
|
2026-02-02 18:07:37 -08:00 |
|
Douglas Yang
|
c8da307d7e
|
feature: adding gpt-oss 120b nightly test (#18134)
|
2026-02-02 17:11:28 -08:00 |
|
Alison Shao
|
812fd47cb4
|
Re-enable test_mla_int8_deepseek_v3.py after HF token fix (#18123)
|
2026-02-02 14:38:42 -08:00 |
|
cctry
|
027f314050
|
[Fix] data race in req_to_token pool (#17850)
|
2026-02-02 14:38:15 -08:00 |
|
Yongfei Xu
|
677f3c49da
|
[DeepSeek V3.2] [Bugfix] slice indexer and padding fa3 when can not run cuda graph (#17076)
|
2026-02-03 01:32:20 +08:00 |
|