Douglas Yang
|
c8da307d7e
|
feature: adding gpt-oss 120b nightly test (#18134)
|
2026-02-02 17:11:28 -08:00 |
|
Alison Shao
|
812fd47cb4
|
Re-enable test_mla_int8_deepseek_v3.py after HF token fix (#18123)
|
2026-02-02 14:38:42 -08:00 |
|
cctry
|
027f314050
|
[Fix] data race in req_to_token pool (#17850)
|
2026-02-02 14:38:15 -08:00 |
|
Yongfei Xu
|
677f3c49da
|
[DeepSeek V3.2] [Bugfix] slice indexer and padding fa3 when can not run cuda graph (#17076)
|
2026-02-03 01:32:20 +08:00 |
|
Sugar920
|
c781db0f6c
|
[NPU] update nightly tests (#17952)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: cy <chenyang08056032@163.com>
|
2026-02-03 00:13:30 +08:00 |
|
Byron Hsu
|
5636d16dec
|
[CI] Add logs to debug TestOpenAIServer.test_completion_stream (#17471)
|
2026-02-02 02:38:08 -08:00 |
|
Yuhao Yang
|
d11ccc0a0a
|
fix: avoid double reduce in VLM dp attention (#17991)
|
2026-02-02 09:44:32 +08:00 |
|
Glen Liu
|
99dad105fd
|
[TestFix] rewrite LoRA overlap loading tests (#18047)
|
2026-02-01 14:52:08 -08:00 |
|
Alison Shao
|
56907cbcb1
|
Move deleted 8-GPU tests to test/manual/ (#18060)
|
2026-02-01 00:21:56 -08:00 |
|
Alison Shao
|
a0bae4c343
|
Migrate 4-GPU/8-GPU workflow jobs to stage-c and add CI registry decorators (#17299)
|
2026-01-31 22:37:22 -08:00 |
|
Alison Shao
|
95180484e9
|
Disable test_mla_int8_deepseek_v3.py temporarily (#18057)
|
2026-01-31 22:33:43 -08:00 |
|
b8zhong
|
398d13a189
|
[Perf] Add Flashinfer DeepGEMM SM90 for SwapAB Optimization (#15514)
Co-authored-by: Brayden Zhong <b8zhong@users.noreply.github.com>
|
2026-02-01 08:56:23 +08:00 |
|
Hudson Xing
|
c72bf50706
|
add reasoning_tokens usage test for tool call (#18022)
|
2026-01-30 21:09:23 -08:00 |
|
JiaruiChang5268
|
e86476acfc
|
[NPU] support llama-3.2-11B-vision-instruct mode for NPU (#17492)
Co-authored-by: McZyWu <zhuoyun.wu.23@ucl.ac.uk>
Co-authored-by: chenyang08056032 <chenyang08056032@163.com>
Co-authored-by: Hexq0210 <893781835@qq.com>
|
2026-01-31 08:49:38 +08:00 |
|
Fan Yin
|
8ce9609fa2
|
fix: fix SHM pointer re-serialization in DP attention (#17930)
|
2026-01-30 17:03:30 +08:00 |
|
McZyWu
|
70db3398d1
|
[NPU] enhance accuracy for model kimi-vl-a3b-instruct (#17480)
Co-authored-by: cy <chenyang08056032@163.com>
|
2026-01-30 15:19:42 +08:00 |
|
jianan-gu
|
c35aa0238c
|
[CPU][INT4] Add INT4 kernels for CPU (#8226)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2026-01-29 22:30:13 -08:00 |
|
Ma Mingfei
|
88f7759402
|
[CPU] optimize flash_attn_varlen_func (#15708)
|
2026-01-29 22:07:05 -08:00 |
|
StonyPort
|
2b3408ff14
|
feat: add forward timeout (#17831)
Co-authored-by: qiuxuan.lzw <qiuxuan.lzw@alibaba-inc.com>
|
2026-01-30 08:52:29 +08:00 |
|
Hudson Xing
|
d417c6809e
|
Add tool call tests for DeepSeek V3.2 in nightly CI (#17951)
|
2026-01-29 09:50:54 -08:00 |
|
22dimensions
|
7b79326751
|
[NPU] support GPTQ quantization on npu (#15203)
Signed-off-by: 22dimensions <waitingwind@foxmail.com>
|
2026-01-29 15:48:18 +08:00 |
|
Niko Ma
|
cbf90d70ff
|
[PD] Support KV transfer with MORI-IO (#14626)
Co-authored-by: cwortman-amd <cwortman@amd.com>
|
2026-01-28 23:22:41 -08:00 |
|
Joe Redmond
|
0ff0d181ca
|
feat: add custom request header logging (#17786)
|
2026-01-28 19:33:08 -08:00 |
|
kk
|
f1384f5293
|
Integration mori backend for EP a2a data communication (#17012)
Co-authored-by: Duyi-Wang <duyi.wang@amd.com>
Co-authored-by: billishyahao <bill.he@amd.com>
Co-authored-by: HaiShaw <hixiao@gmail.com>
|
2026-01-28 19:07:34 -08:00 |
|
Артем Савкин
|
b77b0ffd60
|
[NPU] NZ for non-quantized MOE, Qwen3 MOE double memory consumption fix (#15904)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2026-01-29 00:55:08 +08:00 |
|
Michael
|
f8636fbb25
|
[AMD] Add Kimi-K2, DeepSeek-V3.2 tests to nightly CI (#17523)
Co-authored-by: YC Tseng <yctseng@amd.com>
|
2026-01-28 00:55:46 -08:00 |
|
Even Zhou
|
6077de1237
|
[CI] [NPU] npu ci use existing modelscope model (#17868)
|
2026-01-28 16:51:41 +08:00 |
|
YC Tseng
|
52bca42870
|
[AMD] CI - enable deepseekv3.2 on MI325-8gpu and merge perf/accuracy test suites into stage-b suites (#17633)
Co-authored-by: Bingxu Chen <Bingxu.Chen@amd.com>
|
2026-01-27 18:54:36 -08:00 |
|
Baizhou Zhang
|
1d942e4eef
|
[DeepSeek] Update tests and document for DeepSeek V3.2 NVFP4 checkpoint (#17657)
|
2026-01-27 22:10:57 +08:00 |
|
shuwenn
|
57e432d951
|
fix: preserve disconnect events in api key middleware (#17253)
|
2026-01-26 22:48:24 -05:00 |
|
shuwenn
|
fd3b179ffd
|
[HiCache][HA 1/N] Support HiCache storage runtime attach/detach (#15892)
|
2026-01-26 19:33:19 -08:00 |
|
Shangming Cai
|
3ad3268e06
|
[CI] Skip PD hybrid attention test with different TP temporarily (#17791)
|
2026-01-27 10:59:33 +08:00 |
|
Alison Shao
|
6c0f9b4824
|
Add test_gpt_oss_4gpu.py to B200 test suite (#17743)
|
2026-01-26 21:57:06 +08:00 |
|
sogalin
|
738b1ac988
|
[AMD CI] Add moonshotai/Kimi-K2-Instruct-0905 testcases (#17656)
|
2026-01-26 02:12:34 -08:00 |
|
McZyWu
|
2734b23481
|
accuracy enhancement for baichuan2-13B for npu (#16868)
Co-authored-by: cy <chenyang08056032@163.com>
|
2026-01-26 16:14:35 +08:00 |
|
Alison Shao
|
30b3192039
|
Merge performance/accuracy test suites into regular stage-b suites (#17609)
|
2026-01-25 22:49:19 -08:00 |
|
Kangyan-Zhou
|
5aaedac3c6
|
Add EP=2 to qwen235b nightly tests (#17738)
|
2026-01-25 21:36:04 -08:00 |
|
Kangyan-Zhou
|
592603d77b
|
Fix flaky streaming logprobs test by handling detokenizer text buffering (#17687)
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
|
2026-01-25 15:09:06 -08:00 |
|
Kangyan-Zhou
|
9123491430
|
A few updates to the night tests (#17694)
|
2026-01-25 11:20:17 -08:00 |
|
xjx471258437
|
9bd92ba0f6
|
Support PD disaggregation with different TP/DP size for Qwen3-Next (#16056)
Co-authored-by: xjx392321 <xjx392321@alibaba-inc.com>
|
2026-01-25 15:34:02 +08:00 |
|
Kangyan-Zhou
|
69a7a70e47
|
Temporarily disable lora overlap loading test due to flakiness (#17683)
|
2026-01-24 12:24:44 -08:00 |
|
Kangyan-Zhou
|
137eb5b95c
|
Fix NSA indexer test and move it to pre commit test (#17682)
|
2026-01-24 12:06:18 -08:00 |
|
strgrb
|
176da1bbdd
|
Fix: mistake sigmoid in kda (#17508)
|
2026-01-24 13:35:14 +08:00 |
|
Lianmin Zheng
|
bc6f0b5ce7
|
[Auto Sync] Update logits_processor.py, test_logprobs.py (20260124) (#17664)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: yehu-ux <yehu@x.ai>
|
2026-01-23 17:57:41 -08:00 |
|
McZyWu
|
b4a611fb33
|
[NPU] solve accuracy problem for stablelm-2-1-6b for npu (#17470)
|
2026-01-24 08:27:38 +08:00 |
|
McZyWu
|
8a5ed2434f
|
[NPU]support model MiniCPM3-4B for npu (#16866)
|
2026-01-24 08:25:12 +08:00 |
|
Mansoor
|
bdaa3de075
|
Add return routed experts to the completions and chat/completions endpoints (#17434)
|
2026-01-23 12:12:36 -08:00 |
|
Nicolas Castet
|
48e9daadff
|
Support symmetric memory pre-allocation to avoid fragmentation (#17089)
|
2026-01-23 17:57:04 +08:00 |
|
Even Zhou
|
69ac8b58f7
|
[NPU] [CI] temporarily disable mtp test (#17614)
|
2026-01-23 15:17:31 +08:00 |
|
Alison Shao
|
d7dd0b8832
|
Re-enable unit-test-deepep-8-gpu and unit-test-backend-4-gpu-gb200 (#17438)
|
2026-01-23 14:31:44 +08:00 |
|