Baizhou Zhang
|
3373545b9f
|
[HotFix]Fix dtype mismatch in nsa indexer on AMD device (#17518)
|
2026-01-22 12:17:02 +08:00 |
|
Piotr Mazurek
|
d6e2b88288
|
Add Liquid Foundation Model (LFM2) (#16890)
|
2026-01-22 11:11:20 +08:00 |
|
Kangyan-Zhou
|
f2ae066a6b
|
Update release-branch-cut.yml for actions: write (#17539)
|
2026-01-21 18:03:27 -08:00 |
|
cen121212
|
0c2993eed0
|
Optimize Qwen3-VL video memory usage (#16366)
|
2026-01-22 09:10:08 +08:00 |
|
Xiaoyu Zhang
|
590969ee9c
|
[Diffusion] Support select fa2 backend in hopper (#17514)
|
2026-01-22 08:23:53 +08:00 |
|
Alison Shao
|
e6ccb2949b
|
Increase wait-for-stage timeouts to handle long queue times (#17536)
|
2026-01-21 16:10:52 -08:00 |
|
Alison Shao
|
d6ea2c529c
|
Fix import path for UnquantizedLinearMethod in test (#17529)
|
2026-01-21 15:34:02 -08:00 |
|
Alison Shao
|
9be2a3a9a3
|
Remove test_gpt_oss_4gpu.py from __not_in_ci__ (keep in per-commit-4-gpu) (#17534)
|
2026-01-21 15:31:03 -08:00 |
|
Lianmin Zheng
|
b74a57a8d9
|
[Auto Sync] Update detokenizer_manager.py, io_struct.py, mu... (20260120) (#17442)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Wangfan Fu <wangfan@x.ai>
|
2026-01-21 14:48:32 -08:00 |
|
DarkSharpness
|
95f59c13fd
|
[Chore] include all jit files in building packages (#17493)
|
2026-01-21 14:48:02 -08:00 |
|
Alison Shao
|
85d9af51da
|
Temporarily disable flaky test_gpt_oss_4gpu.py on B200 (#17528)
|
2026-01-21 14:01:04 -08:00 |
|
Jacob Gordon
|
858f317f13
|
ci(codespell): centralizes list of ignorable words (#17524)
|
2026-01-21 12:29:14 -08:00 |
|
Lingjun Wen
|
cf89351691
|
[new-model] Add support for Cohere2ForCausalLM behind Command-A and Command-R Models (#16927)
|
2026-01-21 12:28:33 -08:00 |
|
Lianmin Zheng
|
1fdf5cac39
|
[Auto Sync] Update environ.py, fp8.py (20260121) (#17486)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Binyao Jiang <byjiang1996@gmail.com>
|
2026-01-21 12:04:09 -08:00 |
|
Jacob Gordon
|
cda43ffa4d
|
ci: avoids duplication of codespell config (#17519)
|
2026-01-21 12:02:37 -08:00 |
|
Yunmeng
|
390898545e
|
[Misc] Fix argument help string formatting (#17416)
|
2026-01-21 09:25:32 -08:00 |
|
YC Tseng
|
b827e9d381
|
[AMD] CI - Fix sgl-kernel unittest (#17490)
|
2026-01-21 09:23:28 -08:00 |
|
Ke Bao
|
d725487dc8
|
Disable swa memory for gpt-oss with spec (#17517)
|
2026-01-22 01:04:19 +08:00 |
|
JiaruiChang5268
|
a95c9f5b81
|
[NPU] Remove paged attention & Change fia to default attention (#17394)
Co-authored-by: Liwansi <62291011+Liwansi@users.noreply.github.com>
Co-authored-by: chenxu214 <justin_cc2025@163.com>
Co-authored-by: chenyang08056032 <chenyang08056032@163.com>
|
2026-01-21 23:58:25 +08:00 |
|
Xiaoyu Zhang
|
19089aa431
|
[Diffusion] Refactor diffusion is_cuda check (#17498)
|
2026-01-21 23:02:24 +08:00 |
|
Qiaolin Yu
|
4f6f5d25c8
|
Support fa4 decoding (#16034)
|
2026-01-21 22:54:02 +08:00 |
|
Yi Zhong
|
458fe5a337
|
[docs] Show user the fastAPI docs available (#17510)
Signed-off-by: vincentzed <207368749+vincentzed@users.noreply.github.com>
|
2026-01-21 14:26:25 +00:00 |
|
b8zhong
|
2ff0880a0e
|
[Fix] GLM 4.7 + NVFP4 + MTP (#17166)
|
2026-01-21 21:34:18 +08:00 |
|
Zhu Yuhua
|
2c1b164a92
|
[diffusion] improve: skip negative prompt encoding when guidance_scale <= 1.0 or negative_prompt is None (#16919)
Signed-off-by: zhuyuhua-v <yuhzhu@amd.com>
|
2026-01-21 20:01:55 +08:00 |
|
strgrb
|
bcc6d84f93
|
Use fused_sigmoid_gating_delta_rule_update_kernel for KDA (#17108)
|
2026-01-21 19:24:29 +08:00 |
|
Ke Bao
|
a618202fc7
|
Tiny refine swa kv cache free (#17417)
|
2026-01-21 19:20:05 +08:00 |
|
siyu
|
7520b92927
|
Support EPD error handling (#16670)
Co-authored-by: ZhengWG <zwg0606@gmail.com>
|
2026-01-21 18:47:00 +08:00 |
|
Fan Lin
|
e7224e9681
|
[diffusion] fix: fix the LoRA weights mismatch caused by weights packing (#17355)
|
2026-01-21 18:17:16 +08:00 |
|
HuangJi
|
e776239afd
|
[diffusion] feat: support SageSparseLinearAttention attention backend (#17399)
|
2026-01-21 18:13:51 +08:00 |
|
Yi Zhang
|
1b97fa769b
|
[BUGFIX] fix value oom in radix tree (#17400)
|
2026-01-21 17:12:57 +08:00 |
|
Yi Zhang
|
236772c0e1
|
[RadixTree][2/N Refactor]: swa cache init tiny refactor (#17397)
|
2026-01-21 15:48:30 +08:00 |
|
Sam Shleifer
|
0d49b13fdd
|
Fix circular import in quantization modules (#17372)
|
2026-01-21 15:47:09 +08:00 |
|
blahblah
|
0a7a2017a0
|
[diffusion] refactor: refactor and simplify teacache for cachabledit and wanvideo (#16396)
Co-authored-by: Brain97 <Brain97@users.noreply.github.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
Co-authored-by: blahblah <blahblah>
|
2026-01-21 15:42:45 +08:00 |
|
amote-i
|
0a9099e137
|
update ascend docs (#17457)
|
2026-01-21 15:36:26 +08:00 |
|
Alison Shao
|
0050c476fd
|
Add job-level timeout for weekly test workflow (#17462)
|
2026-01-20 22:45:39 -08:00 |
|
Baizhou Zhang
|
8251a74d5f
|
[Tiny] Backward compatibility for fp4 gemm flags (#17466)
|
2026-01-21 14:34:40 +08:00 |
|
Baizhou Zhang
|
a54d75bf2e
|
[Fix] Set fa3 as default MHA backend on Hopper (#17425)
|
2026-01-21 13:54:09 +08:00 |
|
Alison Shao
|
3321eb4efa
|
Fix pr-test-finish to fail when wait-for-stage jobs fail (#17465)
|
2026-01-20 20:50:27 -08:00 |
|
Kangyan-Zhou
|
be5121b452
|
Fix NSA indexer in the nightly test (#17452)
|
2026-01-20 19:58:13 -08:00 |
|
Baizhou Zhang
|
c3f9c30f99
|
[Minor] Change lora_target_modules to "all" in CI tests (#17386)
|
2026-01-21 11:46:36 +08:00 |
|
Fan Lin
|
54a821794e
|
[diffusion] fix: fix the bug of output_path not taking effect when generate (#17293)
|
2026-01-21 11:39:51 +08:00 |
|
khalilzhk
|
aca354bcb3
|
[NPU] remove features supported on Ascend NPU (#17455)
|
2026-01-21 11:00:04 +08:00 |
|
Yinghai Lu
|
aea57b33c6
|
[scheduler] Clear MM data of finished batch (#17251)
Co-authored-by: Kangyan-Zhou <zky314343421@gmail.com>
|
2026-01-20 17:45:38 -08:00 |
|
Alison Shao
|
823a046e8f
|
Add hybrid parallelism test to nightly CI (#17444)
|
2026-01-20 17:43:50 -08:00 |
|
Alison Shao
|
648aab0ce3
|
Fix wait-for-stage jobs running when call-gate fails (#17443)
|
2026-01-20 17:40:05 -08:00 |
|
ishandhanani
|
1e309030e3
|
update urllib3 and gpgv Dockerfile (#17439)
|
2026-01-20 14:47:20 -08:00 |
|
Binyao Jiang
|
6092721594
|
[Piecewise] Fix PCG issue for multimodal and embedding model that wraps language_model (#17290)
|
2026-01-20 14:06:06 -08:00 |
|
Binyao Jiang
|
38c233fd04
|
[Piecewise] Support PCG weak_ref_tensor cuda kernel on AMD (#17291)
|
2026-01-20 14:05:32 -08:00 |
|
Lianmin Zheng
|
20ed3822bb
|
[Auto Sync] Update piecewise_cuda_graph_runner.py (20260119) (#17313)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Binyao Jiang <byjiang1996@gmail.com>
|
2026-01-20 14:05:05 -08:00 |
|
zijiexia
|
4ecd9afde9
|
[Docs] Rename SGLang Router to SGLang Model Gateway (#17436)
|
2026-01-20 12:31:10 -08:00 |
|