Commit Graph

8057 Commits

Author SHA1 Message Date
Hubert Lu
51e2eaa458 [AMD] Support fast_topk kernels in sgl-kernel (#15172) 2025-12-19 22:19:09 -08:00
yctseng0211
6468cb5823 [AMD] remove the redundant projection (#15178) 2025-12-19 22:10:46 -08:00
Qiaolin Yu
e220da1723 tiny fix sampling seed for completion api (#15498) 2025-12-19 20:32:13 -08:00
Yuxuan Zhang
b82c7a0ae7 [GLM-4.7] GLM-4.7 Tool Parser and Doc Update (#15333) 2025-12-19 20:30:44 -08:00
Mick
c0f9b51992 [diffusion] refactor: deprecate workloadtype (#15267)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2025-12-20 11:24:31 +08:00
Simo Lin
5529ab5895 [model-gateway] Implement RAII load guard with response body attachment (#15507) 2025-12-19 19:14:52 -08:00
Simo Lin
74a3349bea [ci] remove rust benchmark in unit test ci (#15510) 2025-12-19 18:59:41 -08:00
Liangsheng Yin
b9ebf0ed63 Clean hidden_states_before_norm (#15485) 2025-12-20 09:38:46 +08:00
kun-llfl
3c116d5e5a [router] bugfix: cache_aware in grpc inbalance forward (#15473)
Signed-off-by: Kun(llfl) <i@imux.top>
2025-12-19 16:25:46 -08:00
Simo Lin
71a602883a [model-gateway] simplify workflow engine backoff and reduce duplicate reads (#15505) 2025-12-19 16:23:02 -08:00
Simo Lin
61405b3d00 [model-gateway] Run workflow event subscribers concurrently (#15504) 2025-12-19 15:49:51 -08:00
Simo Lin
0adfc42b99 [model-gateway] Optimize workflow engine with pre-computed dependency graph (#15503) 2025-12-19 15:44:57 -08:00
Simo Lin
ba72e759ca [model-gateway] Improve logging across core modules (#15497) 2025-12-19 13:35:28 -08:00
YL D
6afc5d497b [DSv32] Move deep_gemm.get_paged_mqa_logits_metadata to init time as metadata (#15040) 2025-12-19 13:23:33 -08:00
michael-amd
2ee6c810b8 [AMD] Add TP=8 models to nightly test and make TP=2 test stable (#15296) 2025-12-19 13:15:19 -08:00
Simo Lin
bd16244d93 [model-gateway] Improve logging in policies module (#15496) 2025-12-19 13:02:57 -08:00
Simo Lin
b5eb0214ca [model-gateway] Improve logging in data_connector module (#15495) 2025-12-19 12:57:18 -08:00
Simo Lin
d72e908ba5 [model-gateway] refactor: extract common graceful shutdown code before TLS branch (#15494) 2025-12-19 12:56:33 -08:00
Simo Lin
50cad014dc [model-gateway] fix graceful shutdown for TLS/Non-TLS server (#15491) 2025-12-19 12:35:50 -08:00
Khush Gupta
ef908aeb40 fixed trtllm nvfp4 backend for moe (#15022) 2025-12-19 10:49:21 -08:00
Yuwei An
9d0347b33a EP Support for Piecewise Cuda Graph (#14164)
Signed-off-by: Oasis-Git <ayw.sirius19@gmail.com>
2025-12-20 01:59:27 +08:00
Yashika Gandhi - Google
05eb0bcc61 [diffusion] http-server: support vertex generate pathway (#15348) 2025-12-20 01:47:17 +08:00
Liangsheng Yin
5dccd9bdcc Enhance protection rules of code owners (#15406) 2025-12-20 00:25:57 +08:00
Liangsheng Yin
933cef16cc Tiny fix mimo model conflicts with main (#15483) 2025-12-19 23:20:59 +08:00
Hexq0210
241ae17b25 [NPU] bugfix for chunkedprefill (#15166) 2025-12-19 22:41:17 +08:00
Mayyyy
2c5a446006 [NPU]Fix for ipc handle with npu (#14138) 2025-12-19 22:39:04 +08:00
Xiaoyu Zhang
f3705b0115 [diffusion] doc: add doc for attention backends (#15408)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2025-12-19 22:10:29 +08:00
shuwenn
5a0ad7310e fix: update model name after weights update (#15416) 2025-12-19 21:53:14 +08:00
Jackie
5045aa34d3 Add MiDasheng Model Support (#15219) 2025-12-19 21:41:59 +08:00
Mick
ff1e2ce24d [diffusion] fix: fix wrong validation on 2k resolution (#15478) 2025-12-19 21:28:45 +08:00
ryang
1e58248808 [diffusion] multi-platform: add Sage Attention 3 Support for sm 120 (RTX5090) (#15382)
Co-authored-by: Mengxi Li <marcyleemx@gmail.com>
2025-12-19 21:27:02 +08:00
66RING
46be74b4b4 [diffusion] kernel: timestep embedding kernel implementation (#12995)
Co-authored-by: 戚余航 <qiyuhang@bytedance.com>
Co-authored-by: Qi Yuhang <45795032+HydraQYH@users.noreply.github.com>
2025-12-19 20:59:50 +08:00
suyedu
1c65802648 [diffusion] feat: support multiple input images for generate mode (#15394)
Co-authored-by: 也渡 <fangxinlei.fxl@alibaba-inc.com>
2025-12-19 20:08:10 +08:00
Yuhao Yang
ba41080892 [diffusion] model: support qwen-image-edit-2511 (#15458) 2025-12-19 20:06:00 +08:00
HuangJi
89512029f1 [diffusion]Support url image input (#15262) 2025-12-19 19:37:23 +08:00
Lianmin Zheng
92e6b3c30e [Auto Sync] Update scheduler_runtime_checker_mixin.py (20251219) (#15437)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Hanming Lu <69857889+hanming-lu@users.noreply.github.com>
2025-12-19 01:46:03 -08:00
weiliang
6559e43f30 Support FP8 MLA prefill and 128k context. (#14395) 2025-12-19 01:24:00 -08:00
yctseng0211
af780c59f3 [AMD] add unit-test-backend-8-gpu-amd back (#15253)
Co-authored-by: Sai Enduri <saimanas.enduri@amd.com>
2025-12-19 00:59:28 -08:00
yctseng0211
a21aa87ec5 [AMD] Fix and add accuracy-test-2-gpu-amd back (#15415)
Co-authored-by: bingxche <Bingxu.Chen@amd.com>
2025-12-19 00:41:14 -08:00
shuwenn
fb17845723 fix: unreachable error check in retraction (#15433)
Co-authored-by: Liangsheng Yin <lsyincs@gmail.com>
2025-12-19 16:33:34 +08:00
Fan Yin
65c098592d [sgl-kernel] chore: update deepgemm version (#13402) 2025-12-19 00:20:24 -08:00
Yuzhen Zhou
4bf06635fc [diffusion] multi-platform: support diffusion on amd and fix encoder loading on MI325 (#13760)
Co-authored-by: Sabre Shao <sabre.shao@amd.com>
Co-authored-by: Yusheng (Ethan) Su <yushengsu.thu@gmail.com>
Co-authored-by: Hubert Lu <Hubert.Lu@amd.com>
Co-authored-by: xsun <sunxiao04@gmail.com>
2025-12-19 15:38:46 +08:00
sunxxuns
f2d64e6782 [amd] Add deterministic all-reduce kernel for AMD (ROCm) (#15340)
Co-authored-by: Thomas Wang <1am9trash@gmail.com>
2025-12-18 23:36:03 -08:00
coco
0e869f0868 [diffusion] refactor: refactor _build_req_from_sampling to use shallow_asdict (#13782) 2025-12-19 15:19:56 +08:00
Qiaolin Yu
173940927f Add customized sampler registration (#15423) 2025-12-18 23:10:23 -08:00
Lianmin Zheng
f228b662a7 Update readme (#15425) 2025-12-18 23:06:00 -08:00
b8zhong
a36142aa71 Fix Mindspore model import warning (#15287)
Co-authored-by: Even Zhou <even.y.zhou@outlook.com>
2025-12-18 23:15:33 -05:00
Yingchun Lai
160a06cab2 [Feature] Xiaomi MiMo-V2-Flash day0 support (#15207)
Co-authored-by: 谢学扬 <xiexueyang@xiaomi.com>
Co-authored-by: tz <tangzhen3@xiaomi.com>
Co-authored-by: 李家乐 <lijiale10@xiaomi.com>
Co-authored-by: 张晨 <zhangchen50@xiaomi.com>
Co-authored-by: Shaohui Liu <liushaohui3@xiaomi.com>
Co-authored-by: 王晨 <wangchen77@xiaomi.com>
Co-authored-by: jiangzihan <jiangzihan@xiaomi.com>
Co-authored-by: xiexueyang <xyxie_wangyi@163.com>
Co-authored-by: Linghao Zhang <zhanglinghao@xiaomi.com>
Co-authored-by: ispobock <ispobaoke@gmail.com>
Co-authored-by: Liangsheng Yin <lsyincs@gmail.com>
Co-authored-by: JoyFuture <35593546+JoyFuture@users.noreply.github.com>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
Co-authored-by: Qiaolin Yu <liin1211@outlook.com>
Co-authored-by: root <root@bj9-ml-g8h20e-k8s-slave106-20251106.alicn.idc.xiaomi.com>
2025-12-19 11:40:07 +08:00
Mick
a0985dd5e5 [diffusion] profiling: add bench_serving.py and VBench (#15410) 2025-12-19 10:57:39 +08:00
Zehuan Li
f6c9db4bc4 [DLLM] Fix dLLM regression (#15371) 2025-12-19 10:55:46 +08:00