Commit Graph

11116 Commits

Author SHA1 Message Date
Simo Lin
7385834c8d Add reference counting to ModelInstance for parallel test safety (#16672) 2026-01-07 08:28:15 -08:00
Ziwen Zhao
b5a94f8a8e [model-gateway] Fix IGW routing for external OpenAI workers (#16633) 2026-01-07 08:15:23 -08:00
Simo Lin
c356ed03dd refactor(e2e): unify RouterInstance into Gateway class, split conftest.py into modular fixtures (#16671) 2026-01-07 07:50:28 -08:00
Insideyyy
ee4d2287ab Add SwapAB Optimization for triton fused_moe_kernel on SM90. (#15712) 2026-01-07 23:45:35 +08:00
Baizhou Zhang
153c69f63d [CI] Enable dpsk v31 test on nightly H200 (#16660) 2026-01-07 23:21:19 +08:00
Simo Lin
55b7936582 refactor(e2e_test): fix smg ci e2e test code quality (#16664) 2026-01-07 06:59:49 -08:00
Yi Zhang
7fc12e0bfa support page size large than 64 for mamba radix cache (#16657)
Co-authored-by: Hanming Lu <hanming@x.ai>
2026-01-07 22:52:24 +08:00
Simo Lin
8729ad5e6c fix(e2e_test): remove dead code and fix type annotations (#16661) 2026-01-07 06:35:15 -08:00
Simo Lin
e432057381 [smg][ci] preserve model launch order with test collected (#16618) 2026-01-07 06:16:59 -08:00
Li Jinliang
4d902c8211 [diffusion] bench: upgrade multimodal benchmarks for diverse applications and create a prettier, more intuitive logger. (#16179) 2026-01-07 22:10:23 +08:00
Douglas Yang
2ff872311b ci: adding llama4 placeholder test to nightly (#16599) 2026-01-07 21:51:30 +08:00
JiLi
fd16c91cb8 Handle Marlin weight restoration and shape recording (QAT INT4 Rollout Part1) (#15238)
Co-authored-by: Gao016 <yngao016@163.com>
Co-authored-by: yefei12 <xjtu_yefeichen@163.com>
Co-authored-by: yzlnew <yzlnew@gmail.com>
Co-authored-by: Peng Zhang <aniz1905@gmail.com>
2026-01-07 21:11:24 +08:00
Xiaoyu Zhang
32a6540afc [Diffusion] Fix Ulysses/Ring process group construction under TP to enable correct Wan2.2 tensor parallelism (#16532)
Co-authored-by: Mick <mickjagger19@icloud.com>
2026-01-07 20:52:36 +08:00
Xiaoyu Zhang
62d0280f62 Tiny fix readme (#16654) 2026-01-07 20:41:07 +08:00
Even Zhou
d4b717c01e [NPU] update docs (#16651) 2026-01-07 20:20:01 +08:00
Hudson Xing
98a107d491 Re-enable temp_prefill_info assertion after pairing fix (#16203) 2026-01-07 18:05:17 +08:00
Hubert Lu
b86bbf841e [AMD] Add 8-GPU MX35X test running DSR1-MXFP4 model for AMD CI (#13602) 2026-01-07 01:43:11 -08:00
YC Tseng
48381c3b6d [AMD] suppress warning for amd (#16620) 2026-01-07 01:37:40 -08:00
YC Tseng
8bce085321 [AMD] CI - add 2 pp test cases to performance-test-2-gpu-amd (#16514) 2026-01-06 23:52:46 -08:00
Alison Shao
6b8a9d7058 fix: update AMD CI estimated time for test_torch_compile (#16631) 2026-01-06 23:42:18 -08:00
Shangming Cai
973116e6bb [Doc] Optimize pipeline parallelism doc (#16630)
Signed-off-by: Shangming Cai <csmthu@gmail.com>
2026-01-07 14:52:42 +08:00
Thomas Wang
820e97d6c9 Upgrade aiter version (#16619) 2026-01-06 22:20:42 -08:00
Minglei Zhu
4c85f9d039 Only allocate encoder metadata for encoder-decoder models (#16527) 2026-01-06 22:17:04 -08:00
Baizhou Zhang
7d757d6f17 Clean Some Environment Variables for DeepSeek V32 (#15938) 2026-01-07 14:00:16 +08:00
Douglas Yang
5c04088b3a fix: remove performance testing from nightly dpsk v32 cp single node (#16582) 2026-01-07 13:38:31 +08:00
Chi McIsaac
f066036c8b [diffusion] fix: fix ZImage SP sharding for 5D latents and unpad frames (#16418)
Signed-off-by: Chi <chixie.mcisaac@gmail.com>
2026-01-07 13:24:04 +08:00
Alison Shao
52de807dd7 Migrate profiling tests to test/registered/profiling/ (#16459) 2026-01-06 21:04:48 -08:00
Alison Shao
70933f34f1 Migrate attention unit tests to test/registered/attention/ (#16465) 2026-01-06 20:50:54 -08:00
Alison Shao
ce453fa43b Migrate backends tests to CI registration system (#16468) 2026-01-06 20:31:52 -08:00
fzyzcjy
3be1e734ee [model-gateway] extract header extraction in policy and add (#16566) 2026-01-06 20:18:44 -08:00
Alison Shao
38895a0064 ci: adjust partition counts for stage-b and unit tests (#16617) 2026-01-06 20:18:10 -08:00
Simo Lin
d8b8198192 [smg][ci]: migrate benchmarks to e2e_test/benchmarks/, use parent conftest (#16597) 2026-01-06 20:15:20 -08:00
Yingchun Lai
913b688f21 fix: fill a meaningful tool_index (#16504) 2026-01-06 20:00:09 -08:00
Douglas Yang
951d16c890 fix: adjusting vlm accuracy thresholds (#16593) 2026-01-06 19:41:05 -08:00
Yuan Luo
53846746bf [VLM] Fix CUDA IPC OOM (#16118)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
2026-01-07 11:30:35 +08:00
MOHENOO
534ac384db [HiCacheStorage & PD] fix prefill bootstrap request host memory leaks (#15439)
Co-authored-by: yangjia1 <yangjia1@kingsoft.com>
2026-01-07 11:03:57 +08:00
Yuwei An
2a8d5493f3 PCG Unit Test Adjustment (#16609)
Signed-off-by: Oasis-Git <ayw.sirius19@gmail.com>
2026-01-07 10:57:15 +08:00
Alison Shao
90eac38a12 Migrate FP8/TorchAO tests to test/registered/quant/ (#16453) 2026-01-06 18:27:43 -08:00
Mick
badcd02896 [diffusion] chore: automatically enable dit_layerwise_offload for Wan (#16499) 2026-01-07 10:22:08 +08:00
fzyzcjy
d874c8bba4 Tiny support http headers in bench serving (#16606) 2026-01-07 10:15:17 +08:00
fzyzcjy
9a21d89c5b Tiny add metrics for prefill delayer (#16603) 2026-01-07 09:53:52 +08:00
Hexq0210
4c9ac8566c [NPU] fix command in npu best practice (#16576) 2026-01-07 09:37:27 +08:00
Chang Su
fb5b71d015 [router][openai] Rename prepare_mcp_payload_for_streaming and patch_streaming_response_json (#16596) 2026-01-06 16:33:29 -08:00
Chang Su
05b54b6d7b [router][grpc] Replace Vec<(String, String, String)> with ExtractedToolCall (#16598) 2026-01-06 16:32:59 -08:00
Chang Su
4f443f445a [model-gateway][cleanup] Fix wrong comment in manager.rs (#16601) 2026-01-06 16:32:37 -08:00
Simo Lin
dce8b0606c refactor(e2e): keep only benchmark tests in e2e_http, remove redundant tests (#16594) 2026-01-06 15:05:08 -08:00
Alison Shao
399d5283f8 ci: migrate scheduler tests to test/registered/scheduler/ (#16442) 2026-01-06 14:55:52 -08:00
Alison Shao
2e0527dd74 ci: migrate Debug Utils, Ops, and Rotary Embedding tests to test/registered/ (#16422) 2026-01-06 14:35:57 -08:00
Yingchun Lai
6beb50d612 feat: add .dockerignore to ignore files when build images (#16223) 2026-01-06 14:34:16 -08:00
Kangyan-Zhou
18e2ef09d7 Add v1/models endpoint to diffusion model APIs so that they can be discovered by model gateway (#16425) 2026-01-06 14:28:24 -08:00