Commit Graph

8491 Commits

Author SHA1 Message Date
Hudson Xing
f4ab2ec5be Add unified metrics collection framework (v1) (#16064)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-01-03 16:30:38 -08:00
fy
25fa2ac290 Convert cu_seqlens to CPU for npu_flash_attention_unpad operator (#15434) 2026-01-04 08:16:42 +08:00
fzyzcjy
ef5ac6f01e Support shared prefix (gsp) in schedule simulator (#16353) 2026-01-04 08:03:16 +08:00
fzyzcjy
c88aaf22c7 Support in-flight request age metrics for router (#16341) 2026-01-04 07:38:53 +08:00
fzyzcjy
e139d2aa76 Super tiny fix CI (#16351) 2026-01-04 07:24:46 +08:00
fzyzcjy
2337b1bbb0 Refactor and fix prefill delayer (scheduler enhancer) (#16269) 2026-01-04 07:15:06 +08:00
Chang Su
7bc13c906c [grpc] update api to scheduler in grpc request manager (#16350) 2026-01-03 10:54:31 -08:00
fzyzcjy
877c8e3a96 Tiny refactor router test contexts (#16340) 2026-01-03 09:44:10 -08:00
fzyzcjy
66dfb8c156 Tiny fix non-PD router http header missing whitelist (#16339)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-01-03 09:43:02 -08:00
Chi McIsaac
dcacc492d0 [diffusion] fix: fix RuntimeError in SageAttention3 on Blackwell with Qwen-Image (#16335)
Co-authored-by: qimcis <qimcis@users.noreply.github.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
2026-01-03 22:12:53 +08:00
fzyzcjy
87ef05e2e1 Support simple schedule simulator (#16344) 2026-01-03 22:05:40 +08:00
fzyzcjy
b65c9889a0 Fix incorrect running batch size in prefill stats (#15941) 2026-01-03 20:49:54 +08:00
fzyzcjy
c7c0d97fc6 Tiny support whitelisted headers in request logging (#16342) 2026-01-03 20:10:58 +08:00
shengzhaotian
6bc5a52fd2 [NPU] Adapt qwen3-next W8A8 on NPU (#16164) 2026-01-03 19:41:20 +08:00
Mick
2c09de343e [diffusion] improve: skip loading vision module for text encoders (#16304) 2026-01-03 19:30:45 +08:00
Mick
38d48de93d [diffusion] CI: simplify warmup (#16303)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-01-03 18:09:38 +08:00
fzyzcjy
7f2fa2167b Tiny add --log-requests-target (#16338) 2026-01-03 17:28:27 +08:00
Xiaoyu Zhang
d0fb24ee7b [Diffusion] Flux2 tp support (#16219)
Co-authored-by: Mick <mickjagger19@icloud.com>
2026-01-03 17:02:27 +08:00
Simo Lin
65b0b5b248 fix(logging): use Display format for model_id instead of Debug (#16337) 2026-01-03 00:21:39 -08:00
Siyuan Chen
9a414b164c [Performance] Optimze the performance of Qwen25VL (#15640)
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
2026-01-02 23:15:36 -08:00
Alison Shao
5b4f790200 [diffusion] CI: add CI validation for diffusion model downloads (#16311) 2026-01-03 14:12:45 +08:00
Chang Su
d8ac5eecf7 [model-gateway] bug fix on module name (#16332) 2026-01-02 21:56:18 -08:00
Teng Ma
bdde949619 [HiCache] Add PP Support with suffix pp rank (#15175)
Co-authored-by: Xuchun Shang <xuchun.shang@gmail.com>
Co-authored-by: ybyang <10629930+whybeyoung@users.noreply.github.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>
2026-01-03 13:49:21 +08:00
Chang Su
b23e7ed13c refactor(gateway): shorten logging targets from sgl_model_gateway to smg (#16328) 2026-01-02 21:12:33 -08:00
fzyzcjy
9821fae5c6 Tiny add explanations to realtime token metric labels (#16009) 2026-01-03 12:36:55 +08:00
siyu
078d96213a [FEAT] optimize tensor zmq transfer for multimodal inputs (#13592)
Co-authored-by: Shangming Cai <csmthu@gmail.com>
Co-authored-by: Yuhao Yang <47235274+yhyang201@users.noreply.github.com>
2026-01-03 12:23:23 +08:00
Douglas Yang
8b111b20c3 feature: improvements to CI failure monitor (#16272) 2026-01-02 20:09:41 -08:00
b8zhong
74a166cb86 [Fix] Only add SM90 and SM100 to check for auto-enabling TRT Allreduce Fusion (#16283) 2026-01-03 11:43:17 +08:00
triple-mu
888e126ac9 [diffusion] comment: fix typo (#16257)
Co-authored-by: Mick <mickjagger19@icloud.com>
2026-01-03 11:42:04 +08:00
Baizhou Zhang
bb23a8fe77 [Tiny]Remove progress bar for fp8 ue8m0 quant when unneeded (#16177) 2026-01-03 10:54:12 +08:00
sunxxuns
8b869e326c [AMD] feat: add DLLM support for AMD GPUs with LLaDA2 testing (#15560) 2026-01-03 10:41:11 +08:00
Alison Shao
6256936d09 Fix: Allow build-wheels to run when target_stage is set (#16324) 2026-01-02 17:58:11 -08:00
Lianmin Zheng
62f73a8c81 [Auto Sync] Update engine.py (20260102) (#16317)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-01-02 17:29:40 -08:00
Ke Bao
7c1b4b1c4c Support swa allocator page size > 1 (#16296) 2026-01-03 08:49:47 +08:00
Simo Lin
31ed68e7c1 [model-gateway] address embedding similarity threshold in ut (#16321) 2026-01-02 16:46:55 -08:00
Chang Su
24c91001cf [model-gateway] code clean up in tokenizer register step workflow (#16316) 2026-01-02 16:05:10 -08:00
Chang Su
c4edcac6d7 refactor(core): remove get_by_model_fast alias in worker_registry (#16313) 2026-01-02 16:04:23 -08:00
Chang Su
e93433892b refactor(steps): consolidate duplicate strip_protocol function (#16318) 2026-01-02 16:02:42 -08:00
Chang Su
a2d4f58a96 fix(http): use 504 Gateway Timeout for upstream timeouts (#16320) 2026-01-02 16:01:44 -08:00
Chang Su
f66b091699 refactor(http): improve pd_types.rs documentation and style (#16319) 2026-01-02 16:01:11 -08:00
Chang Su
2f623368ab refactor(core): use idiomatic .min() in calculate_delay (#16314) 2026-01-02 15:46:13 -08:00
Chang Su
bd9a2ced47 refactor(core): use thiserror for WorkerError (#16315) 2026-01-02 15:33:07 -08:00
Chang Su
0ca417d9ae chore(core): remove unused serde_json import in worker.rs (#16312) 2026-01-02 15:31:41 -08:00
sunxxuns
30cfb687fa fixed amd multimodal CI failures caused by refactor in #15812 #15813 (#16287) 2026-01-02 14:55:42 -08:00
Chenxi Li
b7c7e03d93 Fix crash dump replay script for image data replay (#16277) 2026-01-02 13:42:22 -08:00
Alison Shao
f0195627a9 Fix sgl-kernel jobs to skip when target_stage is specified (#16308) 2026-01-02 12:49:26 -08:00
Simo Lin
d401d23876 [model-gateway] Add embedding correctness test comparing against HuggingFace (#16092)
Co-authored-by: Chang Su <chang.s.su@oracle.com>
2026-01-02 10:58:20 -08:00
Alison Shao
17041f4673 Enable /rerun-stage slash command to work on fork PRs (#16128) 2026-01-02 10:36:12 -08:00
Baizhou Zhang
f07e76b229 Multiple refactors of DeepSeek V32 and context parallel (#16305) 2026-01-03 02:21:22 +08:00
Xiaoyu Zhang
1cfd2b2ded [diffusion] chore: remove redundant ulysses nccl warmup (#16301) 2026-01-03 00:19:35 +08:00