Arthur Cheng
|
f65fa04748
|
[model-gateway]Enable IGW mode with gRPC router and auto enable IGW when service discovery is turned on (#15459)
|
2025-12-24 00:27:58 -08:00 |
|
Lee Nau
|
7e027691c8
|
update benchmark README to use --fp8-gemm-backend instead of env var (#15689)
|
2025-12-23 23:23:31 -08:00 |
|
Liangsheng Yin
|
cb719c74ad
|
[2/N] clean duplicate code of logprob processing in spec. (#15593)
|
2025-12-24 15:21:13 +08:00 |
|
Lianmin Zheng
|
ff903a7eea
|
Simplify server args (#15704)
|
2025-12-23 22:46:12 -08:00 |
|
Prozac614
|
eee3700d84
|
[diffusion] feat: support lora strength (#15691)
|
2025-12-24 13:31:23 +08:00 |
|
Yuwei An
|
e254cdf326
|
[CI] Remove pcg-omni-ci (#15656)
|
2025-12-24 13:28:05 +08:00 |
|
Mick
|
dfb5357448
|
[diffusion] http-server: relax openai image endpoint's strict content_type limit (#15717)
|
2025-12-24 13:26:13 +08:00 |
|
Simo Lin
|
9665574937
|
[docs] major SGL Model Gateway documentation update (#15715)
|
2025-12-23 20:26:09 -08:00 |
|
vincentzed
|
ac320a6f04
|
Move some quant args to its own section in environ variables doc (#15722)
Signed-off-by: vincentzed <207368749+vincentzed@users.noreply.github.com>
|
2025-12-23 20:11:08 -08:00 |
|
Liangsheng Yin
|
6292c24437
|
Tiny fix test eagle infer b. (#15716)
|
2025-12-24 11:57:15 +08:00 |
|
Liangsheng Yin
|
5f5a567768
|
Tiny add flush in the suite partition status print. (#15719)
|
2025-12-24 11:55:59 +08:00 |
|
Xuchun Shang
|
3bf07c684f
|
[Feature][MM] split the images of one request into multiparts (#11828)
Signed-off-by: Xuchun Shang <xuchun.shang@linux.alibaba.com>
Signed-off-by: Kun(llfl) <i@imux.top>
Signed-off-by: Xuchun Shang <xuchun.shang@gmail.com>
Co-authored-by: Teng Ma <sima.mt@alibaba-inc.com>
Co-authored-by: Kun(llfl) <llfl@linux.alibaba.com>
Co-authored-by: Kun(llfl) <i@imux.top>
Co-authored-by: liusy58 <liusy58@linux.alibaba.com>
|
2025-12-24 11:22:04 +08:00 |
|
michael-amd
|
e7b09efc0a
|
[AMD] Add AMD Nightly Performance & VLMs Accuracy Tests (#15500)
|
2025-12-23 19:03:27 -08:00 |
|
fzyzcjy
|
99d3bcdfed
|
[model-gateway] add back router worker health metric and fix init state (#15622)
|
2025-12-23 18:50:38 -08:00 |
|
fzyzcjy
|
4d64f15086
|
[mode;-gateway] add back fixes of incorrect metrics after worker removal (#15624)
|
2025-12-23 18:50:03 -08:00 |
|
Qiaolin Yu
|
aef7ca7cf2
|
Raise the accept length bar in dpsk-r1-fp4 spec decoding tests (#15705)
|
2025-12-23 18:38:02 -08:00 |
|
Xuchun Shang
|
fe712aa3df
|
[bug fix] fix hicache jit kernel (#15177)
Signed-off-by: Xuchun Shang <xuchun.shang@gmail.com>
|
2025-12-24 10:31:21 +08:00 |
|
Simo Lin
|
846953d9f1
|
[model-gateway] Add tokenize/detokenize HTTP endpoints and tokenizer management (#15702)
|
2025-12-23 17:32:07 -08:00 |
|
Roger Young
|
5c64a20da7
|
Update MiniMax-M2 ToolCall and add MiniMax-M2.1 in Docs (#15538)
Co-authored-by: xuebi <xuebi@minimaxi.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
|
2025-12-23 15:11:52 -08:00 |
|
Lianmin Zheng
|
cf81737693
|
Add kv_transfer_total_mb to metrics (#15667)
|
2025-12-23 11:16:55 -08:00 |
|
Simo Lin
|
aa6ac96635
|
[model-gateway] Fix tokenizer caching and improve error handling (#15695)
|
2025-12-23 10:33:38 -08:00 |
|
Liangsheng Yin
|
0d0367e9d0
|
Tiny fix CI (#15696)
|
2025-12-24 01:51:08 +08:00 |
|
Liangsheng Yin
|
bd572360f3
|
Tiny apply gsm8k mixin to ngram test (#15606)
|
2025-12-24 01:30:26 +08:00 |
|
ratish
|
5f3a47d8a7
|
[model-gateway]: add gRPC router embeddings endpoint implementation (#15273)
|
2025-12-23 09:12:22 -08:00 |
|
Praneth Paruchuri
|
80ae2229d3
|
[model-gateway] Optimize router selection with lock-free snapshots (#15672)
|
2025-12-23 09:09:30 -08:00 |
|
Liangsheng Yin
|
705287b2e5
|
Tiny add more information in retract logging. (#15694)
|
2025-12-24 00:43:36 +08:00 |
|
Yubo Wang
|
762846531f
|
Fix Illegal Memory Access when fa3 + spec + topk + page_size > 1 (#15469)
Co-authored-by: Liangsheng Yin <lsyincs@gmail.com>
|
2025-12-24 00:13:57 +08:00 |
|
Arthur Cheng
|
dd620987d1
|
[model-gateway] Replace tokenizer with tokenizer registry for dynamic tokenizer loading in gRPC router (#12968)
|
2025-12-23 07:58:47 -08:00 |
|
shuwenn
|
53f974b973
|
fix: potential crash for missing stream attribute (#15644)
|
2025-12-23 23:24:31 +08:00 |
|
fzyzcjy
|
6a5764a719
|
Super tiny add test_soft_watchdog to nightly (#15692)
|
2025-12-23 22:46:15 +08:00 |
|
DarkSharpness
|
291f11ae39
|
[Minor] Enhance JIT kernel and add dev docs (#14570)
|
2025-12-23 22:34:59 +08:00 |
|
Teng Ma
|
d7301c89ba
|
[Feature] support fastsafetensors (#15091)
Signed-off-by: Xuchun Shang <xuchun.shang@gmail.com>
Co-authored-by: Xuchun Shang <xuchun.shang@gmail.com>
|
2025-12-23 22:33:56 +08:00 |
|
shuwenn
|
758b9067a0
|
[CI] fix UT assert error in test_tokenizer_manager.py (#15646)
|
2025-12-23 22:24:33 +08:00 |
|
ryang
|
ffc23ef877
|
[diffusion] feat: generalize layerwise offloader to flux1 (#15633)
|
2025-12-23 22:22:40 +08:00 |
|
Shangming Cai
|
b3f83cc1c5
|
Fix pipeline parallelism doc typos (#15688)
Signed-off-by: Shangming Cai <csmthu@gmail.com>
|
2025-12-23 22:19:15 +08:00 |
|
Shangming Cai
|
c15fa1c577
|
[2/N] Update doc of Pipeline Parallelism with case study (#15684)
Signed-off-by: Shangming Cai <csmthu@gmail.com>
|
2025-12-23 20:56:15 +08:00 |
|
Tianyu Guo
|
fa2966983a
|
Support PP for zmq_to_scheduler (#15312)
|
2025-12-23 17:07:55 +08:00 |
|
Douglas Yang
|
f9dd90ac35
|
fix: increasing H200 test timeout (#15600)
|
2025-12-23 01:00:37 -08:00 |
|
Bingxu Chen
|
66902e0f1b
|
[AMD] CI - Improve image discovery with remote registry fallback (#15463)
|
2025-12-22 23:04:58 -08:00 |
|
yctseng0211
|
e50f356f6f
|
[AMD] CI - Detect the aiter version and rebuild if needed (#15460)
|
2025-12-22 22:42:48 -08:00 |
|
Alison Shao
|
ac42797cf7
|
[CI] Enable retry logic for flaky CI tests (#14983)
|
2025-12-22 22:23:42 -08:00 |
|
Alison Shao
|
883747ced1
|
[CI] Migrate Attention Backend tests to test/registered/attention/ (#15563)
|
2025-12-22 22:17:52 -08:00 |
|
Alison Shao
|
989d4b3012
|
[CI] Migrate nightly tests to test/registered/ (#15582)
|
2025-12-22 22:16:32 -08:00 |
|
Baidu-AIAK
|
bc3ca30023
|
[PD] Support fake decode for PD disaggregation without prefill node (#14628)
Co-authored-by: sunhailiang <sunhailiang@baidu.com>
|
2025-12-23 12:43:33 +08:00 |
|
yuchengz816-bot
|
061f41affc
|
MoE: Skip SiLU/GELU activation for masked experts (#15539)
Co-authored-by: Runkai Tao <rt572@physics.rutgers.edu>
|
2025-12-22 19:08:31 -08:00 |
|
Yuxuan Zhang
|
82f1d6157f
|
[GLM-ASR] GLM-ASR Support (#15570)
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
|
2025-12-22 17:37:55 -08:00 |
|
Lianmin Zheng
|
5e1a495c65
|
Improve engine customization interface (#15635)
|
2025-12-22 14:24:16 -08:00 |
|
sglang-bot
|
34013d9d5a
|
chore: bump sgl-kernel version to 0.3.20 (#15590)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
|
2025-12-22 12:32:34 -08:00 |
|
SeanWeiSean
|
7759716786
|
bugfix[schedule]: Refactor sort method and add related UT (#13576)
Co-authored-by: Yuxuan Wei <w1300012920@pku.edu.cn>
Co-authored-by: Yuxuan Wei <w1300012920@gmail.com>
Co-authored-by: Yuxuan Wei🚚 <yuxwei@microsoft.com>
|
2025-12-23 03:13:57 +08:00 |
|
Liangsheng Yin
|
3c882db3ad
|
Adjust wrong mtp meaning introduce by mimo (#15632)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2025-12-23 02:06:46 +08:00 |
|