Commit Graph

  • d879e37f1b Add interface_v1 option for dynamic HiCache backend (#13140) pansicheng 2025-11-18 09:42:17 +08:00
  • e2c9a59023 Update pr-test.yml to fix invalid job name error Kangyan-Zhou 2025-11-17 17:08:02 -08:00
  • a1e37b0258 Temporarily comment out multimodal gen test to recover runners (#13463) Kangyan-Zhou 2025-11-17 16:55:04 -08:00
  • e389f91dec [NVIDIA] Fix broken fp8 MoE of deepseek v3 (#13264) Kaixi Hou 2025-11-17 16:13:28 -08:00
  • aac07bf7fd [Embeddings Performance Testing] Add performance test for embedding models (#12359) Vedant V Jhaveri 2025-11-17 15:35:18 -08:00
  • ea89a3a0c5 Fixes validation errors for Wan-AI models which store model weights in subdirectories (#13461) Kangyan-Zhou 2025-11-17 15:33:02 -08:00
  • 2bc7c5ebef Fix 8-gpu B200 nightly tests (#13457) Kangyan-Zhou 2025-11-17 13:26:56 -08:00
  • a63f433b6f extend sagemaker.Dockerfile serve script to allow all sglang serve flags (#13173) Sirut Buasai 2025-11-17 13:14:17 -08:00
  • 58f8f4e408 Add missing models (#13456) Kangyan-Zhou 2025-11-17 13:10:37 -08:00
  • 25acbbc612 Remove verbs from GET endpoint paths to follow REST standards (#13273) Simo Lin 2025-11-17 12:34:36 -08:00
  • a8fcbf6fe3 [Deepseek V3.2] Use torch.compile to speed up torch.cat in nsa (#13022) hlu1 2025-11-17 12:20:49 -08:00
  • e486308c99 Temporarily disable model hooks CI (#13450) Liangsheng Yin 2025-11-18 02:35:13 +08:00
  • 6042010964 [2/N] CI refactor: sperate some backend-independent CPU tasks. (#13447) Liangsheng Yin 2025-11-18 02:08:19 +08:00
  • ff00b6adbe diffusion: fix loading with local model_path (#13445) Mick 2025-11-18 00:56:30 +08:00
  • 7b44526038 [Ci tiny fix] Lower score threshold in evaluation test (#13443) Xiaoyu Zhang 2025-11-18 00:45:20 +08:00
  • c236d05f8c Fix log time stats (#13418) Rain H 2025-11-17 23:10:49 +08:00
  • df56139226 Adding user defined hooks support (#13217) Carlo Mussolini 2025-11-17 15:07:37 +00:00
  • 15db5497d3 [feature] Custom base path on FastAPI server (#5879) kebyn 2025-11-17 22:40:09 +08:00
  • 9ba3597d2a Fix cache_tokens calculate issue when retracted (#11900) Mike Qiu 2025-11-17 21:25:38 +08:00
  • 80797c2a15 [PD][bug fix] fix memleak when last_batch is none (#13144) Xuchun Shang 2025-11-17 21:22:20 +08:00
  • 7afff8fd1a diffusion: fix wan2.2 ti2v num_frames adjust logic (#13379) Mick 2025-11-17 20:53:15 +08:00
  • ac406d4301 fix uneven PP layer indices (#13282) AlphaBaby 2025-11-17 19:54:23 +08:00
  • b436113fc2 [model-gateway] Add Gateway Release Tooling (#13420) Simo Lin 2025-11-17 03:26:48 -08:00
  • 8b5e2c5368 [Tiny fix] Fix bench_speculative.py run bug (#13416) Xiaoyu Zhang 2025-11-17 18:58:19 +08:00
  • b24235b8bb [model-gateway] update workflow names for gateway and exclude npu (#13415) Simo Lin 2025-11-17 02:04:07 -08:00
  • ae7698fbd5 Remove deprecated scripts (#13399) Liangsheng Yin 2025-11-17 16:54:39 +08:00
  • a3e4fe4b41 refactor linear memory pool (#13004) Yi Zhang 2025-11-17 16:24:29 +08:00
  • f3e9336dcb Support weight update for blackwell DeepGEMM (#13324) fzyzcjy 2025-11-17 15:55:26 +08:00
  • 290fcd8971 [HiCache] add GPU id to IB dev topo for mooncake storage backend (#13112) Teng Ma 2025-11-17 15:48:29 +08:00
  • 1dcde53928 [HiCache] support memory_pool_host page head layout (#11644) huangtingwei 2025-11-17 13:45:17 +08:00
  • 15bc1f5cd7 Update .github/MAINTAINER.md (#13398) Ying Sheng 2025-11-16 21:32:24 -08:00
  • 4d59761623 [Doc] Update CI oncall list (#13396) Lianmin Zheng 2025-11-16 20:43:12 -08:00
  • ab63f3c50b [1/N] CI refactor: introduce CI register. (#13345) Liangsheng Yin 2025-11-17 12:21:20 +08:00
  • 147b782352 [RL] re-abort_request when model_update_lock is still locked (#13338) Zilin Zhu 2025-11-17 12:13:39 +08:00
  • d368c7451a (1/n)support context parallel with deepseekv3.2-DSA (#12065) lixiaolx 2025-11-17 12:12:25 +08:00
  • 7e626d12b7 Update docs (#13391) Lianmin Zheng 2025-11-16 19:36:33 -08:00
  • 2a5773440e refactor: replace worker pool with semaphore-based concurrency in jobqueue (#13383) Rivers 2025-11-17 09:57:59 +08:00
  • 7b2fb3d47c chore: bump SGLang version to 0.5.5.post3 (#13366) sglang-bot 2025-11-16 17:55:38 -08:00
  • d64dd3e18e [Tiny]Fix 1-gpu nightly test bugs (#13389) Baizhou Zhang 2025-11-16 15:54:17 -08:00
  • 3ccd7fa669 [CI] Fix B200 CI (#13387) Baizhou Zhang 2025-11-16 15:13:57 -08:00
  • 254f62d879 Support spec decoding when LoRA is applied to target model (#12903) Lifu Huang 2025-11-16 13:20:23 -08:00
  • 2b8b9d8496 [CI] use cached deepep installation in gb200 CI (#13388) Cheng Wan 2025-11-16 12:49:01 -08:00
  • e970892ffa fix import qwenvl error in RL engine (#12874) DangKai 2025-11-17 04:01:16 +08:00
  • d5fa58c4dd fix nightly docker build (#13386) b8zhong 2025-11-16 11:21:09 -08:00
  • 9f011f617f fix generative_models.md table - remove newlines (#13385) Netanel Haber 2025-11-16 20:33:38 +02:00
  • 3b18fd4cf5 [router] bindings for go (#13384) ybyang 2025-11-17 02:12:10 +08:00
  • 1869f25ca5 Tiny deprecate the --range-begin in run_suite.py (#13381) Liangsheng Yin 2025-11-16 23:15:36 +08:00
  • e019f233f9 Remove unused code / testcases in lang (#13335) Liangsheng Yin 2025-11-16 22:48:35 +08:00
  • b1c688fba2 refactor: cleanup vision attention related codes (#13228) Xinyuan Tong 2025-11-16 06:44:40 -08:00
  • 8e3663d4e8 Add default enable_memory_saver to HybridLinearKVPool (#13371) Zilin Zhu 2025-11-16 22:35:13 +08:00
  • 9edb0e0da3 Add SGLANG_ENABLE_REQ_POOL_LEAK_STRICT_CHECK to bypass mem leak check (#13339) Zilin Zhu 2025-11-16 22:02:22 +08:00
  • 95f43669b5 [2 / 2] apply sgl-kernel weak_ref_tensor (#12978) Xiaoyu Zhang 2025-11-16 21:24:35 +08:00
  • 50691d7b49 [opt kimi k2 2/n] apply kimi k2 thinking moe_fused_gate (#13332) Xiaoyu Zhang 2025-11-16 21:20:08 +08:00
  • 6afe396399 diffusion: support fa4 in fa backend for blackwell (#13263) Yuhao Yang 2025-11-16 21:02:45 +08:00
  • 191f5c7795 diffusion: correct check-changes for multimodal_gen (#13375) Mick 2025-11-16 20:32:33 +08:00
  • d724670873 [diffusion] refactor and added tests for Flux, T2V, TI2V, I2V(#13344) Adarsh Shirawalmath 2025-11-16 17:28:04 +05:30
  • efc5d8f5ed [model-gateway] fix SDist step readme path (#13373) Simo Lin 2025-11-16 03:16:52 -08:00
  • ef32a252d0 [model-gateway] fix model gateway pypi release workflow path (#13372) Simo Lin 2025-11-16 02:33:49 -08:00
  • 9509c4ccd3 [RL] enable offloading hybrid linear attn model (#13336) Zilin Zhu 2025-11-16 17:12:52 +08:00
  • 4e19c1d541 Add missing models (#13369) Kangyan-Zhou 2025-11-16 00:04:47 -08:00
  • 78a4b446c6 Fix dpsk-r1-fp4 tp8 by reverting two commits (#13162 and #13341) (#13348) Qiaolin Yu 2025-11-15 21:31:36 -08:00
  • f969664172 [Performance] Move the contiguous to torch compile region (#13199) DarkSharpness 2025-11-16 12:49:52 +08:00
  • f35f7f1245 [Piecewise CUDA Graph] Support W4A8 (#13179) b8zhong 2025-11-15 19:53:50 -08:00
  • 12c789ebd1 [Ascend]support xgrammar backend for ascend npu (#12310) ash-sigh 2025-11-16 11:39:24 +08:00
  • 597d416070 [feature] Add layerwise NVTX support (#11870) kyleliang-nv 2025-11-15 19:20:56 -08:00
  • 1ca205f6da chore: bump sgl-kernel version to 0.3.17.post1 (#13358) sglang-bot 2025-11-16 11:11:41 +08:00
  • 24a25ffa20 [Piecewise CUDA Graph] Support ModelOpt FP4 (#13101) b8zhong 2025-11-15 19:03:19 -08:00
  • db7299aa30 [Fix] Register custom ops only if they exist (#13321) Lianmin Zheng 2025-11-15 17:48:17 -08:00
  • 13366843a4 [CI] check unit-test-backend-8-gpu-h20 in workflow (#13355) Cheng Wan 2025-11-15 16:44:10 -08:00
  • be353ffd13 Add missing models (#13351) Kangyan-Zhou 2025-11-15 13:51:20 -08:00
  • 20e59f9510 Add FP32 dtype support for RoPE - Part1 (#13181) iLeGend 2025-11-16 03:37:18 +08:00
  • 0d116b9a0b Clean up deprecated tile_tokens_dim for next flashinfer (#13341) Vincent Zhong 2025-11-15 11:50:15 -05:00
  • 4a56fa5cf2 chore: bump sgl-kernel version to 0.3.17.post1 (#13325) sglang-bot 2025-11-15 23:36:32 +08:00
  • 51f9b9628b [optimize] Provide Usrbio compilation and installation commands (#12329) FlyPanda 2025-11-15 23:06:24 +08:00
  • daf494b681 tiny fix lint (#13337) Liangsheng Yin 2025-11-15 23:05:08 +08:00
  • 37e8724ef3 perf: optimize TypeBasedDispatcher using dict for O(1) lookup (#12001) xlzheng 2025-11-15 22:02:42 +08:00
  • 10592e9c08 [Ascend][Feat] Add Ascend sampling backend (#12692) Zhihao Lyu 2025-11-15 21:59:58 +08:00
  • 2aec8b6e1b [Feature] Spec-Overlap supporting DP-ATTN; PD-Disaggregation; npugraph mode (#12443) Even Zhou 2025-11-15 21:51:07 +08:00
  • 0d41ddfbd0 Temporarily disable test_vision_openai_server_a CI (#13331) Ke Bao 2025-11-15 19:46:49 +08:00
  • 37c8761567 [model-gateway] remove grpc feature flag and mark as default (#13330) Simo Lin 2025-11-15 03:16:57 -08:00
  • 3f400f25df [Diffusion] add health endpoints to diffusion server (#13329) Adarsh Shirawalmath 2025-11-15 16:42:37 +05:30
  • d91b16eb16 Opt tp: tp attn support tp reduce scattered input (#10568) Yongfei Xu 2025-11-15 18:08:12 +08:00
  • 4a10e37ba7 [router] Fix flaky router e2e tests (#13306) Xinyue Zhang 2025-11-15 02:00:25 -08:00
  • b051d76dab Fix: add missing get_embed_and_head in MiniMax M2 for Eagle3 (#13297) Charles Chen 2025-11-15 01:36:03 -08:00
  • d52d992a90 Update README (#13326) Yineng Zhang 2025-11-15 01:33:10 -08:00
  • d971f22898 Super tiny expose transform_scale_ue8m0 API for RL frameworks (#13323) fzyzcjy 2025-11-15 17:31:04 +08:00
  • 1d3d42bda0 [opt kimi k2 1 / n] Add kimi k2 moe fused gate (#13287) Xiaoyu Zhang 2025-11-15 17:14:19 +08:00
  • 8e9f05ece1 Update marlin moe kernel interface (#13322) Ke Bao 2025-11-15 17:10:39 +08:00
  • bc083521a8 [RL] support update_weights_from_tensor for mtp (#7415) Zilin Zhu 2025-11-15 16:57:01 +08:00
  • 33f08a98b0 Tiny refactor condition to requant scale ue8m0 (#13286) fzyzcjy 2025-11-15 16:36:00 +08:00
  • 8e6083bfcf Support inverse transform ue8m0 scale (#13285) fzyzcjy 2025-11-15 16:34:32 +08:00
  • 2fbc78a083 Support fast gemm when in batch invariant DeepGEMM fallback (#13259) fzyzcjy 2025-11-15 16:34:15 +08:00
  • f0b5ccf5f5 [RL] Allow bypassing /health check (#13320) Zilin Zhu 2025-11-15 16:33:54 +08:00
  • b732ffa404 [model-gateway] move python to binding folder (#13295) Simo Lin 2025-11-15 00:32:21 -08:00
  • 56fc483073 [RL] support only do cpu backup on draft model (#13318) Zilin Zhu 2025-11-15 16:31:57 +08:00
  • 2a96e302cb Revert moe sum reduce for marlin moe (#13314) Ke Bao 2025-11-15 15:57:41 +08:00
  • 8a4373405e re-submit 12911 but relax the requirement for deepgemm (#13226) Minglei Zhu 2025-11-14 23:37:12 -08:00
  • f0021c0dc8 Add feature flag for mm inputs processing optimization (#13278) Yuan Luo 2025-11-15 14:43:44 +08:00
  • 7ee3e36412 Fix: test_vlm_offline_throughput output throughput (#13279) Douglas Yang 2025-11-14 22:31:48 -08:00
  • 6d5e16fb1c feat: Add FP4 (E2M1) KV Cache Support for MHA (#12612) Ho-Ren (Jack) Chuang 2025-11-14 22:31:35 -08:00