Commit Graph

  • 76b3c698d6 feat: reduce constrained-decoding overhead in TP (#13947) Raayan Dhar 2026-01-09 08:38:32 -08:00
  • 5dcff94791 [nemetron/mtp] fix nemotron mtp (#16275) Hanming Lu 2026-01-09 07:22:21 -08:00
  • 8eeffbe9aa Support token low usage watermark in prefill delayer (#16814) fzyzcjy 2026-01-09 23:07:18 +08:00
  • 71e9c31cb8 Tiny add and enhance tests for prefill delayer (#16813) fzyzcjy 2026-01-09 22:58:38 +08:00
  • 7656d26737 Tiny enhance prefill delayer observability (#16812) fzyzcjy 2026-01-09 22:51:16 +08:00
  • dd24ba9084 Refactor prefill delayer for clarity and extensibility (#16811) fzyzcjy 2026-01-09 22:46:41 +08:00
  • 068abe7e40 add doc for #14386 (#14655) siyu 2026-01-09 22:38:51 +08:00
  • 64a31d4b75 [diffusion] amd: fix SGLANG_DIFFUSION_ATTENTION_BACKEND env var for diffusion attention backend selection (#16325) sunxxuns 2026-01-09 04:16:58 -08:00
  • 2babf88f24 fix spec qwen3 pd error (#16708) ybyang 2026-01-09 19:20:17 +08:00
  • d56d14e566 [Docs] Improve docs for install on gb200 (#16760) Lianmin Zheng 2026-01-09 03:07:13 -08:00
  • 0c4e155a3c chore: bump mooncake version to 0.3.8.post1 (#16792) Shangming Cai 2026-01-09 18:42:27 +08:00
  • d6d5c3fdea [AMD] Clean up vllm dependencies in moe_runner/triton.py (#11349) Hubert Lu 2026-01-09 00:24:04 -08:00
  • e46f79431b Fix external_models import path and migrate model loading tests (#16458) Alison Shao 2026-01-08 23:43:49 -08:00
  • 9d4d57dbfa [Diffusion] Tiny rename parallel_groups (#16743) Xiaoyu Zhang 2026-01-09 15:08:25 +08:00
  • 41609b52fe support page size large than 64 for mamba radix cache with fix (#16768) Hanming Lu 2026-01-08 21:28:37 -08:00
  • ccd0fb3291 [AMD] Change AITER package name (#16721) YC Tseng 2026-01-09 13:17:20 +08:00
  • 87ee6b5eaa Tiny remove auto-triggering for list-active-pr-runs. (#16778) Liangsheng Yin 2026-01-09 13:16:23 +08:00
  • f7c1d24b18 List ci workflows status. (#16774) Liangsheng Yin 2026-01-09 13:02:12 +08:00
  • cceb5e6aa5 add AWS SGLang DLC to docs (#16686) Sirut Buasai 2026-01-08 20:06:31 -08:00
  • c1c13c841a [diffusion] docs: add LoRA support (#16378) Fenglin Yu 2026-01-08 17:31:38 -10:00
  • fcec35dc4a [AMD] Add MI35x nightly CI tests (#16588) Michael 2026-01-08 19:29:15 -08:00
  • 75da784d48 Tiny simplify draft worker init. (#16446) Liangsheng Yin 2026-01-09 11:20:29 +08:00
  • 77d3566555 Tiny fix wording about CI preemption. (#16773) Liangsheng Yin 2026-01-09 11:08:17 +08:00
  • 05dfef92a1 [DeepSeek 3.2] Support and optimize pipeline parallelis when context pipeline enabled (#16380) Yongfei Xu 2026-01-09 11:01:49 +08:00
  • 7460240737 [model-gateway] release 0.3.1 (#16254) Simo Lin 2026-01-08 17:49:37 -08:00
  • f7f5c3896d [diffusion] model: wan tp+usp optimize (#16720) triple-mu 2026-01-09 10:40:31 +09:00
  • 9e3a032ad6 [smg] cleanup router RAII guards (#16560) fzyzcjy 2026-01-09 08:39:53 +08:00
  • 1bc7aa5801 [smg] update gRPC proto to match upstream changes (#16764) Simo Lin 2026-01-08 16:33:09 -08:00
  • 8726d30cb2 [smg] Add Nemotron Nano V3 reasoning parser support (#16763) Simo Lin 2026-01-08 16:29:24 -08:00
  • a1b243d76f fix: extend cpu only install dependency timeout to 20 mins (#16759) Douglas Yang 2026-01-08 16:08:21 -08:00
  • a979927727 Skip causal_conv1d test with padded batches due to Triton kernel bug (#16715) Alison Shao 2026-01-08 15:13:13 -08:00
  • ee71e773e1 [smg] Work around sglang's notorious orphan process problem (#16756) Simo Lin 2026-01-08 14:41:03 -08:00
  • 55a8dd0095 [grpc] Fix protobuf compilation in isolated build environments (#16754) Simo Lin 2026-01-08 13:09:01 -08:00
  • cf24232100 fix(e2e): prevent potential hangs in model pool subprocess handling (#16752) Simo Lin 2026-01-08 12:35:56 -08:00
  • 9d03af916a [smg][ci] delete old chat completion integration tests and workflow step (#16751) Simo Lin 2026-01-08 11:31:54 -08:00
  • 064ae3419e Adjust the cancel PR workflows for better preemption. (#16749) Liangsheng Yin 2026-01-09 03:25:41 +08:00
  • 49305fa1c1 [smg][ci] migrate function calling tests to new infrastructure (#16748) Simo Lin 2026-01-08 11:08:01 -08:00
  • 4e999404c3 [AMD] Fix CI - unit-test-backend-1-gpu-amd-mi35x and unit-test-backend-2-gpu-amd, stage-b-test-small-1-gpu-amd (#16675) YC Tseng 2026-01-09 02:16:27 +08:00
  • fbc24886dd [smg][ci] migrate validation tests to new infrastructure (#16746) Simo Lin 2026-01-08 09:26:31 -08:00
  • 16880235d1 [grpc] Auto-generate protobuf files during wheel build (#16409) Chang Su 2026-01-08 09:09:54 -08:00
  • cda35611d4 [perf] Add two stream norm for Olmo3 speedup 5% (#13681) Yi Zhong 2026-01-08 12:01:35 -05:00
  • 8a45a9c6a9 [smg][ci] fix model pool GPU cleanup and add startup reliability improvements (#16745) Simo Lin 2026-01-08 08:55:18 -08:00
  • aecd5f5f3e [smg][ci] migrate reasoning_content tests to new infrastructure (#16741) Simo Lin 2026-01-08 08:25:39 -08:00
  • 05ab110e02 [Performance] Force split_k=1 for MXFP4 Triton kernels on Hopper (#16014) Mohammad Miadh Angkad 2026-01-08 23:52:22 +08:00
  • c8dc4d2d1f [smg][ci] migrate enable_thinking tests to new infrastructure (#16739) Simo Lin 2026-01-08 06:57:45 -08:00
  • 82a8d77bc0 [Diffusion] model: fix zimage tp (#16719) cheng peng 2026-01-08 22:47:21 +08:00
  • 294ff71d18 [Diffusion] Avoid cpu2gpu sync in flashinfer rope and apply flashinfer rope to wanvideo (#16668) Xiaoyu Zhang 2026-01-08 22:44:38 +08:00
  • f52ae586b6 Remove migrated e2e_grpc/basic tests (#16738) Simo Lin 2026-01-08 06:30:45 -08:00
  • 2f8a36347c [smg][ci] migrate chat completions tests to new infrastructure and build wheel once and share via artifact (#16709) Simo Lin 2026-01-08 06:29:23 -08:00
  • b6e8a0d851 Add file size hints in bitwise model file verifier (#16735) fzyzcjy 2026-01-08 22:28:52 +08:00
  • fb7609f1dd Fix FP8 MoE NaN with DeepGEMM on Blackwell (#16622) luoyuyan 2026-01-08 22:24:12 +08:00
  • d2ea44f775 VLM: enhance VL embedding model with video input support and revise warm-up strategy (#16635) yuhao 2026-01-08 22:12:01 +08:00
  • 20ca2c6e1e [NPU] update model and features supported (#16733) Hexq0210 2026-01-08 21:50:42 +08:00
  • 83abecd0c2 Support pre-generating and using expected checksums (#16730) fzyzcjy 2026-01-08 20:26:13 +08:00
  • d54f0a10b4 Support bitwise weight checksum verifier (#16729) fzyzcjy 2026-01-08 20:19:25 +08:00
  • fb04e7e3c8 [Auto Sync] Update schedule_batch.py, common.py, eagle_info... (20260105) (#16519) Lianmin Zheng 2026-01-08 02:35:46 -08:00
  • 1e5de05e35 fix: adding timeout for install dependencies (#16706) Douglas Yang 2026-01-08 00:30:41 -08:00
  • 7dd679cbb9 [NPU][Bugfix] Fix qwen3 error when enable-dp-lm-head (#16115) chenxu214 2026-01-08 15:15:43 +08:00
  • 3d51ae18a1 [AMD] Turn on AMD CI if rocm.Dockerfile changed (#16634) YC Tseng 2026-01-08 14:51:53 +08:00
  • f9c0426692 [AMD CI] re-enable testcases missed when migrating ci test files (#16535) Bingxu Chen 2026-01-08 14:43:48 +08:00
  • 48b8dcd42e [jit kernel] support dtype as a cpp template parameter (#16452) 陈一涵 2026-01-08 13:54:33 +08:00
  • 41b434a7e6 [diffusion] endpoint: add API endpoint to query loaded LoRA adapters information (#16533) Fan Lin 2026-01-08 13:51:26 +08:00
  • 4935344fcd [AMD] Fix aiter page-size handling, DeepSeek MLA tuple inputs, and HiCache/FA3 decode-backend override (#16531) Hubert Lu 2026-01-07 21:14:32 -08:00
  • 63cc97f4ef ci: migrate 2-GPU tests to test/registered/ (#16529) Alison Shao 2026-01-07 20:28:16 -08:00
  • ab7d5829cd [AMD] Add pip install / wheel build support for ROCm sgl-kernel (#15627) Alan Kao 2026-01-08 12:18:29 +08:00
  • 261860e17b [NPU][Bugfix] move free_page logics to cpu (#16608) hw-csong 2026-01-08 12:00:18 +08:00
  • 154740bd4d Disable PCG TP Unittest (#16693) Yuwei An 2026-01-07 19:59:47 -08:00
  • 1c09cbe3ed [Build] Enable full kernel in aarch64 wheel (#16155) MarcoDWei 2026-01-08 11:40:03 +08:00
  • e14f5ec8a8 [diffusion] refactor: eliminate redundant parameters in req (#16505) Yuhao Yang 2026-01-08 11:14:03 +08:00
  • 8867d24879 Tiny adjust cancel PR workflow. (#16697) Liangsheng Yin 2026-01-08 11:13:46 +08:00
  • 6b3f93c4dd vlm: support SGLANG_MM_SKIP_COMPUTE_HASH for bypassing multimodal feature hashing (#16555) siyu 2026-01-08 11:10:00 +08:00
  • 5a5cece561 [Diffusion] clean useless and buggy set_seq_parallel_pg in yunchang (#16669) Xiaoyu Zhang 2026-01-08 11:08:34 +08:00
  • d566739b65 Add sufeng-buaa into CI_PERMISSION (#16625) Teng Ma 2026-01-08 11:01:32 +08:00
  • 4c46ecde80 [smg][ci] delete old responses api ci (#16695) Simo Lin 2026-01-07 18:24:12 -08:00
  • 12a0292bfd Revert "[sgl-kernel] Update flashmla to include fp8 sparse_mla optimizations" (#16678) hlu1 2026-01-07 18:23:06 -08:00
  • a08dc5aa10 [smg][ci] rename 3rd models from cloud backend and delete dead code (#16692) Simo Lin 2026-01-07 18:19:44 -08:00
  • eec7dbd31e remove redundant max_running_reqs calculation in r3 (#16629) Junrong Lin 2026-01-08 09:45:21 +08:00
  • 109fe03ad1 [smg][ci] Migrate Response API e2e tests to shared infrastructure (#16680) Simo Lin 2026-01-07 17:40:03 -08:00
  • bb798a1c26 [diffusion] fix: reduce default text length for Qwen-Image from 1024 to 512 (#16445) Changyi Yang 2026-01-07 17:32:21 -08:00
  • 38dc5839dd [1/n]deepseek_v2.py Refactor: attention backend handlers and forward method definition (#16306) Baizhou Zhang 2026-01-08 09:22:31 +08:00
  • 5e867f60cf [NPU] Update model and features supported (#16652) Hexq0210 2026-01-08 09:13:30 +08:00
  • 65bed8382b Add google-cloud-storage into Dockerfile (#15343) gongwei-130 2026-01-07 16:54:45 -08:00
  • 156d97b219 Fix KeyError when logprobs=false in completions endpoint (#16095) Harish 2026-01-07 15:49:02 -08:00
  • 24b30f7757 MoE Refactor: Refactor fp8.py -> flashinfer_trllm.py (#15151) b8zhong 2026-01-07 15:35:00 -08:00
  • 3a4767daa3 Fix pytest tests to exit with proper exit code (#16681) Alison Shao 2026-01-07 15:20:46 -08:00
  • 6037267f5b [smg][ci] Add thread safety to ModelPool and GPUAllocator (#16674) Simo Lin 2026-01-07 13:25:41 -08:00
  • 0241e0460f Migrate tokenizer tests to test/registered/tokenizer/ (#16457) Alison Shao 2026-01-07 13:21:16 -08:00
  • 0c474273c5 Fix gpt_oss_common import path and migrate core tests (#16426) Alison Shao 2026-01-07 12:58:32 -08:00
  • 3e73e12458 Revert "Add SwapAB Optimization for triton fused_moe_kernel on SM90." (#16676) Michael 2026-01-07 11:24:47 -08:00
  • f4742558ac fix: 8-gpu-b200 increase timeout length (#16658) Douglas Yang 2026-01-07 09:45:55 -08:00
  • 7385834c8d Add reference counting to ModelInstance for parallel test safety (#16672) Simo Lin 2026-01-07 08:28:15 -08:00
  • b5a94f8a8e [model-gateway] Fix IGW routing for external OpenAI workers (#16633) Ziwen Zhao 2026-01-07 08:15:23 -08:00
  • c356ed03dd refactor(e2e): unify RouterInstance into Gateway class, split conftest.py into modular fixtures (#16671) Simo Lin 2026-01-07 07:50:28 -08:00
  • ee4d2287ab Add SwapAB Optimization for triton fused_moe_kernel on SM90. (#15712) Insideyyy 2026-01-07 23:45:35 +08:00
  • 153c69f63d [CI] Enable dpsk v31 test on nightly H200 (#16660) Baizhou Zhang 2026-01-07 23:21:19 +08:00
  • 55b7936582 refactor(e2e_test): fix smg ci e2e test code quality (#16664) Simo Lin 2026-01-07 06:59:49 -08:00
  • 7fc12e0bfa support page size large than 64 for mamba radix cache (#16657) Yi Zhang 2026-01-07 22:52:24 +08:00
  • 8729ad5e6c fix(e2e_test): remove dead code and fix type annotations (#16661) Simo Lin 2026-01-07 06:35:15 -08:00
  • e432057381 [smg][ci] preserve model launch order with test collected (#16618) Simo Lin 2026-01-07 06:16:59 -08:00
  • 4d902c8211 [diffusion] bench: upgrade multimodal benchmarks for diverse applications and create a prettier, more intuitive logger. (#16179) Li Jinliang 2026-01-07 22:10:23 +08:00