Commit Graph

  • 99cb2ed988 [refactor] Move trtllm_fp8_kv_kernel to triton_ops directory (#15044) Hudson Xing 2025-12-14 15:22:37 -08:00
  • 0612175ce1 [model-gateway] add streaming metrics (TTFT, TPOT, tokens, duration) for gRPC router (#15125) Simo Lin 2025-12-14 14:39:10 -08:00
  • 8c96fcda70 feature: ci failure monitor slack bot (#15110) Douglas Yang 2025-12-14 13:47:15 -08:00
  • ea7c69ce28 [hotfix]: Add missing args for 3FS bench_client.py (#14791) zhangheng 2025-12-15 05:13:19 +08:00
  • 8102e36b5d Add NanoV3 reasoning parser support (#15113) danielafrimi 2025-12-14 22:09:44 +02:00
  • 3f0482174a Fix Mamba2-based models' default attention backend (#15117) roikoren755 2025-12-14 21:08:48 +02:00
  • b11af135eb [model-gateway] feat(metrics): implement Layer 2 router metrics (smg_router_*) (#15124) Simo Lin 2025-12-14 10:34:12 -08:00
  • f9bceea064 [model-gateway] Implement Layer 1 HTTP metrics instrumentation (#15121) Simo Lin 2025-12-14 09:39:15 -08:00
  • 997ea57eaf Fix tensor mismatch error in sepc + topk > 1 + page_size > 1 (#14874) Ziming Huang 2025-12-15 01:10:58 +08:00
  • 47633c191f [model-gateway] Add new SMG metrics architecture with 6 layers (#15106) Simo Lin 2025-12-14 08:37:19 -08:00
  • 64b5c3ab90 [diffusion] refactor: refactor fuse qkv with QKVParallelLinear linear (#15090) Xiaoyu Zhang 2025-12-15 00:33:29 +08:00
  • d277a86dea [CI] Add disaggregation decode PP test (#15114) Shangming Cai 2025-12-14 23:54:25 +08:00
  • 9acb21ae27 feat: support EPD disaggregation (#12263) Tianyu Guo 2025-12-14 22:30:08 +08:00
  • a9ce1623cd [kernel][moe] add moe topk fast (#13969) zyl_keep_moving 2025-12-14 22:26:40 +08:00
  • 6f0c77d7f8 [diffusion] app: support webui (#14961) Li Jinliang 2025-12-14 21:07:36 +08:00
  • e3f51e823e [diffusion] feat: add support for additional sampling parameters in video generation API (#15062) Xiaoyu Zhang 2025-12-14 19:44:03 +08:00
  • fdfabb7afc [diffusion] fix: tiny fix _templated_ring_attention bug (#15053) Xiaoyu Zhang 2025-12-14 19:41:53 +08:00
  • 19c16748ce [diffusion] feat: support resolution check for video model (#14881) blahblah 2025-12-14 17:50:13 +08:00
  • 5c75907e62 Introduce native kv cache move (#15108) Liangsheng Yin 2025-12-14 17:23:28 +08:00
  • 4ea3642250 [NPU] bug fix for multi stream (#15048) liupeng374 2025-12-14 16:38:12 +08:00
  • f50af32d8b [scheduler] remove scheduler allgather for best throughout (#14294) liupeng374 2025-12-14 16:37:49 +08:00
  • 729529190d [ci] Move dpsk-r1-fp4 b200 test to stage b (#15084) Qiaolin Yu 2025-12-13 23:14:08 -08:00
  • 2ae5bed193 [NPU][CI] change de trigger of release image workflow (#14969) monkeyLoveding 2025-12-14 15:11:47 +08:00
  • 54df514bac Avoid confusing zero value metric when worker is removed (#15096) fzyzcjy 2025-12-14 13:57:20 +08:00
  • 681c68cfbf Fix issue not reported when load decrement is incorrect (#15061) fzyzcjy 2025-12-14 13:17:35 +08:00
  • 74ea45cc4b [model-gateway] optimize metric labels to avoid unnecessary allocations (#15095) Simo Lin 2025-12-13 21:10:45 -08:00
  • 0fa044ad63 [model-gateway] Add circuit breaker and discovery watcher metrics (#15094) Simo Lin 2025-12-13 20:47:04 -08:00
  • 7d8e42c979 [model-gateway] Fix metric emission gaps and name mismatch (#15093) Simo Lin 2025-12-13 20:16:30 -08:00
  • 6abdf73f4d [Fix] Environment variable SGL_* is deprecated (#14943) Huang Lin 2025-12-14 11:55:43 +08:00
  • 20ce9938b5 [model-gateway] Remove unused TokenizerMetrics to reduce CPU overhead (#15087) Simo Lin 2025-12-13 18:54:18 -08:00
  • 168a31eb00 Support prefill max requests limitation (#14993) fzyzcjy 2025-12-14 10:52:10 +08:00
  • 69cfb17b9a [Fix] avoid stream sync in _compute_mrope_positions (#14956) narutolhy 2025-12-13 18:50:57 -08:00
  • a7a4b1755d [Doc][TPU]add sglang-jax tpu docs (#15056) Brian 2025-12-14 09:29:41 +08:00
  • fdc93b019c Add sglang:decode_sum_seq_lens metric (#15066) fzyzcjy 2025-12-14 08:42:45 +08:00
  • fd37cc5d38 [model-gateway] Refactor worker steps and add update workflow (#15085) Simo Lin 2025-12-13 15:00:16 -08:00
  • 96705514bd fix: dpskv32 chat history processing, default drop_thinking to true (#15064) Xinyuan Tong 2025-12-13 20:16:04 +00:00
  • ab3ffd1c8e Add nightly accuracy test for DeepSeek V3.2 (#14935) Baizhou Zhang 2025-12-13 12:11:16 -08:00
  • 2285affffa [model-gateway] Avoid MCP Server Initialization Issue (#15065) Wenyi Xu 2025-12-14 03:20:16 +08:00
  • a81cc1b8b3 add transformers version validation for glm-4.6v moe models (#14998) Yuhao Yang 2025-12-14 02:54:08 +08:00
  • 3134d2b2a7 [scheduler] enhance scheduler in dp_attention mixed case with spec (#14201) liupeng374 2025-12-14 02:52:26 +08:00
  • 06b58c5dc5 fix flaky image access in ci by switching to raw content url (#14940) Yuhao Yang 2025-12-14 02:52:06 +08:00
  • ea07a283b2 fix: adding schedule for nightly wheel (#15054) Douglas Yang 2025-12-13 10:48:21 -08:00
  • ea91a720d5 feature: ci failure monitor improvements (#15055) Douglas Yang 2025-12-13 10:47:52 -08:00
  • 5d9c6bac07 [bug] fix grpc secheduler launcher breaking change (#15080) Simo Lin 2025-12-13 10:36:49 -08:00
  • e048ee90fc [model-gateway] Simplify error response creation (#15079) Simo Lin 2025-12-13 10:28:24 -08:00
  • ed52d01b0b Fix spec info's filter when reqs are finished right after prefill (#14742) Liangsheng Yin 2025-12-14 00:32:54 +08:00
  • 90e7d4f78f Tiny adjust CI run suite (#15074) Liangsheng Yin 2025-12-14 00:05:19 +08:00
  • c20d43d2e6 [diffusion] doc: update profiling.md with output location details (#15072) Mick 2025-12-13 23:15:23 +08:00
  • d977dd2e07 Fix IMA with flashinfer + spec + topk & Add radix attention test cases for eagle (#13740) Liangsheng Yin 2025-12-13 23:04:41 +08:00
  • 0c23331e2e [diffusion] doc: add multimodal-gen profiling doc (#15069) Xiaoyu Zhang 2025-12-13 22:26:20 +08:00
  • 993278b488 Fix double decrease load (#15060) fzyzcjy 2025-12-13 22:12:35 +08:00
  • 9e9d910744 Fix load metric not updated when using guard (#15059) fzyzcjy 2025-12-13 22:11:52 +08:00
  • 3b8a824b8b [VLM] Support VLM ViT Piecewise CUDA Graph (#14422) Yuan Luo 2025-12-13 20:49:40 +08:00
  • f1bbd26ff7 Clean up GDN Init (#14855) Stefan He 2025-12-13 02:56:54 -08:00
  • d36299ad77 [NPU] perf update with kvcache nz & w4a8 quant (#14423) liupeng374 2025-12-13 17:39:55 +08:00
  • 0e7d7969d5 [PP Prefill][NIXL] Fix PP mode transfer completion tracking to wait for all ranks (#15027) YAMY 2025-12-13 00:55:28 -08:00
  • 80554598d3 Fix GLM-4.6 tool calls don't support streaming output for arguments i… (#13989) cynial 2025-12-13 16:37:18 +08:00
  • b2e240bc4a feature: adding nightly wheel workflow and indexer (#14924) Douglas Yang 2025-12-13 00:26:10 -08:00
  • dcc5f5c0da [diffusion] feat: Improve LoRA compatibility by adding unified format detection and diffusers-based normalization (#14659) Fenglin Yu 2025-12-13 03:04:36 -05:00
  • 4eda4194f2 [Fix] Disable trtllm moe backend for draft model for a qucik fix (#15002) Sam 2025-12-13 15:58:47 +08:00
  • 875f84db7b [diffusion] fix: use NDRotaryEmbedding in flux_2 (#15034) Mick 2025-12-13 13:42:38 +08:00
  • f6031adf08 Mistral Large 3 NVFP4 support (#14485) Daniel Cámpora 2025-12-13 06:34:42 +01:00
  • 2a39cfe0ff call check_quantized_moe_compatibility after initialize (#13876) Chunyuan WU 2025-12-13 13:32:19 +08:00
  • bf17e769fe Add sgl_router_attempt_http_responses_total for single attempt information (#15037) fzyzcjy 2025-12-13 13:29:18 +08:00
  • 9a5d6a84ab Add error code in prometheus metrics and add X-SMG-Error-Code header (#15036) fzyzcjy 2025-12-13 13:28:32 +08:00
  • 31c23e5fe3 Provide more fine grained error reason for reqwest error (#15032) fzyzcjy 2025-12-13 13:27:06 +08:00
  • 06617a9ec8 Tiny change http router response format to unify (#15031) fzyzcjy 2025-12-13 13:25:45 +08:00
  • e79ca95961 Tiny unify grpc existing error responses into new format (#15030) fzyzcjy 2025-12-13 13:25:10 +08:00
  • 9d3b411c56 Add code field and unify error responses for router (#15028) fzyzcjy 2025-12-13 13:21:34 +08:00
  • 05325db349 Super tiny remove unused log_request (#15035) fzyzcjy 2025-12-13 13:21:24 +08:00
  • 01e3b3f3a3 Fix decode OOM caused by retraction (#14939) Liangsheng Yin 2025-12-13 13:59:17 +09:00
  • 8698867479 [CI]Add gb200 runner back (#15024) Baizhou Zhang 2025-12-12 20:19:34 -08:00
  • 291396544e Add a special label for b200 CI runner that can run kernel tests (#15033) Kangyan-Zhou 2025-12-12 19:41:36 -08:00
  • 665cb02003 Fix regression caused by fa3 block_table (#15009) Shu Wang 2025-12-12 21:11:09 -06:00
  • 7160283800 Tiny remove the duplicate function in spec v2 (#14957) Liangsheng Yin 2025-12-13 12:08:27 +09:00
  • 77873343c4 tiny update: use rope kernel in sgl-kernel for amd (#14955) Qiaolin Yu 2025-12-12 18:48:12 -08:00
  • 267170bf1d Clean up server args and engine startup processes (#15015) Lianmin Zheng 2025-12-12 18:46:07 -08:00
  • 313f59ad80 Add soft watchdogs to debug soft hangs (#15023) fzyzcjy 2025-12-13 10:41:35 +08:00
  • 487cf81a68 Tiny extract SchedulerWatchdog (#15021) fzyzcjy 2025-12-13 10:40:32 +08:00
  • df111bc0fe Super tiny add gsp-fast-prepare (#14992) fzyzcjy 2025-12-13 09:45:21 +08:00
  • 6d2b3324ef Super tiny fix confusing slash_command_handler hint (#14976) fzyzcjy 2025-12-13 09:42:50 +08:00
  • 8cc77261ec Super tiny remove unused argument (#14966) fzyzcjy 2025-12-13 09:41:01 +08:00
  • d143b02097 [registry] Add a strict mode to model registration (#14933) Yinghai Lu 2025-12-12 17:25:48 -08:00
  • 9b9d21312a Feature/Fix multi lora scheduler blocking issue and evict LoRA None lastly (#14795) Chenxi Li 2025-12-12 17:13:05 -08:00
  • 44fd701732 Tune triton fused moe for the case of glm-4.6-fp8 b200 tp4 (#15020) Qiaolin Yu 2025-12-12 17:04:56 -08:00
  • 9a56273ad2 [model-gateway] refactor: unify worker management into modular workflow structure (#15010) Simo Lin 2025-12-12 15:25:35 -08:00
  • b737a125c5 Update ci permission (#15014) Lianmin Zheng 2025-12-12 14:11:40 -08:00
  • 1b5e903480 Refactor of http and engine entrypoints to allow custom override (#14869) Lianmin Zheng 2025-12-12 12:01:08 -08:00
  • 171b442ad3 Add KV4-capable backend flashmla and update server args (#14989) Ho-Ren (Jack) Chuang 2025-12-12 11:50:27 -08:00
  • 4b7b5af36a Revert several PRs (#14958) Yineng Zhang 2025-12-12 11:25:12 -08:00
  • ec242f516e Super tiny extract route_typed_request_once (#14951) fzyzcjy 2025-12-13 03:10:10 +08:00
  • b243154614 Fix CI by reverting incorrect metric check logic (#15004) Kangyan-Zhou 2025-12-12 10:07:38 -08:00
  • 526fd0082f [model-gateway] refactor: workflow engine cleanup and minor optimization (#15001) Simo Lin 2025-12-12 09:03:28 -08:00
  • 56d0ad47f4 [model-gateway] fix: handle workflow deadlock and optimize cycle detection (#15000) Simo Lin 2025-12-12 08:53:45 -08:00
  • 306e5b8d0b [model-gateway] feat: add DAG parallel execution support and workflow optimization (#14999) Simo Lin 2025-12-12 08:31:12 -08:00
  • 10c68f6236 [model-gateway] refactor: extract workflow engine to src/workflow module (#14996) Simo Lin 2025-12-12 06:29:12 -08:00
  • c7c837cd1d Update CODEOWNERS for multimodal_gen (#14995) Mick 2025-12-12 22:08:16 +08:00
  • 3e1e71575c [diffusion] docker: Tiny fix Docker Hub link in installation documentation (#14987) Xiaoyu Zhang 2025-12-12 20:25:36 +08:00
  • 8fa8d9d7e8 [PD] Add decode PP event loop for PD disaggregation (#14945) Kevin Li 2025-12-12 04:15:00 -08:00
  • c8cf1cafdb [1/N] Update doc of Pipeline Parallelism (#14985) Shangming Cai 2025-12-12 19:32:52 +08:00