Commit Graph

  • 8f0b63139e Docs: improve EAGLE docs (#4038) simveit 2025-03-06 07:40:21 +01:00
  • b9b3b098b9 feat: support docs auto live-reload with sphinx-autobuild (#4111) samzong 2025-03-06 14:39:34 +08:00
  • aee30630d8 Add a pointer to the real KV cache pool (#4113) Zhiqiang Xie 2025-03-05 21:39:07 -08:00
  • 286e6540a6 Remove prefill-only-one-req (#4117) Lianmin Zheng 2025-03-05 20:58:48 -08:00
  • 718c391fd7 [Hoxfix] Fix incomplete token_to_kv_pool refactor (#4121) Wenxuan Tan 2025-03-05 21:32:42 -06:00
  • fc671f66c1 chore: bump v0.4.3.post3 (#4114) Yineng Zhang 2025-03-05 17:26:10 -08:00
  • 197751e9a1 fix Non-consecutive header level increase in docs/router/router.md (#4099) samzong 2025-03-06 09:02:32 +08:00
  • d2d0d061d9 fix cross-reference error and spelling mistakes (#4101) samzong 2025-03-06 08:39:02 +08:00
  • 25482edb5c Online serving benchmarks of real datasets for hierarchical KV caching (#3211) Yueyang Pan 2025-03-06 01:16:43 +01:00
  • 62b362b1f1 Debug radixcache: refactor recursive helper methods (#3029) luzengxiangcn 2025-03-06 08:11:42 +08:00
  • 44d7646371 remove testing on PR workflow change (#4110) saienduri 2025-03-05 16:03:18 -08:00
  • cd85b78f94 Create release-docker-amd-nightly.yml (#4105) saienduri 2025-03-05 14:46:26 -08:00
  • 0aaccbbfec revert deepseek docs (#4109) Yineng Zhang 2025-03-05 13:23:11 -08:00
  • 357671e216 Add examples for server token-in-token-out (#4103) Qiaolin Yu 2025-03-05 16:16:31 -05:00
  • e70fa279bc Docs: reorganize dpsk docs (#4108) Chayenne 2025-03-05 13:01:03 -08:00
  • abe74b7b59 Docs: Add DeepSeek optimization ablations documentation (#4107) Tommy Yang 2025-03-06 04:25:51 +08:00
  • 70b3c6eeb1 Add update_weights_from_disk endpoint to Engine (#4102) Jhin 2025-03-05 14:25:18 -06:00
  • ef9d3b3c2c Fix triton kernel illegal memory issue for eagle (#4100) Ke Bao 2025-03-06 03:23:53 +08:00
  • fc91d08a8f [Revision] Add fast decode plan for flashinfer mla (#4012) Baizhou Zhang 2025-03-05 11:20:41 -08:00
  • 71ab0dabe0 Fix the moe padding conditional logic (#4081) HAI 2025-03-05 10:56:51 -08:00
  • d3d4d76758 [Eagle] Refactor eagle speculative decoding (#3986) Ying Sheng 2025-03-05 08:06:07 -08:00
  • 5be8f1ed98 ROCM: AITER BLOCK GEMM (#4075) yigex 2025-03-05 19:10:49 +08:00
  • e5760bc40a bench: add dataset param for bench_multiturn (#3990) Lu Changqi 2025-03-05 17:21:37 +08:00
  • 56a724eba3 [QUANT] Add GPTQModel Dynamic Quantization + lm_head Quantization (#3790) Qubitium-ModelCloud 2025-03-05 17:11:00 +08:00
  • 583d6af71b example: add vlm to token in & out example (#3941) Mick 2025-03-05 14:18:26 +08:00
  • e074d84e5b [Minor] more code cleanup (#4077) Lianmin Zheng 2025-03-04 21:23:47 -08:00
  • 4725e3f652 Add examples for returning hidden states when using the server (#4074) Qiaolin Yu 2025-03-04 22:31:50 -05:00
  • 77a3954bf7 Simplify eagle tests and TP sync in grammar backend (#4066) Lianmin Zheng 2025-03-04 13:40:40 -08:00
  • 03b0364f76 Update nextn ci test (#4071) Ke Bao 2025-03-05 05:01:24 +08:00
  • 2dd7d0c533 Revert "Fix nightly-test CI" (#4065) Lianmin Zheng 2025-03-04 05:38:24 -08:00
  • 0d4e3228cf [Feature] Add test for speculative_token_map (#4016) William 2025-03-04 20:26:24 +08:00
  • 926f8efc0c remove unused max_jobs (#3607) Liu Jinjie 2025-03-04 04:23:39 -08:00
  • 9545bfb28a fix: support gelu_new activation function in gpt2 (#3712) Xiuyu Li 2025-03-04 04:09:52 -08:00
  • 37373ef2bb sgl-router - issues on routing and project build. (#3870) (#3948) Michael Feil 2025-03-04 04:06:30 -08:00
  • 61261b3996 [XCCL] Use xccl for xpu backend since xccl is ready in latest PyTorch. (#3954) Chen Shengzhi 2025-03-04 20:05:56 +08:00
  • 19120f71f3 [Fix & Style] Refactor the grammar backend to reduce human errors and improve readability (#4030) DarkSharpness 2025-03-04 20:56:45 +09:00
  • 2415ec3896 Remove grafana dashboard's datasource uid (#4051) Kebe 2025-03-04 19:44:51 +08:00
  • 87f671ab58 Fix debug_tensor_dump_output_folder optional key missing (#4046) Qubitium-ModelCloud 2025-03-04 19:42:48 +08:00
  • 51d25405a7 ROCm: update aiter and its usage to fused moe (bloat16, fp8, fp8 block-quant) (#4053) HAI 2025-03-04 03:00:46 -08:00
  • e0a2c96308 Fix breakage problem when using custom_ar (#4052) kk 2025-03-04 18:59:03 +08:00
  • 12f2e6c3f1 Fix: #3988 using blockwise_int8 (#4023) Xihuai Wang 2025-03-04 15:49:58 +08:00
  • 95575aa76a Reasoning parser (#4000) Xihuai Wang 2025-03-04 13:16:36 +08:00
  • 11eea69e70 Fix assert options.num_stages != 0 error in the latest ROCm build image (#4049) kk 2025-03-04 12:37:03 +08:00
  • 1baa9e6cf9 docs: update README (#4044) Yineng Zhang 2025-03-03 17:09:18 -08:00
  • 911fcd0910 Update README.md (#4043) Lianmin Zheng 2025-03-03 16:29:46 -08:00
  • 9fafa62db7 Share target model embed and head weights for nextn (#4033) Ke Bao 2025-03-04 05:30:04 +08:00
  • 146ac8df07 Add examples in sampling parameters (#4039) Chayenne 2025-03-03 13:04:32 -08:00
  • 57a404fd55 Remove outdated test utils and fix links for the doc of sampling params (#3999) Qiaolin Yu 2025-03-03 12:41:38 -05:00
  • 2796fbb53d Docs: Fix sampling parameter (#4034) Chayenne 2025-03-03 09:32:36 -08:00
  • 935cda944b Misc clean up; Remove the support of jump forward (#4032) Lianmin Zheng 2025-03-03 07:02:14 -08:00
  • 110e006673 Reorganize python source files in sgl-kernel with multiple files (#4027) Lianmin Zheng 2025-03-03 06:36:40 -08:00
  • 6b45a21d16 Reorganize c++ source files in sgl-kernel with multiple folders (#4025) Lianmin Zheng 2025-03-03 05:32:30 -08:00
  • a7000a7650 Update metrics documentation (#3264) Yudi Xue 2025-03-03 05:03:58 -08:00
  • 1a8f995c46 remove cache configs in model definitions (#4031) Lianmin Zheng 2025-03-03 05:00:50 -08:00
  • a3ab768a2b Clean up custom allreduce (#4029) Lianmin Zheng 2025-03-03 04:59:53 -08:00
  • 66301e124f Improve code styles (#4021) Lianmin Zheng 2025-03-03 03:20:23 -08:00
  • ac2387279e Support penalty in overlap mode; return logprob with chunked prefill; improve benchmark scripts (#3988) Lianmin Zheng 2025-03-03 00:12:04 -08:00
  • 0194948fd9 Optimize Triton Kernel of Group GEMM in DeepGEMM Benchmark (#4014) Stefan He 2025-03-02 23:29:55 -08:00
  • b4d34cd35d Fix nightly-test CI (#3826) yinfan98 2025-03-03 15:14:45 +08:00
  • 728e175fc4 Add examples to token-in-token-out for LLM (#4010) Chayenne 2025-03-02 21:03:49 -08:00
  • 9e1014cf99 Revert "Add fast decode plan for flashinfer mla" (#4008) Lianmin Zheng 2025-03-02 19:29:10 -08:00
  • fa56106731 Add fast decode plan for flashinfer mla (#3987) Baizhou Zhang 2025-03-02 19:16:37 -08:00
  • 7fbab730bd [feat] add small vocab table for eagle's draft model[1]. (#3822) Zhousx 2025-03-03 10:58:45 +08:00
  • b7e274f2d9 Add Benchmark for DeepGEMM Group GEMM (#3993) Stefan He 2025-03-02 17:47:21 -08:00
  • 9cf4077294 Enable custom AR for AMD GPUs and maintain it in sgl-kernel (#3406) Hubert Lu 2025-03-02 15:19:06 -08:00
  • d3fe9bae56 Add accuracy test for TP torch compile (#3994) Ke Bao 2025-03-03 05:18:18 +08:00
  • 00ce7e311c Fix all gather torch compile (#3992) Ke Bao 2025-03-02 16:41:38 +08:00
  • 50f28f65a0 fix typo in deep gemm benchmarking(#3991) Xiaoyu Zhang 2025-03-02 16:34:00 +08:00
  • 90a55e2566 add deepgemm and sglang fp8 block-wise gemm benchmark (#3893) Xiaoyu Zhang 2025-03-02 15:01:58 +08:00
  • 407e2b923d Update CODEOWNERS (#3989) Lianmin Zheng 2025-03-01 21:47:30 -08:00
  • 40782f05d7 Refactor: Move return_hidden_states to the generate input (#3985) Qiaolin Yu 2025-03-01 20:51:29 -05:00
  • 18bb216c28 Revert "[MOE] enable efficient moe_alignment multi-blocks execution (3x~6x)" (#3982) Chayenne 2025-02-28 23:57:17 -08:00
  • 6b859e7ddd Docs: add special warning to engine docs (#3979) Chayenne 2025-02-28 21:59:20 -08:00
  • 930da877c4 rename FunctionCallReqInput to ParseFunctionCallReq (#3976) Chayenne 2025-02-28 18:46:25 -08:00
  • 3f8a441437 Docs: Add redline to highlight main process (#3977) Chayenne 2025-02-28 18:37:15 -08:00
  • aceb420179 Docs: add type hint to smapling parameters (#3975) Chayenne 2025-02-28 18:21:20 -08:00
  • 90a4b7d98a [Feature]Support ragged prefill in flashinfer mla backend (#3967) Baizhou Zhang 2025-02-28 18:13:56 -08:00
  • f3b99f73b3 update flashinfer-python version Yineng Zhang 2025-02-28 16:31:59 -08:00
  • 9e74ee91da Update cutlass dependency (#3966) Elfie Guo 2025-02-28 16:16:31 -08:00
  • 77a6c9d229 Remove unused imports from rocm mla kernel. (#3963) Chaitanya Sri Krishna Lolla 2025-02-28 23:31:08 +05:30
  • e3e0bc50a9 [Feature] SPMD for SGLang + Verl (#3852) fzyzcjy 2025-03-01 01:53:10 +08:00
  • bac414ab53 [Feature] integrate Structural Tag in xgrammar backend for function calling (#3566) mlmz 2025-02-28 15:33:41 +08:00
  • eec3f6d1eb [Bugfix] Fix tokenizer_manager not getting 400 when req is too long (#3678) Chang Su 2025-02-27 22:59:43 -08:00
  • 90bc26a813 set a strict sgl-kernel version (#3950) Chayenne 2025-02-27 22:44:57 -08:00
  • ec0a72c2d9 Fix bench_serving not recognizing OPENAI_API_KEY (#3870) Kebe 2025-02-28 12:18:53 +08:00
  • 1c96fa86cf [MOE] enable efficient moe_alignment multi-blocks execution (3x~6x) (#3613) yiakwy-xpu-ml-framework-team 2025-02-28 11:42:48 +08:00
  • bc20e93f2d [feat] Add Vertex AI compatible prediction route for /generate (#3866) KCFindstr 2025-02-27 19:42:15 -08:00
  • d38878523d Fix the doc link for sampling params (#3861) Qiaolin Yu 2025-02-27 16:31:43 -05:00
  • 564bdf29f7 upgrade flashinfer v0.2.2.post1 (#3934) Yineng Zhang 2025-02-27 09:53:48 -08:00
  • 5d86016855 revert "Docs: Reorngaize dpsk links #3900" (#3933) Yineng Zhang 2025-02-27 08:57:13 -08:00
  • d281587989 Improve: Support xgrammar 0.1.14 (#3593) Enrique Shockwave 2025-02-27 16:42:54 +00:00
  • b0df5d240b Tuning Script for Feature DeepSeek V3/R1 INT8 Quantization (block-wise) (#3922) laixin 2025-02-27 18:59:46 +08:00
  • 3e02526b1f [Doc] Add experimental tag for flashinfer mla (#3925) Baizhou Zhang 2025-02-27 01:55:36 -08:00
  • d8a98a2cad [Docs] Improve DPSK docs in dark mode (#3914) Stefan He 2025-02-27 00:13:04 -08:00
  • 0519269d20 [Docs] Disable notebook CI when merge to main (#3905) Qing 2025-02-26 22:13:33 -08:00
  • d6898dd253 Add return hidden state in the native API (#3897) Qiaolin Yu 2025-02-27 01:06:54 -05:00
  • 71ed01833d [doc] Update document for flashinfer mla (#3907) Baizhou Zhang 2025-02-26 20:40:45 -08:00
  • 8b681d7724 [Rocm] Fix to the rocm_mla_decode_rope.py returning random result (#3898) Tianxing Wu 2025-02-27 03:05:30 +02:00
  • 194eea1774 [doc] update sponsorship (#3903) ybyang 2025-02-27 08:28:15 +08:00
  • acd1a15921 Docs: Implemented frontend docs (#3791) simveit 2025-02-27 00:30:05 +01:00