Commit Graph

  • 6d08ce2aa9 Use Optional with None default (#2770) HAI 2025-01-07 01:35:08 -08:00
  • 380930a959 add benchmark_moe_align_blocks (#2767) Xiaoyu Zhang 2025-01-07 14:20:50 +08:00
  • 9dec582dab Remove --modelopt-config in server_args (#2758) Lianmin Zheng 2025-01-06 16:35:45 -08:00
  • b01febdca0 Update README.md (#2757) Lianmin Zheng 2025-01-06 15:36:23 -08:00
  • 1acbaf1b5a Add generator-style run_batch function (#2513) Xingyao Wang 2025-01-06 18:04:55 -05:00
  • 287427e2e6 Enable Nvidia's ModelOpt fp8 quantized models (#2535) Zhiyu 2025-01-06 14:54:52 -08:00
  • b8574f6953 Clean up eagle code (#2756) Lianmin Zheng 2025-01-06 14:54:18 -08:00
  • 2855caa481 feat: add devcontainer.json for VSCode development (#2745) 王博伟 2025-01-07 06:00:55 +08:00
  • 2329e1ddd0 Support llamafy/Qwen-Qwen2.5-7B-Instruct-llamafied (#2748) Xu-Chen 2025-01-07 05:56:28 +08:00
  • 0f3eb1d294 Support cutlass Int8 gemm (#2752) Ke Bao 2025-01-06 22:51:22 +08:00
  • 06dd2eab84 Remove unused var in moe_align_kernel (#2751) Ke Bao 2025-01-06 22:13:28 +08:00
  • 439f65809f Fix sgl-kernel cu118 compile issue (#2750) Ke Bao 2025-01-06 21:59:31 +08:00
  • 2f0d386496 chore: bump v0.4.1.post4 (#2713) Yineng Zhang 2025-01-06 01:29:54 +08:00
  • 3900a94afe Support twoshot kernel (#2688) yizhang2077 2025-01-06 00:47:16 +08:00
  • ded9fcd09a improve moe_align_kernel for deepseek v3 (#2735) Xiaoyu Zhang 2025-01-06 00:28:22 +08:00
  • bc6ad367c2 fix lint (#2733) Yineng Zhang 2025-01-05 14:45:42 +08:00
  • 3a22a303d1 Revert the GLOO_SOCKET_IFNAME change (#2731) Lianmin Zheng 2025-01-04 20:13:16 -08:00
  • bdb3929dbb Refactor SchedulePolicy to improve code organization (#2571) libra 2025-01-04 00:05:16 +08:00
  • f5d0865b25 feat: Support VLM in reference_hf (#2726) Ce Gao 2025-01-03 22:32:30 +08:00
  • afdee7b1a9 [Docs] fix 404 - Contributor Guide, again (#2727) Ce Gao 2025-01-03 22:21:38 +08:00
  • cb34d848ac Update README.md (#2722) Lianmin Zheng 2025-01-03 00:32:20 -08:00
  • 0f9cc6d8d3 Fix package loss for small models (#2717) Lianmin Zheng 2025-01-02 18:25:26 -08:00
  • c7ae474a49 [Feature, Hardware] Enable DeepseekV3 on AMD GPUs (#2601) yigex 2025-01-03 08:23:19 +08:00
  • bdf946bf81 Support loading pre-sharded moe weights (#2716) Lianmin Zheng 2025-01-02 15:07:37 -08:00
  • 8c8779cd05 [Fix] fix retract error in eagle speculative decoding (#2711) yukavio 2025-01-03 02:28:39 +08:00
  • 1775b963db [Fix] fix incorrectly overwriting the port specified in ServerArgs (#2714) Mick 2025-01-03 02:28:22 +08:00
  • dd2e2d275f Docs: Update documentation workflow and contribution guide (#2704) Shi Shuai 2025-01-02 17:18:31 +00:00
  • a990daff9c Included multi-node DeepSeekv3 example (#2707) Rodrigo Garcia 2025-01-02 15:17:03 +01:00
  • ba5112ff69 feat: support moe_align_block_size_triton (#2712) Yineng Zhang 2025-01-02 21:47:44 +08:00
  • 815dce0554 Eagle speculative decoding part 4: Add EAGLE2 worker (#2150) yukavio 2025-01-02 19:22:34 +08:00
  • ad20b7957e Eagle speculative decoding part 3: small modifications to the general scheduler (#2709) Lianmin Zheng 2025-01-02 02:09:08 -08:00
  • 9183c23eca Speed up update_weights_from_tensor (#2695) fzyzcjy 2025-01-02 18:05:19 +08:00
  • 148254d4db Improve moe reduce sum kernel performance (#2705) kk 2025-01-02 17:11:06 +08:00
  • a4d6d6f1dd [feat]: Add math eval to CI nightly run (#2663) Xiaotong Jiang 2025-01-01 15:29:35 -08:00
  • 062c48d2bd [Docs] Add Support for Pydantic Structured Output Format (#2697) Shi Shuai 2025-01-01 23:08:43 +00:00
  • b6e0cfb5e1 ROCm base image update (#2692) kk 2025-01-01 12:12:19 +08:00
  • 0d8d97b8e6 Doc: Rename contribution_guide.md (#2691) Chayenne 2024-12-31 14:35:48 -08:00
  • 0a765bbccc Docs: Refactor Contribution Guide (#2690) Shi Shuai 2024-12-31 22:11:00 +00:00
  • 286cad3ee3 h200 tuning fused_moe_triton config for Mixtral 8x7B/8x22B and Qwen2 57BA14B (#2689) Xiaoyu Zhang 2024-12-31 23:17:36 +08:00
  • dc7eb01f19 [Fix] fix openai adapter (#2685) Ying Sheng 2024-12-31 02:48:19 -08:00
  • b0524c3789 Eagle speculative decoding part 2: Fix cuda graph + DP attention hanging (#2684) Lianmin Zheng 2024-12-31 02:25:05 -08:00
  • 6c42fa229d Update README.md (#2683) Lianmin Zheng 2024-12-31 00:13:10 -08:00
  • d49b13c6f8 feat: use CUDA 12.4 by default (for FA3) (#2682) Yineng Zhang 2024-12-31 15:52:09 +08:00
  • bedc4c7a50 misc: update CODEOWNERS (#2680) Yineng Zhang 2024-12-31 15:04:50 +08:00
  • f44d143949 Support target model verification in the attention backend (#2678) Lianmin Zheng 2024-12-30 22:58:55 -08:00
  • b6b57fc200 minor: cleanup sgl-kernel (#2679) Yineng Zhang 2024-12-31 14:52:00 +08:00
  • b4403985d0 Add cutlass submodule for sgl-kernel (#2676) Ke Bao 2024-12-31 14:28:29 +08:00
  • 339c69a243 Improve the computation for time_per_output_token Prometheus metrics (#2674) Lianmin Zheng 2024-12-30 21:40:14 -08:00
  • f707470019 CI: Update scripts to fail fast (#2672) fzyzcjy 2024-12-31 11:04:01 +08:00
  • 21ec66e59e Minor follow-up fixes for the logprob refactor (#2670) Lianmin Zheng 2024-12-30 05:42:08 -08:00
  • c5210dfa38 AMD DeepSeek_V3 FP8 Numerical fix (#2667) HAI 2024-12-30 05:31:12 -08:00
  • a29dd9501d Add GemLite caching after each capture (#2669) mobicham 2024-12-30 14:27:29 +01:00
  • 9c6ba2484f Refactor logprob computation to return the real logprob used in sampling (#2664) Lianmin Zheng 2024-12-30 04:51:38 -08:00
  • b02da24a5b Refactor sgl-kernel build (#2642) Ke Bao 2024-12-30 18:07:01 +08:00
  • bdd2827a80 Update structured_outputs.ipynb (#2666) Lianmin Zheng 2024-12-30 00:46:41 -08:00
  • 8c3b420eec [Docs] clean up structured outputs docs (#2654) Lianmin Zheng 2024-12-29 23:57:16 -08:00
  • e6f523b5f2 fix typo in python/sglang/srt/layers/quantization/fp8.py (#2655) HAI 2024-12-29 23:45:02 -08:00
  • 3231817861 Revert "[feat] Add math eval to CI" (#2656) Lianmin Zheng 2024-12-29 23:05:50 -08:00
  • a11f8d5f6a [feat] Add math eval to CI (#2652) Xiaotong Jiang 2024-12-29 22:49:41 -08:00
  • 098d659c0e docs: update README (#2651) Yineng Zhang 2024-12-30 13:33:29 +08:00
  • 76d14f8cb9 add 2*h20 node serving example for deepseek v3 (#2650) Lzhang-hub 2024-12-30 13:04:38 +08:00
  • b08c308ebc Update the timeout in nightly-test.yml (#2649) Lianmin Zheng 2024-12-29 14:51:07 -08:00
  • 03d5fbfd44 Release 0.4.1.post3 - upload the config.json to PyPI (#2647) Lianmin Zheng 2024-12-29 14:25:53 -08:00
  • 1703d766d8 CI: skip special token for engine token ids unit test (#2648) Chayenne 2024-12-29 13:52:50 -08:00
  • 09e6e2aa33 Merge branch 'main' of github.com:sgl-project/sglang zhaochenyang20 2024-12-29 21:48:21 +00:00
  • fad29f7f52 CI: Fix unittest for engine input token ids and output token ids (#2646) Shi Shuai 2024-12-29 21:28:59 +00:00
  • 35bdb48557 [Feature] Get Token IDs with Engine.generate() (#2636) Shi Shuai 2024-12-29 20:28:27 +00:00
  • b085e06b01 docs: add development guide using docker (#2645) Yineng Zhang 2024-12-30 02:22:54 +08:00
  • 763dd55d17 docs: update README (#2644) Yineng Zhang 2024-12-30 01:24:06 +08:00
  • 3ccf566b0d chore: bump v0.4.1.post2 (#2643) Yineng Zhang 2024-12-30 00:11:46 +08:00
  • afa0341e57 Update Triton configs for block fp8 kernels (#2641) HandH1998 2024-12-29 22:53:47 +08:00
  • 30828e7192 AMD: set weights and scaling numbers properly for block FP8 (#2637) HAI 2024-12-29 03:23:39 -08:00
  • e0e09fceeb [Session] Update session control interface (#2635) Ying Sheng 2024-12-29 02:10:27 -08:00
  • 9c05c6898e Add llama_eagle.py (#2640) Lianmin Zheng 2024-12-29 01:45:35 -08:00
  • 3464e57b62 minor: add nsys cli for docker dev (#2639) Yineng Zhang 2024-12-29 17:28:11 +08:00
  • 3815b23ccb Clean up wrapper in flashinfer backend (#2638) Lianmin Zheng 2024-12-29 00:45:57 -08:00
  • fd34f2da35 [Docs] Add EBNF to sampling params docs (#2609) Adarsh Shirawalmath 2024-12-29 13:35:00 +05:30
  • 8ee9a8501a [Feature] Function Calling (#2544) Tanjiro 2024-12-29 11:28:52 +05:30
  • fd28640dc5 Add update_weights_from_tensor (#2631) fzyzcjy 2024-12-29 05:30:27 +08:00
  • 7863e4368a add configs for block fp8 related kernels (#2628) Yineng Zhang 2024-12-28 23:12:04 +08:00
  • 333e3bfde5 [docs]Refactor constrained decoding tutorial (#2633) Shi Shuai 2024-12-28 15:00:38 +00:00
  • 239c9d4d3a Docs: Add constrained decoding tutorial (#2614) Shi Shuai 2024-12-28 07:54:28 +00:00
  • 855d0ba381 [CI] Fix nightly test and raise better error message (#2626) Lianmin Zheng 2024-12-27 22:16:39 -08:00
  • 9254a33ad4 avoid fused_moe_triton padding circular import (#2624) Xiaoyu Zhang 2024-12-28 14:01:35 +08:00
  • 8a2681e26a Update readme (#2625) Ke Bao 2024-12-28 13:39:56 +08:00
  • 5276a675f5 Add more supporting organizations (#2623) Lianmin Zheng 2024-12-27 13:41:41 -08:00
  • 751e5ca273 [minor] clean up docs and eos id (#2622) Lianmin Zheng 2024-12-27 11:23:46 -08:00
  • 7a7ac6bea1 [FIX] Update EOS from config (#2475) Yang Zheng 2024-12-28 02:59:56 +08:00
  • d9e6ee382b docs: update README (#2618) Yineng Zhang 2024-12-28 00:21:53 +08:00
  • ef5b0ff90b chore: bump v0.4.1.post1 (#2616) Yineng Zhang 2024-12-28 00:11:06 +08:00
  • 6e5305158c update sgl_moe_align_block_size usage (#2617) HandH1998 2024-12-28 00:01:13 +08:00
  • 77d1210b36 fix moe_align_block_size (#2615) HandH1998 2024-12-27 23:32:53 +08:00
  • 70dc2fbe2d Change extend attention kernel launch parameter for ROCm platform to … (#2610) kk 2024-12-27 16:32:17 +08:00
  • b438a2e512 Fix triton kernel performance regression (#2611) kk 2024-12-27 15:54:38 +08:00
  • 7ca751ff7d Fused moe triton cfg opt for rocm (#2612) kk 2024-12-27 15:38:22 +08:00
  • c75adfec59 Update CODEOWNERS (#2608) Lianmin Zheng 2024-12-26 20:58:08 -08:00
  • 7722c11c1d Regression fix to AMD/ROCm from recent change (#2606) HAI 2024-12-26 20:22:14 -08:00
  • b2ed5c8ea7 Tiny code cleanup in tokenizer_manager.py (#2586) fzyzcjy 2024-12-27 09:53:09 +08:00
  • f46f394f4d Update README.md (#2605) Lianmin Zheng 2024-12-26 10:58:49 -08:00
  • 2125898af5 Update contributor_guide.md (#2603) Lianmin Zheng 2024-12-26 08:36:13 -08:00