 bjmsongandbjmsong
|
d3024f4fc8
|
support e4m3 kvcache in qwen2 & add kv scaling facotr json (#2894)
Co-authored-by: bjmsong <bjmsong@126.com>
|
2025-01-18 11:43:22 +08:00 |
|
 bjmsongandroot
|
17de02f98d
|
Integration of TurboMind AWQ (#2828)
Co-authored-by: root <bjmsong@126.com>
|
2025-01-13 20:14:16 +08:00 |
|
 bjmsongandroot
|
0bb0f76311
|
Support FP8 E4M3 KV Cache (#2786)
Co-authored-by: root <bjmsong@126.com>
|
2025-01-12 21:17:11 -08:00 |
|
 bjmsongandroot
|
e21026690d
|
benchmark decoding attention kernel with cudnn (#2467)
Co-authored-by: root <bjmsong@126.com>
|
2024-12-17 03:31:57 -08:00 |
|
 bjmsongandroot
|
f67723940d
|
decoding attention kernel benchmark (#2425)
Co-authored-by: root <bjmsong@126.com>
|
2024-12-11 04:46:59 -08:00 |
|
 bjmsongandroot
|
01017d4c20
|
Support LoRA in Completion API (#2243)
Co-authored-by: root <bjmsong@126.com>
|
2024-11-29 16:13:38 -08:00 |
|
 bjmsongandroot
|
91e5dbf554
|
add profile in offline benchmark & update doc (#2123)
Co-authored-by: root <bjmsong@126.com>
|
2024-11-27 14:57:13 -08:00 |
|
 bjmsongandroot
|
ad30d5cf9a
|
Benchmark with Pytorch Profiler easily (#2110)
Co-authored-by: root <bjmsong@126.com>
|
2024-11-21 23:29:50 -08:00 |
|