v4.3.1 update. (#2817)
This commit is contained in:
@@ -1,9 +1,9 @@
|
||||

|
||||
# Overview
|
||||
|
||||
# CUTLASS 4.3.0
|
||||
# CUTLASS 4.3.1
|
||||
|
||||
_CUTLASS 4.3.0 - Nov 2025_
|
||||
_CUTLASS 4.3.1 - Nov 2025_
|
||||
|
||||
CUTLASS is a collection of abstractions for implementing high-performance matrix-matrix multiplication (GEMM)
|
||||
and related computations at all levels and scales within CUDA. It incorporates strategies for
|
||||
@@ -51,6 +51,8 @@ To get started quickly - please refer :
|
||||
- Added fake tensor and stream to decouple compile jit function with "from_dlpack" flow. Now we no longer require users to have real tensor when compile jit function.
|
||||
- Added FastDivmodDivisor with Python operator overloads, new APIs, Cute dialect integration, and optimized static tile scheduler performance for faster index mapping.
|
||||
- Added l2 cache evict priority for tma related ops. Users could do fine-grain l2 cache control.
|
||||
- Added Blackwell SM103 support.
|
||||
- Multiple dependent DSOs in the wheel have been merged into one single DSO.
|
||||
* Debuggability improvements:
|
||||
- Supported source location tracking for DSL APIs (Allow tools like ``nsight`` profiling to correlate perf metrics with Python source code)
|
||||
- Supported dumping PTX and CUBIN code: [Hello World Example](https://github.com/NVIDIA/cutlass/blob/main/examples/python/CuTeDSL/notebooks/hello_world.ipynb)
|
||||
@@ -95,6 +97,8 @@ To get started quickly - please refer :
|
||||
- Fixed TensorSSA.__getitem__ indexing to match CuTe's indexing convention
|
||||
- Fixed an issue with cutlass.max and cutlass.min
|
||||
- Fixed an issue with mark_compact_shape_dynamic
|
||||
- Fixed device reset issue with tvm-ffi
|
||||
- Fixed tvm-ffi export compiled function
|
||||
|
||||
## CUTLASS C++
|
||||
* Further enhance Blackwell SM100 Attention kernels in [example 77](https://github.com/NVIDIA/cutlass/tree/main/examples/77_blackwell_fmha/).
|
||||
|
||||
Reference in New Issue
Block a user