v3.9 update (#2203)
* v3.9 update * voidD --------- Co-authored-by: yuzhai <yuzhai@nvidia.com>
This commit is contained in:
150
media/docs/cpp/blackwell_cluster_launch_control.md
Normal file
150
media/docs/cpp/blackwell_cluster_launch_control.md
Normal file
@@ -0,0 +1,150 @@
|
||||
# Blackwell Cluster Launch Control
|
||||
|
||||
## Overview
|
||||
|
||||
A GEMM workload usually consists of three phases: prologue, mainloop and epilogue. Each SM will process multiple output tiles in series if the number of output tiles are much more than the number of SMs, completely exposing the overhead of prologue and epilogue.
|
||||
|
||||
Consider a GEMM that has `20x20x1` output tiles, running on a GPU with `100` SMs. There is another kernel occupying all the resources of `20` SMs so only `80` SMs can be used. Assume cluster shape is `1x1x1`. The following diagram shows how the schedule would look like for such a kernel.
|
||||
|
||||
<p align="center"><img src=../images/non_persistent.png alt="A beautiful sunset" title="Sunset over the mountains"></p>
|
||||
|
||||
|
||||
### Static Scheduler
|
||||
CUTLASS has adopted a software technique named **persistent kernels**. Persistent clusters, or Workers, can stay on the GPU throughout kernel execution and process multiple tiles, hiding prologue and epilogue costs. The tile scheduler statically determines the next output tile to process with zero overhead.
|
||||
|
||||
However, static scheduler is susceptible to workload imbalance if the resources of some SMs are unavailable. The following diagram illustrates this issue.
|
||||
|
||||
<p align="center"><img src=../images/persistent_static.png alt="A beautiful sunset" title="Sunset over the mountains"></p>
|
||||
|
||||
### Dynamic Scheduler with Cluster Launch Control
|
||||
A fundamental limitation of persistent scheduling is that the number of SMs this kernel can utilize is unknown in real time. Some SMs might be occupied by another kernel and thus their resources are unavailable. This makes it challenging to load-balance work across SMs.
|
||||
|
||||
Blackwell introduces cluster launch control (CLC) for dynamic scheduling. (See https://docs.nvidia.com/cuda/parallel-thread-execution). With this feature, the kernel launches a grid containing as many threadblocks as there are output tiles to compute in the kernel -- just like one would in a non-persistent kernel. Here we define `ClcID` to be a coordinate from the 3D grid launched on GPU.
|
||||
|
||||
Cluster launch control follows the below rules:
|
||||
|
||||
1. A `ClcID` will be launched as a Worker when there are available resources.
|
||||
2. A `ClcID` can be queried by an existing Worker via `clusterlaunchcontrol.try_cancel` instruction.
|
||||
3. Every `ClcID` is guaranteed to be processed by either (1) or (2).
|
||||
4. Each worker uses the `{blockIdx.x, blockIdx.y, blockIdx.z}` coordinate as the first output tile to process and uses the CLC query for subsequent processing of output tiles.
|
||||
5. `clusterlaunchcontrol.try_cancel` instruction returns either a success signal with a `ClcID` or a decline signal. The most common reason of a decline is that all `ClcID`s have been processed.
|
||||
6. Cluster launch control works on the granularity of clusters. For example, a 2x2 persistent worker cluster's query will consume 2x2 `ClcID`s at once.
|
||||
|
||||
The following diagram shows how the schedule would look like with cluster launch control.
|
||||
|
||||
<p align="center"><img src=../images/persistent_clc.png alt="A beautiful sunset" title="Sunset over the mountains"></p>
|
||||
|
||||
## Programming Model
|
||||
### Pseudo Code
|
||||
#### Non-persistent kernel
|
||||
``` c++
|
||||
// Non-persistent kernel
|
||||
__device__ non_persistent_kernel(...) {
|
||||
setup_common_data_structures();
|
||||
dim3 workCoordinates = blockIdx;
|
||||
coordinate_specific_compute(workCoordinates);
|
||||
}
|
||||
```
|
||||
#### Static Persistent Kernel
|
||||
``` c++
|
||||
// Static Persistent Kernel
|
||||
__device__ static_persistent_kernel(...) {
|
||||
setup_common_data_structures(...);
|
||||
dim3 workCoordinates = blockIdx;
|
||||
do {
|
||||
coordinate_specific_compute(workCoordinates);
|
||||
isValidId, workCoordinates = staticTileScheduler.fetch_next_work();
|
||||
} while (isValidId);
|
||||
}
|
||||
```
|
||||
|
||||
|
||||
#### Blackwell Dynamic Persistent Kernel
|
||||
``` c++
|
||||
// Dynamic Persistent Kernel
|
||||
__device__ clc_dynamic_persistent_kernel(...) {
|
||||
setup_common_data_structures(...);
|
||||
dim3 workCoordinates = blockIdx;
|
||||
do {
|
||||
coordinate_specific_compute(workCoordinates);
|
||||
isValidId, newClcID = clcTileScheduler.fetch_next_work();
|
||||
workCoordinates = newClcID;
|
||||
} while (isValidId);
|
||||
}
|
||||
```
|
||||
|
||||
|
||||
### Cluster Launch Control Pipeline Class
|
||||
|
||||
Please refer to the `PipelineCLCFetchAsync` pipeline class defined in [Cluster launch control pipeline class](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/pipeline/sm100_pipeline.hpp). Cluster launch control queries can be pipelined and mananged by an asynchronous pipeline with producer-consumer relationship (See
|
||||
[pipeline](pipeline.md) document). The producer is the scheduler warp of the 0th CTA in the cluster and the consumers are all warps that need `ClcID`s.
|
||||
|
||||
To setup a CLC pipeline correctly, we need to make sure the params are set to the right values:
|
||||
|
||||
* `transaction_bytes` is `16` as CLC will return a 16B response and store it in the specified shared memory address.
|
||||
* `consumer_arv_count` is the thread count of all the consumer warps in the cluster.
|
||||
* `producer_arv_count` is `1` because only one thread from scheduler warp will be elected to issue `clusterlaunchcontrol.try_cancel`.
|
||||
* `producer_blockid` is `0` to denote that the first CTA in the cluster is producing.
|
||||
|
||||
|
||||
### Dynamic tile scheduler class
|
||||
Please refer to `PersistentTileSchedulerSm100` class defined in [sm100 dynamic persistent tile scheduler](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/gemm/kernel/sm100_tile_scheduler.hpp).
|
||||
|
||||
There are two important methods of the CLC scheduler class. The first is `advance_to_next_work`, which is intended to be executed by one elected thread from the scheduler warp. It effectively sends out the CLC query to the CLC. A CLC query response will be broadcast to the same shared memory address of all CTAs in the cluster.
|
||||
|
||||
The other method is named `get_current_work`. It simply loads the CLC response from the shared memory buffer indexed by a pipeline state.
|
||||
|
||||
|
||||
The CLC pipeline and scheduler classes are used together to ensure correct functionality and necessary synchronization of CLC feature. Please refer to [cluster launch control pipeline unit test](https://github.com/NVIDIA/cutlass/tree/main/test/unit/pipeline/pipeline_cluster_launch_control_async_warp_specialized_blackwell.cu).
|
||||
|
||||
## Blackwell Warp-specialized Persistent Kernel
|
||||
|
||||
Now, let's take a look at how CLC feature is used in our [Blackwell dense GEMM kernel](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/gemm/kernel/sm100_gemm_tma_warpspecialized.hpp).
|
||||
|
||||
This particular warp-specialized kernel has the following warp assignment:
|
||||
|
||||
| Warp Role | Warp |
|
||||
|------------------|-------------|
|
||||
| MMA | 0 |
|
||||
| Scheduler | 1 |
|
||||
| Mainloop Load | 2 |
|
||||
| Epilogue Load | 3 |
|
||||
| Epilogue | 4, 5, 6, 7 |
|
||||
|
||||
Scheduler warp is the producer of the CLC pipeline. The consumers are the MMA, Mainloop Load, Epilogue Load and Epilogue warps. In addition, the scheduler warp is its own consumer! This is because it needs the `success` information from the query to terminate the persistent loop on end-of-grid.
|
||||
|
||||
The CLC pipeline has a depth of 3 to overlap the CLC operations of multiple waves for latency hiding. The first `ClcID` is the preloaded `blockIdx`, which does not require CLC query and is fully static.
|
||||
|
||||
|
||||
|
||||
# Copyright
|
||||
|
||||
Copyright (c) 2025 - 2025 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
```
|
||||
Redistribution and use in source and binary forms, with or without
|
||||
modification, are permitted provided that the following conditions are met:
|
||||
|
||||
1. Redistributions of source code must retain the above copyright notice, this
|
||||
list of conditions and the following disclaimer.
|
||||
|
||||
2. Redistributions in binary form must reproduce the above copyright notice,
|
||||
this list of conditions and the following disclaimer in the documentation
|
||||
and/or other materials provided with the distribution.
|
||||
|
||||
3. Neither the name of the copyright holder nor the names of its
|
||||
contributors may be used to endorse or promote products derived from
|
||||
this software without specific prior written permission.
|
||||
|
||||
THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS "AS IS"
|
||||
AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE
|
||||
IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE ARE
|
||||
DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT HOLDER OR CONTRIBUTORS BE LIABLE
|
||||
FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL
|
||||
DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR
|
||||
SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER
|
||||
CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY,
|
||||
OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE
|
||||
OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
|
||||
```
|
||||
647
media/docs/cpp/blackwell_functionality.md
Normal file
647
media/docs/cpp/blackwell_functionality.md
Normal file
@@ -0,0 +1,647 @@
|
||||
# Blackwell SM100 GEMMs
|
||||
|
||||
[**TLDR; jump to block scaled GEMM example**](#detailed_blockscale_example)
|
||||
|
||||
Blackwell SM100 introduces `tcgen05.mma` instructions. `tcgen05.mma` instructions support all legacy types (`tfloat32_t`, `half_t`, `bfloat16_t`, `int8_t`, `uint8_t`) and
|
||||
the new 4, 6, and 8-bits floating point datatypes with and without scale factors.
|
||||
This document explains the new `tcgen05.mma` instructions supported by CUTLASS and how one can leverage CUTLASS to create
|
||||
efficient SM100 GEMM kernels targeting these new mma instructions.
|
||||
|
||||
Blackwell SM100 has 7 new `tcgen05.mma` instructions. These instructions are 2x to 4x faster then Hopper Architecture's WGMMA instructions.
|
||||
|
||||
| Ptx Instruction | Throughput | Notes |
|
||||
|----------------------------------------------------------------------------------|----------------------------|-------|
|
||||
|tcgen05.mma.cta_group::[1\|2].kind::tf32 | 2x Hopper Tf32 Tensor Core | MMA with A={tf32} x B={tf32} TN, NT, TT, NN layouts |
|
||||
|tcgen05.mma.cta_group::[1\|2].kind::f16 | 2x Hopper Fp16 Tensor Core | MMA with A={f16} x B={f16} or A={bf16} x B={bf16} TN, NT, TT, NN layouts |
|
||||
|tcgen05.mma.cta_group::[1\|2].kind::i8 | 2x Hopper I8 Tensor Core | MMA with A={i8} x B={i8} or A={u8} x B={u8} TN, NT, TT, NN layouts |
|
||||
|tcgen05.mma.cta_group::[1\|2].kind::f8f6f4 | 2x Hopper Fp8 Tensor Core | Mixed precision MMA with A={f4,f6,f8} x B={f4,f6,f8} TN, NT, TT, NN layouts |
|
||||
|tcgen05.mma.cta_group::[1\|2].kind::mxf8f6f4.block_scale | 2x Hopper Fp8 Tensor Core | Block scaled mixed precision MMA with A={mxf4,mxf6,mxf8} x B={mxf4,mxf6,mxf8} with TN, NT, TT, NN layouts |
|
||||
|tcgen05.mma.cta_group::[1\|2].kind::mxf4.block_scale | 4x Hopper Fp8 Tensor Core | Block scaled MMA with A={mxf4} x B={mxf4} with TN layouts |
|
||||
|tcgen05.mma.cta_group::[1\|2].kind::mxf4nvf4.block_scale.scale_vec_size::[2X\|4X] | 4x Hopper Fp8 Tensor Core | Block scaled MMA with A={mxf4} x B={mxf4} or A={nvf4} x B={nvf4} with TN layouts |
|
||||
|
||||
For more detailed information see [`tcgen05.mma` PTX documentation](https://docs.nvidia.com/cuda/parallel-thread-execution/index.html#tensorcore-5th-generation-family-instructions).
|
||||
|
||||
## New in Blackwell SM100
|
||||
|
||||
### Block Scaled GEMMs
|
||||
|
||||
Instructions with `kind` modifiers `mxf8f6f4`, `mxf4`, and `nvf4mxf4` perform matrix multiplication operations with scale
|
||||
factors of the form $D = C +( A \times SFA) * (B \times SFB)$. Scale factors are applied to GEMM-K dimension such that
|
||||
every 16 or 32 elements of $A$ and $B$ matrices in K dimension have an associated scale factor. For example, an $M\times K$,
|
||||
$A$ matrix has an associated $M \times \lceil K/32 \rceil$ SFA matrix; and an $N\times K$ $B$, matrix has an associated
|
||||
$N \times \lceil K/32 \rceil$ SFB matrix. For block scaled GEMMs, an entry of output D matrix is
|
||||
$D_{ij} = C_{ij} + \sum_{k} (A_{i,k} \times SFA_{i,k/SV}) \times (B_{j,k}\times SFB_{j,k/SV})$, in index notation, we SV is the scale factor vector size (16 or 32).
|
||||
Further details can be found in
|
||||
[PTX documentation on block scaling](https://docs.nvidia.com/cuda/parallel-thread-execution/index.html#tcgen05-block-scaling).
|
||||
|
||||
### Blackwell Narrow Precision Data Types
|
||||
|
||||
Narrow-precision `tcgen05.mma` instructions can operate on several 4, 6, and 8-bit data types. Blackwell MMAs can operate
|
||||
on five different 8-bit floating point values, of which only two (`float_ue8m0_t` and `float_ue4m3_t`) can be used as scale factor data types.
|
||||
There are two 6-bit floating point types and one 4-bit floating point data type.
|
||||
See [PTX documentation for narrow precision data types](https://docs.nvidia.com/cuda/parallel-thread-execution/index.html#alternate-floating-point-data-formats) for details.
|
||||
|
||||
**Blackwell Narrow Precision Data Types**
|
||||
| Data Type | Exponent Bits | Mantissa Bits | Signed | Bit Size |
|
||||
|-------------------|---------------|---------------|--------|----------|
|
||||
| float_e4m3_t |4 |3 | Yes | 8 |
|
||||
| float_e5m2_t |5 |2 | Yes | 8 |
|
||||
| float_e2m3_t |2 |3 | Yes | 6 |
|
||||
| float_e3m2_t |3 |2 | Yes | 6 |
|
||||
| float_e2m1_t |2 |1 | Yes | 4 |
|
||||
| float_ue8m0_t[^1] |8 |0 | No | 8 |
|
||||
| float_ue4m3_t[^1] |4 |3 | No | 8 |
|
||||
|
||||
[^1]: Only valid as scale factor data types.
|
||||
|
||||
Block scaled MMAs use `mx` and `nv` types which are a pair of float8_t, float6_t, float4_t with 2 of the scale factor data types with a predetermined scale factor vector size. `mx` types follow OCP specification (see [OCP Specification](https://www.opencompute.org/documents/ocp-microscaling-formats-mx-v1-0-spec-final-pdf)). The following types provided by CUTLASS can be used as inputs to collective builders to generate the block scaled kernels:
|
||||
|
||||
**Blackwell Block Scaled Narrow Precision Data Types**
|
||||
| Mx/Nv Data Type |Scale Factor Type | SF Vector Size | OCP Compliant |
|
||||
|----------------------------|------------------|----------------|---------------|
|
||||
| mx_float8_t\<Any F8type\> |float_ue8m0_t |32 | Yes |
|
||||
| mx_float6_t\<Any F6Type\> |float_ue8m0_t |32 | Yes |
|
||||
| mx_float4_t |float_ue8m0_t |32 | Yes |
|
||||
| nv_float4_t |float_ue4m3_t |16 | No |
|
||||
|
||||
## Layouts, Tensor Alignment Requirements to Target `tcgen05.mma` Instructions
|
||||
|
||||
Tables below list valid data type, and AB layout combinations. Note that the alignment is reported as number of elements. A and B matrix layouts are
|
||||
represented with T and N. T represents row-major layouts, and N represents column-major layouts. For instance, TN is
|
||||
row-major A matrix with column-major B matrix.
|
||||
|
||||
For legacy types (`tf32`, `f16`, `bf16`, `i8` and `u8`) alignment requirements for A and B matrices are the same as in Hopper.
|
||||
All four layouts (TT, NN, NT, TT) are supported for all legacy data types.
|
||||
|
||||
**Table 1: Valid Data Type, Alignment, and Layout Combinations For MMAs with Legacy Types** <a id="legacy_gemm_table" name="legacy_gemm_table"></a>
|
||||
| | A Type | B Type | AB Layout | A Alignment | B Alignment | Target tcgen05.mma.kind | Unit Test |
|
||||
|-------------------------------|------------|------------|----------------|-------------|-------------|-------------------------|-----------|
|
||||
|1 | tfloat32_t | tfloat32_t | TN, NN, NT, TT | 4 | 4 | tf32 | |
|
||||
|2 | half_t | half_t | TN, NN, NT, TT | 8 | 8 | f16 | [Unit tests](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/sm100_tensorop_gemm/f16_f16_void_f32.cu)|
|
||||
|3 | bfloat16_t | bfloat16_t | TN, NN, NT, TT | 8 | 8 | f16 | [Similar to half_t unit tests](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/sm100_tensorop_gemm/f16_f16_void_f32.cu)|
|
||||
|4 | int8_t | int8_t | TN, NN, NT, TT | 16 | 16 | i8 | [Unit tests](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/sm100_tensorop_gemm/s8_s8_void_s32.cu)|
|
||||
|5 | uint8_t | uint8_t | TN, NN, NT, TT | 16 | 16 | i8 | [Similar to int8_t unit tests](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/sm100_tensorop_gemm/s8_s8_void_s32.cu)|
|
||||
|
||||
For narrow precision Mmas, not all A/B type, and A/B layout combinations are supported by every `tcgen05.mma` instructions.
|
||||
Furthermore, tensor copy instructions for subbyte types impose additional alignment requirements while loading narrow-precision
|
||||
tensors from global memory to shared memory
|
||||
(see [PTX doc](https://docs.nvidia.com/cuda/parallel-thread-execution/index.html#data-movement-and-conversion-instructions-tensor-copy-restrictions) for details).
|
||||
|
||||
Below tables list valid layout, and alignment values for each A and B data type combination and their target `tcgen05.mma`
|
||||
instructions supported by CUTLASS.
|
||||
|
||||
**Table 2: Valid Data Type, Alignment, and Layout Combinations For Narrow Precision MMAs Without Block Scaling** <a id="non_bs_gemm_table" name="non_bs_gemm_table"></a>
|
||||
| | A Type | B Type | AB Layout | A Alignment | B Alignment | Target tcgen05.mma.kind | Unit Test |
|
||||
|-------------------------------|----------|----------|----------------|-------------|-------------|-------------------------|-----------|
|
||||
|[1](#nonbs_rows_1_2_3_6) | float4_t | float4_t | TN, NN, NT, TT | 128 | 128 | f8f6f4 | [TN unit tests](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/sm100_tensorop_gemm/narrow_precision/f6f4_f6f4_void_f32_tn_layout.cu) <br> [NT unit tests](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/sm100_tensorop_gemm/narrow_precision/f6f4_f6f4_void_f32_nt_layout.cu) <br> [NN unit tests](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/sm100_tensorop_gemm/narrow_precision/f6f4_f6f4_void_f32_nn_layout.cu) <br> [TT unit tests](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/sm100_tensorop_gemm/narrow_precision/f6f4_f6f4_void_f32_tt_layout.cu) |
|
||||
|[2](#nonbs_rows_1_2_3_6) | float4_t | float6_t | TN, NN, NT, TT | 128 | 128 | f8f6f4 | [TN unit tests](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/sm100_tensorop_gemm/narrow_precision/f6f4_f6f4_void_f32_tn_layout.cu) <br> [NT unit tests](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/sm100_tensorop_gemm/narrow_precision/f6f4_f6f4_void_f32_nt_layout.cu) <br> [NN unit tests](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/sm100_tensorop_gemm/narrow_precision/f6f4_f6f4_void_f32_nn_layout.cu) <br> [TT unit tests](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/sm100_tensorop_gemm/narrow_precision/f6f4_f6f4_void_f32_tt_layout.cu) |
|
||||
|[3](#nonbs_rows_1_2_3_6) | float6_t | float4_t | TN, NN, NT, TT | 128 | 128 | f8f6f4 | [TN unit tests](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/sm100_tensorop_gemm/narrow_precision/f6f4_f6f4_void_f32_tn_layout.cu) <br> [NT unit tests](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/sm100_tensorop_gemm/narrow_precision/f6f4_f6f4_void_f32_nt_layout.cu) <br> [NN unit tests](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/sm100_tensorop_gemm/narrow_precision/f6f4_f6f4_void_f32_nn_layout.cu) <br> [TT unit tests](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/sm100_tensorop_gemm/narrow_precision/f6f4_f6f4_void_f32_tt_layout.cu) |
|
||||
|[4](#nonbs_rows_4_7) | float4_t | float8_t | TN, NN, NT, TT | 128 | 16 | f8f6f4 | [TN unit tests](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/sm100_tensorop_gemm/narrow_precision/f6f4_f8_void_f32_tn_layout.cu) <br> [NT unit tests](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/sm100_tensorop_gemm/narrow_precision/f6f4_f8_void_f32_nt_layout.cu) |
|
||||
|[5](#nonbs_rows_5_8) | float8_t | float4_t | TN, NN, NT, TT | 16 | 128 | f8f6f4 | [TN unit tests](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/sm100_tensorop_gemm/narrow_precision/f8_f6f4_void_f32_tn_layout.cu) <br> [NT unit tests](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/sm100_tensorop_gemm/narrow_precision/f8_f6f4_void_f32_nt_layout.cu) |
|
||||
|[6](#nonbs_rows_1_2_3_6) | float6_t | float6_t | TN, NN, NT, TT | 128 | 128 | f8f6f4 | [TN unit tests](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/sm100_tensorop_gemm/narrow_precision/f6f4_f6f4_void_f32_tn_layout.cu) <br> [NT unit tests](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/sm100_tensorop_gemm/narrow_precision/f6f4_f6f4_void_f32_nt_layout.cu) <br> [NN unit tests](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/sm100_tensorop_gemm/narrow_precision/f6f4_f6f4_void_f32_nn_layout.cu) <br> [TT unit tests](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/sm100_tensorop_gemm/narrow_precision/f6f4_f6f4_void_f32_tt_layout.cu) |
|
||||
|[7](#nonbs_rows_4_7) | float6_t | float8_t | TN, NN, NT, TT | 128 | 16 | f8f6f4 | [TN unit tests](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/sm100_tensorop_gemm/narrow_precision/f6f4_f8_void_f32_tn_layout.cu) <br> [NT unit tests](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/sm100_tensorop_gemm/narrow_precision/f6f4_f8_void_f32_nt_layout.cu) |
|
||||
|[8](#nonbs_rows_5_8) | float8_t | float6_t | TN, NN, NT, TT | 16 | 128 | f8f6f4 | [TN unit tests](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/sm100_tensorop_gemm/narrow_precision/f8_f6f4_void_f32_tn_layout.cu) <br> [NT unit tests](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/sm100_tensorop_gemm/narrow_precision/f8_f6f4_void_f32_nt_layout.cu) |
|
||||
|[9](#nonbs_rows_9) | float8_t | float8_t | TN, NN, NT, TT | 16 | 16 | f8f6f4 | [Unit tests](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/sm100_tensorop_gemm/f8_f8_void_f32.cu)|
|
||||
|
||||
|
||||
**Table 3: Valid Data Type, Alignment, and Layout Combinations for Block Scaled Narrow Precision MMAs** <a id="bs_gemm_table" name="bs_gemm_table"></a>
|
||||
| | A Type | B Type | AB Layout | A Alignment | B Alignment | Target tcgen05.mma.kind |Unit Test|
|
||||
|-------------------------|-------------|-------------|----------------|-------------|-------------|-------------------------|------|
|
||||
|[1](#bs_rows_1) | nv_float4_t | nv_float4_t | TN | 32 | 32 | mxf4nvf4 |[TN unit tests](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/sm100_blockscaled_tensorop_gemm/nvf4_nvf4_bf16_bf16.cu)|
|
||||
|[2](#bs_rows_2) | mx_float4_t | mx_float4_t | TN | 32 | 32 | mxf4, mxf4nvf4 |[TN unit tests](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/sm100_blockscaled_tensorop_gemm/mxf4_mxf4_void_f16_tn_layout.cu)|
|
||||
|[3](#bs_rows_3) | mx_float4_t | mx_float4_t | TN, NN, NT, TT | 128 | 128 | mxf8f6f4 |[NT unit tests](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/sm100_blockscaled_tensorop_gemm/mxf4_mxf4_void_f16_nt_layout.cu)|
|
||||
|[4](#bs_rows_4_5_7_8_10) | mx_float4_t | mx_float6_t | TN, NN, NT, TT | 128 | 128 | mxf8f6f4 |[TN unit tests](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/sm100_blockscaled_tensorop_gemm/mxf4_mxf6_f32_f16_tn_layout.cu)<br>[NT unit tests](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/sm100_blockscaled_tensorop_gemm/mxf4_mxf6_f32_f16_nt_layout.cu)|
|
||||
|[5](#bs_rows_4_5_7_8_10) | mx_float6_t | mx_float4_t | TN, NN, NT, TT | 128 | 128 | mxf8f6f4 |[TN unit tests](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/sm100_blockscaled_tensorop_gemm/mxf6_mxf4_f16_f16_tn_layout.cu)<br>[NT unit tests](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/sm100_blockscaled_tensorop_gemm/mxf6_mxf4_f16_f16_nt_layout.cu)|
|
||||
|[6](#bs_rows_6_9_11) | mx_float4_t | mx_float8_t | TN, NN, NT, TT | 128 | 16 | mxf8f6f4 |[TN unit tests](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/sm100_blockscaled_tensorop_gemm/mxf4_mxf8_bf16_bf16_tn_layout.cu)<br>[NT unit tests](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/sm100_blockscaled_tensorop_gemm/mxf4_mxf8_bf16_bf16_nt_layout.cu)|
|
||||
|[7](#bs_rows_4_5_7_8_10) | mx_float8_t | mx_float4_t | TN, NN, NT, TT | 16 | 128 | mxf8f6f4 |[TN unit tests](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/sm100_blockscaled_tensorop_gemm/mxf8_mxf4_f16_bf16_tn_layout.cu)<br>[NT unit tests](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/sm100_blockscaled_tensorop_gemm/mxf8_mxf4_f16_bf16_nt_layout.cu)|
|
||||
|[8](#bs_rows_4_5_7_8_10) | mx_float6_t | mx_float6_t | TN, NN, NT, TT | 128 | 128 | mxf8f6f4 |[TN unit tests](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/sm100_blockscaled_tensorop_gemm/mxf6_mxf6_void_bf16_tn_layout.cu)<br>[NT unit tests](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/sm100_blockscaled_tensorop_gemm/mxf6_mxf6_void_bf16_nt_layout.cu)|
|
||||
|[9](#bs_rows_6_9_11) | mx_float6_t | mx_float8_t | TN, NN, NT, TT | 128 | 16 | mxf8f6f4 |[TN unit tests](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/sm100_blockscaled_tensorop_gemm/mxf6_mxf8_void_f32_tn_layout.cu)<br>[NT unit tests](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/sm100_blockscaled_tensorop_gemm/mxf6_mxf8_void_f32_nt_layout.cu)|
|
||||
|[10](#bs_rows_4_5_7_8_10)| mx_float8_t | mx_float6_t | TN, NN, NT, TT | 16 | 128 | mxf8f6f4 |[TN unit tests](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/sm100_blockscaled_tensorop_gemm/mxf8_mxf6_f16_f8_tn_layout.cu)<br>[NT unit tests](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/sm100_blockscaled_tensorop_gemm/mxf8_mxf6_f16_f8_nt_layout.cu)|
|
||||
|[11](#bs_rows_6_9_11) | mx_float8_t | mx_float8_t | TN, NN, NT, TT | 16 | 16 | mxf8f6f4 |[TN unit tests](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/sm100_blockscaled_tensorop_gemm/mxf8_mxf8_void_f8_tn_layout.cu.cu)<br>[NT unit tests](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/sm100_blockscaled_tensorop_gemm/mxf8_mxf8_void_f8_nt_layout.cu)|
|
||||
|
||||
## MMA tile shapes supported
|
||||
|
||||
The alignment restrictions also limit the options for Mma Tile Shapes. Tables below list the supported/valid `MmaTileShape`,
|
||||
Layout, and Dispatch Policy combinations for each row of [Table 1](#legacy_gemm_table), [Table 2](#non_bs_gemm_table), and [Table 3](#bs_gemm_table).
|
||||
|
||||
**Table 4: Valid Tile Shapes and Dispatch Policies for lagacy types (All rows of Table 1)** <a id="legacy_rows" name="legacy_rows"></a>
|
||||
| 1/2 SM | Mma Tile Shape | TN | TT | NT | NN | Dispatch Policy |
|
||||
|--------|------------------|----|----|----|----|------------------------------------|
|
||||
| 1SM | 64x64x(4*MMA-K) | Y | Y | Y | Y | `KernelTmaWarpSpecialized1SmSm100` |
|
||||
| 1SM | 64x128x(4*MMA-K) | Y | Y | Y | Y | `KernelTmaWarpSpecialized1SmSm100` |
|
||||
| 1SM | 64x192x(4*MMA-K) | Y | Y | Y | Y | `KernelTmaWarpSpecialized1SmSm100` |
|
||||
| 1SM | 64x256x(4*MMA-K) | Y | Y | Y | Y | `KernelTmaWarpSpecialized1SmSm100` |
|
||||
| 1SM | 128x64x(4*MMA-K) | Y | Y | Y | Y | `KernelTmaWarpSpecialized1SmSm100` |
|
||||
| 1SM | 128x128x(4*MMA-K)| Y | Y | Y | Y | `KernelTmaWarpSpecialized1SmSm100` |
|
||||
| 1SM | 128x192x(4*MMA-K)| Y | Y | Y | Y | `KernelTmaWarpSpecialized1SmSm100` |
|
||||
| 1SM | 128x256x(4*MMA-K)| Y | Y | Y | Y | `KernelTmaWarpSpecialized1SmSm100` |
|
||||
| 2SM | 128x64x(4*MMA-K) | Y | Y | Y | Y | `KernelTmaWarpSpecialized2SmSm100` |
|
||||
| 2SM | 128x128x(4*MMA-K)| Y | Y | Y | Y | `KernelTmaWarpSpecialized2SmSm100` |
|
||||
| 2SM | 128x192x(4*MMA-K)| Y | Y | Y | Y | `KernelTmaWarpSpecialized2SmSm100` |
|
||||
| 2SM | 128x256x(4*MMA-K)| Y | Y | Y | Y | `KernelTmaWarpSpecialized2SmSm100` |
|
||||
| 2SM | 256x64x(4*MMA-K) | Y | Y | Y | Y | `KernelTmaWarpSpecialized2SmSm100` |
|
||||
| 2SM | 256x128x(4*MMA-K)| Y | Y | Y | Y | `KernelTmaWarpSpecialized2SmSm100` |
|
||||
| 2SM | 256x192x(4*MMA-K)| Y | Y | Y | Y | `KernelTmaWarpSpecialized2SmSm100` |
|
||||
| 2SM | 256x256x(4*MMA-K)| Y | Y | Y | Y | `KernelTmaWarpSpecialized2SmSm100` |
|
||||
|
||||
**Table 5: Valid Tile Shapes and Dispatch Policies for {float4_t, float6_t} x {float4_t, float6_t} (Rows 1,2,3,6 of Table 2)** <a id="nonbs_rows_1_2_3_6" name="nonbs_rows_1_2_3_6"></a>
|
||||
|
||||
| 1/2 SM | Mma Tile Shape | TN | TT | NT | NN | Dispatch Policy |
|
||||
|--------|----------------|----|----|----|----|------------------------------------|
|
||||
| 1SM | 64x64x128 | Y | N | N | N | `KernelTmaWarpSpecialized1SmSm100` |
|
||||
| 1SM | 64x128x128 | Y | Y | N | N | `KernelTmaWarpSpecialized1SmSm100` |
|
||||
| 1SM | 64x192x128 | Y | N | N | N | `KernelTmaWarpSpecialized1SmSm100` |
|
||||
| 1SM | 64x256x128 | Y | Y | N | N | `KernelTmaWarpSpecialized1SmSm100` |
|
||||
| 1SM | 128x64x128 | Y | N | N | Y | `KernelTmaWarpSpecialized1SmSm100` |
|
||||
| 1SM | 128x128x128 | Y | Y | Y | Y | `KernelTmaWarpSpecialized1SmSm100` |
|
||||
| 1SM | 128x192x128 | Y | N | N | Y | `KernelTmaWarpSpecialized1SmSm100` |
|
||||
| 1SM | 128x256x128 | Y | Y | Y | Y | `KernelTmaWarpSpecialized1SmSm100` |
|
||||
| 2SM | 128x64x128 | Y | N | N | N | `KernelTmaWarpSpecialized2SmSm100` |
|
||||
| 2SM | 128x128x128 | Y | N | N | N | `KernelTmaWarpSpecialized2SmSm100` |
|
||||
| 2SM | 128x192x128 | Y | N | N | N | `KernelTmaWarpSpecialized2SmSm100` |
|
||||
| 2SM | 128x256x128 | Y | Y | N | N | `KernelTmaWarpSpecialized2SmSm100` |
|
||||
| 2SM | 256x64x128 | Y | N | N | Y | `KernelTmaWarpSpecialized2SmSm100` |
|
||||
| 2SM | 256x128x128 | Y | N | N | Y | `KernelTmaWarpSpecialized2SmSm100` |
|
||||
| 2SM | 256x192x128 | Y | N | N | Y | `KernelTmaWarpSpecialized2SmSm100` |
|
||||
| 2SM | 256x256x128 | Y | Y | Y | Y | `KernelTmaWarpSpecialized2SmSm100` |
|
||||
|
||||
**Table 6: Valid Tile Shapes and Dispatch Policies for float8_t x {float4_t, float6_t} (Rows 5,8 of Table 2)** <a id="nonbs_rows_5_8" name="nonbs_rows_5_8"></a>
|
||||
|
||||
| 1/2 SM | Mma Tile Shape | TN | TT | NT | NN | Dispatch Policy |
|
||||
|--------|----------------|----|----|----|----|------------------------------------|
|
||||
| 1SM | 64x64x128 | Y | N | N | Y | `KernelTmaWarpSpecialized1SmSm100` |
|
||||
| 1SM | 64x128x128 | Y | Y | Y | Y | `KernelTmaWarpSpecialized1SmSm100` |
|
||||
| 1SM | 64x192x128 | Y | N | N | Y | `KernelTmaWarpSpecialized1SmSm100` |
|
||||
| 1SM | 64x256x128 | Y | Y | Y | Y | `KernelTmaWarpSpecialized1SmSm100` |
|
||||
| 1SM | 128x64x128 | Y | N | N | Y | `KernelTmaWarpSpecialized1SmSm100` |
|
||||
| 1SM | 128x128x128 | Y | Y | Y | Y | `KernelTmaWarpSpecialized1SmSm100` |
|
||||
| 1SM | 128x192x128 | Y | N | N | Y | `KernelTmaWarpSpecialized1SmSm100` |
|
||||
| 1SM | 128x256x128 | Y | Y | Y | Y | `KernelTmaWarpSpecialized1SmSm100` |
|
||||
| 2SM | 128x64x128 | Y | N | N | Y | `KernelTmaWarpSpecialized2SmSm100` |
|
||||
| 2SM | 128x128x128 | Y | N | N | Y | `KernelTmaWarpSpecialized2SmSm100` |
|
||||
| 2SM | 128x192x128 | Y | N | N | Y | `KernelTmaWarpSpecialized2SmSm100` |
|
||||
| 2SM | 128x256x128 | Y | Y | Y | Y | `KernelTmaWarpSpecialized2SmSm100` |
|
||||
| 2SM | 256x64x128 | Y | N | N | Y | `KernelTmaWarpSpecialized2SmSm100` |
|
||||
| 2SM | 256x128x128 | Y | N | N | Y | `KernelTmaWarpSpecialized2SmSm100` |
|
||||
| 2SM | 256x192x128 | Y | N | N | Y | `KernelTmaWarpSpecialized2SmSm100` |
|
||||
| 2SM | 256x256x128 | Y | Y | Y | Y | `KernelTmaWarpSpecialized2SmSm100` |
|
||||
|
||||
**Table 7: Valid Tile Shapes and Dispatch Policies for {float4_t, float6_t} x float8_t (Rows 4,7 of Table 2)** <a id="nonbs_rows_4_7" name="nonbs_rows_4_7"></a>
|
||||
|
||||
| 1/2 SM | Mma Tile Shape | TN | TT | NT | NN | Dispatch Policy |
|
||||
|--------|----------------|----|----|----|----|------------------------------------|
|
||||
| 1SM | 64x64x128 | Y | Y | N | N | `KernelTmaWarpSpecialized1SmSm100` |
|
||||
| 1SM | 64x128x128 | Y | Y | N | N | `KernelTmaWarpSpecialized1SmSm100` |
|
||||
| 1SM | 64x192x128 | Y | Y | N | N | `KernelTmaWarpSpecialized1SmSm100` |
|
||||
| 1SM | 64x256x128 | Y | Y | N | N | `KernelTmaWarpSpecialized1SmSm100` |
|
||||
| 1SM | 128x64x128 | Y | Y | Y | Y | `KernelTmaWarpSpecialized1SmSm100` |
|
||||
| 1SM | 128x128x128 | Y | Y | Y | Y | `KernelTmaWarpSpecialized1SmSm100` |
|
||||
| 1SM | 128x192x128 | Y | Y | Y | Y | `KernelTmaWarpSpecialized1SmSm100` |
|
||||
| 1SM | 128x256x128 | Y | Y | Y | Y | `KernelTmaWarpSpecialized1SmSm100` |
|
||||
| 2SM | 128x64x128 | Y | Y | Y | Y | `KernelTmaWarpSpecialized2SmSm100` |
|
||||
| 2SM | 128x128x128 | Y | Y | N | N | `KernelTmaWarpSpecialized2SmSm100` |
|
||||
| 2SM | 128x192x128 | Y | Y | N | N | `KernelTmaWarpSpecialized2SmSm100` |
|
||||
| 2SM | 128x256x128 | Y | Y | N | N | `KernelTmaWarpSpecialized2SmSm100` |
|
||||
| 2SM | 256x64x128 | Y | Y | Y | Y | `KernelTmaWarpSpecialized2SmSm100` |
|
||||
| 2SM | 256x128x128 | Y | Y | Y | Y | `KernelTmaWarpSpecialized2SmSm100` |
|
||||
| 2SM | 256x192x128 | Y | Y | Y | Y | `KernelTmaWarpSpecialized2SmSm100` |
|
||||
| 2SM | 256x256x128 | Y | Y | Y | Y | `KernelTmaWarpSpecialized2SmSm100` |
|
||||
|
||||
**Table 8: Valid Tile Shapes and Dispatch Policies for float8_t x float8_t (Row 9 of Table 2)** <a id="nonbs_rows_9" name="nonbs_rows_9"></a>
|
||||
|
||||
| 1/2 SM | Mma Tile Shape | TN | TT | NT | NN | Dispatch Policy |
|
||||
|--------|----------------|----|----|----|----|------------------------------------|
|
||||
| 1SM | 64x64x128 | Y | Y | Y | Y | `KernelTmaWarpSpecialized1SmSm100` |
|
||||
| 1SM | 64x128x128 | Y | Y | Y | Y | `KernelTmaWarpSpecialized1SmSm100` |
|
||||
| 1SM | 64x192x128 | Y | Y | Y | Y | `KernelTmaWarpSpecialized1SmSm100` |
|
||||
| 1SM | 64x256x128 | Y | Y | Y | Y | `KernelTmaWarpSpecialized1SmSm100` |
|
||||
| 1SM | 128x64x128 | Y | Y | Y | Y | `KernelTmaWarpSpecialized1SmSm100` |
|
||||
| 1SM | 128x128x128 | Y | Y | Y | Y | `KernelTmaWarpSpecialized1SmSm100` |
|
||||
| 1SM | 128x192x128 | Y | Y | Y | Y | `KernelTmaWarpSpecialized1SmSm100` |
|
||||
| 1SM | 128x256x128 | Y | Y | Y | Y | `KernelTmaWarpSpecialized1SmSm100` |
|
||||
| 2SM | 128x64x128 | Y | Y | Y | Y | `KernelTmaWarpSpecialized2SmSm100` |
|
||||
| 2SM | 128x128x128 | Y | Y | Y | Y | `KernelTmaWarpSpecialized2SmSm100` |
|
||||
| 2SM | 128x192x128 | Y | Y | Y | Y | `KernelTmaWarpSpecialized2SmSm100` |
|
||||
| 2SM | 128x256x128 | Y | Y | Y | Y | `KernelTmaWarpSpecialized2SmSm100` |
|
||||
| 2SM | 256x64x128 | Y | Y | Y | Y | `KernelTmaWarpSpecialized2SmSm100` |
|
||||
| 2SM | 256x128x128 | Y | Y | Y | Y | `KernelTmaWarpSpecialized2SmSm100` |
|
||||
| 2SM | 256x192x128 | Y | Y | Y | Y | `KernelTmaWarpSpecialized2SmSm100` |
|
||||
| 2SM | 256x256x128 | Y | Y | Y | Y | `KernelTmaWarpSpecialized2SmSm100` |
|
||||
|
||||
|
||||
**Table 9: Valid Tile Shapes for nv_float4_t x nv_float4_t (Row 1 of Table 3)** <a id="bs_rows_1" name="bs_rows_1"></a>
|
||||
| 1/2 SM | Mma Tile Shape | TN | TT | NT | NN | Dispatch Policy |
|
||||
|--------|---------------|----|----|----|----|----------------------------------------|
|
||||
| 1SM | 128x128x256 | Y | N | N | N | `KernelTmaWarpSpecialized1SmNvf4Sm100` |
|
||||
| 1SM | 128x192x256 | Y | N | N | N | `KernelTmaWarpSpecialized1SmNvf4Sm100` |
|
||||
| 1SM | 128x256x256 | Y | N | N | N | `KernelTmaWarpSpecialized1SmNvf4Sm100` |
|
||||
| 2SM | 256x128x256 | Y | N | N | N | `KernelTmaWarpSpecialized2SmNvf4Sm100` |
|
||||
| 2SM | 256x192x256 | Y | N | N | N | `KernelTmaWarpSpecialized2SmNvf4Sm100` |
|
||||
| 2SM | 256x256x256 | Y | N | N | N | `KernelTmaWarpSpecialized2SmNvf4Sm100` |
|
||||
|
||||
**Table 10: Valid Tile Shapes and Dispatch Policies for mx_float4_t x mx_float4_t (Row 2 of Table 3)** <a id="bs_rows_2" name="bs_rows_2"></a>
|
||||
| 1/2 SM | Mma Tile Shape | TN | TT | NT | NN | Dispatch Policy |
|
||||
|--------|---------------|----|----|----|----|----------------------------------------|
|
||||
| 1SM | 128x128x256 | Y | N | N | N | `KernelTmaWarpSpecialized1SmMxf4Sm100` |
|
||||
| 1SM | 128x192x256 | Y | N | N | N | `KernelTmaWarpSpecialized1SmMxf4Sm100` |
|
||||
| 1SM | 128x256x256 | Y | N | N | N | `KernelTmaWarpSpecialized1SmMxf4Sm100` |
|
||||
| 2SM | 256x128x256 | Y | N | N | N | `KernelTmaWarpSpecialized2SmMxf4Sm100` |
|
||||
| 2SM | 256x192x256 | Y | N | N | N | `KernelTmaWarpSpecialized2SmMxf4Sm100` |
|
||||
| 2SM | 256x256x256 | Y | N | N | N | `KernelTmaWarpSpecialized2SmMxf4Sm100` |
|
||||
|
||||
**Table 11: Valid Tile Shapes and Dispatch Policies for mx_float4_t x mx_float4_t (Row 3 of Table 3)** <a id="bs_rows_3" name="bs_rows_3"></a>
|
||||
| 1/2 SM | Mma Tile Shape | TN | TT | NT | NN | Dispatch Policy |
|
||||
|--------|---------------|----|----|----|----|--------------------------------------------|
|
||||
| 1SM | 128x128x128 | Y | Y | Y | Y | `KernelTmaWarpSpecialized1SmMxf8f6f4Sm100` |
|
||||
| 1SM | 128x192x128 | Y | N | N | Y | `KernelTmaWarpSpecialized1SmMxf8f6f4Sm100` |
|
||||
| 1SM | 128x256x128 | Y | Y | Y | Y | `KernelTmaWarpSpecialized1SmMxf8f6f4Sm100` |
|
||||
| 2SM | 256x128x128 | Y | N | N | Y | `KernelTmaWarpSpecialized2SmMxf8f6f4Sm100` |
|
||||
| 2SM | 256x192x128 | Y | N | N | Y | `KernelTmaWarpSpecialized2SmMxf8f6f4Sm100` |
|
||||
| 2SM | 256x256x128 | Y | Y | Y | Y | `KernelTmaWarpSpecialized2SmMxf8f6f4Sm100` |
|
||||
|
||||
**Table 12: Valid Tile Shapes and Dispatch Policies for {mx_float4_t, mx_float6_t, mx_float8_t} x {mx_float4_t, mx_float6_t} (Rows 4, 5, 7, 8, 10 of Table 3)** <a id="bs_rows_4_5_7_8_10" name="bs_rows_4_5_7_8_10"></a>
|
||||
| 1/2 SM | Mma Tile Shape | TN | TT | NT | NN | Dispatch Policy |
|
||||
|--------|---------------|----|----|----|----|--------------------------------------------|
|
||||
| 1SM | 128x128x128 | Y | Y | Y | Y | `KernelTmaWarpSpecialized1SmMxf8f6f4Sm100` |
|
||||
| 1SM | 128x192x128 | Y | N | N | Y | `KernelTmaWarpSpecialized1SmMxf8f6f4Sm100` |
|
||||
| 1SM | 128x256x128 | Y | Y | Y | Y | `KernelTmaWarpSpecialized1SmMxf8f6f4Sm100` |
|
||||
| 2SM | 256x128x128 | Y | N | N | Y | `KernelTmaWarpSpecialized2SmMxf8f6f4Sm100` |
|
||||
| 2SM | 256x192x128 | Y | N | N | Y | `KernelTmaWarpSpecialized2SmMxf8f6f4Sm100` |
|
||||
| 2SM | 256x256x128 | Y | Y | Y | Y | `KernelTmaWarpSpecialized2SmMxf8f6f4Sm100` |
|
||||
|
||||
**Table 13: Valid Tile Shapes and Dispatch Policies for {mx_float4_t, mx_float6_t, mx_float8_t} x mx_float8_t (Rows 6, 9, 11 of Table 3)** <a id="bs_rows_6_9_11" name="bs_rows_6_9_11"></a>
|
||||
| 1/2 SM | Mma Tile Shape | TN| TT | NT | NN | Dispatch Policy |
|
||||
|--------|---------------|----|----|----|----|--------------------------------------------|
|
||||
| 1SM | 128x128x128 | Y | Y | Y | Y | `KernelTmaWarpSpecialized1SmMxf8f6f4Sm100` |
|
||||
| 1SM | 128x192x128 | Y | Y | Y | Y | `KernelTmaWarpSpecialized1SmMxf8f6f4Sm100` |
|
||||
| 1SM | 128x256x128 | Y | Y | Y | Y | `KernelTmaWarpSpecialized1SmMxf8f6f4Sm100` |
|
||||
| 2SM | 256x128x128 | Y | Y | Y | Y | `KernelTmaWarpSpecialized2SmMxf8f6f4Sm100` |
|
||||
| 2SM | 256x192x128 | Y | Y | Y | Y | `KernelTmaWarpSpecialized2SmMxf8f6f4Sm100` |
|
||||
| 2SM | 256x256x128 | Y | Y | Y | Y | `KernelTmaWarpSpecialized2SmMxf8f6f4Sm100` |
|
||||
|
||||
## Epilogue config supported
|
||||
|
||||
**Table 14: Epilogue Dispatch Policy** <a id="epi_dispatch" name="epi_dispatch"></a>
|
||||
| 1/2 SM | Epilogue Dispatch Policy |
|
||||
|--------|------------------------------------------|
|
||||
| 1SM | cutlass::epilogue::TmaWarpSpecialized1Sm |
|
||||
| 1SM | cutlass::epilogue::NoSmemWarpSpecialized1Sm |
|
||||
| 2SM | cutlass::epilogue::TmaWarpSpecialized2Sm |
|
||||
| 2SM | cutlass::epilogue::NoSmemWarpSpecialized2Sm |
|
||||
|
||||
**Table 15: Epilogue PerSmTileShape_MNK** <a id="epi_persmtileshape" name="epi_persmtileshape"></a>
|
||||
| 1/2 SM | MMA tile Shape | PerSmTileShape_MNK |
|
||||
|--------|--------------------------|-------------------------|
|
||||
| 1SM | 64x64xMMA_TileShape_K | 64x64xMMA_TileShape_K |
|
||||
| 1SM | 64x128xMMA_TileShape_K | 64x128xMMA_TileShape_K |
|
||||
| 1SM | 64x192xMMA_TileShape_K | 64x192xMMA_TileShape_K |
|
||||
| 1SM | 64x256xMMA_TileShape_K | 64x256xMMA_TileShape_K |
|
||||
| 1SM | 128x64xMMA_TileShape_K | 128x64xMMA_TileShape_K |
|
||||
| 1SM | 128x128xMMA_TileShape_K | 128x128xMMA_TileShape_K |
|
||||
| 1SM | 128x192xMMA_TileShape_K | 128x192xMMA_TileShape_K |
|
||||
| 1SM | 128x256xMMA_TileShape_K | 128x256xMMA_TileShape_K |
|
||||
| 2SM | 128x64xMMA_TileShape_K | 64x64xMMA_TileShape_K |
|
||||
| 2SM | 128x128xMMA_TileShape_K | 64x128xMMA_TileShape_K |
|
||||
| 2SM | 128x192xMMA_TileShape_K | 64x192xMMA_TileShape_K |
|
||||
| 2SM | 128x256xMMA_TileShape_K | 64x256xMMA_TileShape_K |
|
||||
| 2SM | 256x64xMMA_TileShape_K | 128x64xMMA_TileShape_K |
|
||||
| 2SM | 256x128xMMA_TileShape_K | 128x128xMMA_TileShape_K |
|
||||
| 2SM | 256x192xMMA_TileShape_K | 128x192xMMA_TileShape_K |
|
||||
| 2SM | 256x256xMMA_TileShape_K | 128x256xMMA_TileShape_K |
|
||||
|
||||
MMA_TileShape_K is is generally 4 * MMA-Instruction-K. It depends on the config we defined in MMA tile shapes supported section.
|
||||
|
||||
### Auto Kernel Dispatch Policies
|
||||
|
||||
In addition to direct dispatch policies listed above, the user can also use auto policies for both non-block scaled narrow-precision
|
||||
GEMMs, and block scaled narrow-precision GEMMs.
|
||||
|
||||
CUTLASS will do its best to find the most efficient kernel for given parameters, however, the preferred method for building
|
||||
these kernels is to use direct kernel dispatch policies shown in the above tables.
|
||||
|
||||
* `cutlass::gemm::collective::KernelScheduleAuto`: For a given Mma Tile Size, data type and layout combinations choose instr kind (mxf8f6f4, mxf4, nvf4mxf4) and 1/2 SM `tcgen05.mma`.
|
||||
* `KernelTmaWarpSpecialized1SmBlockScaledSm100`: Use 1 SM `tcgen05.mma` instruction and choose instr kind (mxf8f6f4, mxf4, nvf4mxf4) automatically.
|
||||
* `KernelTmaWarpSpecialized2SmBlockScaledSm100`: Use 2 SM `tcgen05.mma` instruction and choose instr kind (mxf8f6f4, mxf4, nvf4mxf4) automatically.
|
||||
|
||||
Similarly for epilogues, we can use `cutlass::epilogue::collective::EpilogueScheduleAuto`.
|
||||
|
||||
## Building a Block Scaled Kernel <a id="detailed_blockscale_example" name="detailed_blockscale_example"></a>
|
||||
|
||||
For non-blockscaled dense GEMM refer to [quick start page](quickstart.md#instantiating-a-blackwell-sm100-gemm-kernel). An example dense GEMM can be found:
|
||||
1. [Blackwell FP16 GEMM example](https://github.com/NVIDIA/cutlass/tree/main/examples/70_blackwell_gemm/).
|
||||
|
||||
Narrow precision and block scaled narrow precision kernels can be built using CUTLASS 3.x collective builder interface
|
||||
(as described in [CUTLASS 3.0 GEMM API](gemm_api_3x.md#cutlass-30-gemm-api)). However, special attention needs to be given to
|
||||
A and B matrix layouts, alignment requirements, and dispatch policies to obtain a functionally correct and performant kernel
|
||||
which are listed above.
|
||||
|
||||
Several examples of block scaled kernels can be found in [examples/72_blackwell_narrow_precision_gemm](https://github.com/NVIDIA/cutlass/tree/main/examples/72_blackwell_narrow_precision_gemm/) directory:
|
||||
1. [NVF4 Gemm with block scaling](https://github.com/NVIDIA/cutlass/tree/main/examples/72_blackwell_narrow_precision_gemm/72a_blackwell_nvfp4_bf16_gemm.cu)
|
||||
2. [NVF4 Gemm with block scaling and NVF4 output matrix](https://github.com/NVIDIA/cutlass/tree/main/examples/72_blackwell_narrow_precision_gemm/72b_blackwell_nvfp4_nvfp4_gemm.cu)
|
||||
3. [Mixed precision Nvf4 x Mxf8 GEMM with block scaling](https://github.com/NVIDIA/cutlass/tree/main/examples/72_blackwell_narrow_precision_gemm/72c_blackwell_mixed_mxfp8_bf16_gemm.cu)
|
||||
|
||||
Collective builder interface expects the same arguments as any other CUTLASS 3.x kernels as described
|
||||
[here](gemm_api_3x.md#collective-builder-for-collectivemmas) with a small difference for Collective MMA builder interface.
|
||||
As in all Blackwell kernels, the `TileShape_MNK` argument expects the `MmaTileShape_MNK` which is the tile shape needed
|
||||
by 1 or 2 SM `tcgen05.mma` instructions.
|
||||
|
||||
Let's consider building a block scaled GEMM where the A matrix is of type `mx_float4_t` and column-major (N), and the
|
||||
B matrix is of type `mx_float4_t` and row-major (T). We first need to describe the A and B tensors, and find the
|
||||
instruction that can support the selected A and B type and layout pair. Then, we will choose the performance parameters.
|
||||
|
||||
The skeleton C++ code is shown below:
|
||||
|
||||
```cpp
|
||||
///////////////////////////////////////////////////////////
|
||||
// Mainloop Builder Setup
|
||||
///////////////////////////////////////////////////////////
|
||||
|
||||
///////////////////////////////////////////
|
||||
// 1. Describe A and B tensors
|
||||
///////////////////////////////////////////
|
||||
using ElementA = // TBD
|
||||
constexpr int AlignA = // TBD
|
||||
using GmemLayoutA = // TBD
|
||||
using ElementB = // TBD
|
||||
constexpr int AlignB = // TBD
|
||||
using GmemLayoutB = // TBD
|
||||
|
||||
// Mma's accumulator type
|
||||
using ElementAccumulator = float; // Always float for block scaled tcgen05.mma instructions
|
||||
|
||||
//////////////////////////////////////////
|
||||
// 2. Choose Performance Parameters
|
||||
//////////////////////////////////////////
|
||||
|
||||
// Tile and cluster shapes
|
||||
// Collective MMA takes tile shape of the MMA operation as input
|
||||
using KernelMainloopPolicy = // TBD
|
||||
using MmaTileShape_MNK = // TBD
|
||||
using ClusterShape_MNK = // TBD
|
||||
|
||||
using CollectiveMainloop = typename cutlass::gemm::collective::CollectiveBuilder<
|
||||
cutlass::arch::Sm100, cutlass::arch::OpClassBlockScaledTensorOp, // Arch and Tensorop spec
|
||||
ElementA, GmemLayoutA, AlignA, // A tensor elem type, layout and alignment requirement
|
||||
ElementB, GmemLayoutB, AlignB, // B tensor elem type, layout and alignment requirement
|
||||
ElementAccumulator, // Mma instruction accumulator type
|
||||
MmaTileShape_MNK, ClusterShape_MNK, // Mma instruction tile shape, cluster shape
|
||||
// Epilogue's SMEM usage that needs to be subtracted from overall SMEM capacity
|
||||
cutlass::gemm::collective::StageCountAutoCarveout<static_cast<int>(sizeof(typename CollectiveEpilogue::SharedStorage))>,
|
||||
KernelMainloopPolicy // Kernel schedule policy.
|
||||
// Auto or using targeted scheduling policy
|
||||
>::CollectiveOp;
|
||||
```
|
||||
|
||||
From the valid type and layout combinations [Table 3](#bs_gemm_table), we see that only **row 3** can support `mx_float4_t`x`mx_float4_t`
|
||||
combination with NT layout. As a result, we need to use the `tcgen05.mma.kind:mxf8f6f4` instruction. Additionally, in order
|
||||
to use `tcgen05.mma.kind:mxf8f6f4`, we see that A and B tensors both should be 128-element aligned.
|
||||
Thus, we can describe A and B tensors as follows:
|
||||
|
||||
```cpp
|
||||
///////////////////////////////////////////////////////////
|
||||
// Mainloop Builder Setup
|
||||
///////////////////////////////////////////////////////////
|
||||
|
||||
///////////////////////////////////////////
|
||||
// 1. Describe A and B tensors
|
||||
///////////////////////////////////////////
|
||||
using ElementA = mx_float4_t;
|
||||
constexpr int AlignA = 128;
|
||||
using GmemLayoutA = cutlass::layout::ColumnMajor;
|
||||
using ElementB = mx_float4_t;
|
||||
constexpr int AlignB = 128;
|
||||
using GmemLayoutB = cutlass::layout::RowMajor;
|
||||
```
|
||||
Next, we need to choose the performance parameters such as `MmaTileShape_MNK`, `KernelMainloopPolicy`,
|
||||
and `ClusterShape_MNK`.
|
||||
|
||||
`MmaTileShape_MNK` supported for `mx_float4_t`x`mx_float4_t` with `mxf8f6f4` are listed in [Table 11](#bs_rows_3).
|
||||
For NT layout, we see that 3 `MmaTileShape_MNK` are supported: `128x128x128`, and `128x256x128` with 1SM instruction;
|
||||
and `256x256x128` with 2SM instruction. Let's say, we expect to get the best performance with `256x256x128` MMA tile shape
|
||||
for our GEMM problem. Then, we need to set the `KernelMainloopPolicy` to `KernelTmaWarpSpecialized2SmMxf8f6f4Sm100`.
|
||||
Now, we need to choose the `ClusterShape_MNK`. Since we have selected a 2SM mma instruction, `ClusterShape_MNK` should be
|
||||
compatible and its first mode should be a multiple of 2. `ClusterShape_MNK = cute::Shape<_2, [_1|_2|_4], _1>` or
|
||||
`ClusterShape_MNK = cute::Shape<_4, [_1|_2|_4], _1>` would be valid options. Let's choose `cute::Shape<_4,_4,_1>`.
|
||||
Our performance parameters looks like below:
|
||||
|
||||
```cpp
|
||||
//////////////////////////////////////////
|
||||
// 2. Choose Performance Parameters
|
||||
//////////////////////////////////////////
|
||||
|
||||
// Tile and cluster shapes
|
||||
// Collective MMA takes tile shape of the MMA operation as input
|
||||
using KernelMainloopPolicy = cutlass::gemm::KernelTmaWarpSpecialized2SmMxf8f6f4Sm100;
|
||||
using MmaTileShape_MNK = cute::Shape<_256,_256,_128>;
|
||||
using ClusterShape_MNK = cute::Shape<_4,_4,_1>;
|
||||
```
|
||||
|
||||
After we config the main-loop, let's setup the epilogue.
|
||||
A normal epilogue looks like below, we need to specify the output layout, datatype, alignment and PerSmTileShape_MNK, and let others to be default/auto.
|
||||
|
||||
PerSmTileShape_MNK should be deduced from the mainloop setup. For example, in above mainloop setup, the MmaTileShape_MNK is
|
||||
256x256x128 and the KernelMainloopPolicy is 2sm policy.
|
||||
It means each CTA is doing (256 / 2sm) x 256 x 128 output, so the PerSmTileShape_MNK is 128x256x128. The possible PerSmTileShape_MNK
|
||||
is listed in [Table 15](#epi_persmtileshape)
|
||||
|
||||
The epilogue scheduling policy is configurable, and it is common to set `cutlass::epilogue::collective::EpilogueScheduleAuto`
|
||||
to allow the epilogue builder to automatically select the appropriate policy. However, it can also be explicitly defined to
|
||||
use other policies based on the 1sm or 2sm MMA instruction. The available policies are listed in [Table 14](#epi_dispatch).
|
||||
|
||||
```cpp
|
||||
// Describe C and D tensors
|
||||
using ElementC = cutlass::half_t;
|
||||
constexpr int AlignC = 8;
|
||||
using GmemLayoutC = cutlass::layout::RowMajor;
|
||||
using ElementD = cutlass::float_e2m1_t;
|
||||
constexpr int AlignD = 32;
|
||||
using GmemLayoutD = cutlass::layout::RowMajor;
|
||||
// Mma's accumulator type
|
||||
using ElementAccumulator = float;
|
||||
// Epilogue computation's precision type
|
||||
using ElementCompute = float;
|
||||
|
||||
//
|
||||
// Construct CollectiveEpilogue
|
||||
//
|
||||
|
||||
using CollectiveEpilogue = typename cutlass::epilogue::collective::CollectiveBuilder<
|
||||
cutlass::arch::Sm100, cutlass::arch::OpClassBlockScaledTensorOp, // Arch and Tensorop spec
|
||||
MmaTileShape_MNK, ClusterShape_MNK, // MMA tile shape, and cluster shape
|
||||
cutlass::epilogue::collective::EpilogueTileAuto, // Epilogue subtile shape. Auto will find a suitable tile shape
|
||||
ElementAccumulator, ElementCompute, // Mma instr's accumulator type and compute precision for epilogue
|
||||
ElementC, GmemLayoutC, AlignC, // C tensor description
|
||||
ElementD, GmemLayoutD, AlignD, // D tensor description
|
||||
cutlass::epilogue::TmaWarpSpecialized2Sm // Epilogue schedule policy
|
||||
>::CollectiveOp;
|
||||
|
||||
```
|
||||
|
||||
If we want to let the epilogue generate mxf4/nvf4/mxf6/mxf8 (i.e. elements + block-scalefactor), we need to setup the epilogue fusion into the builder.
|
||||
First, we need to choose a SFDVectorSize indicates how many elements sharing the same block-scalefactor.
|
||||
Then, we need to choose ElementSFD and GmemLayoutSFD which indicates the output datatype and which output-dim is used to generate the block-scalefactor.
|
||||
Typically, GmemLayoutSFD would be same as the GmemLayoutD.
|
||||
|
||||
```cpp
|
||||
//
|
||||
// Construct FusionOperation
|
||||
//
|
||||
constexpr int SFDVectorSize = 16;
|
||||
// Define the fusion operation applied during epilogue
|
||||
using FusionOperation = cutlass::epilogue::fusion::LinCombBlockScaleFactor<
|
||||
SFDVectorSize,
|
||||
ElementD, ElementCompute,
|
||||
ElementSFD, GmemLayoutSFD,
|
||||
ElementC
|
||||
>;
|
||||
|
||||
using CollectiveEpilogue = typename cutlass::epilogue::collective::CollectiveBuilder<
|
||||
cutlass::arch::Sm100, cutlass::arch::OpClassBlockScaledTensorOp, // Arch and Tensorop spec
|
||||
MmaTileShape_MNK, ClusterShape_MNK, // MMA tile shape, and cluster shape
|
||||
cutlass::epilogue::collective::EpilogueTileAuto, // Epilogue subtile shape. Auto will find a suitable tile shape
|
||||
ElementAccumulator, ElementCompute, // Mma instr's accumulator type and compute precision for epilogue
|
||||
ElementC, GmemLayoutC, AlignC, // C tensor description
|
||||
ElementD, GmemLayoutD, AlignD, // D tensor description
|
||||
cutlass::epilogue::TmaWarpSpecialized2Sm // Epilogue schedule policy
|
||||
FusionOperation // <================================== Pass the fusion config into epilogue builder.
|
||||
>::CollectiveOp;
|
||||
```
|
||||
|
||||
Above example made a gentle introduction to using the fusion operations in the epilogue. For more detailed example, see
|
||||
[Blackwell GEMM with collective builder](https://github.com/NVIDIA/cutlass/tree/main/examples/71_blackwell_gemm_with_collective_builder/71_blackwell_gemm_with_collective_builder.cu)
|
||||
|
||||
Note that we have first discussed the CollectiveMainloop, then the CollectiveEpilogue for clarity.
|
||||
However, the CollectiveMainloop needs to know the SMEM utilization of the epilogue. Therefore, it needs to be setup before the CollectiveMainloop. See [examples/72_blackwell_narrow_precision_gemm](https://github.com/NVIDIA/cutlass/tree/main/examples/72_blackwell_narrow_precision_gemm/) directory for full kernel and run setup.
|
||||
|
||||
### Scale Factor Layouts
|
||||
|
||||
The scale factor layout consists of a 512B basic-block structure, as illustrated in the diagram below. Each block contains 128 M/N dimension and 4 scale factors (SF) along the K dimension.
|
||||
The byte order of the basic storage chunk is row-major, meaning that M0SF0 to M0SF3, M32SF0 to M32SF3, M64SF0 to M64SF3, and M96SF0 to M96SF3 are stored consecutively in GMEM.
|
||||
|
||||

|
||||
|
||||
If the scale factor tensor exceeds M128xSF4, it indicates that there are multiple basic blocks along both the M and SFK dimensions. The arrangement of these basic blocks follows a K-major order. Here is a diagram illustrating the scenario where M equals 512 and the SFK is 16.
|
||||
|
||||

|
||||
|
||||
The creation of scale factor tensors' layouts are tedious. CUTLASS provides `Sm1xxBlockScaledConfig` to create these layouts easily
|
||||
(See [sm100_blockscaled_layout.hpp](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/detail/sm100_blockscaled_layout.hpp)).
|
||||
The interface to create SFA and SFB tensor layouts is as follows:
|
||||
|
||||
```cpp
|
||||
auto problem_shape = make_shape(M, N, K, L);
|
||||
using SfConfig = Sm1xxBlockScaledConfig<SFVecSize>;
|
||||
|
||||
// SFA shape: ((32,4), ceil(M/128)), ((SFVecSize,4), ceil(K/4), L)
|
||||
auto layout_sfa = SfConfig::tile_atom_to_shape_SFA(problem_shape);
|
||||
// SFB shape: ((32,4), ceil(N/128)), ((SFVecSize,4), ceil(K/4), L)
|
||||
auto layout_sfb = SfConfig::tile_atom_to_shape_SFB(problem_shape);
|
||||
|
||||
auto tensor_sfa = make_tensor(aptr, layout_sfa);
|
||||
auto tensor_sfb = make_tensor(bptr, layout_sfb);
|
||||
// Access SF for for element m,k of A tensor
|
||||
auto val_a_mk = tensor_sfa(make_coord(m,k,0));
|
||||
```
|
||||
# Blackwell SM120 GEMMs
|
||||
The NVIDIA RTX 5000 Series GPUs introduce support for new narrow precision (4bit and 6bit) block-scaled and non-block-scaled tensor cores. The PTX ISA has extended the `mma` instructions to support these data formats which are 1x to 4x faster than Ada architecture's fp8 tensor cores. For more detailed information see [`mma` PTX documentation](https://docs.nvidia.com/cuda/parallel-thread-execution/index.html#multiply-and-accumulate-instruction-mma).
|
||||
|
||||
CUTLASS 4.0 has added support for these newly introduced narrow precision GEMMs. Similar to the Blackwell SM100 GEMMs, the SM120 GEMMs can be built using the collective builder interface. See examples in [examples/79_blackwell_geforce_gemm/](../../examples/79_blackwell_geforce_gemm/) and unit tests listed below.
|
||||
|
||||
The data types supported and tensor alignment requirements are the same as the Blackwell SM100 GEMMs. The scale factor layout is also the same as SM100 mentioned above. `OpClassTensorOp` is used for non-blockscaled narrow precision GEMMs and `OpClassBlockScaledTensorOp` is used for blockscaled narrow precision GEMMs.
|
||||
|
||||
| Ptx Instruction | Throughput | Notes | Unit Test |
|
||||
|---------------------------------------------------------------------|----------------------------|-------|-----------|
|
||||
|mma.sync.aligned.kind::f8f6f4 | 1x Ada Fp8 Tensor Core(2x for FP32 accumulator) | Mixed precision MMA with A={f4,f6,f8} x B={f4,f6,f8} TN layouts | [unit test](../../test/unit/gemm/device/sm120_tensorop_gemm/) |
|
||||
|mma.sync.aligned.kind::mxf8f6f4.block_scale | 1x Ada Fp8 Tensor Core(2x for FP32 accumulator) | Block scaled mixed precision MMA with A={mxf4,mxf6,mxf8} x B={mxf4,mxf6,mxf8} with TN layouts | [unit test](../../test/unit/gemm/device/sm120_blockscaled_tensorop_gemm/sm120_bs_gemm_mxf6_mxf8_f32_f32.cu) |
|
||||
|mma.sync.aligned.kind::mxf4.block_scale | 2x Ada Fp8 Tensor Core(4x for FP32 accumulator) | Block scaled MMA with A={mxf4} x B={mxf4} with TN layouts | [unit test](../../test/unit/gemm/device/sm120_blockscaled_tensorop_gemm/sm120_bs_gemm_mxf4_mxf4_f32_f32.cu) |
|
||||
|mma.sync.aligned.kind::mxf4nvf4.block_scale.scale_vec::[2X\|4X] | 2x Ada Fp8 Tensor Core(4x for FP32 accumulator) | Block scaled MMA with A={mxf4} x B={mxf4} or A={nvf4} x B={nvf4} with TN layouts | [unit test](../../test/unit/gemm/device/sm120_blockscaled_tensorop_gemm/sm120_bs_gemm_nvf4_nvf4_f32_f32.cu) |
|
||||
|
||||
Besides the similarities, there are some key differences from the Blackwell SM100 GEMMs:
|
||||
|
||||
## Cluster Size
|
||||
|
||||
On Geforce series graphics card, there is no multicast feature therefore the cluster shape is fixed to 1x1x1.
|
||||
|
||||
## Tensor Layout
|
||||
|
||||
Only TN layout is supported. Matrix A is row major and matrix B is column major.
|
||||
|
||||
## Pingpong v.s. cooperative kernel schedule
|
||||
|
||||
Similar to Hopper's warp-group GEMM, SM120 GEMMs support both pingpong and cooperative kernel schedules. Pingpong kernel schedule has two groups of 4 MMA warps working on different output tiles, overlapping the mainloop and epilogue, while the cooperative kernel schedule has only one group of 8 MMA warps working on the same output tile. If `KernelScheduleAuto` is specified, `KernelTmaWarpSpecializedCooperative` will be selected by default.
|
||||
|
||||
## Epilogue schedule:
|
||||
|
||||
`EpilogueScheduleAuto` must be used.
|
||||
|
||||
## Tile size:
|
||||
|
||||
Below are tables that summarize the valid tile shapes and dispatch policies for SM120 GEMMs. If the output is `float_6_t`, the tile size in the leading dimension of output tensor must be 128.
|
||||
|
||||
**Table 16: Valid Tile Shapes and Dispatch Policies for {float8_t, float_6_t, float_4_t} x {float8_t, float_6_t, float_4_t} of SM120 GEMMs**
|
||||
| Mma Tile Shape | TN | TT | NT | NN | Dispatch Policy |
|
||||
|----------------|----|----|----|----|------------------------------------|
|
||||
64x64x128 | Y | N | N | N | `KernelTmaWarpSpecializedPingpong` or `KernelTmaWarpSpecializedCooperative` |
|
||||
64x128x128 | Y | N | N | N | `KernelTmaWarpSpecializedPingpong` or `KernelTmaWarpSpecializedCooperative` |
|
||||
128x64x128 | Y | N | N | N | `KernelTmaWarpSpecializedPingpong` or `KernelTmaWarpSpecializedCooperative` |
|
||||
128x128x128 | Y | N | N | N | `KernelTmaWarpSpecializedPingpong` or `KernelTmaWarpSpecializedCooperative` |
|
||||
|
||||
**Table 17: Valid Tile Shapes for nv_float4_t x nv_float4_t of SM120 GEMMs**
|
||||
| Mma Tile Shape | TN | TT | NT | NN | Dispatch Policy |
|
||||
|----------------|----|----|----|----|------------------------------------|
|
||||
128x128x128 | Y | N | N | N | `KernelTmaWarpSpecializedPingpong` or `KernelTmaWarpSpecializedCooperative` |
|
||||
256x128x128 | Y | N | N | N | `KernelTmaWarpSpecializedCooperative` |
|
||||
128x128x256 | Y | N | N | N | `KernelTmaWarpSpecializedPingpong` or `KernelTmaWarpSpecializedCooperative` |
|
||||
|
||||
**Table 18: Valid Tile Shapes and Dispatch Policies for mx_float4_t x mx_float4_t of SM120 GEMMs**
|
||||
| Mma Tile Shape | TN | TT | NT | NN | Dispatch Policy |
|
||||
|----------------|----|----|----|----|------------------------------------|
|
||||
128x128x128 | Y | N | N | N | `KernelTmaWarpSpecializedPingpong` or `KernelTmaWarpSpecializedCooperative` |
|
||||
256x128x128 | Y | N | N | N | `KernelTmaWarpSpecializedCooperative` |
|
||||
128x128x256 | Y | N | N | N | `KernelTmaWarpSpecializedPingpong` or `KernelTmaWarpSpecializedCooperative` |
|
||||
|
||||
**Table 19: Valid Tile Shapes and Dispatch Policies for mx_float4_t x mx_float4_t of SM120 GEMMs**
|
||||
| Mma Tile Shape | TN | TT | NT | NN | Dispatch Policy |
|
||||
|----------------|----|----|----|----|------------------------------------|
|
||||
128x128x128 | Y | N | N | N | `KernelTmaWarpSpecializedMxf8f6f4Sm120` or `KernelTmaWarpSpecializedPingpongMxf8f6f4Sm120` |
|
||||
256x128x128 | Y | N | N | N | `KernelTmaWarpSpecializedMxf8f6f4Sm120` |
|
||||
128x128x256 | Y | N | N | N | `KernelTmaWarpSpecializedMxf8f6f4Sm120` or `KernelTmaWarpSpecializedPingpongMxf8f6f4Sm120` |
|
||||
|
||||
Specialized policies must be used to generate mixed-input-datatype `mx_float4_t` kernels.
|
||||
|
||||
**Table 20: Valid Tile Shapes and Dispatch Policies for {mx_float4_t, mx_float6_t, mx_float8_t} x {mx_float4_t, mx_float6_t, mx_float8_t}**
|
||||
| Mma Tile Shape | TN | TT | NT | NN | Dispatch Policy |
|
||||
|----------------|----|----|----|----|------------------------------------|
|
||||
128x128x128 | Y | N | N | N | `KernelTmaWarpSpecializedPingpong` or `KernelTmaWarpSpecializedCooperative` |
|
||||
|
||||
# Copyright
|
||||
|
||||
Copyright (c) 2025 - 2025 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
```
|
||||
Redistribution and use in source and binary forms, with or without
|
||||
modification, are permitted provided that the following conditions are met:
|
||||
|
||||
1. Redistributions of source code must retain the above copyright notice, this
|
||||
list of conditions and the following disclaimer.
|
||||
|
||||
2. Redistributions in binary form must reproduce the above copyright notice,
|
||||
this list of conditions and the following disclaimer in the documentation
|
||||
and/or other materials provided with the distribution.
|
||||
|
||||
3. Neither the name of the copyright holder nor the names of its
|
||||
contributors may be used to endorse or promote products derived from
|
||||
this software without specific prior written permission.
|
||||
|
||||
THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS "AS IS"
|
||||
AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE
|
||||
IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE ARE
|
||||
DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT HOLDER OR CONTRIBUTORS BE LIABLE
|
||||
FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL
|
||||
DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR
|
||||
SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER
|
||||
CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY,
|
||||
OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE
|
||||
OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
|
||||
```
|
||||
119
media/docs/cpp/build/building_in_windows_with_visual_studio.md
Normal file
119
media/docs/cpp/build/building_in_windows_with_visual_studio.md
Normal file
@@ -0,0 +1,119 @@
|
||||
# Building on Windows with Visual Studio
|
||||
|
||||
CUTLASS 3.2 reintroduces support for the Microsoft Visual Studio compiler on Windows.
|
||||
Users and developers may build either
|
||||
in Visual Studio's graphical integrated development environment,
|
||||
or on the command line with `cmake --build`.
|
||||
|
||||
# Software prerequisites
|
||||
|
||||
1. Windows 10 or 11
|
||||
|
||||
2. Visual Studio 2019 version 16.11.27, or Visual Studio 2022
|
||||
|
||||
3. CUDA Toolkit (at least 12.2; earlier 12.x versions may work)
|
||||
|
||||
4. CMake (at least 3.18)
|
||||
|
||||
5. git
|
||||
|
||||
6. Python (at least 3.6)
|
||||
|
||||
Visual Studio must be installed *before* the CUDA Toolkit.
|
||||
Otherwise, Visual Studio's build system won't know about CUDA.
|
||||
|
||||
# Operating system settings
|
||||
|
||||
By default, Windows restricts the maximum file path length (`MAX_PATH`) to 260 characters.
|
||||
CUTLASS has many files and directory paths that challenge this requirement.
|
||||
As a result, CUTLASS is unlikely to build with this default setting.
|
||||
The choice of source and build directories affect path lengths,
|
||||
so the kinds of errors and whether they occur may depend on this.
|
||||
Symptoms may vary, from errors when running `cmake`
|
||||
(e.g., during the "generating library instances" step) to build failures.
|
||||
|
||||
CUTLASS recommends changing the maximum file path length setting
|
||||
and rebooting the computer before attempting to clone or build CUTLASS.
|
||||
Windows 10 (as of version 1607) and 11 permit changing this setting
|
||||
by making sure that the following registry key exists,
|
||||
and that its value is set to 1.
|
||||
|
||||
```
|
||||
Computer\HKEY_LOCAL_MACHINE\SYSTEM\CurrentControlSet\Control\FileSystem\LongPathsEnabled
|
||||
```
|
||||
|
||||
After changing the registry key's value, reboot the computer first
|
||||
before attempting to clone or build CUTLASS.
|
||||
|
||||
[This Microsoft help article](https://learn.microsoft.com/en-us/windows/win32/fileio/maximum-file-path-limitation?tabs=registry)
|
||||
explains different ways to change the registry setting.
|
||||
|
||||
# Set up build environment
|
||||
|
||||
1. Run "git bash" to get a familiar command-line interface
|
||||
|
||||
2. Edit `~/.profile` and set the environment variables as needed to access the CUTLASS repository
|
||||
|
||||
3. Clone the CUTLASS repository
|
||||
|
||||
4. Create the `build` subdirectory in the CUTLASS clone directory, and run CMake in it,
|
||||
specifying whatever CMake options are desired, e.g.,
|
||||
`cmake .. -DCUTLASS_NVCC_ARCHS=90a`
|
||||
|
||||
Alternate approaches may rely on the CMake GUI and/or Windows' native command line.
|
||||
|
||||
# Building
|
||||
|
||||
A successful CMake run will create a `CUTLASS.sln` Visual Studio "solution" file in the build directory.
|
||||
One can open this in Visual Studio and build the entire solution or any subset of projects as desired.
|
||||
It may be necessary to limit maximum build parallelism by setting the appropriate Visual Studio option.
|
||||
|
||||
Alternately, one can run `cmake --build . --config Release -j 4` in the build directory.
|
||||
Replace 4 with the desired maximum build parallelism.
|
||||
It's important to put the `--build` option before the period that signifies the build directory.
|
||||
The `--config` option specifies the kind of build;
|
||||
`--config Release` builds a Release build, while `--config Debug` builds a Debug build.
|
||||
Unlike with CMake's Makefile or Ninja generators,
|
||||
`CMAKE_BUILD_TYPE` has no effect on the Visual Studio generator,
|
||||
because the Visual Studio generator creates all build configurations.
|
||||
|
||||
# Tips
|
||||
|
||||
With Windows builds, one may find that CMake reruns unnecessarily.
|
||||
For example, cancelling a build and starting it again may rerun CMake.
|
||||
This will in turn touch build files that result in unnecessary rebuilds.
|
||||
One work-around is to set the CMake option `CMAKE_SUPPRESS_REGENERATION=ON`.
|
||||
However, this turns off CMake's ability to detect on its own when it needs to rerun.
|
||||
As a result, one will need to know when to rerun CMake by hand.
|
||||
|
||||
## Copyright
|
||||
|
||||
Copyright (c) 2017 - 2025 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
```
|
||||
Redistribution and use in source and binary forms, with or without
|
||||
modification, are permitted provided that the following conditions are met:
|
||||
|
||||
1. Redistributions of source code must retain the above copyright notice, this
|
||||
list of conditions and the following disclaimer.
|
||||
|
||||
2. Redistributions in binary form must reproduce the above copyright notice,
|
||||
this list of conditions and the following disclaimer in the documentation
|
||||
and/or other materials provided with the distribution.
|
||||
|
||||
3. Neither the name of the copyright holder nor the names of its
|
||||
contributors may be used to endorse or promote products derived from
|
||||
this software without specific prior written permission.
|
||||
|
||||
THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS "AS IS"
|
||||
AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE
|
||||
IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE ARE
|
||||
DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT HOLDER OR CONTRIBUTORS BE LIABLE
|
||||
FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL
|
||||
DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR
|
||||
SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER
|
||||
CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY,
|
||||
OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE
|
||||
OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
|
||||
```
|
||||
88
media/docs/cpp/build/building_with_clang_as_host_compiler.md
Normal file
88
media/docs/cpp/build/building_with_clang_as_host_compiler.md
Normal file
@@ -0,0 +1,88 @@
|
||||
# Building with Clang as host compiler
|
||||
|
||||
CUTLASS 3.2(.1) reintroduces support for building with
|
||||
Clang as host compiler, and NVCC as device compiler.
|
||||
This is NOT the same as building with
|
||||
Clang as both host and device compiler ("CUDA Clang").
|
||||
|
||||
# Software prerequisites
|
||||
|
||||
1. Clang (regularly tested with Clang 17;
|
||||
occasionally tested with Clang 10 and greater)
|
||||
|
||||
2. CUDA Toolkit (tested with 12.2; other versions likely work)
|
||||
|
||||
3. CMake (at least 3.18)
|
||||
|
||||
4. git
|
||||
|
||||
5. Python (at least 3.6)
|
||||
|
||||
Experience with Ubuntu 22.04 LTS is that
|
||||
clang requires the following packages to be installed.
|
||||
|
||||
```bash
|
||||
$ sudo apt-get install clang cmake ninja-build pkg-config libgtk-3-dev liblzma-dev libstdc++-12-dev
|
||||
```
|
||||
|
||||
A symptom of not installing all needed dependencies
|
||||
is the following error when attempting to use clang:
|
||||
`"/usr/bin/ld: cannot find -lstdc++: No such file or directory"`.
|
||||
|
||||
# Running CMake
|
||||
|
||||
## Required CMake options
|
||||
|
||||
The Clang build requires specifying the following CMake options.
|
||||
Replace `<path-to-clang++>` with the path to your `clang++` executable.
|
||||
You may use `clang++` directly if it is in your `PATH`.
|
||||
|
||||
* `CMAKE_CXX_COMPILER=<path-to-clang++>`
|
||||
* `CMAKE_CUDA_HOST_COMPILER=<path-to-clang++>`
|
||||
|
||||
One must set both! It's not enough just to set the `CXX` environment
|
||||
variable, for example. Symptoms of only setting `CMAKE_CXX_COMPILER`
|
||||
(or only setting the `CXX` environment variable) include `cc1plus`
|
||||
(GCC's compiler executable) reporting build errors due to it not
|
||||
understanding Clang's command-line options.
|
||||
|
||||
Users can also specify a particular CUDA Toolkit version
|
||||
by setting the CMake option `CMAKE_CUDA_COMPILER`
|
||||
to the path to the `nvcc` executable
|
||||
that lives in the CUDA Toolkit's directory. For example,
|
||||
if `${PATH_TO_CUDA_TOOLKIT}` is the CUDA Toolkit directory,
|
||||
then one can set `CMAKE_CUDA_COMPILER` as follows.
|
||||
|
||||
* `CMAKE_CUDA_COMPILER=${PATH_TO_CUDA_TOOLKIT}/bin/nvcc`
|
||||
|
||||
## Copyright
|
||||
|
||||
Copyright (c) 2017 - 2025 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
```
|
||||
Redistribution and use in source and binary forms, with or without
|
||||
modification, are permitted provided that the following conditions are met:
|
||||
|
||||
1. Redistributions of source code must retain the above copyright notice, this
|
||||
list of conditions and the following disclaimer.
|
||||
|
||||
2. Redistributions in binary form must reproduce the above copyright notice,
|
||||
this list of conditions and the following disclaimer in the documentation
|
||||
and/or other materials provided with the distribution.
|
||||
|
||||
3. Neither the name of the copyright holder nor the names of its
|
||||
contributors may be used to endorse or promote products derived from
|
||||
this software without specific prior written permission.
|
||||
|
||||
THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS "AS IS"
|
||||
AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE
|
||||
IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE ARE
|
||||
DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT HOLDER OR CONTRIBUTORS BE LIABLE
|
||||
FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL
|
||||
DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR
|
||||
SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER
|
||||
CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY,
|
||||
OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE
|
||||
OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
|
||||
```
|
||||
267
media/docs/cpp/code_organization.md
Normal file
267
media/docs/cpp/code_organization.md
Normal file
@@ -0,0 +1,267 @@
|
||||

|
||||
|
||||
# CUTLASS Code Organization
|
||||
|
||||
This document describes the layout of the CUTLASS repository. The main components are:
|
||||
|
||||
* **CUTLASS Template Library** - CUDA Templates for Linear Algebra Subroutines and Solvers (header only)
|
||||
* **CuTe Template Library** - CUTLASS's core vocabulary layout type and associated algebra (header only)
|
||||
* **CUTLASS Utilities** - Additional templates
|
||||
* **CUTLASS Instance Library** - instantiations of CUTLASS templates covering the design space
|
||||
* **CUTLASS Profiler** - CUTLASS Library, Profiler, and Utilities
|
||||
* **Examples** - SDK examples of CUTLASS Template Library and components
|
||||
* **Media** - supporting documentation and media content
|
||||
* **Tests** - test components for CUTLASS Template Library and tools
|
||||
|
||||
## CUTLASS Template Library
|
||||
|
||||
CUDA Templates for Linear Algebra Subroutines and Solvers is a library of CUDA C++ template classes for
|
||||
performing efficient matrix computations on NVIDIA GPUs.
|
||||
|
||||
Like NVIDIA CUB, the components of CUTLASS are organized hierarchically based on the scope of cooperative
|
||||
elements. For example, warp-level GEMM components perform a matrix multiply collectively by the
|
||||
set of threads within a warp. The following figure illustrates each layer.
|
||||
|
||||
Components are designed to be usable by client applications accessing functionailty at each scope.
|
||||
|
||||
CUTLASS Templates are implemented by header files in the following directory structure:
|
||||
|
||||
```
|
||||
include/ # Top-level include directory. Client applications should target this path.
|
||||
cutlass/ # CUDA Templates for Linear Algebra Subroutines and Solvers - headers only
|
||||
|
||||
arch/ # direct exposure of architecture features (including instruction-level GEMMs)
|
||||
*
|
||||
gemm/ # code specialized for general matrix product computations
|
||||
thread/ # thread-level operators
|
||||
warp/ # warp-level operators
|
||||
collective/ # 3.x API operators for all threads a tiled mma/copy are built over
|
||||
threadblock/ # CTA-level operators
|
||||
kernel/ # CUDA kernel entry points
|
||||
device/ # launches kernel(s) over a full device
|
||||
* # scope-agnostic components and basic vocabulary type definitions for GEMM
|
||||
|
||||
layout/ # layout definitions for matrices, tensors, and other mathematical objects in memory
|
||||
*
|
||||
|
||||
reduction/ # bandwidth-limited reduction kernels that do not fit the "gemm" models
|
||||
thread/ # thread-level operators
|
||||
warp/ # warp-level operators
|
||||
threadblock/ # CTA-level operators
|
||||
kernel/ # CUDA kernel entry points
|
||||
device/ # launches kernel(s) over a full device
|
||||
* # scope-agnostic components and basic vocabulary type definitions
|
||||
|
||||
transform/ # code specialized for layout, type, and domain transformations
|
||||
thread/ # thread-level operators
|
||||
warp/ # warp-level operators
|
||||
threadblock/ # CTA-level operators
|
||||
kernel/ # CUDA kernel entry points
|
||||
device/ # launches kernel(s) over a full device
|
||||
* # scope-agnostic components and basic vocabulary type definitions
|
||||
|
||||
util/ # miscellaneous CUTLASS components
|
||||
*
|
||||
* # core vocabulary types and fundamental arithmetic operators
|
||||
|
||||
cute / # CuTe Layout, layout algebra, MMA/Copy atoms, tiled MMA/Copy
|
||||
algorithm/ # Definitions of core operations such as copy, gemm, and operations on cute::tuples
|
||||
arch/ # Bare bones PTX wrapper structs for copy and math instructions
|
||||
atom/ # Meta-information either link to or built from arch/ operators
|
||||
mma_atom.hpp # cute::Mma_Atom and cute::TiledMma
|
||||
copy_atom.hpp # cute::Copy_Atom and cute::TiledCopy
|
||||
*sm*.hpp # Arch specific meta-information for copy and math operations
|
||||
container/ # Core container types used across CuTe, namely, cute::tuple
|
||||
numeric/ # CuTe's internal numerics implementation
|
||||
* # Core library types such as Shape, Stride, Layout, Tensor, and associated operations
|
||||
```
|
||||
|
||||
See [Programming Guidelines](programming_guidelines.md) for further details about
|
||||
conventions and design patterns used throughout CUTLASS.
|
||||
|
||||
## CuTe
|
||||
|
||||
CuTe is a collection of C++ CUDA template abstractions for defining and operating on hierarchically multidimensional layouts of threads and data. CuTe provides `Layout` and `Tensor` objects that compactly packages the type, shape, memory space, and layout of data, while performing the complicated indexing for the user. This lets programmers focus on the logical descriptions of their algorithms while CuTe does the mechanical bookkeeping for them. With these tools, we can quickly design, implement, and modify all dense linear algebra operations. More documentation
|
||||
for CuTe can be found in [`cute/`](cute/index).
|
||||
|
||||
## Tools
|
||||
|
||||
The `tools/` directory contains clients of the CUTLASS Template library and includes the following.
|
||||
|
||||
## CUTLASS Instance Library
|
||||
|
||||
The CUTLASS Instance Library contains instantiations of the above CUTLASS templates covering supported configurations,
|
||||
data types, block structure, and tile sizes. These instantiations are procedurally generated using a set of
|
||||
scripts to span the design space.
|
||||
|
||||
```
|
||||
tools/
|
||||
library/ # static/dynamic library containing all kernel instantiations of interest
|
||||
# (with some build-level filter switches to compile specific subsets)
|
||||
|
||||
include/
|
||||
cutlass/
|
||||
library/ # header files for CUTLASS Deliverables Library (in cutlass::library:: namespace)
|
||||
|
||||
handle.h # implements a host-side API for launching kernels, similar to cuBLAS
|
||||
library.h # defines enums and structs to describe the tiled structure of operator instances
|
||||
manifest.h # collection of all instances
|
||||
|
||||
src/
|
||||
|
||||
python/
|
||||
cutlass_library/ # scripts to procedurally generate CUTLASS template instances
|
||||
|
||||
gemm_operations.py
|
||||
library.py
|
||||
generator.py # entry point of procedural generation scripts - invoked by cmake
|
||||
manifest.py
|
||||
```
|
||||
|
||||
When CMake is executed, the CUTLASS Instance Library generator scripts are executed to construct a set of
|
||||
instantiations in `build/tools/library/generated/`.
|
||||
|
||||
### CUTLASS Profiler
|
||||
|
||||
The CUTLASS Profiler is designed to load the CUTLASS Instance Library and execute all operations contained therein.
|
||||
This command-line driven application constructs an execution environment for evaluating functionality and performance.
|
||||
It is implemented in
|
||||
```
|
||||
tools/
|
||||
profiler/
|
||||
```
|
||||
|
||||
and may be built as follows.
|
||||
```
|
||||
$ make cutlass_profiler -j
|
||||
```
|
||||
|
||||
[Further details about the CUTLASS Profiler are described here.](profiler.md)
|
||||
|
||||
### CUTLASS Utilities
|
||||
|
||||
`tools/util/` defines a companion library of headers and sources that support the CUTLASS test programs, examples, and other client applications. Its structure is as follows:
|
||||
|
||||
```
|
||||
tools/
|
||||
util/
|
||||
include/
|
||||
cutlass/
|
||||
util/ # CUTLASS Utility companion library
|
||||
|
||||
reference/ # functional reference implementation of CUTLASS operators
|
||||
# (minimal consideration for performance)
|
||||
|
||||
detail/
|
||||
*
|
||||
|
||||
device/ # device-side reference implementations of CUTLASS operators
|
||||
thread/
|
||||
kernel/
|
||||
*
|
||||
host/ # host-side reference implementations of CUTLASS operators
|
||||
*
|
||||
*
|
||||
```
|
||||
|
||||
[More details about CUTLASS Utilities may be found here.](utilities.md)
|
||||
|
||||
|
||||
## Examples
|
||||
|
||||
To demonstrate CUTLASS components, several SDK examples are implemented in `examples/`.
|
||||
|
||||
CUTLASS SDK examples apply CUTLASS templates to implement basic computations.
|
||||
|
||||
```
|
||||
examples/
|
||||
00_basic_gemm/ # launches a basic GEMM with single precision inputs and outputs
|
||||
|
||||
01_cutlass_utilities/ # demonstrates CUTLASS Utilities for allocating and initializing tensors
|
||||
|
||||
02_dump_reg_smem/ # debugging utilities for printing register and shared memory contents
|
||||
|
||||
03_visualize_layout/ # utility for visualizing all layout functions in CUTLASS
|
||||
|
||||
04_tile_iterator/ # example demonstrating an iterator over tiles in memory
|
||||
|
||||
05_batched_gemm/ # example demonstrating CUTLASS's batched strided GEMM operation
|
||||
|
||||
06_splitK_gemm/ # exmaple demonstrating CUTLASS's Split-K parallel reduction kernel
|
||||
|
||||
07_volta_tensorop_gemm/ # example demonstrating mixed precision GEMM using Volta Tensor Cores
|
||||
|
||||
08_turing_tensorop_gemm/ # example demonstrating integer GEMM using Turing Tensor Cores
|
||||
|
||||
10_planar_complex/ # example demonstrating planar complex GEMM kernels
|
||||
|
||||
11_planar_complex_array/ # example demonstrating planar complex kernels with batch-specific problem sizes
|
||||
|
||||
12_gemm_bias_relu/ # example demonstrating GEMM fused with bias and relu activation function
|
||||
|
||||
13_fused_two_gemms/ # example demonstrating two GEMMs fused into one kernel
|
||||
```
|
||||
|
||||
## Media
|
||||
|
||||
This directory contains documentation, images, and performance result data which accompanies the CUTLASS library and components.
|
||||
|
||||
## Tests
|
||||
|
||||
Test programs for CUTLASS. Tests are organized hierarchically, mirroring the organization of source files.
|
||||
```
|
||||
test/ # unit tests for CUTLASS Template Library
|
||||
unit/
|
||||
arch/
|
||||
core/
|
||||
gemm/
|
||||
device/
|
||||
kernel/
|
||||
thread/
|
||||
threadblock/
|
||||
warp/
|
||||
reduction/
|
||||
kernel/
|
||||
thread/
|
||||
transform/
|
||||
threadblock/
|
||||
*
|
||||
```
|
||||
Tests can be built and run at the top level scope by invoking `make test_unit` or by building
|
||||
and explicitly executing each individual target, e.g. `cutlass_test_unit_gemm_device`.
|
||||
|
||||
Tests are configured to specify appropriate GTest filter strings to avoid running except on
|
||||
architectures where they are expected to pass. Thus, no tests should fail. The actual number
|
||||
of tests run may vary over time as more are added.
|
||||
|
||||
# Copyright
|
||||
|
||||
Copyright (c) 2017 - 2025 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
```
|
||||
Redistribution and use in source and binary forms, with or without
|
||||
modification, are permitted provided that the following conditions are met:
|
||||
|
||||
1. Redistributions of source code must retain the above copyright notice, this
|
||||
list of conditions and the following disclaimer.
|
||||
|
||||
2. Redistributions in binary form must reproduce the above copyright notice,
|
||||
this list of conditions and the following disclaimer in the documentation
|
||||
and/or other materials provided with the distribution.
|
||||
|
||||
3. Neither the name of the copyright holder nor the names of its
|
||||
contributors may be used to endorse or promote products derived from
|
||||
this software without specific prior written permission.
|
||||
|
||||
THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS "AS IS"
|
||||
AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE
|
||||
IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE ARE
|
||||
DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT HOLDER OR CONTRIBUTORS BE LIABLE
|
||||
FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL
|
||||
DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR
|
||||
SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER
|
||||
CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY,
|
||||
OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE
|
||||
OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
|
||||
```
|
||||
151
media/docs/cpp/cute/00_quickstart.md
Normal file
151
media/docs/cpp/cute/00_quickstart.md
Normal file
@@ -0,0 +1,151 @@
|
||||
# Getting Started With CuTe
|
||||
|
||||
CuTe is a collection of C++ CUDA template abstractions for defining and operating on hierarchically multidimensional layouts of threads and data. CuTe provides `Layout` and `Tensor` objects that compactly packages the type, shape, memory space, and layout of data, while performing the complicated indexing for the user. This lets programmers focus on the logical descriptions of their algorithms while CuTe does the mechanical bookkeeping for them. With these tools, we can quickly design, implement, and modify all dense linear algebra operations.
|
||||
|
||||
The core abstraction of CuTe are the hierarchically multidimensional layouts which can be composed with data arrays to represent tensors. The representation of layouts is powerful enough to represent nearly everything we need to implement efficient dense linear algebra. Layouts can also be combined and manipulated via functional composition, on which we build a large set of common operations such as tiling and partitioning.
|
||||
|
||||
## System Requirements
|
||||
|
||||
CuTe shares CUTLASS 3.x's software requirements,
|
||||
including NVCC with a C++17 host compiler.
|
||||
|
||||
## Knowledge prerequisites
|
||||
|
||||
CuTe is a CUDA C++ header-only library. It requires C++17
|
||||
(the revision of the C++ Standard that was released in 2017).
|
||||
|
||||
Throughout this tutorial, we assume intermediate C++ experience.
|
||||
For example, we assume that readers know
|
||||
how to read and write templated functions and classes, and
|
||||
how to use the `auto` keyword to deduce a function's return type.
|
||||
We will be gentle with C++ and explain some things
|
||||
that you might already know.
|
||||
|
||||
We also assume intermediate CUDA experience.
|
||||
For example, readers must know
|
||||
the difference between device and host code,
|
||||
and how to launch kernels.
|
||||
|
||||
## Building Tests and Examples
|
||||
|
||||
CuTe's tests and examples build and run as part of CUTLASS's normal build process.
|
||||
|
||||
CuTe's unit tests live in the [`test/unit/cute`](https://github.com/NVIDIA/cutlass/tree/main/test/unit/cute) subdirectory.
|
||||
|
||||
CuTe's examples live in the [`examples/cute`](https://github.com/NVIDIA/cutlass/tree/main/examples/cute) subdirectory.
|
||||
|
||||
## Library Organization
|
||||
|
||||
CuTe is a header-only C++ library, so there is no source code that needs building. Library headers are contained within the top level [`include/cute`](https://github.com/NVIDIA/cutlass/tree/main/include/cute) directory, with components of the library grouped by directories that represent their semantics.
|
||||
|
||||
| Directory | Contents |
|
||||
|------------------------|------------------------|
|
||||
| [`include/cute`](https://github.com/NVIDIA/cutlass/tree/main/include/cute) | Each header in the top level corresponds to one of the fundamental building blocks of CuTe, such as [`Layout`](https://github.com/NVIDIA/cutlass/tree/main/include/cute/layout.hpp) and [`Tensor`](https://github.com/NVIDIA/cutlass/tree/main/include/cute/tensor.hpp). |
|
||||
| [`include/cute/container`](https://github.com/NVIDIA/cutlass/tree/main/include/cute/container) | Implementations of STL-like objects, such as tuple, array, and aligned array. |
|
||||
| [`include/cute/numeric`](https://github.com/NVIDIA/cutlass/tree/main/include/cute/numeric) | Fundamental numeric data types that include nonstandard floating-point types, nonstandard integer types, complex numbers, and integer sequence. |
|
||||
| [`include/cute/algorithm`](https://github.com/NVIDIA/cutlass/tree/main/include/cute/algorithm) | Implementations of utility algorithms such as copy, fill, and clear that automatically leverage architecture-specific features if available. |
|
||||
| [`include/cute/arch`](https://github.com/NVIDIA/cutlass/tree/main/include/cute/arch) | Wrappers for architecture-specific matrix-matrix multiply and copy instructions. |
|
||||
| [`include/cute/atom`](https://github.com/NVIDIA/cutlass/tree/main/include/cute/atom) | Meta-information for instructions in `arch` and utilities like partitioning and tiling.
|
||||
|
||||
## Tutorial
|
||||
|
||||
This directory contains a CuTe tutorial in Markdown format.
|
||||
The file
|
||||
[`0x_gemm_tutorial.md`](./0x_gemm_tutorial.md)
|
||||
explains how to implement dense matrix-matrix multiply using CuTe components.
|
||||
It gives a broad overview of CuTe and thus would be a good place to start.
|
||||
|
||||
Other files in this directory discuss specific parts of CuTe.
|
||||
|
||||
* [`01_layout.md`](./01_layout.md) describes `Layout`, CuTe's core abstraction.
|
||||
|
||||
* [`02_layout_algebra.md`](./02_layout_algebra.md) describes more advanced `Layout` operations and the CuTe layout algebra.
|
||||
|
||||
* [`03_tensor.md`](./03_tensor.md) describes `Tensor`,
|
||||
a multidimensional array abstraction which composes `Layout`
|
||||
with an array of data.
|
||||
|
||||
* [`04_algorithms.md`](./04_algorithms.md) summarizes CuTe's
|
||||
generic algorithms that operate on `Tensor`s.
|
||||
|
||||
* [`0t_mma_atom.md`](./0t_mma_atom.md) demonstrates CuTe's meta-information and interface to our GPUs'
|
||||
architecture-specific Matrix Multiply-Accumulate (MMA) instructions.
|
||||
|
||||
* [`0x_gemm_tutorial.md`](./0x_gemm_tutorial.md) walks through building a GEMM from scratch using CuTe.
|
||||
|
||||
* [`0y_predication.md`](./0y_predication.md) explains what to do
|
||||
if a tiling doesn't fit evenly into a matrix.
|
||||
|
||||
* [`0z_tma_tensors.md`](./0z_tma_tensors.md) explains an advanced `Tensor` type that CuTe uses to support TMA loads and stores.
|
||||
|
||||
## Quick Tips
|
||||
|
||||
### How do I print CuTe objects on host or device?
|
||||
|
||||
The `cute::print` function has overloads for almost all CuTe types, including Pointers, Integers, Strides, Shapes, Layouts, and Tensors. When in doubt, try calling `print` on it.
|
||||
|
||||
CuTe's print functions work on either host or device.
|
||||
Note that on device, printing is expensive.
|
||||
Even just leaving print code in place on device,
|
||||
even if it is never called
|
||||
(e.g., printing in an `if` branch that is not taken at run time),
|
||||
may generate slower code.
|
||||
Thus, be sure to remove code that prints on device after debugging.
|
||||
|
||||
You might also only want to print on thread 0 of each threadblock, or threadblock 0 of the grid. The `thread0()` function returns true only for global thread 0 of the kernel, that is, for thread 0 of threadblock 0. A common idiom for printing CuTe objects to print only on global thread 0.
|
||||
|
||||
```c++
|
||||
if (thread0()) {
|
||||
print(some_cute_object);
|
||||
}
|
||||
```
|
||||
|
||||
Some algorithms depend on some thread or threadblock,
|
||||
so you may need to print on threads or threadblocks other than zero.
|
||||
The header file
|
||||
[`cute/util/debug.hpp`](https://github.com/NVIDIA/cutlass/tree/main/include/cute/util/debug.hpp),
|
||||
among other utilities,
|
||||
includes the function `bool thread(int tid, int bid)`
|
||||
that returns `true` if running on thread `tid` and threadblock `bid`.
|
||||
|
||||
#### Other output formats
|
||||
|
||||
Some CuTe types have special printing functions that use a different output format.
|
||||
|
||||
The `cute::print_layout` function will display any rank-2 layout in a plain test table. This is excellent for visualizing the map from coordinates to indices.
|
||||
|
||||
The `cute::print_tensor` function will display any rank-1, rank-2, rank-3, or rank-4 tensor in a plain text multidimensional table. The values of the tensor are printed so you can verify the tile of data is what you expect after a copy, for example.
|
||||
|
||||
The `cute::print_latex` function will print LaTeX commands that you can use to build a nicely formatted and colored tables via `pdflatex`. This work for `Layout`, `TiledCopy`, and `TiledMMA`, which can be very useful to get a sense of layout patterns and partitioning patterns within CuTe.
|
||||
|
||||
## Copyright
|
||||
|
||||
Copyright (c) 2017 - 2025 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
```
|
||||
Redistribution and use in source and binary forms, with or without
|
||||
modification, are permitted provided that the following conditions are met:
|
||||
|
||||
1. Redistributions of source code must retain the above copyright notice, this
|
||||
list of conditions and the following disclaimer.
|
||||
|
||||
2. Redistributions in binary form must reproduce the above copyright notice,
|
||||
this list of conditions and the following disclaimer in the documentation
|
||||
and/or other materials provided with the distribution.
|
||||
|
||||
3. Neither the name of the copyright holder nor the names of its
|
||||
contributors may be used to endorse or promote products derived from
|
||||
this software without specific prior written permission.
|
||||
|
||||
THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS "AS IS"
|
||||
AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE
|
||||
IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE ARE
|
||||
DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT HOLDER OR CONTRIBUTORS BE LIABLE
|
||||
FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL
|
||||
DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR
|
||||
SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER
|
||||
CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY,
|
||||
OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE
|
||||
OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
|
||||
```
|
||||
569
media/docs/cpp/cute/01_layout.md
Normal file
569
media/docs/cpp/cute/01_layout.md
Normal file
@@ -0,0 +1,569 @@
|
||||
# CuTe Layouts
|
||||
|
||||
This document describes `Layout`, CuTe's core abstraction.
|
||||
Fundamentally, a `Layout` maps from coordinate space(s)
|
||||
to an index space.
|
||||
|
||||
`Layout`s present a common interface to multidimensional array access
|
||||
that abstracts away the details of how the array's elements are organized in memory.
|
||||
This lets users write algorithms that access multidimensional arrays generically,
|
||||
so that layouts can change, without users' code needing to change. For example, a row-major MxN layout and a column-major MxN layout can be treated identically in software.
|
||||
|
||||
CuTe also provides an "algebra of `Layout`s."
|
||||
`Layout`s can be combined and manipulated
|
||||
to construct more complicated layouts
|
||||
and to tile layouts across other layouts.
|
||||
This can help users do things like partition layouts of data over layouts of threads.
|
||||
|
||||
## Fundamental Types and Concepts
|
||||
|
||||
### Integers
|
||||
|
||||
CuTe makes great use of dynamic (known only at run-time) and static (known at compile-time) integers.
|
||||
|
||||
* Dynamic integers (or "run-time integers") are just ordinary integral types like `int` or `size_t` or `uint16_t`. Anything that is accepted by `std::is_integral<T>` is considered a dynamic integer in CuTe.
|
||||
|
||||
* Static integers (or "compile-time integers") are instantiations of types like `std::integral_constant<Value>`. These types encode the value as a `static constexpr` member. They also support casting to their underlying dynamic types, so they can be used in expressions with dynamic integers. CuTe defines its own CUDA-compatibe static integer types `cute::C<Value>` along with overloaded math operators so that math on static integers results in static integers. CuTe defines shortcut aliases `Int<1>`, `Int<2>`, `Int<3>` and `_1`, `_2`, `_3` as conveniences, which you should see often within examples.
|
||||
|
||||
CuTe attempts to handle static and dynamic integers identically. In the examples that follow, all dynamic integers could be replaced with static integers and vice versa. When we say "integer" in CuTe, we almost always mean a static OR dynamic integer.
|
||||
|
||||
CuTe provides a number of traits to work with integers.
|
||||
* `cute::is_integral<T>`: Checks whether `T` is a static or dynamic integer type.
|
||||
* `cute::is_std_integral<T>`: Checks whether `T` is a dynamic integer type. Equivalent to `std::is_integral<T>`.
|
||||
* `cute::is_static<T>`: Checks whether `T` is an empty type (so instantiations cannot depend on any dynamic information). Equivalent to `std::is_empty`.
|
||||
* `cute::is_constant<N,T>`: Checks that `T` is a static integer AND its value is equivalent to `N`.
|
||||
|
||||
See the [`integral_constant` implementations](https://github.com/NVIDIA/cutlass/tree/main/include/cute/numeric/integral_constant.hpp) for more information.
|
||||
|
||||
### Tuple
|
||||
|
||||
A tuple is a finite ordered list of zero or more elements.
|
||||
The [`cute::tuple` class](https://github.com/NVIDIA/cutlass/tree/main/include/cute/container/tuple.hpp) behaves like `std::tuple`, but works on device and host. It imposes restrictions on its template arguments and strips down the implementation for performance and simplicity.
|
||||
|
||||
### IntTuple
|
||||
|
||||
CuTe defines the IntTuple concept as either an integer, or a tuple of IntTuples. Note the recursive definition.
|
||||
In C++, we define [operations on `IntTuple`](https://github.com/NVIDIA/cutlass/tree/main/include/cute/int_tuple.hpp).
|
||||
|
||||
Examples of `IntTuple`s include:
|
||||
* `int{2}`, the dynamic integer 2.
|
||||
* `Int<3>{}`, the static integer 3.
|
||||
* `make_tuple(int{2}, Int<3>{})`, the tuple of dynamic-2, and static-3.
|
||||
* `make_tuple(uint16_t{42}, make_tuple(Int<1>{}, int32_t{3}), Int<17>{})`, the tuple of dynamic-42, tuple of static-1 and dynamic-3, and static-17.
|
||||
|
||||
CuTe reuses the `IntTuple` concept for many different things,
|
||||
including Shape, Stride, Step, and Coord
|
||||
(see [`include/cute/layout.hpp`](https://github.com/NVIDIA/cutlass/tree/main/include/cute/layout.hpp)).
|
||||
|
||||
Operations defined on `IntTuple`s include the following.
|
||||
|
||||
* `rank(IntTuple)`: The number of elements in an `IntTuple`. A single integer has rank 1, and a tuple has rank `tuple_size`.
|
||||
|
||||
* `get<I>(IntTuple)`: The `I`th element of the `IntTuple`, with `I < rank`. For single integers, `get<0>` is just that integer.
|
||||
|
||||
* `depth(IntTuple)`: The number of hierarchical `IntTuple`s. A single integer has depth 0, a tuple of integers has depth 1, a tuple that contains a tuple of integers has depth 2, etc.
|
||||
|
||||
* `size(IntTuple)`: The product of all elements of the `IntTuple`.
|
||||
|
||||
We write `IntTuple`s with parentheses to denote the hierarchy. For example, `6`, `(2)`, `(4,3)`, and `(3,(6,2),8)` are all `IntTuple`s.
|
||||
|
||||
### Shapes and Strides
|
||||
|
||||
Both `Shape` and `Stride` are `IntTuple` concepts.
|
||||
|
||||
### Layout
|
||||
|
||||
A `Layout` is a tuple of (`Shape`, `Stride`).
|
||||
Semantically, it implements a mapping from
|
||||
any coordinate within the Shape to an index via the Stride.
|
||||
|
||||
### Tensor
|
||||
|
||||
A `Layout` can be composed with data -- e.g., a pointer or an array -- to create a `Tensor`. The index generated by the `Layout` is used to subscript an iterator to retrieve the appropriate data. For details on `Tensor`, please refer to the
|
||||
[`Tensor` section of the tutorial](./03_tensor.md).
|
||||
|
||||
## Layout Creation and Use
|
||||
|
||||
A `Layout` is a pair of `IntTuple`s: the `Shape` and the `Stride`. The first element defines the abstract *shape* of the `Layout`, and the second element defines the *strides*, which map from coordinates within the shape to the index space.
|
||||
|
||||
We define many operations on `Layout`s analogous to those defined on `IntTuple`.
|
||||
|
||||
* `rank(Layout)`: The number of modes in a `Layout`. Equivalent to the tuple size of the `Layout`'s shape.
|
||||
|
||||
* `get<I>(Layout)`: The `I`th sub-layout of the `Layout`, with `I < rank`.
|
||||
|
||||
* `depth(Layout)`: The depth of the `Layout`'s shape. A single integer has depth 0, a tuple of integers has depth 1, a tuple of tuples of integers has depth 2, etc.
|
||||
|
||||
* `shape(Layout)`: The shape of the `Layout`.
|
||||
|
||||
* `stride(Layout)`: The stride of the `Layout`.
|
||||
|
||||
* `size(Layout)`: The size of the `Layout` function's domain. Equivalent to `size(shape(Layout))`.
|
||||
|
||||
* `cosize(Layout)`: The size of the `Layout` function's codomain (not necessarily the range). Equivalent to `A(size(A) - 1) + 1`.
|
||||
|
||||
### Hierarchical access functions
|
||||
|
||||
`IntTuple`s and `Layout`s can be arbitrarily nested.
|
||||
For convenience, we define versions of some of the above functions
|
||||
that take a sequence of integers, instead of just one integer.
|
||||
This makes it possible to access elements
|
||||
inside of nested `IntTuple` or `Layout` more easily.
|
||||
For example, we permit `get<I...>(x)`, where `I...` is a "C++ parameter pack" that denotes zero or more (integer) template arguments. These hierarchical access functions include the following.
|
||||
|
||||
* `get<I0,I1,...,IN>(x) := get<IN>(...(get<I1>(get<I0>(x)))...)`. Extract the `IN`th of the ... of the `I1`st of the `I0`th element of `x`.
|
||||
|
||||
* `rank<I...>(x) := rank(get<I...>(x))`. The rank of the `I...`th element of `x`.
|
||||
|
||||
* `depth<I...>(x) := depth(get<I...>(x))`. The depth of the `I...`th element of `x`.
|
||||
|
||||
* `shape<I...>(x) := shape(get<I...>(x))`. The shape of the `I...`th element of `x`.
|
||||
|
||||
* `size<I...>(x) := size(get<I...>(x))`. The size of the `I...`th element of `x`.
|
||||
|
||||
In the following examples, you'll see use of `size<0>` and `size<1>` to determine loops bounds for the 0th and 1st mode of a layout or tensor.
|
||||
|
||||
### Constructing a Layout
|
||||
|
||||
A `Layout` can be constructed in many different ways.
|
||||
It can include any combination of compile-time (static) integers
|
||||
or run-time (dynamic) integers.
|
||||
|
||||
```c++
|
||||
Layout s8 = make_layout(Int<8>{});
|
||||
Layout d8 = make_layout(8);
|
||||
|
||||
Layout s2xs4 = make_layout(make_shape(Int<2>{},Int<4>{}));
|
||||
Layout s2xd4 = make_layout(make_shape(Int<2>{},4));
|
||||
|
||||
Layout s2xd4_a = make_layout(make_shape (Int< 2>{},4),
|
||||
make_stride(Int<12>{},Int<1>{}));
|
||||
Layout s2xd4_col = make_layout(make_shape(Int<2>{},4),
|
||||
LayoutLeft{});
|
||||
Layout s2xd4_row = make_layout(make_shape(Int<2>{},4),
|
||||
LayoutRight{});
|
||||
|
||||
Layout s2xh4 = make_layout(make_shape (2,make_shape (2,2)),
|
||||
make_stride(4,make_stride(2,1)));
|
||||
Layout s2xh4_col = make_layout(shape(s2xh4),
|
||||
LayoutLeft{});
|
||||
```
|
||||
|
||||
The `make_layout` function returns a `Layout`.
|
||||
It deduces the types of the function's arguments and returns a `Layout` with the appropriate template arguments.
|
||||
Similarly, the `make_shape` and `make_stride` functions
|
||||
return a `Shape` resp. `Stride`.
|
||||
CuTe often uses these `make_*` functions
|
||||
due to restrictions around constructor template argument deduction (CTAD) and to avoid having to repeat static or dynamic integer types.
|
||||
|
||||
When the `Stride` argument is omitted, it is generated from the provided `Shape` with `LayoutLeft` as default. The `LayoutLeft` tag constructs strides as an exclusive prefix product of the `Shape` from left to right, without regard to the `Shape`'s hierarchy. This can be considered a "generalized column-major stride generation". The `LayoutRight` tag constructs strides as an exclusive prefix product of the `Shape` from right to left, without regard to the `Shape`'s hierarchy. For shapes of depth one, this can be considered a "row-major stride generation", but for hierarchical shapes the resulting strides may be surprising. For example, the strides of `s2xh4` above could be generated with `LayoutRight`.
|
||||
|
||||
Calling `print` on each layout above results in the following
|
||||
|
||||
```
|
||||
s8 : _8:_1
|
||||
d8 : 8:_1
|
||||
s2xs4 : (_2,_4):(_1,_2)
|
||||
s2xd4 : (_2,4):(_1,_2)
|
||||
s2xd4_a : (_2,4):(_12,_1)
|
||||
s2xd4_col : (_2,4):(_1,_2)
|
||||
s2xd4_row : (_2,4):(4,_1)
|
||||
s2xh4 : (2,(2,2)):(4,(2,1))
|
||||
s2xh4_col : (2,(2,2)):(_1,(2,4))
|
||||
```
|
||||
|
||||
The `Shape:Stride` notation is used quite often for `Layout`. The `_N` notation is shorthand for a static integer while other integers are dynamic integers. Observe that both `Shape` and `Stride` may be composed of both static and dynamic integers.
|
||||
|
||||
Also note that the `Shape` and `Stride` are assumed to be *congruent*. That is, `Shape` and `Stride` have the same tuple profiles. For every integer in `Shape`, there is a corresponding integer in `Stride`. This can be asserted with
|
||||
```cpp
|
||||
static_assert(congruent(my_shape, my_stride));
|
||||
```
|
||||
|
||||
### Using a Layout
|
||||
|
||||
The fundamental use of a `Layout` is to map between coordinate space(s) defined by the `Shape` and an index space defined by the `Stride`. For example, to print an arbitrary rank-2 layout in a 2-D table, we can write the function
|
||||
|
||||
```c++
|
||||
template <class Shape, class Stride>
|
||||
void print2D(Layout<Shape,Stride> const& layout)
|
||||
{
|
||||
for (int m = 0; m < size<0>(layout); ++m) {
|
||||
for (int n = 0; n < size<1>(layout); ++n) {
|
||||
printf("%3d ", layout(m,n));
|
||||
}
|
||||
printf("\n");
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
which produces the following output for the above examples.
|
||||
|
||||
```
|
||||
> print2D(s2xs4)
|
||||
0 2 4 6
|
||||
1 3 5 7
|
||||
> print2D(s2xd4_a)
|
||||
0 1 2 3
|
||||
12 13 14 15
|
||||
> print2D(s2xh4_col)
|
||||
0 2 4 6
|
||||
1 3 5 7
|
||||
> print2D(s2xh4)
|
||||
0 2 1 3
|
||||
4 6 5 7
|
||||
```
|
||||
|
||||
We can see static, dynamic, row-major, column-major, and hierarchical layouts printed here. The statement `layout(m,n)` provides the mapping of
|
||||
the logical 2-D coordinate (m,n) to the 1-D index.
|
||||
|
||||
Interestingly, the `s2xh4` example isn't row-major or column-major. Furthermore, it has three modes but is still interpreted as rank-2 and we're using a 2-D coordinate. Specifically, `s2xh4` has a 2-D multi-mode in the second mode, but we're still able to use a 1-D coordinate for that mode. More on this in the next section, but first we can generalize this another step. Let's use a 1-D coordinate and treat all of the modes of each layout as a single multi-mode. For instance, the following `print1D` function
|
||||
|
||||
```c++
|
||||
template <class Shape, class Stride>
|
||||
void print1D(Layout<Shape,Stride> const& layout)
|
||||
{
|
||||
for (int i = 0; i < size(layout); ++i) {
|
||||
printf("%3d ", layout(i));
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
produces the following output for the above examples.
|
||||
|
||||
```
|
||||
> print1D(s2xs4)
|
||||
0 1 2 3 4 5 6 7
|
||||
> print1D(s2xd4_a)
|
||||
0 12 1 13 2 14 3 15
|
||||
> print1D(s2xh4_col)
|
||||
0 1 2 3 4 5 6 7
|
||||
> print1D(s2xh4)
|
||||
0 4 2 6 1 5 3 7
|
||||
```
|
||||
|
||||
Any multi-mode of a layout, including the entire layout itself, can accept a 1-D coordinate. More on this in the following sections.
|
||||
|
||||
CuTe provides more printing utilities for visualizing Layouts. The `print_layout` function produces a formatted 2-D table of the Layout's mapping.
|
||||
|
||||
```text
|
||||
> print_layout(s2xh4)
|
||||
(2,(2,2)):(4,(2,1))
|
||||
0 1 2 3
|
||||
+---+---+---+---+
|
||||
0 | 0 | 2 | 1 | 3 |
|
||||
+---+---+---+---+
|
||||
1 | 4 | 6 | 5 | 7 |
|
||||
+---+---+---+---+
|
||||
```
|
||||
|
||||
The `print_latex` function generates LaTeX that can be compiled with `pdflatex` into a color-coded vector graphics image of the same 2-D table.
|
||||
|
||||
### Vector Layouts
|
||||
|
||||
We define a vector as any `Layout` with `rank == 1`.
|
||||
For example, the layout `8:1` can be interpreted as an 8-element vector whose indices are contiguous.
|
||||
|
||||
```
|
||||
Layout: 8:1
|
||||
Coord : 0 1 2 3 4 5 6 7
|
||||
Index : 0 1 2 3 4 5 6 7
|
||||
```
|
||||
|
||||
Similarly,
|
||||
the layout `8:2` can be interpreted as an 8-element vector where the indices of the elements are strided by `2`.
|
||||
|
||||
```
|
||||
Layout: 8:2
|
||||
Coord : 0 1 2 3 4 5 6 7
|
||||
Index : 0 2 4 6 8 10 12 14
|
||||
```
|
||||
|
||||
By the above rank-1 definition, we *also* interpret layout `((4,2)):((2,1))` as a vector, since its shape is rank-1. The inner shape looks like a 4x2 row-major matrix, but the extra pair of parenthesis suggest we can interpret those two modes as a 1-D 8-element vector. The strides tell us that the first `4` elements are strided by `2` and then there are `2` of those first elements strided by `1`.
|
||||
|
||||
```
|
||||
Layout: ((4,2)):((2,1))
|
||||
Coord : 0 1 2 3 4 5 6 7
|
||||
Index : 0 2 4 6 1 3 5 7
|
||||
```
|
||||
|
||||
We can see the second set of `4` elements are duplicates of the first `4` with an extra stride of `1`.
|
||||
|
||||
Consider the layout `((4,2)):((1,4))`. Again, it's `4` elements strided by `1` and then `2` of those first elements strided by `4`.
|
||||
|
||||
```
|
||||
Layout: ((4,2)):((1,4))
|
||||
Coord : 0 1 2 3 4 5 6 7
|
||||
Index : 0 1 2 3 4 5 6 7
|
||||
```
|
||||
|
||||
As a function from integers to integers, it's identical to `8:1`. It's the identity function.
|
||||
|
||||
### Matrix examples
|
||||
|
||||
Generalizing, we define a matrix as any `Layout` that is rank-2. For example,
|
||||
|
||||
```
|
||||
Shape : (4,2)
|
||||
Stride: (1,4)
|
||||
0 4
|
||||
1 5
|
||||
2 6
|
||||
3 7
|
||||
```
|
||||
|
||||
is a 4x2 column-major layout with stride-1 down the columns and stride-4 across the rows, and
|
||||
|
||||
```
|
||||
Shape : (4,2)
|
||||
Stride: (2,1)
|
||||
0 1
|
||||
2 3
|
||||
4 5
|
||||
6 7
|
||||
```
|
||||
|
||||
is a 4x2 row-major layout with stride-2 down the columns and stride-1 across the rows. Majorness is simply which mode has stride-1.
|
||||
|
||||
Just like the vector layouts, each of the modes of the matrix can also be split into *multi-modes*.
|
||||
This lets us express more layouts beyond just row-major and column-major. For example,
|
||||
|
||||
```
|
||||
Shape: ((2,2),2)
|
||||
Stride: ((4,1),2)
|
||||
0 2
|
||||
4 6
|
||||
1 3
|
||||
5 7
|
||||
```
|
||||
|
||||
is also logically 4x2, with stride-2 across the rows but a multi-stride down the columns. The first `2` elements down the column have a stride of `4` and then there is a copy of those with stride-1. Since this layout is logically 4x2,
|
||||
like the column-major and row-major examples above,
|
||||
we can _still_ use 2-D coordinates to index into it.
|
||||
|
||||
## Layout Concepts
|
||||
|
||||
In this section, we'll introduce the coordinate sets that `Layout`s accept and how the coordinate mappings and index mappings are computed.
|
||||
|
||||
### Layout compatibility
|
||||
|
||||
We say that layout A is *compatible* with layout B if the shape of A is compatible with the shape of B.
|
||||
Shape A is compatible with shape B if
|
||||
|
||||
* the size of A is equal to the size of B and
|
||||
* all coordinates within A are valid coordinates within B.
|
||||
|
||||
For example:
|
||||
* Shape 24 is NOT compatible with Shape 32.
|
||||
* Shape 24 is compatible with Shape (4,6).
|
||||
* Shape (4,6) is compatible with Shape ((2,2),6).
|
||||
* Shape ((2,2),6) is compatible with Shape ((2,2),(3,2)).
|
||||
* Shape 24 is compatible with Shape ((2,2),(3,2)).
|
||||
* Shape 24 is compatible with Shape ((2,3),4).
|
||||
* Shape ((2,3),4) is NOT compatible with Shape ((2,2),(3,2)).
|
||||
* Shape ((2,2),(3,2)) is NOT compatible with Shape ((2,3),4).
|
||||
* Shape 24 is compatible with Shape (24).
|
||||
* Shape (24) is NOT compatible with Shape 24.
|
||||
* Shape (24) is NOT compatible with Shape (4,6).
|
||||
|
||||
That is, *compatible* is a weak partial order on Shapes as it is reflexive, antisymmetric, and transitive.
|
||||
|
||||
### Layouts Coordinates
|
||||
|
||||
With the notion of compatibility above, we emphasize that every `Layout` accepts multiple kinds of coordinates. Every `Layout` accepts coordinates for any `Shape` that is compatible with it. CuTe provides mappings between these sets of coordinates via a colexicographical order.
|
||||
|
||||
Thus, all Layouts provide two fundamental mappings:
|
||||
|
||||
* the map from an input coordinate to the corresponding natural coordinate via the `Shape`, and
|
||||
* the map from a natural coordinate to the index via the `Stride`.
|
||||
|
||||
#### Coordinate Mapping
|
||||
|
||||
The map from an input coordinate to a natural coordinate is the application of a colexicographical order (reading right to left, instead of "lexicographical," which reads left to right) within the `Shape`.
|
||||
|
||||
Take the shape `(3,(2,3))`, for example. This shape has three coordinate sets: the 1-D coordinates, the 2-D coordinates, and the natural (h-D) coordinates.
|
||||
|
||||
| 1-D | 2-D | Natural | | 1-D | 2-D | Natural |
|
||||
| ----- | ------- | ----------- |-| ----- | ------- | ----------- |
|
||||
| `0` | `(0,0)` | `(0,(0,0))` | | `9` | `(0,3)` | `(0,(1,1))` |
|
||||
| `1` | `(1,0)` | `(1,(0,0))` | | `10` | `(1,3)` | `(1,(1,1))` |
|
||||
| `2` | `(2,0)` | `(2,(0,0))` | | `11` | `(2,3)` | `(2,(1,1))` |
|
||||
| `3` | `(0,1)` | `(0,(1,0))` | | `12` | `(0,4)` | `(0,(0,2))` |
|
||||
| `4` | `(1,1)` | `(1,(1,0))` | | `13` | `(1,4)` | `(1,(0,2))` |
|
||||
| `5` | `(2,1)` | `(2,(1,0))` | | `14` | `(2,4)` | `(2,(0,2))` |
|
||||
| `6` | `(0,2)` | `(0,(0,1))` | | `15` | `(0,5)` | `(0,(1,2))` |
|
||||
| `7` | `(1,2)` | `(1,(0,1))` | | `16` | `(1,5)` | `(1,(1,2))` |
|
||||
| `8` | `(2,2)` | `(2,(0,1))` | | `17` | `(2,5)` | `(2,(1,2))` |
|
||||
|
||||
Each coordinate into the shape `(3,(2,3))` has two *equivalent* coordinates and all equivalent coordinates map to the same natural coordinate. To emphasize again, because all of the above coordinates are valid inputs, a Layout with Shape `(3,(2,3))` can be used as if it is a 1-D array of 18 elements by using the 1-D coordinates, a 2-D matrix of 3x6 elements by using the 2-D coordinates, or a h-D tensor of 3x(2x3) elements by using the h-D (natural) coordinates.
|
||||
|
||||
The previous 1-D print demonstrates how CuTe identifies 1-D coordinates with a colexicographical ordering of 2-D coordinates. Iterating from `i = 0` to `size(layout)` and indexing into our layout with the single integer coordinate `i`, traverses the 2-D coordinates in this "generalized-column-major" order, even if the layout maps coordinates to indices in a row-major or more complex fashion.
|
||||
|
||||
The function `cute::idx2crd(idx, shape)` is responsible for the coordinate mapping. It will take any coordinate within the shape and compute the equivalent natural coordinate for that shape.
|
||||
```cpp
|
||||
auto shape = Shape<_3,Shape<_2,_3>>{};
|
||||
print(idx2crd( 16, shape)); // (1,(1,2))
|
||||
print(idx2crd(_16{}, shape)); // (_1,(_1,_2))
|
||||
print(idx2crd(make_coord( 1,5), shape)); // (1,(1,2))
|
||||
print(idx2crd(make_coord(_1{},5), shape)); // (_1,(1,2))
|
||||
print(idx2crd(make_coord( 1,make_coord(1, 2)), shape)); // (1,(1,2))
|
||||
print(idx2crd(make_coord(_1{},make_coord(1,_2{})), shape)); // (_1,(1,_2))
|
||||
```
|
||||
|
||||
#### Index Mapping
|
||||
|
||||
The map from a natural coordinate to an index is performed by taking the inner product of the natural coordinate with the `Layout`'s `Stride`.
|
||||
|
||||
Take the layout `(3,(2,3)):(3,(12,1))`, for example. Then a natural coordinate `(i,(j,k))` will result in the index `i*3 + j*12 + k*1`. The indices this layout computes are shown in the 2-D table below where `i` is used as the row coordinate and `(j,k)` is used as the column coordinate.
|
||||
|
||||
```
|
||||
0 1 2 3 4 5 <== 1-D col coord
|
||||
(0,0) (1,0) (0,1) (1,1) (0,2) (1,2) <== 2-D col coord (j,k)
|
||||
+-----+-----+-----+-----+-----+-----+
|
||||
0 | 0 | 12 | 1 | 13 | 2 | 14 |
|
||||
+-----+-----+-----+-----+-----+-----+
|
||||
1 | 3 | 15 | 4 | 16 | 5 | 17 |
|
||||
+-----+-----+-----+-----+-----+-----+
|
||||
2 | 6 | 18 | 7 | 19 | 8 | 20 |
|
||||
+-----+-----+-----+-----+-----+-----+
|
||||
```
|
||||
|
||||
The function `cute::crd2idx(c, shape, stride)` is responsible for the index mapping. It will take any coordinate within the shape, compute the equivalent natural coordinate for that shape (if it is not already), and compute the inner product with the strides.
|
||||
```cpp
|
||||
auto shape = Shape <_3,Shape< _2,_3>>{};
|
||||
auto stride = Stride<_3,Stride<_12,_1>>{};
|
||||
print(crd2idx( 16, shape, stride)); // 17
|
||||
print(crd2idx(_16{}, shape, stride)); // _17
|
||||
print(crd2idx(make_coord( 1, 5), shape, stride)); // 17
|
||||
print(crd2idx(make_coord(_1{}, 5), shape, stride)); // 17
|
||||
print(crd2idx(make_coord(_1{},_5{}), shape, stride)); // _17
|
||||
print(crd2idx(make_coord( 1,make_coord( 1, 2)), shape, stride)); // 17
|
||||
print(crd2idx(make_coord(_1{},make_coord(_1{},_2{})), shape, stride)); // _17
|
||||
```
|
||||
|
||||
## Layout Manipulation
|
||||
|
||||
### Sublayouts
|
||||
|
||||
Sublayouts can be retrieved with `layout<I...>`
|
||||
```cpp
|
||||
Layout a = Layout<Shape<_4,Shape<_3,_6>>>{}; // (4,(3,6)):(1,(4,12))
|
||||
Layout a0 = layout<0>(a); // 4:1
|
||||
Layout a1 = layout<1>(a); // (3,6):(4,12)
|
||||
Layout a10 = layout<1,0>(a); // 3:4
|
||||
Layout a11 = layout<1,1>(a); // 6:12
|
||||
```
|
||||
or `select<I...>`
|
||||
```cpp
|
||||
Layout a = Layout<Shape<_2,_3,_5,_7>>{}; // (2,3,5,7):(1,2,6,30)
|
||||
Layout a13 = select<1,3>(a); // (3,7):(2,30)
|
||||
Layout a01 = select<0,1,3>(a); // (2,3,7):(1,2,30)
|
||||
Layout a2 = select<2>(a); // (5):(6)
|
||||
```
|
||||
or `take<ModeBegin, ModeEnd>`
|
||||
```cpp
|
||||
Layout a = Layout<Shape<_2,_3,_5,_7>>{}; // (2,3,5,7):(1,2,6,30)
|
||||
Layout a13 = take<1,3>(a); // (3,5):(2,6)
|
||||
Layout a14 = take<1,4>(a); // (3,5,7):(2,6,30)
|
||||
// take<1,1> not allowed. Empty layouts not allowed.
|
||||
```
|
||||
|
||||
### Concatenation
|
||||
|
||||
A `Layout` can be provided to `make_layout` to wrap and concatenate
|
||||
```cpp
|
||||
Layout a = Layout<_3,_1>{}; // 3:1
|
||||
Layout b = Layout<_4,_3>{}; // 4:3
|
||||
Layout row = make_layout(a, b); // (3,4):(1,3)
|
||||
Layout col = make_layout(b, a); // (4,3):(3,1)
|
||||
Layout q = make_layout(row, col); // ((3,4),(4,3)):((1,3),(3,1))
|
||||
Layout aa = make_layout(a); // (3):(1)
|
||||
Layout aaa = make_layout(aa); // ((3)):((1))
|
||||
Layout d = make_layout(a, make_layout(a), a); // (3,(3),3):(1,(1),1)
|
||||
```
|
||||
or can be combined with `append`, `prepend`, or `replace`.
|
||||
```cpp
|
||||
Layout a = Layout<_3,_1>{}; // 3:1
|
||||
Layout b = Layout<_4,_3>{}; // 4:3
|
||||
Layout ab = append(a, b); // (3,4):(1,3)
|
||||
Layout ba = prepend(a, b); // (4,3):(3,1)
|
||||
Layout c = append(ab, ab); // (3,4,(3,4)):(1,3,(1,3))
|
||||
Layout d = replace<2>(c, b); // (3,4,4):(1,3,3)
|
||||
```
|
||||
|
||||
### Grouping and flattening
|
||||
|
||||
Layout modes can be grouped with `group<ModeBegin, ModeEnd>` and flattened with `flatten`.
|
||||
```cpp
|
||||
Layout a = Layout<Shape<_2,_3,_5,_7>>{}; // (_2,_3,_5,_7):(_1,_2,_6,_30)
|
||||
Layout b = group<0,2>(a); // ((_2,_3),_5,_7):((_1,_2),_6,_30)
|
||||
Layout c = group<1,3>(b); // ((_2,_3),(_5,_7)):((_1,_2),(_6,_30))
|
||||
Layout f = flatten(b); // (_2,_3,_5,_7):(_1,_2,_6,_30)
|
||||
Layout e = flatten(c); // (_2,_3,_5,_7):(_1,_2,_6,_30)
|
||||
```
|
||||
Grouping, flattening, and reordering modes allows the reinterpretation of tensors in place as matrices, matrices as vectors, vectors as matrices, etc.
|
||||
|
||||
### Slicing
|
||||
|
||||
`Layout`s can be sliced, but slicing is more appropriate to perform on `Tensor`s. See the [`Tensor` section](./03_tensor.md) for slicing details.
|
||||
|
||||
## Summary
|
||||
|
||||
* The `Shape` of a `Layout` defines its coordinate space(s).
|
||||
|
||||
* Every `Layout` has a 1-D coordinate space.
|
||||
This can be used to iterate over the coordinate spaces in a colexicographical order.
|
||||
|
||||
* Every `Layout` has a R-D coordinate space,
|
||||
where R is the rank of the layout.
|
||||
The colexicographical enumeration of the R-D coordinates
|
||||
correspond to the 1-D coordinates above.
|
||||
|
||||
* Every `Layout` has an h-D (natural) coordinate space where h is "hierarchical." These are ordered colexicographically and the enumeration of that order corresponds to the 1-D coordinates above. A natural coordinate is *congruent* to the `Shape` so that each element of the coordinate has a corresponding element of the `Shape`.
|
||||
|
||||
* The `Stride` of a `Layout` maps coordinates to indices.
|
||||
|
||||
* The inner product of the elements of the natural coordinate with the elements of the `Stride` produces the resulting index.
|
||||
|
||||
For each `Layout` there exists an integral `Shape` that is that compatible with that `Layout`. Namely, that integral shape is `size(layout)`. We can then observe that
|
||||
|
||||
> Layouts are functions from integers to integers.
|
||||
|
||||
If you're familiar with the C++23 feature `mdspan`,
|
||||
this is an important difference between
|
||||
`mdspan` layout mappings and CuTe `Layout`s. In CuTe, `Layout` is a first class citizen, is natively hierarchical to naturally represent functions beyond row-major and column-major, and can similarly be indexed with a hierarchy of coordinates.
|
||||
(`mdspan` layout mappings can represent hierarchical functions as well,
|
||||
but this requires defining a custom layout.)
|
||||
Input coordinates for an `mdspan` must have the same shape as the `mdspan`;
|
||||
a multidimensional `mdspan` does not accept 1-D coordinates.
|
||||
|
||||
## Copyright
|
||||
|
||||
Copyright (c) 2017 - 2025 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
```
|
||||
Redistribution and use in source and binary forms, with or without
|
||||
modification, are permitted provided that the following conditions are met:
|
||||
|
||||
1. Redistributions of source code must retain the above copyright notice, this
|
||||
list of conditions and the following disclaimer.
|
||||
|
||||
2. Redistributions in binary form must reproduce the above copyright notice,
|
||||
this list of conditions and the following disclaimer in the documentation
|
||||
and/or other materials provided with the distribution.
|
||||
|
||||
3. Neither the name of the copyright holder nor the names of its
|
||||
contributors may be used to endorse or promote products derived from
|
||||
this software without specific prior written permission.
|
||||
|
||||
THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS "AS IS"
|
||||
AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE
|
||||
IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE ARE
|
||||
DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT HOLDER OR CONTRIBUTORS BE LIABLE
|
||||
FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL
|
||||
DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR
|
||||
SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER
|
||||
CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY,
|
||||
OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE
|
||||
OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
|
||||
```
|
||||
580
media/docs/cpp/cute/02_layout_algebra.md
Normal file
580
media/docs/cpp/cute/02_layout_algebra.md
Normal file
@@ -0,0 +1,580 @@
|
||||
# CuTe Layout Algebra
|
||||
|
||||
CuTe provides an "algebra of `Layout`s" to support combining layouts in different ways. This algebra includes operations such as
|
||||
|
||||
* `Layout` functional composition,
|
||||
* a notion of `Layout` "product" to reproduce one layout according to another, and
|
||||
* a notion of `Layout` "divide" to split one layout according to another.
|
||||
|
||||
Common utilities for building complicated layouts from simpler ones depend on the `Layout` product. Common utilities for partitioning layouts (of data, for example) across other layouts (of threads, for example) depend on the `Layout` divide. All of these utilities rely on the functional composition of `Layout`s.
|
||||
|
||||
In this section, we'll build up the tools of the `Layout` algebra and explain some of these core operations in detail.
|
||||
|
||||
## Coalesce
|
||||
|
||||
In the previous section, we summarized `Layout`s with
|
||||
> Layouts are functions from integers to integers.
|
||||
|
||||
The `coalesce` operation is a "simplify" on functions from integers to integers. If we only care about input integers, then we can manipulate the shape and number of modes of the `Layout` without changing it as a function. The only thing `coalesce` can't change is the `Layout`'s `size`.
|
||||
|
||||
More specifically, you can find the checked post-conditions in [the `coalesce` unit test](https://github.com/NVIDIA/cutlass/tree/main/test/unit/cute/core/coalesce.cpp), which we'll reproduce here:
|
||||
```cpp
|
||||
// @post size(@a result) == size(@a layout)
|
||||
// @post depth(@a result) <= 1
|
||||
// @post for all i, 0 <= i < size(@a layout), @a result(i) == @a layout(i)
|
||||
Layout coalesce(Layout const& layout)
|
||||
```
|
||||
|
||||
For example,
|
||||
|
||||
```cpp
|
||||
auto layout = Layout<Shape <_2,Shape <_1,_6>>,
|
||||
Stride<_1,Stride<_6,_2>>>{};
|
||||
auto result = coalesce(layout); // _12:_1
|
||||
```
|
||||
|
||||
where we can see the result has fewer modes and is "simpler." Indeed, this could save us a few operations in the coordinate mapping and index mapping (if those are performed dynamically).
|
||||
|
||||
So, how do we get there?
|
||||
|
||||
* We've already seen that column-major `Layout`s like `(_2,_4):(_1,_2)` act identically to `_8:_1` for 1-D coordinates.
|
||||
* Modes with size static-1 will always produce a natural coordinate of static-0. They can be ignored no matter the stride.
|
||||
|
||||
Generalizing, consider a layout with just two integral modes, s0:d0 and s1:d1. Denote the result of coalescing this layout as s0:d0 ++ s1:d1. Then, there are four cases:
|
||||
|
||||
1. `s0:d0 ++ _1:d1 => s0:d0`. Ignore modes with size static-1.
|
||||
2. `_1:d0 ++ s1:d1 => s1:d1`. Ignore modes with size static-1.
|
||||
3. `s0:d0 ++ s1:s0*d0 => s0*s1:d0`. If the second mode's stride is the product of the first mode's size and stride, then they can be combined.
|
||||
4. `s0:d0 ++ s1:d1 => (s0,s1):(d0,d1)`. Else, nothing can be done and they must be treated separately.
|
||||
|
||||
That's it! We can flatten any layout and apply the above binary operation to each pair of adjacent modes in order to "coalesce" the modes of the layout.
|
||||
|
||||
### By-mode Coalesce
|
||||
|
||||
Obviously, sometimes we do care about the shape of our `Layout`, but would still like to coalesce. For example, I have a 2-D `Layout` and I would like the result to remain 2-D.
|
||||
|
||||
For this reason, there's an overload of `coalesce` that takes an additional parameter
|
||||
```cpp
|
||||
// Apply coalesce at the terminals of trg_profile
|
||||
Layout coalesce(Layout const& layout, IntTuple const& trg_profile)
|
||||
```
|
||||
|
||||
which can be used as follows
|
||||
|
||||
```cpp
|
||||
auto a = Layout<Shape <_2,Shape <_1,_6>>,
|
||||
Stride<_1,Stride<_6,_2>>>{};
|
||||
auto result = coalesce(a, Step<_1,_1>{}); // (_2,_6):(_1,_2)
|
||||
// Identical to
|
||||
auto same_r = make_layout(coalesce(layout<0>(a)),
|
||||
coalesce(layout<1>(a)));
|
||||
```
|
||||
|
||||
This function is recursing into `Step<_1,_1>{}` and applying `coalesce` to the corresponding sublayout whenever it sees an integer (the values don't matter, they're just flags) rather than a tuple.
|
||||
|
||||
> This theme of defining an operation that treats a `Layout` as a "1-D" function from integers to integers and then generalizing to use it for an arbitrarily shaped layout will be a common one!
|
||||
|
||||
## Composition
|
||||
|
||||
Functional composition of `Layout`s is the core of CuTe and is used in just about every higher-level operation.
|
||||
|
||||
Starting again from the observation that `Layout`s are just functions from integers to integers, we can define functional composition that results in another `Layout`. First, an example.
|
||||
|
||||
```text
|
||||
Functional composition, R := A o B
|
||||
R(c) := (A o B)(c) := A(B(c))
|
||||
|
||||
Example
|
||||
A = (6,2):(8,2)
|
||||
B = (4,3):(3,1)
|
||||
|
||||
R( 0) = A(B( 0)) = A(B(0,0)) = A( 0) = A(0,0) = 0
|
||||
R( 1) = A(B( 1)) = A(B(1,0)) = A( 3) = A(3,0) = 24
|
||||
R( 2) = A(B( 2)) = A(B(2,0)) = A( 6) = A(0,1) = 2
|
||||
R( 3) = A(B( 3)) = A(B(3,0)) = A( 9) = A(3,1) = 26
|
||||
R( 4) = A(B( 4)) = A(B(0,1)) = A( 1) = A(1,0) = 8
|
||||
R( 5) = A(B( 5)) = A(B(1,1)) = A( 4) = A(4,0) = 32
|
||||
R( 6) = A(B( 6)) = A(B(2,1)) = A( 7) = A(1,1) = 10
|
||||
R( 7) = A(B( 7)) = A(B(3,1)) = A(10) = A(4,1) = 34
|
||||
R( 8) = A(B( 8)) = A(B(0,2)) = A( 2) = A(2,0) = 16
|
||||
R( 9) = A(B( 9)) = A(B(1,2)) = A( 5) = A(5,0) = 40
|
||||
R(10) = A(B(10)) = A(B(2,2)) = A( 8) = A(2,1) = 18
|
||||
R(11) = A(B(11)) = A(B(3,2)) = A(11) = A(5,1) = 42
|
||||
```
|
||||
|
||||
The absolutely amazing observation is that the function `R(c) = k` defined above can be written down as another `Layout`
|
||||
|
||||
```
|
||||
R = ((2,2),3):((24,2),8)
|
||||
```
|
||||
|
||||
AND
|
||||
|
||||
```
|
||||
compatible(B, R)
|
||||
```
|
||||
|
||||
That is, every coordinate of `B` can also be used as a coordinate of `R`. This is an expected property of functional composition because `B` defines the *domain* of `R`.
|
||||
|
||||
You can find many examples and checked post-conditions in [the `composition` unit test](https://github.com/NVIDIA/cutlass/tree/main/test/unit/cute/core/composition.cpp). The post-conditions are precisely as we just stated.
|
||||
```cpp
|
||||
// @post compatible(@a layout_b, @a result)
|
||||
// @post for all i, 0 <= i < size(@a layout_b), @a result(i) == @a layout_a(@a layout_b(i)))
|
||||
Layout composition(LayoutA const& layout_a, LayoutB const& layout_b)
|
||||
```
|
||||
|
||||
### Computing Composition
|
||||
|
||||
First, a few observations:
|
||||
|
||||
* `B = (B_0, B_1, ...)`. A layout can be expressed as the concatenation of its sublayouts.
|
||||
|
||||
* `A o B = A o (B_0, B_1, ...) = (A o B_0, A o B_1, ...)`. When `B` is injective, composition is left-distributive with concatenation.
|
||||
|
||||
With the above, we can assume without loss of generality that `B = s:d` is a layout with integral shape and stride. We can also assume that `A` is a flattened, coalesced layout.
|
||||
|
||||
When `A` is integral, `A = a:b`, the result is rather trivial: `R = A o B = a:b o s:d = s:(b*d)`. But when `A` is multimodal, we need to be more careful.
|
||||
|
||||
Put into words, `A o B = A o s:d`, for integral `s` and `d` means that we want (1) "remove" the first `d` elements from `A`, and then (2) "keep" the first `s` of those strided elements.
|
||||
|
||||
1. Removing the first `d` elements of `A` can be computed by progressively "dividing out" the first `d` elements from the shape of `A` starting from the left. For example,
|
||||
* `(6,2) / 2 => (3,2)`
|
||||
* `(6,2) / 3 => (2,2)`
|
||||
* `(6,2) / 6 => (1,2)`
|
||||
* `(6,2) / 12 => (1,1)`
|
||||
* `(3,6,2,8) / 3 => (1,3,2,8)`
|
||||
* `(3,6,2,8) / 6 => (1,3,2,8)`
|
||||
* `(3,6,2,8) / 9 => (1,2,2,8)`
|
||||
* `(3,6,2,8) / 72 => (1,1,1,4)`
|
||||
|
||||
As you may have noticed, we can only divide shapes by certain values and get a sensible result. This is called the **stride divisibility condition** and is statically checked in CuTe when possible.
|
||||
|
||||
2. Keeping the first `s` elements of the strided `A` layout can be computed by "modding out" the first `s` elements from the shape of `A` starting from the left. For example,
|
||||
* `(6,2) % 2 => (2,1)`
|
||||
* `(6,2) % 3 => (3,1)`
|
||||
* `(6,2) % 6 => (6,1)`
|
||||
* `(6,2) % 12 => (6,2)`
|
||||
* `(3,6,2,8) % 6 => (3,2,1,1)`
|
||||
* `(3,6,2,8) % 9 => (3,3,1,1)`
|
||||
* `(1,2,2,8) % 2 => (1,2,1,1)`
|
||||
* `(1,2,2,8) % 16 => (1,2,2,4)`
|
||||
|
||||
Again, this operation must satisfy a **shape divisibility condition** to yield a sensible result and is statically checked in CuTe when possible.
|
||||
|
||||
From the above examples, we can construct the composition `(3,6,2,8):(w,x,y,z) o 16:9 = (1,2,2,4):(3*w,3*x,y,z)`.
|
||||
|
||||
#### Example 1 -- Reshape a layout into a matrix
|
||||
|
||||
`20:2 o (5,4):(4,1)`. Composition formulation.
|
||||
|
||||
This describes interpreting the layout `20:2`
|
||||
as a 5x4 matrix in a row-major order.
|
||||
|
||||
1. ` = 20:2 o (5:4,4:1)`. Layout `(5,4):(4,1)` as concatenation of sublayouts.
|
||||
|
||||
2. ` = (20:2 o 5:4, 20:2 o 4:1)`. Left distributivity.
|
||||
|
||||
* `20:2 o 5:4 => 5:8`. Trivial case.
|
||||
* `20:2 o 4:1 => 4:2`. Trivial case.
|
||||
|
||||
3. ` = (5:8, 4:2)`. Composed Layout as concatenation of sublayouts.
|
||||
|
||||
4. ` = (5,4):(8,2)`. Final composed layout.
|
||||
|
||||
#### Example 2 -- Reshape a layout into a matrix
|
||||
|
||||
`(10,2):(16,4) o (5,4):(1,5)`
|
||||
|
||||
This describes interpreting the layout `(10,2):(16,4)`
|
||||
as a 5x4 matrix in a column-major order.
|
||||
|
||||
1. ` = (10,2):(16,4) o (5:1,4:5)`. Layout `(5,4):(1,5)` as concatenation of sublayouts.
|
||||
|
||||
2. ` = ((10,2):(16,4) o 5:1, (10,2):(16,4) o 4:5)`. Left distributivity.
|
||||
|
||||
* `(10,2):(16,4) o 5:1 => (5,1):(16,4)`. Mod out the shape `5`.
|
||||
* `(10,2):(16,4) o 4:5 => (2,2):(80,4)`. Div out the stride `5`.
|
||||
|
||||
3. ` = ((5,1):(16,4), (2,2):(80,4))`. Composed Layout as concatenation of sublayouts.
|
||||
|
||||
4. ` = (5:16, (2,2):(80,4))`. By-mode coalesce.
|
||||
|
||||
5. ` = (5,(2,2))):(16,(80,4))`. Final composed layout.
|
||||
|
||||
We get exactly this result with CuTe
|
||||
if we use compile-time shapes and strides.
|
||||
The following C++ code prints `(_5,(_2,_2)):(_16,(_80,_4))`.
|
||||
|
||||
```cpp
|
||||
Layout a = make_layout(make_shape (Int<10>{}, Int<2>{}),
|
||||
make_stride(Int<16>{}, Int<4>{}));
|
||||
Layout b = make_layout(make_shape (Int< 5>{}, Int<4>{}),
|
||||
make_stride(Int< 1>{}, Int<5>{}));
|
||||
Layout c = composition(a, b);
|
||||
print(c);
|
||||
```
|
||||
|
||||
If we use dynamic integers, the following C++ code prints `((5,1),(2,2)):((16,4),(80,4))`.
|
||||
|
||||
```cpp
|
||||
Layout a = make_layout(make_shape (10, 2),
|
||||
make_stride(16, 4));
|
||||
Layout b = make_layout(make_shape ( 5, 4),
|
||||
make_stride( 1, 5));
|
||||
Layout c = composition(a, b);
|
||||
print(c);
|
||||
```
|
||||
|
||||
The results may _look_ different but are the mathematically the same. The 1s in the shape don't affect the layout as a mathematical function from 1-D coordinates to integers or as a function from 2-D coordinates to integers. In the dynamic case, CuTe can not coalesce the dynamic size-1 modes to "simplify" the layout due to the static rank and type of the tuples containing them.
|
||||
|
||||
### By-mode Composition
|
||||
|
||||
Similar to by-mode `coalesce` and building up to a generic tiling operation, sometimes we do care about the shape of the `A` layout and would still like to apply `composition` to individual modes. For example, I have a 2-D `Layout` and would like some sublayout of the elements down the columns and another sublayout of elements across the rows.
|
||||
|
||||
For this reason, `composition` also works when its second parameter -- the `B` -- is a `Tiler`. In general, a tiler is a layout or a tuple-of-layouts (note the generalization on `IntTuple`), which can be used as follows
|
||||
```cpp
|
||||
// (12,(4,8)):(59,(13,1))
|
||||
auto a = make_layout(make_shape (12,make_shape ( 4,8)),
|
||||
make_stride(59,make_stride(13,1)));
|
||||
// <3:4, 8:2>
|
||||
auto tiler = make_tile(Layout<_3,_4>{}, // Apply 3:4 to mode-0
|
||||
Layout<_8,_2>{}); // Apply 8:2 to mode-1
|
||||
|
||||
// (_3,(2,4)):(236,(26,1))
|
||||
auto result = composition(a, tiler);
|
||||
// Identical to
|
||||
auto same_r = make_layout(composition(layout<0>(a), get<0>(tiler)),
|
||||
composition(layout<1>(a), get<1>(tiler)));
|
||||
```
|
||||
We often use the `<LayoutA, LayoutB, ...>` notation to distinguish `Tiler`s from the concatenation-of-sublayouts notation `(LayoutA, LayoutB, ...)` that we used previously.
|
||||
|
||||
The `result` in the above code can be depicted as the 3x8 sublayout of the original layout highlighted in the figure below.
|
||||
<p align="center">
|
||||
<img src="../../images/cute/composition1.png" alt="composition1.png" height="250"/>
|
||||
</p>
|
||||
|
||||
For convenience, CuTe also interprets `Shape`s as a tiler as well. A `Shape` is interpreted as tuple-of-layouts-with-stride-1:
|
||||
```cpp
|
||||
// (12,(4,8)):(59,(13,1))
|
||||
auto a = make_layout(make_shape (12,make_shape ( 4,8)),
|
||||
make_stride(59,make_stride(13,1)));
|
||||
// (8, 3)
|
||||
auto tiler = make_shape(Int<3>{}, Int<8>{});
|
||||
// Equivalent to <3:1, 8:1>
|
||||
// auto tiler = make_tile(Layout<_3,_1>{}, // Apply 3:1 to mode-0
|
||||
// Layout<_8,_1>{}); // Apply 8:1 to mode-1
|
||||
|
||||
// (_3,(4,2)):(59,(13,1))
|
||||
auto result = composition(a, tiler);
|
||||
```
|
||||
where `result` can be depicted as the 3x8 sublayout of the original layout highlighted in the figure below.
|
||||
<p align="center">
|
||||
<img src="../../images/cute/composition2.png" alt="composition2.png" height="250"/>
|
||||
</p>
|
||||
|
||||
## Composition Tilers
|
||||
|
||||
In summary, a `Tiler` is one of the following objects.
|
||||
1. A `Layout`.
|
||||
2. A tuple of `Tiler`s.
|
||||
3. A `Shape`, which will be interpreted as a tiler of `Layout`s with stride-1.
|
||||
|
||||
Any of the above can be used as the second argument in `composition`. With (1), we think of the `composition` as between two functions from integers to integers, no matter the ranks of the layouts. With (2) and (3), the `composition` is performed on each pair of corresponding modes of `A` and `B`, until case (1) is found.
|
||||
|
||||
This allows composition to be applied by-mode to retrieve arbitrary sublayouts of specified modes of a tensor ("Give me the 3x5x8 subblock of this MxNxL tensor") but also allows entire tiles of data to be reshaped and reordered as if they were 1-D vectors ("Reorder this 8x16 block of data into a 32x4 block using this weird order of elements"). We will see the by-mode cases appear often when we are tiling for threadblocks in examples that follow. We will see 1-D reshaping and reordering when we want to apply arbitrary partitioning patterns for threads and values in MMAs in examples that follow.
|
||||
|
||||
## Complement
|
||||
|
||||
Before getting to "product" and "divide," we need one more operation. We can think of `composition` as a layout `B` that is "selecting" certain coordinates from another layout `A`. But what about the coordinates that aren't "selected"? To implement generic tiling, we want to be able to select arbitrary elements -- the tile -- and to describe the layout of those tiles -- the leftovers, or the "rest."
|
||||
|
||||
The `complement` of a layout attempts to find another layout that represents the "rest" -- the elements that aren't touched by the layout.
|
||||
|
||||
You can find many examples and checked post-conditions in [the `complement` unit test](https://github.com/NVIDIA/cutlass/tree/main/test/unit/cute/core/complement.cpp). The post-conditions include
|
||||
```cpp
|
||||
// @post cosize(make_layout(@a layout_a, @a result))) >= size(@a cotarget)
|
||||
// @post cosize(@a result) >= round_up(size(@a cotarget), cosize(@a layout_a))
|
||||
// @post for all i, 1 <= i < size(@a result),
|
||||
// @a result(i-1) < @a result(i)
|
||||
// @post for all i, 1 <= i < size(@a result),
|
||||
// for all j, 0 <= j < size(@a layout_a),
|
||||
// @a result(i) != @a layout_a(j)
|
||||
Layout complement(LayoutA const& layout_a, Shape const& cotarget)
|
||||
```
|
||||
That is, the complement `R` of a layout `A` with respect to a Shape (IntTuple) `M` satisfies the following properties.
|
||||
1. The size (and cosize) of `R` is *bounded* by `size(M)`.
|
||||
2. `R` is *ordered*. That is, the strides of `R` are positive and increasing. This means that `R` is unique.
|
||||
3. `A` and `R` have *disjoint* codomains. `R` attempts to "complete" the codomain of `A`.
|
||||
|
||||
The `cotarget` parameter above is most commonly an integer -- you can see we only use `size(cotarget)` above. However, sometimes it is useful to specify an integer that has static properties. For example, `28` is a dynamic integer and `(_4,7)` is a shape with size `28` that is statically known to be divisible by `_4`. Both will produce the same `complement` mathematically, but the extra information can used by `complement` to preserve the staticness of the result as much as possible.
|
||||
|
||||
### Complement Examples
|
||||
|
||||
`complement` is most effective on static shapes and strides, so consider all integers below to be static. Similar examples for dynamic shapes and strides as well as IntTuple `cotarget` can be found in [the unit test](https://github.com/NVIDIA/cutlass/tree/main/test/unit/cute/core/complement.cpp).
|
||||
|
||||
* `complement(4:1, 24)` is `6:4`. Note that `(4,6):(1,4)` has cosize `24`. The layout `4:1` is effectively repeated 6 times with `6:4`.
|
||||
|
||||
* `complement(6:4, 24)` is `4:1`. Note that `(6,4):(4,1)` has cosize `24`. The "hole" in `6:4` is filled with `4:1`.
|
||||
|
||||
* `complement((4,6):(1,4), 24)` is `1:0`. Nothing needs to be appended.
|
||||
|
||||
* `complement(4:2, 24)` is `(2,3):(1,8)`. Note that `(4,(2,3)):(2,(1,8))` has cosize `24`. The "hole" in `4:2` is filled with `2:1` first, then everything is repeated 3 times with `3:8`.
|
||||
|
||||
* `complement((2,4):(1,6), 24)` is `3:2`. Note that `((2,4),3):((1,6),2)` has cosize `24` and produces unique indices.
|
||||
|
||||
* `complement((2,2):(1,6), 24)` is `(3,2):(2,12)`. Note that `((2,2),(3,2)):((1,6),(2,12))` has cosize `24` and produces unique indices.
|
||||
|
||||
<p align="center">
|
||||
<img src="../../images/cute/complement1.png" alt="complement1.png" height="75"/>
|
||||
</p>
|
||||
As a visualization, the above figure depicts the codomain of the last example. The image of the original layout `(2,2):(1,6)` is colored in gray. The complement effectively "repeats" the original layout (displayed in the other colors) such that the codomain size of the result is `24`. The complement `(3,2):(2,12)` can be viewed as the "layout of the repetition."
|
||||
|
||||
## Division (Tiling)
|
||||
|
||||
Finally, we can define the division of a `Layout` by another `Layout`. Functions that divide a layout into components are useful as a basis for tiling and partitioning layouts.
|
||||
|
||||
In this section, we'll define `logical_divide(Layout, Layout)`, which again considers all `Layout`s as 1-D functions from integers to integers, and then use that definition to create multidimensional `Layout` divides.
|
||||
|
||||
Informally, `logical_divide(A, B)` splits a layout `A` into two modes -- in the first mode are all elements pointed to by `B` and in the second mode are all elements not pointed to by `B`.
|
||||
|
||||
Formally, this can be written as
|
||||
|
||||
$A \oslash B := A \circ (B,B^*)$
|
||||
|
||||
and implemented as
|
||||
```cpp
|
||||
template <class LShape, class LStride,
|
||||
class TShape, class TStride>
|
||||
auto logical_divide(Layout<LShape,LStride> const& layout,
|
||||
Layout<TShape,TStride> const& tiler)
|
||||
{
|
||||
return composition(layout, make_layout(tiler, complement(tiler, size(layout))));
|
||||
}
|
||||
```
|
||||
Note that this is defined only in terms of concatenation, composition, and complement.
|
||||
|
||||
So what is that?
|
||||
|
||||
> in the first mode are all elements pointed to by `B`
|
||||
|
||||
This is clearly composition, `A o B`.
|
||||
|
||||
> in the second mode are all elements not pointed to by `B`
|
||||
|
||||
The elements NOT pointed to by `B` sounds like a complement, `B*`, up to the size of `A`. As we've seen above in the `complement` section, this can be described as the "layout of the repetition of `B`." If `B` is the "tiler", then `B*` is the layout of the tiles.
|
||||
|
||||
### Logical Divide 1-D Example
|
||||
|
||||
Consider tiling the 1-D layout `A = (4,2,3):(2,1,8)` with the tiler `B = 4:2`. Informally, this means that we have a 1-D vector of 24 elements in some storage order defined by `A` and we want to extract tiles of 4 elements strided by 2.
|
||||
|
||||
This is computed in the three steps described in the implementation above.
|
||||
* Complement of `B = 4:2` under `size(A) = 24` is `B* = (2,3):(1,8)`.
|
||||
* Concantenation of `(B,B*) = (4,(2,3)):(2,(1,8))`.
|
||||
* Composition of `A = (4,2,3):(2,1,8)` with `(B,B*)` is then `((2,2),(2,3)):((4,1),(2,8))`.
|
||||
|
||||
<p align="center">
|
||||
<img src="../../images/cute/divide1.png" alt="divide1.png" height="150"/>
|
||||
</p>
|
||||
|
||||
The above figure depicts `A` as a 1-D layout with the elements pointed to by `B` highlighted in gray. The layout `B` describes our "tile" of data, and there are six of those tiles in `A` shown by each of the colors. After the divide, the first mode of the result is the tile of data and the second mode of the result iterates over each tile.
|
||||
|
||||
### Logical Divide 2-D Example
|
||||
|
||||
Using the `Tiler` concept defined above, this immediately generalizes to multidimensional tiling. The below example simply applies `layout_divide` by-mode to the cols and rows of a 2-D layout using a `Tiler`.
|
||||
|
||||
Similar to the 2-D composition example above, consider a 2-D layout `A = (9,(4,8)):(59,(13,1))` and want to apply `3:3` down the columns (mode-0) and `(2,4):(1,8)` across the rows (mode-1). This means the tiler can be written as `B = <3:3, (2,4):(1,8)>`.
|
||||
|
||||
<p align="center">
|
||||
<img src="../../images/cute/divide2.png" alt="divide2.png" height="450"/>
|
||||
</p>
|
||||
|
||||
The above figure depicts `A` as a 2-D layout with the elements pointed to by `B` highlighted in gray. The layout `B` describes our "tile" of data, and there are twelve of those tiles in `A` shown by each of the colors. After the divide, the first mode of each mode of the result is the tile of data and the second mode of each mode iterates over each tile. In that sense, this operation can be viewed as a kind of `gather` operation or as simply a permutation on the rows and cols.
|
||||
|
||||
Note that the first mode of each mode of the result is the sublayout `(3,(2,4)):(177,(13,2))` and is precisely the result we would have received if we had applied `composition` instead of `logical_divide`.
|
||||
|
||||
### Zipped, Tiled, Flat Divides
|
||||
|
||||
It's easy to see the tiles when they are highlighted in the images above, but working with them can still be awkward. How would you slice out the `3`rd tile or the `7`th tile or the `(1,2)`th tile so you could continue working on it?
|
||||
|
||||
Enter the convenience flavors of `logical_divide`. Suppose we have a `Layout` and a `Tiler` of some shape, then each operation will apply `logical_divide`, but potentially rearrange the modes into more convenient forms.
|
||||
```text
|
||||
Layout Shape : (M, N, L, ...)
|
||||
Tiler Shape : <TileM, TileN>
|
||||
|
||||
logical_divide : ((TileM,RestM), (TileN,RestN), L, ...)
|
||||
zipped_divide : ((TileM,TileN), (RestM,RestN,L,...))
|
||||
tiled_divide : ((TileM,TileN), RestM, RestN, L, ...)
|
||||
flat_divide : (TileM, TileN, RestM, RestN, L, ...)
|
||||
```
|
||||
|
||||
For example, the `zipped_divide` function applies `logical_divide`, and then gathers the "subtiles" into a single mode and the "rest" into a single mode.
|
||||
```cpp
|
||||
// A: shape is (9,32)
|
||||
auto layout_a = make_layout(make_shape (Int< 9>{}, make_shape (Int< 4>{}, Int<8>{})),
|
||||
make_stride(Int<59>{}, make_stride(Int<13>{}, Int<1>{})));
|
||||
// B: shape is (3,8)
|
||||
auto tiler = make_tile(Layout<_3,_3>{}, // Apply 3:3 to mode-0
|
||||
Layout<Shape <_2,_4>, // Apply (2,4):(1,8) to mode-1
|
||||
Stride<_1,_8>>{});
|
||||
|
||||
// ((TileM,RestM), (TileN,RestN)) with shape ((3,3), (8,4))
|
||||
auto ld = logical_divide(layout_a, tiler);
|
||||
// ((TileM,TileN), (RestM,RestN)) with shape ((3,8), (3,4))
|
||||
auto zd = zipped_divide(layout_a, tiler);
|
||||
```
|
||||
Then, the offset to the `3`rd tile is `zd(0,3)`. The offset to the `7`th tile is `zd(0,7)`. The offset to the `(1,2)`th tile is `zd(0,make_coord(1,2))`. The tile itself always has layout `layout<0>(zd)`. Indeed, it is always the case that
|
||||
|
||||
`layout<0>(zipped_divide(a, b)) == composition(a, b)`.
|
||||
|
||||
We note that `logical_divide` preserves the *semantics* of the modes while permuting the elements within those modes -- the `M`-mode of layout `A` is still the `M`-mode of the result and the `N`-mode of layout `A` is still the `N`-mode of the result.
|
||||
|
||||
This is not the case with `zipped_divide`. The mode-0 in the `zipped_divide` result is the `Tile` itself (of whatever rank the `Tiler` was) and mode-1 is the layout of those tiles. It doesn't always make sense to plot these as 2-D layouts, because the `M`-mode is now more aptly the "tile-mode" and the `N`-mode is more aptly the "rest-mode". Regardless, we still can plot the resulting layout as 2-D as shown below.
|
||||
|
||||
<p align="center">
|
||||
<img src="../../images/cute/divide3.png" alt="divide3.png" height="450"/>
|
||||
</p>
|
||||
|
||||
We've kept each tile as its color in the previous images for clarity. Clearly, iterating across tiles is now equivalent to iterating across a row of this layout and iterating over elements within a tile is equivalent to iterating down a column of this layout. As we'll see in the `Tensor` section, this can be used to great effect in partitioning within or across tiles of data.
|
||||
|
||||
## Product (Tiling)
|
||||
|
||||
Finally, we can define the product of a Layout by another Layout. In this section, we'll define `logical_product(Layout, Layout)`, which again considers all `Layout`s as 1-D functions from integers to integers, and then use that definition to create multidimensional `Layout` products.
|
||||
|
||||
Informally, `logical_product(A, B)` results in a two mode layout where the first mode is the layout `A` and the second mode is the layout `B` but with each element replaced by a "unique replication" of layout `A`.
|
||||
|
||||
Formally, this can be written as
|
||||
|
||||
$A \otimes B := (A, A^* \circ B)$
|
||||
|
||||
and implemented in CuTe as
|
||||
```cpp
|
||||
template <class LShape, class LStride,
|
||||
class TShape, class TStride>
|
||||
auto logical_product(Layout<LShape,LStride> const& layout,
|
||||
Layout<TShape,TStride> const& tiler)
|
||||
{
|
||||
return make_layout(layout, composition(complement(layout, size(layout)*cosize(tiler)), tiler));
|
||||
}
|
||||
```
|
||||
Note that this is defined only in terms of concatenation, composition, and complement.
|
||||
|
||||
So what is that?
|
||||
|
||||
> where the first mode is the layout `A`
|
||||
|
||||
This is clearly just a copy of `A`.
|
||||
|
||||
> the second mode is the layout `B` but with each element replaced by a "unique replication" of layout `A`.
|
||||
|
||||
The "unique replication" of layout `A` sounds like complement, `A*`, up to the cosize of `B`. As we've seen in the `complement` section, this can be described as the "layout of the repetition of `A`". If `A` is the "tile", then `A*` is the layout of repetitions that are available for `B`.
|
||||
|
||||
### Logical Product 1-D Example
|
||||
|
||||
Consider reproducing the 1-D layout `A = (2,2):(4,1)` according to `B = 6:1`. Informally, this means that we have a 1-D layout of 4 elements defined by `A` and we want to reproduce it 6 times.
|
||||
|
||||
This is computed in the three steps described in the implementation above.
|
||||
* Complement of `A = (2,2):(4,1)` under `6*4 = 24` is `A* = (2,3):(2,8)`.
|
||||
* Composition of `A* = (2,3):(2,8)` with `B = 6:1` is then `(2,3):(2,8)`.
|
||||
* Concatenation of `(A,A* o B) = ((2,2),(2,3)):((4,1),(2,8))`.
|
||||
|
||||
<p align="center">
|
||||
<img src="../../images/cute/product1.png" alt="product1.png" height="175"/>
|
||||
</p>
|
||||
|
||||
The above figure depicts `A` and `B` as a 1-D layouts. The layout `B` describes the number and order of repetitions of `A` and they are colored for clarity. After the product, the first mode of the result is the tile of data and the second mode of the result iterates over each tile.
|
||||
|
||||
Note that the result is identical to the result of the 1-D Logical Divide example.
|
||||
|
||||
Of course, we can change the number and order of the tiles in the product by changing `B`.
|
||||
|
||||
<p align="center">
|
||||
<img src="../../images/cute/product2.png" alt="product2.png" height="175"/>
|
||||
</p>
|
||||
|
||||
For example, in the above image with `B = (4,2):(2,1)`, there are 8 repeated tiles instead of 6 and the tiles are in a different order.
|
||||
|
||||
### Logical Product 2-D Example
|
||||
|
||||
We can use the by-mode `tiler` strategies previously developed to write multidimensional products as well.
|
||||
|
||||
<p align="center">
|
||||
<img src="../../images/cute/product2d.png" alt="product2d.png" height="250"/>
|
||||
</p>
|
||||
|
||||
The above image demonstates the use of a `tiler` to apply `logical_product` by-mode. Despite this **not being the recommended approach**, the result is a rank-2 layout consisting of 2x5 row-major block that is tiled across a 3x4 column-major arrangement.
|
||||
|
||||
The reason **this is not the recommended approach** is that the `tiler B` in the above expression is highly unintuitive. In fact, it requires perfect knowledge of the shape and strides of `A` in order to construct. We would like to express "Tile Layout `A` according to Layout `B`" in a way that makes `A` and `B` independent and is much more intuitive.
|
||||
|
||||
#### Blocked and Raked Products
|
||||
|
||||
The `blocked_product(LayoutA, LayoutB)` and `raked_product(LayoutA, LayoutB)` are rank-sensitive transformations on top of 1-D `logical_product` that let us express the more intuitive `Layout` products that we most often want to express.
|
||||
|
||||
A key observation in the implementation of these functions are the compatibility post-conditions of `logical_product`:
|
||||
```
|
||||
// @post rank(result) == 2
|
||||
// @post compatible(layout_a, layout<0>(result))
|
||||
// @post compatible(layout_b, layout<1>(result))
|
||||
```
|
||||
|
||||
Because `A` is always compatible with mode-0 of the result and `B` is always compatible with mode-1 of the result, if we made `A` and `B` the same rank then we could "reassociate" like-modes after the product. That is, the "column" mode in `A` could be combined with the "column" mode in `B` and the "row" mode in `A` could be combined with the "row" mode in `B`, etc.
|
||||
|
||||
This is exactly what `blocked_product` and `raked_product` do and it is why they are called rank-sensitive. Unlike other CuTe functions that take `Layout` arguments, these care about the top-level rank of the arguments so that each mode can be reassociated after the `logical_product`.
|
||||
|
||||
<p align="center">
|
||||
<img src="../../images/cute/productblocked2d.png" alt="productblocked2d.png" height="250"/>
|
||||
</p>
|
||||
|
||||
The above image shows the same result as the `tiler` approach, but with much more intuitive arguments. A 2x5 row-major layout is arranged as a tile in a 3x4 column-major arrangement. Also note that `blocked_product` went ahead and `coalesced` mode-0 for us.
|
||||
|
||||
Similarly, `raked_product` combines the modes slightly differently. Instead of the resulting "column" mode being constructed from the `A` "column" mode then the `B` "column" mode, the resulting "column" mode is constructed from the `B` "column" mode then the `A` "column" mode.
|
||||
|
||||
<p align="center">
|
||||
<img src="../../images/cute/productraked2d.png" alt="productraked2d.png" height="250"/>
|
||||
</p>
|
||||
|
||||
This results in the "tile" `A` now being interleaved or "raked" with the "layout-of-tiles" `B` instead of appearing as blocks. Other references call this a "cyclic distribution."
|
||||
|
||||
### Zipped and Tiled Products
|
||||
|
||||
Similar to `zipped_divide` and `tiled_divide`, the `zipped_product` and `tiled_product` simply rearrange the modes that result from a by-mode `logical_product`.
|
||||
|
||||
```text
|
||||
Layout Shape : (M, N, L, ...)
|
||||
Tiler Shape : <TileM, TileN>
|
||||
|
||||
logical_product : ((M,TileM), (N,TileN), L, ...)
|
||||
zipped_product : ((M,N), (TileM,TileN,L,...))
|
||||
tiled_product : ((M,N), TileM, TileN, L, ...)
|
||||
flat_product : (M, N, TileM, TileN, L, ...)
|
||||
```
|
||||
|
||||
## Copyright
|
||||
|
||||
Copyright (c) 2017 - 2025 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
```
|
||||
Redistribution and use in source and binary forms, with or without
|
||||
modification, are permitted provided that the following conditions are met:
|
||||
|
||||
1. Redistributions of source code must retain the above copyright notice, this
|
||||
list of conditions and the following disclaimer.
|
||||
|
||||
2. Redistributions in binary form must reproduce the above copyright notice,
|
||||
this list of conditions and the following disclaimer in the documentation
|
||||
and/or other materials provided with the distribution.
|
||||
|
||||
3. Neither the name of the copyright holder nor the names of its
|
||||
contributors may be used to endorse or promote products derived from
|
||||
this software without specific prior written permission.
|
||||
|
||||
THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS "AS IS"
|
||||
AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE
|
||||
IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE ARE
|
||||
DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT HOLDER OR CONTRIBUTORS BE LIABLE
|
||||
FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL
|
||||
DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR
|
||||
SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER
|
||||
CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY,
|
||||
OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE
|
||||
OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
|
||||
```
|
||||
425
media/docs/cpp/cute/03_tensor.md
Normal file
425
media/docs/cpp/cute/03_tensor.md
Normal file
@@ -0,0 +1,425 @@
|
||||
# CuTe Tensors
|
||||
|
||||
This document describes `Tensor`, CuTe's core container that deploys the `Layout` concepts previously described.
|
||||
|
||||
Fundamentally, a `Tensor` represents a multidimensional array. `Tensor`s abstracts away the details of how the array's elements are organized and how the array's elements are stored. This lets users write algorithms that access multidimensional arrays generically and potentially specialize algorithms on a `Tensor`s traits. For example, the rank of the `Tensor` can be dispatched against, the `Layout` of data can be inspected, and the type of data can be verified.
|
||||
|
||||
A `Tensor` is represented by two template parameters: `Engine` and `Layout`.
|
||||
For a description of `Layout`, please refer to [the `Layout` section](./01_layout.md).
|
||||
The `Tensor` presents the same shape and access operators as the `Layout` and uses the result of the `Layout` computation to
|
||||
offset and dereference a random-access iterator held by the `Engine`.
|
||||
That is, the layout of the data is provided by `Layout` and the actual data is provided by the iterator. Such data can live in any kind of memory -- global memory, shared memory, register memory -- or can even be transformed or generated on the fly.
|
||||
|
||||
## Fundamental operations
|
||||
|
||||
CuTe `Tensor` provides container-like operations for accessing elements.
|
||||
|
||||
* `.data()`. The iterator this `Tensor` holds.
|
||||
|
||||
* `.size()`. The total logical size of this `Tensor`.
|
||||
|
||||
* `.operator[](Coord)`. Access the element corresponding to the logical coordinate `Coord`.
|
||||
|
||||
* `.operator()(Coord)`. Access the element corresponding to the logical coordinate `Coord`.
|
||||
|
||||
* `.operator()(Coords...)`. Access the element corresponding to the logical coordinate `make_coord(Coords...)`.
|
||||
|
||||
CuTe `Tensor` provides a similar core of hierarchical operations as `Layout`.
|
||||
|
||||
* `rank<I...>(Tensor)`. The rank of the `I...`th mode of the `Tensor`.
|
||||
|
||||
* `depth<I...>(Tensor)`. The depth of the `I...`th mode of the `Tensor`.
|
||||
|
||||
* `shape<I...>(Tensor)`. The shape of the `I...`th mode of the `Tensor`.
|
||||
|
||||
* `size<I...>(Tensor)`. The size of the `I...`th mode of the `Tensor`.
|
||||
|
||||
* `layout<I...>(Tensor)`. The layout of the `I...`th mode of the `Tensor`.
|
||||
|
||||
* `tensor<I...>(Tensor)`. The subtensor corresponding to the the `I...`th mode of the `Tensor`.
|
||||
|
||||
## Tensor Engines
|
||||
|
||||
The `Engine` concept is a wrapper for an iterator or array of data.
|
||||
It uses a stripped-down interface of `std::array` to present the iterator.
|
||||
|
||||
```c++
|
||||
using iterator = // The iterator type
|
||||
using value_type = // The iterator value-type
|
||||
using reference = // The iterator reference-type
|
||||
iterator begin() // The iterator
|
||||
```
|
||||
|
||||
In general, users do not need to construct `Engine`s on their own. When a `Tensor` is constructed,
|
||||
the appropriate engine -- often `ArrayEngine<T,N>`, `ViewEngine<Iter>`, or
|
||||
`ConstViewEngine<Iter>` -- will be constructed.
|
||||
|
||||
### Tagged Iterators
|
||||
|
||||
Any random-access iterator can be used to construct a `Tensor`, but
|
||||
users can also "tag" any iterator with a memory space --
|
||||
e.g., to indicate this iterator is accessing global memory or shared memory.
|
||||
This is done by calling `make_gmem_ptr(g)` or `make_gmem_ptr<T>(g)` to tag `g` as a global memory iterator,
|
||||
and `make_smem_ptr(s)` or `make_smem_ptr<T>(s)` to tag `s` as a shared memory iterator.
|
||||
|
||||
Tagging memory makes it possible for CuTe's `Tensor` algorithms
|
||||
to use the fastest implementation for the specific kind(s) of memory.
|
||||
When calling very specific operations with `Tensor`s, it also allows those
|
||||
operators to verify the tags against what is expected.
|
||||
For example, some kinds of optimized copy operations require
|
||||
the source of the copy to be global memory
|
||||
and the destination of the copy to be shared memory.
|
||||
Tagging makes it possible for CuTe to dispatch
|
||||
to those copy operations and/or verify against those copy operations.
|
||||
|
||||
## Tensor Creation
|
||||
|
||||
`Tensor`s can be constructed as owning or nonowning.
|
||||
|
||||
"Owning" `Tensor`s behave like `std::array`.
|
||||
When you copy the `Tensor`, you (deep-)copy its elements,
|
||||
and the `Tensor`'s destructor deallocates the array of elements.
|
||||
|
||||
"Nonowning" `Tensor`'s behave like a (raw) pointer.
|
||||
Copying the `Tensor` doesn't copy the elements,
|
||||
and destroying the `Tensor` doesn't deallocate the array of elements.
|
||||
|
||||
This has implications for developers of generic `Tensor` algorithms.
|
||||
For example, input `Tensor` parameters of a function
|
||||
should be passed by referece or const reference,
|
||||
because passing a `Tensor` by value
|
||||
may or may not make a deep copy of the `Tensor`'s elements.
|
||||
|
||||
### Nonowning Tensors
|
||||
|
||||
A `Tensor` is usually a nonowning view of existing memory.
|
||||
Nonowning `Tensor`s are created by calling `make_tensor`
|
||||
with two arguments: a random-access iterator, and the `Layout` or arguments to construct a `Layout`.
|
||||
|
||||
Here are some examples of creating `Tensor`s
|
||||
that are nonowning views of existing memory.
|
||||
|
||||
```cpp
|
||||
float* A = ...;
|
||||
|
||||
// Untagged pointers
|
||||
Tensor tensor_8 = make_tensor(A, make_layout(Int<8>{})); // Construct with Layout
|
||||
Tensor tensor_8s = make_tensor(A, Int<8>{}); // Construct with Shape
|
||||
Tensor tensor_8d2 = make_tensor(A, 8, 2); // Construct with Shape and Stride
|
||||
|
||||
// Global memory (static or dynamic layouts)
|
||||
Tensor gmem_8s = make_tensor(make_gmem_ptr(A), Int<8>{});
|
||||
Tensor gmem_8d = make_tensor(make_gmem_ptr(A), 8);
|
||||
Tensor gmem_8sx16d = make_tensor(make_gmem_ptr(A), make_shape(Int<8>{},16));
|
||||
Tensor gmem_8dx16s = make_tensor(make_gmem_ptr(A), make_shape ( 8 ,Int<16>{}),
|
||||
make_stride(Int<16>{},Int< 1>{}));
|
||||
|
||||
// Shared memory (static or dynamic layouts)
|
||||
Layout smem_layout = make_layout(make_shape(Int<4>{},Int<8>{}));
|
||||
__shared__ float smem[decltype(cosize(smem_layout))::value]; // (static-only allocation)
|
||||
Tensor smem_4x8_col = make_tensor(make_smem_ptr(smem), smem_layout);
|
||||
Tensor smem_4x8_row = make_tensor(make_smem_ptr(smem), shape(smem_layout), LayoutRight{});
|
||||
```
|
||||
|
||||
As shown, users wrap the pointer by identifying its memory space:
|
||||
e.g., global memory (via `make_gmem_ptr` or `make_gmem_ptr<T>`) or shared memory (via `make_smem_ptr` or `make_smem_ptr<T>`).
|
||||
`Tensor`s that view existing memory can have either static or dynamic `Layout`s.
|
||||
|
||||
Calling `print` on all of the above tensors displays
|
||||
```
|
||||
tensor_8 : ptr[32b](0x7f42efc00000) o _8:_1
|
||||
tensor_8s : ptr[32b](0x7f42efc00000) o _8:_1
|
||||
tensor_8d2 : ptr[32b](0x7f42efc00000) o 8:2
|
||||
gmem_8s : gmem_ptr[32b](0x7f42efc00000) o _8:_1
|
||||
gmem_8d : gmem_ptr[32b](0x7f42efc00000) o 8:_1
|
||||
gmem_8sx16d : gmem_ptr[32b](0x7f42efc00000) o (_8,16):(_1,_8)
|
||||
gmem_8dx16s : gmem_ptr[32b](0x7f42efc00000) o (8,_16):(_16,_1)
|
||||
smem_4x8_col : smem_ptr[32b](0x7f4316000000) o (_4,_8):(_1,_4)
|
||||
smem_4x8_row : smem_ptr[32b](0x7f4316000000) o (_4,_8):(_8,_1)
|
||||
```
|
||||
|
||||
which displays the pointer type along with any memory space tags, the pointer's `value_type` width, the raw pointer address, and the associated `Layout`.
|
||||
|
||||
### Owning Tensors
|
||||
|
||||
A `Tensor` can also be an owning array of memory.
|
||||
Owning `Tensor`s are created by calling `make_tensor<T>`,
|
||||
where `T` is the type of each element of the array, and
|
||||
a `Layout` or arguments to construct a `Layout`.
|
||||
The array is allocated analogously to `std::array<T,N>` and, therefore, owning `Tensor`s must be constructed with a `Layout` that has static shapes and static strides.
|
||||
CuTe does not perform dynamic memory allocation in `Tensor`s as it is not a common or performant operation within CUDA kernels.
|
||||
|
||||
Here are some examples of creating owning `Tensor`s.
|
||||
|
||||
```c++
|
||||
// Register memory (static layouts only)
|
||||
Tensor rmem_4x8_col = make_tensor<float>(Shape<_4,_8>{});
|
||||
Tensor rmem_4x8_row = make_tensor<float>(Shape<_4,_8>{},
|
||||
LayoutRight{});
|
||||
Tensor rmem_4x8_pad = make_tensor<float>(Shape <_4, _8>{},
|
||||
Stride<_32,_2>{});
|
||||
Tensor rmem_4x8_like = make_tensor_like(rmem_4x8_pad);
|
||||
```
|
||||
|
||||
The `make_tensor_like` function makes an owning Tensor of register memory with the same value type and shape as its input `Tensor` argument and attempts to use the same order of strides as well.
|
||||
|
||||
Calling `print` on each of the above tensors produces similar output
|
||||
|
||||
```
|
||||
rmem_4x8_col : ptr[32b](0x7fff48929460) o (_4,_8):(_1,_4)
|
||||
rmem_4x8_row : ptr[32b](0x7fff489294e0) o (_4,_8):(_8,_1)
|
||||
rmem_4x8_pad : ptr[32b](0x7fff489295e0) o (_4,_8):(_32,_2)
|
||||
rmem_4x8_like : ptr[32b](0x7fff48929560) o (_4,_8):(_8,_1)
|
||||
```
|
||||
|
||||
and we can see that each pointer address is unique indicating that each `Tensor` is a unique array-like allocation.
|
||||
|
||||
## Accessing a Tensor
|
||||
|
||||
Users can access the elements of a `Tensor` via `operator()` and `operator[]`,
|
||||
which take `IntTuple`s of logical coordinates.
|
||||
|
||||
When users access a `Tensor`,
|
||||
the `Tensor` uses its `Layout` to map the logical coordinate
|
||||
to an offset that can be accessed by the iterator.
|
||||
You can see this in `Tensor`'s implementation of `operator[]`.
|
||||
|
||||
```c++
|
||||
template <class Coord>
|
||||
decltype(auto) operator[](Coord const& coord) {
|
||||
return data()[layout()(coord)];
|
||||
}
|
||||
```
|
||||
|
||||
For example, we can read and write to `Tensor`s using natural coordinates, using the variadic `operator()`, or the container-like `operator[]`.
|
||||
|
||||
```c++
|
||||
Tensor A = make_tensor<float>(Shape <Shape < _4,_5>,Int<13>>{},
|
||||
Stride<Stride<_12,_1>, _64>{});
|
||||
float* b_ptr = ...;
|
||||
Tensor B = make_tensor(b_ptr, make_shape(13, 20));
|
||||
|
||||
// Fill A via natural coordinates op[]
|
||||
for (int m0 = 0; m0 < size<0,0>(A); ++m0)
|
||||
for (int m1 = 0; m1 < size<0,1>(A); ++m1)
|
||||
for (int n = 0; n < size<1>(A); ++n)
|
||||
A[make_coord(make_coord(m0,m1),n)] = n + 2 * m0;
|
||||
|
||||
// Transpose A into B using variadic op()
|
||||
for (int m = 0; m < size<0>(A); ++m)
|
||||
for (int n = 0; n < size<1>(A); ++n)
|
||||
B(n,m) = A(m,n);
|
||||
|
||||
// Copy B to A as if they are arrays
|
||||
for (int i = 0; i < A.size(); ++i)
|
||||
A[i] = B[i];
|
||||
```
|
||||
|
||||
## Tiling a Tensor
|
||||
|
||||
Many of the [`Layout` algebra operations](https://github.com/NVIDIA/cutlass/blob/main/media/docs/cute/02_layout_algebra.md) can also be applied to `Tensor`.
|
||||
```cpp
|
||||
composition(Tensor, Tiler)
|
||||
logical_divide(Tensor, Tiler)
|
||||
zipped_divide(Tensor, Tiler)
|
||||
tiled_divide(Tensor, Tiler)
|
||||
flat_divide(Tensor, Tiler)
|
||||
```
|
||||
The above operations allows arbitrary subtensors to be "factored out" of `Tensor`s. This very commonly used in tiling for threadgroups, tiling for MMAs, and reodering tiles of data for threads.
|
||||
|
||||
Note that the `_product` operations are not implemented for `Tensor`s as those would
|
||||
often produce layouts with increased codomain sizes, which means the `Tensor` would
|
||||
require accessing elements unpredictably far outside its previous bounds. `Layout`s can be
|
||||
used in products, but not `Tensor`s.
|
||||
|
||||
## Slicing a Tensor
|
||||
|
||||
Whereas accessing a `Tensor` with a coordinate will return an element of that tensor,
|
||||
slicing a `Tensor` will return a subtensor of all the elements in the sliced mode(s).
|
||||
|
||||
Slices are performed through the same `operator()`
|
||||
that are used for accessing an individual element.
|
||||
Passing in `_` (the underscore character, an instance of the `cute::Underscore` type)
|
||||
has the same effect as `:` (the colon character) in Fortran or Matlab:
|
||||
retain that mode of the tensor as if no coordinate had been used.
|
||||
|
||||
Slicing a tensor performs two operations,
|
||||
* the `Layout` is evaluated on the partial coordinate and the resulting offset is accumulated into the iterator -- the new iterator points to the start of the new tensor.
|
||||
* the `Layout` modes cooresponding to `_`-elements of the coordinate are used to construct a new layout.
|
||||
Together, the new iterator and the new layout construct the new tensor.
|
||||
|
||||
```cpp
|
||||
// ((_3,2),(2,_5,_2)):((4,1),(_2,13,100))
|
||||
Tensor A = make_tensor(ptr, make_shape (make_shape (Int<3>{},2), make_shape ( 2,Int<5>{},Int<2>{})),
|
||||
make_stride(make_stride( 4,1), make_stride(Int<2>{}, 13, 100)));
|
||||
|
||||
// ((2,_5,_2)):((_2,13,100))
|
||||
Tensor B = A(2,_);
|
||||
|
||||
// ((_3,_2)):((4,1))
|
||||
Tensor C = A(_,5);
|
||||
|
||||
// (_3,2):(4,1)
|
||||
Tensor D = A(make_coord(_,_),5);
|
||||
|
||||
// (_3,_5):(4,13)
|
||||
Tensor E = A(make_coord(_,1),make_coord(0,_,1));
|
||||
|
||||
// (2,2,_2):(1,_2,100)
|
||||
Tensor F = A(make_coord(2,_),make_coord(_,3,_));
|
||||
```
|
||||
|
||||
<p align="center">
|
||||
<img src="../../images/cute/slice.png" alt="slice.png" height="300"/>
|
||||
</p>
|
||||
|
||||
In the image above, a `Tensor` is sliced in various ways and the subtensors generated by those slices are highlighted within the original tensor. Note that tensor `C` and `D` contain the same elements, but have different ranks and shapes due to the use of `_` versus the use of `make_coord(_,_)`. In each case, the rank of the result is equal to the number of `Underscore`s in the slicing coordinate.
|
||||
|
||||
## Partitioning a Tensor
|
||||
|
||||
To implement generic partitioning of a `Tensor`, we apply composition or tiling followed by a slicing. This can be performed in many ways, but we have found three ways that are particularly useful: inner-partitioning, outer-partitioning, and TV-layout-partitioning.
|
||||
|
||||
### Inner and outer partitioning
|
||||
|
||||
Let's take a tiled example and look at how we can slice it in useful ways.
|
||||
|
||||
```cpp
|
||||
Tensor A = make_tensor(ptr, make_shape(8,24)); // (8,24)
|
||||
auto tiler = Shape<_4,_8>{}; // (_4,_8)
|
||||
|
||||
Tensor tiled_a = zipped_divide(A, tiler); // ((_4,_8),(2,3))
|
||||
```
|
||||
|
||||
Suppose that we want to give each threadgroup one of these 4x8 tiles of data. Then we can use our threadgroup coordinate to index into the second mode.
|
||||
|
||||
```cpp
|
||||
Tensor cta_a = tiled_a(make_coord(_,_), make_coord(blockIdx.x, blockIdx.y)); // (_4,_8)
|
||||
```
|
||||
|
||||
We call this an *inner-partition* because it keeps the inner "tile" mode. This pattern of applying a tiler and then slicing out that tile by indexing into the remainder mode is common and has been wrapped into its own function `inner_partition(Tensor, Tiler, Coord)`. You'll often see `local_tile(Tensor, Tiler, Coord)` which is just another name for `inner_partition`. The `local_tile` partitioner is very often applied at the threadgroup level to partition tensors into tiles across threadgroups.
|
||||
|
||||
Alternatively, suppose that we have 32 threads and want to give each thread one element of these 4x8 tiles of data. Then we can use our thread to index into the first mode.
|
||||
|
||||
```cpp
|
||||
Tensor thr_a = tiled_a(threadIdx.x, make_coord(_,_)); // (2,3)
|
||||
```
|
||||
|
||||
We call this an *outer-partition* because it keeps the outer "rest" mode. This pattern of applying a tiler and then slicing into that tile by indexing into the tile mode is common and has been wrapped into its own function `outer_partition(Tensor, Tiler, Coord)`. Sometimes you'll see `local_partition(Tensor, Layout, Idx)`, which is a rank-sensitive wrapper around `outer_partition` that transforms the `Idx` into a `Coord` using the inverse of the `Layout` and then constructs a `Tiler` with the same top-level shape of the `Layout`. This allows the user to ask for a row-major, column-major, or arbitrary layout of threads with a given shape that can be used to partition into a tensor.
|
||||
|
||||
To see how these partitioning patterns are used, see the [introductory GEMM tutorial](./0x_gemm_tutorial.md).
|
||||
|
||||
### Thread-Value partitioning
|
||||
|
||||
Another common partitioning strategy is called a thread-value partitioning. In this pattern, we construct a `Layout` that represents the mapping of all threads (or any parallel agent) and all values that each thread will receive to coordinates of the target data. With `composition` the target data layout is transformed according to our TV-layout and then we can simply slice into the thread-mode of the result with our thread index.
|
||||
|
||||
```cpp
|
||||
// Construct a TV-layout that maps 8 thread indices and 4 value indices
|
||||
// to 1D coordinates within a 4x8 tensor
|
||||
// (T8,V4) -> (M4,N8)
|
||||
auto tv_layout = Layout<Shape <Shape <_2,_4>,Shape <_2, _2>>,
|
||||
Stride<Stride<_8,_1>,Stride<_4,_16>>>{}; // (8,4)
|
||||
|
||||
// Construct a 4x8 tensor with any layout
|
||||
Tensor A = make_tensor<float>(Shape<_4,_8>{}, LayoutRight{}); // (4,8)
|
||||
// Compose A with the tv_layout to transform its shape and order
|
||||
Tensor tv = composition(A, tv_layout); // (8,4)
|
||||
// Slice so each thread has 4 values in the shape and order that the tv_layout prescribes
|
||||
Tensor v = tv(threadIdx.x, _); // (4)
|
||||
```
|
||||
|
||||
<p align="center">
|
||||
<img src="../../images/cute/tv_layout.png" alt="tv_layout.png" height="300"/>
|
||||
</p>
|
||||
|
||||
The above image is a visual representation of the above code. An arbitrary 4x8 layout of data is composed with a specific 8x4 TV-layout that represents a partitioning pattern. The result of the composition is on the right where each threads' values are arranged across each row. The bottom layout depicts the inverse TV layout which shows the mapping of 4x8 logical coordinates to the thread id and value id they will be mapped to.
|
||||
|
||||
To see how these partitioning patterns are constructed and used, see the [tutorial on building MMA Traits](./0t_mma_atom.md).
|
||||
|
||||
## Examples
|
||||
|
||||
### Copy a subtile from global memory to registers
|
||||
|
||||
The following example copies rows of a matrix (with any `Layout`)
|
||||
from global memory to register memory,
|
||||
then executes some algorithm `do_something`
|
||||
on the row that lives in register memory.
|
||||
|
||||
```c++
|
||||
Tensor gmem = make_tensor(ptr, make_shape(Int<8>{}, 16)); // (_8,16)
|
||||
Tensor rmem = make_tensor_like(gmem(_, 0)); // (_8)
|
||||
for (int j = 0; j < size<1>(gmem); ++j) {
|
||||
copy(gmem(_, j), rmem);
|
||||
do_something(rmem);
|
||||
}
|
||||
```
|
||||
|
||||
This code does not need to know anything about the `Layout` of `gmem`
|
||||
other than that it is rank-2 and that the first mode has a static size.
|
||||
The following code checks both of those conditions at compile time.
|
||||
|
||||
```c++
|
||||
CUTE_STATIC_ASSERT_V(rank(gmem) == Int<2>{});
|
||||
CUTE_STATIC_ASSERT_V(is_static<decltype(shape<0>(gmem))>{});
|
||||
```
|
||||
|
||||
Extending this example using the tiling utilities detailed in [the `Layout` algebra section](./02_layout_algebra.md), we can copy an arbitrary subtile of a tensor using almost the same code as above.
|
||||
|
||||
```c++
|
||||
Tensor gmem = make_tensor(ptr, make_shape(24, 16)); // (24,16)
|
||||
|
||||
auto tiler = Shape<_8,_4>{}; // 8x4 tiler
|
||||
//auto tiler = Tile<Layout<_8,_3>, Layout<_4,_2>>{}; // 8x4 tiler with stride-3 and stride-2
|
||||
Tensor gmem_tiled = zipped_divide(gmem, tiler); // ((_8,_4),Rest)
|
||||
Tensor rmem = make_tensor_like(gmem_tiled(_, 0)); // ((_8,_4))
|
||||
for (int j = 0; j < size<1>(gmem_tiled); ++j) {
|
||||
copy(gmem_tiled(_, j), rmem);
|
||||
do_something(rmem);
|
||||
}
|
||||
```
|
||||
|
||||
This applies a statically shaped `Tiler` to the global memory `Tensor`, creates an register `Tensor` that is compatible with the shape of that tile, then loops through each tile to copy it into memory and `do_something`.
|
||||
|
||||
## Summary
|
||||
|
||||
* `Tensor` is defined as an `Engine` and a `Layout`.
|
||||
|
||||
* `Engine` is an iterator that can be offset and dereferenced.
|
||||
* `Layout` defines the logical domain of the tensor and maps coordinates to offsets.
|
||||
|
||||
* Tile a `Tensor` using the same methods for tiling `Layout`s.
|
||||
|
||||
* Slice a `Tensor` to retrieve subtensors.
|
||||
|
||||
* Partitioning is tiling and/or composition followed by slicing.
|
||||
|
||||
## Copyright
|
||||
|
||||
Copyright (c) 2017 - 2025 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
```
|
||||
Redistribution and use in source and binary forms, with or without
|
||||
modification, are permitted provided that the following conditions are met:
|
||||
|
||||
1. Redistributions of source code must retain the above copyright notice, this
|
||||
list of conditions and the following disclaimer.
|
||||
|
||||
2. Redistributions in binary form must reproduce the above copyright notice,
|
||||
this list of conditions and the following disclaimer in the documentation
|
||||
and/or other materials provided with the distribution.
|
||||
|
||||
3. Neither the name of the copyright holder nor the names of its
|
||||
contributors may be used to endorse or promote products derived from
|
||||
this software without specific prior written permission.
|
||||
|
||||
THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS "AS IS"
|
||||
AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE
|
||||
IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE ARE
|
||||
DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT HOLDER OR CONTRIBUTORS BE LIABLE
|
||||
FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL
|
||||
DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR
|
||||
SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER
|
||||
CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY,
|
||||
OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE
|
||||
OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
|
||||
```
|
||||
255
media/docs/cpp/cute/04_algorithms.md
Normal file
255
media/docs/cpp/cute/04_algorithms.md
Normal file
@@ -0,0 +1,255 @@
|
||||
# CuTe Tensor algorithms
|
||||
|
||||
This section summarizes the interfaces and implementations
|
||||
of common numerical algorithms performed on `Tensor`s.
|
||||
|
||||
The implementation of these algorithms may be found in the
|
||||
[include/cute/algorithm/](https://github.com/NVIDIA/cutlass/tree/main/include/cute/algorithm/)
|
||||
directory.
|
||||
|
||||
## `copy`
|
||||
|
||||
CuTe's `copy` algorithm copies the elements of a source `Tensor`
|
||||
into the elements of a destination `Tensor`.
|
||||
The various overloads of `copy` can be found in
|
||||
[`include/cute/algorithm/copy.hpp`](https://github.com/NVIDIA/cutlass/tree/main/include/cute/algorithm/copy.hpp).
|
||||
|
||||
### Interface and specialization opportunities
|
||||
|
||||
A `Tensor` encapsulates the data type, data location,
|
||||
and possibly also the shape and stride of the tensor at compile time.
|
||||
As a result, `copy` can and does dispatch,
|
||||
based on the types of its arguments,
|
||||
to use any of various synchronous or asynchronous hardware copy instructions.
|
||||
|
||||
The `copy` algorithm has two main overloads.
|
||||
The first just takes the source `Tensor` and the destination `Tensor`.
|
||||
|
||||
```c++
|
||||
template <class SrcEngine, class SrcLayout,
|
||||
class DstEngine, class DstLayout>
|
||||
CUTE_HOST_DEVICE
|
||||
void
|
||||
copy(Tensor<SrcEngine, SrcLayout> const& src,
|
||||
Tensor<DstEngine, DstLayout> & dst);
|
||||
```
|
||||
|
||||
The second takes those two parameters, plus a `Copy_Atom`.
|
||||
|
||||
```c++
|
||||
template <class... CopyArgs,
|
||||
class SrcEngine, class SrcLayout,
|
||||
class DstEngine, class DstLayout>
|
||||
CUTE_HOST_DEVICE
|
||||
void
|
||||
copy(Copy_Atom<CopyArgs...> const& copy_atom,
|
||||
Tensor<SrcEngine, SrcLayout> const& src,
|
||||
Tensor<DstEngine, DstLayout> & dst);
|
||||
```
|
||||
|
||||
The two-parameter `copy` overload picks a default implementation
|
||||
based only on the types of the two `Tensor` parameters.
|
||||
The `Copy_Atom` overload lets callers override that default
|
||||
by specifying a nondefault `copy` implementation.
|
||||
|
||||
### Parallelism and synchronization depend on parameter types
|
||||
|
||||
Either the default implementation or
|
||||
the implementation selected by a `Copy_Atom` overload
|
||||
may use none or all available parallelism,
|
||||
and may have a variety of synchronization semantics.
|
||||
The behavior depends on `copy`'s parameter types.
|
||||
Users are expected to figure this out based on their knowledge
|
||||
of the architecture on which they are running.
|
||||
(Developers often write a custom optimized kernel
|
||||
for each GPU architecture.)
|
||||
|
||||
The `copy` algorithm may be sequential per thread,
|
||||
or it may be parallel across some collection of threads
|
||||
(e.g., a block or cluster).
|
||||
|
||||
If `copy` is parallel,
|
||||
then the collection of participating threads
|
||||
may need synchronization before any thread in the collection
|
||||
may assume that the copy operation has completed.
|
||||
For example, if the participating threads form a thread block,
|
||||
then users must invoke `__syncthreads()`
|
||||
or the Cooperative Groups equivalent
|
||||
before they may use the results of `copy`.
|
||||
|
||||
The `copy` algorithm may use asynchronous copy instructions,
|
||||
such as `cp.async`, or its C++ interface `memcpy_async`.
|
||||
In that case, users will need to perform
|
||||
the additional synchronization appropriate to that underlying implementation
|
||||
before they may use the results of the `copy` algorithm.
|
||||
[The CuTe GEMM tutorial example](https://github.com/NVIDIA/cutlass/tree/main/examples/cute/tutorial/)
|
||||
shows one such synchronization method.
|
||||
More optimized GEMM implementations use pipelining techniques
|
||||
to overlap asynchronous `copy` operations with other useful work.
|
||||
|
||||
### A generic copy implementation
|
||||
|
||||
A simple example of a generic `copy` implementation
|
||||
for any two `Tensor`s looks like this.
|
||||
|
||||
```c++
|
||||
template <class TA, class ALayout,
|
||||
class TB, class BLayout>
|
||||
CUTE_HOST_DEVICE
|
||||
void
|
||||
copy(Tensor<TA, ALayout> const& src, // Any logical shape
|
||||
Tensor<TB, BLayout> & dst) // Any logical shape
|
||||
{
|
||||
for (int i = 0; i < size(dst); ++i) {
|
||||
dst(i) = src(i);
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
This generic `copy` algorithm addresses both `Tensor`s
|
||||
with 1-D logical coordinates, thus traversing both `Tensor`s
|
||||
in a logical column-major order.
|
||||
Some reasonable architecture-independent optimizations
|
||||
would include the following.
|
||||
|
||||
1. If the two `Tensor`s have known memory spaces with optimized
|
||||
access instructions (like `cp.async`), then dispatch to the
|
||||
custom instruction.
|
||||
|
||||
2. The two `Tensor`s have static layouts and it can be proven
|
||||
that element vectorization is valid -- for example, four `ld.global.b32`s
|
||||
can be combined into a single `ld.global.b128` -- then vectorize the source
|
||||
and destinations tensors.
|
||||
|
||||
3. If possible, validate that the copy instruction to be used is
|
||||
appropriate for the source and destination tensors.
|
||||
|
||||
CuTe's optimized copy implementations can do all of these.
|
||||
|
||||
## `copy_if`
|
||||
|
||||
CuTe's `copy_if` algorithm lives in the same header as `copy`,
|
||||
[`include/cute/algorithm/copy.hpp`](https://github.com/NVIDIA/cutlass/tree/main/include/cute/algorithm/copy.hpp).
|
||||
The algorithm takes source and destination `Tensor` parameters like `copy`,
|
||||
but it also takes a "predication `Tensor`"
|
||||
with the same shape as the input and output.
|
||||
Elements of the source `Tensor` are only copied
|
||||
if the corresponding predication `Tensor` element is nonzero.
|
||||
|
||||
For details on why and how to use `copy_if`,
|
||||
please refer to the
|
||||
["predication" section of the tutorial](./0y_predication.md).
|
||||
|
||||
## `gemm`
|
||||
|
||||
### What `gemm` computes
|
||||
|
||||
The `gemm` algorithm takes three `Tensor`s, A, B, and C.
|
||||
What it does depends on the number of modes
|
||||
that its `Tensor` parameters have.
|
||||
We express these modes using letters.
|
||||
|
||||
* V indicates a "vector," a mode of independent elements.
|
||||
|
||||
* M and N indicate the number of rows resp. columns
|
||||
of the matrix result C of the BLAS's GEMM routine.
|
||||
|
||||
* K indicates the "reduction mode" of GEMM,
|
||||
that is, the mode along which GEMM sums.
|
||||
Please see the [GEMM tutorial](./0x_gemm_tutorial.md) for details.
|
||||
|
||||
We list the modes of the input `Tensor`s A and B,
|
||||
and the output `Tensor` C,
|
||||
using a notation `(...) x (...) => (...)`.
|
||||
The two leftmost `(...)` describe A and B (in that order),
|
||||
and the `(...)` to the right of the `=>` describes C.
|
||||
|
||||
1. `(V) x (V) => (V)`. The element-wise product of vectors: C<sub>v</sub> += A<sub>v</sub> B<sub>v</sub>. Dispatches to FMA or MMA.
|
||||
|
||||
2. `(M) x (N) => (M,N)`. The outer product of vectors: C<sub>mn</sub> += A<sub>m</sub> B<sub>n</sub>. Dispatches to (4) with V=1.
|
||||
|
||||
3. `(M,K) x (N,K) => (M,N)`. The product of matrices: C<sub>mn</sub> += A<sub>mk</sub> B<sub>nk</sub>. Dispatches to (2) for each K.
|
||||
|
||||
4. `(V,M) x (V,N) => (V,M,N)`. The batched outer product of vectors: C<sub>vmn</sub> += A<sub>vm</sub> B<sub>vn</sub>. Optimizes for register reuse and dispatches to (1) for each M, N.
|
||||
|
||||
5. `(V,M,K) x (V,N,K) => (V,M,N)`. The batched product of matrices: C<sub>vmn</sub> += A<sub>vmk</sub> B<sub>vnk</sub>. Dispatches to (4) for each K.
|
||||
|
||||
Please refer to the [GEMM tutorial](./0x_gemm_tutorial.md)
|
||||
for an overview of CuTe's convention for ordering the modes.
|
||||
For example, if K appears, it always appears rightmost ("outermost").
|
||||
If V appears, it always appears leftmost ("innermost").
|
||||
|
||||
### Dispatch to optimized implementations
|
||||
|
||||
Just like with `copy`, CuTe's implementations of `gemm`
|
||||
uses its `Tensor` arguments' types to dispatch
|
||||
to an appropriately optimized implementation.
|
||||
Also like `copy`, `gemm` takes an optional `MMA_Atom` parameter
|
||||
that lets callers override the default `FMA` instruction
|
||||
that CuTe would select based on the `Tensor` arguments' types.
|
||||
|
||||
For more information on `MMA_Atom` and on specialization of `gemm`
|
||||
for different architectures, please refer to the
|
||||
[MMA section of the tutorial](./0t_mma_atom.md).
|
||||
|
||||
## `axpby`
|
||||
|
||||
The `axpby` algorithm lives in the header file
|
||||
[`include/cute/algorithm/axpby.hpp`](https://github.com/NVIDIA/cutlass/tree/main/include/cute/algorithm/axpby.hpp).
|
||||
It assigns to $y$ the result of $\alpha x + \beta y$,
|
||||
where $\alpha$ and $\beta$ are scalars and $x$ and $y$ are `Tensor`s.
|
||||
The name stands for "Alpha times X Plus Beta times Y,"
|
||||
and is a generalization of the original BLAS "AXPY" routine
|
||||
("Alpha times X Plus Y").
|
||||
|
||||
## `fill`
|
||||
|
||||
The `fill` algorithm lives in the header file
|
||||
[`include/cute/algorithm/fill.hpp`](https://github.com/NVIDIA/cutlass/tree/main/include/cute/algorithm/fill.hpp).
|
||||
It overwrites the elements of its `Tensor` output argument
|
||||
with a given scalar value.
|
||||
|
||||
## `clear`
|
||||
|
||||
The `clear` algorithm lives in the header file
|
||||
[`include/cute/algorithm/clear.hpp`](https://github.com/NVIDIA/cutlass/tree/main/include/cute/algorithm/clear.hpp).
|
||||
It overwrites the elements of its `Tensor` output argument with zeros.
|
||||
|
||||
## Other algorithms
|
||||
|
||||
CuTe provides other algorithms.
|
||||
Their header files can be found in the
|
||||
[`include/cute/algorithm`](https://github.com/NVIDIA/cutlass/tree/main/include/cute/algorithm)
|
||||
directory.
|
||||
|
||||
## Copyright
|
||||
|
||||
Copyright (c) 2017 - 2025 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
```
|
||||
Redistribution and use in source and binary forms, with or without
|
||||
modification, are permitted provided that the following conditions are met:
|
||||
|
||||
1. Redistributions of source code must retain the above copyright notice, this
|
||||
list of conditions and the following disclaimer.
|
||||
|
||||
2. Redistributions in binary form must reproduce the above copyright notice,
|
||||
this list of conditions and the following disclaimer in the documentation
|
||||
and/or other materials provided with the distribution.
|
||||
|
||||
3. Neither the name of the copyright holder nor the names of its
|
||||
contributors may be used to endorse or promote products derived from
|
||||
this software without specific prior written permission.
|
||||
|
||||
THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS "AS IS"
|
||||
AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE
|
||||
IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE ARE
|
||||
DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT HOLDER OR CONTRIBUTORS BE LIABLE
|
||||
FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL
|
||||
DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR
|
||||
SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER
|
||||
CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY,
|
||||
OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE
|
||||
OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
|
||||
```
|
||||
559
media/docs/cpp/cute/0t_mma_atom.md
Normal file
559
media/docs/cpp/cute/0t_mma_atom.md
Normal file
@@ -0,0 +1,559 @@
|
||||
# CuTe's support for Matrix Multiply-Accumulate instructions
|
||||
|
||||
In this file, we explain in detail how we support our GPUs'
|
||||
Matrix Multiply-Accumulate (MMA) hardware instructions in CuTe.
|
||||
|
||||
MMAs are architecture-specific.
|
||||
Different generations of GPU architectures
|
||||
introduce different sets of MMA instructions.
|
||||
However, CuTe features such as `Layout`
|
||||
makes it possible to expose MMAs for use in generic CUDA C++ code.
|
||||
We accomplish this in multiple steps.
|
||||
|
||||
1. We wrap each MMA's PTX instruction in an "Operation" struct.
|
||||
|
||||
2. For each Operation struct, we define a "Traits" struct
|
||||
that defines all of the meta-information needed to use the Operation.
|
||||
|
||||
3. Combining the above, an "Atom" is the combination of the PTX Operation struct with the
|
||||
meta-information Traits struct and provides methods to construct
|
||||
`cute::Tensor` "fragments" for that Operation and to use that Operation
|
||||
on existing `cute::Tensor`s.
|
||||
|
||||
4. Combining potentially multiple Atoms, a "TiledMMA" provides utilities for building
|
||||
more complex partitioning patterns by creating layouts and interleavings of Atoms.
|
||||
|
||||
## CuTe MMA Atoms
|
||||
|
||||
CuTe exposes each MMA to generic CUDA C++ code as a pair of structs:
|
||||
an "Operation" struct,
|
||||
and an `MMA_Traits` struct templated on the Operation struct type.
|
||||
|
||||
An "Operation" struct exposes the PTX instruction
|
||||
for that specific operation.
|
||||
It defines the arguments and interface it expects.
|
||||
Operation structs have minimal software dependencies --
|
||||
they do not use layouts, tensors, or non-standard numeric data types -- and
|
||||
describe only the physical inputs and outputs to the instruction.
|
||||
Different structs have different names
|
||||
that describe what the MMA instruction does.
|
||||
We will explain the naming scheme below.
|
||||
|
||||
A corresponding `MMA_Traits` struct specialization
|
||||
defines meta-information about the Operation,
|
||||
such as the logical compute types, the logical shape of the operation,
|
||||
and the `Layout`s of threads and values within the operation.
|
||||
The `MMA_Traits` struct takes the Operation as a template parameter.
|
||||
CuTe specializes `MMA_Traits` for each Operation type that it supports.
|
||||
|
||||
Together, these two types comprise an "Atom" that decouples the complexity of thread and data layouts from the call site of the PTX instruction. The Atom's Traits struct exposes information that is relevant to a single MMA operation, no matter the granularity at which it operates.
|
||||
|
||||
CuTe MMA atoms expose the semantics of a single MMA operation.
|
||||
This is true regardless of the hardware level at which the MMA operates.
|
||||
CuTe supports MMA atoms that operate at a variety of hardware levels,
|
||||
including
|
||||
|
||||
* a single thread (e.g., fused multiply-add (FMA) instruction);
|
||||
|
||||
* a quadpair (Volta);
|
||||
|
||||
* a single warp (Ampere); and
|
||||
|
||||
* a warpgroup (Hopper).
|
||||
|
||||
### Operation structs
|
||||
|
||||
#### Location of files
|
||||
|
||||
CuTe provides its Operations structs in the
|
||||
[`include/cute/arch`](https://github.com/NVIDIA/cutlass/tree/main/include/cute/arch)
|
||||
directory, in header files starting with `mma`.
|
||||
|
||||
#### Operation struct's name
|
||||
|
||||
A CuTe Operation struct's name principally encodes the PTX instruction it wraps.
|
||||
These often include
|
||||
|
||||
* its first supported architecture,
|
||||
|
||||
* the M, N, and K dimensions that it accepts,
|
||||
|
||||
* the types that it takes, and
|
||||
|
||||
* the arrangement of the A and B inputs.
|
||||
|
||||
For example, the Volta section below will refer to the
|
||||
`SM70_8x8x4_F32F16F16F32_NT` Operation struct defined in
|
||||
[`include/cute/arch/mma_sm70.hpp`](https://github.com/NVIDIA/cutlass/tree/main/include/cute/arch/mma_sm70.hpp).
|
||||
|
||||
* "SM70" refers to Volta.
|
||||
|
||||
* "8x8x4" refers to M = 8, N = 8, and K = 4,
|
||||
the dimensions of the MMA operation that the quadpair performs
|
||||
(see below). This is reflected in the PTX as `.m8n8k4.`.
|
||||
|
||||
* "F32F16F16F32" refers to the element types
|
||||
of the four matrix operands A, B, C, and D.
|
||||
An MMA computes D = C + A * B,
|
||||
so we read the types from left to right:
|
||||
D is F32 (`float`), A is F16 (half),
|
||||
B is F16 (half), and C is F32 (`float`). This is reflected in the PTX instruction name as `.f32.f16.f16.f32`.
|
||||
|
||||
* "NT" means that the PTX instruction is designed for inputs A as M-major (not transposed, column-major)
|
||||
and inputs B as N-major (transposed, row-major). This is reflected in the PTX instruction name as `.col.row.`.
|
||||
|
||||
#### Contents
|
||||
|
||||
An Operation struct has the following members.
|
||||
|
||||
##### Type aliases
|
||||
|
||||
An Operation struct has four public type aliases:
|
||||
`DRegisters`, `ARegisters`, `BRegisters`, and `CRegisters`.
|
||||
For example, the `SM70_8x8x4_F32F16F16F32_NT` Operation struct defined in
|
||||
[`include/cute/arch/mma_sm70.hpp`](https://github.com/NVIDIA/cutlass/tree/main/include/cute/arch/mma_sm70.hpp)
|
||||
defines these as follows.
|
||||
|
||||
```c++
|
||||
using DRegisters = float[8];
|
||||
using ARegisters = uint32_t[2];
|
||||
using BRegisters = uint32_t[2];
|
||||
using CRegisters = float[8];
|
||||
```
|
||||
|
||||
This shows how many values each thread will pass into the PTX instruction
|
||||
for each of the matrices A, B, C, and D. For this Operation,
|
||||
each thread passes 8 F32 values each for C and D (hence `float[8]`),
|
||||
and 4 F16 values each for A and B (hence `uint32_t[2]`;
|
||||
the instruction packs two 16-bit F16 values
|
||||
in each of the two 32-bit `uint32_t` values).
|
||||
|
||||
##### `fma` static member device function
|
||||
|
||||
An operation struct defines a public `static void fma` function.
|
||||
It is marked with the `CUTE_HOST_DEVICE` macro,
|
||||
which adds the `__host__ __device__` annotations.
|
||||
Different Operations define `fma` to take different numbers of arguments,
|
||||
depending on the PTX MMA instruction.
|
||||
The implementation protects use of the PTX instruction with a macro,
|
||||
and raises an `assert` if `fma` is called when the macro is not defined.
|
||||
This ensures that tests and examples that use this Operation in an Atom
|
||||
can still compile, even if the PTX instruction is not available.
|
||||
|
||||
### Traits
|
||||
|
||||
#### Location of files
|
||||
|
||||
CuTe provides its Traits structs in the
|
||||
[`include/cute/atom`](https://github.com/NVIDIA/cutlass/tree/main/include/cute/atom)
|
||||
directory, in header files starting with `mma_traits`.
|
||||
|
||||
#### Contents
|
||||
|
||||
An `MMA_Traits` specialization defines the following public type aliases.
|
||||
|
||||
* `ValTypeD`: Logical compute type of the D matrix
|
||||
|
||||
* `ValTypeA`: Logical compute type of the A matrix
|
||||
|
||||
* `ValTypeB`: Logical compute type of the B matrix
|
||||
|
||||
* `ValTypeC`: Logical compute type of the C matrix
|
||||
|
||||
* `Shape_MNK`: Logical MxNxK shape of the MMA operation
|
||||
|
||||
* `ThrID`: Logical thread mapping within the single MMA operation
|
||||
(specifying the thread, quadpair, warp, or warpgroup view)
|
||||
|
||||
* `ALayout`: Mapping of (thread,value) pairs to coordinates in the MxK A matrix
|
||||
|
||||
* `BLayout`: Mapping of (thread,value) pairs to coordinates in the NxK B matrix
|
||||
|
||||
* `CLayout`: Mapping of (thread,value) pairs to coordinates in the MxN C matrix
|
||||
|
||||
#### Example
|
||||
|
||||
The specialization of MMA_Traits for the
|
||||
`SM70_8x8x4_F32F16F16F32_NT` Operation lives in the header file
|
||||
[`include/cute/atom/mma_traits_sm70.hpp`](https://github.com/NVIDIA/cutlass/tree/main/include/cute/atom/mma_traits_sm70.hpp).
|
||||
It looks like this.
|
||||
|
||||
```c++
|
||||
template <>
|
||||
struct MMA_Traits<SM70_8x8x4_F32F16F16F32_NT>
|
||||
{
|
||||
using ValTypeD = float;
|
||||
using ValTypeA = half_t;
|
||||
using ValTypeB = half_t;
|
||||
using ValTypeC = float;
|
||||
|
||||
using Shape_MNK = Shape<_8,_8,_4>;
|
||||
using ThrID = SM70_QuadPair;
|
||||
using ALayout = SM70_8x4_Col;
|
||||
using BLayout = SM70_8x4_Col;
|
||||
using CLayout = SM70_8x8_32b;
|
||||
};
|
||||
```
|
||||
|
||||
The next section will explain these type aliases in detail.
|
||||
|
||||
## Volta
|
||||
|
||||
This and the following sections show examples of how to construct MMA atoms.
|
||||
We don't try to explain this for all GPU architectures and MMAs.
|
||||
Instead, we use selected examples to illustrate the process
|
||||
of developing new atoms.
|
||||
|
||||
Volta architecture implements an HMMA instruction where a group of 8 threads called a quadpair (QP) collaborate to share data and perform an 8x8x4 (fp32 or fp16) matrix multiply-accumulate. (since a warp is 32 threads wide, it would perform an MMA across 4 QPs for a tile size of 16x16x4).
|
||||
|
||||
We first take a look at how we would take the ISA semantics of thread and data partitioning for the HMMA instruction, and encode it in a Traits struct. The HMMA NT instruction has the thread-data layout:
|
||||
|
||||
<p align="center">
|
||||
<img src="../../images/cute/HMMA.8x8x4.NT.png" alt="HMMA.8x8x4.NT.png" height="400"/>
|
||||
</p>
|
||||
|
||||
### Types
|
||||
|
||||
The HMMA NT above uses types:
|
||||
|
||||
```cpp
|
||||
using ValTypeD = float;
|
||||
using ValTypeA = half_t;
|
||||
using ValTypeB = half_t;
|
||||
using ValTypeC = float;
|
||||
```
|
||||
|
||||
The rest of the `MMA_Traits` will be described in units of these types.
|
||||
|
||||
### Shape
|
||||
|
||||
The HMMA NT above has shape 8x8x4:
|
||||
|
||||
```cpp
|
||||
// Logical shape of the MMA
|
||||
using Shape_MNK = Shape <_8,_8,_4>;
|
||||
```
|
||||
|
||||
### Thread ID
|
||||
|
||||
If the 32 threads in a warp are logically indexed by [0 ... 31], then the above image contains threads [0,1,2,3]U[16,17,18,19]. These threads make up the 0th quadpair. We can write a thread mapping that maps eight logical thread ids [0,1,2,3,4,5,6,7] of the MMA to a quadpair thread index [0,1,2,3]U[16,17,18,19] of a warp. The layout function has 4 elements with a stride of 1 and 2 of those with a stride of 16. With this, we write a layout that represents a quadpair:
|
||||
|
||||
```cpp
|
||||
// Mapping from (logical thread id) -> (thread idx)
|
||||
using ThrID = Layout<Shape <_4, _2>,
|
||||
Stride<_1,_16>>;
|
||||
```
|
||||
|
||||
Again, this layout function maps the logical thread id [0,8) of the MMA operation onto the quadpair thread index [0,4)U[16,20) of a warp.
|
||||
|
||||
### Accumulator Mapping
|
||||
|
||||
Let us look at exactly how the 8 threads within a QP are mapped to the A, B and C matrices. For the C and D matrices, the above image is broken down a bit more below. On the left is shown the whole QP level view, and on the right is shown the values owned by just thread 0.
|
||||
|
||||
<p align="center">
|
||||
<img src="../../images/cute/HMMA.8x8x4.quadpair.C.png" alt="HMMA.8x8x4.quadpair.C.png" height="400"/>
|
||||
</p>
|
||||
|
||||
The metainformation of this single instruction level view is what we want to encode in CuTe. Specifically, the QP level view in this diagram corresponds to the four MMA traits for [SM70_F32F16F16F32](https://github.com/NVIDIA/cutlass/tree/main/include/cute/arch/mma_sm70.hpp). These structs contain the `Element` types, the `Shape_MNK`, and the `ThrID` mapping we constructed above. Now, let us take a look at the definition of `CLayout`, the thread-data layout of accumulators. The job of `CLayout` is to construct a mapping between the `(logical_thr_id, logical_val_id)` and `(m, n)` coordinate in the C matrix which can then be used to build up more complicated layouts and operations like the 16x16x4 WMMA.
|
||||
|
||||
We can start constructing a `CLayout` from the picture above. As with any CuTe layout, it is a pair of `Shape` and corresponding `Stride`. Let us just look at the shape for now. We know that the HMMA uses 8 threads each of which own 8 values. Therefore, the shape of our mapping must have a size of 8 along two modes. With this, we have
|
||||
|
||||
```cpp
|
||||
// (T8,V8) -> (m,n)
|
||||
using CLayout = Layout<Shape <_8, _8>,
|
||||
Stride<_?, _?>; // Stride to be filled in below
|
||||
```
|
||||
|
||||
This is not to be confused with the logical 8x8 shape of the C matrix. This is 8-threads by 8-values. We now want to map those to (m,n) coordinates. Since CuTe layouts return indices rather than coordinates, we choose a column-major encoding of the (m,n) coordinates:
|
||||
|
||||
```
|
||||
(logical_thr_id, logical_val_id) -> (m, n) == m + n * M
|
||||
```
|
||||
|
||||
With this in place, we can start thinking about how to construct the strides in `CLayout`. Let's begin by looking at the strides between threads. Note that
|
||||
* `(T0,V0)` is located at `(m,n) = (0,0) = 0`
|
||||
* `(T1,V0)` is located at `(m,n) = (1,0) = 1`
|
||||
* `(T2,V0)` is located at `(m,n) = (0,2) = 16`
|
||||
* `(T3,V0)` is located at `(m,n) = (1,2) = 17`
|
||||
* `(T4,V0)` is located at `(m,n) = (4,0) = 4`
|
||||
* `(T5,V0)` is located at `(m,n) = (5,0) = 5`
|
||||
* `(T6,V0)` is located at `(m,n) = (4,2) = 20`
|
||||
* `(T7,V0)` is located at `(m,n) = (5,2) = 21`
|
||||
|
||||
where `T4`,`T5`,`T6`,`T7` are the 4th,5th,6th,7th logical thread id of the MMA corresponding to thread indices of 16,17,18,19 of the warp (recorded in the `ThrID` mapping!).
|
||||
|
||||
We note that the pattern can be transcribed to a layout. We can find the position of the 8 threads via
|
||||
|
||||
```cpp
|
||||
using CLayout = Layout<Shape <Shape <_2, _2, _2>, _8>,
|
||||
Stride<Stride<_1, _16, _4>, _?>;
|
||||
```
|
||||
|
||||
With the exact same approach, we can construct the stride along the `logical value id` mode.
|
||||
* `(T0,V0)` is located at `(m,n) = (0,0) = 0`
|
||||
* `(T0,V1)` is located at `(m,n) = (0,1) = 8`
|
||||
* `(T0,V2)` is located at `(m,n) = (2,0) = 2`
|
||||
* `(T0,V3)` is located at `(m,n) = (2,1) = 10`
|
||||
* `(T0,V4)` is located at `(m,n) = (0,4) = 32`
|
||||
* `(T0,V5)` is located at `(m,n) = (0,5) = 40`
|
||||
* `(T0,V6)` is located at `(m,n) = (2,4) = 34`
|
||||
* `(T0,V7)` is located at `(m,n) = (2,5) = 42`
|
||||
|
||||
We note that this pattern can also be transcribed to a layout. We can find the position of the 8 values via
|
||||
|
||||
```cpp
|
||||
// (T8,V8) -> (m,n)
|
||||
using CLayout = Layout<Shape <Shape <_2, _2,_2>, Shape <_2,_2, _2>>,
|
||||
Stride<Stride<_1,_16,_4>, Stride<_8,_2,_32>>>;
|
||||
```
|
||||
|
||||
And that's all! We can verify that each `(tid,vid)` coordinate in this layout is reliably mapped to the correct (encoded) `(m,n)` coordinate.
|
||||
|
||||
In the case of F16 accumulators, the layout is way less complex. Each row of accumulators `(m, :)` is held by a single thread, which makes the layout:
|
||||
|
||||
```cpp
|
||||
using CLayout = Layout<Shape <_8,_8>,
|
||||
Stride<_1,_8>>;
|
||||
```
|
||||
|
||||
### A and B Layout Mapping
|
||||
|
||||
A and B matrix layouts depend on whether the sources are transposed or not. The diagram below shows the thread ID to data ownership map for A and B matrices in the case of NT and TN transposes.
|
||||
|
||||
<p align="center">
|
||||
<img src="../../images/cute/HMMA.8x8x4.quadpair.AB.png" alt="HMMA.8x8x4.quadpair.AB.png" height="400"/>
|
||||
</p>
|
||||
|
||||
Let's look at the TN layout for A matrix first (right side in the diagram). Again, there are the same 8 logical threads, but each threads owns only 4 elements this time. The shape of `ALayout` will then be `Shape<_8, _4>`. As for the strides, we again need a similar mapping between `(m, k) == m + k * M`. Looking down the `M` mode, we go from `(T0, V0)` to `(T1, V0)` which is a stride of 1 for all 8 threads. For the `K` mode, as we go across, we go from `(T0, V0)` to `(T0, V1)`, which makes a stride of 8 for all 4 values. Therefore, the A layout is:
|
||||
|
||||
```cpp
|
||||
// (T8,V4) -> (m,k)
|
||||
using ALayout = Layout<Shape <_8,_4>,
|
||||
Stride<_1,_8>>;
|
||||
```
|
||||
|
||||
Source B layout is constructed similarly for the TN HMMA, except that we want write it as `(N,K)` rather than `(K,N)` for convenience. For the strides, as we go across the `N` mode, we go from `(T0, V0)` to `(T1, V0)`, making this a stride of 1 for all 8 threads. As we go down the `K` mode, `(T0, V0)` to `(T0, V1)` which is a stride of 8 for all 4 values. So the B layout is the same as A:
|
||||
|
||||
```cpp
|
||||
// (T8,V4) -> (n,k)
|
||||
using BLayout = Layout<Shape <_8,_4>,
|
||||
Stride<_1,_8>>;
|
||||
```
|
||||
|
||||
The layouts in the case of NT are a bit more complicated (left side of the diagram). Going down the `M` mode of `A`, we see the four values of `T0` first and then we see the four values of `T4`. This means we first have a stride of 1 for 4 values, followed by a stride of 4 from `T0` to `T4`. So we have two sub-strides along the `M` mode. For the `K` mode, as we go across, we simply increment the `thr_id`, keeping `val_id` the same, making the stride 8 for 4 threads. This makes the A layout:
|
||||
|
||||
```cpp
|
||||
// (T8,V4) -> (m,k)
|
||||
using ALayout = Layout<Shape <Shape <_4,_2>,_4>,
|
||||
Stride<Stride<_8,_4>,_1>>;
|
||||
```
|
||||
|
||||
With the `(N,K)` ordering for B, the layout is the same.
|
||||
|
||||
```cpp
|
||||
// (T8,V4) -> (n,k)
|
||||
using BLayout = Layout<Shape <Shape <_4,_2>,_4>,
|
||||
Stride<Stride<_8,_4>,_1>>;
|
||||
```
|
||||
|
||||
For the NN and TT transposes, they are simply combinations of the two layouts we have seen for A and B so far.
|
||||
|
||||
## Hopper
|
||||
|
||||
Now, we are ready to take a look at the much larger GMMA operation (Group MMA) first introduced with Hopper architecture. These MMA instructions operate at the granularity of 128 threads (4 warps), which are collectively referred to as a warpgroup.
|
||||
|
||||
### Thread ID
|
||||
|
||||
In the case of Hopper GMMAs, the thread IDs are assigned based on the simple 1D contiguous layout, which makes `thrID` trivial:
|
||||
|
||||
```cpp
|
||||
using ThrID = Layout<_128, _1>;
|
||||
```
|
||||
|
||||
### Accumulator Mapping
|
||||
|
||||
Accumulators are mapped hierarchically in GMMA, starting from the concept of a core matrix and building up to a layout for the whole C matrix tile. Let's look at this core matrix first. We only consider fp16 accumulators here, but extensions of fp32 accumulators as trivial as we will see later.
|
||||
|
||||
Each core matrix has the layout as shown in the diagram below.
|
||||
<p align="center">
|
||||
<img src="../../images/cute/gmma_coremat_cd_fp16.png" alt="gmma_coremat_cd_fp16.png" height="600"/>
|
||||
</p>
|
||||
|
||||
As in the Volta examples, the thread IDs are logical only, and which of the four warps they belong to in the warpgroup is not important.
|
||||
|
||||
Then GMMA tiles this core matrix first vertically along the M mode, and then repeats that column of core matrices along the N mode to construct the full MxN tile. This tiling is shown in the image below.
|
||||
|
||||
<p align="center">
|
||||
<img src="../../images/cute/gmma_wg_n_slice.png" alt="gmma_wg_n_slice.png" height="600"/>
|
||||
</p>
|
||||
|
||||
With this image, we are again ready to start building the `CLayout` for `SM90_64x128x16_F16F16F16F16_TN` atom. Same as before, we are constructing a mapping between the `(logical_thr_id, logical_val_id) -> (m, n)` coordinate spaces.
|
||||
|
||||
To begin, let's follow the first few threads and values. We immediately see that they are arranged along the `N`-mode with pairs of values and four threads. This gives us
|
||||
|
||||
```cpp
|
||||
// (T128,V4) -> (M64,N8)
|
||||
using CLayout = Layout<Shape <Shape < _4, ...>, Shape < _2, ...>>,
|
||||
Stride<Stride<_128, ...>, Stride<_64, ...>>>;
|
||||
```
|
||||
|
||||
To complete the first 8x8 core matrix, the four threads repeat eight times down the `M`-mode:
|
||||
|
||||
```cpp
|
||||
// (T128,V4) -> (M64,N8)
|
||||
using CLayout = Layout<Shape <Shape < _4, _8, ...>, Shape < _2, ...>>,
|
||||
Stride<Stride<_128, _1, ...>, Stride<_64, ...>>>;
|
||||
```
|
||||
|
||||
Then, as we go to the next core matrix, we wrap back again to `T0`, but this time to `(T0, V2)`.
|
||||
|
||||
```cpp
|
||||
// (T128,V4) -> (M64,N8)
|
||||
using CLayout = Layout<Shape <Shape < _4, _8, ...>, Shape < _2, _2>>,
|
||||
Stride<Stride<_128, _1, ...>, Stride<_64, _8>>>;
|
||||
```
|
||||
|
||||
Finally, we get this entire pattern repeating four times, once for each warp, down the `M`-mode starting at `(m,n) = (16,0) = 16`. where two core matrices that belong to the same warp are stacked on top of each other. This makes the size of the final sub-mode of M 4. As for the stride, this time we go to `(T32, V0)`, which makes it a stride of 32.
|
||||
|
||||
```cpp
|
||||
// (T128,V4) -> (M64,N8)
|
||||
using CLayout = Layout<Shape <Shape < _4, _8, _4>, Shape < _2, _2>>,
|
||||
Stride<Stride<_128, _1, _16>, Stride<_64, _8>>>;
|
||||
```
|
||||
|
||||
This is the full `CLayout` for 64x8 accumulators. The GMMA instructions include 64xN variants with `N = [16,32,64,128,256]` where this 64x8 pattern is repeated giving each thread additional values. As this starts at `(m,n) = (0,8) = 512`, this is easy to account for in our `CLayout`. For example, the 64x128 `CLayout` is
|
||||
|
||||
```cpp
|
||||
// (T128,V64) -> (M64,N128)
|
||||
using CLayout = Layout<Shape <Shape < _4, _8, _4>, Shape < _2, _2, _16>>,
|
||||
Stride<Stride<_128, _1, _16>, Stride<_64, _8, _512>>>;
|
||||
```
|
||||
|
||||
where we see 16 copies of the 64x8 tile.
|
||||
|
||||
### A and B Layout Mapping
|
||||
|
||||
GMMA atoms that consume A and B sources directly from shared memory are a bit interesting. The GMMA Descriptor is constructed on an entire tile of A and/or B data in shared memory rather than being partitioned by threads. That is, every thread sees the entire tile of data and the tile is not reordered so that the descriptor can be constructed on it. In `ALayout` form, this can be expressed
|
||||
|
||||
```cpp
|
||||
// (T128,V64x8) -> (M64,K16)
|
||||
using ALayout = Layout<Shape <_128, Shape <_64,_16>>,
|
||||
Stride< _0, Stride< _1,_64>>>;
|
||||
```
|
||||
|
||||
That is, all threads are mapped the to `(m,k) = (0,0) = 0` element and the values (and shape of the values) remains unchanged. The GMMA Descriptor Constructor can then inspect the `(M,K)` layout of this data and create an appropriate GMMA Descriptor or produce an error message saying the data is in an invalid layout for GMMA.
|
||||
|
||||
## `TiledMMA`s
|
||||
|
||||
We can make more complex patterns by combining and interleaving multiple atoms.
|
||||
|
||||
Let's start with `SM70_8x8x4_F32F16F16F32_NT`.
|
||||
```cpp
|
||||
MMA_Atom mma = MMA_Atom<SM70_8x8x4_F32F16F16F32_NT>{};
|
||||
print_latex(mma);
|
||||
```
|
||||
<p align="center">
|
||||
<img src="../../images/cute/HMMA.8x8x4.NT_Atom.png" alt="HMMA.8x8x4.NT_Atom.png" height="400"/>
|
||||
</p>
|
||||
|
||||
The above is equivalent to
|
||||
```cpp
|
||||
TiledMMA mma = make_tiled_mma(SM70_8x8x4_F32F16F16F32_NT{},
|
||||
Layout<Shape<_1,_1,_1>>{}, // Layout of Atoms
|
||||
Tile<_8,_8,_4>{}); // Tiler
|
||||
print_latex(mma);
|
||||
```
|
||||
as it is a single atom and has a natural tile size of 8x8x4.
|
||||
|
||||
We can create an object akin to a WMMA by using four of these quadpair MMAs:
|
||||
```cpp
|
||||
TiledMMA mma = make_tiled_mma(SM70_8x8x4_F32F16F16F32_NT{},
|
||||
Layout<Shape <_2,_2>,
|
||||
Stride<_2,_1>>{}); // 2x2 n-major layout of Atoms
|
||||
print_latex(mma);
|
||||
```
|
||||
<p align="center">
|
||||
<img src="../../images/cute/HMMA.8x8x4.NT_2x2.png" alt="HMMA.8x8x4.NT_2x2.png" height="400"/>
|
||||
</p>
|
||||
This `TiledMMA` replicates the `MMA_Atom` across threads as we can see the `T4` and `T8` and `T12` threads in the `C`-matrix that were not used before. Each quadrant of the `C`-matrix is a replica of the atom's partitioning pattern for a new quadpair and this replication follows a `(2,2):(2,1)` layout.
|
||||
|
||||
The above represents a 16x16x4 MMA now, but we can immediately expand this "tile size" up to 32x32x4 instead:
|
||||
```cpp
|
||||
TiledMMA mma = make_tiled_mma(SM70_8x8x4_F32F16F16F32_NT{},
|
||||
Layout<Shape <_2,_2>,
|
||||
Stride<_2,_1>>{}, // 2x2 n-major layout of Atoms
|
||||
Tile<_32,_32,_4>{}); // 32x32x4 tiler
|
||||
print_latex(mma);
|
||||
```
|
||||
<p align="center">
|
||||
<img src="../../images/cute/HMMA.8x8x4.NT_2x2_32x32x4.png" alt="HMMA.8x8x4.NT_2x2_32x32x4.png" height="400"/>
|
||||
</p>
|
||||
This `TiledMMA` replicates the previous `TiledMMA` across values instead of threads. We can see the `T0V8` and `T16V8` and `T8V8` values in the `C`-matrix that were not used before. Each quadrant of the `C`-matrix is a replica of the previous `TiledMMA`'s partitioning pattern for a new set of values.
|
||||
|
||||
Continuing, we see that there are eight values that `T0` receives from the `A`-matrix. Those reads occur at coordinates
|
||||
```
|
||||
T0V0 => ( 0,0)
|
||||
T0V1 => ( 1,0)
|
||||
T0V2 => ( 2,0)
|
||||
T0V3 => ( 3,0)
|
||||
T0V4 => (16,0)
|
||||
T0V5 => (17,0)
|
||||
T0V6 => (18,0)
|
||||
T0V7 => (19,0)
|
||||
```
|
||||
which are separate, but we might prefer them to be next to each other. That is we would like to permute the `M`-mode to create another valid `TiledMMA`.
|
||||
|
||||
```cpp
|
||||
TiledMMA mma = make_tiled_mma(SM70_8x8x4_F32F16F16F32_NT{},
|
||||
Layout<Shape <_2,_2>,
|
||||
Stride<_2,_1>>{}, // 2x2 n-major layout of Atoms
|
||||
Tile<Layout<Shape <_4,_4,_2>,
|
||||
Stride<_1,_8,_4>>, // Permutation on M, size 32
|
||||
_32, // Permutation on N, size 32 identity
|
||||
_4>{}); // Permutation on K, size 4 identity
|
||||
print_latex(mma);
|
||||
```
|
||||
<p align="center">
|
||||
<img src="../../images/cute/HMMA.8x8x4.NT_2x2_32Mx32x4.png" alt="HMMA.8x8x4.NT_2x2_32Mx32x4.png" height="400"/>
|
||||
</p>
|
||||
|
||||
That layout `(4,4,2):(1,8,4)` is read like a scatter permutation, telling the m-coords of the original image where to go in the new image.
|
||||
```
|
||||
old m-coord: 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31
|
||||
new m-coord: 0 1 2 3 8 9 10 11 16 17 18 19 24 25 26 27 4 5 6 7 12 13 14 15 20 21 22 23 28 29 30 31
|
||||
```
|
||||
This permutes only the M-mode (in `A` and `C` accordingly) and brings the access of all threads to be contiguous in m-coordinates in the `A`-matrix. This is convenient when designing layouts for shared memory or registers, for example. The MMA instructions contained within the image above are now effectively interleaved in the logical m-coordinates. Of course, permutations in the N-mode and K-mode are also valid.
|
||||
|
||||
To see how these `TiledMMA`s are used to partition data tensors, see the [`0x_gemm_tutorial.md`](./0x_gemm_tutorial.md).
|
||||
|
||||
## Copyright
|
||||
|
||||
Copyright (c) 2017 - 2025 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
```
|
||||
Redistribution and use in source and binary forms, with or without
|
||||
modification, are permitted provided that the following conditions are met:
|
||||
|
||||
1. Redistributions of source code must retain the above copyright notice, this
|
||||
list of conditions and the following disclaimer.
|
||||
|
||||
2. Redistributions in binary form must reproduce the above copyright notice,
|
||||
this list of conditions and the following disclaimer in the documentation
|
||||
and/or other materials provided with the distribution.
|
||||
|
||||
3. Neither the name of the copyright holder nor the names of its
|
||||
contributors may be used to endorse or promote products derived from
|
||||
this software without specific prior written permission.
|
||||
|
||||
THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS "AS IS"
|
||||
AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE
|
||||
IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE ARE
|
||||
DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT HOLDER OR CONTRIBUTORS BE LIABLE
|
||||
FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL
|
||||
DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR
|
||||
SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER
|
||||
CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY,
|
||||
OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE
|
||||
OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
|
||||
```
|
||||
570
media/docs/cpp/cute/0x_gemm_tutorial.md
Normal file
570
media/docs/cpp/cute/0x_gemm_tutorial.md
Normal file
@@ -0,0 +1,570 @@
|
||||
# CuTe dense matrix-matrix multiply tutorial
|
||||
|
||||
In this section, we review
|
||||
[these examples](https://github.com/NVIDIA/cutlass/tree/main/examples/cute/tutorial/),
|
||||
which demonstrate a few self-contained, single-file dense matrix-matrix multiply implementations using only CuTe.
|
||||
|
||||
## `sgemm_1.cu`
|
||||
|
||||
The simplest of the tutorial examples covers the basics of partitioning the global memory into tiles across the CTAs (also called threadblocks in CUDA), partitioning the data tiles across the threads of each CTA, and writing a mainloop using `cute::copy` and `cute::gemm`.
|
||||
|
||||
### High-level interface
|
||||
|
||||
We'll start with the kernel entry point `gemm_device` at the top of the file.
|
||||
|
||||
```c++
|
||||
template <class ProblemShape, class CtaTiler,
|
||||
class TA, class AStride, class ASmemLayout, class AThreadLayout,
|
||||
class TB, class BStride, class BSmemLayout, class BThreadLayout,
|
||||
class TC, class CStride, class CSmemLayout, class CThreadLayout,
|
||||
class Alpha, class Beta>
|
||||
__global__ static
|
||||
__launch_bounds__(decltype(size(CThreadLayout{}))::value)
|
||||
void
|
||||
gemm_device(ProblemShape shape_MNK, CtaTiler cta_tiler,
|
||||
TA const* A, AStride dA, ASmemLayout sA_layout, AThreadLayout tA,
|
||||
TB const* B, BStride dB, BSmemLayout sB_layout, BThreadLayout tB,
|
||||
TC * C, CStride dC, CSmemLayout , CThreadLayout tC,
|
||||
Alpha alpha, Beta beta)
|
||||
```
|
||||
|
||||
There are many template parameters, let's quickly review them and then go into more depth on their uses.
|
||||
|
||||
* `ProblemShape`. The MxNxK problem shape of this matrix multiply.
|
||||
|
||||
* `CtaTiler`. A CuTe [tiler concept](./02_layout_algebra.md#composition-tilers) that determines how to extract a tile of data from the problem shape.
|
||||
|
||||
* `TA const* A`, `TB const* B`, `TC* C`. The types and pointers to the A, B, and C data, respectively.
|
||||
|
||||
* `AStride`, `BStride`, `CStride`. The layout strides corresponding to the `ProblemShape` for each A, B, and C.
|
||||
|
||||
* `ASmemLayout`, `BSmemLayout`, `CSmemLayout`. The layouts, if needed, of shared memory to use for staging A-data, B-data, and C-data within each CTA.
|
||||
|
||||
* `AThreadLayout`, `BThreadLayout`, `CThreadLayout`. The layouts of threads to be used in partitioning each stage.
|
||||
|
||||
* `Alpha alpha`, `Beta beta`. The types and values of the scalar constants to compute GEMM: `C = alpha * A * B + beta * C`.
|
||||
|
||||
### The Full Tensors: Shapes, Strides, and Data
|
||||
|
||||
Most GEMM interfaces list the matrices' dimensions
|
||||
in the order M, N, K. CuTe also uses this convention, but packages them
|
||||
into a single `IntTuple`. In this example, they are dynamic values
|
||||
defined at the top of the `gemm_nt` and `gemm_tn` host functions
|
||||
that invoke the device kernel.
|
||||
```cpp
|
||||
// Define shapes (dynamic)
|
||||
auto M = int(m);
|
||||
auto N = int(n);
|
||||
auto K = int(k);
|
||||
auto prob_shape = make_shape(M, N, K); // (M, N, K)
|
||||
```
|
||||
|
||||
Inside the kernel, the problem shape is checked against the preconditions and then used to construct each of the full matrices.
|
||||
```cpp
|
||||
// Preconditions
|
||||
CUTE_STATIC_ASSERT_V(rank(shape_MNK) == Int<3>{}); // (M, N, K)
|
||||
|
||||
CUTE_STATIC_ASSERT_V(congruent(select<0,2>(shape_MNK), dA)); // dA strides for shape MK
|
||||
CUTE_STATIC_ASSERT_V(congruent(select<1,2>(shape_MNK), dB)); // dB strides for shape NK
|
||||
CUTE_STATIC_ASSERT_V(congruent(select<0,1>(shape_MNK), dC)); // dC strides for shape MN
|
||||
|
||||
// Represent the full tensors
|
||||
Tensor mA = make_tensor(make_gmem_ptr(A), select<0,2>(shape_MNK), dA); // (M,K)
|
||||
Tensor mB = make_tensor(make_gmem_ptr(B), select<1,2>(shape_MNK), dB); // (N,K)
|
||||
Tensor mC = make_tensor(make_gmem_ptr(C), select<0,1>(shape_MNK), dC); // (M,N)
|
||||
```
|
||||
The appropriate modes of the `Shape` are selected to construct each of the tensors. The preconditions make sure that for every integer in the `Shape` there is a corresponding integer in the associated `Stride`.
|
||||
|
||||
Note that the comment after B says `(N,K)` rather than `(K,N)`.
|
||||
This means that B is treated as an NxK matrix instead of a KxN matrix as is typical within BLAS and most other matrix-matrix multiplications.
|
||||
CuTe follows the convention that the semantics of matrix modes is
|
||||
`(M,K)` for `A`, `(N,K)` for `B`, and `(M,N)` for `C`, which we try to record in comments everywhere.
|
||||
|
||||
For each of the `(M,K)`, `(N,K)`, and `(M,N)` tensors, the `gemm_nt` and `gemm_tn` construct the strides those tensors will use. In `gemm_nt` the strides are defined as
|
||||
```cpp
|
||||
// Define NT strides (mixed)
|
||||
auto dA = make_stride(Int<1>{}, ldA); // (dM, dK)
|
||||
auto dB = make_stride(Int<1>{}, ldB); // (dN, dK)
|
||||
auto dC = make_stride(Int<1>{}, ldC); // (dM, dN)
|
||||
```
|
||||
and in `gemm_tn` the strides are defined as
|
||||
```cpp
|
||||
// Define TN strides (mixed)
|
||||
auto dA = make_stride(ldA, Int<1>{}); // (dM, dK)
|
||||
auto dB = make_stride(ldB, Int<1>{}); // (dN, dK)
|
||||
auto dC = make_stride(Int<1>{}, ldC); // (dM, dN)
|
||||
```
|
||||
|
||||
#### Aside: M-major, N-major, K-major
|
||||
|
||||
We've found that the BLAS convention of using "non-transposed" (N) and "transposed" (T) flags in conjunction with the mode conventions of `MxK * KxN` to confuse the core issue of "what layout does this matrix use" and "in which mode does my matrix have a stride-1?". Indeed, the answer to those questions can always be found by inspecting the CuTe `Layout`.
|
||||
|
||||
Instead of row-major or column-major (or Transposed
|
||||
and Not-Transposed), we have found it much more convenient to say that a matrix is "M-major" if it is stride-1 in the M-mode, "N-major" if it is stride-1 in the N-mode, or "K-major" if it is stride-1 in the K-mode. Furthermore, knowing that matrix multiply always performs a reduction in the K-mode, it is very convenient from a software perspective to always have the K-mode in the same place and adopt the mode convention `MxK * NxK`. Implementations will always reduce over the second mode (the K mode) of both input matrices and leads to cases where implementations can treat both input matrices the same way.
|
||||
|
||||
How do we translate this into the BLAS user's experience?
|
||||
|
||||
| BLAS | A Majorness | A Layout | B Majorness | B Layout |
|
||||
| --- | --- | --- | --- | --- |
|
||||
| NT | M-major | `(M,K):(1,ldA)` | N-major | `(N,K):(1,ldB)` |
|
||||
| TN | K-major | `(M,K):(ldA,1)` | K-major | `(N,K):(ldB,1)` |
|
||||
| NN | M-major | `(M,K):(1,ldA)` | K-major | `(N,K):(ldB,1)` |
|
||||
| TT | K-major | `(M,K):(ldA,1)` | N-major | `(N,K):(1,ldB)` |
|
||||
|
||||
Regardless, we'll still use the BLAS "NT" and "TN" notations for high-level descriptions of kernels when it's appropriate.
|
||||
|
||||
### CTA Partitioning
|
||||
|
||||
Now that we have the representations of the full matrices, it's time to tile them and split up the work!
|
||||
|
||||
At the highest level, the work is distributed across CTAs. In principle, each CTA's tile could come from the input tensors in many different ways. Many [CuTe `Tiler`s](./02_layout_algebra.md#composition-tilers) could be used to tile the data, but for these cases it is sufficient to simply use the shape of the desired CTA tile.
|
||||
```cpp
|
||||
// Define CTA tile sizes (static)
|
||||
auto bM = Int<128>{};
|
||||
auto bN = Int<128>{};
|
||||
auto bK = Int< 8>{};
|
||||
auto cta_tiler = make_shape(bM, bN, bK); // (BLK_M, BLK_N, BLK_K)
|
||||
```
|
||||
|
||||
Once the tiler has been defined, we can use it to tile and partition the tensors across the CTAs.
|
||||
|
||||
```cpp
|
||||
// Get the appropriate blocks for this threadblock
|
||||
auto cta_coord = make_coord(blockIdx.x, blockIdx.y, _); // (m,n,k)
|
||||
Tensor gA = local_tile(mA, cta_tiler, cta_coord, Step<_1, X,_1>{}); // (BLK_M,BLK_K,k)
|
||||
Tensor gB = local_tile(mB, cta_tiler, cta_coord, Step< X,_1,_1>{}); // (BLK_N,BLK_K,k)
|
||||
Tensor gC = local_tile(mC, cta_tiler, cta_coord, Step<_1,_1, X>{}); // (BLK_M,BLK_N)
|
||||
```
|
||||
|
||||
First, the CTA coordinate is created.
|
||||
* The `m`-coordinate of this tile is given by `blockIdx.x`.
|
||||
* The `n`-coordinate of this tile is given by `blockIdx.y`.
|
||||
* The `k`-coordinate of this tile is unspecified -- we want all of the tiles in `K` so the coordinate is `_`, the `Underscore` value, to keep that mode.
|
||||
|
||||
Then, `local_tile` is used to remove the modes of the tiler and coord corresponding to the `X`s. That is, the `Step<_1, X,_1>` is just shorthand for
|
||||
```cpp
|
||||
// Use select<0,2> to use only the M- and K-modes of the tiler and coord
|
||||
Tensor gA = local_tile(mA, select<0,2>(cta_tiler), select<0,2>(cta_coord));
|
||||
```
|
||||
This `local_tile` is simply shorthand for
|
||||
1. apply the tiler via [`zipped_divide`](./02_layout_algebra.md#zipped-tiled-flat-divides)
|
||||
```cpp
|
||||
// ((BLK_M,BLK_K),(m,k))
|
||||
Tensor gA_mk = zipped_divide(mA, select<0,2>(cta_tiler));
|
||||
```
|
||||
2. apply the coord to the second mode, the "Rest" mode, to extract out the correct tiles for this CTA.
|
||||
```cpp
|
||||
// (BLK_M,BLK_K,k)
|
||||
Tensor gA = gA_mk(make_coord(_,_), select<0,2>(cta_coord));
|
||||
```
|
||||
Because the projections of the tiler and coord are symmetric and the two steps (apply a tiler and then slice into the rest-mode to produce a partition) are so common, they are wrapped together into the projective `local_tile` interface.
|
||||
|
||||
For tensor `A`, we are left with a rank-3 tensor of shape `(BLK_M,BLK_K,k)`. The first two modes are precisely the modes of the CTA tile and the last mode indexes over all of the tiles that will be reduced by this CTA. In the mainloop section below, this mode is iterated over via the `k_tile` loop.
|
||||
|
||||
### SMEM tensors
|
||||
|
||||
The shared memory layouts that are used to hold the tiles of data for A and B are also passed in as the parameters `ASmemLayout sA_layout` and `BSmemLayout sB_layout`.
|
||||
|
||||
These are defined in `gemm_nt` as
|
||||
```c++
|
||||
// Define the smem layouts (static)
|
||||
auto sA = make_layout(make_shape(bM, bK)); // (m,k) -> smem_idx; m-major
|
||||
auto sB = make_layout(make_shape(bN, bK)); // (n,k) -> smem_idx; n-major
|
||||
```
|
||||
which produces simple M-major and N-major layouts. In `gemm_tn` these are
|
||||
```cpp
|
||||
// Define the smem layouts (static)
|
||||
auto sA = make_layout(make_shape(bM,bK), LayoutRight{}); // (m,k) -> smem_idx; k-major
|
||||
auto sB = make_layout(make_shape(bN,bK), LayoutRight{}); // (n,k) -> smem_idx; k-major
|
||||
```
|
||||
which produces simple K-major layouts.
|
||||
|
||||
As is evident, these smem layouts can be almost anything. Inside the kernel, they are checked for only two properties: the shared memory layouts are static and they are the same top-level shape as the `CtaTiler`.
|
||||
|
||||
```cpp
|
||||
// Preconditions
|
||||
static_assert(is_static<ASmemLayout>::value);
|
||||
static_assert(is_static<BSmemLayout>::value);
|
||||
static_assert(is_static<CSmemLayout>::value);
|
||||
|
||||
CUTE_STATIC_ASSERT_V(size<0>(ASmemLayout{}) == size<0>(cta_tiler)); // BLK_M
|
||||
CUTE_STATIC_ASSERT_V(size<0>(CSmemLayout{}) == size<0>(cta_tiler)); // BLK_M
|
||||
CUTE_STATIC_ASSERT_V(size<0>(BSmemLayout{}) == size<1>(cta_tiler)); // BLK_N
|
||||
CUTE_STATIC_ASSERT_V(size<1>(CSmemLayout{}) == size<1>(cta_tiler)); // BLK_N
|
||||
CUTE_STATIC_ASSERT_V(size<1>(ASmemLayout{}) == size<2>(cta_tiler)); // BLK_K
|
||||
CUTE_STATIC_ASSERT_V(size<1>(BSmemLayout{}) == size<2>(cta_tiler)); // BLK_K
|
||||
```
|
||||
|
||||
Use of static layouts has a few advantages.
|
||||
* Static layouts let us statically allocate shared memory as shown below.
|
||||
* Static layouts are often more efficient and allow CuTe to dispatch to optimized implementations.
|
||||
* Static layouts makes it easier to prove correctness of the algorithm and provide checks like the above -- the smem layout sizes are the same as the CTA tile sizes.
|
||||
|
||||
As stated, the shared memory layouts can be anything that satisfy those conditions. Optimizing kernels like these is often performed by finding a good shared memory layout that provides good access patterns for both the writes to and the reads from shared memory. This includes the ability to vectorize reads and writes as well as avoid shared memory bank conflicts.
|
||||
|
||||
With the static smem layouts, the `gemm_device` kernel can allocate the required shared memory and create the smem `Tensor`s.
|
||||
|
||||
```cpp
|
||||
// Shared memory buffers
|
||||
__shared__ TA smemA[cosize_v<ABlockLayout>];
|
||||
__shared__ TB smemB[cosize_v<BBlockLayout>];
|
||||
Tensor sA = make_tensor(make_smem_ptr(smemA), sA_layout); // (BLK_M,BLK_K)
|
||||
Tensor sB = make_tensor(make_smem_ptr(smemB), sB_layout); // (BLK_N,BLK_K)
|
||||
```
|
||||
|
||||
Note how the shared memory allocation depends only on the data type and the layout. What's a `cosize`? Because a `Layout` is a function, we can speak of its domain and codomain. The `size` of a layout is the size of its domain and the `cosize` of a layout is the size of its codomain. If we want to allocate an array for which all the offsets produced by a layout are valid, then we can use the `cosize` of the layout as the length of the array (in units of elements).
|
||||
|
||||
### Copy partitioning
|
||||
|
||||
The kernel now has tiles of global memory by applying the `CtaTiler` to the full tensors and it also has tiles of shared memory by allocating appropriately. We now want to create an efficient way to copy one tile of global memory to our tile of shared memory. A trivial way to do this would be to use a single thread and copy each element.
|
||||
```cpp
|
||||
if (thread0()) {
|
||||
Tensor gA0 = gA(_,_,0); // (BLK_M,BLK_K), the 0th tile
|
||||
for (int i = 0; i < size(sA); ++i) {
|
||||
sA(i) = gA0(i);
|
||||
}
|
||||
}
|
||||
```
|
||||
This would work, but we have lots of threads to use inside this CTA, so let's use them!
|
||||
|
||||
If we partition the two tiles of data across the threads in the CTA, then each thread can copy its own subtensor of data. There are lots of ways this partitioning could occur, however.
|
||||
|
||||
The `gemm_nt` function defines two layouts of *threads* as
|
||||
```c++
|
||||
// Define thread layouts (static)
|
||||
auto tA = make_layout(make_shape(Int<32>{},Int<8>{})); // (m,k) -> thr_idx
|
||||
auto tB = make_layout(make_shape(Int<32>{},Int<8>{})); // (n,k) -> thr_idx
|
||||
```
|
||||
and the `gemm_tn` functions defines two layouts of *threads* as
|
||||
```c++
|
||||
// Define thread layouts (static)
|
||||
auto tA = make_layout(make_shape(Int<32>{},Int<8>{}), LayoutRight{}); // (m,k) -> thr_idx; k-major
|
||||
auto tB = make_layout(make_shape(Int<32>{},Int<8>{}), LayoutRight{}); // (n,k) -> thr_idx; k-major
|
||||
```
|
||||
Both cases happen to use 32x8 threads, which will be used to partition a 128x8 tile of gmem and smem data into a 4x1 subtensor for each thread. The only difference here is that `gemm_nt` uses M-major and N-major threads to match the order of data in global memory and `gemm_tn` uses K-major threads to match the order of data in global memory.
|
||||
|
||||
Again, the conditions on the thread layouts are checked inside the kernel.
|
||||
```cpp
|
||||
static_assert(is_static<AThreadLayout>::value);
|
||||
static_assert(is_static<BThreadLayout>::value);
|
||||
|
||||
CUTE_STATIC_ASSERT_V(size(tA) == size(tB)); // NumThreads
|
||||
|
||||
CUTE_STATIC_ASSERT_V(size<0>(cta_tiler) % size<0>(tA) == Int<0>{}); // BLK_M / THR_M
|
||||
CUTE_STATIC_ASSERT_V(size<2>(cta_tiler) % size<1>(tA) == Int<0>{}); // BLK_K / THR_K
|
||||
CUTE_STATIC_ASSERT_V(size<1>(cta_tiler) % size<0>(tB) == Int<0>{}); // BLK_N / THR_N
|
||||
CUTE_STATIC_ASSERT_V(size<2>(cta_tiler) % size<1>(tB) == Int<0>{}); // BLK_K / THR_K
|
||||
```
|
||||
|
||||
These thread layouts are then used to partition the global memory tensors data and shared memory tensors
|
||||
```cpp
|
||||
Tensor tAgA = local_partition(gA, tA, threadIdx.x); // (THR_M,THR_K,k)
|
||||
Tensor tAsA = local_partition(sA, tA, threadIdx.x); // (THR_M,THR_K)
|
||||
|
||||
Tensor tBgB = local_partition(gB, tB, threadIdx.x); // (THR_N,THR_K,k)
|
||||
Tensor tBsB = local_partition(sB, tB, threadIdx.x); // (THR_N,THR_K)
|
||||
|
||||
CUTE_STATIC_ASSERT_V(size<0>(tAgA) == size<0>(tAsA)); // THR_M
|
||||
CUTE_STATIC_ASSERT_V(size<1>(tAgA) == size<1>(tAsA)); // THR_K
|
||||
CUTE_STATIC_ASSERT_V(size<0>(tBgB) == size<0>(tBsB)); // THR_N
|
||||
CUTE_STATIC_ASSERT_V(size<1>(tBgB) == size<1>(tBsB)); // THR_K
|
||||
```
|
||||
where `local_partition` is a lot like `local_tile`, except the coordinate slices into the tile-mode (the first mode) of the `zipped_divide` rather than the rest-mode (the second mode). That is, each thread gets one element of data assigned to it per thread tile and that thread tile is repeated to cover the entire data tile.
|
||||
|
||||
The naming convention `tAsA` is pretty typical across CuTe and CUTLASS. This is read as "Partitioning pattern `tA` applied to tensor `sA`". In the next section, we'll see a different partitioner applied to `sA` to produce `tCsA`. By applying the same partitioning pattern, `tA`, to tensors `sA` and `gA`, we preserve the *logical consistency* of those tensors (checked by the assertions above) where logical elements between the two tensors correspond despite any differences in their data layouts. When used in `cute::copy`, for example, this naming convention let's us lexically verify that the two tensors are using the same partitioning pattern.
|
||||
|
||||
With the data partitioned across the threads, *every thread* can now participate in the copy by writing
|
||||
```cpp
|
||||
copy(tAgA(_,_,0), tAsA);
|
||||
```
|
||||
because every thread owns a different subtensor of the tile that will be copied.
|
||||
|
||||
### Math partitioning
|
||||
|
||||
The kernel now has tiles of shared memory copied in from global memory. We now want to create an efficient way to compute and accumulate the matrix product on that tile of shared memory. A trivial way to do this would be to use a single thread and compute directly.
|
||||
```cpp
|
||||
if (thread0()) {
|
||||
for (int m = 0; m < size<0>(gC); ++m) {
|
||||
for (int n = 0; n < size<1>(gC); ++n) {
|
||||
for (int k = 0; k < size<1>(sA); ++k) {
|
||||
gC(m,n) += sA(m,k) * sB(n,k);
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
This would work, but we have lots of threads to use inside this CTA, so let's use them!
|
||||
|
||||
If we partition the output tile `gC` across the threads in the CTA, then each thread can compute its own subtensor. There are lots of ways this partitioning could occur, however.
|
||||
|
||||
The `gemm_nt` and `gemm_tn` functions define one more layout of *threads*:
|
||||
```cpp
|
||||
// Define thread layouts (static)
|
||||
auto tC = make_layout(make_shape(Int<16>{}, Int<16>{})); // (m,n) -> thr_idx; m-major
|
||||
```
|
||||
This is a m-major 16x16 layout of threads which will be used to partition a 128x128 tile of `C`-data, resulting in each thread computing its own 8x8 subtensor of `gC`.
|
||||
|
||||
Again, the conditions on the thread layouts are checked inside the kernel.
|
||||
```cpp
|
||||
static_assert(is_static<CThreadLayout>::value);
|
||||
|
||||
CUTE_STATIC_ASSERT_V(size(tC) == size(tA)); // NumThreads
|
||||
|
||||
CUTE_STATIC_ASSERT_V(size<0>(cta_tiler) % size<0>(tC) == Int<0>{}); // BLK_M / THR_M
|
||||
CUTE_STATIC_ASSERT_V(size<1>(cta_tiler) % size<1>(tC) == Int<0>{}); // BLK_N / THR_N
|
||||
```
|
||||
|
||||
These thread layouts are then used to partition the tiles of data in global memory and shared memory
|
||||
```cpp
|
||||
// Partition sA (M,K) by the rows of tC
|
||||
Tensor tCsA = local_partition(sA, tC, threadIdx.x, Step<_1, X>{}); // (THR_M,BLK_K)
|
||||
// Partition sB (N,K) by the cols of tC
|
||||
Tensor tCsB = local_partition(sB, tC, threadIdx.x, Step< X,_1>{}); // (THR_N,BLK_K)
|
||||
// Partition gC (M,N) by the tile of tC
|
||||
Tensor tCgC = local_partition(gC, tC, threadIdx.x, Step<_1,_1>{}); // (THR_M,THR_N)
|
||||
|
||||
// Allocate the accumulators -- same shape/layout as the partitioned data
|
||||
Tensor tCrC = make_tensor_like(tCgC); // (THR_M,THR_N)
|
||||
|
||||
CUTE_STATIC_ASSERT_V(size<0>(tCrC) == size<0>(tCgC)); // THR_M
|
||||
CUTE_STATIC_ASSERT_V(size<0>(tCrC) == size<0>(tCsA)); // THR_M
|
||||
CUTE_STATIC_ASSERT_V(size<1>(tCrC) == size<1>(tCgC)); // THR_N
|
||||
CUTE_STATIC_ASSERT_V(size<1>(tCrC) == size<0>(tCsB)); // THR_N
|
||||
CUTE_STATIC_ASSERT_V(size<1>(tCsA) == size<1>(tCsB)); // BLK_K
|
||||
```
|
||||
where we've used the same projection-style interface to avoid applying the `N`-mode of `tC` to the `(BLK_M,BLK_K)` shape of `sA` and avoid applying the `M`-mode of `tC` to the `(BLK_N,BLK_K)` shape of `sB`.
|
||||
|
||||
<p align="center">
|
||||
<img src="../../images/cute/tC_partitioning.png" alt="tC_partitioning.png" height="300"/>
|
||||
</p>
|
||||
This diagram shows a `tC` layout, highlights two threads in green and blue, shows the projections of the `tC` layout, and finally highlights the subtensors within `sA`, `sB`, and `gC` that `tCsA`, `tCsB`, and `tCgC` represent.
|
||||
|
||||
With the data partitioned across the threads, *every thread* can now participate in the compute step by writing
|
||||
```cpp
|
||||
gemm(tCsA, tCsB, tCrC);
|
||||
```
|
||||
because every thread owns different subtensors of the data to be computed.
|
||||
|
||||
### Mainloop
|
||||
|
||||
The mainloop iterates over tiles of global memory, reads those tiles into shared memory, and then performs the matrix-multiply and accumulates into the accumulators.
|
||||
|
||||
```c++
|
||||
// TUTORIAL: Example of a very simple compute mainloop
|
||||
// copy(.) operates on the global and shared memory via the tA|tB partitioning
|
||||
// gemm(.) operates on the shared and register memory via the tC partitioning
|
||||
|
||||
auto K_TILE_MAX = size<2>(tAgA);
|
||||
|
||||
for (int k_tile = 0; k_tile < K_TILE_MAX; ++k_tile)
|
||||
{
|
||||
// Copy gmem to smem with tA|tB thread-partitioned tensors
|
||||
copy(tAgA(_,_,k_tile), tAsA); // A (THR_M,THR_K) -> (THR_M,THR_K)
|
||||
copy(tBgB(_,_,k_tile), tBsB); // B (THR_N,THR_K) -> (THR_N,THR_K)
|
||||
|
||||
cp_async_fence(); // Label the end of (potential) cp.async instructions
|
||||
cp_async_wait<0>(); // Sync on all (potential) cp.async instructions
|
||||
__syncthreads(); // Wait for all threads to write to smem
|
||||
|
||||
// Compute gemm on tC thread-partitioned smem
|
||||
gemm(tCsA, tCsB, tCrC); // (THR_M,THR_N) += (THR_M,BLK_K) * (THR_N,BLK_K)
|
||||
__syncthreads(); // Wait for all threads to read from smem
|
||||
}
|
||||
```
|
||||
|
||||
We can see that `k_tile` iterates over each tile of data, the `cute::copy` is performed for the current `k_tile` using the `tA` and `tB` thread-partitioned tensors, and the `cute::gemm` is computed for that current `k_tile` using the `tC` thread-partitioned tensors. Synchronization is provided so that this kernel works on any architecture.
|
||||
|
||||
## `sgemm_2.cu`
|
||||
|
||||
An example that uses more complex `TiledMMA` and `TiledCopy` to perform partitioning in place of the `tA`, `tB`, and `tC` thread layouts. With this example, we try to emphasize that the shared memory layouts, the partitioning patterns, and the PTX instruction to use in each stage can be specified independently.
|
||||
|
||||
### TiledCopy
|
||||
|
||||
First, we can replace the `tA` partitioning and `tB` partitioning with `TiledCopy` partitioning, which provides for more complex partitioning patterns and checked dispatch to specific copy instructions.
|
||||
|
||||
As a first example, lets look at the `TiledCopy` that `gemm_nt` generates.
|
||||
```cpp
|
||||
TiledCopy copyA = make_tiled_copy(Copy_Atom<UniversalCopy<uint128_t>, TA>{}, // Atom: Copy TAs as if they were uint128_t
|
||||
Layout<Shape<_32,_8>>{}, // Thr layout 32x8 m-major
|
||||
Layout<Shape< _4,_1>>{}); // Val layout 4x1 m-major
|
||||
print_latex(copyA);
|
||||
```
|
||||
The easiest way to see what this `TiledCopy` does is to look at the partition pattern in LaTeX.
|
||||
<p align="center">
|
||||
<img src="../../images/cute/TiledCopyA.png" alt="TiledCopyA.png" height="300"/>
|
||||
</p>
|
||||
On the left is the source-tensor partitioning and on the right is the destination-tensor partitioning. The partition patterns are the same for this case, but there exist PTX instructions which require different patterns in the source and destination. The diagram shows that each thread reads 4x1 `TA` elements and there are 32x8 threads. The `UniversalCopy<uint128_t>` forces the instruction to use a 128-bit copy instruction. If the partition (of `sA` or `gA` in this case) does not result in 4 `TA` elements that cannot be vectorized to a 128-bit load/store, then CuTe will statically fail with an error message to that effect.
|
||||
|
||||
To use the `TiledCopy`, the kernel writes
|
||||
```cpp
|
||||
ThrCopy thr_copy_a = copy_a.get_slice(threadIdx.x);
|
||||
Tensor tAgA = thr_copy_a.partition_S(gA); // (CPY,CPY_M,CPY_K,k)
|
||||
Tensor tAsA = thr_copy_a.partition_D(sA); // (CPY,CPY_M,CPY_K)
|
||||
// Allocate registers same shape/layout as partitioned data
|
||||
Tensor tArA = make_fragment_like(tAsA); // (CPY,CPY_M,CPY_K)
|
||||
```
|
||||
which applies the source-tensor partitioning to `gA` via `partition_S` and applies the destination-tensor partitioning to `sA` via `partition_D`. The first mode, `CPY`, of the result tensors hold all of the elements that a single instruction will consume. In this case, that mode should have size-4 since there are four `TA=float` elements in a single 128-bit `uint128_t`.
|
||||
|
||||
Once the partition has been performed, we can execute the `copy` on the thread-partitioned tensors using the provided instruction in `copy_a`.
|
||||
```cpp
|
||||
cute::copy(copy_a, tAgA, tArA);
|
||||
```
|
||||
|
||||
### TiledMMA
|
||||
|
||||
Next, we can replace the `tC` partitioning with `TiledMMA` partitioning, which provides for more complex partitioning patterns and checked dispatch to specific MMA instructions.
|
||||
|
||||
As a first example, lets look at the `TiledMMA` that `gemm_nt` generates.
|
||||
```cpp
|
||||
TiledMMA mmaC = make_tiled_mma(UniversalFMA<TC,TA,TB>{},
|
||||
Layout<Shape<_16,_16,_1>>{}); // 16x16x1 UniversalFMA
|
||||
print_latex(mmaC);
|
||||
```
|
||||
The easiest way to see what this `TiledMMA` does is to look at the partition pattern in LaTeX.
|
||||
<p align="center">
|
||||
<img src="../../images/cute/TiledMmaC.png" alt="TiledMmaC.png" height="300"/>
|
||||
</p>
|
||||
On the left is the A-tensor partitioning, on the top is the B-tensor partitioning, and in the middle is the C-tensor partitioning.Because the `UniversalFMA` is a 1x1x1 MMA instruction, a 16x16x1 tiling of them results in a 16x16x1 `TiledMMA`. Other MMA instructions will have different threads involved and have different instruction sizes. In this case, all threads will read a single element from `A`, `B`, and `C` each.
|
||||
|
||||
To use the `TiledMMA`, the kernel writes
|
||||
```cpp
|
||||
ThrMMA thr_mma = mma.get_slice(threadIdx.x);
|
||||
Tensor tCsA = thr_mma.partition_A(sA); // (MMA,MMA_M,MMA_K)
|
||||
Tensor tCsB = thr_mma.partition_B(sB); // (MMA,MMA_N,MMA_K)
|
||||
Tensor tCgC = thr_mma.partition_C(gC); // (MMA,MMA_M,MMA_N)
|
||||
// Allocate the accumulators -- same size as the projected data
|
||||
Tensor tCrC = thr_mma.make_fragment_C(tCgC); // (MMA,MMA_M,MMA_N)
|
||||
```
|
||||
which applies the A-tensor partitioning to `sA` via `partition_A`, applies the B-tensor partitioning to `sB` via `partition_B`, and applies the C-tensor partitioning to `gC` via `partition_C`. The first mode, `MMA`, of the result tensors hold all of the elements that a single instruction will consume. In this case, that mode should have size-1 since `UniversalFMA` is a 1x1x1 MMA, but in general the size of the first mode can vary and not even be the same across `tCsA`, `tCsB`, and `tCgC` depending on the MMA.
|
||||
|
||||
Once the partition has been performed, we can execute the `gemm` on the thread-partitioned tensors using the provided instruction in `mma`.
|
||||
```cpp
|
||||
cute::gemm(mma, tCsA, tCsB, tCrC);
|
||||
```
|
||||
|
||||
### Other changes
|
||||
|
||||
In this version, we have also updated the shared memory layouts for `gemm_tn` from K-major to
|
||||
```cpp
|
||||
// Define the smem layouts (static)
|
||||
auto sA = make_layout(make_shape ( bM, bK),
|
||||
make_stride(Int<1>{}, bM+Int<1>{})); // (m,k) -> smem_idx; padded m-major
|
||||
auto sB = make_layout(make_shape ( bN, bK),
|
||||
make_stride(Int<1>{}, bN+Int<1>{})); // (n,k) -> smem_idx; padded n-major
|
||||
```
|
||||
which produces M-major and N-major layouts, but they are padded to avoid shared memory bank conflicts. This simply improves the access pattern to and from shared memory and no other changes in the kernel are required.
|
||||
|
||||
## `sgemm_sm70.cu`
|
||||
|
||||
An example that uses an optimized mainloop for Volta SM70 architectures that pipelines shared memory and register memory.
|
||||
|
||||
## `sgemm_sm80.cu`
|
||||
|
||||
An example that uses an optimized mainloop for Ampere SM80 architectures that explicitly pipelines shared memory using asynchronous reads from global memory.
|
||||
|
||||
## Next steps
|
||||
|
||||
All of the above examples assume that the CTA tile size divides the problem size so that global memory loads do no need to be predicated. The
|
||||
[predication section of the tutorial](./0y_predication.md)
|
||||
explains what to do if a matrix tiling
|
||||
doesn't perfectly divide the matrix.
|
||||
|
||||
## GETT as GEMM
|
||||
|
||||
"GETT" here stands for "general(ized) tensor times tensor," a tensor contraction.
|
||||
|
||||
CuTe permits matrices to have nested `Layout`s.
|
||||
This means that we can fold a `Tensor` into a "matrix" by grouping modes according to their categories.
|
||||
|
||||
As a result, we can implement GETT by using
|
||||
our existing GEMM implementation. Included below is a launcher like `gemm_nt` that uses the same device kernel contained in `sgemm_1.cu` to compute a GETT with two m-modes.
|
||||
```cpp
|
||||
// Setup params for a GETT with two m-modes.
|
||||
// The A and C tensors are assumed to be m0-major.
|
||||
// Calls sgemm_1.cu's gemm_device<<<>>> without modification.
|
||||
template <class TA, class TB, class TC,
|
||||
class Alpha, class Beta>
|
||||
void
|
||||
gett(int m0, int m1, int n, int k,
|
||||
Alpha alpha,
|
||||
TA const* A, int ldAm1, int ldAk, // m0-major
|
||||
TB const* B, int ldBk,
|
||||
Beta beta,
|
||||
TC * C, int ldCm1, int ldCn, // m0-major
|
||||
cudaStream_t stream = 0)
|
||||
{
|
||||
using namespace cute;
|
||||
|
||||
// Define shapes (dynamic)
|
||||
auto M = make_shape(m0, m1); // (m0,m1)-multimode M
|
||||
auto N = int(n);
|
||||
auto K = int(k);
|
||||
auto prob_shape = make_shape(M, N, K); // (M, N, K)
|
||||
|
||||
// Define NT strides (mixed)
|
||||
auto dA = make_stride(make_stride(Int<1>{}, ldAm1), ldAk); // (dM, dK)
|
||||
auto dB = make_stride(Int<1>{}, ldB); // (dN, dK)
|
||||
auto dC = make_stride(make_stride(Int<1>{}, ldCm1), ldCn); // (dM, dN)
|
||||
|
||||
// Define CTA tile sizes (static)
|
||||
auto bM = Shape<_64, _2>{}; // Take _64 elements from m0 and _2 elements from m1
|
||||
auto bN = Int<128>{};
|
||||
auto bK = Int< 8>{};
|
||||
auto cta_tiler = make_shape(bM, bN, bK); // (BLK_M, BLK_N, BLK_K)
|
||||
|
||||
// Define the smem layouts (static)
|
||||
auto sA = make_layout(make_shape(bM, bK)); // (m,k) -> smem_idx; m-major
|
||||
auto sB = make_layout(make_shape(bN, bK)); // (n,k) -> smem_idx; n-major
|
||||
auto sC = make_layout(make_shape(bM, bN)); // (m,n) -> smem_idx; m-major
|
||||
|
||||
// Define the thread layouts (static)
|
||||
auto tA = make_layout(make_shape(Int<32>{}, Int< 8>{})); // (m,k) -> thr_idx
|
||||
auto tB = make_layout(make_shape(Int<32>{}, Int< 8>{})); // (n,k) -> thr_idx
|
||||
auto tC = make_layout(make_shape(Int<16>{}, Int<16>{})); // (m,n) -> thr_idx
|
||||
|
||||
dim3 dimBlock(size(tC));
|
||||
dim3 dimGrid(size(ceil_div(M, bM)),
|
||||
size(ceil_div(N, bN)));
|
||||
gemm_device<<<dimGrid, dimBlock, 0, stream>>>
|
||||
(prob_shape, cta_tiler,
|
||||
A, dA, sA, tA,
|
||||
B, dB, sB, tB,
|
||||
C, dC, sC, tC,
|
||||
alpha, beta);
|
||||
}
|
||||
```
|
||||
Note that the only changes are the definition of shape `M`, the definition of strides `dA` and `dC`, and the definition of the CTA Tiler `bM`. The above uses a multimodel problem shape `M = (m0,m1)` and a multimodal CTA Tiler `bM = <_64,_2>` to change which portion of the global memory tensors `A` and `C` each CTA will be responsible for computing.
|
||||
|
||||
Similar examples can be found for CUTLASS 3.x kernels that are based on CuTe, such as [this Hopper GETT example](https://github.com/NVIDIA/cutlass/tree/main/examples/51_hopper_gett).
|
||||
|
||||
## Copyright
|
||||
|
||||
Copyright (c) 2017 - 2025 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
```
|
||||
Redistribution and use in source and binary forms, with or without
|
||||
modification, are permitted provided that the following conditions are met:
|
||||
|
||||
1. Redistributions of source code must retain the above copyright notice, this
|
||||
list of conditions and the following disclaimer.
|
||||
|
||||
2. Redistributions in binary form must reproduce the above copyright notice,
|
||||
this list of conditions and the following disclaimer in the documentation
|
||||
and/or other materials provided with the distribution.
|
||||
|
||||
3. Neither the name of the copyright holder nor the names of its
|
||||
contributors may be used to endorse or promote products derived from
|
||||
this software without specific prior written permission.
|
||||
|
||||
THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS "AS IS"
|
||||
AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE
|
||||
IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE ARE
|
||||
DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT HOLDER OR CONTRIBUTORS BE LIABLE
|
||||
FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL
|
||||
DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR
|
||||
SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER
|
||||
CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY,
|
||||
OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE
|
||||
OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
|
||||
```
|
||||
249
media/docs/cpp/cute/0y_predication.md
Normal file
249
media/docs/cpp/cute/0y_predication.md
Normal file
@@ -0,0 +1,249 @@
|
||||
# Predication: What to do when tiling isn't perfect
|
||||
|
||||
The [GEMM tutorial](./0x_gemm_tutorial.md) shows how
|
||||
we compute a matrix-matrix multiply
|
||||
by iterating over tiles of the input matrices and output matrix.
|
||||
The examples all assume that the tiles fit evenly into the matrices,
|
||||
with no remainder.
|
||||
What do we do if this is not the case?
|
||||
For example, we might want to tile a 41 x 55 matrix into 4 x 8 tiles,
|
||||
but 41 / 4 is 10 remainder 1, and 55 / 8 is 6 remainder 7.
|
||||
What do we do with those "leftover" parts of the matrix?
|
||||
|
||||
Another way to say this, is that `logical_divide`
|
||||
(CuTe's way of tiling layouts) "rounds up."
|
||||
For example, if `N` is the layout (1000, 1) and `B` is the layout (128, 1),
|
||||
then `logical_divide(N, B)` is the layout ((128, 8), (1, 128)).
|
||||
This effectively rounds up the original shape N = 1000
|
||||
into an 128 x 8 matrix (as if N = 1024).
|
||||
What about those last 24 elements,
|
||||
that aren't part of the original data?
|
||||
|
||||
The idiomatic CuTe way to solve this problem is through "predication."
|
||||
Rather than trying to reason about the "remainder tiles,"
|
||||
CuTe instead rounds up, but only tries to access data in each tile
|
||||
that are part of the matrix.
|
||||
This corresponds well with how our GPUs optimize:
|
||||
branches without warp divergence are relatively fast.
|
||||
It also matches the usual CUDA idiom
|
||||
when dividing N work items in 1-D fashion over B thread blocks:
|
||||
first test if "my thread" is out of bounds before doing work.
|
||||
|
||||
There are a few ways to figure out
|
||||
which elements need to be predicated.
|
||||
In-kernel GEMMs like to do this in the following way.
|
||||
|
||||
```c++
|
||||
// Create the predicate tensor
|
||||
Layout idA = make_layout(shape(A)); // e.g. 1000:1
|
||||
Layout idAB = logical_divide(idA, B); // e.g. (128,8):(1,128)
|
||||
|
||||
Tensor pred = make_tensor<bool>(shape(idAB));
|
||||
for (int i = 0; i < size(pred); ++i) {
|
||||
pred(i) = idAB(i) < size(A);
|
||||
}
|
||||
|
||||
// ... intervening code ...
|
||||
|
||||
// Use the predicate tensor. c is some coordinate.
|
||||
// This code would likely live inside some algorithm.
|
||||
if (pred(c)) { copy(idAB(c), smem(c)); }
|
||||
```
|
||||
|
||||
The general procedure is that we
|
||||
|
||||
1. create an "identity" layout (`Layout idA = make_layout(shape(A))`,
|
||||
in the above example) with the same shape as our original data;
|
||||
|
||||
2. repeat the same tiling/partitioning/slicing (possibly rounding up)
|
||||
on that identity layout (`Layout idAB = logical_divide(idA, B)`);
|
||||
|
||||
3. create a "predicate tensor" by comparing the coordinates
|
||||
of that reference layout with the bounds of the original layout;
|
||||
and then
|
||||
|
||||
4. use the predicate tensor to mask off accesses to out-of-bounds elements.
|
||||
|
||||
For example, suppose that we've partitioned A and B tiles
|
||||
across threads as follows.
|
||||
|
||||
```c++
|
||||
Tensor tAgA = local_partition(gA, tA, thread_idx); // (THR_M,THR_K,k)
|
||||
Tensor tAsA = local_partition(sA, tA, thread_idx); // (THR_M,THR_K,PIPE)
|
||||
|
||||
Tensor tBgB = local_partition(gB, tB, thread_idx); // (THR_N,THR_K,k)
|
||||
Tensor tBsB = local_partition(sB, tB, thread_idx); // (THR_N,THR_K,PIPE)
|
||||
```
|
||||
|
||||
`tAgA` and `tBgB` partition the global A resp. B matrices over threads,
|
||||
and `tAsA` and `tBsB` partition the shared memory tiles of A resp. B over threads.
|
||||
|
||||
The following code creates predicate tensors
|
||||
corresponding to `tAgA` and `tBgB`.
|
||||
They will be computed once in the prologue.
|
||||
and will be used to mask off instructions in the inner loop.
|
||||
|
||||
```c++
|
||||
Tensor tApA = make_tensor<bool>(make_shape (size<0>(tAgA), size<1>(tAgA)),
|
||||
make_stride( Int<1>{}, Int<0>{}));
|
||||
Tensor tBpB = make_tensor<bool>(make_shape (size<0>(tBgB), size<1>(tBgB)),
|
||||
make_stride( Int<1>{}, Int<0>{}));
|
||||
```
|
||||
|
||||
We're only thread-parallelizing over the leftmost (row) dimension,
|
||||
so we only need to predicate over the leftmost dimension.
|
||||
Thus, we can make the rightmost (column) stride zero,
|
||||
since we will never actually address the rightmost dimension.
|
||||
|
||||
The following code creates "two-dimensional identity tensors"
|
||||
that map coordinates (m,k) -> (m,k)
|
||||
for the tile of data within the thread block.
|
||||
|
||||
```c++
|
||||
Tensor cA = make_identity_tensor(make_shape(size<0>(sA), size<1>(sA))); // (BLK_M,BLK_K) -> (blk_m,blk_k)
|
||||
Tensor cB = make_identity_tensor(make_shape(size<0>(sB), size<1>(sB))); // (BLK_N,BLK_K) -> (blk_n,blk_k)
|
||||
```
|
||||
|
||||
The following lines then tile and partition
|
||||
the two reference tensors
|
||||
in exactly the same way the data were tiled and partitioned
|
||||
into `tAsA` and `tBsB`.
|
||||
|
||||
```c++
|
||||
Tensor tAcA = local_partition(cA, tA, thread_idx);
|
||||
Tensor tBcB = local_partition(cB, tB, thread_idx);
|
||||
```
|
||||
|
||||
Tiling and partitioning affect the offset and domain,
|
||||
but not the codomain of the tensors,
|
||||
so we're left with tensors that map `(thr_m,thr_k) -> (m,k)`
|
||||
where `(thr_m,thr_k)` is this particular thread's subtensor of the tile
|
||||
and `(m,k)` is the original codomain: a coordinate into the original tile.
|
||||
|
||||
The unrolled loops in the code below then compare
|
||||
the m- and n-coordinates of those tensors with our known maximums
|
||||
to mask off elements we are not allowed to access.
|
||||
|
||||
```c++
|
||||
Tensor cA = make_identity_tensor(make_shape(size<0>(sA), size<1>(sA))); // (BLK_M,BLK_K) -> (blk_m,blk_k)
|
||||
Tensor tAcA = local_partition(cA, tA, thread_idx);
|
||||
|
||||
Tensor cB = make_identity_tensor(make_shape(size<0>(sB), size<1>(sB))); // (BLK_N,BLK_K) -> (blk_n,blk_k)
|
||||
Tensor tBcB = local_partition(cB, tB, thread_idx);
|
||||
|
||||
// Populate
|
||||
CUTE_UNROLL
|
||||
for (int m = 0; m < size<0>(tApA); ++m) {
|
||||
tApA(m,0) = get<0>(tAcA(m,0)) < m_max_coord;
|
||||
}
|
||||
CUTE_UNROLL
|
||||
for (int n = 0; n < size<0>(tBpB); ++n) {
|
||||
tBpB(n,0) = get<0>(tBcB(n,0)) < n_max_coord;
|
||||
}
|
||||
```
|
||||
|
||||
Those last `for` loops fill in the two predicate tensors.
|
||||
In this case, we only need to predicate over the leftmost dimension,
|
||||
so we only address `(m,0)` resp. `(n,0)`.
|
||||
|
||||
We can then use the predicate tensors in `copy_if`
|
||||
to copy only the elements for which the corresponding
|
||||
predicate tensor elements are nonzero.
|
||||
|
||||
```c++
|
||||
// Prefetch k_tile=0, gate these on k_residue as well
|
||||
CUTE_UNROLL
|
||||
for (int k = 0; k < size<1>(tAsA); ++k) {
|
||||
if (get<1>(tAcA(0,k)) >= -k_residue) { // some other condition on the column index
|
||||
copy_if(tApA, tAgA(_,k,0), tAsA(_,k,0));
|
||||
}
|
||||
}
|
||||
|
||||
CUTE_UNROLL
|
||||
for (int k = 0; k < size<1>(tBsB); ++k) {
|
||||
if (get<1>(tBcB(0,k)) >= -k_residue) { // some other condition on the column index
|
||||
copy_if(tBpB, tBgB(_,k,0), tBsB(_,k,0));
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
Here are some advantages of this "reference tensor" approach.
|
||||
|
||||
1. It doesn't depend on the layout/strides of the tensor
|
||||
being predicated, just the logical bounds being imposed.
|
||||
|
||||
2. The partitioning stage can be anything.
|
||||
|
||||
3. It naturally extends to any-dimensional predication.
|
||||
|
||||
4. It's a natural generalization of a typical CUDA 1-D
|
||||
parallel vector access pattern,
|
||||
which computes an access index `k`
|
||||
(e.g., as `blockDim.x * blockIdx.x + threadIdx.x`)
|
||||
and then predicates access to the vector's `k`-th element
|
||||
on whether `k` is in bounds.
|
||||
|
||||
As an example of (3), the epilogue predication does exactly the same thing,
|
||||
|
||||
```c++
|
||||
// Repeat with a tensor of coordinates for predication
|
||||
Tensor cC = make_identity_tensor(make_shape(size<0>(gC), size<1>(gC)));
|
||||
Tensor tCcC = thr_mma.partition_C(cC);
|
||||
|
||||
const bool isBetaZero = (beta == 0);
|
||||
|
||||
CUTE_UNROLL
|
||||
for (int i = 0; i < size(tCrC); ++i) {
|
||||
if (elem_less(tCcC(i), make_coord(m_max_coord,n_max_coord))) {
|
||||
tCgC(i) = isBetaZero ? alpha * tCrC(i) : alpha * tCrC(i) + beta * tCgC(i);
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
but with the mma responsible for the tiling/partitioning `tCcC`
|
||||
so that the reference subtensor matches the accumulator's subtensor.
|
||||
Then, the reference subtensor is predicated against the `if` bounds
|
||||
(in both m- and n-coordinates) inside the `for` loop.
|
||||
|
||||
Another way to explain this is that we don't modify the tiles
|
||||
to give you the "right" extents so that you never overrun.
|
||||
Instead, we let you query the original coordinate
|
||||
to see if that coordinate overruns.
|
||||
This avoids all branching and variable/dynamic loop bounds
|
||||
(thus maintaining load balance and synchronicity,
|
||||
both very important in-kernel) in favor of predication.
|
||||
It's also general enough to extend to all ranks,
|
||||
all layouts of threads and data,
|
||||
and all tiling/partitioning patterns.
|
||||
|
||||
## Copyright
|
||||
|
||||
Copyright (c) 2017 - 2025 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
```
|
||||
Redistribution and use in source and binary forms, with or without
|
||||
modification, are permitted provided that the following conditions are met:
|
||||
|
||||
1. Redistributions of source code must retain the above copyright notice, this
|
||||
list of conditions and the following disclaimer.
|
||||
|
||||
2. Redistributions in binary form must reproduce the above copyright notice,
|
||||
this list of conditions and the following disclaimer in the documentation
|
||||
and/or other materials provided with the distribution.
|
||||
|
||||
3. Neither the name of the copyright holder nor the names of its
|
||||
contributors may be used to endorse or promote products derived from
|
||||
this software without specific prior written permission.
|
||||
|
||||
THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS "AS IS"
|
||||
AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE
|
||||
IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE ARE
|
||||
DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT HOLDER OR CONTRIBUTORS BE LIABLE
|
||||
FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL
|
||||
DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR
|
||||
SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER
|
||||
CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY,
|
||||
OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE
|
||||
OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
|
||||
```
|
||||
263
media/docs/cpp/cute/0z_tma_tensors.md
Normal file
263
media/docs/cpp/cute/0z_tma_tensors.md
Normal file
@@ -0,0 +1,263 @@
|
||||
# CuTe TMA Tensors
|
||||
|
||||
Along your travels, you may find strange looking CuTe Tensors that are printed as something like
|
||||
```
|
||||
ArithTuple(0,_0,_0,_0) o ((_128,_64),2,3,1):((_1@0,_1@1),_64@1,_1@2,_1@3)
|
||||
```
|
||||
What is an `ArithTuple`? Are those tensor strides? What do those mean? What is this for?
|
||||
|
||||
This documentation intends to answer those questions and introduce some of the more advanced features of CuTe.
|
||||
|
||||
# Introduction to TMA instructions
|
||||
|
||||
The Tensor Memory Accelerator (TMA) is a set of instructions for copying possibly multidimensional arrays between global and shared memory. TMA was introduced in the Hopper architecture. A single TMA instruction can copy an entire tile of data all at once. As a result, the hardware no longer needs to compute individual memory addresses and issue a separate copy instruction for each element of the tile.
|
||||
|
||||
To accomplish this, the TMA instruction is given a *TMA descriptor*, which is a packed representation of a multidimensional tensor in global memory with 1, 2, 3, 4, or 5 dimensions. The TMA descriptor holds
|
||||
|
||||
* the base pointer of the tensor;
|
||||
|
||||
* the data type of the tensor's elements (e.g., `int`, `float`, `double`, or `half`);
|
||||
|
||||
* the size of each dimension;
|
||||
|
||||
* the stride within each dimension; and
|
||||
|
||||
* other flags representing the smem box size, smem swizzling patterns, and out-of-bounds access behavior.
|
||||
|
||||
This descriptor must be created on the host before kernel execution.
|
||||
It is shared between all thread blocks that will be issuing TMA instructions.
|
||||
Once inside the kernel, the TMA is executed with the following parameters:
|
||||
|
||||
* pointer to the TMA descriptor;
|
||||
|
||||
* pointer to the SMEM; and
|
||||
|
||||
* coordinates into the GMEM tensor represented within the TMA descriptor.
|
||||
|
||||
For example, the interface for TMA-store with 3-D coordinates looks like this.
|
||||
|
||||
```cpp
|
||||
struct SM90_TMA_STORE_3D {
|
||||
CUTE_DEVICE static void
|
||||
copy(void const* const desc_ptr,
|
||||
void const* const smem_ptr,
|
||||
int32_t const& crd0, int32_t const& crd1, int32_t const& crd2) {
|
||||
// ... invoke CUDA PTX instruction ...
|
||||
}
|
||||
};
|
||||
```
|
||||
|
||||
We observe that the TMA instruction does not directly consume pointers to global memory. Indeed, the global memory pointer is contained in the descriptor, is considered constant, and is NOT a separate parameter to the TMA instruction. Instead, the TMA consumes TMA coordinates into the TMA's view of global memory that is defined in the TMA descriptor.
|
||||
|
||||
That means that an ordinary CuTe Tensor that stores a GMEM pointer and computes offsets and new GMEM pointers is useless to the TMA.
|
||||
|
||||
What do we do?
|
||||
|
||||
# Building a TMA Tensor
|
||||
|
||||
## Implicit CuTe Tensors
|
||||
|
||||
All CuTe Tensors are compositions of Layouts and Iterators. An ordinary global memory tensor's iterator is its global memory pointer. However, a CuTe Tensor's iterator doesn't have to be a pointer; it can be any random-access iterator.
|
||||
|
||||
One example of such an iterator is a *counting iterator*.
|
||||
This represents a possibly infinite sequence of integers that starts at some value.
|
||||
We call the members of this sequence *implicit integers*,
|
||||
because the sequence is not explicitly stored in memory.
|
||||
The iterator just stores its current value.
|
||||
|
||||
We can use a counting iterator to create a tensor of implicit integers,
|
||||
```cpp
|
||||
Tensor A = make_tensor(counting_iterator<int>(42), make_shape(4,5));
|
||||
print_tensor(A);
|
||||
```
|
||||
which outputs
|
||||
```
|
||||
counting_iter(42) o (4,5):(_1,4):
|
||||
42 46 50 54 58
|
||||
43 47 51 55 59
|
||||
44 48 52 56 60
|
||||
45 49 53 57 61
|
||||
```
|
||||
This tensor maps logical coordinates to on-the-fly computed integers. Because it's still a CuTe Tensor, it can still be tiled and partitioned and sliced just like a normal tensor by accumulating integer offsets into the iterator.
|
||||
|
||||
But the TMA doesn't consume pointers or integers, it consumes coordinates. Can we make a tensor of implicit TMA
|
||||
coordinates for the TMA instruction to consume? If so, then we could presumably also tile and partition and slice that tensor of coordinates so that we would always have the right TMA coordinate to give to the instruction.
|
||||
|
||||
## ArithTupleIterators and ArithTuples
|
||||
|
||||
First, we build a `counting_iterator` equivalent for TMA coordinates. It should support
|
||||
|
||||
* dereference to a TMA coordinate, and
|
||||
|
||||
* offset by another TMA coordinate.
|
||||
|
||||
We'll call this an `ArithmeticTupleIterator`. It stores a coordinate (a tuple of integers) that is represented as an `ArithmeticTuple`. The `ArithmeticTuple` is simply a (public subclass of) `cute::tuple` that has an overloaded `operator+` so that it can be offset by another tuple. The sum of two tuples is the tuple of the sum of the elements.
|
||||
|
||||
Now similar to `counting_iterator<int>(42)` we can create an implicit "iterator" (but without increment or other common iterator operations) over tuples that can be dereferenced and offset by other tuples
|
||||
```cpp
|
||||
ArithmeticTupleIterator citer_1 = make_inttuple_iter(42, Int<2>{}, Int<7>{});
|
||||
ArithmeticTupleIterator citer_2 = citer_1 + make_tuple(Int<0>{}, 5, Int<2>{});
|
||||
print(*citer_2);
|
||||
```
|
||||
which outputs
|
||||
```
|
||||
(42,7,_9)
|
||||
```
|
||||
|
||||
A TMA Tensor can use an iterator like this to store the current TMA coordinate "offset". The "offset" here is in quotes because it's clearly not a normal 1-D array offset or pointer.
|
||||
|
||||
In summary, one creates a TMA descriptor for the *whole global memory tensor*. The TMA descriptor defines a view into that tensor and the instruction takes TMA coordinates into that view. In order to generate and track those TMA coordinates, we define an implicit CuTe Tensor of TMA coordinates that can be tiled, sliced, and partitioned the exact same way as an ordinary CuTe Tensor.
|
||||
|
||||
We can now track and offset TMA coordinates with this iterator, but how do we get CuTe Layouts to generate non-integer offsets?
|
||||
|
||||
## Strides aren't just integers
|
||||
|
||||
Ordinary tensors have a layout that maps
|
||||
a logical coordinate `(i,j)` into a 1-D linear index `k`.
|
||||
This mapping is the inner-product of the coordinate with the strides.
|
||||
|
||||
TMA Tensors hold iterators of TMA coordinates.
|
||||
Thus, a TMA Tensor's Layout must map a logical coordinate
|
||||
to a TMA coordinate, rather than to a 1-D linear index.
|
||||
|
||||
To do this, we can abstract what a stride is. Strides need not be integers, but rather any algebraic object that supports inner-product with the integers (the logical coordinate). The obvious choice is the `ArithmeticTuple` we used earlier since they can be added to each other, but this time additionally equipped with an `operator*` so it can also be scaled by an integer.
|
||||
|
||||
### Aside: Integer-module strides
|
||||
|
||||
A group of objects that support addition between elements and product between elements and integers is called an integer-module.
|
||||
|
||||
Formally, an integer-module is an abelian group `(M,+)` equipped with `Z*M -> M`, where `Z` are the integers. That is, an integer-module `M` is
|
||||
a group that supports inner products with the integers.
|
||||
The integers are an integer-module.
|
||||
Rank-R tuples of integers are an integer-module.
|
||||
|
||||
In principle, layout strides may be any integer-module.
|
||||
|
||||
### Basis elements
|
||||
|
||||
CuTe's basis elements live in the header file `cute/numeric/arithmetic_tuple.hpp`.
|
||||
To make it easy to create `ArithmeticTuple`s that can be used as strides, CuTe defines normalized basis elements using the `E` type alias. "Normalized" means that the scaling factor of the basis element is the compile-time integer 1.
|
||||
|
||||
| C++ object | Description | String representation |
|
||||
| --- | --- | --- |
|
||||
| `E<>{}` | `1` | `1` |
|
||||
| `E<0>{}` | `(1,0,...)` | `1@0` |
|
||||
| `E<1>{}` | `(0,1,0,...)` | `1@1` |
|
||||
| `E<0,1>{}` | `((0,1,0,...),0,...)` | `1@1@0` |
|
||||
| `E<1,0>{}` | `(0,(1,0,...),0,...)` | `1@0@1` |
|
||||
|
||||
The "description" column in the above table
|
||||
interprets each basis element as an infinite tuple of integers,
|
||||
where all the tuple's entries not specified by the element's type are zero.
|
||||
We count tuple entries from left to right, starting with zero.
|
||||
For example, `E<1>{}` has a 1 in position 1: `(0,1,0,...)`.
|
||||
`E<3>{}` has a 1 in position 3: `(0,0,0,1,0,...)`.
|
||||
|
||||
Basis elements can be *nested*.
|
||||
For instance, in the above table, `E<0,1>{}` means that
|
||||
in position 0 there is a `E<1>{}`: `((0,1,0,...),0,...)`.
|
||||
|
||||
Basis elements can be *scaled*.
|
||||
That is, they can be multiplied by an integer *scaling factor*.
|
||||
For example, in `5*E<1>{}`, the scaling factor is `5`.
|
||||
`5*E<1>{}` prints as `5@1` and means `(0,5,0,...)`.
|
||||
The scaling factor commutes through any nesting.
|
||||
For instance, `5*E<0,1>{}` prints as `5@1@0`
|
||||
and means `((0,5,0,...),0,...)`.
|
||||
|
||||
Basis elements can also be added together,
|
||||
as long as their hierarchical structures are compatible.
|
||||
For example, `3*E<0>{} + 4*E<1>{}` results in `(3,4,0,...)`.
|
||||
Intuitively, "compatible" means that
|
||||
the nested structure of the two basis elements
|
||||
matches well enough to add the two elements together.
|
||||
|
||||
### Linear combinations of strides
|
||||
|
||||
Layouts work by taking the inner product
|
||||
of the natural coordinate with their strides.
|
||||
For strides made of integer elements, e.g., `(1,100)`,
|
||||
the inner product of the input coordinate `(i,j)`
|
||||
and the stride is `i + 100j`.
|
||||
Offsetting an "ordinary" tensor's pointer and this index
|
||||
gives the pointer to the tensor element at `(i,j)`.
|
||||
|
||||
For strides of basis elements, we still compute the inner product of the natural coordinate with the strides.
|
||||
For example, if the stride is `(1@0,1@1)`,
|
||||
then the inner product of the input coordinate `(i,j)`
|
||||
with the strides is `i@0 + j@1 = (i,j)`.
|
||||
That translates into the (TMA) coordinate `(i,j)`.
|
||||
If we wanted to reverse the coordinates,
|
||||
then we could use `(1@1,1@0)` as the stride.
|
||||
Evaluating the layout would give `i@1 + j@0 = (j,i)`.
|
||||
|
||||
A linear combination of basis elements
|
||||
can be interpreted as a possibly multidimensional and hierarchical coordinate.
|
||||
For instance, `2*2@1@0 + 3*1@1 + 4*5@1 + 7*1@0@0`
|
||||
means `((0,4,...),0,...) + (0,3,0,...) + (0,20,0,...) + ((7,...),...) = ((7,4,...),23,...)`
|
||||
and can be interpreted as the coordinate `((7,4),23)`.
|
||||
|
||||
Thus, linear combinations of these strides can be used to generate TMA coordinates.
|
||||
These coordinates, in turn, can be used to offset TMA coordinate iterators.
|
||||
|
||||
## Application to TMA Tensors
|
||||
|
||||
Now we can build CuTe Tensors like the one seen in the introduction.
|
||||
|
||||
```cpp
|
||||
Tensor a = make_tensor(make_inttuple_iter(0,0),
|
||||
make_shape ( 4, 5),
|
||||
make_stride(E<0>{}, E<1>{}));
|
||||
print_tensor(a);
|
||||
|
||||
Tensor b = make_tensor(make_inttuple_iter(0,0),
|
||||
make_shape ( 4, 5),
|
||||
make_stride(E<1>{}, E<0>{}));
|
||||
print_tensor(b);
|
||||
```
|
||||
prints
|
||||
```
|
||||
ArithTuple(0,0) o (4,5):(_1@0,_1@1):
|
||||
(0,0) (0,1) (0,2) (0,3) (0,4)
|
||||
(1,0) (1,1) (1,2) (1,3) (1,4)
|
||||
(2,0) (2,1) (2,2) (2,3) (2,4)
|
||||
(3,0) (3,1) (3,2) (3,3) (3,4)
|
||||
|
||||
ArithTuple(0,0) o (4,5):(_1@1,_1@0):
|
||||
(0,0) (1,0) (2,0) (3,0) (4,0)
|
||||
(0,1) (1,1) (2,1) (3,1) (4,1)
|
||||
(0,2) (1,2) (2,2) (3,2) (4,2)
|
||||
(0,3) (1,3) (2,3) (3,3) (4,3)
|
||||
```
|
||||
|
||||
## Copyright
|
||||
|
||||
Copyright (c) 2017 - 2025 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
```
|
||||
Redistribution and use in source and binary forms, with or without
|
||||
modification, are permitted provided that the following conditions are met:
|
||||
|
||||
1. Redistributions of source code must retain the above copyright notice, this
|
||||
list of conditions and the following disclaimer.
|
||||
|
||||
2. Redistributions in binary form must reproduce the above copyright notice,
|
||||
this list of conditions and the following disclaimer in the documentation
|
||||
and/or other materials provided with the distribution.
|
||||
|
||||
3. Neither the name of the copyright holder nor the names of its
|
||||
contributors may be used to endorse or promote products derived from
|
||||
this software without specific prior written permission.
|
||||
|
||||
THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS "AS IS"
|
||||
AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE
|
||||
IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE ARE
|
||||
DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT HOLDER OR CONTRIBUTORS BE LIABLE
|
||||
FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL
|
||||
DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR
|
||||
SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER
|
||||
CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY,
|
||||
OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE
|
||||
OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
|
||||
```
|
||||
17
media/docs/cpp/cute/index.rst
Normal file
17
media/docs/cpp/cute/index.rst
Normal file
@@ -0,0 +1,17 @@
|
||||
.. _cpp_cute:
|
||||
|
||||
CuTe
|
||||
====================
|
||||
|
||||
.. toctree::
|
||||
:maxdepth: 2
|
||||
|
||||
00_quickstart<00_quickstart.md>
|
||||
01_layout<01_layout.md>
|
||||
02_layout_algebra<02_layout_algebra.md>
|
||||
03_tensor<03_tensor.md>
|
||||
04_algorithms<04_algorithms.md>
|
||||
0t_mma_atom<0t_mma_atom.md>
|
||||
0x_gemm_tutorial<0x_gemm_tutorial.md>
|
||||
0y_predication<0y_predication.md>
|
||||
0z_tma_tensors<0z_tma_tensors.md>
|
||||
471
media/docs/cpp/cutlass_3x_backwards_compatibility.md
Normal file
471
media/docs/cpp/cutlass_3x_backwards_compatibility.md
Normal file
@@ -0,0 +1,471 @@
|
||||
# CUTLASS 3.0 GEMM Backwards Compatibility
|
||||
|
||||
Although CUTLASS 3.0 restructures the GEMM hierarchy and introduces new types for the
|
||||
threadblock layer and below, we intend the entire source code to be usable in user applications.
|
||||
We expect users to be able to `#include` any source file from CUTLASS 3.0, whether
|
||||
they implement the 2.x or the 3.x API, without breaking user builds. This means that a single
|
||||
translation unit should be able to contain any valid kernel regardless of its API version. The
|
||||
sections below discuss how `device` and `kernel` layer type names are made compatible across the
|
||||
two API versions, and what the users can expect out of the `threadblock` layer API going forward.
|
||||
|
||||
## Compatible Device API
|
||||
|
||||
The entry point for CUTLASS's Device GEMM API
|
||||
is the class
|
||||
`cutlass::gemm::device::GemmUniversalAdapter`.
|
||||
This class lives in the header file
|
||||
[include/cutlass/gemm/device/gemm_universal_adapter.h](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/gemm/device/gemm_universal_adapter.h).
|
||||
|
||||
`GemmUniversalAdapter` is a "universal adapter"
|
||||
and serves as a common device interface
|
||||
for both CUTLASS 3.x and CUTLASS 2.x kernels.
|
||||
Its template parameter `GemmKernel`,
|
||||
the GEMM kernel type, can be any of the following:
|
||||
|
||||
* `cutlass::gemm::kernel::GemmUniversal`,
|
||||
implementing CUTLASS 3.x API kernels;
|
||||
* `cutlass::gemm::kernel::GemmUniversal`,
|
||||
implementing CUTLASS 2.x API kernels;
|
||||
* Any valid CUTLASS 2.x `kernel` layer GEMM that
|
||||
was previously composable with `device::GemmUniversalAdapter`
|
||||
|
||||
Users implementing new kernels in either API should prefer
|
||||
using `kernel::GemmUniversal` as the kernel type
|
||||
and compose it with `device::GemmUniversalAdapter`.
|
||||
Users with existing `kernel::Gemm` kernels
|
||||
can continue to use them as template arguments
|
||||
of `device::GemmUniversalAdapter`. They can adopt
|
||||
`GemmUniversal` as a gradual migration path,
|
||||
since `GemmUniversal` accepts either 3.0 or 2.x collectives.
|
||||
Please see the [next section for `kernel::GemmUniversal`](#compatible-kernel-api) for details.
|
||||
|
||||
`GemmUniversalAdapter` presents a single
|
||||
host-side interface to both 3.0 and 2.x kernels.
|
||||
CUTLASS accomplishes this by
|
||||
specializing `GemmUniversalAdapter`'s implementation
|
||||
on either 2.x API implementing kernel layer GEMMs, or 3.x API
|
||||
implementing kernel layer GEMMs (as detected by `gemm::detail::IsCutlass3GemmKernel`
|
||||
discussed below). As a result, `GemmUniversalAdapter`'s behavior
|
||||
might differ between the two specializations.
|
||||
|
||||
### Device API design differences
|
||||
|
||||
In CUTLASS 2.x, the Device API was more closely tied
|
||||
to the Kernel API. In CUTLASS 3.0, the Device API
|
||||
accepts any kernel type that meets the Kernel API
|
||||
interface requirements. CUTLASS 3.0's Device API code is
|
||||
parameterized by the kernel type, but this code
|
||||
is *generic*; the same code works for any kernel type.
|
||||
|
||||
The device layer compatibility interface, `device::GemmUniversalAdapter`,
|
||||
also provides reflective mappings from 3.0-specific types
|
||||
back to the closest possible 2.x equivalent types. This is [discussed further in the section below](#conversions-between-2x-tags-and-30-types).
|
||||
|
||||
CUTLASS 3.0's `device::GemmUniversalAdapter` also exposes some new APIs that the 2.x `device::GemmUniversalAdapter` implementation does not. Most notably, this includes the ability to bypass the `GemmKernel::Arguments` to `GemmKernel::Params` lowering.
|
||||
|
||||
```c++
|
||||
// Primary run() entry point API that is static allowing users to create and manage their own params.
|
||||
static Status
|
||||
run(Params& params, cudaStream_t stream = nullptr);
|
||||
```
|
||||
|
||||
This new API is useful for the following scenarios.
|
||||
|
||||
* Running again does not require reinvoking `GemmKernel::to_underlying_arguments()`
|
||||
* Manual control over construction of `GemmKernel::Params` for custom kernels with custom stride types
|
||||
* Fully static problem shapes and strides for bespoke kernels where no argument mapping needs to take place
|
||||
|
||||
## Compatible Kernel API
|
||||
|
||||
CUTLASS 3.x API shares the kernel layer API with CUTLASS 2.x
|
||||
through the single entry point type `cutlass::gemm::kernel::GemmUniversal`.
|
||||
All kernel layer GEMMs are viewed as a composition of a collective mainloop
|
||||
and a collective epilogue.
|
||||
|
||||
**`kernel::GemmUniversal` implements both 2.x and 3.x APIs**
|
||||
|
||||
The entry point for CUTLASS's kernel API is the class
|
||||
`cutlass::gemm::kernel::GemmUniversal`.
|
||||
This class' declaration lives in the header file
|
||||
[include/cutlass/gemm/kernel/gemm_universal.hpp](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/gemm/kernel/gemm_universal.hpp).
|
||||
|
||||
```c++
|
||||
/*
|
||||
* Stateless universal device GEMM kernel type that treats GEMM as
|
||||
* a composition of a collective mainloop and a collective epilogue.
|
||||
* SFIANE shims both 2.x and 3.0 API kernels based on ProblemShapeOrThreadblockMma_.
|
||||
**/
|
||||
template <
|
||||
class ProblemShapeOrThreadblockMma_,
|
||||
class CollectiveMainloopOrEpilogue_,
|
||||
class CollectiveEpilogueOrThreadblockSwizzle_,
|
||||
class TileScheduler_ = void,
|
||||
class Enable = void
|
||||
>
|
||||
class GemmUniversal;
|
||||
```
|
||||
|
||||
We call this class "universal" because it can be built
|
||||
using either the CUTLASS 3.0 or the 2.x mainloops and epilogues.
|
||||
If `GemmUniversal`'s first template argument
|
||||
(`ProblemShapeOrThreadblockMma_`) is a `cute::tuple`,
|
||||
then `GemmUniversal` assumes that
|
||||
the remaining three template arguments
|
||||
(the mainloop, epilogue, and grid swizzle)
|
||||
implement the 3.0 APIs.
|
||||
Otherwise, `GemmUniversal` assumes that
|
||||
the remaining three template arguments
|
||||
implement the 2.x APIs.
|
||||
All the template arguments must be either
|
||||
CUTLASS 3.0 or CUTLASS 2.x types. For example,
|
||||
`GemmUniversal` does not permit using
|
||||
a 2.x mainloop with a 3.0 collective epilogue.
|
||||
|
||||
CUTLASS 3.x implements various embodiments of `kernel::GemmUniversal`.
|
||||
Each kernel layer schedule is specialized
|
||||
for a GEMM scheduling algorithm and GPU architecture.
|
||||
Specializations of `kernel::GemmUniversal` for 3.0 APIs live in
|
||||
any of various `gemm_*.hpp` files in the directory
|
||||
[include/cutlass/gemm/kernel/](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/gemm/kernel/).
|
||||
The specialization to which to dispatch is decided through the dispatch policy's `Schedule` type.
|
||||
|
||||
Specializations for 2.x APIs live in the header file
|
||||
[include/cutlass/gemm/kernel/gemm_universal.h](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/gemm/kernel/gemm_universal.h).
|
||||
|
||||
### Kernel API design differences
|
||||
|
||||
The CUTLASS 2.x Kernel API was more closely tied
|
||||
to the Device API, as we mentioned above.
|
||||
In particular, the 2.x Device API specified the grid shape
|
||||
used to launch the Kernel API.
|
||||
In CUTLASS 3.0, the Kernel API controls its own grid shape,
|
||||
while the device adapter simply queries the kernel with which it needs to be launched.
|
||||
|
||||
This change is required to support various kernel schedules
|
||||
that may need their own schedule specific grid planning logic.
|
||||
For example, persistent kernel schedules generally only launch with
|
||||
as many threadblocks as the number of multiprocessors on the GPU.
|
||||
|
||||
All CUTLASS 3 `kernel::GemmUniversal` specializations expose the following (static) API:
|
||||
|
||||
```c++
|
||||
// Returns true if the kernel can execute the provided GEMM arguments.
|
||||
static bool
|
||||
can_implement(Arguments const& args);
|
||||
|
||||
// Returns a dim3 representing the threadblock shape.
|
||||
static dim3
|
||||
get_block_shape();
|
||||
|
||||
// Returns a dim3 representing the grid shape in terms of threadblocks.
|
||||
static dim3
|
||||
get_grid_shape(Params const& params);
|
||||
```
|
||||
|
||||
The device adapter simply queries the kernel for these three before launching it on the device.
|
||||
CUTLASS 3.0 provides a meta-function to detect whether a `cutlass::gemm::kernel::*` implements
|
||||
the 3.x API or 2.x API:
|
||||
|
||||
```c++
|
||||
// include/cutlass/gemm/gemm.h
|
||||
|
||||
namespace cutlass:gemm::detail {
|
||||
|
||||
// The following metafunction is used to detect whether a
|
||||
// `kernel::Gemm` or `kernel::GemmUniversal` implements the CUTLASS 3.x API,
|
||||
// by checking whether the problem shape type is aliased within.
|
||||
template <class GemmKernel, class = void>
|
||||
struct IsCutlass3GemmKernel;
|
||||
|
||||
} // namespace cutlass:gemm::detail
|
||||
```
|
||||
|
||||
Users can dispatch their generic code against 2.x and 3.x specializations with
|
||||
this as a type trait for the kernel API version.
|
||||
|
||||
## Threadblock API and Inner Loops
|
||||
|
||||
Much of the CUTLASS 3 GEMM hierarchy for mainloops and inner loops diverges
|
||||
from that of CUTLASS 2.x. With that also comes the introduction of the
|
||||
`cutlass::gemm::collective` layer as a direct replacement and a superset
|
||||
of the 2.x `cutlass::gemm::threadblock` layer. Going forward,
|
||||
CUTLASS 3.x will discontinue new developments in the following namespaces.
|
||||
|
||||
* `cutlass::*::threadblock::*`
|
||||
* `cutlass::*::warp::*`
|
||||
* `cutlass::gemm::thread::*`
|
||||
* `cutlass::arch::*` (except `barrier.h`)
|
||||
|
||||
`cutlass::gemm::collective`s are a superset of the threadblock layer where
|
||||
all new mainloops will be developed. Users should look to the `CollectiveMma` type
|
||||
if they wish to author custom mainloop code in the 3.x API.
|
||||
|
||||
Similarly, for the GEMM inner loops, `cute::MMA_Atom`s replace the
|
||||
`gemm::warp` and `gemm::thread` layer code. Going forward, all new PTX instructions
|
||||
and associated metadata development will occur directly inside [`cute/arch/*.hpp`](https://github.com/NVIDIA/cutlass/tree/main/include/cute/arch/) and [`cute/atom/*.hpp`](https://github.com/NVIDIA/cutlass/tree/main/include/cute/atom/).
|
||||
|
||||
The desired inner loop MMA iteration order and tiling can be achieved through careful
|
||||
selection of the atom layout, value layout, and permutations of the `cute::TiledMma`.
|
||||
|
||||
For epilogues, the `cutlass::epilogue::collective` layer replaces `cutlass::threadblock::collective`. However, the thread-level epilogue elementwise operations
|
||||
in `cutlass::epilogue::thread` will continue to be used in 3.x kernels as well, albeit, with
|
||||
a more idiomatic epilogue vectorization strategy.
|
||||
[Example 50](https://github.com/NVIDIA/cutlass/tree/main/examples/50_hopper_gemm_with_epilogue_swizzle/50_hopper_gemm_with_epilogue_swizzle.cu)
|
||||
shows how to use 2.x epilogue thread operators with 3.0 API kernels.
|
||||
|
||||
## Porting from 2.x to 3.0 API
|
||||
|
||||
### CUTLASS 2.x layout tags and CUTLASS 3.0 major modes
|
||||
|
||||
CUTLASS 2.x and CUTLASS 3.0 use both
|
||||
different wording and different types
|
||||
to describe the permitted layouts
|
||||
of GEMM's input matrices A and B.
|
||||
|
||||
CUTLASS 3.0 does not use the terms "column major"
|
||||
or "row major" to describe matrix layouts.
|
||||
Starting with CUTLASS 3.0, adoption of CuTe allows us to decouple
|
||||
|
||||
* the coordinate mode order (logical shape) of layouts from
|
||||
|
||||
* the index space stride order of the backing storage.
|
||||
|
||||
In line with our switch to a conceptual GEMM hierarchy, we view the major modes not from a BLAS-3 perspective.
|
||||
Rather, we divide the modes into two categories.
|
||||
|
||||
* "Inner modes" or "K-modes" are contracted over during the GEMM.
|
||||
Therefore, they are not present in the output tensor.
|
||||
|
||||
* "Outer modes" or "MN-modes" are preserved in the output.
|
||||
|
||||
Now, instead of `RowMajor` or `ColumnMajor`, whose major stride depends on whether we are referring to the
|
||||
A or the B matrix, we uniformly employ the "K major" or "MN major" terminology and enforce the convention of all tensors having the shape `[M/N, K, L]` regardless of which mode is major. That is,
|
||||
|
||||
* the input matrix A has shape M x K,
|
||||
* the input matrix B has shape N x K, and
|
||||
* the input/output matrices C/D have shape M x N.
|
||||
|
||||
Note that this convention for B
|
||||
differs from the BLAS's GEMM interface,
|
||||
which specifies that B has shape K x N.
|
||||
|
||||
CUTLASS 3.0 uses these names of the modes
|
||||
to specify which mode of a matrix has stride 1.
|
||||
For the matrix A,
|
||||
|
||||
* "M major" means that the matrix is stride 1
|
||||
in the M mode, and
|
||||
* "K major" means that the matrix is stride 1
|
||||
in the K mode.
|
||||
|
||||
For the matrix B,
|
||||
|
||||
* "N major" means that the matrix is stride 1
|
||||
in the N mode (which for B is mode 0,
|
||||
because the convention is that B is N x K); and
|
||||
* "K major" means that the matrix is stride 1
|
||||
in the K mode (which for B is mode 1).
|
||||
|
||||
CUTLASS 2.x defines "layout tag" classes
|
||||
`cutlass::layout::ColumnMajor` and `cutlass::layout::RowMajor`,
|
||||
that live in the header file
|
||||
[`cutlass/layout/matrix.h`](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/layout/matrix.h).
|
||||
The interpretation of these layouts in GEMM
|
||||
depends on whether they are applied
|
||||
to the input matrix A or B. For the matrix A, "column major" means
|
||||
that mode corresponding to M extent has stride 1,
|
||||
and "row major" means that mode corresponding to K extent has stride 1.
|
||||
This is the usual computer science definition
|
||||
of column major and row major for a rank-2 array.
|
||||
For the matrix B, the opposite holds:
|
||||
"column major" means that mode corresponding to N extent has stride 1,
|
||||
and "row major" means that mode corresponding to K extent has stride 1.
|
||||
|
||||
Using the convention of `[outer, inner, batch]` mode order for tensor logical shapes
|
||||
avoids potential confusion with the meaning of column major and row major
|
||||
changing depending on whether they are applied to A or B.
|
||||
|
||||
The table below summarizes our mode order convention and
|
||||
mapping of 2.x layout tags to corresponding M-major, N-major, or K-major strides.
|
||||
|
||||
| Matrix | CUTLASS 2.x layout | 2.x Shape | Logical major mode| 3.x Shape/Stride | Major ordinal |
|
||||
| --- | --- | --- | --- | --- | --- |
|
||||
| A | `ColumnMajor` | M x K | M major | M x K x L | 0 (outer) |
|
||||
| A | `RowMajor` | M x K | K major | M x K x L | 1 (inner) |
|
||||
| B | `RowMajor` | K x N | N major | N x K x L | 0 (outer) |
|
||||
| B | `ColumnMajor` | K x N | K major | N x K x L | 1 (inner) |
|
||||
| C | `ColumnMajor` | M x N | M major | M x N x L | 0 (outer) |
|
||||
| C | `RowMajor` | M x N | N major | M x N x L | 1 (inner) |
|
||||
|
||||
Notice that in CUTLASS 3.0, interpretation of layouts no longer changes based on
|
||||
whether we are talking about the A or B matrix. M and N major inputs always have a
|
||||
static size-1 stride in their 0th (outer) mode. Similarly, K major inputs
|
||||
always contain the static size-1 stride in their 1st mode. This uniformity in stride order
|
||||
allows us to represent tensor layouts much more cleanly and treat both A and B equally in our interfaces.
|
||||
See for example the following snippet from our [`kernel/sm70_gemm.hpp`](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/gemm/kernel/sm70_gemm.hpp)
|
||||
for Ampere kernel schedules.
|
||||
|
||||
```c++
|
||||
// Represent the full tensors
|
||||
Tensor mA_mkl = make_tensor(make_gmem_ptr(params.mainloop.ptr_A), make_shape(M,K,L), params.mainloop.dA); // (m,k,l)
|
||||
Tensor mB_nkl = make_tensor(make_gmem_ptr(params.mainloop.ptr_B), make_shape(N,K,L), params.mainloop.dB); // (n,k,l)
|
||||
|
||||
// Get batch slice
|
||||
Tensor mA_mk = mA_mkl(_,_,get<3>(blk_coord_mnkl)); // (m,k)
|
||||
Tensor mB_nk = mB_nkl(_,_,get<3>(blk_coord_mnkl)); // (n,k)
|
||||
|
||||
// Slice to get the tiles for which this thread block is responsible
|
||||
Tensor gA = local_tile(mA_mk, blk_shape, take<0,3>(blk_coord_mnkl), Step<_1, X,_1>{}); // (BLK_M,BLK_K,k)
|
||||
Tensor gB = local_tile(mB_nk, blk_shape, take<0,3>(blk_coord_mnkl), Step< X,_1,_1>{}); // (BLK_N,BLK_K,k)
|
||||
```
|
||||
|
||||
As seem in this snippet, all input tensors have the logical shape `[outer, inner, batch]`,
|
||||
and the strides could represent either outer or inner
|
||||
(or any other complex hierarchical stride) major storage.
|
||||
CuTe layouts always maintain the logical consistency of the coordinate spaces regardless of the strides.
|
||||
|
||||
By convention, in CUTLASS 3.0, we treat the M and N mode as the 0th mode,
|
||||
and K mode as the 1st mode of the stride.
|
||||
|
||||
### Conversions between 2.x tags and 3.0 types
|
||||
|
||||
Starting with CUTLASS 3.0, all layouts are described using
|
||||
`cute::Shape` and `cute::Stride` which compose into a `cute::Layout<Shape, Stride>`.
|
||||
In CUTLASS 2.x, various layout tags such as `cutlass::layout::RowMajor` are used to specialize
|
||||
template implementations. These tag types only encode information about the tensor strides,
|
||||
as 2.x layouts did not incorporate any concept of tensor shape in the layout tags themselves.
|
||||
Users may find a need to convert between CUTLASS 2.x layout tags, and 3.0
|
||||
CuTe stride types. CUTLASS 3.0 `gemm::collective::CollectiveBuilder` interfaces
|
||||
also accept these 2.x layout tags as input parameters in their template API as a convenience for users.
|
||||
At every entry point into CUTLASS 3.0, these tags get converted to their corresponding CuTe Stride type with
|
||||
metafunctions that best approximate their corresponding `cute::Stride`.
|
||||
|
||||
* `cutlass::gemm::detail::TagToStrideA_t<LayoutTag>`
|
||||
* `cutlass::gemm::detail::TagToStrideB_t<LayoutTag>`
|
||||
* `cutlass::gemm::detail::TagToStrideC_t<LayoutTag>`
|
||||
|
||||
By convention, and to match user expectations, the `cute::Stride` types that these
|
||||
map onto always contain one static mode corresponding to the layout tag, and two 64-bit
|
||||
dynamic stride modes corresponding to the minor mode and the batch mode. Batch
|
||||
mode is included by default as all CUTLASS 3.0 kernels support packed batch-mode GEMMs
|
||||
out of the box.
|
||||
|
||||
The [`cutlass/gemm/gemm.h#440`](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/gemm/gemm.h#440)
|
||||
header file includes functions
|
||||
that can be useful for converting
|
||||
from CUTLASS 3.0 `cute::Stride`s back to CUTLASS 2.x layout tags.
|
||||
|
||||
* `cutlass::gemm::detail::StrideToLayoutTagA_t<CuteStride>`
|
||||
* `cutlass::gemm::detail::StrideToLayoutTagB_t<CuteStride>`
|
||||
* `cutlass::gemm::detail::StrideToLayoutTagC_t<CuteStride>`
|
||||
|
||||
These metafunctions take the CuTe Stride as a template parameter and
|
||||
attempt to find the size-1 stride in the idiomatic M, N, or K modes
|
||||
to best approximate a corresponding 2.x layout tag type.
|
||||
Note that this may not work in general for any `cute::Stride`
|
||||
as the mapping between the stride and tag type is not bijective.
|
||||
|
||||
These mapping utilities are kept in a `detail` namespace
|
||||
as we do not guarantee stability of their implementation.
|
||||
Their behavior may change in future releases as we add new features.
|
||||
However, we do expect these type names to remain stable. For users who want
|
||||
these 2.x reflective types from an assembled kernel with a more stable API,
|
||||
the specialization of `cutlass::gemm::device::GemmUniversalAdapter`
|
||||
for CUTLASS 3.0 kernel provides all aliases for all 2.x type aliases
|
||||
in addition to the layout tags. You can see how they are used in the header file
|
||||
[`cutlass/gemm/device/gemm_universal_adapter.h`](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/gemm/device/gemm_universal_adapter.h).
|
||||
Here is an excerpt.
|
||||
|
||||
```c++
|
||||
// Map back to 2.x type as best as possible
|
||||
using LayoutA = gemm::detail::StrideToLayoutTagA_t<typename GemmKernel::StrideA>;
|
||||
using LayoutB = gemm::detail::StrideToLayoutTagB_t<typename GemmKernel::StrideB>;
|
||||
using LayoutC = gemm::detail::StrideToLayoutTagC_t<typename GemmKernel::StrideC>;
|
||||
using LayoutD = gemm::detail::StrideToLayoutTagC_t<typename GemmKernel::StrideD>;
|
||||
|
||||
// Legacy: Assume MultiplyAdd only since we do not use this tag type in 3.0
|
||||
using MathOperator = cutlass::arch::OpMultiplyAdd;
|
||||
|
||||
// If our TiledMMA's instruction thread layout size is larger than 1,
|
||||
// we know it's a tensorop
|
||||
using OperatorClass = std::conditional_t<
|
||||
(cute::size(typename GemmKernel::TiledMma::AtomThrID{}) > 1),
|
||||
cutlass::arch::OpClassTensorOp, cutlass::arch::OpClassSimt>;
|
||||
|
||||
// Assume TiledMma's ShapeMNK is the same as 2.x's ThreadblockShape
|
||||
using ThreadblockShape = cutlass::gemm::GemmShape<
|
||||
cute::size<0>(TileShape{}),
|
||||
cute::size<1>(TileShape{}),
|
||||
cute::size<2>(TileShape{})>;
|
||||
|
||||
using ClusterShape = cutlass::gemm::GemmShape<
|
||||
cute::size<0>(typename GemmKernel::DispatchPolicy::ClusterShape{}),
|
||||
cute::size<1>(typename GemmKernel::DispatchPolicy::ClusterShape{}),
|
||||
cute::size<2>(typename GemmKernel::DispatchPolicy::ClusterShape{})>;
|
||||
|
||||
// We get the instruction shape directly from our TiledMma's atom shape
|
||||
using InstructionShape = cutlass::gemm::GemmShape<
|
||||
cute::size<0>(typename CollectiveMainloop::TiledMma::AtomShape_MNK{}),
|
||||
cute::size<1>(typename CollectiveMainloop::TiledMma::AtomShape_MNK{}),
|
||||
cute::size<2>(typename CollectiveMainloop::TiledMma::AtomShape_MNK{})>;
|
||||
|
||||
static int constexpr kStages = CollectiveMainloop::DispatchPolicy::Stages;
|
||||
static int const kThreadCount = GemmKernel::MaxThreadsPerBlock;
|
||||
|
||||
// Warp shape is not a primary API type in 3.x,
|
||||
// but we can best approximate it by inspecting the TiledMma
|
||||
// For this, we make the assumption that we always have 4 warps along M,
|
||||
// and the rest along N, with none along K. We also always round up
|
||||
// the warp count to 4 if the tiled mma is smaller than 128 threads.
|
||||
static constexpr int WarpsInMma = std::max(4, CUTE_STATIC_V(cute::size(typename GemmKernel::TiledMma{})) / 32);
|
||||
static constexpr int WarpsInMmaM = 4;
|
||||
static constexpr int WarpsInMmaN = cute::ceil_div(WarpsInMma, WarpsInMmaM);
|
||||
using WarpCount = cutlass::gemm::GemmShape<WarpsInMmaM, WarpsInMmaN, 1>;
|
||||
using WarpShape = cutlass::gemm::GemmShape<
|
||||
CUTE_STATIC_V(cute::tile_size<0>(typename CollectiveMainloop::TiledMma{})) / WarpsInMmaM,
|
||||
CUTE_STATIC_V(cute::tile_size<1>(typename CollectiveMainloop::TiledMma{})) / WarpsInMmaN,
|
||||
CUTE_STATIC_V(cute::tile_size<2>(typename CollectiveMainloop::TiledMma{}))>;
|
||||
|
||||
// Inspect TiledCopy for A and B to compute the alignment size
|
||||
static int constexpr kAlignmentA = gemm::detail::get_alignment_count_from_gmem_tiled_copy<
|
||||
typename CollectiveMainloop::GmemTiledCopyA, ElementA>();
|
||||
static int constexpr kAlignmentB = gemm::detail::get_alignment_count_from_gmem_tiled_copy<
|
||||
typename CollectiveMainloop::GmemTiledCopyB, ElementB>();
|
||||
```
|
||||
|
||||
CUTLASS's library and profiler use these reflective interfaces to
|
||||
obtain the kernel's configuration parameters. Users can use these to approximate the CUTLASS 2.x types
|
||||
for 3.0 API kernels. However, the reflective interfaces cannot always match the types exactly,
|
||||
as the mappings are not always bijective.
|
||||
|
||||
# Copyright
|
||||
|
||||
Copyright (c) 2023 - 2025 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
```
|
||||
Redistribution and use in source and binary forms, with or without
|
||||
modification, are permitted provided that the following conditions are met:
|
||||
|
||||
1. Redistributions of source code must retain the above copyright notice, this
|
||||
list of conditions and the following disclaimer.
|
||||
|
||||
2. Redistributions in binary form must reproduce the above copyright notice,
|
||||
this list of conditions and the following disclaimer in the documentation
|
||||
and/or other materials provided with the distribution.
|
||||
|
||||
3. Neither the name of the copyright holder nor the names of its
|
||||
contributors may be used to endorse or promote products derived from
|
||||
this software without specific prior written permission.
|
||||
|
||||
THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS "AS IS"
|
||||
AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE
|
||||
IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE ARE
|
||||
DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT HOLDER OR CONTRIBUTORS BE LIABLE
|
||||
FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL
|
||||
DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR
|
||||
SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER
|
||||
CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY,
|
||||
OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE
|
||||
OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
|
||||
```
|
||||
147
media/docs/cpp/cutlass_3x_design.md
Normal file
147
media/docs/cpp/cutlass_3x_design.md
Normal file
@@ -0,0 +1,147 @@
|
||||
# CUTLASS 3.0 Design
|
||||
|
||||
CUTLASS 3.0 is a major enhancement over the abstractions of CUTLASS 2.x
|
||||
and aims to make usage of all layers of the GEMM hierarchy easier and more composable
|
||||
while still achieving peak performance on Hardware.
|
||||
|
||||
## CUTLASS 3.0 design goals
|
||||
|
||||
CUTLASS 3.0 has the following design goals, in no particular order.
|
||||
|
||||
- Simplify expressing and manipulating data and thread layouts across
|
||||
the GEMM hierarchy with CuTe layouts and layout algebra.
|
||||
|
||||
- Improve code readability and learning curve by
|
||||
reducing the number of named types.
|
||||
|
||||
- Functional correctness by default,
|
||||
actionable static asserts otherwise.
|
||||
|
||||
- Single, clear points of performance tuning and custom kernel extensions.
|
||||
|
||||
- Support for NVIDIA Hopper GPUs with great performance using
|
||||
features such as Tensor Cores, tensor memory accelerator, and thread block clusters.
|
||||
|
||||
## A new Conceptual GEMM Hierarchy
|
||||
|
||||
CUTLASS 2.x decomposes the moving parts of a GEMM operation
|
||||
across a hierarchy that closely mirrors the organization of GPU
|
||||
architectures. This discussed in detail within the
|
||||
[CUTLASS 2.x GEMM API documentation](gemm_api.md).
|
||||
This design, however, sometimes results in a coupling that is too tight
|
||||
to extend to newer GPU features that might not fit into the same architectural
|
||||
hierarchy. For instance, Hopper's warp-group wide instructions do not naturally
|
||||
fit into any warp or thread layer GEMM concept in CUTLASS 2.x. Even for Volta tensor cores,
|
||||
instructions that atomically exist at the quad-pair granularity are first tiled at
|
||||
the warp level before use. This hints at the brittleness of the abstraction power.
|
||||
|
||||
CUTLASS 3.0 detaches its interface layers from the hardware,
|
||||
centering them instead around the natural structure of GEMM algorithms
|
||||
not tied to any particular GPU generation.
|
||||
This makes CUTLASS's code more robust to GPU architecture evolution,
|
||||
less prone to implementation detail leakage, and provides users
|
||||
with a consistent interface to hardware acceleration regardless of
|
||||
the architecture specific details.
|
||||
|
||||
The new conceptual GEMM hierarchy is discussed in detail in the dedicated
|
||||
[CUTLASS 3.0 GEMM API documentation readme](gemm_api_3x.md),
|
||||
along with code examples of the core concepts and types.
|
||||
|
||||
## Adoption of CuTe Layout and Tensors
|
||||
|
||||
CUTLASS 3.0 introduces a new core library, CuTe, to describe and manipulate tensors of threads and data.
|
||||
CuTe is a collection of C++ CUDA template abstractions for defining and operating on hierarchically multidimensional layouts of threads and data. CuTe provides `Layout` and `Tensor` objects that compactly packages the type, shape, memory space, and layout of data, while performing the complicated indexing for the user.
|
||||
|
||||
CUTLASS 3.0 adopts CuTe throughout the GEMM hierarchy in its templates, greatly simplifying the design,
|
||||
improving code composability, and readability. More documentation specific to CuTe can be found in its [dedicated documentation directory](cute/00_quickstart.md).
|
||||
|
||||

|
||||
|
||||
Programming massively parallel systems with various layers of logical thread and data hierarchies is not a trivial task.
|
||||
|
||||
- `cute::Layout`s always maintain logical consistency of their coordinates,
|
||||
allowing us to check pre- and post-conditions at compile time for all static inner loops.
|
||||
- Explicit thread to data mapping allows users and kernel authors to inspect and reason about operations
|
||||
from a single point in the source code.
|
||||
- Layouts provide a single point of performance tuning, as most optimizations can be done by careful
|
||||
selection of thread and data layouts.
|
||||
- Formalized algebra makes manipulation of and reasoning about thread->data mapping explicit in source code.
|
||||
- Single vocabulary type (`cute::Layout`) subsumes every iterator and layout in CUTLASS 2.x CUTLASS 2.x uses many bespoke thread maps, iterators, and data layouts. Iterators are fundamentally 1-D, whereas most layouts we encounter in the GPU hierarchy are fundamentally n-D.
|
||||
|
||||
## Reducing the number of named types and iterator concepts
|
||||
|
||||
CUTLASS 2.x design preferred introducing bespoke named types for each
|
||||
architecture specific thread and data layout. For instance, `gemm::treadblock` namespace
|
||||
contains implementation for `MmaMultistage`, `MmaPlanarComplexMultistage`, `MmaPipelined` etc.
|
||||
despite them providing mainloops for GEMMs. To spell these types the same way in generic code,
|
||||
CUTLASS 2.x provides aliases through its `default_x_configuration.h` files, however,
|
||||
these aliases make the code much harder to read as the user has to perform type substitution
|
||||
mentally in order to understand the codebase.
|
||||
|
||||
CUTLASS 3.0 greatly reduces the number of named types used throughout by
|
||||
|
||||
- Replacing all iterator concepts for all memory domains with `cute::Tensor`s
|
||||
- Dispatching mainloop and epilogue implementations on tag-dispatch policies rather than naming new types
|
||||
- Dispatching kernel layer schedules on tag-dispatch policies rather than naming new types
|
||||
|
||||
Reducing the number of named types has many benefits:
|
||||
|
||||
- It *makes writing generic code easier*, as the primary type names share the same lexical
|
||||
without aliasing through configuration providers.
|
||||
- It *flattens the learning curve of CUTLASS* by greatly reducing the mental context required
|
||||
as the library only exposes a handful of named types.
|
||||
- It *provides a clear, singular extension point* for users to plug in their customizations
|
||||
through the dispatch policies.
|
||||
|
||||
## Correctness by default, Performance through clear, individual points of tuning
|
||||
|
||||
CUTLASS 2.x maintained its thread layouts as implicit indexing math implemented
|
||||
as a part of 1D iterators. This meant that the thread to data layout mapping
|
||||
was implicit in the imperative structure of the C++ code itself and did not have
|
||||
a formal algebra we could use to manipulate these mappings. Each iterator
|
||||
had to re-implement its indexing and mapping logic. This made it hard to learn
|
||||
how this mapping was performed for existing iterators, and even harder to
|
||||
implement custom layout functions for the core inner loops of a GEMM.
|
||||
|
||||
CUTLASS 3.0 replaces all iterator concepts from CUTLASS 2.x
|
||||
with a single layout type for thread and data tensors.
|
||||
CuTe's formalized layout algebra is then used at every layer of
|
||||
the GEMM hierarchy to manipulate the mapping between the two.
|
||||
CuTe layouts always maintain logical consistency, and for fully static layouts
|
||||
(such as in the core unrolled inner loops), provide
|
||||
compile time checks that break builds if this consistency is violated.
|
||||
In this way, CuTe reifies the thread-to-data-layout mapping,
|
||||
makes it easier to write code that is "correct by construction".
|
||||
If the code compiles, it's probably correct.
|
||||
|
||||
## Copyright
|
||||
|
||||
Copyright (c) 2017 - 2025 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
```
|
||||
Redistribution and use in source and binary forms, with or without
|
||||
modification, are permitted provided that the following conditions are met:
|
||||
|
||||
1. Redistributions of source code must retain the above copyright notice, this
|
||||
list of conditions and the following disclaimer.
|
||||
|
||||
2. Redistributions in binary form must reproduce the above copyright notice,
|
||||
this list of conditions and the following disclaimer in the documentation
|
||||
and/or other materials provided with the distribution.
|
||||
|
||||
3. Neither the name of the copyright holder nor the names of its
|
||||
contributors may be used to endorse or promote products derived from
|
||||
this software without specific prior written permission.
|
||||
|
||||
THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS "AS IS"
|
||||
AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE
|
||||
IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE ARE
|
||||
DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT HOLDER OR CONTRIBUTORS BE LIABLE
|
||||
FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL
|
||||
DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR
|
||||
SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER
|
||||
CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY,
|
||||
OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE
|
||||
OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
|
||||
```
|
||||
74
media/docs/cpp/dependent_kernel_launch.md
Normal file
74
media/docs/cpp/dependent_kernel_launch.md
Normal file
@@ -0,0 +1,74 @@
|
||||
# Dependent kernel launches
|
||||
|
||||
The Hopper and Blackwell architectures supports a new feature through which two kernels in the same stream can
|
||||
overlap their execution, named
|
||||
[Programmatic Dependent Launch (PDL)](https://docs.nvidia.com/cuda/cuda-c-programming-guide/index.html#programmatic-dependent-launch-and-synchronization).
|
||||
This allows kernels with conflict in global memory to programmatically and safely overlap portions
|
||||
of their execution. Primary kernel can signal it is about to finish execution, and the next kernel is expected to
|
||||
programatically wait on the previous kernel to finish flushing its memory.
|
||||
|
||||
We enable PDL by setting a flag through the extended CUDA launch APIs. All CUTLASS kernels with PDL support
|
||||
will wait on the prior kernel to flush its output to memory and signal the next kernel to start. This means
|
||||
they can safely be dropped in with any other set of kernels using PDL as long as they also adhear to waiting on
|
||||
the prior to flush its memory as well.
|
||||
|
||||
For more information, we refer you to the [PDL section in the CUDA Programming Guide](https://docs.nvidia.com/cuda/cuda-c-programming-guide/index.html#programmatic-dependent-launch-and-synchronization).
|
||||
|
||||
## Using dependent launch in CUTLASS
|
||||
|
||||
When building CUTLASS, you can use the `CUTLASS_ENABLE_GDC_FOR_SM90` and `CUTLASS_ENABLE_GDC_FOR_SM100` macro
|
||||
respectively to enable PDL-related instructions:
|
||||
|
||||
```
|
||||
cmake . -DCUTLASS_ENABLE_GDC_FOR_SM90=1
|
||||
```
|
||||
|
||||
Note that this only adds PDL-related instructions to the _kernels_, but to actually allow a dependent
|
||||
launch, you must also run your GEMM kernel with PDL:
|
||||
|
||||
```
|
||||
gemm.run(
|
||||
/* stream = */ stream,
|
||||
/* cuda_adapter = */ nullptr,
|
||||
/* launch_with_pdl = */ true
|
||||
);_
|
||||
```
|
||||
## Model-Aware Optimizations with PDL
|
||||
|
||||
In [example 63](https://github.com/NVIDIA/cutlass/tree/main/examples/63_hopper_gemm_with_weight_prefetch/README.md), we use PDL to explicitly optimize for
|
||||
performance of kernels where we know that one of the input matricies (our weights) will not be produced by a prior
|
||||
kernel. In that case, we only need to wait on the prior kernels memory flush in order to load the other input matrix
|
||||
(our activations). During our prologue, we can prefetch our weights to improve performance for memory bandwidth-bound
|
||||
problem sizes. For more informations we refer the reader to [the example](https://github.com/NVIDIA/cutlass/tree/main/examples/63_hopper_gemm_with_weight_prefetch/README.md).
|
||||
|
||||
## Copyright
|
||||
|
||||
Copyright (c) 2017 - 2025 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
```
|
||||
Redistribution and use in source and binary forms, with or without
|
||||
modification, are permitted provided that the following conditions are met:
|
||||
|
||||
1. Redistributions of source code must retain the above copyright notice, this
|
||||
list of conditions and the following disclaimer.
|
||||
|
||||
2. Redistributions in binary form must reproduce the above copyright notice,
|
||||
this list of conditions and the following disclaimer in the documentation
|
||||
and/or other materials provided with the distribution.
|
||||
|
||||
3. Neither the name of the copyright holder nor the names of its
|
||||
contributors may be used to endorse or promote products derived from
|
||||
this software without specific prior written permission.
|
||||
|
||||
THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS "AS IS"
|
||||
AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE
|
||||
IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE ARE
|
||||
DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT HOLDER OR CONTRIBUTORS BE LIABLE
|
||||
FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL
|
||||
DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR
|
||||
SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER
|
||||
CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY,
|
||||
OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE
|
||||
OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
|
||||
```
|
||||
69
media/docs/cpp/doxygen_mainpage.md
Normal file
69
media/docs/cpp/doxygen_mainpage.md
Normal file
@@ -0,0 +1,69 @@
|
||||
# CUTLASS 3.0
|
||||
|
||||
_CUTLASS 3.0 - January 2023_
|
||||
|
||||
CUTLASS is a collection of CUDA C++ template abstractions for implementing
|
||||
high-performance matrix-multiplication (GEMM) at all levels and scales within CUDA.
|
||||
It incorporates strategies for hierarchical decomposition and data movement similar
|
||||
to those used to implement cuBLAS. CUTLASS decomposes these "moving parts" into
|
||||
reusable, modular software components abstracted by C++ template classes. These
|
||||
components can be specialized
|
||||
and tuned via custom tiling sizes, data types, and other algorithmic policies. The
|
||||
resulting flexibility simplifies their use as building blocks within custom kernels
|
||||
and applications.
|
||||
|
||||
To support a wide variety of applications, CUTLASS provides extensive support for
|
||||
mixed-precision computations, providing specialized data-movement and
|
||||
multiply-accumulate abstractions for 8-bit integer, half-precision floating
|
||||
point (FP16), single-precision floating point (FP32), and double-precision floating
|
||||
point (FP64) types. Furthermore, CUTLASS exploits the _Tensor Cores_ and asynchronous
|
||||
memory copy operations of the latest NVIDIA GPU architectures.
|
||||
|
||||
# What's New in CUTLASS 3.0
|
||||
|
||||
For an overview of CUTLASS 3.0's GEMM interface levels,
|
||||
please refer to the
|
||||
[CUTLASS 3.0 GEMM API document](./gemm_api_3x.md).
|
||||
To learn how to migrate code using CUTLASS 2.x's interface
|
||||
to CUTLASS 3.0, please refer to the
|
||||
[backwards compatibility document](./cutlass_3x_backwards_compatibility.md).
|
||||
|
||||
# GEMM examples
|
||||
|
||||
For a code example showing how to define
|
||||
a GEMM kernel using CUTLASS, please refer to
|
||||
[the quickstart guide](./quickstart.md).
|
||||
The [`examples` directory](https://github.com/NVIDIA/cutlass/tree/main/examples)
|
||||
has a variety of examples.
|
||||
|
||||
# Copyright
|
||||
|
||||
Copyright (c) 2017 - 2025 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
```
|
||||
Redistribution and use in source and binary forms, with or without
|
||||
modification, are permitted provided that the following conditions are met:
|
||||
|
||||
1. Redistributions of source code must retain the above copyright notice, this
|
||||
list of conditions and the following disclaimer.
|
||||
|
||||
2. Redistributions in binary form must reproduce the above copyright notice,
|
||||
this list of conditions and the following disclaimer in the documentation
|
||||
and/or other materials provided with the distribution.
|
||||
|
||||
3. Neither the name of the copyright holder nor the names of its
|
||||
contributors may be used to endorse or promote products derived from
|
||||
this software without specific prior written permission.
|
||||
|
||||
THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS "AS IS"
|
||||
AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE
|
||||
IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE ARE
|
||||
DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT HOLDER OR CONTRIBUTORS BE LIABLE
|
||||
FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL
|
||||
DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR
|
||||
SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER
|
||||
CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY,
|
||||
OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE
|
||||
OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
|
||||
```
|
||||
290
media/docs/cpp/efficient_gemm.md
Normal file
290
media/docs/cpp/efficient_gemm.md
Normal file
@@ -0,0 +1,290 @@
|
||||

|
||||
|
||||
# Efficient GEMM in CUDA
|
||||
|
||||
CUTLASS implements the hierarchically blocked structure described in
|
||||
[CUTLASS: Fast Linear Algebra in CUDA C++](https://devblogs.nvidia.com/cutlass-linear-algebra-cuda/)
|
||||
and the [CUTLASS GTC2018 talk](http://on-demand.gputechconf.com/gtc/2018/presentation/s8854-cutlass-software-primitives-for-dense-linear-algebra-at-all-levels-and-scales-within-cuda.pdf).
|
||||
|
||||
## Hierarchical Structure
|
||||
|
||||
The basic triple loop nest computing matrix multiply may be blocked and tiled to match
|
||||
concurrency in hardware, memory locality, and parallel programming models. In CUTLASS,
|
||||
GEMM is mapped to NVIDIA GPUs with the structure illustrated by the following loop nest.
|
||||
|
||||
```c++
|
||||
for (int cta_n = 0; cta_n < GemmN; cta_n += CtaTileN) { // for each threadblock_y } threadblock-level concurrency
|
||||
for (int cta_m = 0; cta_m < GemmM; cta_m += CtaTileM) { // for each threadblock_x }
|
||||
|
||||
for (int cta_k = 0; cta_k < GemmK; cta_k += CtaTileK) { // "GEMM mainloop" - no unrolling
|
||||
// - one iteration of this loop is one "stage"
|
||||
//
|
||||
for (int warp_n = 0; warp_n < CtaTileN; warp_n += WarpTileN) { // for each warp_y } warp-level parallelism
|
||||
for (int warp_m = 0; warp_m < CtaTileM; warp_m += WarpTileM) { // for each warp_x }
|
||||
//
|
||||
for (int warp_k = 0; warp_k < CtaTileK; warp_k += WarpTileK) { // fully unroll across CtaTileK
|
||||
// - one iteration of this loop is one "k Group"
|
||||
//
|
||||
for (int mma_k = 0; mma_k < WarpTileK; mma_k += MmaK) { // for each mma instruction } instruction-level parallelism
|
||||
for (int mma_n = 0; mma_n < WarpTileN; mma_n += MmaN) { // for each mma instruction }
|
||||
for (int mma_m = 0; mma_m < WarpTileM; mma_m += MmaM) { // for each mma instruction }
|
||||
//
|
||||
mma_instruction(d, a, b, c); // TensorCore matrix computation
|
||||
|
||||
} // for mma_m
|
||||
} // for mma_n
|
||||
} // for mma_k
|
||||
|
||||
} // for warp_k
|
||||
} // for warp_m
|
||||
} // for warp_n
|
||||
|
||||
} // for cta_k
|
||||
} // for cta_m
|
||||
} // for cta_n
|
||||
```
|
||||
|
||||
This tiled loop nest targets concurrency among
|
||||
- threadblocks,
|
||||
- warps, and
|
||||
- CUDA and Tensor Cores.
|
||||
|
||||
It takes advantage of memory locality within
|
||||
- shared memory and
|
||||
- registers.
|
||||
|
||||
The figure below illustrates the flow of data within this structure.
|
||||
This is the hierarchical GEMM computation embodied by CUTLASS. Each stage depicts a
|
||||
nested level of tiling which corresponds to a layer of concurrency within the CUDA execution model and to a
|
||||
level within the memory hierarchy, becoming increasingly finer moving left to right.
|
||||
|
||||

|
||||
|
||||
|
||||
### Threadblock-level GEMM
|
||||
|
||||
Each threadblock computes its portion of the output GEMM by iteratively loading tiles of input
|
||||
matrices and computing an accumulated matrix product. At the threadblock level, data are loaded from
|
||||
global memory. The blocking strategy in general is key to achieving efficiency. However, the programmer
|
||||
must balance multiple conflicting goals. A
|
||||
larger threadblock means fewer fetches from global memory, thereby ensuring that DRAM bandwidth
|
||||
does not become a bottleneck.
|
||||
However, large threadblock tiles may not match the dimensions of the problem well. If either the
|
||||
GEMM _M_ or _N_ dimension is small, some threads within the threadblock may not perform meaningful
|
||||
work, as the threadblock may be partially outside the bounds of the problem. If both _M_ and _N_
|
||||
are small while _K_ is large, this scheme may launch relatively few threadblocks and fail to
|
||||
make full use of all multiprocessors within the GPU. Strategies to optimize performance for this case,
|
||||
as described in the section [Parallelized Reductions](efficient_gemm.md#parallelized-reductions),
|
||||
partition the GEMM K dimension across multiple threadblocks or multiple warps. These threadblocks
|
||||
or warps compute matrix products in parallel; the products are then reduced to compute the result.
|
||||
|
||||
In CUTLASS, the dimensions of the threadblock tile are specified as `ThreadblockShape::{kM, kN, kK}`
|
||||
and may be tuned to specialize the GEMM computation for the target processor and dimensions of
|
||||
the GEMM problem.
|
||||
|
||||
|
||||
### Warp-level GEMM
|
||||
|
||||
The warp-level GEMM maps to the warp-level parallelism within the CUDA execution model. Multiple
|
||||
warps within a threadblock fetch data from shared memory into registers and perform computations.
|
||||
Warp-level GEMMs may be implemented either by TensorCores issuing
|
||||
[mma.sync](https://docs.nvidia.com/cuda/parallel-thread-execution/index.html#warp-level-matrix-instructions-mma)
|
||||
or [wmma](https://docs.nvidia.com/cuda/parallel-thread-execution/index.html#warp-level-matrix-instructions-wmma-mma)
|
||||
instructions, or by thread-level matrix computations issued to CUDA cores.
|
||||
For maximum performance, access to shared memory should be bank conflict free. To maximize data
|
||||
reuse within the warp, a large warp-level GEMM tile should be chosen.
|
||||
|
||||
|
||||
### Thread-level GEMM
|
||||
|
||||
At the lowest level of blocking, each thread is responsible for processing a certain number of
|
||||
elements. Threads cannot access each other's registers, so we choose an organization that enables
|
||||
reuse of values held in registers for multiple math instructions. This results in a 2D tiled
|
||||
structure within a thread, in which each thread issues a sequence of independent math instructions
|
||||
to the CUDA cores and computes an accumulated outer product.
|
||||
|
||||
SGEMM, IGEMM, HGEMM, and DGEMM are computed by SIMT math instructions issued by thread-level matrix multiply
|
||||
procedures.
|
||||
|
||||
|
||||
## Epilogue
|
||||
|
||||
The above code focuses only on the matrix multiply computation **C = AB** whose result is
|
||||
held in the registers of each thread within the threadblock. The mapping of logical elements
|
||||
in the output tile to each thread is chosen to maximize performance of the matrix multiply
|
||||
computation but does not result in efficient, coalesced loads and stores to global memory.
|
||||
|
||||
The epilogue is a separate phase in which threads exchange data through shared memory then
|
||||
cooperatively access global memory using efficient striped access patterns. It is also
|
||||
the phase in which linear scaling and other elementwise operations may be conveniently
|
||||
computed using the matrix product results as inputs.
|
||||
|
||||
CUTLASS defines several typical epilogue operations such as linear scaling and clamping,
|
||||
but other device-side function call operators may be used to perform custom operations.
|
||||
|
||||
## Optimizations
|
||||
|
||||
The hierarchical structure described above yields an efficient mapping to the CUDA execution model and
|
||||
CUDA/TensorCores in NVIDIA GPUs. The following sections describe strategies for obtaining peak performance
|
||||
for all corners of the design space, maximizing parallelism and exploiting data locality wherever possible.
|
||||
|
||||
### Pipelining
|
||||
|
||||
The blocked structure demands a large storage allocation within the registers of each CUDA thread. The
|
||||
accumulator elements typically occupy at least half a thread's total register budget. Consequently,
|
||||
occupancy -- the number of concurrent threads, warps, and threadblocks -- is relatively low compared
|
||||
to other classes of GPU workloads. This limits the GPU's ability to hide memory latency and other stalls
|
||||
by context switching to other concurrent threads within an SM.
|
||||
|
||||
To mitigate the effects of memory latency, CUTLASS uses *software pipelining* to overlap memory accesses
|
||||
with other computation within a thread. CUTLASS accomplishes this by double buffering at the
|
||||
following scopes.
|
||||
|
||||
- **Threadblock-scoped shared memory tiles:** two tiles are allocated in shared memory.
|
||||
One is used to load data for the current matrix operation,
|
||||
while the other tile is used to buffer data loaded from global memory
|
||||
for the next mainloop iteration.
|
||||
|
||||
- **Warp-scoped matrix fragments:** two fragments are allocated within registers.
|
||||
One fragment is passed to CUDA and TensorCores during the current matrix computation,
|
||||
while the other is used to receive shared memory fetch returns
|
||||
for the next warp-level matrix operation.
|
||||
|
||||
The following diagram illustrates the efficient, pipelined mainloop body used in CUTLASS GEMMs.
|
||||
|
||||

|
||||
|
||||
### Threadblock Rasterization
|
||||
|
||||
To maximize reuse of data held in the last level cache, CUTLASS defines several functions to
|
||||
affect the mapping of threadblocks to logical partitions of the GEMM problem. These map
|
||||
consecutively launched threadblocks to packed two-dimensional regions of the partitioned GEMM
|
||||
problem to increase the probability that these will access the same tiles of global memory at
|
||||
approximately the same time.
|
||||
|
||||
Several functions are defined in [cutlass/gemm/threadblock_swizzle.h](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/gemm/threadblock/threadblock_swizzle.h).
|
||||
|
||||
|
||||
### Parallelized Reductions
|
||||
|
||||
**Split K - reduction across threadblocks**
|
||||
|
||||
Matrix product computations expose parallelism among _O(MN)_ independent inner product
|
||||
computations. For sufficiently large problem sizes, a GEMM kernel in CUTLASS may approach
|
||||
the theoretical maximum computational throughput. For small problems, however, there are
|
||||
too few threadblocks to efficiently occupy the entire GPU.
|
||||
|
||||
As a recourse, parallelizing the reduction performed during the inner product computation
|
||||
enables more threadblocks to execute concurrently while still taking advantage of the throughput
|
||||
benefits of large threadblock-level GEMM tiles.
|
||||
|
||||
CUTLASS implements parallel reductions across threadblocks by partitioning the GEMM _K_ dimension
|
||||
and launching an additional set of threadblocks for each partition. Consequently, we refer to
|
||||
this strategy within CUTLASS as "parallel reduction splitK." The "parallel reduction splitK" strategy
|
||||
requires the execution of 2 kernels: partitionedK GEMM, and batched reduction.
|
||||
|
||||
PartitionedK GEMM resembles one flavor of batched strided GEMM. Instead of requiring users
|
||||
to specify the problem size of each batch, partitionedK GEMM asks for the overall problem size and the
|
||||
number of partitions that will be applied along the K dimension for operands A and B. For example,
|
||||
parameters of m=128, n=128, k=4096 and partition=16 will result in 16 batched strided GEMMs
|
||||
with each batch of m=128, n=128, k=256. PartitionedK also allows scenario where k is not divisible
|
||||
by the partition count.
|
||||
|
||||
For example, parameters of m=128, n=128, k=4096 and partition=20
|
||||
will result in 20 batched strided GEMMs.
|
||||
The first 19 batches will have m=128, n=128, and k=4096/20=204,
|
||||
and the last batch will have m=128, n=128, and k=220.
|
||||
|
||||
The batched reduction kernel takes as input the output (C) of partitionedK GEMM,
|
||||
and performs a reduction along the K-dimension.
|
||||
Users must manage workspace memory to store this intermediate result.
|
||||
|
||||
**Sliced K - reduction across warps**
|
||||
|
||||
Similar to the split-k scenario, sliced-k aims at improving the efficiency of kernels
|
||||
with smaller M and N dimensions, but large K dimension.
|
||||
At the thread-block level, the parameters CtaTileN and CtaTileM expose parallelism
|
||||
by partitioning the work among warps.
|
||||
Larger warpTiles expose better instruction-level parallelism (ILP) and reuse,
|
||||
but also limit the number of warps running per threadblock, which reduces efficiency.
|
||||
|
||||
In order to improve efficiency in such scenarios, partitioning the warpTiles also along ctaTileK
|
||||
helps use the hardware more efficiently by allowing more warps to run concurrently in a CTA.
|
||||
Sliced-k kernels break down a threadblock's computation among participating warps
|
||||
not just among the CtaTileN, CtaTileM dimension, but also the CtaTileK dimension.
|
||||
Thus, sliced-k entails a small cost in form of a reduction
|
||||
which has to happen at the end among the participating warps.
|
||||
This is because each warp computes using only a "slice" of CtaTileK,
|
||||
so each warp only has a partial sum before the reduction.
|
||||
|
||||
### Hopper Warp Specialization
|
||||
|
||||
Note: the following section on warp-specialization contains details that are specific
|
||||
to the Hopper kernel design. Blackwell SM100 kernels have a substantially different warp-specialization structure,
|
||||
however, the concept of separating out producer and consumer agents still applies.
|
||||
|
||||
Starting with Hopper, CUTLASS 3.0 incorporates the concept of [Warp Specialization](https://docs.nvidia.com/cuda/cuda-c-programming-guide/index.html#spatial-partitioning-also-known-as-warp-specialization)
|
||||
as part of the kernel design. A thread block is partitioned into two sets of warps, [*producer* warp group](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/gemm/kernel/sm90_gemm_tma_warpspecialized.hpp) and [*consumer* warp group](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/gemm/kernel/sm90_gemm_tma_warpspecialized.hpp). The *producer* warp group loads data from global memory into shared memory buffers using the new [Tensor Memory Accelerator (TMA)](https://developer.nvidia.com/blog/nvidia-hopper-architecture-in-depth/).
|
||||
|
||||
[*Producer* warp group (DMA)](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/gemm/collective/sm90_mma_tma_gmma_ss_warpspecialized.hpp) waits for the shared memory buffers to be signaled as [empty](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/gemm/collective/sm90_mma_tma_gmma_ss_warpspecialized.hpp) by the *consumer* warp group using the newly added **Async Pipeline class** ([refer](pipeline.md)). Once the data is written into the shared memory, TMA is also updates the barrier associated with that stage to notify affected threads that the buffer has been [filled](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/gemm/collective/sm90_mma_tma_gmma_ss_warpspecialized.hpp). The [*Consumer* warp group (MMA)](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/gemm/collective/sm90_mma_tma_gmma_ss_warpspecialized.hpp) on the other hand waits for the *producer* warp group to signal that the buffer is [filled](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/gemm/collective/sm90_mma_tma_gmma_ss_warpspecialized.hpp) and then launches tensor core MMA operations. Finally, the *consumer* warp group [releases](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/gemm/collective/sm90_mma_tma_gmma_ss_warpspecialized.hpp) the buffers for the next set of TMA loads to happens.
|
||||
|
||||
**Warp-Specialized Persistent Cooperative kernel design**
|
||||
|
||||
Another flavor of Warp-Specialized kernel design being introduced starting with Hopper is the [*Warp-Specialized Persistent Cooperative*](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/gemm/kernel/sm90_gemm_tma_warpspecialized_cooperative.hpp) kernel. Like the Warp-Specialized kernel, the concepts of warp groups and barrier synchronization between warp groups remain the same in the cooperative design.
|
||||
The distinctive feature of the Warp-Specialized Persistent Cooperative kernel are the following :
|
||||
* Persistent thread blocks launched to occupy as many SMs as mentioned in the [KernelHardwareInfo](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/kernel_hardware_info.hpp) struct. These persistent thread blocks are used to tile the output and thus (potentially) compute multiple output tiles through their lifetime. The main benefit this adds is amortization of the thread-block launch and kernel prologue overheads which are typical of all kernels.
|
||||
* Presence of two *consumer* warp groups cooperating on the same output tile by splitting the tile in half across the M dimension. This allows for larger tile sizes to be enabled - since the register pressure per *consumer* warp group is reduced - and hence improving performance.
|
||||
|
||||
Since each thread block now computes multiple output tiles, the shape of the grid launch and the scheduling of tiles to the thread blocks is managed using the new [*Tile Scheduler*](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/gemm/kernel/sm90_tile_scheduler.hpp). The *Tile Scheduler* considers the shape of the *clusters* as well as the available number of available SMs to compute a valid scheduling of the output tiles to launched thread blocks.
|
||||
|
||||
**Warp-Specialized Persistent Ping-Pong kernel design**
|
||||
|
||||
The third kernel design is the [*Warp-Specialized Persistent Ping-Pong*](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/gemm/kernel/sm90_gemm_tma_warpspecialized_pingpong.hpp) kernel.
|
||||
Like the Warp-Specialized Persistent Cooperative, kernel the concepts of warp groups, barrier synchronization between warp groups, and the shape of the grid launch remain the same in the persistent ping-pong design.
|
||||
The distinctive feature of the Warp-Specialized Persistent Ping-Pong kernel is the following :
|
||||
* The two *consumer* warp groups are assigned a different output tile using the Tile Scheduler. This allows for *epilogue* of one *consumer* warp group to be overlapped with the math operations of the other *consumer* warp group - thus maximizing tensor core utilization.
|
||||
* The *producer* warp group synchronizes using the [Ordered Sequence Barrier](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/pipeline/pipeline.hpp) to fill buffers of the two *consumer* warp groups one after the other in order.
|
||||
|
||||
# Resources
|
||||
|
||||
The following additional resources describe design and implementation details of GEMMs
|
||||
targeting NVIDIA GPUs.
|
||||
|
||||
- [Developing CUDA Kernels to Push Tensor Cores to the Absolute Limit on NVIDIA A100.](https://www.nvidia.com/en-us/gtc) (SR 21745)
|
||||
- [CUTLASS: Fast Linear Algebra in CUDA C++](https://devblogs.nvidia.com/cutlass-linear-algebra-cuda/)
|
||||
- [CUTLASS: SOFTWARE PRIMITIVES FOR DENSE LINEAR ALGEBRA AT ALL LEVELS AND SCALES WITHIN CUDA](https://on-demand-gtc.gputechconf.com/gtcnew/sessionview.php?sessionName=s8854-cutlass%3a+software+primitives+for+dense+linear+algebra+at+all+levels+and+scales+within+cuda)
|
||||
- [Programming Tensor Cores: NATIVE VOLTA TENSOR CORES WITH CUTLASS](https://developer.download.nvidia.com/video/gputechconf/gtc/2019/presentation/s9593-cutensor-high-performance-tensor-operations-in-cuda-v2.pdf)
|
||||
- [CUDA Programming Guide: warp matrix functions](https://docs.nvidia.com/cuda/cuda-c-programming-guide/index.html#wmma)
|
||||
- [Matrix Multiply Accumulate Instructions](https://docs.nvidia.com/cuda/parallel-thread-execution/index.html#warp-level-matrix-instructions-mma)
|
||||
|
||||
# Copyright
|
||||
|
||||
Copyright (c) 2017 - 2025 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
```
|
||||
Redistribution and use in source and binary forms, with or without
|
||||
modification, are permitted provided that the following conditions are met:
|
||||
|
||||
1. Redistributions of source code must retain the above copyright notice, this
|
||||
list of conditions and the following disclaimer.
|
||||
|
||||
2. Redistributions in binary form must reproduce the above copyright notice,
|
||||
this list of conditions and the following disclaimer in the documentation
|
||||
and/or other materials provided with the distribution.
|
||||
|
||||
3. Neither the name of the copyright holder nor the names of its
|
||||
contributors may be used to endorse or promote products derived from
|
||||
this software without specific prior written permission.
|
||||
|
||||
THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS "AS IS"
|
||||
AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE
|
||||
IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE ARE
|
||||
DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT HOLDER OR CONTRIBUTORS BE LIABLE
|
||||
FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL
|
||||
DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR
|
||||
SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER
|
||||
CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY,
|
||||
OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE
|
||||
OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
|
||||
```
|
||||
310
media/docs/cpp/functionality.md
Normal file
310
media/docs/cpp/functionality.md
Normal file
@@ -0,0 +1,310 @@
|
||||

|
||||
|
||||
|
||||
# Functionality
|
||||
|
||||
Note : CUTLASS-3 requires users to use CUDA 11.4 or newer, and SM70 or newer, for the target toolkit and architecture, respectively.
|
||||
|
||||
- N - Column Major Matrix
|
||||
- T - Row Major matrix
|
||||
- {N,T} x {N,T} - All combinations, i.e., NN, NT, TN, TT
|
||||
- [NHWC](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/layout/tensor.h#L63-206) - 4 dimension tensor used for convolution
|
||||
- [NCxHWx](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/layout/tensor.h#L290-395) - Interleaved 4 dimension tensor used for convolution
|
||||
- f - floating point
|
||||
- s - signed int
|
||||
- b - bit
|
||||
- cf - complex float
|
||||
- bf16 - bfloat16
|
||||
- tf32 - tfloat32
|
||||
- Simt - Use Simt CUDA Core MMA
|
||||
- TensorOp - Use Tensor Core MMA
|
||||
- SpTensorOp - Use Sparse Tensor Core MMA
|
||||
- WmmaTensorOp - Use WMMA abstraction to use Tensor Core MMA
|
||||
|
||||
## Device-level GEMM
|
||||
|
||||
The following tables summarize device-level GEMM kernels in CUTLASS, organized by opcode class, data type, and layout.
|
||||
Hyperlinks to relevant unit tests demonstrate how specific template instances may be defined.
|
||||
|
||||
### CUTLASS 3.x Kernels
|
||||
|
||||
|**Opcode Class** | **Compute Capability** | **CUDA Toolkit** | **Data Type** | **Layouts** | **Unit Test** |
|
||||
|-----------------|------------------------|------------------|--------------------------------|------------------------|------------------|
|
||||
| **TensorOp** | 90a | 12.0+ | `f16 * f16 + { f16, f32 } => { f16, f32 }` | {N,T} x {N,T} => {N,T} | [example](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/sm90_gemm_f16_f16_f16_tensor_op_f32_cluster_warpspecialized.cu) |
|
||||
| **TensorOp** | 90a | 12.0+ | `bf16 * bf16 + { f16, f32 } => { bf16, f32 }`| {N,T} x {N,T} => {N,T} | [example](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/sm90_gemm_bf16_bf16_bf16_tensor_op_f32.cu) |
|
||||
| **TensorOp** | 90a | 12.0+ | `{f32, tf32} * {f32, tf32} + f32 => f32`| { T } x { N } => {N,T} | [example](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/sm90_gemm_f32_f32_f32_tensor_op_f32.cu) |
|
||||
| **TensorOp** | 90a | 12.0+ | `s8 * s8 + s32 => {s32, s8}` | { T } x { N } => {N,T} | [example](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/sm90_gemm_s8_s8_s8_tensor_op_s32.cu) |
|
||||
|
||||
|
||||
### CUTLASS 2.x Kernels
|
||||
|
||||
|**Opcode Class** | **Compute Capability** | **CUDA Toolkit** | **Data Type** | **Layouts** | **Unit Test** |
|
||||
|-----------------|------------------------|------------------|--------------------------------|------------------------|------------------|
|
||||
| **Simt** | 50+ | 11.4+ | `f32 * f32 + f32 => f32` | {N,T} x {N,T} => {N,T} | [example](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/simt_sgemm_nt_sm50.cu) |
|
||||
| **Simt** | 50+ | 11.4+ | `f64 * f64 + f64 => f64` | {N,T} x {N,T} => {N,T} | [example](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/simt_dgemm_nt_sm50.cu) |
|
||||
| **Simt** | 60+ | 11.4+ | `f16 * f16 + f16 => f16` | {N,T} x {N,T} => {N,T} | [example](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/simt_hgemm_nt_sm50.cu) |
|
||||
| **Simt** | 61+ | 11.4+ | `s8 * s8 + s32 => {s32,s8}` | {N,T} x {N,T} => {N,T} | [example](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/simt_igemm_nt_sm50.cu) |
|
||||
| **WmmaTensorOp** | 70+ | 11.4+ | `f16 * f16 + f16 => f16` | {N,T} x {N,T} => {N,T} | [example](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/gemm_f16t_f16t_f16n_wmma_tensor_op_f16_sm70.cu) |
|
||||
| **WmmaTensorOp** | 70+ | 11.4+ | `f16 * f16 + f32 => {f16, f32}`| {N,T} x {N,T} => {N,T} | [example](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/gemm_f16t_f16t_f16n_wmma_tensor_op_f32_sm70.cu) |
|
||||
| **WmmaTensorOp** | 75+ | 11.4+ | `s8 * s8 + s32 => {s32, s8}` | {N,T} x {N,T} => {N,T} | [example](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/gemm_s8t_s8n_s8t_wmma_tensor_op_s32_sm72.cu) |
|
||||
| **WmmaTensorOp** | 75+ | 11.4+ | `s4 * s4 + s32 => {s32, s4}` | {N,T} x {N,T} => {N,T} | [example](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/gemm_s4t_s4n_s4t_wmma_tensor_op_s32_sm75.cu) |
|
||||
| **WmmaTensorOp** | 75+ | 11.4+ | `b1 ^ b1 + s32 => {s32, b1}` | { T } x { N } => {N,T} | [example](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/gemm_b1t_b1n_b1t_wmma_tensor_op_s32_sm75.cu) |
|
||||
| **TensorOp** | 70+ | 11.4+ | `f16 * f16 + f16 => f16` | {N,T} x {N,T} => {N,T} | [example](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/gemm_f16n_f16t_f16t_volta_tensor_op_f16_sm70.cu) |
|
||||
| **TensorOp** | 70+ | 11.4+ | `f16 * f16 + f32 => {f16, f32}`| {N,T} x {N,T} => {N,T} | [example](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/gemm_f16n_f16t_f16t_volta_tensor_op_f32_sm70.cu) |
|
||||
| **TensorOp** | 75+ | 11.4+ | `f16 * f16 + f16 => f16` | {N,T} x {N,T} => {N,T} | [example](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/gemm_f16n_f16t_f16t_tensor_op_f16_sm75.cu) |
|
||||
| **TensorOp** | 75+ | 11.4+ | `f16 * f16 + f32 => {f16, f32}`| {N,T} x {N,T} => {N,T} | [example](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/gemm_f16n_f16t_f16t_tensor_op_f32_sm75.cu) |
|
||||
| **TensorOp** | 75+ | 11.4+ | `s8 * s8 + s32 => {s32, s8}` | { T } x { N } => {N,T} | [example](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/gemm_s8t_s8n_s32n_tensor_op_s32_sm75.cu) |
|
||||
| **TensorOp** | 75+ | 11.4+ | `s4 * s4 + s32 => {s32, s4}` | { T } x { N } => {N,T} | [example](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/gemm_s4t_s4n_s32n_tensor_op_s32_sm75.cu) |
|
||||
| **TensorOp** | 75+ | 11.4+ | `b1 ^ b1 + s32 => {s32, b1}` | { T } x { N } => {N,T} | [example](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/gemm_b1t_b1n_s32n_tensor_op_s32_sm75.cu) |
|
||||
| **TensorOp** | 80+ | 11.4+ | `f16 * f16 + f16 => f16` | {N,T} x {N,T} => {N,T} | [example](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/gemm_f16n_f16t_f16t_tensor_op_f16_sm80.cu) |
|
||||
| **TensorOp** | 80+ | 11.4+ | `f16 * f16 + f32 => {f16, f32}`| {N,T} x {N,T} => {N,T} | [example](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/gemm_f16n_f16t_f16t_tensor_op_f32_sm80.cu) |
|
||||
| **TensorOp** | 80+ | 11.4+ | `bf16 * bf16 + f32 => {bf16, f32}`| {N,T} x {N,T} => {N,T} | [example](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/gemm_bf16n_bf16t_bf16t_tensor_op_f32_sm80.cu) |
|
||||
| **TensorOp** | 80+ | 11.4+ | `tf32 * tf32 + f32 => f32`| {N,T} x {N,T} => {N,T} | [example](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/gemm_f32n_f32t_f32t_tensor_op_f32_sm80.cu) |
|
||||
| **TensorOp** | 80+ | 11.4+ | `s8 * s8 + s32 => {s32, s8}` | { T } x { N } => {N,T} | [example](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/gemm_s8t_s8n_s32n_tensor_op_s32_sm80.cu) |
|
||||
| **TensorOp** | 80+ | 11.4+ | `s4 * s4 + s32 => {s32, s4}` | { T } x { N } => {N,T} | [example](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/gemm_s4t_s4n_s32n_tensor_op_s32_sm80.cu) |
|
||||
| **TensorOp** | 80+ | 11.4+ | `b1 ^ b1 + s32 => {s32, b1}` | { T } x { N } => {N,T} | [example](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/gemm_b1t_b1n_s32n_tensor_op_s32_sm80.cu) |
|
||||
| **TensorOp** | 80+ | 11.4+ | `f64 * f64 + f64 => f64` | {N,T} x {N,T} => {N,T} | [example](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/gemm_f64n_f64t_f64t_tensor_op_f64_sm80.cu) |
|
||||
| **TensorOp** | 80+ | 11.4+ | `cf32 * cf32 + cf32 => cf32` | {N,T} x {N,T} => {N,T} | [example](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/gemm_cf32n_cf32t_cf32t_tensor_op_tf32_f32_sm80.cu) |
|
||||
| **TensorOp** | 80+ | 11.4+ | `cf64 * cf64 + cf64 => cf64` | {N,T} x {N,T} => {N,T} | [example](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/gemm_cf64n_cf64t_cf64t_tensor_op_f64_sm80.cu), [Gaussian 3m](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/gemm_cf64n_cf64t_cf64t_tensor_op_f64_gaussian_sm80.cu) |
|
||||
| **SpTensorOp** | 80+ | 11.4+ | `f16 * f16 + f32 => {f16, f32}` | {N,T} x {N,T} => {N,T} | [example](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/gemm_f16n_f16n_f32t_tensor_op_f32_sparse_sm80.cu) |
|
||||
| **SpTensorOp** | 80+ | 11.4+ | `bf16 * bf16 + f32 => {bf16, f32}` | {N,T} x {N,T} => {N,T} | [example](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/gemm_f16n_f16n_f32t_tensor_op_f32_sparse_sm80.cu) |
|
||||
| **SpTensorOp** | 80+ | 11.4+ | `tf32 * tf32 + f32 => f32` | {N,T} x {N,T} => {N,T} | [example](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/gemm_f32n_f32n_f32t_tensor_op_f32_sparse_sm80.cu) |
|
||||
| **SpTensorOp** | 80+ | 11.4+ | `s8 * s8 + s32 => {s8, s32}` | {N,T} x {N,T} => {N,T} | [example](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/gemm_s8t_s8n_s32t_tensor_op_s32_sparse_sm80.cu) |
|
||||
| **SpTensorOp** | 80+ | 11.4+ | `s4 * s4 + s32 => {s4, s32}` | {N,T} x {N,T} => {N,T} | [example](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/gemm_s4t_s4n_s32t_tensor_op_s32_sparse_sm80.cu) |
|
||||
| **TensorOp** | 90+ | 11.8+ | `f64 * f64 + f64 => f64` | {N,T} x {N,T} => {N,T} | [example](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/gemm_f64n_f64t_f64t_tensor_op_f64_sm90.cu) |
|
||||
|
||||
|
||||
## Device-level Implicit GEMM convolution
|
||||
|
||||
The following table summarizes device-level implicit GEMM convolution kernels in CUTLASS, organized by opcode class, data type, and layout.
|
||||
Hyperlinks to relevant conv2d fprop unit tests demonstrate how specific template instances may be defined.
|
||||
One can find and/or create equivalent dgrad and wgrad convolutional operators.
|
||||
|
||||
|**Opcode Class** | **Compute Capability** | **CUDA Toolkit** | **Data Type** | **Layouts** | **Unit Test** |
|
||||
|-----------------|------------------------|------------------|--------------------------------|------------------|------------------|
|
||||
| **Simt** | 50+ | 11.4+ | `f32 * f32 + f32 => f32` | NHWC | [example](https://github.com/NVIDIA/cutlass/tree/main/test/unit/conv/device/conv2d_fprop_implicit_gemm_f32nhwc_f32nhwc_f32nhwc_simt_f32_sm50.cu) |
|
||||
| **Simt** | 50+ | 11.4+ | `cf32 * cf32 + cf32 => cf32` | NHWC | [example](https://github.com/NVIDIA/cutlass/tree/main/test/unit/conv/device/conv2d_fprop_implicit_gemm_cf32nhwc_cf32nhwc_cf32nhwc_simt_f32_sm50.cu) |
|
||||
| **TensorOp** | 70+ | 11.4+ | `f16 * f16 + f32 => {f16, f32}`| NHWC | [example](https://github.com/NVIDIA/cutlass/tree/main/test/unit/conv/device/conv2d_fprop_implicit_gemm_f16nhwc_f16nhwc_f32nhwc_tensor_op_f32_sm70.cu) |
|
||||
| **TensorOp** | 75+ | 11.4+ | `f16 * f16 + f32 => {f16, f32}`| NHWC | [example](https://github.com/NVIDIA/cutlass/tree/main/test/unit/conv/device/conv2d_fprop_implicit_gemm_f16nhwc_f16nhwc_f32nhwc_tensor_op_f32_sm75.cu) |
|
||||
| **TensorOp** | 75+ | 11.4+ | `s8 * s8 + s32 => {s32, s8}` | NHWC, NCxHWx | [example](https://github.com/NVIDIA/cutlass/tree/main/test/unit/conv/device/conv2d_fprop_implicit_gemm_s8nhwc_s8nhwc_s32nhwc_tensor_op_s32_sm75.cu), [ncxhwx](https://github.com/NVIDIA/cutlass/tree/main/test/unit/conv/device/conv2d_fprop_implicit_gemm_s8ncxhwx_s8cxrskx_s8ncxhwx_tensor_op_s32_sm75.cu) |
|
||||
| **TensorOp** | 75+ | 11.4+ | `s4 * s4 + s32 => {s32, s4}` | NHWC, NCxHWx | [example](https://github.com/NVIDIA/cutlass/tree/main/test/unit/conv/device/conv2d_fprop_implicit_gemm_s4nhwc_s4nhwc_s32nhwc_tensor_op_s32_sm75.cu), [ncxhwx](https://github.com/NVIDIA/cutlass/tree/main/test/unit/conv/device/conv2d_fprop_implicit_gemm_s4ncxhwx_s4cxrskx_s4ncxhwx_tensor_op_s32_sm75.cu) |
|
||||
| **Simt** | 80+ | 11.4+ | `f32 * f32 + f32 => f32` | NHWC | [example](https://github.com/NVIDIA/cutlass/tree/main/test/unit/conv/device/conv2d_fprop_implicit_gemm_f32nhwc_f32nhwc_f32nhwc_simt_f32_sm80.cu) |
|
||||
| **Simt** | 80+ | 11.4+ | `cf32 * cf32 + cf32 => cf32` | NHWC | [example](https://github.com/NVIDIA/cutlass/tree/main/test/unit/conv/device/conv2d_fprop_implicit_gemm_cf32nhwc_cf32nhwc_cf32nhwc_simt_f32_sm80.cu) |
|
||||
| **TensorOp** | 80+ | 11.4+ | `f16 * f16 + f32 => {f16, f32}`| NHWC | [example](https://github.com/NVIDIA/cutlass/tree/main/test/unit/conv/device/conv2d_fprop_implicit_gemm_f16nhwc_f16nhwc_f32nhwc_tensor_op_f32_sm80.cu) |
|
||||
| **TensorOp** | 80+ | 11.4+ | `f16 * f16 + f16 => f16` | NHWC | [example](https://github.com/NVIDIA/cutlass/tree/main/test/unit/conv/device/conv2d_fprop_implicit_gemm_f16nhwc_f16nhwc_f32nhwc_tensor_op_f32_sm80.cu) |
|
||||
| **TensorOp** | 80+ | 11.4+ | `tf32 * tf32 + f32 => f32` | NHWC | [example](https://github.com/NVIDIA/cutlass/tree/main/test/unit/conv/device/conv2d_fprop_implicit_gemm_tf32nhwc_tf32nhwc_f32nhwc_tensor_op_f32_sm80.cu) |
|
||||
| **TensorOp** | 80+ | 11.4+ | `s8 * s8 + s32 => {s32, s8}` | NHWC, NCxHWx | [example](https://github.com/NVIDIA/cutlass/tree/main/test/unit/conv/device/conv2d_fprop_implicit_gemm_s8nhwc_s8nhwc_s32nhwc_tensor_op_s32_sm80.cu), [ncxhwx](https://github.com/NVIDIA/cutlass/tree/main/test/unit/conv/device/conv2d_fprop_implicit_gemm_s8ncxhwx_s8cxrskx_s8ncxhwx_tensor_op_s32_sm80.cu) |
|
||||
| **TensorOp** | 80+ | 11.4+ | `s4 * s4 + s32 => {s32, s4}` | NHWC, NCxHWx | [example](https://github.com/NVIDIA/cutlass/tree/main/test/unit/conv/device/conv2d_fprop_implicit_gemm_s4nhwc_s4nhwc_s32nhwc_tensor_op_s32_sm80.cu), [ncxhwx](https://github.com/NVIDIA/cutlass/tree/main/test/unit/conv/device/conv2d_fprop_implicit_gemm_s4ncxhwx_s4cxrskx_s4ncxhwx_tensor_op_s32_sm80.cu) |
|
||||
|
||||
|
||||
|
||||
## Warp-level Matrix Multiply with Tensor Cores
|
||||
|
||||
The following table summarizes supported warp level shapes for each TensorOp instruction.
|
||||
|
||||
|**Opcode Class** | **Instruction Shape** | **Warp Shapes** |
|
||||
|-----------------|-----------------------|--------------------------------------------|
|
||||
| **TensorOp** | 8-by-8-by-4 | 32x32x4, 32x64x4, 64x32x4, 64x64x4 |
|
||||
| **TensorOp** | 16-by-8-by-8 | 32x32x8, 32x64x8, 64x32x8, 64x64x8 |
|
||||
| **TensorOp** | 16-by-8-by-16 | 32x32x16, 32x64x16, 64x32x16, 64x64x16 |
|
||||
| **TensorOp** | 8-by-8-by-16 | 32x32x16, 32x64x16, 64x32x16, 64x64x16 |
|
||||
| **TensorOp** | 8-by-8-by-32 | 32x32x32, 32x64x32, 64x32x32, 64x64x32 |
|
||||
| **TensorOp** | 16-by-8-by-32 | 32x32x32, 32x64x32, 64x32x32, 64x64x32 |
|
||||
| **TensorOp** | 16-by-8-by-64 | 32x32x64, 32x64x64, 64x32x64, 64x64x64 |
|
||||
| **TensorOp** | 8-by-8-by-128 | 32x32x128, 32x64x128, 64x32x128, 64x64x128 |
|
||||
| **TensorOp** | 16-by-8-by-256 | 32x32x256, 32x64x256, 64x32x256, 64x64x256 |
|
||||
| **SpTensorOp** | 16-by-8-by-16 | 64x64x16, 64x32x16, 32x64x16, 32x32x16 |
|
||||
| **SpTensorOp** | 16-by-8-by-32 | 64x64x32, 64x32x32, 32x64x32, 32x32x32 |
|
||||
| **SpTensorOp** | 16-by-8-by-64 | 64x64x64, 64x32x64, 32x64x64, 32x32x64 |
|
||||
| **SpTensorOp** | 16-by-8-by-128 | 64x64x128, 64x32x128, 32x64x128, 32x32x128 |
|
||||
|
||||
|
||||
TensorOp instructions depend on a permuted shared memory layout that can be efficiently
|
||||
loaded from. The following tables summarize the destination shared memory layout that
|
||||
can be targeted by matrix operands. It is assumed that each thread loads 128b vectors
|
||||
from global memory with layout specified in the column "GMEM Layout."
|
||||
|
||||
**TensorOp 8-by-8-by-4.**
|
||||
|
||||
|**Operand**|**Element** | **GMEM Layout** | **SMEM Layout** |
|
||||
|-----------|--------------|-----------------|-----------------------------------------|
|
||||
| **A** | `half_t` | `ColumnMajor` | `ColumnMajorVoltaTensorOpCongruous<16>` |
|
||||
| **A** | `half_t` | `RowMajor` | `RowMajorVoltaTensorOpCrosswise<16>` |
|
||||
| **B** | `half_t` | `ColumnMajor` | `ColumnMajorVoltaTensorOpCrosswise<16>` |
|
||||
| **B** | `half_t` | `RowMajor` | `RowMajorVoltaTensorOpCongruous<16>` |
|
||||
| **C** | `half_t` | `RowMajor` | `RowMajor` |
|
||||
| **C** | `float` | `RowMajor` | `RowMajor` |
|
||||
|
||||
**TensorOp 16-by-8-by-8.**
|
||||
|
||||
|**Operand**|**Element** | **GMEM Layout** | **SMEM Layout** |
|
||||
|-----------|--------------|-----------------|------------------------------------|
|
||||
| **A** | `half_t` | `ColumnMajor` | `ColumnMajorTensorOpCongruous<16>` |
|
||||
| **A** | `half_t` | `RowMajor` | `RowMajorTensorOpCrosswise<16>` |
|
||||
| **B** | `half_t` | `ColumnMajor` | `ColumnMajorTensorOpCrosswise<16>` |
|
||||
| **B** | `half_t` | `RowMajor` | `RowMajorTensorOpCongruous<16>` |
|
||||
| **C** | `half_t` | `RowMajor` | `RowMajor` |
|
||||
| **C** | `float` | `RowMajor` | `RowMajor` |
|
||||
|
||||
**TensorOp 16-by-8-by-8.**
|
||||
|
||||
|**Operand**|**Element** | **GMEM Layout** | **SMEM Layout** |
|
||||
|-----------|--------------|-----------------|------------------------------------|
|
||||
| **A** | `tfloat32_t` | `ColumnMajor` | `ColumnMajorTensorOpCongruous<32>` |
|
||||
| **A** | `tfloat32_t` | `RowMajor` | `RowMajorTensorOpCrosswise<32>` |
|
||||
| **B** | `tfloat32_t` | `ColumnMajor` | `ColumnMajorTensorOpCrosswise<32>` |
|
||||
| **B** | `tfloat32_t` | `RowMajor` | `RowMajorTensorOpCongruous<32>` |
|
||||
| **C** | `float` | `RowMajor` | `RowMajor` |
|
||||
|
||||
|
||||
**TensorOp 16-by-8-by-16.**
|
||||
|
||||
|**Operand**|**Element** | **GMEM Layout** | **SMEM Layout** |
|
||||
|-----------|--------------|-----------------|------------------------------------|
|
||||
| **A** | `half_t`, `bfloat16_t` | `ColumnMajor` | `ColumnMajorTensorOpCongruous<16>` |
|
||||
| **A** | `half_t`, `bfloat16_t` | `RowMajor` | `RowMajorTensorOpCrosswise<16>` |
|
||||
| **B** | `half_t`, `bfloat16_t` | `ColumnMajor` | `ColumnMajorTensorOpCrosswise<16>` |
|
||||
| **B** | `half_t`, `bfloat16_t` | `RowMajor` | `RowMajorTensorOpCongruous<16>` |
|
||||
| **C** | `half_t` | `RowMajor` | `RowMajor` |
|
||||
| **C** | `float` | `RowMajor` | `RowMajor` |
|
||||
|
||||
**TensorOp 8-by-8-by-4.**
|
||||
|
||||
|**Operand**|**Element** | **GMEM Layout** | **SMEM Layout** |
|
||||
|-----------|--------------|-----------------|------------------------------------|
|
||||
| **A** | `double` | `ColumnMajor` | `ColumnMajorTensorOpCongruous<64>` |
|
||||
| **A** | `double` | `RowMajor` | `RowMajorTensorOpCrosswise<64>` |
|
||||
| **B** | `double` | `ColumnMajor` | `ColumnMajorTensorOpCrosswise<64>` |
|
||||
| **B** | `double` | `RowMajor` | `RowMajorTensorOpCongruous<64>` |
|
||||
| **C** | `double` | `RowMajor` | `RowMajor` |
|
||||
|
||||
**TensorOp 8-by-8-by-16.**
|
||||
|
||||
|**Operand**|**Element** | **GMEM Layout** | **SMEM Layout** |
|
||||
|-----------|--------------|-----------------|------------------------------------|
|
||||
| **A** | `int8_t` | `RowMajor` | `RowMajorTensorOpCrosswise<8>` |
|
||||
| **B** | `int8_t` | `ColumnMajor` | `ColumnMajorTensorOpCongruous<8>` |
|
||||
| **C** | `int32_t` | `RowMajor` | `RowMajor` |
|
||||
|
||||
**TensorOp 16-by-8-by-32.**
|
||||
|
||||
|**Operand**|**Element** | **GMEM Layout** | **SMEM Layout** |
|
||||
|-----------|--------------|-----------------|------------------------------------|
|
||||
| **A** | `int8_t` | `RowMajor` | `RowMajorTensorOpCrosswise<8>` |
|
||||
| **B** | `int8_t` | `ColumnMajor` | `ColumnMajorTensorOpCongruous<8>` |
|
||||
| **C** | `int32_t` | `RowMajor` | `RowMajor` |
|
||||
|
||||
**TensorOp 8-by-8-by-32.**
|
||||
|
||||
|**Operand**|**Element** | **GMEM Layout** | **SMEM Layout** |
|
||||
|-----------|--------------|-----------------|------------------------------------|
|
||||
| **A** | `int4b_t` | `RowMajor` | `RowMajorTensorOpCrosswise<4>` |
|
||||
| **B** | `int4b_t` | `ColumnMajor` | `ColumnMajorTensorOpCongruous<4>` |
|
||||
| **C** | `int32_t` | `RowMajor` | `RowMajor` |
|
||||
|
||||
**TensorOp 16-by-8-by-64.**
|
||||
|
||||
|**Operand**|**Element** | **GMEM Layout** | **SMEM Layout** |
|
||||
|-----------|--------------|-----------------|------------------------------------|
|
||||
| **A** | `int4b_t` | `RowMajor` | `RowMajorTensorOpCrosswise<4>` |
|
||||
| **B** | `int4b_t` | `ColumnMajor` | `ColumnMajorTensorOpCongruous<4>` |
|
||||
| **C** | `int32_t` | `RowMajor` | `RowMajor` |
|
||||
|
||||
**TensorOp 8-by-8-by-128.**
|
||||
|
||||
|**Operand**|**Element** | **GMEM Layout** | **SMEM Layout** |
|
||||
|-----------|--------------|-----------------|------------------------------------|
|
||||
| **A** | `bin1_t` | `RowMajor` | `RowMajorTensorOpCrosswise<4>` |
|
||||
| **B** | `bin1_t` | `ColumnMajor` | `ColumnMajorTensorOpCongruous<4>` |
|
||||
| **C** | `int32_t` | `RowMajor` | `RowMajor` |
|
||||
|
||||
|
||||
**SpTensorOp 16-by-8-by-16.**
|
||||
|
||||
|**Operand**|**Element** | **GMEM Layout** | **SMEM Layout** |
|
||||
|-----------|--------------|-----------------|------------------------------------|
|
||||
| **A** | `tfloat32_t` | `RowMajor` | `RowMajorTensorOpCrosswise<32, 32>` |
|
||||
| **B** | `tfloat32_t` | `ColumnMajor` | `ColumnMajorTensorOpCrosswise<32, 32>`|
|
||||
| **C** | `float` | `RowMajor` | `RowMajor` |
|
||||
|
||||
**SpTensorOp 16-by-8-by-32.**
|
||||
|
||||
|**Operand**|**Element** | **GMEM Layout** | **SMEM Layout** |
|
||||
|-----------|--------------|-----------------|---------------------------------------|
|
||||
| **A** | `half_t` | `RowMajor` | `RowMajorTensorOpCrosswise<16, 64>` |
|
||||
| **B** | `half_t` | `ColumnMajor` | `ColumnMajorTensorOpCrosswise<16, 64>`|
|
||||
| **C** | `float` | `RowMajor` | `RowMajor` |
|
||||
|
||||
**SpTensorOp 16-by-8-by-64.**
|
||||
|
||||
|**Operand**|**Element** | **GMEM Layout** | **SMEM Layout** |
|
||||
|-----------|--------------|-----------------|---------------------------------------|
|
||||
| **A** | `int8_t` | `RowMajor` | `RowMajorTensorOpCrosswise<8, 128>` |
|
||||
| **B** | `int8_t` | `ColumnMajor` | `ColumnMajorTensorOpCrosswise<8, 128>`|
|
||||
| **C** | `int32_t` | `RowMajor` | `RowMajor` |
|
||||
|
||||
**SpTensorOp 16-by-8-by-128.**
|
||||
|
||||
|**Operand**|**Element** | **GMEM Layout** | **SMEM Layout** |
|
||||
|-----------|--------------|-----------------|------------------------------------|
|
||||
| **A** | `int4b_t` | `RowMajor` | `RowMajorTensorOpCrosswise<4, 256>` |
|
||||
| **B** | `int4b_t` | `ColumnMajor` | `ColumnMajorTensorOpCrosswise<4, 256>`|
|
||||
| **C** | `int32_t` | `RowMajor` | `RowMajor` |
|
||||
|
||||
|
||||
|
||||
## Warp-level Matrix Multiply with CUDA WMMA API
|
||||
|
||||
The following table summarizes supported warp level shapes for each WmmaTensorOp instruction.
|
||||
|
||||
|**Opcode Class** | **Instruction Shape** | **Warp Shapes** |
|
||||
|---------------------|-----------------------|--------------------------------------------|
|
||||
| **WmmaTensorOp** | 16-by-16-by-16 | 32x32x16, 32x64x16, 64x32x16 |
|
||||
| **WmmaTensorOp** | 8-by-32-by-16 | 32x32x16, 32x64x16, 64x32x16 |
|
||||
| **WmmaTensorOp** | 32-by-8-by-16 | 32x32x16, 32x64x16, 64x32x16 |
|
||||
| **WmmaTensorOp** | 8-by-8-by-32 | 32x32x32, 32x64x32, 64x32x32, 64x64x32 |
|
||||
| **WmmaTensorOp** | 8-by-8-by-128 | 32x32x128, 32x64x128, 64x32x128, 64x64x128 |
|
||||
|
||||
|
||||
CUDA exposes warp-level matrix operations in the CUDA C++ WMMA API. The CUDA C++ WMMA API exposes Tensor Cores via a set of functions and types in the `nvcuda::wmma` namespace. The functions and types in `nvcuda::wmma` provide target-independent APIs and implement architecture-specific tensor operation using TensorOp instruction underneath. CUTLASS exposes WMMA API through WmmaTensorOp. The WmmaTensorOp supports canonical shared memory layouts. The following table summarizes the destination shared memory layout that can be targeted by matrix operands. The WMMA API expects that matrices in shared memory loaded by `nvcuda::wmma::load_matrix_sync()` satisfy 128 bit alignment.
|
||||
|
||||
|
||||
**WmmaTensorOp (all matrix sizes and data types).**
|
||||
|
||||
|**Operand** | **GMEM Layout** | **SMEM Layout** |
|
||||
|------------|----------------------------|------------------------------|
|
||||
| **A** | `RowMajor`, `ColumnMajor` | `RowMajor`, `ColumnMajor` |
|
||||
| **B** | `RowMajor`, `ColumnMajor` | `RowMajor`, `ColumnMajor` |
|
||||
| **C** | `RowMajor`, `ColumnMajor` | `RowMajor`, `ColumnMajor` |
|
||||
|
||||
# Copyright
|
||||
|
||||
Copyright (c) 2017 - 2025 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
```
|
||||
Redistribution and use in source and binary forms, with or without
|
||||
modification, are permitted provided that the following conditions are met:
|
||||
|
||||
1. Redistributions of source code must retain the above copyright notice, this
|
||||
list of conditions and the following disclaimer.
|
||||
|
||||
2. Redistributions in binary form must reproduce the above copyright notice,
|
||||
this list of conditions and the following disclaimer in the documentation
|
||||
and/or other materials provided with the distribution.
|
||||
|
||||
3. Neither the name of the copyright holder nor the names of its
|
||||
contributors may be used to endorse or promote products derived from
|
||||
this software without specific prior written permission.
|
||||
|
||||
THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS "AS IS"
|
||||
AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE
|
||||
IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE ARE
|
||||
DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT HOLDER OR CONTRIBUTORS BE LIABLE
|
||||
FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL
|
||||
DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR
|
||||
SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER
|
||||
CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY,
|
||||
OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE
|
||||
OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
|
||||
```
|
||||
388
media/docs/cpp/fundamental_types.md
Normal file
388
media/docs/cpp/fundamental_types.md
Normal file
@@ -0,0 +1,388 @@
|
||||

|
||||
|
||||
# Fundamental Types
|
||||
|
||||
CUTLASS defies several fundamental numeric and container classes upon which computations and
|
||||
algorithms algorithms for linear algebra computations are implemented.
|
||||
|
||||
Where possible, CUTLASS fundamental types mirror the C++ Standard Library. However, there are circumstances that necessitate divergence from the Standard Library's specification. In such cases, the CUTLASS implementation adopts unique capitalization to distinguish that standard vocabulary types may not be safely substituted in all cases.
|
||||
|
||||
Most types in CUTLASS are usable in both host code and device code. Moreover, they are functional regardless of compute capability, but they may only be efficient when hardware support is present.
|
||||
|
||||
## Numeric Types
|
||||
|
||||
CUTLASS defines classes for the following numeric data types.
|
||||
|
||||
* `half_t`: IEEE half-precision floating point (exponent: 5b, mantissa: 10b; literal suffix `_hf`)
|
||||
* `bfloat16_t`: BFloat16 data type (exponent: 8b, mantissa: 7b; literal suffix `_bf16`)
|
||||
* `tfloat32_t`: Tensor Float 32 data type (exponent: 8b, mantissa: 10b; literal suffix `_tf32`)
|
||||
* `int4_t`, `uint4_t`: 4b signed and unsigned integer (literal suffx `_s4`, `_u4`)
|
||||
* `bin1_t`: 1b binary numeric type (literal suffix `_b1`)
|
||||
* `float_e5m2_t`: 8bits signed float (exponent: 5 bits, mantissa: 2 bits)
|
||||
* `float_e4m3_t`: 8bits signed float (exponent: 4 bits, mantissa: 3 bits)
|
||||
* `float_ue4m3_t`: 8bits unsigned float (exponent: 4 bits, mantissa: 3 bits)
|
||||
* `float_ue8m0_t`: 8bits unsigned float (exponent: 8 bits, mantissa: 0 bits)
|
||||
* `float_e3m2_t`: 6bits signed float (exponent: 3 bits, mantissa: 2 bits)
|
||||
* `float_e2m3_t`: 6bits signed float (exponent: 2 bits, mantissa: 3 bits)
|
||||
* `float_e2m1_t`: 4bits signed float (exponent: 2 bits, mantissa: 1 bits)
|
||||
* `type_erased_dynamic_float8_t`: Type agnostic 8 bits signed float allowing the user to provide a specific datatype as runtime argument.
|
||||
* `type_erased_dynamic_float6_t`: Type agnostic 6 bits signed float allowing the user to provide a specific datatype as runtime argument.
|
||||
* `type_erased_dynamic_float4_t`: Type agnostic 4 bits signed float allowing the user to provide a specific datatype as runtime argument.
|
||||
* `mx_float8_t<float_e5m2_t>` or `mx_float8_t<float_e4m3_t>` : Block scaled data type with fp8 element type and float_ue8m0_t scale factor and vector size of 32.
|
||||
* `mx_float6_t<float_e3m2_t>` or `mx_float6_t<float_e2m3_t>` : Block scaled data type with fp6 element type and float_ue8m0_t scale factor and vector size of 32.
|
||||
* `mx_float4_t<float_e2m1_t>` : Block scaled data type with signed e2m1 element type and float_ue8m0_t scale factor and vector size of 32.
|
||||
* `nv_float4_t<float_e2m1_t>` : Block scaled data type with signed e2m1 element type and float_ue4m3_t scale factor and vector size of 16.
|
||||
* `complex<T>`: defines complex-valued data type based on the supplied real-valued numeric type
|
||||
|
||||
Numeric types in CUTLASS may be used in both host and device code and are intended to function
|
||||
like any other plain-old-data type.
|
||||
|
||||
If CUTLASS is compiled with `CUTLASS_F16C_ENABLED`, then hardware conversion is used for
|
||||
half-precision types in host code. Regardless, `cutlass::half_t` uses the most efficient
|
||||
NVIDIA GPU hardware instructions available in device code.
|
||||
|
||||
Example:
|
||||
```c++
|
||||
#include <iostream>
|
||||
#include <cutlass/numeric_types.h>
|
||||
|
||||
__global__ void kernel(cutlass::half_t x) {
|
||||
printf("Device: %f\n", float(x * 2.0_hf));
|
||||
}
|
||||
|
||||
int main() {
|
||||
|
||||
cutlass::half_t x = 0.5_hf;
|
||||
|
||||
std::cin >> x;
|
||||
|
||||
std::cout << "Host: " << 2.0_hf * x << std::endl;
|
||||
|
||||
kernel<<< dim3(1,1), dim3(1,1,1) >>>(x);
|
||||
|
||||
return 0;
|
||||
}
|
||||
```
|
||||
|
||||
## Containers
|
||||
|
||||
CUTLASS uses the following containers extensively for implementing efficient CUDA kernels.
|
||||
|
||||
### Array
|
||||
|
||||
```c++
|
||||
template <
|
||||
typename T, // element type
|
||||
int N // number of elements
|
||||
>
|
||||
struct Array;
|
||||
```
|
||||
|
||||
`Array<class T, int N>` defines a statically sized array of elements of type _T_ and size _N_. This class is similar to
|
||||
[`std::array<>`](https://en.cppreference.com/w/cpp/container/array) in the Standard Library with one notable exception:
|
||||
partial specializations exist to pack or unpack elements smaller than one byte.
|
||||
|
||||
`Array<>` is intended to be a convenient and uniform container class to store arrays of numeric elements regardless of data type or vector length. The storage needed is expected to be the minimum necessary given the logical size of each numeric type in bits (numeric types smaller than one byte are densely packed). Nevertheless, the size reported by `sizeof(Array<T, N>)` is always an integer multiple of bytes.
|
||||
|
||||
Storing numeric elements in a C++ STL-style container class enables useful modern C++ mechanisms such as range-based for loops. For example, to print the elements of `Array<>`, the following range-based for loop syntax is always valid regardless of numeric data type, compute capability, or context in host or device code.
|
||||
|
||||
Example:
|
||||
```c++
|
||||
int const kN;
|
||||
Array<T, kN> elements;
|
||||
|
||||
CUTLASS_PRAGMA_UNROLL // required to ensure array remains in registers
|
||||
for (auto x : elements) {
|
||||
printf("%d, %f", int64_t(x), double(x)); // explictly convert to int64_t or double
|
||||
}
|
||||
```
|
||||
|
||||
When copying `Array<>` objects or passing them as arguments to methods, it is best to avoid accessing individual elements. This enables the use of vector instructions to perform the operation more efficiently. For example, setting all elements to zero is best performed by calling the `clear()` method. Copies should be performed by assigning the entire object.
|
||||
|
||||
Example:
|
||||
```c++
|
||||
#include <cutlass/array.h>
|
||||
|
||||
int const kN;
|
||||
Array<T, kN> source;
|
||||
Array<T, kN> destination;
|
||||
|
||||
source.clear(); // set all elements to value of zero
|
||||
|
||||
destination = source; // copy to `destination`
|
||||
```
|
||||
|
||||
`Array<>` may be used to store elements smaller than one byte such as 4b integers.
|
||||
```c++
|
||||
Array<int4b_t, 2> packed_integers;
|
||||
|
||||
static_assert(
|
||||
sizeof(packed_integers) == 1,
|
||||
"Packed storage of sub-byte data types is compact.");
|
||||
|
||||
// Access array elements using usual indirection and assignment operators
|
||||
packed_integers[0] = 2_s4;
|
||||
packed_integers[1] = 3_s4;
|
||||
|
||||
CUTLASS_PRAGMA_UNROLL
|
||||
for (auto x : elements) {
|
||||
printf("%d", int(x)); // access elements normally
|
||||
}
|
||||
|
||||
```
|
||||
|
||||
### AlignedArray
|
||||
|
||||
```c++
|
||||
template <
|
||||
typename T, // element type
|
||||
int N, // number of elements
|
||||
int Alignment // alignment requirement in bytes
|
||||
>
|
||||
class AlignedArray;
|
||||
```
|
||||
|
||||
`AlignedArray` is derived from `Array<T, N>` and supports an optional alignment field. Pointers to objects of type `AlignedArray<>` reliably yield vectorized memory accesses when dereferenced.
|
||||
|
||||
Example:
|
||||
```c++
|
||||
int const kN = 8;
|
||||
ArrayAligned<half_t, kN> source;
|
||||
ArrayAligned<half_t, kN> const *ptr = ...;
|
||||
|
||||
source = *ptr; // 128b aligned memory access
|
||||
```
|
||||
|
||||
### AlignedBuffer
|
||||
|
||||
```c++
|
||||
template <
|
||||
typename T, // element type
|
||||
int N, // number of elements
|
||||
int Alignment // alignment requirement in bytes
|
||||
>
|
||||
class AlignedBuffer;
|
||||
```
|
||||
|
||||
`AlignedBuffer` provides a uniform way to define aligned memory allocations for all data types. This is particularly
|
||||
useful in defining allocations within shared memory with guaranteed memory alignment needed for vectorized access.
|
||||
Note, constructors of the elements within AlignedBuffer<> are not called, and so the elements are initially in an
|
||||
undefined state.
|
||||
|
||||
Use `AlignedBuffer<>::data()` to obtain a pointer to the first element of the buffer.
|
||||
|
||||
**Example:** Guaranteed aligned shared memory allocation. Note, shared memory contents are uninitialized.
|
||||
```c++
|
||||
int const kN = 32;
|
||||
int const kAlignment = 16; // alignment in bytes
|
||||
|
||||
// Define a shared memory allocation in device code
|
||||
__shared__ AlignedBuffer<complex<half_t>, kN, kAlignment> matrix_tile;
|
||||
|
||||
complex<half_t> *ptr = matrix_tile.data(); // ptr is guaranteed to have 128b (16 Byte) alignment
|
||||
```
|
||||
|
||||
Note, `AlignedBuffer<>` only guarantees that its internal memory allocation is aligned, obtained by `AlignedBuffer<>::data()`. There is no guarantee that the `AlignedBuffer<>` object itself satisfies alignment constraints or that its internal memory allocation is contiguous. Device code performing vectorized memory accesses should use the `AlignedArray<>` type.
|
||||
|
||||
**_Example_:** Vectorized memory access to shared memory allocations.
|
||||
```c++
|
||||
int const kN = 1024;
|
||||
|
||||
__shared__ AlignedBuffer<half_t, kN> smem_buffer;
|
||||
|
||||
AlignedArray<half_t, 8> *ptr = reinterpret_cast<AlignedArray<half_t, 8> *>(smem_buffer.data());
|
||||
|
||||
AlignedArray<half_t, 8> x = ptr[threadIdx.x]; // 128b shared memory load
|
||||
```
|
||||
|
||||
### Numeric Conversion
|
||||
|
||||
CUTLASS defines procedures for performing numeric conversion between data types in `cutlass/numeric_conversion.h`.
|
||||
Where possible, these target hardware acceleration on the target architecture and support multiple rounding modes.
|
||||
|
||||
```c++
|
||||
#include "cutlass/numeric_conversion.h"
|
||||
#include "cutlass/numeric_types.h"
|
||||
|
||||
NumericConverter<half_t, float> convert_f32_to_f16;
|
||||
NumericConverter<tfloat32_t, float> convert_f32_to_tf32;
|
||||
|
||||
half_t x = convert_f32_to_f16(3.14159f);
|
||||
tfloat32_t y = convert_f32_to_tf32(3.14159f);
|
||||
```
|
||||
|
||||
Recent GPU architectures such as NVIDIA Turing and Ampere combine numeric conversion with efficient packing
|
||||
into bit vectors. Consequently, CUTLASS defines conversion on both scalars and `Array<>` objects to implement
|
||||
the optimal code sequence on all architectures.
|
||||
|
||||
```c++
|
||||
//
|
||||
// Example: convert and pack 32b signed integers to a vector of packed signed 8-bit integers.
|
||||
//
|
||||
int const kN = 16;
|
||||
Array<int8_t, kN> destination;
|
||||
Array<int, kN> source;
|
||||
|
||||
NumericConverter<descltype(destination), decltype(source)> convert;
|
||||
|
||||
destination = convert(source);
|
||||
```
|
||||
|
||||
### Coord
|
||||
|
||||
```c++
|
||||
template <
|
||||
int Rank,
|
||||
typename Index = int
|
||||
>
|
||||
class Coord;
|
||||
```
|
||||
|
||||
`Coord<Rank, class T = int>` is a container used explicitly for defining logical coordinates in tensors of known rank. Traditional vector operators are defined such as `+`, `-`, and scalar multiplication `*` to simplify the creation of vector-valued expressions on tensor coordinates.
|
||||
|
||||
**Example:** Vector operations on coordinates.
|
||||
```c++
|
||||
Coord<2> compute_offset(Coord<2> const & base) {
|
||||
|
||||
Coord<2> stride = make_Coord(1, kM);
|
||||
|
||||
return base + stride * make_Coord(threadIdx.x, threadIdx.y);
|
||||
}
|
||||
```
|
||||
|
||||
Instances of `Coord<>` are used throughout CUTLASS to compute indices into tensors. Frequently, the dimensions of tensors of known layouts may be given names such as "rows" or "columns". To clarify the code, we have implemented several classes derived from `Coord<>` with accessors for each coordinate member.
|
||||
|
||||
Such classes include:
|
||||
```c++
|
||||
struct MatrixCoord : public Coord<2> {
|
||||
Index & row();
|
||||
Index & column();
|
||||
};
|
||||
```
|
||||
and
|
||||
|
||||
```c++
|
||||
struct Tensor4DCoord : public Coord<4> {
|
||||
Index & n();
|
||||
Index & h();
|
||||
Index & w();
|
||||
Index & c();
|
||||
};
|
||||
```
|
||||
|
||||
### PredicateVector<int Bits>
|
||||
|
||||
`PredicateVector<int Bits>` contains a statically sized array of hardware predicates packed into registers to enable efficient access within unrolled loops.
|
||||
|
||||
This container is optimized for sequential access through iterators, though these are only efficient when used within fully unrolled loops.
|
||||
|
||||
Moreover, instances of `PredicateVector<>` are not guaranteed to be updated until any non-const iterator objects have gone out of scope. This is because iterators are effectively caches that update the `PredicateVector<>` instance's internal storage as a batch.
|
||||
|
||||
**Example:** Managing an array of predicates.
|
||||
```c++
|
||||
|
||||
unsigned mask;
|
||||
PredicateVector<kBits> predicates;
|
||||
|
||||
// Nested scope to update predicates via an iterator
|
||||
{
|
||||
auto pred_it = predicates.begin();
|
||||
|
||||
CUTLASS_PRAGMA_UNROLL
|
||||
for (int bit = 0; bit < kBits; ++bit, ++pred_it) {
|
||||
bool guard = (mask & (1u << bit));
|
||||
pred_it.set(guard);
|
||||
}
|
||||
}
|
||||
|
||||
// Efficient use of predicates to guard memory instructions
|
||||
T *ptr;
|
||||
Array<T, kAccesses> fragment;
|
||||
|
||||
auto pred_it = predicates.const_begin();
|
||||
for (int access = 0; access < kAccesses; ++access, ++pred_it) {
|
||||
if (*pred_it) {
|
||||
fragment[access] = ptr[access];
|
||||
}
|
||||
}
|
||||
|
||||
```
|
||||
|
||||
Note: `PredicateVector<>` is not efficient when accessed via dynamic random access. If an array of bits is needed with dynamic random access (in contrast with access via _constexpr_ indices), then `Array<bin1_t, N>` should be used instead.
|
||||
|
||||
## Functional
|
||||
|
||||
CUTLASS defines function objects corresponding to basic arithmetic operations modeled after C++ Standard Library's `<functional>` header.
|
||||
|
||||
CUTLASS extends this by defining `multiply_add<T>` which computes `d = a * b + c`. The partial specialization `multiply_add<complex<T>>` computes complex-valued multiplication and addition using four real-valued multiply-add operations; these may correspond to native hardware instructions.
|
||||
|
||||
Example:
|
||||
```c++
|
||||
complex<float> a;
|
||||
complex<float> b;
|
||||
complex<float> c;
|
||||
complex<float> d;
|
||||
|
||||
multiply_add<complex<float>> mad_op;
|
||||
|
||||
d = mad_op(a, b, c); // four single-precision multiply-add instructions
|
||||
```
|
||||
|
||||
CUTLASS defines partial specializations for type `Array<T, N>`, performing elementwise operations on each element. A further partial specialization for `Array<half_t, N>` targets may target native SIMD instructions for compute capability SM60 and beyond.
|
||||
|
||||
**Example:** Fused multiply-add of arrays of half-precision elements.
|
||||
```c++
|
||||
static int const kN = 8;
|
||||
|
||||
Array<half_t, kN> a;
|
||||
Array<half_t, kN> b;
|
||||
Array<half_t, kN> c;
|
||||
Array<half_t, kN> d;
|
||||
|
||||
multiply_add<Array<half_t, kN>> mad_op;
|
||||
|
||||
d = mad_op(a, b, c); // efficient multiply-add for Array of half-precision elements
|
||||
```
|
||||
|
||||
## Numeric Conversion
|
||||
|
||||
Operators are define to convert between numeric types in `numeric_conversion.h`. Conversion operators are defined in
|
||||
terms of individual numeric elements and on arrays which enable the possibility of efficient hardware
|
||||
support on current and future NVIDIA GPUs.
|
||||
|
||||
**Example:** Converting between 32-b and 8-b integers.
|
||||
```c++
|
||||
|
||||
```
|
||||
|
||||
# Copyright
|
||||
|
||||
Copyright (c) 2017 - 2025 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
```
|
||||
Redistribution and use in source and binary forms, with or without
|
||||
modification, are permitted provided that the following conditions are met:
|
||||
|
||||
1. Redistributions of source code must retain the above copyright notice, this
|
||||
list of conditions and the following disclaimer.
|
||||
|
||||
2. Redistributions in binary form must reproduce the above copyright notice,
|
||||
this list of conditions and the following disclaimer in the documentation
|
||||
and/or other materials provided with the distribution.
|
||||
|
||||
3. Neither the name of the copyright holder nor the names of its
|
||||
contributors may be used to endorse or promote products derived from
|
||||
this software without specific prior written permission.
|
||||
|
||||
THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS "AS IS"
|
||||
AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE
|
||||
IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE ARE
|
||||
DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT HOLDER OR CONTRIBUTORS BE LIABLE
|
||||
FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL
|
||||
DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR
|
||||
SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER
|
||||
CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY,
|
||||
OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE
|
||||
OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
|
||||
```
|
||||
571
media/docs/cpp/gemm_api.md
Normal file
571
media/docs/cpp/gemm_api.md
Normal file
@@ -0,0 +1,571 @@
|
||||

|
||||
|
||||
# CUTLASS GEMM API
|
||||
|
||||
CUTLASS presents a uniform programming model for matrix multiply-accumulate operations at each level of the hierarchy. This document
|
||||
focuses on device-level, threadblock-level GEMMs, warp-level GEMMs, thread-level GEMMs, and instruction-level GEMMs.
|
||||
|
||||
# CUTLASS GEMM Model
|
||||
|
||||
CUTLASS implements the basic GEMM triple loop nest with a tiled structure mirroring the execution model hierarchy.
|
||||
|
||||
The following pseudocode describes the model for a GEMM kernel targeting a warp-synchronous matrix multiply instruction like
|
||||
mma.sync. The entire operation is referred to as "Gemm," as it is assumed that an epilogue operation performs the general matrix
|
||||
update similar to BLAS.
|
||||
|
||||
```c++
|
||||
// cutlass::gemm::device::Gemm
|
||||
//
|
||||
for (int cta_n = 0; cta_n < GemmN; cta_n += CtaTileN) { // for each CTA } CTA-level concurrency
|
||||
for (int cta_m = 0; cta_m < GemmM; cta_m += CtaTileM) { // for each CTA }
|
||||
//
|
||||
// cutlass::gemm::threadblock::Mma
|
||||
//
|
||||
for (int cta_k = 0; cta_k < GemmK; cta_k += CtaTileK) { // "GEMM mainloop" - no unrolling - one iteration of this loop is one "stage"
|
||||
//
|
||||
for (int warp_n = 0; warp_n < CtaTileN; warp_n += WarpTileN) { // for each warp } warp-level concurrency
|
||||
for (int warp_m = 0; warp_m < CtaTileM; warp_m += WarpTileM) { // for each warp }
|
||||
//
|
||||
for (int warp_k = 0; warp_k < CtaTileK; warp_k += WarpTileK) { // fully unroll across CtaTileK - one iteration of this loop is one "k Group"
|
||||
//
|
||||
for (int mma_k = 0; mma_k < WarpTileK; mma_k += MmaK) { // cutlass::gemm::warp::Mma
|
||||
for (int mma_n = 0; mma_n < WarpTileN; mma_n += MmaN) { //
|
||||
for (int mma_m = 0; mma_m < WarpTileM; mma_m += MmaM) { //
|
||||
//
|
||||
mma_instruction(d, a, b, c); // cutlass::arch::mma - warp-wide matrix multiply instruction
|
||||
|
||||
} // for mma_m
|
||||
} // for mma_n
|
||||
} // for mma_k
|
||||
|
||||
} // for warp_k
|
||||
} // for warp_m
|
||||
} // for warp_n
|
||||
|
||||
} // for cta_k
|
||||
} // for cta_m
|
||||
} // for cta_n
|
||||
|
||||
```
|
||||
|
||||
The outer-most loops correspond to CTA-level hardware concurrency and are not explicitly written as loops in the code. These
|
||||
are implied by CUDA grid launch semantics.
|
||||
|
||||
The comment `cutlass::gemm::threadblock::Mma` refers to the threadblock-scoped matrix multiply-accumulate concept. This is
|
||||
the computation performed by one threadblock to compute a matrix product in registers. The "GEMM main loop" is listed.
|
||||
|
||||
The comment `cutlass::gemm::warp::Mma` refers to the computation performed by each warp. This is a nested loop executing a
|
||||
sequence of accumulated outer products.
|
||||
|
||||
The inner-most operation corresponds directly to hardware support. In this example, the nested structure terminates with
|
||||
warp-synchronous matrix multiply instructions targeting Tensor Cores.
|
||||
Alternatively, GEMMs targeting single-thread instructions may have an additional series of nested loops corresponding to
|
||||
thread-level concurrency.
|
||||
|
||||
# CUTLASS GEMM Components
|
||||
|
||||
This loop nest is expressed in CUTLASS via the following components which are specialized for data type, layout, and
|
||||
math instruction.
|
||||
|
||||

|
||||
|
||||
These components are described in the following sections.
|
||||
|
||||
## Device-wide GEMM API
|
||||
|
||||
The device-level GEMM API is intended to streamline instantiation and execution of the standard
|
||||
GEMM computation across the GPU. This operator is intended to be used in host-side .cu code and
|
||||
has semantics similar to cuBLAS.
|
||||
|
||||
The device-wide GEMM API is embodied by the following operators:
|
||||
- [cutlass::gemm::device::Gemm](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/gemm/device/gemm.h) - basic GEMM operation
|
||||
- [cutlass::gemm::device::GemmArray](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/gemm/device/gemm_array.h) - batched GEMM operation in which input matrices are read from arrays of pointers
|
||||
- [cutlass::gemm::device::GemmBatched](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/gemm/device/gemm_batched.h) - batched GEMM operation in which input matrices are separated by a constant stride
|
||||
- [cutlass::gemm::device::GemmSplitKParallel](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/gemm/device/gemm_splitk_parallel.h) - GEMM operation that partitions the GEMM K dimension then launches a separate reduction kernel
|
||||
|
||||
**Example:** launch a mixed-precision GEMM targeting Volta Tensor Cores.
|
||||
```c++
|
||||
using Gemm = cutlass::gemm::device::Gemm<
|
||||
cutlass::half_t, // ElementA
|
||||
cutlass::layout::ColumnMajor, // LayoutA
|
||||
cutlass::half_t, // ElementB
|
||||
cutlass::layout::ColumnMajor, // LayoutB
|
||||
cutlass::half_t, // ElementOutput
|
||||
cutlass::layout::ColumnMajor, // LayoutOutput
|
||||
float, // ElementAccumulator
|
||||
cutlass::arch::OpClassTensorOp, // tag indicating Tensor Cores
|
||||
cutlass::arch::Sm70 // tag indicating target GPU compute architecture
|
||||
>;
|
||||
|
||||
Gemm gemm_op;
|
||||
cutlass::Status status;
|
||||
|
||||
//
|
||||
// Launch GEMM on the device
|
||||
//
|
||||
|
||||
status = gemm_op({
|
||||
{m, n, k},
|
||||
{ptrA, lda},
|
||||
{ptrB, ldb},
|
||||
{ptrC, ldc},
|
||||
{ptrD, ldd},
|
||||
{alpha, beta}
|
||||
});
|
||||
|
||||
if (status != cutlass::Status::kSuccess) {
|
||||
return -1;
|
||||
}
|
||||
```
|
||||
|
||||
|
||||
## Threadblock-level GEMM API
|
||||
|
||||
GEMMs at this scope are expected to efficiently load tiles of data from global memory into internal storage and then compute matrix
|
||||
products with warp-level GEMM operators.
|
||||
|
||||
The threadblock-scoped matrix multiply operation is embodied by
|
||||
[cutlass::gemm::threadblock::MmaPipelined](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/gemm/threadblock/mma_pipelined.h).
|
||||
This is a class inspired by [std::transform_reduce()](https://en.cppreference.com/w/cpp/algorithm/transform_reduce)
|
||||
which computes the accumulated matrix product of a range of tiles defined by tile iterators.
|
||||
|
||||

|
||||
|
||||
In the case of GEMM, the tile iterators are
|
||||
[cutlass::transform::threadblock::PredicatedTileIterator](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/transform/threadblock/predicated_tile_iterator.h)
|
||||
to traverse a sequence of tiles in global memory with appropriate predication to avoid out-of-bounds
|
||||
memory accesses.
|
||||
|
||||
*Concept.* Threadblock-level matrix multiply accumulate operators are function objects satisfying the following concept.
|
||||
```c++
|
||||
struct Mma {
|
||||
/// Shape of warp-level matrix operation (concept: GemmShape)
|
||||
struct Shape;
|
||||
|
||||
/// Data type of multiplicand A (concept: numeric type)
|
||||
struct ElementA;
|
||||
|
||||
/// Layout of multiplicand A (concept: Layout)
|
||||
struct LayoutA;
|
||||
|
||||
/// Data type of multiplicand B (concept: numeric type)
|
||||
struct ElementB;
|
||||
|
||||
/// Layout of multiplicand B (concept: Layout)
|
||||
struct LayoutB;
|
||||
|
||||
/// Data type of accumulator matrix C (concept: numeric type)
|
||||
struct ElementC;
|
||||
|
||||
/// Layout of accumulator matrix C (concept: Layout)
|
||||
struct LayoutC;
|
||||
|
||||
/// Iterator of A operand in shared memory - satisfies: ReadableRandomAccessTileIteratorConcept
|
||||
struct IteratorA;
|
||||
|
||||
/// Fragment object loaded from IteratorA (concept: Array<ElementA, ..>)
|
||||
struct FragmentA;
|
||||
|
||||
/// Iterator of B operand in shared memory - satisfies: ReadableRandomAccessTileIteratorConcept
|
||||
struct IteratorB;
|
||||
|
||||
/// Fragment object loaded from IteratorB (concept: Array<ElementB, ..>)
|
||||
struct FragmentB;
|
||||
|
||||
/// Iterator of C operand in shared memory -
|
||||
/// satisfies: ReadableRandomAccessTileIteratorConcept | WriteableRandomAccessTileIteratorConcept
|
||||
struct IteratorC;
|
||||
|
||||
/// Fragment object loaded from IteratorC (concept: Array<ElementC, ..>)
|
||||
struct FragmentC;
|
||||
|
||||
/// Warp-level matrix multiply operator (concept: satisfies gemm::warp::Mma)
|
||||
struct Operator;
|
||||
|
||||
//
|
||||
// Method
|
||||
//
|
||||
|
||||
/// Computes a matrix product accumulated in D
|
||||
CUTLASS_DEVICE
|
||||
void operator()(
|
||||
FragmentC &D,
|
||||
IteratorA iter_A,
|
||||
IteratorB iter_B,
|
||||
FragmentC const &C);
|
||||
};
|
||||
```
|
||||
|
||||
## Warp-level Matrix Multiply API
|
||||
|
||||
Warp-level GEMM operators load tiles from shared memory into registers and then compute matrix multiplies using either
|
||||
Tensor Cores or CUDA Cores. The result is accumulated in a register tile. Iterators are defined for each
|
||||
operand `A`, `B`, and `C`.
|
||||
|
||||
The warp-level GEMM API is a generalization of CUDA's WMMA API to achieve the following objectives:
|
||||
|
||||
- native matrix multiply sizes of Tensor Cores
|
||||
- permuted shared memory layouts to ensure conflict-free accesses
|
||||
- pointer initilization outside of the mainloop
|
||||
- efficient traversal
|
||||
|
||||
Defining a warp-level matrix multiply in CUTLASS is similar to WMMA as shown below.
|
||||
|
||||

|
||||
|
||||
The usage model is also similar. The following example computes a warp-level GEMM operation,
|
||||
accumulating a series of matrix products in a register-backed array. The input to a warp-level
|
||||
GEMM operation in CUTLASS _must_ be data in shared memory loaded by iterators or on
|
||||
register-backed fragments.
|
||||
|
||||

|
||||
|
||||
```c++
|
||||
#include "cutlass/gemm/warp/default_mma_tensor_op.h"
|
||||
|
||||
using LayoutA = cutlass::layout::ColumnMajorTensorOpMultiplicandCongruous<
|
||||
cutlass::sizeof_bits<Element>::value, 64>;
|
||||
|
||||
using LayoutB = cutlass::layout::RowMajorTensorOpMultiplicandCongruous<
|
||||
cutlass::sizeof_bits<Element>::value, 64>;
|
||||
|
||||
using WarpMma = typename cutlass::gemm::warp::DefaultMmaTensorOp<
|
||||
cutlass::gemm::GemmShape<64, 64, 8>, // Overall warp-level GEMM operation
|
||||
cutlass::gemm::GemmShape<16, 8, 8>, // Target instruction
|
||||
cutlass::half_t, LayoutA, // operand A type and layout
|
||||
cutlass::half_t, LayoutB, // operand B type and layout
|
||||
float, // accumulator type
|
||||
cutlass::layout::RowMajor>::Type; // accumulator layout
|
||||
|
||||
//
|
||||
// Define a GEMM operation loading data from shared memory
|
||||
//
|
||||
int const kGemmK = 32;
|
||||
|
||||
__shared__ ElementA smem_buffer_A[WarpMma::Shape::kM * kGemmK];
|
||||
__shared__ ElementB smem_buffer_B[WarpMma::Shape::kN * kGemmK];
|
||||
|
||||
//
|
||||
// Construct iterators into SMEM tiles
|
||||
//
|
||||
|
||||
// leading dimensions inferred from matrix problem size
|
||||
int lda = WarpMma::Shape::kM;
|
||||
int ldb = WarpMma::Shape::kN;
|
||||
|
||||
// iterators into shared memory
|
||||
WarpMma::IteratorA warp_iterator_A({smem_buffer_A, lda});
|
||||
WarpMma::IteratorB warp_iterator_B({smem_buffer_B, ldb});
|
||||
|
||||
// Fragments in registers storing the operands
|
||||
FragmentA frag_A;
|
||||
FragmentB frag_B;
|
||||
FragmentC accum;
|
||||
|
||||
WarpMma mma;
|
||||
|
||||
accum.clear();
|
||||
|
||||
//
|
||||
// Accumulated outer product
|
||||
//
|
||||
|
||||
#pragma unroll 1
|
||||
for (int k = 0; k < kGemmK; k += WarpMma::Shape::kK) {
|
||||
|
||||
|
||||
iter_A.load(frag_A); // Load fragments from A and B matrices
|
||||
iter_B.load(frag_B);
|
||||
|
||||
++iter_A; ++iter_B; // Advance along GEMM K to next tile in A
|
||||
// and B matrices
|
||||
|
||||
// Compute matrix product
|
||||
mma(accum, frag_A, frag_B, accum);
|
||||
}
|
||||
```
|
||||
|
||||
*Concept.* Warp-level Mma operations are function objects satisfying the following concept.
|
||||
|
||||
```c++
|
||||
struct Mma {
|
||||
/// Shape of warp-level matrix operation (concept: GemmShape)
|
||||
struct Shape;
|
||||
|
||||
/// Data type of multiplicand A (concept: numeric type)
|
||||
struct ElementA;
|
||||
|
||||
/// Layout of multiplicand A (concept: Layout)
|
||||
struct LayoutA;
|
||||
|
||||
/// Data type of multiplicand B (concept: numeric type)
|
||||
struct ElementB;
|
||||
|
||||
/// Layout of multiplicand B (concept: Layout)
|
||||
struct LayoutB;
|
||||
|
||||
/// Data type of accumulator matrix C (concept: numeric type)
|
||||
struct ElementC;
|
||||
|
||||
/// Layout of accumulator matrix C (concept: Layout)
|
||||
struct LayoutC;
|
||||
|
||||
/// Iterator of A operand in shared memory - satisfies: ReadableRandomAccessTileIteratorConcept
|
||||
struct IteratorA;
|
||||
|
||||
/// Fragment object loaded from IteratorA (concept: Array<ElementA, ..>)
|
||||
struct FragmentA;
|
||||
|
||||
/// Iterator of B operand in shared memory - satisfies: ReadableRandomAccessTileIteratorConcept
|
||||
struct IteratorB;
|
||||
|
||||
/// Fragment object loaded from IteratorB (concept: Array<ElementB, ..>)
|
||||
struct FragmentB;
|
||||
|
||||
/// Iterator of C operand in shared memory -
|
||||
/// satisfies: ReadableRandomAccessTileIteratorConcept | WriteableRandomAccessTileIteratorConcept
|
||||
struct IteratorC;
|
||||
|
||||
/// Fragment object loaded from IteratorC (concept: Array<ElementC, ..>)
|
||||
struct FragmentC;
|
||||
|
||||
/// Indicates class of matrix operator (arch::OpClassSimt or arch::OpClassTensorOp)
|
||||
struct OperatorClass;
|
||||
|
||||
//
|
||||
// Methods
|
||||
//
|
||||
|
||||
/// Computes a matrix multiply-accumulate
|
||||
CUTLASS_DEVICE
|
||||
void operator()(
|
||||
FragmentC &D,
|
||||
IteratorA A,
|
||||
IteratorB B,
|
||||
FragmentC const &C);
|
||||
};
|
||||
```
|
||||
|
||||
|
||||
|
||||
*Tensor Core Operators.* Warp-level matrix multiply operators targeting Tensor Cores
|
||||
may be defined with the following template arguments. The `Policy` type specifies implementation-level details which may
|
||||
be used to affect performance or internal implementation of the warp-level operator.
|
||||
|
||||
```c++
|
||||
namespace cutlass {
|
||||
namespace gemm {
|
||||
namespace warp {
|
||||
|
||||
/// Structure to compute the matrix product targeting CUDA cores and SIMT math instructions.
|
||||
template <
|
||||
/// Size of the Gemm problem - concept: gemm::GemmShape<>
|
||||
typename Shape_,
|
||||
/// Data type of A elements
|
||||
typename ElementA_,
|
||||
/// Layout of A matrix (concept: MatrixLayout)
|
||||
typename LayoutA_,
|
||||
/// Data type of B elements
|
||||
typename ElementB_,
|
||||
/// Layout of B matrix (concept: MatrixLayout)
|
||||
typename LayoutB_,
|
||||
/// Element type of C matrix
|
||||
typename ElementC_,
|
||||
/// Layout of C matrix (concept: MatrixLayout)
|
||||
typename LayoutC_,
|
||||
/// Shape of the warp in units of thread (concept: MmaSimtPolicy)
|
||||
typename Policy_,
|
||||
/// Used for partial specialization
|
||||
typename Enable = bool
|
||||
>
|
||||
class MmaTensorOp {}
|
||||
|
||||
} // namespace warp
|
||||
} // namespace gemm
|
||||
} // namespace cutlass
|
||||
|
||||
```
|
||||
|
||||
*SIMT Math Instructions.* Warp-level matrix multiply operators targeting CUDA Cores
|
||||
may be defined with the following template arguments. The `Policy` type specifies implementation-level details which may
|
||||
be used to affect performance or internal implementation of the warp-level operator.
|
||||
|
||||
```c++
|
||||
/// Structure to compute the matrix product targeting CUDA cores and SIMT math instructions.
|
||||
template <
|
||||
/// Size of the Gemm problem - concept: gemm::GemmShape<>
|
||||
typename Shape_,
|
||||
/// Data type of A elements
|
||||
typename ElementA_,
|
||||
/// Layout of A matrix (concept: MatrixLayout)
|
||||
typename LayoutA_,
|
||||
/// Data type of B elements
|
||||
typename ElementB_,
|
||||
/// Layout of B matrix (concept: MatrixLayout)
|
||||
typename LayoutB_,
|
||||
/// Element type of C matrix
|
||||
typename ElementC_,
|
||||
/// Layout of C matrix (concept: MatrixLayout)
|
||||
typename LayoutC_,
|
||||
/// Shape of the warp in units of thread (concept: MmaSimtPolicy)
|
||||
typename Policy_,
|
||||
/// Used for partial specialization
|
||||
typename Enable = bool
|
||||
>
|
||||
class MmaSimt;
|
||||
```
|
||||
|
||||
|
||||
## Thread-level GEMM API
|
||||
|
||||
Thread-level GEMM operations perform matrix multiply-accumulate on data held in registers. These target CUDA Cores exclusively.
|
||||
|
||||
*Concept.* Thread-level matrix multiply operations are function objects satisfying the following concept.
|
||||
```c++
|
||||
struct Mma {
|
||||
|
||||
/// Shape of warp-level matrix operation (concept: GemmShape)
|
||||
struct Shape;
|
||||
|
||||
/// Data type of multiplicand A (concept: numeric type)
|
||||
struct ElementA;
|
||||
|
||||
/// Layout of multiplicand A (concept: Layout)
|
||||
struct LayoutA;
|
||||
|
||||
/// Fragment object loaded from IteratorA (concept: Array<ElementA, ..>)
|
||||
struct FragmentA;
|
||||
|
||||
/// Data type of multiplicand B (concept: numeric type)
|
||||
struct ElementB;
|
||||
|
||||
/// Layout of multiplicand B (concept: Layout)
|
||||
struct LayoutB;
|
||||
|
||||
/// Fragment object loaded from IteratorA (concept: Array<ElementB, ..>)
|
||||
struct FragmentB;
|
||||
|
||||
/// Data type of accumulator matrix C (concept: numeric type)
|
||||
struct ElementC;
|
||||
|
||||
/// Layout of accumulator matrix C (concept: Layout)
|
||||
struct LayoutC;
|
||||
|
||||
/// Fragment object loaded from IteratorA (concept: Array<ElementC, ..>)
|
||||
struct FragmentC;
|
||||
|
||||
//
|
||||
// Methods
|
||||
//
|
||||
|
||||
/// Computes a matrix multiply-accumulate
|
||||
CUTLASS_DEVICE
|
||||
void operator()(
|
||||
FragmentC &D,
|
||||
FragmentA const &A,
|
||||
FragmentB const &B,
|
||||
FragmentC const &C);
|
||||
};
|
||||
```
|
||||
|
||||
The CUTLASS thread-level GEMM template accepts the following template arguments.
|
||||
```c++
|
||||
namespace cutlass {
|
||||
namespace gemm {
|
||||
namespace thread {
|
||||
|
||||
/// Structure to compute the matrix product
|
||||
template <
|
||||
/// Size of the Gemm problem - concept: gemm::GemmShape<>
|
||||
typename Shape,
|
||||
/// Data type of A elements
|
||||
typename ElementA,
|
||||
/// Layout of A matrix (concept: MatrixLayout)
|
||||
typename LayoutA,
|
||||
/// Data type of B elements
|
||||
typename ElementB,
|
||||
/// Layout of B matrix (concept: MatrixLayout)
|
||||
typename LayoutB,
|
||||
/// Element type of C matrix
|
||||
typename ElementC,
|
||||
/// Layout of C matrix (concept: MatrixLayout)
|
||||
typename LayoutC,
|
||||
/// Concept: arch::OpMultiplyAdd or arch::Mma<>
|
||||
typename Operator = arch::OpMultiplyAdd,
|
||||
/// Used for partial specialization
|
||||
typename Enable = bool
|
||||
>
|
||||
struct Mma;
|
||||
|
||||
} // namespace thread
|
||||
} // namespace gemm
|
||||
} // namespace cutlass
|
||||
```
|
||||
|
||||
## Efficient Epilogue
|
||||
|
||||
CUTLASS GEMM operators perform mma followed by epilogue operation similar
|
||||
to cuBLAS. CUTLASS implements an efficient row-major epilogue. Thus, to achieve
|
||||
column-major GEMM, operands A & B are transposed and swapped.
|
||||
|
||||
To enable efficient row-major epilogue for both row-major and column-major output layout,
|
||||
CUTLASS' device-level GEMM operators `cutlass::device::Gemm` and `cutlass::device::GemmUniversal`
|
||||
provide two template definitions:
|
||||
- (a) [General definition](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/gemm/device/gemm.h#L217)
|
||||
- (b) [Specialized definition for column-major source/output](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/gemm/device/gemm.h#L545)
|
||||
|
||||
Efficient row-major epilogue for:
|
||||
- (i) GEMM operator on row-major source/output uses template (a). It runs row-major GEMM and
|
||||
an efficient row-major epilogue.
|
||||
- (ii) GEMM operator on column-major source/output uses template (b). It transposes and swaps
|
||||
operands A and B to enable efficient epilogue. `A x B = C => Transpose(B) x Transpose(A) = Transpose(C)`.
|
||||
For column-major source (C) matrix, Transpose(C) is row-major, and efficient epilogue works on
|
||||
row-major.
|
||||
|
||||
Note that cuBLAS typically expects a column-major source (C) and output matrix (D). Thus,
|
||||
CUTLASS library only instantiates and generates GEMM operatos with column-major layout. However,
|
||||
CUTLASS by itself can run both row-major and column-major output layouts for all combinations
|
||||
of input layouts. Thus, CUTLASS supports the following layout combinations for input and output layouts:
|
||||
|
||||
- `{N,T} x {N,T} => {N,T}` - NN, TN, TN, TT GEMM for both row-major and column-major output
|
||||
|
||||
## Instruction-level operations
|
||||
|
||||
CUTLASS defines a template-based interface to Tensor Core operations to avoid resorting
|
||||
to inline PTX.
|
||||
|
||||
- [mma_sm70.h](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/arch/mma_sm70.h) - Volta TensorCore operations
|
||||
- [mma_sm75.h](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/arch/mma_sm75.h) - Turing TensorCore operations
|
||||
|
||||
|
||||
# Copyright
|
||||
|
||||
Copyright (c) 2017 - 2025 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
```
|
||||
Redistribution and use in source and binary forms, with or without
|
||||
modification, are permitted provided that the following conditions are met:
|
||||
|
||||
1. Redistributions of source code must retain the above copyright notice, this
|
||||
list of conditions and the following disclaimer.
|
||||
|
||||
2. Redistributions in binary form must reproduce the above copyright notice,
|
||||
this list of conditions and the following disclaimer in the documentation
|
||||
and/or other materials provided with the distribution.
|
||||
|
||||
3. Neither the name of the copyright holder nor the names of its
|
||||
contributors may be used to endorse or promote products derived from
|
||||
this software without specific prior written permission.
|
||||
|
||||
THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS "AS IS"
|
||||
AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE
|
||||
IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE ARE
|
||||
DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT HOLDER OR CONTRIBUTORS BE LIABLE
|
||||
FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL
|
||||
DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR
|
||||
SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER
|
||||
CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY,
|
||||
OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE
|
||||
OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
|
||||
```
|
||||
703
media/docs/cpp/gemm_api_3x.md
Normal file
703
media/docs/cpp/gemm_api_3x.md
Normal file
@@ -0,0 +1,703 @@
|
||||

|
||||
|
||||
# CUTLASS 3.0 GEMM API
|
||||
|
||||
CUTLASS presents a uniform programming model
|
||||
for matrix multiply-accumulate (MMA) operations
|
||||
at different levels of the GPU system hierarchy.
|
||||
CUTLASS 3.0 has GEMM APIs corresponding to the following levels
|
||||
in order of highest to the lowest level.
|
||||
|
||||
1. Device
|
||||
2. Kernel
|
||||
3. Collective
|
||||
4. Tiled MMA and Copy
|
||||
5. Atom
|
||||
|
||||
This document will cover the first three levels in detail:
|
||||
Device, Kernel, and Collective.
|
||||
It also briefly discusses the Tiled MMA/Copy and Atom level,
|
||||
and then refers readers to CuTe's tutorial for more information.
|
||||
|
||||
# CUTLASS GEMM Model
|
||||
|
||||
CUTLASS implements algorithms that express
|
||||
the classical "triply nested loop" GEMM algorithm
|
||||
with a tiled structure mirroring the above hierarchy.
|
||||
|
||||
The following pseudocode describes the model for a GEMM kernel
|
||||
targeting a warp-synchronous matrix multiply instruction like `mma.sync.`
|
||||
The entire operation is referred to as "Gemm,"
|
||||
as it is assumed that an epilogue operation
|
||||
performs the general matrix update similar to BLAS.
|
||||
This is pseudocode and is only meant to illustrate which parts of the layers
|
||||
correspond to the inner or outer loops of the GEMM.
|
||||
|
||||
```c++
|
||||
// cutlass::gemm::kernel::GemmUniversal: ClusterTileM and ClusterTileN loops
|
||||
// are either rasterized by the hardware or scheduled by the kernel in persistent kernels.
|
||||
// Parallelism over thread block clusters
|
||||
for (int cluster_m = 0; cluster_m < GemmM; cluster_m += ClusterTileM) {
|
||||
for (int cluster_n = 0; cluster_n < GemmN; cluster_n += ClusterTileN) {
|
||||
|
||||
// cutlass::gemm::collective::CollectiveMma: mainloop that iterates over all k-tiles
|
||||
// No loop unrolling is performed at this stage
|
||||
for (int k_tile = 0; k_tile < size<2>(gmem_tensor_A); k_tile++) {
|
||||
|
||||
// loops inside cute::gemm(tiled_mma, a, b, c); Dispatch 5: (V,M,K) x (V,N,K) => (V,M,N)
|
||||
// TiledMma uses the hardware instruction provided through its Mma_Atom
|
||||
// TiledMma's atom layout, value layout, and permutations define the iteration order
|
||||
for (int tiled_mma_k = 0; tiled_mma_k < size<2>(A); tiled_mma_k++) {
|
||||
for (int tiled_mma_m = 0; tiled_mma_m < size<1>(A); tiled_mma_m++) {
|
||||
for (int tiled_mma_n = 0; tiled_mma_n < size<1>(B); tiled_mma_n++) {
|
||||
|
||||
// TiledMma's vector mode dispatches to the underlying instruction.
|
||||
mma.call(d, a, b, c);
|
||||
} // tiled_mma_n
|
||||
} // tiled_mma_m
|
||||
} // tiled_mma_k
|
||||
} // k_tile mainloop
|
||||
} // cluster_m
|
||||
} // cluster_n
|
||||
```
|
||||
|
||||
The first three nested `for` loops
|
||||
correspond to parallelism over thread block clusters.
|
||||
The code does not actually express them as explicit `for` loops.
|
||||
Instead, the parallelization scheme over tiles
|
||||
is implied by CUDA grid launch semantics.
|
||||
However, for persistent kernels,
|
||||
these three loops are expressed in the source code
|
||||
as a single `while` loop that queries the
|
||||
[work tile scheduler](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/gemm/kernel/sm90_tile_scheduler.hpp)
|
||||
for problem tiles on which to compute.
|
||||
|
||||
Inside the three nested `for` loops,
|
||||
one finds code that pulls matrix tiles
|
||||
from global memory into more "local" memory
|
||||
(like shared memory or registers)
|
||||
and computes MMAs.
|
||||
These tiled copy and tiled mma iterations are generally
|
||||
fully static and get fully unrolled.
|
||||
|
||||
# CUTLASS GEMM Components
|
||||
|
||||
CUTLASS expresses the above loop nest
|
||||
with the following components which are specialized for
|
||||
data type, layout, and math instruction.
|
||||
|
||||
| API level | API Class and/or function names |
|
||||
| --- | --- |
|
||||
| Device | `cutlass::gemm::device::GemmUniversalAdapter` |
|
||||
| Kernel | `cutlass::gemm::kernel::GemmUniversal` |
|
||||
| Collective | `cutlass::gemm::collective::CollectiveMma` <br /> `cutlass::epilogue::collective::DefaultEpilogue` <br /> `cutlass::epilogue::collective::Epilogue` <br /> |
|
||||
| Tiled (MMA and Copy) | `cute::TiledMma` and `cute::TiledCopy` <br /> `cute::gemm()` and `cute::copy()` |
|
||||
| Atom | `cute::Mma_Atom` and `cute::Copy_Atom` |
|
||||
|
||||
In CUTLASS 3.0, we assemble kernels
|
||||
by first composing a collective mainloop and collective epilogue
|
||||
together at the kernel layer,
|
||||
and then wrapping them with a host-side adapter
|
||||
to form a GEMM handle to that kernel.
|
||||
|
||||
The following sections describe these components
|
||||
in the order a user should instantiate them
|
||||
in order to assemble a kernel. This order is
|
||||
|
||||
1. assemble the required collective mainloop and epilogues,
|
||||
|
||||
2. compose them together to build a kernel type, and
|
||||
|
||||
3. wrap up the kernel with a device layer adapter.
|
||||
|
||||
This order is also reflected in the [CUTLASS 3.0 Hopper kernel examples](https://github.com/NVIDIA/cutlass/tree/main/examples/48_hopper_warp_specialized_gemm) as seen in the excerpt below.
|
||||
|
||||
```c++
|
||||
// Step 1: Generate the required collective layer mainloop specialization
|
||||
using CollectiveMainloop = typename cutlass::gemm::collective::CollectiveBuilder<
|
||||
ArchTag, OperatorClass,
|
||||
ElementA, LayoutA, AlignmentA,
|
||||
ElementB, LayoutB, AlignmentB,
|
||||
ElementAccumulator,
|
||||
TilesShape, ClusterShape,
|
||||
cutlass::gemm::collective::StageCountAuto,
|
||||
cutlass::gemm::collective::KernelScheduleAuto
|
||||
>::CollectiveOp;
|
||||
|
||||
// Step 2: Specify the collective layer epilogue type
|
||||
using CollectiveEpilogue = cutlass::epilogue::collective::DefaultEpilogue<
|
||||
ElementC,
|
||||
cutlass::gemm::TagToStrideC_t<LayoutC>,
|
||||
cutlass::gemm::TagToStrideC_t<LayoutC>,
|
||||
cutlass::epilogue::thread::LinearCombination<ElementC, 1, ElementAccumulator, ElementAccumulator>>;
|
||||
|
||||
// Step 3: Compose the mainloop and epilogue together at the kernel layer
|
||||
using GemmKernel = cutlass::gemm::kernel::GemmUniversal<
|
||||
cute::Shape<int,int,int,int>, // ProblemShape [M,N,K,L]
|
||||
CollectiveMainloop,
|
||||
CollectiveEpilogue
|
||||
>;
|
||||
|
||||
// Step 4: Wrap up the kernel::GemmUniversal kernel class
|
||||
// with the device adapter to obtain a host-side handle to the kernel
|
||||
using GemmHandle = cutlass::gemm::device::GemmUniversalAdapter<GemmKernel>;
|
||||
```
|
||||
|
||||
Towards the end, we also briefly cover CuTe's tiled mma and copy as well as the atom layer APIs,
|
||||
before redirecting users to CuTe-specific documentation for further details.
|
||||
|
||||
## Collective API
|
||||
|
||||
A Collective is "the largest collection of threads
|
||||
onto which mma atoms and copy atoms are tiled."
|
||||
That is, it is the largest number of threads in a grid
|
||||
that can cooperate by leveraging hardware features
|
||||
for accelerated communication and synchronization.
|
||||
These hardware features include
|
||||
|
||||
* asynchronous array copy
|
||||
(e.g., from global memory to shared memory);
|
||||
|
||||
* MMA instructions
|
||||
for small tiles that live in shared memory;
|
||||
|
||||
* synchronization operations for clusters,
|
||||
thread blocks, and/or warps; and/or
|
||||
|
||||
* hardware acceleration (such as barriers)
|
||||
for ensuring that data dependencies
|
||||
between asynchronous operations are met.
|
||||
|
||||
A Collective uses the `TiledMma` and `TiledCopy` API (see below)
|
||||
to access operations that copy and perform MMA on tiles.
|
||||
|
||||
Different units of parallelism
|
||||
(e.g., threads, warps, or thread blocks)
|
||||
in a Collective might have different roles.
|
||||
For example, in "warp-specialized" algorithms,
|
||||
some warps may be responsible for copying data,
|
||||
while others may be responsible for computation.
|
||||
Nevertheless, the different units of parallelism
|
||||
still need to share data and coordinate access
|
||||
to the shared data. For example,
|
||||
the producer warps in a warp-specialized algorithm
|
||||
that copy input matrix tiles into shared memory
|
||||
need to let the consumer MMA warp(s) know
|
||||
that their MMA inputs are ready.
|
||||
We contrast this with the `kernel::` layer API,
|
||||
which schedules the collectives over *independent* tiles in the grid.
|
||||
|
||||
The Collective API includes both the "mainloop"
|
||||
of matrix multiply-accumulate, and the epilogue.
|
||||
This API is the composition point for optimizations
|
||||
such as mainloop fusions and epilogue fusions.
|
||||
It is responsible for implementing
|
||||
the `k_tile` loop in the above triply nested loop pseudocode.
|
||||
|
||||
### Collective Mainloops
|
||||
|
||||
The `cutlass::gemm::collective::CollectiveMma` class
|
||||
is the primary interface to the collective
|
||||
matrix multiply-accumulate (MMA) mainloops.
|
||||
"Mainloop" refers to the "main loop" over tiles --
|
||||
the "cluster tile k" loop in the pseudocode
|
||||
near the top of this document.
|
||||
Any looping over multiple tiles that
|
||||
the algorithm might need to do would happen here.
|
||||
|
||||
The `CollectiveMma` class is declared in the header
|
||||
[cutlass/gemm/collective/collective_mma.hpp](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/gemm/collective/collective_mma.hpp).
|
||||
|
||||
```c++
|
||||
namespace cutlass::gemm::collective {
|
||||
|
||||
template <
|
||||
class DispatchPolicy,
|
||||
class TileShape,
|
||||
class ElementA,
|
||||
class StrideA,
|
||||
class ElementB,
|
||||
class StrideB,
|
||||
class TiledMma,
|
||||
class GmemTiledCopyA,
|
||||
class SmemLayoutAtomA,
|
||||
class SmemCopyAtomA,
|
||||
class TransformA,
|
||||
class GmemTiledCopyB,
|
||||
class SmemLayoutAtomB,
|
||||
class SmemCopyAtomB,
|
||||
class TransformB
|
||||
>
|
||||
struct CollectiveMma {
|
||||
static_assert(sizeof(ElementA) == 0, "Could not find a mainloop specialization.");
|
||||
};
|
||||
|
||||
} // namespace cutlass::gemm::collective
|
||||
```
|
||||
|
||||
- `DispatchPolicy` is the most important type for a collective, and is
|
||||
[covered in more detail below](#collective-dispatch-policies).
|
||||
|
||||
- `StrideA` and `StrideB` are instances of type `cute::Stride` that represent the global memory layout of A and B tensors. These strides are required to be rank-3, representing the modes `[outer, inner, batch]`. Each of the 3 ranks can be a multi-modal hierarchical stride; this would apply if implementing a tensor contraction.
|
||||
|
||||
- `TiledMma` is an instance of `cute::TiledMma`.
|
||||
|
||||
- `GmemTiledCopyA` and `GmemTiledCopyB` are instances of `cute::TiledCopy` types. Both tiled operation types are [covered in more detail below](#tiled-mma-and-copy).
|
||||
|
||||
- `SmemLayoutAtomA` and `SmemLayoutAtomB` are instances of type `cute::Layout` and represent the smallest
|
||||
layout that will get tiled over the entire collective's shared memory. This layout does _not_ include the
|
||||
pipeline mode, and therefore, both are expected to be rank 2 layouts of shape [`outer`, `inner`].
|
||||
|
||||
- `SmemCopyAtomA` and `SmemCopyAtomB` are `Copy_Atom`s to be used for moving data from shared memory
|
||||
into register memory.
|
||||
|
||||
Notice that CUTLASS 3.0 mainloops do not accept a dedicated accumulator element type.
|
||||
We obtain the accumulator type from the `typename TiledMma::ValTypeC`. Note also that
|
||||
top level API's `ElementA` and `ElementB` can differ from those of the MMA facing
|
||||
`typename TiledMma::ValTypeA` and `typename TiledMma::ValTypeB`, allowing TMA or user
|
||||
supplied transform operations to perform type conversions.
|
||||
|
||||
### Collective Dispatch Policies
|
||||
|
||||
`CollectiveMma` implementations are not generic.
|
||||
Instead, they must be specialized for each algorithm and GPU architecture.
|
||||
Users can dispatch to a `CollectiveMma` specialization
|
||||
by picking template arguments matching that specialization.
|
||||
CUTLASS 3.0 adopts a tag-based dispatch policy type to specialize
|
||||
mainloop implementations and add tuning knobs to them.
|
||||
|
||||
Below is an example of one of the dispatch policies that is used to dispatch to a Hopper TMA
|
||||
warp-specialized mainloop implementation:
|
||||
|
||||
```c++
|
||||
// n-buffer in smem (Hopper TMA),
|
||||
// pipelined with Hopper GMMA and TMA,
|
||||
// warp-specialized dynamic schedule
|
||||
template<
|
||||
int Stages_,
|
||||
class ClusterShape_ = Shape<_1,_1,_1>,
|
||||
class KernelSchedule = KernelTmaWarpSpecializedCooperative
|
||||
>
|
||||
struct MainloopSm90TmaGmmaWarpSpecialized {
|
||||
constexpr static int Stages = Stages_;
|
||||
using ClusterShape = ClusterShape_;
|
||||
using ArchTag = arch::Sm90;
|
||||
using Schedule = KernelSchedule;
|
||||
};
|
||||
```
|
||||
|
||||
The `Stages_` template parameter lets the user freely vary the number of pipeline stages,
|
||||
while the `ClusterShape_` type allows for parameterization over the shape of the threadblock
|
||||
cluster over which TMA multicast will take place.
|
||||
|
||||
The collective dispatch policy is also the primary point of composing various kernel schedules
|
||||
freely with any mainloop. Each mainloop policy either prescribes a `Schedule` with which
|
||||
it needs to be run, or exposes a template API that lets the user pick a subset of the following schedules:
|
||||
|
||||
```c++
|
||||
struct KernelCpAsyncWarpSpecialized { };
|
||||
struct KernelCpAsyncWarpSpecializedPingpong { };
|
||||
struct KernelCpAsyncWarpSpecializedCooperative { };
|
||||
struct KernelTma { };
|
||||
struct KernelTmaWarpSpecialized { };
|
||||
struct KernelTmaWarpSpecializedPingpong { };
|
||||
struct KernelTmaWarpSpecializedCooperative { };
|
||||
```
|
||||
|
||||
- A single kernel schedule can support multiple mainloop implementations. For example,
|
||||
`KernelMultistage` can be composed with many different mainloop implementations across GPU
|
||||
architectures such as `MainloopSm70TwoStage`, `MainloopSm80CpAsyncUnpredicated`, and many more.
|
||||
|
||||
- A single mainloop can be composed with multiple
|
||||
possible kernel schedules. For example, the `MainloopSm90TmaGmmaWarpSpecialized` can be
|
||||
composed with any of the `KernelTmaWarpSpecialized`, `KernelTmaWarpSpecializedPingpong` or `KernelTmaWarpSpecializedCooperative`
|
||||
kernel schedules.
|
||||
|
||||
As [discussed in the CUTLASS 3.0 design documentation](cutlass_3x_design.md), adopting tag
|
||||
dispatch policies for our core vocabulary types allows us to maintain a single type name for
|
||||
all operations that conceptually belong to the same class. This design has the following benefits.
|
||||
|
||||
- It *avoids code duplication* in cases where mainloops can be composed with multiple kernels or vice versa.
|
||||
- It *makes writing generic code easier*, as the primary type name `CollectiveMma` does not change across any implementation.
|
||||
- It *provides a clear, singular extension point* for users to plug in new, custom mainloops implementations specialized on their own dispatch policies.
|
||||
|
||||
### Collective Builder for `CollectiveMma`s
|
||||
|
||||
The primary `CollectiveMma` is intended to be an expert user interface that allows full control over
|
||||
all the properties of the collective's GPU micro-kernel. However, often a user just wants an
|
||||
off-the-shelf GEMM mainloop implementation parameterized on simple configuration parameters. CUTLASS 3.0
|
||||
provides [`cutlass::gemm::collective::CollectiveBuilder`](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/gemm/collective/collective_builder.hpp) for such scenarios.
|
||||
|
||||
```c++
|
||||
namespace cutlass::gemm::collective {
|
||||
template <
|
||||
class ArchTag,
|
||||
class OpClass,
|
||||
class ElementA,
|
||||
class GmemLayoutA,
|
||||
int AlignmentA,
|
||||
class ElementB,
|
||||
class GmemLayoutB,
|
||||
int AlignmentB,
|
||||
class ElementAccumulator,
|
||||
class TileShape_MNK,
|
||||
class ClusterShape_MNK,
|
||||
class StageCountType,
|
||||
class KernelScheduleType,
|
||||
class Enable = void
|
||||
>
|
||||
struct CollectiveBuilder {
|
||||
static_assert(sizeof(ElementA) == 0, "Could not build a collective for given parameters.");
|
||||
};
|
||||
} // namespace cutlass::gemm::collective
|
||||
```
|
||||
|
||||
`CollectiveBuilder` accepts CUTLASS 2.x equivalent input template arguments, and attempts to build
|
||||
the best performing `CollectiveMma` from the given parameters.
|
||||
|
||||
- `ArchTag` is one of the SM architectures tags from `cutlass::arch::Sm*`.
|
||||
- `OpClass` is one of the operator class tags from `cutlass::arch::OpClass*`.
|
||||
- `ElementA` and `ElementB` are the logical value types of the A resp. B tensors.
|
||||
- `ElementAccumulator` is the accumulator type to be used in the instruction.
|
||||
- `GmemLayoutA` and `GmemLayoutB` are CUTLASS 2.x layout tags, `layout::RowMajor` or `layout::ColumnMajor`.
|
||||
- `AlignmentA` and `AlignmentB` are global memory alignments of A and B tensors in terms of element count.
|
||||
- `TileShape_MNK` is an instance of `cute::Shape` that is rank-3, representing the MxNxK collective tile shape.
|
||||
- `ClusterShape_MNK` is an instance of `cute::Shape` that is rank-3, representing the MxNxK threadblock cluster tile shape.
|
||||
- `StageCountType` is either `collective::StageCountAuto` or an instance of `collective::StageCount<N>`.
|
||||
- `KernelScheduleType` is either `collective::KernelScheduleAuto` or one of the specific kernel schedule tags discussed in the [dispatch policy section](#collective-dispatch-policies) above.
|
||||
|
||||
`StageCountAuto` allows the collective builder to compute the size of a single stage's size in shared memory
|
||||
and maximize the shared memory usage assuming 1 threadblock / multiprocessor occupancy.
|
||||
|
||||
`KernelScheduleAuto` allows the collective builder to pick the best kernel schedule available for the
|
||||
given set of parameters, or let's the user override this with a specific kernel schedule type.
|
||||
|
||||
Note that collective builders are still in beta, and their functionality
|
||||
does not map onto the full design space that the primary expert `CollectiveMma` API
|
||||
allows for. We expect their supported mainloop types to expand in future releases, but
|
||||
with 3.0, only SM90 tensorop kernels are supported through the builder API. The builder API
|
||||
may also change in the future as we adopt user feedback.
|
||||
|
||||
If the builder is able to provide a collective mainloop type for the given set of parameters,
|
||||
it will be aliased within as `CollectiveOp`. For more information on how to
|
||||
parameterize kernels conveniently with the collective builder, please see example [49_hopper_gemm_with_collective_builder](https://github.com/NVIDIA/cutlass/tree/main/examples/49_hopper_gemm_with_collective_builder).
|
||||
|
||||
### Epilogue
|
||||
|
||||
The collective epilogue implements element-wise operations
|
||||
involving the output matrix. Users can provide a custom
|
||||
epilogue, or use one of the standard epilogues.
|
||||
These live in the directory
|
||||
[include/cutlass/epilogue/collective/](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/epilogue/collective/),
|
||||
and include classes like
|
||||
`cutlass::epilogue::collective::DefaultEpilogue`
|
||||
and
|
||||
`cutlass::epilogue::collective::Epilogue`.
|
||||
CUTLASS's provided collective epilogues
|
||||
do not live under `include/cutlass/gemm`
|
||||
or in the `cutlass::gemm` namespace,
|
||||
because they can be used for computations
|
||||
other than GEMM.
|
||||
|
||||
## Kernel API
|
||||
|
||||
The kernel is "a collection of all clusters in the grid."
|
||||
The kernel layer schedules have four main responsibilities.
|
||||
|
||||
- Ordering the execution of collectives within the kernel, performing any synchronization between that may be necessary
|
||||
- Marshalling the threads of a warp specialized schedules into their respective roles
|
||||
- Performing any necessary grid swizzling logic
|
||||
- Tiling the input tensors with the threadblock cluster value tile before invoking the collectives on them
|
||||
|
||||
The Kernel API is the entry point for a grid of thread blocks
|
||||
that may or may not be organized in a cluster.
|
||||
It is the composition point for fusing back-to-back GEMMs,
|
||||
epilogues, and/or other operations.
|
||||
|
||||
The entry point API for CUTLASS 3.0 kernel is the class
|
||||
`cutlass::gemm::kernel::GemmUniversal`, found in the header file
|
||||
[include/cutlass/gemm/kernel/gemm_universal.hpp](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/gemm/kernel/gemm_universal.hpp).
|
||||
`GemmUniversal` is a stateless universal device kernel
|
||||
that implements GEMM as the composition of two parts:
|
||||
|
||||
* a collective mainloop, and
|
||||
* a collective epilogue
|
||||
|
||||
```cpp
|
||||
namespace cutlass::gemm::kernel {
|
||||
/*
|
||||
* Stateless universal device GEMM kernel type that treats GEMM as
|
||||
* a composition of a collective mainloop and a collective epilogue.
|
||||
*
|
||||
* Supports both the 2.x and 3.x APIs based on whether the first type is
|
||||
* a cute::tuple<> or not.
|
||||
* 2.x API implementation: cutlass/gemm/kernel/gemm_universal.h
|
||||
* 3.x API implementation: cutlass/gemm/kernel/gemm_*.hpp
|
||||
*
|
||||
* In the following declaration, the name preceding the 'Or' refers to
|
||||
* 3.x API type argument order, and the name succeeding the 'Or' refers to
|
||||
* 2.x API type argument order. Template arguments without two names
|
||||
* belong to the 3.x API only.
|
||||
**/
|
||||
template <
|
||||
class ProblemShapeOrThreadblockMma_, // (m, n, k) or (m, n, k, l)
|
||||
class CollectiveMainloopOrEpilogue_,
|
||||
class CollectiveEpilogueOrThreadblockSwizzle_,
|
||||
class TileScheduler_ = void,
|
||||
class Enable = void
|
||||
>
|
||||
class GemmUniversal;
|
||||
} // namespace cutlass::gemm::kernel
|
||||
```
|
||||
|
||||
*Stateless* means that the caller --
|
||||
for example, the Device API described above --
|
||||
manages the kernel's state.
|
||||
The kernel just takes input and output parameters (`Params`).
|
||||
|
||||
*Universal* means that `GemmUniversal` works
|
||||
for both CUTLASS 3.0 and 2.x interfaces
|
||||
and across a broad range of kernel schedules.
|
||||
If `GemmUniversal`'s first template argument is a `cute::Shape`,
|
||||
then `GemmUniversal` assumes that the remaining template arguments
|
||||
implement the 3.0 APIs. Otherwise, `GemmUniversal` assumes that
|
||||
the remaining template arguments implement the 2.x APIs.
|
||||
Starting with CUTLASS 3.0, the problem shape has been promoted
|
||||
to a top-level template API for the GEMM kernel.
|
||||
This supports fully static GEMM instantiations
|
||||
where the user expects to know some or all
|
||||
of the problem shapes at compile time
|
||||
in order to extract even more performance.
|
||||
|
||||
The *collective mainloop* implements MMA on local tiles.
|
||||
The *collective epilogue* addresses any operations after the MMA,
|
||||
such as applying the `beta * C` part of `C := beta * C + alpha * A * B`.
|
||||
We will explain *collective* in more detail below.
|
||||
|
||||
Specializations of `kernel::GemmUniversal` for 3.0 APIs live in
|
||||
any of various `gemm_*.hpp` files in the directory
|
||||
[include/cutlass/gemm/kernel/](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/gemm/kernel/).
|
||||
Specializations for 2.x APIs can be found in the header file
|
||||
[include/cutlass/gemm/kernel/gemm_universal.h](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/gemm/kernel/gemm_universal.h).
|
||||
|
||||
CUTLASS 3.x implements various embodiments of `kernel::GemmUniversal`.
|
||||
Each kernel layer schedule is specialized
|
||||
for a GEMM scheduling algorithm and GPU architecture.
|
||||
Specializations of `kernel::GemmUniversal` for 3.0 APIs live in
|
||||
any of various `include/cutlass/gemm/kernel/{arch_tag}*.hpp` files in the directory
|
||||
[include/cutlass/gemm/kernel/](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/gemm/kernel/).
|
||||
Which specialization to dispatch to is decided through the dispatch policy's `Schedule` type.
|
||||
|
||||
For example, the header file
|
||||
[include/cutlass/gemm/kernel/sm90_gemm_tma_warpspecialized_pingpong.hpp](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/gemm/kernel/sm90_gemm_tma_warpspecialized_pingpong.hpp)
|
||||
has a specialization of `kernel::GemmUniversal` for Hopper
|
||||
that uses a warp-specialized mainloop with a persistent scheduling algorithm,
|
||||
while the header file
|
||||
[include/cutlass/gemm/kernel/sm90_gemm_tma_warpspecialized.hpp](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/gemm/kernel/sm90_gemm_tma_warpspecialized.hpp)
|
||||
has a specialization of `GemmUniversal` for Hopper
|
||||
that uses a warp-specialized but non-persistent algorithm.
|
||||
|
||||
To support composition between supported kernel schedules and mainloop dispatch policies without having to
|
||||
duplicate collective mainloop implementations, GEMM kernel layer schedules can be composed with
|
||||
any mainloop that specifies their corresponding kernel schedule as their `Schedule` type in the policy.
|
||||
This is discussed in detail in the [collective dispatch policy section](#collective-dispatch-policies) above.
|
||||
|
||||
```c++
|
||||
// An example of the SM90 KernelMultistage kernel's
|
||||
// specialization logic that allows it to be composed
|
||||
// with many mainloops such as `MainloopSm80CpAsync`
|
||||
// and `MainloopSm70TwoStage`.
|
||||
template <
|
||||
class ProblemShape_,
|
||||
class CollectiveMainloop_,
|
||||
class CollectiveEpilogue_,
|
||||
class TileScheduler_
|
||||
>
|
||||
class GemmUniversal<
|
||||
ProblemShape_,
|
||||
CollectiveMainloop_,
|
||||
CollectiveEpilogue_,
|
||||
TileScheduler_,
|
||||
std::enable_if_t<std::is_base_of_v<KernelMultistage, typename CollectiveMainloop_::DispatchPolicy::Schedule>>>
|
||||
```
|
||||
|
||||
## Device API
|
||||
|
||||
The Device API is a universal, kernel-agnostic host interface
|
||||
for kernel launch and managing the lifetime of
|
||||
reusable host-side parameters.
|
||||
|
||||
This API is how users' host-side .cu code
|
||||
invokes CUTLASS's single-GPU GEMM kernels.
|
||||
It serves the same purpose as cuBLAS and behaves similarly.
|
||||
|
||||
The entry point for the Device GEMM API is the class
|
||||
`cutlass::gemm::device::GemmUniversalAdapter`.
|
||||
This class lives in the header file
|
||||
[include/cutlass/gemm/device/gemm_universal_adapter.h](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/gemm/device/gemm_universal_adapter.h).
|
||||
`GemmUniversalAdapter` is a stateful, reusable handle,
|
||||
which is parameterized on the `cutlass::gemm::kernel` type.
|
||||
|
||||
```c++
|
||||
/*!
|
||||
GemmUniversalAdapter is a stateful, reusable GEMM handle built around a kernel
|
||||
of type cutlass::gemm::kernel::*
|
||||
|
||||
It manages the lifetime of the underlying `kernel::Params` struct, and exposes APIs
|
||||
to create it from the host facing arguments. For power users, new static methods
|
||||
are exposed in 3.x APIs that bypass the stateful methods or args->params lowering.
|
||||
|
||||
It supports kernel types that implement both the 2.x and 3.0 APIs,
|
||||
however, this is done by specializing the implementation of GemmUniversalAdapter
|
||||
on the two kernel API types, and thus, GemmUniversalAdapter's behavior might
|
||||
differ between the two specializations.
|
||||
*/
|
||||
template <class GemmKernel_, class Enable = void>
|
||||
class GemmUniversalAdapter;
|
||||
```
|
||||
|
||||
*Stateful* means that the handle instance contains state
|
||||
that the kernel needs to run.
|
||||
This means that the user must initialize the handle first,
|
||||
then use the initialized handle instance to run the kernel.
|
||||
Statefulness also means that the handle can manage the lifetime
|
||||
of the kernel's `Params` -- the parameters of the kernel itself.
|
||||
An important duty of `GemmUniversalAdapter`
|
||||
is to map from the user's `Arguments` --
|
||||
what the user sees as the kernel's parameters --
|
||||
to the `Params` that the kernel actually sees.
|
||||
For power users, the class exposes new static methods
|
||||
in 3.0 APIs that can bypass stateful methods
|
||||
or go directly to `Params` without intermediate `Arguments`.
|
||||
|
||||
*Reusable* means that the handle instance can be used
|
||||
to call the kernel multiple times with different arguments
|
||||
(e.g., different matrices).
|
||||
Reusing the handle may be more efficient than just
|
||||
creating a new handle for each kernel invocation.
|
||||
|
||||
*Parameterized on the kernel type* means that
|
||||
the `GemmUniversalAdapter` class' behavior
|
||||
depends on the GEMM kernel type (see the next section).
|
||||
Specifically, `GemmUniversalAdapter` has a template parameter
|
||||
`GemmKernel`, which is the GEMM kernel type.
|
||||
Valid template arguments for `GemmKernel` are
|
||||
|
||||
* `cutlass::gemm::kernel::GemmUniversal`,
|
||||
implementing CUTLASS 3.x API kernels;
|
||||
* `cutlass::gemm::kernel::GemmUniversal`,
|
||||
implementing CUTLASS 2.x API kernels; or
|
||||
* Any valid CUTLASS 2.x `kernel` layer GEMM that
|
||||
was previously composable with the `device::GemmUniversalAdapter`.
|
||||
|
||||
`GemmUniversalAdapter` presents a single
|
||||
host-side interface to both 3.0 and 2.x kernels.
|
||||
CUTLASS accomplishes this by
|
||||
specializing `GemmUniversalAdapter`'s implementation
|
||||
on either the 2.x API implementing kernel layer GEMMs, or on the 3.x API
|
||||
implementing kernel layer GEMMs. The metafunction [`cutlass::gemm::detail::IsCutlass3GemmKernel`](cutlass_3x_backwards_compatibility.md#kernel-api-design-differences)
|
||||
is what `GemmUniversalAdapter` uses to distinguish between 2.x and 3.x kernels.
|
||||
|
||||
`GemmUniversalAdapter` sets up and launches the kernel, using the
|
||||
CUDA extended launch API for threadblock cluster support if required.
|
||||
Note, `GemmUniversalAdapter` does *not* specify the grid shape.
|
||||
The kernel controls the grid shape
|
||||
and other kernel-specific launch parameters.
|
||||
This makes it possible for all 3.0 kernels
|
||||
to use the same kernel launch code,
|
||||
thus factoring out kernel launch from the actual kernel.
|
||||
|
||||
## Tiled MMA and Copy
|
||||
|
||||
The Tiled MMA or Copy are tilings of MMA atoms resp. Copy atoms
|
||||
across threads and data, with possible permutations applied to the
|
||||
resulting tiling. This layer is most analogous to the warp level
|
||||
tiling of MMA instructions in CUTLASS 2.x. However, it views the tiling
|
||||
from the perspective of all threads participating in the operation
|
||||
and generalizes the concept to copy operations as well. The purpose
|
||||
of this layer is to build composable GPU micro-kernels out of a plethora
|
||||
of hardware accelerated math and data movement operations, each with their
|
||||
unit layouts in threads and data. The tiled MMA and Copy types present
|
||||
all these various hardware accelerated CuTe Atoms with a single, consistent
|
||||
API.
|
||||
|
||||
The resulting tiled operation acts as a single MMA or copy operation
|
||||
that users can invoke in the "inner" loop
|
||||
of the three-nested-loops pseudocode
|
||||
at the top of this document using `cute::gemm()` or `cute::copy()`.
|
||||
|
||||
We call this API "tiled" because it constructs
|
||||
larger operations out of the Atoms provided by CuTe,
|
||||
as if fitting together individual tiles
|
||||
to build a reusable component of a mosaic.
|
||||
For example, CuTe might provide an MMA Atom
|
||||
that users can call on a single warp,
|
||||
for fixed M, N, and K dimensions.
|
||||
CUTLASS can then use CuTe operations like `make_tiled_mma`
|
||||
to turn this Atom into an operation
|
||||
that works on an entire thread block,
|
||||
for larger M, N, and K dimensions.
|
||||
|
||||
## Atom API
|
||||
|
||||
An "Atom" is the smallest collection of threads and data
|
||||
that must participate in the execution of a hardware-accelerated
|
||||
math or copy operation.
|
||||
|
||||
An Atom is "atomic" (indivisible) not in the sense of
|
||||
concurrent memory operations like `atomicAdd`
|
||||
(which are "indivisible in time (causality)"),
|
||||
but in the sense of indivisibility in "space" --
|
||||
the number of values and the groups of parallel workers
|
||||
that must participate in the operation together.
|
||||
|
||||
An Atom uses CuTe Layouts to express the required
|
||||
dimensions and strides of its input and output arrays.
|
||||
Generally these are fixed at compile time.
|
||||
|
||||
The Atom API wraps calls to actual hardware instructions
|
||||
that accelerate MMA or copy operations.
|
||||
Users can ask for GPU architecture-specific implementations,
|
||||
or just pick generic implementations and rely on
|
||||
whatever GPU architectures were enabled.
|
||||
|
||||
For more information about Atoms,
|
||||
please refer to CuTe's tutorial, e.g., the sections on
|
||||
|
||||
* [algorithms](./cute/04_algorithms.md) like `gemm` and `copy`,
|
||||
|
||||
* [MMA Atoms](./cute/0t_mma_atom.md#cute-mma-atoms), and
|
||||
|
||||
* [a GEMM example](./cute/0x_gemm_tutorial.md).
|
||||
|
||||
# Copyright
|
||||
|
||||
Copyright (c) 2023 - 2025 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
```
|
||||
Redistribution and use in source and binary forms, with or without
|
||||
modification, are permitted provided that the following conditions are met:
|
||||
|
||||
1. Redistributions of source code must retain the above copyright notice, this
|
||||
list of conditions and the following disclaimer.
|
||||
|
||||
2. Redistributions in binary form must reproduce the above copyright notice,
|
||||
this list of conditions and the following disclaimer in the documentation
|
||||
and/or other materials provided with the distribution.
|
||||
|
||||
3. Neither the name of the copyright holder nor the names of its
|
||||
contributors may be used to endorse or promote products derived from
|
||||
this software without specific prior written permission.
|
||||
|
||||
THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS "AS IS"
|
||||
AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE
|
||||
IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE ARE
|
||||
DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT HOLDER OR CONTRIBUTORS BE LIABLE
|
||||
FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL
|
||||
DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR
|
||||
SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER
|
||||
CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY,
|
||||
OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE
|
||||
OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
|
||||
```
|
||||
418
media/docs/cpp/grouped_scheduler.md
Normal file
418
media/docs/cpp/grouped_scheduler.md
Normal file
@@ -0,0 +1,418 @@
|
||||

|
||||
|
||||
# CUTLASS Grouped Kernel Schedulers
|
||||
|
||||
CUTLASS's grouped kernel is a persistent kernel which launches multiple problems (e.g., GEMMs, SYR2Ks) within a
|
||||
single CUDA kernel launch.
|
||||
|
||||
Unlike a conventional GEMMs in CUTLASS, which launch a number of threadblocks equal to the number
|
||||
of tiles in the GEMM, CUTLASS grouped kernels typically launch a number of threadblocks that is
|
||||
fewer than the total number of tiles across all problems in the group. Each threadblock is then
|
||||
responsible for computing one or more tiles among the problems in the group. The grouped kernel
|
||||
_scheduler_ (referred to as the _problem visitor_ in code) is responsible for assigning each
|
||||
threadblock the sequence of tiles that it will compute within the group.
|
||||
|
||||
This document provides background on the functionality of the grouped kernel scheduler, and describes
|
||||
various optimizations to the grouped kernel scheduler.
|
||||
|
||||
**Outline**
|
||||
|
||||
* [Introduction to Grouped Kernel Schedulers](grouped_scheduler.md#introduction-to-grouped-kernel-schedulers)
|
||||
* [Grouped GEMM Scheduler](grouped_scheduler.md#grouped-gemm-scheduler)
|
||||
* [Grouped Rank2K Scheduler](grouped_scheduler.md#grouped-rank2k-scheduler)
|
||||
* [Scheduler Modes](grouped_scheduler.md#scheduler-modes)
|
||||
* [Improving Load Balance by Sorting Problems](grouped_scheduler.md#improving-load-balance-by-sorting-problems)
|
||||
|
||||
# Introduction to Grouped Kernel Schedulers
|
||||
Given a group of problem sizes and a grid of threadblocks, the scheduler's job is to assign
|
||||
tiles from problems in the group to threadblocks. Threadblocks in a grouped kernel persistently
|
||||
execute a loop of querying the scheduler for the next tile to compute and performing the
|
||||
kernel-level operations for that tile (e.g., MMA and epilogue). In pseudocode, this looks as
|
||||
follows:
|
||||
```c++
|
||||
ProblemVisitor problem_visitor;
|
||||
|
||||
while (problem_visitor.next_tile()) {
|
||||
//
|
||||
// Get next tile index from scheduler
|
||||
//
|
||||
|
||||
//
|
||||
// Compute MMA and epilogue
|
||||
//
|
||||
|
||||
// Inform the scheduler that we are done with the current tile
|
||||
problem_visitor.advance(gridDim.x);
|
||||
}
|
||||
```
|
||||
|
||||
The key functionality of the grouped kernel scheduler lies in the `next_tile()` method,
|
||||
which determines which tile in the group the calling threadblock should compute next, if any.
|
||||
|
||||
# Grouped GEMM Scheduler
|
||||
The scheduler used by grouped GEMM assigns tiles in the group to threadblocks in a round-robin
|
||||
fashion.
|
||||
|
||||
Consider, for example, the threadblock-to-tile mapping that occurs for a group of four GEMMs
|
||||
each consisting of a grid of 2x2 tiles. Suppose that eight threadblocks are launched. The
|
||||
figure below illustrates the threadblock ID assigned to each tile in each GEMM in the group.
|
||||
|
||||

|
||||
|
||||
A similar mapping for problems that do not have the same number of tiles
|
||||
is shown below:
|
||||
|
||||

|
||||
|
||||
## Computing the schedule for a given block
|
||||
Each threadblock in the grouped GEMM computes its own schedule by calling
|
||||
the `next_tile()` method described above.
|
||||
|
||||
To do this, the threadblock's `ProblemVisitor` maintains a `thread_idx`
|
||||
member that is initialized to `blockIdx.x` and is incremented by
|
||||
`gridDim.x` between each tile computed (only the x dimension is used)
|
||||
in the launch configuration for grouped kernels). The scheduler must
|
||||
then figure out which GEMM in the group `tile_idx` belongs to, and which tile
|
||||
within that problem it maps to.
|
||||
|
||||
1. **Determining which GEMM `tile_idx` maps to:** The scheduler determines
|
||||
the GEMM to which `tile_idx` belongs by iterating through GEMMs starting with
|
||||
the most-recently visited GEMM, and adding the number of tiles within that
|
||||
GEMM to a running variable `problem_tile_start`. The scheduler has found the
|
||||
correct problem for this tile when `problem_tile_start <= tile_idx < problem_tile_start + tiles_in_problem`.
|
||||
|
||||
2. **Determining the tile within a GEMM `tile_idx` maps to:** Once the GEMM
|
||||
to which `tile_idx` maps has been located, the specific tile within that
|
||||
GEMM that this block should compute is given by `tile_idx - problem_tile_start`.
|
||||
Simple rasterization is then performed to map this one-dimensional tile ID
|
||||
into the two-dimensional coordinate of the tile to compute in the GEMM.
|
||||
|
||||
We describe how this search is accelerated in [Scheduler Modes](grouped_scheduler.md#scheduler-modes).
|
||||
|
||||
# Grouped Rank2K Scheduler
|
||||
The previous section described the operation of the scheduler used
|
||||
for grouped GEMM kernels. While this scheduler is sufficient for
|
||||
correctly implementing grouped Rank2K operations (i.e., SYR2K and HER2K), it leads to significant inefficiencies.
|
||||
|
||||
We next describe these inefficiencies as well as how the CUTLASS
|
||||
grouped Rank2K scheduler overcomes them.
|
||||
|
||||
## Inefficiency of grouped GEMM scheduler for grouped Rank2K problems
|
||||
The grouped GEMM scheduler assumes that every tile in every GEMM in the group will
|
||||
ultimately affect the output of the problem. This is not the case for Rank2K
|
||||
problems, for which matrix C is either upper or lower triangular. Using the default
|
||||
grouped GEMM scheduler for such problems will thus lead to threadblocks frequently
|
||||
being assigned to tiles that exit early (e.g., due to being assigned to a tile in the
|
||||
upper-triangular portion of a lower-triangular problem). This further leads to load
|
||||
imbalance among threadblocks, as the grouped GEMM scheduler assigns nearly the same
|
||||
number of tiles to all threadblocks, regardless of how many tiles are truly active.
|
||||
|
||||
Consider an example of a group of four SYR2K problems, each with matrix C consisting
|
||||
of a grid of 2x2 tiles. Matrix C in each problem is lower triangular, indicated by
|
||||
shaded tiles. Consider that eight threadblocks are launched to compute the grouped
|
||||
problem. The default grouped GEMM scheduler will assign threadblocks to tiles in the following order:
|
||||
|
||||

|
||||
|
||||
In this case, threadblocks 1 and 5 are continuously assigned to inactive tiles. In
|
||||
scenarios in which problems within the group have varying size, we have observed
|
||||
this to still lead to significant load imbalance.
|
||||
|
||||
## Specializing the scheduler for triangular problems
|
||||
We seek to design a scheduler that more efficiently maps threadblocks to active tiles
|
||||
for kernels that use triangular output matrices. The scheduler should ideally assign
|
||||
threadblocks only to those tiles within lower-triangular portion of a
|
||||
lower-triangular problem (and vice-versa for upper-triangular problems).
|
||||
|
||||
Using the example above, the resulting assignment of threadblocks to tiles from
|
||||
such a scheduler might be:
|
||||
|
||||

|
||||
|
||||
Achieving this schedule requires mapping from a threadblock ID to tile coordinates
|
||||
`(i, j)`.
|
||||
|
||||
We will illustrate this by mapping a lower-triangular matrix with a 3x3 grid. We
|
||||
first calculate row and column indices assuming one-indexed rows, tiles, and
|
||||
threadblock IDs, and then subtract one to convert to zero-indexed versions. Our
|
||||
description borrows heavily from the mapping described [here](https://stackoverflow.com/a/40954159).
|
||||
|
||||

|
||||
|
||||
### Calculating row `i` given threadblock ID `t`
|
||||
For a given row i, all threadblock IDs t in that row satisfy the following:
|
||||
```
|
||||
t <= 1 + 2 + 3 + ... + (i-1) + i
|
||||
```
|
||||
|
||||
The closed-form equation for the right-hand side is: `i(i+1)/2`.
|
||||
Using this, we can solve for `i` given `t`:
|
||||
```
|
||||
t <= i(i+1)/2
|
||||
2t <= i^2 + i
|
||||
2t <= i^2 + i + 0.25 - 0.25
|
||||
2t + 0.25 <= i^2 + i + 0.25
|
||||
2t + 0.25 <= (i + 0.5)^2
|
||||
sqrt(2t + 0.25) - 0.5 <= i
|
||||
```
|
||||
|
||||
To account for fractional values, we set:
|
||||
```
|
||||
i = ceil(sqrt(2t + 0.25) - 0.5)
|
||||
```
|
||||
|
||||
To turn this into a zero-indexed row and work with zero-indexed `t`, we perform:
|
||||
```
|
||||
i = ceil(sqrt(2(t+1) + 0.25) - 0.5) - 1
|
||||
= ceil(sqrt(2t + 2.25) - 0.5) - 1
|
||||
```
|
||||
|
||||
### Calculating column `j` given threadblock ID `t` and row `i`
|
||||
For a given row `i`, all threadblock IDs `t` in that row also satisfy the following:
|
||||
```
|
||||
t > 1 + 2 + 3 + ... + (i-2) + (i-1)
|
||||
--> t > i(i-1)/2
|
||||
```
|
||||
|
||||
Threadblock IDs within a given row are sequential, so the one-indexed column ID
|
||||
for one-indexed threadblock ID `t` and row `i` is:
|
||||
```
|
||||
j = t - (i(i-1)/2)
|
||||
```
|
||||
|
||||
The zero-indexed version becomes:
|
||||
```
|
||||
j = (t+1) - (i(i+1)/2) -1
|
||||
= t - (i(i+1)/2)
|
||||
```
|
||||
|
||||
### Accounting for non-square grids
|
||||
Though the overall output problem size for Rank2K problems is guaranteed to be square, the
|
||||
grids used in computing may not be square due to using non-square threadblock shapes. For
|
||||
example, a threadblock shape of 64x32 operating on a problem of output size 128x128 would
|
||||
result in a grid of 2x4 tiles.
|
||||
|
||||
This case can be handled by noting that the output resembles a square grid of 2x2 "macro tiles"
|
||||
each of which contains 2 "true tiles." We can thus first map a threadblock ID to its "macro tile"
|
||||
using the equations above, and then map it to the "true tile" within its "macro tile." In the example
|
||||
of a 2x4 grid, this mapping would look as follows:
|
||||
|
||||

|
||||
|
||||
A zero-indexed threadblock ID `t` is mapped to its "macro tile ID" `t_macro` as:
|
||||
```
|
||||
t_macro = t // r
|
||||
```
|
||||
Where `r` is the ratio of the maximum dimension of the grid to the
|
||||
minimum dimension of the grid (i.e., `r = 4 / 2 = 2` in the previous example).
|
||||
|
||||
One uses `t_macro` and the calculations above to find the row and column in the square matrix to
|
||||
obtain `i_macro` and `j_macro` (zero-indexed). The mapping from `(i_macro, j_macro) --> (i, j)`
|
||||
is simply the following:
|
||||
```
|
||||
if (ThreadblockShape::M > ThreadblockShape::N):
|
||||
r = ThreadblockShape::M / ThreadblockShape::N
|
||||
i = i_macro
|
||||
j = (j_macro * r) + (t % r)
|
||||
elif (ThreadblockShape::M < ThreadblockShape::N):
|
||||
r = ThreadblockShape::N / ThreadblockShape::M
|
||||
i = (i_macro * r) + (t % r)
|
||||
j = j_macro
|
||||
else:
|
||||
i = i_macro
|
||||
j = j_macro
|
||||
```
|
||||
|
||||
### Handling cases with grid dimensions that aren't multiples of each other
|
||||
Even though threadblock shapes M and N are typically multiples of one another, the grid
|
||||
for a given problem may not have dimensions of the same ratio as that of the threadblock.
|
||||
For example, a problem of size 132x132 using a threadblock of shape 64x32 will result
|
||||
in a grid of 3x5 tiles. In this case, there is not an integer number of "true tiles"
|
||||
per "macro tile."
|
||||
|
||||
When this scenario arises, we simply pad the larger dimension of the grid such that
|
||||
there are an integer number of "true tiles" per "macro tile." Thus, the 3x5 grid in
|
||||
the example above will be treated as a 3x6 grid. Row and column positions for each
|
||||
tile are calculated as above. Any threadblocks that map to tiles that are outside the
|
||||
problem range or upper/lower triangular portion (e.g., (2, 5)) will exit early from
|
||||
this problem and may proceed to the next problem in the group.
|
||||
|
||||
### Handling upper-triangular matrices
|
||||
The only modification needed for upper-triangular matrices is to swap `i_macro` and `j_macro` in the calculations above.
|
||||
|
||||
# Scheduler modes
|
||||
The grouped kernel schedulers come with two different modes for finding
|
||||
the next tile for a block to compute. These techniques are controlled by
|
||||
the [`cutlass::gemm::kernel::GroupScheduleMode`](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/gemm/kernel/grouped_problem_visitor.h) enum.
|
||||
We describe each mode in greater detail below.
|
||||
|
||||
## `GroupScheduleMode::kDeviceOnly` (default)
|
||||
This scheduler mode performs all scheduling work on the device. It parallelizes
|
||||
the search for the problem that `tile_idx` maps to by having each thread "own"
|
||||
a different problem and determine whether `tile_idx` falls within the range of
|
||||
that problem.
|
||||
|
||||
`GroupScheduleMode::kDeviceOnly` performs this parallelization in a warp-wide
|
||||
fashion. Each thread in the warp loads a problem size indexed by its lane id and
|
||||
computes the number of tiles in that problem. A warp-wide prefix sum is used to find
|
||||
the starting tiles for the set of problems the warp is looking at. At the end of the
|
||||
prefix sum, each thread holds the starting tile index and tile count for a unique
|
||||
problem in the group.
|
||||
|
||||
While `tile_idx` remains within the range of the problems currently hosted by the
|
||||
warp, each thread will check whether `tile_idx` is in the range of its current
|
||||
problem. The matching problem index and its starting tile are then broadcasted to all
|
||||
threads in the warp.
|
||||
|
||||
## Precomputing schedules on the host: `GroupScheduleMode::kHostPrecompute`
|
||||
This scheduler attempts to reduce the amount of scheduling performed on the device
|
||||
by precomputing on the host the sequence of problems that will
|
||||
be accessed by each block. As described above, all that is needed to map tile_idx to
|
||||
the specific tile within a problem to compute is the problem ID and the problem's
|
||||
starting tile (among all of the tiles in the group). Thus, this scheduler precomputes
|
||||
the problem index and problem starting tile for each tile computed by each block.
|
||||
|
||||
The schedule for an individual block is represented as an array of
|
||||
`(problem_idx, problem_starting_tile)` tuples. There is one such array per block.
|
||||
These arrays are produced on the host and copied over to the device. This
|
||||
representation is optimized for the case in which blocks compute at most one
|
||||
tile per problem. When a block computes multiple tiles per problem in the group,
|
||||
the representation above will result in duplicate entries, and thus will be
|
||||
suboptimal (e.g., `[(3, 20), (3, 20)]` for a block that computes two tiles in
|
||||
problem 3, which has starting tile index 20).
|
||||
We have chosen to use the representation described above because grouped kernels
|
||||
themselves are typically most beneficial when problem sizes are small, and, thus,
|
||||
blocks compute at most one tile per problem.
|
||||
|
||||
## Which scheduler mode should I use?
|
||||
Consider the following questions when deciding which scheduling mode to use:
|
||||
|
||||
### How are the parameters used as input to the grouped kernel (e.g., ptrA, lda) set in my application?
|
||||
If these are set by a previous kernel running on
|
||||
the device (rather than by the host), you likely want to use `kDeviceOnly`,
|
||||
as this will minimize additional host-device communication.
|
||||
|
||||
### Can host-side work be overlapped with other device kernels in my application?
|
||||
For example, if a grouped GEMM is used as the Nth layer in a neural network,
|
||||
host-side precomputation for the grouped GEMM can potentially be overlapped
|
||||
with device-side work for layer N-1. In this case `kHostPrecompute` is likely
|
||||
a good fit.
|
||||
|
||||
### How compute-intensive are the problems in my group?
|
||||
The differences in performance between `kHostPrecompute` and `kDeviceOnly` are most
|
||||
noticeable for grouped kernels with low computational intensity, for which time spent in
|
||||
the scheduler accounts for a significant fraction of the grouped kernel's runtime.
|
||||
Intuitively, as problems in a group decrease in computational intensity, a smaller
|
||||
fraction of the overall runtime will be consumed in performing MMA operations, leading
|
||||
to a larger fraction of the overall runtime being consumed by scheduling logic.
|
||||
|
||||
Since the scheduling modes affect only the scheduling logic of the grouped kernels,
|
||||
one expects to see most benefit from `kHostPrecompute` for less computationally-intense
|
||||
groups.
|
||||
|
||||
# Improving Load Balance by Sorting Problems
|
||||
The grouped kernel schedulers assign a nearly equal number
|
||||
of tiles to each block participating in the grouped kernel. Every tile in the
|
||||
group has the same M and N dimensions. However, the K dimension of each
|
||||
tile depends on the K dimension of the problem, so tiles may have different
|
||||
K dimensions. Thus, the K dimension of the
|
||||
tile plays a significant role in determining how long it takes for a given
|
||||
tile to be computed.
|
||||
|
||||
## Potential problems with imbalanced K dimension
|
||||
To ensure that compute load is balanced evenly across blocks, it is important
|
||||
that the sum of the K dimensions among all tiles a block computes be similar
|
||||
to that of other blocks; if one block computes far more tiles with a large
|
||||
value of K than other blocks, it may take longer than the other blocks.
|
||||
|
||||
For example, consider the following group of GEMMs:
|
||||
```
|
||||
0 1152x768x128
|
||||
1 1152x768x1024
|
||||
2 768x1152x128
|
||||
3 768x1152x1024
|
||||
```
|
||||
If a tile size of 128x128 is used, then each problem will have 54 tiles.
|
||||
Thus, there are 216 tiles across the group.
|
||||
|
||||
Suppose this grouped GEMM is run on GA100, which has 108 SMs. Suppose that
|
||||
the occupancy given the parameters of the grouped GEMM is one -- one threadblock
|
||||
can be active at a time on an SM. The grouped GEMM will, thus, run with 108
|
||||
persistent threadblocks, each of which computes (256 / 108) = 2 tiles.
|
||||
|
||||
Under the round-robin assignment of tiles to threadblocks employed by
|
||||
the grouped GEMM scheduler, the assignment of tiles to threadblocks
|
||||
in this GEMM will be as follows:
|
||||
```
|
||||
Threadblocks 0-53: Tiles of size 128x128x128 from problem 0
|
||||
Threadblocks 54-107: Tiles of size 128x128x1024 from problem 1
|
||||
Threadblocks 0-53: Tiles of size 128x128x128 from problem 2
|
||||
Threadblocks 54-107: Tiles of size 128x128x1024 from problem 3
|
||||
```
|
||||
|
||||
Following this assignment, threadblocks 54-107 perform significantly more
|
||||
work than threadblocks 0-53 because they compute two tiles with a K
|
||||
dimension of 1024, whereas threadblocks 0-53 compute two tiles with K
|
||||
dimension of only 128.
|
||||
|
||||
Due to this imbalanced assignment, threadblocks 54-107 will run
|
||||
significantly longer than threadblocks 0-53, leaving threadblocks
|
||||
0-53 idle for a large fraction of time.
|
||||
|
||||
Clearly, a better assignment of tiles to threadblocks for this
|
||||
example would involve all threadblocks computing one tile with
|
||||
a K dimension of 1024 and one tile with a K dimension of 128.
|
||||
This would better balance the workload among threadblocks.
|
||||
|
||||
## Potential for sorting problems to reduce imbalance
|
||||
A simple way to potentially reduce load imbalance is to sort the problems in a group in
|
||||
descending order of their K dimension. This can help to improve load balance
|
||||
because tiles in a group are assigned in a round-robin fashion to blocks
|
||||
sequentially, so every block will always be assigned next the tile with
|
||||
the highest K dimension available.
|
||||
|
||||
Considering the example described above, sorting the problem sizes before
|
||||
executing grouped GEMM improves the runtime of this grouped GEMM on GA100 with each
|
||||
scheduling mode by around 30%.
|
||||
|
||||
To ease the process of sorting groups and their associated metadata in this
|
||||
manner, the device-level grouped kernels provide a `sort_problems()` method.
|
||||
An example of how to use this may be found in the [grouped GEMM example](https://github.com/NVIDIA/cutlass/tree/main/examples/24_gemm_grouped/gemm_grouped.cu).
|
||||
|
||||
Finally, while sorting problems can be helpful in certain scenarios, it is
|
||||
not guaranteed to improve performance. In some cases, performance can
|
||||
decrease when sorting problems due to additional conflicting factors that
|
||||
affect GEMM performance. We recommend profiling your grouped kernel with
|
||||
and without sorting to see whether it helps in your case.
|
||||
|
||||
## Copyright
|
||||
|
||||
Copyright (c) 2017 - 2025 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
```
|
||||
Redistribution and use in source and binary forms, with or without
|
||||
modification, are permitted provided that the following conditions are met:
|
||||
|
||||
1. Redistributions of source code must retain the above copyright notice, this
|
||||
list of conditions and the following disclaimer.
|
||||
|
||||
2. Redistributions in binary form must reproduce the above copyright notice,
|
||||
this list of conditions and the following disclaimer in the documentation
|
||||
and/or other materials provided with the distribution.
|
||||
|
||||
3. Neither the name of the copyright holder nor the names of its
|
||||
contributors may be used to endorse or promote products derived from
|
||||
this software without specific prior written permission.
|
||||
|
||||
THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS "AS IS"
|
||||
AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE
|
||||
IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE ARE
|
||||
DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT HOLDER OR CONTRIBUTORS BE LIABLE
|
||||
FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL
|
||||
DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR
|
||||
SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER
|
||||
CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY,
|
||||
OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE
|
||||
OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
|
||||
```
|
||||
151
media/docs/cpp/ide_setup.md
Normal file
151
media/docs/cpp/ide_setup.md
Normal file
@@ -0,0 +1,151 @@
|
||||
# IDE Setup for CUTLASS Development
|
||||
|
||||
This document outlines instructions and tips for setting up a local editor for CUTLASS development, including support
|
||||
for intellisense, go-to-definition, code formatting, and so on.
|
||||
|
||||
## Overview
|
||||
In order for any intellisense tool to work with CUTLASS, the following things need to be configured with it:
|
||||
* Include paths, i.e. where the compiler (or in this case, the intellisense tool) should look for header files
|
||||
* Compiler flags; especially the C++ standard (`--std`)
|
||||
* Preprocessor variables; especially CUDA-related ones
|
||||
|
||||
One usually needs to configure the above variables in a settings file. Below, two config approaches are described:
|
||||
for VSCode, and for any editor that uses the clangd language server, which includes
|
||||
Vim, Emacs, NeoVim, Sublime Text, and so on. Note that VSCode can also be configured to use clangd.
|
||||
It might be worth setting up clangd for VSCode rather than the default intellisense,
|
||||
and you might see faster responses and more stable performance with clangd.
|
||||
|
||||
## VSCode Setup
|
||||
|
||||
1. Install the [Official C/C++ extension](https://marketplace.visualstudio.com/items?itemName=ms-vscode.cpptools)
|
||||
1. Open settings...
|
||||
1. `Ctrl+Shift+P` to open the command palette
|
||||
1. Enter "C/C++" to filter results
|
||||
1. Select "C/C++ Edit Configurations (UI)" (or "... (JSON)" if you feel like editing the raw JSON)
|
||||
1. View the documentation for these settings
|
||||
[here](https://code.visualstudio.com/docs/cpp/c-cpp-properties-schema-reference)
|
||||
1. Edit "Include Path" to set up **include paths**. For CUTLASS, this includes the following:
|
||||
* `${workspaceFolder}/include`
|
||||
* `${workspaceFolder}/tools/util/include`
|
||||
* `${workspaceFolder}/examples/common`
|
||||
* ...others, depending on which files you edit
|
||||
1. Edit C++ standard to be `c++17`, `gnu++17`, or equivalent.
|
||||
1. Edit `defines` to define preprocessor variables. See
|
||||
[Global Config below](#global-config) for examples. The important
|
||||
ones include `__CUDACC_VER_MAJOR__`, `__CUDA_ARCH__`, `__CUDA_ARCH_FEAT_SM90_ALL__`. But configure
|
||||
them according to your target architecture.
|
||||
1. ...and possible edit any other fields for your specific setup.
|
||||
|
||||
## clangd Setup
|
||||
|
||||
`clangd` is a C++ language server that is part of the LLVM project. You must first set it up your specific IDE:
|
||||
* `clangd` official [documentation](https://clangd.llvm.org/installation#editor-plugins) for editor setup.
|
||||
* NeoVim setup is possible through [lsp](https://neovim.io/doc/user/lsp.html) and either manually installing clangd or
|
||||
using an installation manager like Mason.
|
||||
|
||||
Then, one needs to edit the config ([documentation](https://clangd.llvm.org/config)). One typically has a
|
||||
**global** and a **per-project** config.
|
||||
|
||||
### Global Config
|
||||
|
||||
Here is one example for a global config.
|
||||
On linux this is usually located at `~/.config/clangd/config.yaml` . Here is one example config for CUDA projects on SM90.
|
||||
The key settings here are the preprocessor vars (`-D__CUDACC_VER_MAJOR__` , `-D__CUDA_ARCH__`)
|
||||
|
||||
```
|
||||
CompileFlags:
|
||||
Compiler: /usr/local/cuda/bin/nvcc
|
||||
Add:
|
||||
- --cuda-path=/usr/local/cuda
|
||||
- --cuda-gpu-arch=sm_90a
|
||||
- -I/usr/local/cuda/include
|
||||
- "-xcuda"
|
||||
# report all errors
|
||||
- "-ferror-limit=0"
|
||||
- --cuda-gpu-arch=sm_90a
|
||||
- --std=c++17
|
||||
- "-D__INTELLISENSE__"
|
||||
- "-D__CLANGD__"
|
||||
- "-DCUDA_12_0_SM90_FEATURES_SUPPORTED"
|
||||
- "-DCUTLASS_ARCH_MMA_SM90_SUPPORTED=1"
|
||||
- "-D_LIBCUDACXX_STD_VER=12"
|
||||
- "-D__CUDACC_VER_MAJOR__=12"
|
||||
- "-D__CUDACC_VER_MINOR__=3"
|
||||
- "-D__CUDA_ARCH__=900"
|
||||
- "-D__CUDA_ARCH_FEAT_SM90_ALL"
|
||||
- "-Wno-invalid-constexpr"
|
||||
Remove:
|
||||
# strip CUDA fatbin args
|
||||
- "-Xfatbin*"
|
||||
# strip CUDA arch flags
|
||||
- "-gencode*"
|
||||
- "--generate-code*"
|
||||
# strip CUDA flags unknown to clang
|
||||
- "-ccbin*"
|
||||
- "--compiler-options*"
|
||||
- "--expt-extended-lambda"
|
||||
- "--expt-relaxed-constexpr"
|
||||
- "-forward-unknown-to-host-compiler"
|
||||
- "-Werror=cross-execution-space-call"
|
||||
Hover:
|
||||
ShowAKA: No
|
||||
InlayHints:
|
||||
Enabled: No
|
||||
Diagnostics:
|
||||
Suppress:
|
||||
- "variadic_device_fn"
|
||||
- "attributes_not_allowed"
|
||||
```
|
||||
|
||||
### Local Config
|
||||
Local config is needed to specify per-project settings, especially include paths. An example is:
|
||||
```
|
||||
CompileFlags:
|
||||
Add:
|
||||
- -I</absolute/path/to/cutlass>/include/
|
||||
- -I</absolute/path/to/cutlass>/tools/util/include/
|
||||
- -I</absolute/path/to/cutlass>/cutlass/examples/common/
|
||||
```
|
||||
|
||||
Note that absolute paths are needed since clangd doesn't support relative paths.
|
||||
|
||||
### Note on compile_commands.json
|
||||
For typical C++ projects, clangd can *automatically* configure itself by parsing the `compile_commands.json`
|
||||
generated by your CMake build. The path to such a file is by default `build/compile_commands.json` and is
|
||||
configured by the `CompilationDatabase` config.
|
||||
|
||||
This is usually a convenient way to configure projects, but it's not as simple for CUDA/nvcc projects, since
|
||||
clang doesn't understand many of the compiler flags used by nvcc. Hence, for now, we don't recommend using
|
||||
`compile_commands.json` to configure your CUDA project.
|
||||
|
||||
## Copyright
|
||||
|
||||
Copyright (c) 2017 - 2025 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
```
|
||||
Redistribution and use in source and binary forms, with or without
|
||||
modification, are permitted provided that the following conditions are met:
|
||||
|
||||
1. Redistributions of source code must retain the above copyright notice, this
|
||||
list of conditions and the following disclaimer.
|
||||
|
||||
2. Redistributions in binary form must reproduce the above copyright notice,
|
||||
this list of conditions and the following disclaimer in the documentation
|
||||
and/or other materials provided with the distribution.
|
||||
|
||||
3. Neither the name of the copyright holder nor the names of its
|
||||
contributors may be used to endorse or promote products derived from
|
||||
this software without specific prior written permission.
|
||||
|
||||
THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS "AS IS"
|
||||
AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE
|
||||
IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE ARE
|
||||
DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT HOLDER OR CONTRIBUTORS BE LIABLE
|
||||
FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL
|
||||
DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR
|
||||
SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER
|
||||
CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY,
|
||||
OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE
|
||||
OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
|
||||
```
|
||||
790
media/docs/cpp/implicit_gemm_convolution.md
Normal file
790
media/docs/cpp/implicit_gemm_convolution.md
Normal file
@@ -0,0 +1,790 @@
|
||||

|
||||
|
||||
# CUTLASS Convolution
|
||||
|
||||
Implicit GEMM is the formulation of a convolution operation as a GEMM (generalized matrix-matrix
|
||||
product). Convolution takes an activation tensor and applies a sliding filter on it to produce an
|
||||
output tensor.
|
||||
|
||||
## Introduction
|
||||
|
||||
This release of CUTLASS contains several artifacts related to convolution.
|
||||
|
||||
- [**Implicit GEMM Algorithm**](implicit_gemm_convolution.md#implicit-gemm-algorithm)
|
||||
- [**CUTLASS Convolution Implementation**](implicit_gemm_convolution.md#cutlass-convolution-implementation)
|
||||
- [**Convolution Examples**](implicit_gemm_convolution.md#convolution-example)
|
||||
|
||||
|
||||
# Implicit GEMM Algorithm
|
||||
|
||||
2-D convolution may be mapped to matrix multiply
|
||||
by first forming a _convolution matrix_ containing elements of the activations tensor,
|
||||
then multiplying this by a matrix formed from the filters tensor.
|
||||
The earliest form of this algorithm constructs the convolution matrix explicitly via an operation
|
||||
conventionally referred to as `im2col`. The resulting matrix replicates each activation element by a factor
|
||||
equal to the filter size, consuming additional storage capacity and memory bandwidth.
|
||||
|
||||
The _implicit GEMM_ algorithm is a variation on the blocked, hierarchical GEMM computation in CUDA.
|
||||
Instead of constructing the convolution matrix explicitly,
|
||||
it forms tiles of the convolution matrix on the fly
|
||||
as data are loaded from global memory into Shared Memory
|
||||
by carefully updating pointers and predicates.
|
||||
Once the convolution matrix is formed in Shared Memory,
|
||||
the existing warp-level GEMM components accumulate the result of
|
||||
convolution and update the output tensor.
|
||||
|
||||
This section describes the structure of an efficient Implicit GEMM Convolution CUDA kernel
|
||||
for Turing Tensor Cores.
|
||||
|
||||
## Mapping Convolution to GEMM
|
||||
|
||||
The forward convolutional layer computes an output tensor _y = conv(x, w)_ where x(NHWC), w(KRSC), and y(NPQK)
|
||||
are 4-D tensors.
|
||||
|
||||
This computation may be described by the following analytic function.
|
||||
|
||||
```
|
||||
y[n, p, q, k] = sum_c(sum_r(sum_s( x[n, f(p, r), g(q, s), c] * w[k, r, s, c] )))
|
||||
```
|
||||
where functions _f_ and _g_ are defined as follows.
|
||||
|
||||
```
|
||||
f(p, r) = p * stride_h + R - r - 1 + pad_h
|
||||
g(q, s) = q * stride_w + S - s - 1 + pad_w
|
||||
```
|
||||
|
||||
A [host](https://github.com/NVIDIA/cutlass/tree/main/tools/util/include/cutlass/util/reference/host/convolution.h) and [device](https://github.com/NVIDIA/cutlass/tree/main/tools/util/include/cutlass/util/reference/device/convolution.h)
|
||||
reference implementation are provided in the CUTLASS Utilities.
|
||||
|
||||
This computation may be mapped to the elements of a matrix product as follows.
|
||||
|
||||
```
|
||||
C = gemm(A, B)
|
||||
```
|
||||
where
|
||||
- A is a row-major matrix of extent _NHW_-by-_RSC_ containing activations
|
||||
- B is a column-major matrix of extent _RSC_-by-_K_ containing filters
|
||||
- C is a row-major matrix of extent _NPQ_-by-_K_ containing the output
|
||||
|
||||
Each element of the output matrix _Cij_ corresponds to an element in the output tensor y[n, p, q, k] according to
|
||||
the following relation.
|
||||
```
|
||||
y[n, p, q, k] = Cij
|
||||
```
|
||||
where
|
||||
```
|
||||
i = q + Q * (p + P * n)
|
||||
j = k
|
||||
```
|
||||
|
||||
These relations may be inverted as follows.
|
||||
```
|
||||
k = j
|
||||
|
||||
n = i / (PQ)
|
||||
residual = i % (PQ)
|
||||
|
||||
p = residual / Q
|
||||
q = residual % Q
|
||||
```
|
||||
|
||||
The triple loop nest iterating over CRS to accumulate the result may also be linearized and mapped to the inner
|
||||
GEMM _K_ dimension (not to be confused with the filter tensor dimension _K_) by the following relations.
|
||||
|
||||
```
|
||||
gemm_k = s + S * (r + R * c)
|
||||
```
|
||||
and inverse
|
||||
```
|
||||
c = gemm_k / (RS)
|
||||
residual = gemm_k % (RS)
|
||||
|
||||
r = residual / S
|
||||
s = residual % S
|
||||
```
|
||||
|
||||
Given these equations, a GEMM triple loop nest could be augmented with tensor indexing as follows.
|
||||
```c++
|
||||
int GEMM_M = N * P * Q;
|
||||
int GEMM_N = K;
|
||||
int GEMM_K = C * R * S;
|
||||
|
||||
for (int gemm_i = 0; gemm_i < GEMM_M; ++gemm_i) {
|
||||
for (int gemm_j = 0; gemm_j < GEMM_N; ++gemm_j) {
|
||||
|
||||
int n = gemm_i / (PQ);
|
||||
int npq_residual = gemm_i % (PQ);
|
||||
|
||||
int p = npq_residual / Q;
|
||||
int q = npq_residual % Q;
|
||||
|
||||
Accumulator accum = 0;
|
||||
|
||||
for (int gemm_k = 0; gemm_k < GEMM_K; ++gemm_k) {
|
||||
|
||||
int k = gemm_j;
|
||||
|
||||
int c = gemm_k / (RS);
|
||||
int crs_residual = gemm_k % (RS);
|
||||
|
||||
int r = crs_residual / S;
|
||||
int s = crs_residual % S;
|
||||
|
||||
int h = f(p, r);
|
||||
int w = g(q, s);
|
||||
|
||||
ElementA a = tensor_A.at({n, h, w, c});
|
||||
ElementB b = tensor_B.at({k, r, s, c});
|
||||
|
||||
accum += a * b;
|
||||
}
|
||||
|
||||
C[gemm_i * K + gemm_j] = accum;
|
||||
}
|
||||
}
|
||||
```
|
||||
The [CUTLASS GEMM implementation](efficient_gemm.md) explicitly iterates over tiles. Consequently,
|
||||
a tile iterator could be implemented to compute these functions analytically and load the appropriate
|
||||
elements. However, the resulting modulo arithmetic would be computationally intensive, and overhead would
|
||||
limit performance of a GEMM kernel targeting Turing Tensor Cores.
|
||||
|
||||
The following section describes how an efficient implementation may be implemented within the structure of
|
||||
a hierarchical GEMM kernel targeting Tensor Cores.
|
||||
|
||||
|
||||
# CUTLASS Convolution Implementation
|
||||
|
||||
To get the best performance, the following parameters are recommended.
|
||||
|
||||
- All tensors are 128-bit aligned NHWC tensors
|
||||
- Channel count (C) is a multiple of 32 elements
|
||||
- Filter count (K) is a multiple of 32 elements
|
||||
|
||||
This enables 128-bit vector memory acceses which lead to efficient CUDA kernels. Smaller alignment is supported even on tensor cores by setting AlignmentA and AlignmentB in `conv::kernel::DefaultConv2dFprop`, but the performance is lower than 128-bit aligned tensors.
|
||||
|
||||
# CUTLASS Device-level Convolution Operator
|
||||
|
||||
CUTLASS defines CUDA C++ templates accepting numerous template arguments to specialize the resulting
|
||||
kernel by operation, data type, tile configuration, math instruction, and fused output operation.
|
||||
|
||||
In [turing_tensorop_conv2dfprop.cu](https://github.com/NVIDIA/cutlass/tree/main/examples/09_turing_tensorop_conv2dfprop/turing_tensorop_conv2dfprop.cu), a convolution
|
||||
operation is defined as follows.
|
||||
|
||||
```c++
|
||||
/// Define an Implicit GEMM convolution forward propagation (fprop) kernel
|
||||
using Conv2dFpropKernel = typename cutlass::conv::kernel::DefaultConv2dFprop<
|
||||
ElementInputA, // data type of element a (mapped to activation for fprop)
|
||||
LayoutInputA, // layout of element a (mapped to activation for fprop)
|
||||
ElementInputB, // data type of element b (mapped to filters for fprop)
|
||||
LayoutInputB, // layout of element b (mapped to filters for fprop)
|
||||
ElementC, // data type of element c (mapped to output for fprop)
|
||||
LayoutC, // layout of element c (mapped to output for fprop)
|
||||
ElementAccumulator, // data type of internal accumulation
|
||||
MMAOp, // opcode class tag
|
||||
SmArch, // target SM architecture
|
||||
ThreadblockShape, // shape of threadblock tile
|
||||
WarpShape, // shape of warp-level GEMM tile
|
||||
InstructionShape, // shape of target math instruction
|
||||
EpilogueOp, // epilogue operator
|
||||
SwizzleThreadBlock, // optional function to reorder threadblocks for locality
|
||||
NumStages, // number of pipeline stages in threadblock-scoped GEMM
|
||||
cutlass::arch::OpMultiplyAddSaturate, // math operation on data of element a and b
|
||||
cutlass::conv::IteratorAlgorithm::kOptimized // global memory iterator algorithm
|
||||
>::Kernel
|
||||
```
|
||||
|
||||
This template is intended to be generic and cover all feasible configurations. The example specifies
|
||||
the following concrete data types, layouts, and tile shapes.
|
||||
|
||||
```c++
|
||||
/// Define an Implicit GEMM convolution forward propagation (fprop) kernel
|
||||
using Conv2dFpropKernel = typename cutlass::conv::kernel::DefaultConv2dFprop<
|
||||
cutlass::int4b_t, // data type of element a (mapped to activation for fprop)
|
||||
cutlass::layout::TensorNHWC, // layout of element a (mapped to activation for fprop)
|
||||
cutlass::int4b_t, // data type of element b (mapped to filters for fprop)
|
||||
cutlass::layout::TensorNHWC, // layout of element b (mapped to filters for fprop)
|
||||
int32_t, // data type of element c (mapped to output for fprop)
|
||||
cutlass::layout::TensorNHWC, // layout of element c (mapped to output for fprop)
|
||||
int32_t, // data type of internal accumulation
|
||||
cutlass::arch::OpClassTensorOp, // opcode class tag
|
||||
cutlass::arch::Sm75, // target SM architecture
|
||||
cutlass::gemm::GemmShape<128, 128, 128>, // shape of threadblock tile
|
||||
cutlass::gemm::GemmShape<64, 64, 128>, // shape of warp-level GEMM tile
|
||||
cutlass::gemm::GemmShape<8, 8, 32>, // shape of target math instruction
|
||||
cutlass::epilogue::thread::LinearCombinationClamp<
|
||||
int32_t, // data type of output matrix
|
||||
8, // The number of elements per vectorized
|
||||
// memory access. This becomes the vector width of
|
||||
// math instructions in the epilogue too.
|
||||
int32_t, // Data type of accumulator
|
||||
float>; , // epilogue operator
|
||||
SwizzleThreadBlock, // optional function to reorder threadblocks for locality
|
||||
2, // number of pipeline stages in threadblock-scoped GEMM
|
||||
cutlass::arch::OpMultiplyAddSaturate, // math operation on data of element a and b
|
||||
cutlass::conv::IteratorAlgorithm::kOptimized // global memory iterator algorithm
|
||||
>::Kernel
|
||||
```
|
||||
|
||||
That is, this computes 2D convolutional forward propagation with 4-bit integer inputs and outputs (`cutlass::int4b_t`).
|
||||
Internal accumulation is performed using 32-bit integers (`int32_t`), and an elementwise linear combination operation
|
||||
is performed on the output in single-precision floating point (`float`).
|
||||
|
||||
The threadblock and warp-level tile shapes refer to the hierarchically blocked GEMM computation
|
||||
[described here](gemm_api.md). Larger tiles achieve greater reuse of data loaded through shared memory
|
||||
but launch fewer CTAs and may not fully occupy the GPU for small problem sizes. Smaller tile configurations achieve
|
||||
lower peak utilizations but may better match the number of SMs within the GPU for real-world workloads.
|
||||
|
||||
|
||||
## Launching the convolution
|
||||
|
||||
The following code collects the arguments for an implicit GEMM operation into a structure.
|
||||
|
||||
```c++
|
||||
//
|
||||
// Define arguments for CUTLASS Convolution
|
||||
//
|
||||
|
||||
// mode (kCrossCorrelation or kConvolution)
|
||||
cutlass::conv::Mode mode = cutlass::conv::Mode::kCrossCorrelation;
|
||||
|
||||
// Split K dimension into 1 partitions
|
||||
int split_k_slices = 1;
|
||||
|
||||
cutlass::conv::Conv2dProblemSize problem_size(
|
||||
options.input_size,
|
||||
options.filter_size,
|
||||
options.padding,
|
||||
options.conv_stride,
|
||||
options.dilation,
|
||||
options.output_size(),
|
||||
mode,
|
||||
split_k_slices);
|
||||
|
||||
typename ImplicitGemm::Arguments arguments{
|
||||
problem_size,
|
||||
tensor_a.device_ref(),
|
||||
tensor_b.device_ref(),
|
||||
tensor_c.device_ref(),
|
||||
tensor_c.device_ref(),
|
||||
{options.alpha, options.beta},
|
||||
};
|
||||
```
|
||||
|
||||
The `mode` flag indicates whether to compute cross correlation or convolution. The arguments
|
||||
`input_size`, `filter_size`, `padding`, `conv_stride`, and `dilation` specify the dimensions of the
|
||||
input and output tensors and characterize the problem size.
|
||||
|
||||
The arguments `tensor_a.device_ref()`, `tensor_b.device_ref()`, and `tensor_c.device_ref()` are
|
||||
CUTLASS `TensorRef<>` objects containing a pointer to the tensor data in GPU device memory and stride values.
|
||||
|
||||
The following code initializes and launches the Implicit GEMM operation on the device. After initializing
|
||||
the arguments structure, it is used to query device-side workspace requirements and allocate them
|
||||
in device memory if needed.
|
||||
|
||||
Then, the Implicit GEMM object is initialized with the `arguments` structure and the workspace in
|
||||
device memory. This initialization step precomputes internal lookup tables used by the convolution kernel
|
||||
and may also clear the device-side workspace if needed.
|
||||
|
||||
Finally, the initialized Implicit GEMM object is called, launching a kernel on the device. `tensor_c` now
|
||||
contains the result of the implicit GEMM.
|
||||
|
||||
```c++
|
||||
ImplicitGemm implicit_gemm_op;
|
||||
|
||||
// Query workspace size
|
||||
size_t workspace_size = implicit_gemm_op.get_workspace_size(arguments);
|
||||
|
||||
// Allocate workspace memory
|
||||
cutlass::device_memory::allocation<uint8_t> workspace(workspace_size);
|
||||
|
||||
// Initialize the Implicit GEMM object
|
||||
cutlass::Status status = implicit_gemm_op.initialize(arguments, workspace.get());
|
||||
|
||||
if (status != cutlass::Status::kSuccess) {
|
||||
/* error */
|
||||
}
|
||||
|
||||
//
|
||||
// Launch initialized CUTLASS kernel
|
||||
//
|
||||
|
||||
status = implicit_gemm_op();
|
||||
|
||||
if (status != cutlass::Status::kSuccess) {
|
||||
/* error */
|
||||
}
|
||||
```
|
||||
|
||||
The example demonstrates how the input and output tensors may be written to a file as CSV using
|
||||
`cutlass::HostTensor<>` defined in the [CUTLASS Utilities](utilities.md).
|
||||
|
||||
```c++
|
||||
std::ofstream output_workspace(ss.str());
|
||||
|
||||
output_workspace
|
||||
<< "Input = \n" << tensor_a.host_view() << "\n\n"
|
||||
<< "Filters = \n" << tensor_b.host_view() << "\n\n";
|
||||
|
||||
// Copy device memory to host backing store
|
||||
tensor_c.sync_host();
|
||||
|
||||
output_workspace << "Computed = \n" << tensor_c.host_view() << std::endl;
|
||||
```
|
||||
|
||||
|
||||
## CUTLASS Components
|
||||
|
||||
CUTLASS defines the following CUDA C++ templates to implement Implicit GEMM Convolution which are described in greater detail in subsequent sections.
|
||||
|
||||
**Activations tile iterators** load the activations tile into registers. Two implementations are provided:
|
||||
- [conv2d_fprop_activation_tile_access_iterator_analytic.h](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/conv/threadblock/conv2d_fprop_activation_tile_access_iterator_analytic.h) computes pointer deltas and masks analytically
|
||||
- [conv2d_fprop_activation_tile_access_iterator_optimized.h](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/conv/threadblock/conv2d_fprop_activation_tile_access_iterator_optimized.h) optimizes iterating over global memory and
|
||||
creating GEMM-A tile in shared memory.
|
||||
|
||||
**Filter tile iterators** load filters into registers. Similarly, two implementations are provided:
|
||||
- [conv2d_fprop_filter_tile_access_iterator_analytic.h](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/conv/threadblock/conv2d_fprop_filter_tile_access_iterator_analytic.h) computes pointer deltas and masks analytically
|
||||
- [conv2d_fprop_filter_tile_access_iterator_optimized.h](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/conv/threadblock/conv2d_fprop_filter_tile_access_iterator_optimized.h) optimizes iterating over global memory and
|
||||
creating GEMM-B tile in shared memory.
|
||||
|
||||
The improvements covered by optimized iterators are:
|
||||
|
||||
a. Precomputing kernel-invariant pointer deltas on the host
|
||||
b. Computing cta-invariant mask predicates on device-side iterator ctors
|
||||
c. Use of [fast divmod](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/fast_math.h) to map GEMM dimensions to convolution tensors.
|
||||
|
||||
For example, an _optimized_ activation iterator uses fast divmod to map GEMM _M_ to NPQ.
|
||||
|
||||
**Pipelined mainloop** loads threadblock-scoped tiles from global memory into shared memory and then applies
|
||||
CUTLASS warp-level GEMM operations to load from Shared Memory and issue instructions to Turing Tensor Cores.
|
||||
- [mma_pipelined.h](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/conv/threadblock/implicit_gemm_pipelined.h)
|
||||
|
||||
Operations for storing to shared memory and performing warp-wide matrix multiply operations using
|
||||
Turing Tensor Cores are applied directly from the CUTLASS GEMM components. These include the
|
||||
following components.
|
||||
|
||||
**Regular Tile Iterator** implemented in
|
||||
[transform::threadblock::RegularTileIterator](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/transform/threadblock/regular_tile_iterator.h)
|
||||
stores register-backed fragments to Shared Memory in permuted layouts.
|
||||
|
||||
**Warp-level GEMM** defined in [cutlass::gemm::warp::MmaTensorOp](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/gemm/warp/mma_tensor_op.h)
|
||||
defines tile iterators to load from Shared Memory and issue math instructions to Turing Tensor Cores.
|
||||
Further details are [described in here](gemm_api.md#warp-level-matrix-multiply-api).
|
||||
|
||||
**Epilogue** reorders accumulator elements among threads within a threadblock to efficiently update
|
||||
the output tensor. It is implemented in [epilogue::threadblock::Epilogue](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/epilogue/threadblock/epilogue.h).
|
||||
|
||||
### Loading Activations and Filters
|
||||
|
||||
The Implicit GEMM Convolution algorithm partitions the GEMM _K_ dimension (of extent _CRS_) into
|
||||
threadblock tiles and assigning each threadblock tile to one filter position and an interval
|
||||
of channels. After iterating over all filter positions, the convolution algorithm advances to the
|
||||
next interval of channels and proceeds from filter `r=0, s=0`.
|
||||
|
||||
The matrix product of one threadblock tile is computed per iteration of
|
||||
the mainloop as described in the [CUTLASS GEMM implementation](efficient_gemm.md). To
|
||||
summarize, the threadblock tile of activations and filters are loaded from tensors in global memory
|
||||
and stored to shared memory. Each thread within the threadblock loads one or more vectors and
|
||||
collectively span the entire tile.
|
||||
|
||||
The following figure illustrates one particular iteration of the Implicit GEMM mainloop. Each
|
||||
thread within the threadblock is mapped to several vectors of elements in the Activations and
|
||||
Filters tensors. Each index in the GEMM _M_ dimension corresponds to a unique _(N,P,Q)_
|
||||
index of the output tensor, and pointers may be computed based on this as well as
|
||||
filter position _(r,s)_.
|
||||
|
||||

|
||||
|
||||
The CUTLASS component that embodies this functionality is [Conv2dFpropFilterTileAccessIteratorAnalytic](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/conv/threadblock/conv2d_fprop_activation_tile_access_iterator_analytic.h).
|
||||
Its constructor computes the mapping of GEMM _M_ to _(N, P, Q)_, the `at()` method maps the linear offset into the Activations
|
||||
tensor for each memory access the thread is to perform. Additionally, the method `valid()` computes the valided of the access
|
||||
for each filter position and for each memory access to indicate whether the memory access will be within the bounds of the
|
||||
tensor or out of bounds.
|
||||
|
||||
`operator++()` iterates over memory accesses performed by a thread in both contiguous and strided dimension.
|
||||
|
||||
```c++
|
||||
// cutlass/conv/threadblock/conv2d_fprop_activation_tile_access_iterator_analytic.h
|
||||
|
||||
// Update iterator to thread's next contiguous, strided memory access
|
||||
Conv2dFpropActivationTileAccessIteratorAnalytic &operator++() {
|
||||
++iteration_contiguous_;
|
||||
if (iteration_contiguous_ < ThreadMap::Iterations::kContiguous) {
|
||||
return *this;
|
||||
}
|
||||
iteration_contiguous_ = 0;
|
||||
|
||||
++iteration_strided_;
|
||||
if (iteration_strided_ < ThreadMap::Iterations::kStrided) {
|
||||
return *this;
|
||||
}
|
||||
iteration_strided_ = 0;
|
||||
|
||||
return *this;
|
||||
}
|
||||
```
|
||||
|
||||
After all accesses have been visited for the current threadblock tile, `advance()` updates the pointers to next tile.
|
||||
Offsets added to each pointer follows the traversal of filter positions, performing one of the
|
||||
following:
|
||||
- advance from filter position _(r, s, c)_ to filter position _(r, s+1, c)_
|
||||
- advance from filter position _(r, S-1, c)_ to filter position _(r+1, 0, c)_
|
||||
- advance from filter position _(R-1, S-1, c)_ to filter position _(0, 0, c+32)_
|
||||
|
||||
This logic within method `advance()`'s body computes the above three updates for the activation GEMM-A tile.
|
||||
|
||||
```c++
|
||||
// cutlass/conv/threadblock/conv2d_fprop_activation_tile_access_iterator_analytic.h
|
||||
|
||||
// Advance to the next access
|
||||
void advance() {
|
||||
// moves to the next tile
|
||||
++filter_s_;
|
||||
if (filter_s_ < problem_size_.S) {
|
||||
return;
|
||||
}
|
||||
filter_s_ = 0;
|
||||
|
||||
++filter_r_;
|
||||
if (filter_r_ < problem_size_.R) {
|
||||
return;
|
||||
}
|
||||
filter_r_ = 0;
|
||||
|
||||
filter_c_ += Shape::kRow * problem_size_.split_k_slices;
|
||||
}
|
||||
```
|
||||
|
||||
Similar logic holds for [Conv2dFpropFilterTileAccessIteratorAnalytic](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/conv/threadblock/conv2d_fprop_filter_tile_access_iterator_analytic.h).
|
||||
|
||||
To reduce computational overhead in the mainloop body, the pointer offsets may be precomputed
|
||||
in host code and provided to the CUDA kernel as a lookup table in its `Params` structure.
|
||||
As shown in [Conv2dFpropFilterTileAccessIteratorOptimized](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/conv/threadblock/conv2d_fprop_activation_tile_access_iterator_optimized.h),
|
||||
the logic to compute offsets from filter position has been extracted to the `Params` constructor.
|
||||
|
||||
```c++
|
||||
// cutlass/conv/threadblock/conv2d_params.h
|
||||
struct Conv2dFpropActivationIteratorOptimizedParams<layout::TensorNHWC> {
|
||||
...
|
||||
// next S
|
||||
inc_next[0] = conv_sign * (int64_t(layout.stride()[0]) * problem_size.dilation_w) * element_size_bits / 8;
|
||||
|
||||
// next R
|
||||
inc_next[1] = conv_sign * (
|
||||
int64_t(layout.stride()[1]) * problem_size.dilation_h
|
||||
- (problem_size.S - 1) * layout.stride()[0] * problem_size.dilation_w
|
||||
) * element_size_bits / 8;
|
||||
|
||||
// next C
|
||||
inc_next[2] = (
|
||||
threadblock_shape.column() * problem_size.split_k_slices
|
||||
- conv_sign * int64_t(problem_size.R - 1) * layout.stride()[1] * problem_size.dilation_h
|
||||
- conv_sign * int64_t(problem_size.S - 1) * layout.stride()[0] * problem_size.dilation_w
|
||||
) * element_size_bits / 8;
|
||||
|
||||
...
|
||||
}
|
||||
```
|
||||
|
||||
This allows only a simple lookup from the _delta table_ performed in device code in `Conv2dFpropActivationTileAccessIteratorOptimized::advance()`.
|
||||
|
||||
```c++
|
||||
// cutlass/conv/threadblock/conv2d_fprop_activation_tile_access_iterator_optimized.h
|
||||
CUTLASS_HOST_DEVICE
|
||||
void advance() {
|
||||
|
||||
int next_idx = 0;
|
||||
|
||||
// moves to the next tile
|
||||
++filter_s_;
|
||||
if (filter_s_ == problem_size_.S) {
|
||||
filter_s_ = 0;
|
||||
++filter_r_;
|
||||
|
||||
if (filter_r_ < problem_size_.R) {
|
||||
next_idx = 1;
|
||||
}
|
||||
else {
|
||||
filter_r_ = 0;
|
||||
next_idx = 2;
|
||||
}
|
||||
}
|
||||
|
||||
add_byte_offset_(params_.inc_next[next_idx]); // in addition to Conv2dFpropActivationTileAccessIteratorAnalytic::advance()
|
||||
|
||||
if (next_idx == 2) {
|
||||
filter_c_ += params_.filter_c_delta;
|
||||
}
|
||||
}
|
||||
|
||||
```
|
||||
|
||||
### Making use of Tensor Cores
|
||||
|
||||
Turing Tensor Cores compute matrix multiply-accumulate operations efficiently by sharing data among all
|
||||
threads within a warp. The following operations are supported.
|
||||
|
||||
| **Shape** | **A** | **B** | **C** |
|
||||
|-----------|---------|---------|---------|
|
||||
| 8x8x32 | int4b_t | int4b_t | int32_t |
|
||||
| 8x8x16 | int8b_t | int8b_t | int32_t |
|
||||
| 16x8x8 | half | half | half |
|
||||
| 16x8x8 | half | half | float |
|
||||
|
||||
Functionally, the Turing 8x8x32 matrix multiply operation distributes the _A_, _B_, and _C_ matrix across 32
|
||||
threads within a warp according to the following illustration.
|
||||
|
||||

|
||||
|
||||
This Tensor Core operation is accessible to the CUDA programmer via the PTX instruction
|
||||
[`mma.sync`](https://docs.nvidia.com/cuda/parallel-thread-execution/index.html#warp-level-matrix-fragment-mma-8832).
|
||||
CUTLASS wraps inline PTX with device-side intrinsics defined in [`cutlass/arch/mma_sm75.h`](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/arch/mma_sm75.h)
|
||||
as in the following example.
|
||||
|
||||
```c++
|
||||
unsigned A; // eight packed 4-bit integer elements
|
||||
unsigned B; // eight packed 4-bit integer elements
|
||||
|
||||
int C[2]; // two 32-bit integer elements
|
||||
int D[2]; // two 32-bit integer elements
|
||||
|
||||
asm volatile(
|
||||
"mma.sync.aligned.m8n8k32.row.col.s32.s4.s4.s32 {%0,%1}, {%2}, {%3}, {%4,%5};\n"
|
||||
: "=r"(D[0]), "=r"(D[1])
|
||||
: "r"(A), "r"(B), "r"(C[0]), "r"(C[1]));
|
||||
```
|
||||
|
||||
To load data efficiently from Shared Memory into registers with the distribution among
|
||||
warps matching the above, the Turing GPU architecture introduces
|
||||
[`ldmatrix`](https://docs.nvidia.com/cuda/parallel-thread-execution/index.html#warp-level-matrix-instructions-ldmatrix).
|
||||
`ldmatrix` is the ultimate warp-cooperative instruction, as all threads contribute addresses to up to 32 row vectors of
|
||||
size 128-bits in length. These rows are fetched from Shared Memory and then distributed among groups of four threads
|
||||
per row.
|
||||
|
||||
The arrangement of SMEM pointers and destination registers within threads is illustrated as follows. Thread 0 is highlighted
|
||||
in the illustration to emphasize the mapping.
|
||||
|
||||

|
||||
|
||||
The size of the Turing Tensor Core operation computing matrix multiply-accumulate on INT4 data is 8-by-8-by-32
|
||||
elements. `ldmatrix` fetches up to 32 rows (or columns) per operation. Sixteen Tensor Core operations may be issued
|
||||
to implement a 32-by-32-by-32 matrix product and perfectly consume all data loaded by two `ldmatrix` instructions
|
||||
as shown in the following figure. Larger tiles are possible by increasing the number of memory instructions
|
||||
and issuing more Tensor Core operations, up to warp-level matrix operations of size 64-by-64-by-32. The limit is
|
||||
the number of registers to hold the accumulator elements.
|
||||
|
||||

|
||||
|
||||
### Shared Memory Layouts
|
||||
|
||||
In the previous two sections, we have described how data may be loaded from activations and filters tensors
|
||||
in global memory to compute convolution, and we have described a composition of `ldmatrix` and `mma.sync`
|
||||
to fetch data from Shared Memory and issue Tensor Core operations.
|
||||
|
||||
To ensure this data movement is efficient, care must be taken to ensure bank conflicts are avoided. CUTLASS
|
||||
uses a permuted Shared Memory layout to avoid bank conflicts when storing to Shared Memory and to efficiently
|
||||
load from Shared Memory using `ldmatrix`. The following figure illustrates the thread mapping used for
|
||||
the loading the activations and filters threadblock tiles from global memory and the permuted layout in
|
||||
Shared Memory.
|
||||
|
||||

|
||||
|
||||
In the illustration, one warp-wide memory access is highlighted in blue, with individual threads
|
||||
loading one 128-bit vector. The tile in global memory could correspond either to the activations
|
||||
or filters and is assumed to be 'strip-mined' with four threads loading consecutive channels.
|
||||
|
||||
Shared Memory is visualized as a 'row-major' matrix with eight columns representing
|
||||
the eight 128-bit banks.
|
||||
As described in the CUTLASS GTC 2019 presentation [slides](https://developer.download.nvidia.com/video/gputechconf/gtc/2019/presentation/s9593-cutensor-high-performance-tensor-operations-in-cuda-v2.pdf),
|
||||
[recording](https://developer.nvidia.com/gtc/2019/video/S9593), an access to Shared Memory will be conflict-free if
|
||||
the following conditions are satisfied across each warp:
|
||||
- {T0, T1, .., T7} do not access the same 128-bit bank
|
||||
- {T8, T9, .., T15} do not access the same 128-bit bank
|
||||
- {T16, T17, .., T23} do not access the same 128-bit bank
|
||||
- {T24, T25, .., T31} do not access the same 128-bit bank
|
||||
|
||||
To achieve conflict-free stores, the Shared Memory layout remaps the strip-mined arrangement to transpose
|
||||
the vectors and applies an XOR operation on the column index of each thread's pointer. Specifically,
|
||||
|
||||
```c++
|
||||
int store_column = (lane_id % 8) ^ (lane_id / 8);
|
||||
```
|
||||
|
||||
This transformation on the layout will be instrumental in reading slices of data from Shared Memory
|
||||
to compute the warp-level matrix multiply using Tensor Cores.
|
||||
|
||||
The following figure shows how the first sixteen threads participating in an `ldmatrix` instruction
|
||||
logically map to the c=0..31 slice of a matrix in Shared Memory. This slice is known as a "k-group"
|
||||
within the code because it corresponds to the same K-index of a warp-level matrix multiply.
|
||||
|
||||

|
||||
|
||||
The lower half of the figure shows the physical arrangement in Shared Memory, with threads offset by row and column
|
||||
according to the XOR function. By inspection, we can observe there are no bank conflicts, as _T0 ... T7_ each access unique
|
||||
banks, as do _T8 ... T15_. and beyond.
|
||||
|
||||
To advance to the next "k-group" within Shared Memory, pointers are updated using an XOR operation according to
|
||||
the following sequence:
|
||||
- **^1** advances from _k=0_ to _k=1_
|
||||
- **^3** advances from _k=1_ to _k=2_
|
||||
- **^1** advances from _k=2_ to _k=3_
|
||||
- **^3** advances from _k=3_ to _k=0_
|
||||
|
||||
The first of these transitions is shown below.
|
||||

|
||||
|
||||
The [CUTLASS warp-level GEMM API](gemm_api.md#warp-level-matrix-multiply-api) defines templates for
|
||||
loading slices of data from permuted Shared Memory and issuing operations to Tensor Cores.
|
||||
|
||||
### Updating the Output Tensor
|
||||
|
||||
After the mainloop terminates, the accumulator tile of the warp-level GEMM stores a warp's contribution to the output
|
||||
tensor. However, the distribution of data among threads within the threadblock is specialized for efficient matrix multiply-accumulate
|
||||
operations using Tensor Cores and is not conducive to efficient, coalesced operations to Global Memory. A data rearrangement is
|
||||
needed.
|
||||
|
||||
The **Epilogue** is the component for exchanging accumulator elements through Shared Memory, loading slices of the output
|
||||
matrix or tensor, applying an elementwise operation such as linear scaling or bias, and storing the result to the output tensor.
|
||||
CUTLASS structures this as several components:
|
||||
- [cutlass::epilogue::threadblock::Epilogue](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/epilogue/threadblock/epilogue.h) - the top-level component for looping over the entire threadblock tile
|
||||
- [cutlass::epilogue::warp::TileIteratorTensorOp](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/epilogue/warp/tile_iterator_tensor_op.h) - a specialized component for storing accumulators for Tensor Core to Shared Memory
|
||||
- [cutlass::epilogue::threadblock::SharedLoadIterator](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/epilogue/threadblock/shared_load_iterator.h) - a component for loading elements from a row-major arrangement in Shared Memory
|
||||
- [cutlass::epilogue::threadblock::PredicatedTileIterator](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/epilogue/threadblock/predicated_tile_iterator.h) - a component for loading or storing matrix fragments to Global Memory (with bounds checks)
|
||||
- [cutlass::epilogue::thread::LinearCombination](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/epilogue/thread/linear_combination.h) - an element-wise function computing `alpha * AB + beta * C` to compute the final output
|
||||
|
||||
## Unit Tests
|
||||
|
||||
Unit tests verify the functional behavior of each of the above components in a standalone CUDA kernel. This provides a
|
||||
convenient environment to
|
||||
|
||||
a. inspect the template definition,
|
||||
b. showcase instantiation of use of these templates in device code, and
|
||||
c. assert functional correctness.
|
||||
|
||||
**Convolution unit tests**
|
||||
- Device-wide convolution operator: [conv2d_fprop_implicit_gemm_s4nhwc_s4nhwc_s32nhwc_tensor_op_s32_sm75.cu](https://github.com/NVIDIA/cutlass/tree/main/test/unit/conv/device/conv2d_fprop_implicit_gemm_s4nhwc_s4nhwc_s32nhwc_tensor_op_s32_sm75.cu)
|
||||
|
||||
**GEMM unit tests**
|
||||
- Warp-scoped matrix multiply for Turing Tensor Cores: [gemm_sm75.cu](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/warp/gemm_sm75.cu)
|
||||
|
||||
**Epilogue unit tests**
|
||||
- Epilogue for Turing Tensor Cores: [epilogue_tensor_op.cu](https://github.com/NVIDIA/cutlass/tree/main/test/unit/epilogue/threadblock/epilogue_tensor_op.cu)
|
||||
|
||||
|
||||
# Convolution Example
|
||||
|
||||
This section describes the provided convolution example and is intended to orient the reader to the CUTLASS implementation
|
||||
of Implicit GEMM Convolution.
|
||||
|
||||
## Building and Running the Example
|
||||
|
||||
Example `09_turing_tensorop_conv2dfprop` computes a forward convolutional layer in which inputs and
|
||||
outputs are 4-b integers. The example source is visible in
|
||||
[examples/09_turing_tensorop_conv2dfprop/turing_tensorop_conv2dfprop.cu](https://github.com/NVIDIA/cutlass/tree/main/examples/09_turing_tensorop_conv2dfprop/turing_tensorop_conv2dfprop.cu).
|
||||
|
||||
|
||||
Before building the example, first perform the prerequisite steps for building any CUTLASS component [described here](quickstart.md).
|
||||
Compute capability 7.5 refers to the Turing architecture, and this work requires CUDA 10.2 Toolkit or later to target
|
||||
Turing Tensor Cores using the native `mma` [PTX instruction](https://docs.nvidia.com/cuda/parallel-thread-execution/index.html#warp-level-matrix-fragment-mma-8832).
|
||||
|
||||
```bash
|
||||
$ mkdir build && cd build
|
||||
|
||||
$ cmake .. -DCUTLASS_NVCC_ARCHS=75
|
||||
```
|
||||
|
||||
To build the example, execute `make 09_turing_tensorop_conv2dfprop` from the build directory.
|
||||
```bash
|
||||
$ make 09_turing_tensorop_conv2dfprop
|
||||
|
||||
$ ls examples/09_turing_tensorop_conv2dfprop
|
||||
examples/09_turing_tensorop_conv2dfprop
|
||||
|
||||
```
|
||||
|
||||
This example provides a simple command line interface to specify the extents of 4D tensors of 4-bit integer elements (`cutlass::int4b_t`),
|
||||
initialize them to random values, and compute the result of a convolutional layer. Optionally, the input and output
|
||||
tensors may be saved to .csv files, and the CUTLASS host-side reference check may be executed to verify correctness.
|
||||
|
||||
The complete usage statement is visible by running with `--help`:
|
||||
```
|
||||
$ ./examples/09_turing_tensorop_conv2dfprop/09_turing_tensorop_conv2dfprop --help
|
||||
09_turing_tensorop_conv2dfprop example
|
||||
|
||||
This example uses Turing's Tensor Core operators on int4 data types to compute
|
||||
forward convolution on tensors of layout NHWC.
|
||||
|
||||
Options:
|
||||
|
||||
--help If specified, displays this usage statement.
|
||||
|
||||
--n <int> Input tensor extent N
|
||||
--h <int> Input tensor extent H
|
||||
--w <int> Input tensor extent W
|
||||
--c <int> Input tensor extent C
|
||||
--k <int> Filter extent K
|
||||
--r <int> Filter extent R
|
||||
--s <int> Filter extent S
|
||||
|
||||
--alpha <float> Epilogue scalar alpha
|
||||
--beta <float> Epilogue scalar beta
|
||||
|
||||
--ref-check If set (true), reference check on the host is computed
|
||||
--perf-check If set (true), performance is measured.
|
||||
--benchmark If set (true), performance benchmarking on several layers and batch-size.
|
||||
--iterations <int> Number of profiling iterations to perform.
|
||||
--save-workspace If set, workspace is written to a text file.
|
||||
--tag <string> String to replicate across the first column in the results table
|
||||
|
||||
|
||||
|
||||
Examples:
|
||||
|
||||
$ ./examples/09_turing_tensorop_conv2dfprop/09_turing_tensorop_conv2dfprop --n=32 --h=224 --w=224 --c=128 --k=256 --r=1 --s=1
|
||||
|
||||
$ ./examples/09_turing_tensorop_conv2dfprop/09_turing_tensorop_conv2dfprop --n=1 --h=224 --w=224 --c=32 --k=32 --r=3 --s=3 --ref-check
|
||||
```
|
||||
|
||||
*Note*, this example assumes all tensors are 128b aligned and in format _NHWC_. Consequently, dimension
|
||||
_C_ must be divisible by 32 for activations, filters, and output.
|
||||
|
||||
If the option `--benchmark` is passed, several layers from ResNet50 are profiled for various batch sizes.
|
||||
This sample output was computed on an NVIDIA RTX 2080 compiled with CUDA 10.2.
|
||||
|
||||
```bash
|
||||
build$ ./examples/09_turing_tensorop_conv2dfprop/09_turing_tensorop_conv2dfprop --benchmark
|
||||
```
|
||||
|
||||
Convolution can also be run by the CUTLASS Profiler.
|
||||
|
||||
|
||||
# Copyright
|
||||
|
||||
Copyright (c) 2017 - 2025 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
```
|
||||
Redistribution and use in source and binary forms, with or without
|
||||
modification, are permitted provided that the following conditions are met:
|
||||
|
||||
1. Redistributions of source code must retain the above copyright notice, this
|
||||
list of conditions and the following disclaimer.
|
||||
|
||||
2. Redistributions in binary form must reproduce the above copyright notice,
|
||||
this list of conditions and the following disclaimer in the documentation
|
||||
and/or other materials provided with the distribution.
|
||||
|
||||
3. Neither the name of the copyright holder nor the names of its
|
||||
contributors may be used to endorse or promote products derived from
|
||||
this software without specific prior written permission.
|
||||
|
||||
THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS "AS IS"
|
||||
AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE
|
||||
IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE ARE
|
||||
DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT HOLDER OR CONTRIBUTORS BE LIABLE
|
||||
FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL
|
||||
DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR
|
||||
SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER
|
||||
CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY,
|
||||
OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE
|
||||
OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
|
||||
```
|
||||
301
media/docs/cpp/layout.md
Normal file
301
media/docs/cpp/layout.md
Normal file
@@ -0,0 +1,301 @@
|
||||

|
||||
|
||||
Note: This document talks about CUTLASS 2.x layout tag types.
|
||||
CUTLASS 3.0 deprecates all legacy 2.x layout tags in favour of a single `cute::Layout<Shape, Stride>`
|
||||
vocabulary type for all thread and data tensors. Please refer to the
|
||||
[documentation for cute layouts](cute/01_layout.md) for more details about CUTLASS 3.0's definition of "layout".
|
||||
|
||||
# Layouts and Tensors
|
||||
|
||||
_Tensors_ are mathematical objects represented by a multidimensional array of numeric elements in memory.
|
||||
These may define two dimensional matrices upon which classical linear algebra computations may be defined or
|
||||
higher dimensional objects frequently used to structure data used by Deep Learning applications and frameworks.
|
||||
|
||||
This document describes design patterns used in CUTLASS to map logical index spaces onto memory (Layouts) and to
|
||||
indirectly reference tensors in memory (TensorRef and TensorView objects).
|
||||
|
||||
As described, CUTLASS adheres to the following terminology which is consistent with the C++ Standard Library.
|
||||
|
||||
* *size* (scalar): number of elements in a tensor
|
||||
* *capacity* (scalar): number of elements needed to represent tensor in memory (may be larger than _size_)
|
||||
* *rank* (scalar): number of logical dimensions describing tensor
|
||||
* *extent* (vector): size of each logical dimension in a tensor
|
||||
|
||||
## CUTLASS Layout Concept
|
||||
|
||||
CUTLASS Layouts are a systematic design pattern for the following:
|
||||
* Mapping _logical_ index space to _physical_ offsets in memory
|
||||
* Storing the dynamic state needed in the above computation
|
||||
* Defining a type system for partial specialization of other CUTLASS components
|
||||
|
||||
_Concept:_ layouts satisfy the following concept.
|
||||
```c++
|
||||
/// CUTLASS Layout concept example
|
||||
struct LayoutConcept {
|
||||
|
||||
/// Logical rank of tensor
|
||||
static int const kRank;
|
||||
|
||||
/// Rank of stride vector
|
||||
static int const kStrideRank;
|
||||
|
||||
/// Index type used for coordinates
|
||||
struct Index;
|
||||
|
||||
/// Long index type used for offsets
|
||||
struct LongIndex;
|
||||
|
||||
/// Logical coordinate - satisfies Coord<kRank, ..>
|
||||
struct TensorCoord;
|
||||
|
||||
/// Stride object - satisfies Coord<kStrideRank, ..>
|
||||
struct Stride
|
||||
|
||||
//
|
||||
// Methods
|
||||
//
|
||||
|
||||
/// Constructor
|
||||
CUTLASS_HOST_DEVICE
|
||||
LayoutConcept();
|
||||
|
||||
/// Ctor
|
||||
CUTLASS_HOST_DEVICE
|
||||
LayoutConcept(Stride stride);
|
||||
|
||||
/// Helper returns a layout to a tightly packed tensor
|
||||
CUTLASS_HOST_DEVICE
|
||||
static LayoutConcept packed(TensorCoord const &extent);
|
||||
|
||||
/// Function call operator returns the offset of a coordinate in linear memory.
|
||||
/// Assumes coordinate has convention (row, column)
|
||||
CUTLASS_HOST_DEVICE
|
||||
LongIndex operator()(TensorCoord const &coord) const;
|
||||
|
||||
/// Inverse of layout function, mapping linear offset to logical coordinate
|
||||
CUTLASS_HOST_DEVICE
|
||||
TensorCoord inverse(LongIndex offset) const;
|
||||
|
||||
/// Returns the stride of the layout
|
||||
CUTLASS_HOST_DEVICE
|
||||
Stride stride() const;
|
||||
|
||||
/// Returns the stride of the layout
|
||||
CUTLASS_HOST_DEVICE
|
||||
Stride & stride();
|
||||
|
||||
/// Compute the number of contiguous elements needed to store a tensor with the given size
|
||||
CUTLASS_HOST_DEVICE
|
||||
LongIndex capacity(TensorCoord const &extent) const;
|
||||
};
|
||||
```
|
||||
|
||||
_Layout_ objects generalize leading dimensions of matrices typical in _BLAS_ implementations. For example, cuBLAS assumes
|
||||
Fortran-style _column-major_ layouts of matrices and refers to this as the matrix's "leading dimension."
|
||||
|
||||
```c++
|
||||
cublasGemmEx(
|
||||
...
|
||||
ptr_A, // pointer to first element of matrix A
|
||||
lda, // leading dimension
|
||||
...
|
||||
);
|
||||
```
|
||||
This implies an element at coordinate (_row_, _column_) has offset `row + lda * column`.
|
||||
|
||||
This is equivalently represented by CUTLASS's `layout::ColumnMajor` type as follows.
|
||||
```c++
|
||||
|
||||
layout::ColumnMajor layout(lda);
|
||||
|
||||
int offset = layout({row, column}); // returns row + lda * column
|
||||
```
|
||||
|
||||
Other layout functions are possible such as row-major:
|
||||
```c++
|
||||
|
||||
layout::RowMajor layout(lda);
|
||||
|
||||
int offset = layout({row, column}); // returns lda * row + column
|
||||
```
|
||||
|
||||
In both cases, the _logical_ coordinate (_row_, _column_) is represented by the same object. This enables an algorithm to be
|
||||
implemented as generic template, with locations within tensors always specified in logical space. _Layout_ objects map this to
|
||||
physical offsets in memory.
|
||||
|
||||
The layout's `::packed()` static method may be used to construct a layout object given the extent of a densely packed tensor.
|
||||
This method is needed when an algorithm must define a buffer of arbitrary layout.
|
||||
|
||||
Example:
|
||||
```c++
|
||||
|
||||
typename ArbitraryLayout::TensorCoord extent = make_Coord(...);
|
||||
typename ArbitraryLayout::TensorCoord coord;
|
||||
|
||||
ArbitraryLayout layout = ArbitraryLayout::packed(extent);
|
||||
|
||||
int offset = layout({coord});
|
||||
```
|
||||
|
||||
The layout's `::capacity()` method computes the number of locations in memory needed to represent a tensor. This is
|
||||
useful when allocating memory, as more storage may be needed than what is strictly necessary for a fully packed
|
||||
tensor.
|
||||
|
||||
Example:
|
||||
```c++
|
||||
|
||||
int lda = columns + padding;
|
||||
MatrixCoord extent{rows, columns};
|
||||
|
||||
layout::RowMajor layout(lda);
|
||||
|
||||
auto capacity = layout.capacity(extent); // returns rows * (columns + padding)
|
||||
```
|
||||
|
||||
## Accessing elements within a tensor
|
||||
|
||||
### TensorRef
|
||||
|
||||
`TensorRef<class T, class Layout>` is a structure containing both a pointer to the start of a
|
||||
tensor and a layout object to access its elements. This is a convenient object which may be
|
||||
passed to functions to limit an explosion of arguments when the number of stride elements is
|
||||
numerous.
|
||||
|
||||
Example:
|
||||
```c++
|
||||
int4_t *ptr = ...;
|
||||
int ldm = ...;
|
||||
|
||||
int row = ...;
|
||||
int column = ...;
|
||||
|
||||
layout::ColumnMajor layout(ldm);
|
||||
TensorRef<int4_t, layout::ColumnMajor> ref(ptr, layout);
|
||||
|
||||
int4_t x = ref.at({row, column}); // loads a 4-bit signed integer from the tensor
|
||||
|
||||
ref.at({row, column}) = x * 2_s4; // transforms this quantity and stores it back
|
||||
```
|
||||
|
||||
### TensorView
|
||||
|
||||
Matrices and tensors used in linear algebra computations are invariably finite. `TensorView<class T, class Layout>` extends `TensorRef<>` by
|
||||
adding an `extent` vector to describe the logical extent of the tensor or matrix.
|
||||
|
||||
Example:
|
||||
```c++
|
||||
int4_t *ptr = ...;
|
||||
int ldm = ...;
|
||||
MatrixCoord extent = ...;
|
||||
|
||||
int row = ...;
|
||||
int column = ...;
|
||||
|
||||
layout::ColumnMajor layout(ldm);
|
||||
TensorView<int4_t, layout::ColumnMajor> view(ptr, layout, extent);
|
||||
|
||||
MatrixCoord coord = {row, column};
|
||||
|
||||
if (view.contains(coord)) { // verify coordinate is in bounds before performing access
|
||||
|
||||
int4_t x = ref.at(coord);
|
||||
ref.at({row, column}) = x * 2_s4;
|
||||
}
|
||||
|
||||
```
|
||||
|
||||
A `TensorView<>` may be constructed from a `TensorRef<>` succinctly as follows:
|
||||
```c++
|
||||
layout::ColumnMajor layout(ldm);
|
||||
TensorRef<int4_t, layout::ColumnMajor> ref(ptr, layout);
|
||||
|
||||
TensorView<int4_t, layout::ColumnMajor> view(ref, extent); // construct TensorView from TensorRef and extent
|
||||
```
|
||||
|
||||
Note, computations avoid becoming overdetermined by accepting a single problem size component
|
||||
and `TensorRef` objects for each of the operands whose extents are implied as a precondition of the operation. By avoiding
|
||||
redundant storage of extent quantities, CUTLASS minimizes capacity utilization of precious resources such as constant memory.
|
||||
This is consistent with BLAS conventions.
|
||||
|
||||
# Summary:
|
||||
|
||||
The design patterns described in this document form a hierarchy:
|
||||
* `T *ptr;` is a pointer to a contiguous sequence of elements of type `T`
|
||||
* `Layout layout;` is an object mapping an index space to a linear offset
|
||||
* `TensorRef<T, Layout> ref(ptr, layout);` is an object pointing to an _unbounded_ tensor containing elements of type `T` and a layout of type `Layout`
|
||||
* `TensorView<T, Layout> view(ref, extent);` is an object pointing to a _bounded_ tensor containing elements of type `T` and a layout of type `Layout`
|
||||
|
||||
# Appendix: Existing Layouts
|
||||
|
||||
This section enumerates several existing Layout types defined in CUTLASS.
|
||||
|
||||
Matrix layouts:
|
||||
- `PitchLinear`: data layout defined by _contiguous_ and _strided_ dimensions. _contiguous_ refers to consecutive elements in memory, where as _strided_ refers to data separated by a uniform stride
|
||||
-- Rank: 2
|
||||
-- TensorCoord type: `PitchLinearCoord`
|
||||
-- Shape type: `PitchLinearShape`
|
||||
-- Stride rank: 1
|
||||
|
||||
- `ColumnMajor`: data layout defined by _rows_ and _columns_ dimensions. Can be mapped to `PitchLinear` by: (_contiguous_ = _rows_, _strided_ = _columns_)
|
||||
-- Rank: 2
|
||||
-- TensorCoord type: `MatrixCoord`
|
||||
-- Shape type: `MatrixShape`
|
||||
-- Stride rank: 1
|
||||
|
||||
- `RowMajor`: data layout defined by _rows_ and _columns_ dimensions. Can be mapped to `PitchLinear` by: (_contiguous_ = _columns_, _strided_ = _rows_)
|
||||
-- Rank: 2
|
||||
-- TensorCoord type: `MatrixCoord`
|
||||
-- Shape type: `MatrixShape`
|
||||
-- Stride rank: 1
|
||||
|
||||
- `ColumnMajorInterleaved<k>`: data layout defined by _rows_ and _columns_ dimensions. Data is packed into a 'column-major' arrangement of row vectors of fixed length.
|
||||
-- Rank: 2
|
||||
-- TensorCoord type: `MatrixCoord`
|
||||
-- Shape type: `MatrixShape`
|
||||
-- Stride rank: 1
|
||||
|
||||
- `RowMajorInterleaved<k>`: data layout defined by _rows_ and _columns_ dimensions. Data is packed into a 'row-major' arrangement of column vectors of fixed length.
|
||||
-- Rank: 2
|
||||
-- TensorCoord type: `MatrixCoord`
|
||||
-- Shape type: `MatrixShape`
|
||||
-- Stride rank: 1
|
||||
|
||||
Tensor layouts:
|
||||
- `TensorNHWC`:
|
||||
|
||||
Permuted Shared Memory Layouts:
|
||||
- `TensorOpCongruous<ElementSize>`
|
||||
- `TensorOpCrosswise<ElementSize>`
|
||||
|
||||
|
||||
# Copyright
|
||||
|
||||
Copyright (c) 2017 - 2025 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
```
|
||||
Redistribution and use in source and binary forms, with or without
|
||||
modification, are permitted provided that the following conditions are met:
|
||||
|
||||
1. Redistributions of source code must retain the above copyright notice, this
|
||||
list of conditions and the following disclaimer.
|
||||
|
||||
2. Redistributions in binary form must reproduce the above copyright notice,
|
||||
this list of conditions and the following disclaimer in the documentation
|
||||
and/or other materials provided with the distribution.
|
||||
|
||||
3. Neither the name of the copyright holder nor the names of its
|
||||
contributors may be used to endorse or promote products derived from
|
||||
this software without specific prior written permission.
|
||||
|
||||
THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS "AS IS"
|
||||
AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE
|
||||
IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE ARE
|
||||
DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT HOLDER OR CONTRIBUTORS BE LIABLE
|
||||
FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL
|
||||
DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR
|
||||
SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER
|
||||
CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY,
|
||||
OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE
|
||||
OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
|
||||
```
|
||||
619
media/docs/cpp/overview.md
Normal file
619
media/docs/cpp/overview.md
Normal file
@@ -0,0 +1,619 @@
|
||||

|
||||
|
||||
# Overview
|
||||
|
||||
# CUTLASS 3.9.0
|
||||
|
||||
_CUTLASS 3.9.0 - March 2025_
|
||||
|
||||
CUTLASS is a collection of CUDA C++ template abstractions for implementing
|
||||
high-performance matrix-matrix multiplication (GEMM) and related computations at all levels
|
||||
and scales within CUDA. It incorporates strategies for hierarchical decomposition and
|
||||
data movement similar to those used to implement cuBLAS and cuDNN. CUTLASS decomposes
|
||||
these "moving parts" into reusable, modular software components abstracted by C++ template
|
||||
classes. Primitives for different levels of a conceptual parallelization hierarchy
|
||||
can be specialized and tuned via custom tiling sizes, data types,
|
||||
and other algorithmic policy. The resulting flexibility simplifies their use
|
||||
as building blocks within custom kernels and applications.
|
||||
|
||||
To support a wide variety of applications, CUTLASS provides extensive support for
|
||||
mixed-precision computations, providing specialized data-movement and
|
||||
multiply-accumulate abstractions for FP64, FP32, TF32, FP16, BF16,
|
||||
[FP32 emulation via tensor core instruction](https://github.com/NVIDIA/cutlass/tree/main/examples/27_ampere_3xtf32_fast_accurate_tensorop_gemm),
|
||||
8b floating point types (e5m2 and e4m3),
|
||||
block scaled data types (NVIDIA NVFP4 and OCP standard MXFP4, MXFP6, MXFP8),
|
||||
narrow integer types (4 and 8b signed and unsigned integers),
|
||||
and binary 1b data types (where architectures allow for the
|
||||
native support of such data types).
|
||||
CUTLASS demonstrates optimal matrix multiply operations
|
||||
targeting the programmable, high-throughput _Tensor Cores_ implemented by
|
||||
NVIDIA's Volta, Turing, Ampere, Ada, Hopper, and Blackwell architectures.
|
||||
|
||||
In addition to GEMMs, CUTLASS implements high-performance convolution via
|
||||
the implicit GEMM algorithm. Implicit GEMM is the formulation of a convolution
|
||||
operation as a GEMM thereby taking advantage of CUTLASS's modular GEMM pipeline.
|
||||
This allows CUTLASS to build convolutions by reusing highly-optimized GEMM components.
|
||||
|
||||
See the [Quick Start Guide](quickstart.md) to get started quickly.
|
||||
|
||||
See the [functionality docs](functionality.md) for a more comprehensive
|
||||
list of kernel level features, data types, instructions, and minimum supported by CUTLASS on each GPU
|
||||
architecture.
|
||||
|
||||
# What's New in CUTLASS 3.9
|
||||
|
||||
* Support for Blackwell SM120 kernels for GeForce GPUs in CUTLASS 3.x API:
|
||||
- Collective mainloops that target for:
|
||||
* [Blockscaled datatypes with support for dense GEMM](../../../include/cutlass/gemm/collective/sm120_blockscaled_mma_tma.hpp)
|
||||
* [Blockscaled datatypes with support for sparse GEMM](../../../include/cutlass/gemm/collective/sm120_blockscaled_sparse_mma_tma.hpp)
|
||||
- New [GEMM](../../../include/cutlass/gemm/dispatch_policy.hpp) and [epilogue](../../../include/cutlass/epilogue/dispatch_policy.hpp) dispatch policies for collectives, kernel layers, and builders.
|
||||
- [Blackwell SM120 epilogue](../../../include/cutlass/epilogue/fusion/sm120_visitor_store_tma_warpspecialized.hpp) and [full set of EVT fusions](../../../include/cutlass/epilogue/fusion/sm120_callbacks_tma_warpspecialized.hpp).
|
||||
* Set of examples that demonstrate the usage of the 3.x API for targeting Blackwell SM120 architecture:
|
||||
- [Blockscaled GEMM with NVFP4 input datatype and BF16 output tensor](../../../examples/79_blackwell_geforce_gemm/79a_blackwell_geforce_nvfp4_bf16_gemm.cu).
|
||||
- [Blockscaled GEMM with NVFP4 input datatype and NVFP4 output tensor with scale factor generation](../../../examples/79_blackwell_geforce_gemm/79b_blackwell_geforce_nvfp4_nvfp4_gemm.cu).
|
||||
- [Blockscaled GEMM with mixed input datatype (MXFP8 and MXFP6) and BF16 output tensor](../../../examples/79_blackwell_geforce_gemm/79c_blackwell_geforce_mixed_mxfp8_mxfp6_bf16_gemm.cu).
|
||||
* Set of unit tests that demonstrate the usage of both [sparse](../../../test/unit/gemm/device/sm120_blockscaled_sparse_tensorop_gemm/) and [dense](../../../test/unit/gemm/device/sm120_blockscaled_tensorop_gemm/) Blackwell SM120 blockscaled GEMM.
|
||||
* Enhancement and new support of block-wise and group-wise GEMM for Hopper and Blackwell architectures:
|
||||
- Enhancement of [blockwise GEMM](../../../examples/67_hopper_fp8_warp_specialized_gemm_with_blockwise_scaling/67_hopper_fp8_warp_specialized_gemm_with_blockwise_scaling.cu) for Hopper architecture.
|
||||
- Enhancement of [groupwise GEMM](../../../examples/67_hopper_fp8_warp_specialized_gemm_with_blockwise_scaling/67_hopper_fp8_warp_specialized_gemm_with_groupwise_scaling.cu) for Hopper architecture.
|
||||
- Support for [grouped GEMM with blockwise and groupwise scaling](../../../examples/68_hopper_fp8_warp_specialized_grouped_gemm_with_blockwise_scaling/) for Hopper architecture.
|
||||
- Support for [blockwise GEMM](../../../examples/81_blackwell_gemm_blockwise/81_blackwell_gemm_blockwise.cu) for Blackwell architecture.
|
||||
- Support for [groupwise GEMM](../../../examples/81_blackwell_gemm_blockwise/81_blackwell_gemm_groupwise.cu) for Blackwell architecture.
|
||||
- Support for [grouped GEMM with blockwise](../../../examples/81_blackwell_gemm_blockwise/81_blackwell_grouped_gemm_blockwise.cu) and [groupwise scaling](../../../examples/81_blackwell_gemm_blockwise/81_blackwell_grouped_gemm_groupwise.cu) for Blackwell architecture.
|
||||
* Added support for enhanced kernel performance search (auto-tuning) in CUTLASS profiler:
|
||||
- Sorting performance results by GFLOPs/second: Users can now sort the final performance report based on GFLOPs/second, making it easier to identify the most efficient kernels.
|
||||
- Exhaustive search for best kernel performance in GFLOPs/second: The profiler now searches for the best-performing kernel across a range of problem sizes, swizzle sizes, rasterization orders, and dynamic cluster configurations to maximize performance.
|
||||
- Performance search under a fixed GEMM shape: Enables exhaustive tuning within a fixed GEMM shape, exploring various kernel parameters to find the best configuration.
|
||||
- More detailed introductions and examples to leverage this feature can be found in [profiler.md](./profiler.md#exhaustive-search-mode-and-top-k-output-ranking-according-to-performance-in-gflopss).
|
||||
|
||||
Note: CUTLASS 3.x builds are known to be down on Windows platforms for all CUDA toolkits.
|
||||
CUTLASS team is working on a fix.
|
||||
|
||||
**See the [CHANGELOG](../release_notes.md) for details of all past releases and updates.**
|
||||
|
||||
# Performance
|
||||
|
||||
CUTLASS primitives are very efficient. When used to construct device-wide GEMM kernels,
|
||||
they exhibit nearly optimal utilization of peak theoretical throughput. The figure below
|
||||
shows CUTLASS 3.8's performance as a % of theoretical peak utilization
|
||||
on various input and output data types when run on NVIDIA Blackwell SM100 architecture GPU.
|
||||
|
||||

|
||||
|
||||
The two figures below show the continual CUTLASS performance improvements
|
||||
on an [NVIDIA H100](https://www.nvidia.com/en-us/data-center/h100/) (NVIDIA Hopper architecture) since
|
||||
CUTLASS 3.1.
|
||||
CUTLASS 3.5.1 was compiled with the [CUDA 12.5u1 Toolkit](https://developer.nvidia.com/cuda-downloads).
|
||||
Tensor Core operations are implemented using CUDA's
|
||||
[mma](https://docs.nvidia.com/cuda/parallel-thread-execution/index.html#warp-level-matrix-instructions-mma) and
|
||||
[wgmma](https://docs.nvidia.com/cuda/parallel-thread-execution/index.html#asynchronous-warpgroup-level-matrix-instructions) instructions.
|
||||
|
||||

|
||||

|
||||
|
||||
# CuTe
|
||||
|
||||
CUTLASS 3.0 introduced a new core library, CuTe, to describe and manipulate tensors of threads and data.
|
||||
CuTe is a collection of C++ CUDA template abstractions for
|
||||
defining and operating on hierarchically multidimensional layouts of threads and data.
|
||||
CuTe provides `Layout` and `Tensor` objects that compactly package the type,
|
||||
shape, memory space, and layout of data, while performing the complicated indexing for the user.
|
||||
This lets programmers focus on the logical descriptions of their algorithms while
|
||||
CuTe does the mechanical bookkeeping for them. With these tools, we can quickly design,
|
||||
implement, and modify all dense linear algebra operations.
|
||||
|
||||
The core abstractions of CuTe are hierarchically multidimensional layouts
|
||||
which can be composed with data arrays to represent tensors.
|
||||
The representation of layouts is powerful enough to represent nearly
|
||||
everything we need to implement efficient dense linear algebra.
|
||||
Layouts can also be combined and manipulated via functional composition, on which we build a large set of common operations such as tiling and partitioning.
|
||||
|
||||
CUTLASS 3.0 and beyond adopts CuTe throughout the GEMM hierarchy in its templates.
|
||||
This greatly simplifies the design and improves code composability and readability.
|
||||
More documentation specific to CuTe can be found in its
|
||||
[dedicated documentation directory](cute/00_quickstart.md).
|
||||
|
||||
# Compatibility
|
||||
|
||||
Minimum requirements:
|
||||
|
||||
- Architecture: Volta (compute capability 7.0)
|
||||
- Compiler: Must support at least C++17
|
||||
- CUDA Toolkit version: 11.4
|
||||
|
||||
CUTLASS requires a C++17 host compiler and
|
||||
performs best when built with the [**CUDA 12.8 Toolkit**](https://developer.nvidia.com/cuda-downloads).
|
||||
It is also compatible with CUDA 11.4, CUDA 11.5, CUDA 11.6, CUDA 11.7, CUDA 11.8, and all other CUDA 12.x versions.
|
||||
|
||||
## Operating Systems
|
||||
|
||||
We have tested the following environments.
|
||||
|
||||
|**Operating System** | **Compiler** |
|
||||
|-----------------|----------|
|
||||
| Ubuntu 18.04 | GCC 7.5.0 |
|
||||
| Ubuntu 20.04 | GCC 10.3.0 |
|
||||
| Ubuntu 22.04 | GCC 11.2.0 |
|
||||
|
||||
Note: GCC 8.5.0 has known regressions regarding fold expressions and overloaded operators. Using GCC 7.5.0 or (preferred) GCC >= 9 is recommended.
|
||||
|
||||
Note: CUTLASS 3.x builds are known to be down on Windows platforms for all CUDA toolkits.
|
||||
CUTLASS team is working on a fix.
|
||||
|
||||
## Hardware
|
||||
|
||||
CUTLASS runs successfully on the following NVIDIA GPUs, and it is expected to be efficient on Volta, Turing, Ampere, Ada, and Hopper architecture based NVIDIA GPUs.
|
||||
|
||||
|**GPU**|**CUDA Compute Capability**|**Minimum CUDA Toolkit Required by CUTLASS-3**|
|
||||
|---|---|---|
|
||||
|NVIDIA V100 Tensor Core GPU |7.0|11.4|
|
||||
|NVIDIA TitanV |7.0|11.4|
|
||||
|NVIDIA GeForce RTX 20x0 series |7.5|11.4|
|
||||
|NVIDIA T4 |7.5|11.4|
|
||||
|NVIDIA A100 Tensor Core GPU |8.0|11.4|
|
||||
|NVIDIA A10 |8.6|11.4|
|
||||
|NVIDIA GeForce RTX 30x0 series |8.6|11.4|
|
||||
|NVIDIA GeForce RTX 40x0 series |8.9|11.8|
|
||||
|NVIDIA L40 |8.9|11.8|
|
||||
|NVIDIA H100 Tensor Core GPU |9.0|11.8|
|
||||
|NVIDIA H200 Tensor Core GPU |9.0|11.8|
|
||||
|NVIDIA B200 Tensor Core GPU |10.0|12.8|
|
||||
|NVIDIA GeForce RTX 50x0 series |10.0|12.8|
|
||||
|
||||
## Target Architecture
|
||||
|
||||
In general, PTX code generated for one target architecture can be run on future architectures
|
||||
(i.e., it is forward compatible).
|
||||
However, CUDA 12.0 introduced the concept of "architecture-accelerated features" whose
|
||||
PTX does not have forward compatibility guarantees.
|
||||
Several Hopper and Blackwell PTX instructions fall under this category of
|
||||
architecture-accelerated features, and thus require a `sm_90a` or `sm100a` target architecture
|
||||
(note the "a" appended). For more details on this and other architecture-accelerated instructions,
|
||||
please refer to the [CUDA Documentation](https://docs.nvidia.com/cuda/cuda-c-programming-guide/index.html#feature-availability).
|
||||
|
||||
The target architecture information is passed on to CUTLASS via the cmake flag
|
||||
`CUTLASS_NVCC_ARCHS`. In order to maximize performance on Hopper GH100,
|
||||
users are required to build CUTLASS with `90a` as the target architecture.
|
||||
If a user accidentally builds a kernel which uses SM90a features
|
||||
(e.g. Hopper Tensor Core Instructions), using the SM90 target
|
||||
(note the lack of "a"), with either CUDA Toolkit 12 or 11.8,
|
||||
the kernel is expected to fail with a runtime error.
|
||||
|
||||
```
|
||||
cmake .. -DCUTLASS_NVCC_ARCHS="90a"
|
||||
```
|
||||
Or
|
||||
|
||||
```
|
||||
cmake .. -DCUTLASS_NVCC_ARCHS="100a"
|
||||
```
|
||||
|
||||
Note: The NVIDIA Blackwell SM100 architecture used in the datacenter
|
||||
products has a different compute capability than the one underpinning
|
||||
NVIDIA Blackwell GeForce RTX 50 series GPUs. As a result, kernels
|
||||
compiled for Blackwell SM100 architecture with arch conditional features
|
||||
(using `sm100a`) are not compatible with RTX 50 series GPUs.
|
||||
|
||||
Please refer to the [functionality documentation](functionality.md)
|
||||
for details on which kernels require which target architectures.
|
||||
|
||||
# Documentation
|
||||
|
||||
CUTLASS is described in the following documents and the accompanying
|
||||
[Doxygen documentation](https://nvidia.github.io/cutlass).
|
||||
|
||||
- [Quick Start Guide](quickstart.md) - basics of building and running CUTLASS
|
||||
- [Functionality](functionality.md) - summarizes functionality available in CUTLASS
|
||||
- [Efficient GEMM in CUDA](efficient_gemm.md) - describes how GEMM kernels may be implemented efficiently in CUDA
|
||||
- [CUTLASS 3.x Design](cutlass_3x_design.md) - describes the CUTLASS 3.x design, its benefits, and how CuTe enables us to write much more composable components
|
||||
- [GEMM API 3.x](gemm_api_3x.md) - describes the CUTLASS 3.x GEMM model and C++ template concepts
|
||||
- [GEMM API 2.x](gemm_api.md) - describes the CUTLASS 2.x GEMM model and C++ template concepts
|
||||
- [Implicit GEMM Convolution](implicit_gemm_convolution.md) - describes 2-D and 3-D convolution in CUTLASS
|
||||
- [Code Organization](code_organization.md) - describes the organization and contents of the CUTLASS project
|
||||
- [Terminology](terminology.md) - describes terms used in the code
|
||||
- [Programming Guidelines](programming_guidelines.md) - guidelines for writing efficient modern CUDA C++
|
||||
- [Fundamental types](fundamental_types.md) - describes basic C++ classes used in CUTLASS to represent numeric quantities and arrays
|
||||
- [Layouts](layout.md) - describes layouts of matrices and tensors in memory
|
||||
- [Tile Iterators](tile_iterator_concept.md) - describes C++ concepts for iterating over tiles of matrices in memory
|
||||
- [CUTLASS Profiler](profiler.md) - command-line driven profiling application
|
||||
- [CUTLASS Utilities](utilities.md) - additional templates used to facilitate rapid development
|
||||
- [Dependent kernel launch](dependent_kernel_launch.md) - describes a new feature in Hopper which allows overlapping dependent
|
||||
kernels in the same stream, and how it is used in CUTLASS.
|
||||
|
||||
# Resources
|
||||
We have also described the structure of an efficient GEMM in our talk at the
|
||||
[GPU Technology Conference 2018](http://on-demand.gputechconf.com/gtc/2018/presentation/s8854-cutlass-software-primitives-for-dense-linear-algebra-at-all-levels-and-scales-within-cuda.pdf).
|
||||
|
||||
- [CUTLASS: Software Primitives for Dense Linear Algebra at All Levels and Scales within CUDA](https://www.nvidia.com/en-us/on-demand/session/gtcsiliconvalley2018-s8854/)
|
||||
- [Developing CUDA Kernels to Push Tensor Cores to the Absolute Limit on NVIDIA A100](https://www.nvidia.com/en-us/on-demand/session/gtcsj20-s21745/)
|
||||
- [Accelerating Convolution with Tensor Cores in CUTLASS](https://www.nvidia.com/en-us/on-demand/session/gtcspring21-s31883/)
|
||||
- [Accelerating Backward Data Gradient by Increasing Tensor Core Utilization in CUTLASS](https://www.nvidia.com/en-us/on-demand/session/gtcspring22-s41996/)
|
||||
- [CUTLASS: Python API, Enhancements, and NVIDIA Hopper](https://www.nvidia.com/en-us/on-demand/session/gtcfall22-a41131/)
|
||||
|
||||
# Building CUTLASS
|
||||
|
||||
CUTLASS is a header-only template library and does not need to be built to be used by other
|
||||
projects. Client applications should target CUTLASS's `include/` directory in their include
|
||||
paths.
|
||||
|
||||
CUTLASS unit tests, examples, and utilities can be build with CMake.
|
||||
The minimum version of CMake is given in the [Quickstart guide](quickstart.md).
|
||||
Make sure the `CUDACXX` environment variable points to NVCC in the CUDA Toolkit installed
|
||||
on your system.
|
||||
|
||||
```bash
|
||||
$ export CUDACXX=${CUDA_INSTALL_PATH}/bin/nvcc
|
||||
```
|
||||
|
||||
Create a build directory within the CUTLASS project, then run CMake. By default CUTLASS will build kernels
|
||||
for CUDA architecture versions 5.0, 6.0, 6.1, 7.0, 7.5, 8.0, 8.6, 8.9, and 9.0.
|
||||
To reduce compile time you can specify
|
||||
the architectures to build CUTLASS for by changing the CMake configuration setting
|
||||
`CUTLASS_NVCC_ARCHS`.
|
||||
|
||||
```bash
|
||||
$ mkdir build && cd build
|
||||
|
||||
$ cmake .. -DCUTLASS_NVCC_ARCHS=80 # compiles for NVIDIA's Ampere Architecture
|
||||
```
|
||||
|
||||
From the `build/` directory, compile and run the CUTLASS unit tests by building the target `test_unit` with make.
|
||||
|
||||
The unit tests are organized as several binaries mirroring the top-level namespaces of CUTLASS,
|
||||
and they may be executed in parallel via make's `-j` command line argument.
|
||||
|
||||
```bash
|
||||
$ make test_unit -j
|
||||
...
|
||||
...
|
||||
...
|
||||
[----------] Global test environment tear-down
|
||||
[==========] 946 tests from 57 test cases ran. (10812 ms total)
|
||||
[ PASSED ] 946 tests.
|
||||
```
|
||||
|
||||
All tests should pass on supported platforms, though the exact number of tests may vary over time.
|
||||
|
||||
|
||||
# Project Structure
|
||||
|
||||
CUTLASS is arranged as a header-only library along with Utilities, Tools, Examples, and unit tests.
|
||||
[Doxygen documentation](https://nvidia.github.io/cutlass) provides a complete list of files, classes,
|
||||
and template concepts defined in the CUTLASS project.
|
||||
|
||||
A detailed explanation of the source code organization may be found in the
|
||||
[CUTLASS documentation](code_organization.md), but several main components are summarized below.
|
||||
|
||||
## CUTLASS Template Library
|
||||
|
||||
```
|
||||
include/ # client applications should target this directory in their build's include paths
|
||||
|
||||
cutlass/ # CUDA Templates for Linear Algebra Subroutines and Solvers - headers only
|
||||
|
||||
arch/ # direct exposure of architecture features (including instruction-level GEMMs)
|
||||
|
||||
conv/ # code specialized for convolution
|
||||
|
||||
epilogue/ # code specialized for the epilogue of gemm/convolution
|
||||
|
||||
gemm/ # code specialized for general matrix product computations
|
||||
|
||||
layout/ # layout definitions for matrices, tensors, and other mathematical objects in memory
|
||||
|
||||
platform/ # CUDA-capable Standard Library components
|
||||
|
||||
reduction/ # bandwidth-limited reduction kernels that do not fit the "gemm" model
|
||||
|
||||
thread/ # simt code that can be performed within a CUDA thread
|
||||
|
||||
transform/ # code specialized for layout, type, and domain transformations
|
||||
|
||||
* # core vocabulary types, containers, and basic numeric operations
|
||||
|
||||
cute/ # CuTe Layout, layout algebra, MMA/Copy atoms, tiled MMA/Copy
|
||||
|
||||
algorithm/ # Definitions of core operations such as copy, gemm, and operations on cute::tuples
|
||||
|
||||
arch/ # Bare bones PTX wrapper structs for copy and math instructions
|
||||
|
||||
atom/ # Meta-information either link to or built from arch/ operators
|
||||
|
||||
mma_atom.hpp # cute::Mma_Atom and cute::TiledMma
|
||||
|
||||
copy_atom.hpp # cute::Copy_Atom and cute::TiledCopy
|
||||
|
||||
*sm*.hpp # Arch specific meta-information for copy and math operations
|
||||
|
||||
* # Core library types such as Shape, Stride, Layout, Tensor, and associated operations
|
||||
|
||||
```
|
||||
|
||||
### CUTLASS SDK Examples
|
||||
|
||||
[CUTLASS SDK examples](https://github.com/NVIDIA/cutlass/tree/main/examples) apply CUTLASS templates to implement basic computations.
|
||||
|
||||
### Tools
|
||||
|
||||
```
|
||||
tools/
|
||||
library/ # CUTLASS Instance Library - contains instantiations of all supported CUTLASS templates
|
||||
include/
|
||||
cutlass/
|
||||
library/
|
||||
|
||||
profiler/ # CUTLASS Profiler - command-line utility for executing operations in the
|
||||
# CUTLASS Library
|
||||
|
||||
util/ # CUTLASS Utilities - contains numerous helper classes for
|
||||
include/ # manging tensors in device memory, reference
|
||||
cutlass/ # implementations for GEMM, random initialization
|
||||
util/ # of tensors, and I/O.
|
||||
```
|
||||
|
||||
### Test
|
||||
|
||||
The `test/unit/` directory consist of unit tests implemented with Google Test that demonstrate
|
||||
basic usage of Core API components and complete tests of the CUTLASS GEMM computations.
|
||||
|
||||
Instructions for building and running the Unit tests are described in the [Quickstart guide](quickstart.md).
|
||||
|
||||
# Performance Profiling
|
||||
|
||||
The `tools/profiler/` directory contains a command-line utility for launching each of the GEMM kernels.
|
||||
It can be built as follows:
|
||||
|
||||
```bash
|
||||
$ make cutlass_profiler -j16
|
||||
```
|
||||
## Building all GEMM and Convolution kernels (_long_ build times)
|
||||
|
||||
By default, only one tile size is instantiated for each data type, math instruction, and layout.
|
||||
To instantiate all, set the following environment variable when running CMake from an empty `build/` directory.
|
||||
Beware, this results in *tens of thousands* of kernels and long build times.
|
||||
This would also result in a large binary size and on some platforms linker to fail on building the library.
|
||||
Therefore, it's highly recommended to generate only a subset of kernels as demonstrated in the sub-section below.
|
||||
```bash
|
||||
$ cmake .. -DCUTLASS_NVCC_ARCHS=90a -DCUTLASS_LIBRARY_KERNELS=all
|
||||
...
|
||||
$ make cutlass_profiler -j16
|
||||
```
|
||||
|
||||
## Building a subset of GEMM and Convolution kernels (_reduced_ build times)
|
||||
|
||||
To compile strictly one kernel or a small set of kernels, a comma-delimited list of kernel names with
|
||||
wildcard characters may be used to reduce the set of kernels. The following examples show building exactly one
|
||||
or a subset of kernels for NVIDIA Ampere and Turing architecture:
|
||||
|
||||
### Building a subset Tensor Core GEMM kernels
|
||||
|
||||
To compile a subset of Tensor Core GEMM kernels with FP32 accumulation and FP16 input targeting NVIDIA Ampere and Turing architecture,
|
||||
use the below cmake command line:
|
||||
```bash
|
||||
$ cmake .. -DCUTLASS_NVCC_ARCHS='75;80' -DCUTLASS_LIBRARY_KERNELS=cutlass_tensorop_s*gemm_f16_*_nt_align8
|
||||
...
|
||||
$ make cutlass_profiler -j16
|
||||
```
|
||||
|
||||
Example command line for profiling a subset of Tensor Core GEMM kernels is as follows:
|
||||
```bash
|
||||
./tools/profiler/cutlass_profiler --kernels=cutlass_tensorop_s*gemm_f16_*_nt_align8 --m=3456 --n=4096 --k=4096
|
||||
|
||||
...
|
||||
=============================
|
||||
Problem ID: 1
|
||||
|
||||
Provider: CUTLASS
|
||||
OperationKind: gemm
|
||||
Operation: cutlass_tensorop_s1688gemm_f16_256x128_32x2_nt_align8
|
||||
|
||||
Status: Success
|
||||
Verification: ON
|
||||
Disposition: Passed
|
||||
|
||||
reference_device: Passed
|
||||
cuBLAS: Passed
|
||||
|
||||
Arguments: --gemm_kind=universal --m=3456 --n=4096 --k=4096 --A=f16:column --B=f16:row --C=f32:column --alpha=1 \
|
||||
--beta=0 --split_k_slices=1 --batch_count=1 --op_class=tensorop --accum=f32 --cta_m=256 --cta_n=128 \
|
||||
--cta_k=32 --stages=2 --warps_m=4 --warps_n=2 --warps_k=1 --inst_m=16 --inst_n=8 --inst_k=8 --min_cc=75 \
|
||||
--max_cc=1024
|
||||
|
||||
Bytes: 118489088 bytes
|
||||
FLOPs: 115992428544 flops
|
||||
|
||||
Runtime: 1.55948 ms
|
||||
Memory: 70.7616 GiB/s
|
||||
|
||||
Math: 74378.8 GFLOP/s
|
||||
|
||||
|
||||
|
||||
=============================
|
||||
...
|
||||
```
|
||||
|
||||
### Building one CUDA Core GEMM kernel
|
||||
|
||||
To compile one SGEMM kernel targeting NVIDIA Ampere and Turing architecture, use the below cmake command line:
|
||||
```bash
|
||||
$ cmake .. -DCUTLASS_NVCC_ARCHS='75;80' -DCUTLASS_LIBRARY_KERNELS=cutlass_simt_sgemm_128x128_8x2_nn_align1
|
||||
...
|
||||
$ make cutlass_profiler -j16
|
||||
```
|
||||
|
||||
Example command line for profiling single SGEMM CUDA kernel is as follows:
|
||||
```bash
|
||||
$ ./tools/profiler/cutlass_profiler --kernels=sgemm --m=3456 --n=4096 --k=4096
|
||||
|
||||
=============================
|
||||
Problem ID: 1
|
||||
|
||||
Provider: CUTLASS
|
||||
OperationKind: gemm
|
||||
Operation: cutlass_simt_sgemm_128x128_8x2_nn_align1
|
||||
|
||||
Status: Success
|
||||
Verification: ON
|
||||
Disposition: Passed
|
||||
|
||||
cuBLAS: Passed
|
||||
|
||||
Arguments: --m=3456 --n=4096 --k=4096 --A=f32:column --B=f32:column --C=f32:column --alpha=1 --beta=0 --split_k_slices=1 \
|
||||
--batch_count=1 --op_class=simt --accum=f32 --cta_m=128 --cta_n=128 --cta_k=8 --stages=2 --warps_m=4 \
|
||||
--warps_n=2 --warps_k=1 --inst_m=1 --inst_n=1 --inst_k=1 --min_cc=50 --max_cc=1024
|
||||
|
||||
Bytes: 180355072 bytes
|
||||
FLOPs: 115992428544 flops
|
||||
|
||||
Runtime: 6.73655 ms
|
||||
Memory: 24.934 GiB/s
|
||||
|
||||
Math: 17218.4 GFLOP/s
|
||||
|
||||
=============================
|
||||
```
|
||||
|
||||
### Building a subset of Tensor Core Convolution kernels
|
||||
|
||||
To compile a subset of Tensor core convolution kernels implementing forward propagation (fprop) with FP32 accumulation
|
||||
and FP16 input targeting NVIDIA Ampere and Turing architecture, use the below cmake command line:
|
||||
```bash
|
||||
$ cmake .. -DCUTLASS_NVCC_ARCHS='75;80' -DCUTLASS_LIBRARY_KERNELS=cutlass_tensorop_s*fprop_optimized_f16
|
||||
...
|
||||
$ make cutlass_profiler -j16
|
||||
```
|
||||
|
||||
Example command line for profiling a subset of Tensor Core convolution kernels is as follows:
|
||||
|
||||
```bash
|
||||
$ ./tools/profiler/cutlass_profiler --kernels=cutlass_tensorop_s*fprop_optimized_f16 --n=8 --h=224 --w=224 --c=128 --k=128 --r=3 --s=3
|
||||
|
||||
...
|
||||
=============================
|
||||
Problem ID: 1
|
||||
|
||||
Provider: CUTLASS
|
||||
OperationKind: conv2d
|
||||
Operation: cutlass_tensorop_s16816fprop_optimized_f16_128x128_32x5_nhwc
|
||||
|
||||
Status: Success
|
||||
Verification: ON
|
||||
Disposition: Passed
|
||||
|
||||
reference_device: Passed
|
||||
|
||||
Arguments: --conv_kind=fprop --n=8 --h=224 --w=224 --c=128 --k=128 --r=3 --s=3 --p=224 --q=224 --pad_h=1 --pad_w=1 \
|
||||
--stride_h=1 --stride_w=1 --dilation_h=1 --dilation_w=1 --Activation=f16:nhwc --Filter=f16:nhwc --Output=f32:nhwc \
|
||||
--conv_mode=cross --iterator_algorithm=optimized --alpha=1 --beta=0 --split_k_mode=serial --split_k_slices=1 \
|
||||
--eq_gemm_provider=none --op_class=tensorop --accum=f32 --cta_m=128 --cta_n=128 --cta_k=32 --stages=5 \
|
||||
--warps_m=2 --warps_n=2 --warps_k=1 --inst_m=16 --inst_n=8 --inst_k=16 --min_cc=80 --max_cc=1024
|
||||
|
||||
Bytes: 1130659840 bytes
|
||||
FLOPs: 118482796544 flops
|
||||
|
||||
Runtime: 0.711496 ms
|
||||
Memory: 1479.99 GiB/s
|
||||
|
||||
Math: 166526 GFLOP/s
|
||||
|
||||
=============================
|
||||
...
|
||||
```
|
||||
|
||||
|
||||
### Building one Convolution CUDA kernel
|
||||
|
||||
To compile and run one CUDA Core convolution kernel implementing forward propagation (fprop) with F32 accumulation
|
||||
and FP32 input targeting NVIDIA Ampere and Turing architecture, use the below cmake command line:
|
||||
```bash
|
||||
$ cmake .. -DCUTLASS_NVCC_ARCHS='75;80' -DCUTLASS_LIBRARY_KERNELS=cutlass_simt_sfprop_optimized_128x128_8x2_nhwc
|
||||
...
|
||||
$ make cutlass_profiler -j16
|
||||
```
|
||||
|
||||
Example command line for profiling one CUDA Core convolution kernel:
|
||||
|
||||
```bash
|
||||
$ ./tools/profiler/cutlass_profiler --kernels=cutlass_simt_sfprop_optimized_128x128_8x2_nhwc --n=8 --h=224 --w=224 --c=128 --k=128 --r=3 --s=3
|
||||
|
||||
|
||||
=============================
|
||||
Problem ID: 1
|
||||
|
||||
Provider: CUTLASS
|
||||
OperationKind: conv2d
|
||||
Operation: cutlass_simt_sfprop_optimized_128x128_8x2_nhwc
|
||||
|
||||
Status: Success
|
||||
Verification: ON
|
||||
Disposition: Passed
|
||||
|
||||
reference_device: Passed
|
||||
|
||||
Arguments: --conv_kind=fprop --n=8 --h=224 --w=224 --c=128 --k=128 --r=3 --s=3 --p=224 --q=224 --pad_h=1 --pad_w=1 \
|
||||
--stride_h=1 --stride_w=1 --dilation_h=1 --dilation_w=1 --Activation=f32:nhwc --Filter=f32:nhwc --Output=f32:nhwc \
|
||||
--conv_mode=cross --iterator_algorithm=optimized --alpha=1 --beta=0 --split_k_mode=serial --split_k_slices=1 \
|
||||
--eq_gemm_provider=none --op_class=simt --accum=f32 --cta_m=128 --cta_n=128 --cta_k=8 --stages=2 --warps_m=4 \
|
||||
--warps_n=2 --warps_k=1 --inst_m=1 --inst_n=1 --inst_k=1 --min_cc=50 --max_cc=1024
|
||||
|
||||
Bytes: 2055798784 bytes
|
||||
FLOPs: 118482796544 flops
|
||||
|
||||
Runtime: 7.34266 ms
|
||||
Memory: 260.752 GiB/s
|
||||
|
||||
Math: 16136.2 GFLOP/s
|
||||
|
||||
|
||||
=============================
|
||||
|
||||
```
|
||||
|
||||
## More Details on Compiling CUTLASS Kernels and CUTLASS Profiler
|
||||
- Please follow the links for more CMake examples on selectively compiling CUTLASS kernels:
|
||||
- [GEMM CMake Examples](quickstart.md#gemm-cmake-examples)
|
||||
- [Implicit GEMM convolution CMake Examples](quickstart.md#convolution-cmake-examples)
|
||||
- [Further details about the CUTLASS Profiler are described here.](profiler.md)
|
||||
|
||||
|
||||
# About
|
||||
|
||||
CUTLASS is released by NVIDIA Corporation as Open Source software under the
|
||||
[3-clause "New" BSD license](LICENSE.txt).
|
||||
|
||||
# Contributors
|
||||
|
||||
The official list of CUTLASS developers and contributors is available here: [CONTRIBUTORS](CONTRIBUTORS.md).
|
||||
|
||||
# Copyright
|
||||
|
||||
Copyright (c) 2017 - 2025 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
```
|
||||
Redistribution and use in source and binary forms, with or without
|
||||
modification, are permitted provided that the following conditions are met:
|
||||
|
||||
1. Redistributions of source code must retain the above copyright notice, this
|
||||
list of conditions and the following disclaimer.
|
||||
|
||||
2. Redistributions in binary form must reproduce the above copyright notice,
|
||||
this list of conditions and the following disclaimer in the documentation
|
||||
and/or other materials provided with the distribution.
|
||||
|
||||
3. Neither the name of the copyright holder nor the names of its
|
||||
contributors may be used to endorse or promote products derived from
|
||||
this software without specific prior written permission.
|
||||
|
||||
THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS "AS IS"
|
||||
AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE
|
||||
IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE ARE
|
||||
DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT HOLDER OR CONTRIBUTORS BE LIABLE
|
||||
FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL
|
||||
DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR
|
||||
SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER
|
||||
CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY,
|
||||
OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE
|
||||
OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
|
||||
```
|
||||
210
media/docs/cpp/pipeline.md
Normal file
210
media/docs/cpp/pipeline.md
Normal file
@@ -0,0 +1,210 @@
|
||||
# Synchronization primitives
|
||||
|
||||
## Overview of CUDA's synchronization methods
|
||||
|
||||
The CUDA programming model provides 3 abstractions:
|
||||
|
||||
* hierarchical parallelism -- that is, parallel threads
|
||||
grouped into hierarchical units such as blocks and clusters;
|
||||
|
||||
* shared memory, through which parallel threads that are
|
||||
in the same hierarchical unit can communicate; and
|
||||
|
||||
* synchronization methods for threads.
|
||||
|
||||
These abstractions help developers extract
|
||||
both fine-grained and coarse-grained parallelism,
|
||||
by making it possible for them to subdivide problems
|
||||
into independent components,
|
||||
and to insert synchronization at appropriate points.
|
||||
|
||||
Over the years CUDA has introduced several synchronization primitives
|
||||
that operate at different levels of the hierarchy.
|
||||
These include
|
||||
|
||||
* [thread block - level](https://docs.nvidia.com/cuda/cuda-c-programming-guide/index.html#synchronization-functions) synchronization (e.g., `__syncthreads()`);
|
||||
|
||||
* [warp-level](https://developer.nvidia.com/blog/using-cuda-warp-level-primitives/) synchronization (e.g., `__syncwarp()`); and
|
||||
|
||||
* [thread-level](https://docs.nvidia.com/cuda/cuda-c-programming-guide/#memory-fence-functions) fence operations.
|
||||
|
||||
As an extension to this, starting with the Hopper architecture, CUDA added the following improvements:
|
||||
|
||||
* [thread block clusters](https://docs.nvidia.com/cuda/cuda-c-programming-guide/index.html#thread-block-clusters) --
|
||||
a new level in the thread hierarchy representing
|
||||
a group of thread blocks that can coordinate and share data;
|
||||
|
||||
* synchronization instructions for a thread block cluster and threads within a cluster scope.
|
||||
|
||||
## CUTLASS's abstractions for Hopper features
|
||||
|
||||
CUTLASS now includes abstractions
|
||||
for the following features introduced in Hopper.
|
||||
|
||||
1. Thread block cluster - level synchronization and query
|
||||
[APIs](https://github.com/NVIDIA/cutlass/tree/main/include/cute/arch/cluster_sm90.hpp)
|
||||
|
||||
2. Abstractions for new
|
||||
[barrier instructions](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/arch/barrier.h)
|
||||
which help with efficient synchronization
|
||||
of threads within a thread block cluster.
|
||||
|
||||
### Asynchronous pipelines
|
||||
|
||||
In order to write a performant GEMM Kernel,
|
||||
software pipelining is critical to hide the latency of global memory loads.
|
||||
(Please refer to the
|
||||
[Efficient GEMM](efficient_gemm.md#pipelining) document.)
|
||||
Different threads or groups of threads
|
||||
may have different roles in the pipeline.
|
||||
Some are "producers" that load data or perform computations
|
||||
to satisfy other threads' input data dependencies.
|
||||
The same or different threads may be "consumers"
|
||||
that do other work with those input data dependencies,
|
||||
once they are satisfied.
|
||||
Starting with the Hopper architecture,
|
||||
the presence of hardware-accelerated synchronization instructions
|
||||
make it possible for "producer" and "consumer" threads
|
||||
to communicate with each other efficiently
|
||||
about their data dependencies.
|
||||
|
||||
Implementing a persistent GEMM algorithm calls for managing
|
||||
dozens of different kinds of asynchronously executing operations
|
||||
that synchronize using multiple barriers organized as a circular list.
|
||||
This complexity is too much for human programmers to manage by hand.
|
||||
As a result, we have developed
|
||||
[asynchronous Pipeline classes](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/pipeline/).
|
||||
These classes help developers orchestrate a pipeline
|
||||
of asynchronous producer and consumer threads,
|
||||
without needing to worry about lower-level hardware details.
|
||||
These classes serve a similar function as the various
|
||||
[pipeline abstractions](https://nvidia.github.io/libcudacxx/extended_api/synchronization_primitives/pipeline.html)
|
||||
in libcu++.
|
||||
|
||||
#### Pipeline methods
|
||||
|
||||
##### Producer acquire
|
||||
|
||||
The `producer_acquire` method is to be used by asynchronous producer threads
|
||||
before issuing other instructions associated with a particular pipeline stage
|
||||
(e.g., copy or write).
|
||||
|
||||
This is a blocking instruction
|
||||
which blocks further execution of consumer threads
|
||||
unless the particular stage waiting to be acquired
|
||||
is released by a consumer.
|
||||
|
||||
We say that a pipeline at its start is "empty" if producer threads are free to produce and do not need to wait for a consumer release -- that is, if an acquire operation is expected to succeed. If the pipeline at its start is empty, then we can either skip performing producer acquire operations during the first pass through the pipeline stages, or use the `make_producer_start_state` method. The latter ensures that the acquire operation will succeed at the start of a pipeline.
|
||||
|
||||
##### Producer commit
|
||||
|
||||
The `producer_commit` method is to be issued by asynchronous producer threads
|
||||
after the instructions associated with a particular stage
|
||||
(e.g., shared memory writes) have completed,
|
||||
in order to notify the waiting asynchronous consumer threads.
|
||||
This is a nonblocking instruction.
|
||||
|
||||
This API may result in a No-Op in some cases,
|
||||
if the producer instructions also update the barrier stage associated automatically
|
||||
(e.g., TMA_based producer threads using the `PipelineTmaAsync ` class).
|
||||
|
||||
##### Consumer wait
|
||||
|
||||
The `consumer_wait` method is to be used by consumer threads
|
||||
before consuming data from a particular pipeline stage
|
||||
which is expected to be produced by producer threads.
|
||||
|
||||
This is a blocking instruction. That is,
|
||||
until the producer threads have committed to a particular stage,
|
||||
this instruction is expected to block further execution of consumer threads.
|
||||
|
||||
##### Consumer release
|
||||
|
||||
The `consumer_release` method is to be used by consumer threads
|
||||
to signal waiting producer threads that they have finished consuming data
|
||||
associated with a particular stage of the pipeline.
|
||||
This is a nonblocking instruction.
|
||||
|
||||
#### Pipeline example
|
||||
|
||||
```c++
|
||||
// 4-stage Pipeline
|
||||
static constexpr int NumStages = 4;
|
||||
using MainloopPipeline = typename cutlass::PipelineAsync<NumStages>;
|
||||
using PipelineState = typename cutlass::PipelineState<NumStages>;
|
||||
|
||||
// 2 producer threads and 1 consumer thread
|
||||
typename MainloopPipeline::Params params;
|
||||
params.producer_arv_count = 2;
|
||||
params.consumer_arv_count = 1;
|
||||
MainloopPipeline pipeline(shared_storage.storage, params);
|
||||
|
||||
// Producer threads
|
||||
if (thread_idx == 0 or thread_idx == 1) {
|
||||
PipelineState smem_pipe_write = cutlass::make_producer_start_state<MainloopPipeline>();
|
||||
for ( ; iter > 0; --iter) {
|
||||
pipeline.producer_acquire(smem_pipe_write);
|
||||
|
||||
// Producer ops
|
||||
// If any memory operations are involved, then we also need
|
||||
// to guarantee that writes are completed and visible to consumer(s).
|
||||
|
||||
pipeline.producer_commit(smem_pipe_write);
|
||||
++smem_pipe_write;
|
||||
}
|
||||
}
|
||||
else if (thread_idx == 2) {
|
||||
PipelineState smem_pipe_read;
|
||||
for (; iter > 0; --iter) {
|
||||
pipeline.consumer_wait(smem_pipe_read);
|
||||
|
||||
// Consumer ops
|
||||
|
||||
pipeline.consumer_release(smem_pipe_read);
|
||||
++smem_pipe_read;
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
In this example, we create an instance of the asynchronous pipeline class `PipelineSync`,
|
||||
and then synchronize among 3 asynchronously executing threads:
|
||||
2 producer threads and 1 consumer thread.
|
||||
|
||||
Please note that this is a basic example.
|
||||
There are different versions possible,
|
||||
depending on what the producer and consumer threads are doing.
|
||||
Please refer to our [unit tests](https://github.com/NVIDIA/cutlass/tree/main/test/unit/pipeline)
|
||||
and the other [pipeline classes](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/pipeline/pipeline.hpp)
|
||||
for more details.
|
||||
|
||||
# Copyright
|
||||
|
||||
Copyright (c) 2023 - 2025 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
```
|
||||
Redistribution and use in source and binary forms, with or without
|
||||
modification, are permitted provided that the following conditions are met:
|
||||
|
||||
1. Redistributions of source code must retain the above copyright notice, this
|
||||
list of conditions and the following disclaimer.
|
||||
|
||||
2. Redistributions in binary form must reproduce the above copyright notice,
|
||||
this list of conditions and the following disclaimer in the documentation
|
||||
and/or other materials provided with the distribution.
|
||||
|
||||
3. Neither the name of the copyright holder nor the names of its
|
||||
contributors may be used to endorse or promote products derived from
|
||||
this software without specific prior written permission.
|
||||
|
||||
THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS "AS IS"
|
||||
AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE
|
||||
IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE ARE
|
||||
DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT HOLDER OR CONTRIBUTORS BE LIABLE
|
||||
FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL
|
||||
DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR
|
||||
SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER
|
||||
CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY,
|
||||
OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE
|
||||
OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
|
||||
```
|
||||
756
media/docs/cpp/profiler.md
Normal file
756
media/docs/cpp/profiler.md
Normal file
@@ -0,0 +1,756 @@
|
||||

|
||||
|
||||
# CUTLASS Profiler
|
||||
|
||||
The CUTLASS Profiler is a command-line driven test and profiling environment for CUTLASS computations
|
||||
defined in the CUTLASS Instance Library. The CUTLASS Profiler is capable of executing each GEMM, Sparse Gemm,
|
||||
Conv2d, and Conv3d kernel.
|
||||
|
||||
The CUTLASS Profiler may be compiled with:
|
||||
```bash
|
||||
$ make cutlass_profiler -j
|
||||
```
|
||||
|
||||
To limit compilation time, only one tile size (typically 128x128) and threadblock cluster size (typically 2x1x1) is instantiated for each data type,
|
||||
math instruction, and layout. To instantiate all sizes, set the following environment variable when running CMake from an
|
||||
empty `build/` directory.
|
||||
```bash
|
||||
$ cmake .. -DCUTLASS_NVCC_ARCHS="70;75;80" -DCUTLASS_LIBRARY_KERNELS=all -DCUTLASS_UNITY_BUILD_ENABLED=ON
|
||||
...
|
||||
$ make cutlass_profiler -j
|
||||
```
|
||||
Enabling the unity build places multiple kernel instances in one compilation unit, thereby reducing size of the compiled
|
||||
binary and avoiding linker limitations on some platforms.
|
||||
|
||||
The CUTLASS Profiler sources are stored in:
|
||||
|
||||
```bash
|
||||
tools/
|
||||
profiler/
|
||||
```
|
||||
|
||||
# Emitting kernels via `emit_kernel_listing.py`
|
||||
|
||||
We provide a Python script `emit_kernel_listing.py` that allows a user to selectively test a subset of profiler-based kernels stamped out in `generator.py`. A unique benefit to generate kernels and test via this script is that it can feed a series of runtime arguments, such as different `M`/`N`/`K` and `alpha`/`beta`, to each kernel, instead of relying on a single default value. It also properly generates runtime datatype and cluster shapes for certain kernels to help reduce the generated kernel count and accordingly the total compilation time. An interested user may refer to [emit_kernel_listing.py](https://github.com/NVIDIA/cutlass/tree/main/python/cutlass_library/emit_kernel_listing.py) for details. To enable this new feature, a user should add `-DCUTLASS_BUILD_FOR_PROFILER_REGRESSIONS=ON` when building CUTLASS profiler.
|
||||
|
||||
## Instantiating more kernels with Hopper
|
||||
With Hopper (SM90), you will need to use an additional flag,
|
||||
`CUTLASS_LIBRARY_INSTANTIATION_LEVEL`, in order to instantiate all possible combinations,
|
||||
which unlike previous architectures, will be in the order of millions of kernels.
|
||||
Due to this, `CUTLASS_LIBRARY_KERNELS` must be non-empty, since generating and filtering these
|
||||
kernels alone can take hours.
|
||||
You must also exercise caution, because not all of these configs are tested, and some may fail to
|
||||
compile or fail to launch at runtime.
|
||||
|
||||
```bash
|
||||
$ cmake .. \
|
||||
-DCUTLASS_NVCC_ARCHS="90a" \
|
||||
-DCUTLASS_LIBRARY_KERNELS="cutlass3x_sm90_tensorop_s64x64x16gemm_f16_f16_f32_void_f32_*" \
|
||||
-DCUTLASS_LIBRARY_INSTANTIATION_LEVEL="max" \
|
||||
-DCUTLASS_UNITY_BUILD_ENABLED=ON
|
||||
```
|
||||
|
||||
The CUTLASS profiler employs a four-digit integer level (global instantiation level) mechanism to manage the generation of kernel configurations. This global instantiation level decides the behavior of multiple "generators" by defining how many and which combinations of configurations are produced. If a global instantiation level contains fewer than four digits, it can be padded with leading zeros to ensure it is four digits long. Each of the four digits in the global level corresponds to a specific category that influences kernel generation, from right to left:
|
||||
|
||||
0. **Instruction Shape**
|
||||
1. **MMA Shape Multiplier**
|
||||
2. **Cluster Shape**
|
||||
3. **Schedule Pruning**
|
||||
|
||||
Cluster shape levels define the number of CTAs (Cooperative Thread Arrays) included in the kernel generation:
|
||||
|
||||
- **Level 0**: Only `(1, 2, 1)` cluster shape.
|
||||
- **Level 1**: Clusters with 2 CTAs.
|
||||
- **Level 2**: Clusters with 1 or 2 CTAs.
|
||||
- **Level 3**: Clusters with 1, 2, or 4 CTAs.
|
||||
- **Level 4**: Clusters with 1, 2, 4, or 8 CTAs.
|
||||
- **Level 5**: Clusters with 1, 2, 4, 8, or 16 CTAs.
|
||||
|
||||
The MMA multipliers are combined with MMA instruction shapes (WGMMA shapes) to form CTA shapes. The levels for MMA multipliers determine the configurations generated for different data types.
|
||||
- **Levels [0, 3]**: Control the specific configurations generated for various data types.
|
||||
- **Level 9**: Activates exhaustive mode, generating all possible configurations.
|
||||
|
||||
Higher levels encompass a broader range of CTA configurations, resulting in more comprehensive kernel generation.
|
||||
|
||||
Instruction shape levels control the selection of WGMMA shapes used in kernel generation:
|
||||
|
||||
- **Level 0**: Generates the "default" shape only.
|
||||
- **Level 1**: Includes additional shapes for unpruned cases, specifically for TF32 data type.
|
||||
- **Level 2**: Includes shapes that are powers of 2.
|
||||
- **Level 3**: Includes all other shapes.
|
||||
|
||||
The detailed defination of the three instantiation levels controlling cluster shape, MMA shape multiplier, and instruction shape can be found in [sm90_shapes.py](https://github.com/NVIDIA/cutlass/tree/main/python/cutlass_library/sm90_shapes.py).
|
||||
|
||||
Schedule pruning levels decide the epilogue schedule and mainloop schedule to stamp out a kernel instance. As defined in `get_valid_schedules` in [sm90_utils.py](https://github.com/NVIDIA/cutlass/tree/main/python/cutlass_library/sm90_utils.py),
|
||||
|
||||
- **Level >= 1**: Indicates that no pruning is being applied.
|
||||
- **Level 0**: Indicates pruning according to existing [generator.py](https://github.com/NVIDIA/cutlass/tree/main/python/cutlass_library/generator.py) behavior.
|
||||
|
||||
An instantiation level `500`, which is padded to `0500`, thus indicates:
|
||||
|
||||
- **Instruction Shapes**: At level 0, generating only the "default" shape.
|
||||
- **MMA Multipliers**: At level 0, generating only one multiplier, `(2, 1, 4)`.
|
||||
- **Cluster Sizes**: At level 5, allowing for clusters with 1, 2, 4, 8, or 16 CTAs.
|
||||
- **Schedule Pruning**: At level 0, where pruning is applied according to the existing `generator.py` behavior.
|
||||
|
||||
## Mixed input data type kernels for Hopper
|
||||
|
||||
With Hopper (SM90), the kernel generator will generate the following combinations of mixed input data types ("mixed dtype"):
|
||||
|
||||
| dtype(A) | dtype(B) |
|
||||
| -------- | ---------- |
|
||||
| e4m3 | f16, bf16 |
|
||||
| e5m2 | f16, bf16 |
|
||||
| int8 | f16, bf16 |
|
||||
| uint8 | f16, bf16 |
|
||||
| int4 | f16, bf16 |
|
||||
| int4 | e4m3, e5m2 |
|
||||
| uint4 | f16, bf16 |
|
||||
| int2 | f16, bf16 |
|
||||
| uint2 | f16, bf16 |
|
||||
|
||||
For each mixed dtype kernel, the kernel generator will generate combinations of three different running modes:
|
||||
* Convert-only
|
||||
* Scale-only
|
||||
* Scale-with-zero-point-shifting
|
||||
|
||||
For {4-bits-dtype, 8-bits-dtype} x 16-bits-dtype, the kernel generator will further generate kernels using shuffled layouts for the narrow data type matrix, which may have a better performance compared to its non-shuffle counter parts.
|
||||
|
||||
## CUTLASS Profiler usage
|
||||
|
||||
The CUTLASS Profiler usage statement may be obtained by executing `cutlass_profiler --help` and appears as follows.
|
||||
```bash
|
||||
CUTLASS Performance Tool
|
||||
usage:
|
||||
|
||||
cutlass_profiler [options]
|
||||
|
||||
--help
|
||||
|
||||
--mode=<string> Cutlass profiler execution mode.
|
||||
--mode=profile regular verification and profiling (default)
|
||||
--mode=dry_run no kernels are launched or workspaces allocated
|
||||
--mode=enumerate lists all operation kind and operations
|
||||
--mode=trace executes a single device-side computation with
|
||||
no other kernel launches
|
||||
|
||||
--device-info Prints information on all GPUs present in the system
|
||||
|
||||
--operation=<operation_kind> CUTLASS operation to profile.
|
||||
|
||||
--kernels=<string_list> Filter operations by kernel names. For example, call all kernels with
|
||||
("s1688" and "nt") or ("s844" and "tn" and "align8") in their
|
||||
operation name using --kernels="s1688*nt, s884*tn*align8"
|
||||
|
||||
--kernels-file=<path> Same behavior as `kernels`, but kernel names are specified in a file with
|
||||
one kernel name on each line. Set of profiled kernels is the union of kernels
|
||||
specified here and those specified in `kernels`.
|
||||
|
||||
--ignore-kernels=<string_list> Excludes kernels whose names match anything in this list.
|
||||
|
||||
Device:
|
||||
--device=<int> CUDA Device ID
|
||||
|
||||
--compute-capability=<int> Override the compute capability.
|
||||
|
||||
--llc-capacity=<capacity in KiB> Capacity of last-level cache in kilobytes. If this is non-zero,
|
||||
profiling phases cycle through different input tensors to induce
|
||||
capacity misses in the L2.
|
||||
|
||||
--allocations=<name>:<device>,<name>:<device> Pairs of allocation names to devices. If <device> is negative,
|
||||
the execution device is used
|
||||
|
||||
|
||||
Initialization:
|
||||
--initialization=<bool> Enables initialization (default: true). If false, device memory is
|
||||
not initialized after allocation.
|
||||
|
||||
--initialization-provider=<provider> Selects initialization provider {host, device*}. (default: '*')
|
||||
|
||||
--dist=<distribution> Data distribution of input tensors {uniform*, gaussian, identity, sequential}
|
||||
--dist=uniform,min:<double>,max:<double>,scale:<integer>
|
||||
--dist=gaussian,mean:<double>,stddev:<double>,scale:<integer>
|
||||
--dist=sequential,start:<double>,delta:<double>,scale:<integer>
|
||||
--dist=identity
|
||||
|
||||
--seed=<int> Random number generator seed. Used to enforce deterministic
|
||||
initialization.
|
||||
|
||||
|
||||
Library:
|
||||
--library-algo-mode=<mode> Indicates algorithm mode used to call libraries such as cuBLAS and cuDNN.
|
||||
mode={default*,matching,best}
|
||||
|
||||
--library-algos=<range-list> If --algorithm-mode=best, permits specifying a selection of algorithms.
|
||||
|
||||
|
||||
Profiling:
|
||||
--workspace-count=<workspace count> Number of discrete workspaces maintained to avoid cache-resident
|
||||
If zero (default), the amount is chosen for each workload based on
|
||||
capacity of the last-level cache.
|
||||
|
||||
--profiling-iterations=<iterations> Number of iterations to profile each kernel. If zero, kernels
|
||||
are launched up to the profiling duration. If non-zero, this
|
||||
overrides `profiling-duration` and `min-iterations`.
|
||||
|
||||
--profiling-duration=<duration> Time to spend profiling each kernel (ms). Overriden by
|
||||
`profiling-iterations` when `profiling-iterations` != 0.
|
||||
Note that `min-iterations` must also be satisfied.
|
||||
|
||||
--min-iterations=<iterations> Minimum number of iterations to spend profiling each kernel, even if
|
||||
`profiling-duration` has been met.
|
||||
|
||||
--warmup-iterations=<iterations> Number of iterations to execute each kernel prior to profiling (default: 10).
|
||||
|
||||
--use-cuda-graphs=<bool> If true, kernels are launched in a CUDA graph. Useful when the kernel launch time is a bottleneck.
|
||||
|
||||
--sleep-duration=<duration> Number of ms to sleep between profiling periods (ms).
|
||||
|
||||
--profiling-enabled=<bool> If true, profiling is actually conducted.
|
||||
|
||||
Verification:
|
||||
--verification-enabled=<bool> Whether to perform verification checks.
|
||||
|
||||
--epsilon=<error> Error threshold. Setting to zero (default) requires
|
||||
bit-level equivalence.
|
||||
|
||||
--nonzero-floor=<floor> Results whose absolute value is less than this quantity
|
||||
are treated as zero for comparisons.
|
||||
|
||||
--save-workspace=<string> Specifies when to save the GEMM inputs and results to the filesystem.
|
||||
--save-workspace=never never save workspace (default)
|
||||
--save-workspace=incorrect save workspace for incorrect results
|
||||
--save-workspace=always always save workspace
|
||||
|
||||
--verification-providers=<providers> List of providers used to verify result. (default: '*')
|
||||
Gemm verification-providers {cublas*}
|
||||
Conv2d verification-providers {cudnn*, device*, host}
|
||||
|
||||
|
||||
Report:
|
||||
--append=<bool> If true, result is appended to possibly existing file. Otherwise,
|
||||
any existing file is overwritten.
|
||||
|
||||
--output=<path> Path to output file for machine readable results. Operation kind and '.csv' is appended.
|
||||
|
||||
--junit-output=<path> Path to junit output file for result reporting. Operation kind and '.junit.xml' is appended.
|
||||
|
||||
--report-not-run=<bool> If true, reports the status of all kernels including those that
|
||||
do not satisfy the given arguments.
|
||||
|
||||
--tags=<column:tag,...> Inserts leading columns in output table and uniform values for each
|
||||
column. Useful for generating pivot tables.
|
||||
|
||||
--verbose=<bool> Prints human-readable text to stdout. If false, nothing is written to stdout.
|
||||
|
||||
|
||||
About:
|
||||
--version CUTLASS 2.4.0 built on Nov 19 2020 at 11:59:00
|
||||
|
||||
|
||||
Operations:
|
||||
|
||||
gemm General matrix-matrix product. D = alpha * A*B + beta * C
|
||||
spgemm Structured sparse GEMM. D = alpha * A*B + beta * C
|
||||
conv2d Conv2d operation. Output(Tensor4D) = alpha * Input(Tensor4D) * Filter(Tensor4D) + beta * Input(Tensor4D)
|
||||
conv3d Conv3d operation. Output(Tensor5D) = alpha * Input(Tensor5D) * Filter(Tensor5D) + beta * Input(Tensor5D)
|
||||
|
||||
|
||||
For details about a particular function, specify the function name with --help.
|
||||
|
||||
Example:
|
||||
|
||||
$ cutlass_profiler --operation=Gemm --help
|
||||
|
||||
$ cutlass_profiler --operation=Conv3d --help
|
||||
|
||||
$ cutlass_profiler --operation=Conv2d --help
|
||||
|
||||
```
|
||||
|
||||
# GEMM
|
||||
|
||||
The CUTLASS Profiler is capable of executing GEMM and Sparse GEMM problems.
|
||||
|
||||
The CUTLASS Profiler can be built with cuBLAS enabled to use as a reference implementation. If CMake detects
|
||||
the cuBLAS library available in the system, it is included as a dependency. This may be explicitly overridden
|
||||
with CMake flag `CUTLASS_ENABLE_CUBLAS`.
|
||||
|
||||
## GEMM Arguments
|
||||
|
||||
The complete set of arguments available to each operation may be viewed by specifying the operation name
|
||||
in addition to `--help`. The argument flags and their aliases usable for GEMM appear as follows.
|
||||
|
||||
```bash
|
||||
$ ./tools/profiler/cutlass_profiler --operation=gemm --help
|
||||
|
||||
GEMM
|
||||
|
||||
[enum] --gemm_kind Variant of GEMM (e.g. universal, gemm, planar_complex, planar_complex_array)
|
||||
[int] --m,--problem-size::m M dimension of the GEMM problem space
|
||||
[int] --n,--problem-size::n N dimension of the GEMM problem space
|
||||
[int] --k,--problem-size::k K dimension of the GEMM problem space
|
||||
[tensor] --A Tensor storing the A operand
|
||||
[tensor] --B Tensor storing the B operand
|
||||
[tensor] --C Tensor storing the C operand
|
||||
[scalar] --alpha,--epilogue::alpha Epilogue scalar alpha
|
||||
[scalar] --beta,--epilogue::beta Epilogue scalar beta
|
||||
[enum] --split_k_mode,--split-k-mode Variant of split K mode(serial, parallel)
|
||||
[int] --split_k_slices,--split-k-slices Number of partitions of K dimension
|
||||
[int] --batch_count,--batch-count Number of GEMMs computed in one batch
|
||||
[enum] --op_class,--opcode-class Class of math instruction (simt, tensorop, wmmatensorop, wmma).
|
||||
[enum] --accum,--accumulator-type Math instruction accumulator data type
|
||||
[int] --cta_m,--threadblock-shape::m Threadblock shape in the M dimension
|
||||
[int] --cta_n,--threadblock-shape::n Threadblock shape in the N dimension
|
||||
[int] --cta_k,--threadblock-shape::k Threadblock shape in the K dimension
|
||||
[int] --cluster_m,--cluster-shape::m Cluster shape in the M dimension
|
||||
[int] --cluster_n,--cluster-shape::n Cluster shape in the N dimension
|
||||
[int] --cluster_k,--cluster-shape::k Cluster shape in the K dimension
|
||||
[int] --cluster_m_fallback,--cluster-shape-fallback::m Fallback cluster shape in the M dimension
|
||||
[int] --cluster_n_fallback,--cluster-shape-fallback::n Fallback cluster shape in the N dimension
|
||||
[int] --cluster_k_fallback,--cluster-shape-fallback::k Fallback cluster shape in the K dimension
|
||||
[int] --stages,--threadblock-stages Number of stages of threadblock-scoped matrix multiply
|
||||
[int] --warps_m,--warp-count::m Number of warps within threadblock along the M dimension
|
||||
[int] --warps_n,--warp-count::n Number of warps within threadblock along the N dimension
|
||||
[int] --warps_k,--warp-count::k Number of warps within threadblock along the K dimension
|
||||
[int] --inst_m,--instruction-shape::m Math instruction shape in the M dimension
|
||||
[int] --inst_n,--instruction-shape::n Math instruction shape in the N dimension
|
||||
[int] --inst_k,--instruction-shape::k Math instruction shape in the K dimension
|
||||
[int] --min_cc,--minimum-compute-capability Minimum device compute capability
|
||||
[int] --max_cc,--maximum-compute-capability Maximum device compute capability
|
||||
[enum] --raster_order={heuristic|H|along_m|M|along_n|N} If supported by kernel, sets the tile raster direction
|
||||
[int] --swizzle_size={1,2,4,8} If supported by kernel, sets the 2D tile swizzle extent (In Hopper, other values will be rounded down to the nearest supported value)
|
||||
[int] --use_pdl,--use-pdl Use PDL (true, false)
|
||||
[int] --enable_sm90_mixed_dtype_shuffle_test If true, the profiler will test SM90 mixed input kernels that can use shuffled input layouts for better performance
|
||||
[enum] --runtime_input_datatype_a Runtime data type for A matrix, narrow-precision only (e4m3, e5m2, e3m2, e2m3, e2m1)
|
||||
[enum] --runtime_input_datatype_b Runtime data type for B matrix, narrow-precision only (e4m3, e5m2, e3m2, e2m3, e2m1)
|
||||
|
||||
Examples:
|
||||
|
||||
Profile a particular problem size:
|
||||
$ cutlass_profiler --operation=Gemm --m=1024 --n=1024 --k=128
|
||||
|
||||
Schmoo over problem size and beta:
|
||||
$ cutlass_profiler --operation=Gemm --m=1024:4096:256 --n=1024:4096:256 --k=128:8192:128 --beta=0,1,2.5
|
||||
|
||||
Schmoo over accumulator types:
|
||||
$ cutlass_profiler --operation=Gemm --accumulator-type=f16,f32
|
||||
|
||||
Run when A is f16 with column-major and B is any datatype with row-major (For column major, use column, col, or n. For row major use, row or t):
|
||||
$ cutlass_profiler --operation=Gemm --A=f16:column --B=*:row
|
||||
|
||||
Using various input value distribution:
|
||||
$ cutlass_profiler --operation=Gemm --dist=uniform,min:0,max:3
|
||||
$ cutlass_profiler --operation=Gemm --dist=gaussian,mean:0,stddev:3
|
||||
$ cutlass_profiler --operation=Gemm --dist=sequential,start:0,delta:1
|
||||
|
||||
Using CUTLASS 3.x GEMM kernel with a tile scheduler that supports runtime tile remapping and raster mode order:
|
||||
$ cutlass_profiler --operation=Gemm --m=2048 --n=2048 --k=2048 --raster_order=M --swizzle_size=2
|
||||
|
||||
Run a kernel with cta tile size of 256x128x32 and save workspace if results are incorrect (note that --cta-tile::k=32 is default cta-tile size):
|
||||
$ cutlass_profiler --operation=Gemm --cta_m=256 --cta_n=128 --cta_k=32 --save-workspace=incorrect
|
||||
|
||||
Test your changes to gemm kernels with a quick functional test and save results in functional-test.csv:
|
||||
$ cutlass_profiler --operation=Gemm \
|
||||
--m=8,56,120,136,256,264,512,520,1024,1032,4096,8192,16384 \
|
||||
--n=8,56,120,136,256,264,512,520,1024,1032,4096,8192,16384 \
|
||||
--k=8,16,32,64,128,256,288,384,504,512,520 \
|
||||
--beta=0,1,2 --profiling-iterations=1 \
|
||||
--providers=cutlass --output=functional-test.csv
|
||||
|
||||
Profile when execution is performed on device 0 and the C tensor is located on a device 1 and D on device 2:
|
||||
$ cutlass_profiler --device=0 --allocations=C:1,D:2 --operation=Gemm --m=1024 --n=1024 --k=128
|
||||
```
|
||||
|
||||
The format of tensor argument is followed by `<type>:<layout>`. The type could be `f32` as 32-bit floating point, `s8` as 8-bit signed integer, etc. The available types can be referred to the `NumericTypeID_enumerants` in [util.cu](https://github.com/NVIDIA/cutlass/tree/main/tools/library/src/util.cu). The layout could be `row` or `column`. If `--enable_sm90_mixed_dtype_shuffle_test=true` is used, the actual layout of the narrow data type matrix is a shuffled layout, neither `row` nor `column`.
|
||||
|
||||
In addition to encoded data types, CUTLASS profiler allows non-encoded generic data types, namely `f8`, `f6`, and `f4`, with corresponding encoding specified through GEMM input argument: `--runtime_input_datatype_a` and `--runtime_input_datatype_b`. Currently, six encoding schemes are supported: `e4m3`, `e5m2`, `e3m2`, `e2m3`, and `e2m1`.
|
||||
|
||||
Cluster shapes can be statically set to `Shape<int,int,_1>;` and specified via runtime arguments: `cluster_m`, `cluster_n` and `cluster_k` in CUTLASS profiler. In addition to preferred cluster shapes, a user can also specify fallback cluster shapes via runtime arguments: `cluster_m_fallback`, `cluster_n_fallback` and `cluster_k_fallback` in CUTLASS profiler. Those fallback cluster shapes are smaller shapes than the preferred ones for the hardware to assign when there is no chance to issue a larger preferred CGA cluster to the GPU. There are several rules for using a flexible CGA: 1) Preferred CGA size should be divisible by fallback CGA size. 2) Grid dim should be divisible by preferred CGA size. 3) Preferred CGA and fallback CGA must have the same depth (cluster_dim.z must be equal). One may refer to our CUTLASS Example [73_blackwell_gemm_flexible_cluster](https://github.com/NVIDIA/cutlass/tree/main/examples/73_blackwell_gemm_preferred_cluster/blackwell_gemm_preferred_cluster.cu) for more details of the this feature.
|
||||
Please be noted that this feature (flexible cluster shapes within a single grid) is only applicable to `sm100a` kernels. The hardware will rasterize into a single cluster shape for those kernels that do not support this feature even with preferred or fallback cluster shapes assigned.
|
||||
|
||||
CUTLASS 3.x kernels for Hopper and Blackwell also support a new feature called programatic dependent launch (PDL). This can be enabled with `--use-pdl`, and can overlap the epilogue of the prior kernel with the prologue of the next kernel. This can effectively hide kernel prologues. Using PDL can improve performance for back to back GEMMs. See [dependent kernel launch](dependent_kernel_launch.md) for more information. CUDA graphs can also be used (`--use-cuda-graphs`) with PDL to ensure that smaller kernels are enqueued back-to-back on a stream.
|
||||
|
||||
## Exhaustive search mode and top-k output ranking according to performance in GFLOPS/s
|
||||
|
||||
CUTLASS also allows a few options to enable searching best performing kernel in a broader parameter space.
|
||||
|
||||
1. **Sorting Performance Results by GFLOPs/second**
|
||||
A new option enables users to sort the final performance report based on GFLOPs/second, making it easier to identify the most efficient kernels.
|
||||
|
||||
2. **Exhaustive Search for Best Kernel Performance in GFLOPs/second**
|
||||
This feature allows the profiler to search for the best-performing kernel across a range of problem sizes, swizzle sizes, rasterization orders, and dynamic cluster configurations. It ensures that all viable configurations are considered to maximize performance.
|
||||
|
||||
3. **Performance Search Under a Fixed GEMM Shape**
|
||||
This option enables exhaustive performance tuning for a specific problem size. Unlike the previous feature, this restricts the search to a fixed GEMM shape while still exploring various kernel parameters to find the best configuration.
|
||||
|
||||
### Usage Examples
|
||||
|
||||
#### 1. Finding the Best Performing Kernel
|
||||
|
||||
Use the following command to conduct an exhaustive search and sort results by GFLOPs/second:
|
||||
|
||||
```bash
|
||||
cutlass_profiler --kernels=*gemm* --enable-kernel-performance-search --sort-results-flops-per-sec
|
||||
```
|
||||
|
||||
#### 2. Performance Optimization for a Fixed GEMM Shape
|
||||
|
||||
To optimize kernel performance for a specific GEMM problem size:
|
||||
|
||||
```bash
|
||||
cutlass_profiler --kernels=*gemm* --enable-best-kernel-for-fixed-shape --m=6144 --n=6144 --k=6144 --sort-results-flops-per-sec
|
||||
```
|
||||
|
||||
To search optimized kernel performance for a series of GEMM shapes (m, n, k = 1024, 2048):
|
||||
|
||||
```bash
|
||||
cutlass_profiler --kernels=*gemm* --enable-best-kernel-for-fixed-shape --m=1024,2048 --n=1024,2048 --k=1024,2048 --sort-results-flops-per-sec
|
||||
```
|
||||
|
||||
It is worth noting that by enabling exhaustive performance search via `--enable-kernel-performance-search`, a user is still able and responsible to decide parameters like data distribution in argument list, for which a user can choose `--dist=uniform,min:-1,max:1,scale:-1` to initialize a tensor with floating point numbers in uniform distribution. Otherwise, those parameters will be initialized to their default values.
|
||||
|
||||
For examples above, one can change the kernel filtering regex according to their own use cases.
|
||||
|
||||
## Example CUDA Core GEMM Operation
|
||||
|
||||
Example command line for profiling SGEMM kernels is as follows:
|
||||
```bash
|
||||
$ ./tools/profiler/cutlass_profiler --kernels=sgemm --m=3456 --n=4096 --k=4096
|
||||
|
||||
|
||||
|
||||
=============================
|
||||
Problem ID: 1
|
||||
|
||||
Provider: CUTLASS
|
||||
OperationKind: gemm
|
||||
Operation: cutlass_simt_sgemm_128x128_8x2_nn_align1
|
||||
|
||||
Status: Success
|
||||
Verification: ON
|
||||
Disposition: Passed
|
||||
|
||||
cuBLAS: Passed
|
||||
|
||||
Arguments: --m=3456 --n=4096 --k=4096 --A=f32:column --B=f32:column --C=f32:column --alpha=1 --beta=0 --split_k_slices=1 \
|
||||
--batch_count=1 --op_class=simt --accum=f32 --cta_m=128 --cta_n=128 --cta_k=8 --stages=2 --warps_m=4 \
|
||||
--warps_n=2 --warps_k=1 --inst_m=1 --inst_n=1 --inst_k=1 --min_cc=50 --max_cc=1024
|
||||
|
||||
Bytes: 180355072 bytes
|
||||
FLOPs: 115992428544 flops
|
||||
|
||||
Runtime: 6.73655 ms
|
||||
Memory: 24.934 GiB/s
|
||||
|
||||
Math: 17218.4 GFLOP/s
|
||||
```
|
||||
|
||||
Note, the arguments which appear in the output may be used as command line parameters for subsequent invocations.
|
||||
|
||||
|
||||
## Example Tensor Core GEMM Operations
|
||||
|
||||
To execute kernels targeting Tensor Core operations, supply the flag `--op_class=tensorop` in the command line.
|
||||
```bash
|
||||
$ ./tools/profiler/cutlass_profiler --op_class=tensorop --m=3456 --n=4096 --k=8192
|
||||
|
||||
|
||||
|
||||
=============================
|
||||
Problem ID: 1
|
||||
|
||||
Provider: CUTLASS
|
||||
OperationKind: gemm
|
||||
Operation: cutlass_tensorop_s16816gemm_f16_256x128_32x3_nn_align8
|
||||
|
||||
Status: Success
|
||||
Verification: ON
|
||||
Disposition: Passed
|
||||
|
||||
cuBLAS: Passed
|
||||
|
||||
Arguments: --m=3456 --n=4096 --k=8192 --A=f16:column --B=f16:column --C=f32:column --alpha=1 --beta=0 --split_k_slices=1 \
|
||||
--batch_count=1 --op_class=tensorop --accum=f32 --cta_m=256 --cta_n=128 --cta_k=32 --stages=3 --warps_m=4 \
|
||||
--warps_n=2 --warps_k=1 --inst_m=16 --inst_n=8 --inst_k=16 --min_cc=80 --max_cc=1024
|
||||
|
||||
Bytes: 180355072 bytes
|
||||
FLOPs: 231956545536 flops
|
||||
|
||||
Runtime: 0.98647 ms
|
||||
Memory: 170.272 GiB/s
|
||||
|
||||
Math: 235138 GFLOP/s
|
||||
```
|
||||
|
||||
## Covering the problem space
|
||||
|
||||
All arguments may have single values or comma-delimited set of values. Integers may also be specified
|
||||
as an inclusive range with the following syntax `start:end:increment` or simply `start:end`.
|
||||
|
||||
For example, the following sweeps over the range of the GEMM K dimension from 8 to 4096 in increments
|
||||
of 8 elements.
|
||||
|
||||
```bash
|
||||
$ ./tools/profiler/cutlass_profiler --kernels=cutlass_simt_sgemm_128x128_nn --m=4352 --n=4096 --k=8:4096:8
|
||||
```
|
||||
|
||||
## Output
|
||||
|
||||
By default, runtime and computed GFLOP/s are reported for each operation and problem size. Additionally,
|
||||
a table of comma separated values are reported at the end of the execution. This may be output to a file
|
||||
with the `--output=<filename.csv>` command line option as shown:
|
||||
|
||||
```bash
|
||||
$ ./tools/profiler/cutlass_profiler --kernels=cutlass_simt_sgemm_128x128_nn \
|
||||
--m=3456 --n=4096 --k=8:4096:8 --output=report.csv
|
||||
```
|
||||
|
||||
To faclitate generation of pivot tables and charts, additional columns may be prepended with the
|
||||
`--tags=<column>:<value>` option. One or more tags may be specified using a comma-delimited list.
|
||||
|
||||
```bash
|
||||
$ ./tools/profiler/cutlass_profiler --kernels=cutlass_simt_sgemm_128x128_nn \
|
||||
--m=3456 --n=4096 --k=8:4096:8 --output=report.csv \
|
||||
--tags=cutlass:2.2,date:2020-06-08
|
||||
```
|
||||
|
||||
## CUTLASS 3.0 GEMM procedural names
|
||||
|
||||
CUTLASS 3.0 introduces a new naming convention for GEMMs used by the profiler targeting the NVIDIA
|
||||
Hopper architecture and beyond so as to indicate new features of the kernel within the name
|
||||
(e.g., the cluster shape).
|
||||
|
||||
To best illustrate this naming convention, we will walk through the meaning of each of the components
|
||||
in a GEMM kernel used by the profiler:
|
||||
|
||||
```
|
||||
cutlass3x_sm90_tensorop_s64x128x16gemm_f16_f16_f32_f16_f32_{optional-mixed-dtype-config}_128x128x64_2x1x1_0_ntn_align8
|
||||
```
|
||||
|
||||
The components within this name are as follows:
|
||||
|
||||
* `cutlass3x`: indicates that the kernel was generated through the CUTLASS 3.0 API
|
||||
* `sm90`: indicates that the kernel targets NVIDIA GPUs with compute capability 90
|
||||
* `tensorop`: indicates that the kernel makes use of NVIDIA Tensor Cores
|
||||
(as opposed to `simt`, which indicates the use of "CUDA cores")
|
||||
* `s`: indicates that the Tensor Core instruction being used accumulates in single precision
|
||||
(as opposed to `h`, which indicates half precision)
|
||||
* `64x128x16gemm`: indicates that the shape of the Tensor Core instruction being used (MxNxK) is 64x128x16
|
||||
* `f16_f16_f32_f16_f16`: indicates that the data types for operands A, B, Accumulator, C and D (in that order).
|
||||
* `optional-mixed-dtype-config`: optional, will be empty if this is not a mixed dtype kernel. For mixed dtype kernels, it contains `_cvt`, `_scl`, `_sclzr`, respectively, for convert-only, scale-only, scale-with-zero-point running modes. It further contains `_shfl` if the kernel uses a shuffled layout for the narrow data type input matrix.
|
||||
* `128x128x64`: indicates that the thread block shape used in the GEMM (MxNxK) is 128x128x64
|
||||
* `2x1x1`: indicates that the cluster shape being used is 2x1x1
|
||||
* `0`: indicates that the kernel uses the CollectiveBuilder's automatic stage calculation to determine the
|
||||
number of pipeline stages in the kernel. Note that `0` does not mean that no stages are used. A nonzero value indicates that automatic stage calculation is not performed and indicates the number of pipeline stages to be used.
|
||||
This 0 is only added to the kernel's procedural name, the profiler will still report the actual stage count
|
||||
when printing the kernel argument details (`--stages=N`) and kernel discovery will still support filtering through the `--stages` argument.
|
||||
* `ntn`: indicates that the layouts for operands A, B, and C are column major ("n"; non-transposed),
|
||||
row major ("t"; transposed), and column major, respectively.
|
||||
* `align8`: indicates that the maximum alignment between operands A and B is 8.
|
||||
|
||||
Note that in some special cases where the input A/B types do not match that of the MMA
|
||||
instruction's, the MMA facing input type is added to the instruction string as well.
|
||||
|
||||
```
|
||||
cutlass3x_sm90_tensorop_s64x128x8tf32gemm_f32_f32_f32_f32_f32_128x128x32_2x1x1_0_tnn_align4
|
||||
```
|
||||
|
||||
* `s64x128x8tf32gemm`: indicates that the MMA consumes inputs in `tf32` format, and therefore
|
||||
the kernel performs rounding of the `f32` values in global memory while loading them into shared memory.
|
||||
|
||||
For custom mainloop or epilogue schedules, details of the opted-in schedule are appended to the end of the
|
||||
kernel name. For example,
|
||||
|
||||
```
|
||||
cutlass3x_sm90_tensorop_h64x128x16gemm_f16_f16_f16_void_f16_128x128x64_1x1x1_0_nnn_align8_warpspecialized_cooperative_epi_tma
|
||||
```
|
||||
|
||||
* `warpspecialized_cooperative`: Mainloop employs a persistent warp-specialized mainloop and kernel schedule.
|
||||
* `epi_tma`: Kernel epilogue employs TMA based vectorization.
|
||||
* `f16_f16_f16_void_f16`: In this case, C type is set to `void`, indicating that residual matrix support
|
||||
is disabled.
|
||||
|
||||
# Convolution
|
||||
|
||||
The CUTLASS Profiler is capable of executing 2-D and 3-D convolution problems for forwards and backwards
|
||||
operator variants.
|
||||
|
||||
The CUTLASS Profiler can be built with cuDNN enabled to use as a reference implementation. If CMake detects
|
||||
the cuDNN library available in the system, it is included as a dependency. This may be explicitly overridden
|
||||
with CMake flag `CUTLASS_ENABLE_CUDNN`.
|
||||
|
||||
```bash
|
||||
$ cmake .. -DCUTLASS_LIBRARY_OPERATIONS=conv2d -DCUTLASS_ENABLE_CUDNN=OFF
|
||||
...
|
||||
$ make -j16 cutlass_profiler
|
||||
```
|
||||
|
||||
|
||||
## Convolution Arguments
|
||||
|
||||
```bash
|
||||
$ ./tools/profiler/cutlass_profiler --help --operation=Conv2d
|
||||
|
||||
Conv2d
|
||||
|
||||
[enum] --conv_kind Convolutional operator (fprop, dgrad, wgrad)
|
||||
[int] --n,--input_n Input N dimension of the Conv2d problem space
|
||||
[int] --h,--input_h Input H dimension of the Conv2d problem space
|
||||
[int] --w,--input_w Input W dimension of the Conv2d problem space
|
||||
[int] --c,--input_c Input C dimension of the Conv2d problem space
|
||||
[int] --k,--filter_k Filter K dimension of the Conv2d problem space
|
||||
[int] --r,--filter_r Filter R dimension of the Conv2d problem space
|
||||
[int] --s,--filter_s Filter S dimension of the Conv2d problem space
|
||||
[int] --p,--output_p Output P dimension of the Conv2d problem space
|
||||
[int] --q,--output_q Output Q dimension of the Conv2d problem space
|
||||
[int] --g,--groups Number of convolution groups
|
||||
[int] --pad_h Padding in H direction
|
||||
[int] --pad_w Padding in W direction
|
||||
[int] --stride_h Stride in H direction
|
||||
[int] --stride_w Stride in W direction
|
||||
[int] --dilation_h Dilation in H direction
|
||||
[int] --dilation_w Dilation in W direction
|
||||
[tensor] --Activation Tensor storing the Activation operand
|
||||
[tensor] --Filter Tensor storing the Filter operand
|
||||
[tensor] --Output Tensor storing the Output operand
|
||||
[enum] --conv_mode Convolution filter mode (conv, cross)
|
||||
[enum] --iterator_algorithm,--iterator_algo Convolution iterator algorithm (analytic, optimized)
|
||||
[scalar] --alpha,--epilogue::alpha Epilogue scalar alpha
|
||||
[scalar] --beta,--epilogue::beta Epilogue scalar beta
|
||||
[enum] --split_k_mode,--split-k-mode SplitK mode for serial or parallel reduction (serial, parallel)
|
||||
[int] --split_k_slices,--split-k-slices Number of partitions of K dimension
|
||||
[enum] --eq_gemm_provider,--eq-gemm-provider Enable profiling equivalent gemm by the following providers (cutlass)
|
||||
[enum] --op_class,--opcode-class Class of math instruction (simt, tensorop, wmmatensorop, wmma)
|
||||
[enum] --accum,--accumulator-type Math instruction accumulator data type
|
||||
[int] --cta_m,--threadblock-shape::m Threadblock shape in the M dimension
|
||||
[int] --cta_n,--threadblock-shape::n Threadblock shape in the N dimension
|
||||
[int] --cta_k,--threadblock-shape::k Threadblock shape in the K dimension
|
||||
[int] --cluster_m,--cluster-shape::m Cluster shape in the M dimension
|
||||
[int] --cluster_n,--cluster-shape::n Cluster shape in the N dimension
|
||||
[int] --cluster_k,--cluster-shape::k Cluster shape in the K dimension
|
||||
[int] --cluster_m_fallback,--cluster-shape-fallback::m Fallback cluster shape in the M dimension
|
||||
[int] --cluster_n_fallback,--cluster-shape-fallback::n Fallback cluster shape in the N dimension
|
||||
[int] --cluster_k_fallback,--cluster-shape-fallback::k Fallback cluster shape in the K dimension
|
||||
[int] --stages,--threadblock-stages Number of stages of threadblock-scoped matrix multiply
|
||||
[int] --warps_m,--warp-count::m Number of warps within threadblock along the M dimension
|
||||
[int] --warps_n,--warp-count::n Number of warps within threadblock along the N dimension
|
||||
[int] --warps_k,--warp-count::k Number of warps within threadblock along the K dimension
|
||||
[int] --inst_m,--instruction-shape::m Math instruction shape in the M dimension
|
||||
[int] --inst_n,--instruction-shape::n Math instruction shape in the N dimension
|
||||
[int] --inst_k,--instruction-shape::k Math instruction shape in the K dimension
|
||||
[int] --min_cc,--minimum-compute-capability Minimum device compute capability
|
||||
[int] --max_cc,--maximum-compute-capability Maximum device compute capability
|
||||
|
||||
Examples:
|
||||
|
||||
Profile a particular convolution (specify all the convolution parameters):
|
||||
$ cutlass_profiler --operation=Conv2d --Activation=f16:nhwc --Filter=f16:nhwc --Output=f16 --accumulator-type=f32 --n=32 --h=14 --w=14 --c=8 --k=64 --r=3 --s=3 --pad_h=1 --pad_w=1 --stride_h=1 --stride_w=1 --dilation_h=1 --dilation_w=1
|
||||
|
||||
```
|
||||
|
||||
## Example CUDA Core Convolution Operation
|
||||
|
||||
Example command line for profiling forward propagation convolution kernels on CUDA cores is as follows:
|
||||
```bash
|
||||
$ ./tools/profiler/cutlass_profiler --kernels=simt_sfprop --verification-providers=device --n=8 --h=224 --w=224 --c=128 --k=128 --r=3 --s=3
|
||||
|
||||
|
||||
=============================
|
||||
Problem ID: 1
|
||||
|
||||
Provider: CUTLASS
|
||||
OperationKind: conv2d
|
||||
Operation: cutlass_simt_sfprop_optimized_128x128_8x2_nhwc
|
||||
|
||||
Status: Success
|
||||
Verification: ON
|
||||
Disposition: Passed
|
||||
|
||||
reference_device: Passed
|
||||
|
||||
Arguments: --conv_kind=fprop --n=8 --h=224 --w=224 --c=128 --k=128 --r=3 --s=3 --p=224 --q=224 --pad_h=1 --pad_w=1 \
|
||||
--stride_h=1 --stride_w=1 --dilation_h=1 --dilation_w=1 --Activation=f32:nhwc --Filter=f32:nhwc --Output=f32:nhwc \
|
||||
--conv_mode=cross --iterator_algorithm=optimized --alpha=1 --beta=0 --split_k_mode=serial --split_k_slices=1 \
|
||||
--eq_gemm_provider=none --op_class=simt --accum=f32 --cta_m=128 --cta_n=128 --cta_k=8 --stages=2 --warps_m=4 \
|
||||
--warps_n=2 --warps_k=1 --inst_m=1 --inst_n=1 --inst_k=1 --min_cc=50 --max_cc=1024
|
||||
|
||||
Bytes: 2055798784 bytes
|
||||
FLOPs: 118482796544 flops
|
||||
|
||||
Runtime: 8.13237 ms
|
||||
Memory: 235.431 GiB/s
|
||||
|
||||
Math: 14569.3 GFLOP/s
|
||||
|
||||
```
|
||||
|
||||
## Example Tensor Core Convolution Operation
|
||||
|
||||
Example command line for profiling forward propagation convolution kernels runing on Tensor Cores is as follows:
|
||||
```bash
|
||||
$ ./tools/profiler/cutlass_profiler --kernels=tensorop*fprop --verification-providers=device --n=8 --h=224 --w=224 --c=128 --k=128 --r=3 --s=3
|
||||
|
||||
|
||||
|
||||
=============================
|
||||
Problem ID: 1
|
||||
|
||||
Provider: CUTLASS
|
||||
OperationKind: conv2d
|
||||
Operation: cutlass_tensorop_s16816fprop_optimized_f16_128x128_64x4_nhwc
|
||||
|
||||
Status: Success
|
||||
Verification: ON
|
||||
Disposition: Passed
|
||||
|
||||
reference_device: Passed
|
||||
|
||||
Arguments: --conv_kind=fprop --n=8 --h=224 --w=224 --c=128 --k=128 --r=3 --s=3 --p=224 --q=224 --pad_h=1 --pad_w=1 \
|
||||
--stride_h=1 --stride_w=1 --dilation_h=1 --dilation_w=1 --Activation=f16:nhwc --Filter=f16:nhwc --Output=f32:nhwc \
|
||||
--conv_mode=cross --iterator_algorithm=optimized --alpha=1 --beta=0 --split_k_mode=serial --split_k_slices=1 \
|
||||
--eq_gemm_provider=none --op_class=tensorop --accum=f32 --cta_m=128 --cta_n=128 --cta_k=64 --stages=4 \
|
||||
--warps_m=2 --warps_n=2 --warps_k=1 --inst_m=16 --inst_n=8 --inst_k=16 --min_cc=80 --max_cc=1024
|
||||
|
||||
Bytes: 1130659840 bytes
|
||||
FLOPs: 118482796544 flops
|
||||
|
||||
Runtime: 0.945071 ms
|
||||
Memory: 1114.21 GiB/s
|
||||
|
||||
Math: 125369 GFLOP/s
|
||||
|
||||
|
||||
```
|
||||
|
||||
# Copyright
|
||||
|
||||
Copyright (c) 2017 - 2025 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
```
|
||||
Redistribution and use in source and binary forms, with or without
|
||||
modification, are permitted provided that the following conditions are met:
|
||||
|
||||
1. Redistributions of source code must retain the above copyright notice, this
|
||||
list of conditions and the following disclaimer.
|
||||
|
||||
2. Redistributions in binary form must reproduce the above copyright notice,
|
||||
this list of conditions and the following disclaimer in the documentation
|
||||
and/or other materials provided with the distribution.
|
||||
|
||||
3. Neither the name of the copyright holder nor the names of its
|
||||
contributors may be used to endorse or promote products derived from
|
||||
this software without specific prior written permission.
|
||||
|
||||
THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS "AS IS"
|
||||
AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE
|
||||
IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE ARE
|
||||
DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT HOLDER OR CONTRIBUTORS BE LIABLE
|
||||
FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL
|
||||
DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR
|
||||
SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER
|
||||
CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY,
|
||||
OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE
|
||||
OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
|
||||
```
|
||||
1190
media/docs/cpp/programming_guidelines.md
Normal file
1190
media/docs/cpp/programming_guidelines.md
Normal file
File diff suppressed because it is too large
Load Diff
783
media/docs/cpp/quickstart.md
Normal file
783
media/docs/cpp/quickstart.md
Normal file
@@ -0,0 +1,783 @@
|
||||

|
||||
|
||||
# Quickstart
|
||||
|
||||
## Prerequisites
|
||||
|
||||
CUTLASS requires:
|
||||
- NVIDIA CUDA Toolkit (11.4 or later required, [12.0](https://developer.nvidia.com/cuda-toolkit) recommended)
|
||||
- CMake 3.18+
|
||||
- host compiler supporting C++17 or greater (minimum g++ 7.5.0)
|
||||
- Python 3.6+
|
||||
|
||||
CUTLASS may be optionally compiled and linked with
|
||||
- cuBLAS
|
||||
- cuDNN v7.6 or later
|
||||
|
||||
## Initial build steps
|
||||
|
||||
Construct a build directory and run CMake.
|
||||
```bash
|
||||
$ export CUDACXX=${CUDA_INSTALL_PATH}/bin/nvcc
|
||||
|
||||
$ mkdir build && cd build
|
||||
|
||||
$ cmake .. -DCUTLASS_NVCC_ARCHS=90a # compiles for NVIDIA Hopper GPU architecture
|
||||
$ cmake .. -DCUTLASS_NVCC_ARCHS=100a # compiles for NVIDIA Blackwell SM100 GPU architecture
|
||||
```
|
||||
|
||||
If your goal is strictly to build only the CUTLASS Profiler and to minimize compilation time, we suggest
|
||||
executing the following CMake command in an empty `build/` directory.
|
||||
```bash
|
||||
$ cmake .. -DCUTLASS_NVCC_ARCHS=90a -DCUTLASS_ENABLE_TESTS=OFF -DCUTLASS_UNITY_BUILD_ENABLED=ON
|
||||
```
|
||||
|
||||
This reduces overall compilation time by excluding unit tests and enabling the unity build.
|
||||
|
||||
You may reduce build times by compiling only certain operations by setting the `CUTLASS_LIBRARY_OPERATIONS` flag as shown below,
|
||||
executed from an empty `build/` directory. This only compiles 2-D convolution kernels.
|
||||
|
||||
```bash
|
||||
$ cmake .. -DCUTLASS_NVCC_ARCHS=90a -DCUTLASS_LIBRARY_OPERATIONS=conv2d
|
||||
```
|
||||
|
||||
You may also filter kernels by name by supplying a filter string with flag `CUTLASS_LIBRARY_KERNELS`. For example the below command selects only CUTLASS-3 kernels.
|
||||
|
||||
```bash
|
||||
$ cmake .. -DCUTLASS_NVCC_ARCHS=90a -DCUTLASS_LIBRARY_KERNELS=cutlass3x*
|
||||
```
|
||||
See more examples on selectively compiling CUTLASS GEMM and convolution kernels [here](quickstart.md#example-cmake-commands).
|
||||
|
||||
You may explicitly exclude cuBLAS and cuDNN as dependencies with the following CMake flags.
|
||||
- `-DCUTLASS_ENABLE_CUBLAS=OFF`
|
||||
- `-DCUTLASS_ENABLE_CUDNN=OFF`
|
||||
|
||||
|
||||
## Build and run the CUTLASS Profiler
|
||||
|
||||
From the `build/` directory created above, compile the CUTLASS Profiler.
|
||||
```bash
|
||||
$ make cutlass_profiler -j12
|
||||
```
|
||||
|
||||
Then execute the CUTLASS Profiler computing GEMM, execute the following command.
|
||||
```bash
|
||||
$ ./tools/profiler/cutlass_profiler --kernels=sgemm --m=4352 --n=4096 --k=4096
|
||||
|
||||
=============================
|
||||
Problem ID: 1
|
||||
|
||||
Provider: CUTLASS
|
||||
Operation: cutlass_simt_sgemm_128x128_nn
|
||||
|
||||
Disposition: Passed
|
||||
Status: Success
|
||||
|
||||
Arguments: --m=4352 --n=4096 --k=4096 --A=f32:column --B=f32:column --C=f32:column --alpha=1 --beta=0 \
|
||||
--split_k_slices=1 --batch_count=1 --op_class=simt --accum=f32 --cta_m=128 --cta_n=128 --cta_k=8 \
|
||||
--stages=2 --warps_m=2 --warps_n=2 --warps_k=1 --inst_m=1 --inst_n=1 --inst_k=1 --min_cc=50 \
|
||||
--max_cc=1024
|
||||
|
||||
Bytes: 52428800 bytes
|
||||
FLOPs: 146064539648 flops
|
||||
|
||||
Runtime: 10.5424 ms
|
||||
Memory: 4.63158 GiB/s
|
||||
|
||||
Math: 13854.9 GFLOP/s
|
||||
```
|
||||
|
||||
To execute the CUTLASS Profiler for convolution, run the following example.
|
||||
```bash
|
||||
$ ./tools/profiler/cutlass_profiler --kernels=s1688fprop --n=8 --h=224 --w=224 --c=128 --k=128 --r=3 --s=3 --pad_h=1 --pad_w=1
|
||||
```
|
||||
|
||||
To execute all CUTLASS 2-D convolution operators, execute the following.
|
||||
```bash
|
||||
$ ./tools/profiler/cutlass_profiler --operation=conv2d --n=8 --h=224 --w=224 --c=128 --k=128 --r=3 --s=3
|
||||
|
||||
|
||||
=============================
|
||||
Problem ID: 1
|
||||
|
||||
Provider: CUTLASS
|
||||
OperationKind: conv2d
|
||||
Operation: cutlass_simt_sfprop_optimized_128x128_8x2_nhwc
|
||||
|
||||
Status: Success
|
||||
Verification: ON
|
||||
Disposition: Passed
|
||||
|
||||
reference_device: Passed
|
||||
|
||||
Arguments: --conv_kind=fprop --n=8 --h=224 --w=224 --c=128 --k=128 --r=3 --s=3 --p=224 --q=224 --pad_h=1 --pad_w=1 \
|
||||
--stride_h=1 --stride_w=1 --dilation_h=1 --dilation_w=1 --Activation=f32:nhwc --Filter=f32:nhwc --Output=f32:nhwc \
|
||||
--conv_mode=cross --iterator_algorithm=optimized --alpha=1 --beta=0 --split_k_mode=serial --split_k_slices=1 \
|
||||
--eq_gemm_provider=none --op_class=simt --accum=f32 --cta_m=128 --cta_n=128 --cta_k=8 --stages=2 --warps_m=4 \
|
||||
--warps_n=2 --warps_k=1 --inst_m=1 --inst_n=1 --inst_k=1 --min_cc=50 --max_cc=1024
|
||||
|
||||
Bytes: 2055798784 bytes
|
||||
FLOPs: 118482796544 flops
|
||||
|
||||
Runtime: 8.13237 ms
|
||||
Memory: 235.431 GiB/s
|
||||
|
||||
Math: 14569.3 GFLOP/s
|
||||
|
||||
```
|
||||
|
||||
See [documentation for the CUTLASS Profiler](profiler.md) for more details.
|
||||
|
||||
## Build and run CUTLASS Unit Tests
|
||||
|
||||
From the `build/` directory created above, simply build the target `test_unit` to compile and run
|
||||
all unit tests.
|
||||
|
||||
```bash
|
||||
$ make test_unit -j
|
||||
...
|
||||
...
|
||||
...
|
||||
[----------] Global test environment tear-down
|
||||
[==========] 946 tests from 57 test cases ran. (10812 ms total)
|
||||
[ PASSED ] 946 tests.
|
||||
$
|
||||
```
|
||||
The exact number of tests run is subject to change as we add more functionality.
|
||||
|
||||
No tests should fail. Unit tests automatically construct the appropriate runtime filters
|
||||
to avoid executing on architectures that do not support all features under test.
|
||||
|
||||
The unit tests are arranged hierarchically mirroring the CUTLASS Template Library. This enables
|
||||
parallelism in building and running tests as well as reducing compilation times when a specific
|
||||
set of tests are desired.
|
||||
|
||||
For example, the following executes strictly the warp-level GEMM tests.
|
||||
```bash
|
||||
$ make test_unit_gemm_warp -j
|
||||
...
|
||||
...
|
||||
[----------] 3 tests from SM75_warp_gemm_tensor_op_congruous_f16
|
||||
[ RUN ] SM75_warp_gemm_tensor_op_congruous_f16.128x128x8_32x128x8_16x8x8
|
||||
[ OK ] SM75_warp_gemm_tensor_op_congruous_f16.128x128x8_32x128x8_16x8x8 (0 ms)
|
||||
[ RUN ] SM75_warp_gemm_tensor_op_congruous_f16.128x128x32_64x64x32_16x8x8
|
||||
[ OK ] SM75_warp_gemm_tensor_op_congruous_f16.128x128x32_64x64x32_16x8x8 (2 ms)
|
||||
[ RUN ] SM75_warp_gemm_tensor_op_congruous_f16.128x128x32_32x32x32_16x8x8
|
||||
[ OK ] SM75_warp_gemm_tensor_op_congruous_f16.128x128x32_32x32x32_16x8x8 (1 ms)
|
||||
[----------] 3 tests from SM75_warp_gemm_tensor_op_congruous_f16 (3 ms total)
|
||||
...
|
||||
...
|
||||
[----------] Global test environment tear-down
|
||||
[==========] 104 tests from 32 test cases ran. (294 ms total)
|
||||
[ PASSED ] 104 tests.
|
||||
[100%] Built target test_unit_gemm_warp
|
||||
```
|
||||
|
||||
## Building for Multiple Architectures
|
||||
|
||||
To minimize compilation time, specific GPU architectures can be enabled via the CMake command,
|
||||
selected by [CUDA Compute Capability.](https://docs.nvidia.com/cuda/cuda-c-programming-guide/index.html#compute-capabilities)
|
||||
|
||||
**NVIDIA Blackwell Architecture.**
|
||||
```bash
|
||||
$ cmake .. -DCUTLASS_NVCC_ARCHS=100a # compiles for NVIDIA Blackwell GPU architecture
|
||||
```
|
||||
|
||||
**NVIDIA Hopper Architecture.**
|
||||
```bash
|
||||
$ cmake .. -DCUTLASS_NVCC_ARCHS=90a # compiles for NVIDIA Hopper GPU architecture
|
||||
```
|
||||
|
||||
**NVIDIA Ampere Architecture.**
|
||||
```bash
|
||||
$ cmake .. -DCUTLASS_NVCC_ARCHS=80 # compiles for NVIDIA Ampere GPU architecture
|
||||
```
|
||||
|
||||
**NVIDIA Turing Architecture.**
|
||||
```bash
|
||||
$ cmake .. -DCUTLASS_NVCC_ARCHS=75 # compiles for NVIDIA Turing GPU architecture
|
||||
```
|
||||
|
||||
**NVIDIA Volta Architecture.**
|
||||
```bash
|
||||
$ cmake .. -DCUTLASS_NVCC_ARCHS=70 # compiles for NVIDIA Volta GPU architecture
|
||||
```
|
||||
|
||||
**NVIDIA Pascal Architecture.**
|
||||
```bash
|
||||
$ cmake .. -DCUTLASS_NVCC_ARCHS="60;61" # compiles for NVIDIA Pascal GPU architecture
|
||||
```
|
||||
|
||||
**NVIDIA Maxwell Architecture.**
|
||||
```bash
|
||||
$ cmake .. -DCUTLASS_NVCC_ARCHS="50;53" # compiles for NVIDIA Maxwell GPU architecture
|
||||
```
|
||||
|
||||
## Using CUTLASS within other applications
|
||||
|
||||
Applications should list [`/include`](https://github.com/NVIDIA/cutlass/tree/main/include) within their include paths. They must be
|
||||
compiled as C++17 or greater.
|
||||
|
||||
**Example:** print the contents of a variable storing half-precision data.
|
||||
```c++
|
||||
#include <iostream>
|
||||
#include <cutlass/cutlass.h>
|
||||
#include <cutlass/numeric_types.h>
|
||||
#include <cutlass/core_io.h>
|
||||
|
||||
int main() {
|
||||
|
||||
cutlass::half_t x = 2.25_hf;
|
||||
|
||||
std::cout << x << std::endl;
|
||||
|
||||
return 0;
|
||||
}
|
||||
```
|
||||
|
||||
## Launching a GEMM kernel in CUDA
|
||||
|
||||
**Example:** launch a mixed-precision GEMM targeting Turing Tensor Cores.
|
||||
|
||||
_Note, this example uses CUTLASS Utilities. Be sure `tools/util/include` is listed as an include path._
|
||||
```c++
|
||||
#include <cutlass/numeric_types.h>
|
||||
#include <cutlass/gemm/device/gemm.h>
|
||||
|
||||
#include <cutlass/util/host_tensor.h>
|
||||
|
||||
int main() {
|
||||
|
||||
// Define the GEMM operation
|
||||
using Gemm = cutlass::gemm::device::Gemm<
|
||||
cutlass::half_t, // ElementA
|
||||
cutlass::layout::ColumnMajor, // LayoutA
|
||||
cutlass::half_t, // ElementB
|
||||
cutlass::layout::ColumnMajor, // LayoutB
|
||||
cutlass::half_t, // ElementOutput
|
||||
cutlass::layout::ColumnMajor, // LayoutOutput
|
||||
float, // ElementAccumulator
|
||||
cutlass::arch::OpClassTensorOp, // tag indicating Tensor Cores
|
||||
cutlass::arch::Sm75 // tag indicating target GPU compute architecture
|
||||
>;
|
||||
|
||||
Gemm gemm_op;
|
||||
cutlass::Status status;
|
||||
|
||||
//
|
||||
// Define the problem size
|
||||
//
|
||||
int M = 512;
|
||||
int N = 256;
|
||||
int K = 128;
|
||||
|
||||
float alpha = 1.25f;
|
||||
float beta = -1.25f;
|
||||
|
||||
//
|
||||
// Allocate device memory
|
||||
//
|
||||
|
||||
cutlass::HostTensor<cutlass::half_t, cutlass::layout::ColumnMajor> A({M, K});
|
||||
cutlass::HostTensor<cutlass::half_t, cutlass::layout::ColumnMajor> B({K, N});
|
||||
cutlass::HostTensor<cutlass::half_t, cutlass::layout::ColumnMajor> C({M, N});
|
||||
|
||||
cutlass::half_t const *ptrA = A.device_data();
|
||||
cutlass::half_t const *ptrB = B.device_data();
|
||||
cutlass::half_t const *ptrC = C.device_data();
|
||||
cutlass::half_t *ptrD = C.device_data();
|
||||
|
||||
int lda = A.device_ref().stride(0);
|
||||
int ldb = B.device_ref().stride(0);
|
||||
int ldc = C.device_ref().stride(0);
|
||||
int ldd = C.device_ref().stride(0);
|
||||
//
|
||||
// Launch GEMM on the device
|
||||
//
|
||||
|
||||
status = gemm_op({
|
||||
{M, N, K},
|
||||
{ptrA, lda}, // TensorRef to A device tensor
|
||||
{ptrB, ldb}, // TensorRef to B device tensor
|
||||
{ptrC, ldc}, // TensorRef to C device tensor
|
||||
{ptrD, ldd}, // TensorRef to D device tensor - may be the same as C
|
||||
{alpha, beta} // epilogue operation arguments
|
||||
});
|
||||
|
||||
if (status != cutlass::Status::kSuccess) {
|
||||
return -1;
|
||||
}
|
||||
|
||||
return 0;
|
||||
}
|
||||
```
|
||||
|
||||
Note, the above could be simplified as follows using helper methods defined in `HostTensor`.
|
||||
```c++
|
||||
cutlass::HostTensor<cutlass::half_t, cutlass::layout::ColumnMajor> A({M, K});
|
||||
cutlass::HostTensor<cutlass::half_t, cutlass::layout::ColumnMajor> B({K, N});
|
||||
cutlass::HostTensor<cutlass::half_t, cutlass::layout::ColumnMajor> C({M, N});
|
||||
|
||||
//
|
||||
// Use the TensorRef returned by HostTensor::device_ref().
|
||||
//
|
||||
|
||||
status = gemm_op({
|
||||
{M, N, K},
|
||||
A.device_ref(), // TensorRef to A device tensor
|
||||
B.device_ref(), // TensorRef to B device tensor
|
||||
C.device_ref(), // TensorRef to C device tensor
|
||||
C.device_ref(), // TensorRef to D device tensor - may be the same as C
|
||||
{alpha, beta} // epilogue operation arguments
|
||||
});
|
||||
```
|
||||
|
||||
## Launching a GEMM kernel using CUTLASS 3.0 or newer
|
||||
|
||||
**Example:** launch a mixed-precision GEMM targeting Hopper Tensor Cores.
|
||||
|
||||
```c++
|
||||
#include "cutlass/cutlass.h"
|
||||
#include "cutlass/epilogue/collective/default_epilogue.hpp"
|
||||
#include "cutlass/epilogue/thread/linear_combination.h"
|
||||
#include "cutlass/gemm/collective/collective_builder.hpp"
|
||||
#include "cutlass/gemm/device/gemm_universal_adapter.h"
|
||||
#include "cutlass/gemm/kernel/gemm_universal.hpp"
|
||||
|
||||
#include "cutlass/util/host_tensor.h"
|
||||
#include "cutlass/util/packed_stride.hpp"
|
||||
|
||||
using namespace cute;
|
||||
|
||||
int main(int argc, char const **args) {
|
||||
|
||||
// A matrix configuration
|
||||
using ElementA = cutlass::half_t; // Element type for A matrix operand
|
||||
using LayoutA = cutlass::layout::RowMajor; // Layout type for A matrix operand
|
||||
constexpr int AlignmentA = 128 / cutlass::sizeof_bits<ElementA>::value; // Memory access granularity/alignment of A matrix in units of elements (up to 16 bytes)
|
||||
|
||||
// B matrix configuration
|
||||
using ElementB = cutlass::half_t; // Element type for B matrix operand
|
||||
using LayoutB = cutlass::layout::ColumnMajor; // Layout type for B matrix operand
|
||||
constexpr int AlignmentB = 128 / cutlass::sizeof_bits<ElementB>::value; // Memory access granularity/alignment of B matrix in units of elements (up to 16 bytes)
|
||||
|
||||
// C/D matrix configuration
|
||||
using ElementC = cutlass::half_t; // Element type for C and D matrix operands
|
||||
using LayoutC = cutlass::layout::ColumnMajor; // Layout type for C and D matrix operands
|
||||
|
||||
// Core kernel configurations
|
||||
using ElementAccumulator = float; // Element type for internal accumulation
|
||||
using ArchTag = cutlass::arch::Sm90; // Tag indicating the minimum SM that supports the intended feature
|
||||
using OperatorClass = cutlass::arch::OpClassTensorOp; // Operator class tag
|
||||
using TilesShape = Shape<_128,_128,_64>; // Threadblock-level tile size
|
||||
using ClusterShape = Shape<_1,_2,_1>; // Shape of the threadblocks in a cluster
|
||||
using StageCountType = cutlass::gemm::collective::StageCountAuto; // Stage count maximized based on the tile size
|
||||
using KernelSchedule = cutlass::gemm::collective::KernelScheduleAuto; // Kernel to launch based on the default setting in the Collective Builder
|
||||
|
||||
using CollectiveMainloop = typename cutlass::gemm::collective::CollectiveBuilder<
|
||||
ArchTag, OperatorClass,
|
||||
ElementA, LayoutA, AlignmentA,
|
||||
ElementB, LayoutB, AlignmentB,
|
||||
ElementAccumulator,
|
||||
TilesShape, ClusterShape,
|
||||
cutlass::gemm::collective::StageCountAuto,
|
||||
cutlass::gemm::collective::KernelScheduleAuto
|
||||
>::CollectiveOp;
|
||||
|
||||
using CollectiveEpilogue = cutlass::epilogue::collective::DefaultEpilogue<
|
||||
cutlass::gemm::TagToStrideC_t<LayoutC>,
|
||||
cutlass::gemm::TagToStrideC_t<LayoutC>,
|
||||
cutlass::epilogue::thread::LinearCombination<ElementC, 1, ElementAccumulator, ElementAccumulator>>;
|
||||
|
||||
using GemmKernel = cutlass::gemm::kernel::GemmUniversal<
|
||||
Shape<int,int,int>, // Indicates ProblemShape
|
||||
CollectiveMainloop,
|
||||
CollectiveEpilogue
|
||||
>;
|
||||
|
||||
using Gemm = cutlass::gemm::device::GemmUniversalAdapter<GemmKernel>;
|
||||
|
||||
Gemm gemm_op;
|
||||
cutlass::Status status;
|
||||
|
||||
//
|
||||
// Define the problem size
|
||||
//
|
||||
|
||||
int M = 512;
|
||||
int N = 256;
|
||||
int K = 128;
|
||||
|
||||
float alpha = 1.25f;
|
||||
float beta = -1.25f;
|
||||
|
||||
//
|
||||
// Allocate device memory
|
||||
//
|
||||
|
||||
cutlass::DeviceAllocation<typename Gemm::ElementA> block_A;
|
||||
cutlass::DeviceAllocation<typename Gemm::ElementB> block_B;
|
||||
cutlass::DeviceAllocation<typename Gemm::ElementC> block_C;
|
||||
cutlass::DeviceAllocation<typename Gemm::EpilogueOutputOp::ElementOutput> block_D;
|
||||
|
||||
using StrideA = typename Gemm::GemmKernel::StrideA;
|
||||
using StrideB = typename Gemm::GemmKernel::StrideB;
|
||||
using StrideC = typename Gemm::GemmKernel::StrideC;
|
||||
using StrideD = typename Gemm::GemmKernel::StrideD;
|
||||
|
||||
StrideA stride_A;
|
||||
StrideB stride_B;
|
||||
StrideC stride_C;
|
||||
StrideD stride_D;
|
||||
|
||||
stride_A = cutlass::make_cute_packed_stride(StrideA{}, {M, K, 1});
|
||||
stride_B = cutlass::make_cute_packed_stride(StrideB{}, {N, K, 1});
|
||||
stride_C = cutlass::make_cute_packed_stride(StrideC{}, {M, N, 1});
|
||||
stride_D = cutlass::make_cute_packed_stride(StrideD{}, {M, N, 1});
|
||||
|
||||
block_A.reset(M * K);
|
||||
block_B.reset(K * N);
|
||||
block_C.reset(M * N);
|
||||
block_D.reset(M * N);
|
||||
|
||||
//
|
||||
// Launch GEMM on the device
|
||||
//
|
||||
|
||||
status = gemm_op({
|
||||
cutlass::gemm::GemmUniversalMode::kGemm,
|
||||
{M, N, K},
|
||||
block_A.get(),
|
||||
stride_A,
|
||||
block_B.get(),
|
||||
stride_B,
|
||||
{block_C.get(), stride_C, block_D.get(), stride_D, {alpha, beta}}
|
||||
});
|
||||
|
||||
if (status != cutlass::Status::kSuccess) {
|
||||
return -1;
|
||||
}
|
||||
|
||||
return 0;
|
||||
}
|
||||
```
|
||||
|
||||
# CUTLASS Library
|
||||
|
||||
The [CUTLASS Library](https://github.com/NVIDIA/cutlass/tree/main/tools/library) defines an API for managing and executing collections of compiled
|
||||
kernel instances and launching them from host code without template instantiations in client code.
|
||||
|
||||
The host-side launch API is designed to be analogous to BLAS implementations for convenience, though its
|
||||
kernel selection procedure is intended only to be functionally sufficient. It may not launch the
|
||||
optimal tile size for a given problem. It chooses the first available kernel whose data types,
|
||||
layouts, and alignment constraints satisfy the given problem. Kernel instances and a data structure
|
||||
describing them are completely available to client applications which may choose to implement their
|
||||
own selection logic.
|
||||
|
||||
[cuBLAS](https://developer.nvidia.com/cublas) offers the best performance and functional coverage
|
||||
for dense matrix computations on NVIDIA GPUs.
|
||||
|
||||
The CUTLASS Library is used by the CUTLASS Profiler to manage kernel instances, and it is also used
|
||||
by several SDK examples.
|
||||
|
||||
* [10_planar_complex](https://github.com/NVIDIA/cutlass/tree/main/examples/10_planar_complex/planar_complex.cu)
|
||||
* [11_planar_complex_array](https://github.com/NVIDIA/cutlass/tree/main/examples/11_planar_complex_array/planar_complex_array.cu)
|
||||
|
||||
The CUTLASS Library defines enumerated types describing numeric data types, matrix and tensor
|
||||
layouts, math operation classes, complex transformations, and more.
|
||||
|
||||
Client applications should specify [`tools/library/include`](https://github.com/NVIDIA/cutlass/tree/main/tools/library/include) in their
|
||||
include paths and link against libcutlas_lib.so.
|
||||
|
||||
The CUTLASS SDK example [10_planar_complex](https://github.com/NVIDIA/cutlass/tree/main/examples/10_planar_complex/CMakeLists.txt) specifies
|
||||
its dependency on the CUTLASS Library with the following CMake command.
|
||||
```
|
||||
target_link_libraries(
|
||||
10_planar_complex
|
||||
PRIVATE
|
||||
cutlass_lib
|
||||
cutlass_tools_util_includes
|
||||
)
|
||||
```
|
||||
|
||||
A sample kernel launch from host-side C++ is shown as follows.
|
||||
|
||||
```c++
|
||||
#include "cutlass/library/library.h"
|
||||
#include "cutlass/library/handle.h"
|
||||
|
||||
int main() {
|
||||
|
||||
//
|
||||
// Define the problem size
|
||||
//
|
||||
int M = 512;
|
||||
int N = 256;
|
||||
int K = 128;
|
||||
|
||||
float alpha = 1.25f;
|
||||
float beta = -1.25f;
|
||||
|
||||
//
|
||||
// Allocate device memory
|
||||
//
|
||||
|
||||
cutlass::HostTensor<float, cutlass::layout::ColumnMajor> A({M, K});
|
||||
cutlass::HostTensor<float, cutlass::layout::ColumnMajor> B({K, N});
|
||||
cutlass::HostTensor<float, cutlass::layout::ColumnMajor> C({M, N});
|
||||
|
||||
float const *ptrA = A.device_data();
|
||||
float const *ptrB = B.device_data();
|
||||
float const *ptrC = C.device_data();
|
||||
float *ptrD = C.device_data();
|
||||
|
||||
int lda = A.device_ref().stride(0);
|
||||
int ldb = B.device_ref().stride(0);
|
||||
int ldc = C.device_ref().stride(0);
|
||||
int ldd = D.device_ref().stride(0);
|
||||
|
||||
//
|
||||
// CUTLASS Library call to execute device GEMM
|
||||
//
|
||||
|
||||
cutlass::library::Handle handle;
|
||||
|
||||
//
|
||||
// Launch GEMM on CUDA device.
|
||||
//
|
||||
|
||||
cutlass::Status status = handle.gemm(
|
||||
M,
|
||||
N,
|
||||
K,
|
||||
|
||||
cutlass::library::NumericTypeID::kF32, // data type of internal accumulation
|
||||
cutlass::library::NumericTypeID::kF32, // data type of alpha/beta scalars
|
||||
|
||||
&alpha, // pointer to alpha scalar
|
||||
|
||||
cutlass::library::NumericTypeID::kF32, // data type of A matrix
|
||||
cutlass::library::LayoutTypeID::kColumnMajor, // layout of A matrix
|
||||
ptrA, // pointer to A matrix in device memory
|
||||
lda, // leading dimension of A matrix
|
||||
|
||||
cutlass::library::NumericTypeID::kF32, // data type of B matrix
|
||||
cutlass::library::LayoutTypeID::kColumnMajor, // layout of B matrix
|
||||
ptrB, // pointer to B matrix in device memory
|
||||
ldb, // leading dimension of B matrix
|
||||
|
||||
&beta, // pointer to beta scalar
|
||||
|
||||
cutlass::library::NumericTypeID::kF32, // data type of C and D matrix
|
||||
|
||||
ptrC, // pointer to C matrix in device memory
|
||||
ldc, // leading dimension fo C matrix
|
||||
|
||||
ptrD, // pointer to D matrix in device memory
|
||||
ldd // leading dimension of D matrix
|
||||
);
|
||||
|
||||
if (status != cutlass::Status::kSuccess) {
|
||||
return -1;
|
||||
}
|
||||
|
||||
return 0;
|
||||
}
|
||||
```
|
||||
|
||||
# Example CMake Commands
|
||||
|
||||
To instantiate all operations supporting all tile sizes, data types, and alignment constraints, specify
|
||||
`-DCUTLASS_LIBRARY_KERNELS=all` when running `cmake`.
|
||||
```bash
|
||||
$ cmake .. -DCUTLASS_NVCC_ARCHS='70;75;80' -DCUTLASS_LIBRARY_KERNELS=all
|
||||
```
|
||||
The above command line generates about twenty thousand kernels targeting NVIDIA Ampere, Turing, and Volta architectures.
|
||||
Compiling thousands of kernels for three different architectures is time-consuming. Additionally, this would also result
|
||||
in a large binary size and on some platforms linker to fail on building the library.
|
||||
|
||||
Enabling the "unity build" instantiates multiple kernel instances in each compilation unit, thereby reducing binary size
|
||||
and avoiding linker limitations on some platforms.
|
||||
```bash
|
||||
$ cmake .. -DCUTLASS_NVCC_ARCHS="70;75;80" -DCUTLASS_LIBRARY_KERNELS=all -DCUTLASS_UNITY_BUILD_ENABLED=ON
|
||||
```
|
||||
|
||||
It is advised to only compile CUTLASS kernels for NVIDIA architectures one plans on running. Furthermore, kernels
|
||||
can be selectively included in the CUTLASS Library by specifying filter strings and wildcard characters when executing CMake.
|
||||
|
||||
Several examples are defined below for convenience. They may be combined as a comma-delimited list.
|
||||
Compling only the kernels desired reduces compilation time.
|
||||
|
||||
|
||||
## GEMM CMake Examples
|
||||
**Example.** All GEMM kernels targeting NVIDIA Ampere Tensor Cores.
|
||||
```bash
|
||||
$ cmake .. -DCUTLASS_NVCC_ARCHS=80 -DCUTLASS_LIBRARY_KERNELS=tensorop*gemm
|
||||
```
|
||||
|
||||
**Example.** All GEMM kernels targeting NVIDIA Turing Tensor Cores.
|
||||
```bash
|
||||
$ cmake .. -DCUTLASS_NVCC_ARCHS=75 -DCUTLASS_LIBRARY_KERNELS=tensorop*gemm
|
||||
```
|
||||
|
||||
**Example.** All GEMM kernels with FP32 accumulation targeting NVIDIA Ampere, Turing, and Volta architectures.
|
||||
```bash
|
||||
$ cmake .. -DCUTLASS_NVCC_ARCHS="70;75;80" -DCUTLASS_LIBRARY_KERNELS=s*gemm
|
||||
```
|
||||
|
||||
**Example.** All kernels which expect A and B to be column-major or row-major targeting NVIDIA Ampere, Turing, and Volta architectures.
|
||||
```bash
|
||||
$ cmake .. -DCUTLASS_NVCC_ARCHS="70;75;80" -DCUTLASS_LIBRARY_KERNELS=gemm*nn,gemm*tt
|
||||
```
|
||||
|
||||
**Example.** All planar complex GEMM variants targeting NVIDIA Ampere, Turing, and Volta architectures.
|
||||
```bash
|
||||
$ cmake .. -DCUTLASS_NVCC_ARCHS="70;75;80" -DCUTLASS_LIBRARY_KERNELS=planar_complex
|
||||
```
|
||||
|
||||
## Convolution CMake Examples
|
||||
**Example.** All convolution kernels targeting NVIDIA Ampere's 16816 Tensor Core operation
|
||||
```bash
|
||||
$ cmake .. -DCUTLASS_NVCC_ARCHS='80' -DCUTLASS_LIBRARY_KERNELS=s16816fprop,s16816dgrad,s16816wgrad
|
||||
```
|
||||
|
||||
**Example.** All forward propagation (fprop) convolution kernels targeting CUDA Cores for multiple NVIDIA architectures
|
||||
```bash
|
||||
$ cmake .. -DCUTLASS_NVCC_ARCHS='50;60;61;70;75;80' -DCUTLASS_LIBRARY_KERNELS=sfprop
|
||||
```
|
||||
|
||||
**Example.** All forward propagation (fprop) convolution kernels with FP32 accumulation and FP16 input targeting NVIDIA Ampere's 16816 Tensor Core operation
|
||||
```bash
|
||||
$ cmake .. -DCUTLASS_NVCC_ARCHS='80' -DCUTLASS_LIBRARY_KERNELS=s16816fprop_*_f16
|
||||
```
|
||||
|
||||
**Example.** All backward weight gradient (wgrad) convolution kernels with FP32 accumulation, FP16 input, and optimized global memory iterator
|
||||
targeting NVIDIA Ampere, Turing, and Volta Tensor Core operations
|
||||
```bash
|
||||
$ cmake .. -DCUTLASS_NVCC_ARCHS='70;75;80' -DCUTLASS_LIBRARY_KERNELS=tensorop*s*wgrad_optimized_f16
|
||||
```
|
||||
|
||||
## Instantiating a Blackwell SM100 GEMM kernel
|
||||
|
||||
Blackwell SM100 kernels are instantiated very similarly to Hopper kernels. Let us start with an
|
||||
[FP8 GEMM without blockscaling](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/sm100_gemm_f8_f8_f8_tensor_op_s32_batch_alpha_beta.cu)
|
||||
as an example.
|
||||
|
||||
The kernel starts with setting up datatypes and cluster shapes.
|
||||
```c++
|
||||
using LayoutA = cutlass::layout::RowMajor;
|
||||
using LayoutB = cutlass::layout::ColumnMajor;
|
||||
using LayoutC = cutlass::layout::ColumnMajor;
|
||||
using ElementA = cutlass::float_e4m3_t;
|
||||
using ElementB = cutlass::float_e4m3_t;
|
||||
using ElementC = cutlass::float_e4m3_t;
|
||||
using ElementD = cutlass::float_e4m3_t;
|
||||
using ElementAccumulator = float;
|
||||
using ElementCompute = float;
|
||||
using ElementBias = cutlass::half_t;
|
||||
using MmaTileShape = cute::Shape<_128,_64,Int<128 / sizeof(ElementA)>>;
|
||||
using ClusterShape = cute::Shape<_1,_1,_1>;
|
||||
```
|
||||
|
||||
The epilogue needs to be instantiated first as the mainloop collective builder takes the shared memory budget of epilogue in the template parameter list. The 3.x epilogue collective builder API has not changed
|
||||
for Blackwell, so the epilogue fusion is built in a same way as an SM90 epilogue.
|
||||
|
||||
```c++
|
||||
using EpilogueSchedule = cutlass::epilogue::TmaWarpSpecialized1Sm;
|
||||
|
||||
using FusionOperation = cutlass::epilogue::fusion::LinearCombination<
|
||||
ElementD,
|
||||
ElementCompute,
|
||||
ElementC
|
||||
>;
|
||||
|
||||
using CollectiveEpilogue = typename cutlass::epilogue::collective::CollectiveBuilder<
|
||||
cutlass::arch::Sm100, cutlass::arch::OpClassTensorOp,
|
||||
MmaTileShape, ClusterShape,
|
||||
cutlass::epilogue::collective::EpilogueTileAuto,
|
||||
ElementAccumulator, ElementCompute,
|
||||
ElementC, LayoutC, 16 / sizeof(ElementC),
|
||||
ElementD, LayoutC, 16 / sizeof(ElementD),
|
||||
EpilogueSchedule,
|
||||
FusionOperation
|
||||
>::CollectiveOp;
|
||||
```
|
||||
|
||||
One can refer to our Sm100 unit tests as examples of how to correctly
|
||||
choose mainloop schedules. All of our dispatch policies can be found in [dispatch_policy.hpp](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/gemm/dispatch_policy.hpp)
|
||||
and more comprehensive Blackwell specific documentation for valid
|
||||
dispatch policies can be in [blackwell_functionality.md](./blackwell_functionality.md).
|
||||
|
||||
```c++
|
||||
using MainloopSchedule = cutlass::gemm::KernelTmaWarpSpecialized1SmSm100;
|
||||
using CollectiveMainloop = typename cutlass::gemm::collective::CollectiveBuilder<
|
||||
cutlass::arch::Sm100, cutlass::arch::OpClassTensorOp,
|
||||
ElementA, LayoutA, 16 / sizeof(ElementA),
|
||||
ElementB, LayoutB, 16 / sizeof(ElementB),
|
||||
ElementAccumulator,
|
||||
MmaTileShape, ClusterShape,
|
||||
cutlass::gemm::collective::StageCountAutoCarveout<static_cast<int>(sizeof(typename CollectiveEpilogue::SharedStorage))>,
|
||||
MainloopSchedule
|
||||
>::CollectiveOp;
|
||||
|
||||
using GemmKernel = cutlass::gemm::kernel::GemmUniversal<
|
||||
Shape<int,int,int,int>,
|
||||
CollectiveMainloop,
|
||||
CollectiveEpilogue
|
||||
>;
|
||||
```
|
||||
|
||||
Instantiating a blockscaled GEMM kernel is slightly different. Referring to an [MXFP8 GEMM](https://github.com/NVIDIA/cutlass/tree/main/test/unit/gemm/device/sm100_gemm_mxf8_mxf8_mxf8_tensor_op_f32_auto.cu) sample unit test, it takes a different tensor operation class:
|
||||
|
||||
```c++
|
||||
using ElementA = cutlass::mx_float8_t<cutlass::float_e4m3_t>;
|
||||
using ElementB = cutlass::mx_float8_t<cutlass::float_e4m3_t>;
|
||||
```
|
||||
|
||||
are needed in the mainloop builder:
|
||||
|
||||
```c++
|
||||
using CollectiveMainloop = typename cutlass::gemm::collective::CollectiveBuilder<
|
||||
cutlass::arch::Sm100, cutlass::arch::OpClassTensorOp,
|
||||
ElementA, LayoutA, 16,
|
||||
ElementB, LayoutB, 16,
|
||||
ElementAccumulator,
|
||||
MmaTileShape, ClusterShape,
|
||||
cutlass::gemm::collective::StageCountAutoCarveout<static_cast<int>(sizeof(typename CollectiveEpilogue::SharedStorage))>,
|
||||
cutlass::gemm::KernelScheduleAuto
|
||||
>::CollectiveOp;
|
||||
```
|
||||
|
||||
We encourage a user to refer to Sm100 unit tests and the generated profiler-based kernels as more comprehensive samples.
|
||||
|
||||
# Copyright
|
||||
|
||||
Copyright (c) 2017 - 2025 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
```
|
||||
Redistribution and use in source and binary forms, with or without
|
||||
modification, are permitted provided that the following conditions are met:
|
||||
|
||||
1. Redistributions of source code must retain the above copyright notice, this
|
||||
list of conditions and the following disclaimer.
|
||||
|
||||
2. Redistributions in binary form must reproduce the above copyright notice,
|
||||
this list of conditions and the following disclaimer in the documentation
|
||||
and/or other materials provided with the distribution.
|
||||
|
||||
3. Neither the name of the copyright holder nor the names of its
|
||||
contributors may be used to endorse or promote products derived from
|
||||
this software without specific prior written permission.
|
||||
|
||||
THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS "AS IS"
|
||||
AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE
|
||||
IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE ARE
|
||||
DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT HOLDER OR CONTRIBUTORS BE LIABLE
|
||||
FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL
|
||||
DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR
|
||||
SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER
|
||||
CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY,
|
||||
OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE
|
||||
OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
|
||||
```
|
||||
111
media/docs/cpp/terminology.md
Normal file
111
media/docs/cpp/terminology.md
Normal file
@@ -0,0 +1,111 @@
|
||||

|
||||
|
||||
# CUTLASS Terminology
|
||||
|
||||
**cute::Layout**: A `cute::Layout` vocabulary type composed of the hierarchical `cute::Shape` and `cute::Stride`
|
||||
tuples that is used throughout CUTLASS 3.0 to represent and manipulate thread and data layouts. More details are included in the [CuTe specific tensor type documentation](cute/03_tensor.md).
|
||||
|
||||
**cute::Tensor**: A pointer backed by a `cute::Layout` used to represent a tensor. More details are included in the [CuTe specific tensor type documentation](cute/03_tensor.md).
|
||||
|
||||
**Capacity**: (scalar) physical number of elements in memory required to store a multidimensional object; expressed as the type's LongIndex type
|
||||
- example: the capacity of a column-major matrix is `lda * N`
|
||||
|
||||
**Element**: data type describing one item in a multidimensional tensor, array, or matrix
|
||||
|
||||
**Extent**: (vector-valued quantity) the logical size of each dimension of a multidimensional index space. Consistent with the [C++ Standard Library](https://en.cppreference.com/w/cpp/types/extent).
|
||||
- `Coord<N> extent()`
|
||||
- `Index extent(int dim)`
|
||||
|
||||
**Fragment**: a register-backed array of elements used to store a thread's part of a tile
|
||||
|
||||
**Index**: signed integer representing quantities aligned with a logical dimension
|
||||
|
||||
**Layout**: functor mapping logical coordinates of a tensor to linear offset (as LongIndex); owns stride vectors, if any.
|
||||
|
||||
**LongIndex**: signed integer representing offsets in memory; typically wider than Index type
|
||||
|
||||
**Numeric Type**: a CUTLASS data type used to represent real-valued quantities; is trivially copyable.
|
||||
|
||||
**Pitch Linear**: linear memory allocation obtained from a user-defined 2-D size, which specifies the
|
||||
contiguous and strided dimensions of a tile.
|
||||
|
||||
**Planar Complex**: representation of complex tensors as two real-valued tensors, with real elements in one part and imaginary elements in another part of identical layout, separated by an offset
|
||||
|
||||
**Policy**: additional details extending the interface of a template guiding internal implementation;
|
||||
typically used to target specific design points known to be efficient
|
||||
|
||||
**Rank**: number of dimensions in a multidimensional index space, array, tensor, or matrix. Consistent with
|
||||
[C++ Standard Library](https://en.cppreference.com/w/cpp/types/rank)
|
||||
|
||||
**Register**: in device code, registers are the most efficient storage for statically sized arrays of elements.
|
||||
Arrays may be expected to be stored in registers if all accesses are made via constexpr indices or within
|
||||
fully unrolled loops.
|
||||
|
||||
**Residue**: partial tile or matrix computation which may require special accommodation for functional correctness or performance
|
||||
|
||||
**Size**: (scalar) number of logical elements in a tensor; equal to the product of each member of `extent()`
|
||||
- `LongIndex size()`
|
||||
|
||||
`sizeof_bits<T>::value` - template pattern returning the size of a numeric type or array in units of bits
|
||||
|
||||
**Storage**: when appropriate, refers to some alternative type used to store a packed collection of elements;
|
||||
may be used to handle bit-level packing or make types safe for use in unions
|
||||
|
||||
**TensorRef**: contains base pointer and _Layout_ object for referencing infinitely-sized tensor object
|
||||
|
||||
**TensorView**: contains _TensorRef_ and extent of a finite mathematical object
|
||||
|
||||
**Tile**: partitions of a tensor that have constant extents and layout known at compile time
|
||||
|
||||
**Trait**: characteristics of a fully-specialized type, typically used in metaprogramming reflection
|
||||
|
||||
**View**: an object containing references to a data structure that it does not own; typically, construction of views is lightweight
|
||||
|
||||
**Warp**: a collection of hardware threads executing in lock-step; warp-level operations typically rely on cooperation among the threads within the warp
|
||||
|
||||
`AlignedBuffer<T, N>`: statically sized array type; union-safe, no construction guarantee for elements
|
||||
|
||||
`Array<T, N>`: container for holding numeric types - handles bit packing for small numeric types (e.g. int4_t, uint4_t, bin1_t)
|
||||
`sizeof(Array<T, N>)` - gives expected value in units of bytes with minimum storage of `1 B`: (sizeof_bits<T>::value * N) / 8
|
||||
|
||||
**Operator**: an object performing a computation on matrix or tensor objects. May be further refined by scope within the execution model hierarchy. Deprecated starting CUTLASS 3.0,
|
||||
replaced by [MMA and Copy atoms from CuTe](cute/0t_mma_atom.md).
|
||||
|
||||
**Tile Iterator**: abstraction for accessing and traversing a sequence of tiles in a tensor; CUTLASS specifies
|
||||
[formal concepts for tile iterators](tile_iterator_concept.md). Deprecated starting CUTLASS 3.0.
|
||||
Replaced by `cute::Layout` in equivalent usage scenarios to represent data tensors.
|
||||
|
||||
**Thread Map**: abstraction for defining how threads are mapped to a given tile. Deprecated starting CUTLASS 3.0.
|
||||
Replaced by `cute::Layout` in equivalent usage scenarios to represent thread tensors.
|
||||
|
||||
# Copyright
|
||||
|
||||
Copyright (c) 2017 - 2025 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
```
|
||||
Redistribution and use in source and binary forms, with or without
|
||||
modification, are permitted provided that the following conditions are met:
|
||||
|
||||
1. Redistributions of source code must retain the above copyright notice, this
|
||||
list of conditions and the following disclaimer.
|
||||
|
||||
2. Redistributions in binary form must reproduce the above copyright notice,
|
||||
this list of conditions and the following disclaimer in the documentation
|
||||
and/or other materials provided with the distribution.
|
||||
|
||||
3. Neither the name of the copyright holder nor the names of its
|
||||
contributors may be used to endorse or promote products derived from
|
||||
this software without specific prior written permission.
|
||||
|
||||
THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS "AS IS"
|
||||
AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE
|
||||
IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE ARE
|
||||
DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT HOLDER OR CONTRIBUTORS BE LIABLE
|
||||
FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL
|
||||
DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR
|
||||
SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER
|
||||
CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY,
|
||||
OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE
|
||||
OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
|
||||
```
|
||||
502
media/docs/cpp/tile_iterator_concept.md
Normal file
502
media/docs/cpp/tile_iterator_concept.md
Normal file
@@ -0,0 +1,502 @@
|
||||

|
||||
|
||||
# Tile Iterator Concepts
|
||||
|
||||
Note: CUTLASS 3.0 deprecates all tile access iterators in favour of CuTe's single
|
||||
vocabulary type `cute::Tensor`, which is parameterized on `cute::Layout`.
|
||||
`cute::Tensor`s can therefore be manipulated with the same layout algebra as all CuTe layouts.
|
||||
This removes the need for bespoke types that encapsulate iterator properties.
|
||||
The following text thus only applies to legacy CUTLASS 2.x API and related types.
|
||||
|
||||
CUTLASS 2.x implements generic algorithms on tiles of matrix or tensors of constant size. These may
|
||||
be considered as partitions of tensors of infinite size, with a range of partitions accessible
|
||||
by _tile iterators_.
|
||||
|
||||
Various data structures may make operations such as random access to tiles inexpensive,
|
||||
while data structures may not offer random access at all. For example, iterating over a linked
|
||||
list of matrices requires sequential traversal. Algorithms implemented in terms of sequences of tiles
|
||||
should require only the minimum set of operators be defined for tile iterators.
|
||||
|
||||
This document describes a set of C++ concepts which may be used to define tile iterators used
|
||||
by CUTLASS algorithms. ("Concept" here does not refer to a C++20 concept that uses the `concept` keyword.
|
||||
Rather, it refers to a set of requirements on a type.)
|
||||
Each concept specifies members and type definitions that a tile iterator
|
||||
must implement. Frequently, a tile iterator implements several concepts, and its members are
|
||||
the union of the members from each individual concept. These definitions were inspired by
|
||||
[Boost "New style" iterator concepts](https://www.boost.org/doc/libs/1_40_0/libs/iterator/doc/new-iter-concepts.html).
|
||||
|
||||
The set of all possible combinations of these concepts is quite large, however most tile iterator
|
||||
templates can be described by one of several combinations. The section
|
||||
Frequently Used Tile Iterator Concepts describes several common interfaces used throughout CUTLASS.
|
||||
|
||||
## Definitions
|
||||
|
||||
**_Base Tile Iterator Concept_.** All tile iterators must describe an _Element_ type as well as a _Shape_.
|
||||
```c++
|
||||
/// Base concept for all tile iterators
|
||||
struct TileIteratorConcept {
|
||||
using Element; ///< Element type composing tile (concept: numeric type or Array<>)
|
||||
using Shape; ///< Shape type describing extent of tile. The shape concept depends
|
||||
/// on iterator implementation.
|
||||
};
|
||||
```
|
||||
|
||||
**_Contiguous Memory Tile Iterator Concept_.** Iterators over tiles stored arbitrarily within
|
||||
a continuous block of data in memory. Linear offset in units of _Element_ may be added to
|
||||
internally held pointers to 'move' the iterator in memory.
|
||||
|
||||
```c++
|
||||
/// Tile iterator over partitions of a tensor in contiguous memory which may be referenced via a
|
||||
/// TensorRef object.
|
||||
struct ContiguousMemoryTileIterator : public TileIteratorConcept {
|
||||
|
||||
using Index; ///< index type used to add pointer offsets
|
||||
|
||||
/// Adds a linear offset in units of Element to internal pointer(s) into tensor
|
||||
CUTLASS_DEVICE
|
||||
void add_pointer_offset(Index pointer_offset);
|
||||
};
|
||||
```
|
||||
|
||||
**_Readable Tile Iterator Concept_.** Iterators that may be read from define a `Fragment` type holding
|
||||
each thread's part of the data to be loaded. An explicit `load()` method reads the tile from memory,
|
||||
and places each thread's part in its `Fragment` object.
|
||||
|
||||
```c++
|
||||
/// Tile iterator capable of loading tiles from memory into fragments
|
||||
struct ReadableTileIteratorConcept {
|
||||
|
||||
using Fragment; ///< fragment object derived from cutlass::Array<Element, N>
|
||||
|
||||
CUTLASS_DEVICE
|
||||
void load(Fragment &frag); ///< loads a fragment from memory
|
||||
};
|
||||
```
|
||||
|
||||
**_Readable Contiguous Tile Iterator Concept_.** Iterators reading from contiguous memory
|
||||
support an optional pointer offset that is added to any internally managed pointers before
|
||||
performing the load. This provides a convenient method to fold an offset in with load
|
||||
operations.
|
||||
|
||||
```c++
|
||||
/// Union of the following tile iterator concepts:
|
||||
///
|
||||
/// - ReadableTileIteratorConcept
|
||||
/// - ContiguousMemoryTileIterator
|
||||
///
|
||||
struct ReadableContiguousTileIteratorConcept :
|
||||
public ReadableTileIteratorConcept,
|
||||
public ContiguousMemoryTileIterator {
|
||||
|
||||
/// Loads a fragment from memory with additional logical offset
|
||||
CUTLASS_DEVICE
|
||||
void load_with_pointer_offset(
|
||||
Fragment &frag, ///< fragment to load from the tensor
|
||||
Index pointer_offset); ///< loads a tile with a linear offset
|
||||
};
|
||||
```
|
||||
|
||||
**_Writeable Tile Iterator Concept_.** Iterators that may write to memory define a `Fragment` type holding
|
||||
each thread's part of the data to be written. An explicit `store()` method writes the tile to memory.
|
||||
|
||||
```c++
|
||||
/// Tile iterator capable of storing tiles from memory
|
||||
struct WriteableTileIteratorConcept {
|
||||
|
||||
using Fragment; ///< fragment object derived from cutlass::Array<Element, N>
|
||||
|
||||
/// Stores a fragment to memory
|
||||
CUTLASS_DEVICE
|
||||
void store(Fragment const &frag); ///< stores a fragment to memory
|
||||
};
|
||||
```
|
||||
|
||||
**_Writeable Contiguous Tile Iterator Concept_.** Iterators writing to contiguous memory
|
||||
support an optional pointer offset that is added to any internally managed pointers before
|
||||
performing the store operation. This provides a convenient method to fold an offset into the
|
||||
store.
|
||||
```c++
|
||||
/// Union of the following tile iterator concepts:
|
||||
///
|
||||
/// - WriteableTileIteratorConcept
|
||||
/// - ContiguousMemoryTileIterator
|
||||
///
|
||||
struct WriteableContiguousTileIteratorConcept :
|
||||
public WriteableTileIteratorConcept,
|
||||
public ContiguousMemoryTileIterator {
|
||||
|
||||
/// Loads a fragment from memory with additional logical offset
|
||||
CUTLASS_DEVICE
|
||||
void store_with_pointer_offset(
|
||||
Fragment const &frag, ///< fragment to store to the tensor
|
||||
Index pointer_offset); ///< stores a tile with a linear offset
|
||||
};
|
||||
```
|
||||
|
||||
**_Forward Tile Iterator Concept_.** This concept offers traversal "forward" by one tile in
|
||||
a pre-defined sequence. Often, this sequence is relevant to the context in which the iterator
|
||||
was defined, such as along the _K_ dimension of a GEMM operation. Equality operators are defined
|
||||
to determine whether two iterators point to the same tile.
|
||||
```c++
|
||||
/// Tile iterator that may be incremented along a traversal sequence.
|
||||
struct ForwardTileIteratorConcept {
|
||||
|
||||
CUTLASS_DEVICE bool operator==(TileIterator const &it); ///< true if iterators point to same tile, false if otherwise
|
||||
CUTLASS_DEVICE bool operator!=(TileIterator const &it); ///< false if iterators point to same tile, true if otherwise
|
||||
|
||||
CUTLASS_DEVICE ForwardTileIteratorConcept & operator++(); ///< pre-increment - advance to next tile in sequence
|
||||
CUTLASS_DEVICE ForwardTileIteratorConcept operator++(int); ///< post-increment - advance to next tile in sequence
|
||||
};
|
||||
```
|
||||
|
||||
**_Bidirectional Tile Iterator Concept_.** This concept permits traversal both forward and backward.
|
||||
```c++
|
||||
/// Tile iterator which may be traverse in both directions along a defined sequence.
|
||||
struct BidirectionalTileIteratorConcept : public ForwardTileIteratorConcept {
|
||||
|
||||
CUTLASS_DEVICE
|
||||
BidirectionalTileIteratorConcept & operator--(); ///< pre-decrement - traverse to previous tile in sequence
|
||||
|
||||
CUTLASS_DEVICE
|
||||
BidirectionalTileIteratorConcept operator--(int); ///< post-decrement - traverse to previous tile in sequence
|
||||
};
|
||||
```
|
||||
|
||||
**_Random Access Tile Iterator Concept_.** This iterator defines random access operations in the logical
|
||||
coordinate system of the underlying tensor. Thus, tensors must have a defined _Layout_ with associated
|
||||
_TensorCoord_ coordinate describing logical position within the tensor and _TensorRef_ reference type.
|
||||
It may be advanced forward or backwards by an offset specified as units of whole tiles along each dimension.
|
||||
```c++
|
||||
/// Tile iterator offering random access to tiles in contiguous memory.
|
||||
struct RandomAccessTileIteratorConcept :
|
||||
public BidirectionalTileIteratorConcept,
|
||||
public ContiguousMemoryTileIterator {
|
||||
|
||||
using Layout; ///< Layout object mapping
|
||||
using TensorRef; ///< Tensor Reference object
|
||||
using TensorCoord; ///< Logical coordinate in referenced tensor
|
||||
|
||||
///< advances in units of whole tiles along the logical coordinate space of the tensor
|
||||
CUTLASS_DEVICE
|
||||
RandomAccessTileIteratorConcept & add_tile_offset(TensorCoord const &tile_offset);
|
||||
|
||||
///< advances in units of whole tiles along the logical coordinate space of the tensor
|
||||
CUTLASS_DEVICE
|
||||
RandomAccessTileIteratorConcept & operator+=(TensorCoord const &tile_offset);
|
||||
|
||||
///< advances in units of whole tiles along the logical coordinate space of the tensor
|
||||
CUTLASS_DEVICE
|
||||
RandomAccessTileIteratorConcept & operator-=(TensorCoord const &tile_offset);
|
||||
};
|
||||
```
|
||||
|
||||
**_Readable Random Access Tile Iterator Concept_.** Readable random access iterators
|
||||
accept an additional tile offset in logical coordinate space when loading fragments.
|
||||
```c++
|
||||
/// Loads a fragment with a logical coordinate offset in units of whole tiles.
|
||||
struct ReadableRandomAccessTileIteratorConcept :
|
||||
public RandomAccessTileIteratorConcept,
|
||||
public ReadableTileIteratorConcept {
|
||||
|
||||
/// Loads a fragment from memory with logical offset in units of whole tiles.
|
||||
CUTLASS_DEVICE
|
||||
void load(
|
||||
Fragment &frag, ///< fragment to load from the tensor
|
||||
TensorCoord const &tile_offset); ///< loads a tile with a logical offset in units of whole tiles
|
||||
};
|
||||
```
|
||||
|
||||
**_Readable Random Access Contiguous Tile Iterator Concept_.** Readable random access iterators
|
||||
accept an additional tile offset in logical coordinate space when loading fragments.
|
||||
```c++
|
||||
/// Loads a fragment with a logical coordinate offset in units of whole tiles.
|
||||
struct ReadableRandomAccessContiguousTileIteratorConcept :
|
||||
public ReadableRandomAccessTileIteratorConcept,
|
||||
ReadableContiguousTileIteratorConcept {
|
||||
|
||||
/// Loads a fragment from memory with logical offset in units of whole tiles.
|
||||
CUTLASS_DEVICE
|
||||
void load(
|
||||
Fragment &frag, ///< fragment to load from the tensor
|
||||
TensorCoord const &tile_offset, ///< loads a tile with a logical offset in units of whole tiles
|
||||
Index pointer_offset); ///< loads a tile with a logical offset AND a pointer offset
|
||||
};
|
||||
```
|
||||
**_Writeable Random Access Tile Iterator Concept_.** Writeable random access iterators
|
||||
accept an additional tile offset in logical coordinate space when storing fragments.
|
||||
```c++
|
||||
/// Stores a fragment with a logical coordinate offset in units of whole tiles.
|
||||
struct WriteableRandomAccessTileIteratorConcept :
|
||||
public RandomAccessTileIteratorConcept,
|
||||
public WriteableContiguousTileIteratorConcept {
|
||||
|
||||
/// Stores a fragment from memory with logical offset in units of whole tiles.
|
||||
CUTLASS_DEVICE
|
||||
void store(
|
||||
Fragment const &frag, ///< fragment to store to the location pointed to by the tensor
|
||||
TensorCoord const &tile_offset); ///< stores a tile with a given offset from the current iterator
|
||||
};
|
||||
```
|
||||
|
||||
**_Writeable Random Access Contiguous Tile Iterator Concept_.** Writeable random access iterators
|
||||
accept an additional tile offset in logical coordinate space when storing fragments.
|
||||
```c++
|
||||
/// Stores a fragment with a logical coordinate offset in units of whole tiles.
|
||||
struct WriteableRandomAccessContiguousTileIteratorConcept :
|
||||
public WriteableRandomAccessTileIteratorConcept,
|
||||
public WriteableContiguousTileIteratorConcept {
|
||||
|
||||
/// Stores a fragment from memory with logical offset in units of whole tiles.
|
||||
CUTLASS_DEVICE
|
||||
void store(
|
||||
Fragment const &frag, ///< fragment to store to the location pointed to by the tensor
|
||||
TensorCoord const &tile_offset, ///< stores a tile with a logical offset in units of whole tiles
|
||||
Index pointer_offset); ///< stores a tile witha logical offset AND a pointer offset
|
||||
};
|
||||
```
|
||||
|
||||
**_Masked Tile Iterator Concept_.** Matrix and tensors may not always be multiples of whole tiles.
|
||||
Masked tile iterators define a `Mask` type which may be used to guard accesses to memory. The
|
||||
semantics and interface of this `Mask` are implementation-defined details of each tile iterator,
|
||||
but several convenience methods are defined for interacting with the mask such as efficiently
|
||||
clearing or enabling all guarded memory accesses.
|
||||
```c++
|
||||
/// Supports iterating over tiles that are not 'whole' in memory. Iterator maintains a mask object
|
||||
/// which guards against out-of-bounds access.
|
||||
///
|
||||
/// Note, this concept definition does not formally define operations on the mask or methods it
|
||||
/// supports. These remain implementation-dependent details of iterators implementing this concept.
|
||||
struct MaskedTileIteratorConcept {
|
||||
|
||||
using Mask; ///< mask object used to guard against acceses.
|
||||
|
||||
CUTLASS_DEVICE void clear_mask(); ///< efficiently disables all accesses guarded by mask
|
||||
CUTLASS_DEVICE void enable_mask(); ///< efficiently enables all accesses guarded by mask
|
||||
|
||||
CUTLASS_DEVICE void get_mask(Mask &mask); ///< gets the mask
|
||||
CUTLASS_DEVICE void set_mask(Mask const &mask); ///< sets the mask
|
||||
};
|
||||
```
|
||||
|
||||
## Frequently Used Tile Iterator Concepts
|
||||
|
||||
This section describes several frequently used compositions of the basic tile iterator concepts. They are
|
||||
listed here as complete type declarations for convenience of the reader.
|
||||
|
||||
**_Writeable, Readable, Forward, Contiguous Memory Tile Iterator Concept_.**
|
||||
This combines several of the basic iterator concepts to
|
||||
yield a tile iterator capable of loading and storing tiles as well as advancing forward along a traversal sequence.
|
||||
```c++
|
||||
/// This tile iterator embodies several of the above:
|
||||
///
|
||||
/// - ForwardTileIteratorConcept
|
||||
/// - ReadableContiguousTileIteratorConcept
|
||||
/// - WriteableContiguousTileIteratorConcept
|
||||
///
|
||||
/// It is restated explicitly for convenience of the reader.
|
||||
///
|
||||
struct WriteableReadableForwardContiguousTileIteratorConcept {
|
||||
|
||||
//
|
||||
// Data types
|
||||
//
|
||||
|
||||
using Element; ///< Element type composing tile.
|
||||
using Shape; ///< Shape type describing extent of tile. The shape concept depends
|
||||
/// on iterator implementation
|
||||
using Index; ///< index type used as base for TensorCoord
|
||||
using Fragment; ///< fragment object derived from cutlass::Array<Element, N>
|
||||
|
||||
//
|
||||
// Methods
|
||||
//
|
||||
|
||||
/// Adds a linear offset in units of Element to internal pointer(s) into tensor
|
||||
CUTLASS_DEVICE
|
||||
void add_pointer_offset(Index offset);
|
||||
|
||||
/// true if iterators point to same tile, false if otherwise
|
||||
CUTLASS_DEVICE bool operator==(WriteableReadableForwardContiguousTileIteratorConcept const &it);
|
||||
|
||||
///< false if iterators point to same tile, true if otherwise
|
||||
CUTLASS_DEVICE bool operator!=(WriteableReadableForwardContiguousTileIteratorConcept const &it);
|
||||
|
||||
/// pre-increment - traverse to next tile in sequence
|
||||
CUTLASS_DEVICE
|
||||
WriteableReadableForwardContiguousTileIteratorConcept &
|
||||
operator++();
|
||||
|
||||
///< post-increment - traverse to next tile in sequence
|
||||
CUTLASS_DEVICE
|
||||
WriteableReadableForwardContiguousTileIteratorConcept
|
||||
operator++(int);
|
||||
|
||||
/// Loads a fragment from memory
|
||||
CUTLASS_DEVICE
|
||||
void load(Fragment &frag); ///< fragment to be loaded from memory
|
||||
|
||||
/// Loads a fragment from memory with additional logical offset
|
||||
CUTLASS_DEVICE
|
||||
void load_with_pointer_offset(
|
||||
Fragment &frag, ///< fragment to be loaded from memory
|
||||
Index pointer_offset); ///< linear offset (in units of Element) when loading
|
||||
|
||||
/// Stores a fragment to memory
|
||||
CUTLASS_DEVICE
|
||||
void store(Fragment const &frag); ///< fragment to store to memory
|
||||
|
||||
/// Stores a fragment from memory with additional logical offset
|
||||
CUTLASS_DEVICE
|
||||
void store_with_pointer_offset(
|
||||
Fragment const &frag, ///< fragment to store to memory
|
||||
Index pointer_offset); ///< linear offset (in units of Element) when storing
|
||||
};
|
||||
```
|
||||
|
||||
**_Writeable, Readable, Random Access, Contiguous Memory Tile Iterator Concept_.**
|
||||
This combines several of the basic iterator concepts to
|
||||
yield a tile iterator with random access suitable for loading matrix operands for GEMM.
|
||||
```c++
|
||||
/// This tile iterator embodies several of the above:
|
||||
///
|
||||
/// - ReadableRandomAccessContiguousTileIteratorConcept
|
||||
/// - WriteableRandomAccessContiguousTileIteratorConcept
|
||||
///
|
||||
/// It is restated explicitly for convenience of the reader.
|
||||
///
|
||||
struct WriteableReadableRandomAccessContiguousTileIteratorConcept {
|
||||
|
||||
//
|
||||
// Data types
|
||||
//
|
||||
|
||||
using Element; ///< Element type composing tile.
|
||||
using Shape; ///< Shape type describing extent of tile. The shape concept depends
|
||||
/// on iterator implementation
|
||||
using Layout; ///< Layout object mapping
|
||||
using TensorRef; ///< Tensor Reference object
|
||||
using TensorCoord; ///< Logical coordinate in referenced tensor
|
||||
using Index; ///< index type used as base for TensorCoord
|
||||
using Fragment; ///< fragment object derived from cutlass::Array<Element, N>
|
||||
|
||||
//
|
||||
// Methods
|
||||
//
|
||||
|
||||
/// Adds a linear offset in units of Element to internal pointer(s) into tensor
|
||||
CUTLASS_DEVICE
|
||||
void add_pointer_offset(Index pointer_offset);
|
||||
|
||||
/// true if iterators point to same tile, false if otherwise
|
||||
CUTLASS_DEVICE bool operator==(WriteableReadableRandomAccessContiguousTileIteratorConcept const &it);
|
||||
|
||||
///< false if iterators point to same tile, true if otherwise
|
||||
CUTLASS_DEVICE bool operator!=(WriteableReadableRandomAccessContiguousTileIteratorConcept const &it);
|
||||
|
||||
/// pre-increment - traverse to next tile in sequence
|
||||
CUTLASS_DEVICE
|
||||
WriteableReadableRandomAccessContiguousTileIteratorConcept &
|
||||
operator++();
|
||||
|
||||
///< post-increment - traverse to next tile in sequence
|
||||
CUTLASS_DEVICE
|
||||
WriteableReadableRandomAccessContiguousTileIteratorConcept
|
||||
operator++(int);
|
||||
|
||||
/// pre-decrement - traverse to previous tile in sequence
|
||||
CUTLASS_DEVICE
|
||||
WriteableReadableRandomAccessContiguousTileIteratorConcept &
|
||||
operator--();
|
||||
|
||||
///< post-decrement - traverse to previous tile in sequence
|
||||
CUTLASS_DEVICE
|
||||
WriteableReadableRandomAccessContiguousTileIteratorConcept
|
||||
operator--(int);
|
||||
|
||||
///< advances in units of whole tiles along the logical coordinate space of the tensor
|
||||
CUTLASS_DEVICE
|
||||
WriteableReadableRandomAccessContiguousTileIteratorConcept & operator+=(TensorCoord const &tile_offset);
|
||||
|
||||
///< advances in units of whole tiles along the logical coordinate space of the tensor
|
||||
CUTLASS_DEVICE
|
||||
WriteableReadableRandomAccessContiguousTileIteratorConcept & operator-=(TensorCoord const &tile_offset);
|
||||
|
||||
/// Loads a fragment from memory
|
||||
CUTLASS_DEVICE
|
||||
void load(Fragment &frag); ///< fragment to be loaded from memory
|
||||
|
||||
/// Loads a fragment from memory with additional logical offset
|
||||
CUTLASS_DEVICE
|
||||
void load_with_pointer_offset(
|
||||
Fragment &frag, ///< fragment to be loaded from memory
|
||||
Index pointer_offset); ///< linear offset (in units of Element) when loading
|
||||
|
||||
/// Loads a fragment from memory with logical offset in units of whole tiles.
|
||||
CUTLASS_DEVICE
|
||||
void load(
|
||||
Fragment &frag, ///< fragment to be loaded from memory
|
||||
TensorCoord const &tile_offset); ///< loads a tile with a logical offset in units of whole tiles
|
||||
|
||||
/// Loads a fragment from memory with logical offset in units of whole tiles.
|
||||
CUTLASS_DEVICE
|
||||
void load(
|
||||
Fragment &frag, ///< fragment to be loaded from memory
|
||||
TensorCoord const &tile_offset, ///< loads a tile with a logical offset in units of whole tiles
|
||||
Index pointer_offset); ///< loads a tile with a logical offset AND a pointer offset
|
||||
|
||||
/// Stores a fragment to memory
|
||||
CUTLASS_DEVICE
|
||||
void store(Fragment const &frag); ///< fragment to store to memory
|
||||
|
||||
/// Loads a fragment from memory with additional logical offset
|
||||
CUTLASS_DEVICE
|
||||
void store_with_pointer_offset(
|
||||
Fragment const &frag, ///< fragment to store to memory
|
||||
Index pointer_offset); ///< linear offset (in units of Element) when loading
|
||||
|
||||
/// Stores a fragment from memory with logical offset in units of whole tiles.
|
||||
CUTLASS_DEVICE
|
||||
void store(
|
||||
Fragment const &frag, ///< fragment to store to memory
|
||||
TensorCoord const &tile_offset); ///< stores with logical offset in units of whole tiles
|
||||
|
||||
/// Stores a fragment from memory with logical offset in units of whole tiles.
|
||||
CUTLASS_DEVICE
|
||||
void store(
|
||||
Fragment const &frag, ///< fragment to store to memory
|
||||
TensorCoord const &tile_offset, ///< stores with logical offset in units of whole tiles
|
||||
Index pointer_offset);
|
||||
};
|
||||
```
|
||||
|
||||
# Copyright
|
||||
|
||||
Copyright (c) 2017 - 2025 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
```
|
||||
Redistribution and use in source and binary forms, with or without
|
||||
modification, are permitted provided that the following conditions are met:
|
||||
|
||||
1. Redistributions of source code must retain the above copyright notice, this
|
||||
list of conditions and the following disclaimer.
|
||||
|
||||
2. Redistributions in binary form must reproduce the above copyright notice,
|
||||
this list of conditions and the following disclaimer in the documentation
|
||||
and/or other materials provided with the distribution.
|
||||
|
||||
3. Neither the name of the copyright holder nor the names of its
|
||||
contributors may be used to endorse or promote products derived from
|
||||
this software without specific prior written permission.
|
||||
|
||||
THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS "AS IS"
|
||||
AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE
|
||||
IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE ARE
|
||||
DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT HOLDER OR CONTRIBUTORS BE LIABLE
|
||||
FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL
|
||||
DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR
|
||||
SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER
|
||||
CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY,
|
||||
OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE
|
||||
OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
|
||||
```
|
||||
464
media/docs/cpp/utilities.md
Normal file
464
media/docs/cpp/utilities.md
Normal file
@@ -0,0 +1,464 @@
|
||||

|
||||
|
||||
|
||||
Note: This document discusses utilities commonly used with code that targets CUTLASS 2.x.
|
||||
Although CUTLASS 3.0's primary entry point APIs do not transact in these `cutlass::*` tensor types anymore,
|
||||
users can still find them convenient for managing allocations with trivial affine layouts.
|
||||
For more advanced host side tensor management, [`cute::Tensor`](cute/03_tensor.md)s
|
||||
can be used on either host or device for any memory space and full expressive power of
|
||||
[`cute::Layout`](cute/01_layout.md)s.
|
||||
|
||||
# CUTLASS Utilities
|
||||
|
||||
CUTLASS utilities are additional template classes that facilitate recurring tasks. These are
|
||||
flexible implementations of needed functionality, but they are not expected to be efficient.
|
||||
|
||||
Applications should configure their builds to list `/tools/util/include` in their include
|
||||
paths.
|
||||
|
||||
Source code is in [`/tools/util/include/cutlass/util/`](https://github.com/NVIDIA/cutlass/tree/main/tools/util/include/cutlass/util).
|
||||
|
||||
## Tensor Allocation and I/O
|
||||
|
||||
To allocate a tensor with storage in both host and device memory, use `HostTensor` in
|
||||
[`cutlass/util/host_tensor.h`](https://github.com/NVIDIA/cutlass/tree/main/tools/util/include/cutlass/util/host_tensor.h)
|
||||
|
||||
```c++
|
||||
template <typename Element, typename Layout>
|
||||
class HostTensor;
|
||||
```
|
||||
|
||||
This class is compatible with all CUTLASS numeric data types and layouts.
|
||||
|
||||
**Example:** column-major matrix storage of single-precision elements.
|
||||
```c++
|
||||
#include <cutlass/layout/matrix.h>
|
||||
#include <cutlass/util/host_tensor.h>
|
||||
|
||||
int main() {
|
||||
int rows = 32;
|
||||
int columns = 16;
|
||||
|
||||
cutlass::HostTensor<float, cutlass::layout::ColumnMajor> tensor({rows, columns});
|
||||
|
||||
return 0;
|
||||
}
|
||||
```
|
||||
|
||||
Internal host-side storage may be accessed via the following methods.
|
||||
```c++
|
||||
float *host_ptr = tensor.host_data();
|
||||
cutlass::TensorRef<float, cutlass::layout::ColumnMajor> host_ref = tensor.host_ref();
|
||||
cutlass::TensorView<float, cutlass::layout::ColumnMajor> host_view = tensor.host_view();
|
||||
```
|
||||
|
||||
Device memory may be accessed similarly.
|
||||
```c++
|
||||
float *device_ptr = tensor.device_data();
|
||||
cutlass::TensorRef<float, cutlass::layout::ColumnMajor> device_ref = tensor.device_ref();
|
||||
cutlass::TensorView<float, cutlass::layout::ColumnMajor> device_view = tensor.device_view();
|
||||
```
|
||||
|
||||
Printing to human-readable CSV output is accoplished with `std::ostream::operator<<()` defined in
|
||||
[`cutlass/util/tensor_view_io.h`](https://github.com/NVIDIA/cutlass/tree/main/tools/util/include/cutlass/util/tensor_view_io.h).
|
||||
Note, this assumes all views refer to host memory.
|
||||
```c++
|
||||
#include <cutlass/util/tensor_view_io.h>
|
||||
|
||||
int main() {
|
||||
// Obtain a TensorView into host memory
|
||||
cutlass::TensorView<float, cutlass::layout::ColumnMajor> view = tensor.host_view();
|
||||
|
||||
// Print to std::cout
|
||||
std::cout << view << std::endl;
|
||||
|
||||
return 0;
|
||||
}
|
||||
```
|
||||
|
||||
Host and device memory must be explicitly synchronized by the application.
|
||||
```c++
|
||||
float idx = 0;
|
||||
|
||||
for (int i = 0; i < rows; ++i) {
|
||||
for (int j = 0; j < columns; ++j) {
|
||||
|
||||
// Write the element at location {i, j} in host memory
|
||||
tensor.host_ref().at({i, j}) = idx;
|
||||
|
||||
idx += 0.5f;
|
||||
}
|
||||
}
|
||||
|
||||
// Copy host memory to device memory
|
||||
tensor.sync_device();
|
||||
|
||||
// Obtain a device pointer usable in CUDA kernels
|
||||
float *device_ptr = tensor.device_data();
|
||||
```
|
||||
|
||||
`HostTensor<>` is usable by all CUTLASS layouts including interleaved layouts.
|
||||
```c++
|
||||
int rows = 4;
|
||||
int columns = 3;
|
||||
|
||||
cutlass::HostTensor<float, cutlass::layout::ColumnMajorInterleaved<4>> tensor({rows, columns});
|
||||
|
||||
for (int i = 0; i < rows; ++i) {
|
||||
for (int j = 0; j < columns; ++j) {
|
||||
|
||||
// Write the element at location {i, j} in host memory
|
||||
tensor.host_ref().at({i, j}) = float(i) * 1.5f - float(j) * 2.25f;
|
||||
}
|
||||
}
|
||||
|
||||
std::cout << tensor.host_view() << std::endl;
|
||||
```
|
||||
|
||||
## Device Allocations
|
||||
|
||||
To strictly allocate memory on the device using the smart pointer pattern to manage allocation and deallocation,
|
||||
use `cutlass::DeviceAllocation<>`.
|
||||
|
||||
**Example:** allocating an array in device memory.
|
||||
```c++
|
||||
#include <cutlass/layout/matrix.h>
|
||||
#include <cutlass/layout/tensor_view.h>
|
||||
#include <cutlass/util/device_memory.h>
|
||||
|
||||
__global__ void kernel(float *device_ptr) {
|
||||
|
||||
}
|
||||
|
||||
int main() {
|
||||
|
||||
size_t N = 1024;
|
||||
|
||||
cutlass::DeviceAllocation<float> device_alloc(N);
|
||||
|
||||
// Call a CUDA kernel passing device memory as a pointer argument
|
||||
kernel<<< grid, block >>>(alloc.get());
|
||||
|
||||
if (cudaGetLastError() != cudaSuccess) {
|
||||
return -1;
|
||||
}
|
||||
|
||||
// Device memory is automatically freed when device_alloc goes out of scope
|
||||
|
||||
return 0;
|
||||
}
|
||||
```
|
||||
|
||||
## Tensor Initialization
|
||||
|
||||
CUTLASS defines several utility functions to initialize tensors to uniform, procedural,
|
||||
or randomly generated elements. These have implementations using strictly host code and
|
||||
implementations using strictly CUDA device code.
|
||||
|
||||
`TensorFill()` for uniform elements throughout a tensor.
|
||||
```c++
|
||||
#include <cutlass/layout/matrix.h>
|
||||
#include <cutlass/util/reference/host/tensor_fill.h>
|
||||
#include <cutlass/util/reference/device/tensor_fill.h>
|
||||
#include <cutlass/util/host_tensor.h>
|
||||
|
||||
int main() {
|
||||
int rows = 128;
|
||||
int columns = 64;
|
||||
|
||||
float x = 3.14159f;
|
||||
|
||||
cutlass::HostTensor<float, cutlass::layout::ColumnMajor> tensor({rows, columns});
|
||||
|
||||
// Initialize in host memory
|
||||
cutlass::reference::host::TensorFill(tensor.host_view(), x);
|
||||
|
||||
// Initialize in device memory
|
||||
cutlass::reference::device::TensorFill(tensor.device_view(), x);
|
||||
|
||||
return 0;
|
||||
}
|
||||
```
|
||||
|
||||
`TensorFillRandomUniform()` for initializing elements to a random uniform distribution.
|
||||
The device-side implementation uses CURAND to generate random numbers.
|
||||
```c++
|
||||
#include <cutlass/layout/matrix.h>
|
||||
#include <cutlass/util/reference/host/tensor_fill.h>
|
||||
#include <cutlass/util/reference/device/tensor_fill.h>
|
||||
#include <cutlass/util/host_tensor.h>
|
||||
|
||||
int main() {
|
||||
int rows = 128;
|
||||
int columns = 64;
|
||||
|
||||
double maximum = 4;
|
||||
double minimum = -4;
|
||||
uint64_t seed = 0x2019;
|
||||
|
||||
cutlass::HostTensor<float, cutlass::layout::ColumnMajor> tensor({rows, columns});
|
||||
|
||||
// Initialize in host memory
|
||||
cutlass::reference::host::TensorFillRandomUniform(
|
||||
tensor.host_view(),
|
||||
seed,
|
||||
maximum,
|
||||
minimum);
|
||||
|
||||
// Initialize in device memory
|
||||
cutlass::reference::device::TensorFillRandomUniform(
|
||||
tensor.device_view(),
|
||||
seed,
|
||||
maximum,
|
||||
minimum);
|
||||
|
||||
return 0;
|
||||
}
|
||||
```
|
||||
|
||||
|
||||
`TensorFillRandomGaussian()` for initializing elements to a random gaussian distribution.
|
||||
The device-side implementation uses CURAND to generate random numbers.
|
||||
```c++
|
||||
#include <cutlass/layout/matrix.h>
|
||||
#include <cutlass/util/reference/host/tensor_fill.h>
|
||||
#include <cutlass/util/reference/device/tensor_fill.h>
|
||||
#include <cutlass/util/host_tensor.h>
|
||||
|
||||
int main() {
|
||||
|
||||
int rows = 128;
|
||||
int columns = 64;
|
||||
|
||||
double mean = 0.5;
|
||||
double stddev = 2.0;
|
||||
uint64_t seed = 0x2019;
|
||||
|
||||
cutlass::HostTensor<float, cutlass::layout::ColumnMajor> tensor({rows, columns});
|
||||
|
||||
// Initialize in host memory
|
||||
cutlass::reference::host::TensorFillRandomGaussian(
|
||||
tensor.host_view(),
|
||||
seed,
|
||||
mean,
|
||||
stddev);
|
||||
|
||||
// Initialize in device memory
|
||||
cutlass::reference::device::TensorFillRandomGaussian(
|
||||
tensor.device_view(),
|
||||
seed,
|
||||
mean,
|
||||
stddev);
|
||||
|
||||
return 0;
|
||||
}
|
||||
```
|
||||
|
||||
Each of these functions accepts an additional argument to specify how many bits of
|
||||
the mantissa less than 1 are non-zero. This simplifies functional comparisons when
|
||||
exact random distributions are not necessary, since elements may be restricted to
|
||||
integers or values with exact fixed-point representations.
|
||||
|
||||
```c++
|
||||
#include <cutlass/layout/matrix.h>
|
||||
#include <cutlass/util/reference/host/tensor_fill.h>
|
||||
#include <cutlass/util/reference/device/tensor_fill.h>
|
||||
#include <cutlass/util/host_tensor.h>
|
||||
|
||||
int main() {
|
||||
|
||||
int rows = 128;
|
||||
int columns = 64;
|
||||
|
||||
double mean = 0.5;
|
||||
double stddev = 2.0;
|
||||
uint64_t seed = 0x2019;
|
||||
|
||||
int bits_right_of_binary_decimal = 2;
|
||||
|
||||
cutlass::HostTensor<float, cutlass::layout::ColumnMajor> tensor({rows, columns});
|
||||
|
||||
// Initialize in host memory
|
||||
cutlass::reference::host::TensorFillRandomGaussian(
|
||||
tensor.host_view(),
|
||||
seed,
|
||||
mean,
|
||||
stddev,
|
||||
bits_right_of_binary_decimal);
|
||||
|
||||
// Initialize in device memory
|
||||
cutlass::reference::device::TensorFillRandomGaussian(
|
||||
tensor.device_view(),
|
||||
seed,
|
||||
mean,
|
||||
stddev,
|
||||
bits_right_of_binary_decimal);
|
||||
|
||||
return 0;
|
||||
}
|
||||
```
|
||||
|
||||
These utilities may be used for all data types.
|
||||
|
||||
**Example:** random half-precision tensor with Gaussian distribution.
|
||||
```c++
|
||||
#include <cutlass/numeric_types.h>
|
||||
#include <cutlass/layout/matrix.h>
|
||||
#include <cutlass/util/reference/host/tensor_fill.h>
|
||||
#include <cutlass/util/reference/device/tensor_fill.h>
|
||||
#include <cutlass/util/host_tensor.h>
|
||||
|
||||
int main() {
|
||||
int rows = 128;
|
||||
int columns = 64;
|
||||
|
||||
double mean = 0.5;
|
||||
double stddev = 2.0;
|
||||
uint64_t seed = 0x2019;
|
||||
|
||||
// Allocate a column-major tensor with half-precision elements
|
||||
cutlass::HostTensor<cutlass::half_t, cutlass::layout::ColumnMajor> tensor({rows, columns});
|
||||
|
||||
// Initialize in host memory
|
||||
cutlass::reference::host::TensorFillRandomGaussian(
|
||||
tensor.host_view(),
|
||||
seed,
|
||||
mean,
|
||||
stddev);
|
||||
|
||||
// Initialize in device memory
|
||||
cutlass::reference::device::TensorFillRandomGaussian(
|
||||
tensor.device_view(),
|
||||
seed,
|
||||
mean,
|
||||
stddev);
|
||||
|
||||
return 0;
|
||||
}
|
||||
```
|
||||
|
||||
## Reference Implementations
|
||||
|
||||
CUTLASS defines reference implementations usable with all data types and layouts. These are
|
||||
used throughout the unit tests.
|
||||
|
||||
**Example:** Reference GEMM implementation with mixed precision internal computation.
|
||||
```c++
|
||||
#include <cutlass/numeric_types.h>
|
||||
#include <cutlass/layout/matrix.h>
|
||||
|
||||
#include <cutlass/util/host_tensor.h>
|
||||
#include <cutlass/util/reference/host/gemm.h>
|
||||
|
||||
int main() {
|
||||
|
||||
int M = 64;
|
||||
int N = 32;
|
||||
int K = 16;
|
||||
|
||||
float alpha = 1.5f;
|
||||
float beta = -1.25f;
|
||||
|
||||
cutlass::HostTensor<cutlass::half_t, cutlass::layout::ColumnMajor> A({M, K});
|
||||
cutlass::HostTensor<cutlass::half_t, cutlass::layout::ColumnMajor> B({K, N});
|
||||
cutlass::HostTensor<cutlass::half_t, cutlass::layout::ColumnMajor> C({M, N});
|
||||
|
||||
cutlass::reference::host::Gemm<
|
||||
cutlass::half_t, cutlass::layout::ColumnMajor, // ElementA and LayoutA
|
||||
cutlass::half_t, cutlass::layout::ColumnMajor, // ElementB and LayoutB
|
||||
cutlass::half_t, cutlass::layout::ColumnMajor, // ElementC and LayoutC
|
||||
float, // scalar type (alpha and beta)
|
||||
float> gemm_op; // internal accumulation type
|
||||
|
||||
gemm_op(
|
||||
{M, N, K}, // problem size
|
||||
alpha, // alpha scalar
|
||||
A.host_view(), // TensorView to host memory
|
||||
B.host_view(), // TensorView to host memory
|
||||
beta, // beta scalar
|
||||
C.host_view(), // TensorView to host memory
|
||||
D.host_view()); // TensorView to device memory
|
||||
|
||||
return 0;
|
||||
}
|
||||
```
|
||||
|
||||
## Debugging Asynchronous Kernels with CUTLASS's Built-in `synclog` Tool
|
||||
|
||||
CUTLASS provides a built-in tool called `synclog` that enables printing runtime information useful for debugging asynchronous CUTLASS kernels. With the introduction of Warp Specialization in CUTLASS 3.0 for Hopper GPUs, kernel designs now incorporate synchronization among warps. The `synclog` tool simplifies debugging efforts for these asynchronous programs by recording and displaying timing information for synchronization events.
|
||||
|
||||
### Enabling `synclog`
|
||||
To enable `synclog`, add the -DCUTLASS_ENABLE_SYNCLOG=1 flag during compilation. From the CUTLASS root directory:
|
||||
|
||||
```
|
||||
$ mkdir build && cd build &&
|
||||
$ cmake .. -DCUTLASS_NVCC_ARCHS=90a -DCUTLASS_ENABLE_SYNCLOG=1
|
||||
```
|
||||
|
||||
### Building and Running with `synclog`
|
||||
After enabling `synclog`, build your CUTLASS example. For instance, to build example 54:
|
||||
|
||||
```
|
||||
$ cd examples/54_hopper_fp8_warp_specialized_gemm
|
||||
$ make
|
||||
```
|
||||
|
||||
Run the example, setting the profiling iteration count to 0 to ensure `synclog` information is printed only for the reference run:
|
||||
|
||||
```
|
||||
$ ./54_hopper_fp8_warp_specialized_gemm --iterations=0 &> synclog.txt
|
||||
```
|
||||
|
||||
### Interpreting `synclog` output
|
||||
The synclog.txt file will contain runtime information about synchronization events. Here's a sample output snippet:
|
||||
|
||||
```
|
||||
synclog start
|
||||
synclog at 1: cluster_barrier_init line=281 time=1725400116233388736 thread=0,0,0 block=0,0,0 smem_addr=197632 arrive_count=1
|
||||
synclog at 13: fence_barrier_init line=583 time=1725400116233388768 thread=32,0,0 block=0,0,0
|
||||
...
|
||||
```
|
||||
|
||||
Each line in the main body follows this format:
|
||||
```
|
||||
synclog at [synclog_at]: [header] line=[line] thread=[threadIdx.xyz] block=[blockIdx.xyz]
|
||||
```
|
||||
* `synclog at`: Address in the `synclog` output buffer (in bytes). Output exceeding 2^26 bytes is discarded.
|
||||
* `header`: Name of the synchronization event.
|
||||
* `line`: Code line number of the synchronization operation calling into `synclog`.
|
||||
|
||||
Additional information may appear at the end of each line, such as shared memory address, phase bit, and arrive count. For more detailed information on `synclog` output, refer to [synclog.hpp](https://github.com/NVIDIA/cutlass/tree/main/include/cutlass/arch/synclog.hpp) in the CUTLASS source code.
|
||||
|
||||
Please note that `synclog` is an experimental feature, and its functionality is not always guaranteed. We encourage its use in custom kernels and CUTLASS examples, though it is known to be incompatible with profiler kernels.
|
||||
|
||||
# Copyright
|
||||
|
||||
Copyright (c) 2017 - 2025 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
SPDX-License-Identifier: BSD-3-Clause
|
||||
|
||||
```
|
||||
Redistribution and use in source and binary forms, with or without
|
||||
modification, are permitted provided that the following conditions are met:
|
||||
|
||||
1. Redistributions of source code must retain the above copyright notice, this
|
||||
list of conditions and the following disclaimer.
|
||||
|
||||
2. Redistributions in binary form must reproduce the above copyright notice,
|
||||
this list of conditions and the following disclaimer in the documentation
|
||||
and/or other materials provided with the distribution.
|
||||
|
||||
3. Neither the name of the copyright holder nor the names of its
|
||||
contributors may be used to endorse or promote products derived from
|
||||
this software without specific prior written permission.
|
||||
|
||||
THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS "AS IS"
|
||||
AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE
|
||||
IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE ARE
|
||||
DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT HOLDER OR CONTRIBUTORS BE LIABLE
|
||||
FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL
|
||||
DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR
|
||||
SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER
|
||||
CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY,
|
||||
OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE
|
||||
OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
|
||||
```
|
||||
Reference in New Issue
Block a user