v4.3 update. (#2709)
* v4.3 update. * Update the cute_dsl_api changelog's doc link * Update version to 4.3.0 * Update the example link * Update doc to encourage user to install DSL from requirements.txt --------- Co-authored-by: Larry Wu <larwu@nvidia.com>
This commit is contained in:
@@ -15,6 +15,14 @@ Understanding these limitations will help you avoid potential pitfalls from the
|
||||
Please refer to :doc:`../limitations` for more details.
|
||||
|
||||
|
||||
Source Code Correlation
|
||||
-----------------------
|
||||
|
||||
CuTe DSL provides Python code to PTX/SASS correlation to enable the profiling/debugging of generated kernels with debug symbols by generating line info when compiling the kernel.
|
||||
|
||||
You can enable that globally via the environment variable CUTE_DSL_LINEINFO=1. Alternative, you can use compilation options to enable that per kernel. Please refer to :doc:`./dsl_jit_compilation_options` for more details.
|
||||
|
||||
|
||||
DSL Debugging
|
||||
-------------
|
||||
|
||||
@@ -75,6 +83,48 @@ This helps you verify whether the IR is generated as expected.
|
||||
export CUTE_DSL_KEEP_IR=1
|
||||
|
||||
|
||||
Dump the generated PTX & CUBIN
|
||||
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||||
|
||||
For users familiar with PTX and SASS, CuTe DSL supports dumping the generated PTX and CUBIN.
|
||||
|
||||
.. code:: bash
|
||||
|
||||
# Dump generated PTX in a .ptx file (default: False)
|
||||
export CUTE_DSL_KEEP_PTX=1
|
||||
|
||||
# Dump generated cubin in a .cubin file (default: False)
|
||||
export CUTE_DSL_KEEP_CUBIN=1
|
||||
|
||||
To further get SASS from cubin, users can use ``nvdisasm`` (usually installed with CUDA toolkit) to disassemble the cubin.
|
||||
|
||||
.. code:: bash
|
||||
|
||||
nvdisasm your_dsl_code.cubin > your_dsl_code.sass
|
||||
|
||||
|
||||
Access the dumped contents programmatically
|
||||
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||||
|
||||
For compiled kernels, the generated PTX/CUBIN/IR can be accessed programmatically as well through following attributes:
|
||||
|
||||
- ``__ptx__``: The generated PTX code of the compiled kernel.
|
||||
- ``__cubin__``: The generated CUBIN data of the compiled kernel.
|
||||
- ``__mlir__``: The generated IR code of the compiled kernel.
|
||||
|
||||
.. code:: python
|
||||
|
||||
compiled_foo = cute.compile(foo, ...)
|
||||
print(f"PTX: {compiled_foo.__ptx__}")
|
||||
with open("foo.cubin", "wb") as f:
|
||||
f.write(compiled_foo.__cubin__)
|
||||
|
||||
|
||||
Change the dump directory
|
||||
~~~~~~~~~~~~~~~~~~~~~~~~~
|
||||
|
||||
By default, all dumped files are saved in the current working directory. To specify a different directory for the dumped files, please set the environment variable CUTE_DSL_DUMP_DIR accordingly.
|
||||
|
||||
|
||||
Kernel Functional Debugging
|
||||
----------------------------
|
||||
@@ -122,6 +172,7 @@ For detecting memory errors and race conditions:
|
||||
|
||||
Please refer to the `compute-sanitizer documentation <https://developer.nvidia.com/compute-sanitizer>`_ for more details.
|
||||
|
||||
|
||||
Conclusion
|
||||
----------
|
||||
|
||||
|
||||
@@ -124,7 +124,7 @@ JIT function arguments with |CUSTOM_TYPES|
|
||||
- ``__extract_mlir_values__``: Generate a dynamic expression for the current object.
|
||||
- ``__new_from_mlir_values__``: Create a new object from MLIR values.
|
||||
|
||||
Refer to `typing.py <https://github.com/NVIDIA/cutlass/tree/main/python/CuTeDSL/cutlass/base_dsl/typing.py>`__ for more details on these protocol APIs.
|
||||
Refer to `typing.py <https://github.com/NVIDIA/cutlass/tree/main/python/CuTeDSL/base_dsl/typing.py>`__ for more details on these protocol APIs.
|
||||
|
||||
Depending on different cases of the |CUSTOM_TYPES|, |DSL| provides easy ways to adopt |CUSTOM_TYPES| for JIT function arguments.
|
||||
|
||||
|
||||
@@ -18,9 +18,11 @@ Compilation options allow you to customize how your JIT-compiled functions are b
|
||||
|
||||
These options can be passed as keyword arguments to ``cute.compile`` or set globally for all JIT compilations. The available options and their effects are described in the following sections, along with usage examples to help you get started.
|
||||
|
||||
The |DSL| provides multiple ways to specify compilation options - either by specifying additional arguments to ``cute.compile`` or by using a more Pythonic approach with separate Python types for ``cute.compile``.
|
||||
|
||||
``cute.compile`` Compilation Options
|
||||
------------------------------------
|
||||
|
||||
``cute.compile`` Compilation Options as strings
|
||||
-----------------------------------------------
|
||||
|
||||
You can provide additional compilation options as a string when calling ``cute.compile``. The |DSL| uses ``argparse`` to parse these options and will raise an error if any invalid options are specified.
|
||||
|
||||
@@ -36,10 +38,30 @@ You can provide additional compilation options as a string when calling ``cute.c
|
||||
- Optimization level of compilation. The higher the level, the more optimizations are applied. The valid value range is [0, 3].
|
||||
- 3 (highest level of optimization)
|
||||
- int
|
||||
* - ``enable-device-assertions``
|
||||
- Enable device code assertions.
|
||||
* - ``enable-assertions``
|
||||
- Enable host and device code assertions.
|
||||
- False
|
||||
- bool
|
||||
* - ``keep-cubin``
|
||||
- Keep the generated CUBIN file.
|
||||
- False
|
||||
- bool
|
||||
* - ``keep-ptx``
|
||||
- Keep the generated PTX file.
|
||||
- False
|
||||
- bool
|
||||
* - ``ptxas-options``
|
||||
- The options to pass to the PTX Compiler library.
|
||||
- ""
|
||||
- str
|
||||
* - ``generate-line-info``
|
||||
- Generate line information for debugging.
|
||||
- False
|
||||
- bool
|
||||
* - ``gpu-arch``
|
||||
- The GPU architecture to compile for.
|
||||
- ""
|
||||
- str
|
||||
|
||||
You can use the following code to specify compilation options:
|
||||
|
||||
@@ -47,4 +69,34 @@ You can use the following code to specify compilation options:
|
||||
|
||||
jit_executor_with_opt_level_2 = cute.compile(add, 1, 2, options="--opt-level 2")
|
||||
jit_executor_with_opt_level_1 = cute.compile(add, 1, 2, options="--opt-level 1")
|
||||
jit_executor_with_enable_device_assertions = cute.compile(add, 1, 2, options="--enable-device-assertions")
|
||||
jit_executor_with_enable_device_assertions = cute.compile(add, 1, 2, options="--enable-assertions")
|
||||
jit_executor_with_keep_cubin = cute.compile(add, 1, 2, options="--keep-cubin")
|
||||
jit_executor_with_keep_ptx = cute.compile(add, 1, 2, options="--keep-ptx")
|
||||
jit_executor_with_ptxas_options = cute.compile(add, 1, 2, options="--ptxas-options '--opt-level=2'")
|
||||
|
||||
|
||||
``cute.compile`` Compilation Options as separate Python types
|
||||
-------------------------------------------------------------
|
||||
|
||||
Alternatively, you can also use a more Pythonic way to specify compilation options with separate Python types.
|
||||
Compilation options can be programmatically composed using tuple and passed to ``cute.compile`` separately.
|
||||
|
||||
.. code-block:: python
|
||||
|
||||
from cutlass.cute import OptLevel, EnableAssertions, GenerateLineInfo, KeepCUBIN, KeepPTX
|
||||
|
||||
my_debugging_options = (OptLevel(1), EnableAssertions, GenerateLineInfo, KeepCUBIN, KeepPTX)
|
||||
compiled_kernel_1 = cute.compile[my_debugging_options](my_kernel_1, ...)
|
||||
compiled_kernel_2 = cute.compile[my_debugging_options](my_kernel_2, ...)
|
||||
|
||||
This approach causes invalid options to raise errors immediately, making it much easier to detect typos when specifying multiple options.
|
||||
Notebly, boolean options are automatically converted to True instances of the option type for convenience.
|
||||
|
||||
.. code-block:: python
|
||||
|
||||
jit_executor_with_opt_level_2 = cute.compile[OptLevel(2)](add, 1, 2)
|
||||
jit_executor_with_opt_level_1 = cute.compile[OptLevel(1)](add, 1, 2)
|
||||
jit_executor_with_enable_device_assertions = cute.compile[EnableAssertions](add, 1, 2)
|
||||
jit_executor_with_keep_cubin = cute.compile[KeepCUBIN](add, 1, 2)
|
||||
jit_executor_with_keep_ptx = cute.compile[KeepPTX](add, 1, 2)
|
||||
jit_executor_with_ptxas_options = cute.compile[PtxasOptions("--opt-level=2")](add, 1, 2)
|
||||
@@ -63,7 +63,7 @@ The full signature of from_dlpack is as follows:
|
||||
|
||||
.. code-block:: python
|
||||
|
||||
def from_dlpack(tensor, assumed_align=None):
|
||||
def from_dlpack(tensor, assumed_align=None, use_32bit_stride=False):
|
||||
|
||||
The ``assumed_align`` integer parameter specifies the alignment of the tensor in unit of bytes.
|
||||
The tensor's base address must be divisible by ``assumed_align``. When not provided explicitly,
|
||||
@@ -72,6 +72,13 @@ information is part of the pointer type in the generated IR. Therefore, programs
|
||||
alignments have a different IR and identical IRs are required for hitting the kernel caching
|
||||
mechanism of |DSL|.
|
||||
|
||||
The ``use_32bit_stride`` parameter determines whether to use 32-bit stride for the tensor's dynamic stride values.
|
||||
By default, it is set to False (64bit) to ensure that address calculations do not risk overflow. For smaller
|
||||
problem sizes (where ``cosize(layout_of_tensor) <= Int32_MAX``), users may set it to True (32bit) to improve performance
|
||||
by reducing register usage and the number of address calculation instructions. When ``use_32bit_stride`` is set
|
||||
to True, a runtime check is performed to ensure that the layout does not overflow. Please note that this parameter
|
||||
only has an effect when the tensor's layout is marked as dynamic.
|
||||
|
||||
Code Example
|
||||
~~~~~~~~~~~~
|
||||
|
||||
@@ -242,6 +249,10 @@ The following example demonstrates how to use ``mark_layout_dynamic`` to specify
|
||||
t7 = from_dlpack(b).mark_layout_dynamic(leading_dim=3)
|
||||
# Expected strides[leading_dim] == 1, but got 4
|
||||
|
||||
c = torch.empty(1000000000, 1000000000)
|
||||
t8 = from_dlpack(c, use_32bit_stride=True).mark_layout_dynamic()
|
||||
# Layout in DLTensorWrapper has int32 overflow risk. Please set use_32bit_stride to False.
|
||||
|
||||
Mark the Tensor's Layout as Dynamic with ``mark_compact_shape_dynamic``
|
||||
-----------------------------------------------------------------------
|
||||
|
||||
@@ -398,6 +409,12 @@ The following example demonstrates how to use ``mark_compact_shape_dynamic`` to
|
||||
)
|
||||
# The stride_order is not consistent with the layout
|
||||
|
||||
c = torch.empty(1000000000, 1000000000)
|
||||
t13 = from_dlpack(c, use_32bit_stride=True).mark_compact_shape_dynamic(
|
||||
mode=0, divisibility=1
|
||||
)
|
||||
# Layout in DLTensorWrapper has int32 overflow risk. Please set use_32bit_stride to False.
|
||||
|
||||
|
||||
Bypass the DLPack Protocol
|
||||
--------------------------
|
||||
|
||||
Reference in New Issue
Block a user