diff --git a/docs/build/eps.md b/docs/build/eps.md index 6ddc4e6a07..2c6e2c8948 100644 --- a/docs/build/eps.md +++ b/docs/build/eps.md @@ -250,7 +250,7 @@ See more information on the OpenVINO™ Execution Provider [here](../execution-p * [Windows - CPU, GPU](https://www.intel.com/content/www/us/en/developer/tools/openvino-toolkit/download.html?VERSION=v_2023_1_0&OP_SYSTEM=WINDOWS&DISTRIBUTION=ARCHIVE). * [Linux - CPU, GPU](https://www.intel.com/content/www/us/en/developer/tools/openvino-toolkit/download.html?VERSION=v_2023_1_0&OP_SYSTEM=LINUX&DISTRIBUTION=ARCHIVE) - Follow [documentation](https://docs.openvino.ai/2023.1/index.html) for detailed instructions. + Follow [documentation](https://docs.openvino.ai/latest/index.html) for detailed instructions. *2023.1 is the recommended OpenVINO™ version. [OpenVINO™ 2022.1](https://docs.openvino.ai/archive/2022.1/index.html) is minimal OpenVINO™ version requirement.* *The minimum ubuntu version to support 2023.1 is 18.04.* @@ -291,7 +291,7 @@ See more information on the OpenVINO™ Execution Provider [here](../execution-p * `--use_openvino` builds the OpenVINO™ Execution Provider in ONNX Runtime. * ``: Specifies the default hardware target for building OpenVINO™ Execution Provider. This can be overriden dynamically at runtime with another option (refer to [OpenVINO™-ExecutionProvider](../execution-providers/OpenVINO-ExecutionProvider.md#summary-of-options) for more details on dynamic device selection). Below are the options for different Intel target devices. -Refer to [Intel GPU device naming convention](https://docs.openvino.ai/2023.0/openvino_docs_OV_UG_supported_plugins_GPU.html#device-naming-convention) for specifying the correct hardware target in cases where both integrated and discrete GPU's co-exist. +Refer to [Intel GPU device naming convention](https://docs.openvino.ai/latest/openvino_docs_OV_UG_supported_plugins_GPU.html#device-naming-convention) for specifying the correct hardware target in cases where both integrated and discrete GPU's co-exist. | Hardware Option | Target Device | | --------------- | ------------------------| diff --git a/docs/execution-providers/CUDA-ExecutionProvider.md b/docs/execution-providers/CUDA-ExecutionProvider.md index 33861a896b..05732687d1 100644 --- a/docs/execution-providers/CUDA-ExecutionProvider.md +++ b/docs/execution-providers/CUDA-ExecutionProvider.md @@ -7,131 +7,160 @@ redirect_from: /docs/reference/execution-providers/CUDA-ExecutionProvider --- # CUDA Execution Provider + {: .no_toc } The CUDA Execution Provider enables hardware accelerated computation on Nvidia CUDA-enabled GPUs. - ## Contents + {: .no_toc } * TOC placeholder -{:toc} + {:toc} ## Install -Pre-built binaries of ONNX Runtime with CUDA EP are published for most language bindings. Please reference [Install ORT](../install). +Pre-built binaries of ONNX Runtime with CUDA EP are published for most language bindings. Please +reference [Install ORT](../install). ## Requirements -Please reference table below for official GPU packages dependencies for the ONNX Runtime inferencing package. Note that ONNX Runtime Training is aligned with PyTorch CUDA versions; refer to the Training tab on [onnxruntime.ai](https://onnxruntime.ai/) for supported versions. +Please reference table below for official GPU packages dependencies for the ONNX Runtime inferencing package. Note that +ONNX Runtime Training is aligned with PyTorch CUDA versions; refer to the Training tab +on [onnxruntime.ai](https://onnxruntime.ai/) for supported versions. -Note: Because of CUDA Minor Version Compatibility, Onnx Runtime built with CUDA 11.4 should be compatible with any CUDA 11.x version. -Please reference [Nvidia CUDA Minor Version Compatibility](https://docs.nvidia.com/deploy/cuda-compatibility/#minor-version-compatibility). +Note: Because of CUDA Minor Version Compatibility, Onnx Runtime built with CUDA 11.4 should be compatible with any CUDA +11.x version. +Please +reference [Nvidia CUDA Minor Version Compatibility](https://docs.nvidia.com/deploy/cuda-compatibility/#minor-version-compatibility). -|ONNX Runtime|CUDA|cuDNN|Notes| -|---|---|---|---| -|1.15
1.16|11.8|8.2.4 (Linux)
8.5.0.96 (Windows)|Tested with CUDA versions from 11.6 up to 11.8, and cuDNN from 8.2.4 up to 8.7.0| -|1.14
1.13.1
1.13|11.6|8.2.4 (Linux)
8.5.0.96 (Windows)|libcudart 11.4.43
libcufft 10.5.2.100
libcurand 10.2.5.120
libcublasLt 11.6.5.2
libcublas 11.6.5.2
libcudnn 8.2.4| -|1.12
1.11|11.4|8.2.4 (Linux)
8.2.2.26 (Windows)|libcudart 11.4.43
libcufft 10.5.2.100
libcurand 10.2.5.120
libcublasLt 11.6.5.2
libcublas 11.6.5.2
libcudnn 8.2.4| -|1.10|11.4|8.2.4 (Linux)
8.2.2.26 (Windows)|libcudart 11.4.43
libcufft 10.5.2.100
libcurand 10.2.5.120
libcublasLt 11.6.1.51
libcublas 11.6.1.51
libcudnn 8.2.4| -|1.9|11.4|8.2.4 (Linux)
8.2.2.26 (Windows)|libcudart 11.4.43
libcufft 10.5.2.100
libcurand 10.2.5.120
libcublasLt 11.6.1.51
libcublas 11.6.1.51
libcudnn 8.2.4| -|1.8|11.0.3|8.0.4 (Linux)
8.0.2.39 (Windows)|libcudart 11.0.221
libcufft 10.2.1.245
libcurand 10.2.1.245
libcublasLt 11.2.0.252
libcublas 11.2.0.252
libcudnn 8.0.4| -|1.7|11.0.3|8.0.4 (Linux)
8.0.2.39 (Windows)|libcudart 11.0.221
libcufft 10.2.1.245
libcurand 10.2.1.245
libcublasLt 11.2.0.252
libcublas 11.2.0.252
libcudnn 8.0.4| -|1.5-1.6|10.2|8.0.3|CUDA 11 can be built from source| -|1.2-1.4|10.1|7.6.5|Requires cublas10-10.2.1.243; cublas 10.1.x will not work| -|1.0-1.1|10.0|7.6.4|CUDA versions from 9.1 up to 10.1, and cuDNN versions from 7.1 up to 7.4 should also work with Visual Studio 2017| +| ONNX Runtime | CUDA | cuDNN | Notes | +|--------------------------|--------|-----------------------------------------|-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| +| 1.17 | 12.2 | 8.9.2.26 (Linux)
8.9.2.26 (Windows) | The default CUDA version for ORT 1.17 is CUDA 11.8. To install CUDA 12 package, please look at [Install ORT](../install).
Due to low demend on Java GPU package, only C++/C# Nuget and Python packages are released with CUDA 12.2 | +| 1.15
1.16
1.17 | 11.8 | 8.2.4 (Linux)
8.5.0.96 (Windows) | Tested with CUDA versions from 11.6 up to 11.8, and cuDNN from 8.2.4 up to 8.7.0 | +| 1.14
1.13.1
1.13 | 11.6 | 8.2.4 (Linux)
8.5.0.96 (Windows) | libcudart 11.4.43
libcufft 10.5.2.100
libcurand 10.2.5.120
libcublasLt 11.6.5.2
libcublas 11.6.5.2
libcudnn 8.2.4 | +| 1.12
1.11 | 11.4 | 8.2.4 (Linux)
8.2.2.26 (Windows) | libcudart 11.4.43
libcufft 10.5.2.100
libcurand 10.2.5.120
libcublasLt 11.6.5.2
libcublas 11.6.5.2
libcudnn 8.2.4 | +| 1.10 | 11.4 | 8.2.4 (Linux)
8.2.2.26 (Windows) | libcudart 11.4.43
libcufft 10.5.2.100
libcurand 10.2.5.120
libcublasLt 11.6.1.51
libcublas 11.6.1.51
libcudnn 8.2.4 | +| 1.9 | 11.4 | 8.2.4 (Linux)
8.2.2.26 (Windows) | libcudart 11.4.43
libcufft 10.5.2.100
libcurand 10.2.5.120
libcublasLt 11.6.1.51
libcublas 11.6.1.51
libcudnn 8.2.4 | +| 1.8 | 11.0.3 | 8.0.4 (Linux)
8.0.2.39 (Windows) | libcudart 11.0.221
libcufft 10.2.1.245
libcurand 10.2.1.245
libcublasLt 11.2.0.252
libcublas 11.2.0.252
libcudnn 8.0.4 | +| 1.7 | 11.0.3 | 8.0.4 (Linux)
8.0.2.39 (Windows) | libcudart 11.0.221
libcufft 10.2.1.245
libcurand 10.2.1.245
libcublasLt 11.2.0.252
libcublas 11.2.0.252
libcudnn 8.0.4 | +| 1.5-1.6 | 10.2 | 8.0.3 | CUDA 11 can be built from source | +| 1.2-1.4 | 10.1 | 7.6.5 | Requires cublas10-10.2.1.243; cublas 10.1.x will not work | +| 1.0-1.1 | 10.0 | 7.6.4 | CUDA versions from 9.1 up to 10.1, and cuDNN versions from 7.1 up to 7.4 should also work with Visual Studio 2017 | For older versions, please reference the readme and build pages on the release branch. -For Windows, [Microsoft C and C++ (MSVC) runtime libraries](https://learn.microsoft.com/en-us/cpp/windows/latest-supported-vc-redist) is also required. +For +Windows, [Microsoft C and C++ (MSVC) runtime libraries](https://learn.microsoft.com/en-us/cpp/windows/latest-supported-vc-redist) +is also required. ## Build + For build instructions, please see the [BUILD page](../build/eps.md#cuda). ## Configuration Options + The CUDA Execution Provider supports the following configuration options. ### device_id + The device ID. Default value: 0 ### user_compute_stream + Defines the compute stream for the inference to run on. -It implicitly sets the `has_user_compute_stream` option. It cannot be set through `UpdateCUDAProviderOptions`, but rather `UpdateCUDAProviderOptionsWithValue`. +It implicitly sets the `has_user_compute_stream` option. It cannot be set through `UpdateCUDAProviderOptions`, but +rather `UpdateCUDAProviderOptionsWithValue`. This cannot be used in combination with an external allocator. Example python usage: + ```python -providers = [("CUDAExecutionProvider", {"device_id":torch.cuda.current_device(), "user_compute_stream": str(torch.cuda.current_stream().cuda_stream)})] +providers = [("CUDAExecutionProvider", {"device_id": torch.cuda.current_device(), + "user_compute_stream": str(torch.cuda.current_stream().cuda_stream)})] sess_options = ort.SessionOptions() sess = ort.InferenceSession("my_model.onnx", sess_options=sess_options, providers=providers) ``` -To take advantage of user compute stream, it is recommended to use [I/O Binding](../api/python/api_summary.html#data-on-device) to bind inputs and outputs to tensors in device. +To take advantage of user compute stream, it is recommended to +use [I/O Binding](../api/python/api_summary.html#data-on-device) to bind inputs and outputs to tensors in device. ### do_copy_in_default_stream -Whether to do copies in the default stream or use separate streams. The recommended setting is true. If false, there are race conditions and possibly better performance. + +Whether to do copies in the default stream or use separate streams. The recommended setting is true. If false, there are +race conditions and possibly better performance. Default value: true ### use_ep_level_unified_stream -Uses the same CUDA stream for all threads of the CUDA EP. This is implicitly enabled by `has_user_compute_stream`, `enable_cuda_graph` or when using an external allocator. + +Uses the same CUDA stream for all threads of the CUDA EP. This is implicitly enabled +by `has_user_compute_stream`, `enable_cuda_graph` or when using an external allocator. Default value: false ### gpu_mem_limit -The size limit of the device memory arena in bytes. This size limit is only for the execution provider's arena. The total device memory usage may be higher. + +The size limit of the device memory arena in bytes. This size limit is only for the execution provider's arena. The +total device memory usage may be higher. s: max value of C++ size_t type (effectively unlimited) _Note:_ Will be over-ridden by contents of `default_memory_arena_cfg` (if specified) ### arena_extend_strategy + The strategy for extending the device memory arena. -Value | Description --|- -kNextPowerOfTwo (0) | subsequent extensions extend by larger amounts (multiplied by powers of two) -kSameAsRequested (1) | extend by the requested amount + Value | Description +----------------------|------------------------------------------------------------------------------ + kNextPowerOfTwo (0) | subsequent extensions extend by larger amounts (multiplied by powers of two) + kSameAsRequested (1) | extend by the requested amount Default value: kNextPowerOfTwo _Note:_ Will be over-ridden by contents of `default_memory_arena_cfg` (if specified) ### cudnn_conv_algo_search + The type of search done for cuDNN convolution algorithms. -Value | Description --|- -EXHAUSTIVE (0) | expensive exhaustive benchmarking using cudnnFindConvolutionForwardAlgorithmEx -HEURISTIC (1) | lightweight heuristic based search using cudnnGetConvolutionForwardAlgorithm_v7 -DEFAULT (2) | default algorithm using CUDNN_CONVOLUTION_FWD_ALGO_IMPLICIT_PRECOMP_GEMM + Value | Description +----------------|--------------------------------------------------------------------------------- + EXHAUSTIVE (0) | expensive exhaustive benchmarking using cudnnFindConvolutionForwardAlgorithmEx + HEURISTIC (1) | lightweight heuristic based search using cudnnGetConvolutionForwardAlgorithm_v7 + DEFAULT (2) | default algorithm using CUDNN_CONVOLUTION_FWD_ALGO_IMPLICIT_PRECOMP_GEMM Default value: EXHAUSTIVE - ### cudnn_conv_use_max_workspace + Check [tuning performance for convolution heavy models](#convolution-heavy-models) for details on what this flag does. This flag is only supported from the V2 version of the provider options struct when used using the C API.(sample below) Default value: 1, for versions 1.14 and later - 0, for previous versions +0, for previous versions ### cudnn_conv1d_pad_to_nc1d + Check [convolution input padding in the CUDA EP](#convolution-input-padding) for details on what this flag does. This flag is only supported from the V2 version of the provider options struct when used using the C API. (sample below) Default value: 0 ### enable_cuda_graph + Check [using CUDA Graphs in the CUDA EP](#using-cuda-graphs-preview) for details on what this flag does. This flag is only supported from the V2 version of the provider options struct when used using the C API. (sample below) Default value: 0 ### enable_skip_layer_norm_strict_mode -Whether to use strict mode in SkipLayerNormalization cuda implementation. The default and recommanded setting is false. If enabled, accuracy improvement and performance drop can be expected. + +Whether to use strict mode in SkipLayerNormalization cuda implementation. The default and recommanded setting is false. +If enabled, accuracy improvement and performance drop can be expected. This flag is only supported from the V2 version of the provider options struct when used using the C API. (sample below) Default value: 0 @@ -140,8 +169,10 @@ Default value: 0 gpu_external_* is used to pass external allocators. Example python usage: + ```python from onnxruntime.training.ortmodule.torch_cpp_extensions import torch_gpu_allocator + provider_option_map["gpu_external_alloc"] = str(torch_gpu_allocator.gpu_caching_allocator_raw_alloc_address()) provider_option_map["gpu_external_free"] = str(torch_gpu_allocator.gpu_caching_allocator_raw_delete_address()) provider_option_map["gpu_external_empty_cache"] = str(torch_gpu_allocator.gpu_caching_allocator_empty_cache_address()) @@ -150,29 +181,56 @@ provider_option_map["gpu_external_empty_cache"] = str(torch_gpu_allocator.gpu_ca Default value: 0 ### prefer_nhwc -This option is not available in default builds ! One has to compile ONNX Runtime with `onnxruntime_USE_CUDA_NHWC_OPS=ON`. -If this is enabled the EP prefers NHWC operators over NCHW. Needed transforms will be added to the model. As NVIDIA tensor cores can only work on NHWC layout this can increase performance if the model consists of many supported operators and does not need too many new transpose nodes. Wider operator support is planned in the future. -This flag is only supported from the V2 version of the provider options struct when used using the C API. The V2 provider options struct can be created using [CreateCUDAProviderOptions](https://onnxruntime.ai/docs/api/c/struct_ort_api.html#a0d29cbf555aa806c050748cf8d2dc172) and updated using [UpdateCUDAProviderOptions](https://onnxruntime.ai/docs/api/c/struct_ort_api.html#a4710fc51f75a4b9a75bde20acbfa0783). + +This option is not available in default builds ! One has to compile ONNX Runtime +with `onnxruntime_USE_CUDA_NHWC_OPS=ON`. +If this is enabled the EP prefers NHWC operators over NCHW. Needed transforms will be added to the model. As NVIDIA +tensor cores can only work on NHWC layout this can increase performance if the model consists of many supported +operators and does not need too many new transpose nodes. Wider operator support is planned in the future. +This flag is only supported from the V2 version of the provider options struct when used using the C API. The V2 +provider options struct can be created +using [CreateCUDAProviderOptions](https://onnxruntime.ai/docs/api/c/struct_ort_api.html#a0d29cbf555aa806c050748cf8d2dc172) +and updated +using [UpdateCUDAProviderOptions](https://onnxruntime.ai/docs/api/c/struct_ort_api.html#a4710fc51f75a4b9a75bde20acbfa0783). Default value: 0 ## Performance Tuning -The [I/O Binding feature](../performance/tune-performance/iobinding.md) should be utilized to avoid overhead resulting from copies on inputs and outputs. Ideally up and downloads for inputs can be hidden behind the inference. This can be achieved by doing asynchronous copies while running inference. This is demonstrated in this [PR](https://github.com/microsoft/onnxruntime/pull/14088) + +The [I/O Binding feature](../performance/tune-performance/iobinding.md) should be utilized to avoid overhead resulting +from copies on inputs and outputs. Ideally up and downloads for inputs can be hidden behind the inference. This can be +achieved by doing asynchronous copies while running inference. This is demonstrated in +this [PR](https://github.com/microsoft/onnxruntime/pull/14088) + ```c++ Ort::RunOptions run_options; run_options.AddConfigEntry("disable_synchronize_execution_providers", "1"); session->Run(run_options, io_binding); ``` -By disabling the synchronization on the inference the user has to take care of synchronizing the compute stream after execution. -This feature should only be used with device local memory or an ORT Value allocated in [pinned memory](https://developer.nvidia.com/blog/how-optimize-data-transfers-cuda-cc/), otherwise the issued download will be blocking and not behave as desired. + +By disabling the synchronization on the inference the user has to take care of synchronizing the compute stream after +execution. +This feature should only be used with device local memory or an ORT Value allocated +in [pinned memory](https://developer.nvidia.com/blog/how-optimize-data-transfers-cuda-cc/), otherwise the issued +download will be blocking and not behave as desired. ### Convolution-heavy models -ORT leverages CuDNN for convolution operations and the first step in this process is to determine which "optimal" convolution algorithm to use while performing the convolution operation for the given input configuration (input shape, filter shape, etc.) in each `Conv` node. This sub-step involves querying CuDNN for a "workspace" memory size and have this allocated so that CuDNN can use this auxiliary memory while determining the "optimal" convolution algorithm to use. +ORT leverages CuDNN for convolution operations and the first step in this process is to determine which "optimal" +convolution algorithm to use while performing the convolution operation for the given input configuration (input shape, +filter shape, etc.) in each `Conv` node. This sub-step involves querying CuDNN for a "workspace" memory size and have +this allocated so that CuDNN can use this auxiliary memory while determining the "optimal" convolution algorithm to use. -The default value of `cudnn_conv_use_max_workspace` is 1 for versions 1.14 or later, and 0 for previous versions. When its value is 0, ORT clamps the workspace size to 32 MB which may lead to a sub-optimal convolution algorithm getting picked by CuDNN. To allow ORT to allocate the maximum possible workspace as determined by CuDNN, a provider option named `cudnn_conv_use_max_workspace` needs to get set (as shown below). +The default value of `cudnn_conv_use_max_workspace` is 1 for versions 1.14 or later, and 0 for previous versions. When +its value is 0, ORT clamps the workspace size to 32 MB which may lead to a sub-optimal convolution algorithm getting +picked by CuDNN. To allow ORT to allocate the maximum possible workspace as determined by CuDNN, a provider option +named `cudnn_conv_use_max_workspace` needs to get set (as shown below). -Keep in mind that using this flag may increase the peak memory usage by a factor (sometimes a few GBs) but this does help CuDNN pick the best convolution algorithm for the given input. We have found that this is an important flag to use while using an fp16 model as this allows CuDNN to pick tensor core algorithms for the convolution operations (if the hardware supports tensor core operations). This flag may or may not result in performance gains for other data types (`float` and `double`). +Keep in mind that using this flag may increase the peak memory usage by a factor (sometimes a few GBs) but this does +help CuDNN pick the best convolution algorithm for the given input. We have found that this is an important flag to use +while using an fp16 model as this allows CuDNN to pick tensor core algorithms for the convolution operations (if the +hardware supports tensor core operations). This flag may or may not result in performance gains for other data +types (`float` and `double`). * Python ```python @@ -182,7 +240,6 @@ Keep in mind that using this flag may increase the peak memory usage by a factor ``` - * C/C++ ```c++ OrtCUDAProviderOptionsV2* cuda_options = nullptr; @@ -200,22 +257,26 @@ Keep in mind that using this flag may increase the peak memory usage by a factor ReleaseCUDAProviderOptions(cuda_options); ``` - * C# - ```csharp - var cudaProviderOptions = new OrtCUDAProviderOptions(); // Dispose this finally +* C# + ```csharp + var cudaProviderOptions = new OrtCUDAProviderOptions(); // Dispose this finally - var providerOptionsDict = new Dictionary(); - providerOptionsDict["cudnn_conv_use_max_workspace"] = "1"; + var providerOptionsDict = new Dictionary(); + providerOptionsDict["cudnn_conv_use_max_workspace"] = "1"; - cudaProviderOptions.UpdateOptions(providerOptionsDict); - - SessionOptions options = SessionOptions.MakeSessionOptionWithCudaProvider(cudaProviderOptions); // Dispose this finally - ``` + cudaProviderOptions.UpdateOptions(providerOptionsDict); + SessionOptions options = SessionOptions.MakeSessionOptionWithCudaProvider(cudaProviderOptions); // Dispose this finally + ``` ### Convolution Input Padding -ORT leverages CuDNN for convolution operations. While CuDNN only takes 4-D or 5-D tensor as input for convolution operations, dimension padding is needed if the input is 3-D tensor. Given an input tensor of shape [N, C, D], it can be padded to [N, C, D, 1] or [N, C, 1, D]. While both of these two padding ways produce same output, the performance may be a lot different because different convolution algorithms are selected, especially on some devices such as A100. By default the input is padded to [N, C, D, 1]. A provider option named `cudnn_conv1d_pad_to_nc1d` needs to get set (as shown below) if [N, C, 1, D] is preferred. +ORT leverages CuDNN for convolution operations. While CuDNN only takes 4-D or 5-D tensor as input for convolution +operations, dimension padding is needed if the input is 3-D tensor. Given an input tensor of shape [N, C, D], it can be +padded to [N, C, D, 1] or [N, C, 1, D]. While both of these two padding ways produce same output, the performance may be +a lot different because different convolution algorithms are selected, especially on some devices such as A100. By +default the input is padded to [N, C, D, 1]. A provider option named `cudnn_conv1d_pad_to_nc1d` needs to get set (as +shown below) if [N, C, 1, D] is preferred. * Python ```python @@ -253,7 +314,10 @@ ORT leverages CuDNN for convolution operations. While CuDNN only takes 4-D or 5- ### Using CUDA Graphs (Preview) -While using the CUDA EP, ORT supports the usage of [CUDA Graphs](https://developer.nvidia.com/blog/cuda-10-features-revealed/) to remove CPU overhead associated with launching CUDA kernels sequentially. To enable the usage of CUDA Graphs, use the provider option as shown in the samples below. +While using the CUDA EP, ORT supports the usage +of [CUDA Graphs](https://developer.nvidia.com/blog/cuda-10-features-revealed/) to remove CPU overhead associated with +launching CUDA kernels sequentially. To enable the usage of CUDA Graphs, use the provider option as shown in the samples +below. Currently, there are some constraints with regards to using the CUDA Graphs feature: * Models with control-flow ops (i.e. `If`, `Loop` and `Scan` ops) are not supported. @@ -262,16 +326,26 @@ Currently, there are some constraints with regards to using the CUDA Graphs feat * The input/output types of models need to be tensors. -* Shapes of inputs/outputs cannot change across inference calls. Dynamic shape models are supported - the only constraint is that the input/output shapes should be the same across all inference calls. +* Shapes of inputs/outputs cannot change across inference calls. Dynamic shape models are supported - the only + constraint is that the input/output shapes should be the same across all inference calls. -* By design, [CUDA Graphs](https://developer.nvidia.com/blog/cuda-10-features-revealed/) is designed to read from/write to the same CUDA virtual memory addresses during the graph replaying step as it does during the graph capturing step. Due to this requirement, usage of this feature requires using IOBinding so as to bind memory which will be used as input(s)/output(s) for the CUDA Graph machinery to read from/write to (please see samples below). +* By design, [CUDA Graphs](https://developer.nvidia.com/blog/cuda-10-features-revealed/) is designed to read from/write + to the same CUDA virtual memory addresses during the graph replaying step as it does during the graph capturing step. + Due to this requirement, usage of this feature requires using IOBinding so as to bind memory which will be used as + input(s)/output(s) for the CUDA Graph machinery to read from/write to (please see samples below). -* While updating the input(s) for subsequent inference calls, the fresh input(s) need to be copied over to the corresponding CUDA memory location(s) of the bound `OrtValue` input(s) (please see samples below to see how this can be achieved). This is due to the fact that the "graph replay" will require reading inputs from the same CUDA virtual memory addresses. +* While updating the input(s) for subsequent inference calls, the fresh input(s) need to be copied over to the + corresponding CUDA memory location(s) of the bound `OrtValue` input(s) (please see samples below to see how this can + be achieved). This is due to the fact that the "graph replay" will require reading inputs from the same CUDA virtual + memory addresses. -* Multi-threaded usage is currently not supported, i.e. `Run()` MAY NOT be invoked on the same `InferenceSession` object from multiple threads while using CUDA Graphs. - -NOTE: The very first `Run()` performs a variety of tasks under the hood like making CUDA memory allocations, capturing the CUDA graph for the model, and then performing a graph replay to ensure that the graph runs. Due to this, the latency associated with the first `Run()` is bound to be high. Subsequent `Run()`s only perform graph replays of the graph captured and cached in the first `Run()`. +* Multi-threaded usage is currently not supported, i.e. `Run()` MAY NOT be invoked on the same `InferenceSession` object + from multiple threads while using CUDA Graphs. +NOTE: The very first `Run()` performs a variety of tasks under the hood like making CUDA memory allocations, capturing +the CUDA graph for the model, and then performing a graph replay to ensure that the graph runs. Due to this, the latency +associated with the first `Run()` is bound to be high. Subsequent `Run()`s only perform graph replays of the graph +captured and cached in the first `Run()`. * Python @@ -376,7 +450,6 @@ NOTE: The very first `Run()` performs a variety of tasks under the hood like mak * C# (future) - ## Samples ### Python @@ -457,6 +530,7 @@ cudaProviderOptions.UpdateOptions(providerOptionsDict); SessionOptions options = SessionOptions.MakeSessionOptionWithCudaProvider(cudaProviderOptions); // Dispose this finally ``` + Also see the tutorial here on how to [configure CUDA for C# on Windows](../tutorials/csharp/csharp-gpu.md). ### Java diff --git a/docs/install/index.md b/docs/install/index.md index a283637a30..2ad321769a 100644 --- a/docs/install/index.md +++ b/docs/install/index.md @@ -8,33 +8,77 @@ redirect_from: /docs/how-to/install # Install ONNX Runtime (ORT) +See the [installation matrix](https://onnxruntime.ai) for recommended instructions for desired combinations of target +operating system, hardware, accelerator, and language. -See the [installation matrix](https://onnxruntime.ai) for recommended instructions for desired combinations of target operating system, hardware, accelerator, and language. - -Details on OS versions, compilers, language versions, dependent libraries, etc can be found under [Compatibility](../reference/compatibility). - +Details on OS versions, compilers, language versions, dependent libraries, etc can be found +under [Compatibility](../reference/compatibility). ## Contents + {: .no_toc } * TOC placeholder -{:toc} + {:toc} ## Requirements -* All builds require the English language package with `en_US.UTF-8` locale. On Linux, install [language-pack-en package](https://packages.ubuntu.com/search?keywords=language-pack-en) -by running `locale-gen en_US.UTF-8` and `update-locale LANG=en_US.UTF-8` +* All builds require the English language package with `en_US.UTF-8` locale. On Linux, + install [language-pack-en package](https://packages.ubuntu.com/search?keywords=language-pack-en) + by running `locale-gen en_US.UTF-8` and `update-locale LANG=en_US.UTF-8` -* Windows builds require [Visual C++ 2019 runtime](https://support.microsoft.com/en-us/help/2977003/the-latest-supported-visual-c-downloads). The latest version is recommended. +* Windows builds + require [Visual C++ 2019 runtime](https://support.microsoft.com/en-us/help/2977003/the-latest-supported-visual-c-downloads). + The latest version is recommended. ## Python Installs ### Install ONNX Runtime (ORT) +#### Install ONNX Runtime CPU + ```bash pip install onnxruntime ``` +#### Install ONNX Runtime GPU (CUDA 11.x) + +The default CUDA version for ORT is 11.8 + +```bash +pip install onnxruntime-gpu +``` + +#### Install ONNX Runtime GPU (CUDA 12.x) + +For Cuda 12.x please using the following instructions to download the instruction +from [ORT Azure Devops Feed](https://aiinfra.visualstudio.com/PublicPackages/_artifacts/feed/onnxruntime-cuda-12/) + +1. Install keyring for Azure Artifacts + +```bash +python -m pip install --upgrade pip +pip install keyring artifacts-keyring +``` + +If you're using Linux, ensure you've installed the [prerequisites](https://go.microsoft.com/fwlink/?linkid=2103695), +which are required for artifacts-keyring. + +2. Project setup + +Ensure you have installed the latest version of the Azure Artifacts keyring from the "Get the tools" menu. +If you don't already have one, create a virtualenv +using [these instructions](https://go.microsoft.com/fwlink/?linkid=2103878) from the official Python documentation. +Per the instructions, "it is always recommended to use a virtualenv while developing Python applications." + +Add a pip.ini (Windows) or pip.conf (Linux) file to your virtualenv + +```bash +pip install onnxruntime-gpu --index-url https://aiinfra.pkgs.visualstudio.com/PublicPackages/_packaging/onnxruntime-cuda-12/pypi/simple/ +``` + +3. Install onnxruntime-gpu + ```bash pip install onnxruntime-gpu ``` @@ -45,7 +89,8 @@ pip install onnxruntime-gpu ## ONNX is built into PyTorch pip install torch ``` -```python + +```bash ## tensorflow pip install tf2onnx ``` @@ -59,33 +104,80 @@ pip install skl2onnx ### Install ONNX Runtime (ORT) +#### Install ONNX Runtime CPU + ```bash # CPU dotnet add package Microsoft.ML.OnnxRuntime ``` + +#### Install ONNX Runtime GPU (CUDA 11.x) + +The default CUDA version for ORT is 11.8 + ```bash # GPU dotnet add package Microsoft.ML.OnnxRuntime.Gpu ``` + +#### Install ONNX Runtime GPU (CUDA 12.x) + +1. Project Setup + +Ensure you have installed the latest version of the Azure Artifacts keyring from the +its [Github Repo](https://github.com/microsoft/artifacts-credprovider#azure-artifacts-credential-provider).
Add +a nuget.config file to your project in the same directory as your .csproj file. + +```xml + + + + + + + +``` + +2. Restore packages + +Restore packages (using the interactive flag, which allows dotnet to prompt you for credentials) + +```bash +dotnet add package Microsoft.ML.OnnxRuntime.Gpu +``` + +Note: You don't need --interactive every time. dotnet will prompt you to add --interactive if it needs updated +credentials. + +#### DirectML + ```bash -# DirectML dotnet add package Microsoft.ML.OnnxRuntime.DirectML ``` +#### WinML + ```bash -# WinML dotnet add package Microsoft.AI.MachineLearning ``` ## Install on web and mobile -Unless stated otherwise, the installation instructions in this section refer to pre-built packages that include support for selected operators and ONNX opset versions based on the requirements of popular models. These packages may be referred to as "mobile packages". If you use mobile packages, your model must only use the supported [opsets and operators](../reference/operators/mobile_package_op_type_support_1.14.md). +Unless stated otherwise, the installation instructions in this section refer to pre-built packages that include support +for selected operators and ONNX opset versions based on the requirements of popular models. These packages may be +referred to as "mobile packages". If you use mobile packages, your model must only use the +supported [opsets and operators](../reference/operators/mobile_package_op_type_support_1.14.md). -Another type of pre-built package has full support for all ONNX opsets and operators, at the cost of larger binary size. These packages are referred to as "full packages". +Another type of pre-built package has full support for all ONNX opsets and operators, at the cost of larger binary size. +These packages are referred to as "full packages". -If the pre-built mobile package supports your model/s but is too large, you can create a [custom build](../build/custom.md). A custom build can include just the opsets and operators in your model/s to reduce the size. +If the pre-built mobile package supports your model/s but is too large, you can create +a [custom build](../build/custom.md). A custom build can include just the opsets and operators in your model/s to reduce +the size. -If the pre-built mobile package does not include the opsets or operators in your model/s, you can either use the full package if available, or create a custom build. +If the pre-built mobile package does not include the opsets or operators in your model/s, you can either use the full +package if available, or create a custom build. ### JavaScript Installs @@ -108,7 +200,6 @@ npm install onnxruntime-node #### Install ONNX Runtime for React Native - ```bash # install latest release version npm install onnxruntime-react-native @@ -116,7 +207,9 @@ npm install onnxruntime-react-native ### Install on iOS -In your CocoaPods `Podfile`, add the `onnxruntime-c`, `onnxruntime-mobile-c`, `onnxruntime-objc`, or `onnxruntime-mobile-objc` pod, depending on whether you want to use a full or mobile package and which API you want to use. +In your CocoaPods `Podfile`, add the `onnxruntime-c`, `onnxruntime-mobile-c`, `onnxruntime-objc`, +or `onnxruntime-mobile-objc` pod, depending on whether you want to use a full or mobile package and which API you want +to use. #### C/C++ @@ -170,7 +263,11 @@ In your Android Studio Project, make the following changes to: #### C/C++ -Download the [onnxruntime-android](https://mvnrepository.com/artifact/com.microsoft.onnxruntime/onnxruntime-android) (full package) or [onnxruntime-mobile](https://mvnrepository.com/artifact/com.microsoft.onnxruntime/onnxruntime-mobile) (mobile package) AAR hosted at MavenCentral, change the file extension from `.aar` to `.zip`, and unzip it. Include the header files from the `headers` folder, and the relevant `libonnxruntime.so` dynamic library from the `jni` folder in your NDK project. +Download the [onnxruntime-android](https://mvnrepository.com/artifact/com.microsoft.onnxruntime/onnxruntime-android) ( +full package) or [onnxruntime-mobile](https://mvnrepository.com/artifact/com.microsoft.onnxruntime/onnxruntime-mobile) ( +mobile package) AAR hosted at MavenCentral, change the file extension from `.aar` to `.zip`, and unzip it. Include the +header files from the `headers` folder, and the relevant `libonnxruntime.so` dynamic library from the `jni` folder in +your NDK project. #### Custom build @@ -178,9 +275,11 @@ Refer to the instructions for creating a [custom Android package](../build/custo ## Install for On-Device Training -Unless stated otherwise, the installation instructions in this section refer to pre-built packages designed to perform on-device training. +Unless stated otherwise, the installation instructions in this section refer to pre-built packages designed to perform +on-device training. -If the pre-built training package supports your model but is too large, you can create a [custom training build](../build/custom.md). +If the pre-built training package supports your model but is too large, you can create +a [custom training build](../build/custom.md). ### Offline Phase - Prepare for Training @@ -316,43 +415,44 @@ pip install torch-ort python -m torch_ort.configure ``` -**Note**: This installs the default version of the `torch-ort` and `onnxruntime-training` packages that are mapped to specific versions of the CUDA libraries. Refer to the install options in [onnxruntime.ai](https://onnxruntime.ai). +**Note**: This installs the default version of the `torch-ort` and `onnxruntime-training` packages that are mapped to +specific versions of the CUDA libraries. Refer to the install options in [onnxruntime.ai](https://onnxruntime.ai). ## Inference install table for all languages -The table below lists the build variants available as officially supported packages. Others can be [built from source](../build/inferencing) from each [release branch](https://github.com/microsoft/onnxruntime/tags). - -In addition to general [requirements](#requirements), please note additional requirements and dependencies in the table below: - -||Official build|Nightly build|Reqs| -|---|---|---|---| -|Python|If using pip, run `pip install --upgrade pip` prior to downloading.||| -||CPU: [**onnxruntime**](https://pypi.org/project/onnxruntime)| [ort-nightly (dev)](https://aiinfra.visualstudio.com/PublicPackages/_artifacts/feed/ORT-Nightly/PyPI/ort-nightly/overview)|| -||GPU (CUDA/TensorRT): [**onnxruntime-gpu**](https://pypi.org/project/onnxruntime-gpu) | [ort-nightly-gpu (dev)](https://aiinfra.visualstudio.com/PublicPackages/_artifacts/feed/ORT-Nightly/PyPI/ort-nightly-gpu/overview/)|[View](../execution-providers/CUDA-ExecutionProvider.md#requirements)| -||GPU (DirectML): [**onnxruntime-directml**](https://pypi.org/project/onnxruntime-directml/) | [ort-nightly-directml (dev)](https://aiinfra.visualstudio.com/PublicPackages/_artifacts/feed/ORT-Nightly/PyPI/ort-nightly-directml/overview/)|[View](../execution-providers/DirectML-ExecutionProvider.md#requirements)| -||OpenVINO: [**intel/onnxruntime**](https://github.com/intel/onnxruntime/releases/latest) - *Intel managed*||[View](../build/eps.md#openvino)| -||TensorRT (Jetson): [**Jetson Zoo**](https://elinux.org/Jetson_Zoo#ONNX_Runtime) - *NVIDIA managed*||| -||Azure (Cloud): [**onnxruntime-azure**](https://pypi.org/project/onnxruntime-azure/)||| -|C#/C/C++|CPU: [**Microsoft.ML.OnnxRuntime**](https://www.nuget.org/packages/Microsoft.ML.OnnxRuntime) |[ort-nightly (dev)](https://aiinfra.visualstudio.com/PublicPackages/_packaging?_a=feed&feed=ORT-Nightly)|| -||GPU (CUDA/TensorRT): [**Microsoft.ML.OnnxRuntime.Gpu**](https://www.nuget.org/packages/Microsoft.ML.OnnxRuntime.gpu)|[ort-nightly (dev)](https://aiinfra.visualstudio.com/PublicPackages/_packaging?_a=feed&feed=ORT-Nightly)|[View](../execution-providers/CUDA-ExecutionProvider)| -||GPU (DirectML): [**Microsoft.ML.OnnxRuntime.DirectML**](https://www.nuget.org/packages/Microsoft.ML.OnnxRuntime.DirectML)|[ort-nightly (dev)](https://aiinfra.visualstudio.com/PublicPackages/_artifacts/feed/ORT-Nightly/PyPI/ort-nightly-directml/overview)|[View](../execution-providers/DirectML-ExecutionProvider)| -|WinML|[**Microsoft.AI.MachineLearning**](https://www.nuget.org/packages/Microsoft.AI.MachineLearning)|[ort-nightly (dev)](https://aiinfra.visualstudio.com/PublicPackages/_artifacts/feed/ORT-Nightly/NuGet/Microsoft.AI.MachineLearning/overview)|[View](https://docs.microsoft.com/en-us/windows/ai/windows-ml/port-app-to-nuget#prerequisites)| -|Java|CPU: [**com.microsoft.onnxruntime:onnxruntime**](https://search.maven.org/artifact/com.microsoft.onnxruntime/onnxruntime)||[View](../api/java)| -||GPU (CUDA/TensorRT): [**com.microsoft.onnxruntime:onnxruntime_gpu**](https://search.maven.org/artifact/com.microsoft.onnxruntime/onnxruntime_gpu)||[View](../api/java)| -|Android|[**com.microsoft.onnxruntime:onnxruntime-mobile**](https://search.maven.org/artifact/com.microsoft.onnxruntime/onnxruntime-mobile) ||[View](../install/index.md#install-on-ios)| -|iOS (C/C++)|CocoaPods: **onnxruntime-mobile-c**||[View](../install/index.md#install-on-ios)| -|Objective-C|CocoaPods: **onnxruntime-mobile-objc**||[View](../install/index.md#install-on-ios)| -|React Native|[**onnxruntime-react-native** (latest)](https://www.npmjs.com/package/onnxruntime-react-native)|[onnxruntime-react-native (dev)](https://www.npmjs.com/package/onnxruntime-react-native?activeTab=versions)|[View](../api/js)| -|Node.js|[**onnxruntime-node** (latest)](https://www.npmjs.com/package/onnxruntime-node)|[onnxruntime-node (dev)](https://www.npmjs.com/package/onnxruntime-node?activeTab=versions)|[View](../api/js)| -|Web|[**onnxruntime-web** (latest)](https://www.npmjs.com/package/onnxruntime-web)|[onnxruntime-web (dev)](https://www.npmjs.com/package/onnxruntime-web?activeTab=versions)|[View](../api/js)| - - - -*Note: Dev builds created from the master branch are available for testing newer changes between official releases. Please use these at your own risk. We strongly advise against deploying these to production workloads as support is limited for dev builds.* +The table below lists the build variants available as officially supported packages. Others can +be [built from source](../build/inferencing) from each [release branch](https://github.com/microsoft/onnxruntime/tags). +In addition to general [requirements](#requirements), please note additional requirements and dependencies in the table +below: +| | Official build | Nightly build | Reqs | +|--------------|---------------------------------------------------------------------------------------------------------------------------------------------------|-----------------------------------------------------------------------------------------------------------------------------------------------|------------------------------------------------------------------------------------------------| +| Python | If using pip, run `pip install --upgrade pip` prior to downloading. | | | +| | CPU: [**onnxruntime**](https://pypi.org/project/onnxruntime) | [ort-nightly (dev)](https://aiinfra.visualstudio.com/PublicPackages/_artifacts/feed/ORT-Nightly/PyPI/ort-nightly/overview) | | +| | GPU (CUDA/TensorRT): [**onnxruntime-gpu**](https://pypi.org/project/onnxruntime-gpu) | [ort-nightly-gpu (dev)](https://aiinfra.visualstudio.com/PublicPackages/_artifacts/feed/ORT-Nightly/PyPI/ort-nightly-gpu/overview/) | [View](../execution-providers/CUDA-ExecutionProvider.md#requirements) | +| | GPU (DirectML): [**onnxruntime-directml**](https://pypi.org/project/onnxruntime-directml/) | [ort-nightly-directml (dev)](https://aiinfra.visualstudio.com/PublicPackages/_artifacts/feed/ORT-Nightly/PyPI/ort-nightly-directml/overview/) | [View](../execution-providers/DirectML-ExecutionProvider.md#requirements) | +| | OpenVINO: [**intel/onnxruntime**](https://github.com/intel/onnxruntime/releases/latest) - *Intel managed* | | [View](../build/eps.md#openvino) | +| | TensorRT (Jetson): [**Jetson Zoo**](https://elinux.org/Jetson_Zoo#ONNX_Runtime) - *NVIDIA managed* | | | +| | Azure (Cloud): [**onnxruntime-azure**](https://pypi.org/project/onnxruntime-azure/) | | | +| C#/C/C++ | CPU: [**Microsoft.ML.OnnxRuntime**](https://www.nuget.org/packages/Microsoft.ML.OnnxRuntime) | [ort-nightly (dev)](https://aiinfra.visualstudio.com/PublicPackages/_packaging?_a=feed&feed=ORT-Nightly) | | +| | GPU (CUDA/TensorRT): [**Microsoft.ML.OnnxRuntime.Gpu**](https://www.nuget.org/packages/Microsoft.ML.OnnxRuntime.gpu) | [ort-nightly (dev)](https://aiinfra.visualstudio.com/PublicPackages/_packaging?_a=feed&feed=ORT-Nightly) | [View](../execution-providers/CUDA-ExecutionProvider) | +| | GPU (DirectML): [**Microsoft.ML.OnnxRuntime.DirectML**](https://www.nuget.org/packages/Microsoft.ML.OnnxRuntime.DirectML) | [ort-nightly (dev)](https://aiinfra.visualstudio.com/PublicPackages/_artifacts/feed/ORT-Nightly/PyPI/ort-nightly-directml/overview) | [View](../execution-providers/DirectML-ExecutionProvider) | +| WinML | [**Microsoft.AI.MachineLearning**](https://www.nuget.org/packages/Microsoft.AI.MachineLearning) | [ort-nightly (dev)](https://aiinfra.visualstudio.com/PublicPackages/_artifacts/feed/ORT-Nightly/NuGet/Microsoft.AI.MachineLearning/overview) | [View](https://docs.microsoft.com/en-us/windows/ai/windows-ml/port-app-to-nuget#prerequisites) | +| Java | CPU: [**com.microsoft.onnxruntime:onnxruntime**](https://search.maven.org/artifact/com.microsoft.onnxruntime/onnxruntime) | | [View](../api/java) | +| | GPU (CUDA/TensorRT): [**com.microsoft.onnxruntime:onnxruntime_gpu**](https://search.maven.org/artifact/com.microsoft.onnxruntime/onnxruntime_gpu) | | [View](../api/java) | +| Android | [**com.microsoft.onnxruntime:onnxruntime-mobile**](https://search.maven.org/artifact/com.microsoft.onnxruntime/onnxruntime-mobile) | | [View](../install/index.md#install-on-ios) | +| iOS (C/C++) | CocoaPods: **onnxruntime-mobile-c** | | [View](../install/index.md#install-on-ios) | +| Objective-C | CocoaPods: **onnxruntime-mobile-objc** | | [View](../install/index.md#install-on-ios) | +| React Native | [**onnxruntime-react-native** (latest)](https://www.npmjs.com/package/onnxruntime-react-native) | [onnxruntime-react-native (dev)](https://www.npmjs.com/package/onnxruntime-react-native?activeTab=versions) | [View](../api/js) | +| Node.js | [**onnxruntime-node** (latest)](https://www.npmjs.com/package/onnxruntime-node) | [onnxruntime-node (dev)](https://www.npmjs.com/package/onnxruntime-node?activeTab=versions) | [View](../api/js) | +| Web | [**onnxruntime-web** (latest)](https://www.npmjs.com/package/onnxruntime-web) | [onnxruntime-web (dev)](https://www.npmjs.com/package/onnxruntime-web?activeTab=versions) | [View](../api/js) | +*Note: Dev builds created from the master branch are available for testing newer changes between official releases. +Please use these at your own risk. We strongly advise against deploying these to production workloads as support is +limited for dev builds.* ## Training install table for all languages -Refer to the getting started with [Optimized Training](https://onnxruntime.ai/getting-started) page for more fine-grained installation instructions. +Refer to the getting started with [Optimized Training](https://onnxruntime.ai/getting-started) page for more +fine-grained installation instructions.