onnxruntime

mirror of https://github.com/saymrwulf/onnxruntime.git synced 2026-07-04 04:07:22 +00:00

Author	SHA1	Message	Date
Jing Fang	7fa69461fd	[ARM] MatMulNBits FP16 support - kernels only (#22806 ) ### Description A break down PR of https://github.com/microsoft/onnxruntime/pull/22651 Add fp16 kernels. ### Motivation and Context <!-- - Why is this change required? What problem does it solve? - If it fixes an open issue, please link to the issue here. -->	2024-11-12 14:28:47 -08:00
zz002	d3ad76b2cf	[VitisAI] Cache node subgraph when necessary (#22073 ) ### Description <!-- Describe your changes. --> [VitisAI] Cache node subgraph when necessary ### Motivation and Context <!-- - Why is this change required? What problem does it solve? - If it fixes an open issue, please link to the issue here. --> --------- Co-authored-by: Zhenze Wang <zhenzew@xilinx.com> Co-authored-by: zhenzew <zhenzew@amd.com>	2024-11-08 23:17:16 -08:00
Ranjit Ranjan	193671295e	[AIX] Fix for AIX build break (#22745 ) ### Description With recent changes, below build error is found under AIX. ``` ld: 0706-012 The -p flag is not recognized. ld: 0706-012 The -a flag is not recognized. ld: 0706-012 The -t flag is not recognized. ld: 0706-012 The -h flag is not recognized. ld: 0706-012 The -= flag is not recognized. ld: 0706-012 The -$ flag is not recognized. ld: 0706-012 The -$ flag is not recognized. ld: 0706-012 The -O flag is not recognized. ld: 0706-027 The -R IGIN flag is ignored. collect2: error: ld returned 255 exit status ``` ### Motivation and Context AIX linker doesn't support -rpath option , so blocking this option under AIX.	2024-11-07 13:22:22 -08:00
Yifan Li	3b7a6eba69	[TensorRT EP] support TensorRT 10.6-GA (#22644 ) ### Description <!-- Describe your changes. --> * Update CI with TRT 10.6 * Update oss parser to [10.6-GA-ORT-DDS ](https://github.com/onnx/onnx-tensorrt/tree/10.6-GA-ORT-DDS) and update dependency version * Update Py-cuda11 CI to use trt10.6 ### Motivation and Context <!-- - Why is this change required? What problem does it solve? - If it fixes an open issue, please link to the issue here. --> (There will be 3rd PR to further reduce trt_version hardcoding)	2024-11-06 14:33:46 -08:00
Tianlei Wu	72186bbb71	[CUDA] Build nhwc ops by default (#22648 ) ### Description * Build cuda nhwc ops by default. * Deprecate `--enable_cuda_nhwc_ops` in build.py and add `--disable_cuda_nhwc_ops` option Note that it requires cuDNN 9.x. If you build with cuDNN 8, NHWC ops will be disabled automatically. ### Motivation and Context In general, NHWC is faster than NCHW for convolution in Nvidia GPUs with Tensor Cores, and this could improve performance for vision models. This is the first step to prefer NHWC for CUDA in 1.21 release. Next step is to do some tests on popular vision models. If it help in most models and devices, set `prefer_nhwc=1` as default cuda provider option.	2024-11-06 09:54:55 -08:00
Changming Sun	66980e4646	Refactor the cmake code that is related to delay loading (#22646 ) ### Description Refactor the cmake code that is related to delay loading. Provide a cmake option to control if delay loading should be enabled or not. Disabling the option when python is enabled, due to a known issue. ### Motivation and Context ONNX Runtime's python package depends on DirectML.dll, but supposedly the DLL should be delay loaded. This PR only refactor the code. It doesn't change the behavior.	2024-11-04 16:30:50 -08:00
Yulong Wang	7a8fa12850	Add implementation of WebGPU EP (#22591 ) ### Description This PR adds the actual implementation of the WebGPU EP based on https://github.com/microsoft/onnxruntime/pull/22318. This change includes the following: <details> <summary><b>core framework of WebGPU EP</b></summary> - WebGPU EP factory classes for: - handling WebGPU options - creating WebGPU EP instance - creating WebGPU context - WebGPU Execution Provider classes - GPU Buffer allocator - data transfer - Buffer management classes - Buffer Manager - BufferCacheManager - DisabledCacheManager - SimpleCacheManager - LazyReleaseCacheManager - BucketCacheManager - Program classes - Program (base) - Program Cache Key - Program Manager - Shader helper classes - Shader Helper - ShaderIndicesHelper - ShaderVariableHelper - Utils - GPU Query based profiler - compute context - string utils - Miscs - Python binding webgpu support (basic) </details> <details> <summary><b>Kernel implementation</b></summary> - onnx.ai (default opset): - Elementwise (math): Abs, Neg, Floor, Ceil, Reciprocal, Sqrt, Exp, Erf, Log, Sin, Cos, Tan, Asin, Acos, Atan, Sinh, Cosh, Asinh, Acosh, Atanh, Tanh, Not, Cast - Elementwise (activation): Sigmoid, HardSigmoid, Clip, Elu, Relu, LeakyRelu, ThresholdedRelu, Gelu - Binary (math): Add, Sub, Mul, Div, Pow, Equal, Greater, GreaterOrEqual, Less, LessOrEqual - (Tensors): Shape, Reshape, Squeeze, Unsqueeze - Where - Transpose - Concat - Expand - Gather - Tile - Range - LayerNormalization - com.microsoft - FastGelu - MatMulNBits - MultiHeadAttention - RotaryEmbedding - SkipLayerNormalization - LayerNormalization - SimplifiedLayerNormalization - SkipSimplifiedLayerNormalization </details> <details> <summary><b>Build, test and CI pipeline integration</b></summary> - build works for Windows, macOS and iOS - support onnxruntime_test_all and python node test - added a new unit test for `--use_external_dawn` build flag. - updated MacOS pipeline to build with WebGPU support - added a new pipeline for WebGPU Windows </details> This change does not include: - Node.js binding support for WebGPU (will be a separate PR)	2024-10-29 18:29:40 -07:00
Indy Zhu	e2e837584f	[DML EP] Update DML to 1.15.4 (#22635 ) ### Description [DML EP] Update DML to 1.15.4 ### Motivation and Context <!-- - Why is this change required? What problem does it solve? - If it fixes an open issue, please link to the issue here. --> We want the customer to use the latest DirectML.	2024-10-29 17:13:57 -07:00
Tianlei Wu	b4afc6266f	[ROCm] Python 3.10 in ROCm CI, and ROCm 6.2.3 in MigraphX CI (#22527 ) ### Description Upgrade python from 3.9 to 3.10 in ROCm and MigraphX docker files and CI pipelines. Upgrade ROCm version to 6.2.3 in most places except ROCm CI, see comment below. Some improvements/upgrades on ROCm/Migraphx docker or pipeline: * rocm 6.0/6.1.3 => 6.2.3 * python 3.9 => 3.10 * Ubuntu 20.04 => 22.04 * Also upgrade ml_dtypes, numpy and scipy packages. * Fix message "ROCm version from ..." with correct file path in CMakeList.txt * Exclude some NHWC tests since ROCm EP lacks support for NHWC convolution. #### ROCm CI Pipeline: ROCm 6.1.3 is kept in the pipeline for now. - Failed after upgrading to ROCm 6.2.3: `HIPBLAS_STATUS_INVALID_VALUE ; GPU=0 ; hostname=76123b390aed ; file=/onnxruntime_src/onnxruntime/core/providers/rocm/rocm_execution_provider.cc ; line=170 ; expr=hipblasSetStream(hipblas_handle_, stream);` . It need further investigation. - cupy issues: (1) It currently supports numpy < 1.27, might not work with numpy 2.x. So we locked numpy==1.26.4 for now. (2) cupy support of ROCm 6.2 is still in progress: https://github.com/cupy/cupy/issues/8606. Note that miniconda issues: its libstdc++.so.6 and libgcc_s.so.1 might have conflict with the system ones. So we created links to use the system ones. #### MigraphX CI pipeline MigraphX CI does not use cupy, and we are able to use ROCm 6.2.3 and numpy 2.x in the pipeline. #### Other attempts Other things that I've tried which might help in the future: Attempt to use a single docker file for both ROCm and Migraphx: https://github.com/microsoft/onnxruntime/pull/22478 Upgrade to ubuntu 24.04 and python 3.12, and use venv like [this](`27903e7ff1/tools/ci_build/github/linux/docker/rocm-ci-pipeline-env.Dockerfile`). ### Motivation and Context In 1.20 release, ROCm nuget packaging pipeline will use 6.2: https://github.com/microsoft/onnxruntime/pull/22461. This upgrades rocm to 6.2.3 in CI pipelines to be consistent.	2024-10-25 11:47:16 -07:00
Satya Kumar Jandhyala	4ed5bec2e7	[JS/WebGPU] Support WASM64 (#21836 ) ### Description Support wasm64 ### Motivation and Context Overcome memory limitations --------- Co-authored-by: Yulong Wang <7679871+fs-eire@users.noreply.github.com>	2024-10-24 20:21:51 -07:00
Changming Sun	88676e62b9	Remove nsync (#20413 ) ### Description 1. Remove the onnxruntime::OrtMutex class and replace it with ~absl::Mutex~ std::mutex. 2. After this change, most source files will not include <Windows.h> indirectly. ### Motivation and Context To reduce the number of deps we have, and address some Github issues that are related to build ONNX Runtime from source. In PR #3000 , I added a custom implementation of std::mutex . It was mainly because at that time std::mutex's default constructor was not trivial on Windows. If you had such a mutex as a global var, it could not be initialized at compile time. Then VC++ team fixed this issue. Therefore we don't need this custom implementation anymore. This PR also removes nsync. I ran several models tests on Linux. I didn't see any perf difference. This PR also reverts PR #21005 , which is no longer needed since conda has updated its msvc runtime DLL. This PR unblocks #22173 and resolves #22092 . We have a lot of open issues with nsync. This PR can resolve all of them.	2024-10-21 15:32:14 -07:00
Jeff Daily	5aabc53121	[ROCm] redo hipify of version controlled files (#22449 ) ### Description Updates the ROCm EP opsets to match the current CUDA EP opsets. Also enable the test CApiTest.basic_cuda_graph_with_annotation. Note that some changes are whitespace-only. These changes were made to improve the comparison of corresponding ROCm and CUDA EP source files when using a side by side diff tool. ### Motivation and Context The ROCm EP derives from the CUDA EP. Many source files are shared between the EPs and "hipified" during the ROCm EP build, however quite a few files within the ROCm EP are under source control after their initial hipification. Over time these ROCm EP files get stale relative to their CUDA EP counterparts. It becomes necessary to re-hipify these otherwise static files in order to pick up important changes such as opset differences.	2024-10-18 12:40:54 -07:00
Edward Chen	7964d3aef6	Specify iOS simulator runtime version (#22474 ) - Allow specification of iOS simulator runtime version to use. - Pick simulator runtime version (iphonesimulator 16.4) that is supported by the Xcode version (14.3.1) that we use. - Disable CoreML EP's DepthToSpace op support for CoreML version less than 7, with DCR mode, and FP16 input. It doesn't produce the correct output in this case. - Some cleanup of iOS test infrastructure.	2024-10-18 09:26:06 -07:00
Jeff Daily	8c21680ffc	[ROCm] prefer hip interfaces over roc during hipify (#22394 ) ### Description Change the hipify step to remove the -roc option to hipify-perl. This will prefer hipblas over rocblas. rocblas can still be called directly such as in TunableOp. ### Motivation and Context hip interfaces are preferred over roc for porting from cuda to hip. Calling roc interfaces is meant for ROCm-specific enhancements or extensions.	2024-10-14 20:34:03 -07:00
amarin16	7d17c466ec	Add microbenchmark for layer normalization and improve latency (#22223 ) - Added a microbenchmark for the `LayerNormalization` MLFloat16 support added in https://github.com/microsoft/onnxruntime/pull/22063. - Updated the `LayerNormalization` MLFloat16 implementation to improve the latency. ``` ---------------------------------------------------------------------------------------------- Original MLFloat16 support Time CPU Iterations ---------------------------------------------------------------------------------------------- BM_LayerNormalization<MLFloat16, float>/1/real_time 15599 us 15625 us 47 BM_LayerNormalization<MLFloat16, float>/1/real_time 14714 us 14824 us 39 BM_LayerNormalization<MLFloat16, float>/1/real_time 14634 us 14688 us 50 ---------------------------------------------------------------------------------------------- Updated MLFloat16 support Time CPU Iterations ---------------------------------------------------------------------------------------------- BM_LayerNormalization<MLFloat16, float>/1/real_time 7276 us 7254 us 84 BM_LayerNormalization<MLFloat16, float>/1/real_time 6820 us 6720 us 93 BM_LayerNormalization<MLFloat16, float>/1/real_time 6840 us 6882 us 84 ```	2024-10-14 18:47:27 -07:00
Tianlei Wu	de93f40240	[CUDA] Lean Attention (#22352 ) ### Description Add [Lean Attention](https://arxiv.org/abs/2405.10480) and the integration with MultiHeadAttention operator for LLM in GPU. LeanAttention speeds up self-attention for the token-generation phase (decode-phase) of decoder-only transformer models, especially on long context lengths. - [x] Initial implementation of Lean Attention (by Srikant Bharadwaj) - [x] Integration with MultiHeadAttention operator - [x] Add parity tests - [x] Add benchmark #### Implementation Details (1) Lean Attention is enabled in build for Linux, and disabled for Windows (2) Lean Attention is disabled by default. Need enable it through cuda provider option sdpa_kernel, or use environment variable `ORT_ENABLE_LEAN_ATTENTION=1` (3) It only works for token-generation (sequence_length==1, past_sequence_length > 0). (4) Like flash attention, it only works in Ampere or newer GPU. We can revisit #1 and #2 after comparing with DecoderMaskedMultiHeadAttention and XQA kernels. #### Benchmark ``` cd onnxruntime/test/python/transformers /bin/bash benchmark_mha.sh lean ``` Example outputs in H100: Note that past and present does not share buffer for MHA for now, so we can see low tflops. The relative ratio will change after buffer sharing is enabled. But we expect that the order (kernel A is faster than B) will remain the same after buffer sharing is enabled. Note that common settings `sequence_length=1; causal=True;attn_bias=None;cuda_graph=False` are not shown in the below table. batch_size \| past_sequence_length \| num_heads \| head_size \| average_latency \| tflops \| kernel -- \| -- \| -- \| -- \| -- \| -- \| -- 1 \| 512 \| 16 \| 64 \| 0.000059 \| 0.0178 \| ort:flash 1 \| 512 \| 16 \| 64 \| 0.000068 \| 0.0155 \| ort:efficient 1 \| 512 \| 16 \| 64 \| 0.000065 \| 0.0161 \| ort:math 1 \| 512 \| 16 \| 64 \| 0.000060 \| 0.0176 \| ort:lean 1 \| 512 \| 32 \| 128 \| 0.000062 \| 0.0674 \| ort:flash 1 \| 512 \| 32 \| 128 \| 0.000064 \| 0.0661 \| ort:efficient 1 \| 512 \| 32 \| 128 \| 0.000067 \| 0.0625 \| ort:math 1 \| 512 \| 32 \| 128 \| 0.000062 \| 0.0678 \| ort:lean 1 \| 1024 \| 16 \| 64 \| 0.000061 \| 0.0345 \| ort:flash 1 \| 1024 \| 16 \| 64 \| 0.000086 \| 0.0244 \| ort:efficient 1 \| 1024 \| 16 \| 64 \| 0.000065 \| 0.0322 \| ort:math 1 \| 1024 \| 16 \| 64 \| 0.000063 \| 0.0332 \| ort:lean 1 \| 1024 \| 32 \| 128 \| 0.000075 \| 0.1125 \| ort:flash 1 \| 1024 \| 32 \| 128 \| 0.000088 \| 0.0951 \| ort:efficient 1 \| 1024 \| 32 \| 128 \| 0.000079 \| 0.1068 \| ort:math 1 \| 1024 \| 32 \| 128 \| 0.000072 \| 0.1171 \| ort:lean 1 \| 2048 \| 16 \| 64 \| 0.000069 \| 0.0606 \| ort:flash 1 \| 2048 \| 16 \| 64 \| 0.000125 \| 0.0336 \| ort:efficient 1 \| 2048 \| 16 \| 64 \| 0.000064 \| 0.0655 \| ort:lean 1 \| 2048 \| 32 \| 128 \| 0.000098 \| 0.1720 \| ort:flash 1 \| 2048 \| 32 \| 128 \| 0.000132 \| 0.1270 \| ort:efficient 1 \| 2048 \| 32 \| 128 \| 0.000092 \| 0.1828 \| ort:lean 1 \| 4096 \| 16 \| 64 \| 0.000076 \| 0.1097 \| ort:flash 1 \| 4096 \| 16 \| 64 \| 0.000207 \| 0.0406 \| ort:efficient 1 \| 4096 \| 16 \| 64 \| 0.000069 \| 0.1209 \| ort:lean 1 \| 4096 \| 32 \| 128 \| 0.000140 \| 0.2394 \| ort:flash 1 \| 4096 \| 32 \| 128 \| 0.000213 \| 0.1575 \| ort:efficient 1 \| 4096 \| 32 \| 128 \| 0.000139 \| 0.2419 \| ort:lean 1 \| 8192 \| 16 \| 64 \| 0.000104 \| 0.1609 \| ort:flash 1 \| 8192 \| 16 \| 64 \| 0.000392 \| 0.0428 \| ort:efficient 1 \| 8192 \| 16 \| 64 \| 0.000093 \| 0.1809 \| ort:lean 1 \| 8192 \| 32 \| 128 \| 0.000212 \| 0.3160 \| ort:flash 1 \| 8192 \| 32 \| 128 \| 0.000360 \| 0.1866 \| ort:efficient 1 \| 8192 \| 32 \| 128 \| 0.000212 \| 0.3162 \| ort:lean 1 \| 16384 \| 16 \| 64 \| 0.000139 \| 0.2410 \| ort:flash 1 \| 16384 \| 16 \| 64 \| 0.000731 \| 0.0459 \| ort:efficient 1 \| 16384 \| 16 \| 64 \| 0.000136 \| 0.2465 \| ort:lean 1 \| 16384 \| 32 \| 128 \| 0.000361 \| 0.3722 \| ort:flash 1 \| 16384 \| 32 \| 128 \| 0.000667 \| 0.2014 \| ort:efficient 1 \| 16384 \| 32 \| 128 \| 0.000357 \| 0.3765 \| ort:lean 1 \| 32768 \| 16 \| 64 \| 0.000210 \| 0.3194 \| ort:flash 1 \| 32768 \| 16 \| 64 \| 0.001428 \| 0.0470 \| ort:efficient 1 \| 32768 \| 16 \| 64 \| 0.000209 \| 0.3211 \| ort:lean 1 \| 32768 \| 32 \| 128 \| 0.000659 \| 0.4074 \| ort:flash 1 \| 32768 \| 32 \| 128 \| 0.001270 \| 0.2114 \| ort:efficient 1 \| 32768 \| 32 \| 128 \| 0.000651 \| 0.4123 \| ort:lean 1 \| 65536 \| 16 \| 64 \| 0.000355 \| 0.3785 \| ort:flash 1 \| 65536 \| 16 \| 64 \| 0.002736 \| 0.0491 \| ort:efficient 1 \| 65536 \| 16 \| 64 \| 0.000349 \| 0.3845 \| ort:lean 1 \| 65536 \| 32 \| 128 \| 0.001251 \| 0.4290 \| ort:flash 1 \| 65536 \| 32 \| 128 \| 0.002480 \| 0.2165 \| ort:efficient 1 \| 65536 \| 32 \| 128 \| 0.001239 \| 0.4333 \| ort:lean 4 \| 512 \| 16 \| 64 \| 0.000063 \| 0.0665 \| ort:flash 4 \| 512 \| 16 \| 64 \| 0.000069 \| 0.0607 \| ort:efficient 4 \| 512 \| 16 \| 64 \| 0.000066 \| 0.0634 \| ort:math 4 \| 512 \| 16 \| 64 \| 0.000062 \| 0.0674 \| ort:lean 4 \| 512 \| 32 \| 128 \| 0.000100 \| 0.1677 \| ort:flash 4 \| 512 \| 32 \| 128 \| 0.000099 \| 0.1703 \| ort:efficient 4 \| 512 \| 32 \| 128 \| 0.000108 \| 0.1557 \| ort:math 4 \| 512 \| 32 \| 128 \| 0.000092 \| 0.1818 \| ort:lean 4 \| 1024 \| 16 \| 64 \| 0.000077 \| 0.1094 \| ort:flash 4 \| 1024 \| 16 \| 64 \| 0.000099 \| 0.0850 \| ort:efficient 4 \| 1024 \| 16 \| 64 \| 0.000081 \| 0.1038 \| ort:math 4 \| 1024 \| 16 \| 64 \| 0.000072 \| 0.1161 \| ort:lean 4 \| 1024 \| 32 \| 128 \| 0.000143 \| 0.2343 \| ort:flash 4 \| 1024 \| 32 \| 128 \| 0.000137 \| 0.2447 \| ort:efficient 4 \| 1024 \| 32 \| 128 \| 0.000150 \| 0.2245 \| ort:math 4 \| 1024 \| 32 \| 128 \| 0.000135 \| 0.2496 \| ort:lean 4 \| 2048 \| 16 \| 64 \| 0.000096 \| 0.1757 \| ort:flash 4 \| 2048 \| 16 \| 64 \| 0.000156 \| 0.1078 \| ort:efficient 4 \| 2048 \| 16 \| 64 \| 0.000089 \| 0.1892 \| ort:lean 4 \| 2048 \| 32 \| 128 \| 0.000223 \| 0.3010 \| ort:flash 4 \| 2048 \| 32 \| 128 \| 0.000217 \| 0.3101 \| ort:efficient 4 \| 2048 \| 32 \| 128 \| 0.000209 \| 0.3209 \| ort:lean 4 \| 4096 \| 16 \| 64 \| 0.000137 \| 0.2448 \| ort:flash 4 \| 4096 \| 16 \| 64 \| 0.000256 \| 0.1312 \| ort:efficient 4 \| 4096 \| 16 \| 64 \| 0.000133 \| 0.2530 \| ort:lean 4 \| 4096 \| 32 \| 128 \| 0.000389 \| 0.3450 \| ort:flash 4 \| 4096 \| 32 \| 128 \| 0.000376 \| 0.3574 \| ort:efficient 4 \| 4096 \| 32 \| 128 \| 0.000354 \| 0.3794 \| ort:lean 4 \| 8192 \| 16 \| 64 \| 0.000210 \| 0.3198 \| ort:flash 4 \| 8192 \| 16 \| 64 \| 0.000453 \| 0.1480 \| ort:efficient 4 \| 8192 \| 16 \| 64 \| 0.000206 \| 0.3260 \| ort:lean 4 \| 8192 \| 32 \| 128 \| 0.000725 \| 0.3705 \| ort:flash 4 \| 8192 \| 32 \| 128 \| 0.000693 \| 0.3874 \| ort:efficient 4 \| 8192 \| 32 \| 128 \| 0.000653 \| 0.4114 \| ort:lean 4 \| 16384 \| 16 \| 64 \| 0.000355 \| 0.3782 \| ort:flash 4 \| 16384 \| 16 \| 64 \| 0.000849 \| 0.1581 \| ort:efficient 4 \| 16384 \| 16 \| 64 \| 0.000346 \| 0.3874 \| ort:lean 4 \| 16384 \| 32 \| 128 \| 0.001395 \| 0.3848 \| ort:flash 4 \| 16384 \| 32 \| 128 \| 0.001337 \| 0.4017 \| ort:efficient 4 \| 16384 \| 32 \| 128 \| 0.001252 \| 0.4288 \| ort:lean 4 \| 32768 \| 16 \| 64 \| 0.000647 \| 0.4146 \| ort:flash 4 \| 32768 \| 16 \| 64 \| 0.001649 \| 0.1628 \| ort:efficient 4 \| 32768 \| 16 \| 64 \| 0.000639 \| 0.4204 \| ort:lean 4 \| 32768 \| 32 \| 128 \| 0.002721 \| 0.3947 \| ort:flash 4 \| 32768 \| 32 \| 128 \| 0.002601 \| 0.4128 \| ort:efficient 4 \| 32768 \| 32 \| 128 \| 0.002434 \| 0.4411 \| ort:lean 4 \| 65536 \| 16 \| 64 \| 0.001231 \| 0.4361 \| ort:flash 4 \| 65536 \| 16 \| 64 \| 0.003238 \| 0.1658 \| ort:efficient 4 \| 65536 \| 16 \| 64 \| 0.001217 \| 0.4412 \| ort:lean 4 \| 65536 \| 32 \| 128 \| 0.005357 \| 0.4009 \| ort:flash 4 \| 65536 \| 32 \| 128 \| 0.005118 \| 0.4196 \| ort:efficient 4 \| 65536 \| 32 \| 128 \| 0.004781 \| 0.4492 \| ort:lean 16 \| 512 \| 16 \| 64 \| 0.000098 \| 0.1724 \| ort:flash 16 \| 512 \| 16 \| 64 \| 0.000104 \| 0.1616 \| ort:efficient 16 \| 512 \| 16 \| 64 \| 0.000118 \| 0.1420 \| ort:math 16 \| 512 \| 16 \| 64 \| 0.000087 \| 0.1926 \| ort:lean 16 \| 512 \| 32 \| 128 \| 0.000220 \| 0.3062 \| ort:flash 16 \| 512 \| 32 \| 128 \| 0.000208 \| 0.3237 \| ort:efficient 16 \| 512 \| 32 \| 128 \| 0.000237 \| 0.2838 \| ort:math 16 \| 512 \| 32 \| 128 \| 0.000209 \| 0.3216 \| ort:lean 16 \| 1024 \| 16 \| 64 \| 0.000136 \| 0.2465 \| ort:flash 16 \| 1024 \| 16 \| 64 \| 0.000150 \| 0.2235 \| ort:efficient 16 \| 1024 \| 16 \| 64 \| 0.000148 \| 0.2266 \| ort:math 16 \| 1024 \| 16 \| 64 \| 0.000129 \| 0.2611 \| ort:lean 16 \| 1024 \| 32 \| 128 \| 0.000367 \| 0.3663 \| ort:flash 16 \| 1024 \| 32 \| 128 \| 0.000351 \| 0.3829 \| ort:efficient 16 \| 1024 \| 32 \| 128 \| 0.000400 \| 0.3357 \| ort:math 16 \| 1024 \| 32 \| 128 \| 0.000349 \| 0.3853 \| ort:lean 16 \| 2048 \| 16 \| 64 \| 0.000209 \| 0.3206 \| ort:flash 16 \| 2048 \| 16 \| 64 \| 0.000243 \| 0.2762 \| ort:efficient 16 \| 2048 \| 16 \| 64 \| 0.000201 \| 0.3338 \| ort:lean 16 \| 2048 \| 32 \| 128 \| 0.000671 \| 0.4002 \| ort:flash 16 \| 2048 \| 32 \| 128 \| 0.000645 \| 0.4163 \| ort:efficient 16 \| 2048 \| 32 \| 128 \| 0.000642 \| 0.4185 \| ort:lean 16 \| 4096 \| 16 \| 64 \| 0.000360 \| 0.3732 \| ort:flash 16 \| 4096 \| 16 \| 64 \| 0.000425 \| 0.3162 \| ort:efficient 16 \| 4096 \| 16 \| 64 \| 0.000341 \| 0.3933 \| ort:lean 16 \| 4096 \| 32 \| 128 \| 0.001292 \| 0.4156 \| ort:flash 16 \| 4096 \| 32 \| 128 \| 0.001251 \| 0.4291 \| ort:efficient 16 \| 4096 \| 32 \| 128 \| 0.001241 \| 0.4327 \| ort:lean 16 \| 8192 \| 16 \| 64 \| 0.000666 \| 0.4030 \| ort:flash 16 \| 8192 \| 16 \| 64 \| 0.000804 \| 0.3339 \| ort:efficient 16 \| 8192 \| 16 \| 64 \| 0.000627 \| 0.4283 \| ort:lean 16 \| 8192 \| 32 \| 128 \| 0.002541 \| 0.4226 \| ort:flash 16 \| 8192 \| 32 \| 128 \| 0.002454 \| 0.4376 \| ort:efficient 16 \| 8192 \| 32 \| 128 \| 0.002438 \| 0.4405 \| ort:lean 16 \| 16384 \| 16 \| 64 \| 0.001292 \| 0.4156 \| ort:flash 16 \| 16384 \| 16 \| 64 \| 0.001571 \| 0.3417 \| ort:efficient 16 \| 16384 \| 16 \| 64 \| 0.001217 \| 0.4411 \| ort:lean 16 \| 16384 \| 32 \| 128 \| 0.005042 \| 0.4260 \| ort:flash 16 \| 16384 \| 32 \| 128 \| 0.004859 \| 0.4420 \| ort:efficient 16 \| 16384 \| 32 \| 128 \| 0.004827 \| 0.4449 \| ort:lean 16 \| 32768 \| 16 \| 64 \| 0.002537 \| 0.4233 \| ort:flash 16 \| 32768 \| 16 \| 64 \| 0.003103 \| 0.3461 \| ort:efficient 16 \| 32768 \| 16 \| 64 \| 0.002385 \| 0.4501 \| ort:lean 16 \| 32768 \| 32 \| 128 \| 0.009961 \| 0.4312 \| ort:flash 16 \| 32768 \| 32 \| 128 \| 0.009605 \| 0.4472 \| ort:efficient 16 \| 32768 \| 32 \| 128 \| 0.009524 \| 0.4510 \| ort:lean 16 \| 65536 \| 16 \| 64 \| 0.005019 \| 0.4279 \| ort:flash 16 \| 65536 \| 16 \| 64 \| 0.006133 \| 0.3502 \| ort:efficient 16 \| 65536 \| 16 \| 64 \| 0.004703 \| 0.4566 \| ort:lean 16 \| 65536 \| 32 \| 128 \| 0.019746 \| 0.4350 \| ort:flash 16 \| 65536 \| 32 \| 128 \| 0.019027 \| 0.4515 \| ort:efficient 16 \| 65536 \| 32 \| 128 \| 0.018864 \| 0.4554 \| ort:lean ### Motivation and Context <!-- - Why is this change required? What problem does it solve? - If it fixes an open issue, please link to the issue here. -->	2024-10-14 14:49:37 -07:00
Vishnudas Thaniel S	35adba21c7	Ovep develop lnl 1.2 (#22424 ) ### Description Support OV2024.4 Refactor tensor initialization check for external weights Support loading OV Config OVEP: Tensor Caching fix, Fix accuracy issues Refactor device memory implementation to make it more generic ### Motivation and Context The changes are required to fix accuracy issues, support loading of OV config, support OV2024.4 --------- Co-authored-by: Eric Crawford <eric.r.crawford@intel.com> Co-authored-by: saurabhkale17 <saurabh1.kale@intel.com> Co-authored-by: Javier E. Martinez <javier.e.martinez@intel.com> Co-authored-by: sfatimar <sahar.fatima@intel.com> Co-authored-by: ankitm3k <ankit.maheshkar@intel.com> Co-authored-by: Preetha Veeramalai <preetha.veeramalai@intel.com> Co-authored-by: n1harika <niharika.sathish@intel.com> Co-authored-by: jatinwadhwa921 <110383850+jatinwadhwa921@users.noreply.github.com>	2024-10-14 12:10:01 -07:00
Edward Chen	04404ea482	Fix Xcode 16 iOS build issues (#22379 ) - Work around Xcode 16 iOS test build issue: `error: Multiple commands produce '.../PlugIns'`. - Fix link error in iOS static framework test. - Update build.py to check for the right kind of build before running iOS tests on the simulator. - Update Xcode 16 build images to 'macos-15' because that's the only image that will have Xcode 16 soon. See https://github.com/actions/runner-images/issues/10703.	2024-10-14 09:24:38 -07:00
Ted Themistokleous	572e43c5d7	[MIGraphX EP/ ROCm EP] add gfx1200, gfx1201 to CMAKE_HIP_ARCHITECTURES (#22348 ) ### Description Add additonal gfx targets for AMD GPU support ### Motivation and Context Required to integrate mainline onnxruntime support for AMD GPUs --------- Co-authored-by: Stefan Sokolovic <stsokolo@amd.com> Co-authored-by: Jeff Daily <jeff.daily@amd.com>	2024-10-11 17:31:36 -07:00
Indy Zhu	b4fb32d80d	Pick changes from onnx/onnx#6010 to support EinSum shape inference (#22376 ) ### Description <!-- Describe your changes. --> Pick up onnx/onnx#6010 to support EinSum shape inference ### Motivation and Context <!-- - Why is this change required? What problem does it solve? - If it fixes an open issue, please link to the issue here. --> This change allows EinSum operator's output shape to be inferenced so that it can run on accelerators.	2024-10-10 13:24:08 -07:00
Changming Sun	2bef89c171	Upgrade absl to the latest released version (#22365 ) ### Description Resolve #21976 . ABSL generally does not have forward/backward compatibility. Our code is only compatible with one fixed LTS version. So it's important to fix the version number there when using find_package to detect an installed version.	2024-10-09 20:21:40 -07:00
Yulong Wang	c5d28cac4d	Initial WebGPU EP checkin (#22318 ) ### Description This change introduces the WebGPU EP into ONNX Runtime. To make the PR as simple as possible, this PR excluded the following: - C API changes for WebGPU EP - actual implementation of WebGPU EP. Currently in this PR, WebGPU is a stub implementation that does not register any kernel. - Python IO Binding update - Node.js IO Binding update This PR now contains only 43 file changes (while the working branch contains 130+) and hopefully this makes it easier to review. There is going to be separated PRs for each mentioned above. Current working branch: #21904	2024-10-08 16:10:46 -07:00
Tianlei Wu	f3f33bfa05	Upgrade cutlass to 3.5.1 and cudnn frontend to 1.7.0 (#22316 ) ### Description Upgrade cutlass to 3.5.1 Upgrade cudnn_frontend to 1.7.0	2024-10-04 11:48:50 -07:00
Changming Sun	f25f3868a7	Auto regenerate LORA's fbs files (#22313 ) ### Description A left-over of PR #22046 ### Motivation and Context Right now our VCPKG pipelines are broken.	2024-10-04 10:01:19 -07:00
Ranjit Ranjan	d0ddfa9b9e	[AIX] build fix for using system install protobuf/onnx (#22302 ) ### Description Fixing merge issue occurred in https://github.com/microsoft/onnxruntime/pull/22272 ### Motivation and Context To build onnxruntime using system installed protobuf/onnx.	2024-10-03 19:29:42 -07:00
Edward Chen	f1be92faf0	Patch fp16 to fix Xcode 16 builds with XNNPACK EP targeting x86_64. (#22294 )	2024-10-03 14:17:15 -07:00
Dmitri Smirnov	224f0651d0	[C#] Expose Multi-Lora support in C# (#22281 ) ### Description ### Motivation and Context https://github.com/microsoft/onnxruntime/pull/22046	2024-10-02 10:00:43 -07:00
Edward Chen	c24e55b1f1	[Java] Add API for appending QNN EP (#22208 ) - Add Java API for appending QNN EP - Update Java unit test setup - Fix issues with setting system properties for tests - Unify Windows/non-Windows setup to simplify	2024-10-01 10:18:04 -07:00
Yufeng Li	96e9c99dce	remove neural-speed (#22236 ) ### Description <!-- Describe your changes. --> NS is not developed anymore and ORT doesn't use it for int4 inference either. Remove it to clean up the code ### Motivation and Context <!-- - Why is this change required? What problem does it solve? - If it fixes an open issue, please link to the issue here. -->	2024-10-01 09:50:44 -07:00
Dmitri Smirnov	d9de054eb5	Multi-Lora support (#22046 ) ### Description <!-- Describe your changes. --> ### Motivation and Context <!-- - Why is this change required? What problem does it solve? - If it fixes an open issue, please link to the issue here. -->	2024-09-30 15:59:07 -07:00
Ranjit Ranjan	812075731c	[AIX] Build fix for using system installed protobuf/onnx (#22272 ) ### Description To fix the build issues for AIX OS while using system installed protobuf/onnx. ### Motivation and Context Code changes in this PR contains: 1. Fix for below compilation issue. ``` collect2: fatal error: library liblibprotobuf-lite not found compilation terminated. ``` 2. Adding onnx library into dependency list for test applicaitons.	2024-09-30 12:36:21 -07:00
Sumit Agarwal	529835cc46	[DML EP] Update DML to 1.15.2 (#22247 ) ### Description Update DML binary to the current latest redist version [1.15.2](https://www.nuget.org/packages/Microsoft.AI.DirectML/1.15.2).	2024-09-27 13:20:29 -07:00
Jing Fang	1942e40e05	[ARM64] MatMulNBits: use neon instrinsics to convert between fp16 and fp32 (#22195 ) ### Description For fp16 Atype, the fallback operation is convert the data to fp32 and calculate. Added neon intrinsics version to speed up the conversion. Store address alignment and loop unrolling have insignificant impact on latency so they are omitted. \|Benchmark \| Time \| CPU \| \|--------------\|---------------------------------------------\|--------------------\| \|M_ConvertF16ToF32/baseline/real_time \| 1076961 ns \| 1083398 ns \| \|M_ConvertF16ToF32/aligned:0/real_time \| 46785 ns \| 46516 ns \| \|M_ConvertF16ToF32/aligned:1/real_time \| 46631 ns \| 46391 ns \| \|M_ConvertF16ToF32_unroll2/aligned:0/real_time \| 44074 ns \| 44392 ns \| \|M_ConvertF16ToF32_unroll2/aligned:1/real_time \| 44726 ns \| 45226 ns \| \|M_ConvertF32ToF16/baseline/real_time \| 520109 ns \| 527329 ns \| \|M_ConvertF32ToF16/aligned:0/real_time \| 73610 ns \| 74015 ns \| \|M_ConvertF32ToF16/aligned:1/real_time \| 71557 ns \| 71525 ns \| \|M_ConvertF32ToF16_unroll2/aligned:0/real_time \| 64227 ns \| 63374 ns \| \|M_ConvertF32ToF16_unroll2/aligned:1/real_time \| 67428 ns \| 67989 ns \| ### Motivation and Context speed up fallback implementation of Fp16 MatMulNBits	2024-09-26 13:55:40 -07:00
jingyanwangms	d0b0ecfdb9	[Running CI] Update TensorRT to 10.4 (#22049 ) ### Description TensorRT 10.4 is GA now, update to 10.4 ### Motivation and Context <!-- - Why is this change required? What problem does it solve? - If it fixes an open issue, please link to the issue here. -->	2024-09-26 11:10:52 -07:00
Edward Chen	209ff86d52	Get build working on Xcode 16 (#22168 )	2024-09-24 08:33:03 -07:00
Hann Wang	7a782b7213	[ROCm] fix rocm-6.2 build issues (#21993 ) Composable Kernel build fails under ROCm 6.2. This PR patches Composable Kernel the same way as https://github.com/ROCm/composable_kernel/pull/1346 * fix buffer resource to match "s" constraint * add missing memory clobber	2024-09-23 14:01:54 -07:00
Chester Liu	9b37b3ea44	Specify the paths of system tools when building Apple framework (#22056 ) ### Description <!-- Describe your changes. --> Specify the path of `ar`, `ld` and `libtool` when building apple framework. ### Motivation and Context <!-- - Why is this change required? What problem does it solve? - If it fixes an open issue, please link to the issue here. --> Sometimes non-system executables will comes before the system-provided ones. This PR intends to prevent it from happening.	2024-09-23 17:19:30 +08:00
Yi Zhang	8d2d40781c	set CMAKE_SYSTEM_PROCESSOR in xnnpack.cmake (#22155 ) ### Description <!-- Describe your changes. --> ### Motivation and Context By default, CMAKE_SYSTEM_PROCESSOR is same CMAKE_HOST_SYSTEM_PROCESSOR https://cmake.org/cmake/help/latest/variable/CMAKE_SYSTEM_PROCESSOR.html KleidiAI uses CMAKE_SYSTEM_PROCESSOR to determine whether to include some arm64 ukernels. https://gitlab.arm.com/kleidi/kleidiai/-/blob/main/CMakeLists.txt#L134 We use Mac with Intel CPU to cross compile MAC with ARM in ios packaging pipeline So we need to make CMAKE_SYSTEM_PROCESSOR same with ORT_TARGET_PROCESSOR	2024-09-20 15:19:26 -07:00
Scott McKay	bd60add8ce	Update nuget.exe used in WindowsAI nuget packaging so `readme` property is supported. (#22141 ) ### Description <!-- Describe your changes. --> Use the latest nuget.exe for the `readme` property to be supported. ### Motivation and Context <!-- - Why is this change required? What problem does it solve? - If it fixes an open issue, please link to the issue here. --> #22137	2024-09-19 19:06:47 +10:00
Scott McKay	99ee6eeca2	Enable Android 16 KB page size support (#22076 ) ### Description <!-- Describe your changes. --> Add linker flags to support 16KB page size support on Android. See https://source.android.com/docs/core/architecture/16kb-page-size/16kb#build-lib-16kb-alignment ### Motivation and Context <!-- - Why is this change required? What problem does it solve? - If it fixes an open issue, please link to the issue here. --> #21837	2024-09-19 18:53:57 +10:00
George Wu	944d87381d	[QNN EP] set up py packaging pipeline for Linux x64 (#22132 ) set up a pipeline to produce nightly Linux x64 whls for onnxruntime-qnn this can be used for offline context binary generation.	2024-09-18 23:24:32 -07:00
Tianlei Wu	a9740d6f96	Add onnx export script for segment anything v2 (#22119 ) ### Description Add ONNX export script for segment anything v2 (SAM2). ### Limitations * Does not support video. Only support image right now. * The decoder does not support batch inference. ### Credits The demo that is based on [SAM2 notebook](https://github.com/facebookresearch/segment-anything-2/blob/main/notebooks/image_predictor_example.ipynb), and modified to run with ORT. The export of decoder is inspired by https://github.com/vietanhdev/samexporter. ### Demo Example output of demo: ![sam2_demo](https://github.com/user-attachments/assets/9a9fa360-8c20-482e-9935-a7aba9cf15de) ### Motivation and Context For support optimization of SAM2 image segmentation.	2024-09-18 14:31:59 -07:00
Yi Zhang	b94ba09e4f	Upgrade XNNPACK to latest version (#22012 ) ### Description Update XNNPack to latest version (Sep 4) - Some op outputs are changed, channel or stride paras are moved into reshape func. e.g. `96962a602d` - input params of xnnpack's resize related function are changed a lot - KleidiAI is added as a dependency in ARM64 - The latest XNNPACK includes 2 static libs microkernels-prod and xnnpack. Without microkernels-prod, it throws the exception of Undefined symbols. - Add ORT_TARGET_PROCESSOR to get the real processor target in CMake	2024-09-17 10:12:16 -07:00
liqun Fu	a89bddd5c2	Matmul_nbits kernel for mlas sqnbits to support Fp16 inputs (#21807 )	2024-09-13 14:55:08 -07:00
Michael Tyler	904b850b44	Update Arm Compute Library Execution Provider (#22032 ) ### Description This PR makes the following updates to the Arm Compute Library execution provider: - Target Arm Compute Library 24.07 - Add support for the following operators: - Conv (FP16) - NhwcConv - QLinearConv - MatMul - FusedMatMul - MatMulIntegerToFloat - Optimize memory usage and performance - Expose the enable_fast_math setting - Use the main runtime thread pool ### Motivation and Context These updates improve performance and memory usage, and enable use of a more recent version of Arm Compute Library. @microsoft-github-policy-service agree company="Arm Ltd" --------- Signed-off-by: Michael Tyler <michael.tyler@arm.com>	2024-09-12 20:51:59 -07:00
0xdr3dd	5c361106e6	[Fuzzer] Add two new ORT libfuzzer (Linux clang support for now) (#22055 ) ### Description This PR adds two new libfuzzer in fuzzer project. 1. Binary libfuzzer 2. libprotobuf-fuzzer To compile run below cmd on linux: ``` LLVM_PROFILE_FILE="%p.profraw" CFLAGS="-g -fsanitize=address,fuzzer-no-link -shared-libasan -fprofile-instr-generate -fcoverage-mapping" CXXFLAGS="-g -shared-libasan -fsanitize=address,fuzzer-no-link -fprofile-instr-generate -fcoverage-mapping" CC=clang CXX=clang++ ./build.sh --update --build --config Debug --compile_no_warning_as_error --build_shared_lib --skip_submodule_sync --use_full_protobuf --parallel --fuzz_testing --build_dir build/ ``` Run fuzzer: ``` LD_PRELOAD=$(clang -print-file-name=libclang_rt.asan-x86_64.so) build/Debug/onnxruntime_libfuzzer_fuzz testinput -rss_limit_mb=8196 -max_total_time=472800 -fork=2 -jobs=4 -workers=4 -ignore_crashes=1 -max_len=2097152 2>&1 \| grep -v "\[libprotobuf ERROR" ``` ### Motivation and Context The existing custom fuzzer is not coverage guided and it's slow and it will work on one model mutation at a time. The new fuzzers are coverage guided, and we can use more models' files as a corpus to increase the coverage.	2024-09-12 11:50:34 -07:00
wangshuai09	d539c27de8	Fix version check for using -mavxvnni (#21616 ) ### Description <!-- Describe your changes. --> Change the `CMAKE_CXX_COMPILER_VERSION` greater than `11` for using '-mavxvnni'. ### Motivation and Context <!-- - Why is this change required? What problem does it solve? - If it fixes an open issue, please link to the issue here. --> `CMakeFiles/onnxruntime_mlas.dir/root/Git.d/onnxruntime/onnxruntime/core/mlas/lib/x86_64/QgemmU8S8KernelAvx2.S.o cc: error: unrecognized command-line option ‘-mavxvnni’; did you mean ‘-mavx512vnni’?` using `gcc (GCC) 10.3.1`. `-mavxnni` is supported since [GCC 11 Release](https://gcc.gnu.org/gcc-11/changes.html), this PR change the version check.	2024-09-12 11:42:17 -07:00
sfatimar	0309c5f02f	Ovep release lnl 1.2.1 (#22027 ) Error Codes are added to catch compilation error and signal recompile. Remote Tensors are added to ensure direct memory access for NPU inferencing. UMD Bypass cache enabled with 2024.4 will eliminate need to disk caching ### Motivation and Context The changes are needed to ensure backward compatibility UMD Bypass caching eliminates driver caching Remote Tensors lead to performance improvement with inferencing on NPU --------- Co-authored-by: Preetha Veeramalai <preetha.veeramalai@intel.com> Co-authored-by: Srirammaswamy <srirammaswamy.s@intel.com> Co-authored-by: saurabh <saurabh1.kale@intel.com> Co-authored-by: Javier E. Martinez <javier.e.martinez@intel.com> Co-authored-by: Eric Crawford <eric.r.crawford@intel.com> Co-authored-by: jatinwadhwa921 <jatin.wadhwa@intel.com>	2024-09-11 14:55:40 -07:00
PARK DongHa	f633caa0b1	Create CMake option `onnxruntime_USE_VCPKG` (#21348 ) ### Changes 1. CMake option `onnxruntime_USE_VCPKG`. It will be used in the vcpkg port * Unit test may fail because this option leads to a mixture of unexpected external library versions. Especially ONNX, Protobuf, and Flatbuffers version can be different 2. Overhaul of `onnxruntime_external_deps.cmake` * Make `FetchContent_Declare` to try `find_package`. See https://cmake.org/cmake/help/latest/guide/using-dependencies/index.html * Relocated `FetchContent_Declare` and `FetchContent_MakeAvailable`(or `onnxruntime_fetchcontent_makeavailable`) to closer lines. It was too hard to navigate the entire file to search related sections... * Alias `IMPORTED` targets like build targets (e.g. `ONNX::onnx` --> `onnx`) ```cmake # The script uses `find_package` with the changes. # In this case, use vcpkg to search dependencies # See https://cmake.org/cmake/help/latest/guide/using-dependencies/index.html include(external/onnxruntime_external_deps.cmake) ``` 3. Create CMakePresets.json and presets to [run vcpkg in manifest mode](https://learn.microsoft.com/en-us/vcpkg/concepts/manifest-mode) * Currently, it's NOT for training build * Main triplets are `x64-windows` and `x64-osx` ```pwsh Push-Location "cmake" cmake --preset "x64-windows-vcpkg" cmake --build --preset "x64-windows-vcpkg-debug" Pop-Location ``` ```bash pushd "cmake" cmake --preset "x64-osx-vcpkg" cmake --build --preset "x64-osx-vcpkg-debug" popd ``` 4. Updated tools/ci_build/build.py * `--use_vcpkg` option: it needs `CMAKE_TOOLCHAIN_FILE` with [vcpkg.cmake toolchain script](https://github.com/microsoft/vcpkg/blob/master/scripts/buildsystems/vcpkg.cmake) * `--compile_no_warning_as_error` is recommended because library version differences will cause unexpected compiler warnings ```bash python ./tools/ci_build/build.py \ --compile_no_warning_as_error \ --use_vcpkg \ --cmake_extra_defines "CMAKE_TOOLCHAIN_FILE:FILEPATH=${VCPKG_ROOT}/scripts/buildsystems/vcpkg.cmake" \ --cmake_extra_defines "VCPKG_TARGET_TRIPLET=..." ``` 5. Created Job `Vcpkg` for Windows and macOS * Show how to setup and use vcpkg. Similar to the CMakePresets.json usage ### Motivation and Context * Help #7150 * Help https://github.com/microsoft/vcpkg/pull/36850 * https://github.com/luncliff/vcpkg-registry/pull/212 * https://github.com/microsoft/vcpkg/pull/39881 * https://github.com/luncliff/vcpkg-registry/pull/215 * https://github.com/luncliff/vcpkg-registry/pull/216 * https://github.com/luncliff/vcpkg-registry/pull/227 * https://cmake.org/cmake/help/latest/guide/using-dependencies/index.html * https://github.com/microsoft/vcpkg/blob/master/scripts/buildsystems/vcpkg.cmake ### Future Works? More feature coverage with the vcpkg supported libraries * CUDA feature support * Training feature support	2024-09-10 16:39:27 -07:00
Erick Muñoz	7489bfee53	Enable AVX NE CONVERT for FP16 to FP32 cast (#21183 ) ### Description Implementation of a new cast assembly kernel that uses AVX_NE_CONVERT instructions to accelerate casting from FP16 to FP32. Added CPUID checks to determine support of the ISA. ### Motivation and Context Currently FP16 models executed on systems that lack complete FP16 operator support use single precision on every node to run the model, this means the original FP16 weights have to be casted to FP32 in order to run the model properly, this change aims to accelerate the casting by using upconvert instructions and therefore improve performance.	2024-09-09 21:19:31 -07:00

1 2 3 4 5 ...

1776 commits