onnxruntime

mirror of https://github.com/saymrwulf/onnxruntime.git synced 2026-06-03 23:49:44 +00:00

Author	SHA1	Message	Date
Scott McKay	9372e9a0a3	Support >2GB of Tensor data in training checkpoint (#20077 ) ### Description <!-- Describe your changes. --> Add ability to store initializer data in an external file. Update training checkpoint code to use external file if data > ~2GB. I don't see a way for the flatbuffers 64-bit offsets to be used, as they don't support storing 'table' types with 64-bit offsets (and our Tensor is a 'table' type not a simple struct). `0cfb7eb80b/tests/64bit/test_64bit.fbs (L38-L39)` Allowing a Tensor to have its raw_data in an external file should hopefully work with the least friction. As it's an extra field it's backwards compatible. Please feel free to suggest alternative approaches. Side note: the diffs in the generated *.fbs.h files are unexpectedly large. Maybe they weren't re-generated when the new flatbuffers version was checked in. I updated by running: `python .\compile_schema.py -f <build output dir>\_deps\flatbuffers-build\Debug\flatc.exe` from onnxruntime\core\flatbuffers\schema which I thought was the correct way but maybe that's out of date. I think you can ignore all the diffs in the generated files and just worry about the changes to the .fbs files in onnxruntime/core/flatbuffers/schema. Basically start at the bottom of the files changed and work up as all the 'real' diffs are there. ### Motivation and Context <!-- - Why is this change required? What problem does it solve? - If it fixes an open issue, please link to the issue here. --> --------- Co-authored-by: carzh <wolfivyaura@gmail.com>	2024-04-22 15:17:43 -07:00
Yulong Wang	a457c1df80	upgrade emsdk to 3.1.57 (#20295 ) ### Description upgrade emsdk to 3.1.57	2024-04-19 23:05:18 -07:00
sfatimar	4d1963c2a2	OpenVINO EP Rel 1.18 Changes (#20337 ) ### Description These changes include Support to OpenVINO 2024.1 Import PreCompiled Blobs with EPContext Blob Separate Device/Precision as input Deprecate CPU_FP32 , GPU_FP32 terminology , introduce CPU, GPU AUTO GPU, CPU will only create GPU Blob and not CPU Blob. ### Motivation and Context - OpenVINO 2024.1 will be out soon - Import Precompiled Blob can greatly reduce FEIL/FIL Time. - Separating Device/Precision will make the input cleaner - --------- Co-authored-by: Suryaprakash Shanmugam <suryaprakash.shanmugam@intel.com> Co-authored-by: Preetha Veeramalai <preetha.veeramalai@intel.com>	2024-04-19 00:31:38 -07:00
Patrice Vignola	12569626cb	Update DML to 1.14.1 (#20380 ) ### Description <!-- Describe your changes. --> ### Motivation and Context <!-- - Why is this change required? What problem does it solve? - If it fixes an open issue, please link to the issue here. -->	2024-04-18 22:43:41 -07:00
Patrice Vignola	745b426c60	[DML] Update DML to 1.14 (#20304 ) I am prefiring this change to pre-run the non-dml checks, and also to give folks the time to review it before DML gets released. When DML 1.14 officially releases, we'll only need to run the DML pipeline to automatically pick up the nuget package. This should save us some valuable time. Note that DML 1.14 is the release needed for ORT 1.17.4, and DML 1.15 will come soon after.	2024-04-18 16:22:57 -07:00
Adam Louly	ee74fb6908	Introducing ORTPipelineModule - DeepSpeed Parallel Pipeline Support. (#20287 ) ### Description Introducing a new class ORTPipelineModule to handle wrapping layers in DeepSpeed pipeline parallel. ### Motivation and Context To support pipeline parallelism on ORTModule. This PR will include an initial support of deepspeed Pipeline parallelism. - [x] Support Pipeline parallel where layers are nn Modules in Sequential. - [ ] Support LayerSpec and TiedLayerSpec - [ ] Enable partitioning to accept List - [ ] Full-GPU Graph Consolidation - [ ] Subgraph Merging for Inference	2024-04-18 11:30:15 -07:00
Sumit Agarwal	f664f91298	[DML EP] Expose NPU macro via build command (#20306 ) ### Description This fixes following things: - Expose `ENABLE_NPU_ADAPTER_ENUMERATION` macro via build command, so that a user can enable NPU support for DML EP seamlessly. - Add keyword `_dmlEp_` as part of the node name, which would be useful for debugging purpose. ### Motivation and Context <!-- - Why is this change required? What problem does it solve? - If it fixes an open issue, please link to the issue here. -->	2024-04-18 11:23:13 -07:00
Patrice Vignola	76434907fb	[DML EP] Add graph capture (#20257 ) This adds a new "Graph Capture" option to the DML ep, similar to the cuda graph functionality. Here's how graph capture works: - A user can enable graph capture in the session options by setting `ep.dml.enable_graph_capture` to `true` - When they want to capture a run, they set `gpu_graph_id` in their `RunOptions` to a number bigger than 0 (0 is reserved for internal use according to the cuda graph documentation). - Then, when they start the inference, the graph will be captured and stored in the DML EP for future use - When they execute the run for a second time with the same id, the `ReplayGraph` function in the DML EP will be called instead of executing the kernels, resulting in very low overhead and avoiding kernel recompilation. This feature can give up-to-par or even better performance than specifying the static dimensions at session creation time, but is also much more flexible.	2024-04-18 10:15:00 -07:00
Adrian Lizarraga	0a1902525f	Add patch for ONNX 1.16.0 shape inference bug (#20316 ) ### Description - Adds a patch that fixes a shape inference bug that caused a segfault: https://github.com/onnx/onnx/pull/6080 - Fix documentation describing why QLinearMatMul tests are currently being skipped. ### Motivation and Context The [PR for integrating with ONNX 1.16.0](https://github.com/microsoft/onnxruntime/pull/19745) disabled various python quantization tests due to a shape inference bug. This PR applies the ONNX fix as a patch. We still can't enable the tests because some of our CIs pip install onnx-1.16.0, which doesn't include the fix.	2024-04-17 10:23:22 -07:00
Yi-Hong Lyu	6b6a62fb40	Add vectorized AVX512F kernel for ReduceMaximumF32Kernel (#20268 ) ### Description <!-- Describe your changes. --> This commit introduces a new vectorized AVX512F kernel, MlasReduceMaximumF32KernelAvx512F, which efficiently computes the maximum value of the supplied buffer. Additionally, microbenchmarks have been added for MlasComputeSoftmax (inplace), MlasReduceMaximumF32KernelAvx, MlasComputeSumExpF32KernelAvx512F, and MlasComputeSoftmaxOutputF32KernelAvx. ### Motivation and Context <!-- - Why is this change required? What problem does it solve? - If it fixes an open issue, please link to the issue here. --> The goal of this commit is to enhance the performance of ReduceMaximumF32Kernel on CPUs with AVX512F instruction support. \| AVX \| \| \| AVX512 \| \| \| -- \| -- \| -- \| -- \| -- \| -- \| -- \| -- name \| iterations \| real_time \| cpu_time \| iterations \| real_time \| cpu_time \| time_unit REDUCEMAXIMUMF32KERNEL[]/ByteAligned:4/D:3/real_time \| 271277304 \| 2.58095 \| 2.58091 \| 263338132 \| 2.65661 \| 2.65661 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:8/D:3/real_time \| 271220477 \| 2.58095 \| 2.58095 \| 263509929 \| 2.65652 \| 2.65649 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:16/D:3/real_time \| 271240587 \| 2.58064 \| 2.58064 \| 263479542 \| 2.65671 \| 2.65665 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:32/D:3/real_time \| 271227745 \| 2.58083 \| 2.58079 \| 263402506 \| 2.65657 \| 2.65657 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:64/D:3/real_time \| 271255069 \| 2.58073 \| 2.58071 \| 263463858 \| 2.65682 \| 2.65682 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:128/D:3/real_time \| 271257174 \| 2.58058 \| 2.58052 \| 263460120 \| 2.65682 \| 2.65682 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:4/D:4/real_time \| 174395051 \| 4.01401 \| 4.01401 \| 197330481 \| 3.5465 \| 3.54636 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:8/D:4/real_time \| 174645502 \| 3.99691 \| 3.99691 \| 197474831 \| 3.54298 \| 3.54278 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:16/D:4/real_time \| 174523308 \| 4.01391 \| 4.01386 \| 197389981 \| 3.54518 \| 3.54506 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:32/D:4/real_time \| 174779200 \| 3.99874 \| 3.99874 \| 197519075 \| 3.54227 \| 3.54209 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:64/D:4/real_time \| 174642874 \| 4.00645 \| 4.00641 \| 197642101 \| 3.54195 \| 3.54188 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:128/D:4/real_time \| 174546754 \| 4.0061 \| 4.00608 \| 197621033 \| 3.54296 \| 3.54281 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:4/D:5/real_time \| 162752651 \| 4.30119 \| 4.30114 \| 215552503 \| 3.24767 \| 3.24752 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:8/D:5/real_time \| 162717463 \| 4.30123 \| 4.30116 \| 215541082 \| 3.24711 \| 3.24695 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:16/D:5/real_time \| 162718819 \| 4.3016 \| 4.30153 \| 215589239 \| 3.24725 \| 3.24708 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:32/D:5/real_time \| 162719596 \| 4.30151 \| 4.30145 \| 215563846 \| 3.24956 \| 3.24949 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:64/D:5/real_time \| 162753333 \| 4.30125 \| 4.30125 \| 215537315 \| 3.24924 \| 3.24908 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:128/D:5/real_time \| 162752258 \| 4.3014 \| 4.30141 \| 215526482 \| 3.24744 \| 3.24735 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:4/D:7/real_time \| 143579660 \| 4.87526 \| 4.87516 \| 100000000 \| 5.25767 \| 5.25752 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:8/D:7/real_time \| 143585097 \| 4.87476 \| 4.87467 \| 100000000 \| 5.41583 \| 5.41567 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:16/D:7/real_time \| 143571011 \| 4.87506 \| 4.87503 \| 182359467 \| 3.83773 \| 3.83764 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:32/D:7/real_time \| 143587142 \| 4.87487 \| 4.8748 \| 182397261 \| 3.83807 \| 3.8379 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:64/D:7/real_time \| 143578465 \| 4.87525 \| 4.87521 \| 182428602 \| 3.83777 \| 3.83768 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:128/D:7/real_time \| 143588555 \| 4.87491 \| 4.87488 \| 125280452 \| 5.59791 \| 5.59766 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:4/D:9/real_time \| 284851058 \| 2.43476 \| 2.43476 \| 156879863 \| 4.42895 \| 4.42884 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:8/D:9/real_time \| 270700898 \| 2.59031 \| 2.59024 \| 157953114 \| 4.42995 \| 4.42968 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:16/D:9/real_time \| 282871172 \| 2.45385 \| 2.45385 \| 157801156 \| 4.42817 \| 4.42804 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:32/D:9/real_time \| 285307738 \| 2.47009 \| 2.47005 \| 158058507 \| 4.4279 \| 4.42786 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:64/D:9/real_time \| 285709536 \| 2.45481 \| 2.45476 \| 158070961 \| 4.42809 \| 4.42799 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:128/D:9/real_time \| 285449733 \| 2.47495 \| 2.47491 \| 158069718 \| 4.45026 \| 4.45017 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:4/D:11/real_time \| 189213618 \| 3.79684 \| 3.79676 \| 139459497 \| 5.01882 \| 5.01871 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:8/D:11/real_time \| 185600468 \| 3.76394 \| 3.76376 \| 139444892 \| 5.01922 \| 5.01905 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:16/D:11/real_time \| 184968668 \| 3.80636 \| 3.80636 \| 139470834 \| 5.01948 \| 5.01936 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:32/D:11/real_time \| 183867226 \| 3.80432 \| 3.80427 \| 139481986 \| 5.01975 \| 5.01944 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:64/D:11/real_time \| 184301650 \| 3.81634 \| 3.81634 \| 139452846 \| 5.01983 \| 5.01972 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:128/D:11/real_time \| 186215795 \| 3.82659 \| 3.82654 \| 139497736 \| 5.02119 \| 5.02113 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:4/D:13/real_time \| 135622415 \| 5.16256 \| 5.16252 \| 124661337 \| 5.61227 \| 5.61194 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:8/D:13/real_time \| 135618907 \| 5.15967 \| 5.1596 \| 124805224 \| 5.6088 \| 5.60854 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:16/D:13/real_time \| 135612192 \| 5.15506 \| 5.15501 \| 124803221 \| 5.60901 \| 5.60869 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:32/D:13/real_time \| 135906082 \| 5.15818 \| 5.15818 \| 124776601 \| 5.60898 \| 5.60886 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:64/D:13/real_time \| 135369523 \| 5.15709 \| 5.15682 \| 124790370 \| 5.60927 \| 5.60902 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:128/D:13/real_time \| 135596827 \| 5.1603 \| 5.1603 \| 124792145 \| 5.61637 \| 5.61614 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:4/D:15/real_time \| 110947137 \| 5.96511 \| 5.96495 \| 112861522 \| 6.20035 \| 6.20014 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:8/D:15/real_time \| 118004792 \| 6.22645 \| 6.22628 \| 112909900 \| 6.20073 \| 6.20073 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:16/D:15/real_time \| 112630319 \| 6.25564 \| 6.25552 \| 112874563 \| 6.19932 \| 6.19924 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:32/D:15/real_time \| 117403034 \| 6.17263 \| 6.17258 \| 112927318 \| 6.19866 \| 6.19842 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:64/D:15/real_time \| 108921863 \| 6.48624 \| 6.48612 \| 112927746 \| 6.20057 \| 6.20026 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:128/D:15/real_time \| 110358148 \| 6.66805 \| 6.66789 \| 112907312 \| 6.19938 \| 6.19908 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:4/D:16/real_time \| 203419574 \| 3.4415 \| 3.44137 \| 237134525 \| 2.95649 \| 2.95638 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:8/D:16/real_time \| 203414035 \| 3.4411 \| 3.44099 \| 237129564 \| 2.95178 \| 2.95171 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:16/D:16/real_time \| 203404068 \| 3.44157 \| 3.44151 \| 236981704 \| 2.9518 \| 2.95167 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:32/D:16/real_time \| 203391471 \| 3.44146 \| 3.44137 \| 237108807 \| 2.95203 \| 2.95196 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:64/D:16/real_time \| 203393801 \| 3.44131 \| 3.44127 \| 237126460 \| 2.95278 \| 2.95272 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:128/D:16/real_time \| 203407476 \| 3.44181 \| 3.44162 \| 237154444 \| 2.95293 \| 2.9528 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:4/D:500/real_time \| 37551439 \| 18.6407 \| 18.6407 \| 39222534 \| 17.858 \| 17.8571 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:8/D:500/real_time \| 37544097 \| 18.6404 \| 18.6401 \| 39174151 \| 17.8539 \| 17.8536 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:16/D:500/real_time \| 37549837 \| 18.6391 \| 18.6391 \| 39233956 \| 17.8507 \| 17.8505 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:32/D:500/real_time \| 45996345 \| 15.2157 \| 15.2153 \| 39285929 \| 17.848 \| 17.8474 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:64/D:500/real_time \| 46012429 \| 15.2184 \| 15.2179 \| 65664865 \| 10.7366 \| 10.7364 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:128/D:500/real_time \| 45912375 \| 15.2349 \| 15.2346 \| 65205908 \| 10.8498 \| 10.8492 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:4/D:2000/real_time \| 9493955 \| 73.7232 \| 73.7203 \| 10188090 \| 68.7931 \| 68.7908 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:8/D:2000/real_time \| 9495562 \| 73.7173 \| 73.7173 \| 10180895 \| 68.7533 \| 68.7511 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:16/D:2000/real_time \| 9487371 \| 73.7852 \| 73.7831 \| 10164473 \| 68.7279 \| 68.725 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:32/D:2000/real_time \| 10816047 \| 64.7322 \| 64.7287 \| 10168481 \| 68.8109 \| 68.8096 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:64/D:2000/real_time \| 10808802 \| 64.7232 \| 64.721 \| 19478320 \| 36.1471 \| 36.1461 \| ns REDUCEMAXIMUMF32KERNEL[]/ByteAligned:128/D:2000/real_time \| 10818192 \| 64.7304 \| 64.728 \| 19419672 \| 35.9635 \| 35.9635 \| ns	2024-04-16 13:52:43 -07:00
George Wu	08d208b969	[QNN EP] refactor QNN deps/copy logic. start copying deps to target python loc… (#20317 ) copy QNN deps when building python bindings as well. tweak the wildcard to only copy QNN related files. latest sdk from Qualcomm (>= 2.21) also include SNPE dll's which we don't want to include.	2024-04-15 22:33:12 -07:00
liqun Fu	cd7112f800	Integration with ONNX 1.16.0 (#19745 ) ### Description update with ONNX 1.16.0 branch according to https://github.com/microsoft/onnxruntime/blob/main/docs/How_To_Update_ONNX_Dev_Notes.md ONNX 1.16.0 release notes: https://github.com/onnx/onnx/releases/tag/v1.16.0 #### Updated ops for CPU EP: - DequantizeLinear(21) - Added int16 and uint16 support + various optimizer tests - Missing int4 and uint4 support - Missing block dequantization support - QuantizeLinear(21) - Added int16 and uint16 support + various optimizer tests - Missing int4 and uint4 support - Missing block quantization support - Cast(21) - Missing int4 and uint4 support - CastLike(21) - Missing int4 and uint4 support - ConstantOfShape(21) - Missing int4 and uint4 support - Identity(21) - Missing int4 and uint4 support - If(21) - Missing int4 and uint4 support - Loop(21) - Missing int4 and uint4 support - Reshape(21) - Missing int4 and uint4 support - Scan(21) - Missing int4 and uint4 support - Shape(21) - Missing int4 and uint4 support - Size(21) - Missing int4 and uint4 support - Flatten(21) - Missing float8e4m3fnuz, float8e5m2, float8e5m2fnuz, int4, and uint4 support - Pad(21) - Missing float8e4m3fnuz, float8e5m2, float8e5m2fnuz, int4, and uint4 support - Squeeze(21) - Missing float8e4m3fnuz, float8e5m2, float8e5m2fnuz, int4, and uint4 support - Transpose(21) - Missing float8e4m3fnuz, float8e5m2, float8e5m2fnuz, int4, and uint4 support - Unsqueeze(21) - Missing float8e4m3fnuz, float8e5m2, float8e5m2fnuz, int4, and uint4 support #### Unimplemented opset 21 features/ops - int4 and uint4 data type - QLinearMatMul(21) - GroupNormalization(21) - ai.onnx.ml.TreeEnsemble(5) ### Motivation and Context <!-- - Why is this change required? What problem does it solve? - If it fixes an open issue, please link to the issue here. --> ### Disabled tests #### ORT Training orttraining/orttraining/test/python/orttraining_test_ort_apis_py_bindings.py - test_ort_custom_ops: Potential shape inference bug for custom ops #### Python quantization unit tests test/onnx/python/quantization (shape inference bug) - test_op_conv_transpose.py: test_quantize_conv_transpose_u8u8_fp16 - test_op_conv_transpose.py: test_quantize_conv_transpose_s8s8_fp16 - test_op_gemm.py: test_quantize_qop_gemm_s8s8 - test_op_gemm.py: test_quantize_qop_gemm_e4m3fn_same - test_op_gemm.py: test_quantize_qop_gemm_e4m3fn_p3 - test_op_matmul.py: test_quantize_matmul_u8u8_f16 - test_op_matmul.py: test_quantize_matmul_s8s8_f16 - test_op_matmul.py: test_quantize_matmul_s8s8_f16_entropy - test_op_matmul.py: test_quantize_matmul_s8s8_f16_percentile - test_op_matmul.py: test_quantize_matmul_s8s8_f16_distribution - test_op_relu.py: test_quantize_qop_relu_s8s8 #### ONNX tests - test_maxpool_2d_ceil_output_size_reduce_by_one: ONNX 1.16.0 fixed a maxpool output size bug and added this test. Enable this test when [ORT PR](https://github.com/microsoft/onnxruntime/pull/18377) is merged. Refer to original [ONNX PR](https://github.com/onnx/onnx/pull/5741). - test_ai_onnx_ml_tree_ensemble_set_membership_cpu: new unimplemented op ai.onnx.ml.TreeEnsemble - test_ai_onnx_ml_tree_ensemble_single_tree_cpu: same - test_ai_onnx_ml_tree_ensemble_set_membership_cuda: same - test_ai_onnx_ml_tree_ensemble_single_tree_cuda: same - test_cast_INT4_to_FLOAT_cpu: ORT Cast(21) impl doesn't support int4 yet - test_cast_INT4_to_INT8_cpu: same - test_cast_UINT4_to_FLOAT_cpu: same - test_cast_UINT4_to_UINT8_cpu: same - test_cast_INT4_to_FLOAT_cuda - test_cast_INT4_to_INT8_cuda - test_cast_UINT4_to_FLOAT_cuda - test_cast_UINT4_to_UINT8_cuda - test_constantofshape_float_ones_cuda: ConstantOfShape(21) not implemented for cuda - test_constantofshape_int_shape_zero_cuda: same - test_constantofshape_int_zeros_cuda: same - test_flatten_axis0_cuda: Flatten(21) not implemented for cuda - test_flatten_axis1_cuda: same - test_flatten_axis2_cuda: same - test_flatten_axis3_cuda: same - test_flatten_default_axis_cuda: same - test_flatten_negative_axis1_cuda: same - test_flatten_negative_axis2_cuda: same - test_flatten_negative_axis3_cuda: same - test_flatten_negative_axis4_cuda: same - test_qlinearmatmul_2D_int8_float16_cpu: QLinearMatMul(21) for onnx not implemented in ORT yet - test_qlinearmatmul_2D_int8_float32_cpu: same - test_qlinearmatmul_2D_uint8_float16_cpu: same - test_qlinearmatmul_2D_uint8_float32_cpu: same - test_qlinearmatmul_3D_int8_float16_cpu: same - test_qlinearmatmul_3D_int8_float32_cpu: same - test_qlinearmatmul_3D_uint8_float16_cpu: same - test_qlinearmatmul_3D_uint8_float32_cpu: same - test_qlinearmatmul_2D_int8_float16_cuda: same - test_qlinearmatmul_2D_int8_float32_cuda: same - test_qlinearmatmul_2D_uint8_float16_cuda: same - test_qlinearmatmul_2D_uint8_float32_cuda: same - test_qlinearmatmul_3D_int8_float16_cuda: same - test_qlinearmatmul_3D_int8_float32_cuda: same - test_qlinearmatmul_3D_uint8_float16_cuda: same - test_qlinearmatmul_3D_uint8_float32_cuda: same - test_size_cuda: Size(21) not implemented for cuda - test_size_example_cuda: same - test_dequantizelinear_blocked: Missing implementation for block dequant for DequantizeLinear(21) - test_quantizelinear_blocked_asymmetric: Missing implementation for block quant for QuantizeLinear(21) - test_quantizelinear_blocked_symmetric: Missing implementation for block quant for QuantizeLinear(21) --------- Signed-off-by: liqunfu <liqun.fu@microsoft.com> Signed-off-by: Ganesan Ramalingam <grama@microsoft.com> Co-authored-by: Ganesan Ramalingam <grama@microsoft.com> Co-authored-by: George Wu <jywu@microsoft.com> Co-authored-by: adrianlizarraga <adlizarraga@microsoft.com>	2024-04-12 09:46:49 -07:00
Andrew Fantino	7303a90f49	Fix build errors from date/date.h C++20 compatibility (#20139 ) ### Description For C++ standards >= 20, use `std::chrono::operator<<` in place of `date::operator<<` to fix ambiguous operator compile error. ### Motivation and Context The external dependency HowardHinnant/date has a conflict with std::chrono for >=C++20. Solves #20137	2024-04-02 22:10:25 -07:00
Yi Zhang	dae77e6014	Support building Windows CUDA with Ninja (#20176 ) ### How to run it locally 1. conda install ninja 2. "C:\Program Files\Microsoft Visual Studio\2022\Enterprise\VC\Auxiliary\Build\vcvarsall.bat" x64 3. python.exe {ort_repo}\tools\ci_build\build.py --config RelWithDebInfo --build_dir {ort_repo}\build_cuda --skip_submodule_sync --build_csharp --update --parallel --cmake_generator "Ninja" --build_shared_lib --enable_onnx_tests --enable_pybind --build_java --build_nodejs --use_cuda "--cuda_home=C:\Program Files\NVIDIA GPU Computing Toolkit\CUDA\v11.8" --enable_cuda_profiling --cmake_extra_defines CMAKE_CUDA_ARCHITECTURES=60 4. cd build_cuda\RelWithDebInfo 5. cmake --build . j16 ### Motivation and Context In packaging pipelines, we often come across a random issue that the building with CUDA on Windows takes too much time. Although it has been reduced much by moving the building to the CPU machine. We're planning to build with Ninja instead of msbuild in Packaging pipelines, thus, nvcc can run parallelly. It's the first step to support it locally.	2024-04-03 11:19:31 +08:00
Jeff Bloomfield	2f31560430	Enable generic feature level devices in DML EP (#20114 ) ### Description Enable NPUs supporting DXCORE_ADAPTER_ATTRIBUTE_D3D12_GENERIC_ML and D3D_FEATURE_LEVEL_1_0_GENERIC with DML EP. This also begins ingesting DX headers through the DirectX-Headers repo. Note that this includes an update to cgamanifest.json for onnx-tensorrt which is triggered during re-generation due to a prior changes to deps.txt. ### Motivation and Context <!-- - Why is this change required? What problem does it solve? - If it fixes an open issue, please link to the issue here. -->	2024-03-29 14:37:30 -07:00
Ye Wang	17919717b5	add QMoE (#20108 ) ### Description <!-- Describe your changes. --> 1. Introduce latest cutlass extension from TRTLLM that gives us cutlass upgrade(to 3.4) opportunity from MoE side. 2. Fix Windows build issue 3. Add Int4 MoE op and ut ### Motivation and Context <!-- - Why is this change required? What problem does it solve? - If it fixes an open issue, please link to the issue here. -->	2024-03-29 10:24:19 -07:00
Dmitri Smirnov	b95fd4e644	Enable CUDA EP unit testing on Windows (#20039 ) ### Description Address build issues and source code discrepancies. Fix cuda_test_provider gtest argument stack corruption. ### Motivation and Context `OpTester` class that is widely used for kernel testing is not suitable for testing internal classes for EPs that are built as shared objects. Currently, CUDA EP tests run only on Linux. We want to enable testing and developments on Windows, and create a usable pattern for testing of other EPs internals. Alternatives considered: Abstracting EP unit tests into separate test executable such as `onnxruntime_test_all`. This alternative was rejected as it would create a lot more changes in the established patterns, and potentially interfere with CUDA functionality with more complex source code maintanence.	2024-03-27 13:32:36 -07:00
Dmitri Smirnov	3076b56947	Make MS Debug engine SymInitialize() called as needed. (#20036 ) ### Description <!-- Describe your changes. --> Initialize Symbol engine as needed with no duplicate calls. ### Motivation and Context <!-- - Why is this change required? What problem does it solve? - If it fixes an open issue, please link to the issue here. --> Currently absel library may call SymInitialize more than once when shared libraries are involved. However, this can only be called only once per process. Our debug_alloc also may call it when enabled. This change enables intialization to proceed only when needed with no duplicate effort.	2024-03-22 16:17:47 -07:00
sfatimar	eab35c20fc	Ort openvino npu 1.17 master (#19966 ) ### Description Add NPU to list of device supported. Added changes for Support to OV 2024.0 Nuget packages removes packaging of OpenVINO DLL Bug Fixes with Python API Reverted Dockerfiles not being maintained. ### Motivation and Context NPU Device has been introduced by Intel in latest client systems OpenVINO 2024.0 release is out. --------- Co-authored-by: Suryaprakash Shanmugam <suryaprakash.shanmugam@intel.com> Co-authored-by: Preetha Veeramalai <preetha.veeramalai@intel.com> Co-authored-by: Ubuntu <ubuntu@ubuntu-118727.iind.intel.com> Co-authored-by: hmamidix <hemax.sowjanya.mamidi@intel.com> Co-authored-by: vthaniel <vishnudas.thaniel.s@intel.com> Co-authored-by: saurabhkale17 <saurabh1.kale@intel.com>	2024-03-21 18:44:00 -07:00
Changming Sun	dafbef3a21	CMake: support reading dependency zip files from a local mirror (#20005 ) ### Description To test this feature, run ```bat python cmake\deps_update_and_upload.py --root-path mirror ``` Then run build.py as usual. The zip files will be cached local. To avoid being downloaded again and again.	2024-03-21 17:58:59 -07:00
Yufeng Li	15219e2e71	turn on neural_speed by default (#19627 ) ### Description <!-- Describe your changes. --> the crash caused by the neural_speed turns out to be a very corn case. Turn it on by default. ### Motivation and Context <!-- - Why is this change required? What problem does it solve? - If it fixes an open issue, please link to the issue here. -->	2024-03-20 12:49:58 -07:00
Rachel Guo	6b305f95e0	Support xcframework for mac catalyst builds. (#19534 ) ### Description <!-- Describe your changes. --> ### Motivation and Context <!-- - Why is this change required? What problem does it solve? - If it fixes an open issue, please link to the issue here. --> MAUI on macOS uses mac-catalyst which requires a different native binary. --------- Co-authored-by: rachguo <rachguo@rachguos-Mini.attlocal.net> Co-authored-by: Scott McKay <skottmckay@gmail.com>	2024-03-20 10:55:19 -07:00
mindest	3dfe4a5e6d	[ROCm] Remove MPI dependency and collectives to use NCCL (#19830 ) ### Description * Remove MPI dependency to use NCCL AllReduce, etc. * Exclude unsupported collectives in hipify	2024-03-19 17:35:18 -07:00
Ted Themistokleous	6bb64683f8	Use version instead of version-dev for ROCm (#19967 )	2024-03-19 10:40:40 +08:00
Adam Louly	32558134a9	[On-Device-Training] Upgrade Flatbuffers to Support 2GB+ Checkpoints. (#19770 ) ### Description Modifications to support 2GB+ checkpoint & Upgrading Flatbuffers ### Motivation and Context This PR includes changes that will make ort handle 2GB+ checkpoints. To do that we need to upgrade flatbuffers to 23.5.9 - https://github.com/google/flatbuffers/pull/7945 - Modified the commitHash and the hash for the new version - Removed the patch for rust generator's unused variable warning as it is no longer producing this - [Check it out here](`d121e09d89/src/idl_gen_rust.cpp`) - Updated the VerifyField calls with alignment values that were introduced in the new version. --------- Co-authored-by: Sumit Agarwal <sumitagarwal@microsoft.com>	2024-03-14 16:36:24 -07:00
Changming Sun	1fb6cbddee	Add a build patch for Windows ARM64EC (#19898 ) ### Description Add a patch for Windows ARM64EC ### Motivation and Context Will need more changes in onnxruntime/core/common/cpuid_arch_definition.h and onnxruntime/core/common/cpuid_info.cc	2024-03-14 08:50:42 -07:00
Jeff Daily	9443366009	[ROCm] fix build failure when nccl is enabled (#19900 ) Building onnxruntime ROCm EP with --enable_nccl --use_mpi fails due to inclusion of MOE source files but MOE is not supported. The error observed is `error: contrib_ops/rocm/moe/ft_moe/moe_kernel.h: No such file or directory` The fix is to exclude collective/sharded_moe.* files when nccl is requested.	2024-03-13 21:16:54 -07:00
Adrian Lizarraga	9c3242ab70	[QNN EP] Copy security catalog file for HtpV73Skel.so from QNN SDK (#19903 ) ### Description Copies the `QNN_HOME/lib/hexagon-v73/unsigned/libqnnhtpv73.cat` file from QNN SDK to the unittest build directory. This is necessary in order to be able to load the `libQnnHtpV73Skel.so` file on Windows for modern versions of QNN SDK. ### Motivation and Context A [digitally-signed catalog file](https://learn.microsoft.com/en-us/windows-hardware/drivers/install/catalog-files) (.cat) can be used as a digital signature for an arbitrary collection of files.	2024-03-13 20:52:59 -07:00
Jake Mathern	18ad8587a6	[CP] Fix for xfgcheck and Fix WAI ARM64 build (#19634 ) (#19644 ) ### Description Fix WAI build by only conditionally copying linker flags ### Motivation and Context I broke the WAI build that contains ORT on ARM64	2024-03-13 17:54:06 -07:00
Edward Chen	860eb762c2	[Apple framework] Fix minimal build with training enabled. (#19858 ) Fix some linker errors that come up when integrating the onnxruntime-training-c pod into another Xcode project. The problematic configuration is a minimal build with training APIs enabled. - training_op_defs.o had some unresolved references to ONNX functions. It should not be included at all in a minimal build. - tree_ensemble_helper.o also had unresolved references to ONNX ParseData. The containing function is unused in a minimal build. Added a test to cover this configuration.	2024-03-12 11:33:30 -07:00
Scott McKay	978c40d853	Make partitioning utils QDQ aware so it does not break up QDQ node units (#19723 ) ### Description <!-- Describe your changes. --> If the EP handles QDQ node units, we need to make sure we do not split those into different partitions. Update the partitioning utils to be QDQ aware. If there are node units we process the logical nodes they represent instead of individual nodes. This ensure we process all nodes in a QDQ node unit at the same time so that they are always in the same partition. ### Motivation and Context <!-- - Why is this change required? What problem does it solve? - If it fixes an open issue, please link to the issue here. --> Fix one of the issues in #19590 --------- Co-authored-by: Edward Chen <18449977+edgchen1@users.noreply.github.com>	2024-03-12 10:55:49 +10:00
Changming Sun	efad5bbc5a	Replace some old file system calls with C++17 std::filesystem APIs. (#19196 ) ### Description 1. Replace some old file system calls to use C++17 std::filesystem APIs. 2. Remove tensorflow_C_PACKAGE_PATH cmake option, which was only used in onnxruntime_perf_test and the code is out of maintain. 3. Excludes onnx_test_runner and onnxruntime_perf_test from iOS build because C++17 filesystem library is not available there	2024-03-09 09:17:36 -08:00
Scott McKay	db59cec82f	Don't reduce warning level for CUDA build on Windows (#19663 ) ### Description <!-- Describe your changes. --> Address warnings so all the ORT projects build with /W4 on Windows. Mainly - unused parameters - variables shadowing other ones ### Motivation and Context <!-- - Why is this change required? What problem does it solve? - If it fixes an open issue, please link to the issue here. --> #19588 started on this.	2024-03-06 15:03:55 +10:00
Chi Lo	d9730c7f43	[TensorRT EP] Fix bug for DDS output handling for empty tensor (#19575 ) When the DDS output is empty tensor (i.e. any of the dimension is 0), TRT EP won't perform either cudaMemcpyAsync() nor cuda::Impl_Cast(), to prevent accidentally overwriting other location that might belong to other tensors. This PR also refactors the code to only allocate single bytes for all empty tensors. #TODO: add unit tests to cover the DDS code paths or doing more testing with concurrent,sequential, threaded faster-rcnn using onnx_test_runner and verifying outputs --------- Co-authored-by: Chi Lo <lochi@microsoft.com>	2024-03-05 14:39:36 -08:00
Chen Fu	06e684c9f2	Adding cuda kernel (optimized for sm80) for block-wise 4b quantized float 16 GEMM. (#18619 ) ### Description Adding CUDA kernel for block-wise 4b quantized float 16 GEMM, this is specially optimized for Nvidia Ampere GPUs. ### Motivation and Context Trying to improve quantized LLM inference performance on Nvidia Ampere GPUs ### Note: This is implemented by extending CUTLASS, so it has a hard dependency on CUTLASS. However, in current build system, loading of CUTLASS dependency is guarded with: (onnxruntime_USE_FLASH_ATTENTION OR onnxruntime_USE_MEMORY_EFFICIENT_ATTENTION) If both of these options are turned off, then compilation will fail. Why CUTLASS dependency is guarded at all? It's a header file only library that does not introduce any binary if not instantiated. What's the downside of removing all the guards and just include CUTLASS unconditionally?	2024-03-05 09:37:45 -08:00
Changming Sun	a0521f899e	Enable CPUINFO for all Windows build (#19655 ) ### Description It was disabled in PR #9065. And the reason was: " api-ms-win-core-kernel32-legacy-*.dll wasn't available in Windows 8 and was added in Windows 10, so cpuinfo breaks our Windows 8 support. I'm disabling it again." We no longer support Windows 8. Therefore we can add CPUINFO back. ### Motivation and Context To make the code simpler. If in any case the library doesn't work as expected, we can submit a PR to their code base and fix it.	2024-03-01 16:23:20 -08:00
Edward Chen	5672cdebdf	Update google benchmark to 1.8.3. (#19734 ) Update google benchmark to 1.8.3. Update deps_update_and_upload.py script to make it easier to use.	2024-03-01 11:01:58 -08:00
Scott McKay	2a857d9a86	Add ML Program support for more operators (#19527 ) ### Description <!-- Describe your changes. --> Add support for: - Clip/Relu/Relu6 - Add/Mul/Div/Sub/Pow - GlobalAveragePool/GlobalMaxPool/AveragePool/MaxPool - Reshape - Gemm/MatMul Fix some build issues/warnings from changes. Fix a couple of potential issues with the Resize op as well (noticed due to change to reject inputs with empty data at a higher level). ### Motivation and Context <!-- - Why is this change required? What problem does it solve? - If it fixes an open issue, please link to the issue here. --> Enable mobilenetv2 with ML Program	2024-03-01 10:23:29 +10:00
Maximilian Müller	c20ced4132	Use CMake's find package for CUDA libs (#19673 ) ### Description Answers issue #19640 More details are in the issue, basically I am changing all the include directory and link directory usage to CMake's `CUDA::*` targets	2024-02-27 11:26:48 -08:00
cloudhan	1e69b61238	Make version string detection more robust (#19615 ) `/opt/rocm/.info/version-dev` is only available if the `rocm-dev` metapackage is installed. This will bring a lot of unused packages which are not needed by the users, they may opt for fine grained control. Fallback to `rocm_version.h` in case `rocm-dev` is not installed.	2024-02-27 16:06:06 +08:00
Changming Sun	9ccdc4961a	Stop using apiset in OneCore build: use onecoreuap.lib instead of onecoreuap_apiset.lib (#19632 ) ### Description Stop using apiset in OneCore build: use onecoreuap.lib instead of onecoreuap_apiset.lib in onecore build. ### Motivation and Context 1. Now all Windows Editions come with Reverse Forwarders. We should just use the normal onecore libs. 2. Many new Windows APIs are only available in [windows umbrella libraries](https://learn.microsoft.com/en-us/windows/win32/apiindex/windows-umbrella-libraries). So these libraries are not specific for Windows CoreOS or Onecore. 3. Going forward we should use "IsApiSetImplemented" to guard our API usages: https://learn.microsoft.com/en-us/windows/win32/apiindex/detect-api-set-availability . After this change, our built binaries can pass apivalidator's check. ``` C:\local\apivalidator>apivalidator.exe -BinaryPath:C:\src\onnxruntime\b\Debug\Debug\onnxruntime.dll -SupportedApiXmlFiles:onecoreuap_DDIs.xml ApiValidation: Summary: "C:\src\onnxruntime\b\Debug\Debug\onnxruntime.dll" is Universal ApiValidation: All binaries are Universal ``` So it will give an easy way to test ONNX Runtime's compatibility to Windows versions.	2024-02-23 22:31:57 -08:00
cao lei	f430600432	Enable streams for DML EP. This change is to revert PR 19481 since the bug 19480 is fixed by PR 19515 (#19609 ) ### Description <!-- Describe your changes. --> Enable streams for DML EP. This change is to revert PR 19481 since the bug 19480 is fixed by PR 19515 ### Motivation and Context <!-- - Why is this change required? What problem does it solve? - If it fixes an open issue, please link to the issue here. --> Enable streams for DML EP. This change is to revert PR 19481 since the bug 19480 is fixed by PR 19515	2024-02-23 06:02:05 -08:00
pengwa	ae92d593c0	ONNX Gelu Op in Opset 20 (#19560 ) ### ONNX Gelu Op in Opset 20 Refactor code to support MSDomain Gelu and ONNX Gelu-opset20 Op 1. Move CPU-GELU implmentation from `onnxruntime/contrib_ops/cpu/activations.h/cc` to `onnxruntime/core/providers/cpu/tensor/gelu.h/cc`, as the implementation for approximate attribute to be 'none'. 2. Dumplicate some logic from `onnxruntime/contrib_ops/cpu/bert/bias_gelu.cc` to `onnxruntime/core/providers/cpu/tensor/gelu.h/cc`, as the implementation for approximate attribute to be 'tanh'. 3. Register ONNX domain Gelu CPU kernel from opset 20 in `onnxruntime/core/providers/cpu/cpu_execution_provider.cc`. 4. Move `onnxruntime/contrib_ops/cuda/bert/fast_gelu_impl.h/cu` to `onnxruntime/core/providers/cuda/tensor/gelu_impl.h` and `onnxruntime/core/providers/cuda/tensor/gelu_approximate_impl.cu` respectively, as the implementation for approximate attribute to be 'tanh'. 5. Implement the logic for approximate attribute to be 'none' in `onnxruntime/core/providers/cuda/tensor/gelu_impl.cu`. 6. Register ONNX domain Gelu CUDA kernel from opset 20 in `onnxruntime/core/providers/cuda/cuda_execution_provider.cc`. 7. ROCM ep related changes. 8. Enrich the tests for ONNX domain Gelu in `onnxruntime/test/providers/cpu/activation/activation_op_test.cc`.	2024-02-23 11:05:16 +08:00
PeixuanZuo	6226c5f62f	[ROCm] Add SkipGroupNorm for ROCm EP (#19303 ) Add SkipGroupNorm for ROCm EP. --------- Co-authored-by: Peixuan Zuo <peixuanzuo@microsoft.com@orttrainingdev7.d32nl1ml4oruzj4qz3bqlggovf.px.internal.cloudapp.net>	2024-02-21 11:08:48 +08:00
Jake Mathern	7a5860e490	Fix cmake function duplicate lib (#19547 ) ### Description Fixes cmake function definition in winml.cmake to copy link flags. ### Motivation and Context XFGCheck errors in WindowsAI because this function does not transfer linker flags	2024-02-20 13:41:40 -08:00
pengwa	b55260d076	Minor fix for cmake (#19552 ) ### Minor fix for cmake When build on Linux, get a warning saying " CMake Warning at CMakeLists.txt:1603 (message): MPI and NCCL disabled on Win build. " This message is not correct. So have such a fix to avoid any misunderstanding from users. ![image](https://github.com/microsoft/onnxruntime/assets/10530022/848c2d77-a538-4e31-8e0d-4b539233e515) ### Motivation and Context <!-- - Why is this change required? What problem does it solve? - If it fixes an open issue, please link to the issue here. -->	2024-02-19 10:21:19 +08:00
Scott McKay	4e5119760d	Add initial support for CoreML ML Program to the CoreML EP. (#19347 ) ### Description <!-- Describe your changes. --> Adds infrastructure to create an ML Package containing the Model using ML Program. Updated coremltools files to v7.1 to bring in new protobuf definitions along with the tools to write the weight.bin file and create an ML Package correctly. Enables building a CoreML Model on all platforms which means all the operator builder code can be debugged anywhere. Execution of the generated CoreML model is obviously limited to Apple platforms. The Conv operator builder has been updated to be able to generate an ML Program Operation. ### Motivation and Context <!-- - Why is this change required? What problem does it solve? - If it fixes an open issue, please link to the issue here. --> NeuralNetwork is no longer being developed and ML Program is the replacement going forward.	2024-02-15 08:46:03 +10:00
George Wu	5e70c6b3a6	allow protobuf lite build for TRT EP (#19498 ) allow protobuf-lite builds with TensorRT EP as long as it's built with the trt built-in parser and not the oss-parser. This is because trt built-in parser statically links protobuf so there aren't any conflicts for protobuf-lite.	2024-02-12 22:53:04 -08:00
Patrice Vignola	1182b5509b	Disable streams for the DML EP (#19481 ) There's currently a bug in the allocation planner when reusing buffers and more than one streams are used that make it possible (although rarely) to reach a reference count of 0 for a buffer that is still being used. Since DML doesn't benefit from multiple streams, disabling it is the safest option for now. This is a high priority issue that we need to fix for 1.17.1 since it breaks stable diffusion. Identifying the perfect fix and fixing the underlying issue would be too risky for a patch release, especially given the limited time that we have. https://github.com/microsoft/onnxruntime/issues/19480	2024-02-10 00:34:34 -08:00
Changming Sun	1007d8f3d1	Revert "Revert NeuralSpeed code for x64 MatMulNBits (#19382 )" (#19474 ) This reverts commit `0d10c7f3c1`.	2024-02-09 09:24:54 -08:00

1 2 3 4 5 ...

1628 commits