onnxruntime

mirror of https://github.com/saymrwulf/onnxruntime.git synced 2026-07-08 17:17:15 +00:00

Author	SHA1	Message	Date
Olivia Jain	60089f7093	Cuda11.4 (#8709 ) * initial update from 11.1 to 11.4 * change 11.4.1 to 11.4.0 * adjusting to match nvidia/cuda image tags * adjusting to match nvidia/cuda image tags centos7 * correction to 11.4.0 * correction to 11.4.0 * update to cuda 11.4 * change training back to 11.1 * change training back to 11.1 * point to correct nvcr.io/nvidia/cuda 11.4.1 image * change centos8 to centos7 * correct cudnn path * Update linux-gpu-ci-pipeline.yml for Azure Pipelines * Update c-api-noopenmp-packaging-pipelines.yml * need to resolve centos images but remove space and change to 11.4 * Update linux-gpu-ci-pipeline.yml * add cudnn to docker image * bump devtoolset to 10 * revert cuda 11.4 change to setup_env_trt * orttraining back to 11.1 * use nvcr.io * Fix previous change back to cuda 11.1 * update cudnn path * use cudnn image (revert if failure)	2021-08-17 16:36:26 -07:00
ashbhandare	cc275e7529	Gradient Accumulation optimization verified for correctness (#8273 ) * Fetching frontier tensors to frontend * Move before session initialize call * Fetch tensor and add to cache * Rest of the changes for using cache * Review comments * Review changes * Review comments * switch to shared_ptr * Fix bug after rebase * FE docstring change	2021-08-17 16:24:44 -07:00
Chen Fu	224380448d	Expand Qgemm UDOT kernel to 8x8 block (#8562 ) Create a new M8 loop processing A[8x8] B[8x8] per iteration. Avoid saving registers on paths that are not needed. Adjusted M2 and M1 loop, using more registers to relax the loop carrying dependencies. Nearly 7% improvement observed on Surface Pro X 2 with model ssd_mobilenet_v2_300 About 4.5% improvement on resnet50 on Surface Pro X 2.	2021-08-17 14:36:46 -07:00
baijumeswani	871eeb4dbd	Support dicts as inputs to ORTModule (#8718 )	2021-08-17 13:40:55 -07:00
Thiago Crepaldi	ed254c283f	Add support for experimental json config for fallback (#8759 )	2021-08-17 13:35:42 -07:00
KeDengMS	6ecf626a9c	[Nuphar] Parse node doc_string for quantize info (#8746 )	2021-08-17 11:29:03 -07:00
Wei-Sheng Chin	47b3ecb53b	Packaging pipeline now builds with PythonOp (aka running autograd.Function) (#8652 ) This PR disable UTs in training's package pipelines for building packages with PythonOp (torch.autograd.Function).	2021-08-17 10:55:13 -07:00
liqun Fu	2b1f0816f8	to build cpu training packages for multiple multiple python versions (#8750 ) Co-authored-by: liqun <liqun@OrtTrainingDev4.af05slrtruoetgaxwwjv5nsq5e.px.internal.cloudapp.net>	2021-08-17 10:49:44 -07:00
Thiago Crepaldi	419834d285	Add PyTorch fallback for ORTModule forward exceptions (#8346 )	2021-08-17 10:41:15 -07:00
stevenlix	11a618b2ec	Add engine encryption in TensorRT EP (#8732 ) * add engine encryption * Update tensorrt_execution_provider.cc * Update tensorrt_execution_provider.h * Update tensorrt_execution_provider.cc * Update tensorrt_execution_provider.h * clean up * update encryption signature	2021-08-17 08:34:22 -07:00
Yulong Wang	f668a79532	[js/web] fix perf mode in test (#8748 )	2021-08-16 23:18:42 -07:00
Yulong Wang	4ceedbe933	[js/web] add SharedArrayBuffer check for wasm multi-thread (#8749 )	2021-08-16 23:17:54 -07:00
Changming Sun	ae6fdd3333	Bring code coverage dashboard back (#8394 )	2021-08-16 20:54:39 -07:00
M. Zeeshan Siddiqui	0fb82f0f8a	Memory aware gradient builder. (#8582 )	2021-08-16 19:01:22 -07:00
Nat Kershaw (MSFT)	aa12d68c37	Update ORTModule API docstrings (#8309 )	2021-08-16 16:53:01 -07:00
Dmitri Smirnov	8713d76dd1	Introduce C and C++ APIs for Sparse Tensors (#8621 ) Add IsSparseTensor Add CreateSparseTensor Add utilities and test fully sparse instantiation Fully sparse blocksparse Add test and docs for fully sparse tensor instantiation Rework creation API Use API Non string API Retrofit of existing String API Add tests Add documentation Address build issues (Winml pending) Add inference test Bump binary size Add ifdef DISABLE CONTRIB	2021-08-16 16:33:47 -07:00
Changming Sun	8335d3dc0b	Fix Python Packaging Pipeline (Training Torch 1.9.0 Cuda 11.4) (#8738 )	2021-08-16 14:46:43 -07:00
Olivia Jain	9cefd1303b	Integrate Anubis (#8603 ) * copy over changes * Update build_image.sh * allow for configurable trt * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * reflect previous changes * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * model_list.json * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * checkout trt 7.1 * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update post.py * Update post.py * Update post.py * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update model_list.json * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update post.py * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update post.py * Update post.py * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update model_list.json * Update post.py * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update post.py * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update post.py * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update start_job.ps1 * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update run_mem_test_docker.sh * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Separate anubis files * revert to old pipeline * Update post.py * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * build off master Dockerfile * Delete Dockerfile.custom-trt-perf * Delete install_common_deps.sh * uncomment * Update linux-gpu-tensorrt-ci-perf-pipeline.yml * pass in trt container version * Update linux-gpu-tensorrt-ci-perf-pipeline.yml * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update post.py * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update post.py * remove sudo * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * add back build number * allow python 3.8 * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * python 3.8 fix trtexec * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * remove prev py38 * Update linux-gpu-tensorrt-ci-perf-pipeline.yml for Azure Pipelines * add perf dependencies * Update start_job.ps1	2021-08-16 13:20:28 -07:00
Nick Kreeger	93e1e1dfa1	Drop quant_util.h and move helper function into quantization.h (#8747 )	2021-08-16 15:08:25 -05:00
KeDengMS	d0ff2621ee	[Nuphar] Fix Windows build in VS 2019 (#8728 ) Update TVM to fix c++17 build break in VS 2019 Remove tvm::nnvm from build	2021-08-13 16:13:34 -07:00
Chen Fu	8f7422be69	Limiting platforms where cpuinfo is included (#8716 ) * Limiting platforms where cpuinfo is included * Suppress strncpy warning during msvc build Co-authored-by: Chen Fu <fuchen@microsoft.com>	2021-08-13 14:46:21 -07:00
George Nash	e695cd304a	Dnnl refactor (#8627 ) * dnnl ep rework rework DnnlTensor,DnnlNode,DnnlSubgraph to support arbitrary graph topology and tensor data types rework GetCapability to claim nodes in graph greedily from node topological ordering and delay creation of DnnlSubgraph until Compile rework compile to have DnnlSubgraphPrimitive as the object to handle primitive creation and execution instead of thread local primitive pool which duplicates intermediate memory allocated by the EP across threads DnnlSubgraphPrimitive provides helpers to handle many common functions for each dnnl primitive builder and become the centralized place to store input, output, intermediate memories, initializer memories and etc it provides functions to obtain input memories with automatic reordering/reshaping and moving between engines it provides interfaces to add primitive, set output memory for single node and etc add CONCURRENT_EXEC compile flag for dnnl library as without it, convolution primitive cannot be created and executed on different threads enable unit tests to run on dnnl ep as well if built with dnnl ep add dnnl ep support for Matmulinteger * Add Relu to the DNNL refactor Signed-off-by: George Nash <george.nash@intel.com> * Add Convolution op to the DNNL rework Signed-off-by: George Nash <george.nash@intel.com> * Add Pooling ops to the DNNL rework This adds the following ops: - AveragePool - GlobalAveragePool - GlobalMaxPool - MaxPool Note: Pooling with dilation is not yet supported. Note: GlobalLpPool, LpPool, MaxRoiPool, and MaxUnpool are not supported yet. Signed-off-by: George Nash <george.nash@intel.com> * Add Sum op to the DNNL rework Signed-off-by: George Nash <george.nash@intel.com> * Add ConvGrad op to the DNNL rework Signed-off-by: George Nash <george.nash@intel.com> * Add MaxPoolGrad and AveragePoolGrad ops to DNNL rework Signed-off-by: George Nash <george.nash@intel.com> * Added lrn operator to the refactored code Signed-off by chethan.palangoutu.keshava@intel.com * Added ReduceMean DNNL op to the refactor code Signed-off-by: Chethan Palangotu Keshava <chethan.palangotu.keshava@intel.com> * Added Softmax DNNL op for the refactored code Signed-off-by: Chethan Palangotu Keshava <chethan.palangotu.keshava@intel.com> * Added BatchNorm DNNL op inference-only for refactored code Signed-off-by: Chethan Palangotu Keshava <chethan.palangotu.keshava@intel.com> * Added Binary Ops to DNNL rework Signed-off-by: Wang <zhaoyang.wang@intel.com> * Added ReluGrad to DNNL Rework Signed-off-by: Wang <zhaoyang.wang@intel.com> * Update OneDNN tag to v2.3 Signed-off-by: Wang <zhaoyang.wang@intel.com> * Added support for memory upto dim size 12 this is to fix the CI test cases that contain binary ops of input dim size > 5 Signed-off-by: Wang <zhaoyang.wang@intel.com> * Prevent claiming support for float16 and bfloat16 when only float is suppoted By using The string.find used was causing the code to claiming support for float16 and bfloat16 when we only supported float. We now explicitly check the code for the data type or the data type with a 7 letter prefix basically prefixed with "tensor(" Signed-off-by: George Nash <george.nash@intel.com> * Disable uint8 mul and div, improve type conversion Disable mul_uint8 and div_uint8 test cases as they use modulo for overflow handling while onednn uses saturation improve ype conversion using enum instead of string comparsion as well as adding more types Signed-off-by: Wang <zhaoyang.wang@intel.com> Co-authored-by: Wang <zhaoyang.wang@intel.com> Co-authored-by: Chethan Palangotu Keshava <chethan.palangotu.keshava@intel.com>	2021-08-13 14:15:43 -07:00
Changming Sun	f04a235c77	Update manylinux build scripts (#8724 ) Update manylinux build scripts. Sync it with the latest upstream.	2021-08-13 12:04:00 -07:00
Ye Wang	385b571824	handle a corner case in removing useless reshape (#8669 ) * fix a corner case * optiimize * review comments	2021-08-13 09:24:06 -07:00
Nick Kreeger	d7baef765a	Fix 4byte to 8byte static warnings in attention_quant.cc (#8715 )	2021-08-13 09:31:27 -05:00
Changming Sun	436ac6dd5f	Rename ml_value.h to ort_value.h (#8726 )	2021-08-13 07:04:56 -07:00
Vincent Wang	606b6271fa	fix build (#8725 )	2021-08-13 16:24:24 +08:00
Guoyu Wang	59e59b9e0e	Fix unknown warning "-Wformat-truncation" build failure for arm (#8721 ) * fix clang arm build failure * address CR comments * Change to cmake check flag option * Add missing __aarch64__	2021-08-12 23:47:03 -07:00
Pranav Sharma	26a7886a5d	Improve error reporting for posix system errors. (#8723 )	2021-08-12 23:19:02 -07:00
baijumeswani	217b2c9f93	Removing filelock import from ORTModule (#8722 )	2021-08-12 21:19:49 -07:00
stevenlix	f00933c41a	Update TensorRT parser to the latest (#8712 ) * update trt parser to the latest * update cgmanifest * update cgmanifest * update setup_env_trt to cuda11.4 * Update setup_env_trt.bat	2021-08-12 18:10:51 -07:00
Edward Chen	76d21bbeb2	Update Android API level to 30. (#8717 )	2021-08-12 18:01:13 -07:00
KeDengMS	d9d0228d0b	[Symbolic shape infer] fix a bug in loop/scan (#8694 ) In Loop/Scan, subgraph output may be partly in subgraph input	2021-08-12 17:57:28 -07:00
Dmitri Smirnov	1a8adb96fe	Reduce templatization of C API and refactor for InitOrtValue (#8700 ) Refactor for OrtInit Simplify C API Add ort_provider bridge interfaces	2021-08-12 16:51:18 -07:00
Edward Chen	89601ee6b3	[EP Partitioning Utils] Add check for assigned node. (#8473 ) Adds a check that a node is not already assigned to an EP before adding it to an EP partition.	2021-08-12 16:08:25 -07:00
fredster33	8d3c372dc9	Fix typo	2021-08-12 15:57:15 -07:00
Nathaniel McVicar	ce6675a74e	Avoid setting compile options on system libs for protobuf on Windows Signed-off-by: Nathaniel McVicar <namcvica@microsoft.com>	2021-08-12 15:56:39 -07:00
Changming Sun	5f74f198c1	Merge CPU/GPU nuget pipeline (#8683 ) Merge CPU/GPU nuget pipeline. The old GPU nuget pipeline will be only for DML. TODO: the result GPU package contains PDB files for some of the DLLs, but not all. It is due to the refactoring of CUDA EP to pluggable DLLs. At that time we forgot to copy the PDB files. However, I can't add them in now. Because currently the package is already 220MB large. If the missed PDB files were added, then it will be oversize. nuget.org doesn't accept >250MB packages.	2021-08-12 13:21:29 -07:00
Yulong Wang	3e8cabbc3e	[js/web] WebGL backend refactor (#8586 )	2021-08-12 12:30:49 -07:00
George Wu	7ff6a5e503	work around build warning on jetson (#8701 )	2021-08-12 08:37:25 -07:00
dependabot[bot]	333ef3c089	Bump path-parse from 1.0.6 to 1.0.7 in /js/common Bumps [path-parse](https://github.com/jbgutierrez/path-parse) from 1.0.6 to 1.0.7. - [Release notes](https://github.com/jbgutierrez/path-parse/releases) - [Commits](https://github.com/jbgutierrez/path-parse/commits/v1.0.7) --- updated-dependencies: - dependency-name: path-parse dependency-type: indirect ... Signed-off-by: dependabot[bot] <support@github.com>	2021-08-11 23:35:34 -07:00
Zhang Lei	76dfe8108b	Optimize quantized LSTM (#8634 ) * optimize some lstm gate computation. Remove no need string constructions. * change gcc optimization flags for computation bound logics in rnn_helpers * better qgemm for M=1 * Some improve on avx512 * add condition to limit GCC related marcros * Correct QGemm assembly for M=1 AVX2 optimization to pass mlas_test. * Fix rnn_helper build issue for wasm. * better asm code here according to feedbacks. * Remove customized vectorize and unroll option for GCC. Using restrict on some function to help GCC to correctly vectorize it. Rewrite clip_add_bias() to let GCC correctly vectorize it. * Better restrict semantic for merge_lstm_gates_to_memory() by adding in_place(). Add MSC __restrict for the clip_add_bias() mthod to vectorize correctly. * Force CI restart as it stucked by the onnxruntime-python-checks-ci-pipeline which can not restart.	2021-08-11 22:02:18 -07:00
Adrian Tsai	caacf249c5	Disable candy_opset9 WinML model test on Qualcomm Adreno (#8647 ) Bug #31652854 also repros on Qualcomm Adreno (down to the exact same pixel). This change disables this model test for Qualcomm, in addition to the existing disablement for Intel.	2021-08-11 17:57:12 -07:00
liqun Fu	bec24ca4c1	create packaging pipeline to support cuda11.4 (#8663 )	2021-08-11 17:44:57 -07:00
Zhang Lei	c6ef6b5bc8	Subgraph support for quantization tools (#8012 ) By default, not do enable subgraph quantization to make it consistent with existing behavior. It should be OK to enable it at quantize_dynamic mode with extra_options.	2021-08-11 16:35:52 -07:00
Changming Sun	c5c5d3499b	Rewrite dockerfiles/Dockerfile.arm32v7 (#8686 )	2021-08-11 15:25:04 -07:00
Tang, Cheng	de2a53e46d	[eager mode] fix build and support customize shared provider entry point (#8680 ) * fix build break * support customize the name of shared provide lib's entry point * fix non training build * check error code * check return code	2021-08-11 15:10:35 -07:00
Tianlei Wu	f661c18654	Fix attention perf regression (#8682 ) * undo change in attention cpu * fix perf regression * disable persistent softmax by default	2021-08-11 12:07:18 -07:00
harshithapv	c24335246b	Support bool type for Pad Op and fix Unsqueeze in Tile grad for Opset 13 (#8602 ) * changes * tile grad unsqueeze fix for opset 13 * clean up * remove bool support for opset 2 to 12 for Pad as it is not supported. * Copy OperatorKernels.md from artifacts of Windows CI build.	2021-08-11 11:21:02 -07:00
Guoyu Wang	a13daf550b	iOS Coacopods spec fix (#8678 )	2021-08-11 10:11:44 -07:00

1 2 3 4 5 ...

5375 commits