Introduce I64toI32 export pass (#3727) Summary: Pull Request resolved: https://github.com/pytorch/executorch/pull/3727 ## Context A number of `nn.Module`s targeting the Vulkan delegate use `i64` dtype for operators, inputs, and outputs. This is because `i64` is the default for many `torch` functions. Since [`i64` dtype is a Vulkan extension](https://registry.khronos.org/vulkan/specs/1.3-extensions/man/html/VK_KHR_shader_atomic_int64.html), it is not supported on all Vulkan devices. Hence, we introduce an export pass that converts the majority of the graph to `i32` dtype: 1. For each `i64` dtype input, append a node to convert the output tensor to `i32` dtype. 2. For each operator yielding `i64` dtype, replace it with `i32` dtype. 3. For each `i64` dtype output, prepend a node to convert the input tensor to `i64` dtype. ## Example Take this simple model compiled to one Clamp operation with input `x = (torch.randint(low=-100, high=100, size=(5, 5)),)`. By default, `torch.randint` uses `i64` dtype. ``` class ClampModule(torch.nn.Module): def __init__(self): super().__init__() def forward(self, x): x = torch.clamp(x, min=-3) return x ``` If it's compiled with such `i64` input, the export pass rewrites the graph to be the equivalent of the following. ``` class GoalModule(torch.nn.Module): def __init__(self): super().__init__() def forward(self, x): x = x.to(torch.int32) x = torch.clamp(x, min=-3) x = x.to(torch.int64) return x ``` ghstack-source-id: 229431025 exported-using-ghexport Reviewed By: SS-JIA Differential Revision: D57649650 fbshipit-source-id: ba36f1a563c2d97021b8f7d1ff1929d29223c43d
ExecuTorch is an end-to-end solution for enabling on-device inference capabilities across mobile and edge devices including wearables, embedded devices and microcontrollers. It is part of the PyTorch Edge ecosystem and enables efficient deployment of PyTorch models to edge devices.
Key value propositions of ExecuTorch are:
For a comprehensive technical overview of ExecuTorch and step-by-step tutorials, please visit our documentation website for the latest release (or the main branch).
We welcome any feedback, suggestions, and bug reports from the community to help us improve our technology. Please use the PyTorch Forums for discussion and feedback about ExecuTorch using the ExecuTorch category, and our GitHub repository for bug reporting.
We recommend using the latest release tag from the Releases page when developing.
executorch ├── backends # Backend delegate implementations. ├── build # Utilities for managing the build system. ├── codegen # Tooling to autogenerate bindings between kernels and the runtime. ├── configurations ├── docs # Static docs tooling. ├── examples # Examples of various user flows, such as model export, delegates, and runtime execution. ├── exir # Ahead-of-time library: model capture and lowering APIs. | ├── _serialize # Serialize final export artifact. | ├── backend # Backend delegate ahead of time APIs | ├── capture # Program capture. | ├── dialects # Op sets for various dialects in the export process. | ├── emit # Conversion from ExportedProgram to ExecuTorch execution instructions. | ├── operator # Operator node manipulation utilities. | ├── passes # Built-in compiler passes. | ├── program # Export artifacts. | ├── serde # Graph module serialization/deserialization. | ├── verification # IR verification. ├── extension # Extensions built on top of the runtime. | ├── android # ExecuTorch wrappers for Android apps. | ├── apple # ExecuTorch wrappers for iOS apps. | ├── aten_util # Converts to and from PyTorch ATen types. | ├── data_loader # 1st party data loader implementations. | ├── evalue_util # Helpers for working with EValue objects. | ├── gguf_util # Tools to convert from the GGUF format. | ├── kernel_util # Helpers for registering kernels. | ├── memory_allocator # 1st party memory allocator implementations. | ├── module # A simplified C++ wrapper for the runtime. | ├── parallel # C++ threadpool integration. | ├── pybindings # Python API for executorch runtime. | ├── pytree # C++ and Python flattening and unflattening lib for pytrees. | ├── runner_util # Helpers for writing C++ PTE-execution tools. | ├── testing_util # Helpers for writing C++ tests. | ├── training # Experimental libraries for on-device training ├── kernels # 1st party kernel implementations. | ├── aten | ├── optimized | ├── portable # Reference implementations of ATen operators. | ├── prim_ops # Special ops used in executorch runtime for control flow and symbolic primitives. | ├── quantized ├── profiler # Utilities for profiling runtime execution. ├── runtime # Core C++ runtime. | ├── backend # Backend delegate runtime APIs. | ├── core # Core structures used across all levels of the runtime. | ├── executor # Model loading, initalization, and execution. | ├── kernel # Kernel registration and management. | ├── platform # Layer between architecture specific code and portable C++. ├── schema # ExecuTorch PTE file format flatbuffer schemas. ├── scripts # Utility scripts for size management, dependency management, etc. ├── sdk # Model profiling, debugging, and introspection. ├── shim # Compatibility layer between OSS and Internal builds ├── test # Broad scoped end-to-end tests. ├── third-party # Third-party dependencies. ├── util # Various helpers and scripts.
ExecuTorch is BSD licensed, as found in the LICENSE file.