Add new mps runtime with support for FP16, lifted and unlifted graphs (iOS15+, macOS12+) (#1655) Summary: This PR changes the MPS Backend runtime to support for **iOS15+/macOS12+** (previous runtime was limited to iOS17/macOS14 only). Additionally, this PR contains changes such as support for both lifted and unlifted graphs, support for torch.export API, optimizations for FP16 and faster model loading during runtime (more information in the summary). **Summary of changes:** - Add support for running the models in FP16 (https://github.com/DenisVieriu97/executorch/pull/7 georgepaw) - Replace the previous MPS runtime from ExecuTorch which was relying on `iOS17` / `macOS` Sonoma APIs for serialization of the MPSGraphExecutable. Instead of creating the MPSGraph nodes and serializing them during AOT, create the corresponding entries of the EdgeIR nodes in the FlatBuffer, and parse them during runtime to construct the graph. This removes any dependency on iOS17 / macOS 14.0 APIs. - Add support for node visitor pattern: - Each node visitor class visits an op and serializes the data into MPSTensors and MPSNodes which get appended in the flatbuffer - The entries from the FlatBuffer are parsed in the runtime and based on them the MPSGraph is constructed (for more info see `MPSGraphBuilder` class and corresponding ops from `operators/` folder. - This method removes the additional read and writes to disk of the MPSGraphExecutable (once during AOT and once during runtime). **Models summary:** | Model | FP16 | FP32 | | :---: | :---: | :---: | mul | <ul><li>- [x] </li> | <ul><li>- [x] </li> | add_mul | <ul><li>- [x] </li> | <ul><li>- [x] </li> | linear | <ul><li>- [x] </li> | <ul><li>- [x] </li> | edsr | <ul><li>- [x] </li> | <ul><li>- [x] </li> | Mobilebert | <ul><li>- [ ] </li> | <ul><li>- [x] </li> | mv2 | <ul><li>- [x] </li> | <ul><li>- [x] </li> | mv3 | <ul><li>- [x] </li> | <ul><li>- [x] </li> | vit | <ul><li>- [x] </li> | <ul><li>- [x] </li> | w2l | <ul><li>- [x] </li> | <ul><li>- [x] </li> | ic3 | <ul><li>- [x] </li> | <ul><li>- [x] </li> | ic4 | <ul><li>- [x] </li> | <ul><li>- [x] </li> | resnet18 | <ul><li>- [x] </li> | <ul><li>- [x] </li> | resnet50 | <ul><li>- [x] </li> | <ul><li>- [x] </li> | Llama2 | <ul><li>- [ ] </li> | <ul><li>- [x] </li> | emformer_join | <ul><li>- [x] </li> | <ul><li>- [x] </li> | emformer_predict | <ul><li>- [ ] </li> | <ul><li>- [ ] </li> | emformer_transcribe | <ul><li>- [x] </li> | <ul><li>- [x] </li> | dl3 | <ul><li>- [ ] </li> | <ul><li>- [ ] </li> | Pull Request resolved: https://github.com/pytorch/executorch/pull/1655 Reviewed By: cccclai Differential Revision: D52929916 Pulled By: shoumikhin fbshipit-source-id: 8bd2ed124311744ebe19fc17eb0ff508621f974a
ExecuTorch is an end-to-end solution for enabling on-device inference capabilities across mobile and edge devices including wearables, embedded devices and microcontrollers. It is part of the PyTorch Edge ecosystem and enables efficient deployment of PyTorch models to edge devices.
Key value propositions of ExecuTorch are:
For a comprehensive technical overview of ExecuTorch and step-by-step tutorials, please visit our documentation website.
This is a preview version of ExecuTorch and should be used for testing and evaluation purposes only. It is not recommended for use in production settings. We welcome any feedback, suggestions, and bug reports from the community to help us improve the technology. Please use the PyTorch Forums for discussion and feedback about ExecuTorch using the ExecuTorch category, and our GitHub repository for bug reporting.
The ExecuTorch code and APIs are still changing quickly, and there are not yet any guarantees about forward/backward source compatibility. We recommend using the latest v#.#.# release tag from the Releases page when experimenting with this preview release.
executorch ├── backends # Backend delegate implementations. ├── build # Utilities for managing the build system. ├── bundled_program # Utilities for attaching reference inputs and outputs to models. TODO move to extension ├── codegen # Tooling to autogenerate bindings between kernels and the runtime. TODO move to tool ├── configurations # TODO delete this ├── docs # Static docs tooling ├── examples # Examples of various user flows, such as model export, delegates, and runtime execution. ├── exir # Ahead of time library, model capture and lowering apis. | ├── _serialize # Serialize final export artifact. | ├── backend # Backend delegate ahead of time APIs | ├── capture # Program capture. | ├── dialects # Op sets for various dialects in the export process. | ├── emit # Conversion from ExportedProgram to ExecuTorch execution instructions. | ├── passes # Built-in compiler passes. | ├── program # Export artifacts. | ├── verification # IR verification. ├── extension # Extensions built on top of the runtime. | ├── aten_util | ├── data_loader # 1st party data loader implementations. | ├── memory_allocator # 1st party memory allocator implementations. | ├── pybindings # Python api for executorch runtime. | ├── pytree # C++ and Python flattening and unflattening lib for pytrees. | ├── testing_util ├── kernels # 1st party kernel implementations. | ├── aten | ├── optimized | ├── portable # Reference implementations of ATen operators. | ├── prim_ops # Special ops used in executorch runtime for control flow and symbolic primitives. | ├── quantized ├── profiler # Utilities for profiling. TODO delete in favor of ETDump in sdk/ ├── runtime # core cpp runtime of executorch | ├── backend # Backend delegate runtime APIs | ├── core # Core structures used across all levels of the runtime | ├── executor # Model loading, initalization, and execution. | ├── kernel # Kernel registration and management. | ├── platform # Layer between architecture specific code and user calls. ├── schema # ExecuTorch program definition, TODO move under serialization/ ├── scripts # Utility scripts for size management, dependency management, etc. ├── sdk # Model profiling, debugging, and introspection. ├── shim # Compatibility layer between OSS and Internal builds ├── test # Broad scoped end2end tests ├── third-party # third-party dependencies ├── util # TODO delete this
ExecuTorch is BSD licensed, as found in the LICENSE file.