Add custom ops registration to examples

Summary:
## Context
I plan to add 2 (or 3) examples for different custom ops registration mechanisms. User should be able to use any of these options to use their custom ops. A proper README.md will be added.

## Solution
For the first option, we support the traditional PyTorch op registration python API. This requires users to write python implementations of both functional op and out variant op, like demonstrated in this diff. Note that those ops are only being registered into PyTorch JIT runtime for EXIR to consume. We also use buck2 target macro `executorch_generated_lib` to register custom ops to Executorch runtime.

For the second option, we want to leverage the C++ kernel user wrote for Executorch runtime, treat it as a valid PyTorch op kernel and register it into PyTorch JIT runtime. This way we don't have to write any python kernel. This can be done through CMake build, by pulling in PyTorch C++ dependency, then enabling ATen mode. This will be done once CMake diff D47927863 is landed.

The third option will be the same as the first but on CMake build system.

Note that CMake and Buck2 will then have different capabilities because pulling PyTorch C++ lib in Buck2 can't reuse the existing BUCK files.

Reviewed By: cccclai

Differential Revision: D48054313

fbshipit-source-id: 15fe77a4a69f3260b8fe09d7ce51b2f1e92cce68
8 files changed
tree: b5e6a71fa8c329e7f8b2ec0521eef6a410264472
  1. .ci/
  2. .github/
  3. backends/
  4. build/
  5. bundled_program/
  6. codegen/
  7. configurations/
  8. docs/
  9. examples/
  10. exir/
  11. extension/
  12. kernels/
  13. profiler/
  14. runtime/
  15. schema/
  16. scripts/
  17. sdk/
  18. shim/
  19. test/
  20. third-party/
  21. util/
  22. .buckconfig
  23. .clang-tidy
  24. .gitignore
  25. .gitmodules
  26. CMakeLists.txt
  27. CODE_OF_CONDUCT.md
  28. LICENSE
  29. pyproject.toml
  30. README.md
  31. setup.py
README.md

ExecuTorch

A unified ML software stack within the PyTorch platform for edge devices. It defines new compiler entry points as well as a state-of-art runtime.

Why ExecuTorch?

Compared to the legacy Lite Interpreter, there are some major benefits:

  • Performance wins compared to Lite Interpreter
    • Faster (orders of magnitude lower framework tax in both DSP and CPU)
    • Much smaller binary size, 1.5 MB vs 30 KB without operators.
    • Smaller memory footprint because we do ahead of time memory planning in ExecuTorch and also have clear granular control over where the runtime allocations are done.
  • Long term alignment with the direction of PyTorch infrastructure
    • Lite Interpreter relies on TorchScript, which is being phased out; ExecuTorch is the planned replacement for Lite Interpreter.
  • Model Authoring & Productivity gains
    • More and better defined entry points to perform model, device, and/or use-case specific optimizations (e.g. better backend delegation, user-defined compiler transformations, default or user-defined memory planning, etc)
    • Ability to lower constructs like dynamic control flow to run on device.

Design goals

  • Minimal binary size (< 50KB not including kernels)
  • Minimal framework tax: loading program, initializing executor, kernel and backend-delegate dispatch, runtime memory utilization
  • Portable (cross-compile across many toolchains)
  • Executes ATen kernels (or ATen custom kernels)
  • Executes custom op kernels
  • Supports inter op asynchronous execution
  • Supports static memory allocation (heapless)
  • Supports custom allocation across memory hierarchies
  • Supports control flow needed by models
  • Allows selective build of kernels
  • Allows backend delegation with lightweight interface

Quick Links

Quick Links for Partners

Directory Structure [WIP]

executorch
├── backends                        #  1st party backend implementations.
|   ├── xnnpack
|   ├── vulkan
|   ├── backend_api.py              # TODO move to exir/backend
|   ├── backend_details.py          # TODO move to exir/backend
|   ├── partioner.py                # TODO move to exir/backend
├── build                           #  Utilities for managing the build system.
├── bundled_program                 #  Utilities for attaching reference inputs and outputs to models. TODO move to extension
├── codegen                         #  Tooling to autogenerate bindings between kernels and the runtime. TODO move to tool
├── configurations                  #  TODO delete this
├── docs                            #  Static docs tooling
├── examples                        #  Examples of various user flows, such as model export, delegates, and runtime execution.
|   ├── executor_runner
|   ├── export
|   ├── models
├── exir                            #  Ahead of time library, model capture and lowering apis.
|   ├── capture                     #  Program capture.
|   ├── dialects                    #  Op sets for various dialects in the export process.
|   ├── emit                        #  Conversion from ExportedProgram to Executorch execution instructions.
|   ├── program                     #  Export artifacts.
|   ├── serialize                   #  Serialize final export artifact.
├── extension                       #  Extensions built on top of the runtime.
|   ├── aten_util
|   ├── data_loader                 # 1st party data loader implementations.
|   ├── memory_allocator            # 1st party memory allocator implementations.
|   ├── pybindings                  # Python api for executorch runtime.
|   ├── pytree                      # C++ and Python flattening and unflattening lib for pytrees.
|   ├── testing_util
├── kernels                         #  1st party kernel implementations.
|   ├── aten
|   ├── optimized
|   ├── portable                    #  Reference implementations of ATen operators.
|   ├── prim_ops                    #  Special ops used in executorch runtime for control flow and symbolic primitives.
|   ├── quantized
├── profiler                        #  Utilities for profiling. TODO delete in favor of ETDump in sdk/
├── runtime                         #  core cpp runtime of executorch
|   ├── backend                     #  Backend definition and registration.
|   ├── core                        #  Core structures used across all levels of the runtime
|   ├── executor                    #  Model loading, initalization, and execution.
|   ├── kernel                      #  Kernel registration and management.
|   ├── platform                    #  Layer between architecture specific code and user calls.
├── schema                          #  Executorch program definition, TODO move under serialization/
├── scripts                         #  Utility scripts for size management, dependency management, etc.
├── sdk                             #  Model profiling, debugging, and introspection: NOT READY YET FOR OSS USE
├── shim                            #  Compatibility layer between OSS and Internal builds
├── test                            #  Broad scoped end2end tests
├── third-party                     #  third-party dependencies
├── util                            #  TODO delete this

License

ExecuTorch is BSD licensed, as found in the LICENSE file.