Skip to content

Automatic Code Optimization

July 2024 – Oct. 2024 · New York University Abu Dhabi (NYUAD)

Role
Research Intern, Automatic Code Optimization
Organization
New York University Abu Dhabi (NYUAD)

Compilers that optimise deep-learning code automatically need the model in a form they can work on, such as MLIR. During my remote research internship at NYU Abu Dhabi, my task was to find a reliable way to translate PyTorch models into MLIR. The long-term aim was large language models; in the time I had, I got it working for NLP models and other simpler models.

What I built

I worked on it alone. The pipeline imports a PyTorch model through torch-mlir's FX importer, the only torch-mlir importer still enabled at the time, and emits an MLIR module in one of three forms: the Torch dialect, Linalg-on-tensors or TOSA. The conversions themselves are torch-mlir's; my part was finding the path that worked and turning it into a pipeline with examples for BERT, RoBERTa, ResNet-18 and a linear regression. Documentation for that path was scarce, and people in the LLVM community on Discord pointed me in the right direction.

What I took from it

Sometimes the hardest part of the work is not the code but information that is scattered and hard to understand, and knowing whom to ask.

PyTorch modelFX importertorch-mlirMLIR output, one of:Torch dialectLinalg-on-tensorsTOSAPyTorch modelFX importertorch-mlirMLIR output, one of:Torch dialectLinalg-on-tensorsTOSA
Fig. 1 The conversion path in Model_To_MLIR: a PyTorch model is imported through torch-mlir’s FX importer and emitted as an MLIR module in the Torch dialect, Linalg-on-tensors or TOSA, depending on the selected output type. Drawn from the public repository’s examples. The dialect conversions themselves are torch-mlir’s. No performance figures.