top of page

How to Convert Models to DLA for MediaTek Genio 700

Writer: Joel Gomez Araya
Joel Gomez Araya
14 hours ago
8 min read
Diagram showing the conversion of an ONNX model into DLA format for deployment on MediaTek Genio 700.
This tutorial will guide you through converting AI models to DLA for deployment in MediaTek Genio 700.

The MediaTek Genio 700 is an edge AI platform that integrates a dedicated Neural Processing Unit (NPU) for accelerating deep learning workloads efficiently on-device. To run neural networks on this hardware, models must be converted into hardware-specific binaries, known as Deep Learning Archive (DLA), which can be executed directly by the MediaTek NPU. If you already have a trained model, the next step is to convert it into a DLA file using the MediaTek NeuroPilot toolchain. If you’re new to the MediaTek Genio 700, check out our First Steps with the Mediatek Genio 700 guide.


In this guide, we will explain how to convert models to DLA deployment on the MediaTek Genio 700, presenting a step-by-step workflow, using a practical example of the optimization process required for deployment on the MediaTek NPU.



NeuroPilot Environment for DLA Conversion


Before converting a model to DLA, it is important to understand the NeuroPilot environment. NeuroPilot is a collection of developer tools and APIs designed to build efficient AI applications on MediaTek platforms. It enables developers to deploy AI workloads on edge devices with high performance.


NeuroPilot evolves across multiple SDK generations. For example, the MediaTek Genio 700 platform is associated with NeuroPilot 6, while newer versions such as NeuroPilot 8 provide updated tooling. Many tools maintain backward compatibility with previous versions, making them suitable even when targeting platforms designed for earlier releases.


Required Tools for the Conversion Workflow


To perform the conversion to DLA, we will use two main tools from the NeuroPilot suite: the Converter Tool and the Neuron SDK.

Diagram illustrating the NeuroPilot deployment workflow. The host PC uses the Converter Tool and Neuron SDK compiler to prepare AI models, while the target device uses the Neuron Runtime to execute the compiled model.
NeuroPilot toolchain overview. Taken from: MediaTek

MTK Converter


The MTK Converter enables model conversion from multiple frameworks into TensorFlow Lite. It supports frameworks such as TensorFlow (v1 and v2), PyTorch (v1 and v2), ONNX, and Caffe. Additionally the converter tool is capable of quantizing the model with different configurations, such as 8bit asymmetric quantization or 16-bit symmetric quantization. It maintains backward compatibility with previous versions of NeuroPilot.


Neuron SDK


Neuron SDK is a set of tools and APIs that allow developers to compile their models into DLA format for deployment on MediaTek platforms. The most relevant tools for this workflow are:


  • Neuron Compiler (ncc-tflite): Performs offline compilation by converting TensorFlow Lite models into hardware-specific DLA binaries. This tool is intended to run on the host, although it can also be executed on the device in certain scenarios.


  • Neuron Runtime (neuronrt): This runtime is deployed on the target device. It is responsible for loading and executing .dla models on the NPU. It is used for inference and validation of the model.


ONNX to DLA Conversion Process on MediaTek Genio 700


The process to convert a model to a DLA file is divided into three main stages:

  1. Environment setup

  2. Model conversion from ONNX to TensorFlow Lite

  3. Compilation from TensorFlow Lite to DLA


Environment Setup for DLA Model Conversion


The first step consists of preparing the working environment, including the required tools, dependencies, and project structure.


  1. Create Project Workspace for Model Conversion

Start by creating a dedicated directory to organize all the files involved in the conversion process:



  1. Set Up Your Python Environment

Create and activate a virtual environment to isolate dependencies:



This step is necessary to avoid errors in dependencies.

Additionally, you could use PyEnv to manage your virtual environment instead of venv.


  1. Install Dependencies for Model Conversion Pipeline

Install the dependencies needed for model conversion:



Use the following command to install it:

  1. Download and Install Converter Tool

The Converter Tool can be downloaded from the Public NeuroPilot Downloads page. At the time of this writing, the latest version of the public converter tool is 8.13.0. Once downloaded, extract the package:

Inside the extracted directory, you will find multiple .whl packages. You must install the one that matches your Python version:


Replace <version> with version in the name of files and x with your Python version.

So copy the .whl to your work directory:

And install the package. For example, if you are using Python 3.10, select the package that includes cp310 in its filename.

To check if installation was successful, use:



  1. Download Neuron SDK


Next, download the Neuron SDK from the NeuroPilot 6 downloads page. Navigate to the 7.4 Neuron SDK section and locate the release compatible with the MediaTek Genio 700 (MT8390). The package is named as follows:


After downloading, extract the archive:

Copy the NeuronSDK directory to your work directory:

  1. Convert Your Model to ONNX

Once the NeuroPilot tools are ready, the next step is to ensure that the model is in ONNX format. If your model is not already in ONNX format, you will need to convert it first.

Add your model to your work directory:

In this example, we use a YOLO pose PyThorch model, so we need to convert it to ONNX format:


These options were chosen to improve compatibility of the YOLO Pose model with the MediaTek conversion pipeline. In particular:

  • dynamic=False generates a model with fixed input dimensions, which ensures compatibility for the converter to process

  • nms=False disables the Non-Maximum Suppression (NMS) operation, which is not supported by the MediaTek NPU and must be handled outside the model.

  • The imgsz parameter specifies the model input resolution and can be adjusted according to your application requirements.

  • The batch=1 setting is used in this example because it matches the intended deployment scenario, although other batch sizes may be supported depending on the model and deployment pipeline.


  1. Prepare Calibration Dataset

Download COCO 2017 dataset for yolo pose model calibration using the following bash script:


Execute with:

Then format the images of the dataset with the following script:


Then execute:


Convert ONNX Model to TensorFlow Lite for MediaTek Genio 700


The conversion of ONNX model to TFLite requires creating a script with the following steps:

  • Prepare the model for conversion 

  • Initialize and configure the MediaTek converter

  • Apply quantization using a calibration dataset

  • Convert the model and generate the final TFLite output


To convert ONNX models to TensorFlow Lite for deployment on the MediaTek Genio 700, create a new file export_onnx_to_tflite.py with the following script:


We will now study each portion of the code.


  1. Prepare the ONNX Model For Conversion

The ONNX model is first loaded and adapted to meet the requirements of the MediaTek toolchain. This process involves adjusting the model's ONNX Intermediate Representation (IR) version to ensure compatibility with the converter, modifying the input tensor format from NCHW to NHWC (as expected by TensorFlow Lite), and inserting a transpose operation at the beginning of the graph so that the original model behavior is preserved. 


TensorFlow Lite expects input tensors in NHWC format, while many frameworks such as PyTorch export models in NCHW format. To bridge this mismatch without modifying the original model structure, a transpose operation is inserted at the input layer. If the model is already in NHWC format, this transpose can be omitted.


To determine which format your model uses, you can inspect its input tensor shape using tools such as Netron. For example, an input shape of [N, C, H, W] indicates NCHW format, while [N, H, W, C] indicates NHWC format.


For example, a YOLO model exported from PyTorch commonly uses an input shape such as [1, 3, 640, 640] (NCHW), whereas TensorFlow Lite models typically use [1, 640, 640, 3] (NHWC).

The IR version is also adjusted to ensure compatibility with the MediaTek converter, which may not support newer ONNX IR versions.


Once these changes are applied and validated, the model is structurally ready for the subsequent conversion steps.


  1. Initialize and Configure the MediaTek Converter

A converter object is created using the prepared model. The converter is configured to handle operators that are not directly supported by the NPU of MediaTek Genio 700. For example, some operations (such as SiLU) are decomposed into compatible operations to ensure compatibility.


This step ensures the model can be executed correctly on MediaTek hardware.



If you need to verify the compatible operations, use the following command on MediaTek Genio 700:

If you have an error related to some specific operation, you could find a decompose for the operator you need by going to Converter API Documentation Onnx Converter and search for decompose operations properties.


  1. Apply Model Quantization for NPU Deployment

Quantization reduces model size and improves performance on the NPU. If it is enabled, a calibration dataset is used to estimate value ranges inside the model and the model is converted from floating point to lower precision. For example, the 16W16A configuration specifies 16-bit precision for both weights and activations, which provides a balance between performance and accuracy on the target NPU. 

The calibration dataset is loaded from .npy files stored in a calibration directory.



  1. Convert Model to TensorFlow Lite

The configured converter is used to generate a .tflite model, with parameter tflite_op_export_spec="builtin_ignore_version" to ensure that the TFLite model generated is compatible with the MediaTek Genio 700, avoiding some MediaTek Custom Operations (MTKEXT) that may cause problems.


  1. Execute Converter Script

Use the following command to export the ONNX model to TFLite:

Where:

  • --quantize enables model quantization.

  • --dataset_dir specifies the calibration dataset directory.

  • --input_shapes defines the model input shape in NHWC format.



Convert TensorFlow Lite Model to DLA for MediaTek Genio 700

Once the TensorFlow Lite model is generated, it can be compiled into a DLA format using the Neuron Compiler provided by the Neuron SDK. This process involves two main steps. 

  • First, the environment must be configured by setting the LD_LIBRARY_PATH to include the Neuron SDK libraries, so the compiler can locate the required runtime libraries.

  • Then, the ncc-tflite compiler is used to convert the .tflite model into a DLA binary, specifying the target architecture. For example, the mdla3.0 architecture corresponds to the NPU version available on MediaTek Genio 700.

It is important to note that the MediaTek Genio 700 does not support FP32 models. If the input model is in FP32 precision, the --relax-fp32 flag should be used to allow the compiler to automatically convert operations to FP16, ensuring compatibility with the hardware. 


Validate DLA Model on MediaTek Genio 700 NPU


Once the .dla model is generated, the next step is to validate that it can be executed correctly on the MediaTek Genio 700 NPU using the Neuron Runtime.


  1. Transfer DLA Model to MediaTek Genio 700


Copy your DLA file to the target device using scp (Secure Copy Protocol), a command-line tool for transferring files over SSH. Replace x.x.x.x with the IP address of your MediaTek Genio 700 board:

  1. Inspect DLA Model Input and Output

Use neuronrt to inspect the model’s expected input and output tensors:



As shown above, the model expects an input of 2457600 bytes, so the input provided during inference must match this size and layout.


  1. Generate a Test Input


Create a dummy input file that matches the model required input size, this input file is only for validation. Use representative data for accurate inference results.

  1. Run DLA Inference on MediaTek NPU


Run inference using Neuron Runtime for testing DLA functionality:


At this point, the model has been successfully converted and executed on the MediaTek Genio 700 NPU, confirming that the DLA generation process was completed correctly and the model is ready for deployment. By following this workflow, you can take trained AI models and prepare them for efficient execution on the MediaTek platform, from ONNX conversion and quantization to final compilation and validation on the target hardware.


Need Help Bringing AI Applications to MediaTek Genio 700?


The MediaTek Genio 700 makes it possible to deploy AI models efficiently at the edge, but achieving production-grade performance, reliability, and maintainability requires more than successfully converting a model to DLA format. RidgeRun.ai specializes in optimizing embedded AI deployments, streamlining integration workflows, and building scalable edge AI solutions on MediaTek platforms.


We can assist with AI model optimization and quantization, efficient GStreamer integration, custom model deployment, hardware bring-up, benchmarking and performance measurements, and end-to-end integration of AI applications on embedded platforms.


Whether you're deploying object detection, pose estimation, computer vision, or other AI workloads for robotics, industrial automation, smart retail, or edge analytics, our team can help accelerate your path from prototype to production.


Contact RidgeRun's AI Engineering Services to learn how we can help you unlock the full potential of the MediaTek Genio 700 platform. Contact us at contactus@ridgerun.ai — let’s collaborate!





bottom of page