Skip to main content

Lite Transformer

This document describes how to run the Lite Transformer English-to-Chinese example on the NPU.

info

Refer to Model Zoo Download for the example.

Lite Transformer uses encoder and decoder sub-models.

Lite Transformer example directory structure:

$ tree ./
./
├── CMakeLists.txt
├── convert_model_decoder
├── convert_model_encoder
├── include
├── model
│ ├── bpe_order.txt
│ ├── cw_token_map_order.txt
│ ├── dict_order.txt
│ ├── lite_transformer_decoder_16_int16_a733.nb
│ ├── lite_transformer_encoder_16_int16_a733.nb
│ ├── position_embed.bin
│ └── token_embed.bin
├── src
└── README.md

Model Conversion

Enter the container development environment first. See Create and Start Container in the Model Zoo download page.

info

Select the Docker image that matches the NPU:

  • A733: ubuntu-npu:v2.0.10.2
  • T527: ubuntu-npu:v1.8.13

Download the floating-point ONNX models from the Allwinner netdisk:

Extract the BPE and dictionary files into the lite_transformer directory.

X86 Linux PC
docker exec -it model-zoo /bin/bash

Convert encoder

X86 Linux PC
cd /workspace/examples/lite_transformer/convert_model_encoder/
./convert_model_env.sh
./pegasus_import.sh lite_transformer_encoder_16
./pegasus_quantize.sh lite_transformer_encoder_16 int16 2
X86 Linux PC
./pegasus_export_ovx_nbg.sh lite_transformer_encoder_16 int16 a733

Convert decoder

X86 Linux PC
cd /workspace/examples/lite_transformer/convert_model_decoder/
./convert_model_env.sh
./pegasus_import.sh lite_transformer_decoder_16
./pegasus_quantize.sh lite_transformer_decoder_16 int16 4
X86 Linux PC
./pegasus_export_ovx_nbg.sh lite_transformer_decoder_16 int16 a733

The exported models are stored in the ../model directory.

Build the Example

Then compile the example. Exit the container first, then run the commands below.

Configure the cross-compilation toolchain first.

info

Skip this step if you have already configured it in another example.

X86 Linux PC
cd ../../../0-toolchains/

Download the toolchain from this link, put it in 0-toolchains/, then run:

X86 Linux PC
tar -xvf gcc-arm-10.2-2020.11-x86_64-aarch64-none-linux-gnu.tar.xz
cd ../examples/lite_transformer/
X86 Linux PC
../build_linux.sh -t a733 -s debian11

Model Deployment

After compilation, the example will be installed in the install directory. You can use scp to transfer it to the board.

Configure NPU Driver

info

You can skip this step if you have already configured NPU driver in other examples.

Transfer the driver library to the board's lib directory via scp.

  • A733 corresponds to the common/npuruntime/lib_linux_aarch64/A733 directory
  • T527 corresponds to the common/npuruntime/lib_linux_aarch64/T527 directory

Then execute the following command to export to environment variables.

Radxa SBC
echo 'export LD_LIBRARY_PATH=$HOME/lib:$LD_LIBRARY_PATH' >> ~/.bashrc

Run Example

After configuring the driver, you can run the example.

tip

For T527 platform, you need to first enable NPU by referring to the A5E's "Enable NPU on Board" documentation, then use the following command to grant the current user permission to use /dev/vipcore.

Radxa SBC
sudo chmod 777 /dev/vipcore
Radxa SBC
cd lite_transformer_demo_linux_a733/
Radxa SBC
chmod +x ./lite_transformer_demo_a733
./lite_transformer_demo_a733 -nb0 model/lite_transformer_encoder_16_int16_a733.nb -nb1 model/lite_transformer_decoder_16_int16_a733.nb -i "so big"

The running result is as follows:

$ ./lite_transformer_demo_a733 -nb0 model/lite_transformer_encoder_16_int16_a733.nb -nb1 model/lite_transformer_decoder_16_int16_a733.nb -i "so big"
encoder_path=model/lite_transformer_encoder_16_int16_a733.nb, decoder_path=model/lite_transformer_decoder_16_int16_a733.nb, input_strings=so big
VIPLite driver software version 2.0.3.2-AW-2024-08-30
nbg name=model/lite_transformer_encoder_16_int16_a733.nb, size: 3838272.
create network 0: 2533 us.
prepare network: 442 us.
nbg name=model/lite_transformer_decoder_16_int16_a733.nb, size: 19421952.
create network 1: 11027 us.
prepare network: 406 us.

input sentence:so big
output token: 2 83 139 676 84 2
output_strings: 如此巨大
inference time: 24.000 ms
destroy npu finished.
~NpuUint.

This performance data only calculates the time consumption of model inference. Unless otherwise specified, it does not include the time consumption of pre-processing and post-processing.

SoCNPUModelInput ResolutionNetwork Creation TimeNetwork Preparation TimeSingle Frame Inference TimePost-processing TimeTotal TimeFrame Rate
Allwinner A733Vivante VIP9000lite-transformer16 tokens13.6 ms0.8 ms24.0 ms38.4 ms41.7 FPS

    You need to be logged into GitHub to post a comment. If you are already logged in, please ignore this message.

    Radxa-docs © 2026 by Radxa Computer (Shenzhen) Co.,Ltd. is licensed under CC BY 4.0