Skip to main content

Whisper-Base Example

This document explains how to use the QAI AppBuilder Python API to perform inference with the Whisper-Base speech recognition model using the Qualcomm® Hexagon™ Processor (NPU).

Supported devices

DeviceSoC
Fogwise® AIRbox Q900QCS9075

Install QAI AppBuilder​

tip
  1. Install QAI AppBuilder by following the QAI AppBuilder installation guide.

  2. Configure ADSP environment variables as described in Create ADSP environment variables.

Run the sample​

Install dependencies​

Install sample dependencies in the activated virtual environment:

Device
pip3 install requests tqdm qai-hub py3-wget Pillow torch torchvision opencv-python-headless audio2numpy

Run the script​

  • Enter the upstream samples directory

    Device
    cd qai-appbuilder/samples
  • Prepare input data (use the sample input if provided, or pass script arguments)

input audio

  • Run inference

    Device
    python3 audio/Speech_Recognition/whisper_base_en/whisper_base_en.py --chipset 9075

The default input is the bundled jfk.wav. On success the terminal prints a transcription such as:

Transcription: And so my fellow Americans, ask not what your country can do for you, ask what you can do for your country.
tip

The first run downloads encoder/decoder models via Qualcomm® AI Hub. Decoding needs the tiktoken vocabulary; if the device cannot reach openaipublic.blob.core.windows.net, pre-warm tiktoken.get_encoding("gpt2") on a machine with network access.

python3 run_inference.py --list
python3 run_inference.py --model whisper_base_en --args "--chipset 9075"

    You need to be logged into GitHub to post a comment. If you are already logged in, please ignore this message.

    Radxa-docs © 2026 by Radxa Computer (Shenzhen) Co.,Ltd. is licensed under CC BY 4.0