GGUF, widely used in local LLMs, is commonly run on CPUs and NVIDIA GPUs along with llama.cpp. GGUF itself is a format for storing model weights and metadata, and where the actual processing is performed (CPU, GPU, or NPU) is determined by the runtime and backend. Just as backends such as CUDA are used, there are also backends that run GGUF on Qualcomm's Hexagon NPU.
Qualcomm's GenieX offers two paths: the llama.cpp path for handling GGUFs, and the QAIRT (Qualcomm AI Runtime) path compiled with Qualcomm AI Hub. Both can utilize the Hexagon NPU, but their supported model formats, deployment methods, and performance characteristics are not the same.
This article explains the differences between these two paths and the performance characteristics observed in each, based on actual measurements of the same Qwen3-4B series running on the Dragonwing IQ-9075 EVK with different distribution formats and quantization. For comparison, CPU execution was added, and measurements were taken using three paths: llama.cpp CPU, llama.cpp NPU, and QAIRT NPU.
Writer
Hidetaka Iwakura
Macnica Macnica Finesse Company
Technical Management Department, Technical Division 5, Section 2
What you will learn from this article
- A mechanism that allows GGUF to run on the Hexagon NPU
- Differences between the two NPU execution paths that GenieX offers
- Reasons for selecting Q4_0 when using the NPU with GGUF
- Measured differences between the three paths (prefill, response start time, CPU usage)
- Ease of bringing in items and NPU performance vary depending on the route.
Verification environment
• Board: Dragonwing IQ-9075 EVK (CPU has 8 cores)
• Model: Qwen3-4B (llama.cpp route is unsloth/Qwen3-4B-GGUF:Q4_0, QAIRT route is qualcomm/Qwen3-4B:W4A16)
Runtime: GenieX v0.4.0, QAIRT 2.45, llama.cpp Hexagon backend (included with GenieX v0.4.0)
The measurements primarily focus on speed and CPU utilization. This is not a comparison to demonstrate that the output quality of Q4_0 and W4A16 are equivalent.
Dragonwing IQ-9075 Specifications
This article discusses a high-performance edge AI SoC for IoT and embedded systems, specifically the EVK version (IQ9075-AA) with 100 TOPS. Key specifications are as follows:
|
item |
content |
|
SoCs |
Qualcomm QCS9075 |
|
NPU |
Hexagon V73 x 2 (Dual CDSP) |
|
Arithmetic performance |
100 TOPS (IQ9075-AA, total across 2 NPUs; 1 NPU used for single-model execution) |
|
operating temperature |
-40〜+115℃ |
|
Evaluation unit memory |
36GB |
Click here for product details
GenieX has two inference runtimes.
GenieX is an inference environment for running LLM and VLM on Qualcomm devices. Two runtimes are available through a common interface including a CLI, Python API, and OpenAI-compatible API. llama_cpp uses the llama.cpp engine, with its internal GGML Hexagon backend handling NPU execution. qairt is another runtime that uses Qualcomm AI Engine Direct (QNN/QAIRT).
The official documentation distinguishes between the runtime that handles GGUF as llama_cpp and the runtime that handles pre-compiled bundles from Qualcomm AI Hub as qairt.
|
item |
llama.cpp route |
QAIRT route |
|
Main model types |
GGUF |
QAIRT bundles by chipset |
|
Where to obtain the model |
Hugging Face, ModelScope, local files, etc. |
Qualcomm AI Hub |
|
Execution destination |
CPU / GPU / Hexagon NPU |
Hexagon NPU |
|
Typical quantization for Hexagon NPU |
Q4_0 |
W4A16 is the main one. |
|
Ease of bringing in your own model |
high |
Exporting is required for the target SoC. |
|
feature |
A wide range of GGUFs are easy to try out. |
Easy to optimize for chipsets |
Neither of these two approaches is always superior. The best approach differs depending on whether you want to quickly test any model or aim for high NPU performance on your target SoC.
The runtime backend determines where GGUF is executed.
GGUF is a file format and does not specify the execution destination itself. If GenieX supports the GGUF file, it can be executed on the CPU, GPU, or Hexagon NPU using the same file.
The GGUF stores information such as the model's weights, quantization format, and network structure. However, the llama.cpp backend, not the GGUF itself, determines which processor performs the matrix operations.
The GenieX llama.cpp runtime supports CPUs, Adreno GPUs, and Hexagon NPUs. Therefore, you can use the same file to change the processing destination as long as the GGUF supports the GenieX target architecture, quantization, and computation. This does not mean that any GGUF will work on all Qualcomm devices.
GenieX can retrieve models from various sources using the --model-hub option, supporting Qualcomm AI Hub, Hugging Face, ModelScope, Docker registries, and local file systems. It's also possible to load GGUF models that have already been retrieved from the local machine.
The following is a minimal example using an IQ-9075 EVK with GenieX v0.4.0 already installed. A network connection and sufficient storage space are required for the initial model download. For GenieX installation and supported platforms, please see below.Official QuickstartPlease check.
# Hexagon NPUで実行 geniex infer unsloth/Qwen3-4B-GGUF:Q4_0 \ --compute npu \ --ngl -1 \ --think=false # CPUで実行 geniex infer unsloth/Qwen3-4B-GGUF:Q4_0 \ --compute cpu \ --ngl 0 \ --think=false # 対応tensorをNPU、残りをCPUへ割り当てるhybrid経路(`--think=false`はQwen3の思考モードを無効化する指定) geniex infer unsloth/Qwen3-4B-GGUF:Q4_0 \ --compute hybrid \ --ngl -1 \ --think=false
--ngl -1 is a specification that requests llama.cpp to offload all layers. The name originates from llama.cpp and is an abbreviation for GPU layers, but the same specification is used for offloading to the Hexagon NPU. --compute npu is a fixed path to the HTP (Hexagon Tensor Processor, the computational part of the Hexagon NPU), and model loading or inference may fail due to unsupported calculations or memory allocation failures on the NPU side.
On the other hand, --compute hybrid allocates the corresponding tensors to the NPU and processes the rest with the CPU. Hybrids that include a fixed NPU and CPU fallback should be evaluated separately.
The quantization formats that can be implemented on the Hexagon NPU are limited.
GGUF has several quantization formats, including Q4_0, Q8_0, and Q4_K_M. K-quant is commonly used for CPU applications, but the official GenieX documentation recommends Q4_0 for the Hexagon NPU backend. This is because the Hexagon backend can only handle a limited number of quantization formats.
Upon examining the source code of upstream llama.cpp at the time of this verification, it was found that the Hexagon backend accepts F16, F32, IQ4_NL, MXFP4, Q4_0, Q4_1, and Q8_0, but K-quant is not supported. Furthermore, only Q4_0, Q8_0, and MXFP4 are repacked on the NPU side.
This is where the practical pitfall lies. Most GGUFs distributed by the public are Q4_K_M, and if you use them as is, you won't get an error, but they will run on the CPU without being loaded onto the NPU. It doesn't manifest as "it won't run on the NPU," but rather as "it's running on the CPU without you realizing it."
The quantization name is read as follows:
|
Notation |
meaning |
|
Q4_0 |
Basic quantization centered on 4 bits. The trailing 0 is the system number within the same series. |
|
Q4_K_M |
A 4-bit-centric, medium-level recipe for K-quantization. Mixed quantization using high precision for important tensors. |
|
W4A16 |
Weight (4 bits), Activation (16 bits) |
Q4_K_M does not mean that all tensors are uniformly 4 bits; the effective bits/weight varies depending on the model. In the Llama 3.1 8B example in llama.cpp, it is 4.8944 bits/weight.
|
quantization format |
Handling in the Hexagon backend |
remarks |
|
Q4_0 |
Accepted and repacked for NPU |
The format recommended by the official source. Used in this verification. |
|
Q8_0、MXFP4 |
Accepted and repacked for NPU |
Q8_0 has a larger file size than Q4_0. |
|
F16、F32、IQ4_NL、Q4_1 |
It will be accepted but will not be repacked. |
This was not measured in this verification. |
|
K-quantities such as Q4_K_M and Q5_K_M |
Not accepted |
Executed on the CPU |
Only Q4_0 was measured in this verification. Models that have a Q4_0 version in their distribution can run that file directly via the NPU.
If a Q4_0 version is not available, the standard tools in llama.cpp involve a two-step process: first, creating a GGUF from the original BF16 or FP16 checkpoint, and then quantizing that GGUF to Q4_0.
python convert_hf_to_gguf.py \
--remote <organization>/<model> \
--outfile model-bf16.gguf \
--outtype bf16
./build/bin/llama-quantize \
model-bf16.gguf \
model-Q4_0.gguf \
Q4_0
Double quantization, which involves converting already quantized GGUFs such as Q4_K_M to Q4_0, is not recommended. The official llama.cpp documentation explains that requantizing a quantized model can significantly degrade quality compared to quantizing from 16-bit or 32-bit source data.
QAIRT routes run on W4A16 bundles.
Another QAIRT route retrieves model bundles compiled for the target chipset from Qualcomm AI Hub. LLM primarily uses W4A16, which combines 4-bit weights and 16-bit activations.
# GenieX v0.4.0での実測時モデルID geniex infer qualcomm/Qwen3-4B:W4A16 \ --compute npu \ --think=false
The model namespace varies depending on the version of GenieX. At the time of this article's verification, it is qualcomm/Qwen3-4B, while the official documentation refers to similar AI Hub models in the format ai-hub-models/Qwen3-4B. You can check the actual notation in the model guide for the version you are using and in the output of geniex infer -h.
QAIRT bundles are for NPUs only. Unlike GGUF, which switches to the CPU at runtime, the quantization method, target chipset, context length, etc., are determined when the bundle is created.
While using a pre-configured model makes implementation easy, bundling an arbitrary LLM yourself requires considering model-specific export pipelines, quantization data, KV caches, partitioned graphs, and context binaries. GGUF cannot be directly converted to W4A16; you must revert to the original model checkpoint and corresponding export procedure.
Three pathways were measured using IQ-9075.
In subsequent measurements, the terminology will be used with the following meanings:
|
term |
meaning |
|
prefill |
The stage where input tokens are processed in batches. |
|
decode |
The stage where output is generated one token at a time. |
|
TTFT |
Time from the start of the request until the first output token is received. |
|
E2E |
Time from the start of the request until the completion of receiving the 128 output token. |
|
tok/s |
Number of tokens processed or generated per second |
On the Dragonwing IQ-9075 EVK, the same Qwen3-4B sequence was measured using the following three paths: QAIRT is qualcomm/Qwen3-4B:W4A16, and llama.cpp is unsloth/Qwen3-4B-GGUF:Q4_0.
・llama.cpp CPU
・llama.cpp Q4_0 / Hexagon NPU
・QAIRT W4A16 / Hexagon NPU
The NPU route in llama.cpp is fixed to --compute npu execution. The hybrid route is not included in this measurement.
First, let's compare the speed of prefilling the input. Comparing the CPU and NPU of llama.cpp at the same measurement boundary in llama-bench, the input processing equivalent to 512 tokens was 444.69 tok/s for the NPU and 75.82 tok/s for the CPU, meaning llama.cpp NPU was approximately 5.9 times faster. Since QAIRT cannot be measured in llama-bench, this 5.9 times is a comparison within the llama.cpp path.
Note that this value represents the pure input processing throughput measured by llama-bench and does not match the value calculated inversely from TTFT, which will be discussed later. This is because TTFT includes time other than input processing, and the processing cost per token increases as the input length increases. This 5.9x is the value obtained by running the distributed GGUF as is. Hexagon NPU achieves this difference with prefills without needing to prepare a dedicated bundle.
Next, we look at how much CPU was used for the same process. We generated 128 tokens from a short prompt and converted the process CPU time during the measurement interval to the number of logical cores by dividing it by the actual time.
|
Execution path |
Generation speed |
Average CPU occupancy |
|
QAIRT W4A16 NPU |
13.5 tok/s |
0.83 cores |
|
llama.cpp Q4_0 NPU |
11.3 tok/s |
0.99 cores |
|
llama.cpp CPU |
14.4 tok/s |
6.64 cores |
The generation speed of short responses was similar across all three paths, with differences ranging from a few percent to 20%. The biggest difference was in CPU utilization. The QAIRT NPU path maintained approximately 94% of the generation speed of the CPU path while reducing the average CPU utilization from 6.64 cores to 0.83 cores. This means that approximately 5.8 out of 8 cores are freed up for other processing.
In other words, even comparing only the generation speed of short responses, the difference between the three paths is small, and this metric alone does not make it clear why NPUs are useful. NPUs are effective in edge LLMs for processing long inputs, which we will discuss next, and for the aforementioned CPU utilization.
The longer the input, the more effective the NPU becomes.
Next, we compared three paths using inputs of 256 to 3,840 tokens across the same /v1/completions API boundary in GenieX. The conditions were temperature 0, top-k 1, seed 42, and 128 output tokens. We ran each of the six input lengths five times, alternating between ascending and descending order, with a 45-second cooldown between trials.
Because switching routes does not fully restore the memory allocated by the previous route, each route is measured with only one inference server running from the state immediately after the board restart. The first request, including the initial model load, is excluded from the aggregation, and all remaining 90 requests were completed successfully.
The median TTFT (Time to Receive the First Token) and E2E (Time to Receive 128 Tokens) for 3,840 token inputs were as follows:
|
Execution path |
TTFT median |
Comparison of TTFT with CPU |
E2E median |
|
QAIRT W4A16 NPU |
4.511 seconds |
17.6x faster |
15.674 seconds |
|
llama.cpp Q4_0 NPU |
11.230 seconds |
7.1x faster |
32.797 seconds |
|
llama.cpp CPU |
79.499 seconds |
standard |
91.779 seconds |
The difference persists not only in TTFT but also in E2E (end-to-end) response time, which is the time it takes for all responses to be received. With 3,840 tokens input, the CPU took 91.779 seconds, compared to 15.674 seconds for the QAIRT NPU.
This difference is evident even with short inputs. For a 256-token input, the QAIRT NPU had the shortest E2E time at 8.195 seconds, while the CPU and llama.cpp NPU were nearly identical at 11.488 seconds and 11.243 seconds, respectively. The difference widens further as the input length increases. Evaluating only short prompts may underestimate the value of the NPU's prefill performance.
While longer inputs did decrease the decoding speed itself, the degree of decrease varied depending on the path. When the input was extended from 256 tokens to 3,840 tokens, the smallest decrease was observed on the QAIRT NPU (16.3 → 11.5 tok/s, approximately -29%), followed by the CPU (16.0 → 10.4 tok/s, approximately -35%), and the llama.cpp NPU (12.1 → 5.9 tok/s, approximately -51%).
Note that the decoding ranking changes depending on the measurement boundary. In the short sentence condition in the previous section, the CPU achieved 14.4 tok/s and the QAIRT NPU achieved 13.5 tok/s, but in the 256 token condition of /v1/completions in this section, the QAIRT NPU achieved 16.3 tok/s and the CPU achieved 16.0 tok/s, so the rankings are reversed. When comparing decoding speeds, it is necessary to standardize which API boundary was used for measurement.
The value of NPUs lies in long text input and CPU resource management.
The LLM process is divided into two main parts: prefill, which reads the input, and decode, which generates tokens one by one.
Prefill is a process that easily leverages the parallel processing capabilities of the NPU because it processes large matrix operations in batches. On the other hand, decoding involves repeatedly performing small calculations and continuously reading large model weights, so it is strongly affected by memory bandwidth.
Therefore, for applications involving inputting long prompts, search results, and conversation history, speeding up prefill directly translates to a reduction in response time. On the other hand, decoding is heavily influenced by memory bandwidth, and using an NPU does not necessarily guarantee faster decoding.
Furthermore, the difference in CPU utilization becomes significant when running other processes concurrently on the same device. In the measurements in the previous section, the CPU path consistently utilized an average of 6.64 out of 8 cores, while the QAIRT NPU path utilized only 0.83 cores.
In a configuration where most of the 8 cores are dedicated to LLM inference, there is less capacity left for the following processes:
- Processing of camera and sensor inputs
- Speech recognition and speech synthesis
• Database search
- Execute agent tools
- Responses from Web APIs and UIs
This verification measured the CPU usage of the LLM inference process, and did not measure the performance when these processes are combined.
The NPU is not simply a computing unit that speeds up LLMs; it also plays a role in freeing up the CPU from LLM inference, creating processing headroom for the entire device.
Differences in personality revealed in the two paths
Characteristics of the GGUF + llama.cpp route
- You can import publicly available GGUF files as they are (if a Q4_0 version exists).
- The execution target can be switched between CPU, GPU, and NPU using the same model file.
- Context length and other parameters can be specified at runtime.
- Can be tested even on models that do not have a QAIRT bundle for the target SoC.
In this verification, we checked the presence or absence of the Q4_0 version and whether loading and inference were successful using a fixed NPU path by examining the actual system logs. Whether loading and inference are successful depends on the model and configuration.
Characteristics of the QAIRT pathway
Models with bundles for the target SoC available in Qualcomm AI Hub can be used as is.
- Models without bundles can be exported from the original checkpoint and prepared by the user.
- Runs in a pre-compiled state for the chipset.
In this test, the TTFT was the shortest and CPU usage was the lowest.
- The quantization method, target chipset, and context length are fixed when the bundle is created.
Under the conditions of this verification, the QAIRT path showed the best values in both TTFT and CPU occupancy for models for which bundles are provided for the target SoC. However, quality is outside the scope of this verification.
Because Q4_0 and W4A16 use different quantization methods, the speed ranking does not directly correspond to the output quality ranking. Even models for which a bundle is not provided can be placed on the QAIRT path by exporting from the original checkpoint. On the other hand, when comparing configurations while changing them at runtime, the GGUF path requires fewer steps. Which path is more suitable depends on the target SoC, model, and requirements.
Summary
Where a GGUF is executed is determined by the runtime backend, not its format. Using the GenieX llama.cpp runtime, you can run the corresponding GGUF on a Hexagon NPU. GenieX has two NPU paths:
• GGUF + llama.cpp Hexagon: A versatile path that makes it easy to import a wide range of models.
Qualcomm AI Hub + QAIRT: A pre-compiled NPU path for the target chipset.
In actual measurements with the IQ-9075, the llama.cpp NPU prefilled 512 tokens approximately 5.9 times faster than the same llama.cpp CPU path. In a TTFT with 3,840 tokens input, the QAIRT NPU was 17.6 times faster than the llama.cpp CPU, while the llama.cpp NPU was 7.1 times faster. On the other hand, for short decoding, the CPU was sometimes faster. Looking only at generation speed, the difference between paths was small, but the ranking changed when input length, response start time, and CPU usage were also included.
The above are the results of actual measurements under the conditions of IQ-9075 and Qwen3-4B. The values and rankings between paths will change if the target SoC, model, and bundle availability change.
Inquiry
If you are interested in building on-device LLMs using Qualcomm Dragonwing SoCs, or implementing models using GenieX and Qualcomm AI Hub, please feel free to contact us.
