abn. · Model Evaluation· Backend benchmark · lemonade bench · gemma4-it-e4b

FastFlowLM on AMD NPU — XRT vs HRX backend

The same FastFlowLM 1.0.4 build, packaged twice: one subpackage on the Xilinx Runtime (XRT), one on Hip Runtime Extended (HRX). Same model, same machine, same lemonade bench scenarios, one run per backend in each of three box states: with another workload resident, idle on the balanced power profile, and idle on the performance profile.

Evaluated 2026-09-03Updated 2026-09-04
Verdict · Decode: HRX 4–7% slower in every state · first token: state-dependent, HRX 4–6% later on the idle performance box · memory equal within 2 GB

In all three states HRX decodes slower than XRT in every scenario with no overlap between three-run ranges: −4.1% to −5.4% on the idle box with the performance profile, −5.3% to −7.2% on the balanced profile, −4.9% to −7.0% with a workload resident. Time to first token is where the box state matters more than the runtime: with a workload resident the two tie (−1.0% to +1.8%); on the idle box HRX is later by +4.1% to +10.6% (balanced) and +4.3% to +6.3% (performance). Both backends reach the first token 10–20% later on the idle balanced box than with the workload resident, and the performance profile recovers 4–9% of that for both. Peak memory as reported by lemonade is 16.7–17.2 GB (XRT) against 17.3–18.4 GB (HRX) on the performance box; the 30.0–30.5 GB HRX figure from the first state was the resident workload and does not recur. Zero failed runs in 90.

Model card

Base model
gemma4-it:e4b — lemonade model gemma4-it-e4b-FLM (recipe flm, backend npu)
Native context
131,072 (ctx_size as run)
Provenance
Checkpoint as served by FastFlowLM through lemonade; identical across all six runs

Configuration

exact serving config — reproducible
Command
lemonade bench gemma4-it-e4b-FLM --json
lemonade
11.9.0
Recipe / backend
flm / npu (reported by lemonade as flm/npu v1.0.3)
Package under test
fastflowlm-1.0.4-2.git.1.bf7fe0e.fc44: one SRPM, two backend subpackages
XRT run
fastflowlm-xrt on xrt-npu 2.26.95-2.fc44 + xrt-plugin-amdxdna 2.26.95-2.fc44 (xrt-smi reports XRT 2.26.0)
HRX run
fastflowlm-hrx on hrx 0.3.0-6.fc44: hrx-system built from the amdxdna-hal-native-rel branch at 6867cba (v0.3.0-6360), not a release; FastFlowLM’s own pin is the flm-hrx-amdxdna-v2026.07.30 release on the jtuyls/hrx fork, built at eb0b39f
Backend selection
both subpackages installed side by side; /usr/bin/flm is an alternatives link (flm-hrx priority 20, flm-xrt priority 10); lemonade runs the binary at ~/.cache/lemonade/bin/flm/npu/flm, symlinked to the alternatives-managed /usr/bin/flm, and the backend was switched with alternatives between runs (owner-stated)
Context
131,072
Runs
measurement_runs 3 · warmup_runs 0 · memory_tracking on · model reloaded before every run (15 loads per run, per the logs)
State 1, 3 Sep
another workload resident on the box (owner-stated). XRT 19:17:38 UTC, HRX 20:26:56 UTC
State 2, 4 Sep
idle box, balanced power profile (owner-stated). HRX 17:50:59 UTC, XRT 18:03:33 UTC
State 3, 4 Sep
idle box, performance power profile (owner-stated). XRT 18:11:25 UTC, HRX 18:17:55 UTC. All times converted from the JSON’s local-time stamps, see caveats
Scenarios
chat-short · chat-long-output · code-short · code-explain · code-debug; 13 imagegen scenarios skipped by lemonade bench on every run; 4 long-context excluded by the default; 4 embedding apply to embedding models only

Hardware / environment

public transparency block
GPU
AMD Ryzen AI MAX+ PRO 395 (Strix Halo) w/ Radeon 8060S, 128 GB unified memory
Arch
XDNA 2 NPU, xrt-smi: RyzenAI-npu5, aie2p, 6x8, power mode default · gfx1151 iGPU
Cards
1 APU · 8 × 16 GiB DDR5 at 8000 MT/s = 128 GiB installed
Topology
Unified memory: 32 GiB BIOS carve-out (mem_info_vram_total) + 56 GiB GTT pool (ttm.pages_limit=14680064); the kernel sees 94 GB, which is what lemonade reports as ram_gb
ROCm
hrx 0.3.0-6.fc44
OS
Fedora 44 Workstation
Kernel
7.1.12-200.fc44.x86_64 · in-tree amdxdna module · NPU firmware 1.1.2.65
Engine
FastFlowLM via lemonade 11.9.0
Build
1.0.4-2.git.1.bf7fe0e.fc44 (x86_64)

Evaluation results

measured · state 3 (idle box, performance profile) is the comparison of record; states 1 and 2 are kept beside it

Headline, state 3: idle box, performance profile

Averages of the five scenario means are derived on this page; every other tile is read from the JSON.

11.87tok/s
Decode, XRT
mean of 5 scenario means · 11.60–12.07
11.33tok/s
Decode, HRX
mean of 5 scenario means · 11.07–11.47
−4.06% to −5.36%
Decode Δ (HRX vs XRT)
slower on HRX in all 5 scenarios; −5.3% to −7.2% in state 2, −4.9% to −7.0% in state 1
+4.32% to +6.32%
TTFT Δ (HRX vs XRT)
later on HRX in all 5; ranges overlap in 2 of 5. State 2: +4.1% to +10.6%. State 1: −1.0% to +1.8%
16.7–17.2 vs 17.3–18.4GB
Peak memory
XRT vs HRX · memory_peak_gb · state 1 HRX read 30.0–30.5
0 / 90
Failed runs
15 runs per backend per state

Decode throughput by scenario, state 3, mean tok/s

Three-run means from the performance-profile JSON, rounded to two decimals. The y-axis starts at zero. The state ladder below carries all three states; the range table carries the min–max spread.

03.787.5611.3415.1212.0711.42chat-short11.8911.4chat-long-output12.0311.47code-short11.7511.27code-explain11.611.07code-debug
XRTHRX

Decode throughput, state 3: mean and three-run range

tok/s as reported by lemonade (tps), 4 September, idle box, performance profile. Δ is derived: (HRX − XRT) / XRT on the unrounded means. In every scenario the HRX maximum is below the XRT minimum.

ScenarioOutput tokensXRT meanXRT min–maxHRX meanHRX min–maxΔ
chat-short2012.0712.01–12.1711.4211.28–11.62−5.36%
chat-long-output25611.8911.88–11.9011.411.31–11.57−4.08%
code-short6012.0311.98–12.1011.4711.47–11.48−4.60%
code-explain12811.7511.75–11.7511.2711.27–11.28−4.06%
code-debug10011.611.58–11.6211.0711.06–11.08−4.55%

Decode throughput, the state ladder: means and Δ per state

One XRT and one HRX run per state; Δ within the state, derived on unrounded means. State 1: 3 Sep, workload resident. State 2: 4 Sep, idle, balanced profile. State 3: 4 Sep, idle, performance profile. The three-run ranges never overlap between backends in any state or scenario.

ScenarioS1 XRTS1 HRXS1 ΔS2 XRTS2 HRXS2 ΔS3 XRTS3 HRXS3 Δ
chat-short12.2711.41−7.02%12.0511.18−7.15%12.0711.42−5.36%
chat-long-output12.0811.41−5.58%11.911.27−5.30%11.8911.4−4.08%
code-short12.1911.57−5.04%1211.29−5.92%12.0311.47−4.60%
code-explain11.9411.36−4.85%11.7611.06−5.95%11.7511.27−4.06%
code-debug11.7711.14−5.39%11.6110.95−5.69%11.611.07−4.55%

Time to first token by scenario, state 3, mean ms

Three-run means from the performance-profile JSON. Both backends sit at 1,570–1,712 ms for 28–208 input tokens and at 1,900–1,982 ms for the 363-token code-debug prompt.

06211242186324841585.81658.4chat-short15701645.4chat-long-output1589.21665.8code-short1609.81711.5code-explain1900.41982.4code-debug
XRTHRX

Time to first token, state 3: mean and three-run range

ms as reported by lemonade (ttft_ms), rounded to one decimal. Δ is derived on the unrounded means. The ranges overlap in chat-short and chat-long-output, where HRX’s three runs spread 6.9% and 9.2%; in the three coding scenarios the HRX minimum is above the XRT maximum.

ScenarioInput tokensXRT meanXRT min–maxHRX meanHRX min–maxΔ
chat-short281585.81567.6–1602.91658.41592.1–1705.9+4.57%
chat-long-output4415701558.1–1580.61645.41570.2–1721.6+4.80%
code-short351589.21581.6–1593.71665.81657.2–1672.0+4.82%
code-explain2081609.81604.4–1618.31711.51693.7–1735.8+6.32%
code-debug3631900.41892.0–1912.21982.41957.9–2026.0+4.32%

Time to first token, the state ladder: means and Δ per state

ms, one run per backend per state, Δ within the state on unrounded means. In state 1 the ranges overlap in four scenarios (code-short has the whole HRX range below XRT); in state 2 they overlap in none; in state 3 they overlap in the two chat scenarios.

ScenarioS1 XRTS1 HRXS1 ΔS2 XRTS2 HRXS2 ΔS3 XRTS3 HRXS3 Δ
chat-short1496.51496.5−0.00%1650.41725.8+4.57%1585.81658.4+4.57%
chat-long-output1480.71507.1+1.78%1638.21812.4+10.63%15701645.4+4.80%
code-short1488.51473.8−0.99%1677.51756.7+4.72%1589.21665.8+4.82%
code-explain1508.41514.6+0.41%1676.21775.7+5.93%1609.81711.5+6.32%
code-debug18051817.7+0.70%1993.92075.4+4.09%1900.41982.4+4.32%

Between the states

Every number here is a within-backend delta between states, derived from the unrounded means. Decode barely moves: idle balanced against workload resident, XRT −1.4% to −1.8% and HRX −1.2% to −2.6%; performance against balanced, XRT −0.1% to +0.2% and HRX +1.1% to +2.1%. Time to first token moves a lot: idle balanced against workload resident, XRT +10.3% to +12.7% and HRX +14.2% to +20.3%, with no range overlap for either backend in any scenario; performance against balanced, XRT −3.9% to −5.3% and HRX −3.6% to −9.2%. So the power profile accounts for roughly 4–5% of the shift on XRT and 4–9% on HRX, and both backends still reach the first token 5–10% later on the idle performance box than they did with the other workload resident. What that workload was doing to clocks or memory residency is not recorded; these outputs only show that it made both backends faster to the first token, HRX by more.

Wall-clock duration per run, state 3, mean

ms as reported by lemonade (duration_ms), rounded to whole milliseconds. Δ columns are derived. Summed over the five scenarios: XRT 56,443 ms, HRX 59,274 ms (+5.02%). State 2: 56,957 against 60,736 ms (+6.64%). State 1: 55,322 against 58,413 ms (+5.59%).

ScenarioXRT meanHRX meanΔ msΔ
chat-short32873466+179+5.47%
chat-long-output2328824407+1,119+4.81%
code-short66466997+351+5.28%
code-explain1261013244+634+5.03%
code-debug1061311159+546+5.15%

Peak memory, state 3, highest value across the five scenarios

lemonade’s memory_peak_gb with memory_tracking on. The bars are scaled to the larger value, not to any capacity figure. State 1’s HRX value, 30.5 GB, is in the table below and not on this chart: the owner states another workload was resident during that run, and neither idle-box run reproduces it.

XRT17.2GB16.7–17.2 across scenarios
HRX18.4GB17.3–18.4 across scenarios

Memory by scenario, all six runs

memory_peak_gb / vram_peak_gb as reported, VRAM rounded to two decimals. Both state-1 runs carry 0.7–1.2 GB more VRAM than the idle-box runs, on both backends, so that difference travels with the box state, not the runtime. On the idle box HRX reports 0.15 GB less VRAM than XRT (balanced) or the same within 0.05 GB (performance).

ScenarioS1 XRTS1 HRXS2 XRTS2 HRXS3 XRTS3 HRX
chat-short17.2 / 1.9630.0 / 2.2716.9 / 1.2216.7 / 1.0817.2 / 1.2717.5 / 1.22
chat-long-output17.0 / 1.9530.5 / 2.2916.5 / 1.2017.6 / 1.0516.7 / 1.2718.4 / 1.22
code-short17.1 / 1.9130.0 / 2.3316.7 / 1.1917.0 / 1.0416.8 / 1.2117.3 / 1.22
code-explain17.1 / 1.8930.3 / 2.3116.8 / 1.1917.2 / 1.0416.8 / 1.2117.5 / 1.22
code-debug17.1 / 1.9530.1 / 2.2716.8 / 1.1917.3 / 1.0416.8 / 1.2117.8 / 1.22

Run-to-run spread, state 3

Spread is (max − min) / mean of the three runs, derived from the JSON. Decode: XRT 0.05–1.29% per scenario, HRX 0.05–2.96%. TTFT: XRT 0.76–2.23%, HRX 0.89–9.20%. The decode gap between backends (4.06–5.36%) exceeds every within-backend spread. The TTFT gap (4.32–6.32%) exceeds the XRT spread everywhere, but the HRX spread in chat-short (6.86%) and chat-long-output (9.20%) exceeds the gap in those two scenarios, which is why their ranges overlap. The widest single spread across all six runs is HRX chat-long-output TTFT in state 3, 1,570.2 to 1,721.6 ms.

Per-run values, all six runs

TTFT (ms) and TPS per run as printed by lemonade bench, in run order. S1: 3 Sep, workload resident. S2: 4 Sep, idle, balanced. S3: 4 Sep, idle, performance. The log prints TPS to one decimal; the JSON carries the unrounded values used in the tables above.

ScenarioRunRun 1 TTFTRun 2 TTFTRun 3 TTFTRun 1 TPSRun 2 TPSRun 3 TPS
chat-shortS1 XRT1488.51505149612.212.312.3
chat-shortS1 HRX1504.61476.51508.311.411.511.3
chat-shortS2 XRT1636.61631.11683.6121212.1
chat-shortS2 HRX1717.41730.41729.711.211.211.2
chat-shortS3 XRT1586.91602.91567.6121212.2
chat-shortS3 HRX1677.11705.91592.111.411.311.6
chat-long-outputS1 XRT1502.41461.61478.212.112.112.1
chat-long-outputS1 HRX14931528.61499.611.411.411.5
chat-long-outputS2 XRT16271644.91642.611.911.911.9
chat-long-outputS2 HRX1844.51737.21855.411.411.211.2
chat-long-outputS3 XRT1571.41558.11580.611.911.911.9
chat-long-outputS3 HRX1570.21644.61721.611.611.311.3
code-shortS1 XRT1489.51488.1148812.212.212.2
code-shortS1 HRX1486.21466.81468.411.511.611.6
code-shortS2 XRT1694.91671.51666.1121212
code-shortS2 HRX1775.51723.4177111.311.311.3
code-shortS3 XRT1592.31581.61593.71212.112
code-shortS3 HRX1668.21657.2167211.511.511.5
code-explainS1 XRT1509.31515.9150011.911.912
code-explainS1 HRX1535.31518.41490.211.311.411.4
code-explainS2 XRT1653.11684.81690.811.811.811.8
code-explainS2 HRX1730.518341762.511.111.111.1
code-explainS3 XRT1618.31604.41606.611.711.811.8
code-explainS3 HRX1693.71705.11735.811.311.311.3
code-debugS1 XRT1806.81810.81797.511.811.711.8
code-debugS1 HRX1777.91793.81881.311.21111.2
code-debugS2 XRT1992.32003.31986.211.611.611.6
code-debugS2 HRX2042.12091.42092.810.91110.9
code-debugS3 XRT18921896.91912.211.611.611.6
code-debugS3 HRX20261963.31957.911.111.111.1

Reproducing this run

The packages are published in the abn/amd-npu COPR for Fedora: fastflowlm, fastflowlm-xrt, fastflowlm-hrx, xrt-npu, xrt-plugin-amdxdna, hrx and lemonade-server. lemonade does not call the system flm; it runs its own copy at ~/.cache/lemonade/bin/flm/npu/flm. Point that path at the packaged binary once with ln -sf $(which flm) ~/.cache/lemonade/bin/flm/npu/flm, then alternatives --set flm /usr/bin/flm-xrt or /usr/bin/flm-hrx selects the runtime for every subsequent lemonade run. lemonade bench gemma4-it-e4b-FLM --json then produces the same JSON and log this page is built from. Run both backends in the same box state, idle and on one power profile: the state ladder above shows what a resident workload does to the memory column and what the profile does to the first token. Kernel 7.1.12 with the in-tree amdxdna module was used here; other kernels were not tested.

What these runs do not measure

One model, one context size, one machine, three runs per scenario per state with no warm-up and a model reload before each run. lemonade bench skipped its 13 imagegen scenarios, reporting the model as not suitable for them; the four long-context scenarios are excluded by its default and the four embedding scenarios apply to embedding models only. No power, energy or thermal data was collected, and the power profile is a system setting whose effect on NPU clocks was not read back. The XRT and HRX runtime versions are not recorded in the output files; they were read from the host with rpm and xrt-smi. Output quality was not compared between backends. Nothing here explains why both backends reach the first token sooner with another workload resident than on an idle box under either profile.

Provenance & caveats

empirical — tagged, n-counted, honest

Provenance tags

measuredread from the lemonade bench JSON output (means, min/max, memory)
logread from the printed run log (per-run TTFT / TPS)
derivedcomputed on this page from measured values (deltas, ratios, spreads, averages of means)
packagefrom the RPM metadata of the packages under test
hostread from the benchmark host after the runs (rpm, xrt-smi, sysfs, dmidecode, /proc/cmdline); not present in the benchmark output
owner-statedstated by the machine’s owner and not verifiable from the outputs: which runs had another workload resident, the power profile per state, and that the backend switch was done with alternatives

Caveats

  • State 1 (3 September) ran with another workload resident on the box, owner-stated. Its HRX memory_peak_gb of 30.0–30.5 GB is that workload, not HRX: both idle-box HRX runs report 16.7–18.4 GB, within 2 GB of XRT. State 1 is kept on this page because its decode gap matches the idle box and its first-token tie does not.
  • The states are not interchangeable. Both backends decode 1.2–2.6% faster and reach the first token 10–20% sooner in state 1 than in state 2, and the performance profile (state 3) brings the first token 4–9% back down without moving decode. Compare within a state only; the page never sets a run from one state against a run from another as if they were one pair.
  • Run order within a state was not controlled: state 1 XRT then HRX 69 minutes later; state 2 HRX then XRT 12.5 minutes later; state 3 XRT then HRX 6.5 minutes later.
  • lemonade reports the backend as flm/npu v1.0.3 in every JSON, while the packages under test are FastFlowLM 1.0.4-2.git.1.bf7fe0e. Both values are kept as found. The v1.0.3 is the version lemonade tracks for the FLM tree it manages: the flm.npu pin in its resources/backend_versions.json and the version.txt beside its wrapper, both v1.0.3. It is not a reading from the packaged binary, which reports FLM v1.0.4 with --version.
  • lemonade stamps local time with a Z suffix. The JSON stamps are 2026-09-03T21:17:38Z, 22:26:56Z, 2026-09-04T19:50:59Z, 20:03:33Z, 20:11:25Z and 20:17:55Z; the host clock was CEST (UTC+2), so the page shows 19:17:38, 20:26:56, 17:50:59, 18:03:33, 18:11:25 and 18:17:55 UTC. The state-1 HRX JSON also carries two stamps one second apart (22:26:55Z at the top level, 22:26:56Z on the model entry).
  • Percent deltas, ratios and spreads are computed on this page from the unrounded JSON means; the tables show rounded values, so recomputing from the rounded numbers can differ in the last digit.
  • n = 3 per scenario per run with warmup_runs 0. The logs show the model being reloaded before every run; whether the reload affects the recorded TTFT is not determined from these outputs.
  • memory_peak_gb and vram_peak_gb are lemonade’s fields with memory_tracking enabled. The lemond binary’s strings show the tracker reading MemTotal and MemAvailable from /proc/meminfo and mem_info_vram_used / mem_info_gtt_used from the amdgpu sysfs, so memory_peak_gb is most likely system-wide used memory, which is exactly why a resident workload lands in it, and vram_peak_gb the iGPU’s usage, carve-out and possibly GTT. The tracker itself is in the lemonade CLI, polling lemond’s system-stats endpoint. lemonade reads kilobytes; its ram_gb: 94.05 matches the host’s 96,308 MiB only as GiB, so every memory figure here is GiB with a GB label.
  • The JSON’s vram_gb: 32.0 and ram_gb: 94.05 are what lemonade read at run time on a 128 GiB unified-memory APU. The 32 GiB is the BIOS UMA carve-out (mem_info_vram_total), hidden from the kernel; on top of it the amdgpu driver can map up to 56 GiB of system RAM through GTT (mem_info_gtt_total, ttm.pages_limit=14680064). So 32 GB is neither a device capacity nor a ceiling on GPU-addressable memory, and the two reported numbers do not sum to the installed 128 GiB.
  • The two backend subpackages come from one SRPM and one build, so the FastFlowLM source is identical; the runtime beneath it (XRT 2.26.95 with the amdxdna plugin, or HRX 0.3.0-6) is the only intended variable. The owner states the runs were switched through alternatives with lemonade’s ~/.cache/lemonade/bin/flm/npu/flm symlinked to /usr/bin/flm; the outputs themselves do not record the active backend, so the attribution rests on that statement, the file names, and the decode gap holding its size and sign across all three states.
5 scenarios × 3 runs × 2 backends × 3 states = 90 measured runs · 0 failed · 13 imagegen scenarios skipped per run