5080 16GB vs 3090TI 24GB Generative AI benchmarking!

tl;dr;

5080 16GB is about 20% faster prompt processing and 10-15% faster token generation than 3090TI FE 24GB for the case where both LLM weights and 8k context can fit entirely in 16 GB VRAM.

5080 16 GB is about 10% faster generating images with stable diffusion than 3090TI FE 24GB.

However, the extra 8GB VRAM on the 3090TI FE 24GB offers flexibility to run larger models, longer context, and larger batch sizes as well as general quality of life for driving a desktop rig running xwindows and a browser etc.

Motivation

I’ve been enjoying the recent GPU review content of the 5060TI 16GB and today’s video on the ASUS 5080 16GB and wondering how these new 16GB cards compare to older cards for common home lab ai workloads like LLM inferencing and Stable Diffusion image generation?

Rambling Background

I don’t have a lot of data points yet beyond Wendell’s Procyon AI Text Generation Benchmark that uses ONNXRUNTIME. This is still useful comparison, but in my understanding, ONNX runtime packages a model in such a way that it can run on a variety of systems e.g. Linux, Windows, CPUs, GPUs, etc. However, that can come at a cost to performance over a more optimized runtime built with optimizations for the target hardware.

Just recently, with Wendell’s hardware support, I was able to release an experimental quantization of Google’s gemma-3-27b-it-qat-GGUF LLM model on huggingface designed to work with ik_llama.cpp fork. These quants are perfect for playing around with the best quality gemma-3 in under 16GB VRAM (CUDA GPUs anyway :sweat_smile:).

Unfortunately, these models will NOT run on mainline llama.cpp, ollama, lm studio, koboldcpp, etc. (If you want to run on those engines, check out bartowski/google_gemma-3-27b-it-qat-GGUF. I’ve actually been in touch with bartowski, he’s doing some great stuff to support ai home lab enthusiasts!)

LLM Benchmark

Interestingly, redditor u/Maxious just posted benchmarks running the ubergarm/gemma-3-27b-it-qat-GGUF/gemma-3-27b-it-qat-mix-iq3_k.gguf with 8k context on a Inno3D RTX 5080 X3 OC 16GB! So I ran the exact same command on my 3090TI 24GB and here are the results.

Higher numbers are better for both graphs.

The graph on the left is prompt processing (sometimes called ā€œprefillā€) which is sort of how fast the LLM can read what you give it. You can see that the more input you give the model, the slower it is given there are more computations.

The graph on the right is token generation or how fast the LLM can write in reply. Similarly, the more text that it had to process the slower it is given each new word it generates depends on every word that came before it so it slows down for longer generations.

Just for comparison, I ran a couple others configurations on my 3090TI 24GB to show how if you have more than 16GB VRAM you could not compress the kv-cache and leave it at f16 which is slightly faster than q4_0 on GPU inferencing.

I also ran in a configuration that only takes 12GB VRAM by offloading just the attention and kv cache onto CPU RAM and let my 9950X with 96GB DDR5-6400 and overclocked infinity fabric handle those calculations. This allows a person with low VRAM to use a much larger context but at the penalty of slower performance.

Finally, my 3090TI 24GB could fit up to 32k context with this model which would allow processing of longer text, longer code generation etc. So keep that in mind when comparing.

Personally, I’d probably still choose a used 3090TI 24GB over a 5080 16GB if local LLM inferencing is an important part of your daily activities. While it may be about 10% slower for jobs that fit into 8k context, keeping everything in up to 32k context will out perform the 5080 16GB given you would have to offload to CPU RAM… Also having a little extra VRAM to run xwindows display and firefox gives quality of life :pinched_fingers:

Image Diffusion

The other big home lab use case for GPUs is running ComfyUI for Stable Diffusion image generation. I have run SD1.5, SDXL 1.0, PonyXL, FLUX.1-dev, Illustrious, the newest HiDream, and even Wan2.1-I2V-14B to generate short video clips from a starting image. That whole scene is really impressive with a lot of optimizations allowing complete workflows to fit in under 16GB VRAM. You can find most of those models and workflows on civit.ai and huggingface as well.

I don’t have any direct comparisons myself, but just saw a new open bechmark project pop up today as posted on reddit by u/yachty66.

It runs stable diffusion generations for 5 minutes and measures number of images generated, average temperature, and eventually power. Seems to take fixed 4GB VRAM and kept the utilization pegged near 100% and power maxed out just below 450 watt limit.

==================================================
BENCHMARK RESULTS:
Images Generated: 141
Max GPU Temperature: 79°C
Avg GPU Temperature: 73.9°C
==================================================

Okie, if anyone is interested holler and I can give exact commands if you want to try any of this on your rig to compare! I would love to see how the 5060TI 16GB results stack up!

5 Likes

Run LLM Benchmark

Steps to build and run ik_llama.cpp fork with above gemma-3 LLM quant.

# Install build dependencies and cuda toolkit as needed
git clone https://github.com/ikawrakow/ik_llama.cpp
cd ik_llama.cpp
# Configure CUDA+CPU Backend
cmake -B ./build -DGGML_CUDA=ON -DGGML_BLAS=OFF
# Build
cmake --build ./build --config Release -j $(nproc)
# Confirm
./build/bin/llama-server --version
version: 3640 (93cd77b6)
built with cc (GCC) 14.2.1 20250128 for x86_64-pc-linux-gnu
# Download 13.7GB model
wget https://huggingface.co/ubergarm/gemma-3-27b-it-qat-GGUF/resolve/main/gemma-3-27b-it-qat-mix-iq3_k.gguf
# Benchmark
CUDA_VISIBLE_DEVICES=0 \
./build/bin/llama-sweep-bench \
    --model gemma-3-27b-it-qat-mix-iq3_k.gguf \
    -ctk q4_0 -ctv q4_0 \
    -fa \
    -amb 512 \
    -fmoe -c 8192 \
    -ub 512 \
    -ngl 99 \
    --threads 4

Run Stable Diffusion Benchmark

This downloads about ~6GB into ls ~/.cache/huggingface/ and requires you to login with your huggingface settings → create new token

# 1. Create huggingface account and token with above link
mkdir gpu-benchmark
cd gpu-benchmark/
# install uv (better python pip) locally into ~/.cache/uv
# https://docs.astral.sh/uv/getting-started/installation/
curl -LsSf https://astral.sh/uv/install.sh | sh
uv venv ./venv --python 3.12 --python-preference=only-managed
source venv/bin/activate
uv pip install gpu-benchmark
uv pip install huggingface_hub[cli]
huggingface-cli login
# benchmark takes 5 minutes and 4GB VRAM
gpu-benchmark
# cleanup everything when complete
cd ..
rm -rf gpu-benchmark
huggingface-cli delete-cache
# select sd-legacy/stable-diffusion-v1-5 5.5GB, hit space, enter confirm Yes
# delete huggingface token if desired
rm ~/.cache/huggingface/*tok*
1 Like

Thanks for those benchmarks!

I guess the 3090 is still pretty competitive, specially considering that you can get it used for reasonable prices (depending on your region).

I think I’ll be keeping the couple ones I have for a long time still, or at least until I can get a pair of reasonably-priced 32GB GPUs, which I doubt will happen anytime soon haha

1 Like

Great benchmarks @ubergarm, thanks for sharing the methodology too — reproducibility matters.

The ~20% prompt processing advantage for the 5080 makes sense given the higher memory bandwidth per GB (960 GB/s over 16 GB vs 912 GB/s over 24 GB on the 3090 Ti). For batch-1 token generation, bandwidth-per-capacity is what drives tok/s once the model fits in VRAM.

But as you noted, the 3090 Ti’s extra 8 GB is the real differentiator for practical use. At IQ3 quants you can squeeze ~30B models comfortably on 24 GB with full KV cache for 8K context. On 16 GB you’re either dropping to smaller models or aggressively quantizing the KV cache (q4_0 like in your bench), which works but starts to matter at longer contexts.

@igormp same here — the 3090/3090 Ti is hard to beat on price-to-VRAM ratio. Used prices have stabilized around $700-800 for 24 GB, while the 5080 is $1000+ for 16 GB. Unless you specifically need the power efficiency or newer CUDA features (like FP4 on Blackwell… whenever that actually lands in consumer cards), the 3090 Ti remains the sweet spot for local inference.

One thing worth testing: try running with --flash-attn and different KV cache quantization levels (q8_0 vs q4_0). In my experience the quality delta between q8_0 and q4_0 KV cache is negligible for short contexts but becomes noticeable past 4K tokens, especially for coherence in longer conversations.

2 Likes

Thanks, and thanks for pointing out the ā€œbandwidth-per-capacityā€ measure I hadn’t thought of that. In general PP is compute bottlenecked and TG is memory bandwidth bottlenecked so I assumed the 5080 had more compute, but honestly I didn’t dig toooo deep.

ik_llama.cpp has some interesting options too e.g. -khad -ctk q6_0 -ctv q8_0 type stuff. The traditional wisdom is the keys can be quantized more than the values so for someone minmaxing their rig it could be useful. I don’t think mainline offers the q6_0 format. You can even do -ctk q4_K etc but generally the non-legacy quant types take too much CPU for on the fly quantization like this is doing for kv-cache.

The syntax has changed and in the past -fa meant ā€œuse flash attentionā€. Now I believe it is -fa on …

Though yeah 3090TI dialed in with LACT so it isn’t thermal throttling and oscillating the GPU clock is still a great value proposition. Especially with ik_llama.cpp’s -sm graph ā€œtensor parallelā€ allowing you to span two 3090s and get full use out of them approach vLLM speeds for single batch inference. I’ve heard rumors that mainline llama.cpp is working on tensor parallel and NUMA optimizations.

Its a good time for local enthusiasts (except for the horrible prices lmao… :joy: :sob: ))

Yep - it’s this PR here:

It’s not currently working for Vulkan, but CUDA-capable cards should allow you to play with it. I must admit, having moved on from 2 x 3090 to 2 x R9700, I’ve found myself obsessively checking that PR to see when the guys pick it up again :smiley:

I concluded much the same as you, that past a certain performance level, extra capacity is more important than eval rate. At least for my uses, which is to say single-user agentic coding. The main problem I found with the 3090s was cooling; there aren’t many consumer boards that give sufficient physical separation between them, and even imposing 250W limits left them throttling.

1 Like

Hi Ubergarm. I’ve been using your Q4_0 quantized Qwen 3.5 35B A3B model with an AMD 7900 XTX on a computer that has a 14 years old CPU. It works Great!

I’m trying to learn more about LLM inference; all the hardware factors and what determines how fast or slow it is.

I found this thread because I was looking for some information on how the 4-bit ā€œfloatā€ (if you can call it that) / NVFP4 thing (new in blackwell) impacts performance.

Main issue I’m having with the 7900 XTX is that prompt processing aka ā€œprefillā€ takes a long time. and I have been told that usually the way to make that ā€œmore fasterā€ would be a GPU with more FLOPS.

If I understand correctly, the point of nvidia doing NVFP4 was to support more total flops per watt or flops per amount of transistors.

I’m interested in benchmarks that show that difference in prompt processing speeds, for example a 4-bit quantized model that’s running on a GPU like mine or a pre-blackwell whose smallest floating point number is 16 bits, vs a 4-bit quantized model running on a blackwell GPU like 5080 and using NVFP4.

Does this benchmark show that? Was it using the 4-bit multiply / add instructions that are on the 5080 but not on the older one?

And if no, do you know of any existing benchmarks that could show that ?

1 Like

Greetings @forestj you’re asking a lot of great questions trying to get the most out of your equipment!

tl;dr;

try running llama-server -ub 4096 -b 4096 to increase batch sizes to see if it helps improve your prompt processing (prefill) speed.

detailed rambling

A 7900mXTX is a solid value card for both gaming and ai inference workloads. So 24GB VRAM running at 960 GB/s lpddr6 memory bandwidth.

Yes, in general, Prompt Processing (prefill) tends to be compute bottlenecked because it has all of the context available and just needs to do a ton of matrix multiplications as fast as possible. So more FLOPS can help as well as efficient implemented kernels (e.g. llama.cpp Vulkan backend kernels for matmul etc).

In general, Token Generation (decode), is memory bandwidth bottlenecked because it is autoregressive. This means that for each token it generates (half a word or so on average in english), it has to fetch all the weights from memory over and over and over for each token. So no longer a problem of compute, but of accessing those weights from memory fast enough.

So nvfp4 is kind of confusing and a new beast. It is a quantization type not just a dtype. That means that it is a data structure with say 128 numbers ā€œweightsā€ along with some scaling factors and such per block of 128. This is what most quantizations kind of look like. A dtype ā€œdata typeā€ is more simple like uint8 0-255, or fp16 ieee standard floating point 16bit type.

nvidia added hardware support for nvfp4 quantization type. This is a new thing, before this GPUs would have to have a ā€œkernelā€ to unpack/multiply handle the quantization data structure using the limited dtype support they have e.g. fp16 or newer fp4 hardware registers.

Now like everything, there are trade-offs. nvfp4 might be fairly fast on newer blackwell GPUs that support it, assuming the inference engine has the correct backend support and optimized kernels to actually make use of the hardware registers.

But, as a quantization type, now you’re locked in and it may not be as good as some other GGUF types e.g. iq4_xs , iq4_kss, iq4_kt across all models and tensors.

In general I’d suggest using either mainline llama.cpp llama-bench -d to test speeds or even better ik_llama.cpp llama-sweep-bench … so you can test how changing parameters effects your actual model in the same runtime you’ll be using it for llama-server.

It isn’t quite as simple as you describe for all models, because MoE models tend to keep attn/shexp/dense layers at full 8bpw quality and only use the 4bpw quality for the sparse routed experts.

No, this benchmark was before llama.cpp even supported nvfp4, and the quant is labeled as an IQ3_Kmix.

./build/bin/llama-sweep-bench \
  --model "$model" \
  -c 36864 \
  -muge \
  --merge-qkv \
  -ub 512 -b 2024 \
  -ngl 999 \
  --threads 1 \
  --warmup-batch \
  -n 128

So plug in your Q4_0 model and test it like so. Increase to -ub 2048 -b 2048 and see again. Then increase to -ub 4096 -b 4096 and if you don’t OOM (due to larger buffer sizes) you might see better PP speeds, but I’m not sure on Vulkan backend.

You’ll probably want to build my branch for llama-sweep-bench for mainline llama.cpp which I keep here: GitHub - ubergarm/llama.cpp at ug/port-sweep-bench Ā· GitHub

Cheers!

1 Like