Sanity check my Qwen3.8 5090 Results; new to Local AI

Hey Everyone! First post, kinda shy, be gentle :wink:

So - I don’t have a TON of experience in Local AI - I do work in the AI Adoption / Governance / Cybersecurity Space, and am lucky enough to have active subscriptions across every major frontier model - however, I haven’t messed with Local AI in about 8 months. But - I’ve been hearing a ton about Qwen3.8-27b, so I loaded it up on my rig at the house in LM Studio, and ran a few tests with a harness doing some pentesting on a dev vm, and was pretty happy with the initial results, though it did get stuck in a few loops which ate through my context. So - I’m now in the process of building a 20-40b focused MCP server specifically for red teaming / blue teaming with Kali.

Before I get too deep in that process (I’m around 15ish hours in so far) - I want to get some opinions on if my initial testing is “Stupid” or “The old way”.

This is my current setup:

HW Spec
CPU Intel Core i9-13900K
GPU NVIDIA RTX 5090
RAM 64 GB 7200 MT/s DDR5
OS Windows

I’m using LM Studio right now - which I’m sure isn’t great - but this is the only thing I’ve used for tasks like these - open to suggestions if this is the wrong setup.

Below is the performance I have been getting - this is another area I want a sanity check on. Before that, a bit of context for my decisions - this test was scoped specifically for a phased, agentic red team assessment on a single device. Below are my results:

Model / GGUF Quant MTP Max Decode Prompt Eval Approx. VRAM Relative Speed
chimingw/Qwen3.8-27B-Uncensored-OrcaRouter-GGUF Q6_K 2 12.38 tok/s 1,016.55 tok/s ~28 GB 1.00×
JonathanColetti/Qwen3.8-27B-Uncensored-GGUF Q6_K 1 ~22.3 tok/s — ~27.5–28.5 GB 1.80×
JonathanColetti/Qwen3.8-27B-Uncensored-GGUF Q6_K 2 31.60 tok/s ~1,844.91 tok/s ~27.5–28.5 GB 2.55×
JonathanColetti/Qwen3.8-27B-Uncensored-GGUF Q5_K_M 2 ~44.7 tok/s ~2,045.19 tok/s ~26.3–26.5 GB 3.61×
JonathanColetti/Qwen3.8-27B-Uncensored-GGUF Q4_K_M 2 ~47.4 tok/s ~2,267.78 tok/s ~22.9 GB 3.83×

So - open ended question; give me your feedback. Should I switch away from LM Studio? Are my speeds decent? Am I leaving some performance on the table?

You can ask your agent to help you set these up:

With gittensor-model-hub/Qwen3.8-27B-DSpark-NVFP4 · Hugging Face or z-lab/Qwen3.8-27B-DFlash2 · Hugging Face you should be able to get some crazy speeds. Personally, I think FP8 is the lowest KV you want to go w/o long-context degradation, but should be able to get 256K (or close) full context anyway.

Short answer: you should be able to get 5X faster prefill, 3X faster decode on a 5090.

Interesting; looks like I need to dive into vLLM - I’ve never used this before. I’ll likely need to use orcarouter/Qwen3.8-27B-Uncensored-NVFP4 - but - looks like I can see a major improvement in speed.

Here’s a datapoint from my 5090; I’ve been running qwen3.8 with GitHub - Neroued/ninfer: High-performance single-GPU inference for selected model checkpoints and GPUs. · GitHub . Similar to you some of this local model stuff has outpaced me since I last checked in, don’t completely grasp the impact of MTP and KVcache quant yet. But this ninfer seemed to have some optimizations for single gpu 5090 arch, but it may not work with the finetunes you need. I also pulled one of the uncensored models with lmstudio ran with llama-benchy as a comparison.

specs
9950x3d
5090
96GB 6400mhz
linux

Model / Runtime Quant MTP Max Decode Prompt Eval Approx. VRAM Relative Speed
JonathanColetti/Qwen3.8-27B-Uncensored-GGUF via LM Studio/llama.cpp Q6_K 2 72.03 ± 5.66 tok/s 2,430.18 ± 24.45 tok/s ~28.5 GB total GPU use 1.00×
neroued/Qwen3.8-27B-nvfp4-NInfer via NInfer mixed NVFP4/FP8 3 124.08 ± 10.56 tok/s 8,143.82 ± 52.68 tok/s ~28.5 GB total GPU use 1.72×

Looks like these models tok/s can vary pretty heavily based on the type of output being generated. So, hard to say how useful this will be for you without comparing identical prompts. I’m interested in your usecase though, uncensored model + red teaming on 5090. Let me know if you find some optimal config with vLLM

This, I suspect, is your problem. I don’t know much about optimising for Blackwell, but this is my single-request result on 2 x R9700s, with Qwen 3.8 27B FP8 under vLLM:

| model         |             test |             t/s |      peak t/s |        ttfr (ms) |     est_ppt (ms) |    e2e_ttft (ms) |
|:--------------|-----------------:|----------------:|--------------:|-----------------:|-----------------:|-----------------:|
| primary_agent |           pp8192 | 3375.21 ± 22.63 |               |  2206.71 ± 10.85 |  2203.94 ± 10.85 |  2206.71 ± 10.85 |
| primary_agent |            tg256 |    69.49 ± 5.32 |  83.67 ± 4.11 |                  |                  |                  |
| primary_agent |  pp8192 @ d65535 |  2720.21 ± 5.57 |               |  24564.75 ± 6.45 |  24561.98 ± 6.45 |  24564.75 ± 6.45 |
| primary_agent |   tg256 @ d65535 |    53.88 ± 6.17 |  63.25 ± 2.47 |                  |                  |                  |
| primary_agent | pp8192 @ d131072 |  2347.35 ± 0.80 |               | 53829.53 ± 49.04 | 53826.76 ± 49.04 | 53829.53 ± 49.04 |
| primary_agent |  tg256 @ d131072 |   67.99 ± 34.74 | 90.67 ± 51.91 |                  |                  |                  |

And this is the same thing on my CMP 170HX Ampere server, except INT8 instead of FP8:

| model     |             test |             t/s |      peak t/s |         ttfr (ms) |      est_ppt (ms) |     e2e_ttft (ms) |
|:----------|-----------------:|----------------:|--------------:|------------------:|------------------:|------------------:|
| agent_27B |           pp8192 | 4226.18 ± 21.34 |               |    1749.25 ± 9.27 |    1746.66 ± 9.27 |    1749.25 ± 9.27 |
| agent_27B |            tg256 |    65.66 ± 5.00 |  74.67 ± 8.26 |                   |                   |                   |
| agent_27B |  pp8192 @ d65535 | 2685.51 ± 22.38 |               | 24879.52 ± 253.31 | 24876.93 ± 253.31 | 24879.52 ± 253.31 |
| agent_27B |   tg256 @ d65535 |    55.94 ± 7.34 |  65.00 ± 4.24 |                   |                   |                   |
| agent_27B | pp8192 @ d131072 | 1809.29 ± 13.63 |               | 69808.59 ± 586.46 | 69805.99 ± 586.46 | 69808.59 ± 586.46 |
| agent_27B |  tg256 @ d131072 |   53.31 ± 18.90 | 59.53 ± 17.59 |                   |                   |                   |

Note that those results were at context depths of 0, 64k and 128k. Also, the CMP 170HX was thermal throttling quite a bit towards the end, so prefill dropped off significantly. I really need to sort out a better cooling solution.

Anyway, the point is that both of these solutions get better numbers than you’re seeing from your 5090, with much higher-quality models so something is off - they shouldn’t even be close. You should at least start with moving to llama.cpp instead of using LM Studio’s neutered version. If you can stomach the hassle of dealing with vLLM, go for it - for the sake of expediency you’re better off finding a decent docker-compose.yml to get you started, rather than descending into dependency hell (plenty of other people have done the hard work here, so just find one that suits your architecture).

Just wanted to pass my updated results, got vLLM set up and done some additional benchmarking:

Model / GGUF Quant MTP Max Decode Prompt Eval Approx. VRAM Relative Speed
chimingw/Qwen3.8-27B-Uncensored-OrcaRouter-GGUF Q6_K 2 12.38 tok/s 1,016.55 tok/s ~28 GB 1.00×
JonathanColetti/Qwen3.8-27B-Uncensored-GGUF Q6_K 1 ~22.3 tok/s — ~27.5–28.5 GB 1.80×
JonathanColetti/Qwen3.8-27B-Uncensored-GGUF Q6_K 2 31.60 tok/s ~1,844.91 tok/s ~27.5–28.5 GB 2.55×
JonathanColetti/Qwen3.8-27B-Uncensored-GGUF Q5_K_M 2 ~44.7 tok/s ~2,045.19 tok/s ~26.3–26.5 GB 3.61×
JonathanColetti/Qwen3.8-27B-Uncensored-GGUF Q4_K_M 2 ~47.4 tok/s ~2,267.78 tok/s ~22.9 GB 3.83×
orcarouter/Qwen3.8-27B-Uncensored-NVFP4 / vLLM NVFP4 2 40.15 tok/s — ~30.4 GB 3.24×
orcarouter/Qwen3.8-27B-Uncensored-NVFP4 / vLLM NVFP4 OFF 67.74 tok/s — ~29.8 GB 5.47×
orcarouter/Qwen3.8-27B-Uncensored-NVFP4 / vLLM NVFP4 1 71.23 tok/s — ~30.2 GB 5.75×
1 Like

I’ve been running the cyankiwi nvfp4 ninfer+fp8 build someone else posted in another thread here. What i also noticed coming from LM studio is that LM studio automatically cuts out context which doesn’t happen in Vllm.

This is my command for ubuntu vllm to use with openhands that pretty much fills the vram almost completely.
(don’t forget to install cuda container toolkit and docker to use this)

#!/usr/bin/env bash

# Exit on errors, unset variables, and failed commands in pipelines.
set -Eeuo pipefail

# Print the command that is about to run.
set -x

IMAGE='vllm/vllm-openai@sha256:0db5553091d59de260f67696fff89026e7376ec0d326111ba78fb004a27778a0'
ATTENTION_BACKEND='FLASHINFER'   # TRITON_ATTN | FLASH_ATTN | FLASHINFER
KV_DTYPE='fp8'                   # bfloat16 | fp8 | fp8_per_token_head

docker run --rm --gpus all --network host --ipc host --shm-size 32g \
  -v /home/user/qwency:/models/qwen38-cyanwiki:ro \
  -e HF_HUB_OFFLINE=1 \
  -e TRANSFORMERS_OFFLINE=1 \
  -e VLLM_NO_USAGE_STATS=1 \
  "$IMAGE" /models/qwen38-cyanwiki \
  --served-model-name qwen3.6-27b \
  --tensor-parallel-size 1 \
  --dtype bfloat16 \
  --kv-cache-dtype "$KV_DTYPE" \
  --attention-backend "$ATTENTION_BACKEND" \
  --gdn-prefill-backend triton \
  --max-model-len 180000 \
  --max-num-seqs 4 \
  --max-num-batched-tokens 8192 \
  --gpu-memory-utilization 0.95 \
  --enable-prefix-caching \
  --seed 1 \
  --trust-remote-code \
  --host 0.0.0.0 \
  --language-model-only \
  --port 1234 \
  --enable-auto-tool-choice \
  --tool-call-parser qwen3_coder \
  --reasoning-parser qwen3 \
  --speculative-config '{"method":"mtp","num_speculative_tokens":1}'

I just installed and started using Open Harness today, you can find it on GitHub.. Qwen had been running very slowly in Hermes.. (think, think, think) Open Harness changed all of that.. Just wanted to share that.. I am very much pleased..