B70 Info
A B70?!? In this economy? And yet this may be the “best” path to 128gb vram for around/less than the cost of DGX spark. Assuming you have a machine to plug them into.
B70 32gb AI llm-scaler (vLLM) testing
vllm serve /llm/models/hub/models--Qwen--Qwen3.5-27B/snapshots/b7ca741b86de18df552fd2cc952861e04621a4bd --served-model-name Qwen/Qwen3.5-27B --port 8000 --no-enable-prefix-caching --enable-chunked-prefill --max-num-seqs 128 --block-size 64 --enforce-eager --dtype bfloat16 --disable-custom-all-reduce --tensor-parallel-size 4
Avg generation throughput: 540.0 tokens/s, Running: 50 reqs,
DANG!
============ Serving Benchmark Result ============
Successful requests: 50
Failed requests: 0
Benchmark duration (s): 69.22
Total input tokens: 51200
Total generated tokens: 25600
Request throughput (req/s): 0.72
Output token throughput (tok/s): 369.83
Peak output token throughput (tok/s): 550.00
Peak concurrent requests: 50.00
Total token throughput (tok/s): 1109.48
---------------Time to First Token----------------
Mean TTFT (ms): 11467.51
Median TTFT (ms): 11316.84
P99 TTFT (ms): 21193.65
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms): 110.70
Median TPOT (ms): 111.14
P99 TPOT (ms): 121.26
---------------Inter-token Latency----------------
Mean ITL (ms): 110.70
Median ITL (ms): 92.52
P99 ITL (ms): 567.33
==================================================
Keep in mind this is 50 requests at once, however.
For a single request the floor performance with dynamic FP8 quant of qwen 27b on a single gpu was
Avg prompt throughput: 85.4 tokens/s, , Avg generation throughput: 13.4 tokens/s,
HOWEVER: I would not recommend a single B70 for Qwen 27B dense in fp8 dynamic quant. For vLLM benchmarking I had to lower the context and set the max gpu memory utilization to 0.8 or it was unstable. Two B70s for the Q8 Qwen 3.5 27b was fine. Similarly, there was simply no room to work with Qwen 27b bf16 on two B70s.
The model perplexity/stability was reasonable in dynamic Q8. The native FP8 model crashed on startup, however. So if you have trouble, download the native bf16 and use the dynamic quant? That seems like an odd happenstance to me, but I assume it’s a me problem.
Two B70s and Qwen 3.5 27b in dynamic FP8 mode
vllm bench serve --model Qwen/Qwen3.5-27B --dataset-name random --random-input-len 1024 --random-output-len 512 --num-prompts 8 --request-rate inf --port 8000
============ Serving Benchmark Result ============
Successful requests: 8
Failed requests: 0
Benchmark duration (s): 41.86
Total input tokens: 8192
Total generated tokens: 4096
Request throughput (req/s): 0.19
Output token throughput (tok/s): 97.84
Peak output token throughput (tok/s): 112.00
Peak concurrent requests: 8.00
Total token throughput (tok/s): 293.53
---------------Time to First Token----------------
Mean TTFT (ms): 1961.21
Median TTFT (ms): 2113.37
P99 TTFT (ms): 2826.61
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms): 77.84
Median TPOT (ms): 77.56
P99 TPOT (ms): 80.33
---------------Inter-token Latency----------------
Mean ITL (ms): 77.84
Median ITL (ms): 76.27
P99 ITL (ms): 78.92
==================================================
Interestingly, this is the type of latency I was looking for (hoping for) in the video. 8 requests at once, though.
Single request is still capped around 14 t/s
============ Serving Benchmark Result ============
Successful requests: 1
Failed requests: 0
Benchmark duration (s): 38.63
Total input tokens: 1024
Total generated tokens: 512
Request throughput (req/s): 0.03
Output token throughput (tok/s): 13.25
Peak output token throughput (tok/s): 14.00
Peak concurrent requests: 1.00
Total token throughput (tok/s): 39.76
---------------Time to First Token----------------
Mean TTFT (ms): 443.08
Median TTFT (ms): 443.08
P99 TTFT (ms): 443.08
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms): 74.73
Median TPOT (ms): 74.73
P99 TPOT (ms): 74.73
---------------Inter-token Latency----------------
Mean ITL (ms): 74.73
Median ITL (ms): 74.87
P99 ITL (ms): 76.95
==================================================
all other parameters being equal.
Gaming?
Watch this space. Currently (3/25) we have
UPDATE 3/26: yay driver works now! Here are the results
Cyberpunk 2077: Phantom Liberty - Ultra
Shadow of the Tomb Raider - Highest
Monster Hunter Wilds - Medium







