B70 or R9700? (help decide)

Incorrect. When the R9700 is running in a slot attached to the chipset, the extra hop causes 30-40% performance degradation. In real terms, it acts like a <100W GPU in that use case.

Talking to the llama.cpp devs, the explanation is that the kernel dispatch from the CPU is very sensitive to PCIE latency.

You absolutely don’t want to do this; it’s the reason I downgraded my AI server from a 13th gen Intel CPU to a 5700X and ROG Strix X570 - to get the x8/x8 split from the CPU without spending a fortune.