LLM Homelab patchwork - need advice

tl;dr - this is probably too much for me to dump into a single post, but I appreciate anyone that feels like reading all of this!

Hey everyone,

I’ve been building out a multi-machine local AI setup over the past several months as an experiment: can I get a “chief of staff” agent running on my own hardware that actually feels useful? One of them now basically functions as a personal assistant, which is wild but not quite what I need it to be.

The one thing I think is genuinely interesting: STT/TTS over Telnyx SIP + LiveKit for sort of a “Jarvis” experience. I built a rudimentary Android app for two-way comms, though usually I just call the phone number from my cell and have the agent in my bluetooth earbud when I dare go outside.

But here’s where I’m stuck: it feels like I’ve been building things that don’t lead anywhere productive. The whole stack breaks roughly monthly and I waste a week recovering and updating. I want to build something that actually works, not keep patching it together.

What I’m Trying to Achieve

I was recently laid off and I’m trying to build something practical that gets me back in the industry, or helps me pivot careers if that’s where this goes. The goal is to turn this hardware into a portfolio-worthy solution for:

  • Software QA: automated testing workflows using local LLMs
  • Accessibility auditing: tooling that helps catch WCAG issues before they reach production
  • Other relevant tasks that help me build a demonstrable body of work on GitHub
  • Agentic Harness Consulting: I am doing some consulting with a friend’s nonprofit to help them have their own in-house “Claude” due to data privacy concerns with their clients.
  • Open to other suggestions

The hardware’s here and the agents are running. I just need direction on how to channel all this into something that gets me back in the industry, or helps me pivot if that’s where it goes.


My Current Fleet

Machine 1: Primary Inference Workstation

  • CPU: AMD Ryzen 9 9950X3D
  • RAM: 128GB DDR5
  • Motherboard PCIe lanes: MSI x870e Pro: Slot 1 = x16 Gen5; Other x16 slot is only PCIe 3.0 x1 (I know that’s limiting)
  • OS: Windows 11: originally for work requiring Windows, used to get acquainted with LM Studio
  • Installed GPUs: RTX 5090 (32GB) + RTX 5060 Ti (16GB)
  • Role: Primary inference box. Slated for Proxmox swap soon to run Llama.cpp and host more agent VMs (CPU/RAM idle most of the time)

Machine 2: Delegation / Fallback Node

  • CPU: AMD Ryzen 9 5900X (12 cores / 24 threads)
  • RAM: 64GB DDR4
  • OS: Kubuntu 26.04 LTS
  • Installed GPUs: RTX 3090 (24GB) x2 (dual GPU setup = 48GB total)
  • Role: Delegation tasks, llama.cpp GGUF model serving

Machine 3: Asus Strix Halo Tablet/Laptop

  • SoC: AMD Radeon 8060S (Strix Halo)
  • RAM: 128GB LPDDR5X (96 GiB carved out for GPU via BIOS UMA split)
  • OS: Linux-only (no Windows partition)
  • Installed GPUs: AMD Radeon 8060S integrated
  • Role: Intended for larger models and slower coding work, but struggling to find its role in my setup

Machine 4: Inference Node

  • CPU: AMD Threadripper Pro 3955WX
  • RAM: 128GB
  • Motherboard PCIe lanes: Gigabyte MC62-G40 (EPYC workstation board): multiple slots full x16, one slot electrically x8; most available slots are Gen4 x16
  • OS: Ubuntu Server 26.04 LTS
  • Installed GPUs: RTX 5090 (32GB) + RTX 4090 modded to 48GB + RTX 3080 modded to 20GB
  • Role: Primary CUDA inference node. Has a second 20A circuit nearby, so I could add more GPUs if needed.

Additional PC Parts (Disassembled)

I have three disassembled PCs in inventory that I could piece together if needed:

  • Threadripper 2950X build: Threadripper 2950X CPU + 128GB DDR4 (PCIe 3.0 x64 lanes): no motherboard or GPU currently installed
  • Two Z690 builds: Two Z690 motherboards with 32GB DDR5 each; two Intel LGA 1700 CPUs available (i7-13600K and i7-12900K): no GPU installed for either

GPU Inventory

Currently Installed / Active (in fleet machines above)

  • RTX 5090 (32GB) ×2: one in Machine 1, one in Machine 4
  • RTX 5060 Ti (16GB): Machine 1
  • RTX 3090 (24GB) ×2: Machine 2
  • RTX 4090 modded to 48GB: Machine 4
  • RTX 3080 modded to 20GB: Machine 4
  • AMD Radeon 8060S integrated (~16GB shared): Machine 3

Spare / Unused GPUs in Inventory

  • Intel Arc Pro B70 (32GB VRAM each) ×2: 64GB total, PCIe workstation cards not currently deployed. Curious if they’re viable for local AI/LLM inference workloads
  • RTX 5060 Ti (16GB): A second one I haven’t installed yet.

What I’m Running

Hermes Agent as my primary framework. Openclaw was an earlier attempt but proved too much maintenance overhead to sustain.

My Questions

1. Model Selection and GPU Fleet Optimization

Currently running mostly Qwen3.6-35b-a3b for speed (works great on the 5090 + my SIP setup) and Qwen3.8-27B Q5 on the other 5090 for more thinking. The rest is a revolving door of trying out models for STT, TTS, image gen, and video gen: I’m just playing with those last two.

Given that I have both NVIDIA (RTX 5090, 3090s, 4090 mods, 5060 Ti) and AMD (Radeon 8060S on Strix Halo), what model architectures work best across this mix? Should I be looking at models that can split computation across heterogeneous GPUs, or is it better to specialize each box for specific model types? Any recommendations for quantization strategies that work well on both CUDA and Vulkan/ROCm backends?

2. Multi-Machine Delegation Strategy

My goal is to run my own models on my own hardware as much as possible. I only use cloud providers when I hit a wall and if what I’m working on doesn’t have any PII concerns with the data.

How do others handle delegation across different hardware profiles with varying compute? I’m using Hermes’s built-in delegation with provider routing, but wondering if there are better patterns: model sharding vs specialized roles per machine, or workload distribution that accounts for GPU type and VRAM availability.

3. Agent Framework Comparison (Turnstone)

For the coding side, I’m considering Turnstone agent. Has anyone used it for software QA or accessibility auditing workflows? How does it compare to Hermes for multi-machine setups where you need reliable task delegation?

4. Hardware Planning Before Next Purchase

I’m considering adding more hardware but want to be strategic: should I prioritize NVIDIA for CUDA ecosystem consistency, or invest in AMD/ROCm given Strix Halo’s capabilities? Any recommendations for inter-machine networking setup to optimize model serving across nodes? Is there value in a dedicated orchestration/memory machine, or can the current host handle it better?

Any advice, recommendations, or war stories from others running similar setups would be greatly appreciated. Thanks in advance!