The Beginning
I had the change to get into Home Assistant at the end of last year and found myself in the perfect storm of new development in voice integration.
But, what I had in mind, was far from what most would deem reasonable: I wanted a cheap (ish) box to run everything off of, including Home Assistant and all it’s voice integrations.
I didn’t really know what I was getting into, but it turned out a great learning experience at least. Fueled by hopes and dreams I ventured on!
Hardware setup
To get into the numbers and move along the post quickly here’s what I settled on for the voice chain (there’s no doubt way more on the internet for everything else but, if you’re interested in what else I use in my home just shoot me a DM):
- Minisforum/Kodlix GD90 (Intel i9 12900HK, Iris Xe 96EU, 2x32GB 3200MHz CL20 RAM)
- Raspberry Pi Zero 2W
- Raspiaudio Mic+ V3
Software Setup
And for software here’s what I went for:
- Proxmox 8.4 host + i915 SRIOV DKMS
- Home Assistant OS VM (easier to manage and backup)
- Ubuntu Server 24.04 VM + Docker + ipex-llm
- Wyoming Satellite on the Pi Zero 2W + Open Wake Word
- Piper (TTS add on for Home Asssitant with custom trained voice)
- Whisper (STT add on for Home Assistant)
I didn’t have any specific goal in mind beside being able to “talk” to my house and get some things done hands free. I didn’t need the knowledge of the whole world in reach at all times. I much prefer searching the internet and sit down instead of having quick bits of informations thrown at me.
Challenges
Running Accelerated LLMs
The first challenge I had to face was getting any kind of LLM to work well with the iGPU in my mini PC.
As many of you have figured out by my hardware list there isn’t much love for the Intel iGPUs so I had to dig until I found something that worked with what I have and allowed me to put to use that iGPU I so hard managed to assign to different VMs.
I made my life harder because using the SRIOV driver makes working with the iGPU inside VMs a bit difficult for two reasons:
- on Linux you’re mostly guaranteed success with the guest drivers on Debian based distros
- not all functions are supported so using Intel LLM tool chains isn’t really possible
What solved my issues was finding the ipex-llm container in a dark corner of the internet, at least for me.
This solution is a packaged version of the Intel tool chain for LLM that works as a “one click” solution on almost any system you can think of. The Docker overhead is very minimal and it’s easy to deploy. But they’re slowly updating it so the chance of running the newest models at all times might not be there.
Performance
Once I managed to find a suitable platform I moved to testing what size models would allow me to get a decent balance of performance and functionality. It didn’t really matter what model it was, that part will be addressed later on.
I started with 7B parameters models since I had RAM to spare and I was familiar with the reponses I could get out of the most popular like Llama3 and Gemma 2. I fed them the same prompt I had run on my desktop (2070 Super) multiple times and recorded speed and quality of the answer. To cut it short, those models were too big for the computing power at my disposal and that caused them to allucinate sometimes. So I had to cut in half or less the amout of parameters I could work with.
That meant moving to around 3B parameter models and the one I had some experience with is Llama3.2.
I started testing it out the same way I did with the others and, after a first warm up, I could get around 16 t/s which I figured it wasn’t too bad.
Trying to speed things up I started messing with layer offloading on the CPU and sizing my VM accordingly.
In doing so I discovered that 6P Cores had the same performance as the iGPU, with higher power consumption and that offloading any layer to the CPU only made the prompts generate slower, around 13 or less t/s.
That crushed my spirits a little because I was hoping to get a bit better performance putting those big cores to use, but that wasn’t the case.
The Right Model For The Job
Once I had the foundation solid I started experimenting with Home Assistant and it’s interactions with LLMs to see what would work, what it could do, how fast and so on.
Starting with Llama3.2 I quickly learned it wasn’t the best model to work with in Home Assistant because it couldn’t really do much of anything, beside annoying me with suggestions on .json requests I should use to do what I wanted to do. That wasn’t gonna cut it.
All the issues I had stemmed from the fact that I needed to find a model with strong tools support. If they don’t have that it’s not even worth trying it.
While I was searching and trying the suggestion to use Qwen2.5 popped up and decided to try it.
So far it’s the best small model I ever tried with Home Assistant and I don’t think there are any that can rival it, unfortunately.
But, even if i found the right model, the answers still take quite some time to be generated (even a minute or more sometimes) and need a large context to be interpreted properly. I had to set mine to 32768 to get the model to do stuff. Yet I can’t ask the model to turn on two lights at once because the context window isn’t big enough in my case. What makes it even harder is that Home Assistant always passes to the LLM the initial prompt you’ve given it, so that takes space in the context window and slows the model down. Text generation cannot be compared to performances in Home Assistant.
I’d say that in use for Home Assistant a model is about 50% as fast as it is on it’s own.
I’m not too happy about it but I also don’t feel like complaining because parsing through hundreds of pages, messages and guides really improved my understanding of how things work, what may be the needs and what can be expected.
A fun aspect of using an LLM is that, with the right prompt, it can have the personality of your choosing and sell it pretty well with the right voice.
Update!
The Home Assistant team has made great strides in voice control and has implemented voice generation during text generation. So, as soon as there’s some text buffer the system will start generating and playing back voice even if the output hasn’t been fully completed. This is available from 2025.8.0 and has sped up everything by quite a lot.
Patching Up The Holes
Once I figured out I can’t wait a minute for this model to turn on a light or trying to give me weather forecast details I decided to rely more on the Home Assistant local parsing of commands.
Home Assistant allows users to interact through voice commands in two ways: through a local parser or through the LLM. It’s possible to use the voice assistance in “cascade mode” so that if a command is not parsed by Home Assistant it goes through to the LLM that should answer.
Local parsing is way faster, but it requires the user to ask things in very specific ways so that Home Assistant can understand you. Even the smallest deviation will cause the command to go to the LLM and, in my case, quite the wait for an answer.
To make sure the local parsing is as useful as possible I decided to give all my lights and sensors aliases so that if I call them in a slightly different way I’d have a better chance of receiving immediate response.
This proved extremely useful with numbered devices. For example: if you have “patio light 1” in Home Assistant and ask it to turn it on it might have troubles putting together that “one” and “1” mean the same thing. So giving the “patio light 1” the alias “patio light one” makes it more likely to work.
Also making sure all devices are part of an area in your home is a good way to be sure your commands will execute promptly and correctly all the times.
All this solved the issues with sensors and lights, but the weather required a bit more finesse to be solved. As much as Home Assistant can read correctly all the details about the weather through integrations it was still slow and cumbersome to use if I want to know if I need to bring and umbrella or not on my way out.
To my help came this brilliant weather blueprint that, with a little tune up, allows proper local parsing of questions on weather forecast through dedicated commands you choose.
This blueprint turns into an automation for you and it’s great. The output can be customized to your content.
Going forward I’m going to set up more and more automations and blue prints to avoid going through the LLM, making use of it in edge cases.
Sure it’s not what I had in mind but I’m having fun so that’s surely a plus. It’s also good for all the people using ChatGPT or other integrations paying per token. The less token you send them, the better.
The Final Touch
I don’t think this requires a whole section but I’d like to dedicate it one because it took me quite a bit to figure it out.
Once I put all this together I was really not feeling the stock voices and the start and end command sounds. So once again a took it to the internet to find out how to make it sound how I wanted.
For the start, end and timer end sounds I just uploaded some different .wav files to the Wyoming satellite and that was done. Pretty easy.
When it comes to voices things get a little more tricky because Piper uses .onnx files and most pre-trained voices are in other formats that cannot be converted to .onnx. There are lots of voice changers that give out voices, but they’re completely useless.
So, if the voice you want for your assistant is not available, is possible to train one from scratch at a cost of your time and patience.
To do so I followed Network Chuck’s tutorial which worked. But I made sure to not make the same mistakes he made to get good results with just 40 minutes of audio:
- Break the audio clip of your choice in full sentences
- Have the cleanest audio possible
- Review the transcriptions Whisper makes of those sentences and correct every single mistake
- Check if too many utterances are skipped once the training begins. 10% or under is good, over you’re risking wasting your time.
This requires an RTX GPU so everyone else is out of the game, unless they’re willing to use a Google Colab project that allows you to use an Nvidia T4 GPU for free. I’ve not gone through that process, neither Network Chuck did (he didn’t even mention it) but it apparently work.
All the software setup is pretty well documented in the video so I’ve not gone through describing it to avoid cluttering the post too much.
There aren’t any particular pain points with the setup that aren’t explained in the video or can’t be solved with a quick search.
Lingering Issues
Home Assistant is constantly changing and evolving so there can be issues that appear and disappear as time goes on.
The one I’m facing right now are weird lock ups of the satellite that remains stuck in a constant “listening” state and requires a reload of the integration to make it work again. I thought I had solved it increasing the listening time but, after a whole week of rock solid performance, it did it again. This can be frustrating at times.
I don’t see any errors in logs so I think this is a bug that needs to be ironed out on the Home Assistant side. Or can be overcome if you run Open Wake Word on Home Assistant because decoupling them makes it easier for the main system to “override” the listening state on the satellite.
Though that’s not a recommended setup anymore for privacy (UDP voice packets flying unencrypted in your network) and general direction of the project.
Conclusions
With everything I got going on it took me months to get to this point, but I’m really happy about it! I guess that if I had access to way more powerful hardware I could now make use of it knowing much better what I’m going to get out of it.
I wanted to post this for quite a while and finally found some time to do it.
My setup is nothing that can be even compared to an Amazon Alexa or a Google Home; it’s not quick and it’s not as smart (mostly) but it’s something I cobbled together myself and it’s not spying on me 24/7. If I wanted to rip it out I could and nobody could stop me. Freedom as I like to experience it with my tech. I asked my dad to use it and it was hilarious to see how he would fire up a question without waiting for the start sound. When he figured that out I could hear the Jeopardy wait music playing as the fan spins up and the LLM cooks up the answer (and the CPU).
But it also gave me a perspective on what it takes to build something that’s immediately ready to answer your every question.
The last thing I want to set up on my satellite is music playback. I think that’s gonna be a game changer and make me use it way more than how much I use it right now. Even though using it for vocal alerts is great and something I’d really miss if I had to move to any other voice assistant.
I love how it’s possible to set up any sort of automation and make it speak when it’s triggered.
For example you could have someone at the door, be recognized by the system and have the assistant say to you “there’s somebody at the door”…
even with this audio clip!
In the end, even when not using an LLM, is worth exploring this world and having a tool to make the daily life more convinient, while staying private!
Cheers! Let me know what you think, what more you’d like to see added, I’ll answer any question to the best of my abilities.