Small LLM LLM backend on Strix Halo Hardware

Hey everyone,

I’ve been tinkering with a Strix Halo Box, Docker, and a handful of mixed services lately. Since I was inspired by the L1 videos and this forum, I thought I’d share my setup with you all.

It is a Self-hosted LLM inference stack: local model Backends exposed through a single Gateway, based on docker.

I’m also planning to add optional Imagegen Mode (ComfyUI)

I already optional Grafana/Prometheus monitoring with a small pre-provisioned dashboard.

more here:

github[dot]com/hgnnmnn/ai-stack ( i’m not allowded to add links) :confused:

feel free to comment, share and help me improve it :slight_smile:

Thanks for your attention.

cheers
H