Take a stab at it, what could go wrong?
Not sure what you’ve been talking about, specifically, but my M3 Pro Macbook at work runs ~16b models at q4 wonderfully. Something like 15tps.
Take a stab at it, what could go wrong?
Not sure what you’ve been talking about, specifically, but my M3 Pro Macbook at work runs ~16b models at q4 wonderfully. Something like 15tps.
I did. I didn’t wait for a reply xD. It worked perfectly, like out of the box. Less configuration for that machine than my Linux machine. It just used the M4 GPU out-of-the-box.
I found a nice little tool on GitHub that lets you more easily manage your Ollama models with a TUI, and it has a feature to link them to LM-Studio. So I get the best of both worlds!
I’ve been talking about my new MacBook for like the past two weeks straight. Seems like I’ve only mentioned it once on this forum though.
Ahh, I don’t read every thread, tbh.
Yeah, Apple’s ecosystem is pretty damn dialed. That was my experience as well.
Ohhh, that thing looks pretty slick!
Who does. I was just trying to imply that I talked about it on this forum less than I thought I had.
Do you have a link to that?
The link is embedded in the text “a nice little tool on Github”.
oh, I missed that, thanks.
These are my installation notes for Ollama and Open-webui with ROCm.
Portainer shows you when updates for the containers are available.
The maintenance of the installation is very easy.
# Ubuntu Server 25.04
##Pre-installation instructions
# 1. Systemaktualisierung
apt update && sudo apt upgrade -y
apt install -y wget curl gnupg software-properties-common build-essential git python3-pip python3-venv python3-setuptools python3-wheel git
##Add user to groups to grant permissions for GPU access
usermod -a -G video,render user
##To add all future users to the render and video groups by default, run the following commands:
echo 'ADD_EXTRA_GROUPS=1' | sudo tee -a /etc/adduser.conf
echo 'EXTRA_GROUPS=video' | sudo tee -a /etc/adduser.conf
echo 'EXTRA_GROUPS=render' | sudo tee -a /etc/adduser.conf
## or grant GPU access to all users on the system
nano /etc/udev/rules.d/70-amdgpu.rules
KERNEL=="kfd", MODE="0666"
SUBSYSTEM=="drm", KERNEL=="renderD*", MODE="0666"
udevadm control --reload-rules && sudo udevadm trigger
####Disable integrated graphics (IGP), if applicable
ROCm doesn’t currently support integrated graphics. Should your system have an AMD IGP installed, disable it in the BIOS prior to using ROCm. If the driver can enumerate the IGP, the ROCm runtime may crash the system, even if told to omit it via HIP_VISIBLE_DEVICES.
###ROCM Installation
mkdir --parents --mode=0755 /etc/apt/keyrings
wget https://repo.radeon.com/rocm/rocm.gpg.key -O - | \
gpg --dearmor | sudo tee /etc/apt/keyrings/rocm.gpg > /dev/null
curl -fsSL https://tasks.portaine# Register kernel-mode driver
echo "deb [arch=amd64 signed-by=/etc/apt/keyrings/rocm.gpg] https://repo.radeon.com/amdgpu/6.4/ubuntu jammy main" \
| sudo tee /etc/apt/sources.list.d/amdgpu.list
sudo apt update
# Register ROCm packages
echo "deb [arch=amd64 signed-by=/etc/apt/keyrings/rocm.gpg] https://repo.radeon.com/rocm/apt/6.4 noble main" \
| sudo tee --append /etc/apt/sources.list.d/rocm.list
echo -e 'Package: *\nPin: release o=repo.radeon.com\nPin-Priority: 600' \
| sudo tee /etc/apt/preferences.d/rocm-pin-600
sudo apt updater.io/install.sh | sudo bash
apt install rocm
####Post-installation instructions
tee --append /etc/ld.so.conf.d/rocm.conf <<EOF
/opt/rocm/lib
/opt/rocm/lib64
EOF
ldconfig
export PATH=$PATH:/opt/rocm-6.4.0/bin
export GFX_ARCH=gfx1100
###check
clinfo
rocminfo
rocm-smi
##GPU load
apt install nvtop
nvtop
####Installation Docker
https://docs.docker.com/engine/install/ubuntu/#prerequisites
####Installation Portainer - Key für BE Edition beantragen - 3 Nodes frei
https://docs.portainer.io/start/install/server/docker/linux
###pull Docker image
docker run -d --restart always --device /dev/kfd --device /dev/dri -v ollama:/root/.ollama -p 11434:11434 --name ollama ollama/ollama:rocm
docker run -d --network=host -v open-webui:/app/backend/data -e OLLAMA_BASE_URL=http://127.0.0.1:11434 --name open-webui --restart always ghcr.io/open-webui/open-webui:main
Honestly, kind of addicted to the awesomeness that is running AI locally. I would kill to have a server in my homelab with a really efficient Intel CPU that is poop in raw power but killer at basically idle and a graphics card with a high amount of RAM. Then I can run one of those web UI frontends to be able to access it over the web. Also, kinda sucks that my mini-pc doesn’t have a decent GPU (it has only the iGPU integrated into it’s Intel Core i7-10510U CPU).
Hey, I am trying to get these buggers working on my dev machine at work. It’s running Fedora 42 with an NVIDIA RTX Ada 3000 Laptop GPU, but I cannot get models to run on the graphics card. They keep using the CPU. The ollama serve output indicates that it recognizes the NVIDIA GPU as being available to run models. I have CUDA and the NVIDIA container toolkit installed, so what am I missing?