I tried to run a ComfyUI Docker container, but the container couldn’t find my GPU.
My Troubleshooting Journey
-
I started with
docker run --gpus all. This gave me aCDIerror because on my system,
/dev/dri/card0is a symbolic link (l) that points to a physical device node (c), which the container runtime can’t use directly.
/dev/dri/card1is the physical device and its present and detected and used. -
I tried to fix this with a
udevrule, but it didn’t work. The problem was that other system rules had higher priority, so they kept creating the symbolic link before my rule could take effect. -
Then I tried using the
--deviceflag to manually pass my GPU’s real path (/dev/dri/card1). This worked for a simplehello-worldcontainer, but it still failed with ComfyUI because I was only providing the display device, not the CUDA-specific devices and libraries needed for deep learning. -
** I tried to combine both methods** (
--gpus alland--device), which led me back to the sameCDIerror. The two flags were conflicting, and the automatic method was failing before my manual one could take effect. -
Then I tried a lot of stuff the AI suggested. Verifying drivers for my card but I already have
580.76.05installed, it prompted me to install thenvidia-container-toolkitwith yay which I then reverted to the pacman installer, as both had the same version. -
I edited the
/etc/nvidia-container-runtime/config.tomlmultiple times with various parameters but to no avail. -
I tried logging the output into a file and search dmesg or journalctl in hope of finding any more clues about what happens here
journalctl -b | grep -i "nvidia"
Aug 21 23:58:11 aldr-am5 sudo[33858]: aldr : TTY=pts/0 ; PWD=/home/aldr/comfyui/git/comfyui-docker ; USER=root ; COMMAND=/usr/bin/nano /etc/nvidia-container-runtime/config.toml
Aug 21 23:59:36 aldr-am5 conmon[34087]: conmon f8b0450fbfeeb3b418ae <nwarn>: runtime stderr: error executing hook `/usr/bin/nvidia-container-runtime-hook` (exit code: 1)
cat /tmp/nvidia-container.log
time="2025-08-21T23:59:36+02:00" level=debug msg="Locating \"docker-runc\" in [/usr/local/sbin /usr/local/bin /usr/sbin /usr/bin /sbin /bin]"
time="2025-08-21T23:59:36+02:00" level=debug msg="Locating \"runc\" in [/usr/local/sbin /usr/local/bin /usr/sbin /usr/bin /sbin /bin]"
time="2025-08-21T23:59:36+02:00" level=debug msg="Locating \"crun\" in [/usr/local/sbin /usr/local/bin /usr/sbin /usr/bin /sbin /bin]"
time="2025-08-21T23:59:36+02:00" level=debug msg="Checking candidate '/usr/sbin/crun'"
time="2025-08-21T23:59:36+02:00" level=debug msg="Found 1 candidates; ignoring further candidates"
-
I attempted to run a canonical container image,
nvidia/cuda:12.4.1-runtime-ubuntu22.04, which also failed with the sameexit code: 1error. -
Asked the AI to summarize everything to me
The
nvidia-smioutput you initially provided was misleading. A real driver on a real system with version580.76.05is not compatible with either CUDA 12.4.1 or the fictional CUDA 13.0. Thenvidia-container-runtime-hookis a low-level check that correctly identified this incompatibility and failed to prevent the container from starting.
So I’m out of ideas what to test or if there is anything left for me to look into. I open for suggestions ![]()




