The RTX5090 is running in ubuntu 26.04 to have as much memory available as possible.
Then i have this script for running vllm with qwen3.8 27b cyankiwi. I am serving the modelname qwen3.6-27b so that the agents use LiteLLM config lookups based on that name.
thanks to @thr3e for testing different quants Qwen 3.8 Quant Selection Guide for RTX 5090
I’m using openhands in cli mode to work on my code. But the issue is that the qwen3.8 model defaults to xHigh, which is not always needed. I can’t really set effort based on task done, but i can use it for subagents like reading files or searching by using a different model.
I had an agent create a new model name which is just the same but a different effort of “medium”.
Because this gets overwritten on an update, I had the agent create a markdownfile to update this as well.
The only other thing to do is, run openhands, set the network location of the pc serving the model and we are go!