What is sweetspot for local llms?

Depends what you’re doing, but in my experience there’s huge progress and significant capability occurring in the 24-30b parameter size. In the past 12-18 months models this sort of size have gone from “meh” to “i can actually use this for tool calls and competent local code generation”. Not just autocomplete.

Which means 24–32GB of VRAM or so. Probably to target the upper end of consumer GPUs and lower-mid range of workstation class cards.

While you can run some of these models in say 16 GB of VRAM (just), you’ll have a very small context window.

I bought a 7900XTX recently (upgraded from 6900XT) specifically to get 24 GB to run ~30b models at a decent speed.

1 Like

There are domestic setups with up to 16 DGX Sparks.

Personally I have 4 DGX Sparks at home just for the fun of it and they are awesome. I wish more people could enjoy this type of setup.

As for the question of the topic, the sweet spot is 2x DGX Spark running Deepseek v4 Flash 0731.

1 Like

I think it really depends on what you plan on doing.

DGX Sparks are really cool, and one of the least expensive ways to get above 64GB of memory for AI workloads, but the memory bandwidth is rather anemic, at only ~273 GB/s, and as everyone knows, local inference performance is all about memory bandwidth (as long as you have enough RAM to contain the model and context window.)

I have been playing with Ollama and various models in an LXC container on the CPU (using 24 cores) on my main Proxmox node for a while now. I knew - of course - that the performance would be unimpressive, but this was just for testing purposes to see if any of the models were actually good enough for my purposes to warrant spending real money on hardware.

Honestly, running models on the CPU on this hardware does much better than I was expecting due to my my Milan Epyc platforms 8x memory channels giving it over 200GB/s memory bandwidth (~75% of a DGX spark). I was able to get ~3.3 tokens per second on Llama 3.3 70B (Q4_K_M) which while still frustratingly slow was much faster than I expected, and OK for my evaluation purposes. (submit prompt, go get a cup of coffee). What brings the CPU processing down isn’t necessarily the response output, but rather the prompt pre-fill, as a CPU can’t churn through that even remotely as fast as a GPU can.

Now you almost certainly know this already, but I am writing the next part for the benefit of OP, but the reason you get away with such slow memory bandwidth is because of the Mixture of Experts (MoE) nature of Deepseek v4, which while it requires more memory than traditional dense models, executes faster because it is selecting and sifting through the parameters of just one of the dedicated experts when its providing your output.

Based on the specs, I’m going to make an educated guess that a single DGX spark gets you ~10-14 tokens/s in DeepSeek-V4-Flash-0731, and that adding a second one ups that to about 14-18 tokens/s. Am I pretty close?

I did much soul searching before deciding what to invest in for my local inference hardware. Did I go for more VRAM and try to use a MoE model, or less VRAM and go for a traditional dense model? And at what price?

Dual RTX 3090’s

The absolute budget king is to get dual RTX 3090’s but this approach has some drawbacks. For one, even with two of them you are limited to only 48GB of VRAM, which is going to be really tight for a more capable model like Llama 3.3 70B (Q4_K_M). It can be done, but the memory gets really tight. The model weights themselves will use about 42.7GB, leaving only just over 5GB for the context window of only ~15k tokens. Other downsides are that most 3090’s require three PCIe slots each, and combined have a massive power envelope my server cannot keep up with. I also would not have enough available slots for even one, let alone two 3 slot GPU’s.

Dual RTX 4090’s

Dual 4090’s are also a great option, but they suffer from the same RAM limitation problem as the 3090’s.

Dual RTX 5090’s

This gets you to 64GB of combined RAM, but this approach is going to cost enough money that there are better ways to get to 64GB.

Single AMD Instinct MI210

In the end this is the route I decided to take. This GPU has 64GB of HBM2e VRAM with a beastly 1.6TB/s memory bandwidth that allows it to rip through models at relatively high token/s rates, provided that model fits in 64GB of VRAM.

It also takes only two PCIe slots in my server, and uses “only” 300W when at full load. The plan is to shove it in the main Proxmox node, and pass it through to a dedicated Ollama VM.

This approach does - unfortunately - not get me enough VRAM to run any of the existing MoE models, but it does allow me to tear through some of the traditional dense 70-75B models out there.

I have been going back and forth between which models to use. I eliminated all things Qwen/Deepseek mostly because this model will mostly serve me from a web assisted knowledge/research perspective, and I’d like to avoid any Chinese authoritarian censorship or bias in the model training.

I could still change my mind, but I think I will use Llama 3.3 70B (Q5_K_M) as my main model. This will use about 50GB of VRAM.

In addition to this I plan on running a smaller model that uses less VRAM assigned as a task model to offload that from the main heavy model. I haven’t decided which one yet though.

One good choice is Llama 3.2 3B which ought to be lightning fast on those task threads that don’t require the reasoning or knowledge of a heavier model, while also using very little VRAM (~2GB) and thus not cutting down to my context window so much. This would leave ~14GB for the context window which gets me about 42K tokens which is quite respectable for what I plan to do with it.

Another option is to go with a secondary task model that is a little bit larger but also has vision capability that can be triggered when the web assist digests webpages with images. I initially considered Llama 3.2 11B Vision, but at ~7.5GB for the weights it is quite large and really would hamper my context window size for the main model.

Since I’ll be using this model for both tasks and as a companion model for image processing, I am still eliminating Chinese models, as I don’t want any censorship/bias in the creation of web search terms for the web assist (which is one of the things the task model does) There are a dizzying array of options here, Gemma 4 4B, Gemma 3 4B, Phi-4-Multimodal and many many more.

One of the enormous differences I expect to see is that the MI210 will absolutely rip through prompt evaluation / prefill at 500-1200 tokens/s, meaning the actual evaluation phase where the model is providing its response will start almost instantaneously.

Once the output phase starts I expect to see between 25 and 30 tokens per second out of Llama 3.3 70B (Q5_K_M) which should be more than sufficient for my desired purposes.

All of this is - however - theoretical at this point. Once I actually get the MI210 set up and start playing with it, I’m sure large parts of this plan will change. Nothing beats actual hands on experience.

It would be nice to be able to play with 128GB or 256GB of VRAM for these large MoE models but at this point that is a little rich for my blood. I’d need something like an Nvidia A100 80GB selling for like $13k right now, or a pair of DGX Sparks for a combined $9k (if I could even get my hands on them in this market) and I’m not spending $10k on a hobby. The $4k for the MI210 is enough of a “holy shit I can’t believe I spent that much” moment. But maybe down the line either an MoE model that fits in 64GB will become available, or once I recover from the sticker shock of the MI210 I could pick up a second one and bump my capability up to the 128GB level.

For now I think my approach gets me quite good performance and quite good reasoning/logic/accuracy capability for a lot less money. (At least so I hope)

I am quite privacy oriented in my online presence, so I don’t use any frontier models that require me to sign in (instead I hide behind VPN’s, use fingerprinting resistance, and use DNS filters to eliminate as many trackers as possible.

I was casually using Gemini 3.5 Flash prior to Google taking it away from users that are not signed in. Now all I can get there without signing in is their new Flash-Lite model.

I expect the capability of Llama v3 (Q5_K_M) to land me somewhere in between the capability of Flash-Lite and Gemini 3.5 Flash, probably closer to 3.5 than to flash lite. That ought to be more than enough capability than I need for web assisted research…

…but I’ll learn more over time. I hope I didn’t spend $4k to be disappointed :sweat_smile:

2 Likes

With two Sparks in a cluster I get 46 tok/s with a single session in Deepseek v4 Flash 0731 NVFP4 with a fully functional 1M context and 1100 tok/s pre-fill.

It is perfectly fine for any task, or in my case just have it drive Hermes.

2 Likes

That is quite a lot better than my guesstimate based on memory bandwidths. I wonder what accounts for my guess being so off.

Have you ever benchmarked it with just the one? What kind of performance did you see then?

I do not know what I am doing: I have spent $3500.00 and have reached 64gb of VRAM,128gb of ddr4 RAM and run 45 gb models at 9.2 t/s.. MoE run way faster. RTX 5060 TI 16gb are not gaming GPU’s but they run AI very well. IMHO. I refuse to buy any more because the prices are climbing out of my reach. $450.00 each and now bumping against $600.00. llama.cpp because ollama doesn’t like multiple GPU’s.

1 Like

I’d say that in 2026, llama 3.3 is outperformed by many models half its size…

What are you using 3.3 for? because i did use it 18 months ago for things but ended up moving to smaller 20-30b class models like qwen, gemini, gpt-oss which seem more capable in far smaller size for what i’ve been doing?

1 Like

Speculative decoding, such as DSpark baked into DSv4 Flash gives a good boost.

The best thing with running one Spark is that you are halfway to getting two sparks. :wink: After all you have spent 1500 on a ConnectX-7 NIC that isn’t used.

There are a lot of very smart people optimising the Sparks, as you can see over here: DGX Spark / GB10 Projects - NVIDIA Developer Forums

You can also look at how most models are performing on any size Spark cluster here: https://spark-arena.com/

1 Like

I haven’t seen this mentioned in the thread because it’s mostly hardware focused but what I’ve been recently experiencing is that using an harness can significantly improve the performance of even small LLMs quite a bit.

Maybe this is nothing new for most of you here, but I thought it was worth mentioning because spending some time in customizing a good harness is time well spent.

I’ve been dipping my toes into Hermes combined with Honcho as it’s memory and I’ve been impressed.
Gemma4 12B, for me, had the tendency of going in circles after a while (with a 256K context window) when debugging code/software issues. Using this harness this behaviour totally stopped, which is great!

2 Likes

Well, I’m still learning.

I don’t code, and have no desire to get into so-called “vibe coding”, so my use case is a little different than what I presume many use it for. My primary use is for web-assisted knowledge gathering and research. Essentially find answers to questions that could take me hours of googling and reading in just 30 seconds.

For this type of use case I understand there are some benefits from using a high capability dense model over an MoE model simply because you benefit from the entire weight of the model in every response, thus minimizing any blind spots.

I’ll likely test different models, but as mentioned, for my purposes I want to avoid any potential bias/censorship in the training of the models, so I am intentionally staying away from anything that comes out of China, like Qwen or DeepSeek.

My understanding is that once you limit yourself to non-authoritarian state models that fit within a 64GB footprint, llama 3.3 70B is still king, especially at the 5bit level, but I will undoubtedly continue to learn and reassess as things go on.

Same could be said for adding a second MI-210 :sweat_smile:

1 Like

Well, thanks to you I have done some more reading up on this, and have learned a lot in the last hour.

I think it has changed my mind to start with Gemma 4 instead.

I’m currently torn between the dense Gemma-4:31b or the MoE Gemma 4-26b.

According to the Hugging Face leader board both outperform Llama 3.3 70B, and both are neck at neck with eachother (1441 vs 1435 Elo points)

So I guess the dense model is technically marginally better, but probably not really in a noticeable way in day to day use, and the MoE model ought to be much much faster…

I’m starting to think I wasted a lot of money on a very capable GPU I don’t need :sweat_smile:

2 Likes

The differences are quite stark…

I did a benchmark of three models, again on my server CPU while awaiting the GPU.

Here was my benchmark prompt:

"Answer the following 4 questions. You must strictly follow these formatting rules:
- Number your answers from 1 to 4.
- Do not include any introductory or concluding text. Just provide the numbered list.

1. Spatial & Logic: A small marble is put into a ceramic coffee mug, and the mug is placed upside down on a wooden dining table. Someone then picks up the mug carefully, carries it across the room, and places it inside the microwave. Where is the marble right now, and why?

2. Multi-step Math & Letters: Take the number of letters in the capital of Australia, multiply it by 4, subtract 6, and then add the number of vowels in the word “ENTHUSIASTIC”. What is the final integer?

3. Temporal/Context Trick: A boy runs down the stairs on a bright December morning, looks into his living room, and sees a decorated pine tree with wrapped boxes underneath it. What specific day of the year is it, and what clues in the text tell you this?

4. Coding/Formatting: Write a python one-liner function called is_palindrome that checks if a string is a palindrome, ignoring spaces and capitalization. Output only the raw python code inside a standard markdown code block."

- Llama 3.3 70B

Concise answers. Ran output at 3.3 tokens/s, but got the first two questions wrong! It claimed Canberra has 7 letters (it has 8: C-A-N-B-E-R-R-A), botched the arithmetic, and for Q1, it claimed the marble “would have fallen out due to gravity” during the pickup - which physically wouldn’t happen if the mug was kept upside down against the table while lifting.

Output
  1. The marble is on the wooden dining table because when the mug was picked up carefully from the table, the marble would have fallen out due to gravity since the mug was upside down.

  2. The capital of Australia is Canberra, which has 7 letters. Multiplying by 4 gives 28, subtracting 6 results in 22. The word “ENTHUSIASTIC” has 5 vowels (E, U, A, I). Adding this to 22 gives a final integer of 27.

  3. It is December 25th, Christmas Day. The clues are the decorated pine tree and wrapped boxes underneath it, which are traditional Christmas decorations, indicating that it is Christmas morning.

  4. 
    

def is_palindrome(s): return s.replace(" “, “”).lower() == s.replace(” ", “”).lower()[::-1]

total duration: 55.49757972s
load duration: 214.372397ms
prompt eval count: 270 token(s)
prompt eval duration: 308.503ms
prompt eval rate: 875.19 tokens/s
eval count: 177 token(s)
eval duration: 54.972064s
eval rate: 3.22 tokens/s

Llama 3.2 3B

I threw this one in here just for comparison sake of how a small fast model might do.

It failed miserably on the actual logic. It fell for the object-permanence trap in Q1 (claiming the marble went into the microwave), butchered the vowel count in Q2 (claiming “ENTHUSIASTIC” has 11 vowels, when it has 5), and got the math wrong (15).

But at least it was fast.

Output
  1. The marble is inside the microwave because when the mug was placed upside down on the table, it rested against the bottom of the mug, and when picked up, its position relative to the surface didn’t change.

  2. The capital of Australia is ‘CANBERRA’, so the number of letters in ‘CANBERRA’ is 8. Multiplying this by 4 gives 32; subtracting 6 results in 26; the word “ENTHUSIASTIC” has 11 vowels, making the final integer 15.

  3. The boy sees a decorated pine tree with wrapped boxes underneath it on Christmas Day, December 25th, because only on this day is it common to decorate pine trees and wrap gifts for Christmas.

  4. def is_palindrome(s): return s.replace(' ', '').lower() == s.replace(' ', '').lower()[::-1];

total duration: 55.49757972s
load duration: 214.372397ms
prompt eval count: 270 token(s)
prompt eval duration: 308.503ms
prompt eval rate: 875.19 tokens/s
eval count: 177 token(s)
eval duration: 54.972064s
eval rate: 3.22 tokens/s

Gemma 4 26B (MoE)

Gemma is very chatty. Not only does it output its internal monologue while thinking through a problem, but it is also much more conversational than the Llama models which are concise to the point of a flaw.

The upside is, Gemma is the only of the three models I tested that passed all of my tests.

Output

Thinking…

  • Constraint 1: Answer 4 specific questions.

    • Constraint 2: Number answers from 1 to 4.

    • Constraint 3: No introductory or concluding text (just the numbered list).

    • Question 1 (Spatial & Logic): A small marble is put into a ceramic coffee mug, and the mug is placed upside down on a wooden dining table. Someone then picks up the mug carefully, carries it across the room, and places
      it inside the microwave. Where is the marble right now, and why?

      • Analysis: The mug is upside down on the table. When a mug is upside down on a surface, anything inside it (the marble) is resting on that surface (the table), trapped by the rim of the mug. When the mug is lifted
        “carefully” and moved to the microwave, the marble remains on the table because it wasn’t attached to the mug; it was simply under it.
      • Answer: The marble is on the wooden dining table. Because the mug was upside down on the table, the marble was resting on the table surface, not attached to the mug itself.
    • Question 2 (Multi-step Math & Letters): Take the number of letters in the capital of Australia, multiply it by 4, subtract 6, and then add the number of vowels in the word “ENTHUSIASTIC”. What is the final integer?

      • Step 1: Capital of Australia = Canberra.
      • Step 2: Number of letters in “Canberra” = 8.
      • Step 3: 8 * 4 = 32.
      • Step 4: 32 - 6 = 26.
      • Step 5: Vowels in “ENTHUSIASTIC” (E, U, I, A, I) → E, U, I, A, I. Let’s count carefully: E(1), N, T, H, U(2), S, I(3), A(4), S, T, I(5), C. Total vowels = 5.
      • Step 6: 26 + 5 = 31.
      • Answer: 31.
    • Question 3 (Temporal/Context Trick): A boy runs down the stairs on a bright December morning, looks into his living room, and sees a decorated pine tree with wrapped boxes underneath it. What specific day of the year is
      it, and what clues in the text tell you this?

      • Analysis: The presence of a “decorated pine tree” and “wrapped boxes” in “December” strongly implies Christmas Day (or the holiday season). However, the prompt asks for a specific day. Is there enough info to be 100%
        sure it’s Dec 25? Usually, these riddles point to Christmas. The clues are: December morning, decorated pine tree, wrapped boxes.
      • Answer: It is likely Christmas Day (December 25th). The clues are the month being December, the presence of a decorated pine tree, and the wrapped boxes underneath it.
    • Question 4 (Coding/Formatting): Write a python one-liner function called is_palindrome that checks if a string is a palindrome, ignoring spaces and capitalization. Output only the raw python code inside a standard
      markdown code block.

      • Requirement: One-liner function.

      • Logic: Remove spaces, lowercase everything, check if string == reverse(string).

      • Draft 1: is_palindrome = lambda s: (s := s.replace(" ", "").lower()) == s[::-1] → This uses assignment expression (walrus operator), which is one-liner compatible in Python 3.8+.

      • Draft 2: is_palindrome = lambda s: [s_clean := s.replace(' ', '').lower()][0] == s_clean[::-1] (A bit messy).

      • Draft 3: is_palindrome = lambda s: all(c == r for c, r in zip(s.replace(' ','').lower(), s.replace(' ','').lower()[::-1])) (Too long).

      • Actually, a standard one-line function definition using def is often expected if not specified as a lambda. But “one-liner function” usually implies lambda or a single line of code. Let’s use the walrus operator for
        simplicity and elegance in a single line. Or just:
        is_palindrome = lambda s: s.replace(' ', '').lower() == s.replace(' ', '').lower()[::-1]

      • Check requirements: “Write a python one-liner function called is_palindrome… Output only the raw python code inside a standard markdown code block.”

      • Code:

        is_palindrome = lambda s: (s := s.replace(" ", "").lower()) == s[::-1]
        
    1. The marble is on the wooden dining table because when the mug was placed upside down, the marble was resting on the table surface, not attached to the mug.
    2. 31
    3. Christmas Day (December 25th); clues include the month being December, the decorated pine tree, and the wrapped boxes.
    4. is_palindrome = lambda s: (s := s.replace(" ", "").lower()) == s[::-1]
      

    Refining Answer 1: To be more precise: The marble is on the table. When an upside-down mug is placed over an object, that object remains on the surface. Moving the mug does not move the marble unless specified.

    Refining Answer 2:
    Canberra = 8 letters.
    8 * 4 = 32.
    32 - 6 = 26.
    ENTHUSIASTIC vowels: E, U, I, A, I (5).
    26 + 5 = 31.

    Refining Answer 3: Is there any other day? It could be any day in December, but usually “decorated tree and wrapped boxes” implies Christmas. The prompt asks for a specific day. I will state Christmas Day.

    Refining Answer 4:
    Let’s check if is_palindrome = lambda s: s.replace(' ', '').lower() == s.replace(' ', '').lower()[::-1] is safer (no walrus operator dependency for older versions, though not strictly necessary unless specified). Actually, the
    walrus one is cleaner for a “one-liner” concept in modern Python.
    Wait, is_palindrome = lambda s: (temp := s.replace(' ', '').lower()) == temp[::-1] also works.
    Let’s use the simplest version that doesn’t rely on assignment expressions just in case:
    is_palindrome = lambda s: s.replace(' ', '').lower() == s.replace(' ', '').lower()[::-1]

    Wait, looking at Answer 4 again. “Output only the raw python code inside a standard markdown code block.” This means I shouldn’t include text inside the block other than the code. But the instruction for the whole response is
    “Do not include any introductory or concluding text”.

    Let’s review all constraints.

    1. Numbered 1-4.
    2. No intro/outro.
    3. Answer 1: Location and why.
    4. Answer 2: Integer.
    5. Answer 3: Day and clues.
    6. Answer 4: Python code in markdown block.

    Self-Correction on Q1: “The marble is on the wooden dining table because when the mug was placed upside down, it was resting on the surface of the table, not inside the mug’s cavity.” → Accurate.

    Self-Correction on Q2:
    Canberra (8) * 4 = 32.
    32 - 6 = 26.
    E N T H U S I A S T I C
    Vowels: E, U, I, A, I. Total: 5.
    26 + 5 = 31.

    Self-Correction on Q3: It’s December, tree is decorated, boxes are wrapped. This is Christmas Day.

    Final check of Q4: is_palindrome = lambda s: s.replace(' ', '').lower() == s.replace(' ', '').lower()[::-1]
    This covers spaces and capitalization.

    • 1 to 4 numbering? Yes.
    • No intro/outro? Yes.
    • All questions answered? Yes.
    • Q2 calculation correct? Yes (31).
    • Q3 clues included? Yes.
    • Q4 code format correct? Yes.

    One detail: The prompt asks for “the final integer” in Q2. 31 is an integer.
    Q3 asks “What specific day… and what clues”. I’ll provide both parts.

    Let’s look at Q1 again. “Where is the marble right now, and why?”.
    The marble is on the table. Why? Because when the mug was placed upside down, the marble was trapped under the rim on the table surface. Lifting the mug doesn’t lift the marble.

    One last check on Python code:
    is_palindrome = lambda s: s.replace(' ', '').lower() == s.replace(' ', '').lower()[::-1]
    Does it ignore spaces? Yes (replace(' ', '')).
    Does it ignore capitalization? Yes (lower()).
    Is it a one-liner? Yes.

    Wait, should I provide the def version?
    def is_palindrome(s): return s.replace(' ', '').lower() == s.replace(' ', '').lower()[::-1]
    The instructions say “Write a python one-liner function”. A lambda expression is technically an anonymous function, but often used for this purpose. A single-line def is also a function. The lambda version is more
    traditionally called a “one-liner”. I’ll go with the lambda as it’s most concise.

    Wait, let me double check “ENTHUSIASTIC” vowels.
    E - 1
    N
    T
    H
    U - 2
    S
    I - 3
    A - 4
    S
    T
    I - 5
    C
    Total = 5. Correct.

    Let me double check “Canberra” letters.
    C-A-N-B-E-R-R-A
    1-2-3-4-5-6-7-8. Correct.

    Let’s re-verify the logic for Q1.
    Mug upside down on table → Marble is between the table and the mug rim.
    Someone picks up the mug carefully… moves it to microwave.
    If they pick it up “carefully” (meaning without tilting or lifting the surface), the marble stays behind.
    The text says: “picks up the mug… carries it… places it inside the microwave”. It does not say the person picked up the marble or used a plate/card to slide it out. Therefore, the marble remains on the table.

    Final Plan for formatting:

    1. [Answer]
    2. [Answer]
    3. [Answer]
    4. [Code]
      

    Wait, I need to make sure I don’t include any text outside of the numbered list elements for questions 1-3 and the code block for 4.

    Let’s refine Q3 answer phrasing: “Christmas Day; the clues are the month of December, the decorated pine tree, and the wrapped boxes.”

    One small check on Python one-liner definition.
    If I use is_palindrome = lambda s: s.replace(' ','').lower() == s.replace(' ','').lower()[::-1], it is a single line that defines the function. This fits “one-liner function”.

    Final check of Q3 clues: “A boy runs down the stairs on a bright December morning, looks into his living room, and sees a decorated pine tree with wrapped boxes underneath it.”
    Clues: “December”, “decorated pine tree”, “wrapped boxes”.

    Let’s ensure Q1 logic is tight. If you put an upside-down mug over something, the something is on the table. When you lift the mug, the thing stays on the table. Correct.

    Check Q2 math one last time.
    Canberra = 8 letters.
    8 * 4 = 32.
    32 - 6 = 26.
    Vowels in ENTHUSIASTIC: E, U, I, A, I (5).
    26 + 5 = 31. Correct.

    Check Q4 requirements again: “Output only the raw python code inside a standard markdown code block.”
    This implies for question 4, the answer is just the code block.

    Final check of all constraints. No intro/outro. Just numbered list.

    One possible pitfall in Q1: Does “picks up the mug” imply picking up the marble? Usually, no. In physics/logic puzzles, if an object is under a container and you lift the container, the object remains.

    Wait, looking at Q2 again: “Take the number of letters in the capital of Australia…”.
    Australia Capital = Canberra (8). Correct.

    Looking at Q3: Is there any other day? It could be New Year’s Eve or Christmas Eve, but “decorated tree and wrapped boxes” is the textbook definition used for “Christmas Day” in logic puzzles. I will state Christmas Day.

    Wait, one detail in Q1: “…the mug is placed upside down on a wooden dining table.” The marble was put into the mug before it was placed upside down. So the marble was inside the mug. Then the mug was flipped. Now the
    marble is trapped between the table and the mug’s bottom (which is now the top). When you lift the mug, the marble stays on the table.

    Final verification of Python one-liner:
    is_palindrome = lambda s: s.replace(' ','').lower() == s.replace(' ', '').lower()[::-1]
    It handles “A man a plan a canal Panama” (spaces/caps).
    Input: “A man a plan a canal Panama”
    Cleaned: “amanaplanacanalpanama”
    Reversed: “amanaplanacanalpanama”
    Result: True.

    One last look at formatting instructions.

    • Number your answers from 1 to 4.
    • Do not include any introductory or concluding text. Just provide the numbered list.

    Everything seems solid.

    Wait, one tiny check on Q2 “ENTHUSIASTIC”.
    E (vowel)
    N (consonant)
    T (consonant)
    H (consonant)
    U (vowel)
    S (consonant)
    I (vowel)
    A (vowel)
    S (consonant)
    T (consonant)
    I (vowel)
    C (consonant)
    Vowels: E, U, I, A, I. Count = 5. Correct.

    Wait, Q1 “Where is the marble right now, and why?”.
    Answer: The marble is on the wooden dining table because when the mug was placed upside down, the marble was resting on the table surface under the rim, and lifting the mug does not move the marble.

    Final assembly check of the response structure:

    1. [Text]
    2. [Text]
    3. [Text]
    4. [Code]
      

    (Self-correction: Ensure no extra spaces or anything in the final output.)
    …done thinking.

  1. The marble is on the wooden dining table because when the mug was placed upside down, the marble was resting on the table surface under the rim; lifting the mug does not move the marble.
  2. 31
  3. Christmas Day; the clues are the month of December, the presence of a decorated pine tree, and the wrapped boxes underneath it.
  4. 
    

is_palindrome = lambda s: s.replace(’ ‘, ‘’).lower() == s.replace(’ ', ‘’).lower()[::-1]

total duration: 8.409432944s
load duration: 212.137181ms
prompt eval count: 285 token(s)
prompt eval duration: 2.195083s
prompt eval rate: 129.84 tokens/s
eval count: 185 token(s)
eval duration: 5.959569s
eval rate: 31.04 tokens/s

Because it is an MoE model it also runs almost as fast as the 3B model, from a tokens/s perspective, but the downside is that its chattiness and internal monologue took a response that was 875 tokens on Llama 3.3 70B and turned it into 3,650 tokens, resulting in its actual total duration in the prompt being the longest of the three, more than doubling over that of Llama 3.3 70B.

Conclusion

My conclusion is that accuracy/logic/factuality is my largest priority, and because of this I’m switching to Gemma 4 26b. I also like the ability to peer behind the curtain, so I don’t mind the chattiness and internal monologue.

And yes, knowing what I know now I definitely overspent on a GPU I don’t really need, but I guess what is done is done. At least it will be VERY fast (I’m guesstimating over 250 tokens/s on the MI210), and I’ll be able to support multiple users in my house easily.

Interestingly, on this server platform with 8 channels of DDR4-3200, the Gemma 4 26b MoE model is ALMOST usable on the CPU. 31 tokens /second is quite good for output. When testing on Olla a straight from the terminal it runs quite well, but this all falls apart when using a frontend like OpenWebUI which injects lots of data into the prompts causing them to balloon in size. This causes running them on the CPU to fall apart as it struggles when it comes to pre-fill compared to even a low end GPU.

1 Like

Agreed! The harness can make a huge difference!

1 Like

I wouldn’t say you necessarily over-spent on GPU - but you can certainly make use of even better models than Llama in the space llama requires!

Also with models like Gemma4 family and Qwen you have plenty of leftover capacity for context. Or multiple models held in VRAM at the same time, potentially.

Just read up some more on your use case.

One thing the more recent models can do well that llama struggles with is accessing the internet for up to date information and using tools to calculate things or access data.

If you’re just getting started, I’d suggest checking out AnythingLLM as it includes the ability to expose your LLM(s) to tools, search, rag, etc. fairly easily.

3 Likes

You can also try the following models -

  1. Qwen 3.6 35B A3B (MoE)
  2. Qwen 3.6 27B (dense)
  3. Gemma4 12B (dense)
2 Likes

I am kinda new to local AI as well. I am using all 3 of my gpu’s over my network. A 2070 8gb, a 4060 8gb and my best card and only new one a 9060 xt with 16gb of vram. I run a bunch of little models on them, basically just to learn and compare CUDA vs ROCm. It’s been super fun and I’m learning a lot and it really ties into my Cybersecurity studies. When eventually I have a bit of money I’ll splurge on a more capable rig but the great thing is that all the skills you learn on these smaller models are instantly transferable once you upgrade. It’s all a lot of fun. Use what you have, learn and maybe one day you can upgrade your equipment for whatever your workload requires rather than just getting a monster and not knowing what to do with it. :grin: Keep at it!

2 Likes

If you have a recent AMD card (Navi3/4) i’d suggest running them under a recent linux with Vulkan, seems to perform better in my experience.

Ollama/llamacpp on CachyOS with vulkan is pretty painless.

1 Like

For my own selfish reasons I kind of hope someone builds a modern MoE model along the lines of Gemma 4 but targeting the 60-75B weight footprint, so I can make the most out of my new GPU :stuck_out_tongue:

Right now everything is rather tiny (~30B) or too big (~120B).

I will have to look into that.

Thus far on my CPU test setup I have been using OpenWebUI and it seems pretty decent. I have gotten it to the point where it can search via my custom SearXNG install, and then fetch and ingest the text from the URL’s it finds (though many crash head first into scraping protection :sad_but_relieved_face:)

Once I get this up and running on my new GPU, my plan is to look into how I canset up a better scraper, and also maybe since the model has vision capability, maybe also ingest images contained on said urls and analyze them as well.

Would you say AnythingLLM is more capable or easier to use than OpenWebUI?

1 Like

With your 64GB card you can either run the model in BF16 or with a decently sized context so I wouldn’t say it is wasted.

vLLM + LibreChat is a good combination, sort of depends what your use case is.

3 Likes