Multiple 9950X's and slurm?

Hi all,

I’m an end-user for HPCs. I do scientific calculations, which I first try at home on my itsy bitsy 9950X3D+96GB ram machine, and then I send it out to HPC’s of a few universities.

  • My calculations are embarassingly parallel, i.e. I can directly trade off time vs number of cores I use (once I’m past, say, 50 cores).
  • I need lots of RAM per core. Usually, university HPCs have no shortage of cores but are limited in memory per core.
  • I don’t need a lot of storage for the work itself. I’d need maybe a GB per job. (It’s like the computer is working for eons and the answer is 42, storeable in a single unsigned int.) Once the work is done and the computer is idling, output can easily be transferred to a dedicated storage solution.
  • GPU computing doesn’t help me, so the GPU can be anything. (I’m not working on bajillion small vectors; I’m inverting ~50M by ~50M sparse matrices with ~3-4% sparsity. I wish I had an easy job such as working on AI models.)

So… How stupid is this idea:

I buy some 9950X desktop computers with lots of desktop ram each. I use tricks like using a PCIE card for more RAM, as Wendell described in one of his videos before. (I need more RAM more than I need fast RAM.) I buy a crappy box to run slurm on, to control all. I employ a sla…I mean hire and train a student for the basic administration of these machines.

Please tell me how many ways this is not a good idea before I ruin the life and existence of a perfectly sane student.

Thanks!

1 Like

Why not just use HPC if you have access to them?

I’ve known of a few university research groups that do/did similar things. Building clusters with consumer hardware, even with hand me down infiniband cards from older clusters, etc.

I‘m very sceptical of the PCIe/cxl memory though as most scientific calculations (and certainly inverting large matrices) are usually memory bound. Also, cxl is not supported on consumer hardware any way (it’s even off label on threadripper, only epyc is officially supported).

Nowadays 256GB ddr5 kits exist, that could be the way to go? If that’s not enough ram, refurbished ddr4 server or workstation hardware. There’s dealers around who sell older ddr4 dual Xeon or threadripper workstations for example.

3 Likes

“Why not just use HPC if you have access to them?”

Their memory available per core is usually lower than what I need, for one specific type of work I’m doing. :frowning:

Thank you for the comments on the memory! I’ll look into that direction.


Perhaps too much info:

I’m either inverting or calculating the Pfaffian of these matrices. Fortunately, for most of my work, I don’t need a full inversion so fast/efficient algorithms exist for those. For calculating Pfaffians (of skew-symmetric matrices) though… Fastest algorith is by Michael Wimmer (arXiv: 1102.3440), and a numpy array implementation is by Bas Nijholt. Numpy arrays work through cython, so essentially they are C arrays whose memory positions are “fixed”, so they are fast to work with. (…and I don’t know much more about them.)

That implementation, in my case, ends up taking ~100GB memory per result. Each result is a pixel in a 100x100 plot. So even if a single run is fast, 256GB memory limit means I could at most get three pixels in my plot at a single time. Each pixel takes about half a day or more to calculate.

I’ve recently modified that implementation to use sparse matrices. I’ve achieved roughly 20x improvement in memory usage at the cost of 3x-5x slower implementation. I intend to put it on github at some point. next thing to do would be to implement just-in-time compilation for speed, but the easy to use module for jit (numba) doesn’t work with sparse matrices out of the box. It could still be done, but at some point I need to remember that I’m a physicist, not a computer engineer, and do the research I’m paid to do rather than dive into another rabbit hole.

So, I want to have my cake and eat it too. I don’t want to be limited by memory, and I don’t want to not work with numpy arrays either. I was wondering if I could do this on the cheap.

Maybe I can… Maybe refurbished ddr4 server or workstation hardware is the way to go.

Thanks again!

The requirements here are cryptic but, for 100 GB/core with dense matrices, I’d start from

  • 9955WX + TRA5256G56O438O
  • 9175F + TR5128G60Q456-12 and leave a core unused

Turn the CPU and RDIMM selections up or down based on the actual GB/core needed but 24 RDIMM boards probably aren’t worth it given 2DPC speed constraints. In two channel tops is 9600X at 42 GB/core. ¯\_(ツ)_/¯

No. Probably less expensive to ditch Python for a compiled language than to throw hardware at this, though, and the sparse savings from avoiding RDIMMs probably pay for ~6x AM5 cores .

3 Likes

Thank you! I’ll have to look into all of these before I write a proposal and a budget.

I’d be supporting at least one student for this work, if a switch to a compiled language is necessary. The most mature and commonly used module for our work is written in python, and if such work is necessary, I’d be needing a physics student who does their main research, and on the side devise and code new-ish algorithms for this computation in C and interface them with python. All on a student income. This quickly starts getting into labor abuse territory.

Well, it’s my problem to solve. Thank you for the inputs.

1 Like

When you BOM it out include the space to put all the kit in. This doesn’t seem to be a particularly power dense situation, partly because it’s CPU bound and not GPU accelerated, but check on the mains and HVAC arrangements as well. You wouldn’t be the first person to find they need to budget for additional circuits and reconfiguring room airflow for a stack of machines.

Systems integrators typically mark up hardware by 30-50%, so with the apparent scale of spend here it’ll probably pay for itself pretty quick to buy and assemble parts. As you’ll see in the threads here, that’s a good bit more turnkey with desktop hardware than than workstation or server. But for desktop to go it sounds like code’s needed that shards Pfaffians into fragments that are more desktop compatible.

From what you’ve said so far I’d be kind of surprised if hardware only approaches end up looking like they’ll yield minimum TCO.

Or you could probably just… not fuck over the prospective graduate student here by putting a realistic project structure in the grant narrative and deliverables. Though that might lead to a kind of mostly not completely awful graduate school experience which fails to contribute to the ~50% PhD dropout rate. Or something. ¯\_(ツ)_/¯

Either way, sounds like the spend here is well into US$ hundreds of thousands. Buying hardware installs a capability that’ll probably age out in a few years after producing a few papers. Better code seems like it’d drive down implementation costs for everybody as well as yield methods papers. So one question I’d look at in putting together a proposal is the extent to which algorithmic improvement might result in coding labor paying for itself in hardware savings.

In my PhD I needed to implement ~50,000x speedups just to be able to get basic research results because I was in a similar situation where the mature modules couldn’t scale to what needed to be done. Half the committee knew enough about code to understand what that meant and half didn’t. I had the votes for a unanimous pass because I spent a fair amount of time talking with one of the committee members who didn’t know.

3 Likes

Thank you. This project is currently in the “kicking around ideas” territory. It seems there is a possibility of this ballooning into a much larger project if what you say ends up being the case for this. The code rewrite becomes comparatively feasible in this case.

…or, this part of the project is dumped.

Either way, I didn’t know how to start thinking about this. Now I have some leads which I can look into.

Thanks!

1 Like

If it’s like my PhD, could just be a tipping point where inadequacies of existing code are no longer avoidable if you’re going to produce useful results in the area of study. FWIW, a lot of the research in the discipline I landed in was, and is, pretty much broken. Most publications are closed source and degree timelines were (and remain) too short for graduate students to develop the needed programming skills. The results are mainly junky papers from students struggling to cobble together something that somehow lets them graduate (yo) and papers from experienced 10+ person teams at research institutes with long term stable support that have good methods and good results but aren’t replicable due to the cost of recreating the code and datasets.

Sounds like you’re in a country taking more the former approach to funding science rather than the latter.

Where I did my graduate work the government was unaware of the mismatch between dispensing short term funding (three years was the max award with most grants being closer to a year) and the amount and type of work needed to reach the funding objectives. So it was failing so hard at creating a workforce that could do the desired work it couldn’t even see postdoc funding difficulties and time limits were consistently forcing out the few people with the tenacity to develop the necessary skills and not switch to other work with better career prospects. Institutionally, there tended to be a couple non-tenure track roles that could get around this but finding and sustaining qualified individuals in them was maybe a 1 in 100 jobs thing.

So good luck. Seems like you might end up needing a lot of it.

4 Likes

Interesting stuff!

You could just use fewer cores per node on HPC? But in any case having at least some local machine for development is useful (depending how high contention/waiting times are on the cluster).

That’s not really true – numpy calls compiled C code to do the actual computation. Cython is an extension for python that lets you compile functions that are written in python when you add C types to variables. In both cases you’re using compiled C but the method of generating that is quite different.

How did you implement the sparse matrices – using scipy? I don’t think jit gives you any gains when you’re already doing the actual compute in scipy or numpy as those calls are to already compiled code.

In case it’s helpful, jax is an interesting library for HPC with python and it seems to have experimental sparse support:

https://docs.jax.dev/en/latest/jax.experimental.sparse.html

Fully agree. Screwing over the student with an unfeasible project is a lose-lose; you wasted the funding and the student wasted their time. But to a motivated student with a knack for programming this sounds like a good project. IMO (as someone who did computational work for their PhD) have clear goals/deliverables and a couple nice-to-haves on the side if there’s time, motivation or talent in surplus.

2 Likes

I agree. I mean, I’d like to think I fit into this category. But I’d also describe my PhD as partially salvaging a lose-lose-lose situation. Since the (lack of) project planning was a total disaster no funding objectives were met (multiple chapters of my thesis ended up showing it was impossible to meet them without investing tens of millions on further data collection and software development), I’m still kind of surprised I actually graduated instead of quitting (I had to fill out a government questionnaire after graduating and IIRC my answer about my program’s strengths was something to the effect of hell if I know), and I had the graduate advisor, department head, and a dean stacked up against my PI and was referring committee members to the department head to discuss their (IMO entirely reasonable) concerns over thesis content and which journals papers were going to.

1 Like

why is that sounding too familiar…

FWIW, a lot of the research in the discipline I landed in was, and is, pretty much broken.

Same here, though not (always) for lack of good intentions. All sorts of constraints makes one push out research, which is interesting and correct, but doesn’t solve the advertised problem. And that’s OK, sorta, because we as a humanity are managing finite resources. Things get cut. Are we managing it well, though?.. Eh.

Sounds like you’re in a country taking more the former approach to funding science rather than the latter.

Because of the country I’ll be in in a few short moths, time is slightly less of a factor, but budgets are a big factor. I need to be clever about this.

2 Likes

I have seen similar things. One factor has also been that research funding agencies are not too keen on funding ‘software development’ as they don’t consider this ‘research’. There’s pros & cons to this; students don’t get trapped in what’s basically an underpaid software developer job. But it’s not good to develop maintainable software over time. Code is usually written by previous students/post-docs and the supervisor/PI doesn’t have to overhead to keep on top of maintaining software on top of teaching, supervising, and writing papers.

Luckily scientific software is moving more to open-source which could alleviate some of these factors.

Computational/software work in research really needs to have an analog to the ‘lab technicians’ who get a long-term position to do exactly this kind of work. But somehow I haven’t seen this pan out as well as it does for experimental or lab work. Probably because it’s underpaid.

3 Likes

But in any case having at least some local machine for development is useful (depending how high contention/waiting times are on the cluster).

This fits my case exactly.

In case it’s helpful, jax is an interesting library for HPC with python and it seems to have experimental sparse support:

Thank you, I’ll look into it.

Calculating determinants and pfaffians is essentially “do some row/col permutation, do some product over rows/cols, update a stored value, now repeat with the remaining row/cols” kind of operation. If your rows/cols are alreadt stored as precompiled and dense, this is faster, at the cost of storing many zeros. I rewrote that to work on the value and location arrays (which I got from scipy) of the nonzero values, and tried a little to be clever about it. Many more optimizations are possible (for example, I didn’t use that my matrices are mostly Toeplitz) … But then the code is the tool, not the goal. I am forced to move on. :frowning:

I guess I might need to bring in a computational person to this project as a co-sponsor to hire a student, simply because I wouldn’t know how hireable it makes you if this is what you focus on for your masters or PhD. I’m not in that field. I know my problem and can map it well, but I wouldn’t know how to guide a student to hireability just doing this.

Uh-oh. Yeah. :laughing: :sob: In my experience it’s mostly smart, motivated people who are either learning to mitigate the system or are doing what they can to work around it. But somehow that just doesn’t seem to get to the decision makers with enough force to create the political will to implement reforms. So it seems set to stay broken for the foreseeable future.

Yes. Same goes for journal editors, peer reviewers, and committee members. The field I was in was getting to where a good paper took close to a decade of labor, roughly 40% of which was software engineering, and the expectation was to publish 2-4 papers a year. Since you’d spend maybe half your time actually doing research the only way that’d conceivably work’d is if every paper had ~60 co-authors.

If the objective’s to actually get research done it’s insane. If the objective’s to attract and retain people who are good at spending time getting their name on as many papers as possible instead of doing any scientific work it’s a brilliant success.

Yes. Also because academic funding manages to make job security even less than in the tech industry, which takes some talent. Just not the good kind. Where I was the kind of within institutional political range compromise that was available was essentially what you’d call a professor of practice. So pay was at least around 30-50% of tech industry and people didn’t have to carry as much of the teaching, mentoring, and publication responsibilities as in a tenure track position. But support was pretty much all soft money, meaning only individuals who were good at writing both grants and code could be successful.

So it’s pretty much set up for failure and few individuals are lucky enough to beat the odds for any length of time.

For typical graduate students, probably not very compared to just getting a coding job. The few people I know who’ve had some success have all been in the space of being curious enough to start a second career by way of a PhD. Tiny recruitment pool.

2 Likes

This topic was automatically closed after 270 days. New replies are no longer allowed.