Error correction has been in use in various formats and technologies for ages. Any storage format or system that does not do error correction is lossy, and the user may not realise until it is too late.
DDR5 only just about works because of built in error correction, like Ethernet (lots of noise, very little signal).
Could try running a rowhammer prove of concept, that should flip some bits (and may corrupt your system in the process, so tread carefully).
You can tell by them having more memory chips. If you check the kingston specs pdfs I posted above, it should be clear. Very rarely shops post wrong specs, putting on-die ECC modules under ECC, so check the manufacturers website to be sure.
Your dimms clearly are proper in-band ECC.
Passmark memtest does a rowhammer test. Don’t know if it is the same variety as the infamous security issue. Passmark memtest also logs ECC errors though I can’t verify this since I don’t have ECC sticks.
Some user here claimed that unstable mem oc on their ECC AM5 system logged ECC errors.
But on a serious note, consumer hardware is still consumer hardware. Things break in real life all the time (PC‘s or otherwise). I think consumer hardware with ECC is a sane compromise for home users, in terms of power, cost, performance, etc.
By the way @wendell, I noticed Hetzner (EU hosting provider) offers servers with intel core CPUs and Ryzens with ECC. So basically consumer ECC platforms. And they have for years. Perhaps a collab with them could put this all to bed? If you ask nicely they might have logs of ECC actually working and get some exposure out of it? They have done some videos with derbauer in the past.
As of agesa 1.2.x.y generally ecc is ACTUALLY WORKING PROPERLY (not just posting) on nearly all boards, including boards that had previously been problematic.
WHEREAS it HAD BEEN the case that ecc was really only working on a few specific boards i.e. hetzner qualified (maybe, as cloud-rented systems from them haven’t even had updated bioses, so I dont really know how much they care about this because if you followed the channel for a while, you’d know I rented cloud services just to see if I could get ECC and mostly the answer up til just before am5 for epyc launch was “no” excpet MAYBE it was silently correcting single bit errors).
It requires updated agesa as well as recent-ish kernels for the other side of the ras plumbing.
btw. I have the literal passmark hardware, the ddr4 version. DDR5 is not available yet afaik. When you have RAS you can inject errrors… ddr5 isn’t quite as clunky as ddr4 used to be AND “on die ecc” adds to some of the confusion around ddr5.
So to restate, generally, rasdaemon will give you the ground truth of what’s going on. Don’t expect it to work if your board doesn’t have agesa 1.2.x.y and thats true of epyc 4004 as well as 7000 or 9000 series AM5.
Not all boards, especially more “value oriented” boards will work with ecc but that will generally show in linux. Asus has exposed specific “Ecc enable? y/n” options in the bios, and asrock steel legend boards are great all-rounders with generally good ecc support whether thats b650 or x series chipsets.
On AM5 this was problematic (imho) until epyc 4004 support was added, now it actually does work as long as the specific board doesn’t gimp it (whereas in the past only 1 or 2 specific boards needed it).
There is the somewhat annoying nit of platform controller firmware does not log ecc errors out of band but technically now rasdaemon can do that so… its a little different than it was in ddr4 days.
ZFS with semi old hardware with a profen track record it is for me
Can’t you answer that with yes or no instead of posting 30min of videos
/s
but seriously, I will take a look after I got some sleep
I take that as an answer that you did no actual hardware test.
Totally agree. But when it comes to ECC, why trust ground truth?
Lots of people will argue that ECC is meaningless, even for ZFS and that chances of bit flipping are like winning the lottery. ECC is for the scared, and paranoid crowd (me included) that don’t wanna take chances. That is why I need an actual hardware test to beliefe.
This is not a critique on you or your work wendell. This is just my personal opinion. And English is not my native language, sorry if I sometimes come off a little harsh.
I really enjoy your work and I would hope that ECC becomes more mainstream. Because in the last few years, lots of things went south for the homelab community. PC cases not having 3,5 HDDs is just one example. I would love to at least see the opposite trend for ECC.
Kay so tl;dw confirmed. hardware testing is covered.
no, no passmark hardware for ddr5 to test but yes it’s easy to test by injecting errors with ras tools. "with hardware. "
which is working. except for the out of band management logging being handled by the platform.
what exactly are you looking for here??
on die ECC for ddr5 is a yuuggeee step up for general randos data integrity expectations which simultaneously reinforces your apparent “don’t worry about it” argument and weakens it because errors without that would be so frequent as to destabilize it for nominal usage.
ECC udimm is also not as good as rdimm. another, different, argument. and there’s two kinds of ddr5 ECC. one or two chip.
I’d like to have ECC udimm if it’s not wildly more expensive and if it does something, but that’s just me
I have lost about 1.5 years of work throughout my career to bad ram. When I have the option I buy computers with ECC.
My MacBook m2 air uses lpddr5, which is ecc capable, and may use ecc.
my pc is an epyc 9124, and has rdimms. When I got it, the current line of thread rippers was not available yet. Threadripper would probably be the better option, but I don’t really care, nothing I have done has stressed the CPU, it is all about cards, and having enough slots to keep adding stuff without having to remove other stuff, Still only using 44 pcie lanes, or less than half of the available lanes.
back to the original post…
if you were only running server tasks and didn’t care about single threaded performance as much and just wanted ecc, a bunch of disks, and room for expansion, take a look at sienna.
epyc 8004 CPUs start around $400, and go in a motherboard with 6 rdimm ecc ddr5 channels and 96 pcie lanes. They are like cut down genoa motherboards and controllers with efficiency CPU cores. ie as fast as a 7xxx series epyc core.
I would also like to mention that the surface area of all of the server chips is huge, 4-8 times the area of a desktop CPU while only consuming up to twice the power, so there is much less demand on the thermal paste. You still need to get the heat off the die, so if you choose a high wattage die, and want to use air cooling, count the thermal pipe ends, multiply by 30, and that is your maximum practical cpu wattage. if you need more get a different heatsink. Water cooling can be augmented with an air conditioner if it runs hot, air cooling with heat pipes hits a wall, a high wall, but a wall none the less.
That is great to hear. Will watch it in the evening.
That hardware testing is covered
Question I still have, that probably will be answered in your vids:
What is the way forward here?
Do we buy consumer grade hardware and just hope that it works?
Do we buy only new hardware that you or someone else in the forum tested before?
Do we run tests after every BIOS update to see if something broke down along the line?
I’m not sure I would rely solely on this test in a system where ECC is critical. As I understand it it uses the APEI Error INJection mechanism, which sets an error flag in the hardware, so that you can make sure that the OS or BMC can see and handle errors correctly. I assume it does not however test that the hardware actually checks for errors and sets the flags when appropriate.
Things like overclocking/undervolting to instability seems like a much more reliable test (or that passmark hardware, when it becomes available and if one can afford it…).
yes, and just observing “organic” correcteds in a borderline oc scenario also confirms it’s working. it’s pretty easy to do all that.
I have always tested this in motherboard reviews since time immemorial but I haven’t encountered a board where I could inject errors but not get organic correcteds/uncorrecteds in an oc scenario
I can see people being reluctant with sweeping claims and given how it’s been handled in the past by manufacturers. I’d also want to see some kind of testing on real hardware rather than AGESA XXX improves in this issue therefore * should work.
I’m certainly not saying wendell should do that, but it would probably make a nice video. wink wink But even so, the results would only be valid for the specific mobo model + UEFI version combos tested.
“Anyone” want to make a video showing how this kind of test is done in general? Perhaps that would be more useful, and people could then report in their own results? (I don’t have any DDR5 ECC UDIMMs, otherwise I’d be happy to test my Ryzen 7600 + ASRock B650M-HDV/M.2 system with latest UEFI.)
I did look at those PDFs, but I was still looking for finer points of distinction between the different memory options – RDIMMs with ECC, UDIMMs with different flavors of ECC, which varieties cover which failure modes via what mechanism…
Wendell’s video I think roughly answers my question – that there’s no ECC difference between registered and unregistered DDR5, even in practice, as you can get either 72-bit or 80-bit UDIMMs.
… or not. What is the ECC-related distinction between registered and unbuffered DDR5? Is it definitely the case that UDIMMs with ECC can give you SECDED? I can’t find it spelled out anywhere.
This doesn’t seem to be true in the case of my RAM. dmidecode reports 80 bits. Also that link’s dead.
Those look cool, though probably too weak on single-thread performance for me, unless diminishing returns kick in a lot earlier than I’m thinking. Pretty sure I’m still leaning toward AM5 and a 79- or 99-something (not leaving “pretty good” territory gaming-wise, decently cost-effective, 12 cores ain’t bad…), unless I can come up with an enticing reason to go bigger. Say, hypothetically, a reasonably cost-effective EPYC or TR system had better ECC and would better support doing computationally intensive things in multiple VMs at once, without costing more than 300% the otherwise analogous Ryzen system. I’m still working on digging into things. The systematic approach is a surprisingly large time sink.
rdimms are 80 bits, comprised of 2 sub channels of 40 bits each, comprised of 32 bits of data + 8 bits of checksum. There are 10 chips total.
ecc udimms are 72 bits, comprised of 2 sub channels of 36 bits each, comprised of 32 bits of data + 4 bits of checksum. Are there 10 chips total? or are the 32 bits coming out of 4 bit chips, making 18 chips total.
For ecc udimms you can’t use a single 8 bit chip and pull 2 channels of 4 bits of memory out of it or your throughput would halve, so the checksum part needs to be a 4 bit chip. Unless there are 18 chips on the board, it will look different from the other chips, and probably have fewer pins.
Though they may just give you 10 identical chips of 8 bits each and know that half of the ram on the checksum chip will never get used. If the on memory ecc chips are more sophisticated they could use it like spare area to to fill in holes, but that would probably affect the latency.
Well that’s… odd, given that (again if I understand correctly) there are only 8 physical pins available for ECC in a DDR5 UDIMM slot (4 per subchannel: CB[3:0]_A and CB[3:0]_B). See e.g. the DDR5 SDRAM UDIMM Core document.
I couldn’t find any info about your memory, unfortunately, but I did manage to find a detailed data sheet for MTC20C2085S1EC48BA1 (same as yours except for the R at the end). That clearly shows 4 check bits per 32-bit channel and 72 bit total width.
I’m not sure at all what’s going on here. Could it be an error in the SPD? I mean, the DIMM in the data sheet does have 80 bits total width, just that 8 of them aren’t connected to anything…
(Sorry for the dead link, that document seems to have been deleted.)
I’m starting to doubt this. Doing some research on “truncated Hamming codes” SECDED requires 72-bit codewords for 64 bits of data, or 39-bit codewords for 32-bit data. But maybe they’ve found improved codes, or they are doing error correction on a complete cache line transfer (edit: or simply on two transfers at a time)?