SSD vs. HDD Power Consumption Testing

So I have a testing environment built and ready to document this very boring task.

Hypotheses:
1.) SSD’s use more power and generate more waste heat (according to most data sheets), though if the workload completes substantially faster than an HDD, the power usage will peak higher but total lower.

2.) Local workloads will benefit more from SSD than network workloads until >1Gbps interfaces are used until a point where SSD’s will be the overall winner.

3.) Power consumption vs. performance increase is relative based on electric costs. The potential savings may not justify the increased expense.

Testing:

Power consumption is tracked over a 24 hour period at the outlet
Usage will be scripted benchmarks simulating standard loading with default Server 2022 power management enabled.

It’s an i5-9500T, 16 GB RAM, H310M M.2 2.0 board, 256 GB Samsung 970 EVO boot drive, with an add in TPM module and onboard gigabit
Open air with the only fan being the CPU fan

An add in 25 or 100 gig card may be used to simulate high throughput network utilization with client power consumption not recorded.

On the agenda:
M.2 NVME - 2TB MLC, 4TB QLC
SATA SSD - 2TB MLC, 4TB QLC
U.2 NVME - Gen 3x4 3.2TB mixed use drive in passive interface card
3.5" HDD - 4TB CMR, 22TB CMR enterprise drive
2.5" HDD - 1TB CMR laptop drive, 1TB CMR enterprise drive, 4TB SMR

What workloads and benchmarks would you like to see?
What network based workloads?
General thoughts?

7 Likes

I expect HDD consumes more electricity when it is active.
If you allow HDD to spin down, I don’t think power consumption is an issue.
To be honest, choosing HDD is mainly to save money. SSD is just better.

2 Likes

(Spoiler alert) not exactly, most SSD’s I’ve seen pulling more power than HDD’s.

Which brings us to where HDD’s still make sense.

This is why I used such a stripped down board: to measure small differences in power consumption.

I am really wanting to know if anyone had specific workloads they wanted to see tested.

1 Like

Most power is used for writing on SSDs. heavy random writes will get you max power draw. be sure to overload or bypass any cache the SSD has. Sequential workloads are much lighter on the controller.

And SSDs benefit greatly from board/PCIe power settings while HDDs can’t really be tweaked that much outside of outdated SMART settings.

SSDs used to be very power-efficient, but as capacity and especially the controllers improved, this isn’t really longer the case. U.2 drives use 6W idle and up to 25W under heavy load. And even modern M.2 orientate towards the 10W spec limit.

My money is on SATA SSDs winning the power game.

HDD are the most expensive when talking IOPS. SSDs are just cheaper there. With capacity it’s the other way around.

6 Likes

CrystalDiskMark’s probably appropriate as a standard reference, but that only covers brief active periods. For longer duration I’m actually not sure if it’s more interesting to test workloads or to characterize use of drive power states. Probably some of both. Given the latter workload response can, in principle, be predicted. But some workload measurement would be needed to validate a model.

One thing I’m curious about is how FanControl (Libre Hardware Monitor) and HWiNFO64 affect drive states. So even just baselines with the drives idling with and without the monitors running would be interesting. If your experience is like mine you may end up measuring the mobo and Windows as much as the drives and software workloads. In particular, it seems like just reading SMART can keep some 3.5s spun up but some of my data suggests chipset drivers influence this.

I don’t know of good data on NGFF NVMe drives’ power draws as a function of power state. Anecdotally, both 2.5 SATA and NGFF NVMe get into the 40-60 mW range on idle, and presumably that’s also true of NGFF SATA. So it might end up being more about what the host does for power management and how the specific drives tested respond. I have the impression some NGFFs don’t like to drop below 0.7-1.0 W but I’m not sure how good the tests are.

I just pulled some consumer drive specs and the few 2.5s still sort of current (MX500, QVO 970, SA510) are around 4-6 W/GB of transfer rate active. Common PCIe 3x4 and 4x4 NGFF NVMes are more like 1 W/GB, plus shorter transfers might mean the drive spends more time in low power states. Presumably it varies some by workload IO intensity and profile. In the past I’ve tried to take a look with an HX1000i but it doesn’t resolve below 1.7 W on 3.3 V.

Within the SLC cache my experience pretty well matches what you describe. Once the cache is full, though, writes throttle to the cache folding rate. So the most active workload can end up being sustained read.

I think the most demanding real-world workloads are basically

  • swap: ~50% write + ~50% read within pSLC, potentially pretty much nonstop for hours, days, or until drive failure at however fast a large model’s moving data in and out of memory (y-cruncher might be convenient for testing here)
  • multi-TB reads: heat up the controller and all the NAND chips but generally only lasts for minutes to tens of minutes per burst (can be done with DiskSpd, though usually I load drives with data that actually needs to be on them and read that)

but maybe there’s another pattern I’m not thinking of. A caution I’d make with this kind of testing (and in real world use) is to monitor drive temperatures and mind heatsinks and airflow. When I started doing it I learnt the hard way some drives don’t throttle and will run themselves up over 100 °C. Makes most M.2 thermal solutions cry, too.

+1 Also, hard drive read speeds collapsing with data age is much less of a concern than with consumer flash. Discussions here usually assume enterprise flash rewrites often enough to avoid that problem, though I’m unsure how well tested the assumption actually is.

3 Likes

So I did deliberately omit NGFF SATA from the test as it is such a deprecated / niche use case.

I included 2.5 SATA HDD in both CMR and SMR as I figured they’d show what the power potential of HDD could be

I am tryin to be as unbiased as possible while maintaining relevance.

That’s not too far off what I’ve been seein.

All of this is gospel
And I have drives that haven’t been powered on in a year+ to test the low power state theories.

Unfortunately they’re all 1TB SATA 870 EVO’s, Though I may have some NGFF/M.2 NVME with last access times approaching a year as well.

Figured that would be outside the scope of this video.

This would be simple usable workloads.
Local - performance benchmarks on a schedule with enough iterations to generate measurable data.

Networked/Server - regular read/write benchmarks over the wire and nightly backups to duplicate long sequential reads, regular restores to duplicate sequential reads. Both are things servers have to deal with on the regular.

Surprised no one has critiqued the testing platform, I have another MoBo coming so I can add in a NIC for the network based file operations. I believe HDD will rule server land over 1 gig, unless it’s a database which are all random iops.

But once over 100 gig connections, the field will change.

Finally, I am running M.2 NVME for the boot drive as this comparison is not for boot drives which are largely “random” iops bound. This is for data drives.

If anyone happens to know a game benchmark I can script from loading to full benchmark run, that would be greatly appreciated.

1 Like

I test with whatever extra I have around, pretty much. ¯\_(ツ)_/¯ Rarely I get some budget that’s for test hardware but even then it’s things we might put in daily use if they do well enough.

I’d expect 3.5s to have a tough time against consumer flash in power terms since minimum power while spinning’s usually 5-8 W and spin down timeouts are usually 10-20 minutes. But with larger deployments the lower cost per TB pays the direct cost of a good bit of uptime. How well that works out overall depends on the power source’s assessed and social cost of carbon, though.

FWIW we have enough users per machine to generally put a 2 TB boot drive to handle all the profiles and a few minor datasets (say 100 GBish). So there’s the IOPSy OS+programs plus usually some large sequential user data component. Even PCIe 3.0 x2’s pretty well fine for that, though, and most of our current machines are at least 3x4. It’s the data drives that handle the terabytes and depending on the night, we might have a several hundred GB backing up over the network or initial syncs pulling a few TB down for local processing the next day.

Aye, so far as I know out of a couple hundred computers we have exactly zero NGFF SATAs deployed. I’ve never seen one in person

Unfortunately furmark’s as far as I’ve gotten in that direction.

1 Like

Great idea. Looking forward to results.

The performance and power consumption of nvme drives has changed quite a bit from gen3 to gen4 and gen5. You need to document what you’re using, maybe even try to represent every generation.

I’d be interested in adding enterprise u.2 drives as well - again the supported PCIe gen will make a difference.

I think finding a desktop workload that allows comparing all drives in a meaningful way is going to be hard.

I’d love to see a time based repeatable db benchmark - e.g. the PostgreSQL bench. Works both locally and network based. Run benchmark for a set amount of time, measure both total transactions and power consumed, create a metric that combines both (e.g. transactions per W) for comparison.

1 Like

Can you post said data sheets? I’ve never seen anything like that! If we’re talking SATA drives SSDs are ahead by half: a 4TB Samsung 870 Evo has an average posted power consumption of 2.5W meanwhile a WD Red CMR 4TB drive is rated at 4.7W. Peak is 5W for the SSD and 21W for the hard drive.
An NVME drive is rated closer and a bit worse compared to an HDD (6W average, 8.5W peak for a Samsung 990 Pro).

So I don’t think HDDs could ever come head even just because an SDD is much faster so it’s gonna take less time to complete the same operations. In Wh/operation terms there’s no beating SSD.

What’s the metric for the heat production? I don’t think that’s the case either. Heat is more concentrated but I don’t think they run hotter. I run four 870 Evos very packed together and I’m seeing temps barely above ambient even when stress testing the array for maximum throughput using fio. I highly doubt four HDDs would fair that well honestly.

But I’m still curious to see the results and thanks for setting up this testing environment.

1 Like

Basically every U.2 runs 20-25W and 5W idle, these buggers need actual proper cooling. The recent NAND generation gets down to 16W per drive under load which is good.

M.2 is power-limited to ~10W so the controller couldn’t go further, thus lesser specs for M.2. Performance drives more and more rely on heat sinks although passive with case airflow is usually plenty.

SATA is a totally different beast. they’re all very low power because they don’t run at NVMe speeds, don’t need fancy controllers,etc.
But SATA SSDs are basically legacy hardware at this point. And NVMe drives use a lot more power.

The largest chunk on energy saving comes from sleep states, which HDDs don’t have unless you use spin-down. Ability to change sleep states in ms and having like 60mW in deep slumber is the real deal with SSD power saving and why we use them in mobile devices.

For home use…you don’t need storage that often, so mostly sleep state. Far better than HDD.
Enterprise Server with 16x U.3 drives == low power? you probably have a 40mm fan row pulling 100W alone just to keep the SSDs cool.

SSD != SSD. Different use cases, different power demands. And old SATA vs. new NVMe is a big difference.

4 Likes

That totally makes sense, but I didn’t see any U.2 in the comparison so I was wondering what he meant to say.

Not all NVME, enterprise gear uses a lot more power.

Well, as I said, the energy saving also comes directly from the fact that they can do an operation much quicker than an HDD can. Sure they use double the power but at thirty or more times the speed of an HDD they don’t need to work on that operation for too long.
If you could measure the Wh/operation an SSD is more efficient.

1 Like

Sure, if you run a benchmark “how many Wh per TB read/written”, SSDs always win and the gap isn’t even funny.

Depends on what performance you want. I run with Samsung 980 that are really low-power but otherwise end up at the tail of every benchmark and writes are horrible. They sleep 99,5% of the time @65-250mW, so these make sense. But they ramp up to ~5W once they see write load.

Then there are my Kingston Fury Renegades…top-performing 4.0 M.2 drives that easily go >70°C under proper load if not cooled, pushing bandwidth close to 4.0 limits, don’t like writing a lot as all consumer drives. 10W+ each.

SSDs aren’t using much power (all things considered), but will, if you want performance out of them.

4 Likes

That definitely came up.

Problem became, I need a whole other rig to measure the faster drives at full throughput. This was just an extra PC I had lying around that I am filling with some extra drives I had as we have all wondered and even bench raced as to which is actually more power efficient.

After the testing, we’ll see what else needs to be factored.

There’s a 3.2 TB Samsung PM1725b PCIe 3.0 x 4 U.2 NVME inside a passive adapter card. It is by FAR the fastest thing in this comparison, but also the biggest power hog running in excess of 20 watts when active. Even idling, it’s hot to the touch (which is why I went open air)

I specifically got a gen 3 U.2 as it does not have the negotiation and compatibility issues with consumer hardware that faster pcie 4 & especially 5 U.2’s have.

My testing rig is most accurate at lower power draws, and substantially easier to measure when processor c states don’t exceed the total drive consumption.

I considered using a new T sku cpu, but the deeper sleep states do very strange things when waking up NVME’s. If 1 test is erroneous during the testing, I have to reset and do it over.

This is already going to take 2 weeks of testing…unless I grab 2 more…

3 Likes

Couple things I wanted to add to the discussion.

  • Since the 7 nm E31T’s not out yet, consumer PCIe 5.0 x4 NGFF NVMes mostly use the 12 nm Phison E26. E26 drives generally are around 12 W, so nominally ~0.9 W per GB/s of transfer rate. E31T engineering samples drop that to ~0.5 W/GB but idle at 2.2 W. Samsung Piccolo is 5 nm but the 990 Evos are 5x2 drives, the 4x4 WD MP16+ and Samsung Pascal are 8 nm. The other controllers I know offhand are all 12+ nm.
  • Current consumer NGFFs aren’t particularly thermally conductive and M.2 thermal solutions typically rely on bursty workloads that don’t overcome their thermal inertia. Most drives use 70 °C flash and, absent active cooling or pretty high end passive with good airflow, all the PCIe 4.0 x4 data I have shows climbs above the flash’s operating temperature range or throttling within a few minutes’ operation at 6-7+ GB/s.

Yeah, getting 25 W out of a U.2 2.5 and 25-70 W from a U.3 EDSFF E3 takes airflow that few desktop cases are set up to provide. At least shroud fans becoming increasingly common helps with providing airflow close to the lower M.2 sockets that’s aligned with transverse fin M.2 heatsinks.

1 Like

Worth adding it to the list in the OP, it’s interesting to see how it fairs in the comparison.

What happens? Could it be related to PCIe power management screwing with the drive? I can’t relate low power states on the CPU with issues on the PCIe bus. Unless there are errors on the bus due to excessive voltage droop, maybe.

1 Like

I think that’s the key. HDDs at idle consume 5W and SSDs at idle consumes almost nothihg. For bursty workloads, there is no match. Only in continous workloads will the HDDs compete.

1 Like

I may be wrong but doesn’t testing for efficiency by measuring power usage over a fixed interval pretty much exclude using benchmarks? The varying intervals between each drive’s low power states would seem to be an important factor that wouldn’t be represented by benchmarks because they tend to either do as much work as possible over a fixed interval or just run until a fixed amount of work is completed, neither of which includes idle periods. Any effects of throttling would also tend towards being over represented.

Benchmarks that feature realistic idle periods both in number and duration would be the exception but AFAIK benchmarks don’t typically idle while running.

A decision may need to be made between either testing workloads that include realistic idle periods and scoring total power usage over the entire test; or, scoring drives on performance and how much power they use to attain that performance using separate metrics for each workload with idle characteristics just being stated so efficiency under varying conditions can be inferred.

Factoring idle power savings into the efficiency scores seems to be the tricky part.

1 Like

I am scripting benchmarks to duplicate workloads.
This is an air gapped network between the 2 machines so no updates can alter results.

This is what I am trying to emulate. Real world is a specific workload that the client waits until completed (unless they cancel and retry a few times before letting it run)

This is where the scripting comes into play. Ideally I will script enough time that the slowest drive can fully complete the benchmark before starting a new one.

Perpetual benchmarking would simply measure the TDP of the device connected, where I want to simulate the real-world utilization. In this metric, I presume 2.5" HDD’s would win as their TDP is lower.

An SSD would complete far more work than an HDD under that load, but by using more power. By scheduling sleep between runs, we can better see the total power utilization as presumably the entire system will go to sleep faster and maybe that additional idle time will balance the scales between HDD and SSD power utilization.

Additionally, I did not include multiple drive configs as an HDD scales better than SSD’s when connected to SATA, and vice versa over PCIe due to platform bandwidth limitations.

Yeah, that’s by design so hopefully I can script benchmarks that auto complete and close so as to make repeatable cycles throughout the day.

I intend to do both with the scripting and total utilization over 24 hours better reflecting multiple runs with low power states between runs. Obviously the HDD will have shorter breaks, but with the lower TDP, I am presuming local workloads will benefit from the SSD storage while network based will favor HDD’s at 1Gbps but to take it to the logical extreme, I’ll run 100gig to see what the breakover point is.

FOR SURE
My most difficult hurdle is the scripted benchmarks.
I remember some reviewers listing gaming benchmark scripting, but I have to limit that pretty significantly due to hardware constraints.

Not to mention deciding to tie up a machine for 2 weeks to finally attempt to answer this question that I see on the forum at least once every other day.

I do appreciate all the feedback from everyone. It reinforces my thought process in building this setup.

4 Likes

Newer processors can enter lower sleep states than many enterprise drives cannot handle. When resuming activity, ERRATIC behavior can occur.

I’m not on an optimization quest, just trying to limit my variables.

The 35 watt CPU is to increase measurement precision and do just that.

When you increase total load measurement, the accuracy percentage (sub 1% < 100 watts) increases as well so a single HDD on a 400 watt EPYC Genoa won’t even register with my current measuring equipment as it’s within error of margin.

2 Likes

ok I see where you’re going… you’re building a benchmark schedule spaced out to run over the course of 24 hours. A tally of total idle time between benchmarks alongside the total power usage figure would be interesting.

I’m not sure gaming benchmarks would be appropriate because they tend to exclude load times and just focus on FPS during the simulated gameplay portion which always runs for the same duration regardless of performance. In terms of efficiency this would skew results in favor of drives that rendered the fewest frames. A gaming benchmark that always renders the same number of frames would be an exception but finding one of these may be difficult.

Going even further, checking to ensure that each test moves the same amount of data for faster vs slower drives may be an easy way to exclude tests that cause faster drives to do more work. For gaming benchmarks this wouldn’t apply but given your methodology I think it’s important to ensure that each test doesn’t end up doing more or less work based on drive performance.

1 Like