Do SSD failures follow the bathtub curve? Ask Backblaze
- Reference: 1695745809
- News link: https://www.theregister.co.uk/2023/09/26/ssd_failure_report_backblaze/
- Source link:
[1]Backblaze uses SSDs as boot drives in the server infrastructure for its Cloud Storage platform, while high-capacity rotating drives are typically used for storing and serving up data.
However, they do more than just boot the storage servers, holding log files and temporary files produced by each server. The volume of data a boot drive will read, write, and delete thus depends on the activity of the storage server itself.
[2]
The company [3]previously reported that its SSDs appeared to be at least as reliable as hard drives, but warned this could change as it has not collected SSD data for as long as hard drives and the accumulation of more data could alter the statistics.
[4]
[5]
Backblaze says it has added 238 SSDs to its infrastructure since the last SSD report, ending in Q4 2022. These comprised 110 Crucial drives (model: CT250MX500SSD1), 62 WDC drives (WD Blue SA510 2.5) and 44 Seagate drives (ZA250NM1000).
Looking at the Q1 2023 and Q2 2023 figures, Backblaze notes that some drives appear to have exceptionally high annualized failure rates, with the Seagate model SSDSCKKB240GZR listed with an annualized failure rate (AFR) of over 800 percent, for example.
[6]
This is a fluke of the statistics because of the low number of drives; in Q1 there were just two of this model, one of which failed shortly after being installed. During Q2, the remaining drive did not fail and thus the AFR for that period was zero.
[7]
These figures illustrate why Backblaze considers at least 100 instances of a specific drive model and 10,000 drive days of operation in a specific quarter as a minimum before the calculated AFR can be considered to be reasonable, according to Backblaze storage cloud evangelist Andy Klein
Looking at the AFR over time, Backblaze reports that the AFR across its SSDs was 0.96 percent during Q1 of 2023 and 1.05 percent during Q2. This failure rate is thus up from the previous quarter, but down slightly from the same quarter a year ago. In fact, a chart of the AFR per quarter over the past three years shows that it has fluctuated between 0.36 percent and 1.72 percent, with no apparent underlying pattern.
[8]
However, Backblaze says that the quarterly data is still vital as it can reveal issues such as one particular drive model that was the primary cause of a jump in AFR from 0.58 percent in Q1 2021 to 1.51 percent in Q2 then 1.72 percent in Q3.
"It happens from time to time that a given drive model is not compatible with our environment, and we will moderate or even remove that drive's effect on the system as a whole," Klein said.
[9]
Backblaze earlier this year calculated the average age at which failure occurred for its entire collection of hard drives, and has repeated the calculation for SSDs in this latest report.
This involved collecting the SMART data for the 63 failed SSD drives the company has had to date, which is not a great dataset size for statistical analysis, as Klein admitted. The resulting figure calculated from the data is 14 months, compared with two years and seven months across all hard drives.
But Backblaze cautions this figure is likely to be unrepresentative, as the average age of the entire fleet of SSDs it has in operation is just 25 months.
[10]A moment of silence for all the drives that died in the making of this Backblaze report
[11]As liquid cooling takes off in the datacenter, fortune favors the brave
[12]New SI prefixes clear the way for quettabytes of storage
[13]Yes, it's true: Hard drive failures creep up as disks age
Looking at three drive models for which the company has a reasonable amount of data, Klein found that the average age of the failed drives increases as the average age of drives in operation increases, and it is therefore reasonable to expect that the average age for an SSD failure will increase with time.
Turning to the lifetime annualized failure rate for all of its SSDs, Backblaze reports a figure of 0.9 percent, covering a period from Q4 2018 through to the end of Q2 2023. This figure is up slightly from the 0.89 percent it found at the end of Q4 2022, but down from the same quarter a year ago, when the figure was 1.08 percent.
[14]
However, this includes those drives which have high apparent failure rates because there is just not enough data to make the calculation reliable.
[15]
If the calculation is limited to just those drive models for which there are 100 units in operation and over 10,000 drive days, and also with a confidence interval of 1 percent or lower between the low and the high values, then it cuts the data down to just three drives and an AFR of just 0.6 percent.
Meanwhile, Backblaze has also produced a graph of SSD failures over time to see how well the data matches the classic bathtub curve used in reliability engineering, as the comparable graph for its hard drives does.
[16]
According to Klein, while the actual curve (blue line) showing the SSD failures over each quarter is a bit "lumpy," the trend line (red) does have "a definite bathtub curve look to it."
The trend line is about a 70 percent match to the actual data, so Backblaze says it cannot be totally confident at this point, but for the limited amount of data available, it would appear that the occurrences of SSD failures are on a path to conform to the tried-and-true bathtub curve.
As ever, Backblaze makes the raw data used in its report available on a [17]Drive Stats Data page for anyone to download and analyze – as long as you cite Backblaze as the source if you use the data, and don't sell it, of course. ®
Get our [18]Tech Resources
[1] https://www.backblaze.com/
[2] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/storage&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2ZRNUhU6Ztr3pLVlLCyHqtwAAAIA&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0
[3] https://www.theregister.com/2022/03/04/backblaze_ssd_stats/
[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/storage&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZRNUhU6Ztr3pLVlLCyHqtwAAAIA&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[5] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/storage&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZRNUhU6Ztr3pLVlLCyHqtwAAAIA&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[6] https://regmedia.co.uk/2023/09/26/q1.jpg
[7] https://regmedia.co.uk/2023/09/26/q2.jpg
[8] https://regmedia.co.uk/2023/09/26/afr_chart.jpg
[9] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/storage&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZRNUhU6Ztr3pLVlLCyHqtwAAAIA&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0
[10] https://www.theregister.com/2023/01/31/backblaze_older_drives_fail_more/
[11] https://www.theregister.com/2022/12/29/liquid_cooling_now/
[12] https://www.theregister.com/2022/11/22/new_si_prefixes_clear_the/
[13] https://www.theregister.com/2022/08/03/hard_drive_failure_rate/
[14] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/storage&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZRNUhU6Ztr3pLVlLCyHqtwAAAIA&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0
[15] https://regmedia.co.uk/2023/09/26/select_ssd.png
[16] https://regmedia.co.uk/2023/09/26/bathtub.jpg
[17] https://www.backblaze.com/cloud-storage/resources/hard-drive-test-data
[18] https://whitepapers.theregister.com/
Re: these are consumer SSDs!
Backblaze has long proved that, at their scale and/or needs, enterprisey storage is not more reliable enough to be worth the price difference for them.
Re: these are consumer SSDs!
And bleh on Seagate and WD, their drives suck, as evidenced by the results...
No, it doesn't show that at all. One particular drive failed after 1.4 months, and they only had one of them. This breaks their system somewhat, as it gives a yearly failure rate of 842%, despite them only having one of them. (which would be an actual yearly failure rate of 100% of that model).
They had ~1800 seagate drives installed and two failures in 170 thousand drive days, compared to the Crucial drives which they had ~550 which they had two failures on in 47 thousand drive days.
The seagate drives actually appear to be their favourites, and the WD ones probably have too low a sample size and runtime to be a useful comparison.
not useful results
I'm guessing they use Seagate SSDs because they probably get a super good deal on them for buying so many Seagate HDDs.
But not many people use them of course so can't really leverage this data in a useful way. Unlike their HDD reports I suspect people are unlikely to switch to the SSDs that Backblaze uses based on their results. Would of been nice had they used the more common vendors.
My own personal data point for SSD reliability over the long term is using Sandisk and Samsung enterprise SAS SSDs in my 3PAR 7450 which has been running since Nov 4 2014. Most recent batch of drives was installed in that array in 2016, not many, just 32 total SSDs. Just 1 SSD failure in the first 9 years of operation. Don't recall if it was Sandisk or Samsung that failed, but in Jan of this year that one SSD just went offline without warning. Even now nothing has less than 86% wear life remaining. I physically re-arranged the disks a few months ago so can't extrapolate exactly which kind of SSD failed. All were HPE OEM of course, and at the time were super expensive something like $20k+/drive list price at least, not something Backblaze would ever use on their stuff.
I've yet to have any of my Samsung or Intel SSDs fail on personal devices, starting with Samsung 850 Pro I think was my first one many years ago. Read reports of 980(?) Pro NVMe SSDs having issues though mine have been fine(just about a year old, and came with the fixed firmware already).
It would be interesting to assess failure rate per £ in some meaningful way, to see how much more reliable expensive ones are, as there's always the choice of cheaper ones being bought more often versus pricier ones slightly less so.
As an extra note I've taken to doing a lot of my work in RAM (on a RAM disk) and then just committing it at intervals (with browser cache switched to memory only) - otherwise working on large files can create a lot of writing. For laptops this is unlikely to be a problem, for PCs with stable electricity supply, vary rarely. I wonder if the resulting reduction in writes has any meaningful benefit in regards to lifespan.
d
I think it'd be mostly determined on what kind of SSD you are using and what it's rated for. My previous laptop which was in regular use from mid 2016 until Oct 2022(very little usage since) has a 1TB SATA Samsung 850 pro in it, rated for 300TB of written data. Samsung Magician software says it has 88.1TB written. Also has two Samsung 950 Pro NVMEs, one has 5.6TBW and one has 18TBW. This laptop spent 98% of it's life booted to Linux but does dual boot with Windows 7.
I checked my current laptop (~1 year old), which has a pair of 2TB 980 Pros, using linux's smartctl. One says 113GB written, and I guess the main one says 17.2TB written. Specs say these are rated for 1,200TB. This laptop has spent 99.9% of it's life in Linux(and runs a Windows 10 LTSC VM 24/7) but does dual boot with Windows 10.
for personal use I've only ever purchased Samsung SSDs, important stuff always gets Pro, and non important stuff often gets Evo(even stuff that almost never gets used, other people may choose super cheap brand SSDs for those systems). I have at least one Intel SSD (with a skull on it) in my PS4. I really don't even consider other brands at this point regardless of price/purpose(as things are cheap enough now for sure). Only exceptions was I bought a few HPE OEM SSDs a while back (they were actually Intel though..).
So in my case I don't have any concerns about write lifespan. Not all SSDs are created equal though, the 980 Pro(?) had some firmware issues which caused major problems for some folks. Fortunately not me though.
these are consumer SSDs!
based on the capacities, these appear to all be consumer-ish SSDs. enterprise SSDs should have an even better AFR. And bleh on Seagate and WD, their drives suck, as evidenced by the results...