News: 1675187346

  ARM Give a man a fire and he's warm for a day, but set fire to him and he's warm for the rest of his life (Terry Pratchett, Jingo)

A moment of silence for all the drives that died in the making of this Backblaze report

(2023/01/31)


Cloud storage and backup provider Backblaze has released a report on its hard drive failure rates for 2022 which appears to verify that the age of a drive is a key metric for predicting potential failure.

Backblaze regularly releases stats covering the failure rates of all the hard drives under its management, and these datasets have proven a treasure trove for others to analyze for their own purposes – provided they cite Backblaze as the source and do not sell the data.

The latest figures published on the [1]Backblaze blog cover the hard drive failure rates it saw for 2022 across its storage portfolio, which stood at 235,608 individual drives as of December 31.

[2]

However, 4,299 of these were boot drives while another 388 drives were removed from consideration because they had been used for testing purposes or were models for which the company did not have at least 60 drives in operation. This still leaves statistics covering 230,921 hard drives used for data storage purposes.

[3]

[4]

According to Backblaze's principal cloud storage evangelist, Andy Klein, there was a notable increase in the annualized failure rate (AFR) during last year, rising from 1.01 percent in 2021 to 1.37 percent during 2022.

"In our Q2 2022 and Q3 2022 quarterly Drive Stats reports, we noted an increase in the overall AFR from the previous quarter and attributed it to the aging fleet of drives, but is that really the case?" he asked.

[5]

The answer, it seems, is yes.

Klein compared 2021 and 2022 annualized failure rates for large drives (which Backblaze defines as 12TB, 14TB and 16TB) against smaller drives (4TB, 6TB, 8TB and 10TB) and found that every size (with the exception of 16TB drives) showed an increase in AFR between 2021 and 2022. The figures show the smaller drives failing more often, but they are also older.

A chart showing the average age of each drive model deployed by Backblaze against size shows clearly that the smaller a drive is in capacity, the older it tends to be, with 16TB spinners less than six months old while some 4TB and 6TB models were over 90 months old.

[6]

Average age of Blackblaze hard drives

This increase as a drive model ages follows the [7]bathtub curve , whereby a typical drive can be expected to exhibit a rash of early failures, followed by a period when there are relatively few, followed by an uptick once more as all the drives approach the end of their lives and wear out.

[8]Barge off: Nautilus to bring floating datacenters to two new sites in US, France

[9]New SI prefixes clear the way for quettabytes of storage

[10]Backblaze thinks SSDs are more reliable than hard drives

[11]Yes, it's true: Hard drive failures creep up as disks age

Looking at a chart of annualized failure rates over time and by manufacturer, it might appear that much of the overall rise over the past year has been driven by Seagate and Toshiba models. However, Klein points out that in the case of Seagate, most of the models from that vendor are significantly older than many others, and there is also the lifecycle cost of a given hard drive model versus its failure rate to consider.

[12]

Backblaze hard drive failure rates

"In general, Seagate drives are less expensive and their failure rates are typically higher in our environment. Their failure rates are typically not high enough to make them less cost effective over their lifetime. You could make a good case that for us, many Seagate drive models are just as cost effective as more expensive drives," he said.

When it comes to the lifetime annualized failure rates across all the drive models in production at Backblaze, the current overall rate is 1.39 percent, which the company reckons is down from that of a year ago (1.40 percent) and also down from the previous quarter (1.41 percent).

[13]

Klein says that in 2023, the company's focus is expected to be on replacing the older drives with 16TB and larger hard drives, which means that its 4TB drives and the 6TB Seagate drives (which have an average age of 92.5 months) are likely to go.

As usual, the complete data set used for Backblaze's report is available from the company's [14]Hard Drive Test Data page. ®

Get our [15]Tech Resources



[1] https://www.backblaze.com/blog/backblaze-drive-stats-for-2022/

[2] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/storage&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2Y9mdkpcAYyHOe0v0i5erUgAAABY&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0

[3] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/storage&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44Y9mdkpcAYyHOe0v0i5erUgAAABY&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/storage&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33Y9mdkpcAYyHOe0v0i5erUgAAABY&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0

[5] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/storage&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44Y9mdkpcAYyHOe0v0i5erUgAAABY&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[6] https://regmedia.co.uk/2023/01/31/averageage.jpg

[7] https://www.backblaze.com/blog/drive-failure-over-time-the-bathtub-curve-is-leaking/

[8] https://www.theregister.com/2022/12/07/nautilus_to_bring_floating_datacenters/

[9] https://www.theregister.com/2022/11/22/new_si_prefixes_clear_the/

[10] https://www.theregister.com/2022/09/13/backblaze_ssds_hdds/

[11] https://www.theregister.com/2022/08/03/hard_drive_failure_rate/

[12] https://regmedia.co.uk/2023/01/31/failurerate.jpg

[13] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_onprem/storage&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33Y9mdkpcAYyHOe0v0i5erUgAAABY&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0

[14] https://www.backblaze.com/b2/hard-drive-test-data.html

[15] https://whitepapers.theregister.com/



1% failure rate

Yet Another Anonymous coward

Anyone stopped to think how amazing that actually is ?

Given the speed these things are spinning at and the precision the heads need to hit a bit of data the fact that only 1% fail each year is incredible (haven't read the report to see if that includes drives that failed so early they should really have been caught by the manufacturer)

I'd like to see a report like this for flash drives and SSDs

An_Old_Dog

The good thing about hard drives is they usually give advance clues they're about to fail (check your S.M.A.R.T. logs). The bad thing about flash drives and SSDs is, In my experience, they simply unrecoverably fail once-and-forever, but I'd like to see some volume stats on this.

Re: I'd like to see a report like this for flash drives and SSDs

alain williams

Backblaze do provide [1]ssd-drive-stats .

[1] https://www.backblaze.com/blog/ssd-drive-stats-mid-2022-review/

They(Backblaze) have done them before

Anonymous Coward

But they broke them out into a separate list.

SSD's are tricky because they are so usage dependent though. But there is also a real problem with SSDs keeling over well before their write endurance limit kills them. There were a ton of mundane controller failures out there a few years ago.

Re: They(Backblaze) have done them before

DS999

Controller failure is a problem for HDDs as well - that's usually what causes those early failures in the first few months of operation.

It would be interesting if they tracked S.M.A.R.T. stats for remaining life for their SSDs along with failure rates. Storage Review did a long term test a few years ago running SSDs flat out for as long as it took for them to fail. Some exceeded their write life by as much as 3x.

EDIT: I read the article after posting this and it looks like they are doing exactly what I wished for above so it'll be interesting to check back in a few years.

Re: I'd like to see a report like this for flash drives and SSDs

Sampler

In an environment like this though, unrecoverably fail isn't really an issue as the data will always exist in multiple copies elsewhere, it's only really an issue for the home users not taught better and badly managed infrastructure that should know better.

At least spinning rust hard drives seem more stable nowadays

Andy Non

I remember back in the 80's. There was a tendency for brand new hard drives to fail within the first week or two of use. If they survived beyond that period they tended to last for several years.

Re: At least spinning rust hard drives seem more stable nowadays

JoeCool

That's the "bathtub curve" comment

Re: At least spinning rust hard drives seem more stable nowadays

Andy Non

Interesting. Not heard the phrase before, but it is quite apt.

https://en.wikipedia.org/wiki/Bathtub_curve

1980s Hard Drives

An_Old_Dog

Back then, we would run a testing program overnight on the drives going into the computers we were building. We caught many bad blocks not on the manufacturer's defect list (less than a dozen per drive), and a few outright drive failures, but by-and-large, they passed, and the computers they went into were not brought back to the shop for repair.

We also had a 100% component testing policy, and I caught a shipment of "Hexa" brand multifunction I/O boards from China in which ALL twenty were defective!

How busy are the devices ?

alain williams

Do we assume that all disks are as busy as the rest of them ? I would have thought that the ones doing more work might fail earlier. I cannot see some sort of I/O count.

They exclude boot devices as presumably they are not that busy.

Re: How busy are the devices ?

DS999

I never observed much difference in life between drives that were active 24x7 in DBs and drives that had a more sedate life as lower utilization file servers. There may be some slight correlation there but it wouldn't be worth them selecting for.

The reason they pull out boot drives could be something else, like choosing a smaller size for those or maybe they recycle veteran drives that were used for the main file store and don't want the physical handling required to count against them.

Re: How busy are the devices ?

Richard 12

Boot drives have a very different usage pattern to the data drives.

Depending on the OS and setup, a boot drive might be read just the once during boot and then spun down until the next update, spinning all the time but almost entirely read-only, or have near-continuous writes to log files.

They also tend to be much smaller.

Either way, it makes sense to monitor the two separately. Might be clues as to drive types which fare better for each type of workload.

Process promptly.