News: 1585162569

  ARM Give a man a fire and he's warm for a day, but set fire to him and he's warm for the rest of his life (Terry Pratchett, Jingo)

HPE fixes another SAS SSD death bug: This time, drives will conk out after 40,000 hours of operation

(2020/03/25)


HPE has told customers that four kinds of solid-state drives (SSDs) in its servers and storage systems may experience failure and data loss at 40,000 hours, or 4.5 years, of operation.

The IT titan said in a [1]bulletin this month the “issue is not unique to HPE and potentially affects all customers that purchased these drives.” HPE has not identified the SSD maker, though, and refused to do so, saying: “We’re not confirming manufacturers.”

Your humble vulture has seen evidence the faulty drives were made by Western-Digital-owned SanDisk. WD declined to comment.

Meanwhile, Dell-EMC issued an urgent firmware [2]update last month that also [3]mentioned SSDs failing after 40,000 operating hours, and specifically identified SanDisk SAS drives. The update included firmware version D417 as a fix to prevent future data loss.

If you're getting deja-vu, you're not alone. HPE separately [4]warned of certain SAS SSDs dying after their 32,768th hour of operation in November last year.

Drawing a line between these two blunders, HPE noted this month: "While the failure mode is similar, this [latest] issue is unrelated to the SSD issue detailed in this [5]customer bulletin released in November 2019, which describes an SSD failure at 32,768 hours of operation."

To avoid this new data death bug, HPE customers should install SSD firmware version HPD7, described as a "critical" fix. From the March bulletin:

HPE was notified by a Solid State Drive (SSD) manufacturer of a firmware defect affecting certain SAS SSD models used in a number of HPE server and Storage products (ie: HPE ProLiant, Synergy, Apollo 4200, Synergy Storage Modules, D3000 Storage Enclosure, StoreEasy 1000 Storage). The issue affects SSDs with an HPE firmware version prior to HPD7 that results in SSD failure at 40,000 hours of operation (ie: 4 years, 206 days, 16 hours).

For more on this story, turn to our enterprise storage sister site, [6]Blocks and Files .



[1] https://support.hpe.com/hpesc/public/docDisplay?docLocale=en_US&docId=a00097382en_us

[2] https://www.dell.com/support/home/uk/en/ukbsdt1/drivers/driversdetails?driverid=8h6hj&oscode=w12r2

[3] https://www.reddit.com/r/sysadmin/comments/f5k95v/dell_emc_urgent_firmware_update/

[4] https://www.theregister.co.uk/2019/11/25/hpe_ssd_32768/

[5] https://support.hpe.com/hpsc/doc/public/display?docId=emr_na-a00092491en_us

[6] https://blocksandfiles.com/2020/03/24/hpe-enterprise-ssd-40k-hours-flaw/

Anonymous Coward

That seems rather specific. Is there some sort of self-destruct code in the firmware they forgot to remove?

Version 1.0

When I started building RAID systems, I quickly discovered that you should not buy a bunch of disks from the same vendor, always use different batches because otherwise you often get every disk failing within a couple of weeks of each other down the line.

Mtbf

JimPoak

4.5 Years of continuances server use or 9 years of playing candy crush. Brings back memories of mtbf.

Brush up your backup skills.

druck

That seems rather specific. Is there some sort of self-destruct code in the firmware they forgot to remove randomise sufficiently? FTFY

Purposely?

HildyJ

For once there's an HPE article where it's someone else's fault.

I assume it's an attempt by SSD manufacturers to hard-code their assumed maximum number of read/writes. And like so many companies, they didn't bother to be honest with their users, because it would hurt sales.

Oh well...

Anonymous Coward

Delete as appropriate

- Brilliant timing given Covid-19

- A great way to screw more dosh out of your customers with planned obsolescence.

- Others not listed above

just like the printer cartridges

Anonymous Coward

Just like the inkjet cartridges. They expire a couple of months after the date even when unused and completly full.

Anonymous Coward

Oh look it's this thread again.

Dell

Anonymous Coward

Dell sent me an email with the service tag of an affected server I look after. Nice email with the tag in a url to the page to download the drivers, and the name of what you needed.

So good on them. But then again they sell servers with 5 years warranty, so would not them all failing in 4 years 200 days. But then I want not want all the drives to fail at the same time either.

I've always said

J__M__M

it's best to build arrays with non-matching drives. And by always I mean never.

As usual, this being a 1.3.x release, I haven't even compiled this
kernel yet. So if it works, you should be doubly impressed.
(Linus Torvalds, announcing kernel 1.3.3 on the linux-kernel mailing list.)