I can't believe it, a 4 TB Crucial this time, they are falling like ninepins. I'll have to check if it's still under warranty.
I only use Samsung, Crucial and WD Blue, only the Samsungs are surviving.
I can't believe it, a 4 TB Crucial this time, they are falling like ninepins. I'll have to check if it's still under warranty.
I only use Samsung, Crucial and WD Blue, only the Samsungs are surviving.
How weird. Ive got altogether about 4SSDs on 24x7 and none have failed. I only ever had one failure about 10 years ago. That was a Kingston. All the rest are I think Kingston's too.
WD is not a name I have much regard for any more
Crucial used to be good for RAM...
4TB is however a bit too bleeding edge for me. I think I have 2TBs in my experimental NAS
I originally used them in a NAS and I wonder if the constant writing has knackered them.
Well look at what SMART has to say.
Are you checking the SMART on your fleet ?
Check all of them, and see if you've been beating the piss out of them.
The NAS may not have (periodic) TRIM running. This helps with wear leveling, and the drive will fall back to its own heuristic wear level method otherwise.
Lots of issues are not documented by utilities.
[Picture]My NS100 flotilla are "mushy" and slow a bit due to error correction after about three months. While the drive could measure the "mush level" by showing the average error rate per 512 byte sector, I don't know of a way or means of extracting that data. Sammy drives have also been mushy, so their "MLC-like" nomenclature may not exactly be on display. The ones I own at least, have not been mushy, and there is a way (in firmware, behind the scenes) to fool the user into thinking a drive is not mushy.
Paul
Something sounds amiss here. I've run maybe a hundred SSDs in desktops and datacentres over 15 years and only had I think three fail. One was an OCZ in the very early days - a brand notorious for failures. More recently I had a Samsung enterprise SSD and a Sabrent consumer drive fail.
How exactly have you decided they have failed?
Theo
Either they run at 2 Mb/s or Windows can read them one minute but not the next!
How do they do in Linux? A few read passes with 'dd' should indicate whether they're readable or having problems, and rules out any Windows problems.
Theo
Windows? There's yer problem right there, lady.
You can do a simple bad block scan (*without* updating $BADCLUS), by using the bad block scanner in HDTune 2.55 free. On gnome-disks, they keep adding functions, otherwise ddrescue and clone to /dev/null can be used to chart the condition of the SSD drive surface.
If an SSD throws an actual CRC Error on a run... you are in serious trouble. The unit will brick in 3...2...1 . This is unlike an HDD. Each storage type "has its own slope at the precipice". Some (cheap) SSDs lack the firmware code to be monitoring the number-of-errors per sector ("correctable"), and rewriting blocks which are mushy-but-still-usable. On TLC based drives, it is likely that every block has one errored bit, and needs some error-corrector-love before the corrected sector is sent to the user computer. The blocks are NOT rewritten when they have a single errored bit. That would wear out the drive in no time -- instead they can allow as many as fifty errors to accumulate in a sector, without caring. Now you know why your cheese-flavoured drive reads at 300MB/sec when the manufacturer claimed it reads at 535MB/sec (no errors in sector...).
If you are seeing SSD failures, then you need to start characterizing the fleet and figuring out the root cause. Does the NAS modulate the supply line on the drive, to save power, or does it use the Deep Sleep command to save power ? The "constant writing to drive" was already mentioned as a potential root cause.
Using CrystalDiskInfo or the Samsung Toolbox (or other tool box for other brands), you can collect basic-but-not-conclusive info about your SSDs.
Other than that, you need a disaster plan. What if my NAS caught fire ? What if the NAS PSU rail voltage, was 3V higher than the allowed value, and all SSDs were ruined ? And so on. Disaster plan. The first time I mentioned the possibility of a common mode overvoltage failure, a poster wrote in and reported that is how he lost his RAID. PSU destroyed it for him. You have no redundancy when all drives burn instantly at the same time. Only your offline/disconnected HDD with the backup image of the RAID, is there as your disaster plan.
*******As for brand reliability, if you ask a home user, they would say "their Samsung was flawless". More than one professional has reported a non-zero failure rate for the Samsung consumer drives. But due to the difficulties of determining a prevalent root cause type, you aren't going to find an explanation for what kills them. Suffice to say, Samsung is good but has a non-zero premature failure rate. Again -- disaster plan.
Other brands, are so shabby, you can tell by their runtime behavior that you own a cup of noodles. Disaster plan -- NOW :-) I've backed up one SSD drive with a valuable download on it, just because I now know the drive is shabby and cannot be trusted (that means it could run for twenty years and be positively saintly, but it does not change the fact that it wobbles all day long). Whereas a Samsung, I would be saying something like "oh, I can back that up next week". Well, maybe you can say that, or maybe not. Things that don't wobble, and then tip over, now isn't that just one awful smell ??? That can happen to you. We hope, not too often. This is where HDD are nicer. They provide all sorts of hints about their declining health (little clicks, SMART stats, benchmark failures).
Paul
OK. I will need to investigate how to do that and and fire up my Linux box, number 1 on my 10 port KVM :-)
I am not inclined to use it for data whatever the outcome.
doesn't any one run the provifed utilities to see how worn their SSDs are Dave
I don't because I have no idea what they are.
I have Crystal Disk Info and it supplies lots of data for spinning disks (although I have no idea how to interpret it) but only raw data for SSD/NVMe.
Most of the manufacturers have a utility that shows drive health. For WD and SanDisk there is a Dashboard here :-
I do every 6 months or so. No errors are ever reported, though the hours reads and writes go up.
Only SSD I ever had fail just stopped working with errors everywhere. Itt was only a few weeks old
I still have its warranty replacement
I have downloaded the Sandisk (incorporates WD), Crucial and Sabrent versions and installed them.
I ran the Sandisk one on the main PC and it reported all is well.
I then put the dodgy Crucial drive in a carrier on my server, not recognised at all. Put it in an external USB3 caddy and it flickered in and out of existence in File Manager.
I ran the Crucial tool, it doesn't show drive letters but it is the only 4 TB Crucial SSD so no problems identifying it. According to Crucial it is healthy with a remaining 100% life expectancy.
Not sure what to believe!
Could be something in yer server... If you have another machine to link it to, try that.
There are 7 drives (mix of NVMe, SSD and spinners) on the server and none of them has issues.
I tried to run Windows file checker, it says it has errors but can't fix them I then told it to optimise the drive and it couldn't find it!
Mmm. I see what you mean
I think its time to run a SMART program on it. But it sounds more like timing issues in the drive hardware than a defunct storage medium. If its under warranty send it back
+1
Have something to add? Share your thoughts — no account required.
Ask the community — no account required