On Wednesday 07 November 2007 16:55:13 Aaron Kulkis wrote:
> Billie Walsh wrote:
> > On 11/07/2007 James Knott wrote:
> >> One thing to bear in mind, is that drives have spare sectors, which get
> >> used as others fail.  The warning is to tell you that the drive is well
> >> on it's way to failing and should be replaced ASAP.  You were lucky that
> >> it didn't fail sooner.  What you did, is comparable to disabling the
> >> engine light on a car, rather than fixing what's causing it to turn on.
> >
> > There was nothing wrong with the drive. There was something in SMART
> > that was wrong. I've seen SMART say that a brand new drive is failing.
> > SMART is nothing like the engine light on a car.
>
> On the basis that SMART was wrong *once*, are you willing
> to bet YOUR DATA that his old drive has NOT run out of
> internal spare sectors?


I thought I'd go back and look up a discussion on Security Now about this. 

SMART, or the Smart Monitoring Analysis and Reporting Technology is not 
actually all that smart. Not only does it occasionally report drives that are 
not failing, it often fails to report a drive that IS failing.  When drives 
went to the IDE interface standard, the controller was moved onto the drive, 
the drive started being smart, and had its own microprocessor on it. And 
Compaq, who was the big leader in the clone market then, got  Seagate, 
Quantum, and Connor Peripherals to give us a way to know what’s going on 
behind the scenes in the drive. In order to get the manufacturers to agree,  
the SMART specification had to be left very loose. So what we’ve ended up 
with, even today, 15 years later, is a specification which is sorrowfully 
weak. 

And a Google paper that looked at over a hundred thousand failed drives said 
they often saw seeing no uniform meaning behind many of these SMART 
parameters across a large install base of hard drives. They also say a lot of 
the failed drives had no SMART errors at all.

The drive can die spontaneously while SMART is completely happy and sees 
nothing going wrong. At the same time there are drives which look like 
they’re on their last throes from a standpoint of the SMART data that just 
keep on going for years. 

 It said in the study, "After our initial attempts to derive such models yield 
relatively unimpressive results, we turn to the question of what might be the 
upper bound of the accuracy of any model based solely on SMART parameters. 
Our results are surprising, if not somewhat disappointing. Out of all failed 
drives, over 50 percent of them have no count in any of the four strongest 
SMART signals, namely scan errors, relocation count, offline relocation, and 
probational count. In other words, models based only on those signals can 
never predict more than half of the failed drives.”

So essentially what Google found is that many drives were failing, more than 
half of theirs were failing where nothing showed up at all in the SMART 
subsystem; and also that exactly the reverse was happening, is that SMART was 
showing things where drives never failed. There were things they found, for 
example, when they would ask the SMART system to scan the drive, if an error 
was found during that scanning, the drive was 39 times more likely to fail in 
the next 60 days than all other drives. Except that it turns out 39 times 
more likely wasn’t predictive enough to say, okay, we should replace the 
drive, because it turns out that there were lots of drives that had scan 
errors that never failed.

In other words, don't count on SMART to tell you whether or not to replace a 
drive. 

Personally I use Spinrite for that, but it's not free, and there is no demo 
version. I find it works, and I have no connection to GRC other than as a 
satisfied user.

There's a whole discussion on the Security Now Podcast about this, which you 
can read at http://www.grc.com/sn/SN-081.htm.


-- 
Bob Smits [EMAIL PROTECTED] 

A: Because it messes up the order in which people normally read text.
Q: Why is top-posting such a bad thing?
A: Top-posting.
Q: What is the most annoying thing on usenet and in e-mail? 
--
To unsubscribe, e-mail: [EMAIL PROTECTED]
For additional commands, e-mail: [EMAIL PROTECTED]

Reply via email to