Here’s our specimen:
=== START OF INFORMATION SECTION ===
Device Model: OOS22000G
LU WWN Device Id: 5 000c50 0e6e3ad7e
Firmware Version: OOS1
User Capacity: 22,000,969,973,760 bytes [22.0 TB]
Sector Sizes: 512 bytes logical, 4096 bytes physical
Rotation Rate: 7200 rpm
Form Factor: 3.5 inches
Device is: Not in smartctl database 7.5/5706
ATA Version is: ACS-4 (minor revision not indicated)
SATA Version is: SATA 3.3, 6.0 Gb/s (current: 6.0 Gb/s)
SMART support is: Available - device has SMART capability.
SMART support is: Enabled
smartctl -a shows:
SMART Self-test log structure revision number 1
Num Test_Description Status Remaining LifeTime(hours) LBA_of_first_error
# 1 Extended offline Completed: read failure 90% 15604 35387280
# 2 Extended offline Completed: read failure 90% 15594 35387280
# 3 Extended offline Completed: read failure 90% 15577 35387280
# 4 Extended offline Completed: read failure 90% 15538 35387280
# 5 Extended offline Interrupted (host reset) 00% 14761 -
Oh no multiple smart long tests with read failures. It’s Dead Jim! … Or.. is it?
Vendor Specific SMART Attributes with Thresholds:
...
5 Reallocated_Sector_Ct 0x0033 100 100 010 Pre-fail Always - 1
187 Reported_Uncorrect 0x0032 055 055 000recover Old_age Always - 45
197 Current_Pending_Sector 0x0012 100 100 000 Old_age Always - 12
...
Those 3 stats are the things we’re worried about. Only one reallocated sector? What gives?
The state of things is that we have this drive which was pulled from an array. It seems fine, except it FAILS the long smart diagnostic. But the behavior here is very VERY firmware specific. This drive does NOT automatically correct a bad sector (most drives don’t, some do) because the correction mechanism involves telling you the sector is bad ONE TIME then zeroing it and moving on. The reality is that this is a media error and that sector LBA is likely to be relocated elsewhere.
How do you force relocation? You wire to that LBA:
So the fix for a sector that can’t be read and throws a scary hardware error is to… write to it?
Yes. That forces the LBA to reallocated. But it can be misleading because you may also need to overwrite adjacent LBAs, too.
hdparm --read-sector 35387280 /dev/sdf
results in:
reading sector 35387280: SG_IO: bad/missing sense data, sb[]: f0 00 03 40 53 02 00 0a 00 96 f7 1b 11 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00
and
hdparm --write-sector 35387280 /dev/sdf
will write. If you do the write operation (which requires a secret --force command because its so destructive) then try to do a read:
succeeded
0000 0000 0000 0000 0000 0000 0000 0000
0000 0000 0000 0000 0000 0000 0000 0000
0000 0000 0000 0000 0000 0000 0000 0000
0000 0000 0000 0000 0000 0000 0000 0000
0000 0000 0000 0000 0000 0000 0000 0000
0000 0000 0000 0000 0000 0000 0000 0000
0000 0000 0000 0000 0000 0000 0000 0000
0000 0000 0000 0000 0000 0000 0000 0000
0000 0000 0000 0000 0000 0000 0000 0000
0000 0000 0000 0000 0000 0000 0000 0000
0000 0000 0000 0000 0000 0000 0000 0000
0000 0000 0000 0000 0000 0000 0000 0000
0000 0000 0000 0000 0000 0000 0000 0000
0000 0000 0000 0000 0000 0000 0000 0000
0000 0000 0000 0000 0000 0000 0000 0000
0000 0000 0000 0000 0000 0000 0000 0000
0000 0000 0000 0000 0000 0000 0000 0000
0000 0000 0000 0000 0000 0000 0000 0000
0000 0000 0000 0000 0000 0000 0000 0000
0000 0000 0000 0000 0000 0000 0000 0000
0000 0000 0000 0000 0000 0000 0000 0000
0000 0000 0000 0000 0000 0000 0000 0000
0000 0000 0000 0000 0000 0000 0000 0000
0000 0000 0000 0000 0000 0000 0000 0000
0000 0000 0000 0000 0000 0000 0000 0000
0000 0000 0000 0000 0000 0000 0000 0000
0000 0000 0000 0000 0000 0000 0000 0000
0000 0000 0000 0000 0000 0000 0000 0000
0000 0000 0000 0000 0000 0000 0000 0000
0000 0000 0000 0000 0000 0000 0000 0000
0000 0000 0000 0000 0000 0000 0000 0000
0000 0000 0000 0000 0000 0000 0000 0000
Ayyyy! It worked!
So you’d do another smart long test… and it’ll fail once again. Why?
Because, really, its sort of a firmware bug. Not every LBA is stored individually either.
bump by one?
hdparm --read-sector 35387281 /dev/sdf
We incremented the sector number… and whats this? another read error? In smartctl mode it’s reading more than one LBA at a time so the error means some LBA failed starting at the LBA in the smart log.
How to fix? I usually use a little script like this one:
for i in $(seq 35387270 35387290); do
if sudo hdparm --read-sector "$i" /dev/sdf >/dev/null 2>&1; then
echo "OK $i"
else
echo "BAD $i"
fi
done
Notice carefully the LBA numbers. 10 less and 10 more than the problem one. (You might even want to go +100 / -100 or +20 / -20). The output looks like:
OK 35387270
OK 35387271
OK 35387272
OK 35387273
OK 35387274
OK 35387275
OK 35387276
OK 35387277
OK 35387278
OK 35387279
OK 35387280
OK 35387281
BAD 35387282
OK 35387283
OK 35387284
So we have “more” sectors we need to write to – #82 up there is ALSO bad and it’s likely the smart test flagged this “group” of LBAs under #80 and still fails.
With that fixed, we can re-run smart long test. Also, this scrip gives you a “feel” for those LBAs. Does i get stuck on a number for a sec before returning OK? If so that might not be a good sign for overall drive health.
# smartctl --test=long /dev/sdf
...
Testing has begun.
Please wait 1808 minutes for test to complete.
Test will complete after Sat Sep 5 21:10:08 2026 EDT
Use smartctl -X to abort test.
Writes to the drive are destructive!
But if you have errors from syslog/dmesg about specific sectors this will let you “surgically” / zero those LBAs which will squelch the physical hardware errors.
hdparm is hugely useful for this type of overwrite surgery.
Here’s what I do to a “fresh” disk (HIGHLY DESTRUCTIVE):
sudo fio --name=fill \
--filename=/dev/you-disk-here \
--rw=write \
--bs=8M \
--direct=1 \
--iodepth=16 \
--size=100% \
--verify=0 \
--fill_device=1
then when that’s done re-check:
5 Reallocated_Sector_Ct 0x0033 100 100 010 Pre-fail Always - 4
That’s the thing to keep an eye on. Is that counting up? Are there more errors? A little bi of media error on a drive is somewhat normal. Increasing media errors/read errors is a much more severe problem.
I’ll report back if this drive survives another 3 months (actually have two of these 22tb drives with just 1-2 sectors of media errors).
A Tale Of Two Firmwares
So I have some “official” Seagate 22tb and some white-label 22tb. The white-label disks are more fothcoming with errors than their official counterparts imho. It’s interesting:
SMART Self-test log structure revision number 1
Num Test_Description Status Remaining LifeTime(hours) LBA_of_first_error
# 1 Extended offline Completed: read failure 00% 13068 -
# 2 Offline Completed without error 00% 13014 -
# 3 Extended offline Completed: read failure 00% 12997 -
# 4 Short offline Completed without error 00% 3 -
# 5 Short offline Completed without error 00% 0 -
These drives have the same problem, media error. I can see some pending sectors (low #) and a low # of unreadable sector errors but the smart test on the “official” seagate drive does not report which LBA failed during the extended offline smart test.
In that case you can look through dmesg for clues or go with the full-drive-overwrite with fio.
Update
And what was the result of the subsequent long smart test?
SMART Self-test log structure revision number 1
Num Test_Description Status Remaining LifeTime(hours) LBA_of_first_error
# 1 Extended offline Completed without error 00% 15676 -
# 2 Extended offline Completed: read failure 90% 15622 35387376
# 3 Extended offline Completed: read failure 90% 15604 35387280
# 4 Extended offline Completed: read failure 90% 15594 35387280
# 5 Extended offline Completed: read failure 90% 15577 35387280
# 6 Extended offline Completed: read failure 90% 15538 35387280
# 7 Extended offline Interrupted (host reset) 00% 14761 -
5 of 5 failed self-tests are outdated by newer successful extended offline self-test # 1
the advice is not “multiple smart read failures? trash!” because our most recent smart test completed without error. But we did need to run the test to fill the drive which takes a couple of days on 22tb.


