Decoding SMART Telemetry with smartctl

Tadios Abebe | Oct 2, 2026 min read

In production environments storage drives are one of the few server components guaranteed to degraded and eventually fail over time. Traditional hard drives suffer from mechanical wear while modern SSDs have flash cells with finite write endurance. When a drive begins to fail silently, it can trigger huge performance penalties or stall storage I/O across your workloads.

S.M.A.R.T. (Self-Monitoring, Analysis and Reporting Technology) is a diagnostic and reporting subsystem implemented on virtually all modern storage devices. The command line tool smartctl which is a part of the smartmontools package provides a direct interface to read this data.

However, if you’ve ever run smartctl -a /dev/xxx, you know the output can be confusing:

  • SATA and SAS report health in completely different formats.
  • Some attribute tables contain confusing raw vs. normalized values.
  • Drives hidden behind enterprise hardware RAID controllers refuse to respond without specialized pass through flags.

The smartctl utility is part of the smartmontools package, which is available in the package repositories of almost most Linux distribution. So it can be installed either by

apt install smartmontools -y

for debian based systems, or On RHEL based systems we can run

dnf install smartmontools -y

To view what storage drives the kernel currently sees, you can run lsblk which list all block devices attached on the system

lsblk

SATA vs. SAS SMART

One of the most confusing things when inspecting disk health is why two drives can return completely different smartctl outputs. The reason comes down to the underlying protocol, So if you inspect the output of smartctl for SATA drives you will get a numbered attribute table compared to a plan text SCSI log page of SAS drives. In addition the health metrics for SATA SSDs are RAW values while they are a percentage values for SAS SSDs

Decoding SATA SSDs

On SATA drives, SMART outputs a structured table of vendor defined attributes. Here is a real output from a 2 TB Crucial SSD:

root@zcloudbole01node05:~# smartctl -a /dev/sdd
smartctl 7.3 2022-02-28 r5338 [x86_64-linux-6.8.12-8-pve] (local build)
Copyright (C) 2002-22, Bruce Allen, Christian Franke, www.smartmontools.org

=== START OF INFORMATION SECTION ===
Model Family:     Crucial/Micron Client SSDs
Device Model:     CT2000MX500SSD1
Serial Number:    2210E61656C5
LU WWN Device Id: 5 00a075 1e61656c5
Firmware Version: M3CR043
User Capacity:    2,000,398,934,016 bytes [2.00 TB]
Sector Sizes:     512 bytes logical, 4096 bytes physical
Rotation Rate:    Solid State Device
Form Factor:      2.5 inches
TRIM Command:     Available
Device is:        In smartctl database 7.3/5319
ATA Version is:   ACS-3 T13/2161-D revision 5
SATA Version is:  SATA 3.3, 6.0 Gb/s (current: 6.0 Gb/s)
Local Time is:    Thu Apr  2 22:48:38 2026 EAT
SMART support is: Available - device has SMART capability.
SMART support is: Enabled

=== START OF READ SMART DATA SECTION ===
SMART overall-health self-assessment test result: PASSED

General SMART Values:
Offline data collection status:  (0x80) Offline data collection activity
                                        was never started.
                                        Auto Offline Data Collection: Enabled.
Self-test execution status:      (   0) The previous self-test routine completed
                                        without error or no self-test has ever 
                                        been run.
Total time to complete Offline 
data collection:                (    0) seconds.
Offline data collection
capabilities:                    (0x7b) SMART execute Offline immediate.
                                        Auto Offline data collection on/off support.
                                        Suspend Offline collection upon new
                                        command.
                                        Offline surface scan supported.
                                        Self-test supported.
                                        Conveyance Self-test supported.
                                        Selective Self-test supported.
SMART capabilities:            (0x0003) Saves SMART data before entering
                                        power-saving mode.
                                        Supports SMART auto save timer.
Error logging capability:        (0x01) Error logging supported.
                                        General Purpose Logging supported.
Short self-test routine 
recommended polling time:        (   2) minutes.
Extended self-test routine
recommended polling time:        (  30) minutes.
Conveyance self-test routine
recommended polling time:        (   2) minutes.
SCT capabilities:              (0x0031) SCT Status supported.
                                        SCT Feature Control supported.
                                        SCT Data Table supported.

SMART Attributes Data Structure revision number: 16
Vendor Specific SMART Attributes with Thresholds:
ID# ATTRIBUTE_NAME          FLAG     VALUE WORST THRESH TYPE      UPDATED  WHEN_FAILED RAW_VALUE
  1 Raw_Read_Error_Rate     0x002f   100   100   000    Pre-fail  Always       -       0
  5 Reallocate_NAND_Blk_Cnt 0x0032   100   100   010    Old_age   Always       -       0
  9 Power_On_Hours          0x0032   100   100   000    Old_age   Always       -       19881
 12 Power_Cycle_Count       0x0032   100   100   000    Old_age   Always       -       798
171 Program_Fail_Count      0x0032   100   100   000    Old_age   Always       -       0
172 Erase_Fail_Count        0x0032   100   100   000    Old_age   Always       -       0
173 Ave_Block-Erase_Count   0x0032   034   034   000    Old_age   Always       -       861
174 Unexpect_Power_Loss_Ct  0x0032   100   100   000    Old_age   Always       -       108
180 Unused_Reserve_NAND_Blk 0x0033   000   000   000    Pre-fail  Always       -       132
183 SATA_Interfac_Downshift 0x0032   100   100   000    Old_age   Always       -       0
184 Error_Correction_Count  0x0032   100   100   000    Old_age   Always       -       0
187 Reported_Uncorrect      0x0032   100   100   000    Old_age   Always       -       0
194 Temperature_Celsius     0x0022   069   060   000    Old_age   Always       -       31 (Min/Max 0/40)
196 Reallocated_Event_Count 0x0032   100   100   000    Old_age   Always       -       0
197 Current_Pending_ECC_Cnt 0x0032   100   100   000    Old_age   Always       -       0
198 Offline_Uncorrectable   0x0030   100   100   000    Old_age   Offline      -       0
199 UDMA_CRC_Error_Count    0x0032   100   100   000    Old_age   Always       -       0
202 Percent_Lifetime_Remain 0x0030   034   034   001    Old_age   Offline      -       66
206 Write_Error_Rate        0x000e   100   100   000    Old_age   Always       -       0
210 Success_RAIN_Recov_Cnt  0x0032   100   100   000    Old_age   Always       -       0
246 Total_LBAs_Written      0x0032   100   100   000    Old_age   Always       -       135603240797
247 Host_Program_Page_Count 0x0032   100   100   000    Old_age   Always       -       10499724170
248 FTL_Program_Page_Count  0x0032   100   100   000    Old_age   Always       -       11090850191

SMART Error Log Version: 1
No Errors Logged

SMART Self-test log structure revision number 1
No self-tests have been logged.  [To run self-tests, use: smartctl -t]

SMART Selective self-test log data structure revision number 1
 SPAN  MIN_LBA  MAX_LBA  CURRENT_TEST_STATUS
    1        0        0  Not_testing
    2        0        0  Not_testing
    3        0        0  Not_testing
    4        0        0  Not_testing
    5        0        0  Completed [00% left] (0-65535)
Selective self-test flags (0x0):
  After scanning selected spans, do NOT read-scan remainder of disk.
If Selective self-test is pending on power-up, resume after 0 minute delay.

Understanding the Attribute Table Columns

  • VALUE: The normalized score calculated by the drive firmware, typically starting at 100 (or 200) on a brand new drive and counting downward as the drive ages.
  • WORST: The lowest normalized value the drive has ever reached in its operational lifetime.
  • THRESH: The manufacturer’s minimum threshold. If VALUE <= THRESH, SMART triggers a failure condition.
  • TYPE:
    • Pre-fail: Indicates an attribute that monitors critical hardware stability. A breach means failure is imminent.
    • Old_age: Indicates normal wear and tear over time.
  • RAW_VALUE: The raw, vendor specific counter (e.g., hours, byte counts, temperature in °C, or retired block counts).

Key SATA Attributes to Watch

  1. Attribute 202 (Percent_Lifetime_Remain):
    • VALUE: 034, THRESH: 001, RAW_VALUE: 66.
    • On the above crucial SSD output example, the raw value indicates 66% estimated lifetime remaining. As write cycles wear out the flash blocks, this counter decrements toward 0.
  2. Attribute 5 (Reallocate_NAND_Blk_Cnt) & 196 (Reallocated_Event_Count):
    • On the above crucial SSD output example, Both are 0. When NAND flash memory cells degrade beyond the ability of ECC to correct them, the SSD controller permanently marks those blocks as bad and remaps them to reserve blocks. In a healthy drive, this must remain zero.
  3. Attribute 174 (Unexpect_Power_Loss_Ct):
    • On the above crucial SSD output example, Raw value 108. Records how many times power was cut abruptly without a clean shutdown command. On consumer SSDs without Power Loss Protection (PLP) capacitors, high unexpected power loss can occasionally cause metadata corruption.
  4. Attribute 199 (UDMA_CRC_Error_Count):
    • On the above crucial SSD output example, Raw value 0. If this number is greater than zero and actively climbing, the drive media itself is almost certainly fine instead, you might have a bad SATA data cable, dirty drive bay backplane connector, or poor SATA controller link.
  5. Attribute 246 (Total_LBAs_Written):
    • On the above crucial SSD output example, Raw value 135603240797. With standard 512-byte logical sectors, you can calculate the total Terabytes Written (TBW)
    • Over 19,881 operating hours (~2.27 years), this drive has written ~69.4 TB of data.

Decoding SAS SSDs

Enterprise servers and storage arrays frequently use SAS drives. SAS SMART output is structured around SCSI logs instead of the ATA attribute table. Here is a real output from a 3.84 TB Toshiba Enterprise SAS SSD and Seagate 3.84 TB Enterprise SAS SSD:

smartctl 7.3 2022-02-28 r5338 [x86_64-linux-6.8.12-8-pve] (local build)
Copyright (C) 2002-22, Bruce Allen, Christian Franke, www.smartmontools.org

=== START OF INFORMATION SECTION ===
Vendor:               TOSHIBA
Product:              KPM5XRUG3T84
Revision:             B026
Compliance:           SPC-4
User Capacity:        3,840,755,982,336 bytes [3.84 TB]
Logical block size:   512 bytes
Physical block size:  4096 bytes
LU is resource provisioned, LBPRZ=1
Rotation Rate:        Solid State Device
Form Factor:          2.5 inches
Logical Unit id:      0x58ce38ee2080d2d1
Serial number:        3960A02LT0AF
Device type:          disk
Transport protocol:   SAS (SPL-4)
Local Time is:        Thu Apr  2 22:46:09 2026 EAT
SMART support is:     Available - device has SMART capability.
SMART support is:     Enabled
Temperature Warning:  Disabled or Not Supported

=== START OF READ SMART DATA SECTION ===
SMART Health Status: OK

Percentage used endurance indicator: 1%
Current Drive Temperature:     39 C
Drive Trip Temperature:        70 C

Accumulated power on time, hours:minutes 3523:28
Manufactured in week 10 of year 2019
Elements in grown defect list: 0

Error counter log:
           Errors Corrected by           Total   Correction     Gigabytes    Total
               ECC          rereads/    errors   algorithm      processed    uncorrected
           fast | delayed   rewrites  corrected  invocations   [10^9 bytes]  errors
read:          0        0         0         0          0     253877.239           0
write:         0        0         0         0          0      64498.226           0

Non-medium error count:       11

No Self-tests have been logged
root@zcloudbole01node04:~# smartctl -a /dev/sdl
smartctl 7.3 2022-02-28 r5338 [x86_64-linux-6.8.12-8-pve] (local build)
Copyright (C) 2002-22, Bruce Allen, Christian Franke, www.smartmontools.org

=== START OF INFORMATION SECTION ===
Vendor:               SEAGATE
Product:              XS3840SE10103
Revision:             0005
Compliance:           SPC-5
User Capacity:        3,840,755,982,336 bytes [3.84 TB]
Logical block size:   512 bytes
Physical block size:  4096 bytes
LU is resource provisioned, LBPRZ=1
Rotation Rate:        Solid State Device
Form Factor:          2.5 inches
Logical Unit id:      0x5000c500bb360fe7
Serial number:        ZCZ012X30000822150Z3
Device type:          disk
Transport protocol:   SAS (SPL-4)
Local Time is:        Thu Apr  2 22:50:34 2026 EAT
SMART support is:     Available - device has SMART capability.
SMART support is:     Enabled
Temperature Warning:  Enabled

=== START OF READ SMART DATA SECTION ===
SMART Health Status: OK

Percentage used endurance indicator: 2%
Current Drive Temperature:     37 C
Drive Trip Temperature:        70 C

Accumulated power on time, hours:minutes 14967:12
Manufactured in week 31 of year 2020
Specified cycle count over device lifetime:  10000
Accumulated start-stop cycles:  449
Elements in grown defect list: 0

Vendor (Seagate Cache) information
  Blocks sent to initiator = 1058337104
  Blocks received from initiator = 2315195197
  Blocks read from cache and sent to initiator = 0
  Number of read and write commands whose size <= segment size = 0
  Number of read and write commands whose size > segment size = 0

Vendor (Seagate/Hitachi) factory information
  number of hours powered up = 14967.20
  number of minutes until next internal SMART test = 19

Error counter log:
           Errors Corrected by           Total   Correction     Gigabytes    Total
               ECC          rereads/    errors   algorithm      processed    uncorrected
           fast | delayed   rewrites  corrected  invocations   [10^9 bytes]  errors
read:          0        0        82        82         82      43041.791           0
write:         0        0         0         0          0      71560.380           0
verify:        0        0         0         0          0      65252.217           0

Non-medium error count:        0

SMART Self-test log
Num  Test              Status                 segment  LifeTime  LBA_first_err [SK ASC ASQ]
     Description                              number   (hours)
# 1  Background short  Completed                   -     440                 - [-   -    -]

Long (extended) Self-test duration: 22144 seconds [6.2 hours]

Key SAS Health Indicators

  1. Percentage used endurance indicator: 1%:
    • SAS percentage used counts up from 0%.
  2. Elements in grown defect list: 0:
    • SAS drives track defects in two lists: the factory Primary Defect List (P-list) and the runtime Grown Defect List (G-list). The G-list tracks sectors that became unreadable or defective after deployment and had to be remapped. In a healthy drive, this must always be 0. A non-zero, rising number indicates physical media degradation.
  3. Non-medium error count: 11:
    • Non-medium errors refer to transport layer issues: SAS link synchronization retries, command timeouts, or bus resets. Small values can accumulate during server reboots or backplane initialization, but a rapidly increasing count points to cable or backplane problems.

Bypassing Hardware RAID Controllers

When drives are connected to dedicated hardware RAID controllers, such as HP Smart Array controllers or Dell PERC controllers, the operating system does not see the physical disks directly. Instead, /dev/sda is a virtual logical volume created by the controller.

If you attempt a direct query without specifying the controller driver:

smartctl -a /dev/sda

The controller will only return an error saying /dev/sda: requires option '-d cciss,N' for HP Smart Array controllers or failed: DELL or MegaRaid controller, please try adding '-d megaraid,N' for Dell PERC controllers

To pass commands through the RAID controller down to individual physical disks, smartctl provides the -d <type> flag:

  • For HP Smart Array: -d cciss,N
  • For Dell PERC / LSI / MegaRAID: -d megaraid,N

N is a placeholder for the zero based index of the physical drive (0, 1, 2, …).

Here is a real output from a 1.6 TB HP SAS SSD and WD 500 GB SATA SSD:

smartctl 7.2 2020-12-30 r5155 [x86_64-linux-5.15.0-133-generic] (local build)
Copyright (C) 2002-20, Bruce Allen, Christian Franke, www.smartmontools.org

/dev/sda: Option -d cciss,N requires N to be a non-negative integer
=======> VALID ARGUMENTS ARE: ata, scsi[+TYPE], nvme[,NSID], sat[,auto][,N][+TYPE], usbcypress[,X], usbjmicron[,p][,x][,N], usbprolific, usbsunplus, sntjmicron[,NSID], sntrealtek, intelliprop,N[+TYPE], jmb39x[-q],N[,sLBA][,force][+TYPE], jms56x,N[,sLBA][,force][+TYPE], marvell, areca,N/E, 3ware,N, hpt,L/M/N, megaraid,N, aacraid,H,L,ID, cciss,N, auto, test <=======

Use smartctl -h to get a usage summary

root@kulunetworks:~# smartctl -a /dev/sda -d cciss,0
smartctl 7.2 2020-12-30 r5155 [x86_64-linux-5.15.0-133-generic] (local build)
Copyright (C) 2002-20, Bruce Allen, Christian Franke, www.smartmontools.org

=== START OF INFORMATION SECTION ===
Vendor:               HP
Product:              VO001600JWZJQ
Revision:             HPD2
Compliance:           SPC-5
User Capacity:        1,600,321,314,816 bytes [1.60 TB]
Logical block size:   512 bytes
Physical block size:  4096 bytes
LU is resource provisioned, LBPRZ=1
Rotation Rate:        Solid State Device
Form Factor:          2.5 inches
Logical Unit id:      0x50000f0b01186c60
Serial number:        S5KWNA0R108209
Device type:          disk
Transport protocol:   SAS (SPL-3)
Local Time is:        Fri Oct  2 02:13:24 2026 UTC
SMART support is:     Available - device has SMART capability.
SMART support is:     Enabled
Temperature Warning:  Enabled

=== START OF READ SMART DATA SECTION ===
SMART Health Status: OK

Percentage used endurance indicator: 1%
Current Drive Temperature:     26 C
Drive Trip Temperature:        60 C

Accumulated power on time, hours:minutes 33969:09
Manufactured in week 04 of year 2021
Accumulated start-stop cycles:  169
Specified load-unload count over device lifetime:  0
Accumulated load-unload cycles:  0
Elements in grown defect list: 0

Error counter log:
           Errors Corrected by           Total   Correction     Gigabytes    Total
               ECC          rereads/    errors   algorithm      processed    uncorrected
           fast | delayed   rewrites  corrected  invocations   [10^9 bytes]  errors
read:          0        0         0         0          0          0.011           0
write:         0        0         0         0          0          0.065           0

Non-medium error count:      273

SMART Self-test log
Num  Test              Status                 segment  LifeTime  LBA_first_err [SK ASC ASQ]
     Description                              number   (hours)
# 1  Background short  Completed                   -   12488                 - [-   -    -]

Long (extended) Self-test duration: 1634 seconds [27.2 minutes]
smartctl 7.3 2022-02-28 r5338 [x86_64-linux-6.8.12-8-pve] (local build)
Copyright (C) 2002-22, Bruce Allen, Christian Franke, www.smartmontools.org

=== START OF INFORMATION SECTION ===
Model Family:     WD Blue / Red / Green SSDs
Device Model:     WDC  WDS500G2B0A-00SM50
Serial Number:    194609A00E1A
LU WWN Device Id: 5 001b44 8b1bbbe6f
Firmware Version: 401020WD
User Capacity:    500,107,862,016 bytes [500 GB]
Sector Size:      512 bytes logical/physical
Rotation Rate:    Solid State Device
Form Factor:      2.5 inches
TRIM Command:     Available, deterministic, zeroed
Device is:        In smartctl database 7.3/5319
ATA Version is:   ACS-4 T13/BSR INCITS 529 revision 5
SATA Version is:  SATA 3.3, 6.0 Gb/s (current: 6.0 Gb/s)
Local Time is:    Fri Oct  2 06:14:57 2026 EAT
SMART support is: Available - device has SMART capability.
SMART support is: Enabled

=== START OF READ SMART DATA SECTION ===
SMART Status not supported: ATA return descriptor not supported by controller firmware
SMART overall-health self-assessment test result: PASSED
Warning: This result is based on an Attribute check.

General SMART Values:
Offline data collection status:  (0x00) Offline data collection activity
                                        was never started.
                                        Auto Offline Data Collection: Disabled.
Self-test execution status:      (   0) The previous self-test routine completed
                                        without error or no self-test has ever 
                                        been run.
Total time to complete Offline 
data collection:                (    0) seconds.
Offline data collection
capabilities:                    (0x11) SMART execute Offline immediate.
                                        No Auto Offline data collection support.
                                        Suspend Offline collection upon new
                                        command.
                                        No Offline surface scan supported.
                                        Self-test supported.
                                        No Conveyance Self-test supported.
                                        No Selective Self-test supported.
SMART capabilities:            (0x0003) Saves SMART data before entering
                                        power-saving mode.
                                        Supports SMART auto save timer.
Error logging capability:        (0x01) Error logging supported.
                                        General Purpose Logging supported.
Short self-test routine 
recommended polling time:        (   2) minutes.
Extended self-test routine
recommended polling time:        (  10) minutes.

SMART Attributes Data Structure revision number: 4
Vendor Specific SMART Attributes with Thresholds:
ID# ATTRIBUTE_NAME          FLAG     VALUE WORST THRESH TYPE      UPDATED  WHEN_FAILED RAW_VALUE
  5 Reallocated_Sector_Ct   0x0032   100   100   ---    Old_age   Always       -       0
  9 Power_On_Hours          0x0032   100   100   ---    Old_age   Always       -       41544
 12 Power_Cycle_Count       0x0032   100   100   ---    Old_age   Always       -       169
165 Block_Erase_Count       0x0032   100   100   ---    Old_age   Always       -       2642815830799
166 Minimum_PE_Cycles_TLC   0x0032   100   100   ---    Old_age   Always       -       302
167 Max_Bad_Blocks_per_Die  0x0032   100   100   ---    Old_age   Always       -       39
168 Maximum_PE_Cycles_TLC   0x0032   100   100   ---    Old_age   Always       -       386
169 Total_Bad_Blocks        0x0032   100   100   ---    Old_age   Always       -       219
170 Grown_Bad_Blocks        0x0032   100   100   ---    Old_age   Always       -       0
171 Program_Fail_Count      0x0032   100   100   ---    Old_age   Always       -       0
172 Erase_Fail_Count        0x0032   100   100   ---    Old_age   Always       -       0
173 Average_PE_Cycles_TLC   0x0032   100   100   ---    Old_age   Always       -       343
174 Unexpected_Power_Loss   0x0032   100   100   ---    Old_age   Always       -       116
184 End-to-End_Error        0x0032   100   100   ---    Old_age   Always       -       0
187 Reported_Uncorrect      0x0032   100   100   ---    Old_age   Always       -       0
188 Command_Timeout         0x0032   100   100   ---    Old_age   Always       -       0
194 Temperature_Celsius     0x0022   069   051   ---    Old_age   Always       -       31 (Min/Max 13/51)
199 UDMA_CRC_Error_Count    0x0032   100   100   ---    Old_age   Always       -       0
230 Media_Wearout_Indicator 0x0032   038   038   ---    Old_age   Always       -       0x262b221e262b
232 Available_Reservd_Space 0x0033   100   100   004    Pre-fail  Always       -       100
233 NAND_GB_Written_TLC     0x0032   100   100   ---    Old_age   Always       -       173843
234 NAND_GB_Written_SLC     0x0032   100   100   ---    Old_age   Always       -       228598
241 Host_Writes_GiB         0x0030   253   253   ---    Old_age   Offline      -       81353
242 Host_Reads_GiB          0x0030   253   253   ---    Old_age   Offline      -       4000
244 Temp_Throttle_Status    0x0032   000   100   ---    Old_age   Always       -       0

SMART Error Log Version: 1
No Errors Logged

SMART Self-test log structure revision number 1
No self-tests have been logged.  [To run self-tests, use: smartctl -t]

Selective Self-tests/Logging not supported

At this point the SMART data can be read as dicribed earlier. But the important thing here is that when dirvers are contained within a hardware RAID controller, we need to obeserve the SMART result of all phsyical drives contained within that RAID pool.

Running Active Drive Self Tests

Passive telemetry only reports events captured during standard reads and writes. If bad sectors develop in areas of a disk that are rarely read passive SMART monitoring won’t detect them until a scrub or backup reads those blocks.

To actively verify the integrity of the drive surface, you can trigger internal self tests.

Types of Self-Tests

  • Short Test (-t short): Runs electrical, mechanical, and basic read checks across a small sample of the drive media. Usually completes in 1 to 2 minutes without interrupting active I/O.
  • Extended / Long Test (-t long): Performs a full read verification across every single block/sector on the drive. Depending on drive speed and capacity, this takes anywhere from 30 minutes to several hours.

To launch a short test:

smartctl -t short /dev/sdd

Self tests execute entirely within the drive firmware in the background. You can check the test status or review past test history with:

smartctl -l selftest /dev/sdd

If you ever need to abort a running test:

smartctl -X /dev/sdd
comments powered by Disqus