We already wrote about an Unraid disk that will not mount and a TrueNAS pool that will not import, about the Rebuild button in general, about a RAID 1 mirror left degraded until the surviving disk was the only copy, and about a RAID 5 that loses a second drive. This page is the pool that is still imported. A scrub found checksum errors. The next click is Replace.
What the scrub actually reported
A ZFS scrub reads allocated blocks and checks them against the checksum ZFS stored with each block. It is not chkdsk, and it is not a resilver. The status that matters is zpool status -v, not the sentence in the alert email.
Each disk line has three counters:
- READ — the disk failed an I/O. The read did not complete.
- WRITE — a write to that disk failed.
- CKSUM — the disk returned data, and the checksum did not match. ZFS got a block. It was the wrong bytes. That can be a sector on its way out, a cable, a backplane, or an HBA. It is still a block you cannot trust.
The scan line says whether the scrub is still running or already finished, how much it repaired, and whether errors remained. "Repaired" is bytes ZFS rewrote from a replica that still matched. On a mirror, that replica is the other side. On raidz it is the rest of the vdev, reconstructed from parity. The rewrite lands on the disk that had the bad copy. It is a repair of those blocks. It is not a new member, and it is not proof the disk has stopped returning bad data.
zpool status -v is the list under permanent errors, when there is one. Those are blocks no remaining copy could repair. The line can be a file path. It can also say metadata, or a hex id, when the bad block is pool structure rather than a document you can see in a folder. A file path is not a complete inventory of the damage. A metadata line is not a file to delete.
Checksums on several disks at once are often the path between the HBA and the platters — a cable, a backplane, a power feed — rather than three disks that all decided to die in the same hour. Replacing one of them does not fix a bad cable. It does start a resilver across that same path.
The pool can still look fine
ONLINE with a non-zero CKSUM means ZFS has not failed the vdev out. The pool is up. Applications are still writing. The scrub may still be reading. DEGRADED means a member is already faulted, missing, or out of the vdev, and the pool is still serving. Fault tolerance on that vdev can already be zero: a mirror down to one disk, a raidz1 down to no parity, a raidz2 that has already lost two. FAULTED is the next state. ZFS stops serving the pool because it no longer has enough replicas to trust.
A pool is often more than one vdev. A stripe of mirrors, or two raidz vdevs side by side, has no parity across the stripe. Both sides of one mirror, or two disks in one raidz1, are enough to fault the whole pool even when every other vdev is healthy. The status screen's pool line is the one that matters. A green vdev next to it does not keep the pool up.
The scrub found the bad blocks by reading them. If the scan line still says the scrub is in progress, it is still reading them. That is the moment people open the status UI and click Replace, or type zpool replace, because the email said checksum and the docs say to replace a disk with checksum errors. Those docs assume the other copies are good and that you can afford the read. A pool that is the only copy of the files, on disks that just failed a checksum, is the case that assumption does not cover.
There is a maintenance case, and it is narrow. The scrub has finished. CKSUM is on one disk. Every other member is zero. Nothing is clicking. A current copy of the files already exists on media that is not this pool. Replacing that one disk is the same kind of job the rebuild page allows when one member failed, the array is online, and the rest are confirmed healthy. This page is the other job: the scrub is still running, or it finished with errors, or you have not confirmed the other disks, or there is no other copy. Replace is how that job gets spent.
A scrub that found checksum errors is a report, not a spare disk. If the pool is the only copy and a replace or a resilver is about to start, stop and get it evaluated. Free evaluation, no data, no fee.
What not to do
Do not race Replace. zpool replace, or Replace on a TrueNAS pool status screen, attaches a new disk and resilvers onto it. The resilver reads the allocated data from the copies still in the vdev and writes the new member. When the pass finishes, the old disk is dropped out of the vdev. Scrub and resilver do not run together. Starting the replace cancels the scrub and makes the resilver the scan. The disks that just returned bad checksums do not get a rest. They get the read that is supposed to build the replacement. If a hot spare is configured, a disk that faults during the scrub can start that resilver without a click. A spare kicking off is the same full read. Do not add a second replace on top of it.
A resilver does not repair a block that has no good replica. Permanent errors stay broken. The new disk holds whatever the remaining members could still return, including a hole where the checksum failed on every copy. The read itself is what finishes a marginal drive. Large disks make that pass long enough for the next member to die in the middle of it.
Do not zpool clear and call it fixed. Clear resets the READ, WRITE, and CKSUM counters. It does not put the bad bytes back, and it does not retire the disk. The status looks quiet. The blocks the scrub could not repair are still unrepaired. People clear, watch a clean pool line, and start the replace as if the scrub had been a false alarm. Clearing a FAULTED device can also make ZFS reopen it. That is the disk coming back into the set, not a test.
Deleting the files named under permanent errors, then clearing, is the same kind of tidy. The names disappear from zpool status -v. The disk that produced the bad reads is still in the vdev. A metadata error has no file to delete. Do not "clean up" the list so the alert stops.
Do not force the bad member back online. zpool online, or a clear that reopens a FAULTED disk, puts a member ZFS already stopped trusting back under I/O. The next thing the pool does with that disk is read it. This is not the Force Online in the RAID battery post, and it is not the Force Online in the RAID 1 page. Those are controller buttons on mirror or cache metadata. This one is a ZFS device the scrub or the pool already faulted. Do not treat them as the same click.
Do not detach the disk to get it out of the way. zpool detach removes a mirror member, a spare, a cache device, or a log device. It does not remove a data disk from a raidz vdev — that temptation is Replace, above. Detach is permanent for that membership. On a mirror that just showed checksum errors, each side can hold the only good copy of a different block. The CKSUM column names the disk whose copy failed that check. It does not promise the other side can read every other block. Detaching the "bad" side, or replacing it and letting the resilver drop it at the end, throws the disk out of the vdev before anyone has an image of it. Offline is not detach. Do not detach because you wanted the disk to stop being used for an hour.
Do not destroy the pool and recreate it. Destroy is a wipe of the pool labels and the pointer to the datasets. "Restore from backup" does not run if the backup was this pool. The sentence TrueNAS prints when a pool will not import — destroy and re-create from a backup source — is the import page. If your pool is imported and the scrub is the problem, you are not on that screen. Do not export it so the sentence can apply.
Do not reach for zpool import -F or -FX. That is not this page. Those flags are for a pool that will not import. They rewind transactions to an older consistent label. A committed rewind is a write to pool metadata. The import page already says a dry run is information and a committed rewind on dying disks is a write. We are not writing the force-import post here. A scrub reporting checksum errors on a pool that is already imported is a read problem and a replace problem. Exporting the pool so -F can "repair" the checksums spends the label you still have.
Do not run the scrub again to see if the errors were real, not while a disk is clicking or the counters are climbing. A scrub is a full read. A second one is another full read. The first pass already answered the question.
Why the resilver takes the pool from DEGRADED to FAULTED
DEGRADED still serves files. The redundancy that made the checksum repair possible is already used up for every block that had only one good copy. The resilver then reads those remaining copies from one end of the allocated space to the other and writes a new disk. That is a heavier, longer version of the read the scrub was doing.
On a mirror, the survivor is the source. There is no parity elsewhere. A sector the survivor cannot return is a hole, unless the disk you are about to detach still has a copy of that block from before it failed the checksum. The resilver will not quietly prefer the disk you are replacing once it has decided that disk is the old member. It hammers the side it is rebuilding from. When that side drops, the vdev faults.
On raidz1, one disk of parity is the whole budget. Checksum errors on one member spend it for those blocks. A second disk that starts throwing READ or CKSUM errors during the resilver leaves the vdev with nothing to reconstruct from. The pool goes FAULTED. That is the ZFS-shaped version of a RAID 5 losing its second drive — except you start the second failure by ordering the full read. Raidz2 and raidz3 have more budget, until the scrub has already used it. Two disks with checksum errors on a raidz2 means the resilver is reading a vdev that is already at its limit.
The new disk is a prefix, not a save. If the source dies halfway, the replacement holds the blocks written so far and garbage or zeros after that. The original members hold whatever they can still give up. Letting the resilver "finish" after a source has left writes that partial reconstruction out to the end of the new disk. Power it down. The half-written replacement is a disk to image later. It is not a reason to keep the pass going.
Same batch, same hours. Disks bought together fail close together. The checksum was the early report. A resilver long enough to rebuild a large vdev is long enough for the next disk to follow. A mirror that looks like a two-disk RAID 1, and a raidz1 that looks like a RAID 5, both die this way. The button is just spelled Replace instead of Rebuild.
zpool clear, force a faulted member back online, detach the disk the scrub flagged, or destroy and recreate the pool. zpool import -F and -FX are not a repair for this screen. Those writes and those reads spend the only good copies. A lab cannot put back a block a resilver already overwrote.This screen is not the other pages
Unraid unmountable, TrueNAS will not import is a disk or a pool the UI cannot bring up. The clicks there are Format, New Config, and destroy-and-recreate, plus a string of zpool import -f, -F, and -FX after the import has already failed. This page is earlier and different. The pool imported. Datasets are mounted, or the pool is DEGRADED and still up. The scrub found checksum errors. If Unraid is offering Format on an unmountable array disk, or TrueNAS will not import at all, use that post. If an Unraid ZFS pool is mounted and the scrub report is the problem, you are here. Do not follow the import page into a destroy, and do not follow this page into a force-import.
Rebuild vs. recovery is the general difference: a rebuild writes a replacement from the members the controller believes are healthy, and a lab images first. A ZFS resilver is that rebuild. Use the other post for the definition. Use this one when the reason you are about to resilver is a scrub that found checksum errors. The healthy-members condition in that post is the thing the CKSUM column just failed.
The RAID 1 page is a mirror that kept the share online after one member failed, and the warning sat there until the survivor or the rebuild spent the only current copy. A ZFS mirror can look like that from the desk: files open, one disk unhappy. The metadata is not Intel RST, Storage Spaces, or a NAS RAID 1. There is no Force Online button in that sense, and the stale-member story is a checksum and a resilver, not a controller that forgot which disk was current. If the firmware says RAID 1 and there is no zpool status, you are on that page. If zpool status is what you photographed, you are here.
RAID 5 with two drives down is a parity array that has already lost a second member and gone offline. Do not rebuild it on the originals. A raidz1 that faults a second disk during a resilver is the same shape of failure on a different stack. If your pool is already FAULTED because two raidz1 members are gone, the "do not rebuild" rule transfers, and the click list does not. That post is a hardware or NAS RAID 5. It will not tell you what zpool clear does. If the pool is still up and the scrub is the screen in front of you, you are not in the two-drives-down case yet. A replace is how you get there.
QuTS hero is ZFS inside a QNAP. If Storage & Snapshots says Volume Not Active and the button is Recover, that is the QNAP page. If the QuTS pool is up and a scrub is reporting checksum errors, the Replace temptation is this page. Do not use either post as a click script for the other.
What to do right now
Stop writing to the pool. Stop the replace if you have not started it. If the data matters and there is no backup on different media, the next step is a photograph and an evaluation, not a resilver.
- Stop. Do not click Replace. Do not run
zpool replace. Do not detach. If a resilver is already running and a disk is clicking, grinding, or the READ and CKSUM counters are climbing, power the machine down. Finishing the pass is more reads of the only good copies, and more writes on the new disk. If a scrub is still running and the disks are marginal, stop it withzpool scrub -s. That stops the read. It does not fix anything. If the pool is up and the disks are quiet, stop the clients that are still writing — VMs, backup jobs, shares — and do not start a resilver as a test. - Photograph
zpool status -vbefore the next reboot or a clear wipes the counters. Pool state (ONLINE, DEGRADED, FAULTED). The scan line: scrub in progress, scrub finished, or resilver in progress, including the repaired amount and the error count. Every disk's READ, WRITE, and CKSUM. The permanent-error list, file paths and metadata lines both. The TrueNAS pool status screen if that is the UI you have. Screenshots help. They are not a diagnosis. - Label every disk before anything is unplugged. Bay or port, and the serial if you can see it without pulling the drive. ZFS tracks members by id, and people do not. Keep the disk the scrub flagged, the other members, any spare that joined, and any new disk a replace already started writing. Do not format the old member because the UI says the replace is using a different slot.
- Write down what you remember: whether the scrub was still running when Replace was clicked, whether anyone ran
zpool clear, detached a mirror side, brought a faulted disk back, destroyed a pool, or tried an import flag. The vdev layout if you know it — mirror, raidz1, raidz2, more than one vdev. And whether a backup exists on media that is not this pool. "The pool is the backup" means it does not. If the pool is encrypted, the key still has to exist later. Write down that you have it. Encryption without the key is still encryption. - Leave the originals alone. Do not save files back onto the pool. Do not reuse a replaced disk as the new member. Copying off is reasonable only while the pool is mounted, the disks stay quiet, and the copy is going to different media. Stop if reads start failing. If the pool is FAULTED, a disk is clicking, or a resilver is mid-pass, that is an evaluation, not a copy.
Pack loose disks the way you would any failed drive — how to ship a failed drive safely. If a member is clicking, start with power it off. A chassis whose disks you cannot pull cleanly can come in as a chassis. Do not open a drive on a table to see which one failed the checksum.
How a lab handles it
We do not resilver your pool on your hardware to see if the checksums clear. We do not zpool clear, detach the flagged disk, force it back online, destroy the pool, or export it and import with -F or -FX on the originals.
Each member is imaged write-blocked first. That includes the disk the scrub named, the other side of a mirror, every raidz member, a spare that started a resilver, and a replacement a replace had already started writing. A drive that only reads in pieces is still imaged as far as it will go. Imaging uses the lab tools already named on this site, including the ACE Lab PC-3000 and DeepSpar Disk Imager. Work happens on the copies.
On those images the pool is imported from the copies, not force-imported on the disks you shipped. Checksums are evidence about which copy of a block matched. They are not a reason to resilver the patient so ZFS can decide in place. A block the scrub called permanent may still be readable on a member the pool had stopped trusting, as whatever bytes are still on that platter. Choosing which copy to keep is work on the images. A wrong guess costs a clone. The same guess during a resilver writes the new disk and then drops the old one.
What imaging cannot do: put back a block a resilver already overwrote, a mirror side that was detached and then wiped, or a pool that was destroyed and recreated on these disks. We will tell you if the current copy is not there.
The public case log has no ZFS pool, no TrueNAS scrub, and no case where checksum errors were followed by a replace or a resilver. We will not invent one.
The closest published array work is a different stack. DDR-2026-0003: a Synology four-drive RAID 5, two members failed, nothing rebuilt in the NAS, each drive imaged, the array reconstructed from the images. Full recovery. RAID 5 tolerates one failure. That set had lost two, so it was offline rather than scrubbing. 21607: twelve HPE 1.8 TB SAS drives on one SmartArray, a five-drive RAID 5 and a seven-drive RAID 6, three RAID 6 members failed, nothing rebuilt on the original controller, every member imaged, the arrays reconstructed virtually. Full recovery. It is a hardware controller, not a zpool. Those outcomes belong to those drives. They are not a rate, and they are not a prediction about a pool whose scrub just reported checksum errors. We don't advertise a success rate.
If the pool will not import, start with Unraid and TrueNAS. If the button is Rebuild on a RAID card or a NAS that is not ZFS, start with rebuild vs. recovery. If a second RAID 5 member is already down, start with two drives down. For the work itself, see RAID data recovery.
ZFS scrub checksum errors FAQ
The pool still says ONLINE and the files open. Can I replace the disk the scrub flagged?
zpool status -v lists permanent errors, or if there is no backup on different media, do not start zpool replace or click Replace.Does a checksum error mean the file is already gone?
zpool status -v then names the file, or it names metadata. A resilver does not invent a block that had no good copy.The scrub says it repaired data. Is the disk fine now?
Should I run zpool clear so the error count goes back to zero?
zpool clear resets the error counters. It does not rewrite the bad blocks, and it does not make a marginal disk healthy. The next status looks clean, which is how people talk themselves into a replace on a pool that was not actually repaired. Clearing a FAULTED device can also make ZFS try to reopen it. That is a member coming back, not a diagnostic. Leave the counters on the screen until you have photographed zpool status -v.The scrub is still running. Will Replace wait until it finishes?
zpool replace, or clicking Replace, makes the resilver the scan. The disks that just returned bad checksums are the ones read end to end so a new member can be written. If those disks are clicking, or the READ and CKSUM counts are climbing, stop the scrub with zpool scrub -s and do not start the replace. Stopping a scrub stops the read. It does not repair the pool.What do Replace, Detach, and bringing a faulted disk back online actually do?
zpool replace, or Replace on the pool status screen) attaches a new disk and resilvers onto it: a full read of the remaining copies, a full write of the new member, and the old disk dropped from the vdev when that pass completes. Detach removes a mirror member, spare, cache, or log device from the pool. It is not a pause, and it is not how a raidz data disk is removed. On a mirror, detaching the disk you think is bad throws away a platter that may still hold the only good copy of blocks the other side missed. Bringing a FAULTED disk back — zpool online, or zpool clear so ZFS reopens it — puts a member the pool already stopped trusting back under that read. None of these copy the files onto different media first.Someone said to destroy the pool and restore from backup, or to run zpool import -F.
zpool import -F and -FX are not this page either. Those flags are a different topic: a pool that will not import, and a rewind of pool metadata. This pool imported. A scrub found checksum errors. Do not export it so you can force-import it as a repair. Do not destroy it and build a new pool on these disks.How is this different from an unmountable pool, a RAID rebuild, a RAID 1 left degraded, or RAID 5 with two disks down?
Is there a ZFS scrub checksum-error case in the public log?
I already started the resilver, cleared the errors, or detached a disk. Is it too late?
zpool clear to tidy the counters, no destroy, no export and force-import. If a disk is clicking or the error counts are still climbing, power the machine down. A resilver you stopped and a resilver that finished are different jobs. Bring every disk, including the one you were replacing, the new member, and any spare that joined on its own. We image first and tell you what is left. We will not guess from a screenshot which blocks the resilver already wrote.A scrub that found checksum errors is the report. Replace is a resilver of the only copies you have left. Stop, and start with a free evaluation before anyone races that resilver.
Request free evaluation →Free evaluation · No data, no fee · Talk directly with a technician.