Why scans are the hardest case
Every text-based redaction technique fails on an image. There is nothing to find and replace, nothing to select, nothing for a search to match. The document looks like text to you and is an undifferentiated grid of pixels to the software. That is why scanned material so often gets excluded from a redaction workflow entirely, or gets a black rectangle drawn over it in an image editor and saved in a format that keeps the original layer.
The other half of the problem is silence. A tool that finds nothing in a scan and reports nothing looks exactly like a tool that found nothing because there was nothing there. That failure mode is worse than not running at all, because it produces false confidence.
What DokMine does
Images are read with OCR, matched against the detectors you selected, and the matching regions are blacked out in the image data itself. The pixels change. There is no separate overlay to remove and no recoverable layer underneath.
This applies to standalone images — PNG, JPG, TIFF, BMP, WEBP — and to images embedded inside other documents. A screenshot of an inbox pasted into a Word file, a scanned signature page inside a PDF, a photographed form in a slide deck: all are opened, OCR'd and redacted in place, and the containing document comes back in its original format.
Pages that arrive rotated or upside down are detected and reoriented before reading, with a confidence floor so a low-quality orientation guess on a small image does not send the whole page through sideways.
When it cannot read something, it says so
If OCR cannot run on an image, or the result is too poor to trust, DokMine reports it rather than handing the file back as though it had been checked. That notice is the most useful output the tool produces on a bad scan: it tells you precisely which pages still need a human.
What affects the result
- Resolution. Roughly 300 DPI is the point below which accuracy falls off sharply. A phone photo of a page at an angle is materially worse than a flatbed scan.
- Skew and rotation. Corrected where detectable, but heavy skew still degrades recognition.
- Contrast and noise. Faxes, photocopies of photocopies and shadowed phone photos all lose characters.
- Handwriting. Not reliably readable. Assume handwritten annotations will be missed.
- Unusual fonts and stamps. Decorative type, dot-matrix output and overlapping stamps are frequent failure points.
- Language. Recognition is tuned for Latin-script text.
How to do it
- Upload the image, or the document containing it. Batches are fine.
- Choose the entity types to look for.
- Run extraction first to see what OCR actually managed to read — on a marginal scan this tells you more than the redaction does.
- Redact, download, and inspect the returned image at full size.
Scans are the format where review matters most. DokMine does not guarantee that every piece of personal data in an image is found or removed; check the output before you rely on it.