DokMine
OCR redaction

Redact scans and screenshots where the text is pixels

A scanned contract has no text layer to search. DokMine runs OCR over the image, finds the personal data in it, and blacks it out in the pixels — not in an overlay that can be peeled off.

Free tier · no card needed · files removed after processing.

Why scans are the hardest case

Every text-based redaction technique fails on an image. There is nothing to find and replace, nothing to select, nothing for a search to match. The document looks like text to you and is an undifferentiated grid of pixels to the software. That is why scanned material so often gets excluded from a redaction workflow entirely, or gets a black rectangle drawn over it in an image editor and saved in a format that keeps the original layer.

The other half of the problem is silence. A tool that finds nothing in a scan and reports nothing looks exactly like a tool that found nothing because there was nothing there. That failure mode is worse than not running at all, because it produces false confidence.

What DokMine does

Images are read with OCR, matched against the detectors you selected, and the matching regions are blacked out in the image data itself. The pixels change. There is no separate overlay to remove and no recoverable layer underneath.

This applies to standalone images — PNG, JPG, TIFF, BMP, WEBP — and to images embedded inside other documents. A screenshot of an inbox pasted into a Word file, a scanned signature page inside a PDF, a photographed form in a slide deck: all are opened, OCR'd and redacted in place, and the containing document comes back in its original format.

Pages that arrive rotated or upside down are detected and reoriented before reading, with a confidence floor so a low-quality orientation guess on a small image does not send the whole page through sideways.

When it cannot read something, it says so

If OCR cannot run on an image, or the result is too poor to trust, DokMine reports it rather than handing the file back as though it had been checked. That notice is the most useful output the tool produces on a bad scan: it tells you precisely which pages still need a human.

What affects the result

  • Resolution. Roughly 300 DPI is the point below which accuracy falls off sharply. A phone photo of a page at an angle is materially worse than a flatbed scan.
  • Skew and rotation. Corrected where detectable, but heavy skew still degrades recognition.
  • Contrast and noise. Faxes, photocopies of photocopies and shadowed phone photos all lose characters.
  • Handwriting. Not reliably readable. Assume handwritten annotations will be missed.
  • Unusual fonts and stamps. Decorative type, dot-matrix output and overlapping stamps are frequent failure points.
  • Language. Recognition is tuned for Latin-script text.

How to do it

  1. Upload the image, or the document containing it. Batches are fine.
  2. Choose the entity types to look for.
  3. Run extraction first to see what OCR actually managed to read — on a marginal scan this tells you more than the redaction does.
  4. Redact, download, and inspect the returned image at full size.

Scans are the format where review matters most. DokMine does not guarantee that every piece of personal data in an image is found or removed; check the output before you rely on it.

Questions

Does it black out the pixels or just draw an overlay?

The image data itself is modified. There is no separate layer to remove and nothing recoverable underneath the blacked-out region.

Can it read a scanned PDF?

Yes. Pages and embedded images are processed with OCR, matches are redacted in the image data, and the file comes back as a PDF.

What happens if the scan is too poor to read?

DokMine reports that the image could not be read rather than returning the file as though it had been checked. That notice tells you which pages still need a human review.

Does it handle rotated or upside-down pages?

Orientation is detected and corrected before reading, with a confidence threshold so a poor guess on a small image does not cause the page to be read the wrong way round.

Will it read handwriting?

Not reliably. Assume handwritten text and signatures will be missed, and review any document containing them yourself.

Other formats

The same engine, the same account — the hazards differ by format.

Try it on your own file

Sign in with Google and process your first document in under a minute.

Start free

See plans & quotas