Selectable text and scanned pages need different handling
A text-based PDF already contains characters, so extraction is usually faster and more accurate. A scan is a picture of text. Auto mode runs local OCR on pages without selectable text; Force OCR can help when a page mixes a text layer and scanned content.
Choose the correct OCR language from English, Spanish, French, or German, and select a page range when you only need part of a document. OCR can confuse similar letters, columns, handwriting, or low-resolution scans.
What the DOCX contains—and what it does not
The output is an editable Word document built from extracted or recognized text. It is not a visual reconstruction of the original PDF. Images, complex tables, fonts, exact line breaks, and page geometry may not survive.
For example, a clean one-column report may yield usable paragraphs, while a magazine page with sidebars may need substantial reordering. Compare the DOCX with the PDF before sending it to someone or editing a consequential document.
Privacy and proofreading
Extraction, OCR, and DOCX creation happen locally in the browser. The selected PDF is not uploaded by this tool. Large scanned documents can take time and use significant memory.
Proofread names, figures, punctuation, and tables against the original. If you need only plain text, use PDF to Text; if you need images of the pages, use PDF to JPG or PNG.