How to extract PDF text into Word and understand the limits

By PDFEditor.ae · Published

What this tool does

PDF to Word reads the text layer of a PDF and writes it into a new DOCX file you can edit. It is designed for reusing words, not for reproducing the page. Expect a plain document: the original layout, fonts, images and table structure are not kept.

Steps

  1. Open PDF to Word and upload the PDF.
  2. Click Convert to Word.
  3. Download the DOCX and open it in Word or another editor.

Open PDF to Word

What the Word file looks like

  • It starts with a short header noting it was converted and how many pages the PDF had.
  • Each line of the PDF becomes its own paragraph, so sentences that wrapped onto a new line are split.
  • Page breaks appear as markers such as "-- 2 of 4 --".
  • Images, colours and columns are not carried over; table cells come out as lines of text.

When it will not work

Scanned PDFs and photos of documents have no text layer. In that case the tool now stops with a message instead of giving you an empty file. This site does not offer OCR (text recognition) at the moment.

If you need the layout as well as the text, open the PDF in Microsoft Word directly (File → Open), which rebuilds the layout, or ask the sender for the original document.

Tools used in this guide

Related guides