What to check when a PDF contains tables

By PDFEditor.ae · Published

How PDFs store tables

A PDF has no idea it contains a table. It stores pieces of text at positions on the page, and the grid lines are just drawn lines. Any tool that "converts a table" has to guess the cells from those positions.

What our PDF to Excel tool produces

Each line of text becomes a row in the spreadsheet. A line is split into columns only where the text contains tabs or wide gaps. In our test with a simple ruled table, every table row arrived in a single cell, with the values separated by spaces. Scanned PDFs produce no text.

Open PDF to Excel

Splitting rows into columns in Excel

  1. Select the column with the table rows.
  2. Go to Data → Text to Columns.
  3. Choose Delimited and tick Space (or another separator your data uses).
  4. Check the preview: values that contain spaces, such as "AED 120", will split too, so adjust before finishing.
  5. Compare totals with the PDF to make sure nothing shifted.

Validate before you calculate

  • Compare the row count with the PDF table, and spot-check several entries.
  • Look for descriptions that wrapped onto two lines in the PDF and became an extra row.
  • Check decimal commas and thousands separators were read correctly for your locale.
  • If numbers are stored as text (left-aligned, no sums), remove stray spaces or currency symbols in a copy, then convert them.
  • Keep the untouched extraction and do your analysis in a separate copy, so you can trace corrections.

Before you rely on the numbers

  • Ask the sender for the original spreadsheet — it is always more reliable.
  • Check negative numbers, decimal separators and currency signs.
  • Re-add any rows that wrapped onto two lines in the PDF.

Tools used in this guide

Related guides