What to check when a PDF contains tables
By PDFEditor.ae · Published
How PDFs store tables
A PDF has no idea it contains a table. It stores pieces of text at positions on the page, and the grid lines are just drawn lines. Any tool that "converts a table" has to guess the cells from those positions.
What our PDF to Excel tool produces
Each line of text becomes a row in the spreadsheet. A line is split into columns only where the text contains tabs or wide gaps. In our test with a simple ruled table, every table row arrived in a single cell, with the values separated by spaces. Scanned PDFs produce no text.
Splitting rows into columns in Excel
- Select the column with the table rows.
- Go to Data → Text to Columns.
- Choose Delimited and tick Space (or another separator your data uses).
- Check the preview: values that contain spaces, such as "AED 120", will split too, so adjust before finishing.
- Compare totals with the PDF to make sure nothing shifted.
Validate before you calculate
- Compare the row count with the PDF table, and spot-check several entries.
- Look for descriptions that wrapped onto two lines in the PDF and became an extra row.
- Check decimal commas and thousands separators were read correctly for your locale.
- If numbers are stored as text (left-aligned, no sums), remove stray spaces or currency symbols in a copy, then convert them.
- Keep the untouched extraction and do your analysis in a separate copy, so you can trace corrections.
Before you rely on the numbers
- Ask the sender for the original spreadsheet — it is always more reliable.
- Check negative numbers, decimal separators and currency signs.
- Re-add any rows that wrapped onto two lines in the PDF.