File4Edit

Extract PDF tables to Excel

Free, and nothing to upload — File4Edit reads each PDF page in your browser and infers the table from where the glyphs sit. Every grid is shown with a confidence score and stays editable until you save it as .xlsx or CSV.

Every table found is shown editable, with a confidence score, before anything is written.

In detail

A PDF stores glyphs at coordinates, not tables. This tool reads those coordinates off each page, groups them into rows by baseline, then looks for vertical strips of white space that run the full height of the block — a gap that survives every row is a column break, and a gap one row happens to have is a word space. Where the page draws its own column rules, those are read as a second opinion and offered as an alternative.

Because the result is inferred, it is shown before it is written. Every detected table arrives with a confidence score built from how regular the grid is, how full it is, how wide the gutters are and how short a typical cell is. A page whose cells hold sentences is refused rather than saved as a spreadsheet of paragraphs. The preview grid is editable, so the file you download is the grid you read.

Extract PDF tables to Excel at a glance

Input
One PDF at a time, with a text layer. Up to 40 of the pages you select are checked per pass.
Output
An .xlsx workbook with one sheet per included table, or CSV — a single file, or a ZIP when more than one table is included.
What it does
  • Editable preview grid before anything is saved
  • A confidence score per table, in words and color
  • Refuses a page of prose instead of inventing a sheet
  • Finds columns in tables drawn with no lines at all
  • Reads the page's own column rules as a second opinion
  • Handles pages stored rotated, like scans
  • Delete a stray row or column in one click
  • .xlsx with a sheet per table, or plain CSV
Who it is for
Anyone with a table locked inside a report, bank statement, invoice or price list who needs it in a spreadsheet.
Limits
  • A scanned page has no text layer, so there is nothing to extract from it; those pages are named and skipped.
  • Detection is best-effort. Read the grid and its confidence score before you trust the numbers.
  • Up to 40 pages are checked in one pass; narrow the page range to reach the rest.
  • The workbook carries values only — no fonts, colors, merged cells or column widths from the PDF.
Privacy
The PDF is read in your browser. No page, no cell and no file is uploaded.

How to extract a table from a PDF to Excel

  1. Add the PDF Drop a PDF on the page or browse for it. Every page up to the first 40 is checked for a table, and any page with no text layer is named and skipped rather than turned into an empty sheet.
  2. Read the confidence score Each page that yielded a table gets a tab with a colored dot, and the selected table shows its band in words: good, fair or low. Low means the columns are a guess and every row needs reading.
  3. Tune the detection Narrowest column gap sets how much white space counts as a column break — raise it when one column splits in two, lower it when two columns merge. Row spacing sets how far apart two baselines can be and still be one row. Press Detect again to re-read with the new settings.
  4. Fix the grid Every cell in the preview is an editable field. Type over anything the detection got wrong, and delete a stray row or column of page furniture with the small cross beside its number or letter.
  5. Save it Choose .xlsx for one workbook with a sheet per table, or .csv for one file per table. Decide whether plain numbers are written as numbers, then save. Nothing is uploaded at any point.

Extract PDF tables to Excel: common questions

Why does it show me a grid instead of just giving me the file?

Because a PDF contains no tables, only glyphs at coordinates, so every result is an inference and a wrong one looks exactly as convincing as a right one. A spreadsheet with the numbers one column over opens cleanly and nobody re-checks it. The preview is the safeguard: what you read is what gets written, and the confidence score tells you how hard to look.

What does the confidence score actually measure?

Four things about the detected grid: how many rows share the same number of filled cells, how much of the rectangle carries text, how wide the column gutters are, and how short a typical cell is. A page whose columns came from rules the document itself draws scores higher still, because then the boundaries were not inferred at all. None of those can be read off the finished file, which is why they are reported here.

It says no table was found, but I can see one. What now?

Try the two detection sliders. Narrowest column gap is the usual culprit: a tightly set table has gutters of only four or five points, and the default asks for eight. If the page draws vertical lines between its columns, switch Column edges to Lines on the page, which reads those directly instead of inferring them. A page whose cells average more than about 45 characters is treated as running text in columns and refused on purpose.

Can it read a table out of a scan?

No. A scanned page is an image, so it has no text layer and no coordinates to cluster — there is nothing for this tool to read. Those pages are listed by number and left out rather than turned into blank sheets.

Are my numbers written as numbers or as text?

Only an unadorned number like 1234, -12.5 or 1,200 becomes a numeric cell, and you can switch even that off. A percentage, a currency amount, a date or a figure in accountants' parentheses stays text, because each of those needs a cell format to keep its meaning and a bare number would silently be a different value from the one on the page.