Extract a results table from a paper

Re-typing a table out of a PDF is where transcription errors come from. This recovers the grid from the page and writes figures a spreadsheet can actually work with.

A results table is one of the hard ones

It usually has a two-deck header, a column of standard errors in brackets under the estimates, footnote markers attached to values, and a merged cell spanning a group of columns. Read as plain text the whole thing becomes a line of words with no way to tell which space was a column boundary, and that is unrecoverable.

The grid here is found from where the words sit on the page and where the white runs down it. Cells are grouped by the white inside a row rather than by the column each word starts in, which is what keeps a two-word heading over one column instead of spreading it across two.

A header band set in the same face and size as the body gives a recogniser nothing, so it is found a third way: labels standing over columns of figures is what a header row is. A figure anywhere in the first row settles it the other way.

Check every value before you use it

A recogniser reading a printed page is confident about digits it has got wrong, and a decimal point moving one place in an estimate looks entirely plausible in a cell. Where the paper is a digital PDF the words are taken straight off the page, which is more accurate than any recogniser, and that is the better input by a long way.

Anything shaped like a number is written as one so a column can be totalled or charted. A leading zero, a long run of digits and anything carrying a letter stay as text.

Does it get the footnote markers?

They come across attached to the value as text. Strip them in the spreadsheet if you are going to compute.

Can I do several tables at once?

Yes. Every table found on every page becomes part of the workbook.

Support and billing: support@gosmartpdf.com. Operated by Mrityunjay Kumar, India. Payments by Dodo Payments, our merchant of record.

Loading the editor…