How are the columns worked out?
A PDF stores characters at coordinates, not a table. The columns are found by looking at where text sits: pieces of text on one line, the gaps between them, and the gaps that line up from row to row. Those shared gaps are the column boundaries.
It is the same reasoning a reader uses. It works well on statements, invoices and reports, where columns are clearly separated, and less well when columns run into each other with no visible gutter.
What about descriptions that wrap onto two lines?
A card payment often takes two lines: the merchant on the first, the location and card number on the second. A naive converter treats the second line as a new transaction with empty amounts.
Here, a line with no date, no amounts and text sitting under an existing column is treated as a continuation and joined to the row above. That is why the row count matches the number of transactions rather than the number of printed lines.
Why do the amounts need to be real numbers?
Text that looks like 1,234.56 does not sum. Amounts are parsed into numbers with the formatting kept — currency symbol, thousands separator, decimals — and negatives written in brackets or marked DR come through negative.
Long codes and account numbers are deliberately left as text, so they keep leading zeros instead of turning into 00123 → 123.
What if the statement is a scan?
Then there is no text to read, and the conversion will say so rather than producing an empty sheet. A scan needs OCR first to add a text layer; that tool is in the next wave here.
If the bank offers CSV or OFX downloads, use those instead — they are the original data rather than a reconstruction of a printed page.