Reading a scanned PDF
Today a PDF with no text layer is refused with a clear explanation and a suggestion to paste the numbers instead. Running optical character recognition over the page images would turn a scanned set of accounts into a usable source. The obstacle is that OCR needs either a binary on the server or an external service, and the second one means content leaving the machine, which needs to be an explicit choice rather than a default.