Somebody scanned two hundred pages and the machine went away for a while. This site is about why, and whether it is worth fixing.
This is a working review, not a proposal. Nothing has been built and nothing is being asked for. The purpose is to agree on what the problem actually is, put a defensible number against it, and write down what would have to be true before anyone spends money — including the ways this could turn out not to be worth doing.
Can a Sharp finish a 200-page searchable-PDF scan in essentially the same time as a normal scan with no OCR at all — regardless of how long the OCR afterwards takes?
If the answer is yes, the specific thing David described is solved, and everything else — automatic naming, classification, filing straight to the right matter — is gravy on top of a problem that is already worth fixing. If the answer is no, we will say so on this site rather than in a proposal.
Three things we think are true
Each one is a claim you should be able to knock over if it is wrong. That is what the questions at the end are for.
Scanner availability is tied to OCR capacity, and it should not be
The paper handling and the OCR are two completely different jobs with completely different constraints, but today they sit on one thread. A long OCR queue turns into a busy scanner, and a busy scanner turns into a person waiting in a corridor with an armful of paper.
See where the time actually goes →Capacity is fixed at the size of the box you bought
When twenty large jobs arrive in an hour, job twenty waits for jobs one through nineteen. There is no way to spend a little money for an hour and make that go away. The only lever available today is telling people to scan at a quieter time.
See the queue behaviour →The user experience is the part that cannot be compromised
Walk up, press a button, walk away, the PDF arrives where it is expected. That is the standard set by what is in place now, and anything that asks the user to do more than that has lost before it starts — however good the engineering underneath is.
See what would change and what would not →What we are deliberately not claiming
The fastest way to lose an argument like this is to overstate it. So here is the ceiling on the claim, stated up front rather than discovered later by someone in a meeting.
- OCR does not get faster. The same pages take the same compute. What changes is who waits for it — a queue rather than a person.
- A single small job on an idle system will not feel dramatically different. The gain shows up under load, which is exactly when people currently complain.
- Classification, naming and automatic filing are worth real money, but they are the second conversation. They should not be used to sell the first one.
- Nothing here removes the need for a good scan. Bad originals still produce bad OCR, on any architecture.
- We have not measured your environment. Every number on this site is a model with the assumptions left visible so you can argue with them.
Nothing here makes OCR faster.
The same two hundred pages need the same amount of work done to them. What changes is that the work stops being coupled to a scanner and a person, and that the amount of capacity available to do it stops being a decision you made when you bought a server.
That is a smaller claim than "faster OCR" and a much easier one to defend — and it happens to be the one that fixes the actual complaint.
Where the time goes →