Was digging through an archive of scanned documents and saw an OCR output from 2003 that got a 95% accuracy rate on messy handwriting, I honestly didn't think that tech existed back then. Does anyone else look at these old milestones and wonder if we're just adding flashy features now instead of real progress?
What kind of messy handwriting are we talking about here, because that detail matters a lot? Cursive from one writer who's consistent is a totally different problem than a doctor's scrawl or a form filled out by fifty different people. I remember reading that those old systems got decent numbers by being trained on very specific datasets, like postal codes or bank checks, where the range of shapes was narrow. So a 95% on that archive might say more about the test set than the tech itself. Do you know what the source documents actually looked like?