Technology

76252 readers

3219 users here now

This is a most excellent place for technology news and articles.

Our Rules

Follow the lemmy.world rules.
Only tech related news or articles.
Be excellent to each other!
Mod approved content bots can post up to 10 articles per day.
Threads asking for personal tech support may be deleted.
Politics threads may be removed.
No memes allowed as posts, OK to post as comments.
Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
Check for duplicates before posting, duplicates may be removed
Accounts 7 days and younger will have their posts automatically removed.

Approved Bots

founded 2 years ago

MODERATORS

L3s@lemmy.world

enu@lemmy.world

technopagan@lemmy.world

L4s@lemmy.world

L3s@hackingne.ws

L4s@hackingne.ws

Google Search could soon let you attach and 'ask anything about a file' (www.androidauthority.com)

submitted 10 months ago by moe90@feddit.nl to c/technology@lemmy.world

15 comments fedilink hide all child comments

you are viewing a single comment's thread
view the rest of the comments

[–] ivanafterall@lemmy.world 11 points 10 months ago* (last edited 10 months ago)

I don't have a specific figure for you. My use-case is I'm trying to write a non-fiction book. I've got a ton of old newspaper articles in PDF format. The Library of Congress' built-in OCR is very helpful, but very lacking and, in some cases, can miss large swaths of pages or generate really unhelpful gibberish that requires painful cleaning. I've had similar results from every other OCR tool I've tried.

Thus far, in using Claude/ChatGPT for transcription of a few dozen articles, I've only had to fix one individual stray word a few times. It's been very close to perfect in my limited testing. High 90%. Impressively, with old newspaper articles where words have worn away or are otherwise very hard to make out even for me, it has done a great job of inferring/recognizing, where OCR would start generating gibberish. I haven't tried hand-writing and suspect that's a different beast, but I know there are tools that have cropped up to that end.