Publishers Always Innovating(mander.xyz)

posted 2 months ago

fossilesque@mander.xyzM

science_memes@mander.xyz

39 commentshide report

Sort:

Hot Top Controversial New Old

You are viewing a single thread.

View all comments View context

[ - ]

keepthepace@slrpnk.net

2 points

2 months ago

Yes, PDFs are much more permissive and may not have any semantic information at all. Hell, some old publications are just scanned images!

PDF -> semantic seems to be a hard problem that basically requires OCR, like these people are doing

permalink

report

parent

[ - ]

thevoidzero@lemmy.world

1 point

2 months ago

Not just semantics. PDFs doesn’t even have segmentations like spaces/lines/paragraph. It’s just text drawn at locations the text processor/any other softwares inserted into. Many pdf editor softwares just detect the closeness of the characters to group them together.

And one step further is you can convert text to path, which basically won’t even have glyph (characters) info and font info, all characters will just be geometric shapes. In that case you can’t even copy the text. OCR is your only choice.

PDF is for finalizing something and printing/sharing without the ability to edit.

permalink

report

parent

[ - ]