Parsing HTML with regex(lemmy.sdf.org)

posted 9 months ago

pmjv@lemmy.sdf.org

linuxmemes@lemmy.world

11 commentshide report

cross-posted from: https://lemmy.sdf.org/post/12950329

Sort:

Hot Top Controversial New Old

You are viewing a single thread.

View all comments

[ - ]

jet@hackertalks.com

1 point

9 months ago

Putting a statement into old calligraphy is nice. But it all boils down to because I said so. If you’re going to go to that effort you might as well put the rationale for why it can’t possibly parse the language into the explanation rather than because I said so

permalink

report

[ - ]

Pfosten@feddit.de

0 points

9 months ago

The text does technically give the reason on the first page:

It is not a regular language and hence cannot be parsed by regular expressions.

Here, “regular language” is a technical term, and the statement is correct.

The text goes on to discuss Perl regexes, which I think are able to parse at least all languages in LL(*). I’m fairly sure that is sufficient to recognize XML, but am not quite certain about HTML5. The WHATWG standard doesn’t define HTML5 syntax with a grammar, but with a stateful parsing procedure which defies normal placement in the Chomsky hierarchy.

This, of course, is the real reason: even if such a regex is technically possible with some regex engines, creating it is extremely exhausting and each time you look into the spec to understand an edge case you suffer 1D6 SAN damage.

permalink

report

parent

[ - ]

TheOneCurly@lemm.ee

0 points

9 months ago

The section about “regular language” is the reason. That’s not being cheeky, that’s a technical term. It immediately dives into some complex set theory stuff but that’s the place to start understanding.

permalink

report

parent

[ - ]

Blue_Morpho@lemmy.world

0 points

9 months ago

English isn’t a regular language either. So that means you can’t use regex to parse text. But everyone does anyway.

permalink

report

parent

[ - ]

2xsaiko@discuss.tchncs.de

0 points

9 months ago

Huh? Show me the regex to parse the English language.

permalink

report

parent

[ - ]

Blue_Morpho@lemmy.world

0 points

9 months ago

Parsing text is the reason regex was created!

Page 1, Chapter 1, “Mastering Regular Expressions”, Friedl, O’Reilly 1997.

" Introduction to Regular Expressions

Here’s the scenario: you’re given the job of checking the pages on a web server for doubled words (such as “this this”), a common problem with documents sub- ject to heavy editing. Your job is to create a solution that will:

Accept any number of files to check, report each line of each file that has doubled words, highlight (using standard ANSI escape sequences) each dou- bled word, and ensure that the source filename appears with each line in the report.

Work across lines, even finding situations where a word at the end of one line is repeated at the beginning of the next.

Find doubled words despite capitalization differences, such as with The the, as well as allow differing amounts of whitespace (spaces, tabs, new- lines, and the like) to lie between the words.

Find doubled words even when separated by HTML tags. HTML tags are for marking up text on World Wide Web pages, for example, to make a word bold: it is <b>very</b> very important

That’s certainly a tall order! But, it’s a real problem that needs to be solved. At one point while working on the manuscript for this book. I ran such a tool on what I’d written so far and was surprised at the way numerous doubled words had crept in. There are many programming languages one could use to solve the problem, but one with regular expression support can make the job substantially easier.

Regular expressions are the key to powerful, flexible, and efficient text processing. Regular expressions themselves, with a general pattern notation almost like a mini programming language, allow you to describe and parse text… With additional sup- port provided by the particular tool being used, regular expressions can add, remove, isolate, and generally fold, spindle, and mutilate all kinds of text and data.

Chapter 1: Introduction to Regular Expressions "

permalink

report

parent

[ - ]

notabot@lemm.ee

0 points

9 months ago

There’s a difference between ‘processing’ the text and ‘parsing’ it. The processing described in the section you posted it fine, and you can manage a similar level of processing on HTML. The tricky/impossible bit is parsing the languages. For instance you can’t write a regex that’ll relibly find the subject, object and verb in any english sentence, and you can’t write a regex that’ll break an HTML document down into a hierarchy of tags as regexs don’t support counting depth of recursion, and HTML is irregular anyway, meaning it can’t be reliably parsed with a regular parser.

permalink

report

parent

[ - ]

Blue_Morpho@lemmy.world

1 point

9 months ago

For instance you can’t write a regex that’ll relibly find the subject, object and verb in any english sentence

Identifying parts of speech isn’t a requirement of the word parse. That’s the linguistic definition. In computer science identifying tokens is parsing.

https://en.m.wikipedia.org/wiki/Parsing

permalink

report

parent

Show more comments

linuxmemes

!linuxmemes@lemmy.world

Create post

Hint: :q!

Sister communities:

LemmyMemes: Memes
LemmyShitpost: Anything and everything goes.
RISA: Star Trek memes and shitposts

Community rules (click to expand)

1. Follow the site-wide rules

Instance-wide TOS: https://legal.lemmy.world/tos/
Lemmy code of conduct: https://join-lemmy.org/docs/code_of_conduct.html

2. Be civil

Understand the difference between a joke and an insult.
Do not harrass or attack members of the community for any reason.
Leave remarks of “peasantry” to the PCMR community. If you dislike an OS/service/application, attack the thing you dislike, not the individuals who use it. Some people may not have a choice.
Bigotry will not be tolerated.
These rules are somewhat loosened when the subject is a public figure. Still, do not attack their person or incite harrassment.

3. Post Linux-related content

Including Unix and BSD.
Non-Linux content is acceptable as long as it makes a reference to Linux. For example, the poorly made mockery of sudo in Windows.
No porn. Even if you watch it on a Linux machine.

4. No recent reposts

Everybody uses Arch btw, can’t quit Vim, and wants to interject for a moment. You can stop now.

Please report posts and comments that break these rules!

Community stats

6.8K
Monthly active users
1K
Posts
20K
Comments

Community stats

Community moderators