Rendered at 10:08:54 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
imoverclocked 51 minutes ago [-]
The hardest thing about writing a parser is cognitively accepting what is going to be considered valid input. You can make the best parser that is fast and well specified but invariably someone will (ab)use it in an unexpected way.
Famous examples: despite so many initial good intentions, html tags don’t need to be closed, JSON numbers are too often encoded as strings, YAML can look like what most people expect or it can look progressively more like JSON… and on and on.
simonask 32 minutes ago [-]
I think the second-hardest thing is to accept that CS spent decades optimizing parsing algorithms and grammars, and this is still a significant part of CS curricula in many places. But the practical reality is that parsing is almost never a bottleneck.
If what you're parsing is within the capacity of humans to interact with (so in the range of tens of kilobytes), a grammar that requires an O(N^2) parser is totally fine.
mrkeen 1 hours ago [-]
If you draw a line from 'ad-hoc byte-wrangling nonsense' to 'parser combinators', this can't be more than 20% along it.
Looking at the linked URL parser, why doesn't it look like
url = do scheme
authority
path
query
fragment
where
scheme = ...
authority = ...
etc.
It looks totally ad-hoc.
f311a 1 hours ago [-]
Unfortunately, simple URL parsing breaks on so many things. There is a reason on why every URL parsing library is at least a few thousand LOCs.
One common way to test it is just to pass ipv6 url: http://[f021:d981:b487:e57d:193e:550e::]/
Retr0id 1 hours ago [-]
> LineReader splits input into lines, handles \n and \r\n, and trims the stray trailing \r that malformed input likes to leave behind
Is there a common source of extra \r in malformed inputs, beyond those existing as part of \r\n? Or is this just a dig at Windows-style line endings? If there's something weird going on I think I'd rather fail loudly.
> Bounding the inner scanner to a single line makes “run past the end of a malformed line” unrepresentable rather than merely unlikely.
I don't really see what makes it "unrepresentable", and this reads more like "if you used the right scanning logic, you can't have used the wrong scanning logic".
inigyou 1 hours ago [-]
Sure, start with \r\n, split on \n, now you have a stray \r at the end of every input.
Retr0id 1 hours ago [-]
But the preceding clause says it handles \r\n. If you're already handling \r\n, what remaining sources of \r are there, that you'd actually want to silently ignore?
inigyou 40 minutes ago [-]
Someone else (possibly you) already split on \n
thesz 48 minutes ago [-]
End of line on Classic Mac is \r.
alexjurkiewicz 12 minutes ago [-]
> if (!line.accept('[').isEmpty() ) // [section] header.
Is this really ergonomic?
tomashubelbauer 21 minutes ago [-]
This post doesn't touch on something that makes parsers complicated no matter how simple the grammar: good error messages. Parsing a well formed input is the easy part, but not just spitting out a byte index but actually telling the user why their input is not good and what they could do to make it conform is super hard.
The Rust compiler is a common example of a compiler that does a good job here, and I think it is one of only a few.
jdw64 19 minutes ago [-]
I'm going to collect this post after 24 hours, extract the methodologies from everyone's comments, and write them down in my notes. The reason I like HN is that people freely share their tips in the comments
speedgoose 1 hours ago [-]
I now use nom to write my parsers. Once you understand it, it’s simple and parsing complex data becomes a _fun_ puzzle. I recommend it.
If you created a format that is so difficult to parse that it cannot be parsed with simple readable C code then the problem is the format not the parser code.
jrimbault 12 minutes ago [-]
Can you feel the irony when typing this? "Simple readable C code" itself not being able to be parsed by "simple readable C code".
Famous examples: despite so many initial good intentions, html tags don’t need to be closed, JSON numbers are too often encoded as strings, YAML can look like what most people expect or it can look progressively more like JSON… and on and on.
If what you're parsing is within the capacity of humans to interact with (so in the range of tens of kilobytes), a grammar that requires an O(N^2) parser is totally fine.
Looking at the linked URL parser, why doesn't it look like
It looks totally ad-hoc.One common way to test it is just to pass ipv6 url: http://[f021:d981:b487:e57d:193e:550e::]/
Is there a common source of extra \r in malformed inputs, beyond those existing as part of \r\n? Or is this just a dig at Windows-style line endings? If there's something weird going on I think I'd rather fail loudly.
> Bounding the inner scanner to a single line makes “run past the end of a malformed line” unrepresentable rather than merely unlikely.
I don't really see what makes it "unrepresentable", and this reads more like "if you used the right scanning logic, you can't have used the wrong scanning logic".
Is this really ergonomic?
The Rust compiler is a common example of a compiler that does a good job here, and I think it is one of only a few.
https://github.com/rust-bakery/nom