Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Perl and eval, not what I'd have expected from the title!


Indeed! Arbitrary text was smuggled into the eval() statement through a weak input sanitization step:

The second quote was not escaped because in the regex $tok =~ /(\\+)$/ the $ will match the end of a string, but also match before a newline at the end of a string, so the code thinks that the quote is being escaped when it’s escaping the newline.

This is why I hate using regexes, or deciphering other people's code which uses them. The syntax just isn't obvious. We need tooling that removes mental overhead, not adds to it. I've always found plain old procedural parsing code easier to formulate, trace and reason about.

I'm not some some fanatic preaching they should be banned from all languages, but I do feel there are places they're inappropriate and input sanitization is one of them. It's also good practice to centralize this sort of logic in some flavor of escape() function to reduce the whack-a-mole when bugs like this are found.


> "will match the end of a string, but also match before a newline at the end of a string."

Is it just me, or is this a poor description? I understood $ to be an anchor for end of line, not end of string, i.e. "match before a newline anywhere". The emphasis "newline at the end of a string" feels misleading; the exploit works because the matched newline isn't at the end of the string and there is exploit code after it.

> "I've always found plain old procedural parsing code easier to formulate, trace and reason about."

A regex is a domain specific language which lets you express a lot of computation in a small amount of code. At the extreme, you're saying "I can write a regex engine easier than I can write a regex". "I can write assembler easier than I can write C". "I can write a list of a thousand items quicker than I can write range(1000)". It doesn't make sense, and for anything more than trivial cases, I don't believe it.

(I'll agree that there are places they are inappropriate, expressions which aren't clear, and maybe the extra effort of doing it by hand is worth it in security sensitive situations, but I claim it is extra effort and less clear what a pile of ifs/loops/switch is trying to achieve vs "[a-f]{3}(\d)" or whatever).


The meaning of ^ and $ in perl regexes can be altered by modifiers at the end of the regex. A 'multi-line' regex, with a /m at the end, makes them match the start and end of any line.


The code doesn't use /m. So indeed the problem is that $ will match before \n at the end of the string the Perl interpreter is working with, which is not the whole thing containing the payload.


I've just worked through it to understand it; it's a more subtle exploit than I thought. My above comment about $ being end-of-line is wrong, the above comment saying you need the /m modifier for that to happen is correct.

In case anyone cares, the bug matches <backslash> <newline> <quote> then the regex in question matches <backslash> <newline> and the code logic is that one backslash must be escaping the next character from the input, but instead of adding on the next character it always adds a quote - after searching for a quote the next character must be a quote, right? But it wasn't, it was a newline, which was missed because of $ behaviour with newline at the end of a string. That acts to shift the quote one to the left <backslash> <quote> <newline> and now that makes an escaped quote which won't break eval() and the loop carries on and reads in the exploit text up to the real end quote.

----

Details: the Perl code takes this pattern in the input (the quotes are part of the input, in the file data):

    "a\
    ""
six characters, describing a quoted string of four characters: <a> <backslash> <newline> <quote>

The Perl code finds the string starting quote and moves past it, and sets $tok empty to hold the quoted string content. Then it searches for the next quote (not the last, the next):

    last Tok unless $$dataPt =~ /"/sg;
    
This will match at the quote after the newline. Then it substrings from the saved opening quote position to 1 before the found quote. So the substring includes the newline char:

    # the closing quote position. (not including the quote).
    $tok .= substr($$dataPt, $pos, pos($$dataPt)-1-$pos);
That gets <a> <backslash> <newline> and not the following <quote> <quote>

Then it does the odd-number-of-backslashes test on the substring <a> <backslash> <newline>:

    last unless $tok =~ /(\\+)$/ and length($1) & 0x01;
And the regex matches for <backslash> <newline> instead of the intended <backslash> ENDOFSTRING so the code thinks there is an escaped character. It doesn't add the next character from the string into the token, it assumes the escaped character must be a quote and always adds a quote to $tok.

    $tok .= '"';    # quote is part of the string
Effectively shifting the input quote to the other side of the newline, from:

    "a\
    ""
to

    "a\"
    "
and in the exploit case:

    "a\
    "exploit code"
to

    "a\"
    exploit code"
Then the inner loop runs again and finds the <exploit code to closing quote> text and adds that on to $tok. Now there's a string with an escaped quote and some exploit code.




Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: