The article briefly addresses the problem, but it's pretty fun how different "plain text" looks throughout history and in different domains.
For one, there are a few different ways to terminate lines. All major operating systems now tend to use just \n, but I have older files that use \r\n (Microsoft), \r (Macintosh) or \n\r (RiscOS).
There are also different opinions on how text files end in different operating systems. In POSIX, all lines are terminated by \n, even the last one. Microsoft software still tends to insist that the last line of a file is special case that doesn't need to be terminated even now that they have otherwise adopted POSIX style line endings. In Microsoft's view, it seems that the line ending sequence separates lines rather than terminate them. Files created according to this view don't play well with tools like cat(1) if your intent is to concatenate the lines of two files, but it seems other Unix clone tools have adapted to the possibility that the last line isn't terminated properly.
Finally there's the encoding problem. I don't know of a good tool that determines the original encoding based on heuristics and re-encodes to UTF-8 but if someone does I'd love to know. If I know that the input language is English for example it shouldn't be too hard to determine what encoding the funny byte used in contractions or the funny bytes used in quotes belong to. Still, in English most of the files that use 8-bit encodings remain quite readable if you just box out the invalid bytes.
I'm surprised by the existence of \n\r line terminations!
\r\n was how lines were terminated on the wire for many early devices (for example, the Teletype ASR-33 seen in the famous photograph of Thompson and Ritchie at Bell Labs).
If the carriage is very far to the right, and an LF is sent followed by the CR, then the carriage will still be returning when the next character arrives, which will print at some random location along the return path.
CR followed by the mechanically much faster LF operation gives enough time for the carriage to be fully homed and in proper position for the printable character to be struck.
Text is wonderful, but note that the one keyword not found in this text is 'parse'. How much extra work is generated writing parers for text that would have been so much easier to deal with if it had just been generated in a structured binary format? You still get the pleasure of designing a new wheel with every format.
It's not about writing a parser, it's about figuring out what the content is supposed to be in the first place. With plain text you can see what's a number, what's a string, what's an expression, etc... I'd rather try to figure out how to parse "2 7 + 3 -" than see a file with
If it were easy to see what's a number and whats a string, type confusions a la YAML's fun wouldn't exist:
Is "no" a string, a country, a boolean?
Is 0 a boolean, a number, or char[48]?
This is specifically an area that "text" is quite bad and is the source of much woe. Your provided example would be parsed by most parsers today as a string, since it mixes spaces (non-numeric) and surrounds it with quotes.
In fact, given well-defined binary, it is rather simpler to turn it into "text" than the reverse. It seems like you've forgotten that "0x32 0x20 0x37 0x20 0x2A..." is literally what is in that file you would view. Plain text is, after all, just one example of a binary format.
Unless it's a different encoding from 50 years ago so you can't figure out how to parse 27+3
And you can easily make the opposite example: you'd rather figure out that .ext means format X that you already have an app for, where date is unambiguous courtesy of the format designers rather that make some common mistake trying to figure out where month is in 5/6/79
It’s not a consistent change, but for example programs like notepad couldn’t handle \n, but do accept it now. Visual Studio expects certain file types to use \r\n, but accepts others like .gitignore using \n.
Same thing for \ and / in path separators: Microsoft never made a consistent push for it, and it depends on the application.
I'm not sure I would call that "adoption", so much as "compatibility with", since the default line ending produced by Windows software is still almost always \r\n for files created from scratch.
FWIW I think it's a fair point. My view of this is as an outsider only seeing what comes in for review from Windows-using colleagues so maybe I overstated the level of adoption.
For one, there are a few different ways to terminate lines. All major operating systems now tend to use just \n, but I have older files that use \r\n (Microsoft), \r (Macintosh) or \n\r (RiscOS).
There are also different opinions on how text files end in different operating systems. In POSIX, all lines are terminated by \n, even the last one. Microsoft software still tends to insist that the last line of a file is special case that doesn't need to be terminated even now that they have otherwise adopted POSIX style line endings. In Microsoft's view, it seems that the line ending sequence separates lines rather than terminate them. Files created according to this view don't play well with tools like cat(1) if your intent is to concatenate the lines of two files, but it seems other Unix clone tools have adapted to the possibility that the last line isn't terminated properly.
Finally there's the encoding problem. I don't know of a good tool that determines the original encoding based on heuristics and re-encodes to UTF-8 but if someone does I'd love to know. If I know that the input language is English for example it shouldn't be too hard to determine what encoding the funny byte used in contractions or the funny bytes used in quotes belong to. Still, in English most of the files that use 8-bit encodings remain quite readable if you just box out the invalid bytes.