

Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- ASCII was amazingly efficient for conveying English text. Then we needed to encode multi languages and emojis, the resulting Unicode is just a mess.by anvuong
- The more easily you can access meaning from symbols without intermediate tools, the more durable and resilient.
Symbols on paper is best.
Plain text in a file that can be opened by any computer comes close, but needs a tool.
Fancy file formats that need not only hardware but also special software are the worst.
However, there's an important tradeoff, which is the fancier formats can present information in ways mere symbols might struggle with, and can also use interactivity to improve understanding.
by andsoitis - I love plain text, but the calculus of being able to access the content in 50 years is much more interesting in the age of agents. They can infer the meaning of structured data without schemas, decompressed archives, find embedded files, etc. It's not perfect, but the durability of binary formats is better now than it has ever been.
It's not baked-in to the weights of any model, but agents can write their own tools to work with arbitrary binary formats and get many of the benefits of off-the-shelf unix utilities.
I love that the Godot game engine has a textual scene description. That's a stark contrast to Unreal Engine's binary format for blueprints (visual scripting).
by shmolyneaux - Fun timing on this plaintext conversation. I built a little browser based plaintext playwriting app this weekend for a little weekend project. Found myself really hating 1. how clunk screenwriting software can be and 2. How unsharable the files are.
With plaintext you can hand the file off to anyone on any device--the caveat being absolutely no one wants to be handed a plaintext script. The software I used ten years ago to write plays is long since deprecated and those files are basically unopenable. Plaintext however remains.
by spcebar - What, no mention of our old friend, ASCII-armored Base64? For shame! Everything can be text with Base64, including things that have absolutely no business being text! Best of all, in light of popular widespread abuse of every available resource, Base64 is comparatively efficient! Bring on the petabytes! Yay and I'm not being completely sarcasticby kerblang
- This is the public dividend of a standard finally winning. The article gives credit to Unicode, but it is the fact that ASCII unambiguously won that gives plain text its portability and longevity. It looks like Unicode is on its way to winning in the same way, but it is not there yet. Most text files I write are still pure ASCII because that's the only way to avoid unexpected glitches [0].
[0] Windows newlines not withstanding.
by drhagen - The article briefly addresses the problem, but it's pretty fun how different "plain text" looks throughout history and in different domains.
For one, there are a few different ways to terminate lines. All major operating systems now tend to use just \n, but I have older files that use \r\n (Microsoft), \r (Macintosh) or \n\r (RiscOS).
There are also different opinions on how text files end in different operating systems. In POSIX, all lines are terminated by \n, even the last one. Microsoft software still tends to insist that the last line of a file is special case that doesn't need to be terminated even now that they have otherwise adopted POSIX style line endings. In Microsoft's view, it seems that the line ending sequence separates lines rather than terminate them. Files created according to this view don't play well with tools like cat(1) if your intent is to concatenate the lines of two files, but it seems other Unix clone tools have adapted to the possibility that the last line isn't terminated properly.
Finally there's the encoding problem. I don't know of a good tool that determines the original encoding based on heuristics and re-encodes to UTF-8 but if someone does I'd love to know. If I know that the input language is English for example it shouldn't be too hard to determine what encoding the funny byte used in contractions or the funny bytes used in quotes belong to. Still, in English most of the files that use 8-bit encodings remain quite readable if you just box out the invalid bytes.
by boomlinde - Graydon Hoare, the creator of the Rust programming language, wrote his seminal piece on text in 2014. In it, he said "text is the most powerful, useful, effective communication technology ever, period."[1] Text is durable.
[1] https://archive.ph/FhG5L (the original either got deleted or login-walled, here is the archived version)
[2] https://hn.algolia.com/?q=always+bet+on+text
[3] https://news.ycombinator.com/item?id=26164001
[4] https://news.ycombinator.com/item?id=8451271
by devy