Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- I wonder how this would look in Chinese Mandarin.by kittikitti
- I typed the first stanza of 'Jabberwocky'.
Worked about as well as expected :D
by Applejinx - The kerning is absolutely awful in Safari. Looks fine in Chrome though...by LoganDark
- Now I'm considering how many tokens I've wasted writing "tihs" instead of "this" - tens! Any how much extra work goes into understanding misspellings?by ascots
- Haha this did help me generate empathy for the assistant. Neat idea.by nxtfari
- This made me realize that my usual practice of typing at 80% accuracy, full of mistakes, when sending input to an llm is probably increasing my input token count.by totetsu
- This page reliably hangs Firefox 155 at 100% CPU for me.by dTal
- I typed some German and it was breaking up words so much more than English. Not really surprising given tokenizers are optimized for most commonly used text.
Here's the token efficiency of a corpus translated into various languages and tokenized with the latest OpenAI one:
Source: "Tokenizer Fairness in 2026", a reproduction/extension of Petrov, La Malfa, Torr & Bibi, "Language Model Tokenizers Introduce Unfairness Between Languages" (NeurIPS 2023), using FLORES-200.Language Relative tokens -------------------------------------- English 1.00x Portuguese 1.23x Chinese (Simplified) 1.25x German 1.31x Spanish 1.32x French 1.37x Arabic 1.38x Chinese (Traditional) 1.42x Korean 1.47x Swahili 1.49x Hindi 1.57x Japanese 1.66x Burmese 3.16x Amharic 5.78x Santali 13.70xby croemer