Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • Consider:

    If you want to record the motion of the planets, naively you have large tables of positions.

    To compress that, you may smoothly interpolate sparse positions.

    To compress that, you encode the laws of gravity and simulate from a starting state.

    Compression is literally understanding.

  • There is Compression done by Prediction by partial matching [0]

    There is the Kolmogorov Complexity [1], Normalized Information Distance [2] and Normalized compression distance [3] that correlates those.

    Finally, there's the Pre-Big Bang Informational Compression and the Delayed Release of Antimatter [4]

    All big {rabbit/black} holes to lose some time, if you have any.

    [0] https://en.wikipedia.org/wiki/Prediction_by_partial_matching

    [1] https://en.wikipedia.org/wiki/Kolmogorov_complexity

    [2] https://homepages.cwi.nl/~paulv/papers/chapter08.pdf

    [3] https://en.wikipedia.org/wiki/Normalized_compression_distanc...

    [4] https://philarchive.org/rec/GREPBI

  • Nope; there is a bit more nuance and the distinction is important.

    Compression is functionally equivalent to prediction when the data distribution is exactly representative of all future problems. The story changes drastically if you want generalization -- because the test distribution could be arbitrarily different, even if it had the same support! Eg: you observe a rare edge case in your training data and (lossy) compression could simply ignore it. But if you wanted generalization in that particular part of the space -- either because an adversary was testing you, or for design freedom where you choose to build in that specific corner -- then you don't just want data compression, but good prediction performance on a test distribution which peaks in that corner.

    Assuming that the training data distribution is exactly the distribution you will ever care for is implicitly doing a lot of the heavy lifting in the claim that compression = prediction, and I'm peeved at how much this statement is unthinkingly repeated like a manifesto.

    There is nothing natural about the training data distribution, especially if the data generation process is exploratory while the downstream usage will be exploitative.

  • The page source appears to contain all the actual text within <p> tags, but structured in a completely illogical way. With JavaScript disabled, there are a bunch of shaded bars where the text should appear, which look like placeholders for something that hasn't loaded yet even though it was there from the beginning. The <p> tags don't even seem to show up in the DOM. (I didn't check closely, but maybe they're embedded in an inline script.)

    This is actively user-hostile. The site is going out of its way to interfere with the most basic possible function of HTML, i.e., the presentation of minimally marked-up plain text. The needless complexity is especially ironic in the context of an article about compression.

  • Intuitively, the idea makes sense to me. You can only compress something when you reduce the content to “what matters” in it. And understanding “what matters” is to understand the patterns in the data. Understanding the patterns in the data IS intelligence.

    There's an important consequence here which I take as a lesson in life and business: it is worth optimizing a process or a workflow in your life or business even when there’s no obvious economic benefit. Because to optimize it is the only way to truly understand it. I am very wary of businesses and software that don’t optimize for performance (not just for profit) because it signals they don’t understand what they are doing. Slow software is poorly understood software. Fast software is also likely to be bug-free and secure because someone understands it.

  • The article is using probability where it really means proportion and prediction where it means evaluation. The mathematical equivalency is both much less surprising and less revealing once reframed.

    If we consider the first example with the arithmetic code, the initial presupposition that only the characters A, B, and C appear in the string already reduces the entropy from 56 ascii bits to 14 bits (A vs Not A and B vs Not B for each character). If you further consider that you only need to distinguish B vs Not B if it's not A, then you can just represent As with a single zero bit and only represent the non-As as two bits (the first of which will necessarily always be a 1 bit). This gets you to 10 bits without even having the proportions of the string. Of course this would be a poor convention if there were say only a single A; in that worst case scenario you would need 13 bits, but simply knowing which character appears the most, without knowing by how much, 11 bits is the worst case scenario for a length 7 string with 3 potential characters. The last bit can be made implicit if you further choose the second conditional appropriately - i.e. if instead of B vs Not B we chose C vs Not C, our last bit would be zero and could simply be dropped meaning both 10 and a single 1 bit encode C - allowing you to encode the example string in just 9 bits and an arbitrary string of that length in 10, again regardless of proportions. That improvement over the arithmetic encoding result in the example is just a case of us cramming a little extra information into the encoding algorithm.

    Arithmetic encoding is more clean and more easily extensible, it makes more sense to use than this custom encoding of 7 trits to binary but the point is the "probability" the article mentions is a superficial quality of life feature, not the secret sauce that is the actual key to compression.

  • Grant Sanderson has an excellent video on the same topic [0]. It's part of a series that is ongoing.

    [0] Compression is Intelligence Part 1 - https://youtu.be/l6DKRf-fAAM?si=yyLWq8x4sSRkWd98

  • This is the thesis behind the "Information Theory, Inference, and Learning Algorithms" course that was taught at Cambridge University.

    > Why unify information theory and machine learning? Because they are two sides of the same coin. In the 1960s, a single field, cybernetics, was populated by information theorists, computer scientists, and neuroscientists, all studying common problems. Information theory and machine learning still belong together. Brains are the ultimate compression and communication systems. And the state-of-the-art algorithms for both data compression and error-correcting codes use the same tools as machine learning.

    Book (creative commons): https://www.inference.org.uk/mackay/itila/book.html

    Lectures: https://m.youtube.com/playlist?list=PLruBu5BI5n4aFpG32iMbdWo...

Explore Birbla archives