On Thu, Jul 23, 2026, 3:02 AM Quan Tesla <[email protected]> wrote:

> I'm criticizing the systems thinking about compression, not compression
> itself. IMO, the Shannon limit functions as an allegorical lighthouse for
> compression, but it's seemingly being treated as contextual noise.
>

The current system thinking is as follows. The Shannon limit applies to
known probability distributions. This is easily solved using arithmetic
coding. The Kolmogorov limit applies to unknown distributions. It is not
computable, so this is where all the research is.

Turing defined machine intelligence in 1950 as the ability to fool humans
into believing that it is human 30% of the time after 5 minutes of text
messaging, given a 50% prior probability. That happened about 2023. I
claimed in 1999 that the Turing test is equivalent to text prediction,
which can be measured by compression ratio by adding an arithmetic coder. I
created a benchmark in 2006 to evaluate text compressors that became the
basis for the Hutter prize.

Shannon estimated in 1950 that the entropy of written English relative to
human level text prediction is very roughly 1 bit per character, which is
about the level achieved by the top Hutter prize compressors and modern
LLMs.

The Hutter prize committee consists of Marcus Hutter, who is funding the
prize (5000 euros per 1% improvement) and James Bowery and I evaluating
entries. We are currently testing two entries that claim consecutive
improvements of 1% each.

James proposes using compression to settle disputes over questions of
social policy. Solomonoff induction tells us that the shortest program that
outputs past observations is the best predictor of future observations.
That is exactly what we are measuring with text compression. This seems to
me a better solution than the one he linked that proposes hand to hand
combat between sovereign citizens with a 25 cm knife and 15 meters of
strong cordage.

My question is how to implement this. I selected 1 GB of Wikipedia text for
my benchmark because that is roughly the amount of language that one can
hear and read in a lifetime. Therefore it should be sufficient for human
level intelligence. If we were to apply the same technique to evaluating
vision, we would need a few decades of uncompressed video. The retina has
137 million rods and cones in each eye and an information rate of 10 bits
per second each. This is about 300 petabytes over 30 years.

But that isn't the problem. The problem is that we can effectively compress
video by asking an AI to describe it and compress the text to about 10 bits
per second. Then you decompress by using the text to prompt the video. This
is lossy, of course, but close enough that you don't notice the difference.
The reason this works is that the human brain has a write speed of 5 to 10
bits per second, the same rate that we can read or speak.

This means that video is 1 part per billion content and the rest can be
safely discarded as noise. If we ran a lossless video benchmark then nearly
all the effort would be going into compressing the noise instead of
understanding the image. This is already a problem for the Hutter prize
where 30% of the text is synthetic or XML, HTML, and Wiki formatting whose
compression does not contribute to language understanding but is
nevertheless required to advance.

I tried to think of examples where we could answer questions about social
policy like future population. If AI can collect all human knowledge, as it
seems to be doing, then it should be able to say what is best for humanity
better than any human could. But most policy questions are about the
allocation of resources, and are ultimately resolved by combat.


-- Matt Mahoney, [email protected]

------------------------------------------
Artificial General Intelligence List: AGI
Permalink: 
https://agi.topicbox.com/groups/agi/T5b58bcc51c493d41-Me9a30d4b5dffacbab1cdfaed
Delivery options: https://agi.topicbox.com/groups/agi/subscription

Reply via email to