Since this critical conversation has been derailed by:

1) Matt bringing in the same argument about "noise" that Ben Rudiak-Gould
originally leveled against lossless compression of Wikipedia (indeed his
was the *first public comment* when the Hutter Prize was announced in 2006)
<https://groups.google.com/g/comp.compression/c/Pwlq6pkyc8s/m/ZdHC8HvgCYAJ> --
something both Matt and I addressed in 2006 with Ben and, now, I with Matt
twice now this forum before and
2) as a result, my willingness to offer an alternative to principled model
selection regarding 10s f trillions of $ conflicts of interest (if not the
fate of Earth), consistent with individual consent (ie: making it practical
to "vote with your feet")
3) that "vote with your feet" conversation isn't really proper for the AGI
list (except perhaps as an alternative to Anthropic's "Constitution")...
4) Even if it were an appropriate topic for the AGI list, it is obviously
futile due to the tl;dr problem that sets in on such topics even if the
length of the "Constitution" has an argument surface of just a few
paragraphs (note Matt's elevation of #5 to the sole dispute processing mode
despite it obviously being only a last resort -- this despite Matt being
among the brightest minds on the planet in my opinion).

I'll now attempt to get it back on track by AGAIN addressing Ben
Rudiak-Gould's 2006 argument which Matt has resurrected 20 years later for
some reason.  Charles Sinclair Smith also used this argument with me. In
Charlie's case, the argument concerned "data cleaning" as the predominant
cost of macrosocial dynamics modeling.

Since Charlie co founded the DoE's Energy Information Administration under
President Carter and focused on information quality, and since he financed
the second neural network summer of the 1980s as a result of that
experience, it's safe to say his argument represents the Steel Man case
against my proposition.

So:

   1. What exactly IS "data cleaning"?
   2. How does it address the "noise" problem in data?
   3. Why do I assert that Charlie's claim—that over 90% of the expense of
   macrosocial data curation is "data cleaning" (and that at some point
   quantity has its own quality)—is not a bug but a feature of lossless
   compression as dynamical macrosocial model selection to reformat the social
   pseudoscience?

#1) Anyone who has worked extensively with datasets knows that such trivial
things as typos during data entry can send subsequent automated analysis
into the weeds.  Indeed, I came to know Charlie because, as a consultant to
his company, I bet him that I could transform some paper insurance tables
into electronic datasets before his mid 1990s in-house bleeding-edge-OCR
experts could -- and guarantee that there would be no errors.  He took my
bet and I hired a bunch of "Kelly Girls" to do the data entry. I just hired
enough redundant human OCRs so that a simple cross-check for equality
reduced the error rate to *effectively* zero.  No subsequent work with
those datasets exposed any errors either in the transcription or in the
original paper tables that contained them.  That was a LOT of work but I
won the bet.  However, that is only one of many kinds of data integrity
issues that arise in the real world.  Charlie told me about one case where
he had to personally track down a guy in a nursing home who was responsible
for a particular measurement that didn't comport with the rest of the data
in one curation.  That guy was the sole repository of the transform used to
generate that data.  This is the kind of HUMINT forensics that goes into
"data cleaning".  One can see how this kind of expense could explode to the
90% level seen in 1970s analysis of the dynamics of the US energy economy.

#2 Data cleaning addresses the "noise" problem but only to the extent that
it corrects for obvious errors in the measurement instrument -- such
"instruments" being entire human organizations with their various standards
in some cases.  It cannot address "noise" that those standards are not even
intended to filter.  That noise is sometimes known as the bewildering world
into which we are thrown.  We are confounded by interactions manifesting as
"random noise" that we, like a dog without a bone, nevertheless try to make
sense of.

#3 When curating a dataset to bring the social pseudosciences to heel, we
are dealing with data forensics in the sense that we must be prepared to
treat the various interests proffering their data as world-class fraud
artists.  The stakes in biasing the "narrative" of the largest issues
facing us are epic if not eschatological -- and we needn't even consider
the possibility that these fraud artists are *consciously* colluding to
perpetrate their deceptions!  A great case in point is Matt's mysteriously
abysmal reading comprehension regarding Sortocracy.  Here we have one of
the brightest minds on the planet suddenly becoming incapable of
comprehending a few sentences!  I have no doubt that he is capable of
comprehending far more than a few sentences and that he had conscious
intention of doing so -- yet when it came to an issue of such intense
conflicts of interest, he acted as though he were trying to defraud the
public regarding Sortocracy!  We're all human and we all do this sort of
thing.  That's why the philosophy of science is largely about keeping us
from lying even to ourselves.  That's why Solomonoff's proofs from the
1960s were such a profound advance in the philosophy of science: They
offered, for the first time, a scientific model selection criterion that
could be used to create the right *incentives* given an agreed-on dataset.

The key word here is *incentives*.

What:

   - were the *incentives* Charlie was under to go hunt down that old man
   in a nursing home?
   - would be the *incentives* to stop wasting our time yammering at each
   other in prose about whose narrative is to "rule them all" in the sense of
   providing predictions of the consequences of policies imposed on
   non-consenting subject populations
   <https://fairchurch.org/AProtestantInstauration.pdf>?
   - would be the *incentives* to publish the scripts people used to clean
   data rather than simply declaring that they had "cleaned the data" and
   described in *prose* their "cleaning policies"?

All of these incentives can align with lossless compression as the metric
for awarding MONEY to finance quality assurance for entities like the DoE's
Energy Information Administration.  Let's take Charlie's example of the
investigation that required tracking down a guy in a nursing home:

The result of that investigation was a *computer program* that encoded the
data standard used to convert the data into a form commensurate with the
rest of that dataset.  Charlie could have published that computer program
along with the original "dirty dataset" and presented the "cleaned" dataset
as simply a tabled (encached) intermediate computation.  This cleaning
algorithm would have had a length and would have obviated any
*arguments* Charlie
might have had with competing interests regarding the notion of "clean" vs
"dirty" data.  This would make it easier to reach a consensus on curating
all data under consideration when deciding which macrosocial "narratives"
to impose on non-consenting peoples without a control group and a phase 1
trial for safety, let alone a phase 2 trial for the *efficacy* of such
"scientific" policies.

That said, I've dealt with the Steel Man argument.

However, there is also the straw man Matt set up regarding "video
compression" which has two aspects:

1) The idea that human perception is the standard for defining actual data
content of video data -- hence lossless compression is of no use in
selecting the best model is beside the point.  I'm not proposing to have a
bunch of Kelly Girls look at a bunch of macrosocial data and decide what
they *feel* is important to predicting the consequence of policies.  This
standard simply does not pertain to the definition of "noise" for
macrosocial model selection.  If it has any relevance at all, it is at the
*data* selection stage rather than at the *model* selection stage.
2) The idea that the  Ben Rudiak-Gould notion of "noise", while invalid for
Wikipedia (as Matt correctly argued in response to Ben), is *valid* for
video data is obviously false since there are papers showing enormous gains
in video compression based on 4D model induction
<https://share.gemini.google/D6o8W3zXEtph>.

This is all beside the point that even if these arguments were valid for
data class X they would be valid for dataclass Y when the two aren't
comparable in either orders of magnitude of quantity nor orders of
magnitude of quality control -- as is the case of comparing video data to
macrosocial data.

Matt may even be correct that the lossy compression of the total text
content of the Internet may have already produced AIs that would be
superior to the governments under which we now suffer.  I'm all for
enabling him to join together with consenting adults of like mind to *run
their experiment on themselves* while we, who don't share their *faith* may
pursue our own confessions.

But so long as we're all suffering under *the one true church of social
pseudoscience*, it is rather inhumane to deny us a means of holding that
theocracy to its own proclaimed standards.

On Thu, Jul 23, 2026 at 6:32 PM Matt Mahoney <[email protected]>
wrote:

> ...But that isn't the problem. The problem is that we can effectively
> compress video by asking an AI to describe it and compress the text to
> about 10 bits per second. Then you decompress by using the text to prompt
> the video. This is lossy, of course, but close enough that you don't notice
> the difference. The reason this works is that the human brain has a write
> speed of 5 to 10 bits per second, the same rate that we can read or speak.
>
> This means that video is 1 part per billion content and the rest can be
> safely discarded as noise. If we ran a lossless video benchmark then nearly
> all the effort would be going into compressing the noise instead of
> understanding the image. This is already a problem for the Hutter prize
> where 30% of the text is synthetic or XML, HTML, and Wiki formatting whose
> compression does not contribute to language understanding but is
> nevertheless required to advance.
>
> I tried to think of examples where we could answer questions about social
> policy like future population. If AI can collect all human knowledge, as it
> seems to be doing, then it should be able to say what is best for humanity
> better than any human could. But most policy questions are about the
> allocation of resources, and are ultimately resolved by combat.
>
>
> -- Matt Mahoney, [email protected]
>
>
> *Artificial General Intelligence List <https://agi.topicbox.com/latest>*
> / AGI / see discussions <https://agi.topicbox.com/groups/agi> +
> participants <https://agi.topicbox.com/groups/agi/members> +
> delivery options <https://agi.topicbox.com/groups/agi/subscription>
> Permalink
> <https://agi.topicbox.com/groups/agi/T5b58bcc51c493d41-Me9a30d4b5dffacbab1cdfaed>
>

------------------------------------------
Artificial General Intelligence List: AGI
Permalink: 
https://agi.topicbox.com/groups/agi/T5b58bcc51c493d41-M59e813cd830043a853984f9e
Delivery options: https://agi.topicbox.com/groups/agi/subscription

Reply via email to