There is about 15 TB of freely downloadable text on the Internet for
training LLMs. That compresses to 10^13 bits. In my 2008 proposal I
estimated 10^14 bits based on the number of web pages returned by Google
when it still gave that information. That was a time when everything on the
Internet was free, uncensored, and available worldwide. Now almost
everything is pay walled or behind a login. Social media consisted of
USENET and mailing lists that usually had public archive servers. Now if
you want that data you have to own a social media company.

My estimate of 10^17 bits of human knowledge to automate all labor is based
on 10^9 bits of human long term memory times 10^10 people times 1% of your
knowledge that isn't known to anyone else. The 1% comes from US Labor
department estimates that it costs on average a year's income to replace an
employee depending on the job. Higher paying jobs take longer. A robot
still has to be trained by the humans it is replacing just like any other
new hire, which means transferring that knowledge at 5 to 10 bits per
second through speech and writing. If it wasn't for this cost, AI would
have already taken your job by now.

As for model selection criteria, there are a lot of them for the same
reason there are lots of different compression algorithms. The general
problem is not computable.

-- Matt Mahoney, [email protected]



On Wed, Jun 10, 2026, 11:50 AM James Bowery <[email protected]> wrote:

>
>
> On Tue, Jun 9, 2026 at 10:12 PM Matt Mahoney <[email protected]>
> wrote:
>
>> ...
>>
> 3. Ben's AGI will never be built for the same reason my 2008 distributed,
>> open source, and public data AGI proposal (
>> https://mattmahoney.net/agi2.html ) will never be built.
>>
>
> 10^9 seems well supported by your website's essays, but not 10^17 or
> 10^13.
>
> Those last two orders of magnitude are highly suspect given that only in
> the last year have figures like Illya (among The Great and The Good)
> started taking parameter compression as seriously as they should have since
> at least 2006.  That's 20 years of wandering in the wilderness devoid of
> Solomonoff.
>
> In actuality, what has really been going on is studious ignorance due to
> the tens of trillions of dollars at stake in maintaining macrosocial models
> swimming around in a pool of stagnant information criteria for model
> selection.  The incentives to control "The Narrative" are so great that no
> one can claim immunity from such conflicts of interest in maintaining
> sociology in the dark ages of statistical information criteria
> <https://en.wikipedia.org/wiki/Model_selection#Criteria>.  No one,
> especially those ensconced in academia where these conflicts of interest
> are most focused, can think rationally about this.
> *Artificial General Intelligence List <https://agi.topicbox.com/latest>*
> / AGI / see discussions <https://agi.topicbox.com/groups/agi> +
> participants <https://agi.topicbox.com/groups/agi/members> +
> delivery options <https://agi.topicbox.com/groups/agi/subscription>
> Permalink
> <https://agi.topicbox.com/groups/agi/T5b182c3fd14f61e9-Mcdf9a5766ef6dabebb602889>
>

------------------------------------------
Artificial General Intelligence List: AGI
Permalink: 
https://agi.topicbox.com/groups/agi/T5b182c3fd14f61e9-M85df3f9895d0f3fc276900f9
Delivery options: https://agi.topicbox.com/groups/agi/subscription

Reply via email to