There is about 15 TB of freely downloadable text on the Internet for training LLMs. That compresses to 10^13 bits. In my 2008 proposal I estimated 10^14 bits based on the number of web pages returned by Google when it still gave that information. That was a time when everything on the Internet was free, uncensored, and available worldwide. Now almost everything is pay walled or behind a login. Social media consisted of USENET and mailing lists that usually had public archive servers. Now if you want that data you have to own a social media company.
My estimate of 10^17 bits of human knowledge to automate all labor is based on 10^9 bits of human long term memory times 10^10 people times 1% of your knowledge that isn't known to anyone else. The 1% comes from US Labor department estimates that it costs on average a year's income to replace an employee depending on the job. Higher paying jobs take longer. A robot still has to be trained by the humans it is replacing just like any other new hire, which means transferring that knowledge at 5 to 10 bits per second through speech and writing. If it wasn't for this cost, AI would have already taken your job by now. As for model selection criteria, there are a lot of them for the same reason there are lots of different compression algorithms. The general problem is not computable. -- Matt Mahoney, [email protected] On Wed, Jun 10, 2026, 11:50 AM James Bowery <[email protected]> wrote: > > > On Tue, Jun 9, 2026 at 10:12 PM Matt Mahoney <[email protected]> > wrote: > >> ... >> > 3. Ben's AGI will never be built for the same reason my 2008 distributed, >> open source, and public data AGI proposal ( >> https://mattmahoney.net/agi2.html ) will never be built. >> > > 10^9 seems well supported by your website's essays, but not 10^17 or > 10^13. > > Those last two orders of magnitude are highly suspect given that only in > the last year have figures like Illya (among The Great and The Good) > started taking parameter compression as seriously as they should have since > at least 2006. That's 20 years of wandering in the wilderness devoid of > Solomonoff. > > In actuality, what has really been going on is studious ignorance due to > the tens of trillions of dollars at stake in maintaining macrosocial models > swimming around in a pool of stagnant information criteria for model > selection. The incentives to control "The Narrative" are so great that no > one can claim immunity from such conflicts of interest in maintaining > sociology in the dark ages of statistical information criteria > <https://en.wikipedia.org/wiki/Model_selection#Criteria>. No one, > especially those ensconced in academia where these conflicts of interest > are most focused, can think rationally about this. > *Artificial General Intelligence List <https://agi.topicbox.com/latest>* > / AGI / see discussions <https://agi.topicbox.com/groups/agi> + > participants <https://agi.topicbox.com/groups/agi/members> + > delivery options <https://agi.topicbox.com/groups/agi/subscription> > Permalink > <https://agi.topicbox.com/groups/agi/T5b182c3fd14f61e9-Mcdf9a5766ef6dabebb602889> > ------------------------------------------ Artificial General Intelligence List: AGI Permalink: https://agi.topicbox.com/groups/agi/T5b182c3fd14f61e9-M85df3f9895d0f3fc276900f9 Delivery options: https://agi.topicbox.com/groups/agi/subscription
