On Sat, Sep 12, 2026 at 9:44 AM James Bowery <[email protected]> wrote:
> From the NiNOR essay:
> > Now we still have a free parameter but is it a choice of Turing machine?  
> > No.  It's an integer, the size of which goes up only as the log2 of the 
> > amount of memory required to expand the executable archive.
>
> The point is that if you can reduce the data-dependent prior to an integer 
> "tape" size, you may have made some progress in defining algorithmic 
> information.

I suppose that would work, but all this does is fix the language used
to estimate Kolmogorov complexity. It needs more specification. How do
you represent the description of the gate logic? Is it a list of gates
in numerical order, each followed by a list of inputs? How do you
represent the numbers? Can you compress it, like to use macros to
describe repeated logic like adders or registers?

>> I expect to see even further improvements by growing the transformer from 6M 
>> parameters to maybe 50M.
>
> And that will be interesting in itself for two reasons:  1) Why would 50M be 
> the optimum?  2) Is there nothing to be learned here regarding the role 
> evolution plays in establishing priors? of

>From information theory, a neural network should have one parameter
per compressed bit of training data, the minimum needed to reproduce
the data without overfitting. The top LLMs use 5-10 trillion
parameters on 20-30 TB of text, which is in the same ballpark.

BTW the LTCB has a new leader. https://mattmahoney.net/dc/text.html
RATA-CMIX is a derivation of fx2-cmix-transformer that follows Hutter
prize limits but is not a submission because the improvement is less
than 1%. It uses the same transformer weights but improves their
compression. It also adds some interesting modeling improvements.

Also, tufazip is an open source derivation of nncp, which is closed
source and held the top spot for 2 years.

The top 5 entries all use transformers. I still believe that there are
better algorithms, because we know that the human brain learns in a
single pass. Or maybe it doesn't. I hope that the LTCB or Hutter prize
will either motivate its discovery or explain why the brain needs
10^14 to 10^15 synapses to represent 10^9 bits of long term memory and
LLMs don't have this limitation.

-- 
-- Matt Mahoney, [email protected]

------------------------------------------
Artificial General Intelligence List: AGI
Permalink: 
https://agi.topicbox.com/groups/agi/T3f8115622f860785-M39531654a5fcfdac1d24ab12
Delivery options: https://agi.topicbox.com/groups/agi/subscription

Reply via email to