On Fri, Sep 11, 2026 at 10:53 PM Matt Mahoney <[email protected]>
wrote:

>
> On Fri, Sep 11, 2026, 5:17 PM James Bowery <[email protected]> wrote:
>
>> On Fri, Sep 11, 2026 at 3:27 PM Matt Mahoney <[email protected]>
>> wrote:
>>
>>> On Fri, Sep 11, 2026, 2:29 PM James Bowery <[email protected]> wrote:
>>>
>>>> On Fri, Sep 11, 2026 at 12:27 PM Matt Mahoney <[email protected]>
>>>> wrote:...
>>>>
>>>>> ...Inference takes a lot less compute than training. We know this
>>>>> works, because that's how the Hutter prize leader does it.
>>>>>
>>>>
>>>> I view Vladimir's winning entry the same way I view Kolmogorov's
>>>> failure to include directed *cyclic* graphs of universal gates in his
>>>> measure of information complexity
>>>> <https://claude.ai/share/f5179150-9eab-470d-9cdc-6a9754bdc38e>:
>>>>
>>>
>>> He did? I am pretty sure that Kolmogorov complexity applies to any
>>> Turing complete language. Aren't cyclic graphs of universal gates Turing
>>> complete?
>>>
>>>
>> Finite state machines are not Turing complete.  My NiNOR complexity
>> argument <https://jimbowery.blogspot.com/2023/10/ninor-complexity.html>
>> is based on the notion that one can maintain a finite fiction of Turing
>> completeness without being Turing complete (ie:  infinite tape).
>>
>
> It's true that the shortest Turing machine that outputs a given finite
> string must halt after finite time and write to finite memory, so there is
> an equivalent finite state machine.
>
> However Kolmogorov's proof of language independence no longer works. For
> Turing machines, any language can be translated into any other language by
> a fixed length program, but that program might use an arbitrary amount of
> memory.
>

>From the NiNOR essay:
> Now we still have a free parameter but is it a choice of Turing machine?
No.  It's an integer, the size of which goes up only as the log2 of the
amount of memory required to expand the executable archive.

The point is that if you can reduce the data-dependent prior to an integer
"tape" size, you may have made some progress in defining algorithmic
information.


> For example there is no finite state machine that can translate unary to
> binary or vice versa because the input and output can be arbitrarily long.
> .
>

The relevance to Kolmogorov Complexity is a stretch since there is no
invariance theorem -- no FSM can emulate any other FSM.  So what?  The
point of the NiNOR Complexity approach is a better analytic foundation for
Solomonoff Induction -- one that admits the underlying hardware must be
priced in bits rather than swept under the rug as does Kolmogorov
Complexity.

Anyway my statement was about the efficiency of offline training of
>>> transformer weights.
>>>
>>
>> You're speaking of computational efficiency as opposed to descriptive
>> efficiency.
>>
>> That's true but the Hutter Prize includes the size of the compressor
>> along with the size of the description of the sample of human knowledge.
>> Vladimir has to pay for those pretrained weights *twice*
>>
>
> I disagree with adding the size of the compressor because it has nothing
> to do with Kolmogorov complexity. But Marcus Hutter is funding it so he
> makes the rules.
>

But does have to do with the program that produces the KC program aka
machine learning.

Anyway, it should not be hard to modify fx2-cmix-transformer to train
> online to remove the 2.9 MB penalty and compress enwik9 to 94.1 GB in 8-10
> days with one GPU.
>

This gets back to the "choice" of prior for machine learning and its
measurement.  As I've suggested before, this can be accomplished by writing
a concise "C" emulator of the GPU and using its size as a proxy without
requiring that the execution be on a single CPU core.  That's what I meant
regarding the modified LTCB that includes compressor size.



> I expect to see even further improvements by growing the transformer from
> 6M parameters to maybe 50M.
>

And that will be interesting in itself for two reasons:  1) Why would 50M
be the optimum?  2) Is there nothing to be learned here regarding the role
evolution plays in establishing priors? of



> But what I would really like to see is a faster, one pass training
> algorithm that we know must exist because that's how our brains work. Our
> brains don't do back propagation. They are recurrent and use Hebb's rule.
> What are we missing?
>

Agreed on the importance of the question but there is nothing about a
general purpose instruction set that biases away from exploring that class
of algorithm -- unlike GPUs.



> *Artificial General Intelligence List <https://agi.topicbox.com/latest>*
> / AGI / see discussions <https://agi.topicbox.com/groups/agi> +
> participants <https://agi.topicbox.com/groups/agi/members> +
> delivery options <https://agi.topicbox.com/groups/agi/subscription>
> Permalink
> <https://agi.topicbox.com/groups/agi/T3f8115622f860785-M621f9cbb7422634a5edf7619>
>

------------------------------------------
Artificial General Intelligence List: AGI
Permalink: 
https://agi.topicbox.com/groups/agi/T3f8115622f860785-Mbb4c3de92b82fbfc6ccfb973
Delivery options: https://agi.topicbox.com/groups/agi/subscription

Reply via email to