I am happy to announce some real research toward AGI on the AGI list.
James Bowery and I have been very busy since late June evaluating 5
separate Hutter prize entries plus one record setting LTCB entry.
Since 2023, the closed source program nncp by Fabrice Ballard held the
LTCB record on enwik9 with a compression ratio of .1072 as the only
transformer model, taking 3 days to compress or decompress 1 GB of
text on a GeForce RTX 3090 with 10,496 Cuda cores and 24 GB memory
using a transformer with 199M parameters.

Then last Tuesday Conrad Lippert-Zajaczkowski improved on the record
by .0003 with altxs, another transformer model that took 66 hours to
compress and 330 hours to decompress on a 40 core AMD EPYC 9v84 with
320 GB and a GH100 GPU with 94 GB under Linux 24.04.

He held the record for less than one day. That is when I announced the
5 Hutter prize entries, one of which (fx2-cmix-transformer by Vladimer
Ivanov) smashed the LTCB record with a .0969 compression ratio as well
as taking the top spot in the Hutter competition and winning EUR 37K
(7.5% improvement) if it holds up during the 30 day public comment
period, even after improving over two others winning EUR 5K each. The
announcement with some discussion is here:
https://encode.su/threads/4539-Hutter-prize-backlog

What makes this more amazing is that it runs within the Hutter prize
limits on a single thread with no GPU using under 10 GB memory within
the time limit of 70000/(Geekbench 5 score), about 50 hours on a
typical laptop. LTCB has no hardware or time limits. Furthermore the
Hutter prize counts both the compressor and decompressor size as
Windows or Linux executables, while LTCB counts only the decompressor
and allows zipped source code or executable, whichever is smaller.

The Hutter prize record of .1107 for fx2-cmix by Kaido Orav and Byron
Knoll stood since Sept. 2024 until it was beaten by 1% (the minimum
allowed) by cmix-lex by Ibrahim Marcouch using work released earlier
by Kaido Orav. When that was released on June 26, 4 new derived
entries were quickly released that improved by another 1%, all
submitted within days of each other, with the first winning 5K and the
rest nothing (pending the 30 day comment period).

fx2-cmix and all the derived entries are context mixing compressors
using PAQ like indirect context models, PPM, and an LSTM neural
network language model. The input is preprocessed to encode the
XML/HTML/Wiki markup, sort the articles by topic, and tokenize the
text using a prebuilt dictionary grouping related words together,
compressed, and passed to the decoder. The dictionary organization and
article sort order is CPU intensive so it is done offline and passed
to the compressor and decompressor, adding a size penalty to LTCB and
a double penalty for the Hutter prize because the compressor size
counts too. To stay within the memory limit, the PPM model is memory
mapped to a 14 GB temporary swap file, which slows down my tests
because Windows reserves 4 of 16 GB for itself when I open an Ubuntu
window to run the self extracting archives, leaving only 2 GB for
buffering the SSD. So my tests usually run over 50 hours. James runs
under native Linux so all of the entries have met the time limit and
qualify for the award.

fx2-cmix-transformer replaces the LSTM neural network with a 6M
parameter transformer, trained offline for 26 hours on 8 RTX 5090
GPUs. The parameters are quantized to 4 bits (-7 to 7) and compressed
to 2.9 MB and passed to both the compressor and decompressor. Even
with this 5.9 MB double penalty, it still improves over the best LSTM
Hutter prize entry by 7.5%. The description is quite readable and well
researched, and probably the best explanation I have seen of how a
transformer works.
https://drive.google.com/file/d/17c8hMeitD2kdJpxOtoNCQRmIwQeG9qKd/view

I am very excited about the possibility of running language models
locally, that learn as you use them instead of needing fixed weights
and long context windows. I created the large text benchmark 20 years
ago to accelerate language modeling research, arguing that compression
measures prediction and prediction is sufficient for passing the
Turing test. Even within the Hutter prize limits, we now have programs
that compress a lifetime's worth of learning into days on a PC.

Of the 222 compressors listed, 9 of the top 14 are Hutter prize
entries. They show a decompressor size of 0 because they are self
extracting archives. All are open source and documented, as required.
The top 3 programs are transformers, followed by LSTM, then CM, then
PPM, and LZ77.
https://mattmahoney.net/dc/text.html

-- 
-- Matt Mahoney, [email protected]

------------------------------------------
Artificial General Intelligence List: AGI
Permalink: 
https://agi.topicbox.com/groups/agi/T8131431f16a9101a-M10c29afe1dc01b08c855762a
Delivery options: https://agi.topicbox.com/groups/agi/subscription

Reply via email to