Anthropic developed a process to see what an AI is "thinking" about. They found that when Claude is trained on text to predict tokens using a deep neural network, that a short term memory emerges that is remarkably similar to the way humans think and reason. Anthropic calls this global workspace the Jacobian or J space, a set of neurons for each word that makes that word more likely to be output.
For example, if I asked you the color of the fourth planet, you think "Mars" and answer "red". When Claude does this, neurons corresponding to "Mars" activate. Anthropic can read and modify these activations. For example they can turn off "Mars" and activate "Earth" and Claude answers "blue". They can use this method as a lie detector for safety testing because words like "fake" or "deception" will activate. The J space is used in less than 10% of processing. If the J space is turned off, then Claude can converse mostly normally but it can't solve problems with intermediate steps and it cannot lie because that requires thinking about what you really mean without saying it. The J space models the global workspace, the part of the brain that forms conscious thoughts that communicates with all the other vital but independent, unconscious processes like breathing, low level visual feature detection, and controlling our 600 muscles in the right sequence when we walk without thinking about it. It raises the question of whether conscious thoughts require language, or are babies and animals conscious as well. Anthropic did not design the J space. It emerged from training. A Jacobian is the technique they use to detect it. Mathematically a Jacobian is a matrix of partial derivatives of a vector of functions over vectors. The paper introduction didn't go into details but I believe that Claude uses a deep feed forward neural network (~100 layers) trained by back propagation with multiple passes. Then the weights are frozen to prevent leaking information between users. This requires a large context window because after every question, Claude forgets the whole conversation and has to play it back. The workspace emerges from the time delays going through a deep network and a separate scratchpad memory holding a few dozen tokens. But human brains don't work that way. Back propagation has no biologically plausible mechanism. Instead, learning is local using Hebb's rule. Brains have feedback loops, lateral inhibition, fatigue and time delays on the order of 0.1 to 1 second. Our short term memory is only about 5 to 9 tokens, about half that of chimpanzees. Hebb's rule only works one layer at a time, but natural language is structured to be learned that way. We learn to recognize phonemes and segment speech by 10 months before we learn our first word, then frozen by age 6. In 2000 I found that you can restore deleted spaces in text using only n-gram statistics with up to 77% accuracy for n = 5. This models the tokenization process, which is hard coded in LLMs. The introduction to the paper is here. https://www.anthropic.com/research/global-workspace Thank to James Bowery posting this link in the Hutter prize list. -- Matt Mahoney, [email protected] ------------------------------------------ Artificial General Intelligence List: AGI Permalink: https://agi.topicbox.com/groups/agi/T906c878e4ba411be-M019f5e6f40cd1aac567816b9 Delivery options: https://agi.topicbox.com/groups/agi/subscription
