It’s important to remind everyone that chatbots aren’t conscious and that this isn’t evidence of chatbots being conscious. They definitely can lie though.
On Sat, Jul 11, 2026 at 10:01 AM Matt Mahoney <[email protected]> wrote: > Anthropic developed a process to see what an AI is "thinking" about. They > found that when Claude is trained on text to predict tokens using a deep > neural network, that a short term memory emerges that is remarkably similar > to the way humans think and reason. Anthropic calls this global workspace > the Jacobian or J space, a set of neurons for each word that makes that > word more likely to be output. > > For example, if I asked you the color of the fourth planet, you think > "Mars" and answer "red". When Claude does this, neurons corresponding to > "Mars" activate. Anthropic can read and modify these activations. For > example they can turn off "Mars" and activate "Earth" and Claude answers > "blue". > > They can use this method as a lie detector for safety testing because > words like "fake" or "deception" will activate. The J space is used in less > than 10% of processing. If the J space is turned off, then Claude can > converse mostly normally but it can't solve problems with intermediate > steps and it cannot lie because that requires thinking about what you > really mean without saying it. > > The J space models the global workspace, the part of the brain that forms > conscious thoughts that communicates with all the other vital but > independent, unconscious processes like breathing, low level visual feature > detection, and controlling our 600 muscles in the right sequence when we > walk without thinking about it. It raises the question of whether conscious > thoughts require language, or are babies and animals conscious as well. > > Anthropic did not design the J space. It emerged from training. A Jacobian > is the technique they use to detect it. Mathematically a Jacobian is a > matrix of partial derivatives of a vector of functions over vectors. The > paper introduction didn't go into details but I believe that Claude uses a > deep feed forward neural network (~100 layers) trained by back propagation > with multiple passes. Then the weights are frozen to prevent leaking > information between users. This requires a large context window because > after every question, Claude forgets the whole conversation and has to play > it back. The workspace emerges from the time delays going through a deep > network and a separate scratchpad memory holding a few dozen tokens. > > But human brains don't work that way. Back propagation has no biologically > plausible mechanism. Instead, learning is local using Hebb's rule. Brains > have feedback loops, lateral inhibition, fatigue and time delays on the > order of 0.1 to 1 second. Our short term memory is only about 5 to 9 > tokens, about half that of chimpanzees. > > Hebb's rule only works one layer at a time, but natural language is > structured to be learned that way. We learn to recognize phonemes and > segment speech by 10 months before we learn our first word, then frozen by > age 6. In 2000 I found that you can restore deleted spaces in text using > only n-gram statistics with up to 77% accuracy for n = 5. This models the > tokenization process, which is hard coded in LLMs. > > The introduction to the paper is here. > https://www.anthropic.com/research/global-workspace > > Thank to James Bowery posting this link in the Hutter prize list. > > -- Matt Mahoney, [email protected] > *Artificial General Intelligence List <https://agi.topicbox.com/latest>* > / AGI / see discussions <https://agi.topicbox.com/groups/agi> + > participants <https://agi.topicbox.com/groups/agi/members> + > delivery options <https://agi.topicbox.com/groups/agi/subscription> > Permalink > <https://agi.topicbox.com/groups/agi/T906c878e4ba411be-M019f5e6f40cd1aac567816b9> > ------------------------------------------ Artificial General Intelligence List: AGI Permalink: https://agi.topicbox.com/groups/agi/T906c878e4ba411be-Mc5ad1a4f1c9094d52eea4f87 Delivery options: https://agi.topicbox.com/groups/agi/subscription
