I would like to make some very general comments on AGI. Marcus Hutter's AIXI shows that in a very general sense, the optimal behavior of a goal seeking agent at each point in time is to guess that the environment is simulated by the shortest Turing machine consistent with all observations so far. This is not a solution to the AGI because the problem is not computable, and is intractable even in the restricted domain of space and time bounded environments. Rather AIXI is a unifying framework for all machine learning algorithms, such as neural networks, SVM, Bayes, decision trees, GA, clustering, whatever. The implied goal of each is to find the simplest hypothesis that fits the data. More formally, they all use tractable algorithms to search a subset of short Turing machines for those consistent with the training data.
Let us consider the subset of learning problems that are important to humans (vision, language, robotics, etc). Ben Goertzel made an important observation, that AGI = pattern recognition + goals. These correspond to the two learning mechanisms in animal brains, classical conditioning + operant conditioning. AGI is unsolved, so we cannot yet say which method should be used. But I should comment that the one working model we do have (the human brain) is best modeled by neural networks. I believe the reason that neural networks have so far failed to solve the problem is lack of computing power (10^13 connections, 10^14 connections/second) and lack of corresponding training data (several years of video, audio, etc). Google has this much computing power and data. It can already answer simple natural language queries like "how many days until Xmas?". There have been lots of smart people working on AI for the last 50 years. By 1960 we have already seen language translation, handwriting recognition, natural language query answering in restricted domains, chess playing, automatic theorem proving, etc. If there was a shortcut to the general problem that didn't require massive computing power, I think we would have found it by now. On 9/4/06, YKY (Yan King Yin) <[EMAIL PROTECTED]> wrote: > I think the essense of Hawkins' theory (his HTM [hierarchical temporal > memory] model) is the compression of sensory experience via pattern > recognition. Sensory experience goes in; condensed episodic memory comes > out. He does this with neural networks. > > I worked with NNs for a while along exactly the same line of thought. After > a while I just decided that NN is too difficult to work with, so I switched > to (predicate) logic as the substrate for pattern recognition. > > Take an example: > > John hits Mary. > Mary kicks John. > Mary kicks John again. > John hits Mary again. > etc, etc. > > The point is to recognize that John and Mary are "fighting", thus achieving > compression. The fighting pattern can be irregular consisting of X hit Y, Y > kick X, etc. With logic I can write down a rule for recognizing this pretty > easily, mainly due to the use of symbolic variables. So you see the > compressive power of logic. NN is just too clumsy to work with. Although > we know that the brain somehow must perform this information compression > with neurons, we just don't understand the mechanisms yet. > > Let's say the goal is to compress visual inputs to the "John hits Mary" > level. I think it can be done using my vision scheme plus a logical > knowledge representation. But with NN, this still seems very very > remote.... There is a statistical language modeling solution to this problem. Counting Google hits: "the" 25,250,000,000 (at least this many English web pages) "fight" 603,000,000 "hit kick" 49,700,000 "hit kick fight" 18,500,000 So "fight" occurs on about 2.3% of all web pages, but 37% of web pages containing "hit" and "kick". In fact, you could get similar numbers if your training sample had 1/1,000,000 as many pages. But to use a smaller training set than that you would have to use techniques like LSA to exploit the transitive property of semantics. If there are no documents containing both "hit" and "fight" then you could still infer the relationship from documents containing both "hit" and "punch" plus documents containing both "punch" and "fight". LSA (latent semantic analysis) is described as the factoring of a word-document matrix into 3 matrices by SVD (singular value decomposition), where the middle matrix is diagonal, then discarding all but a few hundred of the largest diagonal terms. This greatly reduces the storage requirement (i.e. a simpler model). Furthermore, the SVD is equivalent to a 3 layer linear neural network with the layers representing words, an abstract semantic space, and documents. Not that SVD is fast... -- Matt Mahoney, [EMAIL PROTECTED] ------- To unsubscribe, change your address, or temporarily deactivate your subscription, please go to http://v2.listbox.com/member/[EMAIL PROTECTED]
