I would like to make some very general comments on AGI.  Marcus Hutter's AIXI 
shows that in a very general sense, the optimal behavior of a goal seeking 
agent at each point in time is to guess that the environment is simulated by 
the shortest Turing machine consistent with all observations so far.  This is 
not a solution to the AGI because the problem is not computable, and is 
intractable even in the restricted domain of space and time bounded 
environments.  Rather AIXI is a unifying framework for all machine learning 
algorithms, such as neural networks, SVM, Bayes, decision trees, GA, 
clustering, whatever.  The implied goal of each is to find the simplest 
hypothesis that fits the data.  More formally, they all use tractable 
algorithms to search a subset of short Turing machines for those consistent 
with the training data.

Let us consider the subset of learning problems that are important to humans 
(vision, language, robotics, etc).  Ben Goertzel made an important observation, 
that AGI = pattern recognition + goals.  These correspond to the two learning 
mechanisms in animal brains, classical conditioning + operant conditioning.

AGI is unsolved, so we cannot yet say which method should be used.  But I 
should comment that the one working model we do have (the human brain) is best 
modeled by neural networks.  I believe the reason that neural networks have so 
far failed to solve the problem is lack of computing power (10^13 connections, 
10^14 connections/second) and lack of corresponding training data (several 
years of video, audio, etc).  Google has this much computing power and data.  
It can already answer simple natural language queries like "how many days until 
Xmas?".

There have been lots of smart people working on AI for the last 50 years.  By 
1960 we have already seen language translation, handwriting recognition, 
natural language query answering in restricted domains, chess playing, 
automatic theorem proving, etc.  If there was a shortcut to the general problem 
that didn't require massive computing power, I think we would have found it by 
now.


On 9/4/06, YKY (Yan King Yin) <[EMAIL PROTECTED]> wrote:
> I think the essense of Hawkins' theory (his HTM [hierarchical temporal
> memory] model) is the compression of sensory experience via pattern
> recognition.  Sensory experience goes in; condensed episodic memory comes
> out.  He does this with neural networks.
>
> I worked with NNs for a while along exactly the same line of thought.  After
> a while I just decided that NN is too difficult to work with, so I switched
> to (predicate) logic as the substrate for pattern recognition.
>
> Take an example:
>
> John hits Mary.
> Mary kicks John.
> Mary kicks John again.
> John hits Mary again.
> etc, etc.
>
> The point is to recognize that John and Mary are "fighting", thus achieving
> compression.  The fighting pattern can be irregular consisting of X hit Y, Y
> kick X, etc.  With logic I can write down a rule for recognizing this pretty
> easily, mainly due to the use of symbolic variables.  So you see the
> compressive power of logic.  NN is just too clumsy to work with.  Although
> we know that the brain somehow must perform this information compression
> with neurons, we just don't understand the mechanisms yet.
>
> Let's say the goal is to compress visual inputs to the "John hits Mary"
> level.  I think it can be done using my vision scheme plus a logical
> knowledge representation.  But with NN, this still seems very very
> remote....

There is a statistical language modeling solution to this problem.  Counting 
Google hits:

"the" 25,250,000,000 (at least this many English web pages)
"fight" 603,000,000
"hit kick" 49,700,000
"hit kick fight" 18,500,000

So "fight" occurs on about 2.3% of all web pages, but 37% of web pages 
containing "hit" and "kick".  

In fact, you could get similar numbers if your training sample had 1/1,000,000 
as many pages.  But to use a smaller training set than that you would have to 
use techniques like LSA to exploit the transitive property of semantics.  If 
there are no documents containing both "hit" and "fight" then you could still 
infer the relationship from documents containing both "hit" and "punch" plus 
documents containing both "punch" and "fight".

LSA (latent semantic analysis) is described as the factoring of a word-document 
matrix into 3 matrices by SVD (singular value decomposition), where the middle 
matrix is diagonal, then discarding all but a few hundred of the largest 
diagonal terms.  This greatly reduces the storage requirement (i.e. a simpler 
model).  Furthermore, the SVD is equivalent to a 3 layer linear neural network 
with the layers representing words, an abstract semantic space, and documents.  
Not that SVD is fast...

-- Matt Mahoney, [EMAIL PROTECTED]




-------
To unsubscribe, change your address, or temporarily deactivate your 
subscription, 
please go to http://v2.listbox.com/member/[EMAIL PROTECTED]

Reply via email to