next up previous
Next: UNSUPERVISED LEARNING Up: Probability Interpretation of Population Previous: PROBABILITY ANALOGY

SUPERVISED LEARNING

Consider a nonlinear smooth function between two variables y and x given by . If x is a random variable with probability density , then this induces a probability density and we seek the functional that maps into . This immediately yields

The Density Mapping Theorem: For any smooth function there exists a piecewise linear transformation between probability representations for x and y given by

 

where for each monotonic segment of f.

This remarkably simple result shows that any nonlinear function between two analog variables has a representation as a linear functional between their probability densities. In terms of population codes, this provides an important computational property since only linear operations will be needed.

Note that the theorem only applies to the original continuous densities and , not to the smoothed and sampled versions given by population codes and . However, it can be shown that the best approximation to given is linear. The optimal linear transformation can be described by where the matrix A is given by

 

is as above, and are column vectors, and is the vector function that minimizes the Hilbert space norm .

Given population codes and we can easily learn the linear mapping between them using either direct computation (via matrix pseudo-inversion) or iterative learning algorithms. Suitable algorithms include the Widrow-Hoff delta rule [Widrow and Hoff 1960] or recursive least squares estimation [Ljung and Soderstrom 1983].



Terence D. Sanger
Mon Aug 21 18:36:58 EDT 1995