%* Last edited: Jan 18 02:13 2000 (sw)

% dense is about .6 the pages of not
\newif\ifdense
\densefalse

\ifdense
\documentstyle[10pt,psfig]{nih}
\else
\documentstyle[12pt,psfig]{nih}
\fi

% hospital
\topmargin = 0.0in

% mit
%\topmargin = -0.75in

\ifdense
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
% set up for maximum density: 10 pt 
\textheight = 9.36in
\textwidth = 8.0in
\oddsidemargin = -.75in
\evensidemargin = -.75in
\renewcommand{\baselinestretch}{1.1}
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
\else
	\textwidth = 6.8in
	\textheight = 9in
	\oddsidemargin = -0.25in
	\evensidemargin = -0.25in
	% ``double spaced''...
	\renewcommand{\baselinestretch}{1.06}
\fi


\pagestyle{myheadings}

\title{Mutual Information-Based Image Processing for fMRI}

\author{
	John Fisher 
	\and Eric Grimson
	\and Ferenc Jolesz
	\and Ron Kikinis
	\and Lawrence Panych
	\and Martha Shenton
	\and Andy Tsai
	\and William M. Wells 
	\and Cindy Wible 
	\and Alan Willsky 
	\and Kelly Zou}

%. Wells, C. Wible, J. Fisher, K. Zou, A. Willsky
%	?????	R. Kikinis, F. Jolesz, E. Grimson  ???????
%	?????   P. Black ??????



\begin{document}

\maketitle

\abstract

This project aims primarily to further develop and evaluate a
novel method for the detection
of activation in functional Magnetic Resonance Images (fMRI).  

Standard approaches to the detection of activation 
are described, as well as the clinical significance
of detection results for neurosurgical planning.

The new method is
  based on a formulation of the mutual information between two
  waveforms--the {\em f}MRI temporal response of a voxel and the
  experimental protocol timeline. In this method, 
scores based on mutual information
  are generated for all voxels and then used to compute the activation
  map of an experiment.  Mutual information for {\em f}MRI analysis is
  employed because it has been shown to be robust in quantifying the
  relationship between any two waveforms. Our
  technique takes a principled approach toward calculating the brain
  activation map by making few assumptions about the relationship
  between the protocol timeline and the temporal response of a voxel.
This may be significant in {\em f}MRI experiments where little
  is known about the relationship between these two waveforms.
Preliminary 
  experiments are presented to demonstrate this approach of computing
  the brain activation map, and a 
comparison to other more traditional
  analysis techniques is presented.

A plan is described to obtain representative data sets from ongoing
projects and to use them for a comparison of the new method and a
standard method.

\markboth{}{Principal Investigator: Wells, William M.}

\eject

%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
\section{Specific Aims}

This project aims to further develop and evaluate a novel method for
the detection of activation in functional Magnetic Resonance Images
(fMRI), with the ultimate goal of developing a larger research program
based on the results obtained from the present project.

It is our hypothesis that the MI-based approach offers significant
advantages over other approaches, in terms of detection sensitivity,
robustness with respect to any unknown relationships between the
stimulus and neural activity, and resistance to confounding signals
such as cardiac and respiratory cycles.

The new method, Mutual Information (MI) based activation detection,
will be compared to a standard method (GLM) on representative data
sets that are currently being processed in our laboratory for
neurosurgical planning and for neurological research.

The comparison will attempt to determine how significantly the methods
differ, and if they do, how these differences would bear on the
current applications, if the new method were used instead of the
standard method.




%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
\section{Background and Significance}

\subsection{Functional magnetic resonance imaging}

Since its development by several groups in the early 1990's([5]-[8]),
fMRI has been accepted as an effective method for mapping brain
function in humans. Prior to the development of fMRI, positron
emission tomography (PET) was the method of choice for spatial mapping
of cerebral function in humans. Although PET enables highly specific
biochemical studies of cerebral function, functional MRI has several
important advantages that are rapidly making it the preferred method
for functional brain mapping: fMRI is non-invasive; it gives much
higher spatial and temporal resolution than PET; and, of special
importance for the research proposed here, it is routinely used now at
many sites, including ours, to do presurgical mapping of eloquent
cortex in neurosurgical candidates.


This powerful new imaging modality generates images of the
brain that reflect brain tissue hemodynamics.  Brain tissue
hemodynamics are spatially related to the metabolic demands of the
brain tissue caused by neuronal activity. Therefore, indirectly, this
imaging modality can capture brain neuronal dynamics at different
sites while being activated by sensory input, motor performance, or
cognitive activity.

Several different MR techniques have been used for mapping cortical
function. The earliest work involved the use of contrast agents in
combination with single-shot echo-planar imaging (EPI) to detect
changes in relative cerebral blood volume due to functional
stimulation [5]. Later work has used non-invasive spin-labeling
techniques to detect local changes in perfusion[9]. The most common
and flexible method, however, is based on detection of the blood
oxygenation dependent (BOLD) effect using various imaging sequences
that are sensitive to local changes in magnetic
susceptibility([6]-[8]). These microscopic susceptibility changes are
associated with deoxyhemoglobin level changes that occur in response
to heightened oxidative metabolism associated with neuronal
activation. At 1.5 Tesla which is the magnetic field strength that is
most commonly available for clinical imaging, BOLD signal changes are
typically less than 6 percent[10].

\subsection{Applications of functional MRI: Pre-Surgical mapping}

Functional MRI has been applied for basic studies of brain
organization, in clinical neurology, in neuropsychiatry and for
neurosurgical planning [11, 12, 13]. It has been used extensively for
mapping the visual[14], motor[15], and auditory[16] cortex. It has
also been used for mapping of cortical areas involved in language[17]
and working memory[18]. Numerous standard cognitive tasks have been
studied with fMRI. Functional MRI has also been used to study clinical
populations with neurophsychiatric disorders such as schizophrenia and
obsessive-compulsive disorder[19].

Because of the ability to map a wide range of functions, its
relatively high spatial resolution and the fact that it is
non-invasive, fMRI is a logical choice for pre-surgical planning and
intra-operative guidance of surgical interventions. Over the last
several years, a number of groups have reported experiences using fMRI
for presurgical mapping of motor [20, 21, 22] and sensorimotor
function[23, 24] as well as primary auditory and language areas[25].

Functional MRI has been used to map function in patients with
neoplasms and vascular malformations. It has also been used for
mapping speech-related functions in patients with medically
intractable temporal lobe seizures[26]. In most of these studies, the
primary goal was to examine the ability of fMRI to map function and,
in some cases the fMRI results were compared with intra-operative
mapping methods such as electro-cortical stimulation (ECS) and
somatosensory evoked potential (SEP) recording[23, 27]. 
Although the
general conclusion of the studies has been that fMRI has great
potential for use in pre-surgical mapping, fMRI is not yet  widely
used to explicitly guide surgical interventions.

\subsection{Detection of Activation in {\em f}MRI}

The specific area of {\em f}MRI analysis we address in this proposal is
the identification of those voxels in the {\em f}MRI scan which are
functionally related to the experimental stimuli. This entails
determining whether the acquired temporal response of a voxel during
the scan is related to the experimental protocol timeline that is used
during the scan. 

The majority of fMRI experiments infer neural activity from blood
oxygenation level-dependent (BOLD) contrast.  When a brain region is
active, blood flow/volume to that region increases, changing the ratio
of deoxyhemoglobin to oxygenated hemoglobin in the region (see Ogawa
et al., 1993; Rosen et al., 1991).  Deoxygenated hemoglobin is
paramagnetic and causes a local loss of MR signal due to
susceptibility effects; in other words, it interferes with the MR
signal.  Although neural activity requires oxygen from the blood, the
hemodynamic response to activity results in more oxygenated hemoglobin
to the region, an overshoot of blood relative to the oxygen needed,
resulting in increased MR signal.  The hemodynamic response to
increased activity in visual cortex, for example, occurs approximately
2 seconds after a stimulus is perceived, reaches a peak around 4
seconds and is prolonged for a total of 10 to 12 seconds (Buckner,
1998; Burock et al., 1998).
	
The precise relationship between neural activity and the hemodynamic
response has not been systematically examined, and the implicit
assumption is that it is approximately linear.  However, it is
important to note that even single unit (or single cell) responses to
stimuli vary considerably from region to region.  Firing rate, firing
pattern (e.g.  burst, sustained, cyclic), and type of response
(excitation or inhibition) all differ between regions.  For example,
the typical firing rate in response to a stimulus can vary from an
average of 15 hz in a region like the hippocampal formation to as much
as 200 hz in primary sensory/motor regions (e.g.  Shadlen and Newsome,
1998).  Thus, even if the relationship between hemodynamic responses
and neural activity were linear, the BOLD signal would be expected to
differ between active regions.

As cognitive and psychological variables such as habituation and
attention are added to the equation, the relationship between brain
activity and stimuli becomes even more complex \cite{Sch:98}. This,
coupled with the fact that {\em f}MRI measurements--which do not
directly measure brain activitiy-- but rather are based on the
hemodynamic response to that activity, makes the relationship between
neural activity, the bold response, and cognitive function even harder
to establish.

Because of the complex, most certainly nonlinear and perhaps
stochastic, nature of the relationship between the two waveforms, it
has been difficult to find a suitable metric to quantify the
dependencies.  The technique we present in this proposal can be used to
overcome such obstacles.

\subsubsection{Popular Strategies for Analysis of {\em f}MRI data}
Currently, the popular analysis methods used to obtain the activation
map includes direct subtraction \cite{Hen:93}, correlation coefficient
\cite{Ban:93,Woo:94}, and the general linear model \cite{Fri:94}.
Quantitative comparisons of these methods are hard to quantify because
although it is now known in general what regions should be active
during many basic language, motor, and visual tasks, individuals vary
in the precise location of the regions and in the amount of bold
signal produced for a given task.  For example, if two
analysis methods yielded different size functional regions in the same
individual and the same task, it would be difficult to determine which
method was closer to the actual neural region that was activated.


{\bf Also, somewhere you may want to talk a little about applying this
method in the future to event-related FMRI studies.  These studies
DO make a fairly strong assumption that the bold response to a given
stimulus is linear, is of a certain shape and is additive.}



\paragraph{Direct Subtraction (DS)}
This method involves calculating two mean intensities for each
voxel--one mean value calculated based on averaging together all the
temporal responses acquired during the ``task'' period, and the other
mean value calculated based on averaging together all the temporal
responses acquired during the ``rest'' period of an experiment.  To
determine whether a voxel is activated or not, one mean intensity is
subtracted from the other.  Voxels with significant difference in the
mean intensities of the two data groups are identified as being
activated. To yield a statistic to identify significant difference in
the intensities, a Student's t-test is employed. This test determines
whether the means of the two data groups are statistically different
from one another by utilizing the difference between the means
relative to the variabilities of the two data groups. The t-value this
method generates, for a temporal response $y$, is calculated as
\begin{eqnarray}
t=\frac{\bar{y}_{on} - \bar{y}_{off}}{\sqrt{\frac{\sigma^2_{y_{on}}}{N_{on}-1} + 
    \frac{\sigma^2_{y_{off}}}{N_{off}-1}}} \nonumber
\end{eqnarray}
where ${y_{on}}$ and ${y_{off}}$ denote the set of data points in the
temporal measurements that correspond to the ``task'' and the ``rest''
periods, respectively, and ${N_{on}}$ and ${N_{off}}$ denote the
number of time points that corresponds to the ``task'' and the
``rest'' periods, respectively. The mean and variance of the data
group ${y_{on}}$ are denoted as ${\bar{y}_{on}}$ and
${\sigma^2_{y_{on}}}$, respectively. Likewise, the mean and variance
of the data group ${y_{off}}$ are denoted as ${\bar{y}_{off}}$ and
${\sigma^2_{y_{off}}}$, respectively. The major shortcoming associated
with this method is that it relies heavily on the assumption that
temporal measurements of a given voxel can be partitioned into two
data groups, each normally distributed according to a different mean
and variance.

\paragraph{Correlation Coefficient (CC)}
The correlation coefficient ${\rho_{xy}}$ is a normalized measure of
the correlation between the reference waveform $x$ and the measurement
waveform $y$, and is defined by
\begin{eqnarray}
  \rho_{xy} = \frac{\sum(x-\bar{x})(y-\bar{y})}
                   {\sqrt{\sum(x-\bar{x})^2\sum(y-\bar{y})^2}} \nonumber
\end{eqnarray}
where ${\bar{x}}$ and ${\bar{y}}$ denote the means of $x$ and $y$,
respectively. The summation is taken over all the time points in the
waveform. It is easy to establish that ${-1 \leq \rho_{xy} \leq 1}$.
Voxels with large ${|\rho_{xy}|}$s are considered to be activated.
For this method, ${|\rho_{xy}|}$ is used as the test statistic for
statistical inference. This method critically depends on the choice of
the reference waveform. Various waveforms have been used
\cite{Ban:93,Woo:94}; however, in light of the many unknown factors
affecting measurement of brain activation, reference waveform design
poses a serious obstacle, especially for more complicated protocols.

\paragraph{General Linear Model (GLM)}
The statistical models used for parameter modeling in the two
previously described analysis methods are both special cases of the
general linear model.  This model is a framework designed to find the
correct linear combination of explanatory variables (such as
hemodynamic response, respiratory and cardiac dynamics) that can
account for the temporal response observed at each voxel during an
experiment. Assume that there exists $T$ number of time point
measurements per voxel in the {\em f}MRI data set. Let ${y_{t}}$
denote the measurement at some voxel at time $t$, and let
${\epsilon_{t}}$ denote the error term associated with the linear
model fit at that same voxel at time $t$, with ${1 \leq t \leq T}$.
Here, ${\epsilon_t \sim {\cal N} (0,\sigma^2)}$. Suppose there are $J$
number of explanatory variables in the linear model. Let ${x_{jt}}$
denote the value of the $j$th explanatory variable at time $t$ with
${1 \leq j \leq J}$. Also let ${\beta_j}$ denote the scaling parameter
for the $j$th explanatory variable. With these definitions, the
general linear model can be written as
\begin{eqnarray}
  \left[ \begin{array}{c}
        y_1 \\ y_2 \\ \vdots \\ y_T
     \end{array} \right] =
  \left[ \begin{array}{cccc}
        x_{11} & x_{12} & \dots & x_{1J} \\
        x_{21} & x_{22} & \dots & x_{2J} \\
        \vdots & & \ddots & \vdots \\
        x_{T1} & x_{T2} & \dots & x_{TJ}
     \end{array} \right]
  \left[ \begin{array}{c}
        \beta_1 \\ \beta_2 \\ \vdots \\ \beta_J
     \end{array} \right] +
  \left[ \begin{array}{c}
        \epsilon_1 \\ \epsilon_2 \\ \vdots \\ \epsilon_J
     \end{array} \right]. \nonumber
\end{eqnarray}
The above equation can be written succinctly in matrix notation as ${Y
  = X \beta + \epsilon}$. In general, $X$ is full rank and the number
of explanatory variables $J$ is less than the number of observations
$T$ indicating that the method of least squares can be employed to
find the scaling parameters $\beta$. Since ${X^TX}$ is invertible, the
least squares estimate for $\beta$, which we denote by
${\hat{\beta}}$, is ${(X^TX)^{-1} X^TY}$.  Then ${\hat{\beta}}$ is
used to test whether it corresponds to the model of an activation
response (as specified in $X$) or the null hypothesis.  One of the
major problems associated with this method is in the design of $X$.
As mentioned earlier, little is known about the relationship between
{\em f}MRI temporal response and brain stimulation. Hence, it is
difficult to identify the necessary explanatory variables that can
account for the temporal responses seen in {\em f}MRI measurements.

\subsection{Clinical significance of the proposal}

This work addresses a major problem in the treatment of brain lesions
- detecting localized brain function for surgical planning and
intra-operative guidance. We address this problem by proposing the
development of a new method tailored to the task of focusing on
localized changes. Providing the capability to accurately and
precisely determine the function of neural tissue that might be
compromised by impending neurosurgical procedure offers great
potential value to minimize post-operative functional
deficits. Achieving effective mapping of all relevant functions is an
objective that cannot be achieved by one single project. We do not
propose, for example, to solve the problem of which functional tasks
are best suited to the mapping problem, although we will present a
model of how these tasks should be studied. We do not propose to
answer the question as to what degree the cortical areas detected by
functional MRI represent `true' eloquent areas, although we expect
that our results will add to the body of evidence. What we do propose,
is to develop a new adaptive method that, in advancing the state of
the art in fMRI measurement, greatly improves our ability to answer
these and other questions. It is evident that the development of
functional MRI has placed an important new tool in the arsenal of the
modern neurosurgeon. We aim to demonstrate that real-time adaptive
fMRI represents the leading edge of that development.




%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
\section{Preliminary Studies}

\subsection{Mutual Information-Based Detection of Activation in fMRI}

\subsubsection{Introduction} %%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
In \cite{Tsai98}
we presented  a novel method based on an information-theoretic approach
to find the brain activation maps for {\em f}MRI experiments.  In this
method, mutual information is calculated between the temporal response
of a voxel and the protocol timeline of the experiment.  This value
can then be used as a score to quantify the relationship between the
two waveforms.  Mutual information is appropriate for {\em f}MRI
analysis because it has been shown to be more robust than other
methods in identifying complex relationships (i.e. those which are
nonlinear and/or stochastic).  More importantly, our nonparametric
estimator of mutual information requires little ${a}$ ${priori}$
knowledge of the relationship between the temporal response of a voxel
and the protocol timeline.  Over the past few years, mutual
information has been used to solve a variety of problems
\cite{Bel:98,Vio:97,Wel:96}.

\subsubsection{Background} \label{sec:background} %%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%

\subsubsection{Description of Method} \label{sec:description} %%%%%%%%%%%%%%%%%%%%

\paragraph{Mutual Information and Entropy}
Mutual information (MI) and entropy are concepts which underly much of
information theory \cite{Cov:91}. They cannot be adequately described
within the scope of this proposal. Suffice it to say that MI is a measure
of the information that one random variable (RV) conveys about
another, and entropy is a measure of the average uncertainty in a RV.
Both quantities are expressed in terms of bits of information. Here,
we demonstrate the appropriateness of MI for {\em f}MRI analysis.

The mutual information, ${I(u,v)}$, between the RVs $u$ and $v$, is
defined as \cite{Cov:91}
\begin{eqnarray}
  I(u,v) = h(v) - h(v|u) = h(u) - h(u|v),
  \label{eq:MI}
\end{eqnarray}
where the entropy, ${h(v)}$, quantifies the randomness of $v$ and the
conditional entropy, ${h(v|u)}$, quantifies the randomness of $v$
conditioned on observations of $u$. These terms are described by the
following expectations:
\begin{eqnarray}
  h(v) &=& -E_v \left[ \log_2 (P(v)) \right] \nonumber \\
  h(v|u) &=& -E_u \left[ E_v \left[ \log_2 (P(v|u)) \right] \right]. \nonumber
\end{eqnarray}
where $P$ denotes probability density.  It is clear from (\ref{eq:MI})
that MI is symmetric. That is, the information that $u$ conveys about
$v$ is equal to the information that $v$ conveys about $u$.
Furthermore, since $u$ is a discrete RV in our case and conditioning
always reduces uncertainty ${\left( h(u|v) \leq h(u) \right)}$, $v$
can convey at most ${h(u)}$ bits of information about $u$ (and vice
versa). We can therefore lower and upper bound the MI between $u$ and
$v$ by $0$ and ${h(u)}$, respectively.

\paragraph{Calculation of Brain Activation Map by MI}
We present nonparametric MI as a formalism for uncovering dependencies
in calculating the {\em f}MRI activation map.  Recall that in our
specific application, we seek at most one {\em bit} of information
(whether or not a voxel is activated). This impacts our choice of the
reference waveform. The reference waveform need be no more complicated
than our hypothesis space (1 bit). The protocol timeline shown in Fig.
\ref{fig:sig_vs_prot} is the simplest model of our hypothesis space
and is sufficient as the reference waveform when using MI as the basis
for comparison. More elaborate waveforms can be employed, but they
imply more information than is necessary. The consequence of this is
that complicated waveform design in unnecessary; the reference
waveform need only adequately {\em encode} the hypothesis space.

\begin{figure}
 \centerline{\psfig{figure=sig_vs_prot.eps,width=4.75in,silent=1}}
 \caption{Illustration of the Protocol Timeline, ${S_{v|u=0}}$, and ${S_{v|u=1}}$.}
 \label{fig:sig_vs_prot}
\end{figure}

In the following derivation, we will refer to the temporal response of
a voxel as $v$, and the reference waveform as $u$. We have already
established the appropriateness of using the protocol timeline as the
reference waveform $u$.  As such, $u$ only takes on two possible
values, 0 and 1, so we can rewrite equation (\ref{eq:MI}) as
\begin{eqnarray}
  I(u,v) = h(v) - P(u=0) h(v|u=0) - P(u=1) h(v|u=1)
  \label{eq:MIsimp}
\end{eqnarray}
where ${P(u=0)}$ and ${P(u=1)}$ are the $a$ ${priori}$ probabilities
of $u$ taking on the values of 0 and 1, respectively. This reference
waveform is chosen because it is the simplest encoding of the actual
hypothesis (task vs. rest) with equal probabilities reflecting
the relative frequencies of samples during each state.

As an illustrative example, suppose $v$ is a scaled and biased
version of $u$ (i.e. ${v=cu+d}$ where ${c,d \in \Re}$ and ${c \neq
  0}$).  Then
\begin{eqnarray}
   h(v|u=0) = - E_v \left[ \log_2 (P(v|u=0)) \right]
            = - E_v \left[ \log_2 (1) \right]
            = 0 \,\, \mbox{bits}, \nonumber
\end{eqnarray}
\begin{eqnarray}
   h(v|u=1) = - E_v \left[ \log_2 (P(v|u=1)) \right]
            = - E_v \left[ \log_2 (1) \right]
            = 0 \,\, \mbox{bits}, \nonumber
\end{eqnarray}
\begin{eqnarray}
   h(v) = - E_v \left[ \log_2 (P(v)) \right]
        = - E_v \left[ \log_2 (0.5) \right]
        = - \log_2 (0.5)
        = 1 \,\, \mbox{bit}, \nonumber
\end{eqnarray}
so that ${I(u,v) = 1}$ bit.  This is the maximum MI that can be
achieve between the square wave $u$ and {\em any} other waveform $v$.
Since only 1 bit of information is encoded in $u$, only 1 bit of MI
can exist between $u$ and {\em any} $v$.

\paragraph{Estimating Entropies}
Evaluating equation (\ref{eq:MIsimp}) requires computing ${h(v)}$,
${h(v|u=0)}$, and ${h(v|u=1)}$ which are integral functions of the
densities ${P(v)}$, ${P(v|u=0)}$, and ${P(v|u=1)}$. In general, these
must be estimated. We choose the nonparametric Parzen window method
\cite{Par:62} to estimate the densities and the sample means to
estimate the entropy terms. The Parzen density estimate (with
leave-one-out) is defined as
\begin{eqnarray}
  \hat{P}(v) = \frac{1}{(N_{S_v}-1)} \left[ \sum_{v_j \in S_v} 
               G_{\sigma}(v-v_j)-G_{\sigma}(0) \right] \nonumber
\end{eqnarray}

where ${N_{S_v}}$ is the number of data points in the sample set
${S_v}$, ${G_{\sigma}}$ is an admissible kernel function (we use the
Gaussian kernel, other kernels are possible), and ${\sigma}$ is the
standard deviation of the density function. Set ${S_v}$ is composed of
${all}$ the data points from $v$.  Our estimate for the conditional
${P(v|u=0)}$ is

\begin{eqnarray}
  \hat{P}(v|u=0) = \frac{1}{(N_{S_{v|u=0}}-1)} \left[ \sum_{v_j \in S_{v|u=0}}
                   G_{\sigma}(v-v_j)-G_{\sigma}(0) \right] \nonumber
\end{eqnarray}
where ${N_{S_{v|u=0}}}$ is the number of data points in the sample set
${S_{v|u=0}}$. Set ${S_{v|u=0}}$ is composed of the subset of data points
from $v$ with time points corresponding to when ${u=0}$.  Similarly,
the estimate for ${P(v|u=1)}$ is identical to ${\hat{P}(v|u=0)}$ only
taken over the subset of samples from $v$ with time points
corresponding to when ${u=1}$.

Kernel size (in our case the standard deviation, $\sigma$ of the
Gaussian kernel) is an issue for the Parzen window density
estimator. Consistent with our information-theoretic approach, we
choose the kernel size which maximizes the likelihood of the
data\cite{Cov:91}. For example, the kernel size ${\hat{\sigma}_{\tiny ML}}$ used to estimate ${h(v)}$ is
\begin{eqnarray*}
  \hat{\sigma}_{\tiny ML} &=& \arg \max_{\sigma}
                \left\{ \frac{1}{N_{S_v}}
                \sum_{v_j \in S_v} \log \hat{P}(v_j)  \right\}.
\end{eqnarray*}

As direct evaluation of the entropy terms is computationally
prohibitive we approximate them with their sample means
\cite{Vio:97,Wel:96}. For example, ${h(v)}$ is approximated by

%as follows:
\begin{eqnarray}
  h(v) \approx -\frac{1}{N_{S_v}} \sum_{v_j \in S_v}
         \log_2 \left( \frac{1}{N_{S_v}-1} \sum_{v_i \in S_v}
         G_{\sigma}(v_j-v_i) -G_{\sigma}(0)\right) \nonumber
\end{eqnarray}
The conditional terms, ${h(v|u=0)}$ and ${h(v|u=1)}$, are defined
similarly using the previously defined subsets ${S_{v|u=0}}$ and
${S_{v|u=1}}$, respectively.
%\begin{eqnarray}
%  h(v|u=0) \approx -\frac{1}{N_{S_{v|u=0}}} \sum_{v_i \in S_{v|u=0}}
%         \log_2 ( \frac{1}{N_{S_{v|u=0}}} \sum_{v_j \in S_{v|u=0}}
%         G_{\sigma}(v_i-v_j)) \nonumber
%\end{eqnarray}
%\begin{eqnarray}
%  h(v|u=0) \approx -\frac{1}{N_{S_{v|u=1}}} \sum_{v_i \in S_{v|u=1}}
%         \log_2 ( \frac{1}{N_{S_{v|u=1}}} \sum_{v_j \in S_{v|u=1}}
%         G_{\sigma}(v_i-v_j) ). \nonumber
%\end{eqnarray}

\subsubsection{Experimental Results} \label{sec:experiment} %%%%%%%%%%%%%%%%%%%%%%%%

We applied the above described {\em f}MRI analysis method to a single
{\em f}MRI data set that examines right-hand movements. The data set
contains 60 whole brain acquisitions with each whole brain acquisition
containing 21 slice images.

\begin{figure}
 \begin{center}
   \begin{tabular}{c@{\extracolsep{0.1in}}c@{\extracolsep{0.1in}}c@{\extracolsep{0.1in}}c}
    \mbox{\psfig{figure={subtraction.eps},width=1.12in,silent=1}} &
    \mbox{\psfig{figure={xcor.eps},width=1.12in,silent=1}} &
    \mbox{\psfig{figure={spm.eps},width=1.12in,silent=1}} &
    \mbox{\psfig{figure={mlmi.eps},width=1.12in,silent=1}} \\
   (a) DS & (b) CC &
   (c) GLM & (d) MI
   \end{tabular}
 \end{center}
 \caption{Comparison of {\em f}MRI Analysis Techniques.}
 \label{fig:AnalysisResults}
\end{figure}

Only the analysis results from the 10th coronal slice of the whole
brain acquisition are shown in Fig. \ref{fig:AnalysisResults}. The
figure provides a qualitative comparison of our analysis technique
with other techniques previously mentioned in this proposal. A
quantitative comparison of these different methods is difficult since
the ground truth is unknown. In keeping with the fairness of the
comparison, the threshold (which determines whether a voxel is
activated or not) that yields the ``best'' activation map for each
analysis technique is used.  For this particular {\em f}MRI data set,
the ``best'' activation map is judged based on the prior expectation
that brain activation is restricted to the left primary motor cortex
and occurs in clusters.  It is important to point out that MI is
inherently a normalized measure so for our technique, the threshold
can be specified meaningfully in terms of bits of information. Fig.
\ref{fig:AnalysisResults}(d) is obtained using a threshold of $0.75$
bits.  While the qualitative differences between the techniques as
observed in Fig.\ref{fig:AnalysisResults} is small the MI approach
combines very few assumptions about the underlying data {\em and}
an inherently normalized threshold (none of the other techniques have
both of these properties).  We believe that these results
demonstrate the viability and efficacy of the MI approach for {\em
f}MRI data analysis, although more extensive experimentation is warranted.

\subsubsection{Summary} \label{sec:summary} %%%%%%%%%
We have developed a theoretical framework for using MI to calculate
the {\em f}MRI activation map. While there are many existing
approaches to calculate the activation map, all of these techniques
depend on some {\em a priori} assumptions about the relationship
between the protocol timeline and the {\em f}MRI voxel temporal
response. The strength of our approach is that it relies on sound
theoretical principles, it is fairly easy to implement, and does not
require strong assumptions about the nature of the relationships
between the {\em f}MRI temporal measurements and the protocol
timeline, while still retaining the ability to uncover complex
relationships (beyond second-order statistics). In addition,
experimental results confirmed that this information-theoretic
approach can be as effective as other methods of calculating
activation maps. Finally, from the clinical standpoint, nuisance
variables are significantly reduced, that is, the protocol timeline
need only encode the actual hypothesis test (e.g. waveform matching is 
unnecessary).


%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%

\subsection{Preliminary Experiment Related to the Reduction
of Confounding Signals}

\label{LabSecSpeechVideo}

Confounding signals are of particular interest in the fMRI analysis
problem. Figure \ref{fig:visaud} illustrates a problem which is quite
similar to the problem of confounding signals in fMRI. Specifically,
there are a sequence of video frames (left) and an associated audio
signal (right). In the video sequence the person in the foreground is
speaking, while the person in the background is moving. In this
particular problem one might be interested in which video pixels (fMRI
voxels) are related to the audio (confounding cardiac) signal. Both problems
share the following challenges:
\begin{itemize}
 \item The measurements arise from different modalities (video
vs. audio).
 \item The joint statistics of the modalities are not well understood, 
and are certainly not well-modeled by such things as simple Gaussian
densities.
 \item The sampling rates differ significantly (video $\approx$ 30 Hz, audio
$\approx$ 22.5 Khz).
\end{itemize}

Analagous statements can be made with regard to fMRI acquisitions and
confounding signals such as cardiac and respiration. Despite these
challenges information theory and nonparametric statistics provides a
theoretical framework for assessing which video pixels are related to
the associated audio signal measurements.

In Figure \ref{fig:visaud2} we show a single video frame for reference
alongside an image of pixel-wise standard deviations (middle) and the
result of analysis by mutual information (right). A striking feature
is that the most intense changes in the images are due to the video
monitor and moving person. The MI analysis clearly disambiguates
changes which are due to speech from those which are not, despite the
fact the speech related pixel changes are less intense than the
others.

\begin{figure}[h]
 \begin{center} \psfig{file=nih_fmri/eps/visaud.eps,width=5in} \end{center}
 \caption{ Video (left) and associated audio (right) signals
 \label{fig:visaud}}
\end{figure}

\begin{figure}[h]
 \begin{center}
 \psfig{file=nih_fmri/eps/visaud2.eps,width=5in}
 \end{center}
 \caption{ Video frame (left), pixel-wise standard deviation image (middle), and mutual information image (right). \label{fig:visaud2}}
\end{figure}

%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
\section{Research Design and Methods}



\subsection{Algorithm and Software Development}

\subsubsection{Activation Detection}

The MI-based activation detection method that is described above is an
experimental program that is currently implemented in MATLAB[?].


% [This will need beef up 
% in other areas...The method will be refined to use additional signals
% EKG, respirations in order to increase the robustness
% of the detection to these confounding influences]

This method will be refined, re-coded in C++, and integrated
into the ``3D Slicer'', which has become a standard platform
in our institution for medical image analysis applications[?].
We will refer to this system as the {\em new method} in 
the descriptions that follow.

\subsubsection{Reduction of Confounding Signals}

We will develop an experimental implementation of
the MI-based method described above in Section \ref{LabSecSpeechVideo}
for the purpose of identifying those voxels whose intensities
are responding to associated EKG signals.  
If such responding voxels are present, a model
will be regressed to the responding voxels,
and we will attempt to reduce the overall effect
of the confounding signal by subtracting a
appropriately weighted version of the fitted
model. 

\subsubsection{Registration}

Our MI-based rigid registration system will be extended 
to increase its utility for the post-processing of fMRI
data sets.  While the system is currently useful for 
the motion correction of time series acquisitions, it 
does not currently perform reliably for the registration
of echo-planar data to the conventional high-resolution
MRI scans which are used as the anatomical reference
for the activation that is detected in the echo-planar
data.  The primary difficulty seems to be due to the 
substantial dark voids in the EPI data that are not
present in the conventional data.  We plan to accommodate
this aspect of the data by explicitly modeling this
aspect of the joint intensity data, in the manner
of Leventon et al. [?].


The MI-based rigid registration is already integrated into
the ``3D-Slicer''.

\subsection{Data Collection}

Members of our group at 
the Surgical Planning Laboratory are currently analyzing
fMRI data for two projects, neurosurgical planning
(Black et al.) [?] and neuroscientific  research (McCarley
et al.)[?].  The neuroscientific research
also contains a population of normal volunteers.

The t-test and cross-correlation tests are currently routinely done
for all FMRI subjects.  In addition, the group has recently starting
also doing the FMRI analyses using the SPM package that uses the GLM as
it's basis \cite{SPM}

The analysis is currently based on 
a variant of GLM, and our current MI-based
registration technology (the {\em existing technology}).


Over the two years of the project the project
will collect additional
data to result in a final tally of 36 representative data sets from each of the
surgical, neuroscientific, and normal populations.  This
scan data will be organized onto a filesystem hierarchy
for use by the present project.

% [should we propose to also record EKG and respiration?]

\subsection{Evaluation of Methods}

The collected data will be analyzed using the existing technology
as of the current projects described above.  

Within the scope of the present project, the collected data
will be processed with the newly developed technology
for registration and activation detection, and the 
derived results will be analyzed as described below.

\subsubsection{Sample size and statistical power calculations}

The calculations are based on Hypothesis {\bf H2}.  We assume a balanced design
with equal number of cases in each of the three experimental groups.
We test the average proportion of at most 75\% of the activated pixels
agreed upon by both estimation methods using one-sided paired t-tests,
stratified by experimental group.  We adopt a Bonferroni correction in
multiple testing. To have an overall significance level (Type I error)
of at most 5 \%, each individual test needs to meet a significance level
of 0.0167. In order to detect a  one-sided difference of 5\% in each
experimental group, the overall statistical power (1-Type II error) is
calculated at 91 \%, with $n=36$ cases per group, yielding a total sample
size of $N=3n=108$. {\it Projected accrual}: we estimate that it will take six
months to enroll these subjects in the study.

\subsubsection{Comparison of Detection Methods}

%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%

%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%

The existing technology and the new method will be compared using a
statistical analysis that is outlined below.  The following describes
the hypotheses to be test, and the methods 
that will be used in the tests.


The first test will analyze differences between the 
methods in detecting activation in application-specific
regions of interest (ROI).  In the case
of surgical patients, the ROI will be be determined
according to the clinical application, for example, 
pre- and post-central regions for the 
localization of the sensory/motor area
and posterior inferior frontal, posterior superior temporal, and
inferior parietal regions for the localization of 
language areas.
In the case of psychiatric patients, the ROI
are superior temporal gyrus for the experiments
that are studying responses to prepared audio stimuli
\cite{SchizScanProject}.
The ROIs will be determined by the use of 
established techiniques \cite{Cindy pre-frontal, marty stg}
currently in use in our laboratory.

The detection of activation will be marked
in the case that at least 25\% of the voxels
in an ROI are labelled by the detection
algorithm as activated.


\begin{itemize}
\item
{\bf H1}:
The two methods yield the same results on 
presence of activation in 
the application-specific ROIs.


{\bf Method}:
A stratified analysis of contingency tables
based on the three populations (surgical, psychiatric, normal).


\end{itemize}

The next test concerns
an objective asessement of
significant difference between the two
methods, a necessary condition for
added value in the new method.


{\em ALTERNATIVE?}
(or just assess the agreement in terms of
intraclass correlation coefficient to see 
if there is an added value of the new method
in terms)

\begin{itemize}

\item

{\bf H2}:
The agreement between the probability measures
yielded by the two methods is at most 70\%.

{\bf Method}: An
analysis of intraclass correlation coefficient
using two-way analysis of variance (ANOVA).

\end{itemize}

The next test is aimed at analzing differences
among the methods in terms of the three
populations (surgical, psychiactric, and normal).
This is based on the observation that the 
non-normal subjects can be more difficult to 
examine, and may generate more challenging
data due to increased motion artifacts, etc.

This analysis will make use of manually-labelled
activation maps produced by expert raters upon
examination of the PMAPS produced by the two 
methods.

\begin{itemize}
\item

{\bf H3}:
At any given threshold, an ROC curve reduces the probability
data to simple binary decisions, i.e., activation
vs. no activation.  Compare the decision and the 
location indicators with those by the
neurosurgical or neuroscientific expert
client, which will serve as the gold standard.

{\bf This part confuses me:}
The information is found in
the questionaire.


{\bf Method}:
Usage of the pairwise kappa statistic and the ROC method
that is described in \cite{KellyROC}.

\end{itemize}


{\bf This needs work...}


\begin{itemize}
\item

{\bf H4}:
Analysis of other items in the questionaire.

\end{itemize}

This test concerns an evaluation of the
confounding signal reduction method outlined above
in comparison to the existing method used for the
same purpose in
Panych et al.'s grant\cite{LarryR01}.  

\begin{itemize}
\item

{\bf H5}:
The new method for confounding
signal reduction reduces the rate
of false positive activations due
to confounding signals in comparison
to the existing method.

{\bf Method}:

NEEDS WORK.


\end{itemize}


\subsubsection{Analysis of Subjective Evaluation of Detection Methods}

The ultimate validation of fMRI as a method for 
detecting neuronal activity is very difficult and
beyond the scope of the present project.  

We will attempt to gain related information by the
analysis of questionnaires answered by experts.  

Each case will be systematically reviewed 
by an expert (neurosurgeon or research neuroradiologist
respectively) who will fill in a questionnaire 
concerning the image processing results.

The questionnaire will be designed to address the
following items:

\begin{enumerate}

\item
Do the two detection methods produce results that
differ in a significant way? (yes no)

\item
Which method appears to produce results that are
more useful for the application? (In the case
that the first answer is positive).

\item
Does one method appear to detect plausible activation
areas that the other method misses?

\item
Does one method appear to be more resistant to 
confounding signals?

\end{enumerate}

The questionnaires will be statistically analyzed to
evaluate the issues described above.

In the case that answers to question three are
yes (HOW OFTEN), focussed testing will
be carried out to increase to establish
the difference. (BUT THEN WE'D NEED IRB APPROVAL
FOR NEW SCANS, RIGHT?)



The next test will compare the performance
of the methods on repeated experiments.  
five normal subjects will be tested
five times for this test.

\begin{itemize}
\item

{\bf H6}:
The new method has a higher reliability and reproducibility
than the existing method over the
5 repeated experiments under a well controlled
environment.

{\bf Method}:
Analysis of repeated measures.

\end{itemize}


\subsubsection{Evaluation of Registration}

We also plan to perform retrospective fiducial-based validation of
registration using a subset of the data described above.  One measure
we plan to evaluate for this purpose will be the magnitude of the
largest displacement error that occurs within the registered
anatomical volumes. We also plan to characterize the actual
requirements of registration accuracy for fMRI activation detection.

{\bf IF IT STARTS ON TWO, THIS HAS TO BE PAGE ELEVEN (after
we switch to ten point)}

%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
\eject
\section{Human Subjects}

No additional MR scans will be initiated for this research.  This
project involves acquiring study information during the context of
fMRI examinations already being performed at our hosptial within the
scope of an NIH funded research project \cite{SchizGrant}, and an
ongoing effort in neurosurgical planning \cite{BlackApplication}.
Because of its use of only the stored medical records (MRI images)
this project therefore is claimed to be exempt from human subject
regulations based on exemptions quoted in the NIH guidlines. There is
existing hospital human subjects review board internal review board
(IRB) approval for the two projects described above.

All subjects will be adults who will be identified throughout the
project only by a unique study subject number; a table linking subject
number, name and hospital identification code will be stored on media
isolated from all results and reporting. All genders and minorities
will be included, with no exclusions, as cases become available in the
two projects.  Hospital human subjects review board IRB approval for
use of the on-line medical images for the present project is pending.

\section{Vertebrate Animals}

None will be used.

\section{Literature Cited}

soon.

\section{Consortium/Congractual Arrangements}

Not applicable.

\section{Consultants}

\subsection{John Fisher}

Dr. Fisher is an MIT researcher and qualified
expert in applications of information theory.
He will contribute primarily in the 
area of develping algorithms for the detection
of activation in fMRI.  He will be a primary
supervisor of the MIT graduate student.

\subsection{Andy Tsai}
 
Mr. Tsai is compeleting his PhD research under
Prof. Willsky at MIT.  His research experience
includes applications of information theory, 
and has valuable knoledge for the project, as
the primary implementer of the system that
has produced our preliminary results.  

Mr. Tsai will provide expertise
in the area of non-parametric statistics,
as well as guidance in the usage and adaptation
of the existing software that was used for the
preliminary experients.

\subsection{Cindy Wible}

Dr. Wible has a PhD in psychology and is an expert
at functional neuroimaging at Harvard Medical School.

Dr Wible is currently performing fMRI examinations for clinical and
research purposes.  She will be our primary contact to clinical fMRI
data for the project, and will assist in the evaluation of the
detection technology as it is developed.

%\begin{thebibliography}{10}

%\bibitem{Ban:93}
%P.A. Bandettini, A.~Jesmanowicz, E.C. Wong, and J.S. Hyde.
%\newblock Processing strategies for time-course data sets in functional {MRI}
%  of the human brain.
%\newblock {\em Magnetic Resonance in Medicine}, 30:161--173, 1993.

%\bibitem{Bel:98}
%F.~Bello and A.C.F. Colchester.
%\newblock Measuring global and local spatial correspondence using information
%  theory.
%\newblock In {\em Proceedings of the First International Conference on Medical
%  Computing and Computer-Assisted Intervention}, 1998.

%\bibitem{Cov:91}
%T.M. Cover and J.A. Thomas.
%\newblock {\em Elements of Information Theory}.
%\newblock John Wiley and Son Inc., 1st edition, 1991.

%\bibitem{Fri:94}
%K.J. Friston, P.~Jezzard, and R.~Turner.
%\newblock Analysis of functional {MRI} time-series.
%\newblock {\em Human Brain Mapping}, 1:153--171, 1994.

%\bibitem{Hen:93}
%O.~Henriksen, H.B.W. Larsson, P.~Ring, E.~Rostrup, A.~Stensgaard, M.~Stubgaard,
%  F~Stahlberg, L.~Sondergaard, C.~Thomsen, and P.~Toft.
%\newblock Functional {MR} imaging at 1.5{T}.
%\newblock {\em Acta Radiologica}, 34:101--103, 1993.

%\bibitem{Par:62}
%E.~Parzen.
%\newblock On estimation of a probability density and mode.
%\newblock {\em Annals of Mathematical Statistics}, 33:1065--1076, 1962.

%\bibitem{Sch:98}
%D.L. Schacter and R.L. Bucknew.
%\newblock On the relations among priming, conscious recollection,and
%  intentional retrieval: evidence from neuroimaging research.
%\newblock {\em Neurobiology of Learning and Memory}, 70(1):284--303, 1998.

%\bibitem{Vio:97}
%P.~Viola and W.M.~Wells III.
%\newblock Alignment by maximization of mutual information.
%\newblock {\em International Journal of Computer Vision}, 24(2):137--154, 1997.

%\bibitem{Wel:96}
%W.M.~Wells III, P.~Viola, H.~Atsumi, S.~Nakajima, and R.~Kikinis.
%\newblock Multi-modal volume registration by maximization of mutual
%  information.
%\newblock {\em Medical Image Analysis}, 1(1):35--51, 1996.

%\bibitem{Woo:94}
%G.K. Wood, B.A. Berkowitz, and C.A. Wilson.
%\newblock Visualization of subtle contrast-related intensity changes using
%  temporal correlation.
%\newblock {\em Magnetic Resonance Imaging}, 12(7):1013--1020, 1994.

%\end{thebibliography}

\end{document}







