Target detection performance bounds in compressive imaging

Krishnamurthy, Kalyani; Willett, Rebecca; Raginsky, Maxim

doi:10.1186/1687-6180-2012-205

Research
Open access
Published: 25 September 2012

Target detection performance bounds in compressive imaging

Kalyani Krishnamurthy¹,
Rebecca Willett¹ &
Maxim Raginsky²

EURASIP Journal on Advances in Signal Processing volume 2012, Article number: 205 (2012) Cite this article

2311 Accesses
7 Citations
12 Altmetric
Metrics details

Abstract

This article describes computationally efficient approaches and associated theoretical performance guarantees for the detection of known targets and anomalies from few projection measurements of the underlying signals. The proposed approaches accommodate signals of different strengths contaminated by a colored Gaussian background, and perform detection without reconstructing the underlying signals from the observations. The theoretical performance bounds of the target detector highlight fundamental tradeoffs among the number of measurements collected, amount of background signal present, signal-to-noise ratio, and similarity among potential targets coming from a known dictionary. The anomaly detector is designed to control the number of false discoveries. The proposed approach does not depend on a known sparse representation of targets; rather, the theoretical performance bounds exploit the structure of a known dictionary of targets and the distance preservation property of the measurement matrix. Simulation experiments illustrate the practicality and effectiveness of the proposed approaches.

Introduction

The theory of compressive sensing (CS) has shown that it is possible to accurately reconstruct a sparse signal from few (relative to the signal dimension) projection measurements[1, 2]. Though such a reconstruction is crucial to visually inspect the signal, there are many instances where one is solely interested in identifying whether the underlying signal is one of several possible signals of interest. In such situations, a complete reconstruction is computationally expensive and does not optimize the correct performance metric. Recently, CS ideas have been exploited in[3–5] to perform target detection and classification from projection measurements, without reconstructing the underlying signal of interest. In[3, 5], the authors propose nearest-neighbor based methods to classify a signal $f \in R^{N}$ to one of m known signals given projection measurements of the form $y = A f + n \in R^{K}$ for K≤N, where $A \in R^{K \times N}$ is a known projection operator and $n \sim N (0, σ^{2} I)$ is the additive Gaussian noise. This model is simple to analyze, but is impractical, since in reality, a signal is always corrupted by some kind of interference or background noise. Extension of the methods in[3, 5] to handle background noise is nontrivial. Though, Duarte et al.[4] provides a way to account for background contamination, it makes a strong assumption that the signal of interest and the background are sparse in bases that are incoherent. This might not always be true in many applications. Recent works on CS[6, 7] allow for the input signal f to be corrupted by some pre-measurement noise $b \sim N (0, σ_{b}^{2} I)$ such that one observes y=A(f + b) + n, and study reconstruction performance as a function of the number of measurements, pre- and post-measurement noise statistics and the dimension of the input signal. In this work, however, we are interested in performing target detection without an intermediate reconstruction step. Furthermore, the increased utility of high-dimensional imaging techniques such as spectral imaging or videography in applications like remote sensing, biomedical imaging and astronomical imaging[8–15] necessitates the extension of compressive target detection ideas to such imaging modalities to achieve reliable target detection from fewer measurements relative to the ambient signal dimensions.

For example, recent advances in CS have led to the development of new spectral imaging platforms which attempt to address challenges in conventional imaging platforms related to system size, resolution, and noise by acquiring fewer compressive measurements than spatiospectral voxels[16–21]. However, these system designs have a number of degrees of freedom which influence subsequent data analysis. For instance, the single-shot compressive spectral imager discussed in[18] collects one coded projection of each spectrum in the scene. One projection per spectrum is sufficient for reconstructing spatially homogeneous spectral images, since projections of neighboring locations can be combined to infer each spectrum. Significantly more projections are required for detecting targets of unknown strengths without the benefit of spatial homogeneity. We are interested in investigating how several such systems can be used in parallel to reliably detect spectral targets and anomalies from different coded projections.

In general, we consider a broadly applicable framework that allows us to account for background and sensor noise, and perform target detection directly from projection measurements of signals obtained at different spatial or temporal locations. The precise problem formulation is provided below.

Problem formulation

Let us assume access to a dictionary of possible targets of interest $D = {f^{(1)}, f^{(2)}, \dots, f^{(m)}}$ , where $f^{(j)} \in R^{N}$ for j=1,…,m is unit-norm. Our measurements are of the form

z_{i} = Φ (α_{i} f_{i}^{*} + b_{i}) + w_{i}

(1)

where

i∈{1,…,M} indexes the spatial or temporal locations at which data are collected;
α_i≥0 is a measure of the signal-to-noise ratio at location i, which is either known or estimated from observations;
$Φ \in R^{K \times N}$ for K < N, is a measurement matrix to be specified in Section “Whitening compressive observations”;

$b_{i} \in R^{N} \sim N (μ_{b}, Σ_{b})$ • is the background noise vector, and $w_{i} \in R^{K} \sim N (0, σ^{2} I)$ is the i.i.d. sensor noise.

For example, in the case of spectral imaging $f_{i}^{*}$ represents the spectrum at the ith spatial location, and in video sequences $f_{i}^{*}$ represents the vectorized image frame obtained at the ith time interval. In this article we consider the following target detection problems:

(1)
Dictionary signal detection (DSD): Here we assume that each $f_{i}^{*} \in D$ for i∈{1,…,M}, and our task is to detect all instances of one target signal $f^{(j)} \in D$ for some unknown j∈{1,…,m}, i.e., to locate $S = \{i : f_{i}^{*} = f^{(j)}\}$ . DSD is useful in contexts in which we know the makeup of a scene and wish to focus our attention on the locations of a particular signal. For instance, in spectral imaging, DSD is used to study a scene of interest by classifying every spectrum in the scene to different known classes [11, 22]. In a video setup, DSD could be used to classify video segments to one of several categories (such as news, weather, sports, etc.) by projecting the video sequence to an appropriate feature space and comparing the feature vectors to the ones in a known dictionary [23].
(2)
Anomalous signal detection (ASD): Here, our task is to detect all signals which are not members of our dictionary, i.e., detect $S = \{i : f_{i}^{*} \notin D\}$ (this is akin to anomaly detection methods in the literature which are based on nominal, nonanomalous training samples [24, 25]). For instance, ASD may be used when we know most components of a spectral image and wish to identify all spectra which deviate from this model [26].

Our goal is to accurately perform DSD or ASD without reconstructing the spectral input $f_{i}^{*}$ from z_i for i∈{1,…,M}. Accounting for background is a crucial issue. Typically, the background corresponding to the scene of interest and the sensor noise are modeled together by a colored multivariate Gaussian distribution[27]. However, in our case, it is important to distinguish the two because of the presence of the projection operator Φ. The projection operator acts upon the background spectrum in the same way as on the target spectrum, but it does not affect the sensor noise. We assume that b_iand w_iare independent of each other, and the prior probabilities of different targets in the dictionary $p^{(j)} = P (f_{i}^{*} = f^{(j)})$ for j∈{1,⋯,m} are known in advance. If these probabilities are unknown, then the targets can be considered equally likely. Given this setup, our goal is to develop suitable target and anomaly detection approaches, and provide theoretical guarantees on their performances.

In this article, we develop detection performance bounds which show how performance scales with the number of detectors in a compressive setting as a function of SNR, the similarity between potential targets in a known dictionary, and their prior probabilities. Our bounds are based on a detection strategy which operates directly on the collected data as opposed to first reconstructing each $f_{i}^{*}$ and then performing detection on the estimated signals. Reconstruction as an intermediate step in detection may be appealing to end users who wish to visually inspect spectral images instead of relying entirely on an automatic detection algorithm. However, using this intermediate step has two potential pitfalls. First, the Rao–Blackwell theorem[28] tells us that an optimal detection algorithm operating on the processed data (i.e., not sufficient statistics) cannot perform better than an optimal detection algorithm operating on the raw data. In other words, optimal performance is possible on the raw data, but we have no such performance guarantee for the reconstructed signals. Second, the relationship between reconstruction errors and detection performance is not well understood in many settings. Although we do not reconstruct the underlying signals, our performance bounds are intimately related to the signal resolution needed to achieve the signal diversity present in our dictionary. Since we have many fewer observations than the signals at this resolution, we adopt the “compressive” terminology.

Performance metric

To assess the performance of our detection strategies, we consider the false discovery rate (FDR) metric and related quantities developed for multiple hypothesis testing problems[29]. Since we collect M independent observations of potentially different signals, we are simultaneously conducting M hypothesis tests when we search for targets. Unlike the probability of false alarm, which measures the probability of falsely declaring a target for a single test, the FDR measures the fraction of declared targets that are false alarms, that is, it provides information about the entire set of M hypotheses instead of just one. More formally, the FDR is given by,

FDR = E [\frac{V}{R}],

where V is the number of falsely rejected null hypotheses, and R is the total number of rejected null hypotheses. Controlling the FDR in a multiple hypothesis testing framework is akin to designing a constant false alarm rate (CFAR) detector in spectral target detection applications that keeps the false alarm rate at a desired level irrespective of the background interference and sensor noise statistics[22].

Previous investigations

Much of the classical target detection literature[30–34] assume that each target lies in a P-dimensional subspace of $R^{N}$ for P < N. The subspace in which the target lies is often assumed to be known or specified by the user, and the variability of the background is modeled using a probability distribution. Given knowledge of the target subspace, background statistics and sensor noise statistics, detection methods based on LRTs (likelihood ratio tests) and GLRTs (generalized likelihood ratio tests) have been proposed in[30–35]. A subspace model is optimal if the subspace in which targets lie is known in advance. However, in many applications, such subspaces might be hard to characterize. An alternative, and a more flexible option is to assume that the high-dimensional target exhibits some low-dimensional structure that can be exploited to perform efficient target detection. This approach is utilized in this work and in[5] where the target signal in $R^{N}$ is assumed to come from a dictionary of m known signals such that m≪N, and in[3], where the targets are assumed to lie in a low-dimensional manifold embedded in high-dimensional target space.

Recently, several methods for target or anomaly detection that rely on recovering the full spatiospectral data from projection measurements[36, 37] have been proposed. However, they are computationally intensive and the detection performance associated with these reconstructions is unknown. Other researchers have exploited CS to perform target detection and classification without reconstructing the underlying signal[3–5]. Duarte et al.[4] propose a matching pursuit based algorithm, called the incoherent detection and estimation algorithm (IDEA), to detect the presence of a signal of interest against a strong interfering signal from noisy projection measurements. The algorithm is shown to perform well on experimental data sets under some strong assumptions on the sparsity of the signal of interest and the interfering signal. Davenport et al.[3] develop a classification algorithm called the smashed filter to classify an image in $R^{N}$ to one of m known classes from K projections of the signal, where K < N. The underlying image is assumed to lie on a low-dimensional manifold, and the algorithm finds the closest match from the m known classes by performing a nearest neighbor search over the m different manifolds. The projection measurements are chosen to preserve the distances among the manifolds. Though Davenport et al.[3] offers theoretical bounds on the number of measurements necessary to preserve distances among different manifolds, it is not clear how the performance scales with K or how to incorporate background models into this setup. Moreover, this approach may be computationally intensive since it involves learning and searching over different manifolds. Haupt et al.[5] use a nearest-neighbor classifier to classify an N-dimensional signal to one of m equally likely target classes based on K < N random projections, and provide theoretical guarantees on the detector performance. While the method discussed in[5] is computationally efficient, it is nontrivial to extend to the case of target detection with colored background noise and nonequiprobable targets. Furthermore, their performance guarantees cannot be directly extended to our problem since we focus on error measures that let us analyze the performance of multiple hypothesis tests simultaneously as opposed to the above methods that consider compressive classification performance for a single hypothesis test.

The authors of a more recent work[38] extend the classical RX anomaly detector[39] to directly detect anomalies from random, orthonormal projection measurements without an intermediate reconstruction step. They numerically show how the detection probability improves as a function of the signal-to-noise ratio when the number of measurements changes. Though probability of detection is a good performance measure, in many applications controlling the false discoveries below a desired level is more crucial. As a result, in our work, we propose an anomaly detection method that controls the FDR below a desired level.

Contributions

This article makes the following contributions to the above literature:

A compressive target detection approach, which (a) is computationally efficient, (b) allows for the signal strengths of the targets to vary with spatial location, (c) allows for backgrounds mixed with potential targets, (d) considers targets with different a priori probabilities, and (e) yields theoretical guarantees on detector performance. This article unifies preliminary work by the authors[40, 41], presents previously unpublished aspects of the proofs, and contains updated experimental results.
A computationally efficient anomaly detection method that detects anomalies of different strengths from projection measurements and also controls the FDR at a desired level.
A whitening filter approach to compressive measurements of signals with background contamination, and associated analysis leading to bounds on the amount of background to which our detection procedure is robust.

The above theoretical results, which are the main focus of this article, are supported with simulation studies in Section “Experimental results”. Classical detection methods described in[22, 26, 27, 30–35, 39, 42–45] do not establish performance bounds as a function of signal resolution or target dictionary properties and rely on relatively direct observation models which we show to be suboptimal when the detector size is limited. The methods in[3, 4] do not contain performance analysis, and our analysis builds upon the analysis in[5] to account for several specific aspects of the compressive target detection problem.

Whitening compressive observations

Before we present our detection methods for DSD and ASD problems, respectively, we briefly discuss a whitening step that is common to both our problems of interest.

Let us suppose that there are enough background training data available to estimate the background mean μ_b and covariance matrix Σ_b. We can assume without loss of generality that μ_b=0 since Φ μ_b can be subtracted from y. Given the knowledge of the background statistics, we can transform the background and sensor noise model $Φ b_{i} + w_{i} \sim N (0, Φ Σ_{b} Φ^{T} + σ^{2} I)$ discussed in (1) to a simple white Gaussian noise model by multiplying the observations z_i, i∈{1,…,M}, by the whitening filterC_Φ≜(Φ Σ_bΦ^T + σ²I)^−1/2. This whitening transformation reduces the observation model in (1) to

\begin{align} y_{i} = C_{Φ} \underset{z_{i}}{\underset{⏟}{(Φ (α_{i} f_{i}^{*} + b_{i}) + w_{i})}} = α_{i} A f_{i}^{*} + n_{i} \end{align}

(2)

where

\begin{align} A = C_{Φ} Φ, \end{align}

(3)

and $n_{i} = C_{Φ} (Φ b_{i} + w_{i}) \sim N (0, I)$ . To verify that $n_{i} \sim N (0, I)$ , observe that

n_{i} = C_{Φ} (Φ b_{i} + w_{i}) \sim N (0, \underset{I}{\underset{⏟}{C_{Φ} (Φ Σ_{b} Φ^{T} + σ^{2} I) C_{Φ}^{T}}}) .

We can now choose Φ so that the corresponding A has certain desirable properties as detailed in Sections “Dictionary signal detection” and “Anomalous signal detection”.

For a given A, the following theorem provides a construction of Φ that satisfies (3) and a bound on the maximum tolerable background contamination:

Theorem 1

Let B=I−A Σ_bA^T. If the largest eigenvalue of Σ_b satisfies

\begin{align} λ_{max} < \frac{1}{∥ A ∥^{2}}, \end{align}

(4)

where ∥A∥ is the spectral norm of A, then B is positive definite and Φ=σB^−1/2A is a sensing matrix, which can be used in conjunction with a whitening filter to produce observations modeled in (2).

The proof of this theorem is provided in Appendix 1. This theorem draws an interesting relationship between the maximum background perturbation that the system can tolerate and the spectral norm of the measurement matrix, which in turn varies with K and N. Hardware designs such as those in[17, 19] use spatial light modulators and digital micro mirrors, which allow the measurement matrix Φ to be adjusted easily in response to changing background statistics and other operating conditions.

In the sections that follow, we consider collecting measurements of the form $y_{i} = α_{i} A f_{i}^{*} + n_{i}$ given in (2), where $f_{i}^{*}$ is the target of interest for i=1,…,M, and $A \in R^{K \times N}$ is a sensing matrix that satisfies (3). It is assumed that any background contamination has been eliminated with the whitening procedure described in this section.

Dictionary signal detection

Suppose that the end user wants to test for the presence of one known target versus the rest, but it is not known a priori which target from $D$ the user wants to detect. In this case, let us cast the DSD problem as a multiple hypothesis testing problem of the form

\begin{align} ℋ_{0 i}^{(j)} : f_{i}^{*} = f^{(j)} vs. ℋ_{1 i}^{(j)} : f_{i}^{*} \neq f^{(j)} \end{align}

(5)

where $f^{(j)} \in D$ is the target of interest and i=1,…,M.

Decision rule

We define our decision rule corresponding to target $f^{(j)} \in D$ in terms of a set of significance regions $Γ_{i}^{(j)}$ such that one rejects the ith null hypothesis if its test statistic y_i falls in the ith significance region. Specifically, $Γ_{i}^{(j)}$ is defined according to

\begin{align} Γ_{i}^{(j)} = \{y : log P (f_{i}^{*} = f^{(j)} | y_{i}, α_{i}, A) \leq \\ log P (f_{i}^{*} = f^{(ℓ)} | y_{i}, α_{i}, A) for some ℓ \in {1, \dots, m}, ℓ \neq j \}, \end{align}

(6)

where $log P (f_{i}^{*} = f^{(j)} | y_{i}, α_{i}, A) = \frac{K}{2} log (\frac{1}{2 Π}) - \frac{{∥y_{i} - α_{i} A f^{(j)}∥}^{2}}{2} + log p^{(j)}$ is the logarithm of the a posteriori probability density of the target f^(j) at the ith spatial location given the observations y_i, the signal-to-noise ratio α_iand the sensing matrix A, and p^(j) is the a priori probability of target class j. Note that the process of determining these decision regions involves a sequence of nearest-neighbor calculations, so the computational complexity scales with the number of classes m. In this work, we operate under the assumption that m is much smaller than the dimensionality of the datasets we consider. For example, if we consider spectral images, then the number of objects (signal classes) that make up a scene of interest is often smaller than the number of voxels in the image. This assumption is not unrealistic and has been exploited in earlier work such as[22] and the references therein. In most of the prior work we have surveyed[46, 47], the number of signal classes is less than 35, which doesn’t make our approach intractable.

The decision rule can be formally expressed in terms of the significance regions as follows:

\begin{align} reject ℋ_{0 i}^{(j)} if the test statistic y_{i} \in Γ_{i}^{(j)} . \end{align}

(7)

We analyze this detector by extending the positive FDR (pFDR) error measure introduced by Storey to characterize the errors encountered in performing multiple, independent and nonidentical hypothesis tests simultaneously[48]. The pFDR, discussed formally below, is the fraction of falsely rejected null hypotheses among the total number of rejected null hypotheses, subject to the positivity condition that one rejects at least one null hypothesis. The pFDR is similar to the FDR except that the positivity condition is enforced here. In our context, the positivity condition means that we declare at least one signal to be a nontarget, which in turn implies that the scene of interest is composed of more than one object in the case of spectral imaging, or that the scene is not static in the case of video imaging.

Consider a collection of significance regions $Γ = \{Γ_{i}^{(j)} : i = 1, \dots, M\}$ , such that one declares $ℋ_{1 i}^{(j)}$ if the test statistic $y_{i} \in Γ_{i}^{(j)}$ . The pFDR for multiple, nonidentical hypothesis tests can be defined in terms of the significance regions as follows:

{pFDR}^{(j)} (Γ) = E [\frac{V (Γ)}{R (Γ)}| R (Γ) > 0]

(8)

where

V (Γ) = \sum_{i = 1}^{M} I_{\{y_{i} \in Γ_{i}^{(j)}\}} I_{\{ℋ_{0 i}\}}

(9)

is the number of falsely rejected null hypotheses,

R (Γ) = \sum_{i = 1}^{M} I_{\{y_{i} \in Γ_{i}^{(j)}\}}

(10)

is the total number of rejected null hypotheses, and $E_{{E}} = 1$ if event E is true and 0 otherwise. In our setup, the pFDR corresponds to the expected ratio of the number of missed targets to the number of signals declared to be nontargets subject to the condition that at least one signal is declared to be a nontarget (note that this ratio is traditionally referred to as the positive false nondiscovery rate (pFNR), but is technically the pFDR in this context because of our definitions of the null and alternate hypotheses). The theorem below presents our main result:

Theorem 2

Given observations of the form (2), if one performs multiple, independent, nonidentical hypothesis tests of the form (5) and decides according to (7), then the worst-case pFDR given by pFDR_max=max_{j∈{1,…,m}}pFDR^(j)(Γ), satisfies the following bound:

\begin{align} {pFDR}_{max} \leq \min (1, \frac{{(P_{e})}_{max}}{1 - p_{max} - {(P_{e})}_{max}}) \end{align}

(11)

where

\begin{align} p_{max} = max_{j \in {1, \dots, m}} p^{(j)}, \\ {(P_{e})}_{max} = max_{i \in {1, \dots, M}} P ({\hat{f}}_{i} \neq f_{i}^{*}), and \\ {\hat{f}}_{i} = {arg max}_{f \in D} P (f_{i}^{*} = f| y_{i}, α_{i}, A) . \end{align}

(12)

The proof of this theorem is detailed in Appendix 2. A key element of our proof is the adaptation of the techniques from[48] to nonidentical independent hypothesis tests.

An achievable bound on the worst-case pFDR

Theorem 2 in the preceding section shows that, for a given A, the worst-case pFDR is bounded from above by a function of the worst-case misclassification probability. In this section, we use this theorem to establish an achievable bound on the worst-case pFDR that explicitly depends on the number of observations K, signal strengths ${α_{i}}_{i = 1}^{M}$ , similarity among different targets of interest, and a priori target probabilities.

Let us first define the quantities

\begin{align} d_{min} = min_{f^{(i)}, f^{(j)} \in D, i \neq j} ∥ f^{(i)} - f^{(j)} ∥ \\ p_{min} = min_{j \in \{1, \dots, m\}} p^{(j)} \\ α_{min} = min_{i \in {1, \dots, M}} α_{i} . \end{align}

Then we have the following theorem, whose proof is given in Appendix 3:

Theorem 3

Let λ_max denote the largest eigenvalue of Σ_b. For a given 0 < ε < 1−p_max, assume that K and N are sufficiently large so that the following conditions hold:

\begin{align} 1 - p_{max} - ε \geq \frac{1 - p_{min}}{p_{min}} {(1 + \frac{{α_{min}}^{2} d_{min}^{2}}{4 K σ^{2}})}^{- \frac{K}{2}} \\ + 2 exp (- \frac{(K + N) ε^{2}}{2}) \end{align}

(13a)

\begin{align} λ_{max} < \frac{1}{{(1 + ε)}^{2} {(\sqrt{\frac{N}{K}} + 1)}^{2}}, \end{align}

(13b)

\begin{align} K > \frac{2 log (\frac{2}{p_{min}} \frac{1 - p_{\min}}{1 - p_{\max}})}{log (1 + \frac{α_{\min}^{2} d_{\min}^{2}}{4 K})} . \end{align}

(13c)

Then there exists a K×N sensing matrix A that satisfies the condition of Theorem 1, and for which

\begin{align} {pFDR}_{max} \leq \frac{1}{p_{min}} {(\frac{1 - p_{max}}{1 - p_{min}} {(1 + \frac{α_{min}^{2} d_{min}^{2}}{4 K})}^{\frac{K}{2}} - \frac{1}{p_{min}})}^{- 1} \\ + \frac{2 (1 - p_{max})}{ε^{2}} exp (- \frac{(K + N) ε^{2}}{2}) . \end{align}

(14)

This result has the following implications and consequences:

(1)
For a given N, the upper bound (13b) on λ_maxincreases as K increases, which implies that the system can tolerate more background perturbation if we collect more measurements.
(2)
The pFDR bound (14) decays with the increase in the values of K, d_minand α_min, and increases as p_mindecreases. For a fixed p_max, p_min, α_minand d_min, the bound in (14) enables one to choose a value of K to guarantee a desired pFDR value.
(3)
The dominant part of the bound (14) is independent of N, and is only a function of K, p_max, p_min, α_min, and d_min. The lack of dependence on N is not unexpected. Indeed, when we are interested in preserving pairwise distances among the members of a fixed dictionary of size m, the Johnson–Lindenstrauss lemma [49] says that, with high probability, $K = O (log m)$ random Gaussian projections suffice, regardless of the ambient dimension N. This is precisely the regime we are working with here.
(4)
The bound on K given in (13c) increases logarithmically with the increase in the difference between p_max and p_min. This is to be expected since one would need more measurements to detect a less probable target as our decision rule weights each target by its a priori probability. If all targets are equally likely, then p_max=p_min=1/m, and $K = O (log m)$ is sufficient provided $α_{min}^{2} d_{min}^{2}$ is sufficiently large such that
$\begin{align} log (1 + \frac{α_{min}^{2} d_{min}^{2}}{4 K}) > log (1 + \frac{α_{min}^{2} d_{min}^{2}}{4 N}) > 1 \end{align}$

(where the first inequality holds since K < N). In addition, the lower bound on K also illustrates the interplay between the signal strength of the targets, the similarity among different targets in $D$ , and the number of measurements collected. A small value of d_min suggests that the targets in $D$ are very similar to each other, and thus α_minand K need to be high enough so that similar targets can still be distinguished. The experimental results discussed in Section “Experimental results” illustrate the tightness of the theoretical results discussed here.

Inspection of the proof shows that if A is generated according to a Gaussian distribution, then the conditions of Theorem 3 will be met with high probability.

Extension to a manifold-based target detection framework

The DSD problem formulation in Section “ASD problem formulation” is accurate if the signals in the dictionary are faithful representations of the target signals that we observe. In reality, however, the target signals will differ from the dictionary signals owing to the differences in the experimental conditions under which they are collected. For instance, in spectral imaging applications, the observed spectrum of any material will not match the reference spectrum of the same material observed in a laboratory because of the differences in atmospheric and illumination conditions. To overcome this problem, one could form a large dictionary to account for such uncertainties in the target signals and perform target detection according to the approaches discussed in Sections “Whitening compressive observations” and “Dictionary signal detection”. A potential drawback with this approach is that our theoretical performance bound increases with the size of $D$ through p_min and d_min. Instead, one could reasonably model the target signals observed under different experimental conditions to lie in a low-dimensional submanifold of the high-dimensional ambient signal space as shown to be true for spectral images in[50]. We can exploit this result to extend our analysis to a much broader framework that accounts for uncertainties in our dictionary.

Let us consider a dictionary of manifolds $D_{ℳ} = \{ℳ^{(1)}, \dots, ℳ^{(m)}\}$ corresponding to m different target classes, and that $f_{i}^{*}$ for i∈{1,…,M} is in one of the manifolds in $D_{ℳ}$ . Considering an observation model of the form given in (2), our goal is to determine $\{i : f_{i}^{*} \in ℳ^{(j)}\}$ , where j∈{1,…,m} is the target class of interest. Let us assume that all target classes are equally likely to keep the presentation simple, though the analysis extends to the case where the targets classes have different a priori probabilities. Suppose that we collect independent sets of measurements ${\{y_{i}\}}_{i = 1}^{M}$ and ${\{{\tilde{y}}_{i}\}}_{i = 1}^{M}$ . Then, we can use the following two-step procedure to extend our DSD method to this manifold-based framework:

(1)
Given {y _i}, form a data-dependent dictionary $D_{y_{i}} = \{{\tilde{f}}_{i}^{(1)}, \dots, {\tilde{f}}_{i}^{(m)}\}$ corresponding to each y _i by finding its nearest-neighbor in each manifold:
${\tilde{f}}_{i}^{(ℓ)} = {arg max}_{f \in ℳ^{(ℓ)}} P (y_{i}| f_{i}^{*} = f, α_{i}, A)$

for ℓ∈{1,…,m} and i=1,…,M.
(2)
Given $\{{\tilde{y}}_{i}\}$ and corresponding $\{D_{y_{i}}\}$ , find
${\hat{f}}_{i} = {arg max}_{\tilde{f} \in D_{y_{i}}} P ({\tilde{y}}_{i}| f_{i}^{*} = \tilde{f}, α_{i}, A)$

and declare that the ith observed spectrum corresponds to class j if ${\hat{f}}_{i} = {\tilde{f}}_{i}^{(j)}$ .

This two-step procedure is studied in[3] for the case $\{y_{i}\} = \{{\tilde{y}}_{i}\}$ where the authors provide bounds on the number of projection measurements needed to preserve distances among manifolds. However, they do not offer associated target detection performance guarantees. Our analysis and the theoretical performance bounds extend directly to this framework, if we collect two sets of observations as discussed above. Specifically, the hypothesis tests corresponding to the second step can be written as

ℋ_{0 i} : f_{i}^{*} = {\tilde{f}}_{i}^{(j)} vs. ℋ_{1 i} : f_{i}^{*} \neq {\tilde{f}}_{i}^{(j)}

where ${\tilde{f}}_{i}^{(j)} \in D_{y_{i}}$ for i=1,…,M. Since the dictionary in this case changes with i, these tests are nonidentical. This is another instance where our extension of pFDR-based analysis towards simultaneous testing of multiple, independent, and nonidentical hypothesis tests (8) is very significant. Following the proof techniques discussed in the appendix, we can straightforwardly show that the bound in (14) in this manifold setting holds with p_min=p_max=1/m since all target classes are assumed to be equally likely here, and d_min=min_{i∈{1,…,M}}d_iwhere

d_{i} = min_{{\tilde{f}}_{i}^{(ℓ)}, {\tilde{f}}_{i}^{(k)} \in D_{y_{i}}, ℓ \neq k} ∥ {\tilde{f}}_{i}^{(ℓ)} - {\tilde{f}}_{i}^{(k)} ∥ .

Anomalous signal detection

The target detection approach discussed above assumes that the target signal of interest resides in a dictionary that is available to the user. However, in some applications (such as military applications and surveillance), one might be interested in detecting objects not in the dictionary. In other words, the target signals of interest are anomalous and are not available to the user. In this section, we show how the target detection methods discussed above can be extended to anomaly detection. In particular, we exploit the distance preservation property of the sensing matrix A to detect anomalous targets from projection measurements.

ASD problem formulation

Given observations of the form in (2), we are interested in detecting whether $f^{*} \in D$ or f^∗is anomalous. Let us write the anomaly detection problem as the following multiple hypothesis test:

\begin{align} ℋ_{0 i} & : ∥ f_{i}^{*} - f ∥ \leq τ for some f \in D \end{align}

(15a)

\begin{align} ℋ_{1 i} & : ∥ f_{i}^{*} - f ∥ > τ for all f \in D \end{align}

(15b)

where $τ \in [0, \sqrt{2})$ is a user-defined threshold that encapsulates our uncertainty about the accuracy with which we know the dictionary.^a In particular, τ controls how different a signal needs to be from every dictionary element to truly be considered anomalous. In the absence of any prior knowledge on the targets of interest, τ can simply be set to zero. The null hypothesis in this setting models the normal behavior, while the alternative hypothesis models the abnormal or anomalous behavior. This formulation is consistent with the literature[26, 38].

Note that the definition of the hypotheses given in (15a) and (15b) matches the definition in (5) for the special case where the dictionary contains just one signal. In this special case, the signal input f^∗ is in the dictionary under the null hypothesis in both DSD and ASD problem formulations.^b

Anomaly detection approach

Our anomaly detection approach and the associated theoretical analysis are based on a “distance preservation” property of A, which is stated formally in (18). We propose an anomaly detection method that controls the FDR below a desired level δ for different background and sensor noise statistics. In other words, we control the expected ratio of falsely declared anomalies to the total number of signals declared to be anomalous. Note that here we work with the FDR as opposed to the pFDR, since it is possible for a scene to not contain any anomalies at all. We let V/R=0 for R=V=0 since one does not declare any signal to be anomalous in this case. In[29], Benjamini and Hochberg discuss a p-value based procedure, “BH procedure”, that controls the FDR of M independent hypothesis tests below a desired level. Let,

d_{i} = min_{f \in D} ∥ y_{i} - α_{i} A f ∥ = min_{f \in D} ∥ α_{i} A (f_{i}^{*} - f) + n_{i} ∥

(16)

be the test statistic at the ith location. The p-value can be defined in terms of our test statistic as follows:

p_{i} = P ({\tilde{d}}_{i} \geq d_{i} | ℋ_{0 i})

(17)

where ${\tilde{d}}_{i} = {min}_{f \in D} ∥ α_{i} A (f_{i}^{*} - f) + n ∥$ and $n \sim N (0, I)$ is independent of n_i. This is the probability under the null hypothesis, of acquiring a test statistic at least as extreme as the one observed. Let us denote the ordered set of p-values by p₍₁₎≤p₍₂₎≤⋯≤p_(M) and let $ℋ_{(0 i)}$ be the null hypothesis corresponding to (i)^thp-value. The BH procedure says that if we reject all $ℋ_{(0 i)}$ for i=1,…,t where t is the largest i for which p_(i)≤iδ/M, then the FDR is controlled at δ.

To apply this procedure in our setting, we need to find a tractable expression for the p-value at every location. This can be accomplished when A satisfies the distance-preservation condition stated below. Let $V = D ⋃ {f_{i}^{*} : i \in {1, \dots, M}}$ be the set of all signals in the dictionary and the ones whose projections are measured. Note that |V|≤M + m. For a given ε∈(0,1), a projection operator $A \in R^{K \times N}$ , K≤N, is distance-preserving on V if the following holds for all u,v∈V:

(1 - ε) ∥ u - v ∥ \leq ∥ A (u - v) ∥ \leq (1 + ε) ∥ u - v ∥, \forall u, v \in V.

(18)

The existence of such projection operators is guaranteed by the celebrated Johnson and Lindenstrauss (JL) lemma[49], which says that there exists random constructions of A for which (18) holds with probability at least 1−2|V|²e^−Kc(ε)provided $K = O (log | V |) \leq N$ , where c(ε)=ε²/16−ε³/48[51, 52]. Examples of such constructions are: (a) Gaussian matrices whose entries are drawn from $N (0, 1 / K)$ , (b) Bernoulli matrices whose entries are $\pm 1 / \sqrt{N}$ with probability 1/2, (c) random matrices whose entries are $\pm \sqrt{3 / N}$ with probability 1/6 and zero with probability 2/3[51, 52], and (d) matrices that satisfy the restricted isometry property (RIP) where the signs of the entries in each column are randomized[53].

We now state our main theorem that gives a tight upper bound on the p-value at every location when {α_i} are unknown and are estimated from the observations. Let ${{\hat{α}}_{i}}$ be the estimates of {α_i} that satisfy

1 - ζ \leq \frac{α_{i}}{{\hat{α}}_{i}} \leq 1 + ζ

(19)

for i=1,…,M where ζ∈[0,1] is a measure of the accuracy of the estimation procedure.

Theorem 4

If the ith hypothesis test is defined according to (15a) and (15b), the projection matrix A satisfies (18) for a given ε∈(0,1), and the estimates ${{\hat{α}}_{i}}$ satisfy (19) for some ζ∈[0,1], then the bound

p_{i} \leq 1 - F (d_{i}^{2}; K, {(1 + ε)}^{2} {\hat{α}}_{i}^{2} {(ζ + τ)}^{2})

(20)

holds for all i=1,…,M where $F (\cdot; K, ν)$ is the CDF of a noncentral χ²random variable with K degrees of freedom and noncentrality parameter ν[54].

The proof of this theorem is given in Appendix 4. We find the p-value upper bounds at every location and use the BH procedure to perform anomaly detection. The performance of this procedure depends on the values of K, {α_i}, τ and ε. The parameter ε is a measure of the accuracy with which the projection matrix A preserves the distances between any two vectors in $R^{N}$ . A value of ε close to zero implies that the distances are preserved fairly accurately. When {α_i} are unknown and estimated from the observations, the performance depends on the accuracy of the estimation procedure, which is reflected in our bounds in (20) through ζ.

One can easily estimate {α_i} from {y_i} for some choices of A. For instance, if the entries of the projection matrix A are drawn from $N (0, 1 / K)$ , the {α_i} can be estimated using a maximum likelihood estimator (MLE) by exploiting the statistics of the projection matrix and noise. Note that the jth element of the ith measured spectrum is $y_{i, j} = \sum_{k = 1}^{N} α_{i} f_{i, k}^{*} a_{j, k} + n_{i, j} \sim N (0, \sum_{k = 1}^{N} \frac{α_{i}^{2}}{K} {f_{i, k}^{*}}^{2} + 1)$ for j∈{1,…,K}. Since ${∥f_{i}^{*}∥}_{2} = 1$ according to our problem formulation, $y_{i, j} \overset{i.i.d.}{\sim} N (0, \frac{α_{i}^{2}}{K} + 1)$ . The MLE of α_i given by ${\hat{α}}_{i} = {arg max}_{α} P (y_{i} | A, α)$ then reduces to

{\hat{α}}_{i} = \sqrt{(∥ y_{i} ∥^{2} - K)} .

(21)

In practice, we use ${\hat{α}}_{i} = \sqrt{{(∥ y_{i} ∥^{2} - K)}_{+}}$ where the (a)₊ =a if a≥0 and 0 otherwise to ensure that ∥y_i∥²−K is nonnegative. We can use concentration inequalities to show that with high probability, ${∥y_{i}∥}_{2}^{2}$ is tightly concentrated around its mean $E [{∥y_{i}∥}_{2}^{2}] = α_{i}^{2} + K$ . Since $y_{i, j} \overset{i.i.d.}{\sim} N (0, \frac{α_{i}^{2}}{K} + 1)$ , $\frac{K}{α^{2} + K} {∥y_{i}∥}_{2}^{2} \sim χ_{K}^{2}$ . From ([55], Lemma 2.2), and ([56], Proposition 1 and Remark 1), for any t > 0

P (|{∥y_{i}∥}_{2}^{2} - (α_{i}^{2} + K)| \geq t) \leq C exp (- c t^{2})

(22)

for some absolute constants C,c > 0. This result shows that with high probability, ${∥y_{i}∥}_{2}^{2} - K$ is nonnegative.

The experimental results discussed in Section “Experimental results” demonstrate the performance of this detector as a function of K, {α_i} and τ when {α_i} are known and as a function of K, τ and ζ when {α_i} are estimated.

Experimental results

In the experiments that follow, the entries of A are drawn from $N (0, 1 / K)$ .

Dictionary signal detection

To test the effectiveness of our approach, we formed a dictionary $D$ of nine spectra (corresponding to different kinds of trees, grass, water bodies and roads) obtained from a labeled HyMap (Hyperspectral Mapper) remote sensing data set[57], and simulated a realistic dataset using the spectra from this dictionary. Each HyMap spectrum is of length N=106. We generated projection measurements of these data such that $z_{i} = α_{i} Φ (f_{i}^{*} + b_{i}) + w_{i}$ according to (1), where $w_{i} \sim N (0, σ^{2} I)$ , $f_{i}^{*} \in D$ for i=1,…,8100, $b_{i} \sim N (μ_{b}, Σ_{b})$ such that Σ_b satisfies the condition in (4), and $α_{i} = α_{i}^{*} \sqrt{K}$ where $α_{i}^{*} \sim U [21, 25]$ and $U$ denotes uniform distribution. We let σ²=5 and model {α_i} to be proportional to $\sqrt{K}$ to account for the fact that the total observed signal energy increases as the number of detectors increases. We transform the z_iby a series of operations to arrive at a model of the form discussed in (2), which is $y_{i} = α_{i} A f_{i}^{*} + n_{i}$ . For this dataset, p_min=0.04938, p_max=0.1481, and d_min=0.04341.

We evaluate the performance of our detector (7) on the transformed observations, relative to the number of measurements K, by comparing the detection results to the ground truth. Our MAP detector returns a label $L_{i}^{MAP}$ for every observed spectrum which is determined according to

L_{i}^{MAP} = {arg min}_{ℓ \in {1, \dots, m}, f^{(ℓ)} \in D} (\frac{1}{2} | | y_{i} - α_{i} A f^{(ℓ)} | |^{2} - log p^{(ℓ)})

where m is the number of signals in $D$ , and p^(ℓ) is the a priori probability of target class ℓ. In our experiments we evaluate the performance of our classifier when (a) {α_i} are known (AK) and (b) {α_i} are unknown (AU) and must be estimated from y, respectively. The empirical pFDR^(j)for each target spectrum j is calculated as follows:

{pFDR}^{(j)} = \frac{\sum_{i = 1}^{M} I_{\{L_{i}^{GT} = j\}} I_{\{L_{i}^{MAP} \neq j\}}}{\sum_{i = 1}^{M} I_{\{L_{i}^{MAP} \neq j\}}}

where ${L_{i}^{GT}}$ denote the ground truth labels. The empirical pFDR^(·)is the ratio of the number of missed targets to the total number of signals that were declared to be nontargets. The plots in Figure1a show the results obtained using our target detection approach under the AK case (shown by a dark gray dashed line) and the AU case (shown by a light gray dashed line), compared to the theoretical upper bound (shown by a solid line). These results are obtained by averaging the pFDR values obtained over 1000 different noise, sensing matrix and background realizations. Note that theoretical results only apply to the AK case since they were derived under the assumption of {α_i} being known. The experimental results are shown for both AK and AU cases to provide a comparison between the two scenarios. In both these cases, the worst-case empirical pFDR curves decay with the increase in the values of K. In the AK case, in particular, the worst-case empirical pFDR curve decays at the same rate as the upper bound. In this experiment, for a fixed α_minand d_min, we chose K to satisfy (13c). The theory is somewhat conservative, and in practice the method works well even when the values of K are below the bound in (13c).

In the experiment that follows, we let $α_{i}^{*} \sim U [10, 20]$ , where $U$ denotes a uniform random variable, $α_{i} = \sqrt{K} α_{i}^{*}$ and evaluate the performance of our detector for different values of K that are not necessarily chosen to satisfy (13c). In addition, we also compare the performance of our detection method to that of a MAP based target detector operating on downsampled versions of our simulated spectral input image. The reason behind such a comparison is to show what kinds of measurements yield better results given a fixed number of detectors.

For an input spectrum $g \in R^{N}$ , we let $\tilde{g} \in R^{K}$ denote its downsampled approximation. Specifically, the j th element of ${\tilde{g}}_{i}$ is $\sum_{ℓ = 1}^{r} g_{(j - 1) r + ℓ}$ where r=⌈N/K⌉. Let us consider making observations of the form

y_{i} = \frac{{\tilde{g}}_{i}}{c} + n_{i} \in R^{K}

(23)

where ${\tilde{g}}_{i} = α_{i} {\tilde{f}}_{i}^{*} + {\tilde{b}}_{i}$ is the K-dimensional downsampled version of $f_{i}^{*} + b_{i}$ for K≤N, $n_{i} \sim N (0, σ^{2} I)$ for σ²=5 and c is a constant that is chosen to preserve the mean signal-to-noise ratio corresponding to the downsampled and projection measurements. The MAP-based detector operating on the downsampled data returns a label $D_{i}^{MAP}$ for every observed spectrum which is determined according to

\begin{align} D_{i}^{MAP} = \overset{arg min}{ℓ \in {1, \dots, m}, f^{(ℓ)} \in D} {(y_{i} - α_{i} {\tilde{f}}^{(ℓ)})}^{T} G^{- 1} \\ \times (y_{i} - α_{i} {\tilde{f}}^{(ℓ)}) - log p^{(ℓ)} \end{align}

where $G = {\tilde{Σ}}_{b} + σ^{2} I$ and ${\tilde{Σ}}_{b}$ is the covariance matrix obtained from the downsampled versions of the background training data and ${\tilde{f}}^{(ℓ)}$ is the downsampled version of $f^{(ℓ)} \in D$ . The algorithm declares that target spectrum $f^{(j)} \in D$ is present in the i th location if $D_{i}^{MAP} = j$ . In order to illustrate the advantages of using a Φ designed according to (24), we compare the performances of the proposed anomaly detector when Φ is chosen to be a random Gaussian matrix whose entries are drawn from $N (0, 1 / K)$ and when Φ is chosen according to (24). Figure1b shows a comparison of the results obtained using the projection measurements obtained using Φ designed according to (24), Φ chosen at random, and the downsampled measurements under the AK case. These results show that the detection algorithm operating on projection measurements using Φ designed using background and sensor noise statistics yield significantly better results than the one operating on the downsampled data, and that the empirical pFDR values in our method decays with K. The improvement in performance using projection measurements comes from the distance-preservation property of the projection operator A. While a Gaussian sensing matrix A preserves distances between any pair of vectors from a finite collection of vectors with high probability[51, 52], downsampling loses some of the fine differences between similar-looking spectra in the dictionary. Furthermore, when Φ is chosen at random, the resulting whitened transformation matrix is not necessarily distance-preserving. This may worsen the performance as illustrated in Figure1b.

Anomaly detection

In this section, we evaluate the performance of our anomaly detection method on (a) a simulated dataset and provide a comparison of the results obtained using the proposed projection measurements and the ones obtained using downsampled measurements, and (b) real AVIRIS (Airborne Visible InfraRed Imaging Spectrometer) dataset.

Experiments on simulated data

We simulate a spectral image f^∗composed of 8100 spectra, where each of them is either drawn from a dictionary $D = {f^{(1)}, \dots, f^{(5)}}$ consisting of five labeled spectra from the HyMap data that correspond to a natural landscape (trees, grass and lakes) or is anomalous. The anomalous spectrum is extracted from unlabeled AVIRIS data, and the minimum distance between the anomalous spectrum f^(a) and any of the spectra in $D$ is $d_{min} = {min}_{f \in D} ∥ f - f^{(a)} ∥ = 0.5308$ . The simulated data has 625 locations that contain the anomalous spectrum. Our goal is to find the spatial locations that contain the anomalous AVIRIS spectrum given noisy measurements of the form $z_{i} = Φ (α_{i} f_{i}^{*} + b_{i}) + w_{i}$ where b_i∼(μ_b,Σ_b), Φ is designed according to (24), $w_{i} \sim N (0, σ^{2} I)$ and $f_{i}^{*} \in D$ under $ℋ_{0 i}$ . As discussed in Section “Anomalous signal detection”, $f_{i}^{*}$ is anomalous under $ℋ_{1 i}$ , and our goal is to control the FDR below a user-specified false discovery level δ. We simulate ${α_{i}} = \sqrt{K} α_{i}^{*}$ where $α_{i}^{*} \sim U [2, 3]$ . In this experiment we assume the availability of background training data to estimate the background statistics and the sensor noise variance σ². Given the knowledge of the background statistics, we perform the whitening transformation discussed in Section “Whitening compressive observations” and evaluate the detection performance on the preprocessed observations given by (2).

For a fixed τ=0. 1 and ε=0. 1, we evaluate the performance of the detector as the number of measurements K increases under the AK and AU cases respectively, by comparing the pseudo-ROC (receiver operating characteristic) curves obtained by plotting the empirical FDR against 1−FNR, where FNR is the false nondiscovery rate. Note that 1−FNR is the expected ratio of the number of null hypotheses that are correctly rejected to the number of declared null hypotheses. The empirical FDR and FNR are computed according to

\begin{align} FDR = \frac{\sum_{i = 1}^{M} I_{\{L_{i}^{GT} = 0\}} I_{{p_{i} \leq p_{t}}}}{\sum_{i = 1}^{M} I_{{p_{i} \leq p_{t}}}} and \\ FNR = \frac{\sum_{i = 1}^{M} I_{\{L_{i}^{GT} = 1\}} I_{{p_{i} > p_{t}}}}{\sum_{i = 1}^{M} I_{{p_{i} > p_{t}}}} \end{align}

where p_tis the p-value threshold such that the BH procedure rejects all null hypotheses for which p_i≤p_t, and the ground truth label $L_{i}^{GT} = 0$ if the i th spectrum is not anomalous, and 1 otherwise. In this experiment, we consider three different values of K approximately given by K∈{N/6,N/3,N/2} where N=106, and evaluate the performance of our detector for each K. Furthermore, in our experiments with simulated data, we declare a spectrum to be anomalous if d_i≥η where η is a user-specified threshold and d_iis defined in (16). We use the p-value upper bound in (20) in our experiments with real data where the ground truth is unknown.

We compare the performance of our method to a generalized likelihood ratio test (GLRT)-based procedure operating on downsampled data, where we collect measurements of the form in (23) and $f_{i}^{*} \in D$ under $ℋ_{0 i}$ . Observe that $y_{i} | ℋ_{0 i} \sim \sum_{f \in D} P (f_{i}^{*} = f) N (α_{i} \tilde{f}, {\tilde{Σ}}_{b} + I)$ , where $\tilde{f}$ refers to the downsampled version of $f \in D$ . In this experiment we assume that each spectrum in $D$ is equally likely under $ℋ_{0 i}$ for i=1,…,M. The GLRT-based approach declares the i th spectrum to be anomalous if

\begin{align} - log P (y_{i} | ℋ_{0 i}) \overset{ℋ_{1 i}}{\underset{ℋ_{0 i}}{≷}} η \end{align}

for i=1,…,M, where η is a user-specified threshold[26]. While our anomaly detection method is designed to control the FDR below a user-specified threshold, the GLRT-based method is designed to increase the probability of detection while keeping the probability of false alarm as low as possible. To facilitate a fair evaluation of these methods, we compare the pseudo-ROC curves (FDR versus 1−FNR) and the actual ROC curves (probability of false alarm p_fversus probability of detection p_d) corresponding to these methods obtained by averaging the empirical FDR, FNR, p_d and p_f over 1,000 different noise and sensing matrix realizations for different values of K. We also compare the performance of the proposed method when Φ is chosen according to (24) and when it is chosen at random, as discussed in the previous section. Figure2a,e show the pseudo-ROC plots and the conventional ROC plots obtained using the GLRT-based method operating on downsampled data when {α_i} are known. Figure2b,f show the results obtained by using a random Gaussian Φ instead of the Φ in (24). Figure2c,g show the pseudo-ROC plots and the conventional ROC plots obtained using our method when {α_i} are known. These plots show that performing anomaly detection from our designed projection measurements yields better results than performing anomaly detection on downsampled measurements and on measurements obtained using a random Gaussian Φ. This is largely due to the fact that carefully chosen projection measurements preserve distances (up to a constant factor) among pairs of vectors in a finite collection, where as the downsampled measurements fail to preserve distances among vectors that are very similar to each other. Similarly, a random projection matrix Φ is not necessarily distance-preserving post-whitening transformation, which leads to poor performance as illustrated in Figure2b,f. Figure2d,h shows the pseudo-ROC plots and the conventional ROC plots obtained using our method when {α_i} are unknown, and are estimated from the measurements. Note that the value of ζ decreases as K increases since the estimation accuracy of {α_i} increases with increase in K. These plots show that the performance improves as we collect more observations, and that, as expected, the performance under the AK case is better than the performance under the AU case.

Experiments on real AVIRIS data

To test the performance of our anomaly detector on a real dataset, we consider the unlabeled AVIRIS Jasper Ridge dataset $g \in R^{614 \times 512 \times 197}$ , which is publicly available from the NASA AVIRIS website,http://aviris.jpl.nasa.gov/html/aviris.freedata.html. We split this data spatially to form equisized training and validation datasets, g^t and g^v respectively, each of which is of size 128×128×197. Figure3a,b show images of the AVIRIS training and validation data summed through the spectral coordinates. The training data are comprised of a rocky terrain with a small patch of trees. The validation data seems to be made of a similar rocky terrain, but also contain an anomalous lake-like structure. The goal is to evaluate the performance of the detector in detecting the anomalous region in the validation data for different values of K. We cluster the spectral targets in the normalized training data to eight different clusters using the K-means clustering algorithm and form a dictionary $D$ comprising of the cluster centroids. Given the dictionary and the validation data, we find the ground truth by labeling the i th validation spectrum as anomalous if ${min}_{f \in D} ∥f - \frac{g_{i}^{v}}{∥ g_{i}^{v} ∥}∥ > τ$ . Since the statistics of the possible background contamination in the data could not be learned in this experiment because of the lack of labeled training data, the dictionary might be background contaminated as well. The parameter τ encapsulates this uncertainty in our knowledge of the dictionary. In this experiment, we set τ=0. 2.

We generate measurements of the form $y_{i} = \sqrt{K} g_{i}^{v} + n_{i}$ for i=1,…,128×128, where $n_{i} \sim N (0, I)$ . The $\sqrt{K}$ factor indicates that the observed signal strength increases with K. For a fixed FDR control value of 0.01, Figure3c,d shows the results obtained for K≈N/5 and K≈N/2, respectively. Figure3e shows how the probability of error decays as a function of the number of measurements K. The results presented here are obtained by averaging over 1,000 different noise and sensing matrix realizations. From these results, we can see that the number of detected anomalies increases with K and the number of misclassifications decrease with K.

Conclusion

This work presents computationally efficient approaches for detecting known targets and anomalies of different strengths from projection measurements without performing a complete reconstruction of the underlying signals, and offers theoretical bounds on the worst-case target detector performance. This article treats each signal as independent of its spatial or temporal neighbors. This assumption is reasonable in many contexts, especially when the spatial or temporal resolution is low relative to the spatial homogeneity of the environment or the pace with which a scene changes. However, emerging technologies in computational optical systems continue to improve the resolution of spectral imagers. In our future work we will build upon the methods that we have discussed here to exploit the spatial or temporal correlations in the data.

Appendix 1: Proof of Theorem 1

Using linear algebra and matrix theory, it is possible to show that if B=I−A Σ_bA^T is positive definite, then

Φ = σ B^{- 1 / 2} A

(24)

satisfies (3).^c In particular, we can substitute (24) in (3) to verify that the proposed construction of Φ satisfies (3). Observe that C_Φ=(Φ Σ_bΦ^T + σ²I)^−1/2 can be written in terms of (24) as follows:

\begin{align} C_{Φ} & = {([σ B^{- \frac{1}{2}} A] Σ_{b} {[σ B^{- \frac{1}{2}} A]}^{T} + σ^{2} I)}^{- \frac{1}{2}} \\ = {(σ^{2} B^{- 1 / 2} (A Σ_{b} A^{T}) {(B^{- \frac{1}{2}})}^{T} + σ^{2} I)}^{- \frac{1}{2}} \\ = {(σ^{2} B^{- \frac{1}{2}} (I - B) {(B^{- \frac{1}{2}})}^{T} + σ^{2} I)}^{- \frac{1}{2}} \\ = {(σ^{2} B^{- 1})}^{- \frac{1}{2}} = σ^{- 1} B^{\frac{1}{2}} \end{align}

(25)

where the third-to-last equation follows from the definition of B and (25) follows from the fact that B is symmetric and positive definite. If B is positive definite, then B⁻¹ is positive definite as well and can be decomposed as B⁻¹=(B^−1/2)^TB^−1/2, where the matrix square root B^−1/2is symmetric and positive definite. By substituting (25) and (24) in (3), we have C_ΦΦ=σ⁻¹B^1/2σ B^−1/2A=A. A sufficient condition for B to be positive definite can be derived as follows.

To ensure positive definiteness of B, we must have

x^{T} B x = x^{T} x - x^{T} (A Σ_{b} A^{T}) x > 0

(26)

for any nonzero $x \in R^{K}$ . Note that since Σ_b is positive semidefinite, x^T(A Σ_bA^T)x≥0. However, the right hand side of (26) is > 0 only if the spectral norm of A Σ_bA^Tis < 1, since $x^{T} (A Σ_{b} A^{T}) x \leq ∥ x ∥^{2} \cdot ∥ A Σ_{b} A^{T} ∥$ . The norm of A Σ_bA^T is in turn bounded above by

∥ A Σ_{b} A^{T} ∥ \leq ∥ A ∥ ∥ Σ_{b} ∥ ∥ A^{T} ∥ = ∥ A ∥^{2} ∥ Σ_{b} ∥ = ∥ A ∥^{2} λ_{\max}

since ∥A∥=∥A^T∥ and ∥Σ_b∥=λ_max, where λ_max is the largest eigenvalue of Σ_b. To ensure ∥A Σ_bA^T∥ < 1, ∥A∥²λ_max has to be < 1, which leads to the result of Theorem 1.

Appendix 2: Proof of Theorem 2

The proof of Theorem 2 adapts the proof techniques from[48] to nonidentical independent hypothesis tests. We begin by expanding the pFDR definition in (8) as follows:

\begin{align} {pFDR}^{(j)} (Γ) = \sum_{k = 1}^{M} E [\frac{V (Γ)}{R (Γ)}| R (Γ) = k] \\ \times P (R (Γ) = k| R (Γ) > 0) . \end{align}

Observe that R(Γ)=k implies that there exists some subset S_k={u₁,…,u_k}⊆{1,…,M} of size k such that $y_{u_{ℓ}} \in Γ_{u_{ℓ}}^{(j)}$ for ℓ=1,…,k and $y_{i} \notin Γ_{i}^{(j)}$ for all i∉S_k. To simplify the notation, let $Λ_{S_{k}} = \prod_{u \in S_{k}} Γ_{u}^{j} \times \prod_{ℓ \notin S_{k}} {\tilde{Γ}}_{ℓ}^{(j)}$ , where ${\tilde{Γ}}_{ℓ}^{(j)}$ is the complement of $Γ_{ℓ}^{(j)}$ , denote the significance region that corresponds to set S_k, and T=(y₁,…,y_M) be a set of test statistics corresponding to each hypothesis test. Considering all such subsets we have

\begin{align} {pFDR}^{(j)} (Γ) = \sum_{k = 1}^{M} \sum_{S_{k}} E [\frac{V (Γ)}{k}| T \in Λ_{S_{k}}] \\ \times P (T \in Λ_{S_{k}}| R (Γ) > 0) . \end{align}

(27)

By plugging in the definition of V({Γ_i}) from (9), we have

\begin{align} E & [V (Γ)| T \in Λ_{S_{k}}] = E [\sum_{i = 1}^{M} I_{\{y_{i} \in Γ_{i}^{(j)}\}} I_{\{ℋ_{i}^{(j)} = 0\}}| T \in Λ_{S_{k}}] \\ \equiv \sum_{ℓ = 1}^{k} E [I_{\{ℋ_{u_{ℓ}}^{(j)} = 0\}}| y_{u_{ℓ}}] = \sum_{ℓ = 1}^{k} P (ℋ_{u_{ℓ}}^{(j)} = 0| y_{u_{ℓ}} \in Γ_{u_{ℓ}}^{(j)}) \end{align}

(28)

for all u_ℓ∈S_k since the tests are independent of each other given A. The posterior probability $P (ℋ_{i}^{(j)} = 0| y_{i} \in Γ_{i}^{(j)})$ for the i^thhypothesis test can be expanded using Bayes’ rule as

\begin{align} P (ℋ_{0 i}^{(j)} |y_{i} \in Γ_{i}^{(j)}) = \frac{P (y_{i} \in Γ_{i}^{(j)} |ℋ_{0 i}) P (ℋ_{0 i}^{(j)})}{P (y_{i}^{(j)} \in Γ_{i}^{(j)})} \\ \equiv \frac{P ({\hat{f}}_{i} \neq f^{(j)}| f_{i}^{*} = f^{(j)}) P (f_{i}^{*} = f^{(j)})}{P ({\hat{f}}_{i} \neq f^{(j)})}, \end{align}

(29)

where ${\hat{f}}_{i} = {arg max}_{f^{(ℓ)} \in D} P (f_{i}^{*} = f^{(ℓ)}| y_{i}, α_{i}, A)$ . To upper bound the numerator of (29), consider the probability of misclassification given by ${(P_{e})}_{i} = P ({\hat{f}}_{i} \neq f_{i}^{*})$ where $f_{i}^{*} = f^{(j)} \in D$ , which can be expanded as follows:

\begin{align} {(P_{e})}_{i} = P ({\hat{f}}_{i} \neq f_{i}^{*}) \\ = \sum_{ℓ = 1}^{m} P ({\hat{f}}_{i} \neq f_{i}^{*}| f_{i}^{*} = f^{(ℓ)}) P (f_{i}^{*} = f^{(ℓ)}) \\ \equiv \sum_{ℓ = 1}^{m} P ({\hat{f}}_{i} \neq f^{(ℓ)}| f_{i}^{*} = f^{(ℓ)}) P (f_{i}^{*} = f^{(ℓ)}) \\ \geq P ({\hat{f}}_{i} \neq f^{(j)}| f_{i}^{*} = f^{(j)}) P (f_{i}^{*} = f^{(j)}) . \end{align}

(30)

The denominator term in (29) can be expanded as follows:

\begin{align} P ({\hat{f}}_{i} \neq f^{(j)}) = P ({\hat{f}}_{i} \neq f^{(j)}| f_{i}^{*} = f^{(j)}) P (f_{i}^{*} = f^{(j)}) \\ + P ({\hat{f}}_{i} \neq f^{(j)}| f_{i}^{*} \neq f^{(j)}) P (f_{i}^{*} \neq f^{(j)}) . \end{align}

Observe that $P ({\hat{f}}_{i} \neq f^{(j)}| f_{i}^{*} = f^{(j)})$ is nonnegative, and

\begin{align} P ({\hat{f}}_{i} \neq f^{(j)}| f_{i}^{*} \neq f^{(j)}) & = P ({\hat{f}}_{i} \in D ∖ f^{(j)}| f_{i}^{*} \neq f^{(j)}) \\ \geq P ({\hat{f}}_{i} = f_{i}^{*}| f_{i}^{*} \neq f^{(j)}) \\ = 1 - P ({\hat{f}}_{i} \neq f_{i}^{*}| f_{i}^{*} \neq f^{(j)}) \\ = 1 - \frac{P ({\hat{f}}_{i} \neq f_{i}^{*}, f_{i}^{*} \neq f^{(j)})}{P (f_{i}^{*} \neq f^{(j)})} \\ \geq 1 - \frac{P ({\hat{f}}_{i} \neq f_{i}^{*})}{P (f_{i}^{*} \neq f^{(j)})} \\ = 1 - \frac{{(P_{e})}_{i}}{1 - p^{(j)}} . \end{align}

Thus

P ({\hat{f}}_{i} \neq f^{(j)}) \geq (1 - \frac{{(P_{e})}_{i}}{1 - p^{(j)}}) (1 - p^{(j)}) = 1 - p^{(j)} - {(P_{e})}_{i} .

(31)

Substituting (30) and (31) in (29),

\begin{align} P (ℋ_{0 i}^{(j)} |y_{i} \in Γ_{i}^{(j)}) \leq \frac{{(P_{e})}_{i}}{1 - p^{(j)} - {(P_{e})}_{i}} \\ \leq \frac{{(P_{e})}_{\max}}{1 - p^{(j)} - {(P_{e})}_{\max}} . \end{align}

(32)

By substituting (32) in (27) and (28) we have:

\begin{align} {pFDR}^{(j)} (Γ) \leq \sum_{k = 1}^{M} \sum_{S_{k}} \frac{1}{k} (\sum_{ℓ = 1}^{k} \frac{{(P_{e})}_{\max}}{1 - p^{(j)} - {(P_{e})}_{\max}}) \\ \times P (T \in Λ_{S_{k}}| R (Γ) > 0) \\ = \frac{{(P_{e})}_{\max}}{1 - p^{(j)} - {(P_{e})}_{\max}} \\ \times \sum_{k = 1}^{M} \sum_{S_{k}} P (T \in Λ_{S_{k}}| R (Γ) > 0) \\ \leq \frac{{(P_{e})}_{\max}}{1 - p^{(j)} - {(P_{e})}_{\max}} \end{align}

since $\sum_{k = 1}^{M} \sum_{S_{k}} P (T \in Λ_{S_{k}}| R (Γ) > 0) \leq 1$ . The result of Theorem 2 is obtained by finding an upper bound on the worst-case pFDR given by

\begin{align} {pFDR}_{max} = max_{j \in {1, \dots, m}} {pFDR}^{(j)} (Γ) \\ \leq max_{j \in {1, \dots, m}} \frac{{(P_{e})}_{\max}}{1 - p^{(j)} - {(P_{e})}_{\max}} \\ = \frac{{(P_{e})}_{max}}{1 - p_{max} - {(P_{e})}_{max}} \end{align}

where p_max=max_{ℓ∈{1,…,m}}p^(ℓ).

Appendix 3: Proof of Theorem 3

The proof is via a random selection technique, similar to random coding arguments common in information theory. Specifically, we will draw a K×N sensing matrix A at random from a particular distribution and then show that, for ε, N, and K satisfying the conditions of the theorem, the probability that the conclusions of the theorem will fail to hold for this randomly chosen A is strictly smaller than unity. This will imply that the conclusions of the theorem must be true for at least one (deterministic) realization of A.

We begin by specifying all the relevant random variables:

$f_{1}^{*}, \dots, f_{M}^{*}$ are i.i.d. random variables taking values in the dictionary $D = {f^{(1)}, \dots, f^{(m)}}$ with probabilities $p^{(j)} = Pr {f_{i}^{*} = f^{(j)}}, j \in {1, \dots, m}$ ;
$n_{1}, \dots, n_{M} \overset{i.i.d.}{\sim} N (0, I)$ ;
G is a random K×N matrix with i.i.d. $N (0, 1)$ entries.

We assume that ${f_{i}^{*}}_{i = 1}^{M}$ , ${n_{i}}_{i = 1}^{M}$ , and G are mutually independent, and we will denote by P their joint probability distribution. Finally, we let $A = \frac{1}{\sqrt{K}} G$ and consider the observation model

y_{i} = α_{i} A f_{i}^{*} + n_{i}, i \in {1, \dots, M}

(33)

where α₁,…,α_M > 0 are the given signal strengths.

We first consider the case when α₁=⋯=α_M=α. Given ε, N, and K, we define the following two error events:

\begin{align} E_{1} ≜ \{∥ G ∥ \geq (1 + ε) (\sqrt{K} + \sqrt{N})\}, and \\ E_{2} ≜ \{{\hat{f}}_{1} \neq f_{1}^{*}\}, \end{align}

where, for each i∈{1,…,M}, ${\hat{f}}_{i}$ is defined according to (12). Note that, since we have assumed that the α_i’s are equal and all the pairs $(f_{i}^{*}, n_{i}), i \in {1, \dots, M}$ , are i.i.d.,

P ({\hat{f}}_{i} \neq f_{i}^{*} | A) = P (E_{2} | A), \forall i \in {1, \dots, M} .

(34)

We will now prove that

\begin{align} P (E_{1} \cup E_{2}) \leq \frac{1 - p_{min}}{p_{min}} {(1 + \frac{α^{2} d_{min}^{2}}{4 K σ^{2}})}^{- \frac{K}{2}} \\ + 2 exp (- \frac{(K + N) ε^{2}}{2}) . \end{align}

(35)

The union bound gives $P (E_{1} \cup E_{2}) \leq P (E_{1}) + P (E_{2})$ . First, we bound $P (E_{1})$ . To do that, we use the following concentration result for Gaussian random matrices[58]: for any t≥0,

\begin{align} Pr \{∥ G ∥ \geq \sqrt{K} + \sqrt{N} + t\} \leq 2 e^{- t^{2} / 2} . \end{align}

Letting $t = ε (\sqrt{K} + \sqrt{N})$ and using the fact that t²≥(K + N)ε², we get

P (E_{1}) \leq 2 exp (- \frac{(K + N) ε^{2}}{2}) .

(36)

Next, we bound $P (E_{2})$ . To that end, we use the following result, which is a straightforward extension of ([5], Theorem 1) to nonequiprobable dictionary elements:

Lemma 1 (Compressive classification error)

Consider the problem of classifying a signal of interest $f^{*} \in D = {f^{(1)}, \dots, f^{(m)}}$ to one of m known target classes by making observations of the form y=α A f^∗ + n where $n \sim N (0, σ^{2} I)$ , given the knowledge of the dictionary $D$ , prior probabilities p^(j)for j∈{1,⋯,m}, sensing matrix A, and the noise variance σ². If the entries of A are drawn i.i.d. from $N (0, 1 / K)$ independently of f^∗ and n, and the estimate $\hat{f}$ is obtained according to (12), then

\begin{align} P (\hat{f} \neq f^{*}) \leq \frac{1 - p_{\min}}{p_{\min}} {(1 + \frac{α^{2} d_{min}^{2}}{4 K σ^{2}})}^{- \frac{K}{2}} \end{align}

where the probability is taken with respect to the distributions underlying f^∗, A, and n.

Using the above lemma, we have

P (E_{2}) \leq \frac{1 - p_{min}}{p_{min}} {(1 + \frac{α^{2} d_{min}^{2}}{4 K})}^{- \frac{K}{2}} .

(37)

Combining (36) and (37), we get (35).

Because of (13a), the right-hand side of (35) is less than 1−ε−p_max, which is strictly positive by hypothesis. Thus, from the fact that

\begin{align} P (E_{1} \cup E_{2}) = E [P (E_{1} \cup E_{2} | A)] \end{align}

and from (34), it follows that there exists at least one deterministic choice of the K×N sensing matrix A^∗, such that:

\begin{align} ∥ A^{*} ∥ \leq (1 + ε) (1 + \sqrt{\frac{N}{K}}) \end{align}

(38a)

\begin{align} {(P_{e})}_{max} (A^{*}) \leq \frac{1 - p_{min}}{p_{min}} {(1 + \frac{α^{2} d_{min}^{2}}{4 K})}^{- \frac{K}{2}} \\ + 2 exp (- \frac{(K + N) ε^{2}}{2}) \end{align}

(38b)

where, for a given choice of A, (P_e)_max(A) denotes the maximum probability of error defined in Theorem 2.

Next, from (38a) and (13b) it follows that A^∗satisfies the conditions of Theorem 1. Finally, we use (11) to bound the worst-case pFDR achievable with A^∗. First of all, we note that the function $U (x) = \frac{x}{1 - p_{max} - x}$ is twice differentiable and convex on the interval [0,1−p_max]. Therefore, for any x∈[0,1−p_max] and any h > 0 small enough so that x + h∈[0,1−p_max], we have

\begin{align} U (x + h) \leq U (x) + U^{'} (x + h) h = U (x) \\ + \frac{(1 - p_{max}) h}{{(1 - p_{max} - x - h)}^{2}} . \end{align}

(39)

Let us choose

\begin{align} x = \frac{1 - p_{min}}{p_{min}} {(1 + \frac{α^{2} d_{min}^{2}}{4 K})}^{- \frac{K}{2}} and \\ h = 2 exp (- \frac{(K + N) ε^{2}}{2}) . \end{align}

Then from (13a) we have x + h≤1−ε−p_max < 1−p_max, and from (13c) we have x + h≥0. Hence, using (39) and simplifying, we obtain the bound

\begin{align} {pFDR}_{max} (A^{*}) \leq \frac{1}{p_{min}} {(\frac{1 - p_{max}}{1 - p_{min}} {(1 + \frac{α^{2} d_{min}^{2}}{4 K})}^{\frac{K}{2}} - \frac{1}{p_{min}})}^{- 1} \\ + \frac{2 (1 - p_{max})}{ε^{2}} exp (- \frac{(K + N) ε^{2}}{2}) . \end{align}

This proves the theorem for the case α₁=⋯=α_M=α.

To handle the case when the α_i’s are distinct, we simply let

i^{*} ≜ {arg min}_{i \in {1, \dots, M}} α_{i}

and replace the definition of the error event $E_{2}$ with $E_{2}^{'} = {{\hat{f}}_{i^{*}} \neq f_{i^{*}}^{*}}$ . Then the same argument goes through, except that instead of (34) we use the bound

\begin{align} P ({\hat{f}}_{i} \neq f_{i}^{*} | A) \leq P ({\hat{f}}_{i^{*}} \neq f_{i^{*}}^{*} | A) = P (E_{2}^{'} | A), \forall i \neq i^{*} \end{align}

which follows from the following argument. First of all, we can replace the observation model with the equivalent model

\begin{align} {\tilde{y}}_{i} = A f_{i}^{*} + {\tilde{n}}_{i}, i \in {1, \dots, M} \end{align}

where ${\tilde{n}}_{i} = \frac{1}{α_{i}} n_{i} \sim N (0, \frac{1}{α_{i}^{2}} I)$ . Secondly, from the fact that α_i≥α_i∗≡α_minfor any i≠i^∗ it follows that ${\tilde{n}}_{i^{*}}$ is equal in distribution to ${\tilde{n}}_{i} + {\tilde{n}}_{i}^{'}$ , where ${\tilde{n}}_{i}^{'} \sim N (0, (\frac{1}{α_{i}^{2}} - \frac{1}{α_{\min}^{2}}) I)$ is independent of ${\tilde{n}}_{i}$ . This implies that the i^∗th observation is the noisiest, and the corresponding MAP estimate ${\hat{f}}_{i^{*}}$ has the largest probability of error.

Appendix 4: Proof of Theorem 4

We first prove this theorem assuming that {α_i} are known and later extend to the case where ${{\hat{α}}_{i}}$ are estimated from the observations. Let ${\tilde{f}}_{i} = {arg min}_{f \in D} ∥ f_{i}^{*} - f ∥$ . The p-value expression in (17) can be expanded as follows:

\begin{align} p_{i} = P ({\tilde{d}}_{i} \geq d_{i}| ℋ_{0 i}) \\ = P (min_{f \in D} ∥ α_{i} A (f_{i}^{*} - f) + n ∥ \geq d_{i}| ℋ_{0 i}) \\ \leq P (∥ α_{i} A (f_{i}^{*} - {\tilde{f}}_{i}) + n ∥ \geq d_{i}| ℋ_{0 i}) \\ = P (∥ α_{i} A (f_{i}^{*} - {\tilde{f}}_{i}) + n ∥^{2} \geq d_{i}^{2}| ℋ_{0 i}) . \end{align}

(40)

Note that $∥ α_{i} A (f_{i}^{*} - {\tilde{f}}_{i}) + n ∥^{2}$ is a noncentral χ²random variable with K degrees of freedom and a noncentrality parameter $ν_{i} = ∥ α_{i} A (f_{i}^{*} - {\tilde{f}}_{i}) ∥^{2}$ . Thus (40) can be written in terms of a noncentral χ²CDF $F (d_{i}^{2}; K, ν_{i})$ with parameter $d_{i}^{2}$ . The upper and lower bounds on ν_ican be obtained using the properties of the projection matrix A. Applying (18), we see that

\begin{align} α_{i}^{2} {(1 - ε)}^{2} ∥ f_{i}^{*} - {\tilde{f}}_{i} ∥^{2} \leq ν_{i} \leq α_{i}^{2} {(1 + ε)}^{2} ∥ f_{i}^{*} - {\tilde{f}}_{i} ∥^{2} \end{align}

with high probability. Thus,

\begin{align} p_{i} \leq 1 - P (∥ α_{i} A (f_{i}^{*} - {\tilde{f}}_{i}) + n ∥^{2} \leq d_{i}^{2}| ℋ_{0 i}) \\ = 1 - F (d_{i}^{2}; K, ν_{i}) \\ \leq 1 - F (d_{i}^{2}; K, α_{i}^{2} {(1 + ε)}^{2} ∥ f_{i}^{*} - {\tilde{f}}_{i} ∥^{2}) \\ \leq 1 - F (d_{i}^{2}; K, α_{i}^{2} {(1 + ε)}^{2} τ^{2}) \end{align}

(41)

since $∥ f_{i}^{*} - f ∥ \leq τ$ for all $f \in D$ under $ℋ_{0 i}$ .

When {α_i} are estimated from the observations such that ${{\hat{α}}_{i}}$ satisfy (19), we can write the p-value expression in (41) as follows:

\begin{align} p_{i} \leq 1 - F (d_{i}^{2}; K, {∥A (α_{i} f_{i}^{*} - {\hat{α}}_{i} {\tilde{f}}_{i})∥}^{2}) \\ \leq 1 - F (d_{i}^{2}; K, {(1 + ε)}^{2} {\hat{α}}_{i}^{2} {∥\frac{α_{i}}{{\hat{α}}_{i}} f_{i}^{*} - {\tilde{f}}_{i}∥}^{2}) \end{align}

(42)

where (42) is due to the distance preservation property of A given in (18). Observe that ${∥\frac{α_{i}}{{\hat{α}}_{i}} f_{i}^{*} - {\tilde{f}}_{i}∥}^{2}$ can be upper bounded as shown below:

\begin{align} {∥\frac{α_{i}}{{\hat{α}}_{i}} f_{i}^{*} - {\tilde{f}}_{i}∥}^{2} & = {∥(\frac{α_{i}}{{\hat{α}}_{i}} - 1) f_{i}^{*} + f_{i}^{*} - {\tilde{f}}_{i}∥}^{2} \\ \leq {(∥(\frac{α_{i}}{{\hat{α}}_{i}} - 1) f_{i}^{*}∥ + ∥ f_{i}^{*} - {\tilde{f}}_{i} ∥)}^{2} \\ = {(|\frac{α_{i}}{{\hat{α}}_{i}} - 1| + ∥ f_{i}^{*} - {\tilde{f}}_{i} ∥)}^{2} \\ \leq {(ζ + ∥ f_{i}^{*} - {\tilde{f}}_{i} ∥)}^{2} \end{align}

where third-to-last equation is due to the triangle inequality, second-to-last equation comes from the assumption that $∥f_{i}^{*}∥ = 1$ , and the last inequality is due to (19). By applying this result to (42) and exploiting the fact that $∥ f_{i}^{*} - f ∥ \leq τ$ under $ℋ_{0 i}$ for some $f \in D$ , we have

\begin{align} p_{i} \leq 1 - F (d_{i}^{2}; K, {(1 + ε)}^{2} {\hat{α}}_{i}^{2} {(ζ + ∥ f_{i}^{*} - {\tilde{f}}_{i} ∥)}^{2}) \\ \leq 1 - F (d_{i}^{2}; K, {(1 + ε)}^{2} {\hat{α}}_{i}^{2} {(ζ + τ)}^{2}) . \end{align}

Endnotes

^a Note that τ cannot exceed $\sqrt{2}$ because we assume that all targets of interest, including those in $D$ and the actual target f^∗, are unit-norm.^b The anomaly detection problem discussed here is more accurately described as target detection in the classical detection theory vocabulary. However, in recent works[24, 25], the authors assume that the nominal distribution is obtained from training data and a test sample is declared to be anomalous if it falls outside of the nominal distribution learned form the training data. Our work is in a similar spirit where we learn our dictionary from training data and label any test spectrum that does not correspond to our dictionary as being anomalous.^c The authors would like to thank Prof. Roummel Marcia for fruitful discussions related to this point.

References

Candès EJ, Tao T: Near-optimal signal recovery from random projections: universal encoding strategies? IEEE Trans. Inf. Theory 2006, 52(12):5406-5425.
Article MathSciNet MATH Google Scholar
Donoho D: Compressed sensing. IEEE Trans. Info. Theory 2006, 52(4):1289-1306.
Article MathSciNet MATH Google Scholar
Davenport M, Duarte M, Wakin M, Laska J, Takhar D, Kelly K, Baraniuk R: The smashed filter for compressive classification and target recognition. Proceedings of SPIE, vol. 6498 (San Jose, CA, 2007), pp. 142–153
Google Scholar
Duarte MF, Davenport MA, Wakin MB, Baraniuk RG: Sparse signal detection from incoherent projections. IEEE International Conference on Acoustics, Speech and Signal Processing, vol. 3 (Toulouse, France, 2006), pp. 305–308
Google Scholar
Haupt J, Castro R, Nowak R, Fudge G, Yeh A: Compressive sampling for signal classification. Fortieth Asilomar Conference on Signals, Systems and Computers,. 2006, pp. 1430–1434
Google Scholar
Aeron S, Saligrama V, Zhao M: Information theoretic bounds for compressed sensing. Inf. IEEE Trans. Theory 2010, 56(10):5111-5130.
Article MathSciNet Google Scholar
Arias-Castro E, Eldar Y: Noise folding in compressed sensing. IEEE Signal Process. Lett 2011, 18: 478-481.
Article Google Scholar
Han J, Bhanu B: Fusion of color and infrared video for moving human detection. Pattern Recogn 2007, 40(6):1771-1784. 10.1016/j.patcog.2006.11.010
Article MATH Google Scholar
Johnson W, Wilson D, Fink W, Humayun M, Bearman G: Snapshot hyperspectral imaging in ophthalmology. J. Biomed. Optics 2007, 12(1):014036-1–014036-7. 10.1117/1.2434950
Article Google Scholar
Lin R, Dennis B, Benz A: The Reuven Ramaty High-Energy Solar Spectrscopic Imager (RHESSI) - Mission Description and Early Results. (Kluwer Academic Publishers, Dordrecht, 2003)
Book Google Scholar
Martin M, Newman S, Aber J, Congalton R: Determining forest species composition using high spectral resolution remote sensing data. Remote Sens. Envir 1998, 65(3):249-254. 10.1016/S0034-4257(98)00035-2
Article Google Scholar
Martin M, Wabuyele M, Chen K, Kasili P, Panjehpour M, Phan M, Overholt B, Cunningham G, Wilson D, DeNovo R, Vo-Dinh T: Development of an advanced hyperspectral imaging (HSI) system with applications for cancer detection. Ann. Biomed. Eng 2006, 34(6):1061-1068. 10.1007/s10439-006-9121-9
Article Google Scholar
Miller J, Elvidge C, Rock B, Freemantle J: An airborne perspective on vegetation phenology from the analysis of AVRIS data sets over the Jasper ridge biological preserve. Geoscience and Remote Sensing Symposium (IGARSS’90): Remote sensing for the nineties (College Park, MD, 20–24 May 1990), pp. 565–568
Google Scholar
Stellman C, Hazel G, Bucholtz F, Michalowicz J, Stocker A, Schaaf W: Real-time hyperspectral detection and cuing. Opt. Eng 2000, 39: 1928-1935. 10.1117/1.602577
Article Google Scholar
Zuzak K, Naik S, Alexandrakis G, Hawkins D, Behbehani K, Livingston E: Intraoperative bile duct visualization using near-infrared hyperspectral video imaging. Am. J. Surg 2008, 195(4):491-497. 10.1016/j.amjsurg.2007.05.044
Article Google Scholar
Brady D, Gehm M: Compressive imaging spectrometers using coded apertures. Proc. of SPIE, vol. 6246 (Kissimmee, Florida, 2006), pp. 62460A-1–62460A-9
Google Scholar
DeVerse RA, Coifman RR, Coppi AC, Fateley WG, Geshwind F, Hammaker RM, Valenti S, Warner FJ, Davis GL: Application of Spatial Light Modulators for New Modalities in Spectrometry and Imaging. Spectral Imaging: Instrumentation, Applications, and Analysis II, vol. 4959, ed. by RM Levenson, GH Bearman, A Mahadevan-Jansen (2003), pp. 12–22
Chapter Google Scholar
Gehm M, John R, Brady D, Willett R, Schulz T: Single-shot compressive spectral imaging with a dual-disperser architecture. Opt. Express 2007, 15(21):14013-14027. 10.1364/OE.15.014013
Article Google Scholar
Takhar D, Laska J, Wakin MB, Duarte MF, Baron D, Sarvotham S, Kelly K, Baraniuk RG: A new compressive imaging camera architecture using optical-domain compression. Proc. IS&T/SPIE Symposium on Electronic Imaging (San Jose, CA, 2006), pp. 43–52
Google Scholar
Wagadarikar A, John R, Willett R, Brady D: Single disperser design for coded aperture snapshot spectral imaging. Appl. Opt 2008, 47(10):B44-B51. 10.1364/AO.47.000B44
Article Google Scholar
Woolfe F, Maggioni M, Davis G, Warner F, Coifman R, Zucker S: Hyper-spectral microscopic discrimination between normal and cancerous colon biopsies. Manuscript (2006)
Google Scholar
Manolakis D, Marden D, Shaw G: Hyperspectral image processing for automatic target detection applications. Lincoln Laboratory J 2003, 14(1):79-116.
Google Scholar
Wei G, Agnihotri L, Dimitrova N: TV program classification based on face and text processing. 2000 IEEE International Conference on Multimedia and Expo, ICME 2000, vol. 3 (2000), pp. 1345–1348
Google Scholar
Hero AO: Geometric entropy minimization (GEM) for anomaly detection and localization. Proc. Advances in Neural Information Processing Systems (NIPS) (MIT Press, Vancouver, Canada, 2006), pp. 585–592
Google Scholar
Steinwart I, Hush D, Scovel C: A classification framework for anomaly detection. J. Mach. Learn. Res 2005, 6: 211-232.
MathSciNet MATH Google Scholar
Stein D, Beaven S, Hoff L, Winter E, Schaum A, Stocker A: Anomaly detection from hyperspectral imagery. IEEE Signal Process. Mag 2002, 19(1):58-69. 10.1109/79.974730
Article Google Scholar
Manolakis D, Shaw G: Detection algorithms for hyperspectral imaging applications. IEEE Signal Process. Mag 2002, 19(1):29-43. 10.1109/79.974724
Article Google Scholar
Berger JO: Statistical Decision Theory and Bayesian Analysis,. (Springer, New York, 1985)
Book MATH Google Scholar
Benjamini Y, Hochberg Y: Controlling the false discovery rate: a practical and powerful approach to multiple testing. J. R. Stat. Soc. Ser. B (Methodological) 1995, 57(1):289-300.
MathSciNet MATH Google Scholar
Jin X, Paswaters S, Cline H: A comparative study of target detection algorithms for hyperspectral imagery. Proceedings of SPIE, vol. 7334 (2009), p. 73341W
Google Scholar
Kelly E: An adaptive detection algorithm. IEEE Trans. Aerospace Electron. Syst. AES- 1986, 22(2):115-127.
Article Google Scholar
Kraut S, Scharf L, McWhorter L: Adaptive subspace detectors. IEEE Trans. Signal Processing 2001, 49(1):1-16. 10.1109/78.890324
Article Google Scholar
Scharf L, Friedlander B: Matched subspace detectors. IEEE Trans. Signal Process 1994, 42(8):2146-2157. 10.1109/78.301849
Article Google Scholar
Kwon H, Nasrabadi N: Kernel matched subspace detectors for hyperspectral target detection. IEEE Trans. Pattern Anal. Mach. Intell 2006, 28(2):178-194.
Article Google Scholar
Scharf LL, McWhorter LT: Adaptive matched subspace detectors and adaptive coherence estimators. Conference Record of the Thirtieth Asilomar Conference on Signals, Systems and Computers (Pacific Grove, CA, 1996), pp. 1114–1117
Google Scholar
Parmar M, Lansel S, Wandell B: Spatio-spectral reconstruction of the multispectral datacube using sparse recovery. 15th IEEE International Conference on Image Processing (San Diego, CA, 2008), pp. 473–476
Google Scholar
Willett R, Gehm M, Brady D: Multiscale reconstruction for computational spectral imaging. Comput. Imag. V 2007, 6498: 64980L-1–64980L-15.
Google Scholar
Fowler J, Du Q: Anomaly detection and reconstruction from random projections. IEEE Trans. Image Process 2012, 21(1):184-195.
Article MathSciNet Google Scholar
Reed I, Yu X: Adaptive multiple-band CFAR detection of an optical pattern with unknown spectral distribution. IEEE Trans. Acoust. Speech Signal Process 1990, 38(10):1760-1770. 10.1109/29.60107
Article Google Scholar
Krishnamurthy K, Raginsky M, Willett R: Hyperspectral target detection from incoherent projections. IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP) (Dallas, TX, 2010), pp. 3550–3553
Google Scholar
Krishnamurthy K, Raginsky M, Willett R: Hyperspectral target detection from incoherent projections: nonequiprobable targets and inhomogenous SNR. 17th IEEE International Conference on Image Processing (ICIP) (Hongkong, 2010), pp. 1357–1360
Google Scholar
Boardman JW: Spectral Angle Mapping: A Rapid Measure of Spectral Similarity. (AVIRIS, 1993)
Google Scholar
Guo Z, Osher S: Template matching via L1 minimization and its application to hyperspectral data. Accepted to Inverse Problems and Imaging (IPI), 2009
MATH Google Scholar
Kwon H, Nasrabadi N: Kernel RX-algorithm: a nonlinear anomaly detector for hyperspectral imagery. IEEE Trans. Geosci. Remote Sens 2005, 43(2):388-397.
Article Google Scholar
Szlam A, Guo Z, Osher S: A split Bregman method for non-negative sparsity penalized least squares with applications to hyperspectral demixing. IEEE 17th International Conference on Image Processing (ICIP) (Hongkong, 2010), pp. 1917–1920
Google Scholar
Chang C: Virtual dimensionality for hyperspectral imagery. SPIE Newsroom 2009, 10(2.1200909):1749.
Google Scholar
Chang C, Du Q: Estimation of number of spectrally distinct signal sources in hyperspectral imagery. IEEE Trans. Geosci. Remote Sens 2004, 42(3):608-619. 10.1109/TGRS.2003.819189
Article Google Scholar
Storey J: The positive false discovery rate: a Bayesian interpretation and the q-value. Ann. Stat 2003, 2013-2035.
Google Scholar
Johnson W, Lindenstrauss J: Extensions of Lipschitz maps into a Hilbert space. Contemp. Math 1984, 26: 189-206.
Article MathSciNet MATH Google Scholar
Healey G, Slater D: Models and methods for automated material identification in hyperspectral imagery acquired under unknown illumination and atmospheric conditions. IEEE Trans. Geosci. Remote Sens 1999, 37(6):2706-2717. 10.1109/36.803418
Article Google Scholar
Achlioptas D: Database-friendly random projections. Proc. 20th ACM Symp. Principles of Database Systems (ACM Press, New York, NY 2001), pp. 274–281
Google Scholar
Baraniuk R, Davenport M, DeVore R, Wakin M: A simple proof of the restricted isometry property for random matrices. Constructive Approx 2008, 28(3):253-263. 10.1007/s00365-007-9003-x
Article MathSciNet MATH Google Scholar
Krahmer F, Ward R: New and improved johnson-lindenstrauss embeddings via the restricted isometry property. SIAM Journal on Mathematical Analysis 2011, 43(3):1269-1281. Arxiv preprint arXiv:1009.0744, 2010 10.1137/100810447
Article MathSciNet MATH Google Scholar
Wasserman L: All of Statistics: A Concise Course in Statistical Inference. (Springer, New York, NY 2004)
Book MATH Google Scholar
Tao T, Vu V: On random±1 matrices: singularity and determinant. Random Struct. Algor 2006, 28(1):1-23. 10.1002/rsa.20109
Article MathSciNet MATH Google Scholar
Tao T: Talagrand’s concentration inequality,. . Accessed on 08/03/2012 http://terrytao.wordpress.com/2009/06/09/talagrands-concentration-inequality/
Kruse FA, Boardman JW, Lefkoff AB, Young JM, Kierein-Young KS, Cocks TD, Jensen R, Cocks PA: HyMap: an Australian hyperspectral sensor solving global problems-results from USA HyMap data acquisitions. Proc. of the 10th Australasian Remote Sensing and Photogrammetry Conference (Adelaide, Australia, 2000), pp. 18–23
Google Scholar
Davidson KR, Szarek SJ: Local operator theory, random matrices and Banach spaces. (North-Holland, Amsterdam, 2001), pp. 317–366
MATH Google Scholar

Download references

Acknowledgements

This work was supported by the NSF Award No. DMS-08-11062, DARPA Grant No. HR0011-09-1-0036, and AFRL Grant No. FA8650-07-D-1221.

Author information

Authors and Affiliations

Department of Electrical and Computer Engineering, Duke University, Durham, NC, 27708, USA
Kalyani Krishnamurthy & Rebecca Willett
Department of Electrical and Computer Engineering and Coordinated Science Laboratory, University of Illinois at Urbana-Champaign, Urbana, IL, 61801, USA
Maxim Raginsky

Authors

Kalyani Krishnamurthy
View author publications
You can also search for this author in PubMed Google Scholar
Rebecca Willett
View author publications
You can also search for this author in PubMed Google Scholar
Maxim Raginsky
View author publications
You can also search for this author in PubMed Google Scholar

Corresponding author

Correspondence to Kalyani Krishnamurthy.

Additional information

Competing interest

The authors declare that they have no competing interests.

Authors’ original submitted files for images

Below are the links to the authors’ original submitted files for images.

Authors’ original file for figure 1

Authors’ original file for figure 2

Authors’ original file for figure 3

Authors’ original file for figure 4

Authors’ original file for figure 5

Authors’ original file for figure 6

Authors’ original file for figure 7

Authors’ original file for figure 8

Authors’ original file for figure 9

Authors’ original file for figure 10

Authors’ original file for figure 11

Authors’ original file for figure 12

Authors’ original file for figure 13

Authors’ original file for figure 14

Authors’ original file for figure 15

Rights and permissions

Open Access This article is distributed under the terms of the Creative Commons Attribution 2.0 International License (https://creativecommons.org/licenses/by/2.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.

Reprints and permissions

About this article

Cite this article

Krishnamurthy, K., Willett, R. & Raginsky, M. Target detection performance bounds in compressive imaging. EURASIP J. Adv. Signal Process. 2012, 205 (2012). https://doi.org/10.1186/1687-6180-2012-205

Download citation

Received: 02 December 2011
Accepted: 28 August 2012
Published: 25 September 2012
DOI: https://doi.org/10.1186/1687-6180-2012-205

Target detection performance bounds in compressive imaging

Abstract

Introduction

Problem formulation

Performance metric

Previous investigations

Contributions

Whitening compressive observations

Theorem 1

Dictionary signal detection

Decision rule

Theorem 2

An achievable bound on the worst-case pFDR

Theorem 3

Extension to a manifold-based target detection framework

Anomalous signal detection

ASD problem formulation

Anomaly detection approach

Theorem 4

Experimental results

Dictionary signal detection

Anomaly detection

Experiments on simulated data

Experiments on real AVIRIS data

Conclusion

Appendix 1: Proof of Theorem 1

Appendix 2: Proof of Theorem 2

Appendix 3: Proof of Theorem 3

Lemma 1 (Compressive classification error)

Appendix 4: Proof of Theorem 4

Endnotes

References

Acknowledgements

Author information

Authors and Affiliations

Corresponding author

Additional information

Competing interest

Authors’ original submitted files for images

Rights and permissions

About this article

Cite this article

Share this article

Keywords