Sensor array beamformer post-processor

Description

BACKGROUND

Using multiple sensors arranged in an array, for example microphones arranged in a microphone array, to improve the quality of a captured signal, such as an audio signal, is a common practice. Various processing is typically performed to improve the signal captured by the array. For example, beamforming is one way that the captured signal can be improved.

Beamforming operations are applicable to processing the signals of a number of arrays, including microphone arrays, sonar arrays, directional radio antenna arrays, radar arrays, and so forth. In general, a beamformer is basically a spatial filter that operates on the output of an array of sensors, such as microphones, in order to enhance the amplitude of a coherent wave front relative to background noise and directional interference. In the case of a microphone array, beamforming involves processing output audio signals of the microphones of the array in such a way as to make the microphone array act as a highly directional microphone. In other words, beamforming provides a “listening beam” which points to, and receives, a particular sound source while attenuating other sounds and noise, including, for example, reflections, reverberations, interference, and sounds or noise coming from other directions or points outside the primary beam. Beamforming operations make the microphone array listen to given look-up direction, or angular space range. Pointing of such beams to various directions is typically referred to as beamsteering. A typical beamformer employs a set of beams that cover a desired angular space range in order to better capture the target or desired signal. There are, however, limitations to the improvement possible in processing a signal by employing beamforming.

Under real life conditions high reverberation leads to spatial spreading of the sound, even of point sources. For example, in many cases point noise sources are not stationary and have the dynamics of the source speech signal or are speech signals themselves, i.e. interference sources. Conventional time invariant beamformers are usually optimized under the assumption of isotropic ambient noise. Adaptive beamformers, on the other hand, work best under low reverberation conditions and a point noise source. In both cases, however, the improvements possible in noise suppression and signal selection capabilities of these algorithms are nearly exhausted with already existing algorithms.

Therefore, the SNR of the output signal generated by conventional beamformer systems is often further enhanced using post-processing or post-filtering techniques. In general, such techniques operate by applying additional post-filtering algorithms for sensor array outputs to enhance beamformer output signals. For example, microphone array processing algorithms generally use a beamformer to jointly process the signals from all microphones to create a single-channel output signal with increased directivity and thus higher SNR compared to a single microphone. This output signal is then often further enhanced by the use of a single channel post-filter for processing the beamformer output in such a way that the SNR of the output signal is significantly improved relative to the SNR produced by use of the beamformer alone.

Unfortunately, one problem with conventional beamformer post-filtering techniques is that they generally operate on the assumption that any noise present in the signal is either incoherent or diffuse. As such, these conventional post-filtering techniques generally fail to make allowances for point noise sources which may be strongly correlated across the sensor array. Consequently, the SNR of the output signal is not generally improved relative to highly correlated point noise sources.

SUMMARY

This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.

In general, the present beamforming post-processor technique is a novel technique for post-processing a sensor array's (e.g., a microphone array's) beamformer output to achieve better spatial filtering under conditions of noise and reverberation. For each frame (e.g., audio frame) and frequency bin the technique estimates the spatial probability for sound source presence (the probability that the desired sound source is in a particular look-up direction or angular space). It uses the spatial probability for the sound source presence and multiplies it by the beamformer output for each frequency bin to select the desired signal and to suppress undesired signals (i.e. not coming from the likely sound source direction or sector).

The technique uses so called instantaneous direction of arrival space (IDOA) to estimate the probability of the desired or target signal arriving from a given location. In general, for a microphone array, the phase differences at a particular frequency bin between the signals received at a pair of microphones give an indication of the instantaneous direction of arrival (IDOA) of a given sound source. IDOA vectors provide an indication of the direction from which a signal and/or point noise source originates. Non-correlated noise will be evenly spread in this space, while the signal and ambient noise (correlated components) will lie inside a hyper-volume that represents all potential positions of a sound source within the signal field.

In one embodiment the present beamforming post-processor technique is implemented as a real-time post-processor after a time-invariant beamformer. The present technique substantially improves the directivity of the microphone array. It is CPU efficient and adapts quickly when the listening direction changes, even in the presence of ambient and point noise sources. One exemplary embodiment of the present technique improves the performance of a traditional time invariant beamformer 3-9 dB.

It is noted that while the foregoing limitations in existing sensor array beamforming and noise suppression schemes described in the Background section can be resolved by a particular implementation of the present beamforming post-processor technique, this is in no way limited to implementations that just solve any or all of the noted disadvantages. Rather, the present technique has a much wider application as will become evident from the descriptions to follow.

In the following description of embodiments of the present disclosure reference is made to the accompanying drawings which form a part hereof, and in which are shown, by way of illustration, specific embodiments in which the technique may be practiced. It is understood that other embodiments may be utilized and structural changes may be made without departing from the scope of the present disclosure.

DESCRIPTION OF THE DRAWINGS

The specific features, aspects, and advantages of the disclosure will become better understood with regard to the following description, appended claims, and accompanying drawings where:

FIG. 1 is a diagram depicting a general purpose computing device constituting an exemplary system for a implementing a component of the present beamforming post-processor technique.

FIG. 2 is a diagram depicting one exemplary architecture of the present beamforming post-processor technique.

FIG. 3 is a flow diagram depicting one generalized exemplary embodiment of a process employing the present beamforming post-processor technique.

FIG. 4 is a flow diagram depicting one more detailed exemplary embodiment of a process employing the present beamforming post-processor technique.

DETAILED DESCRIPTION

1.0 The Computing Environment

Before providing a description of embodiments of the present Beamforming post-processor technique, a brief, general description of a suitable computing environment in which portions thereof may be implemented will be described. The present technique is operational with numerous general purpose or special purpose computing system environments or configurations. Examples of well known computing systems, environments, and/or configurations that may be suitable include, but are not limited to, personal computers, server computers, hand-held or laptop devices (for example, media players, notebook computers, cellular phones, personal data assistants, voice recorders), multiprocessor systems, microprocessor-based systems, set top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments that include any of the above systems or devices, and the like.

FIG. 1 illustrates an example of a suitable computing system environment. The computing system environment is only one example of a suitable computing environment and is not intended to suggest any limitation as to the scope of use or functionality of the present beamforming post-processor technique. Neither should the computing environment be interpreted as having any dependency or requirement relating to any one or combination of components illustrated in the exemplary operating environment. With reference to FIG. 1, an exemplary system for implementing the present beamforming post-processor technique includes a computing device, such as computing device 100. In its most basic configuration, computing device 100 typically includes at least one processing unit 102 and memory 104. Depending on the exact configuration and type of computing device, memory 104 may be volatile (such as RAM), non-volatile (such as ROM, flash memory, etc.) or some combination of the two. This most basic configuration is illustrated in FIG. 1 by dashed line 106. Additionally, device 100 may also have additional features/functionality. For example, device 100 may also include additional storage (removable and/or non-removable) including, but not limited to, magnetic or optical disks or tape. Such additional storage is illustrated in FIG. 1 by removable storage 108 and non-removable storage 110. Computer storage media includes volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules or other data. Memory 104, removable storage 108 and non-removable storage 110 are all examples of computer storage media. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium which can be used to store the desired information and which can accessed by device 100. Any such computer storage media may be part of device 100.

Device 100 has a sensor array 118, such as, for example, a microphone array, and may also contain communications connection(s) 112 that allow the device to communicate with other devices. Communications connection(s) 112 is an example of communication media. Communication media typically embodies computer readable instructions, data structures, program modules or other data in a modulated data signal such as a carrier wave or other transport mechanism and includes any information delivery media. The term “modulated data signal” means a signal that has one or more of its characteristics set or changed in such a manner as to encode information in the signal. By way of example, and not limitation, communication media includes wired media such as a wired network or direct-wired connection, and wireless media such as acoustic, RF, infrared and other wireless media. The term computer readable media as used herein includes both storage media and communication media.

Device 100 may have various input device(s) 114 such as a keyboard, mouse, pen, camera, touch input device, and so on. Output device(s) 116 such as a display, speakers, a printer, and so on may also be included. All of these devices are well known in the art and need not be discussed at length here.

The present beamforming post-processor technique may be described in the general context of computer-executable instructions, such as program modules, being executed by a computing device. Generally, program modules include routines, programs, objects, components, data structures, and so on, that perform particular tasks or implement particular abstract data types. The present beamforming post-processor technique may also be practiced in distributed computing environments where tasks are performed by remote processing devices that are linked through a communications network. In a distributed computing environment, program modules may be located in both local and remote computer storage media including memory storage devices.

The exemplary operating environment having now been discussed, the remaining parts of this description section will be devoted to a description of the program modules embodying the present beamforming post-processor technique.

2.0 Beamforming Post-Processor Technique

In one embodiment, the present beamforming post-processor technique is a non-linear post-processing technique for sensor arrays, which improves the directivity of the beamformer and separates the desired signal from noise. The technique works in so-called instantaneous direction of arrival space to estimate the probability of the signal coming from a given location (e.g., look-up direction in angular space) and uses this probability to apply a time-varying, gain-based, spatio-temporal filter for suppressing sounds coming from other non-desired directions other than the estimated sound source direction, resulting in minimal artifacts and musical noise.

2.2 Exemplary Architecture of the Present Beamforming Post-Processor Technique.

One exemplary architecture of the present beamforming post-processor technique 200 is shown in FIG. 2. This architecture 200 consists of a conventional beamformer 202 which receives inputs from an array of sensors, such as, for example, an array of microphones 204. The output of the beamformer 202 is input into a post-processor 206, which consists of a spatial filtering module 210 and a spatial probability estimation module 208 which employs an instantaneous direction of arrival computation. The spatial probability estimation module 208 estimates the probability that the desired signal originates from a given direction, θ_S, using the inputs from the array of sensors. This probability is then multiplied by the beamformer output in the spatial filtering module 210, to provide the desired sound source signal with an improved signal to noise ratio 212.

2.3 Exemplary Process Employing the Present Beamforming Post-Processor Technique.

One very general exemplary process employing the present post-processor beamforming technique is shown in FIG. 3. As shown in FIG. 3, box 302, signals of a sensor array in the frequency domain are input into a standard beamformer. A beamformer output is computed as a function of the input signals divided into frequency bins and an index of time frames (box 304). The probability that the desired signal originates a given direction θ_Sis computed using an instantaneous direction of arrival computation (box 306). This probability is multiplied by the beamformer output (box 308) to produce the desired signal with an enhanced signal to noise ratio (box 310).

More particularly, a more detailed exemplary process employing the present beamforming post-processor technique for a microphone is shown in FIG. 4. The audio signals captured by the microphone array x_i(l),i=1 . . . (M−1), where M is the number of microphones, are digitized using conventional analog to digital (A/D) conversion techniques, breaking the audio signals into frames (boxes 402, 404). The present beamforming post-processor technique then converts the time-domain signal x_i(n) to the frequency-domain (box 406). In one embodiment a modulated complex lapped transform (MCLT) is used for this purpose, although other conventional transforms could equally well be used. One can denote the frequency domain transform as x_i⁽ⁿ⁾(k), where k is the frequency bin, n is the index of the time-frame (e.g., frame), and i is the microphone (where i is 1 to M)).

The signals in the frequency domain, x_i⁽ⁿ⁾(k), are then input into a beamformer, whose output represents the optimal solution for capturing an audio signal at a target point using the total microphone array input (box 408). Additionally, the signals in the frequency domain are used to compute the instantaneous direction of arrival of the desired signal for each angular space (defined by incident angle or look-up angle (box 410)). This information is used to compute the spatial variation of the sound source position in presence of Noise (N(0,λ_IDOA(k))), for each frequency bin. The IDOA information and the spatial variation of the sound source in the presence of Noise is then used to compute the probability density that the desired sound source signal comes from a given direction, θ, for each frequency bin (box 412). This probability is used to compute the likelihood that for a frequency bin k of a given frame the desired signal originates from a given direction θ_S(414). If desired this likelihood can also optionally be temporally smoothed (box 416). The likelihood, smoothed or not, is then used to find the estimated probability that the desired signal originates from direction θ_S. Spatial filtering is then performed by multiplying the estimated probability the desired signal comes from a given direction by the beamformer output (box 418), outputting a signal with an enhanced signal to noise ratio (box 420). The final output in the time domain can be obtained by taking the inverse-MCLT (IMCLT) or corresponding inverse transformation of the transformation used to convert to frequency domain (inverse Fourier transformation, for example), of the enhanced signal in the frequency domain (box 422). Other processing such as encoding and transmitting the enhanced signal can also be performed (box 424).

2.4 Exemplary Computations

The following paragraphs provide exemplary models and exemplary computations that can be employed with the present beamforming post-processor technique.

2.4.1 Modeling

A typical beamformer is capable of providing optimized beam design for sensor arrays of any known geometry and operational characteristics. In particular, consider an array of M microphones with a known positions vector p. The microphones in the array sample the signal field in the workspace around the array at locations p_m=(x_m,y_m,z_m):m=0, 1, . . . , M−1. This sampling yields a set of signals that are denotes by the signal vector x(t, p).

Further, each microphone m has a known directivity pattern, U_m(f,c), where f is the frequency and c={φ,θ, ρ} represents the coordinates of a sound source in a radial coordinate system. A similar notation will be used to represent those same coordinates in a rectangular coordinate system, in this case, c={x,y,z}. As is known to those skilled in the art, the directivity pattern of a microphone is a complex function which provides the sensitivity and the phase shift introduced by the microphone for sounds coming from certain locations or directions. For an ideal omni-directional microphone, U_m(f,c)=constant. However, the microphone array can use microphones of different types and directivity patterns without loss of generality of the typical beamformer.

2.4.1.1 Sound Capture Model

Let vector p={p_mm=0, 1, . . . , M−1} denote the positions of the M microphones in the array, where p_m=(x_m,y_m,z_m). This yields a set of signals that one can denote by vector x(t, p). Each sensor m has known directivity pattern U_m(f,c), where c={φ,θ,ρ} represents the coordinates of the sound source in a radial coordinate system and f denotes the signal frequency. It is often preferable to perform signal processing algorithms in the frequency domain because efficient implementations can be employed.

As is known to those skilled in the art, a sound signal originating at a particular location, c, relative to a microphone array is affected by a number of factors. For example, given a sound signal, S(f), originating at point c, the signal actually captured by each microphone can be defined by Equation (1), as illustrated below:

X_m(f,p_m)=D_m(f,c)S(f)+N_m(f) (1)

where the first term on the right-hand side,

$\begin{matrix} D_{m} (f, c) = \frac{ⅇ^{- j 2 π f \frac{\langle c - p_{m} \rangle}{v}}}{ c - p_{m} } A_{m} (f) U_{m} (f, c) & (2) \end{matrix}$

represents the delay and decay due to the distance from the sound source to the microphone ∥c−p_m∥, and ν is the speed of sound. The term A_m(f) is the frequency response of the system preamplifier/ADC circuitry for each microphone, m, S(f) is the source signal, and N_m(f) is the captured noise. The variable U_m(f,c), accounts for microphone directivity relative to point c.

2.4.1.2 Ambient Noise Model

Given the captured signal, X_m(f,p_m), the first task is to compute noise models for modeling various types of noise within the local environment of the microphone array. The noise models described herein distinguish two types of noise: isotropic ambient nose and instrumental noise. Both time and frequency-domain modeling of these noise sources are well known to those skilled in the art. Consequently, the types of noise models considered will only be generally described below.

The captured noise N_m(f,p_m) is considered to contain two noise components: acoustic noise and instrumental noise. The acoustic noise, with spectrum denoted with N_A(f), is correlated across all microphone signals. The instrumental noise, having a spectrum denoted by the term N_l(f), represents electrical circuit noise from the microphone, preamplifier, and ADC (analog/digital conversion) circuitry. The instrumental noise in each channel is incoherent across the channels, and usually has a nearly white noise spectrum N_l(f). Assuming isotropic ambient noise one can represent the signal, captured by any of the microphones, as a sum of infinite number of uncorrelated noise sources randomly spread in space:

$\begin{matrix} N_{m} = N_{A} \sum_{l = 1}^{\infty} D_{m} (c_{l}) ℕ (0, λ_{I} (c_{l})) + N_{I} ℕ (0, λ_{I}) & (3) \end{matrix}$

Indices for frame and frequency are omitted for simplicity. Estimation of all of these noise sources is impossible because one has a finite number of microphones. Therefore, the isotropic ambient noise is modeled as one noise source in different positions in the work volume for each frame, plus a residual incoherent random component, which incorporates the instrumental noise. The noise capture equation changes to:

N_m⁽ⁿ⁾=D_m(c_n)N(0,λ_N(c_n))+N(0,λ_NC) (4)

where c_nis the noise source random position for n^thaudio frame, λ_N(c_n) is the spatially dependent correlated noise variation (λ_N(c_n)=const ∀c_nfor isotropic noise) and λ_NCis the variation of the incoherent component.

2.4.2 Spatio-Temporal Filter

The sound capture model and noise models having been described, the following paragraphs describe the computations performed in one embodiment of the present beamforming post-processor technique to obtain a spatial and temporal post-processor that improves the quality of the beamformer output of the desired signal. The following paragraphs are also referenced with respect to the flow diagram shown in FIG. 4.

2.4.2.1 Instantaneous Direction of Arrival Space

In general, for a microphone array, the phase differences at a particular frequency bin between the signals received at a pair of microphones give an indication of the instantaneous direction of arrival (IDOA) of a given sound source. IDOA vectors provide an indication of the direction from which a signal and/or point noise source originates. Non-correlated noise will be evenly spread in this space, while the signal and ambient noise (correlated components) will lie inside a hyper-volume that represents all potential positions of a sound source within the signal field.

To provide an indication of the direction a signal or noise source originates from (as indicated in FIG. 4, box 410), one can find the Instantaneous Direction of Arrival (IDOA) for each frequency bin based on the phase differences of non-repetitive pairs of input signals. For M microphones these phase differences form a M−1 dimensional space, spanning all potential IDOA. If one defines an IDOA vector in this space as

Δ(f) custom character [δ₁(f),δ₂(f), . . . , δ_M-1(f)] (5)

where δ_i(f) is the phase difference between channels 1 and i+1:

δ_l(f)=arg(X₁(f))−arg(X_l+l(f))l={1, . . . , M−1} (6)

then the non-correlated noise will be evenly spread in this space, while the signal and ambient noise (correlated components) will lay inside a hypervolume that represents all potential positions c={φ,θ,ρ} of a sound source in real three dimensional space. For far field sound capture, this is a M−1 dimensional hypersurface as the distance is presumed to approach infinity. Linear microphone arrays can distinguish only one dimension—the incident angle, and the real space is represented by a M−1 dimensional hyperline. For each frequency, a theoretical line that represents the positions of sound sources in the angular range of −90 degrees to +90 degrees can be computed using Equation (5). The actual distribution of the sound sources is a cloud around the theoretical line due to the presence of an additive non-correlated component. For each point in the real space there is a corresponding point in the IDOA space (which may be not unique). The opposite is not true: there are points in the IDOA space without corresponding point in the real space.

2.4.2.2 Presence of a Sound Source.

For simplicity and without any loss of generality, a linear microphone array is considered, sensitive only to the incident angle θ-direction of arrival in one dimension. The incident angle is defined by a discretization of space. For example, in one embodiment a set of angles is defined that is used to compute various parameters—probability, likelihood, etc. Such set can, for example, be in from −90 to +90 degrees every 5 degrees. Let Ψ_k(θ) denote the function that generates the vector Δ for given incident angle θ and frequency bin k according to equations (1), (5) and (6). In each frame, the k^thbin is represented by one point Δ_kin the IDOA space. Consider a sound source at θ_Swith its correspondence in IDOA space at Δ_S(k)=Ψ_k(θ_S). With additive noise, the resultant point in IDOA space will be spread around Δ_S(k):

Δ_S+N(k)=Δ_S(k)+N(0,λ_IDOA(k)). (7)

where N(0,θ_IDOA(k)) is the spatial movement of Δ_kin the IDOA space, caused by the correlated and non-correlated noises.

2.4.2.3 Space Conversion

The distance from each IDOA point to the theoretical in IDOA space is computed as a function of incident angle space, as shown in FIG. 4, box 412. The conversion from the distance from an IDOA point to the theoretical hyperline in IDOA space into the incident angle space (real world, one dimensional in this case) is given by:

$\begin{matrix} Υ_{k} (θ) = \frac{ Δ_{k} - Ψ_{k} (θ) }{ \frac{ⅆ Ψ_{k} (θ)}{ⅆ θ} } & (8) \end{matrix}$

where ∥Δ_k−Ψ_k(θ)∥ is the Euclidean distance between Δ_kand Ψ_k(θ) in IDOA space,

$\frac{ⅆ Ψ_{k} (θ)}{ⅆ θ}$

are the partial derivatives, and γ_k(θ) is the distance of observed IDOA point to the points in the real world. Note that the dimensions in IDOA space are measured in radians as phase difference, while γ_k(θ) is measured in radians as units of incident angle. This computation provides the distance between each IDOA point and the theoretical line as a function of the incident angle for each frequency bin and each frame.

2.4.2.4 Estimation of the Variance in Real Space

As shown in FIG. 4, box 414, in order to compute the probability that the sound source originates from a given incident angle, one must have the conversion from distance to the theoretical hyperline in IDOA space to distance into the incident angle space given by Equation (7) and the noise properties.

Analytic estimation in real-time of the probability density function for a sound source in every frequency bin is computationally expensive. Therefore the beamforming post-processor technique estimates indirectly the variation λ_k(θ) of the sound source position in presence of noise N(0,λ_IDOA(k)) from Equation (7). Let λ_k(θ) and γ_k(θ) be a K×N matrix, where K is the number of frequency bins and N is the number of discrete values of the incident or direction angle of the microphone. Variation estimation goes through two stages. During the first stage a rough variation estimation matrix λ (θ,k) is built. If θ_minis the angle that minimizes γ_k(θ), only the minimum values in the rough model are updated:

λ_k⁽ⁿ⁾(θ_min)=(1−α)λ_k^(n-1)(θ_min)+αγ_k(θ_min)₂ (9)

where γ is estimated according to Eq. (8),

$α = \frac{T}{τ_{A}}$

(τ_Ais the adaptation time constant, T is the frame duration). During the second stage a direction-frequency smoothing filter H (θ,k) is applied after each update to estimate the spatial variation matrix λ(θ,k)=H(θ,k)*λ(θ,k). Here it is assumed a Gaussian distribution of the non-correlated component, which allows one to assume the same deviation in the real space towards the incident angle, θ.

2.4.2.5 Likelihood Estimation

As shown in FIG. 4, box 416, a likelihood estimation that the desired signal comes from a given incident angle is computed using the IDOA information and the variation due to noise. With known spatial variation λ_k(θ) and the distance of the observed IDOA points to the points in the real world, γ_k(θ), the probability density for frequency bin k to originate from direction θ is given by:

$\begin{matrix} p_{k} (θ) = \frac{1}{\sqrt{2 π λ_{k} (θ)}} \exp {- \frac{{Υ_{k} (θ)}^{2}}{2 λ_{k} (θ)}}, & (10) \end{matrix}$

and for a given direction, θ_S, the likelihood that the sound source originates from this direction for a given frequency bin is:

$\begin{matrix} Λ_{k} (θ_{S}) = \frac{p_{k} (θ_{S})}{p_{k} (θ_{\min})}, & (11) \end{matrix}$

where θ_minis the value which minimizes p_k(θ).

2.4.2.6 Spatio-Temporal Filtering

Besides spatial position, the desired (e.g., speech) signal has temporal characteristics and consecutive frames are highly correlated due to the fact that this signal changes slowly relatively to the frame duration. Rapid change of the estimated spatial filter can cause musical noise and distortions in the same way as in gain based noise suppressors. As shown in FIG. 4, box 418, to reflect the temporal characteristics of the speech signal, temporal smoothing can optionally be applied. For a given direction, the absence/presence of speech can be modeled with two states: S₀and S₁. The sequence of frequency bin states is modeled as first-order Markov process. Then the pseudo-stationarity property of the desired (e.g., speech) signal can be represented by P(q_n=S₁|q_n-1=S₁) with the following constraint: P(q_n=S₁|q_n-1=S₁)>P(q_n=S₁), where q_ndenotes the state of n-th frame as either S₀or S₁. By assuming that the Markov process is time invariant, one can use the notation a_ij custom character P(q_n=H_j|q_{n . . . 1}=H_j). Based on the formulations above, a recursive formula for signal presence likelihood for given look-up direction in n^thframe Λ_k⁽ⁿ⁾is obtained as:

$\begin{matrix} Λ_{k}^{(n)} (θ_{S}) = \frac{a_{01} + a_{11} Λ_{k}^{(n - 1)} (θ_{S})}{a_{00} + a_{10} Λ_{k}^{(n - 1)} (θ_{S})} Λ_{k} (θ_{S}), & (12) \end{matrix}$

where a_ijare the transition probabilities, Λ_k(θ_S) is estimated by Equation (11), and Λ_k⁽ⁿ⁾(θ_S) is the likelihood of having a signal at direction θ_Sfor n^thframe. As shown in FIG. 4, box 420, this likelihood can be converted to a probability and spatial filtering can be performed by multiplying the probability that the desired signal comes form a given direction times the beamformer output. More specifically, conversion to probability gives the estimated probability for the speech signal to originate from this direction:

$\begin{matrix} P_{k}^{(n)} (θ_{S}) = \frac{Λ_{k}^{(n)} (θ_{S})}{1 + Λ_{k}^{(n)} (θ_{S})} . & (13) \end{matrix}$

The spatio-temporal filter to compute the post-processor output Z_k⁽ⁿ⁾(for all frequency bins in the current frame) from the beamformer output Y_k⁽ⁿ⁾is:

Z_k⁽ⁿ⁾=P_k⁽ⁿ⁾(θ_S)·Y_k⁽ⁿ⁾, (14)

i.e., the signal presence probability is used as a suppression.

It should also be noted that any or all of the aforementioned alternate embodiments may be used in any combination desired to form additional hybrid embodiments. For example, even though this disclosure describes the present beamforming post-processor technique with respect to a microphone array, the present technique is equally applicable to sonar arrays, directional radio antenna arrays, radar arrays, and the like. Although the subject matter has been described in language specific to structural features and/or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. The specific features and acts described above are disclosed as example forms of implementing the claims.

Claims

1. A computer-implemented process for improving the directivity and signal to noise ratio of the output of a beamformer employed with a sensor array in an environment, comprising: capturing sound sources dispersed at locations in the environment with sensors of a sensor array;inputting signals of the sound sources and a desired signal captured by the sensors of a sensor array in the frequency domain defined by frequency bins and frames in time;computing a beamformer output as function of the input signals divided into frequency bins and frames in time;dividing a spatial region corresponding to a working space of the sensor array into a plurality of incident angle regions, and for each frequency bin and incident angle region, computing the probability that the desired signal occurs at a given incident angle region using an instantaneous direction of arrival computation; andspatially filtering the beamformer output by multiplying the probability that the desired signal occurs at a given incident angle region by the beamformer output while attenuating signals from the locations of the sound sources.
2. The computer-implemented process of claim 1 wherein the captured sound sources are estimated locations.
3. The computer-implemented process of claim 1 wherein the input signals in the frequency domain are converted from the time domain into the frequency domain prior to inputting them using a Modulated Complex Lapped Transform (MCLT).
4. The computer-implemented process of claim 1 wherein the sensors are microphones and wherein the sensor array is a microphone array.
5. The computer-implemented process of claim 1 wherein the instantaneous direction of arrival computation for each frequency bin is based on the phase differences of the input signals from a pair of sensors.
6. The computer-implemented process of claim 1 wherein spatially filtering the beamformer output attenuates signals originating from directions other than the direction of the desired signal.
7. A system for improving the signal to noise ratio of a desired signal received from a microphone array in an environment, comprising: a general purpose computing device;a computer program comprising program modules executable by the general purpose computing device, wherein the computing device is directed by the program modules of the computer program to,capture audio signals of dispersed sound sources and a desired signal in an environment in the time domain with a microphone array;convert the time-domain signals to frequency-domain and frequency bins using a converter;input the signals in the frequency domain into a beamformer and compute a beamformer output wherein the beamformer output represents the optimal solution for capturing an audio signal at a target point using the total microphone array input;estimate the probability that the desired signal comes from a given incident angle using an instantaneous direction of arrival computation; andoutput an enhanced signal for the desired signal with a greater signal to noise ratio by taking the product of the beamformer output and the probability estimation that the desired signal comes from a given incident angle while attenuating audio signals that come from directions of the dispersed sound sources.
8. The system of claim 7 wherein the instantaneous direction of arrival computation for each frequency bin is based on the phase differences of the input signals from a pair of microphones.
9. The system of claim 7 wherein the beamformer is a time-invariant beamformer.
10. The system of claim 7 wherein the enhanced signal with a greater signal to noise ratio is computed and output in real time.
11. The system of claim 7 wherein the modules to estimate the probability that a desired signal comes from a given incident angle using an instantaneous direction of arrival computation and the module to output an enhanced signal with a greater signal to noise ratio by taking the product of the beamformer output and the probability estimation that the desired signal comes from a given incident angle form a post-processor that attenuates signals originating from directions other than the direction of the desired signal to output a signal with an enhanced signal to noise ratio.
12. A computer-implemented process for improving the signal to noise ratio of a desired signal received from a microphone array in an environment, comprising: capturing audio signals of dispersed sound sources and a desired signal in the environment in the time domain with a microphone array;converting the time-domain signals to frequency-domain and frequency bins using a converter;inputting the signals in the frequency domain into a beamformer and computing a beamformer output wherein the beamformer output represents the optimal solution for capturing an audio signal at a target point using the total microphone array input;estimating the probability that the desired signal comes from a given incident angle using an instantaneous direction of arrival computation; andoutputting an enhanced signal of the desired signal with a greater signal to noise ratio by taking the product of the beamformer output and the probability estimation that the desired signal comes from a given incident angle.
13. The computer-implemented process of claim 12 further comprising attenuating audio signals that come from directions of the sound sources.
14. The computer-implemented process of claim 12 wherein the instantaneous direction of arrival computation for each frequency bin is based on the phase differences of the input signals from a pair of microphones.
15. The computer-implemented process of claim 12 wherein the beamformer is a time-invariant beamformer.
16. The computer-implemented process of claim 12 wherein the enhanced signal with a greater signal to noise ratio is computed and output in real time.
17. The computer-implemented process of claim 12 wherein the instantaneous direction of arrival computation is based on the phase differences of the input signals from a pair of sensors.
18. The computer-implemented process of claim 12 wherein the dispersed sound sources are at estimated locations.

Parent Case Info

The above-identified application is a continuation of a prior application entitled “Sensor Array Beamformer Post-Processor” which was assigned Ser. No. 11/750,319, and was filed on May 17, 2007.

US Referenced Citations (181)

Number	Name	Date	Kind
4627620	Yang	Dec 1986	A
4630910	Ross et al.	Dec 1986	A
4645458	Williams	Feb 1987	A
4695953	Blair et al.	Sep 1987	A
4702475	Elstein et al.	Oct 1987	A
4711543	Blair et al.	Dec 1987	A
4751642	Silva et al.	Jun 1988	A
4796997	Svetkoff et al.	Jan 1989	A
4809065	Harris et al.	Feb 1989	A
4817950	Goo	Apr 1989	A
4843568	Krueger et al.	Jun 1989	A
4893183	Nayar	Jan 1990	A
4901362	Terzian	Feb 1990	A
4925189	Braeunig	May 1990	A
5101444	Wilson et al.	Mar 1992	A
5148154	MacKay et al.	Sep 1992	A
5184295	Mann	Feb 1993	A
5229754	Aoki et al.	Jul 1993	A
5229756	Kosugi et al.	Jul 1993	A
5239463	Blair et al.	Aug 1993	A
5239464	Blair et al.	Aug 1993	A
5288078	Capper et al.	Feb 1994	A
5295491	Gevins	Mar 1994	A
5320538	Baum	Jun 1994	A
5347306	Nitta	Sep 1994	A
5385519	Hsu et al.	Jan 1995	A
5405152	Katanics et al.	Apr 1995	A
5417210	Funda et al.	May 1995	A
5423554	Davis	Jun 1995	A
5454043	Freeman	Sep 1995	A
5469740	French et al.	Nov 1995	A
5495576	Ritchey	Feb 1996	A
5516105	Eisenbrey et al.	May 1996	A
5524637	Erickson et al.	Jun 1996	A
5534917	MacDougall	Jul 1996	A
5563988	Maes et al.	Oct 1996	A
5577981	Jarvik	Nov 1996	A
5580249	Jacobsen et al.	Dec 1996	A
5594469	Freeman et al.	Jan 1997	A
5597309	Riess	Jan 1997	A
5616078	Oh	Apr 1997	A
5617312	Iura et al.	Apr 1997	A
5638300	Johnson	Jun 1997	A
5641288	Zaenglein	Jun 1997	A
5682196	Freeman	Oct 1997	A
5682229	Wangler	Oct 1997	A
5690582	Ulrich et al.	Nov 1997	A
5703367	Hashimoto et al.	Dec 1997	A
5704837	Iwasaki et al.	Jan 1998	A
5715834	Bergamasco et al.	Feb 1998	A
5875108	Hoffberg et al.	Feb 1999	A
5877803	Wee et al.	Mar 1999	A
5913727	Ahdoot	Jun 1999	A
5933125	Fernie	Aug 1999	A
5980256	Carmein	Nov 1999	A
5989157	Walton	Nov 1999	A
5995649	Marugame	Nov 1999	A
6005548	Latypov et al.	Dec 1999	A
6009210	Kang	Dec 1999	A
6054991	Crane et al.	Apr 2000	A
6066075	Poulton	May 2000	A
6072494	Nguyen	Jun 2000	A
6073489	French et al.	Jun 2000	A
6077201	Cheng et al.	Jun 2000	A
6098458	French et al.	Aug 2000	A
6100896	Strohecker et al.	Aug 2000	A
6101289	Kellner	Aug 2000	A
6128003	Smith et al.	Oct 2000	A
6130677	Kunz	Oct 2000	A
6141463	Covell et al.	Oct 2000	A
6147678	Kumar et al.	Nov 2000	A
6152856	Studor et al.	Nov 2000	A
6159100	Smith	Dec 2000	A
6173066	Peurach et al.	Jan 2001	B1
6181343	Lyons	Jan 2001	B1
6188777	Darrell et al.	Feb 2001	B1
6215890	Matsuo et al.	Apr 2001	B1
6215898	Woodfill et al.	Apr 2001	B1
6226396	Marugame	May 2001	B1
6229913	Nayar et al.	May 2001	B1
6256033	Nguyen	Jul 2001	B1
6256400	Takata et al.	Jul 2001	B1
6283860	Lyons et al.	Sep 2001	B1
6289112	Jain et al.	Sep 2001	B1
6299308	Voronka et al.	Oct 2001	B1
6308565	French et al.	Oct 2001	B1
6316934	Amorai-Moriya et al.	Nov 2001	B1
6363160	Bradski et al.	Mar 2002	B1
6384819	Hunter	May 2002	B1
6411744	Edwards	Jun 2002	B1
6430997	French et al.	Aug 2002	B1
6476834	Doval et al.	Nov 2002	B1
6496598	Harman	Dec 2002	B1
6503195	Keller et al.	Jan 2003	B1
6539931	Trajkovic et al.	Apr 2003	B2
6570555	Prevost et al.	May 2003	B1
6633294	Rosenthal et al.	Oct 2003	B1
6640202	Dietz et al.	Oct 2003	B1
6661918	Gordon et al.	Dec 2003	B1
6681031	Cohen et al.	Jan 2004	B2
6714665	Hanna et al.	Mar 2004	B1
6731799	Sun et al.	May 2004	B1
6738066	Nguyen	May 2004	B1
6765726	French et al.	Jul 2004	B2
6788809	Grzeszczuk et al.	Sep 2004	B1
6801637	Voronka et al.	Oct 2004	B2
6873723	Aucsmith et al.	Mar 2005	B1
6876496	French et al.	Apr 2005	B2
6937742	Roberts et al.	Aug 2005	B2
6950534	Cohen et al.	Sep 2005	B2
7003134	Covell et al.	Feb 2006	B1
7036094	Cohen et al.	Apr 2006	B1
7038855	French et al.	May 2006	B2
7039676	Day et al.	May 2006	B1
7042440	Pryor et al.	May 2006	B2
7050606	Paul et al.	May 2006	B2
7058204	Hildreth et al.	Jun 2006	B2
7060957	Lange et al.	Jun 2006	B2
7113918	Ahmad et al.	Sep 2006	B1
7121946	Paul et al.	Oct 2006	B2
7170492	Bell	Jan 2007	B2
7184048	Hunter	Feb 2007	B2
7202898	Braun et al.	Apr 2007	B1
7222078	Abelow	May 2007	B2
7227526	Hildreth et al.	Jun 2007	B2
7259747	Bell	Aug 2007	B2
7308112	Fujimura et al.	Dec 2007	B2
7317836	Fujimura et al.	Jan 2008	B2
7348963	Bell	Mar 2008	B2
7359121	French et al.	Apr 2008	B2
7367887	Watabe et al.	May 2008	B2
7379563	Shamaie	May 2008	B2
7379566	Hildreth	May 2008	B2
7389591	Jaiswal et al.	Jun 2008	B2
7412077	Li et al.	Aug 2008	B2
7421093	Hildreth et al.	Sep 2008	B2
7430312	Gu	Sep 2008	B2
7436496	Kawahito	Oct 2008	B2
7450736	Yang et al.	Nov 2008	B2
7452275	Kuraishi	Nov 2008	B2
7460690	Cohen et al.	Dec 2008	B2
7489812	Fox et al.	Feb 2009	B2
7536032	Bell	May 2009	B2
7555142	Hildreth et al.	Jun 2009	B2
7560701	Oggier et al.	Jul 2009	B2
7570805	Gu	Aug 2009	B2
7574020	Shamaie	Aug 2009	B2
7576727	Bell	Aug 2009	B2
7590262	Fujimura et al.	Sep 2009	B2
7593552	Higaki et al.	Sep 2009	B2
7598942	Underkoffler et al.	Oct 2009	B2
7607509	Schmiz et al.	Oct 2009	B2
7620202	Fujimura et al.	Nov 2009	B2
7668340	Cohen et al.	Feb 2010	B2
7680298	Roberts et al.	Mar 2010	B2
7683954	Ichikawa et al.	Mar 2010	B2
7684592	Paul et al.	Mar 2010	B2
7701439	Hillis et al.	Apr 2010	B2
7702130	Im et al.	Apr 2010	B2
7704135	Harrison, Jr.	Apr 2010	B2
7710391	Bell et al.	May 2010	B2
7729530	Antonov et al.	Jun 2010	B2
7746345	Hunter	Jun 2010	B2
7760182	Ahmad et al.	Jul 2010	B2
7809167	Bell	Oct 2010	B2
7834846	Bell	Nov 2010	B1
7852262	Namineni et al.	Dec 2010	B2
RE42256	Edwards	Mar 2011	E
7898522	Hildreth et al.	Mar 2011	B2
8005238	Tashev et al.	Aug 2011	B2
8035612	Bell et al.	Oct 2011	B2
8035614	Bell et al.	Oct 2011	B2
8035624	Bell et al.	Oct 2011	B2
8072470	Marks	Dec 2011	B2
8818002	Tashev et al.	Aug 2014	B2
20050141731	Hamalainen	Jun 2005	A1
20060147054	Buck et al.	Jul 2006	A1
20060171547	Lokki et al.	Aug 2006	A1
20060233389	Mao et al.	Oct 2006	A1
20080026838	Dunstan et al.	Jan 2008	A1
20080288219	Tashev et al.	Nov 2008	A1

Foreign Referenced Citations (6)

Number	Date	Country
201254344	Jun 2010	CN
0583061	Feb 1994	EP
08044490	Feb 1996	JP
9310708	Jun 1993	WO
9717598	May 1997	WO
9944698	Sep 1999	WO

Non-Patent Literature Citations (27)

Entry
Kanade et al., “A Stereo Machine for Video-rate Dense Depth Mapping and Its New Applications”, IEEE Computer Society Conference on Computer Vision and Pattern Recognition, Jun. 18-20, 1996, pp. 196-202,The Robotics Institute, Carnegie Mellon University, Pittsburgh, PA.
Miyagawa et al., “CCD-Based Range Finding Sensor”, Oct. 1997, pp. 1648-1652, vol. 44 No. 10, IEEE Transactions on Electron Devices.
Rosenhahn et al., “Automatic Human Model Generation”, Sep. 5-8, 2005, pp. 41-48, University of Auckland (CITR), New Zealand.
Aggarwal et al., “Human Motion Analysis: A Review”, IEEE Nonrigid and Articulated Motion Workshop, Jun. 1997, University of Texas at Austin, Austin, TX.
Shao et al., “An Open System Architecture for a Multimedia and Multimodal User Interface”, Aug. 24, 1998, Japanese Society for Rehabilitation of Persons with Disabilities (JSRPD), Japan.
Kohler, “Special Topics of Gesture Recognition Applied in Intelligent Home Environments”, In Proceedings of the Gesture Workshop, Sep. 17-19, 1997, pp. 285-296, Germany.
Kohler, “Vision Based Remote Control in Intelligent Home Environments”, University of Erlangen-Nuremberg/Germany, 1996, pp. 147-154, Germany.
Kohler, “Technical Details and Ergonomical Aspects of Gesture Recognition applied in Intelligent Home Environments”, 1997, Germany.
Hasegawa et al., “Human-Scale Haptic Interaction with a Reactive Virtual Human in a Real-Time Physics Simulator”, Jul. 2006, vol. 4, No. 3, Article 6C, ACM Computers in Entertainment, New York, NY.
Qian et al., “A Gesture-Driven Multimodal Interactive Dance System”, Jun. 2004, pp. 1579-1582, IEEE International Conference on Multimedia and Expo (ICME), Taipei, Taiwan.
Zhao, “Dressed Human Modeling, Detection, and Parts Localization”, Jun. 26, 2001, The Robotics Institute, Carnegie Mellon University, Pittsburgh, PA.
He, “Generation of Human Body Models”, Apr. 2005, University of Auckland, New Zealand.
Isard et al., “Condensation—Conditional Density Propagation for Visual Tracking”, Aug. 1998, pp. 5-28, International Journal of Computer Vision 29(1), Netherlands.
Livingston, “Vision-based Tracking with Dynamic Structured Light for Video See-through Augmented Reality”, Oct. 1998, University of North Carolina at Chapel Hill, North Carolina, USA.
Wren et al., “Pfinder: Real-Time Tracking of the Human Body”, MIT Media Laboratory Perceptual Computing Section Technical Report No. 353, Jul. 1997, vol. 19, No. 7, pp. 780-785, IEEE Transactions on Pattern Analysis and Machine Intelligence, Caimbridge, MA.
Breen et al., “Interactive Occlusion and Collusion of Real and Virtual Objects in Augmented Reality”, Technical Report ECRC-95-02, 1995, European Computer-Industry Research Center GmbH, Munich, Germany.
Freeman et al., “Television Control by Hand Gestures”, Dec. 1994, Mitsubishi Electric Research Laboratories, TR94-24, Caimbridge, MA.
Hongo et al., “Focus of Attention for Face and Hand Gesture Recognition Using Multiple Cameras”, Mar. 2000, pp. 156-161, 4th IEEE International Conference on Automatic Face and Gesture Recognition, Grenoble, France.
Pavlovic et al., “Visual Interpretation of Hand Gestures for Human-Computer Interaction: A Review”, Jul. 1997, pp. 677-695, vol. 19, No. 7, IEEE Transactions on Pattern Analysis and Machine Intelligence.
Azarbayejani et al., “Visually Controlled Graphics”, Jun. 1993, vol. 15, No. 6, IEEE Transactions on Pattern Analysis and Machine Intelligence.
Granieri et al., “Simulating Humans in VR”, The British Computer Society, Oct. 1994, Academic Press.
Brogan et al., “Dynamically Simulated Characters in Virtual Environments”, Sep./Oct. 1998, pp. 2-13, vol. 18, Issue 5, IEEE Computer Graphics and Applications.
Fisher et al., “Virtual Environment Display System”, ACM Workshop on Interactive 3D Graphics, Oct. 1986, Chapel Hill, NC.
“Virtual High Anxiety”, Tech Update, Aug. 1995, pp. 22.
Sheridan et al., “Virtual Reality Check”, Technology Review, Oct. 1993, pp. 22-28, vol. 96, No. 7.
Stevens, “Flights into Virtual Reality Treating Real-World Disorders”, The Washington Post, Mar. 27, 1995, Science Psychology, 2 Pages.
“Simulation and Training”, 1994, Division Incorporated, pp. 1-6.

Related Publications (1)

	Number	Date	Country
	20110274289 A1	Nov 2011	US

Continuations (1)

	Number	Date	Country
Parent	11750319	May 2007	US
Child	13187235		US

Sensor array beamformer post-processor

Information

Patent Number

Date Filed

Date Issued

Inventors

Original Assignees

Examiners

Agents

CPC

Field of Search

US

International Classifications

Disclaimer

Term Extension

Abstract