1. Priority Claim
This application claims the benefit of priority from European Patent Application No. 06 009468.7, filed May 8, 2006, which is incorporated by reference.
2. Technical Field
This disclosure relates to echo reduction. In particular, this disclosure relates to echo reduction and suppression of residual echo signals in communication systems.
3. Related Art
Echo reduction or suppression may be needed in communication systems, such as hands-free sets and speech recognition systems. Communication systems may include a microphone that detects a desired signal, such as a speech signal from a user. However, the microphone may also detect undesirable signals, such as echoes produced by a loudspeaker in the audio environment.
Echoes may occur when signals transmitted by a remote party and received by a near-end are output by loudspeakers at the near end. Such signals may be detected by the near-end microphone and re-transmitted back to the remote party. Echoes may be annoying to the user and may result in a communication failure.
Echo suppression may be particularly difficult if the speaker is moving, such as when a driver is using a hands-free set and also steers a vehicle. In this situation, an impulse response of the “loudspeaker-room-microphone” (LRM) environment may be time-variant. Existing echo suppression systems may be unreliable and may not be particularly effective in a time-varying LRM environment. After application of echo suppression techniques, residual echoes may be present. The severity of the residual echoes may be increased by the large delay paths associated with mobile phone services. Accordingly, a need exists for an echo reduction system capable of reducing echoes and residual echoes.
A filter bank may separate a microphone output signal into sub-band microphone signals. An audio input filter bank may separate audio input signals into of sub-band audio signals. An echo compensation filter receives the audio input signals from each sub-band range and may generate a corresponding estimated echo signal. A speech activity detector may detect the speech of a local speaker by analyzing the power spectrum. The estimated echo signal may be subtracted from the microphone output signal to compensate for echoes. A residual echo reduction circuit may remove any residual echoes in the echo compensated signal based on speech activity.
Other systems, methods, features and advantages will be, or will become, apparent to one with skill in the art upon examination of the following figures and detailed description. It is intended that all such additional systems, methods, features and advantages be included within this description, be within the scope of the invention, and be protected by the following claims.
The system may be better understood with reference to the following drawings and description. The components in the figures are not necessarily to scale, emphasis instead being placed upon illustrating the principles of the invention. Moreover, in the figures, like-referenced numerals designate corresponding parts throughout the different views.
The microphone 120 may detect undesirable signals, such as echo signals produced by the loudspeaker 124. Echoes may occur when signals transmitted by the remote communication party 134 are output by loudspeakers at the near end, and retransmitted via the microphone back to the remote communication party 134. Such echoes may occur in various speech recognition systems or speech dialog systems.
The microphone 120 may detect speech signals s(n) of the local speaker 130 as well as background noise signals b(n). The microphone 120 may further detect a “loudspeaker-room-microphone” (LRM) transfer signal d(n) based on an impulse response h(n) of a LRM environment or system 140. A microphone output signal y(n) may include contributions from the speech signal s(n), the background noise signal b(n) and the LRM transfer signal d(n). The term “n” may denote a discrete time index.
The echo reduction system 110 may provide echo compensation and residual echo suppression. Echo compensation may be performed by an adaptive echo compensation filtering system that may model the loudspeaker-room-microphone system transfer function using an impulse response. Such processing may be performed by dividing the signal spectrum into sub-bands and processing each sub-band separately. Alternatively, the input signals my be processed in the frequency domain using Fast Fourier Transforms of the respective audio input signal x(n) and the microphone output signal y(n).
An audio input filter bank 146 may generate a plurality of sub-band audio signals X(ejΩ
Echo compensation, detection of speech activity (whether the speaker is silent or is speaking) and suppression of residual echo may be performed in the sub-band ranges. After suppression of the residual echo in the sub-bands, a synthesizer 230 may synthesize the desired output signal, which may be transmitted to the remote communication party 134.
The audio sub-band signals X(ejΩ
The echo compensation filter 160 may include hardware and/or software, and may include a digital signal processor (DSP). The DSP may execute instructions that delay an input signal one or more additional times, track frequency components of a signal, filter a signal, and/or attenuate or boost an amplitude of a signal. Alternatively, the echo compensation filter 160 or DSP may be implemented as discrete logic or circuitry, a mix of discrete logic and a processor, or may be distributed over multiple processors or software programs.
The echo compensation filter 160 may generate the estimated echo signal {circumflex over (D)}(ejΩ
A residual echo reduction circuit 170 may process the echo compensated signals E(ejΩ
The residual echo reduction circuit 170 may use a Wiener filter implementation to estimate a power density spectrum of the residual echo Ŝεε(Ωμ,n) and estimate the power density spectrum of the echo compensated sub-band signals Ŝee(Ωμ,n). The Wiener filter may exhibit the following filter response:
A maximum damping value may be determined by the parameter Gmin, and the sensitivity of the filter may be controlled by the parameter β. If β>1, the damping may be greater than desired, which may dampen the desired signal below an predetermined level.
The filter characteristic or frequency response of the residual echo reduction filter 170 may be adjusted so that damping is sensitive (“aggressive”) when local speaker 130 speech is absent. The echo reduction system 110 may distinguish speech signals produced by the local speaker 130 from signals transmitted by the loudspeaker 124 in the time-variant LRM environment 140. Such local speaker speech signals s(n) may be detected when the local speaker 130 is moving. Accordingly, the filter response of the residual echo reduction circuit 170 may adapt to the speech activity of the local speaker 130.
where Ωμ denotes the mid-frequency of the sub-band μ, and the term “smooth” indicates a selected smoothing function. The pre-determined range of sub-bands may cover the range between about 200 Hz and about 3500 Hz. This range of frequencies may show significant speech signal power levels.
The smoothing circuits 420 and 426 may include a first order recursive filter to smooth the magnitudes, or the squares of the magnitudes of the signals. Smoothing may be performed in a positive direction (Ω0 to ΩM−1) or in a negative direction (ΩM−1 to Ω0) with respect to the frequency range. Smoothing may be performed according to the following equations:
in the positive direction and
in the negative direction.
For a sampling rate of about 11025 Hz and M=256 sub-bands, a smoothing parameter may be selected so that 0.2≦λFre≦0.8. Smoothing the microphone sub-band signals Y(ejΩ
The smoothing circuits 420 and 426 may also receive the output Ŝbb(Ωμ,n) of the noise estimating circuit 414, which may be based on the estimated power density spectrum of the microphone sub-band signals Y(ejΩ
where the background noise may be overestimated by Kb.
Satisfactory echo suppression results may be obtained when 2≦Kb≦16. Weighting functions w1S{circumflex over (d)}{circumflex over (d)},mod(Ωμ,n)=w2Syy,mod(Ωμ,n) may be chosen to determine a distance measure, which may indicate speech activity. The values of w1 and w2 may be selected depending on the value of Ωμ. Alternatively, the values of w1 and w2 may be constants.
A flank detection circuit 430 may detect a strong level increase or decrease of the microphone sub-band signals Y(ejΩ
using a detection threshold of 4≦KΔ≦100.
A distance detection circuit 436 may determine a spectrum distance measure based on Δ(Ωμ,n), and the modified spectra S{circumflex over (d)}{circumflex over (d)},mod(Ωμ,n) and Syy,mod(Ωμ,n), calculated by the smoothing circuits 420 and 426 according to the following equations:
The detection threshold values may be selected as: K1=16, K2=4, K3=4, and K4=16. The distances C(Ωμ,n) may be specified by detection parameters C1=−0.4, C2=0.1, C3=0.1, and C4=0.6. Large positive values of the spectrum distance measure C(Ωμ,n) may indicate that the power density of the microphone spectrum dominates the power density of the estimated echo. If the power density of the estimated echo significantly exceeds the power density of the microphone spectrum, strong changes in the LRM system 140 may be indicated, which may result in a high negative cost parameter C1.
The determination of both Δ(Ωμ,n) and C(Ωμ,n) may be restricted to sub-bands that show a significant power of speech signals. This may indicate that the sub-bands may be restricted to μ∈[μstart,μend] where μstart and μend correspond to a frequency range of about 200 Hz to about 300 Hz. A summing circuit 440 may sum the results C(Ωμ,n) for the individual sub-bands, where
Smoothing over a pre-determined time interval may be performed to obtain a smoothed distance measure
Based upon the detected speech activity as measured by
Double-talk may occur when both the local speaker 130 and the remote communication party 134 speak simultaneously, and significant speech activity is detected. When double-talk occurs, a Wiener filter with a time-dependent filter parameter β(n) may be used. The residual echo reduction circuit 170 may process residual echo according to the following equations:
Gmod(ejΩ
The mid-frequency of the sub-band μ may be denoted by Ωμ, and Ŝee(Ωμ,n) and Ŝεε(Ωμ,n) may denote the estimated power density of the echo compensated signal and the estimated power density of the residual echo, respectively. The discrete time index may be denoted by n, and β(n) may be a filter parameter based upon the detected speech activity.
The parameter β(n) may control the sensitivity of the filter based on the smoothed distance measure according to the following equation:
where suitable values for the β-parameters may be β1=1 and β2=1000. A predetermined threshold value, Cthres, may indicate the presence of significant speech activity from the local speaker 130. Echo suppression may be limited by the value of Gmin, e.g., Gmin=0.1.
If the estimated power density of the echo compensated signal greatly exceeds the estimated power density of the residual echo, suppression of the residual echo may be minimized so that the microphone output signal is not significantly modified. In contrast, if the estimated power density of the residual echo exceeds the estimated power density of the echo compensated signal, compensated signal may be filtered aggressively.
Suppressing the residual echo in the echo compensated signal may include filtering the echo compensated signal with a filter having the following frequency response:
where Ŝee(Ωμ,n) and Ŝεε(Ωμ,n) denote the estimated power density of the echo compensated signal. The estimated power density of the residual echo and β(n) may be a filter parameter based on the detected speech activity.
The power density spectrum Ŝee(Ωμ,n) of the echo compensated sub-band microphone signals E(ejΩ
Ŝee(Ωμ, n)=λeŜee(Ωμ,n−1)+(1−λe)|E(ejΩ
where a smoothing parameter may be chosen as 0≦λe≦1.
Further, the power density spectrum Ŝεε(Ωμ,n) of the residual echo may be determined according to the following equation:
Ŝεε(Ωμ,n)=Ŝxx(Ωμ,n)|ĤΔ(ejΩ
where the estimated echo compensation ĤΔ(ejΩ
The estimated power density spectrum Ŝxx(Ωμ,n) of the audio signal output of the loudspeaker 124 may be calculated according to the following equation:
Ŝxx(Ωμ,n)=λxŜxx(Ωμ,n−1)+(1−λx)|X(ejΩ
where the smoothing parameter may be chosen as 0≦λx≦1.
A β-parameter circuit 750 may receive input from the local speech activity detector 720 and determine an appropriate value of β. The residual reduction circuit 170 may utilize the β-parameter value.
A local speech activity detector 720 may include the first and second smoothing circuit 420 and 426 shown in
The artificial noise generator 730 may generate artificial noise with substantially the same statistical power distribution as the background noise. The artificially generated noise signal or “comfort noise” signal B(ejΩ
The desired signal or sub-band output signal Ŝ(ejΩ
where E(ejΩ
The LRM environment 140 was changed by simulating movement of the local speaker 130 at a time t=5 seconds, and again at a time t=15 seconds. Panel D shows the addition of double talk from a time t=20.5 seconds to a time t=27.5 seconds.
Panel C shows the smoothed distance measure
The logic, circuitry, and processing described above may be encoded in a computer-readable medium such as a CDROM, disk, flash memory, RAM or ROM, an electromagnetic signal, or other machine-readable medium as instructions for execution by a processor. Alternatively or additionally, the logic may be implemented as analog or digital logic using hardware, such as one or more integrated circuits (including amplifiers, adders, delays, and filters), or one or more processors executing amplification, adding, delaying, and filtering instructions; or in software in an application programming interface (API) or in a Dynamic Link Library (DLL), functions available in a shared memory, or defined as local or remote procedure calls; or as a combination of hardware and software.
The logic may be represented in (e.g., stored on or in) a computer-readable medium, machine-readable medium, propagated-signal medium, and/or signal-bearing medium. The media may comprise any device that contains, stores, communicates, propagates, or transports executable instructions for use by or in connection with an instruction executable system, apparatus, or device. The machine-readable medium may selectively be, but is not limited to, an electronic, magnetic, optical, electromagnetic, or infrared signal or a semiconductor system, apparatus, device, or propagation medium. A non-exhaustive list of examples of a machine-readable medium includes: a magnetic or optical disk, a volatile memory such as a Random Access Memory “RAM,” a Read-Only Memory “ROM,” an Erasable Programmable Read-Only Memory (i.e., EPROM) or Flash memory, or an optical fiber. A machine-readable medium may also include a tangible medium upon which executable instructions are printed, as the logic may be electronically stored as an image or in another format (e.g., through an optical scan), then compiled, and/or interpreted or otherwise processed. The processed medium may then be stored in a computer and/or machine memory.
The systems may include additional or different logic and may be implemented in many ways. A controller may be implemented as a microprocessor, microcontroller, application specific integrated circuit (ASIC), discrete logic, or a combination of other types of circuits or logic. Similarly, memories may be DRAM, SRAM, Flash, or other types of memory. Parameters (e.g., conditions and thresholds), and other data structures may be separately stored and managed, may be incorporated into a single memory or database, or may be logically and physically organized in many different ways. Programs and instruction sets may be parts of a single program, separate programs, or distributed across several memories and processors. The systems may be included in a wide variety of electronic devices, including a cellular or wireless phone, a headset, a hands-free set, a speakerphone, communication interface, or an infotainment system.
While various embodiments of the invention have been described, it will be apparent to those of ordinary skill in the art that many more embodiments and implementations are possible within the scope of the invention. Accordingly, the invention is not to be restricted except in light of the attached claims and their equivalents.
| Number | Date | Country | Kind |
|---|---|---|---|
| 06009468 | May 2006 | EP | regional |
| Number | Name | Date | Kind |
|---|---|---|---|
| 6049607 | Marash et al. | Apr 2000 | A |
| 6421377 | Langberg et al. | Jul 2002 | B1 |
| 6442275 | Diethorn | Aug 2002 | B1 |
| 6510225 | Robertson et al. | Jan 2003 | B1 |
| 6738480 | Berthault et al. | May 2004 | B1 |
| 6839426 | Kamoi et al. | Jan 2005 | B1 |
| 6895095 | Thomas | May 2005 | B1 |
| 7068798 | Hugas et al. | Jun 2006 | B2 |
| 20020176585 | Egelmeers et al. | Nov 2002 | A1 |
| 20030021389 | Hirai et al. | Jan 2003 | A1 |
| 20030091182 | Marchok et al. | May 2003 | A1 |
| 20030185402 | Benesty et al. | Oct 2003 | A1 |
| 20030235312 | Pessoa et al. | Dec 2003 | A1 |
| 20040018860 | Hoshuyama | Jan 2004 | A1 |
| 20040125942 | Beaucoup et al. | Jul 2004 | A1 |
| 20050213747 | Popovich et al. | Sep 2005 | A1 |
| 20060018459 | McCree | Jan 2006 | A1 |
| 20060062380 | Kim et al. | Mar 2006 | A1 |
| 20060067518 | Klinke et al. | Mar 2006 | A1 |
| 20060233353 | Beaucoup et al. | Oct 2006 | A1 |
| 20070093714 | Beaucoup | Apr 2007 | A1 |
| Number | Date | Country |
|---|---|---|
| 1 404 147 | Mar 2004 | EP |
| 1 406 397 | Apr 2004 | EP |
| 09-307625 | Nov 1997 | JP |
| 10-023172 | Jan 1998 | JP |
| 2003-506924 | Feb 2003 | JP |
| 2003-264483 | Sep 2003 | JP |
| 2003-264483 | Sep 2003 | JP |
| WO 9317510 | Sep 1993 | WO |
| 0110102 | Feb 2001 | WO |
| Number | Date | Country | |
|---|---|---|---|
| 20080031467 A1 | Feb 2008 | US |