Voice Recognition
Voice recognition verifies identity from the physical characteristics of a person's voice, and its ability to reuse existing telephone infrastructure has made it especially attractive for call-center and telephone-banking authentication.
Voice recognition is a technology that allows a user to use their voice as an input device. Voice recognition may be used to dictate text into a computer or to give commands to the computer (such as opening application programs, pulling down menus, or saving work).
Older voice recognition applications require each word to be separated by a distinct pause. This allows the machine to determine where one word ends and the next begins. These kinds of voice recognition applications are still used to navigate a computer's system and operate applications such as web browsers or spreadsheets.
Newer voice recognition applications allow a user to dictate text fluently into the computer. These newer applications can recognize speech at up to 160 words per minute. Applications that support continuous speech are generally designed to recognize text and format it, rather than control the computer system itself.
Voice recognition uses a neural net to "learn" to recognize a person's voice. As you speak, the voice recognition software remembers the way you say each word. This customization allows voice recognition to work even though everyone speaks with varying accents and inflection.
In addition to learning how you pronounce words, voice recognition also uses grammatical context and frequency of use to predict the word you wish to input. These powerful statistical tools allow the software to narrow down the massive language database before you even speak the next word.
While the accuracy of voice recognition has improved over the past few years, some users still experience accuracy problems, either because of the way they speak or the nature of their voice.
Voice recognition technology uses the distinctive aspects of the voice to verify the identity of individuals. Voice recognition is occasionally confused with speech recognition, a technology that translates what a user is saying (a process unrelated to authentication). Voice recognition technology, by contrast, verifies the identity of the individual who is speaking. The two technologies are often bundled — speech recognition is used to translate the spoken word into an account number, and voice recognition verifies the vocal characteristics against those associated with that account.
Voice recognition can use any audio capture device, including mobile and landline telephones and PC microphones. The performance of voice recognition systems can vary according to the quality of the audio signal, as well as variation between enrollment and verification devices, so acquisition normally takes place on the device likely to be used for future verification.
During enrollment, an individual is prompted to select a passphrase or to repeat a sequence of numbers. The passphrase selected should be approximately 1-1.5 seconds long — very short passphrases lack enough identifying data, and long passwords have too much, both resulting in reduced accuracy. The individual is generally prompted to repeat the passphrase or number set a handful of times, making the enrollment process somewhat longer than for most other biometrics.
One of the challenges facing large-scale biometric deployments is the need to distribute new hardware to employees, customers and users. One strength of telephony-based voice recognition implementations is that they can avoid this problem, especially when implemented in call-center and account-access applications. Without additional hardware at the user end, voice recognition systems can be installed as a subroutine through which calls are routed before access to sensitive information is granted. The ability to use existing telephones means voice recognition vendors have hundreds of millions of authentication devices available for transactional use today.
Similarly, voice recognition can leverage existing account access and authentication processes, eliminating the need to introduce unwieldy or confusing authentication scenarios. Automated telephone systems using speech recognition are currently ubiquitous, due to the savings possible from reducing the number of employees needed to operate call centers. Voice recognition and speech recognition can function simultaneously on the same utterance, allowing the technologies to blend seamlessly. Voice recognition can function as a reliable authentication mechanism for automated telephone systems, adding security to automated telephone-based transactions in areas such as financial services and health care.
Though inconsistent with many users' perceptions, certain voice recognition technologies are highly resistant to imposter attacks — even more so than some fingerprint systems. While false non-matching can be a common problem, this resistance to false matching means voice recognition can be used to protect reasonably high-value transactions.
Since the technology has not traditionally been used in law enforcement or tracking applications, where it could be viewed as a "Big Brother" technology, there is less public fear that voice recognition data can be tracked across databases or used to monitor individual behavior. Thus, voice recognition largely avoids one of the biggest hurdles facing other biometric technologies: the perception of invasiveness.
Voice recognition is a strong solution for implementations in which vocal interaction is already present. It is not a strong solution when speech is introduced as a new process. Telephony is the primary growth area for voice recognition, and will likely remain by far the most common area of implementation for the technology. Telephony-based applications for voice recognition include account access for financial services, customer authentication for service calls, and challenge-response implementations for house arrest and probation-related authentication. These solutions route callers through enrollment and verification subroutines, using vendor-specific hardware and software integrated with an institution's existing infrastructure.
Voice recognition has also been implemented in physical access solutions for border crossings, although this is not the technology's ideal deployment environment.
Though revenues from the technology are relatively small today, voice recognition will draw substantially greater revenues through 2007. Most likely to be deployed in telephony-based environments (such as account access for financial services and customer authentication for service calls), voice recognition revenues are projected to grow from $12.2m in 2002 to $142.1m in 2007. Voice recognition revenues are expected to comprise approximately 4% of the entire biometric market.
Telephone banking is increasingly popular with customers, and will become increasingly attractive to banks and other financial institutions as they implement highly cost-effective automated speech recognition technology to handle routine transactions.
But the procedures for verifying customers over the telephone are unsatisfactory, both in terms of customer convenience and, increasingly, from a security point of view.
The Problem
The usual approach to verifying customers — proving that they are who they claim to be — is to use some sort of PIN or password. To avoid the customer having to say the password out loud, they are usually prompted for, say, the second and fourth letters in the password.
There are several problems with this approach:
- First, passwords and PINs are difficult to remember and unwieldy for customers to use in this manner.
- Second, it takes time — identification and verification of the caller is often the lengthiest component of a transaction, and this translates directly to the bottom line.
- Third, the security itself leaves a lot to be desired — many customers write down their passwords or reveal them to the operator (in extreme cases they may self-select the same PIN they use for ATM withdrawals). Many call centers prompt callers for additional "secret" items such as their mother's maiden name, but this only exacerbates the other two problems.
Technology now exists that enables individuals to be reliably, rapidly and cost-effectively verified based on the physical characteristics of their voice.
Several vendors now supply commercial voice verification technology. A good example is Nuance Communications, based in California, using essentially the same technology that underlies their speaker-independent speech recognition software. But in this case recognition is speaker-dependent — the customer is only allowed to use the system if their individual voiceprint matches their identity (normally established through an account number).
A new customer automatically enrolls in the system over the telephone by repeating about 10 four-digit numbers or reading a short piece of text. The software extracts from this a number of physical characteristics unique to that voice. In all subsequent transactions, the caller, once identified, is asked to repeat a couple of randomly generated PINs or, for example, city names (this is to prevent fraudsters from tape-recording a customer saying their password or PIN). If the voiceprint matches the one stored against the account number, the transaction proceeds; if not, the customer is referred to a supervisor.
Pilot tests of the technology are encouraging. A high accuracy of correct verification can be combined with a low probability of false rejection, which is suitable for most banking operations, and the whole procedure is faster, easier and much more cost-effective. Surprisingly, only a few kilobytes of storage are required for each voiceprint, and because the claimed identity of the customer is already established, a single comparison is all that is required, so verification is quite rapid (voice identification using the same technology is, of course, much slower, since the system must find a match out of many voiceprints).
Voice verification is particularly well suited to automated speech recognition dialogues, and a seamless combination of the two technologies is expected to rapidly become the norm for most simple telephone banking transactions.
Of course, voice verification is much less applicable to other delivery channels, such as branch banking or screen-based systems (although pilot systems have been built). Another intriguing approach to customer verification over the Internet is based on face recognition, using systems such as Passfaces.
The speaker-specific characteristics of speech are due to differences in the physiological and behavioral aspects of the human speech-production system. The main physiological aspect of the human speech production system is the vocal tract shape. The vocal tract is generally considered the speech-production organ above the vocal folds, consisting of: (i) the laryngeal pharynx (beneath the epiglottis), (ii) the oral pharynx (behind the tongue, between the epiglottis and velum), (iii) the oral cavity (forward of the velum and bounded by the lips, tongue and palate), (iv) the nasal pharynx (above the velum, rear end of the nasal cavity), and (v) the nasal cavity (above the palate and extending from the pharynx to the nostrils). The shaded area in figure 1 depicts the vocal tract.
The vocal tract modifies the spectral content of an acoustic wave as it passes through, thereby producing speech. It is therefore common in speaker verification systems to use features derived only from the vocal tract. To characterize the features of the vocal tract, the human speech production mechanism is represented as a discrete-time system of the form shown in figure 2.
The acoustic wave is produced when airflow from the lungs is carried by the trachea through the vocal folds. This source of excitation can be characterized as phonation, whispering, frication, compression, vibration, or a combination of these. Phonated excitation occurs when airflow is modulated by the vocal folds. Whispered excitation is produced by airflow rushing through a small triangular opening between the arytenoid cartilages at the rear of the nearly closed vocal folds. Frication excitation is produced by constrictions in the vocal tract. Compression excitation results from releasing a completely closed, pressurized vocal tract. Vibration excitation is caused by air forced through a closure other than the vocal folds, especially at the tongue. Speech produced by phonated excitation is called voiced; that produced by phonated excitation plus frication is called mixed voiced; and that produced by other types of excitation is called unvoiced.
The vocal tract can be represented in a parametric form as the transfer function H(z). To estimate the parameters of H(z) from the observed speech waveform, it is necessary to assume some form for H(z). Ideally, the transfer function would contain poles as well as zeros. However, if only the voiced regions of speech are used, an all-pole model for H(z) is sufficient. Furthermore, linear prediction analysis can efficiently estimate the parameters of an all-pole model. It can also be noted that the all-pole model is the minimum-phase part of the true model and has an identical magnitude spectrum, which contains the bulk of the speaker-dependent information.
This also underlines the text-dependent nature of vocal-tract models. Since the model is derived from the observed speech, it is dependent on the speech itself. Figure 3 illustrates the differences in the models for two speakers saying the same vowel.
LPC features were very popular in early speech-recognition and speaker-verification systems. However, comparing two LPC feature vectors requires computationally expensive similarity measures, such as the Itakura-Saito distance, making LPC features unsuitable for real-time systems. Furui suggested the use of the cepstrum — defined as the inverse Fourier transform of the logarithm of the magnitude spectrum — in speech-recognition applications. The cepstrum allows the similarity between two cepstral feature vectors to be computed as a simple Euclidean distance. Furthermore, Atal demonstrated that the cepstrum derived from LPC features produces the best performance in terms of FAR and FRR for a speaker verification system. Consequently, the LPC-derived cepstrum is used for the speaker verification system described here.
Using cepstral analysis as described in the previous section, an utterance may be represented as a sequence of feature vectors. Utterances spoken by the same person at different times result in similar, yet different, sequences of feature vectors. The purpose of voice modeling is to build a model that captures these variations in the extracted set of features. Two types of models have been used extensively in speaker verification and speech recognition systems: stochastic models and template models. The stochastic model treats the speech-production process as a parametric random process and assumes the parameters of the underlying stochastic process can be estimated in a precise, well-defined manner. The template model attempts to model the speech-production process in a non-parametric way, by retaining a number of feature-vector sequences derived from multiple utterances of the same word by the same person. Template models dominated early work in speaker verification and speech recognition, because the template model is intuitively more reasonable. However, recent work on stochastic models has demonstrated that these models are more flexible and hence allow better modeling of the speech-production process. A very popular stochastic model for the speech-production process is the Hidden Markov Model (HMM). HMMs extend conventional Markov models, in that the observations are a probabilistic function of the state — i.e., the model is a doubly embedded stochastic process where the underlying stochastic process is not directly observable (it is hidden). The HMM can only be viewed through another set of stochastic processes that produce the sequence of observations. Thus, the HMM is a finite-state machine, where a probability density function p(x | s_i) is associated with each state s_i. The states are connected by a transition network, where the state transition probabilities are a_ij = p(s_i | s_j). A fully connected three-state HMM is depicted in figure 4.
For speech signals, another type of HMM, called a left-right model or Bakis model, is found to be more useful. A left-right model has the property that as time increases, the state index increases (or stays the same) — that is, the system states proceed from left to right. Since the properties of a speech signal change over time in a successive manner, this model is very well suited to modeling the speech-production process.
The pattern-matching process compares a given set of input feature vectors against the speaker model for the claimed identity and computes a matching score. For the Hidden Markov Models discussed above, the matching score is the probability that a given set of feature vectors was generated by the model.
Legacy links (no longer active) (4)
- http://www.newscientist.com/news/news.jsp?id=ns99993921
- http://www.speech.cs.cmu.edu/comp.speech/SpeechLinks.html
- http://www.authentify.com/solutions/demos/index.html
- http://www.cybertown.com/qvoice/startrek.html