Dynamic switching between local and remote speech rendering
    1.
    发明授权
    Dynamic switching between local and remote speech rendering 有权
    本地和远程语音呈现之间的动态切换

    公开(公告)号:US08024194B2

    公开(公告)日:2011-09-20

    申请号:US11007830

    申请日:2004-12-08

    IPC分类号: G10L21/00 G10L13/00 G10L15/00

    摘要: A multimodal browser for rendering a multimodal document on an end system defining a host can include a visual browser component for rendering visual content, if any, of the multimodal document, and a voice browser component for rendering voice-based content, if any, of the multimodal document. The voice browser component can determine which of a plurality of speech processing configuration is used by the host in rendering the voice-based content. The determination can be based upon the resources of the host running the application. The determination also can be based upon a processing instruction contained in the application.

    摘要翻译: 用于在定义主机的终端系统上呈现多模式文档的多模式浏览器可以包括用于呈现多模式文档的视觉内容(如果有的话)的视觉浏览器组件,以及用于呈现基于语音的内容(如果有的话)的语音浏览器组件 多模式文件。 语音浏览器组件可以确定主机在渲染基于语音的内容中使用多个语音处理配置中的哪一个。 确定可以基于运行应用程序的主机的资源。 该确定还可以基于应用中包含的处理指令。

    Pausing a VoiceXML dialog of a multimodal application
    2.
    发明授权
    Pausing a VoiceXML dialog of a multimodal application 有权
    暂停多模式应用程序的VoiceXML对话框

    公开(公告)号:US08713542B2

    公开(公告)日:2014-04-29

    申请号:US11679236

    申请日:2007-02-27

    摘要: Pausing a VoiceXML dialog of a multimodal application, including generating by the multimodal application a pause event; responsive to the pause event, temporarily pausing the dialogue by the VoiceXML interpreter; generating by the multimodal application a resume event; and responsive to the resume event, resuming the dialog. Embodiments are implemented with the multimodal application operating on a multimodal device supporting multiple modes of interaction including a voice mode and one or more non-voice modes, the multimodal application is operatively coupled to a VoiceXML interpreter, and the VoiceXML interpreter is interpreting the VoiceXML dialog to be paused.

    摘要翻译: 暂停多模式应用程序的VoiceXML对话框,包括由多模态应用程序生成暂停事件; 响应暂停事件,VoiceXML解释器临时暂停对话; 由多模式应用程序生成一个简历事件; 并响应resume事件,恢复对话。 实施例是通过在多模式设备上操作的多模式应用来实现的,该多模式设备支持包括语音模式和一种或多种非语音模式的多种交互模式,多模式应用可操作地耦合到VoiceXML解释器,并且VoiceXML解释器正在解释VoiceXML对话 暂停

    Enabling speech recognition grammars in web page frames
    3.
    发明授权
    Enabling speech recognition grammars in web page frames 有权
    在网页框架中启用语音识别语法

    公开(公告)号:US08073692B2

    公开(公告)日:2011-12-06

    申请号:US12917741

    申请日:2010-11-02

    IPC分类号: G10L15/22 G06F17/20

    摘要: Enabling grammars in web page frames, including receiving, in a multimodal application on a multimodal device, a frameset document, where the frameset document includes markup defining web page frames; obtaining by the multimodal application content documents for display in each of the web page frames, where the content documents include navigable markup elements; generating by the multimodal application, for each navigable markup element in each content document, a segment of markup defining a speech recognition grammar, including inserting in each such grammar markup identifying content to be displayed when words in the grammar are matched and markup identifying a frame where the content is to be displayed; and enabling by the multimodal application all the generated grammars for speech recognition.

    摘要翻译: 在网页框架中启用语法,包括在多模式设备上的多模式应用程序中接收框架集文档,其中框架集文档包括定义网页框架的标记; 通过多模式应用程序内容文档获取以在每个网页帧中显示,其中内容文档包括可导航标记元素; 由多模式应用为每个内容文档中的每个可导航标记元素生成定义语音识别语法的标记段,包括在每个这样的语法标记中插入标识要在语法中的词匹配时要显示的内容,并且标识标识帧 要显示的内容; 并通过多模式应用程序实现所有生成的语法用于语音识别。

    Web service support for a multimodal client processing a multimodal application
    4.
    发明授权
    Web service support for a multimodal client processing a multimodal application 失效
    处理多模式应用程序的多模式客户端的Web服务支持

    公开(公告)号:US08788620B2

    公开(公告)日:2014-07-22

    申请号:US11696230

    申请日:2007-04-04

    摘要: Web service support for a multimodal client processing a multimodal application, the multimodal client providing an execution environment for the application and operating on a multimodal device supporting multiple modes of user interaction including a voice mode and one or more non-voice modes, the application stored on an application server, includes: receiving, by the server, an application request from the client that specifies the application and device characteristics; determining, by a multimodal adapter of the server, modality requirements for the application; selecting, by the adapter, a modality web service in dependence upon the modality requirements and the characteristics for the device; determining, by the adapter, whether the device supports VoIP in dependence upon the characteristics; providing, by the server, the application to the client; and providing, by the adapter to the client in dependence upon whether the device supports VoIP, access to the modality web service for processing the application.

    摘要翻译: 处理多模式应用程序的多模式客户端的Web服务支持,多模式客户端为应用程序提供执行环境并在支持包括语音模式和一种或多种非语音模式的多种用户交互模式的多模式设备上运行,应用程序存储 在应用服务器上,包括:由服务器接收来自客户端的指定应用和设备特征的应用请求; 通过服务器的多模式适配器确定应用的模态要求; 根据模式要求和设备的特性,由适配器选择模态web服务; 由所述适配器确定所述设备是否根据所述特征支持VoIP; 由服务器将应用程序提供给客户端; 以及根据所述设备是否支持VoIP,通过所述适配器向所述客户端提供对所述模态网络服务的访问以处理所述应用。

    Invoking tapered prompts in a multimodal application
    5.
    发明授权
    Invoking tapered prompts in a multimodal application 有权
    在多模式应用程序中调用渐变提示

    公开(公告)号:US08150698B2

    公开(公告)日:2012-04-03

    申请号:US11678920

    申请日:2007-02-26

    IPC分类号: G10L21/00

    摘要: Methods, apparatus, and computer program products are described for invoking tapered prompts in a multimodal application implemented with a multimodal browser and a multimodal application operating on a multimodal device supporting multiple modes of user interaction with the multimodal application, the modes of user interaction including a voice mode and one or more non-voice modes. Embodiments include identifying, by a multimodal browser, a prompt element in a multimodal application; identifying, by the multimodal browser, one or more attributes associated with the prompt element; and playing a speech prompt according to the one or more attributes associated with the prompt element.

    摘要翻译: 描述了用于在多模式浏览器和多模式应用程序实现的多模式应用程序中调用渐变提示的方法,装置和计算机程序产品,该多模式应用程序在多模式设备上运行,该多模式应用程序支持与多模式应用程序的多种用户交互模式,用户交互模式包括 语音模式和一个或多个非语音模式。 实施例包括通过多模式浏览器识别多模式应用中的提示元素; 通过多模式浏览器识别与提示元素相关联的一个或多个属性; 以及根据与所述提示元素相关联的一个或多个属性播放语音提示。

    VOIP barge-in support for half-duplex DSR client on a full-duplex network
    6.
    发明授权
    VOIP barge-in support for half-duplex DSR client on a full-duplex network 有权
    VOIP在全双工网络上支持半双工DSR客户端

    公开(公告)号:US07848314B2

    公开(公告)日:2010-12-07

    申请号:US11382575

    申请日:2006-05-10

    IPC分类号: H04L12/66

    摘要: Providing VOIP barge-in support for a half-duplex DSR client on a full-duplex network by buffering, in a half-duplex DSR client, input audio from the full-duplex network; playing, through the half-duplex DSR client, the buffered input audio; pausing, during voice activity on the half-duplex DSR client, the playing of the buffered input audio; sending, during voice activity on the half-duplex DSR client, speech for recognition through the full-duplex network to a voice server; receiving in the half-duplex DSR client through the full-duplex network from the voice server notification of speech recognition, the notification bearing a time stamp; and, responsive to receiving the notification, resuming the playing of the buffered input audio, including playing only buffered VOIP audio data bearing time stamps later than the time stamp of the recognition notification.

    摘要翻译: 通过在半双工DSR客户端中通过缓冲从全双工网络输入音频,在全双工网络上提供VOIP支持半双工DSR客户端的支持; 通过半双工DSR客户端播放缓冲输入音频; 在半双工DSR客户端的语音活动期间暂停播放缓冲输入音频; 在半双工DSR客户端的语音活动期间,通过全双工网络将话音发送到语音服务器; 通过语音服务器通过全双工网络在半双工DSR客户端中接收语音识别通知,具有时间戳; 并且响应于接收到通知,恢复播放缓冲的输入音频,包括播放比识别通知的时间戳更晚的缓存的支持时间戳的VOIP音频数据。

    Speech-enabled content navigation and control of a distributed multimodal browser
    7.
    发明授权
    Speech-enabled content navigation and control of a distributed multimodal browser 有权
    语音启用的内容导航和分布式多模式浏览器的控制

    公开(公告)号:US08862475B2

    公开(公告)日:2014-10-14

    申请号:US11734445

    申请日:2007-04-12

    摘要: Speech-enabled content navigation and control of a distributed multimodal browser is disclosed, the browser providing an execution environment for a multimodal application, the browser including a graphical user agent (‘GUA’) and a voice user agent (‘VUA’), the GUA operating on a multimodal device, the VUA operating on a voice server, that includes: transmitting, by the GUA, a link message to the VUA, the link message specifying voice commands that control the browser and an event corresponding to each voice command; receiving, by the GUA, a voice utterance from a user, the voice utterance specifying a particular voice command; transmitting, by the GUA, the voice utterance to the VUA for speech recognition by the VUA; receiving, by the GUA, an event message from the VUA, the event message specifying a particular event corresponding to the particular voice command; and controlling, by the GUA, the browser in dependence upon the particular event.

    摘要翻译: 公开了一种分布式多模式浏览器的语音启用内容导航和控制,浏览器为多模式应用提供执行环境,浏览器包括图形用户代理(“GUA”)和语音用户代理(“VUA”), GUA在多模式设备上操作,VUA在语音服务器上操作,其包括:由GUA向VUA发送链接消息,指定控制浏览器的语音命令的链接消息和与每个语音命令相对应的事件; 由GUA接收来自用户的语音发音,指定特定语音命令的语音话语; 通过GUA向VUA发送语音识别语音识别语音; 由GUA接收来自VUA的事件消息,事件消息指定与特定语音命令对应的特定事件; 并由GUA根据特定事件控制浏览器。

    Automatic speech recognition with a selection list
    8.
    发明授权
    Automatic speech recognition with a selection list 有权
    具有选择列表的自动语音识别

    公开(公告)号:US08612230B2

    公开(公告)日:2013-12-17

    申请号:US11619209

    申请日:2007-01-03

    IPC分类号: G10L21/00

    摘要: Methods, apparatus, and computer program products are described for automatic speech recognition (‘ASR’) that include accepting by the multimodal application speech input and visual input for selecting or deselecting items in a selection list, the speech input enabled by a speech recognition grammar; providing, from the multimodal application to the grammar interpreter, the speech input and the speech recognition grammar; receiving, by the multimodal application from the grammar interpreter, interpretation results including matched words from the grammar that correspond to items in the selection list and a semantic interpretation token that specifies whether to select or deselect items in the selection list; and determining, by the multimodal application in dependence upon the value of the semantic interpretation token, whether to select or deselect items in the selection list that correspond to the matched words.

    摘要翻译: 描述用于自动语音识别(“ASR”)的方法,装置和计算机程序产品,其包括通过多模式应用语音输入的接受和用于在选择列表中选择或取消选择项目的可视输入,由语音识别语法启用的语音输入 ; 从多模式应用程序提供语法解释器,语音输入和语音识别语法; 通过多模式应用从语法解释器接收包括对应于选择列表中的项目的语法的匹配词的解释结果以及指定是否选择或取消选择列表中的项目的语义解释令牌; 以及根据所述语义解释令牌的值由所述多模式应用程序确定是否选择或取消选择列表中对应于所述匹配词的项目。

    Altering behavior of a multimodal application based on location
    9.
    发明授权
    Altering behavior of a multimodal application based on location 有权
    基于位置改变多模式应用程序的行为

    公开(公告)号:US09208783B2

    公开(公告)日:2015-12-08

    申请号:US11679301

    申请日:2007-02-27

    CPC分类号: G10L15/22 G10L15/24

    摘要: Methods, apparatus, and products are disclosed for altering behavior of a multimodal application based on location. The multimodal application operates on a multimodal device supporting multiple modes of user interaction with the multimodal application, including a voice mode and one or more non-voice modes. The voice mode of user interaction with the multimodal application is supported by a voice interpreter. Altering behavior of a multimodal application based on location includes: receiving a location change notification in the voice interpreter from a device location manager, the device location manager operatively coupled to a position detection component of the multimodal device, the location change notification specifying a current location of the multimodal device; updating, by the voice interpreter, location-based environment parameters for the voice interpreter in dependence upon the current location of the multimodal device; and interpreting, by the voice interpreter, the multimodal application in dependence upon the location-based environment parameters.

    摘要翻译: 公开了基于位置改变多模式应用的行为的方法,装置和产品。 多模式应用程序在多模式设备上运行,支持与多模式应用程序的多种用户交互模式,包括语音模式和一种或多种非语音模式。 与多模式应用程序的用户交互的语音模式由语音解释器支持。 基于位置改变多模式应用的行为包括:从设备位置管理器在语音解释器中接收位置改变通知,该设备位置管理器可操作地耦合到多模态设备的位置检测组件,位置变化通知指定当前位置 的多模式设备; 语音解释器根据多模式设备的当前位置更新语音解释器的基于位置的环境参数; 并且由语音解释器根据基于位置的环境参数来解释多模式应用。

    Invoking tapered prompts in a multimodal application
    10.
    发明授权
    Invoking tapered prompts in a multimodal application 有权
    在多模式应用程序中调用渐变提示

    公开(公告)号:US08744861B2

    公开(公告)日:2014-06-03

    申请号:US13410103

    申请日:2012-03-01

    IPC分类号: G10L21/00 G10L25/00

    摘要: Methods, apparatus, and computer program products are described for invoking tapered prompts in a multimodal application implemented with a multimodal browser and a multimodal application operating on a multimodal device supporting multiple modes of user interaction with the multimodal application, the modes of user interaction including a voice mode and one or more non-voice modes. Embodiments include identifying, by a multimodal browser, a prompt element in a multimodal application; identifying, by the multimodal browser, one or more attributes associated with the prompt element; and playing a speech prompt according to the one or more attributes associated with the prompt element.

    摘要翻译: 描述了用于在多模式浏览器和多模式应用程序实现的多模式应用程序中调用渐变提示的方法,装置和计算机程序产品,该多模式应用程序在多模式设备上运行,该多模式应用程序支持与多模式应用程序的多种用户交互模式,用户交互模式包括 语音模式和一个或多个非语音模式。 实施例包括通过多模式浏览器识别多模式应用中的提示元素; 通过多模式浏览器识别与提示元素相关联的一个或多个属性; 以及根据与所述提示元素相关联的一个或多个属性播放语音提示。