###
DOI:
Journal of Software:2000.11(7):971-978

基于模糊训练集的领域相关统计语言模型
陈浪舟,黄泰翼
(中国科学院自动化研究所,北京,100080)
Domain Dependent Language Model Based on Fuzzy Training Subset
CHEN Lang-zhou,HUANG Tai-yi
()
Abstract
Chart / table
Reference
Similar Articles
Article :Browse 2580   Download 2648
Received:February 08, 1999    Revised:June 17, 1999
> 中文摘要: 统计语言模型在语音识别中具有重要作用.对于特定领域的识别系统来说,主题相关的语言模型效果远远优于领域无关的语言模型.传统方法在建立领域相关的语言模型时通常会遇到两个问题,一个是领域相关的语料不像普通语料那样充分,另一个是一篇特定的文章往往与好几个主题相关,而在模型的训练过程中,这种现象没有得到充分的考虑.为解决这两个问题,提出了一种新的领域相关训练语料的组织方法——基于模糊训练集的组织方法,领域相关的语言模型就建立在模糊训练集的基础上.同时,为了增强模型的预测能力,将自组织学习引入到模型的训练过程中,取得了良好的效果.
Abstract:Statistical language model is very important to speech recognition. To a system of special topic, domain dependent language model is much better than the general model. There are two problems in traditional method. (1) The corpus of special topic is not large enough as general corpus. (2) An article is always related to more than one topic, but these phenomena have not been considered during the process of model training. In this paper, the authors try to solve these two problems. They present a new method to organize the corpus——the method based on fuzzy training subset. And the training of domain dependent models is based on these fuzzy subsets. At the same time, self organized learning has been introduced in training process to improve the models' prediction ability. It can improve the performance of models evidently.
文章编号:     中图分类号:    文献标志码:
基金项目:本文研究得到国家自然科学基金(No.69835003)资助. 本文研究得到国家自然科学基金(No.69835003)资助.
Foundation items:
Reference text:

陈浪舟,黄泰翼.基于模糊训练集的领域相关统计语言模型.软件学报,2000,11(7):971-978

CHEN Lang-zhou,HUANG Tai-yi.Domain Dependent Language Model Based on Fuzzy Training Subset.Journal of Software,2000,11(7):971-978