引用本文:刘挺,马金山,李生.基于词汇支配度的汉语依存分析模型.软件学报,2006,17(9):1876-1883
【打印本页】   【下载PDF全文】   查看/发表评论  【EndNote】   【RefMan】   【BibTex】
←前一篇|后一篇→ 过刊浏览    高级检索
本文已被:浏览 5142次   下载 6813 本文二维码信息
码上扫一扫!
分享到: 微信 更多
基于词汇支配度的汉语依存分析模型
刘挺1, 马金山1, 李生1
哈尔滨工业大学,信息检索研究室,黑龙江,哈尔滨,150001
摘要:
如何应用句法结构和词汇化是句法分析建模所面临的两个主要问题,汉语依存分析对这两方面做了初步的探索.首先通过对大规模依存树库的统计学习,获取其中的词汇依存信息,建立了一个词汇化的概率分析模型.然后引入词汇支配度的概念,以充分利用了句子中的结构信息.词汇化方法有效地弥补了以前工作中词性信息的粒度过粗问题.同时,词汇支配度增强了对句法结构的识别,有效地避免了非法结构的生成.在4 000句的测试集上,依存分析获得了约74%的正确率.
关键词:  依存语法  句法分析  支配度  动态规划
DOI:
分类号:
基金项目:Supported by the Key Project of National Natural Science Foundation of China under Grant No.60435020 (国家自然科学基金重点项目); the National Natural Science Foundation of China under Grant Nos.60575042, 60503072 (国家自然科学基金)
Chinese Dependency Parsing Model Based on Lexical Governing Degree
LIU Ting,MA Jin-Shan,LI Sheng
Abstract:
Use of structural information and lexicalization are two of the main challenges facing syntactic analysis, and they are investigated in this paper. First, the probabilities of lexical dependencies are obtained by training a large-scale dependency treebank and used to build the lexical model. Second, the governing degree of words is introduced to utilize the structure information. The lexical method overcomes the weakness of POS dependencies in the past work; meanwhile the governing degree of words is helpful to distinguish the syntactic structures so some ill-formed structures are avoided. Finally, the paper shows a good experimental result of around 74% accuracy on the test set that consists of 4000 sentences.
Key words:  dependency grammar  parsing  governing degree  dynamic programming

引用本文:
【打印本页】   【下载PDF全文】   查看/发表评论  【EndNote】   【RefMan】   【BibTex】
←前一篇|后一篇→ 过刊浏览    高级检索
本文已被:浏览次   下载  
分享到: 微信 更多
摘要:
关键词:  
DOI:
分类号:
基金项目:
Abstract:
Key words: