| 摘要: |
| 有导词义消歧机器学习方法的引入虽然使词义消歧取得了长足的进步,但由于需要大量人力进行词义标注,使其难以适用于大规模词义消歧任务.针对这一问题,提出了一种避免人工词义标注巨大工作量的无导学习方法.在仅需义项词语知识库的支持下,将待消歧多义词与义项词语映射到向量空间中,基于k-NN(k=1)方法,计算二者相似度来实现词义消歧任务.在对10个典型多义词进行词义消歧的测试实验中,采用该方法取得了平均正确率为83.13%的消歧结果. |
| 关键词: 词义消歧 无导方法 义项词语 上下文位置权重计算 向量空间模型 |
| DOI: |
| 分类号: |
| 基金项目:国家自然科学基金资助项目(69773008);国家863高科技发展计划资助项目(863-306-2D02-01-3);国家重点基础研究发展规划973资助项目(G1998030510) |
|
| An Unsuptervised Approach to Word Sense Disambiguation Based on Sense-Words in Vector Space Model |
|
LU Song,BAI Shuo,HUANG Xiong
|
| Abstract: |
| WSD (word sense disambiguation) based on supervised machine learning made a great progress, but it is hard to deal with large-scale WSD because of its 慴ig?labor cost. An unsupervised WSD method is provided in this paper to solve this problem. Only under the knowledge database of sense-words, this method formulates the sense-words and polysemous words in vector space, and based on k-NN (k=1) it calculates the similarity between them to disambiguate polysemous words. The average accuracy is 83.13% for 10 polysemous words in open test by this method. |
| Key words: word sense disambiguation unsupervised approach sense-word weight of context position vector space model |