%0 Journal Article
%T Word Segment Based on Suffix Array
基于后缀数组的分词技术
%A REN Xue-Li
%A DAI Yu-Biao
%A
任雪利
%A 代余彪
%J 计算机系统应用
%D 2010
%I
%X Chinese word segmentation technology is the basis of machine translation, classification, search engines, as well as information retrieval. But the Internet emerging new words have seriously affected the performance of word segmentation. To improve the recognition rate of new words, suffix array is used in this paper, and the number of length of common prefix is calculated. The candidates on their words are filtered out by the threshold. Experimental results show that the new word recognition method has advantages.
%K suffix array
%K word segment
%K LCP
后缀数组
%K 分词
%K 公共前缀长度
%U http://www.alljournals.cn/get_abstract_url.aspx?pcid=5B3AB970F71A803DEACDC0559115BFCF0A068CD97DD29835&cid=8240383F08CE46C8B05036380D75B607&jid=D4F6864C950C88FFCE5B6C948A639E39&aid=BB1A22803481BD634195512F3BCC1D80&yid=140ECF96957D60B2&vid=2A8D03AD8076A2E3&iid=5D311CA918CA9A03&sid=F8035C8B7D8A4264&eid=D6354F61445E9456&journal_id=1003-3254&journal_name=计算机系统应用&referenced_num=0&reference_num=5