%0 Journal Article %T Extraction Model Based on Web Format Information Quantity in Blog Post and Comment Extraction
基于网页格式信息量的博客文章和评论抽取模型 %A CAO Dong-Lin %A LIAO Xiang-Wen %A XU Hong-Bo %A BAI Shuo %A
曹冬林 %A 廖祥文 %A 许洪波 %A 白硕 %J 软件学报 %D 2009 %I %X Based on the information theory, this paper presents a model based on Web format information quantity in blog information extraction. First, the vision information in blog Web page and the effective text information are combined to locate the main text which represents the theme of the blog Web page. Second, the format information of blog Web page is used to calculate the information quantity of each block and the minimal separating information quantity of separate position is used to detect the boundary of posts and comments in the main text. This model is language insensitive and can be used in a lot of blogs which are written in different natural languages. Experimental results show that this method achieves high precision in locating main text and separating the post and comment. %K blog information extraction %K minimal main text subtree %K effective information ratio %K Web format information %K vision information %K information quantity of separate position
博客信息抽取 %K 最小正文子树 %K 有效信息率 %K 网页格式信息 %K 视觉信息 %K 切分位置信息量 %U http://www.alljournals.cn/get_abstract_url.aspx?pcid=5B3AB970F71A803DEACDC0559115BFCF0A068CD97DD29835&cid=8240383F08CE46C8B05036380D75B607&jid=7735F413D429542E610B3D6AC0D5EC59&aid=1B85DAA49F8A84C775F34DF8ADFE4204&yid=DE12191FBD62783C&vid=A04140E723CB732E&iid=94C357A881DFC066&sid=7CF64E95CEC38520&eid=0E37D9F9BC838B8F&journal_id=1000-9825&journal_name=软件学报&referenced_num=3&reference_num=17