%0 Journal Article
%T Examination of Extraction Rules in Web Data Extraction
%A Erdin？ Uzun
%J -
%D 2018
%X Extracting the desired data from the web page is important issue for applications in the fields of data mining and information retrieval. DOM-based methods or regular expressions can be used to extract data from a web page. For this extraction process, multiple extraction rules can be prepared for both DOM-based methods and regular expressions. In this study, the effectiveness of obtaining more than one data with extraction rules is investigated. As a data set, fifteen websites including in the fields of news, film and shopping have been selected. Extraction rule files have been created for data extraction with different extraction techniques for these websites. Web sites are mainly focused on repetitive data such as reviews. Experiments have shown that regular expressions, the creation process is more laborious and time consuming, give better results than DOM-based methods. Among the DOM-based methods, the lxml parser library provided the best results as expected. Experiments indicate that the extraction rules prepared by a developer affect the extraction time. As a result, it is possible to extract the desired data much faster in web pages with the well-prepared regular expressions
%K ？？kar？m y？ntemleri
%K Web veri ？？kar？m？
%K DOM
%K Düzenli ifadeler
%U http://dergipark.org.tr/ejeas/issue/41931/486132