Journal of East China Normal University(Natural Sc ›› 2010, Vol. 2010 ›› Issue (5): 96-102.
• Article • Previous Articles Next Articles
JING Han-xing, CHEN Shao-hong, YU Kun
Received:
Revised:
Online:
Published:
Contact:
Abstract: This paper proposed a new tree alignment algorithm for determining the optimal matching structure of the input web pages, in order to extract web data automatically. Based on the alignment, the trees were merged into one union tree whose nodes record statistical information obtained from multiple web pages. The algorithm detects repeating patterns on the union tree, and a wrapper built on the most probable content block and the repeating patterns extracts data from web pages. Experimental results showed that the proposed algorithm achieves high extraction accuracy and has steady performance.
Key words: wrapper, tree alignment, data extraction, wrapper, tree alignment
CLC Number:
TP311.13
JING Han-xing;CHEN Shao-hong;YU Kun. Automatic web data extraction based on tree alignment[J]. Journal of East China Normal University(Natural Sc, 2010, 2010(5): 96-102.
0 / / Recommend
Add to citation manager EndNote|Reference Manager|ProCite|BibTeX|RefWorks
URL: https://xblk.ecnu.edu.cn/EN/
https://xblk.ecnu.edu.cn/EN/Y2010/V2010/I5/96