ML algorithms for information retrieval from html pages

ML algorithms for information retrieval from html pages

Manage alerts

Loading saved threads...

jottbe · External communityPost link
External question — Cross Validated Stack Exchange Author: jottbe Original post: https://stats.stackexchange.com/questions/663030 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. I wonder if there is a ML algorithm suitable for doing things like extracting information from html or other tagged data documents. In python one would usually create a script using a library like beatiful soup, but this requires knowlege about the structure of the document which often needs to be gained by reverse engineering. I wonder if there is a ML algorithm to relief us from this burden by learning from examples. E.g. if you extract data from pages with the same structure many times, you could use previously scraped html with the desired output as labelled training data. What could be an approach for this? I thought about RNNs but I have my doubts that it is the most efficient way (in terms of learning speed and energy consumption), because they are usually large, the sequence might be long, so learning to extract the right part might take a lot of resources and the architecture changes dramatically if the token set changes and thus probably needs a complete retraining. So is there something else? Ideally the logic should be transferrable to PDF scraping as well (e.g. after applying a customized tokenizer). The kind of documents I want to process are web pages containing information about stock prices. So think of a page regarding the shares of a company with tables and label text pairs or tables with infos like prices, eps figures, balance sheet infos. On the PDF side, a similar problem is extracting infos from papers for annual share holder meetings, or bank papers for trading security papers etc.
Quote
Report

Post Reply

Checking account access…