KnowledgeHub
Questions
Tags
Users
Search
Alex Rivera
|
Logout
Edit Question
Title
Body
I've been googling this all day with out finding the answer, so apologies in advance if this is already answered. I'm trying to get all visible text from a large number of different websites. The reason is that I want to process the text to eventually categorize the websites. After a couple of days of research, I decided that Selenium was my best chance. I've found a way to grab all the text, with Selenium, unfortunately the same text is being grabbed multiple times: from selenium import webdriver import codecs filen = codecs.open('outoput.txt', encoding='utf-8', mode='w+') driver = webdriver.Firefox() driver.get("http://www.examplepage.com") allelements = driver.find_elements_by_xpath("//*") ferdigtxt = [] for i in allelements: if i.text in ferdigtxt: pass else: ferdigtxt.append(i.text) filen.writelines(i.text) filen.close() driver.quit() The if condition inside the for loop is an attempt at eliminating the problem of fetching the same text multiple times - it does not however, only work as planned on some webpages. (it also makes the script A LOT slower) I'm guessing the reason for my problem is that - when asking for the inner text of an element - I also get the inner text of the elements nested inside the element in question. Is there any way around this? Is there some sort of master element I grab the inner text of? Or a completely different way that would enable me to reach my goal? Any help would be greatly appreciated as I'm out of ideas for this one. Edit: the reason I used Selenium and not Mechanize and Beautiful Soup is because I wanted JavaScript tendered text
Tags (comma-separated)
Save Edits
Cancel