KnowledgeHub
Questions
Tags
Users
Search
Alex Rivera
|
Logout
Edit Question
Title
Body
How do I iterate through sites with Scrapy? I'd like to extract the body of all sites that match http://www.saylor.org/site/syllabus.php?cid=NUMBER , where NUMBER is 1 through 400 or so. I've written this spider: from scrapy.contrib.spiders import CrawlSpider, Rule from scrapy.contrib.linkextractors.sgml import SgmlLinkExtractor from scrapy.selector import HtmlXPathSelector from syllabi.items import SyllabiItem class SyllabiSpider(CrawlSpider): name = 'saylor' allowed_domains = ['saylor.org'] start_urls = ['http://www.saylor.org/site/syllabus.php?cid='] rules = [Rule(SgmlLinkExtractor(allow=['\d+']), 'parse_syllabi')] def parse_syllabi(self, response): x = HtmlXPathSelector(response) syllabi = SyllabiItem() syllabi['url'] = response.url syllabi['body'] = x.select("/html/body/text()").extract() return syllabi But it doesn't work. I understand it's looking for links in that start_url, which is not really what I want it to do. I want to iterate through the sites. Make sense? Thanks for the help.
Tags (comma-separated)
Save Edits
Cancel