KnowledgeHub
Questions
Tags
Users
Search
Alex Rivera
|
Logout
Edit Question
Title
Body
I am crawling a site which may contain a lot of start_urls , like: http://www.a.com/list_1_2_3.htm I want to populate start_urls like [list_\d+_\d+_\d+\.htm] , and extract items from URLs like [node_\d+\.htm] during crawling. Can I use CrawlSpider to realize this function? And how can I generate the start_urls dynamically in crawling?
Tags (comma-separated)
Save Edits
Cancel