27
I am crawling a site which may contain a lot of start_urls, like:
http://www.a.com/list_1_2_3.htm
I want to populate start_urls like [list_\d+_\d+_\d+\.htm],
and extract items from URLs like [node_\d+\.htm] during crawling.
Can I use CrawlSpider to realize this function?
And how can I generate the start_urls dynamically in crawling?