Alex Rivera | Logout

scraping the file with html saved in local system

Asked 2012-06-05T10:12:53.740
31

For example i had a site "www.example.com" Actually i want to scrape the html of this site by saving on to local system. so for testing i saved that page on my desktop as example.html

Now i had written the spider code for this as below

class ExampleSpider(BaseSpider):
   name = "example"
   start_urls = ["example.html"]

   def parse(self, response):
       print response
       hxs = HtmlXPathSelector(response)

But when i run the above code i am getting this error as below

ValueError: Missing scheme in request url: example.html

Finally my intension is to scrape the example.html file that consists of www.example.com html code saved in my local system

Can any one suggest me on how to assign that example.html file in start_urls

Thanks in advance

Edit
Report

1 Answer

14

You can use the HTTPCacheMiddleware, which will give you the ability to to a spider run from cache. The documentation for the HTTPCacheMiddleware settings is located here.

Basically, adding the following settings to your settings.py will make it work:

HTTPCACHE_ENABLED = True
HTTPCACHE_EXPIRATION_SECS = 0 # Set to 0 to never expire

This however requires to do an initial spider run from the web to populate the cache.

answered 2012-06-05T12:27:23.410

Your Answer