KnowledgeHub
Questions
Tags
Users
Search
Alex Rivera
|
Logout
Edit Question
Title
Body
The crawler needs to have an extendable architecture to allow changing the internal process, like implementing new steps (pre-parser, parser, etc...) I found the Heritrix Project ( http://crawler.archive.org/ ). But there are other nice projects like that?
Tags (comma-separated)
Save Edits
Cancel