KnowledgeHub
Questions
Tags
Users
Search
Alex Rivera
|
Logout
Edit Question
Title
Body
I have webpage that I need to scrape some data from. The problem is, each page may or may not have specific data, or it may have extra data above or below it in the DOM, and there is no CSS ids to speak of. Typically I could use either CSS ids or XPath to get to the node I'm looking for. I don't have that option in this case. What I'm trying to do is search for the "label" text then grab the data in the next <TD> node: <tr> <td><b>Name:</b></td> <td>Joe Smith <small><a href="/Joe"><img src="/joe.png"></a></small></td> </tr> In the above HTML, I would search for: doc.search("[text()*='Name:']") to get the node just before the data I need, but I'm not sure how to navigate from there.
Tags (comma-separated)
Save Edits
Cancel