Aim
I am looking to scrape 20/20 cricket scorecard data from the Cricinfo website, ideally into CSV form for data analysis in Excel
As an example the current Australian Big Bash 2011/12 scorecards are available from
- Game 1: http://www.espncricinfo.com/big-bash-league-2011/engine/match/524915.html
- Last Game: http://www.espncricinfo.com/big-bash-league-2011/engine/match/524935.html
Background
I am proficient in using VBA (either automating IE or using XMLHTTP and then using regular expressions) to scrape data from websites, ie
Extract values from HTML TD and Tr
In that same question a comment was posted suggesting html parsing - which I hadn't come accross before - so I have taken a look at questions such as RegEx match open tags except XHTML self-contained tags
Query
While I could write a regex to parse the cricket data below I would like advice as to how I could efficiently retrieve these results with html parsing.
Please bear in mind that my preference is a repeatable CSV format containing:
- the date/name of the match
- Team 1 name
- the output should dump up to 11 records for Team 1 (blank records where players haven't batted, ie "Did Not Bat")
- Team 2 name
- the output should dump up to 11 records for Team 2 (blank records where players haven't batted)