KnowledgeHub
Questions
Tags
Users
Search
Alex Rivera
|
Logout
Edit Question
Title
Body
I have implemented MapReduce paradigm based local clustering coefficient algorithm . However I have run into serious troubles for bigger datasets or specific datasets (high average degree of a node). I tried to tune my hadoop platform and the code however the results were unsatisfactory (to say the least). No I have turned my attention to actually change/improve the algorithm. Below is my current algorithm (pseudo code) foreach(Node in Graph) { //Job1 /* Transform edge-based input dataset to node-based dataset */ //Job2 map() { emit(this.Node, this.Node.neighbours) //emit myself data to all my neighbours emit(this.Node, this.Node) //emit myself to myself } reduce() { NodeNeighbourhood nodeNeighbourhood; while(values.hasNext) { if(myself) this.nodeNeighbourhood.setCentralNode(values.next) //store myself data else this.nodeNeighbourhood.addNeighbour(values.next) //store neighbour data } emit(null, this.nodeNeighbourhood) } //Job3 map() { float lcc = calculateLocalCC(this.nodeNeighbourhood) emit(0, lcc) //emit all lcc to specific key, combiners are used } reduce() { float combinedLCC; int numberOfNodes; while(values.hasNext) { combinedLCC += values.next; } emit(null, combinedLCC/numberOfNodes); // store graph average local clustering coefficient } } Little bit more details about the code. For directed graphs neighbour data is restricted to node ID and OUT edges destination IDs (to decrease the data size), for undirected its also node ID and edges destination IDs. Sort and Merge buffers are increased to 1.5 Gb, merge streams 80. It can be clearly seen that Job2 is the actual problem of the whole algorithm. It generates massive amount of data that has to be sorted/copied/merged. This basically kills my al
Tags (comma-separated)
Save Edits
Cancel