Alex Rivera | Logout

Running jobs parallely in hadoop

Asked 2011-09-20T10:22:17.810
11

I am new to hadoop.

I have set up a 2 node cluster.

How to run 2 jobs parallely in hadoop.

When i submit jobs, they are running one by one in FIFO order. I have to run the jobs parallely. How to acheive that.

Thanks MRK

Edit
Report

1 Answer

13

Hadoop can be configured with a number of schedulers and the default is the FIFO scheduler.

FIFO Schedule behaves like this.

Scenario 1: If the cluster has 10 Map Task capacity and job1 needs 15 Map Task, then running job1 takes the complete cluster. As job1 makes progress and there are free slots available which are not used by job1 then job2 runs on the cluster.

Scenario 2: If the cluster has 10 Map Task capacity and job1 needs 6 Map Task, then job1 takes 6 slots and job2 takes 4 slots. job1 and job2 run in parallel.

To run jobs in parallel from the start, you can either configure a Fair Scheduler or a Capacity Scheduler based on your requirements. The mapreduce.jobtracker.taskscheduler and the specific scheduler parameters have to be set for this to take effect in the mapred-site.xml.

Edit: Updated the answer based on the comment from MRK.

answered 2011-09-20T11:55:01.793

Your Answer