Alex Rivera | Logout

Java & RabbitMQ - Queueing & Multithreading - Or Couchbase as Job-Queue

Asked 2012-09-05T08:07:05.423
10

i have one Job Distributor who publishes messages on different Channels.

Further, i want to have two (and more in the future) Consumers who work on different tasks and run on different machines. (Currently i have only one and need to scale it)

Let's name these tasks (just examples):

  • FIBONACCI (generates fibonacci numbers)
  • RANDOMBOOKS (generates random sentences to write a book)

Those tasks run up to 2-3 hours and should be divided equally to each Consumer.

Every Consumer can have x parallel threads for working on these tasks. So i say: (those numbers are just examples and will be replaced by variables)

  • Machine 1 can consume 3 parallel jobs for FIBONACCI and 5 parallel jobs for RANDOMBOOKS
  • Machine 2 can consume 7 parallel jobs for FIBONACCI and 3 parallel jobs for RANDOMBOOKS

How can i achieve this?

Do i have to start x Threads for each Channel to listen on on each Consumer ?

When do i have to ack that?

My current approach for only one Consumer is: Start x Threads for each Task - each Thread is a Defaultconsumer implementing Runnable. In the handleDelivery method, i call basicAck(deliveryTag,false) and then do the work.

Further: I want to send some tasks to a special consumer. How can i achieve that in combination with the fair distribution as mentioned above?

This is my Code for publishing

String QUEUE_NAME = "FIBONACCI";

Channel channel = this.clientManager.getRabbitMQConnection().createChannel();

channel.queueDeclare(QUEUE_NAME, true, false, false, n
Edit
Report

1 Answer

3

Here are my thoughts on your question. As @Daniel mentioned in his answer, I believe this is more a question of architectural principles than of implementation. Once the architecture is clear, the implementation becomes trivial.

First, I would like to address something related to scheduling theory. You have very long-running tasks here, and if they are not scheduled in the proper manner, you will either (a) end up running your servers at less than full capacity or (b) taking much longer to finish the tasks than otherwise possible. So, I have some questions for you related to your scheduling paradigm:

  1. Do you have the ability to estimate how long each job will take?
  2. Do the jobs have a due date associated with them, and if so, how is it determined?

Is RabbitMQ Appropriate in this case?

I do not believe RabbitMQ is the proper solution to dispatch extremely long-running jobs. In fact, I think you are having these questions as a result of the fact that RabbitMQ is not the right tool for the job. By default, you do not have enough insight into the jobs before you remove them from the queue to determine which should be processed next. Second, as mentioned in @Daniel's answer, you probably won't be able to use the built-in ACK mechanism, because it would be probably be bad for a job to get re-queued whenever the connection to the RabbitMQ server fails.

Instead, I would look for something like MongoDB or Couchbase to store your "queue" for jobs. Then, you can have full control over the dispatching logic, rather than rely on the built-in round-robin enforced by RabbitMQ.

Other considerations:

Further, i want to have two (and more in the future) Consumers who work on different tasks and run on different machines. (Currently i have only one and need to scale it)

In this case, I don't think you want to use a push-based consumer. Instead, use a pull-ba

answered 2013-03-14T00:26:54.747

Your Answer