Alex Rivera | Logout

Group by In HBase

Asked 2012-02-03T09:36:39.503
9

I almost do not know anything about HBase. Sorry for basic questions.

Imagine I have a table of 100 billion rows with 10 int, one datetime, and one string column.

  1. Does HBase allow querying this table and Group the result based on key (even a composite key)?
  2. If so, does it have to run a map/reduce job to it?
  3. How do you feed it the query?
  4. Can HBase in general perform real-time like queries on a table?
Edit
Report

1 Answer

5

FAST RANDOM READS = PREPREPARED data sitting in HBase! Use Hbase for what it is...

1. A place to store a lot of data.
2. A place from which you can do super fast reads.
3. A place where SQL is not gonna do you any good (use java).

Although you can read data from HBase and do all sorts of aggregates right in Java data structures before you return your aggregated result, its best to leave the computation to mapreduce. From your questions, it seems as if you want the source data for computation to sit in HBase. If this is the case, the route you want to take is have HBase as the source data for a mapreduce job. Do computations on that and return the aggregated data. But then again, why would you read from Hbase to run a mapreduce job? Just leave the data sitting HDFS/ Hive tables and run mapreduce jobs on them THEN load the data to Hbase tables "pre-prepared" so that you can do super fast random reads from it.

answered 2012-06-16T18:57:26.710

Your Answer