KnowledgeHub
Questions
Tags
Users
Search
Alex Rivera
|
Logout
Edit Question
Title
Body
SOLVED: See Update #2 below for the 'solution' to this issue. ~~~~~~~ In s3, I have some log*.gz files stored in a nested directory structure like: s3://($BUCKET)/y=2012/m=11/d=09/H=10/ I'm attempting to load these into Hive on Elastic Map Reduce (EMR), using a multi-level partition spec like: create external table logs (content string) partitioned by (y string, m string, d string, h string) location 's3://($BUCKET)'; Creation of the table works. I then attempt to recover all of the existing partitions: alter table logs recover partitions; This seems to work and it does drill down through my s3 structure and add all the various levels of directories: hive> show partitions logs; OK y=2012/m=11/d=06/h=08 y=2012/m=11/d=06/h=09 y=2012/m=11/d=06/h=10 y=2012/m=11/d=06/h=11 y=2012/m=11/d=06/h=12 y=2012/m=11/d=06/h=13 y=2012/m=11/d=06/h=14 y=2012/m=11/d=06/h=15 y=2012/m=11/d=06/h=16 ... So it seems that Hive can see and interpret my file layout successfully. However, no actual data ever gets loaded. If I try to do a simple count or select *, I get nothing: hive> select count(*) from logs; ... OK 0 hive> select * from logs limit 10; OK hive> select * from logs where y = '2012' and m = '11' and d = '06' and h='16' limit 10; OK Thoughts? Am I missing some additional command to load data beyond recovering the partitions? If I manually add a partition with an explicit location, then that works: alter table logs2 add partition (y='2012', m='11', d='09', h='10') location 's3://($BUCKET)/y=2012/m=11/d=09/H=10/' I can just write a script to do this, but it feels like I'm missing something fundamental w.r.t 'recover partitions'. UPDATE #1 Thanks to a brilliant and keen observation by Joe K in
Tags (comma-separated)
Save Edits
Cancel