I have 10M+ photos saved on the local file system. Now I want to go through each of them to analyze the binary of the photo to see if it's a dog. I basically want to do the analysis on a clustered hadoop environment. The problem is, how should I design the input for the map method? let's say, in the map method,
new FaceDetection(photoInputStream).isDog() is all the underlying logic for the analysis.
Specifically,
Should I upload all of the photos to HDFS? Assume yes,
how can I use them in the
mapmethod?Is it ok to make the input(to the
map) as a text file containing all of the photo path(inHDFS) with each a line, and in the map method, load the binary like:photoInputStream = getImageFromHDFS(photopath);(Actually, what is the right method to load file from HDFS during the execution of the map method?)
It seems I miss some knowledges about the basic principle for hadoop, map/reduce and hdfs, but can you please point me out in terms of the above question, Thanks!