Why Hadoop Is Not Good for Small Files

Hadoop is not suited for small data. Hadoop distributed file system lacks the ability to efficiently support the random reading of small files because of its high capacity design. … If there are too many small files, then the NameNode will be overloaded since it stores the namespace of HDFS.

Why does Hadoop not effective with large number of small files also suggest how does this issue can be handled by MapReduce program?

HDFS lacks the ability to support the random reading of small due to its high capacity design. Small files are smaller than the HDFS Block size (default 128MB). If you are storing these huge numbers of small files, HDFS cannot handle these lots of small files.

What would happen if you store too many small files in a cluster on HDFS *?

Too many small files can also cause the NameNode to run out of metadata space in memory before the DataNodes run out of data space on disk. The datanodes also report block changes to the NameNode over the network; more blocks means more changes to report over the network.

Chloe Bennett

Chloe Bennett

Culture, Media & Entertainment Columnist

Chloe Bennett explores the intersection of pop culture, streaming entertainment, digital trends, and contemporary lifestyle. Her weekly commentary reaches thousands of culture enthusiasts.

Share this article
Twitter Facebook Pinterest