
Apache Hadoop
0.0
0 ReviewsReliable, scalable, distributed computing.
About Apache Hadoop
Apache Hadoop is the absolute, foundational grandfather of the entire modern "Big Data" industry. Before Hadoop existed, if a company like Yahoo wanted to analyze a petabyte of web logs, it was physically impossible because no single database server on earth was large enough to hold the data. Hadoop solved this with the concept of "Distributed Storage."
It utilizes the Hadoop Distributed File System (HDFS). Instead of buying one $5 million supercomputer, a company buys 1,000 incredibly cheap, standard commodity servers. HDFS takes the massive petabyte file, chops it into tiny blocks, and physically scatters those blocks across all 1,000 cheap servers.
To process the data, Hadoop introduced "MapReduce." Instead of pulling the massive data over the network to the CPU (which would crash the network), MapReduce physically sends the analytical software code to the 1,000 servers. All 1,000 servers process their tiny chunk of data simultaneously (in parallel), and then send the highly condensed answers back to the master server, allowing massive analysis to happen in minutes instead of months.
Deployment
- Linux
- On-Premise
- Cloud, SaaS, Web
Support
- Community Forum
- Email/Help Desk
- Phone Support
Training
- Documentation
- Live Online
Ideal Company Size
Medium, Enterprise Employees
Pricing Overview
Free
LicensingSubscription, Open Source
Supported LanguagesEnglish
Write a Review