# 使用 Homebrew, chocolatey, MacPorts, Nix 安装 hadoop

查看 hadoop 的安装路径、可执行文件、元数据以及面向 AI 代理工作流的安全说明。

## 安装

```sh
sudo av install brew:hadoop
```

其他安装命令:

### macOS

- Homebrew (100%):

```sh
brew install hadoop
```

  证据: local Homebrew formula metadata

- MacPorts (94%):

```sh
sudo port install hadoop
```

  证据: MacPorts ports tree: java/hadoop/Portfile from https://api.github.com/repos/macports/macports-ports/git/trees/master?recursive=1

### Linux

- Nix (92%):

```sh
nix profile install nixpkgs#hadoop
```

  证据: nixpkgs package indexes: hadoop from https://raw.githubusercontent.com/NixOS/nixpkgs/master/pkgs/top-level/all-packages.nix

### Windows

- Chocolatey (92%):

```sh
choco install hadoop
```

  证据: Chocolatey community package catalog: hadoop from http://community.chocolatey.org/api/v2/Packages?$filter=IsLatestVersion&$select=Id&$top=1000&$skiptoken='9.151999','cports'

## 软件包事实

- **软件包键:** brew:hadoop
- **软件包管理器:** Homebrew
- **版本:** 3.5.0
- **来源摘要:** Framework for distributed processing of large data sets
- **主页:** <https://hadoop.apache.org/>
- **最后更新:** 2026-07-30T09:41:21-04:00
- **已生成:** 2026-08-03T19:37:03+00:00

## 可执行文件

- FederationStateStore (别名)
- distribute-exclude.sh (别名)
- hadoop (别名)
- hadoop-daemon.sh (别名)
- hadoop-daemons.sh (别名)
- hdfs (别名)
- httpfs.sh (别名)
- kms.sh (别名)
- mapred (别名)
- mr-jobhistory-daemon.sh (别名)
- refresh-namenodes.sh (别名)
- start-all.sh (别名)
- start-balancer.sh (别名)
- start-dfs.sh (别名)
- start-secure-dns.sh (别名)
- start-yarn.sh (别名)
- stop-all.sh (别名)
- stop-balancer.sh (别名)
- stop-dfs.sh (别名)
- stop-secure-dns.sh (别名)
- stop-yarn.sh (别名)
- workers.sh (别名)
- yarn-daemon.sh (别名)
- yarn-daemons.sh (别名)

## 安装行为

- Bottle: 不可用

## 版本和新鲜度

- 页面生成时间: 2026-08-03
- 管理器版本: 3.5.0
## 项目历史与用法

Apache Hadoop is the Apache Software Foundation framework for reliable, scalable, distributed computing across clusters. The Homebrew package exposes the command-line surface around Hadoop Common, HDFS, YARN, and MapReduce, making it a developer workstation entry point into a historically server-oriented big-data stack.

### 项目历史

Apache's early Hadoop documentation records the project split from Nutch MapReduce and distributed filesystem code in 2006, with Hadoop code moved into its own source tree in February, project approval in March, and release 0.1.0 in April. The project grew around a practical abstraction: store data across many machines with HDFS and run distributed computation with MapReduce.

Hadoop 1.0.0 arrived in December 2011 after what Apache called six years of gestation. That release came from the 0.20-security line and included security, HBase-related HDFS append/sync work, WebHDFS, and performance improvements. The 1.x era established Hadoop as a deployable foundation for large-scale storage and batch processing.

The 2.x line changed Hadoop's architecture by introducing YARN, described in Apache release notes as NextGen MapReduce. YARN separated cluster resource management from the MapReduce programming model, letting Hadoop run a wider set of distributed applications while keeping HDFS as a common storage layer. Apache's 2.0.2-alpha notes also recorded YARN deployment on a 2000-node cluster at release time.

Hadoop 3.x continued the infrastructure evolution. The 3.0.0 documentation highlights erasure coding in HDFS for lower storage overhead than three-way replication, along with a broader set of HDFS, YARN, MapReduce, and operational features. Later Apache site news documents 3.4 and 3.5 lines as large maintenance and enhancement releases.

### 采用历史

Hadoop's adoption history is tied to the rise of commodity-cluster data processing. The Apache homepage describes a framework designed to scale from single servers to thousands of machines, with failure detection and handling in the application layer rather than relying only on hardware high availability.

The broader ecosystem made Hadoop more than one binary. Apache release notes and docs refer to HBase, Pig, Oozie, Hive, Sqoop, WebHDFS, and YARN applications as downstream or companion systems. For package managers, that ecosystem means Hadoop often arrives as a dependency, local test harness, or compatibility target even when production clusters are provisioned by other means.

### 使用方式

On a workstation, the package is commonly used for command-line tools such as `hadoop`, `hdfs`, `yarn`, and `mapred`, for local pseudo-distributed testing, for interacting with HDFS-compatible services, or for building software that expects Hadoop client libraries and scripts. The single-node setup docs are a common starting point because they let developers run Hadoop locally before touching a cluster.

Operational Hadoop usage is configuration-heavy. Apache cluster setup documentation names site-specific XML files under `etc/hadoop`, including core-site.xml, hdfs-site.xml, yarn-site.xml, and mapred-site.xml, plus environment scripts such as hadoop-env.sh and yarn-env.sh. The same docs instruct administrators to distribute configuration to HADOOP_CONF_DIR across machines.

### 为什么软件包爱好者会关心

Hadoop is a heavyweight formula by package-manager standards: a Java-centric distributed-systems stack with many scripts, XML configs, daemon roles, and ecosystem expectations. It matters to package nerds because installing it locally can satisfy client tooling, build-time assumptions, demos, and compatibility tests without standing up a full distribution.

It is also a naming and dependency landmark. Many big-data packages historically needed to say which Hadoop line, HDFS client, YARN API, or MapReduce compatibility level they targeted. That makes the package less about a single executable and more about a versioned contract with the surrounding data ecosystem.

### 时间线

- 2006: Hadoop code moved out of Nutch, the project was approved, and release 0.1.0 was published.
- 2011: Hadoop 1.0.0 was released from the 0.20-security code line.
- 2012: Hadoop 2.0.0-alpha introduced HDFS HA, YARN, HDFS Federation, and protobuf wire compatibility.
- 2012: Hadoop 2.0.2-alpha notes recorded a more stable YARN deployed on a 2000-node cluster.
- 2017: Hadoop 3.0.0 documentation described HDFS erasure coding and other 3.x architecture updates.
- 2024: Apache announced the Hadoop 3.4 line with thousands of fixes, improvements, and enhancements since 3.3.
- 2026: Apache announced the Hadoop 3.5 line with hundreds of fixes, improvements, and enhancements since 3.4.

### Related projects

- Apache Nutch: origin project from which Hadoop's early MapReduce and distributed filesystem code split.
- HDFS: Hadoop Distributed File System storage layer.
- YARN: Hadoop cluster resource-management framework.
- MapReduce: batch-processing framework and command surface in Hadoop.
- Apache HBase, Hive, Pig, Oozie, and Sqoop: common Hadoop-era ecosystem projects referenced in Apache release and setup material.

### 来源

- <https://hadoop.apache.org/>
- <https://hadoop.apache.org/docs/r3.0.0/>
- <https://hadoop.apache.org/docs/stable/hadoop-project-dist/hadoop-common/ClusterSetup.html>
- <https://hadoop.apache.org/docs/stable/hadoop-project-dist/hadoop-common/SingleCluster.html>
- <https://hadoop.apache.org/release/2.0.0-alpha.html>
- <https://hadoop.apache.org/release/page/9.html>
- <https://hadoop.apache.org/version_control.html>
- <https://web.mit.edu/~mriap/hadoop/hadoop-0.13.1/docs/>


## 安全说明

broad file, network, media, or database tool signal.

- **Geiger 风险:** blue / 中
- broad file, network, media, or database tool signal


## Configuration and credential file locations

These source-backed paths show where this package keeps local settings or durable credentials. Automic Vault can use them as review targets for secret scanning, migration, and command approval.


## Configuration files

- Unix: $HADOOP_CONF_DIR/core-site.xml, $HADOOP_CONF_DIR/hdfs-site.xml, $HADOOP_CONF_DIR/mapred-site.xml, $HADOOP_CONF_DIR/yarn-site.xml, etc/hadoop/core-site.xml, etc/hadoop/hdfs-site.xml, etc/hadoop/mapred-site.xml, etc/hadoop/yarn-site.xml
## 其他软件包管理器记录

- Nix - hadoop: normalized package name match | nixpkgs package indexes: hadoop from https://raw.githubusercontent.com/NixOS/nixpkgs/master/pkgs/top-level/all-packages.nix
- MacPorts - hadoop: normalized package name match | MacPorts ports tree: java/hadoop/Portfile from https://api.github.com/repos/macports/macports-ports/git/trees/master?recursive=1
- Chocolatey - hadoop: normalized package name match | Chocolatey community package catalog: hadoop from http://community.chocolatey.org/api/v2/Packages?$filter=IsLatestVersion&$select=Id&$top=1000&$skiptoken='9.151999','cports'


## Combined YAML source

View the package source record on GitHub. [combined/hadoop.yml](https://github.com/mxcl/pkgdb/blob/main/combined/hadoop.yml)


## 来源

- pkg.so package database
- Geiger risk classifier
- curated configuration and credential file locations
- curated package history
- pkgdb category and tag curation
- external package-manager database matches
- cross-ecosystem install command graph
