# trafilatura を Homebrew でインストール

trafilatura のインストール経路、実行ファイル、メタデータ、AI エージェント向けセキュリティノートを確認します。

## インストール

```sh
sudo av install brew:trafilatura
```

追加のインストールコマンド:

### macOS

- Homebrew (100%):

```sh
brew install trafilatura
```

  証拠: local Homebrew formula metadata

## パッケージ情報

- **パッケージキー:** brew:trafilatura
- **パッケージマネージャ:** Homebrew
- **パッケージマネージャページ:** <https://formulae.brew.sh/formula/trafilatura>
- **バージョン:** 2.2.0
- **ソース概要:** Discovery, extraction and processing for Web text
- **ホームページ:** <https://trafilatura.readthedocs.io/en/latest/>
- **上流ドキュメント:** <https://trafilatura.readthedocs.io/en/latest/>
- **ライセンス:** GPL-3.0-or-later
- **ソースアーカイブ:** <https://files.pythonhosted.org/packages/a3/96/737133a93e73e967f9c888e6cfb1f2c31b2083d27263edb19fd65a9aca02/trafilatura-2.2.0.tar.gz>
- **最終更新:** 2026-08-01T18:58:41Z
- **生成日時:** 2026-08-04T22:13:35+00:00

## 実行可能ファイル

- trafilatura (cli)
- trafilatura (エイリアス)

## 依存関係

- certifi
- python@3.14

## macOS 提供ライブラリ

- libxml2
- libxslt

## インストール挙動

- post-install フック: 未定義
- Bottle: 利用可能 対象 arm64_linux, arm64_sequoia, arm64_sonoma, arm64_tahoe, sonoma, x86_64_linux

## バージョンと鮮度

- ページ生成日: 2026-08-04
- マネージャ版: 2.2.0
- マネージャ更新日: 2026-08-01
- ローカルデータ: OK
- 上流リポジトリ: https://trafilatura.readthedocs.io/en/latest/
- 情報: Release/tag comparison is only available for GitHub repositories.
## プロジェクトの歴史と使われ方

Trafilatura is a Python package and command-line tool for discovering, downloading, extracting, and processing web text. It grew from academic web-corpus work into a widely used open-source extraction tool for NLP, data acquisition, and scraping workflows.

### プロジェクトの歴史

The official README says the work started as a PhD project at the crossroads of linguistics and NLP, initially launched to create text databases for research at the Berlin-Brandenburg Academy of Sciences, specifically the DWDS and ZDL units. The repository was created on April 8, 2019.

The project name comes from the Italian word trafilatura, referring to wire drawing and used by the author as a metaphor for refinement and conversion. Official citations connect the project to earlier research on metadata-enhanced web corpora in 2016, generic web content extraction in 2019, and the 2021 ACL/IJCNLP system demonstration paper.

Official GitHub release metadata shows v0.1.0 published on September 25, 2019 and v1.0.0 on November 30, 2021. The README notes that versions prior to v1.8.0 were GPLv3+ and that current versions are distributed under Apache 2.0.

### 採用の歴史

The official README says Trafilatura is widely used and integrated into thousands of projects, naming companies and institutions including Hugging Face, IBM, Microsoft Research, the Allen Institute, Stanford, Tokyo Institute of Technology, and the University of Munich.

The official documentation has a uses and citations page, and the README emphasizes academic citation, benchmarks, and integration into the web data extraction ecosystem. Its adoption is therefore both software-package adoption and research-method adoption.

### 使われ方

Trafilatura can be used as a Python library or as the trafilatura command-line tool. Official docs describe crawling, downloads, scraping, main-text extraction, metadata extraction, comment extraction, language detection, and output as TXT, Markdown, CSV, JSON, HTML, XML, and XML-TEI.

For package-manager users, the CLI is useful for quick extraction from live URLs or stored HTML without writing code, while Python users can embed the same extraction functions in data pipelines.

### パッケージ好きにとっての重要性

Trafilatura matters to package nerds because it packages a research-grade text extraction stack behind one CLI and Python module. It is useful for building local corpora, testing scraping pipelines, and comparing extraction quality without assembling many separate tools.

Its packaging also reflects a common modern pattern: a Python data/NLP library that is simultaneously a CLI, a research artifact with citations, and an operational dependency for downstream software.

### タイムライン

- 2016: Author's related work on metadata-enhanced web corpora appears in the official citation list.
- 2019: Repository created and v0.1.0 released.
- 2021: Trafilatura ACL/IJCNLP system demonstration paper published; v1.0.0 released.
- 2022: v1.2.x releases appear in official GitHub release metadata.
- Current docs: Trafilatura 2.1.0 documentation is published on Read the Docs.

### Related projects

- The official README lists jusText and readability as generic algorithms used in the extractor.
- The official README points to htmldate and related web data extraction packages in the author's software ecosystem.

### ソース

- <https://api.github.com/repos/adbar/trafilatura/releases - official GitHub release metadata.>
- <https://github.com/adbar/trafilatura - official README, repository metadata, history/context, adoption, license, and citations.>
- <https://trafilatura.readthedocs.io/en/latest/ - official documentation overview.>
- <https://trafilatura.readthedocs.io/en/latest/usage-cli.html - official command-line usage documentation.>
- <https://trafilatura.readthedocs.io/en/latest/used-by.html - official uses and citations page.>


## セキュリティノート

trafilatura に一致するローカルシークレット処理マニフェストは見つかりませんでした。将来の対応で安定したパッケージ URL を使えるよう、パッケージメタデータはここに公開されています。


## ソースデータベース詳細

- **Source Database:** Homebrew formula API
- **Tap:** homebrew/core
- **Full Name:** trafilatura
- **Version Scheme:** 0
- **Revision:** 0
- **Bottle Stable Root URL:** <https://ghcr.io/v2/homebrew/core>
- **Deprecated:** no
- **Disabled:** no
- **Keg Only:** no
- **URL Keys:** stable


## 関連リンク

- [Terminal utility packages](https://pkg.so/ja/terminal-utilities/) - Matched terminal and command-line workflow metadata.
- [Text processing packages](https://pkg.so/ja/text-processing-tools/) - Matched text, document, or structured-data processing metadata.
- [Language runtime packages](https://pkg.so/ja/language-runtime-packages/) - Matched language runtime, compiler, or interpreter metadata.
- [Networking and protocol packages](https://pkg.so/ja/networking-protocol-tools/) - Matched network, protocol, or remote-service metadata.
- [python@3.14](https://pkg.so/ja/brew/python-3-14/) - Runtime dependency declared by Homebrew.
- [scrapy](https://pkg.so/ja/brew/scrapy/) - Shares pkgdb curated category or tags: cli, data, python, web-crawling, web-scraping.
- [csvkit](https://pkg.so/ja/brew/csvkit/) - Shares pkgdb curated category or tags: cli, data, python.
- [visidata](https://pkg.so/ja/brew/visidata/) - Shares pkgdb curated category or tags: cli, data, python.
- [sqlite-utils](https://pkg.so/ja/brew/sqlite-utils/) - Shares pkgdb curated category or tags: cli, data, python.
- [iredis](https://pkg.so/ja/brew/iredis/) - Shares pkgdb curated category or tags: cli, data, python.
- [kaskade](https://pkg.so/ja/brew/kaskade/) - Shares pkgdb curated category or tags: cli, data, python.
- [tika](https://pkg.so/ja/brew/tika/) - Shares pkgdb curated category or tags: cli, data, text-extraction.
- [mysql-to-sqlite3](https://pkg.so/ja/brew/mysql-to-sqlite3/) - Shares pkgdb curated category or tags: cli, data, python.
- [datasette](https://pkg.so/ja/brew/datasette/) - Both packages touch the same language runtime or ecosystem. Shared terms: certifi, cli, data, python, python-3-14.
- [fred](https://pkg.so/ja/brew/fred/) - Both packages touch the same language runtime or ecosystem. Shared terms: certifi, cli, data, python, python-3-14.
- [trafilatura](https://pkg.so/ja/cargo/trafilatura/) - Same normalized package name exists in another local package ecosystem.

## Combined YAML source

View the package source record on GitHub. [combined/trafilatura.yml](https://github.com/mxcl/pkgdb/blob/main/combined/trafilatura.yml)


## ソース

- pkg.so package database
- Geiger risk classifier
- package-page enrichment
- curated package history
- package version freshness
- pkgdb category and tag curation
- package relationship graph
- cross-ecosystem install command graph
