macOS
brew install trafilaturalocal Homebrew formula metadata
安装
brew install trafilaturalocal Homebrew formula metadata
概览
Discovery, extraction and processing for Web text
历史
Trafilatura is a Python package and command-line tool for discovering, downloading, extracting, and processing web text. It grew from academic web-corpus work into a widely used open-source extraction tool for NLP, data acquisition, and scraping workflows.
The official README says the work started as a PhD project at the crossroads of linguistics and NLP, initially launched to create text databases for research at the Berlin-Brandenburg Academy of Sciences, specifically the DWDS and ZDL units. The repository was created on April 8, 2019.
The project name comes from the Italian word trafilatura, referring to wire drawing and used by the author as a metaphor for refinement and conversion. Official citations connect the project to earlier research on metadata-enhanced web corpora in 2016, generic web content extraction in 2019, and the 2021 ACL/IJCNLP system demonstration paper.
Official GitHub release metadata shows v0.1.0 published on September 25, 2019 and v1.0.0 on November 30, 2021. The README notes that versions prior to v1.8.0 were GPLv3+ and that current versions are distributed under Apache 2.0.
The official README says Trafilatura is widely used and integrated into thousands of projects, naming companies and institutions including Hugging Face, IBM, Microsoft Research, the Allen Institute, Stanford, Tokyo Institute of Technology, and the University of Munich.
The official documentation has a uses and citations page, and the README emphasizes academic citation, benchmarks, and integration into the web data extraction ecosystem. Its adoption is therefore both software-package adoption and research-method adoption.
Trafilatura can be used as a Python library or as the trafilatura command-line tool. Official docs describe crawling, downloads, scraping, main-text extraction, metadata extraction, comment extraction, language detection, and output as TXT, Markdown, CSV, JSON, HTML, XML, and XML-TEI.
For package-manager users, the CLI is useful for quick extraction from live URLs or stored HTML without writing code, while Python users can embed the same extraction functions in data pipelines.
Trafilatura matters to package nerds because it packages a research-grade text extraction stack behind one CLI and Python module. It is useful for building local corpora, testing scraping pipelines, and comparing extraction quality without assembling many separate tools.
Its packaging also reflects a common modern pattern: a Python data/NLP library that is simultaneously a CLI, a research artifact with citations, and an operational dependency for downstream software.
安全态势
没有找到 trafilatura 的匹配本地密钥处理 manifest。Nucleus 软件包元数据仍在此发布,以便未来覆盖拥有稳定的软件包 URL。
在无人值守的代理使用前,请检查该工具是否读取明文凭据、写入远程状态、发布制品或调用插件。
可执行文件
| 命令 | 类型 | 暴露范围 | 备注 |
|---|---|---|---|
trafilatura | 可执行文件 | 已索引可执行文件 | 从本地可执行文件索引发现。 |
新鲜度
这些信号区分页生成时间、软件包管理器活动和上游发布比较。只有存在证据 URL 和可比较版本时,才会提示版本落后。
安装元数据
| 软件包键 | brew:trafilatura |
|---|---|
| 版本 | 2.2.0 |
| 软件包管理器 | Homebrew |
| 主页 | https://trafilatura.readthedocs.io/en/latest/ |
| 最后更新 | 2026-08-01T18:58:41Z |
| Pulse | updated |
| Bottle | 未记录 |
| 服务 | 未声明 |
来源线索
此页面由 av-web 从 scripts/generate-pkg-sqlite.py 生成的私有软件包 SQLite 工件提供。
View the package source record on GitHub.