pkg.sopackage field notes

brew / 排名 218

使用 Homebrew 安装 llama.cpp

查看 llama.cpp 的安装路径、可执行文件、元数据以及面向 AI 代理工作流的安全说明。

安装

其他安装命令

macOS

Homebrew已验证 · 100%
brew install llama.cpp

provider-native install command

概览

软件包摘要

LLM inference in C/C++

命令和别名

  • llama
  • llama-batched
  • llama-batched-bench
  • llama-bench
  • llama-cli
  • llama-completion
  • llama-debug
  • llama-debug-template-parser
  • llama-diffusion-cli
  • llama-embedding
  • llama-eval-callback
  • llama-finetune
  • llama-fit-params
  • llama-gen-docs
  • llama-gguf
  • llama-gguf-hash
  • llama-gguf-split
  • llama-idle
  • llama-imatrix
  • llama-lookahead
  • llama-lookup
  • llama-lookup-create
  • llama-lookup-merge
  • llama-lookup-stats
  • llama-mtmd-cli
  • llama-parallel
  • llama-passkey
  • llama-perplexity
  • llama-quantize
  • llama-results
  • llama-retrieval
  • llama-server

历史

项目历史与用法

llama.cpp is one of the defining packages of the local-LLM era: a C/C++ inference stack that made it practical to run quantized transformer models on laptops, desktops, servers, and small devices without a heavyweight Python runtime.

项目历史

The repository was created on GitHub on March 10, 2023, shortly after Meta's LLaMA model release changed the center of gravity for local language-model experimentation. The README states the project goal as LLM inference with minimal setup and strong performance across local and cloud hardware.

The project is closely tied to ggml. Its README describes llama.cpp as the main playground for developing new ggml features, and the implementation grew around plain C/C++, integer quantization, CPU backends, and hardware accelerators such as Metal, CUDA, Vulkan, SYCL, HIP, and related GPU paths.

As model support broadened beyond the original LLaMA family, llama.cpp became a runtime and tooling umbrella: converters, quantizers, benchmarking tools, embedding tools, an OpenAI-compatible server, multimodal support, and many model-family loaders are represented in the command set and documentation.

采用历史

Package adoption spread because llama.cpp lowered the cost of trying local inference: build from source, install from Homebrew, Nix, winget, conda-forge, Docker, or download release binaries, then run a model file or fetch one from Hugging Face-oriented workflows.

The README's bindings list shows the surrounding ecosystem that formed around the C/C++ core, including Python, Go, Node.js, Ruby, browser/Wasm, editor-completion plugins, and server clients. That ecosystem made llama.cpp both an end-user CLI and a library/runtime target for other packages.

Its high-frequency build-tag release pattern reflects active downstream pressure: package managers, bindings, model hubs, and local-AI applications all depend on fast propagation of backend, quantization, and model-format changes.

使用方式

Users run llama-cli for local prompts, llama-server for an OpenAI-compatible HTTP API, llama-bench for performance testing, llama-quantize for smaller model files, and many auxiliary tools for embeddings, perplexity, retrieval, tokenization, and model-file manipulation.

The package is especially useful when a developer wants a self-contained inference engine: compile once, point it at a model, and choose a CPU/GPU backend without adopting a full ML framework stack.

为什么软件包爱好者会关心

For package maintainers, llama.cpp is unusually dynamic: hardware backend flags, model-format transitions, CLI renames, bundled tools, and release cadence all matter. It turned local AI into something package managers had to treat like a fast-moving systems tool rather than a single Python application.

It is also a packaging bridge between model hubs and Unix tooling. The same project can be installed as a formula, used as a server daemon, linked by language bindings, wrapped by desktop apps, or embedded in another inference product.

时间线

  • 2023: ggml-org/llama.cpp repository created on March 10.
  • 2023: The project manifesto discussion documented goals and direction for the early community.
  • 2024: llama.cpp issue trackers began maintaining separate API changelogs for libllama and llama-server.
  • 2025: Multimodal support reached llama-server through the upstream pull request and documentation referenced from the README.
  • 2026: Homebrew, Nix, winget, conda-forge, Docker, and release binaries were documented installation paths in the upstream README.

Related projects

  • ggml is the closest related project, because llama.cpp is the main development playground for that tensor library.
  • Important downstream and adjacent projects include llama-cpp-python, node-llama-cpp, go-llama.cpp, wllama, llama.vscode, llama.vim, and applications that wrap llama-server's OpenAI-compatible API.

安全态势

尚未找到受保护工具覆盖

没有找到 llama.cpp 的匹配本地密钥处理 manifest。Nucleus 软件包元数据仍在此发布,以便未来覆盖拥有稳定的软件包 URL。

安装行为

  • 未记录 Homebrew bottle 元数据。

建议审查

在无人值守的代理使用前,请检查该工具是否读取明文凭据、写入远程状态、发布制品或调用插件。

可执行文件

已安装的可执行文件

命令类型暴露范围备注
llama可执行文件已索引可执行文件从本地可执行文件索引发现。
llama-batched可执行文件已索引可执行文件从本地可执行文件索引发现。
llama-batched-bench可执行文件已索引可执行文件从本地可执行文件索引发现。
llama-bench可执行文件已索引可执行文件从本地可执行文件索引发现。
llama-cli可执行文件已索引可执行文件从本地可执行文件索引发现。
llama-completion可执行文件已索引可执行文件从本地可执行文件索引发现。
llama-debug可执行文件已索引可执行文件从本地可执行文件索引发现。
llama-debug-template-parser可执行文件已索引可执行文件从本地可执行文件索引发现。
llama-diffusion-cli可执行文件已索引可执行文件从本地可执行文件索引发现。
llama-embedding可执行文件已索引可执行文件从本地可执行文件索引发现。
llama-eval-callback可执行文件已索引可执行文件从本地可执行文件索引发现。
llama-finetune可执行文件已索引可执行文件从本地可执行文件索引发现。
llama-fit-params可执行文件已索引可执行文件从本地可执行文件索引发现。
llama-gen-docs可执行文件已索引可执行文件从本地可执行文件索引发现。
llama-gguf可执行文件已索引可执行文件从本地可执行文件索引发现。
llama-gguf-hash可执行文件已索引可执行文件从本地可执行文件索引发现。
llama-gguf-split可执行文件已索引可执行文件从本地可执行文件索引发现。
llama-idle可执行文件已索引可执行文件从本地可执行文件索引发现。
llama-imatrix可执行文件已索引可执行文件从本地可执行文件索引发现。
llama-lookahead可执行文件已索引可执行文件从本地可执行文件索引发现。
llama-lookup可执行文件已索引可执行文件从本地可执行文件索引发现。
llama-lookup-create可执行文件已索引可执行文件从本地可执行文件索引发现。
llama-lookup-merge可执行文件已索引可执行文件从本地可执行文件索引发现。
llama-lookup-stats可执行文件已索引可执行文件从本地可执行文件索引发现。
llama-mtmd-cli可执行文件已索引可执行文件从本地可执行文件索引发现。
llama-parallel可执行文件已索引可执行文件从本地可执行文件索引发现。
llama-passkey可执行文件已索引可执行文件从本地可执行文件索引发现。
llama-perplexity可执行文件已索引可执行文件从本地可执行文件索引发现。
llama-quantize可执行文件已索引可执行文件从本地可执行文件索引发现。
llama-results可执行文件已索引可执行文件从本地可执行文件索引发现。
llama-retrieval可执行文件已索引可执行文件从本地可执行文件索引发现。
llama-server可执行文件已索引可执行文件从本地可执行文件索引发现。
llama-simple可执行文件已索引可执行文件从本地可执行文件索引发现。
llama-simple-chat可执行文件已索引可执行文件从本地可执行文件索引发现。
llama-speculative可执行文件已索引可执行文件从本地可执行文件索引发现。
llama-speculative-simple可执行文件已索引可执行文件从本地可执行文件索引发现。
llama-template-analysis可执行文件已索引可执行文件从本地可执行文件索引发现。
llama-tokenize可执行文件已索引可执行文件从本地可执行文件索引发现。
llama-tts可执行文件已索引可执行文件从本地可执行文件索引发现。

新鲜度

版本和新鲜度

这些信号区分页生成时间、软件包管理器活动和上游发布比较。只有存在证据 URL 和可比较版本时,才会提示版本落后。

页面生成时间2026-08-03
管理器版本10210
管理器更新时间2026-07-31
本地数据未知
上游不可用
检测到的最新版本未检测到
  • OK没有生成新鲜度警告。

安装元数据

软件包元数据

软件包键brew:llama.cpp
版本10210
软件包管理器Homebrew
主页https://llama.app
仓库https://github.com/ggml-org/llama.cpp
最后更新2026-07-31T19:04:02Z
Pulseupdated
Bottle未记录
服务未声明

来源线索

由仓库数据生成

此页面由 av-webscripts/generate-pkg-sqlite.py 生成的私有软件包 SQLite 工件提供。

使用的来源

  • Geiger risk classifier
  • Nucleus package database
  • curated package history
  • pkgdb category and tag curation