macOS
brew install llama.cppprovider-native install command
安装
brew install llama.cppprovider-native install command
概览
LLM inference in C/C++
历史
llama.cpp is one of the defining packages of the local-LLM era: a C/C++ inference stack that made it practical to run quantized transformer models on laptops, desktops, servers, and small devices without a heavyweight Python runtime.
The repository was created on GitHub on March 10, 2023, shortly after Meta's LLaMA model release changed the center of gravity for local language-model experimentation. The README states the project goal as LLM inference with minimal setup and strong performance across local and cloud hardware.
The project is closely tied to ggml. Its README describes llama.cpp as the main playground for developing new ggml features, and the implementation grew around plain C/C++, integer quantization, CPU backends, and hardware accelerators such as Metal, CUDA, Vulkan, SYCL, HIP, and related GPU paths.
As model support broadened beyond the original LLaMA family, llama.cpp became a runtime and tooling umbrella: converters, quantizers, benchmarking tools, embedding tools, an OpenAI-compatible server, multimodal support, and many model-family loaders are represented in the command set and documentation.
Package adoption spread because llama.cpp lowered the cost of trying local inference: build from source, install from Homebrew, Nix, winget, conda-forge, Docker, or download release binaries, then run a model file or fetch one from Hugging Face-oriented workflows.
The README's bindings list shows the surrounding ecosystem that formed around the C/C++ core, including Python, Go, Node.js, Ruby, browser/Wasm, editor-completion plugins, and server clients. That ecosystem made llama.cpp both an end-user CLI and a library/runtime target for other packages.
Its high-frequency build-tag release pattern reflects active downstream pressure: package managers, bindings, model hubs, and local-AI applications all depend on fast propagation of backend, quantization, and model-format changes.
Users run llama-cli for local prompts, llama-server for an OpenAI-compatible HTTP API, llama-bench for performance testing, llama-quantize for smaller model files, and many auxiliary tools for embeddings, perplexity, retrieval, tokenization, and model-file manipulation.
The package is especially useful when a developer wants a self-contained inference engine: compile once, point it at a model, and choose a CPU/GPU backend without adopting a full ML framework stack.
For package maintainers, llama.cpp is unusually dynamic: hardware backend flags, model-format transitions, CLI renames, bundled tools, and release cadence all matter. It turned local AI into something package managers had to treat like a fast-moving systems tool rather than a single Python application.
It is also a packaging bridge between model hubs and Unix tooling. The same project can be installed as a formula, used as a server daemon, linked by language bindings, wrapped by desktop apps, or embedded in another inference product.
安全态势
没有找到 llama.cpp 的匹配本地密钥处理 manifest。Nucleus 软件包元数据仍在此发布,以便未来覆盖拥有稳定的软件包 URL。
在无人值守的代理使用前,请检查该工具是否读取明文凭据、写入远程状态、发布制品或调用插件。
可执行文件
| 命令 | 类型 | 暴露范围 | 备注 |
|---|---|---|---|
llama | 可执行文件 | 已索引可执行文件 | 从本地可执行文件索引发现。 |
llama-batched | 可执行文件 | 已索引可执行文件 | 从本地可执行文件索引发现。 |
llama-batched-bench | 可执行文件 | 已索引可执行文件 | 从本地可执行文件索引发现。 |
llama-bench | 可执行文件 | 已索引可执行文件 | 从本地可执行文件索引发现。 |
llama-cli | 可执行文件 | 已索引可执行文件 | 从本地可执行文件索引发现。 |
llama-completion | 可执行文件 | 已索引可执行文件 | 从本地可执行文件索引发现。 |
llama-debug | 可执行文件 | 已索引可执行文件 | 从本地可执行文件索引发现。 |
llama-debug-template-parser | 可执行文件 | 已索引可执行文件 | 从本地可执行文件索引发现。 |
llama-diffusion-cli | 可执行文件 | 已索引可执行文件 | 从本地可执行文件索引发现。 |
llama-embedding | 可执行文件 | 已索引可执行文件 | 从本地可执行文件索引发现。 |
llama-eval-callback | 可执行文件 | 已索引可执行文件 | 从本地可执行文件索引发现。 |
llama-finetune | 可执行文件 | 已索引可执行文件 | 从本地可执行文件索引发现。 |
llama-fit-params | 可执行文件 | 已索引可执行文件 | 从本地可执行文件索引发现。 |
llama-gen-docs | 可执行文件 | 已索引可执行文件 | 从本地可执行文件索引发现。 |
llama-gguf | 可执行文件 | 已索引可执行文件 | 从本地可执行文件索引发现。 |
llama-gguf-hash | 可执行文件 | 已索引可执行文件 | 从本地可执行文件索引发现。 |
llama-gguf-split | 可执行文件 | 已索引可执行文件 | 从本地可执行文件索引发现。 |
llama-idle | 可执行文件 | 已索引可执行文件 | 从本地可执行文件索引发现。 |
llama-imatrix | 可执行文件 | 已索引可执行文件 | 从本地可执行文件索引发现。 |
llama-lookahead | 可执行文件 | 已索引可执行文件 | 从本地可执行文件索引发现。 |
llama-lookup | 可执行文件 | 已索引可执行文件 | 从本地可执行文件索引发现。 |
llama-lookup-create | 可执行文件 | 已索引可执行文件 | 从本地可执行文件索引发现。 |
llama-lookup-merge | 可执行文件 | 已索引可执行文件 | 从本地可执行文件索引发现。 |
llama-lookup-stats | 可执行文件 | 已索引可执行文件 | 从本地可执行文件索引发现。 |
llama-mtmd-cli | 可执行文件 | 已索引可执行文件 | 从本地可执行文件索引发现。 |
llama-parallel | 可执行文件 | 已索引可执行文件 | 从本地可执行文件索引发现。 |
llama-passkey | 可执行文件 | 已索引可执行文件 | 从本地可执行文件索引发现。 |
llama-perplexity | 可执行文件 | 已索引可执行文件 | 从本地可执行文件索引发现。 |
llama-quantize | 可执行文件 | 已索引可执行文件 | 从本地可执行文件索引发现。 |
llama-results | 可执行文件 | 已索引可执行文件 | 从本地可执行文件索引发现。 |
llama-retrieval | 可执行文件 | 已索引可执行文件 | 从本地可执行文件索引发现。 |
llama-server | 可执行文件 | 已索引可执行文件 | 从本地可执行文件索引发现。 |
llama-simple | 可执行文件 | 已索引可执行文件 | 从本地可执行文件索引发现。 |
llama-simple-chat | 可执行文件 | 已索引可执行文件 | 从本地可执行文件索引发现。 |
llama-speculative | 可执行文件 | 已索引可执行文件 | 从本地可执行文件索引发现。 |
llama-speculative-simple | 可执行文件 | 已索引可执行文件 | 从本地可执行文件索引发现。 |
llama-template-analysis | 可执行文件 | 已索引可执行文件 | 从本地可执行文件索引发现。 |
llama-tokenize | 可执行文件 | 已索引可执行文件 | 从本地可执行文件索引发现。 |
llama-tts | 可执行文件 | 已索引可执行文件 | 从本地可执行文件索引发现。 |
新鲜度
这些信号区分页生成时间、软件包管理器活动和上游发布比较。只有存在证据 URL 和可比较版本时,才会提示版本落后。
安装元数据
| 软件包键 | brew:llama.cpp |
|---|---|
| 版本 | 10210 |
| 软件包管理器 | Homebrew |
| 主页 | https://llama.app |
| 仓库 | https://github.com/ggml-org/llama.cpp |
| 最后更新 | 2026-07-31T19:04:02Z |
| Pulse | updated |
| Bottle | 未记录 |
| 服务 | 未声明 |
来源线索
此页面由 av-web 从 scripts/generate-pkg-sqlite.py 生成的私有软件包 SQLite 工件提供。
View the package source record on GitHub.