跳到主内容
@wquguru
精选70Hacker News Best(web_list)产品发布/更新

llama.cpp 项目主页:本地运行开源 AI

llama.cpp 项目主页

原文
发到 X

llama.cpp

llama.cpp

AI that lives on your computer. Open-source, private & always local.

住在你电脑里的 AI。开源、私密、始终本地运行。

Run frontier AI entirely on your machine. No API keys, no telemetry, no limits. Own your models and conversation data.

在你的机器上完全运行前沿 AI。无需 API 密钥,无遥测,无限制。拥有你的模型和对话数据。

curl -LsSf https://llama.app/install.sh | sh

curl -LsSf https://llama.app/install.sh | sh

Prefer Brew or Winget? Package managers · Rather build from source? Follow instructions

更喜欢 Brew 或 Winget?包管理器 · 想从源码构建?查看说明

Pair it with a local coding agent.

搭配本地编码代理使用。

Run llama serve, install the pi-llama plugin and launch Pi. It will automatically discover your local model. No config, no API keys. Files stay on your machine, requests never leave it.

运行 llama serve,安装 pi-llama 插件并启动 Pi。它会自动发现你的本地模型。无需配置,无需 API 密钥。文件留在你的机器上,请求永不离开。

代码 · 6
# 1. Serve a model
llama serve
# 2. Install the pi-llama plugin
pi install git:github.com/huggingface/pi-llama
# 3. Run Pi, everything is set
pi
代码 · 6
# 1. Serve a model
llama serve
# 2. Install the pi-llama plugin
pi install git:github.com/huggingface/pi-llama
# 3. Run Pi, everything is set
pi

Optimized for any hardware.

为任何硬件优化。

From your laptop to a cluster, llama.cpp runs on whatever you have. Same binary, same models, same hand-tuned kernels for every GPU and CPU.

从你的笔记本电脑到集群,llama.cpp 能在你拥有的任何设备上运行。相同的二进制文件,相同的模型,为每个 GPU 和 CPU 手工调优的内核。

Apple Silicon

Apple Silicon

M Ultra

M Ultra

RTX 5090

RTX 5090

CPU

CPU

Jetson

Jetson

H100

H100

MI300

MI300

RTX 4090

RTX 4090

A100

A100

M Pro

M Pro

M Max

M Max

DGX Spark

DGX Spark

T4

T4

Radeon RX

Radeon RX

B200

B200

Intel Arc

Intel Arc

RTX 3090

RTX 3090

Run your first model

运行你的第一个模型

Qwen 3.6

Qwen 3.6

Alibaba's next-gen natively multimodal reasoning models. Dense and MoE variants that rival models many times their size on coding and vision tasks.

阿里巴巴下一代原生多模态推理模型。密集和MoE变体,在编码和视觉任务上可与数倍于其规模的模型相媲美。

Gemma 4

Gemma 4

Google's most capable open models, built from Gemini 3 technology. Supports multimodal reasoning, agentic workflows, and 140+ languages.

谷歌最强大的开放模型,基于Gemini 3技术构建。支持多模态推理、智能体工作流和140多种语言。

GPT-OSS

GPT-OSS

OpenAI's first open-weight models since GPT-2. Built for reasoning, agentic tasks, and developer use with function calling and tool use capabilities.

OpenAI自GPT-2以来首个开放权重模型。专为推理、智能体任务和开发者使用而设计,具备函数调用和工具使用能力。

Gemma 3

Gemma 3

Google's multimodal models built from Gemini technology. Supports 140+ languages, vision, and text tasks with up to 128K context for edge to cloud deployment.

谷歌基于Gemini技术构建的多模态模型。支持140多种语言、视觉和文本任务,上下文长度可达128K,适用于从边缘到云端的部署。

Browse all models

浏览所有模型

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近