开源雷达
返回全部项目

promptfoo

AI 应用与智能体

promptfoo/promptfoo

Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.

cici-cdcicdevaluationevaluation-frameworkllmllm-evalllm-evaluationllm-evaluation-frameworkllmopspentestingprompt-engineering

复现步骤

按顺序执行即可在本地跑起来;具体参数以项目 README 为准。

  1. 1

    克隆仓库到本地

    git clone --depth 1 https://github.com/promptfoo/promptfoo.git
    cd promptfoo
  2. 2

    用 Docker 一键起环境,无需本机装依赖。仓库里没有 compose 文件时,改用 docker build -t app . 再 docker run --rm -it app

    docker compose up -d
  3. 3

    安装 Node 依赖并启动开发服务

    npm install && npm run dev

为什么这个项目容易复现

100
开箱即用
  • README 有明确的安装/快速开始章节
  • 提供 Docker / Compose,开箱即用
  • 有明确的依赖清单,环境可还原
  • 带示例 / demo 目录
  • 有独立文档目录
  • 有测试,质量更有保障
  • MIT 许可证,可放心使用
  • 有正式 Release 版本
  • 两周内仍在活跃更新

项目 README

Promptfoo: LLM evals & red teaming

Promptfoo is now part of OpenAI. Promptfoo remains open source and MIT licensed. Read the company update.

Quick Start

Requires Node.js >=22.22.0 for npm and npx usage. Node.js 24 LTS is recommended; see the runtime support guide.

npm install -g promptfoo
promptfoo init --example getting-started

Also available via brew install promptfoo and pip install promptfoo. You can also use npx promptfoo@latest to run any command without installing.

Most LLM providers require an API key. Set yours as an environment variable:

export OPENAI_API_KEY=sk-abc123

Once you're in the example directory, run an eval and view results:

cd getting-started
promptfoo eval
promptfoo view

See Getting Started (evals) or Red Teaming (vulnerability scanning) for more.

What can you do with Promptfoo?

  • Test your prompts and models with automated evaluations
  • Secure your LLM apps with red teaming and vulnerability scanning
  • Compare models side-by-side (OpenAI, Anthropic, Azure, Bedrock, Ollama, and more)
  • Automate checks in CI/CD
  • Review pull requests for LLM-related security and compliance issues with code scanning
  • Share results with your team

Here's what it looks like in action:

It works on the command line too:

It also can generate security vulnerability reports:

Why Promptfoo?

  • Developer-first: Fast, with features like live reload and caching
  • Private: LLM evals run 100% locally - your prompts never leave your machine
  • Flexible: Works with any LLM API or programming language
  • Battle-tested: Powers LLM apps serving 10M+ users in production
  • Data-driven: Make decisions based on metrics, not gut feel
  • Open source: MIT licensed, with an active community

Learn More

Contributing

We welcome contributions! Check out our contributing guide to get started.

Join our Discord community for help and discussion.

同方向的其他项目

headroom

headroomlabs-ai

AI 应用与智能体

Compress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.

DockerPython 依赖Rust示例代码+3
66k9.0kPythonApache-2.0今天
开箱即用
milvus

milvus-io

AI 应用与智能体

Milvus is a high-performance, cloud-native vector database built for scalable vector ANN search

DockerGo示例代码文档+2
46k542GoApache-2.0今天
开箱即用
AionUi

iOfficeAI

AI 应用与智能体

Open-source 24/7 Cowork app for OpenClaw, Hermes, Claude Code, Codex, OpenCode and 20+ more CLI Agent | Customize your assistants | Team them up|Star if you like it!

DockerNode 依赖示例代码文档+2
32k2.6kTypeScriptApache-2.02 天前
开箱即用