Compress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
promptfoo
AI 应用与智能体promptfoo/promptfoo
Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.
复现步骤
按顺序执行即可在本地跑起来;具体参数以项目 README 为准。
- 1
克隆仓库到本地
git clone --depth 1 https://github.com/promptfoo/promptfoo.git cd promptfoo - 2
用 Docker 一键起环境,无需本机装依赖。仓库里没有 compose 文件时,改用 docker build -t app . 再 docker run --rm -it app
docker compose up -d - 3
安装 Node 依赖并启动开发服务
npm install && npm run dev
为什么这个项目容易复现
- README 有明确的安装/快速开始章节
- 提供 Docker / Compose,开箱即用
- 有明确的依赖清单,环境可还原
- 带示例 / demo 目录
- 有独立文档目录
- 有测试,质量更有保障
- MIT 许可证,可放心使用
- 有正式 Release 版本
- 两周内仍在活跃更新
项目 README
Promptfoo: LLM evals & red teaming
Promptfoo is now part of OpenAI. Promptfoo remains open source and MIT licensed. Read the company update.
Quick Start
Requires Node.js >=22.22.0 for npm and npx usage. Node.js 24 LTS
is recommended; see the runtime support guide.
npm install -g promptfoo
promptfoo init --example getting-started
Also available via brew install promptfoo and pip install promptfoo. You can also use npx promptfoo@latest to run any command without installing.
Most LLM providers require an API key. Set yours as an environment variable:
export OPENAI_API_KEY=sk-abc123
Once you're in the example directory, run an eval and view results:
cd getting-started
promptfoo eval
promptfoo view
See Getting Started (evals) or Red Teaming (vulnerability scanning) for more.
What can you do with Promptfoo?
- Test your prompts and models with automated evaluations
- Secure your LLM apps with red teaming and vulnerability scanning
- Compare models side-by-side (OpenAI, Anthropic, Azure, Bedrock, Ollama, and more)
- Automate checks in CI/CD
- Review pull requests for LLM-related security and compliance issues with code scanning
- Share results with your team
Here's what it looks like in action:
It works on the command line too:
It also can generate security vulnerability reports:
Why Promptfoo?
- Developer-first: Fast, with features like live reload and caching
- Private: LLM evals run 100% locally - your prompts never leave your machine
- Flexible: Works with any LLM API or programming language
- Battle-tested: Powers LLM apps serving 10M+ users in production
- Data-driven: Make decisions based on metrics, not gut feel
- Open source: MIT licensed, with an active community
Learn More
- Getting Started
- Full Documentation
- Red Teaming Guide
- CLI Usage
- Node.js Package
- Supported Models
- Code Scanning Guide
Contributing
We welcome contributions! Check out our contributing guide to get started.
Join our Discord community for help and discussion.
同方向的其他项目
Milvus is a high-performance, cloud-native vector database built for scalable vector ANN search
Open-source 24/7 Cowork app for OpenClaw, Hermes, Claude Code, Codex, OpenCode and 20+ more CLI Agent | Customize your assistants | Team them up|Star if you like it!