A free, browser-based evaluation bench for AI prompts. Pick one of four built-in tasks (support-ticket triage, meeting actions, email classification, changelog writing), see what “good” means as a weighted rubric, score a naive prompt against an output-specified one across six golden examples, and get a paired signal-vs-noise verdict — is the improvement real, or within the example-to-example spread? Then A/B your own task live: a few free runs a day on the house (relayed through this site to Anthropic, nothing stored), or bring your own API key for unlimited runs straight from your browser.
Part of the Sustainable AI Toolkit: all tools · methodology & sources · Model Router · Prompt Efficiency · Water Footprint
This page requires JavaScript to run the interactive bench.