Define the AI task and success criteria
Write the acceptance criteria before grading outputs.
AI Engineering tool
Compare prompt versions using accepted results, retries, review time, and a fixed test set.
Runs in your browser with editable illustrative assumptions. It does not call a model, send entered content to a server, or produce a current market quote.
Interactive calculator
Use this worksheet to compare prompt versions against the same cases. A score is only useful when the test set represents the work users actually send.
Record prompt version, model, settings, tokens, latency, and grader notes beside these results. This worksheet does not call a model or store test content.
Starting values are illustrative defaults you can edit — not live ComputeTape Market benchmark prices. Replace them with a real quote.
How to use it
Use the same representative cases for each prompt version. Quality is only one dimension; retries, tokens, latency, and review time are workload costs too.