Prompt & Model Benchmark Lab
Compare prompts and models on the same cases, then inspect the outputs, scores, cost, and latency behind each result.
The lab compares candidates on the same work
Prompt & Model Benchmark Lab turns a model choice into a repeatable experiment. A run crosses cases with candidates, where each candidate specifies a model, prompt version, and parameters. The lab saves each output alongside its grading result, token use, latency, and estimated cost. That record lets someone ask which candidate worked on this task and inspect the evidence behind the answer.
Prompt comparisons preserve each version’s evidence
The Prompt benchmarks screen lets a user enter text cases and expected answers, name prompt variants, select catalog models, and run a comparison. A saved setup can be copied for a revision without changing the earlier run. The current screen offers exact and contains checks; the worker also supports broader graders for experiments that need them.
Prompt runs remain separate from the general model leaderboard. Pooling prompt variants into one model score would make a prompt improvement look like a model improvement.
The mosaic shows coverage as well as performance
The model-fitness mosaic groups results by work type, complexity, and stack layer. A cell links to its contributing runs so a reader can inspect the cases and grading method. Sparse or low-calibration cells remain visible as limited evidence; an empty cell is a gap to test, not a score of zero.


Recommendations apply constraints to the recorded results
The recommendation view weighs quality against cost and latency limits for a specified task. It shows the selected candidate, alternatives, and the strength of the supporting evidence. Deterministic graders handle cases with known answers. A model judge can score open-ended work, but its composite is calculated from dimension scores in code and its confidence depends on calibration. The lab flags a result when the judge also appears as a candidate.

The execution path keeps long runs out of the dashboard
Next.js serves a trusted local dashboard. PostgreSQL stores experiment definitions and results; Redis and BullMQ run the candidate grid and grading outside the web request. OpenTelemetry traces model calls. The system can compare cataloged cloud models with local Ollama models under the same cases, while retaining the provider, prompt version, and grading method for each result.