# Prompt & Model Benchmark Lab

Compare prompts and models on the same cases, then inspect the outputs, scores, cost, and latency behind each result.

Status: active

## The lab compares candidates on the same work

Prompt & Model Benchmark Lab turns a model choice into a repeatable experiment. A run crosses cases with candidates, where each candidate specifies a model, prompt version, and parameters. The lab saves each output alongside its grading result, token use, latency, and estimated cost. That record lets someone ask which candidate worked on this task and inspect the evidence behind the answer.

## Prompt comparisons preserve each version's evidence

The Prompt benchmarks screen lets a user enter text cases and expected answers, name prompt variants, select catalog models, and run a comparison. A saved setup can be copied for a revision without changing the earlier run. The current screen offers exact and contains checks; the worker also supports broader graders for experiments that need them.

Prompt runs remain separate from the general model leaderboard. Pooling prompt variants into one model score would make a prompt improvement look like a model improvement.

## The mosaic shows coverage as well as performance

The model-fitness mosaic groups results by work type, complexity, and stack layer. A cell links to its contributing runs so a reader can inspect the cases and grading method. Sparse or low-calibration cells remain visible as limited evidence; an empty cell is a gap to test, not a score of zero.

![Model-fitness mosaic ::border](/images/projects/prompt-model-benchmark-lab/mosaic.png)

![Benchmark run evidence ::border](/images/projects/prompt-model-benchmark-lab/evidence.png)

## Recommendations apply constraints to the recorded results

The recommendation view weighs quality against cost and latency limits for a specified task. It shows the selected candidate, alternatives, and the strength of the supporting evidence. Deterministic graders handle cases with known answers. A model judge can score open-ended work, but its composite is calculated from dimension scores in code and its confidence depends on calibration. The lab flags a result when the judge also appears as a candidate.

![Benchmark recommendation view ::border](/images/projects/prompt-model-benchmark-lab/recommend.png)

## The execution path keeps long runs out of the dashboard

Next.js serves a trusted local dashboard. PostgreSQL stores experiment definitions and results; Redis and BullMQ run the candidate grid and grading outside the web request. OpenTelemetry traces model calls. The system can compare cataloged cloud models with local Ollama models under the same cases, while retaining the provider, prompt version, and grading method for each result.

## Links

- [HTML project page](https://rosslabs.ai/projects/prompt-model-benchmark-lab/)
