How to Benchmark LLMs for Research: A Reproducible Workflow
A step-by-step workflow for benchmarking LLMs in research: tasks, metrics, contamination, prompts, seeds, eval harnesses, confidence intervals and cost control.

To benchmark LLMs for research in a way reviewers will trust, you need four things: a clearly defined task and metric, a dataset you have checked for contamination, fully pinned and logged model settings, and statistics that show how uncertain your numbers are. This post walks through a workflow we use for evaluation work, with open-source tools and code you can adapt.
Step 1: Define the research question before the benchmark
Start from the claim you want to make, not from a leaderboard. “Model A is better” is not a research question. These are:
- “Does retrieval improve factual accuracy on questions about our domain corpus?”
- “How does accuracy on task X change with model size within one model family?”
- “Can an open-weight model match a commercial API on our annotation task?”
Then write down, before running anything:
- The task: input, expected output, and what counts as correct.
- The primary metric: exact match, F1, pass@k for code, accuracy against human labels, or a rubric score.
- The comparison: which models, prompts or systems, and what difference would matter in practice.
Writing this down first (even informally, in the repository README) protects you from picking the metric that happens to make your method look best.
Step 2: Choose metrics that match the task
| Task type | Common metrics | Watch out for |
|---|---|---|
| Multiple choice / classification | Accuracy, macro-F1 | Answer extraction errors; class imbalance |
| Short-answer QA | Exact match, token F1 | Formatting differences that are not real errors |
| Code generation | pass@k with unit tests | Weak tests; sandbox security |
| Summarisation / open-ended | Human ratings, LLM-as-judge with a rubric | Judge bias; verbosity preference |
| Extraction / annotation | Agreement with human labels (e.g. Cohen’s kappa) | Low-quality gold labels |
If you use an LLM as a judge, treat it as a measuring instrument you must validate: compare its scores with human ratings on a sample, report the agreement, and keep the judge model and prompt fixed. Our LLM evaluation guide goes deeper on judge design.
Step 3: Pick datasets and check for contamination
Public benchmarks are convenient, but many have been on the internet for years and may be in the training data of the models you test. That can inflate scores.
Practical ways to reduce the risk:
- Prefer held-out or recent data. Items created after a model’s training cutoff are less likely to be memorised.
- Build a private test set for your domain, and do not publish it in plain text (publish a hashed or access-controlled version if needed).
- Probe for memorisation. Ask the model to complete the beginning of a test item word for word; exact continuation is a warning sign.
- Compare with a perturbed version. Rephrase questions or reorder options. A big drop on trivially changed items suggests memorisation rather than skill.
- Report what you checked. Reviewers accept that contamination can’t be fully ruled out. They don’t accept silence about it.
Also record dataset version, split, and hash, so others can confirm they evaluate on the same data.
Step 4: Fix prompts, decoding and seeds
LLM scores are sensitive to small choices. Pin and log all of them:
- Model identifier. For APIs, use a dated snapshot name, not a moving alias like “latest”. For open weights, record the exact repository and commit/revision.
- Prompt template, including the system prompt, few-shot examples and their order, and the chat template.
- Decoding parameters: temperature, top-p, max tokens, stop sequences.
- Seeds for sampling, few-shot selection and data shuffling.
- Inference stack: library versions (transformers, vLLM), precision and quantisation, hardware.
Even with temperature 0, API outputs are not always deterministic, and GPU inference can vary slightly between runs. The answer is to run multiple times and report the variation, not to assume determinism.
Test more than one reasonable prompt when your claim is about the model rather than the prompt. If rankings flip between prompt variants, say so. That is a finding, not a failure. See our prompt engineering best practices for how to write robust prompts.
Step 5: Use an evaluation harness instead of ad-hoc scripts
A harness handles prompt formatting, batching, answer extraction, scoring and logging consistently. As of September 2026, well-known open-source options include:
- lm-evaluation-harness (EleutherAI): a large library of academic benchmark tasks, with backends for Hugging Face models, vLLM and API models. Good for standard benchmarks and comparisons with published numbers.
- Inspect (UK AI Security Institute and Meridian Labs): a Python framework for writing custom evaluations, including agentic and tool-use tasks, with a log viewer.
- Lighteval (Hugging Face): evaluation across several backends, including vLLM and inference endpoints.
- HELM (Stanford CRFM): holistic evaluation across many scenarios and metrics. Note that the project announced it entered maintenance mode on June 1, 2026.
A reproducible lm-evaluation-harness run looks like this:
pip install "lm_eval[hf]"
lm_eval --model hf \
--model_args pretrained=EleutherAI/pythia-1.4b,revision=main,dtype=bfloat16 \
--tasks hellaswag,arc_easy \
--num_fewshot 5 \
--batch_size 8 \
--seed 0,1234,1234,1234 \
--output_path results/pythia-1.4b \
--log_samples
--log_samples saves every prompt and response, which you need for error analysis and for anyone checking your results. The --seed flag sets the Python, NumPy, PyTorch and few-shot seeds. Replace revision=main with a specific commit hash for a truly pinned run.
For a custom task, Inspect keeps the dataset, solver and scorer together in one file:
from inspect_ai import Task, task
from inspect_ai.dataset import json_dataset
from inspect_ai.scorer import match
from inspect_ai.solver import generate, system_message
@task
def domain_qa():
return Task(
dataset=json_dataset("data/domain_qa_v1.jsonl"),
solver=[system_message("Answer with a single word."), generate()],
scorer=match(),
)
inspect eval domain_qa.py --model openai/<dated-model-snapshot>
Step 6: Log everything
A good rule: someone should be able to rebuild your results table from your logs without asking you anything. Keep:
- A config file (YAML or JSON) per run, committed to version control
- Raw model outputs for every item (JSONL)
- Scores per item, not just the aggregate
- Environment info: package versions, GPU type, container image
- Timestamps and API costs
# configs/run_2026-08-01_domainqa.yaml
model: "org/model-name"
model_revision: "a1b2c3d"
dataset: "data/domain_qa_v1.jsonl"
dataset_sha256: "9f86d08..."
prompt_template: "prompts/qa_v3.txt"
temperature: 0.0
max_tokens: 64
n_repeats: 3
seeds: [0, 1, 2]
Store results in a structured way (one row per item, per model, per repeat). That makes the statistics in the next step easy.
Step 7: Report uncertainty with confidence intervals
Benchmarks are samples. With a few hundred questions, a two-point difference between models may be noise. Report:
- Confidence intervals for each score. A bootstrap over test items is simple and works for most metrics.
- Paired comparisons when models answer the same items. Pairing removes item difficulty from the comparison and gives much tighter intervals than comparing two separate intervals.
- Variation across seeds or repeats when you sample.
- Clustered standard errors when items come in groups (for example several questions about the same document).
Evan Miller’s paper “Adding Error Bars to Evals” is a clear, practical reference on this.
A minimal paired bootstrap in Python:
import numpy as np
def paired_bootstrap(a, b, n_boot=10_000, seed=0):
"""a, b: per-item scores (0/1 or floats) for two models on the same items."""
rng = np.random.default_rng(seed)
a, b = np.asarray(a), np.asarray(b)
n = len(a)
diffs = np.empty(n_boot)
for i in range(n_boot):
idx = rng.integers(0, n, n)
diffs[i] = a[idx].mean() - b[idx].mean()
lo, hi = np.percentile(diffs, [2.5, 97.5])
return a.mean() - b.mean(), (lo, hi)
diff, ci = paired_bootstrap(scores_model_a, scores_model_b)
print(f"Difference: {diff:.3f}, 95% CI [{ci[0]:.3f}, {ci[1]:.3f}]")
If the interval includes zero, do not claim one model is better. If you compare many models or prompts, account for multiple comparisons.
Step 8: Control cost and compute
Evaluation budgets get out of hand quickly with API models and large test sets. Some habits that help:
- Develop on a small subset (
--limitin lm-evaluation-harness) and run the full set only when the pipeline is final. - Cache responses keyed by model, prompt and parameters, so reruns of scoring code don’t call the API again.
- Use batch APIs where providers offer them for non-urgent jobs; they are often cheaper.
- Self-host open-weight models with vLLM for large sweeps, if you have GPU access. See our list of open-weight LLMs you can self-host.
- Log tokens and cost per run, and set spending limits in each provider’s console.
- Estimate before you run: items x repeats x (input + output tokens) x price per token.
Our LLM cost optimisation guide covers more techniques.
Step 9: Report so others can reproduce
In the paper or appendix, include:
- Model names with exact versions or snapshot dates, and the date you ran API evaluations
- Full prompts (or links to them) and decoding settings
- Dataset source, version, split, size and any filtering
- The harness and its version, plus a link to your code and configs
- Scores with confidence intervals, number of repeats and seeds
- Contamination checks and their results
- Compute used and approximate cost
- Known limitations and failure cases, with examples from the logs
Release the code, configs and (where licences allow) the raw outputs. Our guide to reproducible research code explains how to package this.
Checklist for benchmarking LLMs for research
- Research question, task and primary metric written down before running
- Dataset version and hash recorded; contamination checked and reported
- Model versions, prompts, decoding settings and seeds pinned in config files
- An established harness used, or custom code tested on a small sample
- Per-item outputs and scores logged
- Multiple runs where sampling is involved
- Confidence intervals and paired comparisons reported
- Token usage and cost tracked, with spending limits
- Code, configs and outputs released with the paper
Get help with the engineering
Designing the experiment is your job. Making it run reliably across many models, GPUs and APIs is engineering work that eats research time. We help researchers build evaluation pipelines, run benchmarks and make results reproducible. See our research engineering services or contact us to discuss your project.