Tag
LLM evaluation
LLM evaluation examines whether models reason, judge, and stay consistent beyond producing a plausible answer. It spans long-horizon benchmarks like LongCoT, ASR quality assessment, and agreement with human labels on tasks where accuracy alone misses real failure modes.
15 articles

Dutch Government LLMs Need More Than Accuracy
A Dutch benchmark suite shows government LLMs trade off quality, cost, energy, bias, and transparency.

Prepare for Gemini 3.5 Pro on launch day
A practical setup guide for testing Gemini 3.5 Pro when it becomes available.

SocietyBench tests social-event forecasting
SocietyBench measures whether LLMs can forecast how real social events unfold, using anonymized counterfactual timelines.

Opus 5 lets you cut cost without losing quality
I break down Claude Opus 5’s pricing and performance claims into a copy-ready rollout template.

Partition, Prompt, Aggregate: Testing LLM Self-Consistency
The paper shows that LLMs often violate basic probabilistic consistency when aggregating subpopulation estimates.

A benchmark for scientific lineage reasoning
IG-Bench tests whether LLMs can trace scientific idea lineage and generate new ideas from it.

BINEVAL uses binary questions to score LLM outputs
BINEVAL splits LLM evals into yes-or-no questions, improving inspectability and matching or beating G-Eval and UniEval on key benchmarks.

Measuring when LLM behavior actually переносится
A new framework tests whether an LLM’s behavior transfers across payoff-equivalent decision environments.

AI Benchmarks 2026: Top Evaluations and Limits
MMLU, HLE, SWE-Bench and agent tests are hitting limits in 2026, while production gaps and contamination keep human review necessary.

Confident AI’s guide to LLM evaluation metrics
Confident AI explains how to score LLMs with metrics that match correctness, relevance, hallucination, and agent task completion.

Cattle Trade benchmarks LLM bluffing and bargaining
Cattle Trade is a multi-agent benchmark for testing how LLMs bluff, bid, and bargain in negotiation tasks.

DeepTest 2026 benchmarks an LLM car manual assistant
DeepTest’s first LLM testing competition compared four tools on car manual retrieval, showing how to benchmark automotive assistants.

Why Databricks RAG Is a Platform Play, Not a Feature
Databricks treats RAG as an end-to-end platform problem, and that is the right way to build it.

LLMs for ASR Evaluation: Beyond WER
This paper tests decoder-based LLMs as ASR evaluators and finds they beat WER on human agreement, with 92–94% on one task.

LongCoT Benchmark: 2,500-Probl. Long-Horizon Reasoning
LongCoT is a 2,500-problem benchmark for measuring whether frontier models can sustain long, interdependent reasoning chains.