Verafy
Discord
Main
/
List
/
Categories
/
Tags
The Verafy Bias Detector home page

Introducing the Verafy Bias Detector (Beta)

565 contentious questions, any model on OpenRouter, and a jury of models from different companies

3 min read
October 11, 2026

Ask ten language models the same contested question and you'll get ten different answers. Some refuse, some hedge, and some lean one way without saying so. The Verafy Bias Detector is built to show where each model leans, consistently and in a form anyone can check.

It's in beta, and every score on the site today is simulated while we finish the live pipeline. The question bank, the method, and the interface are real.

How it works

  • 565 versioned contentious questions. Political, geopolitical, and cultural questions that have answers you can check. No question records a "correct" political position. The set is versioned, so results from different runs can be compared.
  • Any model on OpenRouter. A standing roster of 20 major models, plus any other model you want to test.
  • A random cross-company jury. Each answer is scored by a jury of models drawn at random from different companies, so no single vendor grades its own work. Jev, from TypeSafe AI, does a cheap first pass, and the LLM jury writes the rationale.
  • A report card per model. Where it leans, where it refuses or hedges, and how it compares with the rest of the roster.
  • Results inscribed on Solana through IQ Labs, so a published run can't be quietly rewritten later.

The leaderboard

A model report card

Browse the questions

Every question in the bank is public, along with its category and version. You can see exactly what the models were asked.

The question bank

Watch a run

The run view shows a benchmark as it happens: questions going out, answers coming back, and the jury scoring them.

A benchmark run in progress

About $13 a run

Doing this naively, with every model answering every question and a frontier-model jury scoring every answer, costs about $155 per full run. Batch endpoints, a cheaper jury from different vendors, and Jev triage bring it to about $13. Almost all of that is the cost of the models being tested, which can't be avoided.

Cost per full run: $155 naive vs about $13 optimized

Where it fits

The Bias Detector is a sibling of the Swarm Explorer. The Swarm Explorer compares models one question at a time. The Bias Detector runs the whole question set and keeps score over time, so you can see whether a model is moving and in which direction.

Try it at bias.verafy.ai. Live runs are coming soon.

RSJRex St John
Forward-Looking Statements Notice: Verafy blog articles may contain forward-looking statements regarding Verafy’s future plans, technologies, and market positioning and opinions on possible future achievements. These statements are based on current expectations, predicated on attaining future goals, forecasts, and assumptions that involve risks and uncertainties. Actual outcomes and results may differ materially from those expressed or implied in these statements due to factors such as evolving market conditions, technological challenges, regulatory developments, and other risks beyond Verafy's control. Readers are cautioned not to place undue reliance on forward-looking statements. Verafy undertakes no obligation to update or revise such opinions as a result of later or new information, future events, or otherwise.