
Introducing the Verafy Bias Detector (Beta)
565 contentious questions, any model on OpenRouter, and a jury of models from different companies
Ask ten language models the same contested question and you'll get ten different answers. Some refuse, some hedge, and some lean one way without saying so. The Verafy Bias Detector is built to show where each model leans, consistently and in a form anyone can check.
It's in beta, and every score on the site today is simulated while we finish the live pipeline. The question bank, the method, and the interface are real.
How it works
- 565 versioned contentious questions. Political, geopolitical, and cultural questions that have answers you can check. No question records a "correct" political position. The set is versioned, so results from different runs can be compared.
- Any model on OpenRouter. A standing roster of 20 major models, plus any other model you want to test.
- A random cross-company jury. Each answer is scored by a jury of models drawn at random from different companies, so no single vendor grades its own work. Jev, from TypeSafe AI, does a cheap first pass, and the LLM jury writes the rationale.
- A report card per model. Where it leans, where it refuses or hedges, and how it compares with the rest of the roster.
- Results inscribed on Solana through IQ Labs, so a published run can't be quietly rewritten later.


Browse the questions
Every question in the bank is public, along with its category and version. You can see exactly what the models were asked.

Watch a run
The run view shows a benchmark as it happens: questions going out, answers coming back, and the jury scoring them.

About $13 a run
Doing this naively, with every model answering every question and a frontier-model jury scoring every answer, costs about $155 per full run. Batch endpoints, a cheaper jury from different vendors, and Jev triage bring it to about $13. Almost all of that is the cost of the models being tested, which can't be avoided.

Where it fits
The Bias Detector is a sibling of the Swarm Explorer. The Swarm Explorer compares models one question at a time. The Bias Detector runs the whole question set and keeps score over time, so you can see whether a model is moving and in which direction.
Try it at bias.verafy.ai. Live runs are coming soon.