Skip to main content
Research

We test everything. Then we publish the results.

Studies, benchmarks, and technical articles from the ASURIQ team. The numbers on our product pages have a methodology behind them, and studies in progress to test them. This is where we show our work.

Looking for the scoring framework itself? The three-category scoring model and evidence standard behind every ASURIQ verdict live at Methodology →. This page covers our own studies and benchmarks.

1 active study · reproducible methodology · raw data published
StudyIn progressVerify·2026-06

100 Questions: How Often Does AI Get It Wrong?

100 questions across 10 domains. Run through GPT-4o mini, then verified by the ASURIQ engine at full depth. Methodology defined, test harness built, initial results being validated.

Key finding: Preliminary results: ~37% of AI responses contain claims that are unsupported, exaggerated, or contradicted by evidence databases. Full methodology and raw data will be published with the final report.
37%
flagged (preliminary)
Our methodology

How we run studies.

Every study published here follows the same principles. Were building trust in AI verification. That starts with being transparent about how we verify our own claims.

Reproducible
Every study includes the exact prompts, models, parameters, and database versions used. If you want to run the same test, you can.
Pre-registered hypotheses
We state what we expect to find before running the test. Post-hoc rationalization is the enemy of honest research.
Published negative results
When our systems don’t improve outcomes, we publish that too. Cherry-picking positive results would undermine the product we’re building.
Open to challenge
Every study includes the raw data and our methodology. If you find a flaw, we want to know. Contact research@asuriq.dev.
Reports
This page asks whether the method holds. Reports asks what it found.
Published findings from applying this methodology, one release at a time, separate from the frameworks and studies covered here.
See published findings →
Your data stays yours
Prompts stay with your provider. We see analysis metadata only.
Keys never stored
One-way hash for authentication. Your credentials pass through. Never persist.
~2 second responses
Single-model cognitive tools return in about 2 seconds.
37 of 100 flagged
We ran 100 ChatGPT answers through verification. See the study →

Currently in friends-and-family beta. Built on peer-reviewed cognitive architecture.

Baars — Global Workspace TheoryACT-R — Memory Decay ModelWang et al. 2025 — Silent AgreementLi et al. EMNLP 2024 — Sparse Debate

The evidence behind the product.

Questions about our research? research@asuriq.dev