Skip to main content
Research

We test everything. Then we publish the results.

Studies, benchmarks, and technical articles from the ASURIQ team. Every claim we make on our product pages has a study behind it. Every number has a methodology. This is where we show our work.

1 active study · 5 planned · 4 products · reproducible methodology
StudyIn progressVerify·2026-06

100 Questions: How Often Does AI Get It Wrong?

100 questions across 10 domains. Run through GPT-4o mini, then verified by the ASURIQ engine at full depth. Methodology defined, test harness built, initial results being validated.

Key finding: Preliminary results: ~37% of AI responses contain claims that are unsupported, exaggerated, or contradicted by evidence databases. Full methodology and raw data will be published with the final report.
37%
flagged (preliminary)
BenchmarkPlannedRouting·2026-07

Routing Cost Analysis: Claude Model Family

Planned benchmark: classify 10,000 real API queries by complexity, route within the Claude model family, and compare blended cost against flat Sonnet-for-everything baseline. Will include methodology, raw query classifications, and per-tier cost breakdowns.

Key finding: Study not yet complete. Hypothesis: 60%+ of queries are simple enough for Haiku, yielding 40-55% cost reduction.
WhitepaperPlannedRouting·2026-07

Multi-Model Panel Accuracy: Condorcet in Practice

Planned study: test error independence between Claude, GPT, Gemini, and DeepSeek across 500 factual questions. Will apply the Condorcet Jury Theorem framework to measure whether cross-provider panels actually improve accuracy.

Key finding: Study not yet complete. Theoretical prediction: 3-model panels should achieve ~96.6% accuracy from individual models averaging 85%, if error independence holds.
ArticlePlannedMemory·2026-06

NIGHTSHIFT: Autonomous Memory Maintenance Impact

Planned article: technical deep-dive into the 13 autonomous maintenance passes. Will measure recall precision before and after 30 days of NIGHTSHIFT operation, with controlled comparison against unmaintained memory graphs.

Key finding: Study not yet complete. Will measure recall precision, stale memory accumulation rate, and graph health metrics over time.
ArticlePlannedCode Quality·2026-07

ARGUS: Why Single-Number Quality Scores Fail

Planned article: how ARGUS decomposes code quality into 6 independent structural dimensions and why composite scores are meaningless. Will include case studies from real codebases showing hidden weaknesses that aggregate scores mask.

Key finding: Article not yet written. Will demonstrate with real examples how a file can score well overall while having critical dimensional weaknesses.
BenchmarkPlannedCode Quality·2026-07

Mutation Testing: What Your Test Suite Actually Catches

Planned benchmark: apply WHETSTONE mutation testing to open-source TypeScript projects and measure surviving mutation rates. Will publish full mutation logs, test suite analysis, and recommendations.

Key finding: Study not yet complete. Hypothesis: mutation survival rates will be significantly higher than line coverage metrics suggest.
Our methodology

How we run studies.

Every study published here follows the same principles. Were building trust in AI verification. That starts with being transparent about how we verify our own claims.

Reproducible
Every study includes the exact prompts, models, parameters, and database versions used. If you want to run the same test, you can.
Pre-registered hypotheses
We state what we expect to find before running the test. Post-hoc rationalization is the enemy of honest research.
Published negative results
When our systems don’t improve outcomes, we publish that too. Cherry-picking positive results would undermine the product we’re building.
Open to challenge
Every study includes the raw data and our methodology. If you find a flaw, we want to know. Contact research@asuriq.dev.
Your data stays yours
Prompts stay with your provider. We see analysis metadata only.
Keys never stored
One-way hash for authentication. Your credentials pass through. Never persist.
~2 second responses
Single-model cognitive tools return in about 2 seconds.
37 of 100 flagged
We ran 100 ChatGPT answers through verification. See the study →

Currently in friends-and-family beta. Built on peer-reviewed cognitive architecture.

Baars — Global Workspace TheoryACT-R — Memory Decay ModelWang et al. 2025 — Silent AgreementLi et al. EMNLP 2024 — Sparse Debate

The evidence behind the product.

Questions about our research? research@asuriq.dev