We test everything. Then we publish the results.
Studies, benchmarks, and technical articles from the ASURIQ team. Every claim we make on our product pages has a study behind it. Every number has a methodology. This is where we show our work.
100 Questions: How Often Does AI Get It Wrong?
100 questions across 10 domains. Run through GPT-4o mini, then verified by the ASURIQ engine at full depth. Methodology defined, test harness built, initial results being validated.
Routing Cost Analysis: Claude Model Family
Planned benchmark: classify 10,000 real API queries by complexity, route within the Claude model family, and compare blended cost against flat Sonnet-for-everything baseline. Will include methodology, raw query classifications, and per-tier cost breakdowns.
Multi-Model Panel Accuracy: Condorcet in Practice
Planned study: test error independence between Claude, GPT, Gemini, and DeepSeek across 500 factual questions. Will apply the Condorcet Jury Theorem framework to measure whether cross-provider panels actually improve accuracy.
NIGHTSHIFT: Autonomous Memory Maintenance Impact
Planned article: technical deep-dive into the 13 autonomous maintenance passes. Will measure recall precision before and after 30 days of NIGHTSHIFT operation, with controlled comparison against unmaintained memory graphs.
ARGUS: Why Single-Number Quality Scores Fail
Planned article: how ARGUS decomposes code quality into 6 independent structural dimensions and why composite scores are meaningless. Will include case studies from real codebases showing hidden weaknesses that aggregate scores mask.
Mutation Testing: What Your Test Suite Actually Catches
Planned benchmark: apply WHETSTONE mutation testing to open-source TypeScript projects and measure surviving mutation rates. Will publish full mutation logs, test suite analysis, and recommendations.
How we run studies.
Every study published here follows the same principles. We’re building trust in AI verification. That starts with being transparent about how we verify our own claims.
Currently in friends-and-family beta. Built on peer-reviewed cognitive architecture.
The evidence behind the product.
Questions about our research? research@asuriq.dev