Aibstracts
Screening validation

A published review already contains its own answer sheet

Anyone can demo well on a friendly example. So we test the hard way: published reviews, replayed as written, every included study accounted for.

The method and headline results are public below. The full report contains every denominator, per-review result and named miss.

Get the study-by-study validation report

  • Results for all ten published reviews
  • Every denominator, definition and named miss
  • Per-review failure analysis and evidence boundaries

PDF delivered immediately. Not a newsletter.

  1. 01 Published review
    INCLUDED STUDIES ANSWER KEY

    Its included studies are already known. They become the reference set.

  2. 02 Its own search
    Search strategy PRESERVED ("endovascular"[tiab] OR "angioplasty"[tiab]) AND "peripheral arterial disease"[MeSH] AND randomized[pt]

    Search concepts and eligibility preserved. Syntax was translated only where the destination database required it, and the query was never tuned against the known included studies.

  3. 03 Every study lands somewhere
    • High in the ranking
    • Low in the ranking
    • Missed

    The full report shows how many landed in each group, and identifies every miss.

The reference set is fixed before the run, so nothing can be graded generously after it. The published included-study set is the pre-specified reference standard, not assumed infallible.
Preliminary internal results
99.3% reachable-study retrieval recall
86.9% of retrieved reference studies surfaced in 8.5% of records

Reachable means indexed in a source Aibstracts searches. Embase, CENTRAL and Web of Science are not connected.

Search and title/abstract prioritization only. See the complete scope and limitations below.

The test set

Ten published reviews across eight specialties, three of them Cochrane, six published in 2025 or later.

OncologyGastroenterology ×3PsychiatryPaediatricsGeriatricsOrthopaedicsNeuroimagingNutrition
Which journals the test-set reviews came from

Cochrane Database of Systematic Reviews · JAMA Network Open · American Journal of Psychiatry · Archives of Disease in Childhood · Nutrition Reviews · Journal of Clinical Medicine · Osteoarthritis and Cartilage Open. Journal names describe where the test-set reviews were published. They are not an endorsement.

Not established by this evaluation

  • Reproduction of final included-study sets
  • Precision and agreement, both pending human adjudication
  • Run-to-run screening stability
  • Full-text eligibility performance
  • Performance on diagnostic, qualitative, or complex-intervention reviews
Guided benchmark pilot

The same test, on your own review

Replay a completed review with nothing live at risk, and judge the workflow on evidence your team already knows by heart.

No fee, and no commitment. This is how we show the work, not something we sell.

Opens our scheduler. Thirty minutes with the founder, no sales qualification call first.

Your live review
Carries on. Nothing pauses, nothing is touched.
A review you already finished
Replayed end to end, fully traced
You already know what this review concluded. That is what makes it a test.
How the pilot runs

Three steps, one objective

  1. 01

    Bring one completed review

    Its search strategy, eligibility criteria, and included-study list, as published. Non-confidential.

  2. 02

    We replay it, fully traced

    Retrieval, ranking, and rationales against your protocol, with provenance on every record.

  3. 03

    Read the results, miss by miss

    Coverage, retention, and every individual loss, walked through together with your methodologists.

What you walk away with

Two baselines, measured on your own review

The rows are fixed before the pilot starts. They are blank because the values come from your review, not ours.

For your methodologists

Evidence performance

Does it find what your reviewers found?

  • Retrieval coverage
  • Reference-study retention
  • Reading-list concentration
  • Every miss, itemized

You decide Whether the screening is methodologically acceptable to you.

For your review managers

Operational baseline

What would it cost you in reviewer time?

  • Reviewer minutes per record
  • Person-hours for the priority set
  • Time to surface known included studies
  • Time to first full-text handoff
  • Hours per protocol revision
  • QA and audit-preparation effort

You decide Whether the business case holds at your volume.

Pilot measures are commitments to measure, not results. Performance figures from our own validation, and their limits, travel together in the validation summary.

Results referenced on this page and in the validation summary are interim results from an internal validation study dated 29 July 2026. They are subject to revision and not intended for citation. The study is not independent, peer-reviewed, or externally validated.