98/99 verified shortlist entries
Lev returned 99 eligible entries and got all 4 no-match cases right. Its selected facts passed 646/648 checks.
Lev Agent, GPT-5 and Opus 5 on sales comps, rent comps, sponsor qualification and refinance screening. Same source packets. Original answers and files.
All 120 research conditions were attempted. GPT-5 direct on Rent Comps Elm timed out twice under the frozen retry limit: 119 returned answers, one infrastructure failure. Its 3 shortlist slots and 20 decisions remain unanswered in the headline denominators. Its cost is a lower bound because both timeout charges are unknown. Scoring v1.1 correction and original scores.
These cases test reasoning over the same supplied fictional records. Lev’s internal execution is only partly observable, so equal data access is unverified. They do not measure live comp discovery, lead database coverage, document ingestion, contact deliverability, or customer conversion. Six variants share each of four templates. Lev authored and funded this diagnostic; independent practitioner review is pending.
Lev returned 99 eligible entries and got all 4 no-match cases right. Its selected facts passed 646/648 checks.
Selection precision was 93.4%. The audit identifies evidence-cutoff mistakes and 9 consequential findings across 7 cases. A high yield alone does not mean a clean prospect list.
Lev uses its product workflow. GPT-5 and Opus 5 share arithmetic and file tools. All received the same source packet. Lev’s recorded backend tools show no retrieval, but internal file-creation execution is not visible. Native data-access equivalence is unverified; read this as a product-condition comparison.
Direct controls answer the same packet without tools. Lev’s internal execution has partial audit coverage; equal data access is unverified. File production is outside the direct controls’ task. This compares the finished product condition with model calls, with that difference disclosed.
All twenty records per case require an eligibility decision. Rejected, omitted and unanswered records stay in the denominator. The timed-out GPT-5 direct run contributes 20 unanswered decisions, not 20 attributed factual errors.
| System / condition | Verified shortlist yield | Selection precision | Screening success | Cases with material errors | Inference cost |
|---|---|---|---|---|---|
| Lev Agent24/24 completed | 99.0%98 / 99 | 93.4%99 / 106 | 96.3%462 / 480 | 7 / 24 | $12.61Trace estimate · 24/24 recorded |
| GPT-5 + tools24/24 completed | 99.0%98 / 99 | 98.0%99 / 101 | 99.0%475 / 480 | 2 / 24 | $6.35Gateway reported · 24/24 recorded |
| Opus 5 + tools24/24 completed | 99.0%98 / 99 | 100.0%99 / 99 | 99.8%479 / 480 | 1 / 24 | $9.38Gateway reported · 24/24 recorded |
| GPT-5 direct23/24 completed | 94.9%94 / 99 | 96.9%95 / 98 | 94.8%455 / 480 | 3 / 23 | ≥ $2.32Gateway reported · 23/24 recorded |
| Opus 5 direct24/24 completed | 99.0%98 / 99 | 100.0%99 / 99 | 100.0%480 / 480 | 0 / 24 | $4.49Gateway reported · 24/24 recorded |
At least $22.54 in reported API inference plus $12.61 in estimated Lev inference. All attempts are retained; two timed-out requests have unknown charges, so the API total is a lower bound. No product-credit conversion or implied retail price.
Four delivery-consistency checks per case. Truth is scored separately. The generic CSV writer turned some model-supplied JSON nulls into blank cells. Those delivery failures are attributed to the benchmark harness in each case audit; they are not model factual errors. Direct API controls have no file-generation requirement.
| Condition | Required format | Matching identities / order | Facts / explicit unknowns | Source identifiers |
|---|---|---|---|---|
| Lev Agent | 21/24 | 24/24 | 17/24 | 24/24 |
| GPT-5 + tools | 24/24 | 24/24 | 15/24 | 24/24 |
| Opus 5 + tools | 24/24 | 24/24 | 16/24 | 24/24 |
Read the supplementary file and layout review, including percentage-display and clipped-text issues outside these four checks.