150 attempted · 149 completed · Lev-funded diagnostic release

Where expertise helps.
Where it still falls short.

Commercial real estate workflows, compared with direct model calls and general agents. Accuracy, work products and costs stay visible.

Financing / 30 runs

Lev matches the top financial score.

Lev and both general-agent baselines passed all 132 financial checks. Lev’s workbooks passed 56/60 acceptance checks; its memos passed 46/60. Opus 5 led the artifact and citation checks.

Financing charts & evidence
Research / 119 runs

Strong shortlists. Screening gaps.

Lev verified 98/99 requested eligible entries, with 7 ineligible entries also returned. Compare all five conditions across sales, rents, sponsors and refinance leads.

Research charts & evidence

The direct-model comparison

On six financing cases, Lev passed 132/132 financial checks; GPT-5 direct passed 106/132 and Opus 5 direct passed 132/132. One scanned exhibit was inaccessible to the direct controls. On the five mutually readable packets, Lev and Opus 5 passed 110/110; GPT-5 passed 89/110. Inspect access and artifacts.

Research verified shortlist yield comparing Lev with direct models
Download SVG · Download PNG · Source data

Research performance at a glance

System / conditionVerified shortlist yieldSelection precisionScreening successCases with material errorsInference cost
Lev Agent24/24 completed99.0%98 / 9993.4%99 / 10696.3%462 / 4807 / 24$12.61Trace estimate · 24/24 recorded
GPT-5 + tools24/24 completed99.0%98 / 9998.0%99 / 10199.0%475 / 4802 / 24$6.35Gateway reported · 24/24 recorded
Opus 5 + tools24/24 completed99.0%98 / 99100.0%99 / 9999.8%479 / 4801 / 24$9.38Gateway reported · 24/24 recorded
GPT-5 direct23/24 completed94.9%94 / 9996.9%95 / 9894.8%455 / 4803 / 23≥ $2.32Gateway reported · 23/24 recorded
Opus 5 direct24/24 completed99.0%98 / 99100.0%99 / 99100.0%480 / 4800 / 24$4.49Gateway reported · 24/24 recorded

All 120 research conditions were attempted. GPT-5 direct on Rent Comps Elm timed out twice under the frozen retry limit: 119 returned answers, one infrastructure failure. Its 3 shortlist slots and 20 decisions remain unanswered in the headline denominators. Its cost is a lower bound because both timeout charges are unknown. Scoring v1.1 correction and original scores.

These cases test reasoning over the same supplied fictional records. Lev’s internal execution is only partly observable, so equal data access is unverified. They do not measure live comp discovery, lead database coverage, document ingestion, contact deliverability, or customer conversion. Six variants share each of four templates. Lev authored and funded this diagnostic; independent practitioner review is pending.