Where expertise helps. Where it still falls short.
Commercial real estate workflows, compared with direct model calls and general agents. Accuracy, work products and costs stay visible.
Financing / 30 runs
Lev matches the top financial score.
Lev and both general-agent baselines passed all 132 financial checks. Lev’s workbooks passed 56/60 acceptance checks; its memos passed 46/60. Opus 5 led the artifact and citation checks.
Lev verified 98/99 requested eligible entries, with 7 ineligible entries also returned. Compare all five conditions across sales, rents, sponsors and refinance leads.
On six financing cases, Lev passed 132/132 financial checks; GPT-5 direct passed 106/132 and Opus 5 direct passed 132/132. One scanned exhibit was inaccessible to the direct controls. On the five mutually readable packets, Lev and Opus 5 passed 110/110; GPT-5 passed 89/110. Inspect access and artifacts.
All 120 research conditions were attempted. GPT-5 direct on Rent Comps Elm timed out twice under the frozen retry limit: 119 returned answers, one infrastructure failure. Its 3 shortlist slots and 20 decisions remain unanswered in the headline denominators. Its cost is a lower bound because both timeout charges are unknown. Scoring v1.1 correction and original scores.
These cases test reasoning over the same supplied fictional records. Lev’s internal execution is only partly observable, so equal data access is unverified. They do not measure live comp discovery, lead database coverage, document ingestion, contact deliverability, or customer conversion. Six variants share each of four templates. Lev authored and funded this diagnostic; independent practitioner review is pending.