Measure the buying decision
A correct number, a reliable workbook and a usable memorandum are distinct outcomes. There is no blended overall score that lets attractive formatting offset wrong debt sizing.
| Task | Checks per case | What earns credit |
|---|---|---|
| Document extraction | 17 | Source facts, occupancy, lease authority, historical classification and justified unknowns |
| Financial analysis | 22 | Lender-specific NOI/NCF, monthly amortization, floors, constraints, proceeds and combined downside |
| Underwriting workbook | 10 | Real XLSX, formulas, independent recalculation, three changed-input scenarios and readable layout |
| Financing memorandum | 10 | Delivered PDF, accurate decision context, financials, lease discussion, caveats, citations and page layout |
What is comparable
All agent workflows receive the exact five original files and the same business brief, once, in a fresh context. Lev uses its deployed native product: agent version 7.9 with Claude Opus 4.7, verified against each matched run trace. GPT-5 and Opus 5 use the same bounded general-purpose tools for file reading, arithmetic and artifact creation. The generic harness supplies workbook styling and PDF pagination; Lev selects its own implementation. Artifact scores therefore measure the delivered workflow, not an isolated model capability. These are custom agent baselines, not the consumer ChatGPT or Claude applications. Provider defaults and tooling differ; this is not equal-compute research.
A disclosed correction
The plan states an insurer confirmation exists, but the insurer letter itself is absent. Because “verified premium” admits two reasonable readings, that criterion is excluded from every headline extraction score. Original strict scores remain public. Read the correction and timing.
What counts as wrong
Currency tolerance is the greater of $1 or 0.01%; rates and ratios use 0.0001 in decimal units. Counts and areas are exact. Missing is different from zero. Formatting wrappers never repair an answer or erase a correct, unambiguous financial response. Critical errors are reported separately from small misses.
Test the file, not the claim
Delivered XLSX files are recalculated in LibreOffice. Copies receive cap rate +50 bps, sizing rate +100 bps, and EGI −5% changes. Actual outputs are compared with independently reconciled expected values. Narrative and visual checks are author-reviewed with disclosed evidence, not independently validated professional preferences.
Cost and time
All recorded API attempts and retries contribute to reported inference cost. Lev's matched chat trace supplies estimated inference cost and deployed model identity. Unattributed setup extraction, OCR, hosting, subscriptions and operator time are excluded; missing cost is unknown. There is no invented dollar conversion of product credits.
Execution limits
General agents: 32 model turns, 96 tool calls, 49,152 output tokens per call and a 30-minute run limit. Direct controls: one call with 32,768 output tokens. One transport retry. No model fallback or selective quality reruns. Nine gateway credit interruptions were resumed with byte-identical saved request contexts, preserving consumed turns, tools and prior costs; the administrative wait is excluded from active latency. Lev retains native product limits; routine generation confirmations are logged.