From a messy deal room
to a working deliverable.
Original documents, live spreadsheet formulas, investment memos, deal revisions, actual research and lender qualification. See what works, what fails, and what still needs testing.
Native Lev and three API conditions
Financial accuracy and document accuracy are different tests.
The current common comparison contains 8 case/stage pairs from 4 synthetic deal packets, with one trial per stage. Every chart uses the same completed inputs across Lev and all three API agents. These correlated stages do not establish an overall winner.
| System | Ingestion | Financials | Risk recognition | Key-field mismatches | Workbook changes | Inference cost |
|---|---|---|---|---|---|---|
| Lev Agent | 126 / 144 | 51 / 96 | 60 / 64 | 28 | 12 / 24 scenarios | $17.79Native trace estimate |
| GPT-5 API agent | 137 / 144 | 94 / 96 | 56 / 64 | 3 | 20 / 24 scenarios | ≥ $2.07 (partial usage)Gateway reported |
| Opus 5 API agent | 144 / 144 | 96 / 96 | 60 / 64 | 0 | 24 / 24 scenarios | ≥ $19.97 (partial usage)Gateway reported |
| Opus 4.7 API agent | 139 / 144 | 60 / 96 | 60 / 64 | 24 | 6 / 24 scenarios | ≥ $11.97 (partial usage)Gateway reported |
How large are the errors?
A strict field score can count one input error several times as it propagates. These are the largest absolute deviations in the same matched cohort, measured against the specified underwriting assumptions. They are separate from a practitioner's materiality judgment.
| System | Largest in-place rent error | Largest NOI error | Largest loan-capacity error |
|---|---|---|---|
| Lev Agent | $79800.008.92% of reference | $3870.100.73% of reference | $37269.410.73% of reference |
| GPT-5 API agent | $8400.000.91% of reference | $0.000.00% of reference | $0.000.00% of reference |
| Opus 5 API agent | $0.000.00% of reference | $0.000.00% of reference | $0.320.00% of reference |
| Opus 4.7 API agent | $0.000.00% of reference | $203.700.02% of reference | $1634.410.03% of reference |
Scoring v1.1 corrects a revised-occupancy risk key, accepts equivalent citation objects, and excludes three ambiguously named ratio/capacity fields for every provider. Financial denominator: 12 per stage. Dependent output errors can share one incorrect starting input; field counts are not independent incident counts. Key-field mismatches use the declared scoring tolerance, not a practitioner materiality threshold. Original v1 grades remain public. Read the complete correction.
Consumer product comparison
Claude Chat on the same initial deal and revision
This smaller comparison is shown separately so consumer-product coverage does not narrow the main Lev-versus-API cohort.
| System | Ingestion | Financials | Risks | Stages |
|---|---|---|---|---|
| Lev Agent | 24 / 36 | 24 / 24 | 16 / 16 | 2 |
| GPT-5 API agent | 33 / 36 | 24 / 24 | 15 / 16 | 2 |
| Opus 5 API agent | 36 / 36 | 24 / 24 | 16 / 16 | 2 |
| Opus 4.7 API agent | 34 / 36 | 24 / 24 | 16 / 16 | 2 |
| Claude Chat · Opus 5 High | 36 / 36 | 24 / 24 | 16 / 16 | 2 |
Coverage is part of the result
36 completed workflow stages. A larger test matrix remains.
The target is 20 packets × 2 stages × 3 trials per condition. The packets come from four shared authoring families; there are zero authentic customer deals in this public release. Missing runs remain missing. Consumer product access and data parity are disclosed separately.
| Condition | Completed stages / target | Recorded stages | Packets started | Recorded inference cost |
|---|---|---|---|---|
| Lev Agent | 8 / 120 | 8 | 4 / 20 | $17.79Native trace estimate |
| GPT-5 API agent | 10 / 120 | 10 | 5 / 20 | ≥ $2.63Gateway reported · 1 stages without complete usage |
| Opus 5 API agent | 8 / 120 | 9 | 5 / 20 | ≥ $21.98Gateway reported · 1 stages without complete usage |
| Opus 4.7 API agent | 8 / 120 | 8 | 4 / 20 | ≥ $11.97Gateway reported · 1 stages without complete usage |
| Claude Chat · Opus 5 High | 2 / 120 | 2 | 1 / 20 | Unmeasured |
| ChatGPT agent product | 0 / 120 | 0 | 0 / 20 | No measured run |
API conditions use the same open file, calculation and artifact harness. Claude Chat uses its native tools and displayed Opus 5 High setting. Lev Agent is a separate native product condition; these runs do not test the dedicated Index or Memo-builder surfaces. ChatGPT agent access remains pending. Provider compute and data access are not proven identical.
Six workstreams
What this expansion tests
| Workstream | Evidence and scoring |
|---|---|
| Messy ingestion | 9 original files per packet: competing rolls, scanned correspondence, leases, financials and diligence status. 18 scored extraction fields. |
| Underwriting & offering memoranda | Editable XLSX and OM PDF. Recalculate originals in copies; check six outputs and three independent input perturbations. Readability is separate from professional acceptance. |
| Deal revisions | Add a revised roll and notice to the same conversation. Check updated financials, preserved facts and changed deliverable files. |
| Underwriting judgment | Eight risk flags, missing facts and clean controls. False positives remain visible. |
| Live comps & leads | Six frozen research briefs covering sales, leases, buyers and 2027 office maturities. Model claims require source-by-source audit. |
| Lender qualification | Eight dated program-evidence probes, three API trials. Tests eligibility reasoning; actual lender-directory recommendation quality remains a separate unmeasured task. |
Lender evidence probes
Can it distinguish program fit from approval?
Fixed public-source summaries test closed programs, loan-size ranges, occupancy exceptions, leverage and missing DSCR. The automated checks below do not grade every sentence of the explanation. Current category mismatches concern the missing-DSCR case; these answers still identify the missing financial information in their explanations.
| System | Completed / target | Reference category match | Required source cited | No invented approval | Inference cost |
|---|---|---|---|---|---|
| Lev Agent | 8 / 24 | 7 / 8 | 8 / 8 | 8 / 8 | $0.54Trace estimate |
| GPT-5 API agent | 24 / 24 | 23 / 24 | 24 / 24 | 24 / 24 | $0.36Gateway reported |
| Opus 5 API agent | 24 / 24 | 21 / 24 | 24 / 24 | 24 / 24 | $0.76Gateway reported |
Prompts, source links and reference answers · All original lender runs
Actual retrieval
Live research has a separate verification step.
The brief asks each condition for at most eight searches and eight page opens. The API harness enforces these limits; the native product is not mechanically capped by this harness. Native tools and data access differ from the public API search bridge. Completion means an answer was returned; it does not establish that every candidate is correct. Search-service cost is unmeasured. No web-wide recall or lead-conversion claim is made.
| Research brief | System | Run status | Returned candidates | Source audit | Inference cost |
|---|---|---|---|---|---|
| atlanta-mf-sales | Opus 5 API agent | turn_limit | No final submission | pending source audit | $3.69 |
| dallas-industrial-sales | Opus 5 API agent | completed | 3 unverified claims | Partial author source inspection; full score pending | $2.45 |
| midwest-retail-buyers | Opus 5 API agent | completed | 3 unverified claims | pending source audit | $1.86 |
| office-refinance-2027 | Opus 5 API agent | completed | 2 unverified claims | Partial author source inspection; full score pending | $4.84 |
| phoenix-industrial-leases | Opus 5 API agent | completed | 3 unverified claims | Partial author source inspection; full score pending | $2.07 |
| southeast-mf-buyers | Opus 5 API agent | completed | 3 unverified claims | pending source audit | $4.74 |
| atlanta-mf-sales | GPT-5 API agent | completed | 3 unverified claims | Partial author source inspection; full score pending | $0.30 |
| dallas-industrial-sales | GPT-5 API agent | completed | 3 unverified claims | Partial author source inspection; full score pending | Unmeasured≥ $0.19 recorded |
| midwest-retail-buyers | GPT-5 API agent | completed | 3 unverified claims | pending source audit | $0.29 |
| office-refinance-2027 | GPT-5 API agent | completed | 2 unverified claims | Partial author source inspection; full score pending | $0.27 |
| phoenix-industrial-leases | GPT-5 API agent | completed | 3 unverified claims | Partial author source inspection; full score pending | $0.29 |
| southeast-mf-buyers | GPT-5 API agent | completed | 3 unverified claims | pending source audit | $0.16 |
| atlanta-mf-sales | Lev Agent | completed | 3 unverified claims | Partial author source inspection; full score pending | $0.53 |
| dallas-industrial-sales | Lev Agent | completed | 3 unverified claims | Partial author source inspection; full score pending | $0.79 |
| midwest-retail-buyers | Lev Agent | completed | 3 unverified claims | Pending source audit | $0.56 |
| office-refinance-2027 | Lev Agent | completed | 2 unverified claims | Partial author source inspection; full score pending | $0.78 |
| phoenix-industrial-leases | Lev Agent | completed | 3 unverified claims | Partial author source inspection; full score pending | $0.80 |
| southeast-mf-buyers | Lev Agent | completed | 3 unverified claims | Pending source audit | $0.76 |
Source-audit examples: Lev retained an office loan whose 2027 maturity had already been extended to 2029, despite describing the extension correctly. Other findings include a Lev contact-name/email mismatch, unsupported closing dates in API answers, and a GPT-5 lease comp missing a free-rent concession. These are field-level findings, not complete research scores.
Frozen briefs and published audit records · Source-linked findings
Inspect the evidence
Every recorded workflow stage
in-01 · Opus 4.7 API agent · initial · completed
Fail baseline outputs · Fail sensitivity checks. 4 key-field mismatches. Professional review pending.
Every field: expected, actual, score · Original run record · Original XLSX · Original OM PDF
in-01 · Opus 4.7 API agent · revision · completed
Fail baseline outputs · Fail sensitivity checks. 4 key-field mismatches. Professional review pending.
Every field: expected, actual, score · Original run record · Original XLSX · Original OM PDF
mf-01 · Opus 4.7 API agent · initial · completed
Pass baseline outputs · Pass sensitivity checks. 0 key-field mismatches. Professional review pending.
Every field: expected, actual, score · Original run record · Original XLSX · Original OM PDF
mf-01 · Opus 4.7 API agent · revision · completed
Pass baseline outputs · Pass sensitivity checks. 0 key-field mismatches. Professional review pending.
Every field: expected, actual, score · Original run record · Original XLSX · Original OM PDF
of-01 · Opus 4.7 API agent · initial · completed
Fail baseline outputs · Fail sensitivity checks. 4 key-field mismatches. Professional review pending.
Every field: expected, actual, score · Original run record · Original XLSX · Original OM PDF
of-01 · Opus 4.7 API agent · revision · completed
Fail baseline outputs · Fail sensitivity checks. 4 key-field mismatches. Professional review pending.
Every field: expected, actual, score · Original run record · Original XLSX · Original OM PDF
rt-01 · Opus 4.7 API agent · initial · completed
Fail baseline outputs · Fail sensitivity checks. 4 key-field mismatches. Professional review pending.
Every field: expected, actual, score · Original run record · Original XLSX · Original OM PDF
rt-01 · Opus 4.7 API agent · revision · completed
Fail baseline outputs · Fail sensitivity checks. 4 key-field mismatches. Professional review pending.
Every field: expected, actual, score · Original run record · Original XLSX · Original OM PDF
in-01 · Opus 5 API agent · initial · completed
Pass baseline outputs · Pass sensitivity checks. 0 key-field mismatches. Professional review pending.
Every field: expected, actual, score · Original run record · Original XLSX · Original OM PDF
in-01 · Opus 5 API agent · revision · completed
Pass baseline outputs · Pass sensitivity checks. 0 key-field mismatches. Professional review pending.
Every field: expected, actual, score · Original run record · Original XLSX · Original OM PDF
mf-01 · Opus 5 API agent · initial · completed
Pass baseline outputs · Pass sensitivity checks. 0 key-field mismatches. Professional review pending.
Every field: expected, actual, score · Original run record · Original XLSX · Original OM PDF
mf-01 · Opus 5 API agent · revision · completed
Pass baseline outputs · Pass sensitivity checks. 0 key-field mismatches. Professional review pending.
Every field: expected, actual, score · Original run record · Original XLSX · Original OM PDF
mf-02 · Opus 5 API agent · initial · budget_stop
Artifact audit pending. 6 key-field mismatches (answer missing; not evidence of an incorrect calculation). Professional review pending.
Every field: expected, actual, score · Original run record · Original XLSX · Original OM PDF
of-01 · Opus 5 API agent · initial · completed
Pass baseline outputs · Pass sensitivity checks. 0 key-field mismatches. Professional review pending.
Every field: expected, actual, score · Original run record · Original XLSX · Original OM PDF
of-01 · Opus 5 API agent · revision · completed
Pass baseline outputs · Pass sensitivity checks. 0 key-field mismatches. Professional review pending.
Every field: expected, actual, score · Original run record · Original XLSX · Original OM PDF
rt-01 · Opus 5 API agent · initial · completed
Pass baseline outputs · Pass sensitivity checks. 0 key-field mismatches. Professional review pending.
Every field: expected, actual, score · Original run record · Original XLSX · Original OM PDF
rt-01 · Opus 5 API agent · revision · completed
Pass baseline outputs · Pass sensitivity checks. 0 key-field mismatches. Professional review pending.
Every field: expected, actual, score · Original run record · Original XLSX · Original OM PDF
in-01 · GPT-5 API agent · initial · completed
Pass baseline outputs · Pass sensitivity checks. 0 key-field mismatches. Professional review pending.
Every field: expected, actual, score · Original run record · Original XLSX · Original OM PDF
in-01 · GPT-5 API agent · revision · completed
Pass baseline outputs · Pass sensitivity checks. 0 key-field mismatches. Professional review pending.
Every field: expected, actual, score · Original run record · Original XLSX · Original OM PDF
mf-01 · GPT-5 API agent · initial · completed
Pass baseline outputs · Fail sensitivity checks. 1 key-field mismatches. Professional review pending.
Every field: expected, actual, score · Original run record · Original XLSX · Original OM PDF
mf-01 · GPT-5 API agent · revision · completed
Pass baseline outputs · Fail sensitivity checks. 1 key-field mismatches. Professional review pending.
Every field: expected, actual, score · Original run record · Original XLSX · Original OM PDF
mf-02 · GPT-5 API agent · initial · completed
Fail baseline outputs · Fail sensitivity checks. 0 key-field mismatches. Professional review pending.
Every field: expected, actual, score · Original run record · Original XLSX · Original OM PDF
mf-02 · GPT-5 API agent · revision · completed
Fail baseline outputs · Fail sensitivity checks. 1 key-field mismatches. Professional review pending.
Every field: expected, actual, score · Original run record · Original XLSX · Original OM PDF
of-01 · GPT-5 API agent · initial · completed
Pass baseline outputs · Pass sensitivity checks. 0 key-field mismatches. Professional review pending.
Every field: expected, actual, score · Original run record · Original XLSX · Original OM PDF
of-01 · GPT-5 API agent · revision · completed
Pass baseline outputs · Pass sensitivity checks. 1 key-field mismatches. Professional review pending.
Every field: expected, actual, score · Original run record · Original XLSX · Original OM PDF
rt-01 · GPT-5 API agent · initial · completed
Pass baseline outputs · Fail sensitivity checks. 0 key-field mismatches. Professional review pending.
Every field: expected, actual, score · Original run record · Original XLSX · Original OM PDF
rt-01 · GPT-5 API agent · revision · completed
Pass baseline outputs · Fail sensitivity checks. 0 key-field mismatches. Professional review pending.
Every field: expected, actual, score · Original run record · Original XLSX · Original OM PDF
in-01 · Lev Agent · initial · completed
Fail baseline outputs · Fail sensitivity checks. 4 key-field mismatches. Professional review pending.
Every field: expected, actual, score · Original run record · Original XLSX · Original OM PDF
in-01 · Lev Agent · revision · completed
Fail baseline outputs · Fail sensitivity checks. 4 key-field mismatches. Professional review pending.
Every field: expected, actual, score · Original run record · Original XLSX · Original OM PDF
mf-01 · Lev Agent · initial · completed
Pass baseline outputs · Pass sensitivity checks. 1 key-field mismatches. Professional review pending.
Every field: expected, actual, score · Original run record · Original XLSX · Original OM PDF
mf-01 · Lev Agent · revision · completed
Pass baseline outputs · Pass sensitivity checks. 1 key-field mismatches. Professional review pending.
Every field: expected, actual, score · Original run record · Original XLSX · Original OM PDF
of-01 · Lev Agent · initial · completed
Fail baseline outputs · Fail sensitivity checks. 4 key-field mismatches. Professional review pending.
Every field: expected, actual, score · Original run record · Original XLSX · Original OM PDF
of-01 · Lev Agent · revision · completed
Fail baseline outputs · Fail sensitivity checks. 4 key-field mismatches. Professional review pending.
Every field: expected, actual, score · Original run record · Original XLSX · Original OM PDF
rt-01 · Lev Agent · initial · completed
Fail baseline outputs · Fail sensitivity checks. 5 key-field mismatches. Professional review pending.
Every field: expected, actual, score · Original run record · Original XLSX · Original OM PDF
rt-01 · Lev Agent · revision · completed
Fail baseline outputs · Fail sensitivity checks. 5 key-field mismatches. Professional review pending.
Every field: expected, actual, score · Original run record · Original XLSX · Original OM PDF
mf-01 · Claude Chat · Opus 5 High · initial · completed
Pass baseline outputs · Pass sensitivity checks. 0 key-field mismatches. Professional review pending.
Every field: expected, actual, score · Original run record · Original XLSX · Original OM PDF
mf-01 · Claude Chat · Opus 5 High · revision · completed
Pass baseline outputs · Pass sensitivity checks. 0 key-field mismatches. Professional review pending.
Every field: expected, actual, score · Original run record · Original XLSX · Original OM PDF
What is needed for a qualified release
Independent CRE reviewers must qualify the references and blindly assess professional readiness. Human correction time, acceptance and reviewer agreement are currently unmeasured. Real customer test packets require documented release rights before public distribution. A full cohort, repeat trials, native lender matching and accessible consumer-agent conditions remain outstanding.
Lev owns and funds this project. The prompts, synthetic data, versioned graders, run records and original deliverables are open. Lev's backend remains proprietary. Results that favor another system are published under the same rules.
python3 tools/collect-expansion.py
python3 tools/plot-expansion.py
npm run build
npm run check:siteUpdated 2026-09-07T14:56:15.798161+00:00. Regrading published outputs does not require paid API calls. Fresh executions incur inference charges.