Development preview / September 2026

GPT-5 and Opus 5, tested on a deal

Six completed API runs on one fictional case. Inspect the answers, calculations and mistakes. This is a development pilot, not a leaderboard.

6 of 6 answers scored successfully

Structured outputs and a tolerant parser keep Markdown, number formatting and equivalent labels from obscuring the financial results. No model answers were corrected.

ModelRunFinancial checksReference checksTimeCost, USD
GPT-5r1-125/3027/2746.6s$0.0509
Claude Opus 5r1-229/3027/2714.9s$0.0712
GPT-5r2-129/3027/2742.1s$0.0411
Claude Opus 5r2-228/3027/2715.2s$0.0712
GPT-5r3-129/3027/2752.8s$0.0487
Claude Opus 5r3-229/3027/2713.9s$0.0712

30 financial checks per run: 27 values, one conflict-set check and two consistency checks. Reference checks measure compliance with supplied source locations. These are correlated checks on one case.

What each system cost to run

SystemAttemptsMean cost, USDTotal cost, USD
GPT-53$0.0469$0.1407
Claude Opus 53$0.0712$0.2135
MiMo V2.53$0.0014$0.0041
MiMo V2.5 Pro3$0.0040$0.0119
Ling 3.0 Flash Sante3$0.0000$0.0000
Lev native product1 blocked + 1 recoveryNot recordedNot recorded

API costs are reported inference charges for the complete packet. Zero means the gateway reported zero, not that the full service is free. Missing cost is never treated as zero. Lev product credits and underlying inference dollars require separate, attributable measurements.

Earlier API cohort: all 9 runs and costs
ModelRunFinancial checksTimeCost, USD
MiMo V2.5r1-123/3037.9s$0.0010
MiMo V2.5 Pror1-229/30126.2s$0.0044
Ling 3.0 Flash Santer1-329/3018.4s$0.0000
MiMo V2.5r2-129/30116.9s$0.0017
MiMo V2.5 Pror2-228/30101.2s$0.0034
Ling 3.0 Flash Santer2-329/3028.1s$0.0000
MiMo V2.5r3-129/3065.4s$0.0014
MiMo V2.5 Pror3-229/3090.2s$0.0040
Ling 3.0 Flash Santer3-329/3019.0s$0.0000
Native product / Recovery observation

How Lev did on the same source material

The initial deal-context attempt could not read the files. Reattaching them directly to the chat restored access. The recovered answer reports 24 of 27 public-key fields; 24 agree within the key's tolerances. Three total-unit-by-type fields were not requested by its prompt.

Financial valuePublic keyGPT-5: runs 1 / 2 / 3Opus 5: runs 1 / 2 / 3Lev: recovery
Operating expenses96,00084,000 / 96,000 / 96,00096,000 / 96,000 / 96,00096,000
NOI150,000162,000 / 150,000 / 150,000150,000 / 150,000 / 150,000150,000
Unit occupancy, %7575 / 75 / 7575 / 75 / 7575
Annual in-place rent177,600177,600 / 177,600 / 177,600177,600 / 178,800 / 177,600177,600
DSCR loan limit1,714,2861,851,429 / 1,714,286 / 1,714,2861,714,286 / 1,714,286 / 1,714,2861,714,285.71
Maximum loan1,560,0001,560,000 / 1,560,000 / 1,560,0001,560,000 / 1,560,000 / 1,560,0001,560,000

Selected financial values, not an overall score. Amounts are USD unless marked as percentages. Loan limits allow $1 rounding tolerance. The full 27-field inventory and response evidence are linked below.

Different execution conditions: API models had no tools and used a structured prompt; Lev used product tools and a prose prompt, then recovered in the same conversation. This observation does not establish that Lev outperforms either model.

Cost: Full-run USD cost and attributable credits were not recorded. Work products: Workbook and OM generation remain untested. Review: The response's conflict terminology and “as-reported” NOI label still need assessment.

What went wrong

In its first run, GPT-5 reported $84,000 in operating expenses instead of $96,000. That raised NOI to $162,000 instead of $150,000 and affected two loan limits. Its other two runs got all 27 field values right.

In its second run, Opus 5 reported annual in-place rent of $178,800 instead of $177,600. Its other two runs got all 27 field values right.

Both models flagged historical income differing from current annualized rent as a conflict in every run. These cover different periods; the difference alone is not a discrepancy. This was the remaining financial failure in otherwise correct runs.

Execution and cost

Three repetitions each of openai/gpt-5 and anthropic/claude-opus-5 used the same packet and structured-output schema, with no tools or answer-key access in the request. Provider reasoning defaults were retained and are not equal compute. Gateway-reported inference cost for all six calls: $0.3542. These amounts exclude platform subscriptions, operator time, hosting, and product tooling. Cost is measured for the complete packet; it cannot be allocated to individual task stages from these records.

The grading correction

The original parser rejected nine earlier model responses because of formatting. The corrected scorer can evaluate all nine, while preserving their original responses and execution records. A second correction recognizes equivalent constraint labels such as LTV and ltv_limit. Both corrections apply consistently to every retained answer.

Read the correction and its timing · Inspect the earlier cohort after rescoring

What this pilot cannot establish

The case and answer key are public, practitioner review is pending, and the API inputs are text rather than original PDFs. Repeated calls on one packet do not establish performance on real deals, usable underwriting workbooks or finished OMs. Broader reviewed cases and artifact tests come next.