Every task. Every case.
Compare the native Lev product with two general-purpose agent baselines. Financial correctness and delivered artifact quality remain separate.
Can a reviewer find the supporting source?
Correct source facts and correct source locations are separate checks. Missing locations fail this measure even when the value is right.
The scale runs from 0 to 100%. Final analysis citations, including explicit table context and named source records, are checked against original files. Workbook and memo citations have their own acceptance checks. Review the exact interpretation.
The complete comparison
Each row is one fresh case attempt. Cost covers the entire workflow, not individual checks. Missing or failed deliveries remain in the denominator.
| Case | System | Source facts | Financials | Excel model | Memo | Inference cost | Time |
|---|---|---|---|---|---|---|---|
| Cedar Landing | Lev Agent | 17/17 | 22/22 | 9/10 | 8/10 | $2.167 | 7.7 min |
| Cedar Landing | GPT-5 agent | 11/17 | 22/22 | 9/10 | 9/10 | $0.312 | 3.9 min |
| Cedar Landing | Opus 5 agent | 17/17 | 22/22 | 10/10 | 10/10 | $3.020 | 8.7 min |
| Hawthorn Center | Lev Agent | 17/17 | 22/22 | 10/10 | 9/10 | $4.030 | 11.4 min |
| Hawthorn Center | GPT-5 agent | 17/17 | 22/22 | 9/10 | 9/10 | $0.261 | 3.3 min |
| Hawthorn Center | Opus 5 agent | 17/17 | 22/22 | 10/10 | 10/10 | $3.231 | 9.4 min |
| Juniper Flats | Lev Agent | 17/17 | 22/22 | 9/10 | 6/10 | $2.473 | 8.3 min |
| Juniper Flats | GPT-5 agent | 17/17 | 22/22 | 9/10 | 9/10 | $0.289 | 3.6 min |
| Juniper Flats | Opus 5 agent | 17/17 | 22/22 | 10/10 | 9/10 | $4.750 | 12.3 min |
| Market Row | Lev Agent | 17/17 | 22/22 | 9/10 | 8/10 | $2.626 | 8.2 min |
| Market Row | GPT-5 agent | 17/17 | 22/22 | 9/10 | 10/10 | $0.395 | 5.1 min |
| Market Row | Opus 5 agent | 17/17 | 22/22 | 10/10 | 7/10 | ≥ $3.053 | 8.9 min |
| Meadow Commerce | Lev Agent | 17/17 | 22/22 | 10/10 | 7/10 | $3.306 | 14.4 min |
| Meadow Commerce | GPT-5 agent | 17/17 | 22/22 | 9/10 | 9/10 | ≥ $0.288 | 3.8 min |
| Meadow Commerce | Opus 5 agent | 17/17 | 22/22 | 10/10 | 10/10 | ≥ $2.172 | 6.6 min |
| Stonebridge Logistics | Lev Agent | 17/17 | 22/22 | 9/10 | 8/10 | $2.853 | 9.2 min |
| Stonebridge Logistics | GPT-5 agent | 17/17 | 22/22 | 9/10 | 8/10 | ≥ $0.375 | 3.6 min |
| Stonebridge Logistics | Opus 5 agent | 17/17 | 22/22 | 10/10 | 9/10 | ≥ $4.551 | 11.4 min |
The cost of getting there
Measured inference for the complete workflow. A model call price is not a customer subscription price or a complete cost of delivery.
| System / track | Completed | Six-case inference | Per planned case | Measurement |
|---|---|---|---|---|
| Lev Agent | 6/6 | $17.455 | $2.909 | Matched native chat traces; estimated inference |
| GPT-5 agent | 6/6 | ≥ $1.921 | ≥ $0.320 | Gateway-reported inference; rejected-attempt billing unreported |
| Opus 5 agent | 6/6 | ≥ $20.777 | ≥ $3.463 | Gateway-reported inference; rejected-attempt billing unreported |
| GPT-5 agent · direct | 6/6 | ≥ $0.995 | ≥ $0.166 | Gateway-reported inference; rejected-attempt billing unreported |
| Opus 5 agent · direct | 6/6 | ≥ $3.050 | ≥ $0.508 | Gateway-reported inference; rejected-attempt billing unreported |
A ≥ amount is the known subtotal; unavailable cost is not silently zero. Separate adapter-development costs are excluded from scored runs. Administrative credit waiting is excluded from resumed active latency and preserved in execution records.
Direct API controls
Same business questions and source-text representations, with one API call and no agent tools. The industrial scan has no text layer in this control. Excel and memo delivery are outside its scope.
| Case | API model | Source facts | Financials | Inference cost | Time |
|---|---|---|---|---|---|
| Cedar Landing | GPT-5 | 17/17 | 17/22 | $0.165 | 147.6s |
| Cedar Landing | Claude Opus 5 | 17/17 | 22/22 | $0.523 | 205.5s |
| Hawthorn Center | GPT-5 | 17/17 | 20/22 | $0.137 | 92.6s |
| Hawthorn Center | Claude Opus 5 | 17/17 | 22/22 | $0.548 | 213.7s |
| Juniper Flats | GPT-5 | 17/17 | 17/22 | $0.188 | 146.6s |
| Juniper Flats | Claude Opus 5 | 17/17 | 22/22 | $0.542 | 212.7s |
| Market Row | GPT-5 | 17/17 | 17/22 | $0.169 | 127.0s |
| Market Row | Claude Opus 5 | 17/17 | 22/22 | $0.516 | 201.3s |
| Meadow Commerce | GPT-5 | 17/17 | 18/22 | ≥ $0.173 | 255.5s |
| Meadow Commerce | Claude Opus 5 | 17/17 | 22/22 | ≥ $0.443 | 168.3s |
| Stonebridge Logistics | GPT-5 | 16/17 | 17/22 | ≥ $0.162 | 123.4s |
| Stonebridge Logistics | Claude Opus 5 | 13/17 | 22/22 | ≥ $0.478 | 188.0s |
Earlier one-case development pilot and its corrections · Complete execution records