Research v1.1 scoring · 120 attempted · 119 completed · Synthetic diagnostic

Comps and leads.
Every claim checked.

Lev Agent, GPT-5 and Opus 5 on sales comps, rent comps, sponsor qualification and refinance screening. Same source packets. Original answers and files.

All 120 research conditions were attempted. GPT-5 direct on Rent Comps Elm timed out twice under the frozen retry limit: 119 returned answers, one infrastructure failure. Its 3 shortlist slots and 20 decisions remain unanswered in the headline denominators. Its cost is a lower bound because both timeout charges are unknown. Scoring v1.1 correction and original scores.

These cases test reasoning over the same supplied fictional records. Lev’s internal execution is only partly observable, so equal data access is unverified. They do not measure live comp discovery, lead database coverage, document ingestion, contact deliverability, or customer conversion. Six variants share each of four templates. Lev authored and funded this diagnostic; independent practitioner review is pending.

What works

98/99 verified shortlist entries

Lev returned 99 eligible entries and got all 4 no-match cases right. Its selected facts passed 646/648 checks.

What needs work

7 ineligible entries also returned

Selection precision was 93.4%. The audit identifies evidence-cutoff mistakes and 9 consequential findings across 7 cases. A high yield alone does not mean a clean prospect list.

Native and general agents

Lev uses its product workflow. GPT-5 and Opus 5 share arithmetic and file tools. All received the same source packet. Lev’s recorded backend tools show no retrieval, but internal file-creation execution is not visible. Native data-access equivalence is unverified; read this as a product-condition comparison.

Verified shortlist yield by task for Lev Agent, GPT-5 with tools, and Opus 5 with tools
Download SVG · Download PNG · Source data

Lev versus models called directly

Direct controls answer the same packet without tools. Lev’s internal execution has partial audit coverage; equal data access is unverified. File production is outside the direct controls’ task. This compares the finished product condition with model calls, with that difference disclosed.

Verified shortlist yield by task for Lev Agent and direct GPT-5 and Opus 5 API calls
Download SVG · Download PNG · Source data

Correct inclusions are only half the job

All twenty records per case require an eligibility decision. Rejected, omitted and unanswered records stay in the denominator. The timed-out GPT-5 direct run contributes 20 unanswered decisions, not 20 attributed factual errors.

Candidate screening accuracy across all systems and task families
Download SVG · Download PNG · Source data
Cases with material shortlist errors by condition
Download SVG · Download PNG · Source data
System / conditionVerified shortlist yieldSelection precisionScreening successCases with material errorsInference cost
Lev Agent24/24 completed99.0%98 / 9993.4%99 / 10696.3%462 / 4807 / 24$12.61Trace estimate · 24/24 recorded
GPT-5 + tools24/24 completed99.0%98 / 9998.0%99 / 10199.0%475 / 4802 / 24$6.35Gateway reported · 24/24 recorded
Opus 5 + tools24/24 completed99.0%98 / 99100.0%99 / 9999.8%479 / 4801 / 24$9.38Gateway reported · 24/24 recorded
GPT-5 direct23/24 completed94.9%94 / 9996.9%95 / 9894.8%455 / 4803 / 23≥ $2.32Gateway reported · 23/24 recorded
Opus 5 direct24/24 completed99.0%98 / 99100.0%99 / 99100.0%480 / 4800 / 24$4.49Gateway reported · 24/24 recorded

What the work cost

At least $22.54 in reported API inference plus $12.61 in estimated Lev inference. All attempts are retained; two timed-out requests have unknown charges, so the API total is a lower bound. No product-credit conversion or implied retail price.

Inference cost for all 24 cases by condition
Download SVG · Download PNG · Source data

Do the files match the answer?

Four delivery-consistency checks per case. Truth is scored separately. The generic CSV writer turned some model-supplied JSON nulls into blank cells. Those delivery failures are attributed to the benchmark harness in each case audit; they are not model factual errors. Direct API controls have no file-generation requirement.

ConditionRequired formatMatching identities / orderFacts / explicit unknownsSource identifiers
Lev Agent21/2424/2417/2424/24
GPT-5 + tools24/2424/2415/2424/24
Opus 5 + tools24/2424/2416/2424/24

Read the supplementary file and layout review, including percentage-display and clipped-text issues outside these four checks.

Find the weak spots

Screening accuracy for each of 24 cases and five execution conditions
Download SVG · Download PNG · Source data
Open the case-by-case evidence