What was tested
Each case supplies twenty fictional candidate records, source updates and an as-of date. Sales tasks distinguish executed sales from listings, portfolios and minority interests. Rent tasks normalize concessions, reimbursements and lease terms. Sponsor tasks resolve parents and current decision-makers. Refinance tasks distinguish confirmed maturity from options and missing balances. Four cases have no qualifying records.
How the comparison works
GPT-5 and Opus 5 each run directly and with identical generic calculation, CSV and XLSX tools. Lev runs in fresh chats in the authorized test account. Its traced model is Claude Opus 4.7 with Agent 7.9; it is not relabeled as Opus 5. Complete packet text was verified before submission; the native editor changes paragraph whitespace. Native traces were inspected for tool use and matched to those submissions. Recorded backend tools show no retrieval, but internal file-creation execution is absent from the available trace metadata. Native data access is therefore unverified for a strict common-corpus comparison; the native scores are a separately identified product condition.
What the scores mean
Verified shortlist yield: unique eligible entries with every required fact and supporting base/update IDs, divided by the smaller of the requested count and available eligible entities. Empty cases have a separate abstention score. Selection precision: unique eligible selections divided by every returned row. Screening success: all twenty accept/reject decisions, including duplicates before shortlist deduplication. Missing responses receive no delivery or screening credit; unanswered counts are separate from factual errors. Material findings: ineligible entries, duplicates, invented rows, false categorical values and numeric errors beyond the frozen tolerance.
Evidence checks require relevant source IDs and controlling updates. They do not imply expert review of every sentence. The artifact audit checks readable required format, matching identities/order, matching selected facts with explicit unknowns, and carried source IDs. The criteria were frozen before execution; artifact-mapping code was implemented during the audit after some outputs were visible. Files remain unmodified. Equivalent column labels are mapped uniformly; blank cells fail explicit-unknown annotation rather than being called false financial values. For generic CSV files, the audit attributes blank cells to the benchmark serializer when the original tool arguments prove the model supplied JSON null. This is a harness delivery issue, not a model reasoning failure.
Versioned scoring correction
Scoring v1.1 accepts numeric value/unit wrappers when the unit matches the brief, and accurate parent/SPV name annotations supported by the source relationship. The original fact schema did not restrict these representations. Original v1 scores, answers and frozen code remain public. This correction was introduced after reviewing results and applies uniformly to every condition. Read the full erratum and score deltas.
Cost and repeatability
API costs use gateway-reported inference charges. Two requests timed out without usage records, so GPT-5 direct and the API total are lower bounds; missing cost is never zero. Lev costs sum matched generation and embedding events, excluding duplicate run summaries; these are model-rate estimates, not customer prices. Native session naming, subscriptions, hosting, operator time and non-inference overhead are excluded. Provider defaults differ and are not equal compute. Only transport retries permitted by the frozen protocol were allowed.
python3 tools/collect-research.py
python3 tools/plot-research.py
npm run build
npm run check:siteReplaying scores is free. Fresh model executions require your own API credentials and incur costs. Figure export needs requirements-research.txt. Native trace import uses private captured evidence and is not required to replay published scores. The scored generic XLSX writer used the Codex-bundled artifact runtime; its open wrapper is published, but identical fresh file generation requires that runtime. Numerical regrading and inspection of the published files use public packages. Lev’s backend remains proprietary.
Limits of the evidence
All cases are original synthetic diagnostics. Six variants per task share a generator; these are four correlated task families, not 24 independent customer deals. They are public and not held out. Independent CRE reviewer qualification and real-world discovery validation have not occurred. No web-wide recall, verified-email deliverability, conversion, valuation or production-readiness claim follows.