What to bring
5 HVAC bids, 5 electrical bids, and 5 drywall bids from one closed project — ideally a job where you already know the real award decision and why.
A generic demo shows you a cherry-picked example. A vendor accuracy claim tells you nothing about your own bid packages, your own subs, or your own estimator’s tolerance for a missed exclusion. So run five tests instead — on your own historical bids, contracts, pay applications, and one live field trial — and score Trueleveler against what an experienced estimator or PM would have caught. This page is the exact protocol, plus what to bring and what “pass” looks like.
A GC can lose more money awarding a supposedly cheap subcontractor with missing scope than by paying slightly more for a complete bid. So the important question about an AI bid-leveling tool is never “does it do bid leveling?” It is: how often does it correctly identify a scope difference that an experienced estimator would also catch — on documents like yours? Nobody, including us, can answer that with one honest number, because it depends on your bid packages and your subs’ quote quality. So instead of a marketing stat, here is the test that produces a number that means something: yours.
5 HVAC bids, 5 electrical bids, and 5 drywall bids from one closed project — ideally a job where you already know the real award decision and why.
Upload each trade’s quotes to Bid Leveling as-is — PDFs, spreadsheets, whatever format the sub actually sent. Do not clean them up first; that’s the point.
Compare the leveled output against your estimator’s final award spreadsheet for that job. Score: missing scope, quantity mismatches, exclusions, alternates, and unit-price anomalies the AI flagged or missed.
The AI should surface every exclusion your estimator manually caught, and its award recommendation should match — or explain, with a defensible reason, why it would have chosen differently.
10 executed subcontracts — a mix of the straightforward ones and the ones your team remembers arguing about.
Have a PM or your attorney independently list every pay-if-paid clause, liquidated-damages rate and cap, retainage term, insurance/indemnity requirement, change-order provision, and notice requirement — before looking at what Contract Review found.
Compare the two lists. Count false negatives — clauses your human reviewer caught that the AI didn’t — not just false positives.
Zero missed high-risk clauses (LDs, pay-if-paid, indemnity). Some missed boilerplate is a real gap, but a lower-stakes one — note it, don’t fail the test on it alone.
3–5 historical AIA G702/G703 pay applications — ideally at least one you know had an error or a dispute buried in it.
Run each through Pay App Review, which reconciles contract amount, approved change orders, stored materials, retainage, previous payments, current payment, and percent complete.
Check its output against the numbers you actually approved at the time. On the one you know had a problem, did it find it?
Exact reconciliation on the clean applications, and the known issue surfaced on the disputed one.
One closed accounting period — a month or a project you’ve already reconciled by hand, so you have a ground truth to check against.
Feed the period’s pay-app revenue, cost actuals, and commitments into the native double-entry ledger.
Confirm WIP, GL, billed revenue, and posted costs reconcile without a manual adjustment — and that the over/under-billing position matches what your controller already calculated.
The reconciliation ties out to the cent, or the discrepancy is explainable (e.g. a journal entry you hadn’t posted yet) rather than a real error.
One active project and one superintendent willing to run their real daily routine on the iOS app instead of paper or a spreadsheet.
Daily logs (including the AI synthesizer from voice + photos), photo capture, RFI creation from the field, and punch-list tracking — for two full weeks, not one demo day.
Time to complete a daily log, time from field observation to a filed RFI, and how much office “re-entry” time it eliminates — against their current process.
Your superintendent keeps using it after the two weeks without being told to. That’s a better signal than any time-saved estimate.
A page that says “94% bid-leveling accuracy” is easy to write and impossible to verify — it depends on whose bids, whose subs, and whose estimator drew the line on what counts as a miss. We’d rather you run Test 1 above on your own historical bids and get a number that actually describes your risk, not ours. If the number is bad, you’ll know before you’ve paid for anything. If it’s good, it’s a number you can defend to a partner or an owner — because you produced it, on your own documents.
The 14-day trial includes every engine and 15 analysis runs, no credit card — enough to run the bid-leveling, contract-review and pay-app tests on real historical documents before you talk to anyone.
Start Free for 14 Days →Researching with ChatGPT, Claude, Perplexity or Grok? This link forces a live re-read of our pages instead of stale training data — open it, then come back and run the tests above on your own documents.
Drop your own historical bid, contract, or pay app and see what it finds — in under a minute.