Research
We test the frontier models on real enterprise work and publish what we find, so a team choosing a model for contract review or the monthly close has evidence to put in front of a board.
Research areas:Model benchmarksCost and routingAgent reliabilitySecurity and permissions
Model benchmarks
We run the frontier models head to head on real enterprise tasks, on real data, and publish which ones hold up and where.
Cost and routing
What a correct answer costs per task and per model, and when a cheaper model catches up, so each task runs on the right one.
Agent reliability
How agents fail in production, how to test them on last month's cases, and when an agent should stop and ask a person.
Security and permissions
What it takes to run agents inside a person's permissions, inside a company's cloud, with a record a controller will read.
Contract review: where fourteen models hold up, and where they don't
Fourteen models, one set of real commercial contracts, one question: which of them can be trusted with a first-pass review, and on which clause types.
Recent
Cost and routingCost per correct answer on enterprise tasks
Accuracy alone doesn't tell a finance team which model to run. This paper measures what a correct answer costs, per task, across the models in our benchmarks.
Model benchmarksThe monthly close, model versus model
Invoice matching, accrual drafting, and reconciliation checks, run model against model on a month of real finance data.
Agent reliabilityWhen agents should stop: measuring the flag rate finance teams actually want
An agent that guesses on the hard two percent costs more than one that stops and asks. This paper measures where the line should sit, per task.
Security and permissionsRunning agents inside a person's permissions: what breaks and what doesn't
Most agent platforms run on a service account that can reach everything. We ran ours inside the permissions of the person who started it and recorded what stopped working.
Publications
Get the next paper when it's out.
One email per paper. Nothing else.