Research

We test the frontier models on real enterprise work and publish what we find, so a team choosing a model for contract review or the monthly close has evidence to put in front of a board.

Research areas:Model benchmarksCost and routingAgent reliabilitySecurity and permissions

01

Model benchmarks

We run the frontier models head to head on real enterprise tasks, on real data, and publish which ones hold up and where.

02

Cost and routing

What a correct answer costs per task and per model, and when a cheaper model catches up, so each task runs on the right one.

03

Agent reliability

How agents fail in production, how to test them on last month's cases, and when an agent should stop and ask a person.

04

Security and permissions

What it takes to run agents inside a person's permissions, inside a company's cloud, with a record a controller will read.

Get the next paper when it's out.

One email per paper. Nothing else.