MODEL TOPOGRAPHY

BENCHMARK IN DEVELOPMENT

What can you
hand off?

Explore how AI models handle work, from a small repair to a complete project. Start with the evidence behind each result.

First diagnostic cohort

18 attempts. Three coding tasks.
Full benchmark grading is still pending.

What these results establish ↓

Results explorer

Diagnostic results. An app passing its tests does not establish full benchmark success. These results do not rank models.

Download data ↓

Total elapsed time

All attempts included

Diagnostic task results. Use column buttons to change sorting.
TaskRequested modelApp checksEvidence

READ THE MEASUREMENTS

A result you can inspect.

Each entry records a specific model, client, configuration, and task version. Subscription, API, and local access can all participate.

Read the cohort findings ↗
Correctness
17 of 18 apps passed independent task checks. Two app-passing attempts have execution review holds. Full transcript grading and human audit remain pending.
Speed and cost
Elapsed time includes failed attempts. Cost is a dated API-equivalent token estimate, not a subscription bill. Throughput and cost-efficiency scores remain unavailable.
Trusted workload and efficiency
Workload calibration and efficiency scoring are unfinished. A short repair does not establish reliability on a larger project. The 3D model comparison will follow validated measurements.
External intelligence data
Epoch AI supplies a separate capability measure. Our dated snapshot contains 266 observations. Model identity mapping remains pending. Epoch ECI and other intelligence indices are kept on separate scales. Original CSV. Data: CC BY 4.0.

Run it with your own setup.

The local pilot exports tasks and independently grades returned apps. Public uploads are not open yet.

Get the local runner ↗