BENCHMARK IN DEVELOPMENT
What can you
hand off?
Explore how AI models handle work, from a small repair to a complete project. Start with the evidence behind each result.
First diagnostic cohort
18 attempts. Three coding tasks.
Full benchmark grading is still pending.
Results explorer
Diagnostic results. An app passing its tests does not establish full benchmark success. These results do not rank models.
No results for this tier yet.
Total elapsed time
All attempts included| Task | Requested model | App checks | Evidence |
|---|
Results could not be loaded. Please refresh or read the results in the GitHub repository.
READ THE MEASUREMENTS
A result you can inspect.
Each entry records a specific model, client, configuration, and task version. Subscription, API, and local access can all participate.
Read the cohort findings ↗- Correctness
- 17 of 18 apps passed independent task checks. Two app-passing attempts have execution review holds. Full transcript grading and human audit remain pending.
- Speed and cost
- Elapsed time includes failed attempts. Cost is a dated API-equivalent token estimate, not a subscription bill. Throughput and cost-efficiency scores remain unavailable.
- Trusted workload and efficiency
- Workload calibration and efficiency scoring are unfinished. A short repair does not establish reliability on a larger project. The 3D model comparison will follow validated measurements.
- External intelligence data
- Epoch AI supplies a separate capability measure. Our dated snapshot contains 266 observations. Model identity mapping remains pending. Epoch ECI and other intelligence indices are kept on separate scales. Original CSV. Data: CC BY 4.0.
Run it with your own setup.
The local pilot exports tasks and independently grades returned apps. Public uploads are not open yet.