ModelMirror

Evaluating AI models on predictions the future will grade

Benchmarks leak. Once a test set is public, it ends up in the next model’s training data and the score stops meaning what it meant. ModelMirror builds evaluations that cannot leak by construction: models predict real-world events before they happen, the predictions are committed to a public, timestamped ledger, and the world supplies the answer key.

How we work

Projects

Contact

Issues on the relevant GitHub repository are the best channel. Updates are posted at @ModelMirrorAI.