Skip to content
Vaticinus

Why we’re building Vaticinus

We want to build AI that forecasts as well as the best human forecasters. We haven’t shown that yet.

A forecast is useful when it helps you decide what to do before you know how things turn out. That might mean questioning a construction deadline, or deciding which experiment is worth funding.

We think AI can make the research behind these decisions cheaper. It can help work through evidence and calculate how an estimate changes under different assumptions. But a model can also invent a fact or give a convincing explanation for a bad estimate. Adding more model calls does not, by itself, make a forecast better.

What we’re working on

Vaticinus is the software around the model: it keeps the question fixed, checks the proposed forecast and computes the result. You can use it in chat or call it from your own software. The source code is public, so you can run it with your own model provider and change the parts you disagree with. Model calls still cost money, and each data source has its own terms.

Consider a reporter checking a promise to build more housing. The useful work is in the permits, financing and construction progress. Someone who knows the local planning process may spot an assumption a model misses. We want that person to be able to inspect the forecast and correct it. Making that kind of research affordable for a small newsroom is one reason we want to build this.

How we’ll know if it works

We compare the model on its own with the same model using our software, on the same questions and evidence. We record the cost and keep failed attempts in the results. Some tests have improved the scores; others have exposed mistakes our checks missed.

You can inspect those results in the benchmark record. Historical tests help us find problems now, but a model may have seen the answers during training. Predictions recorded before an event give us a different test. Both belong in the record, clearly separated.

Why the code is open

We want people outside Vaticinus to reproduce a result, find a failure, or try a better method without asking our permission. You should also be able to keep using the software without buying our hosted service.

The business we want to build is running this work for teams that need forecasts kept up to date. That would include maintaining data connections and supporting private workspaces. The reference method stays open; paid hosting covers the work of operating it.

If you know a subject well, you can help by finding where our forecasts get it wrong. The community page explains how to contribute questions, reproduce a benchmark, or work on the code.