Skip to content
Vaticinus

Open Forecasting Initiative

Forecasting
for everyone.

Ask a question in chat, inspect the forecast, or run the open-source code yourself. We publish our tests, including the failures.

We’re trying to build better AI forecasters.

That means testing whether our software improves a model’s predictions, and finding out when it makes them worse. The code and results are public so other people can check our work.

Why we build in the open

What is forecasting?

A forecast estimates how likely a specific event is to happen by a given date. You can check it against what actually happens.

An example, not a live Vaticinus forecast

Will U.S. inflation be below 2.5% in December 2027?

Measured by the Bureau of Labor Statistics’ all-items CPI, year over year, for that month.

60% chance

Roughly 6 out of 10 comparable cases. It is an estimate, not a guarantee. New evidence can change it.

The model, then the decision.

We work on both: estimating what may happen, then understanding what it would mean for an investment or a business.

  1. Forecasting models

    We’re developing models that estimate probabilities from evidence, and testing those estimates against outcomes. This is the forecasting layer.

  2. Tools for decisions

    The harness is the software around the model. It organizes research, compares scenarios, and tracks the assumptions behind an investment or business decision.

How do models compare with people?

ForecastBench compares AI forecasts with human forecasts, including superforecasters: people selected for their forecasting skill. Its tournament leaderboard is embedded below.

Live rankings from ForecastBench

You can also use “Open full leaderboard” above to read the rankings directly.

Swipe horizontally inside the frame to see all score columns.

Higher Brier Index scores are better. The highlighted “Superforecaster median forecast” row is the human reference. This is context for the field, not a claim about Vaticinus’ rank or investment returns.

Human forecasts were collected in July 2024. Later models answered different questions; ForecastBench adjusts for question difficulty. The full leaderboard includes uncertainty intervals and methodology.

The frame loads directly from ForecastBench. If it is unavailable, use “Open full leaderboard” above.

The harness has to earn its cost.

We test a model directly against the same model inside our forecasting pipeline, using the same questions and evidence cutoffs. Our early tests do not establish a reliable accuracy advantage. Historical replay is useful for finding failures, but a stated training cutoff alone cannot rule out leakage.

Read our tests, failures, and evaluation method

Help build it.

Forecasting should be something anyone can use, inspect, and improve. We’re making our work public so others can test it and contribute.

The public forecasting stack includes data collectors, probability tools, a model harness, dated records, and evaluation code. Run it with your own keys. Model-provider charges and source-data terms still apply.

Contribute on GitHub

Bring a question, a dataset, a quantitative baseline, or a failed forecast. World events, weather, science, macroeconomics, and prediction markets all belong here. Better predictions are not automatically profitable trades.

Where to start contributing

Work with us.

Open methods. A service built around your decisions. We’re building managed forecasting for teams that need continuous updates, data connections, private workspaces, and support.

Macro thesis

Which inflation, policy, or growth assumptions would change your view?

Investment research

Where could energy, manufacturing, or supply constraints change costs and company earnings?

Business planning

How would a change in demand, regulation, or input costs affect a planned investment?

The reference method stays open. The business is operating it reliably for a team—not charging for the right to inspect it. Tell us the question, the decision, and the time horizon.