Model release

8 minute read

Tarot-Draw

Three forecasting models trained on the same 81,870 questions. A controlled look at what model size changes, and what it does not.

Layered mountain ridges receding into a blue-grey horizon
Tarot-Draw estimates a probability from the text of a question and its resolution date. It does not retrieve news or observe a market.
Family
1.7B, 4B, 8B
Task
Binary forecasting
Output
Probability in [0, 1]
Recommended
Tarot-Draw 4B

What Tarot-Draw is

Tarot-Draw is a family of probability estimators for binary questions about future events.

Give a model a question, its resolution criteria, a scheduled resolution date, and an as_of date. It returns one number between zero and one. The models are intended to provide a calibrated prior when current evidence is sparse, a market does not exist, or a forecaster needs a consistent place to begin.

This is deliberately narrower than a general forecasting agent. There is no retrieval layer, news feed, market price, or generated explanation. The release isolates a single capability so that the effect of model scale can be measured cleanly.

Input

  • Question
  • Resolution criteria
  • Resolution date
  • As-of date

Output

0.73 estimated probability

One variable: model size

The 1.7B, 4B, and 8B models share the same corpus, targets, seed, learning rate, and evaluation footing.

Holding those choices fixed turns the family into a controlled experiment. The central question is not simply whether a larger model scores better. It is which part of forecast quality changes as capacity increases.

Calibration stayed flat. Resolution did not.

All three models learn when to be cautious. The smallest model is worse at separating easier questions from harder ones. In Murphy terms, its loss is in resolution rather than calibration.

The 4B and 8B results are statistically indistinguishable. The 1.7B trails the 4B by 0.0119 Brier, with a paired 95% interval of +0.0080 to +0.0158.

Development split, n = 3,000. Lower Brier and calibration are better; higher resolution is better.
Model Dev Brier Calibration Resolution
Tarot-Draw 8B 0.1674 0.0010 0.0718
Tarot-Draw 1.7B 0.1803 0.0004 0.0593

Question-level intervals do not include run-to-run training variance. Differences below roughly 0.008 Brier should be treated as unresolved.

Choosing a model

Start with the 4B. It matches the 8B within the resolution of this experiment and requires a smaller deployment footprint.

The other sizes remain useful when hardware is the binding constraint or when reproducing the scaling comparison matters more than choosing a default.

Tarot-Draw 4B

Recommended

General use and most deployments

Tarot-Draw 8B

Largest

Capacity-rich serving and scale studies

Tarot-Draw 1.7B

Smallest

Laptops, CPUs, and tight memory budgets

A result you can inspect

The forecasts, split metadata, metric code, and verification script ship with each release.

The published forward split contains 277 questions that resolved after a freeze date committed in public before it passed. The 4B scored 0.1831 Brier, the 8B scored 0.1893, and the 1.7B scored 0.2030. These are clean point estimates, but the split is too small to support a ranking between the 4B and 8B.

Confidence intervals use 10,000 paired, question-clustered bootstrap samples. Recomputing every headline number requires no model, GPU, or network access.

Verify the published results

uv run python verify.py

Where the prior stops

Tarot-Draw is a starting estimate, not an oracle. Its calibration should be measured again when the venue or question distribution changes.

  1. Current information is absent.

    The models cannot see breaking news, private evidence, or live market prices. A forecast that depends on them needs an external update.

  2. Mechanical questions remain difficult.

    Thresholds, counts, scores, and questions requiring precise arithmetic sit outside the strengths of a judgment prior.

  3. Very short horizons are weak.

    At horizons under seven days, the tested small model does not beat a market crowd at any point in a question's life.

  4. Some domains are underrepresented.

    Sports and obscure local elections are measured weaknesses in the present release.

Models, code, and terms

Adapter weights, forecasts, and evaluation metadata are released under CC BY-NC 4.0. Verification and inference code are Apache 2.0. The training corpus is documented, not redistributed.

Version
1.0
Published
29 August 2026
Authors
Arya Somu; Bruce Nshuti Hirwa
Publisher
Laplace Research

Cite this release

@misc{somu2026tarotdraw,
  title     = {Tarot-Draw},
  author    = {Somu, Arya and Hirwa, Bruce Nshuti},
  year      = {2026},
  publisher = {Laplace Research},
  url       = {https://laplaceresearch.org/tarot-draw/}
}