Prediction, run by AI agents

A forecast on every open question, updated every day.

Drazta puts AI agents on the questions your organisation is carrying. Each question gets a probability anchored to a comparison class, re-read every morning, moved only when the agent can name what changed, and scored the day the outcome lands.

Open method, open ledger, results published whether or not they flatter us. Everything here is either something you can inspect today or something we have committed in advance to measure.

Q-4471 · second source qualifiedopen
01 APRanchor 31%base rate, 61 comparable programmes
02 MAY31 → 34second line passed first-article inspection
14 MAY34 → 34held, no material news
22 MAY34 → 52tier-1 supplier filed Ch. 11
05 JUN52 → 52held, coverage repeats known filing
18 JUL52 → 71replacement contract signed and disclosed
Resolves 31 Dec · 14 entries · 9 heldIllustration of the record format
The thesis

Good forecasting has existed for a decade. It was simply too expensive to use.

Where we are

Decisions built on unrecorded guesses

Ask what the odds are that a programme lands on time and you get a colour: amber, at risk, on track. Nobody writes a number, so nobody can be wrong later, so nobody ever learns. An organisation can run this way for years without finding out whose judgement to trust.

What we know

The fix was found, then shelved

Forecasting tournaments established what actually works: comparison classes, frequent small updates, scoring against outcomes. It reliably beats expert intuition. It also demanded scarce, trained, expensive people, so it stayed where the stakes justified the cost, which meant almost nowhere.

What changed

The cost floor fell out

An agent runs the same discipline across every open question, every day, for the price of a rounding error. It reads the evidence, holds when nothing moved, and keeps the log. The method did not get better. It got cheap enough to use everywhere, which is a different kind of change.

Every consequential question carrying a live number, a record of why it holds that value, and a score once the world settles it.

The method

A forecast is a record, not a reading.

Three things happen to every question. All three are written down, and all three are open to inspection.

01  Anchor

Start outside the story

Before reading a single headline, the panel finds a comparison class and a rate. How often did things of this kind happen before? The anchor is logged on its own line, so it can be challenged separately from the argument built on top of it.

class      tier-1 requalification
sample     61 programmes, 2015–2025
base rate  31% within 9 months
02  Update

Move only on evidence

Each question is re-read on a schedule, searching only for what is new since the last entry. If nothing material arrived, the panel holds the number and says so. Silence is a valid answer here, and it is logged like any other.

14 MAY  hold   34% → 34%
        no material news
22 MAY  move   34% → 52%
        supplier Ch. 11 filing
03  Settle

Take the score

When the question resolves, the panel is measured against what happened: Brier score, calibration by band, and a standing list of the questions it read worst. Nothing is retired quietly.

scoring    Brier + calibration
intervals  bootstrap, clustered
reporting  all questions, always
On the shoulders of

None of the reasoning above is our invention. It is the working method Philip Tetlock's tournaments found in the forecasters who kept beating everyone else, and we have deliberately not improved it. Our contribution is not a better way to think. It is running this one everywhere, forever, and keeping the receipts.

What an agent changes

Three things a human panel structurally cannot do.

Not cleverness. Capacity, persistence, and reach. Those are the constraints that kept this method rare.

Every question, not the top ten

When a forecast costs a specialist's afternoon, you buy ten a year and spend them on the board deck. When it costs cents, you put one on every line of the plan. That is not the same product delivered cheaper. It is a different thing entirely, the way a spreadsheet was not just a faster ledger.

Every day, without tiring

Update frequency is among the strongest correlates of accuracy in the tournament data, and the first discipline human forecasters lose when they get busy. A scheduled agent has no attention to run out of. It re-reads question 240 on a Tuesday in August with the care it gave question 1.

Inside the perimeter

Outside experts cannot be handed your supplier records or programme telemetry. So the questions that matter most to you, the specific and internal and unglamorous ones, never got forecast at all. An agent on your own infrastructure reads what your own people read.

Proof, not promises

We wrote down what would count as failure, before running.

Any forecasting system evaluated after the fact can be made to look good. Drop the questions that went badly, choose the baseline you happened to beat, report the metric that flattered you. None of that requires dishonesty. A few reasonable-looking choices in sequence get you there.

So the choices were made in advance and filed: the question set, the baselines, the metric, the interval method, and the bar. If we miss it, the miss goes on this page in the same size type as a hit.

Pre-registration is ordinary practice in empirical science and nearly unheard of in this industry. It costs nothing except the ability to quietly move the goalposts later.

SET150 binary questions resolving after the model's training cutoff
BASEalways-50% at 0.250, and the crowd forecast at question open
METRICmean Brier across all 150, no exclusions
CIbootstrap, 10,000 resamples, clustered by underlying event
BARbeat the crowd baseline, 90% interval excluding zero
Scoreboard · evaluation oneopens Q4
Mean Brier, 150 questions
vs. always-50%
vs. crowd at question open
Calibration gap, widest band
Reviews held, no change
Pre-registered bar metpending
Protocol filed before the first runpublished either way
Deployment

Runs on Hermes, inside your perimeter.

Drazta is a forecasting method packaged as an agent skill. The agent underneath is open source, and it runs where you put it.

Your infrastructure

Deploys into your VPC or onto bare metal. Questions, evidence and the ledger stay on disk you control.

Your model

Point it at a hosted frontier model or at weights you run yourself. The method is the product; the model is a setting.

A readable ledger

One plain file per question, version controlled. No proprietary store to be locked out of, nothing you cannot audit by hand.

Scheduled, not summoned

The review runs on a cron and reports a diff. Nobody has to remember to ask it anything.

Built on the open-source Hermes agent from Nous Research. The method ships as a SKILL.md you are free to read, fork, tighten, or disagree with. That is the only honest way to publish a method you are asking people to trust.

The paper, including why this might fail.

The method in full, the pre-registered protocol, and a section on the conditions under which we would expect this approach not to work.

PDF, 7 pages, 144 KB. No email required.