Each thesis enters the ledger with a creation timestamp and a fixed horizon of 1 week, 1 month, 3 months or 6 months. Its resolution window begins at commit time. The ledger keeps the original wording permanently; resolved claims are never reworded, re-dated or removed.
Directional claims are scored using market-adjusted return: the realized stock return minus beta multiplied by SPY's return over the same window. Beta is calculated from roughly 120 trading days of prior daily returns. A LONG claim is a HIT above +2%, a MISS below −2%, and PARTIAL inside the band between them. Short claims are scored in the opposite direction.
Scoring v3 (from 2026-07-31). For tickers that ride a commodity or sector complex (uranium, rare earths, energy, precious metals, agriculture, semis, defense), the adjustment also subtracts the complex leg: realized return minus the market beta leg minus a second beta against a disclosed complex ETF (URA, REMX, XLE, GLD, MOO, SMH, ITA), both betas estimated jointly from prior daily returns. Reason: an audit of our own resolved claims showed the single-factor "stock-specific" residual was often mostly the complex moving — up to 90% of it for energy names. Only tickers where the second factor demonstrably explained movement use v3; every claim is stamped with the scoring version it resolved under, and claims resolved before this date keep their original version. Nothing was re-scored retroactively.
Headline counting v2 — net stances (from 2026-08-18). The research cycle sometimes re-affirms or revises a stance while an earlier claim's window is still open. Counting every such claim as an independent sample let a re-affirmation pad the sample and a change of mind add one guaranteed hit and one guaranteed miss — an audit of our own ledger found 94% of decided directional claims entangled this way. From this date the headline counts NET-STANCE VOTES: a claim votes only if no later directional claim on the same stock and horizon was committed inside its window. Superseded claims remain in the ledger, scored and public, forever — they simply no longer count twice. Individual claim scoring is unchanged; this changes only how the headline adds them up.
Confidence updates (from 2026-08-18). Claims are never edited — statement, direction, horizon and the committed probability are immutable. When new evidence arrives (only when the daily news scan maps fresh, material events onto the drivers behind an open claim), a timestamped, evidence-linked probability revision is APPENDED to the claim, forming a visible trajectory: at most one revision per claim per day, none in the final 10% of a claim's window, and a revision can never change the claim's direction. When trajectory calibration is scored, each day's standing probability counts for that day only (time-weighted Brier), so raising confidence the day before resolution moves almost nothing — early conviction carries the weight. Whether updating actually adds skill over the committed probabilities is itself a pre-registered, falsifiable metric, reported only once at least 30 updated claims have resolved. This rule was published before any updated claim resolved.
Range-bound claims are scored against the stock's own volatility. The band is 0.674·σ·√t, the median expected move. This gives a no-skill range-bound caller a long-run hit rate of about 50%, matching the baseline used for directional claims.
Catalyst scenarios are logged as explicit branches, such as “if X, then up” and “if Y, then down.” Once the event resolves, only the branch whose condition occurred is scored. Every other branch is marked VOID and excluded from the statistics. A dead branch is neither a hit nor a miss.
Hits, misses, partials, voids and pending claims all remain visible in the ledger. No resolved claim is unpublished.
Headline accuracy is always shown with a 95% Wilson confidence interval. Every claim also includes an explicit probability of scoring a HIT under these rules. Those probabilities are graded using a Brier score against the 0.25 no-edge baseline, which penalizes overconfidence quadratically.
Each research cycle feeds the resolved scorecard back into the model, including directional accuracy by setup, calibration of stated probabilities and recurring biases. The misses remain public and also become input for the next cycle.
Driver. A force the system tracks for a stock, such as a commodity, policy cycle or supply-chain link. Every driver has a permanent snake_case id, such as pc_insurance_pricing_cycle.
Activation. The driver's current score, from −10 for a strong negative push to +10 for a strong positive push. It combines the standing baseline with recent news pressure and is capped at ±10.
News pressure. The combined effect of recent headlines linked to a driver. Each event is weighted by its size and confidence, then fades with a seven-day half-life so old news gradually stops counting.
Baseline. The driver's standing influence before this week's news is added.
Tripwire. The daily news check. When news pressure on a driver moves sharply, the system can re-run the thesis before the next Monday cycle.
Escalation. A tripwire alert strong enough to trigger that off-cycle review.
Catalyst window. A dated event, such as earnings or a policy decision, with scenarios published before the event happens.
Priced in. The market may already have moved to reflect the thesis. A correct idea that reached the price before the call can still score as a miss.
The ledger is not yet large enough for a headline accuracy number to prove much. Until there are hundreds of resolved claims, the honest way to read the results is to look at the confidence interval as well as the point estimate. That is why both are shown.