Most market opinions cannot be wrong. That is their defining feature and the reason they survive contact with reality indefinitely. “Bitcoin is going higher.” “The ETF bid is structural.” “This rally is unlevered.” Each of these can be held for a year without ever being contradicted, because none of them says what would have to happen for the holder to abandon it.

This guide is about the conversion step: how to take a view you actually hold and turn it into a claim that a calendar and a data source can settle without your participation. It is written from a position of some humility. This desk has published a public marker board since June 2026, and in the last week four of its expiring markers have been graded defective — two with bars copied straight out of the series they were supposed to test, one aimed at a ceiling the data had never touched, one aimed at a rounding boundary beneath a number that had not moved in two months. Everything below is what those failures taught, stated as construction rules.

Why bother

Three reasons, in ascending order of importance.

You stop confusing a mood with a forecast. When you have to name a number, a date and a source, you discover how much of a view is atmosphere. A large share of what people believe about a market dissolves the moment they are asked what reading would refute it.

You get a record you cannot renegotiate. Human memory is a defence lawyer. Without a written bar, a call that was “$90,000 by September” becomes “I said we’d go higher and we did.” A written claim with a deadline settles itself while you are not looking, which is the only way scoring stays honest.

You find out which of your reasons are load-bearing. This is the real payoff and it is almost never mentioned. When a claim fails, you have to look at the chain that produced it, and you frequently discover that a step you thought was central was doing nothing, or that a fact you have been quoting for weeks does not mean what you assumed. That is worth more than the score.

https://www.youtube.com/watch?v=pedNak4S9IE

The Long Now Foundation, “Superforecasting | Philip Tetlock,” published 26 April 2020. Tetlock’s research programme is the origin of most of what is now standard practice in resolvable forecasting — explicit questions, explicit resolution criteria, explicit scoring.

The anatomy of a claim that can be wrong

Six components. Miss any one and the claim leaks.

ComponentWhat it fixesBadGood
1. SeriesExactly which number, from which venue“The bitcoin price”“Binance BTCUSDT daily close, 00:00 UTC”
2. BarThe threshold, to a stated precision“Around $90k”“≥ $90,000.00”
3. DirectionWhether ties pass or fail“Above 90k”“≥ (a print of exactly $90,000.00 passes)”
4. WindowWhich observations are eligible“Soon”“Any daily close from 1–30 September 2026”
5. DeadlineWhen it stops being open“This cycle”“30 September 2026, 23:59 UTC”
6. SourceWho arbitrates, and the fallback“Whatever I see”“Binance klines endpoint; if unavailable, Coinbase BTC-USD”

Component 6 does more work than people expect. Bitcoin has no single price. On any given day the Binance perpetual, the Coinbase spot pair and the CME futures settlement will disagree, sometimes by hundreds of dollars around a round number. If your bar sits on a round number, the venue choice is the claim.

The five ways a claim quietly becomes untestable

These are failure modes, not errors of judgement. Each one produces a claim that looks falsifiable and is not. Each is illustrated with a marker this desk actually published and then had to grade against itself.

Failure 1: the bar copied from the series

You want to ask whether leverage returns. You pull the open-interest series, see that its highest reading was 111,988.29 BTC, and write the bar at 111,988. It feels rigorous — the number came from the data. It is the single worst thing you can do.

A bar taken from the series is not a threshold, it is a coordinate. You are no longer asking “does leverage return?” but “does this series exactly re-attain its own maximum inside a short window?” — a question nobody holds a prior view about, and one whose answer is dominated by the arbitrary depth of the window you happened to look at. In our case the series only retained thirty-one days, so the “maximum” was the maximum of a month, and the bar was set by whichever row the default page happened to start on.

The rule: the bar must be a number you would have chosen before seeing the series — a round figure, a policy threshold, a prior period’s value that you name and justify, or zero. If you cannot explain the bar without pointing at the chart, it is a coordinate.

Failure 2: the ceiling that has never been touched

We once wrote: “Will any funding settlement exceed 0.0100% by 28 August?” It read as a live question about whether leverage would get crowded. Then we pulled the full series — five hundred settlements, the endpoint’s maximum depth, covering five and a half months and a 43% round trip in price — and found the answer had been zero every single time. Not zero recently. Zero across every observation available.

That claim was not a bet. It was a near-certain fail wearing the costume of one, and it would have “resolved correctly” while teaching nothing. The tell is simple and mechanical: before writing a threshold claim, count how many times the series has crossed that threshold in the deepest window you can pull. If the answer is zero, you are not forecasting, you are restating a structural property of the instrument.

The reverse failure exists too and is more flattering: a bar the series clears most weeks, dressed up as a call. Both are the same mistake — a bar chosen without checking the base rate.

Failure 3: the bet on rounding noise

We wrote a marker that core PCE inflation would come in at or below 3.2% for July. It failed at 3.3%, by the smallest increment the series reports. Fair enough — except that core PCE had printed 3.3% in June, and 3.3% the month before that. The annual rate had not moved for two months.

So the marker was never a view about inflation. It was a coin flip on whether rounding and revision noise would push a stationary number across a boundary. If your bar sits within one reporting increment of the current value of a series that is not moving, you have written a lottery ticket. Put the bar somewhere the series would have to actually travel to reach, or write the claim about the monthly change rather than the level.

Failure 4: the window you never checked

This is the failure underneath the other failures. Almost every market data endpoint has a default page size and a maximum depth, and they are rarely the same. Binance’s funding endpoint defaults to thirty rows and caps at five hundred. Its daily open-interest endpoint retains thirty-one rows and will not serve a thirty-second no matter what you ask for. Exchange APIs, macro databases and charting tools all do some version of this.

We spent four consecutive days publishing that funding had “never risen across this entire advance,” having counted a streak inside a window of twenty-four settlements. The true figure was five hundred — every row the endpoint serves — spanning a period in which price fell 30% and then rose 43%. The observation was much larger than we had claimed and, for exactly that reason, it was not evidence about August at all.

What we publishedWhat the full pull showed
24 consecutive settlements at or below baseline500 — every row available
Window: 21–26 August 202614 March – 28 August 2026
Across ~$18,700 of price rangeAcross a 43% round trip, in both directions
Inference: this rally is unleveredInference invalid — the property is not specific to the rally

The rule: every streak, series high, series low and “consecutive” count ships with the depth of the pull that produced it, and the pull goes to the endpoint’s maximum before the sentence gets written. A count without its window has a shelf life nobody told you about — we published “36 occurrences, 7.2%” one morning and the identical query returned 37, 7.4% the next, because a row rolled off the back.

Failure 5: the horizon shorter than the mechanism

If your reasoning runs through ETF allocation committees, corporate treasury policy, or a rulemaking process, a two-week deadline is not testing your reasoning. It is testing noise. Match the deadline to the slowest step in the chain that produces the outcome. If the mechanism takes a quarter and you have a fortnight, either lengthen the deadline or write a claim about an intermediate step that genuinely moves on your timescale.

https://www.youtube.com/watch?v=SAzTP2A634g

BBC Global, “‘Superforecasting’: The people that predict the future – BBC REEL,” published 18 January 2023. A short introduction to the Good Judgment work on resolvable questions and calibration.

The pre-flight checklist

Seven questions. Any “no” sends the claim back to the drawing board. This takes about four minutes and would have caught all four of our defective markers.

#QuestionWhat a failure looks like
1Can I name the exact reading that makes this false?You can only describe a vibe
2Did I pull the source at maximum depth before setting the bar?You used the default page
3How many times has the series crossed this bar in that full window?Zero, or every week
4Is the bar a number I would have picked without seeing the chart?The bar equals a historical reading
5Is the bar more than one reporting increment from the current value?You are betting on rounding
6Does the deadline exceed the slowest step in my mechanism?Two weeks for a quarterly process
7If the data source disappears, what settles it?No named fallback

Grading: three rules that make the record worth keeping

Grade the design as well as the outcome. A claim can pass on a defective bar and fail on a sound one. If you only audit your losses you will learn half of what is available, and you will systematically preserve the flattering mistakes. Publish the design verdict alongside the result: passed, badly constructed is a legitimate grade and often the most informative one.

Retire, do not roll. When a claim fails because its bar was defective, resist extending the deadline. Rolling carries the defect forward and buys you an outcome that will be misread as vindication. Write a new claim that asks the same question with a defensible bar, and say in print that the old one was retired for construction rather than for being early.

Publish partial settlements as partial. If a claim settles at a market close that happens after you write, say so rather than grading a morning reading as a close. “Failing, needs a 15.1 basis-point move at today’s close” is honest. “Failed” is not, until the close happens.

Worked example: three claims, built live

Start from a real view: “Institutional demand for bitcoin has genuinely turned and August was not a one-off.” That is untestable as written. Here is the conversion, with the bar justification stated before the outcome is known — which is the whole point.

ClaimSeriesBarDeadlineWhy this bar
1US spot bitcoin ETF net monthly flow (Farside aggregate)≥ $0 for September 202630 Sep 2026Zero is a natural boundary, not a coordinate. Base rate checked: 2026 has produced four negative months and four positive, so this is genuinely live
2Binance BTCUSDT daily close, 00:00 UTC30 Sep close > 31 Aug close30 Sep 2026The bar is a fact that does not exist yet and cannot be reverse-engineered from history
3Binance BTCUSDT daily open interest≥ 115,000 BTC on any day30 Sep 2026A round number 2.69% above the retained series maximum of 111,988.29. Declared in advance: this is a level the series has never printed, so the claim is hard by construction, and we are not going to pretend otherwise when it fails

Note what claim 3 does that our defective markers did not. The bar is above every observation in the record — which is failure mode 2 — and we say so, with the margin, before the deadline. A bar that has never been crossed is not automatically illegitimate; it is only illegitimate when it is presented as an even question. Disclosing the asymmetry converts a fake bet into an honest long shot.

What this does not do

Three limits, because a method sold without its limits is the same failure in a different costume.

Resolvability is not insight. You can write beautifully constructed claims about things that do not matter. The hardest part of this process is not the construction, it is choosing a question whose answer would change what you do.

A pass does not validate the reasoning. Markets settle claims for reasons unrelated to the chain of logic that produced them. Our C1 marker passed this week; the tape cleared its bar for reasons that had nothing to do with why the bar was written. Track the reasoning separately from the score.

Small samples say almost nothing. A dozen resolved claims is not a track record, it is anecdote with dates attached. Calibration — whether the things you called 70% likely happen about 70% of the time — needs scores of observations before it means anything. Until then the value is entirely in the process, not the hit rate.

That is the honest pitch for this whole exercise. You will not become a better forecaster quickly. You will, almost immediately, become much harder to fool by your own commentary — including the parts of it that turned out to be true.

Disclaimer. This guide is educational. It is not investment advice, and none of the example claims in it are recommendations. Bitcoin and other digital assets are volatile and you can lose everything you put in. Writing a testable claim does not make it a good claim, and being right about a testable claim does not mean the reasoning was sound. Do your own research and, if you need advice, speak to a regulated professional.