CoinVerdict
中文 Scoreboard

How CoinVerdict works

Who

CoinVerdict is an AI analyst desk with a human editor. Every day, a multi-model pipeline deliberates over 100 coins — technicals, sentiment, funding pressure, and an exclusive settled record of KOL calls — and renders one research verdict per coin: overweight, neutral, or underweight. A human (the founder) owns editorial quality, coverage decisions, and this methodology.

How verdicts are made

Each verdict carries four independently-graded dimensions. Grades are per-dimension scales and are never averaged. Prices are a median snapshot across five exchanges (Binance perp/spot, OKX, Bybit, Kraken). Funding pressure comes from perpetual futures. The KOL seat is graded from VeraMind's settled ledger of public calls — facts, not vibes.

How settlement works

Rule settle_v1: a verdict from day T is settled at T+7 days against BTC (BTC itself vs USD; grace to T+9 if a snapshot is missing). The scoreboard counts directional calls only: overweight wins if the coin beats BTC by more than 3%, loses if it lags by more than 3% — inside the band is a push. Neutral verdicts are no-calls and sit outside the win rate.

All settlement prices come from the archived reports themselves — both ends of every settlement are on the record. Settled rows are never edited; rule changes get a new version and never rewrite history. Missing data settles as VOID and is excluded from the win rate.

Why

Crypto produces a million predictions a day and settles roughly zero. We publish our record — wins, losses, and pushes — because a research shop that won't grade itself is asking you to take vibes on faith. Time keeps the receipts.

How the evidence list is read

Coin pages carry a structured evidence list (rule coin_evidence_v1), derived by fixed code from the numbers already on the record — never by a model grading its own prose. Trend items are mechanical: price above a moving average is bull evidence, below is bear; SMA50 above SMA200 is bullish structure. Distance to the 60-day high only counts beyond thresholds (within 2% or a fresh high → bull; more than 10% below → bear), and the 30-day move only counts beyond ±10% — the middle zone is silent rather than guessed. KOL calls count by simple majority of settled-ledger entries.

Sentiment and positioning (Fear & Greed, long/short ratio, taker flow, funding) are listed under Market state without a bull or bear label: contrarian and momentum readings of the same number disagree, so we state the fact and leave the reading to you. Missing data simply drops an item — the list never pads itself. Because evidence is computed at render time from archived numbers, the same rules apply retroactively to every archived report.

The grading scales in detail

Technical and Sentiment share one scale (evidence_rule_v1): an analyst lists independent evidence families from the data pack, and code converts the count into a letter — three or more families with no conflict is A+, two is B, one is D; conflicting evidence caps the grade at C, and insufficient information yields no grade at all. Funding uses its own scale (funding_rule_v1) measuring pressure, not direction: the absolute 8-hour normalized rate maps to letters at 0.5 / 1.5 / 3 / 5 / 10 basis points — under 0.5 bp is A+ (calm), beyond 10 bp is F (extreme crowding). Who pays — longs or shorts — is reported separately.

KOL Consensus (kol_rule_v1) grades the on-record crowd: only authors with at least 20 settled calls qualify. The majority share sets the base letter — 85% or more is A+, then 80% A, 70% B, 60% C, 55% D — but the number of qualified authors caps it: 5-7 authors can reach at most C, 8-11 at most B, 12-19 at most A, and 20 or more unlock A+. With fewer than 5 qualified authors we show the direction and withhold the grade while the record builds. The four letters measure different things — signal strength, crowd volume, rate pressure, consensus quality — and are never averaged into one score.

Reproducibility — why two runs can differ

The pipeline is LLM-driven, so two runs on the same coin and date can produce different prose and, occasionally, a different call. That is a property of language models, not a malfunction — sampling is non-deterministic, and reasoning models vary the most because their internal deliberation is itself sampled.

What we pin down: every published report stores its model name, prompt version and generation date; the price and indicator window is fixed to the analysis date; and once a verdict is settled, the settlement is computed from recorded prices and never re-run. What stays live: news and social inputs reflect the moment of generation.

The discipline that matters is not byte-identical output — it is that every verdict, once published, is frozen, timestamped, settled by objective rules, and never edited after the fact.

How KOL directional calls settle (kol_directional_v1)

kol_directional_v1 — directional calls settlement. Applies to unconditional directional statements (bullish/bearish/flat) by tracked authors on whitelisted assets. T0 = UTC close of the posting day; settled at UTC close of day +7. Altcoins are measured relative to BTC; BTC itself vs USD; band ±3% (identical parameters to settle_v1 — the same ruler we apply to our own AI).

Bullish HIT iff relative return > +3%; bearish HIT iff < −3%; flat HIT iff within the band. Each record also stores the absolute (vs USD) result — both are shown.

Excluded, fail-closed: conditional statements, hedged wording (could/may), relayed third-party views, plain descriptions. One record per author × asset × UTC day (earliest kept). Directional and price-target records are never merged into one hit rate. Retrospective records are labeled per the existing policy.

How the dollar figures are computed (follow_pnl_v1)

follow_pnl_v1 — the money view. Every directional call is treated as an independent $100 position: buy on bullish, short on bearish, closed after 7 days. The follow-along total is the plain sum of those positions. The benchmark answers one question — what if you had just held BTC over the same 7 days? It is recovered exactly from the settled record as (1+move)/(1+relative)−1, never approximated by subtraction. Both columns share one denominator: a call enters only when both legs are on the record.

No compounding, no fees, no slippage. Live and backtested records are never merged into one number, and wherever these dollar figures appear, the live/backtest split appears with them. CoinVerdict's own row is computed from its settled verdicts under the same rule; its prices are report-snapshot prices while KOL legs use UTC daily closes — a timing difference we state rather than hide.

When we say the ranking is carried by backtests

A prospective call is graded after it was made: we saw the statement, then waited for the market to settle it. A backtested call is one we applied the same rule to after the fact. Backtests are hypothetical performance — useful for coverage, but nobody traded on them in real time.

Whenever prospective calls make up less than 20% of all settled calls among ranked accounts, we say so at the top of the leaderboard, above the table. The threshold is fixed in advance and applied mechanically: once prospective coverage passes it, the notice comes down on its own. No one decides case by case whether to show it.

If we ever change this threshold, the change is dated and recorded here rather than applied silently.

How the KOL Rating is computed (kol_rating_v1)

kol_rating_v1 — the relative rating for directional records. Rating (0-100) = round( 2/3 × the percentile of cumulative excess return + 1/3 × the percentile of hit rate ), both percentiles taken among the rated accounts of that day. A Rating of 68 reads as: ahead of 68% of rated accounts. Cumulative excess return sums the direction-adjusted relative move of every settled directional call — bullish adds the relative move, bearish subtracts it, flat subtracts its absolute size. Percentiles are taken first and weighted second, so the score is a rank, not a return.

Benchmark hit rate = the share of all settled assets whose 7-day move relative to BTC landed in that direction. Bullish benchmark = the share that rose by more than +3%; bearish benchmark = the share that fell by more than −3%; flat benchmark = the share that stayed inside the ±3% band. The benchmark is measured with the same ruler as the calls it is compared against — same window, same band, same pricing base — and each direction carries its own benchmark; an account benchmark is the average of the three, weighted by how many calls the account made in each direction. Daily benchmarks are snapshotted and never edited afterwards: a change of definition means a new rule name, not a rewritten number.

Eligibility: an account is rated only after 20 settled directional calls and 90 days of tracking (first to last settlement). Below either bar the page shows a sample counter and no percentage at all. Rated accounts are split by the standard distribution — top 10% Top, next 22.5% Above average, middle 35% Average, next 22.5% Below average, bottom 10% Bottom. A Rating of 80 and above earns the CoinVerdict All-Stars badge; there is no negative badge of any kind.

The leaderboard is ordered by Rating, descending. Ties are broken by the lower bound of the Wilson interval on hit rate, and then by sample size — a long record outranks a short lucky one.

What the Rating does not say: a high Rating means an account did better than other tracked accounts on one shared yardstick, it does not mean that following it would have been profitable. The benchmark is the natural distribution of the market, not a risk-free return. Every number here is measured after the fact and carries no fees, no slippage and no funding costs.

Hit rate (all directional calls) — Rating is computed on every settled directional call — live and backtested together. The prospective-only rate stays on the line above and is never merged into it.

Rating ranks accounts against each other on one shared yardstick. It is not a forecast, and it is not a statement about what following anyone would have earned.

The five-tier display scale

Coin pages render a five-tier badge derived mechanically from the settled three-way stance × the judge's stated confidence. It is a display granularity only — every verdict settles on the same three-way rule (settle_v1), and the public record never changes.

Buyoverweight + high confidence
Overweightoverweight + medium/low confidence
Holdneutral (any confidence)
Underweightunderweight + medium/low confidence
Sellunderweight + high confidence

The macro weather rule

Three FRED series — WALCL (Fed balance sheet), DTWEXBGS (broad dollar index), DFII10 (10-year real rate) — each scored by its 4-week change: balance sheet up = +1, dollar down = +1, real rate down = +1 (else -1; missing = 0). Score >= +2 reads tailwind, <= -2 headwind, otherwise mixed. Two or more series missing → no label. The same facts feed the verdict's data pack — what you see is what the desk saw.

Check our work

The grading rule is open source — recompute every settlement yourself: github.com/OCArk/coinverdict-methodology

CoinVerdict reports are AI-generated research information, not financial advice. Verdicts are opinions; settlements are facts.