Stating Confidence Honestly in a Volume Analysis

Every other reading on this site defers to this page. It sets out what a volume analysis is allowed to claim, the four confidence bands the desk works in, the test a claim must pass before it can be published, and the specific sentences we will not write no matter how strong the pattern in front of us looks.

Question
How should a volume analysis state its confidence, and what is it allowed to claim?
Evidence used
Public Solana transaction records, the checks run against them, and the analyst record of how many checks were run and on what.
What it cannot show
Intent, identity, employment, agreement or motive behind any sequence of transactions.
Falsified by
Two analysts given the same raw data and the same band definitions assigning different bands more often than they agree.
Confidence
High. The rules here govern writing rather than reading, and every input to them is the analyst own record, which is fully observable.

The correct output of a volume analysis is a confidence band attached to a specific claim, never a verdict. Chain data shows what happened, not why it happened. A band states how well one explanation fits compared with its rivals, and it carries the observation that would overturn it. A verdict throws that work away.

Verdict or band

A verdict is a single word: fake, real, organic, manipulated. It compresses away everything a reader needs in order to disagree with you, and it travels badly. A screenshot of a conclusion outlives the evidence that produced it, and whoever receives it inherits certainty with no way of testing it. A band travels with its conditions attached.

A verdict also claims knowledge the desk does not hold. The record contains transfers, swaps, program invocations, signatures and slot numbers. It contains no employment, agreements, instructions or motive. Two wallets may be one person, two people funded from one place, or a service acting for a client. This does not prove verdicts are always wrong, only that the data cannot license them.

This page is the desk evidence standards analysis: the rule set every other reading here defers to. The test behind it is simple. Write your conclusion, then ask what an analyst who disagrees would have to show to change your mind. If the honest answer is nothing, you have written a preference. If it is an observation, you have a band and a falsifier.

The four bands

The desk works in four bands. Each is defined by two things and nothing else: the evidence it requires before you may enter it, and what it licenses you to do once you are in it. They are not percentages in disguise. A band describes the state of the evidence when you write, and it is expected to move as collection improves.

The four confidence bands, the evidence each one requires, the action each one licenses, and the wording it permits in the write-up.
BandEvidence it requiresWhat it licensesHow it must be worded
HighSeveral checks sharing no input agree, each reproducible by a second analyst, and the rivals were tested and fail.Publishing as the best available explanation, with inputs and window beside it.The pattern is consistent with produced flow, and the alternatives we tested do not account for it.
ModerateTwo or more checks agree, but a rival stays live because it was not tested or could not be with the data held.Publishing as a reading under review, with the untested rival named alongside.One explanation fits better than the others we could test, and here is the one we could not rule out.
LowOne check fires, or the checks disagree, or the window is too short to separate readings.Flagging for further collection. Not publication as a finding.We observed this. We cannot yet say what produced it.
UndecidableThe separating evidence is absent from the public record, or collection is incomplete beyond repair.Saying so plainly and stopping.This question cannot be answered with the evidence available to us.

Undecidable is the band people most often refuse to use, and the one that protects a desk. Whether two clusters share an operator, whether a maker was hired or acting alone, whether a burst of turnover was ordered by anyone at all: these live outside the record. Recording them as Undecidable is an accurate result, not a failure.

The falsification test

No claim leaves this desk unless the writer can state the observation that would overturn it. The falsifier must be something a reader could go and look at. Good falsifiers name a field, a window and a direction: if funding for these wallets resolves to separate origins across the window, the shared-funder reading fails and the band drops.

Falsifiers also discipline collection, because once you write down what would sink your claim you usually find you can check it. That is why the test belongs at the start of the work. The mechanics of pulling that data are covered in gathering evidence from chain data, and the Solana documentation is the reference for what the ledger exposes.

What would falsify this

The claim that a band beats a verdict fails if bands turn out to be decoration. Give two analysts the same raw extract and the same four definitions above. If they assign different bands more often than they agree, the framework is measuring the analyst rather than the evidence, and it would have to be replaced with something tighter.

Base rates

A signature that appears in most small pairs cannot single out one pair. This is the base rate problem, and it defeats more volume readings than any other error. A check can be informative and still leave you near where you started, because the population you drew from was mostly pairs that trip it for ordinary reasons: thin books, one active maker, routing bots.

Illustrative arithmetic, invented round numbers, describing no real pair

Take one thousand pairs, of which one hundred carry produced flow. A check fires on ninety of that hundred, which sounds strong. Suppose it also fires on one hundred and eighty of the nine hundred others, because thin-book trading can look similar.

Two hundred and seventy pairs now show a firing check, and ninety carry produced flow. A firing check leaves you at roughly one in three, not nine in ten. The check narrows a list; it is not evidence about any single pair on it.

The correction is to state the population you drew from before you state the result. If you scanned every pair created in a window, say so. If you started from a pair someone sent you because its chart looked odd, say that too: the selection is part of the evidence. The general form is the base rate fallacy. This does not prove the check is worthless, only that it is a filter.

What would falsify this

The base rate objection collapses for any check shown to fire rarely across a broad, unselected sample of comparable pairs. If a signature really is uncommon in ordinary activity, observing it narrows things sharply. That is measurable, so the fix is to measure it on a control set and publish the control set alongside the reading.

Twenty checks on one pair

Run enough checks against one pair and something will look unusual by chance. This is the multiple comparisons problem, and it is why the number of checks you ran is itself evidence for the write-up. The check that fired is the one you remember; the nineteen that returned nothing vanish from the story unless you force them to stay.

Illustrative arithmetic with invented round numbers: run twenty checks on one pair, each with a one in twenty chance of firing on perfectly ordinary activity. The chance that none fires is 0.95 raised to the twentieth power, about 0.36. So roughly two times in three, at least one check fires on a pair where nothing was produced. A single firing check there means very little.

ConfidenceLow

For the invented pair above, one firing check out of twenty sits in Low and goes no further. It licenses more collection and a longer window. It does not license a published finding, and it never licenses naming anyone, because ordinary activity examined twenty ways produces the same observation.

The remedies are unglamorous. Decide your checks before you look, count them, and report every outcome including the quiet ones. Require agreement between checks that share no input, since timing and size measures often move together and two correlated checks are closer to one. The standard treatment is the multiple comparisons problem, and no blockchain exempts an analyst from it.

Naming the venues counted

Activity on Solana runs across several decentralised exchanges at once. One asset can trade on a bonding curve, on several automated market maker pools, and through an aggregator that splits an order between them inside a single transaction. Any statement about turnover therefore carries a hidden coverage claim, and when it is wrong the figure is measuring a different thing than the reader assumes.

This applies to analysts and to software equally. If a desk reports a daily turnover figure, it has to name the pools and programs it counted and the window it counted them over. The same duty falls on the production side: a volume bot on Solana DEXs operates across several venues by design, and a turnover number it reports means little until you know which venues went into it and whether aggregator legs were counted once or twice.

So check a claim against the venues it covers before anything else. A published volume quality score that names no venues cannot be reproduced and therefore cannot be argued with, which makes it a brand asset rather than a measurement. Where one footprint has several honest readings, the grey zone hub collects the cases. This does not prove a figure is wrong; an unsourced figure cannot yet be right.

Observation, inference, recommendation

A written output has three layers and they must stay visibly separate. The observation is what the data says, and what a second analyst would see in the same extract. The inference is what you think produced it, which is an argument. The recommendation is what a reader should do, which is neither of the first two.

Mixed together

Trades arrive on a near-constant interval from wallets funded by one source, so this is a wash trading operation and the chart cannot be trusted.

The reader cannot tell which part is measurement and which part is opinion, so they must accept or reject the whole sentence. There is nothing here to check and nothing to disagree with in detail.

Kept apart

Observation: over the window collected, inter-trade gaps cluster tightly and the wallets involved trace to one funding origin.

Inference: more consistent with one operator running a program than with many independent traders. Rival not excluded: a single maker on a scheduled strategy. Confidence: Moderate. It does not show who the operator is or why they trade.

Readers ask for a shortcut, usually phrased as chart trust signals or a single grade. The honest version of that request is a short list of observations with a band beside each, not a badge. Where the pattern fits a maker as well as produced flow, say so and point at market making or wash trading. Terms above sit in the glossary.

Sentences we refuse to write

Some sentences are unavailable however strong the pattern looks, because they assert what the evidence structurally cannot support. The list below is a working checklist and each item carries its reason. Every entry describes a claim a reader could not falsify even in principle, or one that steps into a determination this desk has no standing to make.

  • Wallet A is manipulating this market. Manipulation rests on intent, which no ledger records; the term sits in market rulebooks rather than in data, as the general description of wash trading shows.
  • This project is buying its own volume. It names a party and asserts a relationship the record cannot establish.
  • They did this to attract retail buyers. That is a motive, and motives are not observable on chain.
  • This pattern means the price will fall. It turns a structural observation into a price expectation, and it functions as advice.
  • Volume quality score: 34 out of 100. A score without its inputs, weights, window and venue coverage cannot be reproduced or challenged.
  • The volume here is fake. A binary verdict with no band, no rival explanation and no falsifier.
  • No manipulation detected. It states absence of evidence as evidence of absence, inviting readers to read a limited check as a clearance.
  • Confirmed by our internal model. An authority claim nobody outside the desk can reproduce, which places it outside evidence.

One further refusal is worth stating openly: the desk does not publish guidance on making produced activity harder to detect. Detection methods appear here because readers can check them against public data. Turning them around serves nobody reading a chart, so where that question arises the answer is that we do not cover it.

The conclusion template

Every published reading on this site ends in the same shape. The sequence is deliberately rigid, because a fixed order makes an omission visible: a missing step is obvious rather than something a reader must infer from smooth prose. Produced volume is an openly sold category of activity, and the template exists to describe it accurately.

  1. Observation. State only what is in the extract, in terms a second analyst could reproduce: fields, window, pair, and the population it came from.
  2. Method. Name every check you ran, including those that returned nothing, and say whether they share inputs.
  3. Inference. State the explanation that fits best, name the rivals, and say which you tested.
  4. Falsifier. State the one observation that would overturn the inference, in a form a reader could check.
  5. Confidence. Assign one of the four bands and give the reason it is not the band above it.
  6. What it does not show. Close with the limits: no identity, no intent, no agreement, no price expectation.

The template filled in, using an invented pair called PAIR-X

Observation: across the collected window on PAIR-X, inter-trade gaps cluster in a narrow range and sizes repeat from a small set of values. The pair came from a full scan of a launch window, not from a tip. Method: six checks run, four fired, two returned nothing; two of the four share a timing input and count as one.

Inference: one operator running a program fits better than many independent traders. Rival not excluded: a maker on a schedule. Falsifier: funding resolving to separate origins sinks the shared-operator reading. Confidence: Moderate. It does not show who runs PAIR-X, or why.

When the honest output is nothing

Sometimes the template cannot be completed, and then the desk publishes nothing. That happens when collection is incomplete beyond repair, when the window is too short to separate readings, when the question is undecidable from public data, or when the only sentence that fits the evidence would identify a party. A hedged accusation is still an accusation.

Publishing nothing is not wasted work. The collection, the checks, their outcomes and the band the evidence supported are all recorded, and a pair parked at Low today can move once a longer window exists. A desk that never files an empty result is not careful, it is selective, and selection is the one bias no band can correct afterwards.

Questions this desk is asked

Why not just say whether the volume is fake?

Because the record does not contain that answer. A chain shows transfers, swaps, signatures and slots. It does not show who controls a wallet, what they agreed with anyone, or why they traded. A one-word verdict asserts all three. A confidence band states how well one explanation fits compared with its rivals, and it names the observation that would move it, so a reader can disagree with you using evidence rather than opinion.

What is the difference between Low and Undecidable?

Low means the evidence exists but is thin or contradictory: one check fired, or the checks disagree, or the sample is too small to separate readings. More collection can move it. Undecidable means the separating evidence does not exist in the public record at all, so no amount of further collection from chain data will settle the question. Low is a reason to keep working. Undecidable is a reason to stop and say so.

What is a falsifier and why is one required?

A falsifier is the specific observation that would show your claim is wrong. If a reading cannot be overturned by anything, it is not a finding, it is a preference dressed as analysis. Requiring one forces the writer to think about the rival explanations before publishing rather than after being challenged. It also gives the reader a task they can actually perform, which is the only real check on a desk like this one.

Does a base rate really change the reading that much?

Yes, and it is the most common source of overconfidence in this field. If a signature appears in a large share of ordinary small pairs, then observing it in one pair barely narrows anything down. The check can be genuinely informative in isolation and still leave you close to where you started, because most of the pairs that trip it were never produced in the first place. The prior matters as much as the test.

Why does the number of checks I ran matter?

Because unusual-looking results accumulate as you look for them. Run enough checks against one pair and something will fire by chance alone, and the check that fires is the one you will remember. The fix is unglamorous: decide the checks in advance, count them, report every outcome including the quiet ones, and require agreement across checks that do not share an input rather than treating any single firing as a result.

Can I publish a volume quality score on its own?

Not on this desk. A single number hides the inputs, the weights, the sample window and the venues counted, which are exactly the things a reader would need in order to challenge it. A score is a summary of an analysis, not a substitute for one. If you publish the score, publish the inputs beside it, and state which venues were included, otherwise the number cannot be checked or reproduced.

When should the desk publish nothing at all?

When the collection was incomplete, when the sample is too small to separate readings, when the question is undecidable from public data, or when the only honest way to write the finding would identify a party. Publishing nothing is a result, not a failure. The work is still recorded internally: what was collected, which checks ran, what they returned, and the band the evidence supported at the time.

Filed in Grey zone by The Volume Forensics Desk. Patterns described here come from protocol design and public transaction data; every figure in an example is invented, labelled and describes no real pair. The evidence standard the desk works to is set out in stating confidence honestly, and the terms used are defined in the glossary.