Organic vs artificial volume: reading the distribution
Turnover is the least informative number a market publishes. It says how much changed hands, never how many minds were behind it. This page sets out what organic and artificial actually mean, why the difference lives in a distribution rather than in any single trade, the six checks this desk runs, the cost floor that constrains anyone producing volume, and what an honest conclusion has to say out loud.
- Question
- Can organic and artificial trading volume be told apart from on-chain data alone?
- Evidence used
- Populations of swaps read from a stated window: counterparty sets, inter-arrival gaps, size distributions, net inventory change, depth response and funding edges.
- What it cannot show
- Intent. Chain data records what an address did, never why an operator did it, and never who instructed the operator.
- Falsified by
- A market whose flow passes several checks yet is later shown, from off-chain records, to be entirely one operator; or a cluster that fails every check and is shown to be many unrelated traders.
- Confidence
- Moderate. The checks are reproducible and mutually independent, but each has a documented innocent explanation, so the combined reading is a lean rather than a finding.
The question this page answers
Organic volume is turnover produced by many independent decisions. Artificial volume is turnover produced deliberately by one operator so that a market looks busier. Neither label belongs to a single trade. The difference only becomes visible across many trades, and even then the honest output is a probability statement rather than a verdict.
The vocabulary gets in the way, so it is worth fixing early. Produced activity is not hidden, not rare, and not automatically against any rule. It is a normal category of market behaviour with an open supply chain and published prices. The organic vs artificial volume question is a measurement problem. Judging it needs documents a transaction record does not contain.
What organic and artificial mean
An operational definition beats a moral one. Call flow organic when the decisions behind it are independent, meaning one participant trading does not mechanically cause another to trade. Call it produced when one decision-maker is behind trades presented as if they came from separate participants. The word artificial is common, but produced describes the mechanism better.
Notice what those definitions leave out. They say nothing about risk, profit or honesty. A market maker quoting both sides is one decision-maker behind many trades, and is a useful participant. A crowd reacting to the same post is many decision-makers whose choices are anything but independent. Both definitions have awkward edges, and the checks below inherit them.
Spontaneous flow
Arrives when something happens. Bursts around news, listings and price moves, then thins out for hours. Sizes are untidy because people type whatever amount they have. Counterparties keep expanding, since each new participant is a new address the market has not seen before.
The signature is irregularity. Gaps between trades vary by orders of magnitude across a day, and the address set has a long tail of one-time visitors that never gets exhausted.
Produced flow
Arrives because a configuration says so. The operator sets a budget, a size band and a cadence, and the process runs until the budget is spent. Because the configuration is fixed, the output inherits its regularity in gaps, in sizes, or in both at once.
The signature is a narrow band. Not identical values, which would be trivial to spot, but a distribution that is tighter than a crowd of independent people would ever generate on its own.
Why this is a distribution question
Take any single Solana swap and write down everything the chain knows: signer, route, amounts in and out, fee paid, slot. Now try to say whether it was organic. The exercise fails at once, because every field is equally consistent with a person acting on a hunch and with a process executing step 417 of a schedule. There is no field for motive.
Move up a level and the problem changes character. One thousand swaps have properties no individual swap has: a distribution of gaps, a distribution of sizes, a graph of who traded with whom, and a net inventory position for each participant when the window closes. Those properties are measurable and unavailable at the level of a single row, which is why organic vs artificial volume is a distribution question.
This is why the useful question is never whether volume is fake. It is how much of the observed turnover is consistent with independent decisions, expressed as a range with the assumptions written next to it. A reading that cannot be stated as a range is usually a reading that was never measured. This does not prove that any particular market is one thing or the other.
What would falsify this
The distribution argument fails if markets with genuinely independent participants routinely produce gap and size distributions as tight as a scripted process. If a broad sample with no plausible single operator showed the same narrow bands, the checks below would be measuring venue mechanics rather than behaviour.
The six checks and their innocent twins
Six checks carry most of the weight. They are chosen because they are close to independent: a market can look unusual on one without looking unusual on the others, which means agreement between them is informative rather than circular. Each is stated below with the innocent explanation that defeats it, because a check without its counter-argument is not evidence, it is decoration.
| Check | Spontaneous flow tends to look like | Produced flow tends to look like | The innocent explanation |
|---|---|---|---|
| Counterparty diversity | A long tail of addresses, most seen once | A short address list recycled, tail missing | A new pair genuinely has few participants, and aggregator routing hides many users behind one program |
| Inter-arrival regularity | Gaps that bunch and stretch, dead hours and bursts | Gaps inside a narrow band whatever the hour | Scheduled execution, time-weighted orders and rebalancers are periodic by design |
| Size dispersion | Sizes spread widely, untidy fractions, outliers | Sizes drawn from a small repeating set of round values | Front ends offer preset amounts, and people type round numbers |
| Net inventory | Participants finish meaningfully long or short | A cluster finishes near where it started after heavy turnover | Market makers end flat on purpose, and arbitrage desks are flat by definition |
| Depth response | Sizes adapt as liquidity thickens and thins | Sizes hold their shape while depth changes underneath | A risk limit can cap size permanently, so indifference may just be policy |
| Funding topology | Wallets funded from many unrelated sources | Wallets tracing back to a few funders in a short span | An exchange withdrawal address funds thousands of strangers, and a treasury funds its own team |
Counterparty diversity and funding topology are really one question asked twice: how many separate parties are here. Both depend on joining addresses that behave as one operator, a method with its own error rate and its own page in how wallet clustering is done. A cluster drawn too generously manufactures the concentration it claims to have found.
Inter-arrival regularity and size dispersion are the checks people reach for first, and the two most often misread. How spacing and quantisation are measured, and where each collapses, belongs to the work on timing and size signatures. What matters here is only that both describe tightness, not identity.
Net inventory is the check with the most economic content. If a set of addresses generated substantial turnover and finished holding roughly what it started with, the turnover moved no risk anywhere, which is the condition the classical wash trade definition turns on. That is the mechanical core of round trips and self-trading, and also what a competent market maker produces on a quiet day. The pattern is shared; only the context differs.
Depth response is the least discussed and often the most useful. Independent traders react to the book in front of them; a fixed configuration frequently does not, because adapting to depth is work nobody specified. This proves nothing on its own, since a disciplined trader with a hard size limit looks identical.
Moderate rather than High because every row of that table has a defeating explanation common in real markets, so no single check survives cross-examination. Moderate rather than Low because the six are close to independent and their innocent explanations are not mutually compatible, so several tripping at once narrows the field.
What producing turnover costs
Volume is not free to manufacture, and the cost floor is the most underrated constraint in the organic vs artificial volume debate. Every unit of displayed turnover pays network fees, venue fees and price impact. Those costs scale with the turnover produced, so an operator faces a real budget, and budgets create the repetition the checks above notice.
Two protocol facts anchor the arithmetic. One SOL is 1,000,000,000 lamports, and the base transaction fee is 5,000 lamports per signature, as documented in the Solana developer documentation. Priority fees sit on top and vary with demand. Everything else in the example below is invented, and is there only to show the shape of the bill.
Illustrative cost floor, invented round numbers describing no real pair
Suppose an operator wants 10,000 SOL of displayed turnover on an invented pair called PAIR-X. Suppose the swap size is 5 SOL, so 2,000 swaps are needed, each in a single-signature transaction. Network cost is 2,000 multiplied by 5,000 lamports, which is 10,000,000 lamports, or 0.01 SOL in total.
Now add the two invented figures. At an invented venue fee of 0.25 percent, the fee bill is 25 SOL. At an invented average slippage of 0.10 percent, price impact costs a further 10 SOL. The floor is therefore about 35.01 SOL for 10,000 SOL of turnover, roughly 0.35 percent of the number displayed.
The instructive part is the ratio. Network fees are under one part in three thousand of that bill; venue fees and impact are essentially all of it. Turnover therefore cannot be produced at negligible cost, and the bill grows in proportion to the headline figure. These numbers describe no real pair and no real fee schedule.
Read that arithmetic backwards and it becomes a detection tool. A cost scaling linearly with turnover pushes an operator towards the cheapest configuration per unit of displayed volume, and towards rerunning it. Repetition is not evidence of dishonesty. It is evidence of optimisation, and optimisation leaves a readable footprint.
Where produced activity comes from
Produced flow is not improvised by individuals. It comes from a tooling layer that is openly sold, priced per campaign or per wallet, and documented like any other software product. A Solana volume bot is a normal part of Solana market structure in the same way that execution algorithms are a normal part of equity market structure, and treating that supply side as scandalous simply makes the reading worse.
Knowing the supply side helps because tools have defaults. A screen offering a size band, a wallet count and an interval produces output shaped by those fields, and the fields operators rarely change become the features that repeat across unrelated campaigns. No inspection of the tooling is needed; the constraint shows up in the output regardless.
It also sets the neutral tone this desk works in. Some produced flow exists to meet a listing threshold, some to keep a chart from looking abandoned, some to support a market maker on a thin book. Those motives differ enormously, and none of them is visible on chain.
What would falsify this
If commercial tooling produced output statistically indistinguishable from spontaneous flow across all six checks, the supply-side argument would collapse. The falsifying observation is a large sample of known-produced flow whose gap, size and inventory distributions match crowd flow within ordinary sampling error.
What makes a measurement checkable
A volume figure means nothing until four things travel with it. The window, as an exact start and end with the time source named. The data source, so a reviewer can pull the same rows. The deduplication rule, since a multi-hop route can be counted once or several times. And the venue coverage, because a figure omitting a venue is unanswerable rather than wrong.
That standard applies to commercial reporting as much as to research. When a platform publishes what a campaign delivered, the useful question is whether it states its window, source, deduplication rule and venue list, which is precisely what a page describing how volume campaigns are measured exists to answer. A number without those four declarations cannot be checked by anyone, in either direction.
The same discipline is why raw fields get stored rather than conclusions. Signature, slot, fee payer, program, pre and post token balances: those can be recomputed and disagreed with. A stored label saying artificial cannot. Explorers such as Solscan are fine for spot checks, but a reproducible reading needs its own captured rows.
What an honest conclusion contains
The output of an organic vs artificial volume reading is a paragraph, not a verdict, and that paragraph has required parts. Anything missing from the list below turns a measurement into an assertion, and assertions about markets get repeated far past the evidence that started them.
- The exact window, with start, end and the time source used.
- The venues and programs covered, and the ones knowingly left out.
- The deduplication rule applied to multi-hop routes.
- Which of the six checks were run, including the ones that returned nothing.
- The innocent explanation considered for each check that was tripped.
- The single observation that would overturn the reading.
- A named confidence band, with the reason for that band rather than the one above it.
- A plain sentence saying what the reading does not claim.
- Anonymised labels wherever intent has not been established.
The confidence band is the part most often skipped, and it does the most work. A band forces the writer to decide in advance what evidence would have been needed for the next band up, which is a much harder question than describing what was found. The standard this desk applies is set out in the note on stating confidence honestly.
Note what the list refuses. It never asks for a conclusion about intent, because no combination of the six checks reaches intent. Whether a footprint is market making or something else stays open on chain data alone, and pretending otherwise is the commonest failure in published volume analysis.
What none of this proves
None of the six checks proves that produced volume exists in any market, and no combination of them proves who produced it. They narrow a space of explanations. A market tripping five checks has fewer plausible innocent stories than one tripping none, which is useful, but fewer stories is not the same as one story.
Three limits are permanent. Clustering can be wrong, so concentration may be an artefact of the join. Coverage can be incomplete, so a quiet venue can invert a reading. And intent is simply absent from the record, which is why a description of behaviour is never promoted into an accusation on these pages, however many checks a cluster trips.
The desk also does not publish evasion guidance. Where a section naturally raises how a footprint could be reshaped to defeat a check, it stops there. Explaining detection helps a reader interpret a chart; explaining concealment helps only someone with a different aim.
What would falsify this
The framework fails if an operator is shown to have produced most of a market whose flow passed all six checks cleanly, or if a cluster failing all six is shown by funding records to be many unrelated traders. Either case would mean the checks measure something other than the number of independent decision-makers.
Questions this desk is asked
Is artificial volume the same thing as wash trading?
No. Wash trading is a narrower idea: trades where the buyer and the seller are the same economic party, so no risk changes hands. Artificial or produced volume is broader and covers any turnover generated deliberately to change how a market looks, including flow that does move real risk between separate wallets. A campaign can produce volume without any single trade meeting the wash definition.
Can a single transaction be identified as artificial?
Not from chain data alone. A transaction records a size, a route, a fee payer and a slot. None of those fields carries intent. The only way a single transaction becomes evidence is by belonging to a population that behaves unusually as a group, and the claim then attaches to the population, not to the transaction. Reading one swap in isolation is guesswork.
Why is turnover treated as a weak number here?
Turnover adds notional value and stops. It cannot distinguish one address trading with itself two hundred times from two hundred addresses trading once each, because both sum to the same figure. Any statistic that collapses a distribution to a single total discards exactly the information needed to read the distribution. That is why this desk starts from counts, gaps and counterparties instead.
Do the six checks have to agree before a reading is written?
They do not have to agree, and it is more informative when they disagree. Each check has its own innocent explanation, so a cluster that trips one check has almost certainly done nothing unusual. A cluster that trips five, where the innocent explanations for those five are mutually inconsistent, is a different situation. The number of checks tripped sets the confidence band, not the verdict.
Is produced volume against the rules on Solana?
Solana is a settlement layer and does not have a rulebook about trading motives. Venues, launchpads and jurisdictions each apply their own standards, and those standards differ widely. This desk describes what patterns look like in data and stops there. Whether a given campaign breaches a venue term or a national regulation is a legal question that needs documents chain data cannot supply.
Why does network fee arithmetic matter to detection?
Because it sets the shape of what an operator can afford. Producing turnover costs network fees, venue fees and price impact, and those costs scale with the turnover produced. Anyone working against a cost constraint tends to repeat whatever configuration is cheapest, and repetition is what the timing and size checks are built to notice. The economics create the regularity that the measurement then reads.
What does this desk refuse to publish?
Two things. It does not attach wrongdoing to a named token, wallet, team or venue, because intent cannot be established from a transaction record and a wrong accusation is not repairable. It also does not publish evasion guidance, meaning it will not describe how produced activity could be shaped to defeat any check on this page. Detection is explained; concealment is not.
Filed in Signatures by The Volume Forensics Desk. Patterns described here come from protocol design and public transaction data; every figure in an example is invented, labelled and describes no real pair. The evidence standard the desk works to is set out in stating confidence honestly, and the terms used are defined in the glossary.