Report · Recursive self-improvement · Empirical
Trench-RSI: recursive self-improvement in a market that verifies instantly and rewards almost nothing
Primary source: Kamat (2026), arXiv:2607.02823 · Article corpus of 495 entries, collected 17 September 2026
Proposals cost nothing. The verdict lands in 1.66 minutes. And the loop converged on being easier to check.
-
3,000
proposals
cost ~0 to make
-
1.66 min
the verifier
median launch → verdict
-
2,994
timed out
99.802% of the batch
-
6
graduated
0.198%, a lower bound
One batch — 3,000 launches. About six clear the gate — the count varies batch to batch, the rate does not.
RSI here means recursive self-improvement — a loop that proposes, tests, and feeds the result back into the next proposal. It is not the trading indicator of the same initials.
Nothing here argues that memecoins are good, useful, or worth buying. The claim is narrower and, I think, more uncomfortable: this is what a fast loop with a cheap verifier actually optimises.
832,941
launches with a usable terminal outcome, in a 34-day window
1,651
of them graduated — 0.198%, and that is a lower bound
1.66 min
median time from launch to verdict; p90 is 2.80 minutes
17.4×
graduation gap between advertising nothing and advertising three channels
01
The loop
Recursive self-improvement needs four parts: something that proposes candidates, something that judges them, a reward that separates the good from the bad, and a channel that carries the judgement back into the next round of proposals. The interesting question is never whether a loop exists. It is what the loop's verifier is cheap enough to measure — because that, and not the designer's intention, is what gets optimised.
The trenches — the retail memecoin market on Solana and its imitators — instantiate all four parts with unusual clarity, and at a scale no laboratory would fund.
- Proposal. Deploying a token on pump.fun costs approximately nothing and takes seconds. The platform has hosted more than 11.9 million tokens since its January 2024 launch.
- Verifier. The market. It does not read the whitepaper; it prices the thing. In the measured window the median verdict landed 1.66 minutes after launch.
- Reward. Graduation — completing the bonding curve, roughly 85 SOL in real reserves, historically near a $69,000 market capitalisation, after which liquidity moves to a decentralised exchange. It happened 1,651 times out of 832,941.
- Update. Here is the twist. The coins do not improve — each is a fresh draw. What improves is the infrastructure: the launchpads, the trading terminals, the copy-trading rails, the fee mechanics. The corpus documents 55 platforms across six generations.
So the loop is real, and it is closed. But the thing being recursively improved is not the product. It is the factory, and — as section 03 argues — the packaging.
This report is a companion to The Verification Frontier, which argued that progress in most domains is rate-limited by how expensive it is to check an answer. The trenches are that argument's natural experiment: a domain where checking is nearly free. If cheap verification were sufficient, this is where we would see it work. What we see instead is worth looking at closely.
02
The cheapest verifier ever built
Verification cost is usually the binding constraint. A clinical trial takes years. A proof needs a referee. A chip needs a fab. The trenches removed the constraint almost entirely, and the resulting iteration speed is genuinely without precedent.
Kamat's survival analysis of 832,941 launches gives us the timing distribution directly. Of those launches, 1,651 graduated and 831,290 timed out. Among the graduations:
Every graduation in the dataset happened inside six minutes. Compare that to the verification latency of any other field of human effort and the gap is not a difference of degree.
The caveat that has to travel with every number on this page. The collector's coverage extended only about six minutes past each launch. So the observed rate is a fast-regime lower bound on the true 24-hour graduation rate — the paper is explicit that no upper bound is established. A slow graduation would simply be invisible to the instrument. Every rate quoted here, including 0.198%, should be read as "at least this, within six minutes."
And the loop ran. It ran at a rate that makes the point on its own: pump.fun became the fastest crypto application ever to reach $100 million in cumulative revenue, hitting the mark in 217 days, and crossed $1 billion in lifetime revenue in March 2026 — the first application in Solana's history to do so. The verifier was cheap, so it was used, relentlessly.
This is the part of the story that should be taken seriously by anyone who thinks cheap verification is the bottleneck. It is a bottleneck, and removing it does buy blistering iteration. It just does not buy what you might hope.
03
What the loop actually optimised
If you run 11.9 million proposals through a verifier, you learn what that verifier rewards. The single strongest published signal is not about the token. It is about how much of a token's surrounding apparatus was visible to whoever was looking.
Kamat splits the 832,941 launches on whether each advertised a Twitter account, a website, and a Telegram group at launch. The gradient is monotone in the number of channels present:
| Channels advertised | Graduation rate | Relative to none |
|---|---|---|
| None of the three | 0.110% | baseline |
| Pooled — all launches | 0.198% | 1.8× lower bound |
| All three | 1.919% | 17.4× |
Channel-count gradient. Source: Kamat (2026), arXiv:2607.02823.
Broken out by individual channel, absence versus presence, the same shape appears — and Telegram dominates:
| Channel | Absent | Present | Ratio | p |
|---|---|---|---|---|
| 0.1486% | 0.2268% | 1.53× | 1.1e−14 | |
| Website | 0.1584% | 0.2636% | 1.66× | 1.2e−25 |
| Telegram | 0.1661% | 1.4850% | 8.94× | <1e−300 |
Absent/present split. Group sizes: Twitter 304,815 / 528,126; Website 517,680 / 315,261; Telegram 812,671 / 20,270. Source: Kamat (2026).
The Cox proportional-hazards model tells the same story with covariates held together. Hazard ratios, with 95% confidence intervals:
| Covariate | Hazard ratio | 95% CI |
|---|---|---|
| has_telegram | 5.402 | [4.733, 6.166] |
| log(1 + initial mcap in SOL) | 4.506 | [4.293, 4.729] |
| has_twitter | 1.305 | [1.192, 1.428] |
| has_website | 1.194 | [1.096, 1.300] |
| log(1 + description length) | 1.054 | — |
Harrell's C = 0.858 [0.850, 0.870]; split-half 0.881 / 0.847. Source: Kamat (2026).
A concordance of 0.858 is a model that discriminates well. And what it discriminates on is presence of apparatus and initial capital. Not the meme. Not the art. Not anything a person would call quality.
This is Goodhart's law, measured in the wild, at n = 832,941. A cheap verifier is necessarily a proxy verifier — it has to be, or it would not be cheap. And a fast loop finds the proxy's holes faster than it finds real value, because the holes are exactly what the proxy is sensitive to. Eleven point nine million proposals is enough iteration to locate every one of them.
What this is not. The paper explicitly declines to make a causal claim, and so do I. A Telegram group may proxy for creator effort, for bot discoverability, or for self-selection by creators who were going to push harder anyway. Adding a Telegram link to a token does not multiply its odds by 8.94. The 17.4× is a correlational gradient and must not be upgraded into a recipe. What it does establish is what the verifier's decisions track — and that is the claim being made here.
04
The world evolves faster than the players
In a normal improvement loop the environment holds still while the candidates get better. Here it is inverted. The individual proposals never improve — every launch is an independent draw — while the environment around them is rebuilt every few months.
The corpus records 55 platforms, of which 44 carry a datable launch or founding year. The distribution is a generational cascade, not a gradual build:
| Year | 2021 | 2022 | 2023 | 2024 | 2025 | 2026 |
|---|---|---|---|---|---|---|
| Platforms | 8 | 3 | 7 | 10 | 8 | 4 |
44 of 55 platforms have a datable year; 4 more fall before 2021. Derived from the article corpus, collected 17 September 2026.
The generations are legible in the names. 2024 brought daos.fun, pump.fun and Virtuals. 2025 brought Bags, Believe, Boop and Thrust. And pump.fun itself kept re-tuning the reward — Project Ascend in September 2025, then BOOST in 2026 — which is the update step operating on the verifier rather than on the candidates.
You can watch that re-tuning move the measured rate around, which is the clearest possible evidence that the rate is a property of the mechanism and not of the memes:
| Period | Graduation rate | Note |
|---|---|---|
| Lifetime, early 2025 | ~1.4% | reported baseline |
| During 2025 | below 1% | Cointelegraph |
| June 2026 | ~0.26% | DEXTools, at the trough |
| Best single day, post-BOOST | ~6.7% | six-month high |
| Cumulative, all tokens hosted | under 2% | of 11.9 M+, per DefiLlama |
Secondary figures as reported in the article corpus, which cites the named primaries. These are measured on different denominators and windows than the arXiv study and are not directly comparable to its 0.198%.
Meanwhile the environment can also collapse. The corpus documents a 2026 memecoin winter triggered on 1 February 2026 — "Black Sunday II" — after which Solana weekly DEX volume fell about 62% in three weeks, SOL traded roughly 67% below its $294 all-time high, and about $2.2 billion in leveraged positions were liquidated in 24 hours. A loop whose verifier is a market inherits the market's regimes.
05
The reward landscape
A loop's behaviour is set by the shape of its reward, not its average. This reward is close to the most extreme shape available: almost everything scores zero, and nearly all of the value that exists sits in a handful of outcomes.
Take the 168 coins in the corpus. 149 of them record a peak market capitalisation. Their distribution:
| Group | Share of total peak market cap |
|---|---|
| Top 1 coin | 27.4% |
| Top 5 | 53.4% |
| Top 10 | 65.2% |
| Top 25% (37 coins) | 87.9% |
n = 149 coins with a recorded peak market cap, of 168 in the corpus. One coin — COPE — states its peak was not reliably recorded and is excluded rather than guessed. Derived from the article corpus, 17 September 2026.
And these are the survivors of the survivors — coins notable enough to have an encyclopaedia article. The base rate underneath them is the 0.198% lower bound. Layer the two together and the reward landscape is: graduate at roughly one in five hundred, and conditional on mattering at all, expect the median outcome to be three orders of magnitude below the top.
The Wilson 95% interval on the pooled rate is tight, so this is not a small-sample artefact:
\[ \hat{p} = \frac{1{,}651}{832{,}941} = 0.001982, \qquad \text{Wilson 95\% CI} = [0.00189,\ 0.00208] \]
Restricting to the steady-state portion of the window (763,091 launches) gives 0.207%, and the top quartile by initial market cap graduated at 0.634% — better, and still a rounding error.
Downstream of all this, the corpus reports that only about 3% of pump.fun accounts have ever realised more than $1,000 in profit. A loop can iterate 11.9 million times and still leave 97% of its participants with nothing. Iteration speed is not the same thing as progress, and it is definitely not the same thing as distributed benefit.
06
The discovery tree
Eight generations, 3,000 launches each. One dot is one launch. Survivors are the blue nodes, and each one seeds the next generation's spray. The only thing that changes between the three modes is the gate rate — the number of launches never does, and neither does the bar scale.
Figure 1 — the discovery tree. Violet is a launch that did not graduate; blue cleared the
market's check. Gate rates are the published channel-count gradient — 0.110%, 0.198%, 1.919%.
Everything else is illustrative: the 3,000-per-generation batch size and the parent attribution are
drawing choices, and the study does not track lineage between launches. The bar ceiling is fixed at
90 survivors in every mode on purpose — normalising each mode to its own maximum would make all
three look identical and delete the 17.4×, which is the whole point of the figure. Switching mode
resets the counters, because a running rate blended across three different gates would not mean
anything. The animation pauses when scrolled out of view and respects
prefers-reduced-motion.
07
The loop, running live
Figure 1 is a simulation. This one is not. Press run and a language model becomes the proposal stage of the loop, a verifier calibrated to the published rates becomes the market, and you watch the model reason its way toward whatever the gate rewards. Its thinking is streamed as it arrives, unedited.
The proposer controls five inputs and is deliberately not told what any of them mean. Three
are binary — s1, s2, s3 — and two are dials from 0 to 10.
It never sees the gate's rule. Each generation it submits four configurations, each configuration
is run through 3,000 launches, and the only thing that comes back is how many were accepted. No
ranking, no gradient, no explanation. Blinding the labels matters: the paper this page is built on
is public, and a model told that s3 meant a Telegram link could recite the answer
instead of discovering it.
Proposer · qwen/qwen3.8-27b
Checking the backend
| Gen | s1 | s2 | s3 | Scale | Detail | Accepted | Rate |
|---|
Figure 2 — the loop, live. The proposer is a real model call per generation and the
reasoning text is its own, streamed as it is produced. The verifier is arithmetic, not a market:
per-launch probability starts at the published 0.110% for none of the three channels and is
multiplied by each present channel's observed absent-to-present rate ratio. Those three
marginals multiply to 22.7×, but the paper's observed channel-count gradient is 17.4×, because
the channels co-occur — a launch with a Telegram tends to have a website too. So the product is
raised to an exponent of 0.9157, which lands none-of-three on 0.110% and all-three on 1.919%
exactly. That calibration is a deliberate choice, disclosed here because the alternative is a
figure whose headline number contradicts the paper it cites. The two dials are illustrative
only: they use the Cox hazard ratios for market cap (4.506) and description length (1.054),
compressed so neither can swamp the channel effect. The verifier is served in full at
api/gate so none of this has to be taken on trust. Outcomes are random draws, so
two runs differ, and the model can read noise as signal — watch for it inferring a precise
optimum for a dial from a single generation. That is not a bug in the demo. It is the Goodhart
failure the loop is being used to illustrate, happening in front of you.
08
Measurements
Two independent bodies of evidence. The survival study supplies the loop's rates and timings; the corpus supplies the structure of what the loop produced. I re-collected and recomputed the corpus myself rather than inheriting its figures, so that every number below is one I can reproduce.
The survival study
| Quantity | Value |
|---|---|
| Launches with usable terminal outcome | 832,941 |
| Graduated | 1,651 |
| Timed out | 831,290 |
| Pooled fast-regime rate | 0.198% |
| Wilson 95% CI | [0.189%, 0.208%] |
| Steady-state rate (763,091 launches) | 0.207% |
| Top quartile by initial market cap | 0.634% |
| Observation window | 2026-05-08 → 2026-06-10 (34 days) |
| Median / p90 / max time to graduation | 1.66 / 2.80 / 5.98 min |
| Harrell's C (Cox model) | 0.858 [0.850, 0.870] |
| Published dataset size (Zenodo) | 860,213 launches |
Kamat (2026), arXiv:2607.02823 v2, incorporating Corrigenda v1.3 and v1.4. Verified against the primary source on 17 September 2026. All rates are fast-regime lower bounds; collector coverage ended about six minutes after each launch.
The corpus
| Measure | Value |
|---|---|
| Articles | 495 |
| Body characters | 2,816,030 |
| Internal links | 3,881 |
| External links | 6,324 |
| Unique external domains | 791 |
| Link graph edges | 3,762 |
| Mean out-degree | 7.6 |
| Orphans (no inbound link) | 100 |
| Images referenced and retrieved | 494 / 494 |
Recomputed from a fresh collection of all 495 articles on 17 September 2026. External links are counted before infobox removal; counting after gives 6,049, which is why the two figures differ depending on method.
| Category | Coins | People | Lexicon | Platforms | Events | Culture | Guides |
|---|---|---|---|---|---|---|---|
| Articles | 168 | 157 | 61 | 55 | 31 | 13 | 10 |
495 articles total. Derived from the corpus.
The link graph is dominated by a few hubs, which is itself a measurement of what the community treats as load-bearing. Inbound link counts:
| Article | Inbound links |
|---|---|
| The trenches | 340 |
| Solana | 275 |
| pump.fun | 273 |
| Meme coin | 112 |
| Crypto Twitter | 84 |
| Kolscan | 65 |
| Ansem | 53 |
| Bonding curve | 49 |
| Copy trading | 48 |
| KOL | 44 |
| Axiom | 41 |
| Robinhood Chain | 38 |
Top 12 by in-degree, of 495 articles. Derived from the corpus.
The proposal rate over time, and where the proposals happen:
| Coin launches | 2023 | 2024 | 2025 | 2026 |
|---|---|---|---|---|
| Corpus coins | 17 | 62 | 31 | 46 |
Nine more coins predate 2023. By chain: Solana 117, Robinhood Chain 16, Ethereum 10, Base 9, BNB Chain 8. Chain names are normalised by first canonical substring match, so variants such as "BNB Smart Chain" resolve separately where the corpus writes them differently. Derived from the corpus.
And the events the community considered worth recording — 31 in total, 26 of them datable — cluster exactly where the platform generations do: 8 in 2024, 12 in 2025, 6 in 2026.
09
Analysis, and what would falsify this
The argument in one paragraph, then the honest account of how it could be wrong.
Near-free verification really does buy iteration speed — 11.9 million proposals and a median verdict in 1.66 minutes is not a small achievement, and nothing else humans do runs a closed loop that fast. But cheap verification is achieved by making the verifier a proxy, and a loop this fast will find the proxy's holes long before it finds anything of value. What eleven point nine million iterations discovered was not better memes. It was legibility: presence of a Telegram group, a website, a Twitter account, and enough starting capital to look real. The measurable improvement accrued to the infrastructure — 55 platforms in six generations — while the candidates themselves never improved at all.
The lesson generalises past crypto, which is the only reason this is worth writing up. If you make your evaluator cheap enough to run millions of times, you have chosen a proxy, and you have built the machine that will exploit it. Speed does not fix that. Speed is what makes it happen sooner.
What would falsify it
- A slow regime that changes the picture. The strongest threat. If graduations beyond the six-minute observation horizon are common and are distributed differently — say they favour tokens with no social apparatus — then the channel gradient is an artefact of what the collector could see. Extending coverage to 24 hours settles this, and until someone does, the lower bound caveat is not decoration.
- A causal test that comes back null. The 17.4× is correlational. If you randomised the presence of a Telegram link across otherwise-matched launches and graduation rates did not move, then legibility is a marker of creator effort rather than a thing the verifier rewards, and section 03's framing is wrong in an important way.
- Evidence that the candidates do improve. I claim each launch is an independent draw. If someone shows a within-creator learning curve — later launches by the same deployer graduating at materially higher rates after controlling for capital and channels — then the loop is doing candidate-level improvement after all.
- Corpus selection bias. The 495 articles are a curated encyclopaedia, so the coins in it are survivors by construction. The concentration figures in section 05 describe notable coins, not all coins, and I have not corrected for that because there is no principled way to. If the curation is skewed in some way I have not detected, the reward-shape argument weakens — though the survival study, which is a census of a window rather than a curation, does not depend on it.
- Mechanism changes that hold. BOOST pushed a single day to ~6.7%. If a redesigned incentive sustained a rate an order of magnitude higher without a matching collapse elsewhere, that would suggest the 0.198% was a tuning choice rather than a property of cheap verification.
Two things this report does not claim
- It does not claim memecoins are good, socially useful, or worth buying. About 3% of accounts ever cleared $1,000 in profit. Read that as the warning it is.
- It does not claim the market verifier is correct. It claims the verifier is fast and cheap, and that we can therefore read off what it rewards with unusual precision. Those are different claims, and conflating them is the mistake this whole page is about.
10
Sources and citation
Every figure on this page comes from one of two places: the survival study, which I verified against the primary source, or the article corpus, which I collected and recomputed myself. The study is linked in full below; the corpus figures were re-derived from the collected text.
Primary source
-
Arati Uday Kamat. Pump.fun Graduation Regime Windows: Survival Analysis of 832,941 Token
Launches and the Social-Presence Effect (Corrected, v2; incorporates Corrigenda v1.3 and
v1.4). arXiv:2607.02823. ORCID 0009-0000-4781-312X.
Full text: arxiv.org/html/2607.02823 · Abstract: arxiv.org/abs/2607.02823
Corpus
- Article corpus — 495 encyclopaedia entries on memecoins, the people around them and the platforms that launched them, collected 17 September 2026. Every figure attributed to the corpus on this page was re-derived from the collected text rather than quoted from a summary.
- Secondary figures quoted in section 04 are reported in the corpus, which attributes them to DefiLlama (cumulative revenue, share of tokens graduated), DEXTools (June 2026 graduation rate), Cointelegraph (2025 rate) and market commentary (the February 2026 drawdown). Those attributions are the corpus's; I have not independently verified each one, and they are labelled secondary throughout for that reason.
Cite this report
@techreport{trenchrsi2026,
title = {Trench-RSI: Recursive Self-Improvement in a Market That
Verifies Instantly and Rewards Almost Nothing},
year = {2026},
month = {September},
type = {Report},
note = {Analysis of 832,941 token launches (Kamat 2026,
arXiv:2607.02823) and a 495-article corpus collected
2026-09-17. All graduation rates are fast-regime lower
bounds; the channel-count gradient is correlational.},
url = {https://memecoinrsi.com/}
}
Companion report: The Verification Frontier, which makes the general argument this page tests against a case where verification is nearly free.