·
AI Made Answers Cheap. It Made Verification the Scarce Skill.
The marginal cost of generating an answer collapsed to near zero. The cost of checking one didn't move. That gap is the starting point for what we're building next.
If you work with AI every day — diligence, filings, code, pulling a piece of research together — you know the move: you ask, you get back something that looks complete, and then nothing happens.
Whether it’s actually right, you don’t know.
To know, you have to go check. And checking, sometimes, costs more than doing the work yourself would have.
I ran into this over and over this summer. It’s the starting point for what we want to build next.
Everyone is comparing models. I think the scarce thing already changed.
The mainstream conversation is about which model is strongest. Leaderboards turn over weekly, parameter counts climb.
But lay out the cost structure and only one thing really changed in the last two years: the marginal cost of generating an answer collapsed to roughly zero.
The cost of verifying that answer didn’t move at all.
It still takes a person opening the filing, finding the page, matching the number. It still takes a person running the code and watching what it does at the edges. None of that got automated. Most of it doesn’t even get written down.
Which leaves us somewhere strange: the world now produces “looks-right” content at almost no cost, while the ability to tell whether it’s right stays manual, one-off, and impossible to accumulate.
Writing the stablecoin book, my rule was that every number goes back to a primary source. At one point another model handed me a figure from a company’s filing that looked entirely reasonable — reasonable enough that I nearly used it. I went and read the 10-Q anyway.
Checking took far longer than generating had.
The bigger waste: nobody but me got anything out of that check. The next person with the same question starts from scratch.
“Share your chat logs” only solves the first half
What I originally wanted was simple: let people publish the whole exchange with AI, not just the answer at the end.
Because what’s valuable usually isn’t a prompt. It’s the process — how you framed it, what context you fed in, how you caught it making things up, how you pulled it back, when you gave up and switched models.
Someone who isn’t fluent with AI learns more from watching one expert’s full process than from a hundred “magic prompts.” I still think that’s true.
But it only answers half the question.
It answers: how did this answer come about.
It doesn’t answer: is this answer correct, or can anyone build on it.
And those two are where the cost gap actually sits.
Three words: traceable, verifiable, continuable
So we widened it into a loop:
Generate → publish → verify → continue → verify again.
Three things:
Traceable — you can see how the answer was arrived at, not just the last paragraph.
Verifiable — others can push back, re-run it on another model, produce evidence that overturns it.
Continuable — anyone can pick up where you stopped instead of starting over.
The third one matters most to me, because today it’s entirely missing.
When you find a genuinely rigorous piece of AI research, all you can do is like it, comment, save it. You cannot continue it.
It’s a GitHub you can only read from, never fork.
One thing about “verification” needs saying plainly
We’re not trying to get everyone to agree.
A majority can be wrong. In finance that’s obvious — consensus is usually the most expensive thing in the room.
What we want is for a conclusion to be traceable, contestable, reproducible. Credibility that accrues from being checked repeatedly, not from a vote about who’s right.
There’s a fine distinction here I’ve already gotten wrong: slapping “verified / disputed / overturned” on a conclusion looks like verification, but it’s still a vote wearing a different hat.
Only three things can actually settle a question:
- Replayable execution — same inputs, same result, no matter who runs it;
- Authoritative documents — the filing itself, the regulation itself, on-chain state;
- Time — the prediction comes due and reality decides.
Anything else should honestly be called a comment, not a verification.
The unit may need to change
The basic unit of the old internet:
Post → comment → like.
Content is finished the moment it’s published. From there it only sinks.
What we’re considering is a different unit:
Conversation → evidence → verification → continuation.
In that structure, content isn’t a dead post. It’s a node that can be checked, corrected, picked up, and carried forward.
I’m not going to make this sound easier than it is
There are three real problems here. I haven’t solved all of them, so I’d rather lay them out.
One: the cost runs the wrong way.
Content platforms normally have near-zero marginal cost — free to publish, free to read. Run the model inside the platform and it becomes the person publishing pays, the person reading doesn’t. That’s brutal for a cold start: supply has to spend money before anything exists to read.
Two: verification is a public good.
The verifier pays the cost; everyone else gets the benefit. Under pure pay-per-use, no rational person does it. That step has to be subsidized, or it’s just a nice-sounding word.
Three: generating is far cheaper than verifying, so an open plaza gets flooded.
As long as that asymmetry holds, supply overwhelms the capacity to check it. What happened to Stack Overflow in the LLM era is this exact mechanism. Which means this has to start in a domain narrow enough that conclusions can actually be settled — not as a general-purpose plaza.
When I handed the full design to another model for adversarial review, the hardest hit it landed was precisely here. I wrote that review up separately; the verdict wasn’t kind: until the verification loop is proven by real data, everything built on top of it has no foundation.
I accept that. So the first version is going to be small.
So the first version is small
Small enough that it barely reads as a new direction: you can talk to AI inside Pickful, then share the whole exchange so others can see it.
No verification. No fork. No model comparison.
Because before any of that, there are three numbers I need: does anyone want to talk to AI here; do they follow up or ask once and leave; and once they’re done, will they publish it.
If the third number is zero, everything above is speculation.
Until I have them, I’m not going to pretend I already know the answer.
The ability to produce answers is becoming something everyone has. The ability to judge whether an answer holds — and to build on it — is not.
AI made answers cheap, and in the same stroke made verification the scarce craft of this era.
We want to take that craft out of everyone’s private, one-off workflow and make it public and cumulative.
If you work with AI daily, I’d like to hear one thing: the last time you caught AI being wrong, how did you catch it?
That “how” is exactly what we want to keep.
轉發此貼文?
與您的關注者分享。
回覆