Proof Used to Be Expensive

AI did not just make better content. It removed the cost of faking things, and cost was what we were trusting.

Rashid Azarang & Jesús Carlos Acosta Rocha18 min read
Proof Used to Be Expensive

Written by Rashid Azarang and Jesús Carlos Acosta Rocha. Published on both of our sites.

Start with something that sounds too simple to be important.

Almost everything you believe about the world, you believe because faking it would have been too much work.

That sentence is the foundation. Everything else in this post is built on top of it, and everything that is about to change in your life comes from that foundation being removed.

1. A signal is only worth something if it costs something to fake

A photo of you standing on a beach was evidence you went to the beach. Not because photos are truthful, but because staging that photo cost more than simply going.

Your voice on the phone was evidence of who you were. Not because voices are unique in some magical way, but because imitating one well was hard.

A long, carefully argued email was evidence that someone cared about the topic. It took an hour to write. Nobody spends an hour on something they do not care about.

A signature, a diploma, a passport stamp, a handwritten note, a video of an event. None of these were ever proof of truth. They were proof of effort. We treated effort as truth because, for all of human history, effort was the cheapest reliable stand-in for honesty.

This is not a new idea. It is the same reason a peacock's tail means something. The tail is expensive to grow. A weak bird cannot afford one. The cost is the message.

Remove the cost and the message disappears.

2. The cost just collapsed

The common way to say this is "AI content is getting better." That framing is too soft, and it points at the wrong thing.

What actually happened is that the price of producing a convincing artifact fell by orders of magnitude, and it keeps falling.

A convincing video used to need a camera, a crew, a location and days of editing. Now it needs a sentence. A convincing voice used to need a talented impersonator or a studio. Now it needs a few seconds of sample audio. A convincing argument used to need someone who understood the subject. Now it needs a prompt.

Notice which number moved. The cost of producing something believable dropped to near zero. The cost of checking whether it is true did not move at all. Checking still takes a human being time and attention, same as in 1995.

Any time that gap opens up in any system, the system floods. That is the entire history of email spam in one sentence. Sending got free, reading did not, so the inbox drowned.

Watch where the flood arrived first. OpenAI opened ChatGPT to the public on November 30, 2022. The businesses disturbed first were not the ones in the think pieces, not radiology, not law. It was copywriting, content farms, SEO filler and low-cost publishing. Those markets had already agreed that plausible was enough. Nobody buying a product description for the price of a sandwich was buying insight, only grammatical English that fit a slot. When a machine could produce that, the market cleared quickly, because the bar had never had anything to do with origin.

Here is what that looks like from the buyer's side. Somebody in your family bought a children's book recently. Thirty pages, a rhyme scheme that mostly works, an illustrated fox in a red coat, four dollars, delivered in two days. The kid asks for it three nights running. Now ask a question that would have sounded strange in 2019: did a person write it? You cannot answer from the object. The rhymes land. The copyright page lists a name you have never heard, which is true of most children's books. Nothing in the artifact tells you whether it came from an author at a kitchen table or from a prompt run forty times until the meter behaved.

The polish of an object no longer carries much information about its origin. You used to reason backward from quality to effort, and from effort to a human. That inference weakened quietly, and it weakened first exactly where nobody was looking.

3. These systems are not aimed at reality. They are aimed at you

Here is the part most people miss, and it is the part that makes "you will not be able to tell" literally true rather than dramatic.

A generated video does not have to match reality. It only has to match your expectations of reality.

Those are different targets, and the second one is much easier to hit. Real life is full of odd lighting, awkward pauses and things that look wrong but are not. Your brain does not compare a clip against the world. It compares the clip against a rough internal model of what such a clip should look like.

These systems are trained on enormous amounts of human-made material, which means they are trained directly on that internal model. They are extremely good at producing exactly what a person expects to see.

So the fake is not competing against the truth. It is competing against your priors, and it was optimized to win that specific fight. Reality was not optimized for anything.

This is also why the fakes will feel more real than real footage. They already do, sometimes.

4. Therefore detection is the wrong layer

The first thing everyone proposes is a detector. Some algorithm that reads a video or a piece of text and tells you the probability that a machine made it.

Detectors are useful and we should build them. They will not save us, for three reasons that come straight from the principles above.

First, any published detector becomes a training target. The moment you can measure "does this fool the detector," you can optimize for it. The forger gets to practice against the referee. This is not hypothetical. Krishna and colleagues built an 11-billion-parameter paraphraser, DIPPER, for exactly this purpose, and running machine text through it dropped DetectGPT's accuracy from 70.3% to 4.6% at a fixed 1% false positive rate. It also evaded watermarking, GPTZero and OpenAI's own classifier.

Second, a detector returns a probability, and humans read probabilities as verdicts. A number that says 12 percent will be read as "it is real." This makes people less careful, not more. And the errors land on real people. Liang and colleagues found that GPT detectors misclassified more than half of TOEFL essays by non-native English speakers as AI-generated, while scoring near-perfectly on essays by US eighth-graders. The detectors were not measuring authorship. They were measuring how closely a sentence resembled fluent native English, and calling the difference fraud.

Third, and this is the deep one: the detector is trying to answer a question that no longer has a general answer. "Is this artifact fake" was only ever answerable because fakes were expensive and therefore rare and therefore sloppy. The artifact itself no longer carries the information.

If the cost is gone from the artifact, then the cost has to come back somewhere else, or trust does not work at all.

The place it comes back is the source. Not "does this look real" but "who is standing behind this, and what do they lose if it turns out to be false."

That is the difference between detection and provenance, and they are two different products. Detection looks at a finished artifact and guesses. Provenance is a record of origin: attached when the thing is made, naming who made it and what happened to it since, with somebody signing for the claim. It carries a history, or it does not exist. Tellingly, the defense Krishna's team proposes against their own paraphraser is for the provider to keep a searchable database of everything it ever generated. That is a provenance system in a detector's clothes.

Provenance is honest about its limits in a way the discourse around it is not. The C2PA specification says plainly that Content Credentials record what happened to a file and who signed for each step, and that provenance alone cannot tell you whether content is true. A signed photograph of a staged scene is a signed photograph of a staged scene. A chain of custody also binds only the parties who agree to be bound. The uncooperative path is a screenshot, a re-encode, a phone pointed at a monitor, or an open-weights model that signs nothing. Meta admitted the seam when it announced its labeling work in February 2024: image generators were starting to cooperate on IPTC and C2PA markers, but audio and video tools were not doing so at the same scale. So for the two formats where synthesis is most convincing, both Meta and YouTube fall back on asking users to disclose.

Even where disclosure exists, look at where it goes. Amazon's KDP rules require publishers to tell Amazon when a tool created the actual text or images of a book. They say nothing about the customer ever seeing that. So the fact may exist, recorded in a system you cannot query, while a parent decides whether to spend four dollars.

Still, provenance works where it works for one reason: a named person with a reputation can be ruined, a company can be sued, a camera can sign its footage at the moment of capture, a friend on the other end of a call can be asked something only they would know. All of those are still expensive. That is the only reason they work.

5. A feed does not die from lies. It dies from not being worth checking

Every claim on a screen carries a verification cost, the effort required to establish whether it is real. Every claim also carries a value, what knowing would be worth to you. People check when the value exceeds the cost, and they are efficient about this without ever announcing it.

Synthesis raises verification cost across the board, including for all the true things. Our guess is that for a typical post the cost crossed the value some time ago, and that the response is to stop treating the category as evidence and start consuming it as entertainment.

We cannot prove that at population scale. One of us can report having stopped reverse-image-searching things he would have checked in 2019, and not having decided to stop. It went the way habits go.

The strongest objection is that feeds were never evidence. People open these apps to see what people they know are doing. That is a social function, and social functions no more require verification than waving at a neighbor requires a signature. Staged photos, filters, purchased engagement and bot networks predate generative models by a decade, and users adapted years ago by reading feeds as mood rather than record.

We take that seriously, and there is one answer. What the people you know are doing is now synthesizable too. The fake that matters is no longer a fabricated news event, which feeds had already stopped being trusted for. It is a thirty-second clip of your cousin, in your cousin's voice, saying something your cousin never said. The social function rests on a minimal evidential floor: that the person shown is the person, and that the thing shown happened in some form. That floor sits far below the one everyone argues about, and it is the part now under pressure.

Since this is a theory, here is what would prove it wrong. We would treat it as falsified if engagement with unverifiable content holds or grows over the next several years while migration into small accountable spaces stays flat. We would treat it as supported if group chats, thirty-person servers, mailing lists, personally owned sites and rooms with actual chairs keep absorbing the attention that used to go to public feeds, while those feeds keep their numbers and lose their function. That second outcome is a split rather than a collapse: the feeds stay profitable and stop being consulted. We notice we are describing our own behavior, and that publishing on personal sites is exactly this move. We are inside the sample, which is a reason to discount us.

6. An interface is the price of a machine's ignorance

Now the second axis, which is where this stops being about content and starts being about your house.

Ask why a screen exists at all.

A keyboard, a mouse, a menu, an app, a form with twelve fields: every one of those is a translation layer. It exists because the machine cannot figure out what you want, so you have to encode your intention into something it can process. You do the translation work, every time, for free.

Now look at the last forty years. Punch cards to command lines to icons to touch to typing in plain language. Every single step removed translation work from the human and moved it into the machine.

Extend that line. The endpoint is not a better screen. The endpoint is no interface at all, because the machine understands the intention directly.

The screen was never the destination. It was scaffolding around the fact that computers did not understand us.

7. And the price of no interface is your context

Here is the uncomfortable consequence, and it is arithmetic, not paranoia.

For a system to act without you translating your intention into taps, it needs to already know your situation. It needs to know what you are looking at, what you just said, who is in the room, what you were doing ten minutes ago.

There is no clever way around this. Removing the interface requires adding context. Context comes from sensors. Sensors have to be near your senses.

That is why the hardware is moving onto the body and into the room. Earbuds, glasses, watches, pendants, cars, appliances, doorbells. The most visible recent example: in August 2026 a demo video accidentally left inside an Apple release build showed earbuds with infrared cameras feeding what they see to an assistant, though Bloomberg reports the product is still aimed at 2027. The specific product does not matter much. The direction is what matters, and the direction is not a marketing choice. It is what the logic demands.

So the deal on the table is this. You get an assistant that knows what you mean before you finish the sentence. It costs a microphone and a camera that live with you.

We are not going to tell you to refuse it. We do not think refusing is realistic, and we are not sure it is even wise, because these tools genuinely make people better at their work and give a lot of people access to help they could never afford before. But it should be a decision you make on purpose, with the price written on the label.

8. Nobody has defined intelligence, so stop arguing about the word

There is a debate that eats a lot of oxygen: is any of this "real" intelligence.

Look at how we have actually used the word. For decades, intelligence meant chess. Then a machine won, and chess became "just search." Then it meant translation, then image recognition, then conversation, then writing code. Each time the machines arrived, we moved the line and said that part was never the real thing.

That is not a definition. It is a list of things machines cannot do yet, and the list keeps getting shorter.

A definition that moves every time it is tested cannot be used to plan anything. So drop it. The label question is unfalsifiable and it decides nothing.

The useful question is about capability. What can these systems do this year. What is the trend line. What would I do differently if they were twice as capable in eighteen months.

We will be honest about the limits of our own knowledge here. Nobody knows the ceiling, and nobody knows the timeline, including the people building these systems. We think it goes further and faster than most people expect. We could be wrong about when. We do not think we are wrong about the direction.

9. What keeps its value

If cheap things lose their signal, then value moves to whatever is still expensive. Not expensive in dollars. Expensive in the sense of hard to fake. Trust has relocated before, from artifact to institution: photography absorbed retouching, print absorbed the forged pamphlet, and society kept working. It is relocating again, and this is the one claim in this post we would bet money on.

Presence. Being physically in a room with someone is still the most expensive signal available. It costs time you cannot recover and attention you cannot copy.

Liability. Someone who can be fired, sued or publicly ruined is carrying a real cost. That cost is what makes their word worth something. Institutions that accept liability get more valuable, not less.

Track record over time. Years of consistent behavior are expensive because time is the one input nobody has learned to synthesize. A ten year history is still a ten year history.

Software somebody else can run and check. A result that ships with its code and data is expensive to fake because anyone can re-run it. Polish is free. Reproducibility is not.

Good questions. If answers are becoming free, then the scarce skill is knowing which question to ask and being able to smell when an answer is wrong.

That last one has a direct product implication that we think somebody should go build.

Nearly every AI tool today is optimized to deliver the answer faster, because speed is what looks impressive in a demo. But if answers are collapsing in price, then a tool that gives your child answers faster is optimizing the exact thing that just became worthless.

The valuable tool for a nine year old is the opposite. One that deliberately withholds the answer and hands back a better question. One that protects the struggle, because the struggle is where the thinking gets built. That is not nostalgia. It follows directly from where the value moved.

10. What to actually do

Each of these comes from a principle above, not from a list of safety tips.

Stop asking whether something is real. Ask who is behind it and what they lose if it is false. The artifact stopped carrying the answer. The source still does.

Set up a second channel for anything that matters. A word only your family knows. A callback to a number you already had. Identity is no longer proven by a voice, so it has to be proven by a shared secret or a separate path.

Treat fast emotion as evidence of manufacture. Content built to match your expectations will find your triggers before it finds your reasoning. If something spikes you in two seconds, that is the moment to stop rather than share.

Prefer signed and sourced over polished. Polish is now free. Provenance is not. Give your attention and your money to people and institutions that put their name and their exposure on the line.

Assume the sensors are on and learn how to see when they are. As the interface disappears, knowing when a camera or microphone is live becomes basic literacy, like knowing where the exits are in a building.

Train the question, in yourself and in your kids. It is the one skill on this list that is going up in value rather than down.

And one for whoever builds the platforms. The cheapest fix in this entire field needs no new technology: show the buyer what the seller already told the retailer. Under current KDP policy, if that four-dollar children's book was machine-generated, the publisher was required to say so to Amazon. The fact would then exist, three clicks and one permission boundary away from the parent holding the money. Nobody has shipped the three clicks.

The short version

For thousands of years, "did someone go to the trouble" was a good enough substitute for "is this true."

That substitute just expired. Not gradually, and not only for videos on your phone.

Everything else follows from that one sentence.

Sources and further reading


Rashid Azarang builds production AI-agent systems and publishes pre-registered research on how they behave. Jesús Carlos Acosta Rocha writes at carlosacostarocha.com and is on LinkedIn. This essay appears on both sites.

More from the blog