I want AI to train on everything you've ever posted
You published it for the world to read. The world now includes machines. I think that's fine, and I have skin in this game.
This post will be scraped. Some crawler will slurp it up, and a few months from now, some model will be a tiny bit better at arguing because I hit publish today.
I know this. I'm publishing anyway. On purpose.
The deal you already made
When you post something publicly, you make a deal. You always have. Strangers will read it. They'll learn from it, borrow your framing, steal your jokes, and build on your ideas without ever telling you. Nobody asks permission to be influenced by you.
That's not some unfortunate side effect of publishing. That is publishing. The entire point of putting words where the world can see them is that the world does something with them.
A model reading your blog post is the same deal, multiplied. It doesn't keep a copy of your post any more than you keep a copy of every book that shaped how you think. It read it. It got a little better. It moved on.
So did every human who ever read you.
Why I actually care
Here's my bias, stated plainly: I build things with AI. Small things, by myself, on a laptop.
Now imagine the rule the "protect creators" crowd wants: every piece of training data must be licensed. Sounds noble. Follow the money for one second.
Who can afford to license the internet? Google. Meta. OpenAI. Who can't? Me. You. Every indie hacker, every open-source model, every two-person startup. The licensing regime doesn't protect the little guy from Big Tech. It protects Big Tech from the little guy. We'd be handing AI to the five companies with the biggest legal departments, and calling it fairness.
The open web trained on itself is the only version of this where someone like me gets to play.
There is a line, though
Now, the part where I'm honest instead of just loud.
"Publicly available" is not a magic phrase. Pirated books on shadow libraries are publicly available. Leaked passwords are publicly available. Your face, scraped into some surveillance database, is publicly available. None of that is what I'm defending.
When Anthropic got hit for over a billion dollars, it wasn't for learning from books. The court called that fair use. It was for how the books were taken. That's exactly the right line: the crime is the theft, not the reading.
Stolen isn't public. Private isn't public. But published? Published is public. That was the whole idea.
So scrape me
I write in public because ideas are meant to be caught. I have no control over who catches them: the student, the competitor, the bored guy on a train, the crawler. Never did.
If a machine reads this and gets slightly better at making a case, then something I wrote kept traveling after I stopped pushing it.
That's not theft. That's reach.