I taught myself to code as a 10-year-old in order to build “AI.” I was awed by artificial life simulations (Polyworld in particular; I learned C from a For Dummies book to hack on it) and wondered what the logical conclusion of it all could be. When would we cede the world to machine intelligence? Sometime in college (~2007), I stumbled onto a website called Overcoming Bias (later LessWrong) and Eliezer Yudkowsky’s Sequences. He made quite a compelling argument that if we build AI, actual true AI, it would quite plausibly kill us all. I was pretty convinced, and in fact still am to this day! A few years later I even donated $20k to his nonprofit to put my money where my mouth was.
And yet, today in 2026, the year of our lords Fable and Astra, I’m not particularly worried about these systems independently outgrowing our ability to control them. Why? I use these models daily and love them, and they keep getting better, but I keep encountering the same kind of failure. Everything starts off great (at least given enough context and a familiar enough domain), but as things get more complex or confusing or prolonged or whatever, they get lost. You know the feeling. It’s when you typically yell at the thing and start a new session. It’s more than context rot. It’s an inability to accelerate.
In general, models pursue work with extraordinary fluency. When they fail, it’s up to me to diagnose the misconception, repair the framing, and send them back to work. The accomplishment enters the machine’s ledger. My contribution disappears into “prompting.” A human approaching a task is often the opposite: we start slow and stupid, but eventually learn enough lessons to pick up speed, then come up with clever new ideas to build upon and accelerate further.
Does more post-training solve this, with more diverse tasks and longer horizons? I’m skeptical. The pattern has been pretty consistent. In any domain where the rules aren’t fully deterministic and known in advance (as they are in math and similar games), the models struggle once they’re outside of training and anything changes. Can you keep training? Sure, but we only know how to do that with lots of data, and the real frontier is an environment where every sample is costly. You can’t do a million rollouts on most actions that matter.
Transformers can dominate humans at every bounded task whose goals and evaluation criteria someone else has supplied, and still not be “intelligent” or able to replace all of our jobs. The real world changes in real time. The rules of the game are constantly being reinvented, and are never truly known to begin with. Requirements change. Feedback is ambiguous. Yesterday’s useful assumption becomes today’s mistake.
Capability ^ Humans / | / | / | AI ____________/____ | __/ / | _/ / | / _/ | / __/ | / ___/ | / ____/ | / ___/ | /_/ +------------------> Experience
When the AI line looks like the human line, that’s how I’ll know we’ve created real AI (or we’ll all be dead by then, probably around the same time). To me, sustained autonomy (the kind intelligence provides) is all about adapting to new things. The world is constantly in motion; anyone who has used these tools over the past year doesn’t need reminding. Their usefulness is already extraordinary. What I’m questioning is how much of their direction they can supply themselves. Does it compound or dissipate?
That dynamic is why I’m skeptical of a smooth extrapolation from today’s agents to systems that improve themselves beyond our control. The crucial question is whether their improvements also reduce their dependence on human judgment. Can today’s AI still cause damage? Of course. Can it behave in misaligned ways and do things its operators didn’t intend? Also of course. That’s to be expected, given how these models are trained and how much they do on the way to accomplishing goals. But does that mean we have AGI or RSI right now? No, and it’s not even clear we’re on the right path. But I remain open-minded, and especially grateful that we’ve invented these wonderful tools.