16 Comments
User's avatar
Amarda Shehu's avatar

Mark, the chasm is real, and from an institutional vantage I would only sharpen one part of it. The chasm is growing fastest inside the organizations that most need to interpret what is happening. University leaders, and I say this as one of them, are making governance decisions calibrated to a version of the technology they encountered two model generations ago (if that!). The lag is now long enough to constitute a structural risk in its own right.

Your deterministic versus non-deterministic question is where I may push back a little. The task categories where agents have already crossed into knowledge work, LMS submissions, legal document review, sales outreach, customer support triage, share a condition that is easy to miss in the story about capability. Each had been templated, quietly, over years of interface design and workflow compression, to the point where an agent could/can now complete it because the deterministic-enough signal was already there. The template preceded the agent. What the automation did was just make visible a prior condition.

This has an implication for your prediction about the next year. If the pattern truly holds, the tasks most immediately vulnerable to agentic automation are the ones we had already rendered measurable and repeatable for reasons unrelated to AI, often to make them easier to manage, scale, or grade. What remains, and what now becomes precious, is the layer of knowledge work that resisted templating: situated judgment, institutional memory, the reading of texture that Polanyi called tacit and that most organizations discover they need only once the templated layer is gone. I have written about one version of this.

Oh, and on Mythos, a small note. It is plausible and indeed widely reported that Anthropic uses its own models heavily in internal engineering. Whether that constitutes recursive self-improvement (RSI) in the technical sense is a separate question, and Anthropic's own public position on Mythos Preview, as of earlier this month, is that they are less confident than they used to be that junior-researcher work is safe from automation but that the answer to the direct RSI question is still probably not. The difference between 'yes' and 'less confident than before' might change how one should read the tempo of what comes next.

Thank you for writing this. The chasm framing is one I had hoped to write next, as I see what my students can do with lower-tier models versus the $200/month ones. Interesting times ahead all in all.

Mark Humphries's avatar

Hi and thanks for the insightful comments. I fully agree that templated work with pre-existing metrics and evals is the lowest hanging fruit for automation in knowledge work. What I am thinking of are the areas you are talking about that have resisted those types of metrics. In these areas, creating a signal is much harder because the signal may also be highly unstable (meaning it will vary from one person or organization to another). Finding it requires a lot more work. There are also a lot of areas (and often these are properly tasks) where we haven’t attempted to measure human performance at all, or with comparing performance to machines, and so have very little data against which to test models. For example, it is maddening difficult to try and find data about how accurate humans are at transcribing handwriting. Jobs are composed of many little tasks many of which we don’t think about very much.

On Mythos and recursive self improvement: fair points and I was trying to cover something highly technical too quickly. With Mythos, Anthropic reported that its teams experienced an average 4x acceleration in productivity when using Mythos and that more than 90% of internal code is now being written by AI. I agree that is not technically RSI for a variety of technical reasons which boil down to: the model is not doing this autonomously. Even so, Andrej Karpathy, an accomplished and now independent AI researcher, recently open sourced a project in which he got an older AI system to recursively improve its own training efficiency itself, autonomously. When you start to put those two things together, at least some form of RSI seems more likely than not in the not too distant future and that is what a lot of the frontier labs have been hinting at. Maybe they are wrong, but if they are right it will happen quickly and so it’s a possibility we need to take seriously.

Tom Scheinfeldt's avatar

Thanks, Mark. It will be incredibly useful to have this post in my pocket when explaining these changes to normie colleagues.

Mark Humphries's avatar

I’m glad to hear it! The AI bubble is a weird place so it’s hard to know where to pitch this stuff.

Stephanie Decker FAcSS FBAM's avatar

Yes, very true. The AI doomers still think that bad GPT2.5 prose and hallucinated references are the apogee of LLM competence. I often cringe when I hear people I like and respect state this confidently.

Stephen Fitzpatrick's avatar

Mark - I really appreciate this post. I've been writing about a similar chasm widening in K-12 secondary education. The gap between the average person's experience with AI models and the frontier is now significant enough that people are often having very different conversations about AI without realizing it. The coding developments are real, but only a fraction of those curious and interested within the humanities are likely to experiment. I'm a (cough) late-50s non-tech person and have managed to build some interesting things using Cowork. As this divide continues - and I suspect it will - these conversations are going to get harder to have intelligently. I ran a workshop today for a group of HS history teachers, and even those with real interest in AI don't know much about tools that were new 18 months ago, let alone the agentic capabilities. Are you finding your colleagues are in the same place, or are you an outlier?

Mark Humphries's avatar

I would say I am very much an outlier on this in my immediate circle, for no other reason I think than I became interested in this earlier and so had a bit of a first adopter advantage. It’s interesting it’s similar with K-12 as I would have expected maybe a bit more adoption.

Richard Pickering's avatar

Outstanding article, thanks Mark. I agree the purpose-built harnesses are driving tremendous progress on the agentic side and contributing to the Saaspocalypse. On the other hand I would be careful about buying in too much to "recursive self-improvement", this is another one of those kernels of truth hyped well beyond its meaningful definition. I just wrote an article about it today :)

Mark Humphries's avatar

Thanks for the kind words! And fair enough on recursive self improvement although my own sense from fine-tuning models is that it is more real than not. But that’s just my experience and I am inferring.

Paul Esau's avatar

This is a great post, Mark. I think your diagnosis of why non-techies aren't able to assess the pace of change (poor free models, tech industry hype, coding ignorance) is completely accurate. And the "Saaspocalypse" of the last few months has been deeply unnerving. The engineers I talk to all say that their jobs are changing almost daily, even as their productivity has dramatically increased. I switched to Claude (and Cowork) a few months ago — it's been a revelation. I think we've got maybe a 2-year window before massive economic turbulence requires a restructuring of social (and work) expectations. And that's just on the labor side. Who knows what madness will prevail when governments can audit (or generally surveil) all their citizens all at once.

Mark Humphries's avatar

Fully agree Paul and thanks for the kind words. What is so unnerving, though I think, is that when people try systems like Cowork they immediately intuit what you are saying and have difficulty seeing any other line of development. But because it *sounds* so implausible people tend to dismiss it. But we need to find a way to have these conversations long before this happens.

Paul Esau's avatar

If there is a venue where historians are having these sorts of conversations, I'd love to join. Seems like a more productive use of my energy than arms control (at the present moment).

Jim Clifford's avatar

I’d also keep track of the Journal Historical Methods. We’re getting lots of papers and will be publishing interesting work in the coming months.

Briana Morrison's avatar

Nicely written and well-summarized. As someone who teaches computer science (specifically programming), I can tell you that most of us are all scrambling to keep up. The changing models, tasks that software engineers are asked to do today are vastly different than 6 - 12 months ago. Meanwhile we're trying to AI-proof our students for their careers.

The other soon-to-be real problem is a new type of digital divide. As you point out, using these agents isn't a cheap endeavor. When higher ed is under fire and losing funding, asking for funds to teach students to use agents that cost real $$ is a non-starter. And even with university "agreements" you don't get the resources for entire classes to pound at it for an entire semester. I predict that soon there will be CS degrees with $$ that train on the most current models, agents, and harnesses, and everyone else who uses the free or close-to-free models - further exacerbating the equity in hiring problems.

PEG's avatar

Great overview, and the accessibility gap you identify is clearly closing. But I’d push back on a few of the moves doing the heavy lifting.

There’s a pattern worth naming: personal revelation generalised to structural claim. ‘I found this transformative’ and ‘this is broadly transformative’ are different claims requiring different evidence. What you’re describing looks more like diffusion: better tooling lowering the entry point for prototype use cases that have been possible for a few years. The ceiling has risen, but it’s still well below production. The more telling question is durability: how many of your students’ apps are still running three months later, and in what form?

On coding specifically: it’s close to an ideal problem for LLMs. Constrained syntax, relatively complete context, clear success criteria. Most knowledge work lacks all three. The apparent breakthrough in coding is a weak proxy for general agentic capability. it tells us more about the structure of the problem than the generality of the solution.

The harness point follows the same logic. A harness doesn’t create an agent so much as stabilise a workflow: it constrains the model into a loop with memory, tools, feedback, and a clear goal. What looks like ‘waking up’ is better scaffolding around an already capable but unstable system. Someone built a better mill; the engine didn’t suddenly get smarter.

The SaaS argument also feels over-attributed. Many of these businesses were priced on growth assumptions that had already failed—SBC-adjusted free cash flow near zero at firms twenty years old. AI provides a compelling narrative, but the correction was structural.

The personal revelation is real. But it’s not that agents have woken up; we’ve become better at packaging models into systems that can act in well-defined contexts. Significant, and many will find it personally transformative, but uneven, and dependent on the nature of the task.

The open question isn’t whether knowledge work changes—it clearly will—but where the constraints actually sit: in model capability, or in the ambiguity, coordination, and durability requirements that surround it.​​​​​​​​​​​​​​​​