With fall approaching, many of us are contemplating how recent developments in AI might affect our classrooms. As is often the case with this tech, it’s hard to know what’s real and what’s hype. We’ve all read stories about highly capable AI agents, models that can autonomously hack into major websites, and frontier LLMs doing new math. But is it all just marketing or has something genuinely changed?
It has. For years we’ve been saying that the models were getting more capable, but as of this fall they can actually replicate human research processes end to end. To be clear, I’m not saying this is a good thing, but it is a reality in the world. Whether we like it or not, this development alone will have serious consequences.
And it’s not only the technology that’s changing. The students arriving this September are the first cohort that completed high school with AI from grades 9-12, bringing with them different expectations about both the technology and their own futures. In combination, it means this year is going to be the inflection point when it’ll no longer be possible to just muddle through.
The Intelligence Revolution of 2025-26
One of the biggest issues with discussions about AI is that there is a ton of hype. So what’s real and what’s not?
Agents first woke up in the winter of 2025-26 when they proved capable of doing useful, autonomous work in coding through Codex and Claude Code. Over the next months, those capabilities began to extend out into knowledge-work more broadly with tools like Claude Cowork and then ChatGPT Work.
This is where it becomes important to use nuance. These systems are now “good enough” to do many of the tasks one routinely does on a computer autonomously (or nearly so). In this context, “good enough” means tasks on which AI makes mistakes in predictable ways and where those mistakes are no more frequent or consequential than the ones made by people.
There are a number of factors that converged to make this possible. First, the underlying LLMs became more capable, meaning they now make fewer mistakes, draw better inferences from evidence, and can do long running tasks for longer periods of time. The release of Anthropic’s Mythos and then Fable 5 this spring and summer demonstrated that scaling pre-training (meaning increasing the number of artificial neurons and the amount of training data) still unlocks new, emergent capabilities at or above long-standing trend lines. This is important because it suggests that even pure scaling (model size) still has a lot of runway.
Meanwhile, new harness architectures have given LLMs more tools with which to do things in the real world while scaled reinforcement learning (RL) trains LLMs to use those harnesses more effectively. RL also trains LLMs to know when they’ve achieved a goal and how to verify work. At the same time, inference scaling, which means giving the models more time and the tokens with which to think, lets AI systems handle longer and more complex tasks.
In the spring of 2026, these factors made AI systems capable of doing new things like autonomous hacking and doing new math. Some argue that this is just marketing hype, but I would suggest that such claims can simultaneously be both accurate and useful hype.
Hence the new math results, which are a different beast altogether as they’ve been verified by the world’s leading mathematicians. If you want to get an objective sense of the actual trajectory we’re on, remember that three years ago LLMs couldn’t reliably multiply multi-digit numbers. Scaled up and strapped into the right harness, they’ve now made the largest advance in nearly forty years on a problem closely tied to the Riemann Hypothesis. That’s a big deal.
The key point is that AI systems don’t need to get any better than they are today to fundamentally change what we do and how we do it in a lot of areas of knowledge-work. That doesn’t mean that the robot butlers have arrived or that AI companies can automate whole jobs today, but they can reliably start to automate the tasks of which jobs are comprised.
As I’ll discuss below, the real problem for historians (and others in the Arts and Social Sciences) is that it’s the process rather than the final product that we claim to teach at university. It’s true that no one wants or needs an AI generated undergraduate paper, which is why its the skills required to produce a paper that actually have value. So if the process of doing research can be automated i, what happens to the value of the equivalent human skills?
AI and the Classroom in 2025
The key insight of the last year is that whereas in 2025 AI could produce whole papers but couldn’t emulate the process, it can today. To start, cast your mind back a year ago. In the fall of 2025, the best LLMs could write capable but boring papers, although they couldn’t cite page numbers—even if you gave them the PDFs. They were also just starting to be able to use a web browser, but unreliably.
Back then, the LLM process of writing a paper bore no resemblance to its human equivalent. The LLM was given a prompt, it maybe thought for two to eight minutes, and then spit out an answer with imperfect citations and hallucinated quotations. In other words, the text was generated ex nihilo in a sort of machine-equivalent to stream of consciousness—hallucinations and all. The result was a pale caricature of what we do as historians or what we want our students to do. I think that’s probably what enraged many of us the most: LLM results lacked rigour, substance, and human insight so they mocked the process of doing history as much as the results.
AI in the Classroom in 2026
Today agents can follow an essentially human-like process, end to end. To be clear: I am not saying this is a desirable thing or something I would want a model to do for me personally or that I want to teach students to do in a history class. My point is that whatever you might think of this capability on an aesthetic, moral or ethical level, it is what’s available in products like Cowork and ChatGPT Work. One of the major barriers to understanding what this means, though, is that many people don’t have the $100-$200 USD per month subscriptions necessary to access the best models, so they’ve never seen it in action.
To demonstrate these abilities, I dumped an old assignment from a First World War class into Claude CoWork with the Fable 5 model selected. The essence of my prompt was simple in the extreme: “complete the assignment” followed by the assignment text from the syllabus.
For ethical reasons, I modified the assignment a bit to limit Claude to using open-access journals in which the articles are legally available AND where the journal’s terms of service do not prohibit AI or bot access. I wanted to see how it would do with books too so I also gave the model the text of six of the books I’ve written or edited on the First World War along with a number of my peer-reviewed essays. It also had to find primary sources online in repositories—ones that also don’t block bots.
Fable 5 responded with some questions about what I wanted, including the level of writing (senior undergraduate or graduate?). Now I’ve written pretty widely on the Great War (shell shock, military operations in the CEF, tactics, self-inflicted wounds, translations of the German Official History, battlefield guides, Arthur Currie’s leadership, etc) so I was curious to see what it would choose based on the folder I provided. After thinking for ten minutes it decided to “go with shell shock in the CEF; the folder supports it best. Staging the relevant PDFs (skipping the very large battlefield guide and Germany’s Western Front volumes)”. Fair enough.
It went away and created a staged plan as follows:
1. Stage and read the folder PDFs on shell shock
2. Locate open-access secondary and primary sources
3. Write proposal and preliminary bibliography
4. Write the 12-15 page research paper
5. Verify citations, quotations, and page count; build PDFs
Next, Fable 5 worked for a long time! You don’t see much in the actual conversation stream anymore with these newer harnesses but the image below shows one aspect of what it looked like from the user’s perspective. First, one subagent read and took notes on the books and journal articles I’d written and put in the folder while a second searched the web for open-access sources.
Watching it work in Chrome specifically is unsettling and fascinating. It operates the browser like a human would, moving a cursor around the screen, clicking on links, and typing text into entry boxes. It will navigate to sites on its own, snap screenshots, try a bunch of different increasingly complex Boolean keyword search terms and evaluate the sources it finds. What’s most fascinating is that it looked at literally dozens of papers but was highly selective in the way I’d want a student to be. It also found my department’s citation guide on the web and used the version of Chicago Style required there.
Just so that readers are aware: Claude will stop and ask you to login to sites like JSTOR or EBSCO. If you have the credentials to do so, you can. It violates terms of service, university policy, and perhaps copyright, to be sure, but rest assured students can and will do this.
The result was an impressive paper that met all the requirements of the assignment. I checked the block quotes against the originals and spot checked page numbers (as I’d do with a student paper) and all the ones I looked at were correct. The writing is not very good nor engaging and the findings are pretty boilerplate but it’s an A to A+ paper.. You can read it by downloading the pdf below.
The Demise of the Research Paper
To my mind, the traditional research paper as a tool of differential assessment is now dead. And I think it should be left to rest in peace.
Very few student papers have ever been read by anyone other than the student and the professor. There are exceptions, to be sure, but they’re rare. The point of the undergrad paper was to teach the student how to go through the research and writing process and, in the course of doing so repeatedly over several years where the requirements became successively more onerous, to hone their analytical and interpretive skills. The idea was that one left a history degree not only with knowledge about the past, but with an understanding of how to find things out about the past in a rigorous way. It was the skill of conceptualizing a problem, amassing the best evidence, and distilling an argument that was valuable to both the student’s intellectual life and to future employers.
One of the most common and persuasive arguments against AI in the history classroom is that students need to go through the act of wrestling with a problem, reading the sources, and the struggle of writing to learn how to think critically and analytically. But I’ve talked to lots of people who want to try to retain the research paper (myself included until very recently) by altering the conditions of its production so as to avoid AI plagiarism. This might mean more paper/less tech focused work, that is having students work on the paper in class, in supervised settings outside of class, to write it out by hand, to use only paper sources, etc. Done well, this might work and Edward Dunsworth reports good results.
But in my view, if it’s the traditional process that’s intrinsically valuable, contorting it to avoid AI at all costs risks reinforcing (or exacerbating) many of the problems posed by the tech in the first place. The risk is that we’ll lose focus on how students have to exist in the world they actually live in, not the one we wish still existed. For reasons I’ll explain below, I don’t think that argument will get any easier when this fall’s cohort arrives.
Devaluing Knowledge-Work Skills
As my colleague Jim Clifford recently argued to much opposition on BlueSky, the significance of the automated AI paper is not that it’s an object of intrinsic scholarship or that it is something of actual value. It’s that the process itself has been automated.
No one wants or needs AI generated undergraduate papers except dishonest students. People often remind me that there are also lots of fields where AI cannot even complete the assignment because of the nature of the sources. Fair enough. But the fact of the matter is that it can complete the research process in a lot of cases, much as a human would do.
This means that it’s also going to start to be useful in many of the tasks that our university and college graduates would normally be paid to do after they graduate. The technical importance about the automated process I described above is that it is intelligible: you can reconstruct how the AI agent completed the task. This makes verification much easier even while the models themselves are also becoming more reliable. From the conversations I’ve had, it’s been the opaqueness of the AI process as much as fears about reliability that have limited more widespread institutional and corporate adoption of AI. And so that will change.
I often forget that while most of my history students love history, few become historians. Most of the time they go into jobs where they process information in some way, utilizing the skills they learned through the process of doing history to develop market assessments, public policy, corporate position statements, not-for-profit grant applications, and too many other tasks to even imagine. These are often the types of functional tasks where aesthetics matter little. Sure, many thing in knowledge-work will still require humans—perhaps most. But it’s not a zero-sum game: if 20-30% of knowledge-work tasks are automatable today, that is going to have a profound effect on the value of a university degree.
The End of a Strange Détente
I think we all understand that we’ve been muddling through and doing our best for the past few years. At least that is how I’d characterize my own approach: I’ve tried to incorporate AI where I can while also maintaining traditional standards. I’ve tried, but it’s not working—at least not for me, anyway. The thing is, though, that I think what I’ve also come to realize is that I’ve only been able to survive thus far on the goodwill of my students.
In general, I think most students I’ve had over the past 3-4 years have understood and accepted that AI threw an unanticipated and unwelcome wrench into the gears of higher education because that’s exactly what happened to most of them too. In our own annoyance with the tech, we often forget that it’s disrupted their career plans (or at least their perception of those plans) as much or more than it’s disrupted our classrooms. The consequences for them personally are also much greater.
Because I’ve been teaching some GenAI focused courses, I get the rare chance to hear from students in something of a safe space. In fact, they often unload their frustrations during class discussions. Many feel AI’s made it a profoundly unfair time to be a student. Leaving aside housing and the cost of education for a moment, they talk openly about the futility of trying to honestly complete assignments without AI and remain in competitive programs as they watch classmates beat them by a long margin when using the tech. Remember that many programs grade on a curve.
Students also wince when instructors brag about being able to spot machine generated text or that they’ve invented “AI proof” assignments because most students know they’re wrong. Some resent classes in which they are made to do all their work “in class by hand”—sometimes outside of normal classroom hours—because it’s not what university was supposed to be about. I personally know I would have hated history if I had to write my papers in class: I reveled in sitting in the library for hours going through the Foreign Relations of the United States volumes and just exploring the sources on my own. To a one, they all worry about being falsely accused of AI use because they understand the difficulty of proving a negative.
Yet they also feel that they need to understand the tech and how to use it to be competitive in the job market. Over the past few years I’ve watched resentment grow as students complain that the things they’re learning to do at university are becoming increasingly detached from the external reality of the world of work. My computer science students are especially distraught because they aren’t allowed to use AI in class to write code, yet prospective employers increasingly treat fluency with AI coding tools as a baseline expectation.
But for all this, I really have gotten the sense that most of this cohort of students share a sense of loss with us: they planned to go to university in a different time and then the world changed around them—and without their consent. In my experience, many students are just as angry at AI as their professors, but for different reasons. But this détente, such as it is, is soon going to end.
The First Post-Pandemic AI Generation
This fall, just as Cowork effectively automates the skills we teach, we are also going to encounter a different group of students whose expectations about the world were honed on different steel than ours. This cohort—the future class of 2030—were born into the age of the smartphone and streaming. They started middle school (grade 7 where I live) six months into the pandemic and don’t share the same normative assumptions about pre-2020 educational standards that most professors still use as a mental benchmark (at least I know I do). Like the last cohort, technology was a central component of their lives from birth, but the pandemic only deepened its association with education: Zoom classes, learning management systems, and online testing are deeply familiar, not exceptional.
They were in grade 9 when ChatGPT appeared in the fall of 2022 and so completed all their high school classes, post-secondary prep, and university applications in an AI-enabled world. Numerous surveys suggest that the vast majority used AI for homework and assignments, at least in some capacity: 70% of American high school students in one survey this month, 84% in another, and 94% of British students. They’ll have a much deeper consumer-level familiarity with the tech, the tools available, and how to use them than many of their professors (which is a different thing than saying they understand how it works). Many will come from high school classrooms in which their teachers used AI themselves and in which they will have been expected to use it too. They’re also going to enter a weird world in which some professors are prepping lectures with AI or using it for marking (yes, it’s happening, just quietly) while others will make them handwrite essays to avoid it.
I’m not sure the goodwill I wrote about above will continue with this AI-native generation. For them, the tech is simply part of the world they’ve come to know in their formative years. It may be one factor amongst many that makes them pessimistic about the future, but it hasn’t changed their fundamental sense of the rules of the game because they learned those rules when AI was already in the world. This is not to say that they necessarily like AI: studies show that most are deeply skeptical about AI and its effects on society while also using it constantly. But unlike the last batch of senior level students I taught, this cohort came of age with that nuance baked into their expectations, so I’m not sure we’re going to see that goodwill continue.
Conclusion
The automation of the research process and the resulting loss of the human skills those processes embody poses an existential threat to universities. Yes, students come to university for a wide range of aesthetic, humanistic, and practical reasons, but there is a reason we’ve been marketing “transferable skills” for decades. We don’t have to like that reality for it to be true.
My problem is that I don’t have any easy answers. I feel fairly confident in my ability to identify the problems AI poses and to intuit some of its probable second and third order effects, but I am at a loss about what to do about them.
I have been writing about this since the spring of 2023, when I argued that if the trends I saw emerging then continued, we’d eventually be “left with a very small number of very dedicated students, probably far fewer than our large network of post-secondary institutions were designed to accommodate.” At the time I thought decline was probably inevitable either way and that the only real question was the slope. Three years on, I would stand by that. What has changed is that the slope is no longer hypothetical and we’re approaching the real inflection point as both cultural changes and technological changes converge.
I wish I could provide a list of AI-proof assignments beyond in-class testing or essays (which still have their place when done well), but I don’t think that’s either possible or desirable at this point. At the end of the day, I think that if the value of a history degree is tied to what students learn through the process of actually doing history, then we are going to have to assess that process directly rather than inferring it from a paper. That means watching students conceptualize a problem, weigh conflicting evidence, and build an argument, perhaps in seminars, in conversation, orally, and across drafts, whether or not they used AI to get there.
Assessing process means smaller classes and more contact hours, and outside of a few well-resourced programs we are not set up to provide them. That is a resourcing and collective action problem more than a pedagogical one, and we should conceptualize it in those terms. Inevitably, people will experiment with using AI to make that more hands-on workload manageable. I’m also not sure that’s a good thing.
We also need to decide, deliberately, which parts of the process we are willing to hand to the machine and which we are not. I have made the case with research that the most useful AI systems are those that reduce monotonous work while augmenting human judgement rather than replacing it. A non-human historical assessment that no one can verify is worth very little.
The same logic applies in the classroom. I have argued before that students need to leave university knowing when to use AI, when to avoid it, and how to get the most out of it, and I still think that’s true. But universities also have value beyond skill building.
There is a version of the future in which humanistic work becomes more valuable, culturally and socially if not economically, precisely because it is not machine driven. But that future will not arrive on its own, and it will not be built by institutions that spend the next few years litigating whether a student’s paragraph was written by a person. That will just turn students off and confirm their worst fears.
Finally, we owe this incoming cohort some candour. They are not responsible for the world they’ve inherited and it will be jarring when they’re told they can’t use the tools the university provides them because they contain AI (IE Word, PowerPoint, Google, Sheets, Adobe Acrobat, JSTOR, etc). The goodwill of the last few years existed on the understanding that we were all muddling through this together. That understanding is ending.
The inflection point, then, is not that AI can write a research paper. It could do that two years ago. It’s that it can now do everything we once said mattered more than the paper itself and we can’t reasonably pretend otherwise.





That sentence ("The writing is not very good nor engaging and the findings are pretty boilerplate but it’s an A to A+ paper.") stood out for me as well, however, I have a different take. What I find with a lot of people who, like Mark, see the enormous transformation that is taking place and try to explain that to a readership that is not receptive, is that they (perhaps subconsciously) throw in phrases like that to "soften the (existential) blow" (perhaps to themselves as well).
There is no longer any reason for the writing to "not be very good." For a long time now, all you have needed to do is to prompt the LLM to write in a certain style, or to upload some text from someone and to instruct it to write in that style. And the savvy students already know this.
Now with Claude Cowork, you just add instructions. I haven't looked, but I'm sure that there must be "good-style-academic-writing.md" instructions that are circulating out there so students don't even need to think of what "good writing" might be. They can just drop those instructions into Claude and "voila"!
Further, from my own experience and experiments (still only using $20-a-month models), we're moving beyond the point of "boilerplate" conclusions. This has become particularly noticeable to me over the past month. I just had a skeptic the other day ask me to prompt an LLM about a rather niche topic regarding Vietnamese music (one that required figuring out different notation systems, etc.). According to my colleague, it produced a conclusion that no one had ever made before, and that he had never thought of before. And that was with zero follow-up prompts, etc. My own experiments on various topics produce the same outcomes.
Perhaps if the assignments are about well-worn topics, then yes, the LLM might produce boilerplate conclusions, but so would 99.99999% of undergraduate students in the pre-AI age have done the same.
I think we are entering a phase where there are people who know how to use LLMs to get the best results (who ask it to come up with the best prompt first, rather than writing it on their own or just dropping the prof's assignment in, and who add all kinds of instructions to Claude Cowork, etc.), and people who don't. Those who don't know will give the impression that LLMs are not capable of doing what historians can do (and that makes us feel like we still have a place/role). However, it's the papers written by the people who DO know what to do that we need to pay attention to.
It will be interesting to see on which side of the divide the AI-native students will fall. Just because they grew up with it does not mean that they are better at it (laziness/disengagement and procrastination also play important roles). But the fact that some cannot produce good papers doesn't mean that the LLMs are incapable. They are totally capable at this point.
As such, I completely agree with Mark that "To my mind, the traditional research paper as a tool of differential assessment is now dead."
And like Mark, I have no idea what to do next. . .
Excellent! I would have given a master’s student a “B-“ and had her/him re-write after reading my lengthy critique. I would much rather a student re-think, re-write, etc. one paper until it earns an “A” than write 3 separate papers in a semester that aren’t good graduate work. It’s that iterative process that forces deeper thinking and mastery. It also moves the student away from dependence on AI in the latter stages of real scholarship. Yes, I even had doctoral students who hated that “polish and polish again” process, but many thanked me later for not giving up on them and for believing they were more capable than they themselves believed.