That sentence ("The writing is not very good nor engaging and the findings are pretty boilerplate but it’s an A to A+ paper.") stood out for me as well, however, I have a different take. What I find with a lot of people who, like Mark, see the enormous transformation that is taking place and try to explain that to a readership that is not receptive, is that they (perhaps subconsciously) throw in phrases like that to "soften the (existential) blow" (perhaps to themselves as well).
There is no longer any reason for the writing to "not be very good." For a long time now, all you have needed to do is to prompt the LLM to write in a certain style, or to upload some text from someone and to instruct it to write in that style. And the savvy students already know this.
Now with Claude Cowork, you just add instructions. I haven't looked, but I'm sure that there must be "good-style-academic-writing.md" instructions that are circulating out there so students don't even need to think of what "good writing" might be. They can just drop those instructions into Claude and "voila"!
Further, from my own experience and experiments (still only using $20-a-month models), we're moving beyond the point of "boilerplate" conclusions. This has become particularly noticeable to me over the past month. I just had a skeptic the other day ask me to prompt an LLM about a rather niche topic regarding Vietnamese music (one that required figuring out different notation systems, etc.). According to my colleague, it produced a conclusion that no one had ever made before, and that he had never thought of before. And that was with zero follow-up prompts, etc. My own experiments on various topics produce the same outcomes.
Perhaps if the assignments are about well-worn topics, then yes, the LLM might produce boilerplate conclusions, but so would 99.99999% of undergraduate students in the pre-AI age have done the same.
I think we are entering a phase where there are people who know how to use LLMs to get the best results (who ask it to come up with the best prompt first, rather than writing it on their own or just dropping the prof's assignment in, and who add all kinds of instructions to Claude Cowork, etc.), and people who don't. Those who don't know will give the impression that LLMs are not capable of doing what historians can do (and that makes us feel like we still have a place/role). However, it's the papers written by the people who DO know what to do that we need to pay attention to.
It will be interesting to see on which side of the divide the AI-native students will fall. Just because they grew up with it does not mean that they are better at it (laziness/disengagement and procrastination also play important roles). But the fact that some cannot produce good papers doesn't mean that the LLMs are incapable. They are totally capable at this point.
As such, I completely agree with Mark that "To my mind, the traditional research paper as a tool of differential assessment is now dead."
And like Mark, I have no idea what to do next. . .
In my view, the value of this current moment is in revealing the importance of historical interpretation and evidence verification. These are not things I would leave to the machines and are what I now emphasize in teaching.
Thanks for the comment. As I sure you saw in my piece, I fully agree…sometimes people assume that because I am pointing out what the models can do, I like it. I actually don’t. But I fear that if we don’t engage with the fact that the models can do these things (which is a different question they whether they should), we’ll cede the argument to economic imperatives that make AI cheaper.
Thank you. I agree. What I had meant by my comment was that even though the machines can now do data selection, interpretation, and so on, we still require humans to ask the original questions and to conduct the ethical inquiries. I don't think even think the efficiency experts want to leave research completely up to AI, though I may be too optimistic.
Unfortunately, I think you’re are being optimistic! I hate to say. Not the case with Univeristy research though, but with knowledge work outside the academy which is where our graduates most often end up.
Excellent! I would have given a master’s student a “B-“ and had her/him re-write after reading my lengthy critique. I would much rather a student re-think, re-write, etc. one paper until it earns an “A” than write 3 separate papers in a semester that aren’t good graduate work. It’s that iterative process that forces deeper thinking and mastery. It also moves the student away from dependence on AI in the latter stages of real scholarship. Yes, I even had doctoral students who hated that “polish and polish again” process, but many thanked me later for not giving up on them and for believing they were more capable than they themselves believed.
Since I don’t know the level of the student, I would give it an A- if freshman or sophomore, but offer the student a “learning opportunity” by re-writing and using my comments. Junior or senior…a B with the same offer. Master’s student B/B- and require a rewrite. Doctoral student B-, with requirement to rewrite until greatly improved. One of the problems is the prompt writer hasn’t told AI WHO THE AUDIENCE IS. Writing for an expert in the field? Writing for a generalist? Writing for a lay audience? This can easily be solved by giving AI that information from the “get-go.” So, in effect, my grading may actually be punishing the student because the given prompt isn’t complete.
Interesting. As I mentioned in my post, it was told to write at a graduate level (although the assignment was an undergrad one).
I always struggle with this because in my experience, what I’d like to do and what I can reasonably do are two very different things.
Time and workload are structural constraints. I will definitely have my PhD students and MA student writer and re-write again and again as the numbers make it manageable: I’ve only had a max of six active PhD students and 4 MAs in any given year. But when I’m teaching a 200 seat first or second year course and/or a 60 seat third year course, it’A impossible for me (at least) to let dozens of students keep re-writing assignments. Even with 20 seat fourth year and MA classes, it would be pretty demanding. In my job, teaching is only 40% of my workload too…despite those large numbers. It’s not a good system but it is the one I find myself in.
I think the larger point is that if a student can graduate with a B- average (and indeed that is normative or even a bit high) and the AI agent is also operating at a B- level too, then they both offer the market a comparable level of skill. Unfortunately, the machine will produce its result much, much faster for a fraction of the price. And here remember we’re not talking about producing actual history papers but whatever type of knowledge synthesis/research task might be required. Very few of my students (even graduate students) have become professional historians although they might love history.
And that’s the tension: we either need to find a way to allow us to teach in the way you describe (which is the ideal but one that is less often realized) or else our students are going to have to compete with machines that offer similar abilities but for pennies. That is a terrible prospect for them to face and I think we owe it to them to find a solution. I wish I had a good one!
"My computer science students are especially distraught because they aren’t allowed to use AI in class to write code, yet prospective employers increasingly treat fluency with AI coding tools as a baseline expectation."
I'm sorry... what? I've long known that CS degrees have at best a tenuous relationship with what being a working SWE is about, but the above is ludicrous. Any program barring CS students from learning to write code with modern AI coding tools is doing those students a disservice.
The reality is, all university curricula -- at least the ones that aren't solely about preparing undergraduates for graduate school -- are going to have to change. In my view, you can distill the objectives of future programs down to three things: (1) AI familiarity and competency -- understanding differences in models, harnesses, and how to get the most out of them for specific types of work (2) subject expertise -- how to know what "good" is, both in a general critical thinking way and in a specific to the discipline way and (3) how to apply that expertise to the tools to shape the result.
I don't pretend to know exactly how to design assignments or learning processes to teach those things, but I suppose that's one of the reasons we pay professors.
A thought that just occurred to me... perhaps the humanities are going to have to invent the equivalent to a mathematical proof?
In ye olden days when I was in school and learning advanced math, the big fear was the more advanced models of graphing calculators which could solve pretty much any equation, even complex ones, given to them. So, math adapted by requiring students to "show your work" to get credit. It wasn't enough to arrive at the right solution, you had to prove understanding of the techniques to derive it.
A naive way to do this would be to require the production of transcripts along with assignments, and to penalize students whose transcripts fail to demonstrate the requisite amount of curiosity and back-and-forth dialogue.
I'd be sympathetic if it was one component of a program... there is value in learning to code "by hand" to be able to read program flow and understand how algorithms work. A student who only knows how to prompt their way to working code would fail most interview loops. The same, however, is going to be true of students who can code by hand and recite algorithms from memory but can't produce output at the speed expected of modern AI software development life cycles.
No! Your role is to teach students - how to historians and to think and research, critique, and all the rest of the historian’s skills - they are AI independent.
Moreover, how to evaluate what AI does and what the use of it produces.
Students need in their area of interest to know the key resources, the past research, and thinking on a topic. They need to be able to evaluate all this too.
All the different varieties of developments in AI tech are too confusing and change daily.
Maybe tell them to keep
abreast of developments but they need to understand the implications of these debts s ask how relevant is this to my research, etc.
How does technology assist me - is it needed?
Teach them statistics!
There are so many questions, and issues here.
And having them do coding is a mistake - they need to work with people who understand what the tech does. It’s a valuable skill for the future because in an academic career they will be working with more people from different disciplines.
Just to clarify, nowhere here did I advocate teaching history students to code. I also explained why I think the research paper and teaching the historical process is so important. I tried to make the point that AI models are now able to emulate those processes and that the expectations of our students are changing. That will have consequences for things like employability and enrolments. I think those are important issues. I would love to find a solution through which we can continue to teach students history without having to go to all in-person essays and things that don’t reflect how the world is actually evolving. Any suggestions would be amazing!
No AI models do not emulate the processes at all, you need to change students expectations. In terms of employment you need to stress skills - how to think critically, evaluating, researching… working with others, people skills. Ask them to do exercises in research at a library and archives.
Set harder questions, etc. Ask them to submit handwritten essays, and perhaps all their research notes. If they are studying a topic, you as a marker should know the key resources- thinkers, arguments and papers and these should appear in their research paper. Why is that court judges seem to be pretty good at evaluating evidence? They seem I know when made up cases are cited which have been written by AI. Your job is to teach critical thinking, presentation of researching, diplomatic analysis, transcription… they have to learn to be historians.
I agree there's no more muddling through, for the reasons you lay out, and it's why all writing in my courses this semester will be done in class in blue books, and will be focused on using writing as a means to think--the process--with relatively little interest in product except as a transcript of thinking. Because I don't think we can be handling a workflow with agentic AI without ourselves having base knowledge and being able to think critically about both the process and product. You don't even know what you want agentic AI to do for you and with you if you aren't able to think, read and plan without its help.
There was only one part of your account where I stopped dead: "The writing is not very good nor engaging and the findings are pretty boilerplate but it’s an A to A+ paper."
I mean, I know I have a reputation for being a relatively easy grader within my institution, but even I wouldn't say "A to A+ for writing that is not stylistic, engaging and has only boilerplate findings". But this isn't just about grading, it's about goals. Our goals still have to be to think about what we want to know, and agentic AI is not thinking--I don't believe it can go beyond "boilerplate" because that's not a technical limit, it's an ontological limit as long as the AI we're talking about is powered at its core by LLMs. I've never given out top marks for writing that stops at a "boilerplate" conclusion that just restates a kind of averaged or synthesis point. That's not what the top of Bloom's Taxonomy looks like.
Thanks for the comment. I admit I may have been a bit loose in my phrasing but I really wish more undergraduate papers were well written and contained more than boilerplate arguments. In my experience, they simply don’t. Grade inflation has been a real problem for years…for context I can tell you that I’ve often had one of the lowest averages in faculty, or so I am told. :)
This is a very timely discussion since as you say, many of us are drafting syllabi! I am eager for ideas for assignments that allow students to do research and follow their curiosity but do not know how to evaluate the process without an outcome/final product…
Mark is right that this is THE semester where you just cannot be doing whatever it was you were doing. Asynchronous online is 100% over, and I honestly think anything that involves teaching 500-1000 people at one time in any format of any kind is over. Or should be. Honestly, anything where you don't have full and direct custody of the process by which students do their work is not viable, which really means "live, in person, and not on any devices of any kind".
Ok. But this then loops us back to "agentic AIs that are given instructions by someone who knows what they're doing produce outcomes (in history and proximate disciplines) that approximate mediocre undergraduate outputs without obvious errors." Which is something but it's not "AIs can do everything, time to just give up and live out our NPC lives now that all intellectual work has been superceded". The output you got can only exist because there's a long-standing corpus of work on shellshock and World War I (scholarly work, literary work, public commentary, etc.). I think you'd get a different output if you tasked an agentic AI to produce a completely novel synthesis on the causes of World War I that comprehensively considers the total historiography but also places causal weight where few if any historians have placed it. Whereas I might actually give an assignment like that to graduate students doing a field in the history of World War I and reasonably expect that some of them would fulfill the assignment.
The problem with moving everything to in-class blue book writing is 15% of the population is dyslexic and technology allowed us to go to university in the past 25 years (I started in 1998).
That sentence ("The writing is not very good nor engaging and the findings are pretty boilerplate but it’s an A to A+ paper.") stood out for me as well, however, I have a different take. What I find with a lot of people who, like Mark, see the enormous transformation that is taking place and try to explain that to a readership that is not receptive, is that they (perhaps subconsciously) throw in phrases like that to "soften the (existential) blow" (perhaps to themselves as well).
There is no longer any reason for the writing to "not be very good." For a long time now, all you have needed to do is to prompt the LLM to write in a certain style, or to upload some text from someone and to instruct it to write in that style. And the savvy students already know this.
Now with Claude Cowork, you just add instructions. I haven't looked, but I'm sure that there must be "good-style-academic-writing.md" instructions that are circulating out there so students don't even need to think of what "good writing" might be. They can just drop those instructions into Claude and "voila"!
Further, from my own experience and experiments (still only using $20-a-month models), we're moving beyond the point of "boilerplate" conclusions. This has become particularly noticeable to me over the past month. I just had a skeptic the other day ask me to prompt an LLM about a rather niche topic regarding Vietnamese music (one that required figuring out different notation systems, etc.). According to my colleague, it produced a conclusion that no one had ever made before, and that he had never thought of before. And that was with zero follow-up prompts, etc. My own experiments on various topics produce the same outcomes.
Perhaps if the assignments are about well-worn topics, then yes, the LLM might produce boilerplate conclusions, but so would 99.99999% of undergraduate students in the pre-AI age have done the same.
I think we are entering a phase where there are people who know how to use LLMs to get the best results (who ask it to come up with the best prompt first, rather than writing it on their own or just dropping the prof's assignment in, and who add all kinds of instructions to Claude Cowork, etc.), and people who don't. Those who don't know will give the impression that LLMs are not capable of doing what historians can do (and that makes us feel like we still have a place/role). However, it's the papers written by the people who DO know what to do that we need to pay attention to.
It will be interesting to see on which side of the divide the AI-native students will fall. Just because they grew up with it does not mean that they are better at it (laziness/disengagement and procrastination also play important roles). But the fact that some cannot produce good papers doesn't mean that the LLMs are incapable. They are totally capable at this point.
As such, I completely agree with Mark that "To my mind, the traditional research paper as a tool of differential assessment is now dead."
And like Mark, I have no idea what to do next. . .
In my view, the value of this current moment is in revealing the importance of historical interpretation and evidence verification. These are not things I would leave to the machines and are what I now emphasize in teaching.
Thanks for the comment. As I sure you saw in my piece, I fully agree…sometimes people assume that because I am pointing out what the models can do, I like it. I actually don’t. But I fear that if we don’t engage with the fact that the models can do these things (which is a different question they whether they should), we’ll cede the argument to economic imperatives that make AI cheaper.
Thank you. I agree. What I had meant by my comment was that even though the machines can now do data selection, interpretation, and so on, we still require humans to ask the original questions and to conduct the ethical inquiries. I don't think even think the efficiency experts want to leave research completely up to AI, though I may be too optimistic.
Unfortunately, I think you’re are being optimistic! I hate to say. Not the case with Univeristy research though, but with knowledge work outside the academy which is where our graduates most often end up.
Donica, I found this video interesting where a computer scientist makes your point. https://youtu.be/l-QPwk_f4eE?si=SX4yNRAKpCurWCq9
Thanks. Have bookmarked.
Excellent! I would have given a master’s student a “B-“ and had her/him re-write after reading my lengthy critique. I would much rather a student re-think, re-write, etc. one paper until it earns an “A” than write 3 separate papers in a semester that aren’t good graduate work. It’s that iterative process that forces deeper thinking and mastery. It also moves the student away from dependence on AI in the latter stages of real scholarship. Yes, I even had doctoral students who hated that “polish and polish again” process, but many thanked me later for not giving up on them and for believing they were more capable than they themselves believed.
What would you ah e give the paper Claude wrote? Just curious…
Since I don’t know the level of the student, I would give it an A- if freshman or sophomore, but offer the student a “learning opportunity” by re-writing and using my comments. Junior or senior…a B with the same offer. Master’s student B/B- and require a rewrite. Doctoral student B-, with requirement to rewrite until greatly improved. One of the problems is the prompt writer hasn’t told AI WHO THE AUDIENCE IS. Writing for an expert in the field? Writing for a generalist? Writing for a lay audience? This can easily be solved by giving AI that information from the “get-go.” So, in effect, my grading may actually be punishing the student because the given prompt isn’t complete.
Interesting. As I mentioned in my post, it was told to write at a graduate level (although the assignment was an undergrad one).
I always struggle with this because in my experience, what I’d like to do and what I can reasonably do are two very different things.
Time and workload are structural constraints. I will definitely have my PhD students and MA student writer and re-write again and again as the numbers make it manageable: I’ve only had a max of six active PhD students and 4 MAs in any given year. But when I’m teaching a 200 seat first or second year course and/or a 60 seat third year course, it’A impossible for me (at least) to let dozens of students keep re-writing assignments. Even with 20 seat fourth year and MA classes, it would be pretty demanding. In my job, teaching is only 40% of my workload too…despite those large numbers. It’s not a good system but it is the one I find myself in.
I think the larger point is that if a student can graduate with a B- average (and indeed that is normative or even a bit high) and the AI agent is also operating at a B- level too, then they both offer the market a comparable level of skill. Unfortunately, the machine will produce its result much, much faster for a fraction of the price. And here remember we’re not talking about producing actual history papers but whatever type of knowledge synthesis/research task might be required. Very few of my students (even graduate students) have become professional historians although they might love history.
And that’s the tension: we either need to find a way to allow us to teach in the way you describe (which is the ideal but one that is less often realized) or else our students are going to have to compete with machines that offer similar abilities but for pennies. That is a terrible prospect for them to face and I think we owe it to them to find a solution. I wish I had a good one!
"My computer science students are especially distraught because they aren’t allowed to use AI in class to write code, yet prospective employers increasingly treat fluency with AI coding tools as a baseline expectation."
I'm sorry... what? I've long known that CS degrees have at best a tenuous relationship with what being a working SWE is about, but the above is ludicrous. Any program barring CS students from learning to write code with modern AI coding tools is doing those students a disservice.
The reality is, all university curricula -- at least the ones that aren't solely about preparing undergraduates for graduate school -- are going to have to change. In my view, you can distill the objectives of future programs down to three things: (1) AI familiarity and competency -- understanding differences in models, harnesses, and how to get the most out of them for specific types of work (2) subject expertise -- how to know what "good" is, both in a general critical thinking way and in a specific to the discipline way and (3) how to apply that expertise to the tools to shape the result.
I don't pretend to know exactly how to design assignments or learning processes to teach those things, but I suppose that's one of the reasons we pay professors.
A thought that just occurred to me... perhaps the humanities are going to have to invent the equivalent to a mathematical proof?
In ye olden days when I was in school and learning advanced math, the big fear was the more advanced models of graphing calculators which could solve pretty much any equation, even complex ones, given to them. So, math adapted by requiring students to "show your work" to get credit. It wasn't enough to arrive at the right solution, you had to prove understanding of the techniques to derive it.
A naive way to do this would be to require the production of transcripts along with assignments, and to penalize students whose transcripts fail to demonstrate the requisite amount of curiosity and back-and-forth dialogue.
Fully agree. I’m not a computer scientist but I can’t get my head around the logic either.
I'd be sympathetic if it was one component of a program... there is value in learning to code "by hand" to be able to read program flow and understand how algorithms work. A student who only knows how to prompt their way to working code would fail most interview loops. The same, however, is going to be true of students who can code by hand and recite algorithms from memory but can't produce output at the speed expected of modern AI software development life cycles.
No! Your role is to teach students - how to historians and to think and research, critique, and all the rest of the historian’s skills - they are AI independent.
Moreover, how to evaluate what AI does and what the use of it produces.
Students need in their area of interest to know the key resources, the past research, and thinking on a topic. They need to be able to evaluate all this too.
All the different varieties of developments in AI tech are too confusing and change daily.
Maybe tell them to keep
abreast of developments but they need to understand the implications of these debts s ask how relevant is this to my research, etc.
How does technology assist me - is it needed?
Teach them statistics!
There are so many questions, and issues here.
And having them do coding is a mistake - they need to work with people who understand what the tech does. It’s a valuable skill for the future because in an academic career they will be working with more people from different disciplines.
Just to clarify, nowhere here did I advocate teaching history students to code. I also explained why I think the research paper and teaching the historical process is so important. I tried to make the point that AI models are now able to emulate those processes and that the expectations of our students are changing. That will have consequences for things like employability and enrolments. I think those are important issues. I would love to find a solution through which we can continue to teach students history without having to go to all in-person essays and things that don’t reflect how the world is actually evolving. Any suggestions would be amazing!
No AI models do not emulate the processes at all, you need to change students expectations. In terms of employment you need to stress skills - how to think critically, evaluating, researching… working with others, people skills. Ask them to do exercises in research at a library and archives.
Set harder questions, etc. Ask them to submit handwritten essays, and perhaps all their research notes. If they are studying a topic, you as a marker should know the key resources- thinkers, arguments and papers and these should appear in their research paper. Why is that court judges seem to be pretty good at evaluating evidence? They seem I know when made up cases are cited which have been written by AI. Your job is to teach critical thinking, presentation of researching, diplomatic analysis, transcription… they have to learn to be historians.
I agree there's no more muddling through, for the reasons you lay out, and it's why all writing in my courses this semester will be done in class in blue books, and will be focused on using writing as a means to think--the process--with relatively little interest in product except as a transcript of thinking. Because I don't think we can be handling a workflow with agentic AI without ourselves having base knowledge and being able to think critically about both the process and product. You don't even know what you want agentic AI to do for you and with you if you aren't able to think, read and plan without its help.
There was only one part of your account where I stopped dead: "The writing is not very good nor engaging and the findings are pretty boilerplate but it’s an A to A+ paper."
I mean, I know I have a reputation for being a relatively easy grader within my institution, but even I wouldn't say "A to A+ for writing that is not stylistic, engaging and has only boilerplate findings". But this isn't just about grading, it's about goals. Our goals still have to be to think about what we want to know, and agentic AI is not thinking--I don't believe it can go beyond "boilerplate" because that's not a technical limit, it's an ontological limit as long as the AI we're talking about is powered at its core by LLMs. I've never given out top marks for writing that stops at a "boilerplate" conclusion that just restates a kind of averaged or synthesis point. That's not what the top of Bloom's Taxonomy looks like.
Thanks for the comment. I admit I may have been a bit loose in my phrasing but I really wish more undergraduate papers were well written and contained more than boilerplate arguments. In my experience, they simply don’t. Grade inflation has been a real problem for years…for context I can tell you that I’ve often had one of the lowest averages in faculty, or so I am told. :)
This is a very timely discussion since as you say, many of us are drafting syllabi! I am eager for ideas for assignments that allow students to do research and follow their curiosity but do not know how to evaluate the process without an outcome/final product…
Mark is right that this is THE semester where you just cannot be doing whatever it was you were doing. Asynchronous online is 100% over, and I honestly think anything that involves teaching 500-1000 people at one time in any format of any kind is over. Or should be. Honestly, anything where you don't have full and direct custody of the process by which students do their work is not viable, which really means "live, in person, and not on any devices of any kind".
Ok. But this then loops us back to "agentic AIs that are given instructions by someone who knows what they're doing produce outcomes (in history and proximate disciplines) that approximate mediocre undergraduate outputs without obvious errors." Which is something but it's not "AIs can do everything, time to just give up and live out our NPC lives now that all intellectual work has been superceded". The output you got can only exist because there's a long-standing corpus of work on shellshock and World War I (scholarly work, literary work, public commentary, etc.). I think you'd get a different output if you tasked an agentic AI to produce a completely novel synthesis on the causes of World War I that comprehensively considers the total historiography but also places causal weight where few if any historians have placed it. Whereas I might actually give an assignment like that to graduate students doing a field in the history of World War I and reasonably expect that some of them would fulfill the assignment.
The problem with moving everything to in-class blue book writing is 15% of the population is dyslexic and technology allowed us to go to university in the past 25 years (I started in 1998).