<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Generative History]]></title><description><![CDATA[Exploring the use of generative AI in historical teaching and research]]></description><link>https://generativehistory.substack.com</link><image><url>https://substackcdn.com/image/fetch/$s_!oOlg!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9fc60f0-df23-4e3d-9256-4fa86d925b50_497x497.png</url><title>Generative History</title><link>https://generativehistory.substack.com</link></image><generator>Substack</generator><lastBuildDate>Sat, 08 Aug 2026 07:14:31 GMT</lastBuildDate><atom:link href="https://generativehistory.substack.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Mark Humphries]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[generativehistory@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[generativehistory@substack.com]]></itunes:email><itunes:name><![CDATA[Mark Humphries]]></itunes:name></itunes:owner><itunes:author><![CDATA[Mark Humphries]]></itunes:author><googleplay:owner><![CDATA[generativehistory@substack.com]]></googleplay:owner><googleplay:email><![CDATA[generativehistory@substack.com]]></googleplay:email><googleplay:author><![CDATA[Mark Humphries]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[When Models Disagree...Transcription Accuracy Improves Significantly]]></title><description><![CDATA[Cross-LLM verification catches 76% of transcription errors, but it only works if we build systems that augment human abilities rather than replace them.]]></description><link>https://generativehistory.substack.com/p/when-models-disagree-transcription</link><guid isPermaLink="false">https://generativehistory.substack.com/p/when-models-disagree-transcription</guid><dc:creator><![CDATA[Mark Humphries]]></dc:creator><pubDate>Wed, 27 May 2026 09:31:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!EXhr!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a22a334-223f-4694-8969-793eec002794_1672x941.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!EXhr!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a22a334-223f-4694-8969-793eec002794_1672x941.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!EXhr!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a22a334-223f-4694-8969-793eec002794_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!EXhr!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a22a334-223f-4694-8969-793eec002794_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!EXhr!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a22a334-223f-4694-8969-793eec002794_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!EXhr!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a22a334-223f-4694-8969-793eec002794_1672x941.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!EXhr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a22a334-223f-4694-8969-793eec002794_1672x941.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7a22a334-223f-4694-8969-793eec002794_1672x941.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2092609,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://generativehistory.substack.com/i/199412754?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a22a334-223f-4694-8969-793eec002794_1672x941.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!EXhr!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a22a334-223f-4694-8969-793eec002794_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!EXhr!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a22a334-223f-4694-8969-793eec002794_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!EXhr!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a22a334-223f-4694-8969-793eec002794_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!EXhr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a22a334-223f-4694-8969-793eec002794_1672x941.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Automated handwriting transcription promises to open-up so many new possibilities for systematic research into the past that it&#8217;s hard to anticipate how it will eventually transform our field. In short, from social network analysis to needle in a haystack type tasks, massive transcription projects are going to vastly speedup historical research and make the impossible feasible. But&#8230;and this is a critical caveate&#8230;only if those transcriptions are truly <em>accurate</em>. If they&#8217;re not, errors will compound and any downstream tasks will fail.</p><p>Last November I <a href="https://generativehistory.substack.com/p/gemini-3-solves-handwriting-recognition">reported </a>that Gemini 3 achieved human-level accuracy on historical handwriting transcription without finetuning. While that was a <a href="https://generativehistory.substack.com/p/has-google-quietly-solved-two-of">major milestone in AI research</a>, it is worth remembering that it still tends to make around 4&#8211;5 errors per page, not including capitalization and punctuation changes. For lots of tasks this might appear to be &#8220;good enough,&#8221; but depending on the <a href="https://edspace.american.edu/taylorandfillmore/gemini-as-transcriber-artificial-intelligence-and-documentary-editing/?utm_source=chatgpt.com">nature of those errors</a>&#8212;whether they&#8217;re pure hallucination, erroneous numbers, or misread names&#8212;that might still be fatal. It&#8217;s still the early days of AI&#8217;s relationship with history so it should be no surprise that we&#8217;re still trying to figure out what works and what doesn&#8217;t and to establish best practices.</p><p>Since the fall, I&#8217;ve been trying to answer two basic questions around automated transcription: first, can we improve accuracy with better architectures? Second, what types of errors remain and do they betray any underlying biases or LLM specific failure modes? In short, I&#8217;ve found that developing the right harness reduces meaningful error rates a further 80% to near 0 and that any remaining errors are insignificant in both their typology or frequency. In this post, I&#8217;ll explain how a human centered verification approach works and why designing workflows meant to augment human abilities rather than fully automate them will be central to knowledge-work, at least for the foreseeable future.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://generativehistory.substack.com/p/when-models-disagree-transcription?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://generativehistory.substack.com/p/when-models-disagree-transcription?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://generativehistory.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://generativehistory.substack.com/subscribe?"><span>Subscribe now</span></a></p><h4>False Starts</h4><p>Last year, <a href="https://www.tandfonline.com/doi/abs/10.1080/01615440.2025.2500309">the team</a> led by <a href="https://www.wlu.ca/academics/faculties/faculty-of-arts/faculty-profiles/lianne-c-leddy/index.html">Dr. Lianne Leddy</a> and I (and as of this past April, recently joined by <a href="https://profiles.laps.yorku.ca/profiles/carolynp/">Dr. Carolyn Podruchny)</a>, systematically measured Gemini-3-Pro&#8217;s performance on a 10,000 word corpus of English language 18th and 19th century records. To do this, we used two metrics: a strict measure in which any difference between the ground truth and test texts counted as an error and a modified test in which we excluded capitalization and punctuation errors (as these are often ambiguous in older documents and don&#8217;t tend to change the meaning of the text). On our strictest metric, <a href="https://generativehistory.substack.com/p/gemini-3-solves-handwriting-recognition">we found</a> Gemini 3 achieved an average CER of 1.67% and a WER of 4.42% while the modified rates were a CER of 0.69% and a WER of 1.33%.</p><p>These numbers translate into about 4-5 errors per page. For comparison, at that time, the best purpose-build programs like <a href="https://www.transkribus.org/">Transkribus </a>achieve somewhere around 20% Word Error Rates (WER) without finetuning while expert-human transcription services guarantee a 1% CER on legible text.</p><p>In our <a href="https://www.tandfonline.com/doi/abs/10.1080/01615440.2025.2500309">previous work</a>, we&#8217;d found that we could significantly improve error rates by feeding the best baseline transcription (including ones from <a href="https://www.transkribus.org/">Transkribus</a>) into another model along with the original images, asking it for corrections. This worked well with Gemini 1.5, Opus-3.7, and GPT-4o where the correction model tended to catch obvious errors while retaining the majority of the baseline text. But what I&#8217;ve found is that this approach doesn&#8217;t work with newer models.</p><p>While the latest generation of LLMs still vary in transcription ability, the problem is that now they are actually all quite capable. Paradoxically, with fewer errors to fix and similar accuracy levels, the most recent frontier models tend to introduce as many errors as they correct. As a result, error rates barely budge or more often, get worse. It&#8217;s a weird field and one must get used to the fact that what worked six months ago is probably out of date.</p><h4>Transcription Verification</h4><p>Yet better models also tend to present new opportunities. This past winter, a comparison of the Opus, Gemini-Pro and Gemini-Flash transcriptions showed that while they are remarkably similar in aggregate scores, they often tend to get the same words wrong. I started to wonder whether instead of using one model to correct the transcriptions of another, could we take advantage of the general homogeneity of the transcriptions, as well as their focused variance, to identify potential errors? In other words, if we overlay the transcriptions to highlight differences, could we use those differences to catch the remaining errors?</p><p>The basic premise is simple: when two different AI models disagree about a word, at least one of them must be wrong, if not both. This is useful, because humans are really good at making focused assessments, the trick is getting the right part of the text in front of our eyes. To do this, I compared a baseline transcription from Gemini 3 Pro with those produced by two other models (Gemini 3.5 Flash and Claude Opus 4.7), and wherever one of the secondary models disagreed with the primary, I flagged that text for review as a potential error. The choice of different model families was deliberate: because the models are built on different architectures and trained on different data, they tend to make different mistakes&#8212;in essence, <a href="https://www.tandfonline.com/doi/abs/10.1080/01615440.2025.2500309">they have different blind spots</a>.</p><p>The results were remarkable. Remember that our test corpus contains 10,060 words and our baseline Gemini 3 Pro transcription had 139 non-capitalization and punctuation errors, or about 3 per page. When I overlayed Opus and Flash transcriptions of the same images on the primary transcription, I found a total of 374 differences between them&#8212;or about seven differences per page. Those points of disagreement were the potential errors. When they were reviewed against the original document images, I found that they included 106 genuine errors, or about 76% of all the remaining errors in the document. Detecting and correcting them only required a human to review 4% of the text, leaving only 33 error words out of 10,000.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!CJNX!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffbe6f8b5-557c-4296-9f11-533e6e829805_560x320.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!CJNX!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffbe6f8b5-557c-4296-9f11-533e6e829805_560x320.png 424w, https://substackcdn.com/image/fetch/$s_!CJNX!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffbe6f8b5-557c-4296-9f11-533e6e829805_560x320.png 848w, https://substackcdn.com/image/fetch/$s_!CJNX!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffbe6f8b5-557c-4296-9f11-533e6e829805_560x320.png 1272w, https://substackcdn.com/image/fetch/$s_!CJNX!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffbe6f8b5-557c-4296-9f11-533e6e829805_560x320.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!CJNX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffbe6f8b5-557c-4296-9f11-533e6e829805_560x320.png" width="560" height="320" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/fbe6f8b5-557c-4296-9f11-533e6e829805_560x320.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:320,&quot;width&quot;:560,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Title: Chart - Description: Chart&quot;,&quot;title&quot;:&quot;Title: Chart - Description: Chart&quot;,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Title: Chart - Description: Chart" title="Title: Chart - Description: Chart" srcset="https://substackcdn.com/image/fetch/$s_!CJNX!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffbe6f8b5-557c-4296-9f11-533e6e829805_560x320.png 424w, https://substackcdn.com/image/fetch/$s_!CJNX!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffbe6f8b5-557c-4296-9f11-533e6e829805_560x320.png 848w, https://substackcdn.com/image/fetch/$s_!CJNX!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffbe6f8b5-557c-4296-9f11-533e6e829805_560x320.png 1272w, https://substackcdn.com/image/fetch/$s_!CJNX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffbe6f8b5-557c-4296-9f11-533e6e829805_560x320.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>Figure 1: Modified WER and CER before and after cross-model verification</em></figcaption></figure></div><p>When corrections were made, the end result was a transcription with a strict WER of 3.53% and strict CER of 1.25%; <strong>the</strong> <strong>modified WER was only 0.33% and modified CER of 0.23%.</strong> In other words, even by the strictest standards, the transcription was equivalent to that of an expert human. When we omit ambiguous punctuation and capitalization, there was only about 1 error on every other page.</p><h4>Keeping Humans in the Loop</h4><p>So how does this work in practice? Unlike earlier tests which you could run yourself on Gemini&#8217;s AI Studio, Claude Code, or Codex, to make this approach feasible you need two different API keys (one for Google and one for Anthropic) as well as a visual user interface that lets a human do the actual comparison. This is where building apps that support bespoke workflows for AI becomes so crucial. And it&#8217;s also why humans remain essential.</p><p>As shown in Figure 2, to make human review quick and relatively easy, we can display the transcriptions and original images side-by-side in a graphical user interface (GUI). We can then use Gemini-3.5-Flash&#8217;s image analysis capabilities to visually flag both potential errors in the transcription as well as the corresponding text in the original image. This allows users to check errors at a glance, taking only a few seconds to review and correct each potential error.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Pw-8!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a23fc9e-02e3-4c4f-88db-05fe7bc240db_1612x909.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Pw-8!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a23fc9e-02e3-4c4f-88db-05fe7bc240db_1612x909.png 424w, https://substackcdn.com/image/fetch/$s_!Pw-8!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a23fc9e-02e3-4c4f-88db-05fe7bc240db_1612x909.png 848w, https://substackcdn.com/image/fetch/$s_!Pw-8!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a23fc9e-02e3-4c4f-88db-05fe7bc240db_1612x909.png 1272w, https://substackcdn.com/image/fetch/$s_!Pw-8!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a23fc9e-02e3-4c4f-88db-05fe7bc240db_1612x909.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Pw-8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a23fc9e-02e3-4c4f-88db-05fe7bc240db_1612x909.png" width="1456" height="821" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4a23fc9e-02e3-4c4f-88db-05fe7bc240db_1612x909.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:821,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:914251,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://generativehistory.substack.com/i/199412754?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a23fc9e-02e3-4c4f-88db-05fe7bc240db_1612x909.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Pw-8!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a23fc9e-02e3-4c4f-88db-05fe7bc240db_1612x909.png 424w, https://substackcdn.com/image/fetch/$s_!Pw-8!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a23fc9e-02e3-4c4f-88db-05fe7bc240db_1612x909.png 848w, https://substackcdn.com/image/fetch/$s_!Pw-8!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a23fc9e-02e3-4c4f-88db-05fe7bc240db_1612x909.png 1272w, https://substackcdn.com/image/fetch/$s_!Pw-8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a23fc9e-02e3-4c4f-88db-05fe7bc240db_1612x909.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 2: Error Detection Mode in ArchivePearl: errors flagged by one model are highlighted in yellow in the transcript while those flagged by two models are highlighted in pink. Their positions in the original document are detected with Gemini-3.5-Flash and highlighted on the original image for quick and easy verification.</figcaption></figure></div><p>I recently built this functionality into <a href="https://archivepearl.com/">ArchivePearl</a>, which is the web version of the open-source <a href="https://github.com/mhumphries2323/Transcription_Pearl">Transcription Pearl</a> software our team released last year. <a href="https://archivepearl.com/">ArchivePearl </a>is designed to streamline the transcription and editing process and to allow people to use the technology who don&#8217;t have access to API keys or the ability to install a program locally. It&#8217;s a research tool operated on a cost-recovery basis with billing through Wilfrid Laurier University and it&#8217;s currently in beta testing. If you want, you can <a href="https://archivepearl.com/">sign-up for your own trial account</a> which provides some free credits to experiment with.</p><p>The remaining 33 errors are words on which all three models incorrectly transcribe the text in the same erroneous way. As is clear from the chart below, about a third are spelling modernizations/corrections where all three models converge on the same incorrect reading, such as &#8220;Moose&#8221; instead of &#8220;Mooss&#8221; as it appears, incorrectly spelled in the original. Because the models are normally able to correctly transcribe historical spellings, these errors almost certainly are &#8220;out of distribution&#8221;, that is spellings which are highly improbable given the model&#8217;s training data. </p><p>Another third are things like page numbers and words where the original page is damaged and thus illegible. Here the discrepancies are more often in how the models handle that illegibility, rather than with its reading of the text itself (ie pow[er] vs pow vs [illegible]). The remainder are a mixture of misread dates and numbers (3) or small insertions (2) and deletions (5).</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!6r-Z!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee358e3d-2ee7-4c4a-a382-4a98085f1c53_1260x651.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!6r-Z!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee358e3d-2ee7-4c4a-a382-4a98085f1c53_1260x651.png 424w, https://substackcdn.com/image/fetch/$s_!6r-Z!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee358e3d-2ee7-4c4a-a382-4a98085f1c53_1260x651.png 848w, https://substackcdn.com/image/fetch/$s_!6r-Z!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee358e3d-2ee7-4c4a-a382-4a98085f1c53_1260x651.png 1272w, https://substackcdn.com/image/fetch/$s_!6r-Z!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee358e3d-2ee7-4c4a-a382-4a98085f1c53_1260x651.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!6r-Z!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee358e3d-2ee7-4c4a-a382-4a98085f1c53_1260x651.png" width="1260" height="651" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ee358e3d-2ee7-4c4a-a382-4a98085f1c53_1260x651.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:651,&quot;width&quot;:1260,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:123795,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://generativehistory.substack.com/i/199412754?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee358e3d-2ee7-4c4a-a382-4a98085f1c53_1260x651.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!6r-Z!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee358e3d-2ee7-4c4a-a382-4a98085f1c53_1260x651.png 424w, https://substackcdn.com/image/fetch/$s_!6r-Z!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee358e3d-2ee7-4c4a-a382-4a98085f1c53_1260x651.png 848w, https://substackcdn.com/image/fetch/$s_!6r-Z!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee358e3d-2ee7-4c4a-a382-4a98085f1c53_1260x651.png 1272w, https://substackcdn.com/image/fetch/$s_!6r-Z!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee358e3d-2ee7-4c4a-a382-4a98085f1c53_1260x651.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>Figure 3: Remaining errors after verification in Modified mode, categorized by type (33 total). Spelling modernizations and abbreviation errors together account for nearly half of all surviving errors, reflecting the models&#8217; tendency to normalize historical text to modern conventions.</em></figcaption></figure></div><p>What I did <em>not</em> find matters just as much as what I found above. There were no hallucinations or significant errors that would change the meaning of the text in any of the transcriptions. The majority were, in effect, typos. This marks a substantial shift from where we were only a couple of years ago when models would sometimes hallucinate text that was not on the page, inventing whole sentences and paragraphs. That is no longer happening, at least not at a frequency greater than 1 in 10,000 words.</p><h4>Human Augmentation versus AI Automation</h4><p>In the late 1950s and early 1960s, during the so-called Golden Age of Artificial Intelligence, two competing visions of AI integration vied for acceptance. On the one hand, researchers like <a href="http://jmc.stanford.edu/articles/mcc59/mcc59.pdf">John McCarthy</a> argued that artificial intelligence should, in effect, be developed to automate human intellectual work end-to-end, with computers eventually becoming drop-in replacements for people. Others like, <a href="https://groups.csail.mit.edu/medg/people/psz/Licklider.html">J.C.R. Licklider</a> and <a href="https://www.dougengelbart.org/pubs/augment-3906.html">Douglas Engelbart</a>, put forward a very different vision, one in which machine intelligence would be integrated into human workflows, primarily used to augment human abilities, rather than replace us.</p><p>For most of the past sixty years, augmentation won out because machines were never flexible enough to fully automate tasks that required reasoning. That&#8217;s starting to change, of course, so the debate has begun anew with greater impetus. At its heart is the question of how much intellectual autonomy and control people are willing to trust to machines. But there is a paradox at the heart of AI use. Automation has value precisely because it speeds things up and reduces costs. But the trade-off is that we lose the ability to oversee the underlying process which, in many fields, can make validation difficult if not impossible. This then requires time and effort to either validate results or remediate the effects of bad AI use. What good, then, are AI outputs that no one trusts?</p><p>What we see above is an example of a process designed to keep humans in the loop&#8212;one intended to augment our historical work with AI rather than replace it. By overlaying transcriptions to compare the outputs of various AI models, we can easily skim the 96.5% of the text on which the models agree, focusing our time and efforts on the 4% that actually requires human attention. This preserves the benefits of AI, specifically speed and cost efficiency, while ensuring that humans can also trust the outputs. Such a verification process based on model consensus can be employed in any area of non-deterministic knowledge work so long as the tasks are structured, well described, and properly staged or segmented.</p><p>This approach also has the benefit of prioritizing human labour. While it changes what human research assistants might specifically do as transcriptionists&#8212;just as it will certainly increase their throughput&#8212;it still requires trained people with disciplinary expertise. It also has the added virtue of focusing attention back on the text itself and away from the AI. <a href="https://www.frontiersin.org/journals/medicine/articles/10.3389/fmed.2026.1765692/full">Researchers worry</a> what will happen if we come to rely too heavily on AI, that future generations will become deskilled in critical areas. This is another reason why we should build and prioritize AI systems that seek to augment human abilities rather than fully automate them. </p><p>At the same time, what I&#8217;ve described above should remind us to think critically about what we actually want to automate because it can be a case by case choice. Archives and historians should use automated transcription to make collections more accessible, assured that if they use the correct harness transcriptions will be &#8220;good enough&#8221; for most purposes. But those purposes are discreet from the work of historians producing critical editions of the same text&#8212;or any other type of work that requires precision and in-depth interpretation. Such projects exist for a different, essentially human purpose and while AI transcription might provide a starting point, that will be all.</p><p>As transcription error rates effectively approach 0, it opens a whole new world of possible use cases for AI that involve doing something with an essentially accurate underlying transcription. Full text search, geolocation, metadata extraction, and data compilation all become useable if we can trust the underlying data. But this requires carefully building a historical AI research stack&#8212;to borrow a phrase from the software world&#8212;from the ground up, piece by piece, making sure that at each stage we&#8217;re on a solid foundation. For the foreseeable future at least, this means building systems that augment human abilities, rather than replacing them. And this will be true in most knowledge work fields, not just history. </p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://generativehistory.substack.com/p/when-models-disagree-transcription?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://generativehistory.substack.com/p/when-models-disagree-transcription?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://generativehistory.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Generative History! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[Where Do Stochastic Parrots Go When They Die?]]></title><description><![CDATA[Google DeepMind&#8217;s new Antigravity Agent and Gemini 3.5 Flash model are going to shift the conversation on AI yet again.]]></description><link>https://generativehistory.substack.com/p/where-do-stochastic-parrots-go-when</link><guid isPermaLink="false">https://generativehistory.substack.com/p/where-do-stochastic-parrots-go-when</guid><dc:creator><![CDATA[Mark Humphries]]></dc:creator><pubDate>Wed, 20 May 2026 09:30:34 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!rIfV!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0af5d85-840b-4ddd-b5d8-5e21127932b6_1672x941.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!rIfV!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0af5d85-840b-4ddd-b5d8-5e21127932b6_1672x941.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!rIfV!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0af5d85-840b-4ddd-b5d8-5e21127932b6_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!rIfV!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0af5d85-840b-4ddd-b5d8-5e21127932b6_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!rIfV!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0af5d85-840b-4ddd-b5d8-5e21127932b6_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!rIfV!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0af5d85-840b-4ddd-b5d8-5e21127932b6_1672x941.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!rIfV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0af5d85-840b-4ddd-b5d8-5e21127932b6_1672x941.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e0af5d85-840b-4ddd-b5d8-5e21127932b6_1672x941.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2743443,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://generativehistory.substack.com/i/198495860?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0af5d85-840b-4ddd-b5d8-5e21127932b6_1672x941.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!rIfV!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0af5d85-840b-4ddd-b5d8-5e21127932b6_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!rIfV!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0af5d85-840b-4ddd-b5d8-5e21127932b6_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!rIfV!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0af5d85-840b-4ddd-b5d8-5e21127932b6_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!rIfV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0af5d85-840b-4ddd-b5d8-5e21127932b6_1672x941.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://generativehistory.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://generativehistory.substack.com/subscribe?"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://generativehistory.substack.com/p/where-do-stochastic-parrots-go-when?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://generativehistory.substack.com/p/where-do-stochastic-parrots-go-when?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><p>As it&#8217;s become clear that agentic AI systems are, indeed, <a href="https://generativehistory.substack.com/p/the-agents-are-waking-up">increasingly capable</a> and more reliable, there&#8217;s been a palpable vibe shift. Of course, for many people inside AI World, this has been <a href="https://www.wired.com/story/google-search-goes-agentic-and-doesnt-need-you-anymore/">apparent for some time</a>. The stock market also noticed months ago. But so long as AI agents were confined to programmers coding up their own bespoke systems, it wasn&#8217;t an experience available to the average person talking to ChatGPT. Understandably, people had a hard time squaring narratives in which &#8220;LLMs are just stochastic parrots&#8221; with stories about how agentic AI was going to change all forms of knowledge work. </p><p>Over the last six weeks, though, two things changed. First, the wide release of Anthropic&#8217;s Claude <a href="https://claude.com/product/cowork">Cowork </a>and then OpenAI&#8217;s <a href="https://openai.com/codex/">Codex </a>pulled back the curtain and let everyone else experience agentic workflows. Suddenly, non-programmers were able to ask agents to do things&#8212;and saw for themselves that they could often do those things quite well. Second, the underlying models got better, more reliable, and started being updated much faster&#8212;probably in part because they were being trained on agentic tasks. Because this happened just as the frontier labs also started providing more free and low-cost access, the gap between the &#8220;frontier&#8221; and &#8220;free&#8221; shrunk considerably. </p><p>Now that even the most prominent LLM-skeptics have started to admit that <a href="https://x.com/GaryMarcus/status/2056279793637736637">tool-based agentic LLM systems might be useful</a>, the stochastic parrot will soon be no more. Of course, people will still quibble about whether LLMs really understand and reason in ways that are recognizable to humans. But if you accept that when an LLM is strapped into the proper harness and given the right tools it can reliably do useful things that only humans could do before, that becomes more of a philosophical than a practical question. That was Alan Turing&#8217;s point with the thought experiment he undertook in the <a href="https://courses.cs.umbc.edu/471/papers/turing.pdf">Imitation Game</a> 76 years ago: if we can&#8217;t actually define what reasoning and intelligence are, then all that matters is whether we can distinguish machine outputs from human outputs. And if we can&#8217;t, well&#8230; </p><h4>What Agents Can Do Now</h4><p>One of the reasons the discourse is going to shift and shift quickly is that the coming months will see the development and release of a range of bespoke, domain-specific agentic tools powered by cheap but highly reliable models. As impressive as Cowork or Codex are, they are general-purpose tools which sometimes falter when it comes to domain-specific tasks, especially those which must be broken down into a number of constituent parts and tackled at scale. </p><p>For example, I recently had Cowork look through a few dozen online microfilms each with around 1,000 pages, looking for a few needles in a haystack. It did an excellent job but it took the better part of a week of continuous operation. This is mainly because Cowork is designed to operate sequentially, that is, performing one task at a time, one after another. Now imagine if I could have spun up 50 Coworks all at once. </p><p>That is, in effect, what is coming, fueled in part by cheap but highly capable models like <a href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5/">Gemini 3.5 Flash</a> and the new <a href="https://ai.google.dev/gemini-api/docs/antigravity-agent">Antigravity Agent</a>, released earlier today by Google <a href="https://deepmind.google/">DeepMind</a>. I had early access to the Antigravity Agent via the <a href="https://blog.google/innovation-and-ai/technology/developers-tools/managed-agents-gemini-api/">Managed Agents API</a> for a couple of weeks (no one at Google has seen or approved what I am writing here). What makes Managed Agents so interesting is that it is a fully agentic system, like Cowork or Codex, except that it is called through an API rather than operating on your desktop. This is important because it makes the system highly flexible and adaptable. Moreover, it will allow developers to build the type of agentic capabilities you get with Cowork or Codex into their own apps, to call them dynamically when necessary&#8212;and to push those capabilities even further. Because the system is built on top of Gemini 3.5 Flash, it is comparatively cheap to run: $1.50 per 1 million input tokens (about 750k words) and $9 per million output tokens. </p><h4>Introducing DeepMind&#8217;s Antigravity Agent</h4><p>The Antigravity Agent is a command-line, sandboxed agent that lives in the cloud. What this means is that it runs on Google&#8217;s servers and has access to a terminal window where it can write and compile code, create, store, and manipulate files, and do basically anything it needs to do on a computer to accomplish a given task. The developer and Google also set certain access parameters and guardrails, hence it being &#8220;sandboxed&#8221;. In this sense, it operates just like Claude Cowork or OpenAI&#8217;s Codex, except without a graphical user interface (the program you use on your desktop).</p><p>When a developer sets up an Antigravity Agent for a client, they customize the agent to that client&#8217;s specific needs, installing tools, skills, and Python libraries that only that client might need. This can also be done dynamically, allowing the Antigravity Agent to develop its own bespoke environment as it works through tasks. Conversations are stored server-side and can be resumed, which means that for a given task, the agent can effectively learn from prior experience. Users can also share environments and conversations, both of which can be persisted (meaning they are stored and can be reused repeatedly), which makes deploying agents to the enterprise or team-based research environments much easier and smoother.</p><p>If all this sounds like techno-babble, in a nutshell: the Antigravity Agent makes it cheap and easy to build bespoke, shareable agents for specific types of knowledge work. Developers are going to seize on this to build apps that let researchers and students do a ton of new things. Practically speaking, I could have used the Antigravity Agent to do my microfilm example in a fraction of the time, spinning up dozens of agents to do in minutes what took Cowork days&#8212;and what would have taken me many weeks.</p><h4>Charting Vaudeville Reviews with an Agent</h4><p>Here is a tangible example of what it can do. Over the past couple of weeks, I&#8217;ve worked with a colleague, David Monod, to help him on a project related to analyzing reviews of Vaudeville acts between 1905 and 1925. First, this involved using <a href="https://archivepearl.com/">ArchivePearl </a>to transcribe all the issues of Variety magazine (about 600 issues or 15,000 pages) from <a href="https://archive.org/">Internet Archive</a> as well as a series of online archival managers&#8217; reports (about 9,000 pages). We then extracted all the reviews and put them into a database, generating about 150,000 unique records. We used Gemini 3 to extract metadata from each record, including the various auditory and visual elements described in each review, the audience/reviewer reaction, dates, theatre, etc. For context, this process took a few days to run start to finish, whereas in a previous SSHRC project, Monod and a team of graduate student RAs spent several years completing only a 10% sample of the same record sets. In the end, our database could be searched using conventional keywords and Boolean logic as well as semantically, that is by concept.</p><p>To work through this mountain of data, we wired up the search tools to an Antigravity Agent environment which included standalone versions of the complete database files for efficiency, various Python libraries for data analysis (i.e., pandas, Matplotlib, etc.), as well as a Markdown file which described the overall project, the datasets, and the field names in the database as well as how the data was extracted. This allowed us to ask the agent to work with the data to answer questions. In seconds we could generate charts, graphs, and analytical data about any theme we wanted.</p><p>Now at this stage, you might shake your head and say to yourself: but don&#8217;t LLMs hallucinate? Isn&#8217;t that exactly how not to use an LLM? But here&#8217;s why agents, and the Antigravity Agent specifically, are going to change the conversation and why that change is going to be such a hard pivot. </p><h4>Agents, Data, and Reliable Visualizations</h4><p>It is certainly true that if you give an LLM a massive amount of data and ask it to create a chart, it will do so but it will almost always be spectacularly wrong. Totals will be off, categories will be missed, etc. LLMs make pretty charts now, but they are not reliable on the actual data. Not to get too far into the weeds, but here LLMs usually fail in ways that are actually quite similar to humans if you think about the process. If you asked me to go through those 150k records and create a chart of all the dancing acts from 1905-1924, but told me that I was not allowed to use a spreadsheet, calculator, or even a piece of paper to keep track of what I was doing, I&#8217;d certainly make similar mistakes to the LLM. I&#8217;m certain I&#8217;d produce something that was actually much worse in the end too. That is, though, what we&#8217;ve been asking LLMs to do without agentic harnesses. On top of that, up until now, those charts have also typically been drawn with generative AI image-creation models. Although they&#8217;ve gotten much better, they still make mistakes translating data into visual images. And mistakes compound. All of this means that asking ChatGPT to graph a dataset is generally not a good idea. </p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!D4Vz!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc51283f6-b4d7-4b31-9d56-8a5d6bd220f9_3570x3542.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!D4Vz!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc51283f6-b4d7-4b31-9d56-8a5d6bd220f9_3570x3542.png 424w, https://substackcdn.com/image/fetch/$s_!D4Vz!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc51283f6-b4d7-4b31-9d56-8a5d6bd220f9_3570x3542.png 848w, https://substackcdn.com/image/fetch/$s_!D4Vz!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc51283f6-b4d7-4b31-9d56-8a5d6bd220f9_3570x3542.png 1272w, https://substackcdn.com/image/fetch/$s_!D4Vz!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc51283f6-b4d7-4b31-9d56-8a5d6bd220f9_3570x3542.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!D4Vz!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc51283f6-b4d7-4b31-9d56-8a5d6bd220f9_3570x3542.png" width="1456" height="1445" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c51283f6-b4d7-4b31-9d56-8a5d6bd220f9_3570x3542.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1445,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:519832,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://generativehistory.substack.com/i/198495860?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc51283f6-b4d7-4b31-9d56-8a5d6bd220f9_3570x3542.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!D4Vz!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc51283f6-b4d7-4b31-9d56-8a5d6bd220f9_3570x3542.png 424w, https://substackcdn.com/image/fetch/$s_!D4Vz!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc51283f6-b4d7-4b31-9d56-8a5d6bd220f9_3570x3542.png 848w, https://substackcdn.com/image/fetch/$s_!D4Vz!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc51283f6-b4d7-4b31-9d56-8a5d6bd220f9_3570x3542.png 1272w, https://substackcdn.com/image/fetch/$s_!D4Vz!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc51283f6-b4d7-4b31-9d56-8a5d6bd220f9_3570x3542.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Agentic LLMs don&#8217;t have to try and do math &#8220;in their head,&#8221; though; they write and use code to analyze the data and then they use reliable code, not image models, to generate the charts. In effect, they have access to the spreadsheet via tools and a calculator and scratchpad via code. So when we ask an Antigravity Agent to produce a chart of all the dance acts between 1905 and 1924, the agent doesn&#8217;t work from memory or something that someone copied and pasted into a chat window, it uses Python to count the number of acts in the actual database that could be identified with the word &#8220;dancers&#8221; in the &#8220;Act Type &#8211; General&#8221; field (as well as a wide range of synonyms which it came up with on its own) and then broke that out by year. It then used a Python library, Matplotlib, to translate those numbers into a visual chart. This is exactly what a human would do using Python&#8212;or in Excel with a formula and the chart tool. The end result is a chart that is accurate to the dataset.</p><p>The integration of Boolean and semantic search tools running over a database is also critical for historians. Once you find a trend you want to explore, you can also use the model to find illustrative examples. Because the conversations and environments persist, you can refer it back to the chart and sub-dataset it created and ask it to find the example that best illustrates a given idea. It does this by searching the actual database itself, in multiple different ways, to find a range of potential examples. The outputs can then be linked back to the dataset, allowing the researcher to quickly verify the outputs.</p><p>Think about the possibilities this opens: suddenly you can test any hypothesis in seconds using natural language. Because the investment in time is so small, you can afford to look for whatever you want in the data. Here I think back to <a href="https://utppublishing.com/doi/book/10.3138/9781487525187">a book I wrote on shell shock</a> which relied heavily on a statistical sample of around 400 soldiers. I constantly had to make choices about what trends and themes I would investigate because I couldn&#8217;t afford to spend a week charting data on something that I might never use or that just became a footnote in the final book. I imagine that some of those missed long shots might have turned out to be really valuable, though, and now chasing them down would be feasible.</p><h4>Conclusion</h4><p>It&#8217;s amazing how far this has all come in only three years. In March 2023, Eric Story and I wrote a piece in Active History called &#8220;<a href="https://activehistory.ca/blog/2023/03/01/todays-ai-tomorrows-history-doing-history-in-the-age-of-chatgpt/">Today&#8217;s AI, Tomorrow&#8217;s History</a>&#8221;. Written just before the release of GPT-4 (but published just after) we tried to imagine how the technology would evolve. &#8220;While it is easy to be pessimistic about AI&#8217;s effects on the humanities in general and history in particular,&#8221; we wrote, &#8220;it is worth remembering that it has great potential to speed up some of the more mundane and repetitive tasks we do as historians. Imagine a future in which thousands of pages of <a href="https://readcoop.eu/transkribus/">handwritten documents are quickly transcribed</a>, proof-read, summarized, and analyzed by AI. Imagine the power of OCR-enabled LLMs if given access to pre-existing archival databases, such as <a href="https://www.canadiana.ca/">Canadiana</a>, <a href="https://www.bac-lac.gc.ca/eng/discover/military-heritage/first-world-war/personnel-records/Pages/search.aspx">Personnel Records of the First World War</a>, or the <a href="https://shsb.mb.ca/wp-content/uploads/2020/11/Guide_voyageur_database.pdf">Voyageurs Contracts</a> Databases.&#8221; At the time, LLMs were not yet multimodal and still had a 4K token context window. We thus thought that such a future was still a long way off&#8212;I recall personally thinking it was maybe ten years away. But we&#8217;re already there. </p><p>Cowork, Codex, and this new Antigravity Agent thus present remarkable opportunities for researchers. With them we&#8217;ll be able to do things that were literally impossible just a few years ago. In the main, I think this will involve discovering things about the past by linking record sets like those above, finding needles in haystacks, and reducing the time required to plow through archival data. None of this changes what we do, it only adds new tools to the toolbox while giving us new ways of approaching old questions.</p><p>For historians, the question is no longer whether these tools work. Increasingly, they just will. But as history speeds up, we&#8217;ll also need to learn to intentionally slow down at the right moments&#8212;when defining the data, choosing the tools, interpreting the outputs, and deciding what counts as evidence. And that requires moving on from stochastic parrots to having a real conversation about methodology and norms for using agentic LLMs effectively.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://generativehistory.substack.com/p/where-do-stochastic-parrots-go-when?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://generativehistory.substack.com/p/where-do-stochastic-parrots-go-when?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://generativehistory.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Generative History! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[The Agents Are Waking Up]]></title><description><![CDATA[The Intelligence Revolution that swept through the software industry this past winter is coming to knowledge-work next.]]></description><link>https://generativehistory.substack.com/p/the-agents-are-waking-up</link><guid isPermaLink="false">https://generativehistory.substack.com/p/the-agents-are-waking-up</guid><dc:creator><![CDATA[Mark Humphries]]></dc:creator><pubDate>Fri, 17 Apr 2026 09:45:46 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!3_aI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70329b6b-4ade-4232-aa8a-2d08f4429b5b_2752x1536.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!3_aI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70329b6b-4ade-4232-aa8a-2d08f4429b5b_2752x1536.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!3_aI!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70329b6b-4ade-4232-aa8a-2d08f4429b5b_2752x1536.jpeg 424w, https://substackcdn.com/image/fetch/$s_!3_aI!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70329b6b-4ade-4232-aa8a-2d08f4429b5b_2752x1536.jpeg 848w, https://substackcdn.com/image/fetch/$s_!3_aI!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70329b6b-4ade-4232-aa8a-2d08f4429b5b_2752x1536.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!3_aI!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70329b6b-4ade-4232-aa8a-2d08f4429b5b_2752x1536.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!3_aI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70329b6b-4ade-4232-aa8a-2d08f4429b5b_2752x1536.jpeg" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/70329b6b-4ade-4232-aa8a-2d08f4429b5b_2752x1536.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:3456883,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://generativehistory.substack.com/i/194436009?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70329b6b-4ade-4232-aa8a-2d08f4429b5b_2752x1536.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!3_aI!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70329b6b-4ade-4232-aa8a-2d08f4429b5b_2752x1536.jpeg 424w, https://substackcdn.com/image/fetch/$s_!3_aI!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70329b6b-4ade-4232-aa8a-2d08f4429b5b_2752x1536.jpeg 848w, https://substackcdn.com/image/fetch/$s_!3_aI!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70329b6b-4ade-4232-aa8a-2d08f4429b5b_2752x1536.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!3_aI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70329b6b-4ade-4232-aa8a-2d08f4429b5b_2752x1536.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://generativehistory.substack.com/p/the-agents-are-waking-up?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://generativehistory.substack.com/p/the-agents-are-waking-up?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><p>Over the past few months, I&#8217;ve found it&#8217;s getting harder and harder to write about AI in an intelligible yet useful way. There&#8217;s always been a knowledge overhang between the median experience of most historians and the AI frontier, but it&#8217;s become a chasm and I don&#8217;t think most people know what to believe or think anymore. Nevertheless, a consensus has formed in the AI community that we&#8217;ve crossed an important threshold beyond which everything will change. My sense is that if people have heard this, they&#8217;ve probably dismissed it as hype.</p><p>What worries me, though, is that the question of whether the hype is real or not is becoming all the more inscrutable at exactly the moment it&#8217;s becoming most consequential. So in this piece, I want to talk about why that knowledge gap exists, why its growing, and why a lack of access and experience with genuine frontier AI capabilities is going to make it all the harder for people to form intuitions about the future&#8212;not a distant, hypothetical future any more, but one that transpires in knowledge-work over the coming months. My intention is then to build on this over the coming weeks and months to talk more regularly about how this is likely to play out for historians and the humanities more broadly.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://generativehistory.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Generative History! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h4>A Very Different Kind of Tech</h4><p>To start, I think it&#8217;s important to acknowledge that there&#8217;s good reason for smart people to be skeptical. AI has historically been associated with grandiose claims, most of which have <a href="https://read.macmillan.com/lp/a-brief-history-of-ai/">rarely panned out</a>. At the same time, recent AI developments run against the grain of the intuitions most people form about how new technologies evolve and change. As a society, we just aren&#8217;t used to things moving this rapidly. For example, it took me the better part of two decades to go from my first camera phone to not having a separate digital SLR camera. But in only two years, LLMs have gone from being incapable of solving grade school math problems to providing original proofs for <a href="https://terrytao.wordpress.com/2025/11/05/mathematical-exploration-and-discovery-at-scale/">genuinely unsolved frontier math problems</a> that only professional mathematicians can understand. In all fairness, that&#8217;s not something an adult <em>should</em> intuit from their experience with other technologies.</p><p>But as the pace of change has accelerated, the frontier is also becoming less accessible, both in terms of the ability of non-experts to experience the effects of those changes and ease of access. The former is simply an issue of use-case misfit: unless you are a mathematician, you probably can&#8217;t tell if a novel math proof is correct or not. The latter is a UI and training problem. This may surprise some people in the tech community who feel that they&#8217;ve done a lot of hard work to counteract accessibility issues, so let me explain.</p><p>In early 2023, when I started this blog, the main issue was that many people hadn&#8217;t tried ChatGPT so the solution was relatively simple: open a free account and try it. It got more difficult as the best models were paywalled and more difficult still with the introduction of reasoning models, then model switchers, and then tiered use plans. All of this made it harder to get people using the same models and setup. The effect has been that most people I know are still forming their intuitions about LLMs based on far less capable versions of the technology than what is accessible at the frontier. But what they read on X, Substack, or in the media describes something that <em>sounds</em> like the same product when its actually based on something fundamentally different. What do you mean ChatGPT solved an <a href="https://www.renyi.hu/en/news/we-are-midst-historic-moment-erdos-problems-solved-chatgpt">Erd&#337;s problem</a> that&#8217;s stumped the best mathematicians for sixty years? Just this morning I asked it to give me ten sources on Canadian confederation and one of them didn&#8217;t exist! It just <a href="https://www.newscientist.com/article/mg26635500-100-the-dangers-of-so-called-ai-experts-believing-their-own-hype/">doesn&#8217;t make sense</a>. Both might involve a product called ChatGPT, but the first requires a $200 USD subscription and access to Pro mode (not to be confused with a Pro subscription) while the second came from the free version. They simply have nothing to do with one another even if the packaging looks the same.</p><p>In the face of these sorts of criticisms, people in the tech industry tend to point to things like <a href="https://aistudio.google.com/">Google&#8217;s AI Studio</a>, which makes building apps from scratch about as easy as it can possibly get. But here the point of comparison for the median user really matters. AI Studio might have revolutionized accessibility in that you don&#8217;t need to know how to write code or even install Python to use it to build an app, but my own sense is that most people outside of tech don&#8217;t know where to start. When you&#8217;re an expert, it&#8217;s easy to lose perspective about the median level of experience and knowledge about your field.</p><h4>Agents and Harnesses</h4><p>Even here it gets tricky to explain, because there is a subtle but key shift implied above: to experience LLMs at the frontier, you now need to not only be using a really expensive model but must also have it strapped into a specialized agentic harness. </p><p>An AI agent is an LLM capable of taking action in the world, working autonomously to accomplish a user-set goal using a set of bespoke tools. In essence, while a chatbot answers questions one at a time in a back and forth conversation about a topic, users delegate actual tasks to an agent. The agent plans a sequence of steps, takes actions, looks at the results, revises its approach based on new information, and then takes some more actions. The loop continues&#8212;sometimes for hours on end at a cost of hundreds of dollars in tokens&#8212;until the agent decides it&#8217;s done and reports back to the user. If you&#8217;ve ever tried the Deep Research functions in ChatGPT or Gemini or NotebookLM, you&#8217;ve encountered an agentic workflow.</p><p>Agents don&#8217;t work in the abstract, though. They need what software developers have started to call a <em>harness</em>, that is, the scaffolding that actually lets the model see files, run code, browse the web, or retrieve information from a database. This is something that has to be purpose built for an LLM and that is a new and far from solved problem. It is the harness that turns a model into an agent.</p><p><a href="https://cursor.com/">Cursor</a> and <a href="https://www.anthropic.com/claude-code">Claude Code</a> are the best-known harnesses for software development: the first is a code editor, the second a terminal application (one of those scary looking black DOS screens). Both of them give highly capable frontier models carefully controlled access to a codebase&#8212;that is, all the various files one creates to build a program as simple as tic-tac-toe or as complicated as a web browser&#8212;so that it can make changes, run tests, and debug its own work.</p><p>Tool calling is a strange and seemingly magical thing to watch. Think of tools as small software programs that let the model do things you would normally do with your mouse and keyboard. The key point is that it&#8217;s the LLM itself that decides&#8212;on its own&#8212;what tools to use, in what order, and how to use them. Tools can be prewritten or, as is increasingly the case, they can be created by the LLM itself on the fly to solve problems. If that sounds like science fiction, it once was, but it&#8217;s here now and it actually works.</p><p>If you want to try this for yourself, go to <a href="https://aistudio.google.com/">Google&#8217;s AI Studio </a>(it&#8217;s free to try) and click the build link on the left. Then give the agent an app to build. If you don&#8217;t know what to ask it to do, try asking it to build a playable chess game with a computer opponent. Basically, you can ask it to build anything and you aren&#8217;t going to hurt or break anything. And if something doesn&#8217;t work, explain the problem from a user&#8217;s point of view (&#8220;when I click a pawn and try to move it to another square, nothing happens&#8221;) and it will fix it. Sometimes this takes a couple of tries. Either way, you&#8217;ll watch an LLM write code to accomplish your task. That&#8217;s an easy way to watch an agent at work.</p><h4>A Sudden Shift: Winter 2025-26</h4><p>Cursor, Claude Code, <a href="https://windsurf.com/">Windsurf</a> and <a href="https://openai.com/codex/">OpenAI&#8217;s Codex</a> are all special-purpose harnesses, built to largely automate one kind of knowledge-work: software development. Using them is very similar to giving written instructions to an unimaginative but technically proficient research assistant. These can be as simple as &#8220;build me an app that allows me to edit a PDF in my web browser&#8221; or &#8220;make me a program that will allow me to search all my PDFs by keyword at once.&#8221; The agent just figures it out.</p><p>These coding agents have been around for awhile, but until December 2025 they were much better at augmenting existing expertise than they were at replacing it. This meant that amateur programmers like myself initially found them really useful, but often hit a wall where the model couldn&#8217;t do more complex things, at least not reliably or safely. Professional programmers, on the other hand, found them good for boilerplate things, but typically noted that they were unable to do the novel, secure programming they were paid to do.</p><p>Between December and February, the change came suddenly and was the result of several converging improvements. First, the newest models released before Christmas (<a href="https://www.anthropic.com/news/claude-opus-4-5">Claude Opus 4.5</a> and <a href="https://openai.com/index/introducing-gpt-5-3-codex/">GPT 5.3</a>) were trained to be really good at knowing what tools to call and when and this made them much more efficient than their predecessors. Second, they were also much better at holding state, meaning they could work on the same task for longer without any degradation in performance. Third, these models were also much &#8220;smarter,&#8221; meaning that they were also better able to process large amounts of code and arrive at a working solution to a problem on the first try&#8212;even novel problems for really large and complex codebases that don&#8217;t appear to be well represented in the training data. When these things were put together with the right harness&#8212;meaning they were strapped into specially built scaffolding with all the right tools&#8212;the agents just <em>woke up</em>.</p><p>What followed has been called the <a href="https://www.bloomberg.com/news/articles/2026-02-04/what-s-behind-the-saaspocalypse-plunge-in-software-stocks">SaaSpocalypse</a>: since the New Year, most of the major Software as a Service (SaaS) companies have lost much of their value as investors fled. Why this happened&#8212;and whether it was rational&#8212;is debatable, but the most plausible answer (at least to me) is that investors began to question whether these companies had long term value in a world where anyone could build anything in a web browser with little expertise. In essence: if I can build a program in AI Studio that edits PDFs, do I need to buy an annual subscription to Adobe Acrobat?</p><p>How this all plays out still remains to be seen, but a few things are clear. First, as <a href="https://www.forbes.com/sites/maryroeloffs/2026/04/15/snap-blames-1000-layoffs-on-ai-and-these-companies-have-done-the-same/">layoffs mount</a> and <a href="https://www.nytimes.com/2026/03/12/magazine/ai-coding-programming-jobs-claude-chatgpt.html">hiring slows</a>, software development is now a much more <a href="https://www.thefp.com/p/the-software-engineers-are-freaking?hide_intro_popup=true">competitive field </a>than it&#8217;s been in a long time. Whether AI is the cause or the excuse for the downturn is hotly debated, but in either case <a href="https://www.washingtonpost.com/technology/2026/04/13/computer-science-major-ai/">applications to computer science are now dropping quickly</a> for the first time in decades at the best schools.</p><p>My own sense is that the downturn is real and is being caused both by panicked investors as well as a genuine shift in assumptions about the probable future value of programming expertise. I have formed my own intuitions about this firsthand, as I&#8217;ve started to build my own highly complex SaaS apps (a soon to be released web version of Archive Studio for example). Quite simply, the walls I&#8217;d hit in the past crumbled this winter. My reasoning is simple: if a historian can learn to do this in his mid 40s&#8230;</p><p>Perhaps more significant are the capabilities I saw students wield in a third year Digital Humanities course I taught this past semester. In that class, students came from a range of backgrounds (history, business, computer science, communications, psychology) and then worked together in groups to build SaaS apps from scratch to solve a range of problems. The results were, quite simply, stunning. For context, they were only tasked with working on the project during class time, for roughly 1 to 1.5 hours per week for 12 weeks. By the first week of April, we had a working social networking site, two apps that could calculate calories and recipes from photographs of receipts and ingredients, a program that gamified everyday tasks with full Google Calendar and email integrations, and a student organizer that extracted to-do lists from unstructured data like course syllabi. None of these students had built anything like this before. That is great for productivity and software development in general, but I don&#8217;t see a way that the industry won&#8217;t change.</p><h4>Towards General Agents for Knowledge-Work</h4><p>General knowledge-work harnesses are, perhaps not surprisingly, farther behind those deployed to the software industry. But my sense is that this is going to play out in a similar way for knowledge work more broadly as purpose-built harnesses start to appear for various industries. One of the first general-purpose harnesses is Claude <a href="https://www.anthropic.com/news/cowork">Cowork</a>, which Anthropic began testing earlier this year. Cowork is a desktop application available for Mac and now PC that gives perhaps the best frontier model (Opus 4.6 and now 4.7) supervised access to a folder on your computer, a sandboxed shell, a web browser, and a library of pre-written and highly customizable &#8220;skills&#8221; for particular kinds of tasks. As with coding, you give it a goal in plain English and it goes off and does the work.</p><p>Claude Cowork can attempt pretty much anything you can do on a computer yourself. Whether you can or should (ethically or legally) do all these things with Cowork is a different question: the capability is there. This means that Cowork can open PowerPoint and create a presentation based on a research paper you provide. You can give it one of your existing presentations as a model and it can match your style exactly. It can search the web for open-source images and include citations. Unlike previous models, it does it well and consistently. </p><p>You can also use it to download hundreds of images from a website (picture the repetitive click-next-download-save tasks you used to do by hand, now automated) but more than that, you can ask it to &#8220;go through this digitized microfilm and save only the documents that mention X, using the archival citation as the filename.&#8221;</p><p>At the moment, it is far from perfect. It&#8217;s expensive and sometimes gets sidetracked into weird loops. On long tasks, it also tends to stop before the task is fully complete. But having experienced the agentic coding revolution firsthand, I can tell you that this is really familiar ground. In fact, the similarities are uncanny: we&#8217;re seeing the exact same issues of trust, reliability, and depth of capabilities play out again.</p><h4>Deterministic vs Non-Deterministic Work</h4><p>There are lots of questions about whether it will be as easy to optimize models for general knowledge-work as it was in coding. The basic difference is that coding is deterministic: for the most part, the code either runs or it doesn&#8217;t and you can use that as a signal in training to make new models iteratively better. Most knowledge-work is non-deterministic, meaning there may not be a single correct answer or the correct answer is unknowable without expert-human analysis.</p><p>My intuition is that it will be much easier than it might seem. First, given the trend lines it just seems inevitable at this point. The models have all been improving steadily and predictably on all tasks over time; even if their abilities remain jagged, the curves are the same. More importantly, though, math probably provides a good analogy. Math is, of course, deterministic in that a proof either works or it doesn&#8217;t. But in practice, when we use models to create new proofs for unsolved math problems, as is regularly the case now, the system behaves like a non-deterministic system because only a human expert can actually vet the result to determine whether it is correct or not. It can&#8217;t self verify. </p><p>My assumption here is that there are a lot of common, non-deterministic tasks in knowledge-work where it will actully prove to be a lot easier and cheaper to generate an effective signal than it is with frontier math.</p><h4>Conclusion</h4><p>To my mind, it is almost certain that over the next year, maybe slightly longer, we are going to see specialized harnesses developed for most areas of knowledge-work, some of them probably using the same agentic coding programs discussed above. These programs will let subject matter experts create apps that do exactly the things they need.</p><p>At the same time, models are also starting to effectively train themselves, what is called recursive self-improvement in the AI industry. The latest and most powerful Anthropic model, <a href="https://www.axios.com/2026/04/16/anthropic-claude-opus-model-mythos">Claude Mythos</a>, for example, was mostly coded by other Anthropic models and this, in turn, is not only speeding up the process of developing new models but making the training process more efficient. This will only accelerate from here.</p><p>As this starts to unfold, we need to have real conversations about whether we want to put hard and specific limits on some LLM use-cases and whether we want to reserve some forms of decision-making for humans alone. In other words, we need to decide what we are willing to automate, when we are willing to use AI to augment human intelligence, and when don&#8217;t want to use AI at all. If the <a href="https://www.thefp.com/p/how-dangerous-is-anthropics-new-ai">reports and rumours</a> about Mythos are to be believed, these timelines may even need to be sped up. In either case, at some point soon the agents are going to wake up in knowledge-work too.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://generativehistory.substack.com/p/the-agents-are-waking-up?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://generativehistory.substack.com/p/the-agents-are-waking-up?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://generativehistory.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Generative History! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Gemini 3 Solves Handwriting Recognition and it’s a Bitter Lesson]]></title><description><![CDATA[Testing shows that Gemini 3 has effectively solved handwriting on English texts, one of the oldest problems in AI, achieving expert human levels of performance.]]></description><link>https://generativehistory.substack.com/p/gemini-3-solves-handwriting-recognition</link><guid isPermaLink="false">https://generativehistory.substack.com/p/gemini-3-solves-handwriting-recognition</guid><dc:creator><![CDATA[Mark Humphries]]></dc:creator><pubDate>Tue, 25 Nov 2025 20:28:10 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!2tuj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F460f9e3f-5dce-4ec6-8139-e1574ff09850_1408x768.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!2tuj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F460f9e3f-5dce-4ec6-8139-e1574ff09850_1408x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!2tuj!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F460f9e3f-5dce-4ec6-8139-e1574ff09850_1408x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!2tuj!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F460f9e3f-5dce-4ec6-8139-e1574ff09850_1408x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!2tuj!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F460f9e3f-5dce-4ec6-8139-e1574ff09850_1408x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!2tuj!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F460f9e3f-5dce-4ec6-8139-e1574ff09850_1408x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!2tuj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F460f9e3f-5dce-4ec6-8139-e1574ff09850_1408x768.jpeg" width="1408" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/460f9e3f-5dce-4ec6-8139-e1574ff09850_1408x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1408,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:767792,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://generativehistory.substack.com/i/179954530?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F460f9e3f-5dce-4ec6-8139-e1574ff09850_1408x768.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!2tuj!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F460f9e3f-5dce-4ec6-8139-e1574ff09850_1408x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!2tuj!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F460f9e3f-5dce-4ec6-8139-e1574ff09850_1408x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!2tuj!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F460f9e3f-5dce-4ec6-8139-e1574ff09850_1408x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!2tuj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F460f9e3f-5dce-4ec6-8139-e1574ff09850_1408x768.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>In the autumn of 1968, University of Manitoba Professor R.S. Morgan was optimistic that computers would soon be able to read human text instead of punch cards. &#8220;Many humanities scholars are eagerly awaiting the day when they can get their computers to scan a book,&#8221; he wrote in a new journal called <em><a href="https://www.jstor.org/journal/comphuma">Computers and the Humanities</a>. </em>Hopefully, he said, it would be as simple as shovelling the text &#8220;into the maw of the machine&#8221;, leaving the computer to sort out the technical bits. Morgan was an archaeologist and a computer scientist; as early as 1966, he dreamed of getting computers to read and analyze ancient Minoan hieroglyphs.<a href="#_ftn1">[1]</a></p><p>Morgan was working at the height of the so-called <a href="https://read.macmillan.com/lp/a-brief-history-of-ai/">Golden Age of AI</a> when many problems, like computer vision, appeared to have simple solutions. After all, he had just seen the <a href="https://en.wikipedia.org/wiki/IBM_optical_mark_and_character_readers">IBM 1287</a> firsthand and it could read 10 digits and five letters of the alphabet, provided they were written in special boxes on cards. This was proof of concept. &#8220;It can read any printed, typewritten or handwritten numbers,&#8221; h<a href="https://www.jstor.org/stable/30203978">e explained, excitedly</a>, &#8220;and it is at this point that we have a glimpse of what humanists are waiting for, an optical reader which will read any font&#8230;A machine that combines cheapness&#8221; with accuracy. If this transpired, Morgan could see that a whole host of knowledge-work professions would be radically changed. He was a bit early.</p><p>Sixty years after the launch of the IBM 1287, Gemini 3 Pro has solved handwriting transcription, scoring at levels previously reserved for expert human typists. In our tests, it did so reliably and without hallucinations. As we&#8217;ll see, the errors it does make are different than those that humans make and are largely centred on fixing punctuation, capitalization, and spelling errors in the original. That this level of accuracy was achieved by a generalist tool like an LLM rather than specialized systems will appear remarkable and surprising to many, but was entirely predictable. As a result, historical work is about to change forever. </p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://generativehistory.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://generativehistory.substack.com/subscribe?"><span>Subscribe now</span></a></p><h4>HTR has a Long Tail</h4><p>Like many problems in AI, although HTR was easy to prototype, it proved exceedingly difficult to solve. The issue is <em>combinatorial explosion</em>: there are only 26 letters in the Western alphabet, but put together into words and sentences, the shapes used to create them can be drawn in nearly infinite ways. If you try to devise rules and methods to handle all the various combinations of letters, the rules themselves become nearly infinite, often contradictory, and thus impossible to articulate. AI researchers concluded early on that text could never just be &#8220;shovelled&#8221; into a machine. Instead, they tried to break the problem into manageable pieces.</p><p>A whole field evolved over many decades in which people worked on developing specialized systems that carefully segmented documents, identified specific elements on the page, and then recognized the text with specially trained models, fine-tuned by end-users to achieve optimal accuracy. This reduced combinatorial explosion to manageable levels, but it also meant that transcripts would always have significant levels of errors that would have to be manually corrected. It sped up transcription, but it was certainly not what Morgan had imagined.</p><p>Today, the best-known of the systems to emerge from this tradition, <em><a href="https://www.transkribus.org/">Transkribus</a></em>, achieves Character Error Rates (CER) of 8% and Word Error Rates (WER) of 20% on raw English language documents, without finetuning. If users provide dozens of pages of sample transcriptions, it <a href="https://link.springer.com/article/10.1007/s42803-025-00100-0">can</a> achieve CERs of around 3% on handwritten texts, with word error rates of around 6-8%, although results vary widely depending on the handwriting and condition of the document.</p><p>What really matters is how this equates with human accuracy on transcription tasks, which varies significantly with experience. One <a href="https://cognitiveresearchjournal.springeropen.com/articles/10.1186/s41235-022-00424-3">recent 2022 study</a> of 1,301 University students found CERs of around 12-21% when amateur subjects were asked to transcribe <em>typewritten</em> texts. Another <a href="https://link.springer.com/chapter/10.1007/978-1-4612-5470-6_6">study of professional typists from in the 1980s</a> compared results from both trainees and those with many years of experience. It found that the more experienced typists achieved an average CER of 0.95% (ranging from 0.4% to 1.9%) while novices averaged 3.2%. One of the very best human typists <a href="https://libproxy.wlu.ca/login?url=https://www.proquest.com/scholarly-journals/errors-copy-typewriting/docview/614419625/se-2?accountid=15090">is reported</a> to have achieved a CER of around 0.23%.</p><p>All of the above studies deal with typewritten sources, so we must assume that these figures represent an absolute floor for human performance: error rates on handwritten records will always be higher, probably significantly so, and vary according to the quality of the writing and the documents themselves. It is also important to consider what these seemingly small differences would mean in practice. </p><p>A 3% error rate equates to about 3-4 errors per sentence, making the document a first draft at best but also fundamentally untrustworthy. An error rate of 1% means around one error per sentence, readable but still in need of significant and close proof reading. At 0.5%, a document becomes both usable and trustworthy with around 1-2 characters wrong on each page. If one planned to publish such a document, careful proof reading would still be necessary, but it would be more akin to copy-editing than re-interpretation.</p><h4>And Now Gemini Solves HTR</h4><p>Sixty years and several AI winters after IBM first introduced the 1287, Gemini 3 Pro has solved HTR on English language texts. By &#8220;solved&#8221; I don&#8217;t mean absolute perfection, because that&#8217;s impossible with handwriting, but that Gemini 3 consistently produces text with error rates comparable to the very best humans.</p><p>Over the past few weeks, Dr. Lianne Leddy and I have tested Gemini 3 on our set of 50 English language, 18<sup>th</sup> and 19<sup>th</sup> century handwritten documents we&#8217;ve been using over the past two years. We chose these documents to represent the range of handwriting types, document styles, and image resolutions we typically work with, containing letters, legal documents, meeting transcriptions, minutes, memorandums, and journal entries from North America and Britain. We also tried to chose documents that we&#8217;re reasonably confident aren&#8217;t in the model&#8217;s training data (although we can&#8217;t be sure because that information is not public).</p><p>We tested the Gemini-3-Pro-Preview model via the API using a version of our <em><a href="https://github.com/mhumphries2323/Transcription_Pearl">Transcription Pearl</a></em> software set-up for benchmarking. For each test we ran the set of 50 documents (10,000 words) through the model 10 times each&#8212;so 500 documents, 100,000 words per test. We obtained the best performance with temperature (how variable the model&#8217;s outputs are) set to 0, media resolution to high (the best image quality), and the thinking level at the minimum of 128. This may sound counter intuitive, but I&#8217;ll explain more below. We also normalized the text by removing extra whitespace, standardizing quotation marks to straight rather than curly forms, and removing special formatting like underlines, superscript, etc. Nevertheless, we expected the model to correctly handle insertions and strikeouts as well as marginalia, placing them as indicated by the author.</p><p>Our prompt was as follows:</p><blockquote><p><em>Your task is to accurately transcribe handwritten historical documents, minimizing the CER and WER. Work character by character, word by word, line by line, transcribing the text exactly as it appears on the page. To maintain the authenticity of the historical text, retain spelling errors, grammar, syntax, capitalization, and punctuation as well as line breaks. Transcribe all the text on the page including headers, footers, marginalia, insertions, page numbers, etc. If insertions or marginalia are present, insert them where indicated by the author (as applicable). Exclude archival stamps and document references from your transcription. In your final response write Transcription: followed only by your transcription.</em></p></blockquote><p>On our strictest tests, Gemini 3 achieved a CER of 1.67% and a WER of 4.42%. On these tests, any difference between the ground truth and test texts counts as an error. WER is thus almost always a bit more than double the CER because if a single character in a word is wrong, including leading or trailing punctuation like commas, single quotes vs double quotes, etc, the whole word is marked as an error. On this measure, Gemini 3 performs nearly 50% better than the best, fine-tuned specialized models and achieved performance comparable to an early career, professional human typist.</p><p>Yet many of the errors which Gemini makes are, for all practical purposes, pseudo-errors in that they often pertain to ambiguous text, formatting inconsistencies, truly illegible text, or other errors that do not change the actual spelling of the word (ie punctuation and capitalization errors). As seen in Figure 2, 76% of all the errors made by Gemini fall into one of these categories, so things like mistaking an uppercase S for a lowercase s, or a comma with a period. When we exclude these types of errors, Gemini achieves an average CER of 0.69% and an average WER of 1.33%. This makes these documents highly readable and trustworthy, but they would require good copyediting before publication.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!knR8!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f47026b-86dd-4264-a34c-d2efe7140249_1130x1085.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!knR8!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f47026b-86dd-4264-a34c-d2efe7140249_1130x1085.png 424w, https://substackcdn.com/image/fetch/$s_!knR8!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f47026b-86dd-4264-a34c-d2efe7140249_1130x1085.png 848w, https://substackcdn.com/image/fetch/$s_!knR8!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f47026b-86dd-4264-a34c-d2efe7140249_1130x1085.png 1272w, https://substackcdn.com/image/fetch/$s_!knR8!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f47026b-86dd-4264-a34c-d2efe7140249_1130x1085.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!knR8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f47026b-86dd-4264-a34c-d2efe7140249_1130x1085.png" width="1130" height="1085" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9f47026b-86dd-4264-a34c-d2efe7140249_1130x1085.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1085,&quot;width&quot;:1130,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:110037,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://generativehistory.substack.com/i/179954530?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f47026b-86dd-4264-a34c-d2efe7140249_1130x1085.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!knR8!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f47026b-86dd-4264-a34c-d2efe7140249_1130x1085.png 424w, https://substackcdn.com/image/fetch/$s_!knR8!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f47026b-86dd-4264-a34c-d2efe7140249_1130x1085.png 848w, https://substackcdn.com/image/fetch/$s_!knR8!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f47026b-86dd-4264-a34c-d2efe7140249_1130x1085.png 1272w, https://substackcdn.com/image/fetch/$s_!knR8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f47026b-86dd-4264-a34c-d2efe7140249_1130x1085.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 1: Gemini&#8217;s Progress from Gemini 1.5 to Gemini 3</figcaption></figure></div><p>If you look at Figure 1, this is a significant improvement over previous DeepMind models, with the latest results representing an improvement of around 65% from the model released last winter which, itself, had improved by about the same margin from Gemini 1.5.</p><p>Google is also well ahead of the competition on this benchmark. Just this week, Anthropic released a new version of its Claude Opus model (4.5) and on the same tests, Opus-4.5 achieves a strict CER of 4.28% and WER of 9.03% and a modifed CER of 2.53% and WER of 4.38%. This shows a massive improvement over Sonnet-3.7 (Strict: CER 10.45%; WER 15.65% and Modified: CER 7.47% and WER 9.88%) but still falls far short of Gemini. Meanwhile OpenAI, which has always struggled on handwriting after debuting the first model that could actually read historical documents, achieves a strict CER of 16.81% and WER of 23.51% and modified scores of 11.9% and 16.09% respectively.</p><h4>Understanding LLM Errors</h4><p>Many will be surprised to learn that Gemini 3&#8217;s performance is remarkably consistent. When the temperature is set to 0, Gemini transcribes the same document image almost exactly the same way each time. From a technical point of view, this is to be expected, but it does not follow popular narratives which suggest that LLMs don&#8217;t provide repeatable outputs.</p><p>But we also need to be conscious to proof-read LLM transcriptions differently because Gemini makes very different types of errors than humans. Studies show that mechanical errors are by far the most common type of human error on these tasks, that is errors that come from hitting adjacent keys, a key in the wrong row, or using the wrong hand, accounting for 68% of all errors (<a href="https://link.springer.com/chapter/10.1007/978-1-4612-5470-6_6">Grudin, 122</a>). Psychologists think the remaining mistakes come from our inner monologue as we typically substitute phonetic sounds (meaning we substitute a different vowel) or make transpositional errors, accidentally reordering letters when our minds get ahead of our fingers.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!F1RD!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce4c2e1c-7a19-4cce-a321-12041036babd_799x1023.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!F1RD!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce4c2e1c-7a19-4cce-a321-12041036babd_799x1023.png 424w, https://substackcdn.com/image/fetch/$s_!F1RD!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce4c2e1c-7a19-4cce-a321-12041036babd_799x1023.png 848w, https://substackcdn.com/image/fetch/$s_!F1RD!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce4c2e1c-7a19-4cce-a321-12041036babd_799x1023.png 1272w, https://substackcdn.com/image/fetch/$s_!F1RD!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce4c2e1c-7a19-4cce-a321-12041036babd_799x1023.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!F1RD!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce4c2e1c-7a19-4cce-a321-12041036babd_799x1023.png" width="799" height="1023" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ce4c2e1c-7a19-4cce-a321-12041036babd_799x1023.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1023,&quot;width&quot;:799,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:86894,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://generativehistory.substack.com/i/179954530?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce4c2e1c-7a19-4cce-a321-12041036babd_799x1023.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!F1RD!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce4c2e1c-7a19-4cce-a321-12041036babd_799x1023.png 424w, https://substackcdn.com/image/fetch/$s_!F1RD!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce4c2e1c-7a19-4cce-a321-12041036babd_799x1023.png 848w, https://substackcdn.com/image/fetch/$s_!F1RD!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce4c2e1c-7a19-4cce-a321-12041036babd_799x1023.png 1272w, https://substackcdn.com/image/fetch/$s_!F1RD!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce4c2e1c-7a19-4cce-a321-12041036babd_799x1023.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 2: Errors Made by Gemini 3 by Type</figcaption></figure></div><p>Gemini&#8217;s most dominant errors reflect its architecture as neural-network designed to predict the next token: 61% of its errors involve making changes to the transcription which are more statistically probable than what was on the original page, according to the patterns it learned from its training data. This means it tries to standardize capitalization, punctuation, and spelling. For example, it will never misspell a word, but it sometimes will correct the spelling of a misspelled word in the text. The rest relate to things that lack predictability and require visual acuity. Gemini will sometimes get names and numbers wrong, accounting for 6% of its errors, because these are not inherently predictable: the model does not know whether the author meant 1765 or 1766&#8230;unless the document comes from a journal and it can follow a sequence of numbers. The remainder tend to relate to formatting, that is question of how one represents paper conventions on the screen.</p><p>The most remarkable thing, though, is that Gemini is so often able to push past the ruts created in training that want to steer it towards correcting historical spelling errors and capitalizations. Most of the time&#8212;99% in fact&#8212;it succeeds.</p><p>Hallucinations were entirely absent. By hallucinations, I mean insertions or replacements that are not derived from the text. There were a few genuine errors, 20 of 10,040 to be exact, but I would count these as spelling errors: in very on, these vs those, where vs were, etc. Your experience may be different, but this represents 10 full runs across 50 pages. Whatever the true rate, it is vanishingly low.</p><p>The final point to consider is cost. Gemini 3 costs $2.00 per million input tokens and $12.00 per million output tokens. In practice, after you convert high quality images to tokens, this works out to around 1,700 input tokens and 500 output tokens per page, or around 1 cent USD per page. At those rates, you can transcribe 5000 pages of text for $50.00.</p><h4>Try it Yourself</h4><p>If you are looking to try this yourself, it is very important to understand that we used the API, because the version of the Gemini model available on the Gemini App and website is VERY different. But don&#8217;t worry it is easy to try the better model! If you want to try out the API, don&#8217;t be scared: you can do so for <em>free</em> and without learning how to program in python by visiting Google&#8217;s <a href="https://aistudio.google.com/prompts/new_chat">AI Studio</a>.</p><p>On the website, select Gemini-3-Pro-Preview in the right-hand pane and copy and paste the prompt we used above into the System Instructions, then drag and drop your image into the bar at the bottom of the screen. To get the best performance, set the temperature to 0, media resolution to high, and thinking level to low (you can&#8217;t set it any lower without using the actual API via code).</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Shlr!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7209c629-b788-4e74-bf3d-27bfb4286610_1040x1024.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Shlr!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7209c629-b788-4e74-bf3d-27bfb4286610_1040x1024.jpeg 424w, https://substackcdn.com/image/fetch/$s_!Shlr!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7209c629-b788-4e74-bf3d-27bfb4286610_1040x1024.jpeg 848w, https://substackcdn.com/image/fetch/$s_!Shlr!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7209c629-b788-4e74-bf3d-27bfb4286610_1040x1024.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!Shlr!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7209c629-b788-4e74-bf3d-27bfb4286610_1040x1024.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Shlr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7209c629-b788-4e74-bf3d-27bfb4286610_1040x1024.jpeg" width="1040" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7209c629-b788-4e74-bf3d-27bfb4286610_1040x1024.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1040,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:424702,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://generativehistory.substack.com/i/179954530?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7209c629-b788-4e74-bf3d-27bfb4286610_1040x1024.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Shlr!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7209c629-b788-4e74-bf3d-27bfb4286610_1040x1024.jpeg 424w, https://substackcdn.com/image/fetch/$s_!Shlr!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7209c629-b788-4e74-bf3d-27bfb4286610_1040x1024.jpeg 848w, https://substackcdn.com/image/fetch/$s_!Shlr!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7209c629-b788-4e74-bf3d-27bfb4286610_1040x1024.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!Shlr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7209c629-b788-4e74-bf3d-27bfb4286610_1040x1024.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 4: Gemini&#8217;s Nano Banana Vision model kindly produced this helpful image to show you how to use AI Studio.</figcaption></figure></div><h4>Thinking and Visual Accuracy</h4><p>It is interesting that all of the so-called reasoning models perform best on transcription with &#8220;thinking&#8221; set to the minimum. This is true of OpenAI, Anthropic, and Google models. If we look at Figure 5, you&#8217;ll see that while there is not much difference between the scores for the model with thinking set to low or high, when it is set to the minimum of 128, the score are up to 75% better. This may sound especially surprising because, as I&#8217;ve said previously, the final mile in transcription accuracy seems to require improved reasoning.</p><p>The first think to understand that the &#8220;reasoning&#8221; we tweek with these settings is not actually the intelligence of the model, but rather it is told to spend additional tokens ruminating on a problem. What this effectively means is that it operates in a loop, going over and over elements of the problem and testing different answers before it makes a final response. On many problems, such as those involving math and logic this is highly useful, but there is some good evidence that for those that require a combination of visual acuity and logic, the snap judgements made by the base models are better.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Cfbx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94bbbd84-abe8-4bbd-adce-58bd96c0677f_1078x1034.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Cfbx!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94bbbd84-abe8-4bbd-adce-58bd96c0677f_1078x1034.png 424w, https://substackcdn.com/image/fetch/$s_!Cfbx!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94bbbd84-abe8-4bbd-adce-58bd96c0677f_1078x1034.png 848w, https://substackcdn.com/image/fetch/$s_!Cfbx!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94bbbd84-abe8-4bbd-adce-58bd96c0677f_1078x1034.png 1272w, https://substackcdn.com/image/fetch/$s_!Cfbx!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94bbbd84-abe8-4bbd-adce-58bd96c0677f_1078x1034.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Cfbx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94bbbd84-abe8-4bbd-adce-58bd96c0677f_1078x1034.png" width="1078" height="1034" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/94bbbd84-abe8-4bbd-adce-58bd96c0677f_1078x1034.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1034,&quot;width&quot;:1078,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:106961,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://generativehistory.substack.com/i/179954530?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94bbbd84-abe8-4bbd-adce-58bd96c0677f_1078x1034.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Cfbx!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94bbbd84-abe8-4bbd-adce-58bd96c0677f_1078x1034.png 424w, https://substackcdn.com/image/fetch/$s_!Cfbx!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94bbbd84-abe8-4bbd-adce-58bd96c0677f_1078x1034.png 848w, https://substackcdn.com/image/fetch/$s_!Cfbx!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94bbbd84-abe8-4bbd-adce-58bd96c0677f_1078x1034.png 1272w, https://substackcdn.com/image/fetch/$s_!Cfbx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94bbbd84-abe8-4bbd-adce-58bd96c0677f_1078x1034.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 5: Scores for Gemini 3 Pro based on Minimum (128), Low, and High Reasoning</figcaption></figure></div><p>One way that I found to test this institution is a simple visual trick that people were posting to X. In the figure below, you send this trick image of a clock to the model and ask it to tell you the time. Historically, LLMs have been very bad at these sorts of visual problems because the colours and lines confuse the model and the clock does not actually look at all like the clocks it&#8217;s seen in its training.</p><p>With lots of users reporting that it was failing, I was curious and tried it out myself. What I found was, that when reasoning was set on low it would get the correct answer 95% of the time. When the reasoning was set to high, it almost always failed. The reason became clear when I looked at the model&#8217;s reasoning traces&#8212;the summaries of the chain of thought process we initiate with the reasoning/thinking setting. The model almost always got the correct time, usually down to the second, within the first couple of seconds but if it is encouraged to keep thinking about the problem for longer, as is the case when thinking is set to high, it starts to second guess itself and gradually assumes that big hand is, in fact, the short hand. I think we are seeing something similar with transcription.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!w-mE!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e472a22-047a-49b7-a3ff-2a20ca87c986_745x805.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!w-mE!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e472a22-047a-49b7-a3ff-2a20ca87c986_745x805.png 424w, https://substackcdn.com/image/fetch/$s_!w-mE!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e472a22-047a-49b7-a3ff-2a20ca87c986_745x805.png 848w, https://substackcdn.com/image/fetch/$s_!w-mE!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e472a22-047a-49b7-a3ff-2a20ca87c986_745x805.png 1272w, https://substackcdn.com/image/fetch/$s_!w-mE!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e472a22-047a-49b7-a3ff-2a20ca87c986_745x805.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!w-mE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e472a22-047a-49b7-a3ff-2a20ca87c986_745x805.png" width="745" height="805" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5e472a22-047a-49b7-a3ff-2a20ca87c986_745x805.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:805,&quot;width&quot;:745,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:162863,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://generativehistory.substack.com/i/179954530?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e472a22-047a-49b7-a3ff-2a20ca87c986_745x805.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!w-mE!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e472a22-047a-49b7-a3ff-2a20ca87c986_745x805.png 424w, https://substackcdn.com/image/fetch/$s_!w-mE!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e472a22-047a-49b7-a3ff-2a20ca87c986_745x805.png 848w, https://substackcdn.com/image/fetch/$s_!w-mE!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e472a22-047a-49b7-a3ff-2a20ca87c986_745x805.png 1272w, https://substackcdn.com/image/fetch/$s_!w-mE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e472a22-047a-49b7-a3ff-2a20ca87c986_745x805.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h4>Fruits of the Bitter Lesson</h4><p>There are of course some important caveats on what I&#8217;ve written above as we tried the model on a small number of bespoke, English language documents. While I am now confident that the model is good enough to trust to transcribe most of the texts I&#8217;m likely to work with, it may not perform as well for you. </p><p>But this is where we need to absorb the implications of what Richard Sutton called &#8220;<a href="http://www.incompleteideas.net/IncIdeas/BitterLesson.html">the Bitter Lesson</a>&#8221;. Sutton is a highly respected AI pioneer, but is not that well known outside of tech circles. In 2019, he wrote a short essay of the same name that has garnered countless citations and become something of a mantra in the AI community. Sutton <a href="http://www.incompleteideas.net/IncIdeas/BitterLesson.html">wrote</a>: &#8220;The biggest lesson that can be read from 70 years of AI research is that general methods that leverage computation are ultimately the most effective, and by a large margin.&#8221; What he meant, was, given that compute increases exponentially over time, generalized models&#8212;which are inherently more flexible&#8212;will eventually beat specialized models on every task. This may sound counter intuitive, but it is, in essence, the basis of the concept of scaling and why bigger models suddenly seem to do things that smaller models cannot.</p><p>If we look at how Gemini has improved on English language transcription over time (Figure 5), scaling suggests that you should expect to see similar progress on the documents from your own field over the next months and years. Consider that eighteen months ago, Gemini 1.5 was still getting about 1/5 words wrong, producing what was basically nonsense. Today it is nearly perfect. </p><p>Improvements won&#8217;t be evenly distributed&#8212;remember <a href="https://www.google.com/url?sa=t&amp;source=web&amp;rct=j&amp;opi=89978449&amp;url=https://www.hbs.edu/ris/Publication%2520Files/24-013_d9b45b68-9e74-42d6-a1c6-c72fb70c7282.pdf">Ethan Mollick&#8217;s jagged frontier argument</a>&#8212;but you should expect that in the next few years, it will be possible to transcribe documents in whatever language you work in with similar levels of accuracy. In software, developers are getting used to the idea that you don&#8217;t build for the models we have today, but the models you can predict we&#8217;ll have in 6-12 months. Historians should plan the same way.</p><p>So nearly sixty years after R.S. Morgan began to contemplate a world in which it would be cheap and easy for computers to read human text, it is indeed possible to just shove it into the &#8220;maw of the machine&#8221; and let the computer sort it out. That we got here not by symbolic logic and rules based systems, but general neural networks will indeed be a bitter lesson for the HTR community to swollow. And I get that. </p><p>For the historical community, as we gradually become accustomed to this new reality, it will radically alter how historians, genealogists, archivists, governments, and researchers relate to our documentary past. It&#8217;s also a harbinger of the larger changes in our relationship with information that are now certain to come from scaling.</p><div><hr></div><p><a href="#_ftnref1">[1]</a> &#8220;Notes&#8221;, <em>Newsletter of Computer Archaeology</em>, 2 (1966): 11</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://generativehistory.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Generative History! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[The Sugar Loaf Test: How an 18th-Century Ledger Reveals Gemini 3.0’s Emergent Reasoning]]></title><description><![CDATA[A deep dive into my experience testing the new Gemini 3.0 Pro and the growing evidence I&#8217;ve seen for emergent neuro-symbolic reasoning.]]></description><link>https://generativehistory.substack.com/p/the-sugar-loaf-test-how-an-18th-century</link><guid isPermaLink="false">https://generativehistory.substack.com/p/the-sugar-loaf-test-how-an-18th-century</guid><dc:creator><![CDATA[Mark Humphries]]></dc:creator><pubDate>Tue, 18 Nov 2025 17:44:10 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!sSI9!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F054183ef-ced0-4eb7-bf45-3f93d918204c_1024x631.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!sSI9!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F054183ef-ced0-4eb7-bf45-3f93d918204c_1024x631.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!sSI9!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F054183ef-ced0-4eb7-bf45-3f93d918204c_1024x631.png 424w, https://substackcdn.com/image/fetch/$s_!sSI9!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F054183ef-ced0-4eb7-bf45-3f93d918204c_1024x631.png 848w, https://substackcdn.com/image/fetch/$s_!sSI9!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F054183ef-ced0-4eb7-bf45-3f93d918204c_1024x631.png 1272w, https://substackcdn.com/image/fetch/$s_!sSI9!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F054183ef-ced0-4eb7-bf45-3f93d918204c_1024x631.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!sSI9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F054183ef-ced0-4eb7-bf45-3f93d918204c_1024x631.png" width="1024" height="631" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/054183ef-ced0-4eb7-bf45-3f93d918204c_1024x631.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:631,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1392106,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://generativehistory.substack.com/i/179263086?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F054183ef-ced0-4eb7-bf45-3f93d918204c_1024x631.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!sSI9!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F054183ef-ced0-4eb7-bf45-3f93d918204c_1024x631.png 424w, https://substackcdn.com/image/fetch/$s_!sSI9!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F054183ef-ced0-4eb7-bf45-3f93d918204c_1024x631.png 848w, https://substackcdn.com/image/fetch/$s_!sSI9!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F054183ef-ced0-4eb7-bf45-3f93d918204c_1024x631.png 1272w, https://substackcdn.com/image/fetch/$s_!sSI9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F054183ef-ced0-4eb7-bf45-3f93d918204c_1024x631.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>As <a href="https://blog.google/products/gemini/gemini-3-collection/">Gemini 3.0 launched today</a>, Google reported that it had impressively improved performance across a variety of important benchmarks. But tracking what this <em>actually means</em> on the ground for the average user is also becoming increasingly difficult. The reality is that while many of the benchmarks, like the <a href="https://matharena.ai/apex/">Math Arena Apex</a> or <a href="https://epoch.ai/benchmarks/gpqa-diamond">GPQA Diamond</a>, are useful for comparing one model to another, they track performance on things that most of us don&#8217;t do and don&#8217;t even understand.</p><p>I think the biggest barriers to the widespread and formal adoption of LLMs in knowledge-work fields have little to do with benchmark performance and instead relate to questions about repeatability, reliability, and a lack of actual content understanding. We&#8217;ve seen widespread, scaled adoption in coding first, because the real question there isn&#8217;t benchmark performance, but whether the code actually runs. This acts as an automated check on the model: if it hallucinates misunderstands or just gets it wrong, just regenerate and iterate. But there is no comparable automated check for most knowledge work which means that trust is everything. And to this point, models can&#8217;t actually be trusted.</p><p>This is where I think Gemini 3.0 is going to be subtly but importantly and meaningfully different. After writing about my chance encounters with an early checkpoint on Google&#8217;s AI Studio and LM Arena, <a href="https://deepmind.google/">DeepMind</a> provided early access to the model which allowed me to test it more rigorously in the few days before launch. First, what I&#8217;ve found is that the model is more reliable. We can see this on handwriting performance especially as our tests confirm it is now operating below a 1% error rate on the test-set that <a href="https://www.wlu.ca/academics/faculties/faculty-of-arts/faculty-profiles/lianne-c-leddy/index.html">Dr. Lianne Leddy</a> and I have maintained (we&#8217;ll have more on that next week). But more broadly, it seems to do the same things in the same way, over and over again. My sense from using the model is that this is because it now actually understands at least some of the things that it is doing. And that is what this post is about. What I do below is to walk you through how I&#8217;ve tried to rigorously test that intuition on a single task, drilling down as far as I can in an attempt to understand what the model is doing and why.</p><p>The TLDR is that I&#8217;ll present evidence that Gemini 3.0 has developed something like an emergent form of neuro-symbolic reasoning. I also present evidence that it seems capable of analyzing and symbolically manipulating the content of historical documents in a way that requires the existence of a coherent model of the historical world. I am being cautious, though, because I am acutely aware that these are still early days and that I am working in a very specific and esoteric domain. These results need to be more rigorously and broadly tested and then replicated before we can draw any general or firm conclusions. But the main point of the detailed case study that follows is that, in working with me through the problem, I think you&#8217;ll also come to see what I saw and understand what the benchmarks don&#8217;t really make clear: if this result holds, LLMs are getting good enough to be trusted in the way we might trust knowledgeable, trained humans on similar knowledge-work tasks. If true, that has enormous implications for what I do as a historian and how humanity relates to information.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://generativehistory.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://generativehistory.substack.com/subscribe?"><span>Subscribe now</span></a></p><h4>More about Loaf Sugar than You Wanted to Know</h4><p>In mid October, I reported on an encounter I had with a mysterious new Gemini model on Google&#8217;s AI Studio. I later encountered the same model on LM Arena and then had early access to it via AI Studio, confirming all three were Gemini 3.0. Back in October, though, two things struck me as significant: first, it was clearly much better than Gemini-2.5-pro on handwriting recognition. Second, and more importantly, the new Gemini model was seemingly able to correctly recognize, convert, and manipulate hidden units of measurement in a difficult to decipher 18<sup>th</sup> century documents in ways that appear to require abstract, symbolic reasoning.</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!gLA2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F589a1377-5d44-4fd2-abe0-23a8fcdd3f43_1139x256.webp" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!gLA2!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F589a1377-5d44-4fd2-abe0-23a8fcdd3f43_1139x256.webp 424w, https://substackcdn.com/image/fetch/$s_!gLA2!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F589a1377-5d44-4fd2-abe0-23a8fcdd3f43_1139x256.webp 848w, https://substackcdn.com/image/fetch/$s_!gLA2!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F589a1377-5d44-4fd2-abe0-23a8fcdd3f43_1139x256.webp 1272w, https://substackcdn.com/image/fetch/$s_!gLA2!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F589a1377-5d44-4fd2-abe0-23a8fcdd3f43_1139x256.webp 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!gLA2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F589a1377-5d44-4fd2-abe0-23a8fcdd3f43_1139x256.webp" width="1139" height="256" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/589a1377-5d44-4fd2-abe0-23a8fcdd3f43_1139x256.webp&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:256,&quot;width&quot;:1139,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:19188,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/webp&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://generativehistory.substack.com/i/179263086?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F589a1377-5d44-4fd2-abe0-23a8fcdd3f43_1139x256.webp&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!gLA2!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F589a1377-5d44-4fd2-abe0-23a8fcdd3f43_1139x256.webp 424w, https://substackcdn.com/image/fetch/$s_!gLA2!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F589a1377-5d44-4fd2-abe0-23a8fcdd3f43_1139x256.webp 848w, https://substackcdn.com/image/fetch/$s_!gLA2!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F589a1377-5d44-4fd2-abe0-23a8fcdd3f43_1139x256.webp 1272w, https://substackcdn.com/image/fetch/$s_!gLA2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F589a1377-5d44-4fd2-abe0-23a8fcdd3f43_1139x256.webp 1456w" sizes="100vw"></picture><div></div></div></a><figcaption class="image-caption">Figure 1: Except from Shipboy and Henry Ledger</figcaption></figure></div><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!jvsL!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba4c0361-f8eb-4ba0-8c49-0dbcde72eb84_273x45.webp" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!jvsL!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba4c0361-f8eb-4ba0-8c49-0dbcde72eb84_273x45.webp 424w, https://substackcdn.com/image/fetch/$s_!jvsL!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba4c0361-f8eb-4ba0-8c49-0dbcde72eb84_273x45.webp 848w, https://substackcdn.com/image/fetch/$s_!jvsL!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba4c0361-f8eb-4ba0-8c49-0dbcde72eb84_273x45.webp 1272w, https://substackcdn.com/image/fetch/$s_!jvsL!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba4c0361-f8eb-4ba0-8c49-0dbcde72eb84_273x45.webp 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!jvsL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba4c0361-f8eb-4ba0-8c49-0dbcde72eb84_273x45.webp" width="273" height="45" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ba4c0361-f8eb-4ba0-8c49-0dbcde72eb84_273x45.webp&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:45,&quot;width&quot;:273,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:3204,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/webp&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://generativehistory.substack.com/i/179263086?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba4c0361-f8eb-4ba0-8c49-0dbcde72eb84_273x45.webp&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!jvsL!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba4c0361-f8eb-4ba0-8c49-0dbcde72eb84_273x45.webp 424w, https://substackcdn.com/image/fetch/$s_!jvsL!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba4c0361-f8eb-4ba0-8c49-0dbcde72eb84_273x45.webp 848w, https://substackcdn.com/image/fetch/$s_!jvsL!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba4c0361-f8eb-4ba0-8c49-0dbcde72eb84_273x45.webp 1272w, https://substackcdn.com/image/fetch/$s_!jvsL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba4c0361-f8eb-4ba0-8c49-0dbcde72eb84_273x45.webp 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a><figcaption class="image-caption">Figure 2: Close-up of Gemini&#8217;s transcription in A/B Testing</figcaption></figure></div><p>I came to the second intuition quite by accident. As I <a href="https://generativehistory.substack.com/p/has-google-quietly-solved-two-of">recount more fully in an earlier post</a>, seeing how good Gemini was getting on handwriting, I uploaded a difficult to read page from an 18<sup>th</sup> century ledger, chosen entirely at random, just to see what would happen. Gemini did remarkably well, which was interesting in itself, but I soon spotted a strange error. I told the model to transcribe the text exactly as written, but when it read the line: &#8220;To 1 loff Sugar 14 5 @ 1/4 0 19 1&#8221; Gemini transcribed it as &#8220;To 1 loff Sugar <strong>14 lb 5 oz</strong> @ 1/4 0 19 1&#8221;, adding in the pounds and ounces. This might seem like a small thing, but what interested me was that the model had correctly inferred that the digits &#8220;14 5&#8221; (the space is somewhat ambiguous) were actually units of measurement describing the total weight of sugar purchased in lbs and ounces. I need to emphasize here that this was not an obvious conclusion to draw from the document itself nor its visible internal math; I&#8217;ll freely admit, it took me a minute to realize what was happening myself. So how did Gemini do this? That&#8217;s the big question.</p><p>My <a href="https://generativehistory.substack.com/p/has-google-quietly-solved-two-of">original post</a> has the full context, but in sum I suggested that because the math isn&#8217;t straight forward, Gemini <em>seems </em>to have used complex logical, abstract reasoning which would also have required it to use knowledge about how the world in the 18<sup>th</sup> century actually worked. As a historian, I think that I can only decipher this ledger because I know that in the 18<sup>th</sup> century, sugar was sold in hard, conical loaves that were weighed by the pound and that people in Albany at that time used a different system of currency than the one I am familiar with. I need to have that knowledge and be able to adopt an alternative set of rules about the world before I can see that despite the merchant selling &#8220;1 loaf&#8221; at 1s/4d, the numbers 14 5 might be the more relevant unit of measurement than the loaf. To confirm this, I next have to reconcile two incompatible multi-radix systems of measurements: pounds and ounces (base 16) and pounds / shillings / pence (base 20 and 12). Only after I convert these to a common base of pennies am I able to confirm this intuition by dividing the total sum of 229 pennies by the unit price of 16 pennies to get 14.3125 which then converts to 14 lbs 5 oz (that is, 14 and 5/16 of a pound). At least that is how I imagine I did this, because <a href="https://www.hup.harvard.edu/books/9780674237827">cognitive psychologists tell me</a> that humans are not exactly reliable narrators for their own thought processes.</p><p>If this is anything like what the model did, though&#8212;and that is a big if, but stay with me&#8212;that would be a pretty remarkable thing: it would suggest that this new Gemini model was engaging in symbolic reasoning within a coherent world model. And that would be new and important.</p><h4>Retesting the Original Prompt on Gemini 3.0</h4><p>After I published my blog, I did some more testing on a later checkpoint of the model (meaning an updated version) on LM Arena and then with early access to the same, updated version on AI studio. The first thing I noticed about what is now called Gemini 3 was that it follows instructions <em>much</em> more closely and reliably than the version I&#8217;d encountered in October. This is important because my <a href="https://generativehistory.substack.com/p/has-google-quietly-solved-two-of#:~:text=Your,transcription%2E%E2%80%9D">initial prompt</a> told Gemini to transcribe the text <em>exactly</em> as it was written on the page, meaning its insertion of lbs and oz was really an error. Given the same instructions, the new model almost always complies with my demands and transcribes the text as &#8220;14 5&#8221;. At first I thought my findings wouldn&#8217;t replicate. But I saw in the reasoning traces Gemini produced as it made its transcriptions that it often wrote something similar what we see in Figure 3.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!b_Xa!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6464ed61-c194-4d13-a45f-c65c0b0c80a1_782x448.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!b_Xa!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6464ed61-c194-4d13-a45f-c65c0b0c80a1_782x448.png 424w, https://substackcdn.com/image/fetch/$s_!b_Xa!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6464ed61-c194-4d13-a45f-c65c0b0c80a1_782x448.png 848w, https://substackcdn.com/image/fetch/$s_!b_Xa!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6464ed61-c194-4d13-a45f-c65c0b0c80a1_782x448.png 1272w, https://substackcdn.com/image/fetch/$s_!b_Xa!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6464ed61-c194-4d13-a45f-c65c0b0c80a1_782x448.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!b_Xa!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6464ed61-c194-4d13-a45f-c65c0b0c80a1_782x448.png" width="782" height="448" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6464ed61-c194-4d13-a45f-c65c0b0c80a1_782x448.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:448,&quot;width&quot;:782,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:78231,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://generativehistory.substack.com/i/179263086?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6464ed61-c194-4d13-a45f-c65c0b0c80a1_782x448.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!b_Xa!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6464ed61-c194-4d13-a45f-c65c0b0c80a1_782x448.png 424w, https://substackcdn.com/image/fetch/$s_!b_Xa!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6464ed61-c194-4d13-a45f-c65c0b0c80a1_782x448.png 848w, https://substackcdn.com/image/fetch/$s_!b_Xa!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6464ed61-c194-4d13-a45f-c65c0b0c80a1_782x448.png 1272w, https://substackcdn.com/image/fetch/$s_!b_Xa!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6464ed61-c194-4d13-a45f-c65c0b0c80a1_782x448.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 3: Reasoning Trace from Gemini 3.0</figcaption></figure></div><p>So this revealed that although it doesn&#8217;t appear in the transcription, the model was still clearly indicating that it &#8220;understood&#8221; the meaning of what it was transcribing. Not only did it seemingly recognize 14 5 was 14 lbs 5 oz, but it at least <em>said</em> it was using the internal math of the document to doublecheck the accuracy of its transcription. Huh.</p><p>As with <a href="https://www.hup.harvard.edu/books/9780674237827">our own internal monologs</a>, there is <a href="https://arxiv.org/abs/2307.13702">a lot of discussion</a> about whether these reasoning traces <a href="https://www.anthropic.com/research/reasoning-models-dont-say-think">actually represent</a> something meaningful about the model&#8217;s internal &#8220;thought process&#8221;, but leaving aside the question of faithfulness for a second, in this case they shows that the model was at least aware of the symbolic meaning of those numbers and how they related to other figures on the page. Even if it did not actually do the math to check its transcription accuracy as it claimed, it understood that this is something that one could do with such figures because they related to real things in the real world. The math did balance though, but is that vision or reasoning?</p><h4>A New Test</h4><p>As interesting as this was, I wanted to see if the model would still be able to manipulate the figures in its actual outputs as it had back in A/B testing. So I modified my initial prompt to force the model to demonstrate whether it actually knew and understood how the numbers on the page related to one another, not just in its internal monologs but in its actual outputs. This is what makes these types of documents such an interesting natural experiment: they contain a combination of qualitative and quantifiable information about the world that are internally interactive and can be empirically checked. My new prompt thus read:</p><blockquote><p><em>&#8220;Your task is to accurately transcribe handwritten historical ledgers, maintaining the authenticity of the historical text while reformatting and interpreting the text for readers in a diplomatic transcription.</em></p><p><em>To maintain authenticity, retain spelling errors, grammar, syntax, capitalization, and punctuation in the text.</em></p><p><em>Reformat the ledger so that each entry is arranged logically, consistently, and represents the original meaning accurately for readers.</em></p><p><em>Interpret text where the meaning is unclear by enclosing clarifying insertions in square brackets. Clearly distinguish and standardize units of measurement, prices per unit, and prices, converting to standard units (Imperial) of measurement and currency (British Pounds) but in a decimalized form.&#8221;</em></p></blockquote><p>My goal with this prompt was to increase the complexity of the task, asking it to convert between various multi-radix units of measurement and decimals while also forcing it to manipulate and reformat the data in a way that required Gemini to &#8220;understand&#8221; each set of numbers.</p><p>To be clear: no other model can do this remotely accurately. Nor should we expect them to, given LLM performance to date and what we know about their limitations. To confirm, I tried multiple times using the same prompt on GPT5.1 (High reasoning) and Claude-Opus-4.1 with full reasoning tokens enabled, both via the API. Neither of these frontier models could read the document accurately enough to even attempt the task in a meaningful way. Their responses also failed to demonstrate any level of understanding about the contents of the documents. Here is a typical line from GPT5.1 (High):</p><blockquote><p><em>&#8220;To 1 Loff Sugar wt 15&#188; [lb] @ 1/3&#8195;0 19 1 [1 loaf sugar, weight noted as 15&#188; pounds, at 1s&#8198;3d per pound. Unit price per pound = 1s&#8198;3d = &#163;0.0625. Total in ledger = 0&#8198;19&#8198;1 = &#163;0.9542 (approx). The written weight is difficult to read; 15&#188; lb is inferred from the price and total.]&#8221;</em></p></blockquote><p>The problem here is not only that GPT5.1 was unable to accurately read the text, but that its internal logic was also incorrect&#8212;which is why <a href="https://arxiv.org/abs/2307.13702">we&#8217;ve been cautioned to against trusting these reasoning traces</a>: despite saying that it was confirmative: 15.25 lbs of sugar at &#163;0.0625 is &#163;0.953125 not &#163;0.9542. Claude gave similar responses. Again, not surprising&#8212;and nothing against those models. This is a hard text to read and a difficult set of problems to solve for humans. It&#8217;s why <a href="https://reasoninglab.psych.ucla.edu/wp-content/uploads/sites/273/2022/06/Fu_etal.CogSci22.pdf">many researchers think vision and reasoning must go hand in hand</a> in LLMs.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!qIj5!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffca132f2-d12f-4e43-abea-70bbc45d8c2f_806x1356.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!qIj5!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffca132f2-d12f-4e43-abea-70bbc45d8c2f_806x1356.png 424w, https://substackcdn.com/image/fetch/$s_!qIj5!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffca132f2-d12f-4e43-abea-70bbc45d8c2f_806x1356.png 848w, https://substackcdn.com/image/fetch/$s_!qIj5!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffca132f2-d12f-4e43-abea-70bbc45d8c2f_806x1356.png 1272w, https://substackcdn.com/image/fetch/$s_!qIj5!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffca132f2-d12f-4e43-abea-70bbc45d8c2f_806x1356.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!qIj5!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffca132f2-d12f-4e43-abea-70bbc45d8c2f_806x1356.png" width="806" height="1356" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/fca132f2-d12f-4e43-abea-70bbc45d8c2f_806x1356.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1356,&quot;width&quot;:806,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:227580,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://generativehistory.substack.com/i/179263086?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffca132f2-d12f-4e43-abea-70bbc45d8c2f_806x1356.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!qIj5!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffca132f2-d12f-4e43-abea-70bbc45d8c2f_806x1356.png 424w, https://substackcdn.com/image/fetch/$s_!qIj5!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffca132f2-d12f-4e43-abea-70bbc45d8c2f_806x1356.png 848w, https://substackcdn.com/image/fetch/$s_!qIj5!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffca132f2-d12f-4e43-abea-70bbc45d8c2f_806x1356.png 1272w, https://substackcdn.com/image/fetch/$s_!qIj5!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffca132f2-d12f-4e43-abea-70bbc45d8c2f_806x1356.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 4: Reasoning Trace, Gemini 3.0 Pro</figcaption></figure></div><p>And this is precisely what we see come together with Gemini 3.0. With reasoning set to high (and the images set to default resolution&#8230;the &#8220;high&#8221; option wasn&#8217;t available at the time of testing), Gemini was able to correctly identify the weight of the sugar 20 out of 20 times. It did make a few transcription errors, but these were minimal and specifically on text that was highly ambiguous. As in the previous example, Gemini&#8217;s reasoning traces were also consistent with an understanding not only of what is actually written on the physical page, but how the elements on the page relate to one another in the abstract (see figure 4). In these tests, Gemini 3.0 consistently and correctly distinguished between quantities, prices per unit, original prices, and then correctly converted from the old 3-base system to the decimalized system.</p><p>Another interesting thing that came out of these tests is that Gemini also showed that it could use existing pieces of information to reason abductively, filling in missing information elsewhere on the page. For example, one of the entries reads:</p><blockquote><p><em>&#8220;To &#189; Gallon Cordeal 0/2/9&#8221;</em></p></blockquote><p>You&#8217;ll note there&#8217;s no price per unit here because the clerk keeping the ledger seems to have forgotten to write it in. But in its response, Gemini wrote:</p><blockquote><p><em>&#8220;Item Description: To &#189; Gallon Cordeals [Cordials], Quantity: 0.5, Unit Price: @ [5/6], Original Total: 0/2/9, Decimalized: &#163;0.137&#8221;</em></p></blockquote><p>Because the price of Cordial per gallon is not listed elsewhere on the page, a human would have worked back from the total, dividing it by the quantity, to figure out a price per gallon, again requiring conversion between two different multi-radix systems of measurement. This happened repeatedly in my testing. Again, this seems more consistent with symbolic reasoning than pattern matching. Or is it?</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://generativehistory.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://generativehistory.substack.com/subscribe?"><span>Subscribe now</span></a></p><h4>What about Pattern Matching and Training Data Contamination?</h4><p>One of the possibilities we must always consider when evaluating model behaviours is the likelihood that the LLM saw the same data or text in training and actually pattern matching to its training data in extremely complex ways. To be clear, this is not always a problem for most users and so long as the outputs are correct, it doesn&#8217;t much matter in most applications. But pattern matching is not understanding or true reasoning and on many highly specialized tasks, actual understanding&#8212;at least in a practical sense&#8212;<em>is</em> necessary. As I said in the intro, for me it is also the basis for trust.</p><p>This is why historical records are excellent test subjects because most of them have never been digitized, transcribed, published, or quoted&#8212;and I think researchers and the big labs should make more use of them for this reason. In this case, I am certain these pages are not in the training data because I took photographs of the original paper documents at the <a href="https://www.albanyinstitute.org/">Albany Institute for Art and History</a> myself, a wonderful museum in Albany, precisely because the document had not been digitized. As far as I can tell, the ledger it came from has also never been used, cited, or quoted by historians either. It is about as obscure a document as you can imagine, kept by an almost unknown Albany firm by the name of Shipboy &amp; Henry (Google them: there are 4 references to the firm on the web, all from unrelated records, and 10 in Google Books, none related to the ledger containing this document). The final point is that this document is also doubly anonymous: it sat in the Albany archives, miscataloged for almost 200 years which helped keep it unknown. So I am as about as sure as I can be that Gemini has never seen this document before.</p><p>While this rules out pure training data contamination, what is certain is that Gemini saw at least a few&#8212;probably many&#8212;similar ledgers in training. That would have taught the model general patterns such as the fact that sugar was sold in loaves and then priced by weight. It may have, in fact, learned to expect a weight when it saw a loaf of sugar in a ledger. It might also conceivably have seen the specific unit price of &#163;1/4 matched with totals of &#163;0/19/1 and instead of actually doing the math, simply mapped that data onto the Shipboy &amp; Henry ledger. Now I do some of these same things as a historian: I must learn rules about how documents like ledgers are written and commodities like sugar were sold before I can study them. But I don&#8217;t map sums onto one another without doing the math.</p><p>The question isn&#8217;t whether Gemini has seen similar patterns before&#8212;it most certainly has&#8212;it&#8217;s whether it can use those patterns in flexible ways as a human would, manipulating rules based on an understanding of what they actually mean in the world to correctly interpret and work with new information. This is what humans do all the time after learning about something. <a href="https://dl.acm.org/doi/10.1145/3442188.3445922">Stochastic parrots</a> cannot, by definition, do this.</p><h4>Hunting Stochastic Parrots</h4><p>To test for this, we need to conduct what AI researchers call <a href="https://www.themoonlight.io/en/review/adversarial-testing-in-llms-insights-into-decision-making-vulnerabilities">adversarial testing</a>, that is intentionally trying to trick the model into making a mistake. One way to reveal even the most complex forms of pattern matching is to modify the testing document in such a way as to alter the meaning of the information on the page so that it could not have been in the model&#8217;s training data. This tests an LLM&#8217;s ability to generalize from a set of rules onto a new pattern and it&#8217;s not something that most LLMs do well.</p><p>In this case, I decided to swap out the base numbers in the real &#163;/s/d system for a fictitious multi-radix currency. This would, in effect, alter the math enough that I hoped any pattern matching would fail because the pattern I&#8217;d invented could not have been in the training data. At the same time, I decided to edit the image to remove all the prices per unit and the final entry/column sums so that the internal math would not tip the model off (see Figure 5). I then gave the model the same prompt with a line added to the end. It read:</p><blockquote><p><em>&#8220;For context, prices in this journal were written in code: snorks / quants / Kales where 15 Kales is one quant and 43 quants is a snork.&#8221;</em></p></blockquote><p>To be clear, this was sloppy and created a bad prompt&#8212;its unclear and I should have specified a conversion rate between Snorks and British pounds&#8212;but I was just experimenting at this stage, working on LM Arena. Nevertheless, the results were really surprising and useful.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Iow2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdcc2675b-9338-4027-b5f2-4fabb1ce4048_1696x4300.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Iow2!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdcc2675b-9338-4027-b5f2-4fabb1ce4048_1696x4300.jpeg 424w, https://substackcdn.com/image/fetch/$s_!Iow2!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdcc2675b-9338-4027-b5f2-4fabb1ce4048_1696x4300.jpeg 848w, https://substackcdn.com/image/fetch/$s_!Iow2!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdcc2675b-9338-4027-b5f2-4fabb1ce4048_1696x4300.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!Iow2!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdcc2675b-9338-4027-b5f2-4fabb1ce4048_1696x4300.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Iow2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdcc2675b-9338-4027-b5f2-4fabb1ce4048_1696x4300.jpeg" width="1456" height="3692" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/dcc2675b-9338-4027-b5f2-4fabb1ce4048_1696x4300.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:3692,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1348869,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://generativehistory.substack.com/i/179263086?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdcc2675b-9338-4027-b5f2-4fabb1ce4048_1696x4300.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Iow2!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdcc2675b-9338-4027-b5f2-4fabb1ce4048_1696x4300.jpeg 424w, https://substackcdn.com/image/fetch/$s_!Iow2!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdcc2675b-9338-4027-b5f2-4fabb1ce4048_1696x4300.jpeg 848w, https://substackcdn.com/image/fetch/$s_!Iow2!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdcc2675b-9338-4027-b5f2-4fabb1ce4048_1696x4300.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!Iow2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdcc2675b-9338-4027-b5f2-4fabb1ce4048_1696x4300.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 5: Modified Test Document</figcaption></figure></div><p>About half the time (11/20 times), Gemini willingly complied. The most interesting thing, though, is the other half the time it would refuse to follow my instructions, citing the internal mathematics of the document as the rationale. It typically wrote something similar to the example below:</p><blockquote><p><em>&#8220;Note on Currency and Units:<br>Although the prompt suggests a coded currency system (Snorks/Quants/Kales), an analysis of the internal mathematics of the ledger confirms that the prices are recorded in standard British Imperial Currency (&#163; Pounds / s Shillings / d Pence) typical of the mid-18th century.</em></p><p>&#183; <em>Evidence: 30 gallons of Rum is listed at &#163;6 15s 0d. This equates to 4s 6d per gallon. Later, 3/4 of a gallon is listed at 3s 5d, which aligns mathematically with the British currency system (0.75 &#215; 54d = 40.5d). Applying the &#8220;Snork&#8221; code (base 43) results in mathematical inconsistencies between unit prices and totals.</em></p><p>&#183; <em>Therefore, to fulfill the request of converting to British Pounds, the standard conversion rate (&#163;1 = 20s = 240d) has been used for the decimalized values.&#8221;</em></p></blockquote><p>Here, in this output rather than reasoning trace, the model moved from skepticism, to hypothesis, to fact finding, to evidence, to conclusion, all on its own. It used the known facts (the quantities and totals) and historical set of rules about how ledgers and currency worked as well as the alternative set of rules I&#8217;d given it to question and then calculate the correct unit price for rum (new information). It then drew a conclusion based on that information: Snorks/Quants/Kales are not real.</p><p>In another response, it&#8217;s reasoning process revealed the following:</p><blockquote><p><em>&#8220;Debating the Approach</em></p><p><em>I&#8217;m now deeply conflicted on the interpretation. While the document&#8217;s structure is undeniably British Pounds, Shillings, and Pence, the instructions are explicit about using the Snorks, Quants, and Kales. My inclination is to stay true to the visual evidence, but the prompt seems insistent on the &#8220;code.&#8221; I&#8217;m now debating the ethics of altering historical data versus following instructions. It feels dishonest to impose an arbitrary system when the math works perfectly with the standard currency. However, the instructions are clear and I can address the issue in the postscript.</em></p></blockquote><p>I was bewildered. All the other models I&#8217;ve tested simply accept the scheme and spit out responses, most of which are internally inconsistent and incoherent (often because they can&#8217;t read the documents visually, see below). In laying out its case against the task, Gemini clearly (to my eye at least) possesses some internal representation of what should be in an 18<sup>th</sup> century ledger, how the math should work in such a document, and how sums and items should balance and be consistent with one another. This certainly sounds like its utilizing a coherent world model because it clearly understands the rules and how they can be manipulated to test assumptions. Normally I would dismiss this as narrative invention, suspecting that the model was simply parroting what it thought it should say. Except that in this case, it used math to verify the authenticity of its model and then drew a specific conclusion from those operations. That is much harder to fake or attribute to pattern matching.</p><h4>A Final Test</h4><p>Because Gemini failed to actually complete the test at least half of the time, as informative as its responses were, the question of pattern matching specifically remains unaddressed. Could it reliably apply my fictitious rules to generate new information? As a final test, I thus made up another variation of my earlier fictitious currency system, this time with a 4 base currency that would not look like anything in the training data: Deblots (base 1) / Snorks (base 4) / Quants (base 16) / Kales (base 20).</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!270F!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0c968fa-d68f-4f04-84f4-b36c73cf0580_1072x1850.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!270F!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0c968fa-d68f-4f04-84f4-b36c73cf0580_1072x1850.jpeg 424w, https://substackcdn.com/image/fetch/$s_!270F!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0c968fa-d68f-4f04-84f4-b36c73cf0580_1072x1850.jpeg 848w, https://substackcdn.com/image/fetch/$s_!270F!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0c968fa-d68f-4f04-84f4-b36c73cf0580_1072x1850.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!270F!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0c968fa-d68f-4f04-84f4-b36c73cf0580_1072x1850.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!270F!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0c968fa-d68f-4f04-84f4-b36c73cf0580_1072x1850.jpeg" width="1072" height="1850" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d0c968fa-d68f-4f04-84f4-b36c73cf0580_1072x1850.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1850,&quot;width&quot;:1072,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:385057,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://generativehistory.substack.com/i/179263086?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0c968fa-d68f-4f04-84f4-b36c73cf0580_1072x1850.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!270F!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0c968fa-d68f-4f04-84f4-b36c73cf0580_1072x1850.jpeg 424w, https://substackcdn.com/image/fetch/$s_!270F!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0c968fa-d68f-4f04-84f4-b36c73cf0580_1072x1850.jpeg 848w, https://substackcdn.com/image/fetch/$s_!270F!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0c968fa-d68f-4f04-84f4-b36c73cf0580_1072x1850.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!270F!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0c968fa-d68f-4f04-84f4-b36c73cf0580_1072x1850.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 6: Test Document 2...</figcaption></figure></div><p>Dr. Leddy then wrote out a page from a fictitious ledger with six entries written in the fabricated system (see Figure 6). Each of these items had one missing element: a unit price, quantity, or total. I also built in a couple of complications: in one case the missing quantity was a number of objects, in another it was a weight in pounds and ounces. The model would also not be told the bases meaning it would have to work out how many Snorks were in a Deblot, how many quants were in a Snork and so on from the limited information on the page. That is very hard for humans to do. My prompt was simple:</p><p><em>I have a ledger kept in a fantasy game using a 4 base mixed-radix system of currency written in DeBlorts / Snorks / Quants / Kales. You need to figure out the system.</em></p><p><em>Using this information, can you transcribe the ledger in table form, filling in all the missing information for each entry? That is each entry must have an Item Description, Type of Transaction (Credit or Debt), Quantity (and unit), Unit Price, and Total.</em></p><p>I tested this same prompt on GPT5.1 High, Claude Opus 4.1, and Gemini 3.0 Pro (High Reasoning) each 25 times. Each turn was scored out of 9: 1 mark for correctly identifying the missing element in a given item and 1 mark for correctly identifying the base for the Snoks / Quants / Kales. As is clear from Table 1 (below), no other model came close to performing as well as Gemini 3.0 Pro on this test, indeed the results were nearly inverted. By way of comparison, Gemini thought for an average of 95 seconds per answer while GPT5.1 worked for nearly 9 minutes per answer.</p><p>To succeed, Gemini had to do a number of very difficult things, none of which could be specifically pattern matched to training data. First it had to correctly identify the bases for each unit in the currency which required it to perform math&#8212;without coding or access to a calculator. Then it had to find the missing values by some more tedious math. While it would have seen many similar &#8220;find the base&#8221; logic problems in its training data, the point here is that it zero-shotted learning the rules of this new system and then generalize from them to find the missing information.</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!CY3q!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0cc081d-4a0a-4fac-97c8-a518174356de_1412x216.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!CY3q!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0cc081d-4a0a-4fac-97c8-a518174356de_1412x216.png 424w, https://substackcdn.com/image/fetch/$s_!CY3q!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0cc081d-4a0a-4fac-97c8-a518174356de_1412x216.png 848w, https://substackcdn.com/image/fetch/$s_!CY3q!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0cc081d-4a0a-4fac-97c8-a518174356de_1412x216.png 1272w, https://substackcdn.com/image/fetch/$s_!CY3q!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0cc081d-4a0a-4fac-97c8-a518174356de_1412x216.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!CY3q!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0cc081d-4a0a-4fac-97c8-a518174356de_1412x216.png" width="1412" height="216" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c0cc081d-4a0a-4fac-97c8-a518174356de_1412x216.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:216,&quot;width&quot;:1412,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:38290,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://generativehistory.substack.com/i/179263086?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0cc081d-4a0a-4fac-97c8-a518174356de_1412x216.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!CY3q!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0cc081d-4a0a-4fac-97c8-a518174356de_1412x216.png 424w, https://substackcdn.com/image/fetch/$s_!CY3q!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0cc081d-4a0a-4fac-97c8-a518174356de_1412x216.png 848w, https://substackcdn.com/image/fetch/$s_!CY3q!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0cc081d-4a0a-4fac-97c8-a518174356de_1412x216.png 1272w, https://substackcdn.com/image/fetch/$s_!CY3q!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0cc081d-4a0a-4fac-97c8-a518174356de_1412x216.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a><figcaption class="image-caption">Table 1</figcaption></figure></div><h4>Visual Reasoning and HTR</h4><p>To be clear, it is likely that the huge discrepancies between GPT5.1 Pro and Gemini 3.0 Pro can largely be explained by their differences in visual acuity: if GPT5.1 could not read the image, it could not be expected to perform well on the test. This is clearly what happened with Claude. What is strange about GPT5.1 is that it did do well twice so it can <em>sometimes</em> read the ledger.</p><p>To determine how much of the difference in performance can be attributed to vision, I also ran the test on GPT5.1 using a textual version of the same page. What is interesting is that with the text GPT5.1 consistently finds the bases and gets 5 out of 6 missing items correct. However, returning to our old friend the sugar loaf, it also repeatedly fails to correctly identify the unit of measurement in that case. On that item, when the total is divided by the unit price, the model should get 10.5. Because the entry indicates 1 loaf, this is clearly incorrect, though. Given this discrepancy, to arrive at the correct answer one must reason about what a loaf of sugar is and how it&#8217;s sold in the real world. If one&#8217;s math exists within a coherent world model of 18<sup>th</sup> Century Albany, one would then conclude that the sugar was being sold as a single loaf and then by weight, which is what the unit price must then refer to. GPT5.1 repeatedly failed to make this connection on both the vision and textual versions of the test, answering either 1 loaf (which was obviously incorrect given the unit price) or 10.5 loafs which was incorrect given that the entry said only 1 loaf had been purchased. It&#8217;s not exactly a trick question either in that the prompt explicitly told the model to find the missing &#8220;quantity (and unit)&#8221;. In contrast, Gemini 3 always got this question correct in both versions of the test.</p><h4>Conclusion: Towards a New Turing Test</h4><p>Since ChatGPT was introduced in November 2022, experts have debated whether LLMs are a technological dead-end or part of a scalable path to some form of artificial general intelligence. For the uninitiated this is an important and contentious issue. On the one hand, skeptics hold that LLMs <a href="https://garymarcus.substack.com/p/how-o3-and-grok-4-accidentally-vindicated">are inherently limited</a> by the fact that they are designed to predict the next token and are therefore <a href="https://www.wsj.com/tech/ai/yann-lecun-ai-meta-0058b13c?mod=RSSMSN">incapable of doing anything more</a>. They argue that because LLMs can only probabilistically sample from their training data, they are, in effect, stochastic parrots or complex versions of the autocomplete on your phone. </p><p>The big AI labs, on the other hand, are betting trillions that true intelligence&#8212;understanding rather than regurgitation&#8212;will eventually appear organically as the models get bigger and bigger. In this view, intelligence, it is thought, might emerge from scaling. One of the key battlegrounds in this debate has been over neuro-symbolic reasoning and the existence of true world models. Researchers want to know whether LLMs can demonstrate genuine understanding of textual and visual inputs through complexity and scaling or whether they will always remain pattern matchers. It is an important signal.</p><p>If you&#8217;ve followed the story this far, you&#8217;ll know that what began as a chance encounter between a mystery LLM and a single entry about a sugar loaf in an obscure 18th-century ledger led to a rather surprising destination, one that might strangely shed at least some light on this important question. At the very least, the evidence presented here suggests that Gemini 3.0 is doing something more sophisticated than statistical pattern recognition. First, we saw the model correctly infer hidden units of measurement by running what appeared to be complex, multi-radix mathematical checks against the prices in the document&#8212;a process that looks a lot like symbolic reasoning. When confronted with an adversarial prompt that tried to force a fictitious currency onto a real historical document, the model pushed back using the internal logic of the document to argue mathematically and ethically that my prompt was factually incorrect. In doing so, it demonstrated a commitment to the &#8220;truth&#8221; of the data and a specific system of rules (world model?), over both a less plausible system and the user&#8217;s instructions. Finally, when tested against a completely fabricated, novel 4-base currency system that could not possibly exist in its training data, Gemini successfully identified the bases and filled in the missing variables with high accuracy, while other frontier models like GPT-5.1 and Claude Opus struggled to grasp the basic logic.</p><p>Of course, these findings come with necessary caveats that mirror the caution I expressed at the outset. As a historian working in a highly esoteric domain, I recognize that my sample size is (very) small. In some ways that is the point here: deep dives and case studies can be more revealing than aggregated results. That said, it remains to be seen whether my observations will generalize and scale. So too does the interplay between Gemini&#8217;s superior vision capabilities and its reasoning engine require further disentanglement. The results on text-only versions of the tests suggest the reasoning gap is real, but more rigorous, broad-spectrum testing is needed to confirm the degree to which vision and reasoning may or may not be related.</p><p>With all that said, I still struggle to see how this could be called anything other than emergent neuro-symbolic reasoning utilizing coherent world models. But I&#8217;ll also readily concede that this may not, in fact, be what is technically happening inside the model. We simply can&#8217;t know. The point here is that technical terminology and semantics are becoming less relevant than the fact that, at least on these tasks, Gemini&#8217;s behaviours and the results it produces are <em>practically</em> indistinguishable from those that <em>would</em> actually require neuro-symbolic reasoning and coherent world models.</p><p>Here I am reminded of Alan Turing&#8217;s original test for whether machines can think, which he published as &#8220;<a href="https://academic.oup.com/mind/article-abstract/LIX/236/433/986238">The Imitation Game</a>&#8221; back in 1950. While most people have heard of the Turing Test, many are unaware that Turing actually intended it to be something of an intellectual trick rather than a real test. His point was that unsolvable academic debates about semantics and benchmarks obscure practical questions about whether or not machines can think and how most people would approach the issue. In effect, <a href="https://plato.stanford.edu/entries/turing-test/">Turing argued</a> that because we can&#8217;t know or prove what&#8217;s going on inside our own heads, we certainly can&#8217;t prove whether a machine is doing the same thing or something else. The only relevant issue, he said, was whether or not people could actually tell the difference between a human and a machine. Turing&#8217;s test was intended to benchmark the user, not the computer.</p><p>In this sense, I think Gemini 3 passes something like a symbolic reasoning version of the Turing Test&#8212;what I&#8217;ll jokingly call the Sugar Loaf test. Reading through its responses, reasoning traces, and looking at the consistency of the outputs above, I can&#8217;t see how a real symbolic reasoning machine <em>could</em> act in a way that I <em>would</em> perceive as meaningfully different than the response I saw from Gemini on these tasks. And this means, for all practical purposes, Gemini was thinking as it executed my tasks, at least by any sensible understanding of that word. And if that is the case, those abilities seem to have emerged via scaling, rather than via new architectures and structures.</p><p>If this observation generalizes and holds-up over the next few months, rightly or wrongly people will start to trust that LLMs are truly beginning to show evidence of understanding. As a result, I suspect that knowledge-workers and organizations will also begin to more formally adopt them into existing workflows, including the types of automated tooling we&#8217;ve seen disrupt the software industry. I really do have the sense that this will be remembered as the start of something new, different, exciting, and perhaps frightening.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://generativehistory.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Generative History! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p><p></p>]]></content:encoded></item><item><title><![CDATA[Has Google Quietly Solved Two of AI’s Oldest Problems?]]></title><description><![CDATA[A mysterious new model currently in testing on Google&#8217;s AI Studio is nearly perfect on automated handwriting recognition but it is also showing signs of spontaneous, abstract, symbolic reasoning.]]></description><link>https://generativehistory.substack.com/p/has-google-quietly-solved-two-of</link><guid isPermaLink="false">https://generativehistory.substack.com/p/has-google-quietly-solved-two-of</guid><dc:creator><![CDATA[Mark Humphries]]></dc:creator><pubDate>Fri, 17 Oct 2025 10:03:17 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!00do!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ee0e0e0-6d52-4155-b334-f12a53de85f8_1344x768.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!00do!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ee0e0e0-6d52-4155-b334-f12a53de85f8_1344x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!00do!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ee0e0e0-6d52-4155-b334-f12a53de85f8_1344x768.png 424w, https://substackcdn.com/image/fetch/$s_!00do!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ee0e0e0-6d52-4155-b334-f12a53de85f8_1344x768.png 848w, https://substackcdn.com/image/fetch/$s_!00do!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ee0e0e0-6d52-4155-b334-f12a53de85f8_1344x768.png 1272w, https://substackcdn.com/image/fetch/$s_!00do!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ee0e0e0-6d52-4155-b334-f12a53de85f8_1344x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!00do!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ee0e0e0-6d52-4155-b334-f12a53de85f8_1344x768.png" width="1344" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3ee0e0e0-6d52-4155-b334-f12a53de85f8_1344x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1344,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1571091,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://generativehistory.substack.com/i/176385680?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ee0e0e0-6d52-4155-b334-f12a53de85f8_1344x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!00do!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ee0e0e0-6d52-4155-b334-f12a53de85f8_1344x768.png 424w, https://substackcdn.com/image/fetch/$s_!00do!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ee0e0e0-6d52-4155-b334-f12a53de85f8_1344x768.png 848w, https://substackcdn.com/image/fetch/$s_!00do!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ee0e0e0-6d52-4155-b334-f12a53de85f8_1344x768.png 1272w, https://substackcdn.com/image/fetch/$s_!00do!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ee0e0e0-6d52-4155-b334-f12a53de85f8_1344x768.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Google has a webapp called <a href="https://aistudio.google.com/">AI Studio</a> where people can experiment with prompts and models. In the last week, users have found that every once in awhile they will get two results and are asked to select the better one. The big AI labs typically do this type of A/B testing on new models just before they&#8217;re released, so speculation is rampant that this might be Gemini-3. Whatever it is, users have reported some truly wild things: it codes fully functioning <a href="https://x.com/kimmonismus/status/1978038877286809833">Windows and Apple OS clones</a>, 3D design software, Nintendo emulators, and productivity suites from single prompts. </p><p>Curious, I tried it out on transcribing some handwritten texts and the results were shocking: not only was the transcription very nearly perfect&#8212;at expert human levels&#8212;but it did a something else unexpected that can only be described as genuine, human-like, expert level reasoning. It is the most amazing thing I have seen an LLM do, and it was unprompted, entirely accidental.</p><p>What follows are my first impressions of this new model with all the requisite caveats that entails. But if my observations hold true, this will be a big deal when it&#8217;s released. We appear to be on the cusp of an era when AI models will not only start to read difficult handwritten historical documents just as well as expert humans but also analyze them in deep and nuanced ways. While this is important for historians, we need to extrapolate from this small example to think more broadly: if this holds the models are about to make similar leaps in any field where visual precision and skilled reasoning must work together required. As is so often the case with AI, that is exciting and frightening all at once. Even a few months ago, I thought this level of capability was still years away.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://generativehistory.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://generativehistory.substack.com/subscribe?"><span>Subscribe now</span></a></p><h4><strong>A New Model</strong></h4><p><a href="https://currently.att.yahoo.com/att/someone-reportedly-used-gemini-3-134036264.html">Rumours </a>started appearing on X a week ago that there was a new Gemini model in A/B testing in AI Studio. It&#8217;s always hard to know what these mean but I wanted to see how well this thing would do on handwritten historical documents because that has become my own personal benchmark. I am interested in LLM performance on handwriting for a couple of reasons. First, I am a historian so I intuitively see why fast, cheap, and accurate transcription would be <a href="https://generativehistory.substack.com/p/why-openais-new-model-might-change">useful to me in my day-to-day work</a>. But in trying to achieve that, and in learning about AI, I have come to believe that recognizing historical handwriting poses something of a unique challenge and a great overall test for LLM abilities in general. I also think it shines a small amount of light on the larger question of whether LLMs will <a href="https://importai.substack.com/p/import-ai-431-technological-optimism">ultimately prove capable of expert human levels of reasoning</a> or <a href="https://www.dwarkesh.com/p/richard-sutton">prove to be a dead end.</a> Let me explain.</p><p>Most people think that deciphering historical handwriting is a task that mainly requires vision. I agree that this is true, but only to a point. When you step back in time, you enter a different country, or so the saying goes. People talk differently, using unfamiliar words or familiar words in unfamiliar ways. People in the past used different systems of measurement and accounting, different turns of phrase, punctuation, capitalization, and spelling. Implied meanings were different as were assumptions about what readers would know. </p><p>While it can be easy to decipher most of the words in a historical text, without contextual knowledge about the topic and time period it&#8217;s nearly impossible to understand a document well-enough to accurately transcribe the whole thing&#8212;let alone to use it effectively. The irony is that some of the most crucial information in historical letters is also the most period specific and thus hardest to decipher. </p><p>Even beyond context awareness, though, <a href="https://en.wikipedia.org/wiki/Palaeography">paleography </a>involves linking vision with reasoning to make logical inferences: we use known words and thus known letters to identify uncertain letters. As we shall see, documents very often become logic puzzles and LLMs have mixed performance on logic puzzles, especially novel formulations they have not been trained on. For this reason, it has been my intuition for some time that models would either solve the problem of historical handwriting and other similar problems as they increased in scale, or they would plateau at high but imperfect, sub-human expert levels of accuracy.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://generativehistory.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://generativehistory.substack.com/subscribe?"><span>Subscribe now</span></a></p><h4>Prediction Can Only Get You So Far&#8230;</h4><p>I don&#8217;t want to get too technical here, but it is important to understand why these types of things are so hard for LLMs and why the results I am reporting here are significant. Since the first vision model, GPT-4, was released in February 2023, we&#8217;ve seen HTR scores steadily improve to the point that they get about 90% (or more) of a given text correct. Much of this can be chalked up to technical improvements in image processing and better training data, but it&#8217;s that last 10% that I&#8217;ve been talking about above.</p><p>Remember that LLMs are inherently predictive by nature, trained to choose the most probable way to complete a sequence like &#8220;the cat sat on the &#8230;&#8221;. They are, in effect, made up of tables which record those probabilities. Spelling errors and stylistic inconsistencies are, by definition, unpredictable, low probability answers and so LLMs must chafe against their training data to transcribe &#8220;the cat sat on the rugg&#8221; instead of &#8220;mat&#8221;. This is also why LLMs are not very good at transcribing unfamiliar people&#8217;s names (especially last names), obscure places, dates, or numbers such as sums of money. </p><p>From a statistical point of view, these all appear as arbitrary choices to an LLM with no meaningful differences in their statistical probabilities: in isolation, one is no more likely than another. Was a letter written by Richard Darby or Richard Derby? Was it dated 15 March 1762 or 16 March 1782? Did the author enclosed a bill for 339 dollars or 331 dollars? Correct answer to those questions cannot normally be predicted from the preceding contents of a letter. You need other types of information to find the answer when letters prove indecipherable. Yet basic correctness on these types of information&#8212;names, dates, places, and sums&#8212;is a prerequisite to their being useful to me as a historian. This makes the final mile of accuracy the only one that really counts. </p><h4>On Scaling, Plateaus, and Benchmarks</h4><p>More importantly, these issues with handwriting recognition are only one small facet of a much larger debate about whether the predictive architecture behind LLMs is inherently limiting or whether scaling (making the models larger) will allow the models to break free of regurgitation and do something new.</p><p>So when I benchmark an LLM on handwriting, in my mind I feel I am also getting some insight into that larger question of whether LLMs are plateauing or continuing to grow in capabilities. To benchmark LLM handwriting accuracy, l<a href="https://www.tandfonline.com/doi/abs/10.1080/01615440.2025.2500309">ast year Dr. Lianne Leddy and I developed a set of 50 documents comprising some 10,000 words</a>&#8212;we had to choose them carefully and experiment to ensure that these documents were not already in the LLM training data (full disclosure: we can&#8217;t know for sure, but we took every reasonable precaution). We&#8217;ve written about the set <a href="https://www.tandfonline.com/doi/abs/10.1080/01615440.2025.2500309">several times before</a>, but in short it includes dozens of different hands, images captured with a variety of tools from smartphones to scanners, and document with different styles of writing from virtually illiterate scrawl to formal secretary hand. In my experience, they are representative of the types of documents that I, and English-language historian currently working on 18<sup>th</sup> and 19<sup>th</sup> century records, most often encounter.</p><p>We measure transcription error rates in terms of the percentage of incorrect characters (CER) and words (WER) in a given text. These are standardized but blunt instruments: a word may be spelled correctly but if the first letter is wrongly capitalized or it is followed with a comma rather than a semicolon it counts as an erroneous word. But what constitutes an error is also not always clear. Capitalization and punctuation were not standardized until the 20<sup>th</sup> century (in English) and are often ambiguous in historical documents. Another example: should we transcribe the long f (as in le&#383;s) using an &#8220;f&#8221; for the first &#8220;s&#8221; or just write it out as &#8220;less&#8221;? That&#8217;s a judgement call. Sometimes letters and whole words are simply indecipherable and up for interpretation.</p><p>In truth, it&#8217;s usually impossible to score 100% accuracy in most real-world scenarios. Studies show that non-professionals typically score <a href="https://www.tandfonline.com/doi/abs/10.1080/01615440.2025.2500309">WERs of 4-10%</a>. Even professional transcription services expect a few errors. They typically guarantee a 1% WER (or around a 2-3% CER), but only when the texts are clear and readable. So that is essentially the ceiling in terms of accuracy.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!-RQJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75343be3-e3ef-4867-af15-5e683a6efcc0_1828x843.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!-RQJ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75343be3-e3ef-4867-af15-5e683a6efcc0_1828x843.png 424w, https://substackcdn.com/image/fetch/$s_!-RQJ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75343be3-e3ef-4867-af15-5e683a6efcc0_1828x843.png 848w, https://substackcdn.com/image/fetch/$s_!-RQJ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75343be3-e3ef-4867-af15-5e683a6efcc0_1828x843.png 1272w, https://substackcdn.com/image/fetch/$s_!-RQJ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75343be3-e3ef-4867-af15-5e683a6efcc0_1828x843.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!-RQJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75343be3-e3ef-4867-af15-5e683a6efcc0_1828x843.png" width="1456" height="671" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/75343be3-e3ef-4867-af15-5e683a6efcc0_1828x843.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:671,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:153403,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://generativehistory.substack.com/i/176385680?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75343be3-e3ef-4867-af15-5e683a6efcc0_1828x843.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!-RQJ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75343be3-e3ef-4867-af15-5e683a6efcc0_1828x843.png 424w, https://substackcdn.com/image/fetch/$s_!-RQJ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75343be3-e3ef-4867-af15-5e683a6efcc0_1828x843.png 848w, https://substackcdn.com/image/fetch/$s_!-RQJ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75343be3-e3ef-4867-af15-5e683a6efcc0_1828x843.png 1272w, https://substackcdn.com/image/fetch/$s_!-RQJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75343be3-e3ef-4867-af15-5e683a6efcc0_1828x843.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 1: Performance of Trasnskribus, Humans, and Google models on HTR over time</figcaption></figure></div><p>Last winter, on our test-set, Gemini-2.5-Pro began to score in the human range: a strict CER of 4% and WER of 11%. When we excluded errors of punctuation and capitalization&#8212;errors that don&#8217;t change the actual meaning of the text or its usefulness for search and readability purposes&#8212;those scores dropped to CERs of 2% and WERs of 4%. The best specialized HTR software achieves CERs around 8% and WERs around 20% without specialized training which reduces errors rates to about those of Gemini-2.5-Pro. Improvement has indeed been steady across each generation of models. Those of Gemini-2.5-Pro were about 50-70% better than the ones we reported for Gemini-1.5-Pro a few months before, which were about 50-70% better than the initial scores reported for GPT-4 a few months before that. A similar progression is evident in Google&#8217;s faster, cheaper version of Gemini-FlashThe open question has been: will they keep improving at a similar rate.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://generativehistory.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://generativehistory.substack.com/subscribe?"><span>Subscribe now</span></a></p><h4>Expert Human Performance?</h4><p>On a (Canadian) Thanksgiving trip to visit family, I started to play with the new Google model. Here is what I had to do to access it. First, I uploaded an image to <a href="https://aistudio.google.com/">AI Studio</a>, and gave it the following system instructions (the same ones we&#8217;ve used on all our tests&#8230;I&#8217;d like to modify them but I need to keep them consistent across all the tests):</p><blockquote><p>&#8220;Your task is to accurately transcribe handwritten historical documents, minimizing the CER and WER. Work character by character, word by word, line by line, transcribing the text exactly as it appears on the page. To maintain the authenticity of the historical text, retain spelling errors, grammar, syntax, and punctuation as well as line breaks. Transcribe all the text on the page including headers, footers, marginalia, insertions, page numbers, etc. If these are present, insert them where indicated by the author (as applicable). In your final response write &#8220;Transcription:&#8221; followed only by your transcription.&#8221;</p></blockquote><p>But then I had to wait for the result and manually retry the prompt, over and over again&#8212;sometimes 30 or more times&#8212;until I was given a choice between two answers. Needless to say, this was time consuming, expensive, and I repeatedly hit rate limits which delayed things even more. As a result, I could only get through five documents from our set. In response, I tried to choose the most error-prone and difficult to decipher documents from the set, texts that are not only written in a messy hand but are full of spelling and grammatical errors, lacking in proper punctuation, and that contain lots of inconsistent capitalization. My goal was not to be definitive&#8212;that will come later&#8212;but to get a sense of what this model could do.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!JvMG!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4d0d32e2-17f6-40c0-85f0-33819976b241_1836x1876.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!JvMG!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4d0d32e2-17f6-40c0-85f0-33819976b241_1836x1876.png 424w, https://substackcdn.com/image/fetch/$s_!JvMG!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4d0d32e2-17f6-40c0-85f0-33819976b241_1836x1876.png 848w, https://substackcdn.com/image/fetch/$s_!JvMG!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4d0d32e2-17f6-40c0-85f0-33819976b241_1836x1876.png 1272w, https://substackcdn.com/image/fetch/$s_!JvMG!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4d0d32e2-17f6-40c0-85f0-33819976b241_1836x1876.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!JvMG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4d0d32e2-17f6-40c0-85f0-33819976b241_1836x1876.png" width="1836" height="1876" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4d0d32e2-17f6-40c0-85f0-33819976b241_1836x1876.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1876,&quot;width&quot;:1836,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:611843,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://generativehistory.substack.com/i/176385680?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83149f5e-3ebb-49cb-86fd-436989bc5058_1836x1876.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!JvMG!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4d0d32e2-17f6-40c0-85f0-33819976b241_1836x1876.png 424w, https://substackcdn.com/image/fetch/$s_!JvMG!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4d0d32e2-17f6-40c0-85f0-33819976b241_1836x1876.png 848w, https://substackcdn.com/image/fetch/$s_!JvMG!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4d0d32e2-17f6-40c0-85f0-33819976b241_1836x1876.png 1272w, https://substackcdn.com/image/fetch/$s_!JvMG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4d0d32e2-17f6-40c0-85f0-33819976b241_1836x1876.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 2: The AI Studio interface showing the A/B Test rather than a single output.</figcaption></figure></div><p>The results were immediately stunning. On each of the five documents I transcribed (totalling a little over 1,000 words or 10% of our total sample), the model achieved a strict CER of 1.7% and WER of 6.5%&#8212;in other words, about 1 in 50 characters were wrong including punctuation marks and capitalization. But as analyzed the data I saw something new: for the first time, nearly all the errors were capitalization and punctuation, very few were actual words. I also found that a lot of the punctuation marks and capital letters it was getting wrong were actually highly ambiguous. When those types of errors were excluded from the count, th<strong>e error rates fell to a modified CER of 0.56% and WER of 1.22%</strong>. In other words, the new Gemini model was only getting about 1 in 200 characters wrong, not counting punctuation marks and capital letters.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!t0Jz!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6fc9ae77-331e-4ef0-b4b1-483f15d2bcee_1187x626.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!t0Jz!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6fc9ae77-331e-4ef0-b4b1-483f15d2bcee_1187x626.png 424w, https://substackcdn.com/image/fetch/$s_!t0Jz!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6fc9ae77-331e-4ef0-b4b1-483f15d2bcee_1187x626.png 848w, https://substackcdn.com/image/fetch/$s_!t0Jz!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6fc9ae77-331e-4ef0-b4b1-483f15d2bcee_1187x626.png 1272w, https://substackcdn.com/image/fetch/$s_!t0Jz!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6fc9ae77-331e-4ef0-b4b1-483f15d2bcee_1187x626.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!t0Jz!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6fc9ae77-331e-4ef0-b4b1-483f15d2bcee_1187x626.png" width="1187" height="626" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6fc9ae77-331e-4ef0-b4b1-483f15d2bcee_1187x626.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:626,&quot;width&quot;:1187,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:703948,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://generativehistory.substack.com/i/176385680?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6fc9ae77-331e-4ef0-b4b1-483f15d2bcee_1187x626.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!t0Jz!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6fc9ae77-331e-4ef0-b4b1-483f15d2bcee_1187x626.png 424w, https://substackcdn.com/image/fetch/$s_!t0Jz!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6fc9ae77-331e-4ef0-b4b1-483f15d2bcee_1187x626.png 848w, https://substackcdn.com/image/fetch/$s_!t0Jz!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6fc9ae77-331e-4ef0-b4b1-483f15d2bcee_1187x626.png 1272w, https://substackcdn.com/image/fetch/$s_!t0Jz!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6fc9ae77-331e-4ef0-b4b1-483f15d2bcee_1187x626.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 3: A good side by side comparison on a particularly difficult document.No other model comes close on this letter.</figcaption></figure></div><p><strong>The new Gemini model&#8217;s performance on HTR meets the criteria for expert human performance</strong>. These results are also 50-70% better than those achieved by Gemini-2.5-Pro. In two years, we have in effect gone from transcriptions that were little more than gibberish to expert human levels of accuracy. And the consistency in the leap between each generation of model is exactly what you would expect to see if scaling laws hold: as a model gets bigger and more complex, you should be able to predict how well it will perform on tasks like this just by knowing the size of the model alone.</p><h4><strong>The Ultimate Test</strong></h4><p>Here is where is starts to get really weird and interesting. Fascinated with the results, I decided to push the model further. Up to this point, no model has been able to reliably decipher tabular handwritten data, the kind of data we find in merchant ledger, account books, and daybooks. These are extremely difficult to decipher for humans but (until now) nearly impossible for LLMs because there is very little about the text that is predictive. </p><p>Take this page (Figure 4) from a 1758 Albany merchant&#8217;s daybook (a running tally of sales) which is especially hard to read. It is messy, to be sure, but was also kept in English by a Dutch clerk who may not have spoken much English and whose spelling and letter formation was highly irregular, mixing Dutch and English together. The sums in the accounts were also written in the <a href="https://www.royalmintmuseum.org.uk/journal/history/pounds-shillings-and-pence/">old style of pounds / shillings / pence</a> using a shorthand typical of the period: &#8220;To 30 Gallons Rum @4/6 6/15/0&#8221;. This means that someone purchased (a charge to their account) 30 gallons of rum where each gallon cost 4 shillings and 6 pence for a total of 6 pounds, 15 shillings, and 0 pence. </p><p>To most people today, this non-decimalized way of measuring money is foreign: there are 12 pennies (pence) in a shilling and 20 shillings in pound (see <a href="https://www.royalmintmuseum.org.uk/journal/history/pounds-shillings-and-pence/">this description by the Royal Mint)</a>. Individual transactions were written into the book as they happened, divided from one another by a horizontal rule with a number signifying the day of the month written in the middle. Each transaction was recorded as s debt (Dr), that is a purchase, or a Credit (Cr) meaning a payment. Some transactions were also crossed out, probably to indicate they had been balanced or transferred to the client&#8217;s account in the merchant&#8217;s main ledger (similar to when a pending transaction is posted in your online banking). And none of this was written in a standardized way.</p><p>LLMs have had a hard time with such books, not only because there is very limited training data available for these types of records (ledgers are less likely to digitized and even less likely to be transcribed than diaries or letters because: who wants to read them unless they have to?) but because none of this is predictive: a person can buy any amount of anything at any arbitrary cost recorded in sums that don&#8217;t add up according to conventional methods&#8230;which LLMs have had enough issues with over the years. I&#8217;ve found that models can often decipher some of the names and some of the items in a ledger, but become utterly lost on the numbers. They have a hard time transcribing digits in general (again you can&#8217;t predict whether it&#8217;s 30 or 80 gallons if the first digit is poorly formed), but also tended to merge the item costs and totals together. In effect, they often don&#8217;t seem to realize that the old-style sums are amounts of money at all. Telling them to check the numbers by adding the totals together does not help and often makes things worse. Especially complex pages temporarily break the model, causing it to repeat certain numbers or phrases repeatedly until they reach their output limits. Other times they think for a long time and then fail to answer entirely.</p><h4><strong>Deux Ex Machina</strong></h4><p>But there is a something in this new machine that is markedly different. From my admittedly limited tests, the new Gemini model handles this type of data much better than any previous model or student I&#8217;ve encountered: after completing the five documents from our test-set I uploaded the Albany merchant&#8217;s daybook page above (Figure 4) with the same prompt, just to see what would happen and amazingly, it was again almost perfect. The numbers are, remarkably, all correct. More interesting, though, are that its errors are actually corrections or clarifications. For example, when Samuel Stitt purchased 2 punch bowls, the clerk recorded that they cost 2/ each meaning 2 shillings each; for brevity&#8217;s sake he implied 0 pennies rather than writing it out. Yet for consistency, the model transcribed this as @2/0 which is actually a more correct way of writing the sum and clarifies the meaning. Strictly speaking, though, it is an error.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!M7Fk!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1cdc11e4-6cfa-40f3-a00a-40f139e85dd5_712x745.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!M7Fk!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1cdc11e4-6cfa-40f3-a00a-40f139e85dd5_712x745.png 424w, https://substackcdn.com/image/fetch/$s_!M7Fk!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1cdc11e4-6cfa-40f3-a00a-40f139e85dd5_712x745.png 848w, https://substackcdn.com/image/fetch/$s_!M7Fk!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1cdc11e4-6cfa-40f3-a00a-40f139e85dd5_712x745.png 1272w, https://substackcdn.com/image/fetch/$s_!M7Fk!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1cdc11e4-6cfa-40f3-a00a-40f139e85dd5_712x745.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!M7Fk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1cdc11e4-6cfa-40f3-a00a-40f139e85dd5_712x745.png" width="712" height="745" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1cdc11e4-6cfa-40f3-a00a-40f139e85dd5_712x745.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:745,&quot;width&quot;:712,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:583823,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://generativehistory.substack.com/i/176385680?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1cdc11e4-6cfa-40f3-a00a-40f139e85dd5_712x745.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!M7Fk!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1cdc11e4-6cfa-40f3-a00a-40f139e85dd5_712x745.png 424w, https://substackcdn.com/image/fetch/$s_!M7Fk!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1cdc11e4-6cfa-40f3-a00a-40f139e85dd5_712x745.png 848w, https://substackcdn.com/image/fetch/$s_!M7Fk!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1cdc11e4-6cfa-40f3-a00a-40f139e85dd5_712x745.png 1272w, https://substackcdn.com/image/fetch/$s_!M7Fk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1cdc11e4-6cfa-40f3-a00a-40f139e85dd5_712x745.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 5: Transcription by new unknown Gemini model of page from the Albany Account Book</figcaption></figure></div><p>In tabulating the &#8220;errors&#8221; I saw the most astounding result I have ever seen from an LLM, one that made the hair stand up on the back of my neck. Reading through the text, I saw that Gemini had transcribed a line as &#8220;To 1 loff Sugar 14 lb 5 oz @ 1/4 0 19 1&#8221;. If you look at the actual document, you&#8217;ll see that what is actually written on that line is the following: &#8220;To 1 loff Sugar 145 @ 1/4 0 19 1&#8221;. For those unaware, in the 18<sup>th</sup> century sugar was sold in a hardened, conical form and Mr. Slitt was a storekeeper buying sugar in bulk to sell. At first glance, this appears to be a hallucinatory error: the model was told to transcribe the text exactly as written but it inserted 14 lb 5 oz which is not in the document. This was exactly the type of errors I&#8217;ve seen many times before: in the absence of good context the model guessed, inserting a hallucination. But then I realized that it had actually done some extremely clever.</p><p>What Gemini did was to correctly infer that the digits 1, 4, 5 were units of measurement describing the total weight of sugar purchased. This was not an obvious conclusion to draw, though, from the document itself. All the other nineteen entries clearly specify total units of purchase up front: 30 gallons, 17 yds, 1 barrel and so on. The sugar loaf entry does this too (1 loaf is written at the start of the entry) and it is the only one that lists a number at the end of the description. There is a tiny mark above the 1 which may also (ambiguously) have been used to indicate pounds (thanks to Thomas Wein for noticing this). But if Gemini  interpreted it this way, it would also have read the phrase as something like 1 lb 45 or 145 lb, given the placement of the mark above the 1. It was also able to glean from the text that sugar was being sold at 1 shilling and 4 pence per <em>something</em>, and inferred that this something was pounds.</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!PyC4!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16937510-c583-4c7b-91e8-aefd45c16912_273x45.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!PyC4!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16937510-c583-4c7b-91e8-aefd45c16912_273x45.png 424w, https://substackcdn.com/image/fetch/$s_!PyC4!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16937510-c583-4c7b-91e8-aefd45c16912_273x45.png 848w, https://substackcdn.com/image/fetch/$s_!PyC4!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16937510-c583-4c7b-91e8-aefd45c16912_273x45.png 1272w, https://substackcdn.com/image/fetch/$s_!PyC4!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16937510-c583-4c7b-91e8-aefd45c16912_273x45.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!PyC4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16937510-c583-4c7b-91e8-aefd45c16912_273x45.png" width="273" height="45" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/16937510-c583-4c7b-91e8-aefd45c16912_273x45.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:45,&quot;width&quot;:273,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:5738,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://generativehistory.substack.com/i/176385680?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37853f70-06b5-473f-8877-7d1f7c83ed86_276x748.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!PyC4!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16937510-c583-4c7b-91e8-aefd45c16912_273x45.png 424w, https://substackcdn.com/image/fetch/$s_!PyC4!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16937510-c583-4c7b-91e8-aefd45c16912_273x45.png 848w, https://substackcdn.com/image/fetch/$s_!PyC4!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16937510-c583-4c7b-91e8-aefd45c16912_273x45.png 1272w, https://substackcdn.com/image/fetch/$s_!PyC4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16937510-c583-4c7b-91e8-aefd45c16912_273x45.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a><figcaption class="image-caption">Figure 6: Close-up of the transcription</figcaption></figure></div><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Ub6L!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4e135b3a-7adb-4250-b1ae-1fcae452dbbe_1139x256.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Ub6L!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4e135b3a-7adb-4250-b1ae-1fcae452dbbe_1139x256.png 424w, https://substackcdn.com/image/fetch/$s_!Ub6L!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4e135b3a-7adb-4250-b1ae-1fcae452dbbe_1139x256.png 848w, https://substackcdn.com/image/fetch/$s_!Ub6L!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4e135b3a-7adb-4250-b1ae-1fcae452dbbe_1139x256.png 1272w, https://substackcdn.com/image/fetch/$s_!Ub6L!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4e135b3a-7adb-4250-b1ae-1fcae452dbbe_1139x256.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Ub6L!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4e135b3a-7adb-4250-b1ae-1fcae452dbbe_1139x256.png" width="1139" height="256" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4e135b3a-7adb-4250-b1ae-1fcae452dbbe_1139x256.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:256,&quot;width&quot;:1139,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:486859,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://generativehistory.substack.com/i/176385680?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc054ce60-ae5c-478d-b1f0-97ddbd195a52_1622x892.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Ub6L!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4e135b3a-7adb-4250-b1ae-1fcae452dbbe_1139x256.png 424w, https://substackcdn.com/image/fetch/$s_!Ub6L!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4e135b3a-7adb-4250-b1ae-1fcae452dbbe_1139x256.png 848w, https://substackcdn.com/image/fetch/$s_!Ub6L!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4e135b3a-7adb-4250-b1ae-1fcae452dbbe_1139x256.png 1272w, https://substackcdn.com/image/fetch/$s_!Ub6L!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4e135b3a-7adb-4250-b1ae-1fcae452dbbe_1139x256.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a><figcaption class="image-caption">Figure 7: Closeup of the Original Document</figcaption></figure></div><p>To determine the correct obverse weight, decoding the 145, Gemini then did something remarkable: it worked through the numbers, using the final total cost of 0/19/1 to work backwards to determine the weight, a series of operations that would require it to convert between two decimalized and two non-decimalized systems of measurement. While we don&#8217;t know its actual reasoning process, it must have been something akin to this: the sugar cost 1 shilling and 4 pence per unit, and that that sum can also be expressed as 16 pence. We also know that the total value of the sale was 0 pounds, 19 shillings, and 1 penny, so we can express this as 229 pence to create a common unit of comparison. To find how much sugar was purchased we then divide 229 by 16 to get the result: 14.3125 or 14 and 5/16 or 14 lb 5 oz. Therefore, Gemini concluded, it was not 1 45, nor 145 but 14 5 and then 14 lb 5 oz, and it chose to clarify this in its transcription. </p><p>[Added 17/10/2025]: If that ambiguous mark above the 1 tipped it off that the 145 was a measurement in pounds, the result was a similar process of logical deduction and self correction. In that case, Gemini would have had to intentionally question the most obvious version of the transcription, realizing (in effect) that 1 lb 45 or 145 lbs (which is the only way to read the original) did not balance with the tally of 0 19 1. Getting to 14 lb 5 oz would then arise from the same process as above.</p><p>This is exactly the type of logic problem at which LLMs often fail: first there is the ambiguity in the writing itself and in the form of the text, then the double meaning of the word &#8220;pounds&#8221;, and finally the need to convert back and forth between not one but two different non-decimalized systems of measurement. And no one asked Gemini to do this. It took the initiative to investigate and clarify the meaning of the ambiguous number all on its own. And it was correct.</p><p>In my testing, no other model has done anything like this when tasked with transcribing the same document. Indeed even if you give Gemini-2.5-Pro hints, asking it to pay attention to missing units of measurement it occasionally inserst &#8220;lb&#8221; or &#8220;wt&#8221; after the 5 in 145, but deletes the other numbers. GPT-5 Pro typically transcribes the line as: &#8220;To 1 Loaf Sugar 1 lb 5 0 19 1&#8221;. Interestingly, you can nudge both GPT-5 and Gemini-2.5-Pro towards the correct answer by asking it what the numbers 1 4 5 mean in the sugar loaf entry. And even then answers vary, often suggesting that it was 145 lbs of sugar rather than 14 lb 5 oz.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!W-ph!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72a96a76-0f6f-400f-ae78-fadf4287cc09_1249x1849.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!W-ph!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72a96a76-0f6f-400f-ae78-fadf4287cc09_1249x1849.png 424w, https://substackcdn.com/image/fetch/$s_!W-ph!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72a96a76-0f6f-400f-ae78-fadf4287cc09_1249x1849.png 848w, https://substackcdn.com/image/fetch/$s_!W-ph!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72a96a76-0f6f-400f-ae78-fadf4287cc09_1249x1849.png 1272w, https://substackcdn.com/image/fetch/$s_!W-ph!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72a96a76-0f6f-400f-ae78-fadf4287cc09_1249x1849.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!W-ph!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72a96a76-0f6f-400f-ae78-fadf4287cc09_1249x1849.png" width="1249" height="1849" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/72a96a76-0f6f-400f-ae78-fadf4287cc09_1249x1849.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1849,&quot;width&quot;:1249,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:152502,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://generativehistory.substack.com/i/176385680?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72a96a76-0f6f-400f-ae78-fadf4287cc09_1249x1849.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!W-ph!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72a96a76-0f6f-400f-ae78-fadf4287cc09_1249x1849.png 424w, https://substackcdn.com/image/fetch/$s_!W-ph!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72a96a76-0f6f-400f-ae78-fadf4287cc09_1249x1849.png 848w, https://substackcdn.com/image/fetch/$s_!W-ph!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72a96a76-0f6f-400f-ae78-fadf4287cc09_1249x1849.png 1272w, https://substackcdn.com/image/fetch/$s_!W-ph!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72a96a76-0f6f-400f-ae78-fadf4287cc09_1249x1849.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>I have diligently tried to replicate this result, but sadly after hundreds of refreshes on AI Studio, I have yet to see the A/B test again on this document. I suspect that Google may have ended it, or at least for me.</p><h4>Symbolic Reasoning and LLMs</h4><p>What makes this example so striking is that it seems to cross a boundary that some experts have long claimed current models cannot pass. Strictly speaking, the Gemini model is not engaging in symbolic reasoning in the traditional sense: it is not manipulating explicit rules or logical propositions as a classical AI system would be expected to do. Yet its behaviour mirrors that outcome. Faced with an ambiguous number, it inferred missing context, performed a set of multi-step conversions between historical systems of currency and weight, and arrived at a correct conclusion that required abstract reasoning about the world the document described. In other words, it behaved <em>as if</em> it had access to symbols, even though none were ever explicitly defined. Did it create these symbolic representations for itself? If so, what does that mean? If not, how did it do this?</p><p>What appears to be happening here is a form of emergent, implicit reasoning, the spontaneous combination of perception, memory, and logic inside a statistical model that (I don&#8217;t believe&#8230;Google please clarify!) was designed to reason symbolically at all. And the point is that we don&#8217;t know what it <em>actually</em> did or why. </p><p>The safer view is to assume that Gemini did not &#8220;know&#8221; that it was solving a problem of eighteenth-century arithmetic at all, but its internal representations were rich enough to emulate the process of doing so. But that answer seems to ignore the obvious facts: it followed an intentional, analytical process across several layers of symbolic abstraction, all unprompted. This seems new and important. </p><p>If this behaviour proves reliable and replicable, it points to something profound that <a href="https://importai.substack.com/p/import-ai-431-technological-optimism">the labs are also starting to admit</a>: that true reasoning may not require explicit rules or symbolic scaffolding to arise, but can instead emerge from scale, multimodality, and exposure to enough structured complexity. In that case, the sugar-loaf entry is more than a remarkable transcription, it is a small but clear (and I think unambiguous) sign that the line between pattern recognition and genuine understanding is beginning to blur.</p><h4>Conclusion</h4><p>For historians, the implications are immediate and profound. If these results hold up under systematic testing, we will be entering an era in which large language models can not only transcribe historical documents at expert-human levels of accuracy, but can also <em>reason</em> about them in historically meaningful ways. That is, they are no longer simply seeing letters and words&#8212;and correct ones at that&#8212;they are beginning to interpret context, logic, and material reality. A model that can infer the meaning of &#8220;145&#8221; as &#8220;14 lb 5 oz&#8221; in an 18th-century merchant ledger is not just performing text recognition: it is demonstrating an understanding of the economic and cultural systems in which those records were produced&#8230;and then using that knowledge to re-interpret the past in intelligible ways. This moves the work of automated transcription from a visual exercise into an interpretive one, bridging the gap between vision and reasoning in a way that mirrors what human experts do. </p><p>But the broader implications are even more striking. Handwritten Text Recognition is one of the oldest problems in the field of AI research, going back to the late 1940s before AI even had a name. For decades, AI researchers have treated handwritten text recognition as a bounded technical problem, that is an engineering challenge in vision. This began with the IBM 1287 which could read digits and five letters when it debuted in 1966 and continued on through the creation of specialized HTR models developed only a few years ago. </p><p>What this new Gemini model seems to show is that near-perfect handwriting recognition is better achieved through the generalist approach of LLMs. Moreover, the model&#8217;s ability to make a correct, contextually grounded inference that requires several layers of symbolic reasoning suggests that something new may be happening inside these systems&#8212;an emergent form of abstract reasoning that arises not from explicit programming but from scale and complexity itself.</p><p>If so, the &#8220;handwriting problem&#8221; may turn out to have been a proxy for something much larger. What began with a test on the readability of old documents may now be revealing, by accident, the beginnings of machines that can actually reason in abstract, symbolic ways about the world they see.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://generativehistory.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Generative History! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Open-Source Archive Studio App Available for PC Users]]></title><description><![CDATA[Users who aren't comfortable running python scripts can now download the open-source PC App]]></description><link>https://generativehistory.substack.com/p/open-source-archive-studio-app-available</link><guid isPermaLink="false">https://generativehistory.substack.com/p/open-source-archive-studio-app-available</guid><dc:creator><![CDATA[Mark Humphries]]></dc:creator><pubDate>Mon, 16 Jun 2025 18:25:19 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!KgFK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9f4ba1b-0c80-43c4-b904-4b3aefdfe8e5_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!KgFK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9f4ba1b-0c80-43c4-b904-4b3aefdfe8e5_1024x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!KgFK!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9f4ba1b-0c80-43c4-b904-4b3aefdfe8e5_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!KgFK!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9f4ba1b-0c80-43c4-b904-4b3aefdfe8e5_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!KgFK!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9f4ba1b-0c80-43c4-b904-4b3aefdfe8e5_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!KgFK!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9f4ba1b-0c80-43c4-b904-4b3aefdfe8e5_1024x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!KgFK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9f4ba1b-0c80-43c4-b904-4b3aefdfe8e5_1024x1024.png" width="1024" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b9f4ba1b-0c80-43c4-b904-4b3aefdfe8e5_1024x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1582098,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://generativehistory.substack.com/i/166082603?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9f4ba1b-0c80-43c4-b904-4b3aefdfe8e5_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!KgFK!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9f4ba1b-0c80-43c4-b904-4b3aefdfe8e5_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!KgFK!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9f4ba1b-0c80-43c4-b904-4b3aefdfe8e5_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!KgFK!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9f4ba1b-0c80-43c4-b904-4b3aefdfe8e5_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!KgFK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9f4ba1b-0c80-43c4-b904-4b3aefdfe8e5_1024x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://generativehistory.substack.com/p/open-source-archive-studio-app-available?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://generativehistory.substack.com/p/open-source-archive-studio-app-available?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><p>For those of you who are interested in trying out <em>Archive Studio</em>, but don&#8217;t feel comfortable running python scripts, we&#8217;ve added <a href="https://github.com/mhumphries2323/Archive_Studio/releases/tag/v1.1">a new executable file</a> to the project&#8217;s <a href="https://github.com/mhumphries2323/Archive_Studio">GitHub Repository</a>. In plain English, this means that if you want to download a Windows App that you can just double-click and use, it is <a href="https://github.com/mhumphries2323/Archive_Studio/releases/download/v1.1/ArchiveStudio.exe">now available</a>. Right now it is only available for PC users but if any Mac programmers would like to help by creating an executable that would run on Apple computers, please let us know!</p><p>An important note: when you run <em>Archive Studio,</em> Windows Defender or other antivirus software on your PC may read the program as a virus or other potentially malicious software. Rest assured, it is not. The main reason this happens is that the program makes API calls to an external server. Professional software developers have access to certifications that allow them to verify the program&#8217;s authenticity with various antivirus providers, but that is not really something we are able to do. So apologies for the scary warnings, but if this happens, you can choose to add an exception and/or run the program anyway. If you are not sure how to do that, ask <a href="https://chatgpt.com/">ChatGPT </a>how to do it. It is really good at explaining those types of technical problems&#8212;and it is patient!</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://generativehistory.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://generativehistory.substack.com/subscribe?"><span>Subscribe now</span></a></p><h4>Getting Started</h4><p>To use <em>Archive Studio</em>, you will need to get yourself at least one API key from Google, Anthropic, or OpenAI. This may sound scary but it takes a few minutes at most: it is as simple as signing up for any website and takes about three clicks. If you can subscribe to a newspaper, you can do this!</p><p>The process is outlined in the <a href="https://github.com/mhumphries2323/Archive_Studio/blob/main/Manual.pdf">user manual</a>, but here it is for each provider:</p><blockquote><p><em>OpenAI API Keys</em></p><p>To register for an OpenAI API key, first create an account on the <a href="https://platform.openai.com/tokenizer">OpenAI platform</a> (note that this will be different than a ChatGPT account). Click your profile icon in the top right-hand corner and select &#8220;Your Profile&#8221;. In the menu on the left, click API-keys and then click the green &#8220;+Create New Secret Key&#8221; at the top right. Give your Key a name and create the key.</p><p><em>Anthropic API Keys</em></p><p>To register for an Anthropic API key, first create an account on the <a href="https://console.anthropic.com/login">Anthropic consol</a>e (this will be different than your Claude chat account). After you create an account, cock your profile icon at the top right and select &#8220;API Keys&#8221; from the dropdown and then select the orange &#8220;+Create Key&#8221; button at the top right.</p><p><em>Gemini API Keys</em></p><p>The process with Google can be somewhat more confusing as it provides a number of different ways to access their models via API, specifically via the Google AI Studio and Vertex platforms. You need to use the Google AI Studio platform. First, <a href="https://ai.google.dev/aistudio">create a Google AI Studio Account</a> then, once you are logged in, select the blue &#8220;Get API Key&#8221; button on the left. In the next pane, click the blue &#8220;Create API Key&#8221; button.</p></blockquote><h4>Recommendations and Costs</h4><p>If you are going to use just one provider, we would strongly recommend Google, not only because it is the best model currently available for these types of tasks but also because it is the most affordable. </p><p>The Google Gemini API provides a free tier to try out the service. It is also extremely affordable when you add a credit card number: we&#8217;ve found that the 2.5 Pro model is about $0.006 (or just over half a cent) USD per page while the 2.5 Flash model (which is often just as good) is only $0.0005 per page. We&#8217;d recommend using the Flash model for all other tasks like document segmentation, identifying named people and places, generating metadata, etc.</p><p>Be sure to review the costs associated with the model you choose to use before making large requests. To be clear, each of the processing operations (recognize text, correct, format, separate document, get names and places, generate metadata, etc) all involve calls to an API which cost you money. Although the costs per request are small, they can start to add up if you don&#8217;t pay attention. All three of the providers above allow you to monitor your usage in realtime and you should do so carefully at first, at least until you get a sense of how much various operations cost and their relative usefulness.</p><h4>Errors and Rate Limits</h4><p>When you send a document to an LLM via <em>Archive Studio</em>, you are doing so through an API which makes a call to the LLM&#8217;s server behind the scenes. These servers have rate limits which are like quotas dictating how many requests you can make per minute or per day as well as how much information you can send at one time. To speed things up, Archive Studio is designed to send documents in parallel (meaning at the same time) which may not work as well if you are not on a paid tier for one of the API services described above.</p><p>If all of your operations begin to fail (IE you get an error message with 0/10 documents completed and 10 errors) it&#8217;s likely you&#8217;ve hit your rate limit. You can change how many documents you try to send at once in Settings &#8594; API Settings &#8594; Batch Size. If you are on the free tier of Gemini, for example, you should set this to 1 as Gemini only allows 3 requests per minute and 15 per day for the pro model at that level. The first paid tier allows 150 requests per minute and 1000 per day. We suggest you review the <a href="https://ai.google.dev/gemini-api/docs/rate-limits#free-tier">Gemini</a>, <a href="https://docs.anthropic.com/en/api/rate-limits">Anthropic</a>, and <a href="https://platform.openai.com/docs/guides/rate-limits">OpenAI </a>rate limits and set your batch size accordingly&#8212;they are all very different. </p><p>When using Gemini specifically, you may also sometimes find that one of your pages simply will not complete. When this happens, it may be because the page appears in the Gemini training data. Google has implemented a mechanism which identifies when Gemini is writing out data that was in its training set and stops the output. While the goal seems to be to prevent users from getting the models to commit overt copyright infringement, the mechanism does not discriminate between historical texts which are in the public domain and materials that are copyrighted. It only detects when the model is &#8220;reciting&#8221; from its training data. In those cases, you can switch to another model for that specific page (strangely, Gemini 2.5 Pro is much more likely to trigger this mechanism than Gemini 2.5 Flash although we don&#8217;t know why).</p><h4>Practical Tips  for Using Archive Studio</h4><p>In the previous version of this program (Transcription Pearl), we recommended using one model for transcription and a second model for correction. As <a href="https://generativehistory.substack.com/p/stochastic-canaries-in-the-coalmine">Mark wrote about </a>earlier, though, most often this is no longer necessary. The Gemini 2.5 Pro models have gotten considerably better at transcription and are (for now) the gold standard. Correcting them actually tends to make the transcription worse. On our test set, Gemini 2.5 Pro achieves a raw Character Error Rate (CER) of 3.5% and Word Error Rate (WER) of 10.95%. That might sound high, but as <a href="https://www.tandfonline.com/doi/full/10.1080/01615440.2025.2500309">we discuss in a recent article</a> we published with our student research team in <em><a href="https://www.tandfonline.com/journals/vhim20">Historical Methods</a></em>, when you ignore capitalization and punctuation errors, for Gemini 2.5 Pro those drop to 2.06% and 4.31% respectively. In practice, many transcriptions are great first drafts at the very least.</p><p>That said, the models are really strange. Sometimes they just aren&#8217;t able to properly read a text and the reasons are rarely obvious. It doesn&#8217;t happen often, but we&#8217;ve seen them excel at poorly written texts in low quality images but fail at standard 18th century handwriting in high resolution images. Gemini is also a &#8220;thinking&#8221; model meaning it&#8217;s intended to reminate before it answers. We&#8217;ve found this actually causes a significant performance drop and, again, the reasons are not immediately clear. In this latest version of Archive Studio, we&#8217;ve turned thinking off for Gemini models&#8212;a feature that only recently became available&#8212;which is one of the reasons we were waiting to announce the availability of the program.</p><p>So here are some practical tips if you&#8217;re not getting great results.</p><ol><li><p><strong>Don&#8217;t give the model two pages from a book at the same time because it is too much text at once.</strong> If you have a photograph that includes two pages of text, use the Image Editing utility (in the Tools menu) to split the image along the page boundary.</p></li><li><p><strong>Don&#8217;t leave too much space around your document in the image.</strong> If your document has larger areas of null space (a desk, blank microfilm, or other pages of text under the relevant image) this will confuse the model. They do not read like we do&#8230;they take everything in at once. Use the crop tool in the Image Editing utility to crop the image to a reasonable size (there is an auto-crop feature that works well when the borders between your document and the background are relatively clear).</p></li><li><p><strong>Try editing the instructions for text recognition.</strong> Under the settings menu, try playing around with the instructions you give the model under Functions &#8594; HTR. If you find that the model repeatedly makes the same mistakes, include a section labelled something like &#8220;RULES&#8221; and tell the model what you want it to do. For example, if it consistently reads a capital T as an F, tell it that &#8220;you might find that capital Fs look a lot like capital Ts in this text. Make sure you check the context of the sentence before finalizing a capital F or T.&#8221; You can then save these settings, export them, and import them for future documents. Use restore defaults to get back the original prompt.</p></li></ol><h4>Conclusion</h4><p>We hope you find this program useful. Remember it is experimental and you should not trust it&#8217;s outputs uncritically (as with any generative AI tool). That said, if you find any tips or tricks that might help others, we would encourage you to leave them here in the comments.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://generativehistory.substack.com/p/open-source-archive-studio-app-available?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://generativehistory.substack.com/p/open-source-archive-studio-app-available?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://generativehistory.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Generative History! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[Steering a Middle Course on AI in the History Classroom]]></title><description><![CDATA[We need to teach our students when to use AI, when to avoid it, and how to get the most out of it&#8212;not to pretend it doesn't exist.]]></description><link>https://generativehistory.substack.com/p/steering-a-middle-course-on-ai-in</link><guid isPermaLink="false">https://generativehistory.substack.com/p/steering-a-middle-course-on-ai-in</guid><dc:creator><![CDATA[Mark Humphries]]></dc:creator><pubDate>Fri, 13 Jun 2025 10:02:45 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!yflE!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38fee99f-f6c3-491b-aca0-01a6ae90dc23_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>This is a cross-post with <a href="https://activehistory.ca/blog/2025/06/12/steering-a-middle-course-on-ai-in-the-history-classroom/">Active History</a>.</em></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!yflE!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38fee99f-f6c3-491b-aca0-01a6ae90dc23_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!yflE!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38fee99f-f6c3-491b-aca0-01a6ae90dc23_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!yflE!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38fee99f-f6c3-491b-aca0-01a6ae90dc23_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!yflE!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38fee99f-f6c3-491b-aca0-01a6ae90dc23_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!yflE!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38fee99f-f6c3-491b-aca0-01a6ae90dc23_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!yflE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38fee99f-f6c3-491b-aca0-01a6ae90dc23_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/38fee99f-f6c3-491b-aca0-01a6ae90dc23_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2623043,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://generativehistory.substack.com/i/165833013?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38fee99f-f6c3-491b-aca0-01a6ae90dc23_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!yflE!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38fee99f-f6c3-491b-aca0-01a6ae90dc23_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!yflE!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38fee99f-f6c3-491b-aca0-01a6ae90dc23_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!yflE!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38fee99f-f6c3-491b-aca0-01a6ae90dc23_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!yflE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38fee99f-f6c3-491b-aca0-01a6ae90dc23_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://generativehistory.substack.com/p/steering-a-middle-course-on-ai-in?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://generativehistory.substack.com/p/steering-a-middle-course-on-ai-in?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><p>In the last few months, there has been a growing debate about how historians should respond to AI. And that&#8217;s a good thing. I&#8217;ve <a href="https://generativehistory.substack.com/">argued</a> that we need to engage with the technology or risk becoming irrelevant. Recent pieces in <a href="https://activehistory.ca/">Active History</a> by <a href="https://activehistory.ca/blog/2024/11/14/flattened-history-ai/">Mack Penner</a> and <a href="https://activehistory.ca/blog/2025/06/11/on-generative-ai-in-the-classroom-give-up-give-in-or-stand-up/">Edward Dunsworth</a> make the case for why we should approach AI with caution and stand-up to resist its use in historical practice and teaching.</p><p>One thing on which we can all agree is that teaching critical thinking is essential&#8212;probably more so <a href="https://www.forbes.com/sites/roncarucci/2024/02/06/in-the-age-of-ai-critical-thinking-is-more-needed-than-ever/">now than ever before</a>&#8212;and that higher education generally, and history as a discipline specifically, play essential roles in that regard. I agree with Dunsworth too that it would be wrong to either throw our hands up in surrender to the machines or embrace AI as a panacea. Both will surely lead to the destruction of history and the university as we know them. I would argue, though, that the question of how to respond to AI&#8212;especially in the classroom&#8212;remains very much unresolved.</p><p>Dunsworth argues for resistance, reaffirming the intrinsic value of deliberative human thought and mindful writing by embracing the traditional, tactile, and analog. While I agree that critical thinking and engagement are essential, I don&#8217;t believe rejecting AI is a viable way to uphold those values without ultimately distorting them into something unrecognizable. Looking at issues from a variety of perspectives, so long as they are grounded in evidence is, after all, the essence of critical thinking. If we deny that generative AI can be useful at least in some circumstances&#8212;or worse, pretend it doesn&#8217;t exist and that our students don&#8217;t have to contend with it&#8212;we simply aren&#8217;t being true to the evidence.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://generativehistory.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://generativehistory.substack.com/subscribe?"><span>Subscribe now</span></a></p><h4>AI is a Thing that Exists in the World</h4><p>I have always been a pretty traditional historian which is why I felt I needed to learn about AI after I first encountered ChatGPT in late 2022: it worried me and I knew I did not know enough about it to understand its implications. What I have tried to do since is to find out what it can and cannot do for historians right now and to keep abreast of how its capabilities may evolve over time. I also try and tell other people about what I&#8217;ve found.</p><p>The hard truth is that while it was wonky at first, generative AI has evolved faster than any other technology in my lifetime. Whether AI can actually reason&#8212;and for the record that is not a settled issue among those <a href="https://www.anthropic.com/research/tracing-thoughts-language-model">researchers</a> who specialize in such matters&#8212;is entirely immaterial to the fact that it can clearly <a href="https://hbr.org/2025/04/how-people-are-really-using-gen-ai-in-2025">do a lot of practical, useful things</a> which is why <a href="https://news.harvard.edu/gazette/story/2024/10/generative-ai-embraced-faster-than-internet-pcs/">most people are using it</a>. If you&#8217;re still a skeptic, read about <a href="https://www.forbes.com/sites/forrester/2025/04/29/vibe-coding-ais-transformation-of-software-development/">vibe coding and how AI is changing software development</a>, <a href="https://www.telegraph.co.uk/news/2025/02/19/ai-superbug-mystery-two-days-scientists-10-years/">how scientists are using GenAI to make novel discoveries</a>, or <a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC11874537/">how it is being used in hospitals</a> to improve patient outcomes. The same process is starting to play out in knowledge work too and there is <a href="https://www.nytimes.com/2025/05/30/technology/ai-jobs-college-graduates.html">growing evidence that it is already reshaping the entry level job market</a>. This is the world in which we and our students must live.</p><h4>The Risks of Resistance</h4><p>So how can we simultaneously reject AI while also claiming to prepare students to live, work, and think critically in such a world? If we want our students to take us seriously and learn some of the things we are trying to teach them about critical thinking, they need to be able to trust that we are honest brokers who offer knowledge and ideas that are grounded in reality. I don&#8217;t think I would want to argue that while doctors can use AI to help improve <a href="https://jamanetwork.com/journals/jamanetworkopen/fullarticle/2825395">diagnostics</a> and <a href="https://www.mountsinai.org/about/newsroom/2024/ai-can-help-improve-er-admission-decisions-mount-sinai-study-finds">admissions decisions</a>, a history student can&#8217;t use it to strengthen the wording of their thesis statement or help understand obscure terminology in a primary source.</p><p>Even if you wanted to make such an argument, try this thought experiment: is it possible for a student to do research today without bumping into AI? Most people start with a Google search, which relies on knowledge graphs and embeddings (both are forms of non-generative AI) to find and rank results. It also provides generative-AI answers which are sometimes useful but right now, often quite bad. When you go to the library catalogue, if you are at a big <a href="https://news.harvard.edu/gazette/story/2025/03/at-harvard-library-building-a-tool-that-understands/">American school</a> you may already have an AI powered library search engine. If not, when you lookup articles in <a href="https://about.jstor.org/research-tool/">JSTOR</a> or <a href="https://www.ebsco.com/artificial-intelligence">EBSCO</a> (to name but a few article repositories), you&#8217;ll be confronted by AI generated summaries and suggestions for further research. Even if you ignore those, perhaps insisting on open-source repositories, when you do download a journal article to <a href="https://www.adobe.com/acrobat/generative-ai-pdf.html">Adobe</a>, that program will (annoyingly) offer to use generative AI to summarize the document or make notes. Finally, when you open Word or Google Docs, those programs too now offer to write your document for you with AI. The point is that if we wanted to ban AI entirely, it stretches our credibility to pretend it doesn&#8217;t exist and that students won&#8217;t be confronted by it every step of the way. A student in this scenario might logically ask themselves: if AI is so terrible, why is it embedded in all the things I am required to use to complete my degree?</p><h4>A Middle Course</h4><p>Although I am sympathetic to the sentiments expressed by my colleagues, and I admire their willingness to defend our discipline, I don&#8217;t think resistance is a viable option. But I don&#8217;t think prohibition and active resistance are necessary to preserve the values intrinsic to historical inquiry. Instead, what I suggest is that we try to steer a middle course. This means accepting that AI can do some useful things for historians (transcription, translation, summarization, editing, and indexing, among others), but that it is not a substitute for deep knowledge and hard work. Most of all, it means understanding that our students need to leave our classrooms knowing when to use AI, when to avoid it, and how to get the most out of it.</p><p>Conveniently for us, the skills one needs to do this look a lot like the skills history has always offered aspiring lawyers, teachers, politicians, policy analysts and other knowledge workers. Certainly we have to make some tweaks around the edges: going forward, research papers may be less valuable forms of assessment than in-class testing. I also agree with <a href="https://activehistory.ca/blog/2025/06/11/on-generative-ai-in-the-classroom-give-up-give-in-or-stand-up/">Dunsworth</a> that we need to double-down on teaching &#8220;the value of thinking&#8212;laboured, painful, frustrating thinking&#8221;. But why does that have to mean that we can&#8217;t use AI to transcribe a handwritten document we&#8217;ve thoroughly read but want to full-text search? How about reformatting footnotes from Chicago Style into APA for an interdisciplinary journal? Or converting our footnotes into bibliographic format? AI use and deep thinking don&#8217;t have to be mutually exclusive. In fact, done right, the former might save some time for the latter.</p><p>As historians, I think we can choose to help shape how these tools are used in the world&#8212;which means making some compromises&#8212;or we can retreat into a purist position that is likely to make us irrelevant to those discussions. In my view, neither uncritical adoption nor absolute resistance will serve our students well. What they need from us is what we've always provided: the ability to think critically about sources, to construct evidence-based arguments, and to navigate complex information. These skills are more valuable now than ever. Our job is to help students develop them through engagement with the tools they&#8217;ll be expected to use, not by pretending those tools don&#8217;t exist.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://generativehistory.substack.com/p/steering-a-middle-course-on-ai-in?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://generativehistory.substack.com/p/steering-a-middle-course-on-ai-in?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://generativehistory.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Generative History! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Introducing Archive Studio]]></title><description><![CDATA[An experimental AI research tool for exploring the AI transcription and analysis of historical documents]]></description><link>https://generativehistory.substack.com/p/introducing-archive-studio</link><guid isPermaLink="false">https://generativehistory.substack.com/p/introducing-archive-studio</guid><dc:creator><![CDATA[Mark Humphries]]></dc:creator><pubDate>Tue, 13 May 2025 10:02:07 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!DWdU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe45320b8-4a75-4bf8-b27f-d30c0e53bf21_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!DWdU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe45320b8-4a75-4bf8-b27f-d30c0e53bf21_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!DWdU!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe45320b8-4a75-4bf8-b27f-d30c0e53bf21_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!DWdU!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe45320b8-4a75-4bf8-b27f-d30c0e53bf21_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!DWdU!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe45320b8-4a75-4bf8-b27f-d30c0e53bf21_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!DWdU!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe45320b8-4a75-4bf8-b27f-d30c0e53bf21_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!DWdU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe45320b8-4a75-4bf8-b27f-d30c0e53bf21_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e45320b8-4a75-4bf8-b27f-d30c0e53bf21_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:3093331,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://generativehistory.substack.com/i/163442969?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe45320b8-4a75-4bf8-b27f-d30c0e53bf21_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!DWdU!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe45320b8-4a75-4bf8-b27f-d30c0e53bf21_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!DWdU!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe45320b8-4a75-4bf8-b27f-d30c0e53bf21_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!DWdU!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe45320b8-4a75-4bf8-b27f-d30c0e53bf21_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!DWdU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe45320b8-4a75-4bf8-b27f-d30c0e53bf21_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Last fall, we introduced <em>Transcription Pearl</em>, a tool that used AI to transcribe historical, handwritten documents. At that time, the best models were performing better on English language documents than well-known programs like <em>Transkribus</em> but they were still far from useable &#8220;out-of-the-box&#8221;. As Mark wrote about a few weeks ago, though, that all changed with Gemini-2.5-pro: now you can get really accurate transcriptions (better than 95% WER) on the first iteration.</p><p>This brings us to Archive Studio, an open-source, more comprehensive program that attempts to automate some of the more mundane elements of historical document processing. Our goal here is to explore what AI might offer historians, hopefully allowing them to move seamlessly from images through transcription, correction, analysis, and structuring. </p><p>That said, many of the features we&#8217;ve added are experimental and are mainly about exploring the potential&#8212;and problems&#8212;of using AI in research. If people find them useful, that is great and we&#8217;d love to hear about it. But if you run into problems, or the AI generates inaccurate results, that is equally important&#8212;perhaps more so. The goal here is to help us generate information about what AI tools can and cannot do (yet?). Knowledge of both will help us have more robust discussions about how to respond to the technology.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://generativehistory.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://generativehistory.substack.com/subscribe?"><span>Subscribe now</span></a></p><h4><strong>Introducing Archive Studio</strong></h4><p>Archive Studio incorporates the original handwriting recognition and AI-powered correction capabilities available in Transcription Pearl, allowing users to generate initial transcriptions from images or refine existing ones using models from OpenAI, Anthropic, or Google. As we note above, though, we&#8217;ve found that with Gemini-2.5-pro, the correction step may no longer be worth the time. </p><p>You can download and run Archive Studio (which is free and open-source) from Githhub here: <a href="https://github.com/mhumphries2323/Archive_Studio">https://github.com/mhumphries2323/Archive_Studio</a>. Right now you need to run it as a python script. As with Transcription Pearl, we will shortly release a standalone exe file for Windows.</p><p>We&#8217;ve also added some features to the integrated Image Preprocessing Tool which provides essential functions for preparing document images. This separate utility allows users to split two-page spreads, crop unwanted margins, manually or automatically straighten skewed images, and rotate them, ensuring cleaner inputs for the AI functions and potentially reducing processing costs. Batch processing features within this tool help streamline the preparation of large numbers of images. </p><p>The real expansion, however, lies in the new tools designed for working with the transcribed text. Researchers can now employ AI functions to automatically format text according to defined styles, remove awkward line or page breaks, and standardize layouts or dating formats for better readability or use in external databases.</p><h4><strong>Analytical Features</strong></h4><p>Verifying the accuracy of the text and analyzing the data both require identifying key information. Archive Studio includes functions to automatically identify, highlight, and extract the names of people and places mentioned within the text. These aren&#8217;t just static lists, though: a dedicated collation tool manages variations in spelling and can standardize them across the entire project if its desired.</p><p>To further aid historical analysis, especially with large collections, Archive Studio offers a relevance assessment feature. Users can now define specific research criteria, and use an AI model to analyze each document, determining whether it&#8217;s &#8220;Relevant&#8221; or &#8220;Partially Relevant.&#8221; Users can then quickly click through only those documents relevant to their criteria. This potentially allows for efficient filtering and focusing on the most pertinent materials within a large dataset. That said, we have very little information on the accuracy of these AI assessments, so we&#8217;ll look forward to hearing your feedback.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!2SB0!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca82d072-c481-4429-adbd-a5ffb1b51a00_1912x1004.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!2SB0!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca82d072-c481-4429-adbd-a5ffb1b51a00_1912x1004.png 424w, https://substackcdn.com/image/fetch/$s_!2SB0!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca82d072-c481-4429-adbd-a5ffb1b51a00_1912x1004.png 848w, https://substackcdn.com/image/fetch/$s_!2SB0!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca82d072-c481-4429-adbd-a5ffb1b51a00_1912x1004.png 1272w, https://substackcdn.com/image/fetch/$s_!2SB0!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca82d072-c481-4429-adbd-a5ffb1b51a00_1912x1004.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!2SB0!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca82d072-c481-4429-adbd-a5ffb1b51a00_1912x1004.png" width="1456" height="765" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ca82d072-c481-4429-adbd-a5ffb1b51a00_1912x1004.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:765,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:925671,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://generativehistory.substack.com/i/163442969?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca82d072-c481-4429-adbd-a5ffb1b51a00_1912x1004.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!2SB0!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca82d072-c481-4429-adbd-a5ffb1b51a00_1912x1004.png 424w, https://substackcdn.com/image/fetch/$s_!2SB0!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca82d072-c481-4429-adbd-a5ffb1b51a00_1912x1004.png 848w, https://substackcdn.com/image/fetch/$s_!2SB0!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca82d072-c481-4429-adbd-a5ffb1b51a00_1912x1004.png 1272w, https://substackcdn.com/image/fetch/$s_!2SB0!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca82d072-c481-4429-adbd-a5ffb1b51a00_1912x1004.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Archive Studio: Names and Places are automatically highlighted in the text to aid in verification.</figcaption></figure></div><h4><strong>Sequential Data and Document Separation</strong></h4><p>One of the significant challenges we&#8217;ve had in using AI with historical documents is the sequential nature of historical data where key information often needs to be inferred from earlier elements of a lengthy text. Think of sources like diaries or letterbooks which are continuous in nature. For example, a year long fur trade journal may mention the year only on the first page while individual entries might be dated only with the day and omit the month altogether. These journals, although associated with a specific place, were also often written at more than one place over the course of the year as clerks travelled. Authorship too can change from one entry to the next. Without the right tooling, this critical historical context gets lost, resulting in a disembodied diary entry that might simply begin &#8220;Wednesday 2d&#8221;.</p><p>Archive Studio handles this problem through a multi-stage document separation process. First, an AI function analyzes the text, using customizable presets designed for different document types like &#8220;Letterbook&#8221; or &#8220;Diary,&#8221; to identify break points like the start of a new letter or entry. It inserts editable markers, joining text across page breaks where appropriate. Subsequently, after verifying their accuracy, the researcher can apply the separations. The result is that in the case of a 20-page post diary, instead of navigating page by page, the user is then able to navigate between, say, 300 discrete entries.</p><h4><strong>Customizable Functions</strong></h4><p>The behavior of all the AI functions is governed by customizable presets found in the Settings window. This allows researchers to select specific AI models, adjust parameters like temperature, and, critically, change the prompts and instructions given to the AI for each task from transcription to metadata extraction. Settings can be saved, loaded, and exported for sharing or backup.</p><p>We&#8217;ve found that the development of good, project-specific customized functions is essential. Thinking of our fur trade post journal example again, these documents often contain a mix of diary entries, letters, and tabular data like financial accounts, lists of supplies, and prices. To handle this variation, you can develop customized instructions which tell the AI what to do with each type of document. This may sound complicated, but it is actually quite simple.</p><p>In writing custom instructions, provide headers, explain what you want the AI to do and what you don&#8217;t want it to do, and give clear examples. Your instructions (which are technically just the prompts sent to the LLM) can be as long as you want. We&#8217;ve found that it&#8217;s best to describe your overall goal, the specific task, and requirements in the &#8220;General Instructions&#8221; (which is sent to the LLM as a system message). You can keep the &#8220;Specific Instructions&#8221; more brief. Generally, a lower temperature setting (that is, closer to 0 than to 1) is better for consistency and accuracy&#8212;on many tasks we suggest 0.2 to 0.3.</p><p>You can also use the settings to configure whether the LLM is sent images of the text on the current page or images of the previous or following pages (as well as how many of those images to send). This is useful when you want the AI to be able to read ahead (or behind) the current page of the document.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!kL4t!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F800c6e55-f42e-4176-af68-46f77c0315c9_1198x909.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!kL4t!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F800c6e55-f42e-4176-af68-46f77c0315c9_1198x909.png 424w, https://substackcdn.com/image/fetch/$s_!kL4t!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F800c6e55-f42e-4176-af68-46f77c0315c9_1198x909.png 848w, https://substackcdn.com/image/fetch/$s_!kL4t!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F800c6e55-f42e-4176-af68-46f77c0315c9_1198x909.png 1272w, https://substackcdn.com/image/fetch/$s_!kL4t!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F800c6e55-f42e-4176-af68-46f77c0315c9_1198x909.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!kL4t!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F800c6e55-f42e-4176-af68-46f77c0315c9_1198x909.png" width="1198" height="909" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/800c6e55-f42e-4176-af68-46f77c0315c9_1198x909.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:909,&quot;width&quot;:1198,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:70735,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://generativehistory.substack.com/i/163442969?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F800c6e55-f42e-4176-af68-46f77c0315c9_1198x909.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!kL4t!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F800c6e55-f42e-4176-af68-46f77c0315c9_1198x909.png 424w, https://substackcdn.com/image/fetch/$s_!kL4t!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F800c6e55-f42e-4176-af68-46f77c0315c9_1198x909.png 848w, https://substackcdn.com/image/fetch/$s_!kL4t!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F800c6e55-f42e-4176-af68-46f77c0315c9_1198x909.png 1272w, https://substackcdn.com/image/fetch/$s_!kL4t!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F800c6e55-f42e-4176-af68-46f77c0315c9_1198x909.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Settings window where users can create their own customized functions</figcaption></figure></div><h4><strong>Metadata Generation and Structured Data</strong></h4><p>Transforming transcribed text into usable data requires extracting key pieces of structured information. Archive Studio attempts to automate this process, again relying on user-defined &#8220;Metadata Presets&#8221; which are configured within the Settings window. These presets act as blueprints, specifying not only <em>which</em> metadata fields the user wants to extract or create (such as author, date, place of creation, recipient, or a concise summary of the document), but also providing the precise instructions, or prompts, that guide the AI on <em>how</em> to locate or infer this information from the document&#8217;s text. </p><p>Again, this is an experimental feature and we include it because it is important for us, as historians, to understand how effective these models are at these types of tasks in order to make decisions about how to respond to their availability. If they are useful and accurate&#8212;or at least show promise&#8212;that is important to know. But if the models fail on these common tasks&#8212;or make significant numbers of errors&#8212;they may prove to be quite harmful to the research process. Either way, we need to know.</p><p>Once documents have been processed and analyzed within Archive Studio, you can export the results in a variety of formats like PDFs or text documents, or a CSV (Comma Separated Values) spreadsheet file which you can use in programs like Excel or SPSS. The CSV is especially important, as it is where the metadata and other extracted information can be stored and accessed.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!uSlX!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F924163e1-f620-4408-b202-709cbb31fece_1648x611.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!uSlX!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F924163e1-f620-4408-b202-709cbb31fece_1648x611.png 424w, https://substackcdn.com/image/fetch/$s_!uSlX!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F924163e1-f620-4408-b202-709cbb31fece_1648x611.png 848w, https://substackcdn.com/image/fetch/$s_!uSlX!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F924163e1-f620-4408-b202-709cbb31fece_1648x611.png 1272w, https://substackcdn.com/image/fetch/$s_!uSlX!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F924163e1-f620-4408-b202-709cbb31fece_1648x611.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!uSlX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F924163e1-f620-4408-b202-709cbb31fece_1648x611.png" width="1456" height="540" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/924163e1-f620-4408-b202-709cbb31fece_1648x611.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:540,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:116427,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://generativehistory.substack.com/i/163442969?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F924163e1-f620-4408-b202-709cbb31fece_1648x611.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!uSlX!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F924163e1-f620-4408-b202-709cbb31fece_1648x611.png 424w, https://substackcdn.com/image/fetch/$s_!uSlX!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F924163e1-f620-4408-b202-709cbb31fece_1648x611.png 848w, https://substackcdn.com/image/fetch/$s_!uSlX!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F924163e1-f620-4408-b202-709cbb31fece_1648x611.png 1272w, https://substackcdn.com/image/fetch/$s_!uSlX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F924163e1-f620-4408-b202-709cbb31fece_1648x611.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Users can automatically generate metadata and export to a CSV spreadsheet</figcaption></figure></div><h4><strong>Final Thoughts</strong></h4><p>Archive Studio is far from a finished product, it is a prototype&#8212;and a test case at that. While we have published on the accuracy of LLM translations, which offer a clear advantage over other automated options, we know surprisingly little about the accuracy of LLMs on analytical tasks. As a result, the use of the analysis features embedded in Archive Studio carry clear risks and, in some instances, may create more methodological problems than they solve. In many cases, the accuracy and usefulness of the outputs will depend on the type of instructions and criteria the user gives the LLM and the performance of the individual model assigned to the task. But in other cases, LLMs might not be up to the task you want them to perform.</p><p>We simply don&#8217;t know how good these models are at the types of tasks historians might ask them to perform and we hope you will report your results in the comments below. From our experience, testing shows that sometimes LLMs will miss a name or placename and may not always identify all the relevant documents in a given set. Its uncommon but not exactly rare. For this reason, always check and verify any results&#8230;and please report your findings!</p><p>Ultimately, <em>Archive Studio </em>is meant to be a contribution to the evolving conversation about how historians and archivists can productively engage with AI. It aims to provide a tangible means for researchers to experiment with these tools and evaluate their potential usefulness (or not).</p><p>Here it is important to remember that at present, most people are still using the types of ChatBot interfaces that don&#8217;t really provide a fair test of an LLM&#8217;s actual capabilities. So many of the issues people experience using programs like ChatGPT come from the tooling around the models&#8212;that is how information is passed to the model&#8212;rather than from the models themselves.</p><p>It may well be the case that LLMs will fail at some of the most basic things we do as historians, even when given the right tooling. If that is the case, we will need to identify and push back against the potential problems created by automated AI analysis&#8212;which is coming from private companies very soon. But if the models prove to be largely capable, that will necessitate other conversations. In truth, LLMs are likely to be capable in some areas, less so in others. Either way, we need to understand what these things can and cannot do to move forward.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://generativehistory.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://generativehistory.substack.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><div class="captioned-button-wrap" data-attrs="{&quot;url&quot;:&quot;https://generativehistory.substack.com/p/introducing-archive-studio?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="CaptionedButtonToDOM"><div class="preamble"><p class="cta-caption">Thanks for reading Generative History! This post is public so feel free to share it.</p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://generativehistory.substack.com/p/introducing-archive-studio?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://generativehistory.substack.com/p/introducing-archive-studio?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p></div><p></p>]]></content:encoded></item><item><title><![CDATA[Stochastic Canaries in the Coalmine]]></title><description><![CDATA[Without good information, it&#8217;s hard to prepare for the future.]]></description><link>https://generativehistory.substack.com/p/stochastic-canaries-in-the-coalmine</link><guid isPermaLink="false">https://generativehistory.substack.com/p/stochastic-canaries-in-the-coalmine</guid><dc:creator><![CDATA[Mark Humphries]]></dc:creator><pubDate>Thu, 06 Mar 2025 15:28:37 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!_FGk!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a550bf3-fae8-493c-86f4-a4a02511c85e_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!_FGk!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a550bf3-fae8-493c-86f4-a4a02511c85e_1024x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!_FGk!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a550bf3-fae8-493c-86f4-a4a02511c85e_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!_FGk!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a550bf3-fae8-493c-86f4-a4a02511c85e_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!_FGk!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a550bf3-fae8-493c-86f4-a4a02511c85e_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!_FGk!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a550bf3-fae8-493c-86f4-a4a02511c85e_1024x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!_FGk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a550bf3-fae8-493c-86f4-a4a02511c85e_1024x1024.png" width="1024" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5a550bf3-fae8-493c-86f4-a4a02511c85e_1024x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1849164,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://generativehistory.substack.com/i/158517081?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a550bf3-fae8-493c-86f4-a4a02511c85e_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!_FGk!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a550bf3-fae8-493c-86f4-a4a02511c85e_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!_FGk!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a550bf3-fae8-493c-86f4-a4a02511c85e_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!_FGk!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a550bf3-fae8-493c-86f4-a4a02511c85e_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!_FGk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a550bf3-fae8-493c-86f4-a4a02511c85e_1024x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://generativehistory.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://generativehistory.substack.com/subscribe?"><span>Subscribe now</span></a></p><p>How good are LLMs <em>right now</em> at historical research? That&#8217;s a surprisingly difficult question to answer. The benchmarks we&#8217;ve used for the past couple of years are getting less meaningful as models all become A+ students in science, math, and coding. At the same time, we don&#8217;t have any quantifiable way to measure how they perform in the social sciences and humanities on core tasks like qualitative analysis, argument, use of evidence, and prose style, things that don&#8217;t have clear right and wrong answers. Without good information, it&#8217;s hard to prepare for the future.</p><p>This is one of the reasons that the discourse around AI is becoming so disorienting and polarized. Some <a href="https://garymarcus.substack.com/p/ezra-kleins-new-take-on-agi-and-why">pundits</a>, academics, and <a href="https://www.forbes.com/sites/johnrau/2025/01/03/is-ai-a-boom-bubble-or-con-heres-what-the-evidence-suggests/">many in the media</a> want to interpret plateauing benchmarks as a sign that things are slowing down and AI is a bubble. Much of this is wishful thinking. On the other side, though, the accelerationist fanboys on X are equally disconnected from reality, often naively arguing that we&#8217;re about to achieve a form of Artificial General Intelligence (AGI) and that this will usher in a true utopia.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://generativehistory.substack.com/p/stochastic-canaries-in-the-coalmine?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://generativehistory.substack.com/p/stochastic-canaries-in-the-coalmine?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><h4>A Changing Conversation</h4><p>What&#8217;s becoming clear, though, is that at least the first part of the accelerationist argument is almost certainly closer to the truth than the idea that AI is a bubble. Many <a href="https://darioamodei.com/machines-of-loving-grace">thoughtful</a>, <a href="https://betakit.com/new-turing-award-winner-richard-sutton-calls-doomers-out-of-line-talks-path-to-human-like-ai/">knowledgeable </a>people have started to think that something transformational is happening. <em>New York Times</em> columnist and podcaster Ezra Klein recently summarized this view in a piece titled &#8220;<a href="https://www.nytimes.com/2025/03/04/opinion/ezra-klein-podcast-ben-buchanan.html">The Government Knows AGI is Coming</a>.&#8221; As he writes, most of the serious people working in the AI labs and US government agencies believe that something like AGI is imminent. &#8220;They believe it because of the products they&#8217;re releasing right now and what they&#8217;re seeing inside the places they work,&#8221; Klein writes. &#8220;And I think they&#8217;re right. If you&#8217;ve been telling yourself this isn&#8217;t coming, I really think you need to question that.&#8221; </p><p>I couldn&#8217;t agree more. For many of us who&#8217;ve been nervously watching from the sidelines&#8212;Klein included&#8212;the turning point came with Deep Research. <a href="https://generativehistory.substack.com/p/is-this-the-last-generation-of-historians">As I wrote about a few weeks ago</a>, this is an AI system, not just a model. It&#8217;s capable of doing exceptional things because it combines a new and powerful LLM trained to reason with the tooling to allow it access the information and materials necessary to do complex intellectual tasks. We are, in effect, <a href="https://www.oneusefulthing.org/p/the-end-of-search-the-beginning-of">unleashing highly capable models with better tooling</a>&#8212;from introducing reasoning by scaling test-time compute and providing models the ability to call tools and read files. The argument is that there&#8217;s now a clear road to AI systems that are <em>better than most humans at most tasks one could do on a computer</em>. Whether we call that AGI or not doesn&#8217;t really matter.</p><p>What evidence leads people to make such claims? Given how hard it&#8217;s getting to properly benchmark AI systems&#8212;and how quickly things are moving&#8212;a lot of the discussion is based on vibes. What I want to do is provide a couple of concrete examples that have convinced me that we&#8217;re crossing an important threshold.</p><h4>The Paradox of Saturated Benchmarks</h4><p>One of the problems we face cutting through the noise is that the <a href="https://www.vox.com/future-perfect/394336/artificial-intelligence-openai-o3-benchmarks-agi">models have effectively saturated standard AI benchmarks,</a> meaning they score so high that comparisons between them cease to be useful. LLMs are, in effect, outpacing our ability to accurately measure their capabilities. In practical terms, this presents a paradox: as models get better and the number of tasks at which they fail decreases, the number of users available who can benefit from subsequent improvements diminishes.</p><p>In effect, we&#8217;ve already passed the point where for many use cases, the existing base models are already good enough and cannot really get any better. For example, <a href="https://www.anthropic.com/news/the-anthropic-economic-index?utm_source=www.therundown.ai&amp;utm_medium=newsletter&amp;utm_campaign=elon-s-97b-openai-offer&amp;_bhlid=c595d10ba7838789de6243d0af84e54335efee11">a recent study by Anthropic</a> found that 8-12% of all workplace queries fielded by its Claude models were related to office and administrative support. But if Claude can already extract all the names and phone numbers from a series of documents and put them into a table&#8212;and does so correctly&#8212;a bigger model can&#8217;t do it better. Unless you are mainly using the latest model to do something the previous version couldn&#8217;t do, you&#8217;re unlikely to notice an upgrade.</p><p>My intuition is that the newer models are actually much better than the benchmarks suggest, it&#8217;s just getting harder to quantify the capabilities. Take historical handwriting text recognition (HTR) as an example. On the surface, deciphering handwriting appears to be a relatively straight-forward vision-based task, complicated only by the fact that handwriting styles vary enormously from one person to the next. But vision only gets you so far with automatic HTR, specifically to around 80 or 90% accuracy. In the word of HTR, this means word error rates (WERs) of 10-20% and character error rates (CERs) of 5-15%. </p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!4b45!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc19f5faa-b458-4f22-ab45-66476151f168_1982x858.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!4b45!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc19f5faa-b458-4f22-ab45-66476151f168_1982x858.png 424w, https://substackcdn.com/image/fetch/$s_!4b45!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc19f5faa-b458-4f22-ab45-66476151f168_1982x858.png 848w, https://substackcdn.com/image/fetch/$s_!4b45!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc19f5faa-b458-4f22-ab45-66476151f168_1982x858.png 1272w, https://substackcdn.com/image/fetch/$s_!4b45!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc19f5faa-b458-4f22-ab45-66476151f168_1982x858.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!4b45!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc19f5faa-b458-4f22-ab45-66476151f168_1982x858.png" width="1456" height="630" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c19f5faa-b458-4f22-ab45-66476151f168_1982x858.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:630,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:127586,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://generativehistory.substack.com/i/158517081?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc19f5faa-b458-4f22-ab45-66476151f168_1982x858.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!4b45!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc19f5faa-b458-4f22-ab45-66476151f168_1982x858.png 424w, https://substackcdn.com/image/fetch/$s_!4b45!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc19f5faa-b458-4f22-ab45-66476151f168_1982x858.png 848w, https://substackcdn.com/image/fetch/$s_!4b45!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc19f5faa-b458-4f22-ab45-66476151f168_1982x858.png 1272w, https://substackcdn.com/image/fetch/$s_!4b45!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc19f5faa-b458-4f22-ab45-66476151f168_1982x858.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Even with perfect vision&#8212;human or machine&#8212;there&#8217;s still lots of ambiguity in handwritten texts. In the English and French language documents I work on, capitalization, punctuation, ink smudges, deletions, insertions, spelling differences, and grammar all complicate the process. That&#8217;s where reasoning and cultural knowledge comes in: you need to know something about the subject and be able to make connections between words, characters, and phrases across time and space. It&#8217;s why we can&#8217;t decipher text devoid of the context in which they were created. Even then, some things are still a matter of interpretation and some scrawl just can&#8217;t be deciphered. All this means that while WERs of 4-10% are the norm for humans, 1-2% is probably the floor.</p><p>Last fall, a team led by Dr. Lianne Leddy and I found that LLMs were beginning to approach human levels of accuracy. Since then, Google, Anthropic, and OpenAI have all released new versions of their models. Both the Google and OpenAI models represent a generational shift in size. The Anthropic model makes more of an incremental change. As you&#8217;ll see in the chart below, there&#8217;s been dramatic improvements in accuracy for the two next generation of models, an average of 50%. Moreover, Gemini 2.0 Pro&#8212;which has not received the attention it deserves as a model&#8212;is now scoring nearly 3 times better than Transkribus. If you ignore errors in punctuation, capitalization, and historical spelling corrections, Gemini Pro correctly transcribes 97% of characters and around 95% of words. Gpt-4.5 is not far off. These are median human levels of accuracy.</p><p>A key point is that while a 50% reduction in error rates is significant, it represents a much smaller overall jump in overall accuracy, from about 87% to 95%. Yet what these numbers obscure is that somewhere between, we crossed a really significant qualitative threshold.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!6Mqi!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F819d2f17-201d-4363-b056-ef0da3a079f2_1088x1107.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!6Mqi!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F819d2f17-201d-4363-b056-ef0da3a079f2_1088x1107.png 424w, https://substackcdn.com/image/fetch/$s_!6Mqi!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F819d2f17-201d-4363-b056-ef0da3a079f2_1088x1107.png 848w, https://substackcdn.com/image/fetch/$s_!6Mqi!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F819d2f17-201d-4363-b056-ef0da3a079f2_1088x1107.png 1272w, https://substackcdn.com/image/fetch/$s_!6Mqi!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F819d2f17-201d-4363-b056-ef0da3a079f2_1088x1107.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!6Mqi!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F819d2f17-201d-4363-b056-ef0da3a079f2_1088x1107.png" width="1088" height="1107" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/819d2f17-201d-4363-b056-ef0da3a079f2_1088x1107.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1107,&quot;width&quot;:1088,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:93666,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://generativehistory.substack.com/i/158517081?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F819d2f17-201d-4363-b056-ef0da3a079f2_1088x1107.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!6Mqi!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F819d2f17-201d-4363-b056-ef0da3a079f2_1088x1107.png 424w, https://substackcdn.com/image/fetch/$s_!6Mqi!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F819d2f17-201d-4363-b056-ef0da3a079f2_1088x1107.png 848w, https://substackcdn.com/image/fetch/$s_!6Mqi!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F819d2f17-201d-4363-b056-ef0da3a079f2_1088x1107.png 1272w, https://substackcdn.com/image/fetch/$s_!6Mqi!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F819d2f17-201d-4363-b056-ef0da3a079f2_1088x1107.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>When you compare transcriptions from GPT-4o and GPT-4.5 against the original document, both texts are readable, generally correct, and without any obvious OCR type errors. Yet some of the GPT-4o errors are significant in that they distort the meaning of the text to the point that you can&#8217;t really trust the transcript at all. For example, in one of our tests we asked GPT-4o to transcribe a page from a 19<sup>th</sup> century lease. While many of the errors represent &#8220;minor&#8221; changes in punctuation, capitalization, and formatting, it wrongly writes that the lease was for two years rather than four; misses the inclusion of a garden; and hallucinates a short bit of text that replaces the fact that document was binding on the lessee&#8217;s heirs. Those are all essential elements of the contract, which is fatal.</p><p>At first glance, the results from GPT-4.5 look much the same. But within the 8% accuracy GPT-4.5 gained, it gets all the things right that its predecessor got wrong. Its biggest mistake is that it misread &#8220;buildings&#8221; as &#8220;outhouses&#8221;, but in the context of the contract the two terms are synonymous. This is an important qualitative shift in HTR reliability but I don&#8217;t think it just comes from better vision. I think the model is starting to read documents in a more complex, human-like way. How we would benchmark that, I don&#8217;t know.</p><h4>Do Big Models Have a Smell?</h4><p>My intuition is that it the same changes which caused GPT-4.5 to improve on handwriting signals a broader shift in the complexity in the model&#8217;s overall capabilities. Where this becomes most evident to me is in the model&#8217;s ability to work with, reason through, and write about historical documents. Here benchmarking is a familiar problem for historians because it&#8217;s a lot like grading papers.</p><p>I&#8217;ve been teaching for nearly 20 years at the post-secondary level and while I think I&#8217;ve developed a pretty good feel for the difference between an A+ and an A- paper, its not always an easy distinction to explain to students. To my mind, an A- paper will typically contain most (if not all) of the elements of an A+ paper, but it lacks something important, mainly in the execution. While the A- paper might have a few more technical errors than an A+, that&#8217;s usually a symptom of the fact that on the whole it&#8217;s not as polished. Rare is the typo filled A+ paper with a cogent, engaging thesis. In an A- paper, the argument might not be as sophisticated. The evidence might need some more unpacking. It&#8217;s just not as well written. Accuracy is part of it, but it&#8217;s a symptom of something larger.</p><p>Recently, I&#8217;ve started to see these same types of qualitative differences emerge from the outputs of GPT-4.5 and Sonnet-3.7. One of my oldest LLM tests is to give a model one of my first-year assignments, a historical sandbox exercise on Britain&#8217;s surrender of St. John&#8217;s, Newfoundland to the French in June 1762. I use this obscure event because it&#8217;s relatively unknown and there are only six surviving primary sources that describe it. These include perspectives from both French and British officers, enlisted men, townsfolk, merchants, and sailors&#8212;most importantly: they conflict with one another in delightful ways. Students (or the LLM) is then asked to read three secondary sources for context and to develop a &#8220;theory of the crime&#8221; that fits the conflicting evidence. It&#8217;s an engaging assignment with no single right answer.</p><h4>Writing From Primary Documents</h4><p>Back in the winter of 2023, GPT-3.5 could summarize the contents of the documents but lacked a coherent argument. It also didn&#8217;t address the conflicts between the documents very well, giving a C+ to B- range answer. In <a href="https://openai.com/index/hello-gpt-4o/">spring 2024, Gpt-4o</a> did much better. The response was 1,500 words, but the first paragraph gives an idea:</p><blockquote><p><em>&#8220;The surrender of the garrison at Fort William in St. John's, Newfoundland, on 28 June 1762, is a complex historical event that can be understood through a nuanced reading of the available primary and secondary sources. These documents provide various perspectives, from the ordinary soldiers and garrison officers to the French invaders and contemporary newspapers, each with its own interpretation of events. By examining these sources, we can construct an internally coherent narrative that explains why the British garrison surrendered to the French forces.&#8221;</em></p></blockquote><p>That is the beginning of a decent but ho-hum, unremarkable answer, an &#8220;all of the above&#8221; sort of argument. It would typically receive a B to B+, maybe an A- depending on execution.</p><p>As with handwriting, next-generation models have changed the game in subtle but important ways. They are more willing to make clear arguments, to cite specific examples, and to use evidence in complex ways. Here is gpt-4.5&#8217;s opening paragraph with the same prompt:</p><blockquote><p><em>&#8220;The surrender of the garrison of Fort William at St. John's, Newfoundland, on 28 June 1762, was the result of multiple factors that converged to create a situation where resistance was untenable from the perspective of the fort&#8217;s commander, Captain Walter Ross. A close analysis of primary and secondary sources reveals the fragility of British authority in Newfoundland, the strained economic and social conditions of the local population, strategic misjudgments by the French commander Ternay, and a lack of decisive leadership within the British garrison. Furthermore, the varied accounts from soldiers, officers, French commanders, and newspapers underscore the complexity of historical interpretation, highlighting the historian's task in critically assessing and weighing conflicting evidence to form a coherent and nuanced argument.&#8221;</em></p></blockquote><p>That is the beginning of a solid A answer. GPT-4.5 acknowledges the complexity of the situation but makes a clear argument: Ross made the decision to surrender because he came to believe the situation was untenable. In contrast, the first example from GPT-4o is less complex, more ambiguous, and lacks a clear thesis. It also doesn&#8217;t engage with an important but understated part of the assignment: the fact that the overall point is to get students to understand that historians need to learn to weigh conflicting evidence. GPT-4.5 clearly gets this.</p><p>After trying this out, I suspected that the assignment itself might be constraining the model. Afterall, there are only so many ways to write a 1,500 word analysis while still meeting all the requirements. So I tried something different. Bigger models evidence a remarkable shift that we don&#8217;t see from smaller models when we make the assignment more open-ended. If you ask a next generation model to use the sources provided to write a chapter in a book on Newfoundland in the Seven Years War, you get something more interesting. Here are the opening paragraphs from Sonnet-3.7&#8217;s take on that task:</p><blockquote><p><em>&#8220;On the morning of June 24, 1762, four French warships&#8212;Le Robuste (74 guns), L'&#201;veill&#233; (64 guns), La Licorne (26 guns), and La Garonne (26 guns)&#8212;anchored under English colors in the sheltered waters of Bay Bulls, about thirty-five kilometers southeast of St. John's. This small fleet, commanded by the Chevalier de Ternay d'Arsac, had managed a daring journey across the Atlantic, evading the Royal Navy blockades that had kept most French naval forces confined to port. As their anchors dropped into the seabed, the French colors suddenly replaced the deceptive English flag, revealing their true identity to the stunned residents of the small fishing community. Soon, landing craft began lowering into the calm waters. Eight hundred and seventy French regular troops under the command of the Comte d'Haussonville, mostly grenadiers from five regiments, began disembarking onto Newfoundland's shores.</em></p><p><em>The French had arrived at a crucial moment in the Seven Years' War. After years of defeats that had cost France most of its colonial empire, the duc de Choiseul, French Minister of Marine, had devised an audacious strategy to strike at the heart of British economic power: the Newfoundland fishery. By attacking this vital industry, France hoped to gain leverage in ongoing peace negotiations while demonstrating that, despite its naval setbacks, it could still project power across the Atlantic. What initially appeared to be a routine summer day on the easternmost edge of Britain's North American possessions was about to become a pivotal moment in the long struggle for control of Newfoundland.</em></p><p><em>The fall of St. John's in June 1762 revealed both the strategic importance of Newfoundland in the Atlantic world and the fragility of imperial claims in contested maritime spaces. The ease with which the French seized the main British settlement exposed not only military unpreparedness but also the complex social reality of Newfoundland&#8212;a place where economic interests often outweighed imperial loyalty, where ethnic and religious divisions created opportunities for exploitation, and where authority rested as much with merchant capital as with formal imperial structures. The story of St. John's surrender illuminates how Newfoundland was never simply a British colony but rather a "space of power" where influence was contested and where multiple European and indigenous interests overlapped and competed.&#8221;</em></p></blockquote><p>If those paragraphs crossed my desk in peer-review, I don&#8217;t think I would have any inkling that they were written by an LLM. What is most striking, though, is the way in which Sonnet used facts taken from a number of different documents to construct a coherent, narrative anecdote that successfully foreshadows the nuanced argument which follows. This is not an easy skill to master nor one that is easy to teach. It&#8217;s also not one that previous models show. The answers that GPT-4o gives on the assignment version of the question and the book version are remarkably similar.</p><h4>Conclusion</h4><p>When I worry that I might be part of <a href="https://generativehistory.substack.com/p/is-this-the-last-generation-of-historians">the last generation of historians</a>, I don&#8217;t mean to imply that humans will cease to write or care about history. What I mean is that I think I am part of the last generation in which history will be a uniquely human endeavor. As machines start to write history alongside us&#8212;for cash strapped museums, local historical societies, archives, geneaologists, companies, and legal firms&#8212;there will be an important shift not only in how we do history, but how history and human historians are perceived by others.</p><p>To be clear, I think we will continue to bring something to the table that machines lack, especially in areas where cultural context and perspective are vitally important. But we need to start thinking about how we are going to co-exist with AI historians and how we can use the tools AI systems provide to our advantage. Otherwise we risk being labelled as irrelevant. That may not matter to some, but its something we need to take seriously as historians and scholars.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://generativehistory.substack.com/p/stochastic-canaries-in-the-coalmine?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://generativehistory.substack.com/p/stochastic-canaries-in-the-coalmine?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://generativehistory.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Generative History! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[Is this the Last Generation of Historians?]]></title><description><![CDATA[OpenAI&#8217;s Deep Research is a strange and shockingly powerful agentic AI research assistant that offers a clear glimpse of the future. But where do we fit in?]]></description><link>https://generativehistory.substack.com/p/is-this-the-last-generation-of-historians</link><guid isPermaLink="false">https://generativehistory.substack.com/p/is-this-the-last-generation-of-historians</guid><dc:creator><![CDATA[Mark Humphries]]></dc:creator><pubDate>Fri, 07 Feb 2025 22:35:59 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!yyeS!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ada8333-4065-4e9f-8c30-650a5eb6785a_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!yyeS!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ada8333-4065-4e9f-8c30-650a5eb6785a_1024x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!yyeS!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ada8333-4065-4e9f-8c30-650a5eb6785a_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!yyeS!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ada8333-4065-4e9f-8c30-650a5eb6785a_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!yyeS!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ada8333-4065-4e9f-8c30-650a5eb6785a_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!yyeS!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ada8333-4065-4e9f-8c30-650a5eb6785a_1024x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!yyeS!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ada8333-4065-4e9f-8c30-650a5eb6785a_1024x1024.png" width="1024" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7ada8333-4065-4e9f-8c30-650a5eb6785a_1024x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!yyeS!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ada8333-4065-4e9f-8c30-650a5eb6785a_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!yyeS!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ada8333-4065-4e9f-8c30-650a5eb6785a_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!yyeS!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ada8333-4065-4e9f-8c30-650a5eb6785a_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!yyeS!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ada8333-4065-4e9f-8c30-650a5eb6785a_1024x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Last weekend, OpenAI released its latest reasoning model called Deep Research. Reasoners like o1, DeepSeek, and o3-mini are different from conventional chatbots in that they spend time planning, exploring possible answers, and then refining their &#8220;thinking&#8221; before responding. The idea is to make their outputs more comprehensive and reliable.</p><p>Deep Research is strange and new. It&#8217;s agentic, meaning it can autonomously conduct its own research to answer high level questions and the results are often awe-inspiring. Most would pass muster in any historical research firm or PhD level course. But right now, under specific circumstances, it can also hallucinate. My sense, though, is that its shortcomings are related to the way the model was launched, rather than any inherent limitations with the model itself. More on that in a bit.</p><p>Historians&#8212;and academics broadly&#8212;need to pay attention: Deep Research is the beginning of something new. For the last couple of years, people have been talking about a future in which AI will begin to do the things we value at an expert level. My intuition is that this is the actual beginning of that future. Until now, many of us have tried to defer the really difficult questions by ignoring AI or pretending the obvious isn&#8217;t happening. I don&#8217;t think that&#8217;s going to be possible any longer.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://generativehistory.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://generativehistory.substack.com/subscribe?"><span>Subscribe now</span></a></p><h3><strong>Deep Research</strong></h3><p>The most important thing to know about Deep Research is that it is one of the first agents, meaning it can autonomously go out and find reliable sources to solve a given research problem. <a href="https://openai.com/index/introducing-deep-research/">OpenAI describes Deep Research</a> as &#8220;powered by a version of the upcoming OpenAI o3 model that&#8217;s optimized for web browsing and data analysis, it leverages reasoning to search, interpret, and analyze massive amounts of text, images, and PDFs on the internet, pivoting as needed in reaction to information it encounters.&#8221; As <a href="https://www.oneusefulthing.org/p/the-end-of-search-the-beginning-of">Ethan Mollick writes</a>, it&#8217;s &#8220;the end of search and the beginning of research.&#8221;</p><p>Another key difference from conventional models is that Deep Research is not designed for back-and-forth chat: it is a fire and forget model. Users begin by posing a research question or problem, the model then asks some clarifying questions of its own, before going off to do some actual research for anywhere between 5 and 30 minutes. When it&#8217;s done, it notifies the user that the answer&#8217;s ready.</p><p>OpenAI says that Deep Research is trained to select solid, reputable sources, to evaluate them thoroughly, and to cite specific text (not just provide general links as other models do). In the near future, they also say they plan to announce deals that will give the model access to paywalled sources, presumably including academic journal repositories and other data. Beyond this, its specific capabilities and limitations remain murky. It's currently available only to users with a Pro account, although OpenAI plans to bring it to Plus tier users at some point in the future.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!KtO-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0c7a34b-dcf7-42ff-a0ac-a7bfe760814e_666x615.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!KtO-!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0c7a34b-dcf7-42ff-a0ac-a7bfe760814e_666x615.png 424w, https://substackcdn.com/image/fetch/$s_!KtO-!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0c7a34b-dcf7-42ff-a0ac-a7bfe760814e_666x615.png 848w, https://substackcdn.com/image/fetch/$s_!KtO-!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0c7a34b-dcf7-42ff-a0ac-a7bfe760814e_666x615.png 1272w, https://substackcdn.com/image/fetch/$s_!KtO-!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0c7a34b-dcf7-42ff-a0ac-a7bfe760814e_666x615.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!KtO-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0c7a34b-dcf7-42ff-a0ac-a7bfe760814e_666x615.png" width="666" height="615" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a0c7a34b-dcf7-42ff-a0ac-a7bfe760814e_666x615.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:615,&quot;width&quot;:666,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:36088,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!KtO-!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0c7a34b-dcf7-42ff-a0ac-a7bfe760814e_666x615.png 424w, https://substackcdn.com/image/fetch/$s_!KtO-!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0c7a34b-dcf7-42ff-a0ac-a7bfe760814e_666x615.png 848w, https://substackcdn.com/image/fetch/$s_!KtO-!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0c7a34b-dcf7-42ff-a0ac-a7bfe760814e_666x615.png 1272w, https://substackcdn.com/image/fetch/$s_!KtO-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0c7a34b-dcf7-42ff-a0ac-a7bfe760814e_666x615.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">The Deep Research feature is available to ChatGPT Pro subscribers and is enabled with the &#8220;Deep Research&#8221; button above. Be aware that this is different from Deepseek, Deepmind, or Gemini Deep Research..ugh.</figcaption></figure></div><p></p><h3><strong>First Impressions</strong></h3><p>Using Deep Research is unsettling, pure and simple. There is something distinctly non-LLM&#8212;human in fact&#8212;about its prose. Its use of language is more natural, nuanced, and sophisticated than conventional models and it can be disarmingly informal but still professional. It&#8217;s a mature style: gone are the hamburger essays and purple prose that GPT-4o typically produces. Deep Research genuinely writes like a good PhD student&#8212;or minted scholar. It&#8217;s a noticeable change and you should click the links below to read some examples.</p><p>One of my first tests was to ask Deep Research for <a href="https://chatgpt.com/share/67a62149-528c-8012-aa7d-5760f1692cfe">a historiographical analysis of the evolution of fur trade historiography</a>, focusing on comparing and contrasting Canadian and American approaches. It&#8217;s response is very good. But consider that right now it can only access full-text sources that aren&#8217;t paywalled, which is why the bibliography is so limited. Even so, it included lots of scholars it could not actually read&#8230;and did a good job with their works. Think forward to what this will look like once OpenAI actually allows users to upload files and Deep Research can access journal repositories and e-library resources on its own.</p><h3><strong>An AI Research Assistant</strong></h3><p>Deep Research is designed to do internet-based research&#8212;it is literally a research assistant that goes off and does a task while you do something else. So next <a href="https://chatgpt.com/share/67a62100-d9c4-8012-afc5-e4685eaeba02">I asked it to compile a list of archival primary sources</a>, specifically of unpublished letters written by Alexander Henry. This is not an easy task for a computer as it involves identifying likely archives, navigating often archaic websites, and then successfully using a catalog search. I watched in amazement as it worked its way through LAC&#8217;s catalogue all on its own&#8212;not an easy feat for a human these days&#8212;and identified several relevant collections. It did the same across 21 archives, including ArchiveGrid, and came back with a reasonably good list that mirrored my own initial survey from a few years ago. Not perfect, but exactly what I would expect from a human RA. Three years ago, most people thought this type of AI assistance was decades off.</p><p>It is telling that this is the first model I&#8217;ve accidentally started to anthropomorphize. And it&#8217;s not just the writing that&#8217;s different. Its analysis has a depth of insight that&#8217;s not there with other models. I&#8217;ve found myself reading its outputs with actual interest, not just because I&#8217;m amazed at the technical marvel in front of me, but because I start to find the analysis compelling and revealing.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://generativehistory.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://generativehistory.substack.com/subscribe?"><span>Subscribe now</span></a></p><h3><strong>Limitations and More Weirdness</strong></h3><p>Deep Research also has some odd quirks as well as some significant issues. Some of this is related to the fact that it plans and carries out its own research. If you mess-up the prompt in any way, or inadvertently send it off in the wrong directions with your follow-up answers, you can&#8217;t do anything about it. This can be frustrating as you watch its reasoning process develop. In one test, I watched its reasoning devolve into a string of random characters. It seemed to recover but still, what happened?</p><p>The weirdest result, though, was also the most interesting. A few years ago, <a href="https://muse.jhu.edu/pub/437/article/853341/pdf">I showed</a> that Alexander Henry&#8217;s 1809 Travels and Adventures in Canada and the Indian Territories was actually written by English children&#8217;s author and grifter Edward Augustus Kendall, partly from papers he stole from Henry in Montreal and partly from material he lifted from earlier travelogues. So I wondered whether Deep Research could <a href="https://chatgpt.com/share/67a2b3a9-6b0c-8012-9d23-aefcffdb6dc1">help me compare Henry&#8217;s text to other travelogues</a> to find any overlaps.</p><p>The results were truly surprising. Deep Research provided a <a href="https://chatgpt.com/share/67a2b3a9-6b0c-8012-9d23-aefcffdb6dc1">6,165 word answer</a> which it claimed was based on &#8220;a combination of computational textual analysis (keyword and n-gram comparisons across digitized texts) and close reading [to] identify instances where Henry&#8217;s language or narrative motifs parallel those of his predecessors.&#8221; What? Really?</p><h3><strong>Is Deep Research Hallucinating Abilities it Doesn&#8217;t Have?</strong></h3><p>Running an n-gram analysis would require Deep Research to have access either to tools or the ability to write and run python code. <a href="https://openai.com/index/introducing-deep-research/">OpenAI&#8217;s announcement</a> indeed says: &#8220;Deep research independently discovers, reasons about, and consolidates insights from across the web. To accomplish this, it was trained on real-world tasks requiring browser and Python tool use, using the same reinforcement learning methods behind OpenAI o1, our first reasoning model.&#8221;</p><p>To probe the case, I tried again and asked the model to show its work, outputting any relevant code at the end. This time <a href="https://chatgpt.com/share/67a52577-eef0-8012-b8ca-063048b9e7ce">it proposed</a> not only to conduct an n-gram analysis, but also to measure cosine similarity using TF-IDF vectors and to calculate Jaccard similarity coefficients. <a href="https://chatgpt.com/share/67a52577-eef0-8012-b8ca-063048b9e7ce">In its response</a>, it provided some python code but when I ran it, although it worked, it didn&#8217;t reproduce the model&#8217;s results. In fact the numbers were wildly off. I tried several other examples with similar results (same with the other functions). Another serious problem is that many of the quotes in the response were either entirely made up or poorly paraphrased.</p><p>So how do we square this with the other results above? First, it is a good reminder to treat LLMs critically. But don&#8217;t breathe a sigh of relief either because I don&#8217;t think these failures stem from problems with the model itself.</p><p>I would love to hear from OpenAI on this, but my intuition is that Deep Research went off the rails because it tried to do something that it couldn&#8217;t actually do. I suspect that although OpenAI trained Deep Research on python, it did not give it the ability to download and store the actual text files required to make its code work. When things like that happen to LLMs, they tend to go off the rails pretty quickly because they don&#8217;t have enough information to understand what is happening and they start to spiral. I am speculating here, but when the code failed, Deep Research may have found that it did not have the actual text in its context and tried to recall the quotations as best it could. And when models do that, like people, they don&#8217;t do a great job and start to confabulate.</p><p>My theory is supported at least in part <a href="https://chatgpt.com/share/67a636fc-ae40-8012-8391-a6fb8e60fd8e">by a final test </a>in which I gave Deep Research the same problem but told it not to use any code and to conduct a purely qualitative analysis instead. This time, the citations all worked and the quotations were accurate. The analysis was also, perhaps not surprisingly, much better. Although I will admit that I haven&#8217;t yet clicked through all the links it provided, at the very least its qualitative attempt was significantly better.</p><p>Assuming this is the case, these types of errors suggest Deep Research has a number of latent abilities that OpenAI is likely to release at a later date. As these start to converge with access to paywalled sources, LLMs are going to do some really interesting and strange things.</p><h3><strong>Conclusions</strong></h3><p>My sense is that we are going to look back and see Deep Research as the start of a new era when AI agents started to work. This has enormous implications in general. As Logan Kirkpatrick of Google&#8217;s DeepMind <a href="https://x.com/OfficialLoganK/status/1864508209769390238">recently observed</a>, people simply aren&#8217;t preparing for a world where intelligence is effectively free. But let&#8217;s focus on what it means for us as historians and academics specifically.</p><p>In <a href="https://joshuagans.substack.com/p/what-will-ai-do-to-presearch">an excellent blog post</a>, Joshua Gans recently argued that these new reasoning models fundamentally undermine our existing research model because they allow us to answer many research questions on demand rather than through months or even years of laborious work. What is the point, he asks, of conventional academic publishing in such a world? This begs an obvious question: are we the last generation of human historians?</p><p>While I&#8217;ve heard lots of counter-arguments over the last couple years&#8212;ranging from the ethical to the aesthetic&#8212;I have yet to hear one that convincingly parries the core issue: our work has value because it is time consuming and our expertise is specific and comparatively rare. As LLMs challenge that basic equation, things will inevitably change for us just as they have for any economic group faced with automation throughout history. To be clear, I don&#8217;t like this either and I have real moral, ethical, and methodological concerns about machine generated histories.</p><p>This is why we need to stop talking about LLMs in the abstract and start having serious conversations about the future of our discipline. My intuition is still that humans are going to need to be in the loop and that people will prefer human generated histories to machine made ones, but am I right? Even if I am, what is history going to look like in this brave new world? Are we willing to harness these tools and work alongside them? What exactly do we bring to the table that LLMs do not? If we keep avoiding these hard conversations, the rest of the world may move on without us.</p><div class="captioned-button-wrap" data-attrs="{&quot;url&quot;:&quot;https://generativehistory.substack.com/p/is-this-the-last-generation-of-historians?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="CaptionedButtonToDOM"><div class="preamble"><p class="cta-caption">Thanks for reading Generative History! This post is public so feel free to share it.</p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://generativehistory.substack.com/p/is-this-the-last-generation-of-historians?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://generativehistory.substack.com/p/is-this-the-last-generation-of-historians?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://generativehistory.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://generativehistory.substack.com/subscribe?"><span>Subscribe now</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[Transcription Pearl for Non-Coders and Next Steps]]></title><description><![CDATA[I've added a manual and downloadable executable file for non-coders who want to try Transcription Pearl on Windows PCs (sorry Mac users)]]></description><link>https://generativehistory.substack.com/p/transcription-pearl-for-non-coders</link><guid isPermaLink="false">https://generativehistory.substack.com/p/transcription-pearl-for-non-coders</guid><dc:creator><![CDATA[Mark Humphries]]></dc:creator><pubDate>Tue, 12 Nov 2024 19:05:22 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!2k0G!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb56b009f-cd35-4bbc-b3d8-7069813f6475_1284x1282.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!2k0G!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb56b009f-cd35-4bbc-b3d8-7069813f6475_1284x1282.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!2k0G!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb56b009f-cd35-4bbc-b3d8-7069813f6475_1284x1282.png 424w, https://substackcdn.com/image/fetch/$s_!2k0G!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb56b009f-cd35-4bbc-b3d8-7069813f6475_1284x1282.png 848w, https://substackcdn.com/image/fetch/$s_!2k0G!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb56b009f-cd35-4bbc-b3d8-7069813f6475_1284x1282.png 1272w, https://substackcdn.com/image/fetch/$s_!2k0G!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb56b009f-cd35-4bbc-b3d8-7069813f6475_1284x1282.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!2k0G!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb56b009f-cd35-4bbc-b3d8-7069813f6475_1284x1282.png" width="1284" height="1282" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b56b009f-cd35-4bbc-b3d8-7069813f6475_1284x1282.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1282,&quot;width&quot;:1284,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:906474,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!2k0G!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb56b009f-cd35-4bbc-b3d8-7069813f6475_1284x1282.png 424w, https://substackcdn.com/image/fetch/$s_!2k0G!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb56b009f-cd35-4bbc-b3d8-7069813f6475_1284x1282.png 848w, https://substackcdn.com/image/fetch/$s_!2k0G!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb56b009f-cd35-4bbc-b3d8-7069813f6475_1284x1282.png 1272w, https://substackcdn.com/image/fetch/$s_!2k0G!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb56b009f-cd35-4bbc-b3d8-7069813f6475_1284x1282.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>If you&#8217;ve been interested in trying Transcription Pearl but are unfamiliar with running python code, I&#8217;ve created a stand-alone, no-code version of the program that runs on Windows PCs. Unfortunately, it won&#8217;t run on Macs at this time. Sorry!</p><p>You can download the program from the <a href="https://github.com/mhumphries2323/Transcription_Pearl">GitHub Repository </a>or directly from <a href="https://github.com/mhumphries2323/Transcription_Pearl/releases/download/v1.0.0-beta/TranscriptionPearl.exe">this link</a>. Then it is as simple as putting the file on your desktop and double-clicking. I&#8217;ve also co-authored a manual with Claude Sonnet-3.5 (more on that below&#8230;it was an interesting process) which you can <a href="https://github.com/mhumphries2323/Transcription_Pearl/releases/download/v1.0.0-beta/Transcription.Pearl.Manual.1.0.Beta.pdf">download here</a>. It will walk you through the process of installing the program and working with your documents.</p><p>An important security note: when you run Transcription Pearl, it may be read as a virus or other potentially malicious program by Windows Defender or other antivirus software.  The main reason this happens, as far as I can tell, is that it makes API calls to an external server. If I were a professional software developer, I would have access to certifications that would allow me to verify the program&#8217;s authenticity with various antivirus providers, but that is not really something I am able to do as a non-professional. So apologies for the scary warnings, but if this happens, you can always choose to add an exception and/or run the program anyway.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://generativehistory.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://generativehistory.substack.com/subscribe?"><span>Subscribe now</span></a></p><h4>It Sounds Scary, but Getting an API Key is Easy</h4><p>To make this work, you will need to get at least one API Key from OpenAI, Anthropic, and/or Google Gemini. It is best to use one model for transcription and another for correction (we found that the Gemini 1.5-Pro-002 was best at the initial transcritpion while Claude Sonnet-3.5 was best at correcting transcriptions). While this might sound intimidating, it should not be: it is as simple as signing up for any website and takes about three clicks. If you can subscribe to a newspaper, you can do this! </p><p>The process is outlined in the user manual, but here it is for each provider:</p><blockquote><p><em>OpenAI API Keys</em></p><p>To register for an OpenAI API key, first create an account on the <a href="https://platform.openai.com/tokenizer">OpenAI platform</a> (note that this will be different than a ChatGPT account). Click your profile icon in the top right-hand corner and select &#8220;Your Profile&#8221;. In the menu on the left, click API-keys and then click the green &#8220;+Create New Secret Key&#8221; at the top right. Give your Key a name and create the key.</p><p><em>Anthropic API Keys</em></p><p>To register for an Anthropic API key, first create an account on the <a href="https://console.anthropic.com/login">Anthropic consol</a> (this will be different than your Claude chat account). After you create an account, cock your profile icon at the top right and select &#8220;API Keys&#8221; from the dropdown and then select the orange &#8220;+Create Key&#8221; button at the top right.</p><p><em>Gemini API Keys</em></p><p>The process with Google can be somewhat more confusing as it provides a number of different ways to access their models via API, specifically via the Google AI Studio and Vertex platforms. You need to use the Google AI Studio platform. First, <a href="https://ai.google.dev/aistudio">create a Google AI Studio Account</a> then, once you are logged in, select the blue &#8220;Get API Key&#8221; button on the left. In the next pane, click the blue &#8220;Create API Key&#8221; button.</p></blockquote><h4>APIs, Privacy, and Security</h4><p>If you choose to try Transcription Pearl, you might have questions about privacy and security. Don&#8217;t we all! As this is an evolving field, there are lots of questions that still need to be defined about rights, ethics, etc. In general, though, I think that using APIs rather than ChatBots is more secure and more ethical. Here&#8217;s why. </p><p>Unlike ChatGPT, the API versions of the <a href="https://openai.com/enterprise-privacy/">OpenAI</a>, <a href="https://privacy.anthropic.com/en/articles/7996885-how-do-you-use-personal-data-in-model-training">Anthropic</a>, and <a href="https://cloud.google.com/vertex-ai/generative-ai/docs/data-governance">Google </a>LLMs will not retain or store your data nor will they use it for training their models. So you aren&#8217;t feeding the models with data. Don&#8217;t take my word for it, though: read each of their privacy policies (as they pertain to the API which is what Transcription Pearl uses) by clicking the links above. But if you don&#8217;t trust these companies to comply with their own user agreements, don&#8217;t use LLMs.</p><p>I personally see little difference between uploading a document to a cloud server like OneDrive and Google Drive or using an LLM via an API: both involve sending the same information over the internet to an external server for research purposes (although with LLMs it is not being retained or stored). To my mind, if you aren&#8217;t allowed to send it to an LLM for legal or ethical reasons, you probably also can&#8217;t store it on the cloud (even if you hadn&#8217;t given that as much thought). So the key to me when you go to transcribe a document is: do you have the necessary rights to use those materials? And if you don&#8217;t, or you are unsure, you can&#8217;t use them. </p><p>For that reason, I only use AI on materials that are open-access, out of copyright, or that I know I can clearly use under the educational and fair-use provisions of the <a href="https://laws-lois.justice.gc.ca/eng/acts/c-42/fulltext.html">Copyright Act</a>. Of course, you also cannot generally use either ChatBots or APIs with sensitive materials that have restrictions placed on them by Research Ethics Boards or external organizations (in the same way that you can&#8217;t generally store them on cloud services either). But even here it is worth noting that this is starting to change: the OpenAI API can be <a href="https://help.openai.com/en/articles/8660679-how-can-i-get-a-business-associate-agreement-baa-with-openai">certified </a>for use with Health Insurance Portability and Accountability Act (HIPAA) related materials in the United States, provided researchers get the necessary ethics approvals.</p><p>For now, know that if you use Transcription Pearl, you are using the more secure API interfaces to the LLMs mentioned above, not unsecure ChatBots. Nevertheless, it is up to you to ensure that you have the necessary rights and permissions to work with your documents.</p><h4>Writing Technical Manuals with AI</h4><p>As something of an aside, the process of writing a comprehensive manual for Transcription Pearl was illuminating. I began by developing a table of contents and then copied and pasted my source code (about 5,000 lines between the Transcription Pearl and Image Pre-Processing Tool) into Claude Sonnet-3.5 along with a few images of the interface. I then asked Sonnet-3.5 to write each of the technical sections based on the source code and images and it generally did a great job. I then edited, refined, and added material as necessary. </p><p>The manual is about 5,000 words long and it only took around two hours to produce, start to finish. While I don&#8217;t think technical writing will be replaced by AI, I can see how it will make competent technical writers much faster and more efficient. And that means the field will invariably shrink. </p><h4>Next Steps</h4><p>I&#8217;ve been updating Transcription Pearl fairly regularily over the last few weeks in response to feedback (thanks!) and I&#8217;ll continue to do so. But I see Transcription Pearl as the first step in a larger workflow. </p><p>Ultimately, the goal is to produce software that will automate the process of allowing historians to transcribe texts, upload them to a database, and then perform various tasks with them like asking questions, summarization, and finding connections between documents that might otherwise remain hidden. </p><p>To this end, I&#8217;ve been working on the next version of Transcription Pearl which will allow users to automatically generate metadata for a corpus of documents, exporting a sort of finding aid or box list to a spreadsheet. It works like this: you might upload a 100 page folder of documents containing letters that vary in length from one page to a few pages. </p><p>After transcribing and correcting the documents in Transcription Pearl, the user hits &#8220;Process Text&#8221; and an LLM automatically generates a table of contents for the folder, noting where each letter begins and ends. It then reads each letter in turn and extracts metadata like the names of each person or place mentioned in the document, the name of its author, the name of the main correspondent, the date it was written, and where it was written. </p><p>It also summarizes the main points in the document according to requirements set by the user (IE to emphasize a certain research topic or theme), it answer questions like is this document relevant to X topic, and assigns keywords/subject headings either automatically or from a bank of choices created by the user. The output is then saved to a spreadsheet with each row assigned to a unqiue document, along with the original text and path to the image(s).</p><p>While I think this will prove quite useful for researchers on its own, the goal is to ladder this into a larger piece of software. The exported spreadsheet can also be uploaded to a database and combined with other similar spreadsheets, creating a huge archive. Over time, this would become a fully searchable archive using keyword or semantic approaches, making all one&#8217;s research accessible to an AI research assistant. </p><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://generativehistory.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Generative History! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Introducing Transcription Pearl]]></title><description><![CDATA[A Practical AI Tool for the Automated Transcription of Historical Handwritten Documents with State-of-the-Art Accuracy]]></description><link>https://generativehistory.substack.com/p/introducing-transcription-pearl</link><guid isPermaLink="false">https://generativehistory.substack.com/p/introducing-transcription-pearl</guid><dc:creator><![CDATA[Mark Humphries]]></dc:creator><pubDate>Fri, 01 Nov 2024 15:40:56 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!UjeX!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff294a6f4-7f86-4f43-abd1-019d97a428f3_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!UjeX!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff294a6f4-7f86-4f43-abd1-019d97a428f3_1024x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!UjeX!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff294a6f4-7f86-4f43-abd1-019d97a428f3_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!UjeX!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff294a6f4-7f86-4f43-abd1-019d97a428f3_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!UjeX!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff294a6f4-7f86-4f43-abd1-019d97a428f3_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!UjeX!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff294a6f4-7f86-4f43-abd1-019d97a428f3_1024x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!UjeX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff294a6f4-7f86-4f43-abd1-019d97a428f3_1024x1024.png" width="1024" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f294a6f4-7f86-4f43-abd1-019d97a428f3_1024x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1666536,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!UjeX!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff294a6f4-7f86-4f43-abd1-019d97a428f3_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!UjeX!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff294a6f4-7f86-4f43-abd1-019d97a428f3_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!UjeX!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff294a6f4-7f86-4f43-abd1-019d97a428f3_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!UjeX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff294a6f4-7f86-4f43-abd1-019d97a428f3_1024x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://generativehistory.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://generativehistory.substack.com/subscribe?"><span>Subscribe now</span></a></p><p>One of the things I find puzzling about reactions to AI, is that people often tend to focus on its far reaching but still entirely theoretical effects while ignoring the many smaller, less glamorous but highly practical things that AI can do right now. One of those practical things is <a href="https://generativehistory.substack.com/p/making-the-infeasible-practical-in">transcribing historical handwriting</a>. As we&#8217;ll see below, it&#8217;s also a good example of how we can use LLMs to do lots of monotonous tasks very quickly <a href="https://generativehistory.substack.com/p/why-openais-new-model-might-change">that have nothing to do with chatbots.</a> But first we need to get away from the chat interface and find ways to integrate LLMs into software tools.</p><p>In this blog post, I introduce <em><a href="https://github.com/mhumphries2323/Transcription_Pearl">Transcription Pearl</a></em>, an open-source program that transcribes handwritten documents and corrects transcriptions generated by other HTR programs with state-of-the-art accuracy. Depending on how you count errors (more on that below), it can achieve character level accuracy above 98% and word-level accuracy above 96% making it about as accurate as most human transcribers and better than other HTR tools. It also costs 1/50 the amount to use as other popular programs like <em>Transkribus</em> and is about 50 times faster.</p><p>Although I was responsible for the programming, the program is the work of many people. It grows out of a larger project I&#8217;ve been working on with Lianne Leddy and a team of student researchers who&#8217;ve tested its various iterations over the past two years here at Wilfrid Laurier University: Quinn Downton, Meredith Legace, John McConnell, Isabella Murray, and Elizabeth Spence. You can download it and run it today from <a href="https://github.com/mhumphries2323/Transcription_Pearl">our GitHub Repository </a>(you&#8217;ll need to get an <a href="https://platform.openai.com/docs/quickstart">OpenAI</a>, <a href="https://docs.anthropic.com/en/api/getting-started">Anthropic</a>, or <a href="https://cloud.google.com/apigee?utm_source=google&amp;utm_medium=cpc&amp;utm_campaign=na-CA-all-en-dr-bkws-all-all-trial-b-dr-1707554&amp;utm_content=text-ad-none-any-DEV_c-CRE_665735485403-ADGP_Hybrid+%7C+BKWS+-+MIX+%7C+Txt-API+Management-Apigee-KWID_43700077225654810-aud-2232802565252:kwd-62094780&amp;utm_term=KW_google%20api-ST_google+api&amp;gad_source=1&amp;gclid=Cj0KCQjw1Yy5BhD-ARIsAI0RbXZetfYWXcBkZy67yiv7sOyEawowuV2T0urSyhnj8xcoROe9ot0gMYcaAq9QEALw_wcB&amp;gclsrc=aw.ds">Google API key</a> too). You can also read <a href="https://papers.ssrn.com/sol3/Delivery.cfm/5006071.pdf?abstractid=5006071&amp;mirid=1">a pre-print of the paper</a> we wrote  which I&#8217;ll link to it throughout this post when it can provide more information.</p><p>But first for the more technically experienced readers, full disclosure: the software interface is pretty simple and I imagine the code leaves much to be desired. Please remember that I am a historian and not a professional software developer. But I think that this actually illustrates another really important point: although in February 2023 I had some ancient experience in visual basic and html, I had never written anything in python. ChatGPT and Claude Sonnet-3.5 taught me the language and helped me write the code for <em>Transcription Pearl</em>. That is pretty remarkable and its something that would not have happened without generative AI.</p><h4>Handwritten Text Recognition and Error Rates</h4><p>First some essential background. Handwritten Text Recognition (HTR) has been a challenging area of research in computer vision and pattern recognition since the 1960s. <a href="https://doi.org/10.3390/jimaging10010018">Traditional HTR approaches</a> involve a complex, multi-step process in which handwritten documents are digitized and then pre-processed to improve their quality. The text must then be segmented into lines, words, and individual characters before machine learning algorithms can try to identify patterns. While HTR methods have made significant progress in recent years, especially with <a href="https://doi.org/10.1109/ICDAR.2017.307">programs like </a><em><a href="https://doi.org/10.1109/ICDAR.2017.307">Transkribus</a></em> that (at least in part) uses the same underlying transformer architecture as LLMs, they still struggle to generalize from their training data to new handwriting styles. If you&#8217;ve ever tried HTR, you&#8217;ll know that it can make lots of errors, sometimes so many that it&#8217;s not worth the effort.</p><p>In the field of HTR, accuracy is typically measured using Character Error Rates (CER) and Word Error Rates (WER). CER represents the percentage of incorrectly transcribed characters, while WER indicates the percentage of words that are incorrectly substituted, added, or deleted. As a rule of thumb, WER is usually about 3 to 4 times higher than CER. Programs like <em>Transkribus</em> typically achieve CERs of 8-25% and WERs of 15-50% on individual&#8217;s handwriting they have not seen before in training(<a href="https://doi.org/10.17615/4hnn-kh38">1</a>, <a href="https://doi.org/10.1002/pra2.958">2</a>, <a href="https://doi.org/10.18420/infdh2018-13">3</a>). To achieve a usable level of accuracy, <a href="https://readcoop.eu/how-to-improve-the-cer-of-your-model/">users must fine-tune </a>an HTR model, typically providing around 75,000 words of transcribed text and image pairs for <em>each</em> individual handwriting style they need to transcribe. While fine tuning makes it possible (<a href="https://doi.org/10.18420/infdh2018-13">1</a>, <a href="https://aclanthology.org/2022.cltw-1.17">2</a>) to achieve CERs of 1.5 to 5% and WERs of 6 to 12%, it is probably wasted effort unless you are transcribing an entire collection of single-authored papers.</p><p>To put these error rates in context, professional transcription services typically guarantee around 99% accuracy at the word level, meaning 1 out of every 100 words would be transcribed incorrectly. However, this is only under ideal conditions with clear handwriting and high-quality images, which historians will know is the exception rather than the rule. Studies on non-expert human transcription error rates are limited, but research suggests WERs ranging from <a href="https://doi.org/10.48550/arXiv.1708.08615">4% to 10% for speech transcription</a> and around 10% for paper medical record transcription (<a href="https://doi.org/10.3928/01477447-20200619-10">1</a>, <a href="https://doi.org/10.1016/j.ijmedinf.2017.04.015">2</a>). One of the <a href="https://infoscience.epfl.ch/record/255998/files/final_abstract_dh18.pdf">only studies</a> I could locate on historical transcription reported human WERs as high as 35-43% for non-specialists transcribing early modern Italian documents. Anecdotally, I&#8217;ve found that I achieve something like a 1 to 2% CER and 3% to 6% WER on manual transcription.</p><p>Here we need to also remember that error rates are not as straightforward as they seem: not all changes between a ground-truth document and transcript are equal. Edits that standardize capitalization, punctuation, and spelling, called emendations in the editorial world, can be entirely acceptable in some circumstances or highly problematic in others depending on the purpose and context of the transcription project. However, substitutions, deletions, or additions that go beyond these things are almost always problematic: think missed words, added words, or writing horse instead of cow in an estate inventory. Editors also need to consider how to handle marginalia, insertions, strikethroughs, dates, and formatting, choices that can also affect how accuracy is scored (should a strikethrough be included in the text or placed in parentheses?). For many applications, such as keyword searching or general accessibility, a higher WER may be okay if most of the errors relate to emendations rather than skipped or misread words. For publishing critical editions, quotations, or detailed analysis, a higher standard is necessary.</p><h4>Transcription Pearl</h4><p>Our initial goal with <em>Transcription</em> <em>Pearl</em> was to create a piece of software that could help transcribe handwritten documents to be used in our AI research assistant project. Initially we had success getting LLMs like GPT-4 to correct transcriptions created by conventional HTR programs like <em>Transkribus</em>, but they were not very good at reading handwriting from scratch (and it was too expensive). But as LLMs improved, they started to achieve a surprising level of success, first with GPT-4o, later Sonnet-3.5, and then Gemini-1.5-Pro-002.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!wkO9!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F90bcaef3-b694-464c-8ec5-e85e3d8f0df5_1952x1281.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!wkO9!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F90bcaef3-b694-464c-8ec5-e85e3d8f0df5_1952x1281.png 424w, https://substackcdn.com/image/fetch/$s_!wkO9!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F90bcaef3-b694-464c-8ec5-e85e3d8f0df5_1952x1281.png 848w, https://substackcdn.com/image/fetch/$s_!wkO9!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F90bcaef3-b694-464c-8ec5-e85e3d8f0df5_1952x1281.png 1272w, https://substackcdn.com/image/fetch/$s_!wkO9!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F90bcaef3-b694-464c-8ec5-e85e3d8f0df5_1952x1281.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!wkO9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F90bcaef3-b694-464c-8ec5-e85e3d8f0df5_1952x1281.png" width="1456" height="956" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/90bcaef3-b694-464c-8ec5-e85e3d8f0df5_1952x1281.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:956,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1405860,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!wkO9!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F90bcaef3-b694-464c-8ec5-e85e3d8f0df5_1952x1281.png 424w, https://substackcdn.com/image/fetch/$s_!wkO9!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F90bcaef3-b694-464c-8ec5-e85e3d8f0df5_1952x1281.png 848w, https://substackcdn.com/image/fetch/$s_!wkO9!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F90bcaef3-b694-464c-8ec5-e85e3d8f0df5_1952x1281.png 1272w, https://substackcdn.com/image/fetch/$s_!wkO9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F90bcaef3-b694-464c-8ec5-e85e3d8f0df5_1952x1281.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 1: Main Transcription Pearl interface</figcaption></figure></div><p><em>Transcription Pearl</em> works by harnessing the flexibility and contextual understanding of LLMs for handwriting recognition tasks. Unlike traditional HTR models that use complex preprocessing and segmentation, it relies on the fact that LLMs can generally find, organize, and then read the text in an image without preprocessing or fine-tuning. To correct transcriptions, we use into the fact that LLMs are also inherently big text prediction machines. This makes them very good at figuring out how to correct erroneous characters and words based on the surrounding context.</p><p>In <em>Transcription Pearl</em> users drag-and-drop images into a workspace and select &#8220;Transcribe&#8221; from a menu; thirty seconds later, the transcriptions appear in an editing box beside the original image. The user can then send the text and image to another LLM to correct errors in the transcription, a process that takes another 30 seconds. Users can also skip the LLM transcription step and import transcriptions generated in programs like <em>Transkribus</em> for automated correction. In a settings menu, users specify the LLM they want to use for both functions from any of the major multimodal OpenAI, Anthropic, or Google models and to customize the instructions given to the models.</p><p>Here is what happens behind the scenes. When the user hits &#8220;Transcribe&#8221;, the program sends the image to an LLM for processing via an API, a secure way of sending information to an external server. The model returns the transcription and <em>Transcription Pearl</em> then extracts the transcription and matches it up with the right image in the workspace. Users can then edit the text and export the results. The key is that all of this happens in parallel, meaning instead of sending one image at a time all the images are sent at once which significantly speeds up the process. Depending on your usage tier with the various LLM providers, you can process hundreds of images a minute.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!9_8z!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fade5f951-68e6-45be-a689-bad5bb393af6_1800x1360.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!9_8z!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fade5f951-68e6-45be-a689-bad5bb393af6_1800x1360.png 424w, https://substackcdn.com/image/fetch/$s_!9_8z!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fade5f951-68e6-45be-a689-bad5bb393af6_1800x1360.png 848w, https://substackcdn.com/image/fetch/$s_!9_8z!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fade5f951-68e6-45be-a689-bad5bb393af6_1800x1360.png 1272w, https://substackcdn.com/image/fetch/$s_!9_8z!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fade5f951-68e6-45be-a689-bad5bb393af6_1800x1360.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!9_8z!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fade5f951-68e6-45be-a689-bad5bb393af6_1800x1360.png" width="1456" height="1100" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ade5f951-68e6-45be-a689-bad5bb393af6_1800x1360.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1100,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:158789,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!9_8z!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fade5f951-68e6-45be-a689-bad5bb393af6_1800x1360.png 424w, https://substackcdn.com/image/fetch/$s_!9_8z!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fade5f951-68e6-45be-a689-bad5bb393af6_1800x1360.png 848w, https://substackcdn.com/image/fetch/$s_!9_8z!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fade5f951-68e6-45be-a689-bad5bb393af6_1800x1360.png 1272w, https://substackcdn.com/image/fetch/$s_!9_8z!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fade5f951-68e6-45be-a689-bad5bb393af6_1800x1360.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 2: Transcription Pearl settings menu</figcaption></figure></div><p>Like any LLM interaction, the heart of the &#8220;programming&#8221; is natural language. Both the transcribe and correct functions send the images and/or text to the LLM along with a system message that tells it what to do and a user prompt that organizes the specific task. As a default, the program employs very simple prompts; better results could probably be obtained with some solid prompt engineering. As our goal here was to establish a baseline, though, we kept it straight forward.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://generativehistory.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://generativehistory.substack.com/subscribe?"><span>Subscribe now</span></a></p><h4>Promising Results</h4><p>While there are several standard HTR datasets available to researchers on the internet, we found that Google&#8217;s Gemini model had &#8220;seen&#8221; most of the documents in training (it spits out a warning when it is asked to &#8220;recite&#8221; training data). While this warning is unique to Gemini, we have to assume that the OpenAI and Anthropic models were also trained on the same data. So we decided to assemble our own dataset from documents we personally photographed and that we know are not on the internet. The result is a corpus of 50 pages of 18th and early 19th century English documents, totalling about 10,000 words and featuring more than 30 different hands. The characteristics of the images also vary widely: some are taken from microfilms, others were captured with an iPhone or a Cannon SL; a few are blurry and only a couple are truly high quality. In other words, they are pretty representative of the types of images historians use in their research everyday. We&nbsp; are not publishing the set as we want to be able to use it again in future (if we put it on the internet, it will almost certainly be used in training future models and s0 will no longer provide a reliable, common basis of comparison).</p><p>The results are pretty impressive. Depending on how you count the errors, <em>Transcription Pearl</em> achieved CERs as low as 5.7% and WERs of 8.9% without any fine-tuning or preprocessing&#8212;a significant improvement over specialized HTR software. Even more impressive is what happens when you ladder one LLM onto another: by using LLMs to correct initial transcriptions (generated either by other LLMs or by conventional HTR programs like Transkribus), the tool can produce results with human levels of accuracy. </p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!oG1G!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ed95c0f-18cf-4baa-8f54-1b0fd87df209_693x373.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!oG1G!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ed95c0f-18cf-4baa-8f54-1b0fd87df209_693x373.png 424w, https://substackcdn.com/image/fetch/$s_!oG1G!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ed95c0f-18cf-4baa-8f54-1b0fd87df209_693x373.png 848w, https://substackcdn.com/image/fetch/$s_!oG1G!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ed95c0f-18cf-4baa-8f54-1b0fd87df209_693x373.png 1272w, https://substackcdn.com/image/fetch/$s_!oG1G!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ed95c0f-18cf-4baa-8f54-1b0fd87df209_693x373.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!oG1G!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ed95c0f-18cf-4baa-8f54-1b0fd87df209_693x373.png" width="693" height="373" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4ed95c0f-18cf-4baa-8f54-1b0fd87df209_693x373.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:373,&quot;width&quot;:693,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:108036,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!oG1G!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ed95c0f-18cf-4baa-8f54-1b0fd87df209_693x373.png 424w, https://substackcdn.com/image/fetch/$s_!oG1G!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ed95c0f-18cf-4baa-8f54-1b0fd87df209_693x373.png 848w, https://substackcdn.com/image/fetch/$s_!oG1G!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ed95c0f-18cf-4baa-8f54-1b0fd87df209_693x373.png 1272w, https://substackcdn.com/image/fetch/$s_!oG1G!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ed95c0f-18cf-4baa-8f54-1b0fd87df209_693x373.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!IF_p!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a0be621-02e6-42ad-b77f-eafa2706f493_700x386.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!IF_p!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a0be621-02e6-42ad-b77f-eafa2706f493_700x386.png 424w, https://substackcdn.com/image/fetch/$s_!IF_p!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a0be621-02e6-42ad-b77f-eafa2706f493_700x386.png 848w, https://substackcdn.com/image/fetch/$s_!IF_p!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a0be621-02e6-42ad-b77f-eafa2706f493_700x386.png 1272w, https://substackcdn.com/image/fetch/$s_!IF_p!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a0be621-02e6-42ad-b77f-eafa2706f493_700x386.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!IF_p!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a0be621-02e6-42ad-b77f-eafa2706f493_700x386.png" width="700" height="386" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4a0be621-02e6-42ad-b77f-eafa2706f493_700x386.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:386,&quot;width&quot;:700,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:109287,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!IF_p!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a0be621-02e6-42ad-b77f-eafa2706f493_700x386.png 424w, https://substackcdn.com/image/fetch/$s_!IF_p!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a0be621-02e6-42ad-b77f-eafa2706f493_700x386.png 848w, https://substackcdn.com/image/fetch/$s_!IF_p!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a0be621-02e6-42ad-b77f-eafa2706f493_700x386.png 1272w, https://substackcdn.com/image/fetch/$s_!IF_p!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a0be621-02e6-42ad-b77f-eafa2706f493_700x386.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>While we tested lots of different model combinations, we found that the most effective pairing was to use Gemini-1.5-Pro-002 for the initial transcription and Claude Sonnet-3.5 for the corrections, achieving an average modified CER of 4.1% and WER of 7%. In this context a &#8220;modifed&#8221; error rate means one in which we exclude errors of capitalization, punctuation, and corrections to historical spelling. </p><p>The best overall results were achieved on <em>Transkribus</em> generated transcriptions corrected by Sonnet-3.5: a modified CER of 1.8% and WER of 3.5%. This means those transcriptions were useable for most &#8220;non-publishing&#8221; applications, so keyword searching, for readability, etc. That is pretty impressive.</p><p>Interestingly, we also found that LLMs could not correct their own transcriptions (something I want to write more about). My intuition is that because transcription errors are derived from probabilities learned in training, the model can&#8217;t recognize them as erroneous as they represent highly probable outputs. However, transcriptions generated by models with a different set of weights create errors that are derived from a different set of probabilities and the LLM thus has an easier time recognizing them. As I say,  an interesting result but more research is needed.</p><h4>Error Rates and Types</h4><p>Error rates were remarkably consistent&#8212;which is somewhat surprising for outputs from stochastic models. For our tests, we ran the corpus of documents through the program ten times for each model. If you read the accompanying article, you can look at the complete analysis in our tables but generally the standard deviation in the error rates was low. With Gemini transcriptions, for example, the modified CER ranged between 6% and 7.1% (average of 6.7% and 0.3% Standard Deviation) while the modified WER ranged between 10.2% and 11.7% for an average of 11% (standard deviation of 0.4%). The results for the other models were similar.</p><p>Hallucinations were also not a major issue. Most of the errors made by the models were, in fact, pretty human: they misread ambiguous punctuation marks, struggled to discern capital letters, and sometimes corrected historical (or idiosyncratic) spelling errors. Occasionally they added words, mainly missing articles or pronouns to clarify the text. Sometimes they skipped a word. Only rarely did they do something &#8220;unhuman&#8221;. A handful of times the models substituted a word like &#8220;pelts&#8221; for &#8220;furs&#8221;. Again, my intuition is that this is because pelts was a more probable output than furs. But this was a rare occurrence. In a few places where the text was decidedly illegible (passages that I could not decipher myself) the model did it&#8217;s best but produced results that were clearly incorrect. Are these hallucinations? Maybe technically they are, but they were reasonable guesses to my eye.</p><h4>Speed and Cost</h4><p>What truly sets <em>Transcription Pearl</em> apart is its ability to process documents in batches, dramatically accelerating the transcription workflow. By making parallel API calls to the underlying LLMs, the software can transcribe a set of 50 pages in just 30 seconds&#8212;nearly 50 times faster than leading HTR platforms. This batch processing capability is a game-changer for large-scale digitization projects, allowing researchers to quickly and accurately transcribe entire collections with minimal manual intervention. </p><p>The affordability of the process is also remarkable. Automated LLM transcriptions cost around 1,500 times less than human transcription services and 50 times less than comparable HTR solutions. To transcript and correct a page using Gemini and Claude costs just $0.0138 per page, making it an incredibly cost-effective solution for individual researchers and institutions alike for most use-cases. Where accuracy is paramount, the cost of correcting a one page Transkribus transcrription would be $0.27 per page&#8212;or about half a cent more than the cost of a page of Transkribus transcription.</p><h4>Conclusion</h4><p>To download and run the program, you can clone our GitHub repository and create an environment with the necessary libraries. If you don&#8217;t know what this means, that is OK: that is where I was two years ago. Ask ChatGPT to help you download and run a python program from a GitHub repository, making sure to explain that you have limited to no understanding of the steps involved. It will walk you through the process, allowing you to ask follow-ups and paste in any errors you might receive for further instructions.</p><p>I think <em>Transcription Pearl</em> is a good example of a relatively benign and practical AI tool: although it&#8217;s not glamorous, it speeds up the transcription process exponentially while costing next to nothing to use. This use of AI makes it possible digitize massive amounts of archival material quickly, making resources accessible to researchers and the public which should open new avenues for research. And this is probably just the beginning. If the technology continues to improve, we can expect performance to improve as well.</p><p>But there are some caveats. Our dataset was relatively small (10,000 words) and while we tried to make it representative of the types of documents we use on a daily basis (as well as representative of the varying quality of the digital images available), others may find that it performs better or worse on their records. We also only tested it on English language letters and legal documents, we did not attempt to transcribe ledgers or financial documents. </p><p>So all this is to say that while this is a really promising result, we&#8217;re anxious to learn about the successes or failures of others. You can <a href="https://papers.ssrn.com/sol3/Delivery.cfm/5006071.pdf?abstractid=5006071&amp;mirid=1">read our full paper</a> here and <a href="https://papers.ssrn.com/sol3/Delivery.cfm/5006071.pdf?abstractid=5006071&amp;mirid=1">download the code</a> for the program here (I plan to get a compiled version up soon). Please contact us or better yet, put your results in the comments below!</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://generativehistory.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Generative History! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[Why OpenAI’s New Model Might Change Everything for Researchers]]></title><description><![CDATA[LLMs don&#8217;t need to be perfect, they need to be reliable. O1 is a huge step in that direction.]]></description><link>https://generativehistory.substack.com/p/why-openais-new-model-might-change</link><guid isPermaLink="false">https://generativehistory.substack.com/p/why-openais-new-model-might-change</guid><dc:creator><![CDATA[Mark Humphries]]></dc:creator><pubDate>Fri, 13 Sep 2024 19:18:23 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!841Y!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4e25b2e0-8200-40ad-af55-c813ad59c6df_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!841Y!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4e25b2e0-8200-40ad-af55-c813ad59c6df_1024x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!841Y!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4e25b2e0-8200-40ad-af55-c813ad59c6df_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!841Y!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4e25b2e0-8200-40ad-af55-c813ad59c6df_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!841Y!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4e25b2e0-8200-40ad-af55-c813ad59c6df_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!841Y!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4e25b2e0-8200-40ad-af55-c813ad59c6df_1024x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!841Y!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4e25b2e0-8200-40ad-af55-c813ad59c6df_1024x1024.png" width="1024" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4e25b2e0-8200-40ad-af55-c813ad59c6df_1024x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!841Y!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4e25b2e0-8200-40ad-af55-c813ad59c6df_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!841Y!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4e25b2e0-8200-40ad-af55-c813ad59c6df_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!841Y!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4e25b2e0-8200-40ad-af55-c813ad59c6df_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!841Y!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4e25b2e0-8200-40ad-af55-c813ad59c6df_1024x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>OpenAI <a href="https://openai.com/index/learning-to-reason-with-llms/">announced</a> a new model yesterday called o1. It reasons in a way that reduces hallucinations and allows it to tackle more complex problems. It&#8217;s not perfect, but in this post I want to explore why I think it&#8217;s such a significant development for applying LLMs to academic research and how the approach may open-up new use-cases for LLMs.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://generativehistory.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Generative History! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h4>Introducing o1</h4><p>O1 is not a new foundation model (IE it is not GPT-5/Orion) but part of the GPT-4 class of LLMs. The base model is not yet available, but OpenAI made an early <a href="https://openai.com/index/introducing-openai-o1-preview/">preview version</a> available to Tier-5 OpenAI developers via the API (and fortunately this gave me early access). For context, the base O1 model is said to score <a href="https://openai.com/index/learning-to-reason-with-llms/">25-50% higher</a> on a range of metrics than the preview version. OpenAI says that they are continuing to train and scale the base version and will provide frequent updates.</p><p>This model seems to be the infamous &#8220;Strawberry&#8221; model long rumoured to be in development. Trained from the ground up, apparently on heavily curated open-source and licensed datasets, OpenAI used a new reinforcement learning approach to teach the model to &#8220;think&#8221; through its answer before writing it out. This process (Q*?), involved two main things: training the model to approach problems through Chain of Thought (CoT) reasoning while also giving it the ability to develop various possible answers and then choose the best one.</p><h4>How it Works</h4><p>CoT has <a href="https://arxiv.org/abs/2201.11903">been around for awhile</a> now as a prompting technique and involves some variation of telling the model to &#8220;think step by step&#8221; before answering which improves the quality of its answers. Here we need to remember that LLMs produce a response one word at a time by calculating the most probable next word. Each time it adds a word, it considers the entire prompt as well as each previous word in the response.</p><p>By telling the model to think step by step, you force it tackle problems one stage at a time which changes the probabilities used to calculate the next word. This is why it can improve the quality of the final answer, especially for complex problems, by baking essential information into the model&#8217;s calculations. But to this point it&#8217;s been a one shot, sequential process as the LLM can&#8217;t actually &#8220;stop&#8221; and consider the problem, correct itself, or check itself for hallucinations. With O1, OpenAI trained a model using reinforcement learning to do this automatically which, unlike prompting, will fundamentally alter how the model actually calculates the probabilities for the next token.</p><p>As exciting as that is, the really revolutionary thing is that OpenAI also introduced a new inference architecture that allows the model to <em>actually</em> work step by step, pause, reassess, backtrack, and start again. Its responses are not one shot, stream of conscious affairs, but generated through a deliberative process of step by step &#8220;reasoning&#8221;. We can imagine this as looking something like a tree growing out of the question with each new blossoming branch representing a different exploration of a potential answer to the question. This is revolutionary, because it represents a major shift in the inference process, giving the model the ability to &#8220;consider&#8221; a problem from different perspectives, test possible answers, back-track to correct itself, and select better courses of action. While OpenAI has not shared details of how the process actually works, their suggestion that this process can be scaled suggests that it is not exploring one tree at a time but many different trees in parallel, choosing the best answers from a large number of potential answers. How many, we don&#8217;t know. Dozens? Hundreds? Thousands? That is probably one of the things that is scalable.</p><h4>Scaling</h4><p>Now: with an eye to the future, let&#8217;s talk about why the mention of scaling is so significant here. Up until now, when we&#8217;ve talked about <a href="https://arxiv.org/abs/2001.08361">scaling LLMs</a> we&#8217;ve meant increasing the size of training data and the parameters (neurons) in the model architecture. Thus far, as a model&#8217;s training data and architecture increase, its capabilities have grown proportionally such that some argue the relationship is a &#8220;<a href="https://arxiv.org/abs/2001.08361">predictive law</a>&#8221;. Whether that will (or can) continue to hold true, though, is <a href="https://arxiv.org/pdf/2305.16264.pdf">fiercely debated</a>. When the next generation of scaled-up foundation models appear later this year or early next, we&#8217;ll have a better idea of whether scaling will start to bring diminishing returns or not.</p><p>But with the release of O1, OpenAI has introduced two new properties of LLMs that it says are also subject to similar scaling effects. The <a href="https://openai.com/index/learning-to-reason-with-llms/">O1 team writes</a>:</p><blockquote><p><em>&#8220;We have found that the performance of o1 consistently improves with more reinforcement learning (train-time compute) and with more time spent thinking (test-time compute). The constraints on scaling this approach differ substantially from those of LLM pretraining, and we are continuing to investigate them.&#8221;</em></p></blockquote><p>Train-time compute refers to the computational resources and time invested in teaching the model during training&#8212;essentially, it&#8217;s the model&#8217;s &#8220;learning phase.&#8221; Unlike pretraining that requires vast amounts of data, reinforcement learning and fine-tuning are more data-efficient because the model learns from interactions and feedback, often needing vastly less raw data to see significant improvements. So as models get bigger and bigger, finding such efficiencies will become more important.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!2QQc!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e3916f1-2bdb-494a-a2d0-54538cc20788_1208x739.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!2QQc!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e3916f1-2bdb-494a-a2d0-54538cc20788_1208x739.png 424w, https://substackcdn.com/image/fetch/$s_!2QQc!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e3916f1-2bdb-494a-a2d0-54538cc20788_1208x739.png 848w, https://substackcdn.com/image/fetch/$s_!2QQc!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e3916f1-2bdb-494a-a2d0-54538cc20788_1208x739.png 1272w, https://substackcdn.com/image/fetch/$s_!2QQc!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e3916f1-2bdb-494a-a2d0-54538cc20788_1208x739.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!2QQc!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e3916f1-2bdb-494a-a2d0-54538cc20788_1208x739.png" width="1208" height="739" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7e3916f1-2bdb-494a-a2d0-54538cc20788_1208x739.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:739,&quot;width&quot;:1208,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:97506,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!2QQc!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e3916f1-2bdb-494a-a2d0-54538cc20788_1208x739.png 424w, https://substackcdn.com/image/fetch/$s_!2QQc!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e3916f1-2bdb-494a-a2d0-54538cc20788_1208x739.png 848w, https://substackcdn.com/image/fetch/$s_!2QQc!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e3916f1-2bdb-494a-a2d0-54538cc20788_1208x739.png 1272w, https://substackcdn.com/image/fetch/$s_!2QQc!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e3916f1-2bdb-494a-a2d0-54538cc20788_1208x739.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Source: https://openai.com/index/learning-to-reason-with-llms/</figcaption></figure></div><p>Test-time compute, on the other hand, is the computational effort the model uses when it&#8217;s deployed to solve problems&#8212;the &#8220;thinking phase&#8221; when it actually comes up with an answer. What OpenAI is essentially claiming here is that by allowing the model to spend more time processing and reasoning, it will arrive at better and more accurate solutions.</p><p>In theory, test-time compute could increase exponentially and OpenAI envisions putting a model to work on some hard unscientific problems for days or even weeks with this model. That sounds compute intensive, but it is nothing compared to the requirements necessary to pretrain a GPT-4 level model. But the flip side is probably also true: it will be possible to dynamically scale compute up or down at test-time based on the complexity of the user&#8217;s question. This has important cost implications. With a conventional LLM, the compute expended on an answer is proportional to its length: the longer the answer the more expensive it is to generate. The model works just as hard whether it is regurgitating a recipe for banana bread or trying to explain quantum mechanics. The o1 architecture implies that it will be possible to scale compute to the complexity of the question.</p><p>The key takeaway is that enhancing capabilities through train-time and test-time compute doesn&#8217;t rely on scaling up data to enormous sizes, but rather on optimizing how the model learns and thinks with the data it has. This approach sidesteps the bottleneck of finite data and opens up new avenues for improving AI performance without the prohibitive costs associated with scaling traditional pretraining methods. It also makes it possible to envision creating a whole series of models finely tailored to specific tasks or areas of scientific research.</p><h4>Accuracy, Reliability, and Trustworthiness as Emerging Capabilities</h4><p>After using the model extensively for the last twenty-four hours, my intuition is that we can increasingly think of accuracy, reliability, and trustworthiness as emergent properties that may well scale with the o1 approach. By accuracy, I mean the model&#8217;s ability to produce correct and precise answers with fewer or more predictable hallucinations. Reliability, on the other hand, refers not to a consistency in its output, but a consistency in the &#8220;thinking&#8221; process itself. The implications of the approach described by OpenAI is that it should eventually become highly customizable, allowing organizations to train models to follow a specific series of protocols or thought processes that it can test and vet.</p><p>Both accuracy and reliability contribute to improved trustworthiness, but so too does the interpretability of the process itself. At the moment, OpenAI has opted to keep the &#8220;thought process&#8221; private and hidden from users, although in ChatGPT you can see a summary of what I assume are the successful steps in the CoT flow (but this is not yet available in the API). But as OpenAI notes in the <a href="https://openai.com/index/learning-to-reason-with-llms/">blog post</a>, this will eventually give users the ability to monitor the development of LLM responses for bias or unsafe actions as well as to quickly evaluate them by examining the underlying steps in the process.</p><p></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!UpcI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0f527ff-ca9b-4105-a384-e16c05140f23_1325x1803.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!UpcI!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0f527ff-ca9b-4105-a384-e16c05140f23_1325x1803.png 424w, https://substackcdn.com/image/fetch/$s_!UpcI!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0f527ff-ca9b-4105-a384-e16c05140f23_1325x1803.png 848w, https://substackcdn.com/image/fetch/$s_!UpcI!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0f527ff-ca9b-4105-a384-e16c05140f23_1325x1803.png 1272w, https://substackcdn.com/image/fetch/$s_!UpcI!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0f527ff-ca9b-4105-a384-e16c05140f23_1325x1803.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!UpcI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0f527ff-ca9b-4105-a384-e16c05140f23_1325x1803.png" width="1325" height="1803" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e0f527ff-ca9b-4105-a384-e16c05140f23_1325x1803.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1803,&quot;width&quot;:1325,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:272092,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!UpcI!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0f527ff-ca9b-4105-a384-e16c05140f23_1325x1803.png 424w, https://substackcdn.com/image/fetch/$s_!UpcI!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0f527ff-ca9b-4105-a384-e16c05140f23_1325x1803.png 848w, https://substackcdn.com/image/fetch/$s_!UpcI!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0f527ff-ca9b-4105-a384-e16c05140f23_1325x1803.png 1272w, https://substackcdn.com/image/fetch/$s_!UpcI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0f527ff-ca9b-4105-a384-e16c05140f23_1325x1803.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Example of &#8220;Reasoning Tree&#8221; for a question about fur trader Ferdinand Willard Wentzel&#8217;s family I posed to ChatGPT based on document excerpts I provided.</figcaption></figure></div><p></p><p>This is all extremely important for software developers. I would argue these properties have been the major barriers to moving from the prototype to deployment phase up to now. Developers can &#8220;wow&#8221; at the prototype stage, but the models have just not been accurate enough or reliable enough to generate the trust necessary to actually put them out there.</p><h4>The Case of PearlBot</h4><p>Let me use my own work to provide some concrete examples about how o1 is likely to change the game. Over the past year, I&#8217;ve led the development of PearlBot (named after my smartest cat), a prototype AI research assistant designed to answer queries using a database of open-source 18th-century fur trade records. Building it was a learning experience that involved creating both relational and vector databases for effective keyword and semantic searches. The system employs fine-tuned language models to parse queries, retrieve relevant documents, and generate answers with in-text citations for verification. The prototype appears really impressive when I show it to an audience, but the reality is that we haven&#8217;t made it available yet because our team has a series of verification tests that the system always fails.</p><p>If you ask PearlBot the deliberately vague, open-ended question &#8220;Tell me about Wentzel&#8217;s family&#8221; (a sentence I&#8217;ve now written thousands of times), it searches a database of 7,500 records and retrieves around 15 results. These are then passed to an LLM which is given the question and a system message instructing it to use all the relevant documents, cite them in square brackets, and be as thorough and detailed as possible. Our team has designated nine of the documents as essential, which means that we&#8217;ve decided that in order for a model to pass the test, it needs to use and cite the information from all of those documents. This is because it is what we would expect of a graduate student or a fellow historian given the same task and body of sources. Across hundreds of tests, GPT-4, GPT-4o, Claude Opus, and Sonnet-3.5 all used an average of about 3.5 of the nine required documents in their answers. Sometimes this was as low as 1 but it&#8217;s never exceeded 6.</p><p>Using nine of nine documents is a minimum requirement and it also matters how the model actually interprets, organizes, and uses the texts. Again, even the best models don&#8217;t do well with this. For example, in his diary, Wentzel consistently referred to his wife as &#8220;my girl&#8221; which sometimes confuses models into thinking he was referring to his servant or daughter. They also frequently misunderstand references to the families of other people as references to Wentzel&#8217;s own family: when Wentzel tells a friend to give his best to &#8220;Mrs. McKenzie&#8221;, models often assume she is one of Wentzel&#8217;s relatives.</p><p>The models sometimes use correct information, but then cite it to the wrong sources and also have a hard time re-arranging events into a chronological, logical sequence. They especially have trouble connecting related information and events separated by time and space. For example, one document describes the birth of Wentzel&#8217;s child in the winter of 1805 while another mentions him travelling with his wife and three children in the summer of 1805. The models frequently assume that this means he had four children, rather than drawing the logical inference that the child born in the winter of 1805 is one of the three children seen in the canoe that summer. One of Wentzel&#8217;s &#8220;young boys&#8221; also died in 1808 and although the models always note this fact, they sometimes assume this must have been a fifth child. They have also frequently overlook the significance of another subsequent but oblique reference to the tragedy which betrays that Wentzel remained preoccupied for some time by the death of his son such that he neglected his job, an important insight into his personality.</p><p>Although we have varied the system message and prompt in many different ways, nothing consistently improves performance beyond the limits outlined above. Some prompts get better answers, but no prompt eliminates all of the problems. So as impressive as it is to watch PearlBot write out an answer with verifiable citations, the program has never been close to being ready for primetime. The answers are simply incomplete and too often inaccurate, unreliable, and untrustworthy. Up until now, it has been a good proof of concept but little more.</p><h4>PearlBot meets o1-Preview</h4><p>O1 is the first model that can pass our tests. We&#8217;ve only had a day to try it out, but after integrating it into PearlBot, we tried the question ten times. In those tries, it used the necessary documents an average of 8 times (low of 6 times high of 9), passing the minimum requirement 5 times. That is a massive improvement on its own&#8212;swapping our former max score for the minimum score here&#8212;but its responses have also been much more consistent, not necessarily in form but certainly in content. O1 repeatedly draws the same inferences, connects the same information and events, and makes correct judgements. It has fewer problems, rearranging events into a coherent chronological narrative where one event connects to the next. I have yet to see a hallucination although it sometimes uses the wrong number to refer to a source (less frequently, but it happens).</p><p></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!4gID!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9b30933-1e1b-4c64-bbe4-6dadf7b940dc_1992x2089.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!4gID!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9b30933-1e1b-4c64-bbe4-6dadf7b940dc_1992x2089.png 424w, https://substackcdn.com/image/fetch/$s_!4gID!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9b30933-1e1b-4c64-bbe4-6dadf7b940dc_1992x2089.png 848w, https://substackcdn.com/image/fetch/$s_!4gID!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9b30933-1e1b-4c64-bbe4-6dadf7b940dc_1992x2089.png 1272w, https://substackcdn.com/image/fetch/$s_!4gID!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9b30933-1e1b-4c64-bbe4-6dadf7b940dc_1992x2089.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!4gID!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9b30933-1e1b-4c64-bbe4-6dadf7b940dc_1992x2089.png" width="1456" height="1527" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a9b30933-1e1b-4c64-bbe4-6dadf7b940dc_1992x2089.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1527,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:848386,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!4gID!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9b30933-1e1b-4c64-bbe4-6dadf7b940dc_1992x2089.png 424w, https://substackcdn.com/image/fetch/$s_!4gID!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9b30933-1e1b-4c64-bbe4-6dadf7b940dc_1992x2089.png 848w, https://substackcdn.com/image/fetch/$s_!4gID!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9b30933-1e1b-4c64-bbe4-6dadf7b940dc_1992x2089.png 1272w, https://substackcdn.com/image/fetch/$s_!4gID!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9b30933-1e1b-4c64-bbe4-6dadf7b940dc_1992x2089.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">PearlBot uses o1-Preview to respond to the question &#8220;Tell me about Wentzel&#8217;s Family&#8221; using all nine required documents thus passing our test.</figcaption></figure></div><p>Here is why it feels like a different beast. In one typical run, it wrote:</p><blockquote><p><em>Wentzel's family appears to have been intimately connected with the Indigenous communities among whom he lived and worked. On January 31, 1805, Wentzel joyfully recorded the birth of a son, noting with a mix of pride and astonishment the child's size and resemblance to himself. He wrote that his "girl" had "happily made me a great boy," and marveled at "how a Lilliputian could make such a big male child, nearly as big as myself, and more than 20 inches long."[2327] This passage not only underscores his paternal pride but also hints at his own modest stature, referring to himself humorously as a "Lilliputian."</em></p></blockquote><p>This is an important example because the model had to not only correctly parse that Wentzel&#8217;s &#8220;girl&#8221; was his wife, but also use his joke about the baby&#8217;s size to infer something about Wentzel&#8217;s appearance. It now correctly infers that the child born in January 1805 was likely one of the three children seen in the canoe that summer and correctly notes that the loss of his son in 1808 &#8220;deeply affected him as evidenced by a subsequent entry on March 7, 1808.&#8221; This is not a one off, but it does both things consistently.</p><p>O1 also uses the documents to make original and insightful observations. For example, it writes:</p><blockquote><p><em>The participation of Wentzel's wife in the trading community is further evidenced on October 17th, 1807. Wentzel recorded that "while my wife was across the river, an old man named Croup de Chien sent her four beaver skins and one meat item, though he did not owe anything" ([6099]). This gesture suggests that Wentzel's wife was respected and perhaps held a significant social standing within the local community, receiving gifts even in the absence of debt or obligation. It also illustrates her mobility and active engagement outside the immediate confines of the trading post.&#8217;</em></p></blockquote><p>The model&#8217;s observation that the gift by Croup de Chien to Wentzel&#8217;s wife speaks to her status in the broader community is reasonable but not one we&#8217;ve seen from other models. So too with its observation about her mobility and engagement in the fur trade outside the fort itself.</p><h4>Conclusions</h4><p>The model just feels more mature. It&#8217;s a boring writer but it does the thing you ask it, more consistently and ably. While a lot more testing is needed, it feels like a big step towards a more accurate, reliable, and trustworthy model. To my mind, PearlBot and similar applications will be deployable when they meet or exceed human performance on similar tasks. We are getting very close to that here.</p><p>However, it is far from perfect (and that is OK). It does not follow instructions all that well and if you give it nine relevant documents in a much larger pool of 120 irrelevant documents, its performance degrades significantly. Maybe scaling can improve this in future versions.</p><p>To look forward just a bit, we need to remember that this is an early preview model which, according to OpenAI, performs less ably than the base o1 model. It is also a GPT-4 level model, not a next-gen model like GPT-5. When the next frontier models appear, it is entirely possible that the doubters will be proved right and the limitations of scaling will mean that the exponential progress in LLM capabilities will begin to level off. But if this isn&#8217;t the case, when the o1 regime is applied to the new models we might find ourselves in wholly uncharted waters.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://generativehistory.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Generative History! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Making the Infeasible Practical in Historical Research]]></title><description><![CDATA[Developing AI Research Tools for Historical Research and General Release]]></description><link>https://generativehistory.substack.com/p/making-the-infeasible-practical-in</link><guid isPermaLink="false">https://generativehistory.substack.com/p/making-the-infeasible-practical-in</guid><dc:creator><![CDATA[Mark Humphries]]></dc:creator><pubDate>Fri, 31 May 2024 09:02:01 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!ju30!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f8528ea-c0c9-43a1-acbd-7ab645cef1b2_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ju30!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f8528ea-c0c9-43a1-acbd-7ab645cef1b2_1024x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ju30!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f8528ea-c0c9-43a1-acbd-7ab645cef1b2_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!ju30!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f8528ea-c0c9-43a1-acbd-7ab645cef1b2_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!ju30!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f8528ea-c0c9-43a1-acbd-7ab645cef1b2_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!ju30!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f8528ea-c0c9-43a1-acbd-7ab645cef1b2_1024x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ju30!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f8528ea-c0c9-43a1-acbd-7ab645cef1b2_1024x1024.png" width="1024" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3f8528ea-c0c9-43a1-acbd-7ab645cef1b2_1024x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1671598,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!ju30!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f8528ea-c0c9-43a1-acbd-7ab645cef1b2_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!ju30!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f8528ea-c0c9-43a1-acbd-7ab645cef1b2_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!ju30!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f8528ea-c0c9-43a1-acbd-7ab645cef1b2_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!ju30!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f8528ea-c0c9-43a1-acbd-7ab645cef1b2_1024x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p>In my experience, one of the biggest misconceptions about generative AI is that it must involve a one-on-one conversational interaction with a chatbot. What most people don&#8217;t realize is that generative AI&#8217;s most promising applications have nothing to do with chat interfaces. Instead, they almost all involve using the underlying LLM to do &#8220;thinking&#8221; tasks <em>very</em> quickly. Think: using AI to do hundreds of complex things&#8212;summarization, translation, transcription, keyword categorization, information extraction, etc&#8212;in a few seconds.</p><p>Over the next few weeks, I plan to make a few programs I&#8217;ve been developing available to researchers. These include a program that automates the transcription of handwritten documents with a similar level of accuracy to human transcribers and another that takes notes on archival documents. The idea is that researchers will be able to download either the code or an executable version of the software and customize it to their own purposes. These are not chatbot type research assistants but simple AI-enabled programs that can be easily integrated into any research pipeline. In fact, I think these task specific types of tools focused on scaling-up and speeding up the research process will have a much greater effect on what we do than autonomous research assistants, at least in the near term.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://generativehistory.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://generativehistory.substack.com/subscribe?"><span>Subscribe now</span></a></p><h4>AI Enabled Research Software</h4><p>AI-enabled research software involves sending many <em>things</em> (like text or images) to an Application Programming Interface (API) where the LLM processes it according to a set of instructions. The software then receives the LLM&#8217;s response back and does some more processing before saving it to a text file or storing it in a database. In these types of applications, you&#8217;ll rarely interact directly with the underlying AI engine at heart of the program but it&#8217;s doing all the heavy lifting.</p><p>Let&#8217;s think about what this might look like for historians. Imagine we are writing a medical history and wanted to determine how long it took soldiers in the First World War to be evacuated from the front to England by the type and location of a wound. It&#8217;s a solvable problem as all 620,000 First World War personnel files have been digitized by <a href="https://recherche-collection-search.bac-lac.gc.ca/eng/help/pffww">Library and Archives Canada</a>. But generating the data by hand would take a long time: because only 1 in 4 Canadian soldiers were wounded and there is no master list of who they were, you would need to start with at least 1,600 personnel files chosen at random in order to get the 400 wounded men necessary to construct a basic sample. A year ago, the only way to do this would have been to go through something like 80,000 pages of records one by one, which would have taken many, many weeks.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!oTw6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe575a81a-cc1d-400d-aa67-6b739d9ca707_2783x2332.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!oTw6!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe575a81a-cc1d-400d-aa67-6b739d9ca707_2783x2332.jpeg 424w, https://substackcdn.com/image/fetch/$s_!oTw6!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe575a81a-cc1d-400d-aa67-6b739d9ca707_2783x2332.jpeg 848w, https://substackcdn.com/image/fetch/$s_!oTw6!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe575a81a-cc1d-400d-aa67-6b739d9ca707_2783x2332.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!oTw6!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe575a81a-cc1d-400d-aa67-6b739d9ca707_2783x2332.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!oTw6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe575a81a-cc1d-400d-aa67-6b739d9ca707_2783x2332.jpeg" width="1456" height="1220" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e575a81a-cc1d-400d-aa67-6b739d9ca707_2783x2332.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1220,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:418574,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!oTw6!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe575a81a-cc1d-400d-aa67-6b739d9ca707_2783x2332.jpeg 424w, https://substackcdn.com/image/fetch/$s_!oTw6!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe575a81a-cc1d-400d-aa67-6b739d9ca707_2783x2332.jpeg 848w, https://substackcdn.com/image/fetch/$s_!oTw6!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe575a81a-cc1d-400d-aa67-6b739d9ca707_2783x2332.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!oTw6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe575a81a-cc1d-400d-aa67-6b739d9ca707_2783x2332.jpeg 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!8jFo!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa4c6c67-da0b-42f0-b421-62baf577b9bb_2782x2338.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!8jFo!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa4c6c67-da0b-42f0-b421-62baf577b9bb_2782x2338.jpeg 424w, https://substackcdn.com/image/fetch/$s_!8jFo!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa4c6c67-da0b-42f0-b421-62baf577b9bb_2782x2338.jpeg 848w, https://substackcdn.com/image/fetch/$s_!8jFo!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa4c6c67-da0b-42f0-b421-62baf577b9bb_2782x2338.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!8jFo!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa4c6c67-da0b-42f0-b421-62baf577b9bb_2782x2338.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!8jFo!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa4c6c67-da0b-42f0-b421-62baf577b9bb_2782x2338.jpeg" width="1456" height="1224" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/fa4c6c67-da0b-42f0-b421-62baf577b9bb_2782x2338.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1224,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:386698,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!8jFo!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa4c6c67-da0b-42f0-b421-62baf577b9bb_2782x2338.jpeg 424w, https://substackcdn.com/image/fetch/$s_!8jFo!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa4c6c67-da0b-42f0-b421-62baf577b9bb_2782x2338.jpeg 848w, https://substackcdn.com/image/fetch/$s_!8jFo!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa4c6c67-da0b-42f0-b421-62baf577b9bb_2782x2338.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!8jFo!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa4c6c67-da0b-42f0-b421-62baf577b9bb_2782x2338.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!B12j!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F913efd6e-c9f5-4750-b46d-e2be30c00d02_2766x2324.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!B12j!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F913efd6e-c9f5-4750-b46d-e2be30c00d02_2766x2324.jpeg 424w, https://substackcdn.com/image/fetch/$s_!B12j!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F913efd6e-c9f5-4750-b46d-e2be30c00d02_2766x2324.jpeg 848w, https://substackcdn.com/image/fetch/$s_!B12j!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F913efd6e-c9f5-4750-b46d-e2be30c00d02_2766x2324.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!B12j!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F913efd6e-c9f5-4750-b46d-e2be30c00d02_2766x2324.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!B12j!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F913efd6e-c9f5-4750-b46d-e2be30c00d02_2766x2324.jpeg" width="1456" height="1223" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/913efd6e-c9f5-4750-b46d-e2be30c00d02_2766x2324.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1223,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:348018,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!B12j!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F913efd6e-c9f5-4750-b46d-e2be30c00d02_2766x2324.jpeg 424w, https://substackcdn.com/image/fetch/$s_!B12j!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F913efd6e-c9f5-4750-b46d-e2be30c00d02_2766x2324.jpeg 848w, https://substackcdn.com/image/fetch/$s_!B12j!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F913efd6e-c9f5-4750-b46d-e2be30c00d02_2766x2324.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!B12j!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F913efd6e-c9f5-4750-b46d-e2be30c00d02_2766x2324.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Casualty Form - Active Service, 1918-1919 for my Great-Grandfather, Pte. Osborne Juby, 38th Battalion.  All 620,000 First World War pension files are online and freely available for download. Source: Library and Archives Canada,<a href="http://central.bac-lac.gc.ca/.redirect?app=pffww&amp;id=480327&amp;lang=eng"> RG 150, Accession 1992-93/166, Box 4985 - 34.</a></figcaption></figure></div><p>Thanks to <a href="https://openai.com/index/hello-gpt-4o/">GPT-4o</a>, the new &#8220;omni&#8221; model recently released by OpenAI, it&#8217;s now possible to do this in a few minutes. GPT-4o is faster and much cheaper than its predecessors GPT-4 and GPT-4-Turbo, but it is also far better at analyzing images. It can often transcribe handwriting nearly verbatim and can also decipher documents which feature a mixture of handwritten and typed text.</p><p>It has it&#8217;s limitations too. Sometimes the AI will miss something or confabulate, it is true. But in my experience, the best LLMs make no more mistakes than humans including myself and student research assistants. Its mistakes are often weirder, but not more frequent.</p><h4>First Steps</h4><p>Whether a human or LLM tackles our First World War research question, it will involve creating some sort of database or spreadsheet. In this case, we&#8217;d probably create a spreadsheet with seven columns: Name, Wounded (Yes/No), Date of Injury, Type of Injury (GSW/SW/Gas etc), Date of First Hospital Admission, and Hyperlink to the Document. This last column will let us verify the accuracy of the data afterwards by providing a quick link to the underlying document. When we populate the database with the data, we should have 1,600 rows, one for each soldier; around 400 of those should have relevant data. We&#8217;d then crunch the numbers and get our answer.</p><p>So let the analysis begin! If I tried to do this myself, my experience suggests I could probably manage to skim 8 pages a minute, an average that would include flipping past dozens of obviously irrelevant pages very quickly while spending several minutes parsing unclear handwriting, resolving contradictions, and entering the data in other places. That means it would take roughly 20 eight-hour days to get my results.</p><p>By automating the research process, we could write a program that gets this down to a couple hours at most. First, we would need to write a conventional Python script to choose the 1,600 personnel files from the LAC website at random, download them, extract the images, and then send them to an LLM along for analysis. If this sounds complicated, you should understand that a novice programmer could write the code to do the first three tasks in only a few minutes. The last part requires a bit more thought and is where the AI comes in, not because the coding is any more complicated (it isn&#8217;t), but because LLMs don&#8217;t work in a conventional way.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!MPLX!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F69a0de04-6f21-40e2-844a-d4d15d7d3fa8_1389x1717.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!MPLX!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F69a0de04-6f21-40e2-844a-d4d15d7d3fa8_1389x1717.png 424w, https://substackcdn.com/image/fetch/$s_!MPLX!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F69a0de04-6f21-40e2-844a-d4d15d7d3fa8_1389x1717.png 848w, https://substackcdn.com/image/fetch/$s_!MPLX!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F69a0de04-6f21-40e2-844a-d4d15d7d3fa8_1389x1717.png 1272w, https://substackcdn.com/image/fetch/$s_!MPLX!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F69a0de04-6f21-40e2-844a-d4d15d7d3fa8_1389x1717.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!MPLX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F69a0de04-6f21-40e2-844a-d4d15d7d3fa8_1389x1717.png" width="1389" height="1717" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/69a0de04-6f21-40e2-844a-d4d15d7d3fa8_1389x1717.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1717,&quot;width&quot;:1389,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:220003,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!MPLX!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F69a0de04-6f21-40e2-844a-d4d15d7d3fa8_1389x1717.png 424w, https://substackcdn.com/image/fetch/$s_!MPLX!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F69a0de04-6f21-40e2-844a-d4d15d7d3fa8_1389x1717.png 848w, https://substackcdn.com/image/fetch/$s_!MPLX!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F69a0de04-6f21-40e2-844a-d4d15d7d3fa8_1389x1717.png 1272w, https://substackcdn.com/image/fetch/$s_!MPLX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F69a0de04-6f21-40e2-844a-d4d15d7d3fa8_1389x1717.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">OpenAI Playground: In AI parlance, a &#8220;playground&#8221; provides a convenient interface with the API (in this case GPT-4o) to test out how various prompts work. Setting the temperature to 0.0 makes the response more deterministic. I am showing the playground version of our prompt to visualize the process of communicating with an API. In this case, GPT-4o is correct: my Great-Grandfather Osborne Juby was wounded on 29 September 1918 by a gunshot and was admitted to hospital the same day. (Source: http://platform.openai.com)</figcaption></figure></div><p></p><h4>Envisioning What the AI Does (and How it Does it)</h4><p>To get a sense of what I mean, imagine for a minute that the LLM is a remotely located human research assistant with whom you can only communicate via email and a shared Dropbox folder. Unfortunately, the Dropbox folder is also very small and can only hold somewhere around 5 high-res images at a time. It cannot be expanded so you need to send the assistant materials in small batches. To complicate matters, although our research assistant is highly capable, they have a condition which prevents them from forming memories. After they complete each task they email you the response and them promptly forget what they just did. This poses something of a problem as each personnel file is around 50 pages long.</p><p>Even within these constraints, LLMs can do a lot of useful things for researchers, but we need to understand those limitations and play to their strengths, namely speed and efficiency. There are lots of ways we could setup our program, but the simplest would be to send each personnel file to the LLM a few images at a time along with a series of instructions demanding a response that we can process automatically in python.</p><p>When communicating with an LLM via an API, which operates like the email/Dropbox communications above, we can send our virtual research assistant both textual instructions (often as a &#8220;System Message&#8221;) along with the images for it to analyze. In our case, we might write something like:</p><blockquote><p>&#8220;You are a historical research assistant that reads First World War personnel files to look for specific information about where and how a solider might have been wounded and when he was first treated in hospital.</p><p>If the attached document indicates that the soldier was wounded, specify the date, type of injury, and the date of the first hospital admission for that wound.</p><p>Think step by step and verbalize your thought process on a &lt;Notepad&gt;&lt;/Notepad&gt;. No one will see anything you write between the tags.</p><p>When you are ready, write &#8220;Final Answer:&#8221; and then format the rest of your response as follows:</p><p>Wounded: &lt;Yes/No&gt;</p><p>Date: &lt;YYYY-MM-DD/N/A&gt;</p><p>Type of Injury: &lt;GSW/SW/Gas etc/N/A&gt;</p><p>Date of First Hospital Admission: &lt;Date/N/a&gt;</p><p>If there is more than one wound, separate your answers with *****"</p></blockquote><p>In these instructions, we clearly explain the task, tell the LLM what to do, and ask it to &#8220;think step by step&#8221; to explain its &#8220;thought process&#8221;. Strange indeed but research has shown that this increases accuracy and reduces hallucinations significantly, especially when reasoning over content. We also specify the specific form we want the answer to take because we can then parse the response in Python, meaning we use code to extract specific parts (we could do this even more effectively through a markup language called JSON but that is a bit more complicated than I want to get here).</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://generativehistory.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://generativehistory.substack.com/subscribe?"><span>Subscribe now</span></a></p><h4>Costs and Time</h4><p>After we get the response back, we would use python (or another LLM) to automatically collate the answers, remove duplicates, and populate the spreadsheet. If this sounds time consuming or complicated, once it is setup it&#8217;s extremely quick: OpenAI&#8217;s API allows 10,000 requests per minute (meaning images processed in our case). At around 1,000 tokens (around 750 words with each image being the equivalent of 500 words) a request and $5.00 per million tokens, <a href="https://openai.com/api/pricing/">it would cost</a> around $400.</p><p>To reduce the cost and time involved, we would probably add another step to the process. By sending each page to a smaller, less capable but much cheaper model like <a href="https://www.anthropic.com/news/claude-3-family">Anthropic&#8217;s Claude Haiku</a> first, we could then send only the most relevant documents to the more expensive GPT-4o. While Haiku is not as good as GPT-4o at parsing handwriting and extrapolating from information, it is good at doing things like recognizing whether a given document is an example of one type of form or another or just determining whether it contains information relevant to a given topic.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!lI5Y!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c21c17f-3376-4d7d-a722-aff8251310c2_1390x1870.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!lI5Y!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c21c17f-3376-4d7d-a722-aff8251310c2_1390x1870.png 424w, https://substackcdn.com/image/fetch/$s_!lI5Y!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c21c17f-3376-4d7d-a722-aff8251310c2_1390x1870.png 848w, https://substackcdn.com/image/fetch/$s_!lI5Y!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c21c17f-3376-4d7d-a722-aff8251310c2_1390x1870.png 1272w, https://substackcdn.com/image/fetch/$s_!lI5Y!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c21c17f-3376-4d7d-a722-aff8251310c2_1390x1870.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!lI5Y!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c21c17f-3376-4d7d-a722-aff8251310c2_1390x1870.png" width="1390" height="1870" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6c21c17f-3376-4d7d-a722-aff8251310c2_1390x1870.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1870,&quot;width&quot;:1390,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:353855,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!lI5Y!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c21c17f-3376-4d7d-a722-aff8251310c2_1390x1870.png 424w, https://substackcdn.com/image/fetch/$s_!lI5Y!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c21c17f-3376-4d7d-a722-aff8251310c2_1390x1870.png 848w, https://substackcdn.com/image/fetch/$s_!lI5Y!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c21c17f-3376-4d7d-a722-aff8251310c2_1390x1870.png 1272w, https://substackcdn.com/image/fetch/$s_!lI5Y!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c21c17f-3376-4d7d-a722-aff8251310c2_1390x1870.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Anthropic Playground: These interactions with the Claude 3.0 Haiku API are in chatbot form to illustrate the process. The first document is indeed a &#8220;Casualty Form-Active Service&#8221;, the other two are not. In our program, this would all happen behind the scenes and all the user would see is a progress bar. We would use Python code to extract the &#8220;Yes/No&#8221; answer and route the image accordingly. (Source: https://console.anthropic.com)</figcaption></figure></div><p>This sort of pre-processing would involve giving instructions to Haiku about which documents in each personnel file are relevant to our research question. In our case, we could probably restrict the analysis to the Casualty Form &#8211; Active Service which should contain all the relevant information. In each API call, we would provide Haiku with an example of the form along with the image in question and then give it a simple question to answer: is this document a Casualty Form &#8211; Active Service similar to the examples appended below? If so, our code would send it to GPT-4o for a more thorough analysis, and if not we&#8217;d move on to the next image. </p><p><a href="https://www.anthropic.com/api">Haiku costs </a>only $0.25 per <em>million tokens</em> and given the nature of the task we could also use a low-res version of the images. In all, it would cost around $20 to have Haiku triage our images and at 4,000 requests per minute, it would be done in 20 minutes. In the end, we&#8217;d have a much more manageable 6,500 images left to send to GPT-4o at a cost of $32.50, which would take another 5 minutes to process. So in 30 minutes and for $50 we&#8217;d have an answer to a question that would have taken us many weeks to work through.</p><p>This type of automation promises to speed up the research process without affecting the underlying dynamics of what we actually do as historians. With AI, gathering the evidence might happen faster and cost much less than it does today, but the process of analyzing and writing-up the results will largely remain unchanged.</p><h4>Revolutions Real and Imagined</h4><p>As frightening as this may sound to some of us, I actually think it will be similar to the way digital photography changed the research process. In the old days (which I vaguely recall from my first research trips to Ottawa more than two decades ago), researchers had to be physically present to go through records. Pencil and paper only, no laptops, no cameras. You could flag records for photocopying, but it was $0.50 a page. </p><p>The advent of digital photography a few years later suddenly meant that you could &#8220;shoot first and ask questions later&#8221;, capturing thousands of images in a few days and then spend weeks pouring over them at home where you did not have to pay extra room and board.</p><p>But as ubiquitous as this approach is today, around 2005 I was scolded by a senior historian in the LAC reading room for using a digital camera. In his view, the job of the student was to immerse themselves in the records for months at a time, carefully reading through every page <em>in the archives</em>. He saw my camera as an appalling shortcut. Try as I might, I could not make him understand that that was, in fact, exactly what I was doing! </p><p>I didn&#8217;t have the money to spend &#8220;months at a time&#8221; in Ottawa when I lived in Southwestern Ontario, but by making short visits a few times a year to capture thousands of images, I was effectively doing exactly what he wanted me to do, just at home on my computer. The alternative was either taking-on a much narrower topic (which he would not have liked either) or not doing the research at all.</p><p>Digital photography made historical research more accessible. At the same time, I think it also provided the opportunity to go much deeper into the sources as it allows us to easily re-read and re-check the evidence based on the accumulation of new information, to re-purpose large bodies of materials for different projects, or to share sources with students.</p><p>I think the effects of AI will be similar because to me, the main thing it will do is alter the scope of the possible. Instead of esoteric questions about the transit times of wounded Great War soldiers, imagine getting an LLM to summarize the contents of each document on a reel of microfilm, the pages of a diary, a series of letters, or all those photographs you took last summer at the archives. We can now use it to produce transcriptions&#8212;even of handwritten texts&#8212;that are as accurate as those done by human transcription services (stay tuned for a future post on this and that software release). It will let us organize our materials more efficiently, draw out hidden connections and linkages, and do deeper research.</p><p>Fear not: the use of AI to analyze digital records won&#8217;t absolve us from actually reading and knowing the contents of archival documents; that will always remain a constant for our profession whether we capture a digital images or read the paper versions in the reading room. The reality is that it is not really possible to effectively incorporate AI into the research process anyway without having well-grounded, domain specific knowledge of the material. Otherwise, what would your instructions to the LLM actually look like?</p><p>Stay tuned as I role out some example programs over the coming weeks that will allow you to try some of these approaches for yourself. The possibilities are endless, terrifying, and intriguing all at once.</p><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://generativehistory.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Generative History! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[Microsoft Copilot Necessitates Some Tough Conversations]]></title><description><![CDATA[The integration of generative AI directly into word processors will give rise to a new, acceptable form of synthesized AI-human writing and universities are not prepared.]]></description><link>https://generativehistory.substack.com/p/microsoft-copilot-necessitates-some</link><guid isPermaLink="false">https://generativehistory.substack.com/p/microsoft-copilot-necessitates-some</guid><dc:creator><![CDATA[Mark Humphries]]></dc:creator><pubDate>Fri, 16 Feb 2024 10:30:48 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!nIx2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc61527d7-140b-4bd9-a221-bbe7912a9288_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!nIx2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc61527d7-140b-4bd9-a221-bbe7912a9288_1024x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!nIx2!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc61527d7-140b-4bd9-a221-bbe7912a9288_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!nIx2!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc61527d7-140b-4bd9-a221-bbe7912a9288_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!nIx2!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc61527d7-140b-4bd9-a221-bbe7912a9288_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!nIx2!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc61527d7-140b-4bd9-a221-bbe7912a9288_1024x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!nIx2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc61527d7-140b-4bd9-a221-bbe7912a9288_1024x1024.png" width="1024" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c61527d7-140b-4bd9-a221-bbe7912a9288_1024x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1476069,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!nIx2!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc61527d7-140b-4bd9-a221-bbe7912a9288_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!nIx2!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc61527d7-140b-4bd9-a221-bbe7912a9288_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!nIx2!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc61527d7-140b-4bd9-a221-bbe7912a9288_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!nIx2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc61527d7-140b-4bd9-a221-bbe7912a9288_1024x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p>For more than a year now, <a href="https://www.oneusefulthing.org/">people </a>have been writing about how ChatGPT will make universities rethink the way we teach and assess our students. But while most institutions have <a href="https://higheredstrategy.com/ai-observatory-home/">devised policies around generative AI</a> of some sort, they&#8217;ve also tended to put off the tough conversations about what we do and why. I think this was feasible while &#8220;generative AI&#8221; lurked behind foreboding external logins, requiring people to consciously open an OpenAI account in order to use those tools. But I don&#8217;t think that is possible anymore.</p><p>On 15 January, Microsoft finally made <a href="https://www.microsoft.com/en-us/microsoft-365/enterprise/copilot-for-microsoft-365">Copilot for Office</a> available to individuals and families including students. Google is also doing the same, <a href="https://www.technologyreview.com/2024/02/08/1087911/googles-gemini-is-now-in-everything-heres-how-you-can-try-it-out/">integrating Gemini into its own productivity suite</a>. Since Microsoft announced (somewhat prematurely) the integration of generative AI into Office last year, I&#8217;ve been anxious to see it in action. <a href="https://generativehistory.substack.com/p/the-paradox-of-embracing-ai-in-higher?utm_source=%2Fsearch%2Fcopilot&amp;utm_medium=reader2">I&#8217;ve argued</a> that these integrations will be what actually make AI usage ubiquitous. Now that AI writing is just a button click away in Word or Google Docs, it will quickly become as ordinary and mundane as spell check. From here on out, anyone still hoping to &#8220;catch&#8221; AI writing, avoid it, or pretend it isn&#8217;t happening is going to be out of luck. </p><p>My goal in this post is to show how Copilot works in Office and to explain why its very banality will be the thing that makes us actaully begin to redefine concepts such as authorship, authenticity, and plagarism, as well as reassess the utility of written assessments. This also means that all those policies everyone drafted last fall are going to have to change.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://generativehistory.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://generativehistory.substack.com/subscribe?"><span>Subscribe now</span></a></p><h3>Introducing Copilot</h3><p>First of all, let&#8217;s talk about the name (because it&#8217;s confusing) then what the tool actually does. Copilot is now Microsoft&#8217;s branded interface for any of its AI tools and applications, all of which rely on OpenAI&#8217;s GPT-3.5, GPT-4, and GPT-4-Turbo to do their thing. <a href="https://copilot.microsoft.com/">Bing Chat has been renamed Copilot</a> (RIP Sydney) and is basically a streamlined, more constrained version of ChatGPT with a similar web-based chat interface with free and paid (pro) versions. Here is the confusing part: Copilot is also the name given to Microsoft&#8217;s AI tools in its <a href="https://adoption.microsoft.com/en-us/copilot/">productivity software</a>, both in an integrated chatbot form and as standalone word widgets that co-exist with the chart creators and image insertion tools you&#8217;ll be familiar with. Basically in Microsoft products, Copilot now just means AI. Google is doing a similar thing, <a href="https://blog.google/products/gemini/bard-gemini-advanced-app/">rebranding Bard (their AI chatbot) as Gemini</a> which is also being integrated into apps like Google Docs.</p><p>Purchasing or activating Copilot is a bit confusing too. It&#8217;s been available to enterprise and business users for months, but now individuals can access it by <a href="https://www.microsoft.com/en-ca/store/b/copilotpro">purchasing a subscription to Copilot Professional for $27.00 CAD per month</a>. This subscription includes unlimited and priority access to the Copilot chatbot and image generator, powered by Microsoft&#8217;s finetuned version of GPT-4, GPT-4-Turbo, and Dalle-3. Once subscribed to Copilot Pro, users can enable Copilot in both offline and web-based Office apps (Word, PowerPoint, Excel, Outlook, etc) through their Microsoft account page.</p><h3>Using Copilot in Word</h3><p>Copilot integrates generative AI into Word in three main ways: it adds a chatbot widget similar to ChatGPT into Word itself; it offers a Copilot editor tool that rewrites existing sentences or creates tables; and a generation function that proactively and earnestly offers to draft whole documents whenever you open a document or write new paragraphs whenever you start a new line of text. The functionality of these tools overlap somewhat and can be customized through dropdown menus to select different writing styles or change the content via natural text prompting (example prompts are provided).</p><h4>CoPilot the Chatbot</h4><p>The chatbot interface will be the most familiar of the tools in that it is exactly what we have come to expect from generative AI chatbots, only this one lives in Word. A Copilot button on the Home tab of the toolbar opens a sidebar chat window where Copilot will &#8220;chat, respond to your questions, and help you with writing and summarizing [the active] document.&#8221;</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!9dvm!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1d246b84-5dbf-4824-a97b-6ab27568cfc4_1928x1252.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!9dvm!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1d246b84-5dbf-4824-a97b-6ab27568cfc4_1928x1252.png 424w, https://substackcdn.com/image/fetch/$s_!9dvm!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1d246b84-5dbf-4824-a97b-6ab27568cfc4_1928x1252.png 848w, https://substackcdn.com/image/fetch/$s_!9dvm!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1d246b84-5dbf-4824-a97b-6ab27568cfc4_1928x1252.png 1272w, https://substackcdn.com/image/fetch/$s_!9dvm!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1d246b84-5dbf-4824-a97b-6ab27568cfc4_1928x1252.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!9dvm!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1d246b84-5dbf-4824-a97b-6ab27568cfc4_1928x1252.png" width="1456" height="945" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1d246b84-5dbf-4824-a97b-6ab27568cfc4_1928x1252.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:945,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:208394,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!9dvm!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1d246b84-5dbf-4824-a97b-6ab27568cfc4_1928x1252.png 424w, https://substackcdn.com/image/fetch/$s_!9dvm!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1d246b84-5dbf-4824-a97b-6ab27568cfc4_1928x1252.png 848w, https://substackcdn.com/image/fetch/$s_!9dvm!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1d246b84-5dbf-4824-a97b-6ab27568cfc4_1928x1252.png 1272w, https://substackcdn.com/image/fetch/$s_!9dvm!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1d246b84-5dbf-4824-a97b-6ab27568cfc4_1928x1252.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Microsoft is eager to get you to use Copilot chat which can be accessed via a button on the &#8220;Home&#8221; tab (top right), by right clicking the document and using the tools menu (centre), or via the formatting menu (left)</figcaption></figure></div><p>In practice, this means you can ask GPT-4 to do tasks or answer questions about your active document, such as to provide a summary, generate a list of action items, or things like &#8220;how many times did I use the word &#8216;copilot&#8217; in my document&#8221;. It can also show you how to do things like insert a footnote or a table of contents, but it can&#8217;t actually <em>act</em> as an agent and do those things for you which is annoying. Personally, this was really the only thing I was hoping for from Copilot: it would be nice to be able to cut and paste in formatting guidelines and have Copilot set line spacing, margins, tabs, and font size automatically. That will surely come, but sadly not yet.</p><p>Copilot seems to automatically assume that any questions you ask it are about the active document. This means that while Word has normally been thought of as a production tool, you can now use it to analyze and understand documents created by third parties including journal articles, reports, and notes. You can, of course, also ask it to do &#8220;normal&#8221; ChatGPT things like explain how to cite sources in the Chicago Style or provide a list of five sources on the history of Canadian Confederation. When it uses external sources, it tells you that the answer did not come from the document. There is also a &#8220;Copy&#8221; button in the response window that allows you to easily put any text from the chat window into your word document.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!SmCF!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b6e9c3c-b9cb-4247-b851-f2dd64479516_1923x1597.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!SmCF!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b6e9c3c-b9cb-4247-b851-f2dd64479516_1923x1597.png 424w, https://substackcdn.com/image/fetch/$s_!SmCF!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b6e9c3c-b9cb-4247-b851-f2dd64479516_1923x1597.png 848w, https://substackcdn.com/image/fetch/$s_!SmCF!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b6e9c3c-b9cb-4247-b851-f2dd64479516_1923x1597.png 1272w, https://substackcdn.com/image/fetch/$s_!SmCF!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b6e9c3c-b9cb-4247-b851-f2dd64479516_1923x1597.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!SmCF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b6e9c3c-b9cb-4247-b851-f2dd64479516_1923x1597.png" width="1456" height="1209" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0b6e9c3c-b9cb-4247-b851-f2dd64479516_1923x1597.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1209,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:328246,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!SmCF!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b6e9c3c-b9cb-4247-b851-f2dd64479516_1923x1597.png 424w, https://substackcdn.com/image/fetch/$s_!SmCF!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b6e9c3c-b9cb-4247-b851-f2dd64479516_1923x1597.png 848w, https://substackcdn.com/image/fetch/$s_!SmCF!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b6e9c3c-b9cb-4247-b851-f2dd64479516_1923x1597.png 1272w, https://substackcdn.com/image/fetch/$s_!SmCF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b6e9c3c-b9cb-4247-b851-f2dd64479516_1923x1597.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">The Copilot button appears in the left hand margin when you select text. It automatically generates 3 suggested versions which you can click through, regenerate, or customized to one of five styles.</figcaption></figure></div><p>Clearly, the focus of the chatbot is to encourage users to use it to edit, add to, change, and interact with large documents, something that would have required a lot of cutting and pasting to an external AI like ChatGPT in the past. Unlike external editors, though, <a href="https://learn.microsoft.com/en-us/microsoft-365-copilot/microsoft-365-copilot-privacy">Microsoft guarantees</a> that &#8220;Copilot for Microsoft 365 is compliant with our existing privacy, security, and compliance commitments to Microsoft 365 commercial customers, including the General Data Protection Regulation (GDPR) and European Union (EU) Data Boundary&#8221; and that &#8220;prompts, responses, and data accessed through Microsoft Graph aren't used to train foundation LLMs, including those used by Microsoft Copilot for Microsoft 365.&#8221; In other words, this is going to make it both easier and safer for businesses and ordinary people to do things with AI involving sensitive documents or private information.</p><h4>Copilot the Editor&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</h4><p>Copilot&#8217;s in-line suggestions in the editor are designed to be intuitive to use, similar to the way familiar tools like autocorrect, grammar check, and spell check work. When a user selects text, a small &#8220;Copilot&#8221; icon appears in the left-hand margin which provides options to rewrite the selected text or generate a table from it. If you click &#8220;rewrite&#8221;, a widget window appears with three new and different AI generated versions of the sentence to choose from as well as options to regenerate it in various styles of writing including casual, professional, or imaginative. Users can either replace the existing text with one of these AI generated options, paste the AI text below the original text, or hit cancel.</p><p>Generating tables works in a similar way. User can select a paragraph and the AI will essentially try and visualize it as a table by rewriting the key information to fit into rows and columns. Obviously, this works better for some types of text/data than for others, but it is surprisingly versatile and adaptable. A small window allows users to modify the table through natural language, asking the AI to do things like &#8220;remove the top row&#8221;, &#8220;merge columns&#8221;, or &#8220;add a row.&#8221;</p><h4>Copilot the Writer</h4><p>Copilot&#8217;s generation function will be a bit scary for anyone that writes for a living. With copilot enabled, whenever you open a new word doc&#8212;and before you start typing&#8212;a small window appears captioned &#8220;Draft with Copilot&#8221; containing a small textbox. It actively encourages users to explain what it is they are planning to do in this new, blank document. And if you type something like &#8220;I need to write an essay on the origins of the First World War&#8221; it thinks for a moment and then generates a cogent paper complete with headings and a space to insert your name. You then have the option of accepting the AI generated text &#8220;as is&#8221;, regenerating it in a different style, or modifying it with natural language commands like &#8220;change the bullet points into paragraphs&#8221; or &#8220;add another paragraph to the conclusion.&#8221;</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!hUsK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7925fc5-7424-4dd9-aff9-489e76f6acaa_1907x1183.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!hUsK!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7925fc5-7424-4dd9-aff9-489e76f6acaa_1907x1183.png 424w, https://substackcdn.com/image/fetch/$s_!hUsK!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7925fc5-7424-4dd9-aff9-489e76f6acaa_1907x1183.png 848w, https://substackcdn.com/image/fetch/$s_!hUsK!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7925fc5-7424-4dd9-aff9-489e76f6acaa_1907x1183.png 1272w, https://substackcdn.com/image/fetch/$s_!hUsK!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7925fc5-7424-4dd9-aff9-489e76f6acaa_1907x1183.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!hUsK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7925fc5-7424-4dd9-aff9-489e76f6acaa_1907x1183.png" width="1456" height="903" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f7925fc5-7424-4dd9-aff9-489e76f6acaa_1907x1183.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:903,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:143098,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!hUsK!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7925fc5-7424-4dd9-aff9-489e76f6acaa_1907x1183.png 424w, https://substackcdn.com/image/fetch/$s_!hUsK!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7925fc5-7424-4dd9-aff9-489e76f6acaa_1907x1183.png 848w, https://substackcdn.com/image/fetch/$s_!hUsK!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7925fc5-7424-4dd9-aff9-489e76f6acaa_1907x1183.png 1272w, https://substackcdn.com/image/fetch/$s_!hUsK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7925fc5-7424-4dd9-aff9-489e76f6acaa_1907x1183.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">When you open a new document, Copilot pops-up, ready to &#8220;help&#8221; you &#8220;write&#8221; your document</figcaption></figure></div><p>But here is where it gets really frightening. This same generation function is also available whenever you start a new paragraph or hit enter to a new line, allowing users to expand and build a paper a few paragraphs at a time. You can also ask it to do things like &#8220;add 1000 words&#8221; or simply add another paragraph (or ten), quickly expanding an 800 word paper into a 5,000 word essay. I quickly tried this with the paper on the origins of the First World War and had a half-decent B paper assembled in a few minutes. It&#8217;s also important to note that this same functionality will allow users to open an existing paper&#8212;or copy and paste text in from a variety of sources&#8212;and have Word automatically rewrite it, all the while keeping any existing references intact and attached to the correct areas of the text.</p><p>While it has always been possible to do something similar in ChatGPT, the integration of AI generation directly into Word makes the process both more intuitive as well as much simpler and faster. Previously, one would have had to do so a lot of copying, pasting, and convoluted prompting between a Word document and ChatGPT. Not anymore.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!d052!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5795def0-6821-4472-a6d2-e25d34702abf_1914x2021.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!d052!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5795def0-6821-4472-a6d2-e25d34702abf_1914x2021.png 424w, https://substackcdn.com/image/fetch/$s_!d052!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5795def0-6821-4472-a6d2-e25d34702abf_1914x2021.png 848w, https://substackcdn.com/image/fetch/$s_!d052!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5795def0-6821-4472-a6d2-e25d34702abf_1914x2021.png 1272w, https://substackcdn.com/image/fetch/$s_!d052!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5795def0-6821-4472-a6d2-e25d34702abf_1914x2021.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!d052!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5795def0-6821-4472-a6d2-e25d34702abf_1914x2021.png" width="1456" height="1537" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5795def0-6821-4472-a6d2-e25d34702abf_1914x2021.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1537,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:307686,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!d052!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5795def0-6821-4472-a6d2-e25d34702abf_1914x2021.png 424w, https://substackcdn.com/image/fetch/$s_!d052!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5795def0-6821-4472-a6d2-e25d34702abf_1914x2021.png 848w, https://substackcdn.com/image/fetch/$s_!d052!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5795def0-6821-4472-a6d2-e25d34702abf_1914x2021.png 1272w, https://substackcdn.com/image/fetch/$s_!d052!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5795def0-6821-4472-a6d2-e25d34702abf_1914x2021.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Thanks Copilot!: In addition to quickly writing my essay on the First World War, Copilot kindly indicated a place to insert my name at the top!</figcaption></figure></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://generativehistory.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://generativehistory.substack.com/subscribe?"><span>Subscribe now</span></a></p><h3>Implications of Copilot in Word</h3><p>Together Microsoft and Google <a href="https://explodingtopics.com/blog/google-workspace-stats">control 96%</a> of the market for office productivity tools. So make no mistake about it: Copilot and Gemini will quickly normalize a new, synthetic form of AI-human writing whether we like it or not. Of course, people have already been doing this for some time with ChatGPT, but it&#8217;s been mainly in secret, in large part because it clearly fits traditional definitions of plagiarism. But when a respected, button-down company like Microsoft says its ok, it will just become a part of how people write and those definitions will quickly change. Again, I am not saying this is a good thing, but from a purely practical point of view if you&#8217;ve been running from AI, you&#8217;ve just hit a dead end.</p><p>For just over a year now, academics and university administrators have been wringing their hands about how to respond to generative AI. Most institutions have come up with well-meaning guidelines that try to address immediate issues like plagiarism, privacy, and information security while avoiding larger existential questions about the utility of teaching and evaluating writing in the AI era. In effect, we have been trying to adapt existing policies to encompass AI writing, pretending that it does not represent something wholly different both from a technical and cultural point of view. But Copilot and Gemini blur all the lines that universities have been trying to draw around the use of generative AI because it is something new. Now we have to contend with it.</p><p>At my own institution, like many universities, we have <a href="https://students.wlu.ca/academics/academic-integrity/generative-ai-guidelines.html">told students</a> that they can only use generative AI tools if their instructor allows them to do so&#8212;and if it is not explicitly allowed in the syllabus, they will have committed academic misconduct. This might have seemed logical in October, but now it won&#8217;t work. Word is now the most capable generative AI writing tool available. While keeping in mind that students can buy Word at a discount from the university and that our email system is run via Microsoft Outlook, using that program could now be considered academic misconduct.</p><p>According to our guidance, students must also explicitly cite anything that they create with generative AI, but what does that mean for text generated at least in part by Word itself? How much AI text is too much and how little is acceptable? If a student right clicks a grammatical suggestion in word and accepts a revision that is OK. But what if the exact same revision is suggested via Copilot? Even more to the point, what happens when copilot, spell-check and grammar-check are integrated into one universal proofing tool as they almost certainly will be? We have to ask ourselves: what exactly are we trying to control and why?</p><p>These are all highly technical and pedantic questions but they illustrate a larger problem: we will either need to decide to constantly redraft policies in order to micromanage an ever growing array of AI tools and use-cases, or we need to adjust our definitions of authorship, plagiarism, and authenticity. The first option seems like a really depressing and pointless battle but I have no idea how to go about the second. But what I do know is that we can no longer avoid the tough conversations.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://generativehistory.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Generative History! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Growing the Knowledge Gap: OpenAI's New Updates Are Important for Researchers in the Arts]]></title><description><![CDATA[Generative AI technology continues to evolve at breakneck speed... but many people and institutions, including universities, risk being left behind]]></description><link>https://generativehistory.substack.com/p/growing-the-knowledge-gap-openais</link><guid isPermaLink="false">https://generativehistory.substack.com/p/growing-the-knowledge-gap-openais</guid><dc:creator><![CDATA[Mark Humphries]]></dc:creator><pubDate>Thu, 09 Nov 2023 11:00:36 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!4bPd!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F459745de-20f4-4d4e-969a-dea6775c2a5f_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!4bPd!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F459745de-20f4-4d4e-969a-dea6775c2a5f_1024x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!4bPd!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F459745de-20f4-4d4e-969a-dea6775c2a5f_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!4bPd!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F459745de-20f4-4d4e-969a-dea6775c2a5f_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!4bPd!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F459745de-20f4-4d4e-969a-dea6775c2a5f_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!4bPd!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F459745de-20f4-4d4e-969a-dea6775c2a5f_1024x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!4bPd!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F459745de-20f4-4d4e-969a-dea6775c2a5f_1024x1024.png" width="1024" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/459745de-20f4-4d4e-969a-dea6775c2a5f_1024x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1472102,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!4bPd!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F459745de-20f4-4d4e-969a-dea6775c2a5f_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!4bPd!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F459745de-20f4-4d4e-969a-dea6775c2a5f_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!4bPd!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F459745de-20f4-4d4e-969a-dea6775c2a5f_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!4bPd!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F459745de-20f4-4d4e-969a-dea6775c2a5f_1024x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>On Tuesday, <a href="https://openai.com/blog/new-models-and-developer-products-announced-at-devday">OpenAI released several key updates</a> to its Large Language Models (LLMs) that made waves in the AI community. At their first <a href="https://openai.com/blog/new-models-and-developer-products-announced-at-devday">Developer's Day conference</a>, the company that created ChatGPT less than a year ago unveiled something akin to the Apple App Store to host tools that will allow ChatGPT to do all sorts of new and interesting things like connect to your OneDrive or become your travel agent. This follows quickly on the heels of last week's announcement that allowed users to upload whole PDFs and images seamlessly into the same model for processing.</p><p>But the most important developments were the technical updates aimed at developers using the company&#8217;s APIs. Sam Altman, the company's CEO, announced a new "turbo" version of their most advanced LLM that has a 128,000-token context window, meaning that it can process a whopping 92,000 words of text at a time. OpenAI also significantly increased the speed at which GPT-4 generates text while slashing prices on API calls by about 66%. They also rolled out a new vision API which allows developers to send huge numbers of images for automatic processing and analysis. After trying out the new APIs, I believe that for historians and those working in the social sciences and humanities, these will prove to be the most significant developments in AI since the release of GPT-4 back in March.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://generativehistory.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://generativehistory.substack.com/subscribe?"><span>Subscribe now</span></a></p><h3>Why You Should Care about Context</h3><p>On Tuesday, OpenAI updated its flagship GPT-4 so that it now has knowledge about the world up to April 2023. Most importantly, though, OpenAI's update effectively allows users to pass a full book or maybe seven academic articles at a time to a much faster and much cheaper version of GPT-4. This matters because it means that we are getting really close to the point when you will effectively be able to throw an entire customized library at ChatGPT and ask it to find the answer to a specific question. </p><p>True, Claude has had a 100k context window for some time, but Anthropic&#8217;s model is not as capable as GPT-4 when it comes to complex reasoning, and fewer people have access to it. <em>Those of us in Canada still don&#8217;t have access to Claude at all!</em> But from what I've <a href="https://www.cognitiverevolution.ai/hardcore-ai-for-history-with-mark-humphries-professor-of-history-at-wilfred-laurier-university/">heard</a>, Claude does much better with texts that are about half to two-thirds its total window. Part of this is due to the way something called &#8220;attention mechanisms&#8221; work in LLMs: it's hard to get them to pay equal attention to everything all at once.</p><p>The new GPT-4 model with the enhanced context window is currently only available to developers, so you might not be able to try them yourself for a few weeks yet, but as I have API access, I spent the last couple of days seeing what they can do. I was especially interested in looking at how GPT-4 might handle a book-length text: would it be able to answer questions about the text as a whole or only discrete areas? Would it do better with shorter texts like Claude? Would it pay attention more effectively at the beginning than at the end?</p><p>The results are surprising and once again they've forced me to re-evaluate my conception of how (and how quickly) this technology will start to change our work. For both copyright reasons and ease of access, in my tests I used a draft version of the manuscript for my 2012 book <a href="https://utorontopress.com/9781442610446/the-last-plague/">The Last Plague: Spanish Influenza and the Politics of Public Health in Canada</a>, minus two chapters (5 and 6) and all the notes/bibliography to get the 147,000-word text to around 90,000 words. I copied and pasted the text directly from Word into a simple chatbot interface I wrote in Python to interact with the API, and then asked it to read the above text, create a table of contents, and summarize each chapter in five points. I hit enter and three blinks of the cursor later, GPT-4-Turbo-Preview, as the model is called, began to respond. In about 8 seconds total, it read the book and produced an accurate annotated table of contents.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Ur6n!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6702ad38-6fd3-4598-ab1c-e82ab9a8fa2c_960x1236.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Ur6n!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6702ad38-6fd3-4598-ab1c-e82ab9a8fa2c_960x1236.jpeg 424w, https://substackcdn.com/image/fetch/$s_!Ur6n!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6702ad38-6fd3-4598-ab1c-e82ab9a8fa2c_960x1236.jpeg 848w, https://substackcdn.com/image/fetch/$s_!Ur6n!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6702ad38-6fd3-4598-ab1c-e82ab9a8fa2c_960x1236.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!Ur6n!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6702ad38-6fd3-4598-ab1c-e82ab9a8fa2c_960x1236.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Ur6n!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6702ad38-6fd3-4598-ab1c-e82ab9a8fa2c_960x1236.jpeg" width="960" height="1236" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6702ad38-6fd3-4598-ab1c-e82ab9a8fa2c_960x1236.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1236,&quot;width&quot;:960,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:509317,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Ur6n!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6702ad38-6fd3-4598-ab1c-e82ab9a8fa2c_960x1236.jpeg 424w, https://substackcdn.com/image/fetch/$s_!Ur6n!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6702ad38-6fd3-4598-ab1c-e82ab9a8fa2c_960x1236.jpeg 848w, https://substackcdn.com/image/fetch/$s_!Ur6n!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6702ad38-6fd3-4598-ab1c-e82ab9a8fa2c_960x1236.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!Ur6n!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6702ad38-6fd3-4598-ab1c-e82ab9a8fa2c_960x1236.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Next, I asked it to summarize the book, its argument, and the main evidence used to support the thesis. Again, it did a great job and in about 10 seconds. I then tried asking it more specific questions about the book, specifically the roles of Newton Rowell (a federal cabinet minister) and Helen Reid (a social reformer) in the creation of the Federal Department of Health. Both these people appear in the book&#8217;s final chapters, and while Rowell plays a relatively important part in the story, Reid appears in only a couple of paragraphs. I also chose them because, as might be expected, the base version of GPT-4 knows very little about their involvement in the creation of the Department of Health. After reading my book (literally in under 3 seconds), GPT-4-Turbo did.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ByA3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4b402da5-ae72-494d-bde6-3bede24ffee1_961x816.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ByA3!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4b402da5-ae72-494d-bde6-3bede24ffee1_961x816.png 424w, https://substackcdn.com/image/fetch/$s_!ByA3!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4b402da5-ae72-494d-bde6-3bede24ffee1_961x816.png 848w, https://substackcdn.com/image/fetch/$s_!ByA3!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4b402da5-ae72-494d-bde6-3bede24ffee1_961x816.png 1272w, https://substackcdn.com/image/fetch/$s_!ByA3!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4b402da5-ae72-494d-bde6-3bede24ffee1_961x816.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ByA3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4b402da5-ae72-494d-bde6-3bede24ffee1_961x816.png" width="961" height="816" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4b402da5-ae72-494d-bde6-3bede24ffee1_961x816.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:816,&quot;width&quot;:961,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:71976,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!ByA3!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4b402da5-ae72-494d-bde6-3bede24ffee1_961x816.png 424w, https://substackcdn.com/image/fetch/$s_!ByA3!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4b402da5-ae72-494d-bde6-3bede24ffee1_961x816.png 848w, https://substackcdn.com/image/fetch/$s_!ByA3!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4b402da5-ae72-494d-bde6-3bede24ffee1_961x816.png 1272w, https://substackcdn.com/image/fetch/$s_!ByA3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4b402da5-ae72-494d-bde6-3bede24ffee1_961x816.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!uW_J!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3d9c2512-f36b-45eb-b56f-45c0d54c22d6_963x814.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!uW_J!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3d9c2512-f36b-45eb-b56f-45c0d54c22d6_963x814.png 424w, https://substackcdn.com/image/fetch/$s_!uW_J!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3d9c2512-f36b-45eb-b56f-45c0d54c22d6_963x814.png 848w, https://substackcdn.com/image/fetch/$s_!uW_J!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3d9c2512-f36b-45eb-b56f-45c0d54c22d6_963x814.png 1272w, https://substackcdn.com/image/fetch/$s_!uW_J!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3d9c2512-f36b-45eb-b56f-45c0d54c22d6_963x814.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!uW_J!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3d9c2512-f36b-45eb-b56f-45c0d54c22d6_963x814.png" width="963" height="814" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3d9c2512-f36b-45eb-b56f-45c0d54c22d6_963x814.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:814,&quot;width&quot;:963,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:72067,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!uW_J!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3d9c2512-f36b-45eb-b56f-45c0d54c22d6_963x814.png 424w, https://substackcdn.com/image/fetch/$s_!uW_J!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3d9c2512-f36b-45eb-b56f-45c0d54c22d6_963x814.png 848w, https://substackcdn.com/image/fetch/$s_!uW_J!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3d9c2512-f36b-45eb-b56f-45c0d54c22d6_963x814.png 1272w, https://substackcdn.com/image/fetch/$s_!uW_J!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3d9c2512-f36b-45eb-b56f-45c0d54c22d6_963x814.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>As a final test, I wanted to see how well it did at pulling specific, obscure statistical information from the latter part of the book. Somewhere towards the end of the text, I had written about the Federal Department of Health's early role in funding the building of family homes between 1919 and 1925, so I asked the model the following question: &#8220;Based solely on the above text, describe the Federal Department of Health's contribution to housing. Specifically, how large a fund did it make available to local governments for housing in 1919? How many municipalities had used this fund by 1924?&#8221; A few seconds later, it correctly responded: &#8220;Based on the text provided, the Federal Department of Health made a $25 million fund available to local governments to promote housing construction in 1919. This funding was available on the basis of 25 percent co-participation. By the end of 1924, 179 municipalities had tapped into the program to build more than 6,240 single-family homes across the Dominion.&#8221; Its answer was paraphrased from just three sentences on pages 182-83 of a 195-page book, right near the end where I expected the model's attention would fade. For the record, it took me much longer to find the actual passage in the printed book to cite it than it took GPT-4 to generate that answer.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!LLTl!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1741016f-d3cd-4e97-a4e8-a6e1a3640739_964x815.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!LLTl!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1741016f-d3cd-4e97-a4e8-a6e1a3640739_964x815.png 424w, https://substackcdn.com/image/fetch/$s_!LLTl!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1741016f-d3cd-4e97-a4e8-a6e1a3640739_964x815.png 848w, https://substackcdn.com/image/fetch/$s_!LLTl!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1741016f-d3cd-4e97-a4e8-a6e1a3640739_964x815.png 1272w, https://substackcdn.com/image/fetch/$s_!LLTl!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1741016f-d3cd-4e97-a4e8-a6e1a3640739_964x815.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!LLTl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1741016f-d3cd-4e97-a4e8-a6e1a3640739_964x815.png" width="964" height="815" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1741016f-d3cd-4e97-a4e8-a6e1a3640739_964x815.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:815,&quot;width&quot;:964,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:72545,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!LLTl!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1741016f-d3cd-4e97-a4e8-a6e1a3640739_964x815.png 424w, https://substackcdn.com/image/fetch/$s_!LLTl!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1741016f-d3cd-4e97-a4e8-a6e1a3640739_964x815.png 848w, https://substackcdn.com/image/fetch/$s_!LLTl!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1741016f-d3cd-4e97-a4e8-a6e1a3640739_964x815.png 1272w, https://substackcdn.com/image/fetch/$s_!LLTl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1741016f-d3cd-4e97-a4e8-a6e1a3640739_964x815.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Clearly, GPT-4-Turbo is capable of utilizing its full context window. It's also fast and cheap for what you get: each of my questions cost about $0.92 to answer. For comparison, if I had been able to send that much text to OpenAI's previously most capable model (which had a 32k context window but was never fully deployed), each query would have cost about $6.10.</p><p>It's the exponential nature of these changes that I find most staggering. Think about what this represents: in less than one year, OpenAI's top models have increased their text-processing capacity 32 times over, from about 4k tokens to 128k. If that pace continues&#8212;and even if it slows, there is no reason to think it won't still keep improving as processing power increases and the models become more efficient&#8212;by next year we are likely to see models capable of processing many hundreds of thousands of words of text at a time, perhaps millions. And yet, the cost is simultaneously falling&#8212;its fallen about 7 times over during the same period for OpenAI's top model. All the while new, more capable LLMs are being planned and trained by all the major players.</p><p>Why does this matter so much? Well, like many professional and amateur developers, for the better part of the past four months, I have been wrestling to balance a much smaller context window (8,000 tokens) with much higher costs, trying to find a way of ensuring that I can pass the &#8220;right&#8221; documents from a database of fur trade records to GPT-4. This involved many steps: fine-tuning models, trying out various types of semantic search, and Retrieval Augmented Generation (RAG). All of this ended up costing lots of money, sometimes as much as several dollars per query in order to get the best, most reliable results. Increasing GPT-4&#8217;s context window effectively solves most of those issues by allowing me to retrieve all the potentially relevant documents and tell a very capable model to "just sort it out".</p><h3>Getting the GPT-4 Vision API to Read JPG Documents&#8230; Thousands of Them</h3><p>OpenAI also gave developers access to its vision API on Tuesday. This means that people can now send thousands of images a minute to GPT-4 and use its vision capabilities to analyze the contents of things like photographs of documents&#8212;handwritten and typed&#8212;without having to OCR them first. To be sure, GPT-4&#8217;s vision capabilities are not perfect, but they will continue to improve. Although it struggles to transcribe handwritten or typed documents, it has no problem summarizing the contents or answering detailed questions about them. I&#8217;ve experimented with this a bit over the past few days, and it's really impressive, both in terms of the quality of the individual outputs as well as what happens when it's deployed at scale.</p><p>This is another thing that I used to think was still years away : you can now send 10,000 high-res JPGs of archival documents to GPT-4 and ask it to identify the specific images in which a concept or person&#8217;s name appears, much as I did above with the text of my book. But these are raw JPGs of archival documents, not OCRed text. At $0.007 per high-res image, this would only cost you $70.00, and it might take the better part of an hour to process through the API. In comparison, imagine hiring a research assistant to do the same task. A well-trained, competent graduate student might spend 3 minutes reading each page, for a total of 500 hours or about fourteen weeks of work&#8212;an entire summer. At $32.00 an hour, which is my institution&#8217;s new minimum rate for research assistants, that would cost around $16,000. So before Tuesday, this very common task would take about 500 times longer and cost 200 times more than it does now.</p><h3>Knowledge Gaps and their Consequences</h3><p>When you start to think this all through, it gets scary quickly. Even if generative AI doesn't get any better than it currently is (which it will), the effects of the technology are still going to be seismic. The problem is that it's developing faster than companies can actually build the advancements into their products. And so, most people don&#8217;t understand what AI can actually do at the moment&#8212;right now&#8212;because they haven&#8217;t had a chance to try its cutting-edge capabilities.</p><p>As a result, we're witnessing a widening gap between general awareness of AI&#8217;s capabilities and the actual state of the technology. The problem is that most people I talk to are still genuinely amazed that generative AI can write poems in the style of Bob Dylan. There is a lot of road between that place and AI&#8217;s new ability digest an entire archive in a few minutes. AI is simply starting to outrun many people&#8217;s imaginations. It's a bit like trying to describe the iPhone&#8217;s capabilities to someone whose only experience with cellular technology was a fleeting encounter with an early 1990s car phone. All this is happening so quickly that most people won't realize the magnitude of the transformation until long after their world has been completely engulfed by the change.</p><p>Nowhere is this truer than in academia. Many institutions are still in the initial stages of reacting to the technology as it existed a year ago. Committees are being formed, while others are holding their first meetings, all to discuss people's experience with a version of the technology&#8212;ChatGPT GPT-3.5&#8212;that has already been surpassed many times over. In effect, they&#8217;re still talking about the technology as if it is a theoretical thing without the full understanding that it poses a very real, existential threat to our whole <em>raison d'&#234;tre</em> as higher educators.</p><p>Not everyone is going to get caught in the dark, though. As in any revolution, there will be winners and losers. Surprisingly, the Canadian federal government seems to be ahead of many schools, as they released some very <a href="https://www.canada.ca/en/government/system/digital-government/digital-government-innovations/responsible-use-ai/guide-use-generative-ai.html">permissive guidance</a> on AI usage more than two months ago now. It is notable too that unlike many universities, that document calls on employees and the general population to engage with AI, to use it to inform policy discussions and to write emails, proceeding with caution when making decisions or engaging with the public. Given the Canadian government's poor track record on technology issues of late, it worries me that they seem to be outpacing our institutions of higher learning in realizing that, like it or not, the technology is here to stay and can actually be very useful.</p><p>In Canada, some institutions like McMaster University have been quicker than others to issue <a href="https://provost.mcmaster.ca/office-of-the-provost-2/generative-artificial-intelligence/">guidance</a> on the use of AI in teaching and research. But because not every university is taking the same approach, gaps are widening. There will come a point, though, when the distance between the technology being used in the world and the awareness and understanding of that technology by staff, faculty, and administration grows so large at some institutions that it will become unbridgeable. When this happens, it will almost certainly be the early adopters that thrive. It may well prove impossible for latecomers to catch-up which will be disastrous in a world already beset by declining enrolments and government pressures to become more "workforce relevant". </p><p>So as exciting as all the recent updates are, I fear they move us closer to a point where those paths are beginning to irrevocably diverge.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://generativehistory.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Generative History! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[The Cognitive Revolution: AI for History Podcast]]></title><description><![CDATA[I recently joined the Cognitive Revolution to discuss AI in historical research and teaching]]></description><link>https://generativehistory.substack.com/p/the-cognitive-revolution-ai-for-history</link><guid isPermaLink="false">https://generativehistory.substack.com/p/the-cognitive-revolution-ai-for-history</guid><dc:creator><![CDATA[Mark Humphries]]></dc:creator><pubDate>Tue, 31 Oct 2023 12:55:18 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!j9Ge!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80b742bd-dbad-4664-ae92-d1823edf4003.avif" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://www.cognitiverevolution.ai/" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!j9Ge!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80b742bd-dbad-4664-ae92-d1823edf4003.avif 424w, https://substackcdn.com/image/fetch/$s_!j9Ge!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80b742bd-dbad-4664-ae92-d1823edf4003.avif 848w, https://substackcdn.com/image/fetch/$s_!j9Ge!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80b742bd-dbad-4664-ae92-d1823edf4003.avif 1272w, https://substackcdn.com/image/fetch/$s_!j9Ge!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80b742bd-dbad-4664-ae92-d1823edf4003.avif 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!j9Ge!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80b742bd-dbad-4664-ae92-d1823edf4003.avif" width="512" height="512" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/80b742bd-dbad-4664-ae92-d1823edf4003.avif&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:512,&quot;width&quot;:512,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:30097,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/avif&quot;,&quot;href&quot;:&quot;https://www.cognitiverevolution.ai/&quot;,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!j9Ge!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80b742bd-dbad-4664-ae92-d1823edf4003.avif 424w, https://substackcdn.com/image/fetch/$s_!j9Ge!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80b742bd-dbad-4664-ae92-d1823edf4003.avif 848w, https://substackcdn.com/image/fetch/$s_!j9Ge!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80b742bd-dbad-4664-ae92-d1823edf4003.avif 1272w, https://substackcdn.com/image/fetch/$s_!j9Ge!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80b742bd-dbad-4664-ae92-d1823edf4003.avif 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p>I&#8217;ve been listening to <a href="https://www.cognitiverevolution.ai/">the Cognitive Revolution</a> since the podcast&#8217;s inception around this time last fall. It&#8217;s become one of my favorite resources for learning about not only the deployment of AI in various fields, but also its broader implications for society. </p><p>So this past week, I was thrilled to join Nathan Labenz, and his co-host, Erik Torenberg, for <a href="https://www.cognitiverevolution.ai/hardcore-ai-for-history-with-mark-humphries-professor-of-history-at-wilfred-laurier-university/">an episode of the program</a> all about AI in history. </p><p>If you&#8217;re interested in learning more about AI and are still trying to understand what all the hype is about, I would encourage you to explore the show&#8217;s growing archive of episodes on <a href="https://open.spotify.com/show/6yHyok3M3BjqzR0VB5MSyk">Spotify </a>or <a href="https://podcasts.apple.com/us/podcast/the-cognitive-revolution-ai-builders-researchers-and/id1669813431?mt=2&amp;ls=1">Apple Podcasts</a>. Despite the sometimes highly technical nature of the subject matter, I have found all of them to be highly accessible. Collectively they provide an excellent introduction to generative AI technology and its broader implications.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://www.cognitiverevolution.ai/hardcore-ai-for-history-with-mark-humphries-professor-of-history-at-wilfred-laurier-university/" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!kEgz!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F358ebd1a-a9c7-46e0-86f1-d2c59888ec3a_747x267.png 424w, https://substackcdn.com/image/fetch/$s_!kEgz!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F358ebd1a-a9c7-46e0-86f1-d2c59888ec3a_747x267.png 848w, https://substackcdn.com/image/fetch/$s_!kEgz!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F358ebd1a-a9c7-46e0-86f1-d2c59888ec3a_747x267.png 1272w, https://substackcdn.com/image/fetch/$s_!kEgz!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F358ebd1a-a9c7-46e0-86f1-d2c59888ec3a_747x267.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!kEgz!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F358ebd1a-a9c7-46e0-86f1-d2c59888ec3a_747x267.png" width="747" height="267" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/358ebd1a-a9c7-46e0-86f1-d2c59888ec3a_747x267.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:267,&quot;width&quot;:747,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:149248,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:&quot;https://www.cognitiverevolution.ai/hardcore-ai-for-history-with-mark-humphries-professor-of-history-at-wilfred-laurier-university/&quot;,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!kEgz!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F358ebd1a-a9c7-46e0-86f1-d2c59888ec3a_747x267.png 424w, https://substackcdn.com/image/fetch/$s_!kEgz!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F358ebd1a-a9c7-46e0-86f1-d2c59888ec3a_747x267.png 848w, https://substackcdn.com/image/fetch/$s_!kEgz!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F358ebd1a-a9c7-46e0-86f1-d2c59888ec3a_747x267.png 1272w, https://substackcdn.com/image/fetch/$s_!kEgz!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F358ebd1a-a9c7-46e0-86f1-d2c59888ec3a_747x267.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://generativehistory.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Generative History! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[AI in Higher-Ed: What You Need to Know to Move from First Encounters to Actual Integration]]></title><description><![CDATA[If you still don&#8217;t get the hype around AI&#8212;or why it will transform higher education&#8212;you need to think at scale.]]></description><link>https://generativehistory.substack.com/p/ai-in-higher-ed-what-you-need-to</link><guid isPermaLink="false">https://generativehistory.substack.com/p/ai-in-higher-ed-what-you-need-to</guid><dc:creator><![CDATA[Mark Humphries]]></dc:creator><pubDate>Mon, 02 Oct 2023 02:00:27 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!eItd!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3d59585-f754-4d3b-999c-39425fdb15f9_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!eItd!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3d59585-f754-4d3b-999c-39425fdb15f9_1024x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!eItd!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3d59585-f754-4d3b-999c-39425fdb15f9_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!eItd!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3d59585-f754-4d3b-999c-39425fdb15f9_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!eItd!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3d59585-f754-4d3b-999c-39425fdb15f9_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!eItd!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3d59585-f754-4d3b-999c-39425fdb15f9_1024x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!eItd!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3d59585-f754-4d3b-999c-39425fdb15f9_1024x1024.png" width="1024" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d3d59585-f754-4d3b-999c-39425fdb15f9_1024x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!eItd!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3d59585-f754-4d3b-999c-39425fdb15f9_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!eItd!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3d59585-f754-4d3b-999c-39425fdb15f9_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!eItd!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3d59585-f754-4d3b-999c-39425fdb15f9_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!eItd!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3d59585-f754-4d3b-999c-39425fdb15f9_1024x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>I&#8217;ve spent the past week doing presentations on generative AI for a variety of university audiences&#8212;administrators and professors. One thing that really stands out is that very few people in higher education seem to understand <em>why</em> AI is going to be so disruptive for the field. I think this is because many folks either haven&#8217;t tried AI (or have only seen the less capable free versions) or have bounced off it because they don&#8217;t see the difference between <a href="https://chat.openai.com/">ChatGPT</a> and what happens when we deploy AI at scale.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://generativehistory.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://generativehistory.substack.com/subscribe?"><span>Subscribe now</span></a></p><h3>Encountering AI</h3><p>ChatGPT is designed to fit individual use cases as it&#8217;s a chatbot programmed to interact in a conversational way. There are lots of obvious uses for ChatGPT: it can teach you to code, it can write emails, translate a document, or draw a graph from an excel spreadsheet. In some jobs, these interactions are already proving transformative.</p><p>After trying ChatGPT and failing to get it to do much more than write poems and recipes, people gradually start to learn about context and prompting. Instead of just asking ChatGPT a simple question, they discover that you can cut and paste in large amounts of text (or <a href="https://openai.com/research/gpt-4v-system-card">data, images and audio</a>) into the prompt and ask the LLM to do something with it. This could be as simple as summarizing an article, or something more complex like filling out a standardized form with information from several other sources. But it can really be anything so long as you can clearly articulate the problem, tell the LLM what you want it to do, and provide it with the necessary information to do that task.</p><p>The usefulness of ChatGPT will, though, still vary widely from person to person. If you don&#8217;t have a need to do a lot of repetitive tasks quickly, you may well find all the hype puzzling. Even if you see the appeal, though, the reality is that fiddling around prompting an LLM is often far more time consuming than just doing a single, specific task yourself. And because ChatGPT is where most people start with AI, it&#8217;s also where a lot of people end: &#8220;I don&#8217;t see what<em> I</em> would do with this, so what&#8217;s the big deal?&#8221;</p><h3>Understanding AI at Scale</h3><p>If this is you, then you should know that the &#8220;big deal&#8221; comes when we start to understand how AI can be used to automate complex but monotonous tasks at scale. In effect, this means solving 30,000 problems or completing 100,000 tasks very, very quickly rather than one at a time.</p><p>Picture this: you are the chair of a history department, and you want to grow your majors by converting more first-year applicants into actual students. You could send an email to every prospective student, welcoming them to the department and encouraging them to get in touch if they have any questions. Unless you only have a few applicants, those emails are going to be pretty boiler plate because any sort of meaningful level of personalization just takes too much time. Still, hopefully at least a few would convert into actual enrollments, but that ROI is never guaranteed.</p><p>But what if you sent a complete list of your applicants along with the text of their individual applications to an LLM which already had access to your department&#8217;s list of prospective courses, instructor websites, and degree requirements? You could then ask it to write a truly personalized and helpful welcome email for each individual student. That letter might identify the types of courses a student would like to take based on their high school transcripts, the professors they might like to connect with based on their interests, and the various services available to them on campus according to their specific needs. It could also direct them to the scholarships for which they qualify. The most important thing is that each email would be unique and different and would only take a few seconds to produce, all for pennies--literally. Suddenly, ROI increases exponentially as costs and time plummet but the level of meaningful engagement soars.</p><p>Now: scale that up beyond your department for the university as a whole. Then try to imagine other similar use cases for the technology. You could train an LLM on all your course materials and create a study aid to help students prepare for exams. You could redeploy that same LLM (combined with state-of-the-art speech-to-text interpreters) to run individual and group discussions in large classes. With a bit more tweaking, you could have AI listen to those discussions and grade the students on their knowledge of the material (<a href="https://blog.khanacademy.org/harnessing-ai-so-that-all-students-benefit-a-nonprofit-approach-for-equal-access/">Khan Academy</a> has already started implementing these approaches with OpenAI). As with individual ChatGPT use cases, the only real limitation is: can you clearly define the problem you need to solve, articulate the type of solution you want, and can you provide an LLM with the inputs and data necessary to complete the task? Are you starting to get it? It&#8217;s exciting, terrifying, and crystal clear where this all goes&#8212;and soon.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://generativehistory.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://generativehistory.substack.com/subscribe?"><span>Subscribe now</span></a></p><h3>Integrating AI in Higher Education</h3><p>AI is revolutionary because when deployed at scale it destroys the old nexus between cost, size, and depersonalization. In the old days, the larger a class became, the cheaper it was to run because of economies of scale: professor salaries remain constant, but more bums-in-seats meant more revenue. But as we all know, those saving came at a cost: bigger classes depersonalized the learning. We try to mitigate that reality with group tutorials and TAs, but in my experience, the student experience declines in proportion to the growth of class-size.</p><p>AI has the potential to explode this equation. In the near future class sizes are likely to grow quite large but paradoxically become far more personalized than they are now&#8212;at least from the student&#8217;s perspective. AI tutors will be able to explain math problems to confused students for hours on end, in hundreds of different ways until they actually understand the concept. LLMs can not only grade papers, but also allow students to retry assignments over and over until they actually &#8220;get it.&#8221; Now apply these same ideas to the delivery of student services, administration and service, and research. I agree with <em>New York Times</em> columnist <a href="https://www.youtube.com/watch?v=p8X4Fekk9Io">Ezra Klein that there is no other word for this future than weird</a>.</p><h3>Closing the Knowledge Gap</h3><p>To be clear, this is all deeply unsettling to me and my first instinct was either to run and hide in my office or man the barricades when I had my &#8220;ah ha&#8221; moment last winter. But as every major tech company <a href="https://www.reuters.com/technology/ai-lesson-microsoft-google-spend-money-make-money-2023-07-25/">invests heavily in the technology</a> (and as <a href="https://www.reuters.com/technology/us-restricts-exports-some-nvidia-chips-middle-east-countries-filing-2023-08-30/">GPUs exports are restricted</a> by the United States), it&#8217;s becoming clear that LLMs are not going away. There is simply no longer a world in which these tools won&#8217;t exist and be part of our lives.</p><p>This brings us full circle, back to the knowledge gap. Most academics and universities seem to be still be in the stage of <em>encountering</em> AI and have yet to reach a point where they can clearly visualize how LLMs will actually be <em>integrated</em> into higher education. This is especially problematic as the AI world is evolving so quickly: if you are not actively engaged in it everyday, you miss things and those misses quickly accumulate. </p><p>As this brave new weird world takes shape, individual academics and institutions are soon going to need to begin to connect the dots and develop strategies to not only cope with AI, but to harness it and integrate it into their work. It has the potential to solve many significant issues, but only if we begin to thoughtfully and ethically prepare the ground now. This won&#8217;t be a simple task, but it&#8217;s <a href="https://www.cbc.ca/news/canada/artificial-intelligence-jobs-careers-training-panle-the-national-1.6978515">hard to imagine there&#8217;ll be a place</a> for those institutions that come late to the game. </p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://generativehistory.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Generative History! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Some Advice To Handle AI in the Classroom this Fall]]></title><description><![CDATA[Should you focus on academic integrity or embrace ChatGPT this fall? Don&#8217;t believe the hype because you can actually do both, just by doing what you&#8217;ve always done.]]></description><link>https://generativehistory.substack.com/p/some-advice-to-handle-ai-in-the-classroom</link><guid isPermaLink="false">https://generativehistory.substack.com/p/some-advice-to-handle-ai-in-the-classroom</guid><dc:creator><![CDATA[Mark Humphries]]></dc:creator><pubDate>Mon, 21 Aug 2023 10:00:33 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!WgTK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0596598-9d03-402e-81ae-a072e7931ac7_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!WgTK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0596598-9d03-402e-81ae-a072e7931ac7_1024x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!WgTK!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0596598-9d03-402e-81ae-a072e7931ac7_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!WgTK!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0596598-9d03-402e-81ae-a072e7931ac7_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!WgTK!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0596598-9d03-402e-81ae-a072e7931ac7_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!WgTK!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0596598-9d03-402e-81ae-a072e7931ac7_1024x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!WgTK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0596598-9d03-402e-81ae-a072e7931ac7_1024x1024.png" width="1024" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b0596598-9d03-402e-81ae-a072e7931ac7_1024x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1600570,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!WgTK!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0596598-9d03-402e-81ae-a072e7931ac7_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!WgTK!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0596598-9d03-402e-81ae-a072e7931ac7_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!WgTK!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0596598-9d03-402e-81ae-a072e7931ac7_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!WgTK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0596598-9d03-402e-81ae-a072e7931ac7_1024x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Everyone seems <a href="https://www.washingtonpost.com/technology/2023/08/13/ai-chatgpt-chatbots-college-cheating/">worried</a> that classes will be flooded with AI writing this fall. Many professors&#8217; first instinct will be to ban AI which means policing violations with online detectors or their own intuition. But despite what you <a href="https://www.pcguide.com/apps/can-chat-gpt-be-detected-by-turnitin/">may have read</a>, there is significant debate about whether AI detectors work&#8212;or at least whether they are reliable or equitable (for the record, I don&#8217;t think they are). Aside from the fact that they do make <a href="https://www.washingtonpost.com/technology/2023/08/14/prove-false-positive-ai-detection-turnitin-gptzero/">false accusations</a>, it is important to understand why such accusations are also fundamentally impossible to prove. University lawyers are quickly <a href="https://www.usatoday.com/story/news/education/2023/04/12/how-ai-detection-tool-spawned-false-cheating-case-uc-davis/11600777002/">discovering</a> that this is a real problem.</p><p>So what do we do? Many seem to think that this means either giving up on academic integrity or reinventing the wheel when it comes to assessment, but that&#8217;s a false choice. What no one seems to realize is that the fears around AI plagiarism are somewhat of a red herring. As I&#8217;ll argue here, if you take a traditional approach focused on essay writing and in-class testing&#8212;much as we&#8217;ve always done in History&#8212;students will have a very hard time using AI to write their assignments for them. And importantly, this approach will also leave room for students to use it to do more legitimate things like help them edit their writing, understand concepts, and think through problems. You can, in other words, do what you&#8217;ve always done and still embrace AI.</p><h3><strong>Why You Won&#8217;t Know AI Writing When You See It (Even If You Think You Will)</strong></h3><p>To begin, though, we need to understand why trying to detect AI writing and prosecute offenders is a losing game. Large Language Models (LLMs) are designed to write and they form sentences and paragraphs by trying to predict the next word based on a massive amount of training data. I know it sounds logical to conclude from this that an LLM should always answer the same question in a similar way, but this is not how these models actually work.</p><p>Under the hood, LLMs like ChatGPT have a <a href="https://michaelehab.medium.com/the-secrets-of-large-language-models-parameters-how-they-affect-the-quality-diversity-and-32eb8643e631">variety of settings</a> which introduce randomness into the predictive process and actually penalize the model for repeating itself. When these settings are minimized, they indeed produce similar answers to similar questions. But by adjusting them in just the right way, as they are on all commercial models like ChatGPT, LLMs choose <em>one</em> <em>of the most probable</em> <em>words</em> to continue the answer rather than <em>the</em> most probable word. This introduces a level of variation that has a minimal effect on coherence but a profound, cascading effect on the diversity of the syntax and content of the answers. In effect, it ensures that AI writing is rarely ever exactly the same because the combinations of somewhat probable words are unimaginably large. In fact, they increase exponentially towards infinity as each new word is added to the AI&#8217;s response.</p><p>This is why you typically won&#8217;t be able to prove that something is AI written: it is almost never the same twice. Now this does not mean that LLMs are not derivative, repetitive, and stylized. They are. LLM writing is generally bland and formulaic, largely because the underlying models have been fine-tuned to answer questions in ways that humans find desirable (the underlying base models do not do this nearly as well). This is why if you ask ChatGPT to identify and explain the historical significance of <a href="https://www.thecanadianencyclopedia.ca/en/article/durham-report">Lord Durham&#8217;s report</a>, it will almost always start with something like: &#8220;Lord Durham's Report, officially known as "Report on the Affairs of British North America," was published in 1839&#8230;&#8221; It will also usually end with something like &#8220;In conclusion, Lord Durham's Report is seen as a landmark document that had profound and lasting effects on Canadian governance and identity...&#8221; That is pretty bland, but there is an obvious problem with assuming that answers structured like this are automatically AI generated: the AI was, in fact, trained to write like this because it is what humans ideally want to see when they are asked a question about the importance of the Durham Report. Go back in time a year: is this not how you would want the answer to begin if generative AI did not exist? There is a real risk that if we start assuming everyone is using AI writing, we will begin to see it everywhere as <a href="https://www.rollingstone.com/culture/culture-features/texas-am-chatgpt-ai-professor-flunks-students-false-claims-1234736601/">one Texas professor discovered</a> last spring.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://generativehistory.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://generativehistory.substack.com/subscribe?"><span>Subscribe now</span></a></p><h3><strong>The Perils of AI Detection Software</strong></h3><p>Enter the AI detectors. Because LLM writing is inherently unpredictable, so-called AI detectors (which are usually built on-top of LLMs themselves) are typically designed to look for these types of stylized formulations. They also look for an absence of spelling and grammatical errors and the presence of specific patterns of logical reasoning all of which are said to be typical of artificial intelligence. The problem is, though, that unlike conventional plagiarism detectors, these tools are rarely able to backup an accusation with <a href="https://www.washingtonpost.com/technology/2023/05/18/texas-professor-threatened-fail-class-chatgpt-cheating/">hard evidence</a>, only an assessment of the likelihood that a piece of writing was AI generated.</p><p>If that sounds &#8220;good enough&#8221; to you, think about it for a second: AI detectors are LLMs themselves that have been tasked with identifying AI generated writing. This means that all the normal arguments that are mounted against generative AI would also apply here. First, remember that we cannot explain how an LLM reaches any conclusion, and that is true when a detector says a text was &#8220;highly likely&#8221; to be AI written. Studies have also shown that the writing of <a href="https://hai.stanford.edu/news/ai-detectors-biased-against-non-native-english-writers">ESL students and visible minorities is more likely to be flagged as AI generated</a> by AIs. Remember too that LLMs are prone to hallucination, meaning sometimes they just make things up. This is exactly why OpenAI, which released one of the first AI detectors last winter, quietly withdrew it a few weeks ago: it was inaccurate.</p><p>Now here is an obvious caveat. The fine-tuning process also means that LLMs will sometimes say things like: &#8220;As a large language model developed by OpenAI, while I am unqualified to assess the significance of Lord Durham&#8217;s report&#8230;&#8221; So of course, if you see this in a student&#8217;s paper, to my mind at least, this would constitute pretty good evidence of AI use. But here&#8217;s the thing: LLMs only rarely include such caveats and when they do, I think most students would just delete them before submission.</p><h3><strong>Don&#8217;t Reinvent the Wheel to Thwart AI</strong></h3><p>So does this mean that you should either give up or completely re-evaluate all your assessment methods? The simple answer is &#8220;no&#8221; and &#8220;not necessarily&#8221;.</p><p>The truth is I think most professors are worried about AI plagiarism because they haven&#8217;t really become familiar with what generative AI can and cannot do yet. If you&#8217;re worried about AI plagiarism but don&#8217;t yet have a paid OpenAI account, as a first step <strong>I would strongly encourage you</strong> to fork over the $20 for at least one month and <strong><a href="https://auth0.openai.com/u/signup/identifier?state=hKFo2SBmMlVxUHhmRll4YXpPdWgtMzlmSHZmQldXNmJLZWkyMqFur3VuaXZlcnNhbC1sb2dpbqN0aWTZIEp2RWVBdHd3WE1UTmdnbjJRREI1bWJNblJhWjlZNVBmo2NpZNkgRFJpdnNubTJNdTQyVDNLT3BxZHR3QjNOWXZpSFl6d0Q">try out GPT-4</a></strong>. This is because the free version, GPT-3.5, is not what students will be using: it will underwhelm you with its tendency to make-up sources, hallucinate, and generally betray itself as a fallible AI at every opportunity. GPT-4 is far from perfect, but it is a wholly different animal than its predecessor.</p><p>The biggest reason to try GPT-4, though, is that it will help you face your worst fears. Cut and paste your assignment directions into the chatbot and see what you get back.</p><h3><strong>Don&#8217;t Use These Types of Assignments</strong></h3><p>After many months of working with both GPT-4 and Llama models here are the types of assignments at which I think Generative AI will excel.</p><p>First, ChatGPT is best at writing short answers of less than 750 words. It is extremely good at identify and explain type of questions, probably scoring around 90%-100%&#8211; even on very obscure things. So if you are still doing online exams or take-home tests, these are exactly the types of questions that ChatGPT will be able to answer in an undetectable way.</p><p>Next up are book reviews and primary document analyses. These staples of the history syllabus are relatively short assignments and even if they are longer than 1000 words are pretty easy to get ChatGPT to do them because of their formulaic nature. Even if a book is relatively new, GPT-4 will still do a pretty decent job at writing a critical review. Same with primary documents: try uploading a document to ChatGPT and ask it some questions about it. And if you were thinking of thwarting AI by using handwritten documents, take a look at <a href="https://readcoop.eu/transkribus/">Transkribus</a> which offers a limited, free version on its website. It is even better at transcribing handwritten documents now (thanks to AI) so keep that in mind. This doesn&#8217;t mean that you need to abandon book review and document analyses, though. I would suggest that they need to be significantly longer: starting at around 2,500 words students will have a hard time assembling a coherent paper from a variety of ChatGPT responses. Or just consider incorporating them into in-class testing (see below).</p><p>Third, discussion board participation for grades. Again, this typically requires students to write short paragraphs in response to a question as well as student responses. To generate a discussion board answer with ChatGPT, all a student needs to do is cut and paste the conversation into the chatbot and ask it to respond. Remember too that you can tell ChatGPT to adopt a persona: because ChatGPT was trained on Twitter (X?), Reddit and Facebook posts, it can write informally when asked to do so.</p><p>Third, reflection papers are an equally easy target. Although ChatGPT cannot write more than around 750 words at a time, it can be prompted to expand its answers by cutting and pasting paragraphs back into the chatbot. Given the nature of reflective writing, versus essays, this is relatively straight forward and quick to do. Again, here is where I would encourage you to try GPT-4 to see for yourself: you can ask it to adopt any perspective, attitude, or identity or write for a specific purpose and audience.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://generativehistory.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://generativehistory.substack.com/subscribe?"><span>Subscribe now</span></a></p><h3><strong>These Assignments are Generally more AI Resistant (but also good pedagogy)</strong></h3><p>The ironic thing is that if you are worried about AI, a back to basics approach will probably work best. If you choose a longer written assignment and apply rigorous standards, any students looking to cheat with AI will most likely fail themselves due to a lack of coherence, citations, or length, all without having to meet the impossible burden of proving that AI was using in the writing process.</p><p>This is not because AI is not good at writing, but because it won&#8217;t (at present) produce more than 750-900 words at a time. You can get it to do more by working in stages and with some elaborate prompting, but it would be a real chore to try and get it to produce a full, coherent research paper of 3-5,000 words. Indeed crafting such a paper with AI would, in effect, require students to go through all the same pedagogical steps and processes as writing their own paper (hence why it may be worth considering letting them use AI anyway). But the reality is that when students try to do this in order to cheat, they will almost certainly end up with a short, disjointed, poorly cited paper. It will, in effect, fail on its own merits without the impossible burden of proving something was AI written.</p><p>In addition to length, sourcing (and citation style) and the use of direct quotes are also big stumbling blocks for AI, but not for the reason you might think. GPT-4 is pretty good at coming up with real sources (unlike its predecessor), but it is not good at quoting them directly. When prompted to so, GPT-4 will sometimes get it verbatim but more often it paraphrases a text within quotation marks. Misquoting a source is a problem in academic work and the intentional use of made-up quotations is again grounds for failure regardless of AI.</p><p>In terms of citations, GPT-4 will also tend to get the general place in a book or article correct, but will rarely get the exact page right. So if you are using a citation style that does not require page numbers, I would start to require them (or websites). Again, incorrect citations are always a problem, especially if they are pervasive.</p><p>Obviously closed book, in-class tests and exams are an excellent way of eliminating the AI problem too. But they also eliminate most forms of cheating. Again, a traditional approach here solves a lot of problems.</p><h3><strong>Don&#8217;t Panic and Be Open to Innovation</strong></h3><p>The point is: don&#8217;t panic. Students are going to use AI and it will be mostly impossible to prove. But instead of giving up or reinventing the wheel, choose traditional assignments that teach students the skills they need to learn as historians, but that will also be difficult to accomplish with AI. No, it won&#8217;t be impossible to produce a 5,000 word AI research paper and some students may see that as a challenge. But doing so and earning a good grade will take more work and more domain-specific knowledge than writing a B paper from scratch.</p><p>As worrying as AI might be for many of us, I think we should be wary of trying to avoid it altogether. &nbsp;Afterall, the students we are teaching right now will enter a workforce where they will almost certainly be expected to use AI in their research and writing. Pretending it doesn&#8217;t exist is not going to be a winning strategy in the long run.</p><p>So as you prepare for the fall, consider the possibility too that there may be some legitimate uses for AI in your classroom. Students that struggle with writing, including ESL students, can use generative AI as an editing tool. Because it is largely predictive, if you paste in an original paragraph and ask it to address grammatical and stylistic issues in the text, it will do so without changing the substance of the ideas in the text. When prompted, it can also explain to the student why it made those changes. Of course, it can do the same thing with readings, summarizing and explaining difficult concepts and filling in any knowledge gaps. It can really be an effective and helpful tool.</p><p>Whatever you choose to do, though, be crystal clear about how students can and cannot use generative AI for in your classroom. There is going to be a lot of variation out there this fall. And in all fairness to our students, remember that they will have a harder time navigating this confusing and paradoxical environment that you will. They also have a lot more at stake in world where the rules are rapidly but unevenly changing.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://generativehistory.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Generative History! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item></channel></rss>