Testing shows that Gemini 3 has effectively solved handwriting on English texts, one of the oldest problems in AI, achieving expert human levels of performance.
Hi Mark, great post as always! I've been experimenting with AI Studio and I think in the end the performance has nothing to do with language, solely with how hard the hands are. It does extraordinarily well with clear French, Spanish, Italian, and Portuguese handwriting, for instance. I therefore assume the main problem is computer vision, not the underlying language. So I'm not so sure LLMs will surpass Transkribus anytime soon for harder hands, but I might be wrong, as I've underestimated LLMs' transcription capabitilies before!
It is certainly possible, but that is actually the point of the Bitter Lesson: given enough compute the generalist approach will always win out. To this point, progress on accuracy has been entirely in line with scaling to this point, directly correlated with compute and model size. So with Gemini, this means 60% improvements every 9 months or so. Claude is similar. I would thus bet that if I had a document with a 25% CER now, in 9 months it will be 10%, in 18 months 4%, and in 27 months 1.6%. The actual floor, though, will be determined by true legibility as even for humans, some parts of some documents are genuinely ambiguous, or subject to interpretation and extrapolation.
I’m skeptical it is that easy, because scaling supposes the availability of training data, which is surely much, much harder to come by for harder hands - or some surprising emerging ability to read hard hands. In the end, humans can read almost anything another human being wrote - I definitely know people who can read just about anything. I’d be curious to see what you deem a hard hand - all those you shared here are actually quite easy for me. I think your analysis is similar to those who think AGI is just a question of scaling, which also seems very doubtful to me. But I hope you’re right - it would be amazing to mass transcribe everything that cheaply in a couple of years!
Fully agree, I would say the Portuguese documents you shared with me were difficult as are a lot of 18th century French nostrils records from Quebec (the writing is not only messy but the pages are full of insertions, marginalia, etc). In terms of training data, in LLMs this is less about training it on difficult hands specifically, but rather providing a sufficient level of diversity of data and ensuring the distribution is representative. The idea with LLMs is that this allows generalizations, whereas with fine tuning on a specific hand a la Transkribus you are specializing the model. At present the latter approach is more effective, in the long run the general is approach is likely to be more efficient.
And how does Gemini fare with these French nostrils records? Anyway, LLMs have evolved so fast that I believe that they will prevail in the long run - the question is, of course, what's the long run - 1, 5, 10 years? I'd bet on the higher end, but again, I was skeptical of LLM transcription in general until Gemini 3, so what do I know?
Mark and Thiago. It would be interesting to test say thirty different hands from the same document type having as humans scored them by our perceived difficulty (as expert palaeographers) to transcribe them. I can offer over 100 scribal hands in the HCA 13/ image collection I have [over 30,000 images]
I would be interested in your views Mark on the ability of Gemini 3 via API to handle highly slanted text - this is an area where Transkribus continues to be dreadful. It would be a big win if Gemini 3 via the Google AI Studio API can handle lines at say 20 or 30 degrees off horizontal. I have some runs of images where every third image has big divergence from the horizontal
All data is good. In my experience Gemini is fine on slanted text. They are fundamentally two very different approaches (LLMs and Transkribus) even through to the user they look very similar. This is a really important point. An LLM is evaluating a document as a whole, all at once. It does not segment it or do any sort of pre-processing. With Transkribus and other similar approaches there is significant preprocessing and the model is evaluating segmented texts. This is why ultimately generalist approaches will prove more efficient as they can handle edge cases through scale. With Transkribus, edge cases must be handled with manual preprocessing or fine tuning.
I agree with you. How complex are the layouts that you have tried? A really good category of images to try will be tabular data with many cells. My gut feel from what you are saying about the approach of LLMs vs Transkribus is that LLMs (with excellent multimodal capabilities) should be very good with tabular data.
Gemini is my go-to for handwritten documents. I find it the best of the options. However, I have identified two issues not on your errors list. First, it can totally skip lines in a document, and second it can substitute words. I do not see those in your chart. It also does a marginal job on things written in the margins or inserted between lines. It is still way better than doing it yourself!
Two questions: Are you using it through the AI Studio or the Gemini App (AI studio looks like the image I posed above)? Second , is this the Gemini 3 model which was released last Tuesday? I found Gemini 2.5 which was the precious model did this all the time.
Interesting. It’s free and I’d be interested to hear what happens if you follow my intuitions above and try out a page that failed previously to see what happens now. Any data is good data!
Wow, this made my day! I am NOT an expert in this field, but rather a retired guy in the middle of a project turning many decades of my parents' weekly family letters into searchable form so my kids and grandkids can learn about our family using NotebookLM. Most are typewritten but many are handwritten. I have just finished successfully processing 600 or so of the typewritten letters but the learning curve to use Transkribus has been too great for me to get started learning it. I just followed your instructions on several handwritten letters and got nearly perfect results, including from the old blue aerogrammes we used when I was abroad in the 70's. This is the solution I've been waiting for! Thank you!
Fascinating development. I see this as welcome news for archives in terms of access and discovery. And Mark, I love your quote from "Computers and Humanities" in the 1960s. Permit me please to offer you another similarly prescient statement from 1869, which goes even further. It's a short article, but if time is an issue the last paragraph, and last line, really hit the mark in my view!
"THE BOOKS OF THE FUTURE
Pall Mall Gazette, 15 Sept 1869
The Daily News is alarmed at the rapid growth of our national library. Every man living has a pen in his hand ; and if only his writing takes tile form of a book entered at Stationers' Hall, it will be preserved for ages in the British Museum. Centuries hence the bookworm will find there, illustrated with woodcuts, verbatim reports of the trials of Palmer and Rush', the Mannings, Madeleine Smith, and Mdme. Rachel.
Centuries hence also he will find those numerous volumes which it is the fashion for tradesmen to issue, and which are but a sublimated form of trade circular. The wine merchant has a volume on his wines, and the hatter on his hats, and the jeweler on his jewels, and the sewing-machine manufacturer on his sewing-machines, and the lock maker on his locks, and the bootmaker on his boots, and the cook on his viands. How are we to stow all these away, and at the same time to keep pace with the literature of foreign countries which is scarcely less productive?
We think of our future librarians as of those children renowned in fairy tales, who have impossible tasks appointed them by malicious godmothers-to collect in a day all the sands of the shore, or to count ere dinner time all the grains of wheat in the kingdom. There will appear no exaggeration in this to anyone who will go to the British Museum and study the catalogue. A man may take a good constitutional walk every day in hunting for half a dozen books in this enormous catalogue, which of itself fills about 1,000 volumes. We find historians of our day like Mr. Carlyle complaining of the immeasurable amount of rubbish which they have to sift in order to get at a few paltry facts.
Must we not pity the historians of the future if they should at any time be conscientious as to turn over the mountains of waste paper which are being shot by cartloads into the Museum? Human eyes and human hands cannot possibly work through a century of such agglomeration. The human mind will despair, perhaps, of power to deal with the illimitable mass. May we hope that when things come to such a crisis, human labour of the literary sort may be in part superseded by machinery ? Machinery has done wonders, and when we think of what literature is becoming it is certainly to be wished that we could read it by machinery, and by machinery digest it."
Another thing that might be worth flagging about AI Studio is the "Grounding with Google Search" control - it's switched off in your screenshots but defaulted to 'on' when I first opened AI Studio, and it seems like it could make a real difference, so it's worth people being aware of the setting.
I got very quick and accurate results on 19th century English handwriting, and quite good on 15th century English, and I then tried it on a medieval Latin manuscript to see how it would do, and while it (unsurprisingly, as that's giving it a tricky script and also a completely different language) struggled a lot more, turning search on and off there made a very clear difference to how it handled problems like names - with search grounding turned off it made a guess based on the appearance of the text (and did understandable things like turning 'Walterus' into 'Valery'), with search grounding on it either guessed right or named real people who might plausibly have been mentioned in another text on the same subject but weren't involved in this one.
That is interesting! I've been playing around with adding in tool in a custom interface that allow the model to lookup names in a database of known names with info like dates they were active, where they were, that sort of thing, and it does help a lot. I wonder if increasing the reasoning to high on Latin would help at all? Or perhaps playing with the system instructions?
I have tried adding JSON dictionaries for the Gemini API to refer to. Why? You might want to normalise names, ships, whatever, at the time of transcription, rather than through post-processing, or you may want to control the handling of expansions (e.g. English and Latin expansions). Of course you can handle expansions in post-processing by running a very strict diplomatic transcription through a Python script on a fast workstation. A couple of test expansion dictionaries I have developed on the fly are here. I don't absolutely attest for the accuracy of the Latin expansion file:
Regarding "grounding", I think it should be possible in the systems prompt you use in the Gemini 3 API via Google AI Studio to give strict instructions as to how "grounding" is handled, and what interests me is whether relatively small (say under 1000 lines of JSON) files can be referred to by the API. You could of course inject them into the system prompt but that would use a lot of tokens and increase the cost of processing.
I’ve been amazed how Gemini Pro 3 deals with record images from scratched up, streaked, and splotched microfilm. It does a 99% correct job “reading” what is behind all that damage. As genealogists we unfortunately have to deal with this in records, so it’s a huge relief to be able to finally have a tool to help. Now I can compare what it saw to what I saw and feel more confident it’s correct.
It’s impressive to see how quickly LLMs are progressing. I’m also closely following and constantly testing the HTR capabilities of LLMs and other tools/models to see how they evolve.
For relatively clean 18th–19th century handwriting in English though, my experience is that HTR has effectively been “solved” by specialised systems for some time already, and actually for a broader range of documents and languages.
For example, last year I ran an HTR project on a mixed collection of documents in French, Dutch, English, Italian, Spanish and Latin from the 17th to 19th centuries. Using the Transkribus Text Titan I model (released summer 2023), we obtained an average CER of about 5% on a random sample of 100 pages. While not perfect, for such a varied and complex collection that is already a strong performance for an out-of-the-box model.
For relatively clean 18th–19th century Dutch material, similar or even better results can be expected with other Transkribus models.
So in this strict sense, text recognition for these kinds of documents has largely been “solved” for a while. The real remaining challenge in large-scale historical projects is robust layout detection and structured output, and it seems LLMs still have some way to go there.
Looking forward to reading your future articles, the field develops so quickly it is hard to keep up by yourself!
Entirely true. The point, though, I think is that a sub 1% error rate is as good as expert humans. There is a lot of data in that 5%, but yes we’ve had useful HTR for awhile.
It means you can also start to do a whole host of things with these documents via LLMs with having to actually transcribe them…just feed them into the machine. Cost is the other issue. 100 pages via Transkribus is about $100. It is $1 for Gemini. This has huge implications for small archives that want to digitize large collections.
That 5% rate was just an example of a by now already somewhat outdated model on a complex and varied collection.
In the end 1% CER on a badly recognized/structured layout is probably not as useable as 5% on a well recognized layout.
In the end I think we'll still be better off for a while using different tools (together) depending on the context and goals.
Regarding the pricing discussion, it isn't per se an apples with apples comparison, you get different potential necessary workflows & outputs with both solutions, each having advantages and disadvantages.
However the price of Transkribus is way below $1 per page, it'd be unaffordable otherwise (and I would have spent $250k+ in the last year!). If you're spending €0.10 on pure HTR, it is already much. And for that you do get proper line & region coordinates, TEI-XML ouput & general easier workflows for large batches, etc. included.
Of course, LLMs are for example way better at information extraction, but then costs go up significantly too. Here BERT models or QWEN can be interesting options too.
Bram, Mark, Thiago, I asked Google AI Studio the following question:
"Can Gemini 3 Pro Preview via the Google AI Studio API deliver TEI/XML output for a batch of say 20 images of C17th English language legal depositions after being given a systems instruction to provide a transcription in multiple formats such as JSON, TEI/XML, etc., and a prompt to execute the transcription?"
Here is the full response, which confirms it is very doable. I emphasise I have not yet tested the TEI/XML option, but have no reason to doubt it is possible, since it is trivial computationally. https://share.google/aimode/C6cftouavxTNeoRgW
It would be interesting to test whether Gemini has any native ability to do layout analysis. It wouldn't surprise me if it could, since there are publicly available open source programs it has probably ingested.
Well I won't go into the details of Transkribus pricing, that is an art in itself! The plans on the website are all cheaper than buying those single packages though. With larger scale projects prices drop close to the 1 cent too actually.
Woops, have just seen that Mark has already made the comparative cost point. But my point about the potential to integrate NER into the workprocess and equally low cost is still valid.
One important point to consider is comparative cost of Titan and a Gemini 3 API driven workprocess. Mark has the numbers, but my own calculations suggest that I could run 20,000 new images at a significant cost saving on Titan and simultaneously run a NER script so that I then have data which I can use to assemble a basic (or complex) knwowledge graph.
Excellent post, thank you! Appreciated the emphasis on the practical impossibility of completely eliminating typos: “I would count these as spelling errors: in very on, these vs those, where vs were, etc.” ;-)
A fun example of how I've been using Gemini 3 Pro's performance on handwriting:
I have a collection of about 200 handwritten recipes from my grandmother (who has been dead for more than a decade). I've wanted to digitize them for years, but the effort required to sit down and transcribe them myself meant the project never got off the ground. My grandmother had perfect cursive handwriting, having been taught in the 1920s. But, ironically, that also makes these recipes harder to read since so few of us regularly read or write classic cursive anymore.
There are additional complications - the recipes are on decades old notebook paper often with faded ink, stains, or rips. And, because my grandmother was often writing notes to herself with the recipes she had lots of bracketing and margin adjustments. Gemini 3 Pro has been pretty perfect on transcribing and formatting the recipes as I've asked. It's also spit out really interesting observations like commenting about how boxes of cake mix used to be bigger in the 1960s so I'd have to adjust the flour content of one of the recipes if I attempted to make it today. At one point it even recognized that a recipe scan I submitted was actually the reverse side of a document I had given it a few turns earlier, which was wild.
Mark, you make an important point about the different nature of Gemini 3 API errors from both human errors and Transkribus errors.
Have you created a typology of Gemini 3 errors and recorded the stats for the different types? Do the frequencies of error types vary by document type/handwriting, or is the distribution of frequencies a fundamental of the way Gemini 3 does its recognition?
I had a shot sometime ago at analysing the type and frequency of Transkribus errors on English High Court of Admiralty material, intending to use the insights somehow in a finetuning experiment. The data are hand collated by me from comparison of a set of Transkribus transcriptions for the volume of depositions named HCA 13/58. I never did the fine-tuning experiment, but here are the data in JSON format.
The key to Gemini 3's superiority over Gemini 2.5 is I agree the radically enhanced vision and multimodal nature of the model. The logic of this is that all general LLM models will show increased HTR capability as they improve vision and multimodality.
To Thiago Krause's point further down the comments about performance in non-English languages, and to another commentator, the logic of Mark's argument about vision capability should be that the HTR capabilities of Gemini 3 are independent of language and more related to the exact nature of the image and the character formation. That being said, I have to think that there must be some work processes running which draw n the prior LLM training, which will be much more extensive in some languages than others, which will affect the quality of the end results by language, so that CER and WER are not language independent.
I haven't done any non-English language testing of Gemini 3. Mark, if you have an Arabic speaking palaographer accessible, it would be interesting to report what Gemini 3's capabilities are with Arabic. Another test (and if it passes it would be outstanding) would be to test on Ottoman Turkish. I would give gold coins for the ability to even get the gist of Ottoman Turkish at low cost and with an easy to use api.
I tested Ottoman Turkish in LLMs and Transkribus a month or so ago for a class visit, with awful results. I tested it again today (as well as Manchu) and sent an email to the professor, asking her to show the results to her students and let me know what they think.
I can add that I tested Gemini 3 on historical German Gothic handwriting, Russian and Latvian (a small language) handwriting and in all three cases the results were surprisingly good.
The other tests I would suggest should be high volume and high-scholarly impact non-Latin alphabets. I asked Google AI Mode for some suggestions and came up with: https://share.google/aimode/Pgt2U4hFwKfu1p4l2
ChatGPT 5.1 did well enough on a Chinese early nineteenth-century land grant, according to a student in said class - I read the translation and he said it got about 85% right.
Good idea. Did the Gemini 3.0 transcription give you Arabic script or a Roman character transliteration. You could try feeding a transliteration back into Gemini and ask it to translate it into English and see if it is intelligible. I had a conversation recently with Transkribus recently which suggested they were making progress with Ottoman Turkish, but I haven't seen any data.
Arabic - although of course you could change the prompt. The translation sounded completely coherent, but that is to be expected, so I don't think it proves anything.
Transkribus, on the other hand, was completely gibberish on the same document.
Thank you Mark for you work with LLMs and handwritten text recognition.
We have been doing some experiments at our institution as well. Most of the time Gemini works brilliantly, but occasionally our results are cluttered with almost endless repetitions? Ran it on a 260 page PDF the other day, and a few pages were full of """"""""""""""-charachters and a few others were repetitions of a single word over and over again.
Has anyone else ran into the same problem and solved it somehow?
Just occasionally it seems to get stuck and go round in circles before reporting an 'internal error' - at which point it doesn't produce an effort at a transcript despite having obviously done nearly all the work. Has anyone found a workaround for this?
I spent three years building a startup whose core product was handwritten text recognition.
We used to work with banks and insurance companies in India and Europe.
We’d have scraped millions of documents online. Also got hundreds of thousands of lines of handwritten text done by data vendors.
We used a modified LSTM architecture.
Circa 2019 our models were better than what GCP had to offer. AWS had come up with a product called “Textract” which while great at document segmentation was still not as good at HTR.
Then covid came and our fledgling startup saw business slow down.
(It’s not easy selling cutting edge tech to legacy banks in India or Europe even otherwise)
We decided to shut shop rather than slog along.
While transformers were there by that time and we were aware of the potential, we were not sure they’d turn out to be as transformative.
(Excuse the pun)
But we did have an inkling that we should rather work on something else.
In hindsight, it was probably the best decision I took. Changed industries completely.
Now while the AI revolution is passing me by, I seem to have done well financially.
Maybe I’ll get back to programming and model development at some point again. But seems unlikely.
If anyone wants access to the dataset then do comment. I’ll try to find it in my old hard disks. My old Mac which had a lot of the data got bricked.
The country boasts more than 1,000 linear kilometres of handwritten manuscripts stored in archives, libraries, museums, parishes, and notary repositories, spanning from the 10th to the mid-20th century. Until now, a mere 0.001% of this incredibly vast and valuable volume of records has been transcribed and published since the dawn of the printing age. AI may well close that gap to 100%.
Over the last fortnight, I have conducted extensive testing of Gemini 3 Pro and Gemini 3 Flash on the transcription of antique manuscripts ranging from the 16th to the 19th century. I have tested thousands of pages via: A) The Gemini 3 Pro chatbot (Pro subscription); B) Gemini 3 Pro via API on the Google AI Studio Platform; and C) Gemini 3 Pro and Flash via the Google Vertex Platform.
Here are my humble takeaways:
The quality of transcriptions across all three channels is exceptionally high and will, in my view, sound the death knell for professional careers in palaeography. To illustrate: just three months ago, the going rate for a palaeographic transcription of 16th-century cursive in continental Europe was approximately €0.50 per word (!). The cost of a transcription with Gemini is a mere fraction of this sum, and it is calculated per page, not per word. Put simply, palaeography as a profession is dead. However, all intellectual jobs are effectively on death row thanks to AI. They will survive not as professions, but as arts or hobbies for the select few who wish to avoid wandering about as state-subsidised idlers in the coming age of imbecility.
The Google 3 Pro chatbot (which requires a monthly subscription fee) can be fed images of handwritten documents and will obligingly return their transcription in moments. The quality is marginally superior to transcriptions obtained by submitting images directly to the Gemini 3 models via AI Studio and Vertex. This is hardly surprising: one may be the finest prompt engineer in the world, yet replicating the bespoke system prompts running under the bonnet of Google’s flagship chatbot is impossible. Consequently, your direct calls to Google models will always rely on prompts that are slightly less efficient and productive than those crafted internally by Google HQ.
The chatbot is of little use, however, if you require transcriptions for hundreds or thousands of pages. In such instances, one must revert to direct calls to the Gemini models and acquire a modicum of programming knowledge (fear not; AI is highly effective at writing small, end-to-end applications for your needs). You need only be patient and willing to heed your LLM's instructions as it teaches you how to test scripts and report errors. Do as the LLM asks, and through a process of trial and error—where you serve as the dull human in the loop—you will get the application working.
Calling Gemini 3 models via Google AI Studio suffers from strict quota limits. If you intend to process a large volume of images, you must revert to Google’s Vertex platform.
When calling Gemini 3 models for HTR, set "thinking" and "max thinking tokens" to the absolute minimum. The more the model is permitted to "think", the more it tends to run riot, consuming vast quantities of tokens without any improvement in transcription, often failing completely. The precise reason for this eludes me.
Additionally, keep the temperature at 0 and set image resolution to high.
Always pre-process images for the best results. Attempt to reduce noise, calibrate colour, and increase brightness slightly. Crop out any parts of the document that do not bear handwriting.
Hi Mark, great post as always! I've been experimenting with AI Studio and I think in the end the performance has nothing to do with language, solely with how hard the hands are. It does extraordinarily well with clear French, Spanish, Italian, and Portuguese handwriting, for instance. I therefore assume the main problem is computer vision, not the underlying language. So I'm not so sure LLMs will surpass Transkribus anytime soon for harder hands, but I might be wrong, as I've underestimated LLMs' transcription capabitilies before!
It is certainly possible, but that is actually the point of the Bitter Lesson: given enough compute the generalist approach will always win out. To this point, progress on accuracy has been entirely in line with scaling to this point, directly correlated with compute and model size. So with Gemini, this means 60% improvements every 9 months or so. Claude is similar. I would thus bet that if I had a document with a 25% CER now, in 9 months it will be 10%, in 18 months 4%, and in 27 months 1.6%. The actual floor, though, will be determined by true legibility as even for humans, some parts of some documents are genuinely ambiguous, or subject to interpretation and extrapolation.
I’m skeptical it is that easy, because scaling supposes the availability of training data, which is surely much, much harder to come by for harder hands - or some surprising emerging ability to read hard hands. In the end, humans can read almost anything another human being wrote - I definitely know people who can read just about anything. I’d be curious to see what you deem a hard hand - all those you shared here are actually quite easy for me. I think your analysis is similar to those who think AGI is just a question of scaling, which also seems very doubtful to me. But I hope you’re right - it would be amazing to mass transcribe everything that cheaply in a couple of years!
Fully agree, I would say the Portuguese documents you shared with me were difficult as are a lot of 18th century French nostrils records from Quebec (the writing is not only messy but the pages are full of insertions, marginalia, etc). In terms of training data, in LLMs this is less about training it on difficult hands specifically, but rather providing a sufficient level of diversity of data and ensuring the distribution is representative. The idea with LLMs is that this allows generalizations, whereas with fine tuning on a specific hand a la Transkribus you are specializing the model. At present the latter approach is more effective, in the long run the general is approach is likely to be more efficient.
And how does Gemini fare with these French nostrils records? Anyway, LLMs have evolved so fast that I believe that they will prevail in the long run - the question is, of course, what's the long run - 1, 5, 10 years? I'd bet on the higher end, but again, I was skeptical of LLM transcription in general until Gemini 3, so what do I know?
Mark and Thiago. It would be interesting to test say thirty different hands from the same document type having as humans scored them by our perceived difficulty (as expert palaeographers) to transcribe them. I can offer over 100 scribal hands in the HCA 13/ image collection I have [over 30,000 images]
I would be interested in your views Mark on the ability of Gemini 3 via API to handle highly slanted text - this is an area where Transkribus continues to be dreadful. It would be a big win if Gemini 3 via the Google AI Studio API can handle lines at say 20 or 30 degrees off horizontal. I have some runs of images where every third image has big divergence from the horizontal
All data is good. In my experience Gemini is fine on slanted text. They are fundamentally two very different approaches (LLMs and Transkribus) even through to the user they look very similar. This is a really important point. An LLM is evaluating a document as a whole, all at once. It does not segment it or do any sort of pre-processing. With Transkribus and other similar approaches there is significant preprocessing and the model is evaluating segmented texts. This is why ultimately generalist approaches will prove more efficient as they can handle edge cases through scale. With Transkribus, edge cases must be handled with manual preprocessing or fine tuning.
I agree with you. How complex are the layouts that you have tried? A really good category of images to try will be tabular data with many cells. My gut feel from what you are saying about the approach of LLMs vs Transkribus is that LLMs (with excellent multimodal capabilities) should be very good with tabular data.
Gemini is my go-to for handwritten documents. I find it the best of the options. However, I have identified two issues not on your errors list. First, it can totally skip lines in a document, and second it can substitute words. I do not see those in your chart. It also does a marginal job on things written in the margins or inserted between lines. It is still way better than doing it yourself!
Two questions: Are you using it through the AI Studio or the Gemini App (AI studio looks like the image I posed above)? Second , is this the Gemini 3 model which was released last Tuesday? I found Gemini 2.5 which was the precious model did this all the time.
Excellent point. I have used 2.5. Glad to hear 3.0 is much better in these areas.
Interesting. It’s free and I’d be interested to hear what happens if you follow my intuitions above and try out a page that failed previously to see what happens now. Any data is good data!
Wow, this made my day! I am NOT an expert in this field, but rather a retired guy in the middle of a project turning many decades of my parents' weekly family letters into searchable form so my kids and grandkids can learn about our family using NotebookLM. Most are typewritten but many are handwritten. I have just finished successfully processing 600 or so of the typewritten letters but the learning curve to use Transkribus has been too great for me to get started learning it. I just followed your instructions on several handwritten letters and got nearly perfect results, including from the old blue aerogrammes we used when I was abroad in the 70's. This is the solution I've been waiting for! Thank you!
That’s great!
Fascinating development. I see this as welcome news for archives in terms of access and discovery. And Mark, I love your quote from "Computers and Humanities" in the 1960s. Permit me please to offer you another similarly prescient statement from 1869, which goes even further. It's a short article, but if time is an issue the last paragraph, and last line, really hit the mark in my view!
"THE BOOKS OF THE FUTURE
Pall Mall Gazette, 15 Sept 1869
The Daily News is alarmed at the rapid growth of our national library. Every man living has a pen in his hand ; and if only his writing takes tile form of a book entered at Stationers' Hall, it will be preserved for ages in the British Museum. Centuries hence the bookworm will find there, illustrated with woodcuts, verbatim reports of the trials of Palmer and Rush', the Mannings, Madeleine Smith, and Mdme. Rachel.
Centuries hence also he will find those numerous volumes which it is the fashion for tradesmen to issue, and which are but a sublimated form of trade circular. The wine merchant has a volume on his wines, and the hatter on his hats, and the jeweler on his jewels, and the sewing-machine manufacturer on his sewing-machines, and the lock maker on his locks, and the bootmaker on his boots, and the cook on his viands. How are we to stow all these away, and at the same time to keep pace with the literature of foreign countries which is scarcely less productive?
We think of our future librarians as of those children renowned in fairy tales, who have impossible tasks appointed them by malicious godmothers-to collect in a day all the sands of the shore, or to count ere dinner time all the grains of wheat in the kingdom. There will appear no exaggeration in this to anyone who will go to the British Museum and study the catalogue. A man may take a good constitutional walk every day in hunting for half a dozen books in this enormous catalogue, which of itself fills about 1,000 volumes. We find historians of our day like Mr. Carlyle complaining of the immeasurable amount of rubbish which they have to sift in order to get at a few paltry facts.
Must we not pity the historians of the future if they should at any time be conscientious as to turn over the mountains of waste paper which are being shot by cartloads into the Museum? Human eyes and human hands cannot possibly work through a century of such agglomeration. The human mind will despair, perhaps, of power to deal with the illimitable mass. May we hope that when things come to such a crisis, human labour of the literary sort may be in part superseded by machinery ? Machinery has done wonders, and when we think of what literature is becoming it is certainly to be wished that we could read it by machinery, and by machinery digest it."
Pall Mall Gazette - Wednesday 15 September 1869
https://www.britishnewspaperarchive.co.uk/titles/pall-mall-gazette
Great stack and I am glad I found my way to it (via Ben & Sara)
Ain’t no way it can read mine.
Another thing that might be worth flagging about AI Studio is the "Grounding with Google Search" control - it's switched off in your screenshots but defaulted to 'on' when I first opened AI Studio, and it seems like it could make a real difference, so it's worth people being aware of the setting.
I got very quick and accurate results on 19th century English handwriting, and quite good on 15th century English, and I then tried it on a medieval Latin manuscript to see how it would do, and while it (unsurprisingly, as that's giving it a tricky script and also a completely different language) struggled a lot more, turning search on and off there made a very clear difference to how it handled problems like names - with search grounding turned off it made a guess based on the appearance of the text (and did understandable things like turning 'Walterus' into 'Valery'), with search grounding on it either guessed right or named real people who might plausibly have been mentioned in another text on the same subject but weren't involved in this one.
That is interesting! I've been playing around with adding in tool in a custom interface that allow the model to lookup names in a database of known names with info like dates they were active, where they were, that sort of thing, and it does help a lot. I wonder if increasing the reasoning to high on Latin would help at all? Or perhaps playing with the system instructions?
I have tried adding JSON dictionaries for the Gemini API to refer to. Why? You might want to normalise names, ships, whatever, at the time of transcription, rather than through post-processing, or you may want to control the handling of expansions (e.g. English and Latin expansions). Of course you can handle expansions in post-processing by running a very strict diplomatic transcription through a Python script on a fast workstation. A couple of test expansion dictionaries I have developed on the fly are here. I don't absolutely attest for the accuracy of the Latin expansion file:
https://huggingface.co/datasets/MarineLives/English-Expansions
https://huggingface.co/datasets/MarineLives/Latin-Expansions
Regarding "grounding", I think it should be possible in the systems prompt you use in the Gemini 3 API via Google AI Studio to give strict instructions as to how "grounding" is handled, and what interests me is whether relatively small (say under 1000 lines of JSON) files can be referred to by the API. You could of course inject them into the system prompt but that would use a lot of tokens and increase the cost of processing.
I’ve been amazed how Gemini Pro 3 deals with record images from scratched up, streaked, and splotched microfilm. It does a 99% correct job “reading” what is behind all that damage. As genealogists we unfortunately have to deal with this in records, so it’s a huge relief to be able to finally have a tool to help. Now I can compare what it saw to what I saw and feel more confident it’s correct.
Thank you for this in-depth analysis Mark.
It’s impressive to see how quickly LLMs are progressing. I’m also closely following and constantly testing the HTR capabilities of LLMs and other tools/models to see how they evolve.
For relatively clean 18th–19th century handwriting in English though, my experience is that HTR has effectively been “solved” by specialised systems for some time already, and actually for a broader range of documents and languages.
For example, last year I ran an HTR project on a mixed collection of documents in French, Dutch, English, Italian, Spanish and Latin from the 17th to 19th centuries. Using the Transkribus Text Titan I model (released summer 2023), we obtained an average CER of about 5% on a random sample of 100 pages. While not perfect, for such a varied and complex collection that is already a strong performance for an out-of-the-box model.
For relatively clean 18th–19th century Dutch material, similar or even better results can be expected with other Transkribus models.
So in this strict sense, text recognition for these kinds of documents has largely been “solved” for a while. The real remaining challenge in large-scale historical projects is robust layout detection and structured output, and it seems LLMs still have some way to go there.
Looking forward to reading your future articles, the field develops so quickly it is hard to keep up by yourself!
Entirely true. The point, though, I think is that a sub 1% error rate is as good as expert humans. There is a lot of data in that 5%, but yes we’ve had useful HTR for awhile.
It means you can also start to do a whole host of things with these documents via LLMs with having to actually transcribe them…just feed them into the machine. Cost is the other issue. 100 pages via Transkribus is about $100. It is $1 for Gemini. This has huge implications for small archives that want to digitize large collections.
Yes a sub 1% CER rate great indeed!
That 5% rate was just an example of a by now already somewhat outdated model on a complex and varied collection.
In the end 1% CER on a badly recognized/structured layout is probably not as useable as 5% on a well recognized layout.
In the end I think we'll still be better off for a while using different tools (together) depending on the context and goals.
Regarding the pricing discussion, it isn't per se an apples with apples comparison, you get different potential necessary workflows & outputs with both solutions, each having advantages and disadvantages.
However the price of Transkribus is way below $1 per page, it'd be unaffordable otherwise (and I would have spent $250k+ in the last year!). If you're spending €0.10 on pure HTR, it is already much. And for that you do get proper line & region coordinates, TEI-XML ouput & general easier workflows for large batches, etc. included.
Of course, LLMs are for example way better at information extraction, but then costs go up significantly too. Here BERT models or QWEN can be interesting options too.
Bram, Mark, Thiago, I asked Google AI Studio the following question:
"Can Gemini 3 Pro Preview via the Google AI Studio API deliver TEI/XML output for a batch of say 20 images of C17th English language legal depositions after being given a systems instruction to provide a transcription in multiple formats such as JSON, TEI/XML, etc., and a prompt to execute the transcription?"
Here is the full response, which confirms it is very doable. I emphasise I have not yet tested the TEI/XML option, but have no reason to doubt it is possible, since it is trivial computationally. https://share.google/aimode/C6cftouavxTNeoRgW
It would be interesting to test whether Gemini has any native ability to do layout analysis. It wouldn't surprise me if it could, since there are publicly available open source programs it has probably ingested.
Interesting. Transkibus is about $0.25 USD per credit so a 100 page document costs $10.
Well I won't go into the details of Transkribus pricing, that is an art in itself! The plans on the website are all cheaper than buying those single packages though. With larger scale projects prices drop close to the 1 cent too actually.
Woops, have just seen that Mark has already made the comparative cost point. But my point about the potential to integrate NER into the workprocess and equally low cost is still valid.
One important point to consider is comparative cost of Titan and a Gemini 3 API driven workprocess. Mark has the numbers, but my own calculations suggest that I could run 20,000 new images at a significant cost saving on Titan and simultaneously run a NER script so that I then have data which I can use to assemble a basic (or complex) knwowledge graph.
Excellent post, thank you! Appreciated the emphasis on the practical impossibility of completely eliminating typos: “I would count these as spelling errors: in very on, these vs those, where vs were, etc.” ;-)
A fun example of how I've been using Gemini 3 Pro's performance on handwriting:
I have a collection of about 200 handwritten recipes from my grandmother (who has been dead for more than a decade). I've wanted to digitize them for years, but the effort required to sit down and transcribe them myself meant the project never got off the ground. My grandmother had perfect cursive handwriting, having been taught in the 1920s. But, ironically, that also makes these recipes harder to read since so few of us regularly read or write classic cursive anymore.
There are additional complications - the recipes are on decades old notebook paper often with faded ink, stains, or rips. And, because my grandmother was often writing notes to herself with the recipes she had lots of bracketing and margin adjustments. Gemini 3 Pro has been pretty perfect on transcribing and formatting the recipes as I've asked. It's also spit out really interesting observations like commenting about how boxes of cake mix used to be bigger in the 1960s so I'd have to adjust the flour content of one of the recipes if I attempted to make it today. At one point it even recognized that a recipe scan I submitted was actually the reverse side of a document I had given it a few turns earlier, which was wild.
Mark, you make an important point about the different nature of Gemini 3 API errors from both human errors and Transkribus errors.
Have you created a typology of Gemini 3 errors and recorded the stats for the different types? Do the frequencies of error types vary by document type/handwriting, or is the distribution of frequencies a fundamental of the way Gemini 3 does its recognition?
I had a shot sometime ago at analysing the type and frequency of Transkribus errors on English High Court of Admiralty material, intending to use the insights somehow in a finetuning experiment. The data are hand collated by me from comparison of a set of Transkribus transcriptions for the volume of depositions named HCA 13/58. I never did the fine-tuning experiment, but here are the data in JSON format.
https://huggingface.co/datasets/MarineLives/HCA-1358-HTR-Errors-In-Phrases
The key to Gemini 3's superiority over Gemini 2.5 is I agree the radically enhanced vision and multimodal nature of the model. The logic of this is that all general LLM models will show increased HTR capability as they improve vision and multimodality.
To Thiago Krause's point further down the comments about performance in non-English languages, and to another commentator, the logic of Mark's argument about vision capability should be that the HTR capabilities of Gemini 3 are independent of language and more related to the exact nature of the image and the character formation. That being said, I have to think that there must be some work processes running which draw n the prior LLM training, which will be much more extensive in some languages than others, which will affect the quality of the end results by language, so that CER and WER are not language independent.
I haven't done any non-English language testing of Gemini 3. Mark, if you have an Arabic speaking palaographer accessible, it would be interesting to report what Gemini 3's capabilities are with Arabic. Another test (and if it passes it would be outstanding) would be to test on Ottoman Turkish. I would give gold coins for the ability to even get the gist of Ottoman Turkish at low cost and with an easy to use api.
I tested Ottoman Turkish in LLMs and Transkribus a month or so ago for a class visit, with awful results. I tested it again today (as well as Manchu) and sent an email to the professor, asking her to show the results to her students and let me know what they think.
Interesting! Let us know what happens. Did you use Gemini 3 via AI studio as above for the new test?
I can add that I tested Gemini 3 on historical German Gothic handwriting, Russian and Latvian (a small language) handwriting and in all three cases the results were surprisingly good.
Yes - temperature set to zero, thinking low, image quality high. Will keep you posted!
The other tests I would suggest should be high volume and high-scholarly impact non-Latin alphabets. I asked Google AI Mode for some suggestions and came up with: https://share.google/aimode/Pgt2U4hFwKfu1p4l2
ChatGPT 5.1 did well enough on a Chinese early nineteenth-century land grant, according to a student in said class - I read the translation and he said it got about 85% right.
Good idea. Did the Gemini 3.0 transcription give you Arabic script or a Roman character transliteration. You could try feeding a transliteration back into Gemini and ask it to translate it into English and see if it is intelligible. I had a conversation recently with Transkribus recently which suggested they were making progress with Ottoman Turkish, but I haven't seen any data.
Arabic - although of course you could change the prompt. The translation sounded completely coherent, but that is to be expected, so I don't think it proves anything.
Transkribus, on the other hand, was completely gibberish on the same document.
Thank you Mark for you work with LLMs and handwritten text recognition.
We have been doing some experiments at our institution as well. Most of the time Gemini works brilliantly, but occasionally our results are cluttered with almost endless repetitions? Ran it on a 260 page PDF the other day, and a few pages were full of """"""""""""""-charachters and a few others were repetitions of a single word over and over again.
Has anyone else ran into the same problem and solved it somehow?
Just got some amazing results with this, on some Victorian English scrawl I'd really been struggling with - thank you!
Just occasionally it seems to get stuck and go round in circles before reporting an 'internal error' - at which point it doesn't produce an effort at a transcript despite having obviously done nearly all the work. Has anyone found a workaround for this?
I spent three years building a startup whose core product was handwritten text recognition.
We used to work with banks and insurance companies in India and Europe.
We’d have scraped millions of documents online. Also got hundreds of thousands of lines of handwritten text done by data vendors.
We used a modified LSTM architecture.
Circa 2019 our models were better than what GCP had to offer. AWS had come up with a product called “Textract” which while great at document segmentation was still not as good at HTR.
Then covid came and our fledgling startup saw business slow down.
(It’s not easy selling cutting edge tech to legacy banks in India or Europe even otherwise)
We decided to shut shop rather than slog along.
While transformers were there by that time and we were aware of the potential, we were not sure they’d turn out to be as transformative.
(Excuse the pun)
But we did have an inkling that we should rather work on something else.
In hindsight, it was probably the best decision I took. Changed industries completely.
Now while the AI revolution is passing me by, I seem to have done well financially.
Maybe I’ll get back to programming and model development at some point again. But seems unlikely.
If anyone wants access to the dataset then do comment. I’ll try to find it in my old hard disks. My old Mac which had a lot of the data got bricked.
Happy 2026 from Italy.
The country boasts more than 1,000 linear kilometres of handwritten manuscripts stored in archives, libraries, museums, parishes, and notary repositories, spanning from the 10th to the mid-20th century. Until now, a mere 0.001% of this incredibly vast and valuable volume of records has been transcribed and published since the dawn of the printing age. AI may well close that gap to 100%.
Over the last fortnight, I have conducted extensive testing of Gemini 3 Pro and Gemini 3 Flash on the transcription of antique manuscripts ranging from the 16th to the 19th century. I have tested thousands of pages via: A) The Gemini 3 Pro chatbot (Pro subscription); B) Gemini 3 Pro via API on the Google AI Studio Platform; and C) Gemini 3 Pro and Flash via the Google Vertex Platform.
Here are my humble takeaways:
The quality of transcriptions across all three channels is exceptionally high and will, in my view, sound the death knell for professional careers in palaeography. To illustrate: just three months ago, the going rate for a palaeographic transcription of 16th-century cursive in continental Europe was approximately €0.50 per word (!). The cost of a transcription with Gemini is a mere fraction of this sum, and it is calculated per page, not per word. Put simply, palaeography as a profession is dead. However, all intellectual jobs are effectively on death row thanks to AI. They will survive not as professions, but as arts or hobbies for the select few who wish to avoid wandering about as state-subsidised idlers in the coming age of imbecility.
The Google 3 Pro chatbot (which requires a monthly subscription fee) can be fed images of handwritten documents and will obligingly return their transcription in moments. The quality is marginally superior to transcriptions obtained by submitting images directly to the Gemini 3 models via AI Studio and Vertex. This is hardly surprising: one may be the finest prompt engineer in the world, yet replicating the bespoke system prompts running under the bonnet of Google’s flagship chatbot is impossible. Consequently, your direct calls to Google models will always rely on prompts that are slightly less efficient and productive than those crafted internally by Google HQ.
The chatbot is of little use, however, if you require transcriptions for hundreds or thousands of pages. In such instances, one must revert to direct calls to the Gemini models and acquire a modicum of programming knowledge (fear not; AI is highly effective at writing small, end-to-end applications for your needs). You need only be patient and willing to heed your LLM's instructions as it teaches you how to test scripts and report errors. Do as the LLM asks, and through a process of trial and error—where you serve as the dull human in the loop—you will get the application working.
Calling Gemini 3 models via Google AI Studio suffers from strict quota limits. If you intend to process a large volume of images, you must revert to Google’s Vertex platform.
When calling Gemini 3 models for HTR, set "thinking" and "max thinking tokens" to the absolute minimum. The more the model is permitted to "think", the more it tends to run riot, consuming vast quantities of tokens without any improvement in transcription, often failing completely. The precise reason for this eludes me.
Additionally, keep the temperature at 0 and set image resolution to high.
Always pre-process images for the best results. Attempt to reduce noise, calibrate colour, and increase brightness slightly. Crop out any parts of the document that do not bear handwriting.
Can it be let loose on the Epstein files please