Machine learning accelerates digitization of Ancestry records

Machine Learning


Over the past 42 years, Ancestry has collected more than 71 billion birth certificates, marriage licenses, and other family records from 88 countries and built 148 million family trees.

For most of the history of genealogy companies, collecting, labeling, and organizing data has been a painstaking and time-consuming process. Ancestry employees and third-party vendors spent months manually entering data and transcribing international family records. International expansion began in 2001 when the company launched a website in the UK, but adding markets came at a high cost, said Sriram Thiagarajan, Ancestry’s chief technology officer since November 2020.

“The time and cost of digitizing this rich content from around the world was a limiting factor for us,” Thiagarajan told Business Insider.

Most of Ancestry’s AI efforts have been led by Thiagarajan, who joined the company in September 2017 as chief information officer. His expanded role comes a month before investment firm Blackstone completed its $4.7 billion acquisition of Ancestry.

Since then, Ancestry’s investments in machine learning and artificial intelligence, as well as advances in traditional and generative AI, have accelerated the digitization process, Thiagarajan said. AI has also given way to new user tools, such as AI-powered facial and handwriting recognition technology, he added.

A timeline of key events in the use of AI at Ancestry

Training an AI model

Back in 2003, data scientist and software research engineer Jackson Reese was scouted by a friend to join Ancestry as head of digital imaging for preservation services. At the time, Ancestry had a one-person imaging department that digitized census data, birth and death records, immigration documents, and other historical records, Reese told Business Insider. He was originally hired to bring the company’s imaging operations in-house and within three years grew the digital imaging team to more than 70 employees.

Reese told Business Insider that the expanded team worked using now-outdated technology, including microfilm scanners that convert government archives and newspaper clippings into digital files.

Starting in 2014, Ancestry turned to early AI projects focused on developing Ancestry’s own machine learning models and computer vision systems to build algorithms to read paper documents, Reese said. This initial work continued into 2016.

His team then worked with BERT, a family of natural language processing models that Google released in October 2018, to build more accurate data extraction tools. Previously, when the Ancestry team received millions of new birth records, domain experts would review the documents and pass them to indexers, who would transcribe and label them, Reese said.

Ancestry then trains its own AI model based on this data, with the hope that the AI ​​model will be more than 90% accurate after several rounds of back and forth between domain experts, indexers, and data scientists.

“That was the best-case scenario. In some cases, it took eight, 10, or 12 iterations before we could actually tune the model,” Reese said.

By 2019, Ancestry incorporated BERT-based models to more quickly process obituary collection and other record extraction efforts. The company also continued to validate training data and kept employees informed to ensure the model was effectively processing records, Reese said.

ChatGPT tipping point

Thiagarajan said the arrival of ChatGPT in November 2022 is “another turning point in terms of finding the art of the possible.” New large-scale language models from OpenAI, Anthropic, and other AI hyperscalers open up the possibility of accelerating the digitization of unstructured data such as user-generated images, scanned documents, and written stories, he said.

With AI able to perform record extraction faster and more accurately, Ancestry can take birth records and other data and apply a mix of proprietary and open-source AI models from OpenAI, Google, and Anthropic to “tweak them a little bit to fit the use case,” Reese said. He added that the company can handle nearly 200 different languages ​​with little to no iterative model training.

By September 2023, Ancestry will also use LLM for user-facing features, according to Thiagarajan. Face Match is an AI-powered facial recognition tool that debuted in July 2024 to help users identify people in family photos.


An example of Ancestry's AI-powered handwritten note transcription tool.

An example of Ancestry’s AI-powered handwritten note transcription tool.

courtesy of ancestors



In April 2025, the company announced a document transcription feature that allows customers to upload scans of JPG and PNG files and generate transcriptions of family members’ handwritten notes. Launched in December 2025, Ancestry’s AI Stories allows customers to hear a narrated audio story of their life read by an AI when they click on an ancestor’s page in the company’s database.

result

By the end of 2025, more than 50% of Ancestry’s historical records published on its website will have been generated using AI, Thiagarajan said. The company says that thanks to AI, content growth tripled from 800 million records in 2021 to 5.2 billion new records in 2022 and 18.6 billion the following year.

Ancestry also continues to launch external AI use cases, including adding language translation to its customer-facing document transcription tools in June 2026, Thiagarajan said.