← All projects

Case study · Theodore Roosevelt Presidential Library · 2024–present

The Living Library: AI-Driven Access to a Presidential Library’s Collections

Theodore Roosevelt’s papers sit scattered across more than forty repositories — partially catalogued, and reachable in practice only by specialists willing to travel and navigate archival finding aids. This is the work of lowering that bar to a question asked in plain language: a governed corpus of roughly 300,000 records, a researcher-facing interface called Campfire, and an exhibit where visitors speak with Roosevelt himself. The framework behind it — the Living Library — is now published, and I’m a co-author.

RoleChief Communications & Marketing Officer
OrganizationTheodore Roosevelt Presidential Library
Years2024–present
StackFour-layer framework · ~300,000-record governed corpus · hybrid dense/semantic retrieval · in-gallery digital human
PublishedarXiv:2609.09368 · co-author
StatusLive in production
The Reading Room — Discover the life & legacy of Theodore Roosevelt

The brief

Cultural institutions sit on enormous, under-indexed collections: letters, speeches, photographs, ephemera. Finding the right object has historically required knowing the finding aid — an archivist’s skill, not a visitor’s. The job: lower that bar from “know the finding aid” to “ask a question.”

What made it hard was never getting an answer. It was getting an answer a presidential library is willing to stand behind.

The access problem is the whole point, and it is not unique to Roosevelt. Most cultural institutions hold collections that are technically public and practically unreachable: catalogued to different standards, split across repositories, and searchable only by people who already know what they are looking for. A conversational interface is not the achievement. The achievement is a corpus governed well enough that a conversational interface can be pointed at it without embarrassing the institution.

~300,000Records in the governed corpus
40+Repositories aggregated
87.6%Groundedness score
2.80sMean response, in gallery

2024 — The Reading Room

The first version was a public interface to the collection. A student, a teacher, a researcher, or a curious citizen could pick from prompts — I want to do research, find images, build a lesson plan, write a paper, test my knowledge — or type their own question, or start from a topic: the Square Deal, the Rough Riders, antitrust, the Bull Moose party, the books TR wrote.

The AI did the retrieval, summarization and synthesis, but always pointed back to the underlying primary source. No hallucinated quotes, no fabricated dates, no paraphrase masquerading as Roosevelt’s voice.

2026 — Campfire

The rename happened for a reason. “Reading room” describes a place where you are quiet and alone with documents. What the tool actually does is closer to sitting at a fire while somebody who knows the material tells you about it — and answers when you interrupt. The name change followed the product.

Live now at campfire.trlibrary.com, it added:

A naming note: the Library’s opening-weekend program was also called Campfire & Prairie Talks. Same metaphor, deliberately — different thing.

The framework: four layers

What started as one institution’s tool turned out to generalize, and the published framework describes it as four layers — each independently useful, and each one a place another institution could reasonably stop.

  1. Digitization and corpus creation. Material moves into institution-controlled storage with source identifiers and rights metadata preserved. The institution owns the corpus; that is not a detail.
  2. AI-powered processing. OCR and structured metadata enrichment, with original metadata kept immutable and separate from anything a model generated. Curators correct machine output in place through the Archivist App.
  3. Retrieval and reasoning. A hybrid dense/semantic index over the governed corpus. This is the layer that produces Campfire, the researcher-facing interface — and the layer where most of the public value actually lives.
  4. An embodied conversational interface — optional. Voice, avatar, physical presence. This produces Talk to TR in the permanent exhibition.

The fourth layer is the one that gets photographed, and it is the one I would tell a peer institution to skip first. Stop after layer three and you still have the thing that matters: a collection your visitors can actually ask questions of. The avatar is a choice about audience and budget, not a prerequisite for access.

The governance inversion

The part I expect to age best is not a model decision at all. Traditional archival workflow gatekeeps: nothing publishes until a human has reviewed it, which is why enormous quantities of digitized material sit unreachable for years waiting on staff time that will never exist.

The Living Library inverts it. Records publish continuously to the index, each one carrying its review status and its OCR confidence score alongside it. Curators exercise quality control through correction rather than through permission. Access starts immediately and improves continuously, instead of waiting for a completeness that never arrives.

That only works because the provenance labelling is rigorous — a record openly says what a machine produced and what an archivist verified. Take the labelling away and this becomes a way to quietly launder machine output into a scholarly catalogue.

Designed around scholarly trust

Four decisions did most of the work, and all four are about what the system refuses to do.

It cites, and the citations are real

Answers average 8.4 citations, and each one carries the creator, recipient, date, collection, repository and a permalink to the source record — paper records, catalogued and resolvable into the Theodore Roosevelt Center’s digital archive at Dickinson State University. Where a field in the underlying catalogue was machine-generated rather than written by a human archivist, the record says so. An institution that publishes AI-assisted metadata without labelling it is quietly degrading its own catalogue; labelling it costs nothing and preserves the distinction permanently.

It does not fall back on general knowledge

If retrieval returns nothing relevant, the system does not answer from the model’s own training. It says it found nothing and asks for clarification. That is a deliberate constraint, and it is the single most important one: a museum-branded assistant that quietly answers from the open web is no longer a collections tool, it is a chatbot wearing a museum’s logo.

It knows Roosevelt died in 1919

The most common failure mode for a historical-figure AI is the anachronism question — what would TR think about social media, about this election, about climate policy? The tempting answer is a plausible-sounding extrapolation. The correct answer is that Roosevelt died on January 6, 1919, and cannot have had a view.

Adding that single rule moved abstention on temporal-impossibility tests from roughly 60% to 100%. One paragraph of instruction, and the system stopped inventing the opinions of a dead president.

Refusal, though, is a blunt instrument in a gallery. A visitor who asks the Roosevelt avatar about social media and gets a lecture on his date of death has been corrected rather than served. The framework’s answer to this is Cross-Era Analogical Grounding: the system reframes a present-day question through a documented historical parallel and answers from that. Asked about social media, it retrieves Roosevelt on the bully pulpit — going over the heads of the press to speak to citizens directly — and answers there, from attested material.

It is a real distinction and worth being precise about it. The system is not extrapolating what Roosevelt would have thought about a thing that did not exist. It is answering the durable question underneath the modern one, out of what he actually said. The paper is candid that the boundary holds imperfectly — whether grounding reliably prevents synthesis from drifting into anachronism is, in its own words, open.

The modes change the voice, not the retrieval

Discovery, Research, For Teachers and For Students adjust tone and reading level. They do not touch the retrieval, the scope check, or the fact-checking stage — so a student and a scholar asking the same question get the same sources and the same groundedness standard, in different registers. The alternative, where a “kids mode” quietly relaxes the evidentiary bar, is how institutions end up teaching children things they would not print.

The hard part of AI in a cultural-heritage setting isn’t getting an answer. It’s getting an answer the institution is willing to stand behind.

Built with curators, not around them

The collections and curatorial teams were partners from day one, with authority over what the tool will and won’t address. The editorial posture on difficult history came from the Library’s own leadership, and it is not the defensive one: respond like a college professor, challenge the question, include the viewpoints and the norms of the period, don’t defend a side, and present the material so the reader can reach their own conclusion.

That is harder to build than a system that simply refuses to discuss Roosevelt’s record on race and empire. It is also the only version worth having at a presidential library.

One thing the corpus taught us: the collection is far stranger and better than a catalogue suggests. It holds roughly a hundred kinds of object — letters and telegrams, essays, speeches, sheet music, diary entries, even napkins. Among them is Roosevelt’s diary entry for the day his wife and his mother died in the same house, which reads, in full, as a single large X. No summarization improves on that. The system’s job is to put you in front of it.

What the evaluation showed

The system is measured rather than asserted, against a fixed test set with published baselines:

The number that matters most to me is the abstention rate, because it measures the thing a library can actually be embarrassed by. And the evaluation is reproducible by an outside party: Microsoft’s AI for Good Lab, a partner on the project, can re-run the suite against the published baselines. Responsible-AI claims that only the vendor can verify are marketing.

Off the screen: AI TR in the gallery

The step that changed the project’s public profile was moving it off the web. In the permanent exhibition, visitors hold a spoken conversation with Theodore Roosevelt — the same grounded retrieval, the same refusal to speculate past 1919, delivered as a person in a room rather than text in a box.

Doris Kearns Goodwin meets Theodore Roosevelt

Microsoft · July 2026 · 1:05

The historian who has spent a career with Roosevelt’s papers, in conversation with the avatar built on them — filmed during a visit with Microsoft Vice Chair and President Brad Smith. She is the hardest audience this system has: she knows when it is wrong.

It became the most-covered single feature of the opening. Forbes ran “AI-Powered Theodore Roosevelt Is Ready To Answer Your Questions.” When the President spoke with it during the dedication tour, The Hill and The New Republic both covered it — the latter under the headline “People Think Trump Hallucinated Teddy Roosevelt. The Truth Is Weirder.” A science-ethics publication used the exhibit to ask whether talking with the dead is ethical at all.

That last one is the fair question, and the reason the guardrails were built first. An institution that animates a historical figure takes on a duty not to put words in his mouth. Every constraint above — cite or say nothing, never fall back on general knowledge, never speculate past the date of death — exists so that the answer to “is this ethical?” can be something more substantial than “we were careful.”

What running it unattended actually takes

A demo works when someone is watching it. An exhibit has to work at 4:40 on a Saturday with a line of people and no staff nearby, and that requirement drove more engineering than the model work did.

Across one exhibition period, 653 push-to-talk releases produced 457 completed answers. The rest is the texture of a real gallery: 156 were superseded by a visitor interrupting or re-asking, and two smaller groups asked for a repeat or held the button too long. On the completed answers, the delay from releasing the button to Roosevelt’s first spoken word averaged 2.80 seconds — median 2.55, with 97% under five seconds and 69% under three.

Underneath that: presence detection admits a visitor, each conversation is an isolated session, responses stream through dialogue, retrieval, avatar and voice, layered watchdogs recover any failed layer, and the session resets in place for the next person — no restart, no staff intervention. In the first two weeks of public operation in July 2026 it engaged close to 5,000 visitors.

Honest framing, and the paper says so plainly: this is observational evidence from a live exhibit, not a controlled study.

Published, so somebody else can build it

In September 2026 the framework was published with Microsoft’s AI for Good Lab: “The Living Library: Transforming Archival Collections into Conversational Knowledge Systems — Lessons from the Theodore Roosevelt Presidential Library” (arXiv:2609.09368). I’m one of eleven authors, and the Library’s voice in it.

Publishing it was the point. A single institution demonstrating a clever exhibit changes nothing for the field; a documented, transferable framework — with the failure modes named — is something a museum with a fraction of our budget can pick up and use at layer three. The paper is unusually candid about limits: OCR errors propagate into retrieval despite expert review, a corpus centred on one man’s correspondence over-represents his perspective, and the boundary between inference and fabrication is inherently imperfect. Those belong in the record. A framework that only reports its wins is a brochure.

The approach is already travelling. A sibling project, Ukuvula, applies the same pipeline to the Nelson Mandela Foundation’s oral-history archive — a different medium, a different continent, the same four layers.

TRPL’s own write-up of the framework: labs.trlibrary.com/living-library

What I’d carry forward

Why it matters

The bar for “research at a presidential library” should be a question, not a finding aid.

The Library opened with this live on day one, in the browser and in the building. It is a working model for conversational AI at an institution that is serious about accessibility and scholarly integrity at the same time — and a demonstration that the second one is achievable if you are willing to let the system say “I don’t know.”

Try it: campfire.trlibrary.com · Read the paper: arXiv:2609.09368 · Related: the media campaign that made it a national story · AI in the Google Ad Grant